Cell culture methods and compositions

The method uses H3K4me3 modified region analysis to identify cell identity genes, addressing the lack of systematic approaches for cell culture and differentiation, thereby improving cell maintenance and transformation efficiency.

JP7765511B2Active Publication Date: 2025-11-06NATIONAL UNIVERSITY OF SINGAPORE +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024007252
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-06-26
Filing Date
2024-01-22
Publication Date
2025-11-06
Estimated Expiration
2040-06-26

AI Technical Summary

Technical Problem

Current methods lack a systematic approach to identify cell culture conditions and differentiation stimuli for various human cell types, particularly in maintaining cells in vitro and converting them into different cell types, despite advances in single-cell sequencing and computational approaches for transcription factor prediction.

Method used

A method involving determining differential breadth and network scores of H3K4me3 modified regions for protein-coding genes to identify cell identity genes, prioritizing these genes based on their regulatory effect, and using them to determine factors necessary for cell maintenance or transformation, utilizing databases like Epigenome Roadmap, BLUEPRINT, and STRING for data.

Benefits of technology

Enables the identification of specific factors for maintaining and converting cell types by accurately predicting cell identity genes and their regulatory effects, enhancing cell culture conditions and differentiation processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007765511000035
    Figure 0007765511000035
  • Figure 0007765511000036
    Figure 0007765511000036
  • Figure 0007765511000037
    Figure 0007765511000037
Patent Text Reader

Abstract

To provide new methods for identifying factors for maintaining cells in culture and for converting cells into different cell types.SOLUTION: A method for maintaining a cell in vitro comprises the steps of: determining a differential breadth of H3K4me3 modified regions for protein-coding genes in a cell of interest; determining a network score for each of the protein-coding genes in the cell on the basis of the differential breadth of the H3K4me3 modified regions and interactions between protein products of each of the protein-coding genes over at least one network; determining a cell identity score for each of the protein-coding genes on the basis of a combination of the differential breadth and network score; and prioritizing each of the protein-coding genes according to its cell identity score; and comprises thereby identifying the cell identity genes for the cell of interest.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to methods for identifying cell culture factors for cell maintenance and cell transformation. Related Applications This application claims priority from Singapore Patent Application No. 10201905939W, the contents of which are incorporated herein by reference in their entirety. [Background technology]

[0002] Since embryonic stem cells were first isolated, our ability to culture and control cell states has increased. However, one of the most challenging areas in this field is finding the ideal cell culture conditions for maintaining cells in vitro. If the correct conditions are not utilized, cells will transform into a different cellular state or die. Therefore, there is a critical need to mimic in vivo microenvironmental conditions as closely as possible in vitro. To increase the specificity of culture media for cell types, the addition of cell-specific factors, such as components of the extracellular matrix (ECM), growth factors, and other environmental factors, is required.

[0003] Furthermore, a major challenge in the advancement of cell therapy is the requirement for chemically defined cell culture conditions, and attempts have been made to develop serum-free media to maintain various cell types. Similarly, with regard to differentiation, many attempts have been made to determine serum-free, chemically defined protocols for obtaining cells for cell therapy. These protocols use external factors, such as signaling molecules, that mimic natural developmental processes, e.g., differentiation of endothelial cells, microglia, and cardiomyocytes. Furthermore, signaling molecules, such as VEGF, have been shown to remodel the epigenetic landscape at master regulatory gene loci during endothelial differentiation. This suggests that signaling molecules can initiate epigenetic regulation to maintain cell states or differentiate. Summary of the Invention [Problem to be solved by the invention]

[0004] There are over 400 known human cell types, and recent advances in single-cell sequencing technology have led to the identification of many new cell types, such as sensory neurons and blood cells, during embryonic development. Therefore, there is an unmet need to systematically identify cell culture conditions and / or differentiation stimuli for any human cell type. Transcription factor (TF)-mediated transdifferentiation is a well-studied field, and a growing number of data-driven computational approaches have been developed to predict TFs using gene expression data, such as CellNet, D'Alessio AC et al., and Mogrify. However, there are currently no computational approaches for systematically identifying signaling molecules for in vitro cell maintenance or cell transformation across multiple cell types.

[0005] There is a need for new methods for maintaining cells in culture and identifying factors for converting cells into different cell types. The reference to any prior art herein is not an admission or suggestion that this prior art forms part of the common general knowledge in any jurisdiction, or that this prior art could reasonably be expected to be understood by, or considered relevant to, and / or incorporated into other pieces of prior art, by a person skilled in the art. [Means for solving the problem]

[0006] The present invention provides a method for determining cell identity genes for a cell of interest, comprising the steps of: - determining the differential breadth of H3K4me3 modified regions for protein-coding genes in the cell of interest; - determining a network score for each of the protein-coding genes based on the differential widths of the H3K4me3 modified regions and interactions between the protein products of each of the protein-coding genes on at least one network, wherein the network contains information on interactions between the products of the protein-coding genes; - determining a cell identity score for each of the protein-coding genes based on a combination of the differential width and the network score; - prioritizing each of the protein-coding genes according to its cellular identity score; thereby identifying a cell identity gene for a cell of interest.

[0007] The present invention provides a method for determining cell identity genes for a cell of interest, comprising the steps of: - determining the differential width of H3K4me3 modified regions for each protein-coding gene in the cell of interest; - determining a network score for each protein-coding gene in the cell of interest based on differential widths of H3K4me3 modified regions and interactions between protein products of each protein-coding gene on at least one network, wherein the network contains information on interactions between products of the protein-coding genes; - determining a cell identity score for each protein-coding gene in the cell of interest based on a combination of the differential width and the network score; - prioritizing each protein-coding gene according to its cellular identity score; thereby identifying a cell identity gene for a cell of interest.

[0008] The present invention also provides a method for determining cell identity genes for a cell of interest, comprising the steps of: - determining a differential broadness score (DBS) for each protein-coding gene in the cell of interest, the DBS being based on the difference in width of H3K4me3 modified regions for all protein-coding genes compared to the median width of H3K4me3 modified regions for the same protein-coding gene in a population of different cell types; - determining a network score for each protein-coding gene in the cell of interest based on the DBS and interactions between products of each protein-coding gene on at least one network, the network including information of interactions between products of protein-coding genes in the cell; - determining a cellular identity score (RegDBS) for each protein-coding gene in the cell of interest based on a combination of the DBS and network scores; - prioritizing each protein-coding gene according to its RegDBS; thereby identifying a cell identity gene for a cell of interest.

[0009] The present invention also provides a method for determining factors necessary to maintain a cell type in vitro, comprising the steps of: - determining a differential breadth score (DBS) for each protein-coding gene in the cells of interest, the DBS being a differential breadth score for the same protein in a population of different cell types; based on the difference in width of H3K4me3 modified regions for all protein-coding genes compared to the median width of H3K4me3 modified regions for coding genes; - determining a network score for each protein-coding gene in the cell of interest based on interactions between protein products of each protein-coding gene on the DBS and at least one network, the network including information of interactions between products of each protein-coding gene and products of other protein-coding genes in the cell; - determining a cellular identity score (RegDBS) for each protein-coding gene in the cell of interest based on a combination of the DBS and network scores; - prioritizing each protein-coding gene according to its RegDBS, thereby identifying cell identity genes for the cell of interest, wherein each cell identity gene encodes a factor associated with the cell identity of the cell of interest; and thereby identifying factors necessary for maintaining a cell of interest in vitro.

[0010] In any embodiment, the differential width of H3K4me3 modified regions may be determined by obtaining information about the width of H3K4me3 modified regions for each protein-coding gene in a cell of interest to obtain a gene peak width score (B) for each protein-coding gene in the cell, and calculating the difference between the gene width score for each gene in the cell and the median gene width score for the same gene across populations of cells of different types, thereby determining a differential width score (DBS) for each protein-coding gene in the cell of interest.

[0011] The width of the H3K4me3 modified region can be determined based on ChIP-seq information, and more preferably, the information is obtained from either the Epigenome Roadmap, BLUEPRINT, or ENCODE database. In any embodiment, determining the gene peak width score (B) includes first defining the region of the genome where significant H3K4me3 modification exists (histone modification reference peak locus, or RPL). Preferably, the RPL is obtained by merging the obtained ChIP-seq peak regions across all cell types. Preferably, the merged peak regions are overlapping peak regions.

[0012] Preferably, the width of the H3K4me3 modified region is determined based on ChIP-seq information, optionally the information is derived from H3K4me3 profiling methods such as CUT&RUN, scChIC-seq, etc., more preferably the information is obtained from the ENCODE database.

[0013] In a preferred embodiment, determining the gene peak width score (B) for each protein-coding gene may include excluding genes in which H3K4me3 and H3K27me3 modified regions are identified (i.e., poised genes).

[0014] In any embodiment, determining DBS may comprise combining information about the difference in width of the H3K4me3 modified region (compared to control / background cells) for each protein-coding gene with the significance of the difference in width.In certain embodiments, determining DBS may comprise multiplying, normalizing, or adding the difference in width and the significance of the difference, and preferably, DBS is determined by adding the difference in peak width and the significance of the difference.More preferably, DBS is determined by the -log of the difference in H3K4me3 peak width and the significance. 10 The value is determined by adding

[0015] The significance of differences in the width of H3K4me3 modified regions can be determined by any standard statistical method. In one example, the method used is a one-sample Wilcoxon test. In any embodiment, determining a network score for each protein-coding gene in the cell of interest includes combining information about the DBS for each protein-coding gene, the number of out-degree nodes for the gene, and the level of relatedness of the genes in the network. In a preferred embodiment, the network score is determined by summing the DBSs of related genes, weighted for the number of out-degree nodes for the gene and the level of relatedness, to obtain a weighted sum of the DBSs for related protein-coding genes in the network.

[0016] Typically, the protein-protein interaction network is the STRING database, but any sub-network mentioned herein that contains information related to protein interactions within a given cell type may be used.

[0017] It is understood that the Cell Identity Score (RegDBS) is a measure of the regulatory effect of each protein-coding gene on the identity of the cell. In any embodiment, determining RegDBS may include preferentially weighting protein-coding genes that encode regulatory factors.

[0018] In any embodiment, determining the RegDBS for each protein-coding gene may include summing the DBS score and a network score normalized across all genes in the cell of interest. Preferably, summing the DBS and network score further includes weighting the DBS relative to the network score by a factor of 2:1, 1:1, 3:1, 4:1, 5:1, and 1:2. More preferably, the RegDBS is determined by weighting the DBS score relative to the network score by a factor of 2.

[0019] In any embodiment, the step of prioritizing each protein-coding gene according to its RegDBS may include ordering the genes based on their RegDBS values. In a preferred embodiment, the method includes selecting cell identity genes encoding receptor-ligand pairs, thereby identifying signaling molecules necessary for maintaining the cell of interest in vitro. Preferably, the factors necessary for maintaining the cell of interest in vitro are selected from the group consisting of cell surface receptors or ligands involved in cell signaling. Preferably, determining the factors necessary for maintaining the cell of interest in vitro includes ranking each protein-coding gene encoding a cell surface receptor according to its RegDBS score. The method may further include prioritizing the ligands associated with the receptor based on the DBS score for each ligand to obtain a combined ranking of receptors and ligands, which identifies ligands for use in supplementing a culture medium for maintaining the cell of interest in vitro.

[0020] In further embodiments, the factors for maintaining the cell of interest in vitro may include transcription factors or epigenetic remodeling factors. Thus, the method may further include selecting a cell identity gene encoding a transcription factor, thereby identifying the transcription factor required to maintain the cell of interest in vitro.

[0021] Accordingly, the present invention also provides a method for determining cell identity genes for a cell of interest, comprising the steps of: - determining an H3K4me3 gene peak width score (B) for each protein-coding gene (g) in the cell of interest (x), wherein the gene peak width score is , which is the sum of the lengths of the regions in the promoters of each protein-coding gene that contain H3K4me3 modifications; and - The difference in gene peak width score (ΔPeak Width) between the cell of interest and the median gene peak width score of a population of cells representing the background gene peak width score x g) and determining the significance of that difference (Pval); - Delta peak width and Pval value are added to give a differential peak width score (DBS) for each protein-coding gene in the cell. x g ) and - Differential Broad Score (DBS) of related genes (r) in the network x g ) to obtain a network score (N x g ), wherein the network includes information on interactions between the products of each protein-coding gene and the products of other protein-coding genes in the cell, and the differential breadth score is corrected for the number of out-degree nodes of the gene (O) and the level of relatedness of the gene (L), and preferably is calculated as a differential breadth score (DBS x g ) is calculated up to the third level of relevance; - Network score Net across all protein-coding genes (g) in the cell of interest (x) x g normalizing - Differential Broad Score (DBS) x g ) and network score (Net x g ), thereby generating a regulatory differential breadth score (RegDBS) for each protein-coding gene in the cell of interest. x g ) determining a RegDBS x g is an indicator of the regulatory effect of each protein-coding gene on the identity of the cell; - That RegDBS x g and prioritizing each gene according to thereby identifying a cell identity gene for a cell of interest.

[0022] Preferably, interactions between gene nodes are selected if their experimental score in the network is greater than zero and their combined score is greater than 500. In a further embodiment, determining the network score comprises removing protein-coding genes from the protein-protein interaction network that do not have associated peak H3K4me3 width.

[0023] Preferably, scoring each protein-coding gene in the cell of interest is performed using a differential broad score (DBS). x g ) and the normalized network score for the gene (Net x g ), and more preferably, the differential broadness score (DBS x g ) is the normalized network score (Net x g ) by a factor of at least 2.

[0024] The present invention also provides a method for determining factors necessary for conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type, comprising the steps of: - determining differential widths of H3K4me3 modified regions for protein-coding genes in the source cell and the target cell type to obtain a score of H3K4me3 modification (DBS) for protein-coding genes in the source and target cell types; - calculating the difference between the DBS of the source cell and the DBS of the target cell to obtain a measure of the difference in H3K4me3 modification between the source cell and the target cell for each of the protein-coding genes (cell conversion DBS); - determining a network score for transforming a source cell type into a target cell type based on the cell transformation DBS and interactions between protein products of each of the protein-coding genes on at least one network, the network containing information on interactions between products of protein-coding genes in the cell; - Cell transformation ΔDBS and network score combination based on target cell tandem determining a cell identity conversion score for each of the protein-encoding genes; - prioritizing each of the protein-coding genes according to its cell identity conversion score to identify cell identity genes for the target cell, each cell identity gene encoding a factor associated with the cell identity of the target cell; and thereby determining the factors necessary for conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type.

[0025] The present invention also provides a method for determining factors necessary for conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type, comprising the steps of: - determining the differential width of H3K4me3 modified regions for each protein-coding gene in the source cell and the target cell type to obtain a score of H3K4me3 modification for each protein-coding gene (DBS) in the source and target cell types; - calculating the difference between the DBS of the source cell and the DBS of the target cell to obtain a measure of the difference in H3K4me3 modification between the source cell and the target cell for each protein-coding gene (cell conversion ΔDBS); - determining a network score for transforming the source cell type into the target cell type based on the cell transformation ΔDBS and interactions between protein products of each protein-coding gene on at least one network, the network including information on interactions between each protein-coding gene product and other gene products in the cell; - determining a cell identity transformation score for each protein-coding gene in the target cell based on a combination of the cell transformation ΔDBS and the network score; - prioritizing each protein-coding gene according to its cell identity conversion score to identify cell identity genes for the target cell, each cell identity gene encoding a factor associated with the cell identity of the target cell; thereby determining the factors necessary for conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type.

[0026] The present invention also provides a method for determining factors necessary for conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type, comprising the steps of: - determining a differential breadth score (DBS) for each protein-coding gene in the source cells and the target cells, the DBS being based on the difference in width of H3K4me3 modified regions for all protein-coding genes compared to the median width of H3K4me3 modified regions for the same protein-coding gene in a population of different cell types; - calculating the difference between the DBS of the source cell and the DBS of the target cell to obtain a cell conversion ΔDBS for each protein-coding gene; - determining a network score for the cell transformation from the source cell type to the target cell type based on the cell transformation ΔDBS and interactions between protein products of each gene on at least one network, the network including information on interactions between the products of each protein-coding gene and other protein-coding gene products in the cell; - determining a cell identity conversion score (RegΔDBS) for each protein-coding gene in the target cell based on a combination of the cell conversion ΔDBS and the network score; - prioritizing each protein-coding gene according to its cellular transformation RegΔDBS to identify cellular identity genes for the target cell, each cellular identity gene encoding a factor associated with the cellular identity of the target cell; thereby determining the factors necessary for conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type.

[0027] In any embodiment, the differential width of the H3K4me3 modified region is determined by obtaining information about the width of the H3K4me3 modified region for each protein-coding gene in the source cells and the target cells to obtain a gene peak width score (B) for each gene in each cell type, and calculating the difference between the gene peak width score for each protein-coding gene in the cell type and the median gene peak width score for the same gene across populations of cells of different types, thereby determining a differential width score (DBS) for each gene in both the source and target cells.

[0028] The width of the H3K4me3 modified region can be determined based on ChIP-seq information, and more preferably, the information is obtained from either the Epigenome Roadmap, BLUEPRINT, or ENCODE database. In any embodiment, determining the gene peak width score (B) includes first defining a region of the genome where significant H3K4me3 modification exists (histone modification reference peak locus, or RPL). Preferably, the RPL is obtained by merging the obtained ChIP-seq peak regions across all cell types. Preferably, the merged peak regions are overlapping peak regions.

[0029] Preferably, the width of the H3K4me3 modified region is determined based on ChIP-seq information, more preferably, the information is obtained from the ENCODE database. In a preferred embodiment, determining the gene peak width score (B) for each protein-coding gene includes excluding genes in which H3K4me3 and H3K27me3 modified regions are identified (i.e., poised genes).

[0030] In any embodiment, determining the DBS of each protein-coding gene in source cell and target cell can comprise combining information about the difference in H3K4me3 width (compared to background H3K4me3 width) of each protein-coding gene with the significance of this difference.In certain embodiments, determining DBS can comprise multiplying, normalizing, or adding the difference in peak width and the significance of this difference, and preferably, DBS is determined by adding the difference in peak width and the significance of this difference.More preferably, DBS is determined by adding the difference in peak width and the -log10 value of significance.

[0031] The significance of differences in the width of H3K4me3 modified regions can be determined by any standard statistical method. In one example, the method used is a one-sample Wilcoxon test. In a preferred embodiment, calculating the difference between the DBS of the source cell and the DBS of the target cell to obtain a cellular conversion ΔDBS for each gene comprises subtracting the DBS of the source cell for each protein-coding gene from the DBS of the target cell for each gene to obtain a cellular conversion ΔDBS for each gene. The cellular conversion ΔDBS for each protein-coding gene provides a measure of the difference in width of the H3K4me3 modified region for a given protein-coding gene between the source cell and the target cell.

[0032] Preferably, the network score for the cell transformation from the source cell to the target cell is a weighted combination of the cell transformation ΔDBS of the related genes in the network. For example, determining the network score for the cell transformation may include combining the cell transformation ΔDBS for each gene, the number of out-degree nodes of the gene, and information about the level of relatedness of the gene in the network. In a preferred embodiment, the network score is determined by summing the cell transformation ΔDBS of the related genes, weighted by the number of out-degree nodes of the gene and the level of relatedness to obtain a weighted sum of the cell transformation ΔDBS for the related genes in the network.

[0033] Typically, the protein-protein interaction network is the STRING database, but any sub-network mentioned herein that contains information about protein interactions within a given cell type may be used.

[0034] The Cell Conversion Identity Score (CellConversionRegΔDBS) is understood to be an indicator of the regulatory effect of each protein-coding gene on changes in cell identity. In any embodiment, determining the cell transformation RegΔDBS may include preferentially weighting protein-coding genes that encode regulatory factors.

[0035] In any embodiment, determining the cellular transformation RegΔDBS for each protein-coding gene comprises adding the cellular transformation ΔDBS score and a normalized cellular transformation network score across all protein-coding genes. Preferably, adding the cellular transformation ΔDBS and the cellular transformation network score further comprises weighting the cellular transformation ΔDBS relative to the cellular transformation network score by a factor of 2:1, 1:1, 3:1, 4:1, 5:1, and 1:2. More preferably, the cellular transformation RegΔDBS is determined by weighting the cellular transformation ΔDBS score relative to the network score by a factor of 2.

[0036] In any embodiment, the step of prioritizing each protein-coding gene according to its Cell Conversion RegΔDBS comprises ordering the genes based on Cell Conversion RegΔDBS values.

[0037] In a preferred embodiment, the cell transformation factor comprises a transcription factor or an epigenetic remodeling factor. Thus, the method may further comprise selecting a gene encoding a transcription factor, thereby identifying the transcription factor required for the transformation of the source cell into a cell exhibiting at least one characteristic of the target cell type. Preferably, the transformation is transdifferentiation of the differentiated source cell into the differentiated target cell.

[0038] In an alternative embodiment, the method includes selecting protein-coding genes encoding receptor-ligand pairs, thereby identifying signaling molecules required for the conversion of source cells into target cells. Preferably, the factors required for the conversion of source cells into target cells are selected from the group consisting of cell surface receptors or ligands involved in cell signaling. Preferably, determining the factors required for the conversion of source cells into target cells includes ranking each gene encoding a cell surface receptor according to its Cell Conversion RegΔDBS Score. The method may further include prioritizing the ligands associated with the receptor based on the Cell Conversion ΔDBS Score for each ligand to obtain a combined ranking of receptors and ligands, which identifies ligands for use in supplementing the culture medium for the conversion of source cells into target cells. Preferably, the conversion is directed differentiation of pluripotent source cells into differentiated target cells.

[0039] Thus, the present invention also provides a method for determining factors necessary for conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type, comprising the steps of: - determining an H3K4me3 gene peak width score (B) for each protein-coding gene (g) in the cell of interest (S) and in the target cell (T), wherein the gene peak width score is the sum of the lengths of the regions in the promoters of each protein-coding gene that contain the H3K4me3 modification; - the normalized difference in gene peak width score between the source cell and the median of the gene peak width scores of a population of cells exhibiting a background gene peak width score (Δpeak width S g ), and the significance of the difference (Pval S g ) and - the normalized difference in gene width score (ΔPeak Width) between the target cell and the median gene width score of a population of cells representing the background gene width score T g ), and the significance of the difference (Pval T g ) and - ΔPeak Width S g and Pval S g The values ​​of and are added to obtain the differential broadening score (DBS) for each protein-coding gene in the source cell. S g ) and - ΔPeak Width T g and Pval T g The values ​​of and are added to obtain the differential broadening score (DBS) for each protein-coding gene in the target cell. T g ) and - Differential Broad Score (DBS) for the same gene in the target cells T g ) to obtain the differential broadening score (DBS) for each protein-coding gene in the source cell. S g ) to obtain the cellular conversion ΔDBS (ΔDBS) for each protein-coding gene in the target cells. T-Sg ) and - Cell-transformed differential broadening score (ΔDBS) of related genes (r) in the network T-S g ) to obtain the network score (N T-S g ), wherein the network includes information of interactions between the products of each protein-coding gene in the cell, and the differential breadth score is corrected for the number of gene out-degree nodes (O) and the level of gene relatedness (L), and preferably is determined as a cell transformed differential breadth score (ΔDBS T-S g ) is calculated up to a third level of relevance; - Cellular transformation network score Net across all protein-coding genes T-S g normalizing - Cellular transformation differential broadening score (ΔDBS T-S g ) and cell transformation network score (Net T-S g ) to obtain a cell transformation regulatory differential broad score (RegΔDBS T-S g ) determining RegΔDBS T-S g is a measure of the difference in regulatory effect of each protein-coding gene on the target cell compared to the source cell; - That RegΔDBS T-S g prioritizing each protein-coding gene according to thereby identifying factors necessary for conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type.

[0040] Preferably, interactions between gene nodes are selected if their experimental score in the network is greater than zero and their combined score is greater than 500. In a further embodiment, determining the network score comprises removing genes from the protein-protein interaction network that do not have an associated peak H3K4me3 width.

[0041] Preferably, scoring each protein-coding gene in the cell of interest is performed using a differential broad score (ΔDBS T-S g ) and the network score for the gene (Net T-S g ) and more preferably, a differential broadness score (ΔDBS T-S g ) is the network score (Net T-S g ) by a factor of at least 2.

[0042] In any embodiment of the above aspects, the method includes selecting a subset of protein-coding genes that encode transcription factors, thereby identifying transcription factors required for conversion of the source cells into cells exhibiting at least one characteristic of the target cell type. Preferably, the identified factors are for transdifferentiation of the differentiated source cells into the differentiated target cells.

[0043] In a further embodiment of the above aspect, the method comprises: The method includes selecting a protein-encoding gene, thereby identifying a signaling molecule required for conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type. Preferably, the identified factor is for directed differentiation of a pluripotent source cell into a differentiated target cell.

[0044] In a further embodiment, the method further comprises determining a combined rank of the gene encoding the receptor and the gene encoding the receptor ligand in a receptor-ligand pair, the ranking being based on RegDBS and DBS for the receptor and ligand, respectively, thereby identifying receptor-ligand pairs required for cell maintenance.

[0045] In further embodiments of the above aspects, the method further comprises removing transcriptionally redundant TFs from the ranked list from each cell type. Still further, in any of the above methods, the method may subsequently include contacting the population of cells of interest, or the source cells, with one or more of the identified factors so as to maintain the cells of interest in vitro or convert the source cells into cells exhibiting at least one characteristic of the target cells. Alternatively, the method may include transfecting cells of the source cells of interest with nucleic acids encoding one or more factors so as to maintain the cells of interest in vitro or convert the source cells into cells exhibiting at least one characteristic of the target cells.

[0046] The present invention also provides a method for maintaining a population of cells in vitro, comprising: - providing a population of cells of interest in cell culture; - determining a differential breadth score (DBS) for each protein-coding gene in the cell of interest, the DBS being based on the difference in width of H3K4me3 modified regions for all protein-coding genes compared to the median width of H3K4me3 modified regions for the same gene in a population of different cell types; - determining a network score for each protein-coding gene in the cell of interest based on interactions between the protein products of each gene on the DBS and at least one network, the network including information on interactions between the products of each protein-coding gene and products of other protein-coding genes in the cell; - scoring each protein-coding gene in the cell of interest based on a combination of the DBS and the network score, thereby determining a RegDBS for each protein-coding gene in the cell of interest, wherein the RegDBS is an indication of the importance of each protein-coding gene with respect to cell identity; - prioritizing each protein-coding gene according to its RegDBS, thereby identifying cell identity genes for the cell of interest, each cell identity gene encoding a factor associated with the cell identity of the cell of interest; - contacting the population of cells of interest with one or more of the factors associated with the cellular identity of the cells of interest; - culturing the population of cells for a time and under conditions sufficient to allow maintenance of the cells of interest in cell culture; thereby maintaining a population of cells in vitro.

[0047] In certain embodiments, contacting the cell with the factor may include transfecting the cell with one or more nucleic acid molecules encoding the factor and expressing the factor in the cell. Alternatively, contacting may include contacting the cell with an agent that increases expression of the one or more factors.

[0048] The present invention also provides a method for maintaining a population of astrocytes in vitro, comprising, in order: - providing a population of astrocytes in cell culture; - contacting the population of astrocytes with a set of factors selected from FN1, COL4A1, LAMB1, ADAM12, WNT5A, COL1A2 and EDIL3, or variants thereof, to maintain at least one characteristic of astrocytes; - culturing the population of astrocytes for a time and under conditions sufficient to allow maintenance of the astrocytes in cell culture; thereby maintaining a population of astrocytes in vitro.

[0049] Preferably, the factors comprise, consist of, or consist essentially of FN1, COL4A1, LAMB1, ADAM12, WNT5A, COL1A2, and EDIL3.

[0050] In certain embodiments, the method includes contacting astrocytes with one or more, two or more, three or more, four or more, five or more, or six of FN1, COL4A1, LAMB1, ADAM12, WNT5A, COL1A2, and EDIL3.

[0051] In a preferred embodiment, the method comprises contacting astrocytes with FN1, COL4A1, LAMB1, ADAM12, WNT5A, COL1A2 and EDIL3.

[0052] In one embodiment, the method comprises contacting the astrocytes with at least FN1, LAMB1, or COL1A2 (optionally contacting the astrocytes with FN1 and LAMB1, or FN1 and COL1A2, or all three of FN1, LAMB1, and COL1A2). Optionally, one or more of COL4A1, ADAM12, and EDIL3 are also used to contact the astrocytes.

[0053] In any embodiment, a functional variant of any one of the factors, for example, FN1, COL4A1, LAMB1, ADAM12, WNT5A, COL1A2 and EDIL3, may be used.

[0054] In any of the methods described herein, the method may further comprise administering to an individual the astrocytes, or population of cells, produced according to the method. The present invention also provides a method for maintaining a population of cardiomyocytes in vitro, comprising: - providing a population of cardiomyocytes in cell culture; - contacting the population of cardiomyocytes with one or more factors selected from FN1, COL3A1, TFPI, FGF7, APOE, C3, COL1A2, SERPINE1, COL6A3, CXCL12 to maintain at least one characteristic of the cardiomyocytes; - culturing the population of cardiomyocytes for a time and under suitable conditions sufficient to maintain the cardiomyocytes in cell culture; thereby maintaining a population of cardiomyocytes in vitro.

[0055] Preferably, the factor comprises, consists of, or consists essentially of FN1, COL3A1, TFPI, FGF7, APOE, C3, COL1A2, SERPINE1, COL6A3, CXCL12.

[0056] In certain embodiments, the method includes contacting cardiomyocytes with one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, or nine of FN1, COL3A1, TFPI, FGF7, APOE, C3, COL1A2, SERPINE1, COL6A3, CXCL12.

[0057] In a preferred embodiment, the method comprises contacting cardiomyocytes with at least FN1, COL3A1 (collagen III), TFP1, FGF7 and APOE, or functional variants thereof.

[0058] In a further preferred embodiment, the method comprises contacting cardiomyocytes with FN1, COL3A1, TFPI, FGF7, APOE, C3, COL1A2, SERPINE1, COL6A3 and CXCL12.

[0059] In any embodiment, a functional variant of any one of the factors, for example, FN1, COL3A1, TFPI, FGF7, APOE, C3, COL1A2, SERPINE1, COL6A3 and CXCL12, may be used.

[0060] In any of the methods described herein, the method may further comprise administering to an individual the cardiomyocytes, or population of cells, produced according to the method. The present invention also provides a method for maintaining a population of smooth muscle cells in vitro, comprising: - providing a population of smooth muscle cells in cell culture; - contacting the population of smooth muscle cells with one or more factors selected from LAMA5, COL4A1, LAMA4, NID1, COL6A3, COL4A6, COL4A5, FGF10, FGF7, GNAS, COL7A1, COL1A1, and THBS1 to maintain at least one characteristic of the smooth muscle cells; - culturing the population of smooth muscle cells for a time and under suitable conditions sufficient to maintain the smooth muscle cells in cell culture; thereby maintaining a population of smooth muscle cells in vitro.

[0061] Preferably, the factors comprise, consist of, or consist essentially of LAMA5, COL4A1, LAMA4, NID1, COL6A3, COL4A6, COL4A5, FGF10, FGF7, GNAS, COL7A1, COL1A1, and THBS1.

[0062] In certain embodiments, the method includes contacting smooth muscle cells with one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, or nine or more of LAMA5, COL4A1, LAMA4, NID1, COL6A3, COL4A6, COL4A5, FGF10, FGF7, GNAS, COL7A1, COL1A1, and THBS1.

[0063] In a preferred embodiment, the method comprises contacting smooth muscle cells with collagen 4, NID1, collagen 6, FGF10, FGF7, collagen 1 and THBS1.

[0064] In any embodiment, a functional variant of any one of the factors, for example, LAMA5, COL4A1, LAMA4, NID1, COL6A3, COL4A6, COL4A5, FGF10, FGF7, GNAS, COL7A1, COL1A1 and THBS1, may be used.

[0065] In any of the methods described herein, the method may further comprise administering to an individual the smooth muscle cells, or population of cells, produced according to the method. The present invention also provides a method for maintaining a population of endothelial cells, preferably aortic endothelial cells, in vitro, comprising: - providing a population of endothelial cells in cell culture; - contacting the population of endothelial cells with one or more factors selected from BMP6, ADAM9, LAMB1, LAMA4, THBS1, CTGF, BMP4, PDGFB, FN1, and CYR61 to maintain at least one characteristic of the endothelial cells; - culturing the population of endothelial cells for a time and under suitable conditions sufficient to maintain the endothelial cells in cell culture; thereby maintaining a population of endothelial cells in vitro.

[0066] Preferably, the factors comprise, consist of, or consist essentially of BMP6, ADAM9, LAMB1, LAMA4, THBS1, CTGF, BMP4, PDGFB, FN1, and CYR61.

[0067] In certain embodiments, the method includes contacting endothelial cells with one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, or nine or more of BMP6, ADAM9, LAMB1, LAMA4, THBS1, CTGF, BMP4, PDGFB, FN1, and CYR61.

[0068] In a preferred embodiment, the method comprises contacting endothelial cells with BMP6, THBS1, CTGF, BMP4, PDGFB, FN1 and CYR61. In any embodiment, a functional variant of any one of the factors, for example, MP6, ADAM9, LAMB1, LAMA4, THBS1, CTGF, BMP4, PDGFB, FN1 and CYR61, may be used.

[0069] In any of the methods described herein, the method may further comprise administering to an individual the endothelial cells, or population of cells, produced according to the method. In any of the above methods for maintaining a cell of interest, the method can include methods for maintaining transdifferentiating cells, undifferentiated cells, differentiated cells and transdifferentiated cells.

[0070] The present invention also provides a method for converting a source cell into a cell exhibiting at least one characteristic of a target cell type, comprising the steps of: - identifying factors necessary for the conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type by any method described herein; - providing source cells; - contacting the source cells with one or more of the factors identified for the conversion of the source cells into target cells; - culturing the source cells for a time and under conditions sufficient to allow conversion of the source cells into cells exhibiting at least one characteristic of the target cell type; thereby converting a source cell into a cell that exhibits at least one characteristic of a target cell type.

[0071] The present invention also provides a method for converting a source cell into a cell exhibiting at least one characteristic of a target cell type, comprising the steps of: - identifying factors necessary for the conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type by any method described herein; - providing source cells; - contacting the source cells with or increasing the amount of one or more identified factors for conversion of the source cells into target cells; - culturing the source cells for a time and under conditions sufficient to allow conversion of the source cells into cells exhibiting at least one characteristic of the target cell type; thereby converting a source cell into a cell that exhibits at least one characteristic of a target cell type.

[0072] In a preferred embodiment, the source cells are H9 embryonic stem cells and the target cell type is any of the cell types listed in Table 3. "Increasing" the amount can include expressing a nucleic acid encoding one or more factors in the source cells or contacting the source cells with an agent to increase expression of the factors by the cells.

[0073] The present invention provides a method for differentiating a source cell, comprising contacting the source cell with one or more factors or variants thereof or increasing protein expression of one or more factors or variants thereof in the source cell, wherein the source cell differentiates to exhibit at least one characteristic of a target cell; - the source cell is a pluripotent stem cell or progenitor cell and the target cell is any of the cells listed in Table 3; the factors are selected from those set forth in Table 3 for a given target cell type. Optionally, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, or nine or more, or ten or more, or eleven or more, or twelve or more, or thirteen or more, or fourteen or more of the factors set forth in Table 3 are used in the method.

[0074] The present invention provides a method for differentiating a source cell, comprising increasing protein expression of one or more factors or variants thereof in the source cell, wherein the source cell differentiates to exhibit at least one characteristic of a target cell; - the source cells are pluripotent stem or progenitor cells and the target cells are astrocytes; - the factor is selected from FN1, COL4A1, LAMB1, ADAM12, WNT5A, COL1A2 and EDIL3.

[0075] The present invention provides a method for generating cells exhibiting at least one characteristic of astrocytes from pluripotent stem or progenitor cells, comprising: - contacting the pluripotent stem or progenitor cells in the source cells with one or more factors selected from FN1, COL4A1, LAMB1, ADAM12, WNT5A, COL1A2 and EDIL3, or variants thereof; - culturing the pluripotent stem or progenitor cells for a time and under conditions sufficient to allow differentiation into astrocytes, thereby generating cells from the pluripotent stem or progenitor cells that exhibit at least one characteristic of astrocytes; The present invention provides a method comprising:

[0076] The present invention also provides a method for differentiating pluripotent stem cells, preferably embryonic stem cells, or progenitor cells, comprising increasing protein expression of one or more of FN1, COL4A1, LAMB1, ADAM12, WNT5A, COL1A2 and EDIL3, or variants thereof, in the stem cells, wherein the stem cells differentiate to exhibit at least one characteristic of an astrocyte.

[0077] The present invention provides a method for differentiating pluripotent stem cells, preferably embryonic stem cells, or progenitor cells, into cells exhibiting at least one characteristic of astrocytes, the method comprising the steps of: i) providing pluripotent stem cells or progenitor cells, or a cell population comprising pluripotent stem cells or progenitor cells; ii) transfecting said pluripotent stem cells with one or more nucleic acids comprising nucleotide sequences encoding one or more factors important for astrocyte cellular identity; and iii) culturing said cells or cell population and optionally monitoring the cells or cell population for at least one characteristic of astrocytes, wherein preferably the factors for differentiating pluripotent stem cells or progenitor cells into astrocytes, or for generating cells exhibiting at least one characteristic of astrocytes, comprise, consist of, or consist essentially of FN1, COL4A1, LAMB1, ADAM12, WNT5A, COL1A2, and EDIL3.

[0078] In certain embodiments, the progenitor cells are preferably neural progenitor cells or cells obtained from a population of neural progenitor cells. Preferably, the pluripotent stem cells are embryonic stem cells.

[0079] The present invention provides a method for generating cells exhibiting at least one characteristic of astrocytes from embryonic stem cells, comprising: - increasing the amount of any one or more of FN1, COL4A1, LAMB1, ADAM12, WNT5A, COL1A2 and EDIL3, or variants thereof, in embryonic stem cells; - culturing the embryonic stem cells for a time and under conditions sufficient to differentiate them into astrocytes, thereby producing cells from the embryonic stem cells that exhibit at least one characteristic of astrocytes; The present invention provides a method comprising:

[0080] In certain embodiments, increasing the amount of one or more of the above-listed factors comprises expressing a nucleic acid encoding one or more of the factors in the source cells (i.e., stem cells). Alternatively, the factors may be provided directly to the cells via the cell culture medium.

[0081] Typically, suitable conditions for targeting cell differentiation include culturing cells for a sufficient time and in a suitable medium. The sufficient time of culturing can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 days. The suitable medium can be one shown in Table 4.

[0082] Preferably, at least one characteristic of astrocyte is the upregulation of any one or more astrocyte markers and / or the change of cell morphology.Relevant markers are described herein and known to those skilled in the art.The exemplary markers for astrocyte include GFAP, S100B, ALDH1L1, CD44 and GLAST1, or any of the markers listed in Table 1.

[0083] Preferably, at least one characteristic of cardiomyocyte is the upregulation of any one or more cardiomyocyte markers and / or the change of cell morphology.Relevant markers are described herein and known to those skilled in the art.The exemplary markers for cardiomyocyte include NKX2.5, GATA4, GATA6, MEF2C, MYH6, ACTN1, CDH2 and GJA1, or any of the markers listed in Table 1.

[0084] Preferably, at least one characteristic of smooth muscle cell is the upregulation of any one or more smooth muscle cell markers and / or changes in cell morphology.Relevant markers are described herein and known to those skilled in the art.Exemplary markers for smooth muscle cell include SM22 and α-SMA, or any of the markers listed in Table 1.

[0085] Preferably, at least one characteristic of endothelial cell is the upregulation of any one or more endothelial cell markers and / or changes in cell morphology.Relevant markers are described herein and known to those skilled in the art.Exemplary markers for endothelial cell include CD31 and VWF, or any of the markers listed in Table 1.

[0086] The present invention also provides cells exhibiting at least one characteristic of an astrocyte, a cardiac muscle cell, a smooth muscle cell, or an endothelial cell produced by the methods described herein. In any of the methods described herein, the method may further include expanding cells that exhibit at least one characteristic of astrocytes to increase the proportion of cells in the population that exhibit at least one characteristic of astrocytes. Expanding the cells may be in culture for a time and under conditions sufficient to generate a population of cells, as described below.

[0087] In any of the methods described herein, the method may further include expanding the cells that exhibit at least one characteristic of cardiomyocytes to increase the proportion of cells in the population that exhibit at least one characteristic of cardiomyocytes. Expanding the cells may be in culture for a time and under conditions sufficient to generate a population of cells, as described below.

[0088] In any of the methods described herein, the method may further include administering to the individual a cell, or a cell population comprising a cell, that exhibits at least one characteristic of an astrocyte, a cardiomyocyte, a smooth muscle cell, or an endothelial cell.

[0089] The present invention also provides compositions comprising a set of factors for (i) maintaining a population of astrocytes in vitro or (ii) generating cells exhibiting at least one characteristic of astrocytes from pluripotent stem cells. Preferably, the composition comprises a set of factors selected from FN1, COL4A1, LAMB1, ADAM12, WNT5A, COL1A2, and EDIL3, or variants thereof. Preferably, the set of factors comprises, consists of, or consists essentially of FN1, COL4A1, LAMB1, ADAM12, WNT5A, COL1A2, and EDIL3, or variants thereof. In certain embodiments, the composition may comprise one or more, two or more, three or more, four or more, five or more, or six of FN1, COL4A1, LAMB1, ADAM12, WNT5A, COL1A2, and EDIL3, or variants thereof. In one embodiment, the composition comprises at least FN1, LAMB1, or COL1A2 (optionally, FN1 and LAMB1, or FN1 and COL1A2, or all three). Optionally, the composition further comprises one or more of COL4A1, ADAM12, and EDIL3.

[0090] The present invention also provides a composition comprising a set of factors for maintaining a population of cardiomyocytes in vitro. Preferably, the composition comprises a set of factors selected from FN1, COL3A1, TFPI, FGF7, APOE, C3, COL1A2, SERPINE1, COL6A3, CXCL12, or variants thereof. Preferably, the set of factors comprises, consists of, or consists essentially of FN1, COL3A1, TFPI, FGF7, APOE, C3, COL1A2, SERPINE1, COL6A3, CXCL12, or variants thereof. In certain embodiments, the composition comprises FN1, COL3A1, TFPI, FGF7, APOE, C3, COL1A2, SERPINE1, COL6A3, CXCL12, or variants thereof. The factors may include one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, or nine of INE1, COL6A3, CXCL12, or variants thereof. Most preferably, the factors include, consist of, or consist essentially of FN1, COL3A1 (collagen III), TFP1, FGF7, and APOE.

[0091] In any embodiment, the composition may further comprise one or more components for supporting the growth or maintenance of cells in culture in vitro. In any aspect, the composition may be a cell culture medium.

[0092] In any embodiment, the composition may comprise a factor (ie, a protein) or a nucleic acid encoding the factor. The present invention also provides a method of producing a cell culture medium for (i) maintaining a population of astrocytes in vitro, (ii) generating cells exhibiting at least one characteristic of astrocytes from pluripotent stem or progenitor cells, or (iii) maintaining a population of cardiomyocytes in vitro, (iv) maintaining a population of smooth muscle cells in vitro, or (v) maintaining a population of endothelial cells in vitro, comprising: Methods are provided that include adding a composition described herein to a cell culture medium.

[0093] The present invention also provides a population of cells, wherein at least 5% of the cells exhibit at least one characteristic of a cardiomyocyte, and the cells are produced by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population exhibit at least one characteristic of a cardiomyocyte.

[0094] The present invention also provides a population of cells, wherein at least 5% of the cells exhibit at least one characteristic of astrocytes, and the cells are produced by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population exhibit at least one characteristic of astrocytes.

[0095] The present invention also provides a population of cells, wherein at least 5% of the cells exhibit at least one characteristic of a smooth muscle cell, and the cells are produced by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the cells in the population exhibit at least one characteristic of a smooth muscle cell.

[0096] The present invention also provides a population of cells, wherein at least 5% of the cells exhibit at least one characteristic of an endothelial cell, and the cells are produced by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population exhibit at least one characteristic of an endothelial cell.

[0097] The present invention also relates to kits for use in maintaining cells in vitro as disclosed herein. In some embodiments, the kits include a factor or factors described herein. The kit includes one or more nucleic acids having one or more nucleic acid sequences encoding the gene or variants thereof. Alternatively, the kit includes one or more protein factors for supplementing cell culture media for use as described herein. Preferably, the kit can be used to maintain astrocytes, cardiomyocytes, smooth muscle cells, or endothelial cells in culture. In some embodiments, the kit further includes instructions for maintaining cells, preferably astrocytes, cardiomyocytes, smooth muscle cells, or endothelial cells, in vitro.

[0098] The present invention also relates to kits for producing target cells, preferably cells exhibiting at least one characteristic of astrocytes or cardiomyocytes, as disclosed herein. In some embodiments, the kits include one or more nucleic acids having one or more nucleic acid sequences encoding a factor or variants thereof described herein. Alternatively, the kits may include one or more proteins for supplementing a culture medium for use in directed differentiation as described herein. Preferably, the kits can be used to produce cells exhibiting at least one characteristic of astrocytes or at least one characteristic of cardiomyocytes. Preferably, the kits can be used with embryonic stem cells. In some embodiments, the kits further include instructions for converting source cells into cells exhibiting at least one characteristic of target cells according to the methods disclosed herein. Preferably, the present invention provides kits for use in the methods of the present invention described herein.

[0099] Preferred Embodiments The following description relates to preferred embodiments of the invention. 1. A method for maintaining cells in vitro, comprising: - providing cells of interest in cell culture; - determining the differential width of H3K4me3 modified regions for each protein-coding gene in the cell; - determining a differential broad score (DBS) for each protein-coding gene in the cell; - determining a network score for each protein-coding gene in the cell based on interactions between protein products of each protein-coding gene on the DBS and at least one network; - determining a cellular identity score (RegDBS) for each protein-coding gene in said cell based on a combination of DBS and network scores; - prioritizing each protein-coding gene according to its RegDBS, thereby identifying cell identity genes for said cell, wherein each cell identity gene encodes a factor associated with the cell identity of the cell of interest; - contacting said cell of interest with one or more factors associated with the cellular identity of said cell; - culturing said cells of interest for a time and under conditions sufficient to allow maintenance of said cells in cell culture; thereby maintaining the cells of interest in vitro.

[0100] 2. A method for maintaining cells in vitro, comprising: i. providing cells of interest in cell culture; ii. identifying a protein-coding gene encoding a factor or a variant thereof that promotes the maintenance of said cells, wherein the protein-coding gene comprises at least one H3K4me3 modified region; iii. contacting said cell of interest with at least two factors or variants thereof to maintain at least one characteristic of said cell; iv. culturing said cells of interest for a time and under conditions sufficient to allow maintenance of said cells in cell culture; thereby maintaining the cells of interest in vitro.

[0101] 3. A method for maintaining cells in vitro, comprising: i. providing cells of interest in cell culture; ii. identifying a protein-coding gene encoding a factor or a variant thereof that promotes the maintenance of said cells, wherein the protein-coding gene comprises at least one H3K4me3 modified region; iii. contacting said cell of interest with one or more factors or variants thereof to maintain at least one characteristic of said cell; iv. culturing said cells of interest for a time and under conditions sufficient to allow maintenance of said cells in cell culture; thereby maintaining the cells of interest in vitro.

[0102] 4. The method of any one of statements 1 to 3, wherein the cell of interest is a differentiated cell, a differentiating cell, an undifferentiated cell, or a transdifferentiated cell. 5. The method of any one of statements 1 to 4, wherein the cells of interest are tissue.

[0103] 6. The method of any one of statements 1 to 5, wherein the cells of interest are selected from cells derived from the ectodermal, mesodermal and endodermal germ layers. 7. The method of any one of statements 1 to 6, wherein the cell of interest is selected from the group consisting of an astrocyte, a neurosphere, a cardiomyocyte, a smooth muscle cell, an endothelial cell, a myocyte, a fibroblast, a melanocyte, an epithelial cell, a keratinocyte, a melanocyte, and a mononuclear cell, or any of the cells or tissues listed in Table 2.

[0104] 8. The method of any one of statements 1 to 7, wherein the factor associated with cell identity is any one or more of the factors listed in Table 2. 9. The method of any one of statements 1 to 8, wherein the cells of interest are cultured for a time and under conditions sufficient to allow maintenance of said cells, including culturing the cells for at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 days or more.

[0105] 10. The differential width of H3K4me3 modified regions is - obtaining information about the width of H3K4me3 modified regions for each protein-coding gene in the cell of interest, and obtaining a gene peak width score (B) for each protein-coding gene in the cell; - calculating the difference between the gene width score for each gene in the cell and the average gene peak width score for the same gene across populations of cells of different types, thereby determining a differential width score (DBS) for each protein-coding gene in the cell of interest; 10. The method of any one of statements 1 to 9, wherein the

[0106] 11. The method of claim 10, wherein the step of determining the gene peak width score (B) comprises first defining regions of the genome in which significant H3K4me3 modifications are present and defining those regions as histone modification reference peak loci (RPLs).

[0107] 12. The method of statement 10, wherein the RPL is obtained by merging the obtained ChIP-seq peak areas across all cell types, preferably the merged peak areas are overlapping peak areas.

[0108] 13. The method of any one of statements 9 to 12, wherein the gene peak width score (B) for each protein-coding gene includes excluding genes in which H3K4me3 and H3K27me3 modified regions are identified, and the genes are identified as poised genes.

[0109] 14. The method of any one of statements 1 to 13, wherein the step of determining a cellular identity score (RegDBS) comprises preferentially weighting protein-coding genes that encode regulatory factors.

[0110] 15. The method of any one of claims 1 to 14, wherein the factor for maintaining the cells of interest in vitro is selected from the group consisting of receptor-ligand pairs involved in cell signaling, preferably ligands of receptor-ligand pairs, transcription factors, and epigenetic remodeling factors.

[0111] 16. A method according to any one of statements 1 to 15, comprising the step of selecting cell identity genes encoding receptor-ligand pairs involved in cell signaling, thereby identifying signaling molecules required to maintain the cell of interest in vitro.

[0112] 17. The method of statement 16, wherein each protein-coding gene encoding a cell surface receptor is ranked according to its RegDBS score, and each ligand associated with the receptor is ranked according to its DBS score to obtain a combined ranking of receptors and ligands, and the combined ranking identifies a ligand for use in supplementing culture medium for maintaining a cell of interest in vitro.

[0113] 18. The method of claim 16 or 17, wherein the cell identity gene encodes a receptor-ligand pair involved in a cell signaling pathway such as WNT, NOTCH, Hegdehog, Hippo, GPCR, integrin, TGFB family (BMP, activin, TGFB receptors), receptor tyrosine kinase (EGFR, FGFR, VEGF, PDGF, MET, MST, SCF-KIT, insulin receptor, ERBB2, NTRK, etc.), non-receptor tyrosine kinase (PTK6), MTOR, and retinoic acid-mediated signaling.

[0114] 19. The method of any one of statements 1 to 18, further comprising the step of selecting cell identity genes encoding transcription factors, thereby identifying transcription factors required to maintain the cells of interest in vitro.

[0115] 20. The method of any one of statements 1 to 19, wherein the step of determining the network score comprises removing protein-coding genes that do not have associated peak H3K4me3 width from the protein-protein interaction network.

[0116] 21. The method of any one of statements 1 to 20, wherein the cells of interest are contacted with two or more of the factors by contacting the cells with an agent that increases expression of the two or more factors.

[0117] 22. The method of statement 21, wherein the agent is selected from the group consisting of nucleotide sequences, proteins, aptamers, and small molecules, and analogs or variants thereof. 23. A method of maintaining a population of cells in vitro by carrying out the method of any one of statements 1 to 22.

[0118] 24. A population of cells, wherein at least 5% of the cells exhibit at least one characteristic of a cell of interest, and the cells are produced by a method according to any one of statements 1 to 22. 25. The population of cells of statement 24, wherein at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population exhibit at least one characteristic of a cell of interest.

[0119] 26. The method of any one of statements 1 to 23, further comprising administering to the individual a cell of interest or a population of cells according to claim 22 or 23.

[0120] 27. A method for converting a source cell into a cell that exhibits at least one characteristic of a target cell type, comprising: - providing source cells; - determining the differential width of H3K4me3 modified regions for each protein-coding gene in the source cell and the target cell type to obtain a score of H3K4me3 modification (DBS) for each protein-coding gene in the source and target cell types; - calculating the difference between the DBS of the source cell and the DBS of the target cell to obtain a measure of the difference in H3K4me3 modification between the source cell and the target cell for each protein-coding gene (cell conversion ΔDBS); - determining a network score for the transformation of the source cell type into the target cell type based on the cell transformation ΔDBS and interactions between protein products of each protein-coding gene on at least one network, the network including information on interactions between each protein-coding gene product and other gene products in the cell; - determining a cell identity conversion score for each protein-coding gene in the target cell based on a combination of the cell conversion DBS and the network score; - prioritizing each protein-coding gene according to its cell identity conversion score to identify cell identity genes for the target cell, wherein each cell identity gene encodes a factor associated with the cell identity of the target cell; - culturing the source cells for a time and under conditions sufficient to allow conversion of the source cells into cells exhibiting at least one characteristic of the target cell type; whereby converting the source cells into cells exhibiting at least one characteristic of the target cell type, and optionally increasing the amount of one or more factors, comprises i) contacting the source cells with the factor or an agent that increases expression of the factor by the cells, or ii) transfecting the source cells with a nucleic acid encoding the factor and expressing the nucleic acid in the cells.

[0121] 28. The method of statement 27, wherein the method of converting is a method of differentiation, reprogramming, or transdifferentiation. 29. The method of statement 27, wherein the source cell is a differentiated cell, a differentiating cell, an undifferentiated cell, or a transdifferentiated cell.

[0122] 30. The method of statement 27, wherein the source cell is an embryonic stem cell and the target cell is one of the cells listed in Table 3. 31. The method of statement 30, wherein the source cells are embryonic stem cells and the factor for transforming the source cells is listed in Table 3.

[0123] As used herein, unless the context otherwise requires, the term "comprise" and variations of that term, such as "comprising," "comprises," and "comprised," shall mean any combination of the following: , is not intended to exclude further additions, components, integers or steps.

[0124] Further aspects of the invention and further embodiments of the aspects described in the preceding paragraphs will become apparent from the following description, given by way of example and with reference to the accompanying drawings, in which: [Brief explanation of the drawings]

[0125] [Figure 1-1]The H3K4me3 histone modification marks cell identity genes. (A) Venn diagram showing the number of cell types used in this study. H3K4me3 and H3K27me3 ChIP-seq histone modification data and RNA-seq gene expression data obtained from the ENCODE repository. (B) Definition of ChIP-seq peak widths and peak heights. (C) Gene annotations for representative ChIP-seq profiles of cell types. To map ChIP-seq profiles to the human genome (GRCh38) and compare peaks across cell types, we defined reference peak loci (RPLs) by merging ChIP-seq profiles across all cell types. This table summarizes the genes assigned to each RPL and the calculated peak width (B) and height (H) values ​​at each locus. (D) Enrichment for cell identity and housekeeping gene sets by genes ordered by associated H3K4me3 peak width and gene expression across 40 common cell types. The x-axis consists of cumulative bins of genes by H3K4me3 width value or gene expression level ranked in descending order, with cumulative bins increasing at 1 percentile intervals. Enrichment scores are computed using Fisher's exact test (FET). (E) Similarly, across 111 cell types, the highest enrichment scores for the cell-identity gene set are achieved for genes ranked by H3K4me3 width value and with associated H3K4me3 widths greater than 87% of the peak width distribution. n refers to the number of cell types. [Figure 1-2]The H3K4me3 histone modification marks cell identity genes. (A) Venn diagram showing the number of cell types used in this study. H3K4me3 and H3K27me3 ChIP-seq histone modification data and RNA-seq gene expression data obtained from the ENCODE repository. (B) Definition of ChIP-seq peak widths and peak heights. (C) Gene annotations for representative ChIP-seq profiles of cell types. To map ChIP-seq profiles to the human genome (GRCh38) and compare peaks across cell types, we defined reference peak loci (RPLs) by merging ChIP-seq profiles across all cell types. This table summarizes the genes assigned to each RPL and the calculated peak width (B) and height (H) values ​​at each locus. (D) Enrichment for cell identity and housekeeping gene sets by genes ordered by associated H3K4me3 peak width and gene expression across 40 common cell types. The x-axis consists of cumulative bins of genes by H3K4me3 width value or gene expression level ranked in descending order, with cumulative bins increasing at 1 percentile intervals. Enrichment scores are computed using Fisher's exact test (FET). (E) Similarly, across 111 cell types, the highest enrichment scores for the cell-identity gene set are achieved for genes ranked by H3K4me3 width value and with associated H3K4me3 widths greater than 87% of the peak width distribution. n refers to the number of cell types. [Figure 2-1]Identification of cell identity genes and prediction of signaling molecules for cell maintenance. (A) Schematic of the data-driven method developed to model cell identity genes using associated broad H3K4me3 peaks. For each gene, a differential broadness score (DBS) is computed by comparing the broad H3K4me3 width values ​​between the target cell type of interest and background cell types. DBS is a composite value calculated as the difference in associated width values ​​(Δpeak width) and the significance of this difference (P-value). Protein-coding genes are ranked by DBS. (B) To prioritize genes with regulatory effects, a regulatory differential broadness score (RegDBS) is computed for each gene as a weighted sum of the gene's DBS and the DBS of related genes in a protein-protein interaction (PPI) network. Cell-specific RegDBS are used to predict (i) cell identity genes and (ii) signaling molecules for cell maintenance. (C) Gene Set Enrichment Analysis (GSEA) computationally computes enrichment for cell identity gene sets by different scoring metrics, such as peak width values ​​(peak width), DBS, and RegDBS. For each scoring metric, the sum of GSEA enrichment scores (NES*-log10p value) across all cell types is shown. (D) EpiMogrify predictions of signaling molecules, such as receptors and ligands, for in vitro cell maintenance. While receptors are produced by the cell type of interest, ligands can be produced by the cell type itself or by supporting cell types. Receptor and ligand pairs are ranked based on the receptor's RegDBS and the corresponding ligand's DBS value. The top predicted ligands from the ranked receptor-ligand pairs are added to the culture conditions for in vitro cell maintenance. [Figure 2-2]Identification of cell identity genes and prediction of signaling molecules for cell maintenance. (A) Schematic of the data-driven method developed to model cell identity genes using associated broad H3K4me3 peaks. For each gene, a differential broadness score (DBS) is computed by comparing the broad H3K4me3 width values ​​between the target cell type of interest and background cell types. DBS is a composite value calculated as the difference in associated width values ​​(Δpeak width) and the significance of this difference (P-value). Protein-coding genes are ranked by DBS. (B) To prioritize genes with regulatory effects, a regulatory differential broadness score (RegDBS) is computed for each gene as a weighted sum of the gene's DBS and the DBS of related genes in a protein-protein interaction (PPI) network. Cell-specific RegDBS are used to predict (i) cell identity genes and (ii) signaling molecules for cell maintenance. (C) Gene Set Enrichment Analysis (GSEA) computationally computes enrichment for cell identity gene sets by different scoring metrics, such as peak width values ​​(peak width), DBS, and RegDBS. For each scoring metric, the sum of GSEA enrichment scores (NES*-log10p value) across all cell types is shown. (D) EpiMogrify predictions of signaling molecules, such as receptors and ligands, for in vitro cell maintenance. While receptors are produced by the cell type of interest, ligands can be produced by the cell type itself or by supporting cell types. Receptor and ligand pairs are ranked based on the receptor's RegDBS and the corresponding ligand's DBS value. The top predicted ligands from the ranked receptor-ligand pairs are added to the culture conditions for in vitro cell maintenance. [Figure 2-3]Identification of cell identity genes and prediction of signaling molecules for cell maintenance. (A) Schematic of the data-driven method developed to model cell identity genes using associated broad H3K4me3 peaks. For each gene, a differential broadness score (DBS) is computed by comparing the broad H3K4me3 width values ​​between the target cell type of interest and background cell types. DBS is a composite value calculated as the difference in associated width values ​​(Δpeak width) and the significance of this difference (P-value). Protein-coding genes are ranked by DBS. (B) To prioritize genes with regulatory effects, a regulatory differential broadness score (RegDBS) is computed for each gene as a weighted sum of the gene's DBS and the DBS of related genes in a protein-protein interaction (PPI) network. Cell-specific RegDBS are used to predict (i) cell identity genes and (ii) signaling molecules for cell maintenance. (C) Gene Set Enrichment Analysis (GSEA) computationally computes enrichment for cell identity gene sets by different scoring metrics, such as peak width values ​​(peak width), DBS, and RegDBS. For each scoring metric, the sum of GSEA enrichment scores (NES*-log10p value) across all cell types is shown. (D) EpiMogrify predictions of signaling molecules, such as receptors and ligands, for in vitro cell maintenance. While receptors are produced by the cell type of interest, ligands can be produced by the cell type itself or by supporting cell types. Receptor and ligand pairs are ranked based on the receptor's RegDBS and the corresponding ligand's DBS value. The top predicted ligands from the ranked receptor-ligand pairs are added to the culture conditions for in vitro cell maintenance. [Figure 3-1]In vitro astrocyte cell maintenance. A) Cell counts of primary astrocyte and neural stem cell-derived astrocytes 3 days after supplementation with predictor ligands. The control is no supplementation, and Matrigel is used as the gold standard. B) Cell proliferation rates of primary astrocytes and neural stem cell-derived astrocytes, as detected by cell proliferation assay (BrdU) 3 days after supplementation with predictor ligands. C) Immunofluorescence (IF) images of astrocyte-specific markers, such as GFAP (red) and S100b (green), on primary astrocytes under the indicated culture conditions. (D) IF images of astrocyte-specific markers on neural stem cell-derived astrocytes under the indicated culture conditions. Cells were counterstained with DAPI. Scale = 25 μm. An unpaired, one-tailed t-test was used to compare predictor conditions with the control. *P<0.05, **P<0.01, ***P<0.001, ****P<0.0001, and ns indicates not significant. [Figure 3-2] In vitro astrocyte cell maintenance. A) Cell counts of primary astrocyte and neural stem cell-derived astrocytes 3 days after supplementation with predictor ligands. The control is no supplementation, and Matrigel is used as the gold standard. B) Cell proliferation rates of primary astrocytes and neural stem cell-derived astrocytes, as detected by cell proliferation assay (BrdU) 3 days after supplementation with predictor ligands. C) Immunofluorescence (IF) images of astrocyte-specific markers, such as GFAP (red) and S100b (green), on primary astrocytes under the indicated culture conditions. (D) IF images of astrocyte-specific markers on neural stem cell-derived astrocytes under the indicated culture conditions. Cells were counterstained with DAPI. Scale = 25 μm. An unpaired, one-tailed t-test was used to compare predictor conditions with the control. *P<0.05, **P<0.01, ***P<0.001, ****P<0.0001, and ns indicates not significant. [Figure 3-3]In vitro astrocyte cell maintenance. A) Cell counts of primary astrocyte and neural stem cell-derived astrocytes 3 days after supplementation with predictor ligands. The control is no supplementation, and Matrigel is used as the gold standard. B) Cell proliferation rates of primary astrocytes and neural stem cell-derived astrocytes, as detected by cell proliferation assay (BrdU) 3 days after supplementation with predictor ligands. C) Immunofluorescence (IF) images of astrocyte-specific markers, such as GFAP (red) and S100b (green), on primary astrocytes under the indicated culture conditions. (D) IF images of astrocyte-specific markers on neural stem cell-derived astrocytes under the indicated culture conditions. Cells were counterstained with DAPI. Scale = 25 μm. An unpaired, one-tailed t-test was used to compare predictor conditions with the control. *P<0.05, **P<0.01, ***P<0.001, ****P<0.0001, and ns indicates not significant. [Figure 4] Astrocyte cell maintenance in vitro. A) Cell counts 3 days after supplementation with cardiomyocyte predictor ligands. Negative control is the condition with no additional ligand, and positive control is the condition with Geltrex only. B) Cardiomyocyte cell proliferation rate detected by cell proliferation assay (BrdU) 3 days after supplementation with predictor ligands. C) Fluorescence-activated cell sorting (FACS) for cardiomyocyte markers GATA4+ and NKX2.5+ cells for all conditions. Unpaired one-tailed t-tests were used to compare predictor conditions with controls. *P<0.05, **P<0.01, ***P<0.001, ****P<0.0001, and ns is not significant. [Figure 5-1]HPASMC smooth muscle cell maintenance in vitro. (A) Cell counts 3 and 6 days after supplementation with the predictive ligand for HPASMC smooth muscle cells. The negative control is the condition without additional ligand, and the positive control is the condition with Geltrex alone. (B) Cell proliferation rates of smooth muscle cells detected by cell proliferation assay (BrdU) 3 days after supplementation with the predictive ligand. (C) Fluorescence-activated cell sorting (FACS) results for the smooth muscle cell-specific marker α-SMA+ cells for all conditions. (D) Immunofluorescence images for smooth muscle cell-specific markers, such as SM22 (green), on smooth muscle cells under the indicated culture conditions. Cells were counterstained with Hoechst. Scale = 50 μm. An unpaired, one-tailed t-test was used to compare predictive conditions with the control. *P<0.05, **P<0.01, ***P<0.001, ****P<0.0001, and ns indicates non-significant. [Figure 5-2] HPASMC smooth muscle cell maintenance in vitro. (A) Cell counts 3 and 6 days after supplementation with the predictive ligand for HPASMC smooth muscle cells. The negative control is the condition without additional ligand, and the positive control is the condition with Geltrex alone. (B) Cell proliferation rates of smooth muscle cells detected by cell proliferation assay (BrdU) 3 days after supplementation with the predictive ligand. (C) Fluorescence-activated cell sorting (FACS) results for the smooth muscle cell-specific marker α-SMA+ cells for all conditions. (D) Immunofluorescence images for smooth muscle cell-specific markers, such as SM22 (green), on smooth muscle cells under the indicated culture conditions. Cells were counterstained with Hoechst. Scale = 50 μm. An unpaired, one-tailed t-test was used to compare predictive conditions with the control. *P<0.05, **P<0.01, ***P<0.001, ****P<0.0001, and ns indicates non-significant. [Figure 5-3]HPASMC smooth muscle cell maintenance in vitro. (A) Cell counts 3 and 6 days after supplementation with the predictive ligand for HPASMC smooth muscle cells. The negative control is the condition without additional ligand, and the positive control is the condition with Geltrex alone. (B) Cell proliferation rates of smooth muscle cells detected by cell proliferation assay (BrdU) 3 days after supplementation with the predictive ligand. (C) Fluorescence-activated cell sorting (FACS) results for the smooth muscle cell-specific marker α-SMA+ cells for all conditions. (D) Immunofluorescence images for smooth muscle cell-specific markers, such as SM22 (green), on smooth muscle cells under the indicated culture conditions. Cells were counterstained with Hoechst. Scale = 50 μm. An unpaired, one-tailed t-test was used to compare predictive conditions with the control. *P<0.05, **P<0.01, ***P<0.001, ****P<0.0001, and ns indicates non-significant. [Figure 6-1] (B) Cell maintenance of HAoEC endothelial cells in vitro. Cell counts 3 and 6 days after supplementation with predicted ligands for HAoEC endothelial cells. Negative control is the condition with no additional ligand, and positive control is the condition with Geltrex only. (A) Cell proliferation rate of endothelial cells detected by cell proliferation assay (BrdU) 3 days after supplementation with predicted ligands. (C) FACS for endothelial cell-specific marker VWF+ cells for all conditions. (D) IF for endothelial cell-specific marker CD31 (green) for all conditions. Cells were counterstained with Hoechst. Scale = 50 μm. An unpaired one-tailed t-test was used to compare predicted conditions with controls. *P<0.05, **P<0.01, ***P<0.001, ****P<0.0001, and ns is not significant. [Figure 6-2](B) Cell maintenance of HAoEC endothelial cells in vitro. Cell counts 3 and 6 days after supplementation with predicted ligands for HAoEC endothelial cells. Negative control is the condition with no additional ligand, and positive control is the condition with Geltrex only. (A) Cell proliferation rate of endothelial cells detected by cell proliferation assay (BrdU) 3 days after supplementation with predicted ligands. (C) FACS for endothelial cell-specific marker VWF+ cells for all conditions. (D) IF for endothelial cell-specific marker CD31 (green) for all conditions. Cells were counterstained with Hoechst. Scale = 50 μm. An unpaired one-tailed t-test was used to compare predicted conditions with controls. *P<0.05, **P<0.01, ***P<0.001, ****P<0.0001, and ns is not significant. [Figure 7-1] EpiMogrify: Prediction of Cell Transformation Factors. (A) For cell transformation from a source cell type to a target cell type, the change in cell identity is computationally calculated as the difference in DBS values ​​between the source and target cell types. Cell Transformation RegΔDBS is then calculated as the sum of the gene's ΔDBS and protein-protein interaction network score. Genes are ranked by Cell Transformation RegΔDBS to predict (i) signaling molecules for directed differentiation or cell transformation and (ii) transcription factors for transdifferentiation. (B) For x number of cell transformations, GSEA is used to calculate the enrichment of EpiMogrify's predicted list of TFs for (i) the top 100 TFs predicted by JSD, (ii) the redundant TF set by Mogrify, and (iii) the common TFs predicted by both JSD and Mogrify. The graph shows the percentage of cell transformations with significant enrichment in each case. (C) For transdifferentiation of fibroblasts (source cell type) into 10 selected target cell types, a table summarizes the overlap between EpiMogrify's predicted TFs and previously published TFs for the same in vitro transdifferentiation. For each transdifferentiation, the percentage of recall of published TFs and their respective ranks are shown. [Figure 7-2] EpiMogrify: Prediction of Cell Transformation Factors. (A) For cell transformation from a source cell type to a target cell type, the change in cell identity is computationally calculated as the difference in DBS values ​​between the source and target cell types. Cell Transformation RegΔDBS is then calculated as the sum of the gene's ΔDBS and protein-protein interaction network score. Genes are ranked by Cell Transformation RegΔDBS to predict (i) signaling molecules for directed differentiation or cell transformation and (ii) transcription factors for transdifferentiation. (B) For x number of cell transformations, GSEA is used to calculate the enrichment of EpiMogrify's predicted list of TFs for (i) the top 100 TFs predicted by JSD, (ii) the redundant TF set by Mogrify, and (iii) the common TFs predicted by both JSD and Mogrify. The graph shows the percentage of cell transformations with significant enrichment in each case. (C) For transdifferentiation of fibroblasts (source cell type) into 10 selected target cell types, a table summarizes the overlap between EpiMogrify's predicted TFs and previously published TFs for the same in vitro transdifferentiation. For each transdifferentiation, the percentage of recall of published TFs and their respective ranks are shown. [Figure 8-1]Directed differentiation of astrocytes in vitro. (A) Schematic of in vitro differentiation of astrocytes from H9 embryonic stem cells (H9 ESCs). After 14 days, neural cell adhesion molecule (NCAM)+ neural progenitor cells were selected, and cells were seeded under different predictive conditions. Immunofluorescence (IF), fluorescence-activated cell sorting (FACS), and RNA-seq analyses were performed. (B) FACS results for the astrocyte marker CD44+ cells across all conditions. (C) IF results for GFAP (red) and S100b (green) at day 21 in Matrigel-only and all ligand conditions. (D) IF was performed to obtain the percentage of DAPI+ cells bearing the GFAP and S100b astrocyte markers across all conditions. For the group of genes exhibiting ε-astrocyte specificity (astrocyte transcriptional signature), the average z-score of gene expression profiles in the TPM is calculated across the three samples sequenced for each time point and condition. The astrocyte transcriptional signature was defined as (i) significantly upregulated genes between H9 ESCs and primary astrocytes obtained from our RNA-seq data, (ii) significantly upregulated genes between H9 ESCs and cerebral cortex-derived astrocytes obtained from the FANTOM5 (F5) database, and (iii) significantly upregulated genes between H9 ESCs and cerebellum-derived astrocytes obtained from the F5 database. (iv) EpiMogrify GRN (gene regulatory network) is a gene set containing cell identity genes of primary astrocytes with a positive RegDBS score and associated genes on the STRING network up to the first regulated neighbor. (v) HumanBase GRN is an astrocyte-specific network obtained from the HumanBase database. Cells were counterstained with DAPI. Scale = 25 μm. Unpaired one-tailed t-tests were used to compare predicted conditions with controls. *P<0.05, **P<0.01, ***P<0.001 and ns is not significant. [Figure 8-2]Directed differentiation of astrocytes in vitro. (A) Schematic of in vitro differentiation of astrocytes from H9 embryonic stem cells (H9 ESCs). After 14 days, neural cell adhesion molecule (NCAM)+ neural progenitor cells were selected, and cells were seeded under different predictive conditions. Immunofluorescence (IF), fluorescence-activated cell sorting (FACS), and RNA-seq analyses were performed. (B) FACS results for the astrocyte marker CD44+ cells across all conditions. (C) IF results for GFAP (red) and S100b (green) at day 21 in Matrigel-only and all ligand conditions. (D) IF was performed to obtain the percentage of DAPI+ cells bearing the GFAP and S100b astrocyte markers across all conditions. For the group of genes exhibiting ε-astrocyte specificity (astrocyte transcriptional signature), the average z-score of gene expression profiles in the TPM is calculated across the three samples sequenced for each time point and condition. The astrocyte transcriptional signature was defined as (i) significantly upregulated genes between H9 ESCs and primary astrocytes obtained from our RNA-seq data, (ii) significantly upregulated genes between H9 ESCs and cerebral cortex-derived astrocytes obtained from the FANTOM5 (F5) database, and (iii) significantly upregulated genes between H9 ESCs and cerebellum-derived astrocytes obtained from the F5 database. (iv) EpiMogrify GRN (gene regulatory network) is a gene set containing cell identity genes of primary astrocytes with a positive RegDBS score and associated genes on the STRING network up to the first regulated neighbor. (v) HumanBase GRN is an astrocyte-specific network obtained from the HumanBase database. Cells were counterstained with DAPI. Scale = 25 μm. Unpaired one-tailed t-tests were used to compare predicted conditions with controls. *P<0.05, **P<0.01, ***P<0.001 and ns is not significant. [Figure 8-3]Directed differentiation of astrocytes in vitro. (A) Schematic of in vitro differentiation of astrocytes from H9 embryonic stem cells (H9 ESCs). After 14 days, neural cell adhesion molecule (NCAM)+ neural progenitor cells were selected, and cells were seeded under different predictive conditions. Immunofluorescence (IF), fluorescence-activated cell sorting (FACS), and RNA-seq analyses were performed. (B) FACS results for the astrocyte marker CD44+ cells across all conditions. (C) IF results for GFAP (red) and S100b (green) at day 21 in Matrigel-only and all ligand conditions. (D) IF was performed to obtain the percentage of DAPI+ cells bearing the GFAP and S100b astrocyte markers across all conditions. For the group of genes exhibiting ε-astrocyte specificity (astrocyte transcriptional signature), the average z-score of gene expression profiles in the TPM is calculated across the three samples sequenced for each time point and condition. The astrocyte transcriptional signature was defined as (i) significantly upregulated genes between H9 ESCs and primary astrocytes obtained from our RNA-seq data, (ii) significantly upregulated genes between H9 ESCs and cerebral cortex-derived astrocytes obtained from the FANTOM5 (F5) database, and (iii) significantly upregulated genes between H9 ESCs and cerebellum-derived astrocytes obtained from the F5 database. (iv) EpiMogrify GRN (gene regulatory network) is a gene set containing cell identity genes of primary astrocytes with a positive RegDBS score and associated genes on the STRING network up to the first regulated neighbor. (v) HumanBase GRN is an astrocyte-specific network obtained from the HumanBase database. Cells were counterstained with DAPI. Scale = 25 μm. Unpaired one-tailed t-tests were used to compare predicted conditions with controls. *P<0.05, **P<0.01, ***P<0.001 and ns is not significant. [Figure 8-4]Directed differentiation of astrocytes in vitro. (A) Schematic of in vitro differentiation of astrocytes from H9 embryonic stem cells (H9 ESCs). After 14 days, neural cell adhesion molecule (NCAM)+ neural progenitor cells were selected, and cells were seeded under different predictive conditions. Immunofluorescence (IF), fluorescence-activated cell sorting (FACS), and RNA-seq analyses were performed. (B) FACS results for the astrocyte marker CD44+ cells across all conditions. (C) IF results for GFAP (red) and S100b (green) at day 21 in Matrigel-only and all ligand conditions. (D) IF was performed to obtain the percentage of DAPI+ cells bearing the GFAP and S100b astrocyte markers across all conditions. For the group of genes exhibiting ε-astrocyte specificity (astrocyte transcriptional signature), the average z-score of gene expression profiles in the TPM is calculated across the three samples sequenced for each time point and condition. The astrocyte transcriptional signature was defined as (i) significantly upregulated genes between H9 ESCs and primary astrocytes obtained from our RNA-seq data, (ii) significantly upregulated genes between H9 ESCs and cerebral cortex-derived astrocytes obtained from the FANTOM5 (F5) database, and (iii) significantly upregulated genes between H9 ESCs and cerebellum-derived astrocytes obtained from the F5 database. (iv) EpiMogrify GRN (gene regulatory network) is a gene set containing cell identity genes of primary astrocytes with a positive RegDBS score and associated genes on the STRING network up to the first regulated neighbor. (v) HumanBase GRN is an astrocyte-specific network obtained from the HumanBase database. Cells were counterstained with DAPI. Scale = 25 μm. Unpaired one-tailed t-tests were used to compare predicted conditions with controls. *P<0.05, **P<0.01, ***P<0.001 and ns is not significant. [Figure 9-1]Directed differentiation of cardiomyocytes in vitro. (A) Schematic of in vitro differentiation of cardiomyocytes from H9 embryonic stem cells (H9 ESCs). After 3 days, H9 ESCs were seeded under different predictive conditions and the positive control Matrigel alone. Immunofluorescence (IF) and fluorescence-activated cell sorting (FACS) were performed on days 12 and 20. (B) FACS for cardiac markers CD82+ and CD13+ cells for all conditions. (C) IF was performed to obtain the percentage of DAPI+ area with the cTnT cardiac marker. (D) IF for cTnT (green) on day 21 for all conditions. Cells were counterstained with DAPI. Scale = 25 μm. An unpaired one-tailed t-test was used to compare predictive conditions with controls. *P<0.05, **P<0.01, and ns is not significant. [Figure 9-2] Directed differentiation of cardiomyocytes in vitro. (A) Schematic of in vitro differentiation of cardiomyocytes from H9 embryonic stem cells (H9 ESCs). After 3 days, H9 ESCs were seeded under different predictive conditions and the positive control Matrigel alone. Immunofluorescence (IF) and fluorescence-activated cell sorting (FACS) were performed on days 12 and 20. (B) FACS for cardiac markers CD82+ and CD13+ cells for all conditions. (C) IF was performed to obtain the percentage of DAPI+ area with the cTnT cardiac marker. (D) IF for cTnT (green) on day 21 for all conditions. Cells were counterstained with DAPI. Scale = 25 μm. An unpaired one-tailed t-test was used to compare predictive conditions with controls. *P<0.05, **P<0.01, and ns is not significant. [Figure 9-3]Directed differentiation of cardiomyocytes in vitro. (A) Schematic of in vitro differentiation of cardiomyocytes from H9 embryonic stem cells (H9 ESCs). After 3 days, H9 ESCs were seeded under different predictive conditions and the positive control Matrigel alone. Immunofluorescence (IF) and fluorescence-activated cell sorting (FACS) were performed on days 12 and 20. (B) FACS for cardiac markers CD82+ and CD13+ cells for all conditions. (C) IF was performed to obtain the percentage of DAPI+ area with the cTnT cardiac marker. (D) IF for cTnT (green) on day 21 for all conditions. Cells were counterstained with DAPI. Scale = 25 μm. An unpaired one-tailed t-test was used to compare predictive conditions with controls. *P<0.05, **P<0.01, and ns is not significant. [Figure 10] 6 is a block diagram of one type of computer processing system 600 for implementing embodiments and / or features of the methods described herein. DETAILED DESCRIPTION OF THE INVENTION

[0126] It is understood that the invention disclosed and defined herein extends to all alternative combinations of two or more of the individual features mentioned or apparent from the text or drawings, all of these different combinations constituting various alternative aspects of the invention.

[0127] Reference will now be made in detail to certain embodiments of the invention. While the invention will be described in conjunction with the embodiments, it will be understood that the intention is not to limit the invention to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents which may be included within the scope of the present invention as defined by the claims.

[0128] Those skilled in the art will recognize many methods and materials similar or equivalent to those described herein that can be used to practice the present invention. The present invention is in no way limited to the methods and materials described. It is understood that the invention disclosed and defined herein extends to all alternative combinations of two or more of the individual features mentioned or apparent from the text or drawings. All of these different combinations constitute various alternative aspects of the present invention.

[0129] For the purposes of this description, terms used in the singular will also include the plural and vice versa. We aimed to systematically identify factors that would facilitate the development of serum-free, chemically defined cell maintenance and differentiation media. To identify factors for cell maintenance or cell transformation, we needed to model changes in cell identity or cellular identity, respectively.

[0130] The present invention utilizes available epigenetic databases (e.g., from the ENCODE and Roadmap consortia) and provides a novel computational approach (EpiMogrify) that uses H3K4me3 and H3K27me3 histone modifications to model the epigenetic state of cells. EpiMogrify uses data-driven thresholds, statistics, and incorporates protein-protein interaction network information to prioritize key genes that regulate cellular identity (thereby identifying cell identity genes for a given cell). The method of the present invention systematically predicts signaling molecules for cell maintenance and directed differentiation in multiple cell types. Furthermore, the method of the present invention can also prioritize other protein classes, such as transcription factors (TFs) or epigenetic remodelers, for cell maintenance or cell transformation.

[0131] EpiMogrify algorithm The present invention provides a method for determining cell identity genes for a cell type, comprising: - selecting a subset of protein-coding genes based on a priority of the protein-coding genes in the cell type according to the effect of the protein-coding genes on cell identity; and the priority for a given protein-coding gene is The present invention provides a method based on a combination of information about the breadth of H3K4me3 modified regions in a cell and information about the regulatory interactions of a protein-coding gene product with other protein-coding gene products in the cell.

[0132] The method of the present invention (referred to herein as "EpiMogrify") models the epigenetic state of a cell to predict factors important for cell identity, cell maintenance, and cell transformation. The main sections of the algorithm are: (i) Defining histone modification reference peak loci; (ii) identification of cell identity and cell maintenance factors; (iii) Identification of cellular transformation factors for directed differentiation and transdifferentiation Includes.

[0133] This approach aims to model cellular states by leveraging epigenetic histone modifications, such as the H3K4me3 modification, a well-known transcriptional activator mark. ChIP-seq data for H3K4me3 and H3K27me3 (a repressor mark) histone modification profiles are available for a variety of human cell types, including from the ENCODE and Roadmap Consortium data repositories.

[0134] The width of a ChIP-seq peak can be defined as the length of the genomic region with histone modification deposition in base pairs (bp), and the height of a ChIP-seq peak can be defined as the average enrichment of the histone modification or signal value. It is understood that the width of a ChIP-seq peak can range from a narrow peak to a wide peak, and the height of a ChIP-seq peak can range from a short peak to a tall peak.

[0135] (i) Definition of histone modification reference peak loci For each cell type, samples can be pooled together to obtain a ChIP-seq profile representative of the cell type. The ChIP-seq peak width representative of the cell type is typically calculated by merging the peak area across the samples with the peak height calculated by the maximum peak height across all samples. To compare ChIP-seq across various cell types, we defined a reference peak locus (RPL), which is a set of regions obtained by combining representative peaks across all cell types.

[0136] Let n be the number of cell types. For each cell type in the range x1, x2, ... xn, let b be the ChIP-seq peak width for the cell type and h be the ChIP-seq peak height. The genomic location of the reference peak locus (RPL) is obtained by merging peak regions across all cell types. For each locus R1, R2, ... Rn in the RPL, the peak width (b) is calculated by merging overlapping peaks across cell types, and the peak height (h) is given by the maximum peak height of the overlapping peaks.

[0137]

number

[0138] Protein-coding genes can be assigned to RPLs based on the genomic location of the peaks and the transcription start site (TSS) of the gene. Because the H3K4me3 histone modification is a well-known promoter mark, we identified the peaks as located in the promoter region of the gene (500 bp from the TSS). In the case of overlap, the gene is typically assigned to the peak locus, however, it is understood that the RPL can be located in another region of the gene relative to its TSS (including extending the region more than 500 bp from the TSS, including 1000 bp, 2000 bp or more from the TSS).

[0139] Preferably, for each cell type (x), the peak width score (B) of a gene (g) is calculated as the sum of the peak widths of the n peaks annotated to the gene, while the peak height value of the gene is calculated as the maximum height of the n peaks annotated to the gene.

[0140]

number

[0141] is the peak width profile of a given cell type x at the RPL (reference peak locus).

[0142]

number

[0143] (ii) Identification of cell identity genes and cell maintenance factors EpiMogrify utilizes H3K4me3 ChIP-seq peak widths to model cellular states.

[0144] EpiMogrify uses a three-step approach to identify cell-type-specific factors. First, a differential breadth score is computed for each RPL based on the H3K4me3 ChIP-seq peak width. Second, the regulatory effect of each gene is determined based on the value of related genes on a protein-protein interaction network. Finally, cell identity genes are predicted based on the ranked protein-coding genes, and EpiMogrify predicts signaling molecules for cell state maintenance.

[0145] As used herein, the term "each" with respect to a protein-coding gene may refer to a plurality of protein-coding genes in a given cell type. The plurality may include all protein-coding genes in the cell type. Alternatively, the plurality may include a subset of the protein-coding genes in the cell type. Thus, "each protein-coding gene" may include "each protein-coding gene in a cell type" or "each protein-coding gene in a cell type that has been further refined to include only a subset of the genes."

[0146] It is further understood that the methods of the present invention require "determining the differential width of H3K4me3 modified regions for each protein-coding gene in a cell of interest," and that if a gene does not contain an H3K4me3 modified region, the relevant gene is excluded from further analysis. Thus, the present invention is limited to evaluating genes in which at least H3K4me3 modification is present.

[0147] In certain embodiments, a subset of the plurality of protein-coding genes may be excluded in one or more of the steps of the method, for example, poised genes in which both H3K4me3 and H3K27me3 ChIP-seq peaks are present at the TSS of the gene. The child is preferably removed from the model before calculating the DBS.

[0148] Those skilled in the art will understand that the greater the number of protein-coding genes included in the analysis, the more robust the prediction of cell identity factors will be. Furthermore, those skilled in the art will recognize the need to balance this with the need to also consider including protein-coding genes that provide the most information regarding cell identity (such as those protein-coding genes that contain only H3K4me3 modifications at or near the TSS).

[0149] Furthermore, it is understood that those genes for which differential H3K4me3 breadth is determined in later steps of the EpiMogrify method will also be included in the subsequent network analysis. In other words, genes considered in the first step of the method (determining the differential breadth score) will typically also be considered in the subsequent calculation of the network score.

[0150] Step 1: Calculate the differential wideness score To obtain a target cell type-specific ChIP-seq profile, the target cell type of interest can be compared to a set of background cell types within the group. For example, if the target cell type of interest is a primary cell type, the background cell types remain primary and stem cell types.

[0151] If the cell types in the background set are dissimilar to the target cell type, the statistical significance of the modeling will be improved. Therefore, in a preferred embodiment, the background cell type is selected only if the Spearman correlation between the ChIP-seq profile and the target cell type is less than 0.9.

[0152] If more than one Reference Peak Locus (RPL) is assigned to a gene (g) in a target cell type of interest (x), then the gene peak width score (B) is given by the union of the peak width (b) values ​​of the assigned RPLs. x g is the set of gene peak width scores for background cell types in gene g, where the background cell types are selected based on their distinctiveness from the target cell type (x).x g is the normalized difference in gene width score between the target cell type and the average gene peak width score of the background cell type. The number of approaches is x g It is understood that gene peak width scores may be averaged to determine the mean, modality, or average of the gene peak width scores. Although the median is used in this example, one skilled in the art will understand that the mean, mode, or other indicators of the mean may also be used.

[0153] The significance of this difference (p-value) can be estimated by any suitable statistical method known to those skilled in the art (one example is the one-sample Wilcoxon test). The differential broadness score (DBS) of a cell type (x) and gene (g) can be measured as the sum of normalized Δpeak width and Pval, as shown below, although it will be understood that any number of methods for combining the peak width difference, and the significance of that difference, can be used (including multiplication, subtraction, normalization, or a combination thereof).

[0154]

number

[0155] Step 2: Calculating the accommodative differential broadness score Information is included to computationally calculate the regulatory effect of genes on cell type cellular identity genes, STRING V10, and protein-protein interaction networks.

[0156] Protein-protein interaction networks such as STRING provide information about the source of information for predicting interactions.For example, interactions can be based on experimental data or in silico data, or both.In STRING, these interactions are also scored, providing a measure of the reliability of predicted interactions.In a specific embodiment of the present invention, if the experimental score provided in the network is greater than zero, and the combined score is greater than 500, the interaction between gene nodes is selected.This ensures that the high-quality network with experimental evidence is used to determine the regulatory effect of genes.

[0157] For every gene (V) in the STRING network, the network score is calculated as the number of genes in the sub-network (V g ) can be computed based on the combination of DBS scores of related genes (r) in the STRING network, corrected for the number of gene outdegree nodes (O) and the level of relatedness (L). To obtain cell-specific networks, genes without associated broad H3K4me3 peaks from the STRING network can be removed. A weighted sum of the DBS scores of related genes can be calculated up to the third level of relatedness.

[0158]

number

[0159] Normalized network score (Net) across all protein-coding genes G in cell type x

[0160]

number

[0161] Different combinations of DBS and Net scores can be used to determine the Cell Identity Score (RegDBS). In a preferred embodiment, a 2:1 ratio of DDBS and Net scores is used, as shown below. The BS vs. network score (Net) is used.

[0162]

number

[0163] Step 3: Prediction of cell identity genes and cell maintenance factors EpiMogrify predicts protein-coding genes that exhibit cellular identity by ranking genes based on their cellular identity score (RegDBS).

[0164] It is understood that protein-protein interaction networks provide information about the connectivity of products (i.e., proteins) of protein-coding genes, rather than the connectivity of the genes themselves. Furthermore, one skilled in the art will understand that there may be more than one "product" of a protein-coding gene. Thus, in the methods of the present invention, products of a protein-coding gene are understood to include any one or all combinations of all products from a given protein-coding gene.

[0165] With regard to cell maintenance, EpiMogrify predicts signaling molecules, such as receptors and ligands, that are essential for cell growth and survival. We recognize that approximately two-thirds of ligands are produced by cells in an autocrine manner, while the remainder are produced by supporting cell types to mimic the microenvironment. Therefore, to incorporate this into the current model, receptors can be prioritized based on cell-specific RegDBS, as they should both be cell-specific and have a regulatory effect on the cell type of interest. Ligands from the cell type of interest and supporting cell types can then be prioritized based on DBS. Finally, receptor-ligand pairs can then be prioritized by the combined rank of the receptor and ligand. This approach allows for the identification and prioritization of ligands of predicted receptor-ligand pairs for use in supplementing cell culture conditions.

[0166] (iii) Identification of cellular transformation factors for directed differentiation and transdifferentiation EpiMogrify utilizes a three-step approach to identify cell transformation factors. First, the change in differential broadening score is computed for cell transformation from a source cell type to a target cell type. Then, the regulatory effect of genes on the change in cell state is determined. Finally, transcription factors are predicted for transdifferentiation, and signaling molecules are predicted for directed differentiation.

[0167] Briefly, the cell transformation differential broadening score is calculated by:

[0168]

number

[0169] where S is the source cell type and T is the target cell type. First, a differential broadening score (DBS) is computed for both the source and target cell types. Next, Then, the difference between the source DBS score and the target DBS score is computed to obtain the cell transformation ΔDBS.

[0170] The regulatory network score (N) for cell transformation from the source cell type to the target cell type is then computed as a weighted sum of the cell transformation ΔDBS scores of the associated nodes. Similar to the calculation of cell-specific RegDBS, the cell transformation RegDBS is calculated as the aggregate value of the cell transformation ΔDBS and the normalized network score (Net) in a 2:1 ratio.

[0171]

number

[0172] For each cell transformation, protein-coding genes are ranked by their RegΔDBS value. For transdifferentiation, EpiMogrify predicts transcription factors (TFs), a subset of protein-coding genes defined based on the TFClass classification. For directed differentiation, EpiMogrify predicts signaling molecules such as receptor and ligand pairs. Receptors are prioritized by the cell transformation RegΔDBS, and corresponding ligands are prioritized by the cell transformation ΔDBS. Finally, receptor-ligand pairs are prioritized by the combined rank of the receptor and ligand. Ligands in predicted receptor-ligand pairs can be supplemented into differentiation protocols.

[0173] In summary, EpiMogrify models a wide range of H3K4me3 histone modification signatures and predicts cell identity protein-coding genes, signaling molecules for cell maintenance, signaling molecules for directed differentiation, and TFs for transdifferentiation.

[0174] definition As used herein, a "factor" for maintaining cells in culture or for converting source cells into target cells can include proteins, small molecules, or other agents for use in supplementing cell culture media. Preferably, the factor is a protein. The protein can be a signaling molecule comprising a ligand that binds to a cognate receptor, preferably on the cell surface, and triggers a signaling cascade within the cell to maintain at least one characteristic of the cell type. Alternatively, the factor can be a transcription factor that promotes expression of genes associated with cell identity, thereby facilitating maintenance of the cell type in vitro. When the factor is a transcription factor, the transcription factor can be provided directly to the cell or, alternatively, can be expressed within the cell by a nucleic acid expression vector or the like.

[0175] In certain embodiments, the methods of the invention involve the use of small molecules that act on cells to replicate the action of the factors identified herein or to increase the amount of the factors described herein (e.g., to increase transcription of the genes encoding the factors).

[0176] In certain embodiments, factors may be provided exogenously, for example, in the form of recombinant or synthetic proteins / peptides to supplement the culture medium. In further embodiments, factors may be provided in a paracrine manner, for example, the factor may be delivered to a target cell from a different cell. The factor is secreted into the cells of interest. Still further, the factor may be provided in an autocrine manner, where the cells of interest are treated to induce expression / production of the factor (e.g., expression of a nucleic acid molecule or vector encoding the factor by the cells of interest or source cells, as the case may be). It is also understood that a combination of exogenous, paracrine, and autocrine approaches to providing factors may be utilized.

[0177] As used herein, the "H3K4" modification refers to an epigenetic modification of chromatin that affects the regulation of gene expression. H3K4 refers to the addition of a methyl group to lysine 4 on histone H3 protein. Thus, H3K4me3 refers to the addition of three methyl groups (trimethylation at the same lysine residue). Histone H3 protein is used to package DNA in eukaryotic cells, and modifications to histones alter the accessibility of genes for transcription. H3K4me3 is usually associated with the activation of transcription of nearby genes. H3K4 trimethylation regulates gene expression through chromatin remodeling by the NURF complex. In bivalent chromatin, H3K4me3 can colocalize with the repressive modification H3K27me3 to control gene regulation.

[0178] The methods of the present invention find utility in identifying factors for use in cell maintenance or collection of cells of interest, transdifferentiation, reprogramming, directed differentiation and / or conversion of source cells to target cells.

[0179] It is understood that "maintaining cells of interest in vitro" can include culturing cells in cell culture conditions such that the cells retain key morphological and biophysical characteristics. This can be measured, for example, by confirming that markers associated with the cells of interest are retained after several passages in cell culture. It is understood that "maintaining cells of interest" and "retaining key morphological and biophysical characteristics" can include maintaining the viability of at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or more of the cells in culture. Preferably, the cells are maintained for at least 2 days, at least 5 days, at least 10 days, at least 20 days, at least 30 days, at least 60 days, or more. Furthermore, "retaining key and biophysical characteristics" can be understood to mean that cells maintained in culture according to the methods of the present invention retain at least 90%, at least 80%, at least 70%, at least 60%, or at least 50% of the markers associated with the cells of interest. Preferably, the cells retain at least 70% of the markers during the culture period. Furthermore, "retaining key biophysical characteristics" may also be understood to mean that cells maintained in culture according to the methods of the invention retain at least 90%, at least 80%, at least 70%, at least 60%, or at least 50% of the morphological properties associated with the cells of interest. Preferably, the cells retain at least 70% of the morphological properties associated with the cells of interest during the culture period.

[0180] It is understood that the term "cell conversion" (e.g., from a source cell to a target cell) can refer to dedifferentiation, transdifferentiation, reprogramming, or directed differentiation. As used herein, dedifferentiation refers to a method of reverting a terminally differentiated cell to a less differentiated stage within its own lineage, thereby allowing it to proliferate. Directed differentiation refers to the differentiation of a pluripotent cell state (e.g., stem cell) into a specific type of differentiated target cell.

[0181] Transdifferentiation typically refers to a process by which cells are modified to the point where they can either switch lineages or develop directly between two different cell types. Cellular reprogramming typically involves converting mature, specialized cells into induced pluripotent stem cells. Refers to the process of returning

[0182] As used herein, a cell of interest is any cell for which it is necessary to identify factors (including receptor / ligand pairs, ligands, transcription factors, or other agents) for maintaining the cell in vitro. A cell of interest can be any cell found in any of the three embryonic germ layers: ectoderm, mesoderm, and endoderm. As used herein, ectoderm refers to extracorporeal tissues such as skin and nails; mesoderm refers to tissues such as muscle, but also blood; and endoderm refers to the inner lining, such as cells of the digestive tract. Examples of various cells of interest are provided in Table 2. Examples of cells of interest derived from mesoderm include cells of skeletal muscle, skeleton, dermis of skin, connective tissue, genitourinary system, heart, blood (lymphocytes), and spleen. Examples of cells derived from endoderm include cells of the stomach, colon, liver, pancreas, and bladder; the epithelial portion of the urethra, trachea, lung, pharynx, thyroid, parathyroid, and intestinal lining. Examples of cells derived from the ectoderm include cells of the central nervous system, retina and lens, cranial and sensory, ganglionic and neural, pigment cells, head connective tissue, epidermis, hair, and mammary gland.

[0183] A "cell of interest" can include any cell suitable for maintenance in tissue culture. Non-limiting examples of cells of interest include the cells listed in Table 2. More specifically, as used herein, a "cell of interest" or "source cell" can be a primary cell (non-immortalized cell), a cell derived from a cell line (immortalized cell), or tissue isolated from an individual. A cell of interest can also be a progenitor cell, e.g., a neural progenitor cell, a hematopoietic progenitor cell, a myogenic progenitor cell, an endothelial progenitor cell, and a general lymphoid or myeloid progenitor cell. A cell of interest can also refer to a population of cells, e.g., a mixed population of cells, such as a tissue or cells obtained from a tissue. In some embodiments, a cell of interest or a source cell is an undifferentiated cell. In other embodiments, a cell of interest or a source cell is a transdifferentiated cell or a cell undergoing transdifferentiation. A transdifferentiated cell is a cell generated from the process of transdifferentiation, i.e., the process of changing the phenotype of a somatic cell into another somatic cell. As used herein, a "somatic cell" is a cell derived from one of the germ layers (ectoderm, endoderm, or mesoderm). A transdifferentiated cell is a cell undergoing the process of transdifferentiation; that is, the phenotype of the somatic cell has changed to that of another somatic cell, and the markers of the definitive somatic cell have not yet been fully established. In other embodiments, the cell of interest or source cell is a differentiated cell, which is a specialized cell type.

[0184] As used herein, a source cell or a cell of interest can be any cell type described herein, including a somatic cell or a diseased cell. A somatic cell can be an adult cell or a cell derived from an adult that exhibits one or more detectable characteristics of an adult or non-embryonic cell. Other examples of source cells include hematopoietic cells, such as progenitor cells, such as lymphocytes and bone marrow cells; buccal mucosa cells, epidermal cells, mesenchymal cells, keratinocytes, hepatocytes, embryonic cells, stem cells, or cells that have been processed to have one or more characteristics of stem cells. In the case of diseased cells, the diseased cells can be cells that exhibit one or more detectable characteristics of a disease or condition, for example, the diseased cells can be cancer cells that exhibit one or more clinical or biochemical markers of cancer.

[0185] The source cells or cells of interest can also be stem cells, including pluripotent stem cells, preferably induced pluripotent stem cells (iPSCs). As used herein, the term "somatic cell" refers to any cell that forms the body of an organism, as opposed to a germline cell. In mammals, germline cells (also known as "gametes") are the sperm and eggs that fuse during fertilization to produce a cell called a zygote, from which the entire mammalian embryo develops. With the exception of sperm and eggs, the cells from which they are generated (gametocytes), and undifferentiated stem cells, all other cell types in the mammalian body are somatic cells: internal organs, skin, bone, blood, and connective tissue, all of which are somatic cells. In some embodiments, the somatic cell is a "non-embryonic somatic cell," which refers to a somatic cell that is not present in or obtained from an embryo and does not result from the propagation of such cells in vitro. In some embodiments, the somatic cell is an "adult somatic cell," which refers to a cell that is present in or obtained from an organism other than an embryo or fetus, or that results from the propagation of such cells in vitro. Somatic cells can be immortalized to provide an unlimited supply of cells, for example, by increasing the level of telomerase reverse transcriptase (TERT). For example, the level of TERT can be increased by increasing the transcription of TERT from an endogenous gene or by introducing a transgene via any gene delivery method or system.

[0186] Unless otherwise indicated, methods for converting somatic cells can be performed both in vivo and in vitro (in vivo, when the somatic cells are present within a subject, and in vitro, when performed using isolated somatic cells maintained in culture).

[0187] Embryonic cells, such as embryonic stem cells, can be cells derived from an embryonic cell line and can be cells that are not directly derived from an embryo or fetus. Alternatively, embryonic cells can be derived from an embryo or fetus, but the cells are obtained or isolated without destroying the embryo or fetus or adversely affecting the development of the embryo or fetus.

[0188] The present invention also contemplates the use of induced pluripotent stem cells (iPSCs). Differentiated somatic cells, including cells from fetal, neonatal, juvenile, or adult primates, including humans, are suitable source cells (or cells of interest) in the methods of the present invention. Suitable somatic cells include, but are not limited to, bone marrow cells, epithelial cells, endothelial cells, fibroblasts, hematopoietic cells, keratinocytes, hepatocytes, intestinal cells, mesenchymal cells, bone marrow progenitor cells, and spleen cells. Alternatively, somatic cells can be cells capable of self-proliferation and differentiation into other cell types, including blood stem cells, muscle / bone stem cells, brain stem cells, and liver stem cells. Suitable somatic cells are receptive to transcription factor uptake, including genetic material encoding the transcription factor, or can be made receptive using methods commonly known in the scientific literature. Methods for enhancing uptake may vary depending on the cell type and expression system. Exemplary conditions used to prepare receptive somatic cells with suitable transduction efficiency are well known to those of skill in the art. Starting somatic cells may have a doubling time of approximately 24 hours.

[0189] As used herein, the term "isolated cell" refers to a cell that has been removed from the organism in which it was originally found or the progeny of such a cell. Optionally, the cell has been cultured in vitro, for example, in the presence of other cells. Optionally, the cell is later introduced into a second organism or reintroduced into the organism from which it was isolated (or the cell from which it was derived).

[0190] As used herein, the term "isolated population," with respect to an isolated population of cells, refers to a population of cells that has been removed and separated from a mixed or heterogeneous population of cells. In some embodiments, an isolated population is a substantially pure population of cells compared to the heterogeneous population from which the cells are isolated or enriched.

[0191] With respect to a particular cell population, the term "substantially pure" refers to a population of cells that is at least about 75%, preferably at least about 85%, more preferably at least about 90%, and most preferably at least about 95% pure, relative to the cells that make up the total cell population. In contrast, with respect to a population of target cells, the term "substantially pure" or "essentially purified" refers to a population of cells that is less than about 20%, more preferably less than about 15%, 10%, 8%, 7%, and most preferably less than about 5%, 4%, 3%, 2%, 1%, or less than 1% pure, as defined herein. "Target cells" refers to a population of cells that contain cells that are not target cells or their progeny as defined by the present invention.

[0192] As used herein, a reference to a "target cell" may be a reference to any one or more of the cells referred to herein as target cells or target cell types. Target cells can be cells that represent any of the three embryonic germ layers: endoderm, mesoderm, and ectoderm. For example, target cells are typically found in skeletal muscle, skeleton, dermis of skin, connective tissue, urogenital system, heart, blood (lymphocytes), and spleen (mesoderm); stomach, colon, liver, pancreas, bladder; urethra, epithelial part of trachea, lung, pharynx, thyroid, parathyroid, intestinal lining (endoderm); or central nervous system, retina and lens, skull and sensory, ganglion and nerve, pigment cell, head connective tissue, epidermis, hair, mammary gland (ectoderm).

[0193] A source cell is determined to be converted into a target cell or become a target-like cell by the method of the present invention if it exhibits at least one characteristic of the target cell type. For example, a human fibroblast is identified as being converted into a keratinocyte-like cell if the cell exhibits at least one characteristic of the target cell type. Typically, the cell exhibits one, two, three, four, five, six, seven, eight, or more characteristics of the target cell type. For example, if the target cell is a keratinocyte cell, the cell is identified or determined to be a keratinocyte-like cell if upregulation of any one or more keratinocyte markers and / or changes in cell morphology are detectable. Preferably, the keratinocyte markers include keratin 1, keratin 14, and involucrin, and the cell morphology has a cobblestone-like appearance. In any embodiment of the present invention, the characteristics of the target cell can be determined by analyzing cell morphology, gene expression profile, activity assay, protein expression profile, surface marker profile, or differentiation potential. Examples of characteristics or markers include those described herein and those known to those skilled in the art. Other examples of relevant markers include, for example, for the conversion of keratinocytes to hematopoietic stem cells (HSCs): CD45 (pan-hematopoietic marker), CD19 / 20 (B cell marker), CD14 / 15 (bone marrow), CD34 (progenitor cell / SC marker), CD90 (SC), and alpha-integrin (keratinocyte marker not expressed by HSCs); for human embryonic stem cells to hematopoietic stem cells: Runx1 (GFP), CD45 (pan-hematopoietic marker), CD19 / 20 (B cell marker), CD14 / 15 (bone marrow), CD34 (progenitor cell / SC marker), CD90 (SC), and alpha-integrin (keratinocyte marker not expressed by HSCs). 5 (bone marrow), CD34 (progenitor cell / SC marker), CD90 (SC), Tra-1-160 (ESC marker not expressed in HSC); for rejuvenating aged or adult HSCs: comparison of the transcriptional signatures of young and aged human HSCs (e.g., using RNA-seq), and functional characterization of "rejuvenated HSCs" by transplanting rejuvenated cells into animals and then evaluating them 1, 3, and 6 months later to determine bone marrow bias (loss of bone marrow bias indicates "rejuvenated" HSCs). Examples of markers for many of the conversions described herein are shown in Table 1 below.

[0194] Table 1

[0195] Table 2-1

[0196] Table 2-2

[0197] Table 2-3

[0198] Table 2-4

[0199] Table 2-5

[0200] Table 2-6

[0201] Table 2-7

[0202] Table 2-8

[0203] In certain embodiments, one, two, three, four, or five or more of the factors (ligands) listed in Table 2 are added to or used in contact with the cells of interest to maintain the cells of interest in vitro. In other embodiments, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen or more, twenty or more, twenty-five or more, thirty or more, or thirty-five or more of the factors (ligands) listed in Table 2 are added to or used in contact with the cells of interest.

[0204] The present invention also provides methods for converting source cells into target cells. Methods for identifying factors for use in such conversion are further described herein. In addition to the embodiments outlined elsewhere in this document, in certain embodiments, the source cells are embryonic stem cells, and the factors for converting embryonic stem cells into target cells are as provided in the table below:

[0205] [Table 3-1]

[0206] [Table 3-2]

[0207] [Table 3-3]

[0208] [Table 3-4]

[0209] [Table 3-5]

[0210] In certain embodiments, one, two, three, four, or five or more of the factors listed in Table 3 are added to the source cells to convert them to target cells in vitro. In an alternative embodiment, the cells are contacted with an agent to increase expression of one or more of the above factors, e.g., as further described herein. In other embodiments, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more, 20 or more, 25 or more, 30 or more, or 35 or more of the factors (ligands) listed in Table 3 are added to or used to contact the source cells for conversion into target cells.

[0211] Those of skill in the art will be familiar with each of the ligands / factors set forth above in Tables 2 and 3. Some particularly preferred examples are provided further below. As used herein, FN1 refers to the protein fibronectin, a high molecular weight (approximately 440 kDa) glycoprotein of the extracellular matrix that binds to transmembrane receptor proteins called integrins. FN1 is also known as CIG, ED-B, FINC, FN, FNZ, GFND, GFND2, LETS, MSF, fibronectin 1, and SMDCF. An exemplary amino acid sequence of fibronectin is provided in Uniprot accession P02751 and is encoded by the nucleic acid sequence exemplified by accession ENSG00000115414.

[0212] As used herein, COL4A1 refers to collagen alpha-1(IV) chain (also known as HANAC, ICH, POREN1, Arresten, BSVD, RATOR, collagen type IV alpha 1, collagen type IV alpha 1 chain, BSVD1). An exemplary amino acid sequence of COL4A1 is provided in Uniprot accession P02462 and is encoded by the nucleic acid sequence exemplified by accession ENSG00000187498.

[0213] As used herein, COL1A2 refers to collagen alpha-2(I) chain (also known as OI4, collagen type I alpha 2, collagen type I alpha 2 chain, EDSCV, EDSARTH2). An exemplary amino acid sequence of COL1A2 is provided in Uniprot accession P08123 and is encoded by the nucleic acid sequence exemplified by accession ENSG00000164692.

[0214] As used herein, EDIL3 refers to the "EGF-like repeat and discoidin domain 3" protein encoded by the EDIL3 gene in humans. The protein encoded by this gene is an integrin ligand. It plays an important role in mediating angiogenesis and may be important in the remodeling and development of the vascular wall. It also influences endothelial cell behavior. An exemplary amino acid sequence of EDIL3 is provided in Uniprot accession O43854 and is encoded by the nucleic acid sequence exemplified by accession ENSG00000164176.

[0215] As used herein, ADAM12 refers to the disintegrin and metalloproteinase domain-containing protein 12 (formerly known as meltrin), an enzyme encoded by the ADAM12 gene in humans. ADAM12 has two splice variants: the long form ADAM12-L has a transmembrane domain, and the short variant ADAM12-S is soluble and lacks the transmembrane and cytoplasmic domains. ADAM12 is also known as ADAM12-OT1, CAR10, MCMP, MCMPMltna, MLTN, MLTNA, and ADAM metallopeptidase domain 12. An exemplary amino acid sequence of ADAM12 is provided in Uniprot accession O43184 and is encoded by the nucleic acid sequence exemplified by accession ENSG00000148848.

[0216] As used herein, WNT5A (also known as hWnt family member 5A) is a protein encoded by the WNT5A gene in humans. An exemplary amino acid sequence of Wnt5a is provided in Uniprot accession number P41221 and is encoded by the nucleic acid sequence exemplified by accession number ENSG00000114251.

[0217] As used herein, LAMB1 refers to the gene encoding laminin subunit beta-1. This protein is also known as CLM, LIS5, and laminin beta 1. Laminins, a family of extracellular matrix glycoproteins, are the major non-collagenous components of basement membranes. They are involved in a wide variety of biological processes, including cell adhesion, differentiation, migration, signal transduction, neurite outgrowth, and metastasis. Laminins are composed of three non-identical chains, laminin alpha, beta, and gamma (formerly laminin A, B1, and B2, respectively), which form a cross structure consisting of three short arms, each formed by a different chain, and a long arm composed entirely of three chains. An exemplary amino acid sequence of LAMB1 is provided in Uniprot accession P07942 and is encoded by the nucleic acid sequence exemplified by accession ENSG00000091136.

[0218] As used herein, COL3A1 refers to type III collagen, a protein composed of a homotrimer, or three identical peptide chains (monomers), each of which is called the alpha 1 chain of type III collagen. Formally, the monomer is called collagen type III, alpha-1 chain, and in humans, is encoded by the COL3A1 gene. Type III collagen is a type of fibrous collagen with a long, inflexible triple-helical domain. COL3A1 is also referred to by the identifiers EDS4A, collagen type III alpha 1, collagen type III alpha 1 chain, EDSVASC, and PMGEDSV. An exemplary amino acid sequence of COL3A1 is provided in Uniprot accession P02461 and is encoded by the nucleic acid sequence exemplified by accession ENSG00000168542.

[0219] As used herein, TFPI refers to tissue factor pathway inhibitor, a single-chain polypeptide that can reversibly inhibit factor Xa (Xa). Although Xa is inhibited, the Xa-TFPI complex can then inhibit the FVIIa-tissue factor complex. TFPI is also referred to by identifiers as EPI, LACI, TFI, and TFPI1. An exemplary amino acid sequence of TFPI is provided in Uniprot accession P10646 and is encoded by the nucleic acid sequence exemplified by accession ENSG00000003436.

[0220] As used herein, FGF7 (also known as HBGF-7 or KGF) refers to keratinocyte growth factor, a protein encoded by the FGF7 gene in humans. Members of the FGF family have a wide range of mitogenic and cell survival activities and are involved in various biological processes, including embryonic development, cell growth, morphogenesis, tissue repair, tumor growth, and invasion. This protein is a potent epithelial cell-specific growth factor, and its mitogenic activity is primarily expressed in keratinocytes, but not in fibroblasts or endothelial cells. An exemplary amino acid sequence of FGF7 is provided in Uniprot accession P21781 and is encoded by the nucleic acid sequence exemplified by accession ENSG00000140285.

[0221] As used herein, FGF10 refers to fibroblast growth factor 10, a protein encoded by the FGF10 gene in humans. Members of the FGF family have broad mitogenic and cell survival activities and are involved in embryonic development, cell growth, morphogenesis, and It is involved in various biological processes, including tissue repair, tumor growth, and invasion. Fibroblast growth factor 10 is a paracrine signaling molecule first found in limb bud and organogenesis. FGF10 signaling is required for epithelial branching. Therefore, all branching morphogenetic organs, such as the lung, skin, ear, and salivary gland, require constant expression of FGF10. This protein exhibits mitogenic activity for keratinizing epidermal cells but has essentially no activity on fibroblasts, similar to the biological activity of FGF7. An exemplary amino acid sequence of FGF10 is provided in Uniprot accession O15520 and is encoded by the nucleic acid sequence exemplified by accession CCDS3950.1.

[0222] As used herein, APOE refers to apolipoprotein E (also known as AD2, APO-E, LDLCQ5, LPG, ApoE4), a protein involved in the metabolism of fat in the body.The exemplary amino acid sequence of APOE is provided in Uniprot accession number P02649 and is encoded by the nucleic acid sequence exemplified by accession number ENSG00000130203.

[0223] As used herein, C3 refers to complement component 3, a protein that plays a central role in the complement system and contributes to innate immunity. This protein may also be referred to by the identifiers AHUS5, ARMD9, ASP, C3a, C3b, CPAMD1, HEL-S-62p, and complement component 3. An exemplary amino acid sequence of C3 is provided in Uniprot accession P01024 and is encoded by the nucleic acid sequence exemplified by accession ENSG00000125730.

[0224] As used herein, SERPINE1 refers to plasminogen activator inhibitor-1 (PAI-1), also known as endothelial plasminogen activator inhibitor or serpin E1, a protein encoded by the SERPINE1 gene in humans. PAI-1 is a serine protease inhibitor (serpin) that functions as a major inhibitor of tissue plasminogen activator (tPA) and urokinase (uPA), plasminogen activators and thus fibrinolysis (the physiological breakdown of blood clots). It is a serine protease inhibitor (serpin) protein (SERPINE1). An exemplary amino acid sequence of SERPINE1 is provided in Uniprot accession P05121 and is encoded by the nucleic acid sequence exemplified by accession ENSG00000106366.

[0225] As used herein, COL6A3 refers to the collagen alpha-3 (VI) chain, a protein encoded by the COL6A3 gene in humans. This gene encodes the alpha-3 chain, one of three alpha chains of type VI collagen, a beaded filament collagen found in most connective tissues. The alpha-3 chain of type VI collagen is much larger than the alpha-1 and alpha-2 chains. This size difference is largely due to the increased number of subdomains, similar to the von Willebrand factor type A domain, found in the amino-terminal globular domain of all alpha chains. In addition to the full-length transcript, four transcript variants have been identified that encode proteins with N-terminal globular domains of various sizes. An exemplary amino acid sequence of COL6A3 is provided in Uniprot accession P12111 and is encoded by the nucleic acid sequence exemplified by accession ENSG00000163359.

[0226] As used herein, CXCL12 refers to stromal cell-derived factor 1 (SDF1), also known as C-X-C motif chemokine 12, IRH, PBSF, SCYB12, SDF1, TLSF, TPAR1, and C-X-C motif chemokine ligand 12. The gene encoding CXCL12 is divided into seven parts by alternative splicing. An exemplary amino acid sequence of CXCL12 is provided in Uniprot accession P48061 and is encoded by the nucleic acid sequence exemplified by accession ENSG00000107562.

[0227] As used herein, TGFB2 and TGFB3 refer to transforming growth factor beta 2 and transforming growth factor beta 3, respectively. Both proteins are cytokines involved in cell differentiation, embryogenesis and development. An exemplary accession number for TGFB2 is P61812. An exemplary accession number for TGFB3 is P10600.

[0228] As used herein, GDF5 refers to growth / differentiation factor 5, and is encoded by GDF5 gene.The protein encoded by this gene is closely related to bone morphogenetic protein (BMP) family and is a member of TGF-beta superfamily.The exemplary accession number for GDF5 is P43026.

[0229] As used herein, NID1 refers to nidogen-1 (NID-1), formerly known as entactin, which is a protein encoded by the NID1 gene in humans. Both nidogen-1 and nidogen-2 are essential components of the basement membrane, along with other components such as type IV collagen, proteoglycans (heparan sulfate and glycosaminoglycans), laminin, and fibronectin. An exemplary accession number for NID1 is P14543.

[0230] As used herein, THBS1 refers to thrombospondin 1, abbreviated as THBS1, which is a protein encoded by the THBS1 gene in humans. Thrombospondin 1 is a subunit of a disulfide-bonded homotrimeric protein. This protein is an adhesive glycoprotein that mediates cell-cell and cell-matrix interactions. This protein can bind to fibrinogen, fibronectin, laminin, collagen types V and VII, and integrin alpha-V / beta-1. An exemplary accession number for THBS1 is P07996.

[0231] As used herein, BMP4 and BMP6 refer to bone morphogenetic protein 4 and bone morphogenetic protein 6, respectively. The proteins encoded by these genes are members of the TGFβ superfamily. Bone morphogenetic proteins are known for their ability to induce bone and cartilage growth. BMP4 is highly conserved evolutionarily. BMP4 is found in the ventral marginal zone and in early embryonic development in the eye, heart blood, and otic vesicle. BMP6 can induce all osteogenic markers in mesenchymal stem cells. An exemplary accession number for BMP6 is P12644. An exemplary accession number for BMP6 is P22004.

[0232] As used herein, CTGF, also known as CCN2 or connective tissue growth factor, is a matricellular protein of the CCN family of extracellular matrix-associated heparin-binding proteins (see also CCN intercellular signaling proteins). CTGF plays an important role in many biological processes, including cell adhesion, migration, proliferation, angiogenesis, skeletal development, and tissue wound repair, and is critically involved in some forms of fibrosis and cancer. An exemplary accession number for CTGF is P29279.

[0233] As used herein, PDGFB refers to platelet-derived growth factor subunit B, a protein that in humans is encoded by the PDGFB gene. The protein encoded by is a member of the platelet-derived growth factor family. The four members of this family are mitogens for cells of mesenchymal origin and are characterized by an eight-cysteine ​​motif. This gene product can exist as a homodimer (PDGF-BB) or as a heterodimer with the platelet-derived growth factor alpha (PDGFA) polypeptide (PDGF-AB), with the dimers linked by a disulfide bond. An exemplary accession number for PDGFB is P01127.

[0234] It is understood that in a preferred example, relevant factors (ligands) for maintaining the cells of interest can be directly provided to the cells so that the cells are "contacted" with the relevant factors. Furthermore, as used herein, the terms "contact," "contacted," "contacting," "treating," "treating," or "introducing" can be used interchangeably and refer to subjecting cells to any type of process or condition, such as transfection of exogenous DNA or introduction of an agent into the cells.

[0235] For example, the culture medium can be supplemented directly with the relevant factors. In an alternative embodiment, nucleic acid molecules encoding the relevant factors (or vectors encoding one or more of the relevant factors) can be transfected into the cells, whereby the factors are then expressed by the cells.

[0236] Additionally, cells can be treated in such a way as to increase the amount of the relevant factor in the cells, again this can be by direct supplementation of the cell culture medium, or by recombinant methods to increase expression of an endogenous gene in the cells, or by providing an exogenous source of nucleic acid encoding the relevant factor and expressing the nucleic acid, thereby increasing the amount of the factor in the cells.

[0237] Similarly, with respect to methods of converting cells (including those for transdifferentiation and directed differentiation), the converting factors may be provided directly in the culture or, alternatively, may be expressed in the source cells to facilitate conversion to the target cells.

[0238] It is understood that factors that are proteins, including transcription factors and other regulatory factors, can be provided to cells according to the methods described herein. Alternatively, variants of the factors can be provided to achieve the same result.

[0239] The term "variant" when referring to a polypeptide refers to a polypeptide that is at least 70%, 80%, 85%, 90%, 95%, 98%, or 99% identical to the full-length polypeptide. The present invention contemplates the use of variants of the factors described herein. A variant may be a fragment of a full-length polypeptide or a naturally occurring splice variant. A variant may be a polypeptide that is at least 70%, 80%, 85%, 90%, 95%, 98%, or 99% identical to a fragment of a polypeptide, where the fragment is at least 50%, 60%, 70%, 80%, 85%, 90%, 95%, 98%, or 99% identical to the full-length wild-type polypeptide or a domain thereof, so long as the fragment retains the desired functional activity, such as the ability to promote conversion of a source cell type to a target cell type. In some embodiments, the domain begins at any amino acid position in the sequence and extends to the C-terminus, and is at least 100, 200, 300, or 400 amino acids in length. Variations known in the art that eliminate or substantially reduce the activity of a protein are preferably avoided. In some embodiments, variants lack the N- and / or C-terminal portions of the full-length polypeptide, e.g., up to 10, 20, or 50 amino acids from either end. In some embodiments, the polypeptide has the sequence of a mature (full-length) polypeptide, which is amplified during normal intracellular proteolytic processing (e.g., co-translationally or post-translationally). "Functional" refers to a polypeptide having one or more portions, such as a signal peptide, removed (during post-translational processing). In some embodiments where the protein is produced other than by purification from naturally-expressing cells, the protein is a chimeric polypeptide, meaning that the protein contains portions derived from two or more different species. In some embodiments where the protein is produced other than by purification from naturally-expressing cells, the protein is a derivative, meaning that the protein contains additional sequences not related to the protein, so long as these sequences do not substantially reduce the biological activity of the protein. Those skilled in the art will know or can easily confirm whether a particular polypeptide variant, fragment, or derivative is functional using assays known in the art. For example, assays such as those disclosed herein in the Examples can be used to evaluate the ability of a transcription factor variant to convert a source cell into a target cell type. Other convenient assays include measuring the ability to activate transcription of a reporter construct containing a transcription factor binding site operably linked to a nucleic acid sequence encoding a detectable marker, such as luciferase. In certain embodiments of the invention, a functional variant or fragment has at least 50%, 60%, 70%, 80%, 90%, 95% or more of the activity of the full-length wild-type polypeptide.

[0240] The term "increasing the amount of" with respect to increasing the amount of a factor refers to increasing the amount of the factor in a cell of interest (e.g., an astrocyte, a cardiomyocyte, a smooth muscle cell, an endothelial cell, or a source cell for conversion into a target cell). In some embodiments, the amount of a factor is "increased" in a cell of interest (e.g., a cell into which an expression cassette directing the expression of a polynucleotide encoding one or more transcription factors has been introduced) if the amount of the factor is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or more relative to a control (e.g., a cell into which none of the expression cassettes has been introduced). However, any method of increasing the amount of a transcription factor is contemplated, including any method that increases the amount, rate, or efficiency of transcription, translation, stability, or activity of the transcription factor (or the pre-mRNA or mRNA encoding it). Additionally, downregulation or interference with negative regulators of transcription expression, increasing the efficiency of existing transcription (e.g., SINEUP) are also contemplated.

[0241] When used in reference to a protein, gene, nucleic acid, or polynucleotide in a cell or organism, the term "exogenous" refers to a protein, gene, nucleic acid, or polynucleotide that has been introduced into the cell or organism by artificial or natural means, or when used in reference to a cell, refers to a cell that has been isolated and subsequently introduced into another cell or organism by artificial or natural means. An exogenous nucleic acid may be derived from a different organism or cell, or may be a copy of one or more additional nucleic acids naturally present in the organism or cell. An exogenous cell may be derived from a different organism or from the same organism. As a non-limiting example, an exogenous nucleic acid is at a different chromosomal location than that of a native cell, or is otherwise flanked by different nucleic acid sequences than those found in nature. An exogenous nucleic acid may also be extrachromosomal, such as an episomal vector.

[0242] The methods of the present disclosure may be "miniaturized" in assay systems through any acceptable method of miniaturization, including, but not limited to, multi-well plates, microchips, or slides, such as 24, 48, 96, or 384 wells per plate. The assay may be reduced in size so that it can be performed on a microchip support, which advantageously involves fewer reagents and other materials. Any miniaturization of the process that leads to high-throughput screening is within the scope of the present invention.

[0243] In any of the methods of the invention, the target cells are cultured in the same mammal from which the source cells were obtained. In other words, the source cells used in the methods of the present invention may be autologous, i.e., obtained from the same individual to which the target cells are administered. Alternatively, the target cells may be allogeneically transferred into another individual. Preferably, the cells are autologous to the subject in methods of treating or preventing a medical condition in an individual.

[0244] As referred to herein, the term "cell culture medium" (also referred to herein as "culture medium" or "culture medium") is a medium for culturing cells, containing nutrients that maintain cell viability and support growth. Cell culture media may contain any of the following in appropriate combinations: salts, buffers, amino acids, glucose or other sugars, antibiotics, serum or serum replacements, and other components such as peptide growth factors. Cell culture media commonly used for particular cell types are known to those of skill in the art. Exemplary cell culture media for use in the methods of the present invention are shown in Table 4.

[0245] The term "expression," as applicable, refers to the cellular processes involved in producing RNA and proteins and, if desired, secreting the proteins, including, but not limited to, transcription, translation, folding, modification, and processing.

[0246] As used herein, the terms "isolated" or "partially purified" refer to a nucleic acid or polypeptide that has been separated from at least one other component (e.g., nucleic acid or polypeptide) that is present with the nucleic acid or polypeptide as found in the natural source and / or that would be present with the nucleic acid or polypeptide when expressed or, in the case of a secreted polypeptide, secreted by a cell. Chemically synthesized nucleic acids or polypeptides or those synthesized using in vitro transcription / translation are considered "isolated."

[0247] The term "vector" refers to a carrier DNA molecule into which a DNA sequence can be inserted for introduction into a host or source cell. Preferred vectors are capable of autonomous replication and / or expression of linked nucleic acids. Vectors capable of directing the expression of genes to which they are operably linked are referred to herein as "expression vectors." Thus, an "expression vector" is a specialized vector containing the necessary regulatory regions required for the expression of a gene of interest in a host cell. In some embodiments, the gene of interest is operably linked to another sequence in the vector. The vector may be a viral vector or a non-viral vector. When a viral vector is used, the viral vector is preferably replication-deficient, which can be achieved, for example, by removing all viral nucleic acid coding for replication. Replication-deficient viral vectors still retain their infectious properties and enter cells in a manner similar to replicating adenoviral vectors, but once internalized by a cell, they do not replicate or propagate. Vectors also encompass liposomes and nanoparticles, as well as other means of delivering DNA molecules to cells.

[0248] The term "operably linked" means that regulatory sequences required for expression of a coding sequence are positioned in a DNA molecule in the appropriate position relative to the coding sequence to achieve expression of the coding sequence. This same definition is sometimes applied to the placement of coding sequences and transcription control elements (e.g., promoters, enhancers, and termination elements) in an expression vector. The term "operably linked" includes having an appropriate initiation signal (e.g., ATG) immediately preceding the polynucleotide sequence to be expressed, and maintaining the correct reading frame to allow expression of the polynucleotide sequence under the control of the expression control sequences and production of the desired polypeptide encoded by the polynucleotide sequence.

[0249] The term "viral vector" refers to the use of a virus or virus-associated vector as a carrier of a nucleic acid construct into a cell. The construct may be incorporated into and packaged within the genome of a non-replicating defective virus, such as adenovirus, adeno-associated virus (AAV), or herpes simplex virus (HSV), including retrovirus and lentivirus vectors, for infection or transduction into cells. The vector may or may not be incorporated into the genome of the cell. The construct may also include viral sequences for transfection, if desired. Alternatively, the construct may be incorporated into an episomal replicable vector, such as EPV and EBV vectors.

[0250] As used herein, the term "adenovirus" refers to viruses of the Adenoviridae family. Adenoviruses are medium-sized (90-100 nm), non-enveloped (naked), icosahedral viruses composed of a nucleocapsid and a double-stranded linear DNA genome.

[0251] As used herein, the term "non-integrating viral vector" refers to a viral vector that does not integrate into the host genome, and the expression of genes delivered by the viral vector is transient. Because there is little or no integration into the host genome, non-integrating viral vectors have the advantage of not producing DNA mutations by inserting at random points in the genome. For example, non-integrating viral vectors remain extrachromosomal and do not insert their genes into the host genome, potentially disrupting the expression of endogenous genes. Non-integrating viral vectors may include, but are not limited to, the following: adenovirus, alphavirus, picornavirus, and vaccinia virus. Although any of these viral vectors may, in some rare circumstances, integrate viral nucleic acid into the genome of a host cell, as the term is used herein, they are "non-integrating" viral vectors. Importantly, the viral vectors used in the methods described herein do not integrate their nucleic acid into the genome of a host cell, generally or as a major part of their life cycle, under the conditions used.

[0252] The vectors described herein can be constructed and engineered using methods commonly known in the scientific literature to increase safety for use in therapy and, if desired, to include selectable and enrichment markers and optimize expression of the nucleotide sequences they contain. The vector must contain structural components that allow the vector to self-replicate in the source cell type. For example, the well-known Epstein-Barr oriP / nuclear antigen-1 (EBNA-I) combination (e.g., Lindner, S.E., and B. Sugden, "The plasmid replicon of Epstein-Barr virus: mechanistic insights into efficient, licensed, extrachromosomal replication in human," incorporated by reference in its entirety as if set forth herein. (See, e.g., Plasmid 58:1 (2007)) are sufficient to support vector self-replication, and other combinations known to function in mammalian, particularly primate, cells may also be utilized. Standard techniques for the construction of expression vectors suitable for use in the present invention are well known to those skilled in the art and can be found in publications such as Sambrook J et al., "Molecular cloning: a laboratory manual," (3rd ed., Cold Spring Harbor Press, Cold Spring Harbor, NY 2001), which is incorporated herein by reference as if set forth in its entirety.

[0253] In the methods of the present invention, genetic material encoding the relevant factors required for transformation is delivered into the source cells via one or more cell transformation vectors. Each factor may be introduced into the source cells as a polynucleotide transgene encoding the factor operably linked to a heterologous promoter capable of driving expression of the polynucleotide in the source cells.

[0254] Suitable cell transformation vectors are any of those described herein, including episomal vectors such as plasmids that do not encode all or part of a viral genome sufficient to produce infectious or replication-competent virus, although the vector may contain structural elements obtained from one or more viruses. One or more cell transformation vectors may be introduced into a single source cell. One or more transgenes may be provided on a single cell transformation vector. A single strong constitutive transcription promoter may provide transcriptional control for multiple transgenes, which may be provided as expression cassettes. Separate expression cassettes on a vector may be under the transcriptional control of separate strong constitutive promoters, which may be copies of the same promoter or different promoters. A variety of heterologous promoters are known in the art and may be used depending on factors such as the desired expression level of transcription factors. As exemplified below, it may be advantageous to use different promoters with different strengths to control the transcription of separate expression cassettes in the source cell. Another consideration in selecting a transcription promoter is the rate of promoter silencing. Those skilled in the art will understand that it may be advantageous to reduce the expression of one or more transgenes or transgene expression cassettes after the gene product has completed or substantially completed its role in the cell transformation method.Exemplary promoters are the human EF1α elongation factor promoter, the CMV cytomegalovirus immediate early promoter, and the CAG chicken albumin promoter, as well as corresponding homologous promoters from other species.In human somatic cells, both EF1α and CMV are strong promoters, but the former is silenced more efficiently than the latter, so that the expression of transgenes under the control of the CMV promoter is stopped earlier than the expression of transgenes under the control of the EF1α promoter.Transcription factors can be expressed in source cells at relative ratios that can be varied to modulate cell transformation efficiency.Preferably, when multiple transgenes are encoded on a single transcript, an internal ribosome entry site is provided upstream of the transgene, distal to the transcription promoter. The relative ratios of factors may vary depending on the factors being delivered, but the skilled artisan armed with this disclosure will be able to determine the optimal ratio of factors.

[0255] Those skilled in the art will appreciate that it is advantageously more efficient to introduce all factors via a single vector rather than via multiple vectors. Thus, the present invention contemplates the use of a single vector (or a small number of vectors), which optionally encodes multiple factors necessary for transformation or cell maintenance.

[0256] After introducing the transformation vector, and while the source cell is being transformed, the vector can remain present in the target cell, during which time the introduced transgene is transcribed and translated. Transgene expression may be advantageously downregulated or silenced in cells that have been transformed into the target cell type. The cell transformation vector may remain extrachromosomal. With extremely low efficiency, the vector may be integrated into the genome of the cell. The following examples are intended to illustrate the present invention and are not intended to limit it in any way.

[0257] Suitable methods for nucleic acid delivery for transformation of cells, tissues, or organisms for use with the present invention are intended to include virtually any method by which a nucleic acid (e.g., DNA) can be introduced into a cell, tissue, or organism, as described herein or known to those of skill in the art (e.g., Stadtfeld and Hochedlinger, Nature Methods 6(5):329-330 (2009); Yusa et al., Nat. Methods 6:363-369 (2009); Woltjen et al., Nature 458, 766-770 (April 9, 2009)). Such methods include, but are not limited to, direct delivery of DNA, such as by ex vivo transfection (Wilson et al., Science 244:1344-1346, 1989; Nabel and Baltimore, Nature 326:711-713, 1987), optionally using lipid-based transfection reagents such as Fugene6 (Roche) or Lipofectamine (Invitrogen), or by injection (U.S. Pat. Nos. 5,994,624, 5,981,274, 5,945,100, 5,780,448, 5,736,524, 5,702,932, 5,656,610, 5,589,466, and 5,580,859, each of which is incorporated herein by reference), including microinjection (Harland and Weintraub, J. Cell 1999, 11, 141-142, 1999, each of which is incorporated herein by reference). Biol., 101:1094-1099, 1985; U.S. Pat. No. 5,789,215); by electroporation (U.S. Pat. No. 5,384,253; Tur-Kaspa et al., Mol. Cell Biol., 6:716-718, 1986; Potter et al., Proc. Nat'l Acad. Sci. USA, 81:7161-7165, 1984, which are incorporated herein by reference); by calcium phosphate precipitation (Graham and Van Der Eb, Virology, 52:456-467, 1973; Chen and Okayama, Mol. Cell Biol., 7(8):2745-2752, 1987; Rippe et al., Mol. Cell Biol., 10:689-695, 1990); by using DEAE-dextran followed by polyethylene glycol (Gopal, Mol. Cell Biol., 5:1188-1190, 1985); by direct ultrasonic loading (Fechheimer et al., Proc. Nat'l Acad. Sci.USA, 84:8463-8467, 1987); or liposome-mediated transfection (Nicolau and Sene, Biochim. Biophys. Acta, 721:185-190, 1982; Fraley et al., Proc. Nat'l Acad. Sci. USA, 76:3348-3352, 1979; Nicolau et al., Methods Enzymol., 149:157-176, 1987; Wong et al., Gene, 10:87-94, 1980; Kaneda et al., Science, 243:375-378, 1989; Kato et al., J. Biol. Chem., 266:3361-3364, 1991), and receptor-mediated transfection (Wu and Wu, Biochemistry, 27:887-892, 1988; Wu and Wu, J. Biol. Chem., 262:4429-4432, 1987); as well as by any combination of such methods, each of which is incorporated herein by reference.

[0258] Several polypeptides capable of mediating the introduction of associated molecules into cells have been previously described and are applicable to the present invention, see, e.g., Langel (2002) Cell Penetrating Peptides: Processes and Applications, CRC Press, Pharmacology and Toxicology Series. Examples of polypeptide sequences that enhance transport across membranes include, but are not limited to, the Drosophila homeoprotein Antennapedia transcription protein (AntHD) (Joliot et al., New Biol. 3:1121-34, 1991; Joliot et al., Proc. Natl. Acad. Sci. USA, 88:1864-8, 1991; Le Roux et al., Proc. Natl. Acad. Sci. USA, 90:9120-4, 1993), herpes simplex virus structural protein VP22 (Elliott and O'Hare, Cell 88:223-33, 1997); the HIV-1 transcriptional activator TAT protein (Green and Loewenstein, Cell 55:1179-1188, 1988; Frankel and Pabo, Cell 55:1 289-1193, 1988); Kaposi's FGF signal sequence (kFGF ); protein transduction domain-4 (PTD4); penetratin, M918, transportan-10; nuclear localization sequences, PEP-I peptides; amphipathic peptides (e.g., MPG peptides); delivery-enhanced transporters such as those described in U.S. Pat. No. 6,730,293 (including, but not limited to, peptide sequences containing at least 5-25 or more consecutive arginines, or 5-25 or more arginines in a contiguous set of 30, 40, or 50 amino acids; including, but not limited to, peptides having sufficient, e.g., at least five guanidino or amidino moieties); and the commercially available Penetratin™ 1 peptide and the Diatos Peptide Vector ("DPV") on the Vectocell® platform, available from Daitos SA, Paris, France. See also WO / 2005 / 084158 and WO / 2007 / 123667, and the additional transporters described therein. Not only are these proteins able to cross cell membranes, but the attachment of other proteins, such as the transcription factors described herein, is sufficient to stimulate cellular uptake of these complexes.

[0259] [Table 4-1]

[0260] [Table 4-2]

[0261] For differentiation methods, the relevant cell type media described above can be used to facilitate differentiation. For cell maintenance, it is possible to culture cells using only the factors identified herein for cell maintenance. For example: Normal human astrocytes (NHA) (Lonza Bioscience) were thawed and maintained for one passage in Matrigel. (Matrigel is the trademarked name for a gelatinous protein mixture secreted by Engelbreth-Holm-Swarm (EHS) mouse sarcoma cells, manufactured and sold by Corning Life Sciences and BD Biosciences. Trevigen, Inc. sells its own version under the trademark Cultrex BME. Matrigel resembles the complex extracellular environment found in many tissues and is used by cell biologists as a substrate (basement membrane matrix) for culturing cells.) The cells were then passaged using Accutase and seeded in the following conditions: no substrate, Matrigel, FN1 (fibronectin), COL4A1 (collagen 4), LAMB1 (LN221), ADAM12, COL1A2 (collagen I), EDIL3, and all factors (the combination of the above substrates and growth factors without Matrigel).

[0262] - Thawing and maintenance medium for cell culture of primary human cardiomyocytes can be purchased from PromoCell (Heidelberg, Germany). Cells are cultured in myocyte growth medium (PromoCell, C-22011) supplemented with 2% fetal bovine serum, EGF (5 ng / ml), FGF2 (2 ng / ml), EGF (0.5 ng / ml) and insulin (0.5 μg / ml). Cells were then passaged using 0.04% trypsin / 0.03% EDTA (Promocell) and seeded under the following conditions: no substrate, Geltrex (Life Technologies), FN1 (fibronectin) (Sigma Aldrich), COL1A2 (collagen I) (Sigma Aldrich), TFPI (tissue factor pathway inhibitor) (Prospec-Tany TechnoGene Ltd), ApoE (apolipoprotein E) (Novus Biologicals), FGF7 (fibroblast growth factor-7) (Stemcell Technologies), and all ligands (the above substrate and ligand combinations without Geltrex). ReN VM cells (Millipore) were cultured in 100% PBS containing 1% Fibroblast Growth Factor (EFG) and 1% B27 supplement (Gibco), 1% heparin (Stemcell Technologies), 1% penicillin and streptomycin (Gibco), 20 ng / ml EFG (Pep Cells were maintained in DMEM / F12 (Gibco) containing 20 ng / ml of erythrocyte growth factor-2 (FGF2) (Miltenyi Biotec). Cells were passaged using Accutase (Stem Cell Technologies) and seeded onto Matrigel (Corning)-coated flasks. To differentiate into astrocytes, cultures were transferred to astrocyte medium (DMEM high glucose, N2, 10% FBS; Gibco) for 2 weeks. Cells were isolated using the surface marker CD44 (103028, Biolegend) and replated in each condition: no substrate, Matrigel, FN1 (fibronectin) (Sigma Aldrich), COL4A1 (collagen IV) (Sigma Aldrich), LAMB1 (laminin 221) (Biolamina), ADAM12 (Sigma Aldrich), COL1A2 (collagen I) (Sigma Aldrich), EDIL3 (R&D Systems), and all ligands (combinations of the above substrates and secreted ligands without Matrigel). Cells were analyzed at passage 2 (days 3 and 6 of passage 2).

[0263] Thawing and maintenance media for primary pulmonary artery smooth muscle cells (HPASMC) and human aortic endothelial cell (HAoEC) cultures can be purchased from Gibco (C0095C, Life Technologies, USA) and PromoCell (C-12271, Heidelberg, Germany), respectively. Smooth muscle cells were cultured in Medium 231 (Life Technologies, USA) supplemented with fetal bovine serum (4.9% v / v), human basic fibroblast growth factor (2 ng / ml), human epidermal growth factor (0.5 ng / ml), heparin (5 ng / ml), recombinant human insulin-like growth factor-I (0.01 μg / ml), and BSA (0.2 μg / ml). Cells were then passaged using TrypLE Express (Life Technologies) and seeded under the following conditions: no substrate, Geltrex (Life Technologies), NID1 (nidogen-1) (Merck), COL1A1 (collagen I) (Southern Biotech), COL4A1 (collagen 4) (Merck), COL6A1 (collagen 6) (Rockland), THBS1 (thrombospondin 1) (Merck), FGF10 (fibroblast growth factor-10) (Stemcell Technologies), FGF7 (fibroblast growth factor-7) (Stemcell Technologies), and all ligands (combinations of the above substrates and ligands without Geltrex).

[0264] - The endothelial cells are cultured in endothelial cell growth medium MV2 (PromoCell, C-22022) supplemented with a supplement mix containing fetal bovine serum (0.05 ml / ml), epidermal growth factor (recombinant human) (5 ng / ml), basic fibroblast growth factor (recombinant human) (10 ng / ml), insulin-like growth factor (Long R3 IGF) (20 ng / ml), vascular endothelial growth factor (recombinant human) (0.5 ng / ml), ascorbic acid (1 μg / ml) and hydrocortisone (0.2 μg / ml). Cells were then passaged using TrypLE Express (Life Technologies) and plated under the following conditions: no substrate, Geltrex (Life Technologies), FN1 (fibronectin) (Merck), THBS1 (thrombospondin 1) (Merck), CTGF (connective tissue growth factor) (Life Technologies), PDGFB (platelet-derived growth factor subunit B) (Stemcell Technologies), CYR61 (cysteine-rich angiogenic factor 61) (Novus), BMP4 (bone morphogenetic protein 4) (Life Technologies), BMP6 (bone morphogenetic protein 6) (Merck), and all ligands (combinations of the above substrates and ligands without Geltrex).

[0265] - H9 embryonic stem cells are maintained in flasks coated with vitronectin (Gibco) and cultured in Essential 8 medium (Gibco). For differentiation of astrocytes from H9 stem cells via neural progenitor cells, cells were cultured at 5 × 10 in neural induction medium containing DMEM / F12, B27 without vitamin A supplement (Gibco), N2 supplement (Gibco), 0.1% 2-mercaptoethanol (Gibco), 0.66% bovine serum albumin (Gibco), 1% sodium pyruvate (Gibco), 1% non-essential amino acids (Gibco), 1% penicillin and streptomycin, and 100 ng / ml LDN193189 (Tocris Bioscience). 4 cells / cm 2 Cells are seeded onto Matrigel-coated flasks at a density of 1000 μg / ml for 14 days. Cells are then selected for NCAM+ cells (with anti-pSA-NCAM antibody, 130-115-809, Miltenyi Biotec) and replated onto Matrigel-coated flasks for an additional 7 days. After one week of growth, cells are passaged and replated in astrocyte medium (Gibco) in the following conditions: Matrigel, FN1 (fibronectin) (Sigma Aldrich), COL4A1 (collagen 4) (Sigma Aldrich), LAMB1 (LN221) (Biolamina), Matrigel + ADAM12 (Sigma Aldrich), COL1A2 (collagen I) (Sigma Aldrich), Matrigel + EDIL3 (R&D Systems), and all factors (the above substrate and growth factor combinations without Matrigel). Cells are then analyzed on days 14 and 21.

[0266] For cardiac differentiation, cells were plated at 5 × 10 in Essential 8 medium (Gibco) on plates coated with Matrigel or COL1A2, depending on the condition. 4 cells / cm 2On day 0, the medium is replaced with RPMI 1640 (Gibco) supplemented with insulin-free B27 supplement (Gibco), 1% penicillin and streptomycin (Gibco), 213 μg / ml ascorbic acid (Sigma-Aldrich), and 0.66% bovine serum albumin (Gibco) for 5 days, with daily medium changes. On days 0-2, the cultures are also supplemented with 3 μM CHIR. On day 3, the CHIR is harvested. On days 4 and 5, the cultures are supplemented with 5 μM IWR. After 5 days, the culture medium was changed every 2 days to RPMI 1640 (Gibco) supplemented with B27 supplement (Gibco), 1% penicillin and streptomycin (Gibco), 213 μg / ml ascorbic acid (Sigma-Aldrich), and 0.66% bovine serum albumin (Gibco). During differentiation, cells were subjected to their respective conditions using differentiation medium: Matrigel (control), FGF7 (Peprotech, 10 ng / ml) (Miyashita et al., 2013), TGFB3 (Stemcell Technologies, 10 ng / ml) (Dahlin et al., 2014), TGFB2 (Stemcell Technologies, 10 ng / ml), GDF5 (Stemcell Technologies, 50 ng / ml), COL1A2 (collagen I) (Sigma-Aldrich), and all ligands (combinations of the substrates and secreted ligands mentioned above, including COL1A2-coated plates). Unless otherwise stated, substrates and ligands are applied according to the manufacturer's recommendations.

[0267] Computer Implementation Method The method embodiments described herein may be implemented in a computer processing system. Thus, in a further aspect, the present invention provides a computer processing system adapted to perform any one of the method embodiments disclosed herein.

[0268] FIG. 10 provides a block diagram of one type of computer processing system 600 for implementing embodiments and / or features of the methods described herein. System 600 is a general-purpose computer processing system. It is understood that FIG. 10 does not illustrate all functional or physical components of a computer processing system. For example, a power supply or power interface is not shown. However, system 600 may either include a power supply or be configured to connect to a power supply (e.g., It is also understood that the particular type of computer processing system will determine the appropriate hardware and architecture, and that alternative computer processing systems suitable for implementing aspects of the present invention may have additional, alternative, or fewer components than those shown.

[0269] The computing system 600 includes at least one processing unit 602. The processing unit 602 may be a single computing device (e.g., a central processing unit, a graphics processing unit, or other computing device) or may include multiple computing devices. In some examples, all processing is performed by the processing unit 602, while in other examples, processing may also be performed by a remote processing device accessible and usable by the system 600 (either in a shared or dedicated manner).

[0270] Via a communication bus 604, the processing unit 602 is in data communication with one or more machine-readable storage (memory) devices that store instructions and / or data for controlling the operation of the processing system 600. In this case, the system 600 comprises a system memory 606 (e.g., a BIOS), a volatile memory 608 (e.g., random access memory such as one or more DRAM modules), and a non-volatile memory 610 (e.g., one or more hard disks or solid-state drives). In some embodiments, any one or more of the data sources used in the methods described herein may be stored in non-volatile memory (e.g., a database storing ChIP-seq information, protein-protein interaction network data, gene sequence databases).

[0271] System 600 also includes one or more interfaces, generally indicated by 612, through which system 600 couples with various devices and / or networks. Generally speaking, the other devices may be integral to system 600 or may be separate. If a device is separate from system 600, the connection between the device and system 600 may be via wired or wireless hardware and communication protocols, and may be a direct or indirect (e.g., networked) connection.

[0272] Wired connections to other devices / networks may be via any suitable standard or proprietary hardware and connection protocol. For example, system 600 may be configured for wired connections to other devices / communications networks via one or more of the following: USB; FireWire; eSATA; Thunderbolt; Ethernet; OS / 2; Parallel; Serial; HDMI®; DVI; VGA; SCSI; AudioPort. Of course, other wired connections are possible.

[0273] Wireless connectivity with other devices / networks may similarly be via any suitable standard or proprietary hardware and communication protocol. For example, system 600 may be configured for wireless connectivity with other devices / communication networks using one or more of infrared; Bluetooth; WiFi; Near Field Communication (NFC); ​​Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Long Term Evolution (LTE), Wideband Code Division Multiple Access (W-CDMA), and Code Division Multiple Access (CDMA). Of course, other wireless connectivity is also possible.

[0274] Generally speaking, the devices to which the system 600 connects, whether by wired or wireless means, include one or more input devices that can input / receive data into / by the system 600 for processing by the processing unit 602; and one or more output devices capable of outputting data by system 600. Examples of devices are described below, although it is understood that not all computer processing systems will include all of the devices mentioned, and that devices in addition to and instead of those mentioned may well be used.

[0275] For example, system 600 may include or be connected to one or more input devices by which information / data is input to (received by) system 600. Such input devices may include a keyboard, mouse, trackpad, microphone, accelerometer, proximity sensor, GPS device, etc. System 600 may also include or be connected to one or more output devices controlled by system 600 to output information. Such output devices may include devices such as a CRT display, LCD display, LED display, plasma display, touchscreen display, speaker, vibration module, LED / other lights, and the like. System 600 may also include or be connected to devices that can operate as both input and output devices, such as memory devices (hard drives, solid state drives, disk drives, CompactFlash cards, SD cards, etc.) from which system 600 can read data and / or to which system 600 can write data, as well as a touchscreen display that can both display data (output) and receive touch signals (input).

[0276] System 600 may also connect to one or more communication networks (e.g., the Internet, a local area network, a wide area network, a personal hotspot, etc.) to communicate data to and receive data from network devices, which may themselves be other computer processing systems.

[0277] System 600 may be any suitable computing system, such as, by way of non-limiting example, a server computing system, a desktop computer, a laptop computer, a netbook computer, a tablet computing device, a mobile / smartphone, a personal digital assistant, a personal media player, a set-top box, a game console, or the like.

[0278] Typically, the system 600 includes at least a user input and output device 614 (such as a touchscreen display) and a communication interface 616 for communicating with the network 150 .

[0279] System 600 may store or have access to computer programs / software (e.g., instructions and data) that, when executed by processing unit 602, configure system 600 to receive, process, and output data. Such programs typically include operating systems such as Microsoft Windows, Apple OSX, Apple IOS, Android, Unix, or Linux.

[0280] System 600 also stores or has access to instructions and data (i.e., software applications) that, when executed by processing unit 602, configure system 600 to perform various computer-implemented processes / methods according to the various embodiments described above. It will be appreciated that in some cases, some or all of a given computer-implemented method is performed by system 600 itself, while in other cases, processing may be performed by other devices in data communication with system 600. .

[0281] The instructions and data are stored on a non-transitory, machine-readable medium accessible to system 600. For example, the instructions and data may be stored in non-transitory memory 610. The instructions may be transmitted to / received by system 600 via a data signal in a transmission channel enabled by (for example) a wired or wireless network connection. In some embodiments, any one or more of the data sources used in the methods described herein may be stored on a remote data storage or server system. In such cases, the data may be provided to system 600 via a data signal in a transmission channel enabled by a wired or wireless network connection. For example, databases storing ChIP-seq information, protein-protein interaction network data, and gene sequence databases may be located remotely from system 600 and accessed via a computer network, e.g., the Internet, to enable implementation of the methods described herein.

[0282] The present invention includes the following non-limiting examples. [Example]

[0283] Example 1 Data preprocessing: histone modification ChIP-seq and RNA-seq data Histone modification data for all available human cell types, including primary cells, stem cells, and tissues, were obtained from the ENCODE repository. H3K4me3 ChIP-seq data for 316 samples representing 116 cell types was downloaded on April 20, 2017, and H3K27me3 ChIP-seq data for 208 samples representing 91 cell types was downloaded on January 20, 2018. The processed ChIP-seq data deposited in ENCODE were obtained using different processing pipeline parameters or mapped to different human genome versions (hg19 or hg38). Therefore, to ensure consistent and cohesive ChIP-seq data processing, histone ChIP-seq and corresponding control / input ChIP-seq sequencing files were downloaded in bam format and the sequencing reads were realigned to the hg38 human genome and annotation version V25 using BWA. For datasets made available by the Roadmap project, tagAlign FASTA-formatted files were provided instead of bam files, and the tagAlign FASTA files were aligned to the hg38 human genome. Sample quality and fragment size were determined using phantompeakqualtools. ChIP-seq peaks were detected based on the relative enrichment of sequencing reads in samples (treated with H3K4me3 or H3K27me3 antibodies) relative to the input (control without antibodies). For both H3K4me3 and H3K27me3 ChIP-seq data, broad peaks were called using MACS2 with a q-value threshold of 0.1 for enrichment of reads in samples with histone antibodies compared to the input sample. For H3K4me3 ChIP-seq data, a systematic peak quality filtering threshold was identified based on the percentage of discarded peaks with increasing q-value stringency, increasing to 10. -5 It was found that the percentage of discarded peaks at q values ​​of 10 was reduced to less than 10% across all samples. -5A q-value of 0.01 was used to remove low-quality peaks in the H3K4me3 ChIP-seq data. For the H3K27me3 ChIP-seq data, a default q-value threshold of 0.1 was used because increasing the q-value stringency to 0.01 discarded many samples. Additionally, low-quality samples with fewer than 10,000 peaks were also removed to ensure that high-quality samples were used to model cell types. 11 for each of the H3K4me3 and H3K27me3 ChIP-seq data. There were 1 and 81 cell types (Fig. 1A).

[0284] Example 2 EpiMogrify algorithm A computational method called "EpiMogrify" has been developed that models the epigenetic state of cells to predict factors important for cell identity, cell maintenance, and cell transformation. The main sections of the algorithm are as follows: (i) Definition of histone modification reference peak loci (ii) Identification of cell identity and cell maintenance factors (iii) Identification of cellular transformation factors for directed differentiation and transdifferentiation (i) Definition of histone modification reference peak loci For each cell type, samples were pooled together to obtain a ChIP-seq profile representative of the cell type. The ChIP-seq peak width representing the cell type was calculated by merging the peak areas across samples, and the peak height was calculated by the maximum peak height across all samples. To compare ChIP-seq peaks across various cell types, a reference peak locus (RPL) was defined as the set of areas obtained by combining representative peaks across all cell types (Figure 1C). Let n be the number of cell types. For each cell type in the range x1, x2, ... xn, let b be the ChIP-seq peak width of the cell type and h be the ChIP-seq peak height. The genomic location of the reference peak locus (RPL) was obtained by merging the peak areas across all cell types. For each locus R1, R2, ... Rn in the RPL, the peak width (b) was calculated by merging the overlapping peaks across cell types, and the peak height (h) was given by the maximum peak height of the overlapping peaks.

[0285]

number

[0286] Protein-coding genes can be assigned to RPLs based on the genomic location of the peak and the gene's transcription start site (TSS). Because the H3K4me3 histone modification is a well-known promoter mark, genes are typically assigned to the peak locus if the peak overlaps the gene's promoter region (500 bp from the TSS). However, it is understood that RPLs can be located in other regions of the gene relative to its TSS (including extending beyond 500 bp from the TSS, including 1000 bp, 2000 bp, or more from the TSS).

[0287] Preferably, for each cell type (x), the peak width value (B) of a gene (g) is calculated as the sum of the peak widths of the n peaks annotated to the gene, while the peak height value of the gene is calculated as the maximum height of the n peaks annotated to the gene.

[0288]

number

[0289] is the peak width profile of a given cell type x at the RPL (reference peak locus).

[0290]

number

[0291] (ii) Identification of cell identity genes and cell maintenance factors Because genes associated with broad H3K4me3 ChIP-seq widths were observed to be enriched for cell identity genes, EpiMogrify utilizes H3K4me3 ChIP-seq peak widths to model cell states. Poised genes, where both H3K4me3 and H3K27me3 ChIP-seq peaks exist at the gene's TSS, are removed from the model. EpiMogrify uses a three-step approach to identify cell-type-specific factors (Figure 2). First, a differential breadth score is computed for each RPL based on the H3K4me3 ChIP-seq peak width. Second, the regulatory role of a gene is determined based on the value of associated genes on a protein-protein interaction network. Finally, cell identity genes are predicted based on the ranked protein-coding genes, and EpiMogrify predicts signaling molecules for cell state maintenance.

[0292] Step 1: Calculate the differential wideness score While primary cells and stem cells are homogeneous cell populations, tissues are composed of heterogeneous cell populations. Therefore, for comparison, primary cells and stem cells were treated as one group, and tissue cell types were treated as another group. To obtain target cell type-specific ChIP-seq profiles, the target cell type was compared to a set of background cell types within the group. For example, if the target cells were a primary cell type, the background cell types would be the remaining primary and stem cell types. Because the cell types in the background set should not be similar to the target cell type, background cell types were selected only if the Spearman correlation between the ChIP-seq profile and the target cell type was less than 0.9.

[0293] If more than one Reference Peak Locus (RPL) is assigned to a gene (g) in a target cell type of interest (x), then the gene peak width score (B) is given by the sum of the peak width (b) values ​​of the assigned RPLs. x g is the set of gene peak width scores for background cell types in gene g, which are selected based on their distinctiveness from the target cell type (x). x g is the normalized difference in gene peak width score between the target cell type and the median gene peak width score of the background cell types. The significance (p-value) of this difference is estimated by a one-sample Wilcoxon test. The differential width score (DBS) of a cell type (x) and gene (g) is measured as the sum of the normalized Δpeak width and Pval.

[0294]

number

[0295] Step 2: Calculating the accommodative differential broadness score Next, information was included to computationally calculate the regulatory effect of genes on cell type cell identity genes, STRING V10, and protein-protein interaction networks. Interactions between gene nodes were selected if the experimental score was greater than zero and the combined score was greater than 500. This ensures that high-quality networks with experimental evidence are used to determine the regulatory effect of genes. A similar network score (Net) calculation was used as in Mogrify. For every gene (V) in the STRING network, the network score was calculated by multiplying the network score by the sub-network of genes (V). g The correlation coefficients were computed based on the sum of the DBS scores of associated genes (r) in the STRING network, corrected for the number of gene outdegree nodes (O) and the level of association (L). To obtain cell-specific networks, genes without associated broad H3K4me3 peaks from the STRING network were removed. A weighted sum of the DBS scores of associated genes was calculated up to the third level of association.

[0296]

number

[0297] Normalized network score (Net) across all protein-coding genes G in cell type x

[0298]

number

[0299] To construct cell-type-specific networks, we removed genes whose peak width was below a threshold width (87%) (Figure 1E). Different combinations of DBS and Net scores were tested for enrichment of cell identity genes, and we determined that a 2:1 ratio of DBS to Network score (Net) showed the highest enrichment for cell identity genes.

[0300]

number

[0301] Step 3: Prediction of cell identity genes and cell maintenance factors EpiMogrify predicts protein-coding genes that indicate cell identity by ranking genes based on RegDBS. For cell maintenance, EpiMogrify predicts signaling molecules, such as receptors and ligands, essential for cell growth and survival. Receptor-ligand interaction pairs were obtained from studies and identified based on gene expression across different cell types. It was found that approximately two-thirds of ligands are produced by cells in an autocrine manner, while the remainder are produced by supporting cell types to mimic the microenvironment. Therefore, to incorporate this into the current model, receptors were prioritized based on cell-specific RegDBS, as they should both be cell-specific and have regulatory effects on the cell type of interest. Ligands from the cell type of interest and supporting cell types were then prioritized based on RegDBS. Finally, receptor-ligand pairs were prioritized by the combined rank of the receptor and ligand. The ligands of predicted receptor-ligand pairs can be supplemented into cell culture conditions.

[0302] (iii) Identification of cellular transformation factors for directed differentiation and transdifferentiation EpiMogrify utilizes a three-step approach to identify cell transformation factors. First, it computes the change in differential broadening score upon cell transformation from a source cell type to a target cell type. Then it determines the regulatory effect of genes on the change in cell state. Finally, it predicts transcription factors for transdifferentiation and signaling molecules for directed differentiation.

[0303] Step 1: Calculating the cell transform differential broadening score For cell conversion, let S be the source cell type and T be the target cell type. First, a differential broadening score (DBS) is computed for both the source and target cell types. Then, the difference between the source DBS score and the target DBS score is computed to obtain the cell conversion ΔDBS.

[0304]

number

[0305] Step 2: Calculating the cell transformation regulation differential breadth score The regulatory network score (N) for cell transformation from the source cell type to the target cell type is then computed as a weighted sum of the cell transformation ΔDBS scores of the associated nodes. Similar to the calculation of cell-specific RegDBS, the cell transformation RegΔDBS is calculated as the aggregate value of cell transformation ΔDBS and normalized network scores in a 2:1 ratio.

[0306]

number

[0307] Step 3: Predicting cellular transformation factors for directed differentiation and transdifferentiation For each cell transformation, protein-coding genes are ranked by their RegΔDBS value. For transdifferentiation, EpiMogrify predicts transcription factors (TFs), a subset of protein-coding genes defined based on the TFClass classification. For directed differentiation, EpiMogrify predicts signaling molecules such as receptor and ligand pairs. Receptors are prioritized by the cell transformation RegΔDBS, and corresponding ligands are prioritized by the cell transformation ΔDBS. Finally, receptor-ligand pairs are prioritized by the combined rank of the receptor and ligand. The ligands of the predicted receptor-ligand pairs can be supplemented into the differentiation protocol.

[0308] In summary, EpiMogrify models a wide range of H3K4me3 histone modification signatures and predicts cell identity protein-coding genes, signaling molecules for cell maintenance, signaling molecules for directed differentiation, and TFs for transdifferentiation.

[0309] Example 3 Characterization of H3K4me3 and H3K27me3 histone modifications Cellular states can be modeled by utilizing epigenetic histone modifications, namely H3K4me3 and H3K27me3, which are well-known transcriptional activator and repressor marks, respectively. ChIP-seq data for H3K4me3 and H3K27me3 histone modification profiles of various human cell types were obtained from the ENCODE and Roadmap Consortium data repositories. To ensure consistent downstream processing of ChIP-seq datasets across different samples and experiments, an analysis pipeline was implemented, and, where necessary, data-driven thresholds were used to call significant peaks in the H3K4me3 and H3K27me3 profiles. After filtering low-quality peaks, samples representing 111 different cell types with H3K4me3 profiles (51 primary cells, 53 tissues, 7 stem cells) and 81 cell types with H3K27me3 profiles (28 primary cells, 45 tissues, 8 stem cells) were retained (Figure 1A). To investigate the effects of histone modifications on transcriptional status, we used gene expression profiles from ENCODE, consisting of 137 cell types (64 primary cells, 68 tissues, and 5 stem cells). There are 40 common cell types for which epigenetic (H3K4me3 and H3K27me3 by ChIP-seq) and transcriptional (RNA-seq) data are available. ChIP-seq peak width was defined as the length of the genomic region with histone modification deposition in base pairs (bp), and ChIP-seq peak height was defined as the average enrichment of the histone modification or signal value (Figure 1B). ChIP-seq peak width ranged from narrow to wide peaks, while ChIP-seq peak height ranged from short to tall peaks, with a distribution across all cell types. Furthermore, compared to H3K4me3, H3K27me3 has a greater number of narrow and wide ChIP-seq peaks, and H3K27me3 is known to be present in intergenic regions and silent genes.

[0310] Next, we determined the genomic location preferences of H3K4me3 and H3K27me3 ChIP-seq profiles across all cell types. We assessed the frequency of ChIP-seq peaks within a 2000-bp window around the transcription start sites (TSSs) of protein-coding genes. We found that, in contrast to H3K27me3 marks, H3K4me3 had a higher frequency of peaks within 500 bp from the TSS.

[0311] Next, protein-coding genes were annotated for H3K4me3 and H3K27me3 ChIP-seq profiles. To this end, gene values ​​were assigned based on the associated ChIP-seq peak width or height. Figure 1C shows a detailed schematic of the protein-coding gene annotations and assigned gene values. For each cell type, a representative H3K4me3 and H3K27me3 ChIP-seq peak profile was obtained by merging multiple samples or replicates (see Methods for details). Next, a reference ChIP-seq peak region, called the reference peak locus (RPL), was defined, which was obtained by merging overlapping peaks across cell types. For each RPL, protein-coding genes were annotated relative to the locus if their promoter region overlapped with the locus. For H3K4me3 marks, the promoter region was defined as 500 bp from the gene's TSS, because this region showed a higher frequency of H3K4me3 ChIP-seq peaks. Regarding H3K27me3 marks, we defined the promoter region as −500 bp from the TSS of a gene to exclude considering H3K27me3 marks in gene introns. The table in Figure 1C shows example gene annotations for RPL and the values ​​for each gene based on ChIP-seq peak width and height for each cell type.

[0312] Example 4 A data-driven approach to model H3K4me3 profiles to identify cell identity genes To determine the optimal features for modeling cell identity, we tested histone modification features, such as ChIP-seq peak width and height, and transcriptional features, such as gene expression levels. First, protein-coding genes were categorized into 5th percentile bins based on the distribution of H3K4me3 and H3K27me3 ChIP-seq peak width and height, as well as gene expression. Next, we estimated whether either histone modifications or gene expression features were enriched exclusively in cell identity genes. Cell identity genes were defined by text mining of GO biological processes associated with cell type, and the significance of enrichment was calculated using Fisher's exact test (FET). We observed that across all cell types, bins with high gene expression showed higher enrichment significance for both cell identity and housekeeping genes. While high H3K4me3 ChIP-seq peaks were enriched for housekeeping genes, wide H3K4me3 ChIP-seq peaks were found to be significantly enriched for cell identity genes, as shown in previous studies. As expected, the H3K27me3 repressor mark was not enriched for cell identity or housekeeping genes across all cell types. Because cell identity genes were significantly enriched in genes with high gene expression or broad H3K4me3 ChIP-seq peaks, these characteristics were further investigated to identify the optimal metric (H3K4me3, RNA-seq expression, or both) to model cell identity genes. For 40 common cell types, genes were ordered in descending order based on their gene expression or H3K4me3 ChIP-seq peak width, and genes were grouped into cumulative windows in 1 percentile increments (Figure 1D). For each cumulative window, the significance of enrichment for cell identity and housekeeping genes across common cell types was computed. High gene expression was associated with both cell identity and housekeeping genes, while genes indicated by broad H3K4me3 ChIP-seq peaks were enriched exclusively for cell identity genes. We found that H3K4me3 ChIP-seq peak widths significantly increased across 111 cell types (Figure 1D). We then examined the significance of enrichment for cellular identity genes to determine a threshold for H3K4me3 ChIP-seq peak widths that could efficiently identify cellular identity genes. H3K4me3 ChIP-seq peak widths above the 87th percentile of the peak width distribution were found to have the highest significance of enrichment across 111 cell types (Figure 1E). For each cell type, this threshold (87th percentile) corresponded to ChIP-seq peak widths ranging from 3,611 to 15,035 base pairs. Collectively, these analyses suggest that wide H3K4me3 ChIP-seq peak widths are a characteristic associated with cellular identity.

[0313] Next, we developed a data-driven method and "gene score" to model cell states and identify cell-specific genes by systematically analyzing the broad H3K4me3 ChIP-seq peak characteristics. To identify cell-specific genes, we compared the target cell type of interest with background cell types. Background cell types were selected based on their distinctiveness from the target cell type, as calculated by Spearman's correlation. Figure 2A shows a schematic diagram of the approach used to prioritize cell-specific genes. For each cell type and protein-coding gene, we computed a differential breadth score (DBS). This is a composite value calculated by the difference in ChIP-seq peak width between the target cell type and background, and the significance of this difference. Protein-coding genes ranked by DBS are cell-specific genes. To prioritize cell-specific genes that exert regulatory effects on cells, we computed a regulatory differential breadth score (RegDBS) (Figure 2B). RegDBS is a composite value calculated by the gene's DBS and protein-protein interaction network score. Using this approach, it was possible to systematically prioritize cell-specific protein-coding genes for 111 cell types, with each gene ranked by its RegDBS score.

[0314] We sought to determine whether this method would improve the accuracy of detecting cellular identity genes compared to previous studies. Different scoring metrics modeling H3K4me3 ChIP-seq peak width were compared by measuring enrichment for cellular identity genes using GSEA (Gene Set Enrichment Analysis) (Figure 2C). Compared to ranking protein-coding genes by H3K4me3 ChIP-seq peak width alone, DBS was found to exhibit 2.5-, 3.5-, and 3.5-fold higher enrichment for cellular identity genes, respectively. RegDBS had the highest enrichment for cellular identity genes across all cell types, with an approximately 5.3-fold increase in enrichment compared to genes ordered by peak width and a 1.5-fold increase compared to DBS. This confirms that the method of the present invention is more accurate in identifying genes with ontologies associated with cellular identity. Notably, RegDBS scores, which utilize regulatory network information, can be used to prioritize cell-specific genes.

[0315] Example 5 EpiMogrify predicts factors for cell maintenance, e.g., astrocytes Since the method of the present invention can efficiently prioritize cell identity genes, we hypothesized that the prioritized signaling molecules, such as receptors and ligands, are essential for cell survival and maintenance of cell state. Figure 2D shows the EpiMogrify approach to predicting signaling molecules for cell maintenance. By prioritizing receptors based on the RegDBS metric, we hypothesized that it would be possible to identify receptors that can induce the activation of cell identity genes, since these receptors are weighted by their regulatory effects. Meanwhile, we hypothesized that the ligands required for activation of these receptors could be identified. Motility signals can be produced by target cell types in an autocrine manner or by supporting cell types in a paracrine manner. Therefore, we prioritized corresponding ligands by DBS and obtained ranked receptor-ligand pairs that were important to target cell types (Figure 2D). When cells are transferred from in vivo systems to in vitro conditions, they express the necessary receptors. However, the ligands required to activate these receptors may be lost. Therefore, for in vitro cell maintenance, we hypothesized that EpiMogrify could predict ligands that could be added to cell maintenance culture media. To evaluate the accuracy of EpiMogrify's predictions, the predicted ligands for growth and maintenance were tested on astrocytes in vitro. Astrocytes are a multifunctional cell type that provide metabolic and structural support for neurons and regulate neurogenesis and neural pathways in the brain. Some dysfunction of astrocytes can lead to major neurological disorders. Therefore, improving our ability to grow and maintain astrocytes is essential for studying neurological disorders.

[0316] The top seven signaling molecules predicted by EpiMogrify for astrocyte maintenance are components of the extracellular matrix (ECM), such as FN1, COL4A1, LAMB1, and COL1A2, as well as secreted ligands, such as ADAM12 and EDIL3. Because LAMB1 is part of the laminin trimer, the astrocyte-specific laminin-211 complex, consisting of LAMA2, was used to predict LAMB1 and LAMC1, respectively. The predicted factors were tested in astrocytes obtained from two different sources: astrocytes differentiated from neural stem cells and primary astrocytes. Astrocytes were cultured in conditions supplemented with the predicted signaling molecules individually and together. Culture medium without additional ligands was used as a negative control and compared with Matrigel as a positive control.

[0317] We found a significant increase in cell growth in culture conditions with EpiMogrify predictor compared to Matrigel (Figures 3A and 3B). We then performed immunofluorescence for astroglial markers, such as S100b and GFAP, to confirm that the increased growth rate did not affect cellular identity, with astrocytes differentiated from neural stem cells and primary astrocytes expressing both markers (Figures 3C and 3D). Cell maintenance was achieved using only a few signaling molecules (e.g., FN1, COL1A2, or LAMB1), and in some cases, growth efficiency was shown to be comparable to that achieved with Matrigel, a chemically undefined complex.

[0318] Example 6 EpiMogrify predicts factors for cell maintenance, e.g., cardiomyocytes The top seven signaling molecules predicted by EpiMogrify for cardiomyocyte (heart cell) maintenance were FN1, COL3A1, TFPI, FGF7, and APOE. Cardiomyocytes were cultured in conditions supplemented with the predicted signaling molecules individually and together. Culture medium without additional ligands was used as a negative control and compared with Geltrex as a positive control.

[0319] We found a significant increase in cell growth in culture conditions with EpiMogrify predictor compared to Geltrex (Figures 4A and 4B). We then performed immunofluorescence for cardiomyocyte markers, such as GATA4 and NKX2.5, to confirm that the increased growth rate did not affect cellular identity in cells expressing both markers (Figure 4C).

[0320] Example 7 EpiMogrify predicts factors for cell maintenance, e.g., smooth muscle cells Next, we compared the EpiMogrify predicted conditions to determine whether they correlate with smooth muscle cell maintenance in vitro. To test the predictions in smooth muscle cells, human pulmonary artery smooth muscle cells (HPASMCs) were cultured in vitro.

[0321] Because H3K4me3 ChIP-seq data for HPASMCs was not available, we used the H3K4me3 profile of the closest available cell type, "tissue: gastric smooth muscle," to predict ligands for smooth muscle cell maintenance. The top 13 ligands predicted by EpiMogrify for smooth muscle cell maintenance are LAMA5, COL4A1, LAMA4, NID1, COL6A3, COL4A6, COL4A5, FGF10, FGF7, GNAS, COL7A1, COL1A1, and THBS1. Several of these predictions are subunits of collagen 4-like complex proteins, including the predicted subunits COL4A1, COL4A6, and COL4A5; collagen 6 is composed of COL6A3; and collagen 1 is composed of COL1A1. Therefore, collagen 4, NID1, collagen 6, FGF10, FGF7, collagen 1, and THBS1 were tested for smooth muscle cell maintenance. Smooth muscle cells were cultured in culture conditions supplemented with the predicted ligands individually or all together, with culture conditions that contained no additional ligands serving as a negative control and culture conditions with a chemically undefined complex mixture (Geltrex) serving as a positive control.

[0322] To evaluate the validity of the predictions, we counted the number of cells in each condition, along with assessing proliferation rate as a surrogate for cell growth. A significant increase in proliferation rate was observed between culture conditions using EpiMogrify predicted ligands (all ligands combined, including collagen 6, collagen 1, NID1, and FGF7) compared to the negative control (Figures 5A and 5B). Furthermore, smooth muscle cells in the maintenance condition containing all EpiMogrify predictions expressed a significantly higher percentage of the smooth muscle marker α-SMA compared to the negative control (Figure 5C). Next, immunofluorescence for the smooth muscle marker SM22 was performed, demonstrating that the EpiMogrify predicted condition preserved cell identity (Figure 5D).

[0323] Example 8: Epimogrify predicts factors for cell maintenance, e.g., endothelial cells To test the predictions in endothelial cells, we cultured primary human aortic endothelial cells (HAoECs) in vitro. Because H3K4me3 ChIP-seq data for HAoECs was unavailable, we used the H3K4me3 profile of the closest available cell type, "primary cells: umbilical vein endothelial cells," to predict ligands for endothelial cell maintenance. The top 10 ligands predicted by EpiMogrify for endothelial cell maintenance are BMP6, ADAM9, LAMB1, LAMA4, THBS1, CTGF, BMP4, PDGFB, FN1, and CYR61. Of these, BMP6, THBS1, CTGF, BMP4, PDGFB, FN1, and CYR61 were tested for HAoEC endothelial cell maintenance.

[0324] Utilizing an experimental design similar to that used for smooth muscle cells, endothelial cells were maintained in EpiMogrify-predicted conditions (combined and individual) and compared to both negative and positive controls (Geltrex). The EpiMogrify-predicted conditions (all ligands combined, as well as individual factors such as PDGFB, THBS1, CTGF, BMP4, PDGFB, and CYR61) were able to maintain significantly improved cell numbers and proliferation rates compared to the negative control (Figure 6A and B). Furthermore, the EpiMogrify-predicted conditions expressed the endothelial markers CD31 and VWF (Figure 6C and D). While there was no significant increase in cells expressing the VWF marker compared to the negative control, there was a significant increase in cells expressing that marker in the EpiMogrify-defined conditions (including individual ligands and all ligands combined) compared to the chemically undefined complex (Geltrex). There was a significant increase in the number of cells.

[0325] Example 9 EpiMogrify predicts factors for cell transformation, e.g., astrocytes Because EpiMogrify predicts factors that can maintain cell state, we hypothesized that by modeling changes in cell state, we could identify factors required for cell transformation. Figure 7A shows a detailed schematic of the cell transformation score calculation for cell state changes from a source cell type to a target cell type. Genes specific to cell state changes can be prioritized based on the difference in DBS of genes between the target and source cell types. Each gene is then weighted based on the regulatory effect that influences the cell state change to derive a RegΔDBS score for each gene. For each cell transformation, EpiMogrify prioritizes protein-coding genes (including TFs and signaling molecules) ranked by RegΔDBS. A few experimentally validated methods exist for modeling the transcriptional state (or changes in transcriptional state) of cells and identifying TFs for cell transformation, including CellNet, D'Alessio AC et al. (also referred to herein as "JSD" because it uses the Jensen-Shannon divergence statistic), and Mogrify. We compared the epigenetic-based predictions of EpiMogrify with those of transcription-based computational methods. There were 29 common cell types between EpiMogrify and JSD, 43 between EpiMogrify and Mogrify, and 25 common cell types between all three methods. Using GSEA, we found that for common cellular transformations, TFs predicted by EpiMogrify had significant enrichment for TFs predicted by JSD (76.4% of all transformations), Mogrify (60.7%), or both JSD and Mogrify (69.3%) (Figure 7B). Significant overlap was found between EpiMogrify-predicted TFs and other methods. The high occurrence of published TFs among the top 10 predicted by EpiMogrify (Figure 7C) demonstrates EpiMogrify's high accuracy in identifying TFs for cellular transformations.

[0326] Next, we applied EpiMogrify to prioritize signaling molecules for directed differentiation of embryonic stem cells into astrocytes. The predicted signaling molecules were tested in the differentiation protocol.

[0327] EpiMogrify predictions were evaluated for astrocyte differentiation, identifying factors such as FN1, COL4A1, LAMB1, COL1A2, EDIL3, ADAM12, and WNT5A. The astrocyte differentiation protocol was applied to H9 embryonic stem cells (Figure 8A). At day 35 of differentiation, a significant increase in cells positive for the astrocyte marker CD44 was observed in the "all ligands" condition (all predictors combined). Subsequently, at day 42, a significant increase in cells expressing CD44 was observed in H9 cells treated with FN1, COL4A1, LAMB1, COL1A2, and "all ligands" compared to Matrigel controls (Figure 8B). Immunofluorescence was performed to obtain the percentage of DAPI+ cells bearing GFAP and S100b astrocyte markers in all conditions (Figure 8D). Under the predicted conditions, the cultures also expressed the astroglial markers S100b and GFAP at day 42 (Figure 8C), suggesting that the predicted factors increase the efficiency of astroglial differentiation.

[0328] For a group of genes exhibiting astrocyte specificity (astrocyte transcriptional signature), the average z-score of gene expression profiles in TPM is calculated across the three samples sequenced for each time point and condition. The astrocyte transcriptional signature is comprised of (i) significantly upregulated genes between H9 ESCs and primary astrocytes from our RNA-seq data, (ii) significantly upregulated genes between H9 ESCs and cortical-derived astrocytes from the FANTOM5 (F5) database, (iii) significantly upregulated genes between H9 ESCs and primary astrocytes from the FANTOM5 (F5) database, and (iv) significantly upregulated genes between H9 ESCs and primary astrocytes from the FANTOM5 (F5) database. (iv) EpiMogrify GRN (gene regulatory network) is a gene set containing cell identity genes of primary astrocytes with positive RegDBS scores and associated genes on the STRING network up to the first regulated neighbor; and (v) HumanBase GRN is an astrocyte-specific network obtained from the HumanBase database (Figure 8D).

[0329] Example 10 EpiMogrify predicts factors for cell transformation, e.g., cardiomyocytes Next, we tested the conversion conditions predicted by EpiMogrify for cardiomyocyte differentiation (FGF7 (fibroblast growth factor 7), COL1A2 (collagen I), TGFB3 (transforming growth factor beta 3), COL6A3 (collagen VI), EFNA1 (ephrin A1), TGFB2 (transforming growth factor beta 2), and GDF5 (growth differentiation factor 5)). To do this, we supplemented the culture medium with ligands individually and together for H9 ESCs during differentiation, using Matrigel as a control (Figure 9A). At day 12 of differentiation, we observed a significant increase in cells expressing the early cardiac lineage markers CD82 and CD13 in the predicted conditions (Figure 9B). By day 20, the remaining conditions also yielded cells expressing early cardiac lineage markers, but the predicted conditions showed significantly higher expression of the late cardiac marker cTNT at both days 12 and 20 compared to the Matrigel control (Figure 9C,D). Furthermore, the phenotype of differentiated cardiomyocytes was determined based on the potential of the differentiated cells to beat, showing that a greater percentage of plates contained beating cells when using conditions predicted by EpiMogrify (combination of all predictors) compared to Matrigel. Collectively, these data indicate increased cardiac differentiation efficiency, resulting in a higher percentage of functional cardiomyocytes.

[0330] Example 11 Experimental Method Cell counting assay 5 × 10 of ReN astrocytes and primary astrocytes 3 Cells were seeded into 48-well plates in the following conditions: no substrate, Matrigel, FN1 (fibronectin) (Sigma-Aldrich), COL4A1 (collagen IV) (Sigma-Aldrich), LAMB1 (laminin 221) (Biolamina), ADAM12 (Sigma-Aldrich), COL1A2 (collagen I) (Sigma-Aldrich), EDIL3 (R&D Systems), and all ligands (the above substrate and growth factor combinations without Matrigel). After 3 days, cells were dissociated using trypsin / EDTA solution (TE) (Gibco) and counted using an EVE automated cell counter (NanoEnTek Inc.).

[0331] For primary cardiomyocytes, 5 × 10 4 Cells / well were seeded into 12-well plates for each condition: no substrate, Geltrex, FN1 (fibronectin), COL1A2 (collagen I), TFPI (tissue factor pathway inhibitor), ApoE (apolipoprotein E), FGF7 (fibroblast growth factor-7), and all ligands (the above substrate and ligand combinations without Geltrex). After 72 hours, cells were dissociated using 0.04% trypsin / 0.03% EDTA (Promocell) and counted using a LUNA-II™ automated cell counter (Logos Biosystem, Inc.).

[0332] For smooth muscle cells, 5 × 10 4 Cells / well were seeded in 12-well plates for each condition: no substrate, Geltrex, NID1 (nidogen-1), COL1A1 (collagen I), COL4A1 (collagen 4), COL6A1 (collagen 6), TH For endothelial cells: no substrate, Geltrex, FN1 (fibronectin), THBS1 (thrombospondin 1), FGF10 (fibroblast growth factor-10), FGF7 (fibroblast growth factor-7), and all ligands (Geltrex). For endothelial cells: no substrate, Geltrex, FN1 (fibronectin), THBS1 (thrombospondin 1), CTGF (connective tissue growth factor), PDGFB (platelet-derived growth factor subunit B), CYR61 (cysteine-rich angiogenic factor 61), BMP4 (bone morphogenetic protein 4), BMP6 (bone morphogenetic protein 6), and all ligands (Geltrex). After 72 hours, cells were dissociated using TrypLE Express (Life Technologies) and counted using a LUNA-II™ automated cell counter (Logos Biosystem, Inc.).

[0333] BrdU proliferation assay 5 × 10 of ReN astrocytes and primary astrocytes 3 Cells were seeded in 96-well black plates with clear bottoms (Corning). After 12 hours, BrdU labeling was added to the wells according to the manufacturer's recommendations (Cell Proliferation ELISA, BrdU (chemiluminescence), Roche). After 72 hours, cells were fixed and labeled with anti-BrdU antibody, followed by a substrate reaction and analysis using a PHERAstar FSX microplate reader (BMG Labtech). For primary cardiomyocytes, the BrdU Cell Proliferation Assay Kit (BioVision Inc.) was used according to the manufacturer's instructions. Briefly, the indicated cells were cultured at 5 × 10 3 Cells / well were seeded into 96-well tissue culture-treated plates (Costar). After 72 hours, BrdU solution was added to each assay well and incubated at 37°C for 3 hours. BrdU incorporated by proliferating cells was detected by enzyme-linked immunosorbent assay.

[0334] For both smooth muscle cells and endothelial cells, the BrdU Cell Proliferation Assay Kit (BioVision Inc.) was used according to the manufacturer's instructions. Briefly, the indicated cells were cultured at 5 × 10 3 Cells / well were seeded into 96-well tissue culture-treated plates (Costar). After 72 hours, BrdU solution was added to each assay well and incubated at 37°C for 4 hours. BrdU incorporated by proliferating cells was detected by enzyme-linked immunosorbent assay.

[0335] Flow cytometry For differentiation experiments, cultures were dissociated using Accutase (Stem Cell Technologies) and pelleted at 400 × g for 5 min. Cells were then incubated with APC-Cy7 CD44 antibody (103028, Biolegend); or PE-Cy7 CD13 (56159, BD Biosciences) and PE in 2% fetal bovine serum (FBS) (Gibco) and phosphate-buffered saline (PBS) (Gibco). The cells were resuspended in CD82 (342104, Biolegend) and incubated at 4°C for 15 minutes. The cell suspension was washed with PBS and pelleted at 400 × g for 5 minutes for analysis. Cell viability was determined using 4',6-diamidino-2-phenylindole (DAPI) (Life Technologies). Cells were analyzed on an LSR IIa analyzer (BD Bioscience) using BD FACSDiva software (BD Bioscience). Samples were analyzed using a chromatograph (Microwave Aid, Microwave Aid, Microwave Aid, Microwave Batch Array ...

[0336] For primary cardiomyocytes, smooth muscle cells, and endothelial cells, cells were dissociated using 0.04% trypsin / 0.03% EDTA (Promocell) and pelleted at 300 × g for 5 minutes. Cells were then suspended and fixed in 4% paraformaldehyde (PFA) at room temperature, washed twice in FACS buffer (1 × PBS, 0.2% BSA), and permeabilized with 0.1% Triton X-100. Primary antibodies for cardiomyocytes were rabbit anti-human Nkx2.5 (Cell Signaling) and rabbit anti-human GATA-4 (Novus The primary antibodies used for smooth muscle cells and endothelial cells were mouse monoclonal antibodies against alpha-smooth muscle actin (Abcam) and rabbit polyclonal antibodies against von Willebrand factor (Abcam), respectively. Secondary antibodies used included Alexa Fluor 488 goat anti-rabbit (Life Technologies) and Alexa Fluor 647 goat anti-rabbit (Life Technologies). Cells were stained with the primary antibodies for 30 minutes at room temperature. After two washes, cells were incubated with the secondary antibodies, washed twice, and resuspended in FACS buffer for flow cytometry analysis using a BD FACS-Fortessa instrument (BD Biosciences).

[0337] Immunofluorescence Cultured cells were fixed with 4% PFA (Sigma Aldrich) in PBS for 10 minutes and then permeabilized in PBS containing 0.3% Triton X-100 (Sigma Aldrich). Cultures were then incubated with primary antibodies followed by secondary antibodies (see dilutions below). 4',6-diamidino-2-phenylindole (DAPI) (1:1000) (Life Technologies) was added to visualize cell nuclei. Images were captured with an inverted fluorescence microscope and attached camera (Nikon Eclipse Ti). The primary antibody used in this study was anti-S100b (6673, Sigma Aldrich). Aldrich, 1:200), anti-GFAP (ab7260, Abcam, 1:500), and cTNT (MS-295-P1, Thermo Fisher, 1:100). Secondary antibodies used in this study (all 1:400) were Alexa Fluor 488 donkey anti-mouse IgG (A21202, Life Technologies), Alexa Fluor 555 donkey anti-rabbit IgG (A31572, Life Technologies), and Alexa Fluor 488 goat anti-mouse IgG (A28175, Life Technologies).

[0338] For endothelial cells and smooth muscle cells, cultured cells were treated with 4% PFA (Sigma) in PBS. Cultures were then fixed with 0.1% Triton X-100 (Sigma-Aldrich) for 10 min and then permeabilized in PBS containing 0.1% Triton X-100 (Sigma-Aldrich). Cultures were then incubated with primary antibodies followed by secondary antibodies (see dilutions below). Hoechst Cell nuclei were visualized by adding 33342 (1:1000) (Life Technologies). Images were taken with an inverted fluorescence microscope. The primary antibodies used for endothelial cells and smooth muscle cells were anti-CD31 (Abcam, 1:200) and anti-Anti-TAGLN / Transgelin (Abcam, 1:200), respectively. The secondary antibodies used in this study (all 1:400) were Alexa Fluor 488 goat anti-mouse (Life Technologies) and Alexa Fluor 488 donkey anti-goat (Life Technologies).

[0339] RNA-seq library preparation and analysis Total RNA was isolated from 27 samples (H9 ESCs, primary astrocytes, and cells from days 7, 14, and 21 during astrocyte differentiation) using TRIzol™ Reagent (Thermo Fisher Scientific). 100–200 ng of total RNA with a RIN of 8.5–10 was induced for RNA-seq library preparation using the NEBNext Ultra II Directional RNA Library Preparation Kit combined with polyA mRNA enrichment and NEBNext® Multiplex Oligos for Illumina® (96 Unique Dual Index Primer Pairs) according to the manufacturer's protocol. PCR products were purified twice using SPRI beads (Beckman Coulter) to remove excess primers and adapter dimers. Agilent Bioanalyzer DNA 1000 chips (Duke-NUS) were used for RNA-seq library preparation. Library quality was assessed using the Genome Facility, and the molar concentration of each library was determined using a KAPA Library Quantification Kit (Roche). Three pooled libraries (each containing nine samples at a pooled concentration of 8 nM in 30 μL) were sent for sequencing on an Illumina HiSeq4000 (NovogeneAIT Genomics Singapore).

[0340] Gene expression analysis was performed on the above RNA-seq data from H9 ESCs and primary astrocytes at different time points during astrocyte differentiation (days 7, 14, and 21), as well as in control Matrigel and conditions containing all ligands predicted by EpiMogrify. Sequencing reads were trimmed using Trimmomatic v0.39 and aligned to the GRCh38 human genome using STAR v2.7.3. Gene expression profiles for the samples were generated using featureCounts v1.6.0, and differential gene expression profiles were computed using DESeq2 v1.26.0. To compare the gene expression profiles of our RNA-seq data with astrocyte transcriptional signatures, the data were normalized using transcripts per million (TPM) normalization. Next, we selected 5,475 genes (TPM > 1) that were expressed in primary astrocytes, and for these genes, we calculated gene z-scores across all samples (H9 ESCs, primary astrocytes, and differentiated astrocytes at different time points and conditions). To obtain astrocyte transcriptional signatures, differential gene expression profiles between H9 ESCs and astrocytes in the cerebellum, and between H9 and astrocytes in the cerebral cortex, were obtained from the FANTOM5 gene expression atlas. The EpiMogrify GRN consists of astrocyte cell identity genes, which are genes with positive RegDBS scores, and genes directly associated with these cell identity genes on the STRING network. Astrocyte-specific GRNs were obtained from the HumanBase database.

[0341] It is understood that the invention disclosed and defined herein extends to all alternative combinations of two or more of the individual features mentioned or apparent from the text or drawings, all of these different combinations constituting various alternative aspects of the invention.

[0342] Description and embodiments of the present invention i. A method for determining cell identity genes for a cell of interest, comprising: - determining the differential width of H3K4me3 modified regions for each protein-coding gene in the cell of interest; - determining a network score for each protein-coding gene in the cell of interest based on differential widths of H3K4me3 modified regions and interactions between protein products of each protein-coding gene on at least one network, wherein the network contains information on interactions between products of the protein-coding genes; - determining a cell identity score for each protein-coding gene in the cell of interest based on a combination of the differential width and the network score; - prioritizing each protein-coding gene according to its cellular identity score; thereby identifying a cell identity gene for a cell of interest.

[0343] ii. The method of statement i, wherein the step of determining the differential width of H3K4me3 modified regions comprises determining a differential width score (DBS) for each protein-coding gene in the cell of interest, wherein the DBS is based on the difference in width of H3K4me3 modified regions for all protein-coding genes compared to the median width of H3K4me3 modified regions for the same gene in a population of different cell types.

[0344] iii. The method of statement ii, wherein determining the network score is based on interactions between DBS and products of each protein-coding gene on the at least one network.

[0345] iv. A method for determining factors necessary to maintain a cell type in vitro, comprising: - determining a differential breadth score (DBS) for each protein-coding gene in the cell of interest, the DBS being based on the difference in width of H3K4me3 modified regions for all protein-coding genes compared to the median width of H3K4me3 modified regions for the same protein-coding gene in a population of different cell types; - determining a network score for each protein-coding gene in the cell of interest based on interactions between protein products of each protein-coding gene on the DBS and at least one network, the network including information of interactions between products of each protein-coding gene and products of other protein-coding genes in the cell; - determining a cellular identity score (RegDBS) for each protein-coding gene in the cell of interest based on a combination of the DBS and the network score; - prioritizing each protein-coding gene according to its RegDBS, thereby identifying cell identity genes for the cell of interest, wherein each cell identity gene encodes a factor associated with the cell identity of the cell of interest; and thereby identifying factors necessary for maintaining a cell of interest in vitro.

[0346] v. The differential width of the H3K4me3 modified region is - obtaining information about the width of H3K4me3 modified regions for each protein-coding gene in the cell of interest, and obtaining a gene peak width score (B) for each protein-coding gene in the cell; - calculating the difference between the gene-wide score for each gene in the cell and the median gene-wide score for the same gene across populations of cells of different types, thereby determining a differential peak-wide score (DBS) for each protein-coding gene in the cell of interest; The method of any one of statements i to iv, wherein

[0347] vi. The method of any one of statements i to v, wherein the width of the H3K4me3 modified region is determined based on ChIP-seq information, preferably the information is obtained from the ENCODE database.

[0348] vii. The method described in description v or vi, wherein the step of determining a gene peak width score (B) comprises first defining regions of the genome in which significant H3K4me3 modifications are present and defining those regions as histone modification reference peak loci (RPLs).

[0349] viii. The method according to statement vii, wherein the RPL is obtained by merging the obtained ChIP-seq peak areas across all cell types, preferably the merged peak areas are overlapping peak areas.

[0350] ix. The method of any one of statements v to viii, wherein the gene peak width score (B) for each protein-coding gene includes excluding genes in which H3K4me3 and H3K27me3 modified regions are identified, and the genes are identified as poised genes.

[0351] x. The H3K4me3 modification is one of statements i to ix, including the H3K4me3 modification. The method described in the first paragraph. xi. The method of any one of statements ii to ix, wherein the step of determining the differential width (DBS) of H3K4me3 modified regions comprises combining information about the difference in width of H3K4me3 modified regions for each protein-coding gene with the significance of that difference in width.

[0352] xii. The method of statement xi, wherein DBS is determined by adding the difference in peak width and the significance of the difference, preferably DBS is determined by adding the difference in H3K4me3 peak width and the -log10P value of the significance.

[0353] xiii. The method of any one of statements i to xii, wherein the step of determining a network score for each protein-coding gene in the cell of interest comprises combining information about the differential width (DBS) of H3K4me3 modified regions for each protein-coding gene, the number of outdegree nodes of the gene, and the level of relatedness of the gene in the network.

[0354] xiv. The method of statement xiii, wherein the network score is determined by summing the DBSs of related genes, weighted for the number of gene outdegree nodes and the level of relatedness, to obtain a weighted sum of the DBSs for related protein-coding genes in the network.

[0355] xv. The method of any one of statements i to xiv, wherein the protein-protein interaction network is a STRING database. xvi. The method of any one of statements i to xv, wherein the Cellular Identity Score (RegDBS) is a measure of the regulatory effect of each protein-coding gene on the identity of the cell.

[0356] xvii. The method of any one of statements i to xvi, wherein the step of determining a cellular identity score (RegDBS) comprises preferentially weighting protein-coding genes that encode regulatory factors.

[0357] xviii. The method of any one of statements i to xvii, wherein the step of determining a cellular identity score (RegDBS) for each protein-coding gene comprises adding the differential width (DBS) scores of H3K4me3 modified regions and the network score normalized across all genes in the cell of interest.

[0358] xix. The method of statement xviii, wherein adding the DBS and the network score comprises weighting the DBS relative to the network score by a factor of 2:1. xx. The method of any one of statements i to xix, wherein prioritizing each protein-coding gene according to its cellular identity score (RegDBS) comprises ordering the genes based on their RegDBS values.

[0359] xxi. A method according to any one of statements i to xx, wherein the factor for maintaining the cells of interest in vitro is selected from the group consisting of a receptor-ligand pair involved in cell signaling, preferably a ligand of the receptor-ligand pair, a transcription factor, and an epigenetic remodeling factor.

[0360] xxii. A method according to any one of statements i to xxi, comprising selecting cell identity genes encoding receptor-ligand pairs involved in cell signaling, thereby identifying signaling molecules required to maintain the cells of interest in vitro.

[0361] xxiii. Each protein-coding gene encoding a cell surface receptor has its RegD xxii. The method of claim xxii, wherein each ligand associated with a receptor is ranked according to its DBS score to obtain a combined ranking of receptors and ligands, the combined ranking identifying a ligand for use in supplementing culture media for maintaining the cells of interest in vitro.

[0362] xxiv. The method of any one of statements i to xxi, further comprising selecting a cell identity gene encoding a transcription factor, thereby identifying a transcription factor required to maintain the cell of interest in vitro.

[0363] xxv. A method for determining cell identity genes for a cell of interest, comprising: - determining an H3K4me3 gene peak width score (B) for each protein-coding gene (g) in a cell of interest (x), wherein the gene peak width score is the sum of the lengths of regions in the promoters of each protein-coding gene that contain H3K4me3 modifications; - The difference in gene peak width score (ΔPeak Width) between the cell of interest and the median gene peak width score of a population of cells representing the background gene peak width score xg ) and determining the significance of that difference (Pval); - A differential peak width score (DBS) for each protein-coding gene in the cell is calculated by adding the Δ peak width and the Pval value. x g ) and - Differential Broad Score (DBS) of related genes (r) in the network x g ), the network comprising information on interactions between the products of each protein-coding gene and other protein-coding gene products in the cell, the differential breadth score being corrected for the number of out-degree nodes of the gene (O) and the level of relatedness of the gene (L), preferably the differential breadth score (DBS x g ) is calculated up to the third level of relevance; - A network score Net across all protein-coding genes (G) in a cell of interest (x). x g normalizing - Differential Broad Score (DBS) x g ) and network score (Net x g Calculate a regulatory differential breadth score (RegDBS) for each protein-coding gene in the cell of interest by scoring each protein-coding gene in the cell of interest based on a combination of x g ) determining a RegDBS x g is an indicator of the regulatory effect of each protein-coding gene on the identity of the cell; - That RegDBS x g and prioritizing each gene according to thereby identifying a cell identity gene for a cell of interest.

[0364] xxvi. The method of any one of statements i to xxv, wherein interactions between nodes of protein-coding gene products are selected for inclusion in the calculation of the network score if the experimental score in the protein-protein network is greater than zero and the combined score is greater than 500.

[0365] xxvii. The method of any one of statements i to xxvi, wherein determining the network score comprises removing protein-coding genes from the protein-protein interaction network that do not have associated peak H3K4me3 width.

[0366] xxviii. A method for determining factors necessary for the conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type, comprising: - determining the differential width of H3K4me3 modified regions for each protein-coding gene in the source cell and the target cell type to obtain a score of H3K4me3 modification (DBS) for each protein-coding gene in the source and target cell types; - calculating the difference between the DBS of the source cell and the DBS of the target cell to obtain a measure of the difference in H3K4me3 modification between the source cell and the target cell for each protein-coding gene (cell conversion ΔDBS); - determining a network score for transforming the source cell type into the target cell type based on the cell transformation ΔDBS and interactions between protein products of each protein-coding gene on at least one network, the network including information on interactions between each protein-coding gene product and other gene products in the cell; - determining a cell identity transformation score for each protein-coding gene in the target cell based on a combination of the cell transformation ΔDBS and the network score; - prioritizing each protein-coding gene according to its cell identity conversion score to identify cell identity genes for the target cell, each cell identity gene encoding a factor associated with the cell identity of the target cell; thereby determining the factors necessary for conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type.

[0367] xxix. A method for determining factors necessary for the conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type, comprising: - determining a differential breadth score (DBS) for each protein-coding gene in the source cells and the target cells, the DBS being based on the difference in width of H3K4me3 modified regions for all protein-coding genes compared to the median width of H3K4me3 modified regions for the same protein-coding gene in a population of different cell types; - calculating the difference between the DBS of the source cell and the DBS of the target cell to obtain a cellular conversion ΔDBS for each protein-coding gene; - determining a network score for cell transformation from the source cell type to the target cell type based on the cell transformation ΔDBS and interactions between protein products of each gene on at least one network, the network including information on interactions between the products of each protein-coding gene and other protein-coding gene products in the cell; - determining a cell identity conversion score (RegΔDBS) for each protein-coding gene in the target cell based on a combination of the cell conversion ΔDBS and the network score; - prioritizing each protein-coding gene according to its cellular transformation RegΔDBS to identify cellular identity genes for the target cell, each cellular identity gene encoding a factor associated with the cellular identity of the target cell; and thereby determining the factors necessary for conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type.

[0368] xxx.The differential width of the H3K4me3 modified region is - obtaining information about the width of H3K4me3 modified regions for each protein-coding gene in the source cells and the target cells to obtain a gene peak width score (B) for each gene in each cell type; - calculating the difference between the gene peak width score for each protein-coding gene in the cell type and the median gene peak width score for the same gene across populations of cells of different types, thereby determining a differential width score (DBS) for each gene in both the source and target cells; 20. The method of claim 19, wherein the method is determined by any one of statements xxviii to xxix.

[0369] xxxi. The width of the H3K4me3 modified region was determined based on ChIP-seq information. Preferably, the method of any one of statements xxvii to xxx, wherein the information is obtained from the ENCODE database.

[0370] xxxii. The method of any one of statements xxx to xxxi, wherein the step of determining the gene peak width score (B) comprises first defining regions of the genome in which significant H3K4me3 modifications (histone modification reference peak loci, or RPLs) are present.

[0371] xxxiii. The method of statement xxxii, wherein the RPL is obtained by merging the obtained ChIP-seq peak areas across all cell types, preferably the peak areas are overlapping peak areas.

[0372] xxxiv. The method of any one of statements xxx to xxxiii, wherein the step of determining a gene peak width score (B) for each protein-coding gene comprises excluding genes in which H3K4me3 and H3K27me3 modified regions are identified, and the genes are identified as poised genes.

[0373] xxxv. The method of any one of statements xxviii to xxxiv, wherein the H3K4me3 modification comprises an H3K4me3 modification. xxxvi. The method of any one of statements xxviii to xxxv, wherein the step of determining DBS for each protein-coding gene in the source cell and the target cell comprises combining information about the difference in H3K4me3 width for each protein-coding gene with the significance of that difference.

[0374] xxxvii. The method of statement xxxvi, wherein the step of determining DBS comprises multiplying, normalizing or adding the difference in peak width and the significance of the difference, preferably the DBS is determined by adding the difference in peak width and the significance of the difference.

[0375] xxxviii. The method of any one of statements xxviii to xxxvii, wherein the step of calculating the difference between the DBS of the source cell and the DBS of the target cell to obtain a cellular conversion ΔDBS for each gene comprises subtracting the DBS of the source cell for each protein-coding gene from the DBS of the target cell for each gene to obtain a cellular conversion ΔDBS for each gene, wherein the cellular conversion ΔDBS for each protein-coding gene provides a measure of the difference in width of the H3K4me3 modified region for a given protein-coding gene between the source cell and the target cell.

[0376] xxxix. The method of any one of statements xxviii to xxxviii, wherein the network score for cellular transformation from a source cell to a target cell is a weighted combination of cellular transformation ΔDBS of associated genes in the network.

[0377] xxxx. The method of statement xxxix, wherein the step of determining a network score for the cellular transformation includes combining information about the cellular transformation ΔDBS for each gene, the number of outdegree nodes of the gene, and the level of relatedness of the gene in the network.

[0378] xli. The method described in description xl, wherein the network score is determined by summing the cellular transformation ΔDBS of related genes, weighted for the number of gene outdegree nodes and the level of relatedness to obtain a weighted sum of cellular transformation ΔDBS for related genes in the network.

[0379] xlii. The protein-protein interaction network is a STRING database. 2. The method according to claim xxviii to xli. xliii. The method of any one of statements xxviii to xlii, wherein determining the cell transformation RegΔDBS comprises preferentially weighting protein-coding genes that encode regulatory factors.

[0380] xliv. The method of any one of statements xxviii to xlii, wherein the step of determining the cellular transformation RegΔDBS for each protein-coding gene comprises adding the cellular transformation ΔDBS score and the cellular transformation network score normalized across all protein-coding genes.

[0381] xlv. The method of clause xliv, wherein adding the cell transformation ΔDBS and the cell transformation network score further comprises weighting the cell transformation ΔDBS relative to the cell transformation network score by a factor of 2:1.

[0382] xlvi. The method of any one of statements xxviii to xlv, wherein prioritizing each protein-coding gene according to its Cell Conversion RegΔDBS comprises ordering the genes based on Cell Conversion RegΔDBS values.

[0383] xlvii. The method of any one of statements xxviii to xlvi, wherein the cellular transformation factor is selected from the group consisting of a receptor-ligand pair, a transcription factor, or an epigenetic remodeling factor.

[0384] xlviii. The method of any one of statements xxviii to xlvii, further comprising the step of selecting a subset of factors that are transcription factors, thereby identifying transcription factors required for conversion of the source cells into cells exhibiting at least one characteristic of the target cell type, preferably wherein the conversion is transdifferentiation of the differentiated source cells into differentiated target cells.

[0385] xlix. A method according to any one of statements xxviii to xlvii, comprising the step of selecting a subset of factors which are receptor-ligand pairs involved in cell signaling, thereby identifying signaling molecules required for the conversion of source cells into target cells, preferably wherein the identified factors are for the directed differentiation of pluripotent source cells into differentiated target cells.

[0386] l. The method of description xlix, wherein each protein-coding gene encoding a cell surface receptor is ranked according to its RegDBS score and each ligand associated with the receptor is ranked according to its DBS score to obtain a combined ranking of receptors and ligands, the combined ranking identifying a ligand for use in supplementing culture medium for maintaining a cell of interest in vitro.

[0387] 11. A method for determining factors necessary for the conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type, comprising: - determining an H3K4me3 gene peak width score (B) for each protein-coding gene (g) in a source cell (S) and a target cell (T) of interest, wherein the gene width score is the sum of the lengths of the regions in the promoters of each protein-coding gene that contain H3K4me3 modifications; - the normalized difference in gene peak width score between the source cell and the median of the gene peak width scores of a population of cells representing the background gene peak width score (ΔPeak Width S g ), and the significance of the difference (Pval Sg ) and - Gene widths for the target cells and a population of cells representing background gene peak width scores The normalized difference in gene peak width scores between the median scores (Δpeak width T g ), and the significance of the difference (Pval S g ) and - ΔPeak Width S g and Pval T g The values ​​of and are added to obtain the differential broadening score (DBS) for each protein-coding gene in the source cell. S g ) and - ΔPeak Width T g and Pval T g The values ​​of and are added to obtain the differential broadening score (DBS) for each protein-coding gene in the target cell. T g ) and - Differential Broad Score (DBS) for the same gene in the target cells T g ) to obtain the differential broadening score (DBS) for each protein-coding gene in the source cell. S g ) to obtain the cellular conversion ΔDBS (ΔDBS) for each protein-coding gene in the target cells. T-S g ) and - Cell-transformed differential broadening score (ΔDBS) of related genes (r) in the network T-S g ) to obtain a network score (Net) for each gene in the target cell. T-S g), wherein the network includes information of interactions between the products of each protein-coding gene in the cell, and the differential breadth score is corrected for the number of gene out-degree nodes (O) and the level of gene relatedness (L), and preferably is determined as a cell transformed differential breadth score (ΔDBS T-S g ) is calculated up to a third level of relevance; - Cellular transformation network score Net across all protein-coding genes T-S g normalizing - Cellular transformation differential broadening score (ΔDBS T-S g ) and cell transformation network score (Net T-S g ) to obtain a cell transformation regulatory differential broad score (RegΔDBS T-S g ) determining RegΔDBS T-S g is a measure of the difference in regulatory effect of each protein-coding gene on the target cell compared to the source cell; - its cell identity score, RegΔDBS T-S g prioritizing each protein-coding gene according to cell to identify cell identity genes for the target cell, each cell identity gene encoding a factor associated with the cell identity of the target cell; thereby identifying factors necessary for conversion of a source cell into a cell exhibiting at least one characteristic of a target cell type.

[0388] lii. The method of any one of statements xxviii to li, wherein interactions between nodes of protein-coding gene products in the protein-protein interaction network are selected if the experimental score in the network is greater than zero and the combined score is greater than 500.

[0389] liii. The method of any one of statements xxviii to lii, wherein the step of determining the network score comprises removing genes from the protein-protein interaction network that do not have associated peak H3K4me3 width.

[0390] liv. The method of statement xlviii, further comprising removing transcriptionally redundant TFs from the ranked list from each cell type. lv. A method for maintaining a population of cells in vitro, comprising: - providing a population of cells of interest in cell culture; - determining a differential breadth score (DBS) for each protein-coding gene in the cell of interest, the DBS being based on the difference in width of H3K4me3 modified regions for all protein-coding genes compared to the median width of H3K4me3 modified regions for the same gene in a population of different cell types; - Correlation between the protein products of each gene in DBS and at least one network determining a network score for each protein-coding gene in the cell of interest based on interactions, the network including information on interactions between the product of each protein-coding gene and other protein-coding gene products in the cell; - scoring each protein-coding gene in the cell of interest based on a combination of the DBS and the network score, thereby determining a RegDBS for each protein-coding gene in the cell of interest, wherein the RegDBS is an indication of the importance of each protein-coding gene with respect to cell identity; - prioritizing each protein-coding gene according to its RegDBS, thereby identifying cell identity genes for the cell of interest, wherein each cell identity gene encodes a factor associated with the cell identity of the cell of interest; and - contacting the population of cells of interest with one or more of the factors associated with the cellular identity of the cells of interest; - culturing the population of cells for a time and under conditions sufficient to allow maintenance of the cells of interest in cell culture; thereby maintaining a population of cells in vitro.

[0391] lvi. A method for maintaining a population of astrocytes in vitro, comprising the steps of, in order: - providing a population of astrocytes in cell culture; - contacting the population of astrocytes with one or more factors to maintain at least one characteristic of astrocytes; - culturing the population of astrocytes for a time and under conditions sufficient to allow maintenance of the astrocytes in cell culture; Including, the factor is selected from FN1, COL4A1, EDIL3, WNT5A, LAMB1, ADAM12 and COL1A2; A method thereby maintaining a population of astrocytes in vitro.

[0392] lvii. The method of statement lvii, wherein the factors comprise, consist of, or consist essentially of FN1, COL4A1, EDIL3, WNT5A, LAMB1, ADAM12, and COL1A2.

[0393] lviii. The method of statement lvvi or lvii, comprising contacting astrocytes with one or more, two or more, three or more, four or more, five or more of FN1, COL4A1, EDIL3, WNT5A, LAMB1, ADAM12 and COL1A2.

[0394] lix. A method according to any one of statements lvi to lix, further comprising administering astrocytes to the individual. lx. A method for maintaining a population of cardiomyocytes in vitro, comprising: - providing a population of cardiomyocytes in cell culture; - contacting the population of cardiomyocytes with one or more factors to maintain at least one characteristic of the cardiomyocytes; - culturing the population of cardiomyocytes for a time and under suitable conditions sufficient to maintain the cardiomyocytes in cell culture; Including, the factor is selected from FN1, COL3A1, TFPI, FGF7, APOE, C3, COL1A2, SERPINE1, COL6A3 and CXCL12; A method thereby maintaining a population of cardiomyocytes in vitro.

[0395] lxi. The method of statement lx, wherein the factors comprise, consist of, or consist essentially of FN1, COL3A1, TFPI, FGF7, APOE, C3, COL1A2, SERPINE1, COL6A3, and CXCL12.

[0396] lxii. The method of statement lx or lxi, comprising contacting cardiomyocytes with one or more of FN1, COL3A1, TFPI, FGF7, APOE, C3, COL1A2, SERPINE1, COL6A3, and CXCL12, optionally two or more, three or more, four or more, five or more, or six or more of FN1, COL3A1, TFPI, FGF7, APOE, C3, COL1A2, SERPINE1, COL6A3, CXCL12, preferably wherein the factors are FN1, COL3A1, TFP1, FGF7, and APOE.

[0397] lxiii. The method of any one of statements lx to lxii, further comprising administering cardiomyocytes, or populations of cells, to the individual. lxiv. A method for converting a source cell into a cell exhibiting at least one characteristic of a target cell type, comprising: - providing source cells; - determining the differential width of H3K4me3 modified regions for each protein-coding gene in the source cell and the target cell type to obtain a score of H3K4me3 modification (DBS) for each protein-coding gene in the source and target cell types; - calculating the difference between the DBS of the source cell and the DBS of the target cell to obtain a measure of the difference in H3K4me3 modification between the source cell and the target cell for each protein-coding gene (cell conversion ΔDBS); - determining a network score for the transformation of the source cell type into the target cell type based on the cell transformation ΔDBS and interactions between protein products of each protein-coding gene on at least one network, the network including information on interactions between each protein-coding gene product and other gene products in the cell; - determining a cell identity transformation score for each protein-coding gene in the target cell based on a combination of the cell transformation ΔDBS and the network score; - prioritizing each protein-coding gene according to its cell identity conversion score to identify cell identity genes for the target cell, wherein each cell identity gene encodes a factor associated with the cell identity of the target cell; - culturing the source cells for a time and under conditions sufficient to allow conversion of the source cells into cells exhibiting at least one characteristic of the target cell type; whereby converting the source cells into cells exhibiting at least one characteristic of the target cell type, and optionally increasing the amount of one or more factors, comprises i) contacting the source cells with the factor or an agent that increases expression of the factor by the cells, or ii) transfecting the source cells with a nucleic acid encoding the factor and expressing the nucleic acid in the cells.

[0398] lxv. A method for differentiating a source cell, comprising increasing protein expression of one or more factors or variants thereof in the source cell, wherein the source cell is differentiated to exhibit at least one characteristic of a target cell; - the source cells are pluripotent stem cells and the target cells are astrocytes; - the factor is one or more of FN1, COL4A1, COL1A2, EDIL3, ADAM12, LAMB1 and WNT5A.

[0399] lxvi. A method for generating cells exhibiting at least one characteristic of astrocytes from pluripotent stem cells or neural progenitor cells, comprising: - contacting the pluripotent stem cells or neural progenitor cells at the source cells with one or more of the factors FN1, COL4A1, COL1A2, EDIL3, ADAM12, LAMB1 and WNT5A, or variants thereof; - culturing the pluripotent stem cells or neural progenitor cells for a time and under conditions sufficient to allow differentiation into astrocytes, thereby generating cells from the pluripotent stem cells or neural progenitor cells that exhibit at least one characteristic of astrocytes; A method comprising:

[0400] lxvii. A method for differentiating pluripotent stem cells, preferably embryonic stem cells or neural progenitor cells, into cells exhibiting at least one characteristic of an astrocyte, comprising the steps of: i) providing pluripotent stem cells or a cell population comprising pluripotent stem cells or neural progenitor cells; ii) transfecting said pluripotent stem cells or neural progenitor cells with one or more nucleic acids comprising nucleotide sequences encoding one or more factors important for astroglial cellular identity; and iii) culturing said cells or cell population and optionally monitoring the cells or cell population for at least one astroglial cellular characteristic; The method, wherein the factors for differentiating pluripotent stem cells into astrocytes or for generating cells exhibiting at least one characteristic of astrocytes comprise, consist of, or consist essentially of FN1, COL4A1, COL1A2, EDIL3, ADAM12, LAMB1, and WNT5A.

[0401] lxviii. A method for generating cells exhibiting at least one characteristic of astrocytes from embryonic stem cells or neural progenitor cells, comprising: - increasing the amount of any one or more of FN1, COL4A1, COL1A2, EDIL3, ADAM12, LAMB1 and WNT5A, or variants thereof, in embryonic stem cells or neural progenitor cells; - culturing the embryonic stem cells for a time and under conditions sufficient to differentiate them into astrocytes, thereby generating cells from the embryonic stem cells or neural progenitor cells that exhibit at least one characteristic of astrocytes; A method comprising:

[0402] lxix. The method of statement lxviii, wherein increasing the amount of one or more of the factors comprises expressing a nucleic acid encoding one or more of the factors in the source cell.

[0403] lxx. The method of any one of statements lxiv to lxix, wherein at least one characteristic of astrocytes is upregulation of any one or more astrocyte markers and / or changes in cell morphology, preferably the astrocyte markers are selected from GFAP, S100B, ALDH1L1, CD44 and GLAST1.

Claims

1. An in vitro method for converting a source cell into a cell exhibiting at least one characteristic of a target cell, comprising: - determining the differential width of H3K4me3 modified regions for each protein-coding gene in the source and target cells to obtain a H3K4me3 modification score (DBS) for each protein-coding gene in said source and target cells; - calculating the difference between the DBS of the source cell and the DBS of the target cell to obtain a measure of the difference in H3K4me3 modification between the source and target cells for each protein-coding gene (ΔDBS); - determining a network score for the conversion of said source cell to said target cell based on interactions between said ΔDBS and protein products of each protein-coding gene on at least one network, said network including information on interactions between each protein-coding gene product and other gene products in the cell; - determining a cell identity transformation score (RegΔDBS) for each protein-coding gene in the target cell based on a combination of the ΔDBS and the network score; - prioritizing each protein-coding gene according to its RegΔDBS to identify cell identity conversion genes for said target cell, each cell identity conversion gene encoding a factor associated with the cell identity of said target cell; - increasing the amount of one or more factors associated with the cellular identity of said target cells in the source cells in the isolated cell culture; - culturing said source cells in said isolated cell culture for a time and under conditions sufficient to allow conversion of said source cells into cells exhibiting at least one characteristic of a target cell; thereby converting a source cell into a cell exhibiting at least one characteristic of a target cell.

2. 2. The method of claim 1, wherein increasing the amount of the one or more factors comprises: i) contacting the source cells with a factor or an agent that increases expression of the factor by the cells; or ii) transfecting the source cells with a nucleic acid encoding the factor and expressing the nucleic acid in the cells.

3. A computer-implemented method for determining factors necessary to convert a source cell into a cell exhibiting at least one characteristic of a target cell, comprising: - determining the differential widths of H3K4me3 modified regions for protein-coding genes in the source and target cells to obtain a score of H3K4me3 modification (DBS) for protein-coding genes in said source and target cells; - calculating the difference between the DBS of the source cell and the DBS of the target cell to obtain a measure of the difference in H3K4me3 modification between the source and target cells for each protein-coding gene (ΔDBS); - determining a network score for the transformation of said source cell to said target cell based on the interaction between said ΔDBS and protein products of each protein-coding gene on at least one network, said network containing information of interactions between protein-coding gene products in a cell; - determining a cell identity transformation score (RegΔDBS) for each protein-coding gene in the target cell based on a combination of the ΔDBS and the network score; - prioritizing each protein-coding gene according to its RegΔDBS to identify cell identity conversion genes for said target cell, each cell identity conversion gene encoding a factor associated with the cell identity of said target cell; thereby determining the factors necessary for conversion of said source cells into cells exhibiting at least one characteristic of a target cell.

4. The method according to any one of claims 1 to 3, wherein the conversion is differentiation, reprogramming or transdifferentiation.

5. The method of claim 4, wherein the source cells are differentiated cells, differentiating cells, undifferentiated cells or transdifferentiated cells.

6. The method according to any one of claims 1 to 5, the DBS in the source cell is based on the difference in width of the H3K4me3 modified region for each protein-coding gene in the source cell compared to the median, mean, or mode of the width of the H3K4me3 modified gene for the same protein-coding gene in a population of cell types different from the cell type of the source cell; and The DBS in the target cell is based on the difference in width of the H3K4me3 modified region for each protein-coding gene in the target cell compared to the median, mean, or mode of the width of the H3K4me3 modified gene for the same protein-coding gene in a population of cell types different from the cell type of the target cell. method.

7. 7. The method of claim 1, wherein the differential widths of the H3K4me3 modified regions for protein-coding genes in the source cell and the target cell, respectively, are: - obtaining information about the width of the H3K4me3 modified regions for protein-coding genes in each of the source and target cells, and obtaining gene peak width scores (B) for the protein-coding genes in each of the source and target cells; - calculating the difference (Δpeak width) between the gene peak width score of said protein-coding gene in each of said source and target cells and the average gene peak width score for the same gene across populations of cells of different types; is determined by method.

8. 8. The method of claim 7, wherein the H3K4me3 modification score (DBS) for each protein-coding gene in each of the source and target cells is determined by combining the calculated difference (Δ peak width) with an estimated significance of Δ peak width; the combining includes one or more of addition, multiplication, subtraction, and normalization; method.

9. 9. The method of claim 7 or 8, wherein the calculated difference (Δ peak width) is a normalized difference, and / or the H3K4me3 modification score (DBS) for each protein-coding gene in the source cell and the target cell, respectively, is calculated by multiplying the calculated difference (Δ peak width) by a -log of the estimated significance of the Δ peak width. 10 The method is determined by adding values.

10. 10. The method of any one of claims 7 to 9, wherein determining the gene peak width score (B) comprises first defining regions of the genome in which significant H3K4me3 modifications are present, and defining those regions as histone modification reference peak loci (RPLs).

11. 11. The method of claim 10, wherein the RPL is obtained by merging overlapping ChIP-seq peak areas obtained across all cell types.

12. 12. The method of any one of claims 7 to 11, wherein obtaining the gene peak width score (B) for each protein-coding gene comprises excluding genes in which H3K4me3 and H3K27me3 modified regions are identified, and / or the gene peak width score (B) for a protein-coding gene is calculated as the sum of the peak widths of one or more peaks annotated to the protein-coding gene.

13. 13. The method of any one of claims 1 to 12, wherein prioritizing each protein-coding gene according to its RegΔDBS comprises ranking each protein-coding gene based on its RegΔDBS.

14. 14. The method of any one of claims 1 to 13, wherein identifying a cell identity conversion gene for the target cell comprises selecting the protein-coding gene that encodes a transcription factor, wherein the conversion may be transdifferentiation of a differentiated source cell to a differentiated target cell.

15. 14. The method of any one of claims 1 to 13, wherein identifying cell identity conversion genes for the target cell comprises selecting the protein-coding genes that encode a receptor-ligand pair, wherein the conversion may be direct differentiation from a pluripotent source cell to a differentiated target cell, and / or the method comprises ranking each gene encoding a cell surface receptor according to its RegΔDBS and prioritizing ligands associated with the receptor based on the RegΔDBS for each ligand to obtain a combined ranking of receptors and ligands.

16. 16. The method of any one of claims 1 to 15, wherein the cell identity transformed score (RegΔDBS) comprises the sum of the ΔDBS and the network score, or the cell identity transformed score (RegΔDBS) comprises a weighted sum of the ΔDBS and the network score, wherein the ΔDBS is weighted by a factor of 2 relative to the network score.

17. 17. The method of any one of claims 1 to 16, wherein determining the network score for a protein-coding gene comprises combining the ΔDBSs of genes related to the protein-coding gene in the network using a weighted sum, wherein the genes may be weighted by their number of outdegree nodes and their level of relatedness to the protein-coding gene; and / or wherein determining the network score for a protein-coding gene comprises dividing the network score for a protein-coding gene by the maximum network score across all protein-coding genes in the network.

18. The method according to any one of claims 1 to 3, - determining the differential widths of H3K4me3 modified regions for protein-coding genes in source cells and target cells and obtaining a score of H3K4me3 modification (DBS) for said protein-coding genes in said source cells and target cells, comprises, for protein-coding genes: determining an H3K4me3 gene peak width score (B) for a protein-coding gene (g) in a source cell (S) and a target cell (T), wherein said H3K4me3 gene peak width score (B) is the sum of the lengths of regions in the promoters of said protein-coding genes that contain H3K4me3 modifications; The normalized difference in gene peak width score (ΔPeak Width) between the source cell and the median, mean, or mode of the gene peak width score of a population of cells exhibiting a background gene peak width score. S g ), and the significance of the difference (Pval S g ) determining the The normalized difference in gene peak width score (ΔPeak Width) between the target cell and the median, mean, or mode of the gene width score of a population of cells representing the background gene peak width score. T g ), and the significance of the difference (Pval T g ) determining the The Δ peak width S g and Pval S g and the differential broad score (DBS) for the protein-coding gene in the source cell S g ) obtaining the The Δ peak width T g and Pval T g and the differential broad score (DBS) for the protein-coding gene in the target cell. T g ) obtaining the Including, - calculating the difference (ΔDBS) between the DBS of the source cell and the DBS of the target cell is calculated by subtracting the differential broadening score (DBT) for the same gene in the target cell T g ) to calculate the differential broad score (DBS) for each protein-coding gene in the source cell. S g ) to obtain the cellular conversion ΔDBS for each protein-coding gene; - determining a network score for the transformation of said source cell into said target cell: A network score (N) for each protein-coding gene in the target cell is obtained by combining the cellular transformations ΔDBS of related genes in the network. T-S g ), wherein the cellular transformation ΔDBS is corrected for the number of outdegree nodes of the gene and the level of relatedness of the gene, and a weighted sum of the ΔDBS may be calculated up to a third level of relatedness; and The network score N across all protein-coding genes T-S g normalizing ; Including, and - determining a cell identity conversion score (RegΔDBS) by calculating the cell identity conversion ΔDBS and N for each protein-coding gene; T-S g wherein said RegΔDBS is a measure of the differential regulatory effect of each protein-coding gene on the target cell compared to the source cell. method.

19. 19. The method of any one of claims 1 to 18, wherein determining the network score comprises removing protein-coding genes from the protein-protein interaction network that do not have an associated peak H3K4me3 width and / or have a peak width below a predetermined width threshold.

20. at least one processing unit; and a memory storing instructions that, when executed by said at least one processing unit, cause said processing unit to perform the method of claim 3; 1. A computer processing system comprising:

21. A non-volatile computer readable medium storing instructions that, when executed, cause a processor to perform the method of claim 3.