Methods for preventing rapid gene silencing in pluripotent stem cells - Patents.com

JP2024520413A5Pending Publication Date: 2025-05-26CELLULAR DYNAMICS INTERNATIONAL +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023572725
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-05-26
Filing Date
2022-05-26
Publication Date
2025-05-26

AI Technical Summary

Technical Problem

Existing methods fail to address clone-to-clone and batch-to-batch variability in induced pluripotent stem cells (iPSCs) due to rapid silencing of transgenes, which is attributed to epigenetic modifications, leading to inconsistent gene expression and differentiation performance.

Method used

Engineering iPSCs to express transgenes with optimized codons that remove CpG motifs and utilizing novel promoters such as HSP90AB1, ACTB, CTNNB1, MYL6, UBA52, CAG, and UBC, along with gene editing techniques like CRISPR/CAS, to stabilize gene expression.

Benefits of technology

Stabilizes transgene expression in iPSCs for extended periods, allowing consistent differentiation into specific cell types like hematopoietic progenitor cells, neural progenitor cells, and endothelial cells, maintaining fluorescence markers for over a year.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000071_0000
    Figure 00000071_0000
  • Figure 00000071_0001
    Figure 00000071_0001
  • Figure 00000071_0002
    Figure 00000071_0002
Patent Text Reader

Abstract

Provided herein are methods for generating cell lines with stable expression of a transgene by removal of CpG motifs. In further methods, methods are provided for cell lines with stable expression of a transgene by driving expression with a novel promoter or by tagging an endogenous gene.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Claiming priority This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 193,472, filed May 26, 2021, the contents of which are incorporated herein by reference.

[0002] Inclusion of sequence listing The sequence listing contained in the file entitled "CDINP0103WO_ST25.TXT", which is 34,000 bytes (as measured in Microsoft Windows®) and was created on May 26, 2022, is submitted herewith by electronic submission and is incorporated herein by reference. [Background technology]

[0003] 1. Field The present disclosure relates generally to the field of stem cell biology. More specifically, the present invention relates to a method for codon optimization of genes in induced pluripotent stem cells to reduce rapid gene silencing.

[0004] 2. Description of Related Art Studies have shown that seemingly identical cell lines can vary widely in their performance (Kyttala, 2016). These differences, detected when comparing multiple clones derived from the same donor, are referred to as "clonal variability." These clones are believed, and in some cases have been confirmed, to contain identical DNA sequences. The inconsistent yield and purity among differentiation batches of the same cell line is referred to as "batch-to-batch variability." In many cases, differences in differentiation performance are due to epigenetic modifications, but there is an unmet need to identify specific epigenetic mechanisms and ways to alter these epigenetic mechanisms to prevent cell line variability. Summary of the Invention [Means for solving the problem]

[0005] In a first embodiment, the disclosure provides an isolated cell line engineered to express at least one transgene, the at least one transgene being (a) under the control of a promoter having at least 90% (e.g., at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity to SEQ ID NOs: 1-12 or 17, (b) under the control of an endogenous gene selected from the group consisting of HSP90AB1, ACTB, CTNNB1, MYL6, UBA52, CAG, RPS, and UBC, and / or (c) encoded by a sequence modified to remove CpG motifs to provide stable expression. In certain aspects, the cell line is an induced pluripotent stem cell (iPSC) line.

[0006] In some aspects, the sequence modified to remove CpG motifs to provide stable expression has at least 90% (e.g., at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity to SEQ ID NO: 14 or SEQ ID NO: 16. In certain aspects, the sequence modified to remove CpG motifs to provide stable expression is SEQ ID NO: 14 or SEQ ID NO: 16.

[0007] In some aspects, at least one transgene is (a) under the control of a promoter having at least 90% (e.g., at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity to SEQ ID NOs: 1-12 or 17, and / or (b) under the control of an endogenous gene selected from the group consisting of HSP90AB1, ACTB, CTNNB1, MYL6, UBA52, CAG, RPS, and UBC. In certain aspects, at least one transgene is encoded by a sequence modified to remove CpG motifs to provide stable expression.

[0008] In certain embodiments, at least one transgene is encoded by a sequence modified to remove CpG motifs to provide stable expression and is under the control of a promoter having at least 90% (e.g., at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity to SEQ ID NO: 1-12 or 17. In some embodiments, at least one transgene is encoded by a sequence modified to remove CpG motifs to provide stable expression and is under the control of an endogenous gene selected from the group consisting of HSP90AB1, ACTB, CTNNB1, MYL6, UBA52, CAG, RPS, and UBC. In certain embodiments, at least one transgene is encoded by a sequence modified to remove CpG motifs to provide stable expression and is under the control of an endogenous gene selected from the group consisting of HSP90AB1, ACTB, CTNNB1, and MYL6.

[0009] In further aspects, the cell line is engineered to express at least a first transgene and a second transgene. In some aspects, the first transgene is under the control of a promoter having at least 90% (e.g., at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity to SEQ ID NOs: 1-12 or 17, and the second transgene is under the control of an endogenous gene selected from the group consisting of HSP90AB1, ACTB, CTNNB1, MYL6, UBA52, CAG, RPS, and UBC. In other aspects, the first transgene is under the control of a promoter having at least 90% (e.g., at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity to SEQ ID NO: 1-12 or 17, and the second transgene is under the control of an endogenous gene selected from the group consisting of HSP90AB1, ACTB, CTNNB1, and MYL6. In some aspects, the first transgene and / or the second transgene are encoded by a sequence modified to remove CpG motifs to provide stable expression. In certain aspects, at least 50 percent, e.g., at least 70 percent, 80 percent, 90 percent, 95 percent, 96 percent, 97 percent, 98 percent, or 99 percent of the CpG motifs are removed. In certain aspects, all CpG motifs are removed. In some aspects, the codons of the CpG motifs are replaced with codons that are not rare and / or do not result in mononucleotide stretches. In certain embodiments, the codons of the CpG motifs are replaced with the corresponding codons of Table 1.

[0010] In some aspects, the promoter is a response element. In certain aspects, the promoter is driven by a response element.

[0011] In some aspects, the transgene is a reporter gene or a selectable marker. In certain aspects, the reporter gene is a fluorescent or luminescent protein, such as luciferase, green fluorescent protein (GFP) or red fluorescent protein (RFP). In certain aspects, the at least one transgene is a selectable marker, such as puromycin, neomycin, or blasticidin. In certain aspects, the at least one transgene is a suicide gene. In some aspects, the at least one transgene is thymidine kinase, TET, or myoblast determination protein 1 (MYOD1).

[0012] In certain embodiments, the cell line has stable expression of the transgene for at least 30 days, e.g., at least 2 months, 3 months, 4 months, 5 months or longer. In certain embodiments, the cell line has stable expression of the transgene for more than 6 months, e.g., more than 1 year, more than 2 years, or more than 3 years.

[0013] In some aspects, the at least one transgene is encoded by an expression cassette. In certain aspects, the at least one transgene is introduced into the cell line by electroporation or lipofection. In certain aspects, the expression cassette is inserted into a genomic safe harbor site, such as the PPP1R12C (AAVS1) locus or the ROSA locus.

[0014] In certain aspects, the promoter has at least 90% (e.g., at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity to SEQ ID NO: 2, 3, 4, 6, or 17. In some aspects, the promoter comprises SEQ ID NO: 2, 3, 4, 6, or 17.

[0015] In certain embodiments, the method comprises gene editing, and in particular, the transgene comprises gene editing, e.g., TALEN-mediated gene editing, CRISPR-mediated gene editing, or ZFN-mediated gene editing.

[0016] A further embodiment provides a method for preventing silencing of expression of a transgene in an engineered cell comprising optimizing the transgene sequence to remove CpG motifs.

[0017] In some embodiments, optimizing comprises replacing the codons of essentially all CpG motifs. In certain embodiments, optimizing comprises replacing at least 50 percent of the CpG motifs, such as at least 70 percent, 80 percent, 90 percent, 95 percent, 96 percent, 97 percent, 98 percent, or 99 percent. In certain embodiments, all CpG motifs are removed. In certain embodiments, the codons of the CpG motifs are replaced with codons that are not rare and / or do not result in mononucleotide stretches. In some embodiments, the codons of the CpG motifs are replaced with the corresponding codons in Table 1. In certain embodiments, the transgene sequence optimized to remove CpG motifs comprises a percentage of GC content substantially similar to the percentage of GC content of the wild-type transgene sequence.

[0018] In some aspects, the sequence of the transgene is that of a reporter gene, such as a fluorescent protein, e.g., GFP or RFP.

[0019] In certain aspects, the transgene is under the control of a constitutive promoter. In some aspects, the constitutive promoter has expression in substantially all cell types. In certain aspects, the constitutive promoter has expression in essentially all cell types. In certain aspects, the constitutive promoter has expression in all cell types.

[0020] In certain aspects, the transgene is under the control of an inducible promoter, hi some aspects, the transgene is under the control of the EEF1A1 promoter.

[0021] In a further aspect, the method further comprises treating the cell line with sodium butyrate, VPA, or TSA, hi a specific aspect, sodium butyrate is added at a concentration of 0.25 mM to 0.5 mM.

[0022] In some embodiments, the cell line is an iPSC line. In certain embodiments, the method further comprises differentiating the iPSC line. In some embodiments, the iPSC line is differentiated into a mature cell, such as, but not limited to, a hematopoietic progenitor cell, a neural progenitor cell, a GABAergic neuron, a macrophage, a microglia, or an endothelial cell.

[0023] Another embodiment provides an expression vector comprising a promoter having at least 90% (e.g., at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity to SEQ ID NO: 1-12 or 17. In some aspects, the promoter has at least 90% (e.g., at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity to SEQ ID NO: 2, 3, 4, 6, or 17. In certain aspects, the promoter comprises SEQ ID NO: 2, 3, 4, 6, or 17. In certain aspects, the expression vector is a pGL3 plasmid vector. In some aspects, the vector encodes a transgene under the control of the promoter. In certain aspects, the transgene is a reporter gene, e.g., a fluorescent protein or a luminescent protein, e.g., luciferase, green fluorescent protein (GFP) or red fluorescent protein (RFP).

[0024] Further embodiments provide methods of generating a cell line with stable transgene expression, comprising engineering the cell line to express a vector of embodiments of the invention (e.g., comprising a promoter having at least 90% (e.g., at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%) sequence identity to SEQ ID NOs: 1-12 or 17), wherein the vector encodes the transgene. In some aspects, the cell line is a pluripotent cell line, e.g., an iPSC line.

[0025] In some embodiments, the method includes integrating the vector into the AAVS1 locus on chromosome 19. In certain embodiments, the integrating includes gene editing, e.g., CRISPR-mediated gene editing, TALEN-mediated gene editing, or ZFN-mediated gene editing.

[0026] In further embodiments, the method further comprises differentiating the cell line. In some embodiments, the cell line is differentiated into hematopoietic progenitor cells, neural progenitor cells, GABAergic neurons, macrophages, microglia, or endothelial cells. In certain embodiments, the cell line is cultured for at least 30 days, e.g., at least 2 months, 3 months, 4 months, 5 months, or longer. In certain embodiments, the cell line is cultured for more than 6 months, e.g., more than 1 year, more than 2 years, or more than 3 years. In certain embodiments, the cell line has stable expression of the transgene for at least 30 days, e.g., at least 2 months, 3 months, 4 months, 5 months, or longer. In certain embodiments, the cell line has stable expression of the transgene for more than 6 months, e.g., more than 1 year, more than 2 years, or more than 3 years. In some embodiments, the cell line is cultured for at least 6 months. In certain embodiments, the cell line has stable expression of the transgene for 6 months.

[0027] Another embodiment provides an isolated pluripotent cell line comprising an expression vector of this embodiment (e.g., comprising a promoter having at least 90% (e.g., at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99%) sequence identity to SEQ ID NO:1 to 12 or 17).

[0028] Further embodiments provide a method of generating a cell line with stable expression of an exogenous transgene, comprising engineering the cell line to express the transgene under the control of an endogenous gene, wherein the endogenous gene is HSP90AB1, ACTB, CTNNB1, MYL6, UBA52, CAG, RPS, and UBC, e.g., HSP90AB1, ACTB, CTNNB1, or MYL6.

[0029] In some embodiments, the manipulating comprises gene editing, e.g., TALEN-mediated gene editing, CRISPR-mediated gene editing, or ZFN-mediated gene editing. In some embodiments, the transgene is a reporter gene, a selection marker, or a suicide gene.

[0030] In certain embodiments, the cell line is a pluripotent cell line, such as an iPSC line.

[0031] Another embodiment provides an isolated cell line comprising endogenous HSP90AB1, ACTB, CTNNB1, MYL6, UBA52, CAG, RPS, and UBC tagged with a transgene. In some aspects, the transgene is a reporter gene, a selectable marker, or a suicide gene. In certain aspects, the cell line is a pluripotent cell line, e.g., an iPSC line.

[0032] Further provided herein is an assay for detecting cells, comprising culturing the cell line of the present invention embodiment and measuring the expression of a reporter gene. Also provided herein is the use of the cell line of the present invention embodiment for a cellular assay, such as a cell viability assay, or an assay for screening a candidate drug. In some aspects, the assay is a high-throughput assay. In certain aspects, the cellular assay comprises measuring the expression of a reporter gene.

[0033] Another embodiment provides a composition comprising the cell line of the present embodiments for use in a cellular assay.

[0034] Other objects, features and advantages of the present disclosure will become apparent from the following detailed description. It should be understood, however, that the detailed description and specific examples, while indicating preferred embodiments of the present invention, are given by way of illustration only, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description.

[0035] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the invention. The invention may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein. [Brief description of the drawings]

[0036] [Figure 1A-1D] Expression of ZsGreen driven by EEF1A1p is patchy in iPSCs. iPSCs 01278.103 (Figure 1A, 1B) and 01279.107 (Figure 1C, 1D) were engineered with EEF1A1p-ZsGreen at the PPP1R12C locus. Bright-field and fluorescent (GFP) microscopy were used to capture GFP expression in cells at passage 11 post-manipulation (Figure 1A-1B) or passage 18 post-manipulation (Figure 1C-1D).

[0037] [Diagram 2] Codon optimization of the AcGFP1 DNA sequence (SEQ ID NO: 13) resulted in a CpG-free AcGFP1 DNA sequence (SEQ ID NO: 14).

[0038] [Figure 3A-3B] Expression of CpG-free AcGFP1 is stable, but AcGFP1 expression is not maintained over time. Percent GFP expression was monitored over time in five clones targeted with EEF1A1p-CpG-free AcGFP1 (Figure 3A) and nine clones targeted with EEF1A1p-AcGFP1 at the AAVS1 locus (Figure 3B).

[0039] [Figure 4A-4C] Rapid silencing of AcGFP1 in iPSCs. Depiction of iPSC engineering at PPP1R12C locus (AAVS1 safe harbor) with three cassettes, EEF1A1p-mRFP1+PGKp-Puro, EEF1A1p-AcGFP1 or EEF1A1p-CpG-free AcGFP1. Marked in the figure are the ID numbers of engineered iPSCs for cell lines 8717 and 9650, which are used in further experiments throughout this document (Figure 4A). AcGFP1 expressing clones were picked and expanded but did not retain consistent expression. After 2 months of culture, AcGFP1 engineered iPSCs were bulk sorted for AcGFP1 expression. Brightfield and fluorescent (GFP) microscopy were used to capture GFP expression in cells 12 days after sorting (Figure 4B) or 23 days after sorting (Figure 4C). Similar silencing was observed with other green fluorescent proteins, including monomeric mNeonGreen and tetrameric ZsGreen.

[0040] [Figure 5A-5B]The silenced transgene is reactivated by NaBut treatment. In CpG-free AcGPF1 cultures, a small number of cells were silenced (less than 3% of cells, Figure 3A). These silenced cells were sorted and expanded cells to further investigate their silencing and to study ways to overcome the silencing. Two months after sorting for no GFP expression, the silenced CpG-free AcGFP1 clones were treated with 1 mM, 0.5 mM, or 1 μM NaBut. Nine days after NaBut treatment, the cells were assayed for % GFP expression by flow cytometry, and a dose-dependent reactivation of CpG-free AcGFP1 was observed. After successful pilot experiments, the NaBut treatment period was extended to 46 days with NaBut treatment doses of 0.25 mM and 0.5 mM, and GFP expression levels were monitored over time by fluorescence microscopy and flow cytometry. The initial results were confirmed (Figures 5A and 5B, dark blue bars: 8 days of treatment). A dose-dependent effect of NaBut treatment was evident throughout the experimental period (FIGS. 5A and 5B).

[0041] [Figure 6] Differentiation of iPSC 9650 (AAVS1 CpG-less AcGFP1). iPSC 9650 maintained GFP expression throughout hepatocyte differentiation (measured by CXCR4, AAT and ALB expression) and induced neuron (iN) differentiation (measured by TUJ expression).

[0042] [Figure 7A-7C] 1069: WT PuroR (Fig. 7A), 1362: CpG-free PuroR1 (Fig. 7B) and 1363: CpG-free PuroR1 (Fig. 7C) plasmids.

[0043] [Figure 8] Schematic description of the protocol for generating endothelial cells from iPSC 9650-GFP engineered with AcGFP1, which contains no CpG in AAVS1.

[0044] [Figure 9] Hypoxic acclimated iPSCs were seeded onto Purecoat Amine plates to initiate blood endothelial cell generation for 6 days. Representative photographs of iPSC-derived blood endothelial cells on day 6 of differentiation revealed the presence of blood endothelial colonies in a two-dimensional format that retained GFP expression.

[0045] [Figure 10] Morphology of endothelial cells derived from 9650-GFP at passage 2 in culture using a 4x objective to reveal GFP / BF overlap.

[0046] [Figure 11] Purity of endothelial cells derived from 9650-GFP iPSCs. Hypoxia-acclimated iPSCs were plated on Purecoat Amine plates to initiate blood endothelial cell generation and then replated to generate pure endothelial cells capable of expansion over multiple passages. Endothelial cell purity was quantified at the time of passaging by staining for co-expression of CD31, CD144 and CD105 by flow cytometry.

[0047] [Figure 12] Hypoxia-acclimated iPSCs were seeded onto Purecoat Amine plates to initiate the generation of blood endothelial cells and then replated to generate pure endothelial cells that could be expanded for multiple passages. The intensity of GFP expression was quantified by flow cytometry over multiple passages.

[0048] [Figure 13] Schematic description of the protocol for generating hematopoietic progenitor cells (HPCs) from iPSCs.

[0049] [Figures 14A-14C]Hypoxia-acclimated iPSCs were harvested and differentiated into HPCs in a 3D aggregate format for 13-15 days. At the end of the HPC differentiation process, cells were harvested and stained for CD34, CD45, CD31, CD41, and CD235 expression along with GFP (Figure 14A) or RFP (Figure 14C) expression to quantify HPC purity, demonstrating retention of fluorescence in end-stage HPCs. Co-expression of GFP and CD34 after MACS separation is greater than 90% (Figure 14B).

[0050] [Figure 15] HPC generation efficiency: From one input iPSC, 8717 and 9650 generated 0.766 and 0.225 HPCs, respectively.

[0051] [Figure 16] Schematic diagram of microglia generation from HPCs.

[0052] [Figure 17A-17B] Phase and fluorescence images of microglial differentiation of the 9650-GFP (Figure 17A) and 8717-RFP (Figure 17B) lines.

[0053] [Figure 18] Efficiency of hematopoietic progenitor cell (HPC) generation. CD34+ MAC-sorted 9650-GFP-derived HPCs and unsorted 8717-RFP-derived HPCs were differentiated into microglia. Total viable numbers of input HPCs and output microglia were quantified. Process efficiency was calculated based on the purity and absolute number of CD34+ positive cells present at day 23 of microglial differentiation divided by the absolute number of viable input HPCs.

[0054] [Figures 19A-19D]Purity profile of day 23 microglia generated from 8717-RFP (Figure 19A) and 9650-GFP (Figure 19C) iPSCs, respectively. End-stage microglia were harvested and stained for the presence of PU.1, IBA, CX3CR, TREM2 and P2RY12 expression and quantified by flow cytometry. Co-expression of markers was quantified along with retention of GFP or RFP in end-stage cells (Figures 19B, 19D).

[0055] [Figure 20] Schematic diagram of end-stage macrophage generation from HPCs.

[0056] [Figure 21] 8717-RFP-derived HPCs were further differentiated to generate end-stage macrophages, whose purity assessment was quantified by staining for the presence of CD68 expression at days 44 and 51 of the differentiation process.

[0057] [Figure 22] Phase and fluorescence images of the 8717-RFP strain at different days of the macrophage differentiation process. Images were captured at 10x magnification.

[0058] [Figure 23] HPCs derived from 8717-RFP iPSCs were differentiated into end-stage macrophages. The total viable numbers of input HPCs and output macrophages were quantified. Process efficiency was calculated based on the purity and absolute number of CD68+ positive cells present at day 51 of macrophage differentiation divided by the absolute number of viable input HPCs.

[0059] [Figure 24] Retention of the presence of engineered fluorescent dyes throughout the differentiation process. 9650-GFP and 8717-RFP iPSCs retained the presence of fluorescent dyes throughout the differentiation of iPSCs to HPCs and through to the generation of pure end-stage microglia and macrophages.

[0060] [Diagram 25] Schematic description of the method for generating neural progenitor cells (NPCs) from iPSCs without the use of dual SMAD inhibition, illustrating the different steps involved and the composition of the media used.

[0061] [Figure 26A-26B] (FIG. 26A) Visualization of red and green fluorescence during the 2D preconditioning stage of the NPC differentiation process. FIG. 26B captures the fluorescence of the 3D NPC culture at the final stage before harvesting. All images were taken with a 4x objective.

[0062] [Figure 27] Quantification of post-thaw purity in 8717-RFP and 9650-GFP derived NPCs. NPCs were thawed and stained for the presence of SSEA4, CD56 and CD15 expression using relevant isotype controls.

[0063] [Figure 28] Differentiation protocol of NPCs into GABAergic neurons. NPCs were subjected to 3D differentiation culture and transitioned to 2D culture on PLO-laminin coated plates. End-stage neurons were harvested on day 18 and quantified for purity of nestin and β-tubulin 3 by flow cytometry.

[0064] [Fig. 29A-28B] Brightfield and fluorescent images taken on day 2 (3D) (FIG. 29A) and day 18 (2D) (FIG. 29B) of GABAergic neuron differentiation. 3D cultures in ULA T25 Flasks and 2D cultures in 6-well PLO-laminin coated plates. All images were taken at 10x magnification.

[0065] [Diagram 30]Retention of GFP and RFP expression in undifferentiated engineered iPSCs and end-stage neuronal cultures at days 13 and 18 of GABAergic neuronal differentiation. Day 13 samples were stained prior to plating on PLO-laminin, and day 18 cultures were stained at the end of GABAergic neuronal differentiation.

[0066] [Diagram 31] GABAergic neurons from 9650-GFP and 8717-RFP iPSC cultures at day 18 of differentiation were harvested and stained for nestin and β-tubulin purity by flow cytometry. Co-expression of GFP or RFP with nestin and tubulin in end-stage cultures was quantified.

[0067] [Fig. 32A-32B] (Figure 32A) Normalized luciferase (firefly / renilla ratio, normalized to EEF1A1=100%) is shown (HSP90AB1del400 promoter and HSP90AB1 promoter expressed approximately 66% and 75% of EEF1A1). (Figure 32B) Plasmid design using CAG promoter as an example to control ZsGreen fluorescent protein and targeting AAVS1 (PPP1R12C) safe harbor locus on chromosome 19 of human iPSCs.

[0068] [Diagram 33] Engineered iPSC lines expressing ZsGreen (ZsG) fluorescent protein were maintained in culture for up to 7 months (E8 medium / vitronectin-coated plates) and periodically checked for green expression using flow cytometry on an Accuri C6 instrument (BD). Most clones maintained consistent flow profiles over time, except for one RPS19 promoter clone (5363), which showed a drop in fluorescence in many cells at August. Graphs show median fluorescence levels normalized to unmanipulated iPSC=1.

[0069] [Fig. 34A-34B](FIG. 34A) Flow cytometry plots of iPSC lines engineered with ZsGreen (ZsG). (FIG. 34B) At day 21 of differentiation, all cells had a visible neuronal phenotype. Flow cytometry shows that the fluorescence of CAG, UBC(v1), and HSP90AB1del400 promoters was diminished in many cells. UBCv2, UBA52, and RPS19 promoters showed tight and stable expression, as did the tag genes HSP90AB1, CTNNB1, and MYL6. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0070] DNA methylation plays an important role in regulating gene expression, including induction of transcriptional repression, prevention of transcription factor binding to DNA, requirement for binding of some transcription factors to DNA, recruitment of HDAC complexes, inactivation of the X chromosome, and immunogenicity of CpG motifs such as TLR9. DNA methylation in mammals occurs when a methyl group is added to the fifth carbon (5-mC) of cytosine in cytosine phosphate guanine (CpG) by methyltransferases. DNMT3A and DNMT3B (DNA methyltransferases) are responsible for de novo methylation (i.e., methylating previously unmethylated DNA), and DNMT3B has been shown to be turned on in iPSCs. DNMT1 is responsible for methylation of hemimethylated DNA after replication and is characterized as a maintenance methyltransferase. Studies of demethylation have emerged more recently, identifying Gadd45a as a key player in DNA demethylation in DNA repair, and TET and TDG as key players in the oxidation and excision of 5-mC in DNA.

[0071] The addition of transgenes by genomic engineering into iPSCs provides the opportunity to monitor transgene expression over time and through the differentiation process. Green fluorescent protein (GFP) and red fluorescent protein (RFP) have been widely used to generate fusion proteins without significantly interfering with the assembly and function of the native proteins, making them powerful tools for in vivo analysis and biomarkers to monitor progenitor populations and determine the dynamics of emerging cell lineages. As evidenced by the lack of commercially available iPSCs or differentiated cells expressing green fluorescent protein (GFP), there is a need for methods to maintain transgene expression over extensive passaging and differentiation. Specifically, Figure 1 shows the patchy expression of GFP in iPSCs.

[0072] Thus, in certain embodiments, the present disclosure provides a method for maintaining transgene expression in cell lines by optimizing the sequence of the transgene to remove CpG motifs and thus prevent rapid silencing of the transgene. Methylation is a major epigenetic mechanism in addition to RNA-associated silencing and histone modifications. In this study, the DNA sequence of Aequorea coerulescens green fluorescent protein (AcGFP1) was modified to remove CpG motifs as shown in Figure 2. As a result, expression of CpG-free AcGFP1 was stable, but expression of wild-type AcGFP1 was not stable (Figure 3). Thus, the present method allows for the prevention of transgene silencing by global methylation or other epigenetic dysregulation.

[0073] In further embodiments, methods are provided for maintaining expression of a transgene in a cell line by driving expression of the transgene with a novel promoter provided herein (e.g., SEQ ID NOs: 1-12 or 17), or by tagging a gene such as HSP90AB1, ACTB, CTNNB1, or MYL6.

[0074] Furthermore, the cell lines of the invention can be differentiated into specific cell types and maintain transgene expression for more than 3 months, 6 months, or even 12 months. In certain embodiments, the cell lines are cultured for at least 30 days, e.g., at least 2 months, 3 months, 4 months, 5 months, or longer. In certain embodiments, the cell lines are cultured for more than 6 months, e.g., more than 1 year, more than 2 years, or more than 3 years. In certain embodiments, the cell lines have stable expression of the transgene for at least 30 days, e.g., at least 2 months, 3 months, 4 months, 5 months, or longer. In certain embodiments, the cell lines have stable expression of the transgene for more than 6 months, e.g., more than 1 year, more than 2 years, or more than 3 years. In additional embodiments, cellular assay methods are provided that use the cell lines of the invention for cell viability assays and screening assays.

[0075] I. Definition The term "purified" does not require absolute purity, but rather is intended as a relative term. Thus, a purified cell population is greater than about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% pure, or most preferably, essentially free of other cell types.

[0076] As used herein, the term "stable expression" refers to expression that is more stable than the unmodified sequence. For example, stable expression can refer to expression that remains unchanged for a period of one month, six months, one year, or more than one year.

[0077] As used herein, "essentially free" with respect to a particular component is used herein to mean that none of the particular components are intentionally incorporated into the composition and / or are only present as contaminants or in trace amounts.Thus, the total amount of the particular component resulting from unintentional contamination of the composition is less than 0.05%, preferably less than 0.01%.Most preferred are compositions in which the amount of the particular component is undetectable by standard analytical methods.

[0078] As used herein, "a" or "an" may mean one or more. When used in the claims, when used in conjunction with the word "comprising," the words "a" or "an" may mean one or more than one.

[0079] Use of the term "or" in the claims is used to mean "and / or," unless expressly indicated to refer to alternatives only or the alternatives are not mutually exclusive, however, the present disclosure supports a definition that refers only to alternatives and "and / or." As used herein, "another" may mean at least a second or more.

[0080] The term "essentially" should be understood to include only those steps or materials specified, and which do not materially affect the basic and novel characteristics of these methods and compositions.

[0081] The term "substantially free" is used for 98% of the listed components and for compositions or particles that are substantially free of less than 2% of the component.

[0082] The terms "substantially" or "approximately," as used herein, may be applied to modify any quantitative comparison, value, measurement, or other representation that may vary within permissible limits without resulting in a change in the basic function to which it pertains.

[0083] The term "about" generally means within the standard deviation of the stated value, as determined using standard analytical techniques to measure the stated value. The term may also be used to refer to plus or minus 5% of the stated value.

[0084] As used herein, a sequence that is "substantially" similar to a wild-type sequence contains a percent GC content within 5% of the percent GC content of the wild-type.

[0085] The term "cell population" is used herein to refer to a group of cells that typically share a common type. A cell population may be derived from a common precursor cell, and may contain more than one type of cell. An "enriched" cell population refers to a cell population derived from a starting cell population (e.g., an unfractionated heterogeneous cell population) that contains a higher percentage of a particular cell type than the percentage of that cell type in the starting population. A cell population may be enriched for one or more cell types, or may be depleted for one or more cell types.

[0086] The term "stem cell" refers to a cell that, under suitable conditions, can differentiate into a diverse range of specialized cell types, but under other conditions, can self-renew and remain essentially in an undifferentiated pluripotent state. The term "stem cell" also encompasses pluripotent, multipotent, precursor and progenitor cells. Exemplary human stem cells can be derived from hematopoietic or mesenchymal stem cells obtained from bone marrow tissue, embryonic stem cells obtained from embryonic tissue, or embryonic germ cells obtained from fetal reproductive tissue. Exemplary pluripotent stem cells can also be generated from somatic cells by reprogramming the somatic cells to a pluripotent state through the expression of certain transcription factors associated with pluripotency, and these cells are referred to as "induced pluripotent stem cells" or "iPSCs."

[0087] The term "pluripotency" refers to the property of a cell to differentiate into all other cell types of an organism, except extraembryonic or placental cells. Pluripotent stem cells are capable of differentiating into cell types of all three germ layers (e.g., ectoderm, mesoderm, and endoderm), even after long-term culture. Pluripotent stem cells may be embryonic stem cells derived from the inner cell mass of a blastocyst or produced by nuclear transfer. In other embodiments, the pluripotent stem cells are induced pluripotent stem cells derived from the reprogramming of somatic cells.

[0088] The term "differentiation" refers to the process by which unspecialized cells become more specialized types, with changes in structural and / or functional properties. Mature cells typically have altered cellular structure and possess tissue-specific proteins.

[0089] As used herein, "undifferentiated" refers to cells that display characteristic markers and morphological features of undifferentiated cells, clearly distinguishable from terminally differentiated cells of embryonic or adult origin.

[0090] "Embryoid bodies (EBs)" are aggregates of pluripotent stem cells that can differentiate into cells of the endoderm, mesoderm, and ectoderm germ layers. When pluripotent stem cells are aggregated under non-adherent culture conditions, they form spheroid structures, thus forming EBs in suspension.

[0091] An "isolated" cell is one that has been substantially separated or purified from other cells in an organism or culture. An isolated cell can be, for example, at least 99%, at least 98% pure, at least 95% pure, or at least 90% pure.

[0092] "Cell line" as used herein refers to a collection of cells originating from a single cell. A cell line may be maintained in a growth medium in a tube, flask, or dish. A cell line arises by clonal growth from a single cell and can be expanded to multiple cells. A cell line may include cells that are genetically identical and can be maintained in culture over time, such as for months or years.

[0093] "Embryo" refers to a mass of cells resulting from one or more divisions of a zygote or activated oocyte having an artificially reprogrammed nucleus.

[0094] "Embryonic stem (ES) cells" are undifferentiated pluripotent cells obtained from earlier stage embryos, such as the inner cell mass of the blastocyst stage, or generated by artificial means (e.g., nuclear transfer), that can give rise to all differentiated cell types of an embryo or an adult, including germ cells (e.g., sperm and eggs).

[0095] "Induced pluripotent stem cells (iPSCs)" are cells generated by reprogramming somatic cells by expressing or inducing the expression of a combination of factors (referred to herein as reprogramming factors). iPSCs can be generated using fetal, postnatal, neonatal, juvenile, or adult somatic cells. In certain embodiments, factors that can be used to reprogram somatic cells into pluripotent stem cells include, for example, Oct4 (sometimes referred to as Oct 3 / 4), Sox2, c-Myc, and Klf4, Nanog, and Lin28. In some embodiments, somatic cells are reprogrammed by expressing at least two reprogramming factors, at least three reprogramming factors, or four reprogramming factors to reprogram somatic cells into pluripotent stem cells.

[0096] "Feeder-free" or "feeder-independent" refers to cultures supplemented with cytokines and growth factors (e.g., TGFβ, bFGF, LIF) as an alternative to a feeder cell layer. Thus, "feeder-free" or feeder-independent culture systems and media may be used to culture and maintain pluripotent cells in an undifferentiated, proliferative state. In some cases, feeder-free cultures utilize animal-based matrices (e.g., MATRIGEL™) or are grown on substrates such as fibronectin, collagen, or vitronectin. These approaches allow human stem cells to maintain an essentially undifferentiated state without the need for a "feeder layer" of mouse fibroblast cells.

[0097] A "feeder layer" is defined herein as a coating layer of cells, such as on the bottom of a culture dish. Feeder cells can release nutrients into the culture medium and provide a surface to which other cells, such as pluripotent stem cells, can attach.

[0098] The term "defined" or "well-defined" when used in relation to a medium, extracellular matrix, or culture condition refers to a medium, extracellular matrix, or culture condition in which the chemical composition and amount of nearly all components are known. For example, a defined medium does not contain undefined factors such as fetal bovine serum, bovine serum albumin, or human serum albumin. Generally, a defined medium includes a basal medium (e.g., Dulbecco's Modified Eagle Medium (DMEM), F12, or Roswell Park Memorial Institute Medium (RPMI) 1640 containing amino acids, vitamins, inorganic salts, buffers, antioxidants, and energy sources) supplemented with recombinant albumin, chemically defined lipids, and recombinant insulin. An example of a well-defined medium is Essential 8™ medium.

[0099] In media, extracellular matrices, or culture systems used with human cells, the term "xeno-free (XF)" refers to conditions in which the materials used are not of non-human animal origin.

[0100] II. Engineered cell lines In some embodiments, provided herein are cell lines engineered to express a transgene with stable expression. Stable expression can be achieved by codon-optimizing the transgene sequence to remove CpG motifs and driving expression with a novel promoter (e.g., SEQ ID NO: 1-12 or 17), or by tagging an endogenous gene (e.g., HSP90AB1, ACTB, CTNNB1, or MYL6) to drive expression.

[0101] As used herein, a "CpG motif" refers to a nucleotide that contains a cysteine ​​"C" followed by a phosphate bond "p" and a guanine "G". Reference to "removal of the CpG motif" means that the C and / or G nucleotides are modified to remove the motif. As used herein, "humanization" in reference to a nucleic acid molecule means that the nucleic acid molecule has a sequence or a portion of a sequence that is similar or closely resembles a human sequence, or is otherwise made to make the molecule more functional in human cells. For example, codons can be optimized for use in humans based on known codon usage in humans to increase the efficiency of expression of the nucleic acid in human cells, e.g., to achieve faster translation rates and higher accuracy. [Table 1]

[0102] The process of gene switching off by methylation is described by a series of cascading events that ultimately lead to changes in chromatin structure and the formation of a state in which transcription is weak. Methylation of the 5'-CpG-3' of a gene results in binding to the methylated DNA sequence and at the same time binding to histone deacetylases (MBD-HDACs) and transcription inhibitory proteins (transcription redresser proteins). Artificial gene synthesis techniques allow the synthesis of any nucleotide sequence selected from this possibility, where the amino acid sequence encoded by the corresponding gene is preferably not changed. The modified target nucleic acid sequence is generated from long oligonucleotides, for example by stepwise PCR as described in the examples, or by specialized suppliers (e.g. Geneart GmbH, Qiagen AG) in the case of conventional gene synthesis.

[0103] In some aspects, all CpGs of the transgene that can be removed within the genetic code are removed. However, fewer CpGs may be removed, for example, 50%, 60%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99%. The codon-optimized construct according to the present disclosure can be prepared, for example, by selecting the same codon distribution as the expression system used. The expression system may be a mammalian system, for example, a human system. Preferably, the codon optimization is compatible with the codon selection of human genes.

[0104] Some codons in homo sapiens are low in abundance, while others are medium or high in abundance. As used herein, "rare codon" refers to a codon with a frequency of less than 0.2 in homo sapiens. In order to avoid using rare codons when modifying DNA sequences, a codon frequency table may be used to select a codon with a frequency of at least 0.3, for example at least 0.4, 0.5, 0.6, 0.7, 0.8, or 0.9. [ka]

[0105] For example, when modifying a sequence encoding L, leucine, there are six codons to evaluate (listed here as codon+frequency): UUA 8%, UUG 13%, CUU 13%, CUC 20%, CUA 7%, or CUG 40%. If the leucine codon CUC is followed by a codon beginning with G, a CG motif is present and modification of the leucine codon to CUG is preferred over the other four codons to remove the CG motif and avoid the use of a rare codon.

[0106] Modification of codons such as CCG for proline to remove the CG motif can be achieved using the codon CCC, but if the protein region contains several prolines, this will generate mononucleotide stretch repeats. Therefore, other codons such as CCU or CCA can be used for proline to avoid mononucleotide stretches. As used herein, a "mononucleotide stretch" refers to a region of at least six identical nucleotides in a row, such as CCCCCC.

[0107] The sequence of the transgene may code for an RNA, a derivative or mimic thereof, a peptide or polypeptide, a modified peptide or polypeptide, a protein or modified protein. The transgene may be a chimeric and / or assembled sequence of different wild-type sequences, for example, encoding a fusion protein or a mosaic assembled polygene construct. The transgene may also comprise a synthetic sequence. In this regard, the nucleic acid sequence may be synthetically modeled, such as by using a computer model.

[0108] The transgene expressed may be any protein, such as a genetic sequence of a recombinant protein, an artificial polypeptide, a fusion protein, and equivalents thereof. In some aspects, the transgene is a diagnostic and / or therapeutic peptide, polypeptide, or protein. In some aspects, the transgene is a reporter gene, including but not limited to, GFP, RFP, luciferase, β-galactosidase, or chloramphenicol acetyltransferase. In some aspects, the transgene is LacZ, mSEAP, or Lucia. Peptides / proteins include, for example, i) human enzymes (e.g., asparaginase, adenosine deaminase, insulin, tPA, clotting factors, vitamin K epoxide reductase), hormones (e.g., erythropoietin, follicle stimulating hormone, estrogen), and other human derived proteins (e.g., bone morphogenetic proteins, antithrombin), ii) viral proteins, bacterial proteins that can be used as vaccines, or proteins from parasites (e.g., HIV, HBV, HCV, influenza, borrelia, hemophilus, meningococcus, anthrax, botulinum toxin, diphtheria toxin, tetanus toxin, protozoa, etc.), or iii) diagnostic agents. The transgene may be a promoter or a selection gene, e.g., blasticidin or neomycin.

[0109] A. Induced pluripotent stem cells In some embodiments, the engineered cell line is an iPSC. Induction of pluripotency was originally achieved by reprogramming somatic cells by introducing pluripotency-related transcription factors, using mouse cells in 2006 (Yamanaka et al., 2006) and human cells in 2007 (Yu et al., 2007; Takahashi et al., 2007). Pluripotent stem cells can be maintained in an undifferentiated state and can be differentiated into any adult cell type.

[0110] Any somatic cell can be used as the starting point for iPSCs, except for germ cells. For example, the cell type can be keratinocytes, fibroblasts, hematopoietic cells, mesenchymal cells, hepatocytes, or gastric cells. T cells may also be used as a source of somatic cells for reprogramming (U.S. Patent No. 8,741,648). There is no limit to the degree of cell differentiation or the age of the animal from which the cells are harvested, and undifferentiated progenitor cells (including somatic stem cells) and even terminally differentiated mature cells can be used as a source of somatic cells in the methods disclosed herein. iPSCs can be grown under conditions known to differentiate human ES cells into specific cell types and express human ES cell markers, including: SSEA-1, SSEA-3, SSEA-4, TRA-1-60, and TRA-1-81.

[0111] A. HLA compatible The major histocompatibility complex (MHC) is the primary cause of immune rejection of allogeneic organ transplants. There are three major class I MHC haplotypes (A, B, and C) and three major MHC class II haplotypes (DR, DP, and DQ).

[0112] The MHC compatibility between donor and recipient is significantly improved when the donor's cells are HLA homozygous, i.e., contain the same alleles for each antigen-presenting protein. Most individuals are heterozygous for MHC class I and II genes, but certain individuals are homozygous for these genes. These homozygous individuals can act as super donors, and grafts obtained from their cells can be transplanted into all individuals who are either homozygous or heterozygous for that haplotype. Furthermore, if homozygous donor cells have a haplotype that is frequently found in the population, these cells may be applicable to transplantation therapy for a large number of individuals.

[0113] Thus, iPSCs can be generated from somatic cells of the subject to be treated or another subject with an HLA type identical or substantially identical to that of the patient. In some cases, the donor's major HLA (e.g., the three major loci HLA-A, HLA-B, and HLA-DR) is identical to the recipient's major HLA. In some cases, the somatic cell donor may be a super donor, and thus iPSCs derived from an MHC homozygous super donor may be used to generate differentiated cells. Thus, iPSCs derived from a super donor may be transplanted into a subject who is homozygous or heterozygous for that haplotype. For example, the iPSCs may be homozygous for two HLA alleles, such as HLA-A and HLA-B. In this way, iPSCs generated from a super donor can be used in the methods disclosed herein to generate differentiated cells that are potentially "compatible" with a large number of potential recipients.

[0114] B. Reprogramming Factors Somatic cells can be reprogrammed to generate induced pluripotent stem cells (iPSCs) using methods known to those skilled in the art. Those skilled in the art can easily generate induced pluripotent stem cells, see, for example, US Patent Application Publication No. 20090246875, US Patent Application Publication No. 2010 / 0210014, US Patent Application Publication No. 20120276636, US Patent No. 8,058,065, US Patent No. 8,129,187, US Patent No. 8,278,620, PCT Publication No. WO2007 / 069666 A1, and US Patent No. 8,268,620, which are incorporated herein by reference. In general, nuclear reprogramming factors are used to generate pluripotent stem cells from somatic cells. In some embodiments, at least two, at least three, or at least four of Klf4, c-Myc, Oct3 / 4, Sox2, Nanog, and Lin28 are utilized. In other embodiments, Oct3 / 4, Sox2, c-Myc, and Klf4 are utilized. In some aspects, 5, 6, 7, or 8 reprogramming factors are used.

[0115] The cell is treated with a nuclear reprogramming agent, and the nuclear reprogramming agent is generally one or more factors capable of inducing iPSC from somatic cells, or the nucleic acid encoding these agents (including the form incorporated in a vector).The nuclear reprogramming agent generally includes at least Oct3 / 4, Klf4 and Sox2, or the nucleic acid encoding these molecules.Functional inhibitors of p53, L-myc or the nucleic acid encoding L-myc, and Lin28 or Lin28b or the nucleic acid encoding Lin28 or Lin28b can be used as additional nuclear reprogramming agents.Nanog can also be used for nuclear reprogramming. As disclosed in U.S. Patent Application Publication No. 20120196360, exemplary reprogramming factors for the generation of iPSCs include: (1) Oct3 / 4, Klf4, Sox2, L-Myc (Sox2 can be replaced with Soxl, Sox3, Soxl5, Soxl7, or Soxl8, and Klf4 can be replaced with Klfl, Klf2, or Klf5); (2) Oct3 / 4, Klf4, Sox2, L-Myc, TERT, SV40 Large T antigen (SV40LT); (3) Oct3 / 4, Klf4, Sox2, L-Myc, TERT, human papilloma virus (HPV) 16 E6; (4) Oct3 / 4, Klf4, Sox2, L-Myc, TERT, HPV16 E7; (5) Oct3 / 4, Klf4, Sox2, L-Myc, TERT, HPV16 E6, HPV16 E7; (6) Oct3 / 4, Klf4, Sox2, L-Myc, TERT, Bmil; (7) Oct3 / 4, Klf4, Sox2, L-Myc, Lin28; (8) Oct3 / 4, Klf4, Sox2, L-Myc, Lin28, SV40LT; (9) Oct3 / 4, Klf4, Sox2, L-Myc, Lin28, TERT, SV40LT; (10) Oct3 / 4, Klf4, Sox2, L-Myc, SV40LT, (11) Oct3 / 4, Esrrb, Sox2, L-Myc (Esrrb can be replaced with Esrrg), (12) Oct3 / 4, Klf4, Sox2; (13) Oct3 / 4, Klf4, Sox2, TERT, SV40LT; (14) Oct3 / 4, Klf4, Sox2, TERT, HP VI 6 E6;(15) Oct3 / 4、Klf4、Sox2、TERT、HPV16 E7; (16) Oct3 / 4、Klf4、Sox2、TERT、HPV16 E6、HPV16 E7; (17) Oct3 / 4、Klf4、Sox2、TERT、Bmil; (18) Oct3 / 4、Klf4、Sox2、Lin28 (19) Oct3 / 4、Klf4、Sox2、Lin28、SV40LT; (20) Oct3 / 4、Klf4、Sox2、Lin28、TERT、SV40LT;(21) Oct3 / 4, Klf4, Sox2, SV40LT, or (22) Oct3 / 4, Esrrb, Sox2 (Esrrb can be replaced with Esrrg). In one non-limiting example, Oct3 / 4, Klf4, Sox2, and c-Myc are utilized. In other embodiments, Oct4, Nanog, and Sox2 are utilized, see, for example, U.S. Patent No. 7,682,828, which is incorporated herein by reference. These factors include, but are not limited to, Oct3 / 4, Klf4, and Sox2. In other examples, factors include, but are not limited to, Oct 3 / 4, Klf4, and Myc. In some non-limiting examples, Oct3 / 4, Klf4, c-Myc, and Sox2 are utilized. In other non-limiting examples, Oct3 / 4, Klf4, Sox2, and Sal4 are utilized. Factors such as Nanog, Lin28, Klf4, or c-Myc can increase reprogramming efficiency and can be expressed from several different expression vectors. For example, integrating vectors such as systems based on EBV elements can be used (US Pat. No. 8,546,140). In a further aspect, the reprogramming proteins can be directly introduced into the somatic cells by protein transfer. Reprogramming can further include contacting the cells with one or more signaling receptors, including glycogen synthase kinase 3 (GSK-3) inhibitors, mitogen-activated protein kinase kinase (MEK) inhibitors, transforming growth factor beta (TGF-β) receptor inhibitors or signal transduction inhibitors, leukemia inhibitory factor (LIF), p53 inhibitors, NF-kappa B inhibitors, or combinations thereof. These regulators can include small molecules, inhibitory nucleotides, expression cassettes, or protein factors. It is expected that virtually any iPS cell or cell line can be used;

[0116] The mouse and human cDNA sequences for these nuclear reprogramming agents are available at NCBI Accession Nos. WO2007 / 069666, which is incorporated herein by reference. Methods for introducing one or more reprogramming agents, or nucleic acids encoding these reprogramming agents, are known in the art and are disclosed, for example, in U.S. Patent Application Publication No. 2012 / 0196360 and U.S. Patent No. 8,071,369.

[0117] The induced iPSCs can be cultured in a medium sufficient to maintain pluripotency. iPSCs can be used with various media and techniques developed for culturing pluripotent stem cells, more particularly embryonic stem cells, as described in U.S. Pat. No. 7,442,548 and U.S. Pat. Publication No. 2003 / 0211603. In the case of mouse cells, culture is performed by adding leukemia inhibitory factor (LIF) as a differentiation suppressor to a normal medium. In the case of human cells, it is desirable to add basic fibroblast growth factor (bFGF) instead of LIF. Other methods for culturing and maintaining iPSCs may be used, as known to those skilled in the art.

[0118] In certain embodiments, undefined conditions may be used, for example, pluripotent cells may be cultured on fibroblast feeder cells or on media exposed to fibroblast feeder cells to maintain stem cells in an undifferentiated state. In some embodiments, cells are cultured in the presence of mouse embryonic fibroblasts as feeder cells that have been treated with radiation or antibiotics to terminate cell division. Alternatively, pluripotent cells may be cultured and maintained in an essentially undifferentiated state using defined feeder-independent culture systems such as TESR™ medium (Ludwig et al., 2006a; Ludwig et al., 2006b) or E8™ medium (Chen et al., 2011).

[0119] C. Plasmids In some embodiments, iPSCs can be modified to express exogenous nucleic acids, for example, to include an enhancer operably linked to a promoter and a nucleic acid sequence encoding a first marker. The construct can also include other elements, such as a ribosome binding site (internal ribosome binding sequence) for translation initiation, a transcription / translation terminator, etc. In general, it is advantageous to transfect the construct into cells. Vectors suitable for stable transfection include, but are not limited to, retroviral vectors, lentiviral vectors, and Sendai virus.

[0120] In some embodiments, the marker-encoding plasmid is comprised of (1) a high copy number origin of replication, (2) a selectable marker, such as, but not limited to, a neo gene for antibiotic selection with kanamycin, (3) a transcription termination sequence including a tyrosinase enhancer, and (4) a multiple cloning site for incorporating various nucleic acid cassettes, and (5) a nucleic acid sequence encoding the marker operably linked to a tyrosinase promoter. Numerous plasmid vectors for inducing protein-encoding nucleic acids are known in the art. These include, but are not limited to, the vectors disclosed in U.S. Pat. Nos. 6,103,470, 7,598,364, 7,989,425, and 6,416,998, which are incorporated herein by reference. In some aspects, the plasmid contains a "suicide gene" that converts the gene product into a compound that kills the host cell upon administration of a prodrug or drug. Examples of suicide gene, prodrug or drug combinations that can be used include, but are not limited to, truncated EGFR and cetuximab, herpes simplex virus-thymidine kinase (HSV-tk) and ganciclovir, acyclovir or FIAU; oxidoreductase and cycloheximide, cytosine deaminase and 5-fluorocytosine, thymidine kinase thymidylate kinase (Tdk::Tmk) and AZT, and deoxycytidine kinase and cytosine arabinoside.

[0121] The viral gene delivery system can be an RNA-based or DNA-based viral vector. The episomal gene delivery system can be a plasmid, an Epstein-Barr Virus (EBV)-based episomal vector, a yeast-based vector, an adenovirus-based vector, a Simian Virus 40 (SV40)-based episomal vector, a bovine papilloma virus (BPV)-based vector, or a lentivirus vector.

[0122] Markers include, but are not limited to, fluorescent proteins (e.g., green fluorescent protein or red fluorescent protein), enzymes (e.g., horseradish peroxidase or alkaline phosphatase or firefly / renilla luciferase or nanoluc), or other proteins. Markers may be proteins (including secreted proteins, cell surface proteins, or internal proteins; synthesized or taken up by the cell), nucleic acids (e.g., mRNA, or enzymatically active nucleic acid molecules), or polysaccharides. Also included are determinants of any such cellular constituents that are detectable by antibodies, lectins, probes, or nucleic acid amplification reactions and are specific for the markers of the cell type of interest. Markers can also be identified by biochemical or enzymatic assays, or biological responses that depend on the function of the gene product. Nucleic acid sequences encoding these markers can be operably linked to a tyrosinase enhancer. In addition, other genes can also be included, such as genes that can affect stem cell differentiation, or cell function, or physiology, or pathology.

[0123] D. Delivery Systems Introduction of nucleic acids, such as DNA or RNA, into the engineered cell lines of the present disclosure may use any suitable method for nucleic acid delivery for cell transformation as described herein or known to those of skill in the art, including, but not limited to, ex vivo transfection (Wilson et al., 1989; Nabel et al., 1989), injection (U.S. Patent Nos. 5,994,624; 5,981,274; 5,945, each of which is incorporated herein by reference), including microinjection (Harland and Weintraub, 1985; U.S. Patent No. 5,789,215), and microinjection (U.S. Patent Nos. 5,994,624, 5,981,274, 5,945, each of which is incorporated herein by reference). ,100, 5,780,448, 5,736,524, 5,702,932, 5,656,610, 5,589,466, and 5,580,859; electroporation (U.S. Pat. No. 5,384,253; Tur-Kaspa et al., 1986; Potter et al., 1984; calcium phosphate precipitation (Graham and Van Der Eb, 1973; Chen and Okayama, 1987; Rippe et al., 1990); using DEAE-dextran followed by polyethylene glycol (Gopal, 1985); direct sonoloading (Fechheimer et al., 1987); liposome-mediated transfection (Fechheimer et al., 1987); liposome-mediated transfection (Nicolau and Sene, 1982; Fraley et al., 1979; Nicolau et al., 1987; Wong et al., 1980; Kaneda et al., 1991); 89; Kato et al., 1991) and receptor-mediated transfection (Wu and Wu, 1987; Wu and Wu, 1988); particle bombardment (PCT Applications Nos. WO 94 / 09699 and 95 / 06128; U.S. Patent Nos. 5,610,042; 5,322,783; 5,563,055; 5,550,318; 5,538,877; and 5,538,880, each of which is incorporated herein by reference); silicon carbide fiber agitation (Kaeppler et al., 1990;Nos. 5,302,523 and 5,464,765; Agrobacterium-mediated transformation (U.S. Pat. Nos. 5,591,616 and 5,563,055, each of which is incorporated herein by reference); desiccation / inhibition-mediated DNA uptake (Potrykus et al., 1985), and direct delivery of DNA, such as by any combination of such methods. By application of techniques such as these, organelles, cells, tissues or organisms may be stably or transiently transformed;

[0124] 1. Viral Vectors Viral vectors may be provided in certain embodiments of the present disclosure. In generating recombinant viral vectors, non-essential genes are typically replaced with genes or coding sequences of heterologous (or non-native) proteins. Viral vectors are a type of expression construct that utilizes viral sequences to introduce nucleic acids and, in some cases, proteins into cells. The ability of certain viruses to infect cells or enter cells via receptor-mediated endocytosis, integrate into the host cell genome, and stably and efficiently express viral genes makes them attractive candidates for transferring foreign nucleic acids into cells (e.g., mammalian cells). Non-limiting examples of viral vectors that may be used to deliver nucleic acids of certain embodiments of the present disclosure are described below.

[0125] Retroviruses have been shown to be promising gene delivery vectors due to their ability to integrate genes into the host genome, to transfer large amounts of foreign genetic material, to infect a wide range of species and cell types, and to be packaged in specialized cell lines (Miller, 1992).

[0126] To construct retroviral vectors, nucleic acids are inserted into the viral genome in place of certain viral sequences, resulting in replication-defective viruses. To produce virions, packaging cell lines are constructed that contain the gag, pol, and env genes, but without LTRs or packaging components (Mann et al., 1983). When recombinant plasmids containing cDNA together with retroviral LTRs and packaging sequences are introduced into specialized cell lines (e.g., by calcium phosphate precipitation), the packaging sequences package the RNA transcripts of the recombinant plasmid into viral particles that are then secreted into the culture medium (Nicolas and Rubenstein, 1988; Temin, 1986; Mann et al., 1983). The medium containing the recombinant retrovirus is then collected, optionally concentrated, and used for gene transfer. Retroviral vectors can infect a wide variety of cell types. However, integration and stable expression require the division of host cells (Paskind et al., 1975).

[0127] Lentiviruses are complex retroviruses that contain, in addition to the common retroviral genes gag, pol, and env, other genes with regulatory or structural functions. Lentiviral vectors are well known in the art (see, e.g., Naldini et al., 1996; Zufferey et al., 1997; Blomer et al., 1997; U.S. Patent Nos. 6,013,516 and 5,994,136).

[0128] Recombinant lentiviral vectors are capable of infecting non-dividing cells and can be used for both in vivo and ex vivo gene transfer and expression of nucleic acid sequences. For example, recombinant lentiviruses capable of infecting non-dividing cells (suitable host cells are transfected with two or more vectors having packaging functions, i.e., gag, pol and env, and rev and tat) are described in U.S. Patent No. 5,994,136, which is incorporated herein by reference.

[0129] 2. Episomal Vectors The use of plasmid or liposome-based extrachromosomal (i.e., episomal) vectors may also be provided in certain embodiments of the present disclosure. Such episomal vectors may include, for example, oriP-based vectors and / or vectors encoding derivatives of EBNA-1. These vectors may allow large fragments of DNA to be introduced into cells, maintained extrachromosomally, replicated once per cell cycle, efficiently distributed to daughter cells, and do not substantially provoke an immune response.

[0130] In particular, EBNA-1, the only viral protein required for the replication of oriP-based expression vectors, has developed an efficient mechanism to circumvent the processing required for its antigen presentation on MHC class I molecules and therefore does not elicit a cellular immune response (Levitskaya et al., 1997). Furthermore, EBNA-1 acts in trans to enhance the expression of cloned genes and can induce the expression of cloned genes up to 100-fold in some cell lines (Langle-Rouault et al., 1998; Evans et al., 1997). Finally, the production of such oriP-based expression vectors is inexpensive.

[0131] In certain aspects, the reprogramming factors are expressed from an expression cassette contained in one or more exogenous episomal genetic elements (see U.S. Patent Publication No. 2010 / 0003757, incorporated herein by reference). Thus, iPSCs can be essentially free of exogenous genetic elements, such as those derived from retroviral or lentiviral vector elements. These iPSCs are prepared by the use of extrachromosomally replicating vectors (i.e., episomal vectors), which are vectors capable of replicating episomally to generate iPSCs that are essentially free of exogenous vector or viral elements (see U.S. Patent No. 8,546,140; Yu et al., 2009, incorporated herein by reference). Numerous DNA viruses, such as adenovirus, Simian vacuolating virus 40 (SV40) or bovine papilloma virus (BPV), or plasmids containing the budding yeast ARS (autonomously replicating sequence), replicate extrachromosomally or episomally in mammalian cells. These episomal plasmids are essentially free of these drawbacks associated with integrative vectors (Bode et al., 2001). For example, those based on lymphotrophic herpesviruses or Epstein-Barr virus (EBV), as defined above, replicate extrachromosomally and are useful for delivering reprogramming genes to somatic cells. Useful EBV elements are OriP and EBNA-1, or variants or functional equivalents thereof. An additional advantage of episomal vectors is that exogenous elements are introduced into cells and then lost over time, resulting in self-sustaining iPSCs that are essentially free of these elements.

[0132] Other extrachromosomal vectors include other lymphotrophic herpesvirus-based vectors. Lymphotrophic herpesviruses are herpesviruses that replicate in lymphoblasts (e.g., human B lymphoblasts) and become plasmids as part of their natural life cycle. Herpes simplex virus (HSV) is not a "lymphotrophic" herpesvirus. Exemplary lymphotrophic herpesviruses include, but are not limited to, EBV, Kaposi's sarcoma herpesvirus (KSHV); herpesvirus saimiri (HS) and Marek's disease virus (MDV). Other sources of episomal-based vectors are also contemplated, such as yeast ARS, adenovirus, SV40, or BPV.

[0133] Those skilled in the art will be well able to construct vectors by standard recombinant techniques (see, for example, Maniatis et al., 1988 and Ausubel et al., 1994, both of which are incorporated herein by reference).

[0134] Vectors may also contain other components or functionalities that further modulate gene delivery and / or gene expression or that provide beneficial properties to targeted cells. Such other components include, for example, components that affect binding or targeting to cells (including components that mediate cell-type or tissue-specific binding), components that affect uptake of vector nucleic acid by cells, components that affect localization of the polynucleotide within the cell after uptake (such as agents that mediate nuclear localization), and components that affect expression of the polynucleotide.

[0135] Such components may also include markers, such as detectable and / or selectable markers that can be used to detect or select cells that have taken up and are expressing the nucleic acid delivered by the vector. Such components may be provided as natural features of the vector (such as the use of certain viral vectors that have components or functionality that mediate binding and uptake), or the vector may be modified to provide such functionality. A wide variety of such vectors are known in the art and are generally available. When a vector is maintained in a host cell, it may either be stably replicated by the cell as an autonomous structure during mitosis, integrated into the genome of the host cell, or maintained in the nucleus or cytoplasm of the host cell.

[0136] 3. Regulatory Elements The expression cassette contained in the reprogramming vectors useful in the present disclosure preferably contains (in the 5' to 3' direction) a eukaryotic transcriptional promoter operably linked to a protein coding sequence, a splice signal with intervening sequences, and a transcription termination / polyadenylation sequence.

[0137] Promoter / Enhancer The expression constructs provided herein include a promoter that drives expression of the programming genes. Promoters generally contain sequences that function to position the start site for RNA synthesis. The best known example of this is the TATA box, but in some promoters that lack a TATA box, such as the mammalian terminal deoxynucleotidyl transferase gene promoter and the SV40 late gene promoter, separate elements above the start site itself help to anchor the start location. Additional promoter elements regulate the frequency of transcription initiation. Typically, these are located in a region 30-110 bp upstream of the start site, although some promoters have been shown to contain functional elements downstream of the start site as well. To place a coding sequence "under the control" of a promoter, the coding sequence is positioned "downstream" of (i.e., 3' to) the selected promoter, 5' to the transcription start site of the transcriptional reading frame. The "upstream" promoter stimulates transcription of the DNA and promotes expression of the encoded RNA.

[0138] Spacing between promoter elements is often flexible, such that promoter function is preserved when elements are inverted or moved relative to one another. In the tk promoter, spacing between promoter elements can be increased by up to 50 bp before activity begins to decline. Depending on the promoter, individual elements appear to be able to function cooperatively or independently to activate transcription. Promoters may or may not be used in conjunction with "enhancers," which refer to cis-acting regulatory sequences involved in the transcriptional activation of a nucleic acid sequence.

[0139] A promoter may be one that is naturally associated with a nucleic acid sequence, as may be obtained by isolating the 5' non-coding sequence located upstream of a coding segment and / or exon. Such a promoter is referred to as "endogenous". Similarly, an enhancer may be one that is naturally associated with a nucleic acid sequence, located either downstream or upstream of that sequence. Alternatively, certain advantages may be obtained by placing the coding nucleic acid segment under the control of a recombinant or heterologous promoter, which also refers to a promoter that is not normally associated with a nucleic acid sequence in its natural environment. A recombinant or heterologous enhancer also refers to an enhancer that is not normally associated with a nucleic acid sequence in its natural environment. Such promoters or enhancers may include promoters or enhancers of other genes, and promoters or enhancers isolated from any other virus, or prokaryotic or eukaryotic cell, and "non-naturally occurring" promoters or enhancers, i.e., promoters or enhancers that contain different elements of different transcriptional regulatory regions and / or mutations that alter expression. For example, promoters most commonly used in recombinant DNA construction include the β-lactamase (penicillinase), lactose and tryptophan (trp) promoter systems. In addition to synthetically producing promoter and enhancer nucleic acid sequences, sequences may be produced using recombinant cloning and / or nucleic acid amplification techniques, including PCR™, in conjunction with the compositions disclosed herein (see U.S. Pat. Nos. 4,683,202 and 5,928,906, each of which is incorporated herein by reference). Furthermore, it is contemplated that control sequences that direct transcription and / or expression of sequences within non-nuclear organelles, such as mitochondria, chloroplasts, etc., may similarly be used.

[0140] Of course, it will be important to use a promoter and / or enhancer that effectively directs the expression of the DNA segment in the organelle, cell type, tissue, organ, or organism selected for expression. Those skilled in the art of molecular biology are generally aware of the use of promoter, enhancer, and cell type combinations for protein expression (see, for example, Sambrook et al. 1989, incorporated herein by reference). The promoter used may be constitutive, tissue-specific, inducible, and / or useful under appropriate conditions to direct high-level expression of the introduced DNA segment, such as is advantageous in large-scale production of recombinant proteins and / or peptides. The promoter may be heterologous or endogenous.

[0141] Additionally, any promoter / enhancer combination (e.g., from the Eukaryotic Promoter Data Base EPDB) can be used to drive expression. Use of the T3, T7, or SP6 cytoplasmic expression systems is another possible embodiment. Eukaryotic cells can support cytoplasmic transcription from certain bacterial promoters if the appropriate bacterial polymerase is provided as part of the delivery complex or as an additional gene expression construct.

[0142] Non-limiting examples of promoters include early or late viral promoters, such as the SV40 early or late promoters, the cytomegalovirus (CMV) immediate early promoter, the Rous sarcoma virus (RSV) early promoter; eukaryotic promoters, such as the beta-actin promoter (Ng, 1989; Quitsche et al., 1989), the GADPH promoter (Alexander et al., 1988; Ercolani et al., 1988), the metallothionein promoter (Karin et al., 1989; Richards et al., 1984); and ligated response element promoters, such as the cyclic AMP response element promoter near the minimal TATA box (cre), the serum response element promoter (sre), the phorbol ester promoter (TPA) and the response element promoter (tre). It is also possible to use the human growth hormone promoter sequence (e.g., the human growth hormone minimal promoter found in Genbank, Accession No. X05244, nucleotides 283-341) or the mouse mammary tumor promoter (available from the ATCC, catalogue number ATCC 45007).

[0143] Tissue-specific expression of transgenes, particularly reporter gene expression in hematopoietic cells and precursors of hematopoietic cells derived from programming, may be desirable as a method to identify derived hematopoietic cells and precursors. To enhance both specificity and activity, the use of cis-acting regulatory elements is contemplated. For example, hematopoietic cell-specific promoters may be used. Many such hematopoietic cell-specific promoters are known in the art.

[0144] In certain aspects, the methods of the present disclosure also relate to enhancer sequences, i.e. nucleic acid sequences that increase the activity of a promoter and have the potential to act in cis and regardless of their orientation, even at relatively long distances (up to several kilobases away from the target promoter), however, enhancer function is not necessarily limited to such long distances and may also function in close proximity to a given promoter.

[0145] Many hematopoietic cell promoter and enhancer sequences have been identified and may be useful in the methods of the invention. See, e.g., U.S. Patent No. 5,556,954; U.S. Patent Application No. 20020055144; U.S. Patent Application No. 20090148425.

[0146] b. Initiation signals and associated expression Specific initiation signals may also be used in the expression constructs provided in this disclosure for efficient translation of coding sequences. These signals include the ATG initiation codon or adjacent sequences. Exogenous translational control signals, including the ATG initiation codon, may need to be provided. One skilled in the art would be able to easily determine this and provide the necessary signals. It is well known that the initiation codon must be "in frame" with the reading frame of the desired coding sequence to ensure translation of the entire insert. Exogenous translational control signals and initiation codons may be natural or synthetic. The efficiency of expression may be improved by including appropriate transcriptional enhancer elements.

[0147] In certain embodiments, internal ribosome entry site (IRES) elements are used to generate multigene, or polycistronic, messages. IRES elements can bypass the ribosome scanning model of 5' methylated Cap-dependent translation and initiate translation at internal sites (Pelletier and Sonenberg, 1988). IRES elements from two members of the picornavirus family (polio and encephalomyocarditis) have also been reported (Pelletier and Sonenberg, 1988), as well as IRESs from mammalian messages (Macejak and Sarnow, 1991). IRES elements can be linked to heterologous open reading frames. Multiple open reading frames can be transcribed together, each separated by an IRES, generating polycistronic messages. IRES elements allow each open reading frame to be accessible to ribosomes for efficient translation. Multiple genes can be efficiently expressed using a single promoter / enhancer to transcribe a single message (see US Pat. Nos. 5,925,565 and 5,935,819).

[0148] Additionally, certain 2A sequence elements can be used in constructs provided in the present disclosure to generate linked or co-expression of programming genes. For example, cleavage sequences can be used to co-express genes by linking open reading frames to form a single cistron. Exemplary cleavage sequences are F2A (foot and mouth disease virus 2A) or "2A-like" sequences (e.g., Thosea asigna virus 2A; T2A) (Minskaia and Ryan, 2013). In certain embodiments, F2A cleavage peptides are used to link expression of genes in multilineage constructs.

[0149] c. Origin of replication To propagate the vector in a host cell, the vector may contain one or more origin of replication sites (often referred to as "ori"), such as a nucleic acid sequence corresponding to the oriP of EBV as described above, or an engineered oriP (which is a specific nucleic acid sequence at which replication is initiated) with a similar or improved function in programming. Alternatively, origins of replication of other extrachromosomally replicating viruses, as described above, or autonomously replicating sequences (ARS) can be used.

[0150] d. Selectable and Screenable Markers In certain embodiments, cells containing the nucleic acid construct can be identified in vitro or in vivo by including a marker in the expression vector. Such a marker will confer an identifiable change to the cell that allows easy identification of cells containing the expression vector. In general, a selection marker is one that confers a property that allows for selection. A positive selection marker is one whose presence allows for selection, and a negative selection marker is one whose presence prevents selection. An example of a positive selection marker is a drug resistance marker.

[0151] Typically, the inclusion of a drug selection marker facilitates cloning and identification of transformants; for example, genes that confer resistance to neomycin, puromycin, hygromycin, DHFR, GPT, zeocin, and histidinol are useful selection markers. In addition to markers that confer a phenotype that allows for the identification of transformants based on the implementation of conditions, other types of markers are contemplated, including colorimetric-based screenable markers, such as GFP. Alternatively, screenable enzymes as negative selection markers, such as herpes simplex virus thymidine kinase (tk) or chloramphenicol acetyltransferase (CAT), may be utilized. Those skilled in the art will also know how to use immunological markers, possibly in conjunction with FACS analysis. The marker used is not believed to be important, so long as it is capable of being expressed simultaneously with the nucleic acid encoding the gene product. Further examples of selection and screenable markers are well known to those skilled in the art.

[0152] E. Gene Editing In some embodiments, the methods include gene editing with sequence-specific or targeted nucleases, including DNA-binding targeted nucleases, such as zinc finger nucleases (ZFNs) and transcription activator-like effector nucleases (TALENs), and RNA-guided nucleases, such as CRISPR-associated nucleases (Cas), that are specifically designed to be targeted to a sequence of a gene or a portion thereof.

[0153] In some embodiments, gene editing is performed by inducing one or more double-strand breaks and / or one or more single-strand breaks in gene, typically in a targeted manner.In some embodiments, double-strand breaks or single-strand breaks are made by nuclease, for example, endonuclease, for example, gene-targeting nuclease.In some aspects, breaks are induced in the coding region of gene, for example, in exon.For example, in some embodiments, induction occurs near the N-terminal part of coding region, for example, in the first exon, the second exon, or the subsequent exon.

[0154] In some aspects, double-stranded or single-stranded breaks are repaired by cellular repair processes, such as by non-homologous end joining (NHEJ) or homology-directed repair (HDR). In some aspects, the repair process is error-prone and can result in gene disruption, such as frameshift mutations, e.g., biallelic frameshift mutations, resulting in a complete knockout of the gene.

[0155] In some embodiments, gene editing is achieved using DNA targeting molecules, such as DNA binding proteins or DNA binding nucleic acids that specifically bind or hybridize to genes, or complexes, compounds or compositions containing them. In some embodiments, the DNA targeting molecules include DNA binding domains, such as zinc finger protein (ZFP) DNA binding domains, transcription activator-like proteins (TAL) or TAL effector (TALE) DNA binding domains, clustered regularly interspaced short palindromic repeats (CRISPR) DNA binding domains, or DNA binding domains from meganucleases. Zinc finger, TALE, and CRISPR system binding domains can be engineered to bind to a given nucleotide sequence, for example, by engineering (changing one or more amino acids) the recognition helix region of a naturally occurring zinc finger or TALE protein. Engineered DNA binding proteins (zinc finger or TALE) are non-naturally occurring proteins. Rational design criteria include the application of substitution rules and computerized algorithms to process information in databases that store information on existing ZFP and / or TALE design and binding data. See, e.g., U.S. Patent Nos. 6,140,081; 6,453,242; and 6,534,261; see also WO98 / 53058; WO98 / 5305; WO98 / 5306; WO02 / 01653 and WO03 / 01649, and U.S. Patent Application Publication No. 2011 / 0301073.

[0156] In some embodiments, the DNA targeting molecule, complex, or combination contains a DNA binding molecule and one or more additional domains, such as an effector domain that facilitates gene suppression or destruction.For example, in some embodiments, gene editing is performed by a fusion protein that comprises a DNA binding protein and a heterologous regulatory domain or a functional fragment thereof.In some aspects, the domain comprises, for example, a transcription factor domain, such as an activator, a repressor, a coactivator, a corepressor, a silencer, an oncogene, a DNA repair enzyme and its associated factors and modifiers, a DNA reorganization enzyme and its associated factors and modifiers, a chromatin-associated protein and its modifiers, such as a kinase, an acetylase and a deacetylase, and a DNA modifying enzyme, such as a methyltransferase, a topoisomerase, a helicase, a ligase, a kinase, a phosphatase, a polymerase, an endonuclease, and its associated factors and modifiers. For details on the fusion of DNA binding domain and nuclease cleavage domain, see, for example, US Patent Application Publication Nos. 2005 / 0064474; 2006 / 0188987 and 2007 / 0218528, which are incorporated herein by reference in their entirety.In some aspects, the additional domain is a nuclease domain.Thus, in some embodiments, gene editing is facilitated by gene or genome editing using engineered proteins, such as nucleases and nuclease-containing complexes or fusion proteins (composed of sequence-specific DNA binding domains that form fusions or complexes with non-specific DNA cleavage molecules, such as nucleases).

[0157] In some aspects, these targeted chimeric nucleases or nuclease-containing complexes perform precise genetic modification by inducing targeted double-stranded or single-stranded breaks, stimulating cellular DNA repair mechanisms, including error-prone non-homologous end joining (NHEJ) and homology-directed repair (HDR). In some embodiments, the nuclease is an endonuclease, such as zinc finger nuclease (ZFN), TALE nuclease (TALEN), and RNA-guided endonuclease (RGEN), such as CRISPR-associated (Cas) protein, or meganuclease.

[0158] In some embodiments, a donor nucleic acid, e.g., a donor plasmid or nucleic acid encoding an engineered antigen receptor, is provided and inserted into the gene editing site after introduction of a DSB by HDR. Thus, in some embodiments, gene disruption and introduction of an antigen receptor, e.g., a CAR, are performed simultaneously, whereby the gene is partially disrupted by knock-in or insertion of a nucleic acid encoding a CAR.

[0159] In some embodiments, no donor nucleic acid is provided. In some aspects, NHEJ-mediated repair following introduction of a DSB results in an insertion or deletion mutation that can cause gene disruption, for example by generating a missense mutation or a frameshift.

[0160] 1. ZFPs and ZFNs In some embodiments, the DNA targeting molecule comprises a DNA binding protein, such as one or more zinc finger proteins (ZFPs) or transcription activator-like proteins (TALs), fused to an effector protein, such as an endonuclease. Examples include ZFNs, TALEs, and TALENs.

[0161] In some embodiments, the DNA targeting molecule comprises one or more zinc finger proteins (ZFPs) or domains thereof that bind to DNA in a sequence-specific manner. A ZFP or domain thereof is a protein or a domain within a larger protein that binds to DNA in a sequence-specific manner by one or more zinc fingers, which are regions of amino acid sequence within the binding domain whose structure is stabilized by the coordination of a zinc ion. The term zinc finger DNA binding protein is often abbreviated as zinc finger protein or ZFP. Among the ZFPs are artificial ZFP domains that are generated by the assembly of individual fingers to target specific DNA sequences, typically 9-18 nucleotides in length.

[0162] ZFPs include those in which a single finger domain is approximately 30 amino acids long, contains two invariant histidine residues that coordinate with two cysteines in a beta turn via zinc, and contains an alpha helix with two, three, four, five, or six fingers. In general, the sequence specificity of a ZFP may be altered by making amino acid substitutions at the four helical positions (-1, 2, 3, and 6) on the zinc finger recognition helix. Thus, in some embodiments, the ZFP or ZFP-containing molecule is not naturally occurring and has been engineered, e.g., to bind to a selected target site.

[0163] In some aspects, disruption of MeCP2 occurs by contacting a first target site in the gene with a first ZFP, thereby disrupting the gene. In some embodiments, the target site in the gene is contacted with a fusion ZFP that includes six fingers and a regulatory domain, thereby inhibiting expression of the gene.

[0164] In some embodiments, the contacting step further comprises contacting a second target site in the gene with a second ZFP. In some aspects, the first and second target sites are adjacent. In some embodiments, the first and second ZFPs are covalently linked. In some aspects, the first ZFP is a fusion protein comprising a regulatory domain or at least two regulatory domains.

[0165] In some embodiments, the first and second ZFPs are fusion proteins each comprising one regulatory domain or each comprising at least two regulatory domains, hi some embodiments, the regulatory domains are transcriptional repressors, transcriptional activators, endonucleases, methyltransferases, histone acetyltransferases, or histone deacetylases.

[0166] In some embodiments, the ZFP is encoded by a ZFP nucleic acid operably linked to a promoter. In some aspects, the method further comprises initially administering the nucleic acid to the cell in a lipid:nucleic acid complex or as a naked nucleic acid. In some embodiments, the ZFP is encoded by an expression vector comprising a ZFP nucleic acid operably linked to a promoter. In some embodiments, the ZFP is encoded by a nucleic acid operably linked to an inducible promoter. In some embodiments, the ZFP is encoded by a nucleic acid operably linked to a weak promoter.

[0167] In some embodiments, the target site is upstream of the transcription start site of the gene. In some aspects, the target site is adjacent to the transcription start site of the gene. In some embodiments, the target site is adjacent to the RNA polymerase pause site downstream of the transcription start site of the gene.

[0168] In some embodiments, the DNA targeting molecule is or comprises a zinc finger DNA binding domain that is fused to a DNA cleavage domain to form a zinc finger nuclease (ZFN). In some embodiments, the fusion protein comprises a cleavage domain (or cleavage half-domain) derived from at least one liS-type restriction enzyme and one or more zinc finger binding domains, which may or may not be engineered. In some embodiments, the cleavage domain is derived from Fok I, a liS-type restriction endonuclease. Fok I catalyzes double-stranded cleavage of DNA, typically 9 nucleotides from the recognition site on one strand and 13 nucleotides from the recognition site on the other strand.

[0169] In some embodiments, ZFN targets genes present in engineered cells. In some aspects, ZFN efficiently generates double-strand breaks (DSBs), for example, at a predetermined site in the coding region of a gene. Typical regions targeted include exons, regions encoding N-terminal regions, first exons, second exons, and promoter or enhancer regions. In some embodiments, transient expression of ZFNs promotes highly efficient and permanent destruction of target genes in engineered cells. Notably, in some embodiments, delivery of ZFNs results in permanent destruction of genes with efficiency of over 50%.

[0170] Many gene-specific engineered zinc fingers are commercially available. For example, Sangamo Biosciences (Richmond, CA, USA) in collaboration with Sigma-Aldrich (St. Louis, MO, USA) has developed a platform for zinc finger construction (CompoZr) that allows researchers to avoid the construction and validation of zinc fingers and provides specifically targeted zinc fingers for thousands of proteins (Gaj et al., Trends in Biotechnology, 2013, vol. 31(7), pp. 397-405). In some embodiments, commercially available zinc fingers are used or custom-made.

[0171] 2. TALs, TALEs and TALENs In some embodiments, the DNA targeting molecule comprises a naturally occurring or engineered (non-naturally occurring) transcription activator-like protein (TAL) DNA binding domain (e.g., as in a transcription activator-like protein effector (TALE) protein), see, e.g., U.S. Patent Application Publication No. 2011 / 0301073, which is incorporated by reference in its entirety.

[0172] A TALE DNA binding domain or TALE is a polypeptide that contains one or more TALE repeat domains / units. The repeat domain is responsible for the binding of a TALE to its cognate target DNA sequence. A single "repeat unit" (also called a "repeat") is typically 33-35 amino acids long and shows at least some sequence homology with other TALE repeat sequences in naturally occurring TALE proteins. Each TALE repeat unit generally contains one or two DNA binding residues at positions 12 and / or 13 of this repeat that constitute the Repeat Variable Diresidue (RVD). The natural (canonical) codes for DNA recognition of these TALEs were determined such that the HD sequence at positions 12 and 13 results in binding to cysteine ​​(C), NG binds to T, NI binds to A, NN binds to G or A, and NG binds to T, with non-canonical (non-classical) RVDs also being known. See US Patent Publication No. 2011 / 0301073. In some embodiments, TALEs can be targeted to any gene by designing TAL arrays with specificity for the target DNA sequence. The target sequence generally begins with a thymidine.

[0173] In some embodiments, the molecule is a DNA-binding endonuclease, such as a TALE nuclease (TALEN). In some aspects, a TALEN is a fusion protein that includes a DNA-binding domain and a nuclease catalytic domain derived from a TALE to cleave a nucleic acid target sequence.

[0174] In some embodiments, TALENs recognize and cleave target sequences within genes. In some aspects, DNA cleavage results in double-strand breaks. In some aspects, this cleavage stimulates the rate of homologous recombination or non-homologous end joining (NHEJ). In general, NHEJ is an incomplete repair process that often results in changes at the cleavage site of the DNA sequence. In some aspects, the repair mechanism involves rejoining the remaining parts of the two DNA ends by direct religation (Critchlow and Jackson, 1998) or by so-called microhomology-mediated end joining. In some embodiments, repair by NHEJ results in small insertions or deletions that can be used to disrupt and thereby silence genes. In some embodiments, the modification can be a substitution, deletion, or addition of at least one nucleotide. In some aspects, cells that have undergone a break-induced mutagenesis event, i.e., a mutagenesis event subsequent to an NHEJ event, can be identified and / or selected by methods well known in the art.

[0175] In some embodiments, TALE repeats are assembled to specifically target genes. A library of TALENs has been constructed that targets 18,740 human protein-coding genes. Custom TALE arrays are commercially available from Cellectis Bioresearch (Paris, France), Transposagen Biopharmaceuticals (Lexington, KY, USA), and Life Technologies (Grand Island, NY, USA).

[0176] In some embodiments, the TALENs are introduced as transgenes encoded by one or more plasmid vectors. In some aspects, the plasmid vectors can contain selectable markers that allow for identification and / or selection of cells that have received the vector.

[0177] 3. RGEN (CRISPR / Cas system) In some embodiments, the disruption, for example, the disruption via RNA-guided endonuclease (RGEN), is carried out using one or more DNA-binding nucleic acids. For example, the disruption can be carried out using CRISPR (clustered regularly interspaced short palindromic repeats) and CRISPR-associated (Cas) proteins. In general, "CRISPR system" refers collectively to the transcripts and other elements involved in the expression or directing the activity of CRISPR-associated ("Cas") genes, including sequences encoding Cas genes, tracr (transactivating CRISPR) sequences (e.g., tracrRNA or active portion tracrRNA), tracr-mate sequences (including "direct repeats" and tracrRNA processing portion direct repeats in the context of endogenous CRISPR systems), guide sequences (also called "spacers" in the context of endogenous CRISPR systems), and / or other sequences and transcripts from CRISPR loci.

[0178] A CRISPR / Cas nuclease or CRISPR / Cas nuclease system can include a non-coding RNA molecule (guide) RNA that binds to DNA in a sequence-specific manner, and a Cas protein (e.g., Cas9) that has a nuclease function (e.g., two nuclease domains). One or more elements of the CRISPR system can be derived from a type I, type II, or type III CRISPR system, for example, a particular organism that contains an endogenous CRISPR system, such as Streptococcus pyogenes.

[0179] In some embodiments, a Cas nuclease and a gRNA (comprising a fusion of a crRNA specific for a target sequence and a fixed tracrRNA) are introduced into a cell. Generally, a target site at the 5' end of the gRNA uses complementary base pairing to target the Cas nuclease to a target site, e.g., a gene. The target site can be selected based on the location immediately 5' of a protospacer adjacent motif (PAM) sequence, e.g., typically NGG or NAG. In this regard, the gRNA is targeted to a desired sequence by modifying the first 20, 19, 18, 17, 16, 15, 14, 12, 11, or 10 nucleotides of the guide RNA to correspond to the target DNA sequence. In general, CRISPR systems feature elements that promote the formation of a CRISPR complex at the site of the target sequence. Typically, a "target sequence" generally refers to a sequence to which the guide sequence is designed to have complementarity, and hybridization between the target sequence and the guide sequence promotes the formation of a CRISPR complex. Absolute complementarity is not required, just sufficient complementarity for hybridization to occur and promote formation of the CRISPR complex.

[0180] The CRISPR system can induce a double-strand break (DSB) at the target site, followed by disruption as discussed herein. In other embodiments, a Cas9 variant that is believed to be a "nickase" is used to nick a single strand at the target site. Paired nickases can be used, for example, to improve specificity, directed by a pair of different gRNAs targeting sequences such that a 5' overhang is introduced when a nick is introduced simultaneously. In other embodiments, a catalytically inactive Cas9 is fused to a heterologous effector domain, such as a transcriptional repressor or activator, to affect gene expression.

[0181] The target sequence may comprise any polynucleotide, for example, a DNA or RNA polynucleotide. The target sequence may be located in the nucleus or cytoplasm of a cell, for example, in an organelle of a cell. In general, a sequence or template that can be used for recombination into a target locus that comprises a target sequence is referred to as an "editing template" or an "editing polynucleotide" or an "editing sequence". In some aspects, an exogenous template polynucleotide may be referred to as an editing template. In some aspects, the recombination is a homologous recombination.

[0182] Typically, in the context of an endogenous CRISPR system, a CRISPR complex (comprising a guide sequence hybridized to a target sequence and complexed with one or more Cas proteins) induces cleavage of one or both strands within or near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs from the target sequence). The tracr sequence may comprise or consist of all or a portion of the wild-type tracr sequence (e.g., about 20, 26, 32, 45, 48, 54, 63, 67, 85, or more nucleotides or more of the wild-type tracr sequence), and may form part of a CRISPR complex by hybridizing to all or a portion of the tracr mate sequence operably linked to the guide sequence, together with at least a portion of the tracr sequence. The tracr sequence has sufficient complementarity to the tracr mate sequence to hybridize and participate in the formation of a CRISPR complex, e.g., at least 50%, 60%, 70%, 80%, 90%, 95% or 99% sequence complementarity along the length of the tracr mate sequence when optimally aligned.

[0183] One or more vectors driving the expression of one or more elements of the CRISPR system can be introduced into a cell such that the expression of the elements of the CRISPR system directs the formation of a CRISPR complex at one or more target sites. The components can also be delivered to the cell as proteins and / or RNA. For example, the Cas enzyme, the guide sequence linked to the tracr mate sequence, and the tracr sequence can each be operably linked to separate control elements on separate vectors. Alternatively, two or more of the elements expressed from the same or different regulatory elements can be combined into a single vector, with one or more additional vectors providing any components of the CRISPR system not included in the first vector. The vector can include one or more insertion sites, such as restriction endonuclease recognition sequences (also called "cloning sites"). In some embodiments, the one or more insertion sites are located upstream and / or downstream of one or more sequence elements of the one or more vectors. When multiple different guide sequences are used, a single expression construct can be used to target CRISPR activity to multiple different corresponding target sequences in a cell.

[0184] The vector may include regulatory elements operably linked to an enzyme coding sequence encoding a CRISPR enzyme, e.g., a Cas protein. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csfl, Csf2, Csf3, Csf4, homologs thereof, or modified versions thereof. These enzymes are known; for example, the amino acid sequence of the Cas9 protein of Streptococcus pyogenes can be found in the SwissProt database under accession number Q99ZW2.

[0185] The CRISPR enzyme can be Cas9 (e.g., from Streptococcus pyogenes or S. pneumonia). The CRISPR enzyme can direct cleavage of one or both strands at the location of the target sequence, such as within the target sequence and / or within the complement of the target sequence. The vectors are CRISPR enzymes mutated relative to the corresponding wild-type enzyme, such that the mutated CRISPR enzyme lacks the ability to cleave one or both strands of the target polynucleotide containing the target sequence. For example, an aspartic acid to alanine substitution (D10A) in the RuvC I catalytic domain of Cas9 from Streptococcus pyogenes converts Cas9 from a nuclease that cleaves both strands to a nickase (that cleaves a single strand). In some embodiments, the Cas9 nickase can be used in combination with a guide sequence, for example, two guide sequences that target the sense and antisense strands of a DNA target, respectively. This combination allows both strands to be nicked and used to induce NHEJ or HDR.

[0186] In some embodiments, the enzyme coding sequence encoding the CRISPR enzyme is codon-optimized for expression in a particular cell, e.g., a eukaryotic cell. The eukaryotic cell may be or be derived from a particular organism, e.g., a mammal, including but not limited to, human, mouse, rat, rabbit, dog, or non-human primate. In general, codon optimization refers to the process of modifying a nucleic acid sequence to enhance expression in a host cell of interest by replacing at least one codon of the native sequence with a codon that is more or most frequently used in the host cell's genes while maintaining the native amino acid sequence. Different species show a particular bias towards certain codons for certain amino acids. Codon bias (differences in codon usage between organisms) often correlates with the efficiency of translation of messenger RNA (mRNA), which in turn is believed to depend, among other things, on the properties of the codon being translated and the availability of certain transfer RNA (tRNA) molecules. The predominance of tRNAs selected in a cell generally reflects the codons most frequently used in peptide synthesis. Thus, genes can be adjusted based on codon optimization for optimal gene expression in a given organism.

[0187] In general, a guide sequence is any polynucleotide sequence that has sufficient complementarity with a target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of a CRISPR complex to the target sequence. In some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence is or exceeds about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or higher percentage when optimally aligned using a suitable alignment algorithm.

[0188] Optimal alignment may be determined using any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (e.g., Burrows Wheeler Aligner), Clustal W, Clustal X, BLAT, Novoalign (Novocraft Technologies), ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net).

[0189] CRISPR enzymes can be part of a fusion protein that includes one or more heterologous protein domains. CRISPR enzyme fusion proteins can include any additional protein sequences and possibly linker sequences between any two domains. Examples of protein domains that can be fused to CRISPR enzymes include, but are not limited to, epitope tags, reporter gene sequences, and protein domains that have one or more of the following activities: methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, and nucleic acid binding activity. Non-limiting examples of epitope tags include histidine (His) tag, V5 tag, FLAG tag, influenza hemagglutinin (HA) tag, Myc tag, VSV-G tag, and thioredoxin (Trx) tag. Examples of reporter genes include, but are not limited to, glutathione-5-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT) beta-galactosidase, beta-glucuronidase, luciferase, green fluorescent protein (GFP), autofluorescent proteins including HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and blue fluorescent protein (BFP). CRISPR enzymes may be fused to genetic sequences encoding proteins or fragments of proteins that bind to DNA molecules or other cellular molecules, including, but not limited to, maltose binding protein (MBP), S-tags, Lex A DNA binding domain (DBD) fusions, GAL4A DNA binding domain fusions, and herpes simplex virus (HSV) BP16 protein fusions.

[0190] F. Differentiation of iPSCs In some embodiments, a method is provided for generating differentiated cells from an essentially single cell suspension of pluripotent stem cells (PSCs), such as human iPSCs. In some embodiments, the PSCs are cultured preconfluently to prevent any cell aggregation. In certain aspects, the PSCs are dissociated by incubation with a cell dissociation enzyme, as exemplified by TRYPSIN™ or TRYPLE™. The PSCs can also be dissociated into an essentially single cell suspension by pipetting. Additionally, blebbistatin (e.g., about 2.5 μM) can be added to the medium to enhance the viability of the PSCs after dissociation into single cells, when the cells are not attached to the culture vessel. A ROCK inhibitor may be used instead of blebbistatin to enhance the viability of the PSCs after dissociation into single cells.

[0191] Once a single cell suspension of PSCs is obtained at a known cell density, the cells are generally seeded into a suitable culture vessel, such as a tissue culture plate, such as a flask, 6-well, 24-well, 96-well plate. Culture vessels used to culture the cells include, but are not limited to, flasks, tissue culture flasks, dishes, petri dishes, tissue culture dishes, multi-dishes, microplates, microwell plates, multi-plates, multi-well plates, microslides, chamber slides, tubes, trays, CELLSTACK® chambers, culture bags, and roller bottles, so long as the stem cells can be cultured therein. Cells may be cultured in a volume of at least or about 0.2, 0.5, 1, 2, 5, 10, 20, 30, 40, 50 ml, 100 ml, 150 ml, 200 ml, 250 ml, 300 ml, 350 ml, 400 ml, 450 ml, 500 ml, 550 ml, 600 ml, 800 ml, 1000 ml, 1500 ml, or any range derivable therein, depending on the needs of the culture. In certain embodiments, the culture vessel may be a bioreactor, which may refer to any ex vivo device or system that supports a biologically active environment in which cells can grow. The bioreactor may have a volume of at least or about 2, 4, 5, 6, 8, 10, 15, 20, 25, 50, 75, 100, 150, 200, 500 liters, 1, 2, 4, 6, 8, 10, 15 cubic meters, or any range derivable therein.

[0192] In certain embodiments, PSCs, e.g., iPSCs, are seeded at a cell density suitable for efficient differentiation. Generally, cells are seeded at a density of about 1,000 to about 75,000 cells / cm. 2 cell density, for example, about 5,000 to about 40,000 cells / cm 2 In a 6-well plate, the cells may be seeded at a cell density of about 50,000 to about 400,000 cells per well. In an exemplary method, the cells are seeded at a cell density of about 100,000, about 150,000, about 200,000, about 250,000, about 300,000, or about 350,000 cells per well, e.g., about 200,000 (200,000) cells per well.

[0193] PSCs, e.g., iPSCs, are generally cultured on culture plates coated with one or more cell adhesion proteins to promote cell attachment while maintaining cell viability. For example, preferred cell adhesion proteins include extracellular matrix proteins, e.g., vitronectin, laminin, collagen, and / or fibronectin, which can be used to coat culture surfaces as a means of providing a solid support for pluripotent cell growth. The term "extracellular matrix" is recognized in the art. Its components include one or more of the following proteins: fibronectin, laminin, vitronectin, tenascin, entactin, thrombospondin, elastin, gelatin, collagen, fibrillin, merosin, anchorin, chondronectin, link protein, bone sialoprotein, osteocalcin, osteopontin, epinectin, hyaluronectin, undrin, epiligrin, and kalinin. In an exemplary method, PSCs are grown on culture plates coated with vitronectin or fibronectin. In some embodiments, the cell adhesion protein is a human protein.

[0194] Extracellular matrix (ECM) proteins may be naturally derived and purified from human or animal tissues, or ECM proteins may be genetically engineered recombinant proteins or synthetic in nature. ECM proteins may be in the form of whole proteins or peptide fragments, and may be natural or engineered. Examples of ECM proteins that may be useful in matrices for cell culture include laminin, collagen I, collagen IV, fibronectin, and vitronectin. In some embodiments, the matrix composition comprises synthetically produced peptide fragments of fibronectin or recombinant fibronectin. In some embodiments, the matrix composition is xeno-free. For example, xeno-free matrices for culturing human cells may use matrix components of human origin and may exclude any components of non-human animals.

[0195] In some aspects, the total protein concentration in the matrix composition can be about 1 ng / mL to about 1 mg / mL. In some preferred embodiments, the total protein concentration in the matrix composition is about 1 μg / mL to about 300 μg / mL. In more preferred embodiments, the total protein concentration in the matrix composition is about 5 μg / mL to about 200 μg / mL.

[0196] Cells can be cultured with the nutrients necessary to direct the growth of each particular cell population. Generally, cells are cultured in a growth medium that includes a carbon source, a nitrogen source, and a buffer to maintain pH. The medium may also contain fatty acids or lipids, amino acids (e.g., non-essential amino acids), vitamins, growth factors, cytokines, antioxidants, pyruvate, buffers, and inorganic salts. Exemplary growth media contain minimal essential media, such as Dulbecco's Modified Eagle Medium (DMEM) or ESSENTIAL 8™ (E8™) medium, supplemented with various nutrients, such as non-essential amino acids and vitamins, to enhance the growth of stem cells. Examples of minimal essential media include, but are not limited to, Minimum Essential Medium Eagle (MEM) Alpha Medium, Dulbecco's Modified Eagle Medium (DMEM), RPMI-1640 Medium, 199 Medium, and F12 Medium. Additionally, the minimal essential medium may be supplemented with additives, such as horse, bovine, or fetal bovine serum. Alternatively, the medium may be serum-free. In other cases, the growth medium may contain "Knockout serum replacement," referred to herein as a serum-free formulation optimized for culturing, growing and maintaining undifferentiated cells, such as stem cells. KNOCKOUT™ serum replacement is disclosed, for example, in U.S. Patent Application Publication No. 2002 / 0076747, which is incorporated herein by reference. Preferably, PSCs are cultured in a well-defined feeder-free medium.

[0197] Thus, PSCs are generally cultured in a well-defined culture medium after seeding. In certain embodiments, about 18-24 hours after seeding, the medium is aspirated and fresh medium, such as E8™ medium, is added to the culture. In certain embodiments, single-cell PSCs are cultured in a well-defined culture medium for about 1, 2, or 3 days after seeding. Preferably, single-cell PSCs are cultured in a well-defined culture medium for about 2 days before proceeding with the differentiation process.

[0198] In some embodiments, the medium may or may not contain any substitute for serum. Substitutes for serum may include albumin (e.g., lipid-rich albumin, albumin substitutes such as recombinant albumin, vegetable starch, dextran, and protein hydrolysates), transferrin (or other iron transporters), fatty acids, insulin, collagen precursors, trace elements, 2-mercaptoethanol, 3'-thiolgiycerol, or materials that suitably contain equivalents thereof. Substitutes for serum may be prepared, for example, by methods disclosed in International Publication No. WO 98 / 30679. Alternatively, more conveniently, any commercially available material may be used. Commercially available materials include KNOCKOUT™ serum substitute (KSR), concentrated chemically defined lipids (Gibco), and GLUTAMAX™ (Gibco).

[0199] Other culture conditions can be defined as appropriate. For example, the culture temperature can be about 30-40°C, for example, at least about 31°C, about 32°C, about 33°C, about 34°C, about 35°C, about 36°C, about 37°C, about 38°C, about 39°C, but is not particularly limited thereto. In one embodiment, the cells are cultured at 37°C. The CO2 concentration can be about 1-10%, for example, about 2-5%, or any range derived therefrom. The oxygen partial pressure can be at least, up to, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20%, or any range derived therefrom.

[0200] G. Cryopreservation of iPSCs or Differentiated Cells The cells produced by the methods disclosed herein can be cryopreserved, see, e.g., PCT Publication No. 2012 / 149484 A2, which is incorporated herein by reference. The cells can be cryopreserved with or without a substrate. In some embodiments, the storage temperature is in the range of about -50°C to about -60°C, about -60°C to about -70°C, about -70°C to about -80°C, about -80°C to about -90°C, about -90°C to about -100°C, and overlapping ranges thereof. In some embodiments, lower temperatures are used for storage (e.g., maintenance) of the cryopreserved cells. In some embodiments, liquid nitrogen (or other similar liquid coolant) is used to store the cells. In further embodiments, the cells are stored for more than about 6 hours. In additional embodiments, the cells are stored for about 72 hours. In some embodiments, the cells are stored for 48 hours to about 1 week. In still other embodiments, the cells are stored for about 1, 2, 3, 4, 5, 6, 7, or 8 weeks. In further embodiments, the cells are stored for 1, 2, 3, 4, 5, 67, 8, 9, 10, 11, or 12 months. The cells can also be stored for longer periods. The cells can be cryopreserved separately or on a substrate, such as any of the substrates disclosed herein.

[0201] In some embodiments, additional cryoprotectants can be used. For example, cells can be cryopreserved in a cryopreservation solution containing one or more cryoprotectants, such as DM80, serum albumin, such as human or bovine serum albumin. In certain embodiments, the solution contains about 1%, about 1.5%, about 2%, about 2.5%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, or about 10% DMSO. In other embodiments, the solution contains about 1% to about 3%, about 2% to about 4%, about 3% to about 5%, about 4% to about 6%, about 5% to about 7%, about 6% to about 8%, about 7% to about 9%, or about 8% to about 10% dimethyl sulfoxide (DMSO) or albumin. In a specific embodiment, the solution contains 2.5% DMSO. In another specific embodiment, the solution contains 10% DMSO.

[0202] The cells may be cooled at, for example, about 1° C. / min during cryopreservation. In some embodiments, the cryopreservation temperature is about −80° C. to about −180° C., or about −125° C. to about −140° C. In some embodiments, the cells are cooled to 4° C. before being cooled at about 1° C. / min. The cryopreserved cells may be transferred to the vapor phase of liquid nitrogen before thawing for use. In some embodiments, for example, once the cells reach about −80° C., they are transferred to liquid nitrogen storage. Cryopreservation may also be performed using a rate-controlled freezer. The cryopreserved cells may be thawed at, for example, a temperature of about 25° C. to about 40° C., typically at a temperature of about 37° C.

[0203] III. Use of Engineered Cell Lines Certain aspects provide methods for generating cell lines with stable transgene expression that can be used for a number of important research, development, and commercial purposes.

[0204] The cell line generated by the method disclosed herein can be used in any method and application currently known in the art for iPSC or differentiated cells.For example, a method for evaluating a compound can be provided, comprising assaying the pharmacological or toxicological properties of the compound in the cell line.Also provided is a method for evaluating a compound for its effect on a cell culture, comprising a) contacting the cell culture provided herein with a compound, and b) assaying the effect of the compound on the cell culture.

[0205] A. Screening of Test Compounds Cell cultures can be used commercially to screen factors (e.g., solvents, small molecule drugs, peptides, oligonucleotides) or environmental conditions (e.g., culture conditions or manipulations) that affect the properties of such cells and their various progeny. For example, test compounds can be chemical compounds, small molecules, polypeptides, growth factors, cytokines, or other biological agents.

[0206] In one embodiment, the method includes contacting the cell culture with a test agent and determining whether the test agent modulates the activity or function of cells in the population. In some applications, the screening assay is used to identify agents that modulate cell proliferation, alter cell differentiation, or affect cell viability. The screening assay can be performed in vitro or in vivo. Methods for screening and identifying candidate agents include those suitable for high-throughput screening. For example, cell cultures can be arranged or placed on culture dishes, flasks, roller bottles, or plates (e.g., single multi-well dishes or dishes, such as 8, 16, 32, 64, 96, 384, and 1536 multi-well plates or dishes), optionally at defined locations, for the identification of potential therapeutic molecules. Libraries that can be screened include, for example, small molecule libraries, siRNA libraries, and adenovirus transfection vector libraries.

[0207] Other screening applications involve testing the effect of pharmaceutical compounds on the maintenance or repair of retinal tissue. Screens may be performed because compounds are designed to have a pharmacological effect on cells, or because compounds designed to have an effect elsewhere may have unintended side effects on cells of this tissue type.

[0208] B. Treatment and Transplantation Other embodiments may provide for the use of the cell lines for the treatment of a disease or disorder. In another aspect, the disclosure provides a method of treating an individual in need thereof comprising administering to said individual a composition comprising engineered cells.

[0209] To determine the suitability of a cell composition for therapeutic administration, the cells can first be tested in a suitable animal model. In one embodiment, the cell line is evaluated for its ability to survive and maintain its phenotype in vivo. The composition is implanted into an immunodeficient animal (e.g., a nude mouse or an animal that has been rendered immunodeficient by chemical or radiation exposure). After a period of growth, tissue is harvested to assess whether pluripotent stem cell-derived cells are still present.

[0210] As used herein, a disease or disorder refers to a pathological condition in an organism, for example due to an infection or genetic defect, characterized by an identifiable symptom. An exemplary disease described herein is a neoplastic disease, for example cancer. As used herein, a neoplastic disease refers to any disorder involving cancer, including tumor initiation, growth, metastasis and progression.

[0211] As used herein, cancer is a term that refers to diseases caused by or characterized by any type of malignancy, including metastatic cancer, lymphatic tumors, and hematological cancer. Exemplary cancers include, but are not limited to, leukemia, lymphoma, pancreatic cancer, lung cancer, ovarian cancer, breast cancer, cervical cancer, bladder cancer, prostate cancer, glioma tumors, adenocarcinoma, liver cancer, and skin cancer. Exemplary cancers in humans include bladder tumors, breast tumors, prostate tumors, basal cell carcinoma, biliary tract cancer, bladder cancer, bone cancer, brain and CNS cancer (e.g., glioma tumors), cervical cancer, choriocarcinoma, colon and rectal cancer, connective tissue cancer, cancer of the digestive system; endometrial cancer, esophageal cancer; eye cancer; head and neck cancer; gastric cancer; intraepithelial neoplasia; kidney cancer; laryngeal cancer; leukemia; liver cancer; lung cancer (e.g., small cell carcinoma and non-small cell carcinoma); Hodgkin's lymphoma and non-Hodgkin's lymphoma; melanoma; myeloma, neuroblastoma, oral cancer (e.g., lip, tongue, oral cavity, and larynx); ovarian cancer, pancreatic cancer, retinoblastoma; rhabdomyosarcoma; rectal cancer, renal cancer, cancer of the respiratory system; sarcoma, skin cancer; stomach cancer; testicular cancer, thyroid cancer, uterine cancer, cancer of the urinary system, and other carcinomas and sarcomas. Exemplary cancers commonly diagnosed in dogs, cats, and other pets include, but are not limited to, lymphosarcoma, osteosarcoma, mammary tumors, mast cell tumors, brain tumors, melanoma, adenosquamous carcinoma, carcinoid lung tumors, bronchial adenocarcinomas, bronchopulmonary adenocarcinomas, fibromas, myxochondromas, pulmonary sarcomas, neurosarcomas, osteomas, papillomas, retinoblastomas, Ewing's sarcoma, Wilms' tumor, Burkitt's lymphoma, microgliomas, neuroblastomas, osteoclastomas, osteoporosis ... Cancers diagnosed in rodents, such as ferrets, include, but are not limited to, insulinoma, lymphoma, sarcoma, sarcoma, schwannoma, pancreatic islet cell tumor, gastric adenocarcinoma, sarcoma ...Exemplary neoplasms affecting agricultural livestock include, but are not limited to, leukemia, hemangiopericytoma, and bovine ocular neoplasms (in cattle); vestibular fibrosarcoma, ulcerating squamous cell carcinoma, vestibular carcinoma, connective tissue neoplasms, and mast cell tumors (in horses); hepatocellular carcinoma (in pigs); lymphoma and pulmonary adenomatosis (in sheep); pulmonary sarcoma, lymphoma, Rous sarcoma, reticuloendothelioma, fibrosarcoma, nephroblastoma, B-cell lymphoma, and lymphocytic leukemia (in avian species); retinoblastoma, hepatic neoplasms, lymphosarcoma (lymphoblastic lymphoma), plasma cell leukemia, and swimbladder sarcoma (in fish), caseous lymphadenitis (CLA): Corynebacterium pseudotuberculosis (Corynebacterium sarcoma), These include a chronic, infectious contagious disease of sheep and goats caused by the bacterium B. pseudotuberculosis, and a contagious lung tumor of sheep caused by B. jaegersii.

[0212] Pharmaceutical compositions of the cell lines produced by the methods disclosed herein are also provided. These compositions contain at least about 1×10 3 cells, approximately 1 x 10 4 cells, approximately 1 x 10 5 cells, approximately 1 x 10 6 cells, approximately 1 x 10 7 cells, approximately 1 x 10 8 cells, or approximately 1 x 10 9The composition may comprise a cell. In certain embodiments, the composition is a substantially purified preparation comprising differentiated cells produced by the methods disclosed herein. Also provided are compositions comprising a scaffold, such as a polymeric carrier and / or extracellular matrix, and an effective amount of cells produced by the methods disclosed herein. The matrix material is generally physiologically acceptable and suitable for use in in vivo applications. For example, physiologically acceptable materials include, but are not limited to, absorbable and / or non-absorbable solid matrix materials, such as small intestinal submucosa (SIS), cross-linked or non-cross-linked alginate, hydrocolloids, foams, collagen gels, collagen sponges, polyglycolic acid (PGA) meshes, fleeces, and bioadhesives.

[0213] Suitable polymeric carriers also include porous meshes or sponges formed from synthetic or natural polymers, and polymer solutions. For example, the matrix is ​​a polymer mesh or sponge, or a polymer hydrogel. Natural polymers that can be used include proteins, such as collagen, albumin, and fibrin; polysaccharides, such as alginate and hyaluronic acid. Synthetic polymers include both biodegradable and non-biodegradable polymers. For example, biodegradable polymers include hydroxy acids, such as polyactic acid (PLA), polyglycolic acid (PGA), and polylactic-glycolic acid (PGLA), polyorthoesters, polyanhydrides, polyphosphazenes, and combinations thereof. Non-biodegradable polymers include polyacrylates, polymethacrylates, ethylene vinyl acetate, and polyvinyl alcohol.

[0214] Polymers that can form malleable hydrogels, crosslinked ionically or covalently, can be used. Hydrogels are materials that form when organic polymers (natural or synthetic) are crosslinked via covalent, ionic, or hydrogen bonds to create a three-dimensional open lattice structure that traps water molecules to form a gel. Examples of materials that can be used to form hydrogels include ionically crosslinked polysaccharides, such as alginates, polyphosphatides, and polyacrylates, or block copolymers, such as PLURON1CS™ or TETRON1CS™, polyethylene oxide-polypropylene glycol block copolymers, crosslinked by temperature or H, respectively. Other materials include proteins, such as fibrin, polymers, such as polyvinylpyrrolidone, hyaluronic acid, and collagen.

[0215] C. Commercial, Therapeutic, and Research Purposes In some embodiments, a reagent system is provided that includes a set or combination of cells present at any time during manufacture, partitioning, or use. The culture set includes any combination of cell populations described herein, often in combination with undifferentiated pluripotent stem cells or other differentiated cell types that share the same genome. Each cell type may be packaged together or in separate containers, in the same facility or at different locations, at the same or different times, under the control of the same entity or different entities that share a business relationship.

[0216] The pharmaceutical composition may optionally be packaged in a suitable container along with written instructions for use for a desired purpose, eg, reconstituting cellular function to ameliorate disease or tissue injury.

[0217] [Example] IV. Working Examples The following examples are included to demonstrate preferred embodiments of the invention. Those skilled in the art should recognize that the techniques disclosed in the examples that follow represent techniques discovered by the inventors to work well in the practice of the invention, and therefore can be considered to constitute preferred modes for its practice. However, those skilled in the art should recognize in light of this disclosure that many changes can be made in the specific embodiments disclosed and still obtain the same or similar results without departing from the spirit and scope of the invention.

[0218] [Example 1] Codon optimization to prevent gene silencing To test whether methylation is responsible for transgene silencing, all CpG motifs were removed from the coding regions of genes known to be silenced (e.g., GFP) or suspected to be silenced (e.g., PuroR and NeoR). This concept was rigorously tested by comparing WT AcGFP1 with CpG-free AcGFP1, as this is easy to visually inspect. There was no change in the amino acid sequence after removal of the CpG motifs between SEQ ID NO:13 and SEQ ID NO:14. AcGFP1 DNA sequence (SEQ ID NO:13): atggtgagcaagggCGcCGagctgttcacCGgcatCGtgcccatcctgatCGagctgaatggCGatgtgaatggccacaagttcagCGtgagCGgCGagggCGagggCGatgccacctaCGgcaagctgaccctgaagttcatctgcaccacCGgcaagctgcctgtgccctggcccaccctggtgaccaccctgagctaCGgCGtgcagtgcttctcaCGctacccCGatcacatgaagcagcaCGacttcttcaagagCGccatgcctgagggctacatccaggagCGcaccatcttcttCGaggatgaCGgcaactacaagtCGCGCGcCGaggtgaagttCGagggCGataccctggtgaatCGcatCGagctgacCGgcacCGatttcaaggaggatggcaacatcctgggcaataagatggagtacaactacaaCGcccacaatgtgtacatcatgacCGacaaggccaagaatggcatcaaggtgaacttcaagatcCGccacaacatCGaggatggcagCGtgcagctggcCGaccactaccagcagaatacccccatCGgCGatggccctgtgctgctgccCGataaccactacctgtccacccagagCGccctgtccaaggaccccaaCGagaagCGCGatcacatgatctacttCGgcttCGtgacCGcCGcCGccatcacccaCGgcatggatgagctgtacaagTAA CpG-free AcGFP1 DNA sequence (SEQ ID NO: 14): ATGGTGAGCAAGGGCGCCGAGCTGTTCACCGGCATCGTGCCCATCCTGATCGAGCTGAATGGCGATGTGAATGGCCACAAGTTCAGCGTGAGCGGCGAGGGCGAGGGCGATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAGCTGCCTGTGCCCTGGCCCAC CCTGGTGACCACCCTGAGCTACGGCGTGCAGTGCTTCTCACGCTACCCCGATCACATGAAGCAGCACGACTTCTTCAAGAGCGCCATGCCTGAGGGCTACATCCAGGAGCGCACCATCTTCTTCGAGGATGACGGCAACTACAAGTCGCGCGCGAGGTGAAGTTCGAGGGCGATACCC TGGTGAATCGCATCGAGCTGACCGGCACCGATTTCAAGGAGGATGGCAACATCCTGGGCAATAAGATGGAGTACAACTACAACGCCCACAATGTGTACATCATGACCGACAAGGCCAAGAATGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGATGGCAGCGTGCAGCTG GCCGACCACTACCAGCAGAATACCCCCATCGGCGATGGCCCTGTGCTGCTGCCCGATAACCACTACCTGTCCACCCAGAGCGCCCTGTCCAAGGACCCCAACGAGAAGCGCGATCACATGATCTACTTCGGCTTCGTGACCGCCGCCGCCATCACCCACGGCATGGATGAGCTGTACAAG

[0219] The DNA sequence of the gene of interest was meticulously modified to remove all CG motifs that are 1) not rare, 2) do not produce stretches of mononucleotide stretches, and 3) replace codons that maintain a similar % GC content as the WT version of the gene. For example, in the case of CpG-free AcGFP1, the WT version of AcGFP1 (SEQ ID NO: 13) has a GC content of 59%, while the new CpG-free AcGFP1 (SEQ ID NO: 14) has a GC content of 52%.

[0220] The EEF1A1 promoter was used to express AcGFP1 in iPSCs at the PPP1R12C locus. Testing CpG-free AcGFP1 compared to WT AcGFP1 revealed that removing CpGs in the protein coding sequence overcomes silencing of gene expression (Figure 3).

[0221] However, after 5 months, we detected a small percentage (3%) of cells that did not express GFP, despite the removal of CpG from AcGFP1. To investigate this population, clones without GFP expression were isolated by single-cell sorting. These cells were treated with sodium butyrate (NaBut), a histone deacetylase (HDAC) inhibitor that can remove chromatin structure and induce demethylation. We observed that NaBut treatment reactivated GFP expression in a dose-dependent manner (Figure 5).

[0222] When CpG-free AcGFP1 iPSCs were differentiated into hepatocytes or neurons, a high percentage of GFP-positive differentiated cells was observed (Figure 6).

[0223] To verify the results, codon optimization of PuroRv1 (synthesized based on the amino acid sequence from Invivogen) in pUC57-KanR(m) was performed. The CpG-free sequence of PuroR (SEQ ID NO: 15) is shown below. CpG-free PuroR of plasmid 1346: ATGACTGAATACAAACCAACTGTTAGACTGGCAACTAGAGATGATGTTCCAAGAGCAGTTAGAACCCTGGCTGCTGCATTTGCTGACTACCCTGCAACCAGACACACTGTGGACCCAGACAGACACATTGAAAGAGTGACTGAACTGCAGGAGCTGTTCCTGACCAGAGTGGGCCTGGACATTGGCAAAGTGTGGGTGGCAGATGATGGTGCTGCTGTGGCAGTGTGGACCACCCCTGAATCTGTTGAAGCTGGTGCAGTGTTTGCTGAGATTGGCCCAAGAATGGCAGAACTGTCTGGCAGCAGACTGGCAGCACAACAGCAGATGGAAGGTCTGCTGGCACCACACAGACCAAAAGAACCTGCTTGGTTCCTGGCAACTGTGGGTGTGAGCCCTGACCACCAGGGTAAGGGCCTGGGCTCTGCAGTGGTGCTGCCTGGTGTGGAAGCAGCTGAAAGAGCAGGTGTGCCTGCTTTCCTGGAGACCTCAGCTCCAAGAAACCTGCCTTTCTATGAAAGACTGGGCTTCACTGTGACTGCTGATGTGGAAGTGCCAGAAGGCCCAAGAACTTGGTGCATGACTAGAAAACCAGGTGCTTGATAATGA(SEQ ID NO:15) CpG-free PuroRv2 in plasmids 1347 and 1363: (SEQ ID NO:16)

[0224] The CpG-free PuroR cassette was introduced into iPSCs by electroporation, and it was observed that cells carrying CpG-free PuroRv1 and PuroRv2 were able to confer drug resistance.

[0225] [Table 2]

[0226] [Table 3]

[0227] [Table 4]

[0228] [Table 5]

[0229] [Table 6]

[0230] These results suggest that CpG plays an important role in transgene silencing in iPSC lines. Furthermore, these results suggest that global methylation or other epigenetic dysregulation plays an important role in iPSCs with impaired differentiation. Thus, the method of the present invention of optimization to remove some or all of the CpG motifs can be used to prevent transgene silencing.

[0231] [Example 2] Differentiation of CpG-optimized iPSCs iPSCs transfected with CpG-free AcGFP1 and mRFP1 continued to constitutively express the fluorescent dye over many passages in culture. The next step was to confirm the retention of the fluorescent dye during differentiation of engineered iPSCs into progenitor and definitive lineages. Engineered iPSCs transfected with CpG-free plasmids were shown to successfully generate pure populations of endothelial cells, hematopoietic cells, macrophages, and microglia.

[0232] Generation of iPSC-derived endothelial cells from 9650 GFP iPSCs: Undifferentiated 9650-GFP were iPSCs maintained on MATRIGEL™ or vitronectin in the presence of E8 and adapted to hypoxia for at least 5-10 passages. To initiate endothelial cell differentiation, subconfluent iPSCs were harvested and seeded at a density of 250,000 cells / well in PureCoat amine culture dishes in the presence of Serum Free Defined (SFD) medium (Table 5) supplemented with 5uM blebbistatin or 1uM H1152 under hypoxic conditions. 24 hours after seeding, the cells were placed in SFD medium supplemented with 50ng / ml BMP4, VEGF and FGF-b, known as SFDEB#1 medium (Table 7). The cells were fed every 48 hours for 4-6 days to generate blood endothelial progenitor cells. These progenitor cells were either cryopreserved or cultured on tissue culture-treated plastic surfaces at 10 k / cm under normoxic conditions. 2 The cells can then be replated at a density of 100x to initiate endothelial cell differentiation in the presence of SFD-based endothelial cell medium containing H1152 (Table 7).

[0233] In an exemplary method, cryopreserved day 6 blood endothelial cells or live cultures were plated at 10k / cm on tissue culture treated plastic surfaces under normoxic conditions in the presence of SFD-based endothelial cell medium containing 1 uM H1152. 2 Cells were fed with fresh endothelial cell medium 24 hours after seeding and every 48 hours until confluence was reached. Cells took 5–6 days in culture to reach confluence. Cells were harvested using TrypLE Select, stained for surface endothelial markers CD31, CD105 and CD144, and plated at 10k / cm with endothelial cell medium. 2 The cells were then replated onto tissue culture treated plastic and placed under normoxic incubator conditions to expand and grow a pure population of endothelial cells.

[0234] [Table 7]

[0235] Generation of hematopoietic progenitor cells (HPCs) from GFP-engineered 9650 and RFP-engineered 8717 iPSCs: GFP-engineered 9650 and RFP-engineered 8717 iPSCs maintained on matrigel or vitronectin in the presence of E8 were adapted to hypoxia for at least 5-10 passages. Cells were split from subconfluent iPSCs and seeded at a density of 250-500 thousand cells / ml in spinner flasks in the presence of Serum Free Defined (SFD) medium supplemented with 5uM blebbistatin or 1uM H1152. 24 hours after seeding, the SFD medium supplemented with 50ng / ml BMP4, VEGF and FGF2 was replaced. On day 5 of the differentiation process, cells were placed in medium containing 50ng / ml Flt-3 ligand, SCF, TP0, IL3 and IL6 and 10U / ml heparin. Cells were fed every 48 hours throughout the differentiation process. All processes were performed under hypoxic conditions. The purity of HPCs was determined by quantification of CD4 and CD34 expression. An overview of the process is shown in Figure 14. HPCs were further purified by magnetic sorting using CD34 antibody.

[0236] HPC purity was assessed starting at day 12 and continued until CD34 expression was greater than 20%, as outlined in Figure 14A. Differentiated HPC cultures retained expression of GFP. Figure 14B. CD34 in line 9650. + MACS purification was performed on day 15. The RFP engineered 8717 line was found to be less efficient at generating HPCs. Nevertheless, the cultures retained RFP expression throughout the differentiation process. On day 17, half of the cultures were digested and plated for microglial differentiation, while the other half was maintained as aggregates for macrophage differentiation. The efficiency of the process for both lines can be seen in Figure 15.

[0237] [Table 8]

[0238] Generation of microglia: Purified HPCs were placed in microglia differentiation medium (MDM) under normoxic conditions. Cultures were fed every 48 hours with 2X MDM, and the differentiation process was completed after 23 days. A schematic of this process is shown in Figure 16. The cell morphology and fluorescence during the microglia differentiation process can be observed in Figures 17A and 17B. The efficiency of the HPC to microglia process can be seen in Figure 18.

[0239] The purity of the end-stage microglial cultures was assessed by flow cytometry for cell surface expression of CD45, CD33, TREM2, and CD11b, and intracellular expression of PU.1, IBA, P2RY12, TREM2, CX3CR1, and TMEM119 (Figures 19A and 19B).

[0240] [Table 9]

[0241] [Table 10]

[0242] Generation of macrophages: Macrophage differentiation was initiated using the 8717-RFP line on day 17 of HPC differentiation. An overview of the HPC to macrophage process is shown in Figure 20. A compilation of media for this part of the differentiation is listed in Table 11. On day 20, aggregates were digested and seeded in CMP medium. At this point, cultures were changed to normoxic environment. After one week, cultures were changed to macrophage medium and fed 2x macrophage medium every 4 days thereafter. CD68 purity was assessed on days 44 and 51 and is shown in Figure 21. On day 52, cells were harvested and cryopreserved. Cell morphology and fluorescence can be seen in Figure 22. The efficiency of the HPC to macrophage process is described in Figure 23. Fluorescence intensity from iPS cells to HPC, microglia and macrophages, measured by flow cytometry, is demonstrated in Figure 24.

[0243] [Table 11]

[0244] Generation of neural progenitor cells (NPCs) from iPSCs engineered with 8717-RFP and 9650-GFP: Neural progenitor cells (NPCs) are self-renewing progenitor cells that have the capacity to generate neurons and glia (Breunig et al., 2011). Many protocols have been established to generate NPCs from primary neural cells and iPSCs, with varying efficiency (Shi et al., 2012a; Shi et al., 2012b). Most of the recent protocols rely on the inhibition of SMAD signaling pathways. The method of the present invention describes a simple protocol to generate NPCs between different iPSC lines, exploiting the natural drift of iPSCs towards the ectoderm, without using dual SMAD inhibition pathways. Schematic description of the method to generate neural progenitor cells (NPCs) from iPSCs without using dual SMAD inhibition. The different steps involved, and the composition of the medium used, are described in Figure 25. Briefly, episomally reprogrammed iPSC lines, 8717-RFP and 9650-GFP, were maintained on Matrigel / Laminin / Vitronectin coated plates and E8 medium. iPSCs were maintained under hypoxic conditions prior to initiation of differentiation to generate NPCs. To initiate differentiation of neural precursors, iPSCs were harvested and plated at 15K / cm on Matrigel, Laminin or Vitronectin plates in the presence of ROCK inhibitors using E8 medium. 2 The cells were plated at 100 nm in E8 medium. The cells were placed in fresh E8 medium in the absence of ROCK inhibitor for the next 48 h. The next step involved a preconditioning step which involved placing the iPSC cultures in DMEMF12 medium supplemented with 3 μM CHIR for 72 h under normoxic conditions with daily medium changes. The cells were harvested at the end of the preconditioning step and cultured at 30 K / cm 2The cells were either replated in 2D format on Matrigel, laminin, or vitronectin plates at 100°C or 3D aggregates were generated using Ultra low Attachment (ULA) plates or spinner flasks at a density of 300,000 cells / ml in the presence of ROCK inhibitor. Cultures were fed with E6 medium supplemented with N2 every other day for the following 8 days under normoxic conditions. Retention of GFP and RFP fluorescence throughout the differentiation process is captured in Figure 26. Cultures were harvested on day 14 of differentiation and individualized using TrypLE. Cells were stained for the presence of SSEA4, CD56, CD15 by cell surface staining. Quantification of NPC purity is shown in Figure 27. CD56 was used as a marker for NPCs obtained by this method. Cells were cryopreserved using CS10 and retained their purity and proliferation potential after thawing.

[0245] Generation of GABAergic neurons from neural precursor cells: The potency of NPCs was tested by thawing NPCs and following the differentiation pathway outlined in Figure 28. Briefly, NPCs were subjected to downstream differentiation protocols to generate GABAergic neurons. NPCs were thawed and plated in DMEM / F12 supplemented with N2 and NEAA at 0.3e6 cells / mL in the presence of 10 μM blebbistatin for 24 hours to form aggregates. Cultures were fed with DMEM / F12 supplemented with N2 and NEAA and complete medium containing sonic hedgehog signaling molecule (SHH) and purmorphamine at 100 ng / mL and 1.5 μM, respectively, with daily changes for 10 days. NPCs were plated at 200,000 cells / cm on PLO-laminin coated plates using DMEM / F12, N2, NEAA, and 10 μM blebbistatin. 2For the next 48 hours of culture, the cultures were fed with DMEM / F12 supplemented with N2, NEAA, and 5 μM DAPT before being seeded with 5 μM DAPT for 24 hours. The cultures were fed with DMEM / F12, N2, NEAA, and 5 μM DAPT every other day thereafter and harvested 5 days after seeding. Emergence of GABA neurons. Retention of fluorescence in the culture is shown in FIG. 29. Quantification of GFP and RFP intensity from the iPSC stage to GABA neuron differentiation at day 18 is captured in FIG. 30. Finally, the purity of the final stage GABA neurons was shown in FIG. 31 by quantifying the purity of nestin and β-tubulin3. These cell differentiations demonstrated that CpG-optimized iPSCs can differentiate into multiple cell types, including but not limited to those mentioned above.

[0246] [Example 3] Promoter for stable expression Stable transgene expression in iPSC lines has been difficult to achieve over time and after differentiation. Many promoters show silencing or variable expression, and previous studies have shown this with promoters such as PGK and EEF1A1. The following studies were performed to identify promoters or taggable gene loci that can be used to provide stable expression in both iPSCs and differentiated cell types. It is often during the differentiation process that DNA methylation changes significantly, affecting expression, so the best promoters need to be active in both dividing cells and quiescent cells with little cell division (e.g., mature, fully differentiated cardiomyocytes).

[0247] Promoters cloned: The following promoters were identified as likely candidates for constitutive expression in all cell types. Some were cloned from existing plasmids (CAG, PGK, UBC-version 1, EEF1A1, ACTB). Other regions were generated de novo (by PCR from genomic DNA or synthetically) with the goal of identifying promoters that would result in stable expression in both iPSCs and differentiated cells. The new promoters include RPS19, UBA52, HSP90AB1, the expanded region of UBC (version 2), UBB, RPSA, NACA, and COX8A. The sequences were cloned into the pGL3 plasmid vector (replacing the SV40 promoter between the MluI and NcoI restriction sites) to allow for comparison of promoter strengths in driving luciferase reporter genes.

[0248] [Table 12] JPEG2024520413000014.jpg250170JPEG2024520413000015.jpg251170JPEG2024520413000016.jpg251170 JPEG2024520413000017.jpg251170JPEG2024520413000018.jpg251170JPEG2024520413000019.jpg202170

[0249] Luciferase expression during transient transfection: Transient transfection of promoter-pGL3 plasmids into iPSCs was performed to determine the strength of expression. Using the 96wp format, 50uL of E8 medium + 10uM blebbistatin was added to each well. Each plasmid was assayed in triplicate with the addition of 16.5uL of the following reagent preparation: One well of 6wp of iPSC line 01279.107 was harvested using Accutase and resuspended in 3.5mL of E8 medium + 10uM blebbistatin and 50uL was added to each well. One day later, cells were assayed using the Dual-Luciferase Reporter Assay System (Promega).

[0250] [Table 13]

[0251] Normalized luciferase (firefly / renilla ratio, normalized to EEF1A1=100%) is shown below (HSP90AB1del400 promoter and HSP90AB1 promoter expressed approximately 66% and 75% of EEF1A1). Since expression values ​​at or above the level of the PGK promoter were desired, RPS19, UBA52, HSP90AB1, and UBC were selected for further study.

[0252] ZsGreen construct integrated into the AAVS1 safe harbor locus: To investigate long-term expression driven by the candidate promoter in a chromosomal context, the candidate promoter was cloned into a plasmid controlling the ZsGreen fluorescent protein and targeted to the AAVS1 safe harbor locus on chromosome 19 of human iPSCs (plasmid design is shown below using the CAG promoter as an example). The plasmid was integrated into iPSC line 01279.107 by CRISPR-mediated gene editing, puromycin selection was applied, and resistant colonies were picked and genotyped by PCR. Correctly targeted heterozygous clones were expanded.

[0253] [Table 14]

[0254] Genomic loci suitable for tagging leading to constitutive expression: In addition to promoter-driven expression from the safe harbor, specific genes expressed in most cell types can be tagged with reporter genes to obtain constitutive expression. The following genes were selected for evaluation and tagged with ZsGreen and F2A cleavage sequences by TALEN-mediated gene editing. Correctly targeted heterozygous clones were expanded.

[0255] [Table 15]

[0256] ZsGreen expression in iPSCs: Engineered iPSC lines expressing ZsGreen fluorescent protein were maintained in culture for up to 7 months (E8 medium / vitronectin-coated plates) and periodically checked for green expression using flow cytometry on an Accuri C6 instrument (BD). Most clones maintained consistent flow profiles over time, with the exception of one RPS19 promoter clone (5363), which showed a decrease in fluorescence in many cells as of August.

[0257] Differentiation: To determine the stability of expression following differentiation, the engineered lines were subjected to differentiation protocols directing them towards either neuronal or cardiac cell types.

[0258] [Table 16]

[0259] At day 21 of differentiation, all cells had a visible neuronal phenotype. Flow cytometry showed that the fluorescence of the CAG, UBC(v1), and HSP90AB1del400 promoters was diminished in many cells. The UBCv2, UBA52, and RPS19 promoters showed tight and stable expression, as did the tag genes HSP90AB1, CTNNB1, and MYL6.

[0260] [Table 17]

[0261] At day 21 of differentiation, the CAG, UBC(v1), RPS19, and HSP90AB1del400 promoter lines showed varying amounts of expression silencing, while the UBC(v2) and UBA52 promoters showed tight and stable expression, as did the tagged genes HSP90AB1, ACTB, CTNNB1, and MYL6.

[0262] The newly generated promoter regions UBCv2, UBA52, RPS19, and HSP90AB1del400 showed stable iPSC expression over 4 months of culture up to 7 months, with only RPS19 showing some silencing at this time point. The HSP90AB1, ACTB, CTNNB1, and MYL6 gene loci showed stable expression of tagged ZsGreen reporters. The UBCv2 and UBA52 reporters were shown to be stable under the two differentiation protocols, as was the expression driven from the HSP90AB1, CTNNB1, and MYL6 genes.

[0263] All of the methods disclosed and claimed herein can be accomplished and executed without undue experimentation in light of the present disclosure. Although the compositions and methods of the present invention have been described in terms of preferred embodiments, it will be apparent to those skilled in the art that modifications may be applied to the methods in the steps or in the sequence of steps of the methods described herein without departing from the concept, spirit and scope of the invention. More specifically, it will be apparent that certain agents that are chemically and physiologically related may be substituted for the agents described herein while the same or similar results would be obtained. All such similar substitutes and modifications apparent to those skilled in the art are deemed to be within the spirit, scope and concept of the invention as defined by the appended claims. References The following references, to the extent that they provide exemplary procedural or other details supplementary to those set forth herein, are specifically incorporated herein by reference. Alexander et al., Proc. Nat. Acad. Sci. USA,85:5092-5096,1988. Ausubel et al., Current Protocols in Molecular Biology, Greene Publ. Assoc. Inc. & John Wiley & Sons, Inc., MA, 1996.Blomer et al., 1997 Chen and Okayama, Mol. Cell Biol., 7(8):2745-2752, 1987. Chen et al., Nature Methods 8:424-429, 2011. Ercolani et al., J. Biol. Chem., 263:15335-15341, 1988. Evans, et al., In: Cancer Principles and Practice of Oncology, Devita et al. (Eds.), Lippincot-Raven, NY, 1054-1087, 1997.Fechheimer et al., Proc Natl. Acad. Sci. USA, 84:8463-8467, 1987. Fraley et al., Proc. Natl. Acad. Sci. USA, 76:3348-3352, 1979. Gaj et al., Trends in Biotechnology, 2013, 31(7), 397-405 Graham and Van Der Eb, Virology, 52:456-467, 1973. International Publication WO02 / 016536 International Publication WO03 / 016496 International Publication WO2003 / 0211603 International Publication WO2007 / 069666 International Publication WO2007 / 069666 International Publication WO2012 / 0196360 International Publication WO94 / 09699 International Publication WO95 / 06128 International Publication WO98 / 30679 International Publication WO98 / 53058 International Publication WO98 / 53059 International Publication WO98 / 53060 Kaeppler et al., Plant Cell Reports 9: 415-418, 1990. Kaneda et al., Science, 243:375-378, 1989. Karin et al. Cell, 36:371-379,1989. Kato et al, J. Biol. Chem., 266:3361-3364, 1991. Kyttala et al., Stem Cell Reports, 6(2):200-12, 2016. Langle-Rouault et al., J. Virol., 72(7):6181- 6185, 1998. Levitskaya et al., Proc. Natl. Acad. Sci. USA, 94(23):12616- 12621 , 1997. Ludwig et al., Nat. Biotechnol., 24:185-187, 2006b. Ludwig et al., Nat. Methods, 3:637-646, 2006a. Macejak and Sarnow, Nature, 353:90-94, 1991. Maniatis, et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press, Cold Spring Harbor, N.Y., 1988. Mann et al., Cell, 33:153-159, 1983. Nabel et al., Science, 244(4910):1342-1344, 1989. Naldini et al., Science, 272(5259):263-267, 1996. Ng, Nuc. Acid Res., 17:601-615, 1989. Nicolas and Rubenstein, In: Vectors: A survey of molecular cloning vectors and their uses, Rodriguez and Denhardt, eds., Stoneham: Butterworth, pp. 494-513, 1988. Nicolau and Sene, Biochim. Biophys. Acta, 721:185-190, 1982. Nicolau et al., Methods Enzymol., 149:157-176, 1987. Paskind et al., Virology, 67:242-248, 1975. Pelletier and Sonenberg, Nature, 334(6180):320-325, 1988. Potrykus et al., Mol. Gen. Genet., 199(2):169-177, 1985. Potter et al., Proc. Natl. Acad. Sci. USA, 81:7161-7165, 1984. Quitsche et al., J. Biol. Chem., 264:9539-9545, 1989. Richards et al., Cell, 37:263-272, 1984. Rippe, et al., Mol. Cell Biol., 10:689-695, 1990. Sambrook and Russell, Molecular Cloning: A Laboratory Manual, 3rd Ed. Cold Spring Harbor 1997. Takahashi et al., Cell, 131:861-872, 2007. Temin, In: Gene Transfer, Kucherlapati (Ed.), NY, Plenum Press, 149-188, 1986. Tur-Kaspa et al., Mol. Cell Biol., 6:716-718, 1986. U.S. Patent 4,683,202 U.S. Patent 5,302,523 U.S. Patent 5,322,783 U.S. Patent 5,384,253 U.S. Patent 5,464,765 U.S. Patent 5,538,877 U.S. Patent 5,538,880 U.S. Patent 5,550,318 U.S. Patent 5,556,954 U.S. Patent 5,563,055 U.S. Patent 5,563,055 U.S. Patent 5,580,859 U.S. Patent 5,589,466 U.S. Patent 5,591,616 U.S. Patent 5,610,042 U.S. Patent 5,656,610 U.S. Patent 5,702,932 U.S. Patent 5,736,524 U.S. Patent 5,780,448 U.S. Patent 5,789,215 U.S. Patent 5,925,565 U.S. Patent 5,928,906 U.S. Patent 5,935,819 U.S. Patent 5,945,100 U.S. Patent 5,981,274 U.S. Patent 5,994,136 U.S. Patent 5,994,136 U.S. Patent 5,994,624 U.S. Patent 6,013,516 U.S. Patent 6,103,470 U.S. Patent 6,140,081 U.S. Patent 6,416,998 U.S. Patent 6,453,242 U.S. Patent 6,534,261 U.S. Patent 7,442,548 U.S. Patent 7,598,364 U.S. Patent 7,989,425 U.S. Patent 8,058,065 U.S. Patent 8,071,369 U.S. Patent 8,129,187 U.S. Patent 8,268,620 U.S. Patent 8,278,620 U.S. Patent 8,546,140 U.S. Patent 8,546,140 U.S. Patent 8,741,648 U.S. Patent Publication 2002 / 0076747 U.S. Patent Publication 2002 / 0055144 U.S. Patent Publication 2005 / 0064474 U.S. Patent Publication 2006 / 0188987 U.S. Patent Publication 2007 / 0218528 U.S. Patent Publication 2009 / 0148425 U.S. Patent Publication 2009 / 0246875 U.S. Patent Publication 2010 / 0003757 U.S. Patent Publication 2010 / 0210014 U.S. Patent Publication 2011 / 0301073 U.S. Patent Publication 2011 / 0301073 US Patent Publication 2012 / 0276636 Wilson et al., Science, 244:1344-1346, 1989. Wong et al., Gene, 10:87-94, 1980. Yamanaka et al., Cell, 131(5):861-72, 2007. Zufferey et al., Nat. Biotechnol., 15(9):871-875, 1997.

Claims

**Claim 1** An isolated cell line engineered to express at least one transgene, wherein the at least one transgene is: (a) under the control of a promoter having at least 90% sequence identity to SEQ ID NOs: 1-12 or 17; (b) under the control of an endogenous gene selected from the group consisting of HSP90AB1, ACTB, CTNNB1, MYL6, UBA52, CAG, RPS, and UBC; and / or (c) encoded by a sequence modified to remove CpG motifs to effect stable expression. **Claim 2** The cell line according to claim 1, wherein the at least one transgene is encoded by a sequence modified to remove CpG motifs to effect stable expression. **Claim 3** The cell line according to claim 2, wherein the sequence modified to remove CpG motifs to effect stable expression has at least 90% sequence identity to or is identical to SEQ ID NO: 14 or SEQ ID NO:

16. **Claim 4** The cell line according to claim 1, wherein the at least one transgene is encoded by a sequence modified to remove CpG motifs to effect stable expression and is under the control of a promoter having at least 90% sequence identity to SEQ ID NOs: 1-12 or 17. **Claim 5** The cell line according to claim 1, wherein the at least one transgene is encoded by a sequence modified to remove CpG motifs to effect stable expression and is under the control of an endogenous gene selected from the group consisting of HSP90AB1, ACTB, CTNNB1, MYL6, UBA52, CAG, RPS, and UBC. **Claim 6** The cell line according to any one of claims 1-5, wherein the cell line is engineered to express at least a first transgene and a second transgene. **Claim 7** The cell line according to claim 6, wherein the first transgene is under the control of a promoter having at least 90% sequence identity to SEQ ID NOs: 1-12 or 17, and the second transgene is under the control of an endogenous gene selected from the group consisting of HSP90AB1, ACTB, CTNNB1, and MYL6. **Claim 8** The cell line according to any one of claims 1-5, wherein at least 50%, 70%, 90% or 100% of the CpG motifs have been removed. **Claim 9** The cell line according to any one of claims 1 to 5, wherein the codons of the CpG motif are replaced with codons that are not rare and / or do not result in mononucleotide stretches, or are replaced with the corresponding codons in Table 1.

10. The cell line according to any one of claims 1 to 5, wherein the cell line is an induced pluripotent stem cell (iPSC) line.

11. The cell line according to any one of claims 1 to 5, wherein the transgene is a reporter gene, a selectable marker, or a suicide gene.

12. The cell line according to any one of claims 1 to 5, wherein at least one transgene is a suicide gene.

13. The cell line according to any one of claims 1 to 5, wherein the cell line has stable expression of the transgene for more than 6 months.

14. The cell line according to any one of claims 1 to 5, wherein the expression cassette is inserted into a genomic safe harbor site selected from the PPP1R12C (AAVS1) locus or the ROSA locus.

15. The cell line according to any one of claims 1 to 5, wherein the promoter has at least 90%, 95% or 100% sequence identity to SEQ ID NO: 2, 3, 4, 6, or 17.

16. A method for preventing silencing of transgene expression in an engineered cell line, the method comprising optimizing the transgene sequence to remove CpG motifs.

17. The method according to claim 16, wherein optimizing comprises replacing at least 50%, 70%, 90% or 100% of the CpG motifs, and / or the optimized transgene sequence for removing CpG motifs comprises a GC content percentage substantially similar to the percentage of the GC content of the wild-type transgene sequence.

18. The method according to claim 16 or 17, further comprising treating the cell line with sodium butyrate, VPA, or TSA.

19. The method according to claim 16, wherein the cell line is an iPSC line differentiated into hematopoietic progenitor cells, neural progenitor cells, GABAergic neurons, macrophages, microglia, or endothelial cells.

20. An expression vector comprising a promoter having at least 90% sequence identity to SEQ ID NOs: 1-12 or 17.

21. A method for generating a cell line having stable transgene expression, comprising engineering a cell line to express the vector according to claim 20, wherein the vector encodes the transgene.

22. The method according to claim 21, comprising integrating the vector into the AAVS1 locus on chromosome 19.

23. The method according to claim 21 or 22, wherein the integration comprises gene editing, the gene editing comprising CRISPR-mediated gene editing, TALEN-mediated gene editing, or ZFN-mediated editing.

24. The method according to claim 21 or 22, further comprising differentiating the cell line into neurons or cardiac cells.

25. The method according to claim 21 or 22, wherein the cell line is cultured for at least 30 days or 6 months.

26. An isolated cell line having endogenous HSP90AB1, ACTB, CTNNB1, MYL6, UBA52, CAG, RPS, or UBC tagged with a transgene.