Novel sites for safe genomic integration and methods of use thereof
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- BLUEROCK THERAPEUTICS LP
- Filing Date
- 2023-04-28
- Publication Date
- 2026-05-11
AI Technical Summary
The prior art When inserting a transgene into the genome, it is difficult to maintain stable expression of the transgene in different cell types and cell states, and may affect the normal function of the host cell.
Identify and utilize intergenic regions (STAPLRs) that maintain transcriptional activity in different cell types and cell states, and insert exogenous nucleotide sequences (such as transgenes) into these regions to ensure stable expression of the transgene.
The durable expression of transgenes in different cell types and cell states is achieved, which avoids the problems of transgene silencing and impaired host cell function, and improves the stability and effectiveness of cell therapy in gene therapy.
Smart Images

Figure 00000048_0000 
Figure 00000048_0001 
Figure 00000048_0002
Abstract
Description
[Technical field]
[0001] The present invention relates to novel sites for safe genomic integration and methods of their use.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 336,248, filed April 28, 2022, the contents of which are incorporated by reference in their entirety.
[0003] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in XML format, which is incorporated herein by reference in its entirety. The electronic copy of the Sequence Listing, created on April 27, 2023, is designated 025450_WO017_SL.xml and has a size of 379,876 bytes. [Background technology]
[0004] background Many efforts to safely integrate transgenes into the genome have been performed at so-called "genomic safe harbor" sites. A safe harbor site in the genome is a site where a nucleic acid (e.g., an exogenous gene) can be introduced without interfering with the expression or regulation of adjacent genes and thus without interfering with the normal function of the cell. Three genomic sites, AAVS1, CCR5, and ROSA26, are traditionally considered to be safe harbor sites and are used in most targeted transgene integrations. AAVS1 is a region of rare genomic integration of the AAV genome and has been found to allow robust expression without interfering with cellular function. CCR5 was identified serendipitously because naturally occurring CCR5-delta-32 mutations result in an HIV-resistant phenotype, and the disposability of the gene makes it an ideal integration site. The ROSA26 locus was originally identified in mouse embryonic stem cells by a lentiviral gene trap approach.
[0005] Although these genomic safe harbor sites allow robust transgene expression under a given cellular context, they may not maintain faithful transgene expression in other cell lineages or after changes in cellular conditions. This is because interactions between the transgene and the genomic context of the host cell may affect transgene expression, resulting in attenuation or complete silencing (e.g., by DNA methylation) of transgene expression. More importantly, these genomic integration sites may also affect the expression of endogenous genes near the insertion site, affecting normal host cell function. Summary of the Invention
[0006] Disclosure Summary The present disclosure is based, at least in part, on the identification of intergenic sites in the genome that remain transcriptionally active in different cell types and different cellular states (including maturation stages), where an integrated exogenous nucleotide sequence of interest (e.g., a transgene encoding a protein or RNA) remains expressed and functional as cells proliferate and cellular states change.
[0007] Thus, in one aspect, the disclosure provides a genetically modified cell, e.g., a mammalian (e.g., human) cell, comprising an exogenous nucleotide sequence integrated into a sustained transcriptionally active payload region (STAPLR) in the genome of the cell, the STAPLR being located in the intergenic region between the RPL34 gene and the OSTC gene, the intergenic region between the ACTB gene and the FSCN1 gene, the intergenic region between the AKIRIN1 gene and the NDUFS5 gene, the intergenic region between the PRDX1 gene and the AKR1A1 gene, the intergenic region between the PTGES3 gene and the NACA gene, the intergenic region between the MLF2 gene and the PTMS gene, the intergenic region between the RAB13 gene and the RPS27 gene, the intergenic region between the JTB gene and the RAB13 gene, the intergenic region between the ALT1 gene and the ATCC1 gene, the intergenic region between the ATCC ... The intergenic region is selected from the group consisting of the intergenic region between the KR1A1 gene and the NASP gene, the intergenic region between the NDUFS5 gene and the MACF1 gene, the intergenic region between the SRSF9 gene and the DYNLL1 gene, the intergenic region between the MYL6B gene and the MYL6 gene, the intergenic region between the GPX1 gene and the RHOA gene, the intergenic region between the HNRNPA2B1 gene and the CBX3 gene, the intergenic region between the ROMO gene and the RBM39 gene, the intergenic region between the PA2G4 gene and the RPL41 gene, and the intergenic region between the NDUFB10 gene and the RPS2 gene.
[0008] In some embodiments, the intergenic region between the RPL34 gene and the OSTC gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO:1, or a nucleotide sequence that is sufficiently similar to SEQ ID NO:1 to the extent necessary for the function of the sequence to remain intact (e.g., have no adverse effects on the cell).
[0009] In some embodiments, the intergenic region between the ACTB gene and the FSCN1 gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO:2, or a nucleotide sequence that is sufficiently similar to SEQ ID NO:2 to the extent necessary for the function of the sequence to remain intact (e.g., without adversely affecting the cell).
[0010] In some embodiments, the intergenic region between the AKIRIN1 gene and the NDUFS5 gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO:3, or a nucleotide sequence that is sufficiently similar to SEQ ID NO:3 to the extent necessary for the function of the sequence to remain intact (e.g., have no adverse effects on the cell).
[0011] In some embodiments, the intergenic region between the PRDX1 gene and the AKR1A1 gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO:4, or a nucleotide sequence that is sufficiently similar to SEQ ID NO:4 to the extent necessary for the function of the sequence to remain intact (e.g., without adversely affecting the cell).
[0012] In some embodiments, the intergenic region between the PTGES3 gene and the NACA gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO:5, or a nucleotide sequence that is sufficiently similar to SEQ ID NO:5 for the function of the sequence to remain intact (e.g., not to have adverse effects on the cell).
[0013] In some embodiments, the intergenic region between the MLF2 gene and the PTMS gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO:6, or a nucleotide sequence that is sufficiently similar to SEQ ID NO:6 to the extent necessary for the function of the sequence to remain intact (e.g., without adversely affecting the cell).
[0014] In some embodiments, the intergenic region between the RAB13 gene and the RPS27 gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO:7, or a nucleotide sequence that is sufficiently similar to SEQ ID NO:7 for the function of the sequence to remain intact (e.g., have no adverse effects on the cell).
[0015] In some embodiments, the intergenic region between the JTB gene and the RAB13 gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO:8, or a nucleotide sequence that is sufficiently similar to SEQ ID NO:8 to the extent necessary for the function of the sequence to remain intact (e.g., without adversely affecting the cell).
[0016] In some embodiments, the intergenic region between the AKR1A1 gene and the NASP gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO:9, or a nucleotide sequence that is sufficiently similar to SEQ ID NO:9 to the extent necessary for the function of the sequence to remain intact (e.g., without adversely affecting the cell).
[0017] In some embodiments, the intergenic region between the NDUFS5 gene and the MACF1 gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO:10, or a nucleotide sequence that is sufficiently similar to SEQ ID NO:10 to the extent necessary for the function of the sequence to remain intact (e.g., have no adverse effects on the cell).
[0018] In some embodiments, the intergenic region between the SRSF9 gene and the DYNLL1 gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO:11, or a nucleotide sequence that is sufficiently similar to SEQ ID NO:11 to the extent necessary for the function of the sequence to remain intact (e.g., not to have adverse effects on the cell).
[0019] In some embodiments, the intergenic region between the MYL6B gene and the MYL6 gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO:12, or a nucleotide sequence that is sufficiently similar to SEQ ID NO:12 to the extent necessary for the function of the sequence to remain intact (e.g., have no adverse effects on the cell).
[0020] In some embodiments, the intergenic region between the GPX1 gene and the RHOA gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO:13, or a nucleotide sequence that is sufficiently similar to SEQ ID NO:13 to the extent necessary for the function of the sequence to remain intact (e.g., without adversely affecting the cell).
[0021] In some embodiments, the intergenic region between the HNRNPA2B1 gene and the CBX3 gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO:14, or a nucleotide sequence that is sufficiently similar to SEQ ID NO:14 to the extent necessary for the function of the sequence to remain intact (e.g., without adversely affecting the cell).
[0022] In some embodiments, the intergenic region between the ROMO gene and the RBM39 gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO:15, or a nucleotide sequence that is sufficiently similar to SEQ ID NO:15 to the extent necessary for the function of the sequence to remain intact (e.g., not to have adverse effects on the cell).
[0023] In some embodiments, the intergenic region between the PA2G4 gene and the RPL41 gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO:16, or a nucleotide sequence that is sufficiently similar to SEQ ID NO:16 to the extent necessary for the function of the sequence to remain intact (e.g., without adversely affecting the cell).
[0024] In some embodiments, the intergenic region between the NDUFB10 gene and the RPS2 gene comprises a nucleotide sequence that is at least 95% (e.g., at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to SEQ ID NO: 16, or a nucleotide sequence that is sufficiently similar to SEQ ID NO: 97 to the extent necessary for the function of the sequence to remain intact (e.g., without adversely affecting the cell).
[0025] Also provided herein are methods for producing these genetically modified mammalian cells (manufacturing methods), and DNA constructs for introducing nucleotide sequences of interest into the novel genomic integration sites herein.Thus, in one aspect, the present disclosure provides a method for modifying mammalian cells, comprising integrating nucleotide sequences of interest (i.e., exogenous nucleotide sequences) into the STAPLR described herein.In some embodiments, the integration step is carried out using CRISPR / Cas system, Cre / Lox system, FLP-FRT system, TALEN system, ZFN system, homing endonuclease, random integration, homologous recombination, transposase, or nuclease-independent viral vector, which is optionally selected from retroviral vector, adeno-associated viral (AAV) vector, and lentiviral vector. In further detailed embodiments, the CRISPR / Cas system comprises a guide RNA, wherein (i) STAPLR is an intergenic region between the RPL34 gene and the OSTC gene, and the gRNA is selected from SEQ ID NOs: 25-32; (ii) STAPLR is an intergenic region between the ACTB gene and the FSCN1 gene, and the gRNA is selected from SEQ ID NOs: 33-54; (iii) STAPLR is an intergenic region between the AKIRIN1 gene and the NDUFS5 gene, and the gRNA is selected from SEQ ID NOs: 55-70; or (iv) STAPLR is an intergenic region between the PRDX1 gene and the AKR1A1 gene, and the gRNA is selected from SEQ ID NOs: 71-92.
[0026] In some embodiments, the CRISPR / Cas system comprises a gRNA-dependent nuclease or variants thereof of type I, II, III, IV, or V. In more particular embodiments, the CRISPR / Cas system comprises a gRNA-dependent nuclease or variants thereof of type I, II, III, IV, or V. In more particular embodiments, the CRISPR / Cas system comprises a gRNA-dependent nuclease or variants thereof of type I, II, III, IV, or V. In more particular embodiments, the CRISPR / Cas system comprises a gRNA-dependent nuclease or variants thereof of type I, II, III, IV, or V. m5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, CasX, CasY, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, CasPhi, MAD7 and Csf4.
[0027] In another aspect, the present disclosure provides a DNA molecule comprising a nucleotide sequence of interest flanked by a 5' homology region (HR) and a 3' HR, wherein the 5' HR and 3' HR are at least 85% (e.g., at least 90, 95, 96, 97, 98 or 99%) homologous or 100% identical to a first genomic region (GR) and a second GR, respectively, in a STAPLR described herein. In some embodiments, each of the 5'HR and 3'HR is independently at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 550, at least about 600, at least about 650, at least about 700, at least about 750, at least about 800, at least about 850, at least about 900, at least about 950, at least about 1000, at least about 1100, at least about 1200, at least about 1300, at least about 1400, at least about 1500, at least about 1600, at least about 1700, at least about 1800, at least about 1900, or at least about 2000 base pairs in length. In some embodiments, each of the HRs is 200 to 2000 (e.g., 300 to 2500, 400 to 2000, or 500 to 1500) base pairs in length. In more specific embodiments, the 5'HR and 3'HR are at least 90% (e.g., at least 95%) homologous to SEQ ID NOs: 17 and 18, 19 and 20, 21 and 22, 23 and 24, 93 and 94, or 95 and 96, respectively.
[0028] In some embodiments, the exogenous nucleotide sequence or nucleotide sequence of interest comprises a transgene. In more detailed embodiments, the transgene comprises a coding sequence (e.g., of a protein or RNA) and one or more regulatory elements. In some embodiments, the one or more regulatory elements include a constitutive or inducible promoter that directs transcription of the coding sequence. In some embodiments, the transgene encodes a therapeutic protein (e.g., a protein whose deficiency or loss causes a disease, such as a genetic disease, a cytokine, or a recombinant antigen receptor), a cell marker, or a protein that modulates the differentiation state or activity of a cell (e.g., a reprogramming factor). In some embodiments, the transgene encodes SOX10, IL-10, IL-12, CD19t, or ThPOK.
[0029] In some embodiments of the present disclosure, the mammalian cell is a human cell. In some embodiments, the mammalian cell (e.g., a human cell) is a pluripotent stem cell (PSC; e.g., an induced PSC (iPSC) or an embryonic stem cell (ESC)]. In some embodiments, the mammalian cell (e.g., human cell) is a) a cell of the immune system (e.g., T cell, natural killer cell, dendritic cell, macrophage / monocyte or hematopoietic progenitor or precursor cell thereof); b) a cell of the cardiovascular system (e.g., ventricular cardiomyocyte, nodal cell or cardiac progenitor or precursor cell thereof); c) a cell of the metabolic system (e.g., hepatocyte or pancreatic beta cell or progenitor or precursor cell thereof); d) a cell of the central nervous system (e.g., sensory neuron, motor neuron, interneuron, microglial cell, oligodendrocyte or progenitor or precursor cell thereof); e) a muscle cell (e.g., a skeletal muscle cell or a smooth muscle cell or progenitor or precursor cell thereof); f) an adipocyte or progenitor or precursor cell thereof; or g) a cell of the ocular system (e.g., a retinal pigment epithelial cell, a photoreceptor cell or progenitor or precursor cell thereof). Additional cell types of the present disclosure include those described below.
[0030] Also provided herein are pharmaceutical compositions comprising the genetically engineered cells herein and a pharma- ceutically acceptable carrier, as well as gene editing systems comprising the DNA molecules disclosed herein and a gene editing system necessary to incorporate a nucleotide sequence of interest (e.g., a nuclease and a gRNA) on the DNA molecule into STAPLR.
[0031] In another aspect, the disclosure provides a method for identifying persistent transcriptionally active payload regions (STAPLRs) in the genome of a mammalian cell, the method comprising: (i) performing a single cell RNA sequencing analysis on a set of two or more mammalian cell types, wherein the sequencing analysis assigns a unique transcriptome to each cell type; (ii) assigning a Prevalence Score to constituent genes in the transcriptomes, wherein the Prevalence Score represents a proportion of mammalian cell types in the set of mammalian cell types that contain at least one transcript of the gene; (iii) identifying neighboring genes of the constituent genes in the genome of the mammalian cell, wherein the neighboring genes do not overlap with the constituent genes; (iv) calculating a Neighbor Score for the pair of non-overlapping genes or for a region containing three or more genes identified in step (iii). The method includes determining a Prevalence Score (wherein the adjacency score is the product of the prevalence scores of the individual genes in the pair or region), (v) ranking the adjacency scores, and (vi) selecting a pair of non-overlapping genes or a region containing three or more non-overlapping genes based on the high ranking, thereby identifying the intergenic region between the genes of the selected pair or region as a STAPLR. In some embodiments, the method further includes (vii) selecting a targetable intergenic subregion in the STAPLR, and (viii) inserting a transgene into the selected subregion, wherein transcription of the transgene or gene circuit is persistent. In some embodiments, the targetable subregion does not contain a known promoter or enhancer region, and / or contains a minimal number of conserved regions, repetitive regions, epigenetic marks, and / or enzymatic hypersensitivity regions, and / or the nuclease is a CRISPR nuclease.In some embodiments, the intergenic region is at least 30 (e.g., at least 40, at least 50, at least 75, or at least 100) base pairs in length and / or does not contain or contains a minimal number of promoter regions, CpG islands, H3K4Me1 epigenetic marks, H3K4Me3 epigenetic marks, H3K27Ac epigenetic marks, DNase I hypersensitive regions, conserved regions, or repetitive regions.
[0032] Other features, objects and advantages of the present invention will become apparent in the following detailed description. However, it should be understood that such detailed description illustrates embodiments and aspects of the present invention and is given for illustrative purposes only and is not limiting. Various changes and modifications within the scope of the present invention will become apparent to those skilled in the art from such detailed description. [Brief description of the drawings]
[0033] [Figure 1] Figure 1 shows dot plots showing the indel editing rates obtained after Sanger sequencing testing using Synthego's ICE analysis tool. For each different STAPLR site, three different gRNAs were tested, and the gRNA showing the highest indel editing rate is circled. The solid horizontal line shows the average indel editing rate of the three different gRNAs per STAPLR site. [Diagram 2] Figure 2 shows the integration of sequences encoding the 2A peptide and the Tet-On 3G form of rtTA at the GAPDH locus. The left and right homology arms were designed to allow in-frame integration of the transgene immediately 5' to the GAPDH stop codon. This allows expression of rtTA under the control of the endogenous GAPDH promoter. iPSCs edited with the targeting construct constitutively express the rtTA protein. [Diagram 3]Figure 3 shows the integration of each of the four STAPLR targeting constructs, containing the pTRE3G-eGFP-Sv40 transgene flanked by left and right homology arms at each STAPLR site, in iPSCs that constitutively express the rtTA protein. Addition of doxycycline allows binding of the rtTA protein and activation of GFP expression from the TRE3G promoter. [Figure 4] 4 is a panel of fluorescent microscopy images showing expression of GFP in a pooled population of cells to which doxycycline was added for 24 hours. Doxycycline was added to the medium 24 hours after nucleofection of iPSCs with the STAPLR targeting construct and the corresponding RNP. No GFP was observed in cells to which doxycycline was not added. Control iPSCs constitutively expressing rtTA that were treated with doxycycline but not nucleofected with the STAPLR targeting construct and RNP also did not express GFP. [Diagram 5] 5 is a panel of fluorescent microscopy images showing expression of GFP in a pooled population of cells to which doxycycline was added for 6 days. Doxycycline was added to the medium 24 hours after nucleofection of iPSCs with the STAPLR targeting construct and the corresponding RNP. No GFP was observed in cells to which doxycycline was not added. Control iPSCs constitutively expressing rtTA that were treated with doxycycline but not nucleofected with the STAPLR targeting construct and RNP also did not express GFP. [Figure 6-1] 6 is a panel of flow cytometry histograms showing the time-dependent induction of GFP expression in four different clonally derived STAPLR iPSC lines under different concentrations of doxycycline. Cells were harvested and analyzed 0, 3, 8, 24, 48, and 68 hours after doxycycline administration. [Figure 6-2]6 is a panel of flow cytometry histograms showing the time-dependent induction of GFP expression in four different clonally derived STAPLR iPSC lines under different concentrations of doxycycline. Cells were harvested and analyzed 0, 3, 8, 24, 48, and 68 hours after doxycycline administration. [Figure 7] 7 is a panel of flow cytometry histograms showing induction of GFP expression over time following treatment with 2 μg / ml doxycycline in four different clonally derived STAPLR iPSC lines. The left panel shows the PRDX1-AKR1A1, ACTB-FSCN1 and RPL34-OSTC STAPLR lines, as well as the wild-type unedited iPSC control line in the absence or presence of 72 hours of doxycycline treatment. The right panel shows the AKIRIN1-NDUFS5 STAPLR line in the absence or presence of 6 days of doxycycline treatment. [Figure 8] 8 is a panel of flow cytometry histograms showing induction of GFP expression following treatment with 2 μg / ml doxycycline in four different clonally derived STAPLR iPSC lines differentiated into myeloid progenitors. Doxycycline was added to the culture medium on day 12 of differentiation. The left panel shows the PRDX1-AKR1A1, ACTB-FSCN1 and RPL34-OSTC STAPLR lines and the wild type unedited iPSC control line after 15 days of myeloid differentiation in either the absence or presence of 72 hours of doxycycline treatment. The right panel shows the AKIRIN1-NDUFS5 STAPLR line after 18 days of myeloid differentiation in either the absence or presence of 6 days of doxycycline treatment. [Figure 9] 9 is a panel of flow cytometry dot plots showing expression of myeloid progenitor markers CD45, CD14 and CX3CR1 in the non-adherent myeloid cell population of STAPLR-targeted iPSC lines that had been differentiated for the past 30 days. The CD14 and CX3CR1 panel of cells was gated on CD45 positive cells. [Figure 10]10 is a panel of flow cytometry histograms showing induction of GFP expression in non-adherent bone marrow progenitor cells in four differentiated clonally derived STAPLR iPSC lines and a wild-type unedited iPSC control line following treatment with 2 μg / ml doxycycline, which was added to the culture medium for 6 days after day 30 of differentiation. [Figure 11] Figure 11 shows the integration of a targeting construct containing the pTRE3G-CD19t-IL12 transgene flanked by left and right homology arms that allow integration at the PRDX1-AKR1A1 STAPLR site. This construct was transfected into iPSCs that constitutively express the rtTA protein from the GAPDH endogenous promoter. [Figure 12] Figure 12 is a panel of photographs showing live cell imaging of CD19t (truncated to prevent intracellular signaling) staining in clonal populations of cells after single cell clonal density seeding, or in pooled samples of cells after targeting with PRDX1-AKR1A1 pTRE3G-CD19t-IL12 donor template, after 48 hours of treatment with 2 μg / mL doxycycline, compared to untreated cells.Panel A shows cells after targeting with Cpf1-based RNP.Panel B shows cells after targeting with Cas9-based RNP. [Figure 13] 13 is a panel of fluorescent microscopy images showing expression of GFP in a pooled population of cells treated with doxycycline for 24 hours. Doxycycline was added to the medium 48 hours after nucleofection of iPSCs with three different RNPs containing three different gRNAs targeting site 2 and the PRDX1-AKR1A1 site 2 targeting construct. No GFP was observed in cells to which doxycycline was not added. [Figure 14]14 is a panel of fluorescent microscopy images showing expression of GFP in a pooled population of cells treated with doxycycline for 24 hours. Doxycycline was added to the medium 24 hours after nucleofection of iPSCs with three different RNPs containing three different gRNAs targeting site 3 and the PRDX1-AKR1A1 site 3 targeting construct. No GFP was observed in cells to which doxycycline was not added. [Figure 15] Figure 15 is a panel of flow cytometry histograms showing induction of GFP expression in a pooled population of cells following treatment with 2 μg / ml doxycycline. Doxycycline was added to the medium 48 hours after nucleofection of iPSCs with three different RNPs containing three different gRNAs targeting site 2 and the PRDX1-AKR1A1 site 2 targeting construct. Flow cytometry analysis was performed 5 days after doxycycline treatment. No GFP was observed in cells to which doxycycline was not added, as well as in parental GAPDH::rtTA iPSCs to which no targeting construct and RNP were added. [Figure 16] Figure 16 is a panel of flow cytometry histograms showing induction of GFP expression in a pooled population of cells after treatment with 2 μg / ml doxycycline. Doxycycline was added to the medium 24 hours after nucleofection of iPSCs with three different RNPs containing three different gRNAs targeting site 3 and the PRDX1-AKR1A1 site 3 targeting construct. Flow cytometry analysis was performed 6 days after doxycycline treatment. No GFP was observed in cells to which doxycycline was not added, as well as in parental GAPDH::rtTA iPSCs to which no targeting construct and RNP were added.
[0034] Detailed Description Genetically engineered cells are an important tool for cell therapy. However, artificial gene circuits in engineered cells are often disrupted over time by transgene silencing as cells proliferate or undergo changes in cellular conditions or in vivo environments. There is therefore a need to identify genomic regions that are safe for transgene integration and that also result in a chromatin landscape that remains open to transcription across cell types, cellular conditions, and in vivo environments. Transgene integration within such sites would allow the transgene to remain transcriptionally active for the life of the cell therapy product.
[0035] Provided herein are compositions (e.g., compositions of nucleic acid molecules and cells) and methods for genomically (genetically) engineering cells to achieve expression of transgenes across various cellular or differentiation states without effects on endogenous gene expression that may be detrimental to the cells or to the therapeutic purpose of the cells in cell therapy. The provided compositions and methods are based, at least in part, on the identification of chromatin landscapes that contain sustained transcriptionally active payload regions (STAPLRs) that maintain transcriptional activity across cell types and differentiated cell states.
[0036] I. STAPLR The inventors have found that certain intergenic regions within the mammalian genome allow for a constant level of expression of transgenes integrated therein, regardless of cell type and / or even as the cell undergoes changes in its state (e.g., differentiation state, maturation, or activity state). This finding significantly expands the repertoire of genomic sites where a transgene can be stably integrated and its expression can be maintained across changing cellular states. This finding therefore solves a long-standing problem in transgene expression, for example in the case of cell therapy. These intergenic regions are referred to herein as "persistent transcriptionally active payload regions" (STAPLRs), where "payload" or "genomic payload" refers to one or more exogenous or heterologous nucleotide sequences introduced into the region. STAPLRs contain an open chromatin landscape for landing the genomic payload. The chromosomal DNA within the STAPLRs is accessible to components of the gene editing machinery and is in a conformation that allows for integration of the genetic material. In some cases, STAPLRs are located in the vicinity of transcriptionally active genes.
[0037] One application of this finding is the efficient generation of cells (e.g., therapeutic cells) where cells are first genetically modified and then the cell state is changed, for example, by differentiation or dedifferentiation. For example, this genetic engineering method can be applied to iPSCs that are subsequently differentiated into various cell types. In conventional methods, when iPSCs are engineered to incorporate a transgene into their genome and then differentiated into a desired cell type, the transgene may become inactive upon differentiation of the iPSCs. However, the transgene incorporated within the STAPLR disclosed herein does not become inactive upon differentiation of the iPSCs. Thus, STAPLR provides a universal "landing pad" for transgene expression.
[0038] This stability in transgene expression is also advantageous after the therapeutic cells in cell therapy are administered to a subject in need thereof (e.g., a human patient), as the therapeutic cells may encounter different and changing environments that would cause the transgene to be rendered inoperative in the case of a transgene integrated at another site.
[0039] Furthermore, integration of the transgene into an intergenic region, rather than intragenic, minimizes disruption of the expression or regulation of adjacent genes, thus allowing normal function of the genetically engineered cells. Transgene integration in STAPLR also reduces the risk of causing undesired effects in cells (e.g., activation of oncogenes, or disruption of essential genes such as tumor suppressor genes). Furthermore, STAPLR, with its consistent transcriptionally active state, allows for the testing and use of a wider range of regulatory elements (e.g., promoters and enhancers).
[0040] An "intergenic region" as used herein is a stretch of nucleotide sequence located between two adjacent genes. Intergenic regions can be of various sizes. For example, an intergenic region can be at least 30, 40, 50, 75, or 100 base pairs long. In some embodiments, an intergenic region can be at least 150, 200, 300, 400, 500, 750, or 1000 base pairs long. In some embodiments, an intergenic region can be at least 1500, 2000, 2500, 3000, 3500, 5000, or 10000 base pairs long. In some embodiments, an intergenic region can be at least 15000, 20000, 30000, 40000, 50000, 75000, or 100000 base pairs long. In some embodiments, an intergenic region is between 30 base pairs long and 100000 base pairs long. In some embodiments, the intergenic region is between 50 base pairs and 75,000 base pairs in length. In some embodiments, the intergenic region is between 75 base pairs and 70,000 base pairs in length.
[0041] STAPLRs of the present disclosure include, but are not limited to, the following (NCBI gene IDs for human genes are shown in parentheses): the intergenic region between the RPL34 gene (gene ID: 6164) and the OSTC gene (gene ID: 58505), the intergenic region between the ACTB gene (gene ID: 60) and the FSCN1 gene (gene ID: 6624), the intergenic region between the AKIRIN1 gene (gene ID: 79647) and the NDUFS5 gene (gene ID: 4725), the intergenic region between the PRDX1 gene (gene ID: 10666), the intergenic region between the PRDX2 gene (gene ID: 10666), the intergenic region between the PRDX1 ... intergenic region between the AKR1A1 gene (gene ID: 10327), the PTGES3 gene (gene ID: 10728) and the NACA gene (gene ID: 4666), the MLF2 gene (gene ID: 8079) and the PTMS gene (gene ID: 5763), the RAB13 gene (gene ID: 5872) and the RPS27 gene (gene ID: 4840565), the JTB gene (gene ID: 10899) and the RAB13 gene (gene ID: 5872), and the AKR1A2 gene (gene ID: 10327). Intergenic region, the intergenic region between the AKR1A1 gene (gene ID: 10327) and the NASP gene (gene ID: 4678), the intergenic region between the NDUFS5 gene (gene ID: 4725) and the MACF1 gene (gene ID: 23499), the intergenic region between the SRSF9 gene (gene ID: 8683) and the DYNLL1 gene (gene ID: 8655), the intergenic region between the MYL6B gene (gene ID: 140465) and the MYL6 gene (gene ID: 4637), the intergenic region between the GPX1 gene (gene ID: 2876) and the R the intergenic region between the HOA gene (gene ID: 387), the intergenic region between the HNRNPA2B1 gene (gene ID: 3181) and the CBX3 gene (gene ID: 11335), the intergenic region between the ROMO gene (gene ID: 140823) and the RBM39 gene (gene ID: 9584), the intergenic region between the PA2G4 gene (gene ID: 5036) and the RPL41 gene (gene ID: 6171), and the intergenic region between the NDUFB10 gene (gene ID: 4716) and the RPS2 gene (gene ID: 6187).In some embodiments, the genes herein refer to human genes and the mammalian cells are human cells.
[0042] The start and end genomic coordinates and sizes of the STAPLR intergenic regions in the human genome are shown below in Table 1. The coordinates are defined by the information available in the RefSeq database of NCBI.
[0043] [Table 1]
[0044] Due to variation among humans and between mammalian species, the intergenic regions between the gene pairs may vary to some extent from the corresponding SEQ ID NOs shown in Table 1.
[0045] In some embodiments, the intergenic region between the RPL34 gene and the OSTC gene comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:1, or is sufficiently similar to SEQ ID NO:1 such that the intergenic region retains the functionality of SEQ ID NO:1, i.e., such that the function (e.g., transcriptional regulation) of the intergenic region between the RPL34 gene and the OSTC gene remains intact (e.g., has no adverse effects on the cell).
[0046] In some embodiments, the intergenic region between the ACTB gene and the FSCN1 gene comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:2, or is sufficiently similar to SEQ ID NO:2 such that the intergenic region retains the functionality of SEQ ID NO:2, i.e., such that the function (e.g., transcriptional regulation) of the intergenic region between the ACTB gene and the FSCN1 gene remains intact (e.g., has no adverse effects on the cell).
[0047] In some embodiments, the intergenic region between the AKIRIN1 gene and the NDUFS5 gene comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:3, or is sufficiently similar to SEQ ID NO:3 such that the intergenic region retains the functionality of SEQ ID NO:3, i.e., such that the function (e.g., transcriptional regulation) of the intergenic region between the AKIRIN1 gene and the NDUFS5 gene remains intact (e.g., has no adverse effects on the cell).
[0048] In some embodiments, the intergenic region between the PRDX1 gene and the AKR1A1 gene comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:4, or is sufficiently similar to SEQ ID NO:4 such that the intergenic region retains the functionality of SEQ ID NO:4, i.e., such that the function (e.g., transcriptional regulation) of the intergenic region between the PRDX1 gene and the AKR1A1 gene remains intact (e.g., has no adverse effects on the cell).
[0049] In some embodiments, the intergenic region between the PTGES3 gene and the NACA gene comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:5, or is sufficiently similar to SEQ ID NO:5 such that the intergenic region retains the functionality of SEQ ID NO:5, i.e., such that the function (e.g., transcriptional regulation) of the intergenic region between the PTGES3 gene and the NACA gene remains intact (e.g., has no adverse effects on the cell).
[0050] In some embodiments, the intergenic region between the MLF2 gene and the PTMS gene comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:6, or is sufficiently similar to SEQ ID NO:6 such that the intergenic region retains the functionality of SEQ ID NO:6, i.e., such that the function (e.g., transcriptional regulation) of the intergenic region between the MLF2 gene and the PTMS gene remains intact (e.g., has no adverse effects on the cell).
[0051] In some embodiments, the intergenic region between the RAB13 gene and the RPS27 gene comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:7, or is sufficiently similar to SEQ ID NO:7 such that the intergenic region retains the functionality of SEQ ID NO:7, i.e., such that the function (e.g., transcriptional regulation) of the intergenic region between the RAB13 gene and the RPS27 gene remains intact (e.g., has no adverse effects on the cell).
[0052] In some embodiments, the intergenic region between the JTB gene and the RAB13 gene comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:8, or is sufficiently similar to SEQ ID NO:8 such that the intergenic region retains the functionality of SEQ ID NO:8, i.e., such that the function (e.g., transcriptional regulation) of the intergenic region between the JTB gene and the RAB13 gene remains intact (e.g., has no adverse effects on the cell).
[0053] In some embodiments, the intergenic region between the AKR1A1 gene and the NASP gene comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:9, or is sufficiently similar to SEQ ID NO:9 such that the intergenic region retains the functionality of SEQ ID NO:9, i.e., such that the function (e.g., transcriptional regulation) of the intergenic region between the AKR1A1 gene and the NASP gene remains intact (e.g., has no adverse effects on the cell).
[0054] In some embodiments, the intergenic region between the NDUFS5 gene and the MACF1 gene comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:10 or is sufficiently similar to SEQ ID NO:10 such that the intergenic region retains the functionality of SEQ ID NO:10, i.e., the function (e.g., transcriptional regulation) of the intergenic region between the NDUFS5 gene and the MACF1 gene remains intact (e.g., does not adversely affect the cell).
[0055] In some embodiments, the intergenic region between the SRSF9 gene and the DYNLL1 gene comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:11, or is sufficiently similar to SEQ ID NO:11 such that the intergenic region retains the functionality of SEQ ID NO:11, i.e., such that the function (e.g., transcriptional regulation) of the intergenic region between the SRSF9 gene and the DYNLL1 gene remains intact (e.g., does not adversely affect the cell).
[0056] In some embodiments, the intergenic region between the MYL6B gene and the MYL6 gene comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:12, or is sufficiently similar to SEQ ID NO:12 such that the intergenic region retains the functionality of SEQ ID NO:12, i.e., such that the function (e.g., transcriptional regulation) of the intergenic region between the MYL6B gene and the MYL6 gene remains intact (e.g., has no adverse effects on the cell).
[0057] In some embodiments, the intergenic region between the GPX1 gene and the RHOA gene comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:13, or is sufficiently similar to SEQ ID NO:13 such that the intergenic region retains the functionality of SEQ ID NO:13, i.e., such that the function (e.g., transcriptional regulation) of the intergenic region between the GPX1 gene and the RHOA gene remains intact (e.g., does not adversely affect the cell).
[0058] In some embodiments, the intergenic region between the HNRNPA2B1 gene and the CBX3 gene comprises a nucleotide sequence that is at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 14, or is sufficiently similar to SEQ ID NO: 14 such that the intergenic region retains the functionality of SEQ ID NO: 14, i.e., such that the function (e.g., transcriptional regulation) of the intergenic region between the HNRNPA2B1 gene and the CBX3 gene remains intact (e.g., has no adverse effects on the cell).
[0059] In some embodiments, the intergenic region between the ROMO gene and the RBM39 gene comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 15, or is sufficiently similar to SEQ ID NO: 15 such that the intergenic region retains the functionality of SEQ ID NO: 15, i.e., such that the function (e.g., transcriptional regulation) of the intergenic region between the ROMO gene and the RBM39 gene remains intact (e.g., does not adversely affect the cell).
[0060] In some embodiments, the intergenic region between the PA2G4 gene and the RPL41 gene comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:16, or is sufficiently similar to SEQ ID NO:16 such that the intergenic region retains the functionality of SEQ ID NO:16, i.e., such that the function (e.g., transcriptional regulation) of the intergenic region between the PA2G4 gene and the RPL41 gene remains intact (e.g., has no adverse effects on the cell).
[0061] In some embodiments, the intergenic region between the NDUFB10 and RPS2 genes comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:16, or is sufficiently similar to SEQ ID NO:16 such that the intergenic region retains the functionality of SEQ ID NO:16, i.e., such that the function (e.g., transcriptional regulation) of the intergenic region between the NDUFB10 and RPS2 genes remains intact (e.g., does not adversely affect the cell).
[0062] The percent identity of two nucleotide sequences can be determined, for example, by BLAST® using default parameters (available at the US National Library of Medicine's National Center for Biotechnology Information website). In some embodiments, the length of a reference sequence that is aligned for comparison purposes is at least 30% (e.g., at least 40, 50, 60, 70, 80, or 90%) of the reference sequence.
[0063] II. Integration of exogenous sequences into STAPLR A. Integration site The exogenous nucleotide sequence of interest can be integrated at any site within STAPLR. For example, the integration site or the junction between the exogenous sequence and the adjacent endogenous sequence can be located in the first or second half of STAPLR; or in the 5', middle or 3' quarter of STAPLR; or in the first, second, third or fourth quarter of STAPLR. In some embodiments, the integration site of the exogenous sequence, or the junction between the exogenous sequence and the adjacent endogenous sequence, is located within STAPLR and is at least 10, 20, 30, 40, 50, 80, 90, 100, 200, 300, 400, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 5000, 10000, 15000 or 20000 base pairs away from the nearest gene, i.e., the 5' or 3' boundary of STAPLR (e.g., from the start or end coordinates shown in Table 1).
[0064] In a single genome, one or more exogenous nucleotide sequences can be integrated into one or more STAPLRs. In some embodiments, one or more (e.g., two, three or four) exogenous nucleotide sequences can be integrated into one or more sites in a single given STAPLR. In some embodiments, multiple STAPLRs in a single genome are targeted for integration of exogenous nucleotide sequences.
[0065] In some embodiments, exogenous sequences are introduced at at least one STAPLR and at least one persistent transgene expression locus (STEL) (described in WO2021 / 072329). STEL sites are loci of endogenous genes that are robustly and consistently expressed during pluripotent states and differentiation [e.g., as examined by single-cell RNA sequencing (scRNAseq) analysis]. STAPLR may be associated with a STEL site, but it need not be associated with a STEL site. STEL sites may be identified from single-cell RNA-seq data. A defining characteristic of desirable STEL sites is ubiquity of expression. STEL sites may be identified by analyzing expression of candidate loci across diverse cell types and cell maturation states, such as PSCs and PSC-derived dopamine neurons (and selected progenitor states), microglia (and selected progenitor states), and cardiomyocytes (and selected cardiomyocyte progenitor states). The addition of publicly available single-cell RNA-seq data of adult human tissues allows for the refinement of such STEL analysis. STELs include, but are not limited to, certain housekeeping genes that are active in multiple cell types, such as those involved in gene expression (e.g., transcription factors and histones), cellular metabolism (e.g., GAPDH and NADH dehydrogenase) or cellular structure (e.g., actin), or those that encode ribosomal proteins (e.g., large or small ribosomal subunits, e.g., RPL13A, RPLP0 and RPL7).Examples of STELs include genes encoding ribosomal proteins, such as the RPL genes (e.g., RPL13A, RPLP0, RPL10, RPL13, RPS18, RPL3, RPLP1, RPL15, RPL41, RPL11, RPL32, RPL18A, RPL19, RPL28, RPL29, RPL9, RPL8, RPL6, RPL18, RPL7, RPL7A, RPL21, RPL37A, RPL12, RPL5, RPL34, RPL35A, RPL30, RPL24, RPL39, RPL37, RPL14, RPL27A, RPLP2, RPL23A, RPL26, RPL36, RPL35, RPL23, RPL4, and RPL22) and the RPS genes (e.g., RP S2, RPS19, RPS14, RPS3A, RPS12, RPS3, RPS6, RPS23, RPS27A, RPS8, RPS4X, RPS7, RPS24, RPS27, RPS15A, RPS9, RPS28, RPS13, RPSA, RPS5, RPS16, RPS25, RPS15, RPS20 and RPS11); genes encoding mitochondrial proteins (e.g., MT-CO1, MT-CO2, MT-ND4, MT-ND1 and MT-ND2); genes encoding actin proteins (ACTG1 and ACTB); genes encoding eukaryotic translation factors (e.g., EEF1A1, EEF2 and EIF1); and genes encoding histones (e.g., H3F3A and H3F3B). Additional STELs include those that encode proteins involved in focal adhesion, cell-substrate adhesion junctions, cell-substrate binding, cell anchors, extracellular exosomes, extracellular vesicles, intracellular organelles or anchor binding. Additional examples of STELs include FTL, FTH1, TPT1, TMSB10, GAPDH, PTMA, GNB2L1, NACA, YBX1, NPM1, FAU, UBA52, HSP90AB1, MYL6, SERF2 and SRP14.
[0066] In some embodiments, exogenous sequences are introduced into STAPLR, such as RPL34-OSTC or PRDX1-AKR1A1 STAPLR, and STEL, such as GAPDH locus, in a single mammalian (e.g., human) genome. In some embodiments, exogenous sequences are introduced into multiple STAPLR, such as RPL34-OSTC STAPLR and PRDX1-AKR1A1 STAPLR, in a single genome.
[0067] The integration site of the exogenous nucleotide sequence may be within STAPLR or within a gene sequence adjacent to STAPLR (e.g., an exon, intron, or UTR of the gene). In some embodiments, the endonuclease creates a DNA nick within STAPLR. In other embodiments, the endonuclease creates a DNA nick in the gene adjacent to STAPLR such that after integration, the exogenous nucleotide sequence remains integrated within STAPLR. In some embodiments, screening for improper integration events can be performed according to the methods described in WO2021 / 226151, where a DNA nick is introduced in an exon of a gene adjacent to STAPLR and required for cell survival, and cells in which integration is not properly achieved do not survive.
[0068] B. Method of incorporation Any genome integration method can be used to utilize the STAPLR described herein. In some embodiments, the integration of exogenous nucleotide sequence in STAPLR is achieved by using a genome editing system selected from the group consisting of CRISPR / Cas system, Cre / Lox system, FLP-FRT system, transcription activator-like effector nuclease (TALEN) system, zinc finger nuclease (ZFN) system, homing endonuclease, sequence-specific endonuclease, random integration (e.g., by transposon), meganuclease, homologous recombination, transposase and non-nuclease-dependent viral vector (e.g., retrovirus, AAV or lentivirus vector). In some embodiments, integration does not cause the deletion of endogenous sequences in the region and / or the addition of nucleotide sequences other than the exogenous donor sequence to be integrated. In some embodiments, integration causes insertion and / or deletion (indel) (of non-donor sequences) at the integration site.
[0069] In some embodiments, the exogenous sequence can be integrated into the STAPLR site by homologous recombination at a DNA nick generated by a suitable endonuclease, such as, for example, a CRISPR-associated endonuclease, which can be, for example, a Cas endonuclease selected from, but not limited to, a type I (e.g., subtype IA, IB, IC, IC variant, ID, IE, IF, IF variant 1 or IF variant 2), type II (e.g., subtype II-A, II-B, II-B or II-C), type III (e.g., subtype III-A, III-B or III-B variant), type IV or type V Cas protein or variants thereof. In some embodiments, the nuclease is Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas12 (e.g., Cas12a or Cpf1 or Cas12b), Cas13, Cas100, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr2, Csm3, Csm4, Csm5, Csm6, Cmr1, Csm ...sm4, Csm4, Csm5, Csm6, Csm4, Csm4, Csm5, Csm6, Csm4, Csm4, Csm5, Csm6, Csm4, Csm4, Csm5, Csm6, Csm4, Csm4, Csm4, Csm5, Csm6, Csm4, Csm4, Csm4, Csm4, Csm4, Csm4, 3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, CasX, CasY, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, CasPhi, MAD7, Csf4 and homologs or modified forms thereof (e.g., truncated forms or mutants of wild-type Cas proteins that have nuclease activity).
[0070] In some embodiments, the Cas endonuclease is a Cpf1 (Cas12a) endonuclease or a variant, derivative or fragment thereof, such as Cpf1 (FnCpf1) from Francisella novicida U112, Cpf1 (AsCpf1, e.g., an improved variant, e.g., enAsCpf1) from Acidaminococcus sp. BV3L6, Cpf1 (LbCpf1) from Lachnospiraceae bacterium ND2006, Cpf1 (Lb2Cpf1) from Lachnospiraceae bacterium MA2020, Cpf1 (Lb3Cpf1) from Lachnospiraceae bacterium MC2017, Cpf1 (Lb4Cpf1) from Moraxella boehmii, Cpf1 (Lb5Cpf1) from Moraxella borbacillus, Cpf1 (Lb6Cpf1) from Moraxella borbacillus, Cpf1 (Lb7Cpf1) from Moraxella borbacillus, Cpf1 (Lb8Cpf1) from Moraxella borbacillus, Cpf1 (Lb9Cpf1) from Moraxella borbacillus, Cpf1 (Lb10Cpf1) from Moraxella borbacillus, Cpf1 (Lb11Cpf1) from Moraxella borbacillus, Cpf1 (Lb12Cpf1) from Moraxella borbacillus, Cpf1 (Lb13Cpf1) from Moraxella borbacillus, Cpf1 (Lb14Cpf1) from Moraxella borbacillus, Cpf1 (Lb15Cpf1) from Moraxella borbacillus, Cpf1 (Lb16Cpf1) from Morax The Cpf1 genes are Cpf1 derived from M. bovoculi237 (MbCpf1) or Cpf1 derived from Prevotella disiens (PdCpf1).
[0071] In some embodiments, the Cas endonuclease is a Cas9 protein or a variant, derivative or fragment thereof. In some embodiments, the Cas9 protein is SaCas9, SpCas9, SpCas9n, Cas9-HF, Cas9-H840A, FokI-dCas9, or D10A nickase.
[0072] In some embodiments, the Cas endonuclease is a V-type RNA programmable nuclease as disclosed in WO2022 / 258753.
[0073] In some embodiments, the Cas endonuclease is a MAD nuclease, such as the MAD7 nuclease disclosed in U.S. Pat. No. 10,337,028.
[0074] Non-limiting examples of suitable endonucleases are provided in Table A below.
[0075] [Table 2] TIFF2025514159000003.tif241161 TIFF2025514159000004.tif110162
[0076] In some embodiments, the CRISPR / Cas system comprises a gRNA-dependent nuclease (or its coding sequence) that targets a selected intergenic region, a gRNA (or its coding sequence), and donor DNA comprising an exogenous nucleotide sequence.
[0077] In some embodiments, STAPLR is the intergenic region between the RPL34 gene and the OSTC gene, and the gRNA is selected from SEQ ID NOs: 25-32.
[0078] In some embodiments, STAPLR is the intergenic region between the ACTB gene and the FSCN1 gene, and the gRNA is selected from SEQ ID NOs: 33-54.
[0079] In some embodiments, STAPLR is the intergenic region between the AKIRIN1 gene and the NDUFS5 gene, and the gRNA is selected from SEQ ID NOs: 55-70.
[0080] In some embodiments, STAPLR is the intergenic region between the PRDX1 gene and the AKR1A1 gene, and the gRNA is selected from SEQ ID NOs: 71-92.
[0081] C. Exogenous Nucleotide Sequences In some embodiments, the exogenous nucleotide sequence of interest for integration may comprise a transgene encoding a protein (as used herein, including a peptide) or encoding an RNA. The transgene may comprise a coding sequence for a gene product and, optionally, one or more transcriptional regulatory elements. In some embodiments, the transgene comprises one or more regulatory elements, wherein the one or more regulatory elements may optionally be operably linked to the coding sequence.
[0082] Non-limiting examples of regulatory elements include promoters, enhancers, silencers, chromatin insulators, intron sequences, Kozak sequences, ubiquitous chromatin opening elements (UCOEs), transcription activator binding elements, sequences that enhance gene expression or RNA stability (e.g., WPRE elements), polyadenylation signal sequences (e.g., SV40 polyA signals), and the like.
[0083] In some embodiments, the promoter directing the expression of the transgene is a constitutive promoter, including, but not limited to, EF1a, EFS, UBC, PGK, CAGGS, CMV, SV40, B2M, and ROSA26 promoters. In some embodiments, the promoter is a cell type-specific, tissue-specific, or lineage-specific promoter. For example, the promoter can be the tyrosine hydroxylase promoter of dopaminergic neurons, the Hb9 promoter of motor neurons, the SIRPA promoter of cardiomyocytes, the CD14, CD33, CD45, or CD11b promoter of myeloid cells, or the CD3, FOXP3, CD25, CD8, or CD4 promoter of T lymphocytes. In some embodiments, the expression of the transgene is under the control of an inducible promoter (e.g., the lac operon, which can be induced by isopropyl β-D-1-thiogalactopyranoside (IPTG); the TRE promoter, which can be induced by tetracycline and its derivatives).
[0084] In some embodiments, the exogenous sequence includes one or more regulatory elements that are responsive to a factor expressed from another site (e.g., from an endogenous gene or from a transgene integrated into the STEL or STAPLR). A non-limiting example of such a regulatory element is a transcription factor binding site. In some embodiments, such a regulatory element is integrated into the STAPLR site adjacent to the coding sequence of the transgene and / or one or more other regulatory elements. For example, a cell can be modified with a DNA molecule disclosed herein that includes an exogenous nucleotide sequence that includes a transcription factor binding site and a transgene, where a transcription factor that can bind to the transcription factor binding site is expressed from an endogenous gene or from another transgene in any part of the genome (e.g., a STAPLR, STEL or another safe harbor site), or is ectopically expressed.
[0085] In some embodiments, the transgene encodes an RNA (e.g., a small interfering RNA or a microRNA) or a protein of interest. The protein of interest (as used herein, includes peptides) can be, for example, a globular protein (e.g., albumin, globulin, glutelin, prolamin, histone, globin, or protamine), a fibrous protein (e.g., a scleroprotein, such as collagen, elastin, keratin, or fibroin), or an intermediate protein. In some embodiments, the protein of interest is a complex protein, such as a metalloprotein, a chromoprotein, a glycoprotein, a mucoprotein, a phosphoprotein, or a lipoprotein. In some embodiments, the protein of interest is a therapeutic protein (e.g., a protein that can ameliorate or prevent symptoms of a disease or condition). Non-limiting examples of therapeutic proteins include proteins that are deficient or defective in genetic diseases, such as hemophilia and lysosomal storage diseases, hormones, enzymes, cytokines that regulate immunity, recombinant antigen receptors (e.g., chimeric antigen receptors), antibodies, proteins that regulate the differentiation or activity of engineered cells (e.g., transcription factors, or proteins that maintain cells in M1 or M2 polarity), and the like. In some embodiments, the protein of interest is a cell marker, a protein used for immune evasion, or a safety or kill switch used in cell therapy. Examples of proteins of interest include, but are not limited to, SOX10, IL-10, IL-12, CD19t, and ThPOK.
[0086] D. Targeting Vectors The present disclosure provides a targeting vector for integrating an exogenous nucleotide sequence into STAPLR. As used herein, a "targeting vector" is a nucleic acid that includes an exogenous nucleotide sequence of interest and a sequence that is homologous to an endogenous chromosomal nucleotide sequence adjacent to a desired integration location in a genome. These adjacent homologous sequences are referred to as "homology arms". The homology arms direct the targeting vector to a specific chromosomal location in a genome by the homology that exists between the homology arms and the corresponding endogenous nucleotide sequence. In some embodiments, the targeting vector is a DNA molecule that includes a nucleotide sequence of interest adjacent to a 5' nucleotide sequence (left homology arm or homology region) and a 3' nucleotide sequence (right homology arm or homology region), where the 5' nucleotide sequence and the 3' nucleotide sequence are homologous to the nucleotide sequence adjacent to an integration site in the genome of a cell, resulting in the integration of the nucleic acid of interest by homologous recombination into the integration site.
[0087] The 5' and 3' sequences are sufficiently similar to the endogenous nucleotide sequence targeted for homologous recombination such that when the homology arms are integrated (in whole or in part), they do not adversely affect the genetic environment of integration (e.g., do not affect the function of adjacent genes). In some embodiments, the homology arms are at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the nucleotide sequence in the STAPLR to be targeted.
[0088] For example, in some embodiments, the intergenic region between the RPL34 gene and the OSTC gene comprises a nucleotide sequence that is at least 80% identical to SEQ ID NO: 1, and the function of the intergenic region between the RPL34 gene and the OSTC gene remains intact following integration. In some embodiments, the intergenic region between the RPL34 gene and the OSTC gene comprises a nucleotide sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO: 1, and the function of the intergenic region between the RPL34 gene and the OSTC gene remains intact following integration.the intergenic region between the ACTB gene and the FSCN1 gene and its identity to SEQ ID NO:2, the intergenic region between the AKIRIN1 gene and the NDUFS5 gene and its identity to SEQ ID NO:3, the intergenic region between the PRDX1 gene and the AKR1A1 gene and its identity to SEQ ID NO:4, the intergenic region between the PTGES3 gene and the NACA gene and its identity to SEQ ID NO:5, the intergenic region between the MLF2 gene and the PTMS gene and its identity to SEQ ID NO:6, the intergenic region between the RAB13 gene and the RPS27 gene and its identity to SEQ ID NO:7, the intergenic region between the JTB gene and the RAB13 gene and its identity to SEQ ID NO:8, the intergenic region between the AKR1A1 gene and the NASP gene and its identity to SEQ ID NO:9, the intergenic region between the NDUFS5 gene and the M The same may be true for the intergenic region between the ACF1 gene and its identity to SEQ ID NO: 10, the intergenic region between the SRSF9 gene and the DYNLL1 gene and its identity to SEQ ID NO: 11, the intergenic region between the MYL6B gene and the MYL6 gene and its identity to SEQ ID NO: 12, the intergenic region between the GPX1 gene and the RHOA gene and its identity to SEQ ID NO: 13, the intergenic region between the HNRNPA2B1 gene and the CBX3 gene and its identity to SEQ ID NO: 14, the intergenic region between the ROMO gene and the RBM39 gene and its identity to SEQ ID NO: 15, the intergenic region between the PA2G4 gene and the RPL41 gene and its identity to SEQ ID NO: 16, and the intergenic region between the NDUFB10 and RPS2 genes and its identity to SEQ ID NO: 97.
[0089] In the disclosed methods, the homology arms can vary in length. In some embodiments, each of the homology arms is independently at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 550, at least about 600, at least about 650, at least about 700, at least about 750, at least about 800, at least about 850, at least about 900, at least about 950, at least about 1000, at least about 1100, at least about 1200, at least about 1300, at least about 1400, at least about 1500, at least about 1600, at least about 1700, at least about 1800, at least about 1900, or at least about 2000 base pairs in length. In some embodiments, each homology arm is independently 50-2000, 50-1500, 100-1900, 150-1800, 200-1700, 250-1600, 300-1500, 350-1400, 400-1300, 450-1200, 500-1100, 550-1000, 600-950, 650-900, 700-850, or 750-800 base pairs in length.
[0090] In the methods of the present disclosure, the homology arms (i.e., the 5' and 3' nucleotide sequences) can be designed to target any site within the disclosed intergenic region. The homology arms can be designed based on genomic sequences available in sequence databases (e.g., the NCBI database).
[0091] In some embodiments, the 5' nucleotide sequence comprises a nucleotide sequence that is sufficiently similar to SEQ ID NO: 17 to the extent necessary for the function of the sequence to remain intact, and the 3' nucleotide sequence comprises a nucleotide sequence that is sufficiently similar to SEQ ID NO: 18 to the extent necessary for the function of the sequence to remain intact. In some embodiments, the 5' nucleotide sequence comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 17 such that the function of SEQ ID NO: 17 remains intact. In some embodiments, the 3' nucleotide sequence comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 18 such that the function of SEQ ID NO: 18 remains intact.
[0092] In some embodiments, the 5' nucleotide sequence comprises a nucleotide sequence that is sufficiently similar to SEQ ID NO: 19 to the extent necessary for the function of the sequence to remain intact, and the 3' nucleotide sequence comprises a nucleotide sequence that is sufficiently similar to SEQ ID NO: 20 to the extent necessary for the function of the sequence to remain intact. In some embodiments, the 5' nucleotide sequence comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 19, such that the function of SEQ ID NO: 19 remains intact. Similarly, in some embodiments, the 3' nucleotide sequence comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 20, such that the function of SEQ ID NO: 20 remains intact.
[0093] In some embodiments, the 5' nucleotide sequence comprises a nucleotide sequence that is sufficiently similar to SEQ ID NO:21 to the extent necessary for the function of the sequence to remain intact, and the 3' nucleotide sequence comprises a nucleotide sequence that is sufficiently similar to SEQ ID NO:22 to the extent necessary for the function of the sequence to remain intact. In some embodiments, the 5' nucleotide sequence comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:21 such that the function of SEQ ID NO:21 remains intact. Similarly, in some embodiments, the 3' nucleotide sequence comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:22 such that the function of SEQ ID NO:22 remains intact.
[0094] In some embodiments, the 5' nucleotide sequence comprises a nucleotide sequence that is sufficiently similar to SEQ ID NO:23 to the extent necessary for the function of the sequence to remain intact, and the 3' nucleotide sequence comprises a nucleotide sequence that is sufficiently similar to SEQ ID NO:24 to the extent necessary for the function of the sequence to remain intact. In some embodiments, the 5' nucleotide sequence comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:23 such that the function of SEQ ID NO:23 remains intact. Similarly, in some embodiments, the 3' nucleotide sequence comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:24 such that the function of SEQ ID NO:24 remains intact.
[0095] In some embodiments, the 5' nucleotide sequence comprises a nucleotide sequence that is sufficiently similar to SEQ ID NO:93 to the extent necessary for the function of the sequence to remain intact, and the 3' nucleotide sequence comprises a nucleotide sequence that is sufficiently similar to SEQ ID NO:94 to the extent necessary for the function of the sequence to remain intact. In some embodiments, the 5' nucleotide sequence comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:93 such that the function of SEQ ID NO:93 remains intact. Similarly, in some embodiments, the 3' nucleotide sequence comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to SEQ ID NO:94 such that the function of SEQ ID NO:94 remains intact.
[0096] In some embodiments, the 5' nucleotide sequence comprises a nucleotide sequence that is sufficiently similar to SEQ ID NO:95 to the extent necessary for the function of the sequence to remain intact, and the 3' nucleotide sequence comprises a nucleotide sequence that is sufficiently similar to SEQ ID NO:96 to the extent necessary for the function of the sequence to remain intact. In some embodiments, the 5' nucleotide sequence comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:95, such that the function of SEQ ID NO:95 remains intact. Similarly, in some embodiments, the 3' nucleotide sequence comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:96, such that the function of SEQ ID NO:96 remains intact.
[0097] In some embodiments, the homology arms are entirely contained within the target STAPLR, in other embodiments, the homology arms may overlap portions of adjacent genes without interfering with their function after integration, and the exogenous sequence is still integrated within the STAPLR.
[0098] In some embodiments, the targeting vector is a circular vector.In some embodiments, the targeting vector is a linear vector.In some embodiments, the targeting vector provided herein comprises one or more endonuclease targeting sequences, for example, for linearizing the vector when used in combination with endonuclease-guiding.In some embodiments, the targeting vector is a viral vector (e.g., AAV vector, adenovirus vector, lentivirus vector, herpes simplex virus vector) or a plasmid vector.
[0099] The present disclosure provides a STAPLR targeting system comprising a targeting vector described herein and a suitable gene editing system (e.g., as described herein) for incorporating a nucleotide sequence of interest on the targeting vector into the STAPLR.
[0100] III. Genetically modified mammalian cells Provided herein is a genetically modified cell that comprises one or more modifications in STAPLR as disclosed herein. The mammalian cell targeted for STAPLR integration can be any cell type or any cell state of interest. For example, the cell can be a pluripotent cell (e.g., pluripotent stem cell) or a differentiated cell. Cells, such as human cells, can be engineered in vitro, in vivo or ex vivo by gene editing methods as described herein. Cells can be non-human cells, such as cells from laboratory animals (e.g., non-human primates, mice, rats and rabbits), livestock (e.g., cows and horses) and pets (e.g., dogs and cats).
[0101] A. Stem cells In some embodiments, the mammalian cells targeted for modification in STAPLR are stem cells, particularly pluripotent stem cells (PSCs), such as induced pluripotent stem cells (iPSCs, e.g., human iPSCs) or embryonic stem cells (ESCs, e.g., human ESCs). The engineered stem cells can then be induced to differentiate into a desired cell type (referred to herein as PSC derivatives, PSC derivative cells, or PSC-derived cells). Stem cells can be the starting point for the potential generation of large numbers of cells of a particular cell type to be delivered for regenerative medicine in patients with various diseases.
[0102] As used herein, the term "pluripotency" or "multipotency" refers to the ability of a cell to self-renew and differentiate into cells of any of the three germ layers, i.e., endoderm, mesoderm, or ectoderm. "Pluripotent stem cells" or "PSCs" include, for example, ESCs derived from the inner cell mass of a blastocyst or from somatic cell nuclear transfer, and iPSCs derived from non-pluripotent cells.
[0103] As used herein, the terms "embryonic stem", "ES" cells and "ESCs" refer to pluripotent stem cells obtained from an early embryo. In some embodiments, the term does not include stem cells that involve the destruction of a human embryo. That is, ESCs are obtained from an already established ESC line.
[0104] The term "induced pluripotent stem cells" or "iPSCs" refers to a type of pluripotent stem cell artificially prepared from non-pluripotent cells, such as adult somatic cells, partially differentiated cells, or terminally differentiated cells, such as fibroblasts, cells of the hematopoietic system, muscle cells, neurons, epidermal cells, etc., by introducing or contacting the cells with one or more reprogramming factors. Methods for producing iPSCs include, for example, inducing expression of one or more genes, such as, but not limited to, POU5F1 / OCT4 (Gene ID: 5460) in combination with SOX2 (Gene ID: 6657), KLF4 (Gene ID: 9314), c-MYC (Gene ID: 4609), NANOG (Gene ID: 79923), and / or LIN28 / LIN28A (Gene ID: 79727). Reprogramming factors can be delivered by various means, such as viral, non-viral, RNA, DNA, or protein delivery. Alternatively, endogenous genes can be activated, for example using CRISPR tools, to reprogram non-pluripotent cells into PSCs. See, for example, WO2013 / 177133 and WO2022 / 204567.
[0105] Methods for inducing the differentiation of PSCs into cells of various lineages are known in the art.For example, methods for inducing the differentiation of PSCs into dendritic cells are described in Slukvin et al., J Imm. (2006) 176: 2924-32; and Su et al., Clin Cancer Res. (2008) 14 (19): 6207-17; and Tseng et al., Regen Med. (2009) 4 (4): 513-26. Methods for inducing the differentiation of PSCs into hematopoietic progenitor cells, myeloid cells and T lymphocytes are described in Kennedy et al., Cell Rep. (2012) 2: 1722-35.
[0106] Recombinant PSCs can be differentiated into cells suitable for therapy, including cells of the endodermal lineage (e.g., lung, thyroid or pancreatic cells or their precursor cells), cells of the ectodermal lineage (e.g., skin, neural or pigment cells or their precursor cells), and cells of the mesodermal lineage (e.g., cardiac cells, skeletal muscle cells, red blood cells, smooth muscle cells or their precursor cells or precursors).
[0107] In some embodiments, the recombinant PSCs are differentiated into cells of an endodermal lineage (e.g., lung, thyroid or pancreatic cells or precursor cells or precursors thereof), cells of an ectodermal lineage (e.g., skin, neural or pigment cells or precursor cells or precursors thereof), or cells of a mesodermal lineage (e.g., cardiac cells, skeletal muscle cells, red blood cells, smooth muscle cells or precursor cells or precursors thereof).
[0108] In some embodiments, the recombinant PSCs of the present disclosure are differentiated into cardiac cells. In various embodiments, the cardiac cells are cardiac progenitor cells or mature or immature (atrial or ventricular) cardiomyocytes. In other embodiments, the cardiac cells are cardiac endothelial cells or nodal cells.
[0109] In some embodiments, the recombinant PSCs of the present disclosure are differentiated into human immune cells, which may optionally be selected from T cells, T cells expressing a chimeric antigen receptor (CAR) or a recombinant TCR, regulatory T cells, myeloid cells, dendritic cells and / or macrophages / monocytes (e.g., immunosuppressive macrophages), or their progenitors or precursors.
[0110] In some embodiments, the recombinant PSCs of the disclosure are differentiated into oligodendrocyte precursor or progenitor cells or oligodendrocytes. In some embodiments, the recombinant PSCs of the disclosure are differentiated into microglial precursor or progenitor cells or microglial cells.
[0111] In some embodiments, the recombinant PSCs of the disclosure are differentiated into neural cells, such as neural crest cells, astrocytes, dopaminergic neuronal progenitor cells, dopaminergic neurons, midbrain dopaminergic neuronal progenitor cells, midbrain dopaminergic neurons, bona fide midbrain dopamine (DA) neurons, dopaminergic neuronal progenitor cells, floor plate midbrain progenitor cells, floor plate midbrain DA neurons, or precursor cells or precursors thereof.
[0112] In some embodiments, recombinant PSCs of the present disclosure are differentiated into cells of the ocular system, such as photoreceptor cells, photoreceptor progenitor or precursor cells, retinal pigment epithelial (RPE) cells or their precursors or precursors, neural retinal cells or their precursors or precursors, In other embodiments, unedited PSCs are differentiated into cells of the ocular system, which are then engineered using the targeting constructs of the present disclosure.
[0113] In further particular embodiments, the recombinant PSCs of the present disclosure are differentiated into microglial cells or microglial progenitor or precursor cells.
[0114] In further particular embodiments, the recombinant PSCs of the present disclosure are differentiated into cells of a human metabolic system, which may optionally be selected from hepatocytes, cholangiocytes and pancreatic β cells or their precursors or precursors.
[0115] In further particular embodiments, the recombinant PSCs of the present disclosure are differentiated into intestinal progenitor or precursor cells or intestinal cells.
[0116] B. Differentiated cells In yet other embodiments, the engineered cells are differentiated cells (e.g., partially or terminally differentiated cells). Partially differentiated cells can be, for example, tissue-specific progenitors or stem cells, such as hematopoietic progenitors or stem cells, skeletal muscle progenitors or stem cells, cardiac progenitors or stem cells, neural progenitors or stem cells, and mesenchymal stem cells.
[0117] Exemplary differentiated cell types that can be engineered in one or more of their STAPLRs include cells of the endodermal lineage (e.g., lung, thyroid or pancreatic cells or their precursors), cells of the ectodermal lineage (e.g., skin, neural or pigment cells or their precursors or precursors) and cells of the mesodermal lineage (e.g., cardiac cells, skeletal muscle cells, red blood cells, smooth muscle cells or their precursors or precursors). Alternatively, PSCs can be differentiated into cells of these lineages and then engineered with the targeting constructs of the present disclosure.
[0118] In some embodiments, cardiac cells are engineered. In some embodiments, the cardiac cells are cardiac progenitor cells or mature or immature (atrial or ventricular) cardiomyocytes. In other embodiments, the cardiac cells are cardiac endothelial cells or nodal cells.
[0119] In some embodiments, human immune cells are engineered. Human immune cells may be optionally selected from T cells (e.g., CD4+ T cells, CD8+ T cells, or Treg cells), T cells expressing a chimeric antigen receptor (CAR) or a recombinant TCR, regulatory T cells, myeloid cells, dendritic cells, and / or macrophages (e.g., immunosuppressive macrophages), or their progenitor or precursor cells, such as hematopoietic stem or progenitor cells.
[0120] In some embodiments, oligodendrocyte progenitor or precursor cells or oligodendrocytes are engineered.
[0121] In some embodiments, neural cells are engineered. In various embodiments, the neural cells are neural crest cells, astrocytes, dopaminergic neuronal progenitor cells, dopaminergic neuronal cells, midbrain dopaminergic neuronal progenitor cells, midbrain dopaminergic neurons, bona fide midbrain dopamine (DA) neurons, dopaminergic neuronal progenitor cells, floor plate midbrain progenitor cells, floor plate midbrain DA neurons, or progenitor cells or precursors thereof.
[0122] In some embodiments, cells of the ocular system are engineered. In various embodiments, the cells of the ocular system are photoreceptor cells, photoreceptor progenitor or precursor cells, retinal pigment epithelial cells or precursor cells or precursor cells thereof, neural retinal cells or precursor cells or precursor cells thereof.
[0123] In more particular embodiments, microglial cells or microglial progenitor or precursor cells are manipulated.
[0124] In more particular embodiments, cells of the human metabolic system are engineered. In various embodiments, the cells of the human metabolic system may be optionally selected from hepatocytes, bile duct cells and pancreatic beta cells or precursor cells or precursors thereof.
[0125] In further particular embodiments, intestinal progenitor or precursor cells or intestinal cells are engineered.
[0126] Additional cell types that can be engineered in the present invention to incorporate exogenous sequences into STAPLR include, but are not limited to, fibroblasts, adipocytes, muscle cells (e.g., skeletal muscle cells or smooth muscle cells), bone cells, bone marrow cells, and bone marrow progenitor cells (e.g., primitive bone marrow progenitor cells).
[0127] The cells can be from an established cell line or they can be primary cells, where "primary cells", "primary cell lines" and "primary cultures" are used interchangeably herein to refer to cells and cell cultures derived from a subject (e.g., a human) and grown in vitro or ex vivo for a limited number of culture passages. For example, primary cultures include cultures that have been passaged 0, 1, 2, 4, 5, 10 or 15 times, but may have been passaged insufficiently to survive a crisis stage. Primary cell lines may be maintained in vitro or ex vivo for less than 10 passages. In some embodiments, the cells are autologous in the context of cell therapy. In some embodiments, the cells are allogeneic in the context of cell therapy.
[0128] Primary cells may be obtained from an individual by any suitable method, for example, white blood cells may be suitably obtained by apheresis, leukapheresis, density gradient separation, etc., while cells from tissues, such as skin, muscle, bone marrow, spleen, liver, pancreas, lung, intestine, stomach, etc., may be best obtained by biopsy.
[0129] Any of the above differentiated cell types can be differentiated from the PSC prior to engineering it.
[0130] The present disclosure provides pharmaceutical compositions comprising the engineered cells herein and a pharma- ceutically acceptable carrier.
[0131] IV. How to identify STAPLR The present disclosure also provides methods for identifying STAPLR as a site for safe genomic integration in mammalian cells (e.g., human cells). In these methods, the first step is to select a set of cell types for single-cell RNA sequencing ("scRNAseq"). Examples of cell types include, but are not limited to, those mentioned herein, including PSCs (e.g., iPSCs), cells of the immune system (e.g., T cells, NK cells, dendritic cells, macrophages / monocytes or their hematopoietic progenitors), cells of the cardiovascular system (e.g., ventricular cardiomyocytes, nodal cells or cardiac progenitor cells), cells of the metabolic system (e.g., hepatocytes and pancreatic beta cells), cells of the central nervous system (e.g., sensory neurons, motor neurons, interneurons, microglial cells, oligodendrocytes or their progenitors), muscle cells (e.g., skeletal muscle cells and smooth muscle cells), adipocytes and cells of the ocular system (e.g., retinal pigment epithelial cells and photoreceptor cells).
[0132] The second step is to perform an scRNAseq assay, where sequencing analysis assigns a unique transcriptome containing transcribed genes to each cell that meets the quality criteria. Transcriptomes are filtered to exclude those with high sparsity or missingness and those likely to originate from multiple cells to meet the quality criteria.
[0133] Each gene is then assigned a prevalence score. Prevalence scores start at "1". The prevalence score represents the percentage of cells that contain at least one transcript of a given gene based on the scRNAseq database of collected datasets. In some embodiments, the scRNAseq datasets are obtained from PSCs, dopaminergic neurons and / or their precursors (e.g., in various selected differentiation states), microglia and / or their precursors (e.g., in various selected differentiation states), cardiomyocytes and / or their precursors (e.g., in various selected differentiation states), oligodendrocyte cells and / or their precursors (e.g., in various selected differentiation states), or macrophages and / or their precursor cells (e.g., in various selected differentiation states).
[0134] After assigning a prevalence score, the location of each gene in the mammalian (eg, human) genome is determined.
[0135] The next step in identifying STAPLR in mammalian cell genome is to identify adjacent non-overlapping genes. "Non-overlapping genes" means that the genes are at least 50 base pairs, at least 75 base pairs, at least 100 base pairs, at least 200 base pairs, at least 300 base pairs, at least 400 base pairs, at least 500 base pairs, at least 1000 base pairs, at least 1500 base pairs, at least 2000 base pairs, at least 2500 base pairs, at least 3000 base pairs, 3500 base pairs, at least 5000 base pairs, at least 10000 base pairs, at least 15000 base pairs or at least 20000 base pairs away from each other on either strand. The transcripts used to calculate the genetic distance to identify non-overlapping genes can be identified by any genome database, such as NCBI's RefSeq database and GENCODE database.
[0136] In some cases, different genome databases contain non-consensus gene boundary annotations, which may lead to different calculated genetic distances and opposite conclusions about whether two genes overlap.In such cases, two genes are considered to be non-overlapping if they are determined to be non-overlapping by using at least one genome database.For example, MLF2 is adjacent to its neighboring gene PTMS downstream.As annotated in the NCBI RefSeq database, these genes are non-overlapping, and the intergenic distance is about 13 kb.However, the GENCODE V38 database reports one MLF2 transcript with a transcription start site located within the first intron of PTMS encoded on the opposite strand.In this case, the RefSeq annotation is taken into account, and the GENCODE annotation is not taken into account, and this gene pair is classified as non-overlapping.
[0137] Once two or more genes are determined to be non-overlapping, a contiguous score is determined for the pair of non-overlapping genes or for a region containing three or more non-overlapping genes. The contiguous score is the product of the individual prevalence scores and reflects the probability that both genes are transcriptionally active in the aggregated scRNAseq dataset. The contiguous score is essentially a ranking of the proximity of transcriptionally active genes.
[0138] The adjacency scores are then sorted to obtain a ranking of pairs of non-overlapping genes or a ranking of regions containing three or more genes. Once the adjacency scores are ranked, the pair of genes or region containing three or more genes with the best adjacency score is selected, and the intergenic region between the genes of the selected pair or region is identified as a potential STAPLR.
[0139] STAPLR can be targeted for safe genetic integration. Intergenic regions with high ranking adjacent scores are then annotated to design homology arms for site-specific integration. In general, sequences to be avoided as integration sites include promoter regions, enhancer regions, CpG islands, epigenetic marks (e.g., H3K4Me1, H3K4Me3 and H3K27Ac), DNase I hypersensitive peaks, conserved regions and repeat regions. UCSC Genome Browser can be used with the following gene annotation tracks (not limited to): GENCODE V32, RefSeq Genes, GTEx RNA-seq, EPDnew Promoters, ENCODE (transcription, H3K4Me1, H3K4Me3, H3K27Ac and DNase clusters), GeneHancer, CpG Islands, Conservation 100 vertebrates and RepeatMasker.
[0140] When selecting targetable intergenic subregions, known promoter and enhancer regions should be avoided. Furthermore, conserved regions, repetitive regions, epigenetic marks and DNase hypersensitive regions are features that should be minimized when selecting targetable regions. In some embodiments, the targetable intergenic subregion comprises a sequence of a CRISPR endonuclease protospacer adjacent motif (PAM) site. A PAM site is a 2-6 base pair DNA sequence immediately following the DNA sequence targeted by Cas (e.g., Cas9 or Cpf1) endonuclease. To perform the function of the tracrRNA-crRNA complex in the CRISPR / Cas gene editing system, a short oligonucleotide known as a guide RNA (gRNA) is synthesized. The gRNA recognizes gene sequences that have a PAM sequence at the 5' or 3' end. The Cas protein can thereby recognize different PAMs. For example, Cas9 from Streptococcus pyogenes recognizes 5'-NGG-3' ('N': any nucleobase). Cas9 from Staphylococcus aureus recognizes 5'-NNGRR(N)-3'. Cas9 from Neisseria meningitidis recognizes 5'-NNNNGATT-3'. Cas9 from Campylobacter jejuni recognizes 5'-NNNNRYAC-3' ('Y': pyrimidine). Cas9 from Streptococcus thermophilus recognizes 5'-NNAGAAW-3' ('W': A or T). Cpf1 (Cas12a) from Lachnospiraceae bacteria and Acidaminococcus sp. recognizes 5'-TTTV-3' ('V': G, A or C). Cas12b from Alicyclobacillus acidiphilus recognizes 5'-TTN-3'.Cas12b v4 from Bacillus hisashii recognizes 5'-ATTN-3', 5'-TTTN-3' and 5'-GTTN-3'.
[0141] Finally, confirmation that the identified intergenic region can safely support exogenous gene payload can be performed by inserting a transgene into the intergenic region at a targeted location using a gene editing system. The gene editing system can be, for example, a CRISPR system (e.g., using the CRISPR endonuclease disclosed above), a Cre / Lox system, a FLP-FRT system, a TALEN system, a ZFN system, a system that uses a homing endonuclease, a system that produces homologous recombination, or a system that uses a non-nuclease-dependent viral vector (e.g., a retrovirus, an AAV, or a lentivirus vector). Constitutive, inducible, tissue-specific, or lineage-specific promoters can be used to induce the expression of the inserted transgene.
[0142] In some embodiments, the target intergenic region is at least 30, 40, 50, 75 or 100 base pairs long. In some embodiments, the intergenic region does not include promoter or enhancer regions. It may be better for the intergenic region to not include conserved regions, repetitive regions, epigenetic marks and / or DNase hypersensitive regions, but in some embodiments, the intergenic region may actually include a minimal amount of conserved regions, repetitive regions, epigenetic marks and / or enzyme hypersensitive regions. For example, in some embodiments, the intergenic region does not include CpG islands, H3K4Me1 epigenetic marks, H3K4Me3 epigenetic marks, H3K27Ac epigenetic marks, DNase I hypersensitive regions, conserved regions or repetitive regions. However, in some embodiments, the intergenic region may comprise CpG islands, H3K4Me1 epigenetic marks, H3K4Me3 epigenetic marks, H3K27Ac epigenetic marks, DNase I hypersensitive regions, conserved regions or repetitive regions. The amount of conserved regions, repetitive regions, epigenetic marks and / or DNase I hypersensitive regions that can be tolerated depends on various factors. These factors include, for example, the size of the intergenic region; the size of conserved regions, repetitive regions and / or hypersensitive regions or epigenetic marks; the presence of gRNA binding sites; or the difficulty of synthesizing 5' and 3' homology arms for targeting.
[0143] After genomic integration, the transcription level of the integrated transgene is measured, and if the integrated transgene exhibits sustained transcription (or if an inducible promoter controlling the transgene exhibits sustained transcription when induced), it is confirmed that the intergenic region between the selected pair or within the selected region is STAPLR.
[0144] Unless otherwise indicated herein, scientific and technical terms used in connection with the present invention shall have the meanings commonly understood by those of ordinary skill in the art. Exemplary methods and materials are described below, although methods and materials similar or equivalent to those described herein may also be used in the practice or testing of the present invention. In the event of a conflict, the present specification, including definitions, will control. Furthermore, unless otherwise required by context, singular terms shall include the plural and plural terms shall include the singular. Throughout this specification and the embodiments, the words "having" and "including" or variations such as "having", "having", "including" or "including" are understood to mean the inclusion of a stated integer or group of integers, but not the exclusion of any other integer or group of integers. All publications and other references cited herein are incorporated herein by reference in their entirety. Although numerous documents are cited herein, this citation is not an admission that any of these documents constitutes part of the general knowledge in the art. As used herein, the term "approximately" or "about" as applied to one or more values of interest refers to a value similar to the stated reference value. In some embodiments, the term refers to a range of values that is included within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less in either direction (greater or less) of the indicated reference value, unless otherwise indicated or clear from the context.
[0145] In this disclosure, a back-reference in a dependent claim is intended as a shorthand for a direct and unambiguous disclosure of any combination of claims that is indicated by the back-reference. Moreover, headings herein are provided for ease of organization and are not intended to limit the scope of the claimed invention in any way.
[0146] In order that this invention may be better understood, the following examples are set forth. These examples are for illustrative purposes only and are not to be construed as in any way limiting the scope of the invention.
[0147] Working Example In order that this invention may be better understood, the following examples are set forth. These examples are for illustrative purposes only and are not to be construed as in any way limiting the scope of the invention.
[0148] Example 1: Design for STAPLR targeting gRNA Selection CRISPOR was used to identify Cas9-based gRNAs that target near the midpoint region where the homology arms of the STAPLR construct flank the desired site of transgene integration. gRNAs with complete off-targets in the genome were excluded. The use of gRNAs with up to 3bp off-target mismatches was minimized. Three gRNAs for each STAPLR site were selected, as shown in Table 2.
[0149] [Table 3]
[0150] A list of additional Cas9 and Cpf1 based gRNAs for STAPLR targeting is shown in Table 3.
[0151] [Table 4] TIFF2025514159000007.tif174162
[0152] For each STAPLR site, human iPSCs were nucleofected with each gRNA complexed with Cas9 nuclease in the form of ribonucleoprotein (RNP). After 3 days, nucleofected cells were harvested, genomic DNA was extracted, and PCR amplification of the genomic regions flanking the intended cleavage site was performed. Purified PCR products were sequenced, and sequencing data were analyzed for overall cleavage efficiency by Synthego's ICE analysis tool (available on the Synthego website) (Figure 1). gRNAs were considered efficient if they showed more than 50% indel editing.
[0153] The data show that there was at least one efficient gRNA (greater than 50% indel editing) per STAPLR site. The gRNA that showed the greatest overall cleavage efficiency was selected for use in future experiments to integrate transgenes into the STAPLR site.
[0154] Design of STAPLR homology arms A list of gene neighbors consisting of genes that are highly expressed together was created. This list was filtered to exclude gene pairs that contain at least one gene that is a known tumor suppressor or oncogene. Initially, gene pairs with an intergenic distance of less than 5 kb between the genes of the gene pair were excluded. However, gene pairs with an intergenic distance of only about 100 bases between adjacent genes can also be annotated and tested. Regions containing promoter regions, enhancer regions, CpG islands, and epigenetic markers were avoided in the design. Small regions that avoid regulatory elements and can be synthesized in the donor plasmid were classified as potential homology arm regions and used as the basis for gRNA search (Table 4).
[0155] [Table 5]
[0156] After selecting the gRNA with the predicted high efficiency, the homology arm sequence is completed so that the selected gRNA is located in the center of the 800bp left homology arm and the 800bp right homology arm adjacent to the intended site of transgene integration.Table 5 shows the intergenic distance in base pair length units between two gene neighbors for each exemplary STAPLR site, together with the coordinates for each set of STAPLR left and right homology arms based on hg38 human reference genome.The genetic distance is calculated using NCBI's RefSeq database.
[0157] [Table 6]
[0158] The sequences of the left and right homology arms of the targeting construct based on the hg38 human reference genome are shown in the table below.
[0159] [Table 7] TIFF2025514159000011.tif248162 TIFF2025514159000012.tif246162 TIFF2025514159000013.tif180161
[0160] Example 2: Testing inducibility and transgene expression in STAPLR in pooled populations of targeted iPSCs To test the robustness of inducibility and transgene expression at each of the four annotated STAPLR sites [PRDX1-AKR1A1 (site 1), ACTB-FSCN1, RPL34-OSTC, AKIRIN1-NDUFS5], we used the Clontech / TakaRa binary doxycycline-inducible rtTA / TRE system Tet-On 3G [described in U.S. Patent No. 9,127,283, which is incorporated herein by reference in its entirety]. For the components, we expressed Tet-On 3GrtTA (reverse tetracycline transactivator) biallelically from the GAPDH locus by our "sustained transgene expression locus" (STEL) approach (FIG. 2), which is described in WO2021 / 072329, which is incorporated herein by reference in its entirety.
[0161] To test the inducibility of the transgene from the STAPLR site, the expression of eGFP cargo was tested using the TRE3G promoter. A Kozak sequence was included to allow translation initiation, and an SV40 polyA sequence was added to allow transcription termination. In the presence of doxycycline, the rtTA protein binds to and activates the tetracycline response element (TRE) minimal promoter (Figure 3). For each STAPLR site, a parental iPSC line carrying biallelic rtTA integration in GAPDH (GAPDH::rtTA iPSC) was nucleofected with the selected highly efficient RNP and the corresponding STAPLR targeting construct (STAPLR left homology arm-TRE3G promoter-eGFP-SV40-STAPLR right homology arm). Pools of cells administered both STAPLR RNP and STAPLR targeting constructs were fed with medium containing 2 μg / ml doxycycline from day 1 post-nucleofection (FIG. 4) to day 7 post-nucleofection (FIG. 5) to induce GFP expression. Parental rtTA iPSC lines were also fed with 2 μg / ml doxycycline medium as a control. GFP expression was monitored over a period of one week by fluorescence microscopy. An increase in GFP intensity was observed when cells were treated with doxycycline for longer periods. Preliminary testing of this rtTA / TRE-based transgene expression system in STAPLR shows robust inducibility and expression of GFP in pooled populations of STAPLR site-targeted iPSCs.
[0162] Example 3: Testing inducibility and transgene expression in STAPLR in clonal populations of targeted iPSCs Parental GAPDH::rtTA iPSCs were nucleofected with RNP and STAPLR targeting constructs at each of the four STAPLR sites, and then each pooled population of STAPLR-targeted iPSCs was plated at clonal density. Individual clones were picked and screened by PCR across the junction of the left and right homology arms to confirm correct integration of TRE3G-eGFP-SV40 at each of the four STAPLR sites. Targeted iPSC clones were expanded and treated from 0 to 68 hours with medium containing doxycycline ranging from 0.1 μg / ml to 5 μg / ml. Over this time course of GFP induction at a time course of 0, 3, 8, 24, 48 and 68 hours, cells were collected and flow cytometry analysis was performed (Figure 6). The results show that maximum GFP induction from all four STAPLR sites can be seen from administration of 0.1 μg / ml doxycycline and after 48 hours of doxycycline. STAPLR sites vary in their maximum expression levels of GFP, with the PRDX1-AKR1A1 site showing the highest expression of GFP in doxycycline-induced iPSCs. The wild-type unedited iPSC control line and one clone-derived line from each STAPLR-targeted site were then treated with medium containing 2 μg / ml doxycycline for 72 hours (FIG. 7). The AKIRIN1-NDUFS5 STAPLR line showed a slightly delayed GFP induction, so treatment with medium containing 2 μg / ml doxycycline was extended to 6 days (FIG. 7). The results show that all four treated STAPLR-targeted iPSC lines can induce high levels of GFP expression, with the PRDX1-AKR1A1 site again showing the highest expression of GFP in doxycycline-induced iPSCs, while the wild-type unedited doxycycline-treated iPSC control line did not express GFP. In all cases, cells that did not receive doxycycline treatment did not express GFP.
[0163] Example 4: Testing inducibility and transgene expression in STAPLR in iPSC-derived myeloid progenitor cells To demonstrate that transgene integration in STAPLR maintains persistent transgene expression in differentiated iPSCs, clonally derived STAPLR iPSC lines were differentiated into myeloid progenitor cells (Douvaras et al., Stem Cell Reports (2017) 8(6):1516-24). 2 μg / ml doxycycline was added to each STAPLR-targeted clonal line on day 12 of differentiation and doxycycline was supplemented daily for 3 days. On day 15 of differentiation, adherent myeloid progenitor cells were harvested for flow cytometric analysis of GFP induction. Three of the four TRE-eGFP-SV40 STAPLR lines (PRDX1-AKR1A1, ACTB-FSCN1, RPL34-OSTC) showed efficient GFP induction in heterologous adherent myeloid progenitor cells compared to differentiated cells to which doxycycline was not added (Figure 8). Wild-type unedited iPSC control lines differentiated using the same protocol and similarly treated with doxycycline showed no GFP induction. One of the TRE-eGFP-SV40 STAPLR lines (AKIRIN1-NDUFS5) showed delayed GFP induction in fluorescent microscopy. This cell line was supplemented with doxycycline for an additional 3 days, and adherent bone marrow progenitor cells were harvested on day 18 of differentiation for flow cytometry analysis. Figure 8 shows the bimodal GFP induction observed from bone marrow progenitor cells harvested on day 18 of differentiation. In all cases, cells that did not receive doxycycline treatment did not express GFP.
[0164] STAPLR-targeted cell lines were further differentiated for more than 30 days until non-adherent bone marrow progenitor cells were harvestable in suspension culture. 2 μg / ml doxycycline was added for 6 days and non-adherent bone marrow progenitor cells were harvested for flow cytometric analysis of GFP induction. All four TRE-eGFP-SV40 STAPLR lines cultured for more than 30 days showed efficient differentiation into triple-positive bone marrow progenitor cells defined by greater than 80% co-expression of cell surface markers CD45, CD14 and CX3CR1 (Figure 9). STAPLR lines treated with doxycycline also showed efficient GFP induction in heterogeneous non-adherent bone marrow progenitor cells compared to wild-type unedited control lines treated with doxycycline, where some variability in maximum GFP expression levels was observed (Figure 10). The data indicate that transgene integration at all four STAPLR sites allowed sustained expression of the transgene under external promoter control during and after differentiation into myeloid progenitor cells.
[0165] Example 5: Derivation of a human induced pluripotent stem cell line with inducible expression of CD19t-IL12 from the PRDX1-AKR1A1 STAPLR site A parental iPSC line with biallelic rtTA integration at GAPDH (GAPDH::rtTA iPSC) was transfected with a STAPLR targeting construct containing a doxycycline-inducible promoter (TRE3G)-driven CD19t-IL12 cassette flanked by PRDX1-AKR1A1 left and right homology arms and a selected highly efficient RNP for the PRDX1-AKR1A1 STAPLR site (site 1). In this case, CD19t was included as a biologically non-functional cargo, which served as an epitope marker for alternative detection of IL-12 transgene integration by flow cytometry. For targeting at the PRDX1-AKR1A1 STAPLR site, two different gRNAs and their corresponding nucleases were used. To generate clonal lines, a Cpf1-based guide RNA with the sequence 5'-GAGACTGGTTCTTGCAGCACT-3' (SEQ ID NO: 83) or a Cas9-based guide RNA with the sequence 5'-CTTGCAGCACTGCCTAGGCT-3' (SEQ ID NO: 71) was selected. GAPDH::rtTA constitutively expresses reverse tetracycline transactivator (rtTA) from the GAPDH locus. In the presence of doxycycline, rtTA binds to the TRE3G promoter and induces the expression of CD19t and IL-12 driven by the TRE3G promoter (Figure 11).
[0166] Single cell suspensions of GAPDH::rtTA iPSCs were prepared for transfection with Cpf1 or Cas9 gRNA RNP complexes and PRDX1-AKR1A1-targeted pTRE3G-CD19t-IL-12 DNA donor templates. Two days after transfection, cells were treated with doxycycline (2 μg / mL) for 48 hours to induce CD19t-IL12 expression, which was analyzed using live cell imaging of AF488-conjugated anti-CD19t antibody staining (FIG. 12, panels A and B). Cells were then dissociated and plated at single cell clonal density. Four days after clonal density plating, growing colonies were treated with 2 μg / mL doxycycline for 48 hours to induce CD19t-IL-12 expression. After 48 hours of doxycycline treatment, colonies were analyzed with live cell imaging using AF488-conjugated antibodies against CD19t. CD19t positive colonies were identified (Figure 12, panels A and B; labeled "clonal density").
[0167] The data show that integration of the CD19t-IL-12 expression cassette at the PRDX1-AKR1A1 STAPLR site allowed sustained expression of the transgene under external promoter control in both pooled and clonal populations of STAPLR-targeted iPSCs following doxycycline treatment.
[0168] Example 6: Induction of reporter transgene expression at various sites within the STAPLR intergenic region in targeted iPSCs To test the robustness of inducibility and transgene expression at two alternative sites (PRDX1-AKR1A1 site 2 and site 3) within the PRDX1-AKR1A1 intergenic region, we again used the binary doxycycline-inducible rtTA / TRE system. To test the expression of EGFP cargo, the TRE3G promoter was used. Following the design of the original PRDX1-AKR1A1 targeting construct, a Kozak sequence was included to allow translation initiation, and an SV40 polyA sequence was added to allow translation termination. In the presence of doxycycline, the rtTA protein binds to and activates the TRE minimal promoter. A parental iPSC line with biallelic rtTA integration in GAPDH (GAPDH::rtTA iPSC) was nucleofected with the selected highly efficient RNPs and the corresponding PRDX1-AKR1A1 targeting constructs (for either site 2 or site 3). Three different gRNAs were tested for PRDX1-AKR1A1 site 2 (SEQ ID NO: 87-89) and three different gRNAs were tested for PRDX1-AKR1A1 site 3 (SEQ ID NO: 90-92). Pools of cells receiving both PRDX1-AKR1A1 site 2 or site 3 RNP and targeting constructs were fed with medium containing 2 μg / ml doxycycline from day 2 (site 2; FIG. 13) or day 1 (site 3; FIG. 14) after nucleofection until day 7 after nucleofection (FIGS. 15 and 16). GFP expression was monitored over 7 days by fluorescence microscopy or flow cytometry. GFP expression was induced from both PRDX1-AKR1A1 site 2 and PRDX1-AKR1A1 site 3. Although all three gRNAs tested for each site showed differences in construct targeting efficiency (as seen in the flow cytometry histograms as peaks of different sizes), all were able to induce GFP expression with similarly high intensity (similar log expression levels) following the addition of doxycycline. ∧ The peak observed at 6 represents edited cells expressing high levels of GFP, whereas at approximately 10 ∧The peak observed at 4 represents transient GFP expressed from a non-integrated targeting construct. The data show that multiple sites within the PRDX1-AKR1A1 intergenic region allow robust inducibility and expression of GFP in pooled populations of STAPLR site-targeted iPSCs.
Claims
1. A genetically modified mammalian cell comprising an exogenous nucleotide sequence incorporated into a sustained transcriptional activity payload region (STAPLR) in the cell's genome, wherein the STAPLR is The intergenetic region between the RPL34 gene and the OSTC gene, The intergenetic region between the ACTB gene and the FSCN1 gene, The intergenetic region between the AKIRIN1 gene and the NDUFS5 gene, The intergenetic region between the PRDX1 gene and the AKR1A1 gene, The intergenetic region between the PTGES3 gene and the NACA gene, The intergenetic region between the MLF2 gene and the PTMS gene, The intergenetic region between the RAB13 gene and the RPS27 gene, The intergenetic region between the JTB gene and the RAB13 gene, The intergenetic region between the AKR1A1 gene and the NASP gene, The intergenetic region between the NDUFS5 gene and the MACF1 gene, The intergenetic region between the SRSF9 gene and the DYNLL1 gene, The intergenetic region between the MYL6B gene and the MYL6 gene, The intergenetic region between the GPX1 gene and the RHOA gene, The intergenetic region between the HNRNPA2B1 gene and the CBX3 gene, The intergenetic region between the ROMO gene and the RBM39 gene, The intergenetic region between the PA2G4 gene and the RPL41 gene, and Intergenetic region between the NDUFB10 gene and the RPS2 gene Genetically modified mammalian cells selected from a group consisting of the following.
2. A method for modifying mammalian cells, comprising incorporating an exogenous nucleotide sequence into a sustained transcriptional activity payload region (STAPLR) in the cell's genome, wherein the STAPLR is The intergenetic region between the RPL34 gene and the OSTC gene, The intergenetic region between the ACTB gene and the FSCN1 gene, The intergenetic region between the AKIRIN1 gene and the NDUFS5 gene, The intergenetic region between the PRDX1 gene and the AKR1A1 gene, The intergenetic region between the PTGES3 gene and the NACA gene, The intergenetic region between the MLF2 gene and the PTMS gene, The intergenetic region between the RAB13 gene and the RPS27 gene, The intergenetic region between the JTB gene and the RAB13 gene, The intergenetic region between the AKR1A1 gene and the NASP gene, The intergenetic region between the NDUFS5 gene and the MACF1 gene, The intergenetic region between the SRSF9 gene and the DYNLL1 gene, The intergenetic region between the MYL6B gene and the MYL6 gene, The intergenetic region between the GPX1 gene and the RHOA gene, The intergenetic region between the HNRNPA2B1 gene and the CBX3 gene, The intergenetic region between the ROMO gene and the RBM39 gene, The intergenetic region between the PA2G4 gene and the RPL41 gene, and Intergenetic region between the NDUFB10 gene and the RPS2 gene A method selected from the group consisting of the following.
3. The method according to claim 2, wherein the integration step is performed by using a CRISPR / Cas system, Cre / Lox system, FLP-FRT system, TALEN system, ZFN system, homing endonuclease, random integration, homologous recombination, transposase, or nuclease-independent viral vector [which is optionally selected from retroviral vectors, adeno-associated virus (AAV) vectors and lentiviral vectors].
4. The integration process is carried out using a CRISPR / Cas system containing guide RNA, where, STAPLR is the intergenetic region between the RPL34 gene and the OSTC gene, and the gRNA is selected from sequence numbers 25-32. STAPLR is the intergenetic region between the ACTB gene and the FSCN1 gene, and the gRNA was selected from sequence numbers 33-54. STAPLR is the intergenetic region between the AKIRIN1 gene and the NDUFS5 gene, and the gRNA is selected from SEQ ID NOs. 55-70, or STAPLR is the intergenetic region between the PRDX1 gene and the AKR1A1 gene, and the gRNA is selected from sequence numbers 71-92. The method according to claim 2.
5. The method according to claim 3, wherein the CRISPR / Cas system comprises a type I, type II, type III, type IV, or type V gRNA-dependent nuclease or a variant thereof.
6. The CRISPR / Cas system includes Cas9, Cpf1, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas12, Cas13, Cas100, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1 The method according to claim 3, comprising a gRNA-dependent nuclease selected from the group consisting of Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, CasX, CasY, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, CasPhi, MAD7, and Csf4.
7. A DNA molecule comprising a target nucleotide sequence adjacent to a 5' homologous region (HR) and a 3'HR, wherein the 5'HR and 3'HR are at least 95% homologous to a first genomic region (GR) and a second GR in a sustained transcriptional activity payload region (STAPLR) in the genome of a mammalian cell, and the STAPLR is The intergenetic region between the RPL34 gene and the OSTC gene, The intergenetic region between the ACTB gene and the FSCN1 gene, The intergenetic region between the AKIRIN1 gene and the NDUFS5 gene, The intergenetic region between the PRDX1 gene and the AKR1A1 gene, The intergenetic region between the PTGES3 gene and the NACA gene, The intergenetic region between the MLF2 gene and the PTMS gene, The intergenetic region between the RAB13 gene and the RPS27 gene, The intergenetic region between the JTB gene and the RAB13 gene, The intergenetic region between the AKR1A1 gene and the NASP gene, The intergenetic region between the NDUFS5 gene and the MACF1 gene, The intergenetic region between the SRSF9 gene and the DYNLL1 gene, The intergenetic region between the MYL6B gene and the MYL6 gene, The intergenetic region between the GPX1 gene and the RHOA gene, The intergenetic region between the HNRNPA2B1 gene and the CBX3 gene, The intergenetic region between the ROMO gene and the RBM39 gene, The intergenetic region between the PA2G4 gene and the RPL41 gene, and Intergenetic region between the NDUFB10 gene and the RPS2 gene A DNA molecule selected from the group consisting of the following.
8. The DNA molecule according to claim 7, wherein each of the 5'HR and 3'HR is independently at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 550, at least about 600, at least about 650, at least about 700, at least about 750, at least about 800, at least about 850, at least about 900, at least about 950, at least about 1000, at least about 1100, at least about 1200, at least about 1300, at least about 1400, at least about 1500, at least about 1600, at least about 1700, at least about 1800, at least about 1900 or at least about 2000 base pairs long, or between 50 and 1500 base pairs long.
9. 5'HR and 3'HR are, respectively, Sequence IDs 17 and 18, Sequence IDs 19 and 20, Sequence IDs 21 and 22, Sequence IDs 23 and 24, Sequence IDs 93 and 94, or Sequence IDs 95 and 96 The DNA molecule according to claim 7, which is at least 95% homologous to the DNA molecule.
10. The intergenetic region between the RPL34 gene and the OSTC gene contains a nucleotide sequence that is at least 95% identical to SEQ ID NO: 1; The intergenetic region between the ACTB gene and the FSCN1 gene contains a nucleotide sequence that is at least 95% identical to SEQ ID NO: 2; The intergenetic region between the AKIRIN1 gene and the NDUFS5 gene contains a nucleotide sequence that is at least 95% identical to SEQ ID NO: 3; The intergenetic region between the PRDX1 gene and the AKR1A1 gene contains a nucleotide sequence that is at least 95% identical to SEQ ID NO: 4; The intergenetic region between the PTGES3 gene and the NACA gene contains a nucleotide sequence that is at least 95% identical to SEQ ID NO: 5; The intergenetic region between the MLF2 gene and the PTMS gene contains a nucleotide sequence that is at least 95% identical to SEQ ID NO: 6; The intergenetic region between the RAB13 gene and the RPS27 gene contains a nucleotide sequence that is at least 95% identical to that of SEQ ID NO: 7; The intergenetic region between the JTB gene and the RAB13 gene contains a nucleotide sequence that is at least 95% identical to SEQ ID NO: 8; The intergenetic region between the AKR1A1 gene and the NASP gene contains a nucleotide sequence that is at least 95% identical to sequence number 9; The intergenetic region between the NDUFS5 gene and the MACF1 gene contains a nucleotide sequence that is at least 95% identical to sequence number 10; The intergenetic region between the SRSF9 gene and the DYNLL1 gene contains a nucleotide sequence that is at least 95% identical to sequence number 11; The intergenetic region between the MYL6B gene and the MYL6 gene contains a nucleotide sequence that is at least 95% identical to sequence number 12; The intergenetic region between the GPX1 gene and the RHOA gene contains a nucleotide sequence that is at least 95% identical to sequence number 13; The intergenetic region between the HNRNPA2B1 gene and the CBX3 gene contains a nucleotide sequence that is at least 95% identical to sequence number 14; The intergenetic region between the ROMO gene and the RBM39 gene contains a nucleotide sequence that is at least 95% identical to sequence number 15; The intergenetic region between the PA2G4 gene and the RPL41 gene contains a nucleotide sequence that is at least 95% identical to SEQ ID NO: 16; and / or The intergenetic region between the NDUFB10 gene and the RPS2 gene contains a nucleotide sequence that is at least 95% identical to sequence number 97. A cell, method, or DNA molecule according to any one of claims 1 to 9.
11. A cell, method, or DNA molecule according to any one of claims 1 to 9, wherein the exogenous nucleotide sequence or the nucleotide sequence of interest comprises a transgene, wherein the transgene may optionally comprise a constitutive or inducible promoter.
12. Transgenes Therapeutic proteins, and, if desired, proteins, cytokines, or recombinant antigen receptors that are deficient or lacking in hereditary diseases; Cell markers; or Proteins that regulate the differentiation state or activity of cells It is coded as, If desired, the transgene here may code SOX10, IL-10, IL-12, CD19t, or ThPOK. The cell, method, or DNA molecule according to claim 11.
13. The cell, method, or DNA molecule according to any one of claims 1 to 9, wherein the cell is a human cell.
14. The cell, method, or DNA molecule according to any one of claims 1 to 9, wherein the cell is a pluripotent stem cell (PSC), and optionally an induced PSC (iPSC).
15. Cells, a) Immune system cells, optionally T cells, natural killer cells, dendritic cells, macrophages / monocytes or their hematopoietic progenitor cells; b) Cardiovascular cells, preferably ventricular cardiomyocytes, nodular cells, or cardiac progenitor cells; c) Metabolic cells, optionally hepatocytes, pancreatic beta cells, or bile duct cells; d) Cells of the central nervous system, preferably sensory neurons, motor neurons, interneurons, microglia, oligodendrocytes, or their precursor cells; e) Muscle cells, and optionally skeletal muscle cells or smooth muscle cells; f) Fat cells; or g) Cells of the ocular system, preferably retinal pigment epithelial cells, photoreceptor cells, or photoreceptor progenitor cells. A cell, method, or DNA molecule according to any one of claims 1 to 9.