Controllable transcription
Patent Information
- Application Number
- PCT/GB2026/050515
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2026-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure GB2026050515_01102026_PF_FP_ABST
Abstract
Description
[0001] BIT-C-P3829PCT
[0002] 1
[0003] CONTROLLABLE TRANSCRIPTION
[0004] FIELD OF THE INVENTION
[0005] The invention relates to methods of controlling transcription of a genetic sequence in a cell using a minigene containing an alternatively spliced exon operably linked to the genetic sequence. The method may be used for any cell type, from any eukaryotic organism, but has a particular application in somatic cells derived from an induced pluripotent stem cells (iPSCs).
[0006] BACKGROUND OF THE INVENTION
[0007] Stem cell research holds great promise for research of human development, regenerative medicine, disease modelling, drug discovery, and cell transplantation. Moreover, stem cell-derived cells enable studying physiological and pathological responses of human cell populations that are not easily accessible. This often entails the study of genes (and other forms of regulatory mechanisms encoded in non-protein-coding RNAs - ncRNAs). Unfortunately, controllable transcription or expression of genetic information in human cells has been proven to be particularly difficult.
[0008] Moreover, for several key aspects of regenerative medicine, disease modelling, drug discovery and cell transplantation, manipulation and manufacture of mature human cell types from easily accessible sources is required. Controlling the expression of transgenes in human cells is the basis of biological research. However, this has proven to be difficult in human cells. Moreover, there is a real need for the in vitro derivation of many highly desirable human cell types in a quantity and quality suitable for drug discovery and regenerative medicine purposes. Because directed differentiation of stem cells into desired cell types is often challenging, other approaches have emerged, including direct programming of cells into the desired cell types. In particular, forward programming, as a method of directly converting stem cells and progenitor cells, such as pluripotent stem cells (PSCs), including human PSCs (hPSCs), to lineage-restricted cell types (such as lineage-restricted stem cells or mature cells) has been recognised as a powerful strategy for the derivation of human cells. Alternatively, lineage-restricted cells may be directly converted into a cell of a different lineage (i.e. without needing a pluripotent intermediary step). Such methods are known as transdifferentiation. These programming methods typically involve the forced expression of key transcription factors (or polypeptides having the activity of said one or more transcription factors) (or non-coding RNAs, including IncRNA and microRNA) in order to convert the cell type into another desired cell type.BIT-C-P3829PCT
[0009] 2
[0010] Apart from inducible expression of transgenes, it is very desirable to be able to control knockdown and knockout of genes or other coding sequences in cells, to allow loss of function studies to be carried out. Loss of function studies in stem cells and mature cell types provide a unique opportunity to study the mechanisms that regulate human development, disease and physiology. However, current techniques do not permit the easy and efficient manipulation of gene expression.
[0011] For many applications, it is desirable to control the transcription of inserted genetic material in a cell, such that an inducible cassette may be turned on as required and transcribed at particular levels, including high levels.
[0012] WO2018 / 096343 (incorporated herein by reference) discloses a method of controlling transcription comprising inserting the genetic material into genomic safe harbour (GSH) sites. The advantage of this method is that transcription could be controlled without impacting on the expression of the inserted genetic material or impacting on the functioning of the cell.
[0013] However, generation of lineage-restricted cells by forward programming is a sensitive and complex process, often requiring temporal expression of multiple genes in order to mimic developmentally regulated gene expression cascades and generate mature cells. Furthermore, the addition of exogenous genetic material needs to be carefully controlled when preparing a cellular therapy to avoid the risk of adverse events. Therefore, there is a need in the art to provide alternative and / or additional controllable transcription methods, particularly in the context of preparing cells by forward programming.
[0014] SUMMARY OF THE INVENTION
[0015] According to a first aspect of the invention, there is provided a method of controlling transcription of a genetic sequence in a somatic cell derived from a pluripotent stem cell (PSC) comprising:
[0016] (i) providing a PSC comprising an inserted minigene operably linked to the genetic sequence, wherein the minigene comprises a sequence with an alternatively spliced exon and wherein expression of the genetic sequence is controlled by alternative splicing of the minigene; and
[0017] (ii) applying a splicing modifier that regulates the splicing of the alternatively spliced exon thereby inducing expression of the minigene to include the alternatively spliced exon, wherein the presence of the alternatively spliced exon initiates translation of the genetic sequence.BIT-C-P3829PCT
[0018] 3
[0019] According to a further aspect of the invention, there is provided a method of controlling transcription of an exogenous genetic sequence in a pluripotent stem cell (PSC), comprising:
[0020] (i) administering to the PSC a chimeric gene comprising a first portion comprising a minigene having an alternatively spliced exon and a second portion that encodes the exogenous genetic sequence, wherein expression of the exogenous genetic sequence is controlled by alternative splicing of the first portion; and
[0021] (ii) applying a splicing modifier that regulates the splicing of the alternatively spliced exon thereby inducing expression of the minigene to include the alternatively spliced exon and initiating translation of the exogenous genetic sequence.
[0022] According to a further aspect of the invention, there is provided a method of forward programming a pluripotent stem cell (PSC) into a somatic cell, comprising:
[0023] (i) administering to the PSC a chimeric gene comprising a first portion comprising a minigene having an alternatively spliced exon and a second portion that encodes one or more transcription factors and / or one or more polypeptides having the activity of said one or more transcription factors, wherein expression of the second portion is controlled by alternative splicing of the first portion;
[0024] (ii) applying a splicing modifier that regulates the splicing of the alternatively spliced exon thereby inducing expression of the minigene to include the alternatively spliced exon and initiating translation of the one or more transcription factors and / or polypeptides encoded by the second portion,
[0025] wherein translation of the one or more transcription factors and / or polypeptides encoded by the second portion induces forward programming of the PSC into the somatic cell.
[0026] According to a further aspect of the invention, there is provided a cell obtainable by any one of the methods defined herein.
[0027] According to a further aspect of the invention, there is provided a cell with a modified genome that comprises:
[0028] (a) a first expression cassette comprising an inserted minigene operably linked to a first genetic sequence, wherein the minigene comprises a sequence with an alternatively spliced exon and wherein expression of the first genetic sequence is controlled by alternative splicing of the first portion;
[0029] (b) a second expression cassette comprising a genetic sequence encoding a transcriptional regulator protein; andBIT-C-P3829PCT
[0030] 4
[0031] (c) a third expression cassette comprising a second genetic sequence operably linked to an inducible promoter, wherein said inducible promoter is regulated by the transcriptional regulator protein encoded by the second expression cassette.
[0032] According to a further aspect of the invention, there is provided a cell modified to comprise a chimeric gene comprising a first portion comprising a minigene having an alternatively spliced exon operably linked to a second portion that encodes one or more transcription factors and / or one or more polypeptides having the activity of said one or more transcription factors.
[0033] According to a further aspect of the invention, there is provided a cell as defined herein, for use in therapy.
[0034] According to a further aspect of the invention, there is provided a cell as defined herein, for use in in vitro diagnostics or drug screening.
[0035] According to a further aspect of the invention, there is provided a nucleic acid molecule comprising a sequence having at least 90% sequence identity with SEQ ID NO: 1.
[0036] BRIEF DESCRIPTION OF THE FIGURES
[0037] Figure 1. Controllable expression of enhanced green fluorescent protein (EGFP) using the alternative splicing system. Figure 1A provides a schematic representation of the constructs tested. Constructs labelled ‘sT include an ATG start codon before the genetic sequence (EGFPd2) and constructs labelled ‘s2’ include both an ATG start codon and T2A self-cleaving peptide sequence before the genetic sequence. “EGFPd2” represents the unstable GFP. “saPuroR” represents the puromycin resistance gene. Figures 1B and 1C provide flow cytometry plots assessing expression control in Spliceln constructs (a) controlled by a CAG promoter (Figure 1B), or (b) controlled by a TRE (doxycycline-induced, ‘Dox’) promoter (Figure 1C).
[0038] Figure 2. Controllable forward programming using the alternative splicing system.
[0039] Figure 2A provides a schematic representation of the constructs tested. Constructs labelled ‘sT include an ATG start codon before the genetic sequence (NGN2). Figure 2B provides immunocytochemistry for NGN2 expression controlled by the alternative splicing system following an 11-day standard programming protocol. Scale bar is 400 pm. Figure 2C providesBIT-C-P3829PCT
[0040] 5
[0041] a quantification of gene expression for genes involved in neuronal or iPSC function. Figure 2D provides Principal Component Analysis (PCA) on the whole transcriptome of the cells.
[0042] Figure 3. Using the alternative splicing system as a safety switch. Figure 3A provides a schematic representation of caspase constructs used to control cell apoptosis. Figure 3B provides plots of cell confluency after induction of the caspase safety switch expression following 2 days of treatment with LMI070.
[0043] Figure 4. Controlled Cas9 expression using alternative splicing. Figure 4A provides a schematic representation of the construct tested. Constructs labelled ‘s2’ include both an ATG start codon and T2A self-cleaving peptide sequence before the genetic sequence. “saPuroR” represents the puromycin resistance gene. Figure 4B provides a schematic of the B2M gene editing strategy used. Figure 4C provides the results showing that LMI070 induction of Cas9 leads to B2M editing.
[0044] DETAILED DESCRIPTION
[0045] The present invention provides methods of controlling transcription of a genetic sequence in cells, particularly somatic cells derived from an induced pluripotent stem cell (iPSC) and / or cells which undergo forward programming.
[0046] General definitions
[0047] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this invention belongs. As used herein, the following terms have the meanings ascribed to them below.
[0048] The term “expression” as used herein, refers to the process by which genetic sequence information is converted into a function. Where the genetic information (in the form of DNA) encodes non-coding RNA, the function may be to, for example, interfere with messenger RNA (as discussed in detail below). Where the genetic information encodes a gene, the function will be in relation to the protein encoded by the genetic sequence. It is understood that various levels of expression are achievable, from weak or low expression through to high or strong expression.
[0049] The term “exogenous genetic sequence” as used herein refers to a genetic sequence that has been introduced from outside the organism. In contrast, the term “endogenous geneticBIT-C-P3829PCT
[0050] 6
[0051] sequence” refers to a genetic sequence already present in the host cell genome. The exogenous genetic sequence can be naturally occurring (“wild type”) or non-naturally occurring (e.g. a variant, a mutant, synthetic etc.). Thus, the exogenous genetic sequence is introduced as a result of human intervention, but may be a naturally occurring nucleic acid sequence. For example, the exogenous genetic sequence may be introduced to encode another copy of a genetic sequence already encoded within the host cell genome, but is referred to as “exogenous” because it is introduced from outside the cell. Exogenous genetic sequences may also include “transgenes” which are sequences that have been derived from another organism, as well as non-gene sequences such as promoters.
[0052] The term “chimeric” as used herein applies to a nucleic acid or polypeptide containing two or more components that are defined by structures, optionally derived from different sources. A “chimeric gene” in the context of the invention includes nucleotide sequences derived from different genetic sequences, namely a first portion containing a minigene having an alternatively spliced exon and a second portion encoding a genetic sequence whose transcription is to be controlled.
[0053] References to “transcription factor” as used herein, refer to proteins that are involved in gene regulation in both prokaryotic and eukaryotic organisms. In one embodiment, transcription factors can have a positive effect on gene expression and, thus, may be referred to as an “activator” or a “transcriptional activation factor”. In another embodiment, a transcription factor can negatively affect gene expression and, thus, may be referred to as “repressors” or a “transcription repression factor”. Activators and repressors are generally used terms and their functions may be discerned by those skilled in the art.
[0054] Methods of the invention may be used in a “cell population”, i.e., a collection of cells which may be programmed into the desired cell type, in particular a lineage-restricted cell. Said cell population may comprise “source cells”, also referred to as “starting cells”, i.e., a cell type prior to forward programming into the desired cell type.
[0055] References herein to “pluripotent”’ refer to cells which have the potential to differentiate into all types of cell found in an organism. One form of pluripotent stem cell, known as induced pluripotent stem cells, are of particular interest to the present invention. “Induced pluripotent stem cells” (iPSCs) are cells that have been programmed to an embryonic stem cell-like state by being forced to express genes and factors important for maintaining the defining properties of embryonic stem cells. In 2006, it was shown that overexpression of four specific transcriptionBIT-C-P3829PCT
[0056] 7
[0057] factors could convert adult cells into pluripotent stem cells. OCT-3 / 4 and certain members of the SOX gene family have been identified as potentially crucial transcriptional regulators involved in the induction process. Additional genes including certain members of the KLF family, the MYC family, NANOG, and LIN28, may increase the induction efficiency. Examples of the genes that may be used to induce pluripotency to generate iPSCs include OCT3 / 4, S0X2, S0X1, S0X3, S0X15, S0X17, KLF4, KLF2, C-MYC, N-MYC, L-MYC, NANOG, LIN28, F0X15, ERAS, ECAT15-2, TCL1, CTNNB1, LIN28B, SALLI4, ESRRB, TBX3 and GLIS1, GATA6 and these factors may be used singly, or in combination of two or more kinds thereof. In particular, the programming factors may comprise at least the Yamanaka factors, i.e., OCT3 / 4, S0X2, KLF4 and C-MYC.
[0058] Forward programming methods of the invention may be for use in generating lineage-restricted cells. By “lineage-restricted” it is understood that the cells are, compared with the starting stem cells, limited in terms of the number of different cell types that they can differentiate into (without any further artificial transcription factor manipulation). This definition aligns with the understanding of forward programming in the art (the guided differentiation of cells through the forced expression of polypeptides with transcription factor activity and / or transcription factors). As such, the lineage-restricted cell may be a stem cell itself, but that resultant stem cell would be lineage-restricted compared to the source cell (for example, the forward programming of an iPSC to a mesenchymal stem cell). In one embodiment, the linage-restricted cell is a somatic cell.
[0059] References herein to “somatic” refer to any type of cell that makes up the body of an organism, excluding germ cells. Somatic cells therefore include, for example, skin, heart, muscle, bone or blood cells and their stem cells. Somatic cells may also be referred to as differentiated cells. In one embodiment, the somatic cell may be an adult cell or a cell derived from an adult which displays one or more detectable characteristics of an adult or non-embryonic cell.
[0060] Methods of the invention may be used for “cellular reprogramming”, “forward programming”, “direct programming” or “direct differentiation”, i.e., the pluripotent stem cell (e.g. iPSC) is differentiated into a lineage-restricted cell type. Furthermore, cellular reprogramming may be used as generic terminology referring to the use of transcription factors to differentiate a source cell into a lineage-restricted cell type.
[0061] References herein to “culturing” in general include the addition of cells e.g., the cell population, i.e., the source cells), to media comprising growth factors and / or essential nutrients. It will beBIT-C-P3829PCT
[0062] 8
[0063] appreciated that such culture conditions may be adapted according to the cells or cell population to be generated according to methods of the invention.
[0064] A “promoter” is a nucleotide sequence which is recognised by proteins involved in initiating and regulating transcription of a polynucleotide sequence. An “inducible promoter” is a nucleotide sequence where expression of a genetic sequence operably linked to the promoter is controlled by an analyte, co-factor, regulatory protein, etc.. It is intended that the term “promoter” or “control element” includes full-length promoter regions and functional {e.g., controls and / or affects transcription or translation) segments of these regions. A “constitutive promoter” is a promoter that is active in initiating and regulating transcription under all circumstances.
[0065] The term “operably linked” refers to an arrangement of elements wherein the components so described are configured so as to perform their usual function. Thus, an expression element (e.g. a promoter or inserted minigene, as relevant herein) operably linked to a genetic sequence is capable of affecting the expression of that sequence when the regulatory factors are present. The expression element need not be contiguous with the sequence, so long as it functions to direct the expression thereof. Thus, for example, intervening untranslated yet transcribed sequences can be present between the expression element (e.g. promoter sequence) and the genetic sequence and the expression element can still be considered “operably linked” to the genetic sequence. Thus, the term “operably linked” is intended to encompass any spacing or orientation of the expression element and the genetic sequence in the cassette which allows for initiation of transcription of the cassette upon recognition of the expression element by a transcription complex.
[0066] The term “exon” means any part of the gene that will form a part of the final mature messenger RNA produced by that gene after introns have been removed by RNA splicing. The exon will include the 5’ and 3’ untranslated regions (UTRs) that are transcribed but not translated. By contrast, the term “intron” means any part of the gene that is not expressed or operative in the final messenger RNA product.
[0067] The term “vector”, as used herein, is intended to refer to a nucleic acid molecule which is used as a vehicle to carry genetic material into a cell. One type of vector is a “plasmid”, which refers to a circular double stranded DNA loop or circle into which additional DNA segments may be ligated. Another type of vector is an infectious but non-pathogenic viral vector, wherein additional DNA segments may be ligated to certain viral genetic elements. Certain vectors areBIT-C-P3829PCT
[0068] 9
[0069] capable of autonomous replication in a host cell into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian and yeast vectors). Other vectors (e.g., non-episomal mammalian vectors) can be integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operatively linked. Such vectors are referred to herein as “recombinant expression vectors” (or simply, “expression vectors”). In general, expression vectors of utility in recombinant DNA techniques are often in the form of plasmids. However, the invention is intended to include such other forms of expression vectors, such as viral vectors (e.g., replication defective retroviruses, lentiviral vectors, adenoviruses, Sendai viruses and adeno-associated viruses), which serve equivalent functions, and also bacteriophage and phagemid systems. Another type of vector includes synthetic and in vitro transcribed RNA molecules, e.g., mRNA and stabilised RNA, to carry coding genetic information to the cells. This also includes synthetic-self- replicating RNA vectors.
[0070] References to “subject”, “patient” or “individual” refer to a subject, in particular a mammalian subject, to be treated. Mammalian subjects include humans, non-human primates, farm animals (such as cows), sports animals, or pet animals, such as dogs, cats, guinea pigs, rabbits, rats or mice. In some embodiments, the subject is a human. In alternative embodiments, the subject is a non-human mammal, such as a mouse.
[0071] The term "sufficient amount" means an amount sufficient to produce a desired effect. The term "therapeutically effective amount" is an amount that is effective to ameliorate a symptom of a disease or disorder. A therapeutically effective amount can be a "prophylactically effective amount" as prophylaxis can be considered therapy.
[0072] As used herein, the term “about” when used herein includes up to and including 10% greater and up to and including 10% lower than the value specified, suitably up to and including 5% greater and up to and including 5% lower than the value specified, especially the value specified. The term “between” includes the values of the specified boundaries.
[0073] It will be understood that any method as described herein may have one or more, or all, steps performed in vitro, ex vivo or in vivo. The methods described herein do not modify the germ line genetic identity of a human being.BIT-C-P3829PCT
[0074] 10
[0075] Controlling transcription of a genetic sequence
[0076] According to a first aspect of the invention, there is provided a method of controlling transcription of a genetic sequence in a somatic cell derived from a pluripotent stem cell (PSC) comprising:
[0077] (i) providing a PSC comprising an inserted minigene operably linked to the genetic sequence, wherein the minigene comprises a sequence with an alternatively spliced exon and wherein expression of the genetic sequence is controlled by alternative splicing of the minigene; and
[0078] (ii) applying a splicing modifier that regulates the splicing of the alternatively spliced exon thereby inducing expression of the minigene to include the alternatively spliced exon, wherein the presence of the alternatively spliced exon initiates translation of the genetic sequence.
[0079] Methods of the invention are used to control, i.e. modulate, expression of a genetic sequence, such as by increasing or decreasing expression of the genetic sequence. Elements of the system (e.g. minigene, alternatively spliced exon, genetic sequence) are included in a transcriptionally controlled cassette which utilises mechanisms of alternative splicing.
[0080] Expression of the genetic sequence is controlled by alternative splicing of the minigene. This mechanism can be achieved in different ways. In some embodiments, the genetic sequence operably linked to the minigene lacks an initiation or start codon, includes a translation stop codon, is not an open reading frame to produce a protein, or encodes only a portion of a protein. In some embodiments, alternative splicing of the exon in the minigene modifies the transcript thereby introducing an initiation or start codon, deleting or nullifying the stop codon, restoring the open reading frame, or providing a missing portion of the protein.
[0081] In a preferred embodiment, the alternatively spliced exon comprises a translation initiation regulatory sequence and the genetic sequence lacks a translation initiation regulatory sequence, thereby initiation of translation of the genetic sequence only occurs when the alternatively spliced exon is present.
[0082] In one embodiment, the genetic sequence is suitable for forward programming or transdifferentiation, preferably forward programming. The invention finds particular use in cells produced using forward programming, i.e. the guided differentiation of cells through the forced expression of polypeptides with transcription factor activity and / or transcription factors.BIT-C-P3829PCT
[0083] 11
[0084] Forward programming has been developed as an alternative, scalable, and faster strategy for the production of lineage-restricted cells. However, such methods require control over transcription factor expression to produce a homogenous population of cells and avoid silencing.
[0085] In one embodiment, the splicing modifier is applied to the somatic cell following forward programming of the PSC.
[0086] In an alternative embodiment, the splicing modifier is applied during the process of forward programming the PSC into the somatic cell.
[0087] According to a further aspect of the invention, there is provided a method of controlling transcription of a genetic sequence in an iPSC, comprising:
[0088] (i) administering to the iPSC a minigene having an alternatively spliced exon operably linked to the genetic sequence, wherein the minigene comprises a sequence with an alternatively spliced exon and wherein expression of the genetic sequence is controlled by alternative splicing of the minigene; and
[0089] (ii) applying a splicing modifier that regulates the splicing of the alternatively spliced exon thereby inducing expression of the minigene to include the alternatively spliced exon, wherein the presence of the alternatively spliced exon initiates translation of the genetic sequence.
[0090] Where the minigene is present in a transcriptionally controlled cassette comprising more than one distinct functional component, such as more than one gene, the cassette may be in the form of a bicistronic or multicistronic cassette. Typically such cassettes comprise an internal ribosome entry site (IRES) that allows for the expression of more than one separate protein from one cassette.
[0091] If the inserted genetic sequence comprises at least two coding regions, the coding regions may be linked to each other by nucleic acids that encode self-cleaving peptides so as to form a single open reading frame. Therefore, in one embodiment, the minigene additionally encodes a self-cleaving peptide. In a further embodiment, the self-cleaving peptide is present between the minigene and the genetic sequence. If the minigene is present in a chimeric gene, it will be understood that the chimeric gene may encode a self-cleaving peptide between the first portion and the second portion.BIT-C-P3829PCT
[0092] 12
[0093] In one embodiment, the self-cleaving peptide is a viral 2A peptide. Self-cleaving peptides are found in members of the Picornaviridae virus family, including aphthoviruses such as foot-and-mouth disease virus (FMDV), equine rhinitis A virus (ERAV), Thosea asigna virus (TaV) and porcine teschovirus-1 (PTV-1). The 2A peptides derived from FMDV, ERAV, PTV-1, and TaV are sometimes referred to as “F2A”, “E2A”, “P2A”, and “T2A”, respectively. The 2A sequence is believed to mediate ‘ribosomal skipping’ between the proline and glycine, impairing normal peptide bond formation between the P and G without affecting downstream translation. In a preferred embodiment, the self-cleaving peptide is T2A.
[0094] Minigenes
[0095] Alternative splicing is a cellular process in which exons from the same gene are joined in different combinations, leading to different, but related, mRNA transcripts. These mRNAs can be translated to produce different proteins with distinct structures and functions — all from a single gene. Methods of the invention utilise alternative splicing as a mechanism to regulate gene expression. Increased expression of the protein encoded by the genetic sequence is provided by inclusion of an alternatively spliced exon present within a minigene. References herein to a “minigene” refer to a minimal gene fragment. They contain at least one exon and the control regions necessary for expression, but are not required to contain all of the components (e.g. introns and exons) of the wild type gene. Minigenes are known in the art and have been used to investigate splicing patterns.
[0096] In one example, the minigene comprises three exons, Exons 1-3, where Exon 2 is skipped in due to the presence of a 3’ and 5’ splicing site. Translation initiation regulatory sequences are located in Exon 2, and thus when Exon 2 is skipped no translation occurs. Alternative mechanisms to prevent expression can also be used. For example, if Exon 2 is skipped, the reading frame of Exon 3 is shifted, resulting in the creation of a nonsense mutation in Exon 3. As such, translation of the encoded protein stops in Exon 3, and nothing downstream is translated. Since the genetic sequence is located downstream of the minigene, the genetic sequence is not expressed in the basal state.
[0097] In one embodiment, the minigene comprises (from 5’ to 3’): Exon 1, Intron 1, Exon 2, Intron 2, and Exon 3, wherein Exon 2 is the alternatively spliced exon. In a further embodiment, Exon 2 comprises a translation initiation regulatory sequence. In this embodiment, preferably the genetic sequence lacks a translation initiation regulatory sequence.BIT-C-P3829PCT
[0098] 13
[0099] In one embodiment, inclusion of Exon 2 causes a frameshift. In this embodiment, preferably the number of nucleotides present in Exon 2 is not divisible by 3. In one embodiment, Exon 3 comprises a stop codon that is in frame when Exon 2 is skipped. In a further embodiment, the genetic sequence is in frame with the translation initiation regulatory sequence in Exon 2.
[0100] Once an alternatively spliced exon is identified, it can be engineered further for use in the transcriptionally controlled cassette. For example, the alternatively spliced exon could be engineered to contain a Kozak sequence and / or AUG start codon. All possible downstream start codons (e.g. in the genetic sequence) are removed to ensure that an AUG start would be included in response to splicing modifier binding only.
[0101] In one embodiment, the minigene is derived from the SF3B3 gene. This gene encodes subunit 3 of the splicing factor 3b protein complex. The SF3B3 gene may be identified by Ensembl Gene ID: ENSG00000189091.
[0102] In a further embodiment, the minigene comprises an upstream exon, pseudoexon and a downstream exon derived from the SF3B3 gene. A pseudoexon in this context is an exon with one or more degenerate splice junctions, meaning that the sequence has lost its original function. In this embodiment, the pseudoexon of the SF3B3 gene is skipped in the basal state. However, the pseudoexon is included in the presence of a splicing modifier (e.g. LMI070 or RG7800 / RG7619). In one embodiment, the minigene comprises Exon 1, pseudoexon 2a and Exon 2 of the SF3B3 gene.
[0103] In one embodiment, the minigene is derived from the SMN2 gene. This gene encodes the SMN2 protein in humans. The SMN2 gene may be identified by Ensembl Gene ID: ENSG00000277773. Spinal Muscular Atrophy (SMA) results from mutations in the SMN1 gene. However, humans have a very similar gene, SMN2 that can serve to modify the severity of SMN1 deficiency depending on the number of SMN2 copies resident in the patient’s genome. SMN2 cannot fully replace SMN1, however, because unlike SMN1, SMN2 has undergone variations impairing exon 7 inclusion. As a consequence, only about 10% of SMN2 is correctly spliced. Clinical drugs have been developed to improve exon 7 inclusion as a therapy for SMA.
[0104] In a further embodiment, the minigene comprises exons 6-8 of the SMN2 gene. In a further embodiment, the minigene comprises exons 6-7 and the 5’ end of exon 8 of the SMN2 gene. Inclusion of SMN2 exon 7 is triggered by the presence of a splicing modifier.BIT-C-P3829PCT
[0105] 14
[0106] It will be understood that alternative minigenes that may be used in the invention can be designed using methods known to a person skilled in the art. The Human Splicing Finder website (Desmet et al. (2009) Nucleic Acid Res. 37:e67) can be used to identify 5’ and 3’ splice site pairs that may be used to select, or incorporate into, appropriate minigenes. The same tools can be used to identify combinations of silencer and enhancer splicing sequences that repress or promote the selection of cryptic or correct splice sites, and can be used to engineer the minigene further. For example, splice-site modifications can be introduced to reduce background levels of inclusion of the alternatively spliced exon. For example, in addition to the SMN2 gene discussed above, a minigene comprising a pseudoexon in intron 49 of a pathogenic mutant huntingtin gene has been identified (and splicing modifiers HTT-D3 and HTT-C2 developed to target the minigene), and introduction of the pseudoexon provides a premature stop codon, leading to nonsense-mediated mRNA decay of the pathogenic gene.
[0107] In one embodiment, the minigene comprises fewer than 2000, fewer than 1900, fewer than 1800, fewer than 1700, fewer than 1600, fewer than 1500, fewer than 1400, fewer than 1300, fewer than 1200, fewer than 1100, fewer than 1000, fewer than 900, fewer than 800, fewer than 700, fewer than 600 or fewer than 500 nucleotides. Preferably, the minigene comprises fewer than 1000 nucleotides, in particular fewer than 650 nucleotides.
[0108] In one embodiment, the minigene comprises between 500 and 2000 nucleotides, such as between 500-1800, 500-1500, 500-1000, 600-1800, 600-1500, 600-1000, 500-800, 600-800, 500-700 or 600-700 nucleotides.
[0109] More details regarding the minigene and chimeric genes are described in W02020 / 033473 and WO2021 / 163556 which are herein incorporated by reference.
[0110] In one embodiment, the minigene comprises a sequence having at least 80% sequence identity, such as at least 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity, with SEQ ID NO: 1. In a further embodiment, the minigene comprises a sequence having at least 90% sequence identity with SEQ ID NO: 1. In a yet further embodiment, the minigene comprises a sequence of SEQ ID NO: 1.BIT-C-P3829PCT
[0111] 15
[0112] SEQ ID NO: 1 actagttccaaggccacagatccgtttaaacaagcttggcaatccggtactgttggtaaatttctgtacaacttaaccttgcagagag ccactggcatcagctttgccattcttggaaacttttctggtaagttctctcgttaccatcttttgaaattttaagtgaattaatacatatcttgc ttagtctcttgtgcaggaaatgttttccatttatgacaaaacaggaattgtgtgaaatttcaaatattggattaggaaatacaaagttact gaaagtgaggtactaatgtttataaaataaaaactttttcttgccatttgcagatttaacatttttgagtcaatccaagtgccaccATG CAGGAGGTTCATGATTGTGTAgagtaagacataattttgttgaggtttaactctgaatacttaatgtggtactgaatactt aatgtggtactgagaggcagcctaactgaccacacagcattcacgcctggctaatttttgtatttttagtagagacggggtttcaccat ggtggccaggctgggttctggttgtttatgatctttattttttggtgatctaggAACCAAACAACAAGAAATTGTTGTTT CCCGTGGGTAAGA
[0113] In SEQ ID NO: 1 listed above, lower case nucleotides represent non-coding sequences (such as 5’IITR and intronic sequences) and the upper case nucleotides represent coding nucleotides. An expression control sequence, such as a promoter, can be included 5’ of the sequence. The genetic sequence to be expressed is encoded 3’ of the sequence.
[0114] According to a further aspect of the invention, there is provided a nucleic acid molecule comprising a sequence having at least 90% sequence identity with SEQ ID NO: 1.
[0115] Suitably, the nucleic acid molecule is isolated. An “isolated” nucleic acid molecule (also referred to as a polynucleotide) is one that is removed from its original environment. For example, a naturally-occurring polynucleotide is isolated if it is separated from some or all of the coexisting materials in the natural system. A polynucleotide is considered to be isolated if, for example, it is cloned into a vector that is not a part of its natural environment or if it is comprised within cDNA.
[0116] For the purposes of comparing two closely-related polynucleotide sequences, the “% sequence identity” between a first nucleotide sequence and a second nucleotide sequence may be calculated using NCBI BLAST, using standard settings for nucleotide sequences (BLASTN).
[0117] Polypeptide or polynucleotide sequences are said to be the same as or “identical” to other polypeptide or polynucleotide sequences, if they share 100% sequence identity over their entire length. Residues in sequences are numbered from left to right, i.e. from N- to C- terminus for polypeptides; from 5’ to 3’ terminus for polynucleotides.BIT-C-P3829PCT
[0118] 16
[0119] In order to turn on expression of the genetic sequence, the inclusion of the skipped exon must be induced. Such can occur as a result of the presence of a splicing modifier.
[0120] In one embodiment, the splicing modifier is a small molecule. A small molecule, or a micromolecule, is commonly known in the art as one that is less than 1000 daltons in weight and with a size on the order of 1 nm.
[0121] In one embodiment, the small molecule splicing modifier is LMI070. The drug LMI070 may also be referred to as Branaplam, HTT-C2 or NVS-SM1. The structure of LMI070 is known in the art, for example see CAS No.: 1562338-42-4. It was developed and previously sold by Novartis. Analogues of splice modifiers such as LMI070 may also be used.
[0122] In one embodiment, the minigene comprises the sequence: AGAGTAAGAC (SEQ ID NO: 2). Monteys et al. (2021) Nature 596:291-295 screened candidate exons and found that LMI070-treated samples share a strong AGAGLIA motif that is consistent with the identified LMI070-targeted U1 RNA binding site. Therefore, exons containing this motif are susceptible to a LMI070-induced splicing event.
[0123] In an alternative embodiment, the small molecule splicing modifier is RG7800. It has been developed and sold by Roche. The structure of RG7800 is known in the art, for example see CAS No.: 1449598-06-4. Analogues of splice modifiers such as RG7800 may also be used, for example RG7916 (Risdiplam).
[0124] In an alternative embodiment, the small molecule splicing modifier is SMN-C5, another modifier developed by Roche. The structure of SMN-C5 is known in the art, for example see CAS No.: 1449592-94-2. Analogues of splice modifiers such as SMN-C5 may also be used.
[0125] These drugs have been shown to improve the proportion of correct SMN2 splicing, and both show efficacy in animal models of spinal muscular atrophy. These modifiers are therefore drugs that can act on an alternative splicing system. RG7800 and LMI070 have been shown to be effective clinically.
[0126] In an alternative embodiment, the small molecule splicing modifier is HTT-D3. This modifier has been developed by PTC Therapeutics. The structure of HTT-D3 is known in the art, for example see CAS No.: 2254502-89-9. Analogues of splicing modifiers such as HTT-D3 may also be used.BIT-C-P3829PCT
[0127] 17
[0128] As discussed above, HTT-D3 and LMI070 have been shown to remove pathogenic huntingtin genes and are beneficial in treating Huntington's disease in animal models.
[0129] Small molecule splicing modifiers increase incorporation of an exon by stabilising the interaction between the spliceosome (the machinery involved with the splicing process) and the pre-mRNA (that still includes the exon of interest). In particular, the small molecule modifiers may interact with the tertiary RNA structure that includes the exonic splicing enhancer sequence of the exon of interest leading to the stabilisation of an unpaired nucleotide in position 1 of the exon-intron junction through a bulge repair mechanism. This allosteric stabilisation provided by the modifier may strengthen the interaction of specifically the I11 small nuclear ribonucleoprotein (snRNP) of the spliceosome machinery with the 5’ spliceosome start site, leading to exon inclusion. This is set out in review articles such as Malard et al. (2022) RNA Biology 19:943-60, which is herein incorporated by reference.
[0130] In one embodiment the splicing modifier is an antisense oligonucleotide (ASO). ASOs are known in the art that are capable of masking c / s-acting regulatory elements in pre-mRNA and by doing so induce changes in the splicing pattern. ASOs, such as nusinersen, can mask silencer sequences, which then promotes exon inclusion. Nusinersen specifically masks splicing silencer N1 , located immediately downstream of the weak 5’ start site of SMN2 exon 7, alleviating repression of the 5’ start site and promoting exon 7 inclusion. In a preferred embodiment, the splicing modifier is nusinersen, a 18 nt 2’-O-(2-methoxyethyl)-oligoribonucleotide with a fully modified phosphorothioate backbone. The structure of nusinersen is known in the art, for example see CAS No: 1258984-36-9. Analogues of nusinersen may also be used.
[0131] Alternatively, inclusion of the skipped exon may be induced in one cell type, but not another. For example, Exons 8 and 9 of FGFR2 are mutually exclusive, with Exon 9 being only included in mesenchymal tissue (Takeuchi et al. (2010) PLOS One, 5:e10946). As such, the minigene may comprise Exons 7-10 of the FGFR2 gene to allow for expression of the genetic sequence only in mesenchymal cells. In this case, a stop codon is engineered into Exon 8 of the minigene to prevent expression of the genetic sequence in non-mesenchymal cells.
[0132] Exogenous genetic sequences
[0133] In one embodiment, the genetic sequence is an exogenous genetic sequence. In such an embodiment, the inserted minigene is preferably present in a chimeric gene comprising a firstBIT-C-P3829PCT
[0134] 18
[0135] portion encoding the minigene and a second portion encoding the exogenous genetic sequence.
[0136] The introduction of a transcriptionally controlled cassette into the genome has the potential to change the phenotype of that cell, either by addition of a genetic sequence that permits gene expression or knockdown / knockout of endogenous expression. The methods of the invention provide for controllable transcription of the genetic sequence(s) within the transcriptionally controlled cassette in the cell.
[0137] In one embodiment, the genetic sequence is a DNA sequence that encodes an RNA molecule. The RNA molecule may be of any sequence, but is preferably coding or non-coding RNA. In one embodiment, the genetic sequence encodes a non-coding RNA, such as an inhibitory RNA.
[0138] Coding or messenger RNA codes for polypeptide sequences, and transcription of such RNA leads to expression of a protein within the cell. Non-coding RNA may be functional and may include without limitation: MicroRNA, Small interfering RNA, Piwi-interacting RNA, Antisense RNA, Small nuclear RNA, Small nucleolar RNA, Small Cajal Body RNA, Y RNA, Enhancer RNAs, Guide RNA, Ribozymes, Small hairpin RNA, Small temporal RNA, Trans-acting RNA, small interfering RNA and Subgenomic messenger RNA. Non-coding RNA may also be known as functional RNA. Several types of RNA are regulatory in nature, and, for example, can downregulate gene expression by being complementary to a part of an mRNA or a gene's DNA. MicroRNAs (miRNA, approximately 21-22 nucleotides in length) are found in eukaryotes and act through RNA interference (RNAi), where an effector complex of miRNA and enzymes can cleave complementary mRNA, block the mRNA from being translated, or accelerate its degradation. Another type of RNA, small interfering RNAs (siRNA, approximately 20-25 nucleotides in length) act through RNA interference in a fashion similar to miRNAs. Some miRNAs and siRNAs can cause genes they target to be methylated, thereby decreasing or increasing transcription of those genes. Animals have Piwi-interacting RNAs (piRNA, approximately 29-30 nucleotides in length) that are active in germline cells and are thought to be a defence against transposons. Many prokaryotes have CRISPR RNAs, a regulatory system similar to RNA interference, and such a system include guide RNA (gRNA). Antisense RNAs are widespread; most downregulate a gene, but a few are activators of transcription. Antisense RNA can act by binding to an mRNA, forming double-stranded RNA that is enzymatically degraded. There are many long non-coding RNAs that regulate genes in eukaryotes, one such RNA is Xist, which coats one X chromosome in female mammals andBIT-C-P3829PCT
[0139] 19
[0140] inactivates it. Thus, there are a multitude of functional RNAs that can be employed in the methods of the present invention.
[0141] In one embodiment, the genetic sequence encodes a protein-coding gene. This gene may be not naturally present in the cell, or may naturally occur in the cell, but controllable expression of that gene is required. Alternatively, the genetic sequence may be a mutated, modified or correct version of a gene present in the cell, particularly for gene therapy purposes or the derivation of disease models. The genetic sequence may thus include a transgene from a different organism of the same species (i.e. a diseased / mutated version of a gene from a human, or a wild-type gene from a human) or be from a different species.
[0142] In any aspect or embodiment, the genetic sequence may be a synthetic sequence.
[0143] The genetic sequence may include any suitable genetic sequence that it is desired to insert into the genome of the cell. Therefore, the genetic sequence may be a gene that codes for a protein product or a sequence that is transcribed into ribonucleic acid (RNA) which has a function (such as small nuclear RNA (snRNA), antisense RNA, micro RNA (miRNA), small interfering RNA (siRNA), transfer RNA (tRNA) and other non-coding RNAs (ncRNA), including CRISPR-RNA (crRNA) and guide RNA (gRNA).
[0144] The genetic sequence may thus include any genetic sequence, the transcription of which it is desired to control within the cell. The genetic sequence chosen will be dependent upon the cell type and the use to which the cell will be put after modification, as discussed further below.
[0145] Methods of the invention can be used to increase the expression of a target protein by using the genetic sequence to encode the target protein, thereby increasing the transcription of the target protein when the alternatively spliced exon is present. Alternatively, the expression of a target protein can be modulated by using the genetic sequence to encode an RNA or protein that will modulate the expression of an endogenous gene when the alternatively spliced exon is present.
[0146] In one embodiment, the genetic sequence reduces or inhibits the expression of an endogenous gene.
[0147] Further, the genetic sequence may encode non-coding RNA whose function is to knockdown the expression of an endogenous gene or DNA sequence encoding non-coding RNA in theBIT-C-P3829PCT
[0148] 20
[0149] cell. Alternatively, the genetic sequence may encode guide RNA for the CRISPR-Cas system to effect endogenous gene knockout. Preferably the gene targeted for knockout or knockdown is different from the endogenous gene where the genetic material is inserted.
[0150] The methods of the invention thus extend to methods of knocking down endogenous gene expression within a cell. This may be the same endogenous gene where the genetic material is inserted, but preferably it is not. The methods are as described previously, wherein the genetic sequence encodes a non-coding RNA, wherein the non-coding RNA suppresses the expression of said endogenous gene. The non-coding RNA may suppress gene expression by any suitable means including RNA interference and antisense RNA. Thus, the genetic sequence may encode a shRNA which can interfere with the messenger RNA for the endogenous gene.
[0151] The reduction in endogenous gene expression may be partial or full - i.e. expression may be 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% reduced compared to the cell prior to induction of the transcription of the noncoding RNA.
[0152] The methods of the invention also extend to methods of knocking out endogenous genes within a cell, by virtue of the CRIPSR-Cas9 system, although any other suitable systems for gene knockout may be used. Guide RNA (gRNA) is a short synthetic RNA composed of a scaffold sequence necessary for Cas9-binding and an approximately 20 nucleotide targeting sequence which defines the genomic target to be modified. Thus, the genomic target of Cas9 can be changed by simply changing the targeting sequence present in the gRNA.
[0153] In one embodiment, the CRISPR enzyme is used to target a transcriptional regulator protein, such as rtTA.
[0154] In one embodiment, the genetic sequence activates the expression of an endogenous gene. The genetic sequence may encode gRNA for a CRISPR-Cas activation system to effect endogenous gene activation. It is known in the art that CRISPR enzyme may be fused to transcriptional activators or repressors, allowing for not only the knockout of a genomic target but also the activation (CRISPRa) or the interference (CRISPRi) of a genomic target. Preferably the gene targeted for activation is different from the endogenous gene where the genetic material is inserted.BIT-C-P3829PCT
[0155] 21
[0156] The increase in endogenous gene expression may be partial or full - i.e. expression may be 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% increased compared to the cell prior to induction of the transcription of the noncoding RNA.
[0157] In one embodiment, the method is used to prepare a somatic cell for use in cell therapy. Cell therapy is the transfer of intact, live cells into a patient to help lessen or cure a disease.
[0158] Cells produced by methods of the invention may be engineered to express a chimeric antigen receptor (CAR). Therefore, in one embodiment, the genetic sequence encodes a chimeric antigen receptor. A CAR typically consists of an extracellular antigen binding domain, a transmembrane domain, and an intracellular domain that propagates an activation signal that activates the cell. The CAR may make the cell specific for malignant cells and therefore useful for cancer immunotherapy. For example, the engineered cells may recognize cancer cells expressing a tumour antigen, such as a tumour associated antigen that is not expressed by normal somatic cells from the subject tissue. Thus, the CAR-modified cells may be used for adoptive cell therapy of, for example, cancer patients.
[0159] While cell therapy can induce an effective immune response to diseased cells, it can lead to severe therapy-associated toxicities including CAR-related hematotoxicity, ON-target OFF-tumour toxicity, cytokine release syndrome (CRS) or immune effector cell-associated neurotoxicity syndrome (ICANS). This has led to the development of genetic safety switches that can be used to control the activity of the cell therapy.
[0160] In one embodiment, the genetic sequence encodes a safety switch. A “safety switch” refers to an engineered protein designed to prevent potential toxicity or otherwise adverse effects of a cell therapy. The safety switch could mediate induction of apoptosis, inhibition of protein synthesis or DNA replication, growth arrest, transcriptional and post-transcriptional genetic regulation and / or antibody-mediated depletion. Examples of safety switch proteins, include, but are not limited to suicide genes such as caspase 9 (or caspase 3 or 7), thymidine kinase, cytosine deaminase, B-cell CD20, modified EGFR, and any combination thereof. One example of a safety switch are engineered dimeric caspases, as described in Chao et al. (2005) PLOS Biol. 3(6): e183; and Yin et al. (2006) Mol. Cell, which are herein incorporated by reference.
[0161] In one embodiment, the method is used to prepare a somatic cell for use in gene therapy. For gene therapy methods, it may be desirable to provide the wild-type gene sequence as theBIT-C-P3829PCT
[0162] 22
[0163] genetic sequence. In this scenario, the genetic sequence may be any human or animal proteincoding gene. Examples of protein-encoding genes include the human B-globin gene, human lipoprotein lipase (LPL) gene, Rab escort protein 1 in humans encoded by the CHM gene and many more. Alternatively, the genetic sequence may encode growth factors, including BDNF, GDF, NGF, IGF, FGF and / or enzymes that can cleave pro-peptides to form active forms. Gene therapy may also be achieved by expression of a genetic sequence encoding an antisense RNA, a miRNA, a siRNA or any type of RNA that interferes with the expression of another gene within the cell.
[0164] Alternatively, or additionally, the genetic sequence may be genes whose function requires investigation, such that controllable expression can look at the effect of expression on the cell; the gene may include transcription factors, growth factors and / or cytokines in order for the cells to be used in cell transplantation; and / or or the gene may be components of a reporter assay.
[0165] In one embodiment, the genetic sequence encodes a DNA-binding enzyme, such as a recombinase, integrase or transposase. In a further embodiment, the DNA-binding enzyme is a recombinase. A recombinase could be used to remove a blocker sequence (preferably flanked by recombinase recognition sites), and so can be used as a genetic activation system. Such systems are known in the art, as discussed in, for example, WO2023 / 235815 (incorporated by reference).
[0166] CRISPR / Cas
[0167] In one embodiment, the genetic sequence encodes a CRISPR enzyme, catalytically-inactive CRISPR enzyme or a derivative thereof. In one embodiment, the genetic sequence encodes a CRISPR enzyme.
[0168] The “CRISPR enzyme” is a major protein component of a clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated protein (Cas) system and forms a complex with a guide RNA (gRNA), thereby forming a CRISPR-Cas system.
[0169] Three types of CRISPR mechanisms have been identified, of which type II is the most studied. The CRISPR / Cas9 system (type II) utilises the Cas9 nuclease to make a double-stranded break (DSB) in DNA at a site determined by a short guide RNA (also referred to herein as simply, guide RNA). The CRISPR / Cas system is a prokaryotic immune system that confers resistance to foreign genetic elements. CRISPR are segments of prokaryotic DNA containingBIT-C-P3829PCT
[0170] 23
[0171] short repetitions of base sequences. Each repetition is followed by short segments of “protospacer DNA” from previous exposures to foreign genetic elements. CRISPR spacers recognize and cut the exogenous genetic elements using RNA interference. The CRISPR immune response occurs through two steps: CRISPR-RNA (crRNA) biogenesis and crRNA-guided interference. CrRNA molecules are composed of a variable sequence transcribed from the protospacer DNA and a CRISPR repeat. Each crRNA molecule then hybridizes with a second RNA, known as the trans-activating CRISPR RNA (tracrRNA) and together these two eventually form a complex with the nuclease Cas9. The protospacer DNA encoded section of the crRNA directs Cas9 to cleave complementary target DNA sequences, if they are adjacent to short sequences known as protospacer adjacent motifs (PAMs). This natural system has been engineered and exploited to introduce DSB breaks in specific sites in genomic DNA, amongst many other applications. In particular, the CRISPR type II system from Streptococcus pyogenes may be used. At its simplest, the CRISPR / Cas9 system comprises two components that are delivered to the cell to provide genome editing: the Cas9 nuclease itself and a gRNA. The gRNA is a fusion of a customised, site-specific crRNA (directed to the target sequence) and a standardised tracrRNA.
[0172] In one embodiment, the CRISPR enzyme is a Cas protein. In a further embodiment, the CRISPR enzyme is Cas9. In another embodiment, the CRISPR enzyme is Cas12a / Cpf1. CRISPR enzymes are known in the art, such as those described in Wang etal. Genome Biol.
[0173] 19, 62 (2018). Other CRISPR enzymes have been utilized in human cells and are known to the skilled person.
[0174] Derivatives of the CRISPR / Cas system are also possible. For example, mutant forms of Cas9 are available, such as Cas9D10A, with only nickase activity. This means it cleaves only one DNA strand, and does not activate NHEJ. Instead, when provided with a homologous repair template, DNA repairs are conducted via the high-fidelity HDR pathway only. Cas9D10A may be used in paired Cas9 complexes designed to generate adjacent DNA nicks in conjunction with two sgRNAs complementary to the adjacent area on opposite strands of the target site.
[0175] When the genetic sequence is a CRISPR enzyme (including a catalytically inactive CRISPR enzyme or derivatives thereof) and the methods involve forward programming or transdifferentiation, the CRISPR enzyme may still be expressed after the change in cell fate, so that a gRNA may be delivered to the resultant cells in order to modulate the expression of an endogenous gene in the resultant cell.BIT-C-P3829PCT
[0176] 24
[0177] In one embodiment, the genetic sequence encoding the CRISPR enzyme is codon optimised. In particular, the genetic sequence is codon optimised for expression in eukaryotic cells, such as human cells. To codon optimise a genetic sequence of interest, codons within the gene that are present at low frequency (or not at all) in the host (which may be referred to as “nonpreferred codons”) are replaced with synonymous codons that are more commonly used by the host (which may be referred to as “preferred codons”). Codon optimisation aims to improve the expression efficiency of genes of interest without altering the sequence of the encoded proteins. In a further embodiment, the Cas9 is a codon optimised Cas9. Examples of such sequences are described in WO2023 / 105212, which is herein incorporated by reference.
[0178] With the development of highly accurate protein structure prediction with artificial intelligence tools such as AlphaFold, it is now straightforward for polypeptides to be developed that have very similar structure and / or activity to a CRISPR enzyme whilst at the same time having an amino acid sequence that has very little resemblance to that of the naturally occurring CRISPR enzyme. In this regard, large language models trained on biological diversity have been used to develop proteins only around 70% identical to CRISPR-Cas proteins that occur in nature and yet with comparable or improved biological activity and specificity (Ruffolo et al. (2024) bioRxiv, doi: https: / / doi.org / 10.1101 / 2024.04.22.590591). Such polypeptides are covered within the scope of the definition of “CRISPR enzyme” herein.
[0179] In one embodiment, the genetic sequence encodes a catalytically-inactive CRISPR enzyme. Catalytically-inactive CRISPR enzymes may also be referred to as ‘dead’ CRISPR enzymes. The CRISPR enzyme may be rendered catalytically inactive through point mutations in the endonuclease domain(s), such as the RuvC and HNH domains in Cas9. Examples of such point mutations in the endonuclease domains of catalytically inactive Cas9 (also known as dCas9) are D10A and H840A which result in its deactivation. However, as will be readily appreciated such mutations do not affect the ability of the CRISPR enzyme to bind gRNA and to the targeted gene, as such binding is affected by other domains.
[0180] In one embodiment, the catalytically-inactive CRISPR enzyme is catalytically inactive Cas9 ( / .e. dead Cas9 or dCas9). In another embodiment, the catalytically-inactive CRISPR enzyme is catalytically inactive Cas12a ( / .e., dCas12a). In further embodiments, the catalytically-inactive CRISPR enzyme (e.g., dCas9 or dCas12a) is a derivative that comprises point mutations in the RuvCI and HNH nuclease domains. In a particular embodiment, the catalytically-inactive CRISPR enzyme is dCas9 and comprises the point mutations D10A and H840A compared to the wild-type sequence of Cas9.BIT-C-P3829PCT
[0181] 25
[0182] The methods presented herein may utilise CRISPR activation (CRISPRa) which uses modified versions of CRISPR effectors which do not have endonuclease activity but comprise added transcriptional activators on the catalytically-inactive CRISPR enzyme {e.g. dCas9 and dCas12a fused to transcriptional activators) and / or the gRNAs. The transcriptional activators fused to the CRISPRa components or which bind thereto (e.g., through the presence of aptamer sequences) therefore increase expression of genes of interest following targeting to the gene by the gRNA.
[0183] Therefore, in one embodiment, expression of a target gene is activated (CRISPR activation). In this embodiment, one or more transcription activator proteins may be fused to the catalytically-inactive CRISPR enzyme. In a further embodiment, the catalytically-inactive CRISPR enzyme is dCas9 and is fused to the transcription activator proteins VP64, p65 and Rta ( / .e. the catalytically inactive programmable nuclease is dCas9-VPR). In another embodiment, the gRNA sequence is engineered to comprise a transcription activation domain which is capable of recruiting one or more transcription activator proteins to upregulate expression of an endogenous gene.
[0184] The methods presented herein may utilise CRISPR interference (CRISPRi) which uses modified versions of CRISPR effectors which do not have endonuclease activity but comprise added transcriptional repressors on the catalytically-inactive CRISPR enzyme (e.g. dCas9 and dCas12a fused to transcriptional repressors) and / or the gRNAs.
[0185] Therefore, in one embodiment, the expression of a target gene is repressed (CRISPR interference). In this embodiment, one or more transcription repressor proteins is fused to the catalytically-inactive CRISPR enzyme. In one embodiment, transcriptional repressors are chosen from the KRAB repressors, in particular the KRAB domain of KOX1 or ZIM3. According to this embodiment, the catalytically-inactive CRISPR enzyme may be dCas9 fused to the KRAB domain of KOX1 or ZIM3 ( / .e. the catalytically-inactive CRISPR enzyme may be dCas9-KOX1 or dCas9-ZIM3). In another embodiment, the gRNA sequence is engineered to comprise a transcription repression domain which is capable of recruiting one or more transcription repressor proteins to downregulate expression of an endogenous gene.
[0186] In one embodiment, the gRNA sequence is complementary to a coding exon of an endogenous gene. Depending on the CRISPR-related protein targeted by the gRNA, affecting the codingBIT-C-P3829PCT
[0187] 26
[0188] exon of an endogenous gene will result in modulation of the expression of said endogenous gene.
[0189] In one embodiment, the gRNA sequence is complementary to a transcription start site (TSS) for an endogenous gene. Depending on the CRISPR-related protein targeted by the gRNA, affecting the TSS of an endogenous gene will result in modulation of the expression of said endogenous gene.
[0190] Once a DSB has been made, a donor template with homology to the targeted locus can be supplied. Distinct cellular repair mechanisms can be exploited to repair the DSB and to introduce the desired sequence, and these are non-homologous end joining repair (NHEJ), which is more prone to error; and homologous recombination repair. The DSB may be repaired by the homology-directed repair (HDR) pathway allowing for precise insertions to be made.
[0191] In one embodiment, the method comprises introducing a guide RNA (gRNA) sequence into the somatic cell, wherein said gRNA sequence is complementary to a target gene or nucleic acid sequence (e.g. an endogenous gene), thereby modulating the expression of the target gene or nucleic acid sequence in the somatic cell.
[0192] A “gRNA” or “guide RNA” refers to an RNA capable of specifically targeting a CRISPR complex, that is, a gRNA-CRISPR enzyme complex, with respect to a target gene or nucleic acid sequence. The gRNA is specific RNA for a target gene or nucleic acid sequence (e.g. an endogenous gene), which may bind to the CRISPR enzyme, and guide the CRISPR enzyme to the target gene or nucleic acid sequence.
[0193] Thus, the gRNA sequence targets the CRISPR enzyme, a catalytically-inactive CRISPR enzyme or a derivative thereof, to a target gene or nucleic acid sequence (e.g. an endogenous gene) to modulate its expression.
[0194] Methods of delivering the RNA into the cell are known in the art. The gRNA may be a synthetic RNA or it may be incorporated in a plasmid or incorporated in a lentiviral donor plasmid. The gRNA may be delivered by a vector, such as a viral vector, as described herein. In one embodiment, the gRNA is introduced by viral transduction. In another embodiment, the gRNA is introduced using lipid-based transfection or electroporation of a gRNA expressing vector. In another embodiment, the gRNA is introduced using lipid-based transfection or electroporation of a synthetic gRNA.BIT-C-P3829PCT
[0195] 27
[0196] In one embodiment, two or more gRNA sequences are introduced. In a further embodiment, each of the gRNA sequences are complementary to alternative or different endogenous genes, or are complementary to more than one alternative sequences of a TSS for an endogenous gene.
[0197] According to embodiments where multiple {e.g., two or more), or a pool of gRNA sequences are introduced, said multiple gRNA sequences may be separated by cleavable sequences. Cleavable sequences are sequences that are recognised by an entity capable of specifically cutting DNA, and include restriction sites, which are the target sequences for restriction enzymes or sequences for recognition by other DNA cleaving entities, such as nucleases, recombinases, ribozymes or artificial constructs. At least one cleavable sequence may be included. These cleavable sequences may be at any suitable point, such that a selected portion of the sequence, or all of the sequence, can be selectively removed. The cleavable sites may thus flank the part / all of the sequence may be subsequently removed. Such cleavable sequences may be recognised by nucleases, recombinases, ribozymes or artificial constructs {e.g., a bacterial DNA endonuclease, such as Csy4). In one embodiment, the multiple gRNA sequences are provided in an array comprising a promoter, two or more gRNA sequences and two or more cleavable sequences. In a particular embodiment, the multiple gRNA sequences are provided in an array comprising a constitutive promoter e.g. the hU6 promoter), two gRNA sequences and two cleavable sequences. In a further embodiment, each guide RNA sequence may be driven by a separate promoter, e.g. a first gRNA driven by the human U6 promoter, a second gRNA driven by the mouse U6 promoter, a third gRNA driven by the bovine U6 promoter, a fourth gRNA driven by the H1 promoter and so on.
[0198] In another embodiment, the two or more gRNA sequences are introduced separately, e.g., by plasmid transfection, electroporation, lentiviral infection or AAV infection. In one embodiment, a lentiviral vector is used that harbours a single gRNA expression cassette. In a further embodiment, a gRNA library is introduced into the lentiviral vector and provided to the target cell for pooled or arrayed CRISPR screening.
[0199] In one embodiment, the one or more gRNA sequences are operably linked to a constitutive promoter. In one embodiment, the promoter is an RNA Polymerase Ill-driven promoter, such as human U6, mouse U6, bovine U6 or the human H1 promoter.BIT-C-P3829PCT
[0200] 28
[0201] In one embodiment, the gRNA is introduced once cells have exited pluripotency, i.e. introduced to the somatic cell following forward programming of the PSC. In an alternative embodiment, the gRNA is introduced during the process of forward programming the PSC into the somatic cell.
[0202] According to a further aspect of the invention, there is provided a method of controlling transcription of an exogenous genetic sequence, in particular a CRISPR enzyme, in an iPSC, comprising:
[0203] (i) administering to the iPSC a chimeric gene comprising a first portion comprising a minigene having an alternatively spliced exon and a second portion that encodes the exogenous genetic sequence, wherein expression of the exogenous genetic sequence is controlled by alternative splicing of the first portion; and
[0204] (ii) applying a splicing modifier that regulates the splicing of the alternatively spliced exon thereby inducing expression of the minigene to include the alternatively spliced exon and initiating translation of the exogenous genetic sequence.
[0205] This aspect, and relevant embodiments, of the invention are useful to perform CRISPR-Cas screens of forward programmed cells. It enables inducing expression of the CRISPR enzyme at desired stages of programming and helps to avoid silencing of the CRISPR enzyme commonly associated with constitutive expression. Therefore, in one embodiment, the method comprises targeting the CRISPR enzyme to an endogenous gene encoding a transcription factor.
[0206] Transcription factors
[0207] In one embodiment, the genetic sequence encodes one or more polypeptides having the activity of one or more transcription factors and / or one or more transcription factors themselves. In one embodiment, the genetic sequence encodes one or more transcription factors.
[0208] The method described herein may comprise increasing the expression (in particular, the protein expression) of a sufficient number of the transcription factors capable of causing differentiation of a cell population to a lineage-restricted cell, therefore differentiating the cell population into a desired cell type, such as a somatic cell. Expression may be of the transcription factors themselves. In the context of the present invention, these factors may also be referred to as “programming factors”. As described herein, the expression of an exogenousBIT-C-P3829PCT
[0209] 29
[0210] or endogenous (in particular an exogenous) transcription factor may be increased. The increased expression may be of one or more exogenous transcription factors, one or more endogenous transcription factors, or a combination of the two. Preferably the increased expression is of exogenous transcription factors.
[0211] Transcription factors used in methods of forward programming to particular cell types have been disclosed in the art. For example, it is known that the forward programming of induced pluripotent stem cells to hepatocytes is possible through the overexpression of polypeptides having transcription factor activity and / or one or more transcription factors, wherein the transcription factors comprise HNF1A, FOXA3, HNF6 and RORc. This is set out in international patent publication WO2023 / 036983 (incorporated herein by reference).
[0212] It is known that the forward programming of induced pluripotent stem cells to microglia is possible through the overexpression of polypeptides having transcription factor activity and / or one or more transcription factors, wherein the transcription factors comprise SPI1 and one additional transcription factor such as a CEB protein, for example CEBPB. This is set out in international patent publication W02020 / 239807 (incorporated herein by reference).
[0213] It is known that the forward programming of induced pluripotent stem cells to GABAergic neurons is possible through the overexpression of polypeptides having transcription factor activity and / or one or more transcription factors, wherein the transcription factors comprise ASCL1 and DLX2. This is set out in international patent publication WO2011 / 091048 (incorporated herein by reference).
[0214] It is known that the forward programming of induced pluripotent stem cells to glutamatergic neurons is possible through the overexpression of polypeptides having Ngn2 activity and / or Ngn2 itself. This is set out in international patent publication WO2011 / 091048 (incorporated herein by reference).
[0215] It is known that the forward programming of induced pluripotent stem cells to sensory neurons, is possible through the overexpression of polypeptides having transcription factor activity and / or one or more transcription factors, wherein the transcription factors comprise NGN1, ISL1 and KLF7.
[0216] It is known that the forward programming of induced pluripotent stem cells to myocytes, such as skeletal myocytes, is possible through the overexpression of polypeptides having MYOD1BIT-C-P3829PCT
[0217] 30
[0218] activity and / or MYOD1 itself. This is set out in international patent publication WO2018 / 096343 (incorporated herein by reference).
[0219] It is known that the forward programming of induced pluripotent stem cells to oligodendrocytes is possible through the overexpression of polypeptides having transcription factor activity and / or one or more transcription factors, wherein the transcription factors comprise OLIG2 and SOX10. This is set out in international patent publication WO2018 / 096343 (incorporated herein by reference).
[0220] The forward programming of induced pluripotent stem cells to adipocytes is possible through the overexpression of polypeptides having transcription factor activity and / or one or more transcription factors, wherein the transcription factors comprise one or more of a PPAR protein, HOXC8, EBF1, EBF2, ZNF467, ZNF423 ora CEB protein. This is set out in international patent application PCT / GB2024 / 051973, published as W020250022129 (incorporated herein by reference).
[0221] The forward programming of induced pluripotent stem cells to pancreatic beta cells is possible through the overexpression of polypeptides having transcription factor activity and / or one or more transcription factors, wherein the transcription factors comprise one or more of GLIS3, PDX1, NEUROD1, NKX6-2 or HNF6. This is set out in international patient application PCT / GB2024 / 052806, published as WO2025099413 (incorporated herein by reference).
[0222] The forward programming of induced pluripotent stem cells to astrocytes is possible through the overexpression of polypeptides having transcription factor activity and / or one or more transcription factors, wherein the transcription factors comprise the combination of SOX9, NFIA, NFIB and one or more transcription factors selected from the group consisting of: FEZF2, TBR1, FOXG1, RORB, LHX2, DBX2.
[0223] The forward programming of induced pluripotent stem cells to regulatory T cells is possible through the overexpression of polypeptides having transcription factor activity and / or one or more transcription factors, wherein the transcription factors comprise GATA3, GFI1, one or more FOXP factors (such as FOXP3 and / or FOXP1, preferably FOXP3), SPIC, one or more IRF factors (such as I RF1 , IRF2, IRF4 and / or IRF5, preferably IRF1), one or more ETS-domain transcription factors (such as ELK3, ETS1, EHF, ELF3, ERG, ETV6 and / or FLI1, preferably ELK3), KLF13, MEF2C, BCL11B, CENPB, IKZF1, MAF, one or more NFATC factors (such as NFATC1 and / or NFATC2), NR4A1, SMAD3, TCF3, TSC22D3, VSX2. This is set out inBIT-C-P3829PCT
[0224] 31
[0225] European patent application EP24213118.3 and international patent application PCT / EP2025 / 052484 (incorporated herein by reference).
[0226] Table 1 below presents names and gene accession numbers (accessed 20 January 2025) for all of the transcription factors mentioned above:
[0227] Table 1:
[0228]
[0229] BIT-C-P3829PCT
[0230] 32
[0231]
[0232] BIT-C-P3829PCT
[0233] 33
[0234]
[0235] In one embodiment, the method comprises expressing one or more polypeptides having the activity of two or more transcription factors, in particular three or more, four or more or five transcription factors and / or increasing the expression of two or more transcription factors, in particular three or more, four or more or five or more transcription factors.
[0236] In one embodiment, the method comprises expressing one or more polypeptides having the activity of less than eight transcription factors, in particular less than seven, less than six, less than five or less than four transcription factors and / or increasing the expression of less than eight transcription factors, in particular less than seven, less than six or less than five transcription factors.
[0237] In one embodiment, the method comprises expressing one or more polypeptides having the activity of between two and eight transcription factors, in particular between three and seven or between four and six transcription factors and / or increasing the expression of between twoBIT-C-P3829PCT
[0238] 34
[0239] and eight transcription factors, in particular between three and seven or between four and six transcription factors.
[0240] Once the activity of a transcription factor is appreciated, the endogenous transcription machinery can be modulated using not only the transcription factors themselves, but also polypeptides engineered to replicate the action of the transcription factor, such as synthetic transcription factors or artificial transcription factors. For example, CRISPR (clustered regularly interspaced palindromic repeats), TALE (transcriptional activator-like effector) or Zinc Finger technologies can be used to modulate the expression of endogenous cellular genes, to allow for faster and more efficient nuclear reprogramming under conditions amenable for clinical and commercial applications. This is set out in, for example, LIS2016 / 362705, incorporated herein by reference.
[0241] Alternatively, as discussed above, with the development of highly accurate protein structure prediction with artificial intelligence tools such as AlphaFold, it is now straightforward for polypeptides to be developed that have very similar structure and / or activity to a transcription factor of interest whilst at the same time having an amino acid sequence that has very little resemblance to that of the transcription factor of interest. For example, large language models trained on biological diversity have been used to develop proteins only around 70% identical to CRISPR-Cas proteins that occur in nature and yet with comparable or improved biological activity and specificity (Ruffolo et al. (2024) bioRxiv, doi: https: / / doi.org / 10.1101 / 2024.04.22.590591). Such polypeptides are covered within the scope of the invention.
[0242] In some embodiments of the present invention, a polypeptide (in particular a single polypeptide) is engineered to mimic the activity of a transcription factor of interest. In a further embodiment, polypeptides having the activity of one or more transcription factors is expressed, in combination with increasing the expression of another transcription factor.
[0243] Methods of the invention encompass the use of variants of the transcription factors. References to the transcription factors also encompasses species variants, isoforms, homologues, allelic forms, mutant forms, and equivalents thereof, including conservative substitutions, additions, deletions therein not adversely affecting the structure and / or function. Changes in the nucleic acid sequence of the transcription factor gene can result in conservative changes or substitutions in the amino acid sequence. Therefore, the invention includes polypeptides having conservative changes or substitutions. The invention includes sequencesBIT-C-P3829PCT
[0244] 35
[0245] where conservative substitutions are made that do not alter the activity of the transcription factor of interest.
[0246] According to a further aspect of the invention, there is provided a method of forward programming a PSC (in particular, an iPSC) into a somatic cell, comprising:
[0247] (i) administering to the PSC a chimeric gene comprising a first portion comprising a minigene having an alternatively spliced exon and a second portion that encodes one or more transcription factors and / or one or more polypeptides having the activity of said one or more transcription factors, wherein expression of the second portion is controlled by alternative splicing of the first portion;
[0248] (ii) applying a splicing modifier that regulates the splicing of the alternatively spliced exon thereby inducing expression of the minigene to include the alternatively spliced exon and initiating translation of the one or more transcription factors and / or polypeptides encoded by the second portion,
[0249] wherein translation of the one or more transcription factors and / or polypeptides encoded by the second portion induces forward programming of the PSC into the somatic cell.
[0250] This aspect, and relevant embodiments, of the invention are useful to orchestrate sequential activations / pauses of transcription for transcription factors, to mimic developmentally regulated gene expression cascades. It enables inducing expression of selected transcription factor(s) at desired stages of programming. In one embodiment, the method comprises using the transcriptionally controlled cassette in combination with an additional controllable transcription method, such as the dual cassette expression system described in WO2018 / 096343, which is incorporated herein by reference and discussed in more detail hereinbelow.
[0251] Endogenous genetic sequences
[0252] Methods of the invention can be used to tune expression levels of endogenous genes. Therefore, in one embodiment, the genetic sequence is an endogenous genetic sequence and the inserted minigene is operably linked to the endogenous genetic sequence present in the PSC genome. This enables expression of the endogenous gene to be controlled through use of the alternatively splicing mechanism.
[0253] Methods of the invention can be used to target genes involved in immune rejection of cellbased therapies. In particular, removal of the human leukocyte antigens (HLA) barrier can be accomplished by one or more of the following: (1) targeting the polymorphic HLA alleles (HLA-BIT-C-P3829PCT
[0254] 36
[0255] A, HLA-B, HLA-C) and major histocompatibility complex (MHC) II genes directly; (2) removal of B2M, which will prevent surface trafficking of all MHC I molecules; and / or (3) deletion of components of the MHC enhanceosomes, such as NLRC5 and CIITA, that are critical for HLA expression. In one embodiment, the method controls the transcription of one or more MHC I and MHC II HLA in the cell.
[0256] In one embodiment, the endogenous gene is the p2-Microglobulin (B2M) gene or a Human Leukocyte Antigen (HLA) gene, optionally selected from HLA-A, HLA-B and HLA-C.
[0257] In one embodiment, the method controls the transcription of one or more genes encoding one or more transcriptional regulators of MHC I or MHC II. In certain embodiments, transcriptional regulators of MHC I or MHC II are selected from the group consisting of B2M, CIITA, NLRC5 and combinations thereof. Beta-2 microglobulin, also known as B2M, is the light chain of MHC class I molecules, and is an integral part of the MHC. In humans, B2M is encoded by the B2M gene which is located on chromosome 15, opposed to the other MHC genes which are located as a gene cluster on chromosome 6. Because B2M is an important structural component of the MHC, inhibition of B2M expression leads to a reduction or elimination of MHC molecules on the surface of the engineered cell (in particular, an engineered T-cell). As a consequence, the engineered cell no longer presents antigens on the surface which are recognized by CD8+ cells. This reduces the risk of rejection of a cell therapy by the host's immune system.
[0258] In the present context, the transcriptionally controlled cassette could be targeted to the B2M gene to disrupt expression. The engineered cell would therefore not express B2M which would prevent recognition of the engineered cell by a patient’s immune system. If the engineered cell caused an adverse event, or the engineered cell needed to be removed from the patient, the splicing modifier could be applied, thereby inducing expression of the B2M protein. This, in turn, would initiate recognition of the engineered cell by the patient’s immune system, which would lead to removal of the engineered cell by the patient’s immune response.
[0259] Humans have three classical MHC la molecules (HLA-A, HLA-B, and HLA-C), which are vital to the detection and elimination of viruses, cancerous cells, and transplanted cells. In addition, there are three non-classical MHC lb molecules (HLA-E, HLA-F, and HLA-G), which have immune regulatory functions. While MHC's serve a vital cellular function, in certain contexts, such as cell-based transplantation therapies, they may also contribute to immune rejection. Therefore, controlling the transcription of a HLA gene, in particular HLA-A, HLA-B and HLA-C, can reduce the risk of rejection of a cell therapy by the host's immune system. Disruption ofBIT-C-P3829PCT
[0260] 37
[0261] the HLA gene can be used in the present context in the same manner as for B2M, as discussed above.
[0262] As described herein, methods of the invention find particular use in forward programming. It will be understood that in one embodiment, the endogenous gene is a transcription factor. Transcription factors are described in more detail hereinbefore. As such, the method can be used to control the transcription of endogenous transcription factors and induce forward programming of the PSC into the somatic cell.
[0263] Delivery of nucleic acid
[0264] Introduction of a nucleic acid, such as DNA or RNA, into cells may use any suitable methods for nucleic acid delivery for transformation of a cell, as described herein or as would be known to one of ordinary skill in the art. Such methods include, but are not limited to, direct delivery of DNA such as by ex vivo transfection, by injection (including microinjection), by electroporation, by calcium phosphate precipitation, by using DEAE-dextran followed by polyethylene glycol, by direct sonic loading, by liposome mediated transfection, by receptor-mediated transfection, by microprojectile bombardment, by agitation with silicon carbide fibers, by Agrobacterium-mediated transformation, and any combination of such methods. Through the application of these techniques, cells may be stably or transiently transformed.
[0265] Further, the genetic material (e.g. the minigene, chimeric gene and / or exogenous genetic sequences as discussed herein) may include cleavable sequences. Such sequences are sequences that are recognised by an entity capable of specifically cutting DNA, and include restriction sites, which are the target sequences for restriction enzymes or sequences for recognition by other DNA cleaving entities, such as nucleases, recombinases, ribozymes or artificial constructs. At least one cleavable sequence may be included, but preferably two or more are present. These cleavable sequences may be at any suitable point in the cassette, such that a selected portion of the cassette, or the entire cassette, can be selectively removed if desired. The cleavable sites may thus flank the part / all of the genetic sequence that it may be desired to remove. The method may therefore also comprise removal of the expression cassette and / or the genetic material.
[0266] VectorsBIT-C-P3829PCT
[0267] 38
[0268] In one embodiment, the genetic material (e.g. the minigene, chimeric gene and / or exogenous genetic sequences as discussed herein) introduced into the cell population using a vector. One of skill in the art would be well equipped to construct a vector through standard recombinant techniques. Vectors include but are not limited to plasmids, cosmids, viruses (bacteriophage, animal viruses, and plant viruses), and artificial chromosomes (e.g., YACs).
[0269] In one embodiment, the vector is a viral vector. The viral gene delivery system may be an RNA-based or DNA-based viral vector. Viral vectors include retroviral vectors, lentiviral vectors (e.g., derived from HIV-1, HIV-2, SIV, BIV, FIV etc.), gammaretroviral vectors, adenoviral (Ad) vectors (including replication competent, replication deficient and gutless forms thereof), adeno-associated virus-derived (AAV) vectors, simian virus 40 (SV-40) vectors, bovine papilloma virus vectors, Epstein-Barr virus vectors, herpes virus vectors, vaccinia virus vectors, Harvey murine sarcoma virus vectors, murine mammary tumour virus vectors, Rous sarcoma virus vectors and Sendai virus vectors. In a further embodiment, the viral vector is selected from: a lentiviral vector, an adeno-associated virus vector or a Sendai virus vector. In a yet further embodiment, the viral vector is a lentiviral vector.
[0270] Lentiviral vectors are well known in the art. Lentiviral vectors are complex retroviruses capable of integrating randomly into the host cell genome, which, in addition to the common retroviral genes gag, pol, and env, contain other genes with regulatory or structural function (e.g., accessory genes Vif, Nef, Vpu, Vpr). Lentiviral vectors have the advantage of being able to infect non-dividing cells and can be used for both in vivo and ex vivo gene transfer and expression of nucleic acid sequences. For example, recombinant lentiviral vector capable of infecting a non-dividing cell wherein a suitable host cell is transfected with two or more vectors carrying the packaging functions, namely gag, pol and env, as well as rev and tat.
[0271] In one embodiment, a nucleic acid sequence encoding the genetic material is introduced into a cell by a plasmid.
[0272] In one embodiment, the plasmid is episomal. Episomal vectors are able to introduce large fragments of DNA into a cell but are maintained extra-chromosomally, replicated once per cell cycle, partitioned to daughter cells efficiently, and elicit substantially no immune response. In alternative embodiments, an Epstein-Barr virus (EBV)-based episomal vector, a yeast-based vector, an adenovirus-based vector, a simian virus 40 (SV40)-based episomal vector, or a bovine papilloma virus (BPV)-based vector may be used.BIT-C-P3829PCT
[0273] 39
[0274] Site-;
[0275]
[0276] Any suitable technique for insertion of a nucleic acid sequence into a specific sequence may be used, and several are described in the art. Suitable techniques include any method which introduces a break at the desired location and permits recombination of the vector into the gap. Thus, a crucial first step for targeted site-specific genomic modification is the creation of a double-strand DNA break (DSB) at the genomic locus to be modified. Distinct cellular repair mechanisms can be exploited to repair the DSB and to introduce the desired sequence, and these are non-homologous end joining repair (NHEJ), which is more prone to error; and homologous recombination repair (HR).
[0277] Several techniques exist to allow customized site-specific generation of DSB in the genome. Many of these involve the use of customized endonucleases, such as zinc finger nucleases, TALENs or the clustered regularly interspaced short palindromic repeats / CRISPR associated protein (CRISPR / Cas, e.g. CRISPR / Cas9) system.
[0278] In one embodiment, the CRISPR / Cas system is used for targeted insertion. Features of the CRISPR / Cas system are described hereinbefore.
[0279] Zinc finger nucleases are artificial enzymes which are generated by fusion of a zinc-finger DNA-binding domain to the nuclease domain of the restriction enzyme Fokl. The latter has a non-specific cleavage domain which must dimerise in order to cleave DNA. This means that two zinc finger nuclease monomers are required to allow dimerisation of the Fokl domains and to cleave the DNA. The DNA binding domain may be designed to target any genomic sequence of interest, is a tandem array of Cys2His2 zinc fingers, each of which recognises three contiguous nucleotides in the target sequence. The two binding sites are separated by 5-7bp to allow optimal dimerization of the Fokl domains. The enzyme thus is able to cleave DNA at a specific site, and target specificity is increased by ensuring that two proximal DNA-binding events must occur to achieve a double-strand break.
[0280] Transcription activator-like effector nucleases, or TALENs, are dimeric transcription factor / nucl eases. They are made by fusing a TAL effector DNA-binding domain to a DNA cleavage domain (a nuclease). Transcription activator-like effectors (TALEs) can be engineered to bind practically any desired DNA sequence, so when combined with a nuclease, DNA can be cut at specific locations. TAL effectors are proteins that are secreted by Xanthomonas bacteria, the DNA binding domain of which contains a repeated highlyBIT-C-P3829PCT
[0281] 40
[0282] conserved 33-34 amino acid sequence with divergent 12th and 13th amino acids. These two positions are highly variable and show a strong correlation with specific nucleotide recognition. This straightforward relationship between amino acid sequence and DNA recognition has allowed for the engineering of specific DNA-binding domains by selecting a combination of repeat segments containing appropriate residues at the two variable positions. TALENs are thus built from arrays of 33 to 35 amino acid modules, each of which targets a single nucleotide. By selecting the array of modules, almost any sequence may be targeted. Again, the nuclease used may be Fokl or a derivative thereof.
[0283] The elements for making the double-strand DNA break may be introduced in one or more vectors, such as plasmids, for expression in the cell.
[0284] Thus, any method of making specific, targeted double strand breaks in the genome in order to effect the insertion of a genetic material may be used in the method of the invention. It may be preferred that the method for inserting the genetic material utilises any one or more of zinc finger nucleases, TALENs and / or CRISPR / Cas9 systems or any derivative thereof.
[0285] Once the DSB has been made by any appropriate means, the cassette for insertion may be supplied in any suitable fashion as described below. The cassette and associated genetic material form the donor DNA for repair of the DNA at the DSB and are inserted using standard cellular repair machinery / pathways. How the break is initiated will alter which pathway is used to repair the damage, as noted above.
[0286] Other methods in the art for site specific delivery include the use of homologous recombination (HR) and recombinase mediated cassette exchange (RMCE). DNA damage mediated site specific insertion methods (such as CRISPR / Cas) can also be used to perform site specific integration of DNA recognition sequences fatt' sites) which in turn mediate site specific insertion via the activity of tyrosine and serine recombinases or integrases. These sites (e.g. attP) once inserted into the genome, can mediate site specific HR and RMCE. Insertion of exogenous nucleic acid sequences occurs through homologous recombination between cognate attP and attB sites mediated by the expression of the appropriate and cognate recombinase (e.g. Flp, Cre) or integrase (PhiC31, Bxb1). Using targeting vectors, as described above, flanked by attB sites, site specific exogenous DNA insertion of transgenes can be achieved.BIT-C-P3829PCT
[0287] 41
[0288] In one embodiment, the minigene is inserted into a genomic safe harbour (GSH) site in the PSC genome.
[0289] A GSH site is a locus within the genome wherein a gene or other genetic material may be inserted without any deleterious effects on the cell or on the inserted genetic material. Most beneficial is a GSH site in which expression of the inserted gene sequence is not perturbed by any read-through expression from neighbouring genes and expression of the inducible cassette minimizes interference with the endogenous transcription programme. More formal criteria have been proposed that assist in the determination of whether a particular locus is a GSH site in future (Papapetrou etal. (2011) Nature Biotechnology, 29(1): 73-8) These criteria include a site that is (i) 50 kb or more from the 5’ end of any gene, (ii) 300 kb or more from any gene related to cancer, (iii) 300 kb or more from any microRNA (miRNA), (iv) located outside a transcription unit and (v) located outside ultraconserved regions (UCR). It may not be necessary to satisfy all of these proposed criteria, since GSH sites already identified do not fulfil all of the criteria. It is thought that a suitable GSH site will satisfy at least 2, 3, 4 or all of these criteria. Any suitable GSH site may be used in the method of the invention, on the basis that the site allows insertion of genetic material without deleterious effects to the cell and permits transcription of the inserted genetic material. Those skilled in the art may use these simplified criteria to identify a suitable GSH site, and / or the more formal criteria set out above.
[0290] Insertion of the genetic material may be carried out through direct delivery methods as described above. It is understood that although such direct delivery methods may lead to the random insertion of the genetic material, screening may be carried out in order to identify clones that show no deleterious effects and are able to express the inserted genetic material, and by doing so one is able to confirm that the genetic material has been inserted into a GSH site.
[0291] In one embodiment the insertion of the genetic material is targeted. “Targeted insertion”, as with site-specific delivery, is understood as the insertion of the genetic material into a prechosen GSH site. As discussed above, this can be carried out using techniques known in the art such as zinc finger nucleases, TALENs, the clustered regularly interspaced short palindromic repeats / CRISPR associated protein (CRISPR / Cas, e.g. CRISPR / Cas9) system or recombinase / integrase-based systems.
[0292] In one embodiment, the GSH sites are selected from the ROSA26 locus, the AAVS1 locus, the CLYBL gene, the CCR5 gene or the HPRT gene. Insertions specifically within GSH sites isBIT-C-P3829PCT
[0293] 42
[0294] preferred over random genome integration, since this is expected to be a safer modification of the genome, and is less likely to lead to unwanted side effects such as silencing natural gene expression or random insertional mutagenesis.
[0295] The adeno-associated virus integration site 1 locus (AAVS1) is located within the protein phosphatase 1, regulatory subunit 12C (PPP1R12C) gene on human chromosome 19, which is expressed uniformly and ubiquitously in human tissues. AAVS1 has been shown to be a favourable environment for transcription, since it comprises an open chromatin structure and native chromosomal insulators that enable resistance of the inducible cassettes against silencing. There are no known adverse effects on the cell resulting from disruption of the PPP1R12C gene. Moreover, an inducible cassette inserted into this site remains transcriptionally active in many diverse cell types.
[0296] The hROSA26 site has been identified on the basis of sequence analogy with a GSH site from mice (ROSA26 - reverse oriented splice acceptor site #26). The hROSA26 locus is on chromosome 3 (3p25.3), and can be found within the Ensembl database (GenBank:CR624523). The integration site lies within the open reading frame (ORF) of the THUMPD3 long non-coding RNA (reverse strand). Since the hROSA26 site has an endogenous promoter, the inserted genetic material may take advantage of that endogenous promoter, or alternatively may be inserted operably linked to a promoter.
[0297] Intron 2 of the Citrate Lyase Beta-like (CLYBL) gene, on the long arm of Chromosome 13, was identified as a suitable GSH site since it is one of the identified integration hot-spots of the phage derived phiC31 integrase. Studies have demonstrated that randomly inserted inducible cassettes into this locus are stable and expressed. It has been shown that insertion of inducible cassettes at this GSH site do not perturb local gene expression (Cerbini et al. (2015) PLOS One, 10(1): e0116032). CLYBL thus provides a GSH site which may be suitable for use in the present invention.
[0298] CCR5, which is located on chromosome 3 (position 3p21.31) is a gene which codes for HIV-1 major co-receptor. Interest in the use of this site as a GSH site arises from the null mutation in this gene that appears to have no adverse effects, but predisposes to HIV-1 infection resistance. Zinc-finger nucleases that target the third exon have been developed, thus allowing for insertion of genetic material at this locus.BIT-C-P3829PCT
[0299] 43
[0300] The hypoxanthine-guanine phosphoribosyltransferase (HPRT) gene encodes a transferase enzyme that plays a central role in the generation of purine nucleotides through the purine salvage pathway.
[0301] Other GSH sites have been described in the art, such as in Sadelain et al. (2012) Nature Reviews 12:51-58 and in WO2021 / 152086, which are herein incorporated by reference.
[0302] GSH sites in other organisms have been identified and include ROSA26, HRPT and Hippl 1 (H11) loci in mice. Mammalian genomes may include GSH sites based upon pseudo attP sites. For such sites, hiC31 integrase, the Streptomyces phage-derived recombinase, has been developed as a non-viral insertion tool, because it has the ability to integrate an inducible cassette-containing plasmid carrying an attB site into pseudo attP sites.
[0303] Technically, the insertion into the GSH site may occur on one chromosome, or on both chromosomes. The GSH site exists at the same genetic loci on both chromosomes of diploid organisms. Insertion within both chromosomes is advantageous since it may enable an increase in the level of transcription from the inserted genetic material within the inducible cassette, thus achieving particularly high levels of transcription.
[0304] Specific insertion of genetic material into the particular GSH site based upon customised sitespecific generation of DNA double-strand breaks at the GSH site may be achieved. The genetic material may then be introduced using any suitable mechanism, such as homologous recombination. Any method of making a specific DSB in the genome may be used, but preferred systems include CRISPR / Cas9 and modified versions thereof, zinc finger nucleases and the TALEN system, or via HR or ROME mediated integration or recombination.
[0305] Controlled expression
[0306] An (exogenous) expression cassette carrying the minigene may comprise an expression control element. In one embodiment, the expression control element is a constitutive promoter, such as a CAG promoter. In an alternative embodiment, the expression control element is an inducible promoter, such as a Tet Responsive Element (TRE).
[0307] In one embodiment, the expression control element is an externally inducible transcriptional regulatory element (i.e., an inducible promoter). This enables rapid induction of transcription in response to extremal stimuli, i.e. inducible gene (or transgene) expression. The presence orBIT-C-P3829PCT
[0308] 44
[0309] addition of the appropriate external stimuli (e.g. protein, compound or chemical) to cell culture media modulates the controlled expression of the genetic sequence within the inducible expression cassette; and may be administered continuously or transiently to modulate transcription as required.
[0310] In one embodiment, the cell additionally comprises an exogenous transcriptional regulator protein. In a further embodiment, the transcriptional regulator protein is rtTA.
[0311] A “transcriptional regulator protein” is a protein that binds to DNA, preferably sequence-specifically to a DNA site, and either facilitating the transcription of the DNA sequence (a transcriptional activator) or blocks this process (a transcriptional repressor).
[0312] The DNA sequence that a transcriptional regulator protein binds to is called a transcription factor-binding site or response element, and these are found in or near the promoter of the regulated DNA sequence. Transcriptional activator proteins bind to the response element and promote gene expression. T ranscriptional repressor proteins bind to the response element and prevent gene expression.
[0313] T ranscriptional regulator proteins may be activated or deactivated by a number of mechanisms including binding of a substance, protein conformation changes controlled by light, interaction with other transcription factors (e.g., homo- or hetero-dimerization) or coregulatory proteins, phosphorylation, and / or methylation. The transcriptional regulator protein may be controlled by activation or deactivation.
[0314] If the transcriptional regulator protein is a transcriptional activator protein, it is preferred that the transcriptional activator protein requires activation. This activation may be through any suitable means, but it is preferred that the transcriptional regulator protein is activated through the addition to the cell of an exogenous substance. The supply of an exogenous substance to the cell can be controlled, and thus the activation of the transcriptional regulator protein can be controlled. Alternatively, an exogenous substance can be supplied in order to deactivate a transcriptional regulator protein, and then withdrawn in order to activate the transcriptional regulator protein.
[0315] If the transcriptional regulator protein is a transcriptional repressor protein, it is preferred that the transcriptional repressor protein requires deactivation. Thus, a substance is supplied toBIT-C-P3829PCT
[0316] 45
[0317] prevent the transcriptional repressor protein repressing transcription, and thus transcription is permitted.
[0318] Any suitable transcriptional regulator protein may be used, preferably one that may be activated or deactivated. It is preferred that an exogenous substance may be supplied to control the transcriptional regulator protein. Such transcriptional regulator proteins are also called inducible transcriptional regulator proteins.
[0319] Tetracycline-Controlled Transcriptional Activation is a method of inducible gene expression where transcription is reversibly turned on or off in the presence of the antibiotic tetracycline or one of its derivatives (e.g., doxycycline which is more stable). In this system, the transcriptional activator protein is reverse tetracycline-controlled transactivator (rtTa, which may also be referred to as tetracycline - responsive transcriptional activator protein) or a derivative thereof. The rtTA protein is able to bind to DNA at specific TetO operator sequences. Several repeats of such TetO sequences are placed upstream of a minimal promoter (such as the CMV promoter), which together form a tetracycline response element (TRE). There are two forms of this system, depending on whether the addition of tetracycline or a derivative activates (Tet-On) or deactivates (Tet-Off) the rtTA protein.
[0320] In a Tet-Off system, tetracycline or a derivative thereof binds rtTA and deactivates the rtTA, rendering it incapable of binding to TRE sequences, thereby preventing transcription of TRE-controlled genes. This system was first described in Gossen etal. (1992) PNAS 89 (12): 5547-5551.
[0321] The Tet-On system is composed of two components; (1) the constitutively expressed reverse tetracycline-controlled transactivator (rtTa) and the rtTa-sensitive inducible promoter (Tet Responsive Element, TRE). This may be bound by tetracycline or its more stable derivatives, including doxycycline (dox), resulting in activation of rtTa, allowing it to bind to TRE sequences and inducing expression of TRE-controlled genes.
[0322] The transcriptional regulator protein may thus be a reverse tetracycline-controlled transactivator (rtTa) protein, which can be activated or deactivated by the antibiotic tetracycline or one of its derivatives, which are supplied exogenously. If the transcriptional regulator protein is rtTA, then the expression control element includes the tetracycline response element (TRE). The exogenously supplied substance is the antibiotic tetracycline or one of its derivatives.BIT-C-P3829PCT
[0323] 46
[0324] Variants and modified rtTa proteins may also be used in the methods of the invention, these include Tet-On Advanced transactivator (also known as rtTA2S-M2) and Tet-On 3G (also known as rtTA-V16, derived from rtTA2S-S2).
[0325] The tetracycline response element (TRE) generally consists of 7 repeats of the 19bp bacterial TetO sequence separated by spacer sequences, together with a minimal promoter. Variants and modifications of the TRE sequence are possible, since the minimal promoter can be any suitable promoter. Preferably the minimal promoter shows no or minimal expression levels in the absence of rtTa binding. The expression control element (in particular, an inducible promoter) may thus comprise a TRE.
[0326] A modified system based upon tetracycline control is the T-REx™ System (ThermoFisher Scientific), in which the transcriptional regulator protein is a transcriptional repressor protein, TetR. The components of this system include (i) an inducible promoter comprising a strong human cytomegalovirus immediate-early (CMV) promoter and two tetracycline operator 2 (TetO2) sites, and a Tet repressor (TetR). The TetO2 sequences consist of 2 copies of the 19 nucleotide sequence, 5’-TCCCTATCAGTGATAGAGA-3’ (SEQ ID NO: 3) separated by a 2 base pair spacer. In the absence of tetracycline, the Tet repressor forms a homodimer that binds with extremely high affinity to each TetO2 sequence in the inducible promoter, and prevents transcription from the promoter. Once added, tetracycline binds with high affinity to each Tet repressor homodimer rendering it unable to bind to the Tet operator. The Tet repressor: tetracycline complex then dissociates from the Tet operator and allows induction of expression. In this instance, the transcriptional regulator protein is TetR and the inducible promoter comprises two TetO2 sites. The exogenously supplied substance is tetracycline or a derivative thereof.
[0327] The invention further relates to a codon-optimised tetR (OPTtetR). This may be used in any method described herein, or for any additional use where inducible promotion is desirable. This entity was generated using multiparameter-optimisation of the bacterial tetR cDNA sequence. OPTtetR allows a ten-fold increase in the tetR expression when compared to the standard sequence (STDtetR). Homozygous OPTtetR expression of tetR was sufficient to prevent shRNA leakiness whilst preserving knockdown induction in the Examples. The sequence for OPTtetR is included here, with the standard sequence shown as a comparison. Sequences with at least 75%, 80%, 85% or 90% homology for this sequence are hereby claimed, more particularly 91, 92, 93, 94, 95, 96, 97 or 99% homology. Residues shown to be changed between STDtetR and OPTtetR have been indicated in the sequences, and it is preferred thatBIT-C-P3829PCT
[0328] 47
[0329] these residues are not changed in any derivative of OPTtetR since these are thought to be important for the improved properties. Any derivative would optionally retain these modifications at the indicated positions.
[0330] Other inducible expression systems are known and can be used in the method of the invention. These include the Complete Control Inducible system from Agilent Technologies. This is based upon the insect hormone ecdysone or its analogue ponasterone A (ponA) which can activate transcription in mammalian cells which are transfected with both the gene for the Drosophila melanogaster ecdysone receptor (EcR) and an inducible promoter comprising a binding site for the ecdysone receptor. The EcR is a member of the retinoid-X-receptor (RXR) family of nuclear receptors. In humans, EcR forms a heterodimer with RXR that binds to the ecdysoneresponsive element (EcRE). In the absence of PonA, transcription is repressed by the heterodimer.
[0331] Thus, the transcriptional regulator protein can be a repressor protein, such as an ecdysone receptor or a derivative thereof. Examples of the latter include the VgEcR synthetic receptor from Agilent technologies which is a fusion of EcR, the DNA binding domain of the glucocorticoid receptor and the transcriptional activation domain of Herpes Simplex Virus VP16. The expression control element (e.g. inducible promoter) comprises the EcRE sequence or modified versions thereof together with a minimal promoter. Modified versions include the E / GRE recognition sequence of Agilent Technologies, in which mutations to the sequence have been made. The E / GRE recognition sequence comprises inverted half-site recognition elements for the retinoid-X-receptor (RXR) and GR binding domains. In all permutations, the exogenously supplied substance is ponasterone A, which removes the repressive effect of EcR or derivatives thereof on the inducible promoter, and allows transcription to take place.
[0332] Alternatively, inducible systems may be based on the synthetic steroid mifepristone as the exogenously supplied substance. In this scenario, a hybrid transcriptional regulator protein is inserted, which is based upon a DNA binding domain from the yeast GAL4 protein, a truncated ligand binding domain (LBD) from the human progesterone receptor and an activation domain (AD) from the human NF-KB. This hybrid transcriptional regulator protein is available from Thermo-Fisher Scientific (Gene Switch™). Mifepristone activates the hybrid protein, and permits transcription from the inducible promoter which comprises GAL4 upstream activating sequences (UAS) and the adenovirus E1b TATA box. This system is described in Wang etal. (1994) PNAS 91: 8180-8184.BIT-C-P3829PCT
[0333] 48
[0334] The transcriptional regulator protein can thus be any suitable regulator protein, either an activator or repressor protein. Suitable transcriptional activator proteins are tetracyclineresponsive transcriptional activator protein or the Gene Switch hybrid transcriptional regulator protein. Suitable repressor proteins include the Tet-Off version of rtTA, TetR or EcR. The transcriptional regulator proteins may be modified or derivatised as required.
[0335] The expression control element can comprise elements which are suitable for binding or interacting with the transcriptional regulator protein. The interaction of the transcriptional regulator protein with the expression control element is preferably controlled by the exogenously supplied substance or light.
[0336] The exogenously supplied substance can be any suitable substance that binds to or interacts with the transcriptional regulator protein. Suitable substances include tetracycline (or derivatives thereof, such as doxycycline), ponasterone A and mifepristone.
[0337] Alternatively, the cumate system may be used as the transcriptional regulator system. The cumate system is derived from the regulatory mechanisms of bacterial operons (cmt and cym) to regulate gene expression. Regulation is mediated by the binding of the repressor (CymR) to the operator site (CuO). Addition of cumate, a small molecule, removes the CymR repressor from the CuO operator site, thus allowing the expression of genetic material to take place.
[0338] Alternatively, the transcriptional regulator system may be based on the hepatitis C virus (HCV) NS3 protease domain. NS3 is a serine c / s-protease that excises itself from the HCV polyprotein by cleaving recognition sites that flank it at either end. Because it is essential for HCV replication, numerous inhibitors targeting the viral protease have been developed, such as danoprevir and grazoprevir. The protease has been used as a ligand-inducible connection to control the association between modular DNA-binding and transcriptional activation domains. In one embodiment, the protease is inserted between minimal DNA-binding and transcriptional activator sequences. In this configuration, the viral protease would serve as a self-immolating connection, excising itself from the fusion construct and, in doing so, separating the DNA-binding and transcriptional activator elements. However, in the presence of an NS3 inhibitor, self-excision of the protease would be blocked, resulting in the preservation of full-length gene capable of activating the expression of targeted genes. These systems may be developed so that one NS3 protease regulates transcription, but may also involve the use of more than one NS3 protease, such as NS3 variants binding danoprevir and / or grazoprevir.BIT-C-P3829PCT
[0339] 49
[0340] The transcriptional regulator protein may be activated or deactivated by light. Such proteins are described in, for example, W02023 / 004031 (incorporated herein by reference). The protein may comprise a light-activatable domain that responds to light of a particular wavelength. In some cases, the light-activatable domain, upon stimulation with light of a particular wavelength or within a particular spectral range, dimerizes or oligomerizes (e.g., with another light activatable domain). In some cases, the light-activatable domain may form a homodimer or a heterodimer (e.g., may dimerize with a second, different light-activatable domain). In some cases, the light-activatable domain may exist in a (e.g., homo or hetero) dimer or (e.g., homo or hetero) oligomer (e.g., in the absence of light), and may dissociate into a monomeric form after exposure to light. The light-activatable domain may be derived from a natural source (e.g., a naturally occurring protein) or may be synthetically produced. The light-activatable domain may comprise or may be a functional domain or portion of a naturally occurring protein, such as, by way of example only, the PHR domain of Arabidopsis cryptochrome 2. The light-activatable domain may comprise an amino acid sequence identical to an amino acid sequence of a wildtype protein, or may comprise one or more variants (e.g., amino acid substitutions, deletions, insertions, etc.) relative to a wild-type protein. This domain may activate or deactivate the transcriptional regulator protein.
[0341] In various aspects, a combination of light-activatable domains (e.g., a first light activatable domain and a second light-activatable domain) may be used. In this scenario, the first light-activatable domain and the second light-activatable domain are binding partners, such that upon illumination with light at a particular wavelength or within a particular spectral range, the first and second light-activatable domains heterodimerize or heterooligomerize. This heterodimerization or heterooligomerization may activate or deactivate the transcriptional regulator protein.
[0342] In various aspects, the light-activatable domain comprises a Light-Oxygen-Voltage (LOV) photoreceptor domain, a LOV2 photoreceptor domain, a Cryptochrome (CRY) domain, Blue-light-using FAD (BLLIF) photoreceptor domain, a Phytochrome (PHY) domain, a CIBN (N-terminal domain of CIB1 (cryptochrome-interacting basic-helix-loop-helix protein 1)) domain, a PIF (phytochrome interacting factor) domain, a Dronpa domain, a LIVR8 photoreceptor domain, a COP1 domain, a BphP1 domain, a QPAS-1 domain, a cobalamin binding domain (CBD), ora combination thereof. In one example, the light-activatable domain is a LOV domain (e.g., such as a LOV domain derived from Vaucheria frigida Aureochrome 1).BIT-C-P3829PCT
[0343] 50
[0344] In some instances, a combination of light-activatable domains is used, wherein the first light-activatable domain is cryptochrome 2 (or a variant or a functional portion thereof) and the second-light activatable domain is CIBN (or a variant or a functional portion thereof). In some instances, a combination of light-activatable domains is used, wherein the first light-activatable domain is BphP1 (or a variant or a functional portion thereof) and the second-light activatable domain is QPAS1 (or a variant or a functional portion thereof).
[0345] The light-activatable domain may be fused with a domain that, upon activation, interacts with an expression control element (e.g. an inducible promoter) to allow expression of the genetic material to take place. Alternatively the light-activatable domain may be fused with a domain that, upon activation, interacts with a recombinase that removes a blocker sequence, allowing expression of the genetic material to take place. The advantage of the combination with a recombinase is that only a short burst of light is necessary in order to switch on expression, rather than constant light exposure.
[0346] Alternatively the transcriptional regulator protein may be activated or deactivated by a change in temperature. In this scenario, the cells may be exposed to a period of temperature increase or temperature decrease in order to either activate or deactivate the transcriptional regulator protein. The extent of the temperature change and the length of time the cells are exposed to the temperature increase / decrease may be optimised so that cell viability is not detrimentally affected. It is known in the art that some heat shock proteins are expressed only when a temperature increase and / or decrease takes place, and this mechanism could be utilised in a transcriptional regulator protein.
[0347] It is preferred that the gene encoding the transcriptional regulator protein is operably linked to a constitutive promoter. Alternatively, the integration site can be selected such that it already has a constitutive promoter that can also drive expression of the transcriptional regulator protein gene and any associated genetic material. Constitutive promoters ensure sustained and high-level gene expression. Commonly used constitutive promoters, including the human P-actin promoter (ACTB), cytomegalovirus (CMV), elongation factor-1a, (EF1a), phosphoglycerate kinase (PGK) and ubiquitin C (UbC). The CAG promoter is a strong synthetic promoter frequently used to drive high levels of gene expression and was constructed from the following sequences: (C) the cytomegalovirus (CMV) early enhancer element, (A) the promoter, the first exon and the first intron of chicken beta-actin gene, and (G) the splice acceptor of the rabbit beta-globin gene.BIT-C-P3829PCT
[0348] 51
[0349] In one embodiment, the cell comprises an additional, exogenous transcriptional regulator protein. In a further embodiment, the transcriptional regulator protein is rtTA.
[0350] The transcriptional regulator protein may be present in a dual cassette expression system, such as the system described in WO2018 / 096343, which is incorporated herein by reference. In this instance, induced transgene over-expression is achieved by using the Tet-ON system components with transgene expression controlled by doxycycline. The components are split between two genomic (or genetic) safe harbour sites (GSH) to reduce the risk of epigenetic gene silencing. The components are (i) transcriptional activator protein (reverse tetracycline trans-activator (rtTA)), which in the presence of doxycycline binds (ii) tetracycline response element (TRE; multiple TetO repeat sequences & minimal Cytomegalovirus (CMV) promoter). TRE binding by rtTA trans-activates transgene expression. Trans-activatable coding sequences for transgenes may be of human origin.
[0351] In one embodiment, the first and second GSH sites are selected from (in particular any two) of the ROSA26 locus, the AAVS1 locus, the CLYBL gene, the CCR5 gene or the HPRT gene. Insertions specifically within GSH sites is preferred over random genome integration, since this is expected to be a safer modification of the genome, and is less likely to lead to unwanted side effects such as silencing natural gene expression or random insertional mutagenesis.
[0352] Technically, the insertions into the first and / or second GSH sites may occur on one chromosome, or on both chromosomes. The GSH site exists at the same genetic loci on both chromosomes of diploid organisms. Insertion within both chromosomes is advantageous since it may enable an increase in the level of transcription from the inserted genetic material within the inducible cassette, thus achieving particularly high levels of transcription.
[0353] Therefore, in one embodiment, the cell additionally comprises:
[0354] (a) an inserted genetic sequence encoding a transcriptional regulator protein at a first genomic safe harbour (GSH) site; and
[0355] (b) an inserted inducible cassette comprising a second exogenous genetic sequence (e.g. one or more transcription factors) operably linked to an inducible promoter at a second GSH site, wherein said inducible promoter is regulated by the transcriptional regulator protein.
[0356] The dual cassette expression may be used to encode one or more transcription factors, in particular for use in forward programming an iPSC to a somatic cell. Therefore, in oneBIT-C-P3829PCT
[0357] 52
[0358] embodiment, a sequence encoding one or more (e.g., two or more or three or more) of the transcription factors is also introduced into the cell population using a method comprising:
[0359] - targeted insertion of a coding sequence for a transcriptional regulator protein into a first genomic safe harbour site of a source cell present in the cell population; and
[0360] - targeted insertion of an inducible cassette into a second genomic safe harbour site of the source cell, wherein said inducible cassette comprises said sequence encoding one or more transcription factors operably linked to an inducible promoter, and said promoter is regulated by the transcriptional regulator protein.
[0361] The inserted genetic sequence encoding a transcriptional regulator protein in the first GSH site provides the control mechanism for the expression of the inducible cassette which is operably linked to the inducible promoter and inserted in a second GSH site. In one embodiment, the first and second GSH site are different.
[0362] Alternatively, the dual expression cassette system utilises different alleles of the same GSH site. In this embodiment, the inducible cassette may be inserted into one allele of the GSH site and the system controlling the expression of the inducible cassette into the other allele of the GSH site (e.g. as described in DeKelver et al., 2010, Genome Res., 20, 1133-43 and Qian et al., 2014, Stem Cells, 32, 1230-8).
[0363] Preferably, the inducible promoter used in the second GSH site is different to the expression control element used in the transcriptionally controlled cassette (i.e. linked to the minigene).
[0364] One or more genetic sequences may be controllably transcribed from within the second and / or further GSH site. Indeed, the inducible cassette may contain 1 , 2, 3, 4, 5, 6, 7, 8, 9 or 10 genetic sequences (e.g., transcription factor sequences) which it is desired to insert into the GSH site and the transcription of which is to be controllably induced. Therefore, the genetic sequences (e.g. one or more transcription factors) may be included within the same cassette introduced into the second genetic safe harbour site. For example, the three or more transcription factors may be included in, for example, three mono-cistronic constructs, one mono-cistronic and one bi-cistronic construct or one tri-cistronic construct. It will be understood that similar combinations of constructs may be used to achieve higher orders of transcription factor expression.
[0365] Alternatively, if a combination of transcription factors is used, the individual transcription factors may be introduced into separate GSH sites, and the transcription of these transcription factorsBIT-C-P3829PCT
[0366] 53
[0367] may be regulated by the transcriptional regulator protein. Therefore, in one embodiment, at least three or more transcription factors are introduced into separate GSH sites. This may be achieved by utilising three or more different GSH sites for the three or more transcription factors (i.e., wherein the transcription factors are introduced as mono-cistronic cassettes). Alternatively, this may be achieved by utilising the fact that a GSH site exists at the same genetic loci on both chromosomes of diploid organisms, e.g., introducing one transcription factor into the GSH site on one chromosome and a different transcription factor into the same GSH site on the other chromosome. This embodiment is advantageous if different expression levels or timing of expression of the transcription factors is desired. In one embodiment, the method comprises targeted insertion of the transcription factors into a second, third and fourth genomic safe harbour site of the source cell. The transcription of the transcription factors may all be regulated by the transcriptional regulator protein.
[0368] Obtaining somatic cells
[0369] In one embodiment, the method additionally comprises monitoring the cell population for at least one characteristic of the desired cell type, e.g. the somatic cell of interest. Cells may be monitored throughout culturing to identify expression of key lineage markers.
[0370] For example, monitoring may be through the use of engineered ‘reporter’ cell lines (i.e. endogenously tagged proteins or positive selection markers under the control of promoters specific to the cells of interest) or immunostaining and detection, using fluorescence microscopy or flow cytometry. Such material includes genes for markers or reporter molecules, such as genes that induce visually identifiable characteristics including fluorescent and luminescent proteins. Examples include the gene that encodes jellyfish green fluorescent protein (GFP), which causes cells that express it to glow green under blue / UV light, luciferase, which catalyses a reaction with luciferin to produce light, and the red fluorescent protein from the gene dsRed.
[0371] The expression cassette containing the minigene may further comprise a positive selection marker and / or selectable reporter expression cassette, e.g., comprising a promoter specific to the cells of interest operably linked to a reporter gene.
[0372] Selectable markers may include resistance genes to antibiotics or other drugs. Examples of drug resistance genes may include: a puromycin resistance gene, an ampicillin resistance gene, a neomycin resistance gene, a tetracycline resistance gene, a kanamycin resistanceBIT-C-P3829PCT
[0373] 54
[0374] gene or a chloramphenicol resistance gene. Cells can be cultured on a medium containing the appropriate drug (i.e., a selection medium) and only those cells which incorporate and express the drug resistance gene will survive. Therefore, by culturing cells using a selection medium, it is possible to select for cells comprising and expressing a drug resistance gene, positively enriching for a target cell population.
[0375] Examples of fluorescent protein genes which may be used as markers include: a green fluorescent protein (GFP) gene, yellow fluorescent protein (YFP) gene, red fluorescent protein (RFP) gene or aequorin gene. Cells expressing the fluorescent protein can be detected using a fluorescence microscope and fluorescence activated cell sorting (FACS) used to identify and select cell populations based on the expression of fluorescent proteins.
[0376] Fluorescent protein genes may be tagged with a nuclear localization signal peptide to confine expression of the fluorescent proteins to the nucleus. This may be helpful in cell types with a high lipid content which may not be suitable for FACS. This allows end-point fluorescence-activated cell sorting to be carried out on either whole cell populations, or purified nuclei which maintain an intact fluorescent signal.
[0377] Examples of chromogenic enzyme genes which may be used as markers, and known in the art, include but are not limited to: p-galactosidase gene, p-glucuronidase gene, alkaline phosphatase gene, or secreted alkaline phosphatase SEAP gene. Cells expressing these chromogenic enzyme genes can be detected by applying the appropriate chromogenic substrate (e.g., X-gal for p galactosidase) so that cells expressing the marker gene will produce a detectable colour (e.g., blue in a blue-white screen test).
[0378] The method may therefore comprise a selection or enrichment step for the desired cells provided from the methods described herein. In one embodiment, the method comprises the step of sorting the cells using fluorescence activated cell sorting (FACS) or immunomagnetic sorting methods based on the expression of cell markers specific to the desired cells and / or absence of cell markers not indicative of the desired cells (such as pluripotency markers). In one embodiment, where the desired cells are lineage-restricted cells or somatic cells, they are selected or enriched by removing cells that express pluripotency markers, such as Ki67, TRA-1-60, SOX2, OCT4, NANOG or SSEA4, in particular Ki67, NANOG or POU5F1.
[0379] A labelled binding agent directed to target cell surface proteins may be used. Any binding agent capable of specific binding to a particular epitope may be used for this purpose, for exampleBIT-C-P3829PCT
[0380] 55
[0381] an antibody or a fragment thereof, a peptide or a synthetic binder such as a plastic antibody, or an aptamer or oligonucleotide, capable of specific binding to an epitope. The binding agent may be labelled with a detectable marker, such as a luminescent, fluorescent (e.g. fluorochrome), enzyme or radioactive marker; alternatively or additionally an affinity tag, e.g. a biotin, avidin, streptavidin or His (e.g. hexa-His) tag. In one embodiment, fluorochrome conjugated antibodies targeting cell surface proteins may be used to sort target cells.
[0382] In another embodiment, the desired cells are enriched by drug-resistance selection from genetically engineered source cells expressing an antibiotic-resistance gene under the control of a cell-specific promoter.
[0383] The method may generate cells (i.e. , converted cells) exhibiting at least one characteristic of the desired cells. One or more characteristics may be used to select for the cells generated by the methods of the invention.
[0384] Characteristics include but are not limited to the detection or quantitation of expressed cell markers, enzymatic activity, and the characterization of morphological features and intercellular signaling. The biological function of the desired cell may also be evaluated, for example using functional assays, e.g. flow cytometry.
[0385] The cell markers indicative of the desired cells may be markers obtained by transcriptome analysis. For example, single cell RNA sequencing has been used to provide detailed transcriptional profiles of human cells obtained from primary human tissues. This information can be used to identify cells generated by the methods described herein. Additional resources, such as Human Cell Atlas and CellTypist may also be used to identify markers of specific cells.
[0386] The method may comprise assaying the converted cells obtained by the method described herein and determining a set of transcribed genes; comparing the set of transcribed genes of the converted cells to one or more reference sets of transcribed genes from one or more reference cells; and identifying a match between the converted cells and a reference cell.
[0387] In one embodiment, the method comprises the step of identifying converted cells as a desired cell by assaying morphological features of the converted cells and matching the morphological features to a reference tissue or cell's morphological features.BIT-C-P3829PCT
[0388] 56
[0389] In one embodiment, the method comprises the step of identifying converted cells as a desired cell by assaying protein marker expression of the converted cells and matching the protein marker expression to a reference cell protein marker expression.
[0390] In one embodiment, the method comprises the step of identifying converted cells as a desired cell by assaying a function and matching the function to a function of a reference cell.
[0391] In one embodiment, the cells obtained by the methods of the invention express a particular cell phenotype.
[0392] Alternatively, certain converted cells may be sorted from other converted cells and from cells on the basis of their expression of a lineage-specific cell surface antigen. Yet another means is by assessing expression at the RNA level, e.g., by RT-qPCR methods or by single cell RNA sequencing without any sorting or pre-selection step. Such techniques are known in the art.
[0393] Cell types
[0394] The method may be used on any cell type, including stem cells.
[0395] Sources of cells suitable for use with the invention may include any cells that are different from the desired (e.g. lineage-restricted or somatic) cells, for example, pluripotent stem cells. For example, the stem cells may be induced pluripotent stem cells (iPSCs), embryonic stem cells (ESCs) or pluripotent stem cells (PSCs) derived by nuclear transfer or cell fusion. It may be preferred that the embryonic stem cell is derived without destruction of the embryo, particularly where the cells are human. In some embodiments, the stem cells are not derived from human or animal embryos, i.e., the invention does not extend to any methods which involve the destruction of human or animal embryos. The stem cells may also include multipotent stem cells, oligopotent stem cells, or unipotent stem cells. The stem cells may also include fetal stem cells or adult stem cells, such as hematopoietic stem cells, mesenchymal stem cells, neural stem cells, epithelial stem cells or skin stem cells. In certain aspects, the stem cells may be isolated from umbilical, placenta, amniotic fluid, chorionic villi, blastocysts, bone marrow, adipose tissue, brain, peripheral blood, cord blood, menstrual blood, blood vessels, skeletal muscle, skin and liver.BIT-C-P3829PCT
[0396] 57
[0397] The cell used in the method of the invention may be any human or animal cell. It is preferably a mammalian cell. The cell is preferably a human cell. In certain aspects, the cell is preferably one from a livestock animal.
[0398] In one embodiment, the cell used in the method of the invention is of human origin. The source cell may be of human origin. As such, in a further embodiment, the PSC is a human PSC. In a further embodiment, the somatic cell is a human somatic cell.
[0399] In one embodiment, the cell used in the method of the invention is of animal origin. It is preferably a mammalian cell, such as a cell from a rodent, such as mice and rats; marsupial such as kangaroos and koalas; non-human primate such as a bonobo, chimpanzee, lemurs, gibbons and apes; camelids such as camels and llamas; livestock animals such as horses, pigs, cattle, buffalo, bison, goats, sheep, deer, reindeer, donkeys, bantengs, yaks, chickens, ducks and turkeys; domestic animals such as cats, dogs, rabbits and guinea pigs. The cell may also be an insect cell, a fish cell, a bird cell or an arthropod cell (such as an arachnid cell). With respect to stem cells, it is well known that, compared with non-human stem cells, genome engineering in human stem cells is challenging due to, for example, partially due to low transfection / transduction efficiency and high apoptosis under stresses such as low-density culture, drug-selection and sorting (Cerbini et al., PLOS ONE, 10(1), e0116032).
[0400] In preferred embodiments, the cell used in the method of the invention is an iPSC. Methods of preparing induced pluripotent stem cells are also known in the art. Induction of iPSCs typically require the expression of or exposure to at least one member from Sox family and at least one member from Oct family. Sox and Oct are thought to be central to the transcriptional regulatory hierarchy that specifies ES cell identity. For example, Sox may be Sox-1, Sox-2, Sox-3, Sox-15, or Sox-18; Oct may be Oct-4. Additional factors may increase the reprogramming efficiency, like Nanog, Lin28, Klf4, or c-Myc; specific sets of reprogramming factors may be a set comprising Sox-2, Oct-4, Nanog and, optionally, Lin-28; or comprising Sox-2, Oct4, Klf and, optionally, c-Myc. In one method, iPSC may be generated by transfecting cells with transcription factors Oct4, Sox2, c-Myc and Klf4 using viral transduction. In an alternative method, iPSC may be generated by transfecting cells with RNA encoding transcription factors inducing the development of stem cell characteristics, such as transcription factors selected from Oct4, Sox2, c-Myc and Klf4.
[0401] In one embodiment, the cell population comprises somatic cells, e.g., differentiated cells such as fibroblasts. The source cell may therefore be a somatic cell and the method may be aBIT-C-P3829PCT
[0402] 58
[0403] transdifferentiation method (i.e. the conversion of one cell type to another). Alternatively, the method may be used to convert a related somatic cell (e.g. a T lymphocyte other than a regulatory T cell to a regulatory T cell).
[0404] In one embodiment, the induced pluripotent stem cells are derived from somatic or germ cells of the patient. Such use of autologous cells would remove the need for matching cells to a recipient. Alternatively, commercially available PSC may be used, such as those available from WICELL (WiCell Research Institute, Inc, Wisconsin, US). Alternatively, the cells may be a tissue-specific stem cell which may also be autologous or donated.
[0405] In one embodiment, where the genetic sequence is suitable for forward programming or transdifferentiation, the methods are used to generate nerve cells (including neural stem cells, peripheral nervous system neurons, migratory enteric neural crest cell, neural crest cell, enteric neurons, glutamatergic neurons, GABAergic neurons, sensory neurons, motor neurons, dopaminergic neurons, medium spiny neurons, cortical neurons, interneurons, brainstem neurons, cholinergic neurons, hippocampal neurons, projection neurons, Lewy bodycontaining neurons, pyramidal neurons, cerebellar neurons, serotonergic neurons, thalamic neurons, neuromuscular junction cells, ganglion cells, oligodendrocyte precursors, oligodendrocytes, astrocytes, Purkinje cells, Muller cells, Granule cells, photoreceptors (rods and cones), retinal ganglion cells, retinal precursor cells, retinal pigment epithelium and Bruch's membrane cells), myocytes (including myoblasts, myogenic progenitors, skeletal myocytes, smooth muscle cells, satellite cells and cardiomyocytes), osteocytes (such as osteoclasts and osteoblasts), chondrocytes, adipocytes (including preadipocytes, brown adipocytes, beige adipocytes, white adipocytes and epicardial adipocytes), pericytes, hepatocytes (including hepatic stellate cells), kidney cells (such as papillary tips cells, podocytes and mesangial cells), respiratory cells (such as airway epithelial cells, club cells and ciliated cells), megakaryocytes, epithelial cells (such as melanocytes and placental villous trophoblasts), mesothelial cells (such as epicardium), endothelial cells (such as endocardial cells, keratinocytes and trabecular meshwork cells), secretary cells (such as pancreatic beta cells, pancreatic alpha cells, pancreatic acinar cells, pancreatic ductal cells and chromaffin cells), gastrointestinal cells (such as intestinal endocrine cell, intestinal epithelial cells, enteroendocrine cells, goblet cells, submucosal gland cells, Paneth cells and enterocytes), fibroblasts, myofibroblasts, mesodermal cells, mesenchymal cells (such as mesenchymal stem cells) and / or blood cells (including hematopoietic stem cells, erythrocytes, platelets, plasma cells and immune cells, such as neutrophils, glial cells, microglia, dendritic cells, T cells, such as CD4+ T helper cells, CD8+ cytotoxic T cells, regulatory T cells and gamma delta T cells, BBIT-C-P3829PCT
[0406] 59
[0407] cells, macrophages, Kupffer cells, innate lymphoid cells, eosinophils, mast cells, monocytes, Langerhans cells and natural killer cells). In one further embodiment the methods are used to generate mature cells selected from the list consisting of hepatocytes, microglia, glutamatergic neurons, GABAergic neurons, sensory neurons, pancreatic beta cells, astrocytes or regulatory T cells. In yet a further embodiment, the methods are used to generate hepatocytes, microglia or GABAergic neurons. Conversely, where the genetic sequence is suitable for transdifferentiation, the source cells may be any of the cell types listed above.
[0408] Cell culturing
[0409] In one embodiment, the method includes culturing the cell for a sufficient time and under conditions to allow the genetic sequence to be induced and expressed. Where the genetic sequence is suitable for forward programming or transdifferentiation, the method includes culturing the cell population for a sufficient time and under conditions to allow the forward programming or transdifferentiation of the cells to take place. Therefore, in one embodiment, the method comprises culturing the PSC to obtain the somatic cell. Generally, cells of the present invention are cultured in a culture medium, which is a nutrient-rich buffered solution capable of sustaining cell growth.
[0410] The cell culture medium may contain any of the following in an appropriate combination: salt(s), buffer(s), amino acids, glucose or other sugar(s), antibiotics, serum or serum replacement, and other components such as peptide growth factors, etc. Cell culture media ordinarily used for particular cell types are known to those skilled in the art. For example, the media may comprise Basal Medium (e.g. DMEM / F12) supplemented with GLUTAMAX, antibiotics (such as penicillin or streptomycin), B27 supplement and / or N2 supplement (all available from Thermo Fisher Scientific). The media may then be further supplemented at different time points during the culturing process. For example, one or more cytokines can be at 2, 4 and / or 10 days during the culturing process.
[0411] In one embodiment, the method comprises culturing in media comprising one or more cytokines. In one embodiment, the culture media comprises one or more components selected from the group consisting of: fibroblast growth factor (FGF2), bone morphogenetic protein (BMP4), Activin-A, CHIR99021, PI-103, A8301, C59, interleukin 2 (IL2), interleukin 7 (IL7), interleukin 15 (IL15), transforming growth factor-p (TGF-P) and retinoic acid.BIT-C-P3829PCT
[0412] 60
[0413] Cells may be obtained using methods of the invention at least about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 days after culturing. In one embodiment, the method comprises culturing under suitable conditions for at least 4 days, such as at least 10 days or at least 14 days. In further embodiments, method comprises culturing cells for a duration (e.g., at least 4 days, at least 5 days, at least 6 days, at least 7 days, at least 8 days, at least 9 days, at least 10 days, at least 11 days, at least 12 days, at least 13 days, at least 14 days, at least 21 days, at least 28 days, or longer, e.g., from 5 days to 40 days, from 7 days to 35 days, from 14 days to 28 days, or about 21 days) which is sufficient to generate a desired cell type. In some embodiments, the cells are cultured for a period of several hours (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 18, or 21 hours) to about 35 days (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 days). In one embodiment, the method comprises culturing the cells for at least about 5, 7, 10, 13, 15 or 20 days to produce a desired cell type. In one embodiment, the cells are cultured fora period of between 4 and 25 days, such as between 7 and 13 days, or 14 and 21 days. In one embodiment, the cells are cultured for at least about 7 days, for example at least about 13 days.
[0414] After culturing, the cell population may comprise more than one cell type. For example, such a cell population may have two cell types including the stem cells and a desired cell type. In one embodiment, the cell population comprises up to 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 99.5% (or any intermediate ranges) of the desired cell type in the resulting cell population.
[0415] The methods described herein may produce a population of cells that is substantially free from other cell types. In other words, the methods described herein may produce a homogenous (or substantially homogenous) population of cells. For example, a population produced by a method described herein may contain 80% or more, 85% or more, 90% or more, or 95% or more of a desired cell type following culture. Preferably the population of a desired cell type is sufficiently free of other cell types that no purification is required. If required, the population of a desired cell type may be purified by any convenient technique including FACS.
[0416] Culturing the cells may either help to induce cells to commit to a more mature phenotype, preferentially promote survival of the mature cells, or have a combination of both these effects.
[0417] According to a further aspect of the invention, there is provided a cell obtainable by any one of the methods defined herein.BIT-C-P3829PCT
[0418] 61
[0419] According to a further aspect of the invention, there is provided a cell with a modified genome that comprises:
[0420] (a) a first expression cassette comprising an inserted minigene operably linked to a first genetic sequence, wherein the minigene comprises an alternatively spliced exon and wherein expression of the first genetic sequence is controlled by alternative splicing of the minigene;
[0421] (b) a second expression cassette comprising a genetic sequence encoding a transcriptional regulator protein; and
[0422] (c) a third expression cassette comprising a second genetic sequence operably linked to an inducible promoter, wherein said inducible promoter is regulated by the transcriptional regulator protein encoded by the second expression cassette.
[0423] In one embodiment, the first genetic sequence and / or second genetic sequence is an exogenous genetic sequence. Other embodiments related to the minigene are as described hereinbefore. For example in one embodiment, the minigene is operably linked to an expression control element
[0424] In a further embodiment, the first, second and / or third expression cassette is inserted at a genomic safe harbour (GSH) site.
[0425] Therefore, according to a further aspect of the invention, there is provided a cell with a modified genome that comprises:
[0426] (a) a first expression cassette comprising a chimeric gene operably linked to an expression control element, wherein the chimeric gene comprises a first portion comprising a minigene having an alternatively spliced exon and a second portion that encodes a first genetic sequence, wherein expression of the first genetic sequence is controlled by alternative splicing of the first portion;
[0427] (b) a second expression cassette comprising a genetic sequence encoding a transcriptional regulator protein, wherein the second expression cassette is inserted at a first GSH site; and
[0428] (c) a third expression cassette comprising a second genetic sequence operably linked to an inducible promoter, wherein said inducible promoter is regulated by the transcriptional regulator protein encoded by the second expression cassette, wherein the third expression cassette is inserted at a second GSH site.BIT-C-P3829PCT
[0429] 62
[0430] In one embodiment, the transcriptional regulator protein in the second expression cassette is different to the expression control element in the first expression cassette.
[0431] This aspect of the invention provides a cell comprising two controllable transcription systems. The first expression cassette is the transcriptionally controlled cassette as described hereinbefore (and all relevant embodiments apply). The second and third expression cassettes are the dual cassette expression system described hereinbefore, and as described in more detail in WO2018 / 096343. The two expression systems can be used in parallel to provide enhanced transcriptional control of multiple genes / genetic sequences of interest.
[0432] The two systems may be used as complementary systems as part of the same process, for example sequential activation of transcription factors during forward programming. Alternatively, the two systems can be used to prepare a cell with enhanced safety, for example where one expression system is to induce forward programming or engineer the cell to express a target protein (e.g. a CAR), while the other expression system is used to encode a safety switch which may be activated if the cell needs to be terminated.
[0433] The two systems may also be used in tandem. For example, one expression system may be used to induce forward programming, while the other expression system is used to encode a gene such as a CRISPR enzyme or a recombinase which targets the first expression system to remove it from the cell once forward programming is complete. This may be particularly useful in the preparation of cells for therapy where any genetic modifications no longer needed as part of the therapy (e.g. the transcription factors and / or expression control elements used to generate the somatic cell) are preferably silenced or excised before administration to a patient.
[0434] In one embodiment, the first genetic sequence is a CRISPR enzyme. This CRISPR enzyme may target any endogenous gene specifically as is known in the art. In a further embodiment, the CRISPR enzyme is used to target the transcriptional regulator protein (such as rtTA) of the second expression cassette. This embodiment enables the silencing of the second expression system by the first expression system.
[0435] Cells of the invention may be modified for particular use in forward programming. Therefore, according to a further aspect of the invention, there is provided a cell modified to comprise a chimeric gene comprising a first portion comprising a minigene having an alternatively spliced exon operably linked to a second portion that encodes one or more transcription factors and / orBIT-C-P3829PCT
[0436] 63
[0437] one or more polypeptides having the activity of said one or more transcription factors. In one embodiment, the chimeric gene is operably linked to an expression control element.
[0438] Embodiments relevant to the chimeric gene are as described hereinbefore, such as for the minigene and other elements present in the transcriptionally controlled cassette. For example, in a preferred embodiment, the alternatively spliced exon comprises a translation initiation regulatory sequence and the second portion lacks a translation initiation regulatory sequence.
[0439] Cell compositions
[0440] According to a further aspect, there is provided a pharmaceutical composition comprising the cells produced by the method as described herein and a pharmaceutically acceptable carrier.
[0441] Pharmaceutical compositions may include cells as described herein in combination with one or more pharmaceutically or physiologically acceptable carrier, diluents, or excipients. Such compositions may include buffers such as neutral buffered saline, phosphate buffered saline and the like; carbohydrates such as glucose, mannose, sucrose or dextrans, mannitol; proteins; polypeptides or amino acids such as glycine; antioxidants; chelating agents such as EDTA or glutathione; adjuvants (e.g., aluminium hydroxide); and preservatives. Cryopreservation solutions which may be used in the pharmaceutical compositions of the invention include, for example, DMSO.
[0442] For purposes of manufacture, distribution, and use, the cells described herein may be supplied in the form of a cell culture or suspension in an isotonic excipient or culture medium, optionally frozen to facilitate transportation or storage.
[0443] Uses of the cells
[0444] The cells produced according to any of the methods of the invention have applications in basic and medical research, diagnostic and therapeutic methods. The cells may be used in vitro to study cellular development, provide test systems for new drugs, enable screening methods to be developed, scrutinise therapeutic regimens, provide diagnostic tests and the like. These uses form part of the present invention. Alternatively, the cells may be transplanted into a human or animal patient for diagnostic or therapeutic purposes. The use of the cells in therapy is also included in the present invention.BIT-C-P3829PCT
[0445] 64
[0446] According to one aspect of the invention, there is provided a cell as defined herein, for use in in vitro diagnostics or drug screening.
[0447] Cells generated by methods of the invention may find particular use in drug screening. Typically the screening would be carried out / n vitro. Therefore, in one embodiment, the method additionally comprises contacting the cells with a test substance and observing a change (e.g., an effect) in the cells induced by the test substance. The change or effect may be observed using methods known in the art, for example using pharmacological or toxicological assays. In one aspect, the cells may be used in a method of assessing a test substance (e.g., a drug, such as a compound), comprising assaying a pharmacological or toxicological property of the test substance on the cells provided by the methods described herein. The method may comprise: a) contacting the cells described herein with the test substance; and b) assaying an effect of the test substance on the cells.
[0448] Assessment of the activity of a candidate molecule may involve combining the cells described herein with the candidate molecule, determining any change in the morphology, phenotype, or metabolic activity of the cells that is attributable to the molecule (i.e. , compared with a control, such as untreated cells or cells treated with an inert compound), and then correlating the effect of the molecule with the observed change. The screening may be done either because the candidate molecule is designed to have a pharmacological effect on the cells or because the molecule is designed to have effects elsewhere but there is a need to determine if it has any unintended side effects.
[0449] Cytotoxicity can be determined in the first instance by the effect on cell viability, survival, morphology, and leakage of enzymes into the culture medium. More detailed analysis may be conducted to determine whether a test substance affects cell function without causing toxicity.
[0450] Alternatively, the cells can be used to assess changes in gene expression patterns caused by a potential drug candidate. In this embodiment, the changes in gene expression pattern from addition of the candidate drug can be compared with the gene expression pattern caused by a control drug with a known effect on the cells.
[0451] Therefore, according to a further aspect, there is provided a method for drug screening (e.g., evaluating drug reactivity), comprising a step of using the cells produced by the method as described herein. In one embodiment, this method is carried out / n vitro. According to a further aspect of the invention, there is provided a method of drug screening comprising contacting aBIT-C-P3829PCT
[0452] 65
[0453] cell generated using the method as defined herein, or a cell as defined herein, with the drug and observing a change in the cell induced by the drug.
[0454] According to a further aspect of the invention, there is provided the cell as defined herein for use in therapy.
[0455] In one embodiment, the method additionally comprises transplanting the cells into a patient. In this aspect of the invention, the cells used to generate the cells may be autologous (i.e., adult stem cells or mature cells removed, modified and returned to the same individual) or from a donor (i.e., allogeneic, including a stem cell line). Forward programming of cells, for example, into lineage-restricted cells is amenable to the production of autologous and allogeneic lineage-restricted cells. Preferably, where the cells are transplanted into a patient, the splicing modifications are carried out on the cell before transplantation, so that the splicing modifier does not need to be administered to the patient.
[0456] Therefore, according to a further aspect of the invention there is provided a method of treating a subject having or at risk of a disease or disorder comprising administering to the subject a therapeutically effective amount of cells generated using the method as defined herein, or cells as defined herein.
[0457] In a different aspect, the cells produced according to any of the methods of the invention may be used in tissue engineering. Tissue engineering requires the generation of tissue which could be used to replace tissues or even whole organs of a human or animal. Methods of tissue engineering are known to those skilled in the art, but include the use of a scaffold (an extracellular matrix) upon which the cells are applied in order to generate tissues / organs. These methods can be used to generate an “artificial” tissue or organ. Methods of generating tissues may include additive manufacturing, otherwise known as three-dimensional (3D) printing, which can involve directly printing cells to make tissues. The present invention thus provides a method for generating tissues using the cells produced as described in any aspect of the invention.
[0458] For the drug screening, diagnostic and therapeutic uses set out above, it is clear that the cells need to meet stringent requirements in terms of batch-to-batch reproducibility and cell viability. Cell viability typically needs to be maintained throughout the drug screening, diagnostic or therapeutic process. When the uses are applied to humans specifically it is also important that the cells are cultured in defined media in the absence of animal-derived constituents, such asBIT-C-P3829PCT
[0459] 66
[0460] fetal bovine serum. Excipients need to not only be compatible with the cells but they also need to not cause any adverse reactions in the patient if the cells are used therapeutically. The donor information of the cells also needs to be taken into account, as that in itself may affect how the cells behave in a drug screening, diagnostic and therapeutic context. All of these requirements add complexity to the culturing process compared with, for example, a culturing process for meat production.
[0461] It will be understood that all embodiments described herein may be applied to all aspects of the invention.
[0462] Other features and advantages of the present invention will be apparent from the description provided herein. It should be understood, however, that the description and the specific examples while indicating preferred embodiments of the invention are given by way of illustration only, since various changes and modifications will become apparent to those skilled in the art. The invention will now be described using the following, non-limiting examples:
[0463] EXAMPLES EXAMPLE 1 - SplicelN system expression
[0464] Design & Constructs
[0465] As set out in Figure 1A, a series of DNA constructs were prepared incorporating the SplicelN sequence (SEQ ID NO: 1) operably linked to a control sequence encoding a green fluorescent protein (unstable EGFP, ‘EGFPd2’). Constructs contained either a TRE (very strong but inducible) or CAG (less strong but constitutive) promoter to drive expression. TRE-driven expression was chosen to help inform on the leakiness of the SplicelN system, i.e. expression in absence of inducing drug. Constructs also optionally contained a 2A peptide coding sequence to test whether the addition of a self-cleaving peptide at the N-terminus of the genetic sequence (EGFPd2) will affect SplicelN induction. Constructs without a 2A peptide are named “s1” and constructs with a 2A peptide are named “s2”.
[0466] Constructs were cloned into plasmids containing AAVS1 homology arms and puromycin resistance genes for targeting into the AAVS1 genomic safe harbour locus of the host cell genome.BIT-C-P3829PCT
[0467] 67
[0468] Methods
[0469] All constructs were integrated at AAVS1 in host cells. The host cell already expressed a transcriptional activator protein (reverse tetracycline trans-activator (rtTA)). Cells with on-target integration were selected with puromycin.
[0470] Selected pools were induced for 3 days with 100 nM LMI070. Result readout was obtained by flow cytometry after a live-dead stain.
[0471] Results
[0472] The results show the SplicelN system successfully expressed EGFP following LMI070 induction. Expression was achieved with both the CAG (Figure 1B) and TRE (Figure 1C) promoters. The expression is comparable between cells showing that induced-expression is robust. Expression of the system was also not indicated to be leaky when read-out by flow cytometry.
[0473] The addition of a T2A peptide did not affect the expression of constructs with a CAG promoter, but was indicated to disrupt the system when transcription of constructs with a TRE promoter (Figure 1C).
[0474] EXAMPLE 2 - System in cellular forward programming
[0475] Design & Constructs
[0476] As shown in Figure 2A, DNA constructs were prepared incorporating either the sequence encoding the NGN2 transcription factor (top, TRE::NGN2) or the SplicelN sequence (SEQ ID NO: 1) operably linked to control the sequence encoding the NGN2 transcription factor (bottom, TRE::s1NGN2). These constructs contained a TRE promoter to drive expression. Constructs were cloned into plasmids containing AAVS1 homology arms and puromycin resistance genes for targeting into the AAVS1 genomic safe harbour locus of the host cell genome.
[0477] Methods & Results
[0478] As shown in Figure 2B, clonal induced Pluripotent Stem Cell (iPSC) lines carrying either construct presented in Figure 2A were generated and tested for forward programming into glutamatergic neurons, following an 11-day standard protocol. All experiments were done in biological duplicates. TRE::NGN2 (control) and TRE::s1NGN2 (inducible expression, with no 2A peptide included) generated glutamatergic neurons with identical morphologies. In contrast,BIT-C-P3829PCT
[0479] 68
[0480] non-induced TRE::s1NGN2 (treated with doxycycline only) generated cells with undesired morphology and confluence.
[0481] mRNA from the forward programmed cells presented in Figure 2A was sequenced (Bulk RNA sequencing) and analysed. The analysis included the quantification of gene expression for genes involved in neuronal or iPSC function (Figure 20) and a Principal Component Analysis (PCA) on the whole transcriptome of the cells (Figure 2D). This analysis shows that TRE::NGN2 glutamatergic neurons express near-identical transcriptomes to induced TRE::s1NGN2 glutamatergic neurons and broadly different transcriptomes to non-induced (doxycycline only) TRE::s1NGN2 cells.
[0482] EXAMPLE 3 - System in safety switch expression
[0483] Design & Constructs
[0484] DNA constructs were prepared incorporating the SplicelN sequence (SEQ ID NO: 1) operably linked to control sequences encoding functional fragments of Caspase enzymes, which can dimerise via their small subunits and induce cell death by apoptosis. Figure 3A depicts the proteins tested and their mechanism of action. The Caspase (Casp) protein tested contain i) both large and small subunits of truncated Caspase 3 (Casp3), ii) both large and small subunits of truncated Caspase 9 (Casp9) or iii) the large subunit of Caspase 9 and the small subunit containing the dimerisation domain of Caspase 3 (Casp9-3DD). These constructs contained a CAG promoter to drive expression and a T2A peptide at the N-terminus of the Caspase proteins. Constructs were cloned into plasmids containing AAVS1 homology arms and puromycin resistance genes for targeting into the AAVS1 genomic safe harbour locus of the host cell genome.
[0485] Methods & Results
[0486] Caspase protein expression and dimerization were expected to lead to cell death following LMI070 induction. This was quantified by live imaging of wild-type cells and cells carrying the constructs detailed above, either non-induced or induced with LMI070. The results presented in Figure 3B show the SplicelN system successfully controlled expression ofthe three different caspase safety switches resulting in cell death.BIT-C-P3829PCT
[0487] 69
[0488] EXAMPLE 4 - System to express Cas9
[0489] Design & Constructs
[0490] As shown in Figure 4A, DNA constructs were prepared incorporating the SplicelN sequence (SEQ ID NO: 1) operably linked to control the sequence encoding the Cas9 nuclease. This construct contained a CAG promoter to drive expression and a T2A peptide at the N-terminus of the Cas9 protein (s2Cas9).
[0491] Methods & Results
[0492] Cells expressing Cas9 sequence were transfected with either i) a ribonucleoprotein (RNP) complex composed of the Cas9 nuclease and a guide RNA targeting the B2M endogenous genomic locus (B2M gRNA) or ii) with the B2M gRNA only. Cells expressing the s2Cas9 sequence were either not induced with LMI070 or induced with LMI070 before and / or after transfection. The results presented in Figure 4C show that, as predicted, the B2M locus was edited in all conditions where the RNP was transfected. Upon LMI070 induction, the SplicelN system successfully controlled s2Cas9 expression, which resulted in B2M locus editing after B2M gRNA-only transfection. Efficient s2Cas9 expression was achieved with both pre- and post-transfection LMI070 induction.
Claims
BIT-C-P3829PCT70CLAIMS1. A method of controlling transcription of a genetic sequence in a somatic cell derived from a pluripotent stem cell (PSC) comprising:(i) providing a PSC comprising an inserted minigene operably linked to the genetic sequence, wherein the minigene comprises a sequence with an alternatively spliced exon and wherein expression of the genetic sequence is controlled by alternative splicing of the minigene; and(ii) applying a splicing modifier that regulates the splicing of the alternatively spliced exon thereby inducing expression of the minigene to include the alternatively spliced exon, wherein the presence of the alternatively spliced exon initiates translation of the genetic sequence.
2. The method of claim 1 , wherein the alternatively spliced exon comprises a translation initiation regulatory sequence and the genetic sequence lacks a translation initiation regulatory sequence, thereby initiation of translation of the genetic sequence only occurs when the alternatively spliced exon is present.
3. The method of claim 1 or claim 2, wherein the splicing modifier is applied to the somatic cell following forward programming of the PSC.
4. The method of claim 1 or claim 2, wherein the splicing modifier is applied during the process of forward programming the PSC into the somatic cell.
5. The method of any preceding claim, wherein the genetic sequence is an exogenous genetic sequence and wherein the inserted minigene is present in a chimeric gene comprising a first portion encoding the minigene and a second portion encoding the exogenous genetic sequence.
6. The method of any preceding claim, wherein the genetic sequence encodes a CRISPR enzyme.
7. The method of claim 6, wherein the CRISPR enzyme is Cas9, in particular codon optimised Cas9.BIT-C-P3829PCT718. The method of any one of claims 1 to 5, wherein the genetic sequence encodes one or more transcription factors.
9. The method of any one of claims 1 to 5, wherein the genetic sequence reduces or inhibits the expression of an endogenous gene.
10. The method of any one of claims 1 to 5 and 9, wherein the genetic sequence encodes a non-coding RNA, such as an inhibitory RNA.
11. The method of any one of claims 1 to 5, wherein the genetic sequence activates the expression of an endogenous gene.
12. The method of any one of claims 1 to 5, wherein the method is used to prepare a somatic cell for use in cell therapy.
13. The method of any one of claims 1 to 5 and 12, wherein the genetic sequence encodes a safety switch or a chimeric antigen receptor.
14. The method of any one of claims 1 to 5, wherein the genetic sequence encodes a DNA-binding enzyme, such as a recombinase, integrase or transposase.
15. The method of any one of claims 1 to 4, wherein the genetic sequence is an endogenous genetic sequence and the inserted minigene is operably linked to the endogenous genetic sequence present in the PSC genome.
16. The method of claim 15, wherein the endogenous gene is the p2-Microglobulin (B2M) gene or a Human Leukocyte Antigen (HLA) gene selected from HLA-A, HLA-B and HLA-C.
17. The method of any preceding claim, wherein the PSC additionally comprises an exogenous transcriptional regulator protein.
18. The method of any preceding claim, wherein the PSC additionally comprises:(a) an inserted genetic sequence encoding a transcriptional regulator protein at a first genomic safe harbour (GSH) site; andBIT-C-P3829PCT72(b) an inserted inducible cassette comprising a second exogenous genetic sequence operably linked to an inducible promoter at a second GSH site, wherein said inducible promoter is regulated by the transcriptional regulator protein.
19. The method of claim 18, wherein the first and second GSH sites are different.
20. The method of any preceding claim, wherein the minigene additionally encodes a selfcleaving peptide.
21. The method of claim 20, wherein the self-cleaving peptide is present between the minigene and the genetic sequence.
22. The method of claim 20 or claim 21 , wherein the self-cleaving peptide is T2A.
23. The method of any preceding claim, wherein the minigene is operably linked to an expression control element.
24. The method of claim 23, wherein the expression control element is a constitutive promoter, such as a CAG promoter.
25. The method of claim 23, wherein the expression control element is an inducible promoter, such as a Tet Responsive Element (TRE).
26. The method of any preceding claim, wherein the splicing modifier is a small molecule.
27. The method of claim 26, wherein the small molecule splicing modifier is LMI070 or RG7800.
28. The method of any preceding claim, wherein the minigene is derived from the SF3B3 gene, in particular comprising an upstream exon, pseudoexon and a downstream exon derived from the SF3B3 gene.
29. The method of any preceding claim, wherein the minigene is derived from the SMN2 gene, in particular comprising exons 6-8 of the SMN2 gene.BIT-C-P3829PCT7330. The method of any preceding claim, wherein the minigene comprises a sequence having at least 90% sequence identity with SEQ ID NO: 1.
31. The method of any preceding claim, wherein the minigene is inserted into a genomic safe harbour in the PSC genome.
32. The method of claim 31, wherein the genomic safe harbour site is selected from the ROSA26 locus, the AAVS1 locus, the CLYBL gene or the CCR5 gene.
33. The method of any preceding claim, wherein the method is performed ex vivo.
34. The method of any preceding claim, wherein the PSC and somatic cell is a human, marsupial, non-human primate, camelid, livestock animal, rodent, domestic animal, bird, fish or insect cell.
35. The method of any preceding claim, wherein the PSC is an induced pluripotent stem cell.
36. The method of any preceding claim, which comprises culturing the PSC to obtain the somatic cell.
37. The method of any preceding claim, which comprises culturing under suitable conditions for at least 4 days, such as at least 10 days, in particular at least 14 days.
38. The method of any preceding claim, which comprises culturing in media comprising one or more cytokines.
39. A method of controlling transcription of an exogenous genetic sequence in an PSC, comprising:(i) administering to the PSC a chimeric gene comprising a first portion comprising a minigene having an alternatively spliced exon and a second portion that encodes the exogenous genetic sequence, wherein expression of the exogenous genetic sequence is controlled by alternative splicing of the first portion; and(ii) applying a splicing modifier that regulates the splicing of the alternatively spliced exon thereby inducing expression of the minigene to include the alternatively spliced exon and initiating translation of the exogenous genetic sequence.BIT-C-P3829PCT7440. A method of forward programming a PSC into a somatic cell, comprising:(i) administering to the PSC a chimeric gene comprising a first portion comprising a minigene having an alternatively spliced exon and a second portion that encodes one or more transcription factors and / or one or more polypeptides having the activity of said one or more transcription factors, wherein expression of the exogenous genetic sequence is controlled by alternative splicing of the first portion;(ii) applying a splicing modifier that regulates the splicing of the alternatively spliced exon thereby inducing expression of the minigene to include the alternatively spliced exon and initiating translation of the one or more transcription factors and / or polypeptides encoded by the second portion,wherein translation of the one or more transcription factors and / or polypeptides encoded by the second portion induces forward programming of the PSC into the somatic cell.
41. The method of claim 40, wherein the second portion encodes between one and seven transcription factors.
42. The method of any one of claims 39 to 41, wherein the alternatively spliced exon comprises a translation initiation regulatory sequence and the second portion lacks a translation initiation regulatory sequence.
43. The method of any one of claims 39 to 42, wherein the PSC is an induced pluripotent stem cell.
44. A cell obtainable by any one of the methods defined in claims 1 to 43.
45. A cell with a modified genome that comprises:(a) a first expression cassette comprising an inserted minigene operably linked to a first genetic sequence, wherein the minigene comprises a sequence with an alternatively spliced exon and wherein expression of the first genetic sequence is controlled by alternative splicing of the minigene;(b) a second expression cassette comprising a genetic sequence encoding a transcriptional regulator protein; and(c) a third expression cassette comprising a second genetic sequence operably linked to an inducible promoter, wherein said inducible promoter is regulated by the transcriptional regulator protein encoded by the second expression cassette.BIT-C-P3829PCT7546. The cell of claim 45, wherein the first genetic sequence and / or second genetic sequence is an exogenous genetic sequence.
47. The cell of claim 45 or claim 46, wherein the first, second and / or third expression cassette is inserted at a GSH site.
48. The cell of any one of claims 45 to 47, wherein the minigene is operably linked to an expression control element.
49. The cell of claim 48, wherein the transcriptional regulator protein in the second expression cassette is different to the expression control element in the first expression cassette.
50. A cell modified to comprise an expression cassette comprising a chimeric gene comprising a first portion comprising a minigene having an alternatively spliced exon operably linked to a second portion that encodes one or more transcription factors and / or one or more polypeptides having the activity of said one or more transcription factors.
51. The cell of claim 50, wherein the alternatively spliced exon comprises a translation initiation regulatory sequence and the second portion lacks a translation initiation regulatory sequence.
52. A cell as defined in any one of claims 44 to 51 , for use in therapy.
53. A cell as defined in any one of claims 44 to 51 , for use in in vitro diagnostics or drug screening.
54. A nucleic acid molecule comprising a sequence having at least 90% sequence identity with SEQ ID NO: 1.