Genomic Safe Harbor
Patent Information
- Application Number
- JP2023571569
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-05-20
- Filing Date
- 2022-05-19
- Publication Date
- 2025-05-26
AI Technical Summary
Current gene therapy methods face challenges due to unpredictable expression and potential harmful effects of randomly inserted genes, particularly in centromeric and subtelomeric regions, which can lead to position effects, silencing, and influence surrounding endogenous genes, potentially causing malignant transformation and disrupting cell differentiation.
Identification and validation of novel genomic safe harbors (GSH) for stable and predictable transgene insertion, using methods such as in silico approaches and functional assays, to ensure safe and efficient integration of marker genes in human cells, with validation through in vitro and in vivo assays.
The novel GSH loci enable stable and predictable expression of transgenes, reducing the risk of insertional mutagenesis and maintaining cell viability and differentiation potential, paving the way for safer and more effective gene therapy applications.
Smart Images

Figure 00000173_0000 
Figure 00000173_0001 
Figure 00000173_0002
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 190,996, filed May 20, 2021, the entire contents of which are incorporated herein by reference in their entirety. [Background technology]
[0002] background The modification of the human genome by stable insertion of functional transgenes and other genetic elements is highly beneficial in biomedical research and medicine (e.g., for gene therapy). Genetically modified human cells are also useful for studying gene function and tracing and lineage analysis using reporter systems. All these applications depend on the reliable function of the introduced genes in their new environment. However, randomly inserted genes are subject to position effects and silencing, which makes their expression uncertain and unpredictable. Centromeric and subtelomeric regions are particularly prone to transgene silencing.
[0003] Reciprocally, newly integrated genes can affect the surrounding endogenous genes and chromatin, potentially altering cell behavior or favoring cell transformation. As a result, despite the success of therapeutic gene transfer, there have been cases of malignant transformation associated with the insertional activation of oncogenes after stem cell gene therapy, highlighting the importance of where the newly integrated DNA is located. In addition, the insertion of foreign DNA into the genome of progenitor cells can have deleterious effects on terminal differentiation into certain cell types.
[0004] Genomic Safe Harbor (GSH) refers to a genetic locus that accepts the insertion of foreign DNA, either constitutively or conditionally / inducibly, without significantly affecting the viability and ontogenetic processes of somatic, progenitor or germline cells. The availability of GSH loci is extremely useful for expressing reporter, suicide, selectable or therapeutic genes. Three intragenic sites have been proposed for GSH (AAVS1, CCR5 and ROSA26, and albumin in mouse cells) (see, e.g., U.S. Patent Nos. 7,951,925, 8,771,985, 8,110,379, 7,951,925, U.S. Patent Application Publication Nos. 20100218264, 20110265198, 20130137104, 20130122591, 20130177983, 20130177960, 20150056705 and 20150159172, all of which are incorporated by reference). However, these proposed GSHs are in relatively gene-rich regions and near genes associated with cancer. Genes flanking AAVS1 can be spared by some promoters, but safety validation in multiple tissues remains to be performed, and the dispensability of the disrupted genes, especially after biallelic disruption, as is often the case with endonuclease-mediated targeting, remains to be further investigated. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] U.S. Patent No. 7,951,925 [Patent Document 2] U.S. Patent No. 8,771,985 [Patent Document 3] U.S. Pat. No. 8,110,379 [Patent Document 4] U.S. Patent No. 7,951,925 [Patent Document 5] US Patent Publication No. 20100218264 [Patent Document 6] US Patent Publication No. 20110265198 [Patent Document 7] US Patent Publication No. 20130137104 [Patent Document 8] US Patent Publication No. 20130122591 [Patent Document 9] US Patent Publication No. 20130177983 [Patent Document 10] US Patent Publication No. 20130177960 [Patent Document 11] US Patent Publication No. 20150056705 [Patent Document 12] US Patent Publication No. 20150159172 Summary of the Invention [Means for solving the problem]
[0006] Thus, there is a great need for the identification and validation of additional GSH loci, as well as compositions and methods for the identified GSH loci.
[0007] Summary of the Invention The present invention is based, at least in part, on the discovery that the novel GSH locus identified herein is particularly useful in the stable insertion and predictable expression of various transgenes, e.g., required to treat patients (e.g., by gene therapy) or prepare pharmaceuticals (e.g., biologics or vaccines).
[0008] In certain embodiments, various methods for identifying novel GSH loci are provided herein, including in silico approaches as well as functional assays. Further provided herein are various in vitro, ex vivo, and in vivo methods for validating the identified GSH, including de novo targeted insertion of a marker gene into the GSH locus in a cell (e.g., a human cell) to assess the insertion efficiency and expression level of the marker gene; targeted insertion of a marker gene into the GSH locus in a progenitor or stem cell to determine its effect on progenitor or stem cell differentiation in vitro; targeted insertion of a marker gene into a locus in a progenitor or stem cell and engrafting the cells into immune-depleted mice to determine marker gene expression in all developmental lineages in vivo; targeted insertion of a marker gene into the GSH locus in a cell and determining a global cellular transcriptional profile (e.g., using RNAseq or microarrays) to determine the effect of the insertion at the GSH locus on the overall transcriptional profile of the cell; and / or generating transgenic knock-in mice in which the genomic DNA of the mouse has the marker gene inserted into the locus.
[0009] In certain aspects, various compositions comprising the GSH locus described herein are provided herein. For example, a nucleic acid vector comprising at least a portion of the GSH nucleic acid described herein is provided herein. In a preferred embodiment, a sequence (5' and 3' homology arms) having homology to the GSH locus is adjacent to at least one non-GSH nucleic acid, such that the homology arms facilitate the integration of at least one non-GSH nucleic acid into the GSH locus. Such non-GSH nucleic acid may comprise a nucleic acid encoding a protein or a fragment thereof, such as a human protein or a fragment thereof; a therapeutic protein or a fragment thereof, an antigen-binding protein, or a peptide; a suicide gene, such as Herpes Simplex Virus-1 Thymidine Kinase (HSV-TK); a viral protein or a fragment thereof; a nuclease; a marker; and / or a drug resistance protein. Also provided herein are viral vectors comprising various nucleic acid vectors of the present disclosure. Further provided herein are cells comprising the nucleic acid vectors of the present disclosure, as well as cells comprising at least one non-GSH nucleic acid integrated into GSH in the genome. Additionally, pharmaceutical compositions comprising nucleic acid vectors, viral vectors, and / or cells are provided, as well as transgenic organisms comprising at least one non-GSH nucleic acid integrated into the GSH in the genome of the cell.
[0010] In certain embodiments, methods of using and producing the compositions described herein are provided herein. Such methods include methods of preventing or treating various diseases; modulating the level and / or activity of proteins in cells or in a subject (e.g., increasing protein levels by introducing additional copies of the gene encoding said protein, or decreasing protein levels by introducing non-coding RNA and / or CRISPR gene editing that downregulates or eliminates the gene encoding said protein); producing biologics, such as antigen-binding proteins and / or therapeutic proteins (e.g., insulin); producing viral vectors, including those for gene therapy. Further provided herein are compositions and methods for incorporating viral surface proteins into the GSH locus of the present disclosure, which allow in vivo immunization by exposing a subject to viral antigens to induce an immune response. Importantly, such viral antigens can be turned on and off intermittently by using the inducible promoter of the present disclosure, which allows for pulsatile expression of viral antigens. [Brief description of the drawings]
[0011] [Figure 1] Figure 1 illustrates the current challenges for safe gene therapy and the potential consequences of promiscuous (random) DNA integration. There is mounting evidence that promiscuous therapeutic integration of genes can lead to insertional mutagenesis, genotoxicity, or affect expression of the gene of interest (e.g., encompassed herein by a non-GSH nucleic acid), presenting a major barrier to realizing the promise of gene therapy.
[0012] [Diagram 2]Figure 2A and Figure 2B show that targeted integration into GSH allows predictable transgene expression and reduces the risk of insertional mutagenesis in the host genome. Figure 2B shows that syntenic GSH provides predictability across relatedness search models, facilitating preclinical and clinical development. The use of safe, well-characterized loci for persistent gene transfer may well become a prerequisite for safe and successful ex vivo and in vivo gene therapy treatments.
[0013] [Diagram 3] FIG. 3 shows a diagram of a representative method for identifying the GSH locus.
[0014] [Figure 4-1] Figure 4A-C show characterization of the novel GSH locus. CFU (colony forming unit) assay to test the differentiation potential of human CD34+ hematopoietic stem cells (HSC). Figure 4A is a schematic diagram showing the assay performed herein. Gene-directed integration into SYNTX-GSH1, a novel GSH locus identified herein, enabled successful HSC differentiation into committed erythroid precursors. Figure 4B shows high transgene expression (GFP) in committed erythroid precursors. Figure 4C shows a diagram illustrating HSC differentiation (erythropoiesis). [Figure 4-2] Same as above.
[0015] [Diagram 5] Figure 5A-5B show gene editing of marker genes into the GSH locus identified herein. Figure 5A shows gene editing efficiency into GSH in CD34+ HSCs identified herein. AAVS1, a previously known GSH locus, was used as a positive control. Figure 5B shows that differentiation of primary CD34+ HSCs into committed CD71+ / CD235a+ erythroblasts was not affected after gene insertion into SYNTX-GSH (SYNTX-GSH1 and SYNTX-GSH2).
[0016] [Figure 6] Figure 6A-6B show the expression of a marker gene (GFP) integrated into the different GSH loci. GFP expression was determined 14 days after gene editing into SYNTX-GSH and AAVS1 (positive control) in CD34+ HSCs. (SYNTX-GSH1 and SYNTX-GSH2). Gene editing into SYNTX-GSH was more efficient than editing into AAVS1. Edited cells stably expressed GFP 2 weeks after gene editing and proceeded to differentiate from CD34+ HSCs into erythroid precursors. SYNTX-GSH1 and 2 edited cells expressed higher levels of the transgene (GFP) than AAVS1 edited cells. (SYNTX-GSH1 and SYNTX-GSH2).
[0017] [Figure 7-1] Figures 7A-7D show the effect of transgene knock-in into SYNTX-GSH on the global transcriptional profile of cells. Figure 7A shows the experimental design of cellular perturbation analysis by RNAseq. Figure 7B shows the RNAseq analysis performed on SYNTX-GSH1 and SYNTX-GSH2 compared to wild-type cells and AAVS1. Figure 7C shows the principal component analysis. Figure 7D shows the expression of the integrated marker gene GFP in the knock-in cell lines. Transgene integration into SYNTX-GSH had less impact on the transcriptional profile of cells than integration into the AAVS1 site. SYNTX-GSH1 and SYNTX-GSH2 showed higher and more stable transgene expression than AAVS1 in human cells. [Figure 7-2] Same as above.
[0018] [Figure 8-1]Figures 8A-8C assess GSH performance by determining the stability of GFP expression over cell passages. Figure 8A shows a schematic of the experiment. Figures 8B and 8C show the expression of a marker gene (GFP) inserted into the SYNTX-GSH locus. Transgene integration into the four different SYNTX-GSH loci resulted in different editing efficiencies and transgene expression. SYNTX-GSH1 and SYNTX-GSH2 showed higher and more stable transgene expression than AAVS1. SYNTX-GSH3 and SYNTX-GSH4 showed lower expression levels and may be useful for the insertion of genes (e.g., lethal genes) that require lower expression levels. The GSH loci identified herein provide a palette of individual GSHs with different characteristics to accommodate specific gene therapy programs. [Figure 8-2] Same as above.
[0019] [Figure 9-1] Figure 9A and Figure 9B show the secondary structure of AAV ITR and a schematic diagram of rolling hairpin replication model. Figure 9A shows the structure of AAV ITR that forms extensive secondary structure. ITR can acquire two conformations (flip and flop). Figure 9B shows a schematic diagram of the rolling hairpin replication model that viral nucleic acid replicates. [Figure 9-2] Same as above.
[0020] [Figure 10] Figure 10 shows a schematic diagram depicting a heterologous nucleic acid / transgene construct containing a β-globin gene operably linked to a β-globin promoter flanked at its 5' end by one or more HS sequences. Mammalian β-globin genes are controlled by a regulatory region called the locus control region (LCR), which contains a series of five DNase I hypersensitive sites (HS1-HS5). HS is required for efficient expression of the β-globin gene. Each transgene construct is positioned between two homology arms (5' and 3' homology arms), which facilitate site-specific integration in the target cell genome by homologous recombination.
[0021] [Figure 11] Figure 11 shows a schematic diagram depicting heterologous nucleic acid / transgene constructs containing various promoters. Each promoter (e.g., CAG promoter, AHSP promoter, MND promoter, WA promoter, PKLR promoter) is operably linked to a transgene of interest, and the entire construct is placed between two homology arms (5' and 3' homology arms), which facilitates site-specific integration at the GSH locus of the target cell genome by homologous recombination.
[0022] [Figure 12] Figure 12 shows a partial DNA sequence of the erythroid-specific promoter of PKLR. A 469 bp region containing the upstream regulatory domain. Elements conserved between the human and rat PK-R promoters are depicted with dotted lines. The cytosines at the PK-R transcription start site are underlined. The GATA-1, CAC / Sp1 motifs, and the regulatory element PKR-RE1 in the upstream 270 bp region are shown in boxes (orientation indicated by arrows).
[0023] [Figure 13-1] Figures 13A and 13B show exemplary miRNAs that can be targeted by recombinant virions described herein. Recombinant virions of erythroparvoviruses can contain miRNA sequences. Alternatively, recombinant virions can contain nucleic acid sequences that inactivate miRNAs. [Figure 13-2] Same as above.
[0024] [Figure 14]Figure 14 shows a pulsed gene expression system. The schematic shows both negative and positive control of expression. Example I (top panel) shows that ASOs (antisense oligonucleotide ASOs or AONs) can negatively regulate gene expression post-transcriptionally. Without the ASO, the primary transcript (left) is spliced (top line) into a translatable mRNA. Addition of an ASO (red line) complementary to the splice acceptor at the 3' end of the intron / 5' end of the second exon interferes with splicing. Thus, in the presence of the ASO, the intron remains in the transcript. Unprocessed RNA is either untranslatable or produces a non-functional protein upon translation. Example II (bottom panel) illustrates that ASOs can positively affect gene expression post-transcriptionally. The primary transcript (left) contains four exons: the first, third and fourth exons code for the therapeutic protein, and the second exon contains either a nonsense mutation or an out-of-frame mutation (OOF). Such a second exon can be engineered into any transgene. Without ASO, the transcript is processed into a mature mRNA containing four exons (bottom line), i.e., the second exon with a nonsense or OOF mutation remains. As a result, the resulting mRNA is translated into a truncated or non-functional protein. In contrast, the addition of ASO interferes with splicing and the mature mRNA consists of the first, third and fourth exons, i.e., the second exon with a nonsense or OOF mutation is excised. As a result, in the default state (without ASO), no therapeutic protein is produced. Only upon addition of ASO, a therapeutic protein is produced, thereby resulting in a positive control.
[0025] [Figure 15]Figure 15 shows ATACseq coverage and peaks. The EVE insertion site is indicated by a black vertical line in the center of the plot. For each donor, ATACseq coverage is shown as a smooth grey line, with called peaks shown as vertical bars color-coded for each donor. The distance from the EVE insertion to the nearest peak across donors is 1,144 base pairs, indicating available chromatin. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0026] Detailed Description of the Invention In certain aspects, provided herein are novel methods for identifying and validating the GSH locus; newly identified GSH loci; compositions comprising the sequence of the GSH locus; and methods of using the GSH locus and compositions comprising same for treating patients (e.g., by gene therapy or cell therapy), preparing medicaments (e.g., biologics or vaccines), and other applications described herein. definition
[0027] The articles "a" and "an" are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, "an element" means one element or more than one element.
[0028] The term "administering" is intended to include routes of administration that allow the therapy to perform its intended function. Examples of routes of administration include injection (intramuscular, subcutaneous, intravenous, parenteral, intraperitoneal, intrathecal, intratumoral, intranasal, intracranial, intravitreal, subretinal, etc.). Routes of administration also include inhalation, as well as direct injection into bone marrow. Injection can be a bolus injection or a continuous infusion. Depending on the route of administration, the agent can be coated with or placed in a selected material to improve absorption or to protect it from natural conditions that may adversely affect its ability to perform its intended function.
[0029] The term "cetacea" refers to a taxonomic (infra) order of aquatic marine mammals that includes, among others, baleen whales, toothed whales, dolphins and porpoises, and related species, having torpedo-shaped, largely hairless bodies, paddle-shaped forelimbs but no hind limbs, one or two externally opening nostrils on the top of the head, and a horizontally flattened tail used for locomotion.
[0030] The term "chiroptera" refers to the taxonomic order of mammals that can truly fly, and includes bats.
[0031] As used herein, "donor sequence" refers to a polynucleotide that will be inserted into a host cell genome or will be used as a repair template for the host cell genome. The donor sequence can include modifications, which are desirable to be made during gene editing. The sequence to be incorporated can be introduced into the target nucleic acid molecule by homology-directed repair in the target sequence, which results in the target sequence being changed from the original target sequence to the sequence contained in the donor sequence. Thus, the sequence contained in the donor sequence can be an insertion, deletion, indel, point mutation, repair of mutation, etc., for the target sequence. The donor sequence can be, for example, a single-stranded DNA molecule; a double-stranded DNA molecule; a DNA / RNA hybrid molecule; and a DNA / modRNA (modified RNA) hybrid molecule. In an embodiment, the donor sequence is independent of the homology arm. Editing can be RNA and DNA editing. The donor sequence can be endogenous or exogenous to the host cell genome, depending on the nature of the desired gene editing.
[0032] The term "endogenous viral elements" or "EVEs" are DNA sequences derived from viruses and present in the germline of non-viral organisms. EVEs can be entire viral genomes (proviruses) or fragments of viral genomes. They arise when viral DNA sequences are integrated into the genome of embryonic cells that go on to produce viable organisms. The newly established EVEs can be inherited as alleles from one generation to the next in the host species and can even lead to fixation.
[0033] The term "homologous recombination" is art-recognized and, when used in reference to nucleic acid insertion in a target genome, is intended to include homology-dependent repair.
[0034] The term "homology" or "homologous," as used herein, is defined as the percentage of nucleotide residues in the homology arms that are identical to nucleotide residues in the corresponding sequence on the target chromosome after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent sequence identity. Identity between regions of nucleic acid sequences can be determined as a percentage of identity using known computer algorithms, such as the "FASTA" program, using default parameters, for example, as in Pearson et al. (1988) Proc. Natl. Acad. Sci. USA 85:2444 (other programs include the GCG program package (Devereux, J., et al., Nucleic Acids Research 12(I):387 (1984)), BLASTP, BLASTN, FASTAAtschul, SF, et al., J Molec Biol 215:403 (1990); Guide to Huge Computers, Martin J. Bishop, ed., Academic Press, San Diego, 1994, and Carillo et al. (1988) SIAM J Applied Math 48:1073). For example, identity can be determined using the BLAST function of the National Center for Biotechnology Information database. Other commercially available or publicly available programs include the DNAStar "MegAlign" program (Madison, Wis.) and the University of Wisconsin Genetics Computer Group (UWG) "Gap" program (Madison Wis.).In some embodiments, the nucleic acid sequence (e.g., DNA sequence), e.g., of the homology arms of the repair template, is at least or about 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 109, 109, 102, 103, 104, 105, 106, 107, 108, 109 ... , 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical.
[0035] As used herein, "homology arm" refers to a polynucleotide suitable for targeting a donor sequence to a genome by homologous recombination.Typically, two homology arms flank the donor sequence, and each homology arm comprises the genomic sequence upstream and downstream of the integration locus.
[0036] The term "lagomorpha" refers to a taxonomic order of rodent herbivorous mammals, consisting of two families (genera Leporidae and Ochotonidae which make up the family Leporidae) that include rabbits, hares, and pikas, having two pairs of anteroposteriorly arranged incisors on the upper jaw, usually soft hair, and a short or vestigial tail.
[0037] The term "Macropodidae" refers to the taxonomic family of diplodont marsupial mammals, including kangaroos, wallabies, and mouse-kangaroos, which are all jumping animals with long hind limbs and weakly developed forelimbs and are generally harmless terrestrial herbivores.
[0038] The term "monotremata" refers to the taxonomic order of egg-laying mammals that includes the platypus and echidnas.
[0039] The term "provirus" refers to the genome of a virus when integrated or inserted into the DNA of a host cell. Provirus refers to the double-stranded DNA form of a retroviral genome that is linked to a cellular chromosome. Provirus is created by reverse transcription of the RNA genome and subsequent integration into the chromosomal DNA of a host cell.
[0040] The term "primate" refers to a taxonomic order of mammals that is particularly characterized by highly developed binocular vision resulting in stereoscopic depth perception, specialized hands and feet for grasping, and enlarged cerebral hemispheres, and includes humans, apes, monkeys, and closely related species (e.g., lemurs and tarsiers).
[0041] As used herein, "Rep" refers to any nonstructural replicase, Rep protein, or combination of Rep proteins that are capable of providing the necessary functions to enable replication of the viral genome.
[0042] The term "Rodentia" refers to a taxonomic order of relatively small rodent mammals (e.g., mice, squirrels, or beavers) that have a single pair of chisel-edged incisors on both jaws. It includes all rodents.
[0043] The term "subject" or "patient" refers to any healthy or diseased animal, mammal, or human, or any animal, mammal, or human. In some embodiments, the subject is suffering from a hematological disease. In various embodiments of the methods of the invention, the subject is not undergoing treatment. In other embodiments, the subject is undergoing treatment.
[0044] The term "synteny" refers to the similar organization or ordering of a set of genes in different species.
[0045] A "therapeutically effective amount" of a substance or cell or virion is an amount that is capable of producing a medically desirable result (e.g., clinical improvement) at an acceptable benefit / risk ratio in a treated patient, preferably a human or non-human mammal.
[0046] The term "taxonomic order" refers to an orderly grouping of plants and animals according to their presumed natural relationships. The relatedness of species, based on the analysis of genomic sequence data, provides a quantitative alternative approach to natural relationships inferred from physical relationships.
[0047] The term "treating" includes preventive and / or therapeutic treatment. The term "preventive or therapeutic" treatment is recognized in the art and includes administering one or more of the compositions described herein to a subject. When administered before the clinical manifestation of an undesirable condition (e.g., disease, or other undesirable condition of a subject), the treatment is preventive (i.e., protects the subject from developing the undesirable condition), whereas when administered after the manifestation of an undesirable condition, the treatment is therapeutic (i.e., intended to reduce, ameliorate, or stabilize an existing undesirable condition or its side effects). Genome Safe Harbor (GSH)
[0048] The term "genomic safe harbor", also interchangeably referred to herein as "GSH" or "safe harbor gene" or "safe harbor locus", refers to a location within a genome, including a region or specific site of genomic DNA that can be used to integrate an exogenous nucleic acid, where integration does not cause any significant adverse effects on the growth of the host cell due to the addition of the exogenous nucleic acid alone. That is, GSH refers to a gene or locus within a genome into which a nucleic acid sequence can be inserted such that the sequence can be integrated and function in a predictable manner (e.g., express a protein of interest) without significant negative consequences on the activity of the endogenous gene or promotion of cancer. For example, GSH is a site in the host cell genome that can accommodate the integration of new genetic material in a manner that ensures that the newly inserted genetic element (i) functions predictably (e.g., expresses predictably), (ii) does not cause significant changes to the host genome, thereby avoiding risks to the host cell or organism, and (iii) preferably, the inserted nucleic acid is not perturbed by any read-through expression from nearby genes, and (iv) does not activate nearby genes. GSH can be a specific site or a region of genomic DNA. GSH can be a chromosomal site that can stably and reliably express a transgene in all tissues of interest without adversely affecting endogenous gene structure or expression. In some embodiments, GSH is a locus or gene where the insertion of foreign nucleic acid does not significantly alter the ability of the cell to differentiate correctly (e.g., differentiate stem cells). In some embodiments, GSH is also a locus or gene where the inserted nucleic acid sequence can be expressed efficiently and at a higher level than non-safe harbor sites.
[0049] Thus, GSH includes intragenic, intergenic or extragenic regions of human and model species genomes that can accommodate predictable expression of newly integrated DNA without significant deleterious effects on the host cell or organism. GSH may include intronic or exonic gene sequences as well as intergenic or extragenic sequences. Without being limited by theory, a useful safe harbor must allow sufficient transgene expression to give rise to the desired levels of transgene-encoded protein or non-coding RNA. GSH must also not predispose cells to malignant transformation, must not interfere with progenitor cell differentiation, and must not significantly alter normal cellular function. What distinguishes GSH from fortuitous successful integration events is the predictability of outcomes based on prior knowledge and validation of GSH.
[0050] In some embodiments, GSH allows for safe targeted gene delivery utilizing highly specific nucleases with minimal off-target activity while having limited off-target activity and minimal risk of causing genotoxicity or insertional carcinogenesis upon integration of foreign DNA. Identification of genomic safe harbors
[0051] Exemplary methods of identifying GSH locus are provided herein.In some embodiments, any one of the exemplary methods is used to identify GSH locus.In some embodiments, a combination of at least two exemplary methods is used to identify GSH locus.In some embodiments, a combination of at least three exemplary methods is used to identify GSH locus.Any one or combination of multiple exemplary methods can optionally further include at least one assay (in vitro, ex vivo, or in vivo) for verifying the identified GSH locus. Method 1: Functional identification of the GSH locus by random insertion of markers
[0052] In certain aspects, a method for identifying genomic safe harbor (GSH) loci is provided herein, comprising: (a) inducing random insertion of at least one marker gene into genome in a cell; (b) determining the stability and / or level of expression of the marker gene; and (c) identifying the genomic locus in which the inserted marker gene shows stable and / or high level of expression as GSH.In a preferred embodiment, the method further comprises: (a) identifying the genomic locus in which the inserted marker gene does not affect cell viability, and / or (b) identifying the genomic locus in which the inserted marker gene does not affect the differentiation ability of the cell.Thus, in some embodiments, the insertion of the marker gene at the GSH locus does not affect the pluripotency, totipotency, or multipotency of the cell (e.g., stem cell or progenitor cell).
[0053] In some embodiments, the cell used in the method is selected from cell line, primary cell, stem cell, or progenitor cell.In some embodiments, the cell is stem cell.In some such embodiments, the stem cell is selected from embryonic stem cell, tissue-specific stem cell, mesenchymal stem cell, and induced pluripotent stem cell (iPSC).
[0054] In some embodiments, the cells used in the method are selected from hematopoietic stem cells, hematopoietic CD34+ cells, and epidermal stem cells, epithelial stem cells, neural stem cells, lung progenitor cells, and liver progenitor cells.
[0055] In some embodiments, the cells used in the methods are mammalian cells. In some such embodiments, the mammalian cells are mouse cells, canine cells, porcine cells, non-human primate (NHP) cells, or human cells.
[0056] In certain embodiments, the random insertion of at least one marker gene into the genome of a cell is induced by (a) transfecting the cell with a nucleic acid molecule containing the marker gene, optionally a plasmid; or (b) transducing the cell with an integrating virus containing the marker gene. In some embodiments, the random insertion is induced by transducing the cell with an integrating virus containing the marker gene, and the integrating virus is a retrovirus. In some embodiments, the retrovirus is a gamma retrovirus.
[0057] In certain embodiments, the method uses at least one marker gene, including a screenable marker and / or a selectable marker. In some embodiments, the screenable marker gene encodes green fluorescent protein (GFP), beta-galactosidase, luciferase, and / or beta-glucuronidase. In some embodiments, the selectable marker gene is an antibiotic resistance gene. In some such embodiments, the antibiotic resistance gene encodes blasticidin S-deaminase or amino 3'-glycosylphosphotransferase (neomycin resistance gene).
[0058] In certain embodiments, the method uses a marker gene that is not operably linked to a promoter.Here, the use of a promoterless marker allows the identification of the GSH locus that uses nearby promoters and control elements to allow the expression of exogenous nucleic acid.In some embodiments, the nearby promoter is a tissue-specific promoter.
[0059] In certain embodiments, the marker gene is operably linked to a promoter. In some embodiments, the promoter is a tissue-specific promoter.
[0060] In some embodiments, the GSH identified is an intragenic (e.g., exonic or intronic) or intergenic form. In preferred embodiments, the GSH identified is an intronic or intergenic form. Method 2: Identification of the GSH locus using endogenous viral elements (EVEs)
[0061] In certain embodiments, methods are provided herein for identifying the GSH locus using evolutionary biology to identify any proviral remnants (e.g., parvoviral remnants), for example, in the genomes of metazoan species, called endogenous viral elements (EVEs). The results described herein demonstrate that EVEs can be acquired in the germline of progenitor species before species divergence, and thus all evolved or descendant species retain the EVE allele. Meanwhile, closely related species that evolved or diverged before the "endogenization" event retain the empty locus. As a purely illustrative example, loci occupied by intergenic EVEs in Macropodidae (kangaroos and related species) can be identified in other marsupials, including Didelphis virginiana (Didelphis virgiana) (North American opossum). These unoccupied loci are identifiable in other taxonomic families and although the EVE open reading frame is disrupted, viral sequences indicate foreign DNA inserted into the genome of totipotent embryonic cells, thus identifying candidate genomic safe harbor loci.The rationale for identifying EVE as a GSH locus is that insertion at the EVE locus did not affect the viability, function, growth, differentiation and speciation of the organism, thereby providing an inactive site that allows for the insertion of foreign nucleic acid.
[0062] In some embodiments, EVE is intragenic or intergenic. In some embodiments, EVE is intragenic. In some embodiments, EVE is intronic or exonic. In some embodiments, EVE is intronic. For example, in some embodiments, GSH locus is an exonic locus that tolerates EVE insertion in evolutionary lineage. In preferred embodiments, GSH is an intronic or exonic locus. Such locus is less likely to disrupt the function and structure of nearby genes or regulatory sequences by inserting actively transcribed exogenous nucleic acid.
[0063] In certain aspects, provided herein are methods for identifying a GSH locus, the methods comprising: (a) determining the presence and location of an endogenous viral element (EVE) in the genome of a metazoan species; (b) determining an intergenic or intronic boundary proximal to the EVE; and (c) identifying an intergenic or intronic locus that contains the EVE as a GSH locus.
[0064] In some embodiments, the presence and location of EVEs are determined by searching in silico for sequences homologous to viral elements. In some embodiments, EVEs in metazoan species are homologous to sequences of viral elements by at least, about or up to 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 143%, 14 , 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% identical to the sequence of the present invention.
[0065] In some embodiments, an intergenic or intronic boundary proximal to an EVE is determined by aligning sequences adjacent to the EVE with its orthologous sequence in one or more species for which intergenic or intronic boundaries are known. ..., with at least, about, or up to 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 102%, 104%, 105%, 106%, 107%, 108%, 109%, 109%, 109%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 109%, 109%, 1 and / or a sequence that is 8%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% or 100% identical to the sequence of the present invention.
[0066] In some embodiments, the method identifies the GSH locus to be within a mammalian genome, optionally wherein the mammalian genome is a mouse genome, a dog genome, a pig genome, a NHP genome, or a human genome.
[0067] In some embodiments, the EVE comprises a provirus, which is a viral genome integrated into the DNA of a non-viral host cell. In some embodiments, the EVE comprises a portion or fragment of a viral genome. In some embodiments, the EVE comprises a provirus from a retrovirus. In some embodiments, the EVE is not from a retrovirus. In some embodiments, the EVE comprises a provirus or a fragment of a viral genome from a non-retrovirus.
[0068] In some embodiments, the EVE comprises a viral nucleic acid, a viral DNA, or a DNA copy of a viral RNA. In some embodiments, the EVE comprises a viral nucleic acid. In some embodiments, the EVE or the viral nucleic acid in the EVE encodes a structural or non-structural viral protein, or a fragment thereof.
[0069] In some embodiments, the EVE comprises viral nucleic acid from a retrovirus. In some embodiments, the EVE comprises viral nucleic acid from a non-retrovirus, parvovirus, and / or circovirus. In some embodiments, the parvovirus is selected from B19, minute virus of mice (mvm), RA-1, AAV, bufavirus, hocovirus, bocavirus, and any one of the parvoviruses described herein (e.g., the parvoviruses listed in Tables 1A-1D). In some embodiments, the parvovirus is AAV. In some embodiments, the viral nucleic acid is from a circovirus. In some embodiments, the circovirus is a porcine circovirus (PCV) (e.g., PCV-1, PCV-2). In some embodiments, the viral nucleic acid in the EVE comprises non-retroviral nucleic acid. In some embodiments, the non-retroviral nucleic acid encodes a non-structural or structural viral protein (e.g., a rep (replication) protein, or a cap (capsid) protein, respectively).
[0070] In some embodiments, the EVE or viral nucleic acid encodes structural or nonstructural viral proteins. In some embodiments, the EVE or viral nucleic acid encodes Rep and assembly-activated nonstructural (NS) proteins (e.g., those required for viral replication, capsid assembly, etc.), and / or structural (S) viral proteins (capsid proteins, e.g., VP). Such proteins include, but are not limited to, Rep (replication) proteins, including but not limited to Rep78, Rep68, Rep52, and Rep40; and Cap (capsid proteins), including but not limited to VP1, VP2, and VP3, for example, from AAV. Structural proteins also include, but are not limited to, structural proteins A, B, and C, for example, from AAV. In some embodiments, the EVE is a nucleic acid encoding all or part of the nonstructural (NS) proteins or structural (S) proteins disclosed in Supplementary Table S2 of Francois et al. "Discovery of parvovirus-related sequences inan unexpected broad range of animals." Nature Scientific reports 6 (2016).
[0071] In some embodiments, methods for identifying GSH in mammalian genomes include initial sequencing and / or in silico analysis of sequences of genomic DNA inferred from progenitor species from multiple species within a taxonomic rank to identify endogenous viral elements (EVEs) or proviral nucleic acid insertions in the genomic DNA.
[0072] In some embodiments, the genome sequence of metazoan species is analyzed for the presence of EVE. The metazoan species can be from any phylogenetic taxon, including but not limited to Cetacea, Chiropetera, Lagomorpha, and Macropodiadae. Thus, in some embodiments, the metazoan species is selected from Cetacea, Chiropetera, Lagomorpha, and Macropodiadae. Other metazoan species, such as rodentia, primate, monotremata, can also be evaluated. Other species can be used, such as those listed in Figures 4A, 4B of Lui et al, J Virology 2011; 9863-9876, the entirety of which is incorporated herein by reference.
[0073] In some embodiments, EVE comprises nucleic acid from Parvovirus, which is a virus of the Parvoviridae family.The Parvoviridae family contains two subfamilies: Parvovirinae, which infect vertebrate hosts, and Densovirinae, which infect invertebrate hosts.Each subfamily is subdivided into several genera.
[0074] In some embodiments, the EVE comprises nucleic acid from any one of the following genera from Densovirinae: ambidensovirus, brevidensovirus, hepandensovirus, iteradensovirus, and penstyldensovirus.
[0075] In some embodiments, the EVE comprises nucleic acid from any one of the following genera from Parvovirinae: amdoparvovirus, aveparvovirus, bocaparvovirus, copiparvovirus, dependoparvovirus, erythroparvovirus, protoparvovirus, and tetraparvovirus. In some embodiments, the EVE comprises nucleic acid from an erythroparvovirus or a dependoparvovirus. In some embodiments, the EVE is from the subfamily Densovirinae, which includes the following genera: Genus Ambidensovirus. Type species: Lepidopteran ambidensovirus 1. The genus contains 11 recognized species. b. Genus Brevidensovirus. Type species: Dipteran brevidensovirus 1. The genus contains two recognized species. c. Genus Hepandensovirus. Type species: Decapod densovirus 1. The genus contains a single recognized species. D. Genus Iteradensovirus. Type species: Lepidopteran iteradensovirus 1. The genus contains five recognized species. Genus e. Penstyldensovirus. Type species: Decapod penstyldensovirus 1. The genus contains a single recognized species. f. Unassigned genus. Type species: Orthopteran densovirus 1. The genus contains a single recognized species.
[0076] In some embodiments, the EVE is from the subfamily Parvovirinae, which includes the following genera: Genus Amdoparvovirus. Type species: Carnivore amdoparvovirus 1. The genus contains four recognized species that infect minks and foxes. b. Genus Aveparvovirus. Type species: Galliform aveparvovirus 1. The genus contains a single species that infects turkeys and chickens. c. Genus Bocaparvovirus. Type species: Ungulate bocaparvovirus 1. The genus contains 21 recognized species that infect mammals from multiple orders, including primates. d. Genus Copiparvovirus. Type species: Ungulate copiparvovirus 1. The genus contains two recognized species that infect pigs and cows. e. Dependoparvovirus genus. Type species: Adeno-associated dependoparvovirus A. The genus contains seven recognized species that infect mammals, birds or reptiles. f. Genus Erythroparvovirus. Type species: Primate erythroparvovirus 1. The genus includes six recognized species that infect mammals, particularly primates, chipmunks or cows. g. Genus Protoparvovirus. Type species: Rodent protoparvovirus 1. The genus contains 11 recognized species that infect mammals from multiple orders, including primates. h. Genus Tetraparvovirus. Type species: Primate tetraparvovirus 1. The genus contains six recognized species that infect primates, bats, pigs, cows and sheep. Table 1A: Exemplary viruses of Erythroparvovirus in the subfamily Parvovirinae [Table 1A]
[0077] Table 1B: Exemplary viruses in the subfamily Parvovirinae [Table 1B] Table 1C: Exemplary viruses of the Protoparvovirus family in the Parvovirinae subfamily. [Table 1C-1] [Table 1C-2]
[0078] Table 1D: Exemplary viruses of Tetraparvovirus in the subfamily Parvovirinae [Table 1D-1] [Table 1D-2]
[0079] The Parvovirinae subfamily is primarily associated with warm-blooded animal hosts. Of these, RA-1 virus in the parvovirus genus, B19 virus in the erythrovirus genus, and adeno-associated viruses (AAV) 1-9 in the dependovirus genus are human viruses. In some embodiments, EVE is recognized in five genera: Bocaparvovirus (human bocavirus 1-4, HboV1-4), Dependoparvovirus (adeno-associated virus; at least 12 serotypes have been identified), Erythroparvovirus (parvovirus B19, B19), Protoparvovirus (bufavirus 1-2, BuV1-2), and Tetraparvovirus (human parvovirus 4 Gl-3, PARV4 Gl-3).
[0080] In some embodiments, the EVE is from a parvovirus, and in some embodiments, the EVE comprises nucleic acid from AAV (adeno-associated virus). Adeno-associated virus (AAV), a member of the Parvovirus family, is a small, non-enveloped, icosahedral virus with a single-stranded linear DNA genome of 4.7 kilobases (kb) to 6 kb. AAV has been assigned to the Dependoparvovirus genus and was originally called an adenovirus-associated (or satellite) virus because the virus was found as a contaminant in purified adenovirus stocks. The life cycle of AAV includes a latent phase, during which the AAV genome may be integrated into the host cell chromosomal DNA after infection, often at a defined locus, such as AAVS1; and a lytic phase, during which cells are co-infected with adenovirus or herpes simplex virus and AAV, or superinfect latently infected cells, after which the integrated genome is rescued, replicated, and packaged into an infectious virus. Based on serological surveillance analyses, exposure to AAV is frequent in humans and other primates, and several serotypes have been isolated from various tissue samples. Serotypes 2, 3, 6, and 13 have been found in cultured human cells, AAV5 has been isolated from clinical specimens, while AAV serotypes 1, 4, and 7–12 have been isolated from non-human primate (NHP) tissue samples or cells. As of 2013, 13 AAV serotypes have been described. Weitzman, et al. (2011). "Adeno-Associated Virus Biology."In Snyder, RO; Moullier, P. Adeno-associated virus methods and protocols.Totowa, NJ: Humana Press. ISBN 978-1- 61779-370-7;Mori S, et al., (2004)."Two novel adeno-associated viruses from cynomolgus monkey: pseudotyping characterization of capsid protein." Virology 330 (2): 375-83).
[0081] In some embodiments, the EVE is a nucleic acid or a portion of a nucleic acid from any of the parvoviruses listed in Tables 1A-1D; or a nucleic acid or a portion of a nucleic acid from any of the parvoviruses listed in Tables 1A-1D that is at least, about, or up to 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, These include nucleic acids comprising a sequence having 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% identity to the nucleic acid.
[0082] In some embodiments, the EVE is at least, about or up to 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 1090%, 1091%, 1092%, 1093%, 1094%, 1095%, 1096%, 1097%, 1098%, 1099%, 1000%, 1001%, 1002%, 103%, 104%, 105%, 106%, 107, 108, 1090, 1091, 1000%, 1095, 1001, 1002, 1003, 1004, 105, 106, 107, 108, 109, 109, 1005, 1006, 1007, 1008, 1009, 1009, 1008, 1 These include nucleic acids containing sequences having 2%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% identity to the nucleic acid. In some embodiments, the AAV is selected from serotypes AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV 10, AAV11, AAV12, or AAV13.
[0083] In some embodiments, the EVE is at least or about 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 109%, 109%, 108%. 2%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% nucleic acid or amino acid sequence identity to the nucleic acid sequence of the present invention. Method 3: A method to identify GSH loci in orthologous organisms
[0084] In certain aspects, provided herein are methods for identifying a GSH locus in orthologous organisms, the methods comprising: (a) identifying a GSH locus in species A according to any one of the methods described herein (e.g., using a functional method (Method 1) or a method utilizing EVE (Method 2)); (b) determining the location of (i) at least one cis-acting element proximal to the GSH locus in species A and (ii) the corresponding cis-acting element in species B; and (c) identifying the locus in species B as a GSH locus, wherein the distance between the locus and the at least one cis-acting element in species B is substantially proportional to the distance between the GSH locus and the corresponding cis-acting element in species A.
[0085] As described herein, at least one cis-acting element proximal to the GSH locus in species A and / or species B may be known, or alternatively, the location of such elements may be determined by sequence analysis (e.g., by aligning sequences adjacent to the GSH locus in one or more organisms in which at least one cis-acting element proximal to the GSH locus is known with their orthologous sequences). In some embodiments, at least one cis-acting element in species A or B is at least or about 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 109, 102%, 104%, 105%, 106%, 107%, 108%, 109, 109, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109, 109, 104%, 105%, 109, 109, 102%, 10 %, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% identical to the sequence of the present invention. In some embodiments, at least one cis-acting element proximal to the GSH locus in species A is at least or about 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109 ... 9%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% or 100% identical.
[0086] Alternatively, one of skill in the art would know how to experimentally determine at least one cis-acting element proximal to the GSH locus (e.g., determining the RNA sequence by RNA seq or by cloning the cDNA; and comparing it to the genomic sequence to map splice donor sites, splice acceptor sites, polyadenylation sites, etc.).
[0087] Many cis-acting elements are known in the art.In some embodiments, at least one cis-acting element is selected from splicing donor site, splicing acceptor site, polypyrimidine tract, polyadenylation signal, enhancer, promoter, terminator, splicing control element, intron splicing enhancer and intron splicing silencer.
[0088] In certain embodiments, the at least one cis-acting element comprises two or more cis-acting elements.
[0089] In some embodiments, the at least one cis-acting element comprises two cis-acting elements, a first cis-acting element located upstream (i.e., 5') of the GSH locus and a second cis-acting element located downstream (i.e., 3') of the GSH locus.
[0090] In some embodiments, the distance between at least one cis-acting element and the GSH locus relative to the distance between the two cis-acting elements in species B is substantially proportional to the distance between the corresponding cis-acting element and the GSH locus relative to the distance between the two cis-acting elements in species A.
[0091] In some embodiments, the distance between the at least one cis-acting element and the GSH locus in species B is at least, about, or at most 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 100%, 100% or 100% of the distance between the at least one cis-acting element and the GSH locus in species A. 10%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410 %, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, 500%, 510%, 520%, 530%, 540%, 550%, 560%, 570%, 580%, 590%, 600%, 610%, 620%, 630%, 640%, 650%, 660%, 670%, 680%, 690%, 700%, 710%, 720%, 730%, 740%, 750%, 760%, 770%, 780%, 790%, 800%, 810%, 820%, 830%, 840%, 850%, 860%, 870%, 880%, 890%, 900%, 910%, 920%, 930%, 940%, 950%, 960%, 970%, 980%, 990%, or 1000%.
[0092] In some embodiments, the distance between the at least one cis-acting element and the GSH locus in species B is at least 20% but up to 500% of the distance between the at least one cis-acting element and the GSH locus in species A.
[0093] In some embodiments, the distance between the at least one cis-acting element and the GSH locus in species B is at least 80% but at most 250% of the distance between the at least one cis-acting element and the GSH locus in species A.
[0094] In some embodiments, the distance between the at least one cis-acting element and the GSH locus in species B is at least 90% but at most 110% of the distance between the at least one cis-acting element and the GSH locus in species A.
[0095] In some embodiments, the method identifies a GSH locus in a mammalian genome. In some embodiments, the mammalian genome is a mouse genome, a dog genome, a pig genome, a NHP genome, or a human genome.
[0096] Any one of the methods of identifying a GSH locus described above may further include steps and / or considerations in any other method, i.e., any number of the methods described herein may be combined in any order. For example, functional identification of a GSH locus by method 1 may further include steps and / or considerations of method 2 (e.g., identifying EVE). Method 1 may further include steps and / or considerations of method 3 (e.g., identifying a GSH locus in an orthologous organism). Similarly, method 2 may further include steps and / or considerations of method 3. Alternatively, method 1 may further include steps and / or considerations of method 2 and method 3. Optional criteria for selecting the GSH locus or nucleic acid region of GSH
[0097] In some embodiments, the GSH identified according to the methods described herein is an extragenic or intergenic site that is far from known genes or genomic regulatory sequences, or an intragenic site (within a gene) that is likely to be amenable to disruption.
[0098] In some embodiments, GSH may comprise a gene, including intragenic DNA, including intron or exon gene sequences.
[0099] In some embodiments, in addition to validating the GSH identified using the functional in vitro and in vivo analyses disclosed herein, the candidate GSH can be assessed, if desired, using bioinformatics to determine, for example, whether the candidate GSH meets certain criteria, for example, but not limited to, by assessing any one or more of the following: proximity to oncogenes or proto-oncogenes, location within a gene or near the 5' end of a gene, location within a selected housekeeping gene, location in an extragenic region, proximity to mRNA, proximity to ultraconserved regions, and proximity to long non-coding RNA and other such genomic regions. By way of example, the previously identified GSH AAVS1 (adeno-associated virus integration site 1) has been identified as an adeno-associated virus consensus integration site on chromosome 19, located at chromosome 19 (position 19q13.42), and was primarily identified as a recurrent recovery site of wild-type AAV integration in the genome of cultured human cell lines infected with AAV in vitro. Integration at the AAVS1 locus disrupts the gene phosphatase 1 regulatory subunit 12C (PPP1R12C; also known as MBS85), which encodes a protein whose function has not been clearly detailed. The biological consequences of disrupting one or both alleles of PPP1R12C are currently unknown. Neither gross abnormalities nor differentiation defects were observed in human and mouse pluripotent stem cells carrying a targeted transgene at AAVS1. Previous evaluations of the AAVS1 site have generally used Rep-mediated targeting, which preserves the functionality of the targeted allele and maintains expression of PPP1R12C at levels comparable to those in non-targeted cells. AAVS1 has also been evaluated using ZFN-mediated recombination into iPSCs or CD34+ cells.
[0100] Initially characterized, the AAVS1 locus is >4kb and was identified as chromosome 19 nucleotides 55,113,873-55,117,983 (human genome assembly GRCh38 / hg38), overlapping with the first exon of the PPP1R12C gene encoding protein phosphatase 1 regulatory subunit 12C. This >4kb region is extremely rich in G+C nucleotide content, a particularly gene-rich region of chromosome 19 (see Fig. 1A in Sadelain et al, Nature Revs Cancer, 2012; 12; 51-58), and some integrated promoters may indeed activate or cis-activate neighboring genes, although the consequences in different tissues are currently unknown.
[0101] AAVS1 GSH was identified by characterizing AAV proviral structures in latently infected human cell lines using a recombinant bacteriophage genomic library generated from a latently infected clonal cell line (Detroit 6 clone 7374 IIID5) (Kotin and Bems 1989), in which Kotin et al. isolated nonviral cellular DNA flanking the provirus and used a subset of the "left" and "right" flanking DNA fragments as probes to screen an independently derived panel of latently infected clonal cell lines. In approximately 70% of the clonal isolates, AAV DNA was detected by the cell-specific probe (Kotin et al. 1991; Kotin et al. 1990). Sequence analysis of the preintegration site identified near homology to a portion of the AAV inverted terminal repeat (Kotin, Linden, and Bems 1992). Although lacking the characteristic interrupted palindrome, the AAVS1 locus retained the p5 Rep protein that binds and nicks, also called the terminal splitting site (Chiorini et al. 1994; Chiorini et al. 1995; Im and Muzyczka 1989, 1990, 1992). Interestingly, the human orthologue functions as a p5 Rep in vitro origin of DNA synthesis, thus supporting earlier inferences that AAVS1 integration is a Rep-dependent process (Kotinet al., 1990; Kotin et al., 1992; Urcelay et al. 1995; Weitzman et al. 1994). Rep binding elements in cis were shown to be required for AAV integration and provide further support for the involvement of Rep proteins in the targeted non-homologous recombination process (Urabe,et al., Linden, Bems). These elements define a minimal origin for Rep-mediated DNA synthesis, as does the arrangement of Rep binding and nicking sites that allows RNA-primer-independent strand-displacement DNA (leading strand) synthesis.
[0102] Wild-type adeno-associated viruses can cause either productive or latent infections, in which the wild-type viral genome often integrates into the AAVS1 locus on human chromosome 19 in cultured cells (Kotin and Bems 1989; Kotin et al. 1990). This unique aspect of AAV has been exploited as one of the first so-called "safe harbors" for iPSC genetic modification. AAVS1, according to its original definition (Kotin et al., 1991), is located between nucleotides 55,113,873 and 55,117,983 on chromosome 19 (human genome assembly GRCh38 / hg38) and overlaps with the first exon of the PPP1R12C gene, which encodes protein phosphatase 1 regulatory subunit 12C. Interestingly, the PPP1R12C first exon, 5' untranslated region, contains a functional AAV origin of DNA synthesis (Urcelay et al. 1995) shown in the following sequence: GCTC Rep binding motif and terminal cleavage site (GGTTGG) are shown in bold font: [ka] .
[0103] Surprisingly, the human chromosome 19 AAVS1 safe harbor is within an exon region of PPP1R12C, a gene encoding protein phosphatase control 1 regulatory subunit 12C. The choice of exon integration site is not obvious and perhaps counterintuitive, since the insertion and expression of foreign DNA would likely disrupt the expression of endogenous genes. Apparently, insertion of the AAV genome into this locus does not adversely affect cell viability or iPSC differentiation (DeKelver et al. 2010;Wang et al. 2012;Zou et al. 201 1). Integration occurs by non-homologous recombination, which requires the presence of AAV Rep proteins in trans on both recombination substrates and a minimal origin of AAV DNA synthesis in cis, thus allowing Rep protein-mediated juxtaposition of AAV and genomic DNA (Weitzman et al. 1994).
[0104] The Rep-dependent minimal origin of DNA synthesis consists of a p5 Rep protein-binding element (RBE) and appropriately positioned terminal cleavage sites (trs), as exemplified by AAV2 trs AGT|TGG and AAV5 trs AGTG|TGG (vertical lines indicate nicking sites). In addition, the involvement of cellular protein complexes has been inferred but has not yet been identified or characterized.
[0105] These viral replication elements must function very efficiently, otherwise the virus will die out due to lack of replication fitness, whereas the small, non-coding, ∼35 bp element in AAVS1 may not function at all in the host. However, the AAVS1 locus is established as a somatic safe harbor, and disruption of the locus in totipotent or germline cells may interfere with ontogeny processes.
[0106] The AAVS1 locus is within the 5'UTR of the highly conserved PPP1R12C gene. A Rep-dependent minimal origin of DNA synthesis is conserved in the 5'UTR of the human, chimpanzee and gorilla PPP1R12C gene. However, in rodent species (mouse and rat), substitutions occur more frequently within the preferred terminal cleavage site compared to the adjacent non-coding DNA. Accidental, rather than selected or acquired, genotypes may affect the efficiency of other species' specific sequences within the 5'UTR.
[0107] In some embodiments, a candidate GSH identified according to embodiments herein is identified as meeting the criteria for GSH if it has limited off-target activity and poses minimal risk of causing genotoxicity or insertional carcinogenesis upon integration of foreign DNA, while being able to achieve safe targeted gene delivery with minimal off-target activity and utilizing highly specific nucleases.
[0108] GSH is validated based on the in vitro and in vivo assays described herein, but in some embodiments, additional selection based on determining whether GSH meets certain criteria may be used. For example, in some embodiments, the GSH locus identified herein is located in the exon, intron, or untranslated region of an unnecessary gene. Analysis shows that the integration site of proviruses in tumors is generally located near the start of transcription, either upstream of the transcription unit or right within the transcription unit, often within the 5' intron. Proviruses at these locations tend to dysregulate expression by increasing the transcription rate, either by viral promoters or by viral enhancer insertion. Thus, in some embodiments, the GSH locus identified herein is selected based on not being proximal to an oncogene. In some embodiments, GSH does not have an integration site located near the start of transcription of an oncogene, for example, upstream of or within the 5' intron of an oncogene or protooncogene. Such cancer genes are well known to those skilled in the art and are disclosed in Table 1 of Sadelain et al., Nature Revs Cancer, 2012; 12; 51-58, which reference is incorporated herein in its entirety. Exemplary databases of genes involved in cancer, such as the Atlas gene set, the CAN gene set, the CIS(RTCGD) gene set, and those listed in Table 2 below, are well known. Table 2: Exemplary database of genes involved in cancer. [Table 2] *Gene lists, and links to sources, are available at the Bushman lab Cancer Gene List website (see World Wide Web at bushmanlab.org / links / genelists). CAN, cancer; CIS, common insertion site; references in the last column represent reference numbers in Sadelain et al., Nature Revs Cancer (2012) 12:51-58.
[0109] In some embodiments, the GSH loci identified herein have one or more characteristics selected from: (i) outside of gene transcription units; (ii) located between 5-50 kilobases (kb) from the 5' end of any gene; (iii) located between 5-300 kb from cancer-associated genes; (iv) located between 5-300 kb from any identified microRNAs; and (v) outside of ultra-conserved regions and long non-coding RNAs. In some embodiments, the GSH loci identified herein have one or more of the following characteristics: (i) outside of gene transcription units; (ii) located >50 kilobases (kb) from the 5' end of any gene; (iii) located >300 kb from cancer-associated genes; (iv) located >300 kb from any identified microRNAs; and (v) outside of ultra-conserved regions and long non-coding RNAs. In a study of lentiviral vector integration in transduced induced pluripotent stem cells, analysis of over 5,000 integration sites revealed that approximately 17% of integrations occurred within safe harbors. Vectors that integrated within these safe harbors were able to express therapeutic levels of β-globin from their transgene without perturbing endogenous gene expression. Homology and sequence alignment
[0110] Homology, as used herein, refers to the percentage of nucleotide sequence identity between two regions of the same nucleic acid strand or between regions of two different nucleic acid strands. If a nucleotide residue position in both regions is occupied by the same nucleotide residue, the regions are homologous at that position. A first region is homologous to a second region if at least one nucleotide residue position in each region is occupied by the same residue. Homology between two regions is expressed by the percentage of nucleotide residue positions in the two regions that are occupied by the same nucleotide residue. As an example, a region having the nucleotide sequence 5'-ATTGCC-3' and a region having the nucleotide sequence 5'-TATGGC-3' share 50% homology. Preferably, the first region comprises the first portion and the second region comprises the second portion, such that at least about 50%, preferably at least about 75%, at least about 90%, or at least about 95% of the nucleotide residue positions of each of the portions are occupied by the same nucleotide residue. More preferably, all nucleotide residue positions of each of the portions are occupied by the same nucleotide residue.
[0111] With respect to nucleic acids, the term "substantial homology" refers to a homology that, when two nucleic acids, or designated sequences thereof, are optimally aligned and compared, includes at least about 60% of the nucleotides, and usually at least or about 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 109% or 109% of the nucleotides, including appropriate nucleotide insertions or deletions. , 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, more preferably at least about 97%, 98%, 99% or more identical. Alternatively, substantial homology exists when the segments hybridize under selective hybridization conditions, to the complement of the strand.
[0112] The percent identity between two sequences is a function of the number of identical positions shared by the sequences (i.e., % identity = # of identical positions / total # of positions x 100), taking into account the number of gaps and the length of each gap that need to be introduced for optimal alignment of the two sequences. The comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm, as described in the non-limiting examples below.
[0113] The percent identity between two nucleotide sequences can be determined using the GAP program of the GCG software package (available at the GCG Corporation website on the world wide web) using the NWSgapdna.CMP matrix and gap weights of 40, 50, 60, 70 or 80, and length weights of 1, 2, 3, 4, 5 or 6. The percent identity between two nucleotide or amino acid sequences can also be determined using the algorithm of E. Meyers and W. Miller (CABIOS, 4:11 17 (1989)) as incorporated into the ALIGN program (version 2.0) using a PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4. Additionally, percent identity between two amino acid sequences can be determined using the Needleman and Wunsch (J. Mol. Biol. (48):444 453 (1970)) algorithm incorporated into the GAP program of the GCG software package (available at the GCG website on the world wide web), using either a Blosum 62 matrix or a PAM250 matrix, and gap weights of 16, 14, 12, 10, 8, 6 or 4, and length weights of 1, 2, 3, 4, 5 or 6.
[0114] The nucleic acid and protein sequences of the present invention can further be used as a "query sequence" to perform searches against public databases, for example to identify related sequences. Such searches can be performed using the NBLAST and XBLAST programs (version 2.0) of Altschul, et al. (1990) J. Mol. Biol. 215:403 10. BLAST nucleotide searches can be performed with the NBLAST program, score=100, wordlength=12 to obtain nucleotide sequences homologous to the nucleic acid molecules of the present invention. BLAST protein searches can be performed with the XBLAST program, score=50, wordlength=3 to obtain amino acid sequences homologous to the protein molecules of the present invention. To obtain gapped alignments for comparison, gapped BLAST can be utilized as described in Altschul et al., (1997) Nucleic Acids Res. 25(17):3389 3402. When utilizing BLAST and Gapped BLAST programs, the default parameters of the respective programs (eg, XBLAST and NBLAST) can be used (available at the NCBI website on the world wide web). Validation of GSH using in vitro and in vivo assays
[0115] Without being limited by theory, a useful GSH region should allow sufficient transgene expression to obtain the desired levels of vector-encoded protein or non-coding RNA and should not predispose cells to malignant transformation or significantly negatively alter cellular function.
[0116] Methods and compositions for validating candidate GSH regions disclosed herein include, but are not limited to, bioinformatics, in vitro gene expression assays, in vitro and in vivo expression arrays to interrogate nearby genes, in vitro directed differentiation or in vivo reconstitution assays in xenotransplantation models, gene transfer at syntenic regions, and analysis of patient databases from individuals. Thus, any one or combination of methods for identifying GSH loci described herein may further include performing at least one in vitro, ex vivo, and / or in vivo.
[0117] In some embodiments, validation of GSH is determined to check for the absence of germline integration of the introduced gene, thereby reducing the risk of germline transmission of the gene therapy vector.
[0118] After identification of the target locus or candidate GSH, a series of in vitro and in vivo assays can be used to demonstrate safety, particularly the absence of oncogenic potential. In vitro oncogenicity assays can be based on experience in previous gene therapy T cell product characterization.
[0119] In some embodiments, GSH can be verified by several assays. In some embodiments, the functional assay is selected from any one or more of: (a) inserting a marker gene into a locus in human cells and measuring marker gene expression in vitro; (b) inserting a marker gene into an orthologous locus in progenitor or stem cells and engrafting the cells into immune-depleted mice and / or evaluating marker gene expression in all developmental lineages; (c) differentiating hematopoietic CD34+ cells with a marker gene inserted into a candidate GSH locus into terminally differentiated cell types; or (d) generating transgenic knock-in mice, whose genomic DNA has a marker gene inserted into a candidate GSH locus, and the marker gene is operably linked to a tissue-specific or inducible promoter.
[0120] In some embodiments, at least one in vitro, ex vivo, and / or in vivo assay comprises (a) de novo targeted insertion of a marker gene into a locus in a cell (e.g., a human cell), and determining (i) cell viability, (ii) insertion efficiency, and / or (iii) marker gene expression; (b) targeted insertion of a marker gene into a locus in a progenitor or stem cell and differentiating in vitro to determine (i) marker gene expression in all developmental lineages, and / or (ii) whether insertion of the marker gene affects differentiation of said progenitor or stem cell; (c) targeted insertion of a marker gene into a locus in progenitor or stem cells and engrafting the cells into immune-depleted mice to assess marker gene expression in all developmental lineages in vivo; d) targeted insertion of a marker gene into a locus in the cell and determining a global cellular transcription profile (e.g., using RNAseq or microarrays); and e) generating transgenic knock-in mice in which the genomic DNA of the mouse has a marker gene inserted into a locus, and optionally the marker gene is operably linked to a tissue-specific or inducible promoter; is selected from.
[0121] In some embodiments, the stem cells used in the validation assay are selected from embryonic stem cells, tissue-specific stem cells, mesenchymal stem cells, and induced pluripotent stem cells (iPSCs).In some embodiments, the cells, progenitor cells, or stem cells are selected from hematopoietic stem cells, hematopoietic CD34+ cells, and epidermal stem cells, epithelial stem cells, neural stem cells, lung progenitor cells, muscle satellite cells, intestinal K cells, and liver progenitor cells. Exemplary in vitro assays for validating GSH
[0122] In some embodiments, the functional assay for validating GSH comprises inserting a marker gene into a locus of a human cell and determining the expression of the marker in vitro. In some embodiments, the marker gene is introduced by homologous recombination. In some embodiments, the marker gene is operably linked to a promoter, e.g., a constitutive promoter or an inducible promoter. The determination and quantification of gene expression of the marker gene can be performed by any method commonly known to those skilled in the art, such as gene expression using, e.g., RT-PCR, Affymetrix gene array, transcriptome analysis; and / or protein expression analysis (e.g., Western blot). In some embodiments, the effect of the integrated marker transgene on the expression of neighboring genes is determined in vitro in cultured cells.
[0123] In some embodiments, the marker gene is introduced into mammalian cells, such as human cells or mouse cells or rat cells. In some embodiments, the cells are cell lines, such as fibroblast cell lines, HEK293 cells, etc. In some embodiments, the cells used in the assay are pluripotent cells, such as iPSCs, or clonable cell types, such as T lymphocytes. In some embodiments, gene expression of the insertion of the marker gene into a variety of different cell populations, including primary cells, is evaluated. In some embodiments, the iPSCs with the introduced marker gene are differentiated into multiple lineages to check the consistent and reliable gene expression of the marker gene in different lineages.
[0124] In some embodiments, a marker gene is inserted into a candidate GSH locus in the genome of hematopoietic cells, such as CD34+ cells, and differentiated into different terminally differentiated cell types.
[0125] In some embodiments, cell populations with marker genes introduced into the candidate GSH can be assessed for possible tissue dysfunction and / or transformation. For example, CD34+ cells or iPSCs are assessed for abnormal differentiation that is distinct from normal lineage differentiation, and / or increased proliferation that indicates cancer risk.
[0126] In some embodiments, the gene expression level of the proximal gene is determined. For example, in some embodiments, if the integrated marker gene results in abnormal gene expression or other abnormal regulation of surrounding or nearby gene expression, such as downregulation or upregulation of gene expression of the nearby gene, the candidate locus is not selected as a suitable GSH. In some embodiments, if no change in the expression level of the nearby gene is detected, the candidate locus is nominated or selected as a GSH. In some embodiments, the gene expression of adjacent, proximal or nearby genes is determined, and the proximal or nearby genes can be within about 350 kb, or about 300 kb, or about 250 kb, or about 200 kb, or about 100 kb, or between about 1 and 10 kb, or less than 1 kb from the insertion site of the marker gene (i.e., adjacent genes or RNA sequences either 5' or 3' of the insertion locus).
[0127] In some embodiments, the epigenetic traits and profiles of the targeted candidate GSH locus are assessed before and after introduction of a marker gene to determine whether introduction of a marker gene affects the epigenetic signature (e.g., histone modifications, DNA modifications, euchromatin or heterochromatin protein association, etc.) of GSH and / or surrounding or nearby genes within about 350 kb upstream or downstream of the integration site.
[0128] In some embodiments, the insertion of a marker gene into a candidate GSH locus is assessed to see if the locus can accommodate different integrated transcription units. In some embodiments, gene expression of a marker gene that is operably linked to a variety of different genetic elements, including promoters, enhancers, and chromatin determinants, including locus control regions, matrix attachment regions, and insulator elements, is assessed, and in some embodiments, gene expression of nearby genes within about 350 kb, or about 300 kb, or about 250 kb, or about 200 kb, or about 100 kb, or between 10 and 100 kb, or between about 1 and 10 kb, or less than 1 kb, (upstream or downstream) of the insertion site of the marker gene is also assessed.
[0129] In some embodiments, a marker gene that is not operably linked to a promoter is inserted into the GSH locus to assess the effect of any promoters and / or other regulatory elements of nearby genes.
[0130] In some embodiments, as demonstrated herein, the insertion of a marker gene into a candidate GSH locus is assessed to see if it alters the global transcription pattern. Such analysis can be accomplished, for example, by next generation sequencing (NGS) of DNA or RNA, Affymetrix gene arrays, and the like.
[0131] In some embodiments, when GSH locus is associated with a particular gene, knockdown of the gene can be evaluated to verify that the gene is not necessary or dispensable.As an illustrative example disclosed herein, SYNTX-GSH2 is surrounded by several different coding and RNA genes.Therefore, in some embodiments, the effect of RNAi knockdown of SYNTX-GSH2 on the cell function and gene expression of neighboring cells can be evaluated, and if knockdown of candidate gene in GSH locus has no significant effect, the gene can be verified as GSH.In addition, the in vitro assay of knockdown of GSH gene using RNAi is important to determine dispensability of the gene, especially due to biallelic destruction, as is often the case with endonuclease-mediated targeting.
[0132] In some embodiments, because cancer chemotherapy cytotoxic agents are potentially genotoxic and carcinogenic, standard in vitro studies for preclinical evaluation of these types of drugs can also be used to assess GSH locus disruption. For example, the ability of primary T cells to grow without cytokines and cell signaling is a hallmark of oncogenic transformation.
[0133] For example, in some embodiments, a marker gene can be introduced into a candidate GSH locus in T cells, e.g., SB-728-T cells, and cultured for several weeks without cytokine support to demonstrate that normal cell death occurs.
[0134] In other embodiments, the classical biological cell transformation assay is the anchorage-independent growth of fibroblasts, which is a rigorous test of carcinogenicity.Based on this, in some embodiments, marker gene can be inserted into the target GSH locus in fibroblasts and evaluated for anchorage-independent growth.Other in vitro assays or tests for evaluating carcinogenicity can be used, such as mouse micronucleus test, anchorage-independent growth, and mouse lymphoma TK gene mutation assay.
[0135] In some embodiments, the marker gene is selected from any of a fluorescent reporter gene, such as GFP, RFP, etc., and a bioluminescent reporter gene. Exemplary marker genes are described herein.
[0136] In some embodiments, marker or reporter gene sequences include, but are not limited to, DNA sequences encoding β-lactamase, β-galactosidase (LacZ), alkaline phosphatase, thymidine kinase, green fluorescent protein (GFP), chloramphenicol acetyltransferase (CAT), luciferase, and others known in the art. When associated with the control elements that drive their expression, reporter sequences provide signals that can be detected by conventional means, including enzymatic assays, radiographic assays, colorimetric assays, fluorescent assays or other spectroscopic assays, fluorescence-activated cell sorting assays, and immunological assays, including enzyme-linked immunosorbent assays (ELISAs), radioimmunoassays (RIAs), and immunohistochemistry. For example, when the marker sequence is the LacZ gene, the presence of the vector that conveys the signal is detected by assaying for β-galactosidase activity. In some embodiments, when the marker gene is green fluorescent protein or luciferase, the vector that conveys the signal can be measured by colorimetric analysis based on visible light absorption or light production in a luminometer, respectively. Such reporters can be useful, for example, to verify tissue-specific targeting capabilities of nucleic acids and tissue-specific promoter controlling activity.
[0137] In some embodiments, bioinformatics can be used to validate GSH, for example by reviewing sequences of a database of autologous patient-derived iPSCs, as described in Papapetrou et al., 2011, Na. Biotechnology, 29; 73-78, which is incorporated herein in its entirety.
[0138] In addition, once GSH and target integration sites within GSH are identified, bioinformatics and / or web-based tools can be used to identify potential off-target sites. For example, bioinformatics tools such as Predicted Report of Genome-wide Nuclease Off-Target Sites (PROGNOS, world wide web at baolab.bme.gatech.edu / Research / BioinformaticTools / prognos.html) and CRISPOR (world wide web at crispor.tefor.net / ) for designing CRISPR / Cas9 targets and predicting off-target sites. CRISPOR and PROGNOS can provide a report of potential genome-wide nuclease target sites for ZFN and TALEN. Once a specific target site is identified, the program can provide a list that ranks potential off-target sites. In vivo assay to verify GSH
[0139] In some embodiments, in vivo assays to functionally validate GSH may be performed. In some embodiments, in vivo evaluation of GSH may be performed in transgenic mice carrying a transgene integrated into the syntenic region.
[0140] In some embodiments, the in vivo functional assay for validating GSH comprises inserting marker gene into iPSC locus and transplanting into immunodeficient mice.In some embodiments, the insertion of marker gene into iPSC and modified iPSC implanted into immunodeficient mice is evaluated over a period of time.Such in vivo assay allows to evaluate any genotoxic events, including atypical or abnormal differentiation (e.g., changes in hematopoietic transformation and / or clonal hematopoietic bias), and to evaluate tumorigenic cell proliferation from rare events.
[0141] Such in vivo methods in immunodeficient mice using hematopoietic cells are well known to those skilled in the art and are disclosed in Zhou, et al. "Mouse transplant models for evaluating the oncogenic risk of a self-inactivating XSCID lentiviral vector." PloS one8.4 (2013): e62333, which is incorporated herein by reference in its entirety, in which the incidence of malignancies from the introduced modified hematopoietic cells or iPSCs can be assessed compared to controls or cells in which a marker gene has not been introduced into the target locus in GSH. In some embodiments, hematopoietic malignancies can be assessed. In some embodiments, the lineage distribution of peripheral blood cells in the recipient immunodeficient mice is assessed to determine myeloid bias and signals of insertional transformation or deleterious effects due to the marker gene inserted into the GSH locus.
[0142] In some embodiments, because recipient mouse strains are immunodeficient, when tumors arise in such mice, these tumors can be characterized and evaluated as being of human origin.If tumors are of human origin, their clonality will need to be further evaluated for the insertion of marker gene in GSH locus, or any abnormal regulation gene expression (upregulation or downregulation) of on- or off-target sites, such as adjacent RNA sequences or genes.However, the clonality observed in marker gene-transduced cells is not necessarily an equal causal relationship, and may instead be a harmless indicator that simply represents the clonal origin of tumors.
[0143] In some embodiments, an in vivo assay can be used that relies on the fact that human T cells can be maintained in immunodeficient NOG mice. Such an assay requires the introduction of marker genes into targeted GSH loci and modified human T cells, allowing them to survive and expand in NOG models for months, and comparing them with unmodified T cells. In some embodiments, a model with human T cell xeno-GVHD can be used, in which the maximum period of cell proliferation before animals die of GVHD is 2 months, and defines the dose and donor that will produce definite GVHD in NOG mice. After 2 months, animals are euthanized, and tissues are evaluated by histological examination for neoplasms, immunostaining to detect human cells, and gene expression analysis (e.g., Affymetrix array or RT-PCR of adjacent genes around the GSH insertion locus) to detect modified gene expression at on-target and off-target sites.
[0144] In some embodiments, another in vivo assay to functionally validate a candidate locus as GSH is to generate knock-in transgenic animals or mice. Testing for successful gene editing of marker genes into GSH in iPSCs or T lymphocytes or other host cells
[0145] The efficiency of insertion of marker genes can be tested in both in vitro and in vivo models using assays well known in the art. The expression of marker genes can be assessed by those skilled in the art by measuring the mRNA and protein levels of the desired transgene (e.g., reverse transcription PCR, Western blot analysis, and enzyme-linked immunosorbent assay (ELISA)). In some embodiments, the expression of marker or reporter protein can be used to assess the expression of the desired transgene, for example, by examining the expression of reporter protein by fluorescence microscopy or luminescence plate reader. For in vivo applications, protein function assays can be used to test the functionality of a given gene and / or gene product to determine whether gene editing has occurred successfully. It is contemplated herein that the effect of gene editing in cells or subjects can last for at least, about, or up to 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 10 months, 12 months, 18 months, 2 years, 5 years, 10 years, 20 years, or can be permanent. Marker / reporter genes
[0146] Marker / reporter genes can be screenable or selectable.
[0147] Exemplary marker genes include, but are not limited to, any of the fluorescent reporter genes, such as GFP, RFP, etc., and bioluminescent reporter genes. Exemplary marker genes include glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, luciferase, green fluorescent protein (e.g., GFP, GFP-2, tagGFP, turboGFP, sfGFP, EGFP, Emerald, Azami Green, Monomeric Azami Green, CopGFP, AceGFP, ZsGreenl), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (e.g., YFP, EYFP, Citrine, Venus YPet, PhiYFP, ZsYellowl), cyan fluorescent protein (e.g., ECFP, Cerulean, CyPet AmCyanl, Midoriishi-Cyan). These include, but are not limited to, autofluorescent proteins, including red fluorescent proteins (e.g., mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, HcRed-Tandem, HcRed 1, AsRed2, eqFP6l 1, mRaspberry, mStrawberry, Jred), orange fluorescent proteins (e.g., mOrange, mKO, Kusabira-Orange, monomeric Kusabira-Orange, mTangerine, tdTomato), and blue fluorescent proteins (BFPs).
[0148] Marker genes can also include, without limitation, DNA sequences encoding β-lactamase, β-galactosidase (LacZ), alkaline phosphatase, thymidine kinase, green fluorescent protein (GFP), chloramphenicol acetyltransferase (CAT), luciferase and others known in the art. When associated with the control elements that drive their expression, reporter sequences provide signals that can be detected by conventional means, including enzymatic assays, radiographic assays, colorimetric assays, fluorescent assays or other spectroscopic assays, fluorescence-activated cell sorting assays, and immunological assays, including enzyme-linked immunosorbent assays (ELISA), radioimmunoassays (RIA) and immunohistochemistry. For example, when the marker sequence is the LacZ gene, the presence of the vector that conveys the signal is detected by assaying for β-galactosidase activity. In some embodiments, when the marker gene is green fluorescent protein or luciferase, the vector that conveys the signal can be measured by colorimetric analysis based on visible light absorption or light production in a luminometer, respectively. Such reporters can be useful, for example, to verify tissue-specific targeting capabilities of nucleic acids and tissue-specific promoter controlling activity.
[0149] Marker genes include, but are not limited to, sequences encoding proteins that mediate antibiotic resistance (e.g., ampicillin resistance, neomycin resistance, G418 resistance, puromycin resistance) (e.g., blasticidin S-deaminase, amino 3'-glycosylphosphotransferase); colored or fluorescent or luminescent proteins (e.g., green fluorescent protein, enhanced green fluorescent protein, red fluorescent protein, luciferase) and proteins that mediate cellular metabolism that result in improved cell growth rate and / or gene amplification (e.g., dihydrofolate reductase). A vector comprising at least a portion of GSH
[0150] In certain aspects, vector compositions (e.g., nucleic acid vectors, viral vectors) are provided herein that include at least a portion or region of GSH identified using the methods disclosed herein. The portion or region of GSH can be modified, e.g., by point mutation to disrupt or knock out the gene function of the GSH gene identified herein. In other embodiments, the portion or region of GSH in the vector can be modified to include an inserted guide RNA (gRNA), e.g., a guide RNA for a nuclease disclosed herein. In some embodiments, the GSH vector can include a target site for the guide RNA (gRNA) disclosed herein, or alternatively, a restricted cloning site for the introduction of a nucleic acid of interest disclosed herein. In other embodiments, a recombinase recognition site, such as loxP, can be introduced to facilitate directional recombination using Cre recombinase expressed from rAAV or other gene transfer vectors. The loxP site inserted in GSH can also be used by breeding tg mice that express Cre in a tissue-specific manner.
[0151] As illustrative examples, the vector composition may be a plasmid, cosmid, or artificial chromosome (e.g., BAC), a minicircle nucleic acid, or a recombinant viral vector (e.g., rAd, AAV, rHSV, BEV, or variants thereof). In some embodiments, the vector may include a recombinase recognition site (RRS), such as a LoxP site, an attP site, an AttB site, or the like.
[0152] In certain embodiments, the nucleic acid in the vector comprises at least a portion of the GSH nucleic acid identified as genomic safe harbor (GSH) in the methods described herein.For example, in some embodiments, the nucleic acid is present in a vector, such as a plasmid, a cosmid, or an artificial chromosome, such as a BAC.In some embodiments, the nucleic acid composition comprises at least the target integration site in GSH and the 5' and 3' portions of the GSH nucleic acid adjacent to the target integration site.
[0153] In some embodiments, the vector composition comprises a GSH nucleic acid sequence that is between 30 and 1000 nucleotides in length, between 1 and 3 kb, between 3 and 5 kb, between 5 and 10 kb, or between 10 and 50 kb, between 50 and 100 kb, or between 100 and 300 kb, or between 100 and 350 kb, or any integer between 10 base pairs and 350 kb.
[0154] In some embodiments, the vector composition comprises a nucleic acid sequence comprising a first nucleic acid sequence comprising a 5' region of GSH and / or a second nucleic acid sequence comprising a 3' region of GSH. In some embodiments, the 5' region is adjacent to and upstream of the target integration site and the 3' region of GSH is adjacent to and downstream of the target integration site.
[0155] Any vector system can be used, including but not limited to plasmid vector, retrovirus vector, lentivirus vector, adenovirus vector, poxvirus vector; herpes virus (HSV) vector and adeno-associated virus vector, vaccinia virus vector, bacteriophage vector, etc. Also see U.S. Patent Nos. 6,534,261, 6,607,882, 6,824,978, 6,933,113, 6,979,539, 7,013,219, and 7,163,824, which are incorporated herein by reference in their entirety. Furthermore, it will be clear that any of these vectors can contain one or more of the sequences required for treatment. Thus, when one or more nucleic acids of interest are introduced into cells, if the nucleic acid of interest is a gene editing nucleic acid of interest, additional nuclease and / or donor sequence can be carried on the same vector or on a different vector. When multiple vectors are used, each vector may contain one or more of the nucleic acids of interest described herein. A nuclear vector containing at least a portion of GSH
[0156] In certain aspects, provided herein is a nucleic acid vector comprising at least a portion of a GSH nucleic acid identified by any one of the methods described herein. In some embodiments, the GSH nucleic acid comprises an untranslated sequence or an intron. In some embodiments, the GSH is at least, about or at most 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 143%, 144%, 14 %, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% identical to the sequence of the present invention. In some embodiments, the GSH is at least, about or up to 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 143%, 144%, 145%, 146%, 147%, 148%, 149%, 150%, 151%, 152%, 153%, 154%, 155%, %, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% identical to the sequence of the present invention.
[0157] In some embodiments, the nucleic acid vector of the present disclosure comprises at least one non-GSH nucleic acid (see below for further description).
[0158] In some embodiments, the nucleic acid vector of the disclosure further comprises (a) a transcriptional control element (e.g., an enhancer, a transcription termination sequence, an untranslated region (5' or 3' UTR), a proximal promoter element, a locus control region (e.g., the β-globin LCR, or the DNase hypersensitive site (HS) of the β-globin LCR), a polyadenylation signal sequence), and / or (b) a translational control element (e.g., a Kozak sequence, a woodchuck hepatitis virus post-transcriptional control element).
[0159] In some embodiments, the nucleic acid vector is selected from a plasmid, a minicircle, a cosmid, an artificial chromosome (e.g., BAC), a linear covalently closed (LCC) DNA vector (e.g., minicircle, minivector and miniknot), a linear covalently closed (LCC) vector (e.g., MIDGE, MiLV, ministring, miniplasmid), a miniintron plasmid, a pDNA expression vector, or a variant thereof.
[0160] In some embodiments, the nucleic acid vector can transform prokaryotic or eukaryotic cells, and can replicate and / or express.The vector can be a prokaryotic vector, such as a plasmid, or a shuttle vector, an insect vector, or a eukaryotic vector.The expression vector can also be for administration to plant cells, animal cells, preferably mammalian or human cells, fungal cells, bacterial cells, or protozoan cells, using standard techniques, for example, as described in Sambrook et al, supra, and in U.S. Patent Application Publication Nos. 20030232410, 20050208489, 20050026157, 20050064474, and 20060188987, and International Publication WO2007 / 014275.
[0161] Nucleic acid vectors of the disclosure include, for example, DNA plasmids, naked nucleic acid, naked phage DNA, minicircle DNA, and linear plasmids (e.g., as disclosed in US2009 / 0263900), as well as nucleic acid complexed with a delivery vehicle such as a liposome or poloxamer. Circular DNA expression vectors or minicircle vectors are disclosed in WO2002 / 083889, WO2014 / 170,238, WO2004 / 099420, WO20 102 / 026099, U.S. Patent Nos. 6,143,530, 5,622,866, 7,622,252, 8,460,924, 6,277,608, U.S. Patent Application Publication Nos. 2003 / 0032092 and 2004 / 0214329, which are incorporated by reference in their entireties.
[0162] Nucleic acid vectors suitable for use in the methods and compositions disclosed herein include linear covalently closed DNA vectors (e.g., as described in Nafissi and Slavcev "Construction and characterization of anin-vivo linear covalently closed DNA vector production system." Microbial cell factories 11.1 (2012): 154), as well as linear covalently closed (UCC) miniplasmids (e.g., as described by Slavcev, Sum, and Nafissi "Optimized production of a safe and efficient genetherapeutic vaccine versus HIV via a linear covalently closed DNAminivector." BMC Infectious Diseases 14. S2 (2014): p74), DNA ministrings (e.g., as described in U.S. Pat. No. 9,290,778; Nafiseh, et al. "DNA ministrings: highly safe and effective gene delivery vectors." Molecular Therapy - Nucleic Acids 3.6 (2014): el65; Wong, Shirley, et al. "Production of of double-stranded DNA ministrings." Journal of visualized experiments: JoVE 108 (2016)), or ceDNA vectors (e.g., UiU, et al, (2013) Production and Characterization of Novel Recombinant Adeno-Associated Virus Replicative-Form Genomes: A Eukaryotic Source of DNA for Gene Transfer. PLoS ONE 8(8): e69879).
[0163] Nucleic acid vectors include, for example, minimized vectors, plasmids (including antibiotic-free plasmids), mini-plasmids, mini-circles, mini-vectors, such as those described in Hardee, Cinnamon L., et al. "Advances in non-viral DNA vectors for gene therapy." Genes 8.2 (2017): 65. Examples of circular covalently closed vectors (CCC vectors) include mini-circles, mini-vectors, and mini-knots. Examples of linear covalently closed (LCC) vectors include MIDGE, MiLV, and mini-string. Mini-intron plasmids can also be used. These are described in Table 2 of Hardee, Cinnamon L., et al. "Advances in non-viral DNA vectors for gene therapy." Genes 8.2 (2017): 65.
[0164] Nucleic acid vectors further include plasmid DNA vectors (pDNA expression vectors), for example as described in the review articles Gill, et al., "Progress and prospects: the design and production of plasmid vectors." Gene therapy 16.2 (2009): 165-171, and Yin, Hao, et al., "Non-viral vectors for gene-based therapy." Nature Reviews Genetics 15.8 (2014): 541- 555. Nucleic acid vector for integration into the GSH locus of a target genome
[0165] In certain aspects, the present specification provides a nucleic acid vector described herein (e.g., a nucleic acid vector comprising at least a portion of GSH) that is used for integration into the GSH locus of a target genome of interest. In some embodiments, the nucleic acid vector (e.g., a nucleic acid vector comprising at least a portion of GSH) further comprises additional sequences or modifications (e.g., a certain orientation of a sequence homologous to the GSH sequence) for integration into the GSH locus of a target genome. Integration into a target genome can be driven by cellular processes such as homologous recombination or non-homologous end joining (NHEJ). Integration can also be initiated and / or facilitated by exogenously introduced nucleases.
[0166] In a preferred embodiment, the nucleic acid vector comprises at least one non-GSH nucleic acid. In some embodiments, the non-GSH nucleic acid is destined to be integrated into the GSH locus of the target genome.
[0167] In some embodiments, at least one non-GSH nucleic acid (in either forward or reverse orientation) is adjacent to a GSH 5' homology arm and / or a GSH 3' homology arm, the homology arm being at least, about or at most 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 143%, 144%, 145%, 146%, 1 %, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% or 100% identical to a nucleic acid sequence.
[0168] In some embodiments, the GSH homology arm is between 10 and 5000 base pairs in length, between 50 and 3000 base pairs, between 100 and 1500 base pairs, or any integer between 10 and 10,000 base pairs in length. In some embodiments, the GSH homology arm is between 100 and 1500 base pairs in length. In some embodiments, the GSH homology arm is at least 30 base pairs in length. In preferred embodiments, the GSH homology arm is of sufficient length to mediate homology-dependent integration into a GSH locus in the genome of the cell.
[0169] In some embodiments, at least one non-GSH nucleic acid flanked by GSH homology arms is oriented for integration in a forward orientation in GSH. In some embodiments, at least one non-GSH nucleic acid is oriented for integration in a reverse orientation in GSH.
[0170] In some embodiments, the nucleic acid comprises a restriction cloning site, which in some embodiments is adjacent to the GSH-5' homology arm and / or the 3' GSH homology arm to facilitate cloning of at least one non-GSH nucleic acid destined to be integrated into the GSH locus of the target genome.
[0171] Thus, in some embodiments, the nucleic acid vector composition comprises (a) a GSH 5' homology arm, (b) a nucleic acid sequence comprising a restriction cloning site, and (c) a GSH 3' homology arm, wherein the 5' and 3' homology arms bind to a target site located at a GSH locus identified according to the methods disclosed herein, and the 5' and 3' homology arms allow insertion (of the nucleic acid located between the homology arms) by homologous recombination into a locus located within a genome safe. In some embodiments, such a nucleic acid vector further comprises at least one non-GSH nucleic acid destined to be integrated into the GSH locus of the target genome.
[0172] The 5' and 3' homology arms can be any sequence that is homologous to the GSH target sequence in the genome of the host cell. In some embodiments, the 5' and 3' homology arms can be homologous to a portion of GSH as described herein. Furthermore, the 5' and 3' homology arms can be non-coding or coding nucleotide sequences.
[0173] In some embodiments, the 5' and / or 3' homology arms can be homologous to sequences immediately upstream and / or downstream of the integration or DNA cleavage site on the chromosome. Alternatively, the 5' and / or 3' homology arms are far away from the integration or DNA cleavage site, e.g., at least, about, or up to 1, 2, 5, 10, 15, 20, 25, 30, 50, 75, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000, 1010, 1020, 1030, 1040, 1050, 1060, 1070 75, 500, 525, 550, 575, 600, 625, 650, 675, 700, 725, 750, 775, 800, 825, 850, 875, 900, 925, 950, 975, 1000, 1025, 1050, 1075, 1100, 1125, 1150, 1175, 1200, 1225, 1250, 1275, 1300, 1325, 1350, 1375, 1400, 1425, 1450, 1475, 1500, 1525, 1550, 1575, 1600, 1625, 1650, 1675, 1700, 1725, 1750, 1775, 1800, 1825, 1850, 1875, 1900, 1925, 1950, 1975, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 300 The nucleotide sequence may be homologous to a sequence that is 0, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000, or more base pairs away, or that overlaps partially or completely with a DNA cleavage site (e.g., a DNA break induced by an exogenously introduced nuclease). In some embodiments, the 3' homology arm of the nucleotide sequence is proximal to the ITR of the viral vector.
[0174] In some embodiments, nucleic acid is integrated into target genome by homologous recombination, and then exogenously introduced nuclease induces DNA break formation.In some embodiments, nuclease is TALEN, ZFN, meganuclease, megaTAL, or CRISPR endonuclease (e.g., Cas9 endonuclease or its variant).In some embodiments, CRISPR endonuclease is in complex with guide RNA.
[0175] Thus, in some embodiments, the nucleic acid vector of the present disclosure further comprises a nucleic acid encoding a nuclease (e.g., Cas9 or a variant thereof, ZFN, TALEN) and / or a guide RNA, and the nuclease or nuclease / gRNA complex generates a DNA break in GSH, which is repaired using a donor nucleic acid, thereby incorporating at least one non-GSH nucleic acid into GSH. In other embodiments, the nucleic acid encoding the nuclease and / or guide RNA is provided in one or more unrelated nucleic acid vectors.
[0176] For integration of the nucleic acid located between the 5' and 3' homology arms, the 5' and / or 3' homology arms should be long enough to target GSH and allow (e.g., guide) integration into the genome by homologous recombination. The 5' and / or 3' homology arms may contain a sufficient number of nucleotides to increase the likelihood of integration at the correct position and improve the probability of homologous recombination. In some embodiments, the 5' and / or 3' homology arms may contain at least 10 base pairs, but not more than 5,000 base pairs, at least 50 base pairs, but not more than 5,000 base pairs, at least 100 base pairs, but not more than 5,000 base pairs, at least 200 base pairs, but not more than 5,000 base pairs, at least 250 base pairs, but not more than 5,000 base pairs, or at least 300 base pairs, but not more than 5,000 base pairs. In some embodiments, the 5' and / or 3' homology arms are about 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 305, 310, 315, 320, 325, 330, 335, 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 395, 400, 405, 410, 415, 420, 425, 430, 435, 440, 445, 450, 455, 460, 465, 470, 475, 480, 485, 490, 495, or 500 base pairs.Detailed information regarding the length of homology arms and recombination frequency is known in the art, see, for example, Zhang et al. "Efficient precise knock in with a double cut HDRdonor after CRISPR / Cas9-mediated double-stranded DNA cleavage." Genomebiology 18.1 (2017): 35, the entire contents of which are incorporated herein by reference.
[0177] The nucleic acid vectors of the present disclosure can be introduced into a target cell for integration into its genome by any method known in the art, e.g., chemical methods, electroporation, fusion with a cell containing the nucleic acid vector, transduction, etc. In some embodiments, the nucleic acid vectors of the present disclosure are integrated into the genome of the target cell upon transduction. non-GSH nucleic acids
[0178] The vectors (e.g., nucleic acid vectors, viral vectors) of the present disclosure may contain at least one non-GSH nucleic acid. Non-GSH nucleic acid may refer to any nucleic acid that does not contain the sequence of GSH identified herein, e.g., a nucleic acid having a sequence that is heterologous to GSH, e.g., a nucleic acid sequence that is not naturally present in the GSH locus, e.g., a transgene. Non-GSH nucleic acid may contain sequences necessary for replication and / or for maintaining the vector, e.g., an origin of replication, a selection marker (e.g., an antibiotic resistance gene, e.g., a marker that helps in selecting or screening for successful integration), etc. In a preferred embodiment, the non-GSH nucleic acid comprises a nucleic acid sequence that is destined to be integrated into the target genome. In a preferred embodiment, such non-GSH nucleic acid may include sequences that serve therapeutic or research purposes, e.g., downregulating a harmful endogenous gene, upregulating a defective gene, etc.
[0179] In certain embodiments, at least one non-GSH nucleic acid is not operably linked to a promoter. In some embodiments, the non-GSH nucleic acid may contain a sequence that is not intended to be expressed. In other embodiments, the non-GSH nucleic acid may contain a sequence that is intended to be expressed, and expression may be driven by an endogenous promoter near the integration site. The use of nearby promoters has been used to express therapeutic genes (see, for example, LogicBio Therapeutic's integration of a gene of interest into the albumin locus, where gene expression is driven by the albumin promoter).
[0180] In certain embodiments, at least one non-GSH nucleic acid is operably linked to a promoter. In some embodiments, at least one non-GSH nucleic acid is operably linked to a promoter, and the promoter is selected from (a) a promoter heterologous to the nucleic acid to which it is operably linked; (b) a promoter that promotes tissue-specific expression of the nucleic acid; (c) a promoter that promotes constitutive expression of the nucleic acid; (d) an inducible promoter; (e) an immediate early promoter of an animal DNA virus; (f) an immediate early promoter of an insect virus; and (g) an insect cell promoter.
[0181] As described herein, in some embodiments, the inducible promoter is modulated by an agent selected from a small molecule, a metabolite, an oligonucleotide, a riboswitch, a peptide, a peptidomimetic, a hormone, a hormone analog, and light. In some embodiments, the agent is selected from tetracycline, cumate, tamoxifen, estrogen, and an antisense oligonucleotide (ASO), rapamycin, FKCsA, blue light, abscisic acid (ABA), and a riboswitch.
[0182] In some embodiments, the promoter promotes tissue-specific expression in hematopoietic stem cells, hematopoietic CD34+ cells, and epidermal stem cells, epithelial stem cells, neural stem cells, lung progenitor cells, muscle satellite cells, intestinal K cells, neuronal cells, airway epithelial cells, or liver progenitor cells.
[0183] In some embodiments, the promoter is selected from a CMV promoter, a β-globin promoter, a CAG promoter, an AHSP promoter, an MND promoter, a Wiskott-Aldrich promoter, a PKLR promoter, a polyhedrin (polh) promoter, and an immediate early 1 gene (IE-1) promoter.
[0184] In some embodiments, the at least one non-GSH nucleic acid increases or restores expression of an endogenous gene in a target cell.
[0185] In other embodiments, the at least one non-GSH nucleic acid reduces or eliminates expression of an endogenous gene in the target cell.
[0186] In some embodiments, at least one non-GSH nucleic acid further comprises additional control elements. In some embodiments, at least one non-GSH nucleic acid comprises (a) a transcriptional control element (e.g., an enhancer, a transcription termination sequence, an untranslated region (5' or 3' UTR), a proximal promoter element, a locus control region (e.g., a β-globin LCR, or a DNase hypersensitive site (HS) of the β-globin LCR), a polyadenylation signal sequence), and / or (b) a translational control element (e.g., a Kozak sequence, a woodchuck hepatitis virus post-transcriptional control element).
[0187] In some embodiments, at least one non-GSH nucleic acid may encode a coding or non-coding RNA as described below.
[0188] Further provided herein is a method of inserting at least one non-GSH nucleic acid into a GSH locus of a cell, comprising introducing any one of the nucleic acid vectors described herein, any one of the viral vectors described herein, or any one of the pharmaceutical compositions described herein into a cell, whereby the non-GSH nucleic acid is integrated into the GSH locus by homologous recombination of the GSH 5' and 3' homology arms flanking the GSH locus and the non-GSH nucleic acid in the genome. In some embodiments, the non-GSH nucleic acid is integrated into the GSH locus in the forward orientation. In other embodiments, the non-GSH nucleic acid is integrated into the GSH locus in the reverse orientation. Non-coding and coding RNA
[0189] In certain aspects, at least one non-GSH nucleic acid is provided herein, wherein the non-GSH nucleic acid comprises a sequence that encodes a coding RNA.
[0190] In some embodiments, the sequence encoding the coding RNA is codon-optimized for expression in target cells.In some embodiments, at least one non-GSH nucleic acid encoding the coding RNA further comprises a sequence encoding a signal peptide, which allows the production of membrane-localized or secreted polypeptides.
[0191] In some embodiments, the at least one non-GSH nucleic acid is selected from the group consisting of (a) a protein or fragment thereof, preferably a human protein or fragment thereof; (b) a therapeutic protein or fragment thereof, antigen-binding protein, or peptide; (c) a suicide gene, optionally herpes simplex virus-1 thymidine kinase (HSV-TK); (d) a viral protein or fragment thereof; (e) a nuclease, optionally a transcription activator-like effector nuclease (TALEN), zinc finger nuclease (ZFN), meganuclease, megaTAL, or CRISPR endonuclease, (e.g., Cas9 endonuclease or a variant thereof); (f) a marker, e.g., luciferase or GFP; and / or (g) a drug resistance protein, e.g., a sequence encoding an antibiotic resistance gene, e.g., neomycin resistance.
[0192] In some embodiments, at least one non-GSH nucleic acid comprises a sequence encoding a viral protein or fragment thereof. In some embodiments, the viral protein or fragment thereof comprises a structural protein (e.g., VP1, VP2, VP3) or a non-structural protein (e.g., Rep protein). Such non-GSH nucleic acid may be useful for engineering cells to produce recombinant viral proteins (e.g., for vaccine production) and / or for engineering cells to produce recombinant viral particles (e.g., AAV, etc.). In some embodiments, the viral protein or fragment thereof comprises: (a) a parvovirus protein or fragment thereof, optionally VP1, VP2, VP3, NS1, or Rep; (b) a retrovirus protein or fragment thereof, optionally envelope protein, gag, pol, or VSV-G; (c) an adenovirus protein or fragment thereof, optionally E1A, E1B, E2A, E2B, E3, E4, or a structural protein (e.g., A, B, C); and / or (d) a herpes simplex virus protein or fragment thereof, optionally ICP27, ICP4, or pac.
[0193] In some embodiments, the at least one non-GSH nucleic acid encoding a viral protein encodes a viral surface protein or a fragment thereof. In some embodiments, (a) the surface protein or a fragment thereof is an immunogenic surface protein that elicits an immune response in the host, (b) the surface protein or a fragment thereof further comprises a signal peptide, (c) the gene encoding the surface protein or a fragment thereof is operably linked to an inducible promoter, and / or (d) the nucleic acid encoding the surface protein or a fragment thereof further comprises a suicide gene. In some embodiments, the surface protein is of a coronavirus (e.g., MERS, SARS), influenza virus, respiratory syncytial virus, hepatitis A, hepatitis B, hepatitis C, hepatitis D, hepatitis E, human papilloma virus, dengue virus serotype 1, dengue virus serotype 2, dengue virus serotype 3, dengue virus serotype 4, zika virus, West Nile virus, yellow fever virus, chikungunya virus, Mayaro virus, Ebola virus, Marburg virus, or Nipah virus. In some embodiments, the surface protein is the spike protein of SARS-CoV-2.
[0194] In some embodiments, at least one non-GSH nucleic acid comprises a sequence that encodes a protein or a fragment thereof. In some embodiments, the at least one non-GSH nucleic acid comprising a sequence encoding a protein or a fragment thereof is selected from the group consisting of a hemoglobin gene (HBA1, HBA2, HBB, HBG1, HBG2, HBD, HBE1, and / or HBZ), alpha-hemoglobin stabilizing protein (AHSP), coagulation factor VIII, coagulation factor IX, von Willebrand factor, dystrophin or truncated dystrophin, microdystrophin, utrophin or truncated utrophin, microutrophin, usherin (USH2A), GBA1, preproinsulin, insulin, GIP, GLP-1, CEP290, ATPB1, ATPB11, ABCB4, CPS1, ATP7B, KRT5, KRT14, PLEC1, Col7A1, ITGB4, ITGA6, LAMA3, LAMB3, LAMC2, KIND1, INS, F8, or a fragment thereof (e.g., a fragment encoding a B domain deleted polypeptide ... SQ, p-VIII)), IRGM, NOD2, ATG2B, ATG9, ATG5, ATG7, ATG16L1, BECN1, EI24 / PIG8, TECPR2, WDR45 / WIP14, CHMP2B, CHMP4B, Dynein, EPG5, HspB8, LAMP2, LC3b Selected from UVRAG, VCP / p97, ZFYVE26, PARK2 / Parkin, PARK6 / PINK1, SQSTM1 / p62, SMURF, AMPK, ULK1, RPE65, CHM, RPGR, PDE6B, CNGA3, GUCY2D, RS1, ABCA4, MYO7A, HFE, hepcidin, genes encoding soluble forms (e.g., of the TNFα receptor, IL-6 receptor, IL-12 receptor or IL-1β receptor), and cystic fibrosis transmembrane conductance regulator (CFTR).
[0195] In some embodiments, at least one non-GSH nucleic acid comprises a sequence encoding an antigen-binding protein. In some embodiments, the antigen-binding protein is an antibody or an antigen-binding fragment thereof, optionally the antibody or antigen-binding fragment thereof is selected from an antibody, Fv, F(ab')2, Fab', dsFv, scFv, sc(Fv)2, half-antibody-scFv, tandem scFv, Fab / scFv-Fc, tandem Fab', single chain diabody, tandem diabody (TandAb), Fab / scFv-Fc, scFv-Fc, heterodimeric IgG (CrossMab), DART, and diabody.
[0196] In some embodiments, the antigen binding protein specifically binds to TNFα, CD20, a cytokine (e.g., IL-1, IL-6, BLyS, APRIL, IFN-gamma, etc.), Her2, RANKL, IL-6R, GM-CSF, CCR5, or a pathogen (e.g., a bacterial toxin, a viral capsid protein, etc.).
[0197] In some embodiments, the antigen binding protein is selected from adalimumab, etanercept, infliximab, certolizumab, golimumab, anakinra, rituximab, abatacept, tocilizumab, natalizumab, canakinumab, atacicept, belimumab, ocrelizumab, ofatumumab, fontolizumab, trastuzumab, denosumab, sarilumab, lenzilumab, gimcilumab, siltuximab, leronlimab, and antigen-binding fragments thereof.
[0198] Thus, in some embodiments, at least one non-GSH nucleic acid encodes a receptor, a toxin, a hormone, an enzyme, a marker protein encoded by a marker gene (see above), or a cell surface protein, or a therapeutic protein, peptide, or antibody or fragment thereof. In some embodiments, the nucleic acid of interest for use in the vector compositions disclosed herein encodes any polypeptide whose expression in a cell is desired, including, but not limited to, an antigen binding protein (e.g., an antibody), an antigen, an enzyme, a receptor (cell surface or nuclear), a hormone, a lymphokine, a cytokine, a marker polypeptide, a growth factor, and a functional fragment of any of the foregoing. The coding sequence may be, for example, a cDNA.
[0199] The coding RNA may further comprise a sequence encoding a tag, such as an epitope tag, so that the tag is fused to the protein of interest to facilitate detection and / or purification. Exemplary tags include, for example, one or more copies of FLAG, His, myc, Tap, HA, or any detectable amino acid sequence.
[0200] It will be understood by those skilled in the art that a protein intended to be secreted will contain a signal peptide, and that a nucleic acid encoding such a protein will contain a nucleic acid sequence encoding the signal peptide.
[0201] In certain embodiments, at least one non-GSH nucleic acid for use in the vector compositions disclosed herein comprises a nucleic acid sequence encoding a marker gene (as described herein) that allows for selection of cells that have undergone targeted integration, and a linked sequence encoding additional functionality.
[0202] In some embodiments, the at least one non-GSH nucleic acid comprises a nucleic acid for use in a method for preventing or treating one or more genetic defects or dysfunctions in a mammal, such as, for example, a polypeptide deficiency or polypeptide excess in a mammal, particularly for preventing, treating, or reducing the severity or extent of the deficiency in a human exhibiting one or more disorders associated with the deficiency of such polypeptide in cells and tissues. The method comprises administration to a subject of a nucleic acid vector, a viral vector, or a cell comprising said nucleic acid vector or viral vector described herein (e.g., a nucleic acid described by the present disclosure) encoding one or more therapeutic peptides, polypeptides, siRNAs, microRNAs, antisense nucleotides, etc., preferably in a pharma- ceutically acceptable composition, in an amount and for a duration sufficient to prevent or treat the deficiency or disorder in a subject suffering from such disorder.
[0203] Thus, in some embodiments, at least one non-GSH nucleic acid for use in the vector compositions disclosed herein may encode one or more peptides, polypeptides or proteins that are useful for treating or preventing disease in a mammalian subject.
[0204] Exemplary non-GSH nucleic acids for use in the compositions and methods disclosed herein include BDNF, CNTF, CSF, EGF, FGF, G-SCF, GM-CSF, gonadotropins, IFN, IFG-1, M-CSF, NGF, PDGF, PEDF, TGF, VEGF, TGF-B2, TNF, prolactin, somatotropin, XIAP1, IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7, IL-8, IL-9, IL-10, IL-11, IL-12, IL-13, IL-14, IL-15, IL-16, IL-17, IL-18, IL-19, IL-20, IL-21, IL-22, IL-23, IL-24, IL-25, IL-26, IL-27, IL-28, IL-29, IL-30, IL-31, IL-32, IL-33, IL-34, IL-35, IL-36, IL-37, IL-38, IL-39, IL-31, IL-32, IL-34, IL-35, IL-39, IL-39, IL-39, IL-39, IL-39, IL-31, IL-32, IL-33, IL-34, IL-35, IL-35, IL-36, IL-37, IL-38, IL-39 ... , IL-6, IL-7, IL-8, IL-9, IL-10, IL-10(187A), viral IL-10, IL-11, IL-12, IL-13, IL-14, IL-15, IL-16, IL-17, IL-18, VEGF, FGF, SDF-1, connexin 40, connexin 43, SCN4a, HIFia, SERCa2a, ADCY1, and ADCY6.
[0205] In some embodiments, the nucleic acid is selected from the group consisting of a mammalian β-globin gene (e.g., HBA1, HBA2, HBB, HBG1, HBG2, HBD, HBE1, and / or HBZ), alpha-hemoglobin stabilizing protein (AHSP), B-cell lymphoma / leukemia 11A (BCL11A) gene, Krüppel-like factor 1 (KLF1) gene, CCR5 gene, CXCR4 gene, PPP1R12C (AAVS1) gene, hypoxanthine phosphoribosyltransferase (HPRT) gene, albumin gene, factor VIII gene, factor IX gene, leucine-rich repeat kinase 2 (LRRK2) gene, huntingtin (HTT) gene, rhodopsin (RHO) gene, cystic fibrosis transmembrane conductance regulator (CFTR) gene, F8 or a fragment thereof (e.g., a fragment encoding a B-domain deleted polypeptide ... The coding sequence or fragment thereof may be selected from the group consisting of a phosphodiesterase (Pd) gene, ...
[0206] In some embodiments, non-GSH nucleic acids can be used to restore expression of genes that are reduced in expression, silenced, or otherwise dysfunctional in a subject (e.g., silenced tumor suppressors in a subject with cancer).Similarly, in some embodiments, non-GSH nucleic acids can be used to knock down expression of genes that are aberrantly expressed in a subject (e.g., cancer genes that are expressed in a subject with cancer).
[0207] In some embodiments, the dysfunctional gene is a tumor suppressor silenced in a subject with cancer.In some embodiments, the dysfunctional gene is a cancer gene abnormally expressed in a subject with cancer.Exemplary genes (cancer genes and tumor suppressors) associated with cancer include, but are not limited to, AARS, ABCB1, ABCC4, ABI2, ABL1, ABL2, ACK1, ACP2, ACY1, ADSL, AK1, AKR1C2, AKT1, ALB, ANPEP, ANXAS, ANXA7, AP2M1, APC, ARHGAPS, ARHGEFS, ARID4A, ASNS, ATF4, ATM, ATPSB, ATPSO, AXL, BARD1, BAX, BCL2, BHLHB2, BLMH, BRAF, BRCA1, BRCA2, BTK, CANX, CAP1, CAPN1 , CAPNS1, CAV1, CBFB, CBLB, CCL2, CCND1, CCND2, CCND3, CCNE1, CCTS, CCYR61, CD24, CD44, CD59, CDC20, CDC25, CDC25A, CDC25B , CDC2LS, CDK10, CDK4, CDK5, CDK9, CDKL1, CDKN1A, CDKN1B, CDKN1C, CDKN2A, CDKN2B, CDKN2D, CEBPG, CENPC1, CGRRF1, CHAF1A, CIB1, CKMT1, CLK1, CLK2, CLK3, CLNS1A, CLTC, COL1A1, COL6A3, COX6C, COX7A2, CRAT, CRHR1, CSF1R, CSK, CSNK1G2, CTNNA1, CTN NB1, CTPS, CTSC, CTSD, CUL1, CYR61, DCC, DCN, DDX10, DEK, DHCR7, DHRS2, DHX8, DLG3, DVL1, DVL3, E2F1, E2F3, E2F5, EGFR, EGR1 , EIF5, EPHA2, ERBB2, ERBB3, ERBB4, ERCC3, ETV1, ETV3, ETV6, F2R, FASTK, FBN1, FBN2, FES, FGFR1, FGR, FKBP8, FN1, FOS, FOSL1 , FOSL2, FOXG1A, FOXO1A, FRAP1, FRZB, FTL, FZD2, FZDS, FZD9, G22P1, GAS6, GCNSL2, GDF1S, GNA13, GNAS, GNB2, GNB2L1, GPR39,<h2 style=";text-align:left;direction:ltr">GRB2、GSK3A、GSPT1、GTF21、HDAC1、HDGF、HMMR、HPRT1、HRB、HSPA4、HSPAS、HSPA8、 HSPB1、HSPH1、HYAL1、HYOU1、ICAM1、ID1、ID2、IDUA、IER3、IFITM1、IGF1R、IGF2R、I GFBP3、IGFBP4、IFFBPS、IL1B、ILK、ING1、IRF3、ITGA3、ITGA6、ITGB4、JAK1、JARID 1A、JUN、JUNB、JUND、K-ALPHA-1、KIT、KITLG、KLK10、KPNA2、KRAS2、KRT18、KRT2A、K RT9、LAMB1、LAMP2、LCK、LCN2、LEP、LITAF、LRPAP1、LTF、LYN、LZTR1、MADH1、MAP2K 2、MAP3K8、MAPK12、MAPK13、MAPKAPK3、MAPRE1、MARS、MAS1、MCC、MCM2、MCM4、MDM2、 MDM4、MET、MGST1、MICB、MLLT3、MME、MMP1、MMP14、MMP17、MMP2、MNDA、MSH2、MSH6、 MT3、MYB、MYBL1、MYBL2、MYC、MYCLI、MYCN、MYD88、MYL9、MYLK、NEO1、NF1、NF2、NFKB I, NFKB2, NFSF7, NID, NINJ1, NMBR, NME1, NME2, NME3, NOTCH1, NOTCH2, NOTCH4, NPM1, NQO1, NR1D1, NR2F1, NR2F6, NRAS, NRG1, NSEP1, OSM, PA 2G4、PABPC1、PCNA、PCTK1、PCTK2、PCTK3、PDGFA、PDGFB、PDGFRA、PDPK 1、PEA15、PFDN4、PFDN5、PGAM1、PHB、PIK3CA、PIK3CB、PIK3CG、PIM1、PK M2、PKMYT1、PLK2、PPARD、PPARG、PPIH、PPP1CA、PPP2RSA、PRDX2、PRDX 4、PRKAR1A、PRKCBP1、PRNP、PRSS15、PSMA1、PTCH、PTEN、PTGS1、PTMA、P TN、PTPRN、RABSA、RAC1、RADSO、RAF1、RALBP1、RAP1A、RARA、RARB、RAS GRF1、RB1、RBBP4、RBL2、REA、REL、RELA、RELB、RET、RFC2、RGS19、RHOA、<h2 style=";text-align:left;direction:ltr">RHOB、RHOC、RHOD、RIPK1、RPN2、RPS6KB1、RRM1、SARS、SELENBP1、SEMA3C、 SEMA4D、SEPP1、SERPINH1、SFN、SFPQ、SFRS7、SHB、SHH、SIAH2、SIVA、SIVA TP53、SKI、SKIL、SLC16A1、SLC1A4、SLC20A1、SMO、SMPD1、SNAI2、SND1、SN RPB2、SOCS1、SOCS3、SOD1、SORT1、SPINT2、SPRY2、SRC、SRPX、STAT1、STAT 2、STAT3、STAT5B、STC1、TAF1、TBL3、TBRG4、TCF1、TCF7L2、TFAP2C、TFDP1 、TFDP2、TGFA、TGFB1、TGFBR1、TGFBR2、TGFBR3、THBS1、TIE、TIMP1、TIMP3、 TJP1、TK1、TLE1、TNF、TNFRSF10A、TNFRSF10B、TNFRSF1A、TNFRSF1B、TNFR SF6、TNFSF7、TNK1、TOB1、TP53、TP53BP2、TP5313、TP73、TPBG、TPT1、TRAD D、TRAM1、TRRAP、TSG101、TUFM、TXNRD1、TYR03、UBC、UBE2L6、UCHL1、USP7 、VDAC1、VEGF、VHL、VIL2、WEE1、WNT1、WNT2、WNT2B、WNT3、WNTSA、WT1、XRCC 1、YES 1、YWHAB、YWHAZ、ZAP70、およびZNF9。、<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0208] <h2 style=";text-align:left;direction:ltr"> In some embodiments, the dysfunctional gene is HBB. In some embodiments, HBB comprises at least one nonsense, frameshift or splicing mutation that reduces or eliminates β-globin production. In some embodiments, HBB comprises at least one mutation in the promoter region or polyadenylation signal of HBB. In some embodiments, the HBB mutation is c.17A>T, c.-1360G, c.92+1G>A, c.92+6T>C, c.93-21G>A, c.1180T, c.316-106OG, c.25_26delAA, c.27_28insG, c.92+5G>C, c.1180T, c.135delC, c.315+1G>A, c.-78A>G, c.52A>T, c. At least one of the following: c.59A>G, c.92+5G>C, c.l24_127delTTCT, c.316-1970T, c.-78A>G, c.52A>T, c.l24_127delTTCT, c.316-197C>T, C.-1380T, c.-79A>G, c.92+5G>C, c.75T>A, c.316-2A>G, and c.316-2A>C.
[0209] In certain embodiments, sickle cell disease is improved by gene therapy (e.g., stem cell gene therapy) that introduces an HBB variant that includes one or more mutations that include anti-sickling activity. In some embodiments, the HBB variant can be a double mutant (βAS2; T87Q and E22A). In other embodiments, the HBB variant can be a triple mutant β-globin variant (βAS3; T87Q, E22A and G16D). The glycine to aspartic acid modification in β16 provides a competitive advantage over sickle globin (βS, HbS) for binding to the α chain. The glutamic acid to alanine modification in β22 partially enhances the axial interaction with α20 histidine. These modifications result in anti-sickling properties that are superior to those of the single T87Q modification variant and comparable to fetal globin. In a mouse model of SCD, transplantation of bone marrow stem cells transduced with SIN lentivirus carrying βAS3 reversed erythroid physiology and SCD clinical symptoms. Based on this, this variant is being tested in a clinical trial (identification number: NCT02247843), Cytotherapy (2018) 20(7): 899-910.
[0210] In some embodiments, the dysfunctional gene is CFTR. In some embodiments, CFTR is selected from the group consisting of ΔF508, R553X, R74W, R668C, S977F, L997F, K1060T, A1067T, R1070Q, R1066H, T3381, R334W, G85E, A46D, I336K, H1054D, M1V, E92K, V520F, H1085R, R560T, L92 and a mutation selected from 7P, R560S, N1303K, M1101K, L1077P, R1066M, R1066C, L1065P, Y569D, A561E, A559T, S492F, L467P, R347P, S341P, I507del, G1061R, G542X, W1282X, and 2184InsA.
[0211] It will be understood by those skilled in the art that the nucleic acid of interest can code for a protein or polypeptide, and that mutations resulting in conservative amino acid substitutions can be made in the transgene to produce functionally equivalent variants or homologs of the protein or polypeptide. In some aspects, the present disclosure encompasses sequence changes resulting in conservative amino acid substitutions in the transgene. In some embodiments, the non-GSH nucleic acid codes for a gene with a dominant-negative mutation. For example, the nucleic acid of interest defined herein codes for a mutant protein that interacts with the same elements as the wild-type protein, thereby blocking some aspects of the function of the wild-type protein.
[0212] In some embodiments, at least one non-GSH nucleic acid may further comprise a suicide gene operably linked to an inducible promoter and / or a tissue-specific promoter. Thus, such vectors can be used to kill cells by signals or induce cells to undergo apoptosis or programmed cell death by specific and distinct signals. Such vectors containing suicide genes can be used as an emergency exit when gene targeting or gene editing systems do not function as expected. Alternatively, suicide genes can be used to kill cancer cells or to increase the sensitivity of cancer cells, for example, to chemotherapy. Exemplary suicide genes are well known in the art and include thymidine kinase (TK, viral), cytosine deaminase (CD, bacterial and yeast), carboxypeptidase G2 (CPG2, bacterial) and nitroreductase (NTR, bacterial). In some embodiments, the suicide gene is herpes simplex virus-1 thymidine kinase (HSV-TK).
[0213] The method of targeted insertion of any sequence of interest into cells is further described herein.In some embodiments, the nucleic acid of interest is a nucleic acid that encodes a gene or a group of genes whose expression is known to be associated with a specific differentiation lineage of stem cells.A sequence that includes a gene involved in cell fate or other markers of stem cell differentiation can also be inserted.For example, a promoterless construct that contains such a gene can be inserted into a specific region (locus), so that the endogenous promoter at that locus drives the expression of the gene product.
[0214] Similarly, in certain embodiments, genomic modifications (e.g., transgene integration) at the GSH locus identified herein allow for the integration of a nucleic acid of interest that can either utilize a promoter found in that safe harbor locus, or allow for control of expression of the transgene by an exogenous promoter or regulatory element as described herein that is fused to the nucleic acid of interest prior to insertion.
[0215] In certain embodiments, at least one non-GSH nucleic acid comprises a sequence encoding a non-coding RNA. In some embodiments, the non-coding RNA comprises an antisense polynucleotide, lncRNA, piRNA, miRNA, shRNA, siRNA, antisense RNA, snoRNA, snRNA, scaRNA, and / or guide RNA. In some embodiments, the non-coding RNA targets a gene selected from the following genes: DMT-1, ferroportin, TNFα receptor, IL-6 receptor, IL-12 receptor, IL-1β receptor, and genes encoding mutant proteins (e.g., mutant HFE, CFTR).
[0216] Small nucleic acids that can modulate the expression of gene products associated with cancer (e.g., oncogenes) can be used to prevent or treat cancer. In some embodiments, the non-GSH nucleic acid encodes a gene product associated with cancer (or a functional RNA that inhibits the expression of a gene associated with cancer) for use, e.g., for treatment, e.g., for research purposes to study cancer or to identify therapeutic agents that prevent or treat cancer.
[0217] It is also understood by those skilled in the art that non-GSH nucleic acids can contain one or more mutations resulting in conservative amino acid substitutions that can result in functionally equivalent variants or homologs of proteins or polypeptides.Furthermore, the present disclosure contemplates nucleic acids of interest integrated into the GSH locus described herein that have dominant negative mutations.For example, the nucleic acid of interest can code for a mutant protein that interacts with the same elements as the wild-type protein, thereby blocking some aspects of the function of the wild-type protein.
[0218] In some embodiments, the at least one non-GSH nucleic acid comprises a non-coding RNA that mediates RNA interference. For example, the non-coding RNA comprises a small interfering RNA. Small interfering RNA (siRNA) is an agent that functions to inhibit expression of a target nucleic acid, for example, by RNAi. siRNA may be chemically synthesized, generated by in vitro transcription, or produced within a host cell. In some embodiments, the siRNA is a double-stranded RNA (dsRNA) molecule about 15 to about 40 nucleotides in length, preferably about 15 to about 28 nucleotides in length, more preferably about 19 to about 25 nucleotides in length, more preferably about 19, 20, 21, or 22 nucleotides in length, and may contain 3' and / or 5' overhangs on each strand having a length of about 0, 1, 2, 3, 4, or 5 nucleotides. The length of the overhangs is independent between the two strands, i.e., the length of the overhang on one strand is independent of the length of the overhang on the second strand. Preferably, the siRNA is capable of promoting RNA interference by degradation or specific post-transcriptional gene silencing (PTGS) of the target messenger RNA (mRNA).
[0219] In other embodiments, the siRNA is a small hairpin (also called stem-loop) RNA (shRNA). In some embodiments, these shRNAs are composed of a short (e.g., 19-25 nucleotide) antisense strand, followed by a 5-9 nucleotide loop, and a similar sense strand. Alternatively, the sense strand may precede the nucleotide loop structure, and the antisense strand may follow. These shRNAs may be contained in plasmids, retroviruses, and lentiviruses and expressed, for example, from the pol III U6 promoter or another promoter (see, e.g., Stewart, et al. (2003) RNA Apr;9(4):493-501, incorporated herein by reference).
[0220] In some embodiments, the non-coding RNA comprises a piRNA. piwi-binding RNAs (piRNAs) are the largest class of small non-coding RNA molecules. piRNAs form RNA-protein complexes by interaction with piwi proteins. These piRNA complexes have been implicated in both epigenetic and post-transcriptional gene silencing of retrotransposons and other genetic elements in germline cells, particularly during spermatogenesis. They are distinctly different from microRNAs (miRNAs) in size (26-31 nt instead of 21-24 nt), lack sequence conservation, and increased complexity. However, like other small RNAs, piRNAs are thought to be involved in gene silencing, particularly silencing of transposons. The majority of piRNAs are antisense to transposon sequences, suggesting that transposons are piRNA targets. In mammals, the activity of piRNAs in transposon silencing is paramount during embryonic development, and in both C. elegans and humans, piRNAs appear to be essential for spermatogenesis. piRNAs are involved in RNA silencing through the formation of RNA-induced silencing complexes (RISCs).
[0221] In some embodiments, the non-coding RNA comprises miRNA. miRNA and other small interfering nucleic acids control gene expression by target RNA transcript cleavage / degradation or translational repression of target messenger RNA (mRNA). miRNAs are naturally expressed, generally as final 19-25 untranslated RNA products. miRNAs exert their activity through sequence-specific interactions with the 3' untranslated region (UTR) of target mRNAs. These endogenously expressed miRNAs form hairpin precursors, which are then processed into miRNA duplexes and further into "mature" single-stranded miRNA molecules. The mature miRNA guides the miRISC, a multiprotein complex that identifies target sites, for example in the 3'UTR region of target mRNAs, based on complementarity with the mature miRNA. Figures 13A and 13B disclose a non-limiting list of miRNA genes and their homologs, or as targets of small interfering nucleic acids (e.g., miRNA sponges, antisense oligonucleotides, TuD RNAs) encoded by the nucleic acids described herein.
[0222] miRNA inhibits the function of the mRNA it targets, and as a result, inhibits the expression of the polypeptide encoded by the mRNA. Thus, blocking (partially or completely) the activity of miRNA (e.g., silencing miRNA) can effectively induce or restore (derepress) the expression of the polypeptide whose expression is inhibited. In some embodiments, derepression of the polypeptide encoded by the mRNA target of miRNA is achieved by inhibiting miRNA activity in cells by any one of a variety of methods. For example, blocking the activity of miRNA can be achieved by blocking the interaction of miRNA with its target mRNA by hybridization with a small interfering nucleic acid (e.g., antisense oligonucleotide, miRNA sponge, TuD RNA) that is complementary or substantially complementary to miRNA. As used herein, a small interfering nucleic acid that is substantially complementary to miRNA is one that can hybridize with miRNA and block the activity of miRNA. In some embodiments, a short interfering nucleic acid that is substantially complementary to an miRNA is a short interfering nucleic acid that is complementary to an miRNA at all but 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 bases. In some embodiments, a short interfering nucleic acid that is substantially complementary to an miRNA is a short interfering nucleic acid that is complementary to an miRNA at at least one base. Gene editing
[0223] In some embodiments, the methods and compositions described herein are used to integrate a nucleic acid into the GSH of the present disclosure in a target genome. In some embodiments, the integration is initiated and / or facilitated by an exogenously introduced nuclease, and the DNA break induced by the nuclease is repaired using the homology arms as a guide for homologous recombination, thereby inserting the nucleic acid flanked by the homology arms into the target genome.
[0224] In some embodiments, gene editing system is introduced into GSH to knock down the expression of endogenous genes by introducing certain modifications into genes or control elements.In some embodiments, gene editing system can be introduced into GSH to knock out or delete all or part of endogenous genes to remove harmful copies of genes.In some embodiments, such negative modulation of gene expression is controlled, for example, gene editing system can be under inducible promoter or tissue-specific promoter, which allows selective gene downregulation (for example, gene can be deleted at a certain stage of differentiation), for example with temporal regulation, and / or tissue-specific knockdown or knockout of genes.
[0225] For example, double-strand breaks (DSBs) can be created by site-specific nucleases, such as zinc finger nucleases (ZFNs) or TAL effector domain nucleases (TALENs).See, for example, Urnov et al. (2010) Nature 435(7042):646-51; U.S. Patent Nos. 8,586,526, 6,534,261, 6,599,692, 6,503,717, 6,689,558, 7,067,317, and 7,262,054, the disclosures of which are incorporated herein by reference.
[0226] Another nuclease system involves the use of the so-called adaptive immune system found in bacteria and archaea, known as the CRISPR / Cas system. CRISPR / Cas systems are found in 40% of bacteria and 90% of archaea, and differ in the complexity of their systems. See, for example, U.S. Pat. No. 8,697,359. CRISPR loci (clustered regularly interspaced short palindromic repeats) are regions in the genome of an organism where short segments of foreign DNA are integrated between short repeated palindromic sequences. These loci are transcribed and the RNA transcripts ("pre-crRNA") are processed into short CRISPR RNAs (crRNAs). There are three types of CRISPR / Cas systems, all of which contain these RNAs and proteins known as "Cas" proteins (CRISPR-associated). Both types I and III have a Cas endonuclease that processes the pre-crRNA, which, once fully processed into the crRNA, assembles a multi-Cas protein complex that can cleave nucleic acids that are complementary to the crRNA.
[0227] In type II systems, crRNA is produced using a different mechanism in which a transactivating RNA (tracrRNA) complementary to a repeat sequence in the pre-crRNA triggers processing by double-strand-specific RNase III in the presence of Cas9 protein or its variants. Cas9 can then cleave the target DNA that is complementary to the mature crRNA, but cleavage by Cas9 depends on both base pairing between the crRNA and the target DNA and the presence of a short motif in the crRNA called the PAM sequence (protospacer adjacent motif) (see Qi et al (2013) Cell 152: 1173). In addition, tracrRNA must also be present as it base pairs with the crRNA at its 3' end and this association triggers Cas9 activity.
[0228] Cas9 protein has at least two nuclease domains, one of which is similar to HNH endonuclease, while the other resembles Ruv endonuclease domain. HNH-type domains are responsible for cleaving DNA strands that are complementary to crRNA, while Ruv domains appear to cleave non-complementary strands. Variants of Cas9, such as Cas9 nickase mutants that reduce off-target activity (see, e.g., Ran et al. (2014) Cell 154(6): 1380-1389), nCas, Cas9-D10A, are recognized in the art.
[0229] The need for a crRNA-tracrRNA complex can be circumvented by the use of engineered "single guide RNAs" (sgRNAs) that contain a hairpin that is normally formed by annealing of the crRNA and tracrRNA (see Jinek et al (2012) Science 337:816 and Cong et al (2013) Sciencexpress / 10.1126 / science.1231143). Thus, an exogenously introduced CRISPR endonuclease (e.g., Cas9 or a variant thereof) and a guide RNA (e.g., sgRNA or gRNA) can induce DNA cleavage at a specific locus in the genome of a target cell. Non-limiting examples of single guide RNA or guide RNA (sgRNA or gRNA) sequences suitable for targeting are shown in Table 1 of US Patent Application Publication No. 2015 / 0056705, which is incorporated herein by reference in its entirety. Additionally, the sgRNA or gRNA can include the sequence of the GSH locus described herein.
[0230] In some embodiments, the gene editing nucleic acid sequence encodes a molecule selected from the group consisting of a sequence-specific nuclease, one or more guide RNAs (gRNAs), CRISPR / Cas, ribonucleoproteins (RNPs), or any combination thereof. In some embodiments, the sequence-specific nuclease comprises a TAL nuclease, a zinc finger nuclease (ZFN), a meganuclease, a megaTAL, or an RNA-guided endonuclease of a CRISPR / Cas system (e.g., a Cas protein, e.g., CAS1-9, Csy, Cse, Cpfl, Cmr, Csx, Csf, cpfl, nCAS, or others). These gene editing systems are known to those skilled in the art, see, for example, the TALENS described in International Patent Application No. PCT / US2013 / 038536 and US Patent Publication No. 2017-0191078-A9, which are incorporated by reference in their entirety. CRISPR cas9 system is known in the art and described in US Patent Application No. 13 / 842,859, filed March 2013, and US Patent Nos. 8,697,359, 8771,945, 8795,965, 8,865,406, and 8,871,445.GSH is also useful for inactivated nuclease systems, such as CRISPRi or CRISPRa dCas, nCas, or Cas13 systems. Guide RNA (gRNA)
[0231] In general, guide sequence is any polynucleotide sequence that has sufficient complementarity with target polynucleotide sequence to hybridize with target sequence and direct sequence-specific targeting of RNA-guided endonuclease complex to selected genomic target sequence.In some embodiments, guide RNA binds to target sequence and, for example, ribonucleoprotein (RNP), for example, CRISPR-associated protein that can form CRISPR / Cas complex.
[0232] In some embodiments, the guide RNA (gRNA) sequence comprises a targeting sequence that directs the gRNA sequence to a desired site in the genome and is fused to a crRNA and / or tracrRNA sequence that allows the guide sequence to associate with an RNA-guided endonuclease. In some embodiments, the degree of complementarity between the guide sequence and its corresponding target sequence is at least, about or up to 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or higher when optimally aligned using a suitable alignment algorithm. Optimal alignment can be determined by use of any suitable algorithm for aligning sequences, such as the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transformation (e.g., the Burrows-Wheeler aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, ELAND (Illumina, San Diego, Calif.), SOAP, and Maq.
[0233] Guide sequence can be selected to target any target sequence.In some embodiments, target sequence is a sequence in the genome of cell or in the GSH disclosed herein.In some embodiments, guide RNA can be complementary to either strand of the DNA sequence to be targeted.For targeted cleavage by RNA-guided endonuclease, it is understood by those skilled in the art that a unique target sequence in genome is more preferable than a target sequence that appears more than once in genome. Bioinformatics software can be used to predict and minimize off-target effects of guide RNAs (see, e.g., Naito et al. "CRISPRdirect: software for designing CRISPR / Casguide RNA with reduced off-target sites" Bioinformatics (2014), epub; Heigwer et al. "E-CRISP: fast CRISPR target site identification" Nat. Methods11:122-123 (2014); Bae et al. "Cas-OFFinder: a fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases" Bioinformatics 30(10): 1473-1475 (2014); Aach et al. "CasFinder: Flexible algorithm for identifying specific Cas9 targets ingenomes" BioRxiv (2014)).
[0234] In general, a "crRNA / tracrRNA fusion sequence," as the term is used herein, refers to a nucleic acid sequence that is fused to a unique targeting sequence and functions to allow the formation of a complex comprising a guide RNA and an RNA-guided endonuclease. Such a sequence may be modeled after a CRISPR RNA (crRNA) sequence in a prokaryote, including (i) a variable sequence called a "protospacer" that corresponds to the target sequence described herein, and (ii) a CRISPR repeat. Similarly, the tracrRNA ("trans-activating CRISPR RNA") portion of the fusion may be designed to include a secondary structure similar to the tracrRNA sequence in a prokaryote (e.g., a hairpin) to allow the formation of an endonuclease complex. In some embodiments, the single transcript further includes a transcription termination sequence, e.g., a poly-T sequence, e.g., six T nucleotides. In some embodiments, the guide RNA may include two RNA molecules, referred to herein as "dual guide RNA" or "dgRNA." In some embodiments, the dgRNA may include a first RNA molecule comprising a crRNA and a second RNA molecule comprising a tracrRNA. The first and second RNA molecules can form an RNA duplex by base pairing with the flagpole on the crRNA and the tracrRNA. When using dgRNA, the flagpole does not need to have an upper limit on length.
[0235] In other embodiments, the guideRNA may comprise a single RNA molecule, referred to herein as "single guide RNA" or "sgRNA". In some embodiments, the sgRNA may comprise a crRNA covalently linked to a tracrRNA. In some embodiments, the crRNA and the tracrRNA may be covalently linked via a linker. In some embodiments, the sgRNA may comprise a stem-loop structure by base pairing between a flagpole on the crRNA and the tracrRNA. In some embodiments, the single guide RNA is at least, about, or up to 50, 60, 70, 80, 90, 100, 110, 120 or more nucleotides in length (e.g., 75-120, 75-110, 75-100, 75-90, 75-80, 80-120, 80-110, 80-100, 80-90, 85-120, 85-110, 85-100, 85-90, 90-120, 90-110, 90-100, 100-120, 100-120 nucleotides in length). In some embodiments, a nucleic acid vector described herein for integration of a nucleic acid of interest into the GSH locus, or a composition thereof, comprises a nucleic acid encoding at least one gRNA. For example, the second polynucleotide sequence may encode between 1 gRNA and 50 gRNAs, or at least, about or up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 gRNAs. Each of the polynucleotide sequences encoding different gRNAs can be operably linked to a promoter. In some embodiments, the promoters operably linked to different gRNAs can be the same promoter. The promoters operably linked to different gRNAs can be different promoters. The promoter can be a constitutive promoter, an inducible promoter, a repressible promoter, or a regulatable promoter.
[0236] In some embodiments, the non-GSH nucleic acid constitutes or is introduced into the target cell together with another vector that includes a nucleic acid encoding Cas nickase (nCas; e.g., Cas9 nickase or Cas9-D10A). It is contemplated herein that such nCas enzymes can be used with guide RNAs that include homology to GSH as described herein, for example, to release physically constrained sequences or to effect untwisting. Releasing physically constrained sequences can, for example, "unwind" the vector to expose homology-directed repair (HDR) template homology arms for interaction with genomic sequences.
[0237] In some embodiments, zinc finger nucleases are used to induce DNA cleavage that facilitates the integration of desired nucleic acid. As used interchangeably herein, "zinc finger nuclease" or "ZFN" refers to a chimeric protein molecule that comprises at least one zinc finger DNA binding domain operatively linked to at least one nuclease or part of a nuclease that can cleave DNA when fully assembled. "Zinc finger" as used herein refers to a protein structure that recognizes and binds to a DNA sequence. Zinc finger domains are the most common DNA binding motifs in the human proteome. A single zinc finger contains approximately 30 amino acids, and the domain generally functions by binding to three consecutive base pairs of DNA through the interaction of a single amino acid side chain per base pair.
[0238] In some embodiments, the nucleic acid for integration described herein is integrated into the target genome by a nuclease-free homology-dependent repair system, for example, as described in Porro et al., Promoterless gene targeting without nucleases rescues lethality of aCrigler-Najjar syndrome mouse model, EMBO Molecular Medicine, (2017). In some embodiments, the in vivo gene targeting approach is suitable for the insertion of donor sequences without the use of nucleases. In some embodiments, the donor sequence can be promoter-free.
[0239] In some embodiments, the nuclease located between restriction sites can be an RNA-guided endonuclease.As used herein, the term "RNA-guided endonuclease" refers to an endonuclease that forms a complex with an RNA molecule that comprises a region that is complementary to a selected target DNA sequence, and thus the RNA molecule binds to the selected sequence and directs endonuclease activity to the selected target DNA sequence in the GSH identified herein. CRISPR / Cas systems
[0240] As recognized in the art and described above, the CRISPR-CAS9 system comprises a combination of proteins and ribonucleic acid ("RNA") that can modify the genetic sequence of an organism (see, for example, US Patent Application Publication No. 2014 / 0170753). CRISPR-Cas9 provides a set of tools for Cas9-mediated genome editing by non-homologous end joining (NHEJ) or homologous recombination in mammalian cells. A person skilled in the art can choose from several known CRISPR systems, such as type I, type II, and type III. In some embodiments, the nucleic acids described herein for integration of a nucleic acid of interest into the GSH locus can be designed to include sequences encoding one or more components of these systems, such as guide RNA, tracrRNA, or Cas (e.g., Cas9 or a variant thereof). In certain embodiments, a single promoter drives expression of the guide sequence and tracrRNA, and another promoter drives Cas (e.g., Cas9 or a variant thereof) expression. It will be understood by those of skill in the art that certain Cas nucleases require the presence of a protospacer adjacent motif (PAM) adjacent to the target nucleic acid sequence.
[0241] RNA-guided nucleases, including Cas (e.g., Cas9 or variants thereof), are suitable for initiating and / or facilitating the integration of the nucleic acids described herein. The guide RNA can be directed to the same or complementary strand of DNA.
[0242] In some embodiments, the methods and compositions described herein can comprise and / or be used to deliver CRISPRi (CRISPR interference) and / or CRISPRa (CRISPR activation) systems to host cells. CRISPRi and CRISPRa systems comprise an inactivated RNA-guided endonuclease (e.g., Cas9 or its variants) that cannot generate double-strand breaks (DSBs). This allows the endonuclease, in combination with the guide RNA, to specifically bind to target sequences in genomes and provide RNA-dependent reverse transcription control.
[0243] Thus, in some embodiments, the nucleic acid compositions and methods described herein for the integration of a nucleic acid of interest into the GSH locus may comprise an inactivated endonuclease, such as an RNA-guided endonuclease and / or Cas9 or its variant, which lacks endonuclease activity but retains the ability to bind DNA in a site-specific manner, for example, in combination with one or more guide RNAs and / or sgRNAs.In some embodiments, the vector may further comprise one or more tracrRNAs, guide RNAs or sgRNAs.In some embodiments, the inactivated endonuclease may further comprise a transcription activation domain.
[0244] In some embodiments, the nucleic acid compositions and methods described herein for the integration of a nucleic acid of interest into the GSH locus may include hybrid recombinases.For example, hybrid recombinases based on the activated catalytic domain from the resolvase / invertase family of serine recombinases fused to Cys2-His2 zinc finger or TAL effector DNA binding domains are a class of reagents that can improve targeting specificity in mammalian cells and achieve excellent site-specific integration rates.Suitable hybrid recombinases include those described in Gaj et al. Enhancing the Specificity of Recombinase -Mediated GenomeEngineering through Dimer Interface Redesign, Journal of the American Chemical Society, (2014).
[0245] The nucleases described herein can be modified, e.g., engineered, to design sequence-specific nucleases (see, e.g., U.S. Patent No. 8,021,867).For example, nucleases can be designed using the methods described in Certo et al. Nature Methods (2012) 9:073-975; U.S. Patent Nos. 8,304,222, 8,021,867, 8,119,381, 8,124,369, 8,129,134, 8,133,697, 8,143,015, 8,143,016, 8,148,098, or 8,163,514, the contents of each of which are incorporated herein by reference in their entirety. Alternatively, nucleases with site-specific cleavage properties can be obtained using commercially available technologies, for example, Precision BioSciences' Directed Nuclease Editor™ genome editing technology. Mega TAL
[0246] In some embodiments, the nuclease described herein can be a megaTAL. A megaTAL is an engineered fusion protein that includes a transcription activator-like (TAL) effector domain and a meganuclease domain. MegaTALs retain the ease of target specificity engineering of TALs, while reducing off-target effects and overall enzyme size and increasing activity. MegaTAL construction and use are described in more detail, for example, in Boissel et al. 2014 Nucleic Acids Research 42(4):2591-601 and Boissel2015 Methods Mol Biol 1239: 171-196. Protocols for megaTAL-mediated gene knockout and gene editing are known in the art, see, for example, Sather et al. Science Translational Medicine 2015 7(307):ra156 and Boissel et al. 2014 Nucleic Acids Research 42(4):2591-601. MegaTALs can be used as surrogate endonucleases in any of the methods and compositions described herein. Control Arrays
[0247] The nucleic acid vectors disclosed herein can also include transcriptional or translational control sequences, such as sequences encoding a promoter, an enhancer, an insulator, an internal ribosome entry site, a 2A peptide, and / or a polyadenylation signal.
[0248] In some embodiments, the control sequence comprises a suitable promoter sequence that can direct the transcription of the gene operably linked to the promoter sequence, such as the nucleic acid of interest described herein.In an embodiment, an enhancer sequence is provided upstream of the promoter to increase the effectiveness of the promoter.In some embodiments, the control sequence comprises an enhancer and a promoter, and the second nucleotide sequence comprises an intron sequence upstream of the nucleotide sequence encoding the nuclease, the intron comprises one or more nuclease cleavage sites, and the promoter is operably linked to the nucleotide sequence encoding the nuclease.
[0249] Suitable promoters, including those described herein, may be derived from viruses, and therefore may be referred to as viral promoters, or may be derived from any organism, including prokaryotic or eukaryotic organisms.In some embodiments, the promoter is derived from insect cells or mammalian cells.Suitable promoters can be used to drive expression by any RNA polymerase (e.g., pol I, pol II, pol III). Exemplary promoters include, but are not limited to, the SV40 early promoter, the mouse mammary tumor virus long terminal repeat (LTR) promoter; the adenovirus major late promoter (Ad MLP); the herpes simplex virus (HSV) promoter, the cytomegalovirus (CMV) promoter, such as the CMV immediate early promoter region (CMVIE), the Rous sarcoma virus (RSV) promoter, the human U6 small nuclear promoter (Miyagishi et al., Nature Biotechnology 20, 497-500 (2002)), the enhanced U6 promoter (e.g., Xia et al., Nucleic Acids Res. 2003 Sep. 1; 31(17)), the human H1 promoter (H1), and the like. In some embodiments, these promoters are modified to contain one or more nuclease cleavage sites.
[0250] A promoter may contain one or more specific transcriptional control sequences to further enhance expression and / or to modify its spatial and / or temporal expression. A promoter may also contain distal enhancer or repressor elements that may be located as many as several thousand base pairs from the start site of transcription. Promoters may be derived from sources including viruses, bacteria, fungi, plants, insects and animals. Promoters may constitutively or differentially control the expression of gene components with respect to the cell, tissue or organ in which expression occurs, or with respect to the developmental stage in which expression occurs, or in response to external stimuli such as physiological stress, pathogens, metal ions or inducers. Representative examples of promoters include bacteriophage T7 promoter, bacteriophage T3 promoter, SP6 promoter, lac operator-promoter, tac promoter, SV40 late promoter, SV40 early promoter, RSV-LTR promoter, CMV IE promoter, SV40 early promoter or SV40 late promoter and CMV IE promoter, and the promoters listed below. Such promoters and / or enhancers can be used, for example, for expression of any gene of interest, e.g., a gene for editing a molecule, a donor sequence, a therapeutic protein, etc. For example, the nucleic acid can include a promoter operably linked to a DNA endonuclease or a CRISPR / Cas9-based system. The promoter operably linked to a CRISPR / Cas9-based system or a site-specific nuclease coding sequence can be a promoter from simian virus 40 (SV40), a CAG promoter, a mouse mammary tumor virus (MMTV) promoter, a human immunodeficiency virus (HIV) promoter, e.g., a bovine immunodeficiency virus (BIV) long terminal repeat (LTR) promoter, a Moloney virus promoter, an avian leukosis virus (ALV) promoter, a cytomegalovirus (CMV) promoter, e.g., a CMV immediate early promoter, an Epstein-Barr virus (EBV) promoter, or a Rous sarcoma virus (RSV) promoter.The promoter can also be a promoter from a human gene, such as human ubiquitin C (hUbC), human actin, human myosin, human hemoglobin, human muscle creatine, or human metalothionein. The promoter can also be a tissue-specific promoter, such as a liver-specific promoter, natural or synthetic. In some embodiments, delivery to the liver can be achieved using endogenous ApoE-specific targeting of the composition comprising the vector to hepatocytes by the low-density lipoprotein (LDL) receptor present on the surface of hepatocytes. In some embodiments, synthetic promoters designed in silico with an assembly of control elements are used. These synthetic promoters do not exist in nature and are designed for either optimal expression in target tissues, controlled expression, or inclusion in a viral capsid.
[0251] In some embodiments, the promoter may be selected from (a) a promoter heterologous to the nucleic acid, (b) a promoter that promotes tissue-specific expression of the nucleic acid, preferably a promoter that promotes hematopoietic cell-specific or erythroid lineage-specific expression, (c) a promoter that promotes constitutive expression of the nucleic acid, and (d) a promoter that is inducibly expressed, optionally in response to a metabolite or a small molecule or chemical. Examples of inducible promoters include those controlled by tetracycline, cumate, rapamycin, FKCsA, ABA, tamoxifen, blue light, and riboswitches. Additional details are provided, for example, in Kallunki et al. (2019) Cells 8:E796, which is incorporated by reference. In some embodiments, the promoter is selected from a CMV promoter, a β-globin promoter, a CAG promoter, an AHSP promoter, an MND promoter, a Wiskott-Aldrich promoter, and a PKLR promoter. See also the section on "Pulsed and Regulatable Gene Expression."
[0252] A significant number of genes and their regulatory elements (promoters and enhancers) are known, which direct the development and lineage-specific expression of endogenous genes.Therefore, the selection of regulatory elements and / or gene products inserted into stem cells will depend on which lineage and which developmental stage is of interest.In addition, as more details are known about the finer mechanistic differences in lineage-specific expression and stem cell differentiation, they can be incorporated into experimental protocols to fully optimize the system for efficient isolation of a wide range of desired stem cells.
[0253] Any lineage-specific or cell fate control element (e.g., promoter) or cell marker gene can be used in the compositions and methods described herein. Lineage-specific and cell fate genes or markers are well known to those of skill in the art and can be readily selected to assess a particular lineage of interest. Non-limiting examples of such genes include genes such as Ang2, Flkl, VEGFR, MHC genes, aP2, GFAP, Otx2 (see, e.g., U.S. Pat. No. 5,639,618), Dlx (Porteus et al. (1991) Neuron 7:221-229), Nix (Price et al. (1991) Nature 351:748-751), Emx (Simeone et al. (1992) EMBO J. 11:2541- 2550), Wnt (Roelink and Nuse (1991) Genes Dev. 5:381-388), En (McMahon et al.), Hox (Chisaka et al. (1991) Nature 350:473-479), acetylcholine receptor beta chain (ACHRP) (Otl et al. (1994) J. Cell.Biochem. Supplement 18A: 177). Other examples of lineage-specific genes from which regulatory elements can be obtained are available at the NCBI-GEO website, which is readily accessible via the internet and is well known to those of skill in the art. array
[0254] As used herein, a coding region refers to a region of a nucleotide sequence that contains codons that are translated into amino acid residues, whereas a non-coding region refers to a region of a nucleotide sequence that is not translated into amino acids. A transcribed non-coding sequence can be upstream (5'-UTR), downstream (3'-UTR), or intron type. A non-transcribed non-coding sequence can have cis-acting control functions, such as enhancers and promoters, or serve as "spacers", non-transcribed DNA that is used to separate functional groups in DNA, such as polylinkers, or "stuffer" DNA that is used to increase the size of a vector genome.
[0255] "Complement of" or "complementary" refers to the broad concept of sequence complementarity between regions of two nucleic acid strands or between two regions of the same nucleic acid strand. It is known that an adenine residue of a first nucleic acid region can form specific hydrogen bonds (base pair) with a residue of a second nucleic acid region that is antiparallel to the first region if the residue is thymidine or uracil. Similarly, it is known that a cytosine residue of a first nucleic acid strand can base pair with a residue of a second nucleic acid strand that is antiparallel to the first strand if the residue is guanine. A first region of a nucleic acid is complementary to a second region of the same or a different nucleic acid if at least one nucleotide residue of the first region can base pair with a residue of the second region when the two regions are arranged in an antiparallel fashion. In some embodiments, the first region comprises a first portion and the second region comprises a second portion, such that when the first and second portions are arranged in an antiparallel fashion, at least, about or up to 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 109, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109, ... %, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% are capable of base pairing with nucleotide residues in the second portion. In other embodiments, all nucleotide residues in the first portion are capable of base pairing with nucleotide residues in the second portion.
[0256] A nucleic acid is operably linked when it is placed in a functional relationship with another nucleic acid sequence. For example, a promoter or enhancer is operably linked to a coding sequence if it affects the transcription of the sequence. With respect to a transcription control sequence, operably linked means that the DNA sequences that are linked are contiguous, and, where necessary to join two protein coding regions, contiguous and in reading frame.
[0257] There is a known and definite correspondence, as defined by the genetic code (shown below), between the amino acid sequence of a particular protein and the nucleotide sequence which can encode that protein. Similarly, there is a known and definite correspondence, as defined by the genetic code, between the nucleotide sequence of a particular nucleic acid and the amino acid sequence encoded by that nucleic acid. Genetic code Alanine (Ala, A) GCA, GCC, GCG, GCT Arginine (Arg, R) AGA, ACG, CGA, CGC, CGG, CGT Asparagine (Asn, N) AAC, AAT Aspartic acid (Asp, D) GAC, GAT Cysteine (Cys, C) TGC, TGT Glutamic acid (Glu, E) GAA, GAG Glutamine (Gln, Q) CAA, CAG Glycine (Gly, G) GGA, GGC, GGG, GGT Histidine (His, H) CAC, CAT Isoleucine (Ile, I) ATA, ATC, ATT Leucine (Leu, L) CTA, CTC, CTG, CTT, TTA, TTG Lysine (Lys, K) AAA, AAG Methionine (Met, M) ATG Phenylalanine (Phe, F) TTC, TTT Proline (Pro, P) CCA, CCC, CCG, CCT Serine (Ser, S) AGC, AGT, TCA, TCC, TCG, TCT Threonine (Thr, T) ACA, ACC, ACG, ACT Tryptophan (Trp, W) TGG Tyrosine (Tyr, Y) TAC, TAT Valin (Val, V) GTA, GTC, GTG, GTT Termination signals (end) TAA, TAG, TGA
[0258] An important and well-known feature of the genetic code is its degeneracy, where more than one coding nucleotide triplet can be used for most of the amino acids used to make proteins (as exemplified above). Thus, several different nucleotide sequences can code for a given amino acid sequence. The universality of the genetic code dictates that such nucleotide sequences are considered functionally equivalent, since they result in the same amino acid sequence in all organisms, although mitochondria and plastids and similar symbiotic organelles have slightly different genetic codes. Not all codons are utilized with similar translational efficiency, but rare codons can reduce protein production due to limiting tRNA pools. In addition, sometimes methylated variants of purines or pyrimidines can be found in a given nucleotide sequence. Such methylation does not affect the coding relationship between the trinucleotide codon and the corresponding amino acid.
[0259] When altering the amino acid sequence of a polypeptide, the hydropathic index of amino acids may be considered. The importance of the hydropathic amino acid index in conferring interactive biological function to a protein is generally understood in the art. It is generally accepted that the relative hydropathicity of amino acids contributes to the secondary structure of the resulting protein, which in turn determines the interaction of the protein with other molecules, such as enzymes, substrates, receptors, DNA, antibodies, antigens, etc. Each amino acid is assigned a hydropathic index based on their hydrophobicity and charge characteristics, and these are: isoleucine (+4.5); valine (+4.2); leucine (+3.8); phenylalanine (+2.8); cysteine / cystine (+2.5); methionine (+1.9); alanine (+1.8); glycine (-0.4); threonine (-0.7); serine (-0.8); tryptophan (-0.9); tyrosine (-1.3); proline (-1.6); histidine (-3.2); glutamate (-3.5); glutamine (-3.5); aspartate ( <RTI 3.5);アスパラギン(-3.5);リシン(-3.9);およびアルギニン(-4.5)である。
[0260] It is known in the art that certain amino acids can be substituted by other amino acids having a similar hydropathic index or score and still result in a protein having a similar biological activity, i.e., a biologically functional equivalent protein can still be obtained.
[0261] Thus, as outlined above, amino acid substitutions are generally based on the relative similarity of the amino acid side-chain substituents, e.g., their hydrophobicity, hydrophilicity, charge, size, etc. Exemplary substitutions that take into account the various above-mentioned characteristics are well known to those of skill in the art and include arginine and lysine, glutamate and aspartate, serine and threonine, glutamine and asparagine, and valine, leucine and isoleucine.
[0262] It is also known in the art that nucleic acids encoding polypeptides can be codon-optimized for a particular host cell without changing amino acid sequence.Codon optimization refers to a genetic engineering approach that uses synonymous codon changes to increase protein production.This is possible because most amino acids are coded by more than one codon.Replacing rare codons with those used in has been shown to increase protein expression.
[0263] In view of the above, the nucleotide sequence of the DNA or RNA encoding the nucleic acid (e.g., therapeutic nucleic acid) described herein (or any portion thereof) can be used to derive a polypeptide amino acid sequence, and the genetic code can be used to translate the DNA or RNA into an amino acid sequence. Similarly, for a polypeptide amino acid sequence, the corresponding nucleotide sequence that can encode the polypeptide can be inferred from the genetic code (which, due to its redundancy, will give rise to multiple nucleic acid sequences for any given amino acid sequence). Thus, any description and / or disclosure herein of a nucleotide sequence encoding a polypeptide should be considered to also include a description and / or disclosure of the amino acid sequence encoded by the nucleotide sequence. Similarly, any description and / or disclosure herein of a polypeptide amino acid sequence should be considered to also include a description and / or disclosure of all possible nucleotide sequences that can encode the amino acid sequence.
[0264] Finally, nucleic acid and amino acid sequence information for nucleic acid and polypeptide molecules useful in the present invention is well known in the art and readily available in public databases such as the National Center for Biotechnology Information (NCBI). Table 3: Exemplary sequences of the GSH locus [Table 3-1] [Table 3-2] [Table 3-3] *Coordinates in Table 3 are from the human genome assembly GRCh38 / hg38. *cDNA, ssDNA, and RNA nucleic acid molecules (e.g., thymidines replaced with uridines), nucleic acid molecules encoding orthologs or variants of the encoded proteins, and nucleic acid sequences of any of the SEQ ID NOs listed above and their full length at least, about or up to 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, The nucleic acid sequence or a part thereof includes the nucleic acid sequence having 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or higher identity. Such nucleic acid molecule can have the function of full-length nucleic acid as further described herein. *See Table 5 in Example 3 for an exemplary characterization of representative GSH loci. Pulsed and Regulatable Gene Expression
[0265] In certain embodiments, the vectors (e.g., nucleic acid vectors, viral vectors), cells, pharmaceutical compositions, and / or methods of the present disclosure utilize pulsed and / or regulatable gene expression. As used herein, regulatable gene expression allows transgene expression to be controlled at will, for example, using small molecules or oligonucleotides (e.g., tetracycline or antisense oligonucleotides (ASO or AON) respectively) to turn on or off transgene expression. Regulatable gene expression is often achieved using inducible or repressible promoters, but regulatable control is intended to include control of gene expression in addition to transcription.
[0266] Therefore, regulatable gene expression is intended to include the temporal control at transcriptional, post-transcriptional, translational and / or post-translational levels.Regulatable expression corresponds to the spatial regulation of gene expression.For example, the spatial regulation of transgene can be promoted by placing transgene under tissue-specific promoter and then combining with expression modulating agent (e.g., tetracycline or ASO) that mediates temporal regulation.
[0267] Pulsed gene expression refers to turning on and off the production of transgene at regular intervals.Any regulatable gene expression system can be used for pulsed gene expression.In addition, it is contemplated herein that any of the gene expression modulations described herein can be used in combination with pulsed gene expression.
[0268] Pulsed gene expression is important for the success of gene therapy. Obtaining physiological and long-term protein expression levels remains a major challenge in gene therapy applications. High-level expression of transgenes induces ER stress and endoplasmic reticulum stress response many months after treatment, which leads to proinflammatory conditions and cell death, undermining the benefits of treatment. Pulsed transgene expression strategies (PTES) can spare target cells from overexpression stress and allow long-term expression of transgenes without a gradual decrease in expression over time. In addition, pulsed and / or regulatable expression can improve, for example, the production efficiency and / or stability of the protein encoded by the transgene.
[0269] In some embodiments, the PTES described herein is an adjustable expression system in which the default state is off until a reagent turns on or de-represses expression, thereby allowing for dosage modification to meet the specific needs of the patient, thus providing greater safety and long-term benefit. The timing of the pulse can be determined from the initial serum level (t0) and half-life (t1 / 2) of the protein of interest (see Example 11). Exemplary Regulatable Expression Systems Tetracycline-regulated operator system
[0270] A bacterial control element, the Tn10 specific tetracycline resistance operon of E. coli, can be used to control gene expression. For example, there are three exemplary configurations of this system: (1) a repression-based configuration, in which the Tet operator (TetO) is inserted between a constitutive promoter and a gene of interest, and the binding of the tet repressor (TetR) to the operator represses downstream gene expression. In this system, the addition of tetracycline results in the disruption of the association of TetR with TetO, thereby inducing TetO-dependent gene expression. (2) a Tet-off configuration, in which tandem TetO sequences are located upstream of a minimal constitutive promoter, followed by the cDNA of the gene of interest. Here, a chimeric protein consisting of TetR and VP16 (tTA), a eukaryotic transactivator derived from herpes simplex virus type 1, is converted into a transcriptional activator, and an expression plasmid is transfected together with the operator plasmid. Thus, culturing cells with tetracycline switches off the expression of the exogenous gene, but removing tetracycline switches it on. (3) Tet-on configuration, in which the exogenous gene is expressed when tetracycline is added to the growth medium. Even though tetracycline is non-toxic to mammalian cells at the low concentrations required to control TetO-dependent gene expression, its continued presence may be undesirable. Therefore, a mutant tTA with four amino acid substitutions, called rtTA, was generated by random mutagenesis of tTA. Unlike tTA, rtTA binds to the TetO sequence in the presence of tetracycline, thereby activating the silent minimal promoter. Operator system controlled by cumate
[0271] Cumate-regulated operators arise from the p-cmt and p-cym operons in Pseudomonas putida. The corresponding repressor contains an N-terminal DNA-binding domain that recognizes the imperfect repeat between the promoter and the start of the first gene in the p-cymene degradation pathway. Similar to the tetracycline-regulated operator system, the cumate operator (CuO) and its repressor (CymR) can be engineered into three configurations: (1) a repressor configuration, achieved by placing CuO downstream of a constitutive promoter, where binding of CymR to CuO effectively represses downstream gene expression. Addition of cumate releases CymR, thereby inducing downstream gene expression. (2) an activator configuration, where a chimeric molecule (cTA) is formed by the fusion of CymR with VP16. In this configuration, a minimal promoter was placed downstream of a multimerized operator binding site (6×CuO). (3) After random mutagenesis and screening, the addition of cumate generates a cTA mutant (rcTA) that binds to CuO, a reverse activator configuration in which the addition of cumate induces downstream gene expression. Chimeric systems based on protein-protein interactions 1. Induction of target genes by modulating the interaction between FKBP12 and mTOR
[0272] Rapamycin and its analog FK506 bind to the cytosolic protein FKBP12. This complex further binds to mTOR, forming a ternary complex. Thus, when FKBP12 and mTOR are fused to the DNA-binding domain of ZFHD1 and the activation domain of the NF-κB p65 protein, respectively, both domains are cross-linked to drive the expression of a gene of interest in a rapamycin-dependent manner. Due to the immunosuppressive and cell cycle inhibitory effects of FK506 and rapamycin, a new synthetic compound, FKCsA, a heterodimer of FK506 and cyclosporin A (an immunosuppressant complexed with the protein cyclophilin), was developed and shown to exhibit neither toxicity nor immunosuppressive effects. To induce gene expression, addition of FKCsA to cells hinges FKBP12 fused to the Gal4 DNA binding domain (Gal4DBD) and cyclophilin fused to VP16, thereby activating expression of the gene of interest downstream of the upstream activation sequence (UAS, Gal4DBD binding site). 2. Induction of target genes by modulating the interaction between PYL1 and ABI1
[0273] Abscisic acid (ABA)-regulated interactions between two plant proteins are used to control gene expression temporally and quantitatively in mammalian cells. The two proteins are PYL1 (abscisic acid receptor) and ABI1 (protein phosphatase 2C56), which are key players in the ABA signaling pathway required for stress responses and developmental decisions in plants. According to the crystal structure of the PYL1-ABA-ABI1 complex, the interacting complementary surfaces of PYL1 (amino acids 33-209) and ABI1 (amino acids 126-423) were selected for chimeric protein construction. Similarly, Gal4DBD was fused with ABI1 and VP16 with PYL1. As a result, after transfecting this ABA-activator cassette and UAS-driven reporter into mammalian cells, ABA significantly induced the production of the reporter. Compared to the rapamycin system, the ABA system has two compelling advantages: first, ABA is present in many foods containing plant extracts and vegetable oils - its lack of toxicity is supported by extensive evaluation by the US Environmental Protection Agency (EPA); second, the ABA signaling pathway is not present in mammalian cells, so there should be no competing endogenous binding proteins as in the case of the rapamycin system. To further avoid any catalysis of possible unexpected substrates by ABI1, mutations essential for its phosphatase activity were introduced into the chimeric protein. 3. Induction of target genes by light-sensitive protein-protein interactions
[0274] Two light-switchable transgene systems have been developed exploiting light-induced protein-protein interactions. The first one was inspired by the molecular basis of fungal circadian rhythms. vivid (VVD), a photoreceptor and light-oxygen-voltage (LOV) domain-containing protein from Neurospora crassa, forms a dimer that rapidly switches upon blue light activation. Thus, a chimeric protein consisting of VVD and Gal4 residues 1-65 dimerizes and becomes a transcriptional activator under blue light illumination, whereas in the absence of blue light the active dimer dissociates. This means that expression of a reporter downstream of the UAS can be switched on and off in a spatiotemporal manner utilizing blue light. Moreover, mutagenesis optimization of VVD further reduced background expression to minimal levels, making the system even more feasible. Another light-switchable transgene system (photoactivatable (PA)-Tet-OFF / ON) utilizes the blue-light-responsive heterodimerization from Arabidopsis thaliana, consisting of the cryptochrome 2 (Cry2) photoreceptor and cryptochrome-interacting basic helix-loop-helix 1 (CIB1). The photolyase homology region (PHR) in the N-terminal part of Cry2 is a chromophore-binding domain that non-covalently binds flavin adenine dinucleotide (FAD). CIB1 interacts with Cry2 in a blue-light-dependent manner. Therefore, to create an inducible expression system, PHR was fused with the transcriptional activation domain of p65, and CIB1 was fused with the DNA-binding, dimerization, and tetracycline-binding domains of TetR (residues 1-206). As a result, the reporter gene can be switched on with blue light illumination, while switching off can be achieved in two ways, either by the absence of blue light or by the addition of tetracycline. On the other hand, the tetracycline-insensitive mutant, H100Y, proved to be completely dependent on illumination. By applying the same chimeric construct but replacing TetR with rtTA, the reporter gene could be switched on by either blue light illumination or tetracycline, and switched off by either the absence of blue light or removal of tetracycline.In general, two advantages of light-switchable transgene systems overwhelm all other systems. One is their rapid on and off cycles. Due to the nature of circadian rhythms, the two abovementioned protein-protein interactions are dynamic, resulting in fast responses and reversals. It has been shown that even a short light pulse of 1-2 minutes is sufficient to induce luciferase expression, which peaks after 1.1 hours and decays to background levels after 3 hours. The other advantage is its precise spatial induction. Illumination within a restricted area or cell population can be achieved with a high-performance illumination source, thereby allowing reporter expression to be selectively induced in certain cells or subcellular regions of interest. These unique features not only greatly facilitate future cell-cell behavior studies, but also bring enormous potential for clinical gene therapy. 4. Systems regulated by tamoxifen
[0275] One of the best-characterized "reversible switch" models, the tamoxifen-inducible system, has several beneficial features (e.g., reviewed by Whitfield et al. (2015) Cold Spring Harb Protoc. 2015(3):227-234). In this system, the hormone-binding domain of the mammalian estrogen receptor is used as the heterologous regulatory domain. Upon ligand binding, the receptor is released from its inhibitory complex and the fusion protein becomes functional. For example, the ligand-binding domain (LBD) of the estrogen receptor (ER) can be fused to a transgene, and the product is a chimeric protein that can be activated by the anti-estrogen tamoxifen or its derivative 4-OH tamoxifen (4-OH-TAM).
[0276] This system has been used in combination with recombinases to generate controllable recombinases that modify genomes. For example, inducible gene expression can be achieved using either a single plasmid system or a two-plasmid system. The first successful case was performed in mouse embryonic cells. Two plasmids were transfected together. One was a Cre-ER constitutive expression plasmid, and the other contained a gene trap sequence flanked by LoxP and followed by a β-galactosidase (LacZ) open reading frame. As a result, expression of LacZ could be restored only when Cre-loxP-mediated recombination was induced and the gene trap sequence was excised. By these means, reporter genes could be induced not only in undifferentiated embryonic stem cells and embryoid bodies, but also in all tissues of 10-day-old chimeric fetuses or in specific differentiated adult tissues. In another example, to induce enhanced green fluorescent protein (EGFP) expression in baby hamster kidney (BHK) cells and to simplify plasmid construction, a Cre-ER cDNA flanked by LoxP sites was inserted between the phosphoglycerate kinase (PGK) promoter and the EGFP coding sequence. In this system, Cre-ER serves as a gene trap to block transcription of EGFP without 4-OH-TAM. Ignition of recombinase activity by 4-OH-TAM melts off the Cre-ER cassette and restores EGFP expression driven by the PGK promoter. To eliminate the effects exerted by endogenous steroids, three distinct ERs are primarily utilized: (1) mouse ERTM with a G525R mutation, (2) human ERT with a G521R mutation, and (3) human ERT2 containing the triple mutations G400V / M543 / L544A. 5. Riboswitch-controlled expression system
[0277] Riboswitch-controllable expression systems utilize bacterial RNA aptamers linked to hammerhead ribozymes (aptazymes). The aptamers act as molecular sensors and transducers for the entire apparatus, while the ribozymes respond to signals involving conformational changes and mRNA cleavage. For example, aptazymes from Gram-positive bacteria directly sense excess glucosamine-6-phosphate (GlcN6P) and cleave the mRNA of the glms gene, the protein product of which is fructose-6-phosphate (Fru6P) and an enzyme that converts glucosamine to GlcN6P. These aptazymes, which respond to tetracycline, theophylline, guanine, etc., have been engineered to either knock down or overexpress genes of interest (e.g., as reviewed by Yokobayashi et al. (2019) Curr Opin Chem Biol 52:72-78). 6. Antisense oligonucleotide (ASO)-controlled expression system
[0278] ASOs can bind to DNA or RNA. ASOs have been demonstrated to act at the RNA level for effective gene regulation, activating RISC complexes and degrading mRNAs or interfering with the recognition of cis-acting elements. ASOs are routinely formulated in lipid nanoparticles that efficiently transfect cells. ASOs are used in "knockdown" applications, either for gain-of-function (i.e., dominant-negative), transcript, or homozygous recessive genetic diseases. In diseases caused by dominant-negative mutations, where ASOs are not specific for transcripts from mutant alleles, such as Huntington's disease and other polyglutamine expansion diseases, restoration of normal cell function can be achieved using gene replacement, using vector-delivered transgenes with alternative synonymous codons that reduce sequence complementarity to the exogenous ASO. Thus, ASOs deplete transcripts from endogenous alleles, but vector-driven transcripts are not affected.
[0279] As illustrated in Figure 14, ASOs can modulate splicing to either negatively or positively control gene expression (see also Havens and Hastings (2016) Nucleic Acids Research 44:6549-6563). Example I in Figure 11 shows that ASOs (antisense oligonucleotide ASOs or AONs) can negatively control gene expression post-transcriptionally. Without an ASO, the primary transcript is spliced into a translatable mRNA. Addition of an ASO (red line) complementary to the splice acceptor at the 3' end of the intron / 5' end of the second exon interferes with splicing. Thus, in the presence of the ASO, the intron remains in the transcript. This unprocessed RNA containing the intron is either untranslatable or translation produces a non-functional protein.
[0280] Example II in FIG. 11 also illustrates that ASOs can positively affect gene expression post-transcriptionally. The primary transcript (left) contains four exons: the first, third and fourth exons code for the therapeutic protein, and the second exon contains either a nonsense mutation or an out-of-frame mutation (OOF). Such a second exon can be engineered into any transgene. Without the ASO, the transcript is processed into a mature mRNA containing four exons, i.e., the second exon with the nonsense mutation or OOF mutation remains. As a result, the resulting mRNA is translated into a truncated or non-functional protein. In contrast, the addition of the ASO interferes with splicing and the mature mRNA consists of the first, third and fourth exons, i.e., the second exon with the nonsense mutation or OOF mutation is excised. As a result, in the default state (without ASO), no therapeutic protein is produced. Only upon addition of the ASO is the therapeutic protein produced, thereby resulting in positive regulation.
[0281] These approaches allow for knockdown of constitutively active transgene expression, i.e., default on. In some embodiments, the default on state is preferred. In other embodiments, the default off condition is preferred. Exemplary Pulsed Gene Expression for Hemophilia A
[0282] In certain aspects, the vectors (e.g., nucleic acid vectors, viral vectors), cells, pharmaceutical compositions, and methods provided herein use pulsed gene expression in gene therapy for subjects suffering from hemophilia A. In some embodiments, an ASO-controlled expression system is used to transduce a gene encoding human coagulation factor VIII (FVIII) into liver cells of subjects suffering from hemophilia A. In some embodiments, pulsed gene expression (the transgene encoding FVIII is turned on and off at certain intervals) is used to control the amount of FVIII produced (see Example 11). The delivery and control of the transgene encoding FVIII or its active fragment (e.g., one lacking its B domain), compositions, and methods described herein address a long-felt and unmet medical need.
[0283] In 2020, the FDA did not approve Biomarin's Biologics License Application (BLA) for valoctocogene loxaparvovec (or BMN270) as a treatment for hemophilia A (HemA). Recombinant adeno-associated virus type 5 (rAAV5) delivered a genetic derivative of human coagulation factor VIII (FVIII) to the liver of HemA patients. At higher doses, FVIII was expressed and secreted into the patient's circulation at levels equal to or higher than physiological levels that effectively "cure" the treated patient. However, long-term expression levels declined from 0.5 to 0.33 per year during the three-year follow-up. Although FVIII expression remained at a level that was clinically beneficial, the FDA expressed concern that patients would revert to their hemophilia phenotype if expression continued to drop at the same rate. No clear explanation for the reduced expression pattern: Previous clinical studies in hemophilia B demonstrated that the lack of FIX expression was primarily due to acute inflammation triggered by processed AAV capsid antigens. However, prophylactic steroid treatment attenuated or abolished the capsid immune response and is now routine for liver-directed rAAV treatments. Several possible explanations for the loss of FVIII expression are contemplated herein.
[0284] FVIII has been a difficult recombinant protein to produce in both microbial and eukaryotic expression systems. The development of a "B domain" deleted version of FVIII reduced the size of the open reading frame and improved expression levels. However, FVIII expression levels were still substantially lower than other proteins. To overcome these low levels, Biomarin increased the vector dose in clinical studies. Patients were treated with 6E+13 vector particles (called vector genomes, or vg) per kg. Based on large animal models, only a small number of hepatocytes take up (transduced) rAAV5-FVIII and subsequently express relatively large amounts of FVIII as a result of the high number of vg per cell. The metabolic demands of FVIII expression likely interfere with the normal requirements of hepatocyte protein expression. FVIII may be overpacked into hepatocyte intracellular compartments normally involved in protein folding and secretion. Endothelial cells that give rise to FVIII production are likely specialized for this activity, producing FVIII from a single allele on the X chromosome under the transcriptional control of the highly regulated native FVIII promoter.
[0285] Therefore, to prevent the expression of the transgene encoding FVIII from gradually decreasing, the transgene is turned on and off at regular intervals to achieve long-term efficacy.The timing of the pulse is determined based on the serum level and half-life of FVIII protein (see Example 11 for details).The ideal state for FVIII to prevent or treat hemophilia A is off until it is transiently activated.ASO can be used to cause either negative or positive effects by interfering with cis-acting elements in the primary transcript, thereby providing flexibility in controlling pulsed gene expression. Viral Vectors
[0286] In certain aspects, provided herein are viral vectors comprising a nucleic acid vector described herein (e.g., comprising at least a portion of the GSH locus of the present disclosure, a nucleic acid vector for integration into the GSH locus of the present disclosure, etc.). In some embodiments, the viral vector is selected from rAd, AAV, rHSV, retroviral vectors, poxvirus vectors, lentivirus, vaccinia virus vectors, HSV type 1 (HSV-1)-AAV hybrid vectors, baculovirus expression vector systems (BEVS), and variants thereof.
[0287] Specifically, viral vector refers to a virus or viral chromosomal material that can insert a piece of foreign DNA for cell transfer.Any virus that contains a DNA stage in its life cycle can be used as a viral vector in the subject method and composition.For example, the virus can be a single-stranded DNA (ssDNA) virus or a double-stranded DNA (dsDNA) virus.Also suitable are RNA viruses that have a DNA stage in their life cycle, such as retroviruses that are reverse transcribed into DNA, e.g., MMLV, lentivirus.The virus can be an integrated or non-integrated virus.
[0288] Viral vectors encompassed for use in the methods and compositions disclosed herein are discussed in the review articles Hendrie, Paul C., and David W. Russell. "Gene targeting with viral vectors." Molecular Therapy 12.1 (2005): 9-17 and Perez-Pinera, "Advances in targeted genome editing." Current opinion in chemical biology 16.3 (2012): 268-277.
[0289] Adeno-associated virus ("AAV") vectors are encompassed for use as nucleic acid vector compositions disclosed herein and are useful in in vivo and ex vivo gene therapy procedures (see, e.g., West et al., Virology 160:38-47 (1987); U.S. Pat. No. 4,797,368; WO93 / 24641; Kotin, Human Gene Therapy 5:793-801 (1994); Muzyczka, J. Clin. Invest. 94:1351 (1994). Construction of recombinant AAV vectors is described in U.S. Pat. No. 5,173,414; Tratschin et al, Mol. Cell. Biol. 5:3251-3260 (1985); Tratschin, et al., Mol. Cell. Biol.4:2072-2081 (1984); Hermonat & Muzyczka, PNAS 81:6466-6470 (1984); and Samulski et al., J. Virol. 63:03822-3828 (1989). At least six viral vector approaches are currently available for gene transfer in clinical trials, which utilize an approach that involves complementation of a defective vector with a gene inserted into a helper cell line to generate the transducing agent.
[0290] In a preferred embodiment, the viral vector is an adeno-associated virus. Adeno-associated virus, or "AAV", refers to the virus itself or its derivatives. The term includes all subtypes and both naturally occurring and recombinant forms, such as AAV type 1 (AAV-1), AAV type 2 (AAV-2), AAV type 3 (AAV-3), AAV type 4 (AAV-4), AAV type 5 (AAV-5), AAV type 6 (AAV-6), AAV type 7 (AAV-7), AAV type 8 (AAV-8), AAV type 9 (AAV-9), AAV type 10 (AAV-10), AAV type 11 (AAV-11), AAV type 12 (AAV-12), AAV type 13 (AAV-13), AAV type 14 (AAV-14), AAV type 15 (AAV-15), AAV type 16 (AAV-16), AAV type 17 (AAV-17), AAV type 18 (AAV-18), AAV type 19 (AAV-19), AAV type 20 (AAV-20), AAV type 21 (AAV-21), AAV type 22 (AAV-22), AAV type 23 (AAV-23), AAV type 24 (AAV-24), AAV type 25 (AAV-25), AAV type 26 (AAV-26), AAV type 27 (AAV-27), AAV type 28 (AAV-28), AAV type 29 (AAV-29), AAV type 30 (AAV-29), AAV type 31 (AAV-29), AAV type 32 (AAV-29), AAV type 33 (AAV-29), AAV type 34 (AAV-29), AAV type 35 (AAV-29), AAV type 36 (AAV-29), AAV type The term covers AAV, including AAV type 9 (AAV-9), AAV type 10 (AAV-10), AAV type 11 (AAV-11), AAV type 12 (AAV-12), AAV type 13 (AAV-13), avian AAV, bovine AAV, canine AAV, equine AAV, primate AAV, non-primate AAV, ovine AAV, hybrid AAV (i.e., an AAV that contains capsid protein of one AAV subtype and genomic material of another subtype), mutant AAV capsid protein, or chimeric AAV capsid (i.e., a capsid protein having regions or domains or individual amino acids derived from two or more different serotypes of AAV, e.g., AAV-DJ, AAV-LK3, AAV-LK19), unless otherwise required. "Primate AAV" refers to AAV that infects primates, "non-primate AAV" refers to AAV that infects non-primate mammals, "bovine AAV" refers to AAV that infects mammals of the bovine subfamily, etc.
[0291] Recombinant AAV vector or rAAV vector refers to an AAV virus or AAV virus chromosomal material that includes a polynucleotide sequence that is not of AAV origin (i.e., a polynucleotide heterologous to AAV), typically a nucleic acid sequence of interest (e.g., a non-GSH nucleic acid) that is to be integrated into a cell. Generally, the heterologous polynucleotide is flanked by at least one, and generally two, AAV inverted terminal repeats (ITRs). In some cases, the recombinant viral vector also includes viral genes that are important for packaging the recombinant viral vector material. "Packaging" refers to the series of intracellular events that lead to the assembly and encapsidation of viral particles, e.g., AAV viral particles. Examples of nucleic acid sequences (i.e., "packaging genes") that are important for AAV packaging include the AAV "rep" and "cap" genes, which code for the replication and encapsidation proteins of adeno-associated virus, respectively. The term rAAV vector encompasses both rAAV vector particles and rAAV vector plasmids.
[0292] Viral particle refers to a single unit of virus, including a capsid that encapsidates a polynucleotide based on virus, such as a viral genome (such as in wild-type virus), or a targeting vector (such as in recombinant virus). AAV viral particle refers to a viral particle that is composed of at least one AAV capsid protein (typically all wild-type AAV capsid proteins) and an encapsidated polynucleotide AAV vector. When a particle includes a heterologous polynucleotide (i.e., a polynucleotide other than a wild-type AAV genome, such as a transgene that is to be delivered to mammalian cells), it is generally referred to as a rAAV vector particle or simply a rAAV vector. Thus, the production of rAAV particles necessarily includes the production of rAAV vector, and therefore the vector is contained within the rAAV particle.
[0293] In some embodiments, recombinant adeno-associated virus ("rAAV") vectors are derived from plasmids that carry only the AAV 145bp inverted terminal repeats flanking the transgene expression cassette. Efficient gene transfer and stable transgene delivery due to integration into the genome of transduced cells are essential features of this vector system. (Wagner et al., Lancet 351:9117 1702-3 (1998); Kearns et al., GeneTher. 9:748-55 (1996)). All AAV serotypes, including AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, and AAVrh.10, as well as any novel AAV serotypes, can also be used according to the present invention.
[0294] Replication-defective recombinant adenovirus vectors (Ad) are also included in the present specification, and can be produced at high titers and easily infect several different cell types. One example of the use of Ad vectors in clinical trials included polynucleotide therapy for anti-tumor immunization by intramuscular injection (Sterman et al., Hum. Gene Ther. 7: 1083-9 (1998)). Further examples of the use of adenoviral vectors for gene transfer in clinical trials include Rosenecker et al., Infection 24:1 5-10 (1996); Sterman et al., Hum. Gene Ther. 9:71083-1089 (1998); Welsh et al., Hum. Gene Ther. 2:205-18 (1995); Alvarez et al., Hum. Gene Ther. 5:597-613 (1997); Topf et al., Gene Ther. 5:507-513 (1998); Stermanet al., Hum. Gene Ther. 7:1083-1089 (1998).
[0295] Retroviral vectors are included for use as the nucleic acid vector compositions disclosed herein.pLASN and MFG-S are examples of retroviral vectors that have been used in clinical trials (Dunbar et al, Blood 85:3048-305 (1995);Kohn et al., Nat. Med.1:1017-102 (1995);Malech et al, PNAS 94:22 12133-12138 (1997)).
[0296] Vectors suitable for the methods and compositions disclosed herein include lentiviral vectors, such as those disclosed in Picanco-Castro. "Advances in lentiviral vectors: a patent review." Recent patents on DNA & gene sequences 6.2 (2012): 82-90. The tropism of retroviruses can be altered by incorporating foreign coat proteins, expanding the potential target population of target cells. Lentiviral vectors are retroviral vectors that can transduce or infect non-dividing cells and generally produce high viral titers. The choice of retroviral gene transfer system depends on the target tissue. Retroviral vectors are composed of cis-acting long terminal repeats (LTRs) capable of packaging up to 6-10 kb of foreign sequences. The minimal cis-acting LTRs are sufficient for vector replication and packaging, and therefore these vectors are used to integrate therapeutic genes into target cells to provide persistent transgene expression. Widely used retroviral vectors include those based on murine leukemia virus (MuLV), gibbon ape leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (see, e.g., Buchscher et al., J. Virol. 66:2731-2739 (1992); Johann et al, J. Virol. 66:1635-1640(1992); Sommerfelt et al, Virol. 176:58-59' (1990); Wilson et al, J. Virol.63:2374-2378 (1989); Miller et al, J. Virol. 65:2220- 2224 (1991); PCT / US94 / 05700).Other retroviral vectors for use herein include foamy viruses, as disclosed in Sweeney, Nathan Paul, et al. "Delivery of large transgene cassettes by foamy virus vector." Scientific reports 7 (2017) 8085.
[0297] Lentiviral transfer vectors can generally be produced by methods well known in the art.See, for example, U.S. Patent Nos. 5,994,136, 6,165,782, and 6,428,953, U.S. Patent Application Publication No. 2014 / 0315294, and described in Merten et al "Production of lentiviral vectors." Molecular Therapy-Methods & Clinical Development 3 (2016): 16017 and Merten, et al. "Large-scale manufacture and characterization of a lentiviral vector produced for clinical ex vivo gene therapy application." Human genetherapy 22.3 (2010): 343-356, each of which is incorporated herein by reference in its entirety.In some embodiments, the lentivirus is an integrase-deficient lentiviral vector (IDLV). IDLV can be produced using lentiviral vectors that contain one or more mutations of the native lentiviral integrase gene, e.g., as described, e.g., as disclosed in Leavittet al. (1996) J. Virol. 70(2):721-728; Philippe et al. (2006) Proc. Nat II Acad.ScL USA 103(47): 17684-17689; and WO06 / 010834. Lentiviruses for use in the methods and compositions disclosed herein are disclosed in Patent Nos. 6,207,455, 5,994,136, 7,250,299, 6,235,522, 6,312,682, 6,485,965, 5,817,491, and 5,591,624.
[0298] Suitable vectors in the methods and compositions disclosed herein include non-integrating lentiviral vectors (IDLV).See, for example, Ory et al. (1996) Proc. Natl. Acad. Sci. USA 93: 11382-1 1388;Dullet et al. (1998) J. Virol. 72:8463-8471;Zuffery et al. (1998) J. Virol.72:9873-9880;Follenzi et al. (2000) Nature Genetics 25:217-222;US Patent Application Publication No. 2009 / 054985.In certain embodiments, IDLV is an HIV lentiviral vector that comprises a mutation at position 64 of integrase protein (D64V), as described in Leavittet et al. (1996) J. Virol. 70(2):721-728. Further IDLV vectors suitable for use herein are described in US patent application Ser. No. 12 / 288,847, which is incorporated herein by reference.
[0299] Vectors suitable for use in the methods and compositions disclosed herein include recombinant HCMV and RHCMV vectors, such as those disclosed in US2013 / 0136,768.
[0300] The nucleic acid vector useful herein for the introduction of the nucleic acid of interest into hematopoietic stem cells, such as CD34+ cells, includes adenovirus type 35. The nucleic acid vector useful herein for the introduction of the nucleic acid of interest into immune cells (e.g., T cells) includes non-integrating lentivirus vectors.See, for example, Ory et al. (1996) Proc. Natl. Acad. Sci. USA 93:11382-11388;Dull et al. (1998) J. Virol. 72:8463- 8471;Zuffery et al. (1998) J. Virol. 72:9873-9880;Follenzi et al. (2000) Nature Genetics 25:217-222.
[0301] Vectors suitable for use in the methods and compositions disclosed herein include the baculovirus expression vector system (BEVS), discussed in Felberbaum, "The baculovirus expression vector system: a commercial manufacturing platform for viral vaccines and gene therapy vectors." Biotechnology journal 10.5(2015): 702-714.
[0302] Vectors suitable for the methods and compositions disclosed herein include HSV type 1 (HSV-1)-AAV hybrid vectors, for example, as disclosed in Heister, Thomas, et al. "Herpes simplex virus type 1 / adeno-associated virus hybrid vectors mediate site-specific integration at the adeno-associated virus preintegration site, AAVS1, on human chromosome 19." Journal of virology 76.14 (2002): 7163-7173, and 5,965,441. Other hybrid vectors, for example, as disclosed in U.S. Patent No. 6,218,186, can be used. Cells containing one or more nucleic acid vectors and / or viral vectors
[0303] In certain aspects, provided herein is a cell comprising at least one nucleic acid vector of the present disclosure or at least one viral vector of the present disclosure.
[0304] In some embodiments, the cells are selected from a cell line or a primary cell.
[0305] In some embodiments, the cell is a mammalian cell, an insect cell, a bacterial cell, a yeast cell, or a plant cell, and optionally, the mammalian cell is a human cell or a rodent cell. In some embodiments, the cell is an insect cell, and the insect cell is from a species of lepidoptera. In some embodiments, the species of lepidoptera is Spodoptera frugiperda, Spodoptera littoralis, Spodoptera exigua, or Trichoplusia ni. In some embodiments, the insect cell is Sf9.
[0306] In some embodiments, the cell is a hematopoietic cell, hematopoietic progenitor cell, hematopoietic stem cell, erythroid lineage cell, megakaryocyte, erythroid progenitor cell (EPC), CD34+ cell, CD44+ cell, red blood cell, CD36+ cell, mesenchymal stem cell, neuronal cell, intestinal cell, intestinal stem cell, gastrointestinal epithelial cell, endothelial cell, enteroendocrine cell, lung cell, lung progenitor cell, enterocyte, liver cell (e.g., hepatocyte, hepatic stellate cell, Kupffer cell (KC), liver sinusoidal endothelial cell (LSEC), liver progenitor cell), stem cell, progenitor cell, induced pluripotent stem cell (iPSC), skin fibroblast, macrophage, brain microvascular endothelial cell. (BMVEC), neural stem cells, muscle satellite cells, epithelial cells, airway epithelial cells, muscle progenitor cells, erythroid progenitor cells, lymphoid progenitor cells, B lymphoblast cells, B cells, T cells, basophilic endemic Burkitt's lymphoma (EBL), polychromatic erythroblasts, epidermal stem cells, epithelial stem cells, embryonic stem cells, P63 positive keratinocyte derived stem cells, keratinocytes, pancreatic beta cells, K cells, L cells, HEK293 cells, HEK293T cells, MDCK cells, Vero cells, CHO, BHK1, NS0, Sp2 / 0, HeLa, A549, and normochromatic erythroblasts. Cells in which at least one non-GSH nucleic acid has been integrated into one or more GSH loci
[0307] Viral vectors include both DNA and RNA viruses, which have either episomal or integrated genomes after delivery to a cell. For reviews of gene therapy procedures, see Anderson, Science 256:808-813 (1992);Nabel & Felgner, TIBTECH11:211-217 (1993);Mitani & Caskey, TIBTECH 11:162-166 (1993);Dillon, TIBTECH 11: 167-175 (1993);Miller, Nature 357:455-460 (1992);Van Brunt, Biotechnology 6(10): 1149-1154 (1988);Vigne, Restorative Neurology andNeuroscience 8:35-36 (1995);Kremer & Perricaudet, British Medical Bulletin5 1(1):3 1-44 (1995);Haddada et al., in Current Topics in Microbiology andImmunology See Doerfler and Bohm (eds.) (1995); and Yu et al., Gene Therapy 1:13-26(1994).
[0308] Thus, in certain aspects, provided herein is a cell comprising at least one non-GSH nucleic acid integrated into GSH in the genome of the cell, wherein GSH is selected from Table 3. In some embodiments, the GSH nucleic acid comprises a non-translated sequence or intron. In some embodiments, the GSH is selected from SYNTX-GSH1, SYNTX-GSH2, SYNTX-GSH3, and SYNTX-GSH4. In some embodiments, the at least one non-GSH nucleic acid is integrated into one or more GSH loci described herein.
[0309] It is contemplated herein that the cell may be integrated with at least one of any one of the nucleic acid vectors described herein. In some embodiments, any one of the nucleic acid vectors is delivered to the cell by any one of the viral vectors described herein.
[0310] In certain embodiments, the cell comprises at least one non-GSH nucleic acid integrated into GSH in a forward orientation. In some embodiments, at least one non-GSH nucleic acid is integrated into GSH in a reverse orientation.
[0311] In certain embodiments, the cell comprises at least one non-GSH nucleic acid incorporated into GSH, and the at least one non-GSH nucleic acid is (a) operably linked to a promoter or (b) not operably linked to a promoter.
[0312] In some embodiments, the at least one non-GSH nucleic acid is operably linked to a promoter, the promoter being selected from: (a) a promoter heterologous to the nucleic acid to which it is operably linked; (b) a promoter that promotes tissue-specific expression of the nucleic acid; (c) a promoter that promotes constitutive expression of the nucleic acid; (d) an inducible promoter; (e) an immediate early promoter of an animal DNA virus; (f) an immediate early promoter of an insect virus; and (g) an insect cell promoter.
[0313] In some embodiments, the inducible promoter operably linked to at least one non-GSH nucleic acid is modulated by an agent selected from a small molecule, a metabolite, an oligonucleotide, a riboswitch, a peptide, a peptidomimetic, a hormone, a hormone analog, and light. In some embodiments, the agent is selected from tetracycline, cumate, tamoxifen, estrogen, and an antisense oligonucleotide (ASO), rapamycin, FKCsA, blue light, abscisic acid (ABA), and a riboswitch.
[0314] In some embodiments, the promoter that promotes tissue-specific expression of at least one non-GSH nucleic acid is a promoter that promotes tissue-specific expression in hematopoietic stem cells, hematopoietic CD34+ cells, and epidermal stem cells, epithelial stem cells, neural stem cells, lung progenitor cells, muscle satellite cells, intestinal K cells, neuronal cells, airway epithelial cells, or liver progenitor cells.
[0315] In some embodiments, the promoter operably linked to at least one non-GSH nucleic acid is selected from a CMV promoter, a β-globin promoter, a CAG promoter, an AHSP promoter, an MND promoter, a Wiskott-Aldrich promoter, a PKLR promoter, a polyhedrin (polh) promoter, and an immediate early 1 gene (IE-1) promoter.
[0316] In certain embodiments, the cell comprises at least one non-GSH nucleic acid incorporated into GSH, and at least one non-GSH nucleic acid comprises a sequence encoding a coding RNA.In some embodiments, the sequence encoding the coding RNA is codon-optimized for expression in target cell.In some embodiments, at least one non-GSH nucleic acid encoding a coding RNA further comprises a sequence encoding a signal peptide.
[0317] In some embodiments, the cells comprise at least one non-GSH nucleic acid incorporated into GSH, wherein the at least one non-GSH nucleic acid encoding a coding RNA comprises a sequence encoding: (a) a protein or a fragment thereof, preferably a human protein or a fragment thereof; (b) a therapeutic protein or a fragment thereof, an antigen-binding protein, or a peptide; (c) a suicide gene, optionally Herpes Simplex Virus-1 Thymidine Kinase (HSV-TK); (d) a viral protein or a fragment thereof; (e) a nuclease, optionally a transcription activator-like effector nuclease (TALEN), zinc finger nuclease (ZFN), meganuclease, megaTAL, or CRISPR endonuclease, (e.g., Cas9 endonuclease or a variant thereof); (f) a marker, e.g., luciferase or GFP; and / or (g) a drug resistance protein, e.g., an antibiotic resistance gene, e.g., neomycin resistance.
[0318] The viral protein or fragment thereof may comprise a structural protein (e.g., VP1, VP2, VP3) or a non-structural protein (e.g., Rep protein). In some embodiments, the viral protein or fragment thereof comprises (a) a parvovirus protein or fragment thereof, optionally VP1, VP2, VP3, NS1, or Rep; (b) a retrovirus protein or fragment thereof, optionally envelope protein, gag, pol, or VSV-G; (c) an adenovirus protein or fragment thereof, optionally E1A, E1B, E2A, E2B, E3, E4, or a structural protein (e.g., A, B, C); and / or (d) a herpes simplex virus protein or fragment thereof, optionally ICP27, ICP4, or pac.
[0319] In some embodiments, the cell comprises at least one non-GSH nucleic acid encoding a viral protein that is a surface protein of the virus. In some embodiments, the at least one non-GSH nucleic acid encoding a viral protein encodes a viral surface protein or a fragment thereof. In some embodiments, (a) the surface protein or a fragment thereof is an immunogenic surface protein that elicits an immune response in the host, (b) the surface protein or a fragment thereof further comprises a signal peptide, (c) the gene encoding the surface protein or a fragment thereof is operably linked to an inducible promoter, and / or (d) the nucleic acid encoding the surface protein or a fragment thereof further comprises a suicide gene. Cells comprising such nucleic acids are useful not only for producing recombinant viral proteins in vitro for use as a vaccine, but also for implantation into a subject for expression of viral proteins for in vivo immunization in vivo. The in vivo production of viral proteins can be under an inducible promoter, and thus the duration of production as well as the amount of immunogen produced in vivo can be fine-tuned using signals or agents that modulate the inducible promoter (see, for example, the section on pulsed expression systems described herein).
[0320] In some embodiments, such cells for producing vaccines in vitro or for in vivo immunization express a viral surface protein, the surface protein being of a coronavirus (e.g., MERS, SARS), influenza virus, respiratory syncytial virus, hepatitis A, hepatitis B, hepatitis C, hepatitis D, hepatitis E, human papillomavirus, dengue virus serotype 1, dengue virus serotype 2, dengue virus serotype 3, dengue virus serotype 4, Zika virus, West Nile virus, yellow fever virus, chikungunya virus, Mayaro virus, Ebola virus, Marburg virus, or Nipah virus. In some embodiments, the surface protein is the spike protein of SARS-CoV-2.
[0321] In some embodiments, the cell comprises at least one non-GSH nucleic acid incorporated into GSH, wherein the at least one non-GSH nucleic acid encodes a polypeptide or a fragment thereof. In a preferred embodiment, such a polypeptide or a fragment thereof is a therapeutic protein or a fragment thereof. In some embodiments, the at least one non-GSH nucleic acid comprising a sequence encoding a protein or a fragment thereof is selected from the group consisting of a hemoglobin gene (HBA1, HBA2, HBB, HBG1, HBG2, HBD, HBE1, and / or HBZ), alpha-hemoglobin stabilizing protein (AHSP), coagulation factor VIII, coagulation factor IX, von Willebrand factor, dystrophin or truncated dystrophin, microdystrophin, utrophin or truncated utrophin, microutrophin, usherin (USH2A), GBA1, preproinsulin, insulin, GIP, GLP-1, CEP290, ATPB1, ATPB11, ABCB4, CPS1, ATP7B, KRT5, KRT14, PLEC1, Col7A1, ITGB4, ITGA6, LAMA3, LAMB3, LAMC2, KIND1, INS, F8, or a fragment thereof (e.g., a fragment encoding a B domain deleted polypeptide ... SQ, p-VIII)), IRGM, NOD2, ATG2B, ATG9, ATG5, ATG7, ATG16L1, BECN1, EI24 / PIG8, TECPR2, WDR45 / WIP14, CHMP2B, CHMP4B, Dynein, EPG5, HspB8, LAMP2, LC3b Selected from UVRAG, VCP / p97, ZFYVE26, PARK2 / Parkin, PARK6 / PINK1, SQSTM1 / p62, SMURF, AMPK, ULK1, RPE65, CHM, RPGR, PDE6B, CNGA3, GUCY2D, RS1, ABCA4, MYO7A, HFE, hepcidin, genes encoding soluble forms (e.g., of the TNFα receptor, IL-6 receptor, IL-12 receptor or IL-1β receptor), and cystic fibrosis transmembrane conductance regulator (CFTR).
[0322] In some embodiments, at least one non-GSH nucleic acid comprises a sequence encoding a suicide protein.
[0323] In some embodiments, the cell comprises at least one non-GSH nucleic acid incorporated into GSH, and the at least one non-GSH nucleic acid encodes an antigen-binding protein. In some embodiments, the antigen-binding protein is an antibody or an antigen-binding fragment thereof, and optionally the antibody or antigen-binding fragment thereof is selected from an antibody, Fv, F(ab')2, Fab', dsFv, scFv, sc(Fv)2, half-antibody-scFv, tandem scFv, Fab / scFv-Fc, tandem Fab', single-chain diabody, tandem diabody (TandAb), Fab / scFv-Fc, scFv-Fc, heterodimeric IgG (CrossMab), DART, and diabody.
[0324] In some embodiments, the antigen binding protein specifically binds to TNFα, CD20, a cytokine (e.g., IL-1, IL-6, BLyS, APRIL, IFN-gamma, etc.), Her2, RANKL, IL-6R, GM-CSF, CCR5, or a pathogen (e.g., a bacterial toxin, a viral capsid protein, etc.).
[0325] In some embodiments, the antigen binding protein is selected from adalimumab, etanercept, infliximab, certolizumab, golimumab, anakinra, rituximab, abatacept, tocilizumab, natalizumab, canakinumab, atacicept, belimumab, ocrelizumab, ofatumumab, fontolizumab, trastuzumab, denosumab, sarilumab, lenzilumab, gimcilumab, siltuximab, leronlimab, and antigen-binding fragments thereof.
[0326] Further contemplated herein is a cell comprising at least one non-GSH nucleic acid incorporated into GSH, wherein the at least one non-GSH nucleic acid comprises a sequence encoding a non-coding RNA.In some embodiments, the non-coding RNA comprises lncRNA, piRNA, miRNA, shRNA, siRNA, antisense RNA, snoRNA, snRNA, scaRNA, and / or guide RNA.In some embodiments, the non-coding RNA targets a gene selected from the following genes: DMT-1, ferroportin, TNFα receptor, IL-6 receptor, IL-12 receptor, IL-1β receptor, and genes encoding mutant proteins (e.g., mutant HFE, CFTR).
[0327] In some embodiments, the cells comprise at least one non-GSH nucleic acid incorporated into GSH, which increases or restores expression of an endogenous gene in the target cell. In some embodiments, the cells comprise at least one non-GSH nucleic acid incorporated into GSH, which decreases or eliminates expression of an endogenous gene in the target cell.
[0328] In some embodiments, the cells comprise at least one non-GSH nucleic acid integrated into GSH, and the at least one non-GSH nucleic acid further comprises (a) a transcriptional control element (e.g., an enhancer, a transcription termination sequence, an untranslated region (5' or 3' UTR), a proximal promoter element, a locus control region (e.g., the β-globin LCR, or a DNase hypersensitive site (HS) of the β-globin LCR), a polyadenylation signal sequence), and / or (b) a translational control element (e.g., a Kozak sequence, a woodchuck hepatitis virus post-transcriptional control element).
[0329] In some embodiments, the cells are selected from a cell line or a primary cell.
[0330] In some embodiments, the cell is a mammalian cell, an insect cell, a bacterial cell, a yeast cell, or a plant cell, and optionally, the mammalian cell is a human cell or a rodent cell. In some embodiments, the cell is an insect cell, and the insect cell is from a species of lepidoptera. In some embodiments, the species of lepidoptera is Spodoptera frugiperda, Spodoptera littoralis, Spodoptera exigua, or Trichoplusia ni. In some embodiments, the insect cell is Sf9.
[0331] In some embodiments, the cell is a hematopoietic cell, hematopoietic progenitor cell, hematopoietic stem cell, erythroid lineage cell, megakaryocyte, erythroid progenitor cell (EPC), CD34+ cell, CD44+ cell, red blood cell, CD36+ cell, mesenchymal stem cell, neuronal cell, intestinal cell, intestinal stem cell, gastrointestinal epithelial cell, endothelial cell, enteroendocrine cell, lung cell, lung progenitor cell, enterocyte, liver cell (e.g., hepatocyte, hepatic stellate cell, Kupffer cell (KC), liver sinusoidal endothelial cell (LSEC), liver progenitor cell), stem cell, progenitor cell, induced pluripotent stem cell (iPSC), skin fibroblast, macrophage, brain microvascular endothelial cell. (BMVEC), neural stem cells, muscle satellite cells, epithelial cells, airway epithelial cells, muscle progenitor cells, erythroid progenitor cells, lymphoid progenitor cells, B lymphoblast cells, B cells, T cells, basophilic endemic Burkitt's lymphoma (EBL), polychromatic erythroblasts, epidermal stem cells, epithelial stem cells, embryonic stem cells, P63 positive keratinocyte derived stem cells, keratinocytes, pancreatic beta cells, K cells, L cells, HEK293 cells, HEK293T cells, MDCK cells, Vero cells, CHO, BHK1, NS0, Sp2 / 0, HeLa, A549, and normochromatic erythroblasts.
[0332] Further description of cells comprising the nucleic acid vectors or viral vectors of the disclosure, or cells comprising at least one non-GSH nucleic acid integrated into GSH, is provided below. cell
[0333] Cells comprising the nucleic acid, nucleic acid vector or viral vector of the disclosure are provided herein. A further object of the present invention relates to cells transfected, infected, transduced or transformed by the nucleic acid, nucleic acid vector and / or viral vector according to the present invention. The term "transformation" refers to the introduction of a "foreign" (i.e., exogenous or extracellular) gene, DNA or RNA sequence into a cell, such that the cell expresses the introduced gene or sequence to produce a desired substance, typically a protein or enzyme encoded by the introduced gene or sequence. A cell that receives and expresses the introduced DNA or RNA is "transduced".
[0334] The nucleic acids or nucleic acid vectors of the invention can be used to produce recombinant polypeptides of the invention in a suitable expression system. The term "expression system" refers, for example, to a cell and a compatible vector under suitable conditions for the expression of a protein encoded by foreign DNA carried by the vector and introduced into the cell.
[0335] Common expression systems include E. coli cells and plasmid vectors, insect cells and baculovirus vectors, and mammalian cells and vectors. Other examples of cells include, without limitation, prokaryotic cells (e.g., bacteria) and eukaryotic cells (e.g., yeast cells, mammalian cells, insect cells, plant cells, etc.). Specific examples include E. coli, Kluyveromyces or Saccharomyces yeast, mammalian cell lines (e.g., Vero cells, CHO cells, 3T3 cells, COS cells, etc.), and primary or established mammalian cell cultures (e.g., those produced from lymphoblasts, fibroblasts, embryonic cells, epithelial cells, neural cells, adipocytes, etc.). Examples include mouse SP2 / 0-Ag14 cells (ATCC CRL1581), mouse P3X63-Ag8.653 cells (ATCC CRL1580), CHO cells lacking the dihydrofolate reductase gene (hereinafter referred to as the "DHFR gene") (Urlaub G et al; 1980), rat YB2 / 3HL.P2.G11.16Ag.20 cells (ATCC CRL 1662, hereinafter referred to as "YB2 / 0 cells"), etc. YB2 / 0 cells are preferred because the ADCC activity of chimeric or humanized antibodies is enhanced when expressed in these cells.
[0336] The present invention also relates to a method for producing a recombinant cell expressing an antibody or a polypeptide of the invention according to the invention, comprising the steps of (i) introducing a recombinant nucleic acid, nucleic acid vector or viral vector described herein into a competent cell in vitro or ex vivo, (ii) culturing the resulting recombinant cell in vitro or ex vivo, and (iii) optionally selecting cells expressing and / or secreting an antigen binding protein (e.g., an antibody) or polypeptide (e.g., insulin). Such recombinant cells can be used for the production of various polypeptides described herein.
[0337] A cell as used herein includes any type of cell that can contain the vector of the present disclosure and produce the expression product (e.g., mRNA, protein) encoded by the nucleic acid. The cells in some embodiments are adherent cells or suspension cells, i.e., cells that grow in suspension. The cells in various embodiments are cultured cells or primary cells, i.e., cells that are directly isolated from an organism, e.g., a human. The cells can be of any cell type, originate from any type of tissue, and be at any developmental stage.
[0338] In certain embodiments, the antigen binding protein is a glycosylated protein and the cell is a glycosylation-competent cell. In various embodiments, the glycosylation-competent cell is a eukaryotic cell, including but not limited to a yeast cell, a filamentous fungal cell, a protozoan cell, an algae cell, an insect cell, or a mammalian cell. Such cells are described in the art. See, for example, Frenzel, et al., Front Immunol 4: 217 (2013). In various embodiments, the eukaryotic cell is a mammalian cell. In various embodiments, the mammalian cell is a non-human mammalian cell. In some embodiments, the cells are selected from the group consisting of Chinese hamster ovary (CHO) cells and derivatives thereof (e.g., CHO-K1, CHO pro-3), mouse myeloma cells (e.g., NS0, GS-NS0, Sp2 / 0), cells engineered to be deficient in dihydrofolate reductase (DHFR) activity (e.g., DUKX-X11, DG44), human embryonic kidney 293 (HEK293) cells or derivatives thereof (e.g., HEK293T, HEK293-EBNA), green African monkey kidney cells (e.g., COS cells, VERO cells), human cervical cancer cells (e.g., HeLa), human bone osteosarcoma epithelial cells U2-OS, adenocarcinoma human alveolar basal epithelial cells A549, human fibrosarcoma cells HT1080, mouse brain tumor cells CAD, embryonal carcinoma cells P19, mouse embryonic fibroblast cells NIH 3T3, mouse fibroblast L929, mouse neuroblastoma N2a, human breast cancer MCF-7, retinoblastoma Y79, human retinoblastoma SO-Rb50, human hepatocellular carcinoma Hep G2, mouse B myeloma J558L, or baby hamster kidney (BHK) cells (Gaillet et al. 2007; Khan, Adv Pharm Bull 3(2): 257-263 (2013)).
[0339] In some embodiments, for purposes of amplifying or replicating the vector, the cell is in some aspects a prokaryotic cell, e.g., a bacterial cell.
[0340] The present disclosure also provides a population of cells comprising at least one cell described herein. The population of cells in some aspects is a heterogeneous population, comprising cells comprising the vectors described in addition to at least one other cell that does not comprise any of the vectors. Alternatively, in some aspects, the population of cells is a substantially homogeneous population, comprising primarily (e.g., consisting essentially of) cells comprising vectors. The population in some aspects is a clonal population of cells, in which all cells in the population are clones of a single cell comprising vectors, and therefore all cells of the population comprise that vector. In various embodiments of the present disclosure, the population of cells is a clonal population comprising cells comprising vectors described herein.
[0341] In certain aspects, the cells are human cells that are autologous or allogeneic to the subject. In some embodiments, the nucleic acid of the present invention is transduced by a viral vector or transformed by other suitable methods (e.g., electroporation, etc.). Such cells are transferred (e.g., transplanted, implanted, etc.) into the subject for long-term treatment of a disease or condition, such as cancer. Transgenic Organisms
[0342] In certain aspects, provided herein is a transgenic organism comprising at least one non-GSH nucleic acid integrated into a GSH in the genome of a cell, wherein the GSH is selected from Table 3. In some embodiments, the GSH is selected from SYNTX-GSH1, SYNTX-GSH2, SYNTX-GSH3, and SYNTX-GSH4.
[0343] In some embodiments, the transgenic organism comprises any one of the nucleic acid vectors, viral vectors, and / or cells of the present disclosure. In some embodiments, the transgenic organism comprises the cells of the present disclosure.
[0344] The transgenic organism may be derived from any organism, including single-celled and multicellular organisms. Such organisms include animals, plants, fungi, bacteria, protists, fish, and the like. In some embodiments, the transgenic organism is a mammal or a plant. In some embodiments, the transgenic organism is a fungus (e.g., yeast), a bacterium, or a protist. In some embodiments, the transgenic organism is a fish. In some embodiments, the transgenic organism is a rodent (e.g., mouse, rat). In some embodiments, the transgenic organism is a rodent or a plant, optionally the rodent is a mouse. In some embodiments, the transgenic organism is a mammal or a plant, optionally the mammal is a rodent (e.g., mouse, rat), goat, sheep, chicken, llama, or rabbit.
[0345] Genetic modification of the germline of an organism to generate a transgenic organism can be accomplished by introducing any one of the nucleic acid vectors and viral vectors of the present disclosure using methods described herein and methods well known in the art. Pharmaceutical Compositions
[0346] In certain embodiments, provided herein is a pharmaceutical composition comprising any one of the nucleic acid vectors of the present disclosure, any one of the viral vectors of the present disclosure, and / or any one of the cells of the present disclosure.Any combination of nucleic acid vector, viral vector, and cell is contemplated herein, and such combination can provide a powerful therapeutic pharmaceutical composition.
[0347] Pharmaceutical compositions may further comprise carriers and / or diluents. As used herein, pharma-ceutically acceptable carriers are intended to include any and all solvents, dispersion media, antibacterial and antifungal agents, isotonic and absorption delaying agents, etc., that are compatible with pharmaceutical administration. The use of such media and agents for pharmaceutical active substances is well known in the art. Except where any conventional media or agent is incompatible with the active compound, its use in the composition is contemplated. To determine compatibility, various related factors, such as osmolality, viscosity, and / or basicity, can be considered. Auxiliary active compounds may also be incorporated into the composition.
[0348] The pharmaceutical composition of the present invention is formulated to be compatible with its intended route of administration.Examples of routes of administration include parenteral, for example, intravenous, intradermal, subcutaneous, oral, intranasal (e.g., inhalation), transdermal, transmucosal, intravascular, intracerebral, parenteral, intraperitoneal, epidural, intraspinal, intrasternal, intraarticular, intrasynovial, intratumoral, intrathecal, intraarterial, intracardiac, intramuscular, intrapulmonary, and rectal administration.In certain embodiments, direct injection into bone marrow is contemplated. Solutions or suspensions used for parenteral, intradermal or subcutaneous application may contain the following components: sterile diluents, such as water for injection, saline solution, fixed oils, polyethylene glycols, glycerin, propylene glycol or other synthetic solvents; antibacterial agents, such as benzyl alcohol or methylparabens; antioxidants, such as ascorbic acid or sodium bisulfite; chelating agents, such as ethylenediaminetetraacetic acid (EDTA); buffers, such as acetates, citrates or phosphates; and agents for adjusting isotonicity, such as sodium chloride or dextrose. pH can be adjusted with acids or bases, such as hydrochloric acid or sodium hydroxide. Parenteral preparations can be enclosed in glass or plastic ampoules, disposable syringes or multiple dose vials.
[0349] Pharmaceutical compositions suitable for injectable use include sterile aqueous solutions (if water soluble) or dispersions, and sterile powders for the extemporaneous preparation of sterile injectable solutions or dispersions. For example, Ringer's solution and lactated Ringer's solution are approved by the USP for the formulation of IV therapeutics, and these solutions are used in some embodiments. In certain embodiments, the compatibility of the excipients and vectors to preserve biological activity is established according to suitable methods. Suitable carriers for intravenous administration or injection into bone marrow include physiological saline, bacteriostatic water, CremoPhor EL™ (BASF, Parsippany, NJ) or phosphate buffered saline (PBS). In all cases, the composition must be sterile and fluid to the extent that easy syringability exists. It must be stable under the conditions of manufacture and storage, and must be preserved against the contaminating action of microorganisms, such as bacteria and fungi. The carrier can be a solvent or dispersion medium, for example, containing water, ethanol, polyol (e.g., glycerol, propylene glycol, and liquid polyethylene glycol, etc.), and suitable mixtures thereof. Proper fluidity can be maintained, for example, by the use of a coating agent such as lecithin, by maintaining the required particle size in the case of dispersions, and by the use of surfactants. Prevention of microbial action can be achieved by various antibacterial and antifungal agents, such as parabens, chlorobutanol, phenol, ascorbic acid, thimerosal, etc., to the extent that they do not affect the integrity / activity of the virus composition described herein. In many cases, it is preferable to include isotonic agents, such as sugars, polyalcohols, such as mannitol, sorbitol, sodium chloride, in the composition.
[0350] Sterile injectable solutions can be prepared by incorporating the active compound in the required amount in a suitable solvent with one or a combination of ingredients enumerated above, as required, followed by filtered sterilization. Generally, dispersions are prepared by incorporating the active compound into a sterile vehicle that contains a basic dispersion medium and the other required ingredients from those enumerated above.
[0351] For administration by inhalation, the viral vectors or nucleic acid vectors described herein are delivered in the form of an aerosol spray from a pressured container or dispenser which contains a suitable propellant, e.g., a gas such as carbon dioxide, or a nebulizer.
[0352] Systemic administration can also be achieved by transmucosal means.For transmucosal administration, a penetrant suitable for permeating the barrier is used in the formulation.Such penetrants are generally known in the art, and include, for example, for transmucosal administration, surfactants, bile salts, and fusidic acid derivatives.Transmucosal administration can be achieved by using nasal sprays or suppositories. Delivery of Nucleic Acid Vectors
[0353] Various techniques and methods for delivering nucleic acids to cells are known in the art and are encompassed for use in delivering the nucleic acid vectors described herein, including non-viral vectors that contain a portion of GSH, or nucleic acid vectors that contain 5' and 3' GSH-specific homology arms. For example, the nucleic acid can be formulated into lipid nanoparticles (LNPs), lipidoids, liposomes, lipid nanoparticles, lipoplexes, or core-shell nanoparticles. Typically, LNPs are composed of a nucleic acid molecule, one or more ionizable or cationic lipids (or salts thereof), one or more nonionic or neutral lipids (e.g., phospholipids), a molecule that prevents aggregation (e.g., PEG or PEG-lipid conjugates), and optionally a sterol (e.g., cholesterol). Exemplary lipid nanoparticles and methods for preparing the same are described, for example, in WO2015 / 074085, WO2016081029, WO2015 / 199952, WO2017 / 117528, WO2017 / 075531, WO2017 / 004143, WO2012 / 040184, WO2012 / 061259, WO2011 / 149733, WO 2013 / 158579, WO2014 / 130607, WO2011 / 022460, WO2013 / 148541, WO2013 / 116126, WO2011 / 1531 20, WO2012 / 044638, WO2012 / 054365, WO2008 / 042973, WO2010 / 129709, WO2010 / 144740, WO2012 / 099755, WO2013 / 049328, WO2013 / 086322, WO2013 / 086354, WO2013 / 086373, WO2014 / 008334, W O2011 / 075656, WO2011 / 071860, WO2009 / 132131, WO2010 / 088537, WO2010 / 054401, WO2010 / 054 384, WO2010 / 054406, WO2010 / 054405, WO2010 / 048536, WO2009 / 082607, WO2012 / 016184, WO201 4 / 152211, WO2017 / 049074, WO1996 / 040964, WO1999 / 018933, WO2009 / 086558, WO2010 / 129687,WO2010 / 147992, WO2010 / 042877, WO2009 / 108235, WO2014 / 081887, W02005 / 12046 1, WO2011 / 000106, WO2011 / 000107, WO2015 / 011633, WO2005 / 120152, WO2011 / 1417 05, WO2016 / 197133, WO2015 / 011633, WO2013 / 126803, WO2012 / 000104, WO2011 / 141 705, WO2006 / 007712, WO2011 / 038160, WO2005 / 121348, WO2005 / 120152, WO2011 / 06 6651, WO2009 / 127060, WO2011 / 141704, WO2006 / 074546, WO2005 / 121348, WO2006 / 069782, WO2009 / 027337, WO2012 / 030901, WO2012 / 031043, WO2012 / 031046, WO2013 / 006825, WO2013 / 033563, WO2013 / 040429, WO2014 / 043544, WO2016 / 130963, WO2017 / 181026, and WO2013 / 089151, the contents of all of which are incorporated herein by reference in their entirety. In some embodiments, the lipid nanoparticles, in addition to the nucleic acid, comprise lipids in the following molar ratio: 50% cationic lipid, 10% non-ionic lipid (e.g., phospholipids, such as distearoylphosphatidylcholine (DSPC)), 38.5% cholesterol, and 1.5% PEG-lipid (e.g., 2-[2-(w-methoxy(polyethylene glycol 2000)ethoxy]-N,N-ditetradecylacetamide (PEG2000-DMA)).
[0354] Another method for delivering nucleic acid to cells is by conjugating the nucleic acid with a ligand that is internalized by the cell.For example, the ligand can bind a receptor on the cell surface and be internalized by endocytosis.The ligand can be covalently linked to the nucleotide in the nucleic acid.Exemplary conjugates for delivering nucleic acid to cells are described in, for example, WO2015 / 006740, WO2014 / 025805, WO2012 / 037254, WO2009 / 082606, WO2009 / 073809, WO2009 / 018332, WO2006 / 112872, WO2004 / 090108, WO2004 / 091515, WO2017 / 177326, the contents of all of which are incorporated herein by reference in their entirety.
[0355] Nucleic acid can also be delivered to cells by electroporation.Generally, electroporation uses pulsed current to increase cell permeability, thereby allowing nucleic acid to cross plasma membrane.Electroporation techniques are well known in the art and are used to deliver nucleic acid in vivo and clinically.See, for example, Andre et al., Curr Gene Ther. 2010 10:267-280;Chiarella et al, CurrGene Ther. 2010 10:281-286;Hojman, Curr Gene Ther. 2010 10: 128-138, the contents of all of which are incorporated herein by reference in their entirety. Electroporation devices are sold worldwide by many companies, including but not limited to BTX® Instruments (Holliston, MA) (e.g., AgilePulse In Vivo System) and Inovio (Blue Bell, PA) (e.g., Inovio SP-5P intramuscular delivery device or CELLECTRA® 3000 intradermal delivery device). Electroporation can be used after, before and / or during administration of the nucleic acid vector. Further exemplary methods and devices for delivering nucleic acids using electroporation are described, for example, in U.S. Patent Nos. 5,273,525, 6,520,950, 6,654,636 and 6,972,013, the contents of all of which are incorporated herein by reference in their entirety.
[0356] Nucleic acids can also be delivered to cells by transfection. Useful transfection methods include, but are not limited to, lipid-mediated transfection, cationic polymer-mediated transfection, or calcium phosphate precipitation. Transfection reagents are well known in the art and include TurboFect Transfection Reagent (Thermo Fisher Scientific), Pro-Ject Reagent (Thermo Fisher Scientific), TRANSPASS™ P Protein Transfection Reagent (New England Biolabs), CHARIOT™ Protein Delivery Reagent (Active Motif), PROTEOJUICE™ Protein Transfection Reagent (EMD Millipore), 293fectin, LIPOFECTAMINE™ 2000, LIPOFECTAMINE™ 3000 (Thermo Fisher Scientific), FIPOFECTAMINE™ (Thermo Fisher Scientific), FIPOFECTIN™ (Thermo Fisher Scientific), DMRIE-C, CEFFFECTIN™ (Thermo Fisher Scientific), OFGOFECTAMINE™ (Thermo Fisher Scientific), FIPOFECTIN ... Scientific), FIPOFECTACE(TM), FUGENE(TM)(Roche, Basel, Switzerland), FUGENE(TM) HD(Roche), TRANSFECTAM(TM)(Transfectam, Promega, Madison, Wis.), TFX-10(TM)(Pr omega), TFX-20(TM)(Promega), TFX-50(TM)(Promega), TRANSFECTIN(TM)(BioRad, Hercules, Calif.), SIFENTFECT(TM)(Bio-Rad), Effectene(TM)(Qiagen, Valencia, Calif.).), DC-chol (Avanti Polar Lipids), GENEPORTER™ (Gene Therapy Systems, San Diego, Calif.), DHARMAFECT 1™ (Dharmacon, Lafayette, Colo), DHARMAFECT 2™ (Dharmacon), DHARMAFECT 3™ (Dharmacon), DHARMAFECT 4™ (Dharmacon), ESCORT™ III (Sigma, St. Louis, Mo.), and ESCORT™ IV (Sigma Chemical Co.). Nucleic acids can also be delivered to cells by microfluidics methods known to those of skill in the art.
[0357] Non-viral methods for delivery of nucleic acids in vivo or ex vivo include electroporation, lipofection (see U.S. Pat. Nos. 5,049,386, 4,946,787, and commercially available reagents, e.g., Transfectam™ and Lipofectin™), microinjection, biolistic techniques, virosomes, liposomes (see, e.g., Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Ther. 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994); Gao et al., Gene Therapy 2:710-722 (1995); Ahmadet al., Cancer Res. 52:4817-4820 (1992); see U.S. Pat. Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787), immunoliposomes, polycations or lipid nucleic acid conjugates, naked DNA, artificial virions, viral vector systems (e.g., retroviral, lentiviral, adenoviral, adeno-associated, vaccinia and herpes simplex viral vectors as described in WO 2007 / 014275), and drug-enhanced DNA uptake. Sonoporation, for example using the Sonitron 2000 system (Rich-Mar), can also be used for delivery of nucleic acids.
[0358] A vector (e.g., retrovirus, adenovirus, liposome, etc.) containing the nucleic acid described herein can also be administered directly to an organism for transduction of cells in vivo. Alternatively, naked DNA can be administered. Administration is by any of the routes normally used to introduce molecules into blood or ultimate contact with tissue, including, but not limited to, injection, infusion, topical application, and electroporation. Suitable methods of administering such nucleic acids are available and well known to those skilled in the art, and more than one route can be used to administer a particular composition, although a particular route can often result in a more rapid and effective response than another route.
[0359] Methods for the introduction of the nucleic acid vector compositions disclosed herein into hematopoietic stem cells are disclosed, for example, in US Pat. No. 5,928,638.
[0360] The nucleic acid vector compositions disclosed herein can be used for ex vivo cell transfection for diagnosis, research, or gene therapy (e.g., by reinjecting the transfected cells into the host organism). In some embodiments, cells are isolated from a subject organism, transfected with the nucleic acid vector compositions disclosed herein, and reinjected back into the subject organism (e.g., a patient or subject). Various cell types suitable for ex vivo transfection are well known to those of skill in the art (e.g., see Freshney et al., Culture of Animal Cells, A Manual of Basic Technique (3rd ed. 1994) and references cited therein for a discussion of how to isolate and culture cells from a patient).
[0361] In some embodiments, stem cells are used in ex vivo procedures for cell transfection and gene therapy.The advantage of using stem cells is that they can be differentiated in vitro into other cell types, or introduced into mammals (e.g., cell donors) where they engraft into bone marrow.Methods are known for in vitro differentiation of CD34+ cells into clinically important immune cell types using cytokines such as GM-CSF, IFN-γ, and TNF-α (see Inaba et al., J. Exp. Med. 176: 1693-1702 (1992)).
[0362] Stem cells are isolated for transduction and differentiation using known methods. For example, stem cells are isolated from bone marrow cells by panning bone marrow cells with antibodies that bind undesirable cells, such as CD4+ and CD8+ (T cells), CD45+ (panb cells), GR-1 (granulocytes) and lad (differentiated antigen-presenting cells) (see Inaba et al., J. Exp. Med. 176:1693-1702 (1992)). In some embodiments, the cells to be used are oocytes. In other embodiments, cells from model organisms may be used. These may include cells from Xenopus, insect cells (e.g., drosophila) and nematode cells. kit
[0363] In certain aspects, provided herein are kits comprising any one of the nucleic acid vectors of the present disclosure, any one of the viral vectors of the present disclosure, any one of the cells of the present disclosure, and / or any one of the pharmaceutical compositions of the present disclosure.
[0364] In some embodiments, a kit for insertion of a gene or nucleic acid sequence into a target GSH identified according to the methods disclosed herein, and a primer set for determining integration of the gene or nucleic acid sequence.
[0365] In some embodiments, the kit comprises (a) a vector composition as described herein and a primer pair for determining integration by homologous recombination of a nucleic acid located between the restriction site located between the 3'GSH-specific homology arm and the 5'GSH-specific homology arm of the vector. In some embodiments, the kit comprises a primer pair spanning an integration site, comprising at least one GSH 5' primer and at least one GSH 3' primer, wherein GSH is identified according to the method disclosed herein, and wherein at least one GSH 5' primer binds to a region of GSH upstream of the integration site and at least one GSH 3' primer binds to a region of GSH downstream of the integration site. Such a primer pair can function to serve as a negative control, producing a short PCR product when integration has not occurred and producing no PCR product incorporating the inserted nucleic acid or a long PCR product incorporating it when nucleic acid insertion has occurred.
[0366] In some embodiments, the kit can include (a) GSH-specific single guide and RNA guide nucleic acid sequences contained in one or more GSH vectors, and (b) a GSH knock-in vector that includes a GSH vector, where one or more of the sequences of (a) or (b) are contained in a vector described herein. In some embodiments, the GSH vector is a GSH-CRISPR-Cas vector or other GSH-gene editing vector, such as one that includes a gene editing gene described herein. In some embodiments, the GSH CRISPR-Cas vector includes a GSH-sgRNA nucleic acid sequence and a Cas9 nucleic acid sequence.
[0367] In other embodiments, the kit can further comprise a GSH knock-in donor vector comprising a GSH 5' homology arm and a GSH 3' homology arm, wherein the GSH 5' homology arm and the GSH 3' homology arm are at least, about or at most 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109 ... %, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% complementary, and the GSH 5' and 3' homology arms enable (i.e., guide) insertion, by homologous recombination, of a nucleic acid sequence located between the GSH 5' homology arm and the GSH 3' homology arm into a locus located within the genomic safe harbor.As an illustrative example, in some embodiments, the GSH Cas9 knock-in donor vector is a SYNTX-GSH1 Cas9 knock-in donor vector comprising a SYNTX-GSH1 5' homology arm and a SYNTX-GSH1 3' homology arm, wherein the SYNTX-GSH1 5' homology arm and the SYNTX-GSH1 3' homology arm are at least, about or at most 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76% , 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% or 100% complementary, and the SYNTX-GSH1 5' and 3' homology arms guide insertion of a nucleic acid located between the GSH 5' homology arm and the GSH 3' homology arm into a locus within the SYNTX-GSH1 genomic safe harbor by homologous recombination.
[0368] In some embodiments, the kit comprises a GSH vector that is a GSH Cas9 knock-in donor vector.
[0369] In some embodiments, the kit further comprises at least one GSH 5' primer and at least one GSH 3' primer, wherein the at least one GSH 5' primer is at least, about or at most 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 143%, 144%, 145%, 146%, 147%, 148%, 14 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% complementary to at least one GSH The 3' primer should be at least, about, or up to 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 143%, 144%, 145%, 146%, 147%, 148%, 149%, 150%, 151%, 152%, 153%, 154%, 155%, , 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% complementary.
[0370] In some embodiments, the kit can include two primer pairs, each primer pair functions as a positive control.For example, in some embodiments, the kit includes (a) at least two GSH 5' primers, including a forward GSH 5' primer that binds to the region of GSH upstream of the integration site and a reverse GSH 5' primer that binds to the sequence of the nucleic acid inserted into the integration site in the GSH sequence; and (b) at least two GSH 3' primers, including a forward GSH 3' primer that binds to the sequence located at the 3' end of the nucleic acid inserted into the integration site in the GSH sequence and a reverse GSH 3' primer that binds to the region of GSH downstream of the integration site.In such an embodiment, the primer pair functions as a positive control, and can only generate PCR products when integration occurs, and does not generate PCR products when integration does not occur.
[0371] In some embodiments, the kit comprises a forward GSH 5' primer that is at least 80% complementary to a region of GSH upstream of the integration site and a forward GSH 5' primer that is at least, about or up to 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 109, 102, 103, 104, 105, 106, 107, 108, 109 ... and a reverse GSH 5' primer that is 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% or 100% complementary to the GSH 5' primer.
[0372] In some embodiments, the kit comprises a GSH sequence that is inserted into an integration site and is at least, about, or up to 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 109, 109, 102, 103, 104, 105, 106, 107, 108, 109 ... 3%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% complementary and a forward GSH 3' primer that is at least, about or up to 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 143%, 144%, 145%, 146%, 147%, 148%, 149%, 150%, 151%, 152%, 153%, 154%, 1 and a reverse GSH 3' primer that is 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% or 100% complementary to the GSH 3' primer.
[0373] In some embodiments, the kit comprises any one of the nucleic acid vectors described herein.
[0374] In some embodiments, the kit comprises any one of the viral vectors described herein.
[0375] In some embodiments, the kit comprises any one of the cells described herein.
[0376] In some embodiments, the kit comprises any one of the pharmaceutical compositions disclosed herein.
[0377] In some embodiments, the kit comprises any combination of a nucleic acid vector, a viral vector, a cell, and a pharmaceutical composition.
[0378] The nucleic acids, viral vectors, cells, and / or pharmaceutical compositions may be packaged in a suitable container. The kits may include additional components to facilitate the particular use for which the kit is designed. In addition, the kits encompassed by the present disclosure may also include instructional materials that disclose or describe the use of the kit. Use of GSH in Biologics Manufacturing
[0379] Provided herein is the use of the GSH locus identified herein for preparing biologics.Notably, the GSH locus identified herein is particularly useful in that it allows stable integration of the gene expressing biologics into cells, thereby enabling large-scale production of biologics.
[0380] Protein-based therapeutics, including antibodies, peptides and recombinant proteins, represent the majority of new products in development by the pharmaceutical industry (Ho & Chien 2014, PMID: 24186148). Such products are produced in a variety of platforms, including non-mammalian (bacteria, yeast, plant and insect cells), and mammalian systems (rodent and human-derived cells). Mammalian expression systems are usually the preferred platform for producing biopharmaceuticals, as these cells or cell lines can produce large complex proteins with post-translational modifications similar to those found in humans. Among the various mammalian cell systems used for biopharmaceutical manufacturing, human-derived cell lines are attractive as substrates for therapeutic glycoprotein production, as their glycosylation mechanisms eliminate the risk of immunogenicity seen in by-products derived from different cells, such as rodent-derived cell lines (e.g., CHO, BHK1, NS0, Sp2 / 0). These non-human cell lines have different post-translational modification pathways that can generate immunogenic glycans such as galactose-α1,3-galactose (α-galactose) and N-glycolylneuraminic acid (NGNA) (Butler and Spearman 2014, PMID: 25005678). As circulating antibodies against both of these N-glycans are prevalent in the human population, such non-human cell lines need to be screened for clones with acceptable glycosylation profiles (Dumont, J. et al PMID: 26383226).
[0381] Chinese hamster ovary (CHO) cells are aneuploid cells that are commonly used for the production of therapeutic proteins. CHO cell chromosomes have structural abnormalities and undergo structural and numerical changes during cell proliferation. During proliferation, they continuously undergo genomic changes, such as mutations, deletions, duplications and other structural changes, due to errors in DNA replication and repair and chromosome segregation errors. As a result, these cells, along with other commonly used cell lines, such as HEK293, MDCK and Vero cells, have a wide distribution of chromosome numbers. Thus, these cell lines are associated with heterogeneity in the form of genomic and epigenomic diversity or changes to cell phenotype or productivity.
[0382] Such heterogeneity, which can affect the production of biologics, is exacerbated by random integration of transgenes expressing biologics. Current processes for human cell line generation are based on random integration of genes of interest into the genome, resulting in recombinant clones with high genomic and phenotypic variability, referred to as clonal diversity. This variability impacts the predictive value of the product and limits the streamlining of processes and the achievement of cost-effective production of therapeutic glycoproteins.
[0383] In addition, expression of randomly integrated transgenes is unpredictable and prone to become unstable over time due to epigenetic effects. Furthermore, random integration often results in multiple integrants per cell, which may result in disruption or activation of host cell genes. The biopharmaceutical industry has invested substantial resources to improve the yield and quality of recombinant proteins, especially monoclonal antibodies. This process often begins with the selection of high-yielding cell clones from a heterogeneous population of stable cells. Clonal diversity can be explained, in part, by the plasticity and epigenetic imprinting of the host cell genome. This is reflected in repeated chromosomal rearrangements, high mutation rates and genomic instability (Vcelar et al. 2018 PMID: 29328552), as well as the suppression of expression of non-essential genes that negatively affect transgene expression. Genomic diversity also occurs due to random integration of vectors that can be inserted in multiple copies at different genomic loci, known as "position effects," highlighting the importance of the surrounding genomic environment (Wilson, C. et al 1990 PMID: 2275824). In addition, epigenetic regulation also affects the expression of transgenes and can be influenced by environmental conditions such as oxygen and nutrient levels or by the accumulation of toxic by-products during the production process. Clonal heterogeneity requires time-consuming, labor-intensive screening to find cell lines with the desired performance. The clonal selection process can involve single-cell cloning using high-throughput screening, which is an inherently random process.
[0384] In contrast, the GSH locus can be used reliably for predictable expression. First, it eliminates genomic heterogeneity induced by random integration of the transgene. Such is mediated by high-fidelity homologous recombination and / or nuclease-initiated recombination (e.g., CRISPR). Second, the transgene is inserted at a genomic location that allows not only stable integration but also stable expression. There is no concern that the transgene will disrupt important genes in the cells selected for biologics production. This stable expression is also predictable. Because GSH provides a known transcriptional environment, there is no "position effect" or silencing of the transgene, for example, by a nearby repressive (e.g., heterochromatic) environment. As a result, transgene insertion at the GSH locus does not affect cell cycle homeostasis, allowing for high bioproduct yields.
[0385] Thus, provided herein is a method of producing a biologic comprising: (a) culturing (i) a cell comprising any one of the nucleic acid vectors described herein, (ii) a cell comprising any one of the viral vectors described herein, or (iii) any one of the cells described herein and recovering the expressed biologic; or (b) recovering the expressed biologic from any one of the transgenic organisms contemplated herein.
[0386] In some embodiments, the biologic is an antigen-binding protein. In some embodiments, the biologic is an antibody or an antigen-binding fragment thereof, optionally the antibody or antigen-binding fragment thereof is selected from an antibody, Fv, F(ab')2, Fab', dsFv, scFv, sc(Fv)2, half-antibody-scFv, tandem scFv, Fab / scFv-Fc, tandem Fab', single chain diabody, tandem diabody (TandAb), Fab / scFv-Fc, scFv-Fc, heterodimeric IgG (CrossMab), DART, and diabody.
[0387] In some embodiments, the biologic specifically binds to TNFα, CD20, a cytokine (e.g., IL-1, IL-6, BLyS, APRIL, IFN-gamma, etc.), Her2, RANKL, IL-6R, GM-CSF, or CCR5.
[0388] In some embodiments, the biologic is selected from adalimumab, etanercept, infliximab, certolizumab, golimumab, anakinra, rituximab, abatacept, tocilizumab, natalizumab, canakinumab, atacicept, belimumab, ocrelizumab, ofatumumab, fontolizumab, trastuzumab, denosumab, sarilumab, lenzilumab, gimcilumab, siltuximab, leronlimab, and antigen-binding fragments thereof.
[0389] In some embodiments, the biologic is a therapeutic protein, and optionally, the therapeutic protein is insulin. Antigen-binding proteins
[0390] The antigen-binding proteins of the present disclosure may take any one of the many forms of antigen-binding proteins known in the art. In various embodiments, the antigen-binding proteins of the present disclosure take the form of an antibody or an antigen-binding antibody fragment, an engineered antibody protein product (e.g., including a fragment of an antibody), a ligand- or receptor-binding protein or fragment thereof, or a fusion protein.
[0391] As used herein, the term "antibody" refers to a protein having a conventional immunoglobulin format, comprising heavy and light chains, and comprising variable and constant regions. For example, an antibody may be an IgG, which is an identical "Y-shaped" structure of two pairs of polypeptide chains, each pair having one "light" chain (typically having a molecular weight of about 25 kDa) and one "heavy" chain (typically having a molecular weight of about 50-70 kDa). An antibody has a variable region and a constant region. The variable region in the IgG format is generally about 100-110 or more amino acids, contains three complementarity determining regions (CDRs), and is primarily responsible for antigen recognition and differs substantially among other antibodies that bind to different antigens. The constant region allows the antibody to recruit cells and molecules of the immune system. The variable region is made up of the N-terminal region of each light and heavy chain, while the constant region is made up of the C-terminal portion of each heavy and light chain. (Janeway et al., "Structure of the Antibody Molecule and the Immunoglobulin Genes", Immunobiology: The Immune System in Health and Disease, 4 th ed. Elsevier Science Ltd. / Garland Publishing, (1999)).
[0392] The general structure and characteristics of antibody CDR are described in the art. Briefly, in antibody scaffold, CDR is embedded in the framework of heavy and light chain variable region, where they constitute the region mainly responsible for antigen binding and recognition. A variable region typically comprises at least three heavy or light chain CDRs (Kabat et al., 1991, Sequences of Proteins of Immunological Interest, Public Health Service NIH, Bethesda, Md.; see also Chothia and Lesk, 1987, J. Mol. Biol. 196:901-917; Chothia et al., 1989, Nature 342: 877-883), within framework regions (designated framework regions 1-4, FR1, FR2, FR3, and FR4 by Kabat et al., 1991; see also Chothia and Lesk, 1987, supra).
[0393] CDR refers to complementarity determining regions (CDRs), three of which constitute the binding properties of the light chain variable region (CDR-L1, CDR-L2 and CDR-L3) and three of which constitute the binding properties of the heavy chain variable region (CDR-H1, CDR-H2 and CDR-H3). CDRs contribute to the functional activity of the antibody molecule and are separated by amino acid sequences that constitute the scaffolding or framework regions. The exact defining CDR boundaries and lengths follow different classification and numbering systems. CDRs are therefore referred to by Kabat, Chothia, contact or any other boundary definition. Despite the different boundaries, each of these systems overlaps to some extent in what constitutes the so-called "hypervariable regions" within the variable sequences. Thus, the CDR definitions according to these systems may differ in length and boundary area with respect to the adjacent framework regions. See, e.g., Kabat, Chothia, and / or MacCallum et al. (Kabat et al., in "Sequences of Proteins of Immunological Interest," 5th Edition, US Department of Health and Human Services, 1992; Chothia et al. (1987) J. Mol. Biol. 196, 901; and MacCallum et al., J. Mol. Biol. (1996)262, 732, each of which is incorporated by reference herein in its entirety.
[0394] The antibody may include any constant region known in the art. Human light chains are classified as kappa and lambda light chains. Heavy chains are classified as mu, delta, gamma, alpha, or epsilon, defining the antibody's isotype as IgM, IgD, IgG, IgA, and IgE, respectively. IgG has several subclasses, including, but not limited to, IgG1, IgG2, IgG3, and IgG4. IgM has subclasses, including, but not limited to, IgM1 and IgM2. The embodiments of the present disclosure include all such classes or isotypes of antibodies. The light chain constant region may be, for example, a kappa- or lambda-type light chain constant region, for example, a human kappa- or lambda-type light chain constant region. The heavy chain constant region may be, for example, an alpha-, delta-, epsilon-, gamma-, or mu-type heavy chain constant region, for example, a human alpha-, delta-, epsilon-, gamma-, or mu-type heavy chain constant region. Thus, in various embodiments, the antibody is of isotype IgA, IgD, IgE, IgG, or IgM, including any one of IgG1, IgG2, IgG3, or IgG4. In various aspects, the antibody comprises a constant region that comprises one or more amino acid modifications compared to its naturally occurring counterpart to improve half-life / stability or to make the antibody more suitable for expression / manufacturability. In various cases, the antibody comprises a constant region in which the C-terminal Lys residue present in the naturally occurring counterpart has been removed or truncated.
[0395] The antibody may be a monoclonal antibody. In some embodiments, the antibody comprises a sequence that is substantially similar to a naturally occurring antibody produced by a mammal, such as a mouse, rabbit, goat, horse, chicken, hamster, human, etc. In this regard, the antibody may be considered a mammalian antibody, such as a mouse antibody, a rabbit antibody, a goat antibody, a horse antibody, a chicken antibody, a hamster antibody, a human antibody, etc. In certain aspects, the antibody binding protein is an antibody, such as a human antibody. In certain aspects, the antibody binding protein is a chimeric antibody or a humanized antibody. The term "chimeric antibody" refers to an antibody that contains domains from two or more different antibodies. A chimeric antibody may, for example, contain a constant domain from one species and a variable domain from a second species, or more commonly, may contain stretches of amino acid sequences from at least two species. A chimeric antibody may also contain domains of two or more different antibodies within the same species. The term "humanized" when used in reference to an antibody refers to an antibody that has at least a CDR region from a non-human source and that has been engineered to have a structure and immunological function more similar to a true human antibody than the original source antibody. For example, humanizing can include grafting CDRs from a non-human antibody, such as a mouse antibody, onto a human antibody. Humanizing can also include selecting amino acid substitutions to make the non-human sequence more similar to a human sequence. Information, including sequence information, on human antibody heavy and light chain constant regions is publicly available through the Uniprot database as well as other databases well known to those skilled in the art of antibody engineering and production. For example, the IgG2 constant region is available from the Uniprot database as Uniprot number P01859, which is incorporated herein by reference.
[0396] Antibodies can be cleaved into fragments by enzymes, such as papain and pepsin. Papain cleaves an antibody to produce two Fab' fragments and a single Fc fragment. Pepsin cleaves an antibody to produce an F(ab') fragment. 2In various aspects of the disclosure, the antigen-binding proteins of the disclosure are antigen-binding fragments of antibodies (also known as antigen-binding antibody fragments, antigen-binding fragments, antigen-binding portions). In various instances, the antigen-binding antibody fragments are Fab' or F(ab') fragments. 2 It is a fragment.
[0397] Antibody architecture has been exploited to generate a variety of alternative antibody formats, referred to herein as "antibody protein products", with a range of valencies (n) from monomers (n=1), to dimers (n=2), to trimers (n=3), to tetramers (n=4), and potentially higher, spanning a molecular weight range of at least about 12-150 kDa. Antibody protein products include those based on the complete antibody structure, as well as those that mimic antibody fragments that retain full antigen-binding capacity, such as scFv, Fab, and VHH / VH (discussed below). The smallest antigen-binding fragment that retains its complete antigen-binding site is the Fv fragment, consisting entirely of the variable (V) region. Soluble, flexible amino acid peptide linkers are used to connect the V region to the scFv (single chain fragment variable) fragment for stabilization of the molecule, or constant (C) domains are added to the V region to generate Fab' fragments. Both scFv and Fab' fragments can be easily produced in host cells, e.g., prokaryotic host cells. Other antibody protein products include disulfide bond stabilized scFv (ds-scFv), single chain Fab' (scFab'), and dimeric and multimeric antibody formats such as diabodies, triabodies and tetrabodies, or minibodies (miniAbs) including different formats consisting of scFv linked to oligomerization domains. The smallest fragments are the VHH / VH of camelid heavy chain Abs, and single domain Abs (sdAbs). The most frequently used building block for generating new antibody formats is the single chain variable (V) domain antibody fragment (scFv), which contains the V domains (VH and VL domains) from heavy and light chains linked by a peptide linker of about 15 amino acid residues. Peptibodies or peptide-Fc fusions are yet another antibody protein product. The structure of a peptibody consists of a bioactive peptide grafted onto the Fc domain. Peptibodies are well described in the art, see, e.g., Shimamoto et al., mAbs 4(5): 586-591 (2012).
[0398] Other antibody protein products include single chain antibodies (SCAs); diabodies; triabodies; tetrabodies, and the like.
[0399] In various embodiments, the antigen binding proteins of the disclosure comprise, consist essentially of, or consist of any one of these antibody protein products.
[0400] In various aspects, the antigen binding proteins of the disclosure comprise, consist essentially of, or consist of any one of: scFv, Fab', F(ab')2, VHH / VH, Fv fragment, ds-scFv, scFab', half antibody-scFv, heterodimeric Fab / scFv-Fc, heterodimeric scFv-Fc, heterodimeric IgG (CrossMab), tandem scFv, tandem biparatopic scFv, Fab / scFv-Fc, tandem Fab', single chain diabody, dimeric antibody, multimeric antibody (e.g., diabody, triabody, tetrabody), miniAb, peptibody VHH / VH of camelid heavy chain antibody, sdAb, diabody (single chain diabody, homodimeric diabody, heterodimeric diabody, tandem diabody (TandAb), self-dimerizing diabody), triabody, tetrabody. It will be understood by those skilled in the art that any bispecific antigen-binding protein format can be used to generate a biparatopic antigen-binding protein format. In some embodiments, the antigen-binding protein is a dual affinity retargeting antibody (DART). In some embodiments, the antigen-binding protein is a bispecific T cell engager (BiTE). Exemplary Biologics
[0401] Exemplary antigen binding proteins include, for example, antibodies that bind to CD40, Toll-like receptors (TLRs), OX40, GITR, CD27, or to 4-1BB, T cell bispecific antibodies, anti-IL-2 receptor antibodies, anti-CD3 antibodies, OKT3 (muromonab), otelixizumab, teplizumab, visilizumab, anti-CD4 antibodies, clenoliximab, keliximab, zanolimumab, anti-CD11 antibodies, efalizumab, anti-CD18 antibodies, erlizumab, rovelizumab, anti-CD20 antibodies, body, afutuzumab, ocrelizumab, ofatumumab, pascolizumab, rituximab, anti-CD23 antibody, rumiliximab, anti-CD40 antibody, teneliximab, toralizumab, anti-CD40L antibody, ruplizumab, anti-CD62L antibody, acelizumab, anti-CD80 antibody, galiximab, anti-CD147 antibody, gavilimomab, B-lymphocyte stimulator (BLyS) inhibitor antibody, belimumab, CTLA4-Ig fusion protein, abatacept, belatacept, anti-CTLA4 antibody, ipilimumab , tremelimumab, anti-eotaxin 1 antibody, bertilimumab, anti-a4-integrin antibody, natalizumab, anti-IL-6R antibody, tocilizumab, anti-LFA-1 antibody, odulimomab, anti-CD25 antibody, basiliximab, daclizumab, inolimomab, anti-CD5 antibody, zolimomab, anti-CD2 antibody, siplizumab, nerelimomab, faralimomab, atlizumab, atolimumab, cedelizumab, dorlimomab alitox, dorlixizumab, fontolizumab, gantenerumab These include , gomilikimab, lebrilizumab, maslimomab, morolimumab, pexelizumab, reslizumab, rovelizumab, talizumab, terimomab alitox, bapaliximab, beparimomab, aflibercept, alefacept, rilonacept, IL-1 receptor antagonists, anakinra, anti-IL-5 antibodies, mepolizumab, IgE inhibitors, omalizumab, talizumab, IL12 inhibitors, IL23 inhibitors, and ustekinumab.
[0402] Exemplary biologics may include any one of the therapeutic proteins described herein or fragments thereof or those known in the art. For example, the biologic may be a hemoglobin gene (HBA1, HBA2, HBB, HBG1, HBG2, HBD, HBE1, and / or HBZ), alpha-hemoglobin stabilizing protein (AHSP), coagulation factor VIII, coagulation factor IX, von Willebrand factor, dystrophin or truncated dystrophin, microdystrophin, utrophin or truncated utrophin, microutrophin, usherin (USH2A), GBA1, preproinsulin, insulin, GIP, GLP-1, CEP290, ATPB1, ATPB11, ABCB4, CPS1, ATP7B, KRT5, KRT14, PLEC1, Col7A1, ITGB4, ITGA6, LAMA3, LAMB3, LAMC2, KIND1, INS, F8, or a fragment thereof (e.g., a fragment encoding a B domain deleted polypeptide ... SQ, p-VIII)), IRGM, NOD2, ATG2B, ATG9, ATG5, ATG7, ATG16L1, BECN1, EI24 / PIG8, TECPR2, WDR45 / WIP14, CHMP2B, CHMP4B, Dynein, EPG5, HspB8, LAMP2, LC3b The polypeptide may comprise a recombinant polypeptide or fragment thereof selected from UVRAG, VCP / p97, ZFYVE26, PARK2 / Parkin, PARK6 / PINK1, SQSTM1 / p62, SMURF, AMPK, ULK1, RPE65, CHM, RPGR, PDE6B, CNGA3, GUCY2D, RS1, ABCA4, MYO7A, HFE, hepcidin, genes encoding soluble forms (e.g., of the TNFα receptor, IL-6 receptor, IL-12 receptor or IL-1β receptor), and the cystic fibrosis transmembrane conductance regulator (CFTR).
[0403] A complete list of biologics approved by the FDA is available on the World Wide Web at fda.gov / vaccines-blood-biologics / development-approval-process-cber / biological-approvals-year and in Purple Book (on the World Wide Web at purplebooksearch.fda.gov / ). As used herein, biologics encompass biosimilars. Manufacturing method
[0404] Methods of producing a biologic are also provided herein. In some embodiments, the methods include culturing a host cell comprising a nucleic acid comprising a nucleotide sequence encoding the biologic in a cell culture medium, and collecting the secreted biologic from the cell culture medium. The host cell can be any of the host cells described herein. In various aspects, the host cell is selected from the group consisting of CHO cells, NS0 cells, COS cells, VERO cells, and BHK cells. In various aspects, culturing the host cell includes culturing the host cell in a growth medium to support the growth and expansion of the host cell. In various aspects, the growth medium increases cell density, cell viability, and productivity in a timely manner. In various aspects, the growth medium includes amino acids, vitamins, inorganic salts, glucose, and serum as a source of growth factors, hormones, and attachment factors. In various aspects, the growth medium is a complete chemically defined medium consisting of amino acids, vitamins, trace elements, inorganic salts, lipids, and insulin or insulin-like growth factors. In addition to nutrients, the growth medium also helps maintain pH and osmolality. Several growth media are commercially available and described in the art, see, e.g., Arora, "Cell Culture Media: A Review" Mater Methods 3:175 (2013).
[0405] In various aspects, the method comprises culturing the host cells in a feed medium. In various aspects, the method comprises culturing in a fed-batch mode in a feed medium. Recombinant protein production methods are known in the art. See, for example, Li et al., "Cell culture processes for monoclonal antibody production" MAbs 2(5): 466-477 (2010).
[0406] The method of making a biologic may include one or more steps for purifying the protein from the cell culture or its supernatant, and preferably for recovering the purified protein. In various aspects, the method includes one or more chromatography steps, such as affinity chromatography (e.g., Protein A affinity chromatography, nickel resin for histidine (His) tags), ion exchange chromatography, hydrophobic interaction chromatography. In various aspects, the method includes purifying the protein using Protein A affinity chromatography resin.
[0407] In various embodiments, the method further comprises formulating the purified protein, etc., to provide a formulation comprising the purified protein. Such steps are described in Formulation and Process Development Strategies for Manufacturing, eds. Jameel and Hershenson, John Wiley & Sons, Inc. (Hoboken, NJ), 2010.
[0408] In various aspects, the biologic is a fusion protein. For example, the biologic can be an antigen-binding protein linked to a polypeptide (e.g., an Fc domain). Thus, the present disclosure further provides a method for producing a fusion protein. In various embodiments, the method includes culturing a host cell comprising a nucleic acid comprising a nucleotide sequence encoding the fusion protein described herein in a cell culture medium, and collecting the fusion protein from the cell culture medium. Use of GSH in viral vector production
[0409] Recombinant viral vectors (e.g., AAV vectors, retroviral vectors, lentiviral vectors, etc.) are important tools in therapy and research. For example, recombinant AAV vectors are clinically validated tools for in vivo gene transfer. Although the application of AAV vectors offers great potential for many genetic diseases, current vector production methods still have room for improvement to meet the demands of human testing as well as basic biology, toxicology, and efficacy preclinical studies, especially those involving certain genetic diseases that require large quantities of high-quality vectors. For example, gene therapy for muscular dystrophy requires systemic gene transfer in muscle, the largest organ in the body. Other genetic diseases that affect many people, such as sickle cell anemia or cystic fibrosis, require large-scale preparation of recombinant vectors.
[0410] One of the most used methods for AAV production is the human embryonic kidney-derived cell (HEK293) platform. The most widely used protocol for vector production is based on a helper virus-free transient transfection method using cis and trans components (vector plasmid and packaging plasmid along with helper genes isolated from adenovirus) in host cells such as HEK293 cells. The transient transfection method is simple to construct vector plasmids and produces high-titer adenovirus-free AAV vectors, but has limited scalability and is not cost-effective to supply clinical studies.
[0411] The second strategy is a recombinant herpes simplex virus (rHSV)-based AAV production system that utilizes a rHSV vector to deliver the AAV vector and the Rep and Cap genes to cells.
[0412] The third method is based on AAV-producing cell lines derived from HeLa or A549, which stably carry AAV Rep / cap genes and genes of interest. AAV vector cassettes were stably integrated into the host genome (Clark et al., 1995, PMID: 8590738) or introduced by adenovirus containing the cassette. Stable cell lines in continuous culture fall into genetic instability with increasing passage numbers. Randomly integrated viral genes may increase cell instability and prematurely reduce stable cell propagation ability, thus affecting vector productivity. Selection of high-producing stable cell clones is costly and can take many months. In addition, cell propagation may alter the homeostasis, post-translational modification and secretion of recombinant proteins.
[0413] The use of GSH to generate AAV vectors that give rise to stable cell lines (e.g., integration of genes at the GSH locus, e.g., encoding viral capsids and / or recombinant proteins (e.g., gag, pol, rep, etc.)) ensures the quality of the producer cells to reach high vector productivity over the intended passages. The use of GSH also minimizes perturbations of cellular proteostasis during propagation, increasing product reproducibility over different production batches. A similar rationale can be applied to the production of other viral vectors, e.g., adenovirus-derived vectors, retrovirus- and lentivirus-derived vectors, herpesvirus-derived vectors, and alphavirus-derived vectors, e.g., Semliki Forest Virus (SFV) vectors, in which one or more components required for vector production are inserted into a defined GSH locus. Expression of these components can be modulated (e.g., using inducible promoters or using early versus late promoters) to mitigate undesirable early expression and reach a certain host cell number before amplification of the vector components and subsequent transgene packaging begins. Vector manufacturing processes in mammalian cell systems can greatly benefit from the use of GSH by lowering the costs associated with manufacturing and quality control while increasing cell stability, productivity, reproducibility and product safety, which directly impact patient benefits. Thus, directed recombination into GSH for rAAV production, as opposed to randomly generated producer cell lines, could accelerate the process by months or even years.
[0414] Therefore, in one embodiment, a method for producing a viral vector is provided herein.For example, the nucleic acid sequence required for viral assembly, for example, encoding one or more viral structural proteins (gag, VP1, VP2, VP3, etc.) and / or one or more replication proteins, which are operably linked to at least one expression control sequence for expression in host cell, can be integrated into the GSH locus of a host cell.Nucleic acid that includes at least one functional viral replication origin and, if necessary, further includes non-GSH nucleic acid for integration at GSH site can be provided to such cell to produce a viral vector.
[0415] Thus, in some embodiments, the method includes the steps of: (1) providing a host cell comprising: (i) a nucleic acid sequence comprising at least one functional viral origin of replication (e.g., at least one ITR nucleotide sequence), and optionally further comprising a nucleic acid operably linked to a promoter for expression in a target cell; (ii) a nucleic acid sequence comprising at least one gene encoding one or more viral structural proteins (e.g., capsid proteins, e.g., gag, VP1, VP2, VP3, variants thereof) operably linked to at least one expression control sequence for expression in the host cell; and (iii) a nucleic acid sequence comprising at least one gene encoding one or more viral replication proteins (e.g., Rep, pol) operably linked to at least one expression control sequence for expression in the host cell. and (2) maintaining the host cell under conditions such that the recombinant viral vector is produced.
[0416] In some embodiments, (ii) or (iii) is incorporated into GSH. In some embodiments, (ii) and (iii) are incorporated into GSH.
[0417] In some embodiments, the at least one functional viral origin of replication (e.g., at least one ITR nucleotide sequence) comprises (a) a dependoparvovirus ITR, and / or (b) an AAV ITR, optionally an AAV2 ITR.
[0418] In certain embodiments, the ITR is a terminal palindrome with a Rep binding element and trs structurally similar to the wild type ITR. The ITR may be selected from any one of AAV1-AAV13 and AAVrh.10. In certain embodiments, the ITR has an AAV2 RBE and trs. In some embodiments, the ITR is a chimera of different AAVs. In some embodiments, the ITR and Rep protein are from AAV5. In some embodiments, the ITR is synthetic and is composed of an RBE motif and trs GGTTGG, AGTTGG, AGTTGA, ...RRTTRR. The typical T-shaped structure of a terminal palindrome consisting of B / B' and C / C' stems can also be synthetically modified with substitutions and insertions that maintain the overall secondary structure based on folding predictions (available at the URL (http) of unafold.rna.albany.edu / ?q=mfold / DNA-Folding-Form). The stability of the ITR secondary structure is indicated by the Gibbs free energy, delta G, with lower, i.e., more negative, values indicating greater stability. The full-length, 145 nt ITR has a calculated ΔG=-69.91 kcal / mol. The B and C stems: GCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCG have ΔG=-22.44 kcal / mol. Substitutions and insertions resulting in structures with ΔG=-15 kcal / mol to -30 kcal / mol are functionally equivalent and do not differ from the wild-type dependoparvovirus ITR.
[0419] In some embodiments, the at least one expression control sequence for expression in a host cell comprises (a) a promoter, and / or (b) a Kozak-like expression control sequence.
[0420] In some embodiments, the promoter comprises (a) an immediate early promoter of an animal DNA virus, (b) an immediate early promoter of an insect virus, (c) an insect cell promoter, or (d) an inducible promoter. In some embodiments, the animal DNA virus is a cytomegalovirus (CMV), a dependoparvovirus, or an AAV. In some embodiments, the insect virus promoter is from a lepidoptera virus, or a baculovirus, optionally, the baculovirus is an Autographa californica multicapsid nuclear polyhedrosis virus (AcMNPV). In some embodiments, the promoter is a polyhedrin (polh) or an immediate early 1 gene (IE-1) promoter.
[0421] ...
Claims
1. A nucleic acid comprising at least a portion of a GSH nucleic acid.
2. The nucleic acid according to claim 1, wherein the GSH nucleic acid comprises an untranslated sequence or an intron.
3. The GSH comprises a sequence that is at least 65% identical to the sequence of any one of SYNTAX-GSH1 to SYNTAX-GSH56 or a fragment thereof, or The GSH comprises a sequence that is at least 65% identical to the sequence of genomic DNA of SYNTAX-GSH1, SYNTAX-GSH2, SYNTAX-GSH3 or SYNTAX-GSH4 or a fragment thereof, The nucleic acid according to claim 1.
4. The nucleic acid according to claim 1, further comprising at least one non-GSH nucleic acid, a nucleic acid having a sequence heterologous to GSH, a nucleic acid sequence not originally present at the GSH locus, or a transgene.
5. The nucleic acid according to claim 4, wherein the at least one non-GSH nucleic acid is adjacent to a GSH 5' homology arm and / or a GSH 3' homology arm, and the homology arm comprises a nucleic acid sequence that is at least about 65% identical to the target GSH nucleic acid.
6. The GSH homology arm is between 10 and 5000 base pairs in length, or The GSH homology arm is between 100 and 1500 base pairs in length, or The GSH homology arm is at least 30 base pairs in length, or The GSH homology arm is of a length sufficient to mediate homology-dependent integration into the GSH locus in the genome of the cell. The nucleic acid according to claim 5.
7. The at least one non-GSH nucleic acid is in an orientation for integration into the GSH in the forward orientation, or The at least one non-GSH nucleic acid is in an orientation for integration into the GSH in the reverse orientation, or The at least one non-GSH nucleic acid is (a) operably linked to a promoter, or (b) not operably linked to a promoter, or The at least one non-GSH nucleic acid is operably linked to a promoter, and the promoter is (a) a promoter heterologous to the nucleic acid to which it is operably linked; (b) a promoter that promotes tissue-specific expression of the nucleic acid; (c) a promoter that promotes constitutive expression of the nucleic acid; (d) an inducible promoter; (e) an early promoter of an animal DNA virus; (f) an early promoter of an insect virus; and (g) an insect cell promoter selected from the nucleic acid according to claim 5 **Claim 8** The nucleic acid according to claim 7, wherein the inducible promoter is modulated by an agent selected from small molecules, metabolites, oligonucleotides, riboswitches, peptides, peptidomimetics, hormones, hormone analogs, and light. **Claim 9** The nucleic acid according to claim 8, wherein the agent is selected from tetracycline, cumate, tamoxifen, estrogen, and antisense oligonucleotide (ASO), rapamycin, FKCsA, blue light, abscisic acid (ABA), and riboswitch. **Claim 10** The promoter promotes tissue-specific expression in hematopoietic stem cells, hematopoietic CD34+ cells, and epidermal stem cells, epithelial stem cells, nervous system stem cells, lung progenitor cells, muscle satellite cells, intestinal K cells, neuron cells, airway epithelial cells, or liver progenitor cells, or The promoter is selected from the CMV promoter, β-globin promoter, CAG promoter, AHSP promoter, MND promoter, Wiskott-Aldrich promoter, PKL-R promoter, polyhedron (polh) promoter, and immediate early 1 gene (IE-1) promoter. The nucleic acid according to claim 7 **Claim 11** The at least one non-GSH nucleic acid comprises a sequence encoding a coding RNA, or The at least one non-GSH nucleic acid is (a) a protein or a fragment thereof, preferably a human protein or a fragment thereof; (b) a therapeutic protein or a fragment thereof, an antigen-binding protein, or a peptide; (c) a suicide gene, optionally herpes simplex virus-1 thymidine kinase (HSV-TK); (d) a viral protein or a fragment thereof; (e) a nuclease, optionally a transcription activator-like effector nuclease (TALEN), zinc finger nuclease (ZFN), meganuclease, megaTAL, CRISPR endonuclease, Cas9 endonuclease or a variant thereof; (f) a marker, luciferase or GFP; and / or (g) a drug resistance protein, an antibiotic resistance gene, neomycin resistance encoding a sequence containing the nucleic acid according to claim 4 **Claim 12** The nucleic acid according to claim 11, wherein the sequence encoding the coding RNA is codon-optimized for expression in a target cell.
13. The nucleic acid according to claim 11, wherein the at least one non-GSH nucleic acid encoding the coding RNA further comprises a sequence encoding a signal peptide.
14. The viral protein or fragment thereof comprises a structural protein, VP1, VP2, VP3, a non-structural protein, or a Rep protein, or The viral protein or fragment thereof is (a) a parvovirus protein or fragment thereof, optionally VP1, VP2, VP3, NS1, or Rep; (b) a retrovirus protein or fragment thereof, optionally an envelope protein, gag, pol, or VSV-G; (c) an adenovirus protein or fragment thereof, optionally E1A, E1B, E2A, E2B, E3, E4, or a structural protein, A, B, C; and / or (d) a herpes simplex virus protein or fragment thereof, optionally ICP27, ICP4, or pac or The at least one non-GSH nucleic acid encoding the viral protein encodes a viral surface protein or fragment thereof. The nucleic acid according to claim 11.
15. (a) the surface protein or fragment thereof is an immunogenic surface protein that elicits an immune response in a host, (b) the surface protein or fragment thereof further comprises a signal peptide, (c) the gene encoding the surface protein or fragment thereof is operably linked to an inducible promoter, and / or (d) the nucleic acid encoding the surface protein or fragment thereof further comprises a suicide gene, or the surface protein is from a coronavirus, MERS, SARS, influenza virus, respiratory syncytial virus, hepatitis A, hepatitis B, hepatitis C, hepatitis D, hepatitis E, human papillomavirus, dengue virus serotype 1, dengue virus serotype 2, dengue virus serotype 3, dengue virus serotype 4, chikungunya virus, Zika virus, yellow fever virus, Mayaro virus, Ebola virus, Marburg virus, or Nipah virus, or wherein the surface protein is the spike protein of SARS-CoV-2 The nucleic acid according to claim 14 [
16. ] wherein the at least one non-GSH nucleic acid comprising a sequence encoding a protein or a fragment thereof is selected from the hemoglobin genes (HBA1, HBA2, HBB, HBG1, HBG2, HBD, HBE1, and / or HBZ), alpha-hemoglobin stabilizing protein (AHSP), coagulation factor VIII, coagulation factor IX, von Willebrand factor, dystrophin or shortened dystrophin, microdystrophin, utrophin or shortened utrophin, micro-utrophin, usherin (USH2A), GBA1, preproinsulin, insulin, GIP, GLP-1, CEP290, ATPB1, ATPB11, ABCB4, CPS1, ATP7B, KRT5, KRT14, PLEC1, Col7A1, ITGB4, ITGA6, LAMA3, LAMB3, LAMC2, KIND1, INS, F8 or a fragment thereof, a fragment encoding a B-domain deleted polypeptide, VIII SQ or p-VIII, IRGM, NOD2, ATG2B, ATG9, ATG5, ATG7, ATG16L1, BECN1, EI24 / PIG8, TECPR2, WDR45 / WIP14, CHMP2B, CHMP4B, Dynein, EPG5, HspB8, LAMP2, LC3b UVrag, VCP / p97, ZFYVE26, PARK2 / Parkin, PARK6 / PINK1, SQSTM1 / p62, SMURF, AMPK, ULK1, RPE65, CHM, RPGR, PDE6B, CNGA3, GUCY2D, RS1, ABCA4, MYO7A, HFE, hepcidin, TNFα receptor, IL-6 receptor, IL-12 receptor or a gene encoding a soluble form of the IL-1β receptor, and cystic fibrosis transmembrane conductance regulator (CFTR), or The antigen-binding protein is an antibody or an antigen-binding fragment thereof, and optionally, the antibody or an antigen-binding fragment thereof is selected from an antibody, Fv, F(ab’)2, Fab’, dsFv, scFv, sc(Fv)2, half antibody-scFv, tandem scFv, Fab / scFv-Fc, tandem Fab’, single-chain diabody, tandem diabody (TanAb), Fab / scFv-Fc, scFv-Fc, heterodimer IgG (CrossMab), DART, and diabody, or the antigen-binding protein specifically binds to TNFα, CD20, cytokine, IL-1, IL-6, BAFF, APRIL, IFN-γ, Her2, RANKL, IL-6R, GM-CSF, CCR5, pathogen, bacterial toxin, or viral capsid protein, or the antigen-binding protein is selected from adalimumab, etanercept, infliximab, certolizumab, golimumab, anakinra, rituximab, abatacept, tocilizumab, natalizumab, canakinumab, atacicept, belimumab, ocrelizumab, ofatumumab, fontolizumab, trastuzumab, denosumab, sarilumab, ranizumab, gemtuzumab, siltuximab, lerotlimab, and antigen-binding fragments thereof, The nucleic acid according to claim 11.
17. The at least one non-GSH nucleic acid includes a sequence encoding a non-coding RNA, and optionally, the non-coding RNA includes an antisense polynucleotide, lncRNA, piRNA, miRNA, shRNA, siRNA, antisense RNA, snoRNA, snRNA, scaRNA, and / or guide RNA. The nucleic acid according to claim 4.
18. The non-coding RNA targets a gene selected from DMT-1, ferroportin, TNFα receptor, IL-6 receptor, IL-12 receptor, IL-1β receptor, a gene encoding a mutant protein, mutant HFE, and CFTR. The nucleic acid according to claim 17.
19. the at least one non-GSH nucleic acid increases or restores the expression of an endogenous gene in a target cell, or the at least one non-GSH nucleic acid decreases or abolishes the expression of an endogenous gene in a target cell, The nucleic acid according to claim 4.
20. (a) a transcription control element, enhancer, transcription termination sequence, 5'UTR, 3'UTR, proximal promoter element, locus control region, β-globin LCR, DNase hypersensitive site (HS) of β-globin LCR, or polyadenylation signal sequence, and / or (b) a translation control element, Kozak sequence, or woodchuck hepatitis virus post-transcriptional control element further comprises, or the nucleic acid is selected from a plasmid, linear plasmid, minicircle, cosmid, artificial chromosome, BAC, linear covalently closed (LCC) DNA vector, minicircle, minivector and mininot, linear covalently closed (LCC) vector, MIDGE, MiLV, ministring, minicircle plasmid, mini-intron plasmid, pDNA expression vector, or a variant thereof, The nucleic acid according to claim 1.
21. At least a part of the GSH nucleic acid; at least a part of GSH within the nucleic acid according to any one of claims 1 to 20; at least a part of any one of SYNTHX-GSH1 to SYNTHX-GSH56; and / or a virus vector comprising the nucleic acid according to any one of claims 1 to 20.
22. The virus vector according to claim 21, selected from rAd, AAV, rHSV, retrovirus vector, poxvirus vector, lentivirus, vaccinia virus vector, HSV type 1 (HSV-1)-AAV hybrid vector, baculovirus expression vector system (BEVS), and variants thereof.
23. A cell comprising the nucleic acid according to any one of claims 1 to 20 or a virus vector comprising the nucleic acid according to any one of claims 1 to 20.
24. The cell is selected from a cell line or primary cell, or the cell is a mammalian cell, insect cell, bacterial cell, yeast cell, or plant cell, and optionally, the mammalian cell is a human cell or rodent cell, or the cell is an insect cell, and the insect cell is derived from a species of lepidoptera, The cell according to claim 23.
25. the species of Lepidoptera is Spodoptera frugiperda, Spodoptera litura, Spodoptera exigua, or Trichoplusia ni, or the insect cell is Sf9, or the cell is selected from hematopoietic cells, hematopoietic progenitor cells, hematopoietic stem cells, erythroid lineage cells, megakaryocytes, erythroid progenitor cells (EPC), CD34+ cells, CD44+ cells, erythrocytes, CD36+ cells, mesenchymal stem cells, nerve cells, intestinal cells, intestinal stem cells, gastrointestinal epithelial cells, endothelial cells, enteroendocrine cells, lung cells, lung progenitor cells, enterocytes, liver cells, hepatocytes, hepatic stellate cells, Kupffer cells (KC), liver sinusoidal endothelial cells (LSEC), liver progenitor cells, stem cells, progenitor cells, induced pluripotent stem cells (iPSC), skin fibroblasts, macrophages, brain microvascular endothelial cells (BMVEC), nervous system stem cells, muscle satellite cells, epithelial cells, airway epithelial cells, muscle progenitor cells, erythroid progenitor cells, lymphoid progenitor cells, B lymphoblast cells, B cells, T cells, basophilic endemic Burkitt lymphoma (EBL), polychromatic erythroblasts, epidermal stem cells, epithelial stem cells, embryonic stem cells, P63-positive keratinocyte-derived stem cells, keratinocytes, pancreatic β cells, K cells, L cells, HEK293 cells, HEK293T cells, MDCK cells, Vero cells, CHO, BHK1, NS0, Sp2 / 0, HeLa, A549, and orthochromatic erythroblasts, The cell according to claim 24.
26. A cell comprising at least one non-GSH nucleic acid integrated into GSH in the genome of the cell, wherein the GSH is selected from SYNTAX-GSH1 to SYNTAX-GSH56, or the GSH is selected from SYNTAX-GSH1, SYNTAX-GSH2, SYNTAX-GSH3 and SYNTAX-GSH4, a cell.
27. The cell according to claim 26, wherein the cell comprises the nucleic acid according to any one of claims 1 to 20.
28. The cell is selected from a cell line or primary cells, or the cell is a mammalian cell, an insect cell, a bacterial cell, a yeast cell, or a plant cell, and optionally, the mammalian cell is a human cell or a rodent cell, or the cell is an insect cell, and the insect cell is derived from a species of Lepidoptera, The cell according to claim 26.
29. The Lepidoptera species is Spodoptera frugiperda, Spodoptera litura, Spodoptera exigua, or Trichoplusia ni, or the insect cell is Sf9, or the cell is selected from hematopoietic cells, hematopoietic progenitor cells, hematopoietic stem cells, erythroid lineage cells, megakaryocytes, erythroid progenitor cells (EPC), CD34+ cells, CD44+ cells, erythrocytes, CD36+ cells, mesenchymal stem cells, nerve cells, intestinal cells, intestinal stem cells, gastrointestinal epithelial cells, endothelial cells, enteroendocrine cells, lung cells, lung progenitor cells, enterocytes, liver cells (e.g., hepatocytes, hepatic stellate cells, Kupffer cells (KC), liver sinusoidal endothelial cells (LSEC), liver progenitor cells), stem cells, progenitor cells, induced pluripotent stem cells (iPSC), skin fibroblasts, macrophages, brain microvascular endothelial cells (BMVEC), nervous system stem cells, muscle satellite cells, epithelial cells, airway epithelial cells, muscle progenitor cells, erythroid progenitor cells, lymphoid progenitor cells, B lymphoblast cells, B cells, T cells, basophilic endemic Burkitt lymphoma (EBL), polychromatic erythroblasts, epidermal stem cells, epithelial stem cells, embryonic stem cells, P63-positive keratinocyte-derived stem cells, keratinocytes, pancreatic β cells, K cells, L cells, HEK293 cells, HEK293T cells, MDCK cells, Vero cells, CHO, BHK1, NS0, Sp2 / 0, HeLa, A549, and orthochromatic erythroblasts, The cell according to claim 28.
30. A pharmaceutical composition comprising the nucleic acid according to any one of claims 1 to 20, a viral vector comprising the nucleic acid, and / or a cell comprising the nucleic acid or the viral vector.
31. A transgenic organism comprising at least one non-GSH nucleic acid integrated into GSH in the genome of a cell, wherein the GSH is selected from SYNTEX-GSH1 to SYNTEX-GSH56, or wherein the GSH is selected from SYNTEX-GSH1, SYNTEX-GSH2, SYNTEX-GSH3, and SYNTEX-GSH4, A transgenic organism.
32. A transgenic organism comprising a cell comprising the nucleic acid according to any one of claims 1 to 20.
33. The transgenic organism according to claim 32, wherein the organism is a mammal or a plant, and optionally, the mammal is a rodent, mouse, rat, goat, sheep, chicken, llama or rabbit.
34. A composition comprising the nucleic acid according to any one of claims 1 to 20, or a viral vector comprising said nucleic acid, or a pharmaceutical composition comprising said nucleic acid or said viral vector, for use in a method of inserting at least one non-GSH nucleic acid into the GSH locus of a cell, said method comprising introducing said nucleic acid, said viral vector, or said composition into said cell, whereby said non-GSH nucleic acid is integrated into said GSH locus by homologous recombination between said GSH locus in the genome and GSH 5' homology arms and GSH 3' homology arms adjacent to said non-GSH nucleic acid.
35. The non-GSH nucleic acid is integrated into the GSH in a forward orientation, or The non-GSH nucleic acid is integrated into the GSH in a reverse orientation, The composition according to claim 34.
36. The pharmaceutical composition according to claim 30, for use in the prevention or treatment of a disease.
37. A composition comprising the nucleic acid according to any one of claims 1 to 20, or a viral vector comprising said nucleic acid, and / or a pharmaceutical composition comprising said nucleic acid or said viral vector, for use in a method of modulating the level and / or activity of a protein in a cell, said method comprising introducing said nucleic acid, said viral vector, and / or said composition into said cell.
38. The level and / or activity increases, or The level and / or activity decreases or disappears, The composition according to claim 37.
39. A method for producing a biologic, comprising (a) (i) culturing a cell comprising the nucleic acid according to any one of claims 1 to 20, or (ii) a cell comprising a viral vector comprising said nucleic acid, and recovering the expressed biologic; or (b) recovering the expressed biologic from a transgenic organism comprising said cell comprising the method.
40. The biologic is an antigen-binding protein, or The biological agent is an antibody or an antigen-binding fragment thereof, and optionally, the antibody or the antigen-binding fragment thereof is selected from an antibody, Fv, F(ab’)2, Fab’, dsFv, scFv, sc(Fv)2, diabody-scFv, tandem scFv, Fab / scFv-Fc, tandem Fab’, single-chain diabody, tandem diabody (TanAb), Fab / scFv-Fc, scFv-Fc, heterodimeric IgG (CrossMab), DART, and diabody, or the biological agent specifically binds to TNFα, CD20, cytokine, IL-1, IL-6, BAFF, APRIL, IFN-γ, Her2, RANKL, IL-6R, GM-CSF, or CCR5, the biological agent is selected from adalimumab, etanercept, infliximab, certolizumab, golimumab, anakinra, rituximab, abatacept, tocilizumab, natalizumab, canakinumab, atacicept, belimumab, ocrelizumab, ofatumumab, fontolizumab, trastuzumab, denosumab, sarilumab, ranizumab, gemtuzumab, siltuximab, lerotlimab, and antigen-binding fragments thereof, or the biological agent is a therapeutic protein, and optionally, the therapeutic protein is insulin, The method according to claim 39.
41. A method for producing a viral vector, comprising: (1) (i) a nucleic acid sequence comprising at least one functional viral origin of replication or at least one ITR nucleotide sequence, and optionally further comprising a nucleic acid operably linked to a promoter for expression in a target cell; (ii) a nucleic acid sequence encoding one or more viral structural proteins, capsid proteins, gag, VP1, VP2, VP3, or variants thereof, operably linked to at least one expression control sequence for expression in a host cell; (iii) a nucleic acid sequence encoding one or more replication proteins, Rep, or pol, operably linked to at least one expression control sequence for expression in a host cell and preparing a host cell comprising the same, Optionally, said at least one replicating protein comprises (a) a Rep52 or Rep40 coding sequence or a fragment thereof encoding a functional replicating protein operably linked to at least one expression control sequence for expression in a host cell, and / or (b) a Rep78 or Rep68 coding sequence operably linked to at least one expression control sequence for expression in a host cell, wherein at least one of (i), (ii) and (iii) is stably integrated into at least one GSH selected from Table 3 within the host cell genome, and at least one vector, if present / when present, comprises the remainder of (i), (ii) and (iii) that is not stably integrated into the host cell genome; and (2) maintaining said host cell under conditions such that a recombinant viral vector is produced A method comprising.
42. either (ii) or (iii) is integrated into GSH, or said at least one functional viral origin of replication, or at least one ITR nucleotide sequence, comprises (a) a dependoparvovirus ITR, and / or (b) an AAV ITR, optionally an AAV2 ITR or said at least one expression control sequence for expression in said host cell comprises (a) a promoter, and / or (b) a Kozak-like expression control sequence The method according to claim 41, comprising.
43. The method according to claim 42, wherein said promoter comprises (a) an early promoter of an animal DNA virus, (b) an early promoter of an insect virus, (c) an insect cell promoter, or (d) an inducible promoter The method according to claim 42, comprising.
44. The method according to claim 43, wherein said animal DNA virus is cytomegalovirus (CMV), dependoparvovirus, or AAV, or said insect virus is a virus of Lepidoptera or baculovirus, and optionally, said baculovirus is Autographa californica multicapsid nuclear polyhedrosis virus (AcMNPV). The method according to claim 43, comprising.
45. The method according to claim 42, wherein said promoter is polyhedrin (polh) or early gene 1 (IE-1) promoter, or said promoter is an inducible promoter. The method according to claim 42, comprising.
46. The method according to claim 45, wherein the inducible promoter is modulated by an agent selected from small molecules, metabolites, oligonucleotides, riboswitches, peptides, peptidomimetics, hormones, hormone analogs, and light.
47. The method according to claim 46, wherein the agent is selected from tetracycline, cumate, tamoxifen, estrogen, and antisense oligonucleotide (ASO), rapamycin, FKCsA, blue light, abscisic acid (ABA), and riboswitches.
48. (a) the viral replication protein is an AAV replication protein, optionally the Rep52 and / or Rep78 protein; and / or (b) the viral structural protein is an AAV capsid protein, The method according to claim 41.
49. The method according to claim 48, wherein the AAV is AAV2.
50. manufacturing a viral vector comprising the nucleic acid according to any one of claims 1 to 20, or the host cell is a mammalian cell or an insect cell, or the host cell is a mammalian cell, the mammalian cell is a human cell or a rodent cell, or the host cell is an insect cell, and the insect cell is derived from a species of lepidoptera, The method according to claim 41.
51. the mammalian cell is selected from HEK293, HEK293T, HeLa, and A549, or the species of lepidoptera is Spodoptera frugiperda, Spodoptera litura, Spodoptera exigua, or Trichoplusia ni, or the insect cell is Sf9, The method according to claim 50.
52. The method according to claim 41, wherein the viral vector is selected from an adenovirus-derived vector, AAV, retrovirus, lentivirus-derived vector, lentivirus, herpesvirus-derived vector, alphavirus-derived vector, and Semliki Forest virus (SFV) vector.
53. A kit comprising the nucleic acid according to any one of claims 1 to 20, a viral vector containing the nucleic acid, a cell containing the nucleic acid or the viral vector, and / or a pharmaceutical composition containing the nucleic acid, the viral vector, or the cell.
54. A method for identifying a genomic safe harbor (GSH) locus, comprising: (a) determining the stability and / or level of expression of a marker gene; and (b) identifying a genomic locus where the inserted marker gene exhibits stable and / or high-level expression as GSH, wherein random insertion of at least one marker gene is induced into the genome within the cell. The method comprising the steps.
55. (a) identifying a genomic locus where the inserted marker gene does not affect cell viability, and / or (b) identifying a genomic locus where the inserted marker does not affect the differentiation ability (e.g., pluripotency, multipotency) of the cell. The method according to claim 54, further comprising the steps.
56. The cell is selected from a cell line, primary cell, stem cell or progenitor cell, and optionally, whether the cell is a stem cell or progenitor cell, or the cell is selected from embryonic stem cells, tissue-specific stem cells, mesenchymal stem cells, induced pluripotent stem cells (iPSCs), hematopoietic stem cells, hematopoietic CD34+ cells, and epidermal stem cells, epithelial stem cells, nervous system stem cells, lung progenitor cells, and liver progenitor cells, or the cell is a mammalian cell, and optionally, the mammalian cell is a mouse cell, dog cell, pig cell, non-human primate (NHP) cell, or human cell. The method according to claim 54.
57. The random insertion is (a) transfecting the cell with a nucleic acid molecule containing the marker gene, wherein the nucleic acid is optionally a plasmid; or (b) transducing the cell with an integrating virus containing the marker gene. induced by or the random insertion is induced by transducing the cell with an integrating virus containing the marker gene, the integrating virus is a retrovirus, and optionally, the retrovirus is a gammaretrovirus. The method according to claim 54. **Claim 58**: The at least one marker gene includes a screenable marker and / or a selectable marker, and optionally, (a) the screenable marker gene encodes green fluorescent protein (GFP), beta-galactosidase, luciferase, and / or beta-glucuronidase, and / or (b) the selectable marker gene is an antibiotic resistance gene, and optionally, the antibiotic resistance gene encodes blastocidin S-deaminase or amino 3'-glycosylphosphotransferase (neomycin resistance gene), or the marker gene is not operably linked to a promoter, or the marker gene is operably linked to a promoter, and optionally, the promoter is a tissue-specific promoter, or the GSH is of the intron type, exon type, or intergenic type, The method according to claim 54. **Claim 59**: A method for identifying a GSH locus, comprising: (a) determining the presence and location of an endogenous viral element (EVE) in the genome of a metazoan species; (b) determining the intergenic or intron boundary proximal to the EVE; and (c) identifying the intergenic or intron locus containing the EVE as the GSH locus The method comprising. **Claim 60**: (a) The presence and location of the EVE are determined by in silico searching for sequences homologous to the viral element; and / or (b) the intergenic or intron boundary proximal to the EVE is determined by aligning the sequence adjacent to the EVE with the orthologous sequences of one or more species in which the intergenic or intron boundary is known, The method according to claim 59. **Claim 61**: A method for identifying a GSH locus in an orthologous organism, comprising: (a) identifying the GSH locus in species A according to the method according to any one of claims 54 to 60; (b) determining the positions of (i) at least one cis-acting element proximal to the GSH locus in species A and (ii) the corresponding cis-acting element in species B; and Step of identifying the locus in species B as the GSH locus, wherein the distance between the locus in species B and the at least one cis - acting element is substantially proportional to the distance between the GSH locus in species A and the corresponding cis - acting element A method comprising the above. **Claim 62**: The at least one cis - acting element is selected from a splicing donor site, a splicing acceptor site, a polypyrimidine tract, a polyadenylation signal, an enhancer, a promoter, a terminator, a splicing control element, an intron splicing enhancer, and an intron splicing silencer, or the at least one cis - acting element comprises two or more cis - acting elements, or the at least one cis - acting element comprises two cis - acting elements, the first cis - acting element is located upstream (i.e., 5' side) of the GSH locus, and the second cis - acting element is located downstream (i.e., 3' side) of the GSH locus The method according to claim 61. **Claim 63**: The ratio of the distance between the at least one cis - acting element and the GSH locus to the distance between two cis - acting elements in species B is substantially proportional to the ratio of the distance between the corresponding cis - acting element and the GSH locus to the distance between two cis - acting elements in species A, or the distance between the at least one cis - acting element and the GSH locus in species B is at least 20% but at most 500% of the distance between the at least one cis - acting element and the GSH locus in species A, or the distance between the at least one cis - acting element and the GSH locus in species B is at least 80% but at most 250% of the distance between the at least one cis - acting element and the GSH locus in species A The method according to claim 62. **Claim 64**: The GSH locus is within the mammalian genome, and optionally, the mammalian genome is a mouse genome, a dog genome, a pig genome, an NHP genome, or a human genome, or the EVE or the viral element is (a) comprising a provirus of a viral genome or a fragment thereof; (b) comprising a DNA copy of a viral nucleic acid, viral DNA, or viral RNA; and / or (c) encoding a structural or non-structural viral protein, or a fragment thereof, or wherein said EVE comprises viral nucleic acid from a retrovirus, non-retrovirus, parvovirus, or circovirus, The method according to claim 59.
65. (a) said parvovirus is selected from any one of B19, murine minute virus (mvm), RA-1, AAV, bufavirus, hokovirus, bocavirus, and parvovirus, and optionally, said parvovirus is AAV; and / or (b) said circovirus is porcine circovirus (PCV) (e.g., PCV-1, PCV-2), The method according to claim 64.
66. The method according to claim 61, wherein said metazoan species is selected from Cetacea, Chiroptera, Lagomorpha, and Macropodidae.
67. Further comprising the method according to claim 59, or further comprising the step of performing at least one in vitro, ex vivo, and / or in vivo assay, wherein said in vivo assay does not include an in vivo assay on humans, The method according to claim 54.
68. Said at least one in vitro, ex vivo, and / or in vivo assay is (a) de novo targeted insertion of a marker gene into a locus in a cell (e.g., a human cell), and determining (i) cell viability, (ii) insertion efficiency, and / or (iii) marker gene expression; (b) targeted insertion of a marker gene into a locus in a progenitor cell or stem cell, differentiating in vitro, and determining (i) marker gene expression in all developmental lineages, and / or (ii) whether said insertion of said marker gene affects the differentiation of said progenitor cell or stem cell; (c) targeted insertion of a marker gene into a locus in a progenitor cell or stem cell, engrafting said cells into immunodeficient mice, and assessing marker gene expression in all developmental lineages in vivo; (d) targeted insertion of a marker gene into a locus in a cell and determination of a comprehensive cell transcription profile (e.g., using RNAseq or microarray); and (e) generating a transgenic knock-in mouse, wherein the genomic DNA of the mouse has a marker gene inserted into the locus, and optionally, the marker gene is operably linked to a tissue-specific or inducible promoter, the method according to claim 67, selected from the group consisting of: [
69. ] The method according to claim 68, wherein the progenitor cell or the stem cell is selected from embryonic stem cells, tissue-specific stem cells, mesenchymal stem cells, induced pluripotent stem cells (iPSCs), hematopoietic stem cells, hematopoietic CD34+ cells, and epidermal stem cells, epithelial stem cells, nervous system stem cells, lung progenitor cells, muscle satellite cells, intestinal K cells, and liver progenitor cells.