Novel CHO integration sites and their applications

By using specific nucleotide sequence loci and recombinase-mediated cassette exchange technology in mammalian cells, the stability and efficiency issues of recombinant protein expression were resolved, achieving both stability and high efficiency in recombinant protein expression.

CN107109434BActive Publication Date: 2026-05-26REGENERON PHARMACEUTICALS INC

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
REGENERON PHARMACEUTICALS INC
Filing Date
2015-10-21
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing mammalian cell expression systems, there are challenges in efficient gene transfer and stability of recombinant protein integration genes. In particular, long-term expression can easily lead to changes and instability in expression levels, which is especially challenging when engineering stable cell lines to accommodate additional genes such as multispecific antibodies.

Method used

By using a locus containing a specific nucleotide sequence, such as SEQ ID NO:1 or SEQ ID NO:4, to integrate exogenous nucleic acid sequences, and by using recombinase recognition sites such as the LoxP site to enhance expression efficiency, combined with recombinase-mediated box exchange (RMCE) technology, stable expression of recombinant proteins is ensured.

Benefits of technology

It achieved a significant improvement in the stability and expression level of recombinant protein expression, with the expression intensity increasing to 1.5 to 3 times that of normal expression, and the recombination efficiency significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN107109434B_ABST
    Figure CN107109434B_ABST
Patent Text Reader

Abstract

This invention provides expression-enhancing nucleotide sequences for eukaryotic expression systems, enabling enhanced and stable expression of recombinant proteins in eukaryotic cells. It also provides genomic integration sites for enhanced expression and methods of use for expressing genes of interest in eukaryotic cells. Furthermore, it provides chromosomal loci, sequences, and vectors for enhanced and stable gene expression in eukaryotic cells.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing of related patent applications

[0002] This application claims priority to U.S. Provisional Application No. 62 / 067,774, filed October 23, 2014, the entire contents of which are incorporated herein by reference.

[0003] The sequence list is incorporated by reference.

[0004] The sequence list in the 28KB ASCII text file named 32353_T0045US01_SequenceListing.txt, created on October 20, 2015 and filed with the U.S. Patent and Trademark Office via EFS-Web, is incorporated herein by reference. Technical Field

[0005] This invention provides stable integration and / or expression of recombinant proteins in eukaryotic cells. Specifically, the invention includes methods and compositions for improving protein expression in eukaryotic cells, particularly Chinese hamster (Cricetulus griseus) cell lines, by employing expression-enhancing nucleotide sequences. The invention includes polynucleotides that facilitate recombination-mediated cassette exchange (RMCE) and modified cells. The method of this invention integrates exogenous nucleic acids into specific chromosomal loci in the genome of Chinese hamster cells to facilitate enhanced and stable expression of recombinant proteins in modified cells. Background Technology

[0006] Cellular expression systems are designed to provide a reliable and efficient source for the preparation of a given protein, whether for research or therapeutic use. Recombinant protein expression in mammalian cells is a preferred method for preparing therapeutic proteins due to the ability of, for example, mammalian expression systems to make appropriate post-translational modifications to recombinant proteins.

[0007] Several cellular systems are available for protein expression, each containing various combinations of cis and, in some cases, trans-regulatory elements to achieve high levels of recombinant protein in short culture times. Despite the availability of numerous systems, efficient gene transfer and stability of the integrative genes used for recombinant protein expression remain challenging. Multiple local genetic factors will determine not only when the target gene of interest is expressed, but also whether the cell can functionally drive the gene toward highly productive transcription, or even whether expression will be long-term. Chromosomal integration sites, such as those in Chinese hamster ovary (CHO) cells, and control regions at specific gene loci or adjacent loci have been characterized in their respective fields (WO2012 / 138887A1; Li, Q. et al., 2002 Blood. 100:3077-3086). Similarly, target regulatory regions are typically identified within endogenous protein-coding regions. However, for long-term expression of the target transgene, a key consideration is minimizing disruption of the cellular genome to avoid changes in cell line phenotype.

[0008] Engineering stable cell lines to accommodate additional genes for expression, such as extra antibody chains in multispecific antibodies, is particularly challenging. Significant variations in the expression levels of integrated genes can occur. The integration of extra genes can lead to substantial variability and instability in expression due to local genetic conditions (i.e., positional effects). Therefore, improved mammalian expression systems are needed in this field. Summary of the Invention

[0009] In one aspect, the present invention provides a cell comprising an exogenous nucleic acid sequence integrated at a specific site within a locus, wherein the locus comprises a nucleotide sequence that is at least 90% identical to SEQ ID NO:1 or SEQ ID NO:4. In some embodiments, the locus comprises a nucleotide sequence that is at least 90% identical to SEQ ID NO:1. In some embodiments, the locus comprises a nucleotide sequence that is at least 90% identical to SEQ ID NO:4.

[0010] In another aspect, the present invention provides a polynucleotide comprising a first nucleic acid sequence integrated into a specific site (e.g., the locus of the present invention) within a second nucleic acid sequence. In one embodiment, the second nucleic acid sequence comprises the nucleotide sequence of SEQ ID NO:1. In another embodiment, the second nucleic acid sequence comprises the nucleotide sequence of SEQ ID NO:4.

[0011] In one embodiment, the second nucleic acid sequence is an expression-enhancing sequence selected from a nucleotide sequence having at least 90% nucleic acid identity with SEQ ID NO:1, or an expression-enhancing fragment thereof. In one embodiment, the second nucleic acid sequence is an expression-enhancing sequence selected from a nucleotide sequence having at least 90% nucleic acid identity with SEQ ID NO:4, or an expression-enhancing fragment thereof. In another embodiment, the expression-enhancing sequence is capable of enhancing the expression of a protein encoded by a foreign nucleic acid sequence. In another embodiment, the expression-enhancing sequence is capable of increasing the expression of a protein encoded by a foreign nucleic acid sequence by at least about 1.5 to at least about 3 times compared to the expression typically observed through random integration into the genome.

[0012] In another embodiment, the exogenous nucleic acid sequence is integrated into a specific site at any location within SEQ ID NO:1 or SEQ ID NO:4.

[0013] In some embodiments, a specific site located within or adjacent to a location within SEQ ID NO:1 is selected from the group consisting of: numbers spanning SEQ ID NO:1: 10-4,000, 100-3,900, 200-3,800, 300-3,700, 400-3,600, 500-3,500, 600-3,400, 700-3,300, 800-3,200, 900-3,100, 1,000-3,000, 1,100-2,900, 1,200-2,800, 1,300-2,700. Nucleotides at positions 1,200-2,600, 1,300-2,500, 1,400-2,400, 1,500-2,300, 1,600-2,200, 1,700-2,100, 1,800-2050, 1,850-2050, 1,900-2040, 1950-2,025, 1990-2021, 2002-2021, and 2,010-2,015. In some embodiments, specific sites located within or adjacent to positions within SEQ ID NO:1 are selected from the group consisting of: nucleotides spanning SEQ ID NO:1. NO:1's serial numbers are 1990-1991, 1991-1992, 1992-1993, 1993-1994, 1995-1996, 1996-1997, 1997-1998, 1999-2000, 2001-2002, 2002-2003, 2003-2004, 2004-2005, 2005-2006, 2006-2007. Nucleotides at positions 2007-2008, 2008-2009, 2009-2010, 2010-2011, 2011-2012, 2012-2013, 2013-2014, 2014-2015, 2015-2016, 2016-2017, 2017-2018, 2018-2019, 2019-2020, and 2020-2021.

[0014] In another embodiment, the specific site located within or adjacent to a location within SEQ ID NO:1 is selected from the group consisting of nucleotides spanning positions 10-500, 500-1,000, 500-2,100, 1,000-1,500, 1,000-2,100, 1,500-2,000, 1,500-2,500, 2,000-2,500, 2,500-3,000, 2,500-3,500, 3,000-3,500, 3,000-4,000, and 3,500-4,000 of SEQ ID NO:1. In some embodiments, the exogenous nucleic acid sequence is integrated at, within, or near any one or more of the specific sites described above.

[0015] In another embodiment, the exogenous nucleic acid sequence contains a recognition site located within the expression enhancement sequence as described above, provided that the expression enhancement sequence contains a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to the expression enhancement sequence of SEQ ID NO:1 or SEQ ID NO:4, and its expression enhancement fragment.

[0016] In one embodiment, the exogenous nucleic acid sequence includes a recombinase recognition site. In some embodiments, the exogenous nucleic acid sequence further includes at least one recombinase recognition site comprising a sequence independently selected from the following: LoxP site, Lox511 site, Lox2272 site, Lox2372, Lox5171, Loxm2, Lox71, Lox66, LoxFas, and frt site. In one embodiment, the recombinase recognition site is integrated within an expression-enhancing sequence. In another embodiment, the recombinase recognition site is located at a terminal nucleotide immediately adjacent to the 5' end of the gene cassette in the 5' direction, or immediately adjacent to the 3' end of the gene cassette in the 3' direction. In some embodiments, the at least one recombinase recognition site and the gene cassette are integrated within an expression-enhancing sequence.

[0017] In one embodiment, at least two recombinase recognition sites are present within the expression-enhancing sequence. In another embodiment, two recombinase recognition sites in opposite directions are integrated into the expression-enhancing sequence. In yet another embodiment, three recombinase recognition sites are integrated into the expression-enhancing sequence.

[0018] In one aspect, isolated Chinese hamster ovary (CHO) cells are provided, comprising an engineered expression-enhancing sequence of SEQ ID NO:1 or a fragment thereof. In one embodiment, the expression-enhancing sequence comprising the nucleotide sequence of SEQ ID NO:1 or SEQ ID NO:4, or a stable variant thereof, is engineered to integrate the exogenous nucleic acid sequence as described above. In other embodiments, the present invention provides isolated CHO cells comprising an exogenous nucleic acid sequence inserted into a locus comprising the expression-enhancing sequence of SEQ ID NO:1 or SEQ ID NO:4, or a stable variant thereof.

[0019] In one embodiment, the CHO cell further includes at least one recombinase recognition sequence within the expression-enhancing sequence. In another embodiment, the at least one recombinase recognition sequence is independently selected from the LoxP site, Lox511 site, Lox2272 site, Lox2372, Lox5171, Loxm2, Lox71, Lox66, LoxFas, and frt site. In another embodiment, the recombinase recognition site is located immediately adjacent to the 5' end of the gene cassette in the 5' direction, or immediately adjacent to the 3' end of the gene cassette in the 3' direction. In some embodiments, the at least one recombinase recognition site and the gene cassette are integrated within the expression-enhancing sequence of the CHO cell genome described herein.

[0020] In another embodiment, the at least one recombination recognition site is located as described above. It should be noted that the gene cassette contains an expression-enhancing sequence or expression-enhancing fragment thereof with at least 90% identity, at least about 91% identity, at least about 92% identity, at least about 93% identity, at least about 94% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, or at least about 99% identity, of nucleotides 1001 to 2001 of SEQ ID NO:1 (SEQ ID NO:2). In another embodiment, the at least one recombination recognition site is located as described above. It should be noted that the gene cassette contains an expression-enhancing sequence or expression-enhancing fragment thereof with at least 90% identity, at least about 91% identity, at least about 92% identity, at least about 93% identity, at least about 94% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, or at least about 99% identity, of nucleotides 2022 to 3022 of SEQ ID NO:1 (SEQ ID NO:3).

[0021] In yet another embodiment, the at least one recombinase recognition site is inserted into nucleotides 1990-1991, 1991-1992, 1992-1993, 1993-1994, 1995-1996, 1996-1997, 1997-1998, 1999-2000, 2001-2002, 2002-2003, 2003-2004, 2004-2005, 2005-2006, 2006-2007, 2007-2008 of SEQ ID NO:1. In the CHO cell genome at or within the nucleotides of 2008-2009, 2009-2010, 2010-2011, 2011-2012, 2012-2013, 2013-2014, 2014-2015, 2015-2016, 2016-2017, 2017-2018, 2018-2019, 2019-2020, 2020-2021 or 2021-2022.

[0022] In another embodiment, exogenous nucleic acids are inserted into nucleotides 1990-1991, 1991-1992, 1992-1993, 1993-1994, 1995-1996, 1996-1997, 1997-1998, 1999-2000, 2001-2002, 2002-2003, 2003-2004, 2004-2005, 2005-2006, 2006-2007, and 2007-2008 of SEQ ID NO:1. In the CHO genome at or within the nucleotides of 2008-2009, 2009-2010, 2010-2011, 2011-2012, 2012-2013, 2013-2014, 2014-2015, 2015-2016, 2016-2017, 2017-2018, 2018-2019, 2019-2020, 2020-2021 or 2021-2022.

[0023] In another embodiment, the exogenous nucleic acid is inserted into the CHO genome at or within nucleotides 2001-2022 of SEQ ID NO:1. In some embodiments, the exogenous nucleic acid is inserted into nucleotides 2001-2002 or 2021-2022 of SEQ ID NO:1, and nucleotides 2002-2021 of SEQ ID NO:1 are deleted due to the insertion. Similarly, the exogenous nucleic acid is inserted into the CHO genome at or within nucleotides 9302-9321 of SEQ ID NO:4. In some embodiments, the exogenous nucleic acid is inserted into nucleotides 9301-9302 or 9321-9322 of SEQ ID NO:4, and nucleotides 9302-9321 of SEQ ID NO:4 are deleted due to the insertion.

[0024] In some embodiments, the exogenous nucleic acid sequence integrated at a specific site within a locus (such as the nucleotide sequence of SEQ ID NO:1 or SEQ ID NO:4) comprises the gene of interest (GOI) (e.g., a nucleotide sequence encoding the protein of interest or "POI"). In some embodiments, the exogenous nucleic acid sequence comprises one or more genes of interest. In some embodiments, one or more genes of interest are selected from the group consisting of a first GOI, a second GOI, and a third GOI.

[0025] In some embodiments, the exogenous nucleic acid sequence integrated at a specific site within a locus (such as the nucleotide sequence of SEQ ID NO:1 or SEQ ID NO:4) comprises a GOI and at least one recombinase recognition site. In one embodiment, a first GOI is inserted as described above into an expression-enhancing sequence of SEQ ID NO:1 or SEQ ID NO:4 or an expression-enhancing sequence having at least 90% nucleotide identity with SEQ ID NO:1 or SEQ ID NO:4 or an expression-enhancing fragment thereof, and the first GOI is optionally operatively linked to a promoter, wherein the 5' flanking of the promoter-linked GOI (or the GOI) is a first recombinase recognition site and the 3' flanking is a second recombinase recognition site. In another embodiment, a second GOI is inserted at the 3' of the second recombinase recognition site, and the 3' flanking of the second GOI is a third recombinase recognition site.

[0026] In yet another embodiment, GOI is operatively linked to a promoter capable of driving GOI expression, wherein the promoter comprises a eukaryotic promoter that can be regulated by an activating factor or a repressor. In other embodiments, the eukaryotic promoter is operatively linked to a prokaryotic operon, and the eukaryotic cell optionally additionally contains a prokaryotic repressor protein.

[0027] In another embodiment, one or more optional markers are included between the first and second and / or the second and third recombinase recognition sites. In some embodiments, the first and / or second genes of interest and / or one or more optional markers are operatively linked to a promoter, wherein the promoters may be the same or different. In another embodiment, the promoter comprises a eukaryotic promoter (such as the CMV promoter or the SV40 late promoter), which is optionally controlled by a prokaryotic operon (such as the tet operon). In other embodiments, the cell additionally comprises a gene encoding a prokaryotic repressor (such as the tet repressor).

[0028] In another embodiment, the cell additionally contains a gene capable of expressing a recombinase. In some embodiments, the recombinase is a Cre recombinase.

[0029] In one aspect, a CHO host cell is provided, comprising an expression-enhancing sequence selected from SEQ ID NO:1 or SEQ ID NO:4, or an expression-enhancing sequence having at least 90% nucleotide identity with SEQ ID NO:1 or SEQ ID NO:4, or an expression-enhancing fragment thereof, comprising a first recombinase recognition site, followed by a first eukaryotic promoter, a first optional marker gene, a second eukaryotic promoter, a second optional marker gene, and a second recombinase recognition site. In further embodiments, the CHO host cell additionally provides a third eukaryotic promoter, a third marker gene, and a third recombinase recognition site. In one embodiment, the expression-enhancing sequence is included in SEQ ID NO:1 or SEQ ID NO:4 as described above.

[0030] In one embodiment, the first, second, and third recombinase recognition sites are different from each other. In some embodiments, the recombinase recognition sites are selected from the LoxP site, Lox511 site, Lox2272 site, Lox2372, Lox5171, Loxm2, Lox71, Lox66, LoxFas, and frt site.

[0031] In one embodiment, the first optional marker gene is a drug resistance gene. In another embodiment, the drug resistance gene is a neomycin resistance gene or a hygromycin resistance gene. In yet another embodiment, the second and third optional marker genes encode two different fluorescent proteins. In one embodiment, the two different fluorescent proteins are selected from the group consisting of: Discosoma coral (DsRed), green fluorescent protein (GFP), enhanced green fluorescent protein (eGFP), blue-green fluorescent protein (CFP), enhanced blue-green fluorescent protein (eCFP), yellow fluorescent protein (YFP), enhanced yellow fluorescent protein (eYFP), and far-infrared fluorescent proteins (e.g., mKate, mKate2, mPlum, mRaspberry, or E2-crimson).

[0032] In one embodiment, the first, second, and third promoters are identical. In another embodiment, the first, second, and third promoters are different from each other. In yet another embodiment, the first promoter is different from the second and third promoters, and the second and third promoters are identical. In more embodiments, the first promoter is an SV40 late promoter, and the second and third promoters are each human CMV promoters. In other embodiments, the first and second promoters are operatively linked to a prokaryotic operon.

[0033] In one embodiment, the host cell line has a gene encoding a recombinase that is exogenously added and integrated into its genome and operatively linked to a promoter. In another embodiment, the recombinase is a Cre recombinase. In yet another embodiment, the host cell has a gene encoding a regulatory protein that is integrated into its genome and operatively linked to a promoter. In many embodiments, the regulatory protein is a tet repressor protein.

[0034] In one embodiment, the first GOI and the second GOI encode an antibody light chain or a fragment thereof, or an antibody heavy chain or a fragment thereof. In another embodiment, the first GOI encodes an antibody light chain and the second GOI encodes an antibody heavy chain.

[0035] In some embodiments, the first, second, and third GOIs encode polypeptides selected from the group consisting of: a first light chain or a fragment thereof, a second light chain or a fragment thereof, and a heavy chain or a fragment thereof. In yet another embodiment, the first, second, and third GOIs encode polypeptides selected from the group consisting of: a light chain or a fragment thereof, a first heavy chain or a fragment thereof, and a second heavy chain or a fragment thereof.

[0036] In one aspect, a method for preparing a protein of interest is provided, comprising (a) introducing a gene of interest (GOI) into a CHO host cell, wherein the GOI is integrated into a specific locus containing a nucleotide sequence that is at least 90% identical to SEQ ID NO:1 or SEQ ID NO:4; (b) culturing the cells of (a) under conditions allowing expression of the GOI; and (c) recovering the protein of interest. In one embodiment, the protein of interest is selected from the group consisting of: subunits of immunoglobulins or fragments thereof, and receptors or ligand-binding fragments thereof. In some embodiments, the protein of interest is selected from the group consisting of: antibody light chains or antigen-binding fragments thereof, and antibody heavy chains or antigen-binding fragments thereof.

[0037] In some embodiments, the GOI is introduced into the cell using a targeting vector for recombinase-mediated cassette exchange (RMCE), and the CHO host cell genome includes at least one exogenous recognition sequence within a specific locus. In other embodiments, the CHO host cell genome includes at least one exogenous recognition sequence and an optional marker, optionally linked to a promoter, IRES, and / or polyadenylation (polyA) sequence within a specific locus.

[0038] In some embodiments, the CHO host cell genome contains one or more recombinase recognition sites as described above, and the GOI is introduced into a specific locus via the action of the recombinase recognition site.

[0039] In another embodiment, a targeting vector for homologous recombination is used to introduce the GOI into the cell, wherein the targeting vector comprises a 5' homologous arm, the GOI, and a 3' homologous arm, both homologous to a sequence present at the specific locus. In another embodiment, the targeting vector further comprises two, three, four, or five or more genes of interest. In yet another embodiment, one or more genes of interest are operatively linked to a promoter.

[0040] In another aspect, a targeting vector is provided, wherein the targeting vector comprises a 5' homologous arm, a GOI, and a 3' homologous arm, which are sequence homologous to sequences present in a locus containing at least 90% identical nucleotide sequences to SEQ ID NO:1 or SEQ ID NO:4. In another embodiment, the targeting vector further comprises two, three, four, or five or more genes of interest.

[0041] In another aspect, a method for modifying the genome of a CHO cell to integrate a foreign nucleic acid sequence is provided, comprising the step of introducing a carrier including a vector into the cell, wherein the carrier contains a foreign nucleic acid sequence, wherein the foreign nucleic acid is integrated into a locus of the genome containing a nucleotide sequence that is at least 90% identical to SEQ ID NO:1 or SEQ ID NO:4.

[0042] In some embodiments, the vector comprises a 5' homologous arm that is sequence-homologous to a locus of the genome containing a nucleotide sequence that is at least 90% identical to that of SEQ ID NO:1 or SEQ ID NO:4, an exogenous nucleic acid sequence, and a 3' homologous arm that is sequence-homologous to a locus of the genome containing a nucleotide sequence that is at least 90% identical to that of SEQ ID NO:1 or SEQ ID NO:4.

[0043] In some embodiments, the exogenous nucleic acid sequence in the vector contains

[0044] One or more recognition sequences. In other embodiments, the exogenous nucleic acid comprises one or more GOIs, such as nucleic acids that optionally label or encode POIs. In yet another embodiment, the exogenous nucleic acid comprises one or more GOIs and one or more recognition sequences.

[0045] In one embodiment, the vector comprises at least one additional vector or mRNA. In another embodiment, the additional vector is selected from the group consisting of: adenovirus, lentivirus, retrovirus, adeno-associated virus, integrative phage vector, nonviral vector, transposon and / or transposase, integrase substrate, and plasmid. In some embodiments, the additional vector comprises a nucleotide sequence encoding a site-specific nuclease for integrating a foreign nucleic acid sequence.

[0046] In some embodiments, the site-specific nuclease includes zinc finger nucleases (ZFNs), ZFN dimers, transcription activator-like effector nucleases (TALENs), TAL effector domain fusion proteins, or RNA-directed DNA endonucleases.

[0047] In another aspect, a vehicle is provided for modifying the genome of a CHO cell to integrate a foreign nucleic acid sequence, wherein the vehicle comprises a vector containing a 5' homologous arm that is sequence-homologous to a locus containing a nucleotide sequence that is at least 90% identical to SEQ ID NO:1 or SEQ ID NO:4, the foreign nucleic acid sequence, and a 3' homologous arm that is sequence-homologous to a locus containing a nucleotide sequence that is at least 90% identical to SEQ ID NO:1 or SEQ ID NO:4.

[0048] In some embodiments, the exogenous nucleic acid sequence includes one or more recognition sequences. In other embodiments, the exogenous nucleic acid includes one or more GOIs, such as nucleic acids that optionally label or encode POIs. In still other embodiments, the exogenous nucleic acid includes one or more GOIs and one or more recognition sequences.

[0049] In another aspect, a method is provided for modifying the genome of a CHO cell to express a therapeutic agent, the therapeutic agent comprising a carrier for introduction into the genome, an exogenous nucleic acid comprising a sequence for expressing the therapeutic agent, wherein the carrier comprises a 5' homologous arm homologous to a sequence present in the nucleotide sequence of SEQ ID NO:1, a nucleic acid encoding the therapeutic agent, and a 3' homologous arm homologous to a sequence present in the nucleotide sequence of SEQ ID NO:1 or SEQ ID NO:4.

[0050] In another aspect, the present invention provides a modified CHO host cell comprising a modified CHO genome, wherein the CHO genome is modified by inserting an exogenous recognition sequence into a locus having a nucleotide sequence that is at least 90% identical to that of SEQ ID NO:1.

[0051] In another aspect, the present invention provides a modified eukaryotic host cell comprising a modified eukaryotic genome, wherein the eukaryotic genome is modified at a target integration site in a non-coding region of the genome to insert a foreign nucleic acid. In some embodiments, the foreign nucleic acid is a recognition sequence. In other embodiments, the host cell is a mammalian host cell, such as a CHO cell. In other embodiments, the target integration site comprises an expression-enhancing sequence such as SEQ ID NO:1, provided that the sequence does not encode any endogenous protein. The present invention also provides a method for preparing such modified eukaryotic host cells.

[0052] In any of the aspects and embodiments described above, the expression enhancement sequence may be arranged in the same orientation as indicated in SEQ ID NO:1 or in the reverse orientation of SEQ ID NO:1.

[0053] Unless otherwise stated or apparent from the context, any aspect and embodiment of the invention may be used in conjunction with any other aspect or embodiment of the invention.

[0054] Other objectives and benefits will become apparent upon reviewing the following detailed description. Attached Figure Description

[0055] Figures 1A and 1B. Figure 1A: Schematic diagram of an operable construct utilizing multiple copies of a nucleic acid molecule expressing a GOI (e.g., a multi-chain antibody) and a selectable marker randomly introduced into a cellular genome (e.g., a CHO genome for identifying a target locus). An exemplary construct includes: a heavy chain (HC); a first copy selectable marker, such as a hygromycin resistance gene (Hyg); a first copy light chain (LC); a second copy selectable marker (e.g., Hyg), a second copy light chain (LC); and a third duplicate selectable marker (e.g., Hyg). Figure 1B: An example donor vector integrated into a natural locus via homologous recombination is identified as SEQ ID NO:1. The 5' and 3' homologous arms are derived from SEQ ID NO:1.

[0056] Figures 2A to 2C illustrate enhanced mRNA expression of the gene of interest (GOI) when the locus of SEQ ID NO:1 (LOCUS 1) is operatively linked to the locus of interest (GOI) compared to the same GOI operatively linked to the control locus instead of LOCUS 1. Figure 2A: Cells encoding the antibody of interest, operatively linked to the control locus, show equal numbers of gene copies of one heavy chain (HC) and two light chains (LC) of LOCUS 1. Figure 2B: Higher mRNA levels of GOI expressed in LOCUS 1 compared to the control locus mRNA. Figure 2C: Protein titer of cells expressing GOI in LOCUS 1 is 3-fold higher than that of cells expressing the same GOI at the control locus.

[0057] Figure 3A and 3B This study compares instance cassettes containing fluorescent markers and GOIs integrated at LOCUS 1 (e.g., mKate at the lox site exchanged with eYFP and GOI) with the same cassettes integrated at control loci (exchanged with different fluorescent markers at the lox site, such as dsRed2), where such integrations utilize Cre recombinase and recombinase-mediated cassette exchange (RMCE). These cassettes were used experimentally to measure GOI recombination efficiency and transcription.

[0058] Figure 4 shows that the mRNA level of the gene of interest (GOI) measured in the CHO cell pool expressing the GOI in LOCUS 1 (SEQ ID NO:1) is higher than that in the CHO cell pool expressing the same GOI under the same regulatory conditions but integrated into the control locus (i.e., EESYR). Detailed Implementation

[0059] Before describing the method of the present invention, it should be understood that the present invention is not limited to the specific methods and experimental conditions described, as such methods and conditions can vary. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting, as the scope of the invention will be limited only by the appended claims.

[0060] As used in this specification and the appended claims, unless the context clearly specifies otherwise, the singular forms “a / an” and “described” include multiple references. Thus, for example, a reference to “a method” includes one or more methods and / or one or more steps of the type described herein and / or that will become apparent to a person skilled in the art upon reading this disclosure.

[0061] Unless otherwise defined or specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0062] Although any methods and materials similar to or equivalent to those described herein may be used in the practice or testing of the invention, specific methods and materials are described hereafter. All publications mentioned herein are incorporated herein by reference in their entirety.

[0063] definition

[0064] DNA regions are operatively linked when they are functionally related. For example, if a promoter is capable of participating in the transcription of a coding sequence, then the promoter is operatively linked to that sequence; if a ribosome binding site is positioned to allow translation, then the ribosome binding site is operatively linked to the coding sequence. Generally, operative linking may include, but does not require, contiguity. For sequences such as secretory leader sequences, contiguity and appropriate placement within the reading frame are typical characteristics. In cases where an expression-enhancing sequence at the locus of interest is functionally related to the gene of interest (GOI), for example, where its presence enhances GOI expression and / or stabilizes integration, it is operatively linked to the GOI.

[0065] The term "enhancement," when used to describe enhanced expression, includes, for example, an enhancement exceeding, typically observed by at least about 1.5-fold to at least about 3-fold, enhancement compared to a pool of random integrators of a single copy of the same expression construct, either by random integration of a foreign sequence into the genome or by integration into a different locus. The doubling of expression enhancement observed using the sequences of the present invention is compared to the expression level of the same gene measured under substantially the same conditions, in the absence of the sequences of the present invention, for example, compared to another locus integrated into the genome of the same species. Enhanced recombination efficiency includes an enhancement of the recombination capacity of the locus (e.g., using a recombinase recognition site). Enhancement refers to a recombination efficiency exceeding that of random recombination (e.g., without using a recombinase recognition site, etc.), which is typically 0.1%. Preferably, the enhanced recombination efficiency exceeds that of random recombination by about 10-fold, or about 1%. Unless otherwise specified, the claimed invention is not limited to a specific recombination efficiency.

[0066] When the phrase "exogenous gene" or "exogenous nucleic acid" is used to refer to the locus of interest, the phrase refers to any DNA sequence or gene that is not present in the locus of interest, which is a locus found in nature. For example, "exogenous gene" in a CHO locus (e.g., a locus containing the sequence SEQ ID NO:1) can be a hamster gene not found in a specific CHO locus in nature (i.e., a hamster gene from another locus in the hamster genome), a gene from any other species (e.g., a human gene), a chimeric gene (e.g., a human / mouse gene), or any other gene not found in nature present in the CHO locus of interest.

[0067] When describing a locus of interest (such as SEQ ID NO:1 or SEQ ID NO:4) or a fragment thereof, the percentage of consistency means including homologous sequences that show consistency along adjacent homologous regions, but the presence of gaps, deletions or insertions that are not homologous in the compared sequences is not included in the calculation of the percentage of consistency.

[0068] As used herein, the determination of “percentage of similarity” between, for example, SEQ ID NO:1 or a fragment thereof and a species homolog will not include sequence comparisons of species homologs in which there is no homologous sequence comparison (i.e., SEQ ID NO:1 or a fragment thereof has an insertion at that point, or the species homolog has a gap or deletion, as the case may be). Therefore, “percentage of similarity” does not include penalties for gaps, deletions, and insertions.

[0069] In the context of nucleic acid sequences, a "homologous sequence" refers to a sequence that is substantially homologous to a reference nucleic acid sequence. In some embodiments, two sequences are considered substantially homologous if at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more of the corresponding nucleotides are identical in the relevant residue sequence segment. In some embodiments, the relevant sequence segment is a complete sequence.

[0070] "Targeted insertion" refers to a gene-targeting method used to guide the insertion or integration of a gene or nucleic acid sequence into a specific location in the genome; that is, to guide DNA to a specific site between two nucleotides in a linked polynucleotide chain. Targeted insertion can also be performed on specific gene cassettes, which include multiple genes, regulatory elements, and / or nucleic acid sequences. "Insertion" and "integration" are used interchangeably. It should be understood that the insertion of a gene or nucleic acid sequence (e.g., a nucleic acid sequence containing an expression cassette) may result in (or may be engineered to) the substitution or deletion of one or more nucleic acids, depending on the gene-editing technology employed.

[0071] A "recognition site" or "recognition sequence" is a specific DNA sequence that is recognized by nucleases or other enzymes to bind to and guide site-specific cleavage of the DNA backbone. Nucleases cleave DNA within the DNA molecule. Recognition sites are also referred to as target recognition sites in their respective fields.

[0072] A "recombinase recognition site" is a specific DNA sequence recognized by a recombinase, such as Cre recombinase (Cre) or flipping enzyme (flp). Site-specific recombinases can perform DNA rearrangements, including deletions, inversions, and translocations, when one or more of their target recognition sequences are strategically placed in an organism's genome. In one example, Cre specifically mediates recombination events at its DNA target recognition site loxP, which consists of two 13-bp inverted repeat sequences separated by an 8-bp spacer. More than one recombinase recognition site can be used, for example, to facilitate recombination-mediated DNA exchange. Variants or mutants of recombinase recognition sites (e.g., lox sites) can also be used (Araki, N. et al., 2002, Nucleic Acids Research, 30:19, e103).

[0073] "Recombinase-mediated cassette exchange" relates to a method for precisely replacing a genomic target cassette with a donor cassette. Typically, the molecular composition used in this method comprises: 1) a genomic target cassette with 5' and 3' side-attached sites specific to the target site for a particular recombinase; 2) a donor cassette with a matching target site side-attached site; and 3) a site-specific recombinase. Recombinase proteins are well-known in the field (Turan, S. and Bode J., 2011, FASEB J., 25, pp. 4088-4107) and are capable of precisely cleaving DNA (the DNA sequence) within a specific target site without adding or losing nucleotides. Common recombinase / site combinations include (but are not limited to) Cre / lox and Flp / frt.

[0074] A “vehicle” is a composition consisting of any polynucleotide or set of polynucleotides carrying exogenous nucleic acids for introduction into a cell. Vehicles include vectors, plasmids, and mRNA molecules delivered to cells via well-known transfection methods. In one instance, the mRNA introduced into the cell may be transient and not integrated into the genome; however, the mRNA may carry exogenous nucleic acids necessary for the integration process.

[0075] General Instructions

[0076] This invention is based, at least in part, on the discovery of unique sequences (i.e., loci) in the genome that exhibit more efficient recombination, insertion stability, and higher levels of expression compared to other regions or sequences in the genome. This invention is also based, at least in part, on the discovery that when such expression-enhancing sequences are identified, suitable genes or constructs can be exogenously added to or near these sequences, and the exogenously added genes can be advantageously expressed or used for further genomic modifications. These sequences, referred to as expression-enhancing sequences, are considered stable and not located within coding regions of the genome. These expression-enhancing and stable regions can be engineered for future cloning or genome editing events. Therefore, reliable expression systems are constructed into the cellular genome backbone.

[0077] This invention is also based on exogenous gene-specific targeting of integration sites. The method of this invention allows for the efficient “conversion” of the cellular genome into a suitable cloning cassette, for example, by employing recombinase-mediated cassette exchange (RMCE). For this purpose, the method of this invention employs cellular genomic recombinase recognition sites to place the gene of interest, thereby generating high-yield cell lines for recombinant protein production.

[0078] The compositions of the present invention can also be included in expression constructs, such as expression vectors for cloning and engineering new cell lines. Expression vectors containing the polynucleotides of the present invention can be used for transient protein expression or can be integrated into the genome via random or targeted recombination, such as homologous recombination or recombination mediated by recombinases that recognize specific recombination sites (e.g., Cre-lox-mediated recombination). Expression vectors containing the polynucleotides of the present invention can also be used to evaluate the efficacy of other DNA sequences, such as cis-regulatory sequences.

[0079] Integration sites are typically identified through random integration or by analyzing retroviral integration events. The CHO integration site described in detail in this article was identified by random integration into DNA encoding multi-stranded antibodies and by observing enhanced expression of the expressed protein.

[0080] An instance multichain antibody containing one heavy chain (HC) and two light chain (LC) copies was randomly integrated into an expression cassette containing an alternating hygromycin resistance gene in the genome (see, for example, the three identical Hyg genes depicted in Figure 1A). A stable, highly expressed clone was generated by integrating the expression cassette into the locus identified as SEQ ID NO:1.

[0081] Compared to integration into another region of the CHO genome (the control integration site), the example multi-chain antibody exhibited higher expression levels when integrated into the locus of SEQ ID NO:1. Interestingly, the antibody expression polynucleotide integrated into SEQ ID NO:1 had a comparable number of gene copies to the control integration site; however, the antibody expression polynucleotide integrated into SEQ ID NO:1 had a 3-fold higher protein titer.

[0082] The CHO cell genome was converted into a clone construct containing a recombinase recognition site using a targeted recombination approach (see, for example...). Figure 3A -B).

[0083] Essentially, after identifying the integration site of SEQ ID NO:1, an expression cassette is introduced using a recombinase recognition site (e.g., a lox site) in a locus, the expression cassette containing an expressible GOI, such as an optional marker (see, e.g.) Figure 3A -B), and any other required elements, such as promoters, enhancers, markers, operons, ribosome binding sites (e.g., internal ribosome entry sites), etc.

[0084] An illustration of an example donor construct for targeted integration of the lox site within SEQ ID NO:1 is shown in Figure 1B. The donor construct comprises an expression cassette driven by a neomycin (neo) resistance gene and an internal ribosome entry site (IRES), wherein the cassette contains a fluorescent marker (mKate) and is flanked at the 5' and 3' ends by recombinase recognition sites and 5' and 3' homologous arms (homological to SEQ ID NO:1). An insertion within the locus of SEQ ID NO:1 is shown, wherein the insertion causes the donor neo / mKate construct to replace the expression cassette containing the hygromycin resistance marker, wherein the expression cassette within the locus of SEQ ID NO:1 is flanked at its 5' and 3' ends by recombinase recognition sites connected to the 5' and 3' homologous arms (homological to SEQ ID NO:1) (see Figure 1B).

[0085] Compositions and methods are provided for stably integrating nucleic acid sequences into eukaryotic cells, wherein the nucleic acid sequences are capable of enhanced expression by integration into SEQ ID NO:1 or its expression-enhancing fragment. A recombinase recognition sequence containing SEQ ID NO:1 is provided to facilitate insertion into cells containing a GOI for expression of the protein of interest by the GOI. Compositions and methods are also provided for targeting integration sites associated with expression constructs (e.g., expression vectors) and for adding exogenous nucleic acids to CHO cells of interest.

[0086] Physical and functional characterization of CHO integration sites

[0087] The nucleic acid sequence of SEQ ID NO:1 (and more broadly, the nucleic acid sequence of SEQ ID NO:4) was empirically identified by upstream and downstream sequences of the integration site of a nucleic acid construct (containing an expression cassette) from a cell line expressing the protein at high levels. The nucleic acid sequences of the present invention provide sequences with novel functions associated with enhanced expression and stability of nucleic acids (e.g., exogenous nucleic acids containing GOI), and can function identically or differently from previously reported cis-acting elements (such as promoters, enhancers, locus control regions, scaffold attachment regions, or matrix attachment regions) without being bound by any single theory. SEQ ID NO:1 appears to lack any open reading frames (ORFs), making it unlikely that the locus encodes a novel trans-activating protein. A putative zinc finger protein has been identified at the 3' (downstream) genomic locus of SEQ ID NO:4.

[0088] The expression enhancement activity of expression cassettes containing coding sequences for the first hygromycin (Hyg) gene, first GOI, second Hyg gene, second GOI, third Hyg gene, and third GOI was identified at unique sites within the non-coding region of CHO genomic DNA. Expression vectors containing, for example, 5'-separated 1kb regions and 3'-separated 1kb regions were identified from the CHO genomic DNA non-coding region. Expression cassettes expressing GOI conferred high levels of recombinant protein expression in CHO cells after transfection with said expression vectors.

[0089] This invention covers expression vectors comprising the reverse SEQ ID NO:1 fragment or the SEQ ID NO:4 fragment. Other combinations of the fragments described herein can also be generated. Examples of other combinations of the fragments described herein include sequences containing multiple copies of the expression-enhancing sequence disclosed herein, or sequences derived by combining the disclosed SEQ ID NO:1 fragment or SEQ ID NO:4 fragment with other nucleotide sequences to achieve an optimal combination of regulatory elements. Such combinations can be sequentially linked or arranged to provide optimal spacing between the SEQ ID NO:1 or SEQ ID NO:4 fragments (e.g., by introducing spacer nucleotides between said fragments). Regulatory elements can also be arranged to provide optimal spacing between the SEQ ID NO:1 fragment and the regulatory element.

[0090] SEQ ID NO:1 and SEQ ID NO:4 disclosed herein were isolated from CHO cells. Limited homology to the identified expression-enhancing regions was found in other mammalian species (such as humans or mice); however, homologous sequences can be found in other tissue types derived from grey hamsters or in cell lines of other homologous species, and can be isolated using techniques well-known in the art. For example, other homologous sequences can be identified by cross-species hybridization or PCR-based techniques. Additionally, alterations can be made to the nucleotide sequences described in SEQ ID NO:1, SEQ ID NO:4, or fragments thereof using site-directed or random mutagenesis techniques well-known in the art. The expression-enhancing activity of the resulting sequence variants can then be tested as described herein. DNA with expression-enhancing activity that is at least approximately 90% identical in nucleic acid identity to SEQ ID NO:1, SEQ ID NO:4, or fragments thereof can be isolated using routine experiments and is expected to exhibit expression-enhancing activity. For fragments of SEQ ID NO:1 or SEQ ID NO:4, the percentage of identity refers to the portion of the reference natural sequence found in the fragment of SEQ ID NO:1 or SEQ ID NO:4. Therefore, homologs and variants of SEQ ID NO:1, SEQ ID NO:4 or fragments thereof are also covered by the embodiments of the present invention.

[0091] In some embodiments, the fragment of SEQ ID NO:1 is selected from the group consisting of: numbers spanning SEQ ID NO:1, such as 10-4,000, 100-3,900, 200-3,800, 300-3,700, 400-3,600, 500-3,500, 600-3,400, 700-3,300, 800-3,200, 900-3,100, 1,000-3,000, 1,100-2,900, 1,200-2,800, and 1,300-2,700. Nucleotides at positions 0, 1,200-2,600, 1,300-2,500, 1,400-2,400, 1,500-2,300, 1,600-2,200, 1,700-2,100, 1,800-2,050, 1,850-2,050, 1,900-2,040, 1,950-2,025, 1,990-2,211, 2,002-2,211 and 2,010-2,015. In another embodiment, the fragment of SEQ ID NO:1 is selected from the group consisting of nucleotides spanning positions 10-500, 500-1,000, 500-2,100, 1,000-1,500, 1,000-2,100, 1,500-2,000, 1,500-2,500, 2,000-2,500, 2,500-3,000, 2,500-3,500, 3,000-3,500, 3,000-4,000, and 3,500-4,000. In some embodiments, the exogenous nucleic acid sequence is integrated at or near a specific site within the fragment described above.

[0092] In another embodiment, the exogenous nucleic acid sequence is located within SEQ ID NO:1 or a fragment thereof as described above, or within a sequence that is at least about 90% identical, at least about 91% identical, at least about 92% identical, at least about 93% identical, at least about 94% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, or at least about 99% identical to the expression enhancement sequence or the expression enhancement fragment thereof of SEQ ID NO:1.

[0093] The methods provided herein can be used to generate cell populations expressing enhanced levels of the protein of interest. The absolute expression level will vary for the specific protein and depends on how efficiently the cells process the protein. Cell pools generated by integrating exogenous sequences into the expression-enhancing sequences of the present invention stabilize over time and can be treated as stable cell lines for most purposes. The recombination step can also be delayed until later in the development of the cell lines of the present invention.

[0094] CHO expression enhancement locus and its fragments

[0095] This invention covers an expression-enhancing fragment whose nucleotide sequence is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to the nucleotide sequence of SEQ ID NO:1 or SEQ ID NO:4. This invention includes vectors comprising fragments intended for transient or stable transfection, spanning SEQ ID NO:4. NO:1's serial numbers are 10-4,000, 100-3,900, 200-3,800, 300-3,700, 400-3,600, 500-3,500, 600-3,400, 700-3,300, 800-3,200, 900-3,100, 1,000-3,000, 1,100-2,900, 1,200-2,800, and 1,300-2,700. Positions 00, 1,200-2,600, 1,300-2,500, 1,400-2,400, 1,500-2,300, 1,600-2,200, 1,700-2,100, 1,800-2,050, 1,850-2,050, 1,900-2,040, 1,950-2,025, 1990-2,2021, 2002-2021, and 2,010-2,015. The invention also includes a eukaryotic cell containing such fragments, wherein the fragments are exogenous to the cell and integrated into the cell's genome, and the cell containing such fragments has at least one recombinase recognition site located within, 5' adjacent to, or 3' adjacent to the fragment.

[0096] In one embodiment, the expression-enhancing fragment of SEQ ID NO:1 is located within SEQ ID NO:1 at positions spanning numbers 10-500, 500-1,000, 500-2,100, 1,000-1,500, 1,000-2,100, 1,500-2,000, 1,500-2,500, 2,000-2,500, 2,500-3,000, 2,500-3,500, 3,000-3,500, 3,000-4,000, or 3,500-4,000.

[0097] In cases supporting stable integration and / or enhanced transcription of the integrated polynucleotide, the precise location of the locus insertion (i.e., integration) site relative to the illustrated site is not essential. In fact, the integration site can be located anywhere within or adjacent to the fragment of SEQ ID NO:1 or SEQ ID NO:4, as described herein. Whether a specific chromosomal location within or adjacent to the locus of interest supports stable integration and efficient transcription of the integrated exogenous gene can be determined according to standard procedures well-known in the art or the methods illustrated herein.

[0098] The integration sites considered herein are located within a locus containing the nucleotide sequence of SEQ ID NO:1 or SEQ ID NO:4, or very close to the locus of interest, for example, less than about 1 kb, 500 base pairs (bp), 250 bp, 100 bp, 50 bp, 25 bp, 10 bp, or less than about 5 bp upstream (5') or downstream (3') of the position of SEQ ID NO:1 on the chromosomal DNA. In some other embodiments, the integration sites used are located approximately 1000, 2500, 5000, or more base pairs upstream (5') or downstream (3') of the position of SEQ ID NO:1 or SEQ ID NO:4 on the chromosomal DNA.

[0099] Within this field, it should be understood that large genomic regions, such as scaffold / matrix attachment regions, are employed for efficient replication and transcription of chromosomal DNA. Scaffold / matrix attachment regions (S / MARs), also known as scaffold attachment regions (SARs) or matrix-associated or matrix-attached regions (MARs), are eukaryotic genomic DNA regions attached to the nuclear matrix. Without being bound by any particular theory, S / MARs are typically located in non-coding regions, separating a given transcribed region (e.g., chromatin domains) from its neighbors and providing a platform for the machining and / or binding of factors that enable transcription (such as recognition sites for DNases or polymerases). Some S / MARs have been characterized as approximately 14–20 kb long (Klar et al., 2005, Gene 364:79–89). Therefore, gene integration at LOCUS 1 (within or near SEQ ID NO:1 or SEQ ID NO:4) is expected to confer enhanced expression.

[0100] Those skilled in the art will recognize that several elements can be optimized to achieve high transcriptional activity at the target locus, thereby resulting in high expression of the inserted gene encoding the protein of interest. Elements to be considered include a strong promoter driving transcription, sufficient transcription machinery, and DNA with an open and accessible conformation. Insertion at the target locus can be optimized within the skill of those skilled in the art by targeting the integration site selected in SEQ ID NO:1 or SEQ ID NO:4.

[0101] In one embodiment, the expression-enhancing sequence of SEQ ID NO:1 was used to enhance GOI expression. Figure 2A shows the results of GOI operatively linked to SEQ ID NO:1 (LOCUS 1) compared to the same GOI integrated into the genome of CHO cells at different loci (control loci). The gene copy numbers measured in each cell line were equal, but the experiment showed that for GOI operatively linked to LOCUS 1, the mRNA level and protein titer of GOI expression in cells were 3-fold higher.

[0102] In various embodiments, GOI expression can be enhanced by placing GOI within SEQ ID NO:1 or SEQ ID NO:4. In various embodiments, the expression enhancement is at least about 1.5 times to about 3 times or more.

[0103] Gene modification target loci

[0104] Genetic engineering of the cellular genome at specific locations (i.e., target loci) can be achieved in several ways. Genetic editing techniques are used to stably integrate nucleic acid sequences into eukaryotic cells, where the nucleic acid sequences are exogenous sequences not typically found in such cells. Clonal expansion is necessary to ensure that progeny cells will possess consistent genotypic and phenotypic characteristics of the engineered cell line. In some instances, natural cells are modified using homologous recombination techniques to integrate exogenous nucleic acid sequences into SEQ ID NO:1 or SEQ ID NO:4. In other instances, cells containing at least one recombinase recognition sequence within SEQ ID NO:1 or SEQ ID NO:4 are provided to facilitate the integration of exogenous nucleic acid sequences or genes of interest.

[0105] In some instances, cells containing a first recombinase recognition sequence and a second recombinase recognition sequence are provided, wherein each of the first and second recombinase recognition sequences is selected from the group comprising: LoxP, Lox511, Lox5171, Lox2272, Lox2372, Loxm2, Lox-FAS, Lox71, Lox66, and mutants thereof. In this case, if recombinase-mediated cassette exchange (RMCE) is required, the site-specific recombinase is Cre recombinase or a derivative thereof. In other instances, each of the first and second recombinase recognition sequences is selected from the group comprising FRT, F3, F5, FRT mutant-10, FRT mutant+10, and mutants thereof, and in this context, if RMCE is required, the site-specific recombinase is Flp recombinase or a derivative thereof. In yet another instance, each of the first and second recombinase recognition sequences is selected from the group comprising attB, attP, and mutants thereof, and in this case, if RMCE is required, the site-specific recombinase is phiC31 integrase or a derivative thereof.

[0106] In one aspect, methods and compositions for stably integrating nucleic acid sequences into SEQ ID NO:1 or SEQ ID NO:4 or their expression-enhancing fragments are via homologous recombination. The nucleic acid molecule of interest, i.e., a gene or polynucleotide, can be inserted into the targeted locus (i.e., SEQ ID NO:1) via homologous recombination or by using a site-specific nuclease method that specifically targets the sequence at the integration site. Regarding homologous recombination, homologous polynucleotide molecules (i.e., homologous arms) align and exchange a segment of their sequence. If the transgene is side-joined with a homologous genomic sequence, then the transgene can be introduced during this exchange. In one instance, a recombinase recognition site can be introduced into the host cell genome at the integration site.

[0107] Homologous recombination in eukaryotic cells can be promoted by introducing breaks at integration sites in chromosomal DNA. Model systems have demonstrated that the frequency of homologous recombination increases during gene targeting if double-strand breaks are introduced into the target chromosomal sequence. This can be achieved by targeting certain nucleases to specific integration sites. DNA-binding proteins that recognize DNA sequences at target loci are known in the field. Gene-targeting vectors are also used to promote homologous recombination. In the absence of gene-targeting vectors for homology-directed repair, cells often close double-strand breaks via non-homologous end joining (NHEJ), which may result in the deletion or insertion of multiple nucleotides at the cleavage site. Insertions or deletions (InDels) should be present, resulting in the random insertion or deletion of small amounts of nucleotides at the break site, and these InDels can shift or disrupt any open reading frames (ORFs) of the gene within the target locus. It should be understood that the locus identified as SEQ ID NO:1 (or SEQ ID NO:4) is not a gene coding region. Therefore, it is assumed that insertions and / or deletions at this locus do not disrupt endogenous gene transcription.

[0108] Homology-guided repair (or homology-guided recombination) (HDR) is particularly suitable for inserting or integrating genes at target loci. The donor construct contains homologous arms derived from SEQ ID NO:1 or SEQ ID NO:4 as described herein.

[0109] The construction of gene-targeting vectors and the selection of nucleases are within the skill of those skilled in the art to which this invention pertains.

[0110] In some instances, zinc finger nucleases (ZFNs) with a modular structure and containing individual zinc finger domains recognize specific 3-nucleotide sequences (e.g., targeting integration sites) in a target sequence. Some embodiments may utilize ZFNs with combinations of individual zinc finger domains targeting multiple target sequences.

[0111] Transcription activator-like (TAL) effector nucleases (TALENs) can also be used for site-specific genome editing. The DNA-binding domain of a TAL effector protein is typically used in combination with the non-specific cleavage domain of a restriction nuclease (such as FokI). In some embodiments, a fusion protein comprising a TAL effector protein DNA-binding domain and a restriction nuclease cleavage domain is used to recognize and cleave DNA at a target sequence within the locus of this invention (Boch J et al., 2009 Science 326:1509-1512).

[0112] RNA-directed endonucleases (RGENs) are programmable genome engineering tools developed from bacterial adaptive immune mechanisms. In this system (Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) / CRISPR-associated (Cas) Immunoreactivity), the protein Cas9 forms a sequence-specific endonuclease when it complexes with two RNAs (one of which directs target selection). RGENs consist of components (Cas9 and tracrRNA) and target-specific CRISPR RNA (crRNA). The efficiency of DNA target cleavage and the location of the cleavage site vary based on the position of the pre-intercalation sequence neighbor motif (PAM), which is an additional requirement for target recognition (Chen, H. et al., J. Biol. Chem. 2014, March 14, published online as manuscript M113.539726).

[0113] Strategies for identifying sequences specific to the target locus of SEQ ID NO:1 are known in the art; however, alignment of many of these sequences with the CHO genome reveals potential off-target sites with 16-17 base pair matches. An example 20 bp guide RNA encoded by the sequence described in SEQ ID NO:5 (corresponding to nucleotides 1990-2001 of SEQ ID NO:1) is suitable for RNA-guided CRISPR / Cas gene editing of SEQ ID NO:1 or SEQ ID NO:4. A promoter containing a small guide RNA and tracrRNA (e.g., SEQ ID NO:6) for driving expression, and a plasmid carrying a suitable Cas9 enzyme under promoter control, can be co-transfected with a donor vector (carrying the gene of interest with 5' and 3' homologous arms) for targeted integration via this method. Various modifications and variants of RNA molecules other than those described above will be apparent to those skilled in the art and are intended to fall within the scope of this invention.

[0114] In some embodiments, the vehicle for introduction into the genome, i.e., an exogenous nucleic acid containing a sequence encoding a gene of interest, a recognition sequence, or a gene cassette, may, depending on the specific circumstances, include a vector carrying the exogenous nucleic acid and one or more additional vectors or mRNAs. In one embodiment, the one or more additional vectors or mRNAs contain a nucleotide sequence encoding a site-specific nuclease, including (but not limited to) zinc finger nucleases (ZFNs), ZFN dimers, transcription activator-like effector nucleases (TALENs), TAL effector domain fusion proteins, and RNA-directed DNA endonucleases. In some embodiments, the one or more vectors or mRNAs contain a first vector having a guide RNA, a tracrRNA, and a nucleotide sequence encoding a Cas enzyme, and a second vector containing a donor (exogenous) nucleotide sequence. Such donor sequences contain a nucleotide sequence encoding a gene of interest, or a recognition sequence, or a gene cassette containing any of these exogenous elements intended for targeted insertion. When using mRNA, the mRNA can be transfected into cells using common transfection methods known to those skilled in the art and can encode enzymes such as transposases or endonucleases. Although the mRNA introduced into the cell may be transient and not integrated into the genome, it may carry exogenous nucleic acids necessary or beneficial for integration. In some cases, if only short-term expression is required to achieve the desired integration of GOI, then mRNA is chosen to eliminate any risk of persistent side effects from the additional polynucleotides.

[0115] Other homologous recombination methods are available to technicians, such as BuD-derived nucleases (BuDN) with precise DNA binding specificity (Stella, S. et al. Acta Cryst. 2014, D70, 2042-2052). Precise genome modification methods are based on the selection of available tools compatible with the unique target sequence within SEQ ID NO:1 to avoid disrupting the cell phenotype.

[0116] Gene-targeted constructs

[0117] The polynucleotide sequence to be integrated into the host genome can be any industrially applicable DNA sequence suitable for generating cell expression systems, such as a recognition sequence. The polynucleotide sequence to be integrated into the host genome can encode any therapeutically or industrially applicable protein as described herein. Identifying the target sequence within the target locus for integration of the exogenous nucleic acid sequence depends on a variety of factors. Depending on the homologous recombination method employed, selecting a sequence homologous to SEQ ID NO:1 or SEQ ID NO:4 is precisely within the skill of a technician. When using site-specific nuclease vectors, it is necessary to identify additional components (sequence compositions) intended for specific sites of DNA cleavage.

[0118] Therefore, gene-targeting constructs typically incorporate such nucleotide sequences to facilitate the targeted integration of exogenous nucleic acid sequences into the locus of interest. In some embodiments, the construct includes a first homologous arm and a second homologous arm. In other embodiments, the construct (e.g., a gene cassette) includes a homologous arm derived from SEQ ID NO:1 or SEQ ID NO:4. In some embodiments, the homologous arm includes a nucleotide sequence homologous to the nucleotide sequence present in SEQ ID NO:1 or SEQ ID NO:4. In a particular embodiment, the construct includes a 5' homologous arm having the nucleotide sequence of SEQ ID NO:2 (corresponding to nucleotides 1001-2001 of SEQ ID NO:1) and a 3' homologous arm having the nucleotide sequence of SEQ ID NO:3 (corresponding to nucleotides 2022-2001 of SEQ ID NO:1). The homologous arms, such as the first homologous arm (also referred to as the 5' homologous arm) and the second homologous arm (also referred to as the 3' homologous arm), are homologous to the target sequence within the locus. The 5' to 3' homologous arms can amplify regions or target sequences within the locus that contain at least 1 kb, or at least about 2 kb, or at least about 3 kb, or at least about 4 kb, or at least 5 kb, or at least about 10 kb. In other embodiments, the total number of nucleotides selected for the target sequences used in the first and second homologous arms comprises at least 1 kb, or at least about 2 kb, or at least about 3 kb, or at least about 4 kb, or at least 5 kb, or at least about 10 kb. In some cases, the distance between the 5' homologous arm and the 3' homologous arm (homological to the target sequence) includes at least 5 bp, 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or at least 1 kb, or at least about 2 kb, or at least about 3 kb, or at least about 4 kb, or at least 5 kb, or at least about 10 kb. When SEQ ID NO:2 and SEQ ID NO:3 are selected as the 5' and 3' homologous arms, the distance between the two homologous arms can be 20 nucleotides (corresponding to nucleotides 2002-2021 of SEQ ID NO:1); and such homologous arms can mediate the integration of exogenous nucleic acid sequences into the locus containing SEQ ID NO:1, for example, within nucleotides 1990-2021 or 2002-2021 of SEQ ID NO:1 and simultaneously with the deletion of nucleotides 2002-2021 of SEQ ID NO:1.

[0119] In other embodiments, the construct includes a first homologous arm and a second homologous arm, wherein the combined first and second homologous arms contain a target sequence that replaces an endogenous sequence within the locus. In yet another embodiment, the first and second homologous arms contain a target sequence integrated into or inserted into an endogenous sequence within the locus.

[0120] The modified cell lines are created by integrating one or more recombinase recognition sites at the location specified in SEQ ID NO:1. These modified cell lines may also contain additional exogenous genes for negative or positive selection of the gene of interest for expression.

[0121] This invention provides a method for modifying the genome of CHO cells, comprising introducing one or more carriers into the cells, wherein the one or more carriers comprise an exogenous nucleic acid having a sequence for integration, a 5' homologous arm homologous to a sequence present in the nucleotide sequence of SEQ ID NO: 1, and a 3' homologous arm homologous to a sequence present in the nucleotide sequence of SEQ ID NO: 1. In some embodiments, the method further provides one or more carriers comprising a nuclease and a composition for site-specific DNA cleavage at the integration site.

[0122] Modified cell lines can serve as convenient and stable expression systems for recombinase-mediated cassette exchange (RMCE). Nucleic acid sequences encoding the protein of interest can be readily integrated into modified cells containing SEQ ID NO:1 or its enhanced expression fragment and having at least one recombinase recognition site, for example via the RMCE method.

[0123] Recombinant expression vectors may contain synthetic or cDNA-derived DNA fragments encoding proteins, operatively linked to suitable transcriptional and / or translational regulatory elements derived from mammalian, viral, or insect genes. These regulatory elements include transcription promoters, enhancers, sequences encoding suitable mRNA ribosome binding sites, and sequences controlling transcription and translation termination, as described in detail below. Mammalian expression vectors may also contain non-transcriptional elements, such as origins of replication, other 5' or 3' flanking non-transcriptional sequences, and 5' or 3' non-translational sequences, such as splice donor and acceptor sites. Optional marker genes to help identify transfectants may also be incorporated.

[0124] Fluorescent markers are optional marker genes suitable for identifying gene cassettes that have been or have not yet been successfully inserted and / or replaced, depending on the specific circumstances. Examples of fluorescent markers are well known in the field and include (but are not limited to) Discosoma coral (DsRed), green fluorescent protein (GFP), enhanced green fluorescent protein (eGFP), blue-green fluorescent protein (CFP), enhanced blue-green fluorescent protein (eCFP), yellow fluorescent protein (YFP), enhanced yellow fluorescent protein (eYFP), and far-infrared fluorescent proteins (e.g., mKate, mKate2, mPlum, mRaspberry, or E2-crimson). See also, for example, Nagai, T. et al., 2002 Nature Biotechnology 20:87-90; Heim, R. et al., February 23, 1995 Nature 373:663-664; and Strack, RL et al., 2009 Biochemistry 48:8279-81.

[0125] Transcriptional and translational control sequences in expression vectors suitable for transfection into vertebrate cells can be provided from viral sources. For example, commonly used promoters and enhancers are derived from viruses such as polyomavirus, adenovirus 2, simian virus 40 (SV40), and human cytomegalovirus (CMV). Viral genome promoters, control, and / or signaling sequences can be used to drive expression; these control sequences are compatible with the chosen host cell. Non-viral cell promoters (e.g., β-globulin and EF-1α promoters) can also be used, depending on the cell type expressing the recombinant protein.

[0126] DNA sequences derived from the SV40 viral genome, such as the SV40 origin, early and late promoters, enhancers, splicing sites, and polyadenylation sites, can be used to provide other genetic elements useful for the expression of heterologous DNA sequences. Early and late promoters are particularly useful because they can be readily obtained from the SV40 virus as a fragment that also contains the SV40 origin of replication (Fiers et al., Nature 273:113, 1978). Smaller or larger SV40 fragments can also be used. Typically, this includes a sequence of approximately 250 bp extending from the Hind III site at the SV40 origin of replication to the BglI site.

[0127] Bicistronic expression vectors for expressing multiple transcripts have been previously described (Kim SK and Wold BJ, Cell 42:129, 1985) and can be used in combination with the expression-enhancing sequences of the present invention (e.g., SEQ ID NO:1) or fragments thereof. Other types of expression vectors will also be useful, such as those described in U.S. Patent No. 4,634,665 (Axel et al.) and U.S. Patent No. 4,656,134 (Ringold et al.).

[0128] Proteins of interest

[0129] Any protein of interest suitable for expression in eukaryotic cells may be used. For example, proteins of interest include (but are not limited to) antibodies or their antigen-binding fragments, chimeric antibodies or their antigen-binding fragments, ScFv or fragments thereof, Fc fusion proteins or fragments thereof, growth factors or fragments thereof, cytokines or fragments thereof, or extracellular domains of cell surface receptors or fragments thereof. Proteins of interest may be simple polypeptides composed of a single subunit or complex multi-subunit proteins containing two or more subunits.

[0130] Host cells and transfection

[0131] The host cells used in the method of this invention are mammalian host cells, including, for example, Chinese hamster ovary (CHO) cells and mouse cells. In a preferred embodiment, the invention provides a nucleic acid sequence fragment of SEQ ID NO:1 that encodes an expression-enhancing sequence in CHO cells. An integration site can be found within any fragment of SEQ ID NO:1 or SEQ ID NO:1. For example, the integration site can be a recombinase recognition site located within any fragment of SEQ ID NO:1 or SEQ ID NO:1. One example of a suitable integration site is the LoxP site. Another example of a suitable integration site is two recombinase recognition sites, for example selected from the group consisting of: LoxP site, Lox511 site, Lox2272 site, Lox2372 site, Loxm2 site, Lox71 site, Lox66 site, and Lox5171 site. In other embodiments, the integration site is located within or adjacent to the sequence, selected from the group consisting of: spanning SEQ ID NO:1. IDNO:1 numbers 10-4,000, 100-3,900, 200-3,800, 300-3,700, 400-3,600, 500-3,500, 600-3,400, 700-3,300, 800-3,200, 900-3,100, 1,000-3,000, 1,100-2,900, 1,200-2,800, 1,300-2,700 Nucleotides at positions 0, 1,200-2,600, 1,300-2,500, 1,400-2,400, 1,500-2,300, 1,600-2,200, 1,700-2,100, 1,800-2,050, 1,850-2,050, 1,900-2,040, 1,950-2,025, 1,990-2,211, 2,002-2,211, and 2,010-2,015.In some embodiments, the integration site located within or adjacent to a location within SEQ ID NO:1 is selected from the group consisting of: numbers spanning SEQ ID NO:1: 1990-1991, 1991-1992, 1992-1993, 1993-1994, 1995-1996, 1996-1997, 1997-1998, 1999-2000, 2001-2002, 2002-2003, 2003-2004, 2004-2005, 2005-2006, 2006-2007. Nucleotides at positions 2007-2008, 2008-2009, 2009-2010, 2010-2011, 2011-2012, 2012-2013, 2013-2014, 2014-2015, 2015-2016, 2016-2017, 2017-2018, 2018-2019, 2019-2020, and 2020-2021.

[0132] This invention includes mammalian host cells transfected with the expression vector or mRNA of this invention. While any mammalian cell can be used, in one particular embodiment, the host cell is a CHO cell.

[0133] Transfected host cells include cells transfected with an expression vector or mRNA molecule containing a sequence encoding a protein or polypeptide. The expressed protein may be secreted into the culture medium, depending on the selected nucleic acid sequence, but may remain in the cell or deposited in the cell membrane. Various mammalian cell culture systems can be used to express recombinant proteins. Other cell lines generated for specific selection or amplification procedures will also be applicable to the methods and compositions provided herein, provided that a target locus with at least 80% homology to SEQ ID NO:1 has been identified. The proposed cell line is a CHO cell line named K1. To obtain high yields of recombinant protein, the host cell line may be pre-acclimated to the bioreactor medium under appropriate conditions.

[0134] Several transfection methods are known in this field, and they are reviewed in Kaufman (1988) Meth. Enzymology 185:537. The chosen transfection protocol will depend on the host cell type and the nature of the GOI, and can be selected based on routine experiments. The basic requirement of any such protocol is to first introduce DNA encoding the protein of interest into a suitable host cell, and then identify and isolate the host cell with the incorporated heterologous DNA in a relatively stable and expressible manner. The mRNA molecules encoding proteins that are suitable for integration into the host cell genome or other functions can be transient and therefore time-limited.

[0135] Transfection protocols and protocols used to introduce peptide or polynucleotide sequences into cells can be modified. Non-restrictive transfection methods include chemical-based methods, including the use of liposomes; nanoparticles; calcium phosphate (Graham et al. (1973). Virology 52(2):456-67, Bacchetti et al. (1977) Proc Natl Acad Sci USA 74(4):1590-4 and Kriegler, M (1991). Transfer and Expression: A Laboratory Manual. New York: WH Freeman and Company. pp. 96-97); dendritic polymers; or cationic polymers such as DEAE-dextran or polyethyleneimine. Non-chemical methods include electroporation; acoustic perforation; and optical transfection. Particle-based transfection includes the use of gene guns, magnet-assisted transfection (Bertram, J. (2006) Current Pharmaceutical Biotechnology 7, 277-28). Viral methods can also be used for transfection. mRNA delivery includes the use of TransMessenger. TM and The method (Bire et al. BMC Biotechnology 2013, 13:75).

[0136] A common method for introducing heterologous DNA into cells is calcium phosphate precipitation, as described by Wigler et al. (Proc. Natl. Acad. Sci. USA 77:3567, 1980). DNA introduced into host cells via this method often undergoes rearrangement, making this procedure suitable for co-transfection of single genes.

[0137] Polyethylene-induced fusion of bacterial protoplasts with mammalian cells (Schaffner et al., (1980) Proc. Natl. Acad. Sci. USA 77:2163) is another suitable method for introducing heterologous DNA. Protoplast fusion protocols often produce multiple copies of plasmid DNA integrated into the mammalian host cell genome, and this technique requires the selection and amplification of markers on the same plasmid as the GOI.

[0138] Electroporation can also be used to directly introduce DNA into the cytoplasm of host cells, as described by Potter et al. (Proc. Natl. Acad. Sci. USA 81:7161, 1988) or Shigekawa et al. (BioTechniques 6:742, 1988). Unlike protoplast fusion, electroporation does not require the selection marker and GOI to be on the same plasmid.

[0139] Other reagents suitable for introducing heterologous DNA into mammalian cells, such as Lipofectin, have been described. TM Reagents and Lipofectamine TM Reagents (Gibco BRL, Gaithersburg, Md.). Both of these commercially available reagents are used to form lipid-nucleic acid complexes (or liposomes), which, when applied to cultured cells, facilitate nucleic acid uptake into the cells.

[0140] In one embodiment, the introduction of one or more polynucleotides into cells is mediated by electroporation, intracytoplasmic injection, viral infection, adenovirus, lentivirus, retrovirus, transfection, lipid-mediated transfection, or via nucleofection. TM Mediator.

[0141] The methods used to amplify GOI are also required for recombinant protein expression and typically involve the use of selection markers (as reviewed in Kaufman above). Resistance to cytotoxic drugs is the most commonly used characteristic for selection markers and can be a dominant trait (e.g., usable independently of host cell type) or a recessive trait (e.g., applicable to specific host cell types lacking any of the selected activities). Several amplifiable markers are suitable for the expression vectors of this invention (e.g., as described in Sambrook, Molecular Biology: A Laboratory Manual, Cold Spring Harbor Laboratory, NY, 1989; pp. 16.9–16.14).

[0142] Optional markers suitable for gene amplification in drug-resistant mammalian cells are shown in Table 1 above by Kaufman, RJ, ibid., and include DHFR-MTX resistance, P-glycoprotein and multidrug resistance (MDR) – various lipophilic cytotoxic agents (e.g., adenomyomycin, colchicine, vincristine) and adenosine deaminase (ADA) – Xyl-A or adenosine and 2'-deoxycofromycin.

[0143] Other dominant selectable markers include antibiotic resistance genes derived from microorganisms, such as neomycin, kanamycin, or hygromycin resistance. However, these selectable markers have not been shown to be amplifiable (Kaufman, RJ, ibid.). Several suitable selection systems exist in mammalian hosts (Sambrook, ibid., pp. 16.9–16.15). Co-transfection protocols using two dominant selectable markers have also been described (Okayama and Berg, Mol. Cell Biol 5:1136, 1985).

[0144] Useful regulatory elements previously described or known in the art may also be included in nucleic acid constructs for transfecting mammalian cells. The chosen transfection protocol and the elements selected therein will depend on the type of host cell used. Those skilled in the art are familiar with many different protocols and host cells and can select an appropriate system for expressing the desired protein based on the requirements of the cell culture system used.

[0145] Other features of the invention will become apparent during the following description of exemplary embodiments, which are given for the purpose of illustrating the invention and are not intended to limit the invention.

[0146] Example

[0147] The following examples are provided to describe to those skilled in the art how to construct and use the methods and compositions of the present invention, and are not intended to limit the scope of the invention. Efforts have been made to ensure the accuracy of the figures used (e.g., quantities, temperatures, etc.), but some experimental errors and deviations should be taken into account. Unless otherwise specified, parts are parts by weight, molecular weights are average molecular weights, temperatures are in degrees Celsius, and pressures are atmospheric pressure or close to atmospheric pressure.

[0148] Example 1. Identification of loci of interest and characterization of integration sites.

[0149] CHO K1 cells were transfected with two plasmids containing antibody sequences and optional antibiotic resistance genes as markers. Stable transfections were selected by expanding cells in the presence of antibiotics. Sorting techniques are used to separate individual cell clones expressing high levels of antibodies (see U.S. Patent No. 8,673,589,B2). Several clones exhibiting the highest antibody expression levels are then identified.

[0150] Using Covaris Adaptive Focused Acoustics (AFA) TMThe technique fragmented the genomic DNA from these clones (Fisher, S. et al. 2011, Genome Biology 12:R1). DNA libraries (Agilent SureSelectXT#G9612A) were generated and cultured using custom-designed biotinylated RNA decoys (Agilent SureSelectXT#5190-4811) tailored to the entire plasmid sequence introduced into CHO cells. Genomic DNA fragments containing the plasmid sequence were enriched with streptavidin magnetic beads and subjected to Illumina MiSeq sequencing to identify plasmid integration sites. Fusion sequences containing the plasmid sequence and the CHO genome sequence were analyzed and aligned to the CHO genome. Individual integration sites were confirmed by Southern Ink Dot analysis and subsequent PCR sequencing. Integration sites with the nucleotide sequence of SEQ ID NO:1 were identified as expression hotspots (see also GenBank locus ID AFTD01150902.1, nt35529:39558). Integration sites were analyzed to determine their suitability for further cell line generation. The desired outcome is that the integration site is located in a non-coding region, so as not to disrupt normal cellular genomic mechanisms (such as protein translation) or alter the cell phenotype.

[0151] Alignment with SEQ ID NO:1 using a BLAT search (Kent WJ., BLAT - the BLAST-like alignment tool. Genome Res. April 2002; 12(4):656-64) revealed very low homology with mouse and human genome sequences. A sequence blast analysis of SEQ ID NO:1 relative to the CHO-1[ATCC]_refseq_transcript (www.chogenome.org) revealed that the identified locus sequence did not contain any coding regions of any known genes. A broader sequence, SEQ ID NO:4, which encompasses SEQ ID NO:1, was also identified as a suitable locus for targeted integration.

[0152] The integration site sequence was determined to be located in the non-coding region of the CHO and mouse genomes and was further used in experiments described below.

[0153] Example 2. Exogenous DNA efficiently integrated into the host cell integration site

[0154] The exogenous gene was targeted and inserted into a specific locus in the CHO genome identified as SEQ ID NO:1 using the TALE nuclease (TALEN). TALEN targeted a construct containing antibody heavy and light chain sequences randomly integrated into the cell genome, as shown in Example 1. TALEN targeted locations within three identical Hyg genes of the antibody expression construct (see Figure 1A). The TALEN-targeted cleavage sites of the Hyg sequences were based on ZiFit.partners.org (ZiFit Targeter version 4.2). TALEN was designed based on a known method (Boch J et al., 2009 Science 326:1509-1512).

[0155] The donor mKate vector (see Figure 1B) and the TALEN-encoding vector were transfected into CHO host cells using a standard liposome protocol (LIPOFECTAMINE, Life Technologies, Gaithersburg, Md.). Cells were cultured and stable clones with the desired characteristics were isolated and sorted by FACS. Single integration at the desired locus was confirmed by Southern Ink Spot and PCR.

[0156] Example 3. Engineered cells undergo targeted recombination at the locus of interest via RMCE.

[0157] CHO cell lines expressing high levels of fluorescent genes (e.g., mKate) were selected for isolation, wherein the gene was side-linked to a lox site within the locus of interest. A second CHO cell line expressing a second fluorescent gene (dsRed) was selected, wherein the gene was side-linked to a lox site located within the control locus (i.e., EESYR) (U.S. Patent No. 8,389,239B2, issued March 5, 2013).

[0158] Transfected CHO cells were adapted for suspension growth in serum-free production medium. Cells were then transfected in 10 cm plates with a donor expression vector and a plasmid encoding Cre recombinase. The donor expression vector contained the gene of interest encoding the Fc fusion protein flanked by a Lox site (see [link to original text]). Figure 3A (Or 3B). Cells were cultured for two weeks in medium containing 400 μg / ml hygromycin after transfection, and cells expressing eYFP but not mKate (or dsRed in the case of EESYR locus integration) were isolated using flow cytometry. Cells expressing eYFP were expanded in suspension culture in serum-free production medium, and the mRNA levels of each cell pool encoding the Fc fusion protein were determined by qRT-PCR using a standard procedure (see Figure 4).

[0159] The recombination exchange efficiency between cell pools was compared (the percentage of surviving cell populations that exchanged expression donor cassette markers, i.e., eYFP, for expression red markers, i.e., mKate or dsRed) (Table 1). High recombination exchange efficiency was observed at each locus.

[0160] Table 1: Recombination Efficiency

[0161]

[0162] A higher transcription rate (1.5-fold higher) was observed in cell pools containing the engineered LOCUS1 locus (compared to the control locus) (Figure 4).

[0163] The scope of this invention is not limited to the specific embodiments described herein. In fact, various modifications to the invention will become apparent to those skilled in the art from the foregoing description and drawings, in addition to those described herein. It is intended that these modifications fall within the scope of the appended claims. sequence list <110> Regeneron Pharmaceuticals, Inc. <120> Novel CHO integration sites and their applications <130> 32353 (T0045US01) <150> 62 / 067,774 <151> 2014-10-23 <160> 6 <170> PatentIn version 3.5 <210> 1 <211> 4001 <212> DNA <213> gray hamster <400> 1 ccaagatgcc catcaactga ttaatagatg ataaaattat tgtacatttc agtgtaatat 60 tattcagttt ttaagaaaaa tgaaattatg taataagcat gtaaatggat atatcttgaa 120 acaaccattc cccattat tacctaaaca ttgaaagtcc aaaatcatat gatcttttta 180 gtggatctac taatctttg ctatatgtat tttattgaac tacccatgga tgtgagataa 240 ttggtaacaa cagcacatgg gagagcatgg gatcattca ggagattag agagaatgca 300 ttttttagga gatatggag gagcaataga aaggattaaa tgaggttact gatgaagtg 360 atggttagag aaggcaat gaggagggat aactaccact taggggcttt tgaaaaagac 420 atagagaaaa tactattgta gaaacttccct atattggtg tatagttata tacaccaaag 480 agctcagatg gagttaccct ataatggaaa tattaactac ttttcac tgtgataaaa 540 catcctgaac agagcacat agattgggaa gcatttactt tggctcag ttctaacggg 600 aaaaatttc aatgaatg aatgaatg tcagcaaca gcagtagca tggcctgaga 660 agcaggtgag agctcacatc tgaagtgta agaatgtagc aggagaaca aactgcaaat 720 gaccagaaaa tgcttttgga tcagagcccca taccctctg actgacttct ccagaaattc 780 tgaacaaata aaactcccca aacagagcca taactgaagg tccagtgtct gagactacta 840 ggggtatttc ttattcaac cactacaatg gggtgggggg agcaatcctc caagtaggca 900 ctacacacag acaaataaaaa actctagtaa ctggaatgga ttgacttatt tgaattactt 960 gccagtggag ctacatagag cacaattat gtatttaaat taccctttat gatcttacaa 1020 aacttgacag taagatcata ttgctaaaga aaccacatat ttgaatcagg gaacatggtg 1080 atatctagtt gttcttcaac tggaaacttc atgctttctg cccagcattc atgttgctgg 1140 aaagagcaat gtacactacc agtgtagaaa ttaaatcatc aatcttatca agatgtggat 1200 cctataagtt acaataaaaa ttagcctgat aagatatccc caccagaaga atattcacat 1260 aaatgctatg ggagcaacaa gctattttct aaattagctt taatcctatt ctacaagaga 1320 gaatccatat ctagaatagt tatagggatc aagaacccat ggcttgattg gtcataggcc 1380 caatgggaga tcctaatatt attgttctac aaaatgaaaa taactcctaa tgacttgttg 1440 ctgcagtaat aagttagtat gttgctcaac tctcacaaga gaagttttgt cttacaataa 1500 atggcaatta aagcagcccc acaagattta tatcataccg atctcctcat ggcctatgca 1560 tctagaagct aggaaacaaa gaggacccta agagagacat acatggtccc cctggagaag 1620 gggaaggggg cagacctcc aaagctatt gggaggcatgg gggagggggg agggagttag 1680 aagaagaga agggataa agggaggagagaggacaag aggagagagg aagatctagt 1740 chaaggaggaggaggaggaggaggaggagagagaggaggaggaggc cttgtatgtt 1800 taaatagaaa actggcacta gggaattgtc caagatcca caagtccaa ctaataatct 1860 aagcaatagt cgagaggcta ccttaaaagc ctttctctga taatgagat gatgactacc 1920 ttatatacca tcctagagcc ttcatccagt agctgatgga agcagaagca gatactaca 1980 gctaacact gagctagttg cagacaggga ggagtgatga gcaagtcaa gaccaggctg 2040 gagaaacaca cagaaacagc agacctgaa aaaatgttgc acatggaccc cagactgata 2100 gctgggagtc cagcatagga ctttctaga aaccctgaat gaggatatca gtttggaggt 2160 ctggttaatc tatggggaca ctggtagtgg atcatattt atccctagtt catgactgga 2220 atttgggtac ccattccaca tggaggattt ctctgtcagc ctagacacat gggggaggtt 2280 ctaggtcctg ctccaaataa tgtgttagac tttgagac tcccttgaga agactcaccc 2340 tccctgggga gcagaaaggg gatgggatga gggttggtga gggacaggag aggagggg 2400 ggtgagggaa ctgggattga caagtaaatg atgcttgttt ctaatttaaa tgataaagg 2460 aaaagtaaaaaaaaaaacaggcca aaagattata aagacagag gtggtgggtg 2520 actataaaga aacactatta tctaataa aacatgtcag agcacacat gaactatag 2580 tgtttatgaa agtatgtata attackacat atctcaagc CAagaaaaaa attackcatctt 2640 tcagtgatga aggtgatttt atttcccca gattaaagc caagaccta atgaaagtaa 2700 ttatctcaa aaggttgaaa atacatactt tgcaatacac agatctgcct agaaatctca 2760 tgttcacaat accatgatg ctcaattgaa ttccattcaa tgttacagtt tagataaca 2820 gtttgtagat aactcacaa tgtatcatttt ctttttatt tttgaccaaa cagctctca 2880 tctgttattc agaataatc ctcgatggca ggatatccat cccaattggg ggaggggag 2940 aatttgaaga aacctagac cacatacata tttgccattg ggaacaag tctaaaatga 3000 tgttgttcac atcttctcta ctagtccctct ccccgtccca aagaaccttg gtatatgtgc 3060 ctcattttac agagagagga aagcaggaac tgagcatccc ttacttgcca tcctcaaccc 3120 aaaatttgca 3180 caaagtccct ctgtcatgtt tgtgtcaat agttataca gatgacttca tgtcttcata 3240 tctaatgtct tatatagatt atattaaac atgttattt ctctaaccac attttaatt 3300 aatttaaaaa tccattaatt gtgtcttaa atgcagaca gagtgctgag acacaata 3360 agcctgatga tctgaatttg aactcacac ccaccacatg gagaatcaac ttccaaaat 3420 tttcctatta cttcacact cacattg spider ataataatga acaaaatgaa 3480 atgaataaaaattaagtc tctgtaggta atgctactgt gcagcaaag taaaatggc 3540 agcttaagct tgctttatgg ttacacttta ccatcttcca ttaattataa ggacttcaat 3600 catggcagaa ctatgctgtt attgtctcag tgtacctaa ccaggtgttc cagatgttct 3660 taatgtggac acctaaacta tttgatattt gggttaagat ctttccctct ttcagaagaa 3720 acctcaggac agagggaatc ttgtctttta attttgagtc tgtagacttt ttccattca 3780 aatatacatg aaacaagtga tgaagaaaat taatcaaaag gtgggaattg caatgatatt 3840 aggttcaata ttaagcttca atattatcat ggaatcgcct gttatacact gagtgtttgg 3900 caataaggga tttttagaag aaggagtttt tattctcaac aggttcctta agtttagctc 3960 aaataaatct aagcaatcca ctctagaatt aaatagtttc c 4001 <210> 2 <211> 1001 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Polynucleotide <400> 2 taccctttat gatcttacaa aacttgacag taagatcata ttgctaaaga aaccacatat 60 ttgaatcagg gaacatggtg atatctagtt gttcttcaac tggaaacttc atgctttctg 120 cccagcattc atgttgctgg aaagagcaat gtacactacc agtgtagaaa ttaaatcatc 180 aatcttatca agatgtggat cctataagtt acaataaaaa ttagcctgat aagatatccc 240 caccagaaga atattcacat aaatgctatg ggagcaacaa gctattttct aaattagctt 300 taatcctatt ctacaagaga gaatccatat ctagaatagt tatagggatc aagaacccat 360 ggcttgattg gtcataggcc caatgggaga tcctaatatt attgttctac aaaatgaaaa 420 taactcctaa tgacttgttg ctgcagtaat aagttagtat gttgctcaac tctcacaaga 480 gaagttttgt cttacaataa atggcaatta aagcagcccc acaagattta tatcataccg 540 atctcctcat ggcctatgca tctagaagct aggaaacaaa gaggacccta agagagacat 600 acatggtccc cctggagaag gggaaggggg caagacctcc aaagctaatt gggagcatgg 660 gggaggggag agggagttag aagaaagaga aggggataaa aggagggaga ggaggacaag 720 agagagaagg aagatctagt caagagaaga tagaggagag caagaaaaga gataccatag 780 tagagggagc cttgtatgtt taaatagaaa actggcacta gggaattgtc caaagatcca 840 caaggtccaa ctaataatct aagcaatagt cgagaggcta ccttaaaagc ctttctctga 900 taatgagatt gatgactacc ttatatacca tcctagagcc ttcatccagt agctgatgga 960 agcagaagca gacatctaca gctaaacact gagctagttg c 1001 <210> 3 <211> 1001 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Polynucleotide <400> 3 haaagtcaag accaggctgg agaacacac agaacagca gacctgaaa aaatgttgca 60 catggacccc agactgatag ctgggaggtcc agcataggac tttctgaa accctgaatg 120 aggatatcag ttggaggtc tggtatct atgggacac tgtagtgga tcaattta 180 tcctagttc atgactggaa ttgggtacc cattccacat ggaggaattc tctgtcagcc 240 tagacacatg ggggaggttc taggtcctgc tccaaataat gtgttagact ttgagaact 300 cccttgagaa gactcaccct cccttgggag cagaaagggg atgggatgag ggttggtgag 360 ggacaggaga ggaggggagg gtgagggac tgggattgac aagtaatga tgcttgtttc 420 tatttaat gataagga aagtaaag aaaaaaaaaaggccaa agattaa 480 aagacagagg tggtgggtga ctataaagaa acacttatta ctaataaaa atatgtcaga 540 agcacacatg aacttatagt gtttatgaa gtatgtataa taactacata atctcaagcc 600 aagaaaaaaa tatcatctt cagtgatgaa ggtgattta tttctcccag aattaaagcc 660 aaagacctaa tgaaagtaat tatctcaa aggttgaaaa tacatacttt gcaatacaca 720 gatctgccta gaaatctcat gttcacaata cacatgatgc tcaattgaat tccattcaat 780 gttacagttt agataaacag tttgtagata aactcacaat gtatcatttc tttttatttt 840 ttgaccaaac agcttctcat ctgttattca gaataattcc tcgatggcag gatatccatc 900 ccaattgggg gaaggggaga atttgaagaa aacctagacc acatacatat ttgccattgg 960 gaaacaaagt ctaaaatgat gttgttcaca tcttctctac t 1001 <210> 4 <211> 14931 <212> DNA <213> Cricetulus migratorius <220> <221> misc_feature <222> (2176)..(2239) <223> n is a, c, g, t or nucleotide deletion <400> 4 catgtacact tatgcaagta tgatatggcc caacacagta ttttacacca atttttatct 60 ataaaatata catgtacatc aaaatatatt attaataata acatcattat tctttctttc 120 caagtaataa acacatacac tgaaattttg gttcttgtgg ataattttaa tgaaacagga 180 aatgcaaatt tatcttagca tgtttacttc actttctttg catagataac cagtaatcac 240 attgatggat catgtagtga aatgtatttt taggtatcta aggaatttg gcttcgtttt 300 gtgcttgttg acactgaatt ctattcctaa cacagtgtg taaggattct gtctgatttc 360 ttttaccagt atttgtccat ttgcattttc atttattc atggctgctg ttcagaaag 420 tggaaggtag tgtgtcaagt ctgtttaca tgttccctg atgatcagtg tcttacacc 480 tctctgagta catgttggcc aatgtcgttt ctagacccat ctattcttgc ttgacttatc 540 ctggtacatg cctgccaaga aatttctct catcctttct gtctctcac tgatttactt 600 gatgtgtgga ttcacattg atcatatgga atagagat acaatttct ttattcacag 660 tttggaagac ttcaatctc atagatcatc attatttt gctactgttc cctatgctat 720 ggtgaaattt ccatttgaat aattgcttaa acaattaca agaaagaatc tatttttact 780 tgcaataact tccattcag aacatttact acactgttac tatatccaa aactagtttt 840 atatatcatg tgagaaatga ctaattcata atttggccat gatatttt tcagaaacag 900 aaaagtgac caatacatac acatgctat aaatattaag acttcagcaa atttaatatt 960 tattcatgat atcacataaa attcatttat tatgtttat ttaaatgtgt ttttaaaaca 1020 gtggtatcac taaatattaa gttagatgtg tttagtgct taatgaattt atattttaga 1080 atgttataag ttgtatatag tcaaatatgt aaaattttt atttttagg tctttctcat 1140 taaggtattt taatttggg tccctttcc agagtgactc tagctcatga tgagttgaca 1200 taaaaactaa acagtacaaa atgtacattg cattcagtat tgcactgat ctttgcactg 1260 aagtttgagt cagttcatac atttagtact tgggagtac attaagcta ctttcattgc 1320 tctggcaaaa tgctcgataa gataagagtc tattgtggaa agccatggca gcaggaagt 1380 aagactgctg atgatgttta atccatagtc aagacgcaga aggagatgaa tgctggtatc 1440 caacatttt tgctgttcat tttcttaga accctagtcc ataagatgt atgacttgca 1500 ttcaaatgc gtccccttca gttgttcac ttttctgta atatcctttc aggcatgtct 1560 agaagattgt ttcgcaata cttctcaatc cattcaagtt gatagtgcag attackcact 1620 gcagataaa agcctgtaac tggctcacg tgcaggaa tatgcacact cctgacacat 1680 caataagtaa atcaaagtgt agcttttgcc tttaacattg ccagacttat gtaatgttct 1740 gcacgttctt cctccatcac ttttattct aatggtgttt ccttgacatt gaatcacgct 1800 gtggaagctg cttagaatta acattgaaat ctactgatat attatgatg cagcaattta 1860 gatttactat tttacttaga attttttata attgagagaa tataatattt tcacagttat 1920 ctatctgctg taatagagg attttaaaaa aaatctctat aacttttttt tacaacacac 1980 agtaaaatta agttaaaatt tataaagtc actatgttga tttcaaagtg tgctacgccc 2040 acggtggtca cgcaggtgta gcagaagatg ccactaaggt gggctaaggc cgatgggttg 2100 gggtctgcgc tccctggaga tgagccccag gcggttccct ggcaatcagc tgcgatcatg 2160 atgcccgatg agccannnnn nnnnnnnnn nnnnnnnnn nnnnnnnnn nnnnnnnnn 2220 nnnnnnnnn nnnnnnnnc tgggtgactt tatggaaaga atttgataga tttcatgatg 2280 tagaagaatt ttattaggct tattttacag gagactaaga ccctgggacc taaagatatc 2340 tgggtcctga gaatcaggaa atgggtagag acgtggttga tggtatgaga cagattttag 2400 agaactctta gatcatgggc aatgaccgca atctgatgct tagatagat catctataaa 2460 caattatgct gttctttttc tttctgttgt atgatctgat gatgtagccc ccttgccaag 2520 ttccctgatc ccccttgcca agttccctga ttgtaacagt atataagcat tgcttgagag 2580 catattcaac tacattgagt gtgtctgtct gtcatttcct cgccgattcc tgatttctcc 2640 ttgagcctt tcccttgttc tccctcggtc ggtggtctcc acgagaggcg gtccgtggca 2700 aaagtgtata aatgttctaa aacatttgaa ctctaaaaca tgcaaaatga aaaattaaaa 2760 taaataaca tgaaaattaa aatatattag ctgctaaaag ttaaacaata ctatatata 2820 ttttgttatt agaattcaaa atcacattag ttggatttaa tttgaacatt gcattctttc 2880 aataataatt tcaataaaaa aagtttcccc atgatagtag aaaataataa catatgtatc 2940 tatctattta tttaactaca catatatagc atttgtttca actaaaataa atgaatgagc 3000 aaagcaccta agtaattggt gtctattata tttatgaagc caatagtttc aaataaatta 3060 tcatgcataa ggaggtattg caaatgttaa accttttttg aaacagatat tcccagttac 3120 agaaattata atttctaatc tttcctataa gtagaatgat gataattaat ataggccatt 3180 tgtaaataat gttcagatta aaatattctc tattcacta gagaagaatg atattaatg 3240 tattatattt tattcccat ttgtttgca tattcctct attackccctca gcagtttaaa 3300 tttgttcac catatgtgtg tgtgtttgta tcttaatat ggcactaaaa ttagaataat 3360 daddyaa tctttagg aaagattatt gatttatttt atgttgatag gaaatatct 3420 tttaattgtc ciegaatact ttttctcta ttttaggact gatcagaccc aggactaata 3480 ttttatatgt actaattcta tgtaccaaaa tatgttatta tctcatgaat tctgtctcaa 3540 tattgaggta aaaaatag tccatcatga actttaaaat taaaataatg attattaat 3600 ttttattcat atttgttg tatgaatggt tatacatcac atgtgtgcct ggtgactgtg 3660 aatgtcagga gaggtag aagccactgg aattggaata agagatata ttgagatgt 3720 tatgtgggtg ctgagaatta gacgcaagcc atcttcaaga atagccagca tactatacca 3780 ctgagtaatc cattcatccc tcaatatta tctttgtaga cagtaaatat atttctaaac 3840 tataaatgac cagaaaaatt aatgtattat taatgaagac attcatctca tgtgacacac 3900 ttcacctgtc taaatcagta acactctctc cactaattaa gattttctaa gtgcatgaca 3960 cttactattt ctaaagctgt ccaatggggg ccagtcccca gtcagcaccc agtgagataa 4020 tccatgaatg catttatatc ttaggaaaaa ttcttatcta tgtagtattt agaacatttt 4080 catgtgaggg gataaacaag gaagcacaga tgctttctga tagaaacttt ctctttaatt 4140 catctagaaa aaaaaaacct ctcaggaaaa tctctcttgc tctcctccca atgctctatt 4200 cagcatcttc tccctactta attctagatc tttttctcta tgcctccttg ctgctgccct 4260 gctggctctg ctctatgcct ccccatgtca cttttctttg ctatctcacc gttaccttct 4320 ctgcctcact ctctgccttc ttctctgctt ctcacatggc caggctctgg acaattatag 4380 ttatatgtta cattctcata acacatgata tgtcacatag tttctctcag gctagggata 4440 tcacaatgac tggccaatga gcaagtggcc ttgcatgtag ctctaagttg gtgatggttc 4500 ccagacagta agtagccatt tggttgaaat ttgaggttgg gtagtacatg aagactgaat 4560 tttcttcaaa ctctggcctt gaaatagtaa aacaacacct atgaaaatga cgacctgtat 4620 ttgtctttag aggcaaccac atattgtctg cagggcctgc tttgaatttg ctctgaagtt 4680 agcttgtttg tgtaaaagga agaatcctat atcagcctga gaaatgtaaa atatcctagc 4740 atttcaagtc atcaaaatta tatggagagt ataaatcatc cttctgacta ttcatagtca 4800 tatttgtgtc caccaagtat aaaacacact accaaagggc tgtggaaaaa atcgccataa 4860 ctgttcttat tagggaggca tagcagtggt acctgaggaa gttacagcaa caaccagtca 4920 tccagtcaat aaccccatgg ctttgccact tggaggtacc caataatgtt tggctttgcc 4980 gagtaggact ccaacaaatt cagagggtca atttttaaat gctggttgtc actgctgaac 5040 agtcccattg ccctctgcat aattccacaa tggaaagctt tttacactga ttgccaatca 5100 ttaaacagcc tactcagcat aaacaggtat gatattattc tgcattttgt tacattacta 5160 gatgaattcc tatttcttcc tacaatagtg gaactgaaaa aagatacaca atcatactac 5220 ccctctacta atcttatgac ttatatcatt tcaattttca gaccataatg caaactattg 5280 accaaaacat gtgagatga aaatagaa tgtagaata tattacatat aaaaagaaaa 5340 ggcggactta tttgtttta ttcttagca tgcatagca tacatgatt gaggtttata 5400 tataaaggg acaataatc ttcaagaac tacccctac tgaattaaaa tattaaagaa 5460 ggtcacacat ttactcaaat atattagact actgggcaaa tagacatgaa aagtagagtt 5520 atattgagg taggccttct gtgaaatgtc taggaatt atgtttcata cagtgtgtaa 5580 ccaagtggga atcatatcag aaagcagtca aaagcttata ttacaagtaa cagatgcttg 5640 gttatatgac ctcccagagc ttgactgtct atacacaaaa agtgtgtta ataaaactgt 5700 aatttgggct atgttttttt aaatggctc accacatga aaggaaggga atgagcatgt 5760 catggatgct tagagatt gcttccagca agagaattg agctttggct cttattacag 5820 aaacatgaca aggtgtgagt tttatttat agaaattata tatatttta agctggggac 5880 taaaaattttt attgaaaaaaaaaagg gataggcatg tactgaagc aaaaatagga 5940 tgtcaatgct gtaatgttat tttggacc aaatagtat ttcctataga atgacaatg 6000 atcttaggtt attattctc aataagatga caagttcaca agatatccta gttcattaaa 6060 atcgttttag tcatttaata gagtgctgtg atagattaca caaggaaag cacttacgat 6120 gagaaataat gatatccaca attattttct taattcttag aacattcta ttgttatatc 6180 tcaatctcag aagccactta ttgctttatt attgaaacat atgaattgt aagttatata 6240 ttgtctatgg tgacattca aagaacatgt gacgtacagt gtagcacaga taagaacat 6300 aactgcagct gatcagtaa ctaaacttac atacattaaa tctgccatgt tggcacagt 6360 gtgtgcacta ccaaggatg tactaatgct cacgacactc ccctatgtca ccctttgttc 6420 atcattacat cataggtcta tttgtttgc ttttgaatc tagaccagt cttttgtgtc 6480 tttccaagca cagagctcat taatttacct catagacttg tttaacttct tctggttcat 6540 aattgaata gaatactca ctactaatta tgtgagaccc tgccagtacc atagcacatg 6600 gataatttt acataaaca tgcatacaag windowgatt cagactgaac atgaattttta 6660 gagaaatcag gaggagtat atgggagtgg tggagtgag actagagaaa tgtattaaa 6720 ctataatctc atacaagc tctaagc aaaaaacatg aaacattgtc attcaagtga 6780 aacatcagtc ttcaattgg aagatattt ttactaggaa atgtctggt agatggttat 6840 tacttagaaaaaaaat tagaaaacgg taacttta taaaaagaat atacaatga 6900 gactacatga aaagttctta actaatgaaa caataatctt gaactttt tcttaaagt 6960 ttaatca taaccatcat ggaattcaa atttaaacta tttacatatt acccctgaaa 7020 taataactaa tacctaa aaatatata aaaaaaat ggcaatgcat gccatcatgg 7080 atttgggaga gagaatgttc attgcagttc tgaatggata ctggtgccac cacggtgaaa 7140 atctctgtat aggtccttcc aaagctgaa atagacata tcacagacc tgccacat 7200 tttcaagca ataccaccaa ggactctacc tgactgcaga gandactttct cataaaatat 7260 tattgttgat ctattcataa tattctggaaa atagaacag ccagatgcc catcaactga 7320 ttaatagatg aaaattat tgtacatttc agtgtatatat tattcagtttt ttaagaaaaa 7380 tgaaattatg taataagcat gtaaatggat atatcttgaa acaaccattc cccattatat 7440 tacctaaca ttgaagtcc aaaatcatat gatcttttta gtggatctac taatctttg 7500 ctatatgtat tttattgaac tacccatgga tgtgagataa ttggtaacaa cagcacatgg 7560 gagagcatgg gatcattca ggagattag agagaatgca ttttgga gatatggag 7620 gagcataga aaggattaaa tgaggttact gatgaagtg atggttagag aaggcaat 7680 gaggagggat aactagcact taggcctttt tgaaaaagac atagagaaaa tactattgta 7740 gaaacttccct atattggtg tatagttata tacaccaag agctcagatg gagttaccct 7800 ataatggaaa tattaactac ttttcac tgtgataaaa catcctgaac agagcaacat 7860 agattgggaa gcatttactt tggctcag ttctaacgggg aaaaatttc atgatgaaag 7920 aatgaatatg tcagcaaca gcagtagcaa tggctgaga agcaggtgag agctcacacatc 7980 ttgaagtgta agaatgtagc aggagaaca aactgcaat gaccagaaaa tgctttttgga 8040 tcagagcccca taccccctg actgacttct ccagaaattc tgacaata aaactcccca 8100 aacagagcca taactgaagg tccagtgtct gagactacta ggggtatttc ttattcaac 8160 cactacaatg gggtgggggg agcaatcctc caagtaggca ctacacacag acaaataaaaa 8220 actctagtaa ctggaatgga ttgacttatt tgaattactt gccagtggag ctacatagag 8280 cacaattatt gtatttaaat taccctttat gatcttacaa aacttgacag taagatcata 8340 ttgctaaaga aaccacatat ttgaatcagg gaacatggtg atatctagtt gttcttcaac 8400 tggaaacttc atgctttctg cccagcattc atgttgctgg aaagagcaat gtacactacc 8460 agtgtagaaa ttaaatcatc aatcttatca agatgtggat cctataagtt acaataaaaa 8520 ttagcctgat aagatatccc caccagaaga atattcacat aaatgctatg ggagcaacaa 8580 gctattttct aaattagctt taatcctatt ctacaagaga gaatccatat ctagaatagt 8640 tatagggatc aagaacccat ggcttgattg gtcataggcc caatgggaga tcctaatatt 8700 attgttctac aaaatgaaaa taactcctaa tgacttgttg ctgcagtaat aagttagtat 8760 gttgctcaac tctcacaaga gaagtttgt cttacaataa atggcaatta aagcagcccc 8820 acaagattta tatcataccg atctcctcat ggcctatgca tctagaagct aggaaacaaa 8880 gaggacccta agagagacat acatggtccc cctggagaag gggaaggggg caaccctcc 8940 aaagctaatt gggagcatgg gggagggg agggagttag agggaga agggataaa 9000 agggagggagggagggagggagggagggagggagggagggagggagggagg 9060 caagaaaaga gatacatag taggaggagc cttgtatgtt taaatagaa actggcacta 9120 gggaattgtc caagatcca caagtcca ctaataatct aagcaatgt cgagaggcta 9180 ccttaaaagc ctttctctga taatgagat gatgactacc ttatatacca tcctagagcc 9240 ttcatccagt agctgatgga agcagaagca gatactaca gctaacact gagctagttg 9300 cagahaggga ggagtgatga gcaagtcaa gaccaggctg gagaacaca cagaaacagc 9360 agacctgaaa aaatgttgc acatggaccc cagactgata gctgggagtc cagcatagga 9420 cttttctga aaccctgaat gaggatatca gtttggaggt ctggttaatc tatggggaca 9480 ctggtagtgg atcatattt atccctagtt catgactgga atttgggtac ccattccaca 9540 tggaggaatt ctctgtcagc ctagacacat gggggaggtt ctaggtcctg ctccaaataa 9600 tgtgttagac tttgagac tcccttgaga agactcaccc tccctgggga gcagaaaggg 9660 gatgggatga gggttggtga gggacaggag aggagggg gggtgagggaa ctgggattga 9720 caagtaatg atgcttgttt ctaatttaaa tgaataagg aaagtaaa gagaaaaga 9780 aaacaggcca aagattatta aagacagag gtggtgggtg actataaaga aacactatta 9840 tctaataaaatgtcag agcacacat gaactatag tgtttatgaa agtatgtata 9900 ataactacat aatctcaagc caaaaaaa attackatctt tcagtgatga aggtgattt 9960 atttctcccca gattaagc caaaccta atgaagtaa ttatctca aaggttgaa 10020 atacatactt tgcaatacac agatctgcct agaatctca tgttcacaat acacatgatg 10080 ctcaattgaa ttccattcaa tgttacagtt tagataaca gtttgtagat aaactcacaa 10140 tgtatcatttt ctttttatt tttgaccaaa cagctctca tctgttattc agaataattc 10200 ctcgatggca ggatatccat cccaattggg ggaggggag aatttgaaga aacctagac 10260 cacatacata tttgccattg ggaacaaag tctaaaatga tgttgttcac atcttctcta 10320 ctagtccctct ccccgtccca agaaccttg gtatatgtgc ctcattttac agagagagga 10380 aagcaggaac tgagcatccc ttacttgcca tcctcaaccc aaaatttgca tcattgctca 10440 gctctgccct tctcatatga cagttacaag tcaggctc caagtccct ctgtcatgtt 10500 tggtgtcaat agtttataca gatgacttca tgtcttcata tctaatgtct tatatagatt 10560 atattaaac atgttattt ctctaaccac atttttaattt atttaaaaa tccattaatt 10620 gtgtctataa atgcagaca gagtgctgag acacatata agcctgatga tctgaatttg 10680 aaactcacac ccaccacatg gagaatcaac ttccaaaat tttcctatta cttccacact 10740 tacaccattg tashaacaca attaatga aaaaatgaa atgaataa aattaagtc 10800 tctgtaggta atgctactgt gcagcaaag taaaatggc agcttaagct tgctttatgg 10860 ttacacttta ccatcttcca ttattataa ggacttcaat catggcagaa ctatgctgtt 10920 attgtctcag tgtaacctaa ccaggtgttc cagatgttct taatgtggac acctaaacta 10980 tttgatattt gggttaagat ctttccctct ttcagagaa acctcaggac agagggaatc 11040 ttgtctttta attttgagtc tgtagacttt ttccattca atatacatg aaacaagtga 11100 tgaagaaat taatcaaag gtgggaatttg caatgatatt aggttcaata ttaagcttca 11160 atattacat ggaatcgcct gttatacact gagtgttttgg caataggga ttttagag 11220 aaggagtttt tattctcaac aggttcctta agttttagctc aaataatct aagcaatcca 11280 ctctagaattt aaatagtttc ctaagggcac agctatgaat agagctcaat ttacatataa 11340 aattttgttc accatttg accattczag accattaaaaaaaaat 11400 ttagatgtca atatcaagtg atagttcat ctcctttttt atatatatc acctaaatca 11460 ccattttctc agaaaaatct ggctgaagt tctgtctgga acttcacat gaaaaatatg 11520 cacagcttgc tattataat cctagttgat ttttaagatt catgtctggt gtctgactca 11580 gaggggccag aggctagaca atatttt gatcttcat tgtgaagatt tttaatgatt 11640 attttatat aaatacaa gatgatgat atgtaactt tgtacagttc atagacgctg 11700 aactactttg tgcttaaaat gttagttccc tatcataaat gataggtgat aagtgtagt 11760 ttaatacttt ccctctgagc tatattcatg tactgagaa ttattaaa catgaaaga 11820 ctgtgtttat agtctcagct cctgagaact ggtcaacct taggcaggtg atgccagga 11880 gcaacgtttt tctctacag aggatgcttt gctgccaagc aacctggttg tgtggaatg 11940 ttccttttt aatcaagttt aaagggtctt catcatgctg tgctccaca tattttcagg 12000 ttagagcttg gtccttggag tattacttt taccagaaaa ttcatagtat tctttcaata 12060 actaacaact aaacttttcg aaaaaaga attggaattt caatttaaa gcctgagtaa 12120 aattcttgtg aatcaggata ttttattt agtcttatct tttaaaagt tattttatt 12180 tttaaaaaat tataatatac ttcatatt tccctccttc acttttcttt acaacactt 12240 ctatagatca ccatgtgtttt ttttttac atttatggcc tctttctgtt cattgttatt 12300 acatacaaat agtcttgcct atagagaac accacaattt gttacctgat aacaaattat 12360 cacccttaa aacctacaa ctattgaat tactgaaag actatactta tagtaaa 12420 gatatagtg tgtgcacata tatagataca catatatgta ggatttttaa ttttagattt 12480 tagacatca attattat atgactgaga aactagacac tataaatgag cattcagtat 12540 tcaaccgt gattttagat attgtcacaa tgacagaaaa tttcttata gaaaatttta 12600 agttttgtga ttgctctgtg cacttagtga agtctcacag aaaagaatc atagtatttt 12660 tagtttata taaaaagtac atattaa atggttggc aaaaacac atttgagcat 12720 tttcctatt tactatcaag tagtatcatt ttgaaataat atttgacta gtttcaaaaa 12780 tgaaaacaaa atttaacta atgcctaat ctagcctgat aacatttta tgaatgaat 12840 tattcaatag tgttatcaat taggggcca aaacttttcc taaaataaa cttttaattt 12900 tttccattt tttttaaat tagaacaaa attgttttac atgtaaatca gagttttccctc 12960 accctcccct tctccctgtc cctcactaac accctactg tcccatacca tttctgctcc 13020 ccagggagggg tgaggctc catggggaaa ctcaggagtc tgtctatcct tcggatagg 13080 gcctaggccc tcacccattt gtctaggcta aggctcacaa agtttactcc tatgctagtg 13140 ataagtactg atctactaca agagacacca tagatttcct aggcttcctc actgacaccc 13200 atgttcatgg ggtctggaac aatcatatgc tagtttccta ggtatcagtc tggggaccat 13260 gagctccccc ttgttcaggt caactgtttc tgtgggtttc accaccctgg tcttgactgc 13320 tttgctcatc actcctccct ttctgtaact gggttccagt acaattccgt gtttagctgt 13380 gggtgtctac ttctactttc atcagcttct gggatggagc ctctaggata gcatacaatt 13440 agtcatcatc tcattatcag ggaagggcat ttaaagtagc ctctccattg ttgcttggat 13500 tgttagttgg tgtcatcttt gtagatctct ggacatttcc ctagtgccag atatctcttt 13560 aaacctacaa gactacctct attatggtat ctcttttctt gctctcgtct attcttccag 13620 acaaaatctt cctgctccct tatattttcc tctcccctcc tcttctcccc ttctcattct 13680 cctagatcca tcttcccttc ccccatgctc ccaagagaga tgttgctcag gagatcttgt 13740 tccttaaccc ttttcttggg gatctgtctc tcttagggtt gtccttgttt cctagcttct 13800 ctggaagtgt ggattgtaag ctggtaatca ttgctccat gtctaaaatc catatatgag 13860 tgatgtttgt ctttttgtga ctgggttacc tcactcaaa tggtttctc catatgtctg 13920 tggattcaa tagcacaac aacatacagt atcttggggc aacactacc aacaagtga 13980 aagaccagta tagcagaac tttgagtttta aagaaaaaaagaaga taccagaaaa 14040 tggaagatc tcccatgctc ttgaatagc agaatcaca tagtaaaat ggcaatcttg 14100 ccaaaatcca tctacagact caatgcaatc cccattaaat accagcacac ttctcacag 14160 acctgaaga atatactta actttatatg gagaaaaaa agacccagga taggccaaac 14220 aaccctgtac atgaggca cttccagagg catccccatc cctgacttca agctctatta 14280 taggtaata atcctgaaaa cagcttggta atggcacaaa atagacagg tagaccaatg 14340 gattgagtt gaaaaccctg atattaaccc acatacttat gaacacctga ctttgacaaa 14400 gaagctagg ttatacaatg tagaagaa agcatctca acaaatcgtg ctggcataac 14460 tggatgctgg catgtagaag actgcagata gatccatgtc taatgccatg cacaaaactt 14520 aagtccaaat ggatcaaaaa cctcaacata aatccagcca cactgaacct catagaagag 14580 aaagtgggaa gtatccttga ataaattggt acaggagacc acatcttgaa cttaacacca 14640 gtagcacaga caatcagatc aataatcaat aaatgggacc tcctgaaact gagaagcttc 14700 tgtaaggcaa tggataagtc aacaggacaa aatggcagcc cacggaatgg gaaaagatat 14760 tcaccaatcc tatatctgac agagggctgc tctctatttg caaagaacac aataagctag 14820 tttttaaaac accaattaat ccgattataa agttgggtag agaactaaat aaagaattgt 14880 taacagagca atctaacttg gcagaaagac acataagaaa gtgctcacca t 14931 <210> 5 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Synthetic polynucleotide <400> 5 tgagctagtt gcagacaggg 20 <210> 6 <211> 79 <212> RNA <213> Artificial sequence <220> <223> Synthetic polynucleotide <400> 6 guuuuagagc uagaauagca aguuaaaaua aggcuagucc guuaucaacu ugaaaaagug 60 gcaccgaguc ggugcuuuu 79

Claims

1. An isolated cell comprising an exogenous nucleic acid integrated into a locus in the genome of the cell, wherein the locus consists of a nucleotide sequence of SEQ ID NO: 1, and the locus is capable of enhancing the expression of a protein encoded by the exogenous nucleotide sequence by at least 1.5 to at least 3 times compared to the expression observed in an exogenous nucleotide sequence integrated into another locus in the genome of the cell.

2. The cell of claim 1, wherein the cell is a CHO cell.

3. The cell of claim 1, wherein the exogenous nucleic acid comprises one or more recombinant recognition sequences.

4. The cell of claim 3, wherein the exogenous nucleic acid sequence comprises at least two recombination recognition sequences and an optional marker placed between the two recombination recognition sequences.

5. The cell of claim 3, wherein the one or more recombination recognition sequences are selected from the group consisting of: LoxP site, Lox511 site, Lox2272 site, Lox2372 site, Lox5171 site, Loxm2 site, Lox71 site, Lox66 site, LoxFas site, and frt site.

6. The cell of claim 1, wherein the exogenous nucleic acid comprises a first exogenous gene of interest (GOI) and a first exogenous promoter, wherein the first exogenous GOI is operatively linked to the first exogenous promoter.

7. The cell of claim 6, wherein the exogenous nucleic acid further comprises a second exogenous GOI and a second exogenous promoter, wherein the second exogenous GOI is located at the 3' of the first GOI and is operatively linked to the second exogenous promoter.

8. The cell of claim 7, wherein the exogenous nucleic acid further comprises a first recombinant recognition sequence at the 5' of the first exogenous GOI and a second recombinant recognition sequence at the 3' of the second exogenous GOI.

9. The cell of claim 8, wherein the first and second recombination recognition sequences are distinct and selected from the group consisting of: LoxP site, Lox511 site, Lox2272 site, Lox2372 site, Lox5171 site, Loxm2 site, Lox71 site, Lox66 site, LoxFas site, and frt site.

10. The cell of claim 8, wherein the first exogenous GOI encodes an antibody light chain or an antigen-binding fragment thereof, and the second exogenous GOI encodes an antibody heavy chain or an antigen-binding fragment thereof.

11. The cell of claim 7, wherein the exogenous nucleic acid further comprises a third exogenous GOI and a third exogenous promoter, wherein the third exogenous GOI is operatively linked to the third exogenous promoter.

12. The cell of claim 11, wherein the third exogenous GOI and the operatively linked third exogenous promoter are located at the 3' of the second exogenous GOI.

13. The cell of claim 11, wherein the first, second, and third GOI encode a polypeptide selected from the group consisting of: a first antibody light chain or an antigen-binding fragment thereof, a second antibody light chain or an antigen-binding fragment thereof, and an antibody heavy chain or an antigen-binding fragment thereof.

14. The cell of claim 11, wherein the first, second, and third GOI encode a polypeptide selected from the group consisting of: an antibody light chain or an antigen-binding fragment thereof, a first antibody heavy chain or an antigen-binding fragment thereof, and a second antibody heavy chain or an antigen-binding fragment thereof.

15. A method comprising: An exogenous nucleic acid is introduced into CHO cells, wherein the exogenous nucleic acid is integrated into a locus in the genome, the locus consisting of the nucleotide sequence of SEQ ID NO: 1, and the locus is capable of enhancing the expression of the protein encoded by the exogenous nucleotide sequence by at least 1.5 to at least 3 times compared to the expression observed when an exogenous nucleotide sequence is integrated into another locus in the genome of the cell.

16. The method of claim 15, wherein the exogenous nucleic acid comprises one or more recombinant recognition sequences.

17. The method of claim 16, wherein the exogenous nucleic acid comprises at least two recombination recognition sequences and an optional marker placed between the two recombination recognition sequences.

18. The method of claim 16, wherein the one or more recombination recognition sequences are selected from the group consisting of: LoxP site, Lox511 site, Lox2272 site, Lox2372 site, Lox5171 site, Loxm2 site, Lox71 site, Lox66 site, LoxFas site, and frt site.

19. The method of claim 15, wherein the exogenous nucleic acid comprises a first exogenous gene of interest (GOI) and a first exogenous promoter, wherein the first exogenous GOI is operatively linked to the first exogenous promoter.

20. The method of claim 19, wherein the exogenous nucleic acid further comprises a second exogenous GOI and a second exogenous promoter, wherein the second exogenous GOI is located at the 3' of the first exogenous GOI and is operatively linked to the second exogenous promoter.

21. The method of claim 20, wherein the first exogenous GOI encodes an antibody light chain or an antigen-binding fragment thereof, and the second exogenous GOI encodes an antibody heavy chain or an antigen-binding fragment thereof.

22. The method of claim 15, wherein the exogenous nucleic acid is introduced using a nucleic acid vector, the nucleic acid vector comprising: a. The 5' homologous arm that is identical to the sequence present in the nucleotide sequence of SEQ ID NO:

1. b. The aforementioned exogenous nucleic acid, and c. The 3' homologous arm that is identical to the sequence present in the nucleotide sequence of SEQ ID NO:

1.

23. The method of claim 22, wherein the exogenous nucleic acid is introduced using at least one additional nucleic acid vector or mRNA molecule.

24. The method of claim 23, wherein the additional vector is selected from the group consisting of: adenovirus, lentivirus, retrovirus, adeno-associated virus, integrative phage vector, nonviral vector, transposon and / or transposase, integrase substrate and plasmid, and / or the additional vector comprises nucleic acid encoding a site-specific nuclease, wherein the site-specific nuclease is optionally selected from the group consisting of: zinc finger nuclease (ZFN), ZFN dimer, transcription activator-like effector nuclease (TALEN), TAL effector domain fusion protein or RNA-directed DNA endonuclease.

25. The method of claim 15, wherein the CHO cell comprises one or more recombination recognition sequences integrated into the locus, and wherein the exogenous nucleic acid is integrated into the one or more recombination recognition sequences in the locus.

26. The method of claim 25, wherein prior to the introduction of the exogenous nucleic acid, the CHO cell comprises at least two recombination recognition sequences integrated into the locus, and an optional marker placed between the two recombination recognition sequences.

27. The method of claim 26, wherein the at least two recombination recognition sequences are distinct and selected from the group consisting of: LoxP site, Lox511 site, Lox2272 site, Lox2372 site, Lox5171 site, Loxm2 site, Lox71 site, Lox66 site, LoxFas site, and frt site.

28. The method of claim 26, wherein the exogenous nucleic acid comprises a first exogenous gene of interest (GOI) operatively linked to a first exogenous promoter, a first recombination recognition sequence located at the 5' of the first exogenous promoter, and a second recombination recognition sequence located at the 3' of the first exogenous GOI.

29. The method of claim 28, wherein the first and second recombination recognition sequences are distinct and selected from the group consisting of: LoxP site, Lox511 site, Lox2272 site, Lox2372 site, Lox5171 site, Loxm2 site, Lox71 site, Lox66 site, LoxFas site, and frt site.

30. The method of claim 28, wherein the exogenous nucleic acid further comprises a second exogenous GOI operatively linked to a second exogenous promoter, wherein the second exogenous GOI is positioned at the 3' of the first exogenous GOI and the 5' of the second recombinant recognition sequence.

31. The method of claim 30, wherein the first exogenous GOI encodes an antibody light chain or an antigen-binding fragment thereof, and the second exogenous GOI encodes an antibody heavy chain or an antigen-binding fragment thereof.

32. A method comprising: a. A cell providing a foreign nucleic acid integrated within a locus in the cell's genome, wherein the locus comprises the nucleotide sequence of SEQ ID NO: 1, and the locus is capable of enhancing the expression of a protein encoded by the foreign nucleotide sequence by at least 1.5 to at least 3 times compared to expression observed in an expression of a foreign nucleotide sequence integrated within another locus in the cell's genome, wherein the foreign nucleic acid comprises a first foreign GOI operatively linked to a first foreign promoter, and b. Culture (a) cells under conditions that allow for the expression of the first exogenous GOI.

33. The method of claim 32, wherein the cell is a CHO cell.

34. The method of claim 32, wherein the first exogenous GOI encodes the first protein of interest (POI), and wherein the method further comprises: c. Recover the first POI.

35. The method of claim 32, wherein the exogenous nucleic acid further comprises a second exogenous GOI operably linked to a second exogenous promoter, and wherein the conditions allow expression of the first exogenous GOI and the second exogenous GOI.

36. The method of claim 35, wherein the first exogenous GOI encodes an antibody light chain and the second exogenous GOI encodes an antibody heavy chain.

37. A vehicle for modifying the genome of CHO cells, comprising a vector, said vector comprising: a. The 5' homologous arm that is identical to the sequence present in the nucleotide sequence of SEQ ID NO:

1. b. Exogenous nucleic acids, and c. The 3' homologous arm that is identical to the sequence present in the nucleotide sequence of SEQ ID NO:

1.

38. The carrier of claim 37, wherein the exogenous nucleic acid comprises one or more recombinant recognition sequences, and / or comprises a first exogenous gene of interest (GOI).

39. The vehicle of claim 37 or 38, wherein the vehicle comprises at least one additional vector or mRNA molecule.

40. The carrier of claim 39, wherein the additional carrier comprises a nucleic acid encoding a site-specific nuclease for integrating a recognition sequence.

41. The carrier of claim 40, wherein the site-specific nuclease is selected from the group consisting of: zinc finger nucleases ZFN, ZFN dimers, transcription activator-like effector nucleases TALEN, TAL effector domain fusion proteins, or RNA-directed DNA endonucleases.