Methods for improved gene insertion and expression in cell lines
By introducing expression constructs with defined ratios into host cell lines with dock sites, the production of therapeutic proteins is optimized, improving yield and purity while ensuring safety and efficacy, addressing the challenges of current manufacturing methods.
Patent Information
- Application Number
- PCT/US2025/040007
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-01
- Filing Date
- 2025-07-31
- Publication Date
- 2026-02-05
AI Technical Summary
Current methods for producing therapeutic proteins are complex, time-consuming, and expensive, with challenges in maintaining product safety and efficacy due to the complexity of protein therapeutics, including large molecular size, post-translational modifications, and the need for biologically active proteins to be manufactured in living cells, which require efficient tools and processes for enhancing functional attributes.
The introduction of expression constructs encoding different gene products at defined ratios into host cell lines with multiple dock sites, using nucleic acid constructs and enzymes to facilitate insertion, allowing for the production of proteins that require expression of at least two gene products, such as an enzyme and its substrate, with specific ratios and compatible insertion elements.
This approach enhances the production of therapeutic proteins by improving yield, purity, and maintaining product safety and efficacy, addressing the complexity and cost issues of current manufacturing processes.
Smart Images

Figure US2025040007_05022026_PF_FP_ABST
Abstract
Description
Attorney Docket No. CATA-40729.601 METHODS FOR IMPROVED GENE INSERTION AND EXPRESSION IN CELL LINES CROSS-REFERENCE TO RELATED APPLICATIONS The present application claims priority to U.S. Provisional Application No. 63 / 678,309, filed August 1, 2024, which is incorporated herein by reference in its entirety. SEQUENCE LISTING The text of the computer readable sequence listing filed herewith, titled “CATA_40729_601_SequenceListing.xml”, created July 25, 2025, having a file size of 67,632 bytes, is hereby incorporated by reference in its entirety. FIELD OF THE INVENTION The present invention relates to the introduction of expression constructs encoding different gene products such as proteins at defined ratios into host cell lines containing multiple dock sites for insertion of the nucleic acid construct, and to the production of proteins that require expression of at least two gene products such as an enzyme and a substrate for the enzyme. BACKGROUND OF THE INVENTION Therapeutic protein drugs are an important class of medicines serving patients most in need of novel therapies. Recombinant protein therapeutics have been developed to treat a wide variety of clinical indications, including cancers, autoimmunity / inflammation, exposure to infectious agents, and genetic disorders. The latest advances in protein-engineering technologies have allowed drug developers and manufacturers to fine-tune and exploit desirable functional characteristics of proteins of interest while maintaining (and in some cases enhancing) product safety or efficacy or both. The manufacturing and production of therapeutic proteins are highly complex processes. For example, a typical protein drug may include in excess of 5,000 critical process steps, many times greater than the number required for manufacturing a small-molecule drug. Similarly, protein therapeutics, which include monoclonal antibodies as well as large or fusion proteins, can be orders-of-magnitude larger in size than small-molecule drugs, having molecular weights exceeding 100 kDa. In addition, protein therapeutics exhibit complex secondary and tertiary structures that must be maintained. Protein therapeutics cannot be completely synthesized by chemical processes and have to be manufactured inAttorney Docket No. CATA-40729.601 living cells or organisms; consequently, the choices of the cell line, species origin, and culture conditions all affect the final product characteristics. Moreover, most biologically active proteins require post-translational modifications that can be compromised when heterologous expression systems are used. Additionally, as the products are synthesized by cells or organisms, complex purification processes are involved. Furthermore, viral clearance processes such as removal of virus particles by using filters or resins, as well as inactivation steps by using low pH or detergents, are implemented to prevent the serious safety issue of viral contamination of protein drug substances. Given the complexity of therapeutic proteins with respect to their large molecular size, post-translational modifications, and the variety of biological materials involved in their manufacturing process, the ability to enhance particular functional attributes while maintaining product safety and efficacy achieved through protein- engineering strategies is highly desirable. While the integration of novel strategies and approaches to modify protein drug products is not a trivial matter, the potential therapeutic advantages have driven the increased use of such strategies during drug development. A number of protein-engineering platform technologies are currently in use to increase the circulating half-life, targeting, and functionality of novel therapeutic protein drugs as well as to increase production yield and product purity. For example, protein conjugation and derivatization approaches, including Fc- fusion, albumin-fusion, and PEGylation, are currently being used to extend a drug’s circulating half-life. The production of protein pharmaceutical (biologics) is expensive and time consuming. What is needed in the art are more efficient tools and processes for producing this important class of drugs. SUMMARY OF THE INVENTION The present invention relates to the introduction of expression constructs encoding different gene products such as proteins at defined ratios into host cell lines containing multiple dock sites for insertion of the nucleic acid construct, and to the production of proteins that require expression of at least two gene products such as an enzyme and a substrate for the enzyme. In some embodiments, the present invention provides methods comprising: introducing at least first nucleic acid constructs encoding a substrate for an enzyme and a second nucleic acid construct encoding the enzyme at a ratio of first nucleic acid constructs to second nucleic acid constructs of from 1:1 to 1000:1 into a host cell having genomeAttorney Docket No. CATA-40729.601 comprising from 1 to 500 integrated docking sites, each docking site comprising at least one dock site insertion element and the nucleic acid constructs each comprising at least one insertion element compatible with the at least one dock site insertion element in the integrated docking sites, under conditions such that the nucleic acid expression constructs are inserted at the dock sites at a ratio of first nucleic acid constructs to second nucleic acid constructs of at least 2:1. In some embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs is from 2:1 to 1000:1. In some embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs is from 5:1 to 500:1. In some embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs is from 10:1 to 200:1. In some embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs is from 10:1 to 100:1. In some embodiments, only first nucleic acid constructs and second nucleic acid constructs are introduced into the host cell. In some embodiments, the methods further comprise introducing a third nucleic acid construct encoding a third protein of interest at a ratio of first nucleic acid construct or second nucleic acid construct to third nucleic acid construct selected from the group consisting of at least 2:1, from 2:1 to 1000:1, from 5:1 to 500:1, from 10:1 to 200:1, and from 10:1 to 100:1. In some embodiments, the methods further comprise introducing a fourth nucleic acid construct encoding a fourth protein of interest at a ratio of first nucleic acid construct, second nucleic acid construct, or third nucleic acid construct to the fourth nucleic construct selected from the group consisting of at least 2:1, from 2:1 to 1000:1, from 5:1 to 500:1, from 10:1 to 200:1, and from 10:1 to 100:1. In some embodiments, the methods further comprise introducing a fifth nucleic acid construct encoding a fifth protein of interest at a ratio of first nucleic acid construct, second nucleic acid construct, third nucleic acid construct or fourth nucleic construct to the fifth nucleic acid construct selected from the group consisting of at least 2:1, from 2:1 to 1000:1, from 5:1 to 500:1, from 10:1 to 200:1, and from 10:1 to 100:1. In some embodiments, the at least first and second nucleic acid constructs further comprise at least the following elements in operable association in 5’ to 3’ order: an internal promoter sequence; a nucleic acid sequence encoding the first protein of interest or second protein that is operably linked to the internal promoter; and a poly A signal sequence. In some embodiments, the at least first and second nucleic acid constructs comprise a selectable marker sequence. In some embodiments, the at least first and second nucleic acidAttorney Docket No. CATA-40729.601 constructs comprise different selectable marker sequences. In some embodiments, one of the first and second nucleic acid constructs comprises a selectable marker sequence and the other of the first and second nucleic acid constructs does not comprise a selectable marker sequence. In some embodiments, the selectable marker sequences are 5’ to the internal promoter sequence and are operably linked to a 5’ promoter sequence. In some embodiments, the nucleic acid construct comprises an extending packaging region (EPR) between the 5’ promoter and the selectable marker. In some embodiments, the EPR comprises multiple potential Kozak sequences and / or ATG translation start sites. In some embodiments, the promoter sequence is selected from the group consisting of SIN- LTR, SV40, EF1α, E. coli lac, E. coli trp, phage lambda PL, phage lambda PR, T3, T7, cytomegalovirus (CMV) immediate early, herpes simplex virus (HSV) thymidine kinase, alpha-lactalbumin, and mouse metallothionein-I promoter sequences. In some embodiments, the first promoter sequence is a weak promoter sequence. In some embodiments, the first promoter sequence is not a retroviral LTR promoter. In some embodiments, the integrated docking sites further comprise an exogenous promoter. In some embodiments, the exogenous promoter is selected from the group consisting of SIN-LTR, SV40, EF1α, E. coli lac, E. coli trp, phage lambda PL, phage lambda PR, T3, T7, cytomegalovirus (CMV) immediate early, herpes simplex virus (HSV) thymidine kinase, alpha-lactalbumin, and mouse metallothionein-I promoter sequences. In some embodiments, the promoter is a retroviral LTR. In some embodiments, the retroviral LTR is a SIN LTR. In some embodiments, the nucleic acid expression constructs are provided in a vector. In some embodiments, the vector is a plasmid vector. In some embodiments, the vector is transiently introduced into the host cell. In some embodiments, the host cell line comprises a nucleic acid construct encoding an enzyme that facilitates insertion of the nucleic acid expression construct at the dock site. In some embodiments, the nucleic acid construct encoding an enzyme that facilitates insertion of the nucleic acid expression construct at the dock site is transiently introduced into the host cell. In some embodiments, the nucleic acid construct encoding an enzyme that facilitates insertion of the nucleic acid expression construct at the dock site is provided in a vector. In some embodiments, the vector is a plasmid vector. In some embodiments, the ratio of the ratio of the nucleic acid constructs encoding an enzyme that facilitates insertion of the nucleic acid expression construct at the dock site to the nucleic acid expression constructs encoding a first protein of interest that are transiently introduced into the host cell is from 1:1000 to 1:10. In some embodiments, theAttorney Docket No. CATA-40729.601 enzyme is selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase. In some embodiments, the nucleic acid construct encoding an enzyme that facilitates insertion of the nucleic acid expression construct at the dock site is provided in a vector. In some embodiments, host cell genome comprises from 5 to 500 integrated docking sites, each docking site comprising at least one dock site insertion element. In some embodiments, the host cell genome comprises from 5 to 250 integrated docking sites, each docking site comprising at least one dock site insertion element. In some embodiments, the host cell genome comprises from 5 to 100 integrated docking sites, each docking site comprising at least one dock site insertion element. In some embodiments, the integrated docking sites are independently positioned throughout the host cell genome. In some embodiments, the dock site insertion element is targeted by enzyme selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase. In some embodiments, the dock site insertion element is selected from the group consisting of a recombinase dock site insertion element and a HDR dock site insertion element. In some embodiments, the dock site insertion element is a recombinase dock site insertion element. In some embodiments, the recombinase dock site insertion element comprises an attachment site (att). In some embodiments, the attachment site (att) is selected from the group consisting of attB and attP and attR and attL. In some embodiments, the recombinase dock site insertion element comprises a LoxP sequence. In some embodiments, the recombinase dock site insertion element is a Flp Recombination Target (FRT) site. In some embodiments, the dock site insertion element is a HDR dock site insertion element. In some embodiments, the HDR dock site insertion element comprises one or two dock site homology arms. In some embodiments, the HDR dock site insertion element further comprises one or more sequences homologous to a guide RNA sequence. In some embodiments, the dock site homology arms are from about 30 to 1000 bases in length. In some embodiments, the integrase dock site insertion element comprises an AAVS1 safe harbor locus sequence. In some embodiments, each docking site is flanked by exogenous integrating vector sequences. In some embodiments, the exogenous integrating vector sequences are selected from the group consisting of viral vector sequences and transposon vector sequences. In some embodiments, the docking sites each further comprise a sequence encoding a selectable maker operably linked to a promoter. In some embodiments, the host cell further comprises an expression construct encoding an exogenous enzyme selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase. In some embodiments, the expression constructAttorney Docket No. CATA-40729.601 encoding an exogenous enzyme selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase is provided in an episomal expression vector. In some embodiments, the expression construct encoding an exogenous enzyme selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase is integrated into the host cell genome. In some embodiments, the dock site insertion element is positioned to facilitate cassette exchange. In some embodiments, each docking site comprises two dock site insertion elements. In some embodiments, the two dock site insertion elements are positioned to facilitate cassette exchange. In some embodiments, the two dock site insertion elements flank sequences encoding a selectable marker, an enzyme, or a combination thereof. In some embodiments, the nucleic acid expression constructs further comprise a signal peptide sequence operably linked to the first protein of interest. In some embodiments, the signal peptide sequence is selected from the group consisting of tissue plasminogen activator, human growth hormone, lactoferrin, alpha-casein and alpha-lactalbumin signal peptide sequences. In some embodiments, the nucleic acid expression constructs further comprise a protein purification marker sequence. In some embodiments, the protein purification marker sequence is a hexahistidine tag or a hemagglutinin (HA) tag. In some embodiments, the host cell is selected from the group consisting of Chinese Hamster Ovary (CHO) cells, HEK 293 cells, CAP cells, bovine mammary epithelial cells, monkey kidney CV1 line transformed by SV40, baby hamster kidney cells, mouse sertoli cells, monkey kidney cells, African green monkey kidney cells, human cervical carcinoma cells, canine kidney cells, buffalo rat liver cells, human lung cells, human liver cells, mouse mammary tumor, TRI cells, MRC 5 cells, FS4 cells, rat fibroblasts, MDBK cells and human hepatoma line cells. In some embodiments, the host cell is selected from the group consisting of a Chinese Hamster Ovary (CHO) cells, a HEK 293 cells and a CAP cells. In some embodiments, the host cell is a GS knockout cell line. In some embodiments, the host cell is a DHFR knockout cell line. In some embodiments, the present invention provides a cell culture comprising host cells made by any of the foregoing methods. In some embodiments, the present invention provides a process for producing a protein of interest comprising culturing the host cells under conditions that the protein of interest is expressed and purifying the protein of interest from the host cell culture. In some embodiments, the host cells are grown in a medium comprising an inhibitor of the selectable marker. In some embodiments, the selectable markerAttorney Docket No. CATA-40729.601 is GS and the inhibitor is phosphinothricin or methionine sulphoximine (Msx). In some embodiments, the selectable marker is DHFR and the inhibitor is methotrexate. In some embodiments, the present invention provides a host cell or host cells comprising: a plurality of docking sites integrated into the genome of the host cell, each docking site comprising at least one dock site insertion element; at least integrated first nucleic acid constructs comprising at least one insertion element compatible with the dock site insertion element and encoding a substrate of an enzyme; and at least integrated second nucleic acid constructs comprising at least one insertion element compatible with the dock site insertion and encoding the enzyme that acts on the substrate, wherein the at least integrated first nucleic acid constructs and the at least second integrated nucleic acid constructs are integrated at the plurality of docking sites at a ratio of first nucleic acid constructs to second nucleic acid constructs of at least 1:1. In some embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs is from 2:1 to 1000:1. In some embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs is from 5:1 to 500:1. In some embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs is from 10:1 to 200:1. In some embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs is from 10:1 to 100:1. In some embodiments, the host cell(s) further comprises a third nucleic acid construct encoding a third protein of interest at a ratio of first nucleic acid construct or second nucleic acid construct to third nucleic acid construct selected from the group consisting of at least 2:1, from 2:1 to 1000:1, from 5:1 to 500:1, from 10:1 to 200:1, and from 10:1 to 100:1. In some embodiments, the host cell(s) further comprises a fourth nucleic acid construct encoding a fourth protein of interest at a ratio of first nucleic acid construct, second nucleic acid construct, or third nucleic acid construct to the fourth nucleic construct selected from the group consisting of at least 2:1, from 2:1 to 1000:1, from 5:1 to 500:1, from 10:1 to 200:1, and from 10:1 to 100:1. In some embodiments, the host cell(s) further comprises a fifth nucleic acid construct encoding a fifth protein of interest at a ratio of first nucleic acid construct, second nucleic acid construct, third nucleic acid construct or fourth nucleic construct to the fifth nucleic acid construct selected from the group consisting of at least 2:1, from 2:1 to 1000:1, from 5:1 to 500:1, from 10:1 to 200:1, and from 10:1 to 100:1. In some embodiments, the at least first and second nucleic acid constructs further comprise at least the following elements in operable association in 5’ to 3’ order: an internalAttorney Docket No. CATA-40729.601 promoter sequence; a nucleic acid sequence encoding the first protein of interest or second protein that is operably linked to the internal promoter; and a poly A signal sequence. In some embodiments, the at least first and second nucleic acid constructs comprise a selectable marker sequence. In some embodiments, the at least first and second nucleic acid constructs comprise different selectable marker sequences. In some embodiments, one of the first and second nucleic acid constructs comprises a selectable marker sequence and the other of the first and second nucleic acid constructs does not comprise a selectable marker sequence. In some embodiments, the selectable marker sequences are 5’ to the internal promoter sequence and are operably linked to a 5’ promoter sequence. In some embodiments, the nucleic acid construct comprises an extending packaging region (EPR) between the 5’ promoter and the selectable marker. In some embodiments, the EPR comprises multiple potential Kozak sequences and / or ATG translation start sites. In some embodiments, the promoter sequence is selected from the group consisting of SIN-LTR, SV40, EF1α, E. coli lac, E. coli trp, phage lambda PL, phage lambda PR, T3, T7, cytomegalovirus (CMV) immediate early, herpes simplex virus (HSV) thymidine kinase, alpha-lactalbumin, and mouse metallothionein-I promoter sequences. In some embodiments, the first promoter sequence is a weak promoter sequence. In some embodiments, the first promoter sequence is not a retroviral LTR promoter. In some embodiments, the integrated docking sites further comprise an exogenous promoter. In some embodiments, the exogenous promoter is selected from the group consisting of SIN-LTR, SV40, EF1α, E. coli lac, E. coli trp, phage lambda PL, phage lambda PR, T3, T7, cytomegalovirus (CMV) immediate early, herpes simplex virus (HSV) thymidine kinase, alpha-lactalbumin, and mouse metallothionein-I promoter sequences. In some embodiments, the promoter is a retroviral LTR. In some embodiments, the retroviral LTR is a SIN LTR. In some embodiments, the nucleic acid expression constructs are provided in a vector. In some embodiments, the vector is a plasmid vector. In some embodiments, the host cell comprises a nucleic acid construct encoding an enzyme that facilitates insertion of the nucleic acid expression construct at the dock site. In some embodiments, the nucleic acid construct encoding an enzyme that facilitates insertion of the nucleic acid expression construct at the dock site is provided in a vector. In some embodiments, the vector is a plasmid vector. In some embodiments, the ratio of the ratio of the nucleic acid constructs encoding an enzyme that facilitates insertion of the nucleic acid expression construct at the dock site to the nucleic acid expression constructs encoding a first protein of interest that are transiently introduced into the host cell is from 1:1000 to 1:10. InAttorney Docket No. CATA-40729.601 some embodiments, the enzyme is selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase. In some embodiments, the nucleic acid construct encoding an enzyme that facilitates insertion of the nucleic acid expression construct at the dock site is provided in a vector. In some embodiments, the host cell genome comprises from 5 to 500 integrated docking sites, each docking site comprising at least one dock site insertion element. In some embodiments, the host cell genome comprises from 5 to 250 integrated docking sites, each docking site comprising at least one dock site insertion element. In some embodiments, the host cell genome comprises from 5 to 100 integrated docking sites, each docking site comprising at least one dock site insertion element. In some embodiments, the integrated docking sites are independently positioned throughout the host cell genome. In some embodiments, the dock site insertion element is targeted by enzyme selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase. In some embodiments, the dock site insertion element is selected from the group consisting of a recombinase dock site insertion element and a HDR dock site insertion element. In some embodiments, the dock site insertion element is a recombinase dock site insertion element. In some embodiments, the recombinase dock site insertion element comprises an attachment site (att). In some embodiments, the attachment site (att) is selected from the group consisting of attB and attP and attR and attL. In some embodiments, the recombinase dock site insertion element comprises a LoxP sequence. In some embodiments, the recombinase dock site insertion element is a Flp Recombination Target (FRT) site. In some embodiments, the dock site insertion element is a HDR dock site insertion element. In some embodiments, the HDR dock site insertion element comprises one or two dock site homology arms. In some embodiments, the HDR dock site insertion element further comprises one or more sequences homologous to a guide RNA sequence. In some embodiments, the dock site homology arms are from about 30 to 1000 bases in length. In some embodiments, the integrase dock site insertion element comprises an AAVS1 safe harbor locus sequence. In some embodiments, each docking site is flanked by exogenous integrating vector sequences. In some embodiments, the exogenous integrating vector sequences are selected from the group consisting of viral vector sequences and transposon vector sequences. In some embodiments, the docking sites each further comprise a sequence encoding a selectable maker operably linked to a promoter. In some embodiments, the host cell further comprises an expression construct encoding an exogenous enzyme selected from the group consisting of an integrase, aAttorney Docket No. CATA-40729.601 recombinase, a nuclease and a nickase. In some embodiments, the expression construct encoding an exogenous enzyme selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase is provided in an episomal expression vector. In some embodiments,the expression construct encoding an exogenous enzyme selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase is integrated into the host cell genome. In some embodiments, the dock site insertion element is positioned to facilitate cassette exchange. In some embodiments, each docking site comprises two dock site insertion elements. In some embodiments, the two dock site insertion elements are positioned to facilitate cassette exchange. In some embodiments, the two dock site insertion elements flank sequences encoding a selectable marker, an enzyme, or a combination thereof. In some embodiments, the nucleic acid expression constructs further comprise a signal peptide sequence operably linked to the first protein of interest. In some embodiments, the signal peptide sequence is selected from the group consisting of tissue plasminogen activator, human growth hormone, lactoferrin, alpha-casein and alpha-lactalbumin signal peptide sequences. In some embodiments, the nucleic acid expression constructs further comprise a protein purification marker sequence. In some embodiments, the protein purification marker sequence is a hexahistidine tag or a hemagglutinin (HA) tag. In some embodiments, the host cell is selected from the group consisting of Chinese Hamster Ovary (CHO) cells, HEK 293 cells, CAP cells, bovine mammary epithelial cells, monkey kidney CV1 line transformed by SV40, baby hamster kidney cells, mouse sertoli cells, monkey kidney cells, African green monkey kidney cells, human cervical carcinoma cells, canine kidney cells, buffalo rat liver cells, human lung cells, human liver cells, mouse mammary tumor, TRI cells, MRC 5 cells, FS4 cells, rat fibroblasts, MDBK cells and human hepatoma line cells. In some embodiments, the host cell is selected from the group consisting of a Chinese Hamster Ovary (CHO) cells, a HEK 293 cells and a CAP cells. In some embodiments, the host cell is a GS knockout cell line. In some embodiments, the host cell is a DHFR knockout cell line. In some embodiments, the present invention provides a cell culture comprising host cells as described above. In some embodiments, the present invention provides a process for producing a protein of interest comprising culturing host cells as described above under conditions that the protein(s) of interest are expressed and purifying the protein(s) of interest from the host cell culture. In some embodiments, the host cells are grown in a medium comprising an inhibitor of the selectable marker. In some embodiments, the selectable markerAttorney Docket No. CATA-40729.601 is GS and the inhibitor is phosphinothricin or methionine sulphoximine (Msx). In some embodiments, the selectable marker is DHFR and the inhibitor is methotrexate. DESCRIPTION OF THE FIGURES FIG.1 is a map of plasmid pFCS-hHGF-WPRE-SIN. FIG.2 shows sequences of the transgene insert sequence for plasmid pFCS-hHGF- WPRE-SIN (nucleotide: SEQ ID NO.35; amino acid: SEQ ID NO: 1). FIG.3 is a map of plasmid pFCS-newhHepsin-WPRE_SIN. FIG.4 shows sequences of the transgene insert sequence for plasmid pFCS- newhHepsin-WPRE_SIN (nucleotide: SEQ ID NO: 36; amino acid: SEQ ID NO: 37). FIG.5 is a map of plasmid pFCS-SPhHGF-WPRE-SIN. FIG.6 shows sequences of the transgene insert sequence for plasmid pFCS-SPhHGF- WPRE-SIN (nucleotide: SEQ ID NO: 38; amino acid: SEQ ID NO: 39). FIG.7 is a map of plasmid 207attB-GS-Hepsin-WPRE. FIG.8 shows sequences of the transgene insert sequence for plasmid 207attB-GS- Hepsin-WPRE (nucleotide: SEQ ID NO: 36; amino acid: SEQ ID NO: 37). FIG.9 is a map of plasmid 207attB-GS-hHGF-WPRE. FIG.10 shows sequences of the transgene insert sequence for plasmid 207attB-GS- hHGF-WPRE (nucleotide: SEQ ID NO: 35; amino acid: SEQ ID NO: 1). FIGS.11A–11D show cell growth (FIG.11A and FIG.11C) and viability (FIG.11B and FIG.11D) curves for pooled cell cultures transduced with pFCS-hHGF-WPRE-SIN (FIG.11A and FIG.11B) and pFCS-SPhHGF-WPRE-SIN (FIG.11C and FIG.11D). FIG.12 shows a non-reduced SDS-PAGE of HGF produced by pooled cell cultures transduced with pFCS-hHGF-WPRE-SIN or pFCS-SPhHGF-WPRE-SIN. Lane 1: Molecular weight standard; Lane 2: blank; Lane 3: pFCS-hHGF-WPRE-SIN 4d; Lane 4: pFCS-hHGF- WPRE-SIN 8d; Lane 5: pFCS-hHGF-WPRE-SIN 12d; Lane 6: pFCS-SPhHGF-WPRE-SIN 4d; Lane 7: pFCS-SPhHGF-WPRE-SIN 8d; Lane 8: pFCS-SPhHGF-WPRE-SIN 10d; Lane 9: HGF standard (0.2 μg load); Lane 10: HGF standard (2 μg load); Lane 11: HGF standard (5 μg load). FIG.13 shows a reduced SDS-PAGE of HGF produced by pooled cell cultures transduced with pFCS-hHGF-WPRE-SIN or pFCS-SPhHGF-WPRE-SIN. Lane 1: Molecular weight standard; Lane 2: blank; Lane 3: pFCS-hHGF-WPRE-SIN 4d; Lane 4: pFCS-hHGF- WPRE-SIN 8d; Lane 5: pFCS-hHGF-WPRE-SIN 12d; Lane 6: pFCS-SPhHGF-WPRE-SIN 4d; Lane 7: pFCS-SPhHGF-WPRE-SIN 8d; Lane 8: pFCS-SPhHGF-WPRE-SIN 10d; LaneAttorney Docket No. CATA-40729.601 9: HGF standard (0.2 μg load); Lane 10: HGF standard (2 μg load); Lane 11: HGF standard (5 μg load). FIGS.14A–14B show non-reduced (FIG.14A) and reduced (FIG.14B) SDS-PAGE gels of cell pool samples from pools containing and not containing hepsin, as well as a control HGF standard. Lane 1: Molecular weight standard; Lane 2: condition media: Lane 3 CI01 std (2 μg load); Lane 4: CI013x / 2x IP; Lane 5: CI011CI013x / 2x IP; Lane 6: CI013x 9d; Lane 8: CI013x / 1x 9d; Lane 9: Ladder (0.2 μg load). FIG.15 shows cell viability curves after glutamine selection for cells transduced with HGF and hepsin at HGF:hepsin ratios (from left to right), 100:1, 10:1, and 1:1. FIGS.16A–16B show non-reduced (FIG.16A) and reduced (FIG.16B) SDS-PAGE gels of samples GPExTMLightning cell pool samples from the pool lacking hepsin and the pool produced using the 100:1 ratio of HGF:Hepsin, as well as a control HGF standard. Lane 1: ladder; Lane 2: HGF Standard (2 μg load); Lane 3: Conditioned Media (Fed Batch +Q); Lane 4: Line 1-1; Lane 5: Line 1-2; Lane 6: Line 1-3; Lane 7: Line 4-1 (100:1); Lane 8: Line 4-2 (100:1); Lane 9: Line 4-3 (100:1). FIGS.17A–17B show viable cell density (FIG.17A) and cell viability (FIG.17B) through the culture for a fed-batch productivity study of Line 1 (squares) and Line 4 (diamonds). FIGS.18A–18D show non-reduced (FIG.18A and FIG.18B) and reduced (FIG.18C and FIG.18D) SDS-PAGE gels of samples from Line 1 and Line 4 cell pools harvested on various days throughout a production run. FIG.18A and FIG.18C: Lane 1: ladder; Lane 2: HGF standard (2 μg load); Lane 3: Conditioned Fed Batch Media (+Q); Lane 4: Fresh Fed Batch Media (-Q); Lane 5: Line 1 d8 (7.5 μl load); Lane 6: Line 2 d8 (7.5 μl load); Lane 7: Line 1 d10 (7.5 μl load); Lane 8: Line 2 d10 (7.5 μl load); Lane 9: Line 1 d8 (1.5 μl load); Lane 10: Line 2 d8 (1.5 μl load); Lane 11: Line 1 d10 (1.5 μ load); Lane 12: Line 2 d10 (1.5 μ load). FIG.18B and FIG.18D: Lane 1: ladder; Lane 2: HGF standard (2 μg load); Lane 3: Conditioned Fed Batch Media (+Q); Lane 4: Fresh Fed Batch Media (-Q); Lane 5: Line 1 d12 (1.3 μl load); Lane 6: Line 2 d12 (1.3 μl load); Lane 7: Line 1 d16 (1.3 μl load); Lane 8: Line 2 d16 (1.3 μl load); Lane 9: Line 1 d17 harvest (1.3 μl load); Lane 10: Line 2 d17 harvest (1.3 μl load); Lane 11: HGF standard (1 μg load). FIG.19 shows non-reduced (left) and reduced (right) SDS-PAGE gels of material purified from cell pools for Lines 1 (middle lane in each) and 4 (right lane in each). Left lane: molecular weight standard.Attorney Docket No. CATA-40729.601 FIG.20 provides a sequence summary for Yourway Light Chain (LC) CDS in vectors GDD1008.0211 and GDD1008.0223 (nucleotide: SEQ ID NO: 40; amino acid: SEQ ID NO: 41). FIG.21 provides a map of plasmid 223attB-GS-Yourway LC-WPRE, GDD1008.0223. FIG.22 provides a sequence summary for Yourway Heavy Chain (HC) CDS in vectors GDD1008.0211 and GDD1008.0224 (nucleotide: SEQ ID NO: 42; amino acid: SEQ ID NO: 43). FIG.23 provides a map of plasmid 224attB-GS-Yourway HC-WPRE, GDD1008.0224. FIG.24 provides a map of plasmid 211attB287-GS-sC-YourwayHC-WPRE-hC-Int- YourwayLC, GDD1008.0211. FIG.25 provides a graph showing cell viability curves after glutamine selection. FIG.26 provides a graph showing pooled cell productivity of the different heavy and light chain ratios as well as the double chain construct (CT36-2). FIG.27 provides a non-reduced and reduced SDS-PAGE of samples all of the heavy chain / light chain ratio cell pools as well as the single gene construct (Historic Yourway HWIL). DEFINITIONS To facilitate understanding of the invention, a number of terms are defined below. As used herein, the term "host cell" refers to any eukaryotic cell (e.g., mammalian cells, avian cells, amphibian cells, plant cells, fish cells, and insect cells), whether located in vitro or in vivo. As used herein, the term "cell culture" refers to any in vitro culture of cells. Included within this term are continuous cell lines (e.g., with an immortal phenotype), primary cell cultures, finite cell lines (e.g., non-transformed cells), and any other cell population maintained in vitro, including oocytes and embryos. As used herein, the term "vector" refers to any genetic element, such as a plasmid, phage, transposon, cosmid, chromosome, virus, virion, etc., which is capable of replication when associated with the proper control elements and which can transfer gene sequences between cells. Thus, the term includes cloning and expression vehicles, as well as viral vectors.Attorney Docket No. CATA-40729.601 As used herein, the term “genome” refers to the genetic material (e.g., chromosomes) of an organism. The term "nucleotide sequence of interest" refers to any nucleotide sequence (e.g., RNA or DNA), the manipulation of which may be deemed desirable for any reason (e.g., treat disease, confer improved qualities, expression of a protein of interest in a host cell, expression of a ribozyme, etc.), by one of ordinary skill in the art. Such nucleotide sequences include, but are not limited to, coding sequences of structural genes (e.g., reporter genes, selection marker genes, oncogenes, drug resistance genes, growth factors, etc.), and non- coding regulatory sequences which do not encode an mRNA or protein product (e.g., promoter sequence, polyadenylation sequence, termination sequence, enhancer sequence, etc.). As used herein, the term “protein of interest” refers to a protein encoded by a nucleic acid of interest. As used herein, the terms "nucleic acid molecule encoding," "DNA sequence encoding," "DNA encoding," "RNA sequence encoding," and "RNA encoding" refer to the order or sequence of deoxyribonucleotides or ribonucleotides along a strand of deoxyribonucleic acid or ribonucleic acid. The order of these deoxyribonucleotides or ribonucleotides determines the order of amino acids along the polypeptide (protein) chain. The DNA or RNA sequence thus codes for the amino acid sequence. The term "promoter," "promoter element," or "promoter sequence" as used herein, refers to a DNA sequence which when ligated to a nucleotide sequence of interest is capable of controlling the transcription of the nucleotide sequence of interest into mRNA. A promoter is typically, though not necessarily, located 5' (i.e., upstream) of a nucleotide sequence of interest whose transcription into mRNA it controls, and provides a site for specific binding by RNA polymerase and other transcription factors for initiation of transcription. Transcriptional control signals in eukaryotes comprise "promoter" and "enhancer" elements. Promoters and enhancers consist of short arrays of DNA sequences that interact specifically with cellular proteins involved in transcription (Maniatis et al., Science 236:1237
[1987] ). Promoter and enhancer elements have been isolated from a variety of eukaryotic sources including genes in yeast, insect and mammalian cells, and viruses (analogous control elements, i.e., promoters, are also found in prokaryotes). The selection of a particular promoter and enhancer depends on what cell type is to be used to express the protein of interest. Some eukaryotic promoters and enhancers have a broad host range while others areAttorney Docket No. CATA-40729.601 functional in a limited subset of cell types (for review see, Voss et al., Trends Biochem. Sci., 11:287
[1986] ; and Maniatis et al., supra). For example, the SV40 early gene enhancer is very active in a wide variety of cell types from many mammalian species and has been widely used for the expression of proteins in mammalian cells (Dijkema et al., EMBO J. 4:761
[1985] ). Two other examples of promoter / enhancer elements active in a broad range of mammalian cell types are those from the human elongation factor 1α gene (Uetsuki et al., J. Biol. Chem., 264:5791
[1989] ; Kim et al., Gene 91:217
[1990] ; and Mizushima and Nagata, Nuc. Acids. Res., 18:5322
[1990] ) and the long terminal repeats of the Rous sarcoma virus (Gorman et al., Proc. Natl. Acad. Sci. USA 79:6777
[1982] ) and the human cytomegalovirus (Boshart et al., Cell 41:521
[1985] ). As used herein, the term "promoter / enhancer" denotes a segment of DNA which contains sequences capable of providing both promoter and enhancer functions (i.e., the functions provided by a promoter element and an enhancer element, see above for a discussion of these functions). For example, the long terminal repeats of retroviruses contain both promoter and enhancer functions. The enhancer / promoter may be "endogenous" or "exogenous" or "heterologous." An "endogenous" enhancer / promoter is one that is naturally linked with a given gene in the genome. An "exogenous" or "heterologous" enhancer / promoter is one that is placed in juxtaposition to a gene by means of genetic manipulation (i.e., molecular biological techniques such as cloning and recombination) such that transcription of that gene is directed by the linked enhancer / promoter. As used herein, the term “long terminal repeat” of "LTR" refers to transcriptional control elements located in or isolated from the U3 region 5' and 3' of a retroviral genome. As is known in the art, long terminal repeats may be used as control elements in retroviral vectors, or isolated from the retroviral genome and used to control expression from other types of vectors. As used herein, the terms "complementary" or "complementarity" are used in reference to polynucleotides (i.e., a sequence of nucleotides) related by the base-pairing rules. For example, the sequence "5'-A-G-T-3'," is complementary to the sequence "3'-T-C-A-5'." Complementarity may be "partial," in which only some of the nucleic acids' bases are matched according to the base pairing rules. Or, there may be "complete" or "total" complementarity between the nucleic acids. The degree of complementarity between nucleic acid strands has significant effects on the efficiency and strength of hybridization between nucleic acid strands. This is of particular importance in amplification reactions, as well as detection methods that depend upon binding between nucleic acids.Attorney Docket No. CATA-40729.601 The terms "homology" and "percent identity" when used in relation to nucleic acids refers to a degree of complementarity. There may be partial homology (i.e., partial identity) or complete homology (i.e., complete identity). A partially complementary sequence is one that at least partially inhibits a completely complementary sequence from hybridizing to a target nucleic acid sequence and is referred to using the functional term "substantially homologous." The inhibition of hybridization of the completely complementary sequence to the target sequence may be examined using a hybridization assay (Southern or Northern blot, solution hybridization and the like) under conditions of low stringency. A substantially homologous sequence or probe (i.e., an oligonucleotide which is capable of hybridizing to another oligonucleotide of interest) will compete for and inhibit the binding (i.e., the hybridization) of a completely homologous sequence to a target sequence under conditions of low stringency. This is not to say that conditions of low stringency are such that non-specific binding is permitted; low stringency conditions require that the binding of two sequences to one another be a specific (i.e., selective) interaction. The absence of non-specific binding may be tested by the use of a second target which lacks even a partial degree of complementarity (e.g., less than about 30% identity); in the absence of non-specific binding the probe will not hybridize to the second non-complementary target. The terms "in operable combination," "in operable order," and "operably linked" as used herein refer to the linkage of nucleic acid sequences in such a manner that a nucleic acid molecule capable of directing the transcription of a given gene and / or the synthesis of a desired protein molecule is produced. The term also refers to the linkage of amino acid sequences in such a manner so that a functional protein is produced. As used herein, the term “selectable marker” refers to a gene that encodes an enzymatic activity or other protein that confers the ability to grow in medium lacking what would otherwise be an essential nutrient; in addition, a selectable marker may confer resistance to an antibiotic or drug upon the cell in which the selectable marker is expressed. As used herein, the term “retrovirus" refers to a retroviral particle which is capable of entering a cell (i.e., the particle contains a membrane-associated protein such as an envelope protein or a viral G glycoprotein which can bind to the host cell surface and facilitate entry of the viral particle into the cytoplasm of the host cell) and integrating the retroviral genome (as a double-stranded provirus) into the genome of the host cell. The term "retrovirus" encompasses Oncovirinae (e.g., Moloney murine leukemia virus (MoMLV), Moloney murine sarcoma virus (MoMSV), and Mouse mammary tumor virus (MMTV), Spumavirinae, amd Lentivirinae (e.g., Human immunodeficiency virus, Simian immunodeficiency virus, EquineAttorney Docket No. CATA-40729.601 infection anemia virus, and Caprine arthritis-encephalitis virus; See, e.g., U.S. Pat. Nos. 5,994,136 and 6,013,516, both of which are incorporated herein by reference). As used herein, the term "retroviral vector" refers to a retrovirus that has been modified to express a gene of interest. Retroviral vectors can be used to transfer genes efficiently into host cells by exploiting the viral infectious process. Foreign or heterologous genes cloned (i.e., inserted using molecular biological techniques) into the retroviral genome can be delivered efficiently to host cells that are susceptible to infection by the retrovirus. Through well-known genetic manipulations, the replicative capacity of the retroviral genome can be destroyed. The resulting replication-defective vectors can be used to introduce new genetic material to a cell but they are unable to replicate. A helper virus or packaging cell line can be used to permit vector particle assembly and egress from the cell. Such retroviral vectors comprise a replication-deficient retroviral genome containing a nucleic acid sequence encoding at least one gene of interest (i.e., a polycistronic nucleic acid sequence can encode more than one gene of interest), a 5' retroviral long terminal repeat (5' LTR); and a 3' retroviral long terminal repeat (3' LTR). As used herein, the term “lentivirus vector” refers to retroviral vectors derived from the Lentiviridae family (e.g., human immunodeficiency virus, simian immunodeficiency virus, equine infectious anemia virus, and caprine arthritis-encephalitis virus) that are capable of integrating into non-dividing cells (See, e.g., U.S. Pat. Nos.5,994,136 and 6,013,516, both of which are incorporated herein by reference). As used herein, the term “transposon” refers to transposable elements (e.g., Tn5, Tn7, and Tn10) that can move or transpose from one position to another in a genome. In general, the transposition is controlled by a transposase. The term "transposon vector," as used herein, refers to a vector encoding a nucleic acid of interest flanked by the terminal ends of transposon. Examples of transposon vectors include, but are not limited to, those described in U.S. Pat. Nos.6,027,722; 5,958,775; 5,968,785; 5,965,443; and 5,719,055, all of which are incorporated herein by reference. As used herein, the term “adeno-associated virus (AAV) vector” refers to a vector derived from an adeno-associated virus serotype, including without limitation, AAV-1, AAV- 2, AAV-3, AAV-4, AAV-5, AAVX7, etc. AAV vectors can have one or more of the AAV wild-type genes deleted in whole or part, preferably the rep and / or cap genes, but retain functional flanking ITR sequences. AAV vectors can be constructed using recombinant techniques that are known in the art to include one or more heterologous nucleotide sequences flanked on both ends (5' and 3')Attorney Docket No. CATA-40729.601 with functional AAV ITRs. In the practice of the invention, an AAV vector can include at least one AAV ITR and a suitable promoter sequence positioned upstream of the heterologous nucleotide sequence and at least one AAV ITR positioned downstream of the heterologous sequence. A "recombinant AAV vector plasmid" refers to one type of recombinant AAV vector wherein the vector comprises a plasmid. As with AAV vectors in general, 5' and 3' ITRs flank the selected heterologous nucleotide sequence. As used herein, the term “adenoviral vector” refers to a non-enveloped double- stranded DNA vector comprising an adenovirus backbone. As used herein, the term "purified" refers to molecules, either nucleic or amino acid sequences, that are removed from their normal environment, isolated or separated. An "isolated nucleic acid sequence" is therefore a purified nucleic acid sequence. "Substantially purified" molecules are at least 60% free, preferably at least 75% free, and more preferably at least 90% free from other components with which they are normally associated. ABBREVIATIONS AmpR= bacterial ampicillin resistance gene attB= Bacterial Attachment Site attP= Phage Attachment Site attR= Recombined Upstream Attachmend Site Backbone=Plasmid Backbone CDS= Coding Sequence EPR=MMLV Extended Packaging Region GCI = Gene Copy Index GS= Glutamine Synthetase H or HC= Heavy Chain hCMV= Human Cytomegalovirus immediate-early Promoter I= intron L or LC= Light Chain MoMuSV 5’LTR= Moloney Murine Sarcoma Virus 5’ Long Terminal Repeat Neo= Neomycin resistance genePA or PolyA= Polyadenylation signal ProV SIN-LTR= Proviral Self- Inactivating Long Terminal Repeat sCMV= Simian Cytomegalovirus immediate-early Promoter SDS-PAGE= Sodium Dodecyl Sulphate- Polyacrylamide Gel Electrophoresis SIN-3’LTR= Self-Inactivation 3’ Long Terminal RepeatAttorney Docket No. CATA-40729.601 SV40= Simian Virus 40 TK= Thymidine Kinase UTR= Untranslated Region W or WPRE= Woodchuck Post-transcriptional Regulatory Element DETAILED DESCRIPTION OF THE INVENTION The present invention relates to the introduction of expression constructs encoding different gene products such as proteins at defined ratios into host cell lines containing multiple dock sites for insertion of the nucleic acid construct, and to the production of proteins that require expression of at least two gene products. Cell lines containing multiple dock sites and expression constructs for use with the cells are described in PCT / US21 / 35403 and PCT / US21 / 35404, both of which are incorporated by reference herein in their entirety. In some preferred embodiments, the present invention provides methods for integrating two or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more) nucleic acid constructs into a host cell line at a desired ratios, and host cells produced by such processes. In some preferred embodiments, the methods comprise introducing at least first nucleic acid constructs encoding a first protein of interest and second nucleic acid constructs encoding a second protein of interest at a ratio of first nucleic acid constructs to second nucleic acid constructs of at least 2:1 into a host cell having genome comprising from 1 to 500 integrated docking sites, each docking site comprising at least one dock site insertion element and the nucleic acid constructs each comprising at least one insertion element compatible with the at least one dock site insertion element in the integrated docking sites, under conditions such that the nucleic acid expression constructs are inserted at the dock sites at a ratio of first nucleic acid constructs to second nucleic acid constructs of from 1:1 to 1000:1. In some preferred embodiments, the nucleic acid expression constructs are inserted at the dock sites at a ratio of first nucleic acid constructs to second nucleic acid constructs of at least 1.1:1. In some preferred embodiments, the nucleic acid expression constructs are inserted at the dock sites at a ratio of first nucleic acid constructs to second nucleic acid constructs of at least 1.2:1. In some preferred embodiments, the nucleic acid expression constructs are inserted at the dock sites at a ratio of first nucleic acid constructs to second nucleic acid constructs of at least 1.3:1. In some preferred embodiments, the nucleic acid expression constructs are inserted at the dock sites at a ratio of first nucleic acid constructs toAttorney Docket No. CATA-40729.601 second nucleic acid constructs of at least 1.4:1. In some preferred embodiments, the nucleic acid expression constructs are inserted at the dock sites at a ratio of first nucleic acid constructs to second nucleic acid constructs of at least 1.5:1. In some further preferred embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs introduced in the host cells is from 1.1:1 to 1000:1. In some further preferred embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs introduced in the host cells is from 1.5:1 to 1000:1. In some further preferred embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs introduced in the host cells is from 2:1 to 1000:1. In some further preferred embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs introduced in the host cells is from 1.1:1 to 100:1. In some further preferred embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs introduced in the host cells is from 1.5:1 to 100:1. In some further preferred embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs introduced in the host cells is from 2:1 to 100:1. In some preferred embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs is from 5:1 to 500:1. In some preferred embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs is from 10:1 to 200:1. In some preferred embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs is from 10:1 to 100:1. In some preferred embodiments, the present invention provides host cells comprising a plurality of docking sites integrated into the genome of the host cell, each docking site comprising at least one dock site insertion element; at least integrated first nucleic acid constructs comprising at least one insertion element compatible with the dock site insertion element and encoding a first protein of interest, and at least integrated second nucleic acid constructs comprising at least one insertion element compatible with the dock site insertion and encoding a second protein of interest, wherein the at least integrated first nucleic acid constructs and the at least second integrated nucleic acid constructs are integrated at the plurality of docking sites at a ratio of first nucleic acid constructs to second nucleic acid constructs of at least 2:1. In some preferred embodiments, the ratio of integrated first nucleic acid constructs to integrated second nucleic acid constructs introduced in the host cells is from 2:1 to 1000:1. In some preferred embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs is from 5:1 to 500:1. In some preferred embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs is from 10:1 to 200:1. In some preferredAttorney Docket No. CATA-40729.601 embodiments, the ratio of first nucleic acid constructs to second nucleic acid constructs is from 10:1 to 100:1. In some further preferred embodiments, first, second and third nucleic acid constructs encoding, respectively, first, second and third proteins of interest are introduced into the host cells at a ratio of from 1:1:1 to 1000:1:1. In some further preferred embodiments, first, second and third nucleic acid constructs encoding, respectively, first, second and third proteins of interest are introduced into the host cells at a ratio of from 1.5:1:1 to 1000:1:1. In some further preferred embodiments, first, second and third nucleic acid constructs encoding, respectively, first, second and third proteins of interest are introduced into the host cells at a ratio of from 2:1:1 to 1000:1:1. In some further preferred embodiments, first, second and third nucleic acid constructs encoding, respectively, first, second and third proteins of interest are introduced into the host cells at a ratio of from 10:1:1 to 1000:1:1. In some further preferred embodiments, first, second and third nucleic acid constructs encoding, respectively, first, second and third proteins of interest are introduced into the host cells at a ratio of from 1:1:1 to 100:1:1. In some further preferred embodiments, first, second and third nucleic acid constructs encoding, respectively, first, second and third proteins of interest are introduced into the host cells at a ratio of from 1.5:1:1 to 100:1:1. In some further preferred embodiments, first, second and third nucleic acid constructs encoding, respectively, first, second and third proteins of interest are introduced into the host cells at a ratio of from 2:1:1 to 100:1:1. In some further preferred embodiments, first, second and third nucleic acid constructs encoding, respectively, first, second and third proteins of interest are introduced into the host cells at a ratio of from 10:1:1 to 100:1:1. In some further preferred embodiments, first, second and third nucleic acid constructs encoding, respectively, first, second and third proteins of interest are introduced into the host cells at a ratio of from 1:1:1 to 1000:100:1. In some further preferred embodiments, first, second and third nucleic acid constructs encoding, respectively, first, second and third proteins of interest are introduced into the host cells at a ratio of from 1.5:1:1 to 1000:100:1. In some further preferred embodiments, first, second and third nucleic acid constructs encoding, respectively, first, second and third proteins of interest are introduced into the host cells at a ratio of from 2:1:1 to 1000:100:1. In some further preferred embodiments, first, second and third nucleic acid constructs encoding, respectively, first, second and third proteins of interest are introduced into the host cells at a ratio of from 10:1:1 to 1000:100:1. In some further preferred embodiments, first, second and third nucleic acid constructs encoding, respectively, first, second and third proteins of interest are introduced into the hostAttorney Docket No. CATA-40729.601 cells at a ratio of from 1:1:1 to 100:10:1. In some further preferred embodiments, first, second and third nucleic acid constructs encoding, respectively, first, second and third proteins of interest are introduced into the host cells at a ratio of from 1.5:1:1 to 100:10:1. In some further preferred embodiments, first, second and third nucleic acid constructs encoding, respectively, first, second and third proteins of interest are introduced into the host cells at a ratio of from 2:1:1 to 100:10:1. In some further preferred embodiments, first, second and third nucleic acid constructs encoding, respectively, first, second and third proteins of interest are introduced into the host cells at a ratio of from 10:1:1 to 100:10:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 1:1:1:1 to 1000:1:1:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 1.5:1:1:1 to 1000:1:1:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 2:1:1:1 to 1000:1:1:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 10:1:1:1 to 1000:1:1:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 1:1:1:1 to 100:1:1:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 1.5:1:1:1 to 100:1:1:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 2:1:1:1 to 100:1:1:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 10:1:1:1 to 100:1:1:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 1:1:1:1 to 1000:100:1:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding,Attorney Docket No. CATA-40729.601 respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 1.5:1:1:1 to 1000:100:1:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 2:1:1:1 to 1000:100:1:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 10:1:1:1 to 1000:100:1:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 1:1:1:1 to 1000:100:100:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 1.5:1:1:1 to 1000:100:100:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 2:1:1:1 to 1000:100:100:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 10:1:1:1 to 1000:100:100:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 1:1:1:1 to 1000:100:100:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 1.5:1:1:1 to 1000:100:100:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 2:1:1:1 to 1000:100:100:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 10:1:1:1 to 1000:100:100:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 1:1:1:1 to 100:100:100:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding,Attorney Docket No. CATA-40729.601 respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 1.5:1:1:1 to 100:100:100:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 2:1:1:1 to 100:100:100:1. In some further preferred embodiments, first, second, third and fourth nucleic acid constructs encoding, respectively, first, second, third and fourth proteins of interest are introduced into the host cells at a ratio of from 10:1:1:1 to 100:100:100:1. In some preferred embodiments, the first protein of interest is a substrate for an enzyme and the second protein of interest is an enzyme that acts on the substrate. Exemplified herein is the co-integration and expression constructs encoding hepsin (enzyme; protease) and hepatocyte growth factor (HGF; substrate for hepsin). Accordingly in some embodiments, the first protein of interest can be a substrate for a protease and the second protein of interest can be a protease. In some embodiments, expression of at least one of the proteins of interest is toxic to the host cell (e.g., expression of a protease). In some embodiments, the present invention allows for co-integration of an expression construct encoding a protein of interest for a protein that is toxic to the host cell with an expression construct for a protein of interest that is not toxic to the host cell where the expression construct encoding the protein of interest that is not toxic to the host cell is provided at an increased ratio as compared to the expression construct for the protein that is toxic to the host cell. Exemplary ratios are provided above. In some preferred embodiments, an expressed exogenous protein is determined to be toxic to a host cell when its expression decreases a) host cell viability and / or b) host cell growth rate. In some preferred embodiments, the first and second proteins of interest are subunits of a multi-subunit protein. In some embodiments, the present invention allows for co- integration of an expression construct encoding a first subunit of a protein of interest with an expression construct encoding a second subunit of a protein of interest where the expression construct encoding one of the subunits is provided at an increased ratio as compared to the expression construct encoding the other subunits(s). Exemplary ratios are provided above. In some preferred embodiments, the first and second proteins of interest are subunits of a viral particle. In some embodiments, the present invention allows for co-integration of an expression construct encoding a first subunit of a viral particle with an expression construct encoding a second subunit of a viral particle where the expression construct encoding one of the subunits is provided at an increased ratio as compared to the expression construct encoding the other subunits(s). Exemplary ratios are provided above.Attorney Docket No. CATA-40729.601 In some preferred the host cells (and cultures of host cells) are engineered to comprise a plurality of integrated docking sites. For example, in some preferred embodiments, the genomes of the host cells of the present invention preferably comprise from 1 to 1000 integrated docking sites, each docking site comprising at least one dock site insertion element. In other preferred embodiments, the genome of the host cells comprises from 1 to 500 integrated docking sites, each docking site comprising at least one dock site insertion element. In other preferred embodiments, the genome of the host cells comprises from 5 to 500 integrated docking sites, each docking site comprising at least one dock site insertion element. In other preferred embodiments, the genome of the host cells comprises from 5 to 250 integrated docking sites, each docking site comprising at least one dock site insertion element. In other preferred embodiments, the genome of the host cell comprises from 5 to 250 integrated docking sites, each docking site comprising at least one dock site insertion element. In other preferred embodiments, the genome of the host cell comprises from 5 to 100 integrated docking sites, each docking site comprising at least one dock site insertion element. In other preferred embodiments, the genome of the host cell comprises from 5 to 50 integrated docking sites, each docking site comprising at least one dock site insertion element. In some preferred embodiments, the integrated docking sites are independent integrated docking sites that are separated from one another and positioned at independent sites within the genome. For example, the integrated docking sites may preferably be spread across a number of chromosomes in the genome. In other embodiments, the integrated docking sites may be present as concatemers which comprise multiple copies of the same DNA sequence linked in series. The integrated docking sites preferably comprise one or more insertion elements (which may be termed a “dock site insertion element.” The dock site insertion elements are preferably nucleic acid sequences that facilitate insertion of a nucleic acid sequence encoding a protein of interest at the dock site. Nucleic acid constructs that can be inserted into the dock sites in the host cells of the present invention are described in detail below. The present invention is not limited to the use of any particular insertion elements. Indeed the use of a variety of insertion elements is contemplated. In some preferred embodiments, the insertion element is a recombinase dock site insertion element. Recombinase dock site insertion elements are nucleic acid sequences that are recognized and utilized by recombinase enzymes. For example, in some preferred embodiments, the recombinase dock site insertion element comprises an attachment site (att). In some particularly preferred embodiments, theAttorney Docket No. CATA-40729.601 attachment site is attP. These attachment sites are utilized by the PhiC31 integrase, which is a recombinase enzyme and which can be provided in the host cell via a vector in preferred embodiments. These dock sites serve as acceptors for integration of nucleic acid constructs comprising an attB attachment site. In other preferred embodiments, attR and attL attachment sites are utilized In other preferred embodiments, the recombinase dock site insertion element comprises an Flp Recombination Target (FRT) site. These sites are utilized by the enzyme flippase, which is a recombinase enzyme and which can be provided in the host cell via a vector in preferred embodiments. These dock sites serve as acceptors for integration of nucleic acid constructs comprising at the FRT site. In other preferred embodiments, the recombinase dock site insertion element comprises a LoxP site. These sites are utilized by the Cre recombinase which can be provided in the host cell via a vector in preferred embodiments. These dock sites serve as acceptors for integration of nucleic acid constructs comprising the LoxP site. In other preferred embodiments, the insertion element is an HDR (homology directed repair) dock site insertion element. HDR dock site insertion elements are nucleic acid sequences that provide an area of homology (a “homology arm”) that base pair with corresponding homology arms on the nucleic acid construct that is inserted at the site. These systems are preferably used with endonucleases that introduce double stranded breaks at a targeted site or sites, preferably flanked by the homology arms. In some embodiments, the HDR dock site insertion element is an AAVS1 safe harbor locus. In these embodiments, the dock site is used utilized by the Rep 78 endonuclease (nickase) which may be introduced into the host cell via a vector. The Rep 78 protein nickase promotes site-specific integration of nucleic acid sequences bearing homology arms corresponding to the AAVS1 safe harbor locus. In other preferred embodiments, the HDR dock site insertion element comprises one or more homology arms that are exogenous sequences of from 30 to 1000 base pairs in length. These dock sites are preferably used in conjunction with CRISPR gene editing systems. In some embodiments, the dock site further comprises one or more sequences that are homologous to guide RNA sequences. In these embodiments, the nucleic acid construct that is inserted at the dock site preferably comprises homology arms that are homologous to and base pair with the homology arms in the dock site. For utilization with CRISPR gene editing systems, a CRISPR gene editing system-compatible nuclease is introduced into the host cell. The CRISPR gene editing system-compatible nuclease may be a wild-typeAttorney Docket No. CATA-40729.601 endonuclease that creates a double-stranded break at a position determined by the guide RNA (and within the docking site) or a mutated nuclease (i.e., a nickase) that creates a single stranded break at a staggered positions within the dock site defined by two guide RNAs. Suitable nucleases are described in detail below in the discussion of nucleic acid expression constructs. In some preferred embodiments, the docking site may preferably comprise a suitable promoter so that a promoter trap scheme is utilized when suitable nucleic acid constructs are introduced at the docking site. Suitable promoters include, but are not limited to, SIN-LTR, SV40, EF1α, E. coli lac, E. coli trp, phage lambda PL, phage lambda PR, T3, T7, cytomegalovirus (CMV) immediate early, herpes simplex virus (HSV) thymidine kinase, alpha-lactalbumin, and mouse metallothionein-I promoter sequences. In some preferred embodiments the promoter sequence is oriented at the dock site so that the promoter will drive expression from an inserted nucleic acid construct. In some preferred embodiments, the promoter is oriented 5’ to the docking site. In some particularly preferred embodiments, the promoter is a SIN LTR. In these embodiments, the SIN-LTR and EPR are positioned 5’ to the dock site and a SIN LTR is positioned 3’ to the dock site. The docking sites may be introduced into any suitable host cell line. Suitable host cell lines include, but are not limited to, Chinese hamster ovary cells (CHO-K1, ATCC CCl-61); bovine mammary epithelial cells (ATCC CRL 10274; bovine mammary epithelial cells); monkey kidney CV1 line transformed by SV40 (COS-7, ATCC CRL 1651); human embryonic kidney line (293 or 293 cells subcloned for growth in suspension culture; see, e.g., Graham et al., J. Gen Virol., 36:59
[1977] ); baby hamster kidney cells (BHK, ATCC CCL 10); mouse sertoli cells (TM4, Mather, Biol. Reprod.23:243-251
[1980] ); monkey kidney cells (CV1 ATCC CCL 70); African green monkey kidney cells (VERO-76, ATCC CRL- 1587); human cervical carcinoma cells (HELA, ATCC CCL 2); canine kidney cells (MDCK, ATCC CCL 34); buffalo rat liver cells (BRL 3A, ATCC CRL 1442); human lung cells (W138, ATCC CCL 75); human liver cells (Hep G2, HB 8065); mouse mammary tumor (MMT 060562, ATCC CCL51); TRI cells (Mather et al., Annals N.Y. Acad. Sci., 383:44-68
[1982] ); MRC 5 cells; FS4 cells; rat fibroblasts (208F cells); MDBK cells (bovine kidney cells); CAP (CEVEC's Amniocyte Production) cells; and a human hepatoma line (Hep G2). In some particularly preferred embodiments, the host cells are modified so that they are deficient, or are naturally deficient, in an enzyme activity that is required for growth or survival of the cells in the presence of a selection agent and which is provided by the selectable marker. For example, Chinese Hamster Ovary (CHO) cells have been modified toAttorney Docket No. CATA-40729.601 be deficient for GS. In some preferred embodiments where vector includes a GS selectable marker, the host cell line is deficient in GS. In some particularly preferred embodiments, the GS deficient host cell line is the CHOZN® GS- / -cell line available from Merck KGaA. In other embodiments, where the selectable marker is, for example, DHFR, the cell line may preferably be deficient for DHFR activity (i.e., DHFR-). Suitable DHFR- cell lines include but are not limited to CHO-DG44 and derivatives thereof. The docking site sequences may be introduced into the host cells by any suitable genome modification system. In some preferred embodiments, the docking sites are incorporated into the host cells via the use of integrating vectors. The use of integrating vectors to introduce high copy numbers of a sequence of interest, such as a docking site, is described in detail in US Pat. Nos.6,852,510 and 7,332,333 as well as US Publ. Nos. 20030092882, 20030224415, 20040235173 and 20050100952, all which are incorporated herein by reference in their entirety. According to the present invention, host cells such as those described above are transduced or transfected with integrating vectors comprising a dock site under conditions such that multiple copies of the dock site are integrated into the genome of the host cell. Examples of integrating vectors include, but are not limited to, retroviral vectors, lentiviral vectors, adeno-associated viral vectors, and transposon vectors. The design, production, and use of these vectors in the present invention is described below. Retroviral Vectors. Retroviruses (family Retroviridae) are divided into three groups: the spumaviruses (e.g., human foamy virus); the lentiviruses (e.g., human immunodeficiency virus and sheep visna virus) and the oncoviruses (e.g., MLV, Rous sarcoma virus). Retroviruses are enveloped (i.e., surrounded by a host cell-derived lipid bilayer membrane) single-stranded RNA viruses which infect animal cells. When a retrovirus infects a cell, its RNA genome is converted into a double-stranded linear DNA form (i.e., it is reverse transcribed). The DNA form of the virus is then integrated into the host cell genome as a provirus. The provirus serves as a template for the production of additional viral genomes and viral mRNAs. Mature viral particles containing two copies of genomic RNA bud from the surface of the infected cell. The viral particle comprises the genomic RNA, reverse transcriptase and other pol gene products inside the viral capsid (which contains the viral gag gene products), which is surrounded by a lipid bilayer membrane derived from the host cell containing the viral envelope glycoproteins (also referred to as membrane-associated proteins).Attorney Docket No. CATA-40729.601 The organization of the genomes of numerous retroviruses is well known to the art and this has allowed the adaptation of the retroviral genome to produce retroviral vectors. The production of a recombinant retroviral vector carrying, for example, a docking site as described above, is typically achieved in two stages. First, the nucleic acid sequence encoding the docking site is inserted into a retroviral vector which contains the sequences necessary for the efficient integration (including promoter and / or enhancer elements which may be provided by the viral long terminal repeats (LTRs) or by an internal promoter / enhancer and relevant splicing signals), sequences required for the efficient packaging of the viral RNA into infectious virions (e.g., the packaging signal (Psi), the tRNA primer binding site (−PBS), the 3′ regulatory sequences required for reverse transcription (+PBS)) and the viral LTRs. The LTRs contain sequences required for the association of viral genomic RNA, reverse transcriptase and integrase functions, and sequences involved in directing the expression of the genomic RNA to be packaged in viral particles. For safety reasons, many recombinant retroviral vectors lack functional copies of the genes that are essential for viral replication (these essential genes are either deleted or disabled); therefore, the resulting virus is said to be replication defective. Second, following the construction of the recombinant vector, the vector DNA is introduced into a packaging cell line. Packaging cell lines provide proteins required in trans for the packaging of the viral genomic RNA into viral particles having the desired host range (i.e., the viral-encoded gag, pol and env proteins). The host range is controlled, in part, by the type of envelope gene product expressed on the surface of the viral particle. Packaging cell lines may express ecotrophic, amphotropic or xenotropic envelope gene products. Alternatively, the packaging cell line may lack sequences encoding a viral envelope (env) protein. In this case the packaging cell line will package the viral genome into particles that lack a membrane-associated protein (e.g., an env protein). In order to produce viral particles containing a membrane associated protein that will permit entry of the virus into a cell, the packaging cell line containing the retroviral sequences is transfected with sequences encoding a membrane-associated protein (e.g., the G protein of vesicular stomatitis virus (VSV)). The transfected packaging cell will then produce viral particles, which contain the membrane- associated protein expressed by the transfected packaging cell line; these viral particles, which contain viral genomic RNA derived from one virus encapsidated by the envelope proteins of another virus are said to be pseudotyped virus particles. The retroviral vectors of the present invention can be further modified to include additional regulatory sequences. As described above, the retroviral vectors of the presentAttorney Docket No. CATA-40729.601 invention include the following elements in operable association: a) a 5′ LTR; b) a packaging signal; c) a 3′ LTR and d) a nucleic acid encoding the docking site located between the 5′ and 3′ LTRs. Viral vectors, including recombinant retroviral vectors, provide a more efficient means of transferring genes into cells as compared to other techniques such as calcium phosphate-DNA co-precipitation or DEAE-dextran-mediated transfection, electroporation or microinjection of nucleic acids. It is believed that the efficiency of viral transfer is due in part to the fact that the transfer of nucleic acid is a receptor-mediated process (i.e., the virus binds to a specific receptor protein on the surface of the cell to be infected). In addition, the virally transferred nucleic acid once inside a cell integrates in controlled manner in contrast to the integration of nucleic acids which are not virally transferred; nucleic acids transferred by other means such as calcium phosphate-DNA co-precipitation are subject to rearrangement and degradation. The most commonly used recombinant retroviral vectors are derived from the amphotropic Moloney murine leukemia virus (MoMuLV) (See e.g., Miller and Baltimore Mol. Cell. Biol.6:2895
[1986] ). The MoMuLV system has several advantages: 1) this specific retrovirus can infect many different cell types, 2) established packaging cell lines are available for the production of recombinant MoMLV viral particles and 3) the transferred genes are permanently integrated into the target cell chromosome. The established MoMuLV vector systems comprise a DNA vector containing a small portion of the retroviral sequence (e.g., the viral long terminal repeat or “LTR” and the packaging or “psi” signal) and a packaging cell line. The gene to be transferred is inserted into the DNA vector. The viral sequences present on the DNA vector provide the signals necessary for the insertion or packaging of the vector RNA into the viral particle and for the expression of the inserted gene. The packaging cell line provides the proteins required for particle assembly (Markowitz et al., J. Virol.62:1120
[1988] ). Despite these advantages, existing retroviral vectors based upon MoMuLV are limited by several intrinsic problems: 1) they do not infect non-dividing cells (Miller et al., Mol. Cell. Biol.10:4239
[1990] ), except, perhaps, oocytes; 2) they produce low titers of the recombinant virus (Miller and Rosman, BioTechniques 7: 980
[1980] and Miller, Nature 357: 455
[1990] ); and 3) they infect certain cell types (e.g., human lymphocytes) with low efficiency (Adams et al., Proc. Natl. Acad. Sci. USA 89:8981
[1992] ). The low titers associated with MoMLV-based vectors have been attributed, at least in part, to the instabilityAttorney Docket No. CATA-40729.601 of the virus-encoded envelope protein. Concentration of retrovirus stocks by physical means (e.g., ultracentrifugation and ultrafiltration) leads to a severe loss of infectious virus. The low titer and inefficient infection of certain cell types by MoMuLV-based vectors has been overcome by the use of pseudotyped retroviral vectors, which contain the G protein of VSV as the membrane associated protein. Unlike retroviral envelope proteins that bind to a specific cell surface protein receptor to gain entry into a cell, the VSV G protein interacts with a phospholipid component of the plasma membrane (Mastromarino et al., J. Gen. Virol. 68:2359
[1977] ). Because entry of VSV into a cell is not dependent upon the presence of specific protein receptors, VSV has an extremely broad host range. Pseudotyped retroviral vectors bearing the VSV G protein have an altered host range characteristic of VSV (i.e., they can infect almost all species of vertebrate, invertebrate and insect cells). Importantly, VSV G- pseudotyped retroviral vectors can be concentrated 2000-fold or more by ultracentrifugation without significant loss of infectivity (Burns et al. Proc. Natl. Acad. Sci. USA 90:8033 The present invention is not limited to the use of the VSV G protein when a viral G protein is employed as the heterologous membrane-associated protein within a viral particle (See, e.g., U.S. Pat. No.5,512,421, which is incorporated herein by reference). The G proteins of viruses in the Vesiculovirus genera other than VSV, such as the Piry and Chandipura viruses, that are highly homologous to the VSV G protein and, like the VSV G protein, contain covalently linked palmitic acid (Brun et al. Intervirol.38:274
[1995] and Masters et al., Virol.171:285 (1990]). Thus, the G protein of the Piry and Chandipura viruses can be used in place of the VSV G protein for the pseudotyping of viral particles. In addition, the VSV G proteins of viruses within the Lyssa virus genera such as Rabies and Mokola viruses show a high degree of conservation (amino acid sequence as well as functional conservation) with the VSV G proteins. For example, the Mokola virus G protein has been shown to function in a manner similar to the VSV G protein (i.e., to mediate membrane fusion) and therefore may be used in place of the VSV G protein for the pseudotyping of viral particles (Mebatsion et al., J. Virol.69:1444
[1995] ). Viral particles may be pseudotyped using either the Piry, Chandipura or Mokola G protein as described in Example 2, with the exception that a plasmid containing sequences encoding either the Piry, Chandipura or Mokola G protein under the transcriptional control of a suitable promoter element (e.g., the CMV intermediate-early promoter; numerous expression vectors containing the CMV IE promoter are available, such as the pcDNA3.1 vectors (Invitrogen)) is used in place of pHCMV-G. Sequences encoding other G proteins derived from other members of theAttorney Docket No. CATA-40729.601 Rhabdoviridae family may be used; sequences encoding numerous rhabdoviral G proteins are available from the GenBank database. The majority of retroviruses can transfer or integrate a double-stranded linear form of the virus (the provirus) into the genome of the recipient cell only if the recipient cell is cycling (i.e., dividing) at the time of infection. Retroviruses that have been shown to infect dividing cells exclusively, or more efficiently, include MLV, spleen necrosis virus, Rous sarcoma virus and human immunodeficiency virus (HIV; while HIV infects dividing cells more efficiently, HIV can infect non-dividing cells). It has been shown that the integration of MLV virus DNA depends upon the host cell's progression through mitosis and it has been postulated that the dependence upon mitosis reflects a requirement for the breakdown of the nuclear envelope in order for the viral integration complex to gain entry into the nucleus (Roe et al., EMBO J.12:2099
[1993] ). However, as integration does not occur in cells arrested in metaphase, the breakdown of the nuclear envelope alone may not be sufficient to permit viral integration; there may be additional requirements such as the state of condensation of the genomic DNA (Roe et al., supra). Lentiviral Vectors. The present invention also contemplates the use of lentiviral vectors to generate cell lines with high numbers of integrated docking sites. The lentiviruses (e.g., equine infectious anemia virus, caprine arthritis-encephalitis virus, human immunodeficiency virus) are a subfamily of retroviruses that are able to integrate into non- dividing cells. The lentiviral genome and the proviral DNA have the three genes found in all retroviruses: gag, pol, and env, which are flanked by two LTR sequences. The gag gene encodes the internal structural proteins (e.g., matrix, capsid, and nucleocapsid proteins); the pol gene encodes the reverse transcriptase, protease, and integrase proteins; and the pol gene encodes the viral envelope glycoproteins. The 5′ and 3′ LTRs control transcription and polyadenylation of the viral RNAs. Additional genes in the lentiviral genome include the vif, vpr, tat, rev, vpu, nef, and vpx genes. A variety of lentiviral vectors and packaging cell lines are known in the art and find use in the present invention (See, e.g., U.S. Pat. Nos.5,994,136 and 6,013,516, both of which are herein incorporated by reference). Furthermore, the VSV G protein has also been used to pseudotype retroviral vectors based upon the human immunodeficiency virus (HIV) (Naldini et al., Science 272:263
[1996] ). Thus, the VSV G protein may be used to generate a variety of pseudotyped retroviral vectors and is not limited to vectors based on MoMLV. The lentiviral vectors may also be modified as described above to contain various regulatory sequencesAttorney Docket No. CATA-40729.601 (e.g., signal peptide sequences, RNA export elements, and IRES's). After the lentiviral vectors are produced, they may be used to transfect host cells as described above for retroviral vectors. Adeno-Associated Viral Vectors. The present invention also contemplates the use of adeno associated virus (AAV) vectors to generate cell lines with high numbers of integrated docking sites. AAV is a human DNA parvovirus, which belongs to the genus Adenovirus. The AAV genome is composed of a linear, single-stranded DNA molecule that contains approximately 4680 bases. The genome includes inverted terminal repeats (ITRs) at each end that function in cis as origins of DNA replication and as packaging signals for the virus. The internal nonrepeated portion of the genome includes two large open reading frames, known as the AAV rep and cap regions, respectively. These regions code for the viral proteins involved in replication and packaging of the virion. A family of at least four viral proteins are synthesized from the AAV rep region, Rep 78, Rep 68, Rep 52 and Rep 40, named according to their apparent molecular weight. The AAV cap region encodes at least three proteins, VP1, VP2 and VP3 (for a detailed description of the AAV genome, see e.g., Muzyczka, Current Topics Microbiol. Immunol.158:97-129
[1992] ; Kotin, Human Gene Therapy 5:793-801
[1994] ). AAV requires coinfection with an unrelated helper virus, such as adenovirus, a herpesvirus or vaccinia, in order for a productive infection to occur. In the absence of such coinfection, AAV establishes a latent state by insertion of its genome into a host cell chromosome. Subsequent infection by a helper virus rescues the integrated copy, which can then replicate to produce infectious viral progeny. Unlike the non-pseudotyped retroviruses, AAV has a wide host range and is able to replicate in cells from any species so long as there is coinfection with a helper virus that will also multiply in that species. Thus, for example, human AAV will replicate in canine cells coinfected with a canine adenovirus. Furthermore, unlike the retroviruses, AAV is not associated with any human or animal disease, does not appear to alter the biological properties of the host cell upon integration and is able to integrate into nondividing cells. It has also recently been found that AAV is capable of site- specific integration into a host cell genome. In light of the above-described properties, a number of recombinant AAV vectors have been developed for gene delivery (See, e.g., U.S. Pat. Nos.5,173,414; 5,139,941; WO 92 / 01070 and WO 93 / 03769, both of which are incorporated herein by reference; Lebkowski et al., Molec. Cell. Biol.8:3988-3996
[1988] ; Carter, Current Opinion in Biotechnology 3:533-539
[1992] ; Muzyczka, Current Topics in Microbiol. and Immunol.158:97-129Attorney Docket No. CATA-40729.601
[1992] ; Kotin, (1994) Human Gene Therapy 5:793-801; Shelling and Smith, Gene Therapy 1:165-169
[1994] ; and Zhou et al., J. Exp. Med.179:1867-1875
[1994] ). Recombinant AAV virions can be produced in a suitable host cell that has been transfected with both an AAV helper plasmid and an AAV vector. An AAV helper plasmid generally includes AAV rep and cap coding regions, but lacks AAV ITRs. Accordingly, the helper plasmid can neither replicate nor package itself. An AAV vector generally includes a selected gene of interest bounded by AAV ITRs that provide for viral replication and packaging functions. Both the helper plasmid and the AAV vector bearing the selected gene are introduced into a suitable host cell by transient transfection. The transfected cell is then infected with a helper virus, such as an adenovirus, which transactivates the AAV promoters present on the helper plasmid that direct the transcription and translation of AAV rep and cap regions. Recombinant AAV virions harboring the selected gene are formed and can be purified from the preparation. Once the AAV vectors are produced, they may be used to transfect (See, e.g., U.S. Pat. No.5,843,742, herein incorporated by reference) host cells at the desired multiplicity of infection to produce high copy number host cells. As will be understood by those skilled in the art, the AAV vectors may also be modified as described above to contain various regulatory sequences. Transposon vectors. The present invention also contemplates the use of transposon vectors to generate cell lines with high numbers of integrated docking sites. Transposons are mobile genetic elements that can move or transpose from one location another in the genome. Transposition within the genome is controlled by a transposase enzyme that is encoded by the transposon. Many examples of transposons are known in the art, including, but not limited to, Tn5 (See e.g., de la Cruz et al., J. Bact.175: 6932-38
[1993] , Tn7 (See e.g., Craig, Curr. Topics Microbiol. Immunol.204: 27-48
[1996] ), and Tn10 (See e.g., Morisato and Kleckner, Cell 51:101-111
[1987] ). The ability of transposons to integrate into genomes has been utilized to create transposon vectors (See, e.g., U.S. Pat. Nos.5,719,055; 5,968,785; 5,958,775; and 6,027,722; all of which are incorporated herein by reference.) Because transposons are not infectious, transposon vectors are introduced into host cells via methods known in the art (e.g., electroporation, lipofection, or microinjection). Therefore, the ratio of transposon vectors to host cells may be adjusted to provide the desired multiplicity of infection to produce the high copy number host cells of the present invention. Transposon vectors suitable for use in the present invention generally comprise a nucleic acid encoding a protein of interest interposed between two transposon insertion sequences. Some vectors also comprise a nucleic acid sequence encoding a transposaseAttorney Docket No. CATA-40729.601 enzyme. In these vectors, the one of the insertion sequences is positioned between the transposase enzyme and the nucleic acid encoding the protein of interest so that it is not incorporated into the genome of the host cell during recombination. Alternatively, the transposase enzyme may be provided by a suitable method (e.g., lipofection or microinjection). As will be understood by those skilled in the art, the transposon vectors may also be modified as described above to contain various regulatory sequences. Transfection at High Multiplicities of Infection. Once integrating vectors (e.g., retroviral vectors) encoding a docking site have been produced, they may be used to transfect or transduce host cells. Preferably, host cells are transfected or transduced with integrating vectors at a multiplicity of infection sufficient to result in the integration of at least 1, and preferably at least 2 or more retroviral vectors. In some embodiments, multiplicities of infection of from 10 to 1,000,000 may be utilized, so that the genomes of the infected host cells contain from 2 to 100 copies of the integrated vectors, and preferably from 5 to 50 copies of the integrated vectors. In other embodiments, a multiplicity of infection of from 10 to 10,000 is utilized. When non-pseudotyped retroviral vectors are utilized for infection, the host cells are incubated with the culture medium from the retroviral producers cells containing the desired titer (i.e., colony forming units, CFUs) of infectious vectors. When pseudotyped retroviral vectors are utilized, the vectors are concentrated to the appropriate titer by ultracentrifugation and then added to the host cell culture. Alternatively, the concentrated vectors can be diluted in a culture medium appropriate for the cell type. In each case, the host cells are exposed to medium containing the infectious retroviral vectors for a sufficient period of time to allow infection and subsequent integration of the vectors. In general, the amount of medium used to overlay the cells should be kept to as small a volume as possible so as to encourage the maximum amount of integration events per cell. As a general guideline, the number of colony forming units (cfu) per milliliter should be about 105to 107cfu / ml, depending upon the number of integration events desired. It is contemplated that the actual integration rate is dependent not only on the multiplicity of infection, but also on the contact time (i.e., the length of time the host cells are exposed to infectious vector), the confluency or geometry of the host cells being transfected, and the volume of media that the vectors are contained in. It is contemplated that these conditions can be varied as taught herein to produce host cell lines containing multiple integrated copies of integrating vectors. In some embodiments, after transfection or transduction, the cells are allowed to multiply, and are then trypsinized and re-plated. Individual colonies are then selected toAttorney Docket No. CATA-40729.601 provide clonally selected cell lines. In still further embodiments, the clonally selected cell lines are screened by Southern blotting or INVADER assay to verify that the desired number of integration events has occurred. It is also contemplated that clonal selection allows the identification of superior protein producing cell lines. In other embodiments, the cells are not clonally selected following transfection. In still further embodiments, cell lines are serially transfected with vectors encoding the same docking site. In some preferred embodiments, the host cells are transfected (e.g., at an MOI of about 10 to 100,000, preferably 100 to 10,000) with an integrating vector encoding a docking site, cell lines containing single or multiple integrated copies of the integrating vector are selected (e.g., clonally selected), and the selected cell line is re- transfected with the vector (e.g., at an MOI of about 10 to 100,000, preferably 100 to 10,000). This process may be repeated multiple times until the desired level of protein expression is obtained and may also be repeated to introduce vectors encoding multiple proteins of interest. The present invention contemplates a variety of serial transfection procedures. In some embodiments, where retroviral vectors are utilized, serial transduction procedures are provided. In preferred embodiments, serial transduction is carried out on a pool of cells. In these embodiments, an initial pool of host cells is contacted with retroviral vectors, preferably at a multiplicity of infection ranging from about 0.5 to about 1000 vectors / host cell. The cells are then cultured for several days in an appropriate medium. An aliquot of the cells in then taken to determine the number of integrated vectors and to freeze for future possible use. The remaining cells are then re-contacted with retroviral vectors, again preferably at a multiplicity of infection ranging from about 0.5 to about 1000 vectors / host cell. This process is repeated until cells with a desired number if integrated vectors are obtained. For example, the process can be repeated up to 10 to 20 or more times. In some embodiments, cells can be clonally selected after any particular transduction step if so desired, however, utilizing a pool of cells in the absence of transduction results in a decreased time to the desired integrated vector copy number. Following the serial transduction process, cell lines are clonally selected and analyzed for integrated vector copy number and protein production characteristics. Superior cell lines are chosen and stored in a master cell bank. Nucleic Acid Expression Constructs. In some preferred embodiments, nucleic acid constructs for expression of a protein of interest are introduced into the host cell lines containing multiple docking sites. As discussed above, in preferred embodiments, the nucleic acid constructs preferably comprise nucleic acid sequences (which may be termedAttorney Docket No. CATA-40729.601 “expression construct insertion elements”) that are compatible with the dock site insertion elements as described above. Accordingly, in some preferred embodiments, the present invention provides nucleic acid expression constructs for use in expressing a protein or proteins of interest in a host cell, and in particular to expression of two or more proteins of interest where the nucleic acid expression constructs encoding the two or more proteins of interest are integrated into the genome of the host cell at desired ratios as described in detail above. In some preferred embodiments, where the dock site does not comprise a promoter, the nucleic acid expression constructs, for example, comprise the following elements in operable association, most preferably in 5’ to 3’ order: first promoter sequence - selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a first protein of interest - poly A signal sequence. first promoter sequence - selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a second protein of interest - poly A signal sequence. first promoter sequence - selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a third protein of interest - poly A signal sequence. first promoter sequence - selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a fourth protein of interest - poly A signal sequence. In some preferred embodiments, where the dock site comprises an exogenous promoter, the nucleic acid expression constructs, for example, comprise the following elements in operable association, most preferably in 5’ to 3’ order: selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a first protein of interest - poly A signal sequence. selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a second protein of interest - poly A signal sequence.Attorney Docket No. CATA-40729.601 selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a third protein of interest - poly A signal sequence. selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a fourth protein of interest - poly A signal sequence. In some preferred embodiments, the constructs of the invention do not comprise a poly A signal sequence between the selectable marker sequence and second promoter sequence. The present invention is not limited to any particular mechanism of action. Indeed, an understanding of the mechanism of action is not necessary to practice the present invention. Nevertheless, constructs which lack a poly A signal sequence after the selectable marker have been found to provide for better selection and production of the protein of interest in host cell cultures. In still other preferred embodiments, the selectable marker is adjacent to the second promoter. In still other preferred embodiments, the second promoter is adjacent to the nucleic acid sequence encoding the first protein of interest. In this context, the term “adjacent” means that there is no intervening functional element or intron between the listed components. In some particularly preferred embodiments, the nucleic acid expression constructs further comprises at least one expression construct insertion element at a position or positions selected from the group consisting of 5’ to the first promoter, 3’ to the poly A signal sequence, between the first promoter and the poly A signal sequence, between the selectable marker and the second promoter sequence, and both 5’ to the first promoter and 3’ to the poly A signal sequence. Suitable constructs are shown in the following non-limiting examples: expression construct insertion element - first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - second (i.e., internal) promoter sequence - nucleic acid sequence encoding a first protein of interest - poly A signal sequence first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a first protein of interest - poly A signal sequence – expression construct insertion elementAttorney Docket No. CATA-40729.601 expression construct insertion element - first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a first protein of interest - poly A signal sequence – expression construct insertion element. first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence)- selectable marker sequence - expression construct insertion element - second promoter sequence - nucleic acid sequence encoding a first protein of interest - poly A signal sequence. expression construct insertion element - first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - second (i.e., internal) promoter sequence - nucleic acid sequence encoding a second protein of interest - poly A signal sequence first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a second protein of interest - poly A signal sequence – expression construct insertion element expression construct insertion element - first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a second protein of interest - poly A signal sequence – expression construct insertion element. first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence)- selectable marker sequence - expression construct insertion element - second promoter sequence - nucleic acid sequence encoding a second protein of interest - poly A signal sequence. In some preferred embodiments, the constructs may include nucleic acid sequences encoding multiple proteins of interest, for example 2, 3 ,4 or 5 (or more) proteins of interest.Attorney Docket No. CATA-40729.601 Suitable constructs for expressing two proteins of interest are shown in the following nonlimiting examples. These expression constructs may be used at different ratios in conjunction with expression constructs encoding an additional third protein of interest, or as exemplified below, third and fourth proteins of interest. expression construct insertion element - first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - second (i.e., internal) promoter sequence - nucleic acid sequence encoding a first protein of interest – WPRE (optional) – poly A signal sequence – third promoter sequence or IRES - nucleic acid sequence encoding a second protein of interest – WPRE (optional) - poly A signal sequence first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a first protein of interest - WPRE (optional) – poly A signal sequence – third promoter sequence – intron (optional) - nucleic acid sequence encoding a second protein of interest – WPRE (optional) - poly A signal sequence – expression construct insertion element expression construct insertion element - first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a first protein of interest - WPRE (optional) – poly A signal sequence – third promoter sequence – intron (optional) - nucleic acid sequence encoding a second protein of interest – WPRE (optional) - poly A signal sequence – expression construct insertion element. first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - expression construct insertion element - second promoter sequence - nucleic acid sequence encoding a first protein of interest - WPRE – poly A signal sequence – third promoter sequence or IRES - nucleic acid sequence encoding a second protein of interest – WPRE - poly A signal sequenceAttorney Docket No. CATA-40729.601 expression construct insertion element - first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a first protein of interest – WPRE (optional) – poly A signal sequence – third promoter sequence - nucleic acid sequence encoding a second protein of interest – WPRE (optional) - poly A signal sequence – expression construct insertion element. expression construct insertion element - first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a first protein of interest - WPRE (optional) – poly A signal sequence – third promoter sequence – intron- nucleic acid sequence encoding a second protein of interest – WPRE (optional) - poly A signal sequence – expression construct insertion element. expression construct insertion element - first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - second (i.e., internal) promoter sequence - nucleic acid sequence encoding a third protein of interest – WPRE (optional) – poly A signal sequence – third promoter sequence or IRES - nucleic acid sequence encoding a fourth protein of interest – WPRE (optional) - poly A signal sequence first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a third protein of interest - WPRE (optional) – poly A signal sequence – third promoter sequence – intron (optional) - nucleic acid sequence encoding a fourth protein of interest – WPRE (optional) - poly A signal sequence – expression construct insertion element expression construct insertion element - first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a third protein of interest - WPRE (optional) – poly A signal sequence –Attorney Docket No. CATA-40729.601 third promoter sequence – intron (optional) - nucleic acid sequence encoding a fourth protein of interest – WPRE (optional) - poly A signal sequence – expression construct insertion element. first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - expression construct insertion element - second promoter sequence - nucleic acid sequence encoding a third protein of interest - WPRE – poly A signal sequence – third promoter sequence or IRES - nucleic acid sequence encoding a fourth protein of interest – WPRE - poly A signal sequence expression construct insertion element - first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a third protein of interest – WPRE (optional) – poly A signal sequence – third promoter sequence - nucleic acid sequence encoding a fourth protein of interest – WPRE (optional) - poly A signal sequence – expression construct insertion element. expression construct insertion element - first promoter sequence (optional depending on whether the dock site already comprises an exogenous promoter sequence) - selectable marker sequence - second promoter sequence - nucleic acid sequence encoding a third protein of interest - WPRE (optional) – poly A signal sequence – third promoter sequence – intron- nucleic acid sequence encoding a fourth protein of interest – WPRE (optional) - poly A signal sequence – expression construct insertion element. In some embodiments, the mixtures of different constructs are utilized. In some preferred embodiments, the mixture of different constructs may comprise constructs as described above and constructs starting with the internal or second promoter (i.e., starting after and not including the selectable marker). It is contemplated that by using mixtures of constructs, some which do not include selectable markers, higher insertion rates may be achieved.Attorney Docket No. CATA-40729.601 However, any suitable proteins of interest may be expressed via the host cells, constructs and systems of the present invention. Exemplary proteins of interest include immunoglobulins, single chain antibodies, anticoagulant proteins, blood factor proteins, bone morphogenetic proteins, engineered protein scaffolds, enzymes, Fc fusion proteins, growth factors, hormones, interferons, interleukins, antigens, and thrombolytic proteins. In other preferred embodiments, the constructs of the present invention may be utilized to express viral vectors. In these embodiments, the protein of interest sequence described in the exemplary vectors above is replaced with a nucleic acid sequence encoding a viral vector backbone. Viral vector expression sequences that may be included in the constructs of the present invention include, but are not limited to, retroviral vectors, lentiviral vectors, adenoviral vectors and AAV vectors as described elsewhere herein In some preferred embodiments, the retroviral vectors themselves include a nucleic acid sequence encoding a protein of interest as described above that is expressed by the vector. In some particularly preferred embodiments, the protein of interest that is expressed by the vector is an antigen sequence for use in a vaccine. In some preferred embodiments, the expression construct insertion elements are elements that find use in conjunction with or are recognized by transposons, integrases, recombinases or CRISPR systems. Suitable insertion elements include, but are not limited to, inverted terminal repeats, integrase attachment sites (att), and homologous recombination arms which in the context of the constructs described herein can be described as homologous recombination insertion elements. The nucleic acid constructs may be utilized with many different vectors and vectors systems. These vectors and vectors system may preferably be used to introduce the nucleic acid expression constructs into the host cells described above. Suitable vectors and vectors systems include, but are not limited to, viral gene insertion technologies such as retroviral, lentiviral and AAV systems as well as non-viral gene insertion technologies such as transposase, recombinase, integrase or CRISPR gene insertion. Specific examples of technologies / enzymes that can be used with nucleic acid constructs of the present invention include piggyback transposase systems, sleeping beauty transposase systems, Mos1 transposase systems, Tol2 transposase systems, Leapin transposase systems, Lambda recombinase systems, FLP / FRT systems, Cre / Lox systems, MMLV integrase systems, Rep 78 integrase systems and CRISPR systems which can include nucleases or nickases as well as guide sequences. In some preferred embodiments, the system is a nucleic acid integrationAttorney Docket No. CATA-40729.601 system with the proviso that the system is not a retroviral or lentiviral systems utilizing a retroviral or lentiviral LTR. As discussed above, in some preferred embodiments, the expression construct insertion element comprises an attachment site (att). In some particular preferred embodiments, the attachment site is attB. These attachment sites are utilized by the PhiC31 integrase, which is a recombinase enzyme and which can be provided in the host cell via a vector in preferred embodiments. These sites facilitate integration of the nucleic acid constructs into a dock site comprising attP attachment site. In other preferred embodiments, attR and attL attachment sites may be utilized. In other preferred embodiments, the expression construct insertion element comprises an Flp Recombination Target (FRT) site. These sites are utilized by the enzyme flippase, which is a recombinase enzyme and which can be provided in the host cell via a vector in preferred embodiments. These sites serve facilitate integration of nucleic acid constructs into dock sites comprising corresponding FRT sites. In other preferred embodiments, the expression construct insertion element comprises a LoxP site. These sites are utilized by the Cre recombinase which can be provided in the host cell via a vector in preferred embodiments. These sites facilitate integration of nucleic acid constructs into dock sites comprising corresponding LoxP sites. In other preferred embodiments, the expression construct insertion element is an HDR (homology directed repair) expression construct insertion element. HDR expression construct insertion elements are nucleic acid sequences that provide an area of homology (a “homology arm”) that base pair with corresponding homology arms in the dock site. These systems are preferably used with endonucleases that introduce double stranded breaks at a targeted site or sites, preferably flanked by the homology arms. In some embodiments, the HDR expression construct insertion element comprises AAVS1 safe harbor locus homology arms. In these embodiments, the expression construct is specifically integrated in a dock site comprising the AAVS1 safe harbor locus. The integration is facilitated by the Rep 78 endonuclease (nickase) which may be introduced into the host cell via a vector. The Rep 78 protein nickase promotes site-specific integration of nucleic acid sequences bearing homology arms corresponding to the AAVS1 safe harbor locus. In other preferred embodiments, the HDR expression construct insertion element comprises one or more homology arms that are exogenous sequences of from 30 to 1000 base pairs in length. These expression constructs are preferably used in conjunction with CRISPR gene editing systems. In these embodiments, the nucleic acid construct is inserted at dockAttorney Docket No. CATA-40729.601 sites that comprise homology arms that are homologous to and base pair with the homology arms in the nucleic acid construct. For utilization with CRISPR gene editing systems, a CRISPR gene editing system-compatible nuclease is introduced into the host cell. The CRISPR gene editing system-compatible nuclease may be a wild-type endonuclease that creates a double-stranded break at a position determined by the guide RNA (and within the docking site) or a mutated nuclease (i.e., a nickase) that creates a single stranded break at a staggered positions within the dock site defined by two guide RNAs. Suitable nucleases are described in detail below in the discussion of nucleic acid expression constructs. As discussed above, integration at the dock sites generally requires expression of an exogenous enzyme in the host cell. Suitable enzymes include, but are not limited to, recombinases (including integrases), endonucleases, and nickases. Accordingly, in some embodiments, host cells of the present invention comprise an exogenous nucleic acid sequence (or expression construct) for expression of a recombinase (including integrases), a endonuclease, and a nickase In some embodiments, constructs for expressing the exogenous enzymes may be stably integrated into the genome of the host cell. In other embodiments, vectors for expressing the exogenous enzymes are transiently introduced into the host cell, for example with an extrachromosomal vector such as a plasmid. In some embodiments, both the vectors comprising exogenous enzyme and the vectors comprising the nucleic acid constructs for expression of the protein of interest are transiently introduced into the host cell, for example by transfection. In these embodiments, the preferred ratio of the vectors encoding the exogenous enzyme to the gene of interest vectors is from 1:1000 to 1:10. In some more preferred embodiments, the ratio is from 1:100 to 1:750. In some still more preferred embodiments, the ratio is from 1:400 to 1:600. This is surprising as the literature for other integrase systems generally indicates that a higher level of vector encoding the exogenous enzyme to the gene of interest construct is required. In some preferred embodiments, the integrase is the phiC31 integrase (BioCat GmbH, Heidelberg, DE or System Biosciences, Palo Alto, CA)). The phiC31 integrase is a sequence-specific recombinase encoded within the genome of the bacteriophage phiC31. The phiC31 integrase mediates recombination between two 34 base pair sequences termed attachment sites (att), one found in the phage and the other in the host. This serine integrase has been shown to function efficiently in many different cell types including mammalian cells. In the presence of phiC31 integrase, an attB- containing donor plasmid can be unidirectional integrated into a target genome through recombination at sites with sequence similarity to the native attP site (termed pseudo-attP sites). phiC31 integrase can integrate aAttorney Docket No. CATA-40729.601 plasmid of any size, as a single copy, and requires no cofactors. The integrated transgenes are stably expressed and heritable. Other suitable recombinase-based systems include CRISPR gene editing systems, CRE-Lox, FLP-FRT, and lambda recombinase systems. Cre-Lox recombination is a site-specific recombinase technology, used to carry out deletions, insertions, translocations and inversions at specific sites in the DNA of cells. It allows the DNA modification to be targeted to a specific cell type or be triggered by a specific external stimulus. It is implemented both in eukaryotic and prokaryotic systems. The Cre-lox recombination system has been particularly useful to help neuroscientists to study the brain in which complex cell types and neural circuits come together to generate cognition and behaviors. The system consists of a single enzyme, Cre recombinase, that recombines a pair of short target sequences called the Lox sequences. This system can be implemented without inserting any extra supporting proteins or sequences. The Cre enzyme and the original Lox site called the LoxP sequence are derived from bacteriophage P1. See, e.g., Targeted integration of DNA using mutant lox sites in embryonic stem cells. Araki, et al. Nucleic Acids Res, Feb 1997, Vol.25, Issue 4, pp.868-872; High-Resolution Labeling and Functional Manipulation of Specific Neuron Types in Mouse Brain by Cre-Activated Viral Gene Expression. Kuhlman, et al. PLos One, Apr 2008, Vol.3, e2005; When reverse genetics meets physiology: the use of site-specific recombinases in mice. Tronche, et al. FEBS Letters, Aug 2002, Vol.529, Issue 1, pp.116-121. The FLP-FRT recombination system is another site-directed recombination technology very conceptually similar to Cre-lox, with flippase (Flp) and the short flippase recognition target (FRT) site being analogous to Cre and loxP, respectively. See, e.g., Candice et al., Cre / loxP, Flp / FRT Systems and Pluripotent Stem Cell Lines (2012) Topics in Current Genetics, vol 23. The FLP-FRT technology can be an effective alternative to Cre-lox, and has also been used in conjunction with it, allowing for two separate recombination events to be controlled in parallel. The nucleic acid constructs of the present invention may be used in conjunction with CRISPR homologous recombination (HDR) systems. HDR is initiated by the presence of double strand breaks (DSBs) in DNA. The CRISPR / Cas9 system is preferably used to create targeted double stranded breaks via a guide RNA sequence so that the nucleic acid construct of the invention can be inserted. See, e.g., Zhang et al., Efficient precise knockin with a double cut HDR donor after CRISPR / Cas9-mediated double-stranded DNA cleavage (2017) Genome Biol.18:35; Mali et al., Cas9 as a versatile tool for engineering biology. NatureAttorney Docket No. CATA-40729.601 Methods10, 957–963 (2013); Mali et al., RNA-Guided Human Genome Engineering via Cas9. Science339(6121), 823-826 (2013); Ran et al., Double nicking by RNA-guided CRISPR Cas9 for enhanced genome editing specificity. Cell, 155(2), 479-480(2013). Suitable guide RNA sequences (gRNAs) may be designed as is known in the art. In some preferred embodiments, CRISPR systems for HDR utilize either one or two guide sequences. When one guide RNA sequence is utilized, it preferred to use a nuclease such as a Cas9 nuclease which makes a single double stranded break guided by the guide RNA sequence. When two guide sequences are utilized, it is preferred to use a nickase, which can be a mutated Cas9 nuclease which only makes single stranded breaks in the target DNA sequence guided by each of the guide RNA sequences. The single stranded breaks are preferably positioned at staggered points on different strands (i.e., the sense and antisense strands) of the target DNA sequence. This arrangement generally improves HDR efficiency. In general, “CRISPR system” refers collectively to transcripts and other elements involved in the expression of or directing the activity of CRISPR-associated (“Cas”) genes, including sequences encoding a Cas gene, a tracr (trans-activating CRISPR) sequence (e.g. tracrRNA or an active partial tracrRNA), a tracr-mate sequence (encompassing a “direct repeat” and a tracrRNA-processed partial direct repeat in the context of an endogenous CRISPR system), a guide sequence (also referred to as a “spacer” in the context of an endogenous CRISPR system), or other sequences and transcripts from a CRISPR locus. In some embodiments, one or more elements of a CRISPR system is derived from a type I, type II, or type III CRISPR system. In some embodiments, one or more elements of a CRISPR system is derived from a particular organism comprising an endogenous CRISPR system, such as Streptococcus pyogenes. In general, a CRISPR system is characterized by elements that promote the formation of a CRISPR complex at the site of a target sequence (also referred to as a protospacer in the context of an endogenous CRISPR system). In the context of formation of a CRISPR complex, “target sequence” refers to a sequence to which a guide sequence is designed to have complementarity, where hybridization between a target sequence and a guide sequence promotes the formation of a CRISPR complex. Full complementarity is not necessarily required, provided there is sufficient complementarity to cause hybridization and promote formation of a CRISPR complex. A target sequence may comprise any polynucleotide, such as DNA or RNA polynucleotides. In some embodiments, a target sequence is located in the nucleus or cytoplasm of a cell. In some embodiments, the target sequence may be within an organelle of a eukaryotic cell, for example, mitochondrion or chloroplast. A sequence or template that may be used for recombination into the targetedAttorney Docket No. CATA-40729.601 locus comprising the target sequences is referred to as an “editing template” or “editing polynucleotide” or “editing sequence”. In aspects of the invention, an exogenous template polynucleotide may be referred to as an editing template. In an aspect of the invention the recombination is homologous recombination. Typically, in the context of an endogenous CRISPR system, formation of a CRISPR complex (comprising a guide sequence hybridized to a target sequence and complexed with one or more Cas proteins) results in cleavage of one or both strands in or near (e.g. within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs from) the target sequence. Without wishing to be bound by theory, the tracr sequence, which may comprise or consist of all or a portion of a wild-type tracr sequence (e.g. about or more than about 20, 26, 32, 45, 48, 54, 63, 67, 85, or more nucleotides of a wild-type tracr sequence), may also form part of a CRISPR complex, such as by hybridization along at least a portion of the tracr sequence to all or a portion of a tracr mate sequence that is operably linked to the guide sequence. In some embodiments, the tracr sequence has sufficient complementarity to a tracr mate sequence to hybridize and participate in formation of a CRISPR complex. As with the target sequence, it is believed that complete complementarity is not needed, provided there is sufficient to be functional. In some embodiments, the tracr sequence has at least 50%, 60%, 70%, 80%, 90%, 95% or 99% of sequence complementarity along the length of the tracr mate sequence when optimally aligned. In some embodiments, one or more vectors driving expression of one or more elements of a CRISPR system are introduced into a host cell such that expression of the elements of the CRISPR system direct formation of a CRISPR complex at one or more target sites. For example, a Cas enzyme, a guide sequence linked to a tracr-mate sequence, and a tracr sequence could each be operably linked to separate regulatory elements on separate vectors. Alternatively, two or more of the elements expressed from the same or different regulatory elements, may be combined in a single vector, with one or more additional vectors providing any components of the CRISPR system not included in the first vector. CRISPR system elements that are combined in a single vector may be arranged in any suitable orientation, such as one element located 5′ with respect to (“upstream” of) or 3′ with respect to (“downstream” of) a second element. The coding sequence of one element may be located on the same or opposite strand of the coding sequence of a second element, and oriented in the same or opposite direction. In some embodiments, a single promoter drives expression of a transcript encoding a CRISPR enzyme and one or more of the guide sequence, tracr mate sequence (optionally operably linked to the guide sequence), and a tracr sequence embedded within one or more intron sequences (e.g. each in a different intron, two or more in at leastAttorney Docket No. CATA-40729.601 one intron, or all in a single intron). In some embodiments, the CRISPR enzyme, guide sequence, tracr mate sequence, and tracr sequence are operably linked to and expressed from the same promoter. Non-limiting examples of Cas proteins useful in the present invention include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified versions thereof. These enzymes are known; for example, the amino acid sequence of S. pyogenes Cas9 protein may be found in the SwissProt database under accession number Q99ZW2. In some embodiments, the unmodified CRISPR enzyme has DNA cleavage activity, such as Cas9. In some embodiments the CRISPR enzyme is Cas9, and may be Cas9 from S. pyogenes or S. pneumoniae. In some embodiments, the CRISPR enzyme directs cleavage of one or both strands at the location of a target sequence, such as within the target sequence and / or within the complement of the target sequence. In some embodiments, the CRISPR enzyme directs cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of a target sequence. In some embodiments, a vector encodes a CRISPR enzyme that is mutated to with respect to a corresponding wild-type enzyme such that the mutated CRISPR enzyme lacks the ability to cleave one or both strands of a target polynucleotide containing a target sequence. For example, an aspartate-to-alanine substitution (D10A) in the RuvC I catalytic domain of Cas9 from S. pyogenes converts Cas9 from a nuclease that cleaves both strands to a nickase (cleaves a single strand). Other examples of mutations that render Cas9 a nickase include, without limitation, H840A, N854A, and N863A. In aspects of the invention, nickases may be used for genome editing via homologous recombination. In some preferred embodiments, the HDR insertion element comprises AAVS1 safe harbor locus homology arms and are used in conjunction with Rep 78 endonuclease (nickase). The adeno-associated virus serotype 2 (AAV2) Rep 78 protein is a strand-specific endonuclease (nickase) that promotes site-specific integration of transgene sequences bearing homology arms corresponding to the AAVS1 safe harbor locus. See, e.g., Ramachandra et al., Efficient recombinase-mediated cassette exchange at the AAVS1 locus in human embryonic stem cells using baculoviral vectors (2011) Nucleic Acids Research, 39(16):e107; WO1998027207).Attorney Docket No. CATA-40729.601 As indicated above, in some preferred embodiments, the nucleic acid constructs of the present invention comprise an optional first and a second promoter sequence. The first and second promoter sequences may be the same or different. Suitable first and second promoter sequences include, but are not limited to the MMLV LTR promoter, the MoMuSV LTR promoter, the RSV LTR promoter, the SIN LTR promoter, the SV40 promoter, cytomegalovirus (CMV) immediate early promoter, herpes simplex virus (HSV) thymidine kinase promoter, alpha-lactalbumin promoter, mouse metallothionein-I promoter, dihydrofolate reductase promoter, the β-actin promoter, phosphoglycerol kinase (PGK) promoter, and the EF1α promoter sequences, and combinations thereof. In some preferred embodiments, the first promoter sequence is not a retroviral LTR promoter, i.e., the first promoter is promoter sequence other than a retroviral LTR promoter sequence. However, when the promoter is a retroviral promoter sequence, it may be a SIN (self-inactivating) LTR promoter sequence. See, e.g., co-pending application PCT / US2019 / 064423, which is incorporated herein by reference in its entirety. Suitable Sin LTR promotors are known in the art and are prepared by removing either all or a portion of the U3 region of the LTR. As described in PCT / US2019 / 064423, in some preferred embodiments the first promoter which drives selectable marker is a weak promoter. In some preferred embodiments, a weak promoter is a promoter, preferably a constitutive promoter, that has activity that equal to or less than the activity of the SIN LTR promoter in a host of interest (e.g., a CHO cell) when operably linked to a selectable maker sequence. In still other preferred embodiments, a weak promoter is a promoter, preferably a constitutive promoter, that has activity that equal to or less than the activity of the human Ubiquitin C (UBC) promoter in a host of interest (e.g., a CHO cell) when operably linked to a selectable maker sequence. Suitable methods for assessing promoter strength are known in the art. See, e.g., Dandindorj et al. (2014) A Comparative Analysis of Constitutive Promoters Located in Adeno-Associated Viral Vectors, PLoS One 9(8): e106472; Zhang and Baum (2005) Evaluation of Viral and Mammalian Promoters for Use in Gene Delivery to Salivary Glands Mol. Ther.12(3):528-536; Qin et al. (2010) Systematic Comparison of Constitutive Promoters and the Doxycycline-Inducible Promoter PLoS 5(5): e10611; Jeyaseelan et al. (2001) Real-time detection of gene promoter activity: quantitation of toxin gene transcription, Nucleic Acids Research.29 (12): 58e–58. In some embodiments, weak promoters have been altered to reduce promoter activity. Accordingly, in some preferred embodiments, the present invention provides vector(s) for expression of a protein of interest comprising a nucleic acid sequence encoding a selectable marker in operable association with a first weakAttorney Docket No. CATA-40729.601 promoter sequence or promoter sequence that has been altered to reduce promoter activity as compared to a non-altered or wild-type version of the first promoter sequence and a nucleic acid sequence encoding the protein of interest operably linked to a second promoter sequence. The SIN LTR promoter sequence is one such example. Other promoter sequences described above may also be altered to reduce activity and provide a weak promoter or the weak promoter may be naturally occurring weak promoter such as the UBC promoter. In some preferred embodiments, the nucleic acid constructs include a selectable marker. Suitable selectable markers include but are not limited to glutamine synthetase (GS), dihydrofolate reductase (DHFR) and the like. These genes are described in U.S. Pat. Nos. 5,770,359; 5,827,739; 4,399,216; 4,634,665; 5,149,636; and 6,455,275; all of which are incorporated herein by reference. In some preferred embodiments, the selectable marker that is utilized is compatible with a host cell line that is deficient in the production of the enzyme encoded by the selectable marker nucleic acid sequence. Suitable host cell lines are described in more detail below. In other embodiments, the selectable marker is an antibiotic resistance marker, i.e., a gene that produces a protein that provides cells expressing this protein with resistance to an antibiotic. Suitable antibiotic resistance markers include genes that provide resistance to neomycin (neomycin resistance gene (neo)), hygromycin (hygromycin B phosphotransferase gene), puromycin (puromycin N-acetyl-transferase), and the like. In other embodiments of the present invention, where secretion of the protein of interest is desired, the nucleic acid constructs include a signal peptide sequence in operable association with the protein of interest. The sequences of several suitable signal peptides are known to those in the art, including, but not limited to, those derived from tissue plasminogen activator, human growth hormone, lactoferrin, alpha-casein, and alpha-lactalbumin. In other embodiments of the present invention, the nucleic acid constructs include an RNA export element (See, e.g., U.S. Pat. Nos.5,914,267; 6,136,597; and 5,686,120; and WO99 / 14310, all of which are incorporated herein by reference) either 3' or 5' to the nucleic acid sequence encoding the protein of interest. It is contemplated that the use of RNA export elements allows high levels of expression of the protein of interest without incorporating splice signals or introns in the nucleic acid sequence encoding the protein of interest. In still other embodiments, the nucleic acid constructs include at least one internal ribosome entry site (IRES) sequence. The sequences of several suitable IRES's are available, including, but not limited to, those derived from foot and mouth disease virus (FDV), encephalomyocarditis virus, and poliovirus. The IRES sequence can be interposed betweenAttorney Docket No. CATA-40729.601 two transcriptional units (e.g., nucleic acids encoding different proteins of interest or subunits of a multi-subunit protein such as an antibody) to form a polycistronic sequence so that the two transcriptional units are transcribed from the same promoter. The present invention is not limited to expression of any particular protein of interest. In some preferred embodiments, the protein of interest is selected from the group consisting of an Fc-fusion protein, an enzyme, an albumin fusion, a growth factor, a protein receptor, a single chain antibody (scFv), a single chain-Fc (scFv-Fc), a diabody, and minibody (scFv- CH3), Fab, single chain Fab (scFab), an immunoglobulin heavy chain, and an immunoglobulin light chain and other antigen binding proteins. In general, the protein or proteins of interest may be any pharmaceutical or industrial protein for which expression and production via a host culture is desired. In some preferred embodiments, the nucleic acid constructs are incorporated into a nucleic acid expression vector. Vectors include, but are not limited to, nucleic acid molecules that are single-stranded, double-stranded, or partially double-stranded; nucleic acid molecules that comprise one or more free ends, no free ends (e.g. circular); nucleic acid molecules that comprise DNA, RNA, or both; and other varieties of polynucleotides known in the art. One type of vector is a “plasmid,” which refers to a circular double stranded DNA loop into which additional DNA segments can be inserted, such as by standard molecular cloning techniques. Another type of vector is a viral vector, wherein virally-derived DNA or RNA sequences are present in the vector for packaging into a virus (e.g. retroviruses, replication defective retroviruses, adenoviruses, replication defective adenoviruses, and adeno-associated viruses). Viral vectors also include polynucleotides carried by a virus for transfection into a host cell. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (e.g. bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operatively-linked. Such vectors are referred to herein as “expression vectors.” Common expression vectors of utility in recombinant DNA techniques are often in the form of plasmids. Other suitable vectors include, but are not limited to, cosmids and Yeast Artificial Chromosomes. Accordingly, suitable nucleic acid expression vectors include, but are not limited to, transposon vectors as described above, as well as plasmid vectors, retroviral vectors, lentiviral vectors, AAV vectors, phage vectors, etc). It is contemplated that any vector mayAttorney Docket No. CATA-40729.601 be used as long as it is replicable and viable in the host. In preferred embodiments, the vectors are mammalian expression vectors that comprise among other elements described herein an origin of replication, a suitable promoter and enhancer, and also any necessary ribosome binding sites, polyadenylation sites, splice donor and acceptor sites, transcriptional termination sequences, and 5' flanking non-transcribed sequences. Suitable plasmid vectors that may be adapted to incorporate the nucleic acid constructs of the present invention include specific plasmids systems for transposon vectors, FLP-FLT systems, Cre-lox systems, CRISPR-Cas9 systems, recombinase systems and integrase systems as well as plasmid vectors derived from pCIneo, pVAX1, pACT, Gateway plamids, pAdvantage, pBIND, pG5luc, pTNT, pTarget, pCat3, pSI, pCMV, pSV and the like. In some embodiments, the present invention provides host cells and host cell culture wherein the host cells express the protein of interest from the nucleic acid constructs described above. In preferred embodiment, the host cells a mammalian host cells. A number of mammalian host cell lines are known in the art. In general, these host cells are capable of growth and survival when placed in either monolayer culture or in suspension culture in a medium containing the appropriate nutrients and growth factors, as is described in more detail below. Typically, the cells are capable of expressing and secreting large quantities of a particular protein of interest into the culture medium. Examples of suitable mammalian host cells include, but are not limited to Chinese hamster ovary cells (CHO-K1, ATCC CCl-61); bovine mammary epithelial cells (ATCC CRL 10274; bovine mammary epithelial cells); monkey kidney CV1 line transformed by SV40 (COS-7, ATCC CRL 1651); human embryonic kidney line (293 or 293 cells subcloned for growth in suspension culture; see, e.g., Graham et al., J. Gen Virol., 36:59
[1977] ); baby hamster kidney cells (BHK, ATCC CCL 10); mouse sertoli cells (TM4, Mather, Biol. Reprod.23:243-251
[1980] ); monkey kidney cells (CV1 ATCC CCL 70); African green monkey kidney cells (VERO-76, ATCC CRL- 1587); human cervical carcinoma cells (HELA, ATCC CCL 2); canine kidney cells (MDCK, ATCC CCL 34); buffalo rat liver cells (BRL 3A, ATCC CRL 1442); human lung cells (W138, ATCC CCL 75); human liver cells (Hep G2, HB 8065); mouse mammary tumor (MMT 060562, ATCC CCL51); TRI cells (Mather et al., Annals N.Y. Acad. Sci., 383:44-68
[1982] ); MRC 5 cells; FS4 cells; rat fibroblasts (208F cells); MDBK cells (bovine kidney cells); CAP (CEVEC's Amniocyte Production) cells; and a human hepatoma line (Hep G2). In some particularly preferred embodiments, the host cells are modified so that they are deficient, or are naturally deficient, in an enzyme activity that is required for growth or survival of the cells in the presence of a selection agent and which is provided by theAttorney Docket No. CATA-40729.601 selectable marker. For example, Chinese Hamster Ovary (CHO) cells have been modified to be deficient for GS. In some preferred embodiments where vector includes a GS selectable marker, the host cell line is deficient in GS. In some particularly preferred embodiments, the GS deficient host cell line is the CHOZN® GS- / -cell line available from Merck KGaA. In other embodiments, where the selectable marker is, for example, DHFR, the cell line may preferably be deficient for DHFR activity (i.e., DHFR-). Suitable DHFR- cell lines include but are not limited to CHO-DG44 and derivatives thereof. The nucleic acid constructs and vectors of the present invention may be introduced into host cells by any suitable means such as by transfection, transformation or transduction. In some embodiments, after transfection or transduction, the cells are allowed to multiply, and are then trypsinized and re-plated. Individual colonies are then selected to provide clonally selected cell lines. In still further embodiments, the clonally selected cell lines are screened by Southern blotting or PCR assays to verify that the desired number of integration events has occurred. It is also contemplated that clonal selection allows the identification of superior protein producing cell lines. In other embodiments, the cells are not clonally selected following transfection. In some embodiments, the nucleic acid constructs encoding different proteins of interest are introduced into the host cells, for example by transfection or electroporation. The nucleic acid constructs encoding different proteins of interest can be introduced into the host cells at the same time or in a serial manner (e.g., a nucleic acid construct encoding a first protein of interest is introduced, a period of time is allowed to pass, and then a nucleic acid construct encoding a second protein of interest is introduced). In some embodiments of the present invention, following transformation of a suitable host strain and growth of the host strain to an appropriate cell density in media, the protein of interest is secreted during culture of the host cells. In some preferred embodiments where amplifiable markers are utilized, it is contemplated that culture of transduced host cells in a medium comprising an inhibitor of the gene. Suitable inhibitors include, but are not limited to methotrexate for inhibition of DHFR and methionine sulphoximine (Msx) or phosphinothricin for inhibition of GS. It is contemplated that as concentrations of these inhibitors are increased in a cell culture system, cells with higher copy numbers of the amplifiable marker (and thus the genes or genes of interest) or which contain higher- producing insertions are selected. Accordingly, the host cells containing vectors as described above are preferably cultured according to methods known in the art. Suitable culture conditions for mammalianAttorney Docket No. CATA-40729.601 cells are well known in the art (See e.g., J. Immunol. Methods (1983) 56:221-234
[1983] , Animal Cell Culture: A Practical Approach 2nd Ed., Rickwood, D. and Hames, B. D., eds. Oxford University Press, New York
[1992] ). The host cell cultures of the present invention are prepared in a media suitable for the particular cell being cultured. Commercially available media such as ActiPro media (HyClone), ExCell Advanced Fed Batch Medium (SAFC), Ham's F10 (Sigma, St. Louis, MO), Minimal Essential Medium (MEM, Sigma), RPMI-1640 (Sigma), and Dulbecco's Modified Eagle's Medium (DMEM, Sigma) are exemplary nutrient solutions. Suitable media are also described in U.S. Pat. Nos.4,767,704; 4,657,866; 4,927,762; 5,122,469; 4,560,655; and WO 90 / 03430 and WO 87 / 00195; the disclosures of which are herein incorporated by reference. Any of these media may be supplemented as necessary with serum, hormones and / or other growth factors (such as insulin, transferrin, or epidermal growth factor), salts (such as sodium chloride, calcium, magnesium, and phosphate), buffers (such as HEPES), nucleosides (such as adenosine and thymidine), antibiotics (such as gentamycin (gentamicin), trace elements (defined as inorganic compounds usually present at final concentrations in the micromolar range) lipids (such as linoleic or other fatty acids) and their suitable carriers, and glucose or an equivalent energy source. In some preferred embodiments where selectable markers such as GS are utilized, for example, the media will lack glutamine. Any other necessary supplements may also be included at appropriate concentrations that would be known to those skilled in the art. The present invention also contemplates the use of a variety of culture systems (e.g., petri dishes, 96 well plates, roller bottles, and bioreactors) for the transfected host cells. For example, the transfected host cells can be cultured in a perfusion system. Perfusion culture refers to providing a continuous flow of culture medium through a culture maintained at high cell density. The cells are suspended and do not require a solid support to grow on. Generally, fresh nutrients must be supplied continuously with concomitant removal of toxic metabolites and, ideally, selective removal of dead cells. Filtering, entrapment and micro- capsulation methods are all suitable for refreshing the culture environment at sufficient rates. As another example, in some embodiments a fed batch culture procedure can be employed. In the preferred fed batch culture the mammalian host, cells and culture medium are supplied to a culturing vessel initially and additional culture nutrients are fed, continuously or in discrete increments, to the culture during culturing, with or without periodic cell and / or product harvest before termination of culture. The fed batch culture can include, for example, a semi-continuous fed batch culture, wherein periodically whole cultureAttorney Docket No. CATA-40729.601 (including cells and medium) is removed and replaced by fresh medium. Fed batch culture is distinguished from simple batch culture in which all components for cell culturing (including the cells and all culture nutrients) are supplied to the culturing vessel at the start of the culturing process. Fed batch culture can be further distinguished from perfusion culturing insofar as the supernatant is not removed from the culturing vessel during the process (in perfusion culturing, the cells are restrained in the culture by, e.g., filtration, encapsulation, anchoring to microcarriers etc. and the culture medium is continuously or intermittently introduced and removed from the culturing vessel). In some particularly preferred embodiments, the batch cultures are performed in roller bottles. Further, the cells of the culture may be propagated according to any scheme or routine that may be suitable for the particular host cell and the particular production plan contemplated. Therefore, the present invention contemplates a single step or multiple step culture procedure. In a single step culture, the host cells are inoculated into a culture environment and the processes of the instant invention are employed during a single production phase of the cell culture. Alternatively, a multi-stage culture is envisioned. In the multi-stage culture cells may be cultivated in a number of steps or phases. For instance, cells may be grown in a first step or growth phase culture wherein cells, possibly removed from storage, are inoculated into a medium suitable for promoting growth and high viability. The cells may be maintained in the growth phase for a suitable period of time by the addition of fresh medium to the host cell culture. Fed batch or continuous cell culture conditions are devised to enhance growth of the mammalian cells in the growth phase of the cell culture. In the growth phase cells are grown under conditions and for a period of time that is maximized for growth. Culture conditions, such as temperature, pH, dissolved oxygen (dO2) and the like, are those used with the particular host and will be apparent to the ordinarily skilled artisan. Generally, the pH is adjusted to a level between about 6.5 and 7.5 using either an acid (e.g., CO2) or a base (e.g., Na2CO3or NaOH). A suitable temperature range for culturing mammalian cells such as CHO cells is between about 30oto 38oC and a suitable dO2 is between 5-90% of air saturation. Following the polypeptide production phase, the polypeptide of interest is recovered from the culture medium using techniques that are well established in the art. The protein of interest preferably is recovered from the culture medium as a secreted polypeptide (e.g., the secretion of the protein of interest is directed by a signal peptide sequence), although it also may be recovered from host cell lysates. As a first step, the culture medium or lysate isAttorney Docket No. CATA-40729.601 centrifuged to remove particulate cell debris. The polypeptide thereafter is purified from contaminant soluble proteins and polypeptides, with the following procedures being exemplary of suitable purification procedures: by fractionation on immunoaffinity or ion- exchange columns; ethanol precipitation; reverse phase HPLC; chromatography on silica or on a cation-exchange resin such as DEAE; chromatofocusing; SDS-PAGE; ammonium sulfate precipitation; gel filtration using, for example, Sephadex G-75; and protein A Sepharose columns to remove contaminants such as IgG. A protease inhibitor such as phenyl methyl sulfonyl fluoride (PMSF) also may be useful to inhibit proteolytic degradation during purification. Additionally, the protein of interest can be fused in frame to a marker sequence that allows for purification of the protein of interest. Non-limiting examples of marker sequences include a hexa-histidine tag, which may be supplied by a vector, preferably a pQE- 9 vector, and a hemagglutinin (HA) tag. The HA tag corresponds to an epitope derived from the influenza hemagglutinin protein (See e.g., Wilson et al., Cell, 37:767
[1984] ). One skilled in the art will appreciate that purification methods suitable for the polypeptide of interest may require modification to account for changes in the character of the polypeptide upon expression in recombinant cell culture. In some preferred embodiments, the nucleic acid constructs are incorporated into systems. In some embodiments, the systems comprise multiple nucleic acid constructs or vectors as described above which are intended for introduction into a host cell. In other preferred embodiments, the systems comprise one or more multiple nucleic acid constructs or vectors as described above which are intended for introduction into a host cell in addition to a nucleic acid or vector that encodes an enzyme that is necessary for incorporation of the nucleic acid constructs into a host cell genome. Exemplary enzymes include, but are not limited to, transposes for use with transposon vector systems, integrases for use in systems which utilize integration sequences such as the PhiC31 system, MMLV systems, and the like, recombinases for use in vector systems such as Cre-loc, FLP-FRT and the like, and Cas9 nucleases for use in CRISPR based systems. EXPERIMENTALAttorney Docket No. CATA-40729.601 Example 1: Plasmids and Transduction Methods used in the Examples Transductions were performed using either GPExTM(Bleck, “An Alternative Method for the Rapid Generation of Stable, High-Expressing Mammalian Cell Lines,” BioProcess J. 5(4):36–42 (2006)) to generate cell pools or GPExTMLightning. For GPExTM, briefly, replication-defective retroviral vectors, derived from Moloney murine leukemia virus (MLV) and pseudotyped with vesicular stomatitis virus G protein (VSV-G), are used to stably insert single copies of genes into dividing cells. Retrovectors deliver genes coded as RNA that, after entering the cell, are reverse transcribed to DNA and integrated stably into the genome of the host cell. Two enzymes, reverse transcriptase and integrase, provided transiently in the vector particle, perform this function. These integrated genes are maintained through subsequent cell divisions as if they were endogenous cellular genes. Regular GPExTMuses a derivative of the CHO-S cell line, whereas the GPExTMLightning process (see, e.g., WO2021247671, which is hereby incorporated by reference in its entirety) uses the CHOZn® GS knockout cell line as the starting base CHO line. GPExTMLightning also utilizes a novel gene insertion technology, allowing better control over the ratio of the HGF gene and the hepsin gene in the pooled cell lines. The GPExTMLightning technology uses a recombinase to specifically insert genes into “dock” sequences recognized by the recombinase and placed into the cell line using the GPEx process. Approximately 150-200 “dock” sequences are present in the cell line and available for gene insertion. Plasmid Platform Figure(s) Associated Notes Name Sequence(s)Attorney Docket No. CATA-40729.601 Plasmid Platform Figure(s) Associated Notes Name Sequence(s) 207 ttB-GS- GPEx FIG 7 SEQ ID NO 36 B vin l h -Transductions of CHO cells was performed in triplicate using GPExTM, as described in Example 1, using three cycles of transduction for both the native signal peptide construct (pFCS-hHGF-WPRE-SIN) as well as the bovine alpha-lactalbumin signal peptide construct (pFCS-SPhHGF-WPRE-SIN). The pooled cell lines for each were expanded and analyzed in fed-batch culture conditions. Each of the pooled cell lines were compared for amount of gene copies using a gene copy index value that is a ratio of the number of copies of the transgene to a single copy endogenous CHO gene using quantitative real-time PCR. The gene copy index, generated by subtracting the transgene Ct from a control Ct, reflects the number of transgene inserts that are present in the genome of the cell line. In general, gene index values are exponential – increasing by +1 reflects approximately 2-fold increase of the target relative to the control.. The higher the value the more copies of the gene, but a value of 1 does not mean a single gene copy is present. The pooled cell line gene copy index after each cycle of transduction was as follows: Transduction Endogenous SP GCIV Alpha-Lactalbumin SP GCIVThe results indicated very similar copy numbers for the cell pools after 3 transductions and they were compared for HGF production. Since the gene copy numbers were similar, fed-batch production was performed with each cell pool to examine the HGF production and processing, using the conditions and sampling methodology as follows:Attorney Docket No. CATA-40729.601 Fed Batch Culture Conditions (Base Medium: G12.1 + 6 mM L-glutamine + 4 g / L PS307) Day of Culture Supplement Day 0 4 g / L of Cell Boost 4 PS307; Seed at 3.0 x 105cells / mLSample Type Day of Culture ll t D il t ti t D 2As shown in FIGs.11A–11D, each of the cultures showed a very similar growth behavior typical of that of GPExTMpooled cell lines. PAGE analysis of expressed HGF was used to measure reduced and non-reduced forms of HGF in comparison to a control HGF standard (serum activated). As shown in FIG.12, the non-reduced SDS-PAGE analysis of the productivity samples on day 4, 8 and end of culture showed that both the pooled cell lines appear to be producing HGF at a similar size to the HGF standard. However, there is much greater production from the cell pool expressing HGF with the native signal peptide. As shown in FIG.13, the reduced SDS-PAGE analysis of the productivity samples on day 4,8 and end of culture showed that the majority of the native signal peptide sample is not fully cleaved to make fully processed HGF. The non-cleaved product is boxed.Attorney Docket No. CATA-40729.601 Example 3: Transduction of HGF producing CHO cells with hepsin The CHO cells transduced with the pFCS-hHGF-WPRE-SIN plasmid (Example 2) were further transduced with pFCS-newhHepsin-WPRE-SIN in triplicate using GPExTM, as described in Example 1. Cell pools were expanded and frozen after each transduction for further analysis. The pooled cell line gene copy index after each cycle of transduction was as follows: Transduction GCIV 1x -3.1 Surprisingly, ved in the cell pools. Typically, one would see much higher gene copy index values with transductions. One would expect to see numbers similar to what was observed for HGF after each cycle of transduction. The numbers got higher with each cycle, but are significantly lower than the normal values observed with HGF. This can occur when the protein product is toxic to cells or inhibits cell growth when expressed at high levels. The cell pools were expanded and run in fed-batch productivity studies to examine HGF production and determine if any improvements were observed in regards to the production of fully processed HGF. As shown in FIG.14A and FIG.14B, no significant improvement in the amount of processed HGF was observed in these pooled cell lines. Non- reduced gels and reduced gels comparing cell pool production from the pool without hepsin, the 1 cycle hepsin transduced pool and the 2 cycle hepsin transduced pool, showed no significant difference in the amount of processed HGF. The red box highlights the un- processed HGF molecule in each of the 3 cell pool samples. Example 3: Co-transductions of GS knockout CHO cells with HGF and hepsin Four different pooled cell lines were produced using the GPExTMLightning process, as described in Example 1, as follows: Cell Line Pool Expression Construct(s) Ratio of HGF:HepsinAttorney Docket No. CATA-40729.601 207attB-GS-hHGF-WPRE; 3 10:1 207attB-GS-Hepsin-WPRE and thetransgene constructs, selection by removal of glutamine from the media occurs to allow only cell lines containing the transgene (and the glutamine synthase) to survive. How fast cell lines recover from the selection can typically indicate how well the process worked as well as if there are any issues caused by transgene expression. Lines 1 and 4 responded “normally” and appeared to be healthy cultures. Lines 2 and 3 showed much slower recovery profiles and needed additional supplementation and attention in order to maintain growth. As shown in FIG.15, the pooled cell line produced using a 100:1 ratio of HGF:Hepsin recovered from selection much faster than the other two cell pools and the recovery was highly correlated to the amount of hepsin the cells were in theory producing, consistent with the data that was observed using the GPEx process. The pools with poor recovery continued to have some growth issues. HGF production was analyzed for Lines 1 and 4. As was previously observed with the GPExTMcell lines only containing HGF, and shown in FIG.16A and FIG.16B, both Lines 1 and 4 showed good levels of HGF expression (FIG.16A). Line 1 showed good HGF production, but very little processing / cleavage, whereas Line 4 showed good HGF production and all the HGF appears to be processed / cleaved and looks very similar to the HGF standard produced using serum on SDS-PAGE gels (FIG.16B). The expression of HGF was good in expansion media obtained from each of the two cell pools examined on the non-reduced gel. As observed with the original GPExTMexperiments that were performed in the CHO-S base cell line and now the GPExTMLightning data from the CHOZN GS knock-out cell line, very little processed HGF was observed in cell lines lacking hepsin. However, unlike the previous GPExTMexperiments that used hepsin, the 100:1 HGF:Hepsin cell pool now showed fully process HGF on the reduced SDS-PAGE gel. The 100:1 cell pool was showing normal cell growth and behavior and was fully processing the HGF. A fed-batch productivity study was carried out in 500 ml shake flasks for the two pooled cell lines (Line 1 and Line 4) with the following culture and sampling conditions:Attorney Docket No. CATA-40729.601 Fed Batch Culture Conditions (Base Medium: ExCell Advance Fed-Batch Media) Day of Culture Supplement Day 0 Seed at 3.0 x 105cells / mLSample Type Day of CultureAs shown in FIG.17A and FIG.17B, both cell pools showed very similar cell growth characteristics through the culture suggesting no issues with hepsin effecting the growth profile in the culture. HGF production was measured for the fed-batch cultures. As shown in FIGS.18A– 18D, the results were similar to those observed in the expansion media analysis, with little fully processed HGF in the HGF only pool and fully processed HGF in the 100:1 pool, except with respect to the presence of apparent extra cleavage in the 100:1 pool. However, the additional cleavage doesn’t seem to affect the behavior of the molecule in the non-reduced gel. Additional cleavage of HGF during the production / processing is selected against during clonal cell line selection process and the development of the cell culture and protein purification conditions.Attorney Docket No. CATA-40729.601 The material from the cell pools was purified using ion exchange chromatography. As shown in FIG.19, purified products from each of the two cultures showed similar results to the non-purified culture media samples. HGF expression was high for both (left panel, non-reduced) and the HGF expressed without hepsin showed a low percentage of processed HGF, while the 100:1 pooled material showed full processing and the additional cleavage (right panel, reduced). Example 4: Peptide Mapping Peptide mapping was carried out on the HGF protein produced by Line 4. Samples were reduced and alkylated and then digested with each of Trypsin, Asp-N, and Glu-C. Searches were predicted digested species were carried out for each enzyme for both the un- cleaved HGF sequence and the hepsin-activated HGF sequence (i.e. cleaved into alpha and beta chains). The overall 3 enzyme sequence coverage was 99.7% (690 / 692 amino acids— 99% coverage for Trypsin, 91% for Asp-N, and 85% for Glu-C). The trypsin digestion was not suitable for confirmation of hepsin cleavage. The lack of an Asp-N digest peptide corresponding to AA 455–503 of SEQ ID NO:32, suggests correct hepsin cleavage at the hepsin cleavage site between R458 and V459 of SEQ ID NO:32. The lack of a Gly-C digest peptide corresponding to AA 435–490 of SEQ ID NO:32 also suggests correct hepsin cleavage at the hepsin cleavage site between R458 and V459 of SEQ ID NO:32. The HGF sequence of SEQ ID NO:32 contains four RV dipeptides—at the cleavage site between the alpha and beta chains (458–459), one within the alpha chain (337–338, “c1”), and two within the beta chain (594–595, “c2”; and 672–673, “c3”) and. Potential undesirable hepsin cleavage was analyzed by calculating the percent Asp-N and Glu-C cleaved at each of the four dinucleotide sites as follows: ^^^^^^ ^^^^^^^^^^^^^^^^^^^^ ^^^^^^ ^^^^^^^^^^^^^^^^ ^^^^^^^^^^^^^^^^ ^^^^^^ ^^^^ ^^^^^^ ^^^^^^ ^^^^^^^^^^^^^^^^^^^^^^ ^^^^^^ ^^^^^^^^^^^^^^^^^^ ^^^^^^^^^^^^^^^^ ^^^^^^^^^^^^^^^^∗ 100Site Asp-N Glu-CAttorney Docket No. CATA-40729.601 EXAMPLE 5. Clonal Selection An initial round of clonal selection was carried out on the 100:1 HGF:Hepsin pool. For 76 of the clones, the number of gene copies of HGF was estimated using two different measures (Gene Copy Index and HGF Gene Copy Index in the table below). The estimated HGF copy number is 2^(the average of these two measures). The presence of a hepsin gene (positive or negative) was also measured: Serial Clone Gene Copy HGF Gene Mean Estimated HGF Hepsin Band Number Number Index Copy Index Copy Number Present (+ / -) 1 1457 4.47 4.68 4.6 24 - - - + - - - + - - + - - - - - - - + - - + - - - + - - - - - - - - - -Attorney Docket No. CATA-40729.601 37 1076 6.43 6.67 6.6 94 - 38 1064 6.48 7.02 6.8 108 - 39 633 627 680 65 93 - - - - - - + - - - - - - - - - + - - - - - - - - - - - - - - - - - - - - +12 clones with 0, 1, or 2 copies of the Hepsin gene were selected for further analysis. The Estimated HGF Copy Number, Hepsin Gene Copy Number, and HGF:Hepsin Copy Number ratio for those clones are as follows:Attorney Docket No. CATA-40729.601 Clone # HGF Mean Estimated Hepsin Hepsin HGF:Hepsin Gene Gene HGF Copy Band Gene Ratio (X:1) Cp p g y p ed the HGF product in a miniature bioreactor fed-batch culture, resulting in unwanted cleavage products. All the clones with no exogenous Hepsin showed significant under-processing of the HGF. Clone #526 tended to be the clone that processed the HGF consistently to the most desirable form. For each of the three levels of hepsin copies (0, 1, and 2), the minimum, maximum, and mean Estimated HGF Copy Number of the corresponding clones as well as the minimum, maximum, and mean HGF:Hepsin Ratio are shown in the table below. Copies of Hepsin HGF min HGF max HGF mean Ratio min Ratio max Ratio mean 96 49Attorney Docket No. CATA-40729.601 SEQUENCES SEQ ID NO:1. Amino acid sequence of Homo sapien HGF Isoform 3; (Identifier: P14210-3, Accession number NP_001010932.1) (dHGF) MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTK KVNTADQCANRCTRNKGLPFTCKAFVFDKARKQCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRNCIIGKGRSYKGTVSITKSGIKCQPWSSMIPHEHSYRGKDLQENYCRNPRGEEGGPWCFT SNPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPERYPDK GFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQGEGYRG TVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNIRVGYC SQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASKLNENY CRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVVNGIPT RTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGDEKCKQ VLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYTGLINY DGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQHKMRMV LGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO:2. Amino acid sequence of Homo sapien HGF Isoform 1; (Identifier: P14210-1, Accession number: NP_000592.3) MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTK KVNTADQCANRCTRNKGLPFTCKAFVFDKARKQCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRNCIIGKGRSYKGTVSITKSGIKCQPWSSMIPHEHSFLPSSYRGKDLQENYCRNPRGEEGG PWCFTSNPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO:3. Amino acid sequence of Homo sapien HGF Isoform 2 (Identifier: P14210-2, Accession Number: NP_001010931.1): MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTK KVNTADQCANRCTRNKGLPFTCKAFVFDKARKQCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRNCIIGKGRSYKGTVSITKSGIKCQPWSSMIPHEHSFLPSSYRGKDLQENYCRNPRGEEGG PWCFTSNPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCET SEQ ID NO:4. Amino acid sequence of Homo sapien HGF Isoform 4 (Identifier: P14210-4): MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTK KVNTADQCANRCTRNKGLPFTCKAFVFDKARKQCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRNCIIGKGRSYKGTVSITKSGIKCQPWSSMIPHEHSFLPSSYRGKDLQENYCRNPRGEEGG PWCFTSNPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKNMRDITWALNAttorney Docket No. CATA-40729.601 SEQ ID NO:5. Amino acid sequence of Homo sapien HGF Isoform 5 (Identifier: P14210-5, NP_001010933.1): MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTK KVNTADQCANRCTRNKGLPFTCKAFVFDKARKQCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRNCIIGKGRSYKGTVSITKSGIKCQPWSSMIPHEHSYRGKDLQENYCRNPRGEEGGPWCFT SNPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPERYPDK GFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCET SEQ ID NO:6. Amino acid sequence of Homo sapien HGF Isoform 6 (Identifier: P14210-6, NP_001010934.1): MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTK KVNTADQCANRCTRNKGLPFTCKAFVFDKARKQCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRNCIIGKGRSYKGTVSITKSGIKCQPWSSMIPHEHSFLPSSYRGKDLQENYCRNPRGEEGG PWCFTSNPEVRYEVCDIPQCSEGK SEQ ID NO: 7. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALEIKTK KANTADQCANRCTRSKGLPFTCKAFVFDKARKRCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRDCIIGKGRSYRGTVSITKSGIKCQPWSAMIPHEHSFLPSSYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 8. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTE KANTADQCANRCIRNKGLPFTCKAFVFDKARKRCLWFPVNSMSSGVKKEFGHEFDLYENKDY TRNCIVGNGRSYRGTVSTTKSGIKCQPWSAMIPHEHSFLPSSYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 9. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTE KANTADQCANRCTRSKGLPFTCKAFVFDKARKRCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRDCIVGNGRSYRGTVSITKSGIECQPWSAMIPHEHSFLPSSYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASKAttorney Docket No. CATA-40729.601 LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 10. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAKGQGKRRNTIHEFKKSAKTTLIKIDPALKIKTE KADTADQCANRCTRSKGLPFTCKAFVFDKARKQCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRDCIVGNGRSYRGTVSVTKSGIKCQPWSSMIPHEHSFLPSSYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 11. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNAIHEFKKSAKATLIKIDPALKIKTE KANTADQCANRCTRSKGLPFTCKAFVFDKARKRCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRDCIVGNGRSYRGTVSITKSGIECQPWSSMIPHEHSFLPSSYRGEDLQENYCRNPWGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 12. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTE KANTADQCANRCTRSKGLPFTCKAFVFDKARKRCLWFPFNSMSSGVKKEFGHEFDLYENKDY TRNCIVGNGRSYRGTVSITKSGIECQPWSAMIPHEHSFLPSSYQGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 13. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTK KVNTADQCANRCTRNKGLPFTCKAFVFDKARKQCLWFPFNSMSSGVKKEFGHEFDLYENKDYAttorney Docket No. CATA-40729.601 IRNCIIGKGRSYKGTVSITKSGIKCQPWSSMIPHEHSFLPSSYRGKDLQENYCRNPRGEEGG PWCFTSNPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 14. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRDAIHECKRSAKTTLIKIDPALKIKTE KANTADQCANRCTRNKGLPSTCKAFVFDKARKRRLRFPFNSMSSGVKKEFGHEFDLYENKDY TRNCIVGKGRSYRGTVSTTKSGIKCQPWSAMIPHEHSFLPSSYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 15. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQGKRRNTIHEFKKSAKTTLIKIDPALKIKTE KVNTADQCANRCTRSKGLPFTCKAFVFDKARKRCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRDCIIGRGRSYRGTVSITKSGIKCQPWSAMIPHEHSFLPSSYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 16. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPHAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTE KANTADQCANRCTRSKGLPFTCKAFVFDKARKRCLWFPFNSMSSGVKKEFGHEFDLYENKDY TRNCIVGNGRSYRGTVSTTKSGIKCQPWSAMIPHEHSFLPSSYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYTAttorney Docket No. CATA-40729.601 GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 17. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTE KANTADQCANRCTRSRGLPFTCKAFVFDKARKQCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRDCIIGNGRSYRGTVSVTKSGIKCQPWSAMIPHEHSFLPSSYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 18. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTE KANTADQCANRCTRNKGLPFTCKAFVFDKARKRCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRDCIVGNGRSYRGTVSITKSGIKCQPWSSMIPHEHSFLPSSYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 19. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSVKTTLIKIDPALKIKTE KANTADQCANRCTRSKGLPFTCKAFVFDKARKRCLWFPVNSMSSGVKKESGHEFDLYENKDY IRDCIVGNGRSYRGTVSTTKSGIKCQPWSAMIPHEHSFLPSSYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 20. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTE KANTADQCANRCTRSKGLPFTCKAFVFDKARKRCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRDCIVGNGRSYRGTVSITKSGIKCQPWSSMIPHEHSFLPSSYRGEDLRENYCRNPWGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQGAttorney Docket No. CATA-40729.601 EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 21. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALRIKTE KANTADQCANRCTRSRGLPFTCKAFVFDKARKRCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRDCIIGNGRSYRGTVSITKSGIKCQPWSSMIPHEHSFLPSSYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 22. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTE KANTADQCANRCTRSRRLPFTCKAFVFDKARKRCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRDCIIGKGRSYRGTVSVTKSGIECQPWSAMIPHEHSFLPSNYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 23. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTE KANTADQCANRCTRSKGLPFTCKAFVFDKARKRCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRDCIVGNGRSYRGTVSITKSGIECQPWSAMIPHEHSFLPSSYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQSAttorney Docket No. CATA-40729.601 SEQ ID NO: 24. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQRKRRNTIHEFKKSAKTTLIKIDPALKIKTK KVDTADQCANRCTRNKGLPFTCKAFVFDKARKRCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRDCIIGNGRSYRGTVSITKSGIKCQPWSSMIPHEHSFLPSSYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 25. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQGKRRNTIHEFKKSAKTTLIKIDPALRIKTE KANTADQCANRCTRSKGLPFTCKAFVFDKARKRCLWFPFNSMSSGVKKEFGHEFDLYENKAY IRDCIIGRGRNYRGTVSITKSGIKCQPWSAMIPHEHSFLPSSYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 26. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAKGQRKRRNTIHEFKKSAKTTLIKIDPALEIKTE KVNTADQCANRCIRNKGLPFTCKAFVFDKARKRCLWFPFNSMSSGVKKEFGHEFDLYENKAY IRDCIIGRGRNYRGTVSITKSGIKCQPWSAMIPHEHSFLPSSYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVV NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO: 27. HGF Variant: MWVTKLLPALLLQHVLLHLLLLPIAIPYAEGQGKRRNTIHEFKKSAKTTLIKIDPALKIKTE KVNTADQCANRCTRSKGLPFTCKAFVFDKARKRCLWFPFNSMSSGVKKEFGHEFDLYENKDY IRDCIIGNGRSYRGTVSITKSGIKCQPWSAMIPHEHSFLPSSYRGEDLRENYCRNPRGEEGG PWCYTSDPEVRYEVCDIPQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPE RYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQG EGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNI RVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASK LNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVVAttorney Docket No. CATA-40729.601 NGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGD EKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYT GLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQH KMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO:28 (>sp|P05981|HEPS_HUMAN Serine protease hepsin OS=Homo sapiens OX=9606 GN=HPN PE=1 SV=1) MAQKEGGRTVPCCSRPKVAALTAGTLLLLTAIGAASWAIVAVLLRSDQEPLYPVQVSSADAR LMVFDKTEGTWRLLCSSRSNARVAGLSCEEMGFLRALTHSELDVRTAGANGTSGFFCVDEGR LPHTQRLLEVISVCDCPRGRFLAAICQDCGRRKLPVDRIVGGRDTSLGRWPWQVSLRYDGAH LCGGSLLSGDWVLTAAHCFPERNRVLSRWRVFAGAVAQASPHGLQLGVQAVVYHGGYLPFRD PNSEENSNDIALVHLSSPLPLTEYIQPVCLPAAGQALVDGKICTVTGWGNTQYYGQQAGVLQ EARVPIISNDVCNGADFYGNQIKPKMFCAGYPEGGIDACQGDSGGPFVCEDSISRTPRWRLC GIVSWGTGCALAQKPGVYTKVSDFREWIFQAIKTHSEASGMVTQL SEQ ID NO:29 Extracellular domain of P05981 RSDQEPLYPVQVSSADARLMVFDKTEGTWRLLCSSRSNARVAGLSCEEMGFLRALTHSELDV RTAGANGTSGFFCVDEGRLPHTQRLLEVISVCDCPRGRFLAAICQDCGRRKLPVDRIVGGRD TSLGRWPWQVSLRYDGAHLCGGSLLSGDWVLTAAHCFPERNRVLSRWRVFAGAVAQASPHGL QLGVQAVVYHGGYLPFRDPNSEENSNDIALVHLSSPLPLTEYIQPVCLPAAGQALVDGKICT VTGWGNTQYYGQQAGVLQEARVPIISNDVCNGADFYGNQIKPKMFCAGYPEGGIDACQGDSG GPFVCEDSISRTPRWRLCGIVSWGTGCALAQKPGVYTKVSDFREWIFQAIKTHSEASGMVTQ L SEQ ID NO:30 Serine protease hepsin catalytic chain of P05981 IVGGRDTSLGRWPWQVSLRYDGAHLCGGSLLSGDWVLTAAHCFPERNRVLSRWRVFAGAVAQ ASPHGLQLGVQAVVYHGGYLPFRDPNSEENSNDIALVHLSSPLPLTEYIQPVCLPAAGQALV DGKICTVTGWGNTQYYGQQAGVLQEARVPIISNDVCNGADFYGNQIKPKMFCAGYPEGGIDA CQGDSGGPFVCEDSISRTPRWRLCGIVSWGTGCALAQKPGVYTKVSDFREWIFQAIKTHSEA SGMVTQL SEQ ID NO:31 Signal Peptide of SEQ ID NO:1 MWVTKLLPALLLQHVLLHLLLLPIAIPYAEG SEQ ID NO:32 Un-cleaved Alpha and Beta chain peptide of SEQ ID NO:1 QRKRRNTIHEFKKSAKTTLIKIDPALKIKTKKVNTADQCANRCTRNKGLPFTCKAFVFDKAR KQCLWFPFNSMSSGVKKEFGHEFDLYENKDYIRNCIIGKGRSYKGTVSITKSGIKCQPWSSM IPHEHSYRGKDLQENYCRNPRGEEGGPWCFTSNPEVRYEVCDIPQCSEVECMTCNGESYRGL MDHTESGKICQRWDHQTPHRHKFLPERYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCA IKTCADNTMNDTDVPLETTECIQGQGEGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKC KDLRENYCRNPDGSESPWCFTTDPNIRVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRS GLTCSMWDKNMEDLHRHIFWEPDASKLNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCE GDTTPTIVNLDHPVISCAKTKQLRVVNGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTA RQCFPSRDLKDYEAWLGIHDVHGRGDEKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFV STIDLPNYGCTIPEKTSCSVYGWGYTGLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESE ICAGAEKIGSGPCEGDYGGPLVCEQHKMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHK IILTYKVPQSAttorney Docket No. CATA-40729.601 SEQ ID NO:33 Alpha chain peptide of SEQ ID NO:1 QRKRRNTIHEFKKSAKTTLIKIDPALKIKTKKVNTADQCANRCTRNKGLPFTCKAFVFDKAR KQCLWFPFNSMSSGVKKEFGHEFDLYENKDYIRNCIIGKGRSYKGTVSITKSGIKCQPWSSM IPHEHSYRGKDLQENYCRNPRGEEGGPWCFTSNPEVRYEVCDIPQCSEVECMTCNGESYRGL MDHTESGKICQRWDHQTPHRHKFLPERYPDKGFDDNYCRNPDGQPRPWCYTLDPHTRWEYCA IKTCADNTMNDTDVPLETTECIQGQGEGYRGTVNTIWNGIPCQRWDSQYPHEHDMTPENFKC KDLRENYCRNPDGSESPWCFTTDPNIRVGYCSQIPNCDMSHGQDCYRGNGKNYMGNLSQTRS GLTCSMWDKNMEDLHRHIFWEPDASKLNENYCRNPDDDAHGPWCYTGNPLIPWDYCPISRCE GDTTPTIVNLDHPVISCAKTKQLR SEQ ID NO:34 Beta chain peptide of SEQ ID NO:1 VVNGIPTRTNIGWMVSLRYRNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGR GDEKCKQVLNVSQLVYGPEGSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWG YTGLINYDGLLRVAHLYIMGNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCE QHKMRMVLGVIVPGRGCAIPNRPGIFVRVAYYAKWIHKIILTYKVPQS SEQ ID NO:35 nucleic acid sequence of HGF CDS in vector pFCS- hHGF-WPRE-SIN (new ori) 1008-146, CDI10.0002 ATGTGGGTGACCAAACTCCTGCCAGCCCTGCTGCTGCAGCATGTCCTCCTGCATCTCCTCCT GCTCCCCATCGCCATCCCCTATGCAGAGGGACAAAGGAAAAGAAGAAATACAATTCATGAAT TCAAAAAATCAGCAAAGACTACCCTAATCAAAATAGATCCAGCACTGAAGATAAAAACCAAA AAAGTGAATACTGCAGACCAATGTGCTAATAGATGTACTAGGAATAAAGGACTTCCATTCAC TTGCAAGGCTTTTGTTTTTGATAAAGCAAGAAAACAATGCCTCTGGTTCCCCTTCAATAGCA TGTCAAGTGGAGTGAAAAAAGAATTTGGCCATGAATTTGACCTCTATGAAAACAAAGACTAC ATTAGAAACTGCATCATTGGTAAAGGACGCAGCTACAAGGGAACAGTATCTATCACTAAGAG TGGCATCAAATGTCAGCCCTGGAGTTCCATGATACCACACGAACACAGCTATCGGGGTAAAG ACCTACAGGAAAACTACTGTCGAAATCCTCGAGGGGAAGAAGGGGGACCCTGGTGTTTCACA AGCAATCCAGAGGTACGCTACGAAGTCTGTGACATTCCTCAGTGTTCAGAAGTTGAATGCAT GACCTGCAATGGGGAGAGTTATCGAGGTCTCATGGATCATACAGAATCAGGCAAGATTTGTC AGCGCTGGGATCATCAGACACCACACCGGCACAAATTCTTGCCTGAAAGATATCCCGACAAG GGCTTTGATGATAATTATTGCCGCAATCCCGATGGCCAGCCGAGGCCATGGTGCTATACTCT TGACCCTCACACCCGCTGGGAGTACTGTGCAATTAAAACATGCGCTGACAATACTATGAATG ACACTGATGTTCCTTTGGAAACAACTGAATGCATCCAAGGTCAAGGAGAAGGCTACAGGGGC ACTGTCAATACCATTTGGAATGGAATTCCATGTCAGCGTTGGGATTCTCAGTATCCTCACGA GCATGACATGACTCCTGAAAATTTCAAGTGCAAGGACCTACGAGAAAATTACTGCCGAAATC CAGATGGGTCTGAATCACCCTGGTGTTTTACCACTGATCCAAACATCCGAGTTGGCTACTGC TCCCAAATTCCAAACTGTGATATGTCACATGGACAAGATTGTTATCGTGGGAATGGCAAAAA TTATATGGGCAACTTATCCCAAACAAGATCTGGACTAACATGTTCAATGTGGGACAAGAACA TGGAAGACTTACATCGTCATATCTTCTGGGAACCAGATGCAAGTAAGCTGAATGAGAATTAC TGCCGAAATCCAGATGATGATGCTCATGGACCCTGGTGCTACACGGGAAATCCACTCATTCC TTGGGATTATTGCCCTATTTCTCGTTGTGAAGGTGATACCACACCTACAATAGTCAATTTAG ACCATCCCGTAATATCTTGTGCCAAAACGAAACAATTGCGAGTTGTAAATGGGATTCCAACA CGAACAAACATAGGATGGATGGTTAGTTTGAGATACAGAAATAAACATATCTGCGGAGGATC ATTGATAAAGGAGAGTTGGGTTCTTACTGCACGACAGTGTTTCCCTTCTCGAGACTTGAAAG ATTATGAAGCTTGGCTTGGAATTCATGATGTCCACGGAAGAGGAGATGAGAAATGCAAACAG GTTCTCAATGTTTCCCAGCTGGTATATGGCCCTGAAGGATCAGATCTGGTTTTAATGAAGCT TGCCAGGCCTGCTGTCCTGGATGATTTTGTTAGTACGATTGATTTACCTAATTATGGATGCA CAATTCCTGAAAAGACCAGTTGCAGTGTTTATGGCTGGGGCTACACTGGATTGATCAACTAT GATGGCCTATTACGAGTGGCACATCTCTATATAATGGGAAATGAGAAATGCAGCCAGCATCA TCGAGGGAAGGTGACTCTGAATGAGTCTGAAATATGTGCTGGGGCTGAAAAGATTGGATCAG GACCATGTGAGGGGGATTATGGTGGCCCACTTGTTTGTGAGCAACATAAAATGAGAATGGTTAttorney Docket No. CATA-40729.601 CTTGGTGTCATTGTTCCTGGTCGTGGATGTGCCATTCCAAATCGTCCTGGTATTTTTGTCCG AGTAGCATATTATGCAAAATGGATACACAAAATTATTTTAACATATAAGGTACCACAGTCAT AG SEQ ID NO:36 Hepsin CDS in vector pFCS-newhHepsin-WPRE-SIN (new ori) 1008-146, CI01.0004 ATGATGTCCTTTGTCTCTCTGCTCCTGGTTGGCATCCTATTCCATGCCACCCAGGCCAGGAG TGACCAGGAGCCGCTGTACCCAGTGCAGGTCAGCTCTGCGGACGCTCGGCTCATGGTCTTTG ACAAGACGGAAGGGACGTGGCGGCTGCTGTGCTCCTCGCGCTCCAACGCCAGGGTAGCCGGA CTCAGCTGCGAGGAGATGGGCTTCCTCAGGGCACTGACCCACTCCGAGCTGGACGTGCGAAC GGCGGGCGCCAATGGCACGTCGGGCTTCTTCTGTGTGGACGAGGGGAGGCTGCCCCACACCC AGAGGCTGCTGGAGGTCATCTCCGTGTGTGATTGCCCCAGAGGCCGTTTCTTGGCCGCCATC TGCCAAGACTGTGGCCGCAGGAAGCTGCCCGTGGACCGCATCGTGGGAGGCCGGGACACCAG CTTGGGCCGGTGGCCGTGGCAAGTCAGCCTTCGCTATGATGGAGCACACCTCTGTGGGGGAT CCCTGCTCTCCGGGGACTGGGTGCTGACAGCCGCCCACTGCTTCCCGGAGCGGAACCGGGTC CTGTCCCGATGGCGAGTGTTTGCCGGTGCCGTGGCCCAGGCCTCTCCCCACGGTCTGCAGCT GGGGGTGCAGGCTGTGGTCTACCACGGGGGCTATCTTCCCTTTCGGGACCCCAACAGCGAGG AGAACAGCAACGATATTGCCCTGGTCCACCTCTCCAGTCCCCTGCCCCTCACAGAATACATC CAGCCTGTGTGCCTCCCAGCTGCCGGCCAGGCCCTGGTGGATGGCAAGATCTGTACCGTGAC GGGCTGGGGCAACACGCAGTACTATGGCCAACAGGCCGGGGTACTCCAGGAGGCTCGAGTCC CCATAATCAGCAATGATGTCTGCAATGGCGCTGACTTCTATGGAAACCAGATCAAGCCCAAG ATGTTCTGTGCTGGCTACCCCGAGGGTGGCATTGATGCCTGCCAGGGCGACAGCGGTGGTCC CTTTGTGTGTGAGGACAGCATCTCTCGGACGCCACGTTGGCGGCTGTGTGGCATTGTGAGTT GGGGCACTGGCTGTGCCCTGGCCCAGAAGCCAGGCGTCTACACCAAAGTCAGTGACTTCCGG GAGTGGATCTTCCAGGCCATAAAGACTCACTCCGAAGCCAGCGGCATGGTGACCCAGCTCTG A SEQ ID NO:37 Hepsin Protein Sequence in vector pFCS- newhHepsin-WPRE-SIN (new ori) 1008-146, CI01.0004 MMSFVSLLLVGILFHATQARSDQEPLYPVQVSSADARLMVFDKTEGTWRLLCSSRSNARVAG LSCEEMGFLRALTHSELDVRTAGANGTSGFFCVDEGRLPHTQRLLEVISVCDCPRGRFLAAI CQDCGRRKLPVDRIVGGRDTSLGRWPWQVSLRYDGAHLCGGSLLSGDWVLTAAHCFPERNRV LSRWRVFAGAVAQASPHGLQLGVQAVVYHGGYLPFRDPNSEENSNDIALVHLSSPLPLTEYI QPVCLPAAGQALVDGKICTVTGWGNTQYYGQQAGVLQEARVPIISNDVCNGADFYGNQIKPK MFCAGYPEGGIDACQGDSGGPFVCEDSISRTPRWRLCGIVSWGTGCALAQKPGVYTKVSDFR EWIFQAIKTHSEASGMVTQL SEQ ID NO:38 HGF CDS in vector pFCS-SPhHGF-WPRE-SIN (new ori) 1008-146, CI02.0002 ATGATGTCCTTTGTCTCTCTGCTCCTGGTTGGCATCCTATTCCATGCCACCCAGGCCCAAAG GAAAAGAAGAAATACAATTCATGAATTCAAAAAATCAGCAAAGACTACCCTAATCAAAATAG ATCCAGCACTGAAGATAAAAACCAAAAAAGTGAATACTGCAGACCAATGTGCTAATAGATGT ACTAGGAATAAAGGACTTCCATTCACTTGCAAGGCTTTTGTTTTTGATAAAGCAAGAAAACA ATGCCTCTGGTTCCCCTTCAATAGCATGTCAAGTGGAGTGAAAAAAGAATTTGGCCATGAAT TTGACCTCTATGAAAACAAAGACTACATTAGAAACTGCATCATTGGTAAAGGACGCAGCTAC AAGGGAACAGTATCTATCACTAAGAGTGGCATCAAATGTCAGCCCTGGAGTTCCATGATACC ACACGAACACAGCTATCGGGGTAAAGACCTACAGGAAAACTACTGTCGAAATCCTCGAGGGG AAGAAGGGGGACCCTGGTGTTTCACAAGCAATCCAGAGGTACGCTACGAAGTCTGTGACATT CCTCAGTGTTCAGAAGTTGAATGCATGACCTGCAATGGGGAGAGTTATCGAGGTCTCATGGA TCATACAGAATCAGGCAAGATTTGTCAGCGCTGGGATCATCAGACACCACACCGGCACAAAT TCTTGCCTGAAAGATATCCCGACAAGGGCTTTGATGATAATTATTGCCGCAATCCCGATGGCAttorney Docket No. CATA-40729.601 CAGCCGAGGCCATGGTGCTATACTCTTGACCCTCACACCCGCTGGGAGTACTGTGCAATTAA AACATGCGCTGACAATACTATGAATGACACTGATGTTCCTTTGGAAACAACTGAATGCATCC AAGGTCAAGGAGAAGGCTACAGGGGCACTGTCAATACCATTTGGAATGGAATTCCATGTCAG CGTTGGGATTCTCAGTATCCTCACGAGCATGACATGACTCCTGAAAATTTCAAGTGCAAGGA CCTACGAGAAAATTACTGCCGAAATCCAGATGGGTCTGAATCACCCTGGTGTTTTACCACTG ATCCAAACATCCGAGTTGGCTACTGCTCCCAAATTCCAAACTGTGATATGTCACATGGACAA GATTGTTATCGTGGGAATGGCAAAAATTATATGGGCAACTTATCCCAAACAAGATCTGGACT AACATGTTCAATGTGGGACAAGAACATGGAAGACTTACATCGTCATATCTTCTGGGAACCAG ATGCAAGTAAGCTGAATGAGAATTACTGCCGAAATCCAGATGATGATGCTCATGGACCCTGG TGCTACACGGGAAATCCACTCATTCCTTGGGATTATTGCCCTATTTCTCGTTGTGAAGGTGA TACCACACCTACAATAGTCAATTTAGACCATCCCGTAATATCTTGTGCCAAAACGAAACAAT TGCGAGTTGTAAATGGGATTCCAACACGAACAAACATAGGATGGATGGTTAGTTTGAGATAC AGAAATAAACATATCTGCGGAGGATCATTGATAAAGGAGAGTTGGGTTCTTACTGCACGACA GTGTTTCCCTTCTCGAGACTTGAAAGATTATGAAGCTTGGCTTGGAATTCATGATGTCCACG GAAGAGGAGATGAGAAATGCAAACAGGTTCTCAATGTTTCCCAGCTGGTATATGGCCCTGAA GGATCAGATCTGGTTTTAATGAAGCTTGCCAGGCCTGCTGTCCTGGATGATTTTGTTAGTAC GATTGATTTACCTAATTATGGATGCACAATTCCTGAAAAGACCAGTTGCAGTGTTTATGGCT GGGGCTACACTGGATTGATCAACTATGATGGCCTATTACGAGTGGCACATCTCTATATAATG GGAAATGAGAAATGCAGCCAGCATCATCGAGGGAAGGTGACTCTGAATGAGTCTGAAATATG TGCTGGGGCTGAAAAGATTGGATCAGGACCATGTGAGGGGGATTATGGTGGCCCACTTGTTT GTGAGCAACATAAAATGAGAATGGTTCTTGGTGTCATTGTTCCTGGTCGTGGATGTGCCATT CCAAATCGTCCTGGTATTTTTGTCCGAGTAGCATATTATGCAAAATGGATACACAAAATTAT TTTAACATATAAGGTACCACAGTCATAG SEQ ID NO: 39 HGF Protein Sequences in vector pFCS-SPhHGF- WPRE-SIN (new ori) 1008-146, CI02.0002 MMSFVSLLLVGILFHATQAQRKRRNTIHEFKKSAKTTLIKIDPALKIKTKKVNTADQCANRC TRNKGLPFTCKAFVFDKARKQCLWFPFNSMSSGVKKEFGHEFDLYENKDYIRNCIIGKGRSY KGTVSITKSGIKCQPWSSMIPHEHSYRGKDLQENYCRNPRGEEGGPWCFTSNPEVRYEVCDI PQCSEVECMTCNGESYRGLMDHTESGKICQRWDHQTPHRHKFLPERYPDKGFDDNYCRNPDG QPRPWCYTLDPHTRWEYCAIKTCADNTMNDTDVPLETTECIQGQGEGYRGTVNTIWNGIPCQ RWDSQYPHEHDMTPENFKCKDLRENYCRNPDGSESPWCFTTDPNIRVGYCSQIPNCDMSHGQ DCYRGNGKNYMGNLSQTRSGLTCSMWDKNMEDLHRHIFWEPDASKLNENYCRNPDDDAHGPW CYTGNPLIPWDYCPISRCEGDTTPTIVNLDHPVISCAKTKQLRVVNGIPTRTNIGWMVSLRY RNKHICGGSLIKESWVLTARQCFPSRDLKDYEAWLGIHDVHGRGDEKCKQVLNVSQLVYGPE GSDLVLMKLARPAVLDDFVSTIDLPNYGCTIPEKTSCSVYGWGYTGLINYDGLLRVAHLYIM GNEKCSQHHRGKVTLNESEICAGAEKIGSGPCEGDYGGPLVCEQHKMRMVLGVIVPGRGCAI PNRPGIFVRVAYYAKWIHKIILTYKVPQS EXAMPLE 6 This example provides data examining the use of the GPEx Lightning process for producing an antibody product “Yourway” using two methods. The first method with both the Yourway heavy and light chain genes on the same expression plasmid and the second with the heavy and light chain genes on different plasmids. We also examined how varying the ratios of those heavy and light chain plasmids during the GPEx Lightning process would affect antibody expression, protein quality and cell growth characteristics. For each of GPExAttorney Docket No. CATA-40729.601 Lightning transfections 2 micrograms of plasmid was used. In the case of the different ratios of heavy to light chain plasmid, the amounts were 1) 1.8 micrograms heavy chain - 0.2 micrograms of light chain; 2) 1.6 micrograms heavy chain - 0.4 micrograms of light chain; 3) 1.4 micrograms heavy chain - 0.6 micrograms of light chain; 4) 1.2 micrograms heavy chain - 0.8 micrograms of light chain; 5) 1.0 micrograms heavy chain - 1.0 micrograms of light chain; 6) 0.8 micrograms heavy chain - 1.2 micrograms of light chain; 7) 0.6 micrograms heavy chain - 1.4 micrograms of light chain; 8) 0.4 micrograms heavy chain - 1.6 micrograms of light chain; 9) 0.2 micrograms heavy chain - 1.8 micrograms of light chain. In general, after the GPEx Lightning process, the number of each of the genes inserted into the pooled cell line followed the amount of plasmid associated with the transfection proportionally. The subsequent protein expression of each chain followed a similar trend. However, since most antibody heavy chains can be toxic to cells when a light chain is not present to bind to it during production, problems with cell growth were seen at the high ratios of heavy chain. That fact, and the GPEx Lightning process that uses a unique GS selection process appears to have skewed the high heavy chain ratio pools slightly towards the cells in that pool producing more light chain, even in the cell lines with a high ratio of heavy chain. The highest producing ratios were those that were producing extra light chain as observed in the non-reducing SDS-PAGE gels and were around the 1:1 ratio of the two chains. Since light chain genes are smaller, in general more light chain mRNA is produced / gene inserted and more light chain protein is translated per mRNA produced. Experimental Design and Results Gene constructs described herein were used as part of the GPEx Lightning process for development of the antibody heavy chain and light chain ratio pooled cell lines. The cell lines were produced using these gene constructs and standard GPEx Lightning methodology. The expression sequence for the Yourway antibody light chain using the bovine alpha-lactalbumin signal peptide for secretion from the cell is provided in FIG.20. The Yourway antibody light coding DNA sequence (CDS) and protein sequence are shown. The signal peptide included into the CDS is set off in a dark blue color. The Yourway light chain expression vector used for the GPEx Lightning transfection process is provided in FIG.21. The expression sequence for the Yourway antibody using the bovine alpha- lactalbumin signal peptide for secretion from the cell is provided in FIG.22. The YourwayAttorney Docket No. CATA-40729.601 antibody heavy coding DNA sequence (CDS) and protein sequence are shown. The signal peptide included into the CDS is set off in a dark blue color. The Yourway heavy chain expression vector used for the GPEx Lightning transfection process is depicted in FIG.23. The Yourway heavy chain and light chain (two chain) expression vector used for the GPEx Lightning transfection process is depicted in FIG.24. The GPEx Lightning technology uses a recombinase to specifically insert genes into “dock” sequences recognized by the recombinase and placed into the cell line using the GPEx process. Approximately 150-200 “dock” sequences are present in the cell line and available for gene insertion. Nine different cell pools were produced with the GPEx Lightning process using the gene constructs shown in figures 1-5. The pooled cell lines contained the following plasmid ratios / amounts . In the case of the different ratios of heavy to light chain plasmid, the amounts were 1) 1.8 micrograms heavy chain - 0.2 micrograms of light chain; 2) 1.6 micrograms heavy chain - 0.4 micrograms of light chain; 3) 1.4 micrograms heavy chain - 0.6 micrograms of light chain; 4) 1.2 micrograms heavy chain - 0.8 micrograms of light chain; 5) 1.0 micrograms heavy chain - 1.0 micrograms of light chain; 6) 0.8 micrograms heavy chain - 1.2 micrograms of light chain; 7) 0.6 micrograms heavy chain - 1.4 micrograms of light chain; 8) 0.4 micrograms heavy chain - 1.6 micrograms of light chain; 9) 0.2 micrograms heavy chain - 1.8 micrograms of light chain. As part of the GPEx Lightning process after transfection of the integrase and the transgene constructs, selection by removal of glutamine from the media occurs to allow only cell lines containing the transgene (and the glutamine synthase) to survive. How fast cell lines recover from the selection can typically indicate how well the process worked as well as if there are any issues caused by transgene expression. The pooled cell lines that contained a higher ratio of heavy to light chain gene showed slower recoveries during the selection process. Highlighting the potential toxicity of the heavy chain when not paired with a light chain. As the ratios became more favorable, the cells recovered faster. Data table for the different heavy chain to light chain ratios.Attorney Docket No. CATA-40729.601The number of gene copies was measured using a value called gene copy index. This is not actual gene copy number but a ratio of each gene to an endogenous CHO control gene. The details of the assay are shown in the appendix. The gene copy index of the pools followed the expect trend of the ratio transfected into the cells. However you can see at high ratios of heavy chain the numbers are slightly skewed away from extra heavy chain genes due to the potential toxicity and the GS selection that occurs with the process when you compare the values when light chain is in significant excess. Fed-batch productivity studies were performed on each of the cell pools in duplicate with the results averaged. The best antibody yields for this particular antibody were associated with transfection ratios of light chain to heavy chain of 1:1 or 1.2:0.8 where light chain is in slight excess. As was observed in the cell growth data, the higher heavy chain ratios grew to lower cell densities and had lower IVCD’s. Pooled cell line productivity data for each of the pool duplicates with the double chain (Light + Heavy chain) construct as the control (CT36-2) is provided in FIG.25. Productivity of some of the ratio pools reached similar or higher levels as the control construct. SDS-PAGE gel analysis of the of the pools showed the expected trends is provided in FIG.26. More product missing light chain in the pools with higher ratios of heavy chain and much more free light chain in the pools with higher light chain ratios. The control pool with both genes on the same gene construct, showed free light chain similar to the 1:1 ratio cell pools, as was expected since the number of each gene should be in a similar balance. This supports the efficiency at which the gene insertion occurs and how the ratio of genes inserting is directly related to the amount of each plasmid in the initial transfection reaction. All publications and patents mentioned in the above specification are herein incorporated by reference. Various modifications and variations of the described method and system of the invention will be apparent to those skilled in the art without departing from the scope and spirit of the invention. Although the invention has been described in connection with specificAttorney Docket No. CATA-40729.601 preferred embodiments, it should be understood that the invention as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the invention that are obvious to those skilled in the field of this invention are intended to be within the scope of the following claims.
Claims
1. Attorney Docket No. CATA-40729.601 CLAIMS What is claimed is:
1. A method comprising: introducing at least first nucleic acid constructs encoding a substrate for an enzyme and a second nucleic acid construct encoding the enzyme at a ratio of first nucleic acid constructs to second nucleic acid constructs of from 1:1 to 1000:1 into a host cell having genome comprising from 1 to 500 integrated docking sites, each docking site comprising at least one dock site insertion element and the nucleic acid constructs each comprising at least one insertion element compatible with the at least one dock site insertion element in the integrated docking sites, under conditions such that the nucleic acid expression constructs are inserted at the dock sites at a ratio of first nucleic acid constructs to second nucleic acid constructs of at least 2:
1.
2. The method of claim 1, wherein the ratio of first nucleic acid constructs to second nucleic acid constructs is from 2:1 to 1000:
1.
3. The method of claim 1, wherein the ratio of first nucleic acid constructs to second nucleic acid constructs is from 5:1 to 500:
1.
4. The method of claim 1, wherein the ratio of first nucleic acid constructs to second nucleic acid constructs is from 10:1 to 200:
1.
5. The method of claim 1, wherein the ratio of first nucleic acid constructs to second nucleic acid constructs is from 10:1 to 100:
1.
6. The method of any one of claims 1 to 5, wherein only first nucleic acid constructs and second nucleic acid constructs are introduced into the host cell.
7. The method of any one of claims 1 to 6, further comprising introducing a third nucleic acid construct encoding a third protein of interest at a ratio of first nucleic acid construct or second nucleic acid construct to third nucleic acid construct selected from the groupAttorney Docket No. CATA-40729.601 consisting of at least 2:1, from 2:1 to 1000:1, from 5:1 to 500:1, from 10:1 to 200:1, and from 10:1 to 100:
1.
8. The method of claim 7, further comprising introducing a fourth nucleic acid construct encoding a fourth protein of interest at a ratio of first nucleic acid construct, second nucleic acid construct, or third nucleic acid construct to the fourth nucleic construct selected from the group consisting of at least 2:1, from 2:1 to 1000:1, from 5:1 to 500:1, from 10:1 to 200:1, and from 10:1 to 100:
1.
9. The method of claim 8, further comprising introducing a fifth nucleic acid construct encoding a fifth protein of interest at a ratio of first nucleic acid construct, second nucleic acid construct, third nucleic acid construct or fourth nucleic construct to the fifth nucleic acid construct selected from the group consisting of at least 2:1, from 2:1 to 1000:1, from 5:1 to 500:1, from 10:1 to 200:1, and from 10:1 to 100:
1.
10. The method of any one of claims 1 to 9, wherein the at least first and second nucleic acid constructs further comprise at least the following elements in operable association in 5’ to 3’ order: an internal promoter sequence; a nucleic acid sequence encoding the first protein of interest or second protein that is operably linked to the internal promoter; and a poly A signal sequence.
11. The method of claim 10, wherein the at least first and second nucleic acid constructs comprise a selectable marker sequence.
12. The method of claim 11, wherein the at least first and second nucleic acid constructs comprise different selectable marker sequences.
13. The method of claim 10, wherein one of the first and second nucleic acid constructs comprises a selectable marker sequence and the other of the first and second nucleic acid constructs does not comprise a selectable marker sequence.Attorney Docket No. CATA-40729.601 14. The method of any one of claims 11 to 12, wherein the selectable marker sequences are 5’ to the internal promoter sequence and are operably linked to a 5’ promoter sequence.
15. The method of any one of claims 11 to 14, wherein the nucleic acid construct comprises an extending packaging region (EPR) between the 5’ promoter and the selectable marker.
16. The method of claim 15, wherein the EPR comprises multiple potential Kozak sequences and / or ATG translation start sites.
17. The method of any one of claims 10 to 16, wherein the promoter sequence is selected from the group consisting of SIN-LTR, SV40, EF1α, E. coli lac, E. coli trp, phage lambda PL, phage lambda PR, T3, T7, cytomegalovirus (CMV) immediate early, herpes simplex virus (HSV) thymidine kinase, alpha-lactalbumin, and mouse metallothionein-I promoter sequences.
18. The method of any one of claims 10 to 17, wherein the first promoter sequence is a weak promoter sequence.
19. The method of any one of claims 10 to 18, wherein the first promoter sequence is not a retroviral LTR promoter.
20. The method of any one of claims 10 to 19, wherein the integrated docking sites further comprise an exogenous promoter.
21. The method of claim 20, wherein the exogenous promoter is selected from the group consisting of SIN-LTR, SV40, EF1α, E. coli lac, E. coli trp, phage lambda PL, phage lambda PR, T3, T7, cytomegalovirus (CMV) immediate early, herpes simplex virus (HSV) thymidine kinase, alpha-lactalbumin, and mouse metallothionein-I promoter sequences.
22. The method of claim 21, wherein the promoter is a retroviral LTR.
23. The method of claim 22, wherein the retroviral LTR is a SIN LTR.Attorney Docket No. CATA-40729.601 24. The method of any one of claims 1 to 23, wherein the nucleic acid expression constructs are provided in a vector.
25. The method of claim 24, wherein the vector is a plasmid vector.
26. The method of any one of claims 24 to 25, wherein the vector is transiently introduced into the host cell.
27. The method of any one of claims 1 to 26, wherein the host cell line comprises a nucleic acid construct encoding an enzyme that facilitates insertion of the nucleic acid expression construct at the dock site.
28. The method of claim 27, wherein the nucleic acid construct encoding an enzyme that facilitates insertion of the nucleic acid expression construct at the dock site is transiently introduced into the host cell.
29. The method of claims 27 or 28, wherein the nucleic acid construct encoding an enzyme that facilitates insertion of the nucleic acid expression construct at the dock site is provided in a vector.
30. The method of claim 29, wherein the vector is a plasmid vector.
31. The method of any one of claims 27 to 30, wherein the ratio of the ratio of the nucleic acid constructs encoding an enzyme that facilitates insertion of the nucleic acid expression construct at the dock site to the nucleic acid expression constructs encoding a first protein of interest that are transiently introduced into the host cell is from 1:1000 to 1:
10.
32. The method of any one of claims 27 to 31, wherein the enzyme is selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase.
33. The method of any one of claims 1 to 32, wherein the host cell genome comprises from 5 to 500 integrated docking sites, each docking site comprising at least one dock site insertion element.Attorney Docket No. CATA-40729.601 34. The method of any one of claims 1 to 33, wherein the host cell genome comprises from 5 to 250 integrated docking sites, each docking site comprising at least one dock site insertion element.
35. The method of any one of claims 1 to 34, wherein the host cell genome comprises from 5 to 100 integrated docking sites, each docking site comprising at least one dock site insertion element.
36. The method of any one of claims 1 to 35, wherein the integrated docking sites are independently positioned throughout the host cell genome.
37. The method of any one of claims 1 to 36, wherein the dock site insertion element is targeted by enzyme selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase.
38. The method of any one of claims 1 to 37, wherein the dock site insertion element is selected from the group consisting of a recombinase dock site insertion element and a HDR dock site insertion element.
39. The method of claim 38, wherein the dock site insertion element is a recombinase dock site insertion element.
40. The method of claim 39, wherein the recombinase dock site insertion element comprises an attachment site (att).
41. The method of claim 40, wherein the attachment site (att) is selected from the group consisting of attB and attP and attR and attL.
42. The method of claim 41, wherein the recombinase dock site insertion element comprises a LoxP sequence.
43. The method of claim 39, wherein the recombinase dock site insertion element is a Flp Recombination Target (FRT) site.Attorney Docket No. CATA-40729.601 44. The method of claim 38, wherein the dock site insertion element is a HDR dock site insertion element.
45. The method of claim 44, wherein the HDR dock site insertion element comprises one or two dock site homology arms.
46. The method of claim 36, wherein the HDR dock site insertion element further comprises one or more sequences homologous to a guide RNA sequence.
47. The method of any of claims 45 to 46, wherein the dock site homology arms are from about 30 to 1000 bases in length.
48. The method of claim 37, wherein the integrase dock site insertion element comprises an AAVS1 safe harbor locus sequence.
49. The method of any one of claims 1 to 48, wherein each docking site is flanked by exogenous integrating vector sequences.
50. The method of claim 49, wherein the exogenous integrating vector sequences are selected from the group consisting of viral vector sequences and transposon vector sequences.
51. The method of any one of claims 1 to 50, wherein the docking sites each further comprise a sequence encoding a selectable maker operably linked to a promoter.
52. The method of any one of claims 1 to 51, wherein the host cell further comprises an expression construct encoding an exogenous enzyme selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase.
53. The method of claim 52, wherein the expression construct encoding an exogenous enzyme selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase is provided in an episomal expression vector.Attorney Docket No. CATA-40729.601 54. The method of claim 52, wherein the expression construct encoding an exogenous enzyme selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase is integrated into the host cell genome.
55. The method of any one of claims 1 to 54, wherein the dock site insertion element is positioned to facilitate cassette exchange.
56. The method of any one of claims 1 to 55, wherein each docking site comprises two dock site insertion elements.
57. The method of claim 56, wherein the two dock site insertion elements are positioned to facilitate cassette exchange.
58. The method of any one of claims 56 to 57, wherein the two dock site insertion elements flank sequences encoding a selectable marker, an enzyme, or a combination thereof.
59. The method of any one of claims 1 to 58, wherein the nucleic acid expression constructs further comprise a signal peptide sequence operably linked to the first protein of interest.
60. The method of claim 59, wherein the signal peptide sequence is selected from the group consisting of tissue plasminogen activator, human growth hormone, lactoferrin, alpha- casein and alpha-lactalbumin signal peptide sequences.
61. The method of any one of claims 1 to 60, wherein the nucleic acid expression constructs further comprise a protein purification marker sequence.
62. The method of claim 61, wherein the protein purification marker sequence is a hexahistidine tag or a hemagglutinin (HA) tag.
63. The method of any one of claims 1 to 62, wherein the host cell is selected from the group consisting of Chinese Hamster Ovary (CHO) cells, HEK 293 cells, CAP cells, bovine mammary epithelial cells, monkey kidney CV1 line transformed by SV40, baby hamster kidney cells, mouse sertoli cells, monkey kidney cells, African green monkey kidney cells,Attorney Docket No. CATA-40729.601 human cervical carcinoma cells, canine kidney cells, buffalo rat liver cells, human lung cells, human liver cells, mouse mammary tumor, TRI cells, MRC 5 cells, FS4 cells, rat fibroblasts, MDBK cells and human hepatoma line cells.
64. The method of claim 63, wherein the host cell is selected from the group consisting of a Chinese Hamster Ovary (CHO) cells, a HEK 293 cells and a CAP cells.
65. The method of any one of claims 63 to 64, wherein the host cell is a GS knockout cell line.
66. The method of any one of claims 63 to 64, wherein the host cell is a DHFR knockout cell line.
67. A cell culture comprising host cells made the method on any of claims 1 to 66.
68. A process for producing a protein of interest comprising culturing host cells according to claim 67 under conditions that the protein of interest is expressed and purifying the protein of interest from the host cell culture.
69. The process of claim 68, wherein the host cells are grown in a medium comprising an inhibitor of the selectable marker.
70. The process of claim 69, wherein the selectable marker is GS and the inhibitor is phosphinothricin or methionine sulphoximine (Msx).
71. The process of claim 69, wherein the selectable marker is DHFR and the inhibitor is methotrexate.
72. A host cell comprising: a plurality of docking sites integrated into the genome of the host cell, each docking site comprising at least one dock site insertion element; at least integrated first nucleic acid constructs comprising at least one insertion element compatible with the dock site insertion element and encoding a substrate of an enzyme; andAttorney Docket No. CATA-40729.601 at least integrated second nucleic acid constructs comprising at least one insertion element compatible with the dock site insertion and encoding the enzyme that acts on the substrate; wherein the at least integrated first nucleic acid constructs and the at least second integrated nucleic acid constructs are integrated at the plurality of docking sites at a ratio of first nucleic acid constructs to second nucleic acid constructs of at least 1:
1.
73. The host cell of claim 72, wherein the ratio of first nucleic acid constructs to second nucleic acid constructs is from 2:1 to 1000:
1.
74. The host cell of claim 72, wherein the ratio of first nucleic acid constructs to second nucleic acid constructs is from 5:1 to 500:
1.
75. The host cell of claim 72, wherein the ratio of first nucleic acid constructs to second nucleic acid constructs is from 10:1 to 200:
1.
76. The host cell of claim 72, wherein the ratio of first nucleic acid constructs to second nucleic acid constructs is from 10:1 to 100:
1.
77. The host cell of any one of claims 72 to 76, further comprising a third nucleic acid construct encoding a third protein of interest at a ratio of first nucleic acid construct or second nucleic acid construct to third nucleic acid construct selected from the group consisting of at least 2:1, from 2:1 to 1000:1, from 5:1 to 500:1, from 10:1 to 200:1, and from 10:1 to 100:
1.
78. The host cell of claim 77, further comprising a fourth nucleic acid construct encoding a fourth protein of interest at a ratio of first nucleic acid construct, second nucleic acid construct, or third nucleic acid construct to the fourth nucleic construct selected from the group consisting of at least 2:1, from 2:1 to 1000:1, from 5:1 to 500:1, from 10:1 to 200:1, and from 10:1 to 100:
1.
79. The host cell of claim 78, further comprising a fifth nucleic acid construct encoding a fifth protein of interest at a ratio of first nucleic acid construct, second nucleic acid construct, third nucleic acid construct or fourth nucleic construct to the fifth nucleic acid constructAttorney Docket No. CATA-40729.601 selected from the group consisting of at least 2:1, from 2:1 to 1000:1, from 5:1 to 500:1, from 10:1 to 200:1, and from 10:1 to 100:
1.
80. The host cell of any one of claims 73 to 89, wherein the at least first and second nucleic acid constructs further comprise at least the following elements in operable association in 5’ to 3’ order: an internal promoter sequence; a nucleic acid sequence encoding the first protein of interest or second protein that is operably linked to the internal promoter; and a poly A signal sequence.
81. The host cell of claim 80, wherein the at least first and second nucleic acid constructs comprise a selectable marker sequence.
82. The host cell of claim 81, wherein the at least first and second nucleic acid constructs comprise different selectable marker sequences.
83. The host cell of claim 80, wherein one of the first and second nucleic acid constructs comprises a selectable marker sequence and the other of the first and second nucleic acid constructs does not comprise a selectable marker sequence.
84. The host cell of any one of claims 81 to 83, wherein the selectable marker sequences are 5’ to the internal promoter sequence and are operably linked to a 5’ promoter sequence.
85. The host cell of any one of claims 81 to 84, wherein the nucleic acid construct comprises an extending packaging region (EPR) between the 5’ promoter and the selectable marker.
86. The host cell of claim 85, wherein the EPR comprises multiple potential Kozak sequences and / or ATG translation start sites.
87. The host cell of any one of claims 80 to 86, wherein the promoter sequence is selected from the group consisting of SIN-LTR, SV40, EF1α, E. coli lac, E. coli trp, phage lambda PL, phage lambda PR, T3, T7, cytomegalovirus (CMV) immediate early, herpes simplexAttorney Docket No. CATA-40729.601 virus (HSV) thymidine kinase, alpha-lactalbumin, and mouse metallothionein-I promoter sequences.
88. The host cell of any one of claims 80 to 87, wherein the first promoter sequence is a weak promoter sequence.
89. The host cell of any one of claims 80 to 88, wherein the first promoter sequence is not a retroviral LTR promoter.
90. The host cell of any one of claims 80 to 89, wherein the integrated docking sites further comprise an exogenous promoter.
91. The host cell of claim 90, wherein the exogenous promoter is selected from the group consisting of SIN-LTR, SV40, EF1α, E. coli lac, E. coli trp, phage lambda PL, phage lambda PR, T3, T7, cytomegalovirus (CMV) immediate early, herpes simplex virus (HSV) thymidine kinase, alpha-lactalbumin, and mouse metallothionein-I promoter sequences.
92. The host cell of claim 91, wherein the promoter is a retroviral LTR.
93. The host cell of claim 92, wherein the retroviral LTR is a SIN LTR.
94. The host cell of any one of claims 73 to 93, wherein the nucleic acid expression constructs are provided in a vector.
95. The host cell of claim 94, wherein the vector is a plasmid vector.
96. The host cell of any one of claims 72 to 95, wherein the host cell comprises a nucleic acid construct encoding an enzyme that facilitates insertion of the nucleic acid expression construct at the dock site.
97. The host cell of claim 96, wherein the nucleic acid construct encoding an enzyme that facilitates insertion of the nucleic acid expression construct at the dock site is provided in a vector.Attorney Docket No. CATA-40729.601 98. The host cell of claim 97, wherein the vector is a plasmid vector.
99. The host cell of any one of claims 95 to 98, wherein the ratio of the nucleic acid constructs encoding an enzyme that facilitates insertion of the nucleic acid expression construct at the dock site to the nucleic acid expression constructs encoding a first protein of interest that are transiently introduced into the host cell is from 1:1000 to 1:
10.
100. The host cell of any one of claims 96 to 99, wherein the enzyme is selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase.
101. The host cell of any one of claims 72 to 100, wherein the host cell genome comprises from 5 to 500 integrated docking sites, each docking site comprising at least one dock site insertion element.
102. The host cell of any one of claims 72 to 101, wherein the host cell genome comprises from 5 to 250 integrated docking sites, each docking site comprising at least one dock site insertion element.
103. The host cell of any one of claims 72 to 102, wherein the host cell genome comprises from 5 to 100 integrated docking sites, each docking site comprising at least one dock site insertion element.
104. The host cell of any one of claims 72 to 103, wherein the integrated docking sites are independently positioned throughout the host cell genome.
105. The host cell of any one of claims 72 to 104, wherein the dock site insertion element is targeted by enzyme selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase.
106. The host cell of any one of claims 72 to 105, wherein the dock site insertion element is selected from the group consisting of a recombinase dock site insertion element and a HDR dock site insertion element.Attorney Docket No. CATA-40729.601 107. The host cell of claim 106, wherein the dock site insertion element is a recombinase dock site insertion element.
108. The host cell of claim 107, wherein the recombinase dock site insertion element comprises an attachment site (att).
109. The host cell of claim 108, wherein the attachment site (att) is selected from the group consisting of attB and attP and attR and attL.
110. The host cell of claim 107, wherein the recombinase dock site insertion element comprises a LoxP sequence.
111. The host cell of claim 107, wherein the recombinase dock site insertion element is a Flp Recombination Target (FRT) site.
112. The host cell of claim 106, wherein the dock site insertion element is a HDR dock site insertion element.
113. The host cell of claim 112, wherein the HDR dock site insertion element comprises one or two dock site homology arms.
114. The host cell of claim 112, wherein the HDR dock site insertion element further comprises one or more sequences homologous to a guide RNA sequence.
115. The host cell of any of claims 113 to 114, wherein the dock site homology arms are from about 30 to 1000 bases in length.
116. The host cell of claim 105, wherein the integrase dock site insertion element comprises an AAVS1 safe harbor locus sequence.
117. The host cell of any one of claims 72 to 116, wherein each docking site is flanked by exogenous integrating vector sequences.Attorney Docket No. CATA-40729.601 118. The host cell of claim 117, wherein the exogenous integrating vector sequences are selected from the group consisting of viral vector sequences and transposon vector sequences.
119. The host cell of any one of claims 72 to 118, wherein the docking sites each further comprise a sequence encoding a selectable maker operably linked to a promoter.
120. The host cell of any one of claims 72 to 119, wherein the host cell further comprises an expression construct encoding an exogenous enzyme selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase.
121. The host cell of claim 120, wherein the expression construct encoding an exogenous enzyme selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase is provided in an episomal expression vector.
122. The host cell of claim 120, wherein the expression construct encoding an exogenous enzyme selected from the group consisting of an integrase, a recombinase, a nuclease and a nickase is integrated into the host cell genome.
123. The host cell of any one of claims 72 to 122, wherein the dock site insertion element is positioned to facilitate cassette exchange.
124. The host cell of any one of claims 72 to 123, wherein each docking site comprises two dock site insertion elements.
125. The host cell of claim 124, wherein the two dock site insertion elements are positioned to facilitate cassette exchange.
126. The host cell of any one of claims 124 to 125, wherein the two dock site insertion elements flank sequences encoding a selectable marker, an enzyme, or a combination thereof.
127. The host cell of any one of claims 72 to 126, wherein the nucleic acid expression constructs further comprise a signal peptide sequence operably linked to the first protein of interest.Attorney Docket No. CATA-40729.601 128. The host cell of claim 127, wherein the signal peptide sequence is selected from the group consisting of tissue plasminogen activator, human growth hormone, lactoferrin, alpha- casein and alpha-lactalbumin signal peptide sequences.
129. The host cell of any one of claims 72 to 128, wherein the nucleic acid expression constructs further comprise a protein purification marker sequence.
130. The host cell of claim 129, wherein the protein purification marker sequence is a hexahistidine tag or a hemagglutinin (HA) tag.
131. The host cell of any one of claims 72 to 130, wherein the host cell is selected from the group consisting of Chinese Hamster Ovary (CHO) cells, HEK 293 cells, CAP cells, bovine mammary epithelial cells, monkey kidney CV1 line transformed by SV40, baby hamster kidney cells, mouse sertoli cells, monkey kidney cells, African green monkey kidney cells, human cervical carcinoma cells, canine kidney cells, buffalo rat liver cells, human lung cells, human liver cells, mouse mammary tumor, TRI cells, MRC 5 cells, FS4 cells, rat fibroblasts, MDBK cells and human hepatoma line cells.
132. The host cell of claim 131, wherein the host cell is selected from the group consisting of a Chinese Hamster Ovary (CHO) cells, a HEK 293 cells and a CAP cells.
133. The host cell of any one of claims 131 to 132, wherein the host cell is a GS knockout cell line.
134. The host cell of any one of claims 131 to 132, wherein the host cell is a DHFR knockout cell line.
135. A cell culture comprising host cells of any of claims 72 to 134.
136. A process for producing a protein of interest comprising culturing host cells according to any of claims 72 to 134 under conditions that the protein(s) of interest are expressed and purifying the protein(s) of interest from the host cell culture.Attorney Docket No. CATA-40729.601 137. The process of claim 136, wherein the host cells are grown in a medium comprising an inhibitor of the selectable marker.
138. The process of claim 137, wherein the selectable marker is GS and the inhibitor is phosphinothricin or methionine sulphoximine (Msx).
139. The process of claim 137, wherein the selectable marker is DHFR and the inhibitor is methotrexate.