Metabolic selection via the alanine biosynthesis pathway

JP2026527540APending Publication Date: 2026-08-14EMD MILLIPORE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2026-08-14

Smart Images

  • Figure 2026527540000001_ABST
    Figure 2026527540000001_ABST
Patent Text Reader

Abstract

This disclosure provides isolated mammalian cells in which the expression of glutamate-pyruvate transaminase 2 (GPT2) is reduced or absent. Furthermore, methods for preparing such cells and methods for using these cells in the production of recombinant proteins are provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims priority to U.S. Provisional Patent Application No. 63 / 516,653, filed on 31 July 2023, the entirety of which is incorporated herein by reference.

[0002] Technical field This disclosure relates to mammalian cell lines for use in biological production systems, wherein the mammalian cell lines are engineered to have reduced or absent expression of components of the alanine biosynthesis pathway in order to create alanine-dependent cell lines. [Background technology]

[0003] background The development of high-yielding clonal cell lines for biomanufacturing typically utilizes one or more well-known selection methods, such as glutamine synthase (GS for glutamine selection), dihydrofolate receptor (DHFR for hypoxatine and thymidine selection), antibiotic selection (puromycin, hygromycin, blastocydin, etc.), and P5C synthase (P5CS-proline selection). While the GS system is the industry standard, there is a need for cell lines that allow for multiple selection methods, enabling the introduction of multiple vectors into the cell line to facilitate the production of molecules such as bispecific antibodies, multispecific antibodies, other multichain enzymes / proteins, or proteins / enzymes that require effector proteins for expression. [Overview of the project]

[0004] overview Various aspects of this disclosure include the provision of mammalian cell lines for use in biological production systems, in which the expression of the endogenous glutamate-pyruvate transaminase 2 or GPT2 gene (also known as alanine aminotransferase 2 or ALT2) is reduced or lost. In the absence of endogenously expressed functional GPT2 protein, cells require an exogenous source of the amino acid alanine. The GPT2 sequence on the chromosome can be inactivated using genomic modification via targeted endonucleases such as the CRISPR ribonucleoprotein (RNP) complex or zinc finger nucleases. Another aspect of this disclosure is the provision of mammalian cell lines, in which the expression of the endogenous GPT2 gene is reduced or lost and the expression of the endogenous glutamine synthase (GS) gene is reduced or lost.

[0005] Another aspect of this disclosure includes a process for selecting cell lines that enhance the productivity of the biotherapeutic protein to be expressed. Another aspect of this disclosure is to provide a more convenient bioproduction system for the expression of bispecific antibodies or biotherapeutic proteins that require the expression of effector proteins by utilizing multiple selection systems. The process includes expressing at least one recombinant protein in any of the mammalian cell lines.

[0006] Other aspects and iterations of this disclosure are described in more detail below. [Brief explanation of the drawing]

[0007] [Figure 1]Figure 1 illustrates the role of glutamate-pyruvate transaminase 2 in the final stage of alanine biosynthesis. The final stage of alanine biosynthesis is completed by a single pathway requiring the non-redundant enzyme glutamate-pyruvate transaminase 2 (GPT2), also known as alanine aminotransferase 2 (ALT2). In the absence of a functional GPT2 protein, cells require an external source of the amino acid alanine. [Figure 2] Figure 2 shows the cDNA sequence of GPT2 in CHOZN® GS- / -CHO cells, including the preferred gRNA binding site (underlined). The figure shows sequence number 32. [Figure 3] Figure 3 shows the alteration of the GPT2 coding sequence resulting from Cas9 cleavage. Genetic changes in the GPT2 coding sequence were detected by clonal amplicon sequencing, resulting in frameshift mutations and inactivation of the GPT2 protein. The figure shows sequence numbers 33-35 in order of appearance. [Figure 4] Figure 4 shows the vectors used in the glutamine and / or alanine-based selection systems. The vectors were designed to enable the selection of -GFP-positive cells using the glutamine-based selection system, the selection of -BFP-positive cells using the alanine-based selection system, and the development of secreted recombinant proteins by developing two similar vectors containing either the mAb heavy chain, light chain, and either the GPT2 or GS coding sequence. [Figure 5] Figure 5 shows effective alanine selection by GPT2 knockout clones. A group of identified GPT2 KO clones were shown to be unable to survive and grow without supplementation of alanine (Ala) in the culture medium. The culture medium was replaced on days 4 and 9. [Figure 6]Figure 6 shows the expression of GFP and / or BFP in GPT knockout clones. CHO cells in which the GS gene and the GPT2 gene have been genetically disrupted were mock-transfected with water and cultured in a medium supplemented with 6 mM glutamine and 4 mM alanine, and neither GFP nor BFP was expressed (A). CHO cells in which the GS gene and the GPT2 gene have been genetically disrupted and transfected with a plasmid containing the GS+GFP expression cassette were selected in glutamine-free medium and expressed GFP, but not BFP (B). CHO cells in which the GS gene and the GPT2 gene have been genetically disrupted and transfected with a plasmid containing the GPT2+BFP expression cassette were selected in alanine-free medium and expressed low levels of BFP, but not GFP (C). CHO cells in which the GS gene and the GPT2 gene have been genetically disrupted and co-transfected with GS+GFP and GPT2+BFP and cultured in alanine- and glutamine-free medium expressed both GFP and BFP (D). [Figure 7] Figure 7 shows the selection of cells transfected with GFP / GS and / or BFP / GPT2 plasmids. CHO cells in which the GS gene and the GPT2 gene have been genetically disrupted were transfected with either water (Mock), a plasmid containing the GS+GFP expression cassette (xGS), a plasmid containing the GPT2+BFP expression cassette (xGPT2), or a plasmid containing GS+GFP and GPT2+BFP (xGS+GPT2). These data show the viability and growth of the transfected populations when selection is performed by growing them in a medium without a selection factor (+Glut+Ala), GS selection only (-Glut+Ala), GPT2 selection only (+Glut-Ala), or GS+GPT2 selection (-Glut-Ala), respectively.

Mode for Carrying Out the Invention

[0008] Detailed Description The present disclosure provides mammalian cell lines engineered to have reduced or absent expression of the endogenous GPT2 gene. Also provided are mammalian cell lines engineered to have reduced or absent expression of the endogenous GS gene and reduced or absent expression of the endogenous GPT2 gene. Methods for producing the engineered cell lines and methods for selecting and using the engineered cell lines to produce recombinant proteins are provided.

[0009] (I) Engineered cell lines One aspect of the present disclosure encompasses mammalian cell lines engineered to have reduced or absent expression of the endogenous GPT2 gene. Alternatively, the mammalian cell lines are engineered to have reduced or absent expression of both the endogenous GPT2 gene and the endogenous GS gene.

[0010] The cell lines disclosed herein having reduced or absent expression of GPT2, or having reduced expression of GPT2 and GS, are genetically engineered to modify the chromosomal sequences encoding the GPT2 or GS proteins. The chromosomal sequences can be modified using genome editing techniques via a targeted endonuclease, which is described in detail in Section (III) below. For example, the chromosomal sequences can be modified to include at least one nucleotide deletion, at least one nucleotide insertion, at least one nucleotide substitution, or combinations thereof such that the reading frame is shifted and no protein product is produced (i.e., the chromosomal sequence is inactivated). Inactivating one allele of the chromosomal sequence encoding either GPT2 or GS reduces protein expression (i.e., knockdown). Inactivating both alleles of the chromosomal sequence encoding either GPT2 or GS results in no protein expression (i.e., knockout).

[0011] In some embodiments, the expression level of GPT2 can be reduced by at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or more than about 99%. In other embodiments, the expression level of GPT2 can be reduced to a level that is undetectable using standard methods in the art (e.g., Western immunoblotting assays, ELISA enzyme assays, SDS polyacrylamide gel electrophoresis).

[0012] Generally, the cell viability, viable cell density, titer, proliferation rate, proliferation response, cell morphology, levels of apoptosis and autophagy, and / or overall cell health of the engineered cell lines disclosed herein are equivalent to those of unengineered parental cells when supplemented with alanine and / or exogenous GPT2 coding sequences.

[0013] (a) Cell type The engineered cell lines disclosed herein are mammalian cell lines. In some embodiments, the engineered cell lines may be derived from human cell lines. Non-limiting examples of suitable human cell lines include human embryonic kidney cells (HEK293, HEK293T); human connective tissue cells (HT-1080); human cervical cancer cells (HELA); human embryonic retinal cells (PER.C6); human kidney cells (HKB-11); human hepatocytes (Huh-7); human lung cells (W138); human hepatocytes (Hep G2); human U2-OS osteosarcoma cells, human A549 lung cells, human A-431 epithelial cells, CACO-2 human colorectal adenocarcinoma cells, human pluripotent stem cells, Jurkat human T lymphocytes, or human K562 bone marrow cells. In other embodiments, the engineered cell lines may be derived from non-human cell lines. Suitable cell lines include: Chinese hamster ovary (CHO) cells; baby hamster kidney (BHK) cells; mouse myeloma NS0 cells; mouse myeloma Sp2 / 0 cells; mouse mammary gland C127 cells; mouse embryonic fibroblast 3T3 cells (NIH3T3); mouse B lymphoma A20 cells; mouse melanoma B16 cells; mouse myoblast C2C12 cells; mouse embryonic mesenchymal C3H-10T1 / 2 cells; mouse cancer CT26 cells; and mouse prostate DuCuP cells. Cells; mouse mammary EMT6 cells; mouse hepatocytoma Hepa1c1c7 cells; mouse myeloma J5582 cells; mouse epithelial MTD-1A cells; mouse cardiomyocyte MyEnd cells; mouse kidney RenCa cells; mouse pancreatic RIN-5F cells; mouse melanoma X64 cells; mouse lymphoma YAC-1 cells; rat glioblastoma 9L cells; rat B lymphoma RBL cells; rat neuroblastoma B35 cells; rat hepatocytes (HTC); buffalo rat liver BRL 3A cells; canine kidney cells (MDCK); canine mammary gland (CMT) cells; rat osteosarcoma D17 cells; rat monocyte / macrophage DH82 cells; monkey kidney SV-40 transformed fibroblast (COS7) cells; monkey kidney CVI-76 cells; or African green monkey kidney (VERO, VERO-76) cells. An extensive list of mammalian cell lines is described in the American Type Culture Collection catalog (ATCC, Manassas, VA). In some embodiments, the cell lines disclosed herein are other than mouse cell lines.In a specific embodiment, the manipulated cell line is a CHO cell line. Suitable CHO cell lines are CHO-K1, CHO-K1SV, and CHO GS. - / - This includes, but is not limited to, CHO S, DG44, DuxB11, and cell lines derived therefrom.

[0014] In various embodiments, the parental cell line may lack glutamine synthase (GS), dihydrofolate reductase (DHFR), hypoxanthine-guanine phosphoribosyltransferase (HPRT), or a combination thereof. For example, chromosomal sequences encoding GS, DHFR, and / or HPRT may be inactivated. In certain embodiments, all chromosomal sequences encoding GS, DHFR, and / or HPRT are inactivated in the parental cell line.

[0015] (b) A choice of nucleic acids that encode a recombinant protein In some embodiments, the engineered cell lines disclosed herein may further comprise at least one nucleic acid encoding a recombinant protein. Generally, recombinant proteins are heterogeneous, meaning that the protein is not native to the cell. Recombinant proteins may be, without limitation, antibodies, antibody fragments, monoclonal antibodies, humanized antibodies, humanized monoclonal antibodies, chimeric antibodies, IgG molecules, IgG heavy chains, IgG light chains, IgA molecules, IgD molecules, IgE molecules, IgM molecules, vaccines, growth factors, cytokines, interferons, interleukins, hormones, coagulation (or blood clotting) factors, blood components, enzymes, therapeutic proteins, nutritional supplement proteins, functional fragments or functional variants of any of the above, or therapeutic proteins selected from fusion proteins and / or functional fragments or variants of any of the above proteins. In certain embodiments, the recombinant protein may be a bispecific antibody or a multispecific antibody, or a protein that requires an effector protein for expression.

[0016] In some embodiments, nucleic acids encoding recombinant proteins may be ligated to sequences encoding GPT2, ASNS, PSPH, SHMT2, HPRT, DHFR, and / or GS, so that GPT2, asparagine synthase (ASNS), hypoxanthine-guanine phosphoribosyltransferase (HPRT), dihydrofolate reductase (DHFR), and / or glutamine synthase (GS) can be used as selection markers. Nucleic acids encoding recombinant proteins may also be ligated to sequences encoding at least one antibiotic resistance gene and / or marker proteins such as fluorescent proteins. In some embodiments, nucleic acids encoding recombinant proteins may be part of an expression construct. The expression construct or vector may include additional expression regulatory sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selection marker sequences, origins of replication, etc. Additional information can be found in "Current Protocols in Molecular Biology" by Ausubel et al., John Wiley & Sons, New York, 2003, or in "Molecular Cloning: A Laboratory Manual" by Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001.

[0017] In some embodiments, nucleic acids encoding recombinant proteins may be located outside the chromosome. That is, nucleic acids encoding recombinant proteins may be transiently expressed from plasmids, cosmids, artificial chromosomes, minichromosomes, or other extrachromosomal constructs. In other embodiments, nucleic acids encoding recombinant proteins may be integrated into the cell genome, onto the chromosome. Integration may be random or targeted. Thus, recombinant proteins may be stably expressed. In some iterations of this embodiment, nucleic acid sequences encoding recombinant proteins may be operably linked to appropriate heterologous expression regulatory sequences (i.e., promoters). In other iterations, nucleic acid sequences encoding recombinant proteins may be located under the control of endogenous expression regulatory sequences. Nucleic acid sequences encoding recombinant proteins may be integrated into the cell line genome using homologous recombination, genome editing via targeted endonucleases, viral vectors, transposons, cassette exchange systems via recombinases, plasmids, and other well-known means. Additional guidance can be found in Ausubel et al. 2003 and Sambrook & Russell, 2001.

[0018] (II) Kit Further aspects of this disclosure provide kits for the production of recombinant proteins, the kits comprising one of the engineered cell lines described in detail in Section (I) above. The kits may further include cell growth media, transfection reagents, plasmid vectors, selective media, recombinant protein purification means, buffers, etc. The kits provided herein generally include instructions for growing the cell lines and using them for the production of recombinant proteins. The instructions included in the kit may be affixed to the packaging or included as accompanying documentation. The instructions are typically, but are not limited to, documents or printed materials. Any medium on which such instructions can be recorded and communicated to an end user is contemplated by this disclosure. Such media include, but are not limited to, electronic storage media (e.g., magnetic disks, tapes, cartridges, chips), optical media (e.g., CD-ROMs), etc. The term “instructions” as used herein may include the address of an internet site providing the instructions.

[0019] (III) Methods for preparing engineered cell lines Another aspect of this disclosure provides methods for preparing or manipulating cell lines having reduced or absent expression of GPT2 and / or GS, as described in Section (I) above. The chromosomal sequences encoding GPT2 and / or GS can be knocked down or knocked out using a variety of techniques. Generally, the manipulated cell lines are prepared using genome modification processes mediated by targeted endonucleases. Those skilled in the art will understand that the manipulated cell lines can also be prepared using site-directed recombination systems, random mutagenesis, or other methods known in the art.

[0020] Generally, the engineered cell line is prepared by a method comprising introducing at least one targeted endonuclease or nucleic acid encoding the targeted endonuclease into the parent cell line of interest, where the targeted endonuclease is targeted to the chromosomal sequences encoding GPT2 and / or GS. The targeted endonuclease recognizes and binds to a specific chromosomal sequence, introducing a double-strand break. In some embodiments, the double-strand break is repaired by a non-homologous end-joining (NHEJ) repair process. NHEJs are error-prone, so deletions, insertions, and / or substitutions of at least one nucleotide may occur, thereby disrupting the reading frame of the chromosomal sequence and preventing the production of protein products, or producing non-functional proteins, for example, through disruption of the enzymatic active site of the protein. In other embodiments, the targeted endonuclease may also be used to alter the chromosomal sequence via homologous recombination by co-introducing a polynucleotide that has substantial sequence identity to a portion of the target chromosomal sequence. In such situations, double-strand breaks introduced by targeted endonucleases are repaired by homology-directed repair processes in which the chromosome sequence is exchanged with the polynucleotide in a manner that alters or modifies the chromosome sequence (e.g., by fusion of exogenous sequences).

[0021] (a) Targeted endonucleases Various targeted endonucleases may be used to modify the chromosomal sequences encoding GPT2 and / or GS. Targeted endonucleases may be naturally occurring proteins or engineered proteins. Suitable targeted endonucleases include, but are not limited to, zinc finger nucleases (ZFNs), CRISPR nucleases, transcription activator-like effector (TALE) nucleases (TALENs), meganucleases, chimeric nucleases, site-specific endonucleases, and artificially targeted DNA double-strand break inducers.

[0022] (i) Zinc finger nuclease In certain embodiments, targeted endonucleases may be a pair of zinc finger nucleases (ZFNs). ZFNs bind to a specific target sequence and introduce double-strand breaks at the targeted cleavage site. Typically, a ZFN comprises a DNA-binding domain (i.e., a zinc finger) and a cleavage domain (i.e., a nuclease), each of which is described below.

[0023] DNA binding domain DNA-binding domains or zinc fingers can be engineered to recognize and bind to any selected nucleic acid sequence. For example, Beerli et al. (2002) Nat. Biotechnol. 20:135-141; Pabo et al. (2001) Ann. Rev. Biochem. 70:313-340; Isalan et al. (2001) Nat. Biotechnol. 19:656-660; Segal et al. (2001) Curr. Opin. Biotechnol. 12:632-637; Choo et al. (2000) Curr. Opin. Struct. Biol. 10:411-416;Zhang et al. (2000) J. Biol. Chem. 275(43):33850-33860;Doyon et al. (2008) Nat. Biotechnol. 26:702-708; and Santiago et al. See (2008) Proc. Natl. Acad. Sci. USA 105:5809-5814. The manipulated zinc finger binding domains may have novel binding specificity compared to naturally occurring zinc finger proteins. Manipulation methods include, but are not limited to, rational design and various types of selection. Rational design includes, for example, using a database containing duplex, triplex, and / or quadriplex nucleotide sequences and individual zinc finger amino acid sequences, where each duplex, triplex, or quadriplex nucleotide sequence is associated with one or more amino acid sequences of zinc fingers that bind to a particular triplex or quadriplex sequence. See, for example, U.S. Patents 6,453,242 and 6,534,261, whose entire disclosure is incorporated herein by reference. As an example, the algorithm described in U.S. Patent 6,453,242 may be used to design zinc finger binding domains that target pre-selected sequences.Alternative methods may be used to design zinc finger binding domains that target specific sequences, such as rational design using a nondegenerate recognition code table (Sera et al. (2002) Biochemistry 41:7074-7081). Publicly available web-based tools for identifying potential target sites in DNA sequences and designing zinc finger binding domains are known in the art. For example, tools for identifying potential target sites in DNA sequences can be found at zincfingertools.org. Tools for designing zinc finger binding domains can be found at zifit.partners.org / ZiFiT (see also Mandell et al. (2006) Nuc. Acid Res. 34:W516-W523; Sander et al. (2007) Nuc. Acid Res. 35:W599-W605).

[0024] A zinc finger-binding domain may be designed to recognize and bind to a DNA sequence in the range of approximately 3 to 21 nucleotides in length. In one embodiment, a zinc finger-binding domain may be designed to recognize and bind to a DNA sequence in the range of approximately 9 to 18 nucleotides in length. Generally, the zinc finger-binding domain of a zinc finger nuclease used herein includes at least three zinc finger recognition regions or zinc fingers, each of which binds to 3 nucleotides. In one embodiment, the zinc finger-binding domain includes four zinc finger recognition regions. In another embodiment, the zinc finger-binding domain includes five zinc finger recognition regions. In yet another embodiment, the zinc finger-binding domain includes six zinc finger recognition regions. A zinc finger-binding domain may be designed to bind to any suitable target DNA sequence. See, for example, U.S. Patents 6,607,882; 6,534,261 and 6,453,242, whose entire disclosure is incorporated herein by reference.

[0025] Exemplary methods for selecting zinc finger recognition regions include phage displays and two-hybrid systems, which are described in U.S. Patents 5,789,538, 5,925,523, 6,007,988, 6,013,453, 6,410,248, 6,140,466, 6,200,759, and 6,242,568, as well as International Publications 98 / 37186, 98 / 53057, 00 / 27878, 01 / 88197, and UK Patent No. 2,338,237, each of which is incorporated herein by reference in its entirety. In addition, enhancement of the binding specificity of zinc finger binding domains is described, for example, in International Publication 02 / 077227, the entire disclosure of which is incorporated herein by reference.

[0026] Methods for designing and constructing zinc finger-binding domains and fusion proteins (and polynucleotides encoding them) are known to those skilled in the art and are described in detail, for example, in U.S. Patent No. 7,888,121, which is incorporated herein by reference in its entirety. Zinc finger recognition domains and / or multi-finger zinc finger proteins may be linked using appropriate linker sequences, for example, including linkers of 5 amino acids or more in length. For non-limiting examples of linker sequences of 6 amino acids or more in length, see U.S. Patents No. 6,479,626, 6,903,185, and 7,153,949, whose entire disclosures are incorporated herein by reference. The zinc finger-binding domains described herein may include appropriate linker combinations between individual zinc fingers of a protein.

[0027] Disconnected domain Zinc finger nucleases also contain cleavage domains. The cleavage domain portion of a zinc finger nuclease can be obtained from any endonuclease or exonuclease. Non-limiting examples of endonucleases from which the cleavage domain can be derived include, but are not limited to, restriction endonucleases and homing endonucleases. See, for example, the New England Biolabs Catalog or Belfort et al. (1997) Nucleic Acids Res. 25:3379-3388. Other enzymes that cleave DNA are also known (e.g., S1 nuclease; mung bean nuclease; pancreatic DNase I; micrococcus nuclease; yeast HO endonuclease). See also Linn et al. (eds.) Nucleases, Cold Spring Harbor Laboratory Press, 1993. One or more of these enzymes (or their functional fragments) can be used as a source of the cleavage domain.

[0028] The cleavage domain can also be derived from an enzyme or a part thereof that requires dimerization for cleavage activity, as described above. If each nuclease contains monomers of the active enzyme dimer, then two zinc finger nucleases may be required for cleavage. Alternatively, a single zinc finger nuclease may contain both monomers and form the active enzyme dimer. As used herein, “active enzyme dimer” refers to an enzyme dimer capable of cleaving nucleic acid molecules. The two cleavage monomers may originate from the same endonuclease (or its functional fragment), or each monomer may originate from a different endonuclease (or its functional fragment).

[0029] When two cleaved monomers are used to form an active enzyme dimer, it is preferable that the recognition sites of the two zinc fingers be positioned such that the cleaved monomers are spatially oriented relative to each other, enabling the two zinc fingers to bind to their respective recognition sites and form an active enzyme dimer (e.g., by dimerization). As a result, the nearest edges of the recognition sites can be separated by about 5 to about 18 nucleotides. For example, the nearest edges can be separated by about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 nucleotides. However, it will be understood that any integer number of nucleotides or nucleotide pairs can be interposed between the two recognition sites (e.g., about 2 to about 50 nucleotide pairs or more). For example, as described in detail herein, the nearest edges of the recognition sites of a zinc finger nuclease can be separated by 6 nucleotides. Generally, the cleavage sites are located between the recognition sites.

[0030] Restriction endonucleases (restriction enzymes) are present in many species and can sequence-specifically bind to DNA (at the recognition site) and cleave DNA at or near the binding site. Certain restriction enzymes (e.g., Type IIS) cleave DNA at a site distant from the recognition site and have separable binding and cleavage domains. For example, the Type IIS enzyme FokI catalyzes a double-strand break of DNA at a position 9 nucleotides from the recognition site on one strand and 13 nucleotides from the recognition site on the other strand. See, for example, U.S. Patents 5,356,802, 5,436,150, and 5,487,994; and Li et al. (1992) Proc. Natl. Acad. Sci. USA 89:4275-4279; Li et al. (1993) Proc. Natl. Acad. Sci. USA 90:2764-2768; Kim et al. (1994a) Proc. Natl. Acad. Sci. USA 91:883-887; Kim et al. (1994b) J. Biol. Chem. 269:31978-31982. Thus, a zinc finger nuclease may contain at least one Type IIS restriction enzyme cleavage domain and one or more zinc finger binding domains, which may or may not be manipulated. Exemplary Type IIS restriction enzymes are described, for example, in International Publication No. 07 / 014,275, the entire disclosure of which is incorporated herein by reference. Additional restriction enzymes also contain separable binding and cleavage domains, which are also contemplated in this disclosure. See, for example, Roberts et al. (2003) Nucleic Acids Res. 31:418-420.

[0031] An exemplary Type IIS restriction enzyme in which the cleavage domain is separable from the binding domain is FokI. This particular enzyme is active as a dimer (Bitinaite et al. (1998) Proc. Natl. Acad. Sci. USA 95: 10, 570-10, 575). Therefore, in this disclosure, the portion of the FokI enzyme used as a zinc finger nuclease is considered a cleavage monomer. Thus, for targeted double-strand cleavage using the Fok cleavage domain, two zinc finger nucleases, each containing a FokI cleavage monomer, may be used to reconstitute the active enzyme dimer. Alternatively, a single polypeptide molecule containing a zinc finger binding domain and two FokI cleavage monomers may also be used.

[0032] In certain embodiments, the cleavage domain comprises one or more manipulated cleavage monomers that minimize or inhibit homodimerization. As a non-limiting example, the amino acid residues at positions 446, 447, 479, 483, 484, 486, 487, 490, 491, 496, 498, 499, 500, 531, 534, 537, and 538 of FokI are all targets for influencing the dimerization of the Fok cleavage half-domain. Exemplary manipulated cleavage monomers of FokI that form obligate heterodimers include pairs in which the first cleavage monomer contains mutations at amino acid residue positions 490 and 538 of FokI, and the second cleavage monomer contains mutations at amino acid residue positions 486 and 499.

[0033] Therefore, in one embodiment of the manipulated cleaved monomer, a mutation at amino acid position 490 replaces Glu(E) with Lys(K); a mutation at amino acid residue 538 replaces Iso(I) with Lys(K); a mutation at amino acid residue 486 replaces Gln(Q) with Glu(E); and a mutation at position 499 replaces Iso(I) with Lys(K). Specifically, the manipulated cleaved monomer can be prepared by mutating a certain cleaved monomer from E to K at position 490 and from I to K at position 538 to produce a manipulated cleaved monomer named "E490K:I538K", and by mutating another cleaved monomer from Q to E at position 486 and from I to K at position 499 to produce a manipulated cleaved monomer named "Q486E:I499K". The manipulated cleaved monomers described above are obligate heterodimer mutants in which abnormal cleavage is minimized or eliminated. Manipulated cleaved monomers can be prepared using appropriate methods, for example, by site-directed mutagenesis of a wild-type cleaved monomer (FokI) as described in U.S. Patent No. 7,888,121, which is incorporated herein by reference.

[0034] Additional domains In some embodiments, zinc finger nucleases further include at least one nuclear localization sequence (NLS). An NLS is an amino acid sequence that facilitates the targeting of a zinc finger nuclease protein into the nucleus to introduce a double-strand break into a target sequence in a chromosome. Nuclear localization signals are known in the art (see, for example, Lange et al., J. Biol. Chem., 2007, 282:5101-5105). Non-limiting examples of nuclear localization signals include PKKKRKV (SEQ ID NO: 1), PKKKRRV (SEQ ID NO: 2), KRPAATKKAGQAKKKK (SEQ ID NO: 3), YGRKKRRQRRR (SEQ ID NO: 4), RKKRRQRRR (SEQ ID NO: 5), PAAKRVKLD (SEQ ID NO: 6), RQRRNELKRSP (SEQ ID NO: 7), VSRKRPRP (SEQ ID NO: 8), PPKKARED (SEQ ID NO: 9), PQPKKKPL (SEQ ID NO: 10), SALIKKKKKKMAP (SEQ ID NO: 11), This includes PKQKKRK (SEQ ID NO: 12), RKLKKKIKKL (SEQ ID NO: 13), REKKKFLKRR (SEQ ID NO: 14), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 15), RKCLQAGMNLEARKTKK (SEQ ID NO: 16), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 17), and RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 18). NLS can be located at the N-terminus, C-terminus, or inside of zinc finger nucleases.

[0035] In a further embodiment, the zinc finger nuclease may also include at least one cell membrane permeability domain. Suitable cell membrane permeability domains include, but are not limited to, GRKKRRQRRRPPQPKKKRKV (SEQ ID NO: 19), PLSSIFSRIGDPPKKKRKV (SEQ ID NO: 20), GALFLGWLGAAGSTMGAPKKKRKV (SEQ ID NO: 21), GALFLGFLGAAGSTMGAWSQPKKKRKV (SEQ ID NO: 22), KETWWETWWTEWSQPKKKRKV (SEQ ID NO: 23), YARAAARQARA (SEQ ID NO: 24), THLPRRRRR (SEQ ID NO: 25), GGRRRRR (SEQ ID NO: 26), RRQRRRTSKLMKR (SEQ ID NO: 27), GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 28), KALAWEAKLAKALAKHLAKALAKALKCEA (SEQ ID NO: 29), and RQIKIWFQNRRMKWKK (SEQ ID NO: 30). The cell membrane permeable domain can be located at the N-terminus, C-terminus, or within the zinc finger nuclease.

[0036] In yet another embodiment, the zinc finger nuclease may further include at least one marker domain. Non-limiting examples of marker domains include fluorescent proteins, purification tags, and epitope tags. In one embodiment, the marker domain may be a fluorescent protein. Non-limiting examples of suitable fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, EGFP, Emerald, Azami Green, Monomeric Azami). Green fluorescent proteins (e.g., CopGFP, AceGFP, ZsGreen1), yellow fluorescent proteins (e.g., YFP, EYFP, Citrine, Venus, YPet, PhiYFP, ZsYellow1), blue fluorescent proteins (e.g., EBFP, EBFP2, Azurite, mKalama1, GFPuv, Sapphire, T-sapphire), cyan fluorescent proteins (e.g., ECFP, Cerulean, CyPet, AmCyan1, Midoriishi-Cyan), red fluorescent proteins (mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRed1, AsRed2, eqFP611, mRaspberry, mStrawberry, Jred), and orange fluorescent proteins (mOrange, mKO, Kusabira-Orange, Monomeric) This includes Kusabira-Orange, mTangerine, tdTomato, or any other suitable fluorescent protein. In another embodiment, the marker domain may be a purification tag and / or an epitope tag.Appropriate tags include, but are not limited to, poly(His) tags, FLAG (or DDK) tags, Halo tags, AcV5 tags, AU1 tags, AU5 tags, biotin carboxyl carrier protein (BCCP), calmodulin-binding protein (CBP), chitin-binding domain (CBD), E tags, E2 tags, ECS tags, eXact tags, Glu-Glu tags, glutathione-S-transferase (GST), HA tags, HSV tags, KT3 tags, maltose-binding protein (MBP), MAP tags, Myc tags, NE tags, NusA tags, PDZ tags, S tags, S1 tags, SBP tags, Softag 1 tags, Softag 3 tags, Spot tags, Strep tags, SUMO tags, T7 tags, tandem affinity purification (TAP) tags, thioredoxin (TRX), V5 tags, VSV-G tags, and Xa tags. The marker domain can be located at the N-terminus, C-terminus, or within the zinc finger nuclease.

[0037] At least one nuclear localization signal, at least one cell membrane permeable domain, and / or at least one marker domain may be directly linked to a zinc finger nuclease via one or more chemical bonds (e.g., covalent bonds). Alternatively, at least one nuclear localization signal, at least one cell membrane permeable domain, and / or at least one marker domain may be indirectly linked to a zinc finger nuclease via one or more linkers. Suitable linkers include amino acids, peptides, nucleotides, nucleic acids, organic linker molecules (e.g., maleimide derivatives, N-ethoxybenzylimidazole, biphenyl-3,4',5-tricarboxylic acid, p-aminobenzyloxycarbonyl, etc.), disulfide linkers, and polymer linkers (e.g., PEG). The linker may include, but is not limited to, alkylene, alkenylene, alkynylene, alkyl, alkenyl, alkynyl, alkoxy, aryl, heteroaryl, aralkyl, aralkenyl, aralquinyl, etc. Linkers can be neutral or have a positive or negative charge. In addition, linkers can be cleavable such that the covalent bonds of the linker linking it to another chemical group can be broken or cleaved under specific conditions, including pH, temperature, salt concentration, light, catalysts, or enzymes. In some embodiments, linkers can be peptide linkers. Peptide linkers can be mobile amino acid linkers or rigid amino acid linkers. Further examples of suitable linkers are well known in the art, and programs for designing linkers are readily available (Crasto et al., Protein Eng., 2000, 13(5):309-312).

[0038] (ii) CRISPR ribonucleoprotein (RNP) In another embodiment, the targeted endonuclease may be a clustered, regularly spaced short-chain palindromic repeat (CRISPR) nuclease. CRISPR nucleases are RNA-induced nucleases derived from the bacterial or archaeal CRISPR / CRIPSR-associated (Cas) system. The CRISPR RNP system includes CRISPR nucleases and guide RNA.

[0039] Nuclease CRISPR nucleases can originate from type I (i.e., IA, IB, IC, ID, IE, or IF), type II (i.e., IIA, IIB, or IIC), type III (i.e., IIIA or IIIB), type V, or type VI CRISPR systems found in various bacteria and archaea. For example, CRISPR nucleases are found in species of the Streptococcus genus (e.g., S. pyogenes, S. thermophilus, S. pasteurianus), Campylobacter genus (e.g., Campylobacter jejuni), Francisella genus (e.g., Francisella novicida), Acaryochloris genus, Acetohalobium genus, Acidaminococcus genus, and Acidithiobacillus genus. sp.), Alicyclobacillus sp., Allochromatium sp., Ammonifex sp., Anabaena sp., Arthrospira sp., Bacillus sp., Burkholderiales sp., Caldicelulosiruptor sp., Candidatus sp., Clostridium sp., Crocosphaera sp., Cyanothece sp., Exiguobacterium sp., Finegoldia sp. Ktedonobacter sp., Lachnospiraceae sp.), Lactobacillus sp., Lyngbya sp., Marinobacter sp., Methanohalobium sp., Microscilla sp., Microcoleus sp., Microcystis sp., Natranaerobius sp., Neisseria sp., Nitrosococcus sp., Nocardiopsis sp., Nodularia sp., Nostoc sp., Oscillatoria sp., Polaromonas It may be derived from species of the genera Pelotomaculum, Pseudoalteromonas, Petrotoga, Prevotella, Staphylococcus, Streptomyces, Streptosporangium, Synechococcus, Thermosipho, or Verrucomicrobia. In other embodiments, CRISPR nucleases may originate from the archaeal CRISPR system, the CRISPR / CasX system, or the CRISPR / CasY system (Burstein et al., Nature, 2017, 542(7640):237-241).

[0040] In some embodiments, CRISPR nucleases may be derived from type II CRISPR nucleases. For example, type II CRISPR nucleases may be Cas9 proteins. Suitable Cas9 nucleases include Streptococcus pyogenes Cas9 (SpCas9), Francisella nobicida Cas9 (FnCas9), Staphylococcus aureus (SaCas9), Streptococcus thermophilus Cas9 (StCas9), Streptococcus pastelianus (SpaCas9), Campylobacter jejuni Cas9 (CjCas9), Neisseria meningitis Cas9 (NmCas9), or Neisseria cinerea Cas9 (NcCas9). In another embodiment, the CRISPR nuclease may be derived from a type V CRISPR nuclease such as the Cpf1 nuclease. Suitable Cpf1 nucleases include Francisella nobicida Cpf1 (FnCpf1), Asinococcus species Cpf1 (AsCpf1), or Lachnospiraceae bacterium ND2006 Cpf1 (LbCpf1). In yet another embodiment, the CRISPR nuclease may be derived from a type VI CRISPR nuclease such as Leptotrichia wadei Cas13a (LwaCas13a) or Leptotrichia shahii Cas13a (LshCas13a).

[0041] CRISPR nucleases may be wild-type CRISPR nucleases, modified CRISPR nucleases, or fragments of wild-type or modified CRISPR nucleases. CRISPR nucleases may be modified to increase nucleic acid binding affinity and / or specificity, alter enzymatic activity, and / or alter other properties of the protein. For example, the nuclease (i.e., DNase, RNase) domain of a CRISPR nuclease may be modified, deleted, or inactivated. CRISPR nucleases may be cleaved to remove domains that are not essential for nuclease function.

[0042] CRISPR nucleases contain two nuclease domains. For example, Cas9 nuclease contains an HNH domain that cleaves the complementary strand of the guide RNA and a RuvC domain that cleaves the non-complementary strand; Cpf1 nuclease contains a RuvC domain and a NUC domain; and Cas13a nuclease contains two HNEPN domains. When both nuclease domains are functional, CRISPR nucleases introduce double-strand breaks. Either nuclease domain can be inactivated by one or more mutations and / or deletions, thereby creating variants that introduce single-strand breaks in one strand of the double-stranded sequence. For example, one or more mutations in the RuvC domain of the Cas9 nuclease (e.g., D10A, D8A, E762A, and / or D986A) result in an HNH nickase that breaks the complementary strand of the guide RNA; and one or more mutations in the HNH domain of the Cas9 nuclease (e.g., H840A, H559A, N854A, N856A, and / or N863A) result in a RuvC nickase that breaks the non-complementary strand of the guide RNA. Equivalent mutations can convert the Cpf1 and Cas13a nucleases into nickases. By using two CRISPR nickases in combination (via a pair of offset guide RNAs) that target opposite strands of a chromosomal sequence, double-strand breaks can be induced in the chromosomal sequence. Double CRISPR nickase RNPs may improve target specificity and reduce off-target effects.

[0043] Additional domains CRISPR nucleases may further contain at least one nuclear localization sequence (NLS). An NLS is an amino acid sequence that facilitates the targeting of zinc finger nuclease proteins into the nucleus and the introduction of double-strand breaks at target sequences in chromosomes. Nuclear localization signals are known in the art (see, e.g., Lange et al., J. Biol. Chem., 2007, 282:5101-5105). Non-limiting examples of nuclear localization signals include PKKKRKV (SEQ ID NO: 1), PKKKRRV (SEQ ID NO: 2), KRPAATKKAGQAKKKK (SEQ ID NO: 3), YGRKKRRQRRR (SEQ ID NO: 4), RKKRRQRRR (SEQ ID NO: 5), PAAKRVKLD (SEQ ID NO: 6), RQRRNELKRSP (SEQ ID NO: 7), VSRKRPRP (SEQ ID NO: 8), PPKKARED (SEQ ID NO: 9), PQPKKKPL (SEQ ID NO: 10), SALIKKKKKKMAP (SEQ ID NO: 11), This includes PKQKKRK (SEQ ID NO: 12), RKLKKKIKKL (SEQ ID NO: 13), REKKKFLKRR (SEQ ID NO: 14), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 15), RKCLQAGMNLEARKTKK (SEQ ID NO: 16), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 17), and RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 18). NLS can be located at the N-terminus, C-terminus, or inside of the CRISPR nuclease.

[0044] In a further embodiment, the CRISPR nuclease may also include at least one cell membrane permeability domain. Suitable cell membrane permeability domains include, but are not limited to, GRKKRRQRRRPPQPKKKRKV (SEQ ID NO: 19), PLSSIFSRIGDPPKKKRKV (SEQ ID NO: 20), GALFLGWLGAAGSTMGAPKKKRKV (SEQ ID NO: 21), GALFLGFLGAAGSTMGAWSQPKKKRKV (SEQ ID NO: 22), KETWWETWWTEWSQPKKKRKV (SEQ ID NO: 23), YARAAARQARA (SEQ ID NO: 24), THLPRRRRR (SEQ ID NO: 25), GGRRRRR (SEQ ID NO: 26), RRQRRRTSKLMKR (SEQ ID NO: 27), GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 28), KALAWEAKLAKALAKHLAKALAKALKCEA (SEQ ID NO: 29), and RQIKIWFQNRRMKWKK (SEQ ID NO: 30). The cell membrane permeability domain can be located at the N-terminus, C-terminus, or within the CRISPR protein.

[0045] In yet another embodiment, a CRISPR nuclease may further include at least one marker domain. Non-limiting examples of marker domains include fluorescent proteins, purification tags, and epitope tags. In one embodiment, the marker domain may be a fluorescent protein. Non-limiting examples of suitable fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, EGFP, Emerald, Azami Green, Monomeric Azami). Green fluorescent proteins (e.g., CopGFP, AceGFP, ZsGreen1), yellow fluorescent proteins (e.g., YFP, EYFP, Citrine, Venus, YPet, PhiYFP, ZsYellow1), blue fluorescent proteins (e.g., EBFP, EBFP2, Azurite, mKalama1, GFPuv, Sapphire, T-sapphire), cyan fluorescent proteins (e.g., ECFP, Cerulean, CyPet, AmCyan1, Midoriishi-Cyan), red fluorescent proteins (mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRed1, AsRed2, eqFP611, mRaspberry, mStrawberry, Jred), and orange fluorescent proteins (mOrange, mKO, Kusabira-Orange, Monomeric) This includes Kusabira-Orange, mTangerine, tdTomato, or any other suitable fluorescent protein. In another embodiment, the marker domain may be a purification tag and / or an epitope tag.Appropriate tags include, but are not limited to, poly(His) tags, FLAG (or DDK) tags, Halo tags, AcV5 tags, AU1 tags, AU5 tags, biotin carboxyl carrier protein (BCCP), calmodulin-binding protein (CBP), chitin-binding domain (CBD), E tags, E2 tags, ECS tags, eXact tags, Glu-Glu tags, glutathione-S-transferase (GST), HA tags, HSV tags, KT3 tags, maltose-binding protein (MBP), MAP tags, Myc tags, NE tags, NusA tags, PDZ tags, S tags, S1 tags, SBP tags, Softag 1 tags, Softag 3 tags, Spot tags, Strep tags, SUMO tags, T7 tags, tandem affinity purification (TAP) tags, thioredoxin (TRX), V5 tags, VSV-G tags, and Xa tags. The marker domain can be located at the N-terminus, C-terminus, or within the CRISPR nuclease.

[0046] At least one nuclear localization signal, at least one cell membrane permeable domain, and / or at least one marker domain may be directly linked to a CRISPR nuclease via one or more chemical bonds (e.g., covalent bonds). Alternatively, at least one nuclear localization signal, at least one cell membrane permeable domain, and / or at least one marker domain may be indirectly linked to a CRISPR nuclease via one or more linkers. Suitable linkers include amino acids, peptides, nucleotides, nucleic acids, organic linker molecules (e.g., maleimide derivatives, N-ethoxybenzylimidazole, biphenyl-3,4',5-tricarboxylic acid, p-aminobenzyloxycarbonyl, etc.), disulfide linkers, and polymer linkers (e.g., PEG). The linker may include, but is not limited to, alkylene, alkenylene, alkynylene, alkyl, alkenyl, alkynyl, alkoxy, aryl, heteroaryl, aralkyl, aralkenyl, aralquinyl, etc. Linkers can be neutral or have a positive or negative charge. In addition, linkers can be cleavable such that the covalent bonds of the linker linking it to another chemical group can be broken or cleaved under specific conditions, including pH, temperature, salt concentration, light, catalysts, or enzymes. In some embodiments, the linker can be a peptide linker. Peptide linkers can be mobile amino acid linkers or rigid amino acid linkers. Further examples of suitable linkers are well known in the art, and programs for designing linkers are readily available in the art.

[0047] Guide RNA CRISPR nucleases are guided to their target site by guide RNA. The guide RNA hybridizes with the target site and interacts with the CRISPR nuclease, directing it to the target site in the chromosomal sequence. The target site has protospacer flanking motifs at its sequence boundary. p rotospacer a djacent mOther than the presence of a PAM, there are no sequence restrictions. CRISPR proteins from different bacterial species recognize different PAM sequences. For example, PAM sequences include 5'-NGG (SpCas9, FnCAs9), 5'-NGRRT (SaCas9), 5'-NNAGAAW (StCas9), 5'-NNNNGATT (NmCas9), 5-NNNNRYAC (CjCas9), and 5'-TTTV (Cpf1), where N is defined as any nucleotide, R as either G or A, W as either A or T, Y as either C or T, and V as A, C or G. The Cas9 PAM is located at 3' of the target site, and the cpf1 PAM is located at 5' of the target site.

[0048] A guide RNA contains three regions: a first region at the 5' end that is complementary to the target site sequence, a second internal region that forms a stem-loop structure, and a third 3' region that remains essentially single-stranded. The first region of each guide RNA differs so that each guide RNA guides the CRISPR nuclease to a specific target site. The second and third regions (also called scaffold regions) of each guide RNA may be the same for all guide RNAs.

[0049] The first region of the guide RNA is complementary to the target site sequence (i.e., the protospacer sequence) so that the first region of the guide RNA can base-pair with the target site sequence. The complementarity between the first region of the guide RNA (i.e., crRNA) and the target sequence can be at least 80%, at least 85%, at least 90%, at least 95%, or higher. Generally, there is no mismatch between the first region of the guide RNA and the target site sequence (i.e., the complementarity is perfect). In various embodiments, the first region of the guide RNA can contain from about 10 nucleotides to more than about 25 nucleotides. For example, the base-pairing region between the first region of the guide RNA and the target site in the chromosomal sequence may be about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25, or more than 25 nucleotides in length. In an exemplary embodiment, the first region of the guide RNA is approximately 19, 20, or 21 nucleotides in length.

[0050] The guide RNA also includes a second region that forms a secondary structure. In some embodiments, the secondary structure includes a stem (or hairpin) and a loop. The lengths of the loop and stem can vary. For example, the loop may range in length from about 3 to about 10 nucleotides, and the stem may range in length from about 6 to about 20 base pairs. The stem may contain one or more bulges of 1 to about 10 nucleotides. That is, the total length of the second region may range from about 16 to 60 nucleotides. In an exemplary embodiment, the loop is about 4 nucleotides long, and the stem contains about 12 base pairs.

[0051] The guide RNA also contains a third region at its 3' end, which remains essentially single-stranded. Therefore, this third region is not complementary to any chromosomal sequence within the target cell, nor to the rest of the guide RNA. The length of the third region can vary. Generally, the third region is more than approximately 4 nucleotides in length. For example, the length of the third region can range from approximately 5 to approximately 60 nucleotides.

[0052] The combined length of the second and third regions (or scaffolds) of the guide RNA can range from approximately 30 to 120 nucleotides. In some contexts, the combined length of the second and third regions of the guide RNA is in the range of approximately 70 to 100 nucleotides.

[0053] In some embodiments, the guide RNA comprises a single molecule containing all three regions. In other embodiments, the guide RNA may comprise two separate molecules. The first RNA molecule may comprise the first (5') region of the guide RNA and half of the "stem" of the second region of the guide RNA. The second RNA molecule may comprise the other half of the "stem" of the second region of the guide RNA and the third region of the guide RNA. Thus, in this embodiment, the first and second RNA molecules each contain a nucleotide sequence complementary to the other. For example, in one embodiment, the first and second RNA molecules each contain a sequence (approximately 6 to approximately 20 nucleotides) that base pairs with the other sequence to form a functional guide RNA.

[0054] (iii) Other targeted endonucleases In a further embodiment, targeted endonucleases may be meganucleases. Meganucleases are endodeoxyribonucleases characterized by a long recognition sequence, which is typically in the range of approximately 12 to 40 base pairs. Due to this requirement, the recognition sequence typically occurs only once in any genome. Among meganucleases, the family of homing endonucleases, named LAGLIDADG, has become a valuable tool for genome research and genome manipulation (see, e.g., Arnould et al., 2011, Protein Eng Des Sel, 24(1-2):27-31). Other suitable meganucleases include I-CreI and I-Dmol. Meganucleases can be targeted to specific chromosomal sequences by modifying their recognition sequences using techniques well known to those skilled in the art.

[0055] In a further embodiment, targeted endonucleases may be transcription activator-like effector (TALE) nucleases. TALE is a transcription factor of the plant pathogen Xanthomonas genus that can be readily manipulated to bind to novel DNA targets. TALE or cleaved versions thereof may be ligated to the catalytic domain of an endonuclease such as FokI to create a targeted endonuclease called a TALE nuclease or TALEN (Sanjana et al., 2012, Nat Protoc, 7(1):171-192) and Arnould et al., 2011, Protein Engineering, Design & Selection, 24(1-2):27-31).

[0056] In another embodiment, the targeted endonuclease may be a chimeric nuclease. Non-limiting examples of chimeric nucleases include ZF-meganucleases, TAL-meganucleases, Cas9-FokI fusions, ZF-Cas9 fusions, TAL-Cas9 fusions, and the like. Those skilled in the art are familiar with means for producing such chimeric nuclease fusions.

[0057] In yet another embodiment, the targeted endonuclease may be a site-specific endonuclease. In particular, the site-specific endonuclease may be a "rare-cutter" endonuclease whose recognition sequence occurs only rarely in the genome. Alternatively, the site-specific endonuclease may be manipulated to cleave a site of interest (Friedhoff et al., 2007, Methods Mol Biol 352:1110123). Generally, the recognition sequence of a site-specific endonuclease occurs only once in the genome. In yet another embodiment, the targeted endonuclease may be an artificially targeted DNA double-strand break inducer.

[0058] (b) Delivery of targeted endonucleases to cells The method involves introducing a targeted endonuclease into a parental cell line of interest. The targeted endonuclease may be introduced into the cell as a purified, isolated protein or as a nucleic acid encoding the targeted endonuclease. The nucleic acid may be DNA or RNA. In the embodiment where the encoding nucleic acid is mRNA, the mRNA may be 5' capped and / or 3' polyadenylated. In the embodiment where the encoding nucleic acid is DNA, the DNA may be linear or circular. The nucleic acid may be part of a plasmid or viral vector, where the encoding DNA may be operablely ligated to a suitable promoter. Those skilled in the art are familiar with suitable vectors, promoters, other regulatory elements, and means for introducing the vector into cells of interest. In the embodiment where the targeted endonuclease is a CRISPR nuclease, the CRISPR nuclease system may be introduced into the cell as a gRNA-protein complex.

[0059] Targeted endonuclease molecules can be introduced into cells by various means. Suitable delivery methods include microinjection, electroporation, sonoporation, microparticle gun, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendrimer transfection, heat shock transfection, nucleofection transfection, magnetofection, lipofection, impalefection, optical transfection, enhancement of nucleic acid uptake by proprietary agents, and delivery via liposomes, immunoliposomes, virosomes, or artificial virions. In certain embodiments, targeted endonuclease molecules are introduced into cells by nucleofection.

[0060] Optionally selected donor polynucleotides A method for targeted genome modification or manipulation further comprises introducing at least one donor polynucleotide into a cell, the donor polynucleotide comprising a sequence having at least one nucleotide change compared to a target chromosome sequence. The donor polynucleotide has substantial sequence identity with respect to a target site or a sequence near the target site in the chromosome sequence, such that double-strand breaks introduced by a targeted endonuclease can be repaired by homology-directed repair processes, and the sequence of the donor polynucleotide can be inserted into or replaced with a chromosome sequence, thereby modifying the chromosome sequence. For example, the donor polynucleotide may comprise a first sequence having substantial sequence identity with respect to a sequence on one side of the target site and a second sequence having substantial sequence identity with respect to a sequence on the other side of the target site. The donor polynucleotide may further comprise a donor sequence for integration into the target chromosome sequence. For example, the donor sequence may be an exogenous sequence (e.g., a marker sequence) such that integration of the exogenous sequence disrupts the leading frame and inactivates the target chromosome sequence.

[0061] The lengths of the first and second sequences of a donor polynucleotide having substantial sequence identity with respect to a target site or a sequence near it in the chromosome sequence can and will vary. Generally, each of the first and second sequences of a donor polynucleotide is at least about 10 nucleotides in length. In various embodiments, a donor polynucleotide sequence having substantial sequence identity with respect to a chromosome sequence may be about 15 nucleotides, about 20 nucleotides, about 25 nucleotides, about 30 nucleotides, about 40 nucleotides, about 50 nucleotides, about 100 nucleotides, or more than 100 nucleotides in length.

[0062] The phrase "substantial sequence identity" means that the sequence of a polynucleotide has at least about 75% sequence identity with respect to the target chromosome sequence. In some embodiments, the sequence of a polynucleotide has about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with respect to the target chromosome sequence.

[0063] The length of donor polynucleotides can and will vary. For example, donor polynucleotides can range in length from approximately 20 nucleotides to approximately 200,000 nucleotides. In various embodiments, donor polynucleotides can range in length from approximately 20 to approximately 100 nucleotides, from approximately 100 to approximately 1,000 nucleotides, from approximately 1,000 to approximately 10,000 nucleotides, from approximately 10,000 to approximately 100,000 nucleotides, or from approximately 100,000 to approximately 200,000 nucleotides.

[0064] Typically, the donor polynucleotide is DNA. DNA can be single-stranded or double-stranded. DNA can be linear or circular. In some embodiments, the donor polynucleotide may be a single-stranded, linear oligonucleotide containing less than approximately 200 nucleotides. In other embodiments, the donor polynucleotide may be part of a vector. Suitable vectors include DNA plasmids, viral vectors, bacterial artificial chromosomes (BACs), and yeast artificial chromosomes (YACs). In yet another embodiment, the donor polynucleotide may be a PCR fragment or nucleic acid compounded into a delivery medium such as a liposome or poloxamer.

[0065] Donor polynucleotides can be introduced into cells simultaneously with targeted endonuclease molecules. Alternatively, donor polynucleotides and targeted endonuclease molecules can be introduced into cells sequentially. The ratio of targeted endonuclease molecules to donor polynucleotides can and will vary. Generally, the ratio of targeted endonuclease molecules to donor polynucleotides ranges from approximately 1:10 to approximately 10:1. In various embodiments, the ratio of targeted endonuclease molecules to polynucleotides can be approximately 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, or 10:1. In one embodiment, the ratio is approximately 1:1.

[0066] (c) Cell culture The method further comprises maintaining cells under suitable conditions such that double-strand breaks introduced by targeted endonucleases can be repaired by (i) a non-homologous end-joining repair process such that the chromosome sequence is modified by the deletion, insertion, and / or substitution of at least one nucleotide, or optionally by (ii) a homology-directed repair process such that the chromosome sequence is exchanged with a polynucleotide sequence so that the chromosome sequence is modified. In embodiments in which a nucleic acid encoding a targeted endonuclease is introduced into cells, the method comprises maintaining cells under suitable conditions such that the cells express the targeted endonuclease.

[0067] Generally, cells are maintained under conditions suitable for cell proliferation and / or maintenance. Appropriate cell culture conditions are well known in the art and are described, for example, in Santiago et al. (2008) PNAS 105:5809-5814; Moehle et al. (2007) PNAS 104:3055-3060; Urnov et al. (2005) Nature 435:646-651; and Lombardo et al (2007) Nat. Biotechnology 25:1298-1306. Those skilled in the art will understand that methods of cell culture are known in the art and may vary depending on the cell type. In all cases, routine optimization may be used to determine the optimal method for a particular cell type.

[0068] During this step of the process, the targeted endonuclease recognizes and binds to a target break site in the chromosome sequence, causing a double-strand break, and in the process of repairing the double-strand break, at least one nucleotide deletion, insertion, and / or substitution are introduced into the target chromosome sequence. In certain embodiments, the target chromosome sequence is inactivated.

[0069] To confirm that the target chromosome sequence has been modified, a single clone may be isolated and genotyped (via DNA sequencing and / or protein analysis). Cells containing one modified chromosome sequence may undergo one or more further rounds of targeted genome modification to modify further chromosome sequences, thereby creating double knockouts, triple knockouts, etc.

[0070] (IV) Production of recombinant proteins Another aspect of this disclosure encompasses methods for producing recombinant proteins in biological production systems. Suitable recombinant proteins are described in Section (I)(c). The method includes expressing the recombinant protein of interest in one of the engineered cell lines described in Section (I) above, and purifying the expressed recombinant protein. Means for producing or manufacturing recombinant proteins are well known in the art (see, for example, “Biopharmaceutical Production Technology”, Subramanian (ed), 2012, Wiley-VCH; ISBN: 978-3-527-33029-4).

[0071] Recombinant proteins can be purified through a process that includes a clarification step, such as filtration, and one or more chromatographic steps, such as affinity chromatography, protein A (or G) chromatography, or ion exchange (i.e., cation and / or anion) chromatography.

[0072] definition Unless otherwise defined, all technical and scientific terms used herein have the meanings generally understood by those skilled in the art in the field to which this invention pertains. The following references provide general definitions of many of the terms used herein: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). Where used herein, the following terms have the meanings associated with them unless otherwise specified.

[0073] When describing elements or preferred embodiments of this disclosure, the articles “a,” “an,” “the,” and “said” are intended to mean that there are one or more elements. The terms “comprising,” “including,” and “having” are intended to mean that they are comprehensive and that there may be additional elements beyond those listed.

[0074] As used herein, the term “endogenous sequence” means a chromosome sequence that is natural to the cell.

[0075] The term "exogenous sequence" refers to a chromosome sequence that is not natural to the cell, or a chromosome sequence that has moved to a different location on a different chromosome.

[0076] A “manipulated” or “genetically modified” cell refers to a cell whose genome has been altered or manipulated, i.e., the cell contains at least one chromosomal sequence that has been manipulated to include at least one nucleotide insertion, at least one nucleotide deletion, and / or at least one nucleotide substitution.

[0077] The terms "genome modification" and "genome editing" refer to the process by which a specific endogenous chromosome sequence is altered so that the chromosome sequence is modified. The chromosome sequence may be modified to include at least one nucleotide insertion, at least one nucleotide deletion, and / or at least one nucleotide substitution. The modified chromosome sequence is inactivated so that no product is produced. Alternatively, the chromosome sequence may be modified so that an altered product is produced.

[0078] As used herein, “gene” refers to a DNA region that codes for a gene product (including exons and introns), as well as a DNA region that regulates the production of a gene product, regardless of whether such regulatory sequences are adjacent to the coding sequence and / or the sequence being transcribed. Thus, a gene includes, but is not limited to, a promoter sequence, a terminator, translation regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, and locus regulatory regions.

[0079] The term "heterogeneous" refers to a substance that is not natural to the target cell or species.

[0080] The terms “nucleic acid” and “polynucleotide” mean polymers of deoxyribonucleotides or ribonucleotides having a linear or cyclic three-dimensional structure. In this disclosure, these terms should not be construed as limitations on the length of the polymer. These terms may include known analogues of natural nucleotides, as well as nucleotides modified in the base, sugar, and / or phosphate moieties. Generally, analogues of a particular nucleotide have the same base-pairing specificity; that is, an analogue of A base-pairs with T. Nucleic acids or polynucleotides may be linked by phosphodiester bonds, phosphothioate bonds, phosphoramidite bonds, phosphorodiamidate bonds, or combinations thereof.

[0081] The term "nucleotide" means deoxyribonucleotide or ribonucleotide. A nucleotide may be a standard nucleotide (i.e., adenosine, guanosine, cytidine, thymidine, and uridine) or a nucleotide analog. A nucleotide analog means a nucleotide having a modified purine or pyrimidine base, or a modified ribose moiety. A nucleotide analog may be a naturally occurring nucleotide (e.g., inosine) or a non-naturally occurring nucleotide. Non-limiting examples of modifications to the sugar or base moiety of a nucleotide include the addition (or removal) of acetyl, amino, carboxyl, carboxymethyl, hydroxyl, methyl, phosphoryl, and thiol groups, as well as the substitution of carbon and nitrogen atoms of a base with other atoms (e.g., 7-deazapurine). Nucleotide analogs also include dideoxynucleotides, 2'-O-methylnucleotides, roch nucleic acids (LNAs), peptide nucleic acids (PNAs), and morpholinos.

[0082] The terms "polypeptide" and "protein" refer to polymers of amino acid residues. They are used interchangeably for this purpose.

[0083] As used herein, the terms “target site” or “target sequence” refer to a portion of a chromosomal sequence to be modified or edited, and which is manipulated to be recognized and bound by a targeted endonuclease, provided that sufficient conditions for binding are present.

[0084] The terms "upstream" and "downstream" refer to relative positions in a nucleic acid sequence to a specific location. Upstream refers to the region that is 5' relative to that location (i.e., closer to the 5' end of the strand), and downstream refers to the region that is 3' relative to that location (i.e., closer to the 3' end of the strand).

[0085] Methods for determining the identity of nucleic acid and amino acid sequences are well known in the art. Typically, such methods involve determining the nucleotide sequence of a gene's mRNA and / or the amino acid sequence encoded therein, and comparing those sequences with a second nucleotide or amino acid sequence. Genomic sequences can also be determined and compared in this manner. Generally, identity means the exact correspondence between nucleotide-nucleotide or amino acid-amino acid sequences of two polynucleotides or polypeptides. Two or more sequences (polynucleotides or amino acids) can be compared by determining their percentage identity. The percentage identity of two sequences, whether nucleic acid sequences or amino acid sequences, is the number of exact matches between the two aligned sequences divided by the length of the shorter sequence and multiplied by 100. Approximate alignment of nucleic acid sequences is provided by the local homology algorithm in Smith and Waterman, Advances in Applied Mathematics 2:482-489 (1981). This algorithm can be applied to amino acid sequences using a score matrix developed by Dayhoff, Atlas of Protein Sequences and Structure, MO Dayhoff ed., 5 suppl. 3:353-358, National Biomedical Research Foundation, Washington, DC, USA, and can be normalized by Gribskov, Nucl. Acids Res. 14(6):6745-6763 (1986). An exemplary implementation of this algorithm for determining the percent identity of sequences is provided by Genetics Computer Group (Madison, Wis.) in a utility application called "BestFit". Other suitable programs for calculating the percent identity or similarity between sequences are commonly known in the art; for example, another alignment program is BLAST, which is used with default parameters.For example, BLASTN and BLASTP can be used with the following default parameters: genetic code=standard;filter=none;strand=both;cutoff=60;expect=10;Matrix=BLOSUM62;Descriptions=50 sequences;sort by=HIGH SCORE;Databases=non-redundant, GenBank+EMBL+DDBJ+PDB+GenBank CDS translations+Swiss protein+Spupdate+PIR. Details of these programs can be found on the GenBank website. With respect to the sequences described herein, the desirable range of sequence identity is approximately 80% to 100% and any integer value in between. Typically, the percentage identity between sequences is at least 70-75%, preferably 80-82%, more preferably 85-90%, even more preferably 92%, even more preferably 95%, and most preferably 98%.

[0086] Since various modifications can be made to the above cells and methods without departing from the scope of the present invention, all matters included in the above description and the following examples are intended to be construed as illustrative and not limiting. [Examples]

[0087] The following examples illustrate specific aspects of the present invention.

[0088] Example 1: Design of a selection system via alanine To develop a metabolic selection system via alanine, a CHO cell line auxotrophic for the non-essential amino acid alanine (Ala) was first developed. An exhaustive search was conducted against the Reactome and KEGG databases to identify all genes related to alanine synthesis. The endogenous glutamate-pyruvate transaminase 2 gene (GPT2) was identified as the only non-redundant gene responsible for Ala synthesis (Figure 1). To generate the Ala-auxotrophic CHO cell line, the glutamine (Gln)-auxotrophic, CHOZN (登録商標) GS - / - cell line (CHOZN (登録商標) ) was utilized. The endogenous GPT2 coding sequence (Figure 2) was revealed via whole-genome sequencing (WGS) of the CHOZN (登録商標) cell line, and as determined by next-generation sequencing (NGS) analysis, this was found to be present in two copies of the CHOZN (登録商標) genome. To disrupt the third exon of the GPT2 gene encoding the active subunit of the GPT2 protein, a gRNA was designed. The gRNA target sequence is underlined in Figure 2. CHOZN (登録商標) cells were cultured in EX-CELL (登録商標) CD CHO Fusion medium (MilliporeSigma 14365C) (Fusion+Gln) supplemented with 6 mM L-glutamine (MilliporeSigma G7513) under shaking conditions, 37 °C, 5% CO2. Cells were seeded at 0.5e6 the day before transfection, and the culture was maintained in the logarithmic growth phase. Using electroporation, 1.0e6 cells were transfected with the Cas9 RNP complex. Cas9 / gRNA cleavage activity was evaluated via the surveyor nuclease S mutation detection assay. The transfected cells were cultured in EX-CELL (登録商標)The cells were transferred to a 6-well culture flask containing 3 mL of CD CHO Fusion medium (MilliporeSigma 14365C) ​​(Fusion + Gln + Ala). The cells were incubated in a static environment at 30°C and 5% CO2 for 48 hours, and then transferred to a static environment at 37°C and 5% CO2 for another 48 hours. 96 hours after transfection, Cas9-modified CHOZN (登録商標) Cells were scaled up to T-75 flasks. Single-cell clones were isolated from the Cas9 / gRNA modification pool into 96-well culture plates via fluorescence-activated cell sorting (FACS). Single-cell clones were evaluated via NGS to identify clones that successfully disrupted the genetic makeup of both copies of GPT2 (Figure 3).

[0089] To demonstrate the effectiveness of the alanine-mediated selection mechanism both as a standalone system and as part of a dual metabolic selection system, a stable selective cell population expressing numerous molecules was created. The molecules used to validate this system included blue fluorescent protein (BFP) and Dasher green fluorescent protein (GFP). CHOZN (登録商標) GS - / - GPT2 - / - Cells were cultured in Fusion + Gln + Ala medium. The expression vector used in this study contained either glutamate-pyruvate transaminase 2 (GPT2) or glutamine synthase (GS) as a selectable marker according to the experimental design. Expression of mouse GPT2 (protein: alanine aminotransferase 2 {glutamate-pyruvate transaminase 2}; gene: GPT2; UniProtKB ID: Q8BGT5) or mouse GS (protein: glutamine synthase {glutamate-ammonia ligase}; gene: Glul; UniProtKB ID: P15105) was driven by a 5' SV40 promoter, with an SV40 polyadenylated sequence positioned at the 3' end of the gene (Figure 4).

[0090] CHOZN (登録商標) GS - / - GPT2 - / -Cells were cultured in Fusion+Gln+Ala medium at 37°C under 5% CO2 with shaking. 0.5e6 cells were seeded the day before transfection, and the culture was maintained in the logarithmic growth phase. 1.0e6 cells per condition were transfected with 12 mg of plasmid DNA using electroporation. Transfected cells were transferred to a T-25 culture flask containing Fusion+Gln+Ala, and 2 mL of medium was added 48 hours after transfection. 96 hours after transfection, the cells were pelleted, the medium was aspirated and removed, and then resuspended in 12 mL of appropriate selective medium at 4e5 viable cells / mL before being transferred to a T-75 flask. Glutamine-based selection was performed using Fusion-Gln, alanine-based selection using Fusion-Ala, and glutamine / alanine double selection using Fusion-Gln and Fusion-Ala. Cell viability and viable cell density of various selective cultures were monitored over time. To recover from selection, and to adapt to scale-up and shaking conditions, stable selective cultures are used in TPP (登録商標) TubeSpin Bioreactor Tube (TPP) (登録商標) Moved to ).

[0091] EX-CELL (登録商標) Advanced CHO Fed-batch medium (MilliporeSigma 14366C), EX-CELL (登録商標) Advanced CHO Feed (MilliporeSigma 24367C), and Cellvento (登録商標) We developed custom formulations for L-alanine (Ala) deficiency in 4-feed (MilliporeSigma 1.03796.0005) (Advanced-Ala, Feed-Ala, and 4-feed-Ala, respectively).

[0092] Example 2: EX-CELL (登録商標) We developed a custom formulation of CD CHO Fusion medium that does not contain L-alanine (Ala) (Fusion-Gln-Ala).(登録商標) GS - / - GPT2 - / - Clones were cultured in Fusion+Gln+Ala or Fusion+Gln-Ala for at least 10 days. Viability and viable cell density were measured twice a week. Figure 8 shows CHOZN (登録商標) GS - / - GPT2 - / - CHOZN cannot grow in the absence of alanine, but when alanine is added to the culture medium, (登録商標) GS - / - GPT2 - / - This indicates that cell proliferation is rescued.

[0093] Example 3: To demonstrate that a stable, selected cell population can produce the target protein using the Ala-mediated selection system described in Example 1, cells were transfected and subcultured under selective pressure in Fusion-Ala. The GPT2 / BFP vector was then processed using CHOZN (登録商標) GS - / - GPT2 - / - Cells were transfected. Cell proliferation and viability were monitored through selection. Cells transfected with the GPT2 / BFP plasmid and cultured in either Fusion+Gln+Ala (non-selective medium) or Fusion+Gln-Ala (Ala-selective medium) were restored from the selected state. The viable cell population was analyzed by FACS, and the percentage of BFP+ cells was measured. Cells grown in Fusion+Gln+Ala showed a lower percentage of BFP-positive cells compared to cells that underwent Ala-selection in Fusion+Gln-Ala (Figure 6).

[0094] Example 4: To test whether a stable cell population could produce two independent intracellular fluorescent proteins under double-selection conditions, we developed two vectors: the first containing sequences encoding GFP and GS, and the second containing sequences encoding BFP and GPT2 (Figure 4). These two plasmids were then processed by CHOZN. (登録商標) GS - / - GPT2- / - Cells were co-transfected (GFP+BFP). As a control, each vector was also used in CHOZN (登録商標) GS - / - GPT2 - / - Cells were transfected separately (GFP alone and BFP alone, respectively). Cells from all three transfections were then subcultured under GS-selective conditions (Fusion-Gln), GPT2-selective conditions (Fusion-Ala), and dual-metabolism-selective conditions (Fusion-Gln-Ala with amino acids added as needed). The conditions used for selection were also applied during recovery, scale-up, and all other assays. Cells transfected with the GFP vector could survive and grow under -Gln conditions but required the addition of Ala to the medium. On the other hand, cells transfected with the BFP vector could survive and grow under -Ala conditions but required the addition of Gln to the medium. Cells co-transfected with both vectors (GFP+BFP) could survive and grow in -Gln medium, -Ala medium, and -Gln-Ala medium (with amino acids added as needed). Cells transfected with the GFP vector that survive and proliferate in -Gln medium are GFP-positive, cells transfected with the BFP vector that survive and proliferate in -Ala medium are BFP-positive, and cells co-transfected with both vectors (GFP+BFP) that survive and proliferate in -Gln-Ala medium are positive for both GFP and BFP (Figure 7). This data suggests that the GS+GPT2 dual metabolic selection system provides a unique opportunity to select cells into which multiple independent vectors encoding intracellular proteins have been introduced, without requiring the addition of any selection substances to the culture medium, such as antibiotics.

Claims

1. A method for producing recombinant protein products, (a) To provide a mammalian cell line that has been engineered to have reduced or absent expression of endogenous glutamate-pyruvate transaminase 2 (GPT2); (b) Introducing polynucleotides into the mammalian cell line, where the polynucleotides encode a functional GPT2 gene and a recombinant protein; (c) culturing the cell line; and (d) Purifying recombinant proteins to form recombinant protein products. Methods that include...

2. The method according to claim 1, wherein the mammalian cell line of (a) further comprises reduced or absent expression of endogenous glutamine synthase (GS), ASNS, PSPH, SHMT2, DHFR, and HPRT.

3. The method according to claim 1, wherein endogenous GPT2 expression is reduced or eliminated by inactivating the endogenous GPT2 gene in a mammalian cell line.

4. The method according to claim 1, wherein the endogenous GPT2 gene is inactivated using a genome modification technique mediated by a targeted endonuclease.

5. The method according to claim 4, wherein the targeted endonuclease is a CRISPR ribonucleoprotein complex or a pair of zinc finger nucleases.

6. The method according to claim 1, wherein the mammalian cell line is a Chinese hamster ovary (CHO) cell line, a baby hamster kidney (BHK) cell line, an NS0 mouse myeloma cell line, a HEK293 cell line, or a Vero African green monkey kidney cell line.

7. The method according to claim 1, wherein the cell line is a CHO cell line.

8. The method according to claim 1, wherein the recombinant protein product is selected from antibodies, antibody fragments, vaccines, growth factors, cytokines, hormones, or coagulation factors.

9. The method according to claim 8, wherein the antibody is a bispecific antibody or a multispecific antibody.

10. A genetically engineered mammalian cell line for use in a biological production system, wherein the mammalian cell line is engineered to have reduced or absent expression of endogenous GPT2.

11. The mammalian cell line according to claim 10, wherein GPT2 expression is reduced or absent via inactivation of at least one allele of the chromosomal sequence encoding GPT2.

12. The mammalian cell line according to claim 11, wherein both alleles of the chromosomal sequence encoding GPT2 are inactivated.

13. The mammalian cell line according to claim 10, which is engineered to have reduced or absent expression of endogenous glutamine synthase (GS).

14. The mammalian cell line according to claim 12, wherein the chromosome sequence is inactivated using genome modification technology mediated by targeted endonucleases.

15. The mammalian cell line according to claim 14, wherein the targeted endonuclease is a ribonucleoprotein complex or a pair of zinc finger nucleases.

16. The mammalian cell line according to claim 15, wherein the non-human cell line is a Chinese hamster ovary (CHO) cell line, a baby hamster kidney (BHK) cell line, an NS0 mouse myeloma cell line, a HEK293 cell line, or a Vero African green monkey kidney cell line.

17. The mammalian cell line according to claim 16, wherein the cell line is a CHO cell line.

18. The mammalian cell line according to claim 17, wherein the cell viability, viable cell density, titer, proliferation rate, proliferation response, cell morphology, and / or general cell health are equivalent to those of an unmanipulated parental mammalian cell line.

19. The mammalian cell line according to claim 10, further comprising at least one nucleic acid encoding a recombinant protein selected from antibodies, antibody fragments, vaccines, growth factors, cytokines, hormones, or coagulation factors.

20. The mammalian cell line according to claim 19, wherein the antibody is a bispecific antibody or a multispecific antibody.

21. A polynucleotide comprising a nucleic acid sequence encoding functional GPT2 and at least one recombinant protein of interest.

22. a) Nucleic acid sequences encoding functional GPT2; b) Functional GS; and / or ASNs; nucleic acid sequences encoding PSPH, SHMT2, DHFR, HPRT, c) Nucleic acid sequence encoding the target recombinant protein A polynucleotide comprising a functional GPT2 gene comprising a promoter sequence, according to claim 1. d) Nucleic acid sequences encoding mutations in sequences encoding GS, ASNS, PSPH, SHMT2 and / or GPT2 that reduce the activity of one or more enzymes.