Metabolic selection through the glycine-formate biosynthetic pathway

By targeting and inactivating the SHMT2 and GS genes in mammalian cells using gene editing technology, glycine-deficient cell lines are created. Using glycine and formic acid as carbon sources, this solves the problem of introducing multiple vectors into cell lines using existing selection methods, thereby improving the production efficiency and selective expression capabilities of the bioproduction system.

JP2025536205APending Publication Date: 2025-11-05EMD MILLIPORE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025518635
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-29
Publication Date
2025-11-05

AI Technical Summary

Technical Problem

Existing technologies struggle to enable multiple selection methods in biomanufacturing to introduce multiple vectors into cell lines for the production of molecules such as bispecific antibodies and multi-chain enzymes/proteins, and there is a need to improve selective expression methods in biomanufacturing systems.

Method used

By using gene editing technologies, particularly CRISPR RNP complexes or zinc finger nucleases, the SHMT2 and GS genes in mammalian cells can be targeted and inactivated to create glycine-deficient cell lines that utilize glycine and/or formic acid as exogenous carbon sources to support cell growth and the expression of multiple selection systems.

Benefits of technology

This enabled the expression of multiple selectable systems in mammalian cell lines, improving the production efficiency of bioproduction systems, particularly the expression of bispecific antibodies and biotherapeutic proteins that require effector proteins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536205000001_ABST
    Figure 2025536205000001_ABST
Patent Text Reader

Abstract

The present invention provides isolated mammalian cells containing reduced or ablated expression of serine hydroxymethyltransferase 2 (SHMT2), as well as methods for preparing such cells and for using such cells in the production of recombinant proteins.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Application No. 63 / 377,874, filed September 30, 2022, the entire contents of which are incorporated herein by reference.

[0002] Technical Field The present invention relates to mammalian cell lines for use in biological production systems, wherein the mammalian cell line is engineered to have reduced or eliminated expression of components of the glycine-formate biosynthetic pathway to create a glycine auxotrophic cell line. [Background technology]

[0003] background The development of high-producing clonal cell lines for biomanufacturing typically utilizes one or more well-known selection methods, such as glutamine synthetase (GS for glutamine selection), dihydrofolate receptor (DHFR for hypoxanthine and thymidine selection), antibiotic selection (puromycin, hygromycin, blasticidin, etc.), or P5C synthetase (P5CS-proline selection). While the GS system has become the industry standard, there is a need for cell lines that allow for multiple selection methods, so that more than one vector can be introduced into the cell line to facilitate the production of molecules such as bispecific antibodies, multispecific antibodies, and other multi-chain enzymes / proteins, or proteins / enzymes that require effector proteins for expression. Summary of the Invention

[0004] overview Various aspects of the present invention provide mammalian cell lines for use in biological production systems, wherein the mammalian cell lines have been engineered to reduce or eliminate expression of the endogenous serine hydroxymethyltransferase 2 (SHMT2) gene. In the absence of endogenously expressed functional SHMT2 protein, the cells require an exogenous source of the amino acid glycine and / or a single carbon source, such as formate, to survive and / or grow. Chromosomal SHMT2 sequences can be targeted and inactivated by endonuclease-mediated genome modification, such as CRISPR ribonucleoprotein (RNP) complexes or zinc finger nucleases. Another aspect of the present invention provides mammalian cell lines, wherein the mammalian cell lines have been engineered to reduce or eliminate expression of the endogenous SHMT2 gene and reduce or eliminate expression of the endogenous glutamine synthetase (GS) gene.

[0005] Another aspect of the present invention involves a process for selecting cell lines with improved productivity of expressed biotherapeutic proteins. In another aspect of the present invention, a biological production system for the expression of bispecific antibodies or biotherapeutic proteins that require the expression of effector proteins is provided that utilizes multiple selection systems to more conveniently express at least one recombinant protein in any of the mammalian cell lines.

[0006] Other aspects and iterations of the invention are described in more detail below. [Brief explanation of the drawings]

[0007] [Figure 1] Serine hydroxyl methyltransferase 2 (Shmt2) converts serine to glycine in mitochondria. [Figure 2] Shmt2 cDNA sequence in CHO. The gRNA target site is underlined and the NGG PAM is shown in bold. [Figure 3]A) Genotype of Shmt2 KO clone was confirmed by NGS. Insertions or deletions and their respective frequencies are indicated in bold. B) KO alleles generated by CRISPR-Cas9 targeting. All of the above base pair modifications generate premature stop codons in the coding sequence. [Figure 4] Vectors were designed to allow selection of GFP-positive cells using a glutamine-based selection system and CFP-positive cells using a glycine-formate-based selection system, as well as development of secreted recombinant proteins through the use of two similar vectors containing the mAb heavy chain, light chain, and either Shmt2 or GS coding sequences. [Figure 5] Shmt2 KO clones 1B4, 5D5, and 8F9 cannot survive without added glycine, whereas the parental cell line GS- / - Shmt2+ / + survives with or without added glycine. [Figure 6] CHO cells with gene-disrupted GS and Shmt2 genes transfected with a plasmid containing a GS + GFP expression cassette were selected in a glutamine-deficient medium and expressed GFP, but not CFP (left). CHO cells with gene-disrupted GS and Shmt2 genes transfected with a plasmid containing a Shmt2 + CFP expression cassette were selected in a glycine-deficient medium and expressed CFP, but not GFP (center). CHO cells with gene-disrupted GS and Shmt2 genes co-transfected with Shmt2 + CFP and GS + GFP were selected in a glycine- and glutamine-deficient medium and expressed high levels of both GFP and CFP (right). [Figure 7] Endogenous Shmt2 copy number in GS- / - Shmt2+ / + cell lines by ddPCR. [Figure 8] Viability and proliferation of populations (pools) expressing GFP, CFP, or GFP and CFP using GS, Shmt2, or GS and Shmt2 (Dual) selection, respectively. [Figure 9]Addition of sodium formate to the culture medium increases the growth rate of Shmt2 KO clones in a dose-dependent manner, and the addition of 200 μM sodium formate results in a growth rate similar to that of SHMT2+ / + cells. [Figure 10] Shmt2 KO cells cannot survive without glycine, even in the presence of sodium formate. [Figure 11] The dual-expressing bulk pool showed the highest viability and viable cell density in the fed-batch assay. [Figure 12] All bulk pools showed mAb production, with the dual expression pool (GS-SO57 + Shmt2-SO57) showing the highest protein production (left panel). GS-SO57 and Shmt2-SO57 showed lower levels of protein production (right panel). DETAILED DESCRIPTION OF THE INVENTION

[0008] Detailed Description The present invention provides mammalian cell lines engineered to reduce or eliminate expression of the endogenous SHMT2 gene. Additionally, mammalian cell lines engineered to reduce or eliminate expression of the endogenous GS gene and reduce or eliminate expression of the endogenous SHMT2 gene are provided. Methods for producing the engineered cell lines and methods for selecting and using the engineered cell lines to produce recombinant proteins are also provided.

[0009] (I) Engineered cell lines One aspect of the present disclosure includes mammalian cell lines engineered to reduce or eliminate expression of the endogenous SHMT2 gene, or alternatively, mammalian cell lines engineered to reduce or eliminate expression of both the endogenous SHMT2 gene and the endogenous GS gene.

[0010] The cell lines described herein with reduced or eliminated SHMT2 expression or reduced SHMT2 and GS expression have been genetically engineered to modify the chromosomal sequence encoding the SHMT2 protein or the GS protein. The chromosomal sequence can be modified using targeted endonuclease-mediated genome editing techniques. For example, the chromosomal sequence can be modified to contain at least one nucleotide deletion, at least one nucleotide insertion, at least one nucleotide substitution, or a combination thereof, resulting in a shift in the reading frame and no protein product or a non-functional protein (i.e., the chromosomal sequence is inactivated). Inactivating one allele of the chromosomal sequence encoding SHMT2 or GS reduces protein expression (i.e., knockdown). Inactivating both alleles of the chromosomal sequence encoding SHMT2 or GS prevents protein expression (i.e., knockout).

[0011] In some embodiments, the expression level of SHMT2 can be reduced by at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or more than about 99%. In other embodiments, the expression level of SHMT2 can be reduced to an undetectable level using techniques standard in the art (e.g., Western immunoblotting assay, ELISA enzyme assay, SDS-polyacrylamide gel electrophoresis, etc.).

[0012] Generally, when glycine, formate, a carbon source, and / or an exogenous SHMT2 coding sequence are added, the cell viability, viable cell density, titer, growth rate, growth response, cell morphology, levels of apoptosis and autophagy, and / or general cell health of the engineered cell lines described herein are similar to those of the unengineered parental cells.

[0013] (a) Cell species The engineered cell lines described herein are mammalian cell lines. In some embodiments, the engineered cell lines can be derived from human cell lines. Non-limiting examples of suitable human cell lines include human embryonic kidney cells (HEK293, HEK293T); human connective tissue cells (HT-1080); human cervical carcinoma cells (HELA); human embryonic retina cells (PER.C6); human kidney cells (HKB-11); human hepatocytes (Huh-7); human lung cells (W138); human liver cells (Hep G2); human U2-OS osteosarcoma cells, human A549 lung cells, human A-431 epidermal cells, CACO-2 human colon adenocarcinoma cells, human pluripotent stem cells, Jurkat human T lymphocyte cells, or human K562 bone marrow cells. In other embodiments, the engineered cell lines can be derived from non-human cell lines. Suitable cell lines also include Chinese hamster ovary (CHO) cells; baby hamster kidney (BHK) cells; mouse myeloma NS0 cells; mouse myeloma Sp2 / 0 cells; mouse mammary C127 cells; mouse embryonic fibroblast 3T3 cells (NIH3T3); mouse B lymphoma A20 cells; mouse melanoma B16 cells; mouse myoblast C2C12 cells; mouse embryonic mesenchymal C3H-10T1 / 2 cells; mouse carcinoma CT26 cells; mouse prostate D1 cells; These include CuP cells; mouse mammary carcinoma EMT6 cells; mouse hepatoma Hepa1c1c7 cells; mouse myeloma J5582 cells; mouse epithelial MTD-1A cells; mouse cardiac MyEnd cells; mouse kidney RenCa cells; mouse pancreatic RIN-5F cells; mouse melanoma X64 cells; mouse lymphoma YAC-1 cells; rat glioblastoma 9L cells; rat B lymphoma RBL cells; rat neuroblastoma B35 cells; rat hepatocytes (HTC); buffalo rat liver BRL 3A cells; canine kidney cells (MDCK); canine mammary tumor (CMT) cells; rat osteosarcoma D17 cells; rat monocyte / macrophage DH82 cells; monkey kidney SV-40 transformed fibroblast (COS7) cells; monkey kidney CVI-76 cells; and African green monkey kidney (VERO, VERO-76) cells. A comprehensive list of mammalian cell lines can be found in the American Type Culture Collection Catalog (ATCC, Manassas, VA). In some embodiments, the cell lines described herein are other than murine cell lines. In particular embodiments, the engineered cell line is a CHO cell line.Suitable CHO cell lines include, but are not limited to, CHO-K1, CHO-K1SV, CHO GS- / -, CHO S, DG44, DuxB11, and derivatives thereof.

[0014] In various embodiments, the parent cell line can be deficient in glutamine synthetase (GS), dihydrofolate reductase (DHFR), hypoxanthine-guanine phosphoribosyltransferase (HPRT), asparagine synthetase (ASNS), phosphoserine phosphatase (PSPH), or a combination thereof. For example, chromosomal sequences encoding GS, DHFR, HPRT, ASNS, and / or PSPH can be inactivated. In certain embodiments, all chromosomal sequences encoding GS, DHFR, HPRT, ASNS, and / or PSPH are inactivated in the parent cell line.

[0015] (b) any nucleic acid encoding a recombinant protein In some embodiments, the engineered cell lines described herein may further comprise at least one nucleic acid encoding a recombinant protein. Generally, recombinant proteins are heterologous, meaning that they are not native to the cell. The recombinant protein may be a therapeutic protein selected from, but not limited to, an antibody, an antibody fragment, a monoclonal antibody, a humanized antibody, a humanized monoclonal antibody, a chimeric antibody, an IgG molecule, an IgG heavy chain, an IgG light chain, an IgA molecule, an IgD molecule, an IgE molecule, an IgM molecule, a vaccine, a growth factor, a cytokine, an interferon, an interleukin, a hormone, a clotting factor, a blood component, an enzyme, a therapeutic protein, a dietary supplement, a growth factor, a cytokine, an interferon, an interleukin, a hormone, a clotting factor, a blood component, an enzyme, a therapeutic protein, a dietary supplement protein, a functional fragment or functional variant of any of the above, or a fusion protein comprising any of the above proteins and / or functional fragments or variants thereof. In certain embodiments, the recombinant protein is a bispecific or multispecific antibody, or a protein that requires an effector protein for expression.

[0016] In some embodiments, the nucleic acid encoding the recombinant protein can be linked to a sequence encoding serine hydroxymethyltransferase 2 (SHMT2), phosphoserine phosphatase (PSPH), asparagine synthetase (ASNS), hypoxanthine-guanine phosphoribosyltransferase (HPRT), dihydrofolate reductase (DHFR), and / or glutamine synthetase (GS), and SHMT2, PSPH, ASNS, HPRT, DHFR, and / or GS can be used as a selectable marker. The nucleic acid encoding the recombinant protein can also be linked to a sequence encoding at least one antibiotic resistance gene and / or a sequence encoding a marker protein such as a fluorescent protein. In some embodiments, the nucleic acid encoding the recombinant protein can be part of an expression construct. The expression construct or vector can include additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selectable marker sequences, origins of replication, etc. Further information can be found in "Current Protocols in Molecular Biology" Ausubel et al., John Wiley & Sons, New York, 2003 or "Molecular Cloning: A Laboratory Manual" Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001.

[0017] In some embodiments, the nucleic acid encoding the recombinant protein can be extrachromosomal. That is, the nucleic acid encoding the recombinant protein can be transiently expressed from a plasmid, cosmid, engineered chromosome, minichromosome, or another extrachromosomal construct. In other embodiments, the nucleic acid encoding the recombinant protein can be chromosomally integrated into the genome of the cell. Integration can be random or targeted. Thus, the recombinant protein can be stably expressed. In some iterations of this embodiment, the nucleic acid sequence encoding the recombinant protein can be operably linked to an appropriate heterologous expression control sequence (i.e., a promoter). In other iterations, the nucleic acid sequence encoding the recombinant protein can be under the control of endogenous expression control sequences. The nucleic acid sequence encoding the recombinant protein can be integrated into the genome of the cell line using homologous recombination, targeted endonuclease-mediated genome editing, viral vectors, transposons, recombinase-mediated cassette replacement systems, plasmids, and other well-known means. Further guidance can be found in Ausubel et al. 2003, supra and Sambrook & Russell, 2001, supra.

[0018] (II) Kit A further aspect of the present invention provides kits for recombinant protein production, wherein the kits comprise any of the engineered cell lines detailed above in section (I). The kits may further comprise cell growth medium, transfection reagents, plasmid vectors, selective media, recombinant protein purification means, buffers, etc. The kits provided herein generally include instructions for propagating the cell lines and using them to produce recombinant proteins. The instructions included in the kits may be affixed to packaging material or may be included as a package insert. The instructions are typically, but are not limited to, written or printed. Any medium capable of recording such instructions and transmitting them to an end user is contemplated by the present invention. Such media include, but are not limited to, electronic storage media (magnetic disks, tapes, cartridges, chips), optical media (CD ROMs), and the like. As used herein, the term "instructions" may include the address of an internet site providing the instructions.

[0019] (III) Preparation of recombinant cell lines Yet another aspect of the present disclosure provides a method for preparing or genetically engineering a cell line with reduced or eliminated expression of SHMT2 and / or GS, as described in section (I) above. The chromosomal sequences encoding SHMT2 and / or GS can be knocked down or knocked out using various techniques. Generally, engineered cell lines are prepared using a genome modification process mediated by a targeted endonuclease. Those skilled in the art will also understand that the engineered cell line can also be prepared using a site-specific recombination system, random mutagenesis, or other methods known in the art.

[0020] Generally, engineered cell lines are prepared by a method comprising introducing at least one targeting endonuclease or a nucleic acid encoding the targeting endonuclease into a parent cell line of interest, wherein the targeting endonuclease is targeted to a chromosomal sequence encoding SHMT2 and / or GS. The targeting endonuclease recognizes and binds to a specific chromosomal sequence, introducing a double-strand break. In some embodiments, the double-strand break is repaired by the non-homologous end joining (NHEJ) repair process. Because NHEJ is error-prone, deletion, insertion, and / or substitution of at least one nucleotide occurs, thereby disrupting the reading frame of the chromosomal sequence and resulting in the production of no protein product or a non-functional protein, for example, through disruption of the enzyme active site of the protein. In other embodiments, the targeting endonuclease can also be used to modify a chromosomal sequence via a homologous recombination reaction by co-introducing a polynucleotide having substantial sequence identity with a portion of the targeted chromosomal sequence. In such situations, the double-strand break introduced by the targeted endonuclease is repaired by a homology-directed repair process in which the chromosomal sequence is exchanged with a polynucleotide in such a way that the chromosomal sequence is changed or modified (e.g., by integration of the foreign sequence).

[0021] (a) Target endonuclease Various targeting endonucleases can be used to modify the chromosomal sequence encoding SHMT2 and / or GS.Targeting endonucleases can be natural proteins or engineered proteins.Suitable targeting endonucleases include zinc finger nucleases (ZFNs), CRISPR nucleases, transcription activator-like effector (TALE) nucleases (TALENs), meganucleases, chimeric nucleases, site-specific endonucleases, and engineered targeting DNA double-strand break inducers.

[0022] (i) Zinc finger nuclease In certain embodiments, the targeting endonuclease can be a pair of zinc finger nucleases (ZFNs). ZFNs bind to specific target sequences and introduce double-strand breaks at the target cleavage site. Generally, ZFNs comprise a DNA binding domain (i.e., zinc finger) and a cleavage domain (i.e., nuclease), and each domain is described below.

[0023] DNA-binding domainDNA binding domains and zinc fingers can be designed to recognize and bind to any nucleic acid sequence. For example, Beerli et al. (2002) Nat.Biotechnol. 20:135-141; Pabo et al. (2001) Ann. Rev. Biochem. 70:313-340; Isalan et al. (2001) Nat. Biotechnol. 19:656-660; Segal et al. (2001) Curr. Opin. Biotechnol. 12:632-637; Choo et al. (2000) Curr. Opin. Struct. Biol. 10:411-416; Zhang et al. (2000) J. Biol. Chem. 275(43):33850-33860; Doyon et al. (2008) Nat. Biotechnol. 26:702-708; and Santiago et al. See, e.g., U.S. Pat. No. 6,453,242 and U.S. Pat. No. 6,534,261, the entire contents of which are incorporated herein by reference. For example, the algorithm described in U.S. Pat. No. 6,453,242 can be used to design zinc finger binding domains that target preselected sequences.Other methods, such as rational design using non-degenerate recognition code tables, can also be used to design zinc finger binding domains that target specific sequences (Sera et al. (2002) Biochemistry 41:7074-7081). Publicly available web-based tools for identifying potential target sites in DNA sequences and designing zinc finger binding domains are known in the art. For example, a tool for identifying potential target sites in DNA sequences can be found at zincfingertools.org. A tool for designing zinc finger binding domains can be found at zifit.partners.org / ZiFiT. (See Mandell et al. (2006) Nuc. Acid Res. 34: W516-W523; Sander et al. (2007) Nuc. Acid Res. 35: W599-W605).

[0024] A zinc finger binding domain can be designed to recognize and bind to a DNA sequence ranging from about 3 nucleotides to about 21 nucleotides in length. In one embodiment, a zinc finger binding domain can be designed to recognize and bind to a DNA sequence ranging from about 9 nucleotides to about 18 nucleotides in length. Generally, a zinc finger binding domain of a zinc finger nuclease used herein comprises at least three zinc finger recognition regions or zinc fingers, with each zinc finger binding to three nucleotides. In one embodiment, a zinc finger binding domain comprises four zinc finger recognition regions. In another embodiment, a zinc finger binding domain comprises five zinc finger recognition regions. In yet another embodiment, a zinc finger binding domain comprises six zinc finger recognition regions. A zinc finger binding domain can be designed to bind to any suitable target DNA sequence. See, e.g., U.S. Patent Nos. 6,607,882; 6,534,261; and 6,453,242. the entire contents of which are incorporated herein by reference.

[0025] U.S. Patent Nos. 5,789,538; 5,925,523; 6,007,988; 6,013,453; 6,410,248; 6,140,466; 6,200,759; and 6,242,568; as well as WO98 / 37186; WO98 / 53057; WO00 / 27878; WO01 / 88197; and GB2,338,237, the entire contents of each of which are incorporated herein by reference. Furthermore, enhancement of the binding specificity of zinc finger binding domains is described, for example, in WO02 / 077227, the entire contents of which are incorporated herein by reference.

[0026] Methods for designing and constructing zinc finger binding domains and fusion proteins (and polynucleotides encoding them) are known to those skilled in the art and are described in detail in, for example, U.S. Patent No. 7,888,121, the entire contents of which are incorporated herein by reference. Zinc finger recognition regions and / or multi-finger zinc finger proteins can be linked using appropriate linker sequences, including, for example, linkers of 5 amino acids or more in length. For non-limiting examples of linker sequences of 6 amino acids or more in length, see U.S. Patent Nos. 6,479,626; 6,903,185; and 7,153,949 (the entire contents of which are incorporated herein by reference). The zinc finger binding domains described herein can include a combination of appropriate linkers between the individual zinc fingers of the protein.

[0027] Cleavage domain.Zinc finger nucleases also include a cleavage domain. The cleavage domain portion of a zinc finger nuclease can be obtained from any endonuclease or exonuclease. Non-limiting examples of endonucleases from which cleavage domains are derived include, but are not limited to, restriction endonucleases and homing endonucleases. See, for example, the New England Biolabs Catalog or Belfort et al. (1997) Nucleic Acids Res. 25:3379-3388. Other enzymes that cleave DNA are known (e.g., S1 nuclease; mung bean nuclease; pancreatic DNase I; micrococcal nuclease; yeast HO endonuclease). See also Linn et al. (eds.) Nucleases, Cold Spring Harbor Laboratory Press, 1993. One or more of these enzymes (or functional fragments thereof) can be used as a source of the cleavage domain.

[0028] The cleavage domain can also be derived from an enzyme or portion thereof that requires dimerization for cleavage activity, as described above. Two zinc finger nucleases are required for cleavage, as each nuclease constitutes a monomer of the active enzyme dimer. Alternatively, a single zinc finger nuclease can contain both monomers to generate an active enzyme dimer. As used herein, an "active enzyme dimer" is an enzyme dimer that can cleave a nucleic acid molecule. The two cleavage monomers can be derived from the same endonuclease (or functional fragments thereof), or each monomer can be derived from a different endonuclease (or functional fragments thereof).

[0029] When two cleavage monomers are used to form an active enzyme dimer, the recognition sites of the two zinc fingers are preferably positioned such that binding of the two zinc fingers to their respective recognition sites allows the cleavage monomers to form an active enzyme dimer. Consequently, the proximity of the recognition sites can be about 5 to about 18 nucleotides apart. For example, the proximity can be about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 nucleotides apart. However, it is understood that any integer number of nucleotides or nucleotide pairs can be interposed between the two recognition sites (e.g., about 2 to about 50 or more nucleotide pairs). For example, the proximity of the recognition sites of zinc finger nucleases as detailed herein can be 6 nucleotides apart. Generally, the cleavage site is located between the recognition sites.

[0030] Restriction endonucleases (restriction enzymes) exist in many species and can bind to DNA (recognition sites) in a sequence-specific manner and cleave the DNA at or near the binding site. Some restriction enzymes (e.g., type IIS restriction enzymes) cleave DNA at sites distant from the recognition site, allowing the binding and cleavage domains to be separated. For example, the type IIS enzyme FokI catalyzes double-stranded cleavage of DNA, cleaving one strand 9 nucleotides from its recognition site and the other strand 13 nucleotides from its recognition site. See, e.g., U.S. Patent Nos. 5,356,802, 5,436,150, and 5,487,994, as well as Li et al. (1992) Proc. Natl. Acad. Sci. USA 89:4275-4279; Li et al. (1993) Proc. Natl. Acad. Sci. USA 90:2764-2768; Kim et al. (1994a) Proc. Natl. Acad. Sci. USA 91:883-887; Kim et al. (1994b) J. Biol. Chem. 269:31978-31982. Thus, a zinc finger nuclease can comprise a cleavage domain from at least one type IIS restriction enzyme and one or more zinc finger binding domains, which may or may not be engineered. Exemplary type IIS restriction enzymes are described, for example, in International Publication WO 07 / 014,275, the entire contents of which are incorporated herein by reference. Additional restriction enzymes also contain separable binding and cleavage domains and are also contemplated herein. See, for example, Roberts et al. (2003) Nucleic Acids Res. 31: 418-420.

[0031] An exemplary type IIS restriction enzyme whose cleavage domain can be separated from its binding domain is FokI. This particular enzyme exhibits activity as a dimer (Bitinaite et al. (1998) Proc. Natl. Acad. Sci. USA 95: 10, 570-10, 575). Therefore, for purposes of this disclosure, the portion of the FokI enzyme used in a zinc finger nuclease is considered a cleavage monomer. Thus, for targeted double-strand cleavage using a FokI cleavage domain, two zinc finger nucleases, each containing a FokI cleavage monomer, can be used to reconstitute an active enzyme dimer. Alternatively, a single polypeptide molecule containing a zinc finger binding domain and two FokI cleavage monomers can be used.

[0032] In certain embodiments, the cleavage domain contains one or more engineered cleavage monomers that minimize or prevent homodimerization. As a non-limiting example, the amino acid residues at positions 446, 447, 479, 483, 484, 486, 487, 490, 491, 496, 498, 499, 500, 531, 534, 537, and 538 of FokI are all targets that affect the dimerization of the FokI cleavage half-domain. Exemplary engineered cleavage monomers of FokI that predominantly form heterodimers include a pair of a first cleavage monomer containing mutations at amino acid residues 490 and 538 of FokI and a second cleavage monomer containing mutations at amino acid residues 486 and 499.

[0033] Thus, in one embodiment of the engineered cleavage monomer, a mutation at amino acid position 490 replaces Glu(E) with Lys(K); a mutation at amino acid residue 538 replaces Iso(I) with Lys(K); a mutation at amino acid residue 486 replaces Gln(Q) with Glu(E); and a mutation at position 499 replaces Iso(I) with Lys(K). Specifically, engineered cleavage monomers can be prepared by mutating position 490 from E to K and position 538 from I to K in one cleavage monomer to create an engineered cleavage monomer designated "E490K:I538K," and by mutating position 486 from Q to E and position 499 from I to K in another cleavage monomer to create an engineered cleavage monomer designated "Q486E:I499K." The engineered cleavage monomers described above are obligate heterodimer mutants in which aberrant cleavage is minimized or eliminated. A suitable method for preparing the truncated monomer may be, for example, by site-directed mutagenesis of the wild-type truncated monomer (FokI), as described in U.S. Pat. No. 7,888,121, the entire contents of which are incorporated herein by reference.

[0034] Additional domains.In some embodiments, the zinc finger nuclease further comprises at least one nuclear localization sequence (NLS). An NLS is an amino acid sequence that targets the zinc finger nuclease protein into the nucleus and facilitates the introduction of a double-strand break in the target sequence in the chromosome. Nuclear localization signals are known in the art (see, for example, Lange et al., J. Biol. Chem., 2007, 282:5101-5105). Non-limiting examples of nuclear localization signals include PKKKRKV (SEQ ID NO: 1), PKKKRRV (SEQ ID NO: 2), KRPAATKKAGQAKKKK (SEQ ID NO: 3), YGRKKRRQRRR (SEQ ID NO: 4), RKKRRQRRR (SEQ ID NO: 5), PAAKRVKLD (SEQ ID NO: 6), RQRRNELKRSP (SEQ ID NO: 7), VSRKRPRP (SEQ ID NO: 8), PPKKARED (SEQ ID NO: 9), PQPKKKPL (SEQ ID NO: 10), and SALIKKKKKMAP (SEQ ID NO: 11). , PKQKKRK (SEQ ID NO: 12), RKLKKKIKKL (SEQ ID NO: 13), REKKKFLKRR (SEQ ID NO: 14), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 15), RKCLQAGMNLEARKTKK (SEQ ID NO: 16), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 17), and RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 18). The NLS can be located at the N-terminus, C-terminus, or internally of the zinc finger nuclease.

[0035] In further embodiments, zinc finger nucleases may also comprise at least one cell membrane permeation domain. Examples of suitable cell permeation domains include, but are not limited to, GRKKRRQRRRPPQPKKKRKV (SEQ ID NO: 19), PLSSIFSRIGDPPKKKRKV (SEQ ID NO: 20), GALFLGWLGAAGSTMGAPKKKRKV (SEQ ID NO: 21), GALFLGFLGAAGSTMGAWSQPKKKRKV (SEQ ID NO: 22), KETWWETWWTEWSQPKKKRKV (SEQ ID NO: 23), YARAAARQARA (SEQ ID NO: 24), THRLPRRRRRR (SEQ ID NO: 25), GGRRARRRRRR (SEQ ID NO: 26), RRQRRTSKLMKR (SEQ ID NO: 27), GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 28), KALAWEAKLAKALAKALAKHLAKALAKALKCEA (SEQ ID NO: 29), and RQIKIWFQNRRMKWKK (SEQ ID NO: 30). The cell membrane permeation domain can be located at the N-terminus, C-terminus, or internally of the zinc finger nuclease.

[0036] In yet other embodiments, the zinc finger nuclease may further comprise at least one marker domain. Non-limiting examples of marker domains include fluorescent proteins, purification tags, epitope tags, and the like. In one embodiment, the marker domain may be a fluorescent protein. Non-limiting examples of suitable fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, EGFP, Emerald, Thistle Green, Monomeric Thistle Green, CopGFP, AceGFP, ZsGreen1), yellow fluorescent proteins (e.g., YFP, EYFP, Citrine, Venus, YPet, PhiYFP, ZsYellow1), blue fluorescent proteins (e.g., EBFP, EBFP2, Azurite, mKalama1, GFPuv, Sapphire, T-Sapphire), cyan fluorescent proteins (e.g., ECFP, Cerulean, CyP Examples of suitable fluorescent proteins include red fluorescent proteins (mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-monomer, HcRed-tandem, HcRed1, AsRed2, eqFP611, mRasberry, mStrawberry, Jred), and orange fluorescent proteins (mOrange, mKO, Kusabira-Orange, monomeric Kusabira-Orange, mTangerine, tdTomatom), or any other suitable fluorescent protein. In another embodiment, the marker domain can be a purification tag and / or an epitope tag.Suitable tags include, but are not limited to, poly(His) tag, FLAG (or DDK) tag, Halo tag, AcV5 tag, AU1 tag, AU5 tag, biotin carboxyl carrier protein (BCCP), calmodulin-binding protein (CBP), chitin-binding domain (CBD), E tag, E2 tag, ECS tag, eXact tag, Glu-Glu tag, glutathione-S-transferase (GST), HA tag, HSV tag, KT3 tag, maltose-binding protein (MBP), MAP tag, Myc tag, NE tag, NusA tag, PDZ tag, S tag, S1 tag, SBP tag, Softag1 tag, Softag3 tag, Spot tag, Strep tag, SUMO tag, T7 tag, tandem affinity purification (TAP) tag, thioredoxin (TRX), V5 tag, VSV-G tag, Xa tag, and the like. The marker domain can be located at the N-terminus, C-terminus, or internally of the zinc finger nuclease.

[0037] At least one nuclear localization signal, at least one cell membrane-permeable domain, and / or at least one marker domain can be directly linked to the zinc finger nuclease via one or more chemical bonds (e.g., covalent bonds). Alternatively, at least one nuclear localization signal, at least one cell membrane-permeable domain, and / or at least one marker domain can be indirectly linked to the zinc finger nuclease via one or more linkers. Suitable linkers include amino acids, peptides, nucleotides, nucleic acids, organic linker molecules (e.g., maleimide derivatives, N-ethoxybenzylimidazole, biphenyl-3,4',5-tricarboxylic acid, p-aminobenzyloxycarbonyl, etc.), disulfide linkers, and polymer linkers (e.g., PEG). The linker can contain one or more spacing groups, including, but not limited to, alkylene, alkenylene, alkynylene, alkyl, alkenyl, alkynyl, alkoxy, aryl, heteroaryl, aralkyl, aralkenyl, aralkynyl, etc. The linker may be neutral or may carry a positive or negative charge. Furthermore, the linker may be cleavable so that the covalent bond of the linker connecting the linker to another chemical group can be broken or cleaved under certain conditions, such as pH, temperature, salt concentration, light, catalyst, enzyme, etc. In some embodiments, the linker may be a peptide linker. The peptide linker may be a flexible or rigid amino acid linker. Further examples of suitable linkers are well known in the art, and programs for designing linkers are readily available (Crasto et al., Protein Eng., 2000, 13(5):309-312).

[0038] (ii) CRISPR ribonucleoproteins (RNPs) In other embodiments, targeting endonuclease can be clustered regularly interspaced short palindromic repeats (CRISPR) nuclease.CRISPR nuclease is the RNA guided nuclease derived from bacterial or archaeal CRISPR / CRISPR-associated (Cas) system.CRISPR RNP system comprises CRISPR nuclease and guide RNA.

[0039] Nuclease.CRISPR nucleases can be derived from type I (i.e., type IA, IB, IC, ID, IE, or IF), type II (i.e., type IIA, IIB, or IIC), type III (i.e., type IIIA or IIIB), type V, or type VI CRISPR systems present in various bacteria and archaea. For example, CRISPR nucleases have been found in Streptococcus (e.g., S. pyogenes, S. thermophilus, S. pasteurianus), Campylobacter (e.g., Campylobacter jejuni), Francisella (e.g., Francisella novicida), Acaryochloris spp., Acetohalobium spp., Acidaminococcus spp., Acidithiobacillus spp., Alicyclobacillus spp., Allochromatium spp., Ammonifex spp., Anabaena spp., Arthrospira spp., Bacillus spp., Burkholderiales spp., Caldicellulosylputor spp., Candidatus spp., Clostridium spp., Crocosphaera spp., Cyanothece spp., Exiguobacterium spp. The bacteria may be from the genera Finegoldia, Ctedonobacter, Lachnospirus, Lactobacillus, Lyngbaia, Marinobacter, Methanohalobium, Microscilla, Microcoleus, Microcystis, Natranerobic, Neisseria, Nitrosococcus, Nocardiopsis, Nodularia, Nostoc, Oscillatoria, Polaromonas, Pelotomaculum sp., Pseudoalteromonas, Petrotoga sp., Prevotella, Staphylococcus, Streptomyces, Streptosporangium, Synechococcus sp., Thermosipho sp., or Verrucomicrobium. In other embodiments, the CRISPR nuclease can be derived from an archaeal CRISPR system, a CRISPR / CasX system, or a CRISPR / CasY system (Burstein et al., Nature, 2017, 542(7640):237-241).

[0040] In some embodiments, the CRISPR nuclease can be derived from a type II CRISPR nuclease. For example, the type II CRISPR nuclease can be a Cas9 protein. Suitable Cas9 nucleases include Streptococcus pyogenes Cas9 (SpCas9), Francisella novicida Cas9 (FnCas9), Staphylococcus aureus (SaCas9), Streptococcus thermophilus Cas9 (StCas9), Streptococcus pasteurianus (SpaCas9), Campylobacter jejuni Cas9 (CjCas9), Neisseria meningitidis Cas9 (NmCas9), or Neisseria cinerea Cas9 (NcCas9). In other embodiments, the CRISPR nuclease can be derived from a type V CRISPR nuclease, such as Cpf1 nuclease. Suitable Cpf1 nucleases include Francisella novicida Cpf1 (FnCpf1), Acidaminococcus sp. Cpf1 (AsCpf1), or Lachnospiraceae bacterium ND2006 Cpf1 (LbCpf1). In yet another embodiment, the CRISPR nuclease can be derived from a Type VI CRISPR nuclease, such as Leptotrichia weidii Cas13a (LwaCas13a) or Leptotrichia sha'ii Cas13a (LshCas13a).

[0041] CRISPR nucleases can be wild-type CRISPR nucleases, modified CRISPR nucleases, or fragments of wild-type or modified CRISPR nucleases. CRISPR nucleases can be modified to increase nucleic acid binding affinity and / or specificity, change enzymatic activity, and / or change other properties of the protein. For example, the nuclease (i.e., DNase, RNase) domain of CRISPR nucleases can be modified, deleted, or inactivated. CRISPR nucleases can be truncated to remove domains that are not essential for nuclease function.

[0042] CRISPR nucleases contain two nuclease domains. For example, Cas9 nuclease contains an HNH domain that cleaves the complementary strand of the guide RNA and a RuvC domain that cleaves the non-complementary strand; Cpf1 nuclease contains a RuvC domain and a NUC domain; and Cas13a nuclease contains two HNEPN domains. When both nuclease domains function, CRISPR nucleases introduce double-strand breaks. Either nuclease domain can be inactivated by one or more mutations and / or deletions, thereby creating a mutant that introduces a single-strand break into one strand of a double-stranded sequence. For example, one or more mutations in the RuvC domain of Cas9 nuclease (e.g., D10A, D8A, E762A, and / or D986A) result in an HNH nickase that introduces nicks in the guide RNA-complementary strand; and one or more mutations in the HNH domain of Cas9 nuclease (e.g., H840A, H559A, N854A, N856A, and / or N863A) result in a RuvC nickase that introduces nicks in the non-complementary strand of the guide RNA. Similar mutations can convert Cpf1 nuclease and Cas13a nuclease into nickases. Two CRISPR nickases targeting opposite strands of a chromosomal sequence (via a pair of offset guide RNAs) can be used in combination to generate double-strand breaks in the chromosomal sequence. Dual CRISPR nickase RNPs can enhance target specificity and reduce off-target effects.

[0043] Additional domains.CRISPR nuclease may further comprise at least one nuclear localization sequence (NLS). NLS is an amino acid sequence that targets zinc finger nuclease protein into the nucleus and facilitates the introduction of double-strand breaks in the target sequence on the chromosome. Nuclear localization signals are known in the art (see, for example, Lange et al., J. Biol. Chem., 2007, 282:5101-5105). Non-limiting examples of nuclear localization signals include PKKKRKV (SEQ ID NO: 1), PKKKRRV (SEQ ID NO: 2), KRPAATKKAGQAKKKK (SEQ ID NO: 3), YGRKKRRQRRR (SEQ ID NO: 4), RKKRRQRRR (SEQ ID NO: 5), PAAKRVKLD (SEQ ID NO: 6), RQRRNELKRSP (SEQ ID NO: 7), VSRKRPRP (SEQ ID NO: 8), PPKKARED (SEQ ID NO: 9), PQPKKKPL (SEQ ID NO: 10), SALIKKKKKMAP (SEQ ID NO: 11). , PKQKKRK (SEQ ID NO: 12), RKLKKKIKKL (SEQ ID NO: 13), REKKKFLKRR (SEQ ID NO: 14), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 15), RKCLQAGMNLEARKTKK (SEQ ID NO: 16), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 17), and RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 18). The NLS can be located at the N-terminus, C-terminus, or internally of the CRISPR nuclease.

[0044] In further embodiments, CRISPR nuclease can also comprise at least one cell membrane permeation domain.Examples of suitable cell permeation domains include but are not limited to GRKKRRQRRRPPQPKKKRKV (SEQ ID NO: 19), PLSSIFSRIGDPPKKKRKV (SEQ ID NO: 20), GALFLGWLGAAGSTMGAPKKKRKV (SEQ ID NO: 21), GALFLGFLGAAGSTMGAWSQPKKKRKV (SEQ ID NO: 22), KETWWETWWTEWSQPKKKRKV (SEQ ID NO: 23), YARAAARQARA (SEQ ID NO: 24), THRLPRRRRRR (SEQ ID NO: 25), GGRRARRRRRR (SEQ ID NO: 26), RRQRRTSKLMKR (SEQ ID NO: 27), GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 28), KALAWEAKLAKALAKALAKHLAKALAKALKCEA (SEQ ID NO: 29) and RQIKIWFQNRRMKWKK (SEQ ID NO: 30). The cell membrane permeation domain can be located at the N-terminus, C-terminus, or internally of the CRISPR protein.

[0045] In yet other embodiments, the CRISPR nuclease may further comprise at least one marker domain. Non-limiting examples of marker domains include fluorescent proteins, purification tags, epitope tags, and the like. In one embodiment, the marker domain may be a fluorescent protein. Non-limiting examples of suitable fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, EGFP, Emerald, Thistle Green, Monomeric Thistle Green, CopGFP, AceGFP, ZsGreen1), yellow fluorescent proteins (e.g., YFP, EYFP, Citrine, Venus, YPet, PhiYFP, ZsYellow1), blue fluorescent proteins (e.g., EBFP, EBFP2, Azurite, mKalama1, GFPuv, Sapphire, T-Sapphire), cyan fluorescent proteins (e.g., ECFP, Cerulean, CyP Examples of suitable fluorescent proteins include red fluorescent proteins (mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-monomer, HcRed-tandem, HcRed1, AsRed2, eqFP611, mRasberry, mStrawberry, Jred), and orange fluorescent proteins (mOrange, mKO, Kusabira-Orange, monomeric Kusabira-Orange, mTangerine, tdTomatom), or any other suitable fluorescent protein. In another embodiment, the marker domain can be a purification tag and / or an epitope tag.Suitable tags include, but are not limited to, poly(His) tag, FLAG (or DDK) tag, Halo tag, AcV5 tag, AU1 tag, AU5 tag, biotin carboxyl carrier protein (BCCP), calmodulin-binding protein (CBP), chitin-binding domain (CBD), E tag, E2 tag, ECS tag, eXact tag, Glu-Glu tag, glutathione-S-transferase (GST), HA tag, HSV tag, KT3 tag, maltose-binding protein (MBP), MAP tag, Myc tag, NE tag, NusA tag, PDZ tag, S tag, S1 tag, SBP tag, Softag1 tag, Softag3 tag, Spot tag, Strep tag, SUMO tag, T7 tag, tandem affinity purification (TAP) tag, thioredoxin (TRX), V5 tag, VSV-G tag, Xa tag, and the like. The marker domain can be located at the N-terminus, C-terminus, or internally of the CRISPR nuclease.

[0046] At least one nuclear localization signal, at least one cell membrane permeation domain, and / or at least one marker domain can be directly linked to CRISPR nuclease via one or more chemical bonds (e.g., covalent bonds). Alternatively, at least one nuclear localization signal, at least one cell membrane permeation domain, and / or at least one marker domain can be indirectly linked to CRISPR nuclease via one or more linkers. Suitable linkers include amino acids, peptides, nucleotides, nucleic acids, organic linker molecules (e.g., maleimide derivatives, N-ethoxybenzylimidazole, biphenyl-3,4',5-tricarboxylic acid, p-aminobenzyloxycarbonyl, etc.), disulfide linkers, and polymer linkers (e.g., PEG). Linkers can include one or more spacing groups, including, but not limited to, alkylene, alkenylene, alkynylene, alkyl, alkenyl, alkynyl, alkoxy, aryl, heteroaryl, aralkyl, aralkenyl, aralkynyl, etc. A linker can be neutral or can carry a positive or negative charge. Furthermore, a linker can be cleavable such that the covalent bond of the linker connecting the linker to another chemical group can be cleaved or broken under certain conditions, such as pH, temperature, salt concentration, light, catalysts, enzymes, etc. In some embodiments, the linker can be a peptide linker. The peptide linker can be a flexible amino acid linker or a rigid amino acid linker. Further examples of suitable linkers are well known in the art, and programs for designing linkers are readily available in the art.

[0047] Guide RNA.CRISPR nucleases are guided to their target sites by guide RNAs. The guide RNA hybridizes to the target site, interacts with the CRISPR nuclease, and guides it to the target site in the chromosomal sequence. There are no sequence restrictions on the target site other than that it must be surrounded by a protospacer adjacent motif (PAM). CRISPR proteins in different bacterial species recognize different PAM sequences. For example, PAM sequences include 5'-NGG (SpCas9, FnCas9), 5'-NGRRT (SaCas9), 5'-NNAGAAW (StCas9), 5'-NNNNGATT (NmCas9), 5-NNNNRYAC (CjCas9), and 5'-TTTV (Cpf1), where N is defined as any nucleotide, R is defined as either G or A, W is defined as either A or T, Y is defined as either C or T, and V is defined as A, C, or G. The Cas9 PAM is located 3' to the target site and the cpf1 PAM is located 5' to the target site.

[0048] Guide RNAs contain three regions: a first region at the 5' end that is complementary to the sequence of the target site, a second internal region that forms a stem-loop structure, and a third 3' region that remains essentially single-stranded. The first region of each guide RNA is different, and each guide RNA guides the CRISPR nuclease to a specific target site. The second and third regions (also called scaffold regions) of each guide RNA can be the same for all guide RNAs.

[0049] The first region of the guide RNA is complementary to the sequence of the target site (i.e., the protospacer sequence), and the first region of the guide RNA can base-pair with the sequence of the target site. The complementarity between the first region of the guide RNA (i.e., the crRNA) and the target sequence can be at least 80%, at least 85%, at least 90%, at least 95%, or more. Generally, there is no mismatch between the sequence of the first region of the guide RNA and the sequence of the target site (i.e., complete complementarity). In various embodiments, the first region of the guide RNA can comprise from about 10 nucleotides to about 25 nucleotides or more. For example, the base-paired region between the first region of the guide RNA and the target site in the chromosomal sequence can be about 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides, 22 nucleotides, 23 nucleotides, 24 nucleotides, 25 nucleotides, or more than 25 nucleotides in length. In exemplary embodiments, the first region of the guide RNA is about 19, 20, or 21 nucleotides in length.

[0050] The guide RNA also comprises a second region that forms a secondary structure. In some embodiments, the secondary structure comprises a stem (or hairpin) and a loop. The lengths of the loop and stem vary. For example, the length of the loop can range from about 3 nucleotides to about 10 nucleotides, and the length of the stem can range from about 6 base pairs to about 20 base pairs. The stem can include one or more bulges of 1 nucleotide to about 10 nucleotides. Thus, the total length of the second region can range from about 16 nucleotides to about 60 nucleotides. In an exemplary embodiment, the length of the loop is about 4 nucleotides, and the stem comprises about 12 base pairs.

[0051] The guide RNA also comprises a third region at the 3' end, which remains substantially single-stranded.Therefore, the third region is not complementary to any chromosomal sequence in the target cell, and is not complementary to the rest of the guide RNA.The length of the third region can vary.Generally, the length of the third region is about 4 nucleotides or more.For example, the length of the third region can range from about 5 nucleotides to about 60 nucleotides.

[0052] The total length of the second and third regions (or scaffold) of the guide RNA can range from about 30 nucleotides to about 120 nucleotides. In one aspect, the total length of the second and third regions of the guide RNA ranges from about 70 nucleotides to about 100 nucleotides.

[0053] In some embodiments, the guide RNA comprises a single molecule containing all three regions. In other embodiments, the guide RNA comprises two separate molecules. The first RNA molecule can comprise the first (5') region of the guide RNA and half of the "stem" of the second region of the guide RNA. The second RNA molecule can comprise the second region of the guide RNA and the remaining half of the "stem" of the third region of the guide RNA. Thus, in this embodiment, the first and second RNA molecules each comprise a sequence of nucleotides that is complementary to each other. For example, in one embodiment, the first and second RNA molecules each comprise a sequence (about 6 to about 20 nucleotides) that base pairs with the other sequence to form a functional guide RNA.

[0054] (iii) Other target endonucleases In a further embodiment, the targeting endonuclease can be a meganuclease. Meganucleases are endodeoxyribonucleases characterized by long recognition sequences, i.e., the recognition sequences generally range from about 12 base pairs to about 40 base pairs. As a result of this requirement, the recognition sequence generally occurs only once in any genome. Among meganucleases, a family of homing endonucleases called LAGLIDADG has become a valuable tool for the study of genomes and genome engineering (see, for example, Arnould et al., 2011, Protein Eng Des Sel, 24(1-2):27-31). Other suitable meganucleases include I-CreI and I-Dmol. Meganucleases can be targeted to specific chromosomal sequences by modifying their recognition sequences using techniques well known to those skilled in the art.

[0055] In a further embodiment, the targeting endonuclease can be a transcription activator-like effector (TALE) nuclease. TALEs are transcription factors derived from the plant pathogen Xanthomonas that can be easily engineered to bind to new DNA targets. TALEs, or truncated forms thereof, can be linked to the catalytic domain of an endonuclease, such as FokI, to create targeting endonucleases called TALE nucleases or TALENs (Sanjana et al., 2012, Nat Protoc, 7(1):171-192, and Arnould et al., 2011, Protein Engineering, Design & Selection, 24(1-2):27-31).

[0056] In another embodiment, the targeting endonuclease can be a chimeric nuclease. Non-limiting examples of chimeric nucleases include ZF-meganucleases, TAL-meganucleases, Cas9-FokI fusions, ZF-Cas9 fusions, TAL-Cas9 fusions, etc. Those skilled in the art are familiar with means for creating such chimeric nuclease fusions.

[0057] In yet another embodiment, the targeting endonuclease can be a site-specific endonuclease. In particular, the site-specific endonuclease can be a "rare-cutter" endonuclease whose recognition sequence rarely occurs in the genome. Alternatively, the site-specific endonuclease can be engineered to cleave a desired site (Friedhoff et al., 2007, Methods Mol Biol 352:1110123). Generally, the recognition sequence of a site-specific endonuclease occurs only once in the genome. In another further embodiment, the targeting endonuclease can be an artificial targeted DNA double-strand break inducer.

[0058] (b) Delivery of the targeted endonuclease to cells This method involves introducing a targeting endonuclease into a parent cell line of interest. The targeting endonuclease can be introduced into the cell as a purified, isolated protein or as a nucleic acid encoding the targeting endonuclease. The nucleic acid can be DNA or RNA. In embodiments where the encoding nucleic acid is mRNA, the mRNA can be 5'-capped and / or 3'-polyadenylated. In embodiments where the encoding nucleic acid is DNA, the DNA can be linear or circular. The nucleic acid can be part of a plasmid or viral vector, where the encoding DNA can be operably linked to a suitable promoter. Those skilled in the art are familiar with suitable vectors, promoters, other control elements, and means for introducing the vector into the cell of interest. In embodiments where the targeting endonuclease is a CRISPR nuclease, the CRISPR nuclease system can be introduced into the cell as a gRNA-protein complex.

[0059] The targeting endonuclease molecule can be introduced into cells by various methods. Suitable delivery methods include microinjection, electroporation, sonoporation, biolistics, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendrimer transfection, heat shock transfection, nucleofection transfection, magnetofection, lipofection, imparefection, optical transfection, nucleic acid uptake enhanced by a unique agent, and delivery via liposome, immunoliposome, virosome, or engineered virion. In a specific embodiment, the targeting endonuclease molecule is introduced into cells by nucleofection.

[0060] Any donor polynucleotide. The method for targeted genome modification or engineering may further comprise introducing into cells at least one donor polynucleotide comprising a sequence with at least one nucleotide change relative to the target chromosomal sequence. The donor polynucleotide has substantial sequence identity with the sequence at or near the target site of the chromosomal sequence, so that the double-strand break introduced by the targeting endonuclease can be repaired by a homology-directed repair process, and the sequence of the donor polynucleotide can be inserted or exchanged into the chromosomal sequence, thereby modifying the chromosomal sequence. For example, the donor polynucleotide may comprise a first sequence that has substantial sequence identity with the sequence on one side of the target site and a second sequence that has substantial sequence identity with the sequence on the other side of the target site. The donor polynucleotide may further comprise a donor sequence for integration into the target chromosomal sequence. For example, the donor sequence may be a foreign sequence (e.g., a marker sequence), and the integration of the foreign sequence disrupts the reading frame and inactivates the targeted chromosomal sequence.

[0061] The length of the first and second sequences in donor polynucleotides that have substantial sequence identity with the sequence at or near the target site in chromosomal sequence can be varied and can vary.Generally, each of the first and second sequences in donor polynucleotides is at least about 10 nucleotides in length.In various embodiments, the donor polynucleotide sequence that has substantial sequence identity with chromosomal sequence can be about 15 nucleotides, about 20 nucleotides, about 25 nucleotides, about 30 nucleotides, about 40 nucleotides, about 50 nucleotides, about 100 nucleotides, or more than 100 nucleotides in length.

[0062] The phrase "substantial sequence identity" means that the sequence in the polynucleotide has at least about 75% sequence identity with the chromosomal sequence of interest. In some embodiments, the sequence in the polynucleotide has about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the chromosomal sequence of interest.

[0063] The length of the donor polynucleotide can and may vary. For example, the donor polynucleotide can range from about 20 nucleotides to about 200,000 nucleotides in length. In various embodiments, the donor polynucleotide can range from about 20 nucleotides to about 100 nucleotides in length, from about 100 nucleotides to about 1,000 nucleotides in length, from about 1,000 nucleotides to about 10,000 nucleotides in length, from about 10,000 nucleotides to about 100,000 nucleotides in length, or from about 100,000 nucleotides to about 200,000 nucleotides in length.

[0064] Typically, the donor polynucleotide is DNA. The DNA can be single-stranded or double-stranded. The DNA can be linear or circular. In some embodiments, the donor polynucleotide can be a single-stranded linear oligonucleotide containing less than about 200 nucleotides. In other embodiments, the donor polynucleotide can be part of a vector. Suitable vectors include DNA plasmids, viral vectors, bacterial engineered chromosomes (BACs), and yeast engineered chromosomes (YACs). In still other embodiments, the donor polynucleotide can be a PCR fragment or a nucleic acid complexed with a delivery vehicle such as a liposome or poloxamer.

[0065] The donor polynucleotide can be introduced into the cell simultaneously with the targeting endonuclease molecule. Alternatively, the donor polynucleotide and the targeting endonuclease molecule can be introduced into the cell sequentially. The ratio of the targeting endonuclease molecule to the donor polynucleotide can vary. Generally, the ratio of the targeting endonuclease molecule to the donor polynucleotide ranges from about 1:10 to about 10:1. In various embodiments, the ratio of the targeting endonuclease molecule(s) to the polynucleotide(s) can be about 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, or 10:1. In one embodiment, the ratio is about 1:1.

[0066] (c) Cell culture The method further includes maintaining the cells under appropriate conditions so that the double-stranded break introduced by the targeting endonuclease can be repaired by (i) a non-homologous end joining repair process, in which the chromosomal sequence is modified by deletion, insertion, and / or substitution of at least one nucleotide, or, optionally, (ii) a homology-directed repair process, in which the chromosomal sequence is modified and replaced with the sequence of the polynucleotide. In embodiments in which a nucleic acid encoding the targeting endonuclease is introduced into the cells, the method includes maintaining the cells under appropriate conditions so that the cells express the targeting endonuclease.

[0067] Generally, cells are maintained under conditions suitable for cell growth and / or maintenance.Suitable cell culture conditions are well known in the art and are described, for example, in Santiago et al. (2008) PNAS 105:5809-5814; Moehle et al. (2007) PNAS 104:3055-3060; Urnov et al. (2005) Nature 435:646-651; and Lombardo et al. (2007) Nat. Biotechnology 25:1298-1306.Those skilled in the art will understand that cell culture methods are known in the art and may vary depending on the type of cell.In any case, routine optimization can be used to determine the optimal technique for a particular cell type.

[0068] During this step of the process, the targeting endonuclease recognizes, binds to, and forms a double-stranded break at the targeted cleavage site in the chromosomal sequence, and during repair of the double-stranded break, a deletion, insertion, and / or substitution of at least one nucleotide is introduced into the targeted chromosomal sequence. In certain embodiments, the targeted chromosomal sequence is inactivated.

[0069] Once the desired chromosomal sequence has been confirmed to be modified, single cell clones can be isolated and genotyped (by DNA sequencing and / or protein analysis). Cells with one modified chromosomal sequence can undergo one or more further targeted genome modifications to modify additional chromosomal sequences, creating double knockouts, triple knockouts, etc.

[0070] (IV) Production of recombinant proteins Another aspect of the present disclosure includes a method for producing a recombinant protein in a biological production system. Suitable recombinant proteins are described in Section (I)(c). The method includes expressing a recombinant protein of interest in any of the engineered cell lines described in Section (I) above and purifying the expressed recombinant protein. Means for producing or manufacturing recombinant proteins are well known in the art (see, e.g., "Biopharmaceutical Production Technology," Subramanian (ed), 2012, Wiley-VCH; ISBN: 978-3-527-33029-4).

[0071] The recombinant protein can be purified through a clarification step, e.g., filtration, and one or more chromatography steps, e.g., affinity chromatography, protein A (or G) chromatography, ion exchange (i.e., cation and / or anion) chromatography.

[0072] definition Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by those skilled in the art to which this invention belongs. The following references provide those skilled in the art with the general definitions of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed., 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The HarperCollins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them unless otherwise specified.

[0073] When introducing elements of the disclosure or preferred embodiment(s) thereof, the articles "a," "an," "the," and "said" are intended to mean that there are one or more elements. The terms "comprising," "including," and "having" are intended to be inclusive and mean that there may be additional elements other than the listed elements.

[0074] As used herein, the term "endogenous sequence" refers to a chromosomal sequence that is endogenous to a cell.

[0075] The term "exogenous sequence" relates to a chromosomal sequence that is not native to the cell or that has been moved to a different chromosomal location.

[0076] An "engineered" or "genetically modified" cell refers to a cell whose genome has been modified or manipulated, i.e., a cell that contains at least one chromosomal sequence that has been manipulated to contain an insertion of at least one nucleotide, a deletion of at least one nucleotide, and / or a substitution of at least one nucleotide.

[0077] The terms "genome modification" and "genome editing" refer to a process in which a specific endogenous chromosomal sequence is modified. The chromosomal sequence may be modified to include at least one nucleotide insertion, at least one nucleotide deletion, and / or at least one nucleotide substitution. The modified chromosomal sequence is inactivated so that no product is produced. Alternatively, the chromosomal sequence can be modified to produce a modified product.

[0078] As used herein, "gene" refers to a DNA region (including exons and introns) that encodes a gene product, as well as all DNA regions that regulate the production of that gene product, whether or not such regulatory sequences are adjacent to the coding and / or transcribed sequence. Thus, a gene includes, but is not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, locus control regions, and the like.

[0079] The term "xenogeneic" refers to material that is not native to the cell or species of interest.

[0080] The terms "nucleic acid" and "polynucleotide" refer to linear or cyclic deoxyribonucleotide or ribonucleotide polymers. For purposes of this disclosure, these terms are not intended to be limiting with respect to the length of the polymer. The terms can encompass known analogs of natural nucleotides as well as nucleotides modified at the base, sugar, and / or phosphate moieties. Generally, an analog of a particular nucleotide has the same base-pairing specificity; i.e., an analog of A can base pair with T. The nucleotides of a nucleic acid or polynucleotide may be linked by phosphodiester, phosphothioate, phosphoramidite, or phosphorodiamidite linkages, or combinations thereof.

[0081] The term "nucleotide" refers to a deoxyribonucleotide or ribonucleotide. A nucleotide may be a standard nucleotide (i.e., adenosine, guanosine, cytidine, thymidine, uridine) or a nucleotide analog. A nucleotide analog refers to a nucleotide having a modified purine or pyrimidine base or a modified ribose moiety. A nucleotide analog may be a natural nucleotide (e.g., inosine) or a non-naturally occurring nucleotide. Non-limiting examples of modifications to the sugar or base moiety of a nucleotide include the addition (or removal) of acetyl, amino, carboxyl, carboxymethyl, hydroxyl, methyl, phosphoryl, and thiol groups, and the substitution of carbon and nitrogen atoms of the base with other atoms (e.g., 7-deazapurines). Nucleotide analogs also include dideoxynucleotides, 2'-O-methyl nucleotides, locked nucleic acids (LNAs), peptide nucleic acids (PNAs), and morpholinos.

[0082] The terms "polypeptide" and "protein" are used interchangeably to refer to a polymer of amino acid residues.

[0083] As used herein, the term "target site" or "target sequence" refers to a nucleic acid sequence that defines a portion of a chromosomal sequence to be modified or edited and that is engineered for a targeting endonuclease to recognize and bind to, provided sufficient conditions for binding exist.

[0084] The terms "upstream" and "downstream" refer to relative positions from a given position in a nucleic acid sequence. Upstream refers to the region 5' from that position (i.e., near the 5' end of the strand), and downstream refers to the region 3' from that position (i.e., near the 3' end of the strand).

[0085] Techniques for determining the identity of nucleic acid and amino acid sequences are known in the art. Generally, such techniques involve determining the nucleotide sequence of a gene's mRNA and / or the amino acid sequence encoded thereby, and comparing these sequences to a second nucleotide or amino acid sequence. Genomic sequences can also be determined and compared in this manner. Generally, identity means that two polynucleotide or polypeptide sequences match exactly, nucleotide-for-nucleotide or amino acid-for-amino acid, respectively. Two or more sequences (polynucleotide or amino acid) can be compared by determining their percent identity. The percent identity of two sequences, whether nucleic acid or amino acid, is calculated by dividing the number of exact matches between the two aligned sequences by the length of the shorter sequence and multiplying by 100. Approximate alignment of nucleic acid sequences is provided by the local homology algorithm of Smith and Waterman, Advances in Applied Mathematics 2:482-489 (1981). To apply this algorithm to amino acid sequences, a scoring matrix developed by Dayhoff, Atlas of Protein Sequences and Structure, MO Dayhoff ed., 5 suppl.3:353-358, National Biomedical Research Foundation, Washington, DC, USA, and normalized by Gribskov, Nucl. Acids Res. 14(6):6745-6763 (1986) is used. An exemplary implementation of this algorithm for determining percent sequence identity is provided by the Genetics Computer Group (Madison, Wis.) in the "BestFit" utility application. Other suitable programs for calculating percent identity or similarity between sequences are generally known in the art; for example, another alignment program is BLAST, used with default parameters.For example, BLASTN and BLASTP can be used with the following default parameters: genetic code = standard, filter = none, strand = both, cutoff = 60, expect = 10, matrix = BLOSUM62, description = 50 sequences, sort order = HIGH SCORE, database = non-redundant, GenBank + EMBL + DDBJ + PDB + GenBank CDS translations + Swiss protein + Spupdate + PIR. Details of these programs can be found on the GenBank website. For the sequences described herein, the desired range of sequence identity is about 80% to 100% and any integer value therebetween. Generally, the percent identity between sequences is at least 70-75%, preferably 80-82%, more preferably 85-90%, even more preferably 92%, even more preferably 95%, and most preferably 98%.

[0086] Since various modifications can be made to the cells and methods described above without departing from the scope of the present invention, it is intended that all matter contained in the above description and in the examples that follow be interpreted as illustrative and not in a limiting sense. [Example]

[0087] Example The following examples illustrate certain aspects of the present invention.

[0088] Example 1: Design of a glycine-formate mediated selection system To develop a glycine-formate-mediated metabolic selection system, we first developed a CHO cell line auxotrophic for the non-essential amino acid glycine (Gly). A comprehensive search was performed against the Reactome and KEGG databases to identify all genes related to the glycine-formate synthesis pathway. The endogenous serine hydroxymethyltransferase 2 gene (SHMT2) was identified as the only non-redundant gene responsible for Gly synthesis (Figure 1). To generate a Gly-auxotrophic CHO cell line, we used glutamine (Gln)-auxotrophic CHOZN. (登録商標) GS- / - Cell line (MilliporeSigma) (CHOZN (登録商標) The endogenous SHMT2 coding sequence (Figure 2) was cloned from CHOZN (登録商標) CHOZN, as elucidated by whole genome sequencing (WGS) of cell lines and determined by digital droplet PCR (ddPCR) analysis (登録商標) It was found to exist in two copies in the genome (Figure 7). CRISPR / Cas9 gene editing reagents were designed to disrupt the SHMT2 gene. The CRISPR / Cas9 target sequence is underlined in Figure 2. CHOZN (登録商標) Cells are EX-CELL (登録商標) CD CHO cells were cultured in Fusion medium (MilliporeSigma 14365C) ​​supplemented with 6 mM L-glutamine (MilliporeSigma G7513) (Fusion + Gln) at 37°C with shaking under 5% CO2. Cells were split into 0.3e6 aliquots 3 days prior to transfection. Cas9 RNP was prepared by mixing 50 pmol of Cas9 (Sigma CAS9PROT-250UG) with 150 pmol of sgRNA (Sigma) and complexing for 15 minutes at room temperature. 4e5 cells were transfected with a total of 200 pmol of complexed RNP using Lonza's 4DX Nucleofector System with program DT-133 and SF Nucleofector solution. Transfected cells were then transferred to EX-CELL 400 cells supplemented with 6 mM L-glutamine (MilliporeSigma G7513). (登録商標) The cells were transferred to a 6-well culture flask containing 2 mL of CD CHO Fusion medium (MilliporeSigma 14365C). The cells were cultured in a static environment at 37°C / 5% CO2. Forty-eight hours after transfection, 20% of the pool (400 μL) was harvested for genomic DNA extraction (gDNA). The gDNA was extracted using QuickExtract (Lucigen QE09050). 2.5 μL of the gDNA was amplified and subjected to next-generation sequencing (NGS) using an Illumina Miseq. NGS confirmed that over 90% of the pool was edited.

[0089] Single-cell clones from the Cas9-modified pool were isolated into 96-well culture plates using a fluorescence-activated cell sorter (FACS). Single-cell clones were evaluated by next-generation sequencing (NGS) to identify clones that successfully disrupted both copies of SHMT2 (Figure 3).

[0090] To demonstrate the efficacy of the glycine-formate-mediated selection mechanism both as a stand-alone system and as part of a dual metabolic selection system, we generated stable selected cell populations expressing a number of molecules, including cyan fluorescent protein (CFP), Dasher green fluorescent protein (GFP), and human IgG1CHOZN. (登録商標) GS - / - SHMT2 - / - Cells were cultured in Fusion + Gln + Gly + Formic Acid (For) medium. The expression vectors used in this study contain selectable markers for serine hydroxymethyltransferase 2 (SHMT2) or glutamine synthetase (GS), depending on the experimental design. Expression of mouse SHMT2 (protein serine hydroxymethyltransferase 2; gene: SHMT2; UniProtKB ID: Q9CZN7) or mouse GS (protein: glutamine synthetase {glutamate ammonia ligase}; gene: Glul; UniProtKB ID: p15105) was driven by a 5' SV40 promoter with an SV40 polyadenylation sequence at the 3' end of the gene (Figure 4).

[0091] CHOZN (登録商標) GS - / - SHMT2 - / -Cells were cultured in Fusion + Gln + Gly + For at 37°C with shaking under 5% CO2. Cells were split into 0.3e6 aliquots 3 days prior to transfection. 1.0e6 cells were transfected with 1µg of plasmid DNA using the Lonza 4DX Nucleofector System with program DT-133 and SF Nucleofector Solution. Transfected cells were transferred to pre-warmed EX-CELL buffer supplemented with 6mM L-glutamine (MilliporeSigma G7513) and 200µM sodium formate (Sigma 456020-25G). (登録商標) The cells were transferred to a 6-well culture flask containing 3 mL of CD CHO Fusion medium (MilliporeSigma 14365C). The cells were cultured in a static environment at 37°C / 5% CO2. Three days after transfection, each sample was expanded into a T-25 flask and 2 mL of medium was added. On day 7, the cells reached approximately 90% viability and were transferred to a 15 mL conical tube and spun down at 1000 rpm (~300 x g) for 5 minutes. The supernatant was aspirated, and the cells were washed with 5 mL of PBS and spun again for 5 minutes. The supernatant was aspirated, and the cells were suspended in 10-15 mL of selective medium and transferred to a T-75 flask. Glutamine-based selection was performed using Fusion -Gln, glycine-formate-based selection was performed using Fusion -Gly-For, and glutamine / glycine-formate double selection was performed using Fusion excluding Gln, Gly, and For. Cell viability and viable cell density of the various selection cultures were monitored over time.

[0092] EX-CELL (登録商標) Advanced CHO Fed-batch medium (MilliporeSigma 14366C) and EX-CELL (登録商標) We developed custom Gly-deficient formulations of Advanced CHO Feed (MilliporeSigma 24367C) (Advanced-Gly and Feed-Gly, respectively). (登録商標)4 Feed (MilliporeSigma 1.03796.0005) was used without modification. Stable selected cultures transfected with the IgG1 expression vector were pelleted, the selective medium was aspirated, and resuspended in Advanced Feed-Gly at 3e5 viable cells / mL for productivity analysis under fed-batch conditions. Viable cell density and viability of each culture were collected every other day starting on day 3 after seeding. Starting on day 3 after seeding, 1.5 mL of a 50 / 50 mix of Advanced Feed-Gly and 4 Feed was added to each culture. D-(+)-glucose (MilliporeSigma G8769) was added to maintain appropriate glucose levels. Productivity was monitored over time, and the titer of the fed batch was recorded every other day starting on day 8 until the culture reached 70% or less viability. Titers were measured using a ForteBio Octet interferometer and confirmed by HPLC Protein A affinity chromatography.

[0093] Example 2: Basic EX-CELL without glycine (Gly) (登録商標) A custom formulation of CD CHO Fusion medium (Fusion -Gln -Gly) was developed. (登録商標) GS - / - SHMT2 - / - Clones were cultured in Fusion + Gln + Gly + Formate or Fusion + Gln - Gly - Formate for at least 10 days. Viability and viable cell density were measured twice weekly. Figure 5 shows that in the absence of glycine and formate, CHOZN (登録商標) GS - / - SHMT2 - / - The cells cannot grow, but when the medium is supplemented with glycine, CHOZN (登録商標) GS - / - SHMT2 - / - Importantly, SHMT2 restores cell proliferation. - / - Cell proliferation rate is SHMT2 + / +Although the growth rate was significantly slower than that of the cell line, this was restored by adding sodium formate to the medium (Figure 9). Addition of formic acid (Fusion -Gly + For) alone was insufficient for cell survival (Figure 10).

[0094] Example 3: To demonstrate that a stable selected cell population can produce the protein of interest using the Gly / For-mediated selection system described in Example 1, cells were transfected and passaged under Fusion-Gly-For selection pressure. The SHMT2 / CFP vector was transfected into CHOZN (登録商標) GS - / - SHMT2 - / - Cells were transfected with SHMT2 / CFP plasmid. Cell growth and viability were monitored during the selection period. Cells transfected with the SHMT2 / CFP plasmid and cultured in Fusion + Gln -Gly -For (Gly selection medium) recovered from selection pressure. The viable cell population was analyzed by FACS to measure mean fluorescence intensity (MFI) and the percentage of CFP+ cells. Cells grown in Fusion + Gln + Gly +For had significantly fewer CFP-positive cells and a lower MFI than cells grown in Fusion + Gln -Gly -For under Gly / For selection. The results are summarized in Figures 6 and 8.

[0095] Example 4: To demonstrate that stably selected cell populations can produce proteins of interest using the GlyFor-mediated selection system described in Example 1, vectors containing IgG heavy chain, IgG light chain, and serine hydroxymethyltransferase 2 (SHMT2) coding sequences were developed. (登録商標) GS - / - SHMT2 - / -The cell line was transfected with the 500-kJ / ml Feed-Gly-For medium. A mock transfection without DNA was used as a control. This population was passaged under selective pressure into Fusion +Gln-Gly-For medium. The conditions used for selection were also applied during recovery, scale-up, and productivity assays. For the fed-batch productivity assay, cells were seeded at 3e5 viable cells / mL in Advanced +Gln-Gly-For medium. Viable cell density and viability of each culture were collected every other day starting on day 3 after seeding. Starting on day 3 after seeding, 1.5 mL of a 50 / 50 mixture of Advanced Feed-Gly-For and 4 Feed-Gly-For was added to each culture. Glucose measurements were performed every other day starting on day 5, and D-(+)-glucose (MilliporeSigma G8769) was added to maintain appropriate glucose levels. Cells transfected with the SHMT2 vector alone reached a peak viable cell density of approximately 12.5e6 cells / mL (Figure 11, left panel, middle graph) and maintained a viability of over 70% for at least 13 days (Figure 11, right panel, middle graph). The titer of the fed batch cells peaked at approximately 200 mg / L (Figure 12, left panel, middle graph and right panel, right graph).

[0096] Example 5: To investigate whether stable cell populations could produce two independent intracellular fluorescent proteins under glutamine- and glycine / formate-free conditions, we developed two vectors: one containing the coding sequences for GFP and GS, and the other containing the coding sequences for CFP and SHMT2 (Figure 4). (登録商標) GS - / - SHMT2 - / - As a control, each vector was co-transfected into CHOZN cells (GFP + CFP). (登録商標) GS - / - SHMT2 - / -Cells were also transfected separately with either GFP or CFP (GFP only or CFP only, respectively). Cells from the three transfections were then passaged under GS selection (Fusion -Gln), SHMT2 selection (Fusion -Gly -For), and dual metabolic selection (Fusion -Gln -Gly -For). The selection conditions were also applied for recovery, scale-up, and all other assays. Figure 8 shows viability data from the selection assays, demonstrating that cells transfected with the GFP vector can survive and grow under -Gln conditions but require the addition of Gly / For to the medium. In contrast, cells transfected with the CFP vector can survive and grow under -Gly conditions but require the addition of Gln to the medium. Cells cotransfected with both vectors (GFP + CFP) can survive and grow in -Gln -Gly -For medium. Figures 6 and 8 show that cells transfected with the GFP vector that survive and grow in -Gln medium are GFP positive, cells transfected with the CFP vector that survive and grow in -Gly-Formate medium are CFP positive, and cells co-transfected with both vectors (GFP + CFP) that survive and grow in -Gln-Gly-Formate medium are both GFP and CFP positive. This data demonstrates that the GS+SHMT2 dual metabolic selection system provides a unique opportunity to select cells carrying multiple independent vectors encoding intracellular proteins without the need for selective agents such as antibiotics.

[0097] Example 6: To test whether stable cell populations could produce secreted proteins under glutamine- and glycine-free conditions, two vectors expressing IgG1 were developed. One contained the coding sequences for IgG heavy chain, IgG light chain, and GS, while the other contained the coding sequences for the same IgG heavy chain, IgG light chain, and SHMT2. These two independent vectors were transfected into CHOZN. (登録商標) GS - / - SHMT2 - / - As a control, each vector was co-transfected into CHOZN cells (GS + SHMT2). (登録商標) GS- / - SHMT2 - / - Cells were also transfected separately with either GS or SHMT2 (GS only or SHMT2 only). Cells were then passaged under selection with Fusion -Gln (GS only transfected cells), Fusion -Gly -For (SHMT2 only transfected cells), or Fusion -Gln -Gly -For (GS + SHMT2 transfected cells). The selection conditions were also applied during recovery, scale-up, and productivity assays. While GS only selection cultures fully recovered after 14–19 days, SHMT2 only and GS + SHMT2 double selection cultures showed similar selection recovery profiles, requiring 19 days for full recovery. In fed-batch assays, GS only and SHMT2 only clones produced similar levels of IgG, but GS + SHMT2 cells produced significantly more IgG and had the highest proliferation and viability. This suggests that the GS and SHMT2 produced by cells through expression of exogenous GS and / or SHMT2 coding sequences are sufficient to express secreted proteins (Figure 11, right graph and Figure 12, left panel, right graph). This enables the operation of large-scale production bioreactors under dual metabolic selection conditions, which is difficult with antibiotic selection methods. This is due to the need to later separate and purify the antibiotic from the secreted protein of interest or the cost of adding antibiotics to large-scale bioreactors. Furthermore, this dual metabolic selection system (GS + SHMT2) provides an opportunity to more efficiently select cells in which multiple large vectors have been introduced, such as when expressing bispecific antibodies or other large and / or complex proteins.

Claims

1. 1. A method for producing a recombinant protein product, comprising: (a) providing a mammalian cell line engineered to reduce or eliminate expression of endogenous serine hydroxymethyltransferase 2 (SHMT2); (b) introducing a polynucleotide into a mammalian cell line, wherein the polynucleotide encodes a functional SHMT2 gene and a recombinant protein; (c) culturing the cell line; and (d) purifying the recombinant protein to form a recombinant protein product. A method comprising:

2. 2. The method of claim 1, wherein the mammalian cell line of (a) further comprises reduced or eliminated expression of endogenous glutamine synthetase (GS), phosphoserine phosphatase (PSPH), dihydrofolate reductase (DHFR), P5C synthetase (P5CS), asparaginase (ASPG), alanine transaminase (ALT), and / or asparagine synthetase (ASNS).

3. 10. The method of claim 1, wherein endogenous SHMT2 expression is reduced or eliminated by inactivating the endogenous SHMT2 gene in the mammalian cell line.

4. 2. The method of claim 1, wherein the endogenous SHMT2 gene is inactivated using targeted endonuclease-mediated genome modification techniques.

5. 5. The method of claim 4, wherein the targeting endonuclease is a CRISPR ribonucleoprotein complex or a pair of zinc finger nucleases.

6. 2. The method of claim 1, wherein the mammalian cell line is a Chinese hamster ovary (CHO) cell line, a baby hamster kidney (BHK) cell line, an NS0 mouse myeloma cell line, an HEK293 cell line, or a Vero African green monkey kidney cell line.

7. 7. The method according to any one of claims 1 to 6, wherein the cell line is a CHO cell line cultured with or without the addition of glycine and / or formic acid.

8. 8. The method of any one of claims 1 to 7, wherein the recombinant protein product is selected from an antibody, an antibody fragment, a vaccine, a growth factor, a cytokine, a hormone, or a clotting factor.

9. The method of claim 8, wherein the antibody is a bispecific or multispecific antibody.

10. A genetically engineered mammalian cell line for use in a biological production system, wherein the mammalian cell line has been engineered to reduce or eliminate expression of endogenous SHMT2.

11. 11. The mammalian cell line of claim 10, wherein expression of SHMT2 is reduced or abolished through inactivation of at least one allele of the chromosomal sequence encoding SHMT2.

12. 12. The mammalian cell line of claim 11, wherein one or more alleles of the chromosomal sequence encoding SHMT2 are inactivated.

13. 11. The mammalian cell line of claim 10, wherein the cell line has been engineered to reduce or eliminate expression of endogenous glutamine synthetase (GS), phosphoserine phosphatase (PSPH), dihydrofolate reductase (DHFR), P5C synthetase (P5CS), asparaginase (ASPG), alanine transaminase (ALT) and / or asparagine synthetase (ASNS).

14. The mammalian cell line of claim 12, wherein the chromosomal sequence has been inactivated using targeted endonuclease-mediated genome modification techniques.

15. The mammalian cell line of claim 14 , wherein the targeting endonuclease is a ribonucleoprotein complex or a pair of zinc finger nucleases.

16. 16. The mammalian cell line of claim 15, wherein the non-human cell line is a Chinese hamster ovary (CHO) cell line, a baby hamster kidney (BHK) cell line, an NS0 mouse myeloma cell line, a HEK293 cell line, or a Vero African green monkey kidney cell line.

17. 17. The mammalian cell line of claim 16, wherein the cell line is a CHO cell line cultured with or without the addition of glycine and / or formic acid.

18. 18. The mammalian cell line of claim 17, wherein the cell viability, viable cell density, titer, growth rate, growth response, cell morphology, and / or general cell health is comparable to that of the non-genetically modified parent mammalian cell line.

19. 19. The mammalian cell line of any one of claims 10 to 18, further comprising at least one nucleic acid encoding a recombinant protein selected from an antibody, an antibody fragment, a vaccine, a growth factor, a cytokine, a hormone, or a clotting factor.

20. 20. The mammalian cell line of claim 19, wherein the antibody is a bispecific or multispecific antibody.

21. A polynucleotide comprising a nucleic acid sequence encoding a functional SHMT2 and at least one recombinant protein of interest.

22. a) a nucleic acid sequence encoding a functional SHMT2; b) a nucleic acid sequence encoding a functional GS and / or ASNS; and c) nucleic acid sequences encoding mutations in the ASNS and / or SHMT2 coding sequences that attenuate the activity of either or both enzymes; d) a nucleic acid sequence encoding the recombinant protein of interest A polynucleotide comprising: