Metabolic selection via glycine-formate biosynthetic pathways
The glycine-formate metabolic selection system was established by targeting endonuclease modification, and the complexity of multiple selection methods in the prior art was solved, and glycine nutrition-deficient cell lines that efficiently express bispecific antibodies and biological therapeutic proteins were achieved.
Patent Information
- Application Number
- CN202380082202.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-30
- Filing Date
- 2023-09-29
- Publication Date
- 2025-07-08
AI Technical Summary
In existing biomanufacturing technologies, the development of high-yield cloned cell lines requires multiple selection methods to introduce more than one vector, resulting in increased complexity and difficulty in efficiently expressing proteins such as bispecific antibodies and multispecific antibodies.
By targeting endonuclease-mediated genomic modifications, the expression of endogenous serine hydroxymethyltransferase 2 (SHMT2) and glutamine synthase (GS) genes in mammalian cell lines is reduced or eliminated, and glycine-formate metabolic selection system is used to establish glycine nutritionally deficient cell lines to achieve exogenous glycine and formate dependent growth.
The efficient expression of bispecific antibodies and biological therapeutic proteins in the absence of glycine and glutamine is achieved, which simplifies the cell line transformation process and improves protein production efficiency.
Smart Images

Figure CN120283058A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit of priority of U.S. Provisional Application No. 63 / 377,874, filed on September 30, 2022, the entire content of which is incorporated herein by reference. Technical field
[0003] The present disclosure relates to mammalian cell lines for use in bioproduction systems, wherein the mammalian cell lines are engineered to reduce or eliminate the expression of components of their glycine - formate biosynthetic pathway to produce glycine auxotrophic cell lines. Background art
[0004] The development of high - yielding clonal cell lines for biomanufacturing typically utilizes one or more well - known selection methods, such as glutamine synthetase (GS for glutamine selection), dihydrofolate receptor (DHFR for hypoxanthine and thymidine selection), antibiotic selection (puromycin, hygromycin, blasticidin, etc.), or P5C synthetase (P5CS - proline selection). The GS system has become the industry standard, but cell lines that allow for multiple selection methods are needed so that more than one vector can be introduced into the cell line to facilitate the production of molecules such as bispecific antibodies, multispecific antibodies, and other multi - chain enzymes / proteins or proteins / enzymes that require effector proteins for expression. Summary of the invention
[0005] In various aspects of the present disclosure, mammalian cell lines for use in bioproduction systems are provided, wherein the mammalian cell lines are engineered to reduce or eliminate the expression of their endogenous serine hydroxymethyltransferase 2 (SHMT2) gene. In the absence of endogenously expressed functional SHMT2 protein, the cells require an exogenous source of the amino acid glycine and / or a single - carbon source, such as formate, to survive and / or grow. Genome modification mediated by targeted endonucleases, such as CRISPR ribonucleoprotein (RNP) complexes or zinc - finger nucleases, can be used to inactivate the chromosomal SHMT2 sequence. In other aspects of the present disclosure, mammalian cell lines are provided, wherein the mammalian cell lines are engineered to reduce or eliminate the expression of the endogenous SHMT2 gene and reduce or eliminate the expression of the endogenous glutamine synthetase (GS) gene.
[0006] Another aspect of the present disclosure encompasses a method for selecting cell lines with enhanced productivity of an expressed biotherapeutic protein. In other aspects of the present disclosure, bioproduction systems for expressing bispecific antibodies or biotherapeutic proteins that require the more convenient expression of effector proteins by utilizing a multiple - selection system are provided. The method includes expressing at least one recombinant protein in any of the mammalian cell lines.
[0007] Other aspects and iterations of the disclosure are described in more detail below. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 Serine hydroxymethyltransferase 2 (Shmt2) converts serine in the mitochondria to glycine.
[0009] Figure 2 The Shmt2 cDNA sequence in CHO. The gRNA target sites are underlined and the NGG PAM is in bold.
[0010] Figure 3 . A) Genotype of Shmt2 KO clones was confirmed by NGS. Insertions or deletions and their respective frequencies are in bold. B) KO alleles generated by CRISPR-Cas9 targeting. All base pair modifications above generate premature stop codons in the coding sequence.
[0011] Figure 4 The vector was designed to allow selection of GFP-positive cells using a glutamine-based selection system, and selection of CFP-positive cells using a glycine-formate-based selection system, and to develop secreted recombinant proteins via use of two similar vectors containing the mAb heavy chain, light chain, and Shmt2 or GS coding sequences.
[0012] Figure 5 Shmt2 KO clones 1B4, 5D5, and 8F9 cannot survive without glycine, while the parental cell line GS- / - Shmt2+ / + survives in the presence or absence of glycine.
[0013] Figure 6 CHO cells with their GS and Shmt2 genes genetically disrupted and transfected with a plasmid containing the GS+GFP expression cassette were selected in a medium lacking glutamine and expressed GFP but not CFP (left). CHO cells with their GS and Shmt2 genes genetically disrupted and transfected with a plasmid containing the Shmt2+CFP expression cassette were selected in a medium lacking glycine and expressed CFP but not GFP (middle). CHO cells with their GS and Shmt2 genes genetically disrupted and co-transfected with Shmt2+CFP and GS+GFP were selected in a medium lacking both glycine and glutamine and expressed high levels of both GFP and CFP (right).
[0014] Figure 7 Copy number of endogenous Shmt2 in the GS- / - Shmt2+ / + cell line via ddPCR.
[0015] Figure 8The viability and growth of pools expressing GFP, CFP, or GFP and CFP were selected using GS, Shmt2, or GS and Shmt2 (double), respectively.
[0016] Figure 9 The addition of formate in the medium increased the growth rate of Shmt2 KO clones in a dose-dependent manner. The addition of 200 uM formate resulted in a growth rate comparable to that of SHMT2 + / + cells.
[0017] Figure 10 Even in the presence of formate, Shmt2 KO cells cannot survive without glycine.
[0018] Figure 11 The double-expressing bulk pool showed the highest viability and the highest live cell density in the fed-batch assay.
[0019] Figure 12 All bulk pools showed mAb production, with the double-expressing pool (GS-SO57 + Shmt2-SO57) showing the highest protein production (left panel). GS-SO57 and Shmt2-SO57 showed lower levels of protein production (right panel). Detailed Description
[0020] The present disclosure provides mammalian cell lines engineered to reduce or eliminate the expression of the endogenous SHMT2 gene. Further provided are mammalian cell lines engineered to reduce or eliminate the expression of the endogenous GS gene and reduce or eliminate the expression of the endogenous SHMT2 gene. Methods for generating the engineered cell lines are provided, as well as methods for selecting and using the engineered cell lines to produce recombinant proteins.
[0021] (I) Engineered Cell Lines
[0022] One aspect of the present disclosure encompasses mammalian cell lines engineered to reduce or eliminate the expression of the endogenous SHMT2 gene. Alternatively, the mammalian cell line is engineered to reduce or eliminate the expression of both the endogenous SHMT2 gene and the endogenous GS gene.
[0023] The cell lines disclosed herein with reduced or eliminated expression of SHMT2, or reduced expression of SHMT2 and GS, are genetically engineered to modify the chromosomal sequences encoding the SHMT2 or GS proteins. The chromosomal sequences can be modified using targeted endonuclease-mediated genome editing techniques, which are detailed in section (III) below. For example, the chromosomal sequences can be modified to include a deletion of at least one nucleotide, an insertion of at least one nucleotide, a substitution of at least one nucleotide, or a combination thereof, such that the reading frame is shifted and no protein product is produced or a non-functional protein is produced (i.e., the chromosomal sequence is inactivated). Inactivation of one allele of the chromosomal sequence encoding SHMT2 or GS results in reduced expression of the protein (i.e., knockdown). Inactivation of both alleles of the chromosomal sequence encoding SHMT2 or GS results in no protein expression (i.e., knockout).
[0024] In some embodiments, the expression level of SHMT2 can be reduced by at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or more than about 99%. In other embodiments, the expression level of SHMT2 can be reduced to undetectable levels using standard techniques in the art (e.g., protein immunoblot assays, ELISA enzyme assays, SDS polyacrylamide gel electrophoresis, etc.).
[0025] Generally, when supplemented with glycine, formate, a carbon source, and / or an exogenous SHMT2-encoding sequence, the cell viability, live cell density, titer, growth rate, proliferative response, cell morphology, apoptosis and autophagy levels, and / or general cell health of the engineered cell lines disclosed herein are similar to those of their non-engineered parental cells.
[0026] (a) Cell type
[0027] The engineered cell lines disclosed herein are mammalian cell lines. In some embodiments, the engineered cell lines can be derived from human cell lines. Non-limiting examples of suitable human cell lines include human embryonic kidney cells (HEK293, HEK293T); human connective tissue cells (HT-1080); human cervical cancer cells (HELA); human embryonic retina cells (PER.C6); human kidney cells (HKB-11); human hepatocytes (Huh-7); human lung cells (W138); human hepatocytes (Hep G2); human U2-OS osteosarcoma cells, human A549 lung cells, human A-431 epidermal cells, CACO-2 human colorectal adenocarcinoma cells, human pluripotent stem cells, Jurkat human T lymphocytes or human K562 myeloid cells. In other embodiments, the engineered cell lines can be derived from non-human cell lines. Suitable cell lines also include Chinese hamster ovary (CHO) cells; baby hamster kidney (BHK) cells; mouse myeloma NS0 cells; mouse myeloma Sp2 / 0 cells; mouse mammary C127 cells; mouse embryonic fibroblast 3T3 cells (NIH3T3); mouse B lymphoma A20 cells; mouse melanoma B16 cells; mouse myoblast C2C12 cells; mouse embryonic mesenchyme C3H-10T1 / 2 cells; mouse carcinoma CT26 cells, mouse prostate DuCuP cells; mouse mammary EMT6 cells; mouse hepatoma Hepa1c1c7 cells; mouse myeloma J5582 cells; mouse epithelium MTD-1A cells; mouse myocardium MyEnd cells; mouse kidney RenCa cells; mouse pancreas RIN-5F cells; mouse melanoma X64 cells; mouse lymphoma YAC-1 cells; rat glioblastoma 9L cells; rat B lymphoma RBL cells; rat neuroblastoma B35 cells; rat hepatoma cells (HTC); buffalo rat liver BRL 3A cells; canine kidney cells (MDCK); canine mammary (CMT) cells; rat osteosarcoma D17 cells; rat monocytes / macrophages DH82 cells; simian kidney SV-40 transformed fibroblasts (COS7) cells; simian kidney CVI-76 cells; or African green monkey kidney (VERO, VERO-76) cells. An exhaustive list of mammalian cell lines can be found in the American Type Culture Collection catalog (ATCC, Manassas, VA). In some embodiments, the cell lines disclosed herein are cell lines other than mouse cell lines. In certain embodiments, the engineered cell line is a CHO cell line. Suitable CHO cell lines include, but are not limited to, CHO-K1, CHO-K1SV, CHO GS- / -, CHO S, DG44, DuxB11 and their derivatives.
[0028] In various embodiments, the parental cell line can be deficient in glutamine synthetase (GS), dihydrofolate reductase (DHFR), hypoxanthine-guanine phosphoribosyltransferase (HPRT), asparagine synthetase (ASNS), phosphoserine phosphatase (PSPH), or a combination thereof. For example, the chromosomal sequences encoding GS, DHFR, HPRT, ASNS, and / or PSPH can be inactivated. In a specific embodiment, all chromosomal sequences encoding GS, DHFR, HPRT, ASNS, and / or PSPH are inactivated in the parental cell line.
[0029] (b) Optional nucleic acid encoding a recombinant protein
[0030] In some embodiments, the engineered cell lines disclosed herein can further comprise at least one nucleic acid encoding a recombinant protein. Generally, the recombinant protein is heterologous, meaning that the protein is not native to the cell. The recombinant protein can be, but is not limited to, a therapeutic protein selected from the group consisting of: antibodies, antibody fragments, monoclonal antibodies, humanized antibodies, humanized monoclonal antibodies, chimeric antibodies, IgG molecules, IgG heavy chains, IgG light chains, IgA molecules, IgD molecules, IgE molecules, IgM molecules, vaccines, growth factors, cytokines, interferons, interleukins, hormones, coagulation (or clotting) factors, blood components, enzymes, therapeutic proteins, nutritional proteins, functional fragments or functional variants of any of the foregoing, or fusion proteins comprising any of the foregoing proteins and / or their functional fragments or variants. In a particular embodiment, the recombinant protein is a bispecific antibody or a multispecific antibody, or a protein that requires an effector protein for expression.
[0031] In some embodiments, the nucleic acid encoding the recombinant protein can be linked to a sequence encoding serine hydroxymethyltransferase 2 (SHMT2), phosphoserine phosphatase (PSPH), asparagine synthetase (ASNS), hypoxanthine-guanine phosphoribosyltransferase (HPRT), dihydrofolate reductase (DHFR), and / or glutamine synthetase (GS) such that SHMT2, PSPH, ASNS, HPRT, DHFR, and / or GS can be used as selectable markers. The nucleic acid encoding the recombinant protein can also be linked to a sequence encoding at least one antibiotic resistance gene and / or a sequence encoding a marker protein such as a fluorescent protein. In some embodiments, the nucleic acid encoding the recombinant protein can be part of an expression construct. The expression construct or vector can contain additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selectable marker sequences, origins of replication, and the like. Additional information can be found in "Current Protocols in Molecular Biology", Ausubel et al., John Wiley & Sons, New York, 2003 or "Molecular Cloning: A Laboratory Manual" Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001.
[0032] In some embodiments, the nucleic acid encoding the recombinant protein can be extrachromosomally located. That is, the nucleic acid encoding the recombinant protein can be transiently expressed from a plasmid, cosmid, artificial chromosome, minichromosome, or another extrachromosomal construct. In other embodiments, the nucleic acid encoding the recombinant protein can be chromosomally integrated into the genome of the cell. The integration can be random or targeted. Accordingly, the recombinant protein can be stably expressed. In some iterations of this embodiment, the nucleic acid sequence encoding the recombinant protein can be operably linked to an appropriate heterologous expression control sequence (i.e., a promoter). In other iterations, the nucleic acid sequence encoding the recombinant protein can be placed under the control of an endogenous expression control sequence. The nucleic acid sequence encoding the recombinant protein can be integrated into the genome of the cell line using homologous recombination, targeted endonuclease-mediated genome editing, viral vectors, transposons, recombinase-mediated cassette exchange systems, plasmids, and other well-known means. Additional guidance can be found in Ausubel et al. 2003, supra and Sambrook & Russell, 2001, supra.
[0033] (II) Kit
[0034] A further aspect of the present disclosure provides a kit for producing a recombinant protein, wherein the kit comprises any of the engineered cell lines detailed in section (I) above. The kit may further comprise a cell growth medium, a transfection reagent, a plasmid vector, a selection medium, means for purifying the recombinant protein, buffers, and the like. The kits provided herein generally include instructions for growing the cell line and using it to produce a recombinant protein. The instructions included in the kit may be adhered to the packaging material or may be included as a package insert. While the instructions are typically written or printed materials, they are not limited thereto. The present disclosure contemplates any medium capable of storing such instructions and communicating them to the end user. Such media include, but are not limited to, electronic storage media (e.g., disk, tape, cassette, chip), optical media (e.g., CD ROM), and the like. As used herein, the term "instructions" may include the address of an Internet website that provides the instructions.
[0035] (III) Method for preparing an engineered cell line
[0036] Another aspect of the present disclosure provides a method for preparing or engineering a cell line that reduces or eliminates the expression of SHMT2 and / or GS, which is described in section (I) above. The chromosomal sequences encoding SHMT2 and / or GS can be knocked down or knocked out using various techniques. Generally, the engineered cell line is prepared using a targeted endonuclease-mediated genome modification process. Those skilled in the art will understand that the engineered cell line can also be prepared using a site-specific recombination system, random mutagenesis, or other methods known in the art.
[0037] Generally, an engineered cell line is prepared by a method comprising introducing at least one targeted endonuclease or a nucleic acid encoding the targeted endonuclease into a target parental cell line, wherein the targeted endonuclease targets the chromosomal sequence encoding SHMT2 and / or GS. The targeted endonuclease recognizes and binds to a specific chromosomal sequence and introduces a double-strand break. In some embodiments, the double-strand break is repaired by the non-homologous end joining (NHEJ) repair process. Since NHEJ is error-prone, deletions, insertions, and / or substitutions of at least one nucleotide may occur, thereby disrupting the reading frame of the chromosomal sequence such that no protein product or a non-functional protein is produced, for example, by disruption of the enzymatic active site of the protein. In other embodiments, the targeted endonuclease may also be used to alter the chromosomal sequence via a homologous recombination reaction by co-introducing a polynucleotide having substantial sequence identity to a portion of the targeted chromosomal sequence. In such cases, the double-strand break introduced by the targeted endonuclease is repaired by a homology-directed repair process such that the chromosomal sequence exchanges with the polynucleotide in a manner that results in an alteration or modification of the chromosomal sequence (e.g., by integration of an exogenous sequence).
[0038] (a) Targeted endonuclease
[0039] A variety of targeted endonucleases can be used to modify the chromosomal sequences encoding SHMT2 and / or GS. The targeted endonucleases can be naturally occurring proteins or engineered proteins. Suitable targeted endonucleases include, but are not limited to, zinc finger nucleases (ZFNs), CRISPR nucleases, transcription activator-like effector (TALE) nucleases (TALENs), meganucleases, chimeric nucleases, site-specific endonucleases, and artificial targeted DNA double-strand break inducers.
[0040] (i) Zinc finger nuclease
[0041] In a specific embodiment, the targeted endonuclease can be a pair of zinc finger nucleases (ZFNs). The ZFNs bind to a specific target sequence and introduce a double-strand break within the target cleavage site. Generally, ZFNs comprise a DNA-binding domain (i.e., zinc fingers) and a cleavage domain (i.e., nuclease), each of which is described below.
[0042] DNA binding domain。DNA binding domains or zinc fingers can be engineered to recognize and bind to any selected nucleic acid sequence. See, e.g., Beerli et al. (2002) Nat. Biotechnol. 20:135-141; Pabo et al. (2001) Ann. Rev. Biochem. 70:313-340; Isalan et al. (2001) Nat. Biotechnol. 19:656-660; Segal et al. (2001) Curr. Opin. Biotechnol. 12:632-637; Choo et al. (2000) Curr. Opin. Struct. Biol. 10:411-416; Zhang et al. (2000) J. Biol. Chem. 275(43):33850-33860; Doyon et al. (2008) Nat. Biotechnol. 26:702-708; and Santiago et al. (2008) Proc. Natl. Acad. Sci. USA 105:5809-5814. Engineered zinc finger binding domains may have novel binding specificities compared to naturally occurring zinc finger proteins. Methods of engineering include, but are not limited to, rational design and various types of selection. Rational design includes, for example, using databases that contain duplex, triplex, and / or quadruplex nucleotide sequences and individual zinc finger amino acid sequences, where each duplex, triplex, or quadruplex nucleotide sequence is related to one or more amino acid sequences of a zinc finger that binds a specific triplex or quadruplex sequence. See, e.g., U.S. Patent Nos. 6,453,242 and 6,534,261, the disclosures of which are incorporated herein by reference in their entirety. For example, the algorithms described in U.S. Patent 6,453,242 can be used to design zinc finger binding domains to target preselected sequences. Alternative methods, such as rational design using non-degenerate recognition code tables, may also be used to design zinc finger binding domains to target specific sequences (Sera et al. (2002) Biochemistry 41:7074-7081). Web-based tools that are publicly available for identifying potential target sites in DNA sequences and for designing zinc finger binding domains are known in the art. For example, a tool for identifying potential target sites in DNA sequences can be found at zincfingertools.org. A tool for designing zinc finger binding domains can be found at zifit.partners.org / ZiFiT. (See also, Mandell et al. (2006) Nuc. Acid Res. 34:W516-W523; Sander et al. (2007) Nuc. Acid Res. 35:W599-W605.)
[0043] Zinc finger binding domains can be designed to recognize and bind DNA sequences ranging in length from about 3 nucleotides to about 21 nucleotides. In one embodiment, the zinc finger binding domain can be designed to recognize and bind DNA sequences ranging in length from about 9 to about 18 nucleotides. Generally, the zinc finger binding domains of zinc finger nucleases used herein contain at least three zinc finger recognition regions or zinc fingers, where each zinc finger binds 3 nucleotides. In one embodiment, the zinc finger binding domain contains four zinc finger recognition regions. In another embodiment, the zinc finger binding domain contains five zinc finger recognition regions. In yet another embodiment, the zinc finger binding domain contains six zinc finger recognition regions. The zinc finger binding domain can be designed to bind any suitable target DNA sequence. See, e.g., U.S. Patent Nos. 6,607,882; 6,534,261 and 6,453,242, the disclosures of which are incorporated herein by reference in their entirety.
[0044] Exemplary methods for selecting zinc finger recognition regions include phage display and two-hybrid systems, which are described in U.S. Patent Nos. 5,789,538; 5,925,523; 6,007,988; 6,013,453; 6,410,248; 6,140,466; 6,200,759; and 6,242,568; and WO 98 / 37186; WO 98 / 53057; WO 00 / 27878; WO 01 / 88197 and GB 2,338,237, each of which is incorporated herein by reference in its entirety. Additionally, enhancement of the binding specificity of zinc finger binding domains has been described, e.g., in WO 02 / 077227, the entire disclosure of which is incorporated herein by reference.
[0045] Zinc finger binding domains and methods for designing and constructing fusion proteins (and polynucleotides encoding them) are known to those of skill in the art and are described in detail, e.g., in U.S. Patent No. 7,888,121, the disclosure of which is incorporated herein by reference in its entirety. Zinc finger recognition regions and / or multi-fingered zinc finger proteins can be joined together using suitable linker sequences, including, e.g., linkers that are five or more amino acids in length. For non-limiting examples of linker sequences that are six or more amino acids in length, see U.S. Patent Nos. 6,479,626; 6,903,185; and 7,153,949, the disclosures of which are incorporated herein by reference in their entirety. The zinc finger binding domains described herein can include combinations of suitable linkers between individual zinc fingers of a protein.
[0046] Cleavage domain.The zinc finger nuclease also includes a cleavage domain. The cleavage domain portion of the zinc finger nuclease can be obtained from any endonuclease or exonuclease. Non-limiting examples of endonucleases from which the cleavage domain can be derived include, but are not limited to, restriction endonucleases and homing endonucleases. See, for example, New England Biolabs Catalog or Belfort et al. (1997) Nucleic Acids Res. 25:3379-3388. Additional enzymes that cleave DNA are known (e.g., S1 nuclease; mung bean nuclease; pancreatic DNase I; micrococcal nuclease; yeast HO endonuclease). See also Linn et al. (eds.) Nucleases, Cold Spring Harbor Laboratory Press, 1993. One or more of these enzymes (or functional fragments thereof) can be used as a source of the cleavage domain.
[0047] The cleavage domain can also be derived from an enzyme or a portion thereof that requires dimerization for cleavage activity as described above. Two zinc finger nucleases may be required for cleavage because each nuclease contains a monomer of the active enzyme dimer. Alternatively, a single zinc finger nuclease can contain two monomers to produce an active enzyme dimer. As used herein, an "active enzyme dimer" is an enzyme dimer capable of cleaving a nucleic acid molecule. The two cleavage monomers can be derived from the same endonuclease (or functional fragment thereof), or each monomer can be derived from a different endonuclease (or functional fragment thereof).
[0048] When two cleavage monomers are used to form an active enzyme dimer, the recognition sites for the two zinc fingers are preferably arranged such that the binding of the two zinc fingers to their respective recognition sites positions the cleavage monomers in a spatial orientation with respect to each other that allows the cleavage monomers to form an active enzyme dimer, for example, by dimerization. As a result, the proximal edges of the recognition sites can be separated by about 5 to about 18 nucleotides. For example, the proximal edges can be separated by about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 nucleotides. However, it should be understood that any integer number of nucleotides or nucleotide pairs can be between the two recognition sites (e.g., about 2 to about 50 nucleotide pairs or more). The proximal edges of the recognition sites of the zinc finger nuclease, such as those described in detail herein, can be separated by 6 nucleotides. Generally, the cleavage site is located between the recognition sites.
[0049] Restriction endonucleases (restriction enzymes) are present in many species and are capable of sequence-specific binding to DNA (at recognition sites) and cleaving the DNA at or near the binding site. Certain restriction enzymes (e.g., type IIS) cleave DNA at sites distant from the recognition site and have separable binding and cleavage domains. For example, the type IIS enzyme FokI catalyzes double-stranded cleavage of DNA 9 nucleotides from its recognition site on one strand and 13 nucleotides from its recognition site on the other strand. See, e.g., U.S. Patent Nos. 5,356,802; 5,436,150 and 5,487,994; and Li et al. (1992) Proc. Natl. Acad. Sci. USA 89:4275-4279; Li et al. (1993) Proc. Natl. Acad. Sci. USA 90:2764-2768; Kim et al. (1994a) Proc. Natl. Acad. Sci. USA 91:883-887; Kim et al. (1994b) J. Biol. Chem. 269:31978-31982. Thus, zinc finger nucleases can comprise a cleavage domain from at least one type IIS restriction enzyme and one or more zinc finger binding domains, which may or may not be engineered. Exemplary type IIS restriction enzymes are described, e.g., in International Publication WO 07 / 014,275, the disclosure of which is incorporated herein by reference in its entirety. Additional restriction enzymes also contain separable binding and cleavage domains and are also contemplated by this disclosure. See, e.g., Roberts et al. (2003) Nucleic Acids Res. 31:418-420.
[0050] An exemplary type IIS restriction enzyme whose cleavage domain can be separated from the binding domain is FokI. This particular enzyme is active as a dimer (Bitinaite et al. (1998) Proc. Natl. Acad. Sci. USA 95:10,570-10,575). Accordingly, for the purposes of this disclosure, the portion of the FokI enzyme used in a zinc finger nuclease is considered a cleavage monomer. Thus, for targeted double-stranded cleavage using the FokI cleavage domain, two zinc finger nucleases each comprising a FokI cleavage monomer can be used to reconstitute an active enzyme dimer. Alternatively, a single polypeptide molecule containing one zinc finger binding domain and two FokI cleavage monomers can also be used.
[0051] In certain embodiments, the cleavage domain comprises one or more engineered cleavage monomers that minimize or prevent homodimerization. As a non-limiting example, the amino acid residues at positions 446, 447, 479, 483, 484, 486, 487, 490, 491, 496, 498, 499, 500, 531, 534, 537, and 538 of FokI are all targets for influencing the dimerization of the FokI cleavage half-domains. Exemplary engineered FokI cleavage monomers that form obligate heterodimers include a pair in which the first cleavage monomer comprises mutations at amino acid residue positions 490 and 538 of FokI, and the second cleavage monomer comprises mutations at amino acid residue positions 486 and 499.
[0052] Thus, in one embodiment of the engineered cleavage monomers, the mutation at amino acid position 490 replaces Glu (E) with Lys (K); the mutation at amino acid residue 538 replaces Iso (I) with Lys (K); the mutation at amino acid residue 486 replaces Gln (Q) with Glu (E); and the mutation at position 499 replaces Iso (I) with Lys (K). Specifically, the engineered cleavage monomers can be prepared by mutating position 490 from E to K and position 538 from I to K in one cleavage monomer to produce an engineered cleavage monomer designated "E490K:I538K", and by mutating position 486 from Q to E and position 499 from I to K in the other cleavage monomer to produce an engineered cleavage monomer designated "Q486E:I499K". The above-described engineered cleavage monomers are obligate heterodimer mutants in which aberrant cleavage is minimized or eliminated. The engineered cleavage monomers can be prepared using suitable methods, such as site-directed mutagenesis of wild-type cleavage monomers (FokI) as described in U.S. Patent No. 7,888,121, which is incorporated herein by reference in its entirety.
[0053] Additional domain.In some embodiments, the zinc finger nuclease further comprises at least one nuclear localization sequence (NLS). An NLS is an amino acid sequence that promotes targeting of the zinc finger nuclease protein to the nucleus to introduce a double-strand break at a target sequence in a chromosome. Nuclear localization signals are known in the art (see, e.g., Lange et al., J. Biol. Chem., 2007, 282:5101-5105). Non-limiting examples of nuclear localization signals include PKKKRKV (SEQ ID NO:1), PKKKRRV (SEQ ID NO:2), KRPAATKKAGQAKKKK (SEQ ID NO:3), YGRKKRRQRRR (SEQ ID NO:4), RKKRRQRRR (SEQ ID NO:5), PAAKRVKLD (SEQ ID NO:6), RQRRNELKRSP (SEQ ID NO:7), VSRKRPRP (SEQ ID NO:8), PPKKARED (SEQ ID NO:9), PQPKKKPL (SEQ ID NO:10), SALIKKKKKMAP (SEQ ID NO:11), PKQKKRK (SEQ ID NO:12), RKLKKKIKKL (SEQ ID NO:13), REKKKFLKRR (SEQ ID NO:14), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO:15), RKCLQAGMNLEARKTKK (SEQ ID NO:16), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO:17), and RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO:18). The NLS can be located at the N-terminus, C-terminus, or internal location of the zinc finger nuclease.
[0054] In additional embodiments, the zinc finger nuclease may further comprise at least one cell penetrating domain. Examples of suitable cell penetrating domains include, but are not limited to, GRKKRRQRRRPPQPKKKRKV (SEQ ID NO:19), PLSSIFSRIGDPPKKKRKV (SEQ ID NO:20), GALFLGWLGAAGSTMGAPKKKRKV (SEQ ID NO:21), GALFLGFLGAAGSTMGAWSQPKKKRKV (SEQ ID NO:22), KETWWETWWTEWSQPKKKRKV (SEQ ID NO:23), YARAAARQARA (SEQ ID NO:24), THRLPRRRRRR (SEQ ID NO:25), GGRRARRRRRR (SEQ ID NO:26), RRQRRTSKLMKR (SEQ ID NO:27), GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO:28), KALAWEAKLAKALAKALAKHLAKALAKALKCEA (SEQ ID NO:29), and RQIKIWFQNRRMKWKK (SEQ ID NO:30). The cell penetrating domain may be located at the N-terminus, C-terminus, or internal location of the zinc finger nuclease.
[0055] In other embodiments, the zinc finger nuclease can further comprise at least one marker domain. Non-limiting examples of marker domains include fluorescent proteins, purification tags, and epitope tags. In one embodiment, the marker domain can be a fluorescent protein. Non-limiting examples of suitable fluorescent proteins include green fluorescent protein (e.g., GFP, GFP-2, tagGFP, turboGFP, EGFP, Emerald, Azami Green, Monomeric Azami Green, CopGFP, AceGFP, ZsGreen1), yellow fluorescent protein (e.g., YFP, EYFP, Citrine, Venus, YPet, PhiYFP, ZsYellow1), blue fluorescent protein (e.g., EBFP, EBFP2, Azurite, mKalama1, GFPuv, Sapphire, T-sapphire), cyan fluorescent protein (e.g., ECFP, Cerulean, CyPet, AmCyan1, Midoriishi-Cyan), red fluorescent protein (mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRed1, AsRed2, eqFP611, mRasberry, mStrawberry, Jred), and orange fluorescent protein (mOrange, mKO, Kusabira-Orange, Monomeric Kusabira-Orange, mTangerine, tdTomato) or any other suitable fluorescent protein. In another embodiment, the marker domain can be a purification tag and / or an epitope tag. Suitable tags include, but are not limited to, poly(His) tag, FLAG (or DDK) tag, Halo tag, AcV5 tag, AU1 tag, AU5 tag, biotin carboxyl carrier protein (BCCP), calmodulin binding protein (CBP), chitin binding domain (CBD), E tag, E2 tag, ECS tag, eXact tag, Glu-Glu tag, glutathione-S-transferase (GST), HA tag, HSV tag, KT3 tag, maltose binding protein (MBP), MAP tag, Myc tag, NE tag, NusA tag, PDZ tag, S tag, S1 tag, SBP tag, Softag 1 tag, Softag 3 tag, Spot tag, Strep tag, SUMO tag, T7 tag, tandem affinity purification (TAP) tag, thioredoxin (TRX), V5 tag, VSV-G tag, and Xa tag.The marker domain can be located at the N-terminus, C-terminus or internal position of the zinc finger nuclease.
[0056] At least one nuclear localization signal, at least one cell-penetrating domain and / or at least one marker domain can be directly linked to the zinc finger nuclease via one or more chemical bonds (e.g., covalent bonds). Alternatively, at least one nuclear localization signal, at least one cell-penetrating domain and / or at least one marker domain can be indirectly linked to the zinc finger nuclease via one or more linkers. Suitable linkers include amino acids, peptides, nucleotides, nucleic acids, organic linker molecules (e.g., maleimide derivatives, N-ethoxybenzylimidazole, biphenyl-3,4',5-tricarboxylic acid, p-aminobenzyloxycarbonyl, etc.), disulfide bond linkers and polymeric linkers (e.g., PEG). The linker can include one or more spacer groups, including but not limited to alkylene, alkenylene, alkynylene, alkyl, alkenyl, alkynyl, alkoxy, aryl, heteroaryl, aralkyl, aralkenyl, aralkynyl, etc. The linker can be neutral, or carry a positive or negative charge. Additionally, the linker can be cleavable such that the covalent bond of the linker that connects the linker to another chemical group can be broken or cleaved under specific conditions, including pH, temperature, salt concentration, light, catalyst or enzyme. In some embodiments, the linker can be a peptide linker. The peptide linker can be a flexible amino acid linker or a rigid amino acid linker. Additional examples of suitable linkers are well known in the art, and procedures for designing linkers are readily available (Crasto et al., Protein Eng., 2000, 13(5):309-312).
[0057] (ii) CRISPR ribonucleoprotein (RNP)
[0058] In other embodiments, the targeted endonuclease can be a clustered regularly interspaced short palindromic repeat (CRISPR) nuclease. CRISPR nucleases are RNA-guided nucleases derived from bacterial or archaeal CRISPR / CRISPR-associated (Cas) systems. The CRISPR RNP system comprises a CRISPR nuclease and a guide RNA.
[0059] Nuclease. CRISPR nucleases can be derived from type I (i.e., IA, IB, IC, ID, IE or IF), type II (i.e., IIA, IIB or IIC), type III (i.e., IIIA or IIIB), type V or type VI CRISPR systems, which are present in various bacteria and archaea. For example, CRISPR nucleases can be from species of the genus Streptococcus (e.g., Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus pasteurianus), species of the genus Campylobacter (e.g., Campylobacter jejuni), species of the genus Francisella (e.g., Francisella novicida), Acaryochloris sp., species of the genus Acetohalobium, species of the genus Acidaminococcus, species of the genus Acidithiobacillus, species of the genus Alicyclobacillus, species of the genus Allochromatium, species of the genus Ammonifex, species of the genus Anabaena, species of the genus Arthrospira, species of the genus Bacillus, species of the order Burkholderiales, species of the genus Caldicelulosiruptor, Candidatus sp., species of the genus Clostridium, species of the genus Crocosphaera, species of the genus Cyanothece, species of the genus Exiguobacterium, species of the genus Finegoldia, species of the genus Ktedonobacter, species of the family Lachnospiraceae, species of the genus Lactobacillus, species of the genus Lyngbya, species of the genus Marinobacter, species of the genus Methanohalobium, species of the genus Microscilla, species of the genus Microcoleus) Microcystis sp., Natranaerobius sp., Neisseria sp., Nitrosococcus sp., Nocardiopsis sp., Nodularia sp., Nostoc sp., Oscillatoria sp., Polaromonas sp., Pelotomaculum sp., Pseudoalteromonas sp., Petrotoga sp., Prevotella sp., Staphylococcus sp., Streptomyces sp., Streptosporangium sp., Synechococcus sp., Thermosipho sp., or Verrucomicrobia sp. In other embodiments, the CRISPR nuclease can be derived from an archaeal CRISPR system, a CRISPR / CasX system, or a CRISPR / CasY system (Burstein et al., Nature, 2017, 542(7640):237-241).
[0060] In some embodiments, the CRISPR nuclease can be derived from a type II CRISPR nuclease. For example, the type II CRISPR nuclease can be a Cas9 protein. Suitable Cas9 nucleases include Streptococcus pyogenes Cas9 (SpCas9), Francisella novicida Cas9 (FnCas9), Staphylococcus aureus (SaCas9), Streptococcus thermophilus Cas9 (StCas9), Streptococcus pasteurianus (SpaCas9), Campylobacter jejuni Cas9 (CjCas9), Neisseria meningitidis Cas9 (NmCas9), or Neisseria cinerea Cas9 (NcCas9). In other embodiments, the CRISPR nuclease can be derived from a type V CRISPR nuclease, such as a Cpf1 nuclease. Suitable Cpf1 nucleases include Francisella novicida Cpf1 (FnCpf1), Acidaminococcus sp. Cpf1 (AsCpf1), or Lachnospiraceae bacterium ND2006 Cpf1 (LbCpf1). In yet another embodiment, the CRISPR nuclease can be derived from a type VI CRISPR nuclease, such as Leptotrichia wadei Cas13a (LwaCas13a) or Leptotrichia shahii Cas13a (LshCas13a).
[0061] The CRISPR nuclease can be a wild-type CRISPR nuclease, a modified CRISPR nuclease, or a fragment of a wild-type or modified CRISPR nuclease. The CRISPR nuclease can be modified to increase nucleic acid binding affinity and / or specificity, alter enzymatic activity, and / or alter another property of the protein. For example, the nuclease (i.e., DNase, RNase) domain of the CRISPR nuclease can be modified, deleted, or inactivated. The CRISPR nuclease can be truncated to remove domains that are not essential for the function of the nuclease.
[0062] CRISPR nucleases contain two nuclease domains. For example, Cas9 nuclease contains an HNH domain that cleaves the complementary strand of the guide RNA and a RuvC domain that cleaves the non-complementary strand; Cpf1 nuclease contains a RuvC domain and a NUC domain; and Cas13a nuclease contains two HNEPN domains. When both nuclease domains are functional, the CRISPR nuclease introduces a double-strand break. Either nuclease domain can be inactivated by one or more mutations and / or deletions, resulting in variants that introduce a single-strand break in one strand of the double-stranded sequence. For example, one or more mutations in the RuvC domain of Cas9 nuclease (e.g., D10A, D8A, E762A, and / or D986A) result in an HNH nickase that cleaves the complementary strand of the guide RNA; and one or more mutations in the HNH domain of Cas9 nuclease (e.g., H840A, H559A, N854A, N856A, and / or N863A) result in a RuvC nickase that cleaves the non-complementary strand of the guide RNA. Similar mutations can convert Cpf1 and Cas13a nucleases into nickases. Two CRISPR nickases targeting opposite strands of a chromosomal sequence (via a pair of offset guide RNAs) can be used in combination to generate a double-strand break in the chromosomal sequence. Dual CRISPR nickase RNPs can increase target specificity and reduce off-target effects.
[0063] Additional domain.The CRISPR nuclease may further comprise at least one nuclear localization sequence (NLS). An NLS is an amino acid sequence that promotes targeting of the zinc finger nuclease protein to the nucleus to introduce a double-strand break at a target sequence in a chromosome. Nuclear localization signals are known in the art (see, e.g., Lange et al., J. Biol. Chem., 2007, 282:5101-5105). Non-limiting examples of nuclear localization signals include PKKKRKV (SEQ ID NO:1), PKKKRRV (SEQ ID NO:2), KRPAATKKAGQAKKKK (SEQ ID NO:3), YGRKKRRQRRR (SEQ ID NO:4), RKKRRQRRR (SEQ ID NO:5), PAAKRVKLD (SEQ ID NO:6), RQRRNELKRSP (SEQ ID NO:7), VSRKRPRP (SEQ ID NO:8), PPKKARED (SEQ ID NO:9), PQPKKKPL (SEQ ID NO:10), SALIKKKKKMAP (SEQ ID NO:11), PKQKKRK (SEQ ID NO:12), RKLKKKIKKL (SEQ ID NO:13), REKKKFLKRR (SEQ ID NO:14), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO:15), RKCLQAGMNLEARKTKK (SEQ ID NO:16), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO:17), and RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO:18). The NLS may be located at the N-terminus, C-terminus, or internally within the CRISPR nuclease.
[0064] In additional embodiments, the CRISPR nuclease may further comprise at least one cell-penetrating domain. Examples of suitable cell-penetrating domains include, but are not limited to, GRKKRRQRRRPPQPKKKRKV (SEQ ID NO:19), PLSSIFSRIGDPPKKKRKV (SEQ ID NO:20), GALFLGWLGAAGSTMGAPKKKRKV(SEQ ID NO:21), GALFLGFLGAAGSTMGAWSQPKKKRKV(SEQ ID NO:22), KETWWETWWTEWSQPKKKRKV(SEQ ID NO:23), YARAAARQARA(SEQ ID NO:24), THRLPRRRRRR(SEQ ID NO:25), GGRRARRRRRR(SEQ IDNO:26), RRQRRTSKLMKR(SEQ ID NO:27), GWTLNSAGYLLGKINLKALAALAKKIL(SEQ ID NO:28), KALAWEAKLAKALAKALAKHLAKALAKALKCEA(SEQ ID NO:29), and RQIKIWFQNRRMKWKK(SEQ ID NO:30). The cell-penetrating domain may be located at the N-terminus, C-terminus, or internal location of the CRISPR protein.
[0065] In other embodiments, the CRISPR nuclease can further comprise at least one marker domain. Non-limiting examples of marker domains include fluorescent proteins, purification tags, and epitope tags. In one embodiment, the marker domain can be a fluorescent protein. Non-limiting examples of suitable fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, EGFP, Emerald, Azami Green, MonomericAzamiGreen, CopGFP, AceGFP, ZsGreen1), yellow fluorescent proteins (e.g., YFP, EYFP, Citrine, Venus, YPet, PhiYFP, ZsYellow1), blue fluorescent proteins (e.g., EBFP, EBFP2, Azurite, mKalama1, GFPuv, Sapphire, T-sapphire), cyan fluorescent proteins (e.g., ECFP, Cerulean, CyPet, AmCyan1, Midoriishi-Cyan), red fluorescent proteins (mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRed1, AsRed2, eqFP611, mRasberry, mStrawberry, Jred), and orange fluorescent proteins (mOrange, mKO, Kusabira-Orange, Monomeric Kusabira-Orange, mTangerine, tdTomato) or any other suitable fluorescent protein. In another embodiment, the marker domain can be a purification tag and / or an epitope tag. Suitable tags include, but are not limited to, poly(His) tag, FLAG (or DDK) tag, Halo tag, AcV5 tag, AU1 tag, AU5 tag, biotin carboxyl carrier protein (BCCP), calmodulin-binding protein (CBP), chitin-binding domain (CBD), E tag, E2 tag, ECS tag, eXact tag, Glu-Glu tag, glutathione-S-transferase (GST), HA tag, HSV tag, KT3 tag, maltose-binding protein (MBP), MAP tag, Myc tag, NE tag, NusA tag, PDZ tag, S tag, S1 tag, SBP tag, Softag 1 tag, Softag 3 tag, Spot tag, Strep tag, SUMO tag, T7 tag, tandem affinity purification (TAP) tag, thioredoxin (TRX), V5 tag, VSV-G tag, and Xa tag.The marker domain can be located at the N-terminus, C-terminus, or an internal position of the CRISPR nuclease.
[0066] At least one nuclear localization signal, at least one cell-penetrating domain, and / or at least one marker domain can be directly linked to the CRISPR nuclease via one or more chemical bonds (e.g., covalent bonds). Alternatively, at least one nuclear localization signal, at least one cell-penetrating domain, and / or at least one marker domain can be indirectly linked to the CRISPR nuclease via one or more linkers. Suitable linkers include amino acids, peptides, nucleotides, nucleic acids, organic linker molecules (e.g., maleimide derivatives, N-ethoxybenzylimidazole, biphenyl-3,4',5-tricarboxylic acid, p-aminobenzyloxycarbonyl, etc.), disulfide bond linkers, and polymeric linkers (e.g., PEG). The linker can include one or more spacer groups, including but not limited to alkylene, alkenylene, alkynylene, alkyl, alkenyl, alkynyl, alkoxy, aryl, heteroaryl, aralkyl, aralkenyl, aralkynyl, etc. The linker can be neutral, or carry a positive or negative charge. Additionally, the linker can be cleavable such that the covalent bond of the linker connecting it to another chemical group can be broken or cleaved under specific conditions, including pH, temperature, salt concentration, light, catalyst, or enzyme. In some embodiments, the linker can be a peptide linker. The peptide linker can be a flexible amino acid linker or a rigid amino acid linker. Additional examples of suitable linkers are well known in the art, and procedures for designing linkers are readily available in the art.
[0067] Guide RNA 。The CRISPR nuclease is guided to its target site by a guide RNA. The guide RNA hybridizes to the target site and interacts with the CRISPR nuclease to direct the CRISPR nuclease to the target site in the chromosomal sequence. There is no sequence restriction for the target site, except that the sequence is delimited by a protospacer adjacent motif (PAM). CRISPR proteins from different bacterial species recognize different PAM sequences. For example, PAM sequences include 5'-NGG (SpCas9, FnCAs9), 5'-NGRRT (SaCas9), 5'-NNAGAAW (StCas9), 5'-NNNNGATT (NmCas9), 5-NNNNRYAC (CjCas9), and 5'-TTTV (Cpf1), where N is defined as any nucleotide, R is defined as G or A, W is defined as A or T, Y is defined as C or T, and V is defined as A, C, or G. The Cas9 PAM is located 3' of the target site, while the cpf1 PAM is located 5' of the target site.
[0068] The guide RNA comprises three regions: a first region at the 5' end that is complementary to the sequence at the target site, a second internal region that forms a stem-loop structure, and a third 3' region that remains substantially single-stranded. The first region of each guide RNA is different such that each guide RNA directs the CRISPR nuclease to a specific target site. The second and third regions (also referred to as the scaffold regions) of each guide RNA can be the same in all guide RNAs.
[0069] The first region of the guide RNA is complementary to the sequence at the target site (i.e., the protospacer sequence) such that the first region of the guide RNA can base pair with the sequence at the target site. The complementarity between the first region of the guide RNA (i.e., the crRNA) and the target sequence can be at least 80%, at least 85%, at least 90%, at least 95% or more. Generally, there are no mismatches between the sequence of the first region of the guide RNA and the sequence at the target site (i.e., the complementarity is complete). In various embodiments, the first region of the guide RNA can comprise from about 10 nucleotides to more than about 25 nucleotides. For example, the length of the base pairing region between the first region of the guide RNA and the target site in a chromosomal sequence can be about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25 or more than 25 nucleotides. In an exemplary embodiment, the length of the first region of the guide RNA is about 19, 20 or 21 nucleotides.
[0070] The guide RNA further comprises a second region that forms a secondary structure. In some embodiments, the secondary structure comprises a stem (or hairpin) and a loop. The lengths of the loop and the stem can vary. For example, the length of the loop can range from about 3 to about 10 nucleotides, while the length of the stem can range from about 6 to about 20 base pairs. The stem can comprise one or more bulges of 1 to about 10 nucleotides. Thus, the overall length of the second region can range from about 16 to about 60 nucleotides. In an exemplary embodiment, the length of the loop is about 4 nucleotides and the stem comprises about 12 base pairs.
[0071] The guide RNA further comprises a third region at the 3' end that remains substantially single-stranded. Thus, the third region has no complementarity with any chromosomal sequence in the target cell and no complementarity with the remainder of the guide RNA. The length of the third region can vary. Generally, the length of the third region is more than about 4 nucleotides. For example, the length of the third region can range from about 5 to about 60 nucleotides.
[0072] The combined length of the second and third regions (or scaffold) of the guide RNA can range from about 30 to about 120 nucleotides. In one aspect, the combined length of the second and third regions of the guide RNA ranges from about 70 to about 100 nucleotides.
[0073] In some embodiments, the guide RNA comprises a single molecule containing all three regions. In other embodiments, the guide RNA can comprise two separate molecules. The first RNA molecule can contain the first (5’) region of the guide RNA and one half of the “stem” of the second region of the guide RNA. The second RNA molecule can contain the other half of the “stem” of the second region of the guide RNA and the third region of the guide RNA. Thus, in this embodiment, the first and second RNA molecules each contain nucleotide sequences that are complementary to each other. For example, in one embodiment, the first and second RNA molecules each comprise a sequence (from about 6 to about 20 nucleotides) that base pairs with the other sequence to form a functional guide RNA.
[0074] (iii) Other targeted endonucleases
[0075] In a further embodiment, the targeted endonuclease can be a meganuclease. Meganucleases are endodeoxyribonucleases characterized by long recognition sequences, i.e., recognition sequences generally range from about 12 base pairs to about 40 base pairs. As a result of this requirement, the recognition sequence generally occurs only once in any given genome. Among meganucleases, the homing endonuclease family named LAGLIDADG has emerged as a valuable tool for studying genomes and genome engineering (see, e.g., Arnould et al., 2011, Arnould et al., 2011, Protein Eng Des Sel, 24(1-2):27-31). Other suitable meganucleases include I-CreI and I-Dmol. By modifying their recognition sequences using techniques well known to those skilled in the art, meganucleases can be targeted to specific chromosomal sequences.
[0076] In additional embodiments, the targeted endonuclease can be a transcription activator-like effector (TALE) nuclease. TALEs are transcription factors from the plant pathogen Xanthomonas that can be readily engineered to bind new DNA targets. A TALE or a truncated form thereof can be linked to the catalytic domain of an endonuclease such as FokI to generate a targeted endonuclease called a TALE nuclease or TALEN (Sanjana et al., 2012, Nat Protoc, 7(1):171-192; and Arnould et al., 2011, Protein Engineering, Design & Selection, 24(1-2):27-31).
[0077] In alternative embodiments, the targeted endonuclease can be a chimeric nuclease. Non-limiting examples of chimeric nucleases include ZF-meganucleases, TAL-meganucleases, Cas9-FokI fusions, ZF-Cas9 fusions, TAL-Cas9 fusions, and the like. Those skilled in the art are familiar with the means for generating such chimeric nuclease fusions.
[0078] In still other embodiments, the targeted endonuclease can be a site-specific endonuclease. In particular, the site-specific endonuclease can be a "rare-cutter" endonuclease, the recognition sequence of which occurs rarely in the genome. Alternatively, the site-specific endonuclease can be engineered to cleave a target site (Friedhoff et al., 2007, Methods Mol Biol 352:111-123). Generally, the recognition sequence of a site-specific endonuclease occurs only once in the genome. In a further alternative embodiment, the targeted endonuclease can be an artificial targeted DNA double-strand break inducer.
[0079] (b) Delivery of the targeted endonuclease to the cell
[0080] The method includes introducing a targeted endonuclease into a target parental cell line. The targeted endonuclease can be introduced into the cell as a purified and isolated protein or as a nucleic acid encoding the targeted endonuclease. The nucleic acid can be DNA or RNA. In embodiments where the encoding nucleic acid is mRNA, the mRNA may be 5'-capped and / or 3'-polyadenylated. In embodiments where the encoding nucleic acid is DNA, the DNA can be linear or circular. The nucleic acid can be part of a plasmid or viral vector, where the encoding DNA can be operably linked to a suitable promoter. Those skilled in the art are familiar with suitable vectors, promoters, other control elements, and means for introducing the vector into the target cell. In embodiments where the targeted endonuclease is a CRISPR nuclease, the CRISPR nuclease system can be introduced into the cell as a gRNA-protein complex.
[0081] The targeted endonuclease molecule can be introduced into the cell by various means. Suitable delivery means include microinjection, electroporation, sonoporation, gene gun, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendrimer transfection, heat shock transfection, nucleofection, magnetic transfection, lipid transfection, impalefection, optical transfection, nucleic acid uptake enhanced by proprietary reagents, and delivery via liposomes, immunoliposomes, virions, or artificial virus particles. In one specific embodiment, the targeted endonuclease molecule is introduced into the cell by nucleofection.
[0082] Optional donor polynucleotide The method for targeted genome modification or engineering can further include introducing at least one donor polynucleotide into the cell, the donor polynucleotide comprising a sequence having at least one nucleotide change relative to the target chromosomal sequence. The donor polynucleotide has substantial sequence identity with the sequence at or near the targeted site in the chromosomal sequence such that the double-strand break introduced by the targeted endonuclease can be repaired by a homology-directed repair process, and the sequence of the donor polynucleotide can be inserted into or exchanged with the chromosomal sequence, thereby modifying the chromosomal sequence. For example, the donor polynucleotide can comprise a first sequence having substantial sequence identity with the sequence on one side of the target site and a second sequence having substantial sequence identity with the sequence on the other side of the target site. The donor polynucleotide can further comprise a donor sequence for integration into the targeted chromosomal sequence. For example, the donor sequence can be a foreign sequence (e.g., a marker sequence) such that the integration of the foreign sequence disrupts the reading frame and inactivates the targeted chromosomal sequence.
[0083] The lengths of the first and second sequences in the donor polynucleotide that have substantial sequence identity to sequences at or near the target site in the chromosomal sequence can and will vary. Generally, the lengths of the first and second sequences in the donor polynucleotide are each at least about 10 nucleotides. In various embodiments, the length of the donor polynucleotide sequence having substantial sequence identity to the chromosomal sequence can be about 15 nucleotides, about 20 nucleotides, about 25 nucleotides, about 30 nucleotides, about 40 nucleotides, about 50 nucleotides, about 100 nucleotides, or more than 100 nucleotides.
[0084] The phrase "substantial sequence identity" means that the sequence in the polynucleotide has at least about 75% sequence identity to the target chromosomal sequence. In some embodiments, the sequence in the polynucleotide has about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to the target chromosomal sequence.
[0085] The length of the donor polynucleotide can and will vary. For example, the donor polynucleotide can range in length from about 20 nucleotides to about 200,000 nucleotides. In various embodiments, the donor polynucleotide can range in length from about 20 nucleotides to about 100 nucleotides, from about 100 nucleotides to about 1000 nucleotides, from about 1000 nucleotides to about 10,000 nucleotides, from about 10,000 nucleotides to about 100,000 nucleotides, or from about 100,000 nucleotides to about 200,000 nucleotides.
[0086] Generally, the donor polynucleotide is DNA. The DNA can be single-stranded or double-stranded. The DNA can be linear or circular. In some embodiments, the donor polynucleotide can be a single-stranded linear oligonucleotide containing fewer than about 200 nucleotides. In other embodiments, the donor polynucleotide can be part of a vector. Suitable vectors include DNA plasmids, viral vectors, bacterial artificial chromosomes (BACs), and yeast artificial chromosomes (YACs). In still other embodiments, the donor polynucleotide can be a PCR fragment or a nucleic acid complexed with a delivery vehicle such as a liposome or poloxamer.
[0087] A donor polynucleotide can be introduced into a cell simultaneously with a targeted endonuclease molecule. Alternatively, the donor polynucleotide and the targeted endonuclease molecule can be introduced into the cell sequentially. The ratio of the targeted endonuclease molecule to the donor polynucleotide can and will vary. Generally, the ratio of the targeted endonuclease molecule to the donor polynucleotide ranges from about 1:10 to about 10:1. In various embodiments, the ratio of the targeted endonuclease molecule to the polynucleotide can be about 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, or 10:1. In one embodiment, the ratio is about 1:1.
[0088] (c) Culturing the cells
[0089] The method further comprises maintaining the cells under appropriate conditions such that the double-strand break introduced by the targeted endonuclease can be repaired by: (i) a non-homologous end-joining repair process such that the chromosomal sequence is modified by a deletion, insertion, and / or substitution of at least one nucleotide, or optionally, (ii) a homology-directed repair process such that the chromosomal sequence exchanges with the sequence of the polynucleotide such that the chromosomal sequence is modified. In embodiments in which a nucleic acid encoding a targeted endonuclease is introduced into the cell, the method comprises maintaining the cells under appropriate conditions such that the cell expresses the targeted endonuclease.
[0090] Generally, the cells are maintained under conditions suitable for cell growth and / or maintenance. Suitable cell culture conditions are well known in the art and are described, for example, in Santiago et al. (2008) PNAS 105:5809-5814; Moehle et al. (2007) PNAS 104:3055-3060; Urnov et al. (2005) Nature 435:646-651; and Lombardo et al. (2007) Nat. Biotechnology 25:1298-1306. Those skilled in the art will appreciate that methods for culturing cells are known in the art and can and will vary depending on the cell type. In all cases, routine optimization may be used to determine the optimal technique for a particular cell type.
[0091] During this step of the method, the targeted endonuclease recognizes, binds, and creates a double-strand break at a targeted cleavage site in the chromosomal sequence, and during the repair of the double-strand break, a deletion, insertion, and / or substitution of at least one nucleotide is introduced into the targeted chromosomal sequence. In a specific embodiment, the targeted chromosomal sequence is inactivated.
[0092] After confirming that the target chromosomal sequence has been modified, individual cell clones can be isolated and genotyped (by DNA sequencing and / or protein analysis). Cells containing a modified chromosomal sequence can undergo one or more additional rounds of targeted genome modification to modify additional chromosomal sequences, resulting in double knockouts, triple knockouts, etc.
[0093] (IV) Production of recombinant proteins
[0094] Another aspect of the present disclosure encompasses methods for producing recombinant proteins in a biological production system. Suitable recombinant proteins are described in section (I)(c). The method includes expressing the target recombinant protein in any of the engineered cell lines described in section (I) above, and purifying the expressed recombinant protein. Means for producing or manufacturing recombinant proteins are well known in the art (see, for example, “Biopharmaceutical Production Technology”, Subramanian (ed.), 2012, Wiley-VCH; ISBN: 978-3-527-33029-4).
[0095] Recombinant proteins can be purified by methods including a clarification step (e.g., filtration) and one or more chromatography steps (e.g., affinity chromatography, protein A (or G) chromatography, ion exchange (i.e., cation and / or anion) chromatography).
[0096] Definitions
[0097] Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. The following references provide one of ordinary skill in the art with general definitions of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed., 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th ed., R. Rieger et al. (eds.), Springer-Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, unless otherwise specified, the following terms have the meanings given to them below.
[0098] When introducing elements of the present disclosure or its preferred embodiments, the articles "a", "an", "the", and "said" are intended to mean that there is one or more elements. The terms "comprising", "including", and "having" are intended to be inclusive and mean that there may be additional elements other than the listed elements.
[0099] As used herein, the term "endogenous sequence" refers to a chromosomal sequence that is native to a cell.
[0100] The term "exogenous sequence" refers to a chromosomal sequence that is not native to a cell or a chromosomal sequence that has moved to a different chromosomal location.
[0101] A "modified" or "genetically modified" cell is a cell in which the genome has been modified or altered, i.e., the cell contains at least one chromosomal sequence that has been modified to contain at least one nucleotide insertion, at least one nucleotide deletion, and / or at least one nucleotide substitution.
[0102] The terms "genome modification" and "genome editing" refer to processes by which a specific endogenous chromosomal sequence is altered such that the chromosomal sequence is modified. The chromosomal sequence can be modified to contain at least one nucleotide insertion, at least one nucleotide deletion, and / or at least one nucleotide substitution. The modified chromosomal sequence is inactivated such that no product is produced. Alternatively, the chromosomal sequence can be modified such that an altered product is produced.
[0103] As used herein, "gene" refers to a DNA region (including exons and introns) that encodes a gene product, as well as all DNA regions that regulate the production of the gene product, whether or not such regulatory sequences are adjacent to the coding sequence and / or the transcriptional sequence. Accordingly, a gene includes but is not necessarily limited to promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, and locus control regions.
[0104] The term "heterologous" refers to an entity that is not native to the target cell or species.
[0105] The terms "nucleic acid" and "polynucleotide" refer to polymers of deoxyribonucleotides or ribonucleotides in a linear or circular conformation. For the purposes of the present disclosure, these terms should not be construed as limiting with respect to the length of the polymer. The term can encompass known analogs of natural nucleotides, as well as nucleotides modified in the base, sugar, and / or phosphate moieties. In general, analogs of a particular nucleotide have the same base pairing specificity; i.e., an analog of A will base pair with T. The nucleotides of a nucleic acid or polynucleotide may be linked by phosphodiester, phosphorothioate, phosphoramidite, phosphorodiamidate bonds, or combinations thereof.
[0106] The term "nucleotide" refers to deoxyribonucleotide or ribonucleotide. Nucleotides may be standard nucleotides (i.e., adenosine, guanosine, cytidine, thymidine, and uridine) or nucleotide analogs. Nucleotide analogs refer to nucleotides having modified purine or pyrimidine bases or modified ribose moieties. Nucleotide analogs may be naturally occurring nucleotides (e.g., inosine) or non-naturally occurring nucleotides. Non-limiting examples of modifications on the sugar or base moiety of nucleotides include the addition (or removal) of acetyl groups, amino groups, carboxyl groups, carboxymethyl groups, hydroxyl groups, methyl groups, phosphoryl groups, and thiol groups, as well as the substitution of carbon and nitrogen atoms of the base by other atoms (e.g., 7-deazapurine). Nucleotide analogs also include dideoxynucleotides, 2'-O-methyl nucleotides, locked nucleic acids (LNA), peptide nucleic acids (PNA), and morpholino oligonucleotides.
[0107] The terms "polypeptide" and "protein" are used interchangeably to refer to polymers of amino acid residues.
[0108] As used herein, the term "target site" or "target sequence" refers to a nucleic acid sequence that defines a portion of a chromosomal sequence to be modified or edited, and to which a targeted endonuclease has been engineered to recognize and bind, provided that sufficient binding conditions exist.
[0109] The terms "upstream" and "downstream" refer to the positioning relative to a fixed position in a nucleic acid sequence. Upstream refers to the region that is 5' (i.e., closer to the 5' end of the strand) relative to that position, while downstream refers to the region that is 3' (i.e., closer to the 3' end of the strand) relative to that position.
[0110] Techniques for determining nucleic acid and amino acid sequence identity are known in the art. In general, such techniques include determining the nucleotide sequence of the mRNA of a gene and / or determining the amino acid sequence encoded thereby, and comparing these sequences to a second nucleotide or amino acid sequence. Genomic sequences can also be determined and compared in this manner. In general, identity refers to the exact nucleotide-to-nucleotide or amino acid-to-amino acid correspondence of two polynucleotide or polypeptide sequences. Two or more sequences (polynucleotide or amino acid) can be compared by determining their percentage identity. The percentage identity of two sequences, whether nucleic acid or amino acid sequences, is the number of exact matches between the two aligned sequences divided by the length of the shorter sequence and multiplied by 100. Approximate alignment of nucleic acid sequences is provided by the local homology algorithm of Smith and Waterman, Advances in Applied Mathematics 2:482-489 (1981). This algorithm can be applied to amino acid sequences by using a scoring matrix developed by Dayhoff, Atlas of Protein Sequences and Structure, M.O. Dayhoff ed., 5 suppl. 3:353-358, National Biomedical Research Foundation, Washington, D.C., USA, and standardized by Gribskov, Nucl. Acids Res. 14(6):6745-6763 (1986). An exemplary implementation of this algorithm for determining the percentage identity of sequences is provided by the "BestFit" utility application of the Genetics Computer Group (Madison, Wis.). Other suitable programs for calculating the percentage identity or similarity between sequences are generally known in the art. For example, another alignment program is BLAST used with default parameters. For example, the following default parameters can be used to use BLASTN and BLASTP: genetic code = standard; filter = none; strand = both; cutoff = 60; expect = 10; matrix = BLOSUM62; descriptions = 50 sequences; sort by = high score; database = non-redundant, GenBank+EMBL+DDBJ+PDB+GenBank CDS translations+SwissProt+Spupdate+PIR. Details of these programs can be found on the GenBank website. For the sequences described herein, the desired degree of sequence identity ranges from about 80% to 100% and any integer value therebetween. Generally, the percentage identity between sequences is at least 70-75%, preferably 80-82%, more preferably 85-90%, even more preferably 92%, still more preferably 95%, and most preferably 98% sequence identity.
[0111] Since various changes can be made in the above cells and methods without departing from the scope of the present invention, all the content included in the above description and the embodiments given below should be construed as exemplary rather than restrictive.
[0112] Examples
[0113] The following examples illustrate certain aspects of the present invention.
[0114] Example 1: Design of the glycine-formate-mediated selection system
[0115] To develop a glycine-formate-mediated metabolic selection system, a CHO cell line auxotrophic for the non-essential amino acid glycine (Gly) was first developed. A comprehensive search of the Reactome and KEGG databases was conducted to identify all genes related to the glycine-formate synthesis pathway. The endogenous serine hydroxymethyltransferase 2 gene (SHMT2) was identified as the only non-redundant gene responsible for Gly synthesis ( Figure 1 ). To generate a Gly-auxotrophic CHO cell line, a glutamine (Gln)-auxotrophic GS - / - cell line via whole-genome sequencing (WGS) of the cell line elucidated the endogenous SHMT2 coding sequence ( Figure 2 ), and as determined by digital droplet PCR (ddPCR) analysis, it was found to be present in two copies in the genome ( Figure 7 ). The CRISPR / Cas9 gene editing reagent was designed to disrupt the SHMT2 gene. The CRISPR / Cas9 target sequences are underlined in Figure 2 . Cells in EX- Culturing was performed at 37 °C with 5% CO2 under shaking conditions in CD CHO Fusion medium (MilliporeSigma 14365C), and the medium was supplemented with 6 mM L-glutamine (MilliporeSigma G7513) (Fusion+Gln). Cells were split at 0.3e6 three days before transfection. Cas9 RNP was complexed for 15 minutes at room temperature by mixing 50 pmol of Cas9 (Sigma CAS9PROT-250UG) with 150 pmol of sgRNA (Sigma). 4e5 cells were transfected with a total of 200 pmol of the complexed RNP using Lonza's 4D Nucleofector system with program DT-133 and SF Nucleofector solution. The transfected cells were transferred to a 6-well culture plate containing 2 mL of pre-warmed EX- CD CHO Fusion medium (MilliporeSigma 14365C), and the medium was supplemented with 6 mM L-glutamine (MilliporeSigma G7513). The cells were incubated in a static environment at 37 °C / 5% CO2. Forty-eight hours after transfection, 20% of the pool (400 μL) was harvested for genomic DNA extraction (gDNA), and gDNA was extracted using QuickExtract (Lucigen QE09050). 2.5 μL of gDNA was amplified for next-generation sequencing (NGS) using illumina Miseq. NGS confirmed that >90% of the pool was edited.
[0116] Single-cell clones from the Cas9-modified pool were isolated into a 96-well culture plate via a fluorescence-activated cell sorter (FACS). The single-cell clones were evaluated via NGS to identify clones that had successfully genetically disrupted both copies of SHMT2 ( Figure 3 ).
[0117] To confirm the efficacy of the glycine-formate-mediated selection mechanism as both an independent system and as part of a dual metabolic selection system, a stable selected cell population expressing multiple molecules was generated. Molecules used in the validation of this system included cyan fluorescent protein (CFP), Dasher green fluorescent protein (GFP), and human IgG1 GS - / - SHMT2 - / -Cells were cultured in Fusion+Gln+Gly+formate (For) medium. The expression vectors used in the current work contained the serine hydroxymethyltransferase 2 (SHMT2) or glutamine synthetase (GS) selection markers, as determined by design of experiment. The expression of murine SHMT2 (protein serine hydroxymethyltransferase 2; gene: SHMT2; UniProtKB ID: Q9CZN7) or murine GS (protein: glutamine synthetase {glutamate-ammonia ligase}; gene: Glul; UniProtKB ID: P15105) was driven by the 5’ SV40 promoter with an SV40 polyadenylation sequence at the 3’ end of the gene ( Figure 4 ).
[0118] GS - / - SHMT2 - / - Cells were cultured in Fusion+Gln+Gly+For under shaking conditions at 37 °C with 5% CO2. Cells were split at 0.3e6 three days before transfection. 1.0e6 cells / condition were transfected with 1 μg of plasmid DNA using Lonza's 4D-Nucleofector system with program DT-133 and SF Nucleofector solution. The transfected cells were transferred into a 6-well culture flask containing 3 mL of pre-warmed EX- CD CHO Fusion medium (MilliporeSigma 14365C) supplemented with 6 mM L-glutamine (MilliporeSigma G7513) and 200 μM formate (Sigma 456020-25G). Cells were incubated at 37 °C / 5% CO2 in a static environment. Three days post transfection, each sample was amplified to T-25 and 2 mL of medium was added. At day 7, cells reached ~90% viability and were transferred to a 15 mL conical tube and spun down at 1000 rpm (~300 x g) for 5 minutes. The supernatant was aspirated, then the cells were washed with 5 mL of PBS and spun down again for 5 minutes. The supernatant was aspirated and the cells were resuspended in 10-15 mL of their respective selection medium and transferred to a T-75 flask. Glutamine-based selection was performed using Fusion-Gln, glycine-formate-based selection was performed using Fusion-Gly-For, and glutamine / glycine-formate dual selection was performed using Fusion without Gln, Gly, and For. The cell viability and live cell density of the various selection cultures were monitored over time.
[0119] EX- Advanced CHO Fed-batch medium (MilliporeSigma 14366C) and EX- Advanced CHO Feed (custom formulations of the Glycine-deficient (Advanced-Gly and Feed-Gly, respectively) of MilliporeSigma 24367C). Using unmodified 4Feed (MilliporeSigma 1.03796.0005). A stable selection culture transfected with an IgG1 expression vector was pelleted, where the selection medium was aspirated and then resuspended in Advanced-Gly at 3e5 viable cells / mL for productivity analysis under fed-batch conditions. Starting on day 3 post-inoculation, viable cell density and viability were collected every other day for each culture. Starting on day 3 post-inoculation, 1.5 mL of a 50 / 50 blend of Advanced Feed-Gly and 4Feed was added to each culture. Starting on day 5, glucose readings were taken from each culture every other day, where D-+-glucose (MilliporeSigma G8769) was added to maintain an appropriate glucose level. Productivity was monitored over time, where fed-batch titers were recorded every other day starting on day 8 until the culture dropped below 70% viability. Titers were determined using interferometry on a ForteBio Octet and subsequently confirmed via HPLC protein A affinity chromatography.
[0120] Example 2:
[0121] A custom formulation (Fusion-Gln-Gly) of a basal EX- CD CHO Fusion medium without glycine (Gly) was developed. GS - / - SHMT2 - / - Clones were cultured in Fusion+Gln+Gly+For or Fusion+Gln-Gly-For for at least ten days. Viability and viable cell density measurements were taken twice a week. Figure 5 It was indicated that in the absence of glycine and formate, GS - / - SHMT2 - / - cells were unable to grow. However, when glycine was supplemented into the medium, GS - / - SHMT2 - / - cell growth was rescued. Importantly, SHMT2 - / - cell growth rate was significantly slower than that of SHMT2 + / + cell line, although this could be rescued by adding sodium formate to the medium ( Figure 9 ). Addition of formate alone (Fusion-Gly+For) was not sufficient for cell survival (Figure 10 )。
[0122] Example 3:
[0123] To confirm that a stably selected cell population can produce a target protein using the Gly / For-mediated selection system described in Example 1, cells were transfected and passaged in Fusion-Gly-For under selection pressure. The SHMT2 / CFP vector was transfected into GS - / - SHMT2 - / - the cells. Cell growth and viability were monitored throughout the selection. Cells transfected with the SHMT2 / CFP plasmid and grown in Fusion+Gln-Gly-For (Gly selective medium) were recovered from the selection pressure. The surviving cell population was analyzed by FACS, and the mean fluorescence intensity (MFI) and the percentage of CFP+ cells were measured. Cells grown in Fusion+Gln+Gly+For showed a very small percentage of CFP-positive cells and a low MFI compared to cells subjected to Gly / For selection in Fusion+Gln-Gly-For. The results are summarized in Figure 6 and 8 .
[0124] Example 4:
[0125] To confirm that a stably selected cell population can produce a target protein using the Gly For-mediated selection system described in Example 1, we developed a vector with the IgG heavy chain, IgG light chain, and serine hydroxymethyltransferase 2 (SHMT2) coding sequences. The vector was transfected into GS - / - SHMT2 - / - the cell line. As a control, mock transfection without DNA was used. The population was passaged in Fusion+Gln-Gly-For under selection pressure. The conditions for selection were also applied during recovery, scale-up, and productivity determination. The fed-batch productivity assay was inoculated at 3e5 viable cells / mL in Advanced+Gln-Gly-For medium. Starting from day 3 post-inoculation, viable cell density and viability for each culture were collected every other day. Starting from day 3 post-inoculation, a 50 / 50 blend of 1.5 mL of AdvancedFeed-Gly-For and 4Feed-Gly-For was added to each culture. Starting from day 5, glucose readings were taken every other day, with D-+-glucose (MilliporeSigma G8769) added to maintain an appropriate glucose level. Cells transfected with only the SHMT2 vector reached a peak viable cell density of ~12.5e6 cells / mL (Figure 11 (left middle curve), and maintained >70% viability for at least 13 days ( Figure 11 (right middle curve). The cell titer from the fed-batch reached a peak of ~200 mg / L ( Figure 12 (left middle curve and right right curve).
[0126] Example 5:
[0127] To test whether a stable cell population can produce two independent intracellular fluorescent proteins in the absence of glutamine and glycine / formate, we developed two vectors, one containing the GFP and GS coding sequences, while the second contained the CFP and SHMT2 coding sequences ( Figure 4 ). These two plasmids were co-transfected into GS - / - SHMT2 - / - cells (GFP+CFP). As a control, each vector was also transfected independently into GS - / - SHMT2 - / - cells (GFP only and CFP only, respectively). Cells from all three transfections were then passaged under GS selection conditions (Fusion-Gln), SHMT2 selection conditions (Fusion-Gly-For), and dual metabolic selection conditions (Fusion-Gln-Gly-For). The conditions used for selection were also applied during recovery, scale-up, and all other assays. Figure 8 Viability data from the selection assays are shown, indicating that cells transfected with the GFP vector can survive and grow under -Gln conditions, but require Gly / For supplementation in the medium. On the other hand, cells transfected with the CFP vector can survive and grow under -Gly conditions, but require Gln supplementation in the medium. Cells co-transfected with both vectors (GFP+CFP) can survive and grow in -Gln-Gly-For medium. Figure 6 and Figure 8 indicate that cells transfected with the GFP vector that survive and grow in -Gln medium are positive for GFP, cells transfected with the CFP vector that survive and grow in -Gly-formate medium are positive for CFP, and cells co-transfected with both vectors (GFP+CFP) that survive and grow in -Gln-Gly-For medium are positive for both GFP and CFP. This data indicates that the GS+SHMT2 dual metabolic selection system provides a unique opportunity to select cells in which multiple independent vectors encoding intracellular proteins have been introduced, without the addition of any selection agents, such as antibiotics, to the medium.
[0128] Example 6:
[0129] To test whether stable cell populations are able to produce secreted proteins in the absence of glutamine and glycine, we developed two vectors expressing IgG1, one containing the IgG heavy chain, IgG light chain, and GS coding sequence, while the second contained the same IgG heavy chain, IgG light chain, and SHMT2 coding sequence. These two independent vectors were co-transfected into GS - / - SHMT2 - / - cells (GS+SHMT2). As a control, each vector was also transfected independently into GS - / - SHMT2 - / - (GS only and SHMT2 only). Cells were then passaged under selection pressure in Fusion-Gln (GS-only transfected cells), Fusion-Gly-For (SHMT2-only transfected cells), or Fusion-Gln-Gly-For (GS+SHMT2 transfected cells). The conditions used for selection were also applied during recovery, scale-up, and productivity assays. The GS-only selection cultures recovered completely after 14 - 19 days, while the SHMT2-only and GS+SHMT2 double-selection cultures had similar selection recovery curves and required 19 days to recover completely. In fed-batch assays, GS-only and SHMT2-only clones produced similar levels of IgG, while GS+SHMT2 cells produced significantly more IgG and had the highest growth and viability. This suggests that, following expression of exogenous GS and / or SHMT2 coding sequences, GS and SHMT2 produced by the cells are sufficient to express secreted proteins ( Figure 11 right graph and Figure 12 left graph). This provides the potential for running large-scale production bioreactors under dual metabolic selection conditions, which may be difficult to perform using antibiotic selection methods, either due to the need to subsequently separate or purify the antibiotic from the desired secreted protein, or due to the cost of adding antibiotics to large-scale bioreactors. Additionally, this dual metabolic selection system (GS+SHMT2) provides the opportunity to more effectively select cells into which multiple large vectors have been introduced, such as when expressing bispecific antibodies or another large and / or complex protein.
Claims
1. A method for producing a recombinant protein product, the method comprising (a) providing a mammalian cell line engineered to reduce or eliminate the expression of endogenous serine hydroxymethyltransferase 2 (SHMT2); (b) introducing a polynucleotide into the mammalian cell line, wherein the polynucleotide encodes a functional SHMT2 gene and a recombinant protein; (c) culturing the cell line; and (d) purifying the recombinant protein to form a recombinant protein product.
2. The method according to claim 1, wherein the mammalian cell line of (a) further comprises a reduced or eliminated expression of endogenous glutamine synthetase (GS), phosphoserine phosphatase (PSPH), dihydrofolate reductase (DHFR), P5C synthase (P5CS), asparaginase (ASPG), alanine transaminase (ALT), and / or asparagine synthetase (ASNS).
3. The method according to claim 1, wherein the endogenous SHMT2 expression is reduced or eliminated by inactivating the endogenous SHMT2 gene of the mammalian cell line.
4. The method according to claim 1, wherein the endogenous SHMT2 gene is inactivated using a targeted endonuclease-mediated genome modification technique.
5. The method according to claim 4, wherein the targeted endonuclease is a CRISPR ribonucleoprotein complex or a pair of zinc finger nucleases.
6. The method according to claim 1, wherein the mammalian cell line is a Chinese hamster ovary (CHO) cell line, a baby hamster kidney (BHK) cell line, an NS0 mouse myeloma cell line, a HEK293 cell line, or a Vero African green monkey kidney cell line.
7. The method according to any one of claims 1 to 6, wherein the cell line is a CHO cell line cultured with or without glycine and / or formate.
8. The method according to any one of claims 1 to 7, wherein the recombinant protein product is selected from antibodies, antibody fragments, vaccines, growth factors, cytokines, hormones, or blood coagulation factors.
9. The method according to claim 8, wherein the antibody is a bispecific or multispecific antibody.
10. A mammalian cell line for genetic engineering in a bioproduction system, wherein the mammalian cell line is engineered to reduce or eliminate the expression of endogenous SHMT2.
11. The mammalian cell line according to claim 10, wherein the expression of SHMT2 is reduced or eliminated via the inactivation of at least one allele of the chromosomal sequence encoding SHMT2.
12. The mammalian cell line according to claim 11, wherein one or more alleles of the chromosomal sequence encoding SHMT2 are inactivated.
13. The mammalian cell line according to claim 10, wherein the cell line is engineered to reduce or eliminate the expression of endogenous glutamine synthetase (GS), phosphoserine phosphatase (PSPH), dihydrofolate reductase (DHFR), P5C synthase (P5CS), asparaginase (ASPG), alanine transaminase (ALT), and / or asparagine synthetase (ASNS).
14. The mammalian cell line according to claim 12, wherein the chromosomal sequence is inactivated using a targeted endonuclease-mediated genome modification technique.
15. The mammalian cell line according to claim 14, wherein the targeted endonuclease is a ribonucleoprotein complex or a pair of zinc finger nucleases.
16. The mammalian cell line according to claim 15, wherein the non-human cell line is a Chinese hamster ovary (CHO) cell line, baby hamster kidney (BHK) cell line, NS0 mouse myeloma cell line, HEK293 cell line, or Vero African green monkey kidney cell line.
17. The mammalian cell line according to claim 16, wherein the cell line is a CHO cell line cultured with or without glycine and / or formate.
18. The mammalian cell line according to claim 17, wherein the cell viability, viable cell density, titer, growth rate, proliferation response, cell morphology, and / or general cell health are comparable to those of the non-engineered parental mammalian cell line.
19. The mammalian cell line according to any one of claims 10 to 18, which further comprises at least one nucleic acid encoding a recombinant protein selected from the group consisting of an antibody, antibody fragment, vaccine, growth factor, cytokine, hormone, or blood coagulation factor.
20. The mammalian cell line according to claim 19, wherein the antibody is a bispecific or multispecific antibody.
21. A polynucleotide comprising a nucleic acid sequence encoding a functional SHMT2 and at least one target recombinant protein.
22. A polynucleotide comprising the following: a) a nucleic acid sequence encoding a functional SHMT2; b) a nucleic acid sequence encoding a functional GS and / or ASNS; and c) a nucleic acid sequence encoding a mutation in the ASNS and / or SHMT2 coding sequence, the mutation weakening the activity of one or both of the enzymes; d) a nucleic acid sequence encoding a target recombinant protein.
Citation Information
Patent Citations
In vitro peptide or protein expression library
GB2338237A
Functional domains in flavobacterium okeanokoites (FokI) restriction endonuclease
US5356802A
Functional domains in flavobacterium okeanokoities (foki) restriction endonuclease
US5436150A
Insertion and deletion mutants of FokI restriction endonuclease
US5487994A
Zinc finger proteins with high affinity new DNA binding specificities
US5789538A