Metabolic selection by serine biosynthetic pathway
By reducing or eliminating the expression of PSPH and GS in mammalian cell lines, combined with CRISPR/RNP technology and serine-mediated selection, the problem of difficulty in efficiently expressing multiple recombinant proteins in the prior art is solved, and efficient and simplified recombinant protein production is achieved.
Patent Information
- Application Number
- CN202380082136.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-30
- Filing Date
- 2023-09-29
- Publication Date
- 2025-07-08
AI Technical Summary
In existing biomanufacturing technologies, it is difficult to efficiently express multiple recombinant proteins through a single selection system, especially bispecific antibodies and multispecific antibodies, and traditional selection methods require the introduction of multiple vectors into cells, which increases operational complexity and cost.
By engineering mammalian cell lines to reduce or eliminate expression of endogenous phosphoserine phosphatase (PSPH) and glutamine synthase (GS), genomic modifications using CRISPR ribonucleoprotein (RNP) complex or zinc finger nuclease, combined with a serine-mediated selection system, a dual metabolic selection system was developed to express recombinant proteins.
It realizes efficient expression of bispecific and multispecific antibodies without the need for exogenous selective agents, simplifies the operation process, reduces costs, and improves production efficiency.
Smart Images

Figure CN120283057A_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 377,876, filed on September 30, 2022, the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present disclosure relates to mammalian cell lines for use in bioproduction systems, wherein the mammalian cell lines are engineered to have reduced or eliminated expression of components of the serine biosynthetic pathway to produce a cell line that is unable to proliferate in the absence of exogenously provided serine or heterologously expressed coding sequences necessary for serine biosynthesis. Background Art
[0003] The development of high-producing clonal cell lines for biomanufacturing typically utilizes one or more well-known selection methods, such as glutamine synthetase (GS, for glutamine selection), dihydrofolate receptor (DHFR, for hypoxanthine and thymidine selection), antibiotic selection (puromycin, hygromycin, blasticidin, etc.), or P5C synthetase (P5CS-proline selection). The GS system has become standard in the industry, but there is a need for cell lines that allow for multiple selection methods so that more than one vector can be introduced into the cell line to facilitate the production of molecules such as bispecific antibodies, multispecific antibodies, and other multi-chain enzymes / proteins or proteins / enzymes that require effector proteins for expression. Summary of the Invention
[0004] Among the various aspects of the present disclosure, mammalian cell lines are provided for use in bioproduction systems, wherein the mammalian cell lines are engineered to have reduced or eliminated expression of an endogenous phosphoserine phosphatase (PSPH) gene. In the absence of an endogenously expressed functional PSPH protein, the cells require an exogenous source of the amino acid serine. Targeted endonuclease-mediated genome modification, such as CRISPR ribonucleoprotein (RNP) complexes or zinc finger nucleases, can be used to inactivate chromosomal PSPH sequences. In other aspects of the present disclosure, mammalian cell lines are provided, wherein the mammalian cell lines are engineered to have reduced or eliminated expression of an endogenous PSPH gene and reduced or eliminated expression of an endogenous glutamine synthetase (GS) gene.
[0005] Another aspect of the present disclosure includes methods for selecting cell lines with enhanced productivity of expressed biotherapeutic proteins. In other aspects of the present disclosure, a bioproduction system for expressing bispecific antibodies or biotherapeutic proteins is provided, wherein the bispecific antibodies or biotherapeutic proteins require more convenient expression of effector proteins by utilizing multiple selection systems. The method comprises expressing at least one recombinant protein in any mammalian cell line.
[0006] Other aspects and iterations of the disclosure are described in more detail below. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 The final step in the de novo serine synthesis pathway catalyzed by phosphoserine phosphatase (PSPH) is shown. (1) Yang, M., Vousden, KH, Serine and one-carbon metabolism in cancer. Nature. 16, 650-660 (2016).
[0008] Figure 2 The PSPH cDNA sequence in CHO is shown. The gRNA target site is bold and underlined.
[0009] Figure 3 Shown is copy number analysis of endogenous phosphoserine phosphatase genes by ddPCR.
[0010] Figure 4 Figure 2 shows Cas9 cleavage activity assessed by supplying Ser to the culture medium. The percentage of edited PSPH within the RNP transfected pool was determined by NGS. Supplying 10 mM Ser resulted in the best cleavage efficiency.
[0011] Figure 5 Shown are cloned KO alleles generated by CRISPR-Cas9 targeting. All of the above indel mutations (indels) produce early stop codons in the coding sequence. The cloned PSPH KO genotype was confirmed by NGS. The KO indel mutation and frequency are highlighted in bold.
[0012] Figure 6 Shown is how to design vectors to allow selection of GFP-positive cells using a glutamine-based selection system, selection of BFP-positive cells using a serine-based selection system, and expression of secreted recombinant proteins by developing two similar vectors containing mAb heavy and light chains and either PSPH or GS coding sequences. The PSPH mAb vector map is shown as an example.
[0013] Figure 7 PSPH KO clones showed greater growth sensitivity to serine starvation and lower maximal VCD than parental PSPH+ / + controls.
[0014] Figure 8 Shown are the growth and viability of pools co-expressing GFP+BFP selected using GS+PSPH.
[0015] Figure 9It is shown that CHO cells with genetically disrupted GS and PSPH genes co-transfected with PSPH+BFP and GS+GFP were selected with medium lacking both glutamine and serine, and expressed both GFP and BFP.
[0016] Figure 10 Shown are CHO cells with genetically disrupted GS and PSPH genes transfected with GS-IgG or PSPH-IgG, as well as co-transfected with GS-IgG and PSPH IgG, and selected with their respective selection media. Shown are bulk selected pool titers. DETAILED DESCRIPTION
[0017] The present disclosure provides mammalian cell lines engineered to have reduced or eliminated expression of an endogenous PSPH gene. Further provided are mammalian cell lines engineered to have reduced or eliminated expression of an endogenous GS gene and reduced or eliminated expression of an endogenous PSPH gene. Methods for generating the engineered cell lines are provided, as well as methods for selecting and using the engineered cell lines to produce recombinant proteins. (I) Engineered cell lines
[0018] One aspect of the present disclosure includes a mammalian cell line engineered to have reduced or eliminated expression of an endogenous PSPH gene. Alternatively, the mammalian cell line is engineered to have reduced or eliminated expression of an endogenous PSPH gene and an endogenous GS gene.
[0019] The cell lines disclosed herein with reduced or eliminated expression of PSPH or reduced expression of PSPH and GS are genetically engineered to modify the chromosomal sequence encoding the PSPH or GS protein. Targeted endonuclease-mediated genome editing techniques can be used to modify the chromosomal sequence, which is described in detail in section (III) below. For example, the chromosomal sequence can be modified to include a deletion of at least one nucleotide, an insertion of at least one nucleotide, a substitution of at least one nucleotide, or a combination thereof, so that the reading frame is shifted and no protein product is produced (i.e., chromosomal sequence inactivation). Inactivation of one allele of the chromosomal sequence encoding PSPH or GS results in reduced expression of the protein (i.e., knocking down). Inactivation of two alleles of the chromosomal sequence encoding PSPH or GS results in non-expression of the protein (i.e., knocking out).
[0020] In some embodiments, the expression level of PSPH can be reduced by at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or more than about 99%. In other embodiments, the expression level of PSPH can be reduced to an undetectable level using standard techniques in the art (e.g., Western blotting assays, ELISA enzyme assays, SDS polyacrylamide gel electrophoresis, etc.).
[0021] Generally, when supplied with serine and / or exogenous PSPH coding sequences, the engineered cell lines disclosed herein have cell viability, viable cell density, titer, growth rate, proliferative response, cell morphology, apoptosis and autophagy levels, and / or overall cell health similar to those of their unengineered parental cells. (a) Cell type
[0022] The engineered cell lines disclosed herein are mammalian cell lines. In some embodiments, the engineered cell lines can be derived from human cell lines. Non-limiting examples of suitable human cell lines include human embryonic kidney cells (HEK293, HEK293T); human connective tissue cells (HT-1080); human cervical cancer cells (HELA); human embryonic retinal cells (PER.C6); human kidney cells (HKB-11); human hepatocytes (Huh-7); human lung cells (W138); human hepatocytes (HepG2); human U2-OS osteosarcoma cells, human A549 lung cells, human A-431 epidermal cells, CACO-2 human colorectal adenocarcinoma cells, human pluripotent stem cells, Jurkat human T lymphocytes, or human K562 bone marrow cells. In other embodiments, the engineered cell lines can be derived from non-human cell lines. Suitable cell lines also include Chinese hamster ovary (CHO) cells; baby hamster kidney (BHK) cells; mouse myeloma NS0 cells; mouse myeloma Sp2 / 0 cells; mouse mammary gland C127 cells; mouse embryonic fibroblast 3T3 cells (NIH3T3); mouse B lymphoma A20 cells; mouse melanoma B16 cells; mouse myoblast C2C12 cells; mouse embryonic mesenchymal C3H-10T1 / 2 cells; mouse carcinoma CT26 cells; mouse prostate DuCuP cells; mouse mammary gland EMT6 cells; mouse hepatoma Hepa1c1c7 cells; mouse myeloma J5582 cells; mouse epithelial MTD-1A cells; mouse myocardial MyEnd cells; mouse kidney RenCa cells; mouse pancreatic RIN-5F cells; mouse melanoma X64 cells; mouse lymphoma YAC-1 cells; rat glioblastoma 9L cells; rat B lymphoma RBL cells; rat neuroblastoma B35 cells; rat hepatoma (HTC); buffalo rat liver BRL3A cells; canine kidney cells (MDCK); canine mammary gland (CMT) cells; rat osteosarcoma D17 cells; rat monocyte / macrophage DH82 cells; monkey kidney SV-40 transformed fibroblasts (COS7); monkey kidney CVI-76 cells; or African green monkey kidney (VERO, VERO-76) cells. An expanded list of mammalian cell lines can be found in the American Type Culture Collection catalog (ATCC, Manassas, VA). In some embodiments, the cell lines disclosed herein are cell lines other than mouse cell lines. In certain embodiments, the engineered cell line is a CHO cell line. Suitable CHO cell lines include, but are not limited to, CHO-K1, CHO-K1SV, CHO GS- / -, CHO S, DG44, DuxB11, and derivatives thereof.
[0023] In various embodiments, the parental cell line can be deficient in glutamine synthetase (GS), dihydrofolate reductase (DHFR), hypoxanthine-guanine phosphoribosyltransferase (HPRT), asparagine synthetase (ASNS), or a combination thereof. For example, chromosomal sequences encoding GS, DHFR, HPRT, and / or ASNS can be inactivated. In specific embodiments, all chromosomal sequences encoding GS, DHFR, HPRT, and / or ASNS are inactivated in the parental cell line. (b) Optional nucleic acid encoding a recombinant protein
[0024] In some embodiments, the engineered cell lines disclosed herein may further include at least one nucleic acid encoding a recombinant protein. Typically, the recombinant protein is heterologous, which means that the protein is not natural to the cell. The recombinant protein can be, but is not limited to, a therapeutic protein selected from an antibody, an antibody fragment, a monoclonal antibody, a humanized antibody, a humanized monoclonal antibody, a chimeric antibody, an IgG molecule, an IgG heavy chain, an IgG light chain, an IgA molecule, an IgD molecule, an IgE molecule, an IgM molecule, a vaccine, a growth factor, a cytokine, an interferon, an interleukin, a hormone, a coagulation (clotting) (or coagulation (coagulation)) factor, a blood component, an enzyme, a therapeutic protein, a nutritional protein, a functional fragment or functional variant of any of the foregoing proteins, or a fusion protein comprising any of the foregoing proteins and / or its functional fragment or variant. In a specific embodiment, the recombinant protein is a bispecific antibody or a multispecific antibody, or requires an effector protein to express the protein.
[0025] In some embodiments, the nucleic acid encoding the recombinant protein can be linked to a sequence encoding phosphoserine phosphatase (PSPH), hypoxanthine-guanine phosphoribosyltransferase (HPRT), dihydrofolate reductase (DHFR) and / or glutamine synthetase (GS), such that PSPH, ASNS, HPRT, DHFR and / or GS can be used as a selective marker. The nucleic acid encoding the recombinant protein can also be linked to a sequence encoding at least one antibiotic resistance gene and / or a sequence encoding a marker protein, such as a fluorescent protein. In some embodiments, the nucleic acid encoding the recombinant protein can be part of an expression construct. The expression construct or vector can contain additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), a selective marker sequence, an origin of replication, and the like. Additional information can be found in “Current Protocols in Molecular Biology” Ausubel et al., John Wiley & Sons, New York, 2003 or “Molecular Cloning: A Laboratory Manual” Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001.
[0026] In some embodiments, the nucleic acid encoding the recombinant protein can be located outside the chromosome. That is to say, the nucleic acid encoding the recombinant protein can be transiently expressed from a plasmid, a cosmid, an artificial chromosome, a minichromosome or another extrachromosomal construct. In other embodiments, the nucleic acid encoding the recombinant protein can be chromosomally integrated into the genome of the cell. Integration can be random or targeted. Therefore, the recombinant protein can be stably expressed. In some iterations of this embodiment, the nucleic acid sequence encoding the recombinant protein can be operably linked to a suitable heterologous expression control sequence (i.e., a promoter). In other iterations, the nucleic acid sequence encoding the recombinant protein can be placed under the control of an endogenous expression control sequence. The nucleic acid sequence encoding the recombinant protein can be integrated into the genome of the cell line using homologous recombination, targeted endonuclease-mediated genome editing, viral vectors, transposons, recombinase-mediated cassette exchange systems, plasmids and other known means. Additional guidance can be found in Ausubel et al., 2003, supra, and Sambrook & Russell, 2001, supra. (II) Kit
[0027] Another aspect of the present disclosure provides a kit for producing a recombinant protein, wherein the kit includes any of the engineered cell lines described in detail in section (I) above. The kit may further include cell growth medium, transfection reagents, plasmid vectors, selective culture medium, recombinant protein purification means, buffers, etc. The kit provided herein typically includes instructions for growing cell lines and using them to produce recombinant proteins. The instructions included in the kit may be attached to the packaging material or may be included as a package insert. Although the instructions are typically written or printed materials, they are not limited thereto. The present disclosure contemplates any medium capable of storing such instructions and communicating them to the end user. Such media include, but are not limited to, electronic storage media (e.g., disks, tapes, boxes, chips), optical media (e.g., CD ROMs), etc. As used herein, the term "instructions" may include the address of an internet site providing the instructions. (III) Method for preparing engineered cell lines
[0028] Yet another aspect of the present disclosure provides methods for preparing or engineering cell lines described in section (I) above that have reduced or eliminated expression of PSPH and / or GS. The chromosomal sequences encoding PSPH and / or GS can be knocked down or knocked out using a variety of techniques. Typically, engineered cell lines are prepared using targeted endonuclease-mediated genomic modification methods. Those skilled in the art will appreciate that the engineered cell lines can also be prepared using site-specific recombination systems, random mutagenesis, or other methods known in the art.
[0029] Typically, engineered cell lines are prepared by methods including the following: at least one targeted endonuclease or a nucleic acid encoding the targeted endonuclease is introduced into the parental cell line of interest, wherein the targeted endonuclease targets the chromosomal sequence encoding PSPH and / or GS. The targeted endonuclease recognizes and binds to a specific chromosomal sequence and introduces a double-strand break. In some embodiments, double-strand breaks are repaired by non-homologous end joining (NHEJ) repair process. Because NHEJ is fallible, deletion, insertion and / or displacement of at least one nucleotide can occur, thereby destroying the reading frame of the chromosomal sequence so that no protein product is produced, or non-functional protein is produced by destroying, for example, the enzyme active site of the protein. In other embodiments, the targeted endonuclease can also be used for changing the chromosomal sequence via homologous recombination reactions by co-introducing a polynucleotide having a substantial sequence identity with a part of the targeted chromosomal sequence. In such cases, the double-strand break introduced by the targeted endonuclease is repaired by a homology-directed repair process so that the chromosomal sequence is exchanged with the polynucleotide in a manner that causes the chromosomal sequence to change or change (for example, by the integration of an exogenous sequence). (a) Targeted endonucleases
[0030] A variety of targeting endonucleases can be used to modify chromosomal sequences encoding PSPH and / or GS. The targeting endonuclease can be a naturally occurring protein or an engineered protein. Suitable targeting endonucleases include, but are not limited to, zinc finger nucleases (ZFNs), CRISPR nucleases, transcription activator-like effector (TALE) nucleases (TALENs), meganucleases, chimeric nucleases, site-specific endonucleases, and artificially targeted DNA double-strand break inducers. (i) Zinc finger nucleases
[0031] In certain embodiments, the targeted endonuclease can be a zinc finger nuclease (ZFN) pair. ZFNs bind to specific targeted sequences and introduce double-strand breaks into the targeted cleavage site. Typically, ZFNs comprise a DNA binding domain (i.e., zinc finger) and a cleavage domain (i.e., nuclease), each of which is described below.
[0032] DNA binding domainDNA binding domains or zinc fingers can be engineered to recognize and bind to any selected nucleic acid sequence. See, for example, Beerli et al. (2002) Nat. Biotechnol. 20:135-141; Pabo et al. (2001) Ann. Rev. Biochem. 70:313-340; Isalan et al. (2001) Nat. Biotechnol. 19:656-660; Segal et al. (2001) Curr. Opin. Biotechnol. 12:632-637. ; Choo et al. (2000) Curr. Opin. Struct. Biol. 10:411-416; Zhang et al. (2000) J. Biol. Chem. 275(43):33850-33860; Doyon et al. (2008) Nat. Biotechnol. 26:702-708; and Santiago et al. (2008) Proc. Natl. Acad. Sci. USA 105:5809-5814. Engineered zinc finger binding domains can have new binding specificities compared to naturally occurring zinc finger proteins. Engineering methods include, but are not limited to, rational design and various types of selection. Rational design includes, for example, the use of a database comprising doublet, triplet, and / or quadruple nucleotide sequences and individual zinc finger amino acid sequences, wherein each doublet, triplet, or quadruple nucleotide sequence is associated with one or more amino acid sequences of a zinc finger that binds to a specific triplet or quadruple sequence. See, for example, U.S. Patent Nos. 6,453,242 and 6,534,261, the disclosures of which are incorporated herein by reference in their entirety. For example, the algorithm described in U.S. Patent No. 6,453,242 can be used to design zinc finger binding domains targeting preselected sequences. Alternative methods can also be used, such as the rational design of nondegenerate recognition codes to design zinc finger binding domains targeting specific sequences (SERA et al. (2002) Biochemistry 41:7074-7081). Publicly available web-based tools for identifying potential target sites in DNA sequences and designing zinc finger binding domains are known in the art. For example, tools for identifying potential target sites in DNA sequences can be found at zincfingertools.org. Tools for designing zinc finger binding domains can be found at zifit.partners.org / ZiFiT. (See also Mandell et al. (2006) Nuc. Acid Res. 34:W516-W523; Sander et al. (2007) Nuc. Acid Res. 35:W599-W605.)
[0033] The zinc finger binding domain can be designed to recognize and bind to a DNA sequence ranging in length from about 3 nucleotides to about 21 nucleotides. In one embodiment, the zinc finger binding domain can be designed to recognize and bind to a DNA sequence ranging in length from about 9 to about 18 nucleotides. Typically, the zinc finger binding domain of the zinc finger nuclease used herein comprises at least three zinc finger recognition regions or zinc fingers, wherein each zinc finger binds to 3 nucleotides. In one embodiment, the zinc finger binding domain comprises four zinc finger recognition regions. In another embodiment, the zinc finger binding domain comprises five zinc finger recognition regions. In yet another embodiment, the zinc finger binding domain comprises six zinc finger recognition regions. The zinc finger binding domain can be designed to bind to any suitable target DNA sequence. See, for example, U.S. Patent Nos. 6,607,882; 6,534,261 and 6,453,242, the disclosures of which are incorporated herein by reference in their entirety.
[0034] Exemplary methods for selecting zinc finger recognition regions include phage display and two-hybrid systems, which are described in U.S. Patent Nos. 5,789,538; 5,925,523; 6,007,988; 6,013,453; 6,410,248; 6,140,466; 6,200,759; 6,242,568; and WO 98 / 37186; WO 98 / 53057; WO 00 / 27878; WO 01 / 88197 and GB 2,338,237, each of which is incorporated herein by reference in its entirety. In addition, enhancement of the binding specificity of zinc finger binding domains has been described, for example, in WO 02 / 077227, the entire disclosure of which is incorporated herein by reference.
[0035] Zinc finger binding domains and methods for designing and constructing fusion proteins (and polynucleotides encoding them) are known to those skilled in the art and are described in detail, for example, in U.S. Patent No. 7,888,121, which is incorporated herein by reference in its entirety. Zinc finger recognition regions and / or multi-finger zinc finger proteins can be linked together using suitable linker sequences, including, for example, linkers having a length of five or more amino acids. See U.S. Patent Nos. 6,479,626; 6,903,185; and 7,153,949 (the disclosures of which are incorporated herein by reference in their entirety) for non-limiting examples of linker sequences having a length of six or more amino acids. The zinc finger binding domains described herein may include a combination of suitable linkers between the individual zinc fingers of a protein.
[0036] Cleavage domain.Zinc finger nucleases also include cleavage domains. The cleavage domain portion of a zinc finger nuclease can be obtained from any endonuclease or exonuclease. Non-limiting examples of endonucleases from which a cleavage domain can be derived include, but are not limited to, restriction endonucleases and homing endonucleases. See, for example, New England Biolabs Catalog or Belfort et al. (1997) Nucleic Acids Res. 25: 3379-3388. Additional enzymes for cutting DNA are known (e.g., S1 nuclease; mung bean nuclease; pancreatic DNase I; micrococcal nuclease; yeast HO endonuclease). See also Linn et al. (eds.) Nucleases, Cold Spring Harbor Laboratory Press, 1993. One or more of these enzymes (or their functional fragments) can be used as a source of a cleavage domain.
[0037] The cleavage domain can also be derived from an enzyme as described above or a portion thereof that requires dimerization to have cleavage activity. Two zinc finger nucleases may be required for cleavage because each nuclease comprises a monomer of an active enzyme dimer. Alternatively, a single zinc finger nuclease may comprise two monomers to produce an active enzyme dimer. As used herein, an "active enzyme dimer" is an enzyme dimer that can cut a nucleic acid molecule. Two cleavage monomers may be derived from the same endonuclease (or its functional fragment), or each monomer may be derived from a different endonuclease (or its functional fragment).
[0038] When two cutting monomers are used to form an active enzyme dimer, the recognition site of the two zinc fingers is preferably arranged so that the combination of the two zinc fingers and their respective recognition sites places the cutting monomer in a spatial orientation relative to each other, which allows the cutting monomer to form an active enzyme dimer, for example, by dimerization. As a result, the near edge of the recognition site can be separated by about 5 to about 18 nucleotides. For example, the near edge can be separated by about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17 or 18 nucleotides. However, it should be understood that any integer number of nucleotides or nucleotide pairs (for example, about 2 to about 50 nucleotide pairs or more) can be inserted between the two recognition sites. The near edge of the recognition site of the zinc finger nuclease, such as those described in detail herein, can be separated by 6 nucleotides. Typically, the cutting site is located between the recognition sites.
[0039] Restriction endonucleases (restriction enzymes) are present in many species and are capable of sequence-specifically binding to DNA (at a recognition site) and cleaving the DNA at or near the site of binding. Certain restriction enzymes (e.g., Type IIS) cleave DNA at sites away from the recognition site and have separable binding and cleavage domains. For example, the Type IIS enzyme FokI catalyzes double-stranded cleavage of DNA at 9 nucleotides from its recognition site on one chain and 13 nucleotides from its recognition site on the other chain. See, e.g., U.S. Patent Nos. 5,356,802; 5,436,150 and 5,487,994; and Li et al. (1992) Proc. Natl. Acad. Sci. USA 89:4275-4279; Li et al. (1993) Proc. Natl. Acad. Sci. USA 90:2764-2768; Kim et al. (1994a) Proc. Natl. Acad. Sci. USA 91:883-887; Kim et al. (1994b) J. Biol. Chem. 269:31978-31982. Thus, a zinc finger nuclease can comprise a cleavage domain from at least one Type IIS restriction enzyme and one or more zinc finger binding domains, which can be engineered or unengineered. Exemplary Type IIS restriction enzymes are described, for example, in International Publication WO 07 / 014,275, the disclosure of which is incorporated herein by reference in its entirety. Additional restriction enzymes also contain separable binding and cleavage domains, and these are also contemplated by the present disclosure. See, for example, Roberts et al. (2003) Nucleic Acids Res. 31:418-420.
[0040] An exemplary type IIS restriction enzyme whose cleavage domain can be separated from the binding domain is FokI. This particular enzyme is active as a dimer (Bitinite et al. (1998) Proc. Natl. Acad. Sci. USA 95: 10,570-10,575). Therefore, for the purposes of this disclosure, the portion of the FokI enzyme used in the zinc finger nuclease is considered to be a cleavage monomer. Therefore, for targeted double-stranded cleavage using the FokI cleavage domain, two zinc finger nucleases (each of which contains a FokI cleavage monomer) can be used to reconstitute an active enzyme dimer. Alternatively, a single polypeptide molecule containing a zinc finger binding domain and two FokI cleavage monomers can also be used.
[0041] In certain embodiments, the cleavage domain comprises one or more engineered cleavage monomers that minimize or prevent homodimerization. As non-limiting examples, amino acid residues at positions 446, 447, 479, 483, 484, 486, 487, 490, 491, 496, 498, 499, 500, 531, 534, 537, and 538 of FokI are all targets for affecting dimerization of the FokI cleavage half-domain. Exemplary engineered cleavage monomers of FokI that form obligate heterodimers include pairs in which the first cleavage monomer comprises mutations at amino acid residue positions 490 and 538 of FokI and the second cleavage monomer comprises mutations at amino acid residue positions 486 and 499.
[0042] Thus, in one embodiment of the engineered cleavage monomer, the mutation at amino acid position 490 replaces Glu (E) with Lys (K); the mutation at amino acid residue 538 replaces Iso (I) with Lys (K); the mutation at amino acid residue 486 replaces Gln (Q) with Glu (E); and the mutation at position 499 replaces Iso (I) with Lys (K). Specifically, the engineered cleavage monomer can be prepared by mutating position 490 from E to K and position 538 from I to K in one cleavage monomer to produce an engineered cleavage monomer designated "E490K:I538K," and mutating position 486 from Q to E and position 499 from I to K in another cleavage monomer to produce an engineered cleavage monomer designated "Q486E:I499K." The above-described engineered cleavage monomers are obligate heterodimer mutants in which aberrant cleavage is minimized or eliminated. Engineered cleavage monomers can be prepared using suitable methods, for example, by site-directed mutagenesis of the wild-type cleavage monomer (FokI) as described in US Patent No. 7,888,121, which is incorporated herein in its entirety.
[0043] Additional domains.In some embodiments, the zinc finger nuclease further comprises at least one nuclear localization sequence (NLS). NLS is an amino acid sequence that promotes targeting of the zinc finger nuclease protein to the nucleus to introduce double-strand breaks at the target sequence in the chromosome. Nuclear localization signals are known in the art (see, for example, Lange et al., J. Biol. Chem., 2007, 282: 5101-5105). Non-limiting examples of nuclear localization signals include PKKKRKV (SEQ ID NO: 1), PKKKRRV (SEQ ID NO: 2), KRPAATKKAGQAKKKK (SEQ ID NO: 3), YGRKKRRQRRR (SEQ ID NO: 4), RKKRRQRRR (SEQ ID NO: 5), PAAKRVKLD (SEQ ID NO: 6), RQRRNELKRSP (SEQ ID NO: 7), VSRKRPRP (SEQ ID NO: 8), PPKKARED (SEQ ID NO: 9), PQPKKKPL (SEQ ID NO: 10), SALIKKKKKMAP (SEQ ID NO: 11), PKQKKRK (SEQ ID NO: 12), RKLKKKIKKL (SEQ ID NO: 13), REKKKFLKRR (SEQ ID NO: 14), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 15), RKCLQAGMNLEARKTKK (SEQ ID NO: 16), ID NO: 16), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 17), and RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 18). The NLS can be located at the N-terminus, C-terminus, or an internal position of the zinc finger nuclease.
[0044] In additional embodiments, the zinc finger nuclease may further comprise at least one cell penetrating domain. Examples of suitable cell penetrating domains include, but are not limited to, GRKKRRQRRRPPQPKKKRKV (SEQ ID NO: 19), PLSSIFSRIGDPPKKKRKV (SEQ ID NO: 20), GALFLGWLGAAGSTMGAPKKKRKV (SEQ ID NO: 21), GALFLGFLGAAGSTMGAWSQPKKKRKV (SEQ ID NO: 22), KETWWETWWTEWSQPKKKRKV (SEQ ID NO: 23), YARAAARQARA (SEQ ID NO: 24), THRLPRRRRRR (SEQ ID NO: 25), GGRRARRRRRR (SEQ ID NO: 26), RRQRRTSKLMKR (SEQ ID NO: 27), GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 28), KALAWEAKLAKALAKALAKHLAKALAKALKCEA (SEQ ID NO: 29), and RQIKIWFQNRRMKWKK (SEQ ID NO: 30). NO:30). The cell penetrating domain can be located at the N-terminus, C-terminus, or an internal position of the zinc finger nuclease.
[0045] In other embodiments, the zinc finger nuclease may further include at least one tag domain. Non-limiting examples of tag domains include fluorescent proteins, purification tags, and epitope tags. In one embodiment, the tag domain can be a fluorescent protein. Non-limiting examples of suitable fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, EGFP, Emerald, Azami Green, Monomeric Azami Green, CopGFP, AceGFP, ZsGreen1), yellow fluorescent proteins (e.g., YFP, EYFP, Citrine, Venus, YPet, PhiYFP, ZsYellow1), blue fluorescent proteins (e.g., EBFP, EBFP2, Azurite, mKalama1, GFPuv, Sapphire, T-sapphire), cyan fluorescent proteins (e.g., ECFP, Cerulean, CyPet, AmCyan1, Midoriishi-Cyan), red fluorescent proteins (mKate, mKate2, mPlum, DsRed), and cyan fluorescent proteins. In another embodiment, the marker domain can be a purification tag and / or an epitope tag. Suitable tags include, but are not limited to, a poly (His) tag, a FLAG (or DDK) tag, a Halo tag, an AcV5 tag, an AU1 tag, an AU5 tag, a biotin carboxyl carrier protein (BCCP), a calmodulin binding protein (CBP), a chitin binding domain (CBD), an E tag, an E2 tag, an ECS tag, an eXact tag, a Glu-Glu tag, a glutathione-S-transferase (GST), an HA tag, an HSV tag, a KT3 tag, a maltose binding protein (MBP), a MAP tag, a Myc tag, a NE tag, a NusA tag, a PDZ tag, an S tag, an S1 tag, an SBP tag, a Softag 1 tag, a Softag 3 tag, a Spot tag, a Strep tag, a SUMO tag, a T7 tag, a tandem affinity purification (TAP) tag, thioredoxin (TRX), a V5 tag, a VSV-G tag, and a Xa tag.The tag domain can be located at the N-terminus, C-terminus, or an internal position of the zinc finger nuclease.
[0046] At least one nuclear localization signal, at least one cell penetrating domain and / or at least one tag domain can be directly connected to the zinc finger nuclease via one or more chemical bonds (e.g., covalent bonds). Alternatively, at least one nuclear localization signal, at least one cell penetrating domain and / or at least one tag domain can be indirectly connected to the zinc finger nuclease via one or more joints. Suitable joints include amino acids, peptides, nucleotides, nucleic acids, organic joint molecules (e.g., maleimide derivatives, N-ethoxybenzylimidazole, biphenyl-3,4',5-tricarboxylic acid, p-aminobenzyloxycarbonyl, etc.), disulfide joints and polymer joints (e.g., PEG). The joint may include one or more spacer groups, including but not limited to alkylene, alkenylene, alkynylene, alkyl, alkenyl, alkynyl, alkoxy, aryl, heteroaryl, aralkyl, aralkenyl, aralkynyl, etc. The joint may be neutral or carry a positive or negative charge. In addition, the linker can be cleavable so that the covalent bond of the linker connecting the linker to another chemical group can be broken or cracked under certain conditions, including pH, temperature, salt concentration, light, catalyst or enzyme. In some embodiments, the linker can be a peptide linker. The peptide linker can be a flexible amino acid linker or a rigid amino acid linker. Other examples of suitable linkers are well known in the art, and programs for designing linkers are readily available (Crasto et al., Protein Eng., 2000, 13 (5): 309-312). (ii) CRISPR ribonucleoprotein (RNP)
[0047] In other embodiments, the targeting endonuclease can be a clustered regularly interspaced short palindromic repeats (CRISPR) nuclease. CRISPR nucleases are RNA-guided nucleases derived from bacterial or archaeal CRISPR / CRISPR-associated (Cas) systems. The CRISPR RNP system comprises a CRISPR nuclease and a guide RNA.
[0048] Nuclease.CRISPR nucleases can be derived from type I (i.e., IA, IB, IC, ID, IE, or IF), type II (i.e., IIA, IIB, or IIC), type III (i.e., IIIA or IIIB), type V, or type VI CRISPR systems, which are found in various bacteria and archaea. For example, the CRISPR nuclease can be from a Streptococcus sp. (e.g., S. pyogenes, S. thermophiles, S. pasteurianus), a Campylobacter sp. (e.g., Campylobacter jejuni), a Francisella sp. (e.g., Francisella novicida), an Acaryochloris sp., an Acetohalobium sp., an Acidaminococcus sp., an Acidithiobacillus sp., an Alicyclobacillus sp., an Allochromatium sp., an Ammonifex sp., an Anabaena sp., an Arthrospira sp., a Bacillus sp. sp.), Burkholderiales sp., Caldicelulosiruptor sp., Candidatus sp., Clostridium sp., Crocosphaerasp., Cyanothece sp., Exiguobacterium sp., Finegoldia sp., Ktedonobacter sp., Lachnospiraceae sp., Lactobacillus sp., Lyngbya sp., Marinobacter sp., Methanohalobium sp., Microscilla sp., Microcoleus sp., Microcystis sp., Natranaerobius sp., Neisseria sp.), Nitrosococcus sp., Nocardiopsis sp., Nodularia sp., Nostoc sp., Oscillatoria sp., Polaromonas sp., Pelotomaculum sp., Pseudoalteromonas sp., Petrotoga sp., Prevotella sp., Staphylococcus sp., Streptomyces sp., Streptosporangium sp., Synechococcus sp., Thermosiphosp., or Verrucomicrobia sp. In other embodiments, the CRISPR nuclease may be derived from an archaeal CRISPR system, a CRISPR / CasX system, or a CRISPR / CasY system (Burstein et al., Nature, 2017, 542(7640):237-241).
[0049] In some embodiments, the CRISPR nuclease can be derived from a type II CRISPR nuclease. For example, a type II CRISPR nuclease can be a Cas9 protein. Suitable Cas9 nucleases include Streptococcus pyogenes Cas9 (SpCas9), Francisella novae Cas9 (FnCas9), Staphylococcus aureus (SaCas9), Streptococcus thermophilus Cas9 (StCas9), Streptococcus pasteurianus (SpaCas9), Campylobacter jejuni Cas9 (CjCas9), Neisseria meningitis Cas9 (NmCas9) or Neisseria cinerea Cas9 (NcCas9). In other embodiments, the CRISPR nuclease can be derived from a type V CRISPR nuclease, such as a Cpf1 nuclease. Suitable Cpf1 nucleases include Francisella novicida Cpf1 (FnCpf1), Acidaminococcus sp. Cpf1 (AsCpf1), or Lachnospiraceae ND2006 Cpf1 (LbCpf1). In another embodiment, the CRISPR nuclease can be derived from a type VI CRISPR nuclease, such as Leptotrichia wadei Cas13a (LwaCas13a) or Leptotrichia shahii Cas13a (LshCas13a).
[0050] The CRISPR nuclease can be a wild-type CRISPR nuclease, a modified CRISPR nuclease, or a fragment of a wild-type CRISPR nuclease or a modified CRISPR nuclease. The CRISPR nuclease can be modified to increase nucleic acid binding affinity and / or specificity, change enzymatic activity and / or change another property of the protein. For example, the nuclease (i.e., DNA enzyme, RNA enzyme) domain of the CRISPR nuclease can be modified, deleted or inactivated. The CRISPR nuclease can be truncated to remove domains that are not essential for the function of the nuclease.
[0051] CRISPR nucleases contain two nuclease domains. For example, the Cas9 nuclease contains an HNH domain that cleaves the complementary strand of the guide RNA and a RuvC domain that cleaves the non-complementary strand; the Cpf1 nuclease contains a RuvC domain and a NUC domain; and the Cas13a nuclease contains two HNEPN domains. When both nuclease domains are functional, the CRISPR nuclease introduces a double-strand break. Either nuclease domain can be inactivated by one or more mutations and / or deletions, generating variants that introduce a single-strand break in one strand of the double-stranded sequence. For example, one or more mutations in the RuvC domain of the Cas9 nuclease (e.g., D10A, D8A, E762A, and / or D986A) produce an HNH nickase that makes a guide RNA complementary strand produce a nick; and one or more mutations in the HNH domain of the Cas9 nuclease (e.g., H840A, H559A, N854A, N856A, and / or N863A) produce a RuvC nickase that makes a guide RNA non-complementary strand produce a nick. Similar mutations can convert Cpf1 and Cas13a nucleases into nickases. Two CRISPR nickases targeting opposite strands of a chromosomal sequence (by misplaced guide RNA (offset guide RNA) pairs) can be used in combination to produce double-strand breaks in the chromosomal sequence. Dual CRISPR nickase RNPs can increase target specificity and reduce off-target effects.
[0052] Additional domains.The CRISPR nuclease may further comprise at least one nuclear localization sequence (NLS). NLS is an amino acid sequence that facilitates targeting of zinc finger nuclease proteins into the nucleus to introduce double-strand breaks at target sequences in chromosomes. Nuclear localization signals are known in the art (see, for example, Lange et al., J. Biol. Chem., 2007, 282: 5101-5105). Non-limiting examples of nuclear localization signals include PKKKRKV (SEQ ID NO: 1), PKKKRRV (SEQ ID NO: 2), KRPAATKKAGQAKKKK (SEQ ID NO: 3), YGRKKRRQRRR (SEQ ID NO: 4), RKKRRQRRR (SEQ ID NO: 5), PAAKRVKLD (SEQ ID NO: 6), RQRRNELKRSP (SEQ ID NO: 7), VSRKRPRP (SEQ ID NO: 8), PPKKARED (SEQ ID NO: 9), PQPKKKPL (SEQ ID NO: 10), SALIKKKKKMAP (SEQ ID NO: 11), PKQKKRK (SEQ ID NO: 12), RKLKKKIKKL (SEQ ID NO: 13), REKKKFLKRR (SEQ ID NO: 14), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 15), RKCLQAGMNLEARKTKK (SEQ ID NO: 16), (SEQ ID NO: 16), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 17), and RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 18). The NLS can be located at the N-terminus, C-terminus, or an internal position of the CRISPR nuclease.
[0053] In additional embodiments, the CRISPR nuclease may further comprise at least one cell penetrating domain. Examples of suitable cell penetrating domains include, but are not limited to, GRKKRRQRRRPPQPKKKRKV (SEQ ID NO: 19), PLSSIFSRIGDPPKKKRKV (SEQ ID NO: 20), GALFLGWLGAAGSTMGAPKKKRKV (SEQ ID NO: 21), GALFLGFLGAAGSTMGAWSQPKKKRKV (SEQ ID NO: 22), KETWWETWWTEWSQPKKKRKV (SEQ ID NO: 23), YARAAARQARA (SEQ ID NO: 24), THRLPRRRRRR (SEQ ID NO: 25), GGRRARRRRRR (SEQ ID NO: 26), RRQRRTSKLMKR (SEQ ID NO: 27), GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 28), KALAWEAKLAKALAKALAKHLAKALAKALKCEA (SEQ ID NO: 29), and RQIKIWFQNRRMKWKK (SEQ ID NO: 30). NO:30). The cell penetrating domain can be located at the N-terminus, C-terminus, or an internal position of the CRISPR protein.
[0054] In other embodiments, the CRISPR nuclease may further comprise at least one tag domain. Non-limiting examples of tag domains include fluorescent proteins, purification tags, and epitope tags. In one embodiment, the tag domain can be a fluorescent protein. Non-limiting examples of suitable fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, EGFP, Emerald, Azami Green, Monomeric Azami Green, CopGFP, AceGFP, ZsGreen1), yellow fluorescent proteins (e.g., YFP, EYFP, Citrine, Venus, YPet, PhiYFP, ZsYellow1), blue fluorescent proteins (e.g., EBFP, EBFP2, Azurite, mKalama1, GFPuv, Sapphire, T-sapphire), cyan fluorescent proteins (e.g., ECFP, Cerulean, CyPet, AmCyan1, Midoriishi-Cyan), red fluorescent proteins (mKate, mKate2, mPlum, DsRed), and cyan fluorescent proteins. In another embodiment, the marker domain can be a purification tag and / or an epitope tag. Suitable tags include, but are not limited to, a poly (His) tag, a FLAG (or DDK) tag, a Halo tag, an AcV5 tag, an AU1 tag, an AU5 tag, a biotin carboxyl carrier protein (BCCP), a calmodulin binding protein (CBP), a chitin binding domain (CBD), an E tag, an E2 tag, an ECS tag, an eXact tag, a Glu-Glu tag, a glutathione-S-transferase (GST), an HA tag, an HSV tag, a KT3 tag, a maltose binding protein (MBP), a MAP tag, a Myc tag, a NE tag, a NusA tag, a PDZ tag, an S tag, an S1 tag, an SBP tag, a Softag 1 tag, a Softag 3 tag, a Spot tag, a Strep tag, a SUMO tag, a T7 tag, a tandem affinity purification (TAP) tag, thioredoxin (TRX), a V5 tag, a VSV-G tag, and a Xa tag.The tag domain can be located at the N-terminus, C-terminus, or an internal location of the CRISPR nuclease.
[0055] At least one nuclear localization signal, at least one cell penetrating domain and / or at least one tag domain can be directly connected to the CRISPR nuclease via one or more chemical bonds (e.g., covalent bonds). Alternatively, at least one nuclear localization signal, at least one cell penetrating domain and / or at least one tag domain can be indirectly connected to the CRISPR nuclease via one or more joints. Suitable joints include amino acids, peptides, nucleotides, nucleic acids, organic joint molecules (e.g., maleimide derivatives, N-ethoxybenzylimidazole, biphenyl-3,4',5-tricarboxylic acid, p-aminobenzyloxycarbonyl, etc.), disulfide joints and polymer joints (e.g., PEG). The joint may include one or more spacer groups, including but not limited to alkylene, alkenylene, alkynylene, alkyl, alkenyl, alkynyl, alkoxy, aryl, heteroaryl, aralkyl, aralkenyl, aralkyne, etc. The joint can be neutral or carry a positive or negative charge. In addition, the joint can be cleavable so that the covalent bond of the joint connecting the joint to another chemical group can be broken or cracked under certain conditions, and the conditions include pH, temperature, salt concentration, light, catalyst or enzyme. In some embodiments, the joint can be a peptide joint. The peptide joint can be a flexible amino acid joint or a rigid amino acid joint. Other examples of suitable joints are well known in the art, and the program for designing joints is easy to get in this area.
[0056] guide RNA. CRISPR nucleases are guided to their target sites by guide RNAs. The guide RNAs hybridize to the target sites and interact with the CRISPR nucleases to direct the CRISPR nucleases to the target sites in the chromosomal sequence. p rotospacer a djacent m There are no sequence restrictions on the target site except for the PAM (PAM) of Cas9. CRISPR proteins from different bacterial species recognize different PAM sequences. For example, PAM sequences include 5'-NGG (SpCas9, FnCAs9), 5'-NGRRT (SaCas9), 5'-NNAGAAW (StCas9), 5'-NNNNGATT (NmCas9), 5-NNNNRYAC (CjCas9), and 5'-TTTV (Cpf1), where N is defined as any nucleotide, R is defined as G or A, W is defined as A or T, Y is defined as C or T, and V is defined as A, C, or G. The Cas9 PAM is located 3' of the target site, and the cpf1 PAM is located 5' of the target site.
[0057] The guide RNA contains three regions: a first region at the 5' end that is complementary to the sequence at the target site, a second internal region that forms a stem-loop structure, and a third 3' region that remains essentially single-stranded. The first region of each guide RNA is different, allowing each guide RNA to guide the CRISPR nuclease to a specific target site. The second and third regions (also called scaffold regions) of each guide RNA can be the same in all guide RNAs.
[0058] The first region of the guide RNA is complementary to the sequence at the target site (i.e., the pre-spacer sequence) so that the first region of the guide RNA can be base-paired with the sequence at the target site. The complementarity between the first region of the guide RNA (i.e., crRNA) and the target sequence can be at least 80%, at least 85%, at least 90%, at least 95% or more. Generally, there is no mismatch (i.e., complementarity is complete) between the sequence in the first region of the guide RNA and the sequence at the target site. In various embodiments, the first region of the guide RNA can include about 10 nucleotides to more than about 25 nucleotides. For example, the length of the base pairing region between the first region of the guide RNA and the target site in the chromosome sequence can be about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25 or more than 25 nucleotides. In an exemplary embodiment, the length of the first region of the guide RNA is about 19, 20 or 21 nucleotides.
[0059] The guide RNA also includes a second region that forms a secondary structure. In some embodiments, the secondary structure includes a stem (or hairpin) and a loop. The length of the loop and the stem can vary. For example, the length range of the loop can be about 3 to about 10 nucleotides, and the length range of the stem can be about 6 to about 20 base pairs. The stem can include one or more bulges of 1 to about 10 nucleotides. Therefore, the length range of the total length of the second region can be about 16 to about 60 nucleotides. In an exemplary embodiment, the length of the loop is about 4 nucleotides, and the stem includes about 12 base pairs.
[0060] The guide RNA also includes a third region at the 3' end that remains substantially single-stranded. Thus, the third region does not have complementarity with any chromosomal sequence in the cell of interest, and does not have complementarity with the rest of the guide RNA. The length of the third region can vary. Typically, the length of the third region is greater than about 4 nucleotides. For example, the length of the third region can range from about 5 to about 60 nucleotides.
[0061] The length of the sum of the lengths of the second and third regions (or scaffolds) of the guide RNA can range from about 30 to about 120 nucleotides. In one aspect, the length of the sum of the lengths of the second and third regions of the guide RNA ranges from about 70 to about 100 nucleotides.
[0062] In some embodiments, the guide RNA comprises a molecule comprising all three regions. In other embodiments, the guide RNA can comprise two separate molecules. The first RNA molecule can comprise the first (5') region of the guide RNA and half of the "stem" of the second region of the guide RNA. The second RNA molecule can comprise the other half of the "stem" of the second region of the guide RNA and the third region of the guide RNA. Therefore, in this embodiment, the first and second RNA molecules each contain a sequence of nucleotides that are complementary to each other. For example, in one embodiment, the first and second RNA molecules each contain a sequence (about 6 to about 20 nucleotides) that base-pairs with another sequence to form a functional guide RNA. (iii) Other targeted endonucleases
[0063] In a further embodiment, the targeting endonuclease can be a meganuclease. A meganuclease is a deoxyribonuclease characterized by a long recognition sequence, that is, a recognition sequence generally ranging from about 12 base pairs to about 40 base pairs. Due to this requirement, the recognition sequence generally only occurs once in any given genome. Due to this requirement, the recognition sequence generally only occurs once in any given genome. Among meganucleases, the homing endonuclease family, referred to as LAGLIDADG, has become a valuable tool for studying genomes and genome engineering (see, for example, Arnould et al., 2011, Protein Eng Des Sel, 24 (1-2): 27-31). Other suitable meganucleases include I-CreI and I-Dmol. A meganuclease can target a specific chromosomal sequence by modifying its recognition sequence using techniques well known to those skilled in the art.
[0064] In another embodiment, the targeting endonuclease can be a transcription activator-like effector (TALE) nuclease. TALE is a transcription factor from the plant pathogen Xanthomonas (Xanthomonas), which can be easily engineered to bind to new DNA targets. TALE or its truncated version can be connected to the catalytic domain of an endonuclease such as FokI to produce a targeting endonuclease called TALE nuclease or TALEN (Sanjana et al., 2012, Nat Protoc, 7 (1): 171-192) and Arnould et al., 2011, Protein Engineering, Design & Selection, 24 (1-2): 27-31).
[0065] In alternative embodiments, the targeting endonuclease can be a chimeric nuclease. Non-limiting examples of chimeric nucleases include ZF-meganucleases, TAL-meganucleases, Cas9-FokI fusions, ZF-Cas9 fusions, TAL-Cas9 fusions, and the like. Those skilled in the art are familiar with the means for generating such chimeric nuclease fusions.
[0066] In other embodiments, the targeting endonuclease can be a site-specific endonuclease. In particular, the site-specific endonuclease can be a "rare cutting" endonuclease, whose recognition sequence rarely occurs in the genome. Alternatively, the site-specific endonuclease can be engineered to cut the site of interest (Friedhoff et al., 2007, Methods Mol Biol 352:1110123). Typically, the recognition sequence of the site-specific endonuclease only occurs once in the genome. In other alternative embodiments, the targeting endonuclease can be an artificially targeted DNA double-strand break inducing agent. (b) Targeted endonuclease delivery to cells
[0067] The method comprises introducing a targeted endonuclease into a parental cell line of interest. The targeted endonuclease can be introduced into the cell as a purified isolated protein or as a nucleic acid encoding the targeted endonuclease. The nucleic acid can be DNA or RNA. In embodiments where the encoding nucleic acid is mRNA, the mRNA can be 5' capped and / or 3' polyadenylated. In embodiments where the encoding nucleic acid is DNA, the DNA can be linear or circular. The nucleic acid can be part of a plasmid or viral vector, where the encoding DNA can be operably linked to a suitable promoter. Those skilled in the art are familiar with suitable vectors, promoters, other control elements, and means for introducing the vector into the cell of interest. In embodiments where the targeted endonuclease is a CRISPR nuclease, the CRISPR nuclease system can be introduced into the cell as a gRNA-protein complex.
[0068] The targeting endonuclease molecule can be introduced into the cell by various means. Suitable delivery means include microinjection, electroporation, sonoporation, biolistics, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendrimer transfection, heat shock transfection, nucleofection transfection, magnetofection, lipofection, puncture transfection (impalefection), optical transfection, nucleic acid uptake and the delivery via liposome, immunoliposome, virion or artificial virion of proprietary reagent enhancement. In a specific embodiment, the targeting endonuclease molecule is introduced into the cell by nucleofection.
[0069] Optional donor polynucleotide. The method for genome modification or through engineering approaches for targeting can further include at least one donor polynucleotide being introduced into cell, and described donor polynucleotide comprises the sequence that has at least one nucleotide change relative to target chromosome sequence.The sequence at the site of the targeting in donor polynucleotide and chromosomal sequence or near it has the sequence identity of substance, makes the double-strand break introduced by targeting endonuclease can be repaired by homology directed repair process, and the sequence of donor polynucleotide can insert in chromosomal sequence or exchange with chromosomal sequence, thus modify chromosomal sequence.For example, donor polynucleotide can comprise the first sequence that has the sequence identity of substance with the sequence on one side of target site and the second sequence that has the sequence identity of substance with the sequence on the other side of target site.Donor polynucleotide can further comprise the donor sequence for being integrated into the chromosomal sequence of targeting.For example, donor sequence can be exogenous sequence (for example, marker sequence) so that the integration of exogenous sequence destroys reading frame and inactivates the chromosomal sequence of targeting.
[0070] The length of the first and second sequences with the sequence at or near the target site in the donor polynucleotide and the chromosome sequence having substantial sequence identity can and will change.Usually, the length of each of the first and second sequences in the donor polynucleotide is at least about 10 nucleotides.In various embodiments, the length of the donor polynucleotide sequence with the sequence identity to the chromosome sequence can be about 15 nucleotides, about 20 nucleotides, about 25 nucleotides, about 30 nucleotides, about 40 nucleotides, about 50 nucleotides, about 100 nucleotides or more than 100 nucleotides.
[0071] The phrase "substantial sequence identity" refers to that the sequence in the polynucleotide has at least about 75% sequence identity to the chromosome sequence of interest. In some embodiments, the sequence in the polynucleotide has about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to the chromosome sequence of interest.
[0072] The length of the donor polynucleotide can and will change.For example, the length range of the donor polynucleotide can be about 20 nucleotides to about 200,000 nucleotides at the most.In various embodiments, the length range of the donor polynucleotide can be about 20 nucleotides to about 100 nucleotides, about 100 nucleotides to about 1000 nucleotides, about 1000 nucleotides to about 10,000 nucleotides, about 10,000 nucleotides to about 100,000 nucleotides or about 100,000 nucleotides to about 200,000 nucleotides.
[0073] Typically, the donor polynucleotide is DNA. The DNA can be single-stranded or double-stranded. The DNA can be linear or circular. In some embodiments, the donor polynucleotide can be a single-stranded, linear oligonucleotide comprising less than about 200 nucleotides. In other embodiments, the donor polynucleotide can be a part of a vector. Suitable vectors include DNA plasmids, viral vectors, bacterial artificial chromosomes (BACs), and yeast artificial chromosomes (YACs). In other embodiments, the donor polynucleotide can be a PCR fragment or nucleic acid compounded with a delivery vehicle such as a liposome or poloxamer.
[0074] In some embodiments, the donor polynucleotide can be introduced into the cell simultaneously with the targeting endonuclease molecule. Alternatively, the donor polynucleotide and the targeting endonuclease molecule can be introduced into the cell sequentially. The ratio of the targeting endonuclease molecule to the donor polynucleotide can and will change. Typically, the scope of the ratio of the targeting endonuclease molecule to the donor polynucleotide is about 1:10 to about 10:1. In various embodiments, the ratio of the targeting endonuclease molecule to the polynucleotide can be about 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1 or 10:1. In one embodiment, the ratio is about 1:1. (c) Cultured cells
[0075] The method further comprises maintaining the cell under appropriate conditions such that the double-strand break introduced by the targeting endonuclease can be repaired by: (i) a non-homologous end joining repair process such that the chromosomal sequence is modified by deletion, insertion and / or substitution of at least one nucleotide, or optionally, (ii) a homology-directed repair process such that the chromosomal sequence is exchanged with the sequence of the polynucleotide such that the chromosomal sequence is modified. In embodiments where a nucleic acid encoding the targeting endonuclease is introduced into the cell, the method comprises maintaining the cell under appropriate conditions such that the cell expresses the targeting endonuclease.
[0076] Typically, cells are maintained under conditions suitable for cell growth and / or maintenance. Suitable cell culture conditions are well known in the art and are described in, for example, Santiago et al. (2008) PNAS 105:5809-5814; Moehle et al. (2007) PNAS 104:3055-3060; Urnov et al. Nature 435:646-651; and Lombardo et al. (2007) Nat. Biotechnology 25:1298-1306. It will be appreciated by those skilled in the art that methods for culturing cells are known in the art and can and will vary depending on the cell type. In all cases, conventional optimization can be used to determine the best technique for a particular cell type.
[0077] During this step of the method, the targeted endonuclease recognizes, binds to, and produces a double-strand break at the targeted cleavage site in the chromosomal sequence, and during the repair of the double-strand break, the deletion, insertion, and / or substitution of at least one Nucleotide are introduced into the targeted chromosomal sequence. In a specific embodiment, the targeted chromosomal sequence is inactivated.
[0078] After confirming that the chromosomal sequence of interest has been modified, single cell clones can be isolated and genotyped (by DNA sequencing and / or protein analysis). Cells containing a modified chromosomal sequence can undergo one or more rounds of additional targeted genome modification to modify additional chromosomal sequences, thereby generating double knockouts, triple knockouts, and the like. (IV) Production of recombinant proteins
[0079] Another aspect of the present disclosure includes a method for producing a recombinant protein in a bioproduction system. Suitable recombinant proteins are described in part (I) (c). The method includes expressing a recombinant protein of interest in any engineered cell line described in the above-mentioned part (I), and purifying the recombinant protein expressed. Means for producing or manufacturing recombinant proteins are well known in the art (see, for example, " Biopharmaceutical Production Technology ", Subramanian (ed.), 2012, Wiley-VCH; ISBN: 978-3-527-33029-4).
[0080] The recombinant protein can be purified by a method comprising a clarification step (eg, filtration) and one or more chromatography steps (eg, affinity chromatography, protein A (or G) chromatography, ion exchange (ie, cation and / or anion) chromatography). definition
[0081] Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by those skilled in the art to which the invention belongs. The following references provide general definitions of many of the terms used in the present invention to those skilled in the art: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker edition, 1988); The Glossary of Genetics, 5th ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have their corresponding meanings unless otherwise indicated.
[0082] When introducing elements of the present disclosure or the preferred embodiments thereof, the articles "a," "an," "the," and "said" are intended to mean that there are one or more of the elements. The terms "comprising," "including," and "having" are intended to be inclusive and mean that there may be additional elements other than the listed elements.
[0083] As used herein, the term "endogenous sequence" refers to a chromosomal sequence that is native to a cell.
[0084] The term "exogenous sequence" refers to a chromosomal sequence that is not native to the cell, or a chromosomal sequence that has been moved to a different chromosomal location.
[0085] An "engineered" or "genetically modified" cell refers to a cell in which the genome has been modified or engineered, i.e., the cell contains at least one chromosomal sequence that has been engineered to contain an insertion of at least one nucleotide, a deletion of at least one nucleotide, and / or a substitution of at least one nucleotide.
[0086] The terms "genomic modification" and "genome editing" refer to the process of changing a specific endogenous chromosomal sequence so that the chromosomal sequence is modified. The chromosomal sequence can be modified to include the insertion of at least one nucleotide, the deletion of at least one nucleotide, and / or the substitution of at least one nucleotide. The modified chromosomal sequence is inactivated so that no product is produced. Alternatively, the chromosomal sequence can be modified to produce an altered product.
[0087] As used herein, "gene" refers to a DNA region encoding a gene product (including exons and introns), as well as all DNA regions that regulate the production of a gene product, whether or not such regulatory sequences are adjacent to the coding and / or transcribed sequence. Thus, a gene includes, but is not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, replication origins, matrix attachment sites, and locus control regions.
[0088] The term "heterologous" refers to an entity that is not native to the cell or species of interest.
[0089] The terms "nucleic acid" and "polynucleotide" refer to deoxyribonucleotide or ribonucleotide polymers, which are in linear or cyclic conformations. For the purposes of this disclosure, these terms should not be construed as limitations on polymer length. The term can include known analogs of natural nucleotides, as well as nucleotides that are modified in the base, sugar, and / or phosphate moieties. Typically, analogs of a particular nucleotide have the same base pairing specificity; that is, an analog of A will base pair with T. The nucleotides of a nucleic acid or polynucleotide can be linked by phosphodiester, phosphorothioate, phosphoramidite, phosphorodiamidate bonds, or combinations thereof.
[0090] The term "nucleotide" refers to a deoxyribonucleotide or a ribonucleotide. Nucleotide can be a standard nucleotide (i.e., adenosine, guanosine, cytidine, thymidine, and uridine) or a nucleotide analog. Nucleotide analogs refer to nucleotides with modified purine or pyrimidine bases or modified ribose moieties. Nucleotide analogs can be naturally occurring nucleotides (e.g., inosine) or non-naturally occurring nucleotides. Non-limiting examples of modifications on the sugar or base moiety of a nucleotide include adding (or removing) acetyl, amino, carboxyl, carboxymethyl, hydroxyl, methyl, phosphoryl, and thiol groups, as well as replacing the carbon and nitrogen atoms of the base with other atoms (e.g., 7-deazapurine). Nucleotide analogs also include dideoxynucleotides, 2'-O-methyl nucleotides, locked nucleic acids (LNA), peptide nucleic acids (PNA), and morpholino oligonucleotides.
[0091] The terms "polypeptide" and "protein" are used interchangeably to refer to a polymer of amino acid residues.
[0092] As used herein, the term "target site" or "target sequence" refers to a nucleic acid sequence that defines a portion of a chromosomal sequence to be modified or edited and to which a targeting endonuclease is engineered to recognize and bind (provided sufficient conditions for binding exist).
[0093] The terms "upstream" and "downstream" refer to positions in a nucleic acid sequence relative to a fixed position. Upstream refers to the region 5' (i.e., near the 5' end of the chain) relative to the position, and downstream refers to the region 3' (i.e., near the 3' end of the chain) relative to the position.
[0094] The technology for determining nucleic acid and amino acid sequence identity is known in the art. Generally, such technology includes determining the nucleotide sequence of the mRNA of a gene and / or determining the amino acid sequence encoded therein, and comparing these sequences with a second nucleotide or amino acid sequence. Genomic sequence can also be determined and compared in this way. Generally, identity refers to the accurate nucleotide and nucleotide or amino acid and amino acid correspondence of two polynucleotides or polypeptide sequences. Two or more sequences (polynucleotides or amino acids) can be compared by determining their percentage identity. The percentage identity of two sequences (no matter it is a nucleotide sequence or an amino acid sequence) is the number of accurate matches between the two aligned sequences divided by the length of the shorter sequence and multiplied by 100. Smith and Waterman, Advances in Applied Mathematics 2:482-489 (1981) local homology algorithm provides an approximate comparison of nucleotide sequences. The algorithm can be applied to amino acid sequences using a scoring matrix developed by Dayhoff, Atlas of Protein Sequences and Structure, MO Dayhoff ed., 5 suppl. 3: 353-358, National Biomedical Research Foundation, Washington, DC, USA and standardized by Gribskov, Nucl. Acids Res. 14 (6): 6745-6763 (1986). An exemplary implementation of this algorithm for determining the percent identity of a sequence is provided by Genetics Computer Group (Madison, Wis.) in the "BestFit" tool application. Other suitable programs for calculating percent identity or similarity between sequences are generally known in the art, for example, another algorithmic program is the BLAST used with default parameters. For example, BLASTN and BLASTP can be used using the following default parameters: genetic code = standard; filter = none; chain = both; cutoff = 60; expectation = 10; matrix = BLOSUM62; description = 50 sequences; sorting = high score; database = non-redundant GenBank + EMBL + DDBJ + PDB + GenBank CDS translations + Swiss protein + Spupdate + PIR. Details of these programs can be found on the GenBank website. For the sequences described herein, the desired degree of sequence identity ranges from about 80% to 100% and any integer value therebetween. Typically, the percent identity between sequences is at least 70-75%, preferably 80-82%, more preferably 85-90%, even more preferably 92%, still more preferably 95%, and most preferably 98% sequence identity.
[0095] As various changes could be made in the above cells and methods without departing from the scope of the invention, it is intended that all matter contained in the above description and the examples provided below shall be interpreted as illustrative and not in a limiting sense. Example
[0096] The following examples illustrate certain aspects of the invention. Example 1: Design of a serine-mediated selection system
[0097] To develop a serine-mediated metabolic selection system, a CHO cell line sensitive to depletion of the non-essential amino acid serine (Ser) was first developed. A comprehensive search against the Reactome and KEGG databases was performed to identify all genes associated with the serine synthesis pathway. The endogenous phosphoserine phosphatase gene (PSPH) was identified as the only non-redundant gene responsible for de novo synthesis of Ser ( Figure 1 To generate Ser-sensitive CHO cell lines, glutamine (Gln) auxotrophs from MilliporeSigma were used. GS - / - cell lines Endogenous PSPH coding sequence ( Figure 2 )pass Whole genome sequencing (WGS) of the cell line revealed that it is present in two copies. genome, as determined by droplet digital PCR (ddPCR) analysis ( Figure 3 ). CRISPR / Cas9 gene editing reagents were designed to disrupt the second exon of the PSPH gene. The CRISPR / Cas9 target sequence is Figure 2 The cells were cultured in a shaking chamber at 37°C with 5% CO2 in a flask supplemented with 6 mM L-glutamine (MilliporeSigma G7513). Cultured in CD CHO Fusion medium (MilliporeSigma 14365C) (Fusion + Gln) cells. One day before transfection, cells were seeded at 0.5E6 to maintain the culture in logarithmic growth phase. Cas9 RNP was complexed for 15 minutes at room temperature by mixing 50 pmol Cas9 (SigmaCAS9PROT-250UG) with 150 pmol sgRNA (Sigma). 4e5 cells were transfected with a total of 200 pmol of complexed RNP using Lonza's 4DX nucleofection system using program DT-133 and SF nucleofection solution. Cas9 cleavage activity was assessed by next generation sequencing (NGS), and it was determined that an additional 10 mM serine was required in the culture medium to allow optimal cell survival when modifying the PSPH gene ( Figure 4 The transfected cells were transferred to a 6-well culture flask containing 3 mL of medium supplemented with 6 mM L-glutamine (MilliporeSigma G7513) and 10 mM serine (MilliporeSigma S4311-100G). CD CHO Fusion medium (MilliporeSigma 14365C) (Fusion + Gln + Ser). After transfection, cells were incubated in a static environment at 37 ° C / 5% CO2 for 96 hours. Cells were scaled up to T-25 flasks. Single-cell clones from the Cas9-modified pools were isolated by fluorescence-activated cell sorting (FACS) into 96-well culture plates. Single-cell clones were evaluated by next-generation sequencing (NGS) to identify clones containing two copies of PSPH that had been successfully genetically disrupted ( Figure 5 ).
[0098] To demonstrate the efficacy of the serine-mediated selection mechanism both as a stand-alone system and as part of a dual metabolic selection system, stable selected cell populations expressing a variety of molecules were generated. The molecules used in the validation of this system included cyan fluorescent protein (BFP), Dasher green fluorescent protein (GFP), and human IgG1. Cells were cultured in Fusion+Gln+Ser medium. GS - / - PSPH - / -Cells. The expression vectors used in this work contained either a phosphoserine phosphatase (PSPH) or a glutamine synthetase (GS) selection marker, as specified by the experimental design. Expression of Chinese hamster (Cricetulus Griseus) PSPH (protein: phosphoserine phosphatase; gene: PSPH; UniProtKB ID: G3I2M2) or mouse GS (protein: glutamine synthetase {glutamate-ammonia ligase}; gene: Glu1; UniProtKB ID: P15105) was driven by the 5' SV40 promoter and the SV40 polyadenylation sequence at the 3' end of the gene ( Figure 6 ).
[0099] Cultured in Fusion+Gln+Ser at 37°C with shaking in 5% CO2 GS - / - PSPH - / - cells. Using electroporation, 1.0e6 cells per condition were transfected with 7.5 μg of plasmid DNA. The transfected cells were transferred to a 6-well plate containing 3 mL of Fusion+Gln+Ser. After the cells had recovered to >98% viability, the cells were pelleted and the medium aspirated, then 5e5 viable cells / mL were resuspended in 10 mL of the appropriate selection medium and transferred to a T-75 flask. Glutamine-based selection was performed using Fusion-Gln, serine-based selection was performed using Fusion-Ser, and glutamine / serine dual selection was performed using Fusion-Gln and Fusion-Ser. Cell viability and viable cell density of the various selection cultures were monitored over time. After recovery from selection, the stable selected cultures were transferred to TubeSpinBioreactor Tubes To scale up and adapt to oscillating conditions.
[0100] Developed Advanced CHO Fed-batch medium (MilliporeSigma 14366C), EX- Advanced CHO Feed (MilliporeSigma 24367C) and 4Feed (MilliporeSigma 1.03796.0005) lacks serine (Ser) custom formulations (Advanced-Ser, Feed-Ser and 4Feed-Ser, respectively). Stable selected cultures transfected with IgG1 expression vectors were precipitated, the selection medium was aspirated, and then resuspended in Advanced-Ser at 3e5 viable cells / mL for productivity analysis under batch feeding conditions. Viable cell density and viability data for each culture were collected every other day starting from the 3rd day after inoculation. Starting from the 3rd day after inoculation, 1.5mL of a 50 / 50 blend of Advanced Feed-Ser and 4Feed-Ser was added to each culture. Glucose readings were taken from each culture every other day starting from the 5th day after adding D-+-glucose (MilliporeSigma G8769) to maintain appropriate glucose levels. Productivity was monitored over time, with fed-batch titers recorded every other day starting from the 9th day until the culture dropped to less than 70% viability. Titers were determined using interferometry on a ForteBio Octet and confirmed by HPLC protein A affinity chromatography. Example 2
[0101] Developed a Serine-free base Custom formulation of CD CHO Fusion medium (Fusion-Gln-Ser). Culture in Fusion+Gln+Ser or Fusion+Gln-Ser GS - / - PSPH - / - Colonies were grown for at least seven days. Viability and viable cell density measurements were obtained twice weekly. Figure 7 In the absence of serine GS - / - PSPH - / - The cells were unable to grow, however, when serine was supplemented to the culture medium, GS - / - PSPH - / - Cell growth was rescued. Example 3
[0102] To demonstrate that a stable selected cell population can produce a protein of interest using the Ser-mediated selection system described in Example 1, we developed a vector with IgG heavy chain, IgG light chain, and phosphoserine phosphatase (PSPH) coding sequences. This vector was transfected into GS - / - PSPH - / -cell line. As a control, a mock transfection without DNA was used. The colony was passaged in Fusion+Gln-Ser under selection pressure. The conditions for selection were also applied during recovery, scale-up, and productivity assays. Batch fed productivity assays were inoculated at 3e5 viable cells / mL in Advanced+Gln-Ser medium. The viable cell density and viability of each culture were collected every other day starting from the 3rd day after inoculation. Starting from the 3rd day after inoculation, 1.5mL of a 50 / 50 blend of Advanced Feed-Ser and 4Feed-Ser was added to each culture. Glucose readings were taken every other day starting from the 5th day after the addition of D-+-glucose (MilliporeSigma G8769) to maintain appropriate glucose levels. The resulting IgG titers are shown in Figure 10 middle. Example 4
[0103] To test whether a stable cell population would be able to produce two independent intracellular fluorescent proteins under glutamine and serine selection conditions, we developed two vectors, one containing the GFP and GS coding sequences and the second containing the BFP and PSPH coding sequences ( Figure 5 These two plasmids were co-transfected into GS - / - PSPH - / - The cells were then passaged under dual metabolic selection conditions (Fusion-Gln-Ser). The conditions used for selection were also applied during recovery, scale-up, and all other assays. Figure 8 Growth and viability data from selection assays are shown, demonstrating that cells co-transfected with both vectors (GFP+BFP) survived and grew in -Gln-Ser medium. Figure 8 and 9 The cells co-transfected with the two vectors (GFP+BFP) that survived and grew in the Gln-Ser culture medium were both positive for both GFP and BFP. These data indicate that the GS+PSPH dual metabolic selection system provides cells that have been selected to have been introduced into independent vectors encoding multiple intracellular proteins, without the need to add any selection agent (e.g., antibiotic) to the culture medium. Example 5
[0104] To test whether a stable cell population could produce secreted protein under glutamine- and serine-free conditions, we developed two IgG1-expressing vectors, one containing the IgG heavy chain, IgG light chain, and GS coding sequences, and a second containing the same IgG heavy chain, IgG light chain, and PSPH coding sequences. These two independent vectors were co-transfected into GS - / -PSPH - / - As a control, each vector was also transfected into GS - / - PSPH - / - The cells were then passaged under selection pressure in Fusion-Gln (GS-only transfected cells), Fusion-Ser (PSPH-only transfected cells), or Fusion-Gln-Ser (GS+PSPH transfected cells). The conditions used for selection were also applied during recovery, scale-up, and productivity assays. After 14 days, the GS-only selection cultures fully recovered, while the PSPH-only and GS+PSPH dual selection cultures required 21 days to recover. In fed-batch assays, the GS-only, PSPH-only, and GS+PSPH selected pools were all able to drive IgG production ( Figure 10 ). This indicates that the GS and PSPH produced by the cells when expressing exogenous GS and / or PSPH coding sequences are sufficient to maintain secretory protein production. This offers the potential to operate large-scale production bioreactors under dual metabolic selection conditions, which can be difficult to perform using antibiotic selection methods due to the need to subsequently separate or purify the antibiotic from the desired secretory protein, or due to the cost of adding antibiotics to large-scale bioreactors. In addition, this dual metabolic selection system (GS+PSPH) provides the opportunity to more efficiently select cells into which multiple large vectors have been introduced, for example when expressing bispecific antibodies or other large and / or complex proteins.
Claims
1. A method for producing a recombinant protein product, the method comprising (a) providing a mammalian cell line engineered to have reduced or eliminated endogenous phosphoserine phosphatase (PSPH) expression; (b) introducing a polynucleotide into the mammalian cell line, wherein the polynucleotide encodes a functional PSPH gene and a recombinant protein; (c) culturing the cell line; and (d) purifying the recombinant protein to form the recombinant protein product.
2. The method of claim 1, wherein the mammalian cell line of (a) further comprises reduced or eliminated expression or activity of endogenous glutamine synthetase (GS) and / or asparagine synthetase (ASNS).
3. The method of claim 1, wherein endogenous PSPH expression is reduced or eliminated by inactivating the endogenous PSPH gene of the mammalian cell line.
4. The method of claim 1, wherein the endogenous PSPH gene is inactivated using a targeted endonuclease-mediated genome modification technique.
5. The method of claim 4, wherein the targeted endonuclease is a CRISPR ribonucleoprotein complex or a pair of zinc finger nucleases.
6. The method of claim 1, wherein the mammalian cell line is a Chinese hamster ovary (CHO) cell line, a baby hamster kidney (BHK) cell line, an NS0 mouse myeloma cell line, a HEK293 cell line, or a Vero African green monkey kidney cell line.
7. The method of any one of claims 1 to 6, wherein the cell line is a CHO cell line.
8. The method of any one of claims 1 to 7, wherein the recombinant protein product is selected from antibodies, antibody fragments, vaccines, growth factors, cytokines, hormones, or blood coagulation factors.
9. The method of claim 8, wherein the antibody is a bispecific or multispecific antibody.
10. A genetically engineered mammalian cell line for use in a biological production system, wherein the mammalian cell line is engineered to have reduced or eliminated endogenous PSPH expression.
11. The mammalian cell line of claim 10, wherein PSPH expression is reduced or eliminated by inactivating at least one allele of the chromosomal sequence encoding PSPH.
12. The mammalian cell line of claim 11, wherein one or more alleles of the chromosomal sequence encoding PSPH are inactivated.
13. The mammalian cell line of claim 10, wherein the cell line is engineered to have reduced or eliminated expression of endogenous glutamine synthetase (GS) and / or asparagine synthetase (ASNS).
14. The mammalian cell line of claim 12, wherein the chromosomal sequence is inactivated using a targeted endonuclease-mediated genome modification technique.
15. The mammalian cell line of claim 14, wherein the targeted endonuclease is a ribonucleoprotein complex or a pair of zinc finger nucleases.
16. The mammalian cell line according to claim 15, wherein the non-human cell line is a Chinese hamster ovary (CHO) cell line, a baby hamster kidney (BHK) cell line, an NS0 mouse myeloma cell line, a HEK293 cell line, or a Vero African green monkey kidney cell line.
17. The mammalian cell line according to claim 16, wherein the cell line is a CHO cell line.
18. The mammalian cell line according to claim 17, wherein the cell viability, viable cell density, titer, growth rate, proliferation response, cell morphology, and / or overall cell health are comparable to those of the unengineered parental mammalian cell line.
19. The mammalian cell line according to any one of claims 10 to 18, further comprising at least one nucleic acid encoding a recombinant protein selected from the group consisting of an antibody, an antibody fragment, a vaccine, a growth factor, a cytokine, a hormone, or a blood coagulation factor.
20. The mammalian cell line according to claim 19, wherein the antibody is a bispecific or multispecific antibody.
21. A polynucleotide comprising a nucleic acid sequence encoding a functional PSPH and at least one recombinant protein of interest.
22. A polynucleotide comprising a) a nucleic acid sequence encoding a functional PSPH; b) a nucleic acid sequence encoding a functional GS and / or ASNS; and c) a nucleic acid sequence encoding a mutation in the ASNS and / or PSPH coding sequence, the mutation attenuating the activity of either or both of the enzymes d) a nucleic acid sequence encoding a recombinant protein of interest.
Citation Information
Patent Citations
In vitro peptide or protein expression library
GB2338237A
Functional domains in flavobacterium okeanokoites (FokI) restriction endonuclease
US5356802A
Functional domains in flavobacterium okeanokoities (foki) restriction endonuclease
US5436150A
Insertion and deletion mutants of FokI restriction endonuclease
US5487994A
Zinc finger proteins with high affinity new DNA binding specificities
US5789538A