Method for selecting cell line expressing high levels of recombinant protein by using dihydrofolate reductase split expression vector

The DHFR split expression vector method improves the selection efficiency of high-expressing cell lines by using DHFR's N-domain and C-domain portions as markers, addressing inefficiencies in current recombinant protein production processes.

WO2025244428A1PCT designated stage Publication Date: 2025-11-27KOREA ADVANCED INST OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/006937
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-23
Filing Date
2025-05-22
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Current methods for selecting high-expressing cell lines for recombinant proteins, particularly glycoproteins like antibodies, are inefficient, labor-intensive, and costly, with low selection efficiency and time-consuming processes, especially when producing therapeutic antibodies in large quantities.

Method used

A method involving a dihydrofolate reductase (DHFR) split expression vector is used, where DHFR is divided into N-domain and C-domain portions, which are introduced into cells as selection markers, followed by culturing in a DHFR inhibitor medium to select high-expressing cells.

Benefits of technology

This approach significantly enhances the selection efficiency of cell lines expressing high levels of recombinant proteins, improving productivity and reducing the time and cost associated with traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025006937_27112025_PF_FP_ABST
    Figure KR2025006937_27112025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a recombinant vector for expressing a target protein or a target gene containing a dihydrofolate reductase split fragment, and a method for selecting a cell line highly expressing a recombinant protein by using same. The present invention can dramatically increase selection efficiency of biopharmaceutical-producing cell lines, and can be freely used without stability problems since dihydrofolate reductase (DHFR), which is a conventionally used selection marker, has been improved.
Need to check novelty before this filing date? Find Prior Art

Description

Method for selecting a cell line expressing a high level of recombinant protein using a dihydrofolate reductase split expression vector

[0001] The present invention relates to a method for selecting a cell line expressing a high level of a recombinant protein using a dihydrofolate reductase split expression vector, and more particularly, to a recombinant vector for expressing a target protein containing a dihydrofolate reductase split fragment and a method for selecting a cell line expressing a high level of a recombinant protein using the same.

[0002]

[0003] The global biopharmaceutical market was valued at approximately KRW 412 trillion in 2020, growing at an average annual rate of 12%. The therapeutic protein market for domestic companies was estimated at KRW 3.3 trillion in 2020, with a production increase of 54.9% compared to 2019, demonstrating the most dynamic growth among pharmaceutical sectors. The global biopharmaceutical market continues to grow, and the number of biopharmaceuticals approved in each country is steadily increasing. Domestic biopharmaceutical manufacturers are recently expanding into the contract development and manufacturing organization (CDMO) market, including development within the contract manufacturing organization (CMO) business.

[0004] Recombinant proteins, a key component of biopharmaceuticals, can be expressed in a variety of host cells, including prokaryotic and eukaryotic cells. However, glycoproteins, such as antibodies and Fc fusion proteins, in which sugars bind to specific amino acids in a polypeptide to influence efficacy and safety, are primarily produced in animal cells, cells capable of glycosylation, during post-translational modification. In particular, when producing recombinant protein pharmaceuticals, the degree of similarity between the sugar chains of human (natural) glycoproteins and their immunogenicity is closely related to the protein pharmaceutical's immunogenicity. Therefore, it is crucial to produce glycoproteins with sugar chains identical or similar to those of human glycoproteins.

[0005] To date, animal cells primarily used for the expression and production of recombinant protein drugs include hybridomas, mouse myeloma, and CHO cells. However, while protein expression using these animal cells offers excellent similarity to in vivo proteins, it suffers from the drawbacks of significantly low expression levels and difficulty in mass cultivation. In particular, therapeutic antibodies typically require expression in kilogram quantities, making it challenging to secure sufficient protein through simple animal cell culture. Therefore, to achieve high expression of a specific gene in these host cells, the desired DNA must first be randomly introduced into the animal cell genome, targeting a region with high transcription rates. This introduced exogenous gene replicates with the host cell genome, but the technology for introducing the gene into a specific region with high transcription rates (homologous recombination technology) is not yet widely used.

[0006] Accordingly, a technology is being developed to introduce an expression vector containing an amplifiable gene and a target protein gene into a cell line lacking dihydrofolate reductase, glutamine synthetase, etc. as an expression system for current protein pharmaceuticals, and this is a method that can increase the copy number of the expression vector by treating the cell line with inhibitors of these genes such as methotrexate (MTX) and methionine sulfoximine (MSX). In this way, as the copy number of the expression vector per unit cell increases, the expression amount of the target protein increases.

[0007] To select only cells that have been introduced with the target protein gene and produce the target protein, the gene amplification process leads to significant differences in the protein expression levels of each clone. Therefore, selecting high-expressing cell lines is a tedious and arduous process that requires sifting through hundreds of clones. The most widely used method for selecting clones is limiting dilution using multi-well plates. Another method is to physically separate clones using glass or metal cloning rings, then enzymatically treat the cells and detach them from the rings using a micropipette. Another method involves isolating clones by swiping a cotton swab over clusters of cells. However, these traditional methods are now being replaced by cell sorting or robotic cell line selection methods in industrial settings.

[0008] These methods significantly reduce the labor required to select candidate cell lines. In the case of monoclonal antibodies, which require high concentrations, clones with productivity below 20 pg / cell / day are already excluded during the initial screening process. Multiple transformation processes are often required to select highly productive cell lines, as the number of cells that successfully introduce foreign genes after a single transformation is limited. It is difficult to generate even 500 to 1,000 stable cell lines from a single transformation plate. When selecting productive cell lines, cell growth rate must also be considered. Typically, a cell line with very high productivity will grow slowly, but if it grows quickly and has high productivity, scaling up production to industrial scale becomes easier (Na Gyu-Heum, BIOSAFETY Vol. 8, 56-69, 2007).

[0009] As previously explained, identifying cell clones with highly expressed target transgenes requires the examination and testing of numerous clones. This process is time-consuming, labor-intensive, and costly. Furthermore, the low selection efficiency of dihydrofolate reductase, used as a marker gene, results in only a small number of selected cells becoming high-expressing cell lines. This leads to time and cost constraints in selecting a small number of highly expressing cells.

[0010] Accordingly, the inventors of the present invention have made great efforts to improve the low selection efficiency of dihydrofolate reductase, which is used as a marker in the selection of recombinant proteins, and as a result, have confirmed that when dihydrofolate reductase is expressed by dividing it into an N-domain portion and a C-domain portion and then combined within cells, the selection efficiency of cells highly expressing recombinant proteins is significantly improved, thereby completing the present invention.

[0011]

[0012] Summary of the invention

[0013] The purpose of the present invention is to provide a first vector including a nucleic acid encoding an N-domain portion of dihydrofolate reductase as a selection marker, in order to effectively select cells highly expressing a target protein or target gene.

[0014] Another object of the present invention is to provide a second vector including a nucleic acid encoding a C-domain region of dihydrofolate reductase as a selection marker, in order to effectively select cells highly expressing a target protein or target gene.

[0015] Another object of the present invention is to provide a cell for producing a target protein or target gene transformed with the first vector and the second vector.

[0016] Another object of the present invention is to provide a method for effectively selecting cells highly expressing a target protein or target gene.

[0017] Another object of the present invention is to provide a method for producing a target protein by culturing cells for producing the target protein or target gene.

[0018] To achieve the above purpose, the present invention provides a first vector including a first marker nucleic acid encoding an N-domain portion of dihydrofolate reductase having an amino acid sequence represented by any one of SEQ ID NOs: 1 to 3 as a selection marker.

[0019] The present invention also provides a second vector comprising a second marker nucleic acid encoding a C-domain portion of dihydrofolate reductase having an amino acid sequence represented by any one of SEQ ID NOs: 4 to 6 as a selection marker.

[0020] The present invention also provides a cell for producing a target protein or a target gene, wherein the first vector and the second vector are introduced, and the first vector and / or the second vector further include a gene encoding a target protein, wherein the first vector and the second vector are introduced such that when the N-domain portion of the first vector and the C-domain portion of the second vector are combined with each other, a complete dihydrofolate reductase is formed, and the first vector and / or the second vector include an intein, a leucine zipper, or a coiled coil.

[0021] The present invention also provides a method for selecting cells expressing a target protein or target gene, comprising the following steps:

[0022] (a) a step of producing a library of cells for producing the target protein or target gene;

[0023] (b) culturing the library in a medium containing a dihydrofolate reductase inhibitor and selecting growing recombinant cells; and

[0024] (C) A step of confirming the expression of a target protein or target gene in the selected recombinant cell and obtaining a cell expressing the target protein or target gene.

[0025] The present invention also provides a method for producing a target protein or target gene, comprising the following steps:

[0026] (a) a step of culturing cells for producing the target protein or target gene to produce the target protein or target gene; and

[0027] (b) A step of obtaining the target protein or target gene generated above.

[0028]

[0029] Figure 1 shows a schematic representation of the degrees of freedom and cleavage positions for each amino acid sequence of dihydrofolate reductase.

[0030] Figure 2 shows the three-dimensional structure of dihydrofolate reductase and the cleavage position applied in the present invention.

[0031] Figure 3 shows a recombinant protein expression vector including a dihydrofolate reductase 2-cleavage fragment according to the present invention, where DHFRN represents an N-domain cleavage fragment of dihydrofolate reductase, and DHFRC represents a C-domain cleavage fragment of dihydrofolate reductase.

[0032] Figure 4 shows the results of confirming cell growth upon selection by introducing (transfecting) dihydrofolate reductase by dividing it into the N-domain portion and the C-domain portion.

[0033] Figure 5 shows the results of confirming cell viability upon selection by introducing (transfecting) dihydrofolate reductase by dividing it into the N-domain portion and the C-domain portion.

[0034] Figure 6 shows the results of qualitative analysis using a fluorescent label (MFI: Mean Fluorescence intensity) to confirm that cells expressing etanercept, an FC fusion protein, were selected using an expression system according to the present invention, and that a high amount of etanercept production was observed in the cell pool.

[0035] Figure 7 shows the results of selecting cells expressing etanercept, an FC fusion protein, using an expression system according to the present invention, and then confirming the ratio of etanercept high-expressing cell lines in the cell pool.

[0036]

[0037] Detailed description of the invention and preferred embodiments

[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains, unless otherwise defined herein. Generally, the nomenclature used herein is well known and commonly used in the art.

[0039] The current method for developing cell lines for biopharmaceutical production has an inefficient selection system, resulting in a very small number of cells (approximately one in 10,000) from the entire cell pool being selected as cell lines with high expression of the target protein. Furthermore, obtaining these very small numbers of high-expressing cell lines requires a significant amount of time and money.

[0040] In the present invention, a method was developed to dramatically increase the selection efficiency of a cell line that highly expresses a target protein by splitting dihydrofolate reductase (DHFR) into two fragments and using them as selection markers.

[0041] When producing medical proteins or other recombinant proteins using a recombinant protein producing cell line selected using the method of the present invention, it was confirmed that the cell line exhibited superior recombinant protein productivity compared to cells selected using a method using the existing wild-type dihydrofolate reductase as a selection marker (Figs. 6 and 7).

[0042] Accordingly, the present invention relates to a first vector comprising, as a selection marker, a first marker nucleic acid encoding an N-domain portion of dihydrofolate reductase having an amino acid sequence represented by any one of SEQ ID NOs: 1 to 3;

[0043] In the present invention, the N-domain portion of the dihydrofolate reductase may be characterized by being represented by sequence number 1.

[0044] From another aspect, the present invention relates to a second vector comprising a second marker nucleic acid encoding a C-domain portion of dihydrofolate reductase having an amino acid sequence represented by any one of SEQ ID NOs: 4 to 6 as a selection marker.

[0045] In the present invention, the C-domain portion of the dihydrofolate reductase may be characterized as represented by sequence number 4.

[0046] Examples of the N-domain sequence and C-domain sequence of dihydrofolate reductase that can be used in the present invention are shown in Table 1.

[0047]

[0048] In the present invention, “dihydrofolate reductase (EC 1.5.1.3)” is an enzyme involved in the reaction of reducing 7,8-dihydrofolate to 5,6,7,8-tetrahydrofolate using NADPH as an electron donor. The dihydrofolate reductase can be used regardless of its source as long as it is an enzyme having the above activity, and can be derived from, for example, a human, a mouse, a Chinese hamster, a microorganism, a bacterium, or a mold (e.g., Tallaromyces).

[0049] In one example, the N-terminal domain of the dihydrofolate reductase may be composed of any one amino acid sequence of SEQ ID NOs: 1 to 3, and the C-terminal domain of the dihydrofolate reductase may be composed of any one amino acid sequence of SEQ ID NOs: 4 to 6, and if the N-terminal domain and the C-terminal domain are spliced ​​or joined to have the same or corresponding conversion activity as the dihydrofolate reductase, a polypeptide having an amino acid sequence in which a part of the amino acid sequence of SEQ ID NOs: 1 to 6 is deleted, modified, substituted or added may also be included within the scope of the present application.

[0050] In addition, regardless of the source of the enzyme, a polypeptide having at least 90%, 92%, 94%, 96%, 98%, or 99%, 99.5%, or 99.8% homology or identity with the amino acid sequence of SEQ ID NOs: 1 to 6, and having a conversion activity identical to or corresponding to dihydrofolate reductase, may also be included within the scope of the present application as the dihydrofolate reductase.

[0051] In the present application, even if it is described as 'a polypeptide (or protein) including an amino acid sequence described by a specific sequence number' or 'a polypeptide (or protein) consisting of an amino acid sequence described by a specific sequence number', if it has the same or corresponding activity as the polypeptide (or protein) consisting of the amino acid sequence of the corresponding sequence number, a polypeptide (or protein) having an amino acid sequence in which a part of the sequence is deleted, modified, substituted or added may also be used as the protein or polypeptide that is the target of mutation in the present application. For example, a polypeptide (or protein) having a sequence in which a part of the amino acid sequence of any one of SEQ ID NOs: 1 to 6 is deleted, modified or substituted or having a sequence in which a part of the amino acid sequence of any one of SEQ ID NOs: 1 to 6 is added may belong to 'a polypeptide consisting of an amino acid sequence of any one of SEQ ID NOs: 1 to 6' if it has the same or corresponding activity as 'a polypeptide consisting of an amino acid sequence of any one of SEQ ID NOs: 1 to 6'.

[0052] The amino acid sequence of the above dihydrofolate reductase and the base sequence of the gene encoding the dihydrofolate reductase can be easily obtained from databases known in the art, such as the National Center for Biotechnology Information (NCBI) and the DNA Data Bank of Japan (DDBJ), and may be, for example, GenBank Accession No. J00140.1.

[0053] In the present application, the phrase “a polynucleotide or polypeptide comprises a specific nucleic acid sequence (base sequence) or amino acid sequence” may mean that the polynucleotide or polypeptide consists of or essentially includes the specific nucleic acid sequence (base sequence) or amino acid sequence, and may be interpreted as including a sequence in which a mutation (deletion, substitution, modification, and / or addition) is added to the specific nucleic acid sequence (base sequence) or amino acid sequence within the scope of maintaining the original function and / or desired function of the polynucleotide or polypeptide (or not excluding the mutation). In one example, a polynucleotide or polypeptide “comprising a particular nucleic acid sequence (base sequence) or amino acid sequence” may mean that the polynucleotide or polypeptide (i) consists of or essentially comprises the particular nucleic acid sequence (base sequence) or amino acid sequence, or (ii) consists of or essentially comprises a nucleic acid sequence or amino acid sequence that has at least 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98%, 99.5%, or 99.9% homology to the particular nucleic acid sequence (base sequence) or amino acid sequence and maintains its original function and / or desired function.

[0054] In this application, “homology” or “identity” refers to the degree of similarity between two given amino acid sequences or base sequences, which may be expressed as a percentage. The terms homology and identity are often used interchangeably.

[0055] Sequence homology or identity of conserved polynucleotides or polypeptides is determined by standard alignment algorithms, and may be combined with default gap penalties established by the program being used. In practice, homologous or identical sequences can generally hybridize with all or part of the sequence under moderate or high stringency conditions. It should be appreciated that hybridization also includes hybridization with polynucleotides containing common codons or codons that take codon degeneracy into account.

[0056] Whether any two polynucleotide or polypeptide sequences are homologous, similar or identical can be determined using known computer algorithms such as the "FASTA" program using default parameters, for example as in Pearson et al (1988) [Proc. Natl. Acad. Sci. USA 85]: 2444. Alternatively, the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, J. Mol. Biol. 48: 443-453) as implemented in the Needleman program of the EMBOSS package (EMBOSS: The European Molecular Biology Open Software Suite, Rice et al., 2000, Trends Genet. 16: 276-277) (version 5.0.0 or later) can be determined using the GCG program package (Devereux, J., et al, Nucleic Acids Research 12: 387 (1984)), BLASTP, BLASTN, FASTA (Atschul, [S.][F.,] [ET AL, J MOLEC BIOL 215]: 403 (1990); Guide to Huge Computers, Martin J. Bishop, [ED.,]Academic Press, San Diego, 1994, and [CARILLO ETA / .](1988) SIAM J Applied Math 48: 1073). For example, homology, similarity, or identity can be determined using BLAST or ClustalW of the National Center for Biotechnology Information Database.

[0057] Homology, similarity, or identity of polynucleotides or polypeptides can be determined by comparing sequence information, for example, using a GAP computer program, such as that of Needleman et al. (1970), J Mol Biol. 48:443, as disclosed, for example, in Smith and Waterman, Adv. Appl. Math (1981) 2:482. In brief, the GAP program can be defined as the total number of symbols in the shorter of the two sequences divided by the number of similarly arranged symbols (i.e., nucleotides or amino acids). Default parameters for the GAP program include (1) a binary comparison matrix (containing values ​​of 1 for identity and 0 for non-identity) and (2) a comparison matrix, as disclosed by Gribskov et al. (1986) Nucl. Acids Res. 48:443, as disclosed by Schwartz and Dayhoff, eds., Atlas Of Protein Sequence And Structure, National Biomedical Research Foundation, pp. 353-358 (1979). 14: 6745 weighted comparison matrix (or EDNAFULL (EMBOSS version of NCBI NUC4.4) permutation matrix); (2) a penalty of 3.0 for each gap and an additional penalty of 0.10 for each symbol in each gap (or a gap opening penalty of 10, a gap extension penalty of 0.5); and (3) no penalty for terminal gaps.

[0058] In addition, it is obvious that a variant having an amino acid sequence in which some sequences are deleted, modified, substituted, conservatively substituted or added is also included in the dihydrofolate reductase used in the present invention, provided that the amino acid sequence has such homology or identity and exhibits an effect corresponding to the dihydrofolate reductase used in the present invention. For example, this is the case where the variant of the present application has an addition or deletion of a sequence, a naturally occurring mutation, a silent mutation or a conservative substitution at the N-terminus, C-terminus and / or within the amino acid sequence that does not alter the function of the variant of the present application.

[0059] The term "conservative substitution" refers to the replacement of one amino acid with another amino acid having similar structural and / or chemical properties. Such amino acid substitutions may generally be based on similarities in the polarity, charge, solubility, hydrophobicity, hydrophilicity, and / or amphipathic nature of the residues. Typically, conservative substitutions may have little or no effect on the activity of a protein or polypeptide.

[0060] In this application, the term "variant" refers to a polypeptide in which one or more amino acids are conservatively substituted and / or modified, thereby differing from the amino acid sequence of the variant before the mutation, but retaining functions or properties. Such variants can generally be identified by modifying one or more amino acids in the amino acid sequence of the polypeptide and evaluating the properties of the modified polypeptide. That is, the ability of the variant may be increased, unchanged, or decreased compared to the polypeptide before the mutation. Furthermore, some variants may include variants in which one or more portions, such as the N-terminal leader sequence or the transmembrane domain, are deleted. Other variants may include variants in which portions are deleted from the N- and / or C-terminus of the mature protein. The above term "variant" may be used interchangeably with terms such as variant, modification, variant polypeptide, variant protein, variant and variant (in English, modification, modified polypeptide, modified protein, mutant, mutein, divergent, etc.), and is not limited thereto if the term is used in the meaning of variant.

[0061] Additionally, variants may include deletions or additions of amino acids that have minimal impact on the properties and secondary structure of the polypeptide. For example, the N-terminus of the variant may be conjugated with a signal (or leader) sequence involved in co-translational or post-translational protein translocation.

[0062] Additionally, the above variants may be conjugated with other sequences or linkers to enable identification, purification, or synthesis.

[0063] In the present application, “the ~th amino acid sequence of SEQ ID NO. ~” may be used with the same meaning as “the ~th amino acid sequence from the N-terminus of a polypeptide consisting of (or made of) the amino acid sequence of SEQ ID NO. ~.”

[0064] As used herein, the term "corresponding to" refers to an amino acid residue at a position listed in a polypeptide, or an amino acid residue that is similar, identical, or homologous to the residue listed in the polypeptide. Identifying an amino acid at a corresponding position may be determining a specific amino acid in a sequence that references a particular sequence. As used herein, the term "corresponding region" generally refers to a similar or corresponding position in a related or reference protein.

[0065] In one embodiment of the present invention, when dihydrofolate reductase consisting of 187 amino acids is expressed by dividing it into an N-domain region / C-domain region, the dividing site is E45:G46, which is a position where the 45th amino acid is cut from the N-terminus, and the dividing site is expressed by dividing it. However, even when using a dihydrofolate reductase variant in which the amino acid corresponding to the cleavage site is conservatively substituted, the same sequential position from the N-terminus can be selected as the dividing site.

[0066] The vector of the present invention may include an intein sequence, a leucine zipper or a coiled-coil sequence in order to join or splice the N-domain fragment of the dihydrofolate reductase gene and the C-domain fragment of the dihydrofolate reductase gene that are expressed in a split manner, and the intein, leucine zipper or coiled-coil may be characterized in that it is selected from the group consisting of gp41-1, SspGyrB, Mja-KlbA, Cth-Ter, NpuDnaE, ​​NrdA-2, SspDnaX, gp41-8, NrdJ-1, IMPDH-1, M86, coiled coil, leucine zipper, SH3 and PRM.

[0067] Intein sequences that can be used in the present invention are shown in Table 2.

[0068]

[0069]

[0070] The vector of the present invention may be characterized by additionally including an internal ribosomal entry site (IRES) between a gene encoding a target protein and a split fragment of a dihydrofolate reductase gene.

[0071] In the present invention, the target protein may be characterized as being a protein selected from the group consisting of an Fc-fusion protein, an antibody, and a bispecific antibody, and the vector of the present invention may be characterized as including a gene encoding the target protein.

[0072] Examples of backbone vectors that can be used in the expression vector of the present invention include, but are not limited to, expression vectors such as pcDNA™, pTargeT™, and UCOE; viral vectors such as lentiviral (LV), pAAV, and pRetro; and transposon vectors such as piggyBac (PB), sleeping beauty (SB), and Leap-In™.

[0073] The vector of the present invention may be characterized by including a target gene.

[0074] In the present invention, the target gene may be a gene encoding a first or second monomer of a protein selected from the group consisting of a protein, an Fc-fusion protein, an antibody, and a bispecific antibody, a transgene, a gene encoding mRNA, and an RNAi, but is not limited thereto.

[0075] In the present invention, a “vector” may include a DNA construct comprising a base sequence of a polynucleotide encoding a target polypeptide operably linked to a suitable expression control region (or expression control sequence) so as to enable expression of the target polypeptide in a suitable host. The expression control region may include a promoter capable of initiating transcription, an optional operator sequence for regulating such transcription, a sequence encoding a suitable mRNA ribosome binding site, and a sequence regulating the termination of transcription and translation. After being transformed into a suitable host cell, the vector may replicate or function independently of the host genome, and may be integrated into the genome itself.

[0076] The vector used in this application is not particularly limited, and any vector known in the art may be used. Examples of commonly used vectors include plasmids, cosmids, viruses, and bacteriophages, either in their natural or recombinant form.

[0077] The term "transformation" in this application refers to introducing a vector containing a polynucleotide encoding a target polypeptide into a host cell, thereby enabling expression of the polypeptide encoded by the polynucleotide within the host cell. The transformed polynucleotide may be located within the chromosome of the host cell or located extrachromosomally, as long as it can be expressed within the host cell. Furthermore, the polynucleotide includes DNA and / or RNA encoding the target polypeptide. The polynucleotide may be introduced in any form as long as it can be introduced into the host cell and expressed. For example, the polynucleotide may be introduced into the host cell in the form of an expression cassette, which is a genetic construct containing all elements necessary for autonomous expression. The expression cassette may typically include a promoter, a transcription termination signal, a ribosome binding site, and a translation termination signal, all of which are operably linked to the polynucleotide. The expression cassette may be in the form of a self-replicating expression vector. Additionally, the polynucleotide may be introduced into a host cell in its own form and operably linked to a sequence necessary for expression in the host cell, but is not limited thereto.

[0078] Additionally, the term "operably linked" as used herein means that the polynucleotide sequence is functionally linked to a promoter sequence that initiates and mediates transcription of the polynucleotide encoding the target variant of the present application.

[0079] In another aspect, the present invention relates to a cell for producing a target protein or a target gene, wherein the first vector and the second vector are introduced, and the first vector and / or the second vector further comprise a gene encoding a target protein, wherein the first vector and the second vector are introduced such that when the N-domain portion of the first vector and the C-domain portion of the second vector are combined with each other, a complete dihydrofolate reductase is formed, and the first vector and / or the second vector comprises an intein, a leucine zipper, or a coiled coil.

[0080] For example, in the cell for producing the target protein, a first vector including the DHFR N-domain region of SEQ ID NO: 1 can be transduced together with a second vector including the DHFR C-domain region of SEQ ID NO: 4,

[0081] A first vector comprising a DHFR N-domain region of SEQ ID NO: 2 can be transduced together with a second vector comprising a DHFR C-domain region of SEQ ID NO: 5,

[0082] A first vector comprising the DHFR N-domain region of SEQ ID NO: 3 can be transduced together with a second vector comprising the DHFR C-domain region of SEQ ID NO: 6.

[0083] In the present invention, all types of cells that can be used for producing a target protein or target gene may be used, and for example, prokaryotes, insect cells, yeast, or animal cells may be used, and preferably, myeloma cell lines, CHO, HEK293, NS0, VERO, BHK, HeLa, Cos, MDCK, 293, 3T3, and WI38 cells, human-derived primary cells, immune cells, stem cells, T-cells, B-cells, NK-cells, macrophages, etc. derived from peripheral blood mononuclear cells may be used.

[0084] In another aspect, the present invention relates to a method for selecting cells expressing a target protein or target gene, comprising the following steps:

[0085] (a) a step of producing a library of cells for producing the target protein or target gene;

[0086] (b) culturing the library in a medium containing a dihydrofolate reductase inhibitor and selecting growing recombinant cells; and

[0087] (C) A step of confirming the expression of the target protein in the selected recombinant cells and obtaining cells expressing the target protein.

[0088] In the present invention, the dihydrofolate reductase inhibitor may be MSX (methionine sulphoximine), but is not limited thereto.

[0089] In the present invention, the target protein may be a protein, an Fc-fusion protein, an antibody, or a bispecific antibody, but is not limited thereto.

[0090] In another aspect, the present invention relates to a method for producing a target protein or target gene, comprising the following steps:

[0091] (a) a step of culturing cells for producing the target protein or target gene to produce the target protein; and

[0092] (b) A step of obtaining the target protein produced above.

[0093] In the present invention, the target protein may be a protein, an Fc-fusion protein, an antibody, or a bispecific antibody, but is not limited thereto.

[0094] In the present invention, the term "dihydrofolate reductase (DHFR)" refers to an enzyme involved in the reaction of reducing 7,8-dehydrofolate to 5,6,7,8-tetrahydrofolate using NADPH as an electron donor, which exists in mammalian organs or microorganisms. The activity of the enzyme is inhibited by folate antagonists such as aminopterin and methotrexate. For the purpose of the present invention, the dihydrofolate reductase may refer to a protein used for the purpose of selecting cells into which a vector including a gene encoding a target protein has been introduced in a target protein expression system, or for the purpose of increasing the expression of a gene encoding a target protein that is in a polycistronic relationship with a gene encoding dihydrofolate reductase by being inhibited by a dihydrofolate reductase inhibitor. Preferably, a protein capable of increasing the expression level of a target protein by inhibiting dihydrofolate reductase inhibitor, but is not limited thereto.

[0095] In the present invention, the term "dihydrofolate reductase inhibitor" refers to an external factor capable of inhibiting the enzymatic activity of the dihydrofolate reductase. Examples of the dihydrofolate reductase inhibitor include, but are not limited to, aminopterin, methotrexate, and MSX (methionine sulphoximine).

[0096] In the present invention, the term "sensitivity" generally refers to the characteristic of accepting external stimuli, and in terms of enzymes, refers to the characteristic of enzyme activity being enhanced or decreased by external factors that regulate enzyme activity. For the purposes of the present invention, the sensitivity is used to refer to the characteristic of DHFR activity being inhibited by a DHFR inhibitor that inhibits DHFR activity, but is not limited thereto.

[0097] In the present invention, the term "target protein" means a protein to be produced in a host cell, and for the purpose of the present invention, means a protein whose expression is amplified by a mutated dihydrofolate reductase, but is not limited thereto, and the type of protein that can be expressed using the vector of the present invention is not particularly limited.

[0098] In the present invention, the term "target gene" means a gene encoding the target protein, a transgene, mRNA, or RNA for RNA interference (RNAi, shRNA, siRNA, etc.).

[0099] As used herein, the term "expression vector" refers to a genetic construct that includes essential regulatory elements operably linked to an insert so that the insert is expressed when present in a cell of a subject. The expression vector can be produced and purified using standard recombinant DNA technology. The type of the expression vector is not particularly limited as long as it functions to express a desired gene and produce a desired protein in various host cells of prokaryotic and eukaryotic cells. However, a vector that exhibits a strong activity and possesses a strong expression ability while being able to produce a large amount of a foreign protein in a form similar to that in the natural state is preferred. The expression vector preferably includes at least a promoter, an initiation codon, a gene encoding a desired protein, and a stop codon terminator. In addition, it may also appropriately include DNA encoding a signal peptide, an enhancer sequence, untranslated regions on the 5' and 3' sides of a desired gene, a selectable marker region, or a replicative unit. In addition, the form of the expression vector may be a monocistronic vector including a polynucleotide encoding one protein, a polycistronic vector including a polynucleotide encoding two or more recombinant proteins, and in particular, a bicistronic vector in which the target protein, the mutant dihydrofolate reductase, is expressed by one promoter as a bicistronic vector. However, for the purpose of the present invention, the expression vector may preferably be a monocistronic vector including an SV40 promoter or a polycistronic vector including an IRES base sequence, and more preferably, an expression cassette in which a gene encoding a target protein, an IRES, and a gene encoding a mutant dihydrofolate reductase are sequentially operably linked; or an expression vector including an expression cassette in which a gene encoding a mutant dihydrofolate reductase, an IRES, and a gene encoding a target protein are sequentially operably linked may be used, but is not limited thereto.

[0100] As used herein, the term "IRES" refers to an internal ribosome entry site that facilitates the initiation of translation of an mRNA from an internal region (i.e., a region other than the 5' end of the mRNA). An example of a suitable IRES is the IRES of encephalomyocarditis virus (ECMV), as described in Jang and Wimmer Genes & Development 4 1560 (1990) and Jang, Davies, Kaufman and Wimmer J. Vir. 63 1651 (1989). Residues 335-848 of EMCV form a suitable IRES; other variants or portions of the ECMV IRES are known and may be suitable for use in the present invention. Suitable portions or variants of the IRES will confer sufficient translation of the second open reading frame (ORF).

[0101] In the expression vector of the present invention, the expression of the target protein can be linked to the DHFR fragment via IRES, and when linked to IRES, the target protein must be expressed in order for the DHFR fragment to be expressed.

[0102] The recombinant vector comprising the split DHFR according to the present invention as a selection marker can improve the selection efficiency of a cell line carrying the recombinant vector, and this improved selection marker can be used as a selection marker for vectors comprising, for example, a secreted protein expression gene, an endogenous gene, a transgene, RNA, etc. Suitable examples thereof include, but are not limited to, the selection of vectors comprising Fc-fusion proteins, antibodies, biosimilars, bispecific antibodies, cytokines, growth factors, reporter genes, T-cell receptors, chimeric antigen receptors, messenger RNA (mRNA), RNA interference, etc.

[0103]

[0104] [Example]

[0105] Hereinafter, the present invention will be described in more detail through examples. These examples are intended solely to illustrate the present invention, and it will be apparent to those skilled in the art that the scope of the present invention is not limited by these examples.

[0106]

[0107] Example 1: Split position selection for a split expression system of dihydrofolate reductase.

[0108] To select a cleavage site for the cleavage expression of dihydrofolate reductase, the root mean square fluctuation, the 3D structural location of the cleavage site, and the evolutionary conservation of the amino acid sequence were analyzed.

[0109] For structural analysis, the following three programs were used.

[0110] 1) RMSF calculation (CABS-flex),

[0111] 2) Confirmation of tertiary structure location (pymol, https: / pymol.org / 2 / ),

[0112] 3) Evolutionary Trace

[0113] As a result, NC(N45 / C46): E45:G46 was selected as the cleavage site in dihydrofolate reductase consisting of 187 amino acids.

[0114] The selected cleavage positions on the three-dimensional structure of dihydrofolate reductase are shown in Figure 1.

[0115]

[0116] Example 2: Construction of a vector for expression of a recombinant protein containing a fragment of dihydrofolate reductase

[0117] The vector for recombinant protein expression was constructed by removing the existing selection marker from the pcDNA vector (pcDNA™3.1 / Hygro(+), ThermoFisher Scientific) through polymerase chain reaction and KLD reaction, and introducing the IRES and dihydrofolate reductase selection marker using Gibson assembly. Afterwards, the target gene to be expressed was introduced between the CMV promoter and IRES using Gibson assembly to construct the expression vector.

[0118] A recombinant protein expression vector containing the dihydrofolate reductase 2-fragment produced by the above method is shown in Fig. 3, where DHFRN represents the N-domain cleavage fragment of dihydrofolate reductase, and DHFR C represents the C-domain cleavage fragment of dihydrofolate reductase.

[0119]

[0120] Example 3: Confirmation of cell growth and cell viability according to DHFR division.

[0121] The cell line used in this example was a DHFR knockout cell line, and cells with the DHFR gene knocked out of CHO-K1 (ATCC CCL-61) were used.

[0122] A cell line was created by transducing the first and second vectors for target protein expression according to the DHFR cleavage produced in Example 2, and cell growth and cell viability were confirmed.

[0123] As a result, as shown in Fig. 4, cells into which both the first vector containing the N-domain cleavage fragment of dihydrofolate reductase and the second vector containing the C-domain cleavage fragment of dihydrofolate reductase were introduced grew normally.

[0124] Similarly, as shown in Fig. 5, cells introduced with both the first vector containing the N-domain cleavage fragment of dihydrofolate reductase and the second vector containing the C-domain cleavage fragment of dihydrofolate reductase showed normal survival rates.

[0125]

[0126] Example 4: Confirmation of recombinant protein expression level in selected cells

[0127] In the same manner as in Example 2, a vector for expression of a recombinant protein was produced, and in order to confirm the amount of recombinant protein expression in a cell line into which the vector for expression of the recombinant protein was introduced, etanercept, an FC fusion protein, was used.

[0128] The vector containing the Etanercept protein coding gene, represented by SEQ ID NO: 29 in Table 3, was introduced into a CHO-K1 DHFR knockout cell line using a transfection reagent. The cell line introduced with the vector was cultured in suspension in PowerCHO™ 2 Serum-free Medium (Lonza) at 37°C and 110 rpm for 2 days. After selection, the protein expression level of the selected cells was measured using a flow cytometer using a fluorescently conjugated antibody. In addition, the amount of Etanercept produced during the culture period was measured using an enzyme-linked immunosorbent assay (ELISA).

[0129]

[0130] After selecting cells expressing etanercept, a FC fusion protein, and confirming the ratio of cell lines highly expressing etanercept within the cell pool, as shown in Fig. 6, the cell lines of the system using the split DHFR of the present invention showed a higher etanercept production compared to the conventional method (WT and WT / Kanamycin).

[0131] Similarly, as shown in Fig. 7, the system using the split DHFR of the present invention showed a higher ratio of high-expressing cell lines compared to the conventional method (WT and WT / Kanamycin).

[0132]

[0133] According to the present invention, the efficiency of selecting a cell line for producing a biopharmaceutical can be dramatically increased, and since it is an improved version of the existing selection marker, dihydrofolate reductase, it can be used freely from stability issues.

[0134]

[0135] While specific aspects of the present invention have been described in detail above, it will be apparent to those skilled in the art that these specific descriptions merely represent preferred embodiments and are not intended to limit the scope of the present invention. Therefore, the substantial scope of the present invention is defined by the appended claims and their equivalents.

[0136]

[0137] Electronic file attached.

Claims

1. A first vector comprising a first marker nucleic acid encoding an N-domain portion of dihydrofolate reductase having an amino acid sequence represented by any one of sequence numbers 1 to 3 as a selection marker.

2. A first vector for producing a recombinant protein, characterized in that the N-domain portion of the dihydrofolate reductase in the first paragraph is represented by sequence number 1.

3. A first vector according to claim 1, characterized in that it further comprises an intein, a leucine zipper or a coiled coil.

4. A first vector, characterized in that in the third paragraph, the intein, leucine zipper or coiled coil is selected from the group consisting of gp41-1, SspGyrB, Mja-KlbA, Cth-Ter, NpuDnaE, ​​NrdA-2, SspDnaX, gp41-8, NrdJ-1, IMPDH-1, M86, coiled coil, leucine zipper, SH3 and PRM.

5. A first vector according to claim 1, characterized in that it further comprises an internal ribosomal entry site (IRES).

6. A first vector characterized in that it comprises a gene encoding a first monomer of a protein selected from the group consisting of a protein, an Fc-fusion protein, an antibody, and a bispecific antibody, and a target gene selected from the group consisting of a transgene, an mRNA, and a gene encoding RNAi.

7. A second vector comprising a second marker nucleic acid encoding a C-domain portion of dihydrofolate reductase having an amino acid sequence represented by any one of SEQ ID NOs: 4 to 6 as a selection marker.

8. A second vector according to claim 7, characterized in that the C-domain portion of the dihydrofolate reductase is represented by sequence number 4.

9. A second vector according to claim 7, characterized in that it further comprises an intein, a leucine zipper or a coiled coil.

10. A second vector according to claim 9, characterized in that the intein, leucine zipper or coiled coil is selected from the group consisting of gp41-1, SspGyrB, Mja-KlbA, Cth-Ter, NpuDnaE, ​​NrdA-2, SspDnaX, gp41-8, NrdJ-1, IMPDH-1, M86, coiled coil, leucine zipper, SH3 and PRM.

11. A second vector according to claim 7, characterized in that it additionally comprises an internal ribosomal entry site (IRES).

12. A second vector according to claim 7, characterized in that it comprises a gene encoding a second monomer of a protein selected from the group consisting of a protein, an Fc-fusion protein, an antibody, and a bispecific antibody, and a target gene selected from the group consisting of a transgene, an mRNA, and a gene encoding RNAi.

13. The first vector of paragraph 1 and the second vector of paragraph 7 are introduced, and the first vector and / or the second vector additionally include a gene encoding a target protein, The first vector and the second vector are introduced so that when the N-domain portion of the first vector and the C-domain portion of the second vector are combined with each other, a complete dihydrofolate reductase is formed. A cell for producing a target protein or target gene, characterized in that the first vector and / or the second vector comprises an intein, a leucine zipper or a coiled coil.

14. A cell for producing a target protein or target gene, characterized in that the cell in claim 13 is a prokaryotic cell, an insect cell, a yeast cell, or an animal cell.

15. A cell for producing a target protein or target gene, characterized in that the cell is selected from the group consisting of myeloma cell lines, CHO, NS0, VERO, BHK, HeLa, Cos, MDCK, 293, 3T3, and WI38 cells, according to claim 13.

16. A method for selecting cells expressing a target protein or target gene, comprising the following steps: (a) a step of producing a library of cells for producing the target protein or target gene of Article 13; (b) culturing the library in a medium containing a dihydrofolate reductase inhibitor and selecting growing recombinant cells; and (C) A step of confirming the expression of a target protein or target gene in the selected recombinant cell and obtaining a cell expressing the target protein or target gene.

17. A method according to claim 16, characterized in that the dihydrofolate reductase inhibitor is MSX (methionine sulphoximine).

18. A method according to claim 16, characterized in that the target gene is selected from the group consisting of a gene encoding a second monomer of a protein selected from the group consisting of a protein, an Fc-fusion protein, an antibody, and a bispecific antibody, a transgene, a gene encoding mRNA, and a RNAi.

19. A method for producing a target protein or target gene comprising the following steps: (a) a step of culturing a cell for producing the target protein of Article 13 to produce the target protein or target gene; and (b) A step of obtaining the target protein or target gene generated above.

20. A method according to claim 19, characterized in that the target gene is selected from the group consisting of a gene encoding a second monomer of a protein selected from the group consisting of a protein, an Fc-fusion protein, an antibody, and a bispecific antibody, a transgene, a gene encoding mRNA, and a gene encoding RNAi.

Citation Information

Patent Citations

  • Universal Fixing Device

    KR1020200088646A

  • Manufacturing method of nurungji and nurangji made by the same that

    KR1020240167332A