Antibody manufacturing methods

By targeting and removing unpaired donor splice sites in the variable domain of antibody heavy chains during codon optimization, the method enhances antibody expression yield and reduces missplicing, specifically for human IgG1 subclass antibodies.

JP7835713B2Active Publication Date: 2026-03-25F HOFFMANN LA ROCHE & CO AG
View PDF 16 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-09-25
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing codon optimization processes for producing antibodies inadvertently generate unpaired donor splice sites, leading to undefined splicing events and reduced expression yields.

Method used

Remove unpaired donor splice sites specifically in the nucleic acid portion encoding the variable domain of the antibody heavy chain, and introduce silent nucleotide changes in the unpaired donor splice site consensus sequence NGGTA(G)AG to improve expression yield.

Benefits of technology

This approach increases the expression yield of antibodies by minimizing missplicing and ensuring correct antibody chain length, particularly for human IgG1 subclass antibodies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007835713000001
    Figure 0007835713000001
  • Figure 0007835713000002
    Figure 0007835713000002
  • Figure 0007835713000003
    Figure 0007835713000003
Patent Text Reader

Abstract

To provide a method for producing an IgG1 antibody by cultivating a CHO cell comprising / transfected with one or more (exogenous) nucleic acids encoding (and expressing) the antibody.SOLUTION: Disclosed is a method for producing a human IgG1 subclass antibody by culturing a CHO cell comprising one or more expression cassettes comprising nucleic acids encoding antibody heavy and light chains, where the nucleic acids are optimized for the codon usage of human cells and / or for the codon usage of CHO cells, where in the nucleic acid encoding the heavy chain variable domain at least one non-paired donor splice site is removed and in the nucleic acid sequence encoding the heavy chain constant region, non-paired donor splice sites are not removed.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification reports a method for the production of antibodies in which the coding nucleic acid is optimized only with respect to donor splice sites in the portion encoding the variable domain.

Background Art

[0002] Cannarozzi, G., et al. reported the role of codon order in translational dynamics (Cell 141 (2010) 355-367 (Non-Patent Document 1)). The causes and consequences of codon bias were reported by Plotkin, J. B. and Kudla, G. (Nat. Rev. Gen. 12 (2011) 32-42 (Non-Patent Document 2)). Weygand-Durasevic, I. and Ibba, M. reported a new role for codon usage (Science 329 (2010) 14...

[0003] WO 97 / 11086 (Patent Document 1) reports high-level expression of proteins. Plant polypeptide production is reported in WO 03 / 70957 (Patent Document 2). WO 03 / 85114 (Patent Document 3) describes a method for designing synthetic nucleic acid sequences for optimal protein expression in host cells. Codon pair optimization is reported in US Pat. No. 5,082,767 (Patent Document 4). WO 2008 / 000632 (Patent Document 5) pamphlet reports a method for achieving improved polypeptide expression. Codon optimization methods are reported in WO 2007 / 142954 (Patent Document 6) and US Pat. No. 8,128,938 (Patent Document 7).

[0004] Watkins, NE, et al. have reported the nearest neighbor thermodynamics of deoxyinosine pairs in DNA double helix (Nucl. Acids Res. 33 (2005) 6258-6267 (Non-Patent Literature 5)).

[0005] International Publication No. 2013 / 156443 (Patent Document 8) reports a method for expressing polypeptides using modified nucleic acids.

[0006] Zhang, MQ reported the statistical characteristics of human exons and their adjacent regions (Hum.Mol.Genet.7(1998)919-932(Non-Patent Literature 6)).

[0007] mRNA splicing is regulated by the development of donor splice sites, which are combined with acceptor splice sites located at the 5' and 3' ends of introns, respectively. According to Watson et al. (Watson et al. (Eds), Recombinant DNA: Short course, Scientific American Books, distributed by WH Freeman and Company, New York, New York, USA (1983) (Non-Patent Literature 7)), the 5' donor splice site ag|gtragt (exon|intron) and the 3' acceptor splice site (y) n This is the consensus sequence for Ncag|g (intron|exon) (r = purine base; y = pyrimidine base; n = integer; N = any native base).

[0008] In 1980, the first paper dealing with the origins of the secretory and membrane-bound forms of immunoglobulins was published. The formation of secretory (sIg) and membrane-bound (mIg) isoforms arises from alternative splicing of heavy chain pre-mRNA. In the mIg isoform, the secretory form (i.e., C, respectively) H 3 or C HThe donor splice site in the exon encoding the C-terminal domain of the 4 domains, and the acceptor splice site located downstream thereof, are used to connect the constant region to the downstream exon encoding the transmembrane domain.

[0009] A method for preparing synthetic nucleic acid molecules with undesirable or unintended reduced transcriptional properties when expressed in specific host cells is reported in International Publication No. 2002 / 016944 (Patent Document 9). International Publication No. 2006 / 042158 (Patent Document 10) reports nucleic acid molecules modified to enhance recombinant protein expression and / or reduce or eliminate misspliced ​​and / or introns read by the product. Magistrelli, G., et al. reported the optimization of native bispecific antibody assembly and production by codon de-optimization (MABS 9(2016)231-239 (Non-Patent Document 8)).

[0010] International Publication No. 2015 / 128509 (Patent Document 11) reported an expression construct and method for selecting host cells that express a polypeptide.

[0011] International Publication No. 2009 / 003623 (Patent Document 12) reported a heavy chain variant that resulted in improved immunoglobulin production. [Prior art documents] [Patent Documents]

[0012] [Patent Document 1] International Publication No. 97 / 11086 [Patent Document 2] International Publication No. 03 / 70957 [Patent Document 3] International Publication No. 03 / 85114 [Patent Document 4] U.S. Patent No. 5,082,767 [Patent Document 5] International Publication No. 2008 / 000632

Patent document 6

Patent document 7

Patent document 8

Patent Document 9

Patent document 10

Patent document 11

Patent document 12

Non-licensed literature

[0013] [Non-licensed document 1] Cell 141(2010)355-367 [Non-licensed document 2] Nat.Rev.Gen.12(2011)32-42 [Non-licensed document 3] Science 329(2010)1473-1474

Non-licensed Document 4

Non-licensed Document 5

Non-licensed Document 6

Non-licensed Document 7

Summary of the Invention

[0014] For the production of therapeutic or diagnostic antibodies, a high expression yield is targeted. One common option for achieving a good expression rate is first to optimize the codon usage of the coding nucleic acid by adjusting it to the codon usage of cells intended to express the exogenous nucleic acid. Such codon adaptation or optimization can be carried out based on various established protocols.

[0015] However, during such codon adaptation and optimization processes, for example, unpaired splice sites, especially unpaired donor splice sites, can be newly generated inadvertently. That is, for example, during codon optimization, a new donor splice site sequence is generated within the codon-optimized nucleic acid by inadvertently generating a sequence motif within the codon-optimized nucleic acid following the donor splice site consensus sequence. Such events can occur independently of the compilation of the codon-optimized nucleic acid, that is, for both cDNA and genomically compiled nucleic acids. In fact, this is an unintended secondary result of the codon usage optimization process. Such new donor splice sites are additional artificial donor splice sites and do not have the associated target acceptor splice sites. Therefore, such unpaired donor splice sites can cause splicing events with random, that is, undefined acceptor splice sites somewhere in the transcribed mRNA. This results in the formation of by-products, which reduces the expression yield.

[0016] The present invention is at least in part based on the unexpected discovery that the removal of unpaired donor splice sites in antibody heavy chains with optimized codon usage encoding nucleic acids only needs to be performed in the nucleic acid portion encoding the variable domain of the heavy chain, and not in the portion encoding the constant region; i.e., in the case of the constant region, for example, germline or wild-type human nucleic acid sequences can be used. This makes it possible, or even possible, to increase the expression yield of antibody heavy chains with the correct length.

[0017] The present invention is at least in part based on the finding that, only in codon-optimized nucleic acids encoding the variable domain of an antibody heavy chain, the introduction of silent nucleotide changes (mutations) in the unpaired donor splice site consensus sequence NGGTA(G)AG (SEQ ID NO: 01) is sufficient to improve the expression yield.

[0018] One aspect of the present invention is a method for producing antibodies by culturing mammalian cells containing / transfected with one or more (exogenous) nucleic acids that encode (and express antibodies) antibody heavy chains and antibody light chains. One or more (exogenous) nucleic acids are codon uses optimized for codon use in human cells and / or mammalian cells. In nucleic acids encoding the heavy chain variable domain, at least one (artificial) unpaired donor splice site has been removed, and optionally, in nucleic acid sequences encoding the heavy chain constant region (optimized for human wild-type or human or hamster codon use), the (artificial) unpaired donor splice site has not been removed.

[0019] In one embodiment, the antibody is an antibody of the human IgG1 subclass. In another embodiment, the antibody is a humanized antibody of the human IgG1 subclass. In another embodiment, the constant region of the antibody contains mutations suitable for inducing heterodimerization or modifying Fc receptor binding.

[0020] In one embodiment, the mammalian cells are CHO cells.

[0021] In one embodiment, the transfection is transient.

[0022] In one embodiment, all of the one or more (exogenous) nucleic acids encoding the antibody are cDNA.

[0023] In one embodiment, one or more (exogenous) nucleic acids encoding an antibody heavy chain and / or antibody light chain are genomically organized DNA, i.e., having an intron-exon configuration.

[0024] In one embodiment, the removal of the unpaired donor splice site is achieved by introducing an amino acid silent change (mutation) into the amino acid sequence NGGTA(G)AG (SEQ ID NO: 01). In one embodiment, the amino acid sequence silent nucleotide change is introduced into codon NGG, codon GGT, or codon GTA(G) of SEQ ID NO: 01.

[0025] In one embodiment, codon usage optimization is performed based on human codon usage or Chinese hamster codon usage.

[0026] In one embodiment, the nucleic acid encoding the antibody light chain is optimized for codon use, i.e., the variable domain and constant region are optimized for codon use.

[0027] In one embodiment, the unpaired donor splice site is removed from the nucleic acid encoding the complete light chain.

[0028] In one embodiment, the transfection is a stable transfection.

[0029] In one embodiment, the method is: a) A process of culturing mammalian cells, b) a step of recovering the antibody from the cells or the culture medium.

[0030] One aspect of the present invention is the use of removing unpaired donor splice sites in only the portion of the antibody-encoding human or hamster codon-use optimized nucleic acid sequence to reduce missplicing and / or increase antibody expression yield when nucleic acids are used to produce antibodies in CHO cells, the portion being the portion encoding the heavy chain variable domain.

[0031] In one embodiment, the unpaired donor splice site is not removed in the nucleic acid portion encoding the heavy chain constant region.

[0032] In one embodiment, the unpaired donor splice site is further removed from the nucleic acid portion encoding the light chain.

[0033] In one embodiment, the removal of the unpaired splice site is achieved by introducing an amino acid silent mutation into the nucleotide sequence NGGTA(G)AG (SEQ ID NO: 01).

[0034] In one embodiment, the removal of the unpaired splice site is achieved by introducing an amino acid silent mutation in the nucleotide sequence NGGTA(G)AG (SEQ ID NO: 01) at codon NGG, codon GGT, or codon GTA(G).

[0035] In all aspects and in one embodiment of the embodiments, the unpaired (donor) splice site is an artificial unpaired (donor) splice site.

[0036] In all aspects and in one embodiment of the embodiments, the unpaired (donor) splice site is an artificial unpaired (donor) splice site that is generated during codon optimization. [Invention 1001] A method for producing antibodies of a human IgG1 subclass by culturing CHO cells transfected with one or more expression cassettes containing nucleic acids encoding the heavy and light chains of an antibody, The nucleic acids encoding the heavy and light chains of the antibody are codon-optimized for codon use in human cells and / or CHO cells. A method wherein at least one unpaired donor splice site is removed from the portion of the nucleic acid that encodes the heavy chain variable domain. [Invention 1002] The method of the present invention 1001, wherein the unpaired donor splice site is not removed in the portion of the nucleic acid that encodes the heavy chain constant region. [Invention 1003] The method of the present invention 1001 or 1002, wherein the unpaired splice site is removed from the nucleic acid encoding the antibody light chain. [Invention 1004] A method according to any of the present invention 1001 to 1003, wherein the transfection is transient transfection. [Invention 1005] The method according to any of the present invention 1001 to 1004, wherein the nucleic acid encoding the antibody is cDNA. [Invention 1006] The method according to any one of the present invention 1001 to 1004, wherein the nucleic acid encoding the light chain and / or heavy chain of the antibody is genome-organized DNA. [Invention 1007] The removal of the aforementioned unpaired splice portion is Removal of the unpaired donor splice site by introducing silent nucleotide changes in the amino acid sequence of the aforementioned unpaired donor splice site nucleic acid sequence. A method according to any of the present invention 1001 to 1006. [Invention 1008] The removal of the unpaired splice site is achieved by introducing an amino acid silent mutation into the nucleotide sequence NGGTA(G)AG (SEQ ID NO: 01), in any of the methods described in items 1001 to 1007 of the present invention. [Invention 1009] The removal of the unpaired splice site is achieved by introducing an amino acid silent mutation in the nucleotide sequence NGGTA(G)AG (SEQ ID NO: 01) at codon NGG, codon GGT, or codon GTA(G), according to any method 1001 to 1008 of the present invention. [Invention 1010] The aforementioned method is as follows: a) A step of culturing the CHO cells, and b) A step of recovering the antibody from the CHO cells or the culture medium. A method of the present invention, including any of the methods described in items 1001 to 1009. [Invention 1011] A method for producing antibodies of a human IgG1 subclass by culturing CHO cells containing one or more nucleic acids encoding the heavy and light chains of an antibody, In nucleic acids encoding heavy chain variable domains, the unpaired donor splice site is removed by introducing silent nucleotide changes in the amino acid sequence at codon NGG, codon GGT, or codon GTA(G) into the unpaired donor splice site consensus sequence NGGTA(G)AG (SEQ ID NO: 01). A method in which the unpaired donor splice site by the sequence of SEQ ID NO: 01 is not removed in a nucleic acid encoding a heavy chain constant region. [Invention 1012] The method of the present invention 1011, wherein the one or more nucleic acids encoding the heavy chain and light chain of the antibody are genome-organized DNA. [Invention 1013] The method of the present invention 1011 or 1012, wherein the one or more nucleic acids encoding the heavy and light chains of the antibody are codon-used optimized for codon-used human cells and / or codon-used CHO cells. [Invention 1014] A method of the present invention, any of items 1011 to 1013, wherein, in the nucleic acid encoding the antibody light chain, the unpaired splice site is removed by introducing a silent nucleotide change in the amino acid sequence at codon NGG, codon GGT, or codon GTA(G) in the unpaired donor splice site consensus sequence NGGTA(G)AG (SEQ ID NO: 01). [Invention 1015] The use of removing unpaired donor splice sites in portions of antibody-coding nucleic acid sequences, optimized for the use of human or hamster codons, Therefore, the aforementioned portion is the part that codes for the heavy chain variable domain, The use of the nucleic acid in CHO cells to reduce missplicing and / or increase antibody expression yield. [Invention 1016] Use of the present invention 1015, wherein the unpaired splice site has not been removed in the nucleic acid portion encoding the heavy chain constant region. [Invention 1017] Use of the present invention 1015 or 1016, wherein the unpaired splice site is removed in the nucleic acid portion encoding the light chain. [Invention 1018] The removal of the unpaired splice site is achieved by introducing an amino acid silent mutation into the nucleotide sequence NGGTA(G)AG (SEQ ID NO: 01), using any of the inventions 1015 to 1017. [Invention 1019] The removal of the unpaired splice site is achieved by introducing an amino acid silent mutation in the nucleotide sequence NGGTA(G)AG (SEQ ID NO: 01) at codon NGG, codon GGT, or codon GTA(G), using any of the invention 1015 to 1018. [Invention 1020] The method or use of any of the present invention 1001 to 1019, wherein the unpaired splice site is an artificial unpaired splice site. [Invention 1021] The method or use of any of the present invention 1001 to 1019, wherein the unpaired donor splice site is an artificial donor splice site that is generated during codon optimization. [Modes for carrying out the invention]

[0037] Detailed description of embodiments of the present invention The goal is to achieve high expression yields for the production of therapeutic or diagnostic antibodies. One option for achieving good expression rates is to first optimize the codon usage of the coding nucleic acid and then adapt it to the codon usage of cells intended to express the exogenous nucleic acid. This codon matching or optimization can be performed based on various established protocols.

[0038] However, during such codon matching and optimization processes, for example, unpaired splice sites can be generated de novo. That is, during codon optimization, a new donor splice site sequence is generated in the codon-use optimized nucleic acid. This is independent of the organization of the codon-use optimized nucleic acid, i.e., it is possible for both cDNA and genome-organized nucleic acids. In fact, this is an unintended by-effect of the codon-use optimization process. Such a new donor splice site is an additional artificial donor splice site and does not have an associated target acceptor splice site. Therefore, such an unpaired donor splice site can result in a splicing event with a random, i.e., undefined acceptor splice site located somewhere in the transcribed mRNA. This leads to a decrease in expression yield.

[0039] The present invention is at least in part based on the unexpected finding that the removal of unpaired donor splice sites in the nucleic acid encoding the heavy chain of a codon-use optimized antibody needs to be performed only in the portion of the nucleic acid encoding the variable region of the heavy chain, and not in the portion encoding the constant region; that is, in the case of the constant region, for example, germline or wild-type human nucleic acid sequences can be used.

[0040] The present invention is at least in part based on the finding that, only in codon-optimized nucleic acids encoding the variable domain of an antibody heavy chain, the introduction of silent nucleotide changes (mutations) in the unpaired donor splice site consensus sequence NGGTA(G)AG (SEQ ID NO: 01) is sufficient to improve the expression yield.

[0041] definition Methods and techniques useful for carrying out the present invention are known to those skilled in the art and are described, for example, in Ausubel, FM, ed., Current Protocols in Molecular Biology, Volumes I to III (1997), and Sambrook et al., Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (1989). As is known to those skilled in the art, the use of recombinant DNA technology allows for the production of numerous derivatives of nucleic acids and / or polypeptides. Such derivatives can be modified at one or more sites, for example, by substitution, alteration, exchange, deletion, or insertion. Such modification or derivatization can be carried out, for example, by site-directed mutagenesis. Such modifications can be easily performed by those skilled in the art (see, for example, Sambrook, J., et al., Molecular Cloning: A Laboratory Manual (1999), Cold Spring Harbor Laboratory Press, New York, USA). The use of recombination techniques allows those skilled in the art to transform various host cells with different nucleic acids. Although the transcription and translation, i.e., expression, mechanisms of different cells use the same elements, cells belonging to different species may have, in particular, different so-called codon usages. Thus, the same polypeptide (with respect to the amino acid sequence) may be encoded by different nucleic acids. Also, due to the degeneracy of the genetic code, different nucleic acids may encode the same polypeptide.

[0042] The term "approximately" indicates that the following value is not an exact value, but rather the center point of a range that is + / - 10%, + / - 5%, + / - 2%, or + / - 1% of the value. If the value is a relative value given as a percentage, the term "approximately" also indicates that the following value is not an exact value, but rather the center point of a range that is + / - 10%, + / - 5%, + / - 2%, or + / - 1%, thereby indicating that the upper limit of the range cannot exceed 100%.

[0043] As used in this application, the term “amino acid” refers to a group of carboxy-α amino acids that can be encoded by nucleic acids, either directly or in the form of precursors. Each amino acid is encoded by a nucleic acid consisting of three nucleotides, a so-called codon or base triplet. Each amino acid is encoded by at least one codon. The encoding of the same amino acid by different codons is known as “modification of the genetic code.” As used in this application, the term "amino acid" refers to naturally occurring carboxy-α-amino acids, and includes the group of naturally occurring carboxy-α-amino acids, including alanine (3-letter code: ala, 1-letter code: A), arginine (arg, R), asparagine (asn, N), aspartic acid (asp, D), cysteine ​​(cys, C), glutamine (gln, Q), glutamic acid (glu, E), glycine (gly, G), histidine (his, H), isoleucine (ile, I), leucine (leu, L), lysine (lys, K), methionine (met, M), phenylalanine (phe, F), proline (pro, P), serine (ser, S), threonine (thr, T), tryptophan (trp, W), tyrosine (tyr, Y), and valine (val, V).

[0044] The term "immunoglobulin" as used herein is used in its broadest sense and encompasses a variety of immunoglobulin structures, including, but not limited to, monoclonal antibodies, polyclonal antibodies, and multispecific antibodies (e.g., bispecific antibodies) or fragments thereof that include at least a portion of a constant domain or region.

[0045] As used herein, the term “immunoglobulin” refers to a protein consisting of one or more polypeptides substantially encoded by an immunoglobulin gene. This definition includes variants such as mutants, i.e., those having one or more amino acid substitutions, deletions, and insertions, N-terminal cleavage, fusion, chimeric, and humanized forms. Recognized immunoglobulin genes include, for example, various constant-region genes from primates and rodents, including humans, as well as numerous immunoglobulin variable-region genes. Monoclonal immunoglobulins are preferred. Each of the heavy and light polypeptide chains of an immunoglobulin may contain a constant region (generally a carboxyl-terminal region).

[0046] As used herein, the term “monoclonal immunoglobulin” refers to an immunoglobulin obtained from a substantially homogeneous population of immunoglobulins; that is, the individual immunoglobulins within the population are identical except for any naturally occurring variations that may be present in small amounts. Monoclonal immunoglobulins are highly specific and directed to a single antigenic site. Furthermore, in contrast to polyclonal immunoglobulin preparations, which contain different immunoglobulins for different antigenic sites (determinants or epitopes), each monoclonal immunoglobulin is directed to a single antigenic site on an antigen. In addition to specificity, monoclonal immunoglobulins have the advantage of being able to be synthesized without contamination by other immunoglobulins. The modifier “monoclonal” indicates a characteristic of immunoglobulins that they are obtained from a substantially homogeneous population of immunoglobulins and should not be interpreted as requiring the production of immunoglobulins by any particular method.

[0047] The term "codon" refers to an oligonucleotide consisting of three nucleotides that codes for a given amino acid. Due to the degeneracy of the genetic code, most amino acids are coded by two or more codons. These different codons that code for the same amino acid have different relative usage frequencies in individual host cells. Thus, a particular amino acid is coded by either exactly one codon or a different group of codons. Similarly, the amino acid sequence of a polypeptide can be coded by different nucleic acids. Therefore, a particular amino acid (residue) in a polypeptide can be coded by a different group of codons, thereby each of these codons having a usage frequency in a given host cell.

[0048] Since a large number of gene sequences are available to a large number of frequently used host cells, the relative frequency of codon use can be calculated. The calculated codon use tables are available, for example, from the "Codon Use Database" (www.kazusa.or.jp / codon / ), Nakamura, Y., et al., Nucl. Acids Res. 28(2000) 292.

[0049] The codon usage tables for Homo sapiens and hamsters were reproduced from "EMBOSS: The European Molecular Biology Open Software Suite" (Rice, P., et al., Trends Gen.16(2000)276-277, Release 6.0.1, 15.07.2009) and are shown in the table below. The codon usage frequencies of 20 native amino acids in E. coli, yeast, human cells, and CHO cells were calculated for each amino acid, rather than for all 64 codons.

[0050] (Table) Overall codon usage frequency of Homo sapiens (Code amino acid | Codon | Frequency of use [%]) TIFF0007835713000001.tif127128

[0051] (Table) Overall codon usage frequency in hamsters (Code amino acid | Codon | Frequency of use [%]) TIFF0007835713000002.tif127128

[0052] As used herein, the term “expression” refers to the transcription and / or translation process that occurs within a cell. The transcription level of a target nucleic acid sequence within a cell can be determined based on the amount of corresponding mRNA present within the cell. For example, mRNA transcribed from a target sequence can be quantified by RT-PCR (qRT-PCR) or Northern hybridization (see Sambrook, J., et al., 1989). The polypeptide encoded by the target nucleic acid can be quantified by various methods, for example, by assaying the biological activity of the polypeptide by ELISA, or by assays independent of such activity, such as Western blotting or radioimmunoassays using immunoglobulins that recognize and bind the polypeptide (see Sambrook, J., et al., 1989).

[0053] An "expression cassette" refers to a construct that includes regulatory elements necessary for expressing at least one nucleic acid present in a cell, such as a promoter and a polyadenylation site.

[0054] Gene expression is carried out either as transient or persistent expression. The target polypeptide(s) are generally secretory polypeptides and therefore contain an N-terminal extension (also known as a signal sequence) necessary for the transport / secretion of the polypeptide into the extracellular medium across the cell membrane. Generally, the signal sequence can be derived from any gene encoding a secretory polypeptide. If a heterologous signal sequence is used, it is preferably one that is recognized and processed by the host cell (i.e., cleaved by a signal peptidase). For secretion in yeast, for example, the native signal sequences of the heterologous gene to be expressed may be replaced by homologous yeast signal sequences derived from secretory genes such as yeast invertase signal sequences, alpha-factor leaders (including those from Saccharomyces, Kluyveromyces, Pichia, and Hansenula (the second one described in U.S. Patent No. 5,010,182)), acid phosphatase signal sequences, or C. albicans glucoamylase signal sequences (European Patent No. 0362179). In mammalian cell expression, the native signal sequences of the protein of interest are satisfactory, but other mammalian signal sequences, such as signal sequences from secretory polypeptides of the same or related species, e.g., human or mouse immunoglobulins, and viral secretory signal sequences, e.g., herpes simplex glycoprotein D signal sequence, may be preferred. Such a DNA fragment encoding a presegment is ligated in-frame, i.e., operably ligated, to the DNA fragment encoding the target polypeptide.

[0055] The terms “cell” or “host cell” refer to a cell that can or will be transfected with, for example, a nucleic acid encoding a heterologous polypeptide. The term “cell” includes both prokaryotic cells used for nucleic acid expression and plasmid proliferation to produce encoded polypeptides, and eukaryotic cells used for nucleic acid expression and encoding polypeptide production. In one embodiment, the eukaryotic cell is a mammalian cell. In one embodiment, the mammalian cell is a CHO cell, optionally a CHO K1 cell (ATCC CCL-61 or DSM ACC 110), or a CHO DG44 cell (also known as CHO-DHFR[-], DSM ACC 126), or a CHO XL99 cell, a CHO-T cell (see, e.g., Morgan, D., et al., Biochemistry 26(1987) 2959-2963), or a CHO-S cell, or a Super-CHO cell (Pak, SCO, et al., Cytotechnology 22(1996) 139-146). If these cells are not adapted to growth in serum-free medium or suspension, adaptation should be performed before use in this method. As used herein, the term "expressing cells" includes the target cells and their offspring. Therefore, the terms "transformed organism" and "transformed cells" include primary target cells and cultures derived therefrom, regardless of the number of transfers or passages. It should also be understood that all offspring may not be strictly identical in DNA content due to intentional or unintended mutations. This includes mutant offspring with the same function or biological activity as those screened for the original transformed cells.

[0056] The term "codon-optimized nucleic acid" refers to a nucleic acid encoding a polypeptide that has been adapted to improve expression in cells, such as mammalian cells, by replacing one, at least one, or more codons in the parent polypeptide encoding the nucleic acid with codons encoding the same amino acid residues, for example, codons with different relative frequencies of use in cells.

[0057] As used herein, the term “unpaired donor splice site” refers, on the one hand, to a donor splice site artificially generated within a nucleic acid sequence, for example, by codon optimization of the nucleic acid sequence, and on the other hand, to a donor splice site that, due to its artificial introduction into the nucleic acid sequence, lacks an acceptor splice site linked downstream of the nucleic acid sequence. The donor splice site is consistent with the consensus sequence but does not follow the (biological) splicing principle, i.e., the excision of undesirable portions of the nucleic acid during processing.

[0058] A "gene" means a nucleic acid, for example, a segment on a chromosome or plasmid, that can influence the expression of a peptide, polypeptide, or protein. In addition to the coding region, i.e., the structural gene, a gene also includes other functional elements, such as signal sequences, promoters, introns, and / or terminators.

[0059] The term "codon group" and its semantic equivalents refer to a predetermined number of different codons that code for one (i.e., the same) amino acid residue. Each codon in a group differs in their overall frequency of use within the cell's genome. Each codon within a codon group has an intrinsic frequency of use within the group, which depends on the number of codons in the group. This intrinsic frequency within a group may differ from, but is dependent on (related to), the overall frequency of use in the cell's genome. A codon group may contain only one codon, or it may contain up to six codons.

[0060] The term "overall usage frequency in the cell's genome" refers to the frequency of occurrence of a particular codon throughout the cell's entire genome.

[0061] The term "intrinsic frequency" of a codon in a codon group represents how often a single (i.e., specific) codon of a codon group can be found for all codons in that group in nucleic acids encoding polypeptides obtained by the methods reported herein. The value of intrinsic frequency depends on the overall frequency of a particular codon in the cell's genome and the number of codons in that group. Since a codon group does not necessarily contain all possible codons encoding a particular amino acid residue, the intrinsic frequency of a codon in a codon group is at least the same as, and at most 100%, its overall frequency in the cell's genome, i.e., if a particular codon with low frequency is excluded from the group, it is at least the same but may be higher than its overall frequency in the cell's genome. The sum of the specific codon frequencies of all members of a codon group is always approximately 100%.

[0062] The term "amino acid codon motif" refers to a sequence of codons that are all members of the same codon group and therefore encode the same amino acid residue. The number of different codons in an amino acid codon motif is the same as the number of different codons in a codon group, but each codon can appear more than once in the amino acid codon motif. Furthermore, each codon exists within the amino acid codon motif with its intrinsic frequency of use. Thus, an amino acid codon motif represents a sequence of different codons encoding the same amino acid residue, each of which exists with its intrinsic frequency of use, the sequence begins with the codon with the highest intrinsic frequency, and the codons are placed in a defined sequence. For example, the group of codons encoding the amino acid residue alanine contains four codons GCG, GCT, GCA, and GCC, with intrinsic frequencies of 32%, 28%, 24%, and 16%, respectively (corresponding to a ratio of 4:3:3:2). The amino acid codon motif of the amino acid residue alanine is defined by containing four codons GCG, GCT, GCA, and GCC in a ratio of 4:3:3:2, with the first codon being GCG. One exemplary amino acid codon motif of alanine is gcg gct gca gcc gcg gct gca gcc gcg gct gca gcg (SEQ ID NO: 06). This motif consists of 12 consecutive codons (4+3+3+2=12). When the amino acid residue alanine first appears in the amino acid sequence of a polypeptide, the first codon of the amino acid codon motif is used in the corresponding coding nucleic acid. When alanine appears a second time, the second codon of the amino acid codon motif is used, and so on. When alanine appears 13th in the amino acid sequence of a polypeptide, the 13th (i.e., last) codon of the amino acid codon motif is used in the corresponding coding nucleic acid. When the amino acid alanine appears at the 13th position in the polypeptide amino acid sequence, the first codon of the amino acid codon motif is used, and so on.

[0063] In this application, the terms “nucleic acid” or “nucleic acid sequence” as used interchangeably refer to polymer molecules consisting of individual nucleotides (also called bases) a, c, g, and t (or u in RNA), such as DNA, RNA, or modifications and mixtures thereof. This polynucleotide molecule may be a naturally occurring polynucleotide molecule, a synthetic polynucleotide molecule, or a combination of one or more naturally occurring polynucleotide molecules and one or more synthetic polynucleotide molecules. Naturally occurring polynucleotide molecules in which one or more nucleotides have been altered (e.g., by mutagenesis), deleted, or added are also included in this definition. Nucleic acids can be isolated or integrated into another nucleic acid, such as an expression cassette, plasmid, or host cell chromosome. A nucleic acid is characterized by its nucleic acid sequence consisting of individual nucleotides.

[0064] For those skilled in the art, procedures and methods for converting, for example, the amino acid sequence of a polypeptide to a corresponding nucleic acid sequence encoding that amino acid sequence are well known. Thus, a nucleic acid is characterized by its nucleic acid sequence, which consists of individual nucleotides, and similarly by the amino acid sequence of the polypeptide encoded by that nucleic acid sequence.

[0065] A "structural gene" refers to a region of a gene that lacks a signal sequence, i.e., the coding region.

[0066] A "transfection vector" is a nucleic acid (also called a nucleic acid molecule) that provides all the elements necessary for the expression of a coding nucleic acid / structural gene(s) in a host cell. A transfection vector includes, for example, a prokaryotic plasmid growth unit for E. coli, which sequentially includes a prokaryotic origin of replication and a nucleic acid conferring resistance to a prokaryotic selective substance, and further includes one or more nucleic acids conferring resistance to a eukaryotic selective substance and one or more nucleic acids encoding the target polypeptide. Preferably, the nucleic acid conferring resistance to the selective substance and the nucleic acid encoding the target polypeptide are each placed in an expression cassette, each expression cassette containing a promoter, a coding nucleic acid, and a transcription terminator including a polyadenylation signal. Gene expression is typically under the control of a promoter, and such structural genes are said to be "operably linked" to the promoter. Similarly, if a regulatory element modulates the activity of a core promoter, the regulatory element and the core promoter are operably linked.

[0067] As used herein, the term “vector” refers to a nucleic acid molecule capable of replicating another nucleic acid it is linked to. This term includes vectors as self-replicating nucleic acid structures, as well as vectors integrated into the genome of a host cell into which they are introduced. Certain types of vectors can direct the expression of a functionally linked nucleic acid. Such vectors are referred to herein as “expression vectors.”

[0068] The term "full-length antibody" refers to an antibody that has a structure substantially similar to that of a natural antibody. A full-length antibody comprises two full-length antibody light chains, each containing a light chain variable region and a light chain constant domain from the N-terminus to the C-terminus, and two full-length antibody heavy chains, each containing a heavy chain variable region, a first heavy chain constant domain, a hinge region, a second heavy chain constant domain, and a third heavy chain constant domain from the N-terminus to the C-terminus. In contrast to natural antibodies, full-length antibodies may contain additional immunoglobulin domains, such as one or more additional scFvs, or heavy or light chain Fab fragments, or scFab conjugated to one or more ends of different chains of the full-length antibody, but with only a single fragment at each end. These conjugates are also encompassed by the term full-length antibody.

[0069] The "class" of an antibody refers to the type of constant domain or constant region, preferably the Fc region of the heavy chain. There are five main classes of antibodies: IgA, IgD, IgE, IgG, and IgM, some of which may be further divided into subclasses (isotypes), such as IgG1, IgG2, IgG3, IgG4, IgA1, and IgA2. The heavy chain constant domains corresponding to different classes of immunoglobulins are called α, δ, ε, γ, and μ, respectively.

[0070] The term "heavy chain constant region" refers to a region of the immunoglobulin heavy chain that includes the constant domain, i.e., the CH1 domain, hinge region, CH2 domain, and CH3 domain. In one embodiment, the human IgG constant region extends from Ala118 to the carboxyl terminus of the heavy chain (numbering follows the Kabat EU index). However, the C-terminal lysine (Lys447) of the constant region may or may not be present (numbering follows the Kabat EU index). The term "heavy chain constant region" refers to a dimer containing two heavy chain constant regions that can be covalently bonded to each other via hinge region cysteine ​​residues that form an interchain disulfide bond.

[0071] The term "light chain constant region" refers to the region of the immunoglobulin light chain that contains the constant domain, i.e., the CL domain.

[0072] The term "steady-state region" encompasses both the "heavy chain steady-state region" and the "light chain steady-state region."

[0073] The term "heavy chain Fc region" refers to the C-terminal region of an immunoglobulin heavy chain, including at least a portion of the hinge region (intermediate and lower hinge regions), the CH2 domain, and the CH3 domain. In one embodiment, the human IgG heavy chain Fc region extends from Asp221 or Cys226 or Pro230 to the carboxyl terminus of the heavy chain (numbering follows the Kabat EU index). Thus, the Fc region is smaller than a certain region but is located in the same C-terminal region. However, the C-terminal lysine (Lys447) of the heavy chain Fc region may or may not be present (numbering follows the Kabat EU index). The term "Fc region" refers to a dimer containing two heavy chain Fc regions that can be covalently bonded to each other via hinge region cysteine ​​residues that form an interchain disulfide bond.

[0074] The constant region of an antibody, more precisely the Fc region (and the constant region as well), is directly involved in complement activation, C1q binding, C3 activation, and Fc receptor binding. While the antibody's effect on the complement system is condition-dependent, binding to C1q is triggered by a specific binding site in the Fc region. Such binding sites are known from prior art, for example: Lukas, TJ, et al., J.Immunol. 127 (1981) 2555-2560; Brunhouse, R., and Cebra, JJ, Mol.Immunol. 16 (1979) 907-917; Burton, DR, et al., Nature 288 (1980) 338-344; Thommesen, JE, et al., Mol.Immunol. 37 (2000) 995-1004; Idusogie, EE, et al., J.Immunol. 164 (2000) 4178-4184; Hezareh, M., et al., J.Virol. 75 (2001) 12161-12168; Morgan, A., et al., Immunology This is described in 86(1995)319-324 and European Patent No. 0307434. Such binding sites are, for example, L234, L235, D270, N297, E318, K320, K322, P331 and P329 (numbered according to Kabat's EU index). Antibodies of subclasses IgG1, IgG2, and IgG3 typically exhibit complement activation, C1q binding, and C3 activation, whereas IgG4 does not activate the complement system, does not bind to C1q, and does not activate C3. The term "Fc region of an antibody" is well known to those skilled in the art and is defined based on papain cleavage of the antibody.

[0075] Splicing The constant region amino acid sequences of different human immunoglobulins are encoded by corresponding DNA sequences. In the genome, these DNA sequences contain coding (exon) sequences and non-coding (intron) sequences. After transcription of DNA into pre-mRNA, the pre-mRNA also contains these intron and exon sequences. Before translation, non-coding intron sequences are removed during mRNA processing by splicing them from the primary mRNA transcript to produce mature mRNA. Splicing of primary mRNA is controlled by donor splice sites in combination with appropriately spaced acceptor splice sites. The donor splice site is located at the 5' end, and the acceptor splice site is located at the 3' end of the intron sequence.

[0076] The term "properly spaced" means that the donor splice site and acceptor splice site in the nucleic acid are positioned in the appropriate location so that all the elements necessary for the splicing process are available and the splicing process can take place.

[0077] The donor splice site (5' splice site) is a nucleic acid sequence motif that represents the 5' end of an intron.

[0078] The acceptor splice site (3' splice site) is a nucleic acid sequence motif that represents the 3' end of an intron.

[0079] Codon optimization: Recombinant antibody production is based at least on a constant region of the sequence of the naturally occurring wild-type or germline. To overcome the biological limitations of these DNA templates, such as RNA instability, inefficient nuclear export, secondary structure, or insufficient translation rate, gene optimization and subsequent de novo synthesis of genes based on protein sequences are performed using bioinformatics techniques

[10] . Thus, the complex and time-consuming cloning process can be avoided, and translation rates can be increased by such adjustment of codon use to the production system, as has been shown in recent years

[10] . For codon optimization, high GC content, avoidance of splice sites, and adaptation of codon use to the producing organism play a central role. Several methods have already been established to investigate the use and frequency of specific codons. tRNAs encoding the same amino acid are known to compete with each other. On the one hand, polyvalent tRNAs that recognize multiple codons are more commonly available, and on the other hand, different tRNA species are expressed at varying frequencies. As a result, codons encoded by more abundant tRNAs also occur more frequently in the coding sequence

[11] . Table 1 shows three possible methods for applying codon use adaptation / optimization.

[0080] (Table 1) TIFF0007835713000003.tif50139

[0081] One method of codon use aims to make the entire tRNA pool available for translation, i.e., Method 2. In contrast to the “high” method, which uses only the codons most abundant for translation into amino acids, all available codons are used. Method 2 also takes into account the distribution of codons used in each organism's codon use compared to Method 1, and distributes codons within the sequence

[11] .

[0082] Splice location: Almost all eukaryotic protein-coding genes are divided into exons and introns. Currently, 12 exon variants are known

[12] . Splicing is the removal of introns from pre-mRNA during transcription. Correct splicing is based on conserved consensus sequences at the 5' and 3' ends of introns and so-called branching sites, respectively. Branching sites are approximately 20–50 nucleotides upstream of the 3' end of introns

[12] . At the 5' end of introns is a donor splice site with a characteristic dinucleotide GT. At the 3' end is an acceptor splice site located with the base AG

[13] . This pattern is called the canonical pattern of splice sites. Nevertheless, 3.7% of annotated splice sites do not follow this pattern

[14] . Splice sites with GC-AG, GG-AG, GT-TG, GT-CG, or CT-AG dinucleotides within introns have also been observed

[14] . Some of these non-canonical splice sites may be involved in immunoglobulin expression

[15] . According to Burset et al., dinucleotide GC-AG is the most common non-standard pattern of splice sites. Furthermore, the consensus sequence is strongly dependent on the GC content

[12] . When the GC content is high, the 5'-terminus donor consensus sequence is described as AG / GTRAGT (SEQ ID NO: 02) rather than AG / GTAAGT (SEQ ID NO: 03) when the GC content is low

[12] .

[0083] Introns cleave spliceosomes. Those introns adjacent to standard GT-AG pairs are released from premRNA by spliceosomes having subunits U1, U2, U4 / U6, and U5

[16] . Furthermore, eukaryotic genomes are known to have various hidden splice sites that negatively affect proper splicing. These are distinct from the consensus motif of splice sites. Typically, these sites are inactive or hardly used by cellular mechanisms

[17] ,

[18] . Hidden splice sites exist in both introns and exons. The goal of bioinformatics programs is to recognize the location of these potential splice sites. Many of these programs are useful, but not as useful as the complexity of nucleotide information due to the many possibilities

[22] .

[0084] Specific Embodiments of the Method According to the Present Invention The goal is to achieve high expression yields for the production of therapeutic or diagnostic antibodies. One option for achieving good expression rates is to first optimize the codon usage of the coding nucleic acid and then adapt it to the codon usage of cells intended to express the exogenous nucleic acid. This codon matching or optimization can be performed based on various established protocols.

[0085] However, during such codon matching and optimization processes, for example, unpaired splice sites can be generated de novo. That is, during codon optimization, a new donor splice site sequence is generated in the codon-use optimized nucleic acid. This is independent of the organization of the codon-use optimized nucleic acid, i.e., it is possible for both cDNA and genome-organized nucleic acids. In fact, this is an unintended by-effect of the codon-use optimization process. Such a new donor splice site is an additional artificial donor splice site and does not have an associated target acceptor splice site. Therefore, such an unpaired donor splice site can result in a splicing event with a random, i.e., undefined acceptor splice site located somewhere in the transcribed mRNA. This leads to a decrease in expression yield.

[0086] The present invention is at least in part based on the unexpected finding that the removal of unpaired donor splice sites in the nucleic acid encoding the heavy chain of a codon-use optimized antibody needs to be performed only in the portion of the nucleic acid encoding the variable region of the heavy chain, and not in the portion encoding the constant region; that is, in the case of the constant region, for example, germline or wild-type human nucleic acid sequences can be used.

[0087] The present invention is at least in part based on the finding that, in codon-optimized nucleic acids encoding the variable domain of an antibody heavy chain, the introduction of silent nucleotide changes (mutations) in the unpaired donor splice site consensus sequence NGGTA(G)AG (SEQ ID NO: 01) is entirely sufficient to improve the expression yield or to express the antibody.

[0088] The present invention is at least in part based on the finding that removal of unpaired donor splice sites needs to be performed only in nucleic acids encoding the variable domain of the heavy chain, and not in the constant region. That is, for the constant region, either a wild-type sequence, a germline sequence, or a sequence optimized by a standard method such as reported in International Publication No. 2013 / 15644 can be used.

[0089] This invention is at least in part based on the finding that expression yield can be increased by using the method according to the present invention, which utilizes light chains.

[0090] This invention is based on the use of the donor consensus sequence SEQ ID NO: 01: NGGTA(G)AG. Although this sequence was previously identified by Zhang et al. 1998, no entry into a codon optimization protocol had been found.

[0091] The dinucleotide GT indicates the initiation of an intron and the resulting splice site. The following base may be adenine or guanine.

[0092] Another parameter that may be considered is the number of mismatches allowed in this array to adjust the sensitivity and stringency of the method.

[0093] The consensus acceptor splice sequence has the sequence SEQ ID NO: 04:[CT]n N[CT]AG, where N is any base and n is the number of CT dinucleotides. This sequence has also been identified by Zhang 1998.

[0094] Methods according to the present invention are illustrated below using specific antibodies. These examples should not be understood as limitations of the present invention. They are presented merely as examples of generally applicable methods according to the present invention.

[0095] Table 2 below summarizes the expression yields of antibody heavy chain and / or light chain coding nucleic acids treated with different methods. The results were obtained by transient expression in HEK293 cells using a two-plasmid system.

[0096] The construct "00" is the starting nucleic acid.

[0097] Constructs 00'~06' and 16' are, The heavy chain variable domain, Codon optimization is performed using methods known in the art (=referenced prior art methods) as described in International Publication No. 2013 / 156443.

[0098] Constructs 07'~11' and 17' are processed according to the prior art method referenced in International Publication No. 2013 / 156443, which supplements the removal of the donor splice portion based on the consensus motif of Sequence ID No. 01.

[0099] The construct, 12'~15', has been codon-optimized by the commercial provider Geneart using a different approach than that of the second reference prior art method, International Publication 2013 / 156443.

[0100] The term "hu" means that the homocodon codon was used, and the term "CHO" means that the Chinese hamster codon was used.

[0101] (Table 2) TIFF0007835713000004.tif200170

[0102] The data reveals the following: In the case of a light chain: -When using genome organization, the processing according to the present invention yielded equivalent expression yields. TIFF0007835713000005.tif21170

[0103] - The use of cDNA generally results in increased expression yield. TIFF0007835713000006.tif33170

[0104] -When using cDNA, the processing according to the present invention results in an increase in expression yield in most cases, while in other cases the expression yield remains equivalent. TIFF0007835713000007.tif96170

[0105] In the case of heavy chains: -When genome engineering is used, the processing according to the present invention results in an increase in expression yield or increased expression, provided that at least the heavy chain variable domains are optimized. TIFF0007835713000008.tif50170

[0106] - The use of cDNA generally results in increased expression yield. TIFF0007835713000009.tif115170

[0107] -When using cDNA, the processing according to the present invention for the constant region of the heavy chain does not appear to further increase the expression yield.

[0108] -When using cDNA, the processing according to the present invention results in a further increase in expression yield. TIFF0007835713000010.tif21170

[0109] Therefore, one embodiment described herein is a method for producing immunoglobulin: - A step of culturing mammalian cells, preferably CHO cells, containing nucleic acids having an intron-exon structure encoding the heavy chain and light chain of human IgG1 subclass immunoglobulins, so that immunoglobulins are expressed. - The process includes a step of recovering immunoglobulins from cells or culture medium and thereby producing immunoglobulins, In this method, the unpaired donor splice site indicated by SEQ ID NO: 01 in the nucleic acid portion encoding the immunoglobulin heavy chain variable domain is removed by introducing a silent nucleotide change in the amino acid sequence of the unpaired donor splice site consensus sequence NGGTA(G)AG (SEQ ID NO: 01) into the codon NGG, codon GGT, or codon GTA(G).

[0110] Reference codon optimization method Nucleic acids encoding immunoglobulins can be optimized, for example, by adapting common codon usage according to the method reported in International Publication No. 2013 / 156443.

[0111] The above reference method is based on the finding that, for polypeptide expression in cells, each amino acid is encoded by a group of codons, and thereafter, each codon in the group is defined by its intrinsic frequency within the group, which is related to the overall frequency of use of that codon in the cell's genome, and thereafter, the frequency of use of codons in the (whole) polypeptide encoding the nucleic acid is approximately the same as the frequency of use within each group.

[0112] The reference method is a method for recombinantly producing polypeptides in mammalian cells, comprising the steps of culturing cells containing nucleic acids encoding polypeptides and recovering polypeptides from mammalian cells or culture medium, Each amino acid residue in a polypeptide is encoded by one or more (at least one) codons, so that (different) codons encoding the same amino acid residue are grouped together, and each codon in a group is defined by its intrinsic frequency of use within that group, which is the frequency with respect to all codons in a group that a single codon in that group can be found in the nucleic acid encoding the polypeptide, so that the sum of the intrinsic frequencies of all codons in a group is 100%. This method ensures that the overall frequency of use of each codon in a polypeptide-coding nucleic acid is approximately the same as its intrinsic frequency within that group.

[0113] In one embodiment, amino acid residues G, A, V, L, I, P, F, S, T, N, Q, Y, C, K, R, H, D, and E are each encoded by a group of codons, while amino acid residues M and W are encoded by a single codon.

[0114] In one embodiment, amino acid residues G, A, V, L, I, P, F, S, T, N, Q, Y, C, K, R, H, D, and E are each encoded by a group of codons containing at least two codons, while amino acid residues M and W are encoded by a single codon.

[0115] In one embodiment, if an amino acid residue is encoded by exactly one codon, the intrinsic frequency of use of the codon is 100%.

[0116] In one embodiment, amino acid residue G is encoded by a group of up to four codons. In one embodiment, amino acid residue A is encoded by a group of up to four codons. In one embodiment, amino acid residue V is encoded by a group of up to four codons. In one embodiment, amino acid residue L is encoded by a group of up to six codons. In one embodiment, amino acid residue I is encoded by a group of up to three codons. In one embodiment, amino acid residue M is encoded by exactly one codon. In one embodiment, amino acid residue P is encoded by a group of up to four codons. In one embodiment, amino acid residue F is encoded by a group of up to two codons. In one embodiment, amino acid residue W is encoded by exactly one codon. In one embodiment, amino acid residue S is encoded by a group of up to six codons. In one embodiment, amino acid residue T is encoded by a group of up to four codons. In one embodiment, amino acid residue N is encoded by a group of up to two codons. In one embodiment, amino acid residue Q is encoded by a group of up to two codons. In one embodiment, amino acid residue Y is encoded by a group of up to two codons. In one embodiment, amino acid residue C is encoded by a group of up to two codons. In one embodiment, amino acid residue K is encoded by a group of up to two codons. In one embodiment, amino acid residue R is encoded by a group of up to six codons. In one embodiment, amino acid residue H is encoded by a group of up to two codons. In one embodiment, amino acid residue D is encoded by a group of up to two codons. In one embodiment, amino acid residue E is encoded by a group of up to two codons.

[0117] In one embodiment, amino acid residue G is encoded by a group of 1 to 4 codons. In one embodiment, amino acid residue A is encoded by a group of 1 to 4 codons. In one embodiment, amino acid residue V is encoded by a group of 1 to 4 codons. In one embodiment, amino acid residue L is encoded by a group of 1 to 6 codons. In one embodiment, amino acid residue I is encoded by a group of 1 to 3 codons. In one embodiment, amino acid residue M is encoded by a group of 1 codon, i.e., exactly 1 codon. In one embodiment, amino acid residue P is encoded by a group of 1 to 4 codons. In one embodiment, amino acid residue F is encoded by a group of 1 to 2 codons. In one embodiment, amino acid residue W is encoded by a group of 1 codon, i.e., exactly 1 codon. In one embodiment, amino acid residue S is encoded by a group of 1 to 6 codons. In one embodiment, amino acid residue T is encoded by a group of 1 to 4 codons. In one embodiment, amino acid residue N is encoded by a group of 1 to 2 codons. In one embodiment, amino acid residue Q is encoded by a group of 1 to 2 codons. In one embodiment, amino acid residue Y is encoded by a group of 1 to 2 codons. In one embodiment, amino acid residue C is encoded by a group of 1 to 2 codons. In one embodiment, amino acid residue K is encoded by a group of 1 to 2 codons. In one embodiment, amino acid residue R is encoded by a group of 1 to 6 codons. In one embodiment, amino acid residue H is encoded by a group of 1 to 2 codons. In one embodiment, amino acid residue D is encoded by a group of 1 to 2 codons. In one embodiment, amino acid residue E is encoded by a group of 1 to 2 codons.

[0118] In one embodiment, each group comprises only codons that have an overall usage frequency of 5% or more of the cell genome. In one embodiment, each group comprises only codons that have an overall usage frequency of 8% or more of the cell genome. In one embodiment, each group comprises only codons that have an overall usage frequency of 10% or more of the cell genome. In one embodiment, each group comprises only codons that have an overall usage frequency of 15% or more of the cell genome.

[0119] In one embodiment, the sequence of codons in a nucleic acid encoding a polypeptide for a specific amino acid residue in the 5' to 3' direction is the sequence of codons in each amino acid codon motif, i.e., it corresponds to that sequence.

[0120] In one embodiment, for each consecutive occurrence of a specific amino acid in the polypeptide, starting from the N-terminus, the coding nucleic acid includes a codon that is the same as the codon at the corresponding consecutive position in the amino acid codon motif of each specific amino acid, where the first occurrence of an amino acid residue in the polypeptide's amino acid sequence uses the first codon of the amino acid codon motif for the corresponding coding nucleic acid, the second occurrence of an amino acid residue uses the second codon of the amino acid codon motif for the second occurrence, and so on.

[0121] In one embodiment, the frequency of use of a codon in an amino acid codon motif is approximately the same as its intrinsic frequency within that group.

[0122] In one embodiment, after reaching the final codon of the amino acid codon motif upon the next occurrence of a particular amino acid in the polypeptide, the coding nucleic acid includes the codon at the first position of the amino acid codon motif.

[0123] In one embodiment, the codons in the amino acid codon motif are randomly distributed throughout the entire amino acid codon motif.

[0124] In one embodiment, an amino acid codon motif is selected from a group of amino acid codon motifs that include all possible amino acid codon motifs obtained by rearranging the codons within it, all of which motifs have the same number of codons and each codon in each motif has the same unique frequency of use.

[0125] In one embodiment, codons in an amino acid codon motif are arranged such that their intrinsic usage frequency decreases, thereby ensuring that all codons of a single usage frequency are directly consecutive to each other. In one embodiment, codons of a single usage frequency are grouped together.

[0126] In one embodiment, the (different) codons within the amino acid codon motif are uniformly distributed throughout the entire amino acid codon motif.

[0127] In one embodiment, the codons in the amino acid codon motif are arranged such that their intrinsic usage frequency decreases, so that the codon with the highest intrinsic usage frequency is located after the codon with the lowest intrinsic usage frequency or the codon with the second lowest intrinsic usage frequency.

[0128] In one embodiment, the codons in the amino acid codon motif are arranged such that the intrinsic frequency of use decreases, with the codon with the lowest relative frequency followed by the codon with the highest relative frequency (used).

[0129] Therefore, nucleic acids encoding polypeptides are characterized in that each amino acid residue of the polypeptide is encoded by one or more (at least one) codons. Thus, different codons encoding the same amino acid residue are combined into one group, and each codon within a group is defined by its intrinsic frequency of use within that group, which is the frequency with respect to all codons in a group of codons that a single codon in that group can be found in the nucleic acid encoding a polypeptide, and so the sum of the intrinsic frequencies of all codons within a group is 100%. The frequency of use of codons in nucleic acids encoding polypeptides is approximately the same as their intrinsic frequency within that group.

[0130] In one embodiment, amino acid residues G, A, V, L, I, P, F, S, T, N, Q, Y, C, K, R, H, D, and E are each encoded by a group of codons, while amino acid residues M and W are encoded by a single codon.

[0131] In one embodiment, amino acid residues G, A, V, L, I, P, F, S, T, N, Q, Y, C, K, R, H, D, and E are each encoded by a group of codons containing at least two codons, while amino acid residues M and W are encoded by a single codon.

[0132] In one embodiment, if an amino acid residue is encoded by exactly one codon, the intrinsic frequency of use of the codon is 100%.

[0133] In one embodiment, amino acid residue G is encoded by a group of up to four codons. In one embodiment, amino acid residue A is encoded by a group of up to four codons. In one embodiment, amino acid residue V is encoded by a group of up to four codons. In one embodiment, amino acid residue L is encoded by a group of up to six codons. In one embodiment, amino acid residue I is encoded by a group of up to three codons. In one embodiment, amino acid residue M is encoded by exactly one codon. In one embodiment, amino acid residue P is encoded by a group of up to four codons. In one embodiment, amino acid residue F is encoded by a group of up to two codons. In one embodiment, amino acid residue W is encoded by exactly one codon. In one embodiment, amino acid residue S is encoded by a group of up to six codons. In one embodiment, amino acid residue T is encoded by a group of up to four codons. In one embodiment, amino acid residue N is encoded by a group of up to two codons. In one embodiment, amino acid residue Q is encoded by a group of up to two codons. In one embodiment, amino acid residue Y is encoded by a group of up to two codons. In one embodiment, amino acid residue C is encoded by a group of up to two codons. In one embodiment, amino acid residue K is encoded by a group of up to two codons. In one embodiment, amino acid residue R is encoded by a group of up to six codons. In one embodiment, amino acid residue H is encoded by a group of up to two codons. In one embodiment, amino acid residue D is encoded by a group of up to two codons. In one embodiment, amino acid residue E is encoded by a group of up to two codons.

[0134] In one embodiment, amino acid residue G is encoded by a group of 1 to 4 codons. In one embodiment, amino acid residue A is encoded by a group of 1 to 4 codons. In one embodiment, amino acid residue V is encoded by a group of 1 to 4 codons. In one embodiment, amino acid residue L is encoded by a group of 1 to 6 codons. In one embodiment, amino acid residue I is encoded by a group of 1 to 3 codons. In one embodiment, amino acid residue M is encoded by a group of 1 codon. In one embodiment, amino acid residue P is encoded by a group of 1 to 4 codons. In one embodiment, amino acid residue F is encoded by a group of 1 to 2 codons. In one embodiment, amino acid residue W is encoded by a group of 1 codon. In one embodiment, amino acid residue S is encoded by a group of 1 to 6 codons. In one embodiment, amino acid residue T is encoded by a group of 1 to 4 codons. In one embodiment, amino acid residue N is encoded by a group of 1 to 2 codons. In one embodiment, amino acid residue Q is encoded by a group of 1 to 2 codons. In one embodiment, amino acid residue Y is encoded by a group of 1 to 2 codons. In one embodiment, amino acid residue C is encoded by a group of 1 to 2 codons. In one embodiment, amino acid residue K is encoded by a group of 1 to 2 codons. In one embodiment, amino acid residue R is encoded by a group of 1 to 6 codons. In one embodiment, amino acid residue H is encoded by a group of 1 to 2 codons. In one embodiment, amino acid residue D is encoded by a group of 1 to 2 codons. In one embodiment, amino acid residue E is encoded by a group of 1 to 2 codons.

[0135] In one embodiment, each group comprises only codons that have an overall usage frequency of 5% or more of the cell genome. In one embodiment, each group comprises only codons that have an overall usage frequency of 8% or more of the cell genome. In one embodiment, each group comprises only codons that have an overall usage frequency of 10% or more of the cell genome. In one embodiment, each group comprises only codons that have an overall usage frequency of 15% or more of the cell genome.

[0136] In one embodiment, the sequence of codons in a nucleic acid encoding a polypeptide for a specific amino acid residue in the 5' to 3' direction is the sequence of codons in each amino acid codon motif, i.e., it corresponds to that sequence.

[0137] In one embodiment, for each consecutive occurrence of a specific amino acid in the polypeptide, starting from the N-terminus, the coding nucleic acid includes a codon that is the same as the codon at the corresponding consecutive position in the amino acid codon motif of each specific amino acid, where the first occurrence of an amino acid residue in the polypeptide's amino acid sequence uses the first codon of the amino acid codon motif for the corresponding coding nucleic acid, the second occurrence of an amino acid residue uses the second codon of the amino acid codon motif for the second occurrence, and so on.

[0138] In one embodiment, the frequency of use of a codon in an amino acid codon motif is approximately the same as its intrinsic frequency within that group.

[0139] In one embodiment, after reaching the final codon of the amino acid codon motif upon the next occurrence of a particular amino acid in the polypeptide, the coding nucleic acid includes the codon at the first position of the amino acid codon motif.

[0140] In one embodiment, each codon in the amino acid codon motif is randomly distributed throughout the entire amino acid codon motif.

[0141] In one embodiment, each codon in the amino acid codon motif is uniformly distributed throughout the entire amino acid codon motif.

[0142] In one embodiment, the codons in the amino acid codon motif are arranged such that their intrinsic frequency decreases, thereby the codon with the highest intrinsic frequency is used after the codon with the lowest intrinsic frequency or the codon with the second lowest intrinsic frequency.

[0143] In one embodiment, the codons in the amino acid codon motif are arranged so that their intrinsic frequency decreases, with the codon with the highest relative frequency being used after the codon with the lowest relative frequency.

[0144] Therefore, the method according to the present invention is a method for increasing polypeptide expression in eukaryotic cells, - A step of providing nucleic acids that encode a polypeptide, Each amino acid residue in a polypeptide is encoded by at least one codon, and thereafter different codons encoding the same amino acid residue are combined into one group, and each codon in a group is defined by its intrinsic frequency of use within the group, and thereafter the sum of the intrinsic frequencies of all codons in a group is 100%, The frequency of use of codons in nucleic acids encoding polypeptides is approximately the same as their intrinsic frequency within that group. This method involves removing the donor splice site, based on the consensus sequence of Sequence ID No. 01, from the nucleic acid encoding the heavy chain variable domain.

[0145] Recombination method Antibodies can be produced using recombinant methods and compositions, for example, as described in U.S. Patent No. 4,816,567. Antibody-coding nucleic acids may encode amino acid sequences containing the VL of the antibody and / or amino acid sequences containing the VH of the antibody (e.g., the light chain and / or heavy chain of the antibody). In one embodiment, cells expressing an immunoglobulin constant region-containing polypeptide are transfected with one or more vectors (e.g., expression vectors) containing such nucleic acids. In one embodiment, cells containing such nucleic acids modified by the method described herein are provided. In one embodiment, the cells include (e.g., transformed using): (1) a vector containing a nucleic acid encoding an amino acid sequence containing the VL of the antibody and a nucleic acid encoding an amino acid sequence containing the VH of the antibody, or (2) a vector containing a first vector containing a nucleic acid encoding an amino acid sequence containing the VL of the antibody and a second vector containing a nucleic acid encoding an amino acid sequence containing the VH of the antibody. In one embodiment, the cells are eukaryotic cells, for example, Chinese hamster ovary (CHO) cells or lymphocytes (e.g., Y0, NS0, Sp2 / 0 cells). In one embodiment, a method for producing an antibody is provided, which includes culturing cells containing nucleic acids encoding an antibody provided herein under conditions suitable for antibody expression, and optionally recovering the antibody from the cells (or cell culture medium).

[0146] For recombinant antibody production, for example, nucleic acids encoding the antibodies reported herein are generated and inserted into one or more vectors for further cloning and / or expression in cells. Such nucleic acids can be readily isolated and sequenced using conventional procedures (e.g., by using oligonucleotide probes that can specifically bind to the genes encoding the heavy and light chains of the antibody).

[0147] Suitable cells for cloning or expressing antibody-encoding vectors include eukaryotic cells as described herein.

[0148] Furthermore, eukaryotic microorganisms such as filamentous fungi or yeasts are suitable cloning or expression hosts for antibody-encoding vectors, and their glycosylation pathways are "humanized," resulting in the production of antibodies with a partially or completely human glycosylation pattern, including fungal and yeast strains (see Gerngross, TU, Nat. Biotech. 22(2004) 1409-1414; and Li, H., et al., Nat. Biotech. 24(2006) 210-215).

[0149] Furthermore, suitable host cells for expressing glycosylated antibodies are derived from multicellular organisms (invertebrates and vertebrates). Examples of invertebrate cells include plant cells and insect cells. Numerous baculovirus strains have been identified that can be used in connection with the transfection of insect cells, particularly Spodoptera frugiperda cells.

[0150] Plant cell cultures can also be used as hosts (see, for example, U.S. Patents 5,959,177, 6,040,498, 6,420,548, 7,125,978, and 6,417,429, which describe PLANTIBODIES® technology for antibody production in transgenic plants).

[0151] Vertebrate cells can also be used as hosts. For example, mammalian cell lines adapted to grow in suspensions may be useful. Other examples of useful mammalian cell lines include the CV1 monkey kidney cell line transformed with SV40 (COS-7), human embryonic kidney cells (e.g., HEK293 cells as described in Graham, F. Let al., J. Gen Virol. 36 (1977) 59-74, baby hamster kidney cells (BHK), mouse Sertoli cells (e.g., TM4 cells as described in Mather, JP, Biol. Reprod. 23 (1980) 243-252), monkey kidney cells (CV1), African green monkey kidney cells (VERO-76), human cervical tumor cells (HELA), canine kidney cells (MDCK), buffalo rat liver cells (BRL 3A), human lung cells (W138), human liver cells (Hep G2), mouse mammary tumor cells (MMT 060562), TRI cells (e.g., Mather, JP et al., Annals) These include MRC5 cells and FS4 cells, as described in NYAcad.Sci.383(1982)44-68. Other useful mammalian cell lines include DHFR - Examples include Chinese hamster ovary (CHO) cells (Urlaub, G., et al., Proc. Natl. Acad. Sci. USA 77(1980) 4216-4220), including CHO cells, and myeloma cell lines such as Y0, NS0, and Sp2 / 0. For a review of specific mammalian host cells suitable for antibody production, see, for example, Yazaki, P. and Wu, AM, Methods in Molecular Biology, Vol. 248, Lo, BKC (ed.), Humana Press, Totowa, NJ (2004) pp. 255-268.

[0152] purification Various methods have been well established and widely used for the recovery and purification of proteins, including affinity chromatography using microbial proteins (e.g., protein A or protein G affinity chromatography), ion exchange chromatography (e.g., cation exchange (carboxymethyl resin), anion exchange (aminoethyl resin), and mixed-mode exchange), thiophene adsorption (e.g., using β-mercaptoethanol and other SH ligands), hydrophobic interaction or aromatic adsorption chromatography (e.g., using phenyl-sepharose, aza-aromatic resin, or m-aminophenylboronic acid), metal chelate affinity chromatography (e.g., using Ni(II) affinity materials and Cu(II) affinity materials), size exclusion chromatography, and electrophoresis (e.g., gel electrophoresis, capillary electrophoresis) (Vijayalakshmi, MA, Appl. Biochem. Biotech. 75(1998) 93-102).

[0153] Codon usage Codon usage tables (see the table above for an example) are readily available, for example, in the "Codon Usage Database" at http: / / www.kazusa.or.jp / codon / , and these tables can be adapted in several ways (Nakamura, Y., et al., Nucl. Acids Res. 28(2000) 292).

[0154] For the high-yield expression of recombinant polypeptides, coding nucleic acids play a crucial role. Naturally occurring and naturally isolated coding nucleic acids are generally not optimized for high-yield expression, especially when expressed in heterologous host cells. Due to genetic code denaturation, a single amino acid residue can be coded by two or more nucleotide triplets (codons), with the exception of the amino acids tryptophan and methionine. Therefore, for a single amino acid sequence, different coding codons (=corresponding coding nucleic acid sequences) are possible.

[0155] Different codons encoding a single amino acid residue are used by different organisms at different relative frequencies (codon usage). Generally, one particular codon is used more frequently than other possible codons.

[0156] International Publication No. 2001 / 088141 reports on the optimization of reading frames using codons found in highly expressed mammalian genes. For this purpose, a matrix was generated considering almost exclusively the most frequently used codons and the less desirable but second most frequently used codons in highly expressed mammalian genes, as shown in the table below. Using these codons derived from highly expressed human genes, a completely synthetic reading frame, not found in nature, was constructed, which encodes the exact same product as the original wild-type gene construct.

[0157] U.S. Patent No. 8,128,938 reports various codon optimization methods that use the usage frequency of individual codons, including uniform optimization, complete optimization, and minimal optimization.

[0158] The table below shows the most frequently used codon (codon 1) and the second most frequently used codon (codon 2) in highly expressed mammalian genes.

[0159] (table.) TIFF0007835713000011.tif142128 (Ausubel, FM, et al., Current Protocols in Molecular Biology 2(1994), A1.8-A1.9).

[0160] To enable continuous PCR amplification and sequencing of synthetic gene products, there is little deviation from strict adherence to the use of the most frequently found codons for (i) introducing or removing specific restriction sites, and (ii) cleaving G or C extensions exceeding 7 base pairs.

[0161] The following embodiments are provided to aid in understanding the present invention, and its true scope is specified in the claims. It is understood that modifications to the prescribed procedures may be made without departing from the spirit of the invention.

[0162] literature

[10] M. Graf, L. Deml, and R. Wagner, ``Codon-optimized genes that enable increased heterologous expression in mammalian cells and elicit immune responses in mice after vaccination of efficient naked DNA.,'' Methods Mol.Med., vol. 94, pp. 197-210, 2004.

[11] International Publication No. 2013 / 156443

[12] MQZhang,''Statistical features of human exons and their flanking regions,'' vol.7, no.5, pp.919-932, 1998.

[13] R. Breathnach, C. Benoist, K. O'Hare, F. Gannon, and P. Chambon, ``Ovalbumin gene: evidence for a leader sequence in mRNA and DNA sequences at the exon-intron boundaries.'' Proc. Natl. Acad. Sci. USA, vol. 75, no. 10, pp. 4853-4857, 1978.

[14] M. Burset, IASeledtsov, and VV Solovyev, ``Analysis of canonical and non-canonical splice sites in mammalian genomes.,'' Nucleic Acids Res., vol. 28, no. 21, pp. 4364-4375, 2000.

[15] MBShapiro and P.Senapathy, ``RNA splice junctions of different classes of eukaryotes: Sequence statistics and functional implications in gene expression,'' Nucleic Acids Res., vol. 15, no. 17, pp. 7155-7174, 1987.

[16] TWNilsen,''Twenty years of RNA:then and now,'' pp.471-473,2015.

[17] RAPadgetr, PJ Grabowski, MM Konarska, S. Seiler, and PASharp, ``Splicing of messenger RNA precursors,'' 1986.

[18] MRGreen,''Pre-mRNA splicing,'' Annu.Rev.Genet.,vol.20,pp.671-708,1986.

[22] Y. Kapustin, E. Chan, R. Sarkar, F. Wong, I. Vorechovsky, R. M. Winston, T. Tatusova, and N. J. Dibb, ``Cryptic splice sites and split genes,'' Nucleic Acids Res., vol. 39, no. 14, pp. 5837-5844, 2011. [Examples]

[0163] Protein determination: The protein concentration was determined by calculating the optical density (OD) at 280 nm using the molar extinction coefficient calculated based on the amino acid sequence.

[0164] Recombinant DNA technology DNA was manipulated using standard methods, as described in Sambrook, J., et al., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York (1989). Molecular biological reagents were used according to the manufacturer's instructions.

[0165] Example 1 Antibody expression and purification from different codon-optimized nucleic acids Different non-optimized variable domains or codon-optimized variable domains were combined with nucleic acids encoding the wild-type human constant region or CHO codon-optimized nucleic acids encoding the human constant region.

[0166] Expression plasmid Each expression plasmid contained one expression cassette for either heavy chain or light chain expression. These were assembled separately in mammalian cell expression vectors.

[0167] General information regarding human light and heavy chain nucleotide sequences that can be used to estimate codon usage is provided below: Kabat et al., Sequences of Proteins of Immunological Interest, 5th Ed., Public Health Service, National Institutes of Health, Bethesda, MD (1991), NIH Publication No 91-3242.

[0168] In addition to the light chain or heavy chain expression cassette, these plasmids include the following: - Hygromycin resistance gene, - Origin of replication of the Epstein-Barr virus (EBV), oriP - A replication origin from the vector pUC18 that enables replication of this plasmid in E. coli, and - The β-lactamase gene that confers ampicillin resistance in E. coli.

[0169] Recombinant DNA technology: Cloning was performed using standard cloning techniques described in Sambrook, J., et al., Molecular Cloning: A Laboratory Manual, second edition, Cold Spring Harbor Laboratory Press (1989). All molecular biological reagents were commercially available (unless otherwise specified) and were used according to the manufacturer's instructions.

[0170] DNA and protein sequence analysis, and sequence data management: Vector NTI Advance Suite version 9.0 was used for sequence creation, mapping, analysis, annotation, and illustration.

[0171] Antibody expression: Human fetal kidney cells (HEK) 293F were used for antibody expression. HEK293 cells are primarily used for transient gene expression. The desired protein can be recovered after several days. HEK293F cells were cultured in a shaking flask at 7% CO2, 85% humidity, and 37°C in FreeStyle® 293 expression medium (Gibco, Invitrogen®, Life Technologies) supplemented with penicillin and streptomycin (PenStrep) and protein-free. Cells were passaged every 3-4 days and divided according to cell density. The cell count was always 3 × 10⁶. 5 Cell count / mL was used. Cell count and vitality were measured using a CASEY Cell Counter (Roche) with 50 μl of culture suspended in 10 ml of Casyton, using an appropriate program.

[0172] Due to transient transfection, the number of cells was reduced to 2 × 10⁶ on the same day. 6The cell volume was adjusted to 4 cells / mL. The number of transient transfections should not be less than 4 passages and should not exceed 22 passages. For transfection, the transfection reagent PEIpro was used and the Fed Batch process was carried out for 7 days. For transfection, the following volumes and amounts of DNA were used in the transfection mixture (culture volume 20 mL): TIFF0007835713000012.tif49128

[0173] For transfection, appropriate amounts of F17, DNA, and transfection reagents were combined in the appropriate order, mixed, and incubated for 10 minutes. Subsequently, the transfection mix was added to the cells. After approximately 3–5 hours, the corresponding pre-diluted VPA was added. After 16 hours, the culture was supplemented with 0.6% glucose and 12% feed, and incubated further. TIFF0007835713000013.tif44128

[0174] One week later, the supernatant could be collected. For this purpose, the mixture was centrifuged at 1200 rpm for 20 minutes, and the supernatant was sterilized and filtered through a 0.22 μm filter.

[0175] Determination of expression yield: The antibodies expressed from the supernatant were quantified using an HPLC column packed with protein A. Protein A is a cell wall-related protein derived from Staphylococcus aureus that specifically binds to the constant region of IgG antibodies. Protein A is immobilized on a polymer support and then binds to the target molecule. By changing various parameters such as pH and temperature, the antibodies can be eluted and detected.

[0176] Seven days after transfection, the HEK293 cell supernatant was collected. The recombinant antibody contained therein was purified from the supernatant by affinity chromatography using Protein A-Sepharose® affinity chromatography (GE Healthcare, Sweden). Briefly, the clarified culture supernatant containing the antibody was applied to a MabSelectSuRe Protein A (5-50 ml) column equilibrated with PBS buffer (10 mM Na2HPO4, 1 mM KH2PO4, 137 mM NaCl and 2.7 mM KCl, pH 7.4). Unbound proteins were washed away with the equilibration buffer. The antibody (or derivative) was eluted with 50 mM citrate buffer (pH 3.2). The protein-containing fraction was neutralized with 0.1 ml of 2 M Tris buffer (pH 9.0).

[0177] Sequence information SEQUENCE LISTING <110> F. Hoffmann-La Roche AG <120> Method for the production of an antibody <150> EP19183171.8 <151> 2019-06-28 <160> 6 <170> PatentIn version 3.5 <210> 1 <211> 6 <212> DNA <213> Artificial Sequence <220> <223> splice donor consensus sequence <400> 1 ggtrag 6 <210> 2 <211> 8 <212> DNA <213> Artificial Sequence <220> <223> high GC content donor consensus sequence <400> 2 aggtragt 8 <210> 3 <211> 8 <212> DNA <213> Artificial Sequence <220> <223> low GC content donor consensus sequence <400> 3 aggtaagt 8 <210> 4 <211> 7 <212> DNA <213> Artificial Sequence <220> <223> splice acceptor sequence <220> <221> misc_feature <222> (3)..(3) <223> n is a, c, g, or t <400> 4 ctnctag 7 <210> 5 <211> 39 <212> DNA <213> Artificial Sequence <220> <223> codon motif <400> 5 gcggcggcgg cggctgctgc tgctgcagca gcagccgcc 39 <210> 6 <211> 36 <212> DNA <213> Artificial Sequence <220> <223> codon motif <400> 6 gcggctgcag ccgcggctgc agccgcggct gcagcg 36

Claims

1. A method for producing antibodies of a human IgG1 subclass by culturing CHO cells transfected with one or more expression cassettes containing nucleic acids encoding the heavy and light chains of an antibody, The nucleic acids encoding the heavy and light chains of the antibody are codon-optimized for codon use in human cells and / or CHO cells. In the portion of the nucleic acid encoding the light chain, at least one unpaired donor splice site has been removed. The removal of the aforementioned unpaired donor splice site is achieved by introducing an amino acid silent mutation into the nucleotide sequence NGGTA(G)AG. The aforementioned method.

2. The method according to claim 1, wherein the portion of the nucleic acid encoding the light chain constant region has not had the unpaired donor splice site removed.

3. The method according to claim 1 or 2, wherein the transfection is transient transfection.

4. The method according to any one of claims 1 to 3, wherein the nucleic acid encoding the antibody is cDNA.

5. The method according to any one of claims 1 to 4, wherein the nucleic acid encoding the light chain and / or heavy chain of the antibody is genome-organized DNA.

6. The method according to any one of claims 1 to 5, wherein the removal of the unpaired donor splice site is achieved by introducing an amino acid silent mutation at codon NGG, codon GGT, or codon GTA(G) in the nucleotide sequence NGGTA(G)AG.

7. The aforementioned method is as follows: a) A step of culturing the CHO cells in a culture medium, and b) A step of recovering the antibody from the CHO cells or culture medium after step a). The method according to any one of claims 1 to 6, including the method described in any one of claims 1 to 6.

8. A method for producing antibodies of the human IgG1 subclass by culturing CHO cells containing one or more nucleic acids encoding the heavy and light chains of an antibody, In nucleic acids encoding the light chain, the unpaired donor splice site is removed by introducing a silent nucleotide change in the amino acid sequence at codon NGG, codon GGT, or codon GTA(G) into the unpaired donor splice site consensus sequence NGGTA(G)AG. In nucleic acids encoding the light chain constant region, the unpaired donor splice site by the sequence NGGTA(G)AG has not been removed. The aforementioned method.

9. The method according to claim 8, wherein the one or more nucleic acids encoding the heavy chain and light chain of the antibody are genome-organized DNA.

10. The method according to claim 8 or 9, wherein the one or more nucleic acids encoding the heavy and light chains of the antibody are codon-optimized for human cell codon use and / or CHO cell codon use.

11. The method according to any one of claims 8 to 10, wherein in the nucleic acid encoding the light chain of the antibody, the unpaired donor splice site is removed by introducing a silent nucleotide change in the amino acid sequence at codon NGG, codon GGT, or codon GTA(G) in the unpaired donor splice site consensus sequence NGGTA(G)AG.

12. The use of removing unpaired donor splice sites in portions of antibody-coding nucleic acid sequences, optimized for the use of human or hamster codons, Therefore, the aforementioned part is the part that codes for the light chain, The removal of the aforementioned unpaired donor splice site is achieved by introducing an amino acid silent mutation into the nucleotide sequence NGGTA(G)AG. The nucleic acid encoding the light chain of the antibody is genome-organized DNA. The use described above is intended to reduce missplicing when the nucleic acid is used to produce antibodies in CHO cells. The aforementioned use.

13. The use according to claim 12, wherein the unpaired donor splice site has not been removed from the nucleic acid portion encoding the light chain constant region.

14. The use according to claim 12 or 13, wherein the unpaired donor splice site is removed in the nucleic acid portion encoding the heavy chain.

15. The use according to any one of claims 12 to 14, wherein the removal of the unpaired donor splice site is achieved by introducing an amino acid silent mutation in the nucleotide sequence NGGTA(G)AG at codon NGG, codon GGT, or codon GTA(G).

16. The method according to any one of claims 1 to 11, wherein the unpaired donor splice site is an artificial unpaired donor splice site.

17. The method according to any one of claims 1 to 11, wherein the unpaired donor splice site is an artificial donor splice site that is generated during codon optimization.

18. The use according to any one of claims 12 to 15, wherein the unpaired donor splice site is an artificial unpaired donor splice site.

19. The use according to any one of claims 12 to 15, wherein the unpaired donor splice site is an artificial donor splice site that is generated during codon optimization.

Citation Information

Patent Citations

  • Heavy chain variants resulting in improved immunoglobulin production

    JP2010531648A

  • Method for expressing polypeptides using modified nucleic acids

    JP2015514406A

  • Expression constructs and methods for selecting host cells that express polypeptides

    JP2017508459A

  • Immunoglobulin display vectors

    US20090136950A1

  • Codon pair utilization

    US5082767A