Methods of producing antibodies

By removing non-paired donor splicing sites from the nucleic acid of the antibody heavy chain variable domain and optimizing codon usage, the problem of reduced expression yield during codon optimization was solved, and the antibody expression yield was improved.

CN113993888BActive Publication Date: 2026-01-06F HOFFMANN LA ROCHE & CO AG
View PDF 23 Cites 0 Cited by

Patent Information

Application Number
CN202080045482.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-28
Filing Date
2020-06-25
Publication Date
2026-01-06
Estimated Expiration
2040-06-25

AI Technical Summary

Technical Problem

During codon optimization, the generation of unpaired donor splicing sites leads to a decrease in antibody expression yield, which is difficult to effectively solve with existing technologies.

Method used

By removing non-paired donor splicing sites from nucleic acids encoding the variable domain of the antibody heavy chain and not removing these sites from the coding constant region, codon usage is optimized by introducing amino acid sequence silencing nucleotide changes.

Benefits of technology

It increased antibody expression yield, reduced missplicing events, and improved expression rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003423434240000081
    Figure BDA0003423434240000081
  • Figure BDA0003423434240000091
    Figure BDA0003423434240000091
  • Figure BDA0003423434240000161
    Figure BDA0003423434240000161
Patent Text Reader

Abstract

This disclosure discloses a method for generating IgG1 antibodies by culturing CHO cells containing / transfected with one or more (exogenous) nucleic acids encoding (and expressing) an antibody, wherein unpaired splicing sites are removed from the nucleic acid encoding a variable domain of the heavy chain, and these unpaired splicing sites are not removed from the nucleic acid encoding a constant region of the heavy chain.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This article reports a method for producing antibodies in which the encoding nucleic acid is optimized for donor splicing sites only in the portion encoding the variable structural domain. Background Technology

[0002] Cannarozzi, G. et al. reported the role of codon order in translation dynamics (Cell 141(2010)355-367). Plotkin, JB and Kudla, G. (Nat. Rev. Gen. 12(2011)32-42) reported the causes and consequences of codon bias. Weygand-Durasevic, I. and Ibba, M. reported new roles in codon usage (Science 329(2010)1473-1474). Itzkovitz, S. et al. reported overlapping codons in protein-coding sequences (Gen. Res. 20(2010)1582-1589).

[0003] High-level protein expression was reported in WO 97 / 11086. The production of plant polypeptides was reported in WO 03 / 70957. A method for designing synthetic nucleic acid sequences to optimize protein expression in host cells was described in WO 03 / 85114. Codon pair optimization was reported in US 5,082,767. A method for achieving improved polypeptide expression was reported in WO2008 / 000632. Codon optimization methods were reported in WO 2007 / 142954 and US 8,128,938.

[0004] Watkins, NE et al. reported the nearest-neighbor thermodynamics of deoxyinosine pairs in DNA duplexes (Nucl. Acids Res. 33 (2005) 6258-6267).

[0005] A method for expressing peptides using modified nucleic acids was reported in WO 2013 / 156443.

[0006] Zhang, MQ reported the statistical characteristics of human exons and their flanking regions (Hum.Mol.Genet.7(1998)919-932).

[0007] mRNA splicing is jointly regulated by the presence of donor and acceptor splicing sites, which are located at the 5' and 3' ends of introns, respectively. According to Watson et al. (Eds, Recombinant DNA: A Shortcourse, Scientific American Books, distributed by WH Freeman and Company, New York, USA (1983)), the 5' donor splicing site is ag|gtragt (exon|intron) and the 3' acceptor splicing site is y. n The common sequence of Ncag|g (intron|exon) (r = purine base; y = pyrimidine base; n = integer; N = any natural base).

[0008] The first articles on the origin of secretory and membrane-bound immunoglobulins were published in 1980. The formation of secretory (sIg) and membrane-bound (mIg) isoforms originates from the alternative splicing of heavy chain precursor mRNA. In the mIg isoform, the encoding of the secretory C-terminal domain (i.e., Cs, Cm ... H 3 or C H The donor splicing site in the exon of the 4-domain structure and the acceptor splicing site located at a certain distance downstream are used to connect the constant region to the downstream exon encoding the transmembrane domain.

[0009] WO 2002 / 016944 reports a method for preparing synthetic nucleic acid molecules that exhibit reduced inappropriate or unintended transcriptional signatures when expressed in specific host cells. WO 2006 / 042158 reports modified nucleic acid molecules to enhance recombinant protein expression and / or reduce or eliminate missplicing and / or intron readthrough byproducts. Magistrelli, G. et al. reported optimizing the assembly and production of natural bispecific antibodies through codon de-optimization (MABS 9(2016)231-239).

[0010] WO 2015 / 128509 reports expression constructs and methods for selecting host cells to express peptides.

[0011] WO 2009 / 003623 reported a heavy chain mutant that resulted in increased immunoglobulin production. Summary of the Invention

[0012] For the production of therapeutic or diagnostic antibodies, high expression yield is the goal. A common approach to achieving good expression rates is to first optimize the codon usage of the nucleic acid encoding nucleic acid by adjusting it to the codon usage of cells intended to express exogenous nucleic acids. This codon adjustment or optimization can be based on different established protocols.

[0013] However, during such codon adjustment and optimization processes, such as unpaired splicing sites, especially unpaired donor splicing sites, can be unintentionally generated from scratch. That is, for example, during codon optimization, new donor splicing site sequences can be generated in codon-optimized nucleic acids by unintentionally generating sequence motifs that follow the shared sequence of donor splicing sites. This event can occur independently of the organization of codon-optimized nucleic acids, meaning it is possible for both cDNA and genomically organized nucleic acids. It is essentially an unintended side effect of the codon optimization process. This new donor splicing site is an additional artificial donor splicing site without an associated target receptor splicing site. Therefore, this unpaired donor splicing site can induce a splicing event where a random, i.e., unrestricted receptor splicing site exists somewhere in the transcribed mRNA. Consequently, expression yield decreases due to the formation of byproducts.

[0014] This invention is at least partly based on the unexpected discovery that the removal of unpaired donor splicing sites in antibody heavy chain-encoded nucleic acids using optimized codons only needs to be performed in the portion of the nucleic acid encoding the variable structural domain of the heavy chain, but not in the portion encoding the constant region. For the constant region, germline or wild-type human nucleic acid sequences can be used, for example. Thus, the expression yield of antibody heavy chains of the correct length can be increased or made possible.

[0015] This invention is based at least in part on the discovery that introducing an amino acid sequence silencing nucleotide change (mutation) in the unpaired donor splicing site concordant sequence NGGTA(G)AG (SEQ ID NO:01) in a codon-optimized nucleic acid encoding the variable domain of the antibody heavy chain is sufficient to increase expression yield.

[0016] One aspect of the present invention is a method for producing antibodies by culturing mammalian cells containing / transfected with one or more (exogenous) nucleic acids encoding antibody heavy and light chains (and expressing antibodies).

[0017] The one or more (exogenous) nucleic acids mentioned herein are codons optimized for codon usage in human cells and / or mammalian cells.

[0018] In the nucleic acid encoding the variable structural domain of the heavy strand, at least one (artificial) unpaired donor splicing site is removed, and optionally, unpaired donor splicing sites are not removed in the nucleic acid sequence encoding the constant region of the heavy strand (artificial) (human wild-type or human or hamster codon optimized).

[0019] In one embodiment, the antibody is an antibody against a human IgG1 subclass. In one embodiment, the antibody is a humanized antibody against a human IgG1 subclass. In one embodiment, the constant region of the antibody contains mutations suitable for inducing heterodimerization or modifying Fc-receptor binding.

[0020] In one embodiment, the mammalian cell is a CHO cell.

[0021] In one embodiment, the transfection is a transient transfection.

[0022] In one embodiment, one or more (exogenous) nucleic acids encoding antibodies are cDNA.

[0023] In one embodiment, one or more (exogenous) nucleic acids encoding antibody heavy chains and / or antibody light chains are DNA in a genomically organized form, i.e., having an intron-exon organized form.

[0024] In one embodiment, removal of the unpaired donor splicing site is achieved by introducing an amino acid silencing change (mutation) in the amino acid sequence NGGTA(G)AG (SEQ ID NO:01). In another embodiment, an amino acid sequence silencing nucleotide change is introduced in the codon NGG, codon GGT, or codon GTA(G) of SEQ ID NO:01.

[0025] In one embodiment, codon usage optimization is performed based on human codon usage or Chinese hamster codon usage.

[0026] In one embodiment, the nucleic acid encoding the antibody light chain is codon-optimized, meaning that both the variable domains and constant regions are codon-optimized.

[0027] In one embodiment, unpaired donor splicing sites are removed from the whole light chain encoding nucleic acid.

[0028] In one embodiment, the transfection is a stable transfection.

[0029] In one embodiment, the method includes the following steps:

[0030] a) Culturing mammalian cells, and

[0031] b) Recover the antibody from the cells or culture medium.

[0032] One aspect of the invention is the use of removing non-paired donor splicing sites only in a portion of a nucleic acid sequence encoding an antibody using an optimized human or hamster codon, for reducing missplicing and / or increasing antibody expression yield when said nucleic acid is used to generate antibodies in CHO cells, wherein the portion is the portion encoding a heavy chain variable domain.

[0033] In one embodiment, the unpaired donor splicing site in the portion of the coding heavy strand constant region of the nucleic acid is not removed.

[0034] In one embodiment, a non-paired donor splicing site is further removed from the portion of the nucleic acid encoding the light strand.

[0035] In one embodiment, the removal of unpaired splicing sites is achieved by introducing an amino acid silencing mutation in the nucleotide sequence NGGTA(G)AG (SEQ ID NO: 01).

[0036] In one embodiment, the removal of the unpaired splicing site is achieved by introducing an amino acid silencing mutation at the codon NGG, codon GGT, or codon GTA(G) in the nucleotide sequence NGGTA(G)AG (SEQ ID NO: 01).

[0037] In one embodiment of all aspects and embodiments, the unpaired (donor) splice site is an artificial unpaired (donor) splice site.

[0038] In one embodiment of all aspects and examples, the unpaired (donor) splice site is an artificial unpaired (donor) splice site and has been generated during codon optimization. Detailed Implementation

[0039] For the production of therapeutic or diagnostic antibodies, high expression yield is the goal. One option for achieving good expression rates is to first optimize the codon usage encoding the nucleic acid and then adapt it to the codon usage of cells designed to express the foreign nucleic acid. This codon tuning or optimization can be done based on various established protocols.

[0040] However, during such codon adjustment and optimization processes, unpaired splicing sites can be generated de novo. That is, during codon optimization, new donor splicing site sequences are generated in the codon-optimized nucleic acid. This is independent of the organization of the codon-optimized nucleic acid; either cDNA or genomically organized nucleic acid is possible. It is actually an unintended side effect of the codon optimization process. Because such a new donor splicing site is an additional artificial donor splicing site, it has no associated target receptor splicing site. Therefore, this unpaired donor splicing site can induce random, i.e., unrestricted receptor splicing sites present somewhere in the transcribed mRNA, thereby reducing expression yield.

[0041] This invention is based, at least in part, on the unexpected discovery that the removal of unpaired donor splice sites in nucleic acids encoded by optimized antibody heavy chains only needs to be performed in the portion of the nucleic acid encoding the variable structural domain of the heavy chain, but not in the portion encoding the constant region. For the constant region, germline or wild-type human nucleic acid sequences can be used, for example.

[0042] This invention is based at least in part on the discovery that introducing an amino acid sequence silencing nucleotide change (mutation) in the unpaired donor splicing site concordant sequence NGGTA(G)AG (SEQ ID NO:01) in a codon-optimized nucleic acid encoding the variable domain of the antibody heavy chain is sufficient to increase expression yield.

[0043] definition

[0044] The methods and techniques used to carry out the present invention are known to those skilled in the art and are described, for example, in Ausubel, FM, ed., Current Protocols in Molecular Biology, Volumes I to III (1997), and Sambrook et al., Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (1989). As is known to those skilled in the art, various derivatives of nucleic acids and / or polypeptides can be produced using recombinant DNA technology. Such derivatives can be modified, for example, at one or more sites by substitution, alteration, exchange, deletion, or insertion. Modification or derivatization can be performed, for example, by site-directed mutagenesis. Such modifications can be readily performed by those skilled in the art (see, for example, Sambrook, J. et al., Molecular Cloning: A laboratory manual (1999), Cold Spring Harbor Laboratory Press, New York, USA). The use of recombinant technology enables those skilled in the art to transform various host cells with heterologous nucleic acids. Although transcription and translation (i.e., expression mechanisms) in different cells use the same elements, cells belonging to different species may have different so-called codon usage, etc. Therefore, the same polypeptide (for an amino acid sequence) can be encoded by different nucleic acids. Furthermore, due to the degeneracy of the genetic code, different nucleic acids can encode the same polypeptide.

[0045] The term "approximately" indicates that the value following it is not an exact value, but rather the midpoint of a range of + / -10%, + / -5%, + / -2%, or + / -1% of the value. If the value is a relative value given as a percentage, the term "approximately" also indicates that the value following it is not an exact value, but rather the midpoint of a range of + / -10%, + / -5%, + / -2%, or + / -1% of the value, where the upper limit of the range cannot exceed 100% of the value.

[0046] As used in this application, the term "amino acid" refers to the carboxyl-α amino acid group, which can be encoded by nucleic acids directly or in precursor form. A single amino acid is encoded by a nucleic acid consisting of three nucleotides, a so-called codon or base triplet. Each amino acid is encoded by at least one codon. The encoding of the same amino acid by different codons is referred to as "degeneracy of the genetic code." As used in this application, the term "amino acid" refers to naturally occurring carboxyl α-amino acids, including alanine (three-letter code: ala, single-letter code: A), arginine (arg, R), asparagine (asn, N), aspartic acid (asp, D), cysteine ​​(cys, C), glutamine (gln, Q), glutamic acid (glu, E), glycine (gly, G), histidine (his, H), isoleucine (ile, I), leucine (leu, L), lysine (lys, K), methionine (met, M), phenylalanine (phe, F), proline (pro, P), serine (ser, S), threonine (thr, T), tryptophan (trp, W), tyrosine (tyr, Y), and valine (val, V).

[0047] The term “immunoglobulin” in this document is used in the broadest sense and covers a variety of immunoglobulin structures, including but not limited to monoclonal antibodies, polyclonal antibodies, and multispecific antibodies (e.g., bispecific antibodies) or fragments thereof containing at least a portion of a constant domain or region.

[0048] As used herein, the term "immunoglobulin" refers to a protein composed of one or more polypeptides essentially encoded by immunoglobulin genes. This definition includes variants such as mutant forms (i.e., forms with substitutions, deletions, and insertions of one or more amino acids), N-terminal truncations, fusion forms, chimeric forms, and humanized forms. Recognized immunoglobulin genes include distinct constant region genes and numerous immunoglobulin variable region genes from, for example, primates (including humans) and rodents. Monoclonal immunoglobulins are preferred. Each of the heavy and light polypeptide chains of an immunoglobulin contains a constant region (typically the carboxyl-terminal portion).

[0049] As used herein, the term "monoclonal immunoglobulin" refers to an immunoglobulin derived from a substantially homogeneous group of immunoglobulins, meaning that the individual immunoglobulins contained in this group are identical, except for a small number of naturally occurring mutations that may be present. Monoclonal immunoglobulins exhibit high specificity for a single antigenic site. Furthermore, unlike polyclonal immunoglobulin formulations which include different antibodies targeting different antigenic sites (determinants or epitopes), each monoclonal immunoglobulin targets a single antigenic site on the antigen. In addition to specificity, a key advantage of monoclonal immunoglobulins is that they can be synthesized without contamination from other immunoglobulins. The modifier "monoclonal" indicates that the immunoglobulin is derived from a substantially homogeneous group of immunoglobulins and should not be interpreted as requiring any particular method of immunoglobulin production.

[0050] The term "codon" refers to an oligonucleotide consisting of three nucleotides that encode a specific amino acid. Due to the degeneracy of the genetic code, most amino acids are encoded by more than one codon. These different codons encoding the same amino acid have different relative frequencies of use in a single host cell. Therefore, a particular amino acid may be encoded by exactly one codon or by a different set of codons. Similarly, the amino acid sequence of a polypeptide can be encoded by different nucleic acids. Therefore, a specific amino acid (residue) in a polypeptide may be encoded by a different set of codons, each of which has a frequency of use within a given host cell.

[0051] Since a large number of gene sequences are available for many common host cells, the relative frequencies of codon usage can be calculated. Tables of calculated codon usage can be obtained, for example, from the "Codon Usage Database" (www.kazusa.or.jp / codon / ), Nakamura, Y., et al., Nucl. Acids Res. 28 (2000) 292.

[0052] The codon usage tables for Homo sapiens and hamsters are obtained from "EMBOSS: The European Molecular Biology Open Software Suite" (Rice, P., et al., Trends Gen. 16 (2000) 276-277, Release 6.0.1, 15.07. 2009), as shown in the table below. The different codon usage frequencies for 20 natural amino acids in E. coli, yeast, human cells, and CHO cells are calculated for each amino acid, not for all 64 codons.

[0053] Table: Frequency of Overall Codon Use in Homo sapiens

[0054] (Encoded amino acid | codon | frequency of use [%))

[0055]

[0056] Table: Overall Codon Usage Frequency in Hamsters

[0057] (Encoded amino acid | codon | frequency of use [%))

[0058]

[0059] As used herein, the term “expression” refers to the transcription and / or translation process that occurs within a cell. The transcriptional level of a target nucleic acid sequence in a cell can be determined based on the amount of the corresponding mRNA present in the cell. For example, mRNA transcribed from a target sequence can be quantified by RT-PCR (qRT-PCR) or by Northern hybridization (see Sambrook, J., et al., 1989, ibid.). Polypeptides encoded by the target nucleic acid can be quantified by a variety of methods, such as by ELISA, by measuring the biological activity of the polypeptide, or by using assays unrelated to this activity, such as Western blotting or radioimmunoassay, using immunoglobulins that recognize and bind to the polypeptide (see Sambrook, J., et al., 1989, ibid.).

[0060] An "expression cassette" is a construct that contains necessary regulatory elements, such as promoters and polyadenylation sites, for expressing at least the contained nucleic acids in a cell.

[0061] Gene expression can occur transiently or permanently. The target polypeptide is typically a secreted polypeptide and therefore contains an N-terminal extension (also known as a signal sequence), which is essential for the polypeptide to be transported across the cell membrane / secreted into the extracellular medium. Generally, the signal sequence can be derived from any gene encoding the secreted polypeptide. If a heterologous signal sequence is used, it is preferably a heterologous signal sequence that is recognized and processed by the host cell (i.e., cleaved by a signal peptidase). For example, for secretion in yeast, the native signal sequence of the heterologous gene to be expressed can be replaced by a homologous yeast signal sequence derived from the secreting gene, such as a yeast invertase signal sequence, an α-factor leader sequence (including α-factor leader sequences from *Saccharomyces*, *Kluyveromyces*, *Pichia pastoris*, and *Hansenula polymorpha*, the second described in US 5,010,182), an acid phosphatase signal sequence, or a *Candida albicans* glucosylamylase signal sequence (EP 0 362 179). In mammalian cell expression, the native signal sequence of the target protein is satisfactory, although other mammalian signal sequences may be appropriate, such as signal sequences of secreted peptides from the same or related species (e.g., for immunoglobulins of human or mouse origin), and viral secretion signal sequences, such as the herpes simplex glycoprotein D signal sequence. The DNA fragment encoding this preamble is operatively linked within the frame to the DNA fragment encoding the target peptide.

[0062] The term “cell” or “host cell” refers to a cell in which nucleic acids (e.g., nucleic acids encoding heterologous polypeptides) can be transfected or transfected. The term “cell” includes prokaryotic cells for expressing nucleic acids and producing encoded polypeptides (including the proliferation of plasmids) and eukaryotic cells for expressing nucleic acids and producing encoded polypeptides. In one embodiment, the eukaryotic cell is a mammalian cell. In one embodiment, the mammalian cell is a CHO cell, optionally a CHO K1 cell (ATCC CCL-61 or DSMAC 110), or a CHO DG44 cell (also known as CHO-DHFR[-], DSM ACC 126), or a CHO XL99 cell, a CHO-T cell (see, for example, Morgan, D., et al., Biochemistry 26 (1987) 2959-2963), or a CHO-S cell, or a Super-CHO cell (Pak, SCO, et al., Cytotechnology 22 (1996) 139-146). If these cells are not suitable for growth in serum-free media or suspensions, adjustments should be made before using them in the current method. As used herein, the term "cell" includes the test cell and its progeny. Therefore, the terms "transformer" and "transformed cell" include the primary test cell and cultures derived from that cell, regardless of the number of transfers or subcultures. It should also be understood that all progeny may not be exactly identical in DNA content due to intentional or unintentional mutations. This includes variant progeny with the same function or biological activity, such as those screened from the original transformed cells.

[0063] The term "codon-optimized nucleic acid" refers to a nucleic acid that has been modified to improve the expression of a polypeptide in cells (e.g., mammalian cells) by replacing one, at least one, or more codons of the parent polypeptide-encoded nucleic acid with codons that encode the same amino acid residues (e.g., with different relative frequencies of use in the cell).

[0064] As used herein, the term "unpaired donor splicing site" refers to a donor splicing site that, on the one hand, has been artificially generated within the nucleic acid sequence (e.g., through codon optimization of the nucleic acid sequence), and on the other hand, has no recipient splicing site downstream of the nucleic acid sequence due to its artificial introduction into the nucleic acid sequence. Although it conforms to the common sequence of donor splicing sites, it does not follow the (biological) splicing principle, i.e., the removal of unwanted nucleic acid portions during processing.

[0065] A "gene" refers to a segment of nucleic acid, such as those found on chromosomes or plasmids, that can influence the expression of peptides, polypeptides, or proteins. In addition to coding regions, i.e., structural genes, genes also contain other functional elements, such as signaling sequences, promoters, introns, and / or terminators.

[0066] The term "codon set" and its semantic equivalents refer to a limited number of different codons encoding one (i.e., the same) amino acid residue. Individual codons within a set have varying overall frequencies of use in the cellular genome. Each codon in a set has a specific frequency of use within the set, depending on the number of codons in the set. This specific frequency of use within a set may differ from the overall frequency of use in the cellular genome, but depends on (and is related to) the overall frequency of use. A codon set can contain only one codon, or it can contain up to six codons.

[0067] The term "overall frequency of use in the cell genome" refers to the frequency of a particular codon's appearance throughout the cell's genome.

[0068] The term "specific frequency of use" of a codon within a codon set refers to the frequency with which a single (i.e., a specific) codon within a codon set can be found in a nucleic acid encoding a polypeptide obtained using the methods reported herein, relative to all codons in the codon set. The value of the specific frequency of use depends on the overall frequency of use of the specific codon in the cellular genome and the number of codons within the set. Therefore, since a codon set does not necessarily contain all possible codons encoding a specific amino acid residue, the specific frequency of use of a codon within a codon set is at least the same as its overall frequency of use in the cellular genome, and at most 100%, meaning it is at least identical, but it can be higher than the overall frequency of use in the cellular genome if some less frequently used codons are excluded from the set. The sum of the specific frequency of use of all members of a codon set is always approximately 100%.

[0069] The term "amino acid codon motif" refers to a sequence of codons that are all members of the same set of codons and therefore encode the same amino acid residue. The number of different codons in an amino acid codon motif is the same as the number of different codons in a set of codons, but each codon can appear multiple times in the amino acid codon motif. Furthermore, each codon exists in the amino acid codon motif with its specific frequency of use. Therefore, an amino acid codon motif represents a sequence of different codons encoding the same amino acid residue, where each different codon exists with its specific frequency of use, where the sequence begins with the codon with the highest specific frequency of use, and where the codons are arranged in a defined order. For example, the codon set encoding the amino acid residue alanine includes four codons GCG, GCT, GCA, and GCC, with specific frequencies of use of 32%, 28%, 24%, and 16% (corresponding to a 4:3:3:2 ratio), respectively. The amino acid codon motif for the amino acid residue alanine is defined as containing four codons GCG, GCT, GCA, and GCC in a 4:3:3:2 ratio, where the first codon is GCG. An exemplary amino acid codon motif for alanine is gcg gct gca gcc gcg gct gca gcc gcg gct gca gcg (SEQ ID NO:06). This motif consists of twelve consecutive codons (4+3+3+2=12). When the amino acid residue alanine first appears in the amino acid sequence of a polypeptide, the first codon of the amino acid codon motif is used for the corresponding coding nucleic acid. Upon the second appearance of alanine, the second codon of the amino acid codon motif is used, and so on. When alanine appears for the thirteenth time in the amino acid sequence of the polypeptide, the thirteenth codon, i.e., the last position, is used in the corresponding coding nucleic acid. Upon the thirteenth appearance of alanine in the amino acid sequence of the polypeptide, the first codon of the amino acid codon motif is used again, and so on.

[0070] The terms “nucleic acid” or “nucleic acid sequence,” used interchangeably in this application, refer to a polymeric molecule composed of individual nucleotides (also called bases) a, c, g, and t (or u in RNA), such as DNA, RNA, or modifications and mixtures thereof. The polynucleotide molecule can be a naturally occurring polynucleotide molecule, a synthetic polynucleotide molecule, or a combination of one or more naturally occurring polynucleotide molecules with one or more synthetic polynucleotide molecules. This definition also includes naturally occurring polynucleotide molecules in which one or more nucleotides have been altered (e.g., by mutagenesis), deleted, or added. Nucleic acids can be isolated or integrated into another nucleic acid, such as in an expression cassette, plasmid, or the chromosome of a host cell. Nucleic acids are characterized by their nucleic acid sequence consisting of individual nucleotides.

[0071] For those skilled in the art, procedures and methods for converting an amino acid sequence (e.g., of a polypeptide) into a corresponding nucleic acid sequence encoding that amino acid sequence are well known. Thus, nucleic acids are characterized by their nucleic acid sequences consisting of individual nucleotides, and are equally characterized by the amino acid sequence of the polypeptide they encode.

[0072] "Structural genes" refer to gene regions without signal sequences, i.e., coding regions.

[0073] A transfection vector is a nucleic acid (also called a nucleic acid molecule) that provides all the necessary elements for the expression of a transfection vector containing a nucleic acid / structural gene in a host cell. A transfection vector contains a prokaryotic plasmid proliferation unit (e.g., for *E. coli*), a prokaryotic origin of replication and a nucleic acid conferring resistance to a prokaryotic selector, and further comprises the transfection vector, one or more nucleic acids conferring resistance to a eukaryotic selector, and one or more nucleic acids encoding a target polypeptide. Preferably, the selector-resistant nucleic acid and the nucleic acid encoding the target polypeptide are each housed within an expression cassette, whereby each expression cassette contains a promoter, a nucleic acid encoding the nucleic acid, and a transcription terminator including a polyadenylation signal. Gene expression is typically under the control of a promoter; such structural genes are referred to as being "operably linked" to the promoter. Similarly, if a regulatory element modulates the activity of a core promoter, the regulatory element and the core promoter are operably linked.

[0074] As used herein, the term "vector" refers to a nucleic acid molecule capable of carrying another nucleic acid linked to it. This term includes vectors that function as self-replicating nucleic acid structures, as well as vectors incorporated into the genome of a host cell into which they have been introduced. Some vectors are capable of directing the expression of nucleic acids operatively linked to them. Such vectors are referred to herein as "expression vectors."

[0075] The term "full-length antibody" refers to an antibody with a structure substantially similar to that of a natural antibody. A full-length antibody comprises two full-length antibody light chains and two full-length antibody heavy chains. Each full-length antibody light chain includes a light chain variable region and a light chain constant domain in the N-to-C-terminal direction. Each full-length antibody heavy chain includes a heavy chain variable region, a first heavy chain constant domain, a hinge region, a second heavy chain constant domain, and a third heavy chain constant domain in the N-to-C-terminal direction. In contrast to natural antibodies, full-length antibodies may contain additional immunoglobulin domains conjugated to one or more ends of chains different from the full-length antibody, such as one or more additional scFv, or heavy chain or light chain Fab fragments, or scFab, but only a single fragment is conjugated to each end. These conjugates are also covered by the term full-length antibody.

[0076] An antibody "class" refers to the type of constant domain or constant region possessed by the antibody's heavy chain, preferably the Fc-region. There are five major classes of antibodies: IgA, IgD, IgE, IgG, and IgM, and some of them can be further subdivided into subclasses (isotypes), such as IgG1, IgG2, IgG3, IgG4, IgA1, and IgA2. The constant domains of the heavy chain corresponding to different classes of immunoglobulins are respectively called α, δ, ε, γ, and μ.

[0077] The term "heavy chain constant region" refers to the region in the immunoglobulin heavy chain that contains constant structural domains, namely the CH1 domain, hinge region, CH2 domain, and CH3 domain. In one embodiment, the human IgG constant region extends from Ala118 to the carboxyl terminus of the heavy chain (according to the Kabat EU index number). However, the C-terminal lysine residue (Lys447) of the constant region may or may not be present (according to the Kabat EU index number). The term "heavy chain constant region" also refers to a dimer containing two heavy chain constant regions that are covalently linked to each other via cysteine ​​residues in the hinge region that form interchain disulfide bonds.

[0078] The term "light chain constant region" refers to the region in the immunoglobulin light chain that contains a constant structural domain, namely the CL domain.

[0079] The term "constant region" includes "heavy chain constant region" and "light chain constant region".

[0080] The term "heavy chain Fc-region" refers to the C-terminal region of an immunoglobulin heavy chain, which includes at least a portion of the hinge region (middle and lower hinge region), a CH2 domain, and a CH3 domain. In one embodiment, the human IgG heavy chain Fc-region extends from Asp221 or from Cys226 or from Pro230 to the C-terminus of the heavy chain (according to the Kabat EU index number). Thus, the Fc-region is smaller than the constant region but identical to it at the C-terminal portion. However, the C-terminal lysine (Lys447) of the heavy chain Fc-region may or may not be present (according to the Kabat EU index number). The term "Fc-region" also refers to a dimer containing two heavy chain Fc-regions that are covalently linked to each other via hinge region cysteine ​​residues forming interchain disulfide bonds.

[0081] The constant region of an antibody, more precisely the Fc-region (and similarly constant regions), is directly involved in complement activation, C1q binding, C3 activation, and Fc receptor binding. Although the effect of an antibody on the complement system depends on certain conditions, binding to C1q is caused by binding sites defined in the Fc region. Such binding sites are known in the prior art, and for example, by Lukas, TJ et al., J. Immunol. 127 (1981) 2555-2560; Brunhouse, R. and Cebra, JJ, Mol. Immunol. 16 (1979) 907-917; Burton, DR et al., Nature 288 (1980) 338-344; Thommesen, JE et al., Mol. Immunol. 37 (2000) 995-1004; Idusogie, EE et al., J. Immunol. 164 (2000) 4178-4184; Hezareh, M. et al., J. Virol. 75 (2001) 12161-12168; Morgan, A. et al., Immunology 86 (1995) 319-324; and EP As described in 0 307 434. Such binding sites are, for example, L234, L235, D270, N297, E318, K320, K322, P331, and P329 (according to Kabat EU index numbers). Antibodies against subclasses IgG1, IgG2, and IgG3 typically exhibit complement activation, C1q binding, and C3 activation, while IgG4 does not activate the complement system, does not bind C1q, and does not activate C3. "The Fc region of an antibody" is a term well-known to those skilled in the art and is defined based on papain cleavage of the antibody.

[0082] splicing

[0083] The constant region amino acid sequences of different human immunoglobulins are encoded by corresponding DNA sequences. In the genome, these DNA sequences contain coding (exons) and non-coding (intron) sequences. After DNA is transcribed into precursor mRNA, the precursor mRNA also contains these intron and exon sequences. Before translation, non-coding intron sequences are removed from the primary mRNA transcript during mRNA processing by splicing to produce mature mRNA. Splicing of primary mRNA is controlled by donor splicing sites and appropriately spaced recipient splicing sites. The donor splicing site is located at the 5' end of the intron sequence, and the recipient splicing site is located at the 3' end of the intron sequence.

[0084] The term "appropriately spaced" refers to the arrangement of donor and recipient splice sites in nucleic acids such that all the elements required for the splicing process are available and in the right positions to allow the splicing process to occur.

[0085] The donor splice site (5' splice site) is a nucleic acid sequence motif representing the 5' end of an intron.

[0086] The receptor splice site (3' splice site) is a nucleic acid sequence motif representing the 3' end of an intron.

[0087] Codon optimization:

[0088] The production of recombinant antibodies is based at least on constant regions of naturally occurring wild-type or germline sequences. To prevent biological limitations of these DNA templates, such as RNA instability, low nuclear output efficiency, secondary structure, or insufficient translation, bioinformatics techniques are used for gene optimization and subsequent de novo gene synthesis based on protein sequences

[10] . Thus, complex and time-consuming cloning steps can be avoided, and as has been shown in recent years, translation efficiency can be improved by such adjustments to codon usage in the production system

[10] . For codon optimization, high GC content, avoidance of splice sites, and codon usage adjustments play a central role in the production organism. Several methods have been established to examine the use and frequency of specific codons. It is well known that tRNAs encoding the same amino acids compete with each other. On the one hand, multivalent tRNAs that recognize multiple codons are more common, and on the other hand, different tRNA species are expressed at different frequencies. Therefore, codons encoded by more abundant tRNAs also appear more frequently in the coding sequence

[11] . Three possible methods for applying codon usage adjustments / optimization are shown in Table 1.

[0089] Table 1:

[0090]

[0091] One method of codon usage aims to make the entire tRNA library available for translation, namely Method 2. Compared to the “high” method, which uses only the most abundant codons to translate amino acids, all available codons are used. In addition to Method 1, Method 2 also takes into account the distribution of codons used in codon usage in each organism and assigns codons within the sequence

[11] .

[0092] Splice site:

[0093] Almost all protein-coding genes in eukaryotes are divided into exons and introns. Currently, 12 exon variants are known

[12] . Splicing is the removal of introns from precursor mRNA during transcription. Correct splicing is based on conserved common sequences located at the 5' and 3' ends of introns and so-called branching sites, respectively. The branching site is located about 20-50 nucleotides upstream of the 3' end of the intron

[12] . At the 5' end of the intron is the donor splice site with the characteristic dinucleotide GT. At the 3' end is the acceptor splice site located at the base AG

[13] . This pattern is called the classical pattern of splice sites. However, 3.7% of annotated splice sites do not follow this pattern

[14] . Splice sites with GC-AG, GG-AG, GT-TG, GT-CG, or CT-AG dinucleotides have also been observed within introns

[14] . Some of these non-classical splice sites may be associated with the expression of immunoglobulins

[15] . According to Burset et al., the dinucleotide GC-AG splice site is the most common non-classical pattern. In addition, the co-occurrence sequence is strongly dependent on GC content

[12] . At high GC content, the 5' end donor co-occurrence sequence is described as AG / GTRAGT (SEQ ID NO:02), rather than as AG / GTAAGT (SEQ ID NO:03) at low GC content

[12] .

[0094] Introns are cleaved by spliceosomes. Introns flanking the classic GT-AG pair are released from the precursor mRNA by spliceosomes with U1, U2, U4 / U6, and U5 subunits

[16] . Furthermore, eukaryotic genomes are known to have a variety of cryptic splicing sites that can negatively impact proper splicing. These are different from the shared motifs of splicing sites. Typically, these sites are inactive or rarely used by cellular machinery

[17] ,

[18] . Cryptonic splicing sites are present in both introns and exons. The goal of bioinformatics programs is to identify the location of these cryptic splicing sites. Many of these programs provide information, but due to the multiple possibilities, they are less complex than nucleotide information

[22] .

[0095] Specific embodiments of the method according to the present invention

[0096] For the production of therapeutic or diagnostic antibodies, high expression yield is the goal. One option for achieving good expression rates is to first optimize the codon usage encoding the nucleic acid and then adapt it to the codon usage of cells designed to express the foreign nucleic acid. This codon tuning or optimization can be done based on various established protocols.

[0097] However, during such codon adjustment and optimization processes, unpaired splicing sites can be generated de novo. That is, during codon optimization, new donor splicing site sequences are generated in the codon-optimized nucleic acid. This is independent of the organization of the codon-optimized nucleic acid; either cDNA or genomically organized nucleic acid is possible. It is actually an unintended side effect of the codon optimization process. Because such a new donor splicing site is an additional artificial donor splicing site, it has no associated target receptor splicing site. Therefore, this unpaired donor splicing site can induce random, i.e., unrestricted receptor splicing sites present somewhere in the transcribed mRNA, thereby reducing expression yield.

[0098] This invention is based, at least in part, on the unexpected discovery that the removal of unpaired donor splice sites in nucleic acids encoded by optimized antibody heavy chains only needs to be performed in the portion of the nucleic acid encoding the variable structural domain of the heavy chain, but not in the portion encoding the constant region. For the constant region, germline or wild-type human nucleic acid sequences can be used, for example.

[0099] This invention is based in at least part on the discovery that introducing an amino acid sequence silencing nucleotide change (mutation) in the unpaired donor splicing site concordant sequence NGGTA(G)AG (SEQ ID NO:01) in a codon-optimized nucleic acid encoding the variable domain of the antibody heavy chain is sufficient to increase expression yield or allow expression of the antibody.

[0100] This invention is based at least in part on the discovery that the removal of unpaired donor splice sites can be performed only in nucleic acids encoding variable domains of heavy strands rather than constant regions. That is, for constant regions, wild-type sequences, germline sequences, or sequences optimized using standard methods (e.g., as reported in WO 2013 / 15644) can be used.

[0101] This invention is based at least in part on the discovery that expression yield can be improved by using the method according to the invention with a light chain.

[0102] This invention is based on the use of the donor-shared sequence SEQ ID NO:01: NGGTA(G)AG. This sequence was identified by Zhang et al. 1998, but no use for codon optimization protocols was found.

[0103] The dinucleotide GT marks the start of an intron, thus marking the splicing site. The following base can be adenine or guanine.

[0104] Another parameter to consider is the number of mismatches allowed in the sequence, in order to adjust the sensitivity and strictness of the method.

[0105] The common receptor splice sequence has the sequence SEQ ID NO:04: [CT]n N[CT]AG, where N = any base and n = the number of CT dinucleotides. This sequence was also identified by Zhang in 1998.

[0106] The following examples of specific antibodies illustrate the method according to the invention. These examples should not be construed as limiting the invention. They are presented merely as examples of methods generally applicable to the invention.

[0107] Table 2 below summarizes the expression yields of antibody heavy and / or light chain encoded nucleic acids under different treatments. The results were obtained through transient expression in HEK293 cells using two plasmid systems.

[0108] The construct "00" is the starting nucleic acid.

[0109] The constructs “00”-“06” and “16” have been codon optimized using methods known in the art from WO 2013 / 156443, referring to prior art methods.

[0110] Constructs “07”-“11” and “17” have been processed according to the present invention, i.e., the prior art method from WO2013 / 156443, and donor splice sites have been removed based on the common motif of SEQ ID NO:01.

[0111] Constructs “12”-“15” have been codon optimized by the commercial vendor Geneart using a method different from WO 2013 / 156443 as a second reference prior art method.

[0112] The term "hu" indicates the use of the human codon, while the term "CHO" indicates the use of the Chinese hamster codon.

[0113] Table 2:

[0114]

[0115]

[0116] The data reveals the following points:

[0117] For light chains:

[0118] - When using a genome organization form, the treatment according to the invention results in a considerable expression yield.

[0119]

[0120] Using cDNA typically increases expression yield.

[0121]

[0122] When using cDNA, the treatment according to the invention results in increased expression yield in most cases, and comparable expression yield in others.

[0123]

[0124]

[0125] For heavy chains:

[0126] - When using a genome organization form, if at least the heavy chain variable domain is optimized, the treatment according to the invention results in increased expression yield or allows expression.

[0127]

[0128] Using cDNA typically increases expression yield.

[0129]

[0130]

[0131] When using cDNA, the treatment of the heavy chain constant region according to the present invention does not appear to further increase expression yield.

[0132] When using cDNA, the treatment according to the invention results in a further increase in expression yield.

[0133]

[0134] Therefore, one aspect, as reported herein, is a method for producing immunoglobulins, comprising the following steps:

[0135] - Culture mammalian cells, preferably CHO cells, containing nucleic acids in tissue form with intron-exon structures encoding immunoglobulin heavy and light chains of human IgG1 subclasses, thereby expressing immunoglobulins.

[0136] - To produce immunoglobulins by recovering immunoglobulins from cells or culture media.

[0137] In the portion of the nucleic acid encoding the variable domain of the immunoglobulin heavy chain, the unpaired donor splicing site according to SEQ ID NO:01 is removed by introducing an amino acid sequence silencing nucleotide change at codon NGG, codon GGT, or codon GTA(G) in the unpaired donor splicing site concordant sequence NGGTA(G)AG (SEQ ID NO:01).

[0138] Reference codon optimization methods

[0139] Nucleic acids encoding immunoglobulins can be optimized, for example by adjusting the use of general codons according to the method reported in WO 2013 / 156443.

[0140] The reference method is based on the discovery that, in order to express a polypeptide in a cell, a polypeptide-encoded nucleic acid is used, characterized in that each amino acid is encoded by a set of codons, wherein each codon in the set of codons is defined by a specific frequency of use within the set, the specific frequency of use within the set being related to the overall frequency of use of that codon in the cell genome, and thus the frequency of use of the codon in the (total) polypeptide-encoded nucleic acid is approximately the same as the frequency of use within its respective set.

[0141] The reference method is a method for recombinantly producing polypeptides in mammalian cells, including the steps of culturing cells containing nucleic acids encoding the polypeptide and recovering the polypeptide from the mammalian cells or culture medium.

[0142] Each amino acid residue of the polypeptide is encoded by one or more (at least one) codons, thereby combining (different) codons encoding the same amino acid residue into a group, and each codon in a group is defined by a specific frequency of use within the group, which is the frequency of a single codon in a group of codons that can be found in the nucleic acid encoding the polypeptide relative to all codons in the group, where the sum of the specific frequencies of use of all codons in the group is 100%.

[0143] The overall frequency of use of each codon in a polypeptide-encoded nucleic acid is roughly the same as its specific frequency of use within its group.

[0144] In one embodiment, amino acid residues G, A, V, L, I, P, F, S, T, N, Q, Y, C, K, R, H, D, and E are each encoded by a set of codons, and amino acid residues M and W are encoded by a single codon.

[0145] In one embodiment, amino acid residues G, A, V, L, I, P, F, S, T, N, Q, Y, C, K, R, H, D, and E are each encoded by a set of codons (containing at least two codons), and amino acid residues M and W are encoded by a single codon.

[0146] In one embodiment, if an amino acid residue happens to be encoded by a single codon, then the specific frequency of use of that codon is 100%.

[0147] In one embodiment, amino acid residue G is encoded by a set of up to four codons. In one embodiment, amino acid residue A is encoded by a set of up to four codons. In one embodiment, amino acid residue V is encoded by a set of up to four codons. In one embodiment, amino acid residue L is encoded by a set of up to six codons. In one embodiment, amino acid residue I is encoded by a set of up to three codons. In one embodiment, amino acid residue M is encoded by exactly one codon. In one embodiment, amino acid residue P is encoded by a set of up to four codons. In one embodiment, amino acid residue F is encoded by a set of up to two codons. In one embodiment, amino acid residue W is encoded by exactly one codon. In one embodiment, amino acid residue S is encoded by a set of up to six codons. In one embodiment, amino acid residue T is encoded by a set of up to four codons. In one embodiment, amino acid residue N is encoded by a set of up to two codons. In one embodiment, amino acid residue Q is encoded by a set of up to two codons. In one embodiment, amino acid residue Y is encoded by a set of up to two codons. In one embodiment, amino acid residue C is encoded by a set of up to two codons. In one embodiment, amino acid residue K is encoded by a set of up to two codons. In one embodiment, amino acid residue R is encoded by a set of up to 6 codons. In one embodiment, amino acid residue H is encoded by a set of up to 2 codons. In one embodiment, amino acid residue D is encoded by a set of up to 2 codons. In one embodiment, amino acid residue E is encoded by a set of up to 2 codons.

[0148] In one embodiment, amino acid residue G is encoded by a set of 1 to 4 codons. In one embodiment, amino acid residue A is encoded by a set of 1 to 4 codons. In one embodiment, amino acid residue V is encoded by a set of 1 to 4 codons. In one embodiment, amino acid residue L is encoded by a set of 1 to 6 codons. In one embodiment, amino acid residue I is encoded by a set of 1 to 3 codons. In one embodiment, amino acid residue M is encoded by a set of 1 codon, i.e., exactly 1 codon. In one embodiment, amino acid residue P is encoded by a set of 1 to 4 codons. In one embodiment, amino acid residue F is encoded by a set of 1 to 2 codons. In one embodiment, amino acid residue W is encoded by a set of 1 codon, i.e. exactly 1 codon. In one embodiment, amino acid residue S is encoded by a set of 1 to 6 codons. In one embodiment, amino acid residue T is encoded by a set of 1 to 4 codons. In one embodiment, amino acid residue N is encoded by a set of 1 to 2 codons. In one embodiment, amino acid residue Q is encoded by a set of 1 to 2 codons. In one embodiment, amino acid residue Y is encoded by a set of 1 to 2 codons. In one embodiment, amino acid residue C is encoded by a set of 1 to 2 codons. In one embodiment, amino acid residue K is encoded by a set of 1 to 2 codons. In one embodiment, amino acid residue R is encoded by a set of 1 to 6 codons. In one embodiment, amino acid residue H is encoded by a set of 1 to 2 codons. In one embodiment, amino acid residue D is encoded by a set of 1 to 2 codons. In one embodiment, amino acid residue E is encoded by a set of 1 to 2 codons.

[0149] In one embodiment, each of these groups contains only codons with a total usage frequency greater than 5% within the cell genome. In one embodiment, each of these groups contains only codons with a total usage frequency of 8% or greater within the cell genome. In one embodiment, each of these groups contains only codons with a total usage frequency of 10% or greater within the cell genome. In one embodiment, each of these groups contains only codons with a total usage frequency of 15% or greater within the cell genome.

[0150] In one embodiment, the codon sequence of the nucleic acid encoding a polypeptide of a specific amino acid residue in the 5' to 3' direction is, that is, the codon sequence corresponding to the codon motif of each amino acid.

[0151] In one embodiment, for each consecutive occurrence of a specific amino acid in the polypeptide starting from the N-terminus, the codon encoding the nucleic acid is the same as the codon at the corresponding consecutive position in the codon motif of the specific amino acid. Specifically, when the amino acid residue first appears in the amino acid sequence of the polypeptide, the first codon of the amino acid codon motif is used in the corresponding encoding nucleic acid, and when the amino acid residue appears for the second time, the second codon of the amino acid codon motif is used, and so on.

[0152] In one embodiment, the frequency of codon usage in an amino acid codon motif is approximately the same as its specific frequency of usage within its group.

[0153] In one embodiment, when a specific amino acid appears in a polypeptide again after reaching the final codon of the amino acid codon motif, the encoding nucleic acid contains a codon located at the first position of the amino acid codon motif.

[0154] In one embodiment, the codons in the amino acid codon motif are randomly distributed throughout the entire amino acid codon motif.

[0155] In one embodiment, the amino acid codon motif is selected from a set of amino acid codon motifs that includes all possible amino acid codon motifs that can be obtained by swapping the codons therein, wherein all motifs have the same number of codons and the codons in each motif have the same specific frequency of use.

[0156] In one embodiment, the codons in the amino acid codon motif are arranged in descending order of a specific frequency of use, whereby all codons with the same frequency of use directly replace each other. In another embodiment, codons with the same frequency of use are grouped together.

[0157] In one embodiment, the (different) codons in the amino acid codon motif are evenly distributed throughout the entire amino acid codon motif.

[0158] In one embodiment, the codons in the amino acid codon motif are arranged in descending order of a specific frequency of use, such that a codon with the highest specific frequency of use is present (used) after a codon with the lowest specific frequency of use or a codon with the second lowest specific frequency of use.

[0159] In one embodiment, the codons in the amino acid codon motif are arranged in descending order of a specific frequency of use, such that a codon with the highest specific frequency of use is present (used) after the codon with the lowest specific frequency of use.

[0160] Therefore, a characteristic of nucleic acids encoding polypeptides is that each amino acid residue of the polypeptide is encoded by one or more (at least one) codons.

[0161] Thus, different codons encoding the same amino acid residue are grouped into a set, and each codon in a set is defined by a specific frequency of use within the set. This frequency is the frequency of a single codon in a set of codons relative to all codons in the set, which can be found in nucleic acids encoding polypeptides, where the sum of the specific frequencies of use of all codons in the set is 100%.

[0162] The frequency of codon usage in the polypeptide-encoded nucleic acid is roughly the same as its specific usage frequency within its group.

[0163] In one embodiment, amino acid residues G, A, V, L, I, P, F, S, T, N, Q, Y, C, K, R, H, D, and E are each encoded by a set of codons, and amino acid residues M and W are encoded by a single codon.

[0164] In one embodiment, amino acid residues G, A, V, L, I, P, F, S, T, N, Q, Y, C, K, R, H, D, and E are each encoded by a set of codons (containing at least two codons), and amino acid residues M and W are encoded by a single codon.

[0165] In one embodiment, if an amino acid residue happens to be encoded by a single codon, then the specific frequency of use of that codon is 100%.

[0166] In one embodiment, amino acid residue G is encoded by a set of up to four codons. In one embodiment, amino acid residue A is encoded by a set of up to four codons. In one embodiment, amino acid residue V is encoded by a set of up to four codons. In one embodiment, amino acid residue L is encoded by a set of up to six codons. In one embodiment, amino acid residue I is encoded by a set of up to three codons. In one embodiment, amino acid residue M is encoded by exactly one codon. In one embodiment, amino acid residue P is encoded by a set of up to four codons. In one embodiment, amino acid residue F is encoded by a set of up to two codons. In one embodiment, amino acid residue W is encoded by exactly one codon. In one embodiment, amino acid residue S is encoded by a set of up to six codons. In one embodiment, amino acid residue T is encoded by a set of up to four codons. In one embodiment, amino acid residue N is encoded by a set of up to two codons. In one embodiment, amino acid residue Q is encoded by a set of up to two codons. In one embodiment, amino acid residue Y is encoded by a set of up to two codons. In one embodiment, amino acid residue C is encoded by a set of up to two codons. In one embodiment, amino acid residue K is encoded by a set of up to two codons. In one embodiment, amino acid residue R is encoded by a set of up to 6 codons. In one embodiment, amino acid residue H is encoded by a set of up to 2 codons. In one embodiment, amino acid residue D is encoded by a set of up to 2 codons. In one embodiment, amino acid residue E is encoded by a set of up to 2 codons.

[0167] In one embodiment, amino acid residue G is encoded by a set of 1 to 4 codons. In one embodiment, amino acid residue A is encoded by a set of 1 to 4 codons. In one embodiment, amino acid residue V is encoded by a set of 1 to 4 codons. In one embodiment, amino acid residue L is encoded by a set of 1 to 6 codons. In one embodiment, amino acid residue I is encoded by a set of 1 to 3 codons. In one embodiment, amino acid residue M is encoded by a set of 1 codon. In one embodiment, amino acid residue P is encoded by a set of 1 to 4 codons. In one embodiment, amino acid residue F is encoded by a set of 1 to 2 codons. In one embodiment, amino acid residue W is encoded by a set of 1 codon. In one embodiment, amino acid residue S is encoded by a set of 1 to 6 codons. In one embodiment, amino acid residue T is encoded by a set of 1 to 4 codons. In one embodiment, amino acid residue N is encoded by a set of 1 to 2 codons. In one embodiment, amino acid residue Q is encoded by a set of 1 to 2 codons. In one embodiment, amino acid residue Y is encoded by a set of 1 to 2 codons. In one embodiment, amino acid residue C is encoded by a set of 1 to 2 codons. In one embodiment, amino acid residue K is encoded by a set of 1 to 2 codons. In one embodiment, amino acid residue R is encoded by a set of 1 to 6 codons. In one embodiment, amino acid residue H is encoded by a set of 1 to 2 codons. In one embodiment, amino acid residue D is encoded by a set of 1 to 2 codons. In one embodiment, amino acid residue E is encoded by a set of 1 to 2 codons.

[0168] In one embodiment, each of these groups contains only codons with a total usage frequency greater than 5% within the cell genome. In one embodiment, each of these groups contains only codons with a total usage frequency of 8% or greater within the cell genome. In one embodiment, each of these groups contains only codons with a total usage frequency of 10% or greater within the cell genome. In one embodiment, each of these groups contains only codons with a total usage frequency of 15% or greater within the cell genome.

[0169] In one embodiment, the codon sequence of the nucleic acid encoding a polypeptide of a specific amino acid residue in the 5' to 3' direction is, that is, the codon sequence corresponding to the codon motif of each amino acid.

[0170] In one embodiment, for each consecutive occurrence of a specific amino acid in the polypeptide starting from the N-terminus, the codon encoding the nucleic acid is the same as the codon at the corresponding consecutive position in the codon motif of the specific amino acid. Specifically, when the amino acid residue first appears in the amino acid sequence of the polypeptide, the first codon of the amino acid codon motif is used in the corresponding encoding nucleic acid, and when the amino acid residue appears for the second time, the second codon of the amino acid codon motif is used, and so on.

[0171] In one embodiment, the frequency of codon usage in an amino acid codon motif is approximately the same as its specific frequency of usage within its group.

[0172] In one embodiment, when a specific amino acid appears in a polypeptide again after reaching the final codon of the amino acid codon motif, the encoding nucleic acid contains a codon located at the first position of the amino acid codon motif.

[0173] In one embodiment, each codon in the amino acid codon motif is randomly distributed throughout the amino acid codon motif.

[0174] In one embodiment, each codon in the amino acid codon motif is uniformly distributed throughout the entire amino acid codon motif.

[0175] In one embodiment, the codons in the amino acid codon motif are arranged in descending order of a specific frequency of use, such that a codon with the highest specific frequency of use is used after a codon with the lowest specific frequency of use or a codon with the second lowest specific frequency of use.

[0176] In one embodiment, the codons in the amino acid codon motif are arranged in descending order of a specific frequency of use, such that a codon with the highest specific frequency of use is used after the codon with the lowest specific frequency of use.

[0177] Therefore, the method according to the present invention is a method for increasing the expression of polypeptides in eukaryotic cells, comprising the following steps:

[0178] - Provides nucleic acids encoding polypeptides,

[0179] Each amino acid residue of the polypeptide is encoded by at least one codon, thereby grouping different codons encoding the same amino acid residue into a set, and each codon in a set is defined by a specific frequency of use within that set, wherein the sum of the specific frequencies of use of all codons in a set is 100%.

[0180] The frequency of codon usage in the polypeptide-encoded nucleic acid is roughly the same as its specific usage frequency within its group.

[0181] This involves removing the donor splicing site of the common sequence according to SEQ ID NO: 01 from the nucleic acid encoding the variable structural domain of the heavy chain.

[0182] Recombination method

[0183] Antibodies can be generated using recombinant methods and compositions, for example, as described in US 4,816,567. The antibody-encoding nucleic acid may encode an amino acid sequence constituting the VL of the antibody and / or an amino acid sequence constituting the VH of the antibody (e.g., the light chain and / or heavy chain of the antibody). In one embodiment, a cell expressing a polypeptide containing an immunoglobulin constant region has been transfected with one or more vectors containing such nucleic acid (e.g., expression vectors). In one embodiment, a cell is provided containing such nucleic acid modified using the methods reported herein. In one embodiment, the cell comprises (e.g., having been transformed with): (1) a vector containing nucleic acid encoding an amino acid sequence containing the VL of the antibody and an amino acid sequence containing the VH of the antibody; or (2) a first vector and a second vector, the first vector containing nucleic acid encoding an amino acid sequence containing the VL of the antibody, and the second vector containing nucleic acid encoding an amino acid sequence containing the VH of the antibody. In one embodiment, the cell is a eukaryotic cell, such as Chinese hamster ovary (CHO) cells or lymphocytes (e.g., Y0, NSO, Sp2 / O cells). In one embodiment, a method for preparing an antibody is provided, wherein the method includes culturing cells comprising nucleic acids encoding the antibody provided herein under conditions suitable for antibody expression, and optionally recovering the antibody from the cells (or culture medium).

[0184] For antibody recombinant production, nucleic acids encoding antibodies (e.g., as described herein) are generated and inserted into one or more vectors for further cloning and / or expression in cells. Such nucleic acids can be readily isolated and sequenced using routine procedures (e.g., by using oligonucleotide probes capable of specifically binding to genes encoding the heavy and light chains of antibodies).

[0185] Suitable cells for cloning or expressing vectors encoding antibodies include eukaryotic cells as described herein.

[0186] In addition, eukaryotic microorganisms such as filamentous fungi or yeast are also suitable cloning or expression hosts for antibody-encoding vectors. These eukaryotic microorganisms include fungal and yeast strains whose glycosylation pathways have been “humanized”, resulting in the production of antibodies with partial or complete human glycosylation patterns (see Gerngross, TU, Nat. Biotech. 22 (2004) 1409-1414 and Li, H., et al., Nat. Biotech. 24 (2006) 210-215).

[0187] Suitable host cells for expressing glycosylated antibodies also originate from multicellular organisms (invertebrates and vertebrates). Examples of invertebrate cells include plant cells and insect cells. Many baculovirus strains have been identified that can be used in conjunction with insect cells, particularly for transfecting Spodoptera frugiperda cells.

[0188] Plant cell cultures can also be used as hosts (see, for example, US 5,959,177, US 6,040,498, US 6,420,548, US 7,125,978 and US 6,417,429, which describe PLATNIBODIES for producing antibodies in transgenic plants). TM technology).

[0189] Vertebrate cells can also be used as hosts. For example, mammalian cell lines adapted for growth in suspension may be useful. Other examples of useful mammalian cell lines include the monkey kidney CV1 line (COS-7) transformed with SV40; human embryonic kidney cell lines (such as HEK293 cells described, for example, in Graham, FL et al., J. Gen Virol. 36 (1977) 59-74); hamster kidney cells (BHK); mouse Sertoli cells (such as TM4 cells described, for example, in Mather, JP, Biol. Reprod. 23 (1980) 243-252); monkey kidney cells (CV1); African green monkey kidney cells (VERO-76); human cervical cancer cells (HELA); canine kidney cells (MDCK); Buffalo rat hepatocytes (BRL 3A); human lung cells (W138); human hepatocytes (HepG2); and mouse mammary tumors (MMT). 060562); TRI cells, as described, for example, in Mather, JP et al., Annals N.Y. Acad. Sci. 383 (1982) 44-68; MRC 5 cells; and FS4 cells. Other useful mammalian cell lines include Chinese hamster ovary (CHO) cells, including DHFR - CHO cells (Urlaub, G. et al., Proc. Natl. Acad. Sci. USA 77 (1980) 4216-4220); and myeloma cell lines such as Y0, NSO, and Sp2 / 0. For reviews of certain mammalian host cell lines suitable for antibody production, see, for example, Yazaki, P. and Wu, AM, Methods in Molecular Biology, Vol. 248, Lo, BKC (ed.), Humana Press, Totowa, NJ (2004), pp. 255-268.

[0190] purification

[0191] Various methods for protein recovery and purification are well-established and widely used, such as microbial protein affinity chromatography (e.g., protein A or protein G affinity chromatography), ion exchange chromatography (e.g., cation exchange (carboxymethyl resin), anion exchange (aminoethyl resin), and mixed-mode exchange), thiophilic adsorption (e.g., with β-mercaptoethanol and other SH ligands), hydrophobic interaction or aromatic adsorption chromatography (e.g., using phenyl-agarose, azazo-aryl resins, or m-aminophenylboronic acid), metal chelate affinity chromatography (e.g., using Ni(II)-affinity materials and Cu(II)-affinity materials), size exclusion chromatography, and electrophoresis methods (e.g., gel electrophoresis, capillary electrophoresis) (Vijayalakshmi, MA Appl. Biochem. Biotech. 75 (1998) 93-102).

[0192] Codon usage

[0193] Codon usage tables (see the example above) are readily available, for example, in the “Codon Usage Database” provided at http: / / www.kazusa.or.jp / codon / , and these tables can be modified in many ways (Nakamura, Y., et al., Nucl. Acids Res. 28 (2000) 292).

[0194] Encoding nucleic acids play a crucial role in the high-yield expression of recombinant peptides. Naturally occurring and naturally isolated encoding nucleic acids are often not optimized for high-yield expression, especially in heterologous host cells. Due to the degeneracy of the genetic code, an amino acid residue can be encoded by more than one nucleotide triplet (codon) besides tryptophan and methionine. Therefore, for a given amino acid sequence, different encoding codons (corresponding encoding nucleic acid sequences) are possible.

[0195] Different organisms use different codons that encode a single amino acid residue at different relative frequencies (codon usage). Generally, a particular codon is used more frequently than other possible codons.

[0196] In WO 2001 / 088141, an optimization of reading frames based on the use of codons found in highly expressed mammalian genes was reported. For this purpose, the matrix was generated considering almost only the most frequently used codons in highly expressed mammalian genes, with the second most frequently used codons being used secondarily, as shown in the table below. Using these codons from highly expressed human genes, a fully synthetic reading frame not found in nature was created, but it encodes a product identical to the original wild-type gene construct.

[0197] Different methods for codon optimization using the frequency of use of a single codon are reported in US 8,128,938, such as uniform optimization, full optimization, and minimum optimization.

[0198] The table below shows the most frequently used codon (codon 1) and the second most frequently used codon (codon 2) found in highly expressed mammalian genes.

[0199] surface.

[0200] amino acids codon 1 codon 2 Ala GCC GCT Arg AGG AGA Asn AAC AAT Asp GAC GAT Cys TGC TGT End TGA TAA Gln CAG CAT Glu GAG GAA Gly GGC GGA His CAC CAT Ile ATC ATT Leu CTG CTC Lys AAG AAT Met ATG ATG Phe TTC TTT Pro CCC CCT Ser AGC TCC Thr ACC ACA Trp TGG TGG Tyr TAC TAT Val GTG GTC

[0201] (Ausubel, FM, et al., Current Protocols in Molecular Biology 2 (1994), A1.8-A1.9).

[0202] There are few deviations from the strict adherence to the use of the most common codons: (i) to adjust the introduction or removal of unique restriction sites, and (ii) to break G or C segments that extend more than 7 base pairs to allow for sequential PCR amplification and sequencing of synthetic gene products.

[0203] The following examples are provided to aid in understanding the invention, the true scope of which is set forth in the appended claims. It should be understood that modifications may be made to the illustrated procedures without departing from the spirit of the invention.

[0204] literature

[0205]

[10] M.Graf, L.Deml, and R.Wagner, "Codon-optimized genes that enable increased heterologous expression in mammalian cells and elicit efficientimmune responses in mice after vaccination of naked DNA.," Methods Mol. Med., vol.94, pp.197–210, 2004.

[0206]

[11] WO 2013 / 156443.

[0207]

[12] MQZhang, "Statistical features of human exons and their flanking regions," vol.7, no.5, pp.919–932, 1998.

[0208]

[13] R.Breathnach,C.Benoist,K.O’Hare,F.Gannon,and P.Chambon,“Ovalbumingene:evidence for a leader sequence in mRNA and DNA sequences at the exon-intron boundaries.,”Proc.Natl.Acad.Sci.U.S.A.,vol.75,no.10,pp.4853–4857,1978.

[0209]

[14] M.Burset,I.A.Seledtsov,and V.V Solovyev,“Analysis of canonicaland non-canonical splice sites in mammalian genomes.,”Nucleic Acids Res.,vol.28,no.21,pp.4364–4375,2000.

[0210]

[15] M.B.Shapiro and P.Senapathy,“RNA splice junctions of differentclasses of eukaryotes:Sequence statistics and functional implications in geneexpression,”Nucleic Acids Res.,vol.15,no.17,pp.7155–7174,1987.

[0211]

[16] T.W.Nilsen,“Twenty years of RNA:then and now,”pp.471–473,2015.

[0212]

[17] R.A.Padgetr,P.J.Grabowski,M.M.Konarska,S.Seiler,and P.A.Sharp,“Splicing of messenger RNA precursors,”1986.

[0213]

[18] M.R.Green,“Pre-mRNA splicing,”Annu.Rev.Genet.,vol.20,pp.671–708,1986.

[0214]

[22] Y.Kapustin, E.Chan, R.Sarkar, F.Wong, I.Vorechovsky, RMWinston, T.Tatusova, and NJDibb, "Cryptic splice sites and split genes," Nucleic AcidsRes., vol.39, no.14, pp.5837–5844, 2011.

[0215] Example

[0216] Protein assay:

[0217] Protein concentration was determined by measuring the optical density (OD) at 280 nm using the molar extinction coefficient calculated based on the amino acid sequence.

[0218] Recombinant DNA technology

[0219] Use standard methods to manipulate DNA, as described in Sambrook, J., et al., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, 1989. Use molecular biology reagents according to the manufacturer's instructions.

[0220] Example 1

[0221] Antibody expression and purification of nucleic acids optimized based on different codons

[0222] Different unoptimized or codon-optimized variable domains encode nucleic acids with wild-type human constant regions or with nucleic acids that use the CHO codon to encode human constant regions.

[0223] Expression plasmid:

[0224] Each expression plasmid contains an expression cassette for expressing either the heavy or light chain. These are assembled separately into mammalian cell expression vectors.

[0225] General information about the nucleotide sequences from which the codons used by humans can be inferred is given in: Kabat, E.A. et al., Sequences of Proteins of Immunological Interest, 5th edition, Public Health Service, National Institutes of Health, Bethesda, MD (1991), NIH Publication No. 91-3242.

[0226] In addition to light or heavy chain expression cassettes, these plasmids also contain

[0227] - Hygromycin resistance gene,

[0228] -Epstein-Barr virus (EBV) origin of replication, oriP

[0229] - The origin of replication from the pUC18 vector, which allows the plasmid to replicate in E. coli, and

[0230] The β-lactamase gene confers ampicillin resistance in Escherichia coli.

[0231] Recombinant DNA technology:

[0232] Cloning was performed using standard cloning techniques, as described, for example, in Sambrook, J., et al., Molecular Cloning: A Laboratory Manual, second edition, Cold Spring Harbor Laboratory Press (1989). All molecular biology reagents were commercially available (unless otherwise stated) and used according to the manufacturer's instructions.

[0233] DNA and protein sequence analysis and sequence data management

[0234] Vector NTI Advance Suite version 9.0 is used for sequence creation, mapping, analysis, annotation, and diagramming.

[0235] Antibody expression:

[0236] For antibody expression, human embryonic kidney cells (HEK) 293F were used. HEK293 cells are primarily used for transient gene expression. The desired protein can be harvested after a few days. HEK293F cells were cultured in serum-free and protein-free free-style culture medium supplemented with penicillin and streptomycin (PenStrep). TM 293 Expression Medium (Gibco, Invitrogen) TM(Life Technologies), in shake flasks at 7% CO2, 85% humidity and 37°C. Cells were passaged every 3–4 days and divided into trays according to confluence. Cell counts were always set to 3 × 10⁻⁶. 5 Cells / mL. Cell count and viability were measured using a CASEY cell counter (Roche). 50 μl of culture was suspended in 10 ml of Casyton and measured using the appropriate procedure.

[0237] For transient transfection, the daily cell count should be adjusted to 2 × 10⁻⁶. 6 Cells / mL. The transient transfection number should not be less than 4 passages or more than 22 passages. For transfection, use the transfection reagent PEIpro and run the batch feed for 7 days. For transfection, the following volumes and DNA amounts are used for the transfection mixture (culture volume 20 mL):

[0238]

[0239] For transfection, mix appropriate amounts of F17, DNA, and transfection reagent in their respective orders, and incubate for 10 minutes. Then, add the transfection mixture to the cells. After approximately 3-5 hours, add the appropriate pre-diluted VPA solution. After 16 hours, supplement the culture with 0.6% glucose and 12% feed and incubate further.

[0240]

[0241]

[0242] One week later, the supernatant can be harvested. To do this, the mixture is centrifuged at 1200 rpm for 20 minutes and the supernatant is aseptically filtered out using a 0.22 μm filter.

[0243] Expression yield determination:

[0244] The antibody expressed in the supernatant was quantified using an HPLC column packed with protein A. Protein A is a cell wall-associated protein from Staphylococcus aureus that specifically binds to the constant region of IgG antibodies. Protein A was immobilized on a polymer carrier and then bound to the target molecule. The antibody could be eluted and detected by varying various parameters, such as pH and temperature.

[0245] Seven days after transfection, HEK 293 cell supernatant was harvested. Protein A-Sepharose was used. TMAffinity chromatography (GE Healthcare, Sweden) purifies recombinant antibodies contained in the supernatant from the culture supernatant using affinity chromatography. In short, the antibody containing the clarified culture supernatant is placed on a MabSelect SuRe Protein A column (5-50 ml) equilibrated with PBS buffer (10 mM Na₂HPO₄, 1 mM KH₂PO₄, 137 mM NaCl, and 2.7 mM KCl, pH 7.4). Unbound proteins are washed away with equilibration buffer. The antibody (or derivative) is eluted with 50 mM citrate buffer at pH 3.2. The protein fraction is neutralized with 0.1 ml of 2 M Tris buffer at pH 9.0.

Claims

1. A method for producing a human IgGl subclass antibody by culturing CHO cells that have been transfected with one or more expression cassettes comprising nucleic acids encoding a heavy chain and a light chain of an antibody, wherein the nucleic acids encoding the antibody heavy chain and the antibody light chain are codon usage optimized for the codon usage of human cells and / or the codon usage of CHO cells, wherein in the part of the nucleic acids encoding the heavy chain variable domain, a non- paired donor splice site is removed by introducing an amino acid silent mutation in the nucleotide sequence NGGTA(G)AG as set forth in SEQ ID NO: 01, and wherein one or more nucleic acids encoding the antibody heavy chain is DNA in genomic organization.

2. The method according to claim 1, wherein in the part of the nucleic acids encoding the heavy chain constant region, a non-paired donor splice site is not removed.

3. The method according to any one of claims 1 to 2, wherein in the nucleic acid encoding the antibody light chain, a non-paired splice site is removed.

4. The method according to any one of claims 1 to 2, wherein the transfection is a transient transfection.

5. The method according to claim 3, wherein the transfection is a transient transfection.

6. The method according to any one of claims 1, 2, 5, wherein the nucleic acid encoding the light chain of the antibody is DNA in genomic organization.

7. The method according to claim 3, wherein the nucleic acid encoding the light chain of the antibody is DNA in genomic organization.

8. The method according to claim 4, wherein the nucleic acid encoding the light chain of the antibody is DNA in genomic organization.

9. The method according to any one of claims 1, 2, 5, 7, 8, wherein removing the non- paired splice site is by introducing an amino acid silent mutation in the nucleotide sequence NGGTA(G)AG as set forth in SEQ ID NO: 01 at the codon NGG or the codon GGT or the codon GTA(G).

10. The method according to claim 3, wherein removing the non-paired splice site is by introducing an amino acid silent mutation in the nucleotide sequence NGGTA(G)AG as set forth in SEQ ID NO: 01 at the codon NGG or the codon GGT or the codon GTA(G).

11. The method according to claim 4, wherein removing the non-paired splice site is by introducing an amino acid silent mutation in the nucleotide sequence NGGTA(G)AG as set forth in SEQ ID NO: 01 at the codon NGG or the codon GGT or the codon GTA(G).

12. The method according to claim 6, wherein removing the non-paired splice site is by introducing an amino acid silent mutation in the nucleotide sequence NGGTA(G)AG as set forth in SEQ ID NO: 01 at the codon NGG or the codon GGT or the codon GTA(G).

13. The method according to any one of claims 1, 2, 5, 7, 8, 10, 11, 12, wherein the method comprises the following steps: a) culturing the CHO cells, and b) isolating the antibody from the culture medium. b) recovering the antibody from the CHO cell or culture medium.

14. The method of claim 3, wherein the method comprises the steps of: a) culturing the CHO cell, and b) recovering the antibody from the CHO cell or culture medium.

15. The method of claim 4, wherein the method comprises the steps of: a) culturing the CHO cell, and b) recovering the antibody from the CHO cell or culture medium.

16. The method of claim 6, wherein the method comprises the steps of: a) culturing the CHO cell, and b) recovering the antibody from the CHO cell or culture medium.

17. The method of claim 9, wherein the method comprises the steps of: a) culturing the CHO cell, and b) recovering the antibody from the CHO cell or culture medium.

18. A method for producing a human IgGl subclass antibody by culturing a CHO cell comprising one or more nucleic acids encoding a heavy chain and a light chain of an antibody, wherein in the nucleic acid encoding the variable domain of the heavy chain, the unpaired donor splice site is removed by introducing an amino acid sequence silent nucleotide change at codon NGG or codon GGT or codon GTA(G) in the unpaired donor splice site consensus sequence NGGTA(G)AG as set forth in SEQ ID NO: 01, wherein in the nucleic acid encoding the constant region of the heavy chain, the unpaired donor splice site comprising the nucleotide sequence NGGTA(G)AG as set forth in SEQ ID NO: 01 is not removed, wherein the one or more nucleic acids encoding the heavy chain of the antibody are DNA in genomic organization.

19. The method of claim 18, wherein the one or more nucleic acids encoding the light chain of the antibody are DNA in genomic organization.

20. The method of any one of claims 18 to 19, wherein the one or more nucleic acids encoding the heavy chain of the antibody and the light chain of the antibody are codon usage optimized for codon usage of human cells and / or codon usage of CHO cells.

21. The method of any one of claims 18 to 19, wherein in the nucleic acid encoding the light chain of the antibody, the unpaired splice site is removed by introducing an amino acid sequence silent nucleotide change at codon NGG or codon GGT or codon GTA(G) in the unpaired donor splice site consensus sequence NGGTA(G)AG as set forth in SEQ ID NO:

01.

22. The method of claim 20, wherein in the nucleic acid encoding the light chain of the antibody, the unpaired splice site is removed by introducing an amino acid sequence silent nucleotide change at codon NGG or codon GGT or codon GTA(G) in the unpaired donor splice site consensus sequence NGGTA(G)AG as set forth in SEQ ID NO:

01.

23. Use of removing an unpaired donor splice site in a part of a nucleic acid sequence encoding an antibody which is optimized for human or hamster codon usage, wherein removing the unpaired donor splice site is by introducing an amino acid silent mutation in the nucleotide sequence NGGTA(G)AG as shown in SEQ ID NO: 01, wherein the part is a part encoding a heavy chain variable domain and wherein one or more nucleic acids encoding a heavy chain of the antibody is DNA in a genomic organization, for reducing mis-splicing or / and increasing antibody expression yield when the nucleic acid is used for producing the antibody in CHO cells.

24. Use according to claim 23, wherein an unpaired splice site is not removed in a part of the nucleic acid encoding a heavy chain constant region.

25. Use according to any one of claims 23 to 24, wherein an unpaired splice site is removed in a part of the nucleic acid encoding a light chain.

26. Use according to any one of claims 23 to 24, wherein removing the unpaired splice site is by introducing an amino acid silent mutation in the nucleotide sequence NGGTA(G)AG as shown in SEQ ID NO: 01 at the codon NGG or at the codon GGT or at the codon GTA(G).

27. Use according to claim 25, wherein removing the unpaired splice site is by introducing an amino acid silent mutation in the nucleotide sequence NGGTA(G)AG as shown in SEQ ID NO: 01 at the codon NGG or at the codon GGT or at the codon GTA(G).

Citation Information

Patent Citations

  • Altered antibodies

    EP0307434A1

  • Recombinant saccharomyces

    EP0362179A2

  • Recombinant immunoglobin preparations

    US4816567A

  • DNA constructs containing a Kluyveromyces alpha factor leader sequence for directing secretion of heterologous polypeptides

    US5010182A

  • Codon pair utilization

    US5082767A