Polypeptide expression system

The polypeptide expression system addresses the inefficiency of traditional systems by using a modular design with linked nucleic acid molecules, enhancing the expression and production of recombinant polypeptides while minimizing construct redundancy.

JP7863949B2Active Publication Date: 2026-05-22GENENTECH INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
GENENTECH INC
Filing Date
2020-03-10
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing recombinant polypeptide expression systems require multiple constructs for each combination of modules, leading to a geometric increase in the number of constructs, which is inefficient and resource-intensive, especially for high-throughput systems.

Method used

A polypeptide expression system comprising a first and second nucleic acid molecule with specific components operably linked in a 5' to 3' direction, including eukaryotic and prokaryotic promoters, splice sites, hybridization sequences, and polyadenylation sites, allowing for modular expression and production of recombinant polypeptides.

Benefits of technology

Enables efficient and modular expression of recombinant polypeptides, reducing the number of constructs needed and optimizing resource utilization in high-throughput systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007863949000002
    Figure 0007863949000002
  • Figure 0007863949000003
    Figure 0007863949000003
  • Figure 0007863949000004
    Figure 0007863949000004
Patent Text Reader

Abstract

To provide a polypeptide expression system and an application method of the same.SOLUTION: A polypeptide expression system comprises a first nucleic acid molecule and a second nucleic acid molecule. The first nucleic acid molecule contains a first expression cassette including elements: a first eukaryotic promoter (P1Euk1); a first polypeptide code array (PES11); a first 5' splice site (5'ss11); and hybridized array (HS1), wherein the elements are operatively connected to each other as P1Euk1-PES11-5'ss11-HS1 in 5'-3'direction. The second nucleic acid molecule contains elements: a eukaryotic promoter (P2Euk); an array capable of hybridizing with HS2; a 3' splice site (3'ss2); a polypeptide code array (PES2); and a polyadenylated site (pA2), wherein the elements are operatively connected to each other as P2Euk-HS2-3'ss2-PES2-pA2 in the 5'-3'direction.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in ASCII format and is hereby incorporated by reference in its entirety. The name of this ASCII copy, created on June 25, 2015, is P05833-WO_SL.txt, and the size is 24,421 bytes.

[0002] The present invention relates to a polypeptide expression system for modular expression and production of polypeptides.

Background Art

[0003] Recombinant polypeptides are sometimes expressed as fusions of individual domains or tags for functional or purification purposes. Recombinant DNA methods have traditionally been used to join sequences encoding each module, requiring different constructs for each combination. This poses a problem for techniques related to the expression of many protein collections consisting of repetitive modules joined in various combinations, as the number of constructs increases geometrically as a function of the number of modules used.

[0004] High-throughput systems for subcloning can handle multiple inserts in parallel, but they are usually source-intensive and generate a large number of constructs that are ultimately not needed after the initial characterization steps. Thus, there remains a need in the field of development of polypeptide expression systems that enable modular expression and production of recombinant polypeptides that has not yet been addressed.

Summary of the Invention

[0005] The present invention relates to a polypeptide expression system for modular expression and production of polypeptides.

[0006] In one aspect, the present invention features a polypeptide expression system comprising a first nucleic acid molecule and a second nucleic acid molecule, wherein (a) the first nucleic acid molecule comprises the following components: (i) a first eukaryotic promoter (P1 Euk1 ), (ii) a first polypeptide coding sequence (PES11), (iii) a first 5' splice site (5'ss11), and (iv) a hybridization sequence (HS1), and these components are operably linked to each other in the 5' to 3' direction as P1 Euk1 -PES11-5'ss11-HS1, (b) the second nucleic acid molecule comprises the following components: (i) a eukaryotic promoter (P2 Euk ), (ii) a hybridization sequence (HS2) capable of hybridizing to HS1, (iii) a 3' splice site (3'ss2), (iv) a polypeptide coding sequence (PES2), and (v) a polyadenylation site (pA2), and these components are operably linked to each other in the 5' to 3' direction as P2 Euk -HS2-3'ss2-PES2-pA2. In some embodiments, P1 Euk1 is a cytomegalovirus (CMV) promoter or a simian virus 40 (SV40) promoter. In some embodiments, P2 Euk is a CMV promoter or an SV40 promoter. In some embodiments, the first expression cassette further comprises a first nucleic acid sequence encoding a eukaryotic signal sequence (ESS11), and ESS11 is positioned between P1 Euk1 and PES11. In some embodiments, ESS11 is derived from a variable heavy chain (VH) gene.

[0007] In some embodiments, the first expression cassette further comprises an excisable prokaryotic promoter module (ePPM1) comprising the following components: (i) a 5' splice site (5'ss12), (ii) a prokaryotic promoter (P1 Prok1 ), and (iii) a 3' splice site (3'ss11), wherein these components are 5'ss12-P1 Prok1-3'ss11 are operably connected to each other in the 5'~3' direction, and ePPM1 is P1 Euk1 It is positioned between and PES11. In some embodiments, P1 Prok1 The promoter is selected from the group consisting of the PhoA promoter, the Tac promoter, the Lac promoter, and the Tphac promoter. In some embodiments, ePPM1 further comprises a first nucleic acid sequence encoding a prokaryotic signal sequence (PSS11). In some embodiments, PSS11 is derived from the thermostable enterotoxin II (stII) gene. In some embodiments, the polypeptide expression system further comprises a polypyrimidine region (PPT11) located between PSS11 and 3'ss11. In some embodiments, PPT11 comprises the nucleic acid sequence TTCCTTTTTTCTCTTTCC (SEQ ID NO: 1). In some embodiments, PES11 does not contain a latent 5' splice site. In some embodiments, HS1 is a gene encoding all or part of a coat protein or adapter protein. In some embodiments, the coat protein is selected from the group consisting of pI, pII, pIII, pIV, pV, pVI, pVII, pVIII, pIX, and pX of bacteriophage M13, f1, or fd. In some embodiments, the coat protein is the pIII protein of bacteriophage M13. In some embodiments, the pIII fragment comprises amino acid residues 267-421 or 262-418 of the pIII protein. In some embodiments, the adapter protein is a leucine zipper. In some embodiments, the leucine zipper comprises the amino acid sequence of SEQ ID NO: 4 or 5.

[0008] In some embodiments, the first nucleic acid molecule is a second eukaryotic promoter (P1 Euk2 ), (ii) a second polypeptide coding sequence (PES12), and (iii) a polyadenylation site (pA1), further comprising a second expression cassette, wherein these components are P1 Euk2-PES12-pA1 are operably connected to each other in the 5'~3' direction. Euk2 This is a CMV promoter or an SV40 promoter. In some embodiments, the second expression cassette further comprises a second nucleic acid sequence encoding a eukaryotic signal sequence (ESS12). In some embodiments, ESS12 is derived from a mouse-conjugated immunoglobulin protein (mBiP) gene. In some embodiments, ESS12 comprises the nucleic acid sequence ATG AAN TTN ACN GTN GTN GCN GCN GCN CTN CTN CTN CTN GGN (SEQ ID NO: 6), where N is A, T, C, or G.

[0009] In some embodiments, the second expression cassette comprises the following components: (i) 5' splice site (5'ss13), (ii) prokaryotic promoter (P1 Prok2 (iii) a resectable prokaryotic promoter module (ePPM2) comprising (iii) a 3' splice site (3'ss12), wherein these components are 5'ss13-P1 Prok2 -3'ss12 are operably connected to each other in the 5'~3' direction, and ePPM2 is P1 Euk2 It is positioned between and PES12. In some embodiments, P1 Prok2The promoter is selected from the group consisting of the PhoA promoter, the Tac promoter, and the Lac promoter. In some embodiments, ePPM2 further comprises a nucleic acid sequence encoding a prokaryotic signal sequence (PSS12). In some embodiments, PSS12 is derived from the thermostable enterotoxin II (stII) gene. In some embodiments, the polypeptide expression system further comprises a polypyrimidine region (PPT12) located between PSS12 and 3'ss12. In some embodiments, PPT12 comprises the nucleic acid sequence TTCCTTTTTTCTCTTTCC (SEQ ID NO: 1). In some embodiments, the second expression cassette is located 5' relative to the first expression cassette. In some embodiments, the polypeptide expression system further comprises an intron splice enhancer (ISE) (ISE1) located between 5'ss11 and HS1. In some embodiments, ISE1 comprises a G-run comprising three or more consecutive guanine residues. In some embodiments, ISE1 comprises a G-run comprising nine consecutive guanine residues. In some embodiments, the polypeptide expression system further includes a polypyrimidine region (PPT2) located between HS2 and 3'ss2. In some embodiments, PPT2 includes the nucleic acid sequence TTCCTCTTTCCCTTTCTCTCC (SEQ ID NO: 7). In some embodiments, the polypeptide expression system further includes an ISE (ISE2) located between HS2 and 3'ss2. In some embodiments, ISE2 includes a G-run comprising three or more consecutive guanine residues. In some embodiments, ISE2 includes a G-run comprising nine consecutive guanine residues. In some embodiments, 5'ss11 includes the nucleic acid sequence GTAAGA (SEQ ID NO: 8).

[0010] In some embodiments, expression mediated by a eukaryotic promoter occurs in mammalian cells. In some embodiments, the mammalian cells are Expi293F cells, CHO cells, 293T cells, or NSO cells. In some embodiments, the mammalian cells are Expi293F cells. In some embodiments, expression mediated by a prokaryotic promoter occurs in bacterial cells. In some embodiments, the bacterial cells are E. coli cells. In some embodiments, PES11 encodes all or part of the antibody. In some embodiments, PES11 encodes a polypeptide containing a VH domain. In some embodiments, the polypeptide further contains a CH1 domain. In some embodiments, PES2 encodes all or part of the antibody. In some embodiments, PES2 encodes a polypeptide containing a CH2 domain and a CH3 domain. In some embodiments, PES12 encodes all or part of the antibody. In some embodiments, PES12 encodes a polypeptide containing a VL domain and a CL domain.

[0011] In another aspect, the present invention comprises the following components: (a) a first eukaryotic promoter (P1 Euk1 (b) The following components: (i) 5' splice site (5'ss12), (ii) prokaryotic promoter (P1 Prok1 (iii) a first resectable prokaryotic promoter module (ePPM1) comprising a 3' splice site (3'ss11), wherein the components of ePPM1 are 5'ss12-P1 Prok1 The nucleic acid molecule comprises a first expression cassette comprising: (c) a first polypeptide coding sequence (PES11); (d) a first 5' splice site (5'ss11); and (e) a useful peptide coding sequence (UPES), the components of which are P1 Euk1-ePPM1-PES11-5'ss11-UPES are operably linked to each other in the 5'~3' direction. In some embodiments, the first expression cassette further comprises a first nucleic acid sequence encoding a eukaryotic signal sequence (ESS11), where ESS11 is P1 Euk1 It is positioned between and ePPM1. In some embodiments, ePPM1 further comprises a first nucleic acid sequence encoding a prokaryotic signal sequence (PSS11), where PSS11 is P1 Prok1 It is positioned between and 3'ss11. In some embodiments, the nucleic acid molecule is a second eukaryotic promoter (P1 Euk2 ), (ii) a second polypeptide coding sequence (PES12), and (iii) a polyadenylation site (pA1), further comprising a second expression cassette, wherein these components are P1 Euk2 -PES12-pA1 is operably linked to each other in the 5'-3' direction. In some embodiments, the second expression cassette further comprises a second nucleic acid sequence encoding a eukaryotic signal sequence (ESS12), where ESS12 is P1 Prok2 It is located between and 3'ss12. In some embodiments, the second expression cassette consists of the following components: (i) 5' splice site (5'ss13), (ii) prokaryotic promoter (P1 Prok2 (iii) a nucleic acid sequence encoding a prokaryotic signal sequence (PSS12), and (iv) a resectable prokaryotic promoter module (ePPM2) comprising a 3' splice site (3'ss12), wherein these components are 5'ss13-P1 Prok2 -PSS12-3'ss12 are operably connected to each other in the 5'~3' direction, and ePPM2 is P1 Euk2It is positioned between and PES12. In some embodiments, UPES encodes all or part of a useful peptide selected from the group consisting of a tag, label, coat protein, and adapter protein. In some embodiments, the coat protein is selected from the group consisting of pI, pII, pIII, pIV, pV, pVI, pVII, pVIII, pIX, and pX of bacteriophage M13, f1, or fd. In some embodiments, the coat protein is pIII of bacteriophage M13.

[0012] In another embodiment, the present invention is characterized by a vector comprising any one of the nucleic acid molecules described herein. In yet another embodiment, the present invention is characterized by a vector set comprising a first vector and a second vector, wherein the first and second vectors each comprise the first and second nucleic acid molecules of any polypeptide expression system disclosed herein.

[0013] In another embodiment, the present invention features a host cell comprising the aforementioned nucleic acid, vector, and / or vector set. In some embodiments, the host cell is a prokaryotic cell. In some embodiments, the prokaryotic cell is a bacterial cell. In some embodiments, the bacterial cell is an Escherichia coli cell. In other embodiments, the host cell is a eukaryotic cell. In some embodiments, the eukaryotic cell is a mammalian cell. In some embodiments, the mammalian cell is an Expi293F cell, a CHO cell, a 293T cell, or an NSO cell. In one embodiment, the mammalian cell is an Expi293F cell.

[0014] In further embodiments, the present invention is characterized by a method for generating polypeptides, comprising culturing host cells containing one or more of the aforementioned nucleic acids, vectors, and / or vector sets in a culture medium. In some embodiments, the method further comprises recovering polypeptides from host cells or culture medium. [Brief explanation of the drawing]

[0015] [Figure 1]This is a schematic diagram showing the relative configuration of pDV2 and pRK-Fc nucleic acid molecules in a polypeptide expression system for modular protein expression. The diagram also shows the typical premRNA product after transcription of nucleic acid molecules in eukaryotic cells, the expected trans-splicing event between the two generated premRNA products, and the product obtained after translation of the spliced ​​mRNA molecule. [Figure 2A] Figures 2A and 2B, encompassing the partial sequence diagram of the pDV2 vector, show the 5'ss, 3'ss, and polypyrimidine regions (PPTs) in bold and underlined. In Figure 2B, the region encoding the 150-nt gene III sequence that hybridizes to pRK-Fc and pRK-Fc2-derived transcripts is in italics and underlined. Mutations from the wild-type PPT are in bold, italics, and underlined. The AATAAA potential polyadenylation site of gene III in the pDV2 vector is shown above the sequence of the silent mutation introduced into the mutant pDV2 vectors pDV2c and pDC2d (in bold, italics, and underlined). Figures 2A and 2B disclose sequence numbers 2, 3, 9, 10, and 19–22, respectively, in order of appearance. [Figure 2B] Figures 2A and 2B, encompassing the partial sequence diagram of the pDV2 vector, show the 5'ss, 3'ss, and polypyrimidine regions (PPTs) in bold and underlined. In Figure 2B, the region encoding the 150-nt gene III sequence that hybridizes to pRK-Fc and pRK-Fc2-derived transcripts is in italics and underlined. Mutations from the wild-type PPT are in bold, italics, and underlined. The AATAAA potential polyadenylation site of gene III in the pDV2 vector is shown above the sequence of the silent mutation introduced into the mutant pDV2 vectors pDV2c and pDC2d (in bold, italics, and underlined). Figures 2A and 2B disclose sequence numbers 2, 3, 9, 10, and 19–22, respectively, in order of appearance. [Figure 3]This is a partial sequence diagram of the pRK-Fc vector. The branching point consensus sequence (BP), polypyrimidine region, and 3'ss are shown in bold or bold and underlined in Figure 3. The 150 bp antisense gene III sequence is in italics and underlined. The first in-frame ATG codon after the CMV promoter is in bold, italics, and underlined. Figure 3 discloses sequence numbers 23 and 24, respectively, in order of appearance. [Figure 4A] This graph shows the effects of adding the ISE sequence or removing the latent polyadenylation motif of gene III in pDV2, as well as the effects of complementary pRK-Fc and pRK-Fc2 vectors on IgG expression levels (in μg / ml) in Expi293F cells. [Figure 4B] This graph shows the effect of the plasmid ratios of pDV2c and pRK-Fc2 on IgG expression levels (in μg / ml) in Expi293F cells. The values ​​shown are the mean and standard error of representative experiments from two independent, triplicate experiments. [Figure 5A] Figures 5A and 5B encompass a partial sequence diagram of the pDV2b vector. The 5'ss and 3'ss are shown, and these sequences are in bold. The polypyrimidine region (PPT) and 9-nt G-run ISE are shown, in bold and highlighted, respectively. In Figure 5B, the region encoding the 150-nt gene III sequence that hybridizes to transcripts derived from pRK-Fc and pRK-Fc2 is in italics. Mutations from the wild-type stII signal sequence and M13 gene III are in bold and italics, with wild-type nucleotide residues shown above the sequence. Potential AATAAA polyadenylation site motifs are shown above the sequence. The amino acids in parentheses are encoded by codons created by splicing in both E. coli and mammalian cells. The BsiWI and RsrII restriction enzyme sites at the 3' end of the signal sequence used for variable region sequence cloning are shown above the sequence. Figures 5A and 5B disclose sequence numbers 2, 25, 9, 10, 19, 20, 26, and 27, respectively, in order of appearance. [Figure 5B]Figures 5A and 5B encompass a partial sequence diagram of the pDV2b vector. The 5'ss and 3'ss are shown, and these sequences are in bold. The polypyrimidine region (PPT) and 9-nt G-run ISE are shown, in bold and highlighted, respectively. In Figure 5B, the region encoding the 150-nt gene III sequence that hybridizes to transcripts derived from pRK-Fc and pRK-Fc2 is in italics. Mutations from the wild-type stII signal sequence and M13 gene III are in bold and italics, with wild-type nucleotide residues shown above the sequence. Potential AATAAA polyadenylation site motifs are shown above the sequence. The amino acids in parentheses are encoded by codons created by splicing in both E. coli and mammalian cells. The BsiWI and RsrII restriction enzyme sites at the 3' end of the signal sequence used for variable region sequence cloning are shown above the sequence. Figures 5A and 5B disclose sequence numbers 2, 25, 9, 10, 19, 20, 26, and 27, respectively, in order of appearance. [Figure 6] This is a partial sequence diagram of the pRK-Fc2 vector. The branching point consensus sequence (BP), polypyrimidine region, and 3'ss are in bold. The 9-nt G-run ISE is highlighted. The 150 bp antisense gene III sequence is in italics. The first ATG triplet and in-frame stop codon are underlined. The first in-frame ATG codon after the CMV promoter is in bold and italics. The CMV promoter TATA box and transcription start site are shown above the sequence. Glutamate residues in parentheses are encoded by codons created by trans-splicing in mammalian cells. Figure 6 discloses sequence numbers 28 and 29, respectively, in order of appearance. [Figure 7A] This is a series of graphs showing the masses obtained by unconvolution of the heavy chain (left panel) and light chain (right panel) of purified IgG expressed in Expi293F cells, based on mass spectrometry. [Figure 7B] This table shows the predicted and observed masses for both the heavy and light chains in Figure 7A. [Figure 8A]This graph shows the yields (in mg) of five different specific IgG molecules purified from the supernatant (30 ml) of Expi293F cell cultures co-transfected with pDV2d (containing ISE but lacking the AATAAA motif of gene III) and the pRK-Fc2 vector. n=3. Error bars indicate the standard error of the mean. [Figure 8B] This graph shows the yields (in mg) of IgG molecules of five different specificities purified from the supernatant (30 ml) of 293T and CHO cell cultures co-transfected with pDV2d (containing ISE and lacking the AATAAA motif of gene III) and the pRK-Fc2 vector. n=4. Error bars indicate the standard error of the mean. [Figure 9] This is a partial sequence diagram of the pRK-Fab-Flag vector showing the region between the CMV promoter TATA box fused to the Flag tag sequence and the human IgG1 upper hinge region. The hinge and Flag tag sequences are followed by an SV40 polyadenylation signal (not shown). The 3'ss, including the polypyrimidine region and consensus branch point (BP), is shown in bold or bold and underlined. The ISE sequence is shown in bold, underlined, and italicized. The antisense gene III sequence, which mediates hybridization with the donor transcript, is italicized and underlined. Figure 9 discloses Sequence IDs 30 and 31, in order of appearance, respectively. [Figure 10] This is a schematic diagram showing the relative configuration of the first and second nucleic acid molecules possible for the modular expression of a typical polypeptide product. The diagram also shows a typical premRNA product after transcription of the nucleic acid molecule in a eukaryotic cell, the expected trans-splicing event between the two generated premRNA products, and the product obtained after translation of the spliced ​​mRNA molecule. [Figure 11] This is a schematic diagram showing the relative configuration of first and second nucleic acid molecules that are possible for the modular expression of two or more typical polypeptide products. The diagram also shows typical premRNA products after transcription of nucleic acid molecules in eukaryotic cells, the expected trans-splicing event between the two generated premRNA products, and the product obtained after translation of the spliced ​​mRNA molecule. [Figure 12] This is a set of Western blots showing the expression of Mab1 heavy chain and Fd-cP3 fusion protein in Expi293F cells co-transfected with the pDV2 mutant and pRK-Fc2. Transfected Expi293F lysates were reduced with dithiothreitol (DTT) and analyzed by Western blotting using anti-IgG1 Fc (upper panel) or anti-M13 p3 (lower panel) antibodies. GFP and HC control vectors express green fluorescent protein and human IgG1 heavy chain, respectively. HC represents the full-length human IgG1 heavy chain. Gene III AATAAA indicates the presence of a potential polyadenylation site in gene III. FC* indicates the expression product of a putative cytoplasmic N-terminal cleaved Fc fragment. NA, not applicable. [Figure 13] This is a sodium dodecyl sulfate polyacrylamide gel electrophoresis (SDS-PAGE) gel showing the analysis of purified IgG and Fab fragments expressed in Expi293F cells. IgG and Fab were expressed and purified from the supernatant of Expi293F cells co-transfected with pDV2d and pRK-Fc2 (IgG) or pRK-FAB-F (Fab fragment). These were purified, separated by 4-20% gradient SDS-PAGE under reducing or non-reducing conditions, and stained with Coomassie brilliant blue. Band identity is shown on the right. HC: heavy chain. LC: light chain. Fd: heavy chain Fd fragment (VH+CH1+ upper hinge). The HC and LC (non-reducing) bands represent heavy and light chains that do not form interchain disulfide bonds, although intrachain disulfide bonds may be present in IgG samples. The approximately 25 kDa band in the non-reducing Fab sample did not form interchain disulfide bonds, but it possesses cotransitioning heavy and light chains that may have intrachain disulfide bonds. [Figure 14]This graph shows the display of Fab fragments on phages containing the phagemide pDV2, as detected by phage enzyme-linked immunosorbent assay (ELISA). Fab-zip phages were generated by infecting E. coli cells containing the pFab-zip phagemide with the M13KO7 helper phage. pDV2 phages were generated by infecting E. coli cells containing the pDV2d vector with the Amber-2614 KO7 phage. [Modes for carrying out the invention]

[0016] I. Definition The term "antibody" as used herein is used in its broadest sense and encompasses, but is not limited to, monoclonal antibodies, polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies), and antibody fragments, insofar as they exhibit a variety of antibody structures and the desired antigen-binding activity.

[0017] The Kabat numbering system is commonly used when referring to residues within the variable domain (approximately residues 1-107 of the light chain and residues 1-113 of the heavy chain) (e.g., Kabat et al., Sequences of Immunological Interest. 5th edition. Public Health Service, National Institutes of Health, Bethesda, Md. (1991)). The "EU numbering system" or "EU index" is commonly used when referring to residues within the constant region of the immunoglobulin heavy chain (e.g., the EU index reported by Kabat et al. above). "Kabat's EU index" refers to the residue numbering of human IgG1 EU antibodies. Unless otherwise specified herein, references to residue numbers within the variable domain of an antibody mean residue numbering according to the Kabat numbering system. Unless otherwise specified herein, references to residue numbers within the constant region of the heavy chain of an antibody mean residue numbering according to the EU numbering system.

[0018] The basic natural quadrivalent antibody unit is a heterotetrameric glycoprotein consisting of two identical light chains (LCs) and two identical heavy chains (HCs). (IgM antibodies consist of five basic heterotetrameric units with an additional polypeptide called a J chain, and therefore contain 10 antigen-binding sites. On the other hand, secreted IgA antibodies can polymerize to form a polyvalent assembly containing 2 to 5 basic quadrivalent units with J chains.) In the case of IgG, the quadrivalent unit is generally about 150,000 daltons. Two HCs are linked to each other by one or more disulfide bonds depending on the HC isotype, while each LC is linked to an HC by one covalent disulfide bond. Each HC and LC also has intrachain disulfide crosslinks at regular intervals. Each HC has a variable domain (VH) at its N-terminus, followed by three constant domains (CH1, CH2, CH3) for the α and γ chains, respectively, and four Cj domains for the μ and ε isotypes. Each LC has a variable domain (VL) at the N-terminus followed by a constant domain (CL) at the other end. The VL aligns with the VH, and the CL aligns with the first constant domain (CH1) of the heavy chain. CH1 can be connected to the second constant domain (CH2) of the heavy chain via a hinge region. Certain amino acid residues are thought to form junctions between the light chain and the heavy chain variable domain. The VH and VL form a pair and together form a single antigen-binding site. For the structures and properties of various classes of antibodies, see, for example, Basic and Clinical Immunology, 8th edition, Daniel P. Stites, Abba I. Terr and Tristram G. Parslow (eds.), Appleton & Lange, Norwalk, CT, 1994, page 71 and Chapter 6.

[0019] The "CH2 domain" in the human IgG Fc region typically extends from approximately IgG residues 231 to 340. The CH2 domain is unique in that it is not closely paired with another domain. Rather, two N-linked branched carbohydrate chains interpose between the two CH2 domains of an intact native IgG molecule. It is hypothesized that the carbohydrates act as an alternative to domain-domain pairing, potentially helping to stabilize the CH2 domain. (Burton, Molec. Immunol. 22:161-206 (1985)).

[0020] The "CH3 domain" contains a sequence of residues that are C-terminus of the CH2 domain in the Fc region (i.e., amino acid residues approximately 341 to 447 of IgG).

[0021] Light chains (LCs) from any vertebrate species can be assigned to one of two distinct types, called kappa and lambda, based on the amino acid sequence of their constant domains. Based on the amino acid sequence of the constant domains (CHs) of their heavy chains, immunoglobulins can be assigned to various classes or isotypes. There are five classes of immunoglobulins: IgA, IgD, IgE, IgG, and IgM, each with heavy chains named α, δ, γ, ε, and μ, respectively. The γ and α classes are further divided into subclasses based on relatively small differences in CH sequences and function; for example, humans express the following subclasses: IgG1, IgG2, IgG3, IgG4, IgA1, and IgA2.

[0022] The term "variable" refers to the fact that certain parts of the variable domain have widely different sequences among antibodies. The V domain mediates antigen binding and determines the specificity of a particular antibody to that particular antigen. However, the variability is not uniformly distributed across the entire length of the 110 amino acids of the variable domain. Instead, the V region consists of a relatively invariant sequence of 15-30 amino acids called the framework region (FR), and shorter, extremely variable regions of 9-12 amino acids each, called "hypervariable regions," that separate it. Each of the native heavy and light chain variable domains contains four FRs, which primarily take the form of a beta-sheet structure, and these are connected by three hypervariable regions that connect this beta-sheet structure and, in some cases, form loops that make up parts of the beta-sheet structure. The hypervariable regions within each chain are held together in close proximity by the fiber-retaining domain (FR), and together with the hypervariable regions from other chains, they contribute to the formation of the antibody's antigen-binding site (see Kabat et al., Sequences of Proteins of Immunological Interest, 5th edition. Public Health Service, National Institutes of Health, Bethesda, MD, 1991). The constant domain does not directly participate in antibody binding to the antigen, but exhibits various effector functions, such as the involvement of antibodies in antibody-dependent cytotoxicity (ADCC).

[0023] An "antibody fragment" refers to a molecule other than the intact antibody, which contains a portion of the intact antibody that binds to the antigen to which the intact antibody binds. Examples of antibody fragments include, but are not limited to, Fv, Fab, Fab', Fab'-SH, F(ab')2; bispecific antibodies; linear antibodies; single-chain antibody molecules (e.g., scFv); and multispecific antibodies formed from antibody fragments.

[0024] The "Fab" fragment is an antigen-binding fragment produced by papain digestion of an antibody, consisting of the entire light chain with a variable region domain (VH) of the heavy chain and a first constant domain (CH1) of one heavy chain. Papain digestion of an antibody produces two identical Fab fragments. Pepsin treatment of an antibody produces a single large F(ab')2 fragment, which roughly corresponds to two disulfide-bonded Fab fragments with divalent antigen-binding activity and can still crosslink antigens. The Fab' fragment differs from the Fab fragment by having several additional residues, including one or more cysteines from the antibody hinge region, at the carboxyl terminus of the CH1 domain. Fab'-SH is the herein nomenclature for Fab' fragments in which the cysteine ​​residue(s) of the constant domain have a free thiol group. The F(ab')2 antibody fragment was originally produced as a pair of Fab' fragments with a hinge cysteine ​​between them. Other chemical bonds of antibody fragments are also known.

[0025] As used herein, “adapter protein” refers to a protein sequence that specifically interacts with another adapter protein sequence in solution. In one embodiment, the “adapter protein” includes a heteromultimerizing domain. Such an adapter protein is a leucine zipper protein, or SEQ ID NO: 4 (cJUN(R):ASIARL[E]E[K]V KTL[K]A[Q]NYEL [A]S[T]ANMLRE[Q] VAQLGGC) or SEQ ID NO: 5 (FosW(E):AS[I]DEL[Q]AE[V] EQLEE[R]NYAL [R]KE[V]EDL[Q]K[Q] [A]EKLGGC) amino acid sequence or variations thereof (amino acids of SEQ ID NOs: 4 and 5, which may be modified to include underlined and bolded ones, in addition to the amino acids of the underlined and bolded sequences), wherein the variation has an amino acid modification that maintains or increases the affinity of the adapter protein to another adapter protein, or a polypeptide comprising an amino acid sequence selected from the group consisting of SEQ ID NOs: 11 (ASIARLRERVKTLRARNYELRSRANMLRERVAQLGGC) or SEQ ID NOs: 12 (ASLDELEAEIEQLEEENYALEKEIEDLEKELEKLGGC), or a polypeptide comprising the amino acid sequence of SEQ ID NOs: 13 (GABA-R1:EEKSRLLEKE NRELEKIIAE KEERVSELRH QLQSVGGC) or SEQ ID NOs: 14 (GABA-R2:TSRLEGLQSE NHRLRMKITE LDKDLEEV™ QLQDVGGC) or SEQ ID NOs: 15 (Cys:AGSC) or SEQ ID NOs: 16 (Hinge:CPPCPG). Nucleic acid molecules encoding coat proteins or adapter proteins are contained within synthetic introns.

[0026] As used herein, “heteromultimerizing domain” refers to a modification or addition to a biological molecule that promotes heteromultimerization and inhibits homomultimerization. Any heterodimerizing domain having a strong preference for forming heterodimers over homodimers is within the scope of the present invention. Exemplary examples include, but are not limited to, U.S. Patent Application No. 20030078385 (Arathoon et al., Genentech, describing a knob to a hole), International Patent Publication No. WO2007147901 (Kjargaard et al., Novo Nordisk, describing an ionic interaction), International Patent Publication No. WO2009089004 (Kannan et al., Amgen, describing an electrostatic steering effect), and International Patent Publication No. WO2011 / 034605 (Christensen et al., Genentech, describing a multicoil). See also, for example, Pack, P. & Plueckthun, A., Biochemistry 31, 1579-1584 (1992), describing the leucine zipper, or Pack et al., Bio / Technology 11, 1271-1277 (1993), describing the helix-turn-helix motif. The terms “heteromultimerizing domain” and “heterodimerizing domain” are used interchangeably herein.

[0027] As used herein, the term “cloning site” refers to a nucleic acid sequence containing a restriction enzyme site for restriction endonuclease-mediated cloning by ligation of a nucleic acid sequence containing compatible adherent or blunt ends, a region of a nucleic acid sequence that acts as a priming site for PCR-mediated cloning of insert DNA by homology and extension known as “overlap PCR stitching”, a recombinant site for recombinase-mediated insertion of a target nucleic acid sequence by recombinase-exchange reactions, a mosaic end for transposon-mediated insertion of a target nucleic acid sequence, and other nucleic acid sequences that include other techniques common in the art.

[0028] As used herein, “coat protein” refers to any of the five capsid proteins that are components of a phage particle, including pIII, pVI, pVII, pVIII, and pIX. In one embodiment, “coat protein” may be used to display a protein or peptide (see Phage Display, A Practical Approach, Oxford University Press, edited by Clackson and Lowman, 2004, pp. 1–26). In one embodiment, the coat protein may be a variation, part, and / or derivative of the pIII protein or some of its derivatives. For example, the C-terminal portion of the M13 bacteriophage pIII coat protein (cP3), e.g., the sequence encoding C-terminal residues 267–421 of protein III of the M13 phage, may be used. In one embodiment, the pIII sequence is the amino acid sequence of SEQ ID NO: 17 (AEDIEFASGGGSGAETVESCLAKPHTENSFTNVWKDDKTLDRYANYEGCLWNATGVVVCTGDETQCYGTWVPIGLAIPENEGGGSEGGGSEGGGSEGGGTKPPEYGDTPIPGYTYINPLDGTYPPGTEQNPANPNPSLEESQPLNTFMFQNNRFRNRQGALTVYTGTVTQGTDPVKTYYQYTPVSSKAMYDAYWNG (Includes KFRDCAFHSGFNEDPFVCEYQGQSSDLPQPPVNAGGGSGGGSGGGSEGGGSEGGGSEGGGSEGGGSGGGSGSGDFDYEKMANANKGAMTENADENALQSDAKGKLDSVATDYGAAIDGFIGDVSGLANGNGATGDFAGSNSQMAVGDGDNSPLMNNFRQYLPSLPQSVECRPFVFSAGKPYEFSIDCDKINLFRGVFAFLLYVATFMYVFSTFANILRNKES)In one embodiment, the pIII fragment includes the amino acid sequence of SEQ ID NO: 18 (SGGGSGSGDFDYEKMANANKGAMTENADENALQSDAKGKLDSVATDYGAAIDGFIGDVSGLANGNGATGDFAGSNSQMAQVGDGDNSPLMNNFRQYLPSLPQSVECRPFVFGAGKPYEFSIDCDKINLFRGVFAFLLYVATFMYVFSTFANILRNKES).

[0029] As used herein, “expression cassette” means a nucleic acid fragment (e.g., a DNA fragment) containing a specific nucleic acid sequence having particular biological and / or biochemical activity. The terms “cassette,” “gene cassette,” and “DNA cassette” can be used interchangeably and may have the same meaning.

[0030] The terms “host cell,” “host cell line,” and “host cell culture” are used interconvertibly to refer to cells into which exogenous nucleic acids have been introduced (including the offspring of such cells). Host cells include “transformers” and “transformed cells,” including primary transformed cells and their offspring, regardless of passage number. Offspring may contain mutations, but their nucleic acid content may not be exactly the same as that of the parent cells. Mutant offspring having the same function or biological activity as those screened or selected in the original transformed cells are included herein.

[0031] As used herein, “linked,” “links,” or “links” means a covalent bond between two amino acid sequences or two nucleic acid sequences, respectively, via a peptide or phosphodiester bond, which may include any number of additional amino acid sequences or nucleic acid sequences between the two linked amino acid sequences or nucleic acid sequences.

[0032] "Nucleic acid" or "polynucleotide," when used interconvertibly herein, refers to a polymer of nucleotides of any length, including DNA and RNA. The nucleotides may be deoxyribonucleotides, ribonucleotides, modified nucleotides or bases, and / or analogs thereof, or any substrate that can be incorporated into the polymer by DNA or RNA polymerase or by synthetic reaction. Polynucleotides may include modified nucleotides, such as methylated nucleotides and their analogs. Modifications to the nucleotide structure, if present, may be made before or after the organization of the polymer. Nucleic acid sequences may be interrupted by non-nucleotide components. Polynucleotides may be further modified after synthesis, such as by conjugation with labels. Other types of modifications include, for example, the substitution of one or more natural nucleotides with analogues as "caps," internucleotide modifications such as those having non-charged links (e.g., methylphosphonates, phosphotriesters, phosphoamidates, carbamates, etc.) and those having charged links (e.g., phosphorothioates, phosphorodithioates, etc.), those containing pendant moieties such as proteins (e.g., nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), those containing intercalators (e.g., acridine, psoralens, etc.), those containing chelating agents (e.g., metals, radioactive metals, boron, metal oxides, etc.), those containing alkylating agents, those having modified links (e.g., alpha-aromatic nucleic acids, etc.), and unmodified forms of polynucleotides. Furthermore, any of the hydroxyl groups originally present in the sugar may be substituted with, for example, a phosphonic acid group, a phosphate group, protected with a standard protecting group, or activated to prepare additional linkages to additional nucleotides, or bonded to a solid or semi-solid support. The 5' and 3' terminal OH groups may be phosphorylated or substituted with amines or organic capping groups of 1 to 20 carbon atoms. Other hydroxyls may be derivatized to standard protecting groups.Furthermore, the polynucleotide may contain, for example, analogous forms of ribose or deoxyribose sugars commonly known in the art, including 2'-O-methyl-, 2'-O-allyl, 2'-fluoro-, or 2'-azidol-ribose; carbocyclic sugar analogs; alpha-anomeric sugars; epimeric sugars such as arabinose, xylose, or lyxose; pyranose sugars; furanose sugars; sedoheptulose; acyclic analogs; and basic nucleoside analogs such as methylriboside. One or more phosphodiester bonds may be substituted with alternative linking groups. These alternative linking groups include, but are not limited to, embodiments in which the phosphate is substituted with P(O)S ("thioate"), P(S)S ("dithioate"), (O)NR2 ("amidate"), P(O)R, P(O)OR', CO, or CH2 ("formacetal"), where each R or R' is independently H, or optionally an ether (-O-) linkage, aryl, alkenyl, cycloalkyl, cycloalkenyl, or araldyl-containing substituted or unsubstituted alkyl (1-20C). Not all linkings within the polynucleotide are to be identical. The above description applies to all polynucleotides referred to herein, including RNA and DNA.

[0033] When nucleic acids are placed in a structural or functional relationship with another nucleic acid sequence, they are "operably linked." For example, one segment of DNA and another segment of DNA can be operably linked to another segment of DNA if they have structural or functional relationships such as a promoter or enhancer positioned relative to the coding sequence to promote transcription of the coding sequence; a ribosome binding site positioned relative to the coding sequence to promote translation; or a presequence or secretion leader positioned relative to the coding sequence to promote the expression of a preprotein (e.g., a preprotein involved in the secretion of an encoded polypeptide). In other examples, operably linked nucleic acid sequences are not contiguous, but they are positioned to have a functional relationship with each other as nucleic acids or as proteins expressed by them. For example, enhancers do not have to be contiguous. Linking can be achieved by ligation at convenient restriction enzyme sites or by using synthetic oligonucleotide adapters or linkers.

[0034] The terms “polyadenylation signal” or “polyadenylation site” are used herein to mean a sequence sufficient to direct the addition of polyadenosine ribonucleic acid to an RNA molecule expressed in a cell.

[0035] A "promoter" is a nucleic acid sequence that enables the initiation of transcription of a gene sequence in messenger RNA, and such transcription is initiated by the binding of RNA polymerase on or near the promoter.

[0036] The term "3' splice site" is intended to mean the nucleic acid sequence, such as a premRNA sequence, at the 3' intron / exon boundary that can be recognized and bound by the splicing mechanism.

[0037] The term "5' splice site" is intended to mean the nucleic acid sequence, such as a premRNA sequence, at the 5' exon / intron boundary that can be recognized and bound by the splicing mechanism.

[0038] The term "latent splice site" is intended to refer to a normally dormant 5' or 3' splice site that can be activated by mutation or otherwise and function as a splicing element. For example, a mutation may activate a 5' splice site downstream of a native or dominant 5' splice site. The use of this "latent" splice site results in the generation of a distinct mRNA splicing product that would not be produced by the use of the native or dominant splice site.

[0039] As used herein, the term “trans-splicing” means the joining of exons contained on separate, discontinuous RNA molecules.

[0040] The term "variable region" or "variable domain" refers to a domain in the antibody heavy or light chain that is involved in the binding of an antibody to an antigen. The variable domains of the heavy and light chains of native antibodies (VH and VL, respectively) generally have similar structures, and each domain contains four conserved framework regions (FRs) and three hypervariable regions (HVRs). (See, for example, Kindt et al., Kuby Immunology, 6th edition, WH Freeman and Co., page 91 (2007).) A single VH or VL domain may be sufficient to confer antigen-binding specificity. Furthermore, antibodies that bind to a specific antigen can be isolated from antigen-binding antibodies using the VH or VL domain to screen libraries of complementary VL or VH domains, respectively. See, for example, Portolano et al., J.Immunol. 150:880-887 (1993); Clarkson et al., Nature 352:624-628 (1991).

[0041] As used herein, the term “vector” refers to a nucleic acid molecule capable of amplifying another nucleic acid into which it is ligated. This term includes vectors as self-replicating nucleic acid structures, and vectors incorporated into the genome of a host cell into which they are introduced. Certain vectors can direct the expression of nucleic acids into which they are operably ligated. Such vectors are referred herein to as “expression vectors.”

[0042] II. Modular Polypeptide Expression System This invention is at least in part based on the discovery that pre-mRNA trans-splicing can be utilized in mammalian cells to enable modular recombinant protein expression. The concept of modular, flexible protein expression makes it possible to precisely join any two protein-coding sequences encoded by two different constructs into a single mRNA encoding a polypeptide chain, without any of the requirements and constraints of other protein-protein splicing methods. This concept can be adapted to simplify and extend other techniques that require the expression of many collections of proteins with various combinations of repeating modules in mammalian cells.

[0043] This section describes the generation of numerous polypeptide expression systems that enable modular expression of various antibody formats in relation to phage display expression systems. The necessary nucleic acid components, vectors, host cells, and methods for using the polypeptide expression systems of the present invention are described herein.

[0044] A. Embodiment of the invention Unless otherwise indicated, the implementation of this invention will utilize prior art in molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry, and immunology, which is within the scope of the art. Such techniques are fully described in literature such as "Molecular Cloning: A Laboratory Manual" 2nd edition (Sambrook et al., 1989), "Oligonucleotide Synthesis" (MJ Gait, edition, 1984), "Animal Cell Culture" (RI Freshney, edition, 1987), "Methods in Enzymology" (Academic Press, Inc.), "Handbook of Experimental Immunology" 4th edition (DM Weir & C.C. Blackwell, eds., Blackwell Science Inc., 1987), "Gene Transfer Vectors for Mammalian Cells" (JMMiller & M.P. Calos, eds., 1987), "Current Protocols in Molecular Biology" (FMAusubel et al., eds., 1987), "PCR: The Polymerase Chain Reaction" (Mullis et al., eds., 1994), and "Current Protocols in Immunology" (JEColigan et al., eds., 1991).

[0045] B. Modular protein expression system The polypeptide expression system of the present invention can support the expression of polypeptides (e.g., fusion proteins) in the same or different (e.g., reformatted) forms. The present invention provides means for generating such polypeptide expression systems for the modular expression and generation of various forms (e.g., various formats or various fusion forms) of a target protein in a host cell-dependent manner, by utilizing a trans-splicing process.

[0046] 1. Nucleic acid components of modular protein expression systems a. Structure of nucleic acid components in modular protein expression systems The protein expression system uses at least two nucleic acid molecules that together enable the flexible modular expression of any desired polypeptide through the process of directed premRNA trans-splicing. The first nucleic acid molecule is a eukaryotic promoter (P1 Euk1 The first expression cassette includes a polypeptide coding sequence (PES11) (for example, a cytomegalovirus (CMV) promoter, a simian virus 40 (SV40) promoter, a Moloney mouse leukemia virus U3 region, a canine arthritis encephalitis virus U3 region, a visnavirus U3 region, or a retrovirus U3 region sequence), which is operably linked to a polypeptide coding sequence (PES11). In some cases, the polypeptide coding sequence codes only a portion of the desired polypeptide, with the remainder supplied by a polypeptide coding sequence (PES2) contained on a second nucleic acid molecule. The first nucleic acid molecule may include a 5'ss (5'ss11) (for example, GTAAGA (SEQ ID NO: 8)) located downstream (3') of PES11 but upstream (5') of the hybridize sequence (HS1).

[0047] The HS1 sequence may contain a gene encoding all or part of a polypeptide tag, label, coat protein, and / or adapter protein that can be positioned in frame with PES11, so that its expression results in a PES11-encoded protein fused to the HS1-encoded protein. In some cases, HS1 is a gene encoding all or part of a coat protein selected from the group consisting of pI, pII, pIII, pIV, pV, pVI, pVII, pVIII, pIX, and pX of bacteriophage M13, f1, or fd. For example, PES11 may encode all or part of an antibody or its Fab fragment, and the HS1 sequence may encode a coat protein (e.g., all or part of the pIII protein of bacteriophage M13, e.g., the pIII fragment, containing amino acid residues 267-421 or 262-418 of the pIII protein), resulting in an antibody-or Fab fragment-pIII protein fusion product. In another case, HS1 is a gene that codes for all or part of an adapter protein, such as a leucine zipper, which contains the amino acid sequence of SEQ ID NO: 4 or 5.

[0048] In addition, the first nucleic acid molecule is P1 Euk1 It can encode the eukaryotic signal sequence (ESS11) located at 3' relative to P1 and 5' relative to PES11. Therefore, the first nucleic acid molecule is P1 Euk1 -ESS11-PES11-5'ss11-HS1 may include the above components which are connected to each other in the 5'~3' direction (for example, operably connected).

[0049] The second nucleic acid molecule of the protein expression system is a eukaryotic promoter (P2) operably linked to the polypeptide coding sequence (PES2). Euk)(for example, a cytomegalovirus (CMV) promoter or a monkey virus 40 (SV40) promoter). In some cases, the polypeptide coding sequence codes only a portion of the desired polypeptide, with the remainder supplied by a polypeptide coding sequence (PES11) contained on the first nucleic acid molecule. The second nucleic acid molecule may contain a 3' splice site (3'ss2) located 5' relative to PES2. Euk The second nucleic acid molecule may include a hybridization sequence (HS2) that can hybridize to HS1 located between and 3'ss2. Furthermore, the second nucleic acid molecule may include a polyadenylation site (pA2), where the components of the second nucleic acid molecule are P2 Euk -HS2-3'ss2-PES2-pA2 are operably connected to each other in the 5'~3' direction.

[0050] Therefore, trans-splicing between the first and second nucleic acid premRNA products within a eukaryotic cell (e.g., a mammalian cell) is induced by hybridization of complementary sequences (i.e., HS1 and HS2) located on separate mRNA molecules. This brings the isolated 5' splice site (5'ss11) of the first molecule and the isolated 3' splice site (3'ss2) of the second molecule into proximity, causing trans-splicing to occur and supporting the formation of the desired trans-spliced ​​mRNA transcript. In addition, to facilitate trans-splicing, the first nucleic acid molecule may include an intron splice enhancer (ISE) (ISE1) positioned between 5'ss11 and HS1. ISE1 includes, for example, a G-run having three or more consecutive guanine residues, such as a G-run having nine consecutive guanine residues. Furthermore, trans-splicing between the first and second nucleic acid premRNA products can be induced during their transcription in eukaryotic cells (e.g., mammalian cells, e.g., Expi293F, 293T, or CHO cells) by genetically engineering the first nucleic acid molecule to lack the standard polyadenylation sites downstream of the PES11 and / or HS1 components. This would minimize the formation of a mature mRNA transcript that can be transported into the cytoplasm before trans-splicing with the mRNA transcript of the second nucleic acid molecule can occur.

[0051] In some cases, it may be desirable to express separate polypeptide products simultaneously. For example, it may be desirable to express a first polypeptide product encoded by both the first and second nucleic acid molecules and a second polypeptide product that can self-assemble to form a desired heteromultimeric protein product (e.g., an antibody consisting of both heavy and light chains). For this purpose, the first and / or second nucleic acid molecules may further include a second expression cassette. For example, when the first nucleic acid molecule includes a second expression cassette, the second expression cassette includes a second eukaryotic promoter (P1 Euk2), (ii) a second nucleic acid sequence encoding a eukaryotic signal sequence (ESS12), (iii) a second polypeptide coding sequence (PES12), and (iv) a polyadenylation site (pA1), and these components may include P1 Euk2 -ESS12-PES12-pA1 are operably linked to each other in the 5'-3' direction. In some cases, the second expression cassette may not contain the ESS12 component (e.g., if secretion of the expressed polypeptide is not required or desired). Thus, the first nucleic acid molecule encodes two polypeptide products under separate promoters, thereby forming one mRNA transcript encoding one of the polypeptide products of the first nucleic acid molecule via oriented transsplicing with the mRNA transcript encoded by the second nucleic acid molecule. In some cases, the second expression cassette is positioned 5' relative to the first expression cassette. In other cases, the second expression cassette is positioned 3' relative to the first expression cassette.

[0052] b. Polypeptide expression in both prokaryotic and eukaryotic cells In some cases, polypeptide expression systems can be genetically engineered for polypeptide expression in both prokaryotic and eukaryotic cells. Therefore, the first nucleic acid molecule is PES11, or in some cases PES11 and HS1, if expression of polypeptide products encoded by PES11 is desired. Euk1 It may include a resectable prokaryotic promoter module (ePPM1) positioned between and PES11. ePPM1 includes a 5' splice site (5'ss12) and a prokaryotic promoter (P1 Prok1 ), may include a nucleic acid sequence encoding a prokaryotic signal sequence (PSS11), and a 3' splice site (3'ss11), which are 5'ss12-P1 Prok1-PSS11-3'ss11 are positioned relative to each other in the 5'-3' direction and are operably linked to drive the transcription of PES11, or polypeptides encoded by PES11 and HS1. In some cases, ePPM1 may not contain PSS11 components (e.g., when secretion of the expressed polypeptide is not required or desired). Thus, ePPM1 can drive the transcription of PES11-encoded polypeptides of the first nucleic acid molecule in prokaryotic cells. On the other hand, in eukaryotic cells (e.g., mammalian cells), P1 Euk1 ePPM1 drives the transcription and expression of the polypeptide encoded by PES11 of the first nucleic acid molecule, and ePPM1 can be removed from the premRNA transcript by cis-splicing due to collisions with the 5'ss12 and 3'ss11 components.

[0053] In some cases, ePPM1 also includes a polypyrimidine region (PPT11) located between PSS11 and 3'ss11. PPT11 may include, for example, the sequence TTCCTTTTTTCTCTTTCC (SEQ ID NO: 1). The second nucleic acid molecule may also include a polypyrimidine region (PPT2) located between HS2 and 3'ss2. PPT2 may include, for example, the sequence TTCCTCTTTCCCTTTCTCTCCC (SEQ ID NO: 7). In addition, the second nucleic acid molecule may further include an ISE (ISE2) located between HS2 and 3'ss2. ISE2 may include a G-run having three or more consecutive guanine residues, such as a G-run having nine consecutive guanine residues.

[0054] In some embodiments, the first nucleic acid molecule of the polypeptide expression system includes a second expression cassette, where the second expression cassette is P1 Euk2 The module may further include a resectable prokaryotic promoter module (ePPM2) positioned between and PES12, comprising the following components: (i) a 5' splice site (5'ss13), and (ii) a prokaryotic promoter (P1 Prok2(iii) a nucleic acid sequence encoding a prokaryotic signal sequence (PSS12), and (iv) a 3' splice site (3'ss12), thereby comprising these components 5'ss13-P1 Prok2 -PSS12-3'ss12 are positioned relative to each other in the 5'-3' direction and are operably linked to drive the transcription of the polypeptide encoded by PES12. In some cases, ePPM2 may not contain the PSS12 component (e.g., if the secretion of the expressed polypeptide is not required or desired). The second excisable prokaryotic promoter module will function in a manner similar to that of the first excisable prokaryotic promoter module described above.

[0055] The prokaryotic promoter(s) of the resectable prokaryotic promoter module(s) may be the phoA, Tac, Lac, or Tphac promoter (see, for example, Kim et al. PLoS One.7(4):e35844), or another prokaryotic promoter known in this art.

[0056] An additional challenge in constructing vectors capable of expressing a target protein in both prokaryotic cells (e.g., E. coli cells) and eukaryotic cells (mammalian cells, e.g., Expi293F cells) arises from differences in signal sequences found in these cell types. Certain features of signal sequences are generally conserved in both prokaryotic and eukaryotic cells (e.g., hydrophobic residues located in the middle of the sequence, and patches of polar / charged residues adjacent to the cleavage site at the N-terminus of mature polypeptides), while others are more characteristic of one cell type than the other. Furthermore, it is known in the art that different signal sequences can have a significant impact on expression levels in mammalian cells, even if all sequences are of mammalian origin (Hall et al., J of Biological Chemistry, 265:19996-19999 (1990), Humphreys et al., Protein Expression and Purification, 20:252-264 (2000)). For example, bacterial signal sequences typically have a positively charged residue (most commonly lysine) immediately following the initiating methionine, whereas these are not always present in mammalian signal sequences.

[0057] If the secretion of the expressed protein is required or desired, any signal sequence (including consensus signal sequences) that targets the polypeptide of interest to the periplasm in prokaryotes and the endoplasmic reticulum in eukaryotes may be used. For example, a eukaryotic signal sequence (e.g., ESS11 or ESS12) may be derived from, or include in whole or in part, a mouse-conjugated immunoglobulin protein (mBiP) signal sequence (UniProtKB: accession number P20029) or an antibody heavy or light chain signal sequence (e.g., a mouse VH gene signal sequence). In some embodiments, a prokaryotic signal sequence (e.g., PSS11 or PSS12) may be derived from, or include in whole or in part, a thermostable enterotoxin II (stII) gene. Other signal sequences that may be used include those from human growth hormone (hGH) (UniProtKB: accession number BIA4G6), Gaussia princepus luciferase (UniProtKB: accession number Q9BLZ2), and yeast endo-1,3-glucanase (yBGL2) (UniProtKB: accession number P15703). The signal sequences may be native or synthetic. In some embodiments, the synthetic signal sequence is an optimized signal secretion sequence that drives an optimized level of display compared to its non-optimized native signal sequence.

[0058] 2. Vector, host cell, and method of production The present invention is characterized by a vector or vector set comprising one or more of the nucleic acid molecules described above. Accordingly, the present invention is also characterized by a vector set comprising a first vector and a second vector, wherein the first and second vectors each comprise the first and second nucleic acid molecules of the polypeptide expression system described above.

[0059] In addition to the nucleic acid molecular components described in detail above, a vector or vector set may contain nucleic acids encoding polypeptides useful as bacterial origins of replication, mammalian origins of replication, and / or controls (e.g., gD proteins) or for activity (e.g., protein purification, protein tagging, or protein labeling).

[0060] A method for generating polypeptides is also provided, comprising culturing host cells containing one or more of the above-mentioned vectors or vector sets in a culture medium, and optionally recovering antibodies from the host cells (or the culture medium of the host cells).

[0061] C. Phage display vector systems for modular antibody expression and reformatting In some embodiments, antibodies (e.g., full-length antibodies, e.g., full-length IgG antibodies, or fragments thereof, e.g., Fab fragments) can be generated using the polypeptide expression system of the present invention. The application of a modular protein expression system is demonstrated by designing a phage display vector system that enables the expression of different antibody formats within human cells from the same clone. The heavy chain antigen-binding region and a portion of the constant region encoded by the phage display vector were directly and precisely fused to the sequence encoded in a second complementary construct by pre-mRNA trans-splicing during intracellular expression, thereby binding the sequence encoding a different portion of the polypeptide.

[0062] The use of polypeptide expression systems aimed at enabling the direct expression of IgG in mammalian cells without requiring subcloning of the phage Fab sequence is described in Examples 1 and 2 below. In some cases, the first nucleic acid molecule of the polypeptide expression system may be designed to encode the entire Fab fragment component. Thus, the first nucleic acid molecule may include a PES11 component encoding the polypeptide having the VH domain and CH1 domain of Fab. The first nucleic acid molecule may also include a PES12 component encoding the VL domain and CL domain. Transcription of the first nucleic acid molecule yields two discontinuous premRNA products, which together form a Fab fragment that can be appropriately tagged (e.g., fused to pIII of M13) for the purpose of phage display.

[0063] The process of reformatting the Fab fragment into a full-length IgG antibody can then be achieved by expressing the first nucleic acid molecule in eukaryotic cells (e.g., mammalian cells, e.g., Expi293F cells) accompanied by a second nucleic acid molecule providing the rest of the antibody (i.e., the CH2 and CH3 domains). For example, the second nucleic acid molecule may contain a PES2 component encoding a polypeptide having CH2 and CH3 domains. Transcription of the first and second molecules in eukaryotic cells will result in the generation of three premRNA transcripts, and the heavy chains encoding the premRNA transcripts will be induced to undergo trans-splicing with one another, generating the reformatted full-length heavy chain of the desired IgG antibody. The processed mRNA will then be translated, resulting in the generation of both the light and heavy chains of the IgG molecule, such generation will not require the need for labor-intensive subcloning.

[0064] When different antibody formats, such as wild-type IgG, Fab fragments, or IgG with Fc modifications for bispecificity formats, are required for different screening assays, the ability to express different antibody formats from the same clone is useful in antibody discovery. The polypeptide expression system of the present invention enables any of these or additional formats by simply cloning a suitable sequence to be added after the CH1 region in the complementary plasmid. Furthermore, the modular structure of this system enables the expression of new antibody formats without requiring the recreation of a stock of phage display libraries, as it only requires the construction of a new complementary plasmid. The nucleic acid can also be adapted to allow the use of any CH1 region by transferring 5′ss from downstream of the CH1 coding region to the VH or the J region (FR4) of the J-CH1 junction, thereby separating the entire constant region of the VH and heavy chain in two different nucleic acids. The nucleic acid molecule can be adapted to conventional methods for the expression of Fab fragments in E. coli by simply adding a stop codon after the sequence encoding the upper hinge. However, amber stop codons at the junction of heavy chains and gene III sequences in Fab phage display libraries typically result in significantly lower levels of display, thus requiring clonal reformatting after selection, at least in the case of inexperienced repertoire libraries (Lee et al., Journal of Immunological Methods, 284:119-132, 2004). Expression of Fab fragments in mammalian cells using the same methods as those used for IgG expression avoids this need for reformatting, with yields comparable to those typically obtained in E. coli.

[0065] Antibodies produced by this polypeptide expression system may include recombinantly produced chimeric, humanized, and / or human antibodies. In some cases, the antibody is an antibody fragment, e.g., Fab, Fv, Fab′, scFv, a bispecific antibody, or an F(ab′)2 fragment. In other cases, the antibody is a full-length antibody, e.g., an intact IgG1, IgG2, IgG3, or IgG4 antibody as defined herein, or another antibody of a different class or isotype.

[0066] The expressed antibodies may incorporate any of the following characteristics, either individually or in combination, as described in sections 1-7 below.

[0067] 1. Antibody affinity Antibodies produced by polypeptide expression systems described herein (e.g., Fab or full-length IgG antibodies) are defined as having concentrations of ≤1 μM, ≤100 nM, ≤10 nM, ≤1 nM, ≤0.1 nM, ≤0.01 nM, or ≤0.001 nM (e.g., 10 -8 M or less, for example, 10 -8 M~10 -13 M, for example 10 -9 M~10 -13 It may have a dissociation constant (Kd) of M.

[0068] In one embodiment, Kd is measured by radiolabeled antigen-binding assay (RIA) performed on the Fab version of the antibody and its antigen, as described by the assay method below. The solution binding affinity of Fab to the antigen is determined in the presence of a titration series of the unlabeled antigen, ( 125I) Measurement is performed by equilibrating Fab with the minimum concentration of labeled antigen, and then capturing the antigen bound to the anti-Fab antibody coated plate (see, e.g., Chen et al., J.Mol.Biol.293:865-881 (1999)). To establish conditions for the assay, MICROTITER® multi-well plate (Thermo Scientific) is coated overnight with 5 μg / ml of capture anti-Fab antibody (Cappel Labs) in 50 mM sodium carbonate (pH 9.6), and then blocked with 2% (w / v) bovine serum albumin in PBS for 2-5 hours at room temperature (approx. 23°C). In non-adsorbent plates (Nunc#269620), 100 pM or 26 pM [ 125 I) The antigen is mixed with serial dilutions of the Fab of interest (e.g., consistent with the evaluation of anti-VEGF antibody, Fab-12, in Presta et al., Cancer Res. 57:4593-4599 (1997)). The Fab of interest is then cultured overnight, but the culture may be continued for a longer period (e.g., about 65 hours) to ensure that equilibrium is reached. The mixture is then transferred to a capture plate for incubation at room temperature (e.g., 1 hour). The solution is then removed, and the plate is washed eight times with 0.1% polysorbate 20 (TWEEN-20®) in PBS. Once the plate is dry, 150 μl / well of scintillant (MICROSCINT-20®; Packard) is added, and the plate is counted for 10 minutes using a TOPCOUNT® gamma counter (Packard). The concentration of each Fab that provides a maximum binding of 20% or less is selected for use in competitive binding assays.

[0069] According to another embodiment, Kd is measured at 25°C using a CM5 immobilized antigen chip with approximately 10 response units (RUs) and surface plasmon resonance assay using BIACORE®-2000 or BIACORE®-3000 (BIAcore, Inc., Piscataway, NJ). In short, the carboxymethylated dextran biosensor chip (CM5, BIAcore Inc.) is activated with N-ethyl-N'-(3-dimethylaminopropyl)-carbodimide hydrochloride (EDC) and N-hydroxysuccinimide (NHS) according to the supplier's instructions. The antigen is diluted to 5 μg / ml (approximately 0.2 μM) in 10 mM sodium acetate at pH 4.8 and injected at a flow rate of 5 μl / min to achieve approximately 10 bound protein response units (RUs). After antigen injection, 1 M ethanolamine is injected to block the unresponsive group. For dynamic measurement, Fab (0.78 nM to 500 nM), serially diluted 2-fold, is injected into PBS containing 0.05% polysorbate 20 (TWEEN-20®) surfactant (PBST) at a flow rate of approximately 25 μl / min at 25°C. The association rate (k on ) and dissociation rate (k off The equilibrium dissociation constant (Kd) is calculated using a simple one-to-one Langmuir coupled model (BIACORE® evaluation software version 3.2) by simultaneously fitting the association and dissociation sensorgrams. off / k on The calculation was performed as a ratio. For example, see Chen et al., J.Mol.Biol.293:865-881 (1999). The coupling rate obtained by the above surface plasmon resonance assay was 10 6 M -1 s -1If it exceeds this, the binding rate can be determined by using fluorescence quenching techniques to measure the increase or decrease in the fluorescence emission intensity (excitation = 295 nM, emission = 340 nM, 16 nM band-pass) of a 20 nM anti-antigen antibody (Fab type) in PBS at pH 7.2 at 25°C in the presence of increasing concentrations of antigen, measured with a spectrophotometer, for example, a spectrophotometer with stop flow (Aviv Instruments) or an 8000 series SLM-AMINCO™ spectrophotometer with a stirring cuvette (ThermoSpectronic).

[0070] 2. Antibody fragment In certain embodiments, the antibodies produced by the polypeptide expression systems described herein are antibody fragments. Antibody fragments include, but are not limited to, Fab, Fab′, Fab′-SH, F(ab′)2, Fv, and scFv fragments, as well as other fragments described below. For a review of a particular antibody fragment, see Hudson et al. Nat. Med. 9:129-134 (2003). For a review of scFv fragments, see, for example, Pluckthun, in The Pharmacology of Monoclonal Antibodies, vol. 113, Rosenburg and Moore (eds.), (Springer-Verlag, New York), pp. 269-315 (1994), as well as International Patent Publication WO93 / 16185, and U.S. Patents 5,571,894 and 5,587,458. For consideration of Fab and F(ab′)2 fragments containing salvage receptor-binding epitope residues that increase in vivo half-life, see, for example, U.S. Patent No. 5,869,046.

[0071] A bispecific antibody is an antibody fragment having two antigen-binding sites, which may be bivalent or bispecific. See, for example, European Patent No. 404,097, International Patent Publication No. WO1993 / 01161, Hudson et al., Nat. Med. 9:129-134 (2003), and Hollinger et al., Proc. Natl. Acad. Sci. USA 90:6444-6448 (1993). Trispecific and tetraspecific antibodies are also described in Hudson et al., Nat. Med. 9:129-134 (2003).

[0072] A single-domain antibody is an antibody fragment containing all or part of the heavy chain variable domains or all or part of the light chain variable domains of an antibody. In certain embodiments, the single-domain antibody is a human single-domain antibody (see, for example, Domantis, Inc., Waltham, MA; U.S. Patent No. 6,248,516 B1).

[0073] 3. Chimeric and humanized antibodies In certain embodiments, the antibodies produced by the polypeptide expression systems described herein (e.g., Fab or full-length IgG antibodies) are chimeric antibodies. Certain chimeric antibodies are described, for example, in U.S. Patent No. 4,816,567 and Morrison et al., Proc. Natl. Acad. Sci. USA, 81:6851-6855 (1984). In one example, a chimeric antibody includes a non-human variable region (e.g., a variable region derived from mouse, rat, hamster, rabbit, or non-human primate, e.g., monkey) and a human constant region. In further examples, a chimeric antibody is a “class-transverted” antibody in which the class or subclass has changed from that of the parent antibody. Chimeric antibodies include their antigen-binding fragments.

[0074] In certain embodiments, a chimeric antibody is a humanized antibody. Typically, a non-human antibody is humanized to reduce its immunogenicity to humans while maintaining the specificity and affinity of the parent non-human antibody. Generally, a humanized antibody contains one or more variable domains in which the HVR, e.g., CDR (or portions thereof), is derived from the non-human antibody and the FR (or portions thereof) is derived from the human antibody sequence. The humanized antibody may optionally also contain at least a portion of the human constant region. In some embodiments, some FR residues in the humanized antibody are replaced with corresponding residues from the non-human antibody (e.g., the antibody from which the HVR residues are derived) to restore or improve, for example, the specificity or affinity of the antibody.

[0075] Humanized antibodies and their production methods have been confirmed, for example, in Almagro and Fransson, Biosci. 13:1619-1633 (2008), and in other publications such as Riechmann et al., Nature 332:323-329 (1988), Queen et al., Proc. Nat'l Acad. Sci. USA 86:10029-10033 (1989), U.S. Patents No. 5,821,337, No. 7,527,791, No. 6,982,321, and No. 7,087,409, Kashmiri et al., Methods 36:25-34 (2005) (explaining SDR (a-CDR) grafting), and Padlan, Mol. Immunol. 28:489-498 (1991) ("Surface re- This is further explained in Dall'Acqua et al., Methods36:43-60 (2005) (explaining "FR shuffling"), as well as in Osbourn et al., Methods36:61-68 (2005) and Klimka et al., Br.J.Cancer, 83:252-260 (2000) (explaining the "guided selection" method for FR shuffling).

[0076] Human framework regions that can be used for humanization include framework regions selected using the "best fit" method (see, e.g., Sims et al. J.Immunol.151:2296 (1993)), framework regions derived from consensus sequences of special subgroups of human antibodies in light chain or heavy chain variable regions (see, e.g., Carter et al. Proc.Natl.Acad.Sci.USA,89:4285 (1992) and Presta et al. J.Immunol.,151:2623 (1993)), human mature (somatically mutated) framework regions, or human germline framework regions (e.g., Almagro and This includes, but is not limited to, Fransson, Front. Biosci. 13:1619-1633 (2008), as well as framework areas derived from screening FR libraries (see, for example, Baca et al., J. Biol. Chem. 272:10678-10684 (1997) and Rosok et al., J. Biol. Chem. 271:22611-22618 (1996)).

[0077] 4. Human antibodies In certain embodiments, the antibodies produced by the polypeptide expression systems described herein (e.g., Fab or full-length IgG antibodies) are human antibodies. Human antibodies may be recombinant human antibodies independently prepared using various techniques known in the art, and then having their sequences identified. Human antibodies are generally described in van Dijk and van de Winkel, Curr. Opin. Pharmacol. 5:368-74 (2001) and Lonberg, Curr. Opin. Immunol. 20:450-459 (2008).

[0078] 5. Library-derived antibodies By utilizing the polypeptide expression systems described herein, which are useful in phage display systems, antibodies (e.g., Fab or full-length IgG antibodies) produced by the polypeptide expression system of the present invention can be isolated by screening a combinatorial library for antibodies with desired activity(s). For example, Hoogenboom et al. in Methods in Molecular Biology 178:1-37 (O'Brien et al., ed., Human Press, Totowa, NJ, 2001), and for example, in the McCafferty et al., Nature 348:552-554; Clackson et al., Nature 352:624-628 (1991); Marks et al., J.Mol.Biol.222:581-597 (1992), Marks and Bradbury, in Methods in Molecular Biology 248:161-175 (Lo, ed., Human Press). See Press, Totowa, NJ, 2003; Sidhu et al., J.Mol.Biol.338(2):299-310(2004); Lee et al., J.Mol.Biol.340(5):1073-1093(2004); Fellouse, Proc.Natl.Acad.Sci.USA 101(34):12467-12472(2004); and Lee et al., J.Immunol.Methods 284(1-2):119-132(2004).

[0079] 6. Multispecific antibodies In certain embodiments, the antibodies produced by the polypeptide expression systems described herein (e.g., Fab or full-length IgG antibodies) are multispecific antibodies, such as bispecific antibodies. A multispecific antibody is a monoclonal antibody that has binding specificity to at least two different sites. In certain embodiments, one of the binding specificities is with respect to a first antigen and the other is with respect to any other antigen. In certain embodiments, the bispecific antibody may bind to two different epitopes of the first antigen. Bispecific antibodies can also be used to localize cytotoxicity to cells expressing the first antigen. Bispecific antibodies can be prepared as full-length antibodies or antibody fragments.

[0080] This specification also includes engineered antibodies having three or more functional antigen-binding sites, including "octopus antibodies" (see, for example, U.S. Patent No. 2006 / 0025576A1).

[0081] The antibodies or fragments described herein also include “dual-acting FAbs” or “DAFs” that include an antigen-binding site that binds to a first antigen as well as another different antigen (see, for example, U.S. Patent No. 2008 / 0069820).

[0082] 7. Antibody variants In certain embodiments, amino acid sequence variants of antibodies provided herein are considered. For example, it may be desirable to improve the binding affinity and / or other biological properties of the antibody. Amino acid sequence variants of antibodies can be prepared by introducing appropriate modifications to one or more of the nucleic acid molecular sequences encoding all or part of the antibody. Such modifications include, for example, the deletion of residues from the amino acid sequence of the antibody, and / or the insertion of residues into the amino acid sequence, and / or the substitution of residues within the amino acid sequence. Any combination of deletions, insertions, and substitutions may be prepared to achieve the final construct, provided that the final construct possesses the desired characteristics, such as antigen-binding ability.

[0083] In certain embodiments, a collection of antibody variants having one or more amino acid substitutions relative to each other can be generated by the expression system and method of the present invention. Target sites for substitutional mutagenesis include HVR and FR. Conservative substitutions are shown under the heading "Conservative Substitutions" in Table 1. More substantial modifications are provided under the heading "Typical Substitutions" in Table 1 and are further described below in relation to amino acid side chain classes. Amino acid substitutions can be introduced into the target antibody and the screened product with respect to desired activity, e.g., maintained / improved antigen binding, reduced immunogenicity, or improved ADCC or CDC. TIFF0007863949000001.tif208170

[0084] Amino acids can be classified according to their common side-chain properties: (1) Hydrophobic: norleucine, Met, Ala, Val, Leu, Ile; (2) Neutral hydrophilic: Cys, Ser, Thr, Asn, Gln; (3) Acidic: Asp, Glu; (4) Basicity: His, Lys, Arg; (5) Residues that affect chain orientation: Gly, Pro; (6) Aromatic: Trp, Tyr, Phe

[0085] Non-conservative substitution would involve swapping one member of one of these classes with one of another.

[0086] One type of substitution mutant involves substituting one or more hypervariable region residues of a parent antibody (e.g., a humanized antibody or a human antibody). Generally, the resulting mutant(s) selected for further study will have modifications (e.g., improvements) to certain biological properties of the parent antibody (e.g., increased affinity, decreased immunogenicity) and / or substantially maintain certain biological properties of the parent antibody. A typical substitution mutant is an affinity-matured antibody, which can be readily generated using phage display-based affinity maturation methods, such as those described herein. In short, one or more HVR residues are mutated, the mutant antibody is displayed on a phage, and screened for specific biological activity (e.g., binding affinity).

[0087] Modifications (e.g., substitutions) may be made in HVR to improve antibody affinity, for example. Such modifications may be made in HVR "hot spots," i.e., residues encoded by codons that are frequently mutated during the somatic cell maturation process (see, e.g., Chowdhury, PS, Methods Mol. Biol. 207:179-196 (2008)), and / or in SDR (a-CDR), and the resulting mutant VH or VL is tested for binding affinity. Affinity maturation methods, which involve constructing secondary libraries and re-selecting from them, are described, for example, in Hoogenboom, HR et al. in Methods in Molecular Biology 178 1-37 (2001) (O'Brien et al., Human Press, Totowa, NJ). In some embodiments of affinity maturation methods, diversity is introduced into the variable region genes selected for maturation by one of various methods (e.g., error-prone PCR, chain shuffling, or oligonucleotide-directed mutagenesis). A secondary library is then created. The library is then screened to identify any antibody variant with the desired affinity. Another method for introducing diversity involves an HVR-directed method, in which several HVR residues (e.g., 4-6 residues at a time) are randomized. The HVR residues associated with antigen binding can be specifically identified, for example, using alanine scanning mutagenesis or modeling. CDR-H3 and CDR-L3 are often targeted in particular.

[0088] In certain embodiments, substitutions, insertions, or deletions may occur within one or more HVRs, provided that such modifications do not substantially reduce the antibody's ability to bind the antigen. For example, conservative modifications that do not substantially reduce binding affinity (e.g., conservative substitutions provided herein) may be made within an HVR. Such modifications may be located outside of HVR "hot spots" or SDRs. In certain embodiments of the variant VH and VL sequences provided above, each HVR is either unchanged or does not contain one, two, or more than three amino acid substitutions.

[0089] A useful method for identifying antibody residues or regions that can be targeted for mutagenesis is called "alanine scanning mutagenesis," as described by Cunningham and Wells (1989) Science, 244:1081-1085. In this method, residues or groups of target residues (e.g., charged residues such as arg, asp, his, lys, and glu) are identified and replaced with neutral or negatively charged amino acids (e.g., alanine or polyalanine) to determine whether the antibody-antigen interaction has been affected. Further substitutions may be introduced at amino acid positions that are functionally sensitive to the initial substitution. Alternatively, or in addition, the crystal structure of the antigen-antibody complex may be used to identify contact sites between the antibody and antigen. Such contact residues and adjacent residues may be targeted as candidates for substitution or excluded. Mutants may be screened to determine whether they contain the desired properties.

[0090] Amino acid sequence insertions include amino-terminal and / or carboxyl-terminal fusions extending to polypeptides containing 1 to 100 or more residues in length, as well as intersequential insertions of single or multiple amino acid residues. Examples of terminal insertions include antibodies with an N-terminal methionyl residue. Other insertion variants of antibody molecules include fusion to the N-terminus or C-terminus of an antibody to an enzyme (e.g., for ADEPT) or polypeptide that increases the serum half-life of the antibody.

[0091] While the concept of modular protein expression by premRNA trans-splicing is described in detail herein in relation to phage antibody display vector systems, the application of the concept exemplified by the use of nucleic acid molecules, vectors, vector sets, host cells, and methods described herein can be adapted and extended to other techniques that require the expression of many collections of proteins having different combinations of repeating modules in mammalian cells, for example.

[0092] Examples The following are embodiments of the present invention. It should be understood that various other embodiments may be practiced based on the general description provided above.

[0093] Example 1. Generation of a modular protein expression system for phage display vectors and associated antibody reformatting. The creation of polypeptide expression systems for modular expression and generation of polypeptides is described. This invention is at least in part based on experimental findings demonstrating that pre-mRNA trans-splicing can be utilized in mammalian cells to enable modular recombinant protein expression. The concept of modular protein expression makes it possible to precisely combine any two protein-coding sequences encoded by two different constructs into a single mRNA encoding a polypeptide chain, without any of the requirements and constraints of other protein-protein splicing methods. The concept of modular protein expression by pre-mRNA trans-splicing can be adapted to simplify and extend other techniques that require the expression of many collections of proteins with different combinations of repeating modules in mammalian cells. For example, this concept would find applications in other settings requiring the expression of fusion protein partners or combinations of mutations in a single polypeptide. This technique is both simple and effective, making it applicable at any scale and giving it broad importance to the field of recombinant protein expression in mammalian cells, which is the basis of much of modern biotechnology.

[0094] This section describes the creation of polypeptide expression systems that enable the modular expression of different antibody formats in relation to phage display expression systems. Phage display is widely used in the discovery and engineering of antibody fragments in relation to the development of therapeutic and reagent antibodies (McCafferty et al. Nature. 348:552-554, 1990; Sidhu. Current opinion in biotechnology. 11:610-616, 2000; Smith. Science. 228:1315-1317, 1985). While phage display has traditionally enabled rapid selection of antigen-specific conjugates, it has limited the screening of selected antibody fragments. Detailed characterization of antibody fragments often requires the expression of full-length immunoglobulin G (IgG), which is normally expressed in mammalian cells. However, one limiting step in this process is reformatting phage clones into mammalian expression vectors for IgG expression. High-throughput subcloning methods can be used to reformat a large number of clones, but these methods are typically relatively labor-intensive and produce many clones that will not be used beyond the screening stage.

[0095] To avoid the need for subcloning and enable modular protein expression, a first nucleic acid molecule: a dual-host vector, pDV2, was created (Figure 1). Unlike the aforementioned dual-vector pDV (Tesar et al., Protein engineering, design & selection: PEDS. 26: 655-662, 2013), which contains an IgG expression cassette with an engineered signal sequence for heavy chain expression in either bacterial or mammalian cells and requires simultaneous transfection of mammalian cells with a mammalian expression vector expressing the light chain for full IgG expression, pDV2 contains most of the stII signal sequence embedded in the intron, which is removed by splicing in the bacterial promoter and mammalian cells.

[0096] The stII signal sequence in pDV2 was modified to include both the 3′ splice site (3′ss) and an optimized polypyrimidine region (PPT) prior to the 3′ss. This required introducing three relatively conserved amino acid substitutions into the stII signal sequence, which did not affect the display of the Fab fragment on the phage (Figure 2). To enable modular and flexible expression of antibody formats from the same clone, we did not add complete introns and exons encoding the constant region downstream from the region encoding the CH1 domain. Instead, we attempted to add these heavy chain sequences trans from a second nucleic acid molecule. To achieve this, we leveraged the process of premRNA trans-splicing, which combines two different premRNAs to form a single mature mRNA. By hybridizing complementary sequences downstream from 5′ss and upstream from 3′ss, trans-splicing in mammalian cells can be induced to form a single non-covalent premRNA by yielding premRNAs containing both 5′ss and 3′ss sequences, which can then be spliced ​​as a regular premRNA (Konarska et al. Cell. 42:165-171, 1985; Puttaraju et al. Nature biotechnology. 17:246-252, 1999; Solnick. Cell. 42:157-164, 1985). In this special polypeptide expression system, a 150-bp fragment of the M13 gene III (gIII) was used as the hybridizing sequence (Figure 1). This gene III sequence follows the aforementioned optimized GTAAGA 5′ss at the 3′ boundary of the sequence encoding CH1 (Tesar et al., Protein Engineering, design & selection: PEDS. 26: 655-662, 2013).

[0097] To complete the polypeptide expression system, a second nucleic acid molecule, pRK-Fc, was generated, which is a complementary plasmid expressing a premRNA containing a linker sequence, consensus branching point, and a 150-nt antisense gene III sequence followed by the PPT, as well as a hinge, CH2 and CH3 regions in one exon, and a 3′ss followed by an SV40 polyadenylation signal (Figures 1 and 3). This transcript does not encode a signal sequence, and the first two potential start codons are located outside the frame, in the antisense gene III sequence and in the hinge region. Thus, with the exception of the 5′ss, all other sequences required for splicing are encoded by pRK-Fc rather than pDV2. Simultaneous transfection of Expi293F cells (in vitrogen) with pDV2 and pRK-Fc resulted in baseline but detectable levels of IgG expression (Figure 4A).

[0098] Example 2. Generation of an optimized modular protein expression system for reformatting phage display vectors and associated antibodies. The baseline IgG yields achieved with pDV2 and pRK-Fc may be due to the absence of sequences required for efficient trans-splicing or sequences in the vector that inhibit trans-splicing. Nucleotide motifs in both exons and introns can act as splicing enhancers, suppressors, or both, depending on their position. In terms of vector design objectives, intron splice enhancers (ISEs) can be easily added because they are unlikely to affect the coding sequence in mammalian cell expression. One well-described ISE consists of a sequence of three or more consecutive guanine residues or G-runs located close to the intron boundary, which are bound by heteronuclear ribonucleoprotein H or F to enhance splicing (Wang et al. Nature structural & molecular biology. 19:1044-1052, 2012; Xiao et al. Nature structural & molecular biology. 16:1094-1100, 2009). In addition, it has been shown that purine-rich intron sequences adjacent to 5′ss, not limited to G-run, also enhance splicing (Hastings et al. RNA.7:859-874, 2001).

[0099] Therefore, we created variants of pDV2 and pDV2b that contain a 9-nt G-run in the linker-coding region between the upper hinge and the C-terminus of the M13 bacteriophage pIII coat protein (cP3), and a 10-nt second 4-nt G-run downstream, along with a 26-bp purine-rich region of 23 base pairs (bp) downstream from 5′ss (Figure 5). This variant alters the Gly-Arg-Pro linker between the upper hinge and cP3 to three Gly residues. The vector did not contain a standard polyadenylation site for the heavy chain cassette. This was done to minimize the formation of mature heavy chain mRNA from the vector, which would then be transported to the cytoplasm, undergo trans-splicing, and potentially trigger the expression of the Fab-cP3 fusion protein. The pRK-Fc molecule was also optimized. The intronic G-run near 3′ss has been shown to stimulate splicing in vitro (Martinez-Contreras. PLoS biology. 4:e21, 2006). Therefore, an optimized complementary plasmid pRK-Fc2 was generated by adding a 9-nt ISE upstream from the branching site (Figure 6). Co-transfection of human Expi293F cells with pDV2 and pRK-Fc(ISE-) or pRK-Fc2(ISE+) resulted in baseline-level IgG expression (Figure 4A). Co-transfection of Expi293F cells with the ISE+pDV2b plasmid and pRK-Fc or pRK-Fc2 resulted in higher levels of IgG expression, and the highest expression level of up to 25 μg / ml produced by co-transfection of the ISE+ plasmid pDV2b and pRK-Fc2 indicates that the ISE sequence in both transcripts enhances the efficiency of trans-splicing.

[0100] Baseline IgG expression levels in transfected Expi293F cells were associated with apparent cell lysis 7 days after transfection, and this was observed when pDV2 or pDV2b were transfected alone, rather than pRK-Fc or pRK-Fc2. Analysis of transfected cell lysates by Western blotting with anti-M13p3 antibody revealed a polypeptide with an apparent molecular weight of approximately 41 kDa, consistent with the expression of an IgG1Fd fragment (VH-CH1-upper hinge) fused to the M13cP3 peptide (Figure 12, bottom panel, columns 3-6). Expression of this polypeptide was higher in cells transfected with pDV2 or pDV2b without complementary plasmids. The results indicate that both pDV2 and pDV2b plasmids can express mature mRNA encoding a potentially toxic product, despite the fact that both plasmids lack mammalian polyadenylation sites downstream of the vector from the heavy chain cassette.

[0101] Visual inspection of the gene III sequence encoding cP3 revealed an AATAAA motif capable of acting as a polyadenylation site (Figure 2). Two silent mutations were introduced at this site to generate plasmids pDV2c(ISE-) and pDV2d(ISE+) to test whether this reduces toxicity and improves protein expression in mammalian cells. Co-transfection of Expi293F cells with pRK-Fc and either pDV2c or pDV2d resulted in approximately 6-fold higher IgG expression compared to pDV2 and pDV2b vectors due to the potential polyadenylation site in gene III (Figure 4A). This increased IgG expression level was associated with higher viability of transfected cells and significantly reduced or undetectable expression of the Fd-cP3 fusion protein within the transfected cells (Figure 12, lower panel, columns 7-10). This indicates that the presence of a potential polyadenylation site in the donor vector within gene III causes unwanted protein expression from the donor plasmid alone, significantly negatively impacting protein expression. Co-transfection of Expi293F cells with pDV2c or pDV2d and the ISE+pRK-Fc2 complementary vector resulted in a further twofold increase in IgG expression compared to co-transfection with the ISE-pRK-Fc vector (Figure 4A). These results suggest that the primary factor determining baseline protein expression in the pDV2 vector is the presence of a potential polyadenylation site in gene III, while the addition of ISE has only a minor effect on protein expression when the potential gene III polyadenylation motif is absent. In contrast, the addition of ISE in the complementary pRK-Fc2 plasmid results in approximately twofold higher IgG yield when co-transfecting pDV2 mutants without a potential polyadenylation site in gene III (Figure 4A).

[0102] Further optimization of protein expression was achieved by determining the optimal DNA ratio for transfection. Using a 2:1 excess complementary plasmid pRK-Fc2 to pDV2d yielded the highest IgG expression yield in this system (Figure 4B). Using pDV2d and pRK-Fc2 with the optimized DNA ratio, the yield of purified IgG from 30 ml of supernatant of transfected Expi293F cells was 3.2 ± 1.2 mg (n=3). Purified IgG from Expi293F cells co-transfected with these plasmids was indistinguishable from the same IgG expressed by conventional expression vectors by mass spectrometry and SDS-PAGE (Figures 7A-7B and 13). Co-transfection of Expi293F cells with pDV2d, which encodes a variable region with different specificity, and pRK-Fc2, which has an optimized DNA ratio, resulted in high IgG expression of 2.5–5.5 mg of purified IgG from 30 ml of supernatant of transfected Expi293F cells (Figure 8A). Polypeptide expression systems are not limited to the use of Expi293F cells to achieve high expression levels. Other mammalian cell lines widely used for IgG expression, such as 293T and CHO cells, were also effective. Co-transfection of 293T or CHO cells with pDV2d and pRK-Fc2, which express variable regions with different specificity, resulted in high IgG expression (Figure 8B).

[0103] The pRK-Fc2 vector was co-transfected with the pDV2 plasmid and modified for Fab fragment expression. Sequences encoding lower hinges and Fc regions in pRK-Fc2 were removed and replaced with a Flag tag to produce the pRK-Fab-Flag vector (Figure 9). The yield of purified Fab fragments from 30 ml of supernatant of Expi293F cells co-transfected with pDV2d and pRK-Fab-Flag was 0.8 ± 0.06 mg (mean ± standard deviation, n=3). The structural precision of the purified Flag-tagged Fab fragments was confirmed by mass spectrometry and SDS-PAGE (Figure 13). The observed heavy chain mass, excluding the clipped C-terminal lysine, was 25,169 Da, close to the expected mass of 25,172 Da.

[0104] Expression of N-terminal cleaved proteins from complementary transcripts has been observed in transsplicing systems for gene therapy (Monjaret et al. Molecular Therapy 22:1176-1187, 2014). This may be due to translation from the internal start codon by a complementary transcript encoding a 3′ exon containing all the elements necessary for the formation of mature mRNA. Western blotting of lysates from cells transfected with pRK-Fc2 revealed the expression of a polypeptide consistent with the Fc fragment translated from the first in-frame ATG codon (Figure 12, top panel, column 11). This polypeptide likely lacks a secretory signaling sequence and should be expressed only in the cytoplasm. Although this product may be released into the culture medium upon cell lysis, it was not observed in purified IgG samples by SDS-PAGE (Figure 13, column 2) and mass spectrometry. When pDV2c or pDV2d was co-transfected into cells, the expression of this cleaved product was reduced but not eliminated (Figure 12, top panel, columns 8 and 10). Insertion of an out-of-frame open reading frame with an optimal translation initiation site upstream of the intron region from a potential Fc start codon did not significantly reduce the expression of the cleaved Fc product.

[0105] A key characteristic of phage display vectors that determines selection efficiency is the level of antibody fragment display achieved on the phage particles. Using the aforementioned Amber-2614KO7 helper phage, along with reduced p3 expression in the E. coli SupE suppressor strain, the level of Fab fragment display achieved with the pDV2d vector was comparable to the Fab display level achieved with the specialized Fab display vector, Fab-Zip-Phage, using the standard M13KO7 helper phage (Figure 14).

[0106] When different antibody formats, such as wild-type IgG, Fab fragments, or IgG with Fc modifications for bispecificity formats, are required for different screening assays, the ability to express different antibody formats from the same clone is useful in antibody discovery. The vector set, in principle, enables any of these or additional formats simply by cloning a suitable sequence to be added after the CH1 region in the complementary plasmid. Furthermore, the modular structure of this system enables the expression of new antibody formats without requiring the re-creation of a stock of phage display libraries, as it only requires the construction of a novel complementary plasmid. The bi-vector can also be adapted to allow the use of any CH1 region by transferring 5′ss from downstream of the CH1 coding region to the VH or J-CH1 junction J region (FR4), thus separating the entire constant regions of the VH and heavy chain in two different plasmids. With the knowledge that amber stop codons at the junction of the heavy chain and gene III sequence in Fab phage display libraries typically result in significantly lower levels of display, and therefore require clonal reformatting after selection, at least in the case of inexperienced repertoire libraries, the pDV2 vector can be adapted to conventional methods for expressing Fab fragments in E. coli by simply adding a stop codon after the sequence encoding the upper hinge (Lee et al. Journal of immunological methods. 284:119-132, 2004). Expression of Fab fragments in mammalian cells using the same methods used for IgG expression avoids this need for reformatting with yields comparable to those typically obtained in E. coli.

[0107] Example 3. Modular protein expression system The polypeptide expression systems generated and characterized in Examples 1 and 2 demonstrate that modular and flexible polypeptide expression of any desired protein can be directly achieved by using polypeptide expression systems such as the optimized expression systems described above with respect to protein reformatting in relation to phage display. Therefore, this expression system comprises two nucleic acid molecular components (polypeptide coding sequences PES11 and PES2) each encoding a portion of a single desired polypeptide product, where the split-coding regions of these proteins are precisely linked in vivo by premRNA trans-splicing without requiring subcloning of the protein-coding nucleic acid. As shown in Figure 10, the first nucleic acid molecule comprises an expression cassette having PES11, with a eukaryotic promoter (P1) upstream of the PES11 component. Euk1 ) and the eukaryotic signal sequence (ESS11), as well as the 5′ splice site (5′ss11) and hybridize sequence (HS1) located downstream of PES11. The complementary second nucleic acid molecule is the eukaryotic promoter (P2 Euk The first premRNA molecule will contain a hybridization sequence that can hybridize to HS1 (HS2), and a 3′ splice site (3′ss2) upstream of the PES2 component. In addition, the second nucleic acid molecule will contain a polyadenylation site (pA2) downstream of the PES2 component. Thus, when copied in mammalian cells, the two generated premRNA molecules, one with a single 5′ and the other with a single 3′, will be oriented together by their complementary hybridization sequences (HS1 and HS2) and undergo trans-splicing to form a single continuous mRNA that can subsequently translate and encode the desired protein product.

[0108] In prokaryotic cells, if the expression of polypeptide products encoded by PES11 and optionally by the HS1 region is also desirable, the first nucleic acid molecule is P1 Euk1 It may further include a resectable prokaryotic promoter module (ePPM1) positioned between and PES11. ePPM1 includes a 5′ splice site (5′ss12) and a prokaryotic promoter (P1 Prok1), may include a first nucleic acid sequence encoding a prokaryotic signal sequence (PSS11), and a 3′ splice site (3′ss11), and 5′ss12-P1 Prok1 -PSS11-3′ss11 are operably linked to each other in the 5′-3′ direction. ePPM1 will drive the transcription of the polypeptide encoding the first nucleic acid molecule in prokaryotic cells. On the other hand, in eukaryotic cells (e.g., mammalian cells), P1 Euk1 ePPM1 drives the transcription and expression of the polypeptide encoding the first nucleic acid molecule, and ePPM1 will be removed from the premRNA transcript by cis-splicing due to collisions with the 5′ss12 and 3′ss11 components.

[0109] In some cases, it may be desirable to express a second polypeptide. Therefore, the first nucleic acid molecule in a modular protein expression system may be designed to include a second expression cassette. As shown in Figure 11, the second expression cassette encoding a second protein product (PES12) is designed in a similar manner to the first expression cassette, but will include a polyadenylation site (pA1) downstream of the PES12 sequence to ensure the generation of a separate premRNA molecule after transcription. In other cases, the second expression cassette can be designed within the second nucleic acid molecule of the polypeptide expression system.

[0110] Other Embodiments Although the invention described herein is explained in detail by means of examples and embodiments for the purpose of clear understanding, the explanation and embodiments should not be construed as limiting the scope of the invention. All patent and scientific literature disclosures described herein are explicitly incorporated in their entirety by reference. [Embodiment 1] A polypeptide expression system comprising a first nucleic acid molecule and a second nucleic acid molecule, (a) The first nucleic acid molecule comprises the following components: (i) the first eukaryotic promoter (P1 Euk1 (ii) the first polypeptide coding sequence (PES1 1 ), (iii) the first 5' splice site (5'ss1 1 (iv) comprising a first expression cassette comprising a hybridized sequence (HS1), wherein the components are P1 Euk1 -PES1 1 -5'ss1 1 -HS1 is operably connected to each other in the 5'~3' direction. (b) The second nucleic acid molecule comprises the following components: (i) eukaryotic promoter (P2 Euk (ii) a hybridization sequence (HS2) that can hybridize to HS1, (iii) a 3' splice site (3'ss2), (iv) a polypeptide coding sequence (PES2), and (v) a polyadenylation site (pA2), wherein the constituent elements are P2 Euk The polypeptide expression system, in which the polypeptides are operably linked to each other in the 5'~3' direction as -HS2-3'ss2-PES2-pA2. [Embodiment 2] The P1 Euk1 The polypeptide expression system according to Embodiment 1, wherein the promoter is a cytomegalovirus (CMV) promoter or a simian virus 40 (SV40) promoter. [Embodiment 3] The P2 Euk The polypeptide expression system according to Embodiment 1 or 2, wherein the promoter is either a CMV promoter or an SV40 promoter. [Embodiment 4] The first expression cassette is a eukaryotic signal sequence (ESS1 1 The ESS1 further comprises a first nucleic acid sequence encoding ). 1 However, P1 Euk1 and the aforementioned PES1 1 A polypeptide expression system according to any one of Embodiments 1 to 3, positioned between the two. [Embodiment 5] The ESS1 1 However, the polypeptide expression system described in Embodiment 4 is derived from a variable heavy chain (VH) gene. [Embodiment 6] The first expression cassette comprises the following components: (i) 5' splice site (5'ss1 2 ), (ii) Prokaryotic promoter (P1 Prok1 ), and (iii) 3' splice site (3's splice site 1 ) including the resectable prokaryotic promoter module (ePPM) 1 ) further includes, and the said components are 5'ss1 2 -P1 Prok1 -3'ss1 1 They are operably connected to each other in the 5'~3' direction, and the ePPM 1 However, P1 Euk1 and the aforementioned PES1 1 A polypeptide expression system according to any one of Embodiments 1 to 5, positioned between the two. [Embodiment 7] The P1 Prok1 However, the polypeptide expression system according to Embodiment 6 is selected from the group consisting of the PhoA promoter, the Tac promoter, the Lac promoter, and the Tphac promoter. [Embodiment 8] The ePPM 1 However, prokaryotic signal sequence (PSS1 1 A polypeptide expression system according to embodiment 6 or 7, further comprising a first nucleic acid sequence encoding ). [Embodiment 9] The PSS1 1 However, the polypeptide expression system according to any one of Embodiments 6 to 8, derived from the heat-stable enterotoxin II (stII) gene. [Embodiment 10] The PSS1 1 and the aforementioned 3'ss1 1 The polypyrimidine region (PPT1) is located between these two regions. 1 A polypeptide expression system according to any one of embodiments 6 to 9, further comprising ). [Embodiment 11] The PPT1 1 However, the polypeptide expression system according to Embodiment 10 includes the nucleic acid sequence TTCCTTTTTTCTCTTTCC (SEQ ID NO: 1). [Embodiment 12] The PES1 1 However, the polypeptide expression system according to any one of Embodiments 1 to 11, which does not contain a latent 5' splice site. [Embodiment 13] The polypeptide expression system according to any one of Embodiments 1 to 12, wherein HS1 is a gene encoding all or part of a coat protein or adapter protein. [Embodiment 14] The polypeptide expression system according to Embodiment 13, wherein the coat protein is selected from the group consisting of pI, pII, pIII, pIV, pV, pVI, pVII, pVIII, pIX, and pX of bacteriophage M13, f1, or fd. [Embodiment 15] The polypeptide expression system according to Embodiment 14, wherein the coat protein is the pIII protein of bacteriophage M13. [Embodiment 16] The polypeptide expression system according to Embodiment 15, wherein the pIII fragment comprises amino acid residues 267-421 of the pIII protein or amino acid residues 262-418 of the pIII protein. [Embodiment 17] The polypeptide expression system according to Embodiment 13, wherein the adapter protein is a leucine zipper. [Embodiment 18] The polypeptide expression system according to Embodiment 17, wherein the leucine zipper comprises the amino acid sequence of SEQ ID NO: 4 or 5. [Embodiment 19] The first nucleic acid molecule is a second eukaryotic promoter (P1 Euk2 (ii) second polypeptide coding sequence (PES1 2 ), and (iii) a second expression cassette comprising a polyadenylation site (pA1), wherein the component is P1Euk2 -PES1 2 A polypeptide expression system according to any one of embodiments 1 to 18, wherein the polypeptides are operably linked to each other in the 5'-3' direction as -pA1. [Embodiment 20] The P1 Euk2 However, the polypeptide expression system according to Embodiment 19 is a CMV promoter or an SV40 promoter. [Embodiment 21] The second expression cassette is a eukaryotic signal sequence (ESS1 2 A polypeptide expression system according to embodiment 19 or 20, further comprising a nucleic acid sequence encoding ). [Embodiment 22] The ESS1 2 However, the polypeptide expression system according to Embodiment 21 is derived from the mouse-conjugated immunoglobulin protein (mBiP) gene. [Embodiment 23] The ESS1 2 The polypeptide expression system according to any one of Embodiments 19 to 22, wherein the nucleic acid sequence (SEQ ID NO: 6) is ATG AAN TTN ACN GTN GTN GCN GCN GCN CTN CTN CTN CTN GGN, and in the sequence, N is A, T, C, or G. [Embodiment 24] The second expression cassette comprises the following components: (i) 5' splice site (5'ss1 3 ), (ii) Prokaryotic promoter (P1 Prok2 ), and (iii) 3' splice site (3's splice site 2 ) including the resectable prokaryotic promoter module (ePPM) 2 ) further includes, and the said components are 5'ss1 3 -P1 Prok2 -3'ss1 2 They are operably connected to each other in the 5'~3' direction, and the ePPM 2 However, P1 Euk2 and the aforementioned PES1 2 A polypeptide expression system according to any one of embodiments 19 to 23, positioned between the above. [Embodiment 25] The P1 Prok2 However, the polypeptide expression system according to Embodiment 24 is selected from the group consisting of a PhoA promoter, a Tac promoter, and a Lac promoter. [Embodiment 26] The ePPM 2 However, prokaryotic signal sequence (PSS1 2 A polypeptide expression system according to embodiment 24 or 25, further comprising a nucleic acid sequence encoding ). [Embodiment 27] The PSS1 2 However, the polypeptide expression system according to any one of embodiments 24 to 26, derived from the heat-stable enterotoxin II (stII) gene. [Embodiment 28] The PSS1 2 and the aforementioned 3'ss1 2 The polypyrimidine region (PPT1) is located between these two regions. 2 A polypeptide expression system according to any one of embodiments 24 to 27, further comprising ). [Embodiment 29] The PPT1 2 However, the polypeptide expression system according to Embodiment 28 includes the nucleic acid sequence TTCCTTTTTTCTCTTTCC (Sequence ID 1). [Embodiment 30] The polypeptide expression system according to any one of Embodiments 19 to 29, wherein the second expression cassette is positioned 5' relative to the first expression cassette. [Embodiment 31] The 5'ss1 1 A polypeptide expression system according to any one of embodiments 1 to 30, further comprising an intron splice enhancer (ISE) (ISE1) positioned between and the HS1. [Embodiment 32] The polypeptide expression system according to Embodiment 31, wherein the ISE1 comprises a G-run containing three or more consecutive guanine residues. [Embodiment 33] The polypeptide expression system according to Embodiment 32, wherein the ISE1 comprises a G-run containing nine consecutive guanine residues. [Embodiment 34] A polypeptide expression system according to any one of Embodiments 1 to 33, further comprising a polypyrimidine region (PPT2) located between the HS2 and the 3'ss2. [Embodiment 35] The polypeptide expression system according to Embodiment 34, wherein the PPT2 contains the nucleic acid sequence TTCCTCTTTCCCTTTCTCTCC (SEQ ID NO: 7). [Embodiment 36] The polypeptide expression system according to Embodiment 35, further comprising ISE(ISE2) positioned between HS2 and 3'ss2. [Embodiment 37] The polypeptide expression system according to Embodiment 36, wherein the ISE2 comprises a G-run containing three or more consecutive guanine residues. [Embodiment 38] The polypeptide expression system according to Embodiment 37, wherein the ISE2 comprises a G-run containing nine consecutive guanine residues. [Embodiment 39] The 5'ss1 1 However, the polypeptide expression system according to any one of Embodiments 1 to 38, comprising the nucleic acid sequence of GTAAGA (SEQ ID NO: 8). [Embodiment 40] A polypeptide expression system according to any one of Embodiments 1 to 39, wherein expression by a eukaryotic promoter occurs in mammalian cells. [Embodiment 41] The polypeptide expression system according to Embodiment 40, wherein the mammalian cells are Expi293F cells, CHO cells, 293T cells, or NSO cells. [Embodiment 42] The polypeptide expression system according to Embodiment 41, wherein the mammalian cells are Expi293F cells. [Embodiment 43] A polypeptide expression system according to any one of Embodiments 6 to 42, wherein expression by a prokaryotic promoter occurs within a bacterial cell. [Embodiment 44] The polypeptide expression system according to Embodiment 43, wherein the bacterial cell is an Escherichia coli cell. [Embodiment 45] The PES1 1 However, a polypeptide expression system according to any one of Embodiments 1 to 44, which encodes all or part of the antibody. [Embodiment 46] The PES1 1 However, the polypeptide expression system according to Embodiment 45 encodes a polypeptide containing a VH domain. [Embodiment 47] The polypeptide expression system according to Embodiment 46, wherein the polypeptide further comprises a CH1 domain. [Embodiment 48] The polypeptide expression system according to any one of Embodiments 45 to 47, wherein the PES2 encodes all or part of the antibody. [Embodiment 49] The polypeptide expression system according to Embodiment 48, wherein PES2 encodes a polypeptide including a CH2 domain and a CH3 domain. [Embodiment 50] The PES1 2 However, a polypeptide expression system according to any one of embodiments 19 to 49, which encodes all or part of the antibody. [Embodiment 51] The PES1 2 The polypeptide expression system according to Embodiment 50, wherein the polypeptide encodes a polypeptide containing a VL domain and a CL domain. [Embodiment 52] The following components: (a) First eukaryotic promoter (P1 Euk1 )and, (b) The following components: (i) 5' splice area (5'ss1 2 )、 (ii) Prokaryotic promoter (P1 Prok1 ), and (iii) 3' splice site (3'ss1 1 ) The first resectable prokaryotic promoter module (ePPM) includes 1 ) and the ePPM 1 The aforementioned component is 5'ss1 2 -P1 Prok1 -3'ss1 1 A first excisable prokaryotic promoter module is operably linked to each other in the 5'-3' direction, (c) First polypeptide coding sequence (PES1 1 )and, (d) First 5' splice site (5's splice 1 1 )and, (e) Useful peptide coding sequence (UPES), A first expression cassette comprising, wherein the components of the first expression cassette are P1 Euk1 -ePPM 1 -PES1 1 -5'ss1 1 - The nucleic acid molecules, which are operably linked to each other in the 5'-3' direction as UPES. [Embodiment 53] The first expression cassette is a eukaryotic signal sequence (ESS1 1 The ESS1 further comprises a first nucleic acid sequence encoding ). 1 However, P1 Euk1 and the aforementioned ePPM 1 A nucleic acid molecule according to embodiment 52, positioned between the two. [Embodiment 54] The ePPM 1 However, prokaryotic signal sequence (PSS1 1 The PSS1 further comprises a first nucleic acid sequence encoding ). 1 However, P1 Prok1 and the aforementioned 3'ss1 1 A nucleic acid molecule according to embodiment 52 or 53, positioned between the two. [Embodiment 55] Second eukaryotic promoter (P1 Euk2 (ii) second polypeptide coding sequence (PES1 2 (iii) a second expression cassette comprising a polyadenylation site (pA1), wherein the component is P1 Euk2 -PES1 2 Nucleic acid molecules according to any one of embodiments 52 to 54, which are operably linked to one another in the 5' to 3' directions as -pA1. [Embodiment 56] The second expression cassette is a eukaryotic signal sequence (ESS1 2 The ESS1 further comprises a second nucleic acid sequence encoding ). 2 However, P1 Euk2 and the aforementioned PES1 2 A nucleic acid molecule according to embodiment 55, positioned between the two. [Embodiment 57] The second expression cassette comprises the following components: (i) 5' splice site (5'ss1 3 ), (ii) Prokaryotic promoter (P1 Prok2 ), and (iii) 3' splice site (3's splice site 2 ) including the resectable prokaryotic promoter module (ePPM) 2 ) further includes, and the said components are 5'ss1 3 -P1 Prok2 -3'ss1 2 They are operably connected to each other in the 5'~3' direction, and the ePPM 2 However, P1 Euk2 and the aforementioned PES1 2 A nucleic acid molecule according to embodiment 55 or 56, positioned between the two. [Embodiment 58] The ePPM 2 However, prokaryotic signal sequence (PSS1 2 The PSS1 further includes a nucleic acid sequence encoding 2 However, P1 Prok2 and the aforementioned 3'ss1 2 A nucleic acid molecule according to embodiment 57, positioned between the two. [Embodiment 59] A nucleic acid molecule according to any one of embodiments 52 to 58, wherein the UPES encodes all or part of a useful peptide selected from the group consisting of a tag, label, coat protein, and adapter protein. [Embodiment 60] The nucleic acid molecule according to Embodiment 59, wherein the coat protein is selected from the group consisting of pI, pII, pIII, pIV, pV, pVI, pVII, pVIII, pIX, and pX of bacteriophage M13, f1, or fd. [Embodiment 61] The nucleic acid molecule according to Embodiment 60, wherein the coat protein is the pIII of bacteriophage M13. [Embodiment 62] A vector comprising the nucleic acid molecule described in any one of Embodiments 52 to 61. [Embodiment 63] A vector set comprising a first vector and a second vector, wherein the first and second vectors each comprise the first and second nucleic acid molecules of the polypeptide expression system described in any one of Embodiments 1 to 51. [Embodiment 64] A host cell comprising the vector described in Embodiment 62 or the vector set described in Embodiment 63. [Embodiment 65] The host cell according to Embodiment 64, wherein the host cell is a prokaryotic cell. [Embodiment 66] The host cell according to Embodiment 65, wherein the prokaryotic cell is a bacterial cell. [Embodiment 67] The host cell according to Embodiment 66, wherein the bacterial cell is an Escherichia coli cell. [Embodiment 68] The host cell according to Embodiment 64, wherein the host cell is a eukaryotic cell. [Embodiment 69] The host cell according to Embodiment 68, wherein the eukaryotic cell is a mammalian cell. [Embodiment 70] The host cell according to Embodiment 69, wherein the mammalian cell is an Expi293F cell, a CHO cell, a 293T cell, or an NSO cell. [Embodiment 71] The host cell according to Embodiment 70, wherein the mammalian cell is an Expi293F cell. [Embodiment 72] A method for producing polypeptides, comprising culturing a host cell containing the vector described in Embodiment 62 or the vector set described in Embodiment 63 in a culture medium. [Embodiment 73] The method according to Embodiment 72, further comprising recovering the polypeptide from the host cells or the culture medium.

Claims

1. The following components: (a) First eukaryotic promoter (P1 Euk1 )and, (b) The following components: (i) 5' splice site (5's splice 1 2 ), (ii) Prokaryotic promoter (P1 Prok1 ), and (iii) 3' splice area (3' ss1 1 ) A first excisable prokaryotic promoter module (ePPM) comprising 1 ), wherein said ePPM 1 said components of which are 5'ss1 2 -P1 Prok1 -3'ss1 1 operatively linked to each other in the 5' to 3' direction as, a first excisable prokaryotic promoter module; (c) First polypeptide coding sequence (PES1 1 )and, (d) First 5' splice site (5's splice 1 1 )and, (e) A useful peptide coding sequence (UPES) encoding all or part of the pIII coat protein of the M13 bacteriophage, wherein the portion of the pIII coat protein of the M13 bacteriophage comprises the amino acid residue of SEQ ID NO: 17 or 18, A nucleic acid molecule comprising a first expression cassette, The component of the first expression cassette is P1 Euk1 - ePPM 1 - PES1 1 -5'ss1 1 - As UPES, they are operably connected to each other in the 5' to 3' directions, The first expression cassette is PES1 1 It lacks a polyadenylation site downstream of it. The nucleic acid molecule further comprises a second expression cassette comprising (i) a second eukaryotic promoter (P1 Euk2), (ii) a second polypeptide coding sequence (PES1 2), and (iii) a polyadenylation site (pA1), wherein the components of the second expression cassette are operably linked to each other in the 5'-3' direction as P1 Euk2 - PES1 2 - pA1. Nucleic acid molecule.

2. The first expression cassette contains a eukaryotic signal sequence (ESS1 1 The ESS1 further comprises a first nucleic acid sequence encoding ). 1 However, P1 Euk1 and the aforementioned ePPM 1 A nucleic acid molecule according to claim 1, positioned between the two.

3. The aforementioned ePPM 1 However, prokaryotic signal sequence (PSS1 1 The PSS1 further comprises a first nucleic acid sequence encoding 1 However, P1 Prok1 and the aforementioned 3'ss1 1 A nucleic acid molecule according to claim 1 or 2, positioned between the two.

4. The second expression cassette contains a eukaryotic signal sequence (ESS1 2 It further includes a second nucleic acid sequence that codes for ), The aforementioned ESS1 2 However, P1 Euk2 and the aforementioned PES1 2 A nucleic acid molecule according to claim 1, positioned between the two.

5. The second expression cassette comprises the following components: (i) 5' splice site (5' ss1 3 ), (ii) Prokaryotic promoter (P1 Prok2 ), and (iii) 3' splice portion (3' ss1 2 ) including a resectable prokaryotic promoter module (ePPM) 2 ) further includes, The aforementioned component is 5'ss1 3 -P1 Prok2 -3'ss1 2 They are operably connected to each other in the 5' to 3' directions, and the ePPM 2 However, P1 Euk2 and the aforementioned PES1 2 A nucleic acid molecule according to claim 1, positioned between the two.

6. The aforementioned ePPM 2 However, prokaryotic signal sequence (PSS1 2 The PSS1 further comprises a nucleic acid sequence encoding ) 2 However, P1 Prok2 and the aforementioned 3'ss1 2 A nucleic acid molecule according to claim 5, positioned between the two.

7. A vector comprising the nucleic acid molecule according to any one of claims 1 to 6.

8. A host cell comprising the vector described in claim 7.

9. The aforementioned host cells (a) Prokaryotic cells, or (b) Bacterial cells, preferably E. coli cells, prokaryotic cells, (c) Eukaryotic cells, or (d) Mammalian cells, preferably mammalian cells are Expi293F cells, CHO cells, 293T cells, or NSO cells, eukaryotic cells, or (e) Expi293F cells, which are eukaryotic cells The host cell according to claim 8.

10. A method for producing a polypeptide, comprising culturing a host cell containing the vector described in claim 7 in a culture medium.

11. The method according to claim 10, further comprising recovering the polypeptide from the host cells or the culture medium.