Polypeptide expression systems
The polypeptide expression system addresses the inefficiencies of traditional systems by enabling modular production of recombinant polypeptides through a novel nucleic acid design, enhancing efficiency and reducing construct requirements.
Patent Information
- Application Number
- JP2025066875
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2014-11-17
- Filing Date
- 2025-04-15
- Publication Date
- 2025-08-13
AI Technical Summary
Existing recombinant polypeptide expression systems require numerous constructs for expressing large protein collections with varying combinations of modules, leading to resource-intensive and inefficient production processes.
A polypeptide expression system comprising first and second nucleic acid molecules with specific components operably linked, allowing for modular expression and production of recombinant polypeptides, including eukaryotic and prokaryotic promoters, splice sites, and hybridizing sequences, enabling efficient production in mammalian and bacterial cells.
Facilitates efficient and modular production of recombinant polypeptides, reducing the number of constructs needed and optimizing resource utilization in high-throughput systems.
Smart Images

Figure 2025118659000001_ABST
Abstract
Description
[Technical Field]
[0001] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in ASCII format and is incorporated herein by reference in its entirety. This ASCII copy, created on June 25, 2015, is named P05833-WO_SL.txt and is 24,421 bytes in size.
[0002] The present invention relates to a polypeptide expression system for the modular expression and production of polypeptides. [Background technology]
[0003] Recombinant polypeptides are sometimes expressed as fusions of individual domains or tags for functional or purification purposes. Recombinant DNA methods have traditionally been used to combine sequences encoding each module, requiring a different construct for each combination. This poses a challenge to the technology for expressing large protein collections consisting of repeating modules combined in various combinations, as the number of constructs increases geometrically as a function of the number of modules used.
[0004] While high-throughput systems for subcloning can handle large numbers of inserts in parallel, they are typically resource-intensive and generate large numbers of constructs that are ultimately not needed after the initial characterization steps. Thus, there is an unmet need in the art for the development of polypeptide expression systems that allow for modular expression and production of recombinant polypeptides. Summary of the Invention
[0005] The present invention relates to a polypeptide expression system for the modular expression and production of polypeptides.
[0006] In one aspect, the invention features a polypeptide expression system that includes a first nucleic acid molecule and a second nucleic acid molecule, wherein (a) the first nucleic acid molecule includes the following components: (i) a first eukaryotic promoter (P1 Euk1 ), (ii) a first polypeptide coding sequence (PES11), (iii) a first 5' splice site (5'ss11), and (iv) a hybridizing sequence (HS1), these components being referred to as P1 Euk1 -PES11-5'ss11-HS1, and (b) a second nucleic acid molecule comprising the following components: (i) a eukaryotic promoter (P2 Euk ), (ii) a hybridizing sequence (HS2) capable of hybridizing to HS1, (iii) a 3' splice site (3'ss2), (iv) a polypeptide coding sequence (PES2), and (v) a polyadenylation site (pA2), and these components are referred to as P2 Euk In some embodiments, P1 is operably linked to P1 in the 5' to 3' direction as P1-HS2-3'ss2-PES2-pA2. Euk1 is a cytomegalovirus (CMV) promoter or a simian virus 40 (SV40) promoter. In some embodiments, P2 Euk is a CMV promoter or an SV40 promoter. In some embodiments, the first expression cassette further comprises a first nucleic acid sequence encoding a eukaryotic signal sequence (ESS11), wherein ESS11 is a P1 Euk1 and PES11. In some embodiments, ESS11 is derived from a variable heavy (VH) gene.
[0007] In some embodiments, the first expression cassette comprises the following components: (i) a 5' splice site (5'ss12), (ii) a prokaryotic promoter (P1 Prok1 ), and (iii) a 3' splice site (3'ss11), wherein these components are Prok1-3'ss11 are operably linked to each other in the 5' to 3' direction, and ePPM1 is P1 Euk1 In some embodiments, P1 Prok1 is selected from the group consisting of a PhoA promoter, a Tac promoter, a Lac promoter, and a Tphac promoter. In some embodiments, ePPM1 further comprises a first nucleic acid sequence encoding a prokaryotic signal sequence (PSS11). In some embodiments, PSS11 is derived from the heat-stable enterotoxin II (stII) gene. In some embodiments, the polypeptide expression system further comprises a polypyrimidine tract (PPT11) located between PSS11 and 3'ss11. In some embodiments, PPT11 comprises the nucleic acid sequence TTCCTTTTTTCTCTTTCC (SEQ ID NO: 1). In some embodiments, PPS11 does not contain a cryptic 5' splice site. In some embodiments, HS1 is a gene encoding all or part of a coat protein or an adaptor protein. In some embodiments, the coat protein is selected from the group consisting of pI, pII, pIII, pIV, pV, pVI, pVII, pVIII, pIX, and pX of bacteriophage M13, f1, or fd. In some embodiments, the coat protein is the pill protein of bacteriophage M13. In some embodiments, the pill fragment comprises amino acid residues 267-421 of the pill protein or amino acid residues 262-418 of the pill protein. In some embodiments, the adapter protein is a leucine zipper. In some embodiments, the leucine zipper comprises the amino acid sequence of SEQ ID NO: 4 or 5.
[0008] In some embodiments, the first nucleic acid molecule is a second eukaryotic promoter (P1 Euk2 ), (ii) a second polypeptide coding sequence (PES12), and (iii) a polyadenylation site (pA1), wherein these components are Euk2In some embodiments, P1 Euk2 is a CMV promoter or an SV40 promoter. In some embodiments, the second expression cassette further comprises a second nucleic acid sequence encoding a eukaryotic signal sequence (ESS12). In some embodiments, ESS12 is derived from the mouse binding immunoglobulin protein (mBiP) gene. In some embodiments, ESS12 comprises the nucleic acid sequence ATG AAN TTN ACN GTN GTN GCN GCN GCN CTN CTN CTN GGN (SEQ ID NO: 6), where N is A, T, C, or G.
[0009] In some embodiments, the second expression cassette comprises the following components: (i) a 5' splice site (5'ss13), (ii) a prokaryotic promoter (P1 Prok2 ), and (iii) a 3' splice site (3'ss12), wherein these components are Prok2 -3'ss12 are operably linked to each other in the 5' to 3' direction, and ePPM2 is P1 Euk2 In some embodiments, P1 Prok2is selected from the group consisting of a PhoA promoter, a Tac promoter, and a Lac promoter. In some embodiments, ePPM2 further comprises a nucleic acid sequence encoding a prokaryotic signal sequence (PSS12). In some embodiments, PSS12 is derived from the heat-stable enterotoxin II (stII) gene. In some embodiments, the polypeptide expression system further comprises a polypyrimidine tract (PPT12) located between PSS12 and 3'ss12. In some embodiments, PPT12 comprises the nucleic acid sequence TTCCTTTTTTCTCTTTCC (SEQ ID NO: 1). In some embodiments, the second expression cassette is located 5' to the first expression cassette. In some embodiments, the polypeptide expression system further comprises an intron splice enhancer (ISE) (ISE1) located between 5'ss11 and HS1. In some embodiments, ISE1 comprises a G-run containing three or more consecutive guanine residues. In some embodiments, ISE1 comprises a G-run containing nine consecutive guanine residues. In some embodiments, the polypeptide expression system further comprises a polypyrimidine tract (PPT2) positioned between HS2 and 3'ss2. In some embodiments, PPT2 comprises the nucleic acid sequence TTCCTCTTTCCCTTTCTCTCCC (SEQ ID NO: 7). In some embodiments, the polypeptide expression system further comprises an ISE (ISE2) positioned between HS2 and 3'ss2. In some embodiments, ISE2 comprises a G-run containing three or more consecutive guanine residues. In some embodiments, ISE2 comprises a G-run containing nine consecutive guanine residues. In some embodiments, 5'ss11 comprises the nucleic acid sequence GTAAGA (SEQ ID NO: 8).
[0010] In some embodiments, expression driven by a eukaryotic promoter occurs in a mammalian cell. In some embodiments, the mammalian cell is an Expi293F cell, a CHO cell, a 293T cell, or an NSO cell. In some embodiments, the mammalian cell is an Expi293F cell. In some embodiments, expression driven by a prokaryotic promoter occurs in a bacterial cell. In some embodiments, the bacterial cell is an E. coli cell. In some embodiments, PES11 encodes all or a portion of an antibody. In some embodiments, PES11 encodes a polypeptide comprising a VH domain. In some embodiments, the polypeptide further comprises a CH1 domain. In some embodiments, PES2 encodes all or a portion of an antibody. In some embodiments, PES2 encodes a polypeptide comprising a CH2 domain and a CH3 domain. In some embodiments, PES12 encodes all or a portion of an antibody. In some embodiments, PES12 encodes a polypeptide comprising a VL domain and a CL domain.
[0011] In another aspect, the present invention provides a method for producing a gene encoding ... Euk1 (b) the following components: (i) a 5' splice site (5'ss12), (ii) a prokaryotic promoter (P1 Prok1 ), and (iii) a first excisable prokaryotic promoter module (ePPM1) comprising a 3' splice site (3'ss11), wherein the components of ePPM1 are: 5'ss12-P1 Prok1 (c) a first excisable prokaryotic promoter module; (d) a first polypeptide coding sequence (PES11); (e) a useful peptide coding sequence (UPES), operably linked to each other in a 5' to 3' direction as a 5'-3'ss11; and (f) a first polypeptide coding sequence (PES11). Euk1-ePPM1-PES11-5'ss11-UPES are operably linked to each other in a 5' to 3' direction. In some embodiments, the first expression cassette further comprises a first nucleic acid sequence encoding a eukaryotic signal sequence (ESS11), wherein ESS11 is a P1 Euk1 and ePPM1. In some embodiments, ePPM1 further comprises a first nucleic acid sequence encoding a prokaryotic signal sequence (PSS11), wherein PSS11 is located between P1 Prok1 and 3'ss11. In some embodiments, the nucleic acid molecule is located between a second eukaryotic promoter (P1 Euk2 ), (ii) a second polypeptide coding sequence (PES12), and (iii) a polyadenylation site (pA1), wherein these components are Euk2 In some embodiments, the second expression cassette further comprises a second nucleic acid sequence encoding a eukaryotic signal sequence (ESS12), wherein ESS12 is a sequence encoding a P1 Prok2 and 3'ss12. In some embodiments, the second expression cassette comprises the following components: (i) a 5' splice site (5'ss13), (ii) a prokaryotic promoter (P1 Prok2 ), (iii) a nucleic acid sequence encoding a prokaryotic signal sequence (PSS12), and (iv) a 3' splice site (3'ss12), wherein these components are Prok2 -PSS12-3'ss12 are operably linked to each other in the 5' to 3' direction, and ePPM2 is P1 Euk2and PES12. In some embodiments, the UPES encodes all or a portion of a useful peptide selected from the group consisting of a tag, a label, a coat protein, and an adaptor protein. In some embodiments, the coat protein is selected from the group consisting of pI, pII, pIII, pIV, pV, pVI, pVII, pVIII, pIX, and pX of bacteriophage M13, f1, or fd. In some embodiments, the coat protein is pIII of bacteriophage M13.
[0012] In another aspect, the invention features a vector comprising any one of the foregoing nucleic acid molecules. In another aspect, the invention features a vector set comprising a first vector and a second vector, wherein the first and second vectors comprise the first and second nucleic acid molecules, respectively, of any of the polypeptide expression systems disclosed herein.
[0013] In another aspect, the invention features a host cell including the aforementioned nucleic acid, vector, and / or vector set. In some embodiments, the host cell is a prokaryotic cell. In some embodiments, the prokaryotic cell is a bacterial cell. In some embodiments, the bacterial cell is an E. coli cell. In other embodiments, the host cell is a eukaryotic cell. In some embodiments, the eukaryotic cell is a mammalian cell. In some embodiments, the mammalian cell is an Expi293F cell, a CHO cell, a 293T cell, or an NSO cell. In one embodiment, the mammalian cell is an Expi293F cell.
[0014] In a further aspect, the invention features a method for producing a polypeptide, comprising culturing a host cell comprising one or more of the foregoing nucleic acids, vectors, and / or vector sets in a culture medium. In some embodiments, the method further comprises recovering the polypeptide from the host cell or culture medium. [Brief explanation of the drawings]
[0015] [Figure 1]Schematic diagram showing the relative configuration of the pDV2 and pRK-Fc nucleic acid molecules of the polypeptide expression system for modular protein expression. The diagram also shows the typical pre-mRNA product after transcription of the nucleic acid molecule in a eukaryotic cell, the expected trans-splicing event between the two pre-mRNA products generated, and the product obtained after translation of the spliced mRNA molecule. [Figure 2A] Figures 2A and 2B include partial sequence diagrams of the pDV2 vector. The 5'ss, 3'ss, and polypyrimidine tract (PPT) are bold and underlined. In Figure 2B, the region encoding the 150-nt gene III sequence that hybridizes to pRK-Fc and pRK-Fc2-derived transcripts is italicized and underlined. Mutations from the wild-type PPT are bold, italicized, and underlined. The AATAAA potential polyadenylation site in gene III of the pDV2 vector is shown above the sequence of the silent mutation introduced into the mutant pDV2 vectors pDV2c and pDC2d (bold, italicized, and underlined). Figures 2A and 2B disclose SEQ ID NOs: 2, 3, 9, 10, and 19-22, respectively, in order of appearance. [Figure 2B] Figures 2A and 2B include partial sequence diagrams of the pDV2 vector. The 5'ss, 3'ss, and polypyrimidine tract (PPT) are bold and underlined. In Figure 2B, the region encoding the 150-nt gene III sequence that hybridizes to pRK-Fc and pRK-Fc2-derived transcripts is italicized and underlined. Mutations from the wild-type PPT are bold, italicized, and underlined. The AATAAA potential polyadenylation site in gene III of the pDV2 vector is shown above the sequence of the silent mutation introduced into the mutant pDV2 vectors pDV2c and pDC2d (bold, italicized, and underlined). Figures 2A and 2B disclose SEQ ID NOs: 2, 3, 9, 10, and 19-22, respectively, in order of appearance. [Figure 3]
[0033] Figure 3 is a partial sequence diagram of the pRK-Fc vector. The branch point consensus sequence (BP), polypyrimidine tract, and 3'ss are in bold or bold and underlined and are shown in Figure 3. The 150 bp antisense gene III sequence is in italics and underlined. The first in-frame ATG codon after the CMV promoter is in bold, italics, and underlined. Figure 3 discloses SEQ ID NOs: 23 and 24, respectively, in order of appearance. [Figure 4A] Graph showing the effect of adding an ISE sequence or removing a potential polyadenylation motif in gene III in pDV2 and the complementing pRK-Fc and pRK-Fc2 vectors on the expression levels of IgG (in μg / ml) in Expi293F cells. [Figure 4B] 1 is a graph showing the effect of the plasmid ratio of pDV2c and pRK-Fc2 on the expression level of IgG (in μg / ml) in Expi293F cells. Values shown are the mean and standard error of the mean of a representative experiment of two independent experiments performed in triplicate. [Figure 5A] Figures 5A and 5B comprise a partial sequence diagram of the pDV2b vector. The 5'ss and 3'ss are indicated, and these sequences are in bold. The polypyrimidine tract (PPT) and 9-nt G-run ISE are indicated, bold, and highlighted, respectively. In Figure 5B, the region encoding the 150-nt gene III sequence that hybridizes to transcripts from pRK-Fc and pRK-Fc2 is in italics. The wild-type stII signal sequence and mutations from M13 gene III are in bold and italics, with the wild-type nucleotide residues indicated above the sequence. A potential AATAAA polyadenylation site motif is indicated above the sequence. Amino acids in brackets are encoded by both E. coli and codons created by splicing in mammalian cells. BsiWI and RsrII restriction enzyme sites at the 3' end of the signal sequence used for cloning variable region sequences are indicated above the sequence. Figures 5A and 5B disclose, in order of appearance, SEQ ID NOs: 2, 25, 9, 10, 19, 20, 26, and 27, respectively. [Figure 5B]Figures 5A and 5B comprise a partial sequence diagram of the pDV2b vector. The 5'ss and 3'ss are indicated, and these sequences are in bold. The polypyrimidine tract (PPT) and 9-nt G-run ISE are indicated, bold, and highlighted, respectively. In Figure 5B, the region encoding the 150-nt gene III sequence that hybridizes to transcripts from pRK-Fc and pRK-Fc2 is in italics. The wild-type stII signal sequence and mutations from M13 gene III are in bold and italics, with the wild-type nucleotide residues indicated above the sequence. A potential AATAAA polyadenylation site motif is indicated above the sequence. Amino acids in brackets are encoded by both E. coli and codons created by splicing in mammalian cells. BsiWI and RsrII restriction enzyme sites at the 3' end of the signal sequence used for cloning variable region sequences are indicated above the sequence. Figures 5A and 5B disclose, in order of appearance, SEQ ID NOs: 2, 25, 9, 10, 19, 20, 26, and 27, respectively. [Figure 6] 6 is a partial sequence diagram of the pRK-Fc2 vector. The branchpoint consensus sequence (BP), polypyrimidine tract, and 3'ss are in bold. The 9-nt G-run ISE is highlighted. The 150 bp antisense gene III sequence is in italics. The first ATG triplet and in-frame stop codon are underlined. The first in-frame ATG codon after the CMV promoter is in bold and italics. The CMV promoter TATA box and transcription start site are indicated above the sequence. The glutamic acid residue in brackets is encoded by a codon created by trans-splicing in mammalian cells. FIG. 6 discloses SEQ ID NOs: 28 and 29, respectively, in order of appearance. [Figure 7A] 1 is a series of graphs showing deconvoluted masses from mass spectrometry analysis of the heavy (left panel) and light (right panel) chains of IgG expressed and purified in Expi293F cells. [Figure 7B] FIG. 7B is a table showing the predicted and observed masses for both the heavy and light chains of FIG. 7A. [Figure 8A]Figure 1 is a graph showing the yields (in mg) of IgG molecules of five different specificities purified from the supernatant (30 ml) of Expi293F cell cultures co-transfected with pDV2d (containing the ISE and lacking the AATAAA motif of gene III) and pRK-Fc2 vector. n=3. Error bars indicate the standard error of the mean. [Figure 8B] Figure 1 shows the yields (in mg) of IgG molecules of five different specificities purified from supernatants (30 ml) of 293T and CHO cell cultures co-transfected with pDV2d (containing the ISE and lacking the AATAAA motif of gene III) and pRK-Fc2 vector. n=4. Error bars indicate the standard error of the mean. [Figure 9] 9 is a partial sequence diagram of the pRK-Fab-Flag vector showing the region between the CMV promoter TATA box fused to the Flag tag sequence and the human IgG1 upper hinge region. The hinge and Flag tag sequences are followed by an SV40 polyadenylation signal (not shown). The 3'ss, including the polypyrimidine tract and consensus branch point (BP), is shown and is in bold or bold and underlined. The ISE sequence is shown and is in bold, underlined, and italic. The antisense gene III sequence, which mediates hybridization with the donor transcript, is italic and underlined. FIG. 9 discloses SEQ ID NOs: 30 and 31, respectively, in order of appearance. [Figure 10] 1 is a schematic diagram showing the relative configurations of first and second nucleic acid molecules possible for modular expression of a typical polypeptide product. The diagram also shows a typical pre-mRNA product following transcription of the nucleic acid molecule in a eukaryotic cell, the expected trans-splicing events between the two pre-mRNA products generated, and the resulting product following translation of the spliced mRNA molecule. [Figure 11] 1 is a schematic diagram showing the relative configuration of first and second nucleic acid molecules possible for the modular expression of a typical two or more polypeptide products. The diagram also shows a typical pre-mRNA product following transcription of the nucleic acid molecule in a eukaryotic cell, the expected trans-splicing event between the two pre-mRNA products generated, and the resulting product following translation of the spliced mRNA molecule. [Figure 12] Figure 1 shows a pair of Western blots showing the expression of Mab1 heavy chain and Fd-cP3 fusion protein in Expi293F cells cotransfected with pDV2 mutants and pRK-Fc2. Transfected Expi293F lysates were reduced with dithiothreitol (DTT) and analyzed by Western blotting using anti-IgG1 Fc (upper panel) or anti-M13 p3 (lower panel) antibodies. GFP and HC control vectors express green fluorescent protein and human IgG1 heavy chain, respectively. HC indicates full-length human IgG1 heavy chain. Gene III AATAAA indicates the presence of a potential polyadenylation site in gene III. FC* indicates the putative cytoplasmic N-terminally truncated Fc fragment expression product. NA, not applicable. [Figure 13] This is a sodium dodecyl sulfate polyacrylamide gel electrophoresis (SDS-PAGE) gel showing the analysis of IgG and Fab fragments expressed and purified in Expi293F cells. Expressed and purified IgG and Fab fragments from the supernatant of Expi293F cells cotransfected with pDV2d and pRK-Fc2 (IgG) or pRK-FAB-F (Fab fragment) were separated by 4-20% gradient SDS-PAGE under reducing or non-reducing conditions and stained with Coomassie brilliant blue. The identities of the bands are indicated on the right. HC, heavy chain; LC, light chain; Fd, heavy chain Fd fragment (VH + CH1 + upper hinge). The HC and LC (non-reduced) bands are heavy and light chains that do not form interchain disulfide bonds but may have intrachain disulfide bonds in the IgG sample. The approximately 25 kDa band in the non-reduced Fab sample has co-migrating heavy and light chains that did not form interchain disulfide bonds but may have intrachain disulfide bonds. [Figure 14]1 is a graph showing the display of Fab fragments on phage with phagemid pDV2, as detected by phage enzyme-linked immunosorbent assay (ELISA). Fab-zip-phage was generated by infecting E. coli cells carrying pFab-zip phagemid with M13KO7 helper phage. pDV2 phage was generated by infecting E. coli cells carrying pDV2d vector with Amber-2614 KO7 phage. DETAILED DESCRIPTION OF THE INVENTION
[0016] I. Definition The term "antibody" herein is used in the broadest sense and encompasses a variety of antibody structures, including, but not limited to, monoclonal antibodies, polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies), and antibody fragments, so long as they exhibit the desired antigen-binding activity.
[0017] The Kabat numbering system is commonly used when referring to residues within the variable domain (approximately residues 1-107 of the light chain and residues 1-113 of the heavy chain) (e.g., Kabat et al., Sequences of Immunological Interest. 5th ed. Public Health Service, National Institutes of Health, Bethesda, Md. (1991)). The "EU numbering system" or "EU index" is commonly used when referring to residues within an immunoglobulin heavy chain constant region (e.g., the EU index reported in Kabat et al., supra). "Kabat's EU index" refers to the residue numbering of the human IgG1 EU antibody. Unless otherwise stated herein, reference to a residue number within the variable domain of an antibody refers to residue numbering according to the Kabat numbering system. Unless otherwise stated herein, reference to a residue number within the heavy chain constant domain of an antibody refers to residue numbering according to the EU numbering system.
[0018] The natural, basic four-chain antibody unit is a heterotetrameric glycoprotein consisting of two identical light chains (LC) and two identical heavy chains (HC). (IgM antibodies consist of five basic heterotetrameric units with an additional polypeptide called the J chain and therefore contain 10 antigen-binding sites, whereas secreted IgA antibodies can polymerize to form multivalent assemblies containing two to five basic four-chain units with J chains.) In the case of IgG, the four-chain unit is generally approximately 150,000 daltons. The two HCs are linked to each other by one or more disulfide bonds depending on the HC isotype, while each LC is linked to an HC by one covalent disulfide bond. Each HC and LC also have regularly spaced intrachain disulfide bridges. Each HC has a variable domain (VH) at its N-terminus, followed by three constant domains (CH1, CH2, CH3) for each of the α and γ chains, and four Cj domains for the μ and ε isotypes. Each LC has a variable domain (VL) at its N-terminus followed by a constant domain (CL) at its other end. The VL aligns with the VH, and the CL aligns with the first constant domain of the heavy chain (CH1). CH1 can be connected to the second constant domain of the heavy chain (CH2) by a hinge region. Specific amino acid residues are thought to form the interface between the light and heavy chain variable domains. The VH and VL pair together to form a single antigen-binding site. For the structure and properties of various classes of antibodies, see, for example, Basic and Clinical Immunology, 8th ed., Daniel P. Stites, Abba I. Terr and Tristram G. Parslow (eds.), Appleton & Lange, Norwalk, CT, 1994, page 71 and Chapter 6.
[0019] The "CH2 domain" of the human IgG Fc region typically extends from approximately residue 231 to approximately residue 340 of IgG. The CH2 domain is unique in that it is not tightly paired with another domain. Rather, two N-linked branched carbohydrate chains are interposed between the two CH2 domains in an intact native IgG molecule. It has been speculated that the carbohydrates may act as a surrogate for domain-domain pairing and help stabilize the CH2 domain. Burton, Molec. Immunol. 22:161-206 (1985).
[0020] The "CH3 domain" comprises the stretch of residues C-terminal to the CH2 domain in the Fc region (ie, from about amino acid residue 341 to about amino acid residue 447 of IgG).
[0021] Light chains (LC) from any vertebrate species can be assigned to one of two clearly distinct types, called kappa and lambda, based on the amino acid sequence of their constant domains. Depending on the amino acid sequence of the constant domains (CH) of their heavy chains, immunoglobulins can be assigned to various classes or isotypes. There are five classes of immunoglobulins: IgA, IgD, IgE, IgG, and IgM, with heavy chains designated α, δ, γ, ε, and μ, respectively. The γ and α classes are further divided into subclasses based on relatively minor differences in CH sequence and function; for example, humans express the following subclasses: IgG1, IgG2, IgG3, IgG4, IgA1, and IgA2.
[0022] The term "variable" refers to the fact that certain portions of variable domains differ extensively in sequence among antibodies. V domains mediate antigen binding and define the specificity of a particular antibody for its particular antigen. However, variability is not evenly distributed across the 110-amino acid span of the variable domains. Instead, V regions consist of relatively invariant stretches of 15-30 amino acids called framework regions (FRs), separated by shorter regions of extreme variability, each 9-12 amino acids long, called "hypervariable regions." Native heavy and light chain variable domains each contain four FRs that adopt a primarily beta-sheet structure, connected by three hypervariable regions that form loops that connect, and in some cases form part of, this beta-sheet structure. The hypervariable regions in each chain are held together in close proximity by FRs and, together with the hypervariable regions from the other chain, contribute to the formation of the antigen-binding site of antibodies (see Kabat et al., Sequences of Proteins of Immunological Interest, 5th ed. Public Health Service, National Institutes of Health, Bethesda, MD, 1991). The constant domains are not directly involved in binding the antibody to an antigen, but exhibit various effector functions, such as participation of the antibody in antibody-dependent cellular cytotoxicity (ADCC).
[0023] "Antibody fragment" refers to a molecule other than an intact antibody that contains a portion of an intact antibody that binds the antigen to which the intact antibody binds. Examples of antibody fragments include, but are not limited to, Fv, Fab, Fab', Fab'-SH, F(ab')2; diabodies; linear antibodies; single-chain antibody molecules (e.g., scFv); and multispecific antibodies formed from antibody fragments.
[0024] A "Fab" fragment is an antigen-binding fragment produced by papain digestion of an antibody and consists of an entire L chain along with the variable region domain of the H chain (VH) and the first constant domain (CH1) of one heavy chain. Papain digestion of an antibody produces two identical Fab fragments. Pepsin treatment of an antibody produces a single large F(ab')2 fragment, which roughly corresponds to two disulfide-linked Fab fragments with bivalent antigen-binding activity and is still capable of cross-linking antigen. Fab' fragments differ from Fab fragments by having several additional residues at the carboxy terminus of the CH1 domain, including one or more cysteines from the antibody hinge region. Fab'-SH is the designation herein for Fab' in which the cysteine residue(s) in the constant domain bear a free thiol group. F(ab')2 antibody fragments were originally produced as pairs of Fab' fragments with hinge cysteines between them. Other chemical linkages of antibody fragments are also known.
[0025] As used herein, "adapter protein" refers to a protein sequence that specifically interacts with another adaptor protein sequence in solution. In one embodiment, an "adapter protein" comprises a heteromultimerization domain. Such an adaptor protein may be a leucine zipper protein, or a protein sequence similar to SEQ ID NO:4 (cJUN(R):ASIARL[E]E[K]V KTL[K]A[Q]NYEL [A]S[T]ANMLRE[Q] VAQLGGC) or SEQ ID NO:5 (FosW(E):AS[I]DEL[Q]AE[V] EQLEE[R]NYAL [R]KE[V]EDL[Q]K[Q] [A]EKLGGC) or variants thereof (including, but not limited to, the amino acids of SEQ ID NO:4 and SEQ ID NO:5, which may be modified to include underlined and bolded), which variants have amino acid modifications that maintain or increase the affinity of the adaptor protein for another adaptor protein, or a polypeptide comprising an amino acid sequence selected from the group consisting of SEQ ID NO:11 (ASIARLRERVKTLRARNYELRSRANMLRERVAQLGGC) or SEQ ID NO:12 (ASLDELEAEIEQLEEENYALEKEIEDLEKELEKLGGC), or a polypeptide comprising the amino acid sequence of SEQ ID NO:13 (GABA-R1:EEKSRLLEKE NRELEKIIAE KEERVSELRH QLQSVGGC) or SEQ ID NO:14 (GABA-R2:TSRLEGLQSE NHRLRMKITE LDKDLEEV™ QLQDVGGC) or SEQ ID NO:15 (Cys:AGSC) or SEQ ID NO:16 (Hinge:CPPCPG). The nucleic acid molecule encoding the coat protein or adaptor protein is contained within a synthetic intron.
[0026] As used herein, a "heteromultimerization domain" refers to a modification or addition to a biological molecule to promote heteromultimer formation and hinder homomultimer formation. Any heterodimerization domain that has a strong preference for forming heterodimers over homodimers is within the scope of the present invention. Illustrative examples include, but are not limited to, U.S. Patent Application No. 20030078385 (Arathoon et al., Genentech, describing knobs into holes), International Patent Publication No. WO2007147901 (Kjargaard et al., Novo Nordisk, describing ionic interactions), International Patent Publication No. WO2009089004 (Kannan et al., Amgen, describing electrostatic steering effects), and International Patent Publication No. WO2011 / 034605 (Christensen et al., Genentech, describing coiled coils). See also, for example, Pack, P. & Plueckthun, A., Biochemistry 31, 1579-1584 (1992), which describes leucine zippers, or Pack et al., Bio / Technology 11, 1271-1277 (1993), which describes helix-turn-helix motifs. The phrases "heteromultimerization domain" and "heterodimerization domain" are used interchangeably herein.
[0027] As used herein, the term "cloning site" refers to a nucleic acid sequence that contains a restriction enzyme site for restriction endonuclease-mediated cloning by ligation of nucleic acid sequences containing compatible cohesive or blunt ends, a region of a nucleic acid sequence that serves as a priming site for PCR-mediated cloning of insert DNA by homology and extension known as "overlap PCR stitching," or a recombination site for recombinase-mediated insertion of a target nucleic acid sequence by a recombination-exchange reaction, or mosaic ends for transposon-mediated insertion of a target nucleic acid sequence, and other techniques common in the art.
[0028] As used herein, "coat protein" refers to any of the five capsid proteins that are components of phage particles, including pIII, pVI, pVII, pVIII, and pIX. In one embodiment, a "coat protein" can be used to display proteins or peptides (see, "Phage Display, A Practical Approach," Oxford University Press, edited by Clackson and Lowman, 2004, pp. 1-26). In one embodiment, the coat protein can be the pIII protein or some variant, portion, and / or derivative thereof. For example, the C-terminal portion of the M13 bacteriophage pIII coat protein (cP3), e.g., the sequence encoding the C-terminal residues 267-421 of protein III of M13 phage, can be used. In one embodiment, the pill sequence comprises the amino acid sequence of SEQ ID NO: 17 (AEDIEFASGGGSGAETVESCLAKPHTENSFTNVWKDDKTLDRYANYEGCLWNATGVVVCTGDETQCYGTWVPIGLAIPENEGGGSEGGGSEGGGSEGGGTKPPEYGDTPIPGYTYINPLDGTYPPGTEQNPANPNPSLEESQPLNTFMFQNNRFRNRQGALTVYTGTVTQGTDPVKTYYQYTPVSSKAMYDAYWNG Includes KFRDCAFHSGFNEDPFVCEYQGQSSDLPQPPVNAGGGSGGGSGGGSEGGGSEGGGSEGGGSEGGGSGGGSGSGDFDYEKMANANKGAMTENADENALQSDAKGKLDSVATDYGAAIDGFIGDVSGLANGNGATGDFAGSNSQMAVGDGDNSPLMNNFRQYLPSLPQSVECRPFVFSAGKPYEFSIDCDKINLFRGVFAFLLYVATFMYVFSTFANILRNKES).In one embodiment, the pill fragment comprises the amino acid sequence of SEQ ID NO: 18 (SGGGSGSGDFDYEKMANANKGAMTENADENALQSDAKGKLDSVATDYGAAIDGFIGDVSGLANGNGATGDFAGSNSQMAQVGDGDNSPLMNNFRQYLPSLPQSVECRPFVFGAGKPYEFSIDCDKINLFRGVFAFLLYVATFMYVFSTFANILRNKES).
[0029] As used herein, "expression cassette" refers to a nucleic acid fragment (e.g., a DNA fragment) that contains a specific nucleic acid sequence that has a particular biological and / or biochemical activity. The terms "cassette," "gene cassette," or "DNA cassette" can be used interchangeably and can have the same meaning.
[0030] The terms "host cell," "host cell line," and "host cell culture" are used interchangeably and refer to cells into which exogenous nucleic acid has been introduced, including the progeny of such cells. Host cells include "transformants" and "transformed cells," including the primary transformed cell and its progeny regardless of the number of transfers. The progeny may not be completely identical in nucleic acid content to the parent cell, but may contain mutations. Mutant progeny that have the same function or biological activity as screened or selected for in the originally transformed cell are included herein.
[0031] As used herein, "linked" or "links" or "link" is meant to refer to the covalent attachment of two amino acid sequences or two nucleic acid sequences, respectively, via a peptide or phosphodiester bond, which may include any number of additional amino acids or nucleic acid sequences between the two amino acid or nucleic acid sequences being attached.
[0032] "Nucleic acid" or "polynucleotide," as used interchangeably herein, refer to a polymer of nucleotides of any length, and include DNA and RNA. The nucleotides can be deoxyribonucleotides, ribonucleotides, modified nucleotides or bases, and / or their analogs, or any substrate that can be incorporated into a polymer by DNA or RNA polymerase or by a synthetic reaction. A polynucleotide can comprise modified nucleotides, such as methylated nucleotides and their analogs. Modifications to the nucleotide structure, if present, can be made before or after assembly of the polymer. The sequence of nucleotides can be interrupted by non-nucleotide components. A polynucleotide can be further modified after synthesis, such as by conjugation with a label. Other types of modifications include, for example, "cap" substitution of one or more of the naturally occurring nucleotides with an analog; internucleotide modifications, such as those with uncharged linkages (e.g., methylphosphonates, phosphotriesters, phosphoamidates, carbamates, etc.) and those with charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.); those containing pendant moieties, such as proteins (e.g., nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.); those containing intercalators (e.g., acridine, psoralens, etc.); those containing chelators (e.g., metals, radioactive metals, boron, metal oxides, etc.); those containing alkylators; those with modified linkages (e.g., alpha-aromatic nucleic acids, etc.); and unmodified forms of the polynucleotide(s). Additionally, any of the hydroxyl groups originally present in the sugar may be replaced, for example, by phosphonate groups, phosphate groups, protected by standard protecting groups, or activated to prepare additional linkages to additional nucleotides, or attached to a solid or semi-solid support. The 5' and 3' terminal OH may be phosphorylated or substituted with amines or organic capping group moieties of 1 to 20 carbon atoms, and other hydroxyls may be derivatized to standard protecting groups.Polynucleotides may also contain analogous forms of ribose or deoxyribose sugars commonly known in the art, including, for example, 2'-O-methyl-, 2'-O-allyl, 2'-fluoro-, or 2'-azido-ribose; carbocyclic sugar analogs; alpha-anomeric sugars; epimeric sugars such as arabinose, xylose, or lyxose; pyranose sugars; furanose sugars; sedoheptuloses; acyclic analogs; and basic nucleoside analogs such as methyl riboside. One or more phosphodiester linkages may be replaced by alternative linking groups. These alternative linking groups include, but are not limited to, embodiments in which phosphate is replaced by P(O)S ("thioate"), P(S)S ("dithioate"), (O)NR2 ("amidate"), P(O)R, P(O)OR', CO, or CH2 ("formacetal"), where each R or R' is independently H or substituted or unsubstituted alkyl (1-20C) optionally containing an ether (-O-) linkage, aryl, alkenyl, cycloalkyl, cycloalkenyl, or araldyl. Not all linkages within a polynucleotide need be identical. The above description applies to all polynucleotides referred to herein, including RNA and DNA.
[0033] A nucleic acid is "operably linked" when it is placed into a structural or functional relationship with another nucleic acid sequence. For example, a segment of DNA can be operably linked to another segment of DNA when they are positioned together on the same contiguous DNA molecule and have a structural or functional relationship, such as a promoter or enhancer positioned relative to a coding sequence to promote transcription of the coding sequence; a ribosome binding site positioned relative to a coding sequence to promote translation; or a presequence or secretory leader positioned relative to a coding sequence to promote expression of a preprotein (e.g., a preprotein involved in the secretion of the encoded polypeptide). In other examples, operably linked nucleic acid sequences are not contiguous, but are positioned such that they have a functional relationship to each other as nucleic acids or as proteins expressed by them. For example, enhancers need not be contiguous. Linking can be achieved by ligation at convenient restriction enzyme sites or by use of synthetic oligonucleotide adapters or linkers.
[0034] The term "polyadenylation signal" or "polyadenylation site" is used herein to mean a sequence sufficient to direct the addition of polyadenosine ribonucleic acid to an expressed RNA molecule in a cell.
[0035] A "promoter" is a nucleic acid sequence that enables the initiation of transcription of a gene sequence into messenger RNA; such transcription is initiated upon the binding of RNA polymerase on or near the promoter.
[0036] The term "3' splice site" is intended to mean a nucleic acid sequence, eg a pre-mRNA sequence, at a 3' intron / exon boundary that can be recognized and bound by the splicing machinery.
[0037] The term "5' splice site" is intended to mean a nucleic acid sequence, eg a pre-mRNA sequence, at a 5' exon / intron boundary that can be recognized and bound by the splicing machinery.
[0038] The term "cryptic splice site" is intended to mean a normally dormant 5' or 3' splice site that can be activated by mutation or otherwise to serve as a splicing element. For example, a mutation may activate a 5' splice site downstream of a native or dominant 5' splice site. Use of this "cryptic" splice site results in the production of a distinct mRNA splicing product that would not be produced by use of the native or dominant splice site.
[0039] As used herein, the term "trans-splicing" refers to the joining of exons contained on separate, non-contiguous RNA molecules.
[0040] The term "variable region" or "variable domain" refers to the domain of an antibody heavy or light chain that is involved in binding the antibody to an antigen. The heavy and light chain variable domains (VH and VL, respectively) of native antibodies generally have similar structures, with each domain containing four conserved framework regions (FR) and three hypervariable regions (HVR). (See, e.g., Kindt et al., Kuby Immunology, 6th ed., W.H. Freeman and Co., page 91 (2007)). A single VH or VL domain may be sufficient to confer antigen-binding specificity. Furthermore, antibodies that bind to a specific antigen can be isolated from the antigen-binding antibodies by using a VH or VL domain to screen a library of complementary VL or VH domains, respectively. See, e.g., Portolano et al., J. Immunol. 150:880-887 (1993); Clarkson et al., Nature 352:624-628 (1991).
[0041] As used herein, the term "vector" refers to a nucleic acid molecule capable of amplifying another nucleic acid to which it is linked. The term includes vectors as self-replicating nucleic acid structures and vectors that integrate into the genome of a host cell into which they are introduced. Certain vectors are capable of directing the expression of nucleic acids to which they are operatively linked. Such vectors are referred to herein as "expression vectors."
[0042] II. Modular Polypeptide Expression Systems The present invention is based, at least in part, on the discovery that pre-mRNA trans-splicing can be utilized in mammalian cells to enable modular recombinant protein expression. The modular, flexible protein expression concept allows for the precise joining of any two protein-coding sequences encoded by two different constructs into a single mRNA encoding a polypeptide chain, without any of the requirements and constraints of other protein-protein splicing methods. This concept can be adapted to simplify and extend other technologies that require the expression in mammalian cells of large collections of proteins with various combinations of repeating modules.
[0043] Described herein is the generation of a number of polypeptide expression systems that allow for the modular expression of a variety of antibody formats in the context of a phage display expression system. The necessary nucleic acid components, vectors, host cells, and methods for using the polypeptide expression systems of the invention are described herein.
[0044] A. Mode of carrying out the invention The practice of the present invention will employ, unless otherwise indicated, conventional techniques of molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry, and immunology, which are within the skill of the art. Such techniques are fully explained in such references as "Molecular Cloning: A Laboratory Manual," 2nd Edition (Sambrook et al., 1989), "Oligonucleotide Synthesis" (M.J. Gait, ed., 1984), "Animal Cell Culture" (R.I. Freshney, ed., 1987), "Methods in Enzymology" (Academic Press, Inc.), "Handbook of Experimental Immunology," 4th Edition (D.M. Weir & C.C. Blackwell, eds., Blackwell Science Inc., 1987), "Gene Transfer Vectors for Mammalian Cells" (J.M. Miller & M.P. Calos, eds., 1987), "Current Protocols in Molecular Biology" (F.M.A. Usubel et al., eds., 1987), "PCR: The Polymerase Chain Reaction" (Mullis et al., eds., 1994), and "Current Protocols in Immunology" (J.E. Coligan et al., eds., 1991).
[0045] B. Modular Protein Expression System The polypeptide expression systems of the present invention can support the expression of the same or different (e.g., reformatted) forms of a polypeptide (e.g., fusion proteins). The present invention provides a means to generate such polypeptide expression systems for modular expression and production of various forms (e.g., various formats or various fusion forms) of a protein of interest in a host cell-dependent manner by using the process of trans-splicing.
[0046] 1. Nucleic Acid Components of the Modular Protein Expression System a. Structure of the nucleic acid component of the modular protein expression system The protein expression system uses at least two nucleic acid molecules that together allow for flexible, modular expression of any desired polypeptide through the process of directed pre-mRNA trans-splicing. The first nucleic acid molecule is driven by a eukaryotic promoter (P1 Euk1 ) (e.g., a cytomegalovirus (CMV) promoter, a simian virus 40 (SV40) promoter, a Moloney murine leukemia virus U3 region, a caprine arthritis-encephalitis virus U3 region, a Visna virus U3 region, or a retroviral U3 region sequence), operably linked to a polypeptide coding sequence (PES11). In some cases, the polypeptide coding sequence encodes only a portion of a desired polypeptide, with the remainder being supplied by a polypeptide coding sequence (PES2) contained on a second nucleic acid molecule. The first nucleic acid molecule may include a 5'ss (5'ss11) (e.g., GTAAGA (SEQ ID NO: 8)) located downstream (3') of PES11 but upstream (5') of the hybridizing sequence (HS1).
[0047] The HS1 sequence may contain a gene encoding all or part of a polypeptide tag, label, coat protein, and / or adaptor protein, which may be positioned in frame with PES11, such that its expression results in the PES11-encoded protein fused to the HS1-encoded protein. In some cases, HS1 is a gene encoding all or part of a coat protein selected from the group consisting of pI, pII, pIII, pIV, pV, pVI, pVII, pVIII, pIX, and pX of bacteriophage M13, f1, or fd. For example, PES11 can encode all or part of an antibody or Fab fragment thereof, and the HS1 sequence can encode the coat protein (e.g., all or part of the pIII protein of bacteriophage M13, e.g., a pIII fragment including amino acid residues 267-421 of the pIII protein or amino acid residues 262-418 of the pIII protein), resulting in an antibody- or Fab fragment-pIII protein fusion product. Alternatively, HS1 is a gene encoding all or part of an adaptor protein, such as a leucine zipper, wherein the leucine zipper comprises the amino acid sequence of SEQ ID NO:4 or SEQ ID NO:5.
[0048] In addition, the first nucleic acid molecule comprises P1 Euk1 The first nucleic acid molecule can encode a eukaryotic signal sequence (ESS11) located 3' to P1 and 5' to PES11. Euk1 -ESS11-PES11-5'ss11-HS1 may include the above components linked to each other in the 5' to 3' direction (e.g., operably linked).
[0049] The second nucleic acid molecule of the protein expression system comprises a eukaryotic promoter (P2 Euk) (e.g., a cytomegalovirus (CMV) promoter or a simian virus 40 (SV40) promoter). In some cases, the polypeptide coding sequence encodes only a portion of the desired polypeptide, with the remainder being supplied by a polypeptide coding sequence (PES11) contained on the first nucleic acid molecule. The second nucleic acid molecule may include a 3' splice site (3'ss2) located 5' to PES2. The second nucleic acid molecule may include a 3' splice site (3'ss2) located 5' to PES2. Euk The second nucleic acid molecule may further comprise a polyadenylation site (pA2), wherein the second nucleic acid molecule comprises a hybridizing sequence (HS2) capable of hybridizing to HS1 located between P2 and 3'ss2. Euk -HS2-3'ss2-PES2-pA2 are operably linked to each other in the 5' to 3' direction.
[0050] Thus, trans-splicing between first and second nucleic acid pre-mRNA products in eukaryotic cells (e.g., mammalian cells) is triggered by hybridization of complementary sequences (i.e., HS1 and HS2) located on separate mRNA molecules, thereby bringing the isolated 5' splice site (5'ss11) of the first molecule and the isolated 3' splice site (3'ss2) of the second molecule into close proximity, resulting in trans-splicing and supporting the formation of the desired trans-spliced mRNA transcript. Additionally, to facilitate trans-splicing, the first nucleic acid molecule may contain an intron splice enhancer (ISE) (ISE1) located between 5'ss11 and HS1. ISE1 may include a G-run with three or more consecutive guanine residues, such as a G-run with nine consecutive guanine residues. Furthermore, trans-splicing between the first and second nucleic acid pre-mRNA products can be induced during their transcription in eukaryotic cells (e.g., mammalian cells, e.g., Expi293F, 293T, or CHO cells) by genetically engineering the first nucleic acid molecule to lack a canonical polyadenylation site downstream of the PES11 and / or HS1 components. This will minimize the formation of mature mRNA transcripts that can be transported to the cytoplasm before trans-splicing with the mRNA transcript of the second nucleic acid molecule can occur.
[0051] In some cases, it may be desirable to simultaneously express separate polypeptide products. For example, it may be desirable to express a second polypeptide product that can self-assemble with a first polypeptide product encoded by both the first and second nucleic acid molecules to form a desired heteromultimeric protein product (e.g., an antibody consisting of both a heavy chain and a light chain). To this end, the first and / or second nucleic acid molecule may further comprise a second expression cassette. For example, in cases where the first nucleic acid molecule comprises a second expression cassette, the second expression cassette may be driven by a second eukaryotic promoter (P1 Euk2), (ii) a second nucleic acid sequence encoding a eukaryotic signal sequence (ESS12), (iii) a second polypeptide coding sequence (PES12), and (iv) a polyadenylation site (pA1), and these components may comprise P1 Euk2 The first nucleic acid molecule encodes two polypeptide products under separate promoters, whereby one of the mRNA transcripts encoding one of the polypeptide products of the first nucleic acid molecule is formed via directed trans-splicing with the mRNA transcript encoded by the second nucleic acid molecule. In some cases, the second expression cassette is located 5' to 3' relative to the first expression cassette. In other cases, the second expression cassette is located 3' to the first expression cassette.
[0052] b. Polypeptide expression in both prokaryotic and eukaryotic cells In some cases, polypeptide expression systems can be engineered for polypeptide expression in both prokaryotic and eukaryotic cells. Thus, the first nucleic acid molecule can encode PES11, or in some cases PES11 and HS1, if expression of a polypeptide product encoded by PES11 and HS1 is desired. Euk1 and PES11. ePPM1 may contain an excisable prokaryotic promoter module (ePPM1) located between the 5' splice site (5'ss12), the prokaryotic promoter (P1 Prok1 ), a nucleic acid sequence encoding a prokaryotic signal sequence (PSS11), and a 3' splice site (3'ss11), which may comprise 5'ss12-P1 Prok1-PSS11-3'ss11 are positioned relative to each other in the 5' to 3' direction and are operably linked to drive transcription of a polypeptide encoded by PES11, or PES11 and HS1. In some cases, ePPM1 may not contain the PSS11 component (e.g., when secretion of the expressed polypeptide is not required or desired). Thus, ePPM1 may drive transcription of a polypeptide encoded by PES11 of the first nucleic acid molecule in prokaryotic cells. Meanwhile, in eukaryotic cells (e.g., mammalian cells), P1 Euk1 drives the expression of the transcription of the polypeptide encoded by PES11 of the first nucleic acid molecule, and ePPM1 can be removed from the pre-mRNA transcript by cis-splicing due to the collision of adjacent 5'ss12 and 3'ss11 elements.
[0053] In some cases, ePPM1 also includes a polypyrimidine tract (PPT11) located between PSS11 and 3'ss11. PPT11 may include, for example, the sequence TTCCTTTTTTCTCTTTCC (SEQ ID NO: 1). The second nucleic acid molecule may also include a polypyrimidine tract (PPT2) located, for example, between HS2 and 3'ss2. PPT2 may include, for example, the sequence TTCCTCTTTCCCTTTCTCTCCC (SEQ ID NO: 7). In addition, the second nucleic acid molecule may further include an ISE (ISE2) located between HS2 and 3'ss2. ISE2 may include, for example, a G-run having three or more consecutive guanine residues, such as a G-run having nine consecutive guanine residues.
[0054] In some embodiments where the first nucleic acid molecule of the polypeptide expression system comprises a second expression cassette, the second expression cassette is Euk2 and PES12, and may further comprise an excisable prokaryotic promoter module (ePPM2) located between the 5' splice site (5'ss13), (ii) a prokaryotic promoter (P1 Prok2), (iii) a nucleic acid sequence encoding a prokaryotic signal sequence (PSS12), and (iv) a 3' splice site (3'ss12), whereby these components Prok2 -PSS12-3'ss12 are positioned relative to each other in a 5' to 3' direction and are operably linked to drive transcription of the polypeptide encoded by PSS12. In some cases, ePPM2 may not include the PSS12 component (e.g., if secretion of the expressed polypeptide is not required or desirable). The second excisable prokaryotic promoter module will function in a manner similar to that of the first excisable prokaryotic promoter module described above.
[0055] The prokaryotic promoter(s) of the excisable prokaryotic promoter module(s) may be the phoA, Tac, Lac, or Tphac promoter (see, e.g., Kim et al. PLoS One. 7(4):e35844), or another prokaryotic promoter known in the art.
[0056] An additional challenge in constructing vectors capable of expressing a protein of interest in both prokaryotic (e.g., E. coli) and eukaryotic (mammalian, e.g., Expi293F) cells arises from differences in the signal sequences found in these cell types. While certain features of signal sequences are generally conserved in both prokaryotic and eukaryotic cells (e.g., hydrophobic residues located in the center of the sequence and a patch of polar / charged residues adjacent to the cleavage site at the N-terminus of the mature polypeptide), others are more characteristic of one cell type than another. Furthermore, it is known in the art that different signal sequences can have a significant effect on expression levels in mammalian cells, even if the sequences are all of mammalian origin (Hall et al., J of Biological Chemistry, 265:19996-19999 (1990); Humphreys et al., Protein Expression and Purification, 20:252-264 (2000)). For example, bacterial signal sequences typically have a positively charged residue (most commonly a lysine) immediately following the initiating methionine, whereas these are not always present in mammalian signal sequences.
[0057] If secretion of the expressed protein is necessary or desired, any signal sequence (including consensus signal sequences) that targets the polypeptide of interest to the periplasm in prokaryotes and the endoplasmic reticulum in eukaryotes can be used. For example, a eukaryotic signal sequence (e.g., ESS11 or ESS12) can be derived from or include all or a portion of the mouse binding immunoglobulin protein (mBiP) signal sequence (UniProtKB: Accession No. P20029) or an antibody heavy or light chain signal sequence (e.g., a mouse VH gene signal sequence). In some embodiments, a prokaryotic signal sequence (e.g., PSS11 or PSS12) can be derived from or include all or a portion of the heat-stable enterotoxin II (stII) gene. Other signal sequences that can be utilized include signal sequences from human growth hormone (hGH) (UniProtKB: Accession No. BIA4G6), Gaussia princeps luciferase (UniProtKB: Accession No. Q9BLZ2), and yeast endo-1,3-glucanase (yBGL2) (UniProtKB: Accession No. P15703). The signal sequence can be a natural or synthetic signal sequence. In some embodiments, the synthetic signal sequence is an optimized signal secretion sequence that drives an optimized level of display compared to its non-optimized natural signal sequence.
[0058] 2. Vectors, Host Cells, and Production Methods The present invention features a vector or vector set comprising one or more of the above-described nucleic acid molecules. Accordingly, the present invention also features a vector set comprising a first vector and a second vector, wherein the first and second vectors comprise the first and second nucleic acid molecules, respectively, of the above-described polypeptide expression system.
[0059] In addition to the nucleic acid molecule components detailed above, the vector or vector set can include a bacterial origin of replication, a mammalian origin of replication, and / or a nucleic acid encoding a polypeptide useful as a control (e.g., gD protein) or activity (e.g., protein purification, protein tagging, or protein labeling).
[0060] Also provided are methods for producing the polypeptides, comprising culturing host cells containing one or more of the above vector(s) or vector set(s) in a culture medium, and optionally recovering the antibody from the host cells (or the host cell culture medium).
[0061] C. Phage Display Vector System for Modular Antibody Expression and Reformatting In some embodiments, antibodies (e.g., full-length antibodies, e.g., full-length IgG antibodies, or fragments thereof, e.g., Fab fragments) can be produced using the polypeptide expression system of the present invention. The application of a modular protein expression system is demonstrated by designing a phage display vector system that allows the expression of different antibody formats within human cells from the same clone. The heavy chain antigen-binding region and a portion of the constant region encoded by the phage display vector were directly and precisely fused to sequences encoded in a second complementing construct by joining sequences encoding different portions of the polypeptide via pre-mRNA trans-splicing during expression within the cell.
[0062] The use of a polypeptide expression system intended to enable direct expression of IgG in mammalian cells without the need for subcloning of phage Fab sequences is described in Examples 1 and 2 below. In some cases, the first nucleic acid molecule of the polypeptide expression system can be designed to encode the entire Fab fragment component. Thus, the first nucleic acid molecule can include a PES11 component encoding a polypeptide having the VH and CH1 domains of the Fab. The first nucleic acid molecule can also include a PES12 component encoding the VL and CL domains. Transcription of the first nucleic acid molecule results in two non-contiguous pre-mRNA products, which together form the Fab fragment, which can be appropriately tagged (e.g., fused to M13 pIII) for phage display purposes.
[0063] The process of reformatting the Fab fragment into a full-length IgG antibody can be subsequently accomplished by expressing a first nucleic acid molecule in a eukaryotic cell (e.g., a mammalian cell, e.g., Expi293F cell) along with a second nucleic acid molecule that provides the remaining portion of the antibody (i.e., the CH2 and CH3 domains). For example, the second nucleic acid molecule can include a PES2 component that encodes a polypeptide having a CH2 domain and a CH3 domain. Transcription of the first and second molecules in the eukaryotic cell results in the production of three pre-mRNA transcripts, and the heavy chain-encoding pre-mRNA transcripts will be induced to undergo trans-splicing with each other to generate the reformatted full-length heavy chain of the desired IgG antibody. The processed mRNAs are then translated, resulting in the production of both the light and heavy chains of the IgG molecule; such production would not require the need for labor-intensive subcloning.
[0064] The ability to express different antibody formats from the same clone is useful in antibody discovery when different antibody formats, such as wild-type IgG, Fab fragments, or IgG with Fc modifications for bispecific formats, are required for different screening assays. The polypeptide expression system of the present invention allows for any of these or additional formats, in principle, by simply cloning the appropriate sequence added after the CH1 region in the complementing plasmid. Furthermore, the modular organization of the system allows for the expression of new antibody formats without the need to recreate phage display library stocks, since only the construction of a new complementing plasmid is required. The nucleic acid can also be adapted to allow the use of any CH1 region by transferring the 5' sequence from the CH1-coding region downstream to the VH or J region (FR4) at the J-CH1 junction, thus separating the entire constant regions of the VH and heavy chains in two different nucleic acids. The nucleic acid molecule is compatible with conventional methods for expressing Fab fragments in E. coli by simply adding a stop codon after the sequence encoding the upper hinge. However, the amber stop codon at the junction of the heavy chain and gene III sequences in Fab phage display libraries usually results in significantly lower levels of display, thus requiring reformatting of clones after selection, at least in the case of naive repertoire libraries (Lee et al., Journal of immunological methods. 284:119-132, 2004). Expression of Fab fragments in mammalian cells using the same methods used for IgG expression circumvents this need for reformatting, with yields comparable to those typically obtained in E. coli.
[0065] Antibodies produced by this polypeptide expression system can include recombinantly produced chimeric, humanized, and / or human antibodies. In some cases, the antibody is an antibody fragment, such as a Fab, Fv, Fab', scFv, bispecific, or F(ab')2 fragment. In other cases, the antibody is a full-length antibody, such as an intact IgG1, IgG2, IgG3, or IgG4 antibody as defined herein, or other antibodies of another class or isotype.
[0066] The expressed antibody may incorporate any of the features, alone or in combination, as described in sections 1-7 below.
[0067] 1. Antibody affinity Antibodies (e.g., Fab or full-length IgG antibodies) produced by the polypeptide expression systems described herein may have a denaturing activity of ≦1 μM, ≦100 nM, ≦10 nM, ≦1 nM, ≦0.1 nM, ≦0.01 nM, or ≦0.001 nM (e.g., 10 -8 M or less, e.g. 10 -8 M~10 -13 M, e.g. 10 -9 M~10 -13 The antibody may have a dissociation constant (Kd) of 0.05 M.
[0068] In one embodiment, Kd is measured by a radiolabeled antigen binding assay (RIA) performed with the antibody of interest and a Fab version of its antigen, as illustrated by the following assay: The solution binding affinity of the Fab for the antigen is determined in the presence of a titration series of unlabeled antigen ( 125I) It is measured by equilibrating Fab with a minimal concentration of labeled antigen, then capturing the bound antigen with an anti-Fab antibody-coated plate (see, e.g., Chen et al., J. Mol. Biol. 293:865-881 (1999)). To establish the conditions for the assay, MICROTITER® multi-well plates (Thermo Scientific) are coated overnight with 5 μg / ml of capture anti-Fab antibody (Cappel Labs) in 50 mM sodium carbonate (pH 9.6), followed by blocking with 2% (w / v) bovine serum albumin in PBS for 2-5 hours at room temperature (approximately 23°C). For non-adsorbent plates (Nunc #269620), 100 pM or 26 pM [ 125 [I] The antigen is mixed with serial dilutions of the Fab of interest (e.g., consistent with the evaluation of the anti-VEGF antibody, Fab-12, in Presta et al., Cancer Res. 57:4593-4599 (1997)). The Fab of interest is then incubated overnight, although incubation can continue for longer periods (e.g., approximately 65 hours) to ensure equilibrium has been reached. The mixture is then transferred to a capture plate for incubation at room temperature (e.g., 1 hour). The solution is then removed, and the plate is washed eight times with 0.1% polysorbate 20 (TWEEN-20®) in PBS. Once the plate has dried, 150 μl / well of scintillant (MICROSCINT-20™; Packard) is added, and the plate is counted in a TOPCOUNT™ gamma counter (Packard) for 10 minutes. The concentration of each Fab that provides 20% or less of maximum binding is selected for use in the competitive binding assay.
[0069] According to another embodiment, Kd is measured using a surface plasmon resonance assay with a BIACORE®-2000 or BIACORE®-3000 (BIAcore, Inc., Piscataway, NJ) at 25°C using an immobilized antigen CM5 chip at approximately 10 response units (RU). Briefly, a carboxymethylated dextran biosensor chip (CM5, BIAcore Inc.) is activated with N-ethyl-N'-(3-dimethylaminopropyl)-carbodiimide hydrochloride (EDC) and N-hydroxysuccinimide (NHS) according to the supplier's instructions. The antigen is diluted to 5 μg / ml (approximately 0.2 μM) in 10 mM sodium acetate, pH 4.8, and injected at a flow rate of 5 μl / min to achieve approximately 10 response units (RU) of bound protein. After antigen injection, 1 M ethanolamine is injected to block unreacted groups. For kinetic measurements, two-fold serial dilutions of Fab (0.78 nM to 500 nM) are injected in PBS with 0.05% polysorbate 20 (TWEEN-20™) surfactant (PBST) at a flow rate of approximately 25 μl / min at 25 °C. The association rate (k on ) and dissociation rate (k off The equilibrium dissociation constant (Kd) is calculated using a simple one-to-one Langmuir binding model (BIACORE® evaluation software version 3.2) by simultaneously fitting the association and dissociation sensorgrams. off / k on The binding rate was calculated as a ratio. See, e.g., Chen et al., J. Mol. Biol. 293:865-881 (1999). 6 M -1 s -1If the binding rate exceeds 100 kJ / s, the binding rate can be determined by using a fluorescence quenching technique to measure the increase or decrease in the fluorescence emission intensity (excitation = 295 nM, emission = 340 nM, 16 nM bandpass) of 20 nM anti-antigen antibody (Fab form) in PBS pH 7.2 at 25°C in the presence of increasing concentrations of antigen, as measured in a spectrometer, for example, a spectrophotometer equipped with stopped flow (Aviv Instruments) or an 8000 series SLM-AMINCO™ spectrophotometer (ThermoSpectronic) equipped with a stirred cuvette.
[0070] 2. Antibody fragment In certain embodiments, the antibodies produced by the polypeptide expression systems described herein are antibody fragments. Antibody fragments include, but are not limited to, Fab, Fab', Fab'-SH, F(ab')2, Fv, and scFv fragments, as well as other fragments described below. For a review of certain antibody fragments, see Hudson et al., Nat. Med. 9:129-134 (2003). For a review of scFv fragments, see, e.g., Pluckthun, in The Pharmacology of Monoclonal Antibodies, vol. 113, edited by Rosenburg and Moore, (Springer-Verlag, New York), pp. 269-315 (1994); see also International Patent Publication No. WO 93 / 16185, and U.S. Pat. Nos. 5,571,894 and 5,587,458. See, eg, US Pat. No. 5,869,046 for a discussion of Fab and F(ab')2 fragments which contain salvage receptor binding epitope residues and have increased in vivo half-lives.
[0071] Bispecific antibodies are antibody fragments with two antigen-binding sites, which may be bivalent or bispecific. See, e.g., European Patent No. 404,097, International Patent Publication No. WO 1993 / 01161, Hudson et al., Nat. Med. 9:129-134 (2003), and Hollinger et al., Proc. Natl. Acad. Sci. USA 90:6444-6448 (1993). Trispecific and tetraspecific antibodies are also described in Hudson et al., Nat. Med. 9:129-134 (2003).
[0072] Single-domain antibodies are antibody fragments that contain all or part of the heavy chain variable domain or all or part of the light chain variable domain of an antibody. In certain embodiments, single-domain antibodies are human single-domain antibodies (Domantis, Inc., Waltham, MA; see, e.g., U.S. Patent No. 6,248,516 B1).
[0073] 3. Chimeric and Humanized Antibodies In certain embodiments, antibodies (e.g., Fab or full-length IgG antibodies) produced by the polypeptide expression systems described herein are chimeric antibodies. Certain chimeric antibodies are described, for example, in U.S. Pat. No. 4,816,567 and Morrison et al., Proc. Natl. Acad. Sci. USA, 81:6851-6855 (1984). In one example, a chimeric antibody comprises a non-human variable region (e.g., a variable region derived from a mouse, rat, hamster, rabbit, or non-human primate, such as a monkey) and a human constant region. In a further example, a chimeric antibody is a "class-switched" antibody in which the class or subclass has been changed from that of the parent antibody. Chimeric antibodies include antigen-binding fragments thereof.
[0074] In certain embodiments, a chimeric antibody is a humanized antibody. Typically, a non-human antibody is humanized to reduce immunogenicity in humans while maintaining the specificity and affinity of the parent non-human antibody. Generally, a humanized antibody comprises one or more variable domains in which the HVRs, e.g., CDRs (or portions thereof), are derived from a non-human antibody and the FRs (or portions thereof) are derived from human antibody sequences. Optionally, a humanized antibody will also comprise at least a portion of a human constant region. In some embodiments, some FR residues in a humanized antibody are substituted with corresponding residues from the non-human antibody (e.g., the antibody from which the HVR residues are derived), e.g., to restore or improve the specificity or affinity of the antibody.
[0075] Humanized antibodies and methods for their production are reviewed, e.g., in Almagro and Fransson, Biosci. 13:1619-1633 (2008), and in, e.g., Riechmann et al., Nature 332:323-329 (1988), Queen et al., Proc. Nat'l Acad. Sci. USA 86:10029-10033 (1989), U.S. Pat. Nos. 5,821,337, 7,527,791, 6,982,321, and 7,087,409, Kashmiri et al., Methods 36:25-34 (2005) (describing SDR (a-CDR) grafting), Padlan, Mol. Immunol. 28:489-498 (1991) (describing "surface regeneration"). These techniques are further described in Dall'Acqua et al., Methods 36:43-60 (2005) (describing "FR shuffling"), Osbourn et al., Methods 36:61-68 (2005) and Klimka et al., Br. J. Cancer, 83:252-260 (2000) (describing a "guided selection" approach to FR shuffling).
[0076] Human framework regions that can be used for humanization include framework regions selected using the "best fit" method (see, e.g., Sims et al., J. Immunol. 151:2296 (1993)), framework regions derived from consensus sequences of human antibodies of particular subgroups of light or heavy chain variable regions (see, e.g., Carter et al., Proc. Natl. Acad. Sci. USA, 89:4285 (1992) and Presta et al., J. Immunol., 151:2623 (1993)), human mature (somatically mutated) framework regions, or human germline framework regions (see, e.g., Almagro and Fransson, Front. Biosci. 13:1619-1633 (2008)), and framework regions derived from screening FR libraries (see, e.g., Baca et al., J. Biol. Chem. 272:10678-10684 (1997) and Rosok et al., J. Biol. Chem. 271:22611-22618 (1996)).
[0077] 4. Human antibodies In certain embodiments, the antibodies (e.g., Fab or full-length IgG antibodies) produced by the polypeptide expression systems described herein are human antibodies. Human antibodies can be recombinant human antibodies that are independently prepared using various techniques known in the art and then sequence-identified. Human antibodies are generally described in van Dijk and van de Winkel, Curr. Opin. Pharmacol. 5:368-74 (2001) and Lonberg, Curr. Opin. Immunol. 20:450-459 (2008).
[0078] 5. Library-derived antibodies By utilizing the polypeptide expression systems described herein that are useful in phage display systems, antibodies (e.g., Fab or full-length IgG antibodies) generated by the polypeptide expression systems of the present invention can be isolated by screening combinatorial libraries for antibodies with the desired activity(ies). See, for example, Hoogenboom et al. in Methods in Molecular Biology 178:1-37 (O'Brien et al., ed., Human Press, Totowa, NJ, 2001), as well as, for example, McCafferty et al., Nature 348:552-554; Clackson et al., Nature 352:624-628 (1991); Marks et al., J. Mol. Biol. 222:581-597 (1992); Marks and Bradbury, in Methods in Molecular Biology 248:161-175 (Lo, ed., Human Press, Totowa, NJ, 2003), Sidhu et al., J. Mol. Biol. 338(2):299-310 (2004), Lee et al., J. Mol. Biol. 340(5):1073-1093 (2004), Fellouse, Proc. Natl. Acad. Sci. USA 101(34):12467-12472 (2004), and Lee et al., J. Immunol. Methods 284(1-2):119-132 (2004).
[0079] 6. Multispecific antibodies In certain embodiments, the antibody (e.g., Fab or full-length IgG antibody) produced by the polypeptide expression system described herein is a multispecific antibody, e.g., a bispecific antibody. A multispecific antibody is a monoclonal antibody that has binding specificities for at least two different sites. In certain embodiments, one of the binding specificities is for a first antigen and the other is for any other antigen. In certain embodiments, a bispecific antibody can bind to two different epitopes of a first antigen. Bispecific antibodies can also be used to localize cytotoxins to cells expressing the first antigen. Bispecific antibodies can be prepared as full-length antibodies or antibody fragments.
[0080] Engineered antibodies with three or more functional antigen binding sites, including "octopus antibodies," are also included herein (see, e.g., US 2006 / 0025576A1).
[0081] The antibodies or fragments herein also include "dual acting FAbs" or "DABs" that contain an antigen binding site that binds to a first antigen as well as another, different antigen (see, e.g., US 2008 / 0069820).
[0082] 7. Antibody Variants In certain embodiments, amino acid sequence variants of the antibodies provided herein are contemplated. For example, it may be desirable to improve the binding affinity and / or other biological properties of the antibody. Amino acid sequence variants of an antibody can be prepared by introducing appropriate modifications into one or more sequences of a nucleic acid molecule encoding all or a portion of the antibody. Such modifications include, for example, deletion of, and / or insertion of, and / or substitution of, residues within the amino acid sequence of the antibody. Any combination of deletion, insertion, and substitution can be made to arrive at the final construct, provided that the final construct possesses the desired characteristics, e.g., antigen binding.
[0083] In certain embodiments, a collection of antibody variants having one or more amino acid substitutions relative to each other can be generated by the expression systems and methods of the invention. Sites of interest for substitutional mutagenesis include HVRs and FRs. Conservative substitutions are shown in Table 1 under the heading "Conservative Substitutions." More substantial changes are provided in Table 1 under the heading "Exemplary Substitutions" and are further described below in relation to amino acid side chain classes. Amino acid substitutions can be introduced into the antibody of interest and products screened for a desired activity, e.g., maintained / improved antigen binding, reduced immunogenicity, or improved ADCC or CDC. TIFF2025118659000002.tif208170
[0084] Amino acids can be classified according to common side chain properties: (1) Hydrophobic: Norleucine, Met, Ala, Val, Leu, Ile; (2) Neutral hydrophilic: Cys, Ser, Thr, Asn, Gln; (3) Acidic: Asp, Glu; (4) basic: His, Lys, Arg; (5) residues that affect chain orientation: Gly, Pro; (6) Aromatic: Trp, Tyr, Phe
[0085] Non-conservative substitutions would involve exchanging a member of one of these classes for another class.
[0086] One type of substitutional variant involves substituting one or more hypervariable region residues of a parent antibody (e.g., a humanized or human antibody). Generally, the resulting variant(s) selected for further study will have a modified (e.g., improved) biological property relative to the parent antibody (e.g., increased affinity, reduced immunogenicity) and / or will substantially retain a biological property of the parent antibody. A typical substitutional variant is an affinity-matured antibody, which can be conveniently generated using phage-display-based affinity maturation methods such as those described herein. Briefly, one or more HVR residues are mutated, and the variant antibodies are displayed on phage and screened for a particular biological activity (e.g., binding affinity).
[0087] Alterations (e.g., substitutions) can be made in HVRs, e.g., to improve antibody affinity. Such alterations can be made in HVR "hotspots," i.e., residues encoded by codons that undergo frequent mutation during the somatic maturation process (see, e.g., Chowdhury, P.S., Methods Mol. Biol. 207:179-196 (2008)), and / or in SDRs (a-CDRs), and the resulting variant VH or VL are tested for binding affinity. Affinity maturation methods by constructing and reselecting from secondary libraries are described, for example, in Hoogenboom, HR et al. in Methods in Molecular Biology 178:1-37 (2001) (O'Brien et al., ed., Human Press, Totowa, NJ). In some embodiments of affinity maturation methods, diversity is introduced into the variable region genes chosen for maturation by any of a variety of methods (e.g., error-prone PCR, chain shuffling, or oligonucleotide-directed mutagenesis). A secondary library is then created. The library is then screened to identify any antibody variants with the desired affinity. Another method for introducing diversity involves an HVR-directed approach, in which several HVR residues (e.g., 4-6 residues at a time) are randomized. HVR residues involved in antigen binding can be specifically identified, for example, using alanine scanning mutagenesis or modeling. CDR-H3 and CDR-L3 in particular are often targeted.
[0088] In certain embodiments, substitutions, insertions, or deletions may be made within one or more HVRs as long as such changes do not substantially reduce the antibody's ability to bind antigen. For example, conservative changes (e.g., conservative substitutions provided herein) that do not substantially reduce binding affinity may be made in HVRs. Such changes may be located outside of HVR "hot spots" or SDRs. In certain embodiments of the variant VH and VL sequences provided above, each HVR is either unaltered or contains no more than one, two, or three amino acid substitutions.
[0089] A useful method for identifying antibody residues or regions that can be targeted for mutagenesis is called "alanine scanning mutagenesis," described by Cunningham and Wells (1989) Science, 244:1081-1085. In this method, a residue or group of target residues (e.g., charged residues such as arg, asp, his, lys, and glu) is identified and replaced with neutral or negatively charged amino acids (e.g., alanine or polyalanine) to determine whether the interaction of the antibody with the antigen is affected. Further substitutions can be introduced at amino acid positions that demonstrate functional sensitivity to the initial substitutions. Alternatively, or in addition, a crystal structure of the antigen-antibody complex can be used to identify contact points between the antibody and antigen. Such contact residues and neighboring residues can be targeted as candidates for substitution or eliminated. Mutants can be screened to determine whether they contain the desired properties.
[0090] Amino acid sequence insertions include amino- and / or carboxyl-terminal fusions ranging in length from one residue to polypeptides containing 100 or more residues, as well as intrasequence insertions of single or multiple amino acid residues. An example of a terminal insertion includes an antibody with an N-terminal methionyl residue. Other insertional variants of the antibody molecule include the fusion to the N- or C-terminus of the antibody to an enzyme (e.g., for ADEPT) or a polypeptide which increases the serum half-life of the antibody.
[0091] Although the concept of modular protein expression by pre-mRNA trans-splicing is described in detail herein in the context of a phage antibody display vector system, the application of the concept exemplified by the use of the nucleic acid molecules, vectors, vector sets, host cells and methods described herein can be adapted and extended to other technologies that require, for example, the expression in mammalian cells of a large collection of proteins with different combinations of repeat modules.
[0092] III. Working Examples The following are examples of the present invention: It should be understood that various other embodiments may be practiced based on the general description provided above.
[0093] Example 1. Generation of a modular protein expression system for antibody reformatting in conjunction with phage display vectors The creation of a polypeptide expression system for modular expression and production of polypeptides is described. The invention is based, at least in part, on experimental findings showing that pre-mRNA trans-splicing can be exploited in mammalian cells to enable modular recombinant protein expression. The concept of modular protein expression allows for the precise joining of any two protein-coding sequences encoded by two different constructs into a single mRNA encoding a polypeptide chain, without any of the requirements and constraints of other protein-protein splicing methods. The concept of modular protein expression by pre-mRNA trans-splicing can be adapted to simplify and extend other technologies requiring the expression in mammalian cells of large collections of proteins with different combinations of repeating modules. For example, this concept will find use in other settings requiring the expression of fusion protein partners or combinations of mutations in a single polypeptide. This technology is both simple and effective, allows for application at any scale, and has broad significance for the field of recombinant protein expression in mammalian cells, the foundation of much modern biotechnology.
[0094] Here, we describe the creation of such a polypeptide expression system that allows for modular expression of different antibody formats in the context of a phage display expression system. Phage display is widely used in the discovery and engineering of antibody fragments for the development of therapeutic and reagent antibodies (McCafferty et al., Nature. 348:552-554, 1990; Sidhu, Current Opinion in Biotechnology. 11:610-616, 2000; Smith, Science. 228:1315-1317, 1985). While phage display traditionally allows for rapid selection of antigen-specific binders, screening of selected antibody fragments is limited. Detailed characterization of antibody fragments often requires expression of full-length immunoglobulin G (IgG), which is typically expressed in mammalian cells. However, one limiting step in this process is the reformatting of phage clones into mammalian expression vectors for IgG expression. High-throughput subcloning methods can be used to reformat large numbers of clones, but these methods are usually relatively labor-intensive and produce many clones that will not be used beyond the screening stage.
[0095] To circumvent the need for subcloning and enable modular protein expression, a first nucleic acid molecule:dual host vector, pDV2, was generated (Figure 1). Unlike the previously described dual vector pDV, which contains an IgG expression cassette with an engineered signal sequence for expression of the heavy chain in either bacteria or mammalian cells and requires co-transfection of mammalian cells with a mammalian expression vector expressing the light chain for full IgG expression (Tesar et al., Protein engineering, design & selection: PEDS. 26:655-662, 2013), pDV2 contains most of the stII signal sequence embedded in a bacterial promoter and intron that is removed by splicing in mammalian cells.
[0096] The stII signal sequence in pDV2 was modified to include both a 3′ splice site (3′ ss) and an optimized polypyrimidine tract (PPT) preceding the 3′ ss. This required the introduction of three relatively conservative amino acid substitutions in the stII signal sequence, which did not affect the display of Fab fragments on phage (Figure 2). To allow for modular and flexible expression of antibody formats from the same clone, we did not add a complete intron and exon encoding the constant region downstream from the region encoding the CH1 domain. Instead, we attempted to add these heavy chain sequences in trans from a second nucleic acid molecule. To achieve this, we exploited the process of pre-mRNA trans-splicing, which joins two different pre-mRNAs to form a single mature mRNA. Hybridization of complementary sequences downstream from the 5' ss and upstream from the 3' ss can induce trans-splicing in mammalian cells, resulting in a pre-mRNA with a unique 5' ss and 3' ss, forming a single, noncovalently linked pre-mRNA, which is then spliced as a regular pre-mRNA (Konarska et al., Cell. 42:165-171, 1985; Puttaraju et al., Nature biotechnology. 17:246-252, 1999; Solnick, Cell. 42:157-164, 1985). In this particular polypeptide expression system, a 150-bp fragment of M13 gene III (gIII) was used as the hybridizing sequence (Figure 1). This gene III sequence follows a previously described optimized GTAAGA 5′ sequence at the 3′ boundary of the CH1 coding sequence (Tesar et al., Protein engineering, design & selection: PEDS. 26:655-662, 2013).
[0097] To complete the polypeptide expression system, we generated a second nucleic acid molecule, pRK-Fc, a complementary plasmid that expresses a pre-mRNA containing a 150-nt antisense gene III sequence followed by a linker sequence, consensus branch point, and PPT, as well as a 3′ ss followed by a hinge, CH2 and CH3 regions in one exon, and an SV40 polyadenylation signal (Figures 1 and 3). This transcript does not encode a signal sequence, and the first two potential initiation codons are located out of frame in the antisense gene III sequence and the hinge region. Thus, with the exception of the 5′ ss, all other sequences required for splicing are encoded by pRK-Fc rather than pDV2. Cotransfection of Expi293F cells (Invitrogen) with pDV2 and pRK-Fc resulted in baseline but detectable expression levels of IgG (Figure 4A).
[0098] Example 2. Generation of an optimized modular protein expression system for antibody reformatting in conjunction with phage display vectors The baseline IgG yields achieved with pDV2 and pRK-Fc may have been due to the absence of sequences required for efficient trans-splicing or sequences in the vector that inhibit trans-splicing. Nucleotide motifs in both exons and introns can act as splicing enhancers, suppressors, or both, depending on their location. For vector design purposes, intron splice enhancers (ISEs) are easily added because they are unlikely to affect coding sequences in mammalian cell expression. One well-described ISE consists of a sequence, or G-run, of three or more consecutive guanine residues located close to an intron boundary, which is bound by heterogeneous nuclear ribonucleoproteins H or F to enhance splicing (Wang et al., Nature structural & molecular biology. 19:1044-1052, 2012; Xiao et al., Nature structural & molecular biology. 16:1094-1100, 2009). In addition, purine-rich intronic sequences adjacent to the 5'ss, not limited to G-runs, have also been shown to enhance splicing (Hastings et al. RNA. 7:859-874, 2001).
[0099] Therefore, we created a variant of pDV2, pDV2b, containing a 23-base pair (bp) purine-rich region (26 bp) downstream from the 5′ ss, with a 9-nt G-run in the region encoding the linker between the upper hinge and the C-terminal end of the M13 bacteriophage pIII coat protein (cP3), and a second 4-nt G-run 10 nt downstream (Figure 5). This variant changes the Gly-Arg-Pro linker between the upper hinge and cP3 to three Gly residues. The vector did not contain a standard polyadenylation site for the heavy chain cassette. This was done to minimize the formation of mature heavy chain mRNA from the vector, which could then be transported to the cytoplasm and undergo trans-splicing, potentially leading to the expression of the Fab-cP3 fusion protein. The pRK-Fc molecule was also optimized. An intron G-run near the 3′ split has been shown to stimulate splicing in vitro (Martinez-Contreras, PLoS biology. 4:e21, 2006). Therefore, a 9-nt ISE was added upstream from the branch site to generate the optimized complementation plasmid pRK-Fc2 (Figure 6). Cotransfection of human Expi293F cells with pDV2 and pRK-Fc (ISE-) or pRK-Fc2 (ISE+) resulted in baseline levels of IgG expression (Figure 4A). Cotransfection of Expi293F cells with the ISE+ pDV2b plasmid and pRK-Fc or pRK-Fc2 resulted in higher levels of IgG expression, with the highest expression level of up to 25 μg / ml produced by cotransfecting the ISE+ plasmids pDV2b and pRK-Fc2, indicating that the ISE sequence in both transcripts enhances the efficiency of trans-splicing.
[0100] Baseline IgG expression levels in transfected Expi293F cells correlated with apparent cell lysis 7 days after transfection, observed when pDV2 or pDV2b, but not pRK-Fc or pRK-Fc2, was transfected alone. Analysis of transfected cell lysates by Western blotting using anti-M13p3 antibody revealed a polypeptide with an apparent molecular weight of approximately 41 kDa, consistent with expression of the IgG1Fd fragment (VH-CH1-upper hinge) fused to the M13cP3 peptide (Figure 12, lower panel, columns 3-6). Expression of this polypeptide was higher in cells transfected with pDV2 or pDV2b without a complementing plasmid. The results demonstrated that both pDV2 and pDV2b plasmids were capable of expressing mature mRNA encoding potentially toxic products, despite the fact that the vector lacked a mammalian polyadenylation site downstream from the heavy chain cassette.
[0101] Visual inspection of the gene III sequence encoding cP3 revealed an AATAAA motif that could act as a polyadenylation site (Figure 2). Two silent mutations were introduced into this site to generate the plasmids pDV2c(ISE-) and pDV2d(ISE+) to test whether this would reduce toxicity and improve protein expression in mammalian cells. Cotransfection of Expi293F cells with pRK-Fc and pDV2c or pDV2d resulted in approximately sixfold higher levels of IgG expression relative to the pDV2 and pDV2b vectors (Figure 4A), due to the potential polyadenylation site in gene III. This increased IgG expression level was associated with high viability of the transfected cells and significantly reduced or undetectable expression of the Fd-cP3 fusion protein in the transfected cells (Figure 12, lower panel, rows 7–10). This indicates that the presence of a potential polyadenylation site in gene III in the donor vector causes unwanted protein expression from the donor plasmid alone, which has a significant negative impact on protein expression. Cotransfection of Expi293F cells with pDV2c or pDV2d and the ISE+pRK-Fc2 complementing vector resulted in a further 2-fold increase in IgG expression compared to cotransfection with the ISE-pRK-Fc vector (Figure 4A). These results indicate that the primary factor determining baseline protein expression in the pDV2 vector is the presence of a potential polyadenylation site in gene III, whereas the addition of an ISE has a minor effect on protein expression when a potential gene III polyadenylation motif is absent. In contrast, the addition of an ISE in the complementing pRK-Fc2 plasmid results in approximately 2-fold higher IgG yields when cotransfecting a pDV2 mutant without a potential polyadenylation site in gene III (Figure 4A).
[0102] Further optimization of protein expression was achieved by determining the optimal DNA ratio for transfection. Using a 2:1 excess of the complementary plasmid pRK-Fc2 relative to pDV2d resulted in the highest IgG expression yield in this system (Figure 4B). Using the optimized DNA ratio of pDV2d and pRK-Fc2, the yield of IgG purified from 30 ml of supernatant of transfected Expi293F cells was 3.2 ± 1.2 mg (n = 3). The IgG purified from Expi293F cells cotransfected with these plasmids was indistinguishable from the same IgG expressed by conventional expression vectors by mass spectrometry and SDS-PAGE (Figures 7A-7B and 13). Cotransfection of Expi293F cells with pDV2d encoding variable regions of different specificities and pRK-Fc2 with an optimized DNA ratio resulted in high IgG expression, with 2.5 to 5.5 mg of IgG purified from 30 ml of supernatant from transfected Expi293F cells (Figure 8A). The polypeptide expression system is not limited to the use of Expi293F cells to achieve high expression levels. Other mammalian cell lines widely used for IgG expression, such as 293T and CHO cells, were also effective. Cotransfection of 293T or CHO cells with pDV2d and pRK-Fc2 expressing variable regions of different specificities resulted in high IgG expression (Figure 8B).
[0103] The pRK-Fc2 vector was modified for expression of Fab fragments when cotransfected with the pDV2 plasmid. The sequences encoding the lower hinge and Fc region in pRK-Fc2 were removed and replaced with a Flag tag to produce the pRK-Fab-Flag vector (Figure 9). The yield of purified Fab fragments purified from 30 ml of supernatant of Expi293F cells cotransfected with pDV2d and pRK-Fab-Flag was 0.8 ± 0.06 mg (mean ± standard deviation, n = 3). The structural accuracy of the purified Flag-tagged Fab fragments was confirmed by mass spectrometry and SDS-PAGE (Figure 13). The observed heavy chain mass, excluding the clipped C-terminal lysine, was 25,169 Da, close to the expected mass of 25,172 Da.
[0104] Expression of N-terminally truncated proteins from complementary transcripts has been observed in trans-splicing systems for gene therapy (Monjaret et al., Molecular therapy 22:1176-1187, 2014). This may be due to the complementary transcript encoding a 3′ exon containing all the elements necessary for the formation of mature mRNA, leading to translation from an internal start codon. Western blotting of lysates from cells transfected with pRK-Fc2 revealed the expression of a polypeptide corresponding to the Fc fragment translated from the first in-frame ATG codon (Figure 12, top panel, column 11). This polypeptide likely lacks a secretory signal sequence and is expressed exclusively in the cytoplasm. Although this product may be released into the culture medium upon cell lysis, we did not observe it in purified IgG samples by SDS-PAGE (Figure 13, column 2) and mass spectrometry. When pDV2c or pDV2d was cotransfected into cells, expression of this truncated product was reduced, but not eliminated (Figure 12, top panel, columns 8 and 10). Insertion of an out-of-frame open reading frame with an optimal translation initiation site into an intronic region upstream from the potential Fc start codon did not significantly reduce expression of the truncated Fc product.
[0105] A key characteristic of phage display vectors that determines selection efficiency is the level of antibody fragment display on phage particles that is achieved. Using the Amber-2614KO7 helper phage described above, along with reduced p3 expression in the E. coli SupE suppressor strain, the level of Fab fragment display achieved with the pDV2d vector was comparable to that achieved with the specialized Fab display vector, Fab-Zip-phage, using standard M13KO7 helper phage (Figure 14).
[0106] The ability to express different antibody formats from the same clone is useful in antibody discovery when different antibody formats, such as wild-type IgG, Fab fragments, or IgG with Fc modifications for bispecific formats, are required for different screening assays. The vector set allows for any of these or additional formats, in principle, by simply cloning the appropriate sequence added after the CH1 region in the complementing plasmid. Furthermore, the modular organization of this system allows for the expression of new antibody formats without the need to recreate phage display library stocks, since it only requires the construction of a new complementing plasmid. The dual vector can also be adapted to allow the use of any CH1 region by transferring the 5' sequence from the CH1 coding region downstream to the VH or J region (FR4) at the J-CH1 junction, thus separating the entire constant regions of the VH and heavy chains in two different plasmids. With the knowledge that an amber stop codon at the junction of the heavy chain and gene III sequences in Fab phage display libraries usually results in significantly lower levels of display and therefore requires reformatting of clones after selection, at least for naive repertoire libraries, the pDV2 vector is compatible with conventional methods for expression of Fab fragments in E. coli by simply adding a stop codon after the sequence encoding the upper hinge (Lee et al. Journal of immunological methods. 284:119-132, 2004). Expression of Fab fragments in mammalian cells using the same methods used for IgG expression circumvents this need for reformatting, with yields comparable to those typically obtained in E. coli.
[0107] Example 3. Modular protein expression system The polypeptide expression systems generated and characterized in Examples 1 and 2 demonstrate that modular and flexible polypeptide expression of any desired protein can be achieved straightforwardly through the use of polypeptide expression systems such as the optimized expression systems described above for protein reformatting in the context of phage display. Thus, the expression system comprises two nucleic acid molecule components (polypeptide coding sequences PES11 and PES2), each encoding a portion of a single desired polypeptide product, where the split coding regions of these proteins are precisely joined together in vivo by pre-mRNA trans-splicing, without the need for subcloning of the protein-encoding nucleic acid. As shown in Figure 10, a first nucleic acid molecule comprises an expression cassette with PES11 and a eukaryotic promoter (P1) upstream of the PES11 component. Euk1 The complementary second nucleic acid molecule also contains a eukaryotic promoter (P2) and a eukaryotic signal sequence (ESS11), as well as a 5' splice site (5'ss11) and a hybridizing sequence (HS1) located downstream of PES11. Euk The second nucleic acid molecule will contain a hybridizing sequence capable of hybridizing to HS1 (HS2), as well as a 3' splice site (3'ss2) upstream of the PES2 component. In addition, the second nucleic acid molecule will contain a polyadenylation site (pA2) downstream of the PES2 component. Thus, when transcribed in mammalian cells, the two generated pre-mRNA molecules, one with a unique 5' splice and the other with a unique 3' splice, will be directed together by their complementary hybridizing sequences (HS1 and HS2) and undergo trans-splicing to form a single continuous mRNA that can subsequently be translated and encode the desired protein product.
[0108] In prokaryotic cells, when expression of a polypeptide product encoded by the PES11 and optionally the HS1 region is also desired, the first nucleic acid molecule may be P1 Euk1 and PES11. ePPM1 may further comprise an excisable prokaryotic promoter module (ePPM1) located between PES11 and PES12. ePPM1 may comprise a 5' splice site (5'ss12), a prokaryotic promoter (P1 Prok1), a first nucleic acid sequence encoding a prokaryotic signal sequence (PSS11), and a 3' splice site (3'ss11), and 5'ss12-P1 Prok1 -PSS11-3'ss11 are operably linked to each other in the 5' to 3' direction. ePPM1 will drive transcription of the polypeptide encoded by the first nucleic acid molecule in prokaryotic cells. Meanwhile, in eukaryotic cells (e.g., mammalian cells), P1 Euk1 drives the expression of the transcription of the polypeptide encoded by the first nucleic acid molecule, and ePPM1 will be removed from the pre-mRNA transcript by cis-splicing due to the collisions with the flanking 5'ss12 and 3'ss11 elements.
[0109] In some cases, it may be desirable to express a second polypeptide. Thus, the first nucleic acid molecule of the modular protein expression system can be designed to include a second expression cassette. As shown in Figure 11, the second expression cassette encoding the second protein product (PES12) will be designed in a similar manner to the first expression cassette, but will contain a polyadenylation site (pA1) downstream of the PES12 sequence to ensure the production of a separate pre-mRNA molecule after transcription. In other cases, a second expression cassette can be designed into the second nucleic acid molecule of the polypeptide expression system.
[0110] Other embodiments The foregoing invention has been described in detail by way of illustration and example for purposes of clarity of understanding, but the descriptions and examples should not be construed as limiting the scope of the invention. The disclosures of all patent and scientific literature set forth herein are expressly incorporated by reference in their entirety.
Claims
1. 1. A polypeptide expression system comprising a first nucleic acid molecule and a second nucleic acid molecule, (a) the first nucleic acid molecule comprises the following components: (i) a first eukaryotic promoter (P1 Euk1 ), (ii) a first polypeptide coding sequence (PES1 1 ), (iii) the first 5' splice site (5'ss1 1 ), and (iv) a hybridizing sequence (HS1), wherein the components comprise: P1 Euk1 -PES1 1 -5'ss1 1 -HS1, and are operably linked to each other in the 5' to 3' direction; (b) the second nucleic acid molecule comprises the following components: (i) a eukaryotic promoter (P2 Euk ), (ii) a hybridizing sequence (HS2) capable of hybridizing to HS1, (iii) a 3' splice site (3'ss2), (iv) a polypeptide coding sequence (PES2), and (v) a polyadenylation site (pA2), wherein the components are P2 Euk -HS2-3'ss2-PES2-pA2, which are operably linked to each other in the 5' to 3' direction.
2. P1 Euk1 2. The polypeptide expression system of claim 1, wherein the promoter is a cytomegalovirus (CMV) promoter or a simian virus 40 (SV40) promoter.
3. P2 Euk The polypeptide expression system according to claim 1 or 2, wherein the promoter is a CMV promoter or an SV40 promoter.
4. The first expression cassette contains a eukaryotic signal sequence (ESS1 1 ESS1), 1 However, the P1 Euk1 and the PES1 1 The polypeptide expression system according to any one of claims 1 to 3, wherein the polypeptide expression system is located between
5. The ESS1 1 5. The polypeptide expression system of claim 4, wherein said VH gene is derived from a variable heavy chain (VH) gene.
6. The first expression cassette comprises the following components: (i) a 5' splice site (5'ss1 2 ), (ii) prokaryotic promoter (P1 Prok1 ), and (iii) the 3' splice site (3'ss1 1 ) an excisable prokaryotic promoter module (ePPM) 1 ) wherein the component further comprises 5'ss1 2 -P1 Prok1 -3'ss1 1 and the ePPM is operably linked to each other in the 5'-3' direction as 1 However, the P1 Euk1 and the PES1 1 The polypeptide expression system according to any one of claims 1 to 5, wherein the polypeptide expression system is located between
7. P1 Prok1 The polypeptide expression system of claim 6, wherein the promoter is selected from the group consisting of a PhoA promoter, a Tac promoter, a Lac promoter, and a Tphac promoter.
8. The ePPM 1 However, the prokaryotic signal sequence (PSS1 1 8. The polypeptide expression system of claim 6 or 7, further comprising a first nucleic acid sequence encoding a polypeptide of the present invention.
9. PSS1 1 The polypeptide expression system according to any one of claims 6 to 8, wherein is derived from the heat-stable enterotoxin II (stII) gene.
10. PSS1 1 and the 3'ss1 1 The polypyrimidine region (PPT1) is located between 1 The polypeptide expression system according to any one of claims 6 to 9, further comprising:
11. PPT1 1 The polypeptide expression system of claim 10, wherein said nucleic acid sequence is TTCCTTTTTTTCTCTTTCC (SEQ ID NO: 1).
12. PES1 1 The polypeptide expression system according to any one of claims 1 to 11, wherein the expression system does not contain a cryptic 5' splice site.
13. The polypeptide expression system according to any one of claims 1 to 12, wherein the HS1 is a gene encoding all or part of a coat protein or an adaptor protein.
14. 14. The polypeptide expression system of claim 13, wherein the coat protein is selected from the group consisting of pI, pII, pIII, pIV, pV, pVI, pVII, pVIII, pIX, and pX of bacteriophage M13, f1, or fd.
15. 15. The polypeptide expression system of claim 14, wherein the coat protein is the pIII protein of bacteriophage M13.
16. 16. The polypeptide expression system of claim 15, wherein the pill fragment comprises amino acid residues 267 to 421 of the pill protein or amino acid residues 262 to 418 of the pill protein.
17. The polypeptide expression system of claim 13, wherein the adapter protein is a leucine zipper.
18. 18. The polypeptide expression system of claim 17, wherein the leucine zipper comprises the amino acid sequence of SEQ ID NO: 4 or 5.
19. The first nucleic acid molecule is a second eukaryotic promoter (P1 Euk2 ), (ii) a second polypeptide coding sequence (PES1 2 ), and (iii) a second expression cassette comprising a polyadenylation site (pA1), wherein said components are Euk2 -PES1 2 The polypeptide expression system according to any one of claims 1 to 18, wherein the polypeptides are operably linked to each other in the 5' to 3' direction as -pA1.
20. P1 Euk2 The polypeptide expression system of claim 19, wherein the promoter is a CMV promoter or an SV40 promoter.
21. The second expression cassette contains a eukaryotic signal sequence (ESS1 2 21. The polypeptide expression system of claim 19 or 20, further comprising a nucleic acid sequence encoding a polypeptide.
22. The ESS1 2 22. The polypeptide expression system of claim 21, wherein said gene is derived from the mouse binding immunoglobulin protein (mBiP) gene.
23. The ESS1 2 23. The polypeptide expression system of any one of claims 19 to 22, wherein said nucleic acid sequence comprises the nucleic acid sequence ATG AAN TTN ACN GTN GTN GCN GCN GCN GCN CTN CTN CTN CTN GGN (SEQ ID NO: 6), wherein N is A, T, C, or G.
24. The second expression cassette comprises the following components: (i) a 5' splice site (5'ss1 3 ), (ii) prokaryotic promoter (P1 Prok2 ), and (iii) the 3' splice site (3'ss1 2 ) an excisable prokaryotic promoter module (ePPM) 2 ) wherein the component further comprises 5'ss1 3 -P1 Prok2 -3'ss1 2 and the ePPM is operably linked to each other in the 5'-3' direction as 2 However, the P1 Euk2 and the PES1 2 The polypeptide expression system according to any one of claims 19 to 23, which is located between
25. P1 Prok2 25. The polypeptide expression system of claim 24, wherein the promoter is selected from the group consisting of a PhoA promoter, a Tac promoter, and a Lac promoter.
26. The ePPM 2 However, the prokaryotic signal sequence (PSS1 2 26. The polypeptide expression system of claim 24 or 25, further comprising a nucleic acid sequence encoding
27. PSS1 2 The polypeptide expression system according to any one of claims 24 to 26, wherein is derived from the heat-stable enterotoxin II (stII) gene.
28. PSS1 2 and the 3'ss1 2 The polypyrimidine region (PPT1) is located between 2 28. The polypeptide expression system according to any one of claims 24 to 27, further comprising:
29. PPT1 2 29. The polypeptide expression system of claim 28, wherein said nucleic acid sequence is TTCCTTTTTTTCTCTTTCC (SEQ ID NO: 1).
30. 30. The polypeptide expression system of any one of claims 19 to 29, wherein the second expression cassette is positioned 5' to the first expression cassette.
31. The 5'ss1 1 31. The polypeptide expression system of claim 1, further comprising an intron splice enhancer (ISE) (ISE1) located between said HS1 and said HS2.
32. The polypeptide expression system of claim 31, wherein the ISE1 comprises a G-run containing three or more consecutive guanine residues.
33. The polypeptide expression system of claim 32, wherein the ISE1 comprises a G-run containing nine consecutive guanine residues.
34. 34. The polypeptide expression system of any one of claims 1 to 33, further comprising a polypyrimidine tract (PPT2) located between the HS2 and the 3'ss2.
35. 35. The polypeptide expression system of claim 34, wherein the PPT2 comprises the nucleic acid sequence TTCCTCTTTCCCTTTCTCTCCC (SEQ ID NO: 7).
36. 36. The polypeptide expression system of claim 35, further comprising an ISE (ISE2) located between the HS2 and the 3'ss2.
37. The polypeptide expression system of claim 36, wherein the ISE2 comprises a G-run containing three or more consecutive guanine residues.
38. The polypeptide expression system of claim 37, wherein the ISE2 comprises a G-run containing nine consecutive guanine residues.
39. The 5'ss1 1 39. The polypeptide expression system of any one of claims 1 to 38, wherein said expression vector comprises the nucleic acid sequence GTAAGA (SEQ ID NO: 8).
40. 40. The polypeptide expression system of any one of claims 1 to 39, wherein expression by a eukaryotic promoter occurs in mammalian cells.
41. 41. The polypeptide expression system of claim 40, wherein the mammalian cells are Expi293F cells, CHO cells, 293T cells, or NSO cells.
42. 42. The polypeptide expression system of claim 41, wherein the mammalian cells are Expi293F cells.
43. 43. The polypeptide expression system according to any one of claims 6 to 42, wherein expression by a prokaryotic promoter occurs in a bacterial cell.
44. 44. The polypeptide expression system of claim 43, wherein the bacterial cell is an E. coli cell.
45. PES1 1 45. A polypeptide expression system according to any one of claims 1 to 44, wherein said expression system encodes all or part of an antibody.
46. PES1 1 46. The polypeptide expression system of claim 45, wherein said VH domain is a polypeptide comprising said VH domain.
47. 47. The polypeptide expression system of claim 46, wherein the polypeptide further comprises a CH1 domain.
48. 48. The polypeptide expression system of any one of claims 45 to 47, wherein the PES2 encodes all or part of an antibody.
49. 49. The polypeptide expression system of claim 48, wherein the PES2 encodes a polypeptide comprising a CH2 domain and a CH3 domain.
50. PES1 2 50. The polypeptide expression system of any one of claims 19 to 49, wherein said expression system encodes all or part of an antibody.
51. PES1 2 51. The polypeptide expression system of claim 50, wherein said expression system encodes a polypeptide comprising a VL domain and a CL domain.
52. Components include: (a) a first eukaryotic promoter (P1 Euk1 )and, (b) the following components: (i) 5' splice site (5'ss1 2 ), (ii) Prokaryotic promoter (P1 Prok1 ), and (iii) 3' splice site (3'ss1 1 ) a first excisable prokaryotic promoter module (ePPM) comprising 1 ) wherein the ePPM 1 The component of 5'ss1 2 -P1 Prok1 -3'ss1 1 a first excisable prokaryotic promoter module operably linked to each other in a 5' to 3' direction as (c) a first polypeptide coding sequence (PES1 1 )and, (d) the first 5' splice site (5'ss1 1 )and, (e) a useful peptide coding sequence (UPES); and wherein the components of the first expression cassette comprise: P1 Euk1 -ePPM 1 -PES1 1 -5'ss1 1 - the nucleic acid molecules are operably linked to each other in the 5' to 3' direction as UPES.
53. The first expression cassette contains a eukaryotic signal sequence (ESS1 1 ESS1), 1 However, the P1 Euk1 and the ePPM 1 53. The nucleic acid molecule of claim 52, wherein said nucleic acid molecule is located between
54. The ePPM 1 However, the prokaryotic signal sequence (PSS1 1 ), wherein said PSS1 1 However, the P1 Prok1 and the 3'ss1 1 54. The nucleic acid molecule of claim 52 or 53, wherein the nucleic acid molecule is located between
55. A second eukaryotic promoter (P1 Euk2 ), (ii) a second polypeptide coding sequence (PES1 2 ), and (iii) a second expression cassette comprising a polyadenylation site (pA1), wherein said components are Euk2 -PES1 2 -pA1. The nucleic acid molecules according to any one of claims 52 to 54, which are operably linked to each other in the 5' to 3' direction as pA1.
56. The second expression cassette contains a eukaryotic signal sequence (ESS1 2 ESS1), 2 However, the P1 Euk2 and the PES1 2 56. The nucleic acid molecule of claim 55, wherein said nucleic acid molecule is located between
57. The second expression cassette comprises the following components: (i) a 5' splice site (5'ss1 3 ), (ii) prokaryotic promoter (P1 Prok2 ), and (iii) the 3' splice site (3'ss1 2 ) an excisable prokaryotic promoter module (ePPM) 2 ) wherein the component further comprises 5'ss1 3 -P1 Prok2 -3'ss1 2 and the ePPM is operably linked to each other in the 5'-3' direction as 2 However, the P1 Euk2 and the PES1 2 57. The nucleic acid molecule of claim 55 or 56, wherein said nucleic acid molecule is located between
58. The ePPM 2 However, the prokaryotic signal sequence (PSS1 2 ) and further comprising a nucleic acid sequence encoding said PSS1 2 However, the P1 Prok2 and the 3'ss1 2 58. The nucleic acid molecule of claim 57, wherein said nucleic acid molecule is located between
59. 59. The nucleic acid molecule of any one of claims 52 to 58, wherein the UPES encodes all or part of a useful peptide selected from the group consisting of tags, labels, coat proteins, and adaptor proteins.
60. 60. The nucleic acid molecule of claim 59, wherein the coat protein is selected from the group consisting of pI, pII, pIII, pIV, pV, pVI, pVII, pVIII, pIX, and pX of bacteriophage M13, f1, or fd.
61. 61. The nucleic acid molecule of claim 60, wherein the coat protein is the pIII of bacteriophage M13.
62. A vector comprising the nucleic acid molecule of any one of claims 52 to 61.
63. A vector set comprising a first vector and a second vector, wherein the first and second vectors comprise the first and second nucleic acid molecules, respectively, of the polypeptide expression system described in any one of claims 1 to 51.
64. 64. A host cell comprising the vector of claim 62 or the vector set of claim 63.
65. 65. The host cell of claim 64, wherein the host cell is a prokaryotic cell.
66. 66. The host cell of claim 65, wherein the prokaryotic cell is a bacterial cell.
67. 67. The host cell of claim 66, wherein the bacterial cell is an E. coli cell.
68. 65. The host cell of claim 64, wherein the host cell is a eukaryotic cell.
69. 69. The host cell of claim 68, wherein the eukaryotic cell is a mammalian cell.
70. 70. The host cell of claim 69, wherein the mammalian cell is an Expi293F cell, a CHO cell, a 293T cell, or an NSO cell.
71. 71. The host cell of claim 70, wherein the mammalian cell is an Expi293F cell.
72. 64. A method for producing a polypeptide, comprising culturing a host cell comprising the vector of claim 62 or the vector set of claim 63 in a culture medium.
73. 73. The method of claim 72, wherein the method further comprises recovering the polypeptide from the host cell or the culture medium.