Polypeptide expression systems
The polypeptide expression system uses pre-mRNA trans-splicing to combine protein-coding sequences in eukaryotic cells, addressing the inefficiencies of existing systems by enabling modular and flexible expression of polypeptides without subcloning, thereby optimizing construct usage and expression efficiency.
Patent Information
- Application Number
- JP2025027512
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2014-11-17
- Filing Date
- 2025-02-25
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing polypeptide expression systems require multiple constructs for each combination of modules, leading to a geometric increase in the number of constructs needed, and high-throughput methods are source-intensive and generate unnecessary constructs.
A polypeptide expression system comprising a first and second nucleic acid molecule, each with specific components operably linked, enabling modular expression through pre-mRNA trans-splicing in eukaryotic cells, allowing the precise joining of protein-coding sequences without the need for subcloning.
Enables efficient and flexible expression of various polypeptides, including antibodies, by combining modules into a single mRNA chain, reducing the need for labor-intensive subcloning and minimizing unnecessary construct generation.
Smart Images

Figure 2025097990000001_ABST
Abstract
Description
Technical Field
[0001] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in ASCII format and is hereby incorporated by reference in its entirety. The name of this ASCII copy is P05833-WO_SL.txt, created on June 25, 2015, and is 24,421 bytes in size.
[0002] The present invention relates to a polypeptide expression system for modular expression and production of polypeptides.
Background Art
[0003] Recombinant polypeptides are sometimes expressed as fusions of individual domains or tags for functional or purification purposes. Recombinant DNA methods have traditionally been used to join sequences encoding each module, requiring different constructs for each combination. This poses a problem for techniques related to the expression of many protein collections consisting of repetitive modules joined in various combinations, as the number of constructs increases geometrically as a function of the number of modules used.
[0004] High-throughput systems for subcloning can handle a large number of inserts in parallel, but they are usually source-intensive and generate a large number of constructs that are ultimately not needed after the initial characterization step. Therefore, there is a need that has not yet been addressed in the field of developing polypeptide expression systems that enable modular expression and production of recombinant polypeptides.
Summary of the Invention
[0005] The present invention relates to a polypeptide expression system for modular expression and production of polypeptides.
[0006] In one aspect, the present invention features a polypeptide expression system comprising a first nucleic acid molecule and a second nucleic acid molecule, where (a) the first nucleic acid molecule comprises the following components: (i) a first eukaryotic promoter (P1 Euk1 ), (ii) a first polypeptide coding sequence (PES11), (iii) a first 5' splice site (5'ss11), and (iv) a hybridization sequence (HS1), and these components are operably linked to each other in the 5' to 3' direction as P1 Euk1 -PES11-5'ss11-HS1, (b) the second nucleic acid molecule comprises the following components: (i) a eukaryotic promoter (P2 Euk ), (ii) a hybridization sequence (HS2) capable of hybridizing to HS1, (iii) a 3' splice site (3'ss2), (iv) a polypeptide coding sequence (PES2), and (v) a polyadenylation site (pA2), and these components are operably linked to each other in the 5' to 3' direction as P2 Euk -HS2-3'ss2-PES2-pA2. In some embodiments, P1 Euk1 is a cytomegalovirus (CMV) promoter or a simian virus 40 (SV40) promoter. In some embodiments, P2 Euk is a CMV promoter or an SV40 promoter. In some embodiments, the first expression cassette further comprises a first nucleic acid sequence encoding a eukaryotic signal sequence (ESS11), and ESS11 is positioned between P1 Euk1 and PES11. In some embodiments, ESS11 is derived from a variable heavy chain (VH) gene.
[0007] In some embodiments, the first expression cassette further comprises an excisable prokaryotic promoter module (ePPM1) comprising the following components: (i) a 5' splice site (5'ss12), (ii) a prokaryotic promoter (P1 Prok1 ), and (iii) a 3' splice site (3'ss11), where these components are 5'ss12-P1 Prok1-3'ss11 are operably linked to each other in the 5' to 3' direction, and ePPM1 is located between P1 Euk1 and PES11. In some embodiments, P1 Prok1 is selected from the group consisting of the PhoA promoter, the Tac promoter, Lac, and the Tphac promoter. In some embodiments, ePPM1 further comprises a first nucleic acid sequence encoding a prokaryotic signal sequence (PSS11). In some embodiments, PSS11 is derived from the heat-stable enterotoxin II (stII) gene. In some embodiments, the polypeptide expression system further comprises a polypyrimidine tract (PPT11) located between PSS11 and 3'ss11. In some embodiments, PPT11 comprises the nucleic acid sequence TTCCTTTTTTCTCTTTCC (SEQ ID NO: 1). In some embodiments, PES11 does not contain a potential 5' splice site. In some embodiments, HS1 is a gene encoding all or part of a coat protein or an adapter protein. In some embodiments, the coat protein is selected from the group consisting of pI, pII, pIII, pIV, pV, pVI, pVII, pVIII, pIX, and pX of bacteriophage M13, f1, or fd. In some embodiments, the coat protein is the pIII protein of bacteriophage M13. In some embodiments, the pIII fragment comprises amino acid residues 267-421 of the pIII protein or amino acid residues 262-418 of the pIII protein. In some embodiments, the adapter protein is a leucine zipper. In some embodiments, the leucine zipper comprises the amino acid sequence of SEQ ID NO: 4 or 5.
[0008] In some embodiments, the first nucleic acid molecule further comprises a second eukaryotic promoter (P1 Euk2 ), (ii) a second polypeptide coding sequence (PES12), and (iii) a second expression cassette comprising a polyadenylation site (pA1), wherein these components are P1 Euk2-PES12-pA1 are operably linked to each other in the 5' to 3' direction. In some embodiments, P1 Euk2 is a CMV promoter or an SV40 promoter. In some embodiments, the second expression cassette further comprises a second nucleic acid sequence encoding a eukaryotic signal sequence (ESS12). In some embodiments, ESS12 is derived from the mouse binding immunoglobulin protein (mBiP) gene. In some embodiments, ESS12 comprises the nucleic acid sequence of ATG AAN TTN ACN GTN GTN GCN GCN GCN CTN CTN CTN CTN GGN (SEQ ID NO: 6), where N is A, T, C, or G.
[0009] In some embodiments, the second expression cassette further comprises an excisable prokaryotic promoter module (ePPM2) comprising the following components: (i) a 5' splice site (5'ss13), (ii) a prokaryotic promoter (P1 Prok2 ), and (iii) a 3' splice site (3'ss12), wherein these components are operably linked to each other in the 5' to 3' direction as 5'ss13-P1 Prok2 -3'ss12, and ePPM2 is positioned between P1 Euk2 and PES12. In some embodiments, P1 Prok2is selected from the group consisting of the PhoA promoter, the Tac promoter, and the Lac promoter. In some embodiments, ePPM2 further comprises a nucleic acid sequence encoding a prokaryotic signal sequence (PSS12). In some embodiments, PSS12 is derived from the heat-stable enterotoxin II (stII) gene. In some embodiments, the polypeptide expression system further comprises a polypyrimidine tract (PPT12) positioned between PSS12 and 3’ss12. In some embodiments, PPT12 comprises the nucleic acid sequence of TTCCTTTTTTCTCTTTCC (SEQ ID NO: 1). In some embodiments, the second expression cassette is positioned 5’ to the first expression cassette. In some embodiments, the polypeptide expression system further comprises an intron splice enhancer (ISE) (ISE1) positioned between 5’ss11 and HS1. In some embodiments, ISE1 comprises a G-run containing three or more consecutive guanine residues. In some embodiments, ISE1 comprises a G-run containing nine consecutive guanine residues. In some embodiments, the polypeptide expression system further comprises a polypyrimidine tract (PPT2) positioned between HS2 and 3’ss2. In some embodiments, PPT2 comprises the nucleic acid sequence of TTCCTCTTTCCCTTTCTCTCC (SEQ ID NO: 7). In some embodiments, the polypeptide expression system further comprises an ISE (ISE2) positioned between HS2 and 3’ss2. In some embodiments, ISE2 comprises a G-run containing three or more consecutive guanine residues. In some embodiments, ISE2 comprises a G-run containing nine consecutive guanine residues. In some embodiments, 5’ss11 comprises the nucleic acid sequence of GTAAGA (SEQ ID NO: 8).
[0010] In some embodiments, expression by a eukaryotic promoter occurs in mammalian cells. In some embodiments, the mammalian cells are Expi293F cells, CHO cells, 293T cells, or NSO cells. In some embodiments, the mammalian cells are Expi293F cells. In some embodiments, expression by a prokaryotic promoter occurs in bacterial cells. In some embodiments, the bacterial cells are E. coli cells. In some embodiments, PES11 encodes all or part of an antibody. In some embodiments, PES11 encodes a polypeptide comprising a VH domain. In some embodiments, the polypeptide further comprises a CH1 domain. In some embodiments, PES2 encodes all or part of an antibody. In some embodiments, PES2 encodes a polypeptide comprising a CH2 domain and a CH3 domain. In some embodiments, PES12 encodes all or part of an antibody. In some embodiments, PES12 encodes a polypeptide comprising a VL domain and a CL domain.
[0011] In another aspect, the invention provides the following components: (a) a first eukaryotic promoter (P1 Euk1 ); (b) a first excisable prokaryotic promoter module (ePPM1) comprising: (i) a 5' splice site (5'ss12), (ii) a prokaryotic promoter (P1 Prok1 ), and (iii) a 3' splice site (3'ss11), wherein the components of ePPM1 are operably linked to each other in the 5' to 3' direction as 5'ss12-P1 Prok1 -3'ss11; (c) a first polypeptide coding sequence (PES11); (d) a first 5' splice site (5'ss11); and (e) a useful peptide coding sequence (UPES), and the components of the first expression cassette are P1 Euk1-ePPM1-PES11-5’ss11-UPES are operably linked to each other in the 5' to 3' direction. In some embodiments, the first expression cassette further comprises a first nucleic acid sequence encoding a eukaryotic signal sequence (ESS11), where ESS11 is located between P1 Euk1 and ePPM1. In some embodiments, ePPM1 further comprises a first nucleic acid sequence encoding a prokaryotic signal sequence (PSS11), where PSS11 is located between P1 Prok1 and 3’ss11. In some embodiments, the nucleic acid molecule further comprises a second expression cassette comprising (i) a second eukaryotic promoter (P1 Euk2 ), (ii) a second polypeptide coding sequence (PES12), and (iii) a polyadenylation site (pA1), where these components are operably linked to each other in the 5' to 3' direction as P1 Euk2 -PES12-pA1. In some embodiments, the second expression cassette further comprises a second nucleic acid sequence encoding a eukaryotic signal sequence (ESS12), where ESS12 is located between P1 Prok2 and 3’ss12. In some embodiments, the second expression cassette further comprises an excisable prokaryotic promoter module (ePPM2) comprising the following components: (i) a 5' splice site (5’ss13), (ii) a prokaryotic promoter (P1 Prok2 ), (iii) a nucleic acid sequence encoding a prokaryotic signal sequence (PSS12), and (iv) a 3' splice site (3’ss12), where these components are operably linked to each other in the 5' to 3' direction as 5’ss13-P1 Prok2 -PSS12-3’ss12, and ePPM2 is between P1 Euk2It is positioned between and PES12. In some embodiments, UPES encodes all or part of a useful peptide selected from the group consisting of tags, labels, coat proteins, and adapter proteins. In some embodiments, the coat protein is selected from the group consisting of pI, pII, pIII, pIV, pV, pVI, pVII, pVIII, pIX, and pX of bacteriophage M13, f1, or fd. In some embodiments, the coat protein is pIII of bacteriophage M13.
[0012] In another aspect, the present invention features a vector containing any one of the aforementioned nucleic acid molecules. In another aspect, the present invention features a vector set containing a first vector and a second vector, wherein the first and second vectors each contain the first and second nucleic acid molecules of any of the polypeptide expression systems disclosed herein.
[0013] In another aspect, the present invention features a host cell containing the aforementioned nucleic acid, vector, and / or vector set. In some embodiments, the host cell is a prokaryotic cell. In some embodiments, the prokaryotic cell is a bacterial cell. In some embodiments, the bacterial cell is an E. coli cell. In other embodiments, the host cell is a eukaryotic cell. In some embodiments, the eukaryotic cell is a mammalian cell. In some embodiments, the mammalian cell is an Expi293F cell, a CHO cell, a 293T cell, or an NSO cell. In one embodiment, the mammalian cell is an Expi293F cell.
[0014] In a further aspect, the present invention features a method for producing a polypeptide, which includes culturing a host cell containing one or more of the aforementioned nucleic acid, vector, and / or vector set in a culture medium. In some embodiments, the method further includes recovering the polypeptide from the host cell or the culture medium. BRIEF DESCRIPTION OF THE DRAWINGS
[0015]
Figure 1
Figure 2A
Figure 2B
Figure 3
Figure 4A
Figure 4B
Figure 5A
Figure 5B
Figure 6
Figure 7A
Figure 7B
Figure 8A
Figure 8B
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
DETAILED DESCRIPTION OF THE INVENTION
[0016] I. Definitions As used herein, the term "antibody" is used in the broadest sense and encompasses various antibody structures, including monoclonal antibodies, polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies), and antibody fragments, as long as they exhibit the desired antigen-binding activity, but are not limited thereto.
[0017] The Kabat numbering system is generally used when referring to residues within the variable domains (approximately residues 1-107 of the light chain and residues 1-113 of the heavy chain) (e.g., Kabat et al., Sequences of Immunological Interest. 5th ed. Public Health Service, National Institutes of Health, Bethesda, Md. (1991)). The "EU numbering system" or "EU index" is generally used when referring to residues within the constant region of the immunoglobulin heavy chain (e.g., the EU index reported in the above Kabat et al.). "Kabat's EU index" refers to the residue numbering of human IgG1 EU antibodies. Unless otherwise specified herein, reference to residue numbers within the variable domain of an antibody means residue numbering according to the Kabat numbering system. Unless otherwise specified herein, reference to residue numbers within the constant domain of the heavy chain of an antibody means residue numbering according to the EU numbering system.
[0018] The native basic four-chain antibody unit is a heterotetrameric glycoprotein consisting of two identical light chains (LCs) and two identical heavy chains (HCs). (IgM antibodies consist of five of the basic heterotetrameric units with an additional polypeptide called the J chain and thus contain 10 antigen-binding sites, while secreted IgA antibodies can polymerize to form multivalent aggregates containing 2-5 of the basic four-chain units with the J chain.) In the case of IgG, the four-chain unit is generally about 150,000 daltons. While the two HCs are linked to each other by one or more disulfide bonds depending on the HC isotype, each LC is linked to an HC by one covalent disulfide bond. Also, each HC and LC have regularly spaced intra-chain disulfide bridges. Each HC has a variable domain (VH) at the N-terminus, followed by three constant domains (CH1, CH2, CH3) for each of the α and γ chains, and four Cj domains for the μ and ε isotypes. Each LC has a variable domain (VL) at the N-terminus, followed by a constant domain (CL) at the other end. The VL aligns with the VH, and the CL aligns with the first constant domain (CH1) of the heavy chain. CH1 can be connected to the second constant domain (CH2) of the heavy chain by the hinge region. Certain amino acid residues are thought to form the junctions between the light and heavy chain variable domains. The VH and VL pair to form a single antigen-binding site together. For the structure and properties of various classes of antibodies, see, for example, Basic and Clinical Immunology, 8th edition, Daniel P. Stites, Abba I. Terr and Tristram G. Parslow (eds.), Appleton & Lange, Norwalk, CT, 1994, page 71 and Chapter 6.
[0019] The "CH2 domain" of the human IgG Fc region typically extends from approximately residue 231 to approximately residue 340 of IgG. The CH2 domain is unique in that it is not closely paired with another domain. Instead, two N-linked branched carbohydrate chains intervene between the two CH2 domains of the intact native IgG molecule. The carbohydrate is thought to act as an alternative to domain-domain pairing and may help to stabilize the CH2 domain. Burton, Molec. Immunol. 22:161-206 (1985).
[0020] The "CH3 domain" contains a continuous stretch of residues that are C-terminal to the CH2 domain in the Fc region (i.e., from approximately amino acid residue 341 to approximately amino acid residue 447 of IgG).
[0021] Light chains (LCs) from any vertebrate species can be assigned to one of two clearly distinct types, called kappa and lambda, based on the amino acid sequences of their constant domains. Immunoglobulins can be assigned to various classes or isotypes based on the amino acid sequences of their heavy chain constant domains (CH). There are five classes of immunoglobulins: IgA, IgD, IgE, IgG, and IgM, which have heavy chains named α, δ, γ, ε, and μ, respectively. The γ and α classes are further divided into subclasses based on relatively minor differences in CH sequence and function. For example, humans express the following subclasses: IgG1, IgG2, IgG3, IgG4, IgA1, and IgA2.
[0022] The term "variable" refers to the fact that certain portions of the variable domains have extensive sequence variability among antibodies. The V domains mediate antigen binding and define the specificity of a particular antibody for its particular antigen. However, the variability is not uniformly distributed throughout the entire 110 amino acids of the variable domains. Instead, the V regions consist of a continuous stretch of 15 - 30 relatively invariant amino acids called framework regions (FRs), and shorter regions of extreme variability, each 9 - 12 amino acids in length, called "hypervariable regions" that separate them. Each of the variable domains of the native heavy and light chains contains four FRs that mainly adopt a beta - sheet structure, which are connected by three hypervariable regions that form loops connecting this beta - sheet structure and in some cases forming parts of the beta - sheet structure. The hypervariable regions within each chain are held in close proximity to each other by the FRs and, together with the hypervariable regions from the other chain, contribute to the formation of the antigen - binding site of the antibody (see Kabat et al., Sequences of Proteins of Immunological Interest, 5th ed. Public Health Service, National Institutes of Health, Bethesda, MD, 1991). The constant domains do not directly participate in the binding of the antibody to the antigen but exhibit various effector functions such as the involvement of the antibody in antibody - dependent cell - mediated cytotoxicity (ADCC).
[0023] "Antibody fragment" refers to a molecule other than an intact antibody that contains a part of an intact antibody and binds to an antigen to which the intact antibody binds. Examples of antibody fragments include, but are not limited to, Fv, Fab, Fab’, Fab’ - SH, F(ab’)2; bispecific antibodies; linear antibodies; single - chain antibody molecules (e.g., scFv); and multispecific antibodies formed from antibody fragments.
[0024] The "Fab" fragment is an antigen-binding fragment generated by papain digestion of an antibody and consists of the entire L chain with the variable region domain of the H chain (VH) and the first constant domain (CH1) of one heavy chain. Papain digestion of an antibody produces two identical Fab fragments. Pepsin treatment of an antibody generates a single large F(ab’)2 fragment, which generally corresponds to two disulfide-linked Fab fragments with bivalent antigen-binding activity and can still cross-link antigens. The Fab’ fragment differs from the Fab fragment by having several additional residues including one or more cysteines from the antibody hinge region at the carboxy terminus of the CH1 domain. Fab’-SH is the name in this specification for a Fab’ in which the cysteine residue(s) of the constant domain have free thiol groups. The F(ab’)2 antibody fragment was originally generated as a pair of Fab’ fragments with hinge cysteines between them. Other chemical linkages of antibody fragments are also known.
[0025] As used herein, "adapter protein" refers to a protein sequence that specifically interacts with another adapter protein sequence in solution. In one embodiment, the "adapter protein" includes a heteromultimerization domain. Such adapter proteins include leucine zipper proteins, or the amino acid sequences of SEQ ID NO: 4 (cJUN(R): ASIARL[E]E[K]V KTL[K]A[Q]NYEL [A]S[T]ANMLRE[Q] VAQLGGC) or SEQ ID NO: 5 (FosW(E): AS[I]DEL[Q]AE[V] EQLEE[R]NYAL [R]KE[V]EDL[Q]K[Q] [A]EKLGGC) or variants thereof (non-limiting, amino acids of SEQ ID NO: 4 and SEQ ID NO: 5 that can be modified to include those underlined and in bold), where the variant has an amino acid modification that maintains or increases the affinity of the adapter protein for another adapter protein, a polypeptide comprising a variant, or a polypeptide comprising an amino acid sequence selected from the group consisting of SEQ ID NO: 11 (ASIARLRERVKTLRARNYELRSRANMLRERVAQLGGC) or SEQ ID NO: 12 (ASLDELEAEIEQLEEENYALEKEIEDLEKELEKLGGC), or a polypeptide comprising the amino acid sequence of SEQ ID NO: 13 (GABA-R1: EEKSRLLEKE NRELEKIIAE KEERVSELRH QLQSVGGC) or SEQ ID NO: 14 (GABA-R2: TSRLEGLQSE NHRLRMKITE LDKDLEEVTM QLQDVGGC) or SEQ ID NO: 15 (Cys: AGSC) or SEQ ID NO: 16 (Hinge: CPPCPG). The nucleic acid molecule encoding the coat protein or adapter protein is contained within a synthetic intron.
[0026] As used herein, the term "heteromultimerization domain" refers to a modification or addition to a biological molecule that promotes heteromultimer formation and impedes homomultimer formation. Any heterodimerization domain having a strong preference for forming heterodimers over homodimers is within the scope of the present invention. Exemplary examples include, but are not limited to, U.S. Patent Application No. 20030078385 (Arathoon et al., Genentech, describing knobs into holes), International Patent Publication No. WO2007147901 (Kjargaard et al., Novo Nordisk, describing ionic interactions), International Patent Publication No. WO2009089004 (Kannan et al., Amgen, describing electrostatic steering effects), International Patent Publication No. WO2011 / 034605 (Christensen et al., Genentech, describing coiled coils). Also, see, for example, Pack, P. & Plueckthun, A., Biochemistry 31, 1579-1584 (1992) describing leucine zippers, or Pack et al., Bio / Technology 11, 1271-1277 (1993) describing helix-turn-helix motifs. The terms "heteromultimerization domain" and "heterodimerization domain" are used interchangeably herein.
[0027] As used herein, the term "cloning site" refers to a nucleic acid sequence region that functions as a restriction enzyme site for restriction endonuclease-mediated cloning by ligation of nucleic acid sequences containing compatible cohesive or blunt ends, a priming site for PCR-mediated cloning of insert DNA by homology and extensions such as "overlap PCR stitching", or a recombination site for recombinase-mediated insertion of a target nucleic acid sequence by a recombination-exchange reaction, or a mosaic end for transposon-mediated insertion of a target nucleic acid sequence, and other techniques common in the art, including nucleic acid sequences.
[0028] As used herein, "coat protein" refers to any of five capsid proteins that are components of phage particles, including pIII, pVI, pVII, pVIII, and pIX. In one embodiment, "coat protein" can be used to display a protein or peptide (see Phage Display, A Practical Approach, Oxford University Press, edited by Clackson and Lowman, 2004, pp. 1-26). In one embodiment, the coat protein can be the pIII protein or some variant, portion, and / or derivative thereof. For example, the C-terminal portion of the M13 bacteriophage pIII coat protein (cP3), e.g., the sequence encoding the C-terminal residues 267-421 of protein III of the M13 phage, can be used. In one embodiment, the pIII sequence comprises the amino acid sequence of SEQ ID NO: 17 (AEDIEFASGGGSGAETVESCLAKPHTENSFTNVWKDDKTLDRYANYEGCLWNATGVVVCTGDETQCYGTWVPIGLAIPENEGGGSEGGGSEGGGSEGGGTKPPEYGDTPIPGYTYINPLDGTYPPGTEQNPANPNPSLEESQPLNTFMFQNNRFRNRQGALTVYTGTVTQGTDPVKTYYQYTPVSSKAMYDAYWNGKFRDCAFHSGFNEDPFVCEYQGQSSDLPQPPVNAGGGSGGGSGGGSEGGGSEGGGSEGGGSEGGGSGGGSGSGDFDYEKMANANKGAMTENADENALQSDAKGKLDSVATDYGAAIDGFIGDVSGLANGNGATGDFAGSNSQMAVGDGDNSPLMNNFRQYLPSLPQSVECRPFVFSAGKPYEFSIDCDKINLFRGVFAFLLYVATFMYVFSTFANILRNKES).In one embodiment, the pIII fragment comprises the amino acid sequence of SEQ ID NO: 18 (SGGGSGSGDFDYEKMANANKGAMTENADENALQSDAKGKLDSVATDYGAAIDGFIGDVSGLANGNGATGDFAGSNSQMAQVGDGDNSPLMNNFRQYLPSLPQSVECRPFVFGAGKPYEFSIDCDKINLFRGVFAFLLYVATFMYVFSTFANILRNKES).
[0029] As used herein, "expression cassette" means a nucleic acid fragment (e.g., a DNA fragment) that contains a specific nucleic acid sequence having a specific biological and / or biochemical activity. The expressions "cassette", "gene cassette", or "DNA cassette" can be used interchangeably and can have the same meaning.
[0030] The terms "host cell", "host cell line", and "host cell culture" are used interchangeably and refer to a cell (including progeny of such cell) into which an exogenous nucleic acid has been introduced. Host cells include "transformants" and "transformed cells", including primary transformed cells and progeny regardless of the number of passages. Progeny may not have exactly the same nucleic acid content as the parental cell and may contain mutations. Mutant progeny having the same function or biological activity as screened or selected in the original transformed cell are included herein.
[0031] As used herein, "linked" or "link(s)" or "linkage" means a covalent bond between two amino acid sequences or two nucleic acid sequences via a peptide or phosphodiester bond, and such bond may include any number of additional amino acids or nucleic acid sequences between the two amino acid sequences or nucleic acid sequences being joined.
[0032] "Nucleic acid" or "polynucleotide", when used interchangeably herein, refers to a polymer of nucleotides of any length, including DNA and RNA. Nucleotides can be deoxyribonucleotides, ribonucleotides, modified nucleotides or bases, and / or their analogs, or any substrate that can be incorporated into a polymer by DNA or RNA polymerase or by a synthetic reaction. Polynucleotides may include modified nucleotides, such as methylated nucleotides and their analogs. If there are modifications to the nucleotide structure, they may be made before or after the organization of the polymer. The nucleotide sequence may be interrupted by non-nucleotide components. Polynucleotides may be further modified after synthesis, such as by conjugation with a label. Other types of modifications include, for example, "capping" substitution of one or more of the natural nucleotides with analogs, internucleotide modifications, such as those with non-charged linkages (e.g., methylphosphonate, phosphotriester, phosphoramidate, carbamate, etc.) and those with charged linkages (e.g., phosphorothioate, phosphorodithioate, etc.), those containing pendant moieties, such as proteins (e.g., nuclease, toxin, antibody, signal peptide, poly-L-lysine, etc.), those with intercalators (e.g., acridine, psoralen, etc.), those containing chelating agents (e.g., metal, radioactive metal, boron, metal oxide, etc.), those containing alkylating agents, those with modified linkages (e.g., alpha-aromatic nucleic acid, etc.), and the unmodified forms of the polynucleotide(s). Furthermore, any of the hydroxyl groups originally present in the sugar may be substituted, for example, with a phosphonic acid group, a phosphate group, protected with a standard protecting group, or activated to prepare an additional linkage to an additional nucleotide, or conjugated to a solid or semi-solid support. The OH groups at the 5' and 3' ends may be phosphorylated or substituted with an amine or an organic capping group moiety of 1 to 20 carbon atoms. Also, other hydroxyls may be derivatized with standard protecting groups.In addition, the polynucleotide may contain similar forms of ribose or deoxyribose sugars commonly known in the art, such as 2'-O-methyl-, 2'-O-allyl, 2'-fluoro-, or 2'-azido-ribose; carbocyclic sugar analogs; alpha-anomer sugars; epimeric sugars such as arabinose, xylose, or lyxose; pyranose sugars; furanose sugars; sedoheptulose; acyclic analogs; and basic nucleoside analogs such as methyl riboside. One or more phosphodiester bonds may be replaced by alternative linking groups. These alternative linking groups include, but are not limited to, embodiments in which the phosphate is replaced by P(O)S (“thioate”), P(S)S (“dithioate”), “(O)NR2 (“amidate”), P(O)R, P(O)OR’, CO, or CH2 (“formacetal”), where each R or R’ is independently H, or a substituted or unsubstituted alkyl (1-20C) optionally containing an ether (-O-) bond, aryl, alkenyl, cycloalkyl, cycloalkenyl, or araldyl. Not all linkages within the polynucleotide need to be the same. The above description applies to all polynucleotides mentioned herein, including RNA and DNA.
[0033] A nucleic acid is "operably linked" when it is placed into a structural or functional relationship with another nucleic acid sequence. For example, one segment of DNA and another segment of DNA are positioned on the same continuous DNA molecule relative to each other and relative to a coding sequence so as to promote transcription of the coding sequence, a promoter or enhancer; a ribosome binding site positioned relative to a coding sequence so as to promote translation; or a presequence or secretory leader positioned relative to a coding sequence so as to promote expression of a preprotein (e.g., a preprotein involved in secretion of the encoded polypeptide), one segment of DNA can be operably linked to another segment of DNA when they have a structural or functional relationship such as this. In other instances, operably linked nucleic acid sequences are not contiguous, but are positioned relative to each other as if they were contiguous or as if they occur in the same nucleic acid or protein which they express. For example, an enhancer can be non-contiguous. Ligation can be achieved by ligation at convenient restriction enzyme sites, or by use of synthetic oligonucleotide adapters or linkers.
[0034] The terms "polyadenylation signal" or "polyadenylation site" as used herein refer to a sequence sufficient to direct the addition of polyadenosine ribonucleotides to an RNA molecule expressed in a cell.
[0035] A "promoter" is a nucleic acid sequence that enables initiation of transcription of a gene sequence in a messenger RNA, such transcription being initiated by binding of RNA polymerase to or near the promoter.
[0036] The term "3' splice site" is intended to mean a nucleic acid sequence, e.g., a pre-mRNA sequence, of a 3' intron / exon boundary that can be recognized and bound by a splicing mechanism.
[0037] The term "5' splice site" is intended to mean a nucleic acid sequence, such as a pre-mRNA sequence, of a 5' exon / intron boundary that can be recognized and bound by a splicing mechanism.
[0038] The term "cryptic splice site" is intended to mean a normally silent 5' or 3' splice site that can be activated by mutation or otherwise and can function as a splicing element. For example, a mutation can activate a 5' splice site downstream of a native or dominant 5' splice site. Use of this "cryptic" splice site results in the production of distinct mRNA splicing products not produced by use of the native or dominant splice site.
[0039] As used herein, the term "trans-splicing" means the joining of exons contained on separate, non-contiguous RNA molecules.
[0040] The term "variable region" or "variable domain" refers to a domain of an antibody heavy or light chain that is involved in binding of the antibody to an antigen. The variable domains of the heavy and light chains of a native antibody (VH and VL, respectively) generally have a similar structure, and each domain contains four conserved framework regions (FRs) and three hypervariable regions (HVRs). (See, e.g., Kindt et al. Kuby Immunology, 6th ed., W.H. Freeman and Co., page 91 (2007).) A single VH or VL domain may be sufficient to confer antigen-binding specificity. Furthermore, antibodies that bind a particular antigen can be isolated from antibodies that bind the antigen for screening libraries of complementary VL or VH domains, respectively, using the VH or VL domain. See, e.g., Portolano et al., J. Immunol. 150:880-887 (1993); Clarkson et al., Nature 352:624-628 (1991).
[0041] As used herein, the term "vector" refers to a nucleic acid molecule capable of amplifying another nucleic acid to which it is ligated. This term includes vectors as self-replicating nucleic acid structures and vectors integrated into the genome of a host cell into which the vector has been introduced. Certain vectors can direct the expression of nucleic acids to which they are operably linked. Such vectors are referred to herein as "expression vectors."
[0042] II. Modular Polypeptide Expression System The present invention is based, at least in part, on the discovery that pre-mRNA trans-splicing can be utilized in mammalian cells to enable modular recombinant protein expression. The concept of modular, flexible protein expression allows for the precise joining of any two protein-coding sequences encoded by two different constructs into a single mRNA encoding a polypeptide chain without any of the requirements and constraints of other protein-protein splicing methods. This concept can be adapted to simplify and expand other technologies that require the expression in mammalian cells of many collections of proteins having various combinations of repeating modules.
[0043] Here, the generation of a number of polypeptide expression systems that enable the modular expression of various antibody formats in the context of a phage display expression system is described. The necessary nucleic acid components, vectors, host cells, and methods of using the polypeptide expression systems of the present invention are described herein.
[0044] A. Modes of Carrying Out the Invention The practice of the present invention, unless otherwise indicated, uses conventional techniques within the skill of the art in molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry, and immunology. Such techniques are fully explained in the literature such as "Molecular Cloning: A Laboratory Manual", 2nd Edition (Sambrook et al., 1989), "Oligonucleotide Synthesis" (M.J. Gait, ed., 1984), "Animal Cell Culture" (R.I. Freshney, ed., 1987), "Methods in Enzymology" (Academic Press, Inc.), "Handbook of Experimental Immunology", 4th Edition (D.M. Weir & C.C. Blackwell, eds., Blackwell Science Inc., 1987), "Gene Transfer Vectors for Mammalian Cells" (J.M. Miller & M.P. Calos, eds., 1987), "Current Protocols in Molecular Biology" (F.M. Ausubel et al., eds., 1987), "PCR: The Polymerase Chain Reaction" (Mullis et al., eds., 1994), and "Current Protocols in Immunology" (J.E. Coligan et al., eds., 1991).
[0045] B. Module protein expression system The polypeptide expression system of the present invention can support the expression of polypeptides (such as fusion proteins) in the same or different (e.g., reformatted) forms. The present invention provides means for generating such a polypeptide expression system for the modular expression and production of various forms (e.g., various formats or various fusion forms) of a protein of interest in a host cell-dependent manner by using the process of trans-splicing.
[0046] 1. Nucleic acid components of the modular protein expression system a. Structure of the nucleic acid components of the modular protein expression system The protein expression system uses at least two nucleic acid molecules that together enable flexible modular expression of any desired polypeptide via the process of directed pre-mRNA trans-splicing. The first nucleic acid molecule contains a first expression cassette that includes a eukaryotic promoter (P1 Euk1 )(e.g., cytomegalovirus (CMV) promoter, simian virus 40 (SV40) promoter, Moloney murine leukemia virus U3 region, caprine arthritis encephalitis virus U3 region, visna virus U3 region, or retrovirus U3 region sequence), which is operably linked to a polypeptide coding sequence (PES11). In some cases, the polypeptide coding sequence encodes only a portion of the desired polypeptide, and the remainder is supplied by a polypeptide coding sequence (PES2) contained on a second nucleic acid molecule. The first nucleic acid molecule may include a 5'ss (5'ss11) (e.g., GTAAGA (SEQ ID NO: 8)) that is downstream (3') of PES11 but upstream (5') of the hybridization sequence (HS1).
[0047] The HS1 array may contain a gene encoding all or part of a polypeptide tag, label, coat protein, and / or adapter protein that can be positioned in-frame with PES11, such that its expression results in a protein encoded by PES11 fused to a protein encoded by HS1. In some cases, HS1 is a gene encoding all or part of a coat protein selected from the group consisting of pI, pII, pIII, pIV, pV, pVI, pVII, pVIII, pIX, and pX of bacteriophage M13, f1, or fd. For example, PES11 can encode all or part of an antibody or its Fab fragment, and the HS1 sequence can encode a coat protein (e.g., all or part of the pIII protein of bacteriophage M13, e.g., the pIII fragment contains amino acid residues 267-421 of the pIII protein or amino acid residues 262-418 of the pIII protein), resulting in an antibody- or Fab fragment-pIII protein fusion product. In another case, HS1 is a gene encoding all or part of an adapter protein such as a leucine zipper, and the leucine zipper contains the amino acid sequence of SEQ ID NO: 4 or 5.
[0048] In addition, the first nucleic acid molecule can encode a eukaryotic signal sequence (ESS11) located 3' relative to P1 Euk1 and 5' relative to PES11. Thus, the first nucleic acid molecule may contain the above components linked to each other (e.g., operably linked) in the 5' to 3' direction as P1 Euk1 -ESS11-PES11-5'ss11-HS1.
[0049] The second nucleic acid molecule of the protein expression system is a eukaryotic promoter (P2 operably linked to a polypeptide coding sequence (PES2 Euk)(e.g., a cytomegalovirus (CMV) promoter or a simian virus 40 (SV40) promoter). In some cases, the polypeptide coding sequence encodes only a part of the desired polypeptide, and the remaining part is supplied by a polypeptide coding sequence (PES11) contained on the first nucleic acid molecule. The second nucleic acid molecule may include a 3' splice site (3'ss2) located 5' to PES2. The second nucleic acid molecule Euk may include a hybridizing sequence (HS2) that can hybridize to HS1 located between P2 Euk and 3'ss2. Further, the second nucleic acid molecule may include a polyadenylation site (pA2), where the components of the second nucleic acid molecule are operably linked to each other in the 5' to 3' direction as P2
[0050] Therefore, trans-splicing between the first and second nucleic acid pre-mRNA products in eukaryotic cells (e.g., mammalian cells) is induced by hybridization of complementary sequences (i.e., HS1 and HS2) located on separate mRNA molecules, such that the isolated 5' splice site (5'ss11) of the first molecule and the isolated 3' splice site (3'ss2) of the second molecule come into proximity, enabling trans-splicing to occur and supporting the formation of the desired trans-spliced mRNA transcript. Additionally, to facilitate trans-splicing, the first nucleic acid molecule may contain an intron splice enhancer (ISE) (ISE1) positioned between 5'ss11 and HS1. ISE1 contains, for example, a G-run having three or more consecutive guanine residues, such as a G-run having nine consecutive guanine residues. Furthermore, trans-splicing between the first and second nucleic acid pre-mRNA products can be induced during their transcription in eukaryotic cells (e.g., mammalian cells, such as Expi293F, 293T, or CHO cells) by genetically engineering the first nucleic acid molecule to lack a standard polyadenylation site downstream of the PES11 and / or HS1 components. This will minimize the formation of mature mRNA transcripts that can be transported to the cytoplasm before trans-splicing with the mRNA transcript of the second nucleic acid molecule can occur.
[0051] In some cases, it may be desirable to co-express separate polypeptide products. For example, it may be desirable to express a first polypeptide product encoded by both the first and second nucleic acid molecules and a second polypeptide product that can self-organize with the first to form a desired heteromultimeric protein product (e.g., an antibody consisting of both a heavy chain and a light chain). For this purpose, the first and / or second nucleic acid molecule may further comprise a second expression cassette. For example, in the case where the first nucleic acid molecule contains the second expression cassette, the second expression cassette comprises a second eukaryotic promoter (P1 Euk2)、(ii) a second nucleic acid sequence encoding a eukaryotic signal sequence (ESS12), (iii) a second polypeptide coding sequence (PES12), and (iv) may include a polyadenylation site (pA1), and these components are P1 Euk2 -ESS12-PES12-pA1 are operably linked to each other in the 5' to 3' direction. In some cases, the second expression cassette may not contain the ESS12 component (for example, when secretion of the expressed polypeptide is not required or not desired). Thus, the first nucleic acid molecule encodes two polypeptide products under separate promoters, whereby one of the mRNA transcripts encoding one of the polypeptide products of the first nucleic acid molecule is formed via directed trans-splicing with the mRNA transcript encoded by the second nucleic acid molecule. In some cases, the second expression cassette is positioned 5' relative to the first expression cassette. In other cases, the second expression cassette is positioned 3' relative to the first expression cassette.
[0052] b. Polypeptide expression in both prokaryotic and eukaryotic cells In some cases, the polypeptide expression system can be genetically engineered for polypeptide expression in both prokaryotic and eukaryotic cells. Thus, the first nucleic acid molecule may include a removable prokaryotic promoter module (ePPM1) positioned between P1 Euk1 and PES11 when expression of the polypeptide product encoded by PES11 or in some cases PES11 and HS1 is desired. ePPM1 may include a 5' splice site (5'ss12), a prokaryotic promoter (P1 Prok1 ), a nucleic acid sequence encoding a prokaryotic signal sequence (PSS11), and a 3' splice site (3'ss11), and these are 5'ss12-P1 Prok1-PSS11-ss11 are positioned relative to each other in the 5' to 3' direction and are operably linked to drive the transcription of the polypeptide encoded by PES11, or PES11 and HS1. In some cases, ePPM1 may not contain the PSS11 component (e.g., when secretion of the expressed polypeptide is not required or not desired). Thus, ePPM1 can drive the transcription of the polypeptide encoded by PES11 of the first nucleic acid molecule in prokaryotic cells. On the other hand, in eukaryotic cells (e.g., mammalian cells), P1 Euk1 drives the expression of the transcription of the polypeptide encoded by PES11 of the first nucleic acid molecule, and ePPM1 can be removed from the pre-mRNA transcript by cis-splicing due to the collision adjacent to the 5'ss12 and 3'ss11 components.
[0053] In some cases, ePPM1 also contains a polypyrimidine tract (PPT11) positioned between PSS11 and 3'ss11. PPT11 may contain, for example, the sequence TTCCTTTTTTCTCTTTCC (SEQ ID NO: 1). Also, the second nucleic acid molecule may contain a polypyrimidine tract (PPT2) that may be positioned, for example, between HS2 and 3'ss2. PPT2 may contain, for example, the sequence TTCCTCTTTCCCTTTCTCTCC (SEQ ID NO: 7). In addition, the second nucleic acid molecule may further contain an ISE (ISE2) positioned between HS2 and 3'ss2. ISE2 may contain, for example, a G-run having three or more consecutive guanine residues, such as a G-run having nine consecutive guanine residues.
[0054] In some embodiments where the first nucleic acid molecule of the polypeptide expression system contains a second expression cassette, the second expression cassette may further contain an excisable prokaryotic promoter module (ePPM2) positioned between P1 Euk2 and PES12, and the following components: (i) a 5' splice site (5'ss13), (ii) a prokaryotic promoter (P1 Prok2), (iii) a nucleic acid sequence encoding a prokaryotic signal sequence (PSS12), and (iv) a 3' splice site (3'ss12), whereby these components are positioned relative to each other in the 5' to 3' direction as 5'ss13-P1 Prok2 -PSS12-3'ss12 and are operably linked to drive the transcription of the polypeptide encoded by PES12. In some cases, ePPM2 may not include the PSS12 component (e.g., if secretion of the expressed polypeptide is not required or not desired). The second excisable prokaryotic promoter module will function in a manner similar to that of the first excisable prokaryotic promoter module described above.
[0055] The prokaryotic promoter(s) of the excisable prokaryotic promoter module(s) may be the phoA, Tac, Lac, or Tphac promoter (see, e.g., Kim et al. PLoS One. 7(4):e35844), or another prokaryotic promoter known in the art.
[0056] In the construction of vectors capable of expressing a protein of interest in both prokaryotic cells (e.g., E. coli cells) and eukaryotic cells (mammalian cells, e.g., Expi293F cells), additional challenges arise from the differences in signal sequences found in these cell types. Certain features of signal sequences are generally conserved in both prokaryotic and eukaryotic cells (e.g., hydrophobic residues located in the middle of the sequence, and a patch of polar / charged residues adjacent to the cleavage site at the N-terminus of the mature polypeptide), while others are more characteristic of one cell type than the other. Furthermore, it is known in the art that even if the sequences are all of mammalian origin, different signal sequences can have a significant impact on expression levels in mammalian cells (Hall et al., J of Biological Chemistry, 265:19996-19999 (1990), Humphreys et al., Protein Expression and Purification, 20:252-264 (2000)). For example, bacterial signal sequences typically have a positively charged residue (most commonly lysine) immediately following the initiating methionine, while these are not always present in mammalian signal sequences.
[0057] When secretion of the expressed protein is necessary or desired, any signal sequence (including consensus signal sequences) targeting the polypeptide of interest can be used for the periplasm in prokaryotes and the endoplasmic reticulum in eukaryotes. For example, eukaryotic signal sequences (e.g., ESS11 or ESS12) can be derived from all or part of the mouse binding immunoglobulin protein (mBiP) signal sequence (UniProtKB: accession number P20029) or the antibody heavy or light chain signal sequence (e.g., mouse VH gene signal sequence), or can include all or part thereof. In some embodiments, prokaryotic signal sequences (e.g., PSS11 or PSS12) can be derived from all or part of the heat-stable enterotoxin II (stII) gene, or can include all or part thereof. Other signal sequences that can be utilized include signal sequences from human growth hormone (hGH) (UniProtKB: accession number BIA4G6), Gaussia princeps luciferase (UniProtKB: accession number Q9BLZ2), and yeast endo-1,3-glucanase (yBGL2) (UniProtKB: accession number P15703). The signal sequence can be a natural or synthetic signal sequence. In some embodiments, the synthetic signal sequence is an optimized signal secretion sequence that drives an optimized level of display relative to its non-optimized natural signal sequence.
[0058] 2. Vectors, Host Cells, and Production Methods The present invention features a vector or vector set comprising one or more of the nucleic acid molecules described above. Accordingly, the present invention also features a vector set comprising a first vector and a second vector, wherein the first and second vectors each comprise the first and second nucleic acid molecules of the polypeptide expression system described above.
[0059] In addition to the components of the nucleic acid molecule described in detail above, the vector or vector set may include a bacterial origin of replication, a mammalian origin of replication, and / or a nucleic acid encoding a polypeptide useful as a control (e.g., the gD protein) or for an activity useful (e.g., protein purification, protein tagging, or protein labeling).
[0060] Also provided is a method for producing a polypeptide, comprising culturing a host cell comprising one or more of the above vector(s) or vector set(s) in a culture medium and optionally recovering an antibody from the host cell (or the host cell culture medium).
[0061] C. Phage Display Vector Systems for Modular Antibody Expression and Reformatting In some embodiments, antibodies (e.g., full-length antibodies, e.g., full-length IgG antibodies, or fragments thereof, e.g., Fab fragments) can be produced using the polypeptide expression system of the invention. Application of a modular protein expression system is shown by designing a phage display vector system that allows expression of different antibody formats in human cells from the same clone. A portion of the heavy-chain antigen-binding region and the constant region encoded by the phage display vector are directly and precisely fused to sequences encoded in a second complementary construct by binding sequences encoding different portions of the polypeptide by pre-mRNA trans-splicing during expression in the cell.
[0062] The use of a polypeptide expression system aimed at enabling the direct expression of IgG in mammalian cells without the need for subcloning of the phage Fab array is described in Examples 1 and 2 below. In some cases, the first nucleic acid molecule of the polypeptide expression system can be designed to encode the entire Fab fragment component. Thus, the first nucleic acid molecule can include a PES11 component that encodes a polypeptide having the VH and CH1 domains of the Fab. The first nucleic acid molecule can also include a PES12 component that encodes the VL and CL domains. Transcription of the first nucleic acid molecule results in two discontinuous pre-mRNA products, which together form a Fab fragment that can be appropriately tagged (e.g., fused to pIII of M13) for phage display purposes.
[0063] The process of reformatting the Fab fragment into a full-length IgG antibody can then be achieved by the expression of the first nucleic acid molecule in eukaryotic cells (e.g., mammalian cells, such as Expi293F cells), along with a second nucleic acid molecule that provides the remaining portion of the antibody (i.e., the CH2 and CH3 domains). For example, the second nucleic acid molecule can include a PES2 component that encodes a polypeptide having the CH2 and CH3 domains. Transcription of the first and second molecules in eukaryotic cells results in the production of three pre-mRNA transcripts, and the heavy chains encoding the pre-mRNA transcripts are induced to undergo trans-splicing with each other to produce the reformatted full-length heavy chain of the desired IgG antibody. The processed mRNA is then translated, resulting in the production of both the light and heavy chains of the IgG molecule, and such production would not require the need for labor-intensive subcloning.
[0064] When different antibody formats, such as wild-type IgG, Fab fragments, or IgG with Fc modifications for bispecific formats, are needed for different screening assays, the ability to express different antibody formats from the same clone is useful in antibody discovery. The expression system of the polypeptides of the present invention enables any of these or additional formats, in principle, simply by cloning a suitable sequence added after the CH1 region in a complementary plasmid. Furthermore, the modular organization of the system only requires the construction of a new complementary plasmid, thus eliminating the need to recreate the stock of the phage display library and enabling the expression of new antibody formats. The nucleic acid can also be adapted to allow the use of any CH1 region by transferring the 5′ss from downstream of the CH1 coding region to the J region (FR4) of the VH or J-CH1 junction, thus separating the VH and the entire constant region of the heavy chain in two different nucleic acids. The nucleic acid molecule is compatible with conventional methods for the expression of Fab fragments in E. coli by simply adding a stop codon after the sequence encoding the upper hinge. However, the amber stop codon at the junction of the heavy chain and the gene III sequence in the Fab phage display library usually results in a significantly lower level of display, thus requiring recloning of the clones after selection, at least in the case of an inexperienced repertoire library (Lee et al. Journal of immunological methods. 284:119-132, 2004). The expression of Fab fragments in mammalian cells using the same method as used for IgG expression avoids this need for recloning with a yield comparable to that normally obtained in E. coli.
[0065] Antibodies produced by this polypeptide expression system can include recombinantly produced chimeric, humanized, and / or human antibodies. In some cases, the antibody is an antibody fragment, such as Fab, Fv, Fab′, scFv, bispecific antibody, or F(ab′)2 fragment. In other cases, the antibody is a full-length antibody, such as an intact IgG1, IgG2, IgG3, or IgG4 antibody as defined herein, or another antibody of another class or isotype.
[0066] The expressed antibodies can incorporate any of the features, alone or in combination, as described in items 1-7 below.
[0067] 1. Antibody affinity Antibodies (e.g., Fab or full-length IgG antibodies) produced by the polypeptide expression system described herein can have a dissociation constant (Kd) of ≤1 μM, ≤100 nM, ≤10 nM, ≤1 nM, ≤0.1 nM, ≤0.01 nM, or ≤0.001 nM (e.g., 10 -8 M or less, e.g., 10 -8 M to 10 -13 M, e.g., 10 -9 M to 10 -13 M).
[0068] In one embodiment, the Kd is measured by a radiolabeled antigen binding assay (RIA) performed on the Fab version of the antibody of interest and its antigen, as described by the following assay method. The solution binding affinity of the Fab for the antigen is determined in the presence of a titration series of unlabeled antigen, ( 125I) It is measured by equilibrating the Fab with the minimum concentration of the labeled antigen and then capturing the antigen bound to the anti-Fab antibody-coated plate (see, for example, Chen et al., J. Mol. Biol. 293:865-881 (1999)). To establish the conditions for the assay, a MICROTITER® multi-well plate (Thermo Scientific) is coated overnight with 5 μg / ml of the capture anti-Fab antibody (Cappel Labs) in 50 mM sodium carbonate (pH 9.6) and then blocked with 2% (w / v) bovine serum albumin in PBS for 2-5 hours at room temperature (about 23°C). In non-adsorptive plates (Nunc #269620), 100 pM or 26 pM 125 I] The antigen is mixed with a serial dilution of the Fab of interest (for example, without conflicting with the evaluation of the anti-VEGF antibody, Fab-12 in Presta et al., Cancer Res. 57:4593-4599 (1997)). The Fab of interest is then cultured overnight, but the culture can continue for a longer time (for example, about 65 hours) to ensure that equilibrium has been reached. Thereafter, the mixture is transferred to the capture plate for incubation at room temperature (for example, 1 hour). The solution is then removed and the plate is washed 8 times with 0.1% polysorbate 20 (TWEEN-20®) in PBS. When the plate is dry, 150 μl / well of scintillant (MICROSCINT-20™; Packard) is added and the plate is counted in a TOPCOUNT® gamma counter (Packard) for 10 minutes. The concentration of each Fab that provides less than 20% of the maximum binding is selected for use in the competitive binding assay.
[0069] According to another embodiment, Kd is measured at 25° C. using a surface plasmon resonance assay with an immobilized antigen CM5 chip at approximately 10 response units (RU) using a BIACORE®-2000 or BIACORE®-3000 (BIAcore, Inc., Piscataway, N.J.). Briefly, a carboxymethylated dextran biosensor chip (CM5, BIAcore Inc.) is activated with N-ethyl-N'-(3-dimethylaminopropyl)-carbodiimide hydrochloride (EDC) and N-hydroxysuccinimide (NHS) according to the supplier's instructions. The antigen is diluted to 5 μg / ml (approximately 0.2 μM) with 10 mM sodium acetate at pH 4.8 and injected at a flow rate of 5 μl / min to achieve approximately 10 response units (RU) of bound protein. After injection of the antigen, 1 M ethanolamine is injected to block unreacted groups. For kinetic measurements, two-fold serially diluted Fab (0.78 nM to 500 nM) is injected into PBS with 0.05% polysorbate 20 (TWEEN-20™) surfactant (PBST) at a flow rate of approximately 25 μl / min at 25° C. The association rate (k on ) and the dissociation rate (k off ) are calculated using a simple one-to-one Langmuir binding model (BIACORE® evaluation software version 3.2) by simultaneously fitting the association and dissociation sensorgrams. The equilibrium dissociation constant (Kd) is calculated as the ratio of k off / k on . See, for example, Chen et al., J. Mol. Biol. 293:865-881 (1999). The binding rate by the above surface plasmon resonance assay is 10 6 M -1 s -1When it exceeds, the binding rate is measured by a spectrofluorometer, for example, a spectrophotometer equipped with a stop flow (Aviv Instruments) or an 8000 series SLM-AMINCO (trademark) spectrophotometer (ThermoSpectronic) equipped with a stirred cuvette, in the presence of increasing concentrations of antigen, in PBS at pH 7.2, at 25 °C, and can be determined by using a fluorescence quenching technique that measures an increase or decrease in the fluorescence emission intensity (excitation = 295 nM, emission = 340 nM, 16 nM bandpass) of 20 nM anti-antigen antibody (Fab type).
[0070] 2. Antibody fragments In certain embodiments, the antibodies produced by the polypeptide expression systems described herein are antibody fragments. Antibody fragments include, but are not limited to, Fab, Fab′, Fab′-SH, F(ab′)2, Fv, and scFv fragments, as well as other fragments described below. For a review of certain antibody fragments, see Hudson et al Nat. Med. 9:129-134 (2003). For a review of scFv fragments, see, for example, Pluckthun, in The Pharmacology of Monoclonal Antibodies, vol. 113, Rosenberg and Moore eds., (Springer-Verlag, New York), pp. 269-315 (1994), International Patent Publication No. WO93 / 16185, as well as U.S. Patent Nos. 5,571,894 and 5,587,458. For a discussion of Fab and F(ab′)2 fragments that contain salvage receptor binding epitope residues and increase the in vivo half-life, see, for example, U.S. Patent No. 5,869,046.
[0071] Bispecific antibodies are antibody fragments that have two antigen-binding sites and can be bivalent or bispecific. See, for example, European Patent No. 404,097, International Patent Publication No. WO1993 / 01161, Hudson et al., Nat. Med. 9:129-134 (2003), and Hollinger et al., Proc. Natl. Acad. Sci. USA 90:6444-6448 (1993). Trispecific and tetravalent antibodies are also described in Hudson et al., Nat. Med. 9:129-134 (2003).
[0072] Single-domain antibodies are antibody fragments that include all or part of the heavy-chain variable domain of an antibody, or all or part of the light-chain variable domain. In certain embodiments, the single-domain antibody is a human single-domain antibody (Domantis, Inc., Waltham, MA; see, for example, U.S. Patent No. 6,248,516 B1).
[0073] 3. Chimeric and Humanized Antibodies In certain embodiments, antibodies produced by the polypeptide expression systems described herein (e.g., Fab or full-length IgG antibodies) are chimeric antibodies. Certain chimeric antibodies are described, for example, in U.S. Patent No. 4,816,567, and Morrison et al., Proc. Natl. Acad. Sci. USA, 81:6851-6855 (1984). In one example, a chimeric antibody includes a non-human variable region (e.g., a variable region from a mouse, rat, hamster, rabbit, or non-human primate such as a monkey) and a human constant region. In a further example, a chimeric antibody is a "class-switched" antibody in which the class or subclass has been changed from that of the parent antibody. Chimeric antibodies include their antigen-binding fragments.
[0074] In certain embodiments, the chimeric antibody is a humanized antibody. Typically, non-human antibodies are humanized to reduce immunogenicity to humans while maintaining the specificity and affinity of the parental non-human antibody. Generally, a humanized antibody comprises one or more variable domains in which the HVRs, such as CDRs (or portions thereof), are derived from a non-human antibody and the FRs (or portions thereof) are derived from human antibody sequences. A humanized antibody will optionally also include at least a portion of a human constant region. In some embodiments, some FR residues in the humanized antibody are replaced with the corresponding residues from a non-human antibody (e.g., the antibody from which the HVR residues are derived) to, for example, restore or improve the specificity or affinity of the antibody.
[0075] Humanized antibodies and methods of making them are reviewed, for example, in Almagro and Fransson, Biosci. 13:1619-1633 (2008), and further described in, for example, Riechmann et al., Nature 332:323-329 (1988), Queen et al., Proc. Nat′l Acad. Sci. USA 86:10029-10033 (1989), U.S. Pat. Nos. 5,821,337, 7,527,791, 6,982,321, and 7,087,409, Kashmiri et al., Methods 36:25-34 (2005) (describing SDR (α-CDR) grafting), Padlan, Mol. Immunol. 28:489-498 (1991) (describing "resurfacing"), Dall′Acqua et al., Methods 36:43-60 (2005) (describing "FR shuffling"), and Osbourn et al., Methods 36:61-68 (2005) and Klimka et al., Br. J. Cancer, 83:252-260 (2000) (describing "guided selection" approaches to FR shuffling).
[0076] Human framework regions that can be used for humanization are framework regions selected using the "best fit" method (see, e.g., Sims et al., J. Immunol. 151:2296 (1993)), framework regions derived from consensus sequences of human antibodies of particular subgroups of light or heavy chain variable regions (see, e.g., Carter et al., Proc. Natl. Acad. Sci. USA, 89:4285 (1992), and Presta et al., J. Immunol., 151:2623 (1993)), human mature (somatic mutated) framework regions or human germline framework regions (see, e.g., Almagro and Fransson, Front. Biosci. 13:1619-1633 (2008)), and framework regions derived from screening FR libraries (see, e.g., Baca et al., J. Biol. Chem. 272:10678-10684 (1997) and Rosok et al., J. Biol. Chem. 271:22611-22618 (1996)), but are not limited thereto.
[0077] 4. Human Antibodies In certain embodiments, antibodies produced by the polypeptide expression systems described herein (e.g., Fab or full-length IgG antibodies) are human antibodies. Human antibodies can be recombinant human antibodies that are prepared independently using various techniques known in the art and then their sequences are identified. Human antibodies are generally described in van Dijk and van de Winkel, Curr. Opin. Pharmacol. 5:368-74 (2001) and Lonberg, Curr. Opin. Immunol. 20:450-459 (2008).
[0078] 5. Library-Derived Antibodies By utilizing the polypeptide expression system described herein, which is useful in phage display systems, antibodies (e.g., Fab or full-length IgG antibodies) generated by the polypeptide expression system of the invention can be isolated by screening a combinatorial library for antibodies having the desired activity(ies). See, for example, Hoogenboom et al. in Methods in Molecular Biology 178:1-37 (O’Brien et al., ed., Human Press, Totowa, NJ, 2001), and for example in McCafferty et al., Nature 348:552-554; Clackson et al., Nature 352:624-628 (1991); Marks et al., J. Mol. Biol. 222:581-597 (1992), Marks and Bradbury, in Methods in Molecular Biology 248:161-175 (Lo, ed., Human Press, Totowa, NJ, 2003), Sidhu et al., J. Mol. Biol. 338(2):299-310 (2004), Lee et al., J. Mol. Biol. 340(5):1073-1093 (2004), Fellouse, Proc. Natl. Acad. Sci. USA 101(34):12467-12472 (2004), and Lee et al., J. Immunol. Methods 284(1-2):119-132 (2004).
[0079] 6. Multispecific Antibodies In certain embodiments, antibodies (e.g., Fab or full-length IgG antibodies) generated by the polypeptide expression systems described herein are multispecific antibodies, such as bispecific antibodies. Multispecific antibodies are monoclonal antibodies that have binding specificities for at least two different sites. In certain embodiments, one of the binding specificities is with respect to a first antigen and the other is with respect to any other antigen. In certain embodiments, the bispecific antibody can bind to two different epitopes of the first antigen. Bispecific antibodies can also be used to localize a cytotoxic agent to cells expressing the first antigen. Bispecific antibodies can be prepared as full-length antibodies or antibody fragments.
[0080] Engineered antibodies having three or more functional antigen-binding sites, including "octopus antibodies", are also included herein (see, e.g., US 2006 / 0025576A1).
[0081] The antibodies or fragments herein also include "dual action Fab" or "DAF" that contain an antigen-binding site that binds to a first antigen as well as a different antigen (see, e.g., US 2008 / 0069820).
[0082] 7. Antibody Variants In certain embodiments, amino acid sequence variants of the antibodies provided herein are contemplated. For example, it may be desirable to improve the binding affinity and / or other biological properties of the antibody. Amino acid sequence variants of the antibody can be prepared by introducing appropriate modifications into one or more of the nucleic acid molecule sequences encoding all or part of the antibody. Such modifications include, for example, deletions of residues from the amino acid sequence of the antibody, and / or insertions of residues into the amino acid sequence, and / or substitutions of residues in the amino acid sequence. Any combination of deletions, insertions, and substitutions can be made to achieve the final construct, provided that the final construct retains the desired characteristics, such as antigen-binding ability.
[0083] In certain embodiments, collections of antibody variants having one or more amino acid substitutions relative to one another can be generated by the expression systems and methods of the invention. Target sites for substitution mutagenesis include HVRs and FRs. Conservative substitutions are shown under the heading "Conservative Substitutions" in Table 1. More substantial changes are provided under the heading "Typical Substitutions" in Table 1 and are further described below in terms of amino acid side-chain classes. Amino acid substitutions can be introduced into the products screened with respect to the antibody of interest and the desired activity, e.g., maintained / improved antigen binding, reduced immunogenicity, or improved ADCC or CDC. TIFF2025097990000002.tif208170
[0084] Amino acids can be classified according to common side-chain properties: (1) Hydrophobic: norleucine, Met, Ala, Val, Leu, Ile; (2) Neutral hydrophilic: Cys, Ser, Thr, Asn, Gln; (3) Acidic: Asp, Glu; (4) Basic: His, Lys, Arg; (5) Residues that influence chain orientation: Gly, Pro; (6) Aromatic: Trp, Tyr, Phe
[0085] Non-conservative substitutions will involve exchanging one member of one of these classes for another.
[0086] One type of substitution variant involves substituting one or more hypervariable region residues of a parental antibody (e.g., a humanized or human antibody). Generally, the resulting variant(s) selected for further study will have a modification (e.g., improvement) in a particular biological property relative to the parental antibody (e.g., increased affinity, reduced immunogenicity), and / or will substantially maintain a particular biological property of the parental antibody. A typical substitution variant is an affinity matured antibody, which can be readily generated using, for example, affinity maturation methods based on phage display such as those described herein. Briefly, one or more HVR residues are mutated, the variant antibodies are displayed on phage, and screened for a particular biological activity (e.g., binding affinity).
[0087] Alterations (e.g., substitutions) can be made in the HVRs, for example, to improve antibody affinity. Such alterations can be made in HVR “hot spots,” i.e., residues encoded by codons that mutate frequently during the somatic maturation process (see, e.g., Chowdhury, P.S., Methods Mol. Biol. 207:179-196 (2008)), and / or in the SDR (a-CDR), and the resulting mutant VH or VL is tested for binding affinity. Affinity maturation methods by constructing and reselecting from secondary libraries are described, for example, in Hoogenboom, H.R. et al. in Methods in Molecular Biology 178 1-37 (2001) (O’Brien et al., eds., Human Press, Totowa, NJ). In some embodiments of the affinity maturation method, diversity is introduced into the variable region gene selected for maturation by any of a variety of methods (e.g., error-prone PCR, chain shuffling, or oligonucleotide-directed mutagenesis). A secondary library is then created. The library is then screened to identify any antibody variants having the desired affinity. Another method for introducing diversity involves an HVR-directed approach, in which several HVR residues (e.g., 4-6 residues at a time) are randomized. The HVR residues involved in antigen binding can be specifically identified, for example, using alanine scanning mutagenesis or modeling. In particular, CDR-H3 and CDR-L3 are often targeted.
[0088] In certain embodiments, substitutions, insertions, or deletions can occur within one or more HVRs, so long as such alterations do not substantially reduce the ability of the antibody to bind the antigen. For example, conservative alterations (e.g., conservative substitutions provided herein) that do not substantially reduce binding affinity can be made in the HVRs. Such alterations can be outside of HVR “hot spots” or the SDR. In certain embodiments of the mutant VH and VL sequences provided above, each HVR is either unaltered or contains no more than 1, 2, or 3 amino acid substitutions.
[0089] A method useful for identifying residues or regions of an antibody that can be targets for mutagenesis is called "alanine scanning mutagenesis" as described by Cunningham and Wells (1989) Science, 244:1081-1085. In this method, residues or groups of target residues (e.g., charged residues such as arg, asp, his, lys, and glu) are identified and replaced with neutral or negatively charged amino acids (e.g., alanine or polyalanine) to determine whether the interaction between the antibody and the antigen is affected. Further substitutions can be introduced at amino acid positions that show functional sensitivity to the initial substitution. Alternatively, or in addition, the crystal structure of the antigen-antibody complex to identify the contact points between the antibody and the antigen. Such contact residues and adjacent residues can be targeted or excluded as candidates for substitution. Mutants can be screened to determine whether they contain the desired properties.
[0090] Insertions of amino acid sequences include amino-terminal and / or carboxyl-terminal fusions to polypeptides containing from 1 residue to over 100 residues in length, as well as insertions between sequences of single or multiple amino acid residues. Examples of terminal insertions include antibodies having an N-terminal methionyl residue. Other insertion mutants of antibody molecules include fusions of the antibody to an enzyme (e.g., for ADEPT) or polypeptide that increases the serum half-life of the antibody at the N-terminal or C-terminal of the antibody.
[0091] The concept of modular protein expression by pre-mRNA trans-splicing in the context of a phage antibody display vector system is described in detail herein, but the application of the concepts exemplified by the use of the nucleic acid molecules, vectors, vector sets, host cells, and methods described herein can be adapted and extended to other techniques that require, for example, expressing many collections of proteins with different combinations of repeating modules in mammalian cells.
[0092] III. Examples The following are examples of the present invention. It should be understood that various other embodiments can be practiced based on the general description provided above.
[0093] Example 1. Generation of a modular protein expression system for antibody reformatting related to a phage display vector The generation of a polypeptide expression system for modular expression and production of polypeptides is described. The present invention is based, at least in part, on experimental findings showing that pre-mRNA trans-splicing can be utilized in mammalian cells to enable modular recombinant protein expression. The concept of modular protein expression allows for the precise joining of any two protein-coding sequences encoded by two different constructs into a single mRNA encoding a polypeptide chain without any of the requirements and constraints of other protein-protein splicing methods. The concept of modular protein expression by pre-mRNA trans-splicing can be adapted to simplify and expand other techniques that require the expression of many collections of proteins with different combinations of repeating modules in mammalian cells. For example, this concept will find use in other settings that require the expression of fusion protein partners or combinations of mutations in a single polypeptide. This technology is both simple and effective, enables application at any scale, and has broad importance for the field of recombinant protein expression in mammalian cells, which is the basis of much of modern biotechnology.
[0094] Here, the generation of such polypeptide expression systems that enable modular expression of different antibody formats in the context of phage display expression systems is described. Phage display has been widely used in the discovery and engineering of antibody fragments for the development of therapeutic and reagent antibodies (McCafferty et al. Nature. 348:552 - 554, 1990; Sidhu. Current opinion in biotechnology. 11:610 - 616, 2000; Smith. Science. 228:1315 - 1317, 1985). Phage display has conventionally enabled the rapid selection of antigen - specific binders but has limited the screening of the selected antibody fragments. The detailed characterization of antibody fragments often requires the expression of full - length immunoglobulin G (IgG) that is normally expressed in mammalian cells. However, one limiting step in this process is the reformatting of phage clones into mammalian expression vectors for IgG expression. High - throughput subcloning methods can be used to reformat a large number of clones, but these methods are usually relatively labor - intensive and produce many clones that will not be used beyond the screening stage.
[0095] To avoid the need for subcloning and enable modular protein expression, a first nucleic acid molecule, the dual - host vector, pDV2, was generated (Figure 1). Unlike the aforementioned dual - vector pDV (Tesar et al. Protein engineering, design & selection: PEDS. 26:655 - 662, 2013) that requires co - transfection of mammalian cells with an IgG expression cassette containing a signal sequence engineered for expression of the heavy chain in either bacteria or mammalian cells and a mammalian expression vector that expresses the light chain for full IgG expression, pDV2 contains most of the stII signal sequence embedded in an intron that is removed by splicing in bacterial promoters and mammalian cells.
[0096] The stII signal sequence in pDV2 was modified to include both the 3′ splice site (3′ss) and the polypyrimidine tract (PPT) optimized in front of the 3′ss. This required introducing three relatively conserved amino acid substitutions in the stII signal sequence, which did not affect the display of Fab fragments on the phage (Figure 2). To enable modular and flexible expression of antibody formats from the same clone, no complete introns and exons encoding the constant region were added downstream from the region encoding the CH1 domain. Instead, these heavy chain sequences were to be added in trans from a second nucleic acid molecule. To achieve this, the process of pre-mRNA trans-splicing, which binds two different pre-mRNAs to form a single mature mRNA, was utilized. Trans-splicing in mammalian cells was induced by hybridization of complementary sequences downstream from the 5′ss and upstream from the 3′ss, resulting in pre-mRNAs with single 5′ss and 3′ss, which could then form a single non-covalently bound pre-mRNA and be spliced as a normal pre-mRNA (Konarska et al. Cell. 42:165-171, 1985, Puttaraju et al. Nature biotechnology. 17:246-252, 1999, Solnick. Cell. 42:157-164, 1985). In this particular polypeptide expression system, a 150-bp fragment of the M13 gene III (gIII) was used as the hybridizing sequence (Figure 1). This gene III sequence follows the optimized GTAAGA 5′ss described above at the 3′ boundary of the sequence encoding CH1 (Tesar et al. Protein engineering, design & selection: PEDS. 26:655-662, 2013).
[0097] To complete the polypeptide expression system, a second nucleic acid molecule, pRK-Fc, was generated, which is a complementary plasmid that expresses a pre-mRNA containing a linker sequence, a consensus branch point, and a 150-nt antisense gene III sequence followed by a PPT, and a 3′ss followed by a hinge, CH2 and CH3 regions in one exon, and an SV40 polyadenylation signal (Figs. 1 and 3). This transcript does not encode a signal sequence, and the first two potential start codons are out-of-frame and located in the antisense gene III sequence and the hinge region. Thus, with the exception of the 5′ss, all other sequences required for splicing are encoded by pRK-Fc rather than pDV2. Co-transfection of Expi293F cells (Invitrogen) with pDV2 and pRK-Fc resulted in baseline but detectable expression levels of IgG (Fig. 4A).
[0098] Example 2. Generation of an optimized modular protein expression system for antibody reformatting related to phage display vectors The baseline IgG yields achieved by pDV2 and pRK-Fc may have been due to the absence of sequences required for efficient trans-splicing or sequences in vectors that inhibit trans-splicing. Nucleotide motifs in both exons and introns can act as splicing enhancers or suppressors or both, depending on their location. For the purposes of vector design, intron splicing enhancers (ISEs) can be easily added because they are unlikely to affect coding sequences in mammalian cell expression. One well-characterized ISE consists of a sequence of three or more consecutive guanine residues or G-run located near the intron boundary, which are bound by heterogeneous nuclear ribonucleoprotein H or F to enhance splicing (Wang et al. Nature structural&molecular biology. 19:1044-1052, 2012, Xiao et al. Nature structural&molecular biology. 16:1094-1100, 2009). In addition, intron sequences rich in purines that are not limited to G-run and are close to the 5′ss have also been shown to enhance splicing (Hastings et al. RNA. 7:859-874, 2001).
[0099] Therefore, mutants of pDV2 and pDV2b were created that contain a 9-nt G-run in the region encoding the linker between the upper hinge and the C-terminal of the M13 bacteriophage pIII coat protein (cP3), as well as a second 4-nt G-run 10 nt downstream, and a 26-bp purine-rich region 23 base pairs (bp) downstream from the 5′ss (Figure 5). This mutant changes the Gly-Arg-Pro linker between the upper hinge and cP3 to three Gly residues. The vector did not contain a standard polyadenylation site for the heavy chain cassette. This was because an attempt was made to minimize the formation of mature heavy chain mRNA from the vector, which would then be transported to the cytoplasm and potentially cause trans-splicing to occur, leading to the expression of the Fab-cP3 fusion protein. The pRK-Fc molecule was also optimized. The intron G-run near the 3′ss has been shown to stimulate splicing in vitro (Martinez-Contreras. PLoS biology. 4:e21, 2006). Therefore, a 9-nt ISE was added upstream from the branch site to generate the optimized complementary plasmid pRK-Fc2 (Figure 6). Co-transfection of human Expi293F cells with pDV2 and pRK-Fc (ISE-) or pRK-Fc2 (ISE+) resulted in baseline levels of IgG expression (Figure 4A). Co-transfection of Expi293F cells with the ISE+ pDV2b plasmid and pRK-Fc or pRK-Fc2 resulted in higher levels of IgG expression, and the highest expression level of up to 25 μg / ml produced by co-transfecting the ISE+ plasmids pDV2b and pRK-Fc2 indicates that the ISE sequences in both transcripts enhance the efficiency of trans-splicing.
[0100] The baseline IgG expression level in transfected Expi293F cells was associated with apparent cell lysis 7 days after transfection and was observed when pDV2 or pDV2b, but not pRK-Fc or pRK-Fc2, was transfected alone. Analysis of the transfected cell lysates by Western blotting using anti-M13p3 antibody revealed a polypeptide with an apparent molecular weight of approximately 41 kDa that corresponded to the expression of the IgG1 Fd fragment (VH-CH1-upper hinge) fused to the M13cP3 peptide (Figure 12, lower panel, lanes 3-6). The expression of this polypeptide was higher in cells transfected with pDV2 or pDV2b without a complementing plasmid. The results indicated that both the pDV2 and pDV2b plasmids were able to express mature mRNAs encoding potentially toxic products, despite the fact that both lacked a mammalian polyadenylation site downstream of the vector from the heavy chain cassette.
[0101] Visual inspection of the gene III sequence encoding cP3 revealed the AATAAA motif that can act as a polyadenylation site (Figure 2). Two silent mutations were introduced into this site to generate plasmids pDV2c (ISE-) and pDV2d (ISE+) to test whether this reduces toxicity and improves protein expression in mammalian cells. Co-transfection of Expi293F cells with pRK-Fc and pDV2c or pDV2d resulted in approximately 6-fold higher levels of IgG expression compared to the pDV2 and pDV2b vectors due to the potential polyadenylation site in gene III (Figure 4A). This increase in IgG expression level was associated with high viability of the transfected cells and significantly reduced or undetectable expression of the Fd-cP3 fusion protein in the transfected cells (Figure 12, lower panel, columns 7-10). This indicates that the presence of the potential polyadenylation site in the donor vector within gene III causes unwanted protein expression from the donor plasmid alone, which has a significant negative impact on protein expression. Co-transfection of Expi293F cells with pDV2c or pDV2d and the ISE+ pRK-Fc2 complementing vector resulted in an additional 2-fold increase in IgG expression compared to co-transfection with the ISE- pRK-Fc vector (Figure 4A). These results show that the major factor determining baseline protein expression in the pDV2 vector is the presence of the potential polyadenylation site in gene III, while on the other hand, the addition of ISE has a minor effect on protein expression when the potential gene III polyadenylation motif is absent. In contrast, the addition of ISE in the complementing pRK-Fc2 plasmid results in approximately 2-fold higher IgG yields when co-transfecting the pDV2 variant without the potential polyadenylation site in gene III (Figure 4A).
[0102] Further optimization of protein expression was achieved by determining the optimal DNA ratio for transfection. Using a 2:1 excess of the complementing plasmid pRK-Fc2 relative to pDV2d resulted in the highest IgG expression yield in this system (Figure 4B). Using pDV2d and pRK-Fc2 with the optimized DNA ratio, the yield of IgG purified from 30 ml of the supernatant of transfected Expi293F cells was 3.2 ± 1.2 mg (n = 3). IgG purified from Expi293F cells co-transfected with these plasmids could not be distinguished from the same IgG expressed by a conventional expression vector by mass spectrometry and SDS-PAGE (Figures 7A-7B and 13). Co-transfection of pDV2d encoding variable regions of different specificities with pRK-Fc2 having an optimized DNA ratio resulted in high IgG expression of 2.5 - 5.5 mg of IgG purified from 30 ml of the supernatant of transfected Expi293F cells (Figure 8A). The polypeptide expression system is not limited to the use of Expi293F cells to achieve high expression levels. Other mammalian cell lines widely used for IgG expression, such as 293T and CHO cells, were also effective. Co-transfection of pDV2d and pRK-Fc2 expressing variable regions of different specificities with 293T or CHO cells resulted in high IgG expression (Figure 8B).
[0103] When the pRK-Fc2 vector was co-transfected with the pDV2 plasmid, it was modified for the expression of the Fab fragment. The sequences encoding the lower hinge and Fc regions in pRK-Fc2 were removed and replaced with a Flag tag to produce the pRK-Fab-Flag vector (Figure 9). The yield of purified Fab fragment purified from 30 ml of the supernatant of Expi293F cells co-transfected with pDV2d and pRK-Fab-Flag was 0.8 ± 0.06 mg (mean ± standard deviation, n = 3). The structural integrity of the purified Flag-tagged Fab fragment was confirmed by mass spectrometry and SDS-PAGE (Figure 13). The observed heavy chain mass was 25,169 Da, which is close to the predicted mass of 25,172 Da excluding the clipped C-terminal lysine.
[0104] Expression of N-terminally truncated proteins from complementary transcripts has been observed in a trans-splicing system for gene therapy (Monjaret et al. Molecular therapy 22:1176-1187, 2014). This is due to a complementary transcript encoding a 3′ exon with all the elements necessary for the formation of mature mRNA, which may cause translation from an internal start codon. Western blotting of the lysates of cells transfected with pRK-Fc2 revealed the expression of a polypeptide consistent with the Fc fragment translated from the first in-frame ATG codon (Figure 12, upper panel, lane 11). This polypeptide probably lacks a secretion signal sequence and should be expressed only in the cytoplasm. This product may be released into the culture medium by cell lysis, but it was not observed in the purified IgG sample by SDS-PAGE (Figure 13, lane 2) and mass spectrometry. When pDV2c or pDV2d was co-transfected into the cells, the expression of this truncated product was reduced but not eliminated (Figure 12, upper panel, lanes 8 and 10). Insertion of an out-of-frame open reading frame with an optimal translation start site upstream of the intron region from the potential Fc start codon did not significantly reduce the expression of the truncated Fc product.
[0105] An important property of a phage display vector that determines selection efficiency is the level of antibody fragment display on the phage particles achieved. Along with reduced p3 expression in the E. coli SupE suppressor strain, using the aforementioned Amber-2614KO7 helper phage, the level of Fab fragment display achieved with the pDV2d vector was comparable to the Fab display levels achieved with a specialized Fab display vector, Fab-dip-phage, using the standard M13KO7 helper phage (Figure 14).
[0106] When different antibody formats, such as wild-type IgG, Fab fragments, or IgG with Fc modifications for bispecific formats, are required for different screening assays, the ability to express different antibody formats from the same clone is useful in antibody discovery. The vector set can, in principle, enable any of these or additional formats by simply cloning a suitable sequence added after the CH1 region in a complementary plasmid. Furthermore, the modular organization of this system requires only the construction of a new complementary plasmid, eliminating the need to recreate the stock of the phage display library and enabling the expression of new antibody formats. The dual vector can also be adapted to allow the use of any CH1 region by transferring the 5′ss from downstream of the CH1 coding region to the J region (FR4) of the VH or J-CH1 junction, thus separating the VH and the entire constant region of the heavy chain in two different plasmids. The amber stop codon at the junction of the heavy chain and the gene III sequence in the Fab phage display library usually results in a significantly lower level of display, and thus, with the knowledge that at least for inexperienced repertoire libraries, recloning of the clones is required after selection, the pDV2 vector is compatible with conventional methods for the expression of Fab fragments in E. coli by simply adding a stop codon after the sequence encoding the upper hinge (Lee et al., Journal of immunological methods. 284:119-132, 2004). The expression of Fab fragments in mammalian cells using the same method as used for IgG expression circumvents this need for recloning with a yield comparable to that normally obtained in E. coli.
[0107] Example 3. Modular Protein Expression System The polypeptide expression systems generated and characterized in Examples 1 and 2 show that modular and flexible polypeptide expression of any desired protein can be directly achieved by the use of a polypeptide expression system such as the optimized expression system described above for protein reformatting in the context of phage display. Thus, the present expression system comprises two nucleic acid molecule components (polypeptide coding sequences PES11 and PES2) each encoding a part of a single desired polypeptide product, where the split coding regions of these proteins are accurately joined together in vivo by pre-mRNA trans-splicing without the need for subcloning of the protein-coding nucleic acids. As shown in Figure 10, the first nucleic acid molecule contains an expression cassette having PES11, a eukaryotic promoter (P1 Euk1 ) and a eukaryotic signal sequence (ESS11) upstream of the PES11 component, and also includes a 5'splice site (5'ss11) and a hybridizing sequence (HS1) located downstream of PES11. The complementary second nucleic acid molecule will include a eukaryotic promoter (P2 Euk ), as well as a hybridizing sequence capable of hybridizing to HS1 (HS2), and a 3'splice site (3'ss2) upstream of the PES2 component. In addition, the second nucleic acid molecule will include a polyadenylation site (pA2) downstream of the PES2 component. Thus, when replicated in mammalian cells, the two generated pre-mRNA molecules, one having a solitary 5' and the other having a solitary 3', are oriented together by their complementary hybridizing sequences (HS1 and HS2) and undergo trans-splicing to form a single continuous mRNA capable of encoding the desired protein product for subsequent translation.
[0108] In prokaryotic cells, if expression of the polypeptide product encoded by PES11 and optionally the HS1 region is also desired, the first nucleic acid molecule may further comprise a removable prokaryotic promoter module (ePPM1) positioned between P1 Euk1 and PES11. ePPM1 contains a 5'splice site (5'ss12), a prokaryotic promoter (P1 Prok1) A first nucleic acid sequence encoding a prokaryotic signal sequence (PSS11), and may include a 3′ splice site (3′ss11), and 5′ss12-P1 Prok1 -PSS11-3′ss11 are operably linked to each other in the 5′ to 3′ direction. ePPM1 will drive the transcription of the polypeptide encoded by the first nucleic acid molecule in prokaryotic cells. On the other hand, in eukaryotic cells (e.g., mammalian cells), P1 Euk1 drives the expression of the transcription of the polypeptide encoded by the first nucleic acid molecule, and ePPM1 will be removed from the pre-mRNA transcript by cis-splicing due to the collision adjacent to the 5′ss12 and 3′ss11 components.
[0109] In some cases, it may be desirable to express a second polypeptide. Therefore, the first nucleic acid molecule of the modular protein expression system can be designed to include a second expression cassette. As shown in FIG. 11, the second expression cassette encoding a second protein product (PES12) is designed in a manner similar to the first expression cassette, but contains a polyadenylation site (pA1) downstream of the PES12 sequence to ensure the generation of a separate pre-mRNA molecule after transcription. In other cases, the second expression cassette can be designed into the second nucleic acid molecule of the polypeptide expression system.
[0110] Other embodiments The foregoing invention has been described in detail by way of illustration and example for the purpose of clear understanding, but the description and examples should not be construed as limiting the scope of the invention. The disclosures of all patents and scientific literature described herein are hereby expressly incorporated by reference in their entirety.
Claims
1. A polypeptide expression system comprising a first nucleic acid molecule and a second nucleic acid molecule, (a) the first nucleic acid molecule comprises the following components: (i) a first eukaryotic promoter (P1 Euk1 ), (ii) a first polypeptide coding sequence (PES1 1 ), (iii) the first 5' splice site (5'ss1 1 (iv) a hybridizing sequence (HS1), Euk1 -PES1 1 -5'ss1 1 - operably linked to each other in the 5' to 3' direction as HS1; (b) the second nucleic acid molecule comprises the following components: (i) a eukaryotic promoter (P2 Euk (ii) a hybridizing sequence (HS2) capable of hybridizing to HS1; (iii) a 3' splice site (3'ss2); (iv) a polypeptide coding sequence (PES2); and (v) a polyadenylation site (pA2), said components comprising P2 Euk -HS2-3'ss2-PES2-pA2, said polypeptide expression system being operably linked to each other in the 5' to 3' direction.
2. P1 Euk1 The polypeptide expression system of claim 1 , wherein the promoter is a cytomegalovirus (CMV) promoter or a simian virus 40 (SV40) promoter.
3. P2 Euk The polypeptide expression system according to claim 1 or 2, wherein the promoter is a CMV promoter or an SV40 promoter.
4. The first expression cassette comprises a eukaryotic signal sequence (ESS1 1 The method further comprises a first nucleic acid sequence encoding ESS1. 1 However, P1 Euk1 and the PES1 1 The polypeptide expression system according to any one of claims 1 to 3, which is located between
5. The ESS1 1 The polypeptide expression system of claim 4, wherein said VH gene is derived from a variable heavy chain (VH) gene.
6. The first expression cassette comprises the following components: (i) a 5' splice site (5'ss1 2 ), (ii) prokaryotic promoter (P1 Prok1 ), and (iii) the 3' splice site (3'ss1 1 ), an excisable prokaryotic promoter module (ePPM) 1 5'ss1 2 -P1 Prok1 -3'ss1 1 and the ePPM is operably linked to each other in the 5'-3' direction as 1 However, P1 Euk1 and the PES1 1 The polypeptide expression system according to any one of claims 1 to 5, which is located between
7. P1 Prok1 The polypeptide expression system of claim 6, wherein the promoter is selected from the group consisting of PhoA promoter, Tac promoter, Lac, and Tphac promoter.
8. The ePPM 1 However, the prokaryotic signal sequence (PSS1 1 8. The polypeptide expression system of claim 6 or 7, further comprising a first nucleic acid sequence encoding a polypeptide of the present invention.
9. The PSS1 1 The polypeptide expression system according to any one of claims 6 to 8, wherein is derived from the heat-stable enterotoxin II (stII) gene.
10. The PSS1 1 and the 3'ss1 1 A polypyrimidine region (PPT1 1 The polypeptide expression system according to any one of claims 6 to 9, further comprising:
11. PPT1 1 The polypeptide expression system of claim 10, comprising the nucleic acid sequence TTCCTTTTTTTCTCTTTCC (SEQ ID NO: 1).
12. The PES1 1 A polypeptide expression system according to any one of claims 1 to 11, which does not contain a cryptic 5' splice site.
13. The polypeptide expression system according to any one of claims 1 to 12, wherein the HS1 is a gene encoding all or part of a coat protein or an adaptor protein.
14. 14. The polypeptide expression system of claim 13, wherein the coat protein is selected from the group consisting of pI, pII, pIII, pIV, pV, pVI, pVII, pVIII, pIX, and pX of bacteriophage M13, f1, or fd.
15. 15. The polypeptide expression system of claim 14, wherein the coat protein is the pIII protein of bacteriophage M13.
16. 16. The polypeptide expression system of claim 15, wherein the pill fragment comprises amino acid residues 267 to 421 of the pill protein or amino acid residues 262 to 418 of the pill protein.
17. The polypeptide expression system of claim 13, wherein the adapter protein is a leucine zipper.
18. 18. The polypeptide expression system of claim 17, wherein the leucine zipper comprises the amino acid sequence of SEQ ID NO: 4 or 5.
19. The first nucleic acid molecule comprises a second eukaryotic promoter (P1 Euk2 ), (ii) a second polypeptide coding sequence (PES1 2 ), and (iii) a second expression cassette comprising a polyadenylation site (pA1), Euk2 -PES1 2 The polypeptide expression system according to any one of claims 1 to 18, wherein said pA1 and pB2 are operably linked to each other in the 5' to 3' direction as pA2.
20. P1 Euk2 The polypeptide expression system of claim 19, wherein the promoter is a CMV promoter or an SV40 promoter.
21. The second expression cassette comprises a eukaryotic signal sequence (ESS1 2 21. The polypeptide expression system of claim 19 or 20, further comprising a nucleic acid sequence encoding a polypeptide.
22. The ESS1 2 22. The polypeptide expression system of claim 21 , wherein said gene is derived from the mouse binding immunoglobulin protein (mBiP) gene.
23. The ESS1 2 23. The polypeptide expression system of any one of claims 19 to 22, wherein said nucleic acid sequence comprises the nucleic acid sequence ATG AAN TTN ACN GTN GTN GCN GCN GCN CTN CTN CTN CTN GGN (SEQ ID NO: 6), wherein N is A, T, C, or G.
24. The second expression cassette comprises the following components: (i) a 5' splice site (5'ss1 3 ), (ii) prokaryotic promoter (P1 Prok2 ), and (iii) the 3' splice site (3'ss1 2 ), an excisable prokaryotic promoter module (ePPM) 2 5'ss1 3 -P1 Prok2 -3'ss1 2 and the ePPM is operably linked to each other in the 5'-3' direction as 2 However, P1 Euk2 and the PES1 2 The polypeptide expression system according to any one of claims 19 to 23, which is located between
25. P1 Prok2 25. The polypeptide expression system of claim 24, wherein the promoter is selected from the group consisting of a PhoA promoter, a Tac promoter, and a Lac promoter.
26. The ePPM 2 However, the prokaryotic signal sequence (PSS1 2 26. The polypeptide expression system of claim 24 or 25, further comprising a nucleic acid sequence encoding a polypeptide.
27. The PSS1 2 The polypeptide expression system according to any one of claims 24 to 26, wherein is derived from the heat-stable enterotoxin II (stII) gene.
28. The PSS1 2 and the 3'ss1 2 A polypyrimidine region (PPT1 2 28. The polypeptide expression system according to any one of claims 24 to 27, further comprising:
29. PPT1 2 29. The polypeptide expression system of claim 28, comprising the nucleic acid sequence TTCCTTTTTTTCTCTTTCC (SEQ ID NO: 1).
30. 30. A polypeptide expression system according to any one of claims 19 to 29, wherein the second expression cassette is positioned 5' to the first expression cassette.
31. The 5'ss1 1 31. The polypeptide expression system of claim 1, further comprising an intron splice enhancer (ISE) (ISE1) located between said HS1 and said HS2.
32. The polypeptide expression system of claim 31, wherein the ISE1 comprises a G-run that contains three or more consecutive guanine residues.
33. The polypeptide expression system of claim 32, wherein the ISE1 comprises a G-run that includes nine consecutive guanine residues.
34. A polypeptide expression system according to any one of claims 1 to 33, further comprising a polypyrimidine tract (PPT2) located between the HS2 and the 3'ss2.
35. 35. The polypeptide expression system of claim 34, wherein the PPT2 comprises the nucleic acid sequence TTCCTCTTTCCCTTTCTCTCC (SEQ ID NO: 7).
36. 36. The polypeptide expression system of claim 35, further comprising an ISE (ISE2) located between the HS2 and the 3'ss2.
37. The polypeptide expression system of claim 36, wherein the ISE2 comprises a G-run that contains three or more consecutive guanine residues.
38. The polypeptide expression system of claim 37, wherein the ISE2 comprises a G-run that includes nine consecutive guanine residues.
39. The 5'ss1 1 A polypeptide expression system according to any one of claims 1 to 38, comprising the nucleic acid sequence GTAAGA (SEQ ID NO: 8).
40. A polypeptide expression system according to any one of claims 1 to 39, wherein expression by a eukaryotic promoter occurs in a mammalian cell.
41. 41. The polypeptide expression system of claim 40, wherein the mammalian cell is an Expi293F cell, a CHO cell, a 293T cell, or an NSO cell.
42. 42. The polypeptide expression system of claim 41, wherein the mammalian cell is an Expi293F cell.
43. A polypeptide expression system according to any one of claims 6 to 42, wherein expression by a prokaryotic promoter occurs in a bacterial cell.
44. 44. The polypeptide expression system of claim 43, wherein the bacterial cell is an E. coli cell.
45. The PES1 1 A polypeptide expression system according to any one of claims 1 to 44, which encodes all or part of an antibody.
46. The PES1 1 46. The polypeptide expression system of claim 45, wherein said polypeptide expression system encodes a polypeptide comprising a VH domain.
47. 47. The polypeptide expression system of claim 46, wherein the polypeptide further comprises a CH1 domain.
48. 48. A polypeptide expression system according to any one of claims 45 to 47, wherein said PES2 encodes all or part of an antibody.
49. 49. The polypeptide expression system of claim 48, wherein the PES2 encodes a polypeptide comprising a CH2 domain and a CH3 domain.
50. The PES1 2 A polypeptide expression system according to any one of claims 19 to 49, which encodes all or part of an antibody.
51. The PES1 2 51. The polypeptide expression system of claim 50, wherein said polypeptide expression system encodes a polypeptide comprising a VL domain and a CL domain.
52. Components: (a) a first eukaryotic promoter (P1 Euk1 )and, (b) the following components: (i) 5' splice site (5'ss1 2 ), (ii) Prokaryotic promoter (P1 Prok1 ), and (iii) 3' splice site (3'ss1 1 ) A first excisable prokaryotic promoter module (ePPM) comprising 1 ) wherein the ePPM 1 The component of 5'ss1 2 -P1 Prok1 -3'ss1 1 a first excisable prokaryotic promoter module operably linked to each other in a 5' to 3' direction as (c) a first polypeptide coding sequence (PES1 1 )and, (d) the first 5' splice site (5'ss1 1 )and, (e) a useful peptide coding sequence (UPES); and wherein the components of the first expression cassette include P1 Euk1 - ePPM 1 -PES1 1 -5'ss1 1 - the nucleic acid molecules operably linked to each other in the 5' to 3' direction as UPES.
53. The first expression cassette comprises a eukaryotic signal sequence (ESS1 1 The method further comprises a first nucleic acid sequence encoding ESS1. 1 However, P1 Euk1 and the ePPM 1 53. The nucleic acid molecule of claim 52, wherein said nucleic acid molecule is located between
54. The ePPM 1 However, the prokaryotic signal sequence (PSS1 1 The method further comprises the steps of: 1 However, P1 Prok1 and the 3'ss1 1 54. The nucleic acid molecule of claim 52 or 53, wherein said nucleic acid molecule is located between:
55. A second eukaryotic promoter (P1 Euk2 ), (ii) a second polypeptide coding sequence (PES1 2 ), and (iii) a second expression cassette comprising a polyadenylation site (pA1), Euk2 -PES1 2 - The nucleic acid molecule according to any one of claims 52 to 54, which are operably linked to each other in the 5' to 3' direction as pA1.
56. The second expression cassette comprises a eukaryotic signal sequence (ESS1 2 ESS1), 2 However, P1 Euk2 and the PES1 2 56. The nucleic acid molecule of claim 55, wherein said nucleic acid molecule is located between
57. The second expression cassette comprises the following components: (i) a 5' splice site (5'ss1 3 ), (ii) prokaryotic promoter (P1 Prok2 ), and (iii) the 3' splice site (3'ss1 2 ), an excisable prokaryotic promoter module (ePPM) 2 5'ss1 3 -P1 Prok2 -3'ss1 2 and the ePPM is operably linked to each other in the 5'-3' direction as 2 However, P1 Euk2 and the PES1 2 57. The nucleic acid molecule of claim 55 or 56, wherein said nucleic acid molecule is located between
58. The ePPM 2 However, the prokaryotic signal sequence (PSS1 2 The method further comprises the step of: 2 However, P1 Prok2 and the 3'ss1 2 58. The nucleic acid molecule of claim 57, wherein said nucleic acid molecule is located between
59. The nucleic acid molecule of any one of claims 52 to 58, wherein the UPES encodes all or part of a useful peptide selected from the group consisting of tags, labels, coat proteins, and adaptor proteins.
60. 60. The nucleic acid molecule of claim 59, wherein the coat protein is selected from the group consisting of pI, pII, pIII, pIV, pV, pVI, pVII, pVIII, pIX, and pX of bacteriophage M13, f1, or fd.
61. 61. The nucleic acid molecule of claim 60, wherein the coat protein is the pIII of bacteriophage M13.
62. A vector comprising the nucleic acid molecule of any one of claims 52 to 61.
63. A vector set comprising a first vector and a second vector, the first and second vectors comprising the first and second nucleic acid molecules, respectively, of the polypeptide expression system according to any one of claims 1 to 51.
64. 64. A host cell comprising the vector of claim 62 or the vector set of claim 63.
65. 65. The host cell of claim 64, wherein the host cell is a prokaryotic cell.
66. 66. The host cell of claim 65, wherein the prokaryotic cell is a bacterial cell.
67. 67. The host cell of claim 66, wherein the bacterial cell is an E. coli cell.
68. 65. The host cell of claim 64, wherein the host cell is a eukaryotic cell.
69. 69. The host cell of claim 68, wherein the eukaryotic cell is a mammalian cell.
70. 70. The host cell of claim 69, wherein the mammalian cell is an Expi293F cell, a CHO cell, a 293T cell, or an NSO cell.
71. 71. The host cell of claim 70, wherein the mammalian cell is an Expi293F cell.
72. 64. A method for producing a polypeptide comprising culturing a host cell comprising the vector of claim 62 or the vector set of claim 63 in a culture medium.
73. 73. The method of claim 72, wherein the method further comprises recovering the polypeptide from the host cell or the culture medium.