Fusion partners for peptide production

By designing a recombinant fusion protein containing N-terminal bacterial fusion partner and target polypeptide, the problem of difficult to express heterologous recombinant polypeptides in the bacterial expression system is solved, and efficient and economical peptide production is achieved.

CN113444183BActive Publication Date: 2025-06-17PELICAN TECHNOLOGY HOLDINGS INC
View PDF 29 Cites 0 Cited by

Patent Information

Application Number
CN202110725675.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2014-12-01
Filing Date
2015-11-30
Publication Date
2025-06-17
Estimated Expiration
2035-11-30

AI Technical Summary

Technical Problem

Heterologous recombinant peptides are difficult to express in high yields in bacterial expression systems, mainly due to proteolysis, low expression levels, incorrect protein folding and poor secretion of host cells.

Method used

A recombinant fusion protein is designed, including N-terminal bacterial fusion partners such as DnaJ-like proteins, target polypeptides and flexible linker sequences, which are linked by protease cleavage sites, allowing target polypeptides to be expressed in bacterial host cells at high yields.

Benefits of technology

The purification process is simplified by titration of higher than 0.5 g/L in bacterial host cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113444183B_ABST
    Figure CN113444183B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of medicine and, in particular, to the production of large amounts of soluble recombinant polypeptides as part of a fusion protein, said fusion protein comprising an N-terminal fusion partner linked to a target polypeptide.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a patent application filed on November 30, 2015, with application number 201580064973.4 and titled “Fusion Partner for Peptide Production”. Background Art

[0002] Heterologous recombinant polypeptides are often difficult to express in high yields in bacterial expression systems for reasons that include proteolysis, low expression levels, incorrect protein folding (which can lead to poor solubility), and poor secretion from host cells. Summary of the Invention

[0003] The present invention provides recombinant fusion proteins comprising a polypeptide of interest. Expression of the polypeptide of interest as part of the recombinant fusion protein allows for the production of high-quality polypeptides in large quantities. The polypeptide of interest includes small or rapidly degradable peptides, such as the N-terminal fragment of parathyroid hormone (PTH1-34), proteins with easily degradable N-termini, such as GCSF and the Plasmodium falciparum circumsporozoite protein, and proteins that are typically produced in an insoluble form in microbial expression systems, such as proinsulin, GCSF, or interferon-β, which can be processed into insulin or insulin analogs. Figure 1 The recombinant fusion protein illustrated in the figure comprises an N-terminal bacterial fusion partner, for example, a bacterial chaperone or a folding regulator. The target polypeptide and the N-terminal bacterial chaperone or folding regulator are connected by a flexible linker sequence containing a protease cleavage site. Upon cleavage, the target polypeptide is released from the N-terminal fusion partner. The present invention further discloses a vector for expressing the recombinant fusion protein, and a method for producing the recombinant fusion protein in a bacterial host cell with high yield.

[0004] The recombinant fusion construct of the present invention is useful for producing high-yield recombinant target polypeptides (which are difficult to overexpress in bacterial expression systems due to, for example, proteolysis, low expression levels, poor folding and / or poor secretion). In embodiments of the present invention, the recombinant fusion protein of the present invention is produced in bacterial host cells at a titer higher than 0.5 g / L. In some embodiments, the bacterial host cell in which the recombinant target polypeptide is difficult to overexpress is Escherichia coli.

[0005] For example, the PTH 1-34 protein, previously reported as being expressed as part of a fusion protein in inclusion bodies (which required high concentrations of urea (e.g., 7 M) for solubilization), is described herein as being expressed at high titers (above 0.5 g / L) as part of a soluble PTH 1-34 fusion protein. Furthermore, purification can be performed under non-denaturing conditions, e.g., 4 M or lower urea concentrations, or without the use of urea at all. Furthermore, using the methods of the present invention, proteins with easily degradable N-termini, e.g., N-met-GCSF or Plasmodium falciparum circumsporozoite protein, can be produced as part of the fusion protein and separated from the N-terminal fusion partner by cleavage after removal of host cell proteases from the fusion protein preparation. Also as described herein, proinsulin, which is normally produced in an insoluble form, can be produced in large quantities in a soluble form in the recombinant fusion protein of the present invention, eliminating the need for refolding.

[0006] The present invention therefore provides recombinant fusion proteins comprising: an N-terminal fusion partner, wherein the N-terminal fusion partner is a bacterial chaperone protein or a folding regulator; a polypeptide of interest; and a linker comprising a cleavage site between the N-terminal fusion partner and the polypeptide of interest. In some embodiments, the N-terminal fusion partner is selected from: a DnaJ-like protein; an Fk1B protein or a truncate thereof; an EcpD protein or a truncate thereof; or a Skp protein or a truncate thereof. In some embodiments, the N-terminal fusion partner is selected from: a Pseudomonas fluorescens DnaJ-like protein; a Pseudomonas fluorescens Fk1B protein or a C-terminal truncate thereof; a Pseudomonas fluorescens FrnE protein or a truncate thereof; a Pseudomonas fluorescens FkpB2 protein or a C-terminal truncate thereof; or a Pseudomonas fluorescens EcpD protein or a C-terminal truncate thereof. In certain embodiments, the N-terminal fusion partner is a truncation of the Pseudomonas fluorescens Fk1B protein with 1 to 200 amino acids removed from the C-terminus, a truncation of the Pseudomonas fluorescens EcpD protein with 1 to 200 amino acids removed from the C-terminus, or a truncation of the Pseudomonas fluorescens FrnE protein with 1 to 180 amino acids removed from the C-terminus. In some embodiments, the polypeptide of interest is a difficult-to-express protein selected from the group consisting of: small or rapidly degradable peptides; proteins with easily degradable N-termini; and proteins that are typically expressed in an insoluble form in bacterial expression systems. In some embodiments, the target polypeptide is a small or rapidly degradable peptide, wherein the target polypeptide is selected from the group consisting of: hPTH1-34, Glp1, Glp2, IGF-1 Exenatide (SEQ ID NO: 37), Teduglutide (SEQ ID NO: 38), Pramlintide (SEQ ID NO: 39), Ziconotide (SEQ ID NO: 40), Becaplermin (SEQ ID NO: 42), Enfuvirtide (SEQ ID NO: 43), Nesiritide (SEQ ID NO: 44). In some embodiments, the target polypeptide is a protein with an N-terminus that is easily degraded, wherein the target polypeptide is N-met-GCSF or Plasmodium falciparum circumsporozoite protein. In some embodiments, the target polypeptide is a protein that is normally expressed as an insoluble protein in a bacterial expression system, wherein the target polypeptide is proinsulin, GCSF, or IFN-β that can be processed into insulin or an insulin analog. In any of these embodiments, the proinsulin C-peptide has an amino acid sequence selected from SEQ ID NO:97; SEQ ID NO:98; SEQ ID NO:99 or SEQ ID NO:100.In some embodiments, the insulin analog is insulin glargine, insulin aspart, insulin lispro, insulin glulisine, insulin detemir, or insulin degludec. In certain embodiments, the N-terminal fusion partner is a Pseudomonas fluorescens DnaJ-like protein having the amino acid sequence shown in SEQ ID NO: 2. In some embodiments, the N-terminal fusion partner is a Pseudomonas fluorescens Fk1B protein having the amino acid sequence shown in SEQ ID NO: 4, SEQ ID NO: 28, SEQ ID NO: 61, or SEQ ID NO: 62. In some embodiments, the N-terminal fusion partner is a Pseudomonas fluorescens FrnE-like protein having the amino acid sequence shown in SEQ ID NO: 3, SEQ ID NO: 63, or SEQ ID NO: 64. In some embodiments, the N-terminal fusion partner is a Pseudomonas fluorescens EcpD protein having the amino acid sequence set forth in SEQ ID NO: 7, SEQ ID NO: 65, SEQ ID NO: 66, or SEQ ID NO: 67. In some embodiments, the cleavage site in the recombinant fusion protein is recognized by a cleavage enzyme selected from the group consisting of enterokinase, trypsin, factor Xa, and furin. The recombinant fusion protein of any one of claims 1 to 15, wherein the linker comprises an affinity tag. In certain embodiments, the affinity tag is selected from the group consisting of: polyhistidine; FLAG tag; myc tag; GST tag; NBP tag; calmodulin tag; HA tag; E-tag; S-tag; SBP tag; Softtag3; V5 tag; and VSV tag. In some embodiments, the linker has an amino acid sequence selected from the group consisting of: SEQ ID NO: 9; SEQ ID NO: 10; SEQ ID NO: 11; SEQ ID NO: 12; and SEQ ID NO: 226. In some embodiments, the polypeptide of interest is hPTH1-34, and the recombinant fusion protein comprises an amino acid sequence selected from the group consisting of: SEQ ID NO: 45; SEQ ID NO: 46; and SEQ ID NO: 47. In some embodiments, the isoelectric point of the polypeptide of interest is at least about 1.5 times higher than the isoelectric point of the N-terminal fusion partner. In some embodiments, the molecular weight of the polypeptide of interest comprises about 10% to about 50% of the molecular weight of the recombinant fusion protein.

[0007] The present invention also provides an expression vector for expressing a recombinant fusion protein. In some embodiments, the expression vector is used to express the recombinant fusion protein in any of the above embodiments. In some embodiments, the expression vector comprises a nucleotide sequence encoding the recombinant fusion protein in any of the above embodiments.

[0008] The present invention further provides a method for producing a target polypeptide, comprising:

[0009] (i) culturing a microbial host cell transformed with an expression vector comprising an expression construct, wherein the expression construct comprises a nucleotide sequence encoding a recombinant fusion protein; (ii) inducing the host cell of step (i) to express the recombinant fusion protein; (iii) purifying the recombinant fusion protein expressed in the induced host cell of step (ii); and (iv) cleaving the purified recombinant fusion protein of step (iii) by incubating with a cleavage enzyme that recognizes a cleavage site in the linker to release the target polypeptide; thereby obtaining the target polypeptide. In some embodiments, the recombinant fusion protein of step (i) is the recombinant fusion protein described in any of the above embodiments. In some embodiments, the method further comprises measuring the expression level of the fusion protein expressed in step (ii), measuring the amount of the recombinant fusion protein purified in step (iii), or measuring the amount of the correctly released target polypeptide obtained in step (iv), or a combination thereof. In some embodiments, the expression level of the fusion protein expressed in step (ii) is higher than 0.5 g / L. In some embodiments, the expression level of the fusion protein expressed in step (ii) is about 0.5 g / L to about 25 g / L. In some embodiments, the fusion protein expressed in step (ii) is directed to the cytoplasm. In some embodiments, the fusion protein expressed in step (ii) is directed to the periplasm. In some embodiments, the incubation in step (iv) is from about one hour to about 16 hours, and the lytic enzyme is enterokinase.

[0010] In some embodiments, the incubation of step (iv) is about one hour to about 16 hours, and the lytic enzyme is enterokinase, and the amount of the recombinant fusion protein purified in the step (iii) of the correct release in step (iv) is about 90% to about 100%. In some embodiments, the amount of the recombinant fusion protein purified in the step (iii) of the correct release in step (iv) is about 100%. In some embodiments, the amount of the target polypeptide obtained in step (iii) or step (iv) is about 0.1g / L to about 25g / L. In some embodiments, the target polypeptide of the correct release obtained is soluble, complete or both. In some embodiments, step (iii) is carried out under non-denaturing conditions. In some embodiments, the recombinant fusion protein dissolves without using urea. In some embodiments, non-denaturing conditions include lysing the induced cells of step (ii) with a buffer comprising a non-denaturing concentration of a chaotropic agent (chaotropic agent). In some embodiments, the chaotropic agent of the non-denaturing concentration is urea lower than 4M.

[0011] In some embodiments, the microbial host cell is a Pseudomonas or E. coli host cell. In some embodiments, the Pseudomonas host cell is a Pseudomonas host cell. In some embodiments, the Pseudomonas host cell is Pseudomonas fluorescens.

[0012] In certain embodiments, the host cell is deficient in at least one protease selected from the group consisting of Lon (SEQ ID NO: 14); La1 (SEQ ID NO: 15); AprA (SEQ ID NO: 16); HtpX (SEQ ID NO: 17); DegP1 (SEQ ID NO: 18); DegP2 (SEQ ID NO: 19); Npr (SEQ ID NO: 20); Prc1 (SEQ ID NO: 21); Prc2 (SEQ ID NO: 22); M50 (SEQ ID NO: 24); PrlC (SEQ ID NO: 30); Serralysin (RXF04495) (SEQ ID NO: 227); and PrtB (SEQ ID NO: 23). In related embodiments, the host cell is deficient in the proteases Lon (SEQ ID NO: 14), La1 (SEQ ID NO: 15), and AprA (SEQ ID NO: 16). In some embodiments, the host cell is deficient in proteases AprA (SEQ ID NO: 16) and HtpX (SEQ ID NO: 17). In other embodiments, the host cell is deficient in proteases Lon (SEQ ID NO: 14), La1 (SEQ ID NO: 15), and DegP2 (SEQ ID NO: 19). In some embodiments, the host cell is deficient in proteases Npr (SEQ ID NO: 20), DegP1 (SEQ ID NO: 18), and DegP2 (SEQ ID NO: 19). In related embodiments, the host cell is deficient in proteases Serralysin (SEQ ID NO: 227) and AprA (SEQ ID NO: 16).

[0013] Incorporated by reference

[0014] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The novel features of the present invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description of illustrative embodiments in which the principles of the invention are utilized, and to the accompanying drawings:

[0016] Figure 1Schematic representation of a recombinant fusion protein. Domain 1 corresponds to the N-terminal fusion partner, Domain 2 corresponds to the linker, and Domain 3 corresponds to the target polypeptide. Non-limiting examples of N-terminal fusion partners and target polypeptides are listed below each corresponding domain.

[0017] Figures 2A to 2C Three recombinant fusion protein amino acid sequences. The amino acid sequences of three recombinant fusion proteins containing hPTH 1-34 as the target polypeptide are shown. hPTH 1-34 is italicized in each instance, and the linker between the N-terminal fusion partner and PTH 1-34 is underlined. 2A. Recombinant fusion protein containing a DnaJ-like protein N-terminal fusion partner. (DnaJ-like protein, aa 1-77; linker, aa 78-98; hPTH 1-34, aa 99-132) (SEQ ID NO:45). 2B. Recombinant fusion protein containing a Fk1B N-terminal fusion partner. (Fk1B, aa 1-205; linker, aa 206-226; hPTH 1-34, aa 227-260) (SEQ ID NO:46). 2C. Recombinant fusion protein containing a FrnE N-terminal fusion partner. (FrnE, aa 1-216; linker, aa 217-237; hPTH 1-34, aa 238-271) (SEQ ID NO: 47).

[0018] Figure 3 SDS-CGE analysis of shake flask expression samples. Samples are shown in three groups: intact cell culture medium (lanes 1-6); cell-free culture medium (lanes 7-12); and soluble fraction (lanes 13-18), as indicated at the bottom of the figure. Molecular weight markers are indicated on each side of the image (68, 48, 29, 21, 16 kD, from top to bottom). Lanes in each of the three groups represent, from left to right: DNAJ-like protein-PTH 1-34 fusion (STR35970); DNAJ-like protein-PTH 1-34 fusion (STR35984); Fk1B-PTH 1-34 fusion (STR36034); Fk1B-PTH 1-34 fusion (STR36085); FrnE-PTH 1-34 fusion (STR36150); and FrnE-PTH 1-34 fusion (STR36169), as indicated above the lanes. The DnaJ-like-PTH fusion protein band is marked by a solid arrow, and the Fk1B-PTH and FrnE-PTH fusion protein bands are marked by dotted arrows.

[0019] Figure 4Enterokinase cleavage of purified recombinant fusion proteins. Samples are shown in three groups: no enterokinase treatment (lanes 1-6); enterokinase treatment at 40 μg / ml (lanes 7-12); and enterokinase treatment at 10 μg / ml (lanes 13-18). Lanes in each of the three groups represent, from left to right: DNAJ-like protein-PTH 1-34 fusion (STR35970); DNAJ-like protein-PTH 1-34 fusion (STR35984); Fk1B-PTH 1-34 fusion (STR36034); Fk1B-PTH 1-34 fusion (STR36085); FrnE-PTH 1-34 fusion (STR36150); and FrnE-PTH 1-34 fusion (STR36169). The migration of DnaJ-like fusion proteins is indicated by the solid arrow in the lower pair of arrows. The migration of the cleaved DnaJ-like protein N-terminal fusion partner is indicated by the dashed arrow in the lower pair of arrows. The migration of the Fk1B and FrnE fusion proteins is indicated by the solid arrow in the upper pair of arrows. The migration of the Fk1B and FrnE N-terminal fusion partners is indicated by the dashed arrow in the upper pair of arrows. Molecular weight markers are shown on the right side of the image (29, 20, and 16 kD, from top to bottom).

[0020] Figure 5 Intact mass analysis of enterokinase cleavage products. Shown is the deconvoluted mass spectrum of a DnaJ-like protein-PTH 1-34 fusion protein purified from expression strain STR36970 following 1 hour of enterokinase digestion. The peak corresponding to PTH 1-34 is indicated by a solid arrow.

[0021] Figure 6 Enterokinase cleavage of purified fractions of DnaJ-like protein-PTH 1-34 fusion protein. DnaJ-like protein-PTH fusion protein was purified from expression strain STR36005 after growth in a conventional bioreactor. Purified fractions were incubated with enterokinase for 1 hour (lanes 2-4), 16 hours (lanes 6-8), without enterokinase (control) for 1 hour (lane 1), or without enterokinase (control) for 16 hours (lane 5). The fractions analyzed were as follows: fraction 1 (lanes 1, 2, 5, and 6); fraction 2 (lanes 3 and 7); and fraction 3 (lanes 4 and 8). The full-length DnaJ-like protein-PTH 1-34 recombinant fusion protein band is indicated by a solid black arrow. The cleaved DnaJ-like protein-PTH 1-34 fusion partner band is indicated by a dashed arrow. Molecular weight markers are shown on each side of the image (49, 29, 21, and 16 kD, from top to bottom).

[0022] Figures 7A to 7CIntact mass analysis of PTH 1-34 enterokinase cleavage products derived from the Fk1B-PTH 1-34 fusion protein. These figures show deconvoluted mass spectra of purified fractions of the Fk1B-PTH 1-34 fusion protein digested with enterokinase. The peak corresponding to PTH 1-34 is indicated by a solid arrow. 7A. Fk1B-PTH fusion protein purified from STR36034. 7B. Fk1B-PTH fusion protein purified from STR36085. 7C. Fk1B-PTH fusion protein purified from STR36098.

[0023] sequence

[0024] The present invention includes nucleotide sequences SEQ ID NO: 1-237, and these nucleotide sequences are listed in the sequence listing preceding the claims. DETAILED DESCRIPTION

[0025] Overview

[0026] The present invention relates to a recombinant fusion protein for overexpressing a recombinant target polypeptide in a bacterial expression system, a construct for expressing the recombinant fusion protein, and a method for producing a soluble form of the recombinant fusion protein and the recombinant target polypeptide in high yield. In some embodiments, the method of the present invention can produce a recombinant fusion protein higher than 0.5 g / L after purification. In some embodiments, the method of the present invention produces a high yield of recombinant fusion protein without the need to use a denaturing concentration of a chaotropic agent. In some embodiments, the method of the present invention produces a high yield of recombinant fusion protein without the need to use any chaotropic agent.

[0027] As used herein, the term "comprise" or variations thereof, such as "comprises" or "comprising" are understood to mean the inclusion of any recited features but not the exclusion of any other features. Thus, as used herein, the term "comprising" is inclusive and does not exclude other, unrecited features. In any of the embodiments of the compositions and methods provided herein, "comprising" can be replaced with "consisting essentially of" or "consisting of." The phrase "consisting essentially of" is used herein to claim the specified features as well as those features that do not materially affect the properties or functions of the claimed invention. As used herein, the term "consisting of" is used to indicate the presence of only the recited features (e.g., nucleobase sequences) (such that in the case of an antisense oligomer consisting of a specified nucleobase sequence, the presence of other, unrecited nucleobases is excluded).

[0028] Recombinant fusion protein

[0029] The recombinant fusion protein of the present invention comprises three domains, such as Figure 1. From left to right, the fusion protein comprises an N-terminal fusion partner, a linker, and a target polypeptide, wherein the linker is between the N-terminal fusion partners and the target polypeptide is at the C-terminus of the linker. In some embodiments, the linker sequence comprises a protease cleavage site. In some embodiments, the target polypeptide can be released from the recombinant fusion protein by cleavage at the protease cleavage site within the linker.

[0030] In some embodiments, the molecular weight of the recombinant fusion protein is about 2 kDa to about 1000 kDa. In some embodiments, the molecular weight of the recombinant fusion protein is about 2 kDa, about 3 kDa, about 4 kDa, about 5 kDa, about 6 kDa, about 7 kDa, about 8 kDa, about 9 kDa, about 10 kDa, about 11 kDa, about 12 kDa, about 13 kDa, about 14 kDa, about 15 kDa, about 20 kDa, about 25 kDa, about 26 kDa, about 27 kDa, about 28 kDa, about 30 kDa, about 35 kDa, about 40 kDa, about kDa, about 45 kDa, about 50 kDa, about 55 kDa, about 60 kDa, about 65 kDa, about 70 kDa, about 75 kDa, about 80 kDa, about 85 kDa, about 90 kDa, about 95 kDa, about 100 kDa, about 200 kDa, about 300 kDa, about 400 kDa, about 500 kDa, about 550 kDa, about 600 kDa, about 700 kDa, about 800 kDa, about 900 kDa, about 1000 kDa or more.In some embodiments, the recombinant fusion protein has a molecular weight of about 2 kDa to about 1000 kDa, about 2 kDa to about 500 kDa, about 2 kDa to about 250 kDa, about 2 kDa to about 100 kDa, about 2 kDa to about 50 kDa, about 2 kDa to about 25 kDa, about 2 kDa to about 30 kDa, about 2 kDa to about 1000 kDa, about 2 kDa to about 500 kDa, about 2 kDa to about 250 kDa, about 2 kDa to about 100 kDa, about 2 kDa to about 50 kDa, about 2 kDa to about 25 kDa, about 3kDa to about 1000kDa, about 3kDa to about 500kDa, about 3kDa to about 250kDa, about 3kDa to about 100kDa, about 3kDa to about 50kDa, about 3kDa to about 25kDa, about 3kDa to about 30kDa, about 4kDa to about 1000kDa, about 4kDa to about 500kDa, about 4kDa to about 250kDa, about 4kDa to about 100kDa, about 4kDa to about 50kDa, about 4kDa to about 25kDa, about 4kDa to about 30kDa, about 5kDa to about 1000kDa, about 5kDa to about 500kDa, about 5kDa to about 250kDa, about 5kDa to about 100kDa, about 5kDa to about 50kDa, about 5kDa to about 25kDa, about 5kDa to about 30kDa, about 10kDa to about 1000kDa, about 10kDa to about 500kDa, about 10kDa to about 250kDa, about 10kDa to about 100kDa, about 10kDa to about 50kDa, about 10kDa to about 25kDa, about 10kDa to about 30kDa, about 20kDa to about about 1000 kDa, about 20 kDa to about 500 kDa, about 20 kDa to about 250 kDa, about 20 kDa to about 100 kDa, about 20 kDa to about 50 kDa, about 20 kDa to about 25 kDa, about 20 kDa to about 30 kDa, about 25 kDa to about 1000 kDa, about 25 kDa to about 500 kDa, about 25 kDa to about 250 kDa, about 25 kDa to about 100 kDa, about 25 kDa to about 50 kDa, about 25 kDa to about 25 kDa or about 25 kDa to about 30 kDa.

[0031] In some embodiments, the recombinant fusion protein is about 50, 100, 150, 200, 250, 300, 350, 400, 450, 470, 500, 530, 560, 590, 610, 640, 670, 700, 750, 800, 850, 900, 950, 1000, 1200, 1400, 1600, 1800, 2000, 2500 or more amino acids in length. In some embodiments, the recombinant fusion protein is about 50 to 2500, 100 to 2000, 150 to 1800, 200 to 1600, 250 to 1400, 300 to 1200, 350 to 1000, 400 to 950, 450 to 900, 470 to 850, 500 to 800, 530 to 750, 560 to 700, 590 to 670, or 610 to 640 amino acids in length.

[0032] In some embodiments, the recombinant fusion protein comprises an N-terminal fusion partner selected from:

[0033] Pseudomonas fluorescens DnaJ-like protein (e.g., SEQ ID NO: 2), FrnE (SEQ ID NO: 3), FrnE2 (SEQ ID NO: 63), FrnE3 (SEQ ID NO: 64), FklB (SEQ ID NO: 4), FklB3* (SEQ ID NO: 28), FklB2 (SEQ ID NO: 61), FklB3 (SEQ ID NO: 62), FkpB2 (SEQ ID NO: 5), SecB (SEQ ID NO: 6), truncations of SecB, EcpD (SEQ ID NO: 7), EcpD (SEQ ID NO: 65), EcpD2 (SEQ ID NO: 66), and EcpD3 (SEQ ID NO: 67);

[0034] a linker selected from the group consisting of SEQ ID NOs: 9, 10, 11, 12, and 226; and

[0035] A polypeptide of interest selected from the group consisting of hPTH 1-34 (SEQ ID NO: 1), Met-GCSF (SEQ ID NO: 69), rCSP, proinsulin (e.g., human proinsulin SEQ ID NO: 32, insulin glargine, proinsulin SEQ ID NO: 88, 89, 90, or 91), insulin lispro SEQ ID NO: 33, insulin glulisine (SEQ ID NO: 34), insulin C-peptide (SEQ ID NO: 97); mecasermin (SEQ ID NO: 35), Glp-1 (SEQ ID NO: 36), exenatide (SEQ ID NO: 37), teduglutide (SEQ ID NO: 38), pramlintide (SEQ ID NO: 39), ziconotide (SEQ ID NO: 40), becaplermin (SEQ ID NO: 42), enfuvirtide SEQ ID NO: 43), nesiritide (SEQ ID NO: 44), or enterokinase (e.g., SEQ ID NO: 31).

[0036] In some embodiments, the recombinant fusion protein comprises a Pseudomonas fluorescens DnaJ-like protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 101. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 101 is SEQ ID NO: 202.

[0037] In some embodiments, the recombinant fusion protein comprises an N-terminal fusion partner of the Pseudomonas fluorescens EcpD1 protein and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 102 or 103. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 102 or 103 is SEQ ID NO: 202 or 228, respectively.

[0038] In some embodiments, the recombinant fusion protein comprises a Pseudomonas fluorescens EcpD2 protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 104. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 104 is SEQ ID NO: 204.

[0039] In some embodiments, the recombinant fusion protein comprises a Pseudomonas fluorescens EcpD3 protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 105. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 105 is SEQ ID NO: 205.

[0040] In some embodiments, the recombinant fusion protein comprises a Pseudomonas fluorescens Fk1B1 protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 106. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 106 is SEQ ID NO: 206.

[0041] In some embodiments, the recombinant fusion protein comprises the Pseudomonas fluorescens Fk1B2 protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 107. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 107 is SEQ ID NO: 207.

[0042] In some embodiments, the recombinant fusion protein comprises the Pseudomonas fluorescens Fk1B3 protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 108. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 108 is SEQ ID NO: 208.

[0043] In some embodiments, the recombinant fusion protein comprises a Pseudomonas fluorescens FrnE1 protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 109. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 109 is SEQ ID NO: 209.

[0044] In some embodiments, the recombinant fusion protein comprises a Pseudomonas fluorescens FrnE2 protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 110. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 110 is SEQ ID NO: 210.

[0045] In some embodiments, the recombinant fusion protein comprises a Pseudomonas fluorescens FrnE3 protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 111. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 111 is SEQ ID NO: 211.

[0046] In some embodiments, the recombinant fusion protein comprises a Pseudomonas fluorescens DnaJ-like protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 112. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 112 is SEQ ID NO: 212.

[0047] In some embodiments, the recombinant fusion protein comprises the Pseudomonas fluorescens EcpD1 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 113. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 113 is SEQ ID NO: 213, respectively.

[0048] In some embodiments, the recombinant fusion protein comprises a Pseudomonas fluorescens EcpD2 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 114. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 114 is SEQ ID NO: 214.

[0049] In some embodiments, the recombinant fusion protein comprises a Pseudomonas fluorescens EcpD3 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 115. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 115 is SEQ ID NO: 215.

[0050] In some embodiments, the recombinant fusion protein comprises the Pseudomonas fluorescens Fk1B1 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 216. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 116 is SEQ ID NO: 216.

[0051] In some embodiments, the recombinant fusion protein comprises the Pseudomonas fluorescens Fk1B2 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 217. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 117 is SEQ ID NO: 217.

[0052] In some embodiments, the recombinant fusion protein comprises the Pseudomonas fluorescens Fk1B3 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 118. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 118 is SEQ ID NO: 218.

[0053] In some embodiments, the recombinant fusion protein comprises a Pseudomonas fluorescens FrnE1 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 119. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 119 is SEQ ID NO: 219.

[0054] In some embodiments, the recombinant fusion protein comprises a Pseudomonas fluorescens FrnE2 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 120. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 120 is SEQ ID NO: 220.

[0055] In some embodiments, the recombinant fusion protein comprises a Pseudomonas fluorescens FrnE3 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 121. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 121 is SEQ ID NO: 221.

[0056] In some embodiments, the N-terminal fusion partner, linker, and target polypeptide of the recombinant fusion protein are: Pseudomonas fluorescens fold regulator DnaJ-like protein (SEQ ID NO: 2), the linker set forth in SEQ ID NO: 9, and human parathyroid hormone amino acids 1-34 (hPTH 1-34) (SEQ ID NO: 1), respectively. In some embodiments, the N-terminal fusion partner, linker, and target polypeptide of the recombinant fusion protein are: Pseudomonas fluorescens fold regulator FrnE (SEQ ID NO: 3), the linker set forth in SEQ ID NO: 9, and hPTH 1-34 (SEQ ID NO: 1), respectively. In some embodiments, the N-terminal fusion partner, linker, and target polypeptide of the recombinant fusion protein are: Pseudomonas fluorescens fold regulator Fk1B (SEQ ID NO: 4), the linker set forth in SEQ ID NO: 9, and hPTH 1-34 (SEQ ID NO: 1), respectively. In some embodiments, the recombinant hPTH fusion protein has the amino acid sequence set forth in SEQ ID NOs: 45, 46, and 47.

[0057] In some embodiments, the recombinant fusion protein is an insulin fusion protein having the following elements:

[0058] an N-terminal fusion partner selected from Pseudomonas fluorescens: DnaJ-like protein (SEQ ID NO: 2), FrnE (SEQ ID NO: 3), FrnE2 (SEQ ID NO: 63), FrnE3 (SEQ ID NO: 64), Fk1B (SEQ ID NO: 4), Fk1B3* (SEQ ID NO: 28), Fk1B2 (SEQ ID NO: 61), Fk1B3 (SEQ ID NO: 62), FkpB2 (SEQ ID NO: 5), EcpD EcpD (SEQ ID NO: 65), EcpD2 (SEQ ID NO: 66), or EcpD3 (SEQ ID NO: 67);

[0059] a linker having the sequence shown in SEQ ID NO: 226; and

[0060] A target polypeptide selected from proinsulin glargine SEQ ID NO: 88, 89, 90 or 91.

[0061] In some embodiments, the target polypeptide is proinsulin glargine as set forth in SEQ ID NO:88, encoded by the nucleotide sequence set forth in SEQ ID NO:80 or 84. In some embodiments, the target polypeptide is proinsulin glargine as set forth in SEQ ID NO:89, encoded by the nucleotide sequence set forth in SEQ ID NO:81 or 85. In some embodiments, the target polypeptide is proinsulin glargine as set forth in SEQ ID NO:90, encoded by the nucleotide sequence set forth in SEQ ID NO:82 or 86. In some embodiments, the target polypeptide is proinsulin glargine as set forth in SEQ ID NO:91, encoded by the nucleotide sequence set forth in SEQ ID NO:83 or 87.

[0062] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens DnaJ-like protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 101. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 101 is SEQ ID NO: 202.

[0063] In some embodiments, the insulin fusion protein comprises an N-terminal fusion partner of the Pseudomonas fluorescens EcpD1 protein and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 102 or 103. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 102 or 103 is SEQ ID NO: 202 or 228, respectively.

[0064] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens EcpD2 protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 104. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 104 is SEQ ID NO: 204.

[0065] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens EcpD3 protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 105. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 105 is SEQ ID NO: 205.

[0066] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens Fk1B1 protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 106. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 106 is SEQ ID NO: 206.

[0067] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens Fk1B2 protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 107. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 107 is SEQ ID NO: 207.

[0068] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens Fk1B3 protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 108. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 108 is SEQ ID NO: 208.

[0069] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens FrnE1 protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 109. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 109 is SEQ ID NO: 209.

[0070] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens FrnE2 protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 110. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 110 is SEQ ID NO:210.

[0071] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens FrnE3 protein N-terminal fusion partner and a trypsin cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 111. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 111 is SEQ ID NO: 211.

[0072] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens DnaJ-like protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 112. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 112 is SEQ ID NO: 212.

[0073] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens EcpD1 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 113. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 113 is SEQ ID NO: 213, respectively.

[0074] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens EcpD2 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 114. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 114 is SEQ ID NO: 214.

[0075] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens EcpD3 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 115. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 115 is SEQ ID NO: 215.

[0076] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens Fk1B1 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 216. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 116 is SEQ ID NO: 216.

[0077] In some embodiments, the insulin fusion protein comprises the Pseudomonas fluorescens Fk1B2 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 217. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 117 is SEQ ID NO: 217, respectively.

[0078] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens Fk1B3 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 118. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 118 is SEQ ID NO: 218.

[0079] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens FrnE1 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 119. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 119 is SEQ ID NO:219.

[0080] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens FrnE2 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 120. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 120 is SEQ ID NO: 220.

[0081] In some embodiments, the insulin fusion protein comprises a Pseudomonas fluorescens FrnE3 protein N-terminal fusion partner and an enterokinase cleavage site linker, which together have the amino acid sequence of SEQ ID NO: 121. In some embodiments, the nucleotide sequence encoding SEQ ID NO: 121 is SEQ ID NO: 221.

[0082] In some embodiments, the recombinant insulin fusion protein has the amino acid sequence set forth in any one of SEQ ID NOs: 122 to 201.

[0083] In some embodiments, the recombinant fusion protein is a GCSF fusion protein having the following elements:

[0084] an N-terminal fusion partner selected from Pseudomonas fluorescens: DnaJ-like protein (SEQ ID NO: 2), FrnE (SEQ ID NO: 3), FrnE2 (SEQ ID NO: 63), FrnE3 (SEQ ID NO: 64), Fk1B (SEQ ID NO: 4), Fk1B3* (SEQ ID NO: 28), Fk1B2 (SEQ ID NO: 61), Fk1B3 (SEQ ID NO: 62), FkpB2 (SEQ ID NO: 5), EcpD EcpD (SEQ ID NO: 65), EcpD2 (SEQ ID NO: 66), or EcpD3 (SEQ ID NO: 67);

[0085] a linker having the sequence shown in SEQ ID NO: 9; and

[0086] The target polypeptide has the sequence shown in SEQ ID NO:68.

[0087] Target peptide

[0088] The target protein or polypeptide of a recombinant fusion protein, also referred to as a target C-terminal polypeptide, a recombinant target polypeptide, and a C-terminal fusion partner, is a polypeptide that is desired to be expressed in a soluble form and in high yield. In some embodiments, the target polypeptide is a heterologous polypeptide that has been found to be unable to be expressed in high yield in a bacterial expression system due to, for example, proteolysis, low expression levels, incorrect protein folding, and / or poor host cell secretion. Target polypeptides include small or rapidly degradable peptides, proteins with easily degradable N-termini, and proteins that are normally produced in an insoluble form in microbial or bacterial expression systems. In some embodiments, the N-terminus of the target polypeptide is protected from degradation while being fused to an N-terminal fusion partner to form a higher yield of the protein with the N-terminus intact. In some embodiments, heterologous polypeptides have been described as not being expressed in a high yield and in a soluble form in a microbial or bacterial expression system. For example, in some embodiments, heterologous polypeptides have been described as not being expressed in high yield and soluble form in Escherichia coli, Bacillus subtilis (B.subtilis) or Lactobacillus plantarum (L.plantarum), Lactobacillus casei (L.casei), Lactobacillus fermentum (L.fermentum) or Corynebacterium glutamicum (corynebacterium glutamicum) host cells. In some embodiments, the target polypeptide is a eukaryotic polypeptide or is derived from a eukaryotic polypeptide (e.g., an analog thereof). In some embodiments, the target polypeptide is a mammalian polypeptide or is derived from a mammalian polypeptide. In some embodiments, the target polypeptide is a human polypeptide or is derived from a human polypeptide. In some embodiments, the target polypeptide is a prokaryotic polypeptide or is derived from a prokaryotic polypeptide. In some embodiments, the target polypeptide is a microbial polypeptide or is derived from a microbial polypeptide. In some embodiments, the target polypeptide is a bacterial polypeptide or is derived from a bacterial polypeptide. "Heterologous" means that the target polypeptide is derived from an organism other than the expression host cell. In some embodiments, according to the methods of the present invention, the fusion protein and / or target polypeptide is produced in a Pseudomonas host cell (i.e., a host cell of the order Psudomonadales) at a higher yield than in another microbial expression system. In some embodiments, according to the methods of the present invention, under substantially equivalent conditions, the fusion protein or target polypeptide is produced in a Pseudomonas, Pseudomonas, or Pseudomonas fluorescens expression system at a higher yield than in E. coli or other microbial or bacterial expression systems (e.g., those listed above), e.g., about 1.5-fold to about 10-fold, about 1.5-fold, about 2-fold, about 2.5-fold, about 3-fold, about 5-fold, or about 10-fold higher. In some embodiments, the fusion protein or C-terminal polypeptide is produced in an E. coli expression system at a yield of less than 0.5, less than 0.4, less than 0.3, less than 0.2, or less than 0.1 g / L.

[0089] In some embodiments, the target polypeptide is a small and / or rapidly degrading peptide. In some embodiments, the small and / or rapidly degrading peptide is parathyroid hormone (PTH). In some embodiments, the target polypeptide is hPTH 1-34 (SEQ ID NO: 1). PTH is an 84 amino acid (aa) peptide derived from a 115 aa pre-pro-peptide secreted by the parathyroid gland that serves to increase calcium concentrations in the blood and is known to stimulate bone formation. The N-terminal 34 aa peptide is approved for the treatment of osteoporosis ( Eli Lilly and Company; see package insert). The active ingredient in PTH1-34 is produced in E. coli as part of a C-terminal fusion protein (for is NDA 21-319; see, Chemistry Review, Center for Drug Evaluation and Research, 2000-2001; see also Clinical Pharmacology and Biopharmaceutics review, Center for Drug Evaluation and Research, 2000-2001). Purification of (Eli Lilly's LY333334) is described, for example, in Jin et al. ("Crystal Structure of Human Parathyroid Hormone 1-34at Resolution" J. Biol. Chem. 275(35):27238-44, 2000), which is incorporated herein by reference. This report describes the expression of the protein as inclusion bodies and subsequent solubilization in 7M urea.

[0090] In some embodiments, when overexpressed in a bacterial expression system, the target polypeptide is typically produced in an insoluble form. In some embodiments, the target polypeptide that is typically produced in an insoluble form when overexpressed in a bacterial expression system is a eukaryotic polypeptide or a derivative or analog thereof. In some embodiments, the target polypeptide that is typically produced in an insoluble form when overexpressed in a bacterial expression system is proinsulin (a precursor of insulin). Proinsulin consists of three designated segments (from N to C-terminus: BCA). When the internal C-peptide is removed by protease cleavage, proinsulin is processed into insulin (or an insulin analog, depending on the proinsulin). After the C-peptide insulin is removed, the disulfide bond between the A and B-peptides maintains their association. With respect to insulin and insulin analogs herein, "A-peptide" and "A-chain" are used interchangeably, and "B-peptide" and "B-chain" are used interchangeably. Positions within these chains are referred to by the chain and the amino acid number starting from the amino terminus of the chain, for example, "B30" refers to the 30th amino acid in the B-peptide (i.e., B-chain). In some embodiments, the polypeptide of interest is proinsulin that is processed to form a long-acting insulin analog or a rapid-acting insulin analog.

[0091] In some embodiments, the polypeptide of interest is proinsulin that is processed to form a long-acting insulin analog. Long-acting insulin analogs include, for example, insulin glargine, 43 amino acids (6050.41 Da), as Long-acting insulin analogues marketed as degludec and insulin detemir, as Sale. In insulin glargine, the asparagine (Asn21) of N21 is replaced by glycine, and there are two arginines at the C-terminus of the B-peptide. In insulin, these two arginines are present in preinsulin, but are not present in processed mature molecules. In some embodiments, the target polypeptide is processed into insulin glargine, and the target polypeptide is the 87 amino acid preinsulin shown in SEQ ID NO:88, 89, 90 or 91. In non-limiting embodiments, the coding sequence for SEQ ID NO:88 is the nucleotide sequence shown in SEQ ID NO:80 or 84. In non-limiting embodiments, the coding sequence for SEQ ID NO:89 is the nucleotide sequence shown in SEQ ID NO:81 or 85. In non-limiting embodiments, the coding sequence for SEQ ID NO:90 is the nucleotide sequence shown in SEQ ID NO:82 or 86. In non-limiting embodiments, the coding sequence for SEQ ID NO:91 is the nucleotide sequence shown in SEQ ID NO:83 or 87. Each of SEQ ID NOs: 80-87 includes an initial 15 bp cloning site at the 5' end, so in these embodiments, the proinsulin coding sequence referred to is the sequence starting at the first Phe codon, TTT (in SEQ ID NO: 80) or TTC (in SEQ ID NOs: 81-87). Degludec has the threonine at position B30 deleted and is coupled to hexadecanedioic acid at the amino acid lysine at position B29 via a γ-L-glutamyl spacer. Detemir has a fatty acid (myristic acid) bound to the lysine amino acid at position B29.

[0092] In some embodiments, the polypeptide of interest is a proinsulin that is processed to form a fast-acting insulin analog. Fast-acting (or rapid-acting) insulin analogs include, for example, insulin aspart. (SEQ ID NO:94), wherein the proline at position B28 is substituted with aspartic acid, and insulin lispro (pro-insulin lispro, SEQ ID NO: 33), in which the last lysine and proline residues occurring at the C-terminal end of the B-chain are reversed, and glulisine (proinsulin glulisine, SEQ ID NO: 34), wherein the asparagine at position B3 is replaced by lysine and the lysine at position B29 is replaced by glutamic acid. At all other positions, these molecules have the same amino acid sequence as regular insulin (proinsulin, SEQ ID NO: 32; insulin A-peptide, SEQ ID NO: 92; insulin B-peptide, SEQ ID NO: 93).

[0093] In some embodiments, the polypeptide of interest that is typically produced in an insoluble form when overexpressed in a bacterial overexpression system is GCSF, e.g., Met-GCSF. In some embodiments, the polypeptide of interest that is typically produced in an insoluble form when overexpressed in a bacterial overexpression system is IFN-β, e.g., IFN-β-1b. In some embodiments, the bacterial expression system in which the recombinant polypeptide of interest is difficult to overexpress is an E. coli expression system.

[0094] In some embodiments, the target polypeptide is a protein with a susceptible N-terminus. Because the fusion protein produced according to the methods of the present invention is cleaved from host proteases prior to cleavage to release the target polypeptide, the N-terminus of the target polypeptide is protected throughout the purification process. This allows for the production of preparations with up to 100% intact N-terminal target polypeptide.

[0095] In some embodiments, the polypeptide of interest having a degradation-prone N-terminus is filgrastim, an analog of GCSF (granulocyte colony-stimulating factor, or colony-stimulating factor 3 (CSF3)). GCSF is a 174-amino acid glycoprotein that stimulates the bone marrow to produce granulocytes and stem cells and release them into the bloodstream. Filgrastim, which is non-glycosylated and has an N-terminal methionine, acts as

[0014] The amino acid sequence of GCSF (filgrastim) is shown in SEQ ID NO: 69. In some embodiments, the methods of the present invention are used to produce high levels of GCSF (filgrastim) with an intact N-terminus (including an N-terminal methionine). Production of GCSF in protease-deficient host cells is described in U.S. Patent No. 8,455,218, "Methods for G-CSF production in a Pseudomonas host cell," which is incorporated herein by reference in its entirety. In embodiments of the present invention, intact GCSF, including an N-terminal methionine, is produced at high levels within a fusion protein in a non-protease-deficient bacterial host cell (e.g., a Pseudomonas host cell).

[0096] In some embodiments, the polypeptide of interest having a degradation-prone N-terminus is recombinant Plasmodium falciparum circumsporozoite protein (rCSP), described, for example, in U.S. Patent No. 9,169,304, “Process for Purifying Recombinant Plasmodium Falciparum Circumsporozoite Protein,” which is incorporated herein by reference in its entirety.

[0097] In some embodiments, the polypeptide of interest is: a reagent protein; a therapeutic protein; an extracellular receptor or ligand; a protease; a kinase; a blood protein; a chemokine; a cytokine; an antibody; an antibody-based drug; an antibody fragment, e.g., a single-chain antibody, an antigen-binding (ab) fragment, e.g., F(ab), F(ab)', F(ab)'2, Fv, generated from the variable region of IgG or IgM, an Fc fragment generated from the heavy chain constant region of an antibody, a reduced IgG fragment (e.g., generated by reducing the disulfide bonds of the hinge region of IgG), an Fc fusion protein, e.g., comprising the Fc domain of IgG fused to a protein or peptide of interest, or any other antibody fragment described in the art, e.g., U.S. Patent No. 5,648,237, "Expression of Functional Antibody Fragments," incorporated herein by reference in its entirety; an anticoagulant; a blood factor; a bone morphogenetic protein; an engineered protein scaffold; an enzyme; a growth factor; an interferon; an interleukin; a thrombolytic agent; or a hormone. In some embodiments, the polypeptide of interest is selected from the group consisting of: human antihemophilic factor; human antihemophilic factor-von Willebrand factor complex; recombinant antihemophilic factor (Turoctocog Alfa); Ado-trastuzumab emtansine; Albiglutide; alglucosidase alfa; human alpha-1 proteinase inhibitor; Rimabotulinumtoxin B; coagulation factor IX Fc fusion; recombinant coagulation factor IX; recombinant coagulation factor VIIa; recombinant coagulation factor XIII. A-subunit; human coagulation factor VIII-von Willebrand factor complex; Clostridium histolyticum collagenase; human platelet-derived growth factor (cecaplermin); abatacept; abciximab; adalimumab; aflibercept; β-galactosidase; aldesleukin; alefacept; alemtuzumab; alglucosidase alfa; alteplase; anakinra; octocoglin alfa; recombinant human antithrombin; azficel-T; basiliximab; belatacept; belimumab; bevacizumab; botulinum toxin type A; brentuximab Vedotin; Recombinant C1 esterase inhibitor; Canakinumab; Certolizumab Pegol; Cetuximab; Nonacog alfa; Daclizumab;Darbepoetin alfa; Denosumab; Digoxin immune Fab; Dornase alfa; Ecallantide; Eculizumab; Etanercept; Fibrinogen; Filgrastim; Galsulfase; Golimumab; Ibritumomab Tiuxetan; idursulfase; infliximab; interferon alpha; interferon alpha-2b; interferon Alfacon-1; interferon alpha-2a; interferon alpha-n3; interferon beta-1a; interferon beta-1b; interferon gamma-1b; golimumab; laronidase; epoetin alfa; murocoag alfa; muromonab -CD3; natalizumab; ocriplasmin; ofatumumab; omalizumab; oprelleukin; palifermin; palivizumab; panitumumab; pegylated filgrastim; pertuzumab; human papillomavirus (HPV) type 6, 11, 16, and 18-L1 viral protein virus-like particles (VLPs); HPV type 16 and 18L1 protein VLPs; ranibizumab; rasburicase; raxibacumab Recombinant factor IX; Reteplase; Rilonacept; Rituximab; Romiplostim; Sarmostim; Tenecteplase; Tocilizumab; Trastuzumab; Ustekinumab; Abarelix; Cetrorelix; Desirudin; Enfuvirtide; Exenatide; Follitropin beta; Ganirelix; Degarelix; Hyaluronidase; Insulin aspart; Insulin degludec; Insulin detemir; Insulin glargine rDNA injection (long-acting) (a rapid-acting human insulin analog); recombinant insulin glulisine; human insulin; insulin lispro (rapid-acting insulin analog); recombinant insulin lispro protamine; recombinant insulin lispro; lanreotide; liraglutide; Surfaxin (lucinactant; sinapeptide); mecasermin; insulin-like growth factor; nesiritide; pramlintide; recombinant teduglutide; tesmorelin acetate; ziconotide acetate; 10.8 mg goserelin acetate implant; botulinum toxin type A (Abobotulinumtoxin A); galactosidase alfa; alipogene tiparvovec; anciplostatin; anistreplase; aldesparin sodium; avian TB vaccine; batroxobin; bivalirudin; buserelin (gonadotropin-releasing hormone agonist); cabozantinib malate S; caperitide; catumaxomab;Ceruletide; Factor VIII; Coccidiosis vaccine; Daltrioparin sodium; Deferiprone; Defibrotide; Dibotermin Alfa; Drotrecogin Alfa; Edotreotide; Efalizumab; Enoxaparin sodium; Epoetin Delta; Eptifibatide; Eptotermin Alfa; Follicle-stimulating hormone alfa for injection; Fomivirsen; Gemtuzumab ozogamicin; Gonadorelin; Recombinant chronic human gonadorelin; Histrelin acetate (gonadorelin-releasing hormone agonist); HVT IBD vaccine; imiglucerase; oligoprotamine insulin; lepirudin; dog leptospirosis vaccine; leuproprelin; linaclotide; lipegfilgrastim; lixisenatide; luteinizing hormone alfa (human luteinizing hormone); mepolizumab; mifatide; mipomersen sodium; milipasin (macrophage-colony stimulating factor); mogamulizumab; molasprasin (granulocyte-macrophage-colony stimulating factor); monteplase; nadroparin calcium; nafarelin; nebacumab; octreotide; pamiplase; pancreatic lipase; parvaparin sodium; pasireotide daspartate; peginesatide acetate; pegvisomant; pentreotide; poractant alfa; pralmorelin (growth hormone-releasing peptide); prorelin; PTH 1-84; rhBMP-2; rhBMP-7; Eptortermin Alfa; Romotide; Sermorelin; Somatostatin; Somatrem; Vasopressin; Desmopressin; Taliglucerase Alfa; Tatirelin (thyrotropin-releasing hormone analog); Tasonamine; Taslutide; Thrombomodulin Alfa; Thyrotropin alfa; Trifermin; Triptorelin pamoate; Urofollicle-stimulating hormone for injection; Urokinase; Velaglucerase Alfa; Cholera toxin B; Recombinant antihemophilic factor (Efraloctocog Alfa); Human alpha-1 proteinase inhibitor; Erwinia chrysanthemi asparaginase; Caproumab; Denileukin Diftitox; Sheep digoxin immune Fab; Elosulfase Alfa; erythropoietin alpha; factor IX complex; factor XIII concentrate; technetium (Fanolesomab); fibrinogen; thrombin; influenza hemagglutinin and neuraminidase; Glucarpidase; hemin for injection; hepatitis B surface antigen; human albumin; incobotulinumtoxin; novofluvizumab; atezolizumab;L-asparaginase (from Escherichia coli; Erwinia; Pseudomonas; etc.); Pembrolizumab; Protein C concentrate; Ramucirumab; Siltuximab; Tbo-filgrastim; Pertussis toxin subunit AE; Local bovine thrombin; Local human thrombin; Tositumomab; Vedolizumab; Ziv-Aflibercept; Glucagon; Somatropin; Plasmodium falciparum or Plasmodium vivax antigens (e.g., CSP, CelTOS, TRAP, Rh5, AMA-1, LSA-1, LSA-3, Pfs25, MSP-1, MSP-3, STARP, EXP1, pb9, GLURP). The sequences of these polypeptides, including variations, are available in the literature and are known to those skilled in the art. Any known sequence of any of the listed peptides is contemplated for use in the methods of the present invention.

[0098] In some embodiments, the polypeptide of interest is enterokinase (e.g., SEQ ID NO: 31 [bovine]), insulin, proinsulin (e.g., SEQ ID NO: 32), a long-acting insulin analog or proinsulin processed to form a long-acting insulin analog (e.g., insulin glargine, SEQ ID NO: 88, insulin detemir, or insulin degludec), a rapid-acting insulin analog or proinsulin processed to form a rapid-acting insulin analog (e.g., insulin lispro, insulin aspart, or insulin glulisine), insulin C-peptide (e.g., SEQ ID NO: 97), IGF-1 (e.g., mecasermin, SEQ ID NO: 35), Glp-1 (e.g., SEQ ID NO: 36), a Glp-1 analog (e.g., exenatide, SEQ ID NO: 37), Glp-2 (e.g., SEQ ID NO: 38), a Glp-2 analog (e.g., teduglutide, SEQ ID NO: 39), pramlintide (e.g., SEQ ID NO: 40), ziconotide (e.g., SEQ ID NO: 41), or tadalafil (e.g., SEQ ID NO: 43). NO:41), becaplermin (e.g., SEQ ID NO:42), enfuvirtide (e.g., SEQ ID NO:43), or nesiritide (e.g., SEQ ID NO:44).

[0099] In some embodiments, the molecular weight of the polypeptide of interest is about 1 kDa, about 2 kDa, about 3 kDa, about 4 kDa, about 5 kDa, about 6 kDa, about 7 kDa, about 8 kDa, about 9 kDa, about 10 kDa, about 11 kDa, about 12 kDa, about 13 kDa, about 14 kDa, about 15 kDa, about 16 kDa, about 17 kDa, about 18 kDa, about 19 kDa, about 20 kDa, about 30 kDa, about 40 kDa, about 50 kDa, about 60 kDa, about 70 kDa, about 80 kDa, about 90 kDa, about 100 kDa, about 150 kDa, about 200 kDa, about 250 kDa, about 300 kDa, about 350 kDa, about 400 kDa, about 450 kDa, about 500 kDa or more. In some embodiments, the molecular weight of the recombinant polypeptide is about 1 to about 10 kDA, about 1 to about 20 kDA, about 1 to about 30 kDA, about 1 to about 40 kDA, about 1 to about 50 kDA, about 1 to about 60 kDA, about 1 to about 70 kDA, about 1 to about 80 kDA, about 1 to about 90 kDA, about 1 to about 100 kDA, about 1 kDa to about 200 kDa, about 1 kDa to about 300 kDa, about 1 kDa to about 400 kDa, about 1 kDa to about 500 kDa, about 2 to about 10 kDA, about 2 to about 20 kDA, about 2 to about 30 kDA, about 2 to about 40 kDA, about 2 to about 50 kDA, about 2 to about 60 kDA, about 2 to about 70 kDA, kDa, about 2 to about 80 kDA, about 2 to about 90 kDA, about 2 to about 100 kDA, about 2 kDa to about 200 kDa, about 2 kDa to about 300 kDa, about 2 kDa to about 400 kDa, about 2 kDa to about 500 kDa, about 3 to about 10 kDA, about 3 to about 20 kDA, about 3 to about 30 kDA, about 3 to about 40 kDA, about 3 to about 50 kDA, about 3 to about 60 kDA, about 3 to about 70 kDA, about 3 to about 80 kDA, about 3 to about 90 kDA, about 3 to about 100 kDA, about 3 kDa to about 200 kDa, about 3 kDa to about 300 kDa, about 3 kDa to about 400 kDa, or about 3 kDa to about 500 kDa. In some embodiments, the molecular weight of the polypeptide of interest is about 4.1 kDa.

[0100] In some embodiments, the target polypeptide is 25 or more amino acids long. In some embodiments, the target polypeptide is about 25 or about 2000 or more amino acids long. In some embodiments, the target polypeptide is about or at least about 25, 30, 35, 40, 45, 50, 100, 150, 200, 250, 300, 350, 400, 450, 475, 500, 525, 550, 575, 600, 625, 650, 700, 750, 800, 850, 900, 950, 1000, 1200, 1400, 1600, 1800, or 2000 amino acids long. In some embodiments, the target polypeptide is about: 25 to about 2000, 25 to about 1000, 25 to about 500, 25 to about 250, 25 to about 100, or 25 to about 50 amino acids long. In some embodiments, the polypeptide of interest is 32, 36, 39, 71, 109, or 110 amino acids long. In some embodiments, the polypeptide of interest is 34 amino acids long.

[0101] N-terminal fusion partner

[0102] The N-terminal fusion partner of a recombinant fusion protein is a bacterial protein that increases the yield of the recombinant fusion protein obtained using a bacterial expression system. In some embodiments, the N-terminal fusion partner can be stably overexpressed from a recombinant construct in a bacterial host cell. In some embodiments, the yield and / or solubility of the target polypeptide can be increased or improved by the presence of the N-terminal fusion partner. In some embodiments, the N-terminal fusion partner promotes the correct folding of the recombinant fusion protein. In some embodiments, the N-terminal fusion partner is a bacterial folding regulator or chaperone protein.

[0103] In some embodiments, the N-terminal fusion partner is a large affinity tag protein, a folding regulator, a molecular chaperone protein, a ribosomal protein, a translation-related factor, an OB-fold protein (oligopeptide-binding folding protein), or another protein described in the literature, for example, Ahn et al., 2011, "Expression screening of fusion partners from an E. coli genome for soluble expression of recombinant proteins in a cell-free protein synthesis system", PLoS One, 6(11): e26875, incorporated herein by reference. In some embodiments, the N-terminal fusion partner is a large affinity tag protein selected from the N-terminal domain of MBP, GST, NusA, ubiquitin, domain 1 of IF-2, and L9. In some embodiments, the N-terminal fusion partner is a ribosomal protein from the 30S ribosomal subunit, or a ribosomal protein from the 50S ribosomal subunit. In some embodiments, the N-terminal fusion partner is an E. coli or Pseudomonas chaperone protein or a folding regulator protein. In some embodiments, the N-terminal fusion protein is a Pseudomonas fluorescens chaperone protein or folding regulator protein. In some embodiments, the N-terminal fusion partner is a chaperone protein or folding regulator protein selected from Table 1.

[0104] In some embodiments, the N-terminal fusion partner is Pseudomonas fluorescens DnaJ-like protein (SEQ ID NO: 2), FrnE (SEQ ID NO: 3), FrnE2 (SEQ ID NO: 63), FrnE3 (SEQ ID NO: 64), FklB (SEQ ID NO: 4), FklB3* (SEQ ID NO: 28), FklB2 (SEQ ID NO: 61), FklB3 (SEQ ID NO: 62), FkpB2 (SEQ ID NO: 5), SecB (SEQ ID NO: 6), EcpD (RXF04553.1, SEQ ID NO: 7), EcpD (RXF04296.1, SEQ ID NO: 65, also referred to herein as EcpD1), EcpD2 (SEQ ID NO: 66), or EcpD3 (SEQ ID NO: 67). In some embodiments, the N-terminal fusion partner is the Escherichia coli protein Skp (SEQ ID NO: 8).

[0105] In some embodiments, the N-terminal fusion partner is truncated relative to the full-length fusion partner polypeptide. In some embodiments, the N-terminal fusion partner is truncated from the C-terminus to remove at least one C-terminal amino acid. In some embodiments, the N-terminal fusion partner is truncated to remove 1 to 300 amino acids from the C-terminus of the full-length polypeptide. In some embodiments, the N-terminal fusion partner is truncated to remove 300, 290, 280, 270, 260, 250, 240, 230, 220, 210, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, 1 to 300, 1 to 295, 1 to 290, 1 to 280 , 1 to 270, 1 to 260, 1 to 250, 1 to 240, 1 to 230, 1 to 220, 1 to 210, 1 to 200, 1 to 190, 1 to 180, 1 to 170, 1 to 160, 1 to 150, 1 to 140, 1 to 130, 1 to 120, 1 to 110, 1 to 100, 1 to 90, 1 to 80, 1 to 70, 1 to 60, 1 to 50, 1 to 40, 1 to 30, 1 to 20, 1 to 15, 1 to 10 or 1 to 5 amino acids. In some embodiments, the N-terminal fusion partner polypeptide is truncated from the C-terminus to retain the first 300, 290, 280, 270, 260, 250, 240, 230, 220, 210, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 150 to 40, first 150 to 50, first 150 to 75, first 150-100, first 100 to 40, first 100 to 50, first 100 to 75, first 75-40, first 75-50, first 300, first 250, first 200, top 150, top 140, top 130, top 120, top 110, top 100, top 90, top 80, top 75, top 70, top 65, top 60, top 55, top 50, or the first 40 amino acids.

[0106] In some embodiments, the truncated N-terminal fusion partner is Fk1B, FrnE, or EcpD1. 1 to 140, 1 to 130, 1 to 120, 1 to 110, 1 to 100, 1 to 90, 1 to 80, 1 to 70, 1 to 60, 1 to 50, 1 to 40, 1 to 30, 1 to 20, 1 to 15, 1 to 10, or 1 to 5 amino acids. In some embodiments, the truncated N-terminal fusion partner is EcpD, wherein EcpD is truncated from the C-terminus to remove 148, 198, 210, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, 1, 1 to 210, 1 to 200, 1 to 190, 1 to 180, 1 to 170, 1 to 160, 1 to 150, 1 to 140, 1 to 130, 1 to 120, 1 to 110, 1 to 100, 1 to 90, 1 to 80, 1 to 70, 1 to 60, 1 to 50, 1 to 40, 1 to 30, 1 to 20, 1 to 15, 1 to 10, or 1 to 5 amino acids. In some embodiments, the truncated N-terminal fusion partner is FrnE, wherein FrnE is truncated from the C-terminus to remove 118, 168, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, 1, 1 to 190, 1 to 180, 1 to 170, 1 to 160, 1 to 150, 1 to 140, 1 to 130, 1 to 120, 1 to 110, 1 to 100, 1 to 90, 1 to 80, 1 to 70, 1 to 60, 1 to 50, 1 to 40, 1 to 30, 1 to 20, 1 to 15, 1 to 10, or 1 to 5 amino acids.

[0107] In some embodiments, the N-terminal fusion partner is not β-galactosidase. In some embodiments, the N-terminal fusion partner is not thioredoxin. In some embodiments, the N-terminal fusion partner is neither β-galactosidase nor thioredoxin.

[0108] In some embodiments, the molecular weight of the N-terminal fusion partner is about 1 kDa, about 2 kDa, about 3 kDa, about 4 kDa, about 5 kDa, about 6 kDa, about 7 kDa, about 8 kDa, about 9 kDa, about 10 kDa, about 11 kDa, about 12 kDa, about 13 kDa, about 14 kDa, about 15 kDa, about 16 kDa, about 17 kDa, about 18 kDa, about 19 kDa, about 20 kDa, about 30 kDa, about 40 kDa, about 50 kDa, about 60 kDa, about 70 kDa, about 80 kDa, about 90 kDa, about 100 kDa, about 150 kDa, about 200 kDa, about 250 kDa, about 300 kDa, about 350 kDa, about 400 kDa, about 450 kDa, about 500 kDa or more. In some embodiments, the molecular weight of the N-terminal fusion partner is about 1 to about 10 kDA, about 1 to about 20 kDA, about 1 to about 30 kDA, about 1 to about 40 kDA, about 1 to about 50 kDA, about 1 to about 60 kDA, about 1 to about 70 kDA, about 1 to about 80 kDA, about 1 to about 90 kDA, about 1 to about 100 kDA, about 1 kDa to about 200 kDa, about 1 kDa to about 300 kDa, about 1 kDa to about 400 kDa, about 1 kDa to about 500 kDa, about 2 to about 10 kDA, about 2 to about 20 kDA, about 2 to about 30 kDA, about 2 to about 40 kDA, about 2 to about 50 kDA, about 2 to about 60 kDA, about 2 to about 70 kDA A, about 2 to about 80 kDA, about 2 to about 90 kDA, about 2 to about 100 kDA, about 2 kDa to about 200 kDa, about 2 kDa to about 300 kDa, about 2 kDa to about 400 kDa, about 2 kDa to about 500 kDa, about 3 to about 10 kDA, about 3 to about 20 kDA, about 3 to about 30 kDA, about 3 to about 40 kDA, about 3 to about 50 kDA, about 3 to about 60 kDA, about 3 to about 70 kDA, about 3 to about 80 kDA, about 3 to about 90 kDA, about 3 to about 100 kDA, about 3 kDa to about 200 kDa, about 3 kDa to about 300 kDa, about 3 kDa to about 400 kDa, or about 3 kDa to about 500 kDa.

[0109] In some embodiments, the N-terminal fusion partner or truncated N-terminal fusion partner is 25 or more amino acids long. In some embodiments, the N-terminal fusion partner is about 25 to about 2000 or more amino acids long. In some embodiments, the N-terminal fusion partner is about or at least about 25, 35, 40, 45, 50, 100, 150, 200, 250, 300, 350, 400, 450, 470, 500, 530, 560, 590, 610, 640, 670, 700, 750, 800, 850, 900, 950, 1000, 1200, 1400, 1600, 1800, 2000 amino acids long. In some embodiments, the polypeptide of interest is about 25 to about 2000, 25 to about 1000, 25 to about 500, 25 to about 250, 25 to about 100, or 25 to about 50 amino acids in length.

[0110] Relative sizes of target polypeptide and recombinant fusion protein

[0111] The output of target polypeptide is proportional to the output of complete recombinant fusion protein. This ratio depends on the relative size (for example, molecular weight and / or length in amino acids) of target polypeptide and recombinant fusion protein. For example, reducing the size of N-terminal fusion partner in fusion protein will result in a higher proportion of fusion protein produced being target polypeptide. In some embodiments, in order to maximize the output of target polypeptide, N-terminal fusion partner is selected based on its relative size with target polypeptide. In some embodiments, N-terminal fusion partner is selected to be a specific minimum size (for example, MW or length in amino acids) relative to target polypeptide. In some embodiments, recombinant fusion protein is designed to make the molecular weight of target polypeptide constitute about 10% to about 50% of the molecular weight of recombinant fusion protein. In some embodiments, the molecular weight of target polypeptide constitutes about or at least about: 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 35%, 40%, 45%, 50% of the molecular weight of recombinant fusion protein. In some embodiments, the molecular weight of the polypeptide of interest constitutes about or at least about: 10% to about 50%, 11% to about 50%, 12% to about 50%, 13% to about 50%, 14% to about 50%, 15% to about 50%, 20% to about 50%, 25% to about 50%, 30% to about 50%, 35% to about 50%, 40% to about 50%, 13% to about 40%, 14% to about 40%, 15% to about 40%, 20% to about 40%, 25% to about 40%, 30% to about 40%, 35% to about 40%, 13% to about 30%, 14% to about 30%, 15% to about 30%, 20% to about 30%, 25% to about 30%, 13% to about 25%, 14% to about 25%, 15% to about 25%, or 20% to about 25% of the molecular weight of the recombinant fusion protein. In some embodiments, the polypeptide of interest is hPTH and the molecular weight of the polypeptide of interest constitutes about 14.6% of the molecular weight of the recombinant fusion protein. In some embodiments, the target polypeptide is hPTH and the molecular weight of the target polypeptide constitutes about 13.6% of the molecular weight of the recombinant fusion protein. In some embodiments, the target polypeptide is hPTH and the molecular weight of the target polypeptide constitutes about 27.3% of the molecular weight of the recombinant fusion protein. In some embodiments, the target polypeptide is met-GCSF and the molecular weight of the target polypeptide constitutes about 39% to about 72% of the molecular weight of the recombinant fusion protein. In some embodiments, the target polypeptide is proinsulin and the molecular weight of the target polypeptide constitutes about 20% to about 57% of the molecular weight of the recombinant fusion protein.

[0112] In some embodiments, the length of the target polypeptide constitutes about 10% to about 50% of the total length of the recombinant fusion protein. In some embodiments, the length of the target polypeptide constitutes about or at least about: 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 35%, 40%, 45%, 50% of the total length of the recombinant fusion protein. In some embodiments, the length of the polypeptide of interest constitutes about or at least about: 10% to about 50%, 11% to about 50%, 12% to about 50%, 13% to about 50%, 14% to about 50%, 15% to about 50%, 20% to about 50%, 25% to about 50%, 30% to about 50%, 35% to about 50%, 40% to about 50%, 13% to about 40%, 14% to about 40%, 15% to about 40%, 20% to about 40%, 25% to about 40%, 30% to about 40%, 35% to about 40%, 13% to about 30%, 14% to about 30%, 15% to about 30%, 20% to about 30%, 25% to about 30%, 13% to about 25%, 14% to about 25%, 15% to about 25%, or 20% to about 25% of the total length of the recombinant fusion protein. In some embodiments, the polypeptide of interest is hPTH and the length of the polypeptide of interest constitutes about 13.1% of the total length of the recombinant fusion protein. In some embodiments, the polypeptide of interest is hPTH and the length of the polypeptide of interest constitutes about 12.5% ​​of the total length of the recombinant fusion protein. In some embodiments, the polypeptide of interest is hPTH and the length of the polypeptide of interest constitutes about 25.7% of the total length of the recombinant fusion protein. In some embodiments, the polypeptide of interest is met-GCSF and the length of the polypeptide of interest constitutes about 40% to about 72% of the total length of the recombinant fusion protein. In some embodiments, the polypeptide of interest is proinsulin and the length of the polypeptide of interest constitutes about 19% to about 56% of the total length of the recombinant fusion protein.

[0113] Differences in isoelectric points between the target peptide and the N-terminal fusion partner

[0114] The isoelectric point (pi) of a protein is defined as the pH at which the protein carries no net charge. The pi value is known to influence the solubility of a protein at a given pH. At a pH below its pi, a protein carries a net positive charge, and at a pH above its pi, a net negative charge. Proteins can be separated based on their isoelectric point (overall charge). In some embodiments, the pi of the target polypeptide and the pi of the N-terminal fusion protein are substantially different. This can facilitate purification of the target polypeptide from the N-terminal fusion protein. In some embodiments, the pi of the target polypeptide is at least two times higher than that of the N-terminal fusion partner. In some embodiments, the pi of the target polypeptide is 1.5 to 3 times higher than that of the N-terminal fusion partner. In some embodiments, the pi of the target polypeptide is 1.5, 1.6, 1.7, 1.8, 1.9, 2, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, or 3 times higher than that of the N-terminal fusion partner. In some embodiments, the pi of the N-terminal fusion partner is about 4, about 4.1, about 4.2, about 4.3, about 4.4, about 4.5, about 4.6, about 4.7, about 4.8, about 4.9, or about 5. In some embodiments, the pi of the N-terminal fusion partner is about 4 to about 5, about 4.1 to about 4.9, about 4.2 to about 4.8, about 4.3 to about 4.7, about 4.4 to about 4.6.

[0115] In some embodiments, the N-terminal fusion partner is one listed in Tables 8 or 18, having a pI listed therein. In some embodiments, the C-terminal polypeptide of interest is hPTH1-34, having a pI of 8.52 and a molecular weight of 4117.65 Daltons. In some embodiments, the C-terminal polypeptide of interest is Met-GCSF, having a pI of 5.66 and a molecular weight of 18801.9 Daltons. In some embodiments, the C-terminal polypeptide of interest is the proinsulin set forth in SEQ ID NO:88, having a pI of approximately 5.2 and a molecular weight of 9.34 KDa. In some embodiments, the C-terminal polypeptide of interest is the proinsulin set forth in SEQ ID NO:89, having a pI of approximately 6.07 and a molecular weight of approximately 8.81 KDa. In some embodiments, the C-terminal polypeptide of interest is the proinsulin set forth in SEQ ID NO:90, having a pI of approximately 5.52 and a molecular weight of approximately 8.75 KDa. In some embodiments, the C-terminal target polypeptide is proinsulin shown in SEQ ID NO: 91, having a pI of about 6.07 and a molecular weight of about 7.3 KDa. The pI of a protein can be determined according to any method described in the literature and known to those skilled in the art.

[0116] Chaperones and protein folding regulators

[0117] An obstacle to producing heterologous proteins in high yields in non-natural host cells (cells to which the heterologous protein is not natural) is that the cells are often not fully equipped to produce soluble and / or active forms of the heterologous protein. Although the primary structure of a protein is defined by its amino acid sequence, the secondary structure is defined by the presence of α-helices or β-sheets, and the tertiary structure is defined by the interactions of amino acid side chains within the protein (e.g., between protein domains). When expressing heterologous proteins, particularly in large-scale production, the secondary and tertiary structures of the protein itself are very important. Any significant changes in the protein structure can produce functionally inactive molecules, or proteins with significantly reduced biological activity. In many cases, host cells express chaperones or protein folding regulators (PFMs) that are required to properly produce active heterologous proteins. However, typically, under high-level expression conditions that require the production of available, economically satisfactory biotech products, cells often cannot produce enough one or more native protein folding regulators to process the heterologously expressed protein.

[0118] In some expression systems, the overproduction of heterologous proteins can be accompanied by their misfolding and sequestration into insoluble aggregates. In bacterial cells, these aggregates are called inclusion bodies. In some cases, the protein processed into inclusion bodies can be recovered by additional treatment of the insoluble fraction. The protein present in the inclusion bodies must usually be purified through multiple steps, including denaturation and renaturation. The typical renaturation method for inclusion body proteins includes attempting to dissolve the aggregates in concentrated denaturants and subsequently removing the denaturant by dilution. Aggregates are often formed again at this stage. Additional processing increases costs, cannot guarantee that the in vitro refolding produces a biologically active product, and the recovered protein may include a large amount of fragment impurities.

[0119] In vivo protein folding is assisted by molecular chaperones (which promote the correct isomerization and cell targeting of other polypeptides by the transient interaction with folding intermediates) and by foldases (which accelerate the rate limiting step along the folding pathway). In some cases, it has been found that the overexpression of chaperones has improved the soluble protein yield of aggregation-prone proteins (see Baneyx, F., 1999, Curr.Opin.Biotech.10:411-421). The beneficial effect related to the intracellular concentration of these chaperones seems to be highly dependent on the properties of the protein produced excessively, and may not be the overexpression of the same protein folding regulator required for all heterologous proteins. Protein folding regulator, including chaperones, disulfide isomerases and peptidyl-prolyl cis-trans isomerases (PPI enzymes), is a class of proteins present in all cells that help folding, unfolding and degradation of nascent polypeptides.

[0120] Chaperones are by binding nascent polypeptides, making them stable and making them fold correctly to work. Proteins have hydrophobic and hydrophilic residues simultaneously, the former usually exposed on the surface, and the latter embedded in the structure, they interact with other hydrophilic residues here rather than with the water interaction around the molecule. However, in the folded polypeptide chain, hydrophilic residues are often exposed for a period of time because the protein exists in a partially folded or misfolded state. It is during this time period that the polypeptide formed can become permanently misfolded or interact with other misfolded proteins and form large aggregates or inclusion bodies in the cell. Chaperones usually work by binding the hydrophobic region of the partially folded chain and preventing them from misfolding or aggregating with other proteins completely. Chaperones can even bind to the protein in the inclusion body and disaggregate it. The GroES / EL, DnaKJ, Clp, Hsp90 and SecB families of folding regulators are all examples of proteins with chaperone-like activity.

[0121] Disulfide isomerases are another important type of folding regulator. These proteins catalyze a very specific set of reactions that help folded polypeptides form the correct intraprotein disulfide bonds. Any protein with more than two cysteines is at risk for forming disulfide bonds between the incorrect residues. The disulfide-forming family consists of the Dsb proteins, which catalyze disulfide bond formation in the non-reducing environment of the periplasm. When periplasmic polypeptides are misfolded, the disulfide isomerase, DsbC, is able to rearrange the disulfide bonds and allow the protein to reform with the correct linkages.

[0122] The Fk1B and FrnE proteins belong to the peptidyl-prolyl cis-trans isomerase family of folding regulators. This class of enzymes catalyzes the cis-trans isomerization of the proline imine peptide bond in oligopeptides. Proline is unique among amino acids in that the peptide bond immediately preceding it can adopt either a cis or trans conformation. For all other amino acids, this is unfavorable due to steric hindrance. Peptidyl-prolyl cis-trans isomerases (PPI enzymes) catalyze the conversion of this bond from one form to the other. This isomerization can accelerate and / or aid protein folding, refolding, subunit assembly, and transport within the cell.

[0123] In addition to the universal chaperone protein that seems to interact with proteins in a non-specific manner, there is also a chaperone protein that helps specific targets to fold. These protein-specific chaperone proteins form complexes with their targets, thereby preventing aggregation and degradation and leaving the time for them to be assembled into multi-subunit structures. PapD chaperone protein is an example (described in Lombardo et al., 1997, in Escherichia coli PapD, see Guidebook to Molecular Chaperones and Protein-Folding Catalyst, Gething MJ edits, Oxford University Press Inc., New York: 463-465, incorporated to text by reference).

[0124] Folding regulators include, for example, HSP70 proteins, HSP110 / SSE proteins, HSP40 (DnaJ-related) proteins, GRPE-like proteins, HSP90 proteins, CPN60 and CPN10 proteins, cytosolic chaperones, HSP100 proteins, small HSPs, calnexin and calreticulin, PDI and thioredoxin-related proteins, peptidyl-prolyl isomerases, cyclophilin PPI enzymes, FK-506 binding proteins, parvulin PPI enzymes, individual chaperones, protein-specific chaperones or intramolecular chaperones. Folding regulators are generally described in "Guidebook to Molecular Chaperones and Protein-Folding Catalysts," 1997, ed. M. Gething, University of Melbourne, Australia, incorporated herein by reference.

[0125] The best characterized molecular chaperones in the E. coli cytoplasm are the ATP-dependent DnaK-DnaJ-GrpE and GroEL-GroES systems. In E. coli, the network of folding regulators / chaperones includes the Hsp70 family. The main Hsp70 chaperone, DnaK, effectively prevents protein aggregation and supports the refolding of damaged proteins. Heat shock proteins are introduced into protein aggregates and can help disaggregation. Based on in vitro studies and homology considerations, a variety of other cytoplasmic proteins that play a role as molecular chaperones have been proposed in E. coli. These include C1pB, HtpG and IbpA / B, which, like DnaK-DnaJ-GrpE and GroEL-GroES, are heat shock proteins (Hsp) belonging to stress regulators.

[0126] Pseudomonas fluorescens DnaJ-like protein is a molecular chaperone protein belonging to the DnaJ / Hsp40 protein family, characterized in that they have a highly conserved J-domain. The J-domain is a 70-amino acid region located at the C-terminal end of the DnaJ protein. The N-terminal end has a transmembrane (TM) domain that facilitates insertion into the membrane. The A-domain separates the TM domain from the J-domain. By interacting with another chaperone protein, DnaK (as a common chaperone protein), the protein in the DnaJ family plays a key role in protein folding. The highly conserved J-domain is the interaction site between the DnaJ protein and the DnaK protein. It is believed that type I DnaJ protein is a true DnaJ protein, while types II and III are commonly referred to as DnaJ-like proteins. It is also known that DnaJ-like proteins actively participate in the response to hyperosmotic and heat shock by preventing the aggregation of stress-denatured proteins and by making protein depolymerization in two ways, DnaK-dependent and non-DnaK-dependent.

[0127] The trans conformation of the X-Pro bond is energetically favorable in a nascent protein chain; however, approximately 5% of all prolyl peptide bonds are found in the cis conformation in native proteins. The cis-trans isomerization of the X-Pro bond is rate-limiting in the folding of many peptides and is catalyzed in vivo by peptidylprolyl cis / trans isomerases (PPI enzymes). To date, three cytoplasmic PPI enzymes, SlyD, SlpA, and initiation factor (TF), have been identified in Escherichia coli. TF, a 48 kDa protein associated with the 50S ribosomal subunit, has been postulated to cooperate with chaperones in Escherichia coli to ensure the correct folding of newly synthesized proteins. At least five proteins (thioredoxins 1 and 2, and glutaredoxins 1, 2, and 3, products of the trxA, trxC, grxA, grxB, and grxC genes, respectively) are involved in the reduction of disulfide bridges that appear briefly in cytoplasmic enzymes. Therefore, the N-terminal fusion partner can be a disulfide bond-forming protein or chaperone that allows correct disulfide bond formation.

[0128] Examples of folding regulators useful in the methods of the present invention are shown in Table 1. RXF numbering refers to open reading frames. U.S. Patent Application Publication Nos. 200 / 0269070 and 2010 / 0137162 (both invention names are "Method for Rapidly Screening Microbial Hosts to Identify Certain Strains with Improved Yield and / or Quality in the Expression of Heterologous Proteins," which are incorporated herein by reference in their entirety) disclose open reading frame sequences for the proteins listed in Table 1. Proteases and folding regulators are also provided in Tables A to F of U.S. Patent No. 8,603,824, "Process for improved protein expression by strain engineering," which are incorporated herein by reference in their entirety.

[0129] Table 1. Pseudomonas fluorescens folding regulators

[0130]

[0131]

[0132]

[0133]

[0134]

[0135]

[0136] connector

[0137] The recombinant fusion protein of the present invention contains a joint between the N-terminal fusion partner and the C-terminal target polypeptide. In some embodiments, the joint comprises a cleavage site identified by a lytic enzyme (i.e., a proteolytic enzyme that internally cuts proteins). In some embodiments, the cutting of the joint at the cleavage site separates the target polypeptide from the N-terminal fusion partner. The proteolytic enzyme can be any protease known in the art or described in the literature, for example, at PCT Publication No. WO 2003 / 010204, "Process for Preparing Polypeptides of Interest from Fusion Polypeptides," U.S. Patent No. 5,750,374, "Process for Producing Hydrophobic Polypeptides and Proteins, and Fusion Proteins for Use in Producing Same," and U.S. Patent No. 5,935,824, each of which is incorporated herein by reference in its entirety.

[0138] In some embodiments, the linker comprises a cleavage site for cleavage by, for example, a serine protease, a threonine protease, a cysteine ​​protease, an aspartic protease, a glutamic protease, a metalloprotease, an asparaginase, a mixed protease, or a protease of unknown catalytic type. In some embodiments, serine proteases are, for example, trypsin, chymotrypsin, endoproteinase Arg-C, endoproteinase Glu-C, endoproteinase Lys-C, elastase, proteinase K, subtilisin, carboxypeptidase P, carboxypeptidase Y, or acylamino acid releasing enzymes. In some embodiments, metalloproteases are, for example, endoproteinase Asp-N, thermolysin, carboxypeptidase A, or carboxypeptidase B. In some embodiments, cysteine ​​proteases are, for example, papain, clostripain, cathepsin C, or pyroglutamate aminopeptidase. In some embodiments, aspartic proteases are, for example, pepsin, chymase, or cathepsin D. In some embodiments, glutamic proteases are, for example, scytalidoglutamic peptidase. In some embodiments, the asparaginyl protease is, for example, nodavirus peptidase, intein-containing chloroplast ATP-dependent peptidase, intein-containing replicative DNA helicase precursor, or reovirus type 1 coat protein. In some embodiments, the protease of unknown catalyst type is, for example, collagenase, protein P5 murein endopeptidase, homopolypeptidase, microcin-processing peptidase 1, or Dop isopeptidase.

[0139] In some embodiments, the linker comprises a cleavage site for a colorless peptidase, an aminopeptidase, ancrod, angiotensin converting enzyme, bromelain, calpain, calpain I, calpain II, carboxypeptidase A, carboxypeptidase B, carboxypeptidase G, carboxypeptidase P, carboxypeptidase W, carboxypeptidase Y, caspase (universal), caspase 1, caspase 2, caspase 3, caspase 4, caspase 5, caspase 6, caspase 7, caspase 8, caspase 9, caspase 10, caspase 11, caspase 12, caspase 13, cathepsin B, cathepsin C, cathepsin D, cathepsin E, cathepsin G, cathepsin H, cathepsin L, chymopapain, chymase, chymotrypsin , α-clostripain, collagenase, complement C1r, complement C1s, complement factor D, complement factor I, cucumis sativus, dipeptidyl peptidase IV, elastase, leukocyte, elastase, endoproteinase Arg-C, endoproteinase Asp-N, endoproteinase Glu-C, endoproteinase Lys-C, enterokinase, factor Xa, ficin, furin, granzyme A, granzyme B, HIV protease, IG enzyme, tissue kallikrein, leucine aminopeptidase (universal), leucine aminopeptidase, cytosolic, leucine aminopeptidase, microsomal, matrix metalloproteinase, methionine aminopeptidase, neutral protease, papain, pepsin, plasmin, prolyl amino acid dipeptidase, pronase E, prostate-specific antigen, basophilic protease from Streptomyces griseus, protease from Aspergillus zoellularis, protease from Aspergillus zoellularis Saitoi), a protease from Aspergillus sojae, a protease (Bacillus licheniformis) (alkaline), a protease (Bacillus licheniformis) (alkaline protease (Alcalase)), a protease from Bacillus polymyxa (Bacillus polymyxa), a protease from Bacillus (Esperase), a protease from Rhizopus, a protease S, a proteasome, a protease from Aspergillus oryzae, a protease 3, a protease A, a protease K, a protein C, a pyroglutamate aminopeptidase, renin, rennin, streptokinase, subtilisin, thermolysin, thrombin, tissue plasminogen activator, trypsin, trypsin-like enzyme or urokinase. In some embodiments, the linker comprises a cleavage site recognized by enterokinase, factor Xa or furin. In some embodiments, the linker comprises a cleavage site recognized by enterokinase or trypsin. In some embodiments, the linker comprises a cleavage site recognized by bovine enterokinase.These and other proteases useful in the methods of the invention, and their cleavage recognition sites, are known in the art and described in the literature, for example, Harlow and Lane, ANTIBODIES: A LABORATORY MANUAL, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (1988); Walsh, PROTEINS: BIOCHEMISTRY AND BIOTECHNOLOGY, John Wiley & Sons, Ltd., West Sussex, England (2002), which are incorporated herein by reference.

[0140] In some embodiments, the linker comprises an affinity tag. Affinity tags are peptide sequences that can aid in protein purification. Affinity tags are fused to proteins to facilitate purification of proteins from crude biological sources using affinity techniques. Any suitable affinity tag known in the art can be used as desired. In some embodiments, the affinity tag used in the present invention is, for example, a chitin-binding protein, a maltose-binding protein or a glutathione-S-transferase protein, polyhistidine, a FLAG tag (SEQ ID NO: 229), a calmodulin tag (SEQ ID NO: 230), a Myc tag, a BP tag, an HA-tag (SEQ ID NO: 231), an E-tag (SEQ ID NO: 232), an S-tag (SEQ ID NO: 233), an SBP tag (SEQ ID NO: 234), Softag 1, Softag 3 (SEQ ID NO: 235), a V5 tag (SEQ ID NO: 236), an Xpress tag, green fluorescent protein, a Nus tag, a Strep tag, a thioredoxin tag, an MBP tag, a VSV tag (SEQ ID NO: 237), or an Avi tag.

[0141] Affinity tags can be removed by chemical agents or by enzymatic means, such as proteolysis. Methods for using affinity tags in protein purification are described in the literature, for example, Lichty et al., 2005, "Comparison of affinity tags for protein purification", Protein Expression and Purification 41:98-105. Other affinity tags useful in the linkers of the present invention are known in the art and described in the literature, for example, the above-mentioned U.S. Patent No. 5,750,374 and Terpe K., 2003, "Overview of Tag Protein Fusions: from molecular and biochemical fundamentals to commercial systems", Applied Microbiology and Biotechnology (60):523-533, both of which are incorporated by reference in their entirety.

[0142] In some embodiments, the linker is 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more amino acids in length. In some embodiments, the linker is 4 to 50, 4 to 45, 4 to 40, 4 to 35, 4 to 30, 4 to 25, 4 to 20, 4 to 15, 4 to 10, 5 to 50, 5 to 45, 5 to 40, 5 to 35, 5 to 30, 5 to 25, 5 to 20, 5 to 15, 5 to 10, 10 to 50, 10 to 45, 10 to 40, 10 to 35, 10 to 30, 10 to 25, 10 to 20, 10 to 15, 15 to 50, 15 to 45, 15 to 40, 15 to 35, 15 to 30, 15 to 25, 15 to 20, 20 to 50, 20 to 45, 20 to 40, 20 to 35, 20 to 30, or 20 to 25 amino acids in length. In some embodiments, the linker is 18 amino acids in length. In some embodiments, the linker is 19 amino acids in length.

[0143] In some embodiments, the linker includes multiple glycine residues. In some embodiments, the linker includes 1, 2, 3, 4, 5, 6, 7, 8 or more glycine residues. In some embodiments, the linker includes 1-8, 1-7, 1-6, 1-5 or 1-4 glycine residues. In some embodiments, the glycine residues are continuous. In some embodiments, the linker contains at least one serine residue. In some embodiments, the glycine and / or serine residues constitute a spacer. In some embodiments, the spacer is a (G4S)2 spacer with 10 amino acids, as shown in SEQ ID NO:59. In some embodiments, the spacer is a (G4S)1, (G4S)2, (G4S)3, (G4S)4 or (G4S)5 spacer. In some embodiments, the linker contains six histidine residues, or a His-tag. In some embodiments, the linker includes an enterokinase cleavage site, for example, as shown in SEQ ID NO:13 (DDDDK). In some embodiments, the recombinant fusion protein comprises a linker as set forth in any one of SEQ ID NOs: 9 to 12 or 226 listed in Table 2. The enterokinase cleavage site in SEQ ID NO: 9 is underlined. The polyhistidine affinity tag is in italics in each of SEQ ID NOs: 9 to 12 and 226. In some embodiments, the recombinant fusion protein comprises a linker corresponding to SEQ ID NO: 9.

[0144] Table 2: Linker sequences

[0145] SEQ ID NO: Amino acid sequence 9 <![CDATA[GGGGSGGGGHHHHHH DDDDK ]]> 10 GGGGSGGGGHHHHHHRKR 11 GGGGSGGGGHHHHHHRRR 12 GGGGSGGGGHHHHHHLVPR 226 GGGGSGGGGSHHHHHHR

[0146] expression vector

[0147] In some embodiments, a gene segment encoding a recombinant fusion protein is introduced into a suitable expression plasmid to produce an expression vector for expressing the recombinant fusion protein. The expression vector may be, for example, a plasmid. In some embodiments, the plasmid encoding the recombinant fusion protein sequence may contain a selectable marker and allow the host cell maintaining the plasmid to grow under selective conditions. In some embodiments, the plasmid does not contain a selectable marker. In some embodiments, the expression vector may be integrated into the host cell genome. In some embodiments, the expression vector encodes hPTH 1-34 fused to a linker and a protein that directs the expressed fusion protein to the cytoplasm. In some embodiments, the expression vector encodes hPTH 1-34 fused to a linker and a protein that directs the expressed fusion protein to the periplasm. In some embodiments, the expression vector encodes hPTH 1-34 fused to a linker and a Pseudomonas fluorescens DnaJ-like protein. In some embodiments, the expression vector encodes hPTH 1-34 fused to a linker and a Pseudomonas fluorescens Fk1B protein.

[0148] Examples of nucleotide sequences encoding PTH 1-34 fusion proteins are provided in the sequence listing herein. An example of a nucleotide sequence encoding a fusion protein comprising a DnaJ-like protein N-terminal fusion partner is designated Gene ID 126203 (SEQ ID NO: 52), corresponding to a coding sequence optimized for Pseudomonas fluorescens. The sequence designated Gene ID 126206 (SEQ ID NO: 53) corresponds to the native Pseudomonas fluorescens DnaJ coding sequence fused to an optimized linker and the PTH 1-34 coding sequence. Gene sequences 126203 and 126206 are those present in expression plasmids p708-001 and p708-004, respectively. An example of a nucleotide sequence encoding a fusion protein comprising a Fk1B N-terminal fusion partner is designated Gene ID 126204 (SEQ ID NO: 54), corresponding to a coding sequence optimized for Pseudomonas fluorescens. Gene ID 126207 (SEQ ID NO: 55) corresponds to the natural fluorescing Pseudomonas fluorescens Fk1B coding sequence fused to the optimized joint and PTH 1-34 coding sequence. Gene sequences 126204 and 126207 are those present in expression plasmids p708-002 and p708-005, respectively. An example of a nucleotide sequence encoding a fusion protein comprising the FrnE N-terminal fusion partner is the gene ID 126205 (SEQ ID NO: 56) specified, corresponding to the coding sequence optimized for Pseudomonas fluorescens. The sequence of the specified gene ID 126208 (SEQ ID NO: 57) corresponds to the natural fluorescing Pseudomonas fluorescens FrnE coding sequence fused to the optimized joint and PTH 1-34 coding sequence. Gene sequences 126205 and 126208 are those present in expression plasmids p708-003 and p708-006, respectively.

[0149] Codon optimization

[0150] The present invention contemplates the use of any suitable coding sequence for the fusion protein and / or each of its individual components, including any sequence that has been optimized for expression in the host cell used. Methods for optimizing codons to improve expression in bacterial hosts are known in the art and described in the literature. For example, codon optimization for expression in Pseudomonas host strains is described in, for example, U.S. Patent Application Publication No. 2007 / 0292918, "Codon Optimization Method," which is incorporated by reference in its entirety. Codon optimization for expression in Escherichia coli is described in, for example, Welch et al., 2009, PLoS One, "Design Parameters to Control Synthetic Gene Expression in Escherichia coli, 4(9):e7002, which is incorporated by reference in its entirety. Non-limiting examples of coding sequences for fusion protein components are provided herein, however, it will be appreciated that any suitable sequence can be produced as desired according to methods well known to those skilled in the art.

[0151] Expression system

[0152] Based on the teachings herein, those skilled in the art can determine that for producing a suitable bacterial expression system useful for the method according to the present invention, the target polypeptide is expressed in a manner that is suitable for the production of the target polypeptide. In some embodiments, an expression construct comprising a nucleotide sequence encoding a recombinant fusion protein containing the target polypeptide is provided as part of an inducible expression vector. In some embodiments, a host cell transformed with an expression vector is cultivated, and induction is used to express the fusion protein from the expression vector. The expression vector can, for example, be a plasmid. In some embodiments, the expression vector is a plasmid encoding a recombinant fusion protein encoding sequence further comprising a selective marker, and the host cell is grown under selective conditions that allow the plasmid to maintain. In some embodiments, the expression construct is integrated into the host cell genome. In some embodiments, the expression construct encodes a recombinant fusion protein fused to a secretion signal, and the secretion signal can direct the recombinant fusion protein to the periplasm.

[0153] Methods for expressing heterologous proteins, including useful regulatory sequences (e.g., promoters, secretion leaders, and ribosome binding sites), in host cells useful in the methods of the invention, including Pseudomonas host cells, are described, for example, in U.S. Patent Application Publication Nos. 2008 / 0269070 and 2010 / 0137162, U.S. Patent Application Publication No. 2006 / 0040352, "Expression of Mammalian Proteins in Pseudomonas fluorescens," and U.S. Patent No. 8,603,824, each of which is incorporated by reference in its entirety. These publications also describe bacterial host strains useful in practicing the methods of the invention that have been engineered to overexpress folding regulators or into which protease mutations have been introduced, for example, to eliminate, inactivate, or reduce the activity of a protease, thereby increasing heterologous protein expression. Leader sequences are described in detail in U.S. Patent No. 7,618,799, “Bacterial leader sequences for increased expression,” and U.S. Patent No. 7,985,564, “Expression systems with Sec-system Secretion,” both of which are incorporated herein by reference in their entireties, and in previously referenced U.S. Patent Application Publication No. 2010 / 0137162.

[0154] The promoter used according to the present invention can be a constitutive promoter or a regulatable promoter. The example of an inducible promoter includes those (that is, lacZ promoter) derived from the lac promoter family, for example, U.S. Patent No. 4,551,433, tac and trc promoters described in " Microbial Hybrid Promoters ", incorporated by reference, and Ptac16, Ptac17, PtacII, PlacUV5 and T7lac promoters. In some embodiments, the promoter is not derived from a host cell organism. In some embodiments, the promoter is derived from an Escherichia coli organism. In some embodiments, the lac promoter is used to regulate the expression of recombinant fusion proteins from a plasmid. In the case of lac promoter derivatives or family members, for example, tac promoter, inducing agent is IPTG (isopropyl-β-D-1-thiogalactopyranoside, "isopropylthiogalactoside"). In some embodiments, IPTG is added to the host cell culture to induce expression of the recombinant fusion protein from the lac promoter in the Pseudomonas host cells according to methods known in the art and described in the literature (eg, US Patent Publication No. 2006 / 0040352).

[0155] Examples of non-lac promoters useful in the expression system according to the present invention include P R (induced by high temperature), P L (induced by high temperature), P m (induced by alkyl- or halogen-benzoate), P u (by alkyl- or halogen-toluene induction) or P sal (induced by salicylate), as described, for example, in J. Sanchez-Romero & V. De Lorenzo (1999) Manual of Industrial Microbiology and Biotechnology (A. Demain & J. Davies, eds.) pp. 460-74 (ASM Press, Washington, DC); H. Schweizer (2001) Current Opinion in Biotechnology, 12: 439-445; and R. Slater & R. Williams (2000 Molecular Biology and Biotechnology (J. Walker & R. Rapley, eds.) pp. 125-54 (The Royal Society of Chemistry, Cambridge, UK). Promoters having nucleotide sequences of promoters native to selected bacterial host cells can also be used to control expression of expression constructs encoding target polypeptides, for example, Pseudomonas anthranilate or benzoate operon promoters (Pant, Pben). Tandem promoters can also be used, in which more than one promoter is covalently linked to another, whether the sequences are identical or different (e.g., Pant-Pben tandem promoter (promoter hybrid) or Plac-Plac tandem promoter), derived from the same or different organisms. In some embodiments, the promoter is Pmtl, as described, for example, in U.S. Patent Nos. 7,476,532 and 8,017,355, both of which are incorporated by reference in their entirety.

[0156] Regulated (inducible) promoters utilize promoter regulatory proteins to control the transcription of the gene of which the promoter is a part. Where a regulatable promoter is used herein, the corresponding promoter regulatory protein will also be part of the expression system according to the present invention. Examples of promoter regulatory proteins include: activator proteins, e.g., the E. coli catabolite activator protein, MalT protein; AraC family transcriptional activators; repressor proteins, e.g., the E. coli Lad protein; and dual-function regulatory proteins, e.g., the E. coli NagC protein. Many regulatable promoter / promoter-regulatory protein pairs are known in the art.

[0157] The promoter regulatory protein interacts with the effector compound, that is, a compound that is reversibly or irreversibly bound to the regulatory protein so that the protein can release or bind to at least one DNA transcriptional regulatory region of the gene under the control of the promoter, thereby allowing or blocking the effect of the transcriptase initiating gene transcription. Effector compounds are classified as inducers or co-repressors, and these compounds include natural effector compounds and placebo inducer compounds. Many regulated promoter / promoter-regulatory protein / effector compound triplets are known in the art. Although effector compounds can be used throughout the entire process of cell culture or fermentation, in a preferred embodiment using a regulated promoter, after the desired amount or density of host cell biomass is grown, a suitable effector compound is added to the culture to directly or indirectly cause the expression of the desired gene encoding the target protein or polypeptide.

[0158] In embodiments where a lac family promoter is utilized, the lacI gene may also be present in the system. The lacI gene, which is typically a constitutively expressed gene, encodes the Lac repressor protein, LacI protein, which binds to the lac operator of the lac family promoter. Therefore, in instances where a lac family promoter is utilized, the lac gene may also be included and expressed in the expression system.

[0159] Other regulatory elements

[0160] In some embodiments, other regulatory elements are present in the expression construct encoding the recombinant fusion protein. In some embodiments, the soluble recombinant fusion protein is present in the cytoplasm or periplasm of the cell during production. Secretion leader sequences for targeting fusion proteins are described elsewhere herein. In some embodiments, the expression construct encoding of the present invention is fused to a secretion leader sequence that can transport the recombinant fusion protein to the cytoplasm of Pseudomonas cells. In some embodiments, the expression construct encoding is fused to a secretion leader sequence that can transport the recombinant fusion protein to the periplasm of Pseudomonas cells. In some embodiments, the secretion leader sequence is excised from the recombinant fusion protein.

[0161] Other elements include, but are not limited to, transcriptional enhancer sequences, translational enhancer sequences, other promoters, activators, translation start and stop signals, transcriptional terminators, cistron regulators, polycistronic regulators, tag sequences, such as nucleotide sequence "tags" and "tag" polypeptide coding sequences, which facilitate identification, isolation, purification and / or separation of expressed polypeptides, as previously described. In some embodiments, in addition to the protein coding sequence, the expression construct includes any of the following regulatory elements operably linked thereto: a promoter, a ribosome binding site (RBS), a transcription terminator, and translation start and stop signals. Useful RBSs can be obtained from any species that can be used as a host cell in an expression system, as previously mentioned, for example, in U.S. Patent Application Publication Nos. 2008 / 0269070 and 2010 / 0137162. Many specific and multiple consensus RBSs are known, for example, those described and mentioned in D. Frishman et al., Gene 234(2):257-65 (July 8, 1999); and BESuzek et al., Bioinformatics 17(12):1123-30 (December 2011), which are incorporated by reference. In addition, natural or synthetic RBSs can be used, for example, those described in EP 0207459 (synthetic RBS); O. Ikehata et al., Eur. J. Biochem. 181(3):563-70 (1989). In some embodiments, the "Hi" ribosome binding site, aggagg, (SEQ ID NO:60) is used in the construct. Optimization of the ribosome binding site, including the spacing between the RBS and the translation initiation codon, is described in the literature, for example, Chen et al., 1994, "Determination of the optimal aligned spacing between the Shine-Dalgarno sequence and the translation initiation codon of Escherichia coli mRNAs", Nucleic Acids Research 22(23):4953-4957 and Ma et al., 2002, "Correlations between Shine-Dalgarno Sequences and Gene Features Such as Predicted Expression Levels and Operon Structures", J. Bact. 184(20):5733-45, incorporated by reference.

[0162] Other examples of methods, vectors, and translation and transcription elements, and other elements, useful in the present invention are well known in the art and are described, for example, in: U.S. Pat. No. 5,055,294 to Gilroy et al. and U.S. Pat. No. 5,128,130 to Gilroy et al.; U.S. Pat. No. 5,281,532 to Rammler et al.; U.S. Pat. Nos. 4,695,455 and 4,861,595 to Barnes et al.; U.S. Pat. No. 4,755,465 to Gray et al.; and U.S. Pat. No. 5,169,760 to Wilcox et al., all of which are incorporated by reference, as well as in numerous other publications incorporated by reference.

[0163] secretory leader sequence

[0164] In some embodiments, a secretion signal or leader sequence encoding sequence is fused to the N-terminus of the sequence encoding the recombinant fusion protein. The use of a secretion signal sequence can improve the production of recombinant proteins in bacteria. In addition, many types of proteins require secondary modifications that cannot be effectively achieved using known methods. The use of a secretion leader sequence can improve the harvest of correctly folded proteins by secreting proteins from the intracellular environment. In Gram-negative bacteria, proteins secreted from the cytoplasm can end up in the periplasmic space, connecting to the outer membrane or in the extracellular culture fluid. These methods also avoid the formation of inclusion bodies. Protein secretion into the periplasmic space also has the effect of promoting the formation of correct disulfide bonds (Bardwell et al., 1994, Phosphate Microorg, Chapter 45, 270-5 and Manoil, 2000, Methods in Enzymol. 326: 35-47). Other benefits of recombinant protein secretion include more efficient protein isolation, correct protein folding and disulfide bond formation, resulting in increased yield, as measured by, for example, the percentage of protein in active form, reduced inclusion body formation and reduced toxicity to host cells, and an increased percentage of recombinant protein in soluble form. The potential to excrete the target protein into the culture medium can also potentially facilitate continuous culture for protein production rather than batch culture.

[0165] In some embodiments, the recombinant fusion protein or target polypeptide is targeted to the periplasm of the host cell or to the extracellular space. In some embodiments, the expression vector further comprises a nucleotide sequence encoding a secretion signal polypeptide, which is operably linked to the nucleotide sequence encoding the recombinant fusion protein or target polypeptide.

[0166] Thus, in one embodiment, a recombinant fusion protein comprises a secretion signal, an N-terminal fusion partner, a linker, and a polypeptide of interest, wherein the secretion signal is at the N-terminus of the fusion partner. When the protein is targeted to the periplasm, the secretion signal can be cleaved from the fusion recombinant protein. In some embodiments, the linkage between the secretion signal and the protein or polypeptide is modified to enhance cleavage of the secretion signal from the fusion protein.

[0167] Host cells and strains

[0168] Bacterial host cells, including Pseudomonas (i.e., host cells in the order Pseudomonadales) and closely related bacterial organisms, are contemplated for use in practicing the methods of the present invention. In certain embodiments, the Pseudomonas host cell is Pseudomonas fluorescens. The host cell may also be Escherichia coli.

[0169] Host cells and constructs useful in the practice of the methods of the present invention can be identified or prepared using reagents and methods known in the art and described in the literature (e.g., U.S. Patent No. 8,288,127, "Protein Expression Systems," incorporated by reference in its entirety). This patent describes the production of recombinant polypeptides by introducing a nucleic acid construct into an auxotrophic Pseudomonas fluorescens host cell containing an insert of a chromosomal lacI gene. The nucleic acid construct comprises a nucleotide sequence encoding a recombinant polypeptide operably linked to a promoter capable of directing expression of the nucleic acid in the host cell, and further comprises a nucleotide sequence encoding an auxotrophic selection marker. The auxotrophic selection marker is a polypeptide that restores the original nutrition to an auxotrophic host cell. In some embodiments, the cell is auxotrophic for proline, uracil, or a combination thereof. In some embodiments, the host cell is derived from MB101 (ATCC deposit PTA-7841). U.S. Patent No. 8,288,127, “Protein Expression Systems,” and Schneider et al., 2005, “Auxotrophic markers pyrF and proC can replace antibiotic markers on protein production plasmids in high-cell-density Pseudomonas fluorescens fermentation,” Biotechnol. Progress 21(2):343-8, both of which are incorporated by reference in their entirety, describe a production host strain constructed for uracil auxotrophy by deleting the pyrF gene in strain MB101. The pyrF gene was cloned from strain MB214 (ATCC deposit PTA-7840) to generate a plasmid that can compensate for the pyrF deletion and restore auxotrophy. In a specific embodiment, a dual pyrF-proC dual auxotrophic selection marker system in Pseudomonas fluorescens host cells is used. In view of the published literature, one skilled in the art can generate the PyrF production host strain according to standard recombinant methods and use it as a background for introducing other desired genomic changes, including those described herein that can be used to practice the methods of the present invention.

[0170] In some embodiments, the host cell is a Pseudomonas order (referred to herein as "Pseudomonas"). In the case where the host cell is a Pseudomonas order, it can be a member of the Pseudomonas family, including Pseudomonas. γ Proteobacteria (Proteobacteria) hosts include members of Escherichia coli species and members of Pseudomonas fluorescens species. Other Pseudomonas organisms can also be useful. Pseudomonas and closely related species include Gram-negative Proteobacteria subgroup 1, which includes belonging to RE Buchanan and NE Gibbons (editor), Bergey's Manual of Determinative Bacteriology, pp.217-289 (8th edition, 1974) (The Williams & Wilkins Co., Baltimore, Md., USA) (all incorporated by reference) described as "Gram-negative aerobic rods and cocci" and / or the Proteobacteria group of the genus. (i.e., host cells of the Pseudomonas order). Table 3 presents the organisms of these sections and genera.

[0171] Table 3: Families and genera ("Gram-negative aerobic rods and cocci", "Bergey', 1974)

[0172]

[0173] Pseudomonas and closely related bacteria are generally defined as "Gram (-) Proteobacteria subgroup 1" or "Gram-negative aerobic rods and cocci" (Buchanan and Gibbons (eds.) (1974) Bergey's Manual of Determinative Bacteriology, pp. 217-289). Pseudomonas host strains are described in the literature, for example, in U.S. Patent Application Publication No. 2006 / 0040352, which is incorporated by reference in its entirety.

[0174] "Gram-negative Proteobacteria subgroup 1" also includes Proteobacteria that are included in this section according to the criteria used in the classification. This section also includes groups that were previously included in this section but are no longer included therein, such as Acidovorax, Brevundimonas, Burkholderia, Hydrogenophaga, Oceanimonas, Ralstonia and Stenotrophomona, Sphingomonas (and the genus Blastomonas derived therefrom), which was formed by regrouping organisms belonging to the genus Xanthomonas (and formerly known as Xanthomonas species), and Acidomonas, which was formed by regrouping organisms belonging to the genus Acetobacter, as defined in Bergey (1974). Additionally, the host can include cells from the genus Pseudomonas, Pseudomonas enalia (ATCC 14393), Pseudomonas nigrifaciensi (ATCC 19375), and Pseudomonas putrefaciens (ATCC 8071) (which have been reclassified as Alteromonas haloplanktis, Alteromonas nigrifaciens, and Alteromonas putrefaciens, respectively). Similarly, for example, Pseudomonas acidovorans (ATCC 15668) and Pseudomonas testosteroni (ATCC 11996) have been reclassified as Comamonas acidovorans and Comamonas testosteroni, respectively; and Pseudomonas nigrifaciens (ATCC 19375) and Pseudomonas piscicida (ATCC 15057) have been reclassified as Pseudoalteromonas nigrifaciens and Pseudoalteromonas piscicida, respectively."Gram-negative Proteobacteria subgroup 1" also includes Proteobacteria classified as belonging to any of the following families: Pseudomonadaceae, Azotobacteraceae (now often referred to as the "Azotobacter group" with the synonym Pseudomonadaceae), Rhizobacteriaceae, and Methylomonadaceae (now often referred to as the synonym "Methylococcaceae"). Thus, in addition to those genera otherwise described herein, further Proteobacterial genera falling within "Gram-negative Proteobacteria Subgroup 1" include: 1) Azotobacter group bacteria of the genus Azorhizophilus; 2) Pseudomonadaceae bacteria of the genera Cellvibrio, Oligella, and Teredinibacter; 3) Rhizobiaceae bacteria of the genera Chelatobacter, Ensifer, Liberibacter (also known as "Candidatus Liberibacter"), and Sinorhizobium; and 4) Methylobacter, Methylocaldum, Methylomicrobium, Methylosarcina, and Methylosphaera.

[0175] The host cell can be selected from "Gram-negative Proteobacteria subgroup 16." "Gram-negative Proteobacteria subgroup 16" is defined as the group of Proteobacteria of the following Pseudomonas species (with ATCC or other depository numbers of exemplary strains shown in parentheses): Pseudomonas abietaniphila (ATCC 700689); Pseudomonas aeruginosa (ATCC 10145); Pseudomonas alcaligenes (ATCC 14909); Pseudomonas anguilliseptica (ATCC 33660); Pseudomonas citronellolis (ATCC 13674); Pseudomonas flavescens (ATCC 51555); Pseudomonas mendocina (ATCC 51556); Pseudomonas spp. mendocina (ATCC 25411); Pseudomonas nitroreducens (ATCC 33634); Pseudomonas oleovorans (ATCC 8062); Pseudomonas pseudoalcaligenes (ATCC 17440); Pseudomonas resinovorans (ATCC 14235); Pseudomonas straminea (ATCC 33636); Pseudomonas agarici (ATCC 25941); Pseudomonas alcaliphila; Pseudomonas alginovora; Pseudomonasandersonii; Pseudomonas asplenii (ATCC 23835); Pseudomonas azelaica (ATCC 27162); Pseudomonas beyerinckii (ATCC 19372); Pseudomonas borealis; Pseudomonas boreopolis (ATCC 33662); Pseudomonas brassicacearum; Pseudomonas butanovora (ATCC 43655);Pseudomonas cellulosa (ATCC 55703); Pseudomonas aurantiaca (ATCC 33663); Pseudomonas chlororaphis (ATCC 9446, ATCC 13985, ATCC 17418, ATCC 17461); Pseudomonas fragi (ATCC 4973); Pseudomonas lundensis (ATCC 49968); Pseudomonas taetrolens (ATCC 4683); Pseudomonas cissicola (ATCC 33616); Pseudomonas coronafaciens; Pseudomonas diterpeniphila; Pseudomonas elongata (ATCC 10144); Pseudomonas flectens (ATCC 12775); Pseudomonas azotoformans; Pseudomonas brenneri; Pseudomonas cedrella; Pseudomonas corrugata (ATCC 29736); Pseudomonas extremorientalis; Pseudomonas fluorescens (ATCC 35858); Pseudomonas gessardii; Pseudomonas libanensis; Pseudomonas mandelii (ATCC 700871); Pseudomonas marginalis (ATCC 10844); Pseudomonas migulae; Pseudomonas mucidolens (ATCC 4685); Pseudomonas orientalis; Pseudomonas rhodesiae; Pseudomonas synxantha (ATCC 9890); Pseudomonas tolaasii (ATCC 33618);Pseudomonas veronii (ATCC 700474); Pseudomonas frederiksbergensis; Pseudomonas geniculata (ATCC 19374); Pseudomonas gingeri; Pseudomonas graminis; Pseudomonas grimontii; Pseudomonas halodenitrificans; Pseudomonas halophila; Pseudomonas hibiscicola (ATCC 19867); Pseudomonas huttiensis (ATCC 14670); Pseudomonas hydrogenovora; Pseudomonas jessenii (ATCC 700870); Pseudomonas kilonensis; Pseudomonas lanceolata (ATCC 14669); Pseudomonas lini; Pseudomonas marginate (ATCC 25417); Pseudomonas mephitica (ATCC 33665); Pseudomonas denitrificans (ATCC 19244); Pseudomonas pertucinogena (ATCC 190); Pseudomonas pictorum (ATCC 23328); Pseudomonas psychrophila; Pseudomonas filva (ATCC 31418); Pseudomonas monteilii (ATCC 700476); Pseudomonas morganii (ATCC 700477); Pseudomonas spp. mosselii); Pseudomonas oryzihabitans (ATCC 43272); Pseudomonas plecoglossicida (ATCC 700383); Pseudomonas putida (ATCC 12633); Pseudomonas reactans; Pseudomonas spinosa (ATCC 14606);Pseudomonas balearica; Pseudomonas luteola (ATCC 43273); Pseudomonas astutzeri (ATCC 17588); Pseudomonas amygdali (ATCC 33614); Pseudomonas avellanae (ATCC 700331); Pseudomonas carica papaya (ATCC 33615); Pseudomonas cichorii (ATCC 10857); Pseudomonas ficuserectae (ATCC 35104); Pseudomonas fuscovaginae; Pseudomonas meliae (ATCC 33050); Pseudomonas syringae (ATCC 19310); Pseudomonas viridiflava (ATCC 13223); Pseudomonas thermocarboxydovorans (ATCC 35961); Pseudomonas thermotolerans; Pseudomonas thivervalensis; Pseudomonas vancouverensis (ATCC 700688); Pseudomonas wisconsinensis and Pseudomonas xiamenensis. In one embodiment, the host cell is Pseudomonas fluorescens.

[0176] The host cell may also be selected from "Gram-negative Proteobacteria subgroup 17". "Gram-negative Proteobacteria subgroup 17" is defined as the group of Proteobacteria known in the art as "fluorescing Pseudomonas," including those belonging to the following species of the genus Pseudomonas: Pseudomonas azotoformans; Pseudomonas brenneri; Pseudomonas cedrella; Pseudomonas corrugata; Pseudomonas extremorientalis; Pseudomonas fluorescens; Pseudomonas gessardii; Pseudomonas libanensis; Pseudomonas mandelii; Pseudomonas marginalis; Pseudomonas migulae; Pseudomonas mucidolens; Pseudomonas orientalis; Pseudomonas rhodesiae; Pseudomonas xanthophyllids; Pseudomonas spp. synxantha); Pseudomonas tolaasii; and Pseudomonas veronii.

[0177] In some embodiments, the bacterial host cell used in the method of the present invention is defective in protease expression. In some embodiments, the bacterial host cell defective in protease expression is Pseudomonas. In some embodiments, the bacterial host cell defective in protease expression is Pseudomonas. In some embodiments, the bacterial host cell defective in protease expression is Pseudomonas.

[0178] In some embodiments, the bacterial host cell used in the methods of the present invention is not deficient in protease expression. In some embodiments, the bacterial host cell that is not deficient in protease expression is Pseudomonas. In some embodiments, the bacterial host cell that is not deficient in protease expression is Pseudomonas. In some embodiments, the bacterial host cell that is not deficient in protease expression is Pseudomonas. In some embodiments, the bacterial host cell that is not deficient in protease expression is Pseudomonas fluorescens.

[0179] In some embodiments, the Pseudomonas host cell used in the methods of the present invention is deficient in the expression of Lon protease (e.g., SEQ ID NO: 14), La1 protease (e.g., SEQ ID NO: 15), AprA protease (e.g., SEQ ID NO: 16), or a combination thereof. In some embodiments, the Pseudomonas host cell is deficient in the expression of AprA (e.g., SEQ ID NO: 16), HtpX (e.g., SEQ ID NO: 17), or a combination thereof. In some embodiments, the Pseudomonas host cell is deficient in the expression of Lon (e.g., SEQ ID NO: 14), La1 (e.g., SEQ ID NO: 15), AprA (e.g., SEQ ID NO: 16), HtpX (e.g., SEQ ID NO: 17), or a combination thereof. In some embodiments, the Pseudomonas host cell is deficient in the expression of Npr (e.g., SEQ ID NO: 20), DegP1 (e.g., SEQ ID NO: 18), DegP2 (e.g., SEQ ID NO: 19), or a combination thereof. In some embodiments, the Pseudomonas host cell is deficient in expression of La1 (e.g., SEQ ID NO: 15), Prc1 (e.g., SEQ ID NO: 21), Prc2 (e.g., SEQ ID NO: 22), PrtB (e.g., SEQ ID NO: 23), or a combination thereof. These proteases are known in the art and are described, for example, in U.S. Patent No. 8,603,824, "Process for Improved Protein Expression by Strain Engineering," U.S. Patent Application Publication No. 2008 / 0269070, and U.S. Patent Application Publication No. 2010 / 0137162, which disclose open reading frame sequences for the above-mentioned proteases.

[0180] Examples of Pseudomonas fluorescens host strains derived from the base strain MB101 (ATCC deposit PTA-7841) are useful in the methods of the present invention. In some embodiments, the Pseudomonas fluorescens used to express the hPTH fusion protein is, for example, DC454, DC552, DC572, DC1084, DC1106, DC508, DC992.1, PF1201.9, PF1219.9, PF1326.1, PF1331, PF1345.6, or DC1040.1-1. In some embodiments, the Pseudomonas fluorescens host strain is PF1326.1. In some embodiments, the Pseudomonas fluorescens host strain is PF1345.6. Using the information provided herein, recombinant DNA methods known in the art and described in the literature, and available materials, for example, the ATCC deposited Pseudomonas fluorescens strain MB101, as described, one skilled in the art can readily construct these and other strains useful in the methods of the present invention.

[0181] Expression strain

[0182] Can use the method described in this article and open literature to construct the expression strain useful for implementing the method of the present invention.In some embodiments, the expression strain useful in the inventive method comprises the plasmid that overexpresses one or more fluorescing Pseudomonas chaperone proteins or folding regulator protein.For example, DnaJ-like protein, FrnE, Fk1B or EcpD can be overexpressed in the expression strain.In some embodiments, the fluorescing Pseudomonas folding regulator overexpression (FMO) plasmid encodes C1pX, Fk1B3, FrnE, C1pA, Fkbp or ppiA.The example of the expression plasmid encoding Fkbp is pDOW1384-1.In some embodiments, the expression plasmid that does not encode the folding regulator is introduced into the expression strain.In these embodiments, plasmid is, for example, pDOW2247. In some embodiments, the Pseudomonas fluorescens expression strain used to express the hPTH fusion protein in the methods of the invention is STR35970, STR35984, STR36034, STR36085, STR36150, STR36169, STR35949, STR36098, or STR35783, as described elsewhere herein.

[0183] In some embodiments, the Pseudomonas fluorescens host strain used in the methods of the present invention is DC1106 (mtlDYZ knockout mutant ΔpyrFΔproCΔbenAB lsc::lacI Q1), a derivative of the deposited strain MB101, in which the genes pyrF, proC, benA, benB, and mtlDYZ are deleted from the mannitol (mtl) operon, and the E. coli lacI transcriptional repressor is inserted and fused to the levansucrase gene (lsc). The sequences of these genes and methods of using them are known in the art and described in the literature, for example, U.S. Patent Nos. 8,288,127, 8,017,355, "Mannitol induced promoter systems in bacterial host cells," and 7,794,972, "Benzoate-andanthranilate-inducible promoters," each of which is incorporated by reference.

[0184] Host cells equivalent to DC1106 or any of the host cells or expression strains described herein can be constructed from MB101 using methods described herein and in the disclosed literature. In some embodiments, host cells equivalent to DC1106 are used. Host cell DC454 is described in Schneider et al., 2005 (referred to therein as DC206), and in U.S. Patent No. 8,569,015, "rPA Optimization," all of which are incorporated herein by reference. DC206 is the same strain as DC454; after three passages in animal-free culture medium, it was renamed DC454.

[0185] Those of ordinary skill in the art will recognize that, in some embodiments, a deletion plasmid carrying a region flanking a gene to be deleted that is not replicated in Pseudomonas fluorescens can be used to perform genome deletion or mutation (e.g., inactivation or weakening mutation) by, for example, allelic exchange. A deletion plasmid can be constructed by PCR amplification of the gene to be deleted (including upstream and downstream regions of the gene to be deleted). Deletion can be verified by sequencing the PCR product amplified from genomic DNA using analytical primers, observed after electrophoresis separation in an agarose ointment gel, followed by DNA sequencing of the fragment. In some embodiments, gene inactivation is achieved by complete deletion, partial deletion, or mutation, for example, frameshift, point mutation, or insertion mutation.

[0186] In some embodiments, the strain used is transformed with an FMO plasmid according to methods known in the art. For example, DC1106 host cells can be transformed with the FMO plasmid pDOW1384, which overexpresses FkbP (RXF06591.1), a fold regulator belonging to the peptidyl-prolyl cis-trans isomerase family, to produce the expression strain STR36034. The genotypes of specific examples of hPTH fusion protein expression strains and corresponding host cells for expressing hPTH according to the methods of the present invention are listed in Table 4. In some embodiments, host cells equivalent to any of the host cells described in Table 4 are transformed with equivalent FMO plasmids described herein to obtain expression strains equivalent to the strains described herein for expressing hPTH 1-34 using the methods of the present invention. As discussed, suitable expression strains can be similarly obtained according to methods described herein and in the literature.

[0187] Table 4: Pseudomonas fluorescens host cells and expression strains used for PTH 1-34 fusion protein production

[0188]

[0189]

[0190] In some embodiments, the host cells or strains listed in Table 4, or equivalents of any of the host cells or strains described in Table 4, are used to express fusion proteins comprising a polypeptide of interest as described herein using the methods of the present invention. In some embodiments, the host cells or strains listed in Table 4, or equivalents of any of the host cells or strains described in Table 4, are used to express fusion proteins comprising hPTH, GCSF, or an insulin polypeptide (e.g., proinsulin) as described herein using the methods of the present invention. In some embodiments, wild-type host cells, e.g., DC454 or equivalents, are used to express fusion proteins comprising a polypeptide of interest as described herein using the methods of the present invention.

[0191] The sequences of these and other proteases and folding regulators used to generate the host strains of the present invention are known in the art and disclosed in the literature, for example, as provided in Tables A to F of the aforementioned U.S. Patent No. 8,603,824, which is incorporated by reference in its entirety. For example, the M50 S2P protease family membrane metalloprotease open reading frame sequence is provided therein as RXF04692.

[0192] High-throughput screening

[0193] In some embodiments, high throughput screening can be carried out to determine the optimum conditions for expressing soluble recombinant fusion protein.The conditions that can be changed in screening include, for example, host cell, the genetic background of host cell (for example, deletion of different proteases), the promoter type in the expression construct, the type of secretion leader sequence fused with the sequence of the encoding recombinant protein, growth temperature, the OD under induction when using an inducible promoter, the duration of protein induction when using the concentration of the IPTG for induction, the growth temperature after the inducing agent is added to the culture, the agitation rate of the culture, the selection method for plasmid maintenance, the culture volume in the container and the cell lysis method.

[0194] In some embodiments, a library (or "array") of host strains is provided, wherein each strain (or "host cell population") in the library has been genetically modified to regulate the expression of one or more target genes in the host cell. The "best host strain" or "best expression system" can be identified or selected based on the quantity, quality and / or location of the recombinant fusion protein expressed, compared to other host cell populations with significantly different phenotypes in the array. Thus, the best host strain is one that produces the recombinant fusion protein according to the desired specifications. Although the desired specifications will vary depending on the protein to be produced, the specifications include the quality and / or quantity of the protein, for example, whether the protein is sequestered or secreted, and in what quantity, whether the protein is processed and / or folded correctly or as desired, etc. In some embodiments, the improved or ideal quality can be the production of the recombinant fusion protein with high titer expression and low degradation levels. In some embodiments, the optimal host strain or optimal expression system produces a specific absolute level or a specific level of production relative to that produced by an indicator strain (i.e., a strain used for comparison), characterized by the amount or mass of soluble recombinant fusion protein, the amount or mass of recoverable recombinant fusion protein, the amount or mass of correctly processed recombinant fusion protein, the amount or mass of correctly folded recombinant fusion protein, the amount or mass of active recombinant fusion protein, and / or the total amount or mass of recombinant fusion protein.

[0195] Methods for screening microbial hosts to identify strains with improved yield and / or quality in expression of recombinant fusion proteins are described, for example, in U.S. Patent Application Publication No. 2008 / 0269070.

[0196] Fermentation form

[0197] The expression strain of the present invention can be cultured in any fermentation form. For example, batch, fed-batch, semi-continuous and continuous fermentation modes can be used herein.

[0198] In some embodiments, the fermentation medium can be selected from a rich medium, a minimal medium, and a mineral salt medium. In other embodiments, a minimal medium or a mineral salt medium is selected. In certain embodiments, a mineral salt medium is selected.

[0199] Mineral salts culture medium is made up of mineral salts and carbon source, and the carbon source is such as glucose, sucrose or glycerol.The example of mineral salts culture medium includes, for example, M9 culture medium, pseudomonas culture medium (ATCC 179) and Davis and Mingioli culture medium (referring to, Davis, BD and Mingioli, ES, 1950, J.Bact.60:17-28).Mineral salts for the preparation of mineral salts culture medium include those selected from such as potassium phosphate, ammonium sulfate or ammonium chloride, magnesium sulfate or magnesium chloride and trace minerals, such as calcium chloride, borate and the sulfate of iron, copper, manganese and zinc.Usually, organic nitrogen source is not included in mineral salts culture medium, such as peptone, tryptone, amino acid or yeast extract.Alternatively, inorganic nitrogen source is used, and this can be selected from such as ammonium salt, ammoniacal liquor and gaseous ammonia.Mineral salts culture medium usually contains glucose or glycerol as carbon source. In contrast to mineral salts medium, minimal medium may also contain mineral salts and a carbon source, but may be supplemented with, for example, low levels of amino acids, vitamins, peptones, or other ingredients, although these may be added at very low levels. Suitable medium for use in the methods of the present invention may be prepared using methods described in the literature, for example, in U.S. Patent Application Publication No. 2006 / 0040352, referenced above and incorporated by reference. Details of culture procedures and mineral salts medium useful in the methods of the present invention are described in Riesenberg, D et al., 1991, "High cell density cultivation of Escherichia coli at controlled specific growth rate," J. Biotechnol. 20(1): 17-27, incorporated herein by reference.

[0200] In some embodiments, production can be achieved in bioreactor culture. The culture can be grown in, for example, a bioreactor containing a mineral salt medium up to 2 liters and maintained at 32°C and pH 6.5 by adding ammonia. By increasing agitation and sparging air and oxygen into the fermentor, an excess of dissolved oxygen can be maintained. Glycerol can be delivered to the culture to maintain an excess level throughout the fermentation process. In some embodiments, these conditions are maintained until the target culture cell density for induction is reached, for example, an optical density at 575nm (A575), and IPTG is added to start target protein production. It will be understood that the cell density during induction, the concentration of IPTG, pH, temperature, CaCl2 concentration, dissolved oxygen flow rate, can each be changed to determine the optimal conditions for expression. In some embodiments, the cell density during induction can be adjusted from A575 to A575. 575 The absorbance units (AU) were varied from 40 to 200. The IPTG concentration was varied from 0.02 to 1.0 mM, the pH from 6 to 7.5, the temperature from 20 to 35°C, the CaCl2 concentration from 0 to 0.5 g / L, and the dissolved oxygen flow rate from 1 LPM (liters / minute) to 10 LPM. After 6-48 hours, the culture from each bioreactor was harvested by centrifugation and the cell pellet was frozen at -80°C. The samples were then analyzed for product formation, for example, by SDS-CGE.

[0201] Fermentation can be performed at any scale. The expression system according to the present invention is useful for expression of recombinant proteins at any scale. Thus, for example, fermentation volumes on the microliter scale, milliliter scale, centiliter scale, and deciliter scale can be used, and fermentation volumes on the 1 liter scale and larger can be used.

[0202] In some embodiments, the fermentation volume is or is greater than about 1 liter. In some embodiments, the fermentation volume is about 1 liter to about 100 liters. In some embodiments, the fermentation volume is about 1 liter, about 2 liters, about 3 liters, about 4 liters, about 5 liters, about 6 liters, about 7 liters, about 8 liters, about 9 liters, or about 10 liters. In some embodiments, the fermentation volume is about 1 liter to about 5 liters, about 1 liter to about 10 liters, about 1 liter to about 25 liters, about 1 liter to about 50 liters, about 1 liter to about 75 liters, about 10 liters to about 25 liters, about 25 liters to about 50 liters, or about 50 liters to about 100 liters. In other embodiments, the fermentation volume is or is greater than 5 liters, 10 liters, 15 liters, 20 liters, 25 liters, 50 liters, 75 liters, 100 liters, 200 liters, 250 liters, 300 liters, 500 liters, 1,000 liters, 2,000 liters, 5,000 liters, 10,000 liters, or 50,000 liters.

[0203] In some embodiments, the recombinant protein of the present invention can be produced by shaking flasks in a large volume of culture, for example, 0.5mL high throughput screening is cultivated, and by large volume of culture, for example, 50mL shake flask culture, 1L or larger culture produces the amount of recombinant protein that improves.This may not only be due to the increase of culture size, and may be that for example, in large-scale fermentation, cell growth is made to the ability (for example, reflected by culture absorbance) of higher density.For example, the volumetric productivity from same strain can be improved up to ten times from HTP scale to large-scale fermentation.In some embodiments, after large-scale fermentation, the volumetric productivity observed for same expression strain is higher than 2 times to 10 times of HTP scale growth. In some embodiments, the amount of yield observed for the same expression strain after large-scale fermentation is 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 2-fold to 10-fold, 2-fold to 9-fold, 2-fold to 8-fold, 2-fold to 7-fold, 2-fold to 6-fold, 2-fold to 5-fold, 2-fold to 4-fold, 2-fold to 3-fold, 3-fold to 10-fold, 3-fold to 9-fold, 3-fold to 8-fold, 3-fold to 7-fold, 3-fold to 6-fold, 3-fold to 3-fold, 3-fold to 4-fold, 2-fold to 3-fold, 3-fold to 10-fold, 3-fold to 9-fold, 3-fold to 8-fold, fold to 5-fold, 3-fold to 4-fold, 4-fold to 10-fold, 4-fold to 9-fold, 4-fold to 8-fold, 4-fold to 7-fold, 4-fold to 6-fold, 4-fold to 5-fold, 5-fold to 10-fold, 5-fold to 9-fold, 5-fold to 8-fold, 5-fold to 7-fold, 5-fold to 6-fold, 6-fold to 10-fold, 6-fold to 9-fold, 6-fold to 8-fold, 6-fold to 7-fold, 7-fold to 10-fold, 7-fold to 9-fold, 7-fold to 8-fold, 8-fold to 10-fold, 8-fold to 9-fold, 9-fold to 10-fold. See, e.g., Retallack et al., 2012, “Reliable protein production in a Pseudomonasfluorescens expression system,” Prot. Exp. and Purif. 81: 157-165, incorporated herein by reference in its entirety.

[0204] Bacterial growth conditions

[0205] Useful growth conditions in the methods provided herein can include a temperature of about 4° C. to about 42° C. and a pH of about 5.7 to about 8.8. When an expression construct having a lacZ promoter is used, expression can be induced by adding IPTG to the culture at a final concentration of about 0.01 mM to about 1.0 mM.

[0206] The pH of the culture can be maintained using pH buffers and methods known to those skilled in the art. Ammonia can also be used to achieve pH control during the culture. In some embodiments, the pH of the culture is about 5.7 to about 8.8. In some embodiments, the pH is about 5.7, 5.8, 5.9, 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, 7.0, 7.1, 7.2, 7.3, 7.4, 7.5, 7.6, 7.7, 7.8, 7.9, 8.0, 8.1, 8.2, 8.3, 8.4, 8.5, 8.6, 8.7 or 8.8. In some embodiments, the pH is from about 5.7 to about 8.8, from about 5.7 to about 8.5, from about 5.7 to about 8.3, from about 5.7 to about 8, from about 5.7 to about 7.8, from about 5.7 to about 7.6, from about 5.7 to about 7.4, from about 5.7 to about 7.2, from about 5.7 to about 7, from about 5.7 to about 6.8, from about 5.7 to about 6.6, from about 5.7 to about 6.4, from about 5.7 to about 6.2, from about 5.7 to about 6, from about 5.9 to about 8.8, from about 5.9 to about 8.5, from about 5.9 to about 8.3, from about 5.9 to about 8, from about 5.9 to about 7.8, from about 5.9 to about 7.6, about 5.9 to about 7.4, about 5.9 to about 7.2, about 5.9 to about 7, about 5.9 to about 6.8, about 5.9 to about 6.6, about 5.9 to about 6.4, about 5.9 to about 6.2, about 6 to about 8.8, about 6 to about 8.5, about 6 to about 8.3, about 6 to about 8, about 6 to about 7.8, about 6 to about 7.6, about 6 to about 7.4, about 6 to about 7.2, about 6 to about 7, about 6 to about 6.8, about 6 to about 6.6, about 6 to about 6.4, about 6 to about 6.2, about 6.1 to about 8.8, about 6.1 to about 8. 5, about 6.1 to about 8.3, about 6.1 to about 8, about 6.1 to about 7.8, about 6.1 to about 7.6, about 6.1 to about 7.4, about 6.1 to about 7.2, about 6.1 to about 7, about 6.1 to about 6.8, about 6.1 to about 6.6, about 6.1 to about 6.4, about 6.2 to about 8.8, about 6.2 to about 8.5, about 6.2 to about 8.3, about 6.2 to about 8, about 6.2 to about 7.8, about 6.2 to about 7.6, about 6.2 to about 7.4, about 6.2 to about 7, about 6.2 to about 6.8, about 6.2 to about 6.6, about 6.2 to about 6.4, about 6.4 to about 8.8, about 6.4 to about 8.5, about 6.4 to about 8.3, about 6.4 to about 8, about 6.4 to about 7.8, about 6.4 to about 7.6, about 6.4 to about 7.4, about 6.4 to about 7.2, about 6.4 to about 7, about 6.4 to about 6.8, about 6.4 to about 6.6, about 6.6 to about 8.8, about 6.6 to about 8.5, about 6.6 to about 8.3, about 6.6 to about 8, about 6.6 to about 7.8, about 6.6 to about 7.6, about 6.6 to about 7.4, about 6.6 to about 7.2, about 6.6 to about 7, about 6.6 to about 6.8, about 6.8 to about 8.8, about 6.8 to about 8.5, about 6.8 to about 8.3, about 6.8 to about 8, about 6.8 to about 7.8, about 6.8 to about 7.6, about 6.8 to about 7.4, about 6.8 to about 7.2, about 6.8 to about 7, about 7 to about 8.8, about 7 to about 8.5, about 7 to about 8.3, about 7 to about 8, about 7 to about 7.8, about 7 to about 7.6, about 7 to about 7.4, about 7 to about 7.2, about 7.2 to about 8.8, about 7.2 to about 8.5, about 7.2 to about 8.3, about 7 .2 to about 8, about 7.2 to about 7.8, about 7.2 to about 7.6, about 7.2 to about 7.4, about 7.4 to about 8.8, about 7.4 to about 8.5, about 7.4 to about 8.3, about 7.4 to about 8, about 7.4 to about 7.8, about 7.4 to about 7.6, about 7.6 to about 8.8, about 7.6 to about 8.5, about 7.6 to about 8.3, about 7.6 to about 8, about 7.6 to about 7.8, about 7.8 to about 8.8, about 7.8 to about 8.5, about 7.8 to about 8.3, about 7.8 to about 8, about 8 to about 8.8, about 8 to about 8.5, or about 8 to about 8.3. In some embodiments, the pH is about 6.5 to about 7.2.

[0207] In some embodiments, the growth temperature is maintained at about 4° C. to about 42° C. In some embodiments, the growth temperature is about 4° C., about 5° C., about 6° C., about 7° C., about 8° C., about 9° C., about 10° C., about 11° C., about 12° C., about 13° C., about 14° C., about 15° C., about 16° C., about 17° C., about 18° C., about 19° C., about 20° C., about 21° C., about 22° C., about 23° C., about 24° C., about 25° C., about 26° C., about 27° C., about 28° C., about 29° C., about 30° C., about 31° C., about 32° C., about 33° C., about 34° C., about 35° C., about 36° C., about 37° C., about 38° C., about 39° C., about 40° C., about 41° C., or about 42° C. In some embodiments, the growth temperature is about 25° C. to about 32° C. In some embodiments, the growth temperature is maintained at about 22°C to about 27°C, about 22°C to about 28°C, about 22°C to about 29°C, about 22°C to about 30°C, 23°C to about 27°C, about 23°C to about 28°C, about 23°C to about 29°C, about 23°C to about 30°C, about 24°C to about 27°C, about 24°C to about 28°C, about 24°C to about 29°C, about 24°C to about 30°C, about 25°C to about 27°C, about 25°C to about 28°C, about 25°C to about 29°C, about 25°C to about 30°C, about 25°C to about 31°C, about 25°C to about 32°C, about 25°C to about 33°C, about 26°C to about 28°C, about 26°C to about 29°C, about 26°C to about 3 32°C, about 29°C to about 31°C, about 29°C to about 32°C, about 29°C to about 33°C, about 28°C to about 30°C, about 28°C to about 31°C, about 28°C to about 32°C, about 29°C to about 31°C, about 29°C to about 32°C, about 29°C to about 33°C, about 30°C to about 32°C, about 30°C to about 33°C, about 31°C to about 33°C, about 31°C to about 32°C, about 21°C to about 42°C, about 22°C to about 42°C, about 23°C to about 42°C, about 24°C to about 42°C, about 25°C to about 42°C. In some embodiments, the growth temperature is from about 25°C to about 28.5°C. In some embodiments, the growth temperature is greater than about 20°C, greater than about 21°C, greater than about 22°C, greater than about 23°C, greater than about 24°C, greater than about 25°C, greater than about 26°C, greater than about 27°C, greater than about 28°C, greater than about 29°C, or greater than about 30°C.

[0208] In some embodiments, the temperature is changed during the culture process. In some embodiments, before an agent (e.g., IPTG) is added to the culture to induce expression from the construct, the temperature is maintained at about 30°C to about 32°C, and after the inducing agent is added, the temperature is lowered to about 25°C to about 28°C. In some embodiments, before an agent (e.g., IPTG) is added to the culture to induce expression from the construct, the temperature is maintained at about 30°C, and after the inducing agent is added, the temperature is lowered to about 25°C.

[0209] As described elsewhere herein, an inducible promoter can be used in the expression construct to control the expression of the recombinant fusion protein, for example, the lac promoter. In the case of a lac promoter derivative or family member (e.g., tac promoter), the effector compound is an inducer, for example, a placebo inducer such as IPTG. In some embodiments, a lac promoter derivative is used, and the expression of the recombinant fusion protein is controlled by OD values ​​of about 40 to about 180. 575 When the expression of the recombinant protein is at a determined level, the expression of the recombinant protein is induced by adding IPTG to a final concentration of about 0.01 mM to about 1.0 mM. In some embodiments, the OD of the culture at the time of induction of the recombinant protein is 575 It can be about 40, about 50, about 60, about 70, about 80, about 90, about 110, about 120, about 130, about 140, about 150, about 160, about 170, about 180. In other embodiments, OD 575 In other embodiments, OD 575 In other embodiments, OD 575 The OD value of a Pseudomonas fluorescens culture is about 40 to about 140, or about 80 to about 180. Cell density can be measured by other methods and expressed in other units, for example, cells per unit volume. For example, an OD value of about 40 to about 160 for a Pseudomonas fluorescens culture is about 160. 575 Equivalent to approximately 4×10 10 to about 1.6×10 10 Colony forming units / mL or 17.5 to 70 g / L dry cell weight. In some embodiments, the cell density at the time of culture induction is equivalent to the cell density specified herein by absorbance at OD575, regardless of the method or measurement unit used to determine cell density. Those skilled in the art will know how to perform appropriate conversions for any cell culture.

[0210] In some embodiments, the final IPTG concentration of the culture is about 0.01 mM, about 0.02 mM, about 0.03 mM, about 0.04 mM, about 0.05 mM, about 0.06 mM, about 0.07 mM, about 0.08 mM, about 0.09 mM, about 0.1 mM, about 0.2 mM, about 0.3 mM, about 0.4 mM, about 0.5 mM, about 0.6 mM, about 0.7 mM, about 0.8 mM, about 0.9 mM, or about 1 mM. In some embodiments, the final IPTG concentration of the culture is about 0.08 mM to about 0.1 mM, about 0.1 mM to about 0.2 mM, about 0.2 mM to about 0.3 mM, about 0.3 mM to about 0.4 mM, about 0.2 mM to about 0.4 mM, about 0.08 to about 0.2 mM, or about 0.1 to 1 mM.

[0211] As described herein and in the literature, in embodiments where non-lac-type promoters are used, other inducers or effectors may be used. In one embodiment, the promoter is a constitutive promoter.

[0212] After adding the inducing agent, the culture is grown for a period of time, for example, about 24 hours, during which the recombinant protein is expressed. After adding the inducing agent, the culture can be grown for about 1 hr, about 2 hr, about 3 hr, about 4 hr, about 5 hr, about 6 hr, about 7 hr, about 8 hr, about 9 hr, about 10 hr, about 11 hr, about 12 hr, about 13 hr, about 14 hr, about 15 hr, about 16 hr, about 17 hr, about 18 hr, about 19 hr, about 20 hr, about 21 hr, about 22 hr, about 23 hr, about 24 hr, about 36 hr or about 48 hr. After adding the inducing agent to the culture, the culture can be grown for about 1 to 48 hr, about 1 to 24 hr, about 1 to 8 hr, about 10 to 24 hr, about 15 to 24 hr or about 20 to 24 hr. The cell culture can be concentrated by centrifugation, and the culture pellet can be resuspended in a buffer or solution suitable for subsequent lysis procedures.

[0213] In some embodiments, cells are destroyed using equipment for high-pressure mechanical cell destruction (which is commercially available, for example, Microfluidics microfluidizer, Constant cell crusher, Niro-Soavi homogenizer, or APV-Gaulin homogenizer). For example, cells expressing recombinant proteins can be destroyed using ultrasonic treatment. Any suitable method known in the art for lysing cells can be used to release the soluble portion. For example, in some embodiments, chemical and / or enzymatic cell lysis reagents, such as cell wall lysing enzymes and EDTA, can be used. In the methods of the present invention, the use of frozen or previously stored cultures is also contemplated. Before lysis, the culture can be OD-standardized. For example, cells can be standardized to an OD600 of about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20.

[0214] Any suitable apparatus and method can be used for centrifugation. Centrifugation of cell cultures or lysates for the purpose of separating the soluble portion from the insoluble portion is well known in the art. For example, the lysed cells can be centrifuged at 20,800 × g for 20 minutes (at 4° C.), and the supernatant can be removed using manual or automated liquid handling. The cell pellet obtained by centrifugation of the cell culture, or the insoluble portion obtained by centrifugation of the cell lysate, can be resuspended in a buffer solution. Resuspension of the cell pellet or the insoluble portion can be performed using, for example, an apparatus such as an impeller connected to an overhead mixer, a magnetic stirring bar, a seesaw shaker, etc.

[0215] Non-denaturing conditions

[0216] The lysis of the induced host cells is carried out under non-denaturing conditions. In some embodiments, the non-denaturing conditions include using a non-denaturing treatment buffer, for example, to resuspend the cell pellet or cell paste. In some embodiments, the non-denaturing treatment buffer comprises sodium phosphate or Tris buffer, glycerol and sodium chloride. In embodiments where affinity chromatography is performed by immobilized metal affinity chromatography (IMAC), the non-denaturing treatment buffer comprises imidazole. In some embodiments, the non-denaturing treatment buffer comprises 0 to 50mM imidazole. In some embodiments, the non-denaturing treatment buffer does not comprise imidazole. In some embodiments, the non-denaturing treatment buffer comprises 25mM imidazole. In some embodiments, the non-denaturing treatment buffer comprises 10-30mM sodium phosphate or Tris, pH 7 to 9. In some embodiments, the non-denaturing treatment buffer has a pH of 7.3, 7.4 or 7.5. In some embodiments, the non-denaturing treatment buffer comprises 2-10% glycerol. In some embodiments, the non-denaturing treatment buffer comprises 50mM to 750mM NaCl. In some embodiments, the cell paste is resuspended to 10-50% solids. In some embodiments, the non-denaturing processing buffer comprises 20 mM sodium phosphate, 5% glycerol, 500 mM sodium chloride, 20 mM imidazole, pH 7.4, and is resuspended to 20% solids. In some embodiments, the non-denaturing processing buffer comprises 20 mM Tris, 50 mM NaCl, pH 7.5, and is resuspended to 20% solids.

[0217] In some embodiments, the non-denaturing treatment buffer does not contain a chaotropic agent. Chaotropic agents destroy the three-dimensional structure of proteins or nucleic acids, thereby causing denaturation. In some embodiments, the non-denaturing treatment buffer comprises a non-denaturing concentration of a chaotropic agent. In some embodiments, the chaotropic agent is, for example, urea or guanidine hydrochloride. In some embodiments, the non-denaturing treatment buffer comprises 0 to 4M urea or guanidine hydrochloride. In some embodiments, the non-denaturing treatment buffer comprises less than 4M, less than 3.5M, less than 3M, less than 2.5M, less than 2M, less than 1.5M, less than 1M, less than 0.5M, about 0.1M, about 0.2M, about 0.3M, about 0.4M, about 0.5M, about 0.6M, about 0.7M, about 0.8M, about 0.9M, about 1.0M, about 1.1M, about 1.2M, About 1.3M, about 1.4M, about 1.5M, about 1.6M, about 1.7M, about 1.8M, about 1.9M, or about 2.0M, about 2.1M, about 2.2M, about 2.3M, about 2.4M, about 2.5M, about 2.6M, about 2.7M, about 2.8M, about 2.9M, about 3M, about 3.1M, about 3.2M, about 3.3M, about 3.4M, about 3.5M, about 3 .6M, about 3.7M, about 3.8M, about 3.9M, about 4M, about 0.5 to about 3.5M, about 0.5 to about 3M, about 0.5 to about 2.5M, about 0.5 to about 2M, about 0.5 to about 1.5M, about 0.5 to about 1M, about 1 to about 4M, about 1 to about 3.5M, about 1 to about 3M, about 1 to about 2.5M, about 1 to about 2M, about 1 to about 1.5M, about 1.5 to Urea or guanidine hydrochloride at a concentration of about 4 M, about 1.5 to about 3.5 M, about 1.5 to about 3 M, about 1.5 to about 2.5 M, about 1.5 to about 2 M, about 2 to about 4 M, about 2 to about 3.5 M, about 2 to about 3 M, about 2 to about 2.5 M, about 2.5 to about 4 M, about 2.5 to about 3.5 M, about 2.5 to about 3 M, about 3 to about 4 M, about 3 to about 3.5 M, or about 0.5 to about 1 M.

[0218] In embodiments where a non-denaturing treatment buffer is used, the cell paste is slurried at 20% solids in 20 mM Tris, 50 mM NaCl, 4 M urea, pH 7.5 for 1-2.5 hours at 2-8°C. In some embodiments, the cell paste is lysed using a Niro homogenizer, e.g., at 15,000 psi and batch centrifuged at 14,000 × g for 35 minutes or continuously centrifuged at 15,000 × g and 340 mL / min feed, the supe / centrate filtered with a depth filter and a membrane filter, diluted 2X in a resuspension buffer, e.g., 1X PBS pH 7.4, and loaded onto a capture column. In some embodiments, the non-denaturing treatment buffer contains other components, e.g., imidazole for IMAC, as described elsewhere herein.

[0219] Those skilled in the art will appreciate that the denaturing concentration of a chaotropic agent may be affected by pH, and that the level of denaturation depends on the characteristics of the protein. For example, increasing pH may cause protein denaturation despite a low concentration of chaotropic agent.

[0220] Product Evaluation

[0221] The quality of the recombinant fusion protein or target polypeptide produced can be evaluated by any method known in the art or described in the literature. In some embodiments, the denaturation of the protein is evaluated based on its solubility or by lacking or losing biological activity. For many proteins, biological activity tests are commercially available. Biological activity tests can include, for example, antibody binding tests. In some embodiments, physical characterization of the recombinant fusion protein or target polypeptide is carried out using methods available in this area, for example, chromatography and spectrophotometry. The evaluation of the target polypeptide can include determining that it has been correctly released (for example, its N-terminus is intact).

[0222] The activity of hPTH (e.g., hPTH 1-34 or 1-84) can be assessed using any method known in the art or described herein or in the literature, for example, using antibodies that recognize the N-terminus of the protein. Such methods include, for example, intact mass analysis. PTH bioactivity can be measured by, for example, cAMP ELISA, homogeneous time-resolved fluorescence (HTRF) assay (Charles River Laboratories), or as described in Nissenson et al., 1985, "Activation of the Parathyroid Hormone Receptor-Adenylate Cyclase System in Osteosarcoma Cells by a Human Renal Carcinoma Factor," Cancer Res. 45:5358-5363 and U.S. Pat. No. 7,150,974, "Parathyroid Hormone Receptor Binding Method," each of which is incorporated herein by reference. Methods for assessing PTH are also described by Shimizu et al., 2001, “Parathyroid hormone (1-14) and (1-11) analogs conformationally constrained by α-aminoisobutyric acid mediate full agonist responses via the Juxtamembrane region of the PTH-1 receptor,” J. Biol. Chem. 276:49003-49012, which is incorporated herein by reference.

[0223] Purification of recombinant fusion proteins and target peptides

[0224] In some embodiments, the fusion protein of the present invention can be separated or purified by any method known to those skilled in the art or described in the literature, for example, centrifugal method and / or chromatographic method, such as size exclusion, anion or cation exchange, hydrophobic interaction or affinity chromatography, from other proteins and cell debris separation or purification of dissolved recombinant fusion protein or target polypeptide. In some embodiments, fast high performance liquid chromatography (FPLC) is used to purify dissolved protein. FPLC is a liquid chromatography form for separating proteins based on the affinity of various resins. In some embodiments, the affinity tag expressed together with the fusion protein causes the fusion protein dissolved in the solubilization buffer to be bound to the resin, and impurities are carried in the solubilization buffer simultaneously. Subsequently, elution buffer is used, with a gradual gradient or in a stepwise manner to add, to dissociate the fusion protein from the ion exchange resin and to separate the pure fusion protein in the elution buffer.

[0225] In some embodiments, after induction is complete, the fermentation broth is collected by centrifugation, for example, at 15,900 × g for 60 to 90 minutes. The cell paste and supernatant are separated, and the paste is frozen at -80°C. The frozen cell paste is thawed in a buffer (e.g., a non-denaturing buffer or a buffer without urea) as described elsewhere herein. In some embodiments, the frozen cell paste is thawed and resuspended in 20 mM sodium phosphate, 5% glycerol, 500 mM sodium chloride, pH 7.4. In some embodiments, the buffer contains imidazole. In some embodiments, the final volume of the suspension is adjusted to the desired solid percentage, for example, 20% solids. The cells can be lysed chemically or mechanically, for example, the material can then be homogenized by passing through a microfluidizer at 15,000 psi. The lysate is centrifuged, eg, at 12,000 xg for 30 minutes, and filtered, eg, through a Sartorius Sartobran 150 (0.45 / 0.2 μm) capsule filter.

[0226] In some embodiments, fast protein liquid chromatography (FPLC) can be used for purification, for example, using a Frac-950 fraction collector. Explorer 100 chromatography system (GE Healthcare). In embodiments where a His-tag is used, the sample can be loaded onto a HisTrap FF, 10 mL column (two 5 mL HisTrap FF cartridges [GE Healthcare, part number 17-5255-01] connected in series), washed and eluted, e.g., using a 10 column volume linear gradient of elution buffer (by varying the imidazole concentration from 0 mM to 200 mM), and fractions collected.

[0227] In some embodiments, chromatography can be performed as appropriate for the polypeptide of interest. For example, immobilized metal ion affinity chromatography purification (eg, using nickel IMAC) can be performed as described in the Examples herein.

[0228] Cleavage of recombinant fusion proteins

[0229] In some embodiments, the purified recombinant fusion protein fraction is incubated with a lytic enzyme to cut the target polypeptide from the joint and the N-terminal fusion partner. In some embodiments, the lytic enzyme is a protease, for example, a serine protease, for example, bovine enterokinase, porcine enterokinase, trypsin, or any other suitable protease described elsewhere herein. Any suitable protease cleavage method known in the art and described in the literature (including manufacturer's instructions) can be used. Protease can be commercially available, for example, from Sigma-Aldrich (St.Louis, MO), ThermoFisher Scientific (Waltham, MA), and Promega (Madison, WI). For example, in some embodiments, bovine enterokinase (for example, Novagen catalog number #69066-3, batch D00155747) can be concentrated to cut the fusion protein purification fraction and be resuspended in a buffer containing 20mM Tris pH 7.4, 50mM NaCl, and 2mM CaCl2. Two units of bovine enterokinase are added to 100 μg of protein in 100 μL reactions. The mixture of fusion protein purified fraction and enterokinase is hatched for a suitable length of time. In some embodiments, a control reaction without enterokinase is also hatched for comparison. The enzyme reaction can be stopped by adding a complete protease inhibitor cocktail containing hydrochloric acid 4-benzenesulfonyl fluoride (AEBSF, Sigma catalog number # P8465).

[0230] In some embodiments, the lytic enzyme incubation is carried out for about 1 hour to about 24 hours. In some embodiments, the incubation is carried out for about 1 hr, about 2 hr, about 3 hr, about 4 hr, about 5 hr, about 6 hr, about 7 hr, about 8 hr, about 9 hr, about 10 hr, about 11 hr, about 12 hr, about 13 hr, about 14 hr, about 15 hr, about 16 hr, about 17 hr, about 18 hr, about 19 hr, about 20 hr, about 21 hr, about 22 hr, about 23 hr, about 24 hr, about 1 hr to about 24 hr, about 1 hr to about 23 hr, about 1 hr to about 22 hr, about 1 hr to about 21 hr, about 1 hr to about 20 hr, about 1 hr to about 19 hr, about 1 hr to about 18 hr, about 1 hr to about 17 hr, about 1 hr to about 18 hr, about 1 hr to about 19 hr, about 1 hr to about 19 hr, about 1 hr to about 18 ... 6hr, about 1hr to about 15hr, about 1hr to about 14hr, about 1hr to about 13hr, about 1hr to about 12hr, about 1hr to about 11hr, about 1hr to about 10hr, about 1hr to about 9hr, about 1hr to about 8hr, about 1hr to about 7hr, about 1hr to about 6hr, about 1hr to about 5hr, about 1hr to about 4hr, about 1hr to about 3hr, about 1hr to about 2hr, about 2hr to about 24hr, about 2hr to about 23hr, about 2hr to about 22hr, about 2hr to about 21hr, about 2hr to about 20hr, about 2hr to about 19hr, about 2hr to about 18hr, about 2hr to about 17hr, about 2hr to about 16hr, about 2hr to about 15hr, about 2hr to about 14hr, about 2hr to about 13hr, about 2hr to about 12hr, about 2hr to about 11hr, about 2hr to about 10hr, about 2hr to about 9hr, about 2hr to about 8hr, about 2hr to about 7hr, about 2hr to about 6hr, about 2hr to about 5hr, about 2hr to about 4hr, about 2hr to about 3hr, about 3hr to about 24hr, about 3hr to about 23hr, about 3hr to about 22hr, about 3hr to about 21hr, about 3hr to about 20hr, about 3hr to about 19hr, about 3hr to about 18hr, about 3hr to about 17hr, about 3hr to about 16hr, about 3h hr to about 15hr, about 3hr to about 14hr, about 3hr to about 13hr, about 3hr to about 12hr, about 3hr to about 11hr, about 3hr to about 10hr, about 3hr to about 9hr, about 3hr to about 8hr, about 3hr to about 7hr, about 3hr to about 6hr, about 3hr to about 5hr, about 3hr to about 4hr, about 4hr to about 24hr, about 4hr to about 23hr, about 4hr to about 22hr, about 4hr to about 21hr, about 4hr to about 20hr, about 4hr to about 19hr, about 4hr to about 18hr, about 4hr to about 17hr, about 4hr to about 16hr, about 4hr to about 15hr, about 4hr to about 14hr,about 4hr to about 13hr, about 4hr to about 12hr, about 4hr to about 11hr, about 4hr to about 10hr, about 4hr to about 9hr, about 4hr to about 8hr, about 4hr to about 7hr, about 4hr to about 6hr, about 4hr to about 5hr, about 5hr to about 24hr, about 5hr to about 23hr, about 5hr to about 22hr, about 5hr to about 20hr, about 5hr to about 21hr, about 5hr to about 19hr, about 5hr to about 18hr, about 5hr to about 17hr, about 5hr to about 16hr, about 5hr to about 15hr, about 5hr to about 14hr, about 5hr to about 13hr, about 5hr to about 12hr, about 5hr to about 11hr, about 5hr to about 10hr, about 5hr to about 9hr, about 5hr to about 8hr, about 5hr to about 7hr, about 5hr to about 6hr, about 6hr to about 24hr, about 6hr to about 23hr, about 6hr to about 22hr, about 6hr to about 21hr, about 6hr to about 20hr, about 6hr to about 19hr, about 6hr to about 18hr, about 6hr to about 17hr, about 6hr to about 16hr, about 6hr to about 15hr, about 6hr to about 14hr, about 6hr to about 13hr, about 6hr to about 12hr, about 6hr to about 11hr, about 6hr to about 10hr, about 6hr to about 9hr, about 6hr to about 8hr, about 6hr to about 7hr, about 7hr to about 24hr, about 7hr to about 23hr, about 7hr to about 22hr, about 7hr to about 21hr, about 7hr to about 20hr, about 7hr to about 19hr, about 7hr to about 18hr, about 7hr to about 17hr, about 7hr to about 16hr, about 7hr to about 15hr, about 7hr to about 14hr, about 7hr to about 13hr, about 7hr to about 12hr, about 7hr to about 11hr, about 7hr to about 10hr, about 7hr to about 9hr, about 7hr to about 8hr, about 8hr to about 24hr, about 8hr to about 23hr, about 8hr to about 22hr, about 8hr to about 21hr, about 8hr to about 20hr, about 8hr to about 19hr, about 8hr to about 18 ... hr to about 18hr, about 8hr to about 17hr, about 8hr to about 16hr, about 8hr to about 15hr, about 8hr to about 14hr, about 8hr to about 13hr, about 8hr to about 12hr, about 8hr to about 11hr, about 8hr to about 10hr, about 8hr to about 9hr, about 9hr to about 24hr, about 9hr to about 23hr, about 9hr to about 22hr, about 9hr to about 21hr, about 9hr to about 20hr, about 9hr to about 19hr, about 9hr to about 18hr, about 9hr to about 17hr, about 9hr to about 16hr, about 9hr to about 15hr, about 9hr to about 14hr, about 9hr to about 13hr, about 9hr to about 12hr,about 9hr to about 11hr, about 9hr to about 10hr, about 10hr to about 24hr, about 10hr to about 23hr, about 10hr to about 22hr, about 10hr to about 21hr, about 10hr to about 20hr, about 10hr to about 19hr, about 10hr to about 18hr, about 10hr to about 17hr, about 10hr to about 16hr, about 10hr to about 15hr, about 10hr to about 14hr, about 10hr to about 13hr, about 10hr to about 12hr, about 10hr to about 11hr, about 11hr to about 24hr, about 11hr to about 23hr, about 11hr to about 22hr, about 11hr to about 21hr, about 11hr to about 20hr , about 11hr to about 19hr, about 11hr to about 18hr, about 11hr to about 17hr, about 11hr to about 16hr, about 11hr to about 15hr, about 11hr to about 14hr, about 11hr to about 13hr, about 11hr to about 12hr, about 12hr to about 24hr, about 12hr to about 23hr, about 12hr to about 22hr, about 12hr to about 21hr, about 12hr to about 20hr, about 12hr to about 112hr, about 12hr to about 18hr, about 12hr to about 17hr, about 12hr to about 16hr, about 12hr to about 15hr, about 12hr to about 14hr, about 12hr to about 13hr, about 13hr to about 24hr, about 13hr to about 23hr, about 13hr to about 22hr, about 13hr to about 21hr, about 13hr to about 20hr, about 13hr to about 19hr, about 13hr to about 18hr, about 13hr to about 17hr, about 13hr to about 16hr, about 13hr to about 15hr, about 13hr to about 14hr, about 14hr to about 24hr, about 14hr to about 23hr, about 14hr to about 22hr, about 14hr to about 21hr, about 14hr to about 20hr, about 14hr to about 19hr, about 14hr to about 18hr, about 14hr to about 17hr, about 14hr to about 16hr, about 14hr to about 15hr, about 15h 15 hr to about 24 hr, about 15 hr to about 23 hr, about 15 hr to about 22 hr, about 15 hr to about 21 hr, about 15 hr to about 20 hr, about 15 hr to about 19 hr, about 15 hr to about 18 hr, about 15 hr to about 17 hr, about 16 hr to about 24 hr, about 16 hr to about 23 hr, about 16 hr to about 22 hr, about 16 hr to about 21 hr, about 16 hr to about 20 hr, about 16 hr to about 19 hr, about 16 hr to about 18 hr, or about 16 hr to about 17 hr, about 17 hr to about 24 hr, about 17 hr to about 23 hr, about 17 hr to about 22 hr, about 17 hr to about 21 hr, about 17 hr to about 20 hr,From about 17hr to about 19hr, from about 17hr to about 18hr, from about 18hr to about 24hr, from about 18hr to about 23hr, from about 18hr to about 22hr, from about 18hr to about 21hr, from about 18hr to about 20hr, from about 18hr to about 19hr, from about 19hr to about 24hr, from about 19hr to about 23hr, from about 19hr to about 22hr, from about 19hr to about 21hr, from about 19hr to about 20hr, from about 20hr to about 24hr, from about 20hr to about 23hr, from about 20hr to about 22hr, from about 20hr to about 21hr, from about 21hr to about 24hr, from about 21hr to about 23hr, from about 21hr to about 22hr, from about 22hr to about 24hr or from about 22hr to about 23hr.

[0231] In some embodiments, the extent of cleavage of the recombinant fusion protein after incubation with a protease is about 90% to about 100%. In some embodiments, the extent of cleavage after incubation with a protease is about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, about 100%, about 91% to about 100%, about 92% to about 100%, about 93% to about 100%, about 94% to about 100%, about 95% to about 100%, about 96% to about 100%, about 97% to about 100%, about 98% to about 100%, about 99% to about 100%, about 90% to about 99%, about 91% to about 99%, about 92% to about 99%, about 93% to about 99%, about 94% to about 99%, about 95% to about 99%, about 96% to about 99%, about 97% to about 99%, about 98% to about 99%, about 90% to about 98%, about 91% to about 98%, about 92% to about 98%, about 93% to about 98%, about 94% to about 98%, about 95% to about 98%, about 96% to about 98%, about 97% to about 98%, about 90% to about 97%, about 91% to about 97%, about 92% to about 97%, about 93% to about 97%, about 94% to about 97%, about 95% to about 97%, about 96% to about 97%, about 90% to about 96%, about 91% to about 96%, about 92% to about 96%, about 93% to about 96%, about 94% to about 96%, about 95 % to about 96%, about 90% to about 95%, about 91% to about 95%, about 92% to about 95%, about 93% to about 95%, about 94% to about 95%, about 90% to about 94%, about 91% to about 94%, about 92% to about 94%, about 93% to about 94%, about 90% to about 93%, about 91% to about 93%, about 92% to about 93%, about 90% to about 92%, about 91% to about 92% or about 90% to about 91%.

[0232] In some embodiments, protease cleavage causes the target polypeptide to be released from the recombinant fusion protein. In some embodiments, the recombinant fusion protein is correctly cut to correctly release the target polypeptide. In some embodiments, the correct cutting of the recombinant fusion protein forms a correctly released target polypeptide with a complete (undegraded) N-terminus. In some embodiments, the correct cutting of the recombinant fusion protein forms a correctly released target polypeptide containing the first (N-terminal) amino acid. In some embodiments, the amount of polypeptide correctly released after protease cleavage is about 90% to about 100%. In some embodiments, the amount of polypeptide correctly released after protease cleavage is about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, about 100%, about 91% to about 100%, about 92% to about 100%, about 93% to about 100%, about 94% to about 100%, about 95% to about 100%, about 96% to about 100%, about 97% to about 100%, about 98% to about 100%, about 99% to about 100%, about 90% to about 99%, about 91% to about 99%, about 92% to about 99%, about 93% to about 99%, about 94% to about 99%, about 95% to about 99%, about 96% to about 99%, about 97% to about 99%, about 98% to about 99%, about 90% to about 98%, about 91% to about 98%, about 92% to about 98%, about 93% to about 98%, about 94% to about 98%, about 95% to about 98%, about 96% to about 98%, about 97% to about 98%, about 90% to about 97%, about 91% to about 97%, about 92% to about 97%, about 93% to about 97%, about 94% to about 97%, about 95% to about 97%, about 96% to about 97%, about 90% to about 96%, about 91% to about 96%, about 92% to about 96%, about 93% to about 96%, about 94% to about 96%, about 97% to about 97%, about 98% to about 98%, about 99% to about 99%, about 90% to about 99%, about 91% to about 96%, about 92% to about 96%, about 93% to about 96%, about 94% to about 96%, about 97% to about 97%, about 98% to about 99%, about 99% to about 99%, about 90% to about 99%, about 91% to about 96%, about 92% to about 96%, about 93% to about 96%, about 94% to about 96%, about 96% to about 96%, about 97% to about 97%, about 95% to about 97%, about 96% to about 97%, about 96% to about 96%, about 96% to about 96%, about 97% to about 97%, about 95% to about 97%, about 95% to about 97%, about 95% to about 97%, about 96% to about 97%, about 96% to about 96%, about 96% to about 96%, about 96% to about 96%, about 96% to about 96%, 5% to about 96%, about 90% to about 95%, about 91% to about 95%, about 92% to about 95%, about 93% to about 95%, about 94% to about 95%, about 90% to about 94%, about 91% to about 94%, about 92% to about 94%, about 93% to about 94%, about 90% to about 93%, about 91% to about 93%, about 92% to about 93%, about 90% to about 92%, about 91% to about 92% or about 90% to about 91%.

[0233] Recombinant fusion protein evaluation and yield

[0234] The resulting fusion protein and / or polypeptide of interest is characterized in any suitable manner using any suitable assay known in the art or described in the literature for characterizing proteins (eg, for assessing protein yield or quality).

[0235] In some embodiments, LC-MS known in the art or any other suitable method is used to monitor proteolytic shearing (clipping), deamidation, oxidation and fracture, and for confirming that the N-terminus of target polypeptide after linker cutting is complete. The productive rate of recombinant fusion protein or target polypeptide can be measured by methods well known to those skilled in the art, for example, by SDS-PAGE, capillary gel electrophoresis (CGE) or Western blot analysis. In some embodiments, ELISA method is used to measure host cell protein. For example, host protein (HCP) ELISA can be carried out using " Immunoenzymetric Assay for the Measurement of Pseudomonasfluorescens Host Cell Proteins " kit, catalog number F450, from Cygnus Technologies, Inc., according to the manufacturer's experimental protocol. Flat plate can be read on SPECTRAmax Plus (Molecular Devices), using Softmax Pro v3.1.2 software.

[0236] SDS-CGE was performed using a LabChip GXII instrument (Caliper Life Sciences, Hopkinton, MA) with HT Protein Express v2 chips and the corresponding reagents (part numbers 760499 and 760328, Caliper Life Sciences). Samples were prepared according to the manufacturer's protocol (Protein User Guide Document No. 450589, Rev. 3) and electrophoresed on polyacrylamide gels. After separation, the gels were stained, destained, and digitally imaged.

[0237] The concentration of a protein, such as a purified recombinant fusion protein or a target polypeptide as described herein, can be determined by absorbance spectroscopy by methods known to those skilled in the art and described in the literature. In some embodiments, the absorbance of a protein sample at 280 nm is measured (e.g., using an Eppendorf BioPhotometer, Eppendorf, Hamburg, Germany), and the protein concentration is calculated using the Beer-Lambert law. The concentration of a protein can be determined by known methods, such as, for example, Grimsley, GR and Pace, CN, "Spectrophotometric Determination of Protein Concentration," see Current Protocols in Protein Science 3.1.1-3.1.9, 2003 by John Wiley & Sons, Inc., incorporated herein by reference, for calculating the accurate molar absorption coefficient of proteins.

[0238] Table 5 lists the A values ​​at 1 determined using the molar extinction coefficients calculated by Vector NT1, Invitrogen. 280 The concentrations of proteins described herein were as follows.

[0239] Table 5: A of 1 280 Protein concentration

[0240]

[0241]

[0242]

[0243]

[0244]

[0245]

[0246] By transferring the proteins separated on the SDS-PAGE gel to a nitrocellulose membrane and incubating the membrane with a monoclonal antibody specific for the target polypeptide, Western blot analysis for determining the yield or purity of the target polypeptide can be performed according to any suitable method known in the art. Antibodies useful for any analytical method described herein can be generated by suitable procedures known to those skilled in the art.

[0247] Activity assays as described herein and known can also provide information about protein yield.In some embodiments, these or any other methods known in the art are used to assess correct processing of a protein, e.g., correct cleavage of a secretory leader sequence.

[0248] Useful measures of recombinant fusion protein productivity include, for example, the amount of soluble recombinant fusion protein per culture volume (e.g., grams or milligrams of protein per liter of culture), the percentage or fraction of the soluble recombinant fusion protein obtained (e.g., the amount of soluble recombinant fusion protein / total amount of recombinant fusion protein), the percentage or fraction of total cell protein (tcp) and the percentage or ratio of dry biomass. In some embodiments, the measurement of the recombinant fusion protein productivity as described herein is based on the amount of the soluble recombinant fusion protein obtained. In some embodiments, the measurement of the soluble recombinant fusion protein is carried out in the soluble portion obtained after cell lysis, for example, the soluble portion obtained after one or more centrifugation steps or after the purification of the recombinant fusion protein.

[0249] Useful measures of target polypeptide yield include, for example, the amount of soluble target polypeptide obtained per culture volume (e.g., grams or milligrams of protein per liter of culture), the percentage or fraction of soluble target polypeptide obtained (e.g., amount of soluble target polypeptide / amount of total target polypeptide), the percentage or fraction of active target polypeptide obtained (e.g., amount of active target polypeptide in an activity assay / total amount of target polypeptide), the percentage or fraction of total cellular protein (tcp), and the percentage or ratio of dry biomass.

[0250] In embodiments where productivity is expressed according to culture volume, culture cell density can be considered, particularly when comparing productivity between different cultures. In some embodiments, the method of the present invention can be used to obtain a recombinant fusion protein productivity of solubility and / or activity and / or correct processing (e.g., a secretion leader sequence with correct cleavage) of about 0.5 grams per liter to about 25 grams per liter. In some embodiments, the recombinant fusion protein comprises an N-terminal fusion partner, which is a cytoplasmic chaperone protein or a folding regulator from a family of heat shock proteins and, after expression, the fusion protein is directed to the cytoplasm. In some embodiments, the recombinant fusion protein comprises an N-terminal fusion partner, which is a cytoplasmic chaperone protein or a folding regulator from a family of periplasmic peptidylprolyl isomerases and, after expression, the fusion protein is directed to the periplasm. In some embodiments, the yield of the fusion protein (either a cytoplasmically expressed fusion protein or a periplasmically expressed fusion protein) is about 0.5 g / L, about 1 g / L, about 1.5 g / L, about 2 g / L, about 2.5 g / L, about 3 g / L, about 3.5 g / L, about 4 g / L, about 4.5 g / L, about 5 g / L, about 6 g / L, about 7 g / L, about 8 g / L, about 9 g / L, about 10 g / L, about 11 g / L, about 12 g / L, about 13 g / L, about 14 g / L, about 15 g / L, about 16 g / L. / L, about 17g / L, about 18g / L, about 19g / L, about 20g / L, about 21g / L, about 22g / L, about 23g / L, about 24g / L, about 25g / L, about 0.5g / L to about 25g / L, about 0.5g / L to about 23g / L, about 1g / L to about 23g / L, about 1.5g / L to about 23g / L, about 2g / L to about 23g / L, about 2.5g / L to about 23g / L, about 3g / L to about 23g / L, about 3.5g / L to about 23g / L , about 4g / L to about 23g / L, about 4.5g / L to about 23g / L, about 5g / L to about 23g / L, about 6g / L to about 23g / L, about 7g / L to about 23g / L, about 8g / L to about 23g / L, about 9g / L to about 23g / L, about 10g / L to about 23g / L, about 15g / L to about 23g / L, about 20g / L to about 23g / L, about 0.5g / L to about 20g / L, about 1g / L to about 20g / L, about 1.5g / L to about 20g / L, About 2g / L to about 20g / L, about 2.5g / L to about 20g / L, about 3g / L to about 20g / L, about 3.5g / L to about 20g / L, about 4g / L to about 20g / L, about 4.5g / L to about 20g / L, about 5g / L to about 20g / L, about 6g / L to about 20g / L, about 7g / L to about 20g / L, about 8g / L to about 20g / L, about 9g / L to about 20g / L, about 10g / L to about 20g / L, about 15g / L to about 20g / L, about 0.5g / L to about 15g / L, about 1g / L to about 15g / L, about 1.5g / L to about 15g / L, about 2g / L to about 15g / L, about 2.5g / L to about 15g / L, about 3g / L to about 15g / L, about 3.5g / L to about 15g / L, about 4g / L to about 15g / L, about 4.5g / L to about 15g / L, about 5g / L to about 15g / L, about 6g / L to about 15g / L, about 7g / L to about 15g / L, about 8g / L to about 15g / L, about 9g / L to about 15g / L, about 10g / L to about 15g / L, about 0.5g / L to about 12g / L, about 1g / L to about 12g / L, about 1.5g / L to about 12g / L, about 2g / L to about 12g / L, about 2.5g / L to about 12g / L, about 3g / L to about 12g / L, about 3.5g / L to about 12g / L, about 4g / L to about 12g / L, about 4.5g / L to about 12g / L, about 5g / L to about 12g / L, about 6g / L to about 12g / L, about 7g / L to about 12g / L, about 8g / L to about 12g / L, about 9g / L to about 12g / L, about 10g / L to about 12g / L, about 0.5g / L to about 10g / L, about 1g / L to about 10g / L, about 1.5g / L to about 10g / L, about 2g / L to about 10g / L, about 2.5g / L to about 10g / L, about 3g / L to about 1 0g / L, about 3.5g / L to about 10g / L, about 4g / L to about 10g / L, about 4.5g / L to about 10g / L, about 5g / L to about 10g / L, about 6g / L to about 10g / L, about 7g / L to about 10g / L, about 8g / L to about 10g / L, about 9g / L to about 10g / L, about 0.5g / L to about 9g / L, about 1g / L to about 9g / L, about 1.5g / L to about 9g / L, about 2g / L to about 9g / L, about 2.5g / L to about 9g / L, about 3g / L to about 9g / L, about 3.5g / L to about 9g / L, about 4g / L to about 9g / L, about 4.5g / L to about 9g / L, about 5g / L to about 9g / L, about 6g / L to about 9g / L, about 7g / L to about 9g / L, about 8g / L to about 9g / L, about 0.5g / L to about 8g / L, about 1g / L to about 8g / L, about 1.5g / L to about 8g / L, about 2g / L to about 8g / L, about 2.5g / L to about 8g / L, about 3g / L to about 8g / L, about 3.5g / L to about 8g / L, about 4g / L to about 8g / L, about 4.5g / L to about 8g / L, about 5g / L to about 8g / L, about 6g / L to about 8g / L, about 7g / L to about 8g / L, about 0.5g / L to about 7g / L, about 1g / L to about 7g / L, about 1.5g / L to about 7g / L, about 2g / L to about 7g / L, about 2.5g / L to about 7g / L, about 3g / L to about 7g / L, about 3.5g / L to about 7g / L, about 4g / L to about 7g / L, about 4.5g / L to about 7g / L, about 5g / L to about 7g / L, about 6g / L to about 7g / L, about 0.5g / L to about 6g / L, about 1g / L to about 6g / L, about 1.5g / L to about 6g / L, about 2g / L to about 6g / L, about 2.5g / L to about 6g / L, about 3g / L to about 6g / L, about 3.5g / L to about 6g / L, about 4g / L to about 6g / L, about 4.5g / L to about 6g / L, about 5g / L to about 6g / L, about 0.5g / L to about 5g / L, about 1g / L to about 5g / L, about 1.5g / L to about 5g / L, about 2g / L to about 5g / L, about 2.5g / L to about 5g / L, about 3g / L to about 5g / L, about 3.5g / L to about 5g / L, about 4g / L to about 5g / L, about 4.5g / L to about 5g / L, about 0.5g / L to about 4g / L, about 1g / L to about 4g / L, about 1.5g / L to about 4g / L, about 2g / L to about 4g / L, about 2.5g / L to about 4g / L, about 3g / L to about 4g / L, about 0.5g / L to about 3g / L, about 1g / L to about 3g / L, about 1.5g / L to about 3g / L, about 2g / L to about 3g / L, about 0.5g / L to about 2g / L, about 1g / L to about 2g / L or about 0.5g / L to about 1g / L. .

[0251] In some embodiments, the polypeptide of interest is hPTH and the yield of the recombinant fusion protein directed to the cytoplasm is about 0.5 g / L to about 2.4 g / liter.

[0252] In some embodiments, the polypeptide of interest is hPTH and the yield of the recombinant fusion protein directed to the periplasm is about 0.5 g / L to about 6.7 g / liter.

[0253] Yield of target peptide

[0254] In some embodiments, the target polypeptide is released from the fully recombinant fusion protein by protease cleavage within the linker. In some embodiments, the target polypeptide obtained after cleavage with the protease is a correctly released target polypeptide. In some embodiments, the yield of the target polypeptide - based on measurement of correctly released protein or calculated based on the ratio of known target polypeptide to total fusion protein - is about 0.7 g / L to about 25.0 g / L. In some embodiments, the yield of the target polypeptide is about 0.5 g / L (500 mg / L), about 1 g / L, about 1.5 g / L, about 2 g / L, about 2.5 g / L, about 3 g / L, about 3.5 g / L, about 4 g / L, about 4.5 g / L, about 5 g / L, about 6 g / L, about 7 g / L, about 8 g / L, about 9 g / L, about 10 g / L, about 11 g / L, about 12 g / L, about 13 g / L, about 14 g / L, about 15 g / L, about 16 g / L, about 17 g / L, about 18 g / L, about 19 g / L, about 20 g / L, about 21 g / L, about 22 g / L, about 23 g / L, about 24 g / L, about 25 g / L, about 26 g / L, about 27 g / L, about 28 g / L, about 29 g / L, about 30 g / L, about 31 g / L, about 32 g / L, about 33 g / L, about 34 g / L, about 35 g / L, about 36 g / L, about 37 g / L, about 38 g / L, about 39 g / L, about 40 g / L, about 41 g / L About 17g / L, about 18g / L, about 19g / L, about 20g / L, about 21g / L, about 22g / L, about 23g / L, about 24g / L, about 25g / L, about 0.5g / L to about 23g / L, about 1g / L to about 23g / L, about 1.5g / L to about 23g / L, about 2g / L to about 23g / L, about 2.5g / L to about 23g / L, about 3g / L to about 23g / L, about 3.5g / L to about 23g / L, about 4g / L to about 23g / L, about 4.5g / L to about 23g / L, about 5g / L to about 23g / L, about 6g / L to about 2 3g / L, about 7g / L to about 23g / L, about 8g / L to about 23g / L, about 9g / L to about 23g / L, about 10g / L to about 23g / L, about 15g / L to about 23g / L, about 20g / L to about 23g / L, about 0.5g / L to about 20g / L, about 1g / L to about 20g / L, about 1.5g / L to about 20g / L, about 2g / L to about 20g / L, about 2.5g / L to about 20g / L, about 3g / L to about 20g / L, about 3.5g / L to about 20g / L, about 4g / L to about 20g / L, about 4.5g / L to about 20g / L, about 5g / L to about 20g / L, about 6g / L to about 20g / L, about 7g / L to about 20g / L, about 8g / L to about 20g / L, about 9g / L to about 20g / L, about 10g / L to about 20g / L, about 15g / L to about 20g / L, about 0.5g / L to about 15g / L, about 1g / L to about 15g / L, about 1.5g / L to about 15g / L, about 2g / L to about 15g / L, about 2.5g / L to about 15g / L, about 3g / L to about 15g / L, about 3.5g / L to about 15g / L, about 4g / L to about 15g / L, about 4.5g / L to about 15g / L, about 5g / L to about 15g / L, about 6g / L to about 15g / L, about 7g / L to about 15g / L, about 8g / L to about 15g / L, about 9g / L to about 15g / L, about 10g / L to about 15g / L, about 0.5g / L to about 12g / L, about 1g / L to about 12g / L, about 1.5g / L to about 12g / L, about 2g / L to about 12g / L, about 2.5g / L to about 12g / L, about 3g / L to about 12g / L, about 3.5g / L to about 12g / L, about 4g / L to about 12g / L, about 4.5g / L to about 12g / L, about 5g / L to about 12g / L, about 6g / L to about 12g / L, about 7g / L to about 12g / L, about 8g / L to about 12g / L, about 9g / L to about 12g / L, about 10g / L to about 12g / L, about 0.5g / L to about 10g / L, about 1g / L to about 10g / L, about 1.5g / L to about 10g / L, about 2g / L to about 10g / L, about 2.5g / L to about 10g / L, about 3g / L to about 10g / L, about 3.5g / L to about 10g / L, about 4g / L to about 10g / L, about 4.5g / L to about 10g / L, about 5g / L to about 10g / L, about 6g / L to about 10g / L, about 7g / L to about 10g / L, about 8g / L to about 10g / L, about 9g / L to about 10g / L, about 0.5g / L to about 9g / L, about 1g / L to about 9g / L, about 1.5g / L to about 9g / L, about 2g / L to about 9g / L, about 2.5g / L to about 9g / L, about 3g / L to about 9g / L, about 3.5g / L to about 9g / L, about 4g / L to about 9g / L, about 4.5g / L to about 9g / L, about 5g / L to about 9g / L, about 6g / L to about 9g / L, about 7g / L to about 9g / L, about 8g / L to about 9g / L, about 0.5g / L to about 8g / L, about 1g / L to about 8g / L, about 1.5g / L to about 8g / L, about 2g / L to about 8g / L, about 2.5g / L to about 8g / L, about 3g / L to about 8g / L, about 3 .5g / L to about 8g / L, about 4g / L to about 8g / L, about 4.5g / L to about 8g / L, about 5g / L to about 8g / L, about 6g / L to about 8g / L, about 7g / L to about 8g / L, about 0.5g / L to about 7g / L, about 1g / L to about 7g / L, about 1.5g / L to about 7g / L, about 2g / L to about 7g / L, about 2.5g / L to about 7g / L, about 3g / L to about 7g / L, about 3.5g / L to about 7g / L, about 4g / L to about 7g / L, about 4.5g / L to about 7g / L, about 5g / L to about 7g / L, about 6g / L to about 7g / L, about 0.5g / L to about 6g / L, about 1g / L to about 6g / L, about 1.5g / L to about 6g / L, about 2g / L to about 6g / L, about 2.5g / L to about 6g / L, about 3g / L to about 6g / L, about 3.5g / L to about 6g / L, about 4g / L to about 6g / L, about 4.5g / L to about 6g / L, about 5g / L to about 6g / L, about 0.5g / L to about 5g / L, about 1g / L to about 5g / L, about 1.5g / L to about 5g / L, about 2g / L to about 5g / L, about 2.5g / L to about 5g / L, about 3g / L to about 5g / L, about 3.5g / L to about 5g / L, about 4g / L to about 5g / L, about 4.5g / L to about 5g / L, about 0.5g / L to about 4g / L, about 1g / L to about 4g / L, about 1.5g / L to about 4g / L, about 2g / L to about 4g / L, about 2.5g / L to about 4g / L, about 3g / L to about 4g / L, about 0.5g / L to about 3g / L, about 1g / L to about 3g / L, about 1.5g / L to about 3g / L, about 2g / L to about 3g / L, about 0.5g / L to about 2g / L, about 1g / L to about 2g / L or about 0.5g / L to about 1g / L. .

[0255] In some embodiments, hPTH is produced as a fusion protein with an N-terminal fusion partner and hPTH constructs as described in Table 8. In some embodiments, expression of the hPTH fusion protein produces at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, or at least 1000 mg / L of total hPTH fusion protein at a scale of 0.5 mL to 100 L, 0.5 mL, 50 mL, 100 mL, 1 L, 2 L, or more.

[0256] In some embodiments, proinsulin, e.g., for an insulin analog (e.g., insulin glargine), is produced as a proinsulin fusion protein with an N-terminal fusion partner and a proinsulin construct comprising a C-peptide sequence as described in Table 19. In some embodiments, expression of a proinsulin fusion protein according to the methods of the invention produces at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 110, at least about 120, at least about 130, at least about 140, at least about 150, at least about 200, or at least about 250 mg / L of soluble proinsulin at a scale of 0.5 mL to 100 L, 50 mL, 100 mL, 1 L, 2 L, or more, as measured upon proper release or calculated based on a known ratio of the fusion protein.

[0257] In some embodiments, expression of a proinsulin fusion protein according to the methods of the present invention produces about 10 to about 500, about 15 to about 500, about 20 to about 500, about 30 to about 500, about 40 to about 500, about 50 to about 500, about 60 to about 500, about 70 to about 500, about 80 to about 500, about 90 to about 500, or about 100 to about 1000. 0, about 100 to about 500, about 200 to about 500, about 10 to about 400, about 15 to about 400, about 20 to about 400, about 30 to about 400, about 40 to about 400, about 50 to about 400, about 60 to about 400, about 70 to about 400, about 80 to about 400, about 90 to about 400, about 100 to about 400, about 200 to about 400, about 10 to about 300, about 15 to about 300, about 20 to about 300, about 30 to about 300, about 40 to about 300, about 50 to about 300, about 60 to about 300, about 70 to about 300, about 80 to about 300, about 90 to about 300, about 100 to about 300, about 200 to about 300, about 10 to about 250, about 15 to about 250, about 20 to about 250, about 30 to about 250, about 40 to about 250, about 50 to about 250, about 60 to about 250, about 70 to about 250, about 80 to about or about 100 to about 200 mg / L of soluble proinsulin, as measured when properly released or calculated based on the known ratio of the fusion protein thereof.

[0258] In some embodiments, expression of the proinsulin fusion protein produces at least about 100, at least about 125, at least about 150, at least about 175, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 550, at least about 600, at least about 650, or at least about 1000 mg / L of total soluble and insoluble proinsulin. In some embodiments, expression of the proinsulin fusion protein produces about 100 to about 2000 mg / L, about 100 to about 1500 mg / L, about 100 to about 1000 mg / L, about 100 to about 900 mg / L, about 100 to about 800 mg / L, about 100 to about 700 mg / L, about 100 to about 600 mg / L, about 100 to about 500 mg / L, about 100 to about 400 mg / L, about 200 to about 2000 mg / L, about 200 In some embodiments, the amount of proinsulin is from about 1 to about 1500 mg / L, from about 200 to about 1000 mg / L, from about 200 to about 900 mg / L, from about 200 to about 800 mg / L, from about 200 to about 7000 mg / L, from about 200 to about 600 mg / L, from about 200 to about 500 mg / L, from about 300 to about 2000 mg / L, from about 300 to about 1500 mg / L, from about 300 to about 1000 mg / L, from about 300 to about 900 mg / L, from about 300 to about 800 mg / L, from about 300 to about 7000 mg / L or from about 300 to about 600 mg / L of total soluble and insoluble proinsulin. In some embodiments, the proinsulin is cleaved to release the C-peptide and produce mature insulin. In some embodiments, expression of the proinsulin fusion protein produces at least about 100, at least about 200, at least about 250, at least about 300, at least about 400, at least about 500, about 100 to about 2000 mg / L, about 200 to about 2000 mg / L, about 300 to about 2000 mg / L, about 400 to about 2000 mg / L, about 500 to about 2000 mg / L, about 100 to about 1000 mg / L, about 200 to about 1000 mg / L, about 300 to about 1000 mg / L, about 400 to about 1000 mg / L, about 500 to about 1000 mg / L of mature insulin at a scale of 0.5 mL to 100 L, 0.5 mL, 50 mL, 100 mL, 1 L, 2 L or more, as measured when properly released or calculated based on the known ratio of the fusion protein thereof.

[0259] In some embodiments, GCSF is produced as a GCSF fusion protein with an N-terminal fusion partner as described in Table 21. In some embodiments, expression of a GCSF fusion according to the methods of the present invention at a scale of 0.5 mL to 100 L, 0.5 mL, 50 mL, 100 mL, 1 L, 2 L, or more produces a soluble fusion protein comprising at least 100, at least 200, at least 250, at least 300, at least 400, at least 500, or at least 1000, about 100 to about 1000, about 200 to about 1000, about 300 to about 1000, about 400 to about 1000, or about 500 to about 1000 mg / L of soluble GCSF as measured upon proper release or calculated based on a known ratio of the fusion protein thereof. In some embodiments, expression of a GCSF fusion according to the methods of the present invention produces at least 100, at least 200, at least 250, at least 300, at least 400, at least 500, or at least 1000 mg / L of soluble GCSF. In some embodiments, expression of a GCSF fusion produces at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 850, at least 550, at least 600, at least 650, about 100 to about 1000, about 200 to about 1000, about 300 to about 1000, about 400 to about 1000, or about 500 to about 1000 mg / L of total soluble and insoluble GCSF at a scale of 0.5 mL to 100 L, 0.5 mL, 50 mL, 100 mL, 1 L, 2 L, or more.

[0260] In some embodiments, the amount of recombinant fusion protein produced is from about 1% to about 75% of the total cellular protein. In specific embodiments, the amount of recombinant fusion protein produced is from about 1%, about 2%, about 3%, about 4%, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 1% to about 5%, about 1% to about 10%, about 1% to about 20%, about 1% to about 30%, about 1% to about 40%, about 1% to about 50 ...50%, about 1% to about 10%, about 1% to about 20%, about 1% to about 30%, about 1% to about 40%, about %, about 1% to about 60%, about 1% to about 75%, about 2% to about 5%, about 2% to about 10%, about 2% to about 20%, about 2% to about 30%, about 2% to about 40%, about 2% to about 50%, about 2% to about 60%, about 2% to about 75%, about 3% to about 5%, about 3% to about 10%, about 3% to about 20%, about 3% to about 30%, about 3% to about 40%, about 3% to about 50%, about 3% to about 60%, about 3% to about 75%, about 4% to about 10 %, about 4% to about 20%, about 4% to about 30%, about 4% to about 40%, about 4% to about 50%, about 4% to about 60%, about 4% to about 75%, about 5% to about 10%, about 5% to about 20%, about 5% to about 30%, about 5% to about 40%, about 5% to about 50%, about 5% to about 60%, about 5% to about 75%, about 10% to about 20%, about 10% to about 30%, about 10% to about 40%, about 10% to about 50%, about 10% to about 60%, From about 10% to about 75%, from about 20% to about 30%, from about 20% to about 40%, from about 20% to about 50%, from about 20% to about 60%, from about 20% to about 75%, from about 30% to about 40%, from about 30% to about 50%, from about 30% to about 60%, from about 30% to about 75%, from about 40% to about 50%, from about 40% to about 60%, from about 40% to about 75%, from about 50% to about 60%, from about 50% to about 75%, from about 60% to about 75% or from about 70% to about 75%.

[0261] Solubility and activity

[0262] The "solubility" and "activity" of a protein are often determined in different ways, although there are also related qualities. The solubility of a protein, particularly a hydrophobic protein, indicates that the hydrophobic amino acid residues are not correctly located on the outside of the folded protein. Protein activity (which can be assessed by one skilled in the art using methods determined to be applicable to the polypeptide of interest) is another indicator of correct protein conformation. As used herein, "soluble, active, or both" refers to a protein that is determined to be soluble, active, or both soluble and active by methods known to those skilled in the art.

[0263] Generally, about amino acid sequence, the term "modification" includes displacement, insertion, extension, deletion and derivatization alone or in combination. In some embodiments, recombinant fusion protein can include one or more modifications of "non-essential" amino acid residues. In this case, "non-essential" amino acid residues are residues that can change (for example, delete or replace) in the new amino acid sequence without eliminating or substantially reducing the activity (for example, agonist activity) of the recombinant fusion protein. For example, recombinant fusion protein can include 1,2,3,4,5,6,7,8,9,10 or more displacements, in a continuous manner or at intervals throughout the recombinant fusion protein molecule. Alone or in combination with displacement, recombinant fusion protein can include 1,2,3,4,5,6,7,8,9,10 or more insertions, in a continuous manner or at intervals throughout the recombinant fusion protein molecule. Alone or in combination with displacement and / or insertion, recombinant fusion protein can also include 1,2,3,4,5,6,7,8,9,10 or more deletions, in a continuous manner or at intervals throughout the recombinant fusion protein molecule. The recombinant fusion protein may also include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acid additions, alone or in combination with substitutions, insertions and / or deletions.

[0264] Substitution includes conservative amino acid substitutions." conservative amino acid substitutions " are substitutions in which amino acid residues are replaced by amino acids with similar side chains or physicochemical characteristics (e.g., electrostatic, hydrogen bonding, isostericity, hydrophobicity). Amino acids can be naturally occurring or non-mature (normatural) (non-natural). Families of amino acid residues with similar side chains are known in the art. These families include amino acids with basic side chains (e.g., lysine, arginine, histidine), amino acids with acidic side chains (e.g., aspartic acid, glutamic acid), uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, methionine, cysteine), amino acids with non-polar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, tryptophan), amino acids with β-branched side chains (e.g., threonine, valine, isoleucine) and amino acids with aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine). Substitutions can also include non-conservative changes.

[0265] Although preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. Various changes, variations, and substitutions can now be envisioned by those skilled in the art without departing from the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be used to implement the present invention. It is intended that the following claims define the scope of the invention and thus encompass methods and structures within the scope of these claims and their equivalents.

[0266] Example

[0267] Example I: High-throughput screening of strains expressing hPTH 1-34 fusions

[0268] This study was performed to test the levels of recombinant protein produced by P. fluorescens strains expressing hPTH 1-34 fusion proteins containing either DNAJ-like protein, Fk1B, or FrnE as N-terminal fusion partners.

[0269] Materials and methods

[0270] Construction of PTH 1-34 fusion protein expression plasmids: Gene fragments encoding PTH 1-34 fusion proteins were synthesized using DNA2.0, a gene design and synthesis service (Menlo Park, CA). Each gene fragment includes the coding sequence of a Pseudomonas fluorescens folding regulator (DnaJ-like, Fk1B, or FrnE) fused to the coding sequence for PTH 1-34 and a linker. Each gene fragment also includes recognition sequences for the restriction enzymes SpeI and XhoI, a "Hi" ribosome binding site, and an 18-base pair spacer containing a ribosome binding site and restriction site (SEQ ID NO: 58) added upstream of the coding sequence, as well as three stop codons. The nucleotide sequences encoding these PTH 1-34 fusion proteins are provided as SEQ ID NOs: 52-57.

[0271] To generate expression plasmids p708-004, -005, and -006 (listed in Table 6), the PTH 1-34 fusion protein gene fragments were digested with Spel and Xhol restriction enzymes and subcloned into the expression vector pDOW1169, which contains the pTac promoter and the rrnT1T2 transcription terminator. pDOW1169 is described in the literature, for example, in U.S. Patent Application Publication No. 7,833,752, "Bacterial Leader Sequences for Increased Expression" and Schneider et al., 2005, "Auxotrophic markers pyrF and proC can replace antibiotic markers on protein production plasmids in high-cell-density Pseudomonas fluorescensfermentation," Biotechnol. Progress 21(2):343-8, both of which are incorporated herein by reference. The plasmid was electroporated into competent Pseudomonas fluorescens DC454 host cells (pyrF lsc::lacI Q1 )middle.

[0272] Table 6: PTH 1-34 fusion protein plasmid

[0273] Plasmid number N-terminal fusion partner Fusion protein p708-004 DnaJ-like proteins DnaJ-like protein-PTH p708-005 FWf FklB-PTH p708-006 FWf FrnE-PTH

[0274] DNA Sequencing: Use The presence of cloned fragments in the fusion protein expression plasmid was confirmed by DNA sequencing using a Terminator v3.1 cycle sequencing kit (Applied Biosystems, 4337455). A DNA sequencing reaction containing 50 fmol of plasmid DNA to be analyzed was prepared by mixing 1 μL of sequencing premix, 0.5 μL of 100 μM primer stock solution, 3.5 μL of sequencing buffer, and water added to a final volume of 20 μL. Sequencher TM Software (Gene Codes) was used to combine and analyze the results.

[0275] 96-well growth and expression (HTP): The fusion protein expression plasmid was transformed into the Pseudomonas fluorescens host strain in an array format. The transformation reaction was initiated by mixing 35 μL of Pseudomonas fluorescens competent cells and 10 μL of plasmid DNA (2.5 ng). A 25 μL aliquot of the mixture was transferred to a 96-well multiwell plate. Plate (Lonza). Using NucleofectorTM 96-well Shuttle TM System (Lonza AG) was used for electroporation, and the electroporated cells were subsequently transferred to fresh 96-well deep-well plates containing 500 μL of M9 salts supplemented with 1% glucose medium and trace elements. The plates were incubated at 30° C. with shaking for 48 hours to generate seed cultures.

[0276] The seed culture of 10 μ L aliquots is transferred to 96-hole deep well plates in duplicate. Each well contains 500 μ L HTP-YE culture medium (Teknova), supplemented with trace elements and 5% glycerol. The seed culture placed in the HTP culture medium supplemented with glycerol is incubated in a shaking table at 30 DEG C for 24 hours. 0.3 mM final concentration of isopropyl-β-D-1-thiogalactoside (IPTG) is added to each well to induce the expression of PTH 1-34 fusion protein. For the bacterial strain containing the folding regulator overexpression plasmid (see Table 4), IPTG is supplemented with 1% final concentration of mannitol (Sigma, M1902) to induce the expression of the folding regulator. In addition, when inducing, 250 units / μ L of 0.01 μ L are added to each well to reduce the possibility of culture viscosity. After 24 hours of induction, the optical density (OD) at 600 nm is measured. 600 ) was used to calculate the cell density. The cells were then collected and diluted 1:3 with 1× phosphate-buffered saline (PBS) to a final volume of 400 μL and frozen for later processing.

[0277] Preparation of soluble lysate samples for analytical characterization: The collected cell samples were diluted and lysed by sonication using a cell lysis automated sonication system (CLASS, Scinomix) using a 24 probe tip. The lysate was centrifuged at 5,500 × g for 15 minutes at 8 ° C. The supernatant was collected and labeled as the soluble fraction. The precipitate was collected, resuspended in 400 μL 1× PBS pH 7.4 by another round of sonication, and labeled as the insoluble fraction.

[0278] SDS-CGE analysis: Soluble and insoluble fractions were analyzed by HTP microchip SDS capillary gel electrophoresis using a LabChip GXII instrument (Caliper LifeSciences) with HT Protein Express v2 chips and corresponding reagents (part numbers 760499 and 760328, Caliper LifeSciences). Samples were prepared according to the manufacturer's protocol (Protein User Guide Document No.450589, Rev.3). Briefly, 4 μL aliquots of the soluble or insoluble fraction sample were mixed with 14 μL buffer in a 96-well polypropylene conical well PCR plate heated at 95°C for 5 minutes (with or without dithiothreitol (DTT) reducing agent) and diluted with 70 μL deionized water. Lysates from a blank host strain not transformed with the fusion protein expression plasmid were run in parallel with the test samples as a control and quantified using the system's internal standard.

[0279] Shake flask expression: A seed culture of each fusion protein expression strain to be evaluated was grown in M9 Glucose (Teknova) to generate an intermediate culture, and a 5 mL volume of each intermediate culture was used to inoculate one of four 1-liter baffled bottom flasks containing 250 mL of HTP medium (Teknova 3H1129). After growing for 24 hours at 30°C, the culture was induced with 0.3 mM IPTG and 1% mannitol and incubated at 30°C for an additional 24 hours. The flask culture was then centrifuged to collect the cells, and the collected cell paste was frozen for future use.

[0280] Mechanical Release and Purification: 5 or 10 gram quantities of frozen cell paste were thawed and resuspended in 3× PBS, 5% glycerol, 50 mM imidazole, pH 7.4, to a final volume of 50 mL or 100 mL, respectively. The suspension was then homogenized by passing it through a microfluidizer (Microfluidics, Inc., Model M110Y) twice at 15,000 psi. The lysate was centrifuged at 12,000 × g for 30 minutes and filtered through a Sartorius Sartobran 150 (0.45 / 0.2 μm) capsule filter.

[0281] Chromatography: using a Frac-950 fraction collector Explorer 100 chromatography system (GE Healthcare) carries out fast protein liquid chromatography (FPLC) operation.The soluble portion sample prepared from HTP expression culture fluid is loaded on 5mL HisTrap FF column (GE Healthcare, part number 17-5248-02) balanced with 3 × PBS, 5% glycerol, 50mM imidazole pH 7.4 in advance.The column is washed with 4 column volumes of equilibrium buffer, and 10 column volumes of elution buffer are used, and the imidazole linear gradient from 50mM to 200mM is applied, and the fusion protein is eluted from the HisTrap column.The whole process is run with 100cm / h, which is equal to 1.5 minutes residence time.Using the above-mentioned SDS-CGE analytical method, the purified fractions are analyzed by SDS-CGE.

[0282] Enterokinase cleavage: Using a 7000 molecular weight cutoff (MWCO) Slide-A-Lyzer cassette (Pierce), the first set of samples was prepared by dialyzing the purified fraction containing the fusion protein against 1×PBS pH 7.4 supplemented with 2mM CaCl at 4°C overnight. The dialyzed samples were maintained at a concentration of approximately 1 mg / mL. A second set of samples was prepared by diluting the purified fraction containing the fusion protein 2× with water and stored in a buffer containing 1.5×PBS, 2.5% glycerol, and ~30-70mM imidazole at a concentration of 0.5 mg / mL. A stock solution of porcine enterokinase (Sigma E0632-1.5KU) was added to the samples at 5× or 20× dilutions (corresponding to enterokinase concentrations of 40 μg / mL and 10 μg / mL, respectively). CaCl was also added to a final concentration of 2 mM, and the reaction mixture was incubated at room temperature overnight.

[0283] Liquid chromatography-mass spectrometry: A Q-ToF with an electrospray interface (ESI) was coupled to an Agilent 1100 HPLC equipped with an autosampler, column heater, and UV detector. micro A mass spectrometer (Waters) was used for liquid chromatography-mass spectrometry (LC-MS) analysis. A CN-reverse phase column with an inner diameter of 2.1 mm ID, a length of 150 mm, a particle size of 5 μm and a guard column (Agilent, catalog number 820950-923) was used. The pore size of 100 μg was 0.174 nm (Agilent, catalog number 883750-905). HPLC was run at a temperature of 50°C and the flow rate was maintained at 2°C. The HPLC buffer was 0.1% formic acid (mobile phase A) and 90% acetonitrile containing 0.1% formic acid (mobile phase B). Approximately 4 μg of fusion protein sample was loaded onto the HPLC column. The HPLC operating conditions were set to 95% mobile phase A when loading the sample. The fusion protein was eluted using the reversed-phase gradient illustrated in Table 7.

[0284] Table 7: Reversed-phase gradients used for mass spectrometry analysis of purified protein samples

[0285] time % Mobile phase A % mobile phase B Flow rate (ml / min) curve 0.0 95.0 5 0.2 -- 10.0 90 10 0.2 Linear 50.0 35 65 0.2 Linear 52.0 0 100.0 0.2 Linear 57.0 0 100.0 0.2 Keep 57.1 95.0 5.0 0.2 ladder 65.0 95.0 5.0 0.2 Keep

[0286] Before MS, UV absorption spectra were collected from 180 nm to 500 nm. The ESI-MS source was used in positive ion mode at 2.5 kV. MS scans were performed using a 600-2600 m / z range with 2 scans per second. MassLynx software (Waters) was used to analyze MS and UV data. UV chromatograms and MS total ion current (TIC) chromatograms were generated. The MS spectra of the target peaks were summarized. MaxEnt1 (Waters) was used to scan for a 2,800-6,000 molecular weight range (for PTH 1-34, it has a theoretical molecular weight of 4118 kDa, and for fusion proteins or N-terminal fusion proteins, it has a higher window), a resolution of 1 Da / channel, and a Gaussian width of 0.25 Da, and these spectra were deconvoluted.

[0287] result

[0288] Design of PTH 1-34 gene fusion fragments: To promote high-level expression of the PTH 1-34 fusion protein, three folding regulators from Pseudomonas fluorescens were selected based on high soluble expression, a molecular weight below 25 kDa, and an isoelectric point (pI) significantly different from that of PTH 1-34 (which has a pI of 8.52). The characteristics of the folding regulators are shown in Table 8. As shown in Table 8, the pIs of DnaJ-like protein, Fk1B, and FrnE are between 4.6 and 4.8, which are well separated from the pI of PTH 1-34. This allows for separation by ion exchange. To further aid purification of the fusion protein, a hexa-histidine tag was included in the linker. The linker also contains an enterokinase cleavage site (DDDDK) to facilitate separation of the N-terminal fusion partner from the desired PTH 1-34 target polypeptide. The amino acid sequence of the PTH 1-34 fusion protein is shown in Figure 2A (DnaJ-like protein, SEQ ID NO: 45), 2B (Fk1B-PTH, SEQ ID NO: 46) and 2C (FrnE-PTH, SEQ ID NO: 47). Figure 2A In A, B, and C, the amino acids corresponding to the linker are underlined, and those corresponding to PTH 1-34 are italicized.

[0289] Table 8: Physicochemical properties of selected N-terminal fusion partners

[0290]

[0291]

[0292] Construction of PTH Fusion Expression Vectors and HTP Expression: Synthetic gene fragments encoding each of the three PTH fusion proteins listed in Table 6 were synthesized using DNA 2.0. The synthesized gene fragments were digested with SpeI and XhoI and ligated into pDOW1169 (digested with the same enzymes) to generate expression plasmids p708-004, p708-005, and p708-006. After confirmation of the insert, the plasmids were used to electroporate a series of Pseudomonas fluorescens host strains and generate the expression strains listed in Table 4. The resulting transformed strains were grown and induced with IPTG and mannitol according to the procedures described in Materials and Methods. After induction, cells were harvested, sonicated, and centrifuged to separate the soluble and insoluble fractions. The soluble and insoluble fractions were collected. The soluble and insoluble fractions were analyzed using reduced SDS-CGE to measure the expression levels of the PTH 1-34 fusion proteins. A total of six strains, including two high HTP expressing strains for each of the three PTH 1-34 fusion proteins, were selected for shake flask expression. The strains screened using the shake flask expression method are listed in Table 9.

[0293] Shake flask expression: Each of the six strains was grown and induced at a 250 mL culture scale (4 x 250 mL cultures each) as described in the Materials and Methods (Shake flask expression) section. After induction, a sample from each culture (whole cell culture broth, WCB) was retained; a subset of the samples was diluted 3x with PBS, sonicated, and centrifuged to produce soluble and insoluble fractions. The remainder of each culture was centrifuged to produce a cell paste and a cell-free culture broth (CFB) supernatant. The cell paste was retained for purification. The WCB, CFB, and soluble fractions ( Figure 3 ).

[0294] Fusion proteins were observed in the WCB and soluble fractions (a band corresponding to a molecular weight of approximately 14 kDa for the DnaJ-like protein-PTH fusion and a band of approximately 26 kDa for the FrnE-PTH and Fk1B-PTH fusions); no fusion proteins were observed in the CFB. Shake flask expression titers for STR35984, STR36085, and STR36169 were 50% of the HTP expression titers, while shake flask expression titers for strains STR35970, STR36034, and STR36150 were 70-100% of the titers observed at the HTP scale. HTP and shake flask expression titers are listed in Table 9.

[0295] Table 9: Selected PTHs HTP and shake flask expression titers of 1-34 fusion protein expression strains

[0296]

[0297] IMAC purification of PTH fusion protein-expressing strains grown at HTP and shake flask scale to isolate PTH fusion proteins: Cell pastes from six strains were subjected to mechanical lysis and IMAC purification. Each purification run yielded highly enriched fractions. Peak fractions originating from the DnaJ-like protein-PTH-expressing strain STR35970 were 60-80% pure, those from the Fk1B-PTH-expressing strain STR36034 were 60-90% pure, and those from the FrnE-PTH-expressing strain STR36150 were 90-95% pure.

[0298] Enterokinase cleavage of PTH fusion proteins: Highly pure, concentrated fractions containing the fusion proteins from the IMAC purification run were selected for enterokinase cleavage reactions to demonstrate that the N-terminal fusion partner could be cleaved from PTH 1-34. Enterokinase of porcine origin was used for the studies. Since the 4 kDa PTH 1-34 target polypeptide was not readily detected by SDS-PAGE, a molecular weight shift of the total fusion protein from 14 kDa to 10 kDa for the DnaJ-like protein-PTH fusion protein and from 26 kDa to 22 kDa for the Fk1B-PTH and FrnE-PTH fusion proteins was accepted as evidence for enterokinase cleavage. Samples were treated overnight with either 40 μg / mL or 10 μg / mL enterokinase. Following enterokinase treatment, samples were analyzed by SDS-CGE. Figure 4 Complete cleavage of the fusion partner with PTH 1-34 was observed using 40 μg / mL enterokinase (lanes 7-12), and partial cleavage was observed using 10 μg / mL enterokinase (lanes 13-18), as shown by the MW shift in the Figure 5 compared to the uncleaved sample (lanes 1-6).

[0299] Intact mass analysis of PTH fusion protein after enterokinase cleavage: DnaJ-like protein-PTH fusion protein purified from strain STR35970 was used for additional enterokinase cleavage experiments and intact mass analysis. Purified fractions containing DnaJ-like protein-PTH fusion protein from STR35970 were incubated with porcine enterokinase at room temperature for 1 to 3 hours, followed by immediate intact mass analysis. Figure 5As shown in Table 10, the C-terminal PTH 1-34 polypeptide was detected. Details of the intact mass analysis are summarized in Table 10. In addition to full-length PTH 1-34, fragments corresponding to N-terminal deletions of 5 or 8 amino acids were also detected. The observed proteolysis is likely due to host cell protein contaminants or contaminants in the porcine enterokinase preparation. Recombinant enterokinase can also be used to evaluate cleavage using a similar procedure. The observed and theoretical molecular weights (MW) of the major species detected by intact mass analysis are shown in Table 10. The retention time of the uncleaved fusion protein was approximately 33 minutes, compared to an average retention time of 27 minutes for the fusion protein subjected to enterokinase cleavage for 1 to 3 hours.

[0300] Table 10: Complete quality results

[0301]

[0302]

[0303] Example II. Large-Scale Fermentation and Expression of PTH 1-34 Fusion Protein

[0304] The PTH 1-34 fusion proteins described in Example 1 were also evaluated for large-scale expression in Pseudomonas fluorescens to identify highly productive expression strains for large-scale PTH 1-34 production. The Pseudomonas fluorescens strains screened in this study were DnaJ-like protein-PTH fusion expression strains STR35970, STR35984, STR35949, STR36005, and STR35985; FklB-PTH fusion protein expression strains STR36034, STR36085, and STR36098; and FrnE-PTH fusion protein expression strains STR36150 and STR36169, which are listed in Tables 11 and 12.

[0305] Table 11: DnaJ-like protein-PTH fusion expression strains for large-scale fermentation

[0306] strain plasmids Host STR35949 p708-004 DC1084 STR35970 p708-004 DC508 STR35984 p708-004 DC992.1 STR35985 p708-004 PF1201.9 STR36005 p708-004 PF1326.1

[0307] Table 12: FrnE-PTH and Fk1B-PTH fusion expression strains for large-scale fermentation

[0308] strain plasmids Host STR36034 p708-005 DC1106 STR36085 p708-005 PF1326.1 STR36098 p708-005 PF1345.6 STR36150 p708-006 PF1219.9 STR36169 p708-006 PF1331

[0309] Materials and methods

[0310] MBR fermentation: the freezing culture stock of the selected bacterial strain of the shake flask inoculation containing the culture medium supplemented with yeast extract.For micro bioreactor (MBR), use the 250mL shake flask containing the chemical defined medium supplemented with yeast extract of 50mL.The shake flask culture was incubated at 30 ℃ of shaking for 16 to 24 hours.The aliquots from the shake flask culture were used to inoculate MBR (Pall Micro-24).MBR culture was operated with 4mL volume in each 10mL well of disposable micro bioreactor box under the controlled condition of pH, temperature and dissolved oxygen.When the glycerol of the initial content contained in the culture medium was exhausted, the culture was induced with IPTG.Fermentation lasted 16 hours, and samples were collected and frozen for analysis.

[0311] CBR fermentation: Inoculum for 1 liter CBR (conventional bioreactor) fermentor culture was generated by inoculating a shake flask containing 600 mL of chemically defined medium supplemented with yeast extract and glycerol with a frozen culture stock of the selected strain. After 16 to 24 hours of shaking incubation at 32° C., equal aliquots from each shake flask culture were aseptically transferred to each of an 8-unit multiple fermentation system containing a 2-liter bioreactor (1-liter working volume). The fed-batch high cell density fermentation process consisted of a growth phase followed by an induction phase that was initiated by the addition of IPTG once the culture reached the target optical density.

[0312] The induction phase of the fermentation was allowed to proceed for 8 hours and analytical samples were taken from the fermentor to determine the cell density at 575 nm (OD 575 ). Analytical samples were frozen for subsequent analysis to determine fusion protein expression levels. After 8 hours of induction, the entire fermentation broth from each vessel (approximately 0.8 L of broth per 2 L bioreactor) was collected by centrifugation at 15,900 × g for 60 to 90 minutes. The cell paste and supernatant were separated, and the cell paste was frozen at -80°C.

[0313] Mechanical homogenization and purification: Frozen cell paste (20 g) obtained from the CBR fermentation process as described above was thawed and resuspended in 20 mM sodium phosphate, 5% glycerol, 500 mM sodium chloride, 20 mM imidazole, pH 7.4. The final volume of the suspension was adjusted to ensure a solids concentration of 20%. The material was then homogenized by passing it through a microfluidizer (Microfluidics, Inc., Model M110Y) twice at 15,000 psi. The lysate was centrifuged at 12,000 × g for 30 minutes and filtered through a Sartorius Sartobran 150 (0.45 / 0.2 μm) capsule filter.

[0314] Chromatography: using a Frac-950 fraction collector Explorer 100 chromatography system (GE Healthcare) carries out fast protein liquid chromatography (FPLC) operation.Sample is loaded on HisTrap FF, on 10mL post (two 5mL HisTrap FF cartridges [GE Healthcare, part number 17-5255-01] connected in series), washing and by changing imidazole concentration from 0mM to 200mM, using the elution buffer elution of 10 column volumes linear gradient.Collected 2 ml volume fractions.

[0315] Immobilized metal ion affinity chromatography (IMAC) purification was performed using nickel IMAC (GE Healthcare, part number 17-5318-01). The analytical samples collected after CBR fermentation were separated into soluble and insoluble fractions. 600 μL aliquots of the soluble fraction were incubated on a shaker at room temperature for one hour with 100 μL IMAC resin and centrifuged at 12,000 × g for one minute to precipitate the resin. The supernatant was removed and labeled as the flow-through. The resin was then washed three times with 1 mL of a wash buffer containing 20 mM sodium phosphate pH 7.3, 500 mM NaCl, 5% glycerol, and 20 mM imidazole. After the third wash, the resin was resuspended in 200 μL of a wash buffer containing 400 mM imidazole and centrifuged. The supernatant was collected and labeled as the eluate.

[0316] Enterokinase cleavage: The PTH 1-34 fusion protein purified fractions were concentrated and resuspended in a buffer containing 20 mM Tris pH 7.4, 50 mM NaCl, and 2 mM CaCl. Two units of enterokinase (Novagen cat# 69066-3, batch D00155747) were added to 100 μg of protein in 100 μL reactions. The mixture of the fusion protein purified fractions and enterokinase was incubated at room temperature for one hour or overnight. A control reaction without enterokinase was also incubated at room temperature for one hour or overnight. The enzyme reaction was stopped by adding a complete protease inhibitor cocktail containing 4-benzenesulfonyl chloride (AEBSF, Sigma cat# P8465).

[0317] result

[0318] Fermentation evaluation of DnaJ-like protein-PTH, FklB-PTH and FrnE-PTH fusion expressing strains: For fermentation, the first five DnaJ-like protein-PTH fusion expressing strains, three FklB-PTH expressing strains and two FrnE-PTH expressing strains listed in Tables 9 and 10 were each evaluated first in a micro bioreactor (MBR) and then in a conventional bioreactor (CBR).

[0319] The soluble fraction from each MBR fermentation of the DnaJ-like protein-PTH fusion expressing strain was analyzed by SDS-CGE according to the experimental protocol described in the Materials and Methods section of Example 1. The MBR fermentation yields of the DnaJ-like protein-PTH fusion expressing strains are listed in Table 13. Overall, the strain with the highest MBR expression level of soluble fusion protein was STR35949, at 2.1 g / L.

[0320] Table 13: Soluble fusion protein yields of DnaJ-like protein-hPTH fusion strains tested in MBR fermentors

[0321] strain Soluble fusion protein yield STR35949 0.6-2.1g / L STR36005 1.5g / L STR35970 1.1g / L STR35985 0.9g / L

[0322] The DnaJ-like protein PTH fusion strain was evaluated in a conventional bioreactor (CBR) at a 1 L scale fermentation. The CBR expression levels of the DnaJ-like protein-PTH fusion strain were comparable to the MBR levels, as shown in Table 14. At the 8-hour post-induction time point, expression levels were higher than at the 24-hour post-induction time point.

[0323] Table 14: DnaJ-like-hPTH 1-34 evaluated in CBR fermenters 8 (I8) and 24 (I24) hours after induction Soluble fusion protein yield of fusion strain

[0324] strain Soluble fusion protein yield-(I8) Soluble fusion protein yield-(I24) STR35949 1.5-2.4g / L 1.1-1.9g / L STR35970 2.0g / L 0.9g / L STR35985 1.7-2.4g / L 0.3-0.6g / L STR36005 2.1g / L 1.4g / L

[0325] Soluble fractions from MBR fermentations of Fk1B-PTH and FrnE-PTH fusion expression strains were analyzed by SDS-CGE under reducing conditions (results are shown in Table 15).

[0326] Table 15: Solubility of Fk1B-hPTH 1-34 and FrnE-hPTH1-34 fusion strains evaluated in MBR fermentors Fusion protein yield

[0327] strain Soluble fusion protein yield STR36085 6.4g / L STR36034 3.4-5.8g / L STR36098 3.4-4.7g / L STR36150 0.8-2.2g / L

[0328] Overall, the strain with the highest expression level of soluble fusion protein was STR36034, at 6.4 g / L. The same strains were also evaluated for large-scale fermentation in conventional bioreactors (CBRs) (the results are shown in Table 16). In the CBR fermentation, the strain with the greatest yield was STR36034, which expressed the Fk1B-PTH fusion protein at 6.7 g / L after a 24-hour induction period.

[0329] Table 16: Fk1B-hPTH 1-34 and FrnE-hPTH evaluated in CBR fermenters 24 hours after induction (I24) Soluble fusion protein yield of 1-34 fusion strain

[0330] strain Soluble fusion protein yield (I24) STR36034 4.9-6.7g / L STR36085 4.6-4.9g / L STR36098 2.9-5.2g / L STR36150 2.6-3.8g / L

[0331] Purification of DnaJ-like protein-PTH and FklB-PTH fusion proteins and evaluation of enterokinase cleavage: As described in Materials and Methods, cell paste obtained after induction and growth of the DnaJ-like protein-PTH fusion expression strain STR36005 was subjected to mechanical lysis and IMAC purification. Highly enriched fractions were obtained in each purification run. Peak fractions had a purity of 90% or greater.

[0332] A high-purity concentrated fraction of the DnaJ-like protein-PTH fusion protein purified from strain 36005 was used in an enterokinase cleavage assay to confirm that the N-terminal fusion partner could cleave the PTH 1-34 target polypeptide. Recombinant bovine enterokinase was used for the cleavage reaction. The soluble fraction from the analytical-scale sample was used for small-scale batch enrichment of the fusion protein using IMAC resin ( Figure 6 After one hour of incubation with enterokinase, partial cleavage of the DnaJ-like protein fusion partner was observed. After overnight incubation, cleavage was complete (lanes 6-8).

[0333] The Fk1B-PTH fusion strain was shown to be robust at a 1-liter scale. Purified samples were further analyzed to confirm that the fusion protein could be enriched and cleaved with enterokinase. Soluble fractions from analytical-scale samples were used for small-scale batch enrichment of the fusion protein using IMAC resin. One enriched sample from each of the three expression strains, STR36034, STR36085, and STR36098, was treated with enterokinase and subjected to intact mass analysis using the method described in Example 1. For each sample, the PTH 1-34 target polypeptide was confirmed and observed to be the correct mass, ~4118 Da, as shown in Figure 7.

[0334] Example III. Construction of enterokinase fusions

[0335] DnaJ-like protein, Fk1B and FrnE N-terminal fusion partner-enterokinase fusion protein was designed and expression construct was generated for expression of recombinant enterokinase (SEQ ID NO: 31).

[0336] Construction of Enterokinase Fusion Expression Plasmids: The enterokinase (EK) fusion coding regions evaluated are listed in Table 17. The gene fragment encoding the fusion protein was synthesized by DNA2.0. This fragment included Spe1 and Xho1 restriction enzyme sites, a "Hi" ribosome binding site, an 18-base pair spacer (5'-actagtaggaggtctaga-3') added upstream of the coding sequence, and three stop codons.

[0337] Use standard cloning method to build expression plasmid.Use SpeI and XhoI restriction enzyme digestion to contain the plasmid DNA of each enterokinase fusion coding sequence, then subclone into the pDOW1169 expression vector containing pTac promotor and rrnT1T2 transcription terminator of SpeI-XhoI digestion.Connect insert fragment and carrier and spend the night with T4 DNA ligase (Fermentas EL0011), form enterokinase fusion protein expression plasmid.By plasmid electroporation in competent state Pseudomonas fluorescens DC454 host cell.By using Ptac and Term sequence primers (AccuStart II from Quanta, PCR SuperMix, 95137-500) PCR, for the existence screening positive clone of enterokinase fusion protein sequence insert fragment.

[0338] Table 17: Enterokinase fusion proteins

[0339]

[0340] Example IV. Large-Scale Fermentation of Enterokinase Fusion Proteins (DNAJ-like, Fk1B, FrnE N-terminal Chaperone)

[0341] The expression strains described in Example III were tested for expression of recombinant protein by HTP analysis following methods similar to those described in Example I.

[0342] Based on the level of soluble fusion protein expression, expression strains were selected for fermentation studies. Selected strains were grown and induced as described above for the PTH 1-34 fusion protein, and the induced cells were centrifuged, lysed, and centrifuged again. The resulting insoluble and soluble fractions were extracted using the above-described extraction conditions, and the EK fusion protein extract supernatant was quantified using SDS-CGE.

[0343] Example V. High-throughput screening of strains expressing insulin fusion proteins

[0344] This study was performed to test the levels of recombinant protein produced by P. fluorescens strains expressing proinsulin fusion proteins containing DNAJ-like protein, EcpD, Fk1B, FrnE or truncations of EcpD, Fk1B, FrnE as N-terminal fusion partners.

[0345] Materials and methods

[0346] Construction of proinsulin expression vectors: Optimized gene fragments encoding proinsulin (insulin glargine) were synthesized by DNA 2.0 (Menlo Park, CA). The gene fragments and the proinsulin amino acid sequences encoded by the proinsulin coding sequences contained within the gene fragments are listed in Table 18. Each gene fragment contains peptide A and B coding sequences, as well as one of four different glargine C-peptide sequences: CP-A (MW = 9336.94 Da; pI = 5.2; 65% A+B glargine), CP-B (MW = 8806.42 Da; 69% A+B glargine), CP-C (MW = 8749.32 Da; 69% A+B glargine), and CP-D (MW = 7292.67 Da; 83% A+B glargine). The gene fragments were designed to include SapI restriction enzyme sites upstream and downstream of the proinsulin coding sequence to enable rapid cloning of the gene fragments into various expression vectors. The gene segments also include a lysine amino acid codon (AAG) or an arginine amino acid codon (CGA) in the 5' flanking region to facilitate ligation into expression vectors containing an enterokinase cleavage site or a trypsin cleavage site, respectively. In addition, all gene segments include three stop codons (TGA, TAA, TAG) in the 3' flanking region.

[0347] Table 18: Proinsulin gene fragment and C-peptide amino acid sequence

[0348]

[0349] The coding sequence was then ligated into an expression vector using T4 DNA ligase (New England Biolabs, M0202S) and the proinsulin cloned sequence was subcloned into expression vectors containing different fusion partners (Table 19). The ligated vectors were electroporated into competent DC454 Pseudomonas fluorescens cells in a 96-well format.

[0350] Table 19: Vectors used for expression of proinsulin glargine fusion protein

[0351]

[0352]

[0353] Growth and expression in 96-well format (HTP): Transform the plasmid containing the proinsulin coding sequence and fusion partner into the Pseudomonas fluorescens DC454 host strain. Thaw 25 μl of competent cells and transfer them to a 96-well multiwell plate. Plate (Lonza VHNP-1001) and mixed with the ligation mixture prepared in the previous step. TM 96-well Shuttle TMSystem (Lonza AG) was used for electroporation, and the transformed cells were subsequently transferred to 96-well deep-well plates (seed plates) containing 400 μL of M9 salts 1% glucose medium and trace elements. The seed plates were incubated at 30° C. with shaking for 48 hours to generate seed cultures.

[0354] 10 microliters of seed culture were transferred in duplicate to a fresh 96-well deep-well plate containing 500 μL HTP medium (Teknova 3H1129) (supplemented with trace elements and 5% glycerol) in each well and incubated at 30°C for 24 hours. Isopropyl-β-D-1-thiogalactopyranoside (IPTG) was added to each well at a final concentration of 0.3 mM to induce the expression of the proinsulin fusion protein. In addition, upon induction, 0.01 μL of 250 units / μl reserve Benzonase (Novagen, 70746-3) was added to each well to reduce the potential for culture viscosity. 24 hours after induction, the optical density (OD) at 600 nm was measured. 600 ) were used to quantify cell density. 24 hours after induction, cells were collected and diluted 1:3 with 1× PBS to a final volume of 400 μL and then frozen for subsequent processing.

[0355] Preparation of soluble lysate samples for analytical characterization: Culture broth samples prepared and stored frozen as described above were thawed, diluted, and sonicated. The sonicated lysate was centrifuged at 5,500 × g for 15 minutes at 8°C to separate the soluble (supernatant) and insoluble (pellet) fractions. The insoluble fraction was resuspended in PBS using sonication.

[0356] SDS-CGE analysis: Test protein samples prepared as discussed above were analyzed by HTP microchip SDS capillary gel electrophoresis using a LabChip GXII instrument (PerkinElmer) with HT Protein Express v2 chips and the corresponding reagents (part numbers 760499 and 760328, PerkinElmer, respectively). Samples were prepared according to the manufacturer's protocol (Protein User Guide Document No. 450589, Rev. 3). In a 96-well conical well PCR plate, 4 μL of sample was mixed with 14 μL of buffer, with or without dithiothreitol (DTT) reducing agent. The mixture was heated at 95° C. for 5 minutes and diluted by adding 70 μL of deionized water.

[0357] Proinsulin titers at the 96-well scale were determined based on the fusion protein titer multiplied by the percentage of fusion protein consisting of proinsulin.Total titer represents the sum of soluble and insoluble target expression (mg / L).

[0358] result

[0359] As shown in Table 20, the glargine proinsulin fusion protein with DnaJ-sample albumen as the N-terminal fusion partner demonstrates the highest level of proinsulin expression. Surprisingly, the proinsulin fusion protein that contains the EcpD fusion partner of minimum form (50 amino acid whose fusion partner EcpD3) is compared with the truncated form EcpD2 of 100 amino acid whose and full-length fusion partner EcpD1 demonstrates higher expression level. For the proinsulin fusion protein that contains Fk1B or FrnE terminal fusion partner, the expression of the proinsulin that merges with minimum fusion partner fragment Fk1B3 and FrnE3 respectively is equal to or slightly lower than the expression of the construct with longer N-terminal fusion partner. Table 20 has been summarized the proinsulin protein titer that observed in the high throughput expression research process, soluble and total.

[0360] Therefore, mature insulin glargine is determined to be successfully released from the purified fusion protein (and C-peptide) after trypsin cleavage. IMAC enrichment of selected fusion proteins (DnaJ construct G737-031 and Fk1B construct G737-009, purified in the presence of non-denaturing concentrations of urea, and FrnE1 construct G737-018, purified in the absence of urea) followed by trypsin cleavage demonstrated that the fusion protein was cleaved to produce mature insulin, as evaluated by SDS-PAGE or SDS-CGE, compared to insulin glargine standards. Receptor binding assays further demonstrated activity.

[0361] Table 20: HTP expression titers of exemplary proinsulin fusion proteins

[0362]

[0363]

[0364]

[0365]

[0366] Example VI. High-throughput screening of GCSF fusion proteins

[0367] This study was performed to test the levels of recombinant GCSF protein produced by P. fluorescens strains expressing GCSF fusion proteins containing a DnaJ-like protein, varying lengths of Fk1B (Fk1B, Fk1B2, or Fk1B3), FrnE (FrnE, FrnE2, or FrnE3), or EcpD (EcpD1, EcpD3, or EcpD3) as N-terminal fusion partners.

[0368] Materials and methods

[0369] Construction of GCSF Expression Vectors: An AGCSF gene fragment (SEQ ID NO. 68) containing the optimized gcsf coding sequence, recognition sequences for the restriction enzyme Sapl downstream and upstream of the coding sequence, and three terminators downstream of the coding sequence was synthesized by DNA2.0 (Menlo Park, CA). The GCSF gene fragment of plasmid pJ201:207232 was digested with the restriction enzyme Sapl to generate a fragment containing the optimized gcsf coding sequence. The gcsf coding sequence was then subcloned into expression vectors containing various fusion partners by ligating the GCSF gene fragment and expression vector using T4 DNA ligase (Fermentas EL0011) and electroporated into competent Pseudomonas fluorescens DC454 host cells in a 96-well format. A hexahistidine tag was included in the linker between GCSF and each N-terminal fusion partner, along with an enterokinase cleavage site (DDDK) for release of the N-terminal fusion partner from GCSF. The resulting plasmids containing the fusion protein constructs are listed in the third column of Table 21.

[0370] Table 21: Plasmids used for expression of GCSF fusion proteins

[0371]

[0372]

[0373] Growth and Expression (HTP) in 96-well format: Plasmids containing the coding sequences for the gcsf gene and N-terminal fusion partner were transformed into a range of Pseudomonas fluorescens host strains. 35 μL of Pseudomonas fluorescens competent cells were thawed and mixed with 10 μL of 10× diluted plasmid DNA (2.5 ng). Using the Nucleofector TM 96-well Shuttle TM System (Lonza AG) was used to transfer 25 μl of the mixture to a 96-well multiwell Plate (Lonza VHNP-1001) is used for transformation by electroporation, and the transformed cells are subsequently transferred into 96-well deep-well plates (seed plates) containing 500 μL of M9 salt 1% glucose medium and trace elements. The seed plates are incubated at 30° C. with shaking for 48 hours to produce seed cultures.

[0374] 10 microlitre seed cultures were transferred in duplicate to fresh 96-well deep-well plates, each containing 500 μL HTP culture medium (Teknova 3H1129) (supplemented with trace elements and 5% glycerol), and incubated at 30°C for 24 hours. Isopropyl-β-D-1-thiogalactopyranoside (IPTG) was added to each well at a final concentration of 0.3 mM to induce the expression of GCSF fusion protein. In the Pseudomonas strain (FMO strain) overexpressing the folding regulator, mannitol (Sigma, M1902) was added at a final concentration of 1% together with IPTG to induce the expression of the folding regulator. In addition, when inducing, 0.01 μL 250 units / μl reserve Benzonase (Novagen, 70746-3) was added to each well to reduce the possibility of culture viscosity. After 24 hours of induction, the optical density (OD) at 600 nm was measured. 600 ) was used to quantify cell density. 24 hours after induction, cells were harvested and diluted 1:3 with 1× PBS to a final volume of 400 μL, and then frozen for subsequent processing.

[0375] Preparation of soluble lysate samples for analytical characterization: Culture broth samples prepared and frozen as described above were thawed, diluted, and sonicated using a cell lysis automated sonication system (CLASS, Scinomix) with a 24-point probe tip. The sonicated lysate was centrifuged at 5,500 × g for 15 minutes at 8°C to separate the soluble (supernatant) and insoluble (pellet) fractions. The insoluble fraction was resuspended in 400 μL of PBS, pH 7.4, also using sonication.

[0376] SDS-CGE analysis: The test protein samples prepared according to the above discussion were analyzed by HTP microchip SDS capillary gel electrophoresis using a LabChip GXII instrument (Caliper LifeSciences) with HT Protein Express v2 chips and corresponding reagents (part numbers 760499 and 760328, Caliper LifeSciences). Samples were prepared according to the manufacturer's protocol (Protein User Guide Document No.450589, Rev.3). In a 96-well conical well PCR plate, 4 μL of sample was mixed with 14 μL of buffer, with or without dithiothreitol (DTT) reducing agent. The mixture was heated at 95°C for 5 minutes and diluted by adding 70 μL of deionized water. In parallel with the test protein samples, lysates from strains that did not contain the fusion protein (blank strain) were also analyzed. The blank strain lysate was quantified using the system's internal standard without background subtraction. During the HTP screening, one sample of each strain was quantified; typically, the standard deviation of the SDS-CGE method was -10%.

[0377] result

[0378] High levels of GCSF expression were achieved at the 96-well scale using a fusion partner approach that presents an alternative to screening protease-deficient hosts to identify strains capable of high-level expression of N-terminal Met-GCSF. The fusion protein and GCSF titers (calculated based on the percentage of GCSF in the total fusion protein, by MW) are shown in Table 22. Wild-type strain DC454 produced 484 mg / L fusion protein and 305 mg / L GCSF with the dnaJ fusion partner. All fusion partner constructs produced fusion protein titers exceeding 100 mg / L, as shown in Table 22. These high levels observed at the HTP scale show great promise for expression at shake flask or fermentation scales. In addition, a significant increase in volume titer was generally observed between HTP and larger-scale cultures. In previous studies, a prtB protease-deficient strain was shown to be able to express ~247 mg / L Met-GCSF at a 0.5 mL scale (H. Jin et al., 2011, Protein Expression and Purification 78:69-77 and U.S. Patent No. 8,455,218). In the present study, as described, high levels of Met-GCSF expression as part of a fusion protein were observed even in host cells that were not protease-deficient. It is noted that the Met-GCSF preparations obtained by expression as part of either of the fusion proteins and released by protease cleavage contained virtually 100% Met-GCSF (and no des-Met-GCSF) because the cleavage was performed after removal of any protease.

[0379] Table 22: HTP expression titer of GCSF fusion protein

[0380]

[0381] Although preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. It will now be apparent to those skilled in the art that various changes, variations, and substitutions are contemplated without departing from the present invention. It will be appreciated that the various alternatives to the embodiments of the present invention described herein may be used to implement the present invention. It is intended that the following claims define the scope of the present invention and encompass methods and structures within the scope of these claims and their equivalents.

[0382] Sequence Listing

[0383]

[0384]

[0385]

[0386]

[0387]

[0388]

[0389]

[0390]

[0391]

[0392]

[0393]

[0394]

[0395]

[0396]

[0397]

[0398]

[0399]

[0400]

[0401]

[0402]

[0403]

[0404]

[0405]

[0406]

[0407]

[0408]

[0409]

[0410]

[0411]

[0412]

[0413]

[0414]

[0415]

[0416]

[0417]

[0418]

[0419]

[0420]

[0421]

[0422]

[0423]

[0424] SEQUENCE LISTING <110> ϿƼ عɹ ˾ <120> ںϰ <140> 202110725675.2 <141> 2015-11-30 <150> 62 / 086,119 <151> 2014-12-01 <160> 242 <170> PatentIn version 3.5 <210> 1 <211> 34 <212> PRT <213> (Homo sapiens) <400> 1 Ser Val Ser Glu Ile Gln Leu Met His Asn Leu Gly Lys His Leu Asn 1 5 10 15 Ser Met Glu Arg Val Glu Trp Leu Arg Lys Lys Leu Gln Asp Val His 20 25 30 Asn Phe <210> 2 <211> 78 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 2 Met Lys Val Glu Pro Gly Leu Tyr Gln His Tyr Lys Gly Pro Gln Tyr 1 5 10 15 Arg Val Phe Ser Val Ala Arg His Ser Glu Thr Glu Glu Glu Val Val 20 25 30 Phe Tyr Gln Ala Leu Tyr Gly Glu Tyr Gly Phe Trp Val Arg Pro Leu 35 40 45 Ser Met Phe Leu Glu Thr Val Glu Val Asp Gly Glu Gln Val Pro Arg 50 55 60 Phe Ala Leu Val Thr Ala Glu Pro Ser Leu Phe Thr Gly Gln 65 70 75 <210> 3 <211> 217 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 3 Met Ser Thr Pro Leu Lys Ile Asp Phe Val Ser Asp Val Ser Cys Pro 1 5 10 15 Trp Cys Ile Ile Gly Leu Arg Gly Leu Thr Glu Ala Leu Asp Gln Leu 20 25 30 Gly Ser Glu Val Gln Ala Glu Ile His Phe Gln Pro Phe Glu Leu Asn 35 40 45 Pro Asn Met Pro Ala Glu Gly Gln Asn Ile Val Glu His Ile Thr Glu 50 55 60 Lys Tyr Gly Ser Thr Ala Glu Glu Ser Gln Ala Asn Arg Ala Arg Ile 65 70 75 80 Arg Asp Met Gly Ala Ala Leu Gly Phe Ala Phe Arg Thr Asp Gly Gln 85 90 95 Ser Arg Ile Tyr Asn Thr Phe Asp Ala His Arg Leu Leu His Trp Ala 100 105 110 Gly Leu Glu Gly Leu Gln Tyr Asn Leu Lys Glu Ala Leu Phe Lys Ala 115 120 125 Tyr Phe Ser Asp Gly Gln Asp Pro Ser Asp His Ala Thr Leu Ala Ile 130 135 140 Ile Ala Glu Ser Val Gly Leu Asp Leu Ala Arg Ala Ala Glu Ile Leu 145 150 155 160 Ala Ser Asp Glu Tyr Ala Ala Glu Val Arg Glu Gln Glu Gln Leu Trp 165 170 175 Val Ser Arg Gly Val Ser Ser Val Pro Thr Ile Val Phe Asn Asp Gln 180 185 190 Tyr Ala Val Ser Gly Gly Gln Pro Ala Glu Ala Phe Val Gly Ala Ile 195 200 205 Arg Gln Ile Ile Asn Glu Ser Lys Ser 210 215 <210> 4 <211> 205 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 4 Met Ser Glu Val Asn Leu Ser Thr Asp Glu Thr Arg Val Ser Tyr Gly 1 5 10 15 Island Gly Arg Gln Leu Gly Asp Gln Leu Arg Asp Asn Pro Pro Pro Gly 20 25 30 Val Ser Leu Asp Ala Ile Leu Ala Gly Leu Thr Asp Ala Phe Ala Gly 35 40 45 Lys Pro Ser Arg Val Asp Gln Glu Gln Met Ala Ala Ser Phe Lys Val 50 55 60 Ile Arg Glu Ile Met Gln Ala Glu Ala Ala Ala Lys Ala Glu Ala Ala 65 70 75 80 Ala Gly Ala Gly Leu Ala Phe Leu Ala Glu Asn Ala Lys Arg Asp Gly 85 90 95 Ile Thr Thr Leu Ala Ser Gly Leu Gln Phe Glu Val Leu Thr Ala Gly 100 105 110 Thr Gly Ala Lys Pro Thr Arg Glu Asp Gln Val Arg Thr His Tyr His 115 120 125 Gly Thr Leu Ile Asp Gly Thr Val Phe Asp Ser Ser Tyr Glu Arg Gly 130 135 140 Gln Pro Ala Glu Phe Pro Val Gly Gly Val Ile Ala Gly Trp Thr Glu 145 150 155 160 Ala Leu Gln Leu Met Asn Ala Gly Ser Lys Trp Arg Val Tyr Val Pro 165 170 175 Ser Glu Leu Ala Tyr Gly Ala Gln Gly Val Gly Ser Ile Pro Pro His 180 185 190 Ser Val Leu Val Phe Asp Val Glu Leu Leu Asp Val Leu 195 200 205 <210> 5 <211> 225 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 5 Met Ser Arg Tyr Leu Phe Leu Val Phe Gly Leu Ala Ile Cys Val Ala 1 5 10 15 Asp Ala Ser Glu Gln Pro Ser Ser Asn Ile Thr Asp Ala Thr Pro His 20 25 30 Asp Leu Ala Tyr Ser Leu Gly Ala Ser Leu Gly Glu Arg Leu Arg Gln 35 40 45 Glu Val Pro Asp Leu Gln Ile Gln Ala Leu Leu Asp Gly Leu Lys Gln 50 55 60 Ala Tyr Gln Gly Lys Pro Leu Ala Leu Asp Lys Ala Arg Ile Glu Gln 65 70 75 80 Ile Leu Ser Gln His Glu Ala Gln Asn Thr Ala Asp Ala Gln Leu Pro 85 90 95 Gln Ser Glu Lys Ala Leu Ala Ala Glu Gln Gln Phe Leu Thr Arg Glu 100 105 110 Lys Ala Ala Ala Gly Val Arg Gln Leu Ala Asp Gly Ile Leu Leu Thr 115 120 125 Glu Leu Ala Pro Gly Thr Gly Asn Lys Pro Leu Ala Ser Asp Glu Val 130 135 140 Gln Val Lys Tyr Val Gly Arg Leu Pro Asp Gly Thr Val Phe Asp Lys 145 150 155 160 Ser Thr Gln Pro Gln Trp Phe Arg Val Asn Ser Val Ile Ser Gly Trp 165 170 175 Ser Ser Ala Leu Gln Gln Met Pro Val Gly Ala Lys Trp Arg Leu Val 180 185 190 Ile Pro Ser Ala Gln Ala Tyr Gly Ala Asp Gly Ala Gly Glu Leu Ile 195 200 205 Pro Pro Tyr Thr Pro Leu Val Phe Glu Ile Glu Leu Leu Gly Thr Arg 210 215 220 His 225 <210> 6 <211> 159 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 6 Met Thr Asp Gln Gln Asn Thr Glu Ala Ala Gln Asp Gln Gly Pro Gln 1 5 10 15 Phe Ser Leu Gln Arg Ile Tyr Val Arg Asp Leu Ser Phe Glu Ala Pro 20 25 30 Lys Ser Pro Ala Ile Phe Arg Gln Glu Trp Thr Pro Ser Val Ala Leu 35 40 45 Asp Leu Asn Thr Arg Gln Lys Ser Leu Glu Gly Asp Phe His Glu Val 50 55 60 Val Leu Thr Leu Ser Val Thr Val Lys Asn Gly Glu Glu Val Ala Phe 65 70 75 80 Ile Ala Glu Val Gln Gln Ala Gly Ile Phe Leu Ile Gln Gly Leu Asp 85 90 95 Glu Ala Ser Met Ser His Thr Leu Gly Ala Phe Cys Pro Asn Ile Leu 100 105 110 Phe Pro Tyr Ala Arg Glu Thr Leu Asp Ser Leu Val Thr Arg Gly Ser 115 120 125 Phe Pro Ala Leu Met Leu Ala Pro Val Asn Phe Asp Ala Leu Tyr Ala 130 135 140 Gln Glu Leu Gln Arg Met Gln Gln Glu Gly Ala Pro Thr Val Gln 145 150 155 <210> 7 <211> 241 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 7 Met Gly Cys Val Pro Leu Pro Asp His Gly Ile Thr Val Phe Met Phe 1 5 10 15 Leu Leu Arg Met Val Leu Leu Ala Cys Gly Leu Leu Val Leu Ala Pro 20 25 30 Pro Pro Ala Asp Ala Ala Leu Lys Ile Glu Gly Thr Arg Leu Ile Tyr 35 40 45 Phe Gly Gln Asp Lys Ala Ala Gly Ile Ser Val Val Asn Gln Ala Ser 50 55 60 Arg Glu Val Val Val Gln Thr Trp Ile Thr Gly Glu Asp Glu Ser Ala 65 70 75 80 Asp Arg Thr Val Pro Phe Ala Ala Thr Glu Pro Leu Val Gln Leu Gly 85 90 95 Ala Gly Glu His His Lys Leu Arg Ile Leu Tyr Ala Gly Glu Gly Leu 100 105 110 Pro Ser Asp Arg Glu Ser Leu Phe Trp Leu Asn Ile Met Glu Ile Pro 115 120 125 Leu Lys Pro Glu Asp Pro Asn Ser Val Gln Phe Ala Ile Arg Gln Arg 130 135 140 Leu Lys Leu Phe Tyr Arg Pro Pro Ala Leu Gln Gly Gly Ser Ala Glu 145 150 155 160 Ala Val Gln Gln Leu Val Trp Ser Ser Asp Gly Arg Thr Val Thr Val 165 170 175 Asn Asn Pro Ser Ala Phe His Leu Ser Leu Val Asn Leu Arg Ile Asp 180 185 190 Ser Gln Thr Leu Ser Asp Tyr Leu Leu Leu Lys Pro His Glu Arg Lys 195 200 205 Thr Leu Thr Ala Leu Asp Ala Val Pro Lys Gly Ala Thr Leu His Phe 210 215 220 Thr Glu Ile Thr Asp Ile Gly Leu Gln Ala Arg His Ser Thr Ala Leu 225 230 235 240 Asn <210> 8 <211> 141 <212> PRT <213> Escherichia coli <400> 8 Ala Asp Lys Ile Ala Ile Val Asn Met Gly Ser Leu Phe Gln Gln Val 1 5 10 15 Ala Gln Lys Thr Gly Val Ser Asn Thr Leu Glu Asn Glu Phe Lys Gly 20 25 30 Arg Ala Ser Glu Leu Gln Arg Met Glu Thr Asp Leu Gln Ala Lys Met 35 40 45 Lys Lys Leu Gln Ser Met Lys Ala Gly Ser Asp Arg Thr Lys Leu Glu 50 55 60 Lys Asp Val Met Ala Gln Arg Gln Thr Phe Ala Gln Lys Ala Gln Ala 65 70 75 80 Phe Glu Gln Asp Arg Ala Arg Arg Ser Asn Glu Glu Arg Gly Lys Leu 85 90 95 Val Thr Arg Ile Gln Thr Ala Val Lys Ser Val Ala Asn Ser Gln Asp 100 105 110 Ile Asp Leu Val Val Asp Ala Asn Ala Val Ala Tyr Asn Ser Ser Asp 115 120 125 Val Lys Asp Ile Thr Ala Asp Val Leu Lys Gln Val Lys 130 135 140 <210> 9 <211> 20 <212> PRT <213> ˹ (Artificial Sequence) <220> <223> ˹ е ϳ <400> 9 Gly Gly Gly Gly Ser Gly Gly Gly Gly His His His His His His Asp 1 5 10 15 Asp Asp Asp Lys 20 <210> 10 <211> 18 <212> PRT <213> ˹ (Artificial Sequence) <220> <223> ˹ е ϳ <400> 10 Gly Gly Gly Gly Ser Gly Gly Gly Gly His His His His His His Arg 1 5 10 15 Lys Arg <210> 11 <211> 18 <212> PRT <213> ˹ (Artificial Sequence) <220> <223> ˹ е ϳ <400> 11 Gly Gly Gly Gly Ser Gly Gly Gly Gly His His His His His His Arg 1 5 10 15 Arg Arg <210> 12 <211> 19 <212> PRT <213> ˹ (Artificial Sequence) <220> <223> ˹ е ϳ <400> 12 Gly Gly Gly Gly Ser Gly Gly Gly Gly His His His His His His Leu 1 5 10 15 Val Pro Arg <210> 13 <211> 5 <212> PRT <213> δ֪ <220> <223> δ֪ ø и λ <400> 13 Asp Asp Asp Asp Lys 1 5 <210> 14 <211> 798 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 14 Met Lys Thr Thr Ile Glu Leu Pro Leu Leu Pro Leu Arg Asp Val Val 1 5 10 15 Val Tyr Pro His Met Val Ile Pro Leu Phe Val Gly Arg Glu Lys Ser 20 25 30 Ile Glu Ala Leu Glu Ala Ala Met Thr Gly Asp Lys Gln Ile Leu Leu 35 40 45 Leu Ala Gln Lys Asn Pro Ala Asp Asp Asp Pro Gly Glu Asp Ala Leu 50 55 60 Tyr Arg Val Gly Thr Ile Ala Thr Val Leu Gln Leu Leu Lys Leu Pro 65 70 75 80 Asp Gly Thr Val Lys Val Leu Val Glu Gly Glu Gln Arg Gly Ala Val 85 90 95 Glu Arg Phe Met Glu Val Asp Gly His Leu Arg Ala Glu Val Ala Leu 100 105 110 Ile Glu Glu Val Glu Ala Pro Glu Arg Glu Ser Glu Val Phe Val Arg 115 120 125 Ser Leu Leu Ser Gln Phe Glu Gln Tyr Val Gln Leu Gly Lys Lys Val 130 135 140 Pro Ala Glu Val Leu Ser Ser Leu Asn Ser Ile Asp Glu Pro Ser Arg 145 150 155 160 Leu Val Asp Thr Met Ala Ala His Met Ala Leu Lys Ile Glu Gln Lys 165 170 175 Gln Asp Ile Leu Glu Ile Ile Asp Leu Ser Ala Arg Val Glu His Val 180 185 190 Leu Ala Met Leu Asp Gly Glu Ile Asp Leu Leu Gln Val Glu Lys Arg 195 200 205 Ile Arg Gly Arg Val Lys Lys Gln Met Glu Arg Ser Gln Arg Glu Tyr 210 215 220 Tyr Leu Asn Glu Gln Met Lys Ala Ile Gln Lys Glu Leu Gly Asp Gly 225 230 235 240 Glu Glu Gly His Asn Glu Ile Glu Glu Leu Lys Lys Arg Ile Asp Ala 245 250 255 Ala Gly Leu Pro Lys Asp Ala Leu Thr Lys Ala Thr Ala Glu Leu Asn 260 265 270 Lys Leu Lys Gln Met Ser Pro Met Ser Ala Glu Ala Thr Val Val Arg 275 280 285 Ser Tyr Ile Asp Trp Leu Val Gln Val Pro Trp Lys Ala Gln Thr Lys 290 295 300 Val Arg Leu Asp Leu Ala Arg Ala Glu Glu Ile Leu Asp Ala Asp His 305 310 315 320 Tyr Gly Leu Glu Glu Val Lys Glu Arg Ile Leu Glu Tyr Leu Ala Val 325 330 335 Gln Lys Arg Val Lys Lys Ile Arg Gly Pro Val Leu Cys Leu Val Gly 340 345 350 Pro Pro Gly Val Gly Lys Thr Ser Leu Ala Glu Ser Ile Ala Ser Ala 355 360 365 Thr Asn Arg Lys Phe Val Arg Met Ala Leu Gly Gly Val Arg Asp Glu 370 375 380 Ala Glu Ile Arg Gly His Arg Arg Thr Tyr Ile Gly Ser Met Pro Gly 385 390 395 400 Arg Leu Ile Gln Lys Met Thr Lys Val Gly Val Arg Asn Pro Leu Phe 405 410 415 Leu Leu Asp Glu Ile Asp Lys Met Gly Ser Asp Met Arg Gly Asp Pro 420 425 430 Ala Ser Ala Leu Leu Glu Val Leu Asp Pro Glu Gln Asn His Asn Phe 435 440 445 Asn Asp His Tyr Leu Glu Val Asp Tyr Asp Leu Ser Asp Val Met Phe 450 455 460 Leu Cys Thr Ser Asn Ser Met Asn Ile Pro Pro Ala Leu Leu Asp Arg 465 470 475 480 Met Glu Val Ile Arg Leu Pro Gly Tyr Thr Glu Asp Glu Lys Ile Asn 485 490 495 Ile Ala Val Lys Tyr Leu Ala Pro Lys Gln Ile Ser Ala Asn Gly Leu 500 505 510 Lys Lys Gly Glu Ile Glu Phe Glu Val Glu Ala Ile Arg Asp Ile Val 515 520 525 Arg Tyr Tyr Thr Arg Glu Ala Gly Val Arg Gly Leu Glu Arg Gln Ile 530 535 540 Ala Lys Ile Cys Arg Lys Ala Val Lys Glu His Ala Leu Glu Lys Arg 545 550 555 560 Phe Ser Val Lys Val Val Ala Asp Ser Leu Glu His Phe Leu Gly Val 565 570 575 Lys Lys Phe Arg Tyr Gly Leu Ala Glu Gln Gln Asp Gln Val Gly Gln 580 585 590 Val Thr Gly Leu Ala Trp Thr Gln Val Gly Gly Glu Leu Leu Thr Ile 595 600 605 Glu Ala Ala Val Ile Pro Gly Lys Gly Gln Leu Ile Lys Thr Gly Ser 610 615 620 Leu Gly Asp Val Met Val Glu Ser Ile Thr Ala Ala Gln Thr Val Val 625 630 635 640 Arg Ser Arg Ala Arg Ser Leu Gly Ile Pro Leu Asp Phe His Glu Lys 645 650 655 His Asp Thr His Ile His Met Pro Glu Gly Ala Thr Pro Lys Asp Gly 660 665 670 Pro Ser Ala Gly Val Gly Met Cys Thr Ala Leu Val Ser Ala Leu Thr 675 680 685 Gly Ile Pro Val Arg Ala Asp Val Ala Met Thr Gly Glu Ile Thr Leu 690 695 700 Arg Gly Gln Val Leu Ala Ile Gly Gly Leu Lys Glu Lys Leu Leu Ala 705 710 715 720 Ala His Arg Gly Gly Ile Lys Thr Val Ile Ile Pro Glu Glu Asn Val 725 730 735 Arg Asp Leu Lys Glu Ile Pro Asp Asn Ile Lys Gln Asp Leu Gln Ile 740 745 750 Lys Pro Val Lys Trp Ile Asp Glu Val Leu Gln Ile Ala Leu Gln Tyr 755 760 765 Ala Pro Glu Pro Leu Pro Asp Val Ala Pro Glu Ile Val Ala Lys Asp 770 775 780 Glu Lys Arg Glu Ser Asp Ser Lys Glu Arg Ile Ser Thr His 785 790 795 <210> 15 <211> 806 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 15 Met Ser Asp Gln Gln Glu Phe Pro Asp Tyr Asp Leu Asn Asp Tyr Ala 1 5 10 15 Asp Pro Glu Asn Ala Glu Ala Pro Ser Ser Asn Thr Gly Leu Ala Leu 20 25 30 Pro Gly Gln Asn Leu Pro Asp Lys Val Tyr Ile Ile Pro Ile His Asn 35 40 45 Arg Pro Phe Phe Pro Ala Gln Val Leu Pro Val Ile Val Asn Glu Glu 50 55 60 Pro Trp Ala Glu Thr Leu Glu Leu Val Ser Lys Ser Asp His His Ser 65 70 75 80 Leu Ala Leu Phe Phe Met Asp Thr Pro Pro Asp Asp Pro Arg His Phe 85 90 95 Asp Thr Ser Ala Leu Pro Leu Tyr Gly Thr Leu Val Lys Val His His 100 105 110 Ala Ser Arg Glu Asn Gly Lys Leu Gln Phe Val Ala Gln Gly Leu Thr 115 120 125 Arg Val Arg Ile Lys Thr Trp Leu Lys His His Arg Pro Pro Tyr Leu 130 135 140 Val Glu Val Glu Tyr Pro His Gln Pro Ser Glu Pro Thr Asp Glu Val 145 150 155 160 Lys Ala Tyr Gly Met Ala Leu Ile Asn Ala Ile Lys Glu Leu Leu Pro 165 170 175 Leu Asn Pro Leu Tyr Ser Glu Glu Leu Lys Asn Tyr Leu Asn Arg Phe 180 185 190 Ser Pro Asn Asp Pro Ser Pro Leu Thr Asp Phe Ala Ala Ala Leu Thr 195 200 205 Ser Ala Thr Gly Asn Glu Leu Gln Glu Val Leu Asp Cys Val Pro Met 210 215 220 Leu Lys Arg Met Glu Lys Val Leu Pro Met Leu Arg Lys Glu Val Glu 225 230 235 240 Val Ala Arg Leu Gln Lys Glu Leu Ser Ala Glu Val Asn Arg Lys Ile 245 250 255 Gly Glu His Gln Arg Glu Phe Phe Leu Lys Glu Gln Leu Lys Val Ile 260 265 270 Gln Gln Glu Leu Gly Leu Thr Lys Asp Asp Arg Ser Ala Asp Val Glu 275 280 285 Gln Phe Glu Gln Arg Leu Gln Gly Lys Val Leu Pro Ala Gln Ala Gln 290 295 300 Lys Arg Ile Asp Glu Glu Leu Asn Lys Leu Ser Ile Leu Glu Thr Gly 305 310 315 320 Ser Pro Glu Tyr Ala Val Thr Arg Asn Tyr Leu Asp Trp Ala Thr Ser 325 330 335 Val Pro Trp Gly Val Tyr Gly Ala Asp Lys Leu Asp Leu Lys His Ala 340 345 350 Arg Lys Val Leu Asp Lys His His Ala Gly Leu Asp Asp Ile Lys Ser 355 360 365 Arg Ile Leu Glu Phe Leu Ala Val Gly Ala Tyr Lys Gly Glu Val Ala 370 375 380 Gly Ser Ile Val Leu Leu Val Gly Pro Pro Gly Val Gly Lys Thr Ser 385 390 395 400 Val Gly Lys Ser Ile Ala Glu Ser Leu Gly Arg Pro Phe Tyr Arg Phe 405 410 415 Ser Val Gly Gly Met Arg Asp Glu Ala Glu Ile Lys Gly His Arg Arg 420 425 430 Thr Tyr Ile Gly Ala Leu Pro Gly Lys Leu Val Gln Ala Leu Lys Asp 435 440 445 Val Glu Val Met Asn Pro Val Ile Met Leu Asp Glu Ile Asp Lys Met 450 455 460 Gly Gln Ser Phe Gln Gly Asp Pro Ala Ser Ala Leu Leu Glu Thr Leu 465 470 475 480 Asp Pro Glu Gln Asn Val Glu Phe Leu Asp His Tyr Leu Asp Leu Arg 485 490 495 Leu Asp Leu Ser Lys Val Leu Phe Val Cys Thr Ala Asn Thr Leu Asp 500 505 510 Ser Ile Pro Gly Pro Leu Leu Asp Arg Met Glu Val Ile Arg Leu Ser 515 520 525 Gly Tyr Ile Thr Glu Glu Lys Val Ala Ile Ala Lys Arg His Leu Trp 530 535 540 Pro Lys Gln Leu Glu Lys Ala Gly Val Ala Lys Asn Ser Leu Thr Ile 545 550 555 560 Ser Asp Gly Ala Leu Arg Ala Leu Ile Asp Gly Tyr Ala Arg Glu Ala 565 570 575 Gly Val Arg Gln Leu Glu Lys Gln Leu Gly Lys Leu Val Arg Lys Ala 580 585 590 Val Val Lys Leu Leu Asp Glu Pro Asp Ser Val Ile Lys Ile Gly Asn 595 600 605 Lys Asp Leu Glu Ser Ser Leu Gly Met Pro Val Phe Arg Asn Glu Gln 610 615 620 Val Leu Ser Gly Thr Gly Val Ile Thr Gly Leu Ala Trp Thr Ser Met 625 630 635 640 Gly Gly Ala Thr Leu Pro Ile Glu Ala Thr Arg Ile His Thr Leu Asn 645 650 655 Arg Gly Phe Lys Leu Thr Gly Gln Leu Gly Glu Val Met Lys Glu Ser 660 665 670 Ala Glu Ile Ala Tyr Ser Tyr Ile Ser Ser Asn Leu Lys Ser Phe Gly 675 680 685 Gly Asp Ala Lys Phe Phe Asp Glu Ala Phe Val His Leu His Val Pro 690 695 700 Glu Gly Ala Thr Pro Lys Asp Gly Pro Ser Ala Gly Val Thr Met Ala 705 710 715 720 Ser Ala Leu Leu Ser Leu Ala Arg Asn Gln Pro Pro Lys Lys Gly Val 725 730 735 Ala Met Thr Gly Glu Leu Thr Leu Thr Gly His Val Leu Pro Ile Gly 740 745 750 Gly Val Arg Glu Lys Val Ile Ala Ala Arg Arg Gln Lys Ile His Glu 755 760 765 Leu Ile Leu Pro Glu Pro Asn Arg Gly Ser Phe Glu Glu Leu Pro Asp 770 775 780 Tyr Leu Lys Glu Gly Met Thr Val His Phe Ala Lys Arg Phe Ala Asp 785 790 795 800 Val Ala Lys Val Leu Phe 805 <210> 16 <211> 477 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 16 Met Ser Lys Val Lys Asp Lys Ala Ile Val Ser Ala Ala Gln Ala Ser 1 5 10 15 Thr Ala Tyr Ser Gln Ile Asp Ser Phe Ser His Leu Tyr Asp Arg Gly 20 25 30 Gly Asn Leu Thr Val Asn Gly Lys Pro Ser Tyr Thr Val Asp Gln Ala 35 40 45 Ala Thr Gln Leu Leu Arg Asp Gly Ala Ala Tyr Arg Asp Phe Asp Gly 50 55 60 Asn Gly Lys Ile Asp Leu Thr Tyr Thr Phe Leu Thr Ser Ala Thr Gln 65 70 75 80 Ser Thr Met Asn Lys His Gly Ile Ser Gly Phe Ser Gln Phe Asn Thr 85 90 95 Gln Gln Lys Ala Gln Ala Ala Leu Ala Met Gln Ser Trp Ala Asp Val 100 105 110 Ala Asn Val Thr Phe Thr Glu Lys Ala Ser Gly Gly Asp Gly His Met 115 120 125 Thr Phe Gly Asn Tyr Ser Ser Gly Gln Asp Gly Ala Ala Ala Phe Ala 130 135 140 Tyr Leu Pro Gly Thr Gly Ala Gly Tyr Asp Gly Thr Ser Trp Tyr Leu 145 150 155 160 Thr Asn Asn Ser Tyr Thr Pro Asn Lys Thr Pro Asp Leu Asn Asn Tyr 165 170 175 Gly Arg Gln Thr Leu Thr His Glu Ile Gly His Thr Leu Gly Leu Ala 180 185 190 His Pro Gly Asp Tyr Asn Ala Gly Asn Gly Asn Pro Thr Tyr Asn Asp 195 200 205 Ala Thr Tyr Gly Gln Asp Thr Arg Gly Tyr Ser Leu Met Ser Tyr Trp 210 215 220 Ser Glu Ser Asn Thr Asn Gln Asn Phe Ser Lys Gly Gly Val Glu Ala 225 230 235 240 Tyr Ala Ser Gly Pro Leu Ile Asp Asp Ile Ala Ala Ile Gln Lys Leu 245 250 255 Tyr Gly Ala Asn Leu Ser Thr Arg Ala Thr Asp Thr Thr Tyr Gly Phe 260 265 270 Asn Ser Asn Thr Gly Arg Asp Phe Leu Ser Ala Thr Ser Asn Ala Asp 275 280 285 Lys Leu Val Phe Ser Val Trp Asp Gly Gly Gly Asn Asp Thr Leu Asp 290 295 300 Phe Ser Gly Phe Thr Gln Asn Gln Lys Ile Asn Leu Thr Ala Thr Ser 305 310 315 320 Phe Ser Asp Val Gly Gly Leu Val Gly Asn Val Ser Ile Ala Lys Gly 325 330 335 Val Thr Ile Glu Asn Ala Phe Gly Gly Ala Gly Asn Asp Leu Ile Ile 340 345 350 Gly Asn Gln Val Ala Asn Thr Ile Lys Gly Gly Ala Gly Asn Asp Leu 355 360 365 Ile Tyr Gly Gly Gly Gly Ala Asp Gln Leu Trp Gly Gly Ala Gly Ser 370 375 380 Asp Thr Phe Val Tyr Gly Ala Ser Ser Asp Ser Lys Pro Gly Ala Ala 385 390 395 400 Asp Lys Ile Phe Asp Phe Thr Ser Gly Ser Asp Lys Ile Asp Leu Ser 405 410 415 Gly Ile Thr Lys Gly Ala Gly Val Thr Phe Val Asn Ala Phe Thr Gly 420 425 430 His Ala Gly Asp Ala Val Leu Ser Tyr Ala Ser Gly Thr Asn Leu Gly 435 440 445 Thr Leu Ala Val Asp Phe Ser Gly His Gly Val Ala Asp Phe Leu Val 450 455 460 Thr Thr Val Gly Gln Ala Ala Ala Ser Asp Ile Val Ala 465 470 475 <210> 17 <211> 295 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 17 Met Met Arg Ile Leu Leu Phe Leu Ala Thr Asn Leu Ala Val Val Leu 1 5 10 15 Ile Ala Ser Val Thr Leu Ser Leu Phe Gly Phe Asn Gly Phe Met Ala 20 25 30 Ala Asn Gly Val Asp Leu Asn Leu Asn Gln Leu Leu Ile Phe Cys Ala 35 40 45 Val Phe Gly Phe Ala Gly Ser Leu Phe Ser Leu Phe Ile Ser Lys Trp 50 55 60 Met Ala Lys Met Ser Thr Ser Thr Gln Ile Ile Thr Gln Pro Arg Thr 65 70 75 80 Arg His Glu Gln Trp Leu Met Gln Thr Val Glu Gln Leu Ser Gln Glu 85 90 95 Ala Gly Ile Lys Met Pro Glu Val Gly Ile Phe Pro Ala Tyr Glu Ala 100 105 110 Asn Ala Phe Ala Thr Gly Trp Asn Lys Asn Asp Ala Leu Val Ala Val 115 120 125 Ser Gln Gly Leu Leu Glu Arg Phe Ser Pro Asp Glu Val Lys Ala Val 130 135 140 Leu Ala His Glu Ile Gly His Val Ala Asn Gly Asp Met Val Thr Leu 145 150 155 160 Ala Leu Val Gln Gly Val Val Asn Thr Phe Val Met Phe Phe Ala Arg 165 170 175 Ile Ile Gly Asn Phe Val Asp Lys Val Ile Phe Lys Asn Glu Glu Gly 180 185 190 Arg Gly Ile Ala Tyr Phe Val Ala Thr Ile Phe Ala Glu Leu Val Leu 195 200 205 Gly Phe Leu Ala Ser Ala Ile Val Met Trp Phe Ser Arg Lys Arg Glu 210 215 220 Phe Arg Ala Asp Glu Ala Gly Ala Arg Leu Ala Gly Thr Ser Ala Met 225 230 235 240 Ile Gly Ala Leu Gln Arg Leu Arg Ser Glu Gln Gly Leu Pro Val His 245 250 255 Met Pro Asp Ser Leu Thr Ala Phe Gly Ile Asn Gly Gly Ile Lys Gln 260 265 270 Gly Leu Ala Arg Leu Phe Met Ser His Pro Pro Leu Glu Glu Arg Ile 275 280 285 Asp Ala Leu Arg Arg Arg Gly 290 295 <210> 18 <211> 386 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 18 Met Leu Lys Ala Leu Arg Phe Phe Gly Trp Pro Leu Leu Ala Gly Val 1 5 10 15 Leu Ile Ala Met Leu Ile Ile Gln Arg Tyr Pro Gln Trp Val Gly Leu 20 25 30 Pro Thr Leu Asp Val Asn Leu Gln Gln Ala Pro Gln Thr Asn Thr Val 35 40 45 Val Gln Gly Pro Val Thr Tyr Ala Asp Ala Val Val Ile Ala Ala Pro 50 55 60 Ala Val Val Asn Leu Tyr Thr Thr Lys Val Ile Asn Lys Pro Ala His 65 70 75 80 Pro Leu Phe Glu Asp Pro Gln Phe Arg Arg Tyr Phe Gly Asp Asn Gly 85 90 95 Pro Lys Gln Arg Arg Met Glu Ser Ser Leu Gly Ser Gly Val Ile Met 100 105 110 Ser Pro Glu Gly Tyr Ile Leu Thr Asn Asn His Val Thr Thr Gly Ala 115 120 125 Asp Gln Ile Val Val Ala Leu Arg Asp Gly Arg Glu Thr Leu Ala Arg 130 135 140 Val Val Gly Ser Asp Pro Glu Thr Asp Leu Ala Val Leu Lys Ile Asp 145 150 155 160 Leu Lys Asn Leu Pro Ala Ile Thr Leu Gly Arg Ser Asp Gly Leu Arg 165 170 175 Val Gly Asp Val Ala Leu Ala Ile Gly Asn Pro Phe Gly Val Gly Gln 180 185 190 Thr Val Thr Met Gly Ile Ile Ser Ala Thr Gly Arg Asn Gln Leu Gly 195 200 205 Leu Asn Ser Tyr Glu Asp Phe Ile Gln Thr Asp Ala Ala Ile Asn Pro 210 215 220 Gly Asn Ser Gly Gly Ala Leu Val Asp Ala Asn Gly Asn Leu Thr Gly 225 230 235 240 Ile Asn Thr Ala Ile Phe Ser Lys Ser Gly Gly Ser Gln Gly Ile Gly 245 250 255 Phe Ala Ile Pro Val Lys Leu Ala Met Glu Val Met Lys Ser Ile Ile 260 265 270 Glu His Gly Gln Val Ile Arg Gly Trp Leu Gly Ile Glu Val Gln Pro 275 280 285 Leu Thr Lys Glu Leu Ala Glu Ser Phe Gly Leu Thr Gly Arg Pro Gly 290 295 300 Ile Val Val Ala Gly Ile Phe Arg Asp Gly Pro Ala Gln Lys Ala Gly 305 310 315 320 Leu Gln Leu Gly Asp Val Ile Leu Ser Ile Asp Gly Ala Pro Ala Gly 325 330 335 Asp Gly Arg Lys Ser Met Asn Gln Val Ala Arg Ile Lys Pro Thr Asp 340 345 350 Lys Val Ala Ile Leu Val Met Arg Asn Gly Lys Glu Ile Lys Leu Ser 355 360 365 Ala Glu Ile Gly Leu Arg Pro Pro Pro Ala Thr Ala Pro Val Lys Glu 370 375 380 Glu Gln 385 <210> 19 <211> 478 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 19 Met Ser Ile Pro Arg Leu Lys Ser Tyr Leu Ser Ile Val Ala Thr Val 1 5 10 15 Leu Val Leu Gly Gln Ala Leu Pro Ala Gln Ala Val Glu Leu Pro Asp 20 25 30 Phe Thr Gln Leu Val Glu Gln Ala Ser Pro Ala Val Val Asn Ile Ser 35 40 45 Thr Thr Gln Lys Leu Pro Asp Arg Lys Val Ser Asn Gln Gln Met Pro 50 55 60 Asp Leu Glu Gly Leu Pro Pro Met Leu Arg Glu Phe Phe Glu Arg Gly 65 70 75 80 Met Pro Gln Pro Arg Ser Pro Arg Gly Gly Gly Gly Gln Arg Glu Ala 85 90 95 Gln Ser Leu Gly Ser Gly Phe Ile Ile Ser Pro Asp Gly Tyr Ile Leu 100 105 110 Thr Asn Asn His Val Ile Ala Asp Ala Asp Glu Ile Leu Val Arg Leu 115 120 125 Ala Asp Arg Ser Glu Leu Lys Ala Lys Leu Ile Gly Thr Asp Pro Arg 130 135 140 Ser Asp Val Ala Leu Leu Lys Ile Glu Gly Lys Asp Leu Pro Val Leu 145 150 155 160 Lys Leu Gly Lys Ser Gln Asp Leu Lys Ala Gly Gln Trp Val Val Ala 165 170 175 Ile Gly Ser Pro Phe Gly Phe Asp His Thr Val Thr Gln Gly Ile Val 180 185 190 Ser Ala Ile Gly Arg Ser Leu Pro Asn Glu Asn Tyr Val Pro Phe Ile 195 200 205 Gln Thr Asp Val Pro Ile Asn Pro Gly Asn Ser Gly Gly Pro Leu Phe 210 215 220 Asn Leu Ala Gly Glu Val Val Gly Ile Asn Ser Gln Ile Tyr Thr Arg 225 230 235 240 Ser Gly Gly Phe Met Gly Val Ser Phe Ala Ile Pro Ile Asp Val Ala 245 250 255 Met Asp Val Ser Asn Gln Leu Lys Ser Gly Gly Lys Val Ser Arg Gly 260 265 270 Trp Leu Gly Val Val Ile Gln Glu Val Asn Lys Asp Leu Ala Glu Ser 275 280 285 Phe Gly Leu Asp Lys Pro Ala Gly Ala Leu Val Ala Gln Ile Gln Asp 290 295 300 Asn Gly Pro Ala Ala Lys Gly Gly Leu Lys Val Gly Asp Val Ile Leu 305 310 315 320 Ser Met Asn Gly Gln Pro Ile Ile Met Ser Ala Asp Leu Pro His Leu 325 330 335 Val Gly Ala Leu Lys Ala Gly Gly Lys Ala Lys Leu Glu Val Ile Arg 340 345 350 Asp Gly Lys Arg Gln Asn Val Glu Leu Thr Val Gly Ala Ile Pro Glu 355 360 365 Glu Gly Ala Thr Leu Asp Ala Leu Gly Asn Ala Lys Pro Gly Ala Glu 370 375 380 Arg Ser Ser Asn Arg Leu Gly Ile Ala Val Val Glu Leu Thr Ala Glu 385 390 395 400 Gln Lys Lys Thr Phe Asp Leu Gln Ser Gly Val Val Ile Lys Glu Val 405 410 415 Gln Asp Gly Pro Ala Ala Leu Ile Gly Leu Gln Pro Gly Asp Val Ile 420 425 430 Thr His Leu Asn Asn Gln Ala Ile Asp Thr Thr Lys Glu Phe Ala Asp 435 440 445 Ile Ala Lys Ala Leu Pro Lys Asn Arg Ser Val Ser Met Arg Val Leu 450 455 460 Arg Gln Gly Arg Ala Ser Phe Ile Thr Phe Lys Leu Ala Glu 465 470 475 <210> 20 <211> 353 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 20 Met Cys Val Arg Gln Pro Arg Asn Pro Ile Phe Cys Leu Ile Pro Pro 1 5 10 15 Tyr Met Leu Asp Gln Ile Ala Arg His Gly Asp Lys Ala Gln Arg Glu 20 25 30 Val Ala Leu Arg Thr Arg Ala Lys Asp Ser Thr Phe Arg Ser Leu Arg 35 40 45 Met Val Ala Val Pro Ala Lys Gly Pro Ala Arg Met Ala Leu Ala Val 50 55 60 Gly Ala Glu Lys Gln Arg Ser Ile Tyr Ser Ala Glu Asn Thr Asp Ser 65 70 75 80 Leu Pro Gly Lys Leu Ile Arg Gly Glu Gly Gln Pro Ala Ser Gly Asp 85 90 95 Ala Ala Val Asp Glu Ala Tyr Asp Gly Leu Gly Ala Thr Phe Asp Phe 100 105 110 Phe Asp Gln Val Phe Asp Arg Asn Ser Ile Asp Asp Ala Gly Met Ala 115 120 125 Leu Asp Ala Thr Val His Phe Gly Gln Asp Tyr Asn Asn Ala Phe Trp 130 135 140 Asn Ser Thr Gln Met Val Phe Gly Asp Gly Asp Gln Gln Leu Phe Asn 145 150 155 160 Arg Phe Thr Val Ala Leu Asp Val Ile Gly His Glu Leu Ala His Gly 165 170 175 Val Thr Glu Asp Glu Ala Lys Leu Met Tyr Phe Asn Gln Ser Gly Ala 180 185 190 Leu Asn Glu Ser Leu Ser Asp Val Phe Gly Ser Leu Ile Lys Gln Tyr 195 200 205 Ala Leu Lys Gln Thr Ala Glu Asp Ala Asp Trp Leu Ile Gly Lys Gly 210 215 220 Leu Phe Thr Lys Lys Ile Lys Gly Thr Ala Leu Arg Ser Met Lys Ala 225 230 235 240 Pro Gly Thr Ala Phe Asp Asp Lys Leu Leu Gly Lys Asp Pro Gln Pro 245 250 255 Gly His Met Asp Asp Phe Val Gln Thr Tyr Glu Asp Asn Gly Gly Val 260 265 270 His Ile Asn Ser Gly Ile Pro Asn His Ala Phe Tyr Gln Val Ala Ile 275 280 285 Asn Ile Gly Gly Phe Ala Trp Glu Arg Ala Gly Arg Ile Trp Tyr Asp 290 295 300 Ala Leu Arg Asp Ser Arg Leu Arg Pro Asn Ser Gly Phe Leu Arg Phe 305 310 315 320 Ala Arg Ile Thr His Asp Ile Ala Gly Gln Leu Tyr Gly Val Asn Lys 325 330 335 Ala Glu Gln Lys Ala Val Lys Glu Gly Trp Lys Ala Val Gly Ile Asn 340 345 350 Val <210> 21 <211> 704 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 21 Met Arg Tyr Gln Leu Pro Pro Arg Arg Ile Ser Met Lys His Leu Phe 1 5 10 15 Pro Ser Thr Ala Leu Ala Phe Phe Ile Gly Leu Gly Phe Ala Ser Met 20 25 30 Ser Thr Asn Thr Phe Ala Ala Asn Ser Trp Asp Asn Leu Gln Pro Asp 35 40 45 Arg Asp Glu Val Ile Ala Ser Leu Asn Val Val Glu Leu Leu Lys Arg 50 55 60 His His Tyr Ser Lys Pro Pro Leu Asp Asp Ala Arg Ser Val Ile Ile 65 70 75 80 Tyr Asp Ser Tyr Leu Lys Leu Leu Asp Pro Ser Arg Ser Tyr Phe Leu 85 90 95 Ala Ser Asp Ile Ala Glu Phe Asp Lys Trp Lys Thr Gln Phe Asp Asp 100 105 110 Phe Leu Lys Ser Gly Asp Leu Gln Pro Gly Phe Thr Ile Tyr Lys Arg 115 120 125 Tyr Leu Asp Arg Val Lys Ala Arg Leu Asp Phe Ala Leu Gly Glu Leu 130 135 140 Asn Lys Gly Val Asp Lys Leu Asp Phe Thr Gln Lys Glu Thr Leu Leu 145 150 155 160 Val Asp Arg Lys Asp Ala Pro Trp Leu Thr Ser Thr Ala Ala Leu Asp 165 170 175 Asp Leu Trp Arg Lys Arg Val Lys Asp Glu Val Leu Arg Leu Lys Ile 180 185 190 Ala Gly Lys Glu Pro Lys Ala Ile Gln Glu Leu Leu Thr Lys Arg Tyr 195 200 205 Lys Asn Gln Leu Ala Arg Leu Asp Gln Thr Arg Ala Glu Asp Ile Phe 210 215 220 Gln Ala Tyr Ile Asn Thr Phe Ala Met Ser Tyr Asp Pro His Thr Asn 225 230 235 240 Tyr Leu Ser Pro Asp Asn Ala Glu Asn Phe Asp Ile Asn Met Ser Leu 245 250 255 Ser Leu Glu Gly Ile Gly Ala Val Leu Gln Ser Asp Asn Asp Gln Val 260 265 270 Lys Ile Val Arg Leu Val Pro Ala Gly Pro Ala Asp Lys Thr Lys Gln 275 280 285 Val Ala Pro Ala Asp Lys Ile Ile Gly Val Ala Gln Ala Asp Lys Glu 290 295 300 Met Val Asp Val Val Gly Trp Arg Leu Asp Glu Val Val Lys Leu Ile 305 310 315 320 Arg Gly Pro Lys Gly Ser Val Val Arg Leu Glu Val Ile Pro His Thr 325 330 335 Asn Ala Pro Asn Asp Gln Thr Ser Lys Ile Val Ser Ile Thr Arg Glu 340 345 350 Ala Val Lys Leu Glu Asp Gln Ala Val Gln Lys Lys Val Leu Asn Leu 355 360 365 Lys Gln Asp Gly Lys Asp Tyr Lys Leu Gly Val Ile Glu Ile Pro Ala 370 375 380 Phe Tyr Leu Asp Phe Lys Ala Phe Arg Ala Gly Asp Pro Asp Tyr Lys 385 390 395 400 Ser Thr Thr Arg Asp Val Lys Lys Ile Leu Thr Glu Leu Gln Lys Glu 405 410 415 Lys Val Asp Gly Val Val Ile Asp Leu Arg Asn Asn Gly Gly Gly Ser 420 425 430 Leu Gln Glu Ala Thr Glu Leu Thr Ser Leu Phe Ile Asp Lys Gly Pro 435 440 445 Thr Val Leu Val Arg Asn Ala Asp Gly Arg Val Asp Val Leu Glu Asp 450 455 460 Glu Asn Pro Gly Ala Phe Tyr Lys Gly Pro Met Ala Leu Leu Val Asn 465 470 475 480 Arg Leu Ser Ala Ser Ala Ser Glu Ile Phe Ala Gly Ala Met Gln Asp 485 490 495 Tyr His Arg Ala Leu Ile Ile Gly Gly Gln Thr Phe Gly Lys Gly Thr 500 505 510 Val Gln Thr Ile Gln Pro Leu Asn His Gly Glu Leu Lys Leu Thr Leu 515 520 525 Ala Lys Phe Tyr Arg Val Ser Gly Gln Ser Thr Gln His Gln Gly Val 530 535 540 Leu Pro Asp Ile Asp Phe Pro Ser Ile Ile Asp Thr Lys Glu Ile Gly 545 550 555 560 Glu Ser Ala Leu Pro Glu Ala Met Pro Trp Asp Thr Ile Arg Pro Ala 565 570 575 Ile Lys Pro Ala Ser Asp Pro Phe Lys Pro Phe Leu Ala Gln Leu Lys 580 585 590 Ala Asp His Asp Thr Arg Ser Ala Lys Asp Ala Glu Phe Val Phe Ile 595 600 605 Arg Asp Lys Leu Ala Leu Ala Lys Lys Leu Met Glu Glu Lys Thr Val 610 615 620 Ser Leu Asn Glu Ala Asp Arg Arg Ala Gln His Ser Ser Ile Glu Asn 625 630 635 640 Gln Gln Leu Val Leu Glu Asn Thr Arg Arg Lys Ala Lys Gly Glu Asp 645 650 655 Pro Leu Lys Glu Leu Lys Lys Glu Asp Glu Asp Ala Leu Pro Thr Glu 660 665 670 Ala Asp Lys Thr Lys Pro Glu Asp Asp Ala Tyr Leu Ala Glu Thr Gly 675 680 685 Arg Ile Leu Leu Asp Tyr Leu Lys Ile Thr Lys Gln Val Ala Lys Gln 690 695 700 <210> 22 <211> 437 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 22 Met Leu His Leu Ser Arg Leu Thr Ser Leu Ala Leu Thr Ile Ala Leu 1 5 10 15 Val Ile Gly Ala Pro Leu Ala Phe Ala Asp Gln Ala Ala Pro Ala Ala 20 25 30 Pro Ala Thr Ala Ala Thr Thr Lys Ala Pro Leu Pro Leu Asp Glu Leu 35 40 45 Arg Thr Phe Ala Glu Val Met Asp Arg Ile Lys Ala Ala Tyr Val Glu 50 55 60 Pro Val Asp Asp Lys Ala Leu Leu Glu Asn Ala Ile Lys Gly Met Leu 65 70 75 80 Ser Asn Leu Asp Pro His Ser Ala Tyr Leu Gly Pro Glu Asp Phe Ala 85 90 95 Glu Leu Gln Glu Ser Thr Ser Gly Glu Phe Gly Gly Leu Gly Ile Glu 100 105 110 Val Gly Ser Glu Asp Gly Gln Ile Lys Val Val Ser Pro Ile Asp Asp 115 120 125 Thr Pro Ala Ser Lys Ala Gly Ile Gln Ala Gly Asp Leu Ile Val Lys 130 135 140 Ile Asn Gly Gln Pro Thr Arg Gly Gln Thr Met Thr Glu Ala Val Asp 145 150 155 160 Lys Met Arg Gly Lys Leu Gly Gln Lys Ile Thr Leu Thr Leu Val Arg 165 170 175 Asp Gly Gly Asn Pro Phe Asp Val Thr Leu Ala Arg Ala Thr Ile Thr 180 185 190 Val Lys Ser Val Lys Ser Gln Leu Leu Glu Ser Gly Tyr Gly Tyr Ile 195 200 205 Arg Ile Thr Gln Phe Gln Val Lys Thr Gly Asp Glu Val Ala Lys Ala 210 215 220 Leu Ala Lys Leu Arg Lys Asp Asn Gly Lys Lys Leu Asn Gly Ile Val 225 230 235 240 Leu Asp Leu Arg Asn Asn Pro Gly Gly Val Leu Gln Ser Ala Val Glu 245 250 255 Val Val Asp His Phe Val Thr Lys Gly Leu Ile Val Tyr Thr Lys Gly 260 265 270 Arg Ile Ala Asn Ser Glu Leu Arg Phe Ser Ala Thr Gly Asn Asp Leu 275 280 285 Ser Glu Asn Val Pro Leu Ala Val Leu Ile Asn Gly Gly Ser Ala Ser 290 295 300 Ala Ser Glu Ile Val Ala Gly Ala Leu Gln Asp Leu Lys Arg Gly Val 305 310 315 320 Leu Met Gly Thr Thr Ser Phe Gly Lys Gly Ser Val Gln Thr Val Leu 325 330 335 Pro Leu Asn Asn Glu Arg Ala Leu Lys Ile Thr Thr Ala Leu Tyr Tyr 340 345 350 Thr Pro Asn Gly Arg Ser Ile Gln Ala Gln Gly Ile Val Pro Asp Ile 355 360 365 Glu Val Arg Arg Ala Lys Ile Thr Asn Glu Ile Asp Gly Glu Tyr Tyr 370 375 380 Lys Glu Ala Asp Leu Gln Gly His Leu Gly Asn Gly Asn Gly Gly Ala 385 390 395 400 Asp Gln Pro Thr Gly Ser Arg Ala Lys Ala Lys Pro Met Pro Gln Asp 405 410 415 Asp Asp Tyr Gln Leu Ala Gln Ala Leu Ser Leu Leu Lys Gly Leu Ser 420 425 430 Ile Thr Arg Ser Arg 435 <210> 23 <211> 1242 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 23 Met Asp Val Ala Gly Asn Gly Phe Thr Val Ser Gln Arg Asn Arg Thr 1 5 10 15 Pro Arg Phe Lys Thr Thr Pro Leu Thr Pro Ile Ala Leu Gly Leu Ala 20 25 30 Leu Trp Leu Gly His Gly Ser Val Ala Arg Ala Asp Asp Asn Pro Tyr 35 40 45 Thr Pro Gln Val Leu Glu Ser Ala Phe Arg Thr Ala Val Ala Ser Phe 50 55 60 Gly Pro Glu Thr Ala Val Tyr Lys Asn Leu Arg Phe Ala Tyr Ala Asp 65 70 75 80 Ile Val Asp Leu Ala Ala Lys Asp Phe Ala Ala Gln Ser Gly Lys Phe 85 90 95 Asp Ser Ala Leu Lys Gln Asn Tyr Glu Leu Gln Pro Glu Asn Leu Thr 100 105 110 Ile Gly Ala Met Leu Gly Asp Thr Arg Arg Pro Leu Asp Tyr Ala Ser 115 120 125 Arg Leu Asp Tyr Tyr Arg Ser Arg Leu Phe Ser Asn Ser Gly Arg Tyr 130 135 140 Thr Thr Asn Ile Leu Asp Phe Ser Lys Ala Ile Ile Ala Asn Leu Pro 145 150 155 160 Ala Ala Lys Pro Tyr Thr Tyr Val Glu Pro Gly Val Ser Ser Asn Leu 165 170 175 Asn Gly Gln Leu Asn Ala Gly Gln Ser Trp Ala Gly Ala Thr Arg Asp 180 185 190 Trp Ser Ala Asn Ala Gln Thr Trp Lys Thr Pro Glu Ala Gln Val Asn 195 200 205 Ser Gly Leu Asp Arg Thr Asn Ala Tyr Tyr Ala Tyr Ala Leu Gly Ile 210 215 220 Thr Gly Lys Gly Val Asn Val Gly Val Leu Asp Ser Gly Ile Phe Thr 225 230 235 240 Glu His Ser Glu Phe Gln Gly Lys Asn Ala Gln Gly Gln Asp Arg Val 245 250 255 Gln Ala Val Thr Ser Thr Gly Glu Tyr Tyr Ala Thr His Pro Arg Tyr 260 265 270 Arg Leu Glu Val Pro Ser Gly Glu Phe Lys Gln Gly Glu His Phe Ser 275 280 285 Ile Pro Gly Glu Tyr Asp Pro Ala Phe Asn Asp Gly His Gly Thr Glu 290 295 300 Met Ser Gly Val Leu Ala Ala Asn Arg Asn Gly Thr Gly Met His Gly 305 310 315 320 Ile Ala Phe Asp Ala Asn Leu Phe Val Ala Asn Thr Gly Gly Ser Asp 325 330 335 Asn Asp Arg Tyr Gln Gly Ser Asn Asp Leu Asp Tyr Asn Ala Phe Met 340 345 350 Ala Ser Tyr Asn Ala Leu Ala Ala Lys Asn Val Ala Ile Val Asn Gln 355 360 365 Ser Trp Gly Gln Ser Ser Arg Asp Asp Val Glu Asn His Phe Gly Asn 370 375 380 Val Gly Asp Ser Ala Ala Gln Asn Leu Arg Asp Met Thr Ala Ala Tyr 385 390 395 400 Arg Pro Phe Trp Asp Lys Ala His Ala Gly His Lys Thr Trp Met Asp 405 410 415 Ala Met Ala Asp Ala Ala Arg Gln Asn Thr Phe Ile Gln Ile Ile Ser 420 425 430 Ala Gly Asn Asp Ser His Gly Ala Asn Pro Asp Thr Asn Ser Asn Leu 435 440 445 Pro Phe Phe Lys Pro Asp Ile Glu Ala Lys Phe Leu Ser Ile Thr Gly 450 455 460 Tyr Asp Glu Thr Ser Ala Gln Val Tyr Asn Arg Cys Gly Thr Ser Lys 465 470 475 480 Trp Trp Cys Val Met Gly Ile Ser Gly Ile Pro Ser Ala Gly Pro Glu 485 490 495 Gly Glu Ile Ile Pro Asn Ala Asn Gly Thr Ser Ala Ala Ala Pro Ser 500 505 510 Val Ser Gly Ala Leu Ala Leu Val Met Gln Arg Phe Pro Tyr Met Thr 515 520 525 Ala Ser Gln Ala Arg Asp Val Leu Leu Thr Thr Ser Ser Leu Gln Ala 530 535 540 Pro Asp Gly Pro Asp Thr Pro Val Gly Thr Leu Thr Gly Gly Arg Thr 545 550 555 560 Tyr Asp Asn Leu Gln Pro Val His Asp Ala Ala Pro Gly Leu Pro Gln 565 570 575 Val Pro Gly Val Val Ser Gly Trp Gly Leu Pro Asn Leu Gln Lys Ala 580 585 590 Met Gln Gly Pro Gly Gln Phe Leu Gly Ala Val Ala Val Ala Leu Pro 595 600 605 Ser Gly Thr Arg Asp Ile Trp Ala Asn Pro Ile Ser Asp Glu Ala Ile 610 615 620 Arg Ala Arg Arg Val Glu Asp Ala Ala Glu Gln Ala Thr Trp Ala Ala 625 630 635 640 Thr Lys Gln Gln Lys Gly Trp Leu Ser Gly Leu Pro Ala Asn Ala Ser 645 650 655 Ala Asp Asp Gln Phe Glu Tyr Asp Ile Gly His Ala Arg Glu Gln Ala 660 665 670 Thr Leu Thr Arg Gly Gln Asp Val Leu Thr Gly Ser Thr Tyr Val Gly 675 680 685 Ser Leu Val Lys Ser Gly Asp Gly Glu Leu Val Leu Glu Gly Gln Asn 690 695 700 Thr Tyr Ser Gly Ser Thr Trp Val Arg Gly Gly Lys Leu Ser Val Asp 705 710 715 720 Gly Ala Leu Thr Ser Ala Val Thr Val Asp Ser Ser Ala Val Gly Thr 725 730 735 Arg Asn Ala Asp Asn Gly Val Met Thr Thr Leu Gly Gly Thr Leu Ala 740 745 750 Gly Asn Gly Thr Val Gly Ala Leu Thr Val Asn Asn Gly Gly Arg Val 755 760 765 Ala Pro Gly His Ser Ile Gly Thr Leu Arg Thr Gly Asp Val Thr Phe 770 775 780 Asn Pro Gly Ser Val Tyr Ala Val Glu Val Gly Ala Asp Gly Arg Ser 785 790 795 800 Asp Gln Leu Gln Ser Ser Gly Val Ala Thr Leu Asn Gly Gly Val Val 805 810 815 Ser Val Ser Leu Glu Asn Ser Pro Asn Leu Leu Thr Ala Thr Glu Ala 820 825 830 Arg Ser Leu Leu Gly Gln Gln Phe Asn Ile Leu Ser Ala Ser Gln Gly 835 840 845 Ile Gln Gly Gln Phe Ala Ala Phe Ala Pro Asn Tyr Leu Phe Ile Gly 850 855 860 Thr Ala Leu Asn Tyr Gln Pro Asn Gln Leu Thr Leu Ala Ile Ala Arg 865 870 875 880 Asn Gln Thr Thr Phe Ala Ser Val Ala Gln Thr Arg Asn Glu Arg Ser 885 890 895 Val Ala Thr Val Ala Glu Thr Leu Gly Ala Gly Ser Pro Val Tyr Glu 900 905 910 Ser Leu Leu Ala Ser Asp Ser Ala Ala Gln Ala Arg Glu Gly Phe Lys 915 920 925 Gln Leu Ser Gly Gln Leu His Ser Asp Val Ala Ala Ala Gln Met Ala 930 935 940 Asp Ser Arg Tyr Leu Arg Glu Ala Val Asn Ala Arg Leu Gln Gln Ala 945 950 955 960 Gln Ala Leu Asp Ser Ser Ala Gln Ile Asp Ser Arg Asp Asn Gly Gly 965 970 975 Trp Val Gln Leu Leu Gly Gly Arg Asn Asn Val Ser Gly Asp Asn Asn 980 985 990 Ala Ser Gly Tyr Ser Ser Ser Thr Ser Gly Val Leu Leu Gly Leu Asp 995 1000 1005 Thr Glu Val Asn Asp Gly Trp Arg Val Gly Ala Ala Thr Gly Tyr 1010 1015 1020 Thr Gln Ser His Leu Asn Gly Gln Ser Ala Ser Ala Asp Ser Asp 1025 1030 1035 Asn Tyr His Leu Ser Val Tyr Gly Gly Lys Arg Phe Glu Ala Ile 1040 1045 1050 Ala Leu Arg Leu Gly Gly Ala Ser Thr Trp His Arg Leu Asp Thr 1055 1060 1065 Ser Arg Arg Val Ala Tyr Ala Asn Gln Ser Asp His Ala Lys Ala 1070 1075 1080 Asp Tyr Asn Ala Arg Thr Asp Gln Val Phe Ala Glu Ile Gly Tyr 1085 1090 1095 Thr Gln Trp Thr Val Phe Glu Pro Phe Ala Asn Leu Thr Tyr Leu 1100 1105 1110 Asn Tyr Gln Ser Asp Ser Phe Lys Glu Lys Gly Gly Ala Ala Ala 1115 1120 1125 Leu His Ala Ser Gln Gln Ser Gln Asp Ala Thr Leu Ser Thr Leu 1130 1135 1140 Gly Val Arg Gly His Thr Gln Leu Pro Leu Thr Ser Thr Ser Ala 1145 1150 1155 Val Thr Leu Arg Gly Glu Leu Gly Trp Glu His Gln Phe Gly Asp 1160 1165 1170 Thr Asp Arg Glu Ala Ser Leu Lys Phe Ala Gly Ser Asp Thr Ala 1175 1180 1185 Phe Ala Val Asn Ser Val Pro Val Ala Arg Asp Gly Ala Val Ile 1190 1195 1200 Lys Ala Ser Ala Glu Met Ala Leu Thr Lys Asp Thr Leu Val Ser 1205 1210 1215 Leu Asn Tyr Ser Gly Leu Leu Ser Asn Arg Gly Asn Asn Asn Gly 1220 1225 1230 Ile Asn Ala Gly Phe Thr Phe Leu Phe 1235 1240 <210> 24 <211> 450 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 24 Met Ser Ala Leu Tyr Met Ile Val Gly Thr Leu Val Ala Leu Gly Val 1 5 10 15 Leu Val Thr Phe His Glu Phe Gly His Phe Trp Val Ala Arg Arg Cys 20 25 30 Gly Val Lys Val Leu Arg Phe Ser Val Gly Phe Gly Met Pro Leu Leu 35 40 45 Arg Trp His Asp Arg Arg Gly Thr Glu Phe Val Ile Ala Ala Ile Pro 50 55 60 Leu Gly Gly Tyr Val Lys Met Leu Asp Glu Arg Glu Gly Glu Val Pro 65 70 75 80 Ala Asp Gln Leu Asp Gln Ser Phe Asn Arg Lys Thr Val Arg Gln Arg 85 90 95 Ile Ala Ile Val Ala Ala Gly Pro Ile Ala Asn Phe Leu Leu Ala Met 100 105 110 Val Phe Phe Trp Val Leu Ala Met Leu Gly Ser Gln Gln Val Arg Pro 115 120 125 Val Ile Gly Ala Val Glu Ala Asp Ser Ile Ala Ala Lys Ala Gly Leu 130 135 140 Thr Ala Gly Gln Glu Ile Val Ser Ile Asp Gly Glu Pro Thr Thr Gly 145 150 155 160 Trp Gly Ala Val Asn Leu Gln Leu Val Arg Arg Leu Gly Glu Ser Gly 165 170 175 Thr Val Asn Val Val Val Arg Asp Gln Asp Ser Ser Ala Glu Thr Pro 180 185 190 Arg Ala Leu Ala Leu Asp His Trp Leu Lys Gly Ala Asp Glu Pro Asp 195 200 205 Pro Ile Lys Ser Leu Gly Ile Arg Pro Trp Arg Pro Ala Leu Pro Pro 210 215 220 Val Leu Ala Glu Leu Asp Pro Lys Gly Pro Ala Gln Ala Ala Gly Leu 225 230 235 240 Lys Thr Gly Asp Arg Leu Leu Ala Leu Asp Gly Gln Ala Leu Gly Asp 245 250 255 Trp Gln Gln Val Val Asp Leu Val Arg Val Arg Pro Asp Thr Lys Ile 260 265 270 Val Leu Lys Val Glu Arg Glu Gly Ala Gln Ile Asp Val Pro Val Thr 275 280 285 Leu Ser Val Arg Gly Glu Ala Lys Ala Ala Gly Gly Tyr Leu Gly Ala 290 295 300 Gly Val Lys Gly Val Glu Trp Pro Pro Ser Met Val Arg Glu Val Ser 305 310 315 320 Tyr Gly Pro Leu Ala Ala Ile Gly Glu Gly Ala Lys Arg Thr Trp Thr 325 330 335 Met Ser Val Leu Thr Leu Glu Ser Leu Lys Lys Met Leu Phe Gly Glu 340 345 350 Leu Ser Val Lys Asn Leu Ser Gly Pro Ile Thr Ile Ala Lys Val Ala 355 360 365 Gly Ala Ser Ala Gln Ser Gly Val Ala Asp Phe Leu Asn Phe Leu Ala 370 375 380 Tyr Leu Ser Ile Ser Leu Gly Val Leu Asn Leu Leu Pro Ile Pro Val 385 390 395 400 Leu Asp Gly Gly His Leu Leu Phe Tyr Leu Val Glu Trp Val Arg Gly 405 410 415 Arg Pro Leu Ser Asp Arg Val Gln Gly Trp Gly Ile Gln Ile Gly Ile 420 425 430 Serum Leu Val Val Gly Val Met Leu Leu Ala Leu Val Asn Asp Leu Gly 435 440 445 Leo Silver 450 <210> 25 <211> 246 <212> PRT <213> ӫ ٵ (Pseudomonas fluorescens) <400> 25 Met Lys Gln His Arg Leu Ala Ala Ala Val Ala Leu Val Ser Leu Val 1 5 10 15 Leu Ala Gly Cys Asp Ser Gln Thr Ser Val Glu Leu Lys Thr Pro Ala 20 25 30 Gln Lys Ala Ser Tyr Gly Ile Gly Leu Asn Met Gly Lys Ser Leu Ala 35 40 45 Gln Glu Gly Met Asp Asp Leu Asp Ser Lys Ala Val Ala Gln Gly Ile 50 55 60 Glu Asp Ala Val Gly Lys Lys Glu Gln Lys Leu Lys Asp Asp Glu Leu 65 70 75 80 Val Glu Ala Phe Ala Ala Leu Gln Lys Arg Ala Glu Glu Arg Met Thr 85 90 95 Lys Met Ser Glu Glu Ser Ala Ala Ala Gly Lys Lys Phe Leu Glu Asp 100 105 110 Asn Ala Lys Lys Asp Gly Val Val Thr Thr Ala Ser Gly Leu Gln Tyr 115 120 125 Lys Ile Val Lys Lys Ala Asp Gly Ala Gln Pro Lys Pro Thr Asp Val 130 135 140 Val Thr Val His Tyr Thr Gly Lys Leu Thr Asn Gly Thr Th...

Claims

1. A recombinant fusion protein, comprising: (i) an N-terminal fusion partner, wherein the N-terminal fusion partner is a Pseudomonas fluorescens DnaJ-like protein having the amino acid sequence of SEQ ID NO: 2, or a Pseudomonas fluorescens FkbP having the amino acid sequence of SEQ ID NO: 25; (ii) a target polypeptide selected from hPTH1-34, preproinsulin that can be processed into insulin or an insulin analogue, Glp1, Glp2, IGF-1, exenatide SEQ ID NO: 37, teduglutide SEQ ID NO: 38, pramlintide SEQ ID NO: 39, ziconotide SEQ ID NO: 40, becaplermin SEQ ID NO: 42, enfuvirtide SEQ ID NO: 43, nesiritide SEQ ID NO: 44, N-met-GCSF, GCSF, Plasmodium falciparum circumsporozoite protein, and IFN-β, wherein the insulin analogue is insulin glargine, insulin aspart, insulin lispro, insulin glulisine, insulin detemir, or insulin degludec; and (iii) a 10- to 50-amino acid-long linker containing a protease cleavage site between the N-terminal fusion partner and the target polypeptide.

2. The recombinant fusion protein of claim 1, wherein the protease cleavage site is recognized by a lyase selected from the group consisting of enterokinase, trypsin, chymotrypsin, factor Xa, and furin.

3. The recombinant fusion protein of claim 1, wherein the target polypeptide is preproinsulin that can be processed into insulin or an insulin analogue, and wherein the recombinant fusion protein comprises an amino acid sequence selected from SEQ ID NOs: 122-129.

4. The recombinant fusion protein of claim 1, wherein the amino acid sequence of the linker is selected from: (i) from N to C terminus, the (G4S)2 spacer sequence SEQ ID NO: 59, the hexahistidine affinity tag SEQ ID NO: 242, and the protease cleavage site DDDDK SEQ ID NO: 13; (ii) from N to C terminus, the (G4S)2 spacer sequence SEQ ID NO: 59, the hexahistidine affinity tag SEQ ID NO: 242, and the protease cleavage site RKR; (iii) from N to C terminus, the (G4S)2 spacer sequence SEQ ID NO: 59, the hexahistidine affinity tag SEQ ID NO: 242, and the protease cleavage site RRR; (iv) From the N-terminus to the C-terminus, the (G4S)2 spacer sequence SEQ ID NO:59, the hexahistidine affinity tag SEQ ID NO:242, and the protease cleavage site LVPR; and (v) SEQ ID NO:

226.

5. The recombinant fusion protein of claim 1, wherein the target polypeptide is preproinsulin that can be processed into insulin or an insulin analogue, and wherein the preproinsulin comprises a C-peptide having an amino acid sequence selected from SEQ ID NO:97, SEQ ID NO:98, SEQ ID NO:99, and SEQ ID NO:

100.

6. The recombinant fusion protein of claim 5, wherein the protease cleavage site is recognized by a lyase selected from the group consisting of enterokinase, trypsin, chymotrypsin, factor Xa, and furin.

7. The recombinant fusion protein of claim 6, wherein the amino acid sequence of the linker is selected from: (i) From the N-terminus to the C-terminus, the (G4S)2 spacer sequence SEQ ID NO:59, the hexahistidine affinity tag SEQ ID NO:242, and the protease cleavage site DDDDK SEQ ID NO:13; (ii) From the N-terminus to the C-terminus, the (G4S)2 spacer sequence SEQ ID NO:59, the hexahistidine affinity tag SEQ ID NO:242, and the protease cleavage site RKR; (iii) From the N-terminus to the C-terminus, the (G4S)2 spacer sequence SEQ ID NO:59, the hexahistidine affinity tag SEQ ID NO:242, and the protease cleavage site RRR; (iv) From the N-terminus to the C-terminus, the (G4S)2 spacer sequence SEQ ID NO:59, the hexahistidine affinity tag SEQ ID NO:242, and the protease cleavage site LVPR; and (v) SEQ ID NO:

226.

8. The recombinant fusion protein of any one of claims 1, 2, 5, and 6, wherein the linker comprises a spacer and an affinity tag.

9. The recombinant fusion protein of claim 8, wherein the amino acid sequence of the linker is selected from: (i) From the N-terminus to the C-terminus, the (G4S)2 spacer sequence SEQ ID NO:59, the hexahistidine affinity tag SEQ ID NO:242, and the protease cleavage site DDDDK SEQ ID NO:13; (ii) From the N-terminus to the C-terminus, the (G4S)2 spacer sequence SEQ ID NO:59, the hexahistidine affinity tag SEQ ID NO:242, and the protease cleavage site RKR; (iii) From the N-terminus to the C-terminus, the (G4S)2 spacer sequence SEQ ID NO:59, the hexahistidine affinity tag SEQ ID NO:242, and the protease cleavage site RRR; (iv) From the N-terminus to the C-terminus, the (G4S)2 spacer sequence SEQ ID NO:59, the hexahistidine affinity tag SEQ ID NO:242, and the protease cleavage site LVPR; and (v) SEQ ID NO:

226.

10. The recombinant fusion protein of claim 8, wherein the linker comprises: (i) A spacer selected from (G4S)1, (G4S)2, (G4S)3, (G4S)4, and (G4S)5; and (ii) An affinity tag selected from the maltose binding protein tag, the polyhistidine tag, the FLAG tag, the Myc tag, the HA tag, and the Nus tag.

11. The recombinant fusion protein of claim 8, wherein the N-terminal fusion partner is the Pseudomonas fluorescens DnaJ-like protein with the amino acid sequence SEQ ID NO:2, and wherein the target polypeptide is hPTH1-34.

12. The recombinant fusion protein of claim 11, wherein the target polypeptide is hPTH1-34, and wherein the amino acid sequence of the recombinant fusion protein is SEQ ID NO:

45.

13. A method for producing a target polypeptide, comprising: (i) Culturing a microbial host cell transformed with an expression vector comprising an expression construct, wherein the expression construct comprises a nucleotide sequence encoding the recombinant fusion protein of any one of claims 1 to 12; (ii) Inducing the host cell of step (i) to express the recombinant fusion protein; (iii) Purifying the recombinant fusion protein expressed in the host cell induced in step (ii); and (iv) Cleaving the recombinant fusion protein purified in step (iii) by incubating with a lytic enzyme that recognizes the protease cleavage site in the linker to release the target polypeptide; thereby obtaining the target polypeptide.

14. The method of claim 13, further comprising measuring the expression level of the recombinant fusion protein expressed in step (ii), measuring the amount of the recombinant fusion protein purified in step (iii), or measuring the amount of the correctly released target polypeptide obtained in step (iv), or a combination thereof.

15. The method of claim 14, wherein the amount of the target polypeptide obtained in step (iii) or step (iv) is from 0.1 g / L to 25 g / L.

16. The method of claim 14, wherein the correctly released target polypeptide obtained is soluble, intact, or both.

17. An expression vector for expressing the recombinant fusion protein of any one of claims 1 to 12, wherein the expression vector contains a nucleotide sequence encoding the recombinant fusion protein.

Citation Information

Patent Citations

  • New gram-positive expression control sequences

    EP0207459A2

  • Expression of mammalian proteins in Pseudomonas fluorescens

    US20060040352A1

  • Codon optimization method

    US20070292918A1

  • Method for rapidly screening microbial hosts to identify certain strains with improved yield and / or quality in the expression of heterologous proteins

    US20080269070A1

  • Method for Rapidly Screening Microbial Hosts to Identify Certain Strains with Improved Yield and / or Quality in the Expression of Heterologous Proteins

    US20100137162A1