Protection of spy tag-containing periplasmic fusion proteins from protease tsp and ompt degradation

By using protease-deficient E. coli strains and incorporating SpyTag binding motifs into fusion proteins, the challenges of low yields and incomplete folding in periplasmic expression are addressed, resulting in efficient and functional protein production.

JP2025084751AInactive Publication Date: 2025-06-03BIO-RAD ABD SEROTECH GMBH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025016011
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-03-18
Filing Date
2025-02-03
Publication Date
2025-06-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing methods for expressing SpyTag fusion proteins in the periplasm of E. coli often result in low yields or incomplete protein folding due to the reducing conditions in the cytoplasm, which prevent the formation of disulfide bonds.

Method used

A periplasmic fusion protein construct comprising a binding motif (SpyTag, SpyTag002, or SpyTag003) added to or embedded within a first protein, along with a nucleic acid construct and a method for producing this fusion protein in protease-deficient E. coli strains, allowing for proper folding and increased yields.

Benefits of technology

The approach enables the efficient production of fully folded and functional periplasmic fusion proteins, overcoming the limitations of cytoplasmic expression and proteolytic cleavage in the periplasm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025084751000002
    Figure 2025084751000002
  • Figure 2025084751000003
    Figure 2025084751000003
  • Figure 2025084751000004
    Figure 2025084751000004
Patent Text Reader

Abstract

To provide a method of producing periplasmic fusion proteins and a mutant Escherichia coli strain used in the method.SOLUTION: The present invention provides periplasmic fusion proteins comprising a binding motif added to a first protein or embedded within an amino acid sequence of the first protein, nucleic acid constructs encoding the periplasmic fusion proteins, vectors comprising the nucleic acid constructs, and methods of producing the periplasmic fusion proteins. The present invention also provides protease deficient host cells for producing the periplasmic fusion proteins.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-reference to related applications This application claims the benefit of U.S. Provisional Application No. 62 / 819,758, filed Mar. 18, 2019, the contents of which are incorporated herein by reference.

Background Art

[0002] Some techniques enable the covalent attachment or ligation of polypeptides at specific predetermined sites. By way of example, the SpyTag / SpyCatcher (Non-Patent Document 22) system uses the concept of spontaneous isopeptide bond formation in naturally occurring proteins to covalently attach one polypeptide to another. Such proteins containing isopeptide bonds Streptococcus pyogenes The domain derived from protein FbaB is divided into two parts. One part, SpyTag (SEQ ID NO: 1), is a 13-amino acid peptide containing a part of the autocatalytic center (e.g., an aspartic acid residue). The other part, SpyCatcher, is an 116-amino acid protein domain containing the other part of the center (e.g., lysine) and a neighboring catalytic glutamate or aspartic acid residue. By mixing these two polypeptides, the autocatalytic center is restored, leading to the formation of an isopeptide bond, thereby covalently linking SpyTag to SpyCatcher (Non-Patent Document 32). Further engineering led to a short SpyCatcher having only 84 amino acids, an optimized SpyTag002 (SEQ ID NO: 2), and a fast-reacting SpyCatcher002 (Non-Patent Documents 17 and 12, which are hereby incorporated by reference in their entirety). Further engineering led to another optimized SpyTag003 (SEQ ID NO: 36) and SpyCatcher003 due to a reaction close to the diffusion limit (Non-Patent Document 13, which is hereby incorporated by reference in its entirety).

[0003] Two polypeptides to be ligated to each other are typically produced as fusion proteins with each polypeptide having a part of the autocatalytic center from FbaB such that an isopeptide bond is formed when the polypeptides are mixed. For example, a first polypeptide linked to SpyTag and a second polypeptide linked to SpyCatcher are produced separately as fusion proteins, and when the first and second polypeptides are mixed together, an isopeptide bond is formed between SpyTag and SpyCatcher. This type of fusion protein is produced in bacteria. Many fusion proteins are produced in the cytoplasm of bacteria, but due to the reducing conditions, disulfide-bridged proteins do not form disulfide bonds and thus are not properly folded when expressed in the reducing environment of the cytoplasm. However, many important classes of proteins contain disulfide bonds. An example is antibody fragments. The functional expression of antibody Fv and Fab fragments is achieved by directing the transport of the expressed protein chains to the periplasm of Gram-negative bacteria such as Escherichia coli, which have oxidizing conditions and allow the formation of disulfide bonds (Plueckthum A, 1990, Antibodies from Escherichia coli, Nature 347, 497-498), which opened the way to the field of antibody engineering. The bacterial recombinant expression of many SpyTag fusion proteins has been described in the literature, but they have mostly been expressed exclusively in the bacterial cytoplasm (Non-Patent Document 13). There are only a few examples of periplasmic bacterial expression of SpyTag fusion proteins (Non-Patent Documents 4 and 3), and the yields were low or unknown. Non-Patent Documents 13, 4, and 3 are hereby incorporated by reference in their entirety.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

[0005] [Non-Patent Document 1] IMGT definitions according to Lefranc M.-P., De RK, Tomar N. Immunoinformatics of the V, C and G domains: IMGT (registered trademark) definitive system for IG, TR and IgSF, MH and MhSF, Immunoinformatics: From Biology to Informatics, 2014, vol. 1184 2nd edition Springer, NY Humana Press (pg. 59 - 107) [Non-Patent Document 2] Abe, H., Rie, W., Yonemura, H., Yamada, S., Goto, M., and Kamiya, N., (2013), Split Spy0128 as a Potent Scaffold for Protein Cross-Linking and Immobilization. Bioconjugate Chem., 24(2), 242 - 250 [Non-Patent Document 3] Alam et al., 2017, Synthetic Modular Antibody Construction Using the SpyTag / SpyCatcher Protein Ligase System. Chembiochem. 18(22), 2217 - 2221 [Non-Patent Document 4] Alves, N.J., Turner, K.B., Daniele, M.A., Oh, E., Medintz, I.L., Walper, S.A., Bacterial nanobioreactors-directing enzyme packaging into bacterial outer membrane vesicles. ACS Appl Mater Interfaces, 2015; 7: 24963-24972

Non-Patent Document 5

Non-Patent Document 6

Non-Patent Document 7

Non-Patent Document 8

Non-Patent Document 18

Non-Patent Document 19

Non-Patent Document 20

Non-Patent Document 21

Non-Patent Document 22

Non-Patent Document 28

Non-Patent Document 29

Non-Patent Document 30

Non-Patent Document 31

Non-Patent Document 32

[0006] A periplasmic fusion protein comprising a binding motif (e.g., SpyTag, SpyTag002, or SpyTag003) added to a first protein (e.g., an antigen-binding fragment) or embedded within the amino acid sequence of the first protein, a nucleic acid construct encoding the periplasmic fusion protein, a vector comprising the nucleic acid construct, and a method for producing such a periplasmic fusion protein are provided. Also provided are mutant protease-deficient Escherichia coli cells for producing the periplasmic fusion protein.

[0007] In one embodiment, the periplasmic fusion protein comprises a binding motif added to a first protein or embedded within the amino acid sequence of the first protein, the binding motif comprising SEQ ID NO: 1 or a sequence having at least 60% sequence identity with SEQ ID NO: 1. In certain embodiments, the binding motif comprises SEQ ID NO: 2 or a sequence having at least 70% sequence identity with SEQ ID NO: 2. In some embodiments, the binding motif comprises SEQ ID NO: 36 or a sequence having at least 70% sequence identity with SEQ ID NO: 36. In some embodiments, the binding motif is added to the N-terminus or C-terminus of the first protein, either directly or via a linker sequence. In some embodiments, the first protein is a protein structural domain. In certain embodiments, the linker sequence comprises a purification tag. In some embodiments, the binding motif comprises SEQ ID NO: 1 or a sequence having at least 60% sequence identity with SEQ ID NO: 1 and is added to the C-terminus of the first protein, either directly or via a linker sequence. In some embodiments, the binding motif is protease-sensitive. In some embodiments, the binding motif is protease-resistant. In some embodiments, the first protein is an antigen-binding fragment. In some embodiments, the first protein is an antigen-binding fragment and the antigen-binding fragment is a Fab, scFv, or scFab. In some embodiments, the antigen-binding fragment is a Fab. In certain embodiments, the periplasmic fusion protein further comprises a purification tag added to the N-terminus or C-terminus of the binding motif. In certain embodiments, the binding motif is a linker sequence that links the C-terminus of the first protein to the N-terminus of a second protein or the N-terminus of the first protein to the C-terminus of a second protein.

[0008] Also provided is a nucleic acid construct comprising a polynucleotide sequence encoding a periplasmic fusion protein. Also provided is a vector comprising the nucleic acid construct.

[0009] Also provided is a method for producing a periplasmic fusion protein comprising a binding motif added to a first protein or embedded within the amino acid sequence of the first protein, the binding motif comprising a sequence having at least 60% sequence identity to SEQ ID NO: 1 or SEQ ID NO: 1, the binding motif comprising a sequence having at least 70% sequence identity to SEQ ID NO: 2 or SEQ ID NO: 2, or the binding motif comprising a sequence having at least 70% sequence identity to SEQ ID NO: 36 or SEQ ID NO: 36. In some embodiments, the method comprises culturing an E. coli host cell transformed with a vector comprising a nucleic acid encoding the periplasmic fusion protein in a culture medium under conditions effective to express the periplasmic fusion protein, and recovering the periplasmic fusion protein from the E. coli host cell. In some embodiments, the binding motif of the fusion protein expressed in the E. coli host cell is protease resistant. In certain embodiments, the binding motif of the fusion protein expressed in the E. coli host cell is protease sensitive. In such embodiments, the E. coli host cell is a mutant cell lacking one or more periplasmic proteases. In some embodiments, the mutant E. coli cells used in the method lack the functional chromosomal gene tsp encoding protease Tsp (tail-specific protease). In some embodiments, the mutant E. coli cells used in the method lack the functional chromosomal genes tsp and ompT encoding protease Tsp and OmpT (outer membrane protein T), respectively.

[0010] Also provided are Escherichia coli TG1, TG1F-, XL1 Blue, MC1061, SS320, BL21, JM83, JM109, HB2151, W3110, or Cmax5 alpha strains that are deficient in the functional chromosomal gene tsp encoding protease Tsp. In some embodiments, such mutant E. coli strains contain a nucleic acid encoding a periplasmic fusion protein comprising a binding motif, and the binding motif comprises SEQ ID NO: 1 or a sequence having at least 60% sequence identity with SEQ ID NO: 1. Also provided are E. coli TG1, TG1F-, XL1 Blue, MC1061, SS320, BL21, JM83, JM109, HB2151, W3110, or Cmax5 alpha strains that are deficient in the functional chromosomal genes tsp and ompT encoding protease Tsp and ompT, respectively. In some embodiments, such mutant E. coli strains contain a nucleic acid encoding a periplasmic fusion protein comprising a binding motif, and the binding motif comprises SEQ ID NO: 2 or a sequence having at least 70% sequence identity with SEQ ID NO: 2. In certain embodiments, such mutant E. coli strains contain a nucleic acid encoding a periplasmic fusion protein comprising a binding motif, and the binding motif comprises SEQ ID NO: 36 or a sequence having at least 70% sequence identity with SEQ ID NO: 36. BRIEF DESCRIPTION OF THE DRAWINGS

[0011]

Figure 1

Figure 2-1

Figure 2-2

Figure 2-3

Figure 2-4

Figure 3

Figure 4

Figure 5

Figure 6-1

Figure 6-2

Figure 7

Figure 8

Figure 9

Modes for Carrying Out the Invention

[0012] Provided are a periplasmic fusion protein comprising a binding motif (i.e., SpyTag, SpyTag002, or SpyTag003) added to a first protein (e.g., an antigen-binding fragment or a protein structural domain) or embedded within the amino acid sequence of the first protein, a nucleic acid construct encoding the periplasmic fusion protein, a vector comprising the nucleic acid construct, and a method for producing the periplasmic fusion protein. Also provided is a protease-deficient host cell for producing the periplasmic fusion protein.

[0013] Fusion proteins containing SpyTag, SpyTag002, or SpyTag003 binding motifs were found to be digested by periplasmic proteases when expressed in the periplasm of E. coli. When SpyTag is directly linked to the C-terminus or N-terminus of a protein, a protein structural domain, or a protein structural domain fragment without a linker sequence, a fusion protein is obtained in which SpyTag is substantially resistant to periplasmic proteases while being produced in E. coli. When SpyTag, SpyTag002, or SpyTag003 is linked to the N-terminus or C-terminus of a protein or a protein domain by a linker sequence, it has also been found that SpyTag, SpyTag002, or SpyTag003 is sensitive to periplasmic proteases during the expression of the fusion protein in bacteria. For such protease-sensitive fusion proteins, E. coli host cells lacking periplasmic proteases responsible for SpyTag, SpyTag002, or SpyTag003 cleavage have been generated.

[0014] Definitions Unless otherwise indicated, the following terms used in this application, including the specification and claims, have the definitions given below. As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise.

[0015] "Antibody" refers to an immunoglobulin, complex (e.g., fusion), or fragment thereof. This term includes polyclonal or monoclonal antibodies of the IgA, IgD, IgE, IgG, and IgM classes from in vitro antibody libraries, including those derived from antibody-producing cell lines or in vitro in natural or genetically modified or synthetic forms such as humanized, human, single-chain, chimeric, synthetic, recombinant, hybrid, mutant, grafted, and other in vitro-produced antibodies, but is not limited thereto. "Antibody" also includes complex forms including fusion proteins having an immunoglobulin portion, but is not limited thereto.

[0016] As used herein, the term "antigen-binding fragment" refers to a protein that includes the antigen-binding portion of an antibody such as Fab. Other antigen-binding fragments include variable fragments (Fv), disulfide-stabilized Fv fragments (dsFv), single-chain variable fragments (scFv), or single-chain Fab fragments (scFab). Further examples of antigen-binding fragments include the variable domain of a heavy-chain antibody (VHH), single-domain antibodies (sdAb), or monovalent forms of antigen-binding fragments that include antigen-binding sites that include shark variable new antigen receptors (VNAR). Additionally, non-antibody scaffolds such as variable lymphocyte receptors (VLR), affimers, affibodies, darpins, anticalins, monobodies, or antigen-binding peptides can also be considered "antigen-binding fragments".

[0017] The term "binding motif" refers to a protein sequence that is added to a polypeptide and enables the formation of a covalent bond to another polypeptide. Non-limiting examples of binding motifs include SpyTag, SpyTag002, and SpyTag003 sequences. The SpyTag sequence forms a covalent bond with the SpyCatcher sequence. The binding motif can be fused to the N-terminus, C-terminus, or embedded within the amino acid sequence of the polypeptide. One or more linker sequences (e.g., glycine / serine-rich linker) or one or more protein tags can be adjacent to the binding motif to enhance proximity for the reaction, enhance the flexibility of the fusion polypeptide, or purify and / or detect the polypeptide. When the binding motif links two or more proteins, the N-terminus and C-terminus of the binding motif are adjacent by one or more linker sequences to enhance proximity for the reaction, enhance the flexibility of the fusion polypeptide, and purify and / or detect the polypeptide.

[0018] The term "prokaryotic" refers to bacterial cells (e.g., Escherichia or Salmonellarefers to prokaryotic cells such as bacterial cells having it), or prokaryotic phages or bacterial spores. The term "eukaryotic system" refers to eukaryotic cells including cells of animals, plants, fungi and protists, as well as eukaryotic viruses such as retroviruses, adenoviruses, baculoviruses. Prokaryotes and eukaryotic systems may be collectively referred to as "expression systems".

[0019] The term "expression cassette" is used herein to refer to a functional unit constructed in a vector for the purpose of expressing a recombinant polypeptide in the periplasm. The expression cassette contains a promoter, a transcription terminator sequence, a ribosome binding site, and DNA encoding a fusion protein. Depending on the expression system (e.g., enhancers and polyadenylation signals of eukaryotic expression systems), other genetic components can be added to the expression cassette.

[0020] As used herein, the term "vector" preferably refers to a nucleic acid molecule that replicates itself intracellularly, which transfers the inserted nucleic acid molecule within and / or between host cells. Typically, a vector is a circular DNA containing an origin of replication, a selectable marker, and / or a viral packaging signal, as well as other regulatory elements. Vectors, vector DNA, plasmid DNA, phagemid DNA are used interchangeably in the description of the present invention. This term includes vectors that function primarily for the insertion of DNA or RNA into cells, replication vectors that function primarily for the replication of DNA or RNA, and expression vectors that function for the transcription and / or translation of DNA or RNA. Also included are vectors that provide two or more of the above functions.

[0021] As used herein, the term "expression vector" is a polynucleotide that, when introduced into a suitable host cell, directs the transcription and translation of one or more polypeptides under suitable conditions. The term "expression vector" refers to a vector that directs the expression of the relevant polypeptide fused in-frame with a binding motif.

[0022] The terms "nucleic acid" and "polynucleotide" as used herein are used interchangeably. They refer to polymeric forms of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. Non-limiting examples of polynucleotides include the coding or non-coding regions of genes or gene fragments, loci defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Polynucleotides may include modified nucleotides such as methylated nucleotides and nucleotide analogs. Where present, modifications to the nucleotide structure are made either before or after assembly of the nucleotide polymer.

[0023] As used herein, the term "amino acid" refers to natural and / or non-natural or synthetic amino acids, both D and L optical isomers, amino acid analogs, and peptidomimetics.

[0024] The terms "polypeptide", "peptide" and "protein" as used herein are used interchangeably herein to refer to polymers of amino acids of any length.

[0025] As used herein, the term "protein structural domain" or "domain" refers to a semi-autonomous and compact folding unit having a distinct hydrophobic core (Non-Patent Document 9). A domain evolves, functions, and exists independently from other parts of the protein chain and is a conserved part of a given protein sequence and structure. Each domain forms a compact three-dimensional structure that folds and is stable independently. Examples of protein structural domains include, but are not limited to, the constant domains of human IgG1 and maltose-binding protein (MBP) (e.g., CH1 including the first two amino acids of the hinge region). The "C-terminus" of a protein structural domain refers to the last amino acid annotated as being structured (i.e., visible in the crystal structure) in the crystal structure of the same or closely related protein structural domain (i.e., having at least 70% sequence identity) in the Protein Data Bank (Non-Patent Document 5). A protein structural domain can be shortened by up to 10 amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) on the N or C terminus without losing its folding ability, i.e., the ability to remain as a structural domain.

[0026] As used herein, the term "host cell" includes individual cells or cell cultures that can be or have been recipients of the disclosed expression constructs. Host cells include the progeny of a single host cell. The progeny may not be exactly identical to the original parent cell due to natural, accidental, or intentional mutations.

[0027] Two nucleic acid sequences or polypeptides are said to be "identical" if, when the sequences of nucleotides or amino acid residues in the two sequences are aligned to maximize matches as described below, they are the same. The term "identical" or percent "identity" with respect to two or more nucleic or polypeptide sequences refers to sequences or subsequences that are the same or have a specified percentage of the same amino acid residues or nucleotides when compared and aligned to maximize matches over a comparison window, as measured using one of the following sequence comparison algorithms or by manual alignment and visual inspection. When used with respect to a protein or peptide, the percentage of sequence identity often differs at residue positions that are not identical, often due to conservative amino acid substitutions, where an amino acid residue is substituted with another amino acid residue having similar chemical properties (e.g., charge or hydrophobicity), such that the functional properties of the molecule are not changed. When sequences differ in conservative substitutions, the percent sequence identity is adjusted upward to correct for the conservative nature of the substitution. Means for making this adjustment are well known to those of skill in the art. Typically, conservative substitutions are scored as partial mismatches rather than complete mismatches, thereby increasing the percentage of sequence identity. Thus, for example, if a score of 1 is assigned to identical amino acids and a score of 0 is assigned to non-conservative substitutions, a score between 0 and 1 is assigned to conservative substitutions. Scoring of conservative substitutions is calculated, for example, according to the algorithm of Meyers & Miller, Computer Applic. Biol. Sci. 4:11-17 (1988), as implemented in the program PC / GENE (Intelligenetics, Mountain View, CA, USA).

[0028] When aligned and compared to maximize matches across a comparison window, sequences are "substantially identical" to each other if they have a specified percentage of nucleotides or amino acid residues that are the same (e.g., at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity over a specified region, or if no region is specified, over the entire specified sequence).

[0029] For sequence comparison, typically one sequence acts as a reference sequence, against which a test sequence is compared. When using a sequence comparison algorithm, the test and reference sequences are input into a computer, and subsequence coordinates are specified if necessary, as well as the sequence algorithm program parameters. Either the default program parameters can be used or alternative parameters can be specified. The sequence comparison algorithm then calculates the percent sequence identity of the test sequence relative to the reference sequence based on the program parameters.

[0030] As used herein, a "comparison window" refers to any segment of a number of contiguous positions selected from the group consisting of 10 to 600, about 10 to about 300, about 10 to about 150, and the sequences are compared to the reference sequence at the same number of contiguous positions after the two sequences are optimally aligned. A comparison window can also be the entire length of either the reference or test sequence.

[0031] The percent sequence identity and sequence similarity can be determined using the BLAST 2.0 algorithm described in Altschul et al. (J. Mol. Biol. 215:403-10, 1990). Software for performing BLAST 2.0 analysis is publicly available through the National Center for Biotechnology Information (Worldwide Website: ncbi.nlm.nih.gov / ). This algorithm first identifies high-scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with words of the same length in the database sequence. T is referred to as the neighborhood word score threshold (Altschul et al., supra). This initial neighborhood word hit serves as a seed to initiate a search for a long HSP that contains it. The word hits are extended in both directions along each sequence as long as their cumulative alignment score increases. The extension of the word hits in each direction stops when the cumulative alignment score drops from its maximum achieved value by an amount X, or when the cumulative score goes below zero due to the accumulation of one or more negative-scoring residue alignments, or when the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLAST program, by default, uses a word length (W) of 11, the BLOSUM62 scoring matrix (see Henikoff & Henikoff, Proc Natl Acad Sci USA 89:10915 (1989)), an alignment (B) of 50, an expectation value (E) of 10, M = 5, N = -4, and comparison of both strands.

[0032] The BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, e.g., Karlin & Altschul, Proc Nat'l Acad Sci USA 90:5873-5787 (1993)). One measure of similarity provided by the BLAST algorithm is the minimum total probability (P(N)), which provides an indication of the probability that a matching between two nucleotide or amino acid sequences occurs by chance. For example, if the minimum total probability in a comparison of a test nucleic acid and a reference nucleic acid is less than about 0.2, more preferably less than about 0.01, and most preferably less than about 0.001, the nucleic acid is considered to be similar to the reference sequence.

[0033] Periplasmic fusion protein In one embodiment, the periplasmic fusion protein comprises a binding motif (e.g., SpyTag, SpyTag002, or SpyTag003) added to a first protein (Figures 1A - 1N) or embedded within the amino acid sequence of the first protein (e.g., Figure 1P). As used herein, "periplasmic fusion protein" refers to a protein produced in the periplasm of a suitable bacterial host cell and comprises two polypeptides, the first of which can be any protein and the second of which is a SpyTag, SpyTag002, or SpyTag003 binding motif. The binding motif can be added to the N - terminus or C - terminus of the first protein either directly (Figures 1A, 1D, 1G, 1H, 1K, 1N) or via a linker sequence (Figures 1B, 1C, 1E, 1F, 1I, 1J, 1L, 1M). In certain embodiments, the first protein is a protein structural domain. In some embodiments, the first protein is an antigen - binding fragment (e.g., Fab, scFv, or scFab). As used herein, "binding motif" refers to a peptide sequence that, when added to a protein expressed in the periplasm, facilitates the formation of a covalent bond via protein ligation to another binding motif (e.g., SpyCatcher, SpyCatcher002, or SpyCatcher003) added to another polypeptide when the two binding motifs come into contact with each other. The covalent bond between the two binding motifs is formed either spontaneously or with the aid of an enzyme. For example, the binding motif can form a covalent bond with another binding motif on the N - terminus of an Fc fragment or with a multimerization binding motif. Examples of Fc fragments having another binding motif are described in co - pending U.S. Patent Application No. 62 / 819,748 (Antigen - Binding Fragments and Subclasses Bound to Multiple Fc Isotypes, filed March 18, 2019, BRL.123P), and examples of multimerization binding motifs are described in co - pending U.S. Patent Application No. 62 / 819,753 (Antigen - Binding Proteins, filed March 18, 2019, BRL.129P), each of which is incorporated by reference in its entirety.In some embodiments, the binding motif comprises SEQ ID NO: 1 (i.e., SpyTag) or a sequence having at least 60% sequence identity with SEQ ID NO: 1. In some embodiments, the binding motif comprises SEQ ID NO: 2 (i.e., SpyTag002) or a sequence having at least 70% sequence identity with SEQ ID NO: 2. In certain embodiments, the binding motif comprises SEQ ID NO: 36 (i.e., SpyTag003) or a sequence having at least 78% sequence identity with SEQ ID NO: 36.

[0034] In some examples, when a binding motif (e.g., SpyTag) is directly added to the first protein (i.e., without a linker sequence, FIGS. 1A, 1D, 1G, 1H, 1K, 1N), a periplasmic fusion protein that is substantially protease-insensitive, i.e., resistant to cleavage by periplasmic proteases during expression, is produced.

[0035] In some embodiments, the periplasmic fusion protein comprises at least one linker sequence between the first protein and the binding motif. As used herein, "linker sequence" or "linker" refers to a peptide or polypeptide comprising one or more amino acid residues (e.g., 1, 2, 3, 4, 5, 10 or more amino acid residues) linked by peptide bonds. Such linkers provide the rotational freedom that allows each component of the fusion protein to interact with its intended target without being hindered. These linkers are a mixture of glycine and serine such as -(GGGS) n -(where n is 1, 2, 3, 4, or 5). Other suitable peptide / polypeptide linker sequences optionally include naturally occurring or non-naturally occurring peptides or polypeptides. Optionally, the peptide or polypeptide linker sequence is a flexible peptide or polypeptide (FIGS. 1B, 1E, 1I, 1L). Flexible peptides / polypeptides include amino acid sequences Gly-Ser, Ala-Ser, Gly-Ser, Gly 4 -Ser, (Gly 4 -Ser) 2 , (Gly4 -Ser) 3 、(Gly 4 -Ser) 4 、(Gly 4 -Ser)Gly 4 、 2 -Gly-Ala-Gly-SerGly 4 -Ser, Gly-(Gly 4 -Ser) 2 、Gly 4 -Ser-Gly, Gly-Ser 3 、Gly-SerGly 4 -Ser, among others, but not limited thereto. Other suitable peptide linker sequences optionally include the TEV linker ENLYFQG, a linear epitope recognized by tobacco etch virus protease. Examples of the peptide / polypeptide include, but are not limited to, GSENLYFQGSG. Other suitable peptide linker sequences include helix-forming linkers such as Ala-(Glu-Ala-Ala-Lys) n -Ala (n = 1-5). In some embodiments, the linker sequence is a GAP (Gly Ala Pro) sequence. In some embodiments, the linker sequence includes a purification tag (FIGS. 1C, 1F, 1J, 1M). Examples of the purification tag include, but are not limited to, polyhistidine or His-tag and FLAG (registered trademark) tag (i.e., the amino acid sequence DYKDDDDK, where D is aspartic acid, Y is tyrosine, and K is lysine). In certain embodiments, the linker sequence includes a binding motif (FIGS. G and N), and optionally includes a purification tag and / or a flexible linker sequence added to either or both of the C-terminus and N-terminus of the binding motif. In some embodiments, a sequence of 1 to 50 amino acid residues can be used as the linker. In some embodiments, the linker is protease-resistant (i.e., periplasmic expression of the polypeptide having the linker in the host cell occurs without cleavage of the linker by protease). When two or more linker sequences are used between the protein and the binding motif, the two or more linker sequences may be the same or different.

[0036] In some embodiments where the periplasmic fusion protein comprises a linker sequence between the first protein and the binding motif, the binding motifs (i.e., SpyTag, SpyTag002, and SpyTag003) are protease-sensitive, i.e., the binding motifs are cleaved by one or more E. coli proteases during periplasmic expression. In certain embodiments, cleavage of the binding motif by periplasmic proteases depends on the linker length (i.e., the number of amino acids), but not on the linker amino acid composition.

[0037] In some embodiments, the linker sequence comprises a SpyTag, SpyTag002, or SpyTag003 binding motif. In this embodiment, the binding motif links the N-terminus of the first protein to the C-terminus of the second protein (FIG. 1G), or the C-terminus of the first protein to the N-terminus of the second protein (FIG. 1N). The second protein includes, but is not limited to, an antigen-binding fragment, a fluorescent protein such as green fluorescent protein, an enzyme such as horseradish peroxidase or other peroxidase, alkaline phosphatase, luciferase, a split fluorescent protein, and MBP. In certain embodiments, the linker sequence comprising SpyTag, SpyTag002, or SpyTag003 further comprises a purification tag or a flexible linker sequence between the binding motif and either or both of the first and second proteins.

[0038] In certain embodiments, the periplasmic fusion protein has a purification tag added to the N-terminus of the binding motif (FIGS. 1D, 1E), the C-terminus of the binding motif (FIGS. 1K, 1L), or both the N-terminus and C-terminus of the binding motif (FIGS. 1F and 1M).

[0039] Nucleic acid construct Also provided are nucleic acid constructs encoding a periplasmic fusion protein, without or with a linker between the first protein and the binding motif and / or between the binding motif and the second protein of the periplasmic fusion protein. Such nucleic acids are optionally present in an expression vector in a prokaryotic host cell.

[0040] Typically, a polynucleotide sequence encoding a Fab fused to a binding motif at the C-terminus encodes two peptides, namely the L-chain and H-chain of the Fab. A binding motif such as SpyTag is fused either directly or via one or more linkers to either the L-chain or the H-chain. The Fab expression cassette comprises a bicistronic vector that produces one mRNA encoding both the L-chain and the H-chain, at least one of which is fused to a binding motif. Also, both the H-chain and the L-chain have signal peptides for directing their transport to the periplasm.

[0041] Nucleic acid constructs are typically introduced into various vectors. The vectors described herein generally contain the transcriptional or translational control sequences necessary for expressing the fusion protein. Suitable transcriptional or translational control sequences include, but are not limited to, an origin of replication, a promoter, an enhancer, a repressor binding region, a transcription start site, a ribosome binding site, a translation start site, and termination sites for transcription and translation.

[0042] The origin of replication (commonly referred to as the ori sequence) enables replication of the vector in a suitable host cell. The choice of ori depends on the type of host cell used and / or the gene package. When the host cell is prokaryotic, the expression vector typically contains an ori sequence that directs autonomous replication of the vector in the prokaryotic cell. Preferred prokaryotic oris are capable of directing vector replication in bacterial cells. Non-limiting examples of oris in this category include pMB1, pUC, and other E. coli origins.

[0043] As used herein, a "promoter" is a DNA region that can bind to RNA polymerase under specific conditions and initiate transcription of a coding region located downstream (3' direction) of the promoter. It can be either constitutive or inducible. Generally, a promoter sequence is bound at its 3' end by a transcription start site and extends upstream (5' direction) to include the minimum number of bases or elements necessary to initiate transcription at a detectable level above background. Within the promoter sequence, there is a transcription start site as well as protein-binding domains responsible for binding RNA polymerase.

[0044] The selection of a promoter depends largely on the host cell into which the vector is introduced. For prokaryotic cells, various strong promoters are known in the art. Preferred promoters are the lac promoter, Trc promoter, T7 promoter, and pBAD promoter.

[0045] In the construction of the target vector, a termination sequence related to the protein coding sequence can also be inserted at the 3' end of the sequence to be transcribed in order to provide a polyadenylation and / or transcription termination signal for the mRNA. The terminator sequence preferably contains one or more transcription termination sequences (e.g., polyadenylation sequences) and can also be lengthened by including additional DNA sequences to further disrupt transcription read-through. Preferred terminator sequences (or termination sites) of the present invention have a gene followed by a transcription termination sequence of either its own termination sequence or a heterologous termination sequence. Examples of such termination sequences include various yeast transcription termination sequences known in the art and stop codons linked to mammalian polyadenylation sequences that are widely available. When the terminator contains a gene, it is advantageous to use a gene encoding a detectable or selectable marker, thereby providing a means by which the presence and / or absence of the terminator sequence (and thus the corresponding inactivation and / or activation of the transcription unit) can be detected and / or selected.

[0046] In addition to the above elements, the vector contains a selectable marker (e.g., a gene encoding a protein necessary for the survival or growth of a host cell transformed with the vector), and such a marker gene is retained on another polynucleotide sequence co-introduced into the host cell. Only the host cells into which the selectable gene has been introduced will survive and / or grow under selective conditions. Typical selectable genes encode (a) proteins that confer resistance to antibiotics or other toxins, such as ampicillin, kanamycin, neomycin, zeocin, G418, methotrexate, etc., (b) proteins that complement auxotrophic deficiencies, or (c) proteins that supply essential nutrients not available from complex media. The selection of an appropriate marker gene depends on the host cell, and the appropriate genes for different hosts are known in the art.

[0047] In one embodiment, the expression vector is a shuttle vector capable of replicating in at least two unrelated host systems. To facilitate such replication, the vector generally contains at least two origins of replication, one being functional in each host system. Typically, the shuttle vector is replicable in both eukaryotic and prokaryotic host systems. This allows for the detection of protein expression in eukaryotic host (expression cell type) and the amplification of the vector in prokaryotic host (amplification cell type). Preferably, one origin of replication is derived from SV40 or 2u, and one is derived from pUC, but any suitable origin known in the art that directs the replication of the vector can be used. When the vector is a shuttle vector, the vector preferably contains at least two selectable markers, one for the expression cell type and one for the amplification cell type. Any selectable marker known in the art or described herein is used to function in the expression system utilized.

[0048] The vectors encompassed by the present invention can be obtained using recombinant cloning methods and / or by chemical synthesis. A vast number of recombinant cloning techniques such as PCR, restriction endonuclease digestion, and ligation are well known in the art and need not be described in detail herein. One skilled in the art can also obtain the desired vector by any synthetic means available in the art using the sequence data provided herein or sequence data in public or proprietary databases. Further, using well-known restriction and ligation techniques, appropriate sequences can be excised from various DNA sources and incorporated in operable relationship with the exogenous sequences to be expressed according to the embodiments described herein.

[0049] Method for producing a fusion protein Also provided, optionally, is a method for producing a periplasmic fusion protein comprising a binding motif added to a first protein or first and second proteins via one or more linkers. In one embodiment, the method comprises culturing an E. coli host cell transformed with a vector comprising a nucleic acid encoding the periplasmic fusion protein in a culture medium under conditions effective to express the periplasmic fusion protein in the host cell. Any suitable strain of E. coli can be used to produce the periplasmic fusion protein. One such strain of E. coli used for protein (e.g., antibody, antibody fragment, or MBP) expression is the TG1 strain. The TG1 strain is based on E. coli K-12 (genotype gln44Vthi-1Δ(lac-proAB)Δ(mcrB-hsdSM)5(r K -m K -)F´[traD36 proAB+lacIq ZlacΔM15]), which is commonly used for the expression of antibodies and antibody fragments in the periplasm (see, e.g., Non-Patent Document 16 for experiments showing periplasmic expression). In some embodiments, the E. coli host cell strain is TG1F- which is F pilus-deficient (genotype glnV44thi-1Δ(lac-proAB) Δ(mcrB-hsdSM)5(r K -mK-)). In certain embodiments, E. coli host cell lines include, but are not limited to, XL1 Blue, MC1061, SS320, BL21, JM83, JM109, HB2151, W3110, Cmax5 alpha, and any E. coli strain suitable for the functional expression of antibody fragments in E. coli. In some embodiments, the E. coli host cell is a mutant cell lacking one or more periplasmic proteases. In some embodiments, the mutant E. coli cell lacks the functional chromosomal gene tsp encoding protease Tsp (tail-specific protease). In some embodiments, the mutant E. coli cell lacks the functional chromosomal genes tsp and ompT encoding protease Tsp and OmpT (outer membrane protein T), respectively. Genes for proteases are "knocked out" by deleting or replacing the gene with a foreign DNA sequence such as a gene encoding antibiotic resistance. One such process for knocking out protease genes is described in Non-Patent Document 8. In some embodiments, the protease gene is modified to produce a mutant protease with abolished or reduced proteolytic activity (Non-Patent Document 15). In some embodiments, the expression of tsp or ompT protease is inhibited by antisense morpholino, antisense peptide nucleic acid, or other antisense nucleotide oligomers (Non-Patent Document 11), resulting in a decrease or elimination of tsp or ompT protease activity in E. coli, respectively. In some embodiments, both tsp and ompT proteases are inhibited by antisense morpholino, antisense peptide nucleic acid, or other antisense nucleotide oligomers that result in a decrease or elimination of tsp and ompT protease activity. In some embodiments, tsp or ompT protease activity is decreased or eliminated by chemical protease inhibitors including, but not limited to, phenylmethanesulfonyl fluoride or p-toluenesulfonyl fluoride (Non-Patent Document 20), by peptides or small proteins such as aprotinin (Non-Patent Document 6) or metal cations (Non-Patent Document 25).In some embodiments, both the tsp and ompT protease activities are reduced or eliminated by the chemical inhibitors described above.

[0050] The mutant Escherichia coli TG1F strains SK4 (DSM33004) and SK13 (DSM33005) were deposited on January 8, 2019 with the Leibniz Institute DSMZ - German Collection of Microorganisms and Cell Cultures, Inhoffenstraβe 7B, 38124 Braunschweig, Germany. The mutant Escherichia coli TG1 strain SK4 with DSMZ accession number DSM33004 is deficient in tsp, and the mutant Escherichia coli TG1 strain SK13 with DSMZ accession number DSM33005 is deficient in both tsp and ompT.

[0051] Vectors containing nucleic acids encoding periplasmic fusion proteins are used to transform cells using standard techniques, for example, by using chemical methods (Green R, Rogers EJ. Transformation of chemically competent E.coli. Methods Enzymol 2013; 529:329 - 36) or by electroporation. In some embodiments where the binding motif is SEQ ID NO: 1 (SpyTag), the periplasmic fusion protein is transformed into the mutant Escherichia coli TG1F strain with DSMZ accession number DSM33004 that is deficient in tsp. In some embodiments where the binding motif is SEQ ID NO: 2 (SpyTag002) or SEQ ID NO: 36 (SpyTag003), the periplasmic fusion protein is transformed into the mutant Escherichia coli TG1F strain with DSMZ accession number DSM33005 that is deficient in both tsp and ompT.

[0052] Cells that can express one or more markers can survive / grow / proliferate under certain artificially imposed conditions, such as the addition of toxins or antibiotics to the culture medium, due to the properties (e.g., antibiotic resistance) conferred by the polypeptide / gene or polypeptide components of the selection system incorporated therein. Cells that cannot express one or more markers cannot survive / grow / proliferate under the artificially imposed conditions.

[0053] In the methods described herein, any suitable selection system can be used. Typically, the selection system is based on including in the vector one or more genes that confer resistance to known antibiotics, such as tetracycline, chloramphenicol, kanamycin, or ampicillin resistance genes. Cells that grow in the presence of the relevant antibiotic can be selected because they express both the gene conferring resistance to the antibiotic and the desired protein.

[0054] In one embodiment, the method further comprises culturing the transformed cells in a medium, thereby expressing the periplasmic fusion protein.

[0055] The method can also use an inducible expression system or a constitutive promoter to express the periplasmic fusion protein.

[0056] Any suitable medium can be used to culture the transformed cells. The medium can be adapted to a particular selection system, e.g., the medium can contain an antibiotic to allow only cells that have been successfully transformed to grow in the medium.

[0057] Next, the expressed fusion protein is recovered from the periplasm of the host cell by lysing the bacteria, either by whole cell lysis or periplasmic lysis. The method can further include one or more steps for extracting and purifying the periplasmic fusion protein. The periplasmic fusion protein can be separated from the cell extract by suitable purification procedures including, but not limited to, protein A chromatography, protein L chromatography, thiophilic, mixed mode resins, nickel nitrilotriacetic acid (Ni-NTA) resin for His-tag, Strepor-Tactin® or Strepor XT resin for Strep-tag®, FLAG®-tag, hydroxyapatite chromatography, gel electrophoresis, dialysis, ammonium sulfate, ethanol or PEG fractionation / precipitation, ion exchange membranes, expanded bed adsorption chromatography, or simulated moving bed chromatography.

[0058] In some embodiments, the method further includes a step of measuring the expression level of the periplasmic fusion protein after purification.

[0059] Additional disclosure and claimable subject matter Item 1 A periplasmic fusion protein comprising a binding motif added to a first protein or embedded within the amino acid sequence of the first protein, wherein the binding motif comprises SEQ ID NO: 1 or a sequence having at least 60% sequence identity with SEQ ID NO: 1.

[0060] Item 2 The periplasmic fusion protein according to Item 1, wherein the binding motif is added to the N-terminus of the first protein, directly or via a linker sequence.

[0061] Item 3 The periplasmic fusion protein according to Item 1, wherein the binding motif is added to the C-terminus of the first protein, directly or via a linker sequence.

[0062] Item 4 The linker array is the periplasmic fusion protein according to Item - 2 or 3, which contains a purification tag.

[0063] Item 5 The binding motif is directly added to the C - terminus of the protein structure domain of the first protein, and the binding motif is protease - resistant. The periplasmic fusion protein according to Item 3.

[0064] Item 6 The protein structure domain is a human scFv single - chain antibody fragment with a truncated C - terminus within the FR4 region. The periplasmic fusion protein according to Item 5.

[0065] Item 7 The binding motif is added to the C - terminus of the protein structure domain in the first protein via a first or second amino acid linker. The periplasmic fusion protein according to Item 3.

[0066] Item 8 The binding motif is added to the C - terminus at IMGT position 121 of the human heavy - chain CH1 antibody domain via a 2, 3, or 4 - amino - acid linker. The periplasmic fusion protein according to any one of Item 3.

[0067] Item 9 The binding motif is added to the C - terminus at IMGT position 121 of the human light - chain constant antibody domain via a 2, 3, or 4 - amino - acid linker. The periplasmic fusion protein according to Item 3.

[0068] Item 10 The periplasmic fusion protein according to any one of Items 1 - 9, further comprising a purification tag added to the N - terminus or C - terminus of the binding motif.

[0069] Item 11 The binding motif connects the C - terminus of the first protein to the N - terminus of the second protein or the N - terminus of the first protein to the C - terminus of the second protein. The periplasmic fusion protein according to any one of Items 1 - 10.

[0070] A nucleic acid construct comprising a polynucleotide sequence encoding the periplasmic fusion protein according to any one of items 1 to 11 of item 12.

[0071] Item 13 A vector comprising the nucleic acid construct according to item 12.

[0072] Item 14 A method for producing a periplasmic fusion protein, culturing an Escherichia coli host cell transformed with a vector containing a nucleic acid encoding the periplasmic fusion protein in a culture medium under conditions effective for expressing the periplasmic fusion protein, recovering the periplasmic fusion protein from the Escherichia coli host cell, wherein the periplasmic fusion protein comprises a binding motif added to a first protein or embedded within the amino acid sequence of the first protein, wherein the binding motif comprises the sequence of SEQ ID NO: 1 or a sequence having at least 60% sequence identity with SEQ ID NO: 1, wherein the Escherichia coli host cell a) a mutation in the Tsp gene encoding the mutant Tsp protein, a mutation that reduces or abolishes protease activity, or b) a mutation in the Tsp gene or the regulatory sequence of the Tsp gene that reduces or abolishes the expression of the Tsp protein, or c) one or more deletions in a region of the bacterial chromosome that reduce or abolish the Tsp protein activity; or d) an inhibitor or inactivator that reduces or abolishes the Tsp protease activity, or an inhibitor of Tsp protease expression A method in which the Tsp protein activity is reduced or abolished as compared to the wild-type cells resulting therefrom.

[0073] Item 15 A method for producing a periplasmic fusion protein, culturing an Escherichia coli host cell transformed with a vector containing a nucleic acid encoding the periplasmic fusion protein in a culture medium under conditions effective for expressing the periplasmic fusion protein, comprising recovering the periplasmic fusion protein from the E. coli host cell, wherein the periplasmic fusion protein comprises a binding motif added to a first protein or embedded within the amino acid sequence of the first protein, wherein the binding motif comprises SEQ ID NO: 2 or a sequence having at least 70% sequence identity with SEQ ID NO: 2, or SEQ ID NO: 36 or a sequence having at least 78% sequence identity with SEQ ID NO: 36, wherein the E. coli host cell a) is a mutation of the Tsp gene encoding a mutant Tsp protein, a mutation that reduces or abolishes protease activity, or a mutation of the Tsp gene or regulatory sequence of the Tsp gene that reduces or abolishes the expression of the Tsp protein, or one or more deletions in a region of the bacterial chromosome that reduces or abolishes the Tsp protein activity, b) is a mutation of the ompT gene encoding a mutant ompT protein, a mutation that reduces or abolishes protease activity, or a mutation of the ompT gene or regulatory sequence of the ompT gene that reduces or abolishes the expression of the ompT protein, or one or more deletions in a region of the bacterial chromosome that reduces or abolishes the ompT protein activity A method in which the Tsp protein activity and ompT protein activity are reduced or abolished as compared to wild-type cells resulting therefrom.

[0074] Item 16. A method for producing a periplasmic fusion protein, comprising: culturing an E. coli host cell transformed with a vector containing a nucleic acid encoding the periplasmic fusion protein in a culture medium under conditions effective for expressing the periplasmic fusion protein, comprising recovering the periplasmic fusion protein from the E. coli host cell, wherein the periplasmic fusion protein comprises a binding motif added to a first protein or embedded within the amino acid sequence of the first protein, The binding motif includes the sequence of SEQ ID NO: 2 or a sequence having at least 70% sequence identity with SEQ ID NO: 2, or the sequence of SEQ ID NO: 36 or a sequence having at least 78% sequence identity with SEQ ID NO: 36, The Escherichia coli host cell a) an inhibitor or inactivator of Tsp protease, or an inhibitor of Tsp expression, and b) an inhibitor or inactivator of ompT protease, or an inhibitor of ompT expression A method in which the Tsp protein activity and the ompT protein activity are reduced or eliminated as compared with wild-type cells produced thereby.

[0075] Item 17 The method according to any one of Items 14 to 16, wherein the binding motif is added to the N-terminus of the first protein, directly or via a linker sequence.

[0076] Item 18 The method according to any one of Items 14 to 16, wherein the binding motif is added to the C-terminus of the first protein, directly or via a linker sequence.

[0077] Item 19 The method according to any one of Items 14 to 18, wherein the first protein is a protein structure domain.

[0078] Item 20 The method according to any one of Items 17 to 19, wherein the linker sequence includes a purification tag.

[0079] Item 21 The method according to Item 18, wherein the binding motif includes the amino acid sequence shown in SEQ ID NO: 1 or a sequence having at least 60% sequence identity with SEQ ID NO: 1, and is directly added to the C-terminus of the first protein.

[0080] Item 22 The method according to any one of Items 14 to 21, wherein the binding motif is protease-sensitive.

[0081] Item 23 The method according to any one of Items 14 to 21, wherein the binding motif is protease-resistant.

[0082] Item 24 The first protein is an antigen-binding fragment, and the antigen-binding fragment includes Fab, scFv, or scFab. The method according to any one of Items 14 to 23.

[0083] Item 25 The antigen-binding fragment is Fab. The method according to Item 24.

[0084] Item 26 The periplasmic fusion protein further includes a purification tag added to the N-terminus or C-terminus of the binding motif. The method according to any one of Items 14 to 25.

[0085] Item 27 The binding motif links the C-terminus of the first protein to the N-terminus of the second protein, or the N-terminus of the first protein to the C-terminus of the second protein. The method according to any one of Items 14 to 25.

[0086] Item 28 The E. coli host cell is a mutant E. coli TG1F strain having DSM deposit number 33004 deposited on January 8, 2019. The method according to Item 14.

[0087] Item 29 The E. coli host cell is a mutant E. coli TG1F strain having DSM deposit number 33005 deposited on January 8, 2019. The method according to Item 15 or 16.

[0088] Item 30 A mutation of the Tsp gene encoding the mutant Tsp protein, a mutation that reduces or abolishes protease activity, or a mutation of the Tsp gene or regulatory sequence of the Tsp gene that reduces or abolishes the expression of the Tsp protein, or one or more deletions in the region of the bacterial chromosome that reduce or abolish the Tsp protein activity, compared to wild-type cells, the Tsp protein activity is reduced or abolished in E. coli TG1, TG1F-, XL1Blue, MC1061, SS320, BL21, JM83, JM109, HB2151, W3110, or Cmax5 alpha strain.

[0089] An Escherichia coli strain according to item 30, comprising a nucleic acid encoding a periplasmic fusion protein containing a binding motif, wherein the binding motif comprises the sequence of SEQ ID NO: 1 or a sequence having at least 60% sequence identity with SEQ ID NO: 1.

[0090] Item 32 a) A mutation in the Tsp gene encoding a mutant Tsp protein, which is a mutation that reduces or abolishes protease activity, or a mutation in the Tsp gene or regulatory sequence of the Tsp gene that reduces or abolishes the expression of the Tsp protein, or one or more deletions in a region of the bacterial chromosome that reduces or abolishes the Tsp protein activity. b) A mutation in the ompT gene encoding a mutant ompT protein, which is a mutation that reduces or abolishes protease activity, or a mutation in the ompT gene or regulatory sequence of the ompT gene that reduces or abolishes the expression of the ompT protein, or one or more deletions in a region of the bacterial chromosome that reduces or abolishes the ompT protein activity. An Escherichia coli strain TG1, TG1F-, XL1Blue, MC1061, SS320, BL21, JM83, JM109, HB2151, W3110, or Cmax5 alpha strain in which the Tsp protein activity and ompT protein activity are reduced or abolished as compared with the wild-type cells generated by the following:

[0091] Item 33 A nucleic acid encoding a periplasmic fusion protein containing a binding motif comprising the sequence of SEQ ID NO: 2 or a sequence having at least 70% sequence identity with SEQ ID NO: 2, or A nucleic acid encoding a periplasmic fusion protein containing a binding motif comprising the sequence of SEQ ID NO: 36 or a sequence having at least 78% sequence identity with SEQ ID NO: 36. An Escherichia coli strain according to item 32, comprising the same.

[0092] Item 34 An Escherichia coli strain according to any one of items 30 to 33, wherein the binding motif is protease-sensitive.

[0093] Item 35. The E. coli strain according to any one of Items 30 to 33, wherein the binding motif is protease-resistant.

[0094] Item 36 a) A binding motif comprising the sequence of SEQ ID NO: 1 or a sequence having at least 60% sequence identity with SEQ ID NO: 1, which is added to a first protein or embedded within the amino acid sequence of the first protein, for the expression of a periplasmic fusion protein comprising the binding motif, A mutation in the Tsp gene encoding a mutant Tsp protein, which is a mutation that decreases or abolishes protease activity, or a mutation in the Tsp gene or regulatory sequence of the Tsp gene that decreases or abolishes the expression of the Tsp protein, or one or more deletions in a region of the bacterial chromosome that decreases or abolishes the Tsp protein activity, resulting in a decrease or abolition of Tsp protein activity as compared to wild-type cells, b) A binding motif comprising the sequence of SEQ ID NO: 2 or a sequence having at least 70% sequence identity with SEQ ID NO: 2, or the sequence of SEQ ID NO: 36 or a sequence having at least 78% sequence identity with SEQ ID NO: 36, which is added to a first protein or embedded within the amino acid sequence of the first protein, for the expression of a periplasmic fusion protein comprising the binding motif, i) A mutation in the Tsp gene encoding a mutant Tsp protein, which is a mutation that decreases or abolishes protease activity, or a mutation in the Tsp gene or regulatory sequence of the Tsp gene that decreases or abolishes the expression of the Tsp protein, or one or more deletions in a region of the bacterial chromosome that decreases or abolishes the Tsp protein activity, ii) A mutation in the ompT gene encoding a mutant ompT protein, which is a mutation that decreases or abolishes protease activity, or a mutation in the ompT gene or regulatory sequence of the ompT gene that decreases or abolishes the expression of the ompT protein, or one or more deletions in a region of the bacterial chromosome that decreases or abolishes the ompT protein activity A mutant Escherichia coli strain in which the Tsp protein activity and ompT protein activity are decreased or lost as compared with the wild-type cells generated thereby.

[0095] Item 37. Mutant Escherichia coli TG1F strain having DSM deposit number 33004 or 33005, both deposited on January 8, 2019.

[0096] Example The following examples are for illustrative purposes only and are not limiting. Those skilled in the art will readily recognize various non-essential parameters that can be changed or modified to produce essentially the same or similar results.

[0097] Example 1 - Periplasmic expression of the Fab-X-SpyTag fusion protein where X is a linker A gene encoding a human antibody fragment in Fab format having a SpyTag at the C-terminus of the shortened heavy chain (i.e., Fab-SpyTag construct) was cloned into an expression vector having signal sequences in the H and L chains of the Fab. This directs the nascent chain to the periplasm by bacterial transport. The Fab gene used in this example encoded a non-covalent heterodimer of a light chain (without a C-terminal cysteine) with a shortened heavy chain having a VH domain, a CH1 domain, and the first 4 amino acids of the hinge region (up to the first hinge cysteine, which is not included). Escherichia coli TG1F (F-episome free, Bio-Rad) was then transformed with such a vector. Fab constructs derived from 5 different antibodies were tested. The partial sequences of the constructs are shown in Figure 2. The transformants were cultured in 250 mL of 2xYT broth with 0.1% glucose and chloramphenicol. After growing at 37 °C for 1 hour, the culture was induced with 0.8 mM IPTG. Expression was allowed to proceed at 30 °C for about 16 hours. The culture was centrifuged and the cells were frozen at -80 °C. The cells were lysed with BugBuster lysis buffer (Millipore-Sigma). The fusion protein was then purified by affinity chromatography (e.g., by Ni-NTA chromatography for fusion proteins with a hexahistidine tag or by Strep-Tactin® chromatography for fusion proteins with a Strep-tag®) and buffer-exchanged into 3×PBS. The purity of the fusion protein was measured by SDS-PAGE using a 4-20% polyacrylamide gel (Bio-Rad Mini-PROTEAN TGX) and Coomassie® staining under non-reducing conditions. Further, all purified Fab fragments were tested for functionality by ELISA (at 2 μg / ml) using a Fab antigen (5 μg / ml in PBS coated on the surface of microtiter plate wells overnight at 4 °C). Binding of the Fab fragment to its antigen was detected using an HRP-labeled anti-Fab (STAR126P, Bio-Rad) or anti-histidine tag (MCA1396P, Bio-Rad) secondary antibody and a QuantaBlu fluorogenic substrate (Thermo Fisher).

[0098] The first attempt to purify the Fab-SpyTag fusion protein was unsuccessful. A construct containing a FLAG-SpyTag-His peptide sequence at the C-terminus of the Fab heavy chain (SEQ ID NO: 4) could not be purified, which was the same as all other constructs having a purification tag (His-tag or Strep-Tag®) at the C-terminus of SpyTag (SEQ ID NOs: 3-5 and 8-10). On the other hand, constructs containing a His-SpyTag or His-SpyTag-FLAG peptide sequence (SEQ ID NOs: 6 and 7, respectively) at the C-terminus of the Fab heavy chain could be purified, but these constructs were not reactive in subsequent SpyTag-SpyCatcher protein ligation reactions. To test the protein ligation of the SpyTag portion of the fusion protein with SpyCatcher, each fusion protein (final concentration 15 μM) was mixed with SpyCatcher (final concentration 20 μM) in 1xPBS buffer and allowed to couple for 2 hours at room temperature. SpyCatcher was produced by bacterial cytoplasmic expression and purified via Ni NTA as described in Non-Patent Document 32. SDS-PAGE was used to test for the appearance of new bands on the gel corresponding to the SpyTag fusion-SpyCatcher coupling product. Since no new SpyTag fusion protein-SpyCatcher bands of the correct size were observed for any of the fusion proteins, it was suggested that none of the fusion proteins bound to SpyCatcher and that the SpyTag portion of the fusion protein was neither intact nor fully functional. Without being bound by theory, Applicants hypothesized that SpyTag is sensitive to cleavage by one or more periplasmic proteases.

[0099] To determine where cleavage occurred in the fusion protein, the expression products of the Fab-FLAG-SpyTag-His (SEQ ID NO: 4) and Fab-His-SpyTag-FLAG (SEQ ID NO: 7) fusion proteins were analyzed by Western blot analysis (SDS PAGE using reducing sample buffer (Bio-Rad), AnyKD TGX gel (Bio-Rad), transfer to PVDF membrane (Bio-Rad)). The expression products of the fusion proteins were analyzed by Western blot analysis using an HRP-labeled anti-FLAG® antibody (Sigma A8592) or an HRP-labeled anti-histidine tag antibody (Bio-Rad MCA1396P). The results of the Western blot are shown in Figure 3.

[0100] Results: As shown in the Western blot of Figure 3, the expression product of the Fab-FLAG-SpyTag-His fusion protein was recognized by the labeled anti-FLAG® antibody but not by the anti-histidine tag antibody, suggesting that the histidine tag-containing portion was cleaved from the C-terminus of the fusion protein, i.e., cleavage occurred after the FLAG sequence. The expression product of the Fab-His-SpyTag-Flag fusion protein was recognized by the labeled anti-histidine tag antibody but not by the labeled anti-Flag antibody, suggesting that the Flag-containing portion of the fusion protein was cleaved from the C-terminus of the fusion protein, i.e., cleavage occurred after the histidine tag.

[0101] The masses of the light and heavy chain peptides of the expression product of Fab-His-SpyTag (SEQ ID NO: 6) purified by Ni-NTA were determined using MALDI-TOF mass spectrometry (4800 MALDI TOF / TOF Analyzer, AB Sciex). Samples were desalted (ZipTipC4, Merck Millipore) and co-crystallized with sinapinic acid. Masses were determined in linear mode from 5000 to 50000 m / z. 4000 laser shots were applied to one spectrum. Protein Standard I (Bruker) was used for mass calibration. The results of the mass spectrometry are shown below.

[0102] Light chain: Calculated mass (full length) 22691 Da Detected mass (m / z) 22691 Da Heavy chain: Calculated mass (full length) 26437 Da Detected mass (m / z) 25416 Da Calculated mass (-9aa) 25403 Da

[0103] The mass spectrometry results indicated that nine amino acids were cleaved from the C-terminus of the Fab-His-SpyTag fusion protein. Thus, nine amino acid residues from the C-terminus of SpyTag (which is a 13-amino acid length: AHIVMVDAYKPTK, SEQ ID NO: 1) were cleaved by proteases in the Escherichia coli periplasm after the valine at amino acid position 4. Without wishing to be bound by theory, Applicant believes that since SpyTag was cleaved in all constructs, the cleavage of SpyTag is independent of the amino acid sequences before and after SpyTag.

[0104] Example 2 - Periplasmic expression of various maltose-binding protein-SpyTag fusion proteins A gene encoding a maltose-binding protein (MBP) having either FLAG-SpyTag-His tag (SEQ ID NO: 11) or His-SpyTag-FLAG tag (SEQ ID NO: 12) at the C-terminus was cloned into an expression vector for periplasmic expression and transformed into Escherichia coli TG1F-. The MBP used in this example had four amino acids removed from the C-terminus. A partial sequence of the MBP-containing construct is shown in Figure 4. Expression and purification of the construct were performed as described in Example 1.

[0105] The first attempt to purify the periplasm-expressed MBP-SpyTag fusion protein was unsuccessful. Similar to the Fab fragment of Example 1, the periplasmic construct containing the FLAG-SpyTag-His peptide sequence at the C-terminus of MBP could not be purified. A construct containing the His-SpyTag-FLAG peptide sequence at the C-terminus of MBP could be purified, but it was not reactive in the subsequent SpyTag-SpyCatcher protein ligation reaction. As described in Example 1, Western blot analysis of the expression products before purification gave similar results in that the first tag of each construct was detected, but the last tag was not detected.

[0106] MALDI-TOF mass spectrometry of the MBPΔ4aa-His-SpyTag-FLAG construct (SEQ ID NO: 12) was performed as described in Example 1. The results of the mass spectrometry are shown below.

[0107] Full-length protein: Calculated mass (full-length) 44241 Da Detected mass (m / z) 44273 Da (full-length) Detected mass (m / z) 42032 Da (main product) Calculated mass (1-382aa) 42011 Da

[0108] The results of mass spectrometry showed a small amount of full-length protein. However, the main product consisted of amino acids 1 to 382, and it was shown that cleavage occurred after the valine at amino acid position 4 of SpyTag. This is the same position where cleavage occurred for the Fab fragment in Example 1. Without wishing to be bound by theory, the applicants believe that since cleavage occurs even with a structurally completely independent protein (i.e., MBP), SpyTag cleavage does not depend on the Fab amino acid sequence or structure. Mass spectrometry also showed that the ompT signal peptide for transport to the periplasm was cleaved, which occurs after the protein is transferred to the periplasm. Since the expression of the full-length SpyTag fusion protein in the cytoplasm is described in Non-Patent Document 13, the applicants hypothesized that the SpyTag cleavage described in Examples 1 and 2 occurred in the periplasm.

[0109] To test this hypothesis, MBP with FLAG-SpyTag and His-tag at the C-terminus was cloned into an expression vector for cytoplasmic expression without the ompT signal peptide and transformed into E. coli. Expression and purification were performed as described in Example 1. Cytoplasmic expression of the construct resulted in a high yield (about 11 mg / L) of full-length product.

[0110] Periplasmic expression of the Fab and MBP-SpyTag fusion proteins resulted in a truncated protein with a cleaved SpyTag, and cytoplasmic expression of MBP with the same amino acid sequence (without the signal peptide) resulted in a full-length product. Without wishing to be bound by theory, the applicants hypothesized that SpyTag cleavage was caused by one or more periplasmic proteases.

[0111] Periplasmic Expression of the Example 3 - scFv-SpyTag Fusion Protein The gene encoding the scFv having FLAG-SpyTag-His (SEQ ID NO: 34, Figure 9) at the C-terminus was cloned into an expression vector for periplasmic expression and transformed into Escherichia coli TG1F-. The partial sequence of the scFv construct is shown in Figure 9. The expression and purification of the construct were carried out as described in Example 1.

[0112] The first attempt to purify the periplasmically expressed scFv-SpyTag fusion protein was unsuccessful. Similar to the Fab fragment of Example 1 and the MBP construct of Example 2, the periplasmic scFv construct containing the FLAG-SpyTag-His peptide sequence at the C-terminus of scFv could not be purified via the His-tag.

[0113] Example 4 - Periplasmic Expression of FabX-SpyTag Fusion Proteins Where X Is a Linker in Various E. coli Strains The Fab-SpyTag-His (SEQ ID NO: 3) and Fab-FLAG-Spy-His (SEQ ID NO: 4) constructs were each transformed into the following E. coli strains for periplasmic expression to determine which periplasmic protease cleaved the SpyTag.

[0114] 1. TG1F (F- episome-free, Bio-Rad) 2. Jw0157: degP- (Yale Coli Genetic Stock Center) 3. KS476: degP- (Yale Coli Genetic Stock Center) 4. KS1000: prc- (or tsp-) (New England Biolabs) 5. JW3203: degQ- (Yale Coli Genetic Stock Center) 6. 27C2: degP-, ptr3-, ompT (ATCC) 7. HM130: degP-, ptr-, ompT-, tsp-, eda (Austin, Texas, USA)

[0115] The transformants were cultured, expressed, and purified as described in Example 1. The concentrations of the purified fusion proteins with each E. coli strain were determined, and the expression products of each fusion were analyzed by SDS-PAGE under non-reducing conditions. The expression of full-length Fab (including the tag) was visible as heavy and light chains on SDS-PAGE, but SpyTag cleavage resulted in Fab without the purification tag and was not purified. Both types of fusion proteins were expressed as full-length proteins only in the KS1000 (tsp-) strain and the HM130 (degP-, ptr-, ompT-, tsp-, eda) strain, indicating that the tsp protease is involved in the cleavage of SpyTag.

[0116] Generation of knockout cell lines using the TG1F strain for the Example 5 - Fab - SpyTag construct Since all the expression strains used in Example 4, except for TG1F, did not grow well and did not yield a high rate of properly folded soluble Fab, a TG1F protease knockout strain was constructed to increase the yield of Fab fused to SpyTag.

[0117] As described in Non-Patent Document 8, a mutant E. coli TG1F cell line with the tsp gene, degP gene, or both the tsp gene and the depP gene knocked out was prepared. Briefly, the gene was knocked out by transforming a PCR product containing an FRT site and a homologous flanking region containing an antibiotic resistance gene together with a plasmid containing λ recombinase. Clones were selected by antibiotic resistance in which the gene was replaced by recombination with the PCR product. In the next step, these clones were transfected with Flp recombinase, leading to the excision of the resistance gene.

[0118] The Fab-SpyTag-His and Fab-FLAG-Spy-His constructs (i.e., the same constructs as in Example 4) were transformed into the strains constructed above, namely, TG1F-Δtsp (SK4, DSMZ deposit number DSM33004), TG1F-ΔdegP, and the TG1F-ΔtspΔdegP double knockout. The expressed fusion proteins were purified and analyzed by SDS-PAGE as described in Example 1.

[0119] Results: In the TG1F-Δtsp strain (SK4) and the TG1ΔdegP double knockout strain, both types of fusion proteins were expressed as full-length proteins, but not in the TG1F-ΔdegP strain, indicating that degP is not involved in SpyTag cleavage. Furthermore, the expression of TG1F-Δtsp and TG1F-ΔtspΔdegP gave equivalent yields, and it was concluded that only tsp cleaves the SpyTag.

[0120] Next, the Fab-SpyTag-His and Fab-FLAG-SpyTag-His constructs were tested for their ability to form a covalent bond with SpyCatcher by protein ligation. SpyCatcher was produced by bacterial cytoplasmic expression and purified via Ni NTA as described in Non-Patent Document 32. The fusion proteins were expressed and purified using the SK4 strain as described above. Each fusion protein (final concentration 15 μM) was mixed with SpyCatcher (final concentration 20 μM) in 1×PBS buffer and allowed to couple at room temperature for 15 minutes, 30 minutes, 1 hour, 2 hours, 3 hours, and overnight. SDS-PAGE was used to test for the appearance of new bands on the gel corresponding to the SpyTag fusion-SpyCatcher binding product (see Figure 5). New SpyTag fusion-SpyCatcher bands of the correct size were observed for both fusion proteins, indicating that both Fab-SpyTag-H and Fab-FLAG-SpyTag-H are bound to SpyCatcher and that the SpyTag portion of the fusion proteins is intact and fully functional.

[0121] Example 6 - Periplasmic expression of MBP-SpyTag fusion protein in protease knockout strains A gene encoding maltose-binding protein (MBP) with a FLAG-SpyTag-His-tag (SEQ ID NO: 11, Figure 4) at the C-terminus was cloned into an expression vector for periplasmic expression and transformed into Escherichia coli strain SK4. The MBP sequence, which is common in fusion proteins for protein crystallization studies, is truncated by four amino acids at the C-terminus (Non-Patent Document 29). This sequence was used in these experiments. Expression and purification of the construct were performed as described in Example 1. By using the tsp protease knockout strain, a full-length protein with a high yield (about 10 mg / L) was produced.

[0122] Example 7 - Periplasmic expression of scFv-SpyTag fusion protein in protease-deficient cell strains A gene encoding an scFv with a FLAG-SpyTag-His (SEQ ID NO: 34; Figure 9) at the C-terminus was cloned into an expression vector for periplasmic expression and transformed into the TG1F-Δtsp knockout strain. Expression and purification of the construct were performed as described in Example 1. Using the tsp protease knockout strain, a full-length protein with a high yield (about 7 mg / L) was produced.

[0123] Example 8 - Generation of knockout cell lines using the TG1F strain for Fab-SpyTag002 constructs A gene encoding a human antibody fragment in Fab format having SpyTag002 at the C-terminus of the shortened heavy chain (i.e., Fab-SpyTag002 construct, Figure 6) was cloned into an expression vector having signal sequences in the H and L chains of the Fab, and the nascent chains were directed to the periplasm by bacterial transport. Then, Escherichia coli TG1F (without F-episome, Bio-Rad) was transformed with such a vector. Expression and purification of the construct were carried out as described in Example 1. All attempts to purify the Fab-SpyTag002 fusion protein having a C-terminal His-tag (SEQ ID NOs: 13 and 14) were unsuccessful. A construct having a His-tag between the Fab and SpyTag (SEQ ID NO: 15) could be purified, but it did not possess a functional SpyTag002.

[0124] Expression of the Fab-SpyTag002-His (SEQ ID NO: 13 shown in Figure 6) construct (having Fabs derived from various antibodies) was tested in the following knockout strains of cells in which expression was successful with a construct having SpyTag instead of SpyTag002.

[0125] 1. KS1000 (Δtsp) 2. SK4 (TG1F-Δtsp) 3. TG1F-ΔtspΔdeg 4. PHM130 strain (Δtsp, ΔdegP, ΔompT, Δptr)

[0126] After purification by His-tag, the product was analyzed by SDS-PAGE as in Example 4. Expression of the full-length fusion protein failed in KS1000 (Δtsp), SK4 (TG1F-Δtsp), and TG1F-ΔtspΔdegP, and was successful in the HM130 strain (Δtsp, ΔdegP, ΔompT, Δptr), and the yield of the purified antibody was acceptable or "high" (i.e., about 10 mg / L).

[0127] Next, as described above, Fab-His-SpyTag002 (SEQ ID NO: 15, Figure 6) expressed and purified in a non-protease-deficient TG1F- strain (i.e., the strain has both tsp and ompT proteases) and a TG1F-ΔtspΔdegP strain (i.e., the strain has ompT protease) was analyzed by MALDI-TOF mass spectrometry described in Example 1 to determine where the fusion protein was cleaved. The mass spectrometry results are shown below.

[0128] Fab-His-SpyTag002 expression in TG1F Light chain Calculated mass (full length) 22691 Da Detected mass (m / z) 22682 Da Heavy chain Calculated mass (full length) 26643 Da Detected mass (m / z) 25478 Da Calculated mass (-9aa) 25488 Da

[0129] Fab-His-SpyTag002 expression in TG1F-ΔtspΔdegP Light chain Calculated mass (full length) 22691 Da Detected mass (m / z) 22680 Da Heavy chain Calculated mass (full length) 26643 Da Detected mass (m / z) 26626 Da (full length) Detected mass (m / z) 26182 Da (-3aa) Calculated mass (-3aa) 26195 Da

[0130] Results: Based on the mass spectrometry results with non-protease-deficient bacterial strains, the 9-amino acid portion was cleaved from the C-terminus of the fusion protein. This was the same cleavage site as that observed for the SpyTag shown to be cleaved by the tsp protease. Based on the mass spectrometry results with bacterial strains deficient in tsp and degP proteases (and having the ompT protease), the 3-amino acid portion was cleaved from the C-terminus. Based on all the mass spectrometry results, SpyTag002 (length 14 amino acids, VPTIVMVDAYKRYK, SEQ ID NO: 2) was cleaved by the tsp protease after valine at amino acid position 5 and by a second protease after lysine at amino acid position 11. Thus, two different proteases are involved in the cleavage of SpyTag002, one of which is tsp, and the second protease can be hypothesized to be ompT or ptr based on the expression results in four strains.

[0131] The TG1F-knockout strains, TG1F-ΔompT strain and TG1F-ΔtspΔompT strain (SK13, DSMZ deposit number DSM33005), were prepared by the same method as described in Example 4. Subsequently, the expression of the Fab-SpyTag002-His construct (Figure 6, SEQ ID NO: 13) and the Fab-FLAG-SpyTag002-His construct (Figure 6, SEQ ID NO: 14) was tested in the TG1F-ΔompT strain and the TG1F-Δtsp`ompT strain as described above. The expression in the TG1F-ΔompT strain did not result in a significant amount of full-length protein because the tsp cleavage site from the SpyTag still exists in SpyTag002. The expression of the construct in the TG1F-ΔtspΔompT strain (SK13) was successful and resulted in a full-length protein with a high protein yield (i.e., about 10 mg / L) after purification. Thus, both the tsp and ompT proteases are involved in the cleavage of SpyTag002.

[0132] Example 9 - Testing of various strategies for protecting SpyTag during periplasmic expression of Fab-SpyTag fusion proteins in non-protease-deficient bacterial strains Experiments were conducted to investigate whether SpyTag can be protected from cleavage during the expression of SpyTag fusion proteins in non-protease-deficient E. coli TG1F-cell line.

[0133] The following strategies were tested to prevent SpyTag cleavage. 1. ST2 - A linker (CXC) with two cysteine residues was introduced between the FLAG®-tag and SpyTag to generate a disulfide-bridged loop (Non-Patent Document 30). Without being bound by theory, the applicants theorized that such a non-linear linker would prevent protease binding: Fab heavy chain - Flag - CXC - SpyTag - His6. 2. ST3 - A poly-proline linker was introduced into the fusion protein. The poly-proline linker forms a poly-proline helix (Non-Patent Document 21). Without being bound by theory, the applicants theorized that the polyproline helix would prevent protease binding. Two versions of the fusion protein with a poly-proline linker were generated. a. Strong: PPPPPT b. Weak: PLPPPF 3. ST4 - Two amino acids protrude from the globular folding domain of the Fab heavy chain (PDB structure accession number 2JB5). These two amino acids were removed, and SpyTag was added to the shortened Fab heavy chain (ending with conserved valine (IMGT position number 121) and two hinge amino acids, glutamic acid and proline) without an unfolded linker sequence (see "b" below). Without being bound by theory, the applicants theorized that bringing SpyTag closer to the folded domain would prevent protease binding. a. Fab heavy chain:...VEPKS - COOH b. ST4:...VEP - SpyTag - His 4. ST5 - This strategy is similar to ST2, but includes FLAG and SpyTag in the disulfide-bridged loop: Fab heavy chain - C - Flag - Spy - C - His6.

[0134] The above-described construct was expressed in a non-protease-deficient E. coli TG1F- cell strain and purified via a His-tag, and analyzed by SDS-PAGE as described above. Only ST4 (SEQ ID NO: 16, Figure 7) was expressed without SpyTag cleavage. The ligation of ST4 to SpyCatcher was tested as described in Example 5. According to SDS-PAGE analysis, no new band was shown for the SpyTag fusion-SpyCatcher binding product within about 2 hours. Therefore, the fusion protein with two amino acids removed from the heavy chain of the Fab was expressed without SpyTag cleavage but did not ligate to SpyCatcher. Without wishing to be bound by theory, the applicant believes that the proximity of SpyTag to the folded antibody structure prevented the protease from binding to and cleaving SpyTag, but also sterically prevented SpyCatcher, which is required for the SpyTag-SpyCatcher reaction, from binding to SpyTag.

[0135] Constructs based on the ST4 design summarized in Figure 7 (SEQ ID NOs: 17-32) were further prepared and tested as described above to determine whether SpyTag was cleaved. As described in Example 5, the constructs were also tested for their ability to ligate SpyTag to SpyCatcher. The full-length expression yields and ligation results of all ST4 constructs are summarized in Table 1. "High" indicates that the fusion protein expression yield was between about 5-10 mg / L, "low" indicates that the expression yield was about 2-4 mg / L, and "very low" indicates that the expression yield was less than about 2 mg / L.

[0136]

Table 1

[0137] The results in Table 1 show that the fusion protein exhibited high yields of full-length protein (i.e., about 5 - 10 mg / L) and that the SpyTag directly added to the C-terminus of the Fab heavy chain (or ST4+2) gave the best overall performance in terms of being ligated to SpyCatcher. Fusion proteins having 1, 2, or 3 amino acid linkers between the C-terminus of the Fab heavy chain and the SpyTag (i.e., ST4+3 and ST4+4 / 5) were largely cleaved by proteases, resulting in low or very low yields of full-length protein. Without being bound by theory, the applicants believe that the longer the linker between the C-terminus of the folded Fab globular domain and the SpyTag, the more prone the SpyTag is to cleavage and ligation to SpyCatcher. The applicants also believe that a fusion protein having a SpyTag directly added to the C-terminus of the Fab heavy chain represents an optimal compromise between adding the SpyTag in close proximity to the folded domain of the Fab to avoid periplasmic protease cleavage and allowing sufficient space to enable the ligation of the SpyTag to SpyCatcher sterically.

[0138] Example 10 - Testing of strategies for protecting SpyTag during periplasmic expression of maltose-binding protein-SpyTag fusion protein in a non-protease-deficient bacterial strain In maltose-binding protein (MBP), experiments were conducted to test the hypothesis that moving the SpyTag near the folding domain protects the SpyTag from periplasmic protease digestion. MBP was selected because it is a different class of protein from Fab. The MBP sequence most commonly used in crystallization studies has four amino acids truncated at the C-terminus (Non-Patent Document 29). Since increased rigidity of the protein (i.e., folded domain without a flexible linker) is useful for crystallization experiments and structure determination, and the same was expected for the stabilization of the SpyTag from protease cleavage, this sequence was used in this example. A gene encoding MBP with a SpyTag-His-tag (SEQ ID NO: 33, Figure 8) directly fused to the C-terminus was cloned into an expression vector for periplasmic expression and transformed into E. coli TG1F-strain. Expression and purification of the construct were performed as described in Example 1. High yields (about 7 mg / L) of the full-length protein were produced in the non-protease-deficient TG1F strain. This construct was tested as described in Example 5 for the ability of the SpyTag to ligate to SpyCatcher, and this fact was confirmed.

[0139] When SpyTag was directly added to the folded domain of MBP, a fusion protein was produced in which the SpyTag was protected from periplasmic protease cleavage, the SpyTag remained functional, and was similar to the Fab fusion protein of Example 9 in which the SpyTag was directly added to the heavy chain. In contrast, the MBP-Flag-SpyTag-His and MBP-His-SpyTag-Flag fusion proteins in Example 2 contain tags that act as linkers having 13 and 12 amino acids, respectively, that render the SpyTag vulnerable to protease cleavage. Without wishing to be bound by theory, Applicants believe these results indicate that cleavage of the SpyTag fused to the C-terminus of the Fab without a linker (i.e., the Fab-SpyTag fusion protein tested in Example 9) was independent of the Fab structure since cleavage also occurred when the SpyTag was fused to the C-terminus of the MBP via a linker and was structurally unrelated to the Fab.

[0140] Example 11 - Testing of Strategies for Protecting SpyTag During Periplasmic Expression of scFv-SpyTag Fusion Protein in a Non-Protease-Deficient Bacterial Strain In the scFv, experiments were conducted to test the hypothesis that bringing the SpyTag closer to the folded domain protects the SpyTag from periplasmic protease digestion. The scFv was chosen because it is a different protein from Fab or MBP and has a different structure. The scFv used in this example resulted in a C-terminally truncated scFv with 6 amino acids removed from the C-terminus and shortened within its light chain FR4 region (encoded by the J gene). A scFv(Δ6aa)-SpyTag-His fusion construct (SEQ ID NO: 35, Figure 9) in which SpyTag-His was directly fused (i.e., without a linker sequence) to the C-terminus of the C-terminally truncated scFv was cloned into an expression vector for periplasmic expression and transformed into E. coli TG1F-. Expression and purification of the fusion protein were performed as described in Example 1. The full-length protein was produced in high yield (about 8 mg / L), indicating that directly adding the SpyTag to the scFv protected the SpyTag from periplasmic cleavage.

[0141] Example 12 - Periplasmic Expression of Fab-SpyTag003 Fusion Protein in Protease-Deficient Bacterial Strains A gene encoding a human antibody fragment in Fab format having SpyTag003 at the C-terminus of the shortened heavy chain, i.e., a construct of SEQ ID NOs: 3 and 4 having SpyTag002 replaced with SpyTag003 (SEQ ID NO: 36), was cloned into an expression vector for periplasmic expression as described in Example 1. Transformation of the plasmid into E. coli TG1F- or SK4 strains and expression and purification of the construct were performed as described in Example 1. All attempts to purify the Fab-SpyTag003 fusion protein were unsuccessful.

[0142] As described in Example 1, the plasmid was transformed into, expressed in, and purified from the SK13 strain. Expression of the construct in the TG1F-ΔtspΔompT strain (SK13) was successful and resulted in a full-length protein with a good protein yield (i.e., about 6 mg / L) after purification. Thus, both the tsp and ompT proteases are involved in the cleavage of SpyTag003.

Deposit Number

[0143] DSM Deposit Number 33004 DSM Deposit Number 33005

Sequence Listing Free-Text

[0144] All patents, patent applications, and other published reference materials cited herein are hereby incorporated by reference in their entirety. SEQ ID NO: 1 (SpyTag) AHIVMVDAYK PTK SEQ ID NO: 2 (SpyTag002) VPTIVMVDAY KRYK Sequence number 3 (Fab-Spy-His, partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1-EPKSEFGAHIVMVDAYKPTKGAPHHHHHH Sequence number 4 (Fab-FLAG-Spy-His, partial amino acid sequence starting with 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at the defined IMGT position number 121 of human Ig CH1 Human CH1-EPKSEFDYKDDDDKGGSAHIVMVDAYKPTKGAPHHHHHH Sequence number 5 (Fab-X-Spy-His, partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1) Human CH1-EPKSEFGGGSGGGSAHIVMVDAYKPTKGAPHHHHHH Sequence number 6 (Fab-His-Spy, partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1) Human CH1-EPKSEFHHHHHHGAPGAHIVMVDAYKPTK Sequence number 7 (Fab-His-Spy-FLAG, partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1) Human CH1-EPKSEFHHHHHHGAPGAHIVMVDAYKPTKGGSDYKDDDDK Sequence number 8 (Fab-Spy-Sx2, partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1) Human CH1 - EPKSEFGAHIVMVDAYKPTKGAPSAWSHPQFEKGGGSGGGSGGSAWSHPQFEK SEQ ID NO:9 (Fab - FLAG - Spy - Sx2, a partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with the valine conserved at IMGT position number 121 of human Ig CH1 as defined by IMGT) Human CH1 - EPKSEFDYKDDDDKGGSAHIVMVDAYKPTKGAPSAWSHPQFEKGGGSGGGSGGSAWSHPQFEK SEQ ID NO:10 (Fab - X - Spy - Sx2, a partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with the valine conserved at IMGT position number 121 of human Ig CH1 as defined by IMGT) Human CH1 - EPKSEFGGGSGGGSAHIVMVDAYKPTKGAPSAWSHPQFEKGGGSGGGSGGSAWSHPQFEK SEQ ID NO:11 (MBP(Δ4aa) - FLAG - Spy - His) MKKTAIAIAVALAGFATVAQAKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINAASGRQTVDEALKDAQTEFDYKDDDDKGGSAHIVMVDAYKPTKGAPHHHHHH SEQ ID NO:12 (MBP(Δ4aa) - His - Spy - FLAG) MKKTAIAIAVALAGFATVAQAKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINAASGRQTVDEALKDAQTEFHHHHHHGAPGAHIVMVDAYKPTKGGSDYKDDDDK Sequence number 13 (Fab-Spy2-His, partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1-EPKSEFGVPTIVMVDAYKRYKGAPHHHHHH Sequence number 14 (Fab-FLAG-Spy2-His, partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1-EPKSEFDYKDDDDKGGSVPTIVMVDAYKRYKGAPHHHHHH Sequence number 15 (Fab-His-Spy2, partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1-EPKSEFHHHHHHGAPGVPTIVMVDAYKRYK SEQ ID NO: 16 (Fab-Spy-His_ST4(HCΔ2aa), partial amino acid sequence starting with the first two amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1 - EPAHIVMVDAYKPTKGAPHHHHHH SEQ ID NO: 17 (Fab-Spy-His_ST4+1(HCΔ1aa) Partial amino acid sequence starting with the first three amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1 - EPKAHIVMVDAYKPTKGAPHHHHHH SEQ ID NO: 18 (Fab-Spy-His_ST4+1(HCΔ1aa), partial amino acid sequence starting with the first two amino acid residues of the human IgG1 hinge domain, with the third amino acid residue substituted with "E", and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1 - EPEAHIVMVDAYKPTKGAPHHHHHH SEQ ID NO: 19 (Fab-Spy-His_ST4+1(HCΔ1aa), partial amino acid sequence starting with the first two amino acid residues of the human IgG1 hinge domain, with the third amino acid residue substituted with "G", and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1 - EPGAHIVMVDAYKPTKGAPHHHHHH SEQ ID NO: 20 (Fab-Spy-His_ST4+1(HCΔ1aa), partial amino acid sequence starting with the first two amino acid residues of the human IgG1 hinge domain, with the third amino acid residue substituted with "R", and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1 - EPRAHIVMVDAYKPTKGAPHHHHHH Sequence number 21 (Fab-Spy-His_ST4+2(HC), partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1-EPKSAHIVMVDAYKPTKGAPHHHHHH Sequence number 22 (Fab-Spy-His_ST4+2(HC), partial amino acid sequence starting with the first 3 amino acid residues of the human IgG1 hinge domain, with the 4th amino acid residue replaced by "E", and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1-EPKEAHIVMVDAYKPTKGAPHHHHHH Sequence number 23 (Fab-Spy-His_ST4+2(HC), partial amino acid sequence starting with the first 3 amino acid residues of the human IgG1 hinge domain, with the 4th amino acid residue replaced by "K", and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1-EPKKAHIVMVDAYKPTKGAPHHHHHH Sequence number 24 (Fab-Spy-His_ST4+2(HC), partial amino acid sequence starting with the first 3 amino acid residues of the human IgG1 hinge domain, with the 4th amino acid residue replaced by "S", and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1-EPKGAHIVMVDAYKPTKGAPHHHHHH Sequence number 25 (Fab-Spy-His_ST4+3(HC+1), partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1-EPKSDAHIVMVDAYKPTKGAPHHHHHH Sequence number 26 (Fab-Spy-His_ST4+3(HC+1), partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1-EPKSKAHIVMVDAYKPTKGAPHHHHHH Sequence number 27 (Fab-Spy-His_ST4+3(HC+1), partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1-EPKSPAHIVMVDAYKPTKGAPHHHHHH Sequence number 28 (Fab-Spy-His_ST4+3(HC+1), partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1-EPKSSAHIVMVDAYKPTKGAPHHHHHH Sequence number 29 (Fab-Spy-His_ST4+4(HC+2), partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1-EPKSGGAHIVMVDAYKPTKGAPHHHHHH Sequence number 30 (Fab-Spy-His_ST4+4(HC+2), partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1-EPKSGFAHIVMVDAYKPTKGAPHHHHHH Sequence number 31 (Fab-Spy-His_ST4+5(HC+3, partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1 according to the IMGT definition) Human CH1-EPKSEGGAHIVMVDAYKPTKGAPHHHHHH SEQ ID NO: 32 (Fab-Spy-His_ST4+5 (HC+3), partial amino acid sequence starting with the first 4 amino acid residues of the human IgG1 hinge domain and ending with valine conserved at IMGT position number 121 of human Ig CH1) Human CH1-EPKSGGSAHIVMVDAYKPTKGAPHHHHHH SEQ ID NO: 33 (MBP(Δ4aa)-Spy-His_ST4): MKKTAIAIAVALAGFATVAQAKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINAASGRQTVDEALKDAQTAHIVMVDAYKPTKGAPHHHHHH SEQ ID NO: 34 (scFv-F-Spy-H): QVQLVESGGNLVQPGGSLRLSCAASGFTFGSFSMSWVRQAPGGGLEWVAGLSARSSLTHYADSVKGRFTISRDNAKNSVYLQMNSLRVEDTAVYYCARRSYDSSGYWGHFYSYMDVWGQGTLVTVSSGGGGSGGGGSGGGGSQSVLTQPSSVSAAPGQKVTISCSGSTSNIGNNYVSWYQQHPGKAPKLMIYDVSKRPSGVPDRFSGSKSGNSASLDISGLQSEDEADYYCAAWDDSLSEFLFGTGTKLTVLGQEFDYKDDDDKGGSAHIVMVDAYKPTKGAPHHHHHH Sequence number 35 (scFv(Δ6aa)-Spy-H): QVQLVESGGNLVQPGGSLRLSCAASGFTFGSFSMSWVRQAPGGGLEWVAGLSARSSLTHYADSVKGRFTISRDNAKNSVYLQMNSLRVEDTAVYYCARRSYDSSGYWGHFYSYMDVWGQGTLVTVSSGGGGSGGGGSGGGGSQSVLTQPSSVSAAPGQKVTISCSGSTSNIGNNYVSWYQQHPGKAPKLMIYDVSKRPSGVPDRFSGSKSGNSASLDISGLQSEDEADYYCAAWDDSLSEFLFGTGTKAHIVMVDAYKPTKGAPHHHHHH Sequence number 36 (SpyTag003) RGVPHIVMVDAYKRYK

Claims

1. 1. A method for producing a periplasmic fusion protein, comprising: Cultivating in a culture medium a mutant E. coli host cell transformed with a vector containing a nucleic acid encoding the periplasmic fusion protein, thereby expressing the periplasmic fusion protein; recovering said periplasmic fusion protein from said mutant E. coli host cells; the periplasmic fusion protein comprises a binding motif added to an antigen-binding fragment, including a Fab, scFv, or scFab; The binding motif comprises SEQ ID NO:1 (SpyTag) or a sequence having at least 80% sequence identity to SEQ ID NO:1 (SpyTag), each of which is capable of forming a covalent bond via protein ligation to SpyCatcher, SpyCatcher002, or SpyCatcher003; The mutant E. coli host cell comprises: a) a mutation in the Tsp gene that encodes a mutant Tsp (tail specific protease) protein, which mutation reduces or eliminates protease activity, or b) a mutation in the Tsp gene or in a regulatory sequence of the Tsp gene that reduces or eliminates expression of the Tsp protein; or c) one or more deletions in a region of the bacterial chromosome that reduces or eliminates the activity of the Tsp protein; or d) an inhibitor or inactivator that reduces or eliminates the Tsp protease activity, or an inhibitor of Tsp protease expression. Including, A method in which Tsp protein activity is reduced or eliminated compared to wild-type E. coli cells.

2. 1. A method for producing a periplasmic fusion protein, comprising: Cultivating in a culture medium a mutant E. coli host cell transformed with a vector containing a nucleic acid encoding the periplasmic fusion protein, thereby expressing the periplasmic fusion protein; recovering said periplasmic fusion protein from said mutant E. coli host cells; the periplasmic fusion protein comprises a binding motif added to an antigen-binding fragment, including a Fab, scFv, or scFab; the binding motif comprises SEQ ID NO:2 (SpyTag002) or a sequence having at least 80% sequence identity to SEQ ID NO:2 (SpyTag002), or a sequence having at least 80% sequence identity to SEQ ID NO:36 (SpyTag003) or a sequence having at least 80% sequence identity to SEQ ID NO:36 (SpyTag003), each of which is capable of forming a covalent bond via protein ligation to SpyCatcher, SpyCatcher002, or SpyCatcher003; The mutant E. coli host cell comprises: a) a mutation in the Tsp gene encoding a mutant Tsp protein, which reduces or eliminates protease activity, or a mutation in the Tsp gene or a regulatory sequence of the Tsp gene, which reduces or eliminates expression of the Tsp protein, or a deletion in one or more regions of the bacterial chromosome, which reduces or eliminates activity of the Tsp protein; and b) a mutation in the ompT gene encoding a mutant ompT protein, which reduces or eliminates the protease activity, or a mutation in the ompT gene or in a regulatory sequence of the ompT gene, which reduces or eliminates the expression of the ompT protein, or a deletion in one or more regions of the bacterial chromosome, which reduces or eliminates the activity of the ompT protein. Including, A method in which Tsp protein activity and ompT protein activity are reduced or eliminated compared to wild-type E. coli cells.

3. 1. A method for producing a periplasmic fusion protein, comprising: Cultivating in a culture medium a mutant E. coli host cell transformed with a vector containing a nucleic acid encoding the periplasmic fusion protein, thereby expressing the periplasmic fusion protein; recovering said periplasmic fusion protein from said mutant E. coli host cells; the periplasmic fusion protein comprises a binding motif added to an antigen-binding fragment, including a Fab, scFv, or scFab; the binding motif comprises SEQ ID NO:2 (SpyTag002) or a sequence having at least 80% sequence identity to SEQ ID NO:2 (SpyTag002), or a sequence having at least 80% sequence identity to SEQ ID NO:36 (SpyTag003) or a sequence having at least 80% sequence identity to SEQ ID NO:36 (SpyTag003), each of which is capable of forming a covalent bond via protein ligation to SpyCatcher, SpyCatcher002, or SpyCatcher003; The mutant E. coli host cell comprises: a) an inhibitor or inactivator of Tsp protease or an inhibitor of Tsp expression, and b) Inhibitors or inactivators of ompT protease or inhibitors of ompT expression wherein Tsp protein activity and ompT protein activity are reduced or eliminated compared to wild-type E. coli cells.

4. The method of claim 1 , wherein the binding motif is attached to the C-terminus of the antigen-binding fragment directly or via a linker sequence.

5. The method of claim 1 , wherein the antigen-binding fragment is a Fab.

6. 2. The method of claim 1, wherein the mutant E. coli host cell is a mutant E. coli TG1F- strain having DSM accession number 33004, deposited on January 8, 2019.

7. 3. The method of claim 2, wherein the mutant E. coli host cell is the mutant E. coli TG1F- strain having DSM accession number 33005, deposited on January 8, 2019.

8. 1. A mutant E. coli TG1, TG1F-, XL1 Blue, MC1061, SS320, BL21, JM83, JM109, HB2151, W3110, or Cmax5alpha strain comprising a nucleic acid encoding a periplasmic fusion protein comprising an antigen-binding fragment, including a Fab, scFv, or scFab, and a binding motif comprising SEQ ID NO:1 (SpyTag) or a sequence having at least 80% sequence identity to SEQ ID NO:1 (SpyTag), each of which is capable of forming a covalent bond via protein ligation to SpyCatcher, SpyCatcher002, or SpyCatcher003; The mutant E. coli strain comprises a mutation in the Tsp gene encoding a mutant Tsp protein, which reduces or eliminates protease activity, or a mutation in the Tsp gene or a regulatory sequence of the Tsp gene, which reduces or eliminates expression of the Tsp protein, or one or more deletions of regions of the bacterial chromosome that reduce or eliminate Tsp protein activity, and which has reduced or eliminated Tsp protein activity compared to a wild-type E. coli strain.

9. a) a nucleic acid encoding a periplasmic fusion protein comprising an antigen-binding fragment, including a Fab, scFv, or scFab, and a binding motif comprising SEQ ID NO:2 (SpyTag002) or a sequence having at least 80% sequence identity to SEQ ID NO:2 (SpyTag002), each of which is capable of forming a covalent bond via protein ligation to SpyCatcher, SpyCatcher002, or SpyCatcher003; or b) a nucleic acid encoding a periplasmic fusion protein comprising an antigen-binding fragment, including a Fab, scFv, or scFab, and a binding motif comprising SEQ ID NO: 36 (SpyTag003) or a sequence having at least 80% sequence identity to SEQ ID NO: 36 (SpyTag003), each of which is capable of forming a covalent bond via protein ligation to SpyCatcher, SpyCatcher002, or SpyCatcher003. and wherein the E. coli TG1, TG1F-, XL1 Blue, MC1061, SS320, BL21, JM83, JM109, HB2151, W3110, or Cmax5alpha strain comprises The mutant E. coli strain comprises: a) a mutation in the Tsp gene encoding a mutant Tsp protein, which reduces or eliminates protease activity, or a mutation in the Tsp gene or a regulatory sequence of the Tsp gene, which reduces or eliminates expression of the Tsp protein, or a deletion in one or more regions of the bacterial chromosome, which reduces or eliminates activity of the Tsp protein; and b) a mutation in the ompT gene encoding a mutant ompT protein, which reduces or eliminates the protease activity, or a mutation in the ompT gene or in a regulatory sequence of the ompT gene, which reduces or eliminates the expression of the ompT protein, or a deletion in one or more regions of the bacterial chromosome, which reduces or eliminates the activity of the ompT protein. Including, A mutant E. coli strain in which Tsp protein activity and ompT protein activity are reduced or eliminated compared to a wild-type E. coli strain.

10. 1. A mutant E. coli strain comprising: a) an antigen-binding fragment, including a Fab, scFv, or scFab; a binding motif comprising SEQ ID NO:1 (SpyTag) or a sequence having at least 80% sequence identity to SEQ ID NO:1 (SpyTag), each of which is capable of forming a covalent bond via protein ligation to SpyCatcher, SpyCatcher002, or SpyCatcher003, attached to the antigen-binding fragment; For expression of periplasmic fusion proteins comprising 1. A mutant E. coli strain comprising a mutation in a Tsp gene encoding a mutant Tsp protein, which mutation reduces or eliminates protease activity, or a mutation in the Tsp gene or a regulatory sequence of the Tsp gene, which reduces or eliminates expression of the Tsp protein, or a deletion of one or more regions of the bacterial chromosome which reduces or eliminates activity of the Tsp protein, A mutant E. coli strain that has reduced or eliminated Tsp protein activity compared to a wild-type E. coli strain, or b) an antigen-binding fragment, including a Fab, scFv, or scFab; a binding motif comprising SEQ ID NO:2 (SpyTag002) or a sequence having at least 80% sequence identity to SEQ ID NO:2 (SpyTag002), or SEQ ID NO:36 (SpyTag003) or a sequence having at least 80% sequence identity to SEQ ID NO:36 (SpyTag003), respectively, each of which is capable of forming a covalent bond via protein ligation to SpyCatcher, SpyCatcher002, or SpyCatcher003, attached to the antigen-binding fragment; For expression of periplasmic fusion proteins comprising i) a mutation in the Tsp gene encoding a mutant Tsp protein, which reduces or eliminates protease activity, or a mutation in the Tsp gene or a regulatory sequence of the Tsp gene, which reduces or eliminates expression of the Tsp protein, or a deletion in one or more regions of the bacterial chromosome, which reduces or eliminates activity of the Tsp protein; and ii) a mutation in the ompT gene encoding a mutant ompT protein, which reduces or eliminates protease activity, or a mutation in the ompT gene or in a regulatory sequence of the ompT gene, which reduces or eliminates expression of the ompT protein, or a deletion in one or more regions of the bacterial chromosome, which reduces or eliminates the activity of the ompT protein.

1. A mutant E. coli strain comprising: A mutant E. coli strain having reduced or eliminated Tsp and ompT protein activities compared to a wild-type E. coli strain.

11. 2. The method of claim 1, wherein the periplasmic fusion protein comprises a binding motif attached, directly or via a linker, to the C-terminus of a protein structural domain in the antigen-binding fragment, wherein the binding motif comprises SEQ ID NO:1 (SpyTag) or a sequence having at least 80% sequence identity to SEQ ID NO:1 (SpyTag), each of which is capable of forming a covalent bond via protein ligation to SpyCatcher, SpyCatcher002, or SpyCatcher003.

12. The method of claim 11, wherein the binding motif of the periplasmic fusion protein is added directly to the C-terminus of a protein structural domain in the antigen-binding fragment, and the binding motif is resistant to proteolysis.

13. The method of claim 12, wherein the protein structural domain of the periplasmic fusion protein is a human scFv single-chain antibody fragment C-terminally truncated within the FR4 region.

14. The method of claim 11, wherein the binding motif of the periplasmic fusion protein is attached to the C-terminus of the protein structural domain in the antigen-binding fragment via a one or two amino acid linker.

15. 12. The method of claim 11, wherein the binding motif of the periplasmic fusion protein is attached to the C-terminus of IMGT position 121 of the human heavy chain CH1 antibody domain via a 2, 3, or 4 amino acid linker.

16. 12. The method of claim 11, wherein the binding motif of the periplasmic fusion protein is linked to the C-terminus of the human constant light chain antibody domain at IMGT position 121 via a 2, 3 or 4 amino acid linker.

Citation Information

Patent Citations

  • Bacterial host lines containing a mutant spr gene and exhibiting reduced Tsp activity

    JP2013516964A

  • Staged assembly of nanostructures

    US20030198956A1

  • Peptide tag systems that spontaneously form an irreversible link to protein partners via isopeptide bonds

    US9547003B2

  • Compositions and methods for making antibody conjugates

    WO2016183387A1

  • Methods and products for fusion protein synthesis

    WO2016193746A1