Controllable self-cleaving domains

WO2026202065A1PCT designated stage Publication Date: 2026-10-01UNIVERSITE CATHOLIQUE DE LOUVAIN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/058405
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2026-03-24
Publication Date
2026-10-01

Smart Images

  • Figure IMGF000008_0001
    Figure IMGF000008_0001
  • Figure IMGF000009_0001
    Figure IMGF000009_0001
  • Figure IMGF000011_0001
    Figure IMGF000011_0001
Patent Text Reader

Abstract

The present invention relates to a polynucleotide encoding a polypeptide comprising or consisting of SEQ ID NO: 1 or SEQ ID NO: 300, and to the use thereof for producing a chimeric protein.
Need to check novelty before this filing date? Find Prior Art

Description

CONTROLLABLE SELF-CLEAVING DOMAINSFIELD OF INVENTION

[0001] The present invention relates to production of recombinant proteins . In particular, the present invention relates to controllable self-cleaving domains, capable of selfcleaving when placed in specific redox conditions, and to the use thereof in a method for producing chimeric polyproteins.BACKGROUND OF INVENTION

[0002] Hints (Hog / INTein) are auto-proteolytic domains inserted into host proteins and able to catalyze self-splicing and / or self-cleavage and / or ligation reactions. The superfamily includes inteins, Hog domains and bacterial intein-like (BIL) domains. Among Hints, inteins have been the most extensively studied. These self-splicing domains are generally inserted into conserved regions of essential intracellular proteins and most of them are considered as selfish parasitic elements. Hog domains are embedded in animal Hedgehog proteins and participate in a posttranslational maturation process by simultaneously catalyzing self-cleavage and ligation of the Hedge domain to a cholesterol molecule. BIL domains are found in non-conserved proteins. Although much less is known about their activities and biological functions, self-cleavage and sometimes selfsplicing have been reported for the few BILs that have been studied.

[0003] The polypeptide rearrangement reactions that are catalyzed by Hint domains proceed through similar mechanisms involving acyl transfer between alcohol (water, Ser / Thr side chains, cholesterol), thiol (Cys side chain) and amine (main chain amino groups) functions. For example, the well-studied self-splicing of inteins proceeds through successive acyl shifts starting with the attack of the thiol or alcohol at the N-terminal extremity of the intein (Cys or Ser) to the adjacent peptide bond which results in a (thio)ester bond linking the N-flanking polypeptide, called the N-extein, to the side chain;the N-extein is subsequently transferred, by trans(thio)-esterification, to a thiol or alcohol at the N-terminal extremity of the C-extein (Cys, Ser or Thr) forming a branched intermediate; in the third step, cyclization of the highly conserved C-terminal asparagine of the intein releases the intein from the N- and C-exteins that are attached through a (thio)-ester bond; in the final spontaneous acyl shift, the N-extein is transferred, by transpeptidation, to the released amino group of the C-extein, yielding the spliced protein. For other members of the superfamily, the acyl acceptor of the second step can vary: it will be a water molecule in the case of a self-cleavage at the N-terminal extremity of the Hint domain, or the alcohol function of a cholesterol molecule in the case of a hog domain. The asparagine cyclization (third step) can also lead to a simple self-cleavage at the C-terminal extremity.

[0004] The chemically reactive residue side chains are mainly located within the Hint domains with an exception for the first residue flanking the C-terminal end of inteins, a nucleophilic residue that participates in the acyl transfer reactions. Nevertheless, flanking residues may also assist Hint-catalyzed reactions in less direct ways, notably by enabling specific conformations of the polypeptide chain at the junction sites. For example, the last N-extein glycine adjacent to the nucleophilic cysteine at the N-splicing junction of Ssp DnaB intein has been shown as critical for efficient splicing.

[0005] In the present invention, the Inventors characterized the activity of a BIL domain embedded in a large filamentous hemagglutinin (Fha) from Pseudomonas syringae pv. tomato DC3000. They showed that the Psy Fha BIL is not capable of splicing but is undergoing efficient self-cleavage from its host protein. Interestingly, instead of being ligated through a peptide bond, the flanking polypeptides end up connected by a disulfide bridge. These results highlight the potential applications of such auto-cleavable domains in biotechnology, in particular for producing chimeric proteins and a plurality of protein entities connected by a disulfide bridge.SUMMARY

[0006] This invention thus relates a polynucleotide encoding a polypeptide comprising or consisting of SEQ ID NO: 1 or SEQ ID NO: 300, or a sequence with at least 70% identity with SEQ ID NO: 1, preferably wherein the polypeptide comprises or consists of a sequence SEQ ID NO: 2, or a sequence with at least 70% identity with SEQ ID NO: 2.

[0007] In some embodiments, the polypeptide comprises or consists of a sequence selected from the group consisting of SEQ ID NO: 3 to 7, 12, 16 to 23, 26 to 47 or 49 to 102, and sequences with at least 70% identity with SEQ ID NO: 3 to 7, 12, 16 to 23, 26 to 47 or 49 to 102.

[0008] In some embodiments, the polypeptide further comprises in Nter a sequence selected from the group consisting of SEQ ID NO: 203 to 209, preferably a sequence selected from the group consisting of SEQ ID NO: 210 to 216, and more preferably a sequence selected from the group consisting of SEQ ID NO: 217 to 222 or 239.

[0009] In some embodiments, the polypeptide further comprises in Cter a sequence selected from the group consisting of SEQ ID NO: 223 to 227, preferably a sequence selected from the group consisting of SEQ ID NO: 228 to 232.

[0010] In some embodiments, the polynucleotide comprises or consists of a sequence of SEQ ID NO: 233.

[0011] The present invention further relates to a polynucleotide for expressing a chimeric polyprotein, wherein said polynucleotide comprises or consists of, from 5’ to 3’:optionally a sequence encoding a signal peptide,a sequence encoding a first protein entity,a polynucleotide as described herein, anda sequence encoding a second protein entity,wherein the first and second protein entities may be the same or different.

[0012] The present invention further relates to an expression cassette comprising a polynucleotide as described herein.

[0013] The present invention further relates to an expression vector comprising a polynucleotide or an expression cassette as described herein.

[0014] The present invention further relates to a host cell comprising a polynucleotide, an expression cassette, or an expression vector as described herein.

[0015] The present invention further relates to a method for producing a chimeric polyprotein, wherein said polyprotein comprises a first protein entity linked to a second protein entity by a cleavable sequence comprising a disulfide bridge, wherein said method comprises expressing a polynucleotide for expressing a chimeric polyprotein as described herein.

[0016] In some embodiments, the cleavable sequence comprises or consists of two peptides linked by a disulfide bridge, wherein the first peptide comprises or consists of a sequence PC, and wherein the second peptide comprises or consists of a sequence X67GPC (SEQ ID NO: 238, wherein X67is T, I or V).

[0017] In some embodiments, the method comprises a step of culturing a host cell comprising a polynucleotide as described herein, preferably wherein the host cell is a bacterium.

[0018] The present invention further relates to a chimeric polyprotein obtained or susceptible to be obtained by the method as described herein, wherein said chimeric polyprotein comprises a first protein entity and a second protein entity, wherein said first and second protein entities are linked by a disulfide bridge under oxidative conditions and may be dissociated under reductive conditions.

[0019] In some embodiments, the first and second protein entities are a protein of interest and a purification tag.

[0020] In some embodiments, the first and second protein entities are a protein of interest and an anchoring domain.DEFINITIONS

[0021] In the present invention, the following terms have the following meanings:

[0022] “Vector” refers to a vehicle by which a nucleic acid sequence (e.g., a DNA or RNA molecule), for example a nucleic acid encoding a RNA or a polypeptide or a polyprotein, can be introduced into a host cell, so as to transform, transfect or transduce the host cell and promote expression (e.g., transcription and / or translation) of the introduced nucleic acid sequence.

[0023] “Expression vector” refers to a vector capable of directing expression of a polynucleotide of interest (such as, e.g., a polynucleotide as described herein) in an appropriate host cell, comprising a promoter operatively linked to the polynucleotide sequence of interest, itself operatively linked to a termination sequence.

[0024] “Identity” or “identical”, when used herein in a relationship between the sequences of two or more polynucleotides or of two or more polypeptides, refers to the degree of sequence relatedness between polynucleotides or polypeptides (respectively), as determined by the number of matches between strings of two or more nucleotides or of two or more amino acid residues, respectively. “Identity” measures the percent of identical matches between the smaller of two or more sequences with gap alignments (if any) addressed by a particular mathematical model or computer program (i.e., “algorithms”). Identity of related polynucleotide or polypeptide sequences can be readily calculated by known methods. Such methods include, but are not limited to, those described in Computational Molecular Biology, Lesk, A. M., ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, D. W., ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part 1, Griffin, A. M., and Griffin, H. G., eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M. Stockton Press, New York, 1991; and Carillo et al., SIAM J. Applied Math. 48, 1073 (1988). Preferred methods for determining identity are designed to give the largest match between the sequences tested. Methods of determining identity are described in publicly available computer programs. Preferred computer program methods for determining identity between two sequencesinclude the GCG program package, including GAP (Devereux et al., Nucleic Acids Res.1984 Jan 11;12(1 Pt l):387-95; Genetics Computer Group, University of Wisconsin, Madison, Wis.), BLASTP, BLASTN, and FASTA (Altschul et al., J. Mol. Biol. 215, 403-410 (1990)). The BLASTX program is publicly available from the National Center for Biotechnology Information (NCBI) and other sources (BLAST Manual, Altschul et al. NCB / NLM / NIH Bethesda, Md. 20894; Altschul et al., J. Mol. Biol. 215, 403-410 (1990)).

[0025] The term “at least 70% sequence identity'", as in “70% sequence identity with a reference sequence" may thus encompass 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%.

[0026] The term “polypeptide” refers to a linear polymer of amino acids (preferably at least 10 or 20 amino acids) linked together by peptide bonds.

[0027] The term “protein” refers to a functional entity formed of one or more polypeptides, and optionally of non-polypeptides cofactors.

[0028] The term “polyprotein” refers to a functional entity formed of two or more protein entities (wherein a protein entity may be a polypeptide, a protein or a protein fragment), such as, for example, two or more proteins, two or more protein fragments, a protein and a polypeptide, etc., wherein the two or more protein entities are either the same or different and are linked together by a sequence comprising a disulfide bridge. In the present invention, the mixture of protein entities (i.e., proteins, protein fragments and / or polypeptides) recovered after cleavage of the disulfide bridge(s) is referred to as a “plurality of protein entities”.DETAILED DESCRIPTION

[0029] This invention relates to polynucleotides, in particular isolated polynucleotides, encoding polypeptides comprising or consisting of a BIL domain.

[0030] This invention relates to a polynucleotide encoding a polypeptide comprising or consisting of SEQ ID NO: 300. The polypeptide consisting of amino acids 3 to 150 with reference to SEQ ID NO: 1 is a BIL domain.

[0031] This invention relates to a polynucleotide encoding a polypeptide comprising or consisting of SEQ ID NO: 1, or a sequence with at least 70% sequence identity with SEQ ID NO: 1. The polypeptide consisting of amino acids 3 to 150 with reference to SEQ ID NO: 1 is a BIL domain.

[0032] This invention relates, more particularly to polynucleotides susceptible to encode recombinant polypeptides for biotechnological purposes; including fusion proteins and polyproteins.

[0033] When placed in specific conditions, said BIL domain is capable of self-cleavage, and may therefore be referred to as an “auto-cleavable domain”.

[0034] The inventors demonstrated in the present invention that the expression of a genetic construct comprising a polynucleotide of the present invention between the sequences encoding two protein entities leads to the production of a polyprotein comprising said protein entities linked together by a disulfide bridge, and not comprising a polypeptide of SEQ ID NO: 1 and / or SEQ ID NO: 300.

[0035] As demonstrated in the example part, said production involves a reaction between the cysteine residue at position 2 and the cysteine residue at position 154 with reference to SEQ ID NO: 1 or SEQ ID NO: 300. Thus, according to the present invention, the polypeptide encoded by the polynucleotide of the present invention comprises a cysteine residue at position 2 and a cysteine residue at position 154 with reference to SEQ ID NO: 1 or SEQ ID NO: 300.

[0036] The polynucleotide of the invention encodes a polypeptide comprising or consisting of SEQ ID NO: 300.SEQ ID NO: 300PCCFAAGTX1VSTPX2GX3RX4IX5X6LX7X8GDX9VWSKPEX10GGX11PFAAX12X13X 14X15X16X17X18X19X20X21X22X23X24X25X26X27X28X29X30X31X32X33X34X35X36X37X38X39 X40EX41LX42VTPX43HPFYVX44AX45X46X47FX48PVIX49LKX50GDX51LQSLX52DGX5 3X54X55X56X57SSX58VESLELYX59PX60GX6iTX62NLX63VX64X65GHTFYVGX66LKT WVHNX67GPCwherein:Xi is M or K;- X2is D, G, N, S or R;X3 is E or D;X4 is A or T;X5 is D or E;Xe is T or S;- X7is K or N;Xs is V or I;X9 is I or V;- X10 is G, R or K;Xu is K or E;-X12 is A , V, T or K;X13 is I or absent;X14 is L, T or absent;X15 is A or absent;Xi6 is T or absent;X17 is H or absent;Xis is I, V, Q or absent;X19 is R or absent;X20 is T, S, N or absent;X21 is D or absent;X22 is Q or absent;X23 is P or absent;X24 is I or absent;X25 is Y or absent;X26 is R or absent;X27 is L or absent;X28 is K, T or absent;X29 is L or absent;X30 is K or absent;X31 is G, S or absent;X32 is K, V, R, I, S or absent;X33 is Q, R, D or absent;X34 is E, Q, V, A, S, L or absent;X35 is N, D, Y or absent;X36 is G or absent;X37 is any amino acid residue or is absent; preferably X37 is Q, E, K, N, T, I or is absent;X38 is A, V, T or absent;X39 is E, D, A, R, T or absent;- X40is D, E, N, S or G or R;- X41 is S, T or V;X42 is L or M;X43 is G or S;X44 is P or A;- X45is Q, H, R or K;X46 is H or R or K;X47 is G or D;X48 is V or I;X49 is D, N or G;X50 is P or A or L;X51 is R, Q or L;X52 is A, E, or D;X53 is A, E or D;- X54 is S, G or T;X55 is E or D;X56 is N or G;X57 is T, M or A;X58 is E or V;X59 is L, V or A;Xeo is V or E;Xei is K, E or Q or T;X62 is Y or F;X63 is T or S;X64 is D or G;Xes is V or I;X66 is K, D or E;-X67is T, I or V.

[0037] According to a particular embodiment, the polypeptide comprising, or consisting of, SEQ ID NO: 300 may be characterized in that:X12 is A, V, or K;X14 is L, or absent;X20 is T, S, or absent;X28 is K, or absent;X32 is K, V, R, I, or absent;X33 is Q, R, or absent;X34 is E, Q, V, A, S, or absent;X38 is A, V, or absent;X39 is E, D, A, R, or absent;- X40is D, E, N, S or G;X46 is H or R;Xei is K, E or Q.

[0038] According to said particular embodiment, SEQ ID NO: 300 corresponds to SEQ ID NO: 1.

[0039] According to particular embodiment, the polynucleotide of the invention encodes a polypeptide comprising or consisting of SEQ ID NO: 300, or a sequence with at least 70% sequence identity with SEQ ID NO: 300, preferably at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID NO: 300.

[0040] According to particular embodiment, the polynucleotide of the invention encodes a polypeptide comprising or consisting of SEQ ID NO: 1, or a sequence with at least 70% sequence identity with SEQ ID NO: 1, preferably at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID NO: 1.SEQ ID NO: 1PCCFAAGTX1VSTPX2GX3RX4IX5X6LX7X8GDX9VWSKPEX10GGX11PFAAX12X13X 14X15X16X17X18X19X20X21X22X23X24X25X26X27X28X29X30X31X32X33X34X35X36X37X38X39 X40EX41LX42VTPX43HPFYVX44AX45X46X47FX48PVIX49LKX50GDX51LQSLX52DGX5 3X54X55X56X57SSX58VESLELYX59PX60GX6iTX62NLX63VX64X65GHTFYVGX66LKT WVHNX67GPCwherein:Xi is M or K;- X2is D, G, N, S or R;X3 is E or D;X4 is A orT;X5 is D or E;Xe is T or S;- X7is K or N;Xs is V or I;X9 is I or V;- X10 is G, R or K;Xu is K or E;-X12 is A , V or K;X13 is I or absent;X14 is L or absent;X15 is A or absent;Xi6 is T or absent;X17 is H or absent;Xis is I, V, Q or absent;X19 is R or absent;X20 is T, S or absent;X21 is D or absent;X22 is Q or absent;X23 is P or absent;X24 is I or absent;X25 is Y or absent;X26 is R or absent;X27 is L or absent;X28 is K or absent;X29 is L or absent;X30 is K or absent;X31 is G, S or absent;X32 is K, V, R, I or absent;X33 is Q, R or absent;X34 is E, Q, V, A, S or absent;X35 is N, D, Y or absent;X36 is G or absent;X37 is any amino acid residue or is absent; preferably X37 is Q, E, K, N, T, I or is absent;X38 is A, V or absent;X39 is E, D, A, R or absent;- X40is D, E, N, S or G;- X41 is S, T or V;X42 is L or M;X43 is G or S;X44 is P or A;- X45is Q, H, R or K;X46 is H or R;X47 is G or D;X48 is V or I;X49 is D, N or G;X50 is P or A;X51 is R, Q or L;X52 is A, E, or D;X53 is A, E or D;- X54 is S, G or T;X55 is E or D;X56 is N or G;X57 is T, M or A;X58 is E or V;X59 is L, V or A;Xeo is V or E;Xei is K, E or Q;X62 is Y or F;X63 is T or S;X64 is D or G;Xes is V or I;X66 is K, D or E;-X67is T, I or V.

[0041] The polynucleotide may encode a polypeptide comprising or consisting of SEQ ID NO: 2, or a sequence with at least 70% sequence identity with SEQ ID NO: 2, preferably at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity with the amino acid sequence as set forth in SEQ ID NO: 2.SEQ ID NO: 2PCCFAAGTMVSTPDGX3RAIX5TLKX8GDX9VWSKPEGGGKPFAAAILATHIRTD QPIYRLKLKGKQENGX37AEDEX41LX42VTPGHPFYVPAQHGFX48PVIDLKPGDR LQSLADGASX55NTSSEVESLELYLPVGKTX62NLTVDX65GHTFYVGKLKTWVH NX67GPCwherein:X3 is E or D;X5 is D or E;Xs is V or I; X9 is I or V; X37 is any amino acid residue or is absent; preferably X37 is Q, E, K, N, T, I or is absent; - X41 is S, T or V;X42 is L or M; X48 is V or I; X55 is E or D; X62 is Y or F; Xes is V or I; and- X67is T, I or V.

[0042] The polypeptide encoded by the polynucleotide of the invention may be derived from a protein of Pseudomonas sp.

[0043] The polypeptide encoded by the polynucleotide of the invention may be derived from a protein of an organism selected from the group consisting of Pseudomonas syringae (such as, for example, Pseudomonas syringae pv. tomato, Pseudomonas syringae pv. Syringae, Pseudomonas syringae group sp. J248-6, Pseudomonas syringae group sp. J254-4, Pseudomonas syringae group genomosp. 3, Pseudomonas syringae group genomosp. 7, Pseudomonas syringae pv. actinidiae ICMP 18804, Pseudomonas syringae pv. Actinidifoliorum, Pseudomonas syringae pv. Actinidiae, Pseudomonas syringae pv. actinidiae ICMP 19096, Pseudomonas syringae pv. avii, Pseudomonas syringae pv. berberidis, Pseudomonas syringae pv. delphinii, Pseudomonas syringae pv. helianthin, Pseudomonas syringae pv. maculicola, Pseudomonas syringae pv. tagetis, Pseudomonas syringae pv. theae, and Pseudomonas syringae pv. theae ICMP 3923), Pseudomonas amygdali (such as, for example, Pseudomonas amygdali pv. Dendropanacis, Pseudomonas amygdali pv. lachrymans str. M302278, and Pseudomonas amygdali pv. Morsprunorum); Pseudomonas capsica, Pseudomonas caricapapayae, Pseudomonas cichorii, Pseudomonas mediterranea, Pseudomonas quasicaspiana, Pseudomonas sp. FP597, Pseudomonas lijiangensis, Pseudomonas trivialis, Pseudomonas viridiflava, and Pseudomonas sp. NPDC089734.

[0044] The polypeptide encoded by the polynucleotide of the invention may be derived from a protein of an organism selected from the group consisting of Pseudomonas triticifolii or Pseudomonas sp. StFLB209.

[0045] In particular, the polypeptide encoded by the polynucleotide of the invention may be derived from Pseudomonas syringae pv. tomato str. DC3000.

[0046] The polynucleotide of the present invention may encode a polypeptide having a sequence selected from the group comprising or consisting of SEQ ID NOs: 3 to 7, 12, 16 to 23, 26 to 47 or 49 to 102, or a sequence with at least 70% sequence identity with SEQ ID NOs: 3 to 7, 12, 16 to 23, 26 to 47 or 49 to 102 respectively, preferably at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity with the amino acid sequence as set forth in SEQ ID NOs: 3 to 7, 12, 16 to 23, 26 to 47 or 49 to 102 respectively.SEQ ID Origin Amino acid sequenceNO>KPY96595.1:14-183 PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG putative Filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQENGQAED 3 hemagglutinin, intein- ESLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA containing [Pseudomonas DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY syringae pv. tomato] VGKLKTWVHNTGPC>WP_233594450.1:50-219 PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG Hint domain-containing GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD 4 protein, partial ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA [Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY group genomosp. 7] VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_236536788.1 :44-213GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADDHint domain-containing5 ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA protein, partialDGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringae]VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_239690149.1 :74-243GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADDHint domain-containing6 ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA protein, partialDGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringae]VGKLKTWVHNTGPC>KWT08198.1:66-235 PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG 7 hypothetical protein GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD AL046_21435 ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA[Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY pv. avii] VGKLKTWVHNTGPCPCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_236471476.1 :99-268GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADDHint domain-containing ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAprotein, partialDGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringae]VGKLKTWVHNTGPC>RML55265.1:198-367PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGGputative Filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQENGQAEDhemagglutinin, intein- ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining [Pseudomonas DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFYamygdali pv.VGKLKTWVHNTGPCmorsprunorum]PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_191864289.1:116- GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD285 Hint domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas viridiflava]VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_206361494.1 :70-239GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADDHint domain-containing ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAprotein, partialDGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas viridiflava]VGKLKTWVHNTGPC>RMO90454.1:52-221 PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG putative Filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD hemagglutinin, intein- ESLLVTPGHPFYVPAHHGFVPVIDLKPGDRLQSLA containing [Pseudomonas DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY syringae pv. maculicola] VGKLKTWVHNTGPC>WP_183134590.1:59-228 PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG Hint domain-containing GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD protein, partial ESLLVTPGHPFYVPAHHGFVPVIDLKPGDRLQSLA [Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY group genomosp. 3] VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_237613636.1:338- GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD507 Hint domainESLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringae]VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>NAT61618.1:70-235GKPFAAAILATHIRTDQPIYRLKLKGKQENGQAEDhypothetical protein ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA[Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFYpv. actinidifoliorum]VGKLKTWVHNTGPC>WP_235810097.1:506- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG 675 DUF637 domainGKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD containing protein, partial ESLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA [Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFYgroup genomosp. 3] VGKLKTWVHNTGPC>KPZ34019.1:623-792 PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG hypothetical protein GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD AN901_205394 ESLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA [Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY pv. theae] VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_205901011.1:622- GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD791 DUF637 domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas viridiflava]VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>MBL3875093.1:654-823GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADDhypothetical protein ESLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA[Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFYpv. theae]VGKLKTWVHNTGPC>MDU8432958.1:848- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG 1017 DUF637 domainGKPFAAAILATHIRTDQPIYRLKLKGKQENGQAED containing protein ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA [Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY pv. actinidifoliorum] VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>SOS33758.1:813-982GKPFAAAILATHIRTDQPIYRLKLKGKQENGQAEDfilamentous hemagglutinin ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA[Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFYgroup genomosp. 3]VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>MBL3829555.1:826-995GKPFAAAILATHIRTDQPIYRLKLKGKQENGQAEDhypothetical protein ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA[Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFYpv. theae]VGKLKTWVHNTGPC>EPM65432.1:664-833PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGGfilamentousGKPFAAAILATHIRTDQPIYRLKLKGKQENGQADDhemagglutinin, intein- ESLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringaeVGKLKTWVHNTGPCpv. theae ICMP 3923]PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_205932634.1:656- GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD825 DUF637 domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas viridiflava]VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_122311120.1:630- GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD799 DUF637 domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringae]VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_205900893.1:685- GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD854 DUF637 domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas viridiflava]VGKLKTWVHNTGPC>GKS04618.1:1011-1180 PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG hypothetical protein GKPFAAAILATHIRTDQPIYRLKLKGKQENGQAED PSTH1771_06400 ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA [Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY pv. theae] VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_198431190.1:770- GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD939 DUF637 domainESLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringae]VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_191902158.1:752- GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD921 DUF637 domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas viridiflava]VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_205928731.1:896- GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD1065 DUF637 domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas viridiflava]VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_020342126.1:794- GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD963 DUF637 domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringae]VGKLKTWVHNTGPC>EPM96809.1:776-945filamentous PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG hemagglutinin, intein- GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD containing, partial ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA [Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY pv. actinidiae ICMP VGKLKTWVHNTGPC18804]PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_020357456.1:874- GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD1043 DUF637 domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringae]VGKLKTWVHNTGPC>WP_259643145.1:618- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG 787 DUF637 domainGKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD containing protein, partial ESLLVTPGHPFYVPAHHGFVPVIDLKPGDRLQSLA [Pseudomonas syringae DGASENMSSEVESLELYLPVGKTYNLTVGVGHTF group genomosp. 3] YVGKLKTWVHNTGPC>OOK94103.1:1074- 1243 PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG hypothetical protein GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD B0B36_24370 ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA [Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY pv. actinidifoliorum] VGKLKTWVHNTGPC>WP_236445336.1 :59-225PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGGHint domain-containing GEPFAAAILATHIRTDQPIYRLKLKGKQEDGKAEDprotein, partialETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA[Pseudomonas syringae]DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY VGKLKTWVHNTGPCPCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_236474237.1 :44-210GEPFAAAILATHIRTDQPIYRLKLKGKQEDGKAEDHint domain-containing ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAprotein, partialDGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringae]VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_044324420.1:262- GKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAED428 Hint domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas amygdali]VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_240326586.1:9-175GEPFAAAILATHIRTDQPIYRLKLKGKQEDGKAEDHint domain-containing ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAprotein, partialDGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringae]VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_202907079.1:281 - GKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAED447 Hint domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas amygdali]VGKLKTWVHNTGPC>KPX25314.1:284-450PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGGputative Filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAEDhemagglutinin, intein- ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas amygdaliVGKLKTWVHNTGPCpv. dendropanacis]>WP_259640311.1:959- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG 1128 DUF637 domainGKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD containing protein, partial ESLLVTPGHPFYVPAHHGFVPVIDLKPGDRLQSLA [Pseudomonas syringae DGASENMSSEVESLELYLPVGKTYNLTVGVGHTF group genomosp. 3] YVGKLKTWVHNTGPC>WP_284357690.1:5973- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG 6142 filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQENGQAED hemagglutinin N-terminal ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA domain-containing protein DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY [Pseudomonas syringae] VGKLKTWVHNTGPC> WP_284402911.1:5973- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG 6142 filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD hemagglutinin N-terminal ESLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA domain-containing protein DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY [Pseudomonas syringae] VGKLKTWVHNTGPC>WP_200380813.1:5666- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG5835 filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADDhemagglutinin N-terminal ESLLVTPGHPFYVPAHHGFVPVIDLKPGDRLQSLAdomain-containing protein DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringaeVGKLKTWVHNTGPCgroup genomosp. 3]PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG >WP_240326937.1:8-174GEPFAAAILATHIRTDQPIYRLKLKGKQEDGKAEDHint domain-containing ETLLVTPGHPFYVPAQHGFVPVIGLKPGDRLQSLAprotein, partialDGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringae]VGKLKTWVHNTGPC>WP_236693172.1:6293- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG6461 filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADDhemagglutinin N-terminal ESLLVTPGHPFYVPAHHGFVPVIDLKPGDRLQSLAdomain-containing protein DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringaeVGKLKTWVHNTGPCgroup genomosp. 3]PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>MCF5215773.1:345-511 GEPFAAAILATHIRTDQPIYRLKLKGKQEDGKAED hypothetical protein ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA [Pseudomonas syringae] DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>SOS26689.1:3345-3514GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADDfilamentous hemagglutinin ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA[Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFYpv. avii]VGKLKTWVHNTGPC>WP_259639694.1:5970- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG6138 filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADDhemagglutinin N-terminal ESLLVTPGHPFYVPAHHGFVPVIDLKPGDRLQSLAdomain-containing protein DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringaeVGKLKTWVHNTGPCgroup genomosp. 3]>RMM32737.1:5972-6140 PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG Filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD hemagglutinin ESLLVTPGHPFYVPAHHGFVPVIDLKPGDRLQSLA [Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY pv. berberidis] VGKLKTWVHNTGPC>WP_415224733.1:3360- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG 3529 polymorphic toxinGKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD type HINT domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA containing protein DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY [Pseudomonas syringae] VGKLKTWVHNTGPC>WP_235809963.1:4972- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG5140 polymorphic toxinGKPFAAAILATHIRTDQPIYRLKLKGKQENGQADDtype HINT domainESLLVTPGHPFYVPAHHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringaeVGKLKTWVHNTGPCgroup genomosp. 3]>WP_328586564.1:4621- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG4789 polymorphic toxinGKPFAAAILATHIRTDQPIYRLKLKGKQENGQADDtype HINT domainESLLVTPGHPFYVPAHHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringaeVGKLKTWVHNTGPCgroup genomosp. 3]>WP_183135443.1:2515- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG 2684 DUF637 domainGKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD containing protein ESLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA [Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY group genomosp. 3] VGKLKTWVHNTGPC>RMP10307.1:2556-2725 PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG putative Filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQENGQADD hemagglutinin, intein- ESLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA containing [Pseudomonas DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY syringae pv. delphinii] VGKLKTWVHNTGPC>WP_259641617.1:5565- 5734 MULTISPECIES: PCCFAAGTMVSTPGGERAIDTLKVGDIVWSKPEGG filamentous hemagglutinin GKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAED N-terminal domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA containing protein DGASENTSSEVESLELYLPVGKTYNLTVDIGHTFY [Pseudomonas syringae VGKLKTWVHNTGPCgroup]>WP_122239337.1:5631- 5800 MULTISPECIES: PCCFAAGTMVSTPGGERAIDTLKVGDIVWSKPEGG filamentous hemagglutinin GKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAED N-terminal domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA containing protein DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY [Pseudomonas syringae VGKLKTWVHNTGPCgroup]>WP_082427141.1:5631- PCCFAAGTMVSTPGGERAIDTLKVGDIVWSKPEGG5800 filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAEDhemagglutinin N-terminal ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAdomain-containing protein DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringaeVGKLKTWVHNTGPCgroup genomosp. 7]>WP_082441054.1:5631- PCCFAAGTMVSTPGGERAIDTLKVGDIVWSKPEGG5800 filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAEDhemagglutinin N-terminal ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAdomain-containing protein DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringaeVGKLKTWVHNTGPCgroup genomosp. 7]>KPY88866.1:5639-5808 PCCFAAGTMVSTPGGERAIDTLKVGDIVWSKPEGG Filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAED hemagglutinin, intein- ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA containing [Pseudomonas DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY syringae pv. tagetis] VGKLKTWVHNTGPC>RMW09244.1:5639-5808 PCCFAAGTMVSTPGGERAIDTLKVGDIVWSKPEGG putative Filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAED hemagglutinin, intein- ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA containing [Pseudomonas DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY syringae pv. tagetis] VGKLKTWVHNTGPC PCCFAAGTMVSTPGGERAIDTLKVGDIVWSKPEGG>KPX46873.1:5639-5808GKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAEDFilamentousETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAhemagglutinin, intein- DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFYcontaining proteinVGKLKTWVHNTGPC[Pseudomonas syringaepv. helianthi]>RMV43582.1:5608-5777 PCCFAAGTMVSTPGGERAIDTLKVGDIVWSKPEGG putative Filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAED hemagglutinin, intein- ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA containing [Pseudomonas DGASENTSSEVESLELYLPVGKTYNLTVDIGHTFY syringae pv. helianthi] VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_236478565.1:753- GEPFAAAILATHIRTDQPIYRLKLKGKQEDGKAED919 DUF637 domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringae]VGKLKTWVHNTGPC>WP_317845525.1:3678- PCCFAAGTMVSTPGGERAIDTLKVGDIVWSKPEGG3847 polymorphic toxinGKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAEDtype HINT domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[PseudomonasVGKLKTWVHNTGPCcaricapapayae]PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>SOP98770.1:886-1052GEPFAAAILATHIRTDQPIYRLKLKGKQEDGKAEDfilamentous hemagglutinin ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA[Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFYpv. syringae]VGKLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG>WP_236473346.1:943- GEPFAAAILATHIRTDQPIYRLKLKGKQEDGKAED1109 DUF637 domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringae]VGKLKTWVHNTGPC>WP_400698351.1:7-175 PCCFAAGTMVSTPDGDRAIDTLKVGDIVWSKPEKG polymorphic toxin-type GKPFAAAILATHVRTDQPIYRLKLKSVREDGQAEN HINT domain-containing ETLLVTPGHPFYVPAQRDFVPVIDLKPGDRLQSLED protein [Pseudomonas sp. GASENTSSEVESLELYLPEGKTYNLTVDIGHTFYVG NPDC089734] KLKTWVHNTGPC PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG> WP_319803841.1:1997- GKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAED2166 DUF637 domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringae]VGKLKTWVHNTGPC>WP_193601273.1:1801- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG1970 MULTISPECIES:GKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAED DUF637 domainETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining protein DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringaeVGKLKTWVHNTGPCgroup]>WP_316902652.1:1872- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG 2041 DUF637 domainGKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAED containing protein, partial ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA [Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFYgroup sp. J248-6] VGKLKTWVHNTGPC>RMP82444.1:1829-1998PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGGputative Filamentous GKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAEDhemagglutinin, intein- ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLAcontaining, partial DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY[Pseudomonas syringaeVGKLKTWVHNTGPCpv. actinidiae]>MDU8459873.1:1848- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG 2017 DUF637 domainGKPFAAAILATHIRTDQPIYRLKLKGKQEDGKAED containing protein ETLLVTPGHPFYVPAQHGFVPVIDLKPGDRLQSLA [Pseudomonas syringae DGASENTSSEVESLELYLPVGKTYNLTVDVGHTFY group sp. J254-4] VGKLKTWVHNTGPC>WP_198697465.1:4234- PCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGG 4399 filamentous GEPFAAAILATHIRTDQPIYRLKLKGRQEDGKAEDE hemagglutinin N-terminal TLLVTPGHPFYVPARHGFIPVIDLKPGDRLQSLADG domain-containing protein ASENTSSEVESLELYLPVGKTYNLTVDIGHTFYVG [Pseudomonas viridiflava] KLKTWVHNVGPC>WP_122316857.1:4155- PCCFAAGTMVSTPDGDRAIDTLKVGDIVWSKPERG 4323 filamentous GKPFAAAILATHVRTDQPIYRLKLKSVRQDGQAEE hemagglutinin N-terminal ETLLVTPGHPFYVPAQRDFVPVIDLKPGDRLQSLED domain-containing protein GASENTSSEVESLELYLPEGKTYNLTVDVGHTFYV [Pseudomonas cichorii] GKLKTWVHNTGPC>RMQ44380.1:4164-4332 PCCFAAGTMVSTPDGDRAIDTLKVGDIVWSKPERG putative Filamentous GKPFAAAILATHVRTDQPIYRLKLKSVRQDGQAEE hemagglutinin, intein- ETLLVTPGHPFYVPAQRDFVPVIDLKPGDRLQSLED containing [Pseudomonas GASENTSSEVESLELYLPEGKTYNLTVDVGHTFYV cichorii] GKLKTWVHNTGPC PCCFAAGTMVSTPDGDRAIDTLKVGDIVWSKPERG>WP_263948127.1:1703- GKPFAAAILATHVRTDQPIYRLKLKSVRQDGEVEN1872 DUF637 domainETLLVTPGHPFYVPAQRDFVPVIDLKPGDRLQSLEDcontaining protein, partial GASENTSSEVESLELYLPEGKTYNLTVDVGHTFYV[Pseudomonas capsici]GKLKTWVHNTGPC>WP_216705357.1:4404- PCCFAAGTMVSTPDGDRAIDTLKVGDIVWSKPEKG4573 filamentousGKPFAAAILATHVRTDQPIYRLKLKSVRQDGEAENhemagglutinin N-terminal ETLLVTPGHPFYVPAQRDFVPVIDLKPGDRLQSLEDdomain-containing protein GASENTSSEVESLELYLPEGKTYNLTVDVGHTFYV[PseudomonasGKLKTWVHNTGPClijiangensis]>WP_221592387.1:2566- PCCFAAGTMVSTPDGDRAIDTLKVGDIVWSKPEKG 2735 DUF637 domainGKPFAAAILATHVRTDQPIYRLKLKSVRQDGEAEN containing protein, partial ETLLVTPGHPFYVPAQRDFVPVIDLKPGDRLQSLED [Pseudomonas GASENTSSEVESLELYLPEGKTYNLTVDVGHTFYV lijiangensis] GKLKTWVHNTGPC>WP_221582494.1:2565- PCCFAAGTMVSTPDGDRAIDTLKVGDIVWSKPEKG 2734 DUF637 domainGKPFAAAILATHVRTDQPIYRLKLKSVRQDGEAEN containing protein, partial ETLLVTPGHPFYVPAQRDFVPVIDLKPGDRLQSLED [Pseudomonas GASENTSSEVESLELYLPEGKTYNLTVDVGHTFYV lijiangensis] GKLKTWVHNTGPC>WP_263938549.1:4416- PCCFAAGTMVSTPDGDRAIDTLKVGDIVWSKPEKG 4584 filamentous GKPFAAAILATHVRTDQPIYRLKLKSVRQYGEVEN hemagglutinin N-terminal ETLLVTPGHPFYVAAQRDFVPVIDLKPGDRLQSLDdomain-containing protein DGASDNTSSEVESLELYLPEGKTYNLTVDVGHTFY [Pseudomonas capsici] VGKLKTWVHNTGPC>WP_206402144.1:4306- PCCFAAGTMVSTPDGDRAIDTLKVGDIVWSKPERG 4474 filamentous GKPFAAAILATHVRTDQPIYRLKLKSVRQDGEVEN hemagglutinin N-terminal ETLLVTPGHPFYVAAQRDFVPVIDLKPGDRLQSLD domain-containing protein DGASDNTSSEVESLELYVPEGKTYNLTVDVGHTFY [Pseudomonas capsici] VGKLKTWVHNTGPC>WP_221606796.1:4539- PCCFAAGTMVSTPDGDRAIDTLKVGDIVWSKPEKG 4704 filamentous GKPFAAAILATHVRTDQPIYRLKLKSVREDGKAED hemagglutinin N-terminal ETLLVTPGHPFYVPAQRDFVPVIDLKPGDRLQSLED domain-containing protein GASENTSSEVESLELYLPEGKTYNLTVDVGHTFYV [Pseudomonas cichorii] GKLKTWVHNTGPC>EPN29108.1:11-153filamentousPCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEGGhemagglutinin, intein- GKPFAAVDETLLVTPGHPFYVPAQHGFVPVIDLKPcontaining, partial GDRLQSLADGASENTSSEVESLELYLPVGKTYNLT[Pseudomonas syringaeVDVGHTFYVGKLKTWVHNTGPCpv. actinidiae ICMP19096]PCCFAAGTKVSTPSGDRAIESLKVGDVVWSKPEKG>WP_239690305.1:17- 185GKPFAAKILATHQRSDQPIYRLKLKSVRADGTAAGHint domain-containing ETLLVTPSHPFYVPAKRDFIPVINLKPGDLLQSLADprotein [Pseudomonas GDSENTSSEVESLELYLPVGKTYNLSVDVGHTFYVsyringae]GELKTWVHNTGPC PCCFAAGTKVSTPNGDRAIESLKVGDVVWSKPERG>KRP58128.1:134-294GKPFAAKILATHQRSDQPIYRLKLKSVRVDGKAEGhypothetical protein ETLLVTPGHPFYVPAKRDFIPVIDLKPGDLLQSLAD TU79_21765GETDNTSSEVESLELYLPVGETFNLTVDVGHTFYV[Pseudomonas trivialis]GDLKTWVHNTGPC PCCFAAGTKVSTPNGDRAIESLKVGDVVWSKPERG>WP_231986886.1:113- GKPFAAKILATHQRSDQPIYRLKLKSVRVDGKAEG273 Hint domainETLLVTPGHPFYVPAKRDFIPVIDLKPGDLLQSLADcontaining proteinGETDNTSSEVESLELYLPVGETFNLTVDVGHTFYV[Pseudomonas trivialis]GDLKTWVHNTGPC>SDS12561.1:8-168 intein PCCFAAGTKVSTPNGDRAIESLKVGDVVWSKPERG C-terminal splicing GKPFAAKILATHQRSDQPIYRLKLKSVRVDGKAEG region / intein N-terminal ETLLVTPGHPFYVPAKRDFIPVIDLKPGDLLQSLAD splicing region GETDNTSSEVESLELYLPVGETFNLTVDVGHTFYV [Pseudomonas trivialis] GDLKTWVHNTGPC>WP_230167567.1 :26- 184PCCFAAGTMVSTPRGDRAIETLNVGDVVWSKPEKpolymorphic toxin-type GGKPFAAAILATHQRSDQPIYRLKLKGVRSDGIAR HINT domain-containing EETLLVTPSHPFYVPAKRDFIPVIDLKPGDLLQSLAprotein, partialDGDTENTSSEVESLELYAPVGKTYNLTVDIGHTFY[PseudomonasVGELKTWVHNTGPCmediterranea]>MDG6399431.1:161-325 PCCFAAGTMVSTPDGDRAIDTLKIGDIVWSKPEKG polymorphic toxin-type GKPFAAAILATHVRTDQPIYRLKLKSIRQDGNAESE 101 HINT domain-containing VLMVTPSHPFYVPARRDFVPVIDLKAGDQLQSLAD protein [Pseudomonas GAGEGASSVVESLELYLPVGKTYNLTVDVGHTFY quasicaspiana] VGKLKTWVHNTGPC PCCFAAGTKVSTPNGDRTIESLKVGDVVWSKPEKG> WP_305428193.1: 89-249GKPFAAKILATHQRSDQPIYRLKLKSVRADGKAEGHint domain-containing102 ETLLVTPGHPFYVPAKRDFIPVIDLKPGDLLQSLAD protein [Pseudomonas sp.GDTDNTSSEVESLELYLPVGQTFNLTVDVGHTFYV FP597]GDLKTWVHNTGPC

[0047] The polynucleotide of the present invention may encode a polypeptide having a sequence selected from the group comprising or consisting of SEQ ID NOs: 301 or 302, or a sequence with at least 70% sequence identity with SEQ ID NOs: 301 or 302 respectively, preferably at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity with the amino acid sequence as set forth in SEQ ID NOs: 301 or 302 respectively.SEQ ID Origin Amino acid sequenceNO WP_052483384.1 PCCFAAGTKVSTPDGDRAIETLKIGDIVWSKPEKG hemagglutinin repeatGKPFAATITATHVRNDQPIYRLTLKGSDLNGKTTR 301 containing protein ETLLVTPGHPFYVPAQKDFVPVIDLKLGDRLQSLA [Pseudomonas sp. DGATENTSSEVESLELYAPVGTTYNLTVDVGHTFY StFLB209] VGDLKTWVHNTGPC PCCFAAGTKVSTPNGDRTIESLKVGDVVWSKPEK WP_187519014.1 two- GGKPFAAKILATHQRSDQPIYRLKLKSVRADGTAApartner secretion domain302 GETLLVTPSHPFYVPAKHDFIPVIDLKPGDLLQSLA containing proteinDGDSDNTSSEVESLELYLPEGKTYNLTVDVGHTFY[Pseudomonas triticifolii]VGELKTWVHNTGPC

[0048] In some embodiments, the polynucleotide encodes a polypeptide having a sequence SEQ ID NO: 3, or a sequence with at least 70% sequence identity with SEQ ID NO: 3, preferably at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity with the amino acid sequence as set forth in SEQ ID NO: 3.

[0049] The polypeptide encoded by the polynucleotide of the invention may further comprise a sequence in Nter.

[0050] For example, the polypeptide may comprise, in Nter and contiguous to SEQ ID NO: 1, a sequence selected from the group comprising or consisting of G, DG, a dipeptide X69G (wherein X69 is D, K, L, A or E) and SEQ ID NO: 203 to 222, 239, 245, 246, and 248 to 251, and C-terminal fragments of SEQ ID NO: 245, 246, and 248 to 251, wherein X68is A, G, E or V, X69is D, K, L, A or E, X70 is M, E, P, F or S, X71 is F or L, X72 is K, L or N, X73 is R or N and X74 is G, L or N.SEQ ID NO Amino acid sequence203 X70X69G204 X68X70X69G205 X71 X68X70X69G206 X72X71 X68X70X69G207 X73X72X71 X68X70X69G208 X74X73X72X71 X68X70X69G209 AX74X73X72X71 X68X70X69G210 MDG211 X68MDG212 FX68MDG213 KFX68MDG214 RKFX68MDG215 GRKFX68MDG216 AGRKFX68MDG217 AMDG218 FAMDG219 KFAMDG220 RKFAMDG221 GRKFAMDG222 AGRKFAMDG239 GAGRKFAMDG245 ALRKLEPLG246 ANNNFGSEG248 GRKFGEAG249 GRKLGFDG250 GRKLGSEG251 GRLLGEKG

[0051] Preferably, the polypeptide comprises, in Nter and contiguous to SEQ ID NO: 1, a sequence selected from the group comprising or consisting of G, DG, and SEQ ID NOs: 217 to 222, 239, 245, 246, and 248 to 251, and C-terminal fragments of SEQ ID NO: 245, 246, 248, 249, 250 and 251.

[0052] In some embodiments, the polypeptide encoded by the polynucleotide of the invention comprises, from Nter to Cter, (i) a sequence selected from the group comprising or consisting of G, DG, and SEQ ID NOs: 217 to 222, 239, 245, 246, and 248 to 251 andC-terminal fragments of SEQ ID NO: 245, 246, 248, 249, 250 and 251, and (ii) a sequence of SEQ ID NO: 1. Preferably, both sequences are contiguous.

[0053] In some embodiments, the polypeptide encoded by the polynucleotide of the invention comprises, from Nter to Cter, (i) a sequence selected from the group comprising or consisting of G, DG, and SEQ ID NOs: 217 to 222, 239, 245, 246, and 248 to 251 and C-terminal fragments of SEQ ID NO: 245, 246, 248, 249, 250 and 251, and (ii) a sequence of SEQ ID NO: 3. Preferably, both sequences are contiguous.

[0054] The polypeptide encoded by the polynucleotide of the invention may further comprise a sequence in Cter.

[0055] For example, the polypeptide may comprise, in Cter and contiguous to SEQ ID NO: 1, a sequence selected from the group comprising or consisting of E, D, F, A, P, K, Q, T, R, V, a dipeptide EL, a dipeptide X75X76 and SEQ ID NO: 223 to 232, and 252 to 261 and N-terminal fragments of SEQ ID NO: 252 to 261, wherein X75 is E, D, F, A, P, K, Q, T, R or V, X76 is L or I and X77 is E or D.SEQ ID NO Amino acid sequence223 X75X76P224 X75X76PX77225 X75X76PX77G226 X75X76PX77GY227 X75X76PX77GYF228 ELP229 ELPE230 ELPEG231 ELPEGY232 ELPEGYF252 ALPEGYF253 DLPE254 DLPEGYF255 KLPD256 PI257 QLPDGYF258 RLPEGYF259 TLPD260 VLP261 VLPEGYF

[0056] Preferably, the polypeptide comprises, in Cter and contiguous to SEQ ID NO: 1, a sequence selected from the group comprising or consisting of E, EL, and SEQ ID NOs: 228 to 232, and 252 to 261 and N-terminal fragments of SEQ ID NO: 252 to 261.

[0057] In some embodiments, the polypeptide encoded by the polynucleotide of the invention comprises, from Nter to Cter, (i) a sequence of SEQ ID NO: 1 and (ii) a sequence selected from the group comprising or consisting of E, EL, and SEQ ID NOs: 228 to 232, and 252 to 261 and N-terminal fragments of SEQ ID NO: 252 to 261. Preferably, both sequences are contiguous.

[0058] In some embodiments, the polypeptide encoded by the polynucleotide of the invention comprises, from Nter to Cter, (i) a sequence of SEQ ID NO: 300 and (ii) a sequence selected from the group comprising or consisting of E, EL, and SEQ ID NOs: 228 to 232, and 252 to 261 and N-terminal fragments of SEQ ID NO: 252 to 261. Preferably, both sequences are contiguous.

[0059] In some embodiments, the polypeptide encoded by the polynucleotide of the invention comprises, from Nter to Cter, (i) a sequence of SEQ ID NO: 3 and (ii) a sequence selected from the group comprising or consisting of E, EL, and SEQ ID NOs: 228 to 232, and 252 to 261 and N-terminal fragments of SEQ ID NO: 252 to 261. Preferably, both sequences are contiguous.

[0060] In some embodiments, the polypeptide encoded by the polynucleotide of the invention comprises, from Nter to Cter, (i) a sequence selected from the group comprising or consisting of G, DG, and SEQ ID NOs: 217 to 222, 239, 245, 246, and 248 to 251 and C-terminal fragments of SEQ ID NO: 245, 246, and 248 to 251, (ii) a sequence of SEQ ID NO: 1 and (iii) a sequence selected from the group comprising or consisting of E, EL, and SEQ ID NOs: 228 to 232, and 252 to 261 and N-terminal fragments of SEQ ID NO: 252 to 261. Preferably, these sequences are contiguous.

[0061] In some embodiments, the polypeptide encoded by the polynucleotide of the invention comprises, from Nter to Cter, (i) a sequence selected from the group comprising or consisting of G, DG, and SEQ ID NOs: 217 to 222, 239, 245, 246, and 248 to 251 and C-terminal fragments of SEQ ID NO: 245, 246, and 248 to 251, (ii) a sequence of SEQID NO: 3 and (iii) a sequence selected from the group comprising or consisting of E, EL, and SEQ ID NOs: 228 to 232, and 252 to 261 and N-terminal fragments of SEQ ID NO: 252 to 261. Preferably, these sequences are contiguous.

[0062] Examples of polypeptides encoded by the polynucleotide of the invention and comprising additional sequences in Nter and in Cter include, but are not limited to, SEQ ID NOs: 103 to 202 and 264.SEQ ID Origin Amino acid sequenceNO>KPY96595.1:14-183 AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG putative Filamentous DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG 103 hemagglutinin, intein- KQENGQAEDESLLVTPGHPFYVPAQHGFVPVIDLK containing [Pseudomonas PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL syringae pv. tomato] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF >KPY96595.1:14-183 GAGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKV putative Filamentous GDIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLK 264 hemagglutinin, intein- GKQENGQAEDESLLVTPGHPFYVPAQHGFVPVIDL containing [Pseudomonas KPGDRLQSLADGASENTSSEVESLELYLPVGKTYN syringae pv. tomato] LTVDVGHTFYVGKLKTWVHNTGPCELPEGY >WP_233594450.1:50-219 AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG Hint domain-containing DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG 104 protein, partial KQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL group genomosp. 7] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_236536788.1 :44-213DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGHint domain-containing105 KQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLK protein, partialPGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_239690149.1 :74-243DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGHint domain-containing106 KQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLK protein, partialPGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF>KWT08198.1:66-235 AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG hypothetical protein DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG 107 AL046_21435 KQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL pv. avii] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF>WP_201367526.1:71-240 AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG Hint domain-containing DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG protein, partial KQENGQAEDESLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL group genomosp. 3] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF >WP_211270548.1 :36-205 AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG Hint domain-containing DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG protein, partial KQENGQAEDESLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL group genomosp. 3] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF >WP_199726832.1:81 -250 AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG Hint domain-containing DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG protein, partial KQENGQAEDESLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL group genomosp. 7] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>GAB0085919.1:81-250DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGhypothetical protein KQENGQAEDESLLVTPGHPFYVPAQHGFVPVIDLK TOC8172_56410PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_236471476.1 :99-268DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGHint domain-containing KQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLKprotein, partialPGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_259638394.1:ll-180DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGHint domain-containing KQENGQAEDESLLVTPGHPFYVPAQHGFVPVIDLKprotein, partialPGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF>EGH99855.1:4-173filamentous AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG hemagglutinin, intein- DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG containing, putative, KQENGQAEDESLLVTPGHPFYVPAQHGFVPVIDLK partial [Pseudomonas PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL amygdali pv. lachrymans TVDVGHTFYVGKLKTWVHNTGPCELPEGYF str. M302278]>WP_211267573.1:25-194 AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG Hint domain-containing DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG protein, partial KQENGQAEDESLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL group genomosp. 3] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF>RML55265.1:198-367AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVGputative Filamentous DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGhemagglutinin, intein- KQENGQAEDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining [Pseudomonas PGDRLQSLADGASENTSSEVESLELYLPVGKTYNLamygdali pv.TVDVGHTFYVGKLKTWVHNTGPCELPEGYFmorsprunorum]AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_191864289.1:116- DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG285 Hint domainKQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining protein, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas viridiflava]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_206361494.1 :70-239DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGHint domain-containing KQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLKprotein, partialPGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas viridiflava]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF>RMO90454.1:52-221 AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG putative Filamentous DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG hemagglutinin, intein- KQENGQADDESLLVTPGHPFYVPAHHGFVPVIDLK containing [Pseudomonas PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL syringae pv. maculicola] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF >WP_183134590.1:59-228 AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG Hint domain-containing DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG protein, partial KQENGQADDESLLVTPGHPFYVPAHHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL group genomosp. 3] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_237613636.1:338- DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG507 Hint domainKQENGQADDESLLVTPGHPFYVPAQHGFVPVIDLKcontaining protein, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>NAT61618.1:70-235DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGhypothetical protein KQENGQAEDETLLVTPGHPFYVPAQHGFVPVIDLK[Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNLpv. actinidifoliorum]TVDVGHTFYVGKLKTWVHNTGPCELP>WP_235810097.1:506- AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG 675 DUF637 domainDIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG containing protein, partial KQENGQADDESLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL group genomosp. 3] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF>WP_230853343.1:761- AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG 930 DUF637 domainDIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG containing protein, partial KQENGQAEDESLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL group genomosp. 3] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF >KPB95991.1:759-928 AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG Filamentous DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG hemagglutinin KQENGQAEDESLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL pv. maculicola] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF >KPZ34019.1:623-792 AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG hypothetical protein DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG AN901_205394 KQENGQADDESLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL pv. theae] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_205901011.1:622- DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG791 DUF637 domainKQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining protein, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas viridiflava]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>MBL3875093.1:654-823DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGhypothetical protein KQENGQADDESLLVTPGHPFYVPAQHGFVPVIDLK[Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNLpv. theae]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF>MDU8432958.1:848- AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG 1017 DUF637 domainDIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG containing protein KQENGQAEDETLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL pv. actinidifoliorum] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>SOS33758.1:813-982DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGfilamentous hemagglutinin KQENGQAEDETLLVTPGHPFYVPAQHGFVPVIDLK[Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNLgroup genomosp. 3]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>MBL3829555.1:826-995DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGhypothetical protein KQENGQAEDETLLVTPGHPFYVPAQHGFVPVIDLK[Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNLpv. theae]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>EPM65432.1:664-833DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGfilamentousKQENGQADDESLLVTPGHPFYVPAQHGFVPVIDLKhemagglutinin, intein- PGDRLQSLADGASENTSSEVESLELYLPVGKTYNLcontaining, partialTVDVGHTFYVGKLKTWVHNTGPCELPEGYF[Pseudomonas syringaepv. theae ICMP 3923]AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_205932634.1:656- DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG825 DUF637 domainKQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining protein, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas viridiflava]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_122311120.1:630- DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG799 DUF637 domainKQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining protein, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_205900893.1:685- DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG854 DUF637 domainKQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining protein, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas viridiflava]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF>GKS04618.1:1011-1180 AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG hypothetical protein DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG PSTH1771_06400 KQENGQAEDETLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL pv. theae] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_198431190.1:770- DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG939 DUF637 domainKQENGQADDESLLVTPGHPFYVPAQHGFVPVIDLKcontaining protein, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_191902158.1:752- DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG921 DUF637 domainKQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining protein, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas viridiflava]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_205928731.1:896- DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG1065 DUF637 domainKQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining protein, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas viridiflava]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_020342126.1:794- DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG963 DUF637 domainKQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining protein, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF>EPM96809.1:776-945filamentous AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG hemagglutinin, intein- DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG containing, partial KQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL pv. actinidiae ICMP TVDVGHTFYVGKLKTWVHNTGPCELPEGYF 18804]AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_020357456.1:874- DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG1043 DUF637 domainKQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining protein, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF>WP_259643145.1:618- AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG 787 DUF637 domainDIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG containing protein, partial KQENGQADDESLLVTPGHPFYVPAHHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENMSSEVESLELYLPVGKTYNL group genomosp. 3] TVGVGHTFYVGKLKTWVHNTGPCELPEGYF >OOK94103.1:1074- 1243 AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG hypothetical protein DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG B0B36_24370 KQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL pv. actinidifoliorum] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF AGRKFVMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_236445336.1 :59-225DIVWSKPEGGGEPFAAAILATHIRTDQPIYRLKLKGHint domain-containing KQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLKprotein, partialPGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTFYVGKLKTWVHNTGPCTLPD AGRKFVMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_236474237.1 :44-210DIVWSKPEGGGEPFAAAILATHIRTDQPIYRLKLKGHint domain-containing KQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLKprotein, partialPGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTFYVGKLKTWVHNTGPCTLPD AGRKFVMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_044324420.1:262- DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG428 Hint domainKQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining proteinPGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas amygdali]TVDVGHTFYVGKLKTWVHNTGPCKLPD>WP_011104446.1:5976- AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG6145 filamentous DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGhemagglutinin N-terminal KQENGQAEDESLLVTPGHPFYVPAQHGFVPVIDLKdomain-containing protein PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringaeTVDVGHTFYVGKLKTWVHNTGPCELPEGYFgroup genomosp. 3]AGRKFVMDGPCCFAAGTMVSTPDGERAIDTLKVG >WP_240326586.1:9-175DIVWSKPEGGGEPFAAAILATHIRTDQPIYRLKLKGHint domain-containing KQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLKprotein, partialPGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTFYVGKLKTWVHNTGPCTLPD AGRKFVMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_202907079.1:281 - DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG447 Hint domainKQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining protein, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas amygdali]TVDVGHTFYVGKLKTWVHNTGPCKLPD>KPX25314.1:284-450AGRKFVMDGPCCFAAGTMVSTPDGERAIDTLKVGputative Filamentous DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGhemagglutinin, intein- KQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas amygdaliTVDVGHTFYVGKLKTWVHNTGPCKLPDpv. dendropanacis]>WP_259640311.1:959- AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG 1128 DUF637 domainDIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG containing protein, partial KQENGQADDESLLVTPGHPFYVPAHHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENMSSEVESLELYLPVGKTYNL group genomosp. 3] TVGVGHTFYVGKLKTWVHNTGPCELPEGYF >WP_284357690.1:5973- AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG 6142 filamentous DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG hemagglutinin N-terminal KQENGQAEDETLLVTPGHPFYVPAQHGFVPVIDLK domain-containing protein PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL [Pseudomonas syringae] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF > WP_284402911.1:5973- AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG 6142 filamentous DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG hemagglutinin N-terminal KQENGQADDESLLVTPGHPFYVPAQHGFVPVIDLK domain-containing protein PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL [Pseudomonas syringae] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF >WP_200380813.1:5666- AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG5835 filamentous DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGhemagglutinin N-terminal KQENGQADDESLLVTPGHPFYVPAHHGFVPVIDLKdomain-containing protein PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringaeTVDVGHTFYVGKLKTWVHNTGPCELPEGYFgroup genomosp. 3]AGRKFVMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_240326937.1:8-174DIVWSKPEGGGEPFAAAILATHIRTDQPIYRLKLKGHint domain-containing KQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIGLKprotein, partialPGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTFYVGKLKTWVHNTGPCTLPD>WP_236693172.1:6293- GRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVGD6461 filamentousIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGKhemagglutinin N-terminal QENGQADDESLLVTPGHPFYVPAHHGFVPVIDLKPdomain-containing protein GDRLQSLADGASENTSSEVESLELYLPVGKTYNLT[Pseudomonas syringaeVDVGHTFYVGKLKTWVHNTGPCELPEGYFgroup genomosp. 3]AGRKFVMDGPCCFAAGTMVSTPDGERAIDTLKVG>MCF5215773.1:345-511 DIVWSKPEGGGEPFAAAILATHIRTDQPIYRLKLKG hypothetical protein KQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae] PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL TVDVGHTFYVGKLKTWVHNTGPCTLPD AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG>SOS26689.1:3345-3514DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGfilamentous hemagglutinin KQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLK[Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNLpv. avii]TVDVGHTFYVGKLKTWVHNTGPCELPEGYF>WP_259639694.1:5970- GRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVGD6138 filamentousIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGKhemagglutinin N-terminal QENGQADDESLLVTPGHPFYVPAHHGFVPVIDLKPdomain-containing protein GDRLQSLADGASENTSSEVESLELYLPVGKTYNLT[Pseudomonas syringaeVDVGHTFYVGKLKTWVHNTGPCELPEGYFgroup genomosp. 3]>RMM32737.1:5972-6140 GRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVGD Filamentous IVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGK hemagglutinin QENGQADDESLLVTPGHPFYVPAHHGFVPVIDLKP [Pseudomonas syringae GDRLQSLADGASENTSSEVESLELYLPVGKTYNLT pv. berberidis] VDVGHTFYVGKLKTWVHNTGPCELPEGYF >WP_415224733.1:3360- AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG 3529 polymorphic toxinDIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG type HINT domainKQENGQADDETLLVTPGHPFYVPAQHGFVPVIDLK containing protein PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL [Pseudomonas syringae] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF >WP_235809963.1:4972- GRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVGD5140 polymorphic toxinIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGKtype HINT domainQENGQADDESLLVTPGHPFYVPAHHGFVPVIDLKPcontaining protein, partial GDRLQSLADGASENTSSEVESLELYLPVGKTYNLT[Pseudomonas syringaeVDVGHTFYVGKLKTWVHNTGPCELPEGYFgroup genomosp. 3]>WP_328586564.1:4621- GRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVGD4789 polymorphic toxinIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGKtype HINT domainQENGQADDESLLVTPGHPFYVPAHHGFVPVIDLKPcontaining protein, partial GDRLQSLADGASENTSSEVESLELYLPVGKTYNLT[Pseudomonas syringaeVDVGHTFYVGKLKTWVHNTGPCELPEGYFgroup genomosp. 3]>WP_183135443.1:2515- AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG 2684 DUF637 domainDIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG containing protein KQENGQADDESLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL group genomosp. 3] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF >RMP10307.1:2556-2725 AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG putative Filamentous DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG hemagglutinin, intein- KQENGQADDESLLVTPGHPFYVPAQHGFVPVIDLK containing [Pseudomonas PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL syringae pv. delphinii] TVDVGHTFYVGKLKTWVHNTGPCELPEGYF >WP_259641617.1:5565- 5734 MULTISPECIES: AGRKFAMDGPCCFAAGTMVSTPGGERAIDTLKVG filamentous hemagglutinin DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG N-terminal domainKQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLK containing protein PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL [Pseudomonas syringae TVDIGHTFYVGKLKTWVHNTGPCQLPDGYF group]>WP_122239337.1:5631- 5800 MULTISPECIES: AGRKFAMDGPCCFAAGTMVSTPGGERAIDTLKVG filamentous hemagglutinin DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG N-terminal domainKQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLK containing protein PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL [Pseudomonas syringae TVDVGHTFYVGKLKTWVHNTGPCQLPDGYF group]>WP_082427141.1:5631- AGRKFAMDGPCCFAAGTMVSTPGGERAIDTLKVG5800 filamentous DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGhemagglutinin N-terminal KQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLKdomain-containing protein PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringaeTVDVGHTFYVGKLKTWVHNTGPCQLPDGYF group genomosp. 7]>WP_082441054.1:5631- AGRKFAMDGPCCFAAGTMVSTPGGERAIDTLKVG5800 filamentous DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGhemagglutinin N-terminal KQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLKdomain-containing protein PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringaeTVDVGHTFYVGKLKTWVHNTGPCQLPDGYF group genomosp. 7]>KPY88866.1:5639-5808 AGRKFAMDGPCCFAAGTMVSTPGGERAIDTLKVG Filamentous DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG hemagglutinin, intein- KQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLK containing [Pseudomonas PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL syringae pv. tagetis] TVDVGHTFYVGKLKTWVHNTGPCQLPDGYF >RMW09244.1:5639-5808 AGRKFAMDGPCCFAAGTMVSTPGGERAIDTLKVG putative Filamentous DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG hemagglutinin, intein- KQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLK containing [Pseudomonas PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL syringae pv. tagetis] TVDVGHTFYVGKLKTWVHNTGPCQLPDGYF>KPX46873.1:5639-5808AGRKFAMDGPCCFAAGTMVSTPGGERAIDTLKVGFilamentousDIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGhemagglutinin, intein- KQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining proteinPGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringaeTVDVGHTFYVGKLKTWVHNTGPCQLPDGYFpv. helianthi]>RMV43582.1:5608-5777 AGRKFAMDGPCCFAAGTMVSTPGGERAIDTLKVG putative Filamentous DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG hemagglutinin, intein- KQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLK containing [Pseudomonas PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL syringae pv. helianthi] TVDIGHTFYVGKLKTWVHNTGPCQLPDGYF AGRKFVMDGPCCFAAGTMVSTPDGERAIDTLKVG>WP_236478565.1:753- DIVWSKPEGGGEPFAAAILATHIRTDQPIYRLKLKG919 DUF637 domainKQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining protein, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTFYVGKLKTWVHNTGPCTLPD>WP_317845525.1:3678- AGRKFAMDGPCCFAAGTMVSTPGGERAIDTLKVG3847 polymorphic toxinDIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGtype HINT domainKQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining protein, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[PseudomonasTVDVGHTFYVGKLKTWVHNTGPCQLPDGYFcaricapapayae]AGRKFVMDGPCCFAAGTMVSTPDGERAIDTLKVG>SOP98770.1:886-1052DIVWSKPEGGGEPFAAAILATHIRTDQPIYRLKLKGfilamentous hemagglutinin KQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLK[Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNLpv. syringae]TVDVGHTFYVGKLKTWVHNTGPCTLPD AGRKEVMDGPCCEAAGTMVSTPDGERAIDTLKVG>WP_236473346.1:943- DIVWSKPEGGGEPEAAAILATHIRTDQPIYRLKLKG1109 DUF637 domainKQEDGKAEDETLLVTPGHPEYVPAQHGEVPVIDLKcontaining protein, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTEYVGKLKTWVHNTGPCTLPD>WP_400698351.1:7-175 GRKLGEDGPCCEAAGTMVSTPDGDRAIDTLKVGDI polymorphic toxin-type VWSKPEKGGKPEAAAILATHVRTDQPIYRLKLKSV HINT domain-containing REDGQAENETLLVTPGHPEYVPAQRDEVPVIDLKP protein [Pseudomonas sp. GDRLQSLEDGASENTSSEVESLELYLPEGKTYNLT NPDC089734] VDIGHTEYVGKLKTWVHNTGPCRLPEGYE AGRKEVMDGPCCEAAGTMVSTPDGERAIDTLKVG> WP_319803841.1:1997- DIVWSKPEGGGKPEAAAILATHIRTDQPIYRLKLKG2166 DUF637 domainKQEDGKAEDETLLVTPGHPEYVPAQHGEVPVIDLKcontaining protein, partial PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringae]TVDVGHTEYVGKLKTWVHNTGPCQLPDGYE>WP_193601273.1:1801- AGRKFVMDGPCCFAAGTMVSTPDGERAIDTEKVG1970 MULTISPECIES:DIVWSKPEGGGKPFAAAIEATHIRTDQPIYREKEKG DUF637 domainKQEDGKAEDETEEVTPGHPFYVPAQHGFVP VIDEKcontaining proteinPGDREQSEADGASENTSSEVESEEEYEPVGKTYNE[Pseudomonas syringaeTVDVGHTFYVGKEKTWVHNTGPCQEPDGYFgroup]>WP_316902652.1:1872- AGRKFVMDGPCCFAAGTMVSTPDGERAIDTEKVG 2041 DUF637 domainDIVWSKPEGGGKPFAAAIEATHIRTDQPIYREKEKG containing protein, partial KQEDGKAEDETEEVTPGHPFYVPAQHGFVP VIDEK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL group sp. J248-6] TVDVGHTFYVGKEKTWVHNTGPCQEPDGYF >RMP82444.1:1829-1998AGRKFVMDGPCCFAAGTMVSTPDGERAIDTLKVGputative Filamentous DIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGhemagglutinin, intein- KQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLKcontaining, partialPGDRLQSLADGASENTSSEVESLELYLPVGKTYNL[Pseudomonas syringaeTVDVGHTFYVGKEKTWVHNTGPCQEPDGYFpv. actinidiae]>MDU8459873.1:1848- AGRKFVMDGPCCFAAGTMVSTPDGERAIDTLKVG 2017 DUF637 domainDIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKG containing protein KQEDGKAEDETLLVTPGHPFYVPAQHGFVPVIDLK [Pseudomonas syringae PGDRLQSLADGASENTSSEVESLELYLPVGKTYNL group sp. J254-4] TVDVGHTFYVGKEKTWVHNTGPCQEPDGYF >WP_198697465.1:4234- GRKFGEAGPCCFAAGTMVSTPDGERAIDTLKVGDI 4399 filamentous VWSKPEGGGEPFAAAILATHIRTDQPIYRLKLKGR hemagglutinin N-terminal QEDGKAEDETLLVTPGHPFYVPARHGFIPVIDLKPG domain-containing protein DRLQSLADGASENTSSEVESLELYLPVGKTYNLTV [Pseudomonas viridiflava] DIGHTFYVGKLKTWVHNVGPCELPE >WP_122316857.1:4155- GRKLGSEGPCCFAAGTMVSTPDGDRAIDTLKVGDI 4323 filamentous VWSKPERGGKPFAAAILATHVRTDQPIYRLKLKSV hemagglutinin N-terminal RQDGQAEEETLLVTPGHPFYVPAQRDFVPVIDLKP domain-containing protein GDRLQSLEDGASENTSSEVESLELYLPEGKTYNLT [Pseudomonas cichorii] VDVGHTFYVGKLKTWVHNTGPCVLPEGYF >RMQ44380.1:4164-4332 GRKLGSEGPCCFAAGTMVSTPDGDRAIDTLKVGDI putative Filamentous VWSKPERGGKPFAAAILATHVRTDQPIYRLKLKSV hemagglutinin, intein- RQDGQAEEETLLVTPGHPFYVPAQRDFVPVIDLKP containing [Pseudomonas GDRLQSLEDGASENTSSEVESLELYLPEGKTYNLT cichorii] VDVGHTFYVGKLKTWVHNTGPCVLPEGYF ANNNFGSEGPCCFAAGTMVSTPDGDRAIDTLKVG>WP_263948127.1:1703- DIVWSKPERGGKPFAAAILATHVRTDQPIYRLKLK1872 DUF637 domainSVRQDGEVENETLLVTPGHPFYVPAQRDFVPVIDLcontaining protein, partial KPGDRLQSLEDGASENTSSEVESLELYLPEGKTYN[Pseudomonas capsici]LTVDVGHTFYVGKLKTWVHNTGPCVLPEGYF>WP_216705357.1:4404- ANNNFGSEGPCCFAAGTMVSTPDGDRAIDTLKVG4573 filamentousDIVWSKPEKGGKPFAAAILATHVRTDQPIYRLKLKhemagglutinin N-terminal SVRQDGEAENETLLVTPGHPFYVPAQRDFVPVIDLdomain-containing protein KPGDRLQSLEDGASENTSSEVESLELYLPEGKTYN[PseudomonasLTVDVGHTFYVGKLKTWVHNTGPCVLPEGYF lijiangensis]>WP_221592387.1:2566- ANNNFGSEGPCCFAAGTMVSTPDGDRAIDTLKVG 2735 DUF637 domainDIVWSKPEKGGKPFAAAILATHVRTDQPIYRLKLK containing protein, partial SVRQDGEAENETLLVTPGHPFYVPAQRDFVPVIDL [Pseudomonas KPGDRLQSLEDGASENTSSEVESLELYLPEGKTYN lijiangensis] LTVDVGHTFYVGKLKTWVHNTGPCVLPEGYF >WP_221582494.1:2565- ANNNFGSEGPCCFAAGTMVSTPDGDRAIDTLKVG 2734 DUF637 domainDIVWSKPEKGGKPFAAAILATHVRTDQPIYRLKLK containing protein, partial SVRQDGEAENETLLVTPGHPFYVPAQRDFVPVIDL [Pseudomonas KPGDRLQSLEDGASENTSSEVESLELYLPEGKTYN lijiangensis] LTVDVGHTFYVGKLKTWVHNTGPCVLPEGYF >WP_263938549.1:4416- GRKLGSEGPCCFAAGTMVSTPDGDRAIDTLKVGDI 4584 filamentous VWSKPEKGGKPFAAAILATHVRTDQPIYRLKLKSV hemagglutinin N-terminal RQYGEVENETLLVTPGHPFYVAAQRDFVPVIDLKP domain-containing protein GDRLQSLDDGASDNTSSEVESLELYLPEGKTYNLT [Pseudomonas capsici] VDVGHTFYVGKLKTWVHNTGPCVLPEGYF >WP_206402144.1:4306- GRKLGSEGPCCFAAGTMVSTPDGDRAIDTLKVGDI 4474 filamentous VWSKPERGGKPFAAAILATHVRTDQPIYRLKLKSV hemagglutinin N-terminal RQDGEVENETLLVTPGHPFYVAAQRDFVPVIDLKP domain-containing protein GDRLQSLDDGASDNTSSEVESLELYVPEGKTYNLT [Pseudomonas capsici] VDVGHTFYVGKLKTWVHNTGPCVLPEGYF >WP_221606796.1:4539- ANNNFGSEGPCCFAAGTMVSTPDGDRAIDTLKVG 4704 filamentous DIVWSKPEKGGKPFAAAILATHVRTDQPIYRLKLK hemagglutinin N-terminal SVREDGKAEDETLLVTPGHPFYVPAQRDFVPVIDL domain-containing protein KPGDRLQSLEDGASENTSSEVESLELYLPEGKTYN [Pseudomonas cichorii] LTVDVGHTFYVGKLKTWVHNTGPCVLP >EPN29108.1:11-153filamentous AGRKFAMDGPCCFAAGTMVSTPDGERAIDTLKVG hemagglutinin, intein- DIVWSKPEGGGKPFAAVDETLLVTPGHPFYVPAQH containing, partial GFVPVIDLKPGDRLQSLADGASENTSSEVESLELYL [Pseudomonas syringae PVGKTYNLTVDVGHTFYVGKLKTWVHNTGPCELP pv. actinidiae ICMP EGYF19096]GRLLGEKGPCCFAAGTKVSTPSGDRAIESLKVGDV>WP_239690305.1:17- 185VWSKPEKGGKPFAAKILATHQRSDQPIYRLKLKSVHint domain-containing RADGTAAGETLLVTPSHPFYVPAKRDFIPVINLKPGprotein [Pseudomonas DLLQSLADGDSENTSSEVESLELYLPVGKTYNLSVsyringae]DVGHTFYVGELKTWVHNTGPCDLPEGYFPCCFAAGTKVSTPNGDRAIESLKVGDVVWSKPERG >KRP58128.1:134-294GKPFAAKILATHQRSDQPIYRLKLKSVRVDGKAEGhypothetical protein197 ETLLVTPGHPFYVPAKRDFIPVIDLKPGDLLQSLAD TU79_21765GETDNTSSEVESLELYLPVGETFNLTVDVGHTFYV[Pseudomonas trivialis]GDLKTWVHNTGPCALPEGYF PCCFAAGTKVSTPNGDRAIESLKVGDVVWSKPERG>WP_231986886.1:113- GKPFAAKILATHQRSDQPIYRLKLKSVRVDGKAEG273 Hint domain198 ETLLVTPGHPFYVPAKRDFIPVIDLKPGDLLQSLAD containing proteinGETDNTSSEVESLELYLPVGETFNLTVDVGHTFYV[Pseudomonas trivialis]GDLKTWVHNTGPCALPEGYF>SDS12561.1:8-168 intein PCCFAAGTKVSTPNGDRAIESLKVGDVVWSKPERG C-terminal splicing GKPFAAKILATHQRSDQPIYRLKLKSVRVDGKAEG 199 region / intein N-terminal ETLLVTPGHPFYVPAKRDFIPVIDLKPGDLLQSLAD splicing region GETDNTSSEVESLELYLPVGETFNLTVDVGHTFYV [Pseudomonas trivialis] GDLKTWVHNTGPCALPEGYF>WP_230167567.1 :26- 184GPCCFAAGTMVSTPRGDRAIETLNVGDVVWSKPEpolymorphic toxin-type KGGKPFAAAILATHQRSDQPIYRLKLKGVRSDGIA HINT domain-containing200 REETLLVTPSHPFYVPAKRDFIPVIDLKPGDLLQSL protein, partialADGDTENTSSEVESLELYAPVGKTYNLTVDIGHTF[PseudomonasYVGELKTWVHNTGPCDLPEmediterranea]>MDG6399431.1: 161-325 ALRKLEPLGPCCFAAGTMVSTPDGDRAIDTLKIGDI polymorphic toxin-type VWSKPEKGGKPFAAAILATHVRTDQPIYRLKLKSI 201 HINT domain-containing RQDGNAESEVLMVTPSHPFYVPARRDFVPVIDLKA protein [Pseudomonas GDQLQSLADGAGEGASSVVESLELYLPVGKTYNL quasicaspiana] TVDVGHTFYVGKLKTWVHNTGPCPI PCCFAAGTKVSTPNGDRTIESLKVGDVVWSKPEKG> WP_305428193.1: 89-249GKPFAAKILATHQRSDQPIYRLKLKSVRADGKAEGHint domain-containing202 ETLLVTPGHPFYVPAKRDFIPVIDLKPGDLLQSLAD protein [Pseudomonas sp.GDTDNTSSEVESLELYLPVGQTFNLTVDVGHTFYV FP597]GDLKTWVHNTGPCFLPEGYF

[0063] Other examples of polypeptides encoded by the polynucleotide of the invention and comprising additional sequences in Nter and in Cter include, but are not limited to, SEQ ID NOs: 303 and 304.SEQ ID Origin Amino acid sequenceNO ANAAPEIGPCCFAAGTKVSTPDGDRAIETLKIGDIV>WP_052483384.1WSKPEKGGKPFAATITATHVRNDQPIYRLTLKGSDhemagglutinin repeatLNGKTTRETLLVTPGHPFYVPAQKDFVPVIDLKLG303 containing proteinDRLQSLADGATENTSSEVESLELYAPVGTTYNLTV[Pseudomonas sp.DVGHTFYVGDLKTWVHNTGPCDPAARGStFLB209]GRLLGEKGPCCFAAGTKVSTPNGDRTIESLKVGDV >WP_187519014.1 two- VWSKPEKGGKPFAAKILATHQRSDQPIYRLKLKSVpartner secretion domainRADGTAAGETLLVTPSHPFYVPAKHDFIPVIDLKPG304 containing proteinDEEQSEADGDSDNTSSEVESEEEYEPEGKTYNETV[Pseudomonas triticifolii]DVGHTFYVGEEKTWVHNTGPCDEPDGY

[0064] A non-limitative example of a polynucleotide of the present invention is a polynucleotide having a sequence comprising or consisting of SEQ ID NO: 233.SEQ ID NO: 233ccgtgctgcttcgccgctggtacgatggtctcaacgccggatggtgaacgtgctattgatacgctgaaagttggtgatattgtct ggtcaaaaccggaaggcggtggcaaaccgtttgcggccgcaattctggccacccatatccgtacggatcagccgatctatcg cctgaaactgaaaggtaaacaggaaaacggccaagccgaagacgaatcgctgctggtgaccccgggtcatccgttttacgtt ccggcacagcacggcttcgtgccggttattgatctgaaaccgggtgaccgtctgcaaagtctggctgatggcgcgtccgaaa acaccagctctgaagtcgaatctctggaactgtatctgccggttggtaaaacctacaatctgacggtcgacgtgggccacacgt tctatgttggtaaactgaaaacctgggtccataatacgggtccgtgt

[0065] Another object of the present invention is a polynucleotide for expressing a chimeric polyprotein, wherein said polynucleotide comprises or consists of, from 5’ to 3’:optionally a sequence encoding a signal peptide,a sequence encoding a first protein entity,a polynucleotide as disclosed herein and encoding a polypeptide comprising or consisting of SEQ ID NO: 300, or SEQ ID NO: 1, or a sequence with at least 70% sequence identity with SEQ ID NO: 1, preferably at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID NO: 300 or SEQ ID NO: 1, anda sequence encoding a second protein entity,wherein the first and second protein entities may be the same or different.

[0066] According to the present invention, the chimeric polyprotein comprises said first protein entity and said second protein entity, linked together by a disulfide bridge.Accordingly, according to the present invention, the plurality of protein entities comprises said first and second protein entities, wherein said protein entities are not linked together.

[0067] The polynucleotide for expressing a chimeric polyprotein may further comprise linkers between the different entities, such as, for example, between the sequence encoding the first protein entity and the polynucleotide of the present invention, and / or between the polynucleotide of the present invention and the sequence encoding the second protein entity. Examples of linkers include, but are not limited to, GS linkers or a GG linker (that may be encoded by a DNA sequence ggtggt).

[0068] Examples of signal peptides that may be used are well-known in the art, and include, without limitation, the signal peptide from PrtS protein of Streptococcus thermophilus (having for example a sequence SEQ ID NO: 234, that may be encoded by the DNA sequence SEQ ID NO: 235), the signal peptide from DsbA protein of E. coli (having for example a sequence SEQ ID NO: 236, that may be encoded by the DNA sequence SEQ ID NO: 237) or the signal peptide from PelB protein of Pectobacterium carotovorum (having for example a sequence SEQ ID NO: 286, that may be encoded by the DNA sequence SEQ ID NO: 287).SEQSequenceID NO Type234 Amino acid MKKKETFSLRKYKIGTVSVLLGAVFLFAGAPSVAADE atgaaaaagaaagaaactttctcacttcggaagtataaaattggaactgtgtctgttcttttg 235 DNAggtgcagtttttttgtttgcaggtgcaccatcggtagctgcagatgaa236 Amino acid MKKIWLALAGLVLAFSASA237 DNA atgaaaaagatttggctggcgctggctggtttagttttagcgtttagcgcatcggcg286 Amino acid MKYLLPTAAAGLLLLAAQPAMA287 DNA atgaaatacctgctgccgaccgctgctgctggtctgctgcttctcgctgcccagccggcgatggcc

[0069] Another object of the present invention is an expression cassette comprising the polynucleotide as described herein, or the polynucleotide for expressing a chimeric polyprotein as described herein. In particular, said expression cassette may comprise a promoter. Examples of promoters include, but are not limited to, pLac, pT7, pTrc, pTac, pBAD, and pRha.

[0070] Another object of the present invention is a vector, preferably an expression vector comprising the polynucleotide as described herein, the polynucleotide for expressing a chimeric polyprotein as described herein, or the expression cassette as described herein.

[0071] Examples of vectors include, but are not limited to, plasmids, viral vectors, artificial chromosomes, liposomes, and lipid nanoparticles.

[0072] In some embodiments, the vector or expression vector is a plasmid.

[0073] Another object of the present invention is a host cell comprising the polynucleotide as described herein, the polynucleotide for expressing a chimeric polyprotein as described herein, or the expression cassette as described herein.

[0074] The host cell may be a Prokaryote or a Eukaryote. In particular, the host cell may be a bacterium, including a Gram-negative or a Gram-positive bacterium. A non-limitative example of a Gram-negative bacterium is Escherichia coli. A non-limitative example of a Gram-positive bacterium is Streptococcus thermophilus.

[0075] Another object of the present invention is a method for producing a chimeric polyprotein, wherein said chimeric polyprotein comprises a first protein entity linked to a second protein entity by a cleavable sequence comprising a disulfide bridge, wherein said method comprises expressing a polynucleotide for expressing a chimeric polyprotein as described herein. According to the present invention, said polyprotein does not comprise the polypeptide of SEQ ID NO: 1 and / or SEQ ID NO: 300.

[0076] In some embodiments, the method for producing a chimeric polyprotein is an in vivo method and comprises a step of culturing a host cell comprising a polynucleotide for expressing a chimeric polyprotein as described herein, in conditions suitable for expression of said polynucleotide of interest (including transcription and translation thereof). Preferably, said host cell is a bacterium.

[0077] In some embodiments, the host cell is a Gram-negative bacterium, and the polynucleotide for expressing a chimeric polyprotein as described herein comprises a sequence encoding a signal peptide, so that a chimeric polyprotein is secreted in the periplasm of the bacterium.

[0078] In some embodiments, the host cell is a Gram-positive bacterium, and the polynucleotide for expressing a chimeric polyprotein as described herein comprises a sequence encoding a signal peptide, so that a chimeric polyprotein is secreted in the extracellular medium.

[0079] In some embodiments, the method for producing a chimeric polyprotein is an in vitro method and comprises a step of transcription and translation of a polynucleotide for expressing a chimeric polyprotein as described herein. Examples of systems that may be used for in vitro transcribing and translating a DNA sequence include (that may be referred to IVTT systems), but are not limited to, RTS kit (provided by VWR) or TNT kit (provided by Promega).

[0080] Said cleavable sequence consists of two separated peptides, the first one comprising a C-terminal cysteine and the second one comprising a cysteine at position +4 from the N-terminus, both cysteines being linked by a disulfide bridge.

[0081] Said two separated peptides comprise or consist of amino acids derived from the polypeptide as described herein, comprising SEQ ID NO: 1 or SEQ ID NO: 300 or a sequence having at least 70% identity with SEQ ID NO: 1, and wherein said polypeptide optionally further comprises N-terminal and / or C-terminal additional sequences. For the first of the two polypeptides, said amino acids consist of amino acids located N-terminally from the cysteine at position 2 with reference to SEQ ID NO: 1 or SEQ ID NO: 300 (and including said cysteine). For the second of the two polypeptides, said amino acids consist of the four C-terminal amino acids of SEQ ID NO: 1 or SEQ ID NO: 300 and of amino acids located C-terminally from the cysteine at position 154 with reference to SEQ ID NO: 1 or SEQ ID NO: 300 (said amino acid thus include the C-terminal cysteine of SEQ ID NO: 1 or SEQ ID NO: 300). Both cysteines are linked by a disulfide bridge and are not linked by a peptide bond.

[0082] Thus, the cleavable sequence comprises or consists of a first peptide comprising or consisting of a sequence PC, and a second peptide comprising or consisting of a sequence XevGPC (SEQ ID NO: 238, wherein X67 is T, I or V), wherein the cysteine of the first and second polypeptides are linked by a disulfide bridge. Preferably, the second polypeptide comprises or consists of a sequence TGPC, SEQ ID NO: 290).

[0083] Examples of first peptides include, but are not limited to, PC, SEQ ID NO: 240 (GAGRKFAMDGPC) and fragments of 3 to 12 contiguous amino acids of SEQ ID NO: 240 comprising the dipeptide PC.

[0084] Other examples of first peptides include, but are not limited to, SEQ ID NO: 241 or fragments of 3 to 11 contiguous amino acids of SEQ ID NO: 241 (AX74X73X72X71X68X70X69GPC wherein Xes is A, G, E or V, X69 is D, K, L, A or E, X70 is M, E, P, F or S, X71 is F or L, X72 is K, L or N, X73 is R or N and X74 is G, L or N) comprising the dipeptide PC.

[0085] Examples of second peptides include, but are not limited to, X67GPC (SEQ ID NO: 238, wherein X67 is T, I or V), SEQ ID NO: 291 and fragments of 5 to 10 contiguous amino acids of SEQ ID NO: 291 (X67GPCELPEGYF) comprising SEQ ID NO: 238.

[0086] Other examples of second peptides include, but are not limited to, SEQ ID NO: 292 or fragments of 4 to 10 contiguous amino acids of SEQ ID NO: 292 (X67GPCX75X76PX77GYF wherein X67is T, I or V, X75 is E, D, F, A, P, K, Q, T, R or V, X76 is L or I and X77 is E or D) comprising SEQ ID NO: 238.

[0087] The first and / or second peptides may further comprise linker sequences, in situations wherein the polynucleotide for expressing a chimeric polyprotein comprises sequence encoding linkers.

[0088] Another object of the present invention is a method for producing a plurality of protein entities, and includes a first step of producing a chimeric polyprotein using the method as described herein, and a second step of inducing the cleavage of the disulfide bridge between the first and second protein entities, thereby obtaining a plurality of protein entities. According to the present invention, the plurality of protein entitiescorresponds to a mixture of the first and second protein entities and is obtained by cleavage of the disulfide bridge between both protein entities.

[0089] According to the present invention, the first protein entity comprises a C-terminal sequence consisting of a first peptide as described hereinabove, and the second protein entity comprises a N-terminal sequence consisting of a second peptide as described hereinabove.

[0090] In some embodiments, the chimeric polyprotein is soluble, and the method of the present invention allows recovering two soluble protein entities.

[0091] In some embodiments, the chimeric polyprotein is anchored or attached to a surface (examples of surfaces include, but are not limited to, a membrane, a cell wall, a purification column, a nanoparticle, a liposome, an amorphous or crystalline material) and the method of the present invention allows recovering a soluble protein entity, while the second protein entity one stands anchored to the surface.

[0092] Cleavage of the disulfide bridge may be obtained by applying reductive conditions, such as, for example, addition of a reductive agent. Examples of reductive agents include, but are not limited to, beta-mercaptoethanol, dithiothreitol (DTT) and tris(2-carboxyethyl) phosphine (TCEP). These agents may for example be used at a concentration ranging from 1 and 10 mM.

[0093] The present invention further relates to a plurality of protein entities obtained, or susceptible to be obtained, by the method of the present invention. The protein entities of the plurality of protein entities may all be soluble. Some protein entities of the plurality of protein entities may alternatively be anchored or attached to a surface.

[0094] The present invention further relates to a chimeric polyprotein obtained, or susceptible to be obtained, by the method of the present invention.

[0095] A first example of chimeric polyprotein that may be obtained by the method of the present invention is a tagged protein, wherein a protein is linked by a disulfide bridge to a tag. Examples of tags are well-known to the skilled artisan, and include, without limitation, 6His tag (HHHHHH, SEQ ID NO: 242), MycHis tag(ALEQKLISEEDLNSAVDHHHHHH, SEQ ID NO: 243), Strep tag (WSHPQFEK, SEQ ID NO: 288), and Flag tag (DYKDDDDK, SEQ ID NO: 289).

[0096] A second example of chimeric polyprotein that may be obtained by the method of the present invention is an anchored protein, i.e., a protein susceptible to anchor into a biologic membrane or cell wall, wherein a protein is linked by a disulfide bridge to an anchoring domain. Examples of anchoring domains are well-known to the skilled artisan, and include, without limitation, a LPXTG motif (SEQ ID NO: 244, wherein X may be any amino acid, such as, for example, N) recognized by sortase for anchoring polyprotein into the peptidoglycan of Gram-positive bacteria.

[0097] This invention further relates to a polynucleotide encoding a polypeptide comprising or consisting of SEQ ID NO: 1 or SEQ ID NO: 300 wherein the N residue at position 150 is replaced by an A residue, or a sequence with at least 70% sequence identity with SEQ ID NO: 1 wherein the N residue at position 150 is replaced by an A residue.

[0098] This invention further relates to a polynucleotide encoding a polypeptide comprising or consisting of SEQ ID NO: 1 or SEQ ID NO: 300 wherein the C-terminal cysteine residue is replaced by an A residue, or a sequence with at least 70% sequence identity with SEQ ID NO: 1 wherein the C-terminal cysteine residue is replaced by an A residue.

[0099] This invention further relates to a polynucleotide encoding a polypeptide comprising or consisting of SEQ ID NO: 1 or SEQ ID NO: 300 wherein the N residue at position 150 is replaced by an A residue and wherein the C-terminal cysteine residue is replaced by an A residue, or a sequence with at least 70% sequence identity with SEQ ID NO: 1 wherein the N residue at position 150 is replaced by an A residue and wherein the C-terminal cysteine residue is replaced by an A residue.

[0100] All disclosure regarding the content of a polynucleotide encoding a polypeptide of SEQ ID NO: 1 or SEQ ID NO: 300 applies mutatis mutandis to these mutants.

[0101] The present invention further relates to a polynucleotide for expressing a chimeric protein, wherein said polynucleotide comprises or consists of, from 5’ to 3’:optionally a sequence encoding a signal peptide,a sequence encoding a first protein entity,a polynucleotide as disclosed herein and encoding one of the mutants as described hereinabove, anda sequence encoding a second protein entity,wherein the first and second proteins may be the same or different.

[0102] The present invention further relates to a method for producing a chimeric protein, comprising expressing (including transcribing and translating) a polynucleotide encoding one of the mutants as described hereinabove.

[0103] The present invention further relates to a chimeric polyprotein obtained or susceptible to be obtained by the method of the invention, wherein said chimeric polyprotein comprises a first protein entity and a second protein entity, wherein said first and second protein entities are linked by a disulfide bridge under oxidative conditions and may be dissociated under reductive conditions.

[0104] According to particular embodiments, the first and second protein entities are a protein of interest and a purification tag.

[0105] According to particular embodiments, the first and second protein entities are a protein of interest and an anchoring domain.

[0106] According to particular embodiments, the chimeric polyprotein is characterized in that:- the first protein entity comprises a first cysteine forming the disulfide bridge, and the second protein entity comprises a second cysteine forming the disulfide bridge;- the first protein entity comprises or consists of sequence PC as the first cysteine;- the second protein entity comprises or consists of a sequence XevGPC (SEQ ID NO: 238, wherein X67 is T, I or V) as the second cysteine.BRIEF DESCRIPTION OF THE DRAWINGS

[0107] Figure 1 is a schematic representation of the reporter precursor protein designed for studying Psy Fha BIL activity comprising a BIL domain, with the DsbA signal peptide (SP) that will be cleaved off upon secretion, the maltose binding protein (MBP) followed by twelve Psy Fha residues flanking the BIL domain at the N-junction, and ten Psy Fha residues flanking the BIL at the C-junction followed by a MycHis tag. BIL numbering is from 1 to 148, residues before BIL are numbered in backward direction with negative sign (-1, -2...) whereas residues after BIL are numbered with positive signs (+1, +2, etc.).

[0108] Figure 2 is a combination of pictures and graphs showing a characterization of protein Psy Fha BIL, using protein analysis by electrophoresis and by mass spectrometry. Precursors were expressed in the periplasm of E. coli and in vivo processed proteins were affinity-purified using a C-terminal His tag. Proteins were treated or not with DTT (1 mM) prior to denaturation and samples were analyzed either by SDS-PAGE (left) and western blot with His tag detection (middle) or by mass spectrometry after trypsin digestion (right). The results are shown for WT precursor in Figure 2A, for a mutant N148A in Figure 2B, for a mutant N148A / C+4A in Figure 2C, and for a mutant C+4A in Figure 2D.

[0109] Figure 3 is a picture of the result of an SDS PAGE analysis for the protein N148A / C+4A mutant, immobilized on a Ni-NTA resin after incubation with various DTT concentrations (0,1 mM, 0,5 mM and 1 mM) after various time, to detect DTT-triggered self cleavage of the protein.

[0110] Figure 4 is a combination of a graph and of a picture, showing the characterization of purified protein MBP-CC, treated or not with DTT (1 mM) during 60 min, using protein analysis SDS-PAGE (Figure 4A) and by mass spectrometry after trypsin digestion (Figure 4B).

[0111] Figure 5 is a combination of a graph and of a picture, showing the characterization of purified protein MBP-PCC, treated or not with DTT (1 mM) during 60 min, using protein analysis SDS-PAGE (Figure 5A) and by mass spectrometry after trypsin digestion (Figure 5B)

[0112] Figure 6 is a schematic representation of the processing of a chimeric protein obtained by expression of a polynucleotide for expressing a chimeric polyprotein or a plurality of protein entities according to the present invention. In an oxidative environment (such as, for example, the periplasm or the extracellular medium), an initial disulfide bridge is formed between two cysteines (Cys-1 and Cysl on Figure 6, corresponding to cysteines at position 2 and 3 in SEQ ID NO: 1 or 3). Disulfide isomerization with Cysteine at position +4 (Cys+4) releases cysteine at position 1 (Cysl) and triggers acyl transfer reactions leading to BIL self-cleavage (cleavage of first peptide bond between Cys-1 and Cysl, and second peptide bond between N148 and T+l, corresponding to residue N150 and T151 in SEQ ID NO: 1 or 3). In Figure 6, the BIL domain is in white, and two proteins linked by a disulfide bridge are in grey.

[0113] Figure 7 is a combination of (i) a schematic representation of the method for obtaining a polyprotein of the invention, comprising an alpha-amylase protein and an anchoring domain and (ii) a graph showing the alpha - amylase activity measured before or after release of the amylase through the addition of DTT, compared to a construct wherein the sequence of the amylase protein is directly fused to an anchoring domain (without using the polynucleotide of the present invention).EXAMPLES

[0114] The present invention is further illustrated by the following examples.Example 1: Activity and general mechanism of the poly peptide of the inventionConstructs

[0115] The following constructs were used in this example:WT construct (as described in Figure 1) corresponds to a nucleic acid sequence SEQ ID NO: 262, encoding the amino acid sequence SEQ ID NO 263. This construct consists of (i) a polypeptide of the present invention having a sequence SEQ ID NO: 264 (SEQ ID NO: 264 comprises the sequence SEQ ID NO: 3 asdescribed herein. In SEQ ID NO: 264, the domain consisting of amino acids at position 13 to position 161 are referred to as “BIL domain”) fused with sequences of (ii) a N-terminal maltose binding protein (MBP) - SEQ ID NO: 265 coupled with a signal peptide (SEQ ID NO: 236) and (iii) a C-terminal MycHis-Tag (SEQ ID NO: 243), that may be used for affinity purification and western blot analysis;N148A construct: as compared to the WT construct, the N residue at position 148 in the BIL domain is replaced by an alanine residue (SEQ ID NO: 266);- N148A / C+4A construct: as compared to the WT construct, the N residue at position 148 in the BIL domain is replaced by an alanine residue and the cysteine residue at position +4 after the C-terminal end of the BIL domain (corresponding to the cysteine at position 154 in SEQ ID NO: 3) is replaced by an alanine residue (SEQ ID NO: 267);C+4A construct: as compared to the WT construct, the cysteine residue +4 after the end of the BIL domain (corresponding to the cysteine at position 154 in SEQ ID NO: 3) is replaced by an alanine residue (SEQ ID NO: 268)MBP-CC construct (SEQ ID NO: 269, encoded by the nucleic acid sequence SEQ ID NO 270): in SEQ ID NO: 263, the sequence GAGRKFAMDGP (SEQ ID NO: 271) is deleted. Thus, the C terminal lysine of MBP protein is directly fused to the twin cysteine motif at position 2-3 in SEQ ID NO: 3;MBP-PCC construct (SEQ ID NO: 272, encoded by the acid nucleic sequence SEQ ID NO: 273), in SEQ ID NO: 263, the sequence GAGRKFAMDG (SEQ ID NO: 274) is deleted. Thus, the C terminal lysine of MBP is directly fused to the proline P at position 1 in SEQ ID NO: 3.SEQ ID NO: 262 TCATGAAAAAGATTTGGCTGGCGCTGGCTGGTTTAGTTTTAGCGTTTAGCGC ATCGGCGGCCATGGGTAAAATCGAAGAAGGTAAACTGGTAATCTGGATTAA CGGCGATAAAGGCTATAACGGTCTCGCTGAAGTCGGTAAGAAATTCGAGAA AGATACCGGAATTAAAGTCACCGTTGAGCATCCGGATAAACTGGAAGAGAA ATTCCCTCAGGTTGCGGCAACTGGCGATGGCCCTGACATTATCTTCTGGGCGCACGACCGTTTTGGTGGCTACGCTCAATCTGGCCTGTTGGCTGAAATCACCC CGGACAAAGCGTTCCAGGACAAGCTGTATCCGTTTACCTGGGATGCCGTAC GTTACAACGGCAAGCTGATTGCTTACCCGATCGCTGTTGAAGCGTTATCGCT GATTTATAACAAAGATCTGCTGCCGAACCCACCAAAAACCTGGGAAGAGAT CCCGGCGCTGGATAAAGAACTGAAAGCGAAAGGTAAGAGCGCGCTGATGTT CAACCTGCAAGAACCGTACTTCACCTGGCCGCTGATTGCTGCTGACGGGGG TTATGCGTTCAAGTATGAAAACGGCAAGTACGATATTAAAGACGTGGGCGT GGATAACGCTGGCGCGAAAGCGGGTCTGACCTTCCTGGTTGACCTGATTAA AAACAAACACATGAATGCAGACACCGATTACTCCATCGCAGAAGCTGCCTT TAATAAAGGCGAAACAGCGATGACCATCAACGGCCCGTGGGCATGGTCCAA CATCGACACCAGCAAAGTGAATTATGGTGTAACGGTACTGCCGACCTTCAA GGGTCAACCATCTAAACCGTTCGTTGGCGTGCTGAGCGCCGGTATTAACGCC GCCAGTCCGAACAAAGAGCTGGCGAAAGAGTTCCTCGAAAACTATCTGCTG ACTGACGAAGGTCTGGAAGCGGTCAATAAAGACAAACCGCTGGGTGCAGTG GCGCTGAAGTCTTACGAGGAAGAGTTGGCGAAAGACCCACGTATTGCCGCC ACGATGGAAAACGCCCAGAAAGGTGAAATCATGCCAAACATCCCGCAGAT GTCCGCTTTCTGGTATGCCGTACGTACTGCGGTGATCAACGCCGCCAGCGGT CGTCAGACTGTCGATGAAGCCCTGAAAGACGCGCAGACTCGTATCACCAAG GGTGGTGCTGGCCGTAAATTCGCTATGGATGGTCCGTGCTGCTTCGCCGCTG GTACGATGGTCTCAACGCCGGATGGTGAACGTGCTATTGATACGCTGAAAG TTGGTGATATTGTCTGGTCAAAACCGGAAGGCGGTGGCAAACCGTTTGCGG CCGCAATTCTGGCCACCCATATCCGTACGGATCAGCCGATCTATCGCCTGAA ACTGAAAGGTAAACAGGAAAACGGCCAAGCCGAAGACGAATCGCTGCTGG TGACCCCGGGTCATCCGTTTTACGTTCCGGCACAGCACGGCTTCGTGCCGGT TATTGATCTGAAACCGGGTGACCGTCTGCAAAGTCTGGCTGATGGCGCGTCC GAAAACACCAGCTCTGAAGTCGAATCTCTGGAACTGTATCTGCCGGTTGGT AAAACCTACAATCTGACGGTCGACGTGGGCCACACGTTCTATGTTGGTAAA CTGAAAACCTGGGTCCATAATACGGGTCCGTGTGAACTGCCGGAAGGCTAT GCTCTAGA SEQ ID NO: 263 MKKIWLALAGLVLAFASAAMGKIEEGKLVIWINGDKGYNGLAEVGKKFEKDT GIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLLAEITPDKAF QDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKEL KAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKA GLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYG VTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDK PLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINA ASGRQTVDEALKDAQTRITKGGAGRKFAMDGPCCFAAGTMVSTPDGERAIDTL KVGDIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGKQENGQAEDESLLVT PGHPFYVPAQHGFVPVIDLKPGDRLQSLADGASENTSSEVESLELYLPVGKTYN LTVDVGHTFYVGKLKTWVHNTGPCELPEGYALEQKLISEEDLNSAVDHHHHH H SEQ ID NO: 265 AMGKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQV AATGDGPDIIFWAHDRFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALD KELKAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGA KAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVN YGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNK DKPLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVI NAASGRQTVDEALKDAQTRITKG SEQ ID NO: 266 MKKIWLALAGLVLAFASAAMGKIEEGKLVIWINGDKGYNGLAEVGKKFEKDT GIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLLAEITPDKAF QDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKEL KAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKA GLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYG VTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDK PLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINA ASGRQTVDEALKDAQTRITKGGAGRKFAMDGPCCFAAGTMVSTPDGERAIDTL KVGDIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGKQENGQAEDESLLVT PGHPFYVPAQHGFVPVIDLKPGDRLQSLADGASENTSSEVESLELYLPVGKTYN LTVDVGHTFYVGKLKTWVHATGPCELPEGYALEQKLISEEDLNSAVDHHHHH H SEQ ID NO: 267 MKKIWLALAGLVLAFASAAMGKIEEGKLVIWINGDKGYNGLAEVGKKFEKDT GIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLLAEITPDKAF QDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKEL KAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKA GLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYG VTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDK PLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINA ASGRQTVDEALKDAQTRITKGGAGRKFAMDGPCCFAAGTMVSTPDGERAIDTL KVGDIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGKQENGQAEDESLLVT PGHPFYVPAQHGFVPVIDLKPGDRLQSLADGASENTSSEVESLELYLPVGKTYN LTVDVGHTFYVGKLKTWVHATGPAELPEGYALEQKLISEEDLNSAVDHHHHH H SEQ ID NO: 268 MKKIWLALAGLVLAFASAAMGKIEEGKLVIWINGDKGYNGLAEVGKKFEKDT GIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLLAEITPDKAF QDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKEL KAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKA GLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYG VTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDK PLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINA ASGRQTVDEALKDAQTRITKGGAGRKFAMDGPCCFAAGTMVSTPDGERAIDTL KVGDIVWSKPEGGGKPFAAAILATHIRTDQPIYRLKLKGKQENGQAEDESLLVT PGHPFYVPAQHGFVPVIDLKPGDRLQSLADGASENTSSEVESLELYLPVGKTYN LTVDVGHTFYVGKLKTWVHNTGPAELPEGYALEQKLISEEDLNSAVDHHHHH HSEQ ID NO: 269 MKKIWLALAGLVLAFASAAMGKIEEGKLVIWINGDKGYNGLAEVGKKFEKDT GIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLLAEITPDKAF QDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKEL KAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKA GLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYG VTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDK PLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINA ASGRQTVDEALKDAQTRITKCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPEG GGKPFAAAILATHIRTDQPIYRLKLKGKQENGQAEDESLLVTPGHPFYVPAQHG FVPVIDLKPGDRLQSLADGASENTSSEVESLELYLPVGKTYNLTVDVGHTFYVG KLKTWVHNTGPCELPEGYALEQKLISEEDLNSAVDHHHHHH SEQ ID NO: 270 TCATGAAAAAGATTTGGCTGGCGCTGGCTGGTTTAGTTTTAGCGTTTAGCGC ATCGGCGGCCATGGGTAAAATCGAAGAAGGTAAACTGGTAATCTGGATTAA CGGCGATAAAGGCTATAACGGTCTCGCTGAAGTCGGTAAGAAATTCGAGAA AGATACCGGAATTAAAGTCACCGTTGAGCATCCGGATAAACTGGAAGAGAA ATTCCCTCAGGTTGCGGCAACTGGCGATGGCCCTGACATTATCTTCTGGGCG CACGACCGTTTTGGTGGCTACGCTCAATCTGGCCTGTTGGCTGAAATCACCC CGGACAAAGCGTTCCAGGACAAGCTGTATCCGTTTACCTGGGATGCCGTAC GTTACAACGGCAAGCTGATTGCTTACCCGATCGCTGTTGAAGCGTTATCGCT GATTTATAACAAAGATCTGCTGCCGAACCCACCAAAAACCTGGGAAGAGAT CCCGGCGCTGGATAAAGAACTGAAAGCGAAAGGTAAGAGCGCGCTGATGTT CAACCTGCAAGAACCGTACTTCACCTGGCCGCTGATTGCTGCTGACGGGGG TTATGCGTTCAAGTATGAAAACGGCAAGTACGATATTAAAGACGTGGGCGT GGATAACGCTGGCGCGAAAGCGGGTCTGACCTTCCTGGTTGACCTGATTAA AAACAAACACATGAATGCAGACACCGATTACTCCATCGCAGAAGCTGCCTT TAATAAAGGCGAAACAGCGATGACCATCAACGGCCCGTGGGCATGGTCCAA CATCGACACCAGCAAAGTGAATTATGGTGTAACGGTACTGCCGACCTTCAA GGGTCAACCATCTAAACCGTTCGTTGGCGTGCTGAGCGCCGGTATTAACGCC GCCAGTCCGAACAAAGAGCTGGCGAAAGAGTTCCTCGAAAACTATCTGCTG ACTGACGAAGGTCTGGAAGCGGTCAATAAAGACAAACCGCTGGGTGCAGTG GCGCTGAAGTCTTACGAGGAAGAGTTGGCGAAAGACCCACGTATTGCCGCC ACGATGGAAAACGCCCAGAAAGGTGAAATCATGCCAAACATCCCGCAGAT GTCCGCTTTCTGGTATGCCGTACGTACTGCGGTGATCAACGCCGCCAGCGGT CGTCAGACTGTCGATGAAGCCCTGAAAGACGCGCAGACTCGTATCACCAAG TGCTGCTTCGCCGCTGGTACGATGGTCTCAACGCCGGATGGTGAACGTGCTA TTGATACGCTGAAAGTTGGTGATATTGTCTGGTCAAAACCGGAAGGCGGTG GCAAACCGTTTGCGGCCGCAATTCTGGCCACCCATATCCGTACGGATCAGCC GATCTATCGCCTGAAACTGAAAGGTAAACAGGAAAACGGCCAAGCCGAAG ACGAATCGCTGCTGGTGACCCCGGGTCATCCGTTTTACGTTCCGGCACAGCA CGGCTTCGTGCCGGTTATTGATCTGAAACCGGGTGACCGTCTGCAAAGTCTG GCTGATGGCGCGTCCGAAAACACCAGCTCTGAAGTCGAATCTCTGGAACTG TATCTGCCGGTTGGTAAAACCTACAATCTGACGGTCGACGTGGGCCACACG TTCTATGTTGGTAAACTGAAAACCTGGGTCCATGCTACGGGTCCGGCTGAAC TGCCGGAAGGCTATGCTCTAGASEQ ID NO: 272 MKKIWLALAGLVLAFASAAMGKIEEGKLVIWINGDKGYNGLAEVGKKFEKDT GIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLLAEITPDKAF QDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKEL KAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKA GLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYG VTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDK PLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINA ASGRQTVDEALKDAQTRITKPCCFAAGTMVSTPDGERAIDTLKVGDIVWSKPE GGGKPFAAAILATHIRTDQPIYRLKLKGKQENGQAEDESLLVTPGHPFYVPAQH GFVPVIDLKPGDRLQSLADGASENTSSEVESLELYLPVGKTYNLTVDVGHTFYV GKLKTWVHNTGPCELPEGYALEQKLISEEDLNSAVDHHHHHH SEQ ID NO: 273:TCATGAAAAAGATTTGGCTGGCGCTGGCTGGTTTAGTTTTAGCGTTTAGCGC ATCGGCGGCCATGGGTAAAATCGAAGAAGGTAAACTGGTAATCTGGATTAA CGGCGATAAAGGCTATAACGGTCTCGCTGAAGTCGGTAAGAAATTCGAGAA AGATACCGGAATTAAAGTCACCGTTGAGCATCCGGATAAACTGGAAGAGAA ATTCCCTCAGGTTGCGGCAACTGGCGATGGCCCTGACATTATCTTCTGGGCG CACGACCGTTTTGGTGGCTACGCTCAATCTGGCCTGTTGGCTGAAATCACCC CGGACAAAGCGTTCCAGGACAAGCTGTATCCGTTTACCTGGGATGCCGTAC GTTACAACGGCAAGCTGATTGCTTACCCGATCGCTGTTGAAGCGTTATCGCT GATTTATAACAAAGATCTGCTGCCGAACCCACCAAAAACCTGGGAAGAGAT CCCGGCGCTGGATAAAGAACTGAAAGCGAAAGGTAAGAGCGCGCTGATGTT CAACCTGCAAGAACCGTACTTCACCTGGCCGCTGATTGCTGCTGACGGGGG TTATGCGTTCAAGTATGAAAACGGCAAGTACGATATTAAAGACGTGGGCGT GGATAACGCTGGCGCGAAAGCGGGTCTGACCTTCCTGGTTGACCTGATTAA AAACAAACACATGAATGCAGACACCGATTACTCCATCGCAGAAGCTGCCTT TAATAAAGGCGAAACAGCGATGACCATCAACGGCCCGTGGGCATGGTCCAA CATCGACACCAGCAAAGTGAATTATGGTGTAACGGTACTGCCGACCTTCAA GGGTCAACCATCTAAACCGTTCGTTGGCGTGCTGAGCGCCGGTATTAACGCC GCCAGTCCGAACAAAGAGCTGGCGAAAGAGTTCCTCGAAAACTATCTGCTG ACTGACGAAGGTCTGGAAGCGGTCAATAAAGACAAACCGCTGGGTGCAGTG GCGCTGAAGTCTTACGAGGAAGAGTTGGCGAAAGACCCACGTATTGCCGCC ACGATGGAAAACGCCCAGAAAGGTGAAATCATGCCAAACATCCCGCAGAT GTCCGCTTTCTGGTATGCCGTACGTACTGCGGTGATCAACGCCGCCAGCGGT CGTCAGACTGTCGATGAAGCCCTGAAAGACGCGCAGACTCGTATCACCAAG TGCTGCTTCGCCGCTGGTACGATGGTCTCAACGCCGGATGGTGAACGTGCTA TTGATACGCTGAAAGTTGGTGATATTGTCTGGTCAAAACCGGAAGGCGGTG GCAAACCGTTTGCGGCCGCAATTCTGGCCACCCATATCCGTACGGATCAGCC GATCTATCGCCTGAAACTGAAAGGTAAACAGGAAAACGGCCAAGCCGAAG ACGAATCGCTGCTGGTGACCCCGGGTCATCCGTTTTACGTTCCGGCACAGCA CGGCTTCGTGCCGGTTATTGATCTGAAACCGGGTGACCGTCTGCAAAGTCTG GCTGATGGCGCGTCCGAAAACACCAGCTCTGAAGTCGAATCTCTGGAACTG TATCTGCCGGTTGGTAAAACCTACAATCTGACGGTCGACGTGGGCCACACG TTCTATGTTGGTAAACTGAAAACCTGGGTCCATGCTACGGGTCCGGCTGAAC TGCCGGAAGGCTATGCTCTAGAMaterials and MethodsGenetic constructs for the expression of WT or N148A or N148A / C+4A or MBP-CC or MBP-PCC constructs

[0116] A synthetic gene encoding the WT was inserted between Ncol and Xbal restriction sites of the pBAD-Tet plasmid (Deschuyteneer G, Garcia S, Michiels B, Baudoux B, Degand H, Morsomme P, Soumillion, ACS Chem Biol. 2010 Jul 16;5(7):691-700. doi: 10.1021 / cbl00072u), in frame with the sequence encoding the MycHis6 tag. The plasmid also carries a tetracycline resistance cassette. Synthetic genes encoding the MBP-PCC and MBP-CC precursors were cloned similarly. Expression plasmids were introduced into E. coli TOPIO strain (Invitrogen) by electroporation. Point mutations were generated using appropriate pairs of complementary oligonucleotides (N148Af / N148Ar; C+4Afl / C+4Arl; C+4Af2 / C+4Ar2) via site directed mutagenesis with Pfu Ultra DNA polymerase (Agilent Technologies). All plasmids were checked by restriction analysis and by DNA sequencing.

[0117] The oligonucleotides are defined by the following sequences:SEQ NAME SEQUENCE ID NO278 N148Af AACTGAAAACCTGGGTCCATGCTACGGGTCCGTGTGAACTG C279 N148Ar GCAGTTCACACGGACCCGTAGCATGGACCCAGGTTTTCAGTT 280 C+4Afl GGGTCCATAATACGGGTCCGGCTGAACTGCCGGAAGGCTAT G281 C+4 Ari CATAGCCTTCCGGCAGTTCAGCCGGACCCGTATTATGGACCC 282 C+4Af2 GGGTCCATGCTACGGGTCCGGCTGAACTGCCGGAAGGCTAT G283 C+4Ar2 CATAGCCTTCCGGCAGTTCAGCCGGACCCGTAGCATGGACCCwherein C+4Afl / C+4Arl and C+4f2 / C+4r2 are pairs of oligonucleotides (forward / reverse) for mutating C+4 in the wild type and N148A variant respectively.Protein expression and purification

[0118] Cells harboring the expression vector pBAD-Tet plasmid were grown in Luria-Bertani broth medium supplemented with tetracycline (7.5 pg / mL) at 37°C. Overexpression was induced at OD600 = 0.6 with L-(+)-arabinose to a final concentration of 0.2% (w / v) and grown at 30 °C for 3h.

[0119] The proteins are expressed in the oxidative environment of the E. coli periplasm, where disulfide bridges can form. The signal peptide from the DsbA protein of E. coli is used to promote the secretion of the protein into the periplasm, thus preventing folding and splicing in the cytoplasm.

[0120] The protein purification was done on periplasmic extract of an E.coli culture prepared as described in Deschuyteneer G, Garcia S, Michiels B, Baudoux B, Degand H, Morsomme P, Soumillion, ACS Chem Biol. 2010 Jul 16;5(7):691-700. doi: 10.1021 / cbl00072u.

[0121] Overexpressed fusion proteins were purified through their C-terminal His-tags using His*Mag™ Agarose Beads (Novagen) by affinity chromatography using Ni++-NTA resin.Analysis of the purified proteins by electrophoresis

[0122] To study the effect of disulfide bonds reduction on protein, proteins were incubated with 1 mM Dithiothreitol (DTT) in reducing condition, or, without DTT ImM (non-reducing condition) prior to denaturation and samples were analyzed either by SDS-PAGE and western blot with His tag detection, or by mass spectrometry after trypsin digestion.SDS -PAGE:

[0123] Proteins were analyzed by 12.5% SDS-PAGE using Laemmli method with reveal by Coomassie Blue coloration.

[0124] Typically, between 0.1 and 1 pg of protein were heat-denatured in 1% SDS with or without DTT (ImM) before loading on the gel. Samples were incubated with DTT for 1 hour at room temperature prior to protein denaturation.Western Blot:

[0125] Proteins were analyzed by Western Blot. Western Blot analysis was performed after transferring onto a nitrocellulose membrane. The HisProbe™-HRP (Thermo Scientific) was used for the detection of recombinant poly-histidine-tagged fusion proteins.

[0126] Images of immunoblots were taken with a CCD-based device (Kodak image station 4000R) and a standard ECL reagent.Mass spectrometry: MS and MS / MS analysis:

[0127] Proteins were analyzed by mass spectrometry after trypsin digestion. The trypsin digestion was done on Coomassie Blue stained proteins bands excised from the gel or on purified protein solution as following described.

[0128] 2 pl of a solution containing 10 mg / ml of alpha-cyano-4-hydroxycinnamic acid (alpha-cyano MALDI matrix) in 50% (v / v) acetonitrile and 0.1% (v / v) TFA were mixed with 2 pl of each concentrated peptide solution. 0.5 pl of this solution were put down on Applied Biosystems MALDI plate Opti-TOF 384 Well Insert (Foster City, USA).

[0129] MS and MS / MS spectra were acquired using an Applied Bisosystems 4800 MALDI TOF / TOF Analyzer spectrometer using a 200 Hz solid state laser operating at a wavelength of 355 nm.

[0130] MS spectra were obtained using a laser intensity of 3500 and 2000 laser shots by spot in a range of m / z between 600 and 6000. MS / MS spectra were obtained by selecting precursor ions and using a laser intensity of 3800 and 2000 laser shots by precursor. Peptide fragmentation was performed at a collision energy of 1 kV with collision gas air at a pressure of about IxlO6torr.

[0131] Data were collected with the Applied Biosystems 4000 Series Explorer software.Trypsin digestion in gel

[0132] Coomassie Blue stained proteins bands were manually excised from the gel of SDS-PAGE.

[0133] Gels plugs were transferred with 200 pl of HPLC water to 0.5 polypropylene Protein LoBind Eppendorf tubes. Water was removed and gel plugs were incubated under shaking at 20°C for 5 min with 200 pl of a solution containing 50 mM ammonium carbonate (pH 8.0) in 50% acetonitrile.

[0134] This solution was removed, and gel plugs were incubated under shaking at 20°C for 5 min with 200 pl of 100% acetonitrile.

[0135] Acetonitrile was removed and gels plugs were dried under vacuum with a Savant Speed Vac Concentrator.

[0136] 20 pl of a digestion buffer containing 50 mM ammonium carbonate (pH 8.0) and 0.5 pg of Promega Sequencing Grade Modified Trypsin (Madison, USA) were added to each tube to rehydrate gel plugs.

[0137] Proteolysis was allowed to continue overnight at 37°C and stopped by adding 10 ml of 1% (v / v) TFA.

[0138] Supernatants of each tube were recovered and transferred in new 0.5 polypropylene Protein LoBind Eppendorf tubes.

[0139] 50 pl of peptide extraction solution containing 50% (v / v) acetonitrile and 0.1% (v / v) TFA were added to the gel plugs. After incubation for 5 min, extracts were combined with the first. A second extraction with 50 pl of 100% acetonitrile was done for 5 min. Extracts were combined with the first and dried under vacuum with Savant Speed Vac Concentrator.

[0140] Peptides were solubilized in 20 pl of 0.1% TFA, desalted and concentrated with Millipore Cl 8 ZipTip (Billerica, USA) according to the manufacturer’s protocol.Trypsin digestion in solution

[0141] Before trypsin digestion, a chloroform / methanol precipitation was performed according to Wessel, D. and Flugge, U.I. (1984). Methanol and chloroform were stored at -20°C for 1 h before the procedure. Water was stored on ice for 1 h before the procedure. 20 pl of the sample was put in a 1.5 polypropylene Protein LoBind Eppendorf tube and cold water was added to have a final volume of 100 pl. 1 volume of chloroform and 4 volumes of methanol were added to the sample and mixed on a vortex for 30 s. Solution was placed on ice for 15 min. 3 volumes of cold water were added and mixed on a vortex for 30 s to obtain an emulsion. Solution was placed on ice for 5 min and centrifuged at 4°C at 20,000 x g for 15 min in a Model 5417C Eppendorf Centrifuge (Hamburg, Germany) to separate phases. Precipitated proteins were in the interphase between the lower chloroformic phase and the upper methanolic phase. The upper phase was removed and 800 pl of cold methanol were added to dilute the chloroformic phase. The solution was centrifuged at 4°C at 20,000 x g for 15 min in a Model 5417C Eppendorf Centrifuge to pellet precipitated proteins. Supernatant was removed and the pellet was dried with a Savant Speed Vac Concentrator (Ramsey, USA). Pellets were stored at -20°C. Proteins were resuspended with 20 pl of a digestion buffer containing 50 mM ammonium carbonate (pH 8.0) and 0.5 pg of Promega Sequencing Grade Modified Trypsin (Madison, USA). Proteolysis was allowed to continue overnight at 37°C and stopped by adding 10 ml of 1% (v / v) TFA. Peptide solution was dried under vacuum with Savant SpeedVac SC100 (New York, USA) Concentrator. Peptides were solubilised in 20 pl of 0.1% TFA, desalted and concentrated with Millipore Cl 8 ZipTip (Billerica, USA) according to the manufacturer’s protocol.ResultsActivity ofWT, N148A, N 148A / C+4A and C+4A proteinsActivity of WT protein

[0142] Figure 2A shows the purified WT protein, with or without DTT treatment, in SDS PAGE and in Western blot; and in mass spectrometry after trypsin digestion.

[0143] The in vivo processed Psy Pha BIL WT protein was purified by affinitychromatography and starting from a periplasmic extract of an E. coli culture.

[0144] On the SDS-Page picture of Figure 2A, non-reducing SDS-PAGE analysis revealed a protein of around 50 kDa suggesting that the BIL domain was excised. This excision was confirmed by mass spectrometry analysis, since only peptides from the MBP and the C-terminal tag were detected. On the Western Blot image of Figure 2A, the western blot analysis confirmed that the eluted protein was still linked with the C-terminal MycHis-Tag. The result shows that the western blot signal was lost when the protein was treated with DTT prior to electrophoresis and the Coomassie- stained protein appeared around 40 kDa, indicating that a small C-terminal fragment was lost upon reduction. Results of Mass spectrometry shown in Figure 2A further demonstrated that the BIL-flanking polypeptides were connected through a disulfide bridge rather than a peptide bond. Indeed, a tryptic peptide at m / z 2372.1 is observed without DTT treatment, while this peak at m / z 2372.1 disappeared when the sample was treated with DTT. The tryptic peptide at m / z 2372.1 corresponds to the sum of the peptides containing residues -7 to -1 and +1 to +15 minus the 2 daltons resulting from the disulfide bond formation between Cys-1 and Cys+4. MS / MS fragmentation analysis confirmed the nature of the molecule as several fragments could be assigned to this bridged peptide. Finally, peptides that would have been expected in case of a splicing event, i.e., with a peptide bond between residues -1 and +1, could not be detected.

[0145] Thus, the results of Figure 2A indicate that the BIL domain is not a splicing domain but a self-cleaving domain catalyzing both N-terminal and C-terminal cleavages. Activity of mutant N 148 A, mutant N148A / C+4A and mutant C+4A protein

[0146] To investigate the possible interplay between BIL cleavage and disulfide bond formation and get better insight into the mechanism of reaction, purified N148A protein (Figure 2B), N148A / C+4A (Figure 2C), and C+4A protein were prepared and analyzed by SDS-PAGE, western blot and mass spectrometry. Protein preparation consists of expression in vivo and extraction from E.coli culture followed by a step of purification as previous mentioned.

[0147] For the N148A construct, the conserved asparagine residue at the C-terminalposition (residue 148) of the BIL domain is mutated.

[0148] On the SDS-Page picture of Figure 2B, non-reducing SDS-PAGE analysis revealed a protein of around 62 kDa for the purified N148A protein. On the Western Blot image of Figure 2B, the western blot analysis confirmed the C-terminal MycHis-Tag was also detected by western blot and, when treated with DTT, the protein is cleaved into two fragments around 40 and 23 kDa, suggesting cleavage at the N-junction of BIL domain. The smaller fragment contains the MycHis-Tag although the intensity of the western blot signal was strongly reduced. A supplementary mass analysis of the tryptic peptides confirmed that the 40 kDa fragment is the MBP domain while the 23 kDa fragment is the BIL domain connected to the C-terminal MycHis-Tag. Some high molecular weight bands (>80 kDa) observed under non-reducing conditions (without DDT) also disappeared upon DTT treatment, indicating that these were crosslinked forms through disulfide bridges.

[0149] Results of mass spectrometry analysis of the tryptic peptides of Figure 2B confirmed the cleavage of the N-junction since the peptide with the N-terminal Cys+1 from the BIL domain was observed (m / z 1541.5). Furthermore, as previously mentioned, the sequence of this 23 kDa fragment observed upon DTT treatment of the mutant N148A purified protein was confirmed by supplementary MS / MS fragmentation analysis of tryptic peptides. In Mass spectrometry shown in Figure 2B, the peptide, observed at m / z 2966.0 only under non-reducing conditions (without DTT), corresponds to the bridged peptide with a disulfide bond between Cys-1 and Cys+4 and further demonstrates the absence of cleavage at the C-junction (bond between A148 and T+l).

[0150] In the N148A / C+154A construct, the Cys+4Ala mutation was introduced in the N148A construct. This N148A / C+4A construct was analyzed by SDS-PAGE, western blot and mass spectrometry after expression, extraction from E.coli culture and step of purification as previous mentioned. The results are shown in Figure 2C.

[0151] On the SDS-Page picture of Figure 2C, non-reducing condition (without DTT) SDS-PAGE analysis revealed that a molecular weight protein, with around 62 kDa could also be purified. Furthermore, DTT treatment prior to denaturation also generatedfragments of the same size and composition as for the single mutant N148A. The absence of Cys+4 in mutant N148A / C+4A however suggests that the disulfide bridge, if present, is between the two remaining adjacent cysteines flanking the N-junction. Indeed, as illustrated the mass spectrometry results of Figure 2C, in non-reducing conditions (without DTT), a tryptic peptide with m / z 2260.9 was identified as the peptide from residue -7 to residue 15 with a disulfide bridge between Cys-1 and Cysl and without cleavage of the peptide bond between them. No peptide with m / z 2278.9 was observed indicating the absence of peptide cleavage before reduction of the disulfide. In this construct N148A / C+4A, the nucleophilic Cys 1 of the BIL domain is trapped in a disulfide bridge with Cys-1 and reduction is triggering cleavage of the peptide bond at the N-junction. Similar results were observed with the single mutant C+4A mutant although cleavage at the C-junction was also observed as shown by the observation of an additional band around 58 kDa under non-reducing conditions (without DTT) in Figure 2D and by the detection of a tryptic peptide corresponding to the C-terminus of the BIL domain using mass spectrometry.An auto cleavable purification tag study

[0152] As a biotechnological application, triggering peptide cleavage by disulfide bond reduction may be useful for removing the purification tag during affinity chromatography. Thus, the auto cleavage efficiency when the mutant N148A / C+4A fusion protein is immobilized on an affinity column was evaluated.

[0153] In this study, the N148A / C+4A protein are immobilized on a Ni-NTA resin (0,1 mM, 0,5 mM and 1 mM) and are incubated with various DTT concentrations. The results are shown on the SDS-Page picture of Figure 3. In this Figure 3, lanes 1 and 2 shows the protein sample before and after complete cleavage upon DTT (1 mM) treatment in solution; lanes 3 to 14 correspond to elution fractions after various times (1,5 min, 3 min, 6 min and 12 min) of incubation with three different DTT concentrations (0,1 mM, 0,5 mM and 1 mM). Band intensities are compared to the intensity of the 40 kDa band in lane 2 that corresponds to full cleavage.

[0154] The results of Figure 3 show that, with 1 mM, more than 90% of the MBP was eluted in less than 15 minutes, while the BIL-MycHis-Tag remained bound to the resin.

[0155] Since 12 foreign residues from the Psy Fha protein are still fused at the C-terminus of MBP, two new constructs were built with the aim of reducing this C-terminal tail to a minimum.

[0156] The MBP was fused either directly to the twin cysteine motif (MBP-CC clone) or to the preceding proline -2 (MBP-PCC clone).

[0157] In the MBP-CC construct, the C-terminal lysine of MBP is directly fused to the twin cysteine motif at the N-junction with the BIL domain (N148A / C+4A construct).

[0158] In the MBP-PCC construct, the C-terminal lysine of MBP is directly fused to the proline (Pro-2) preceding the twin cysteine motif at the N-junction with the BIL domain (N148A / C+4A).

[0159] For both constructs, the gene was expressed in the periplasm of E. coli and the protein mutant was affinity-purified thanks to the C-terminal His tag. The protein mutant MBP-CC (Figure 4) and MBP-PCC (Figure 5) were treated or not with DTT (1 mM, 60 min) and samples were analyzed either by SDS-PAGE (Figure 4A and Figure 5A) or mass spectrometry after trypsin digestion (Figure 4B and Figure 5B).

[0160] For both constructs, the vicinal disulfide bridge was detected in the purified uncleaved proteins (Figure 4 and Figure 5). More precisely, the peptide containing the vicinal disulfide bridge was detected with a peak m / z = 1642.7 for mutant MBP-CC by mass spectrometry in Figure 4B, or, with a peak m / z = 2082.0 for mutant MBP-CC in Figure 5B. The sequence of the peptide containing the vicinal disulfide bridge was confirmed by MS / MS fragmentation analysis.

[0161] A poorly efficient DTT-triggered activation was observed for MBP-CC (Figure 4) while the MBP-PCC protein underwent almost complete cleavage (Figure 5). When MBP-PCC was immobilized on a Ni-NTA resin, incubation with 2 mM DTT during 100 minutes was necessary for eluting more than 90% of the cleaved MBP. Although slower than the initial construct, only two foreign residues (Pro-2 and Cys-1) remain fused to the MBP after cleaving this BIL-MycHis-Tag.Post -translational processing ofPsy Pha BIL

[0162] All of these results demonstrate that, for the WT protein, the BIL domain is selfexcised extracellularly while the large amino-terminal flanking protein ends up attached to a small C-terminal domain by a disulfide bridge (C-l to C+4). Furthermore, in the processing of the WT construct, a disulfide intermediate state (C- 1 to Cl) is formed before BIL excision as illustrated by Figure 6. These results demonstrate that the pivotal cysteine 1 involved in both disulfide bond and acyl-transfer chemistries affords the coordination of BIL self-cleavage with disulfide bridging of the flanking polypeptides. An intermediate state featuring a vicinal disulfide bridge at the N-scissile junction suggests a non-planar peptide bond and is giving insight into the mechanism of peptide destabilization by BIL domains.2: Use a chimeric in of the present invention in a method of

[0163] In this assay, the propriety of the polypeptide of the invention to create a disulfide bridge between two proteins in which one have an anchoring domain was tested (Figure 7). More precisely, the expression and possible cleavage of alpha-amylase protein, fused or not with a BIL domain was tested, after the creation of an external disulfide bridge connecting the alpha-amylase to the external peptidoglycan of a Gram-positive bacterium. Furthermore, the possibility of recovering the external part of chimeric protein with DTT was evaluated.Materials and MethodsGene construct and protein expression

[0164] Streptococcus thermophilus LMD-9 strain was genetically modified using its natural competence at the chromosomal locus of the prtS gene encoding a peptidoglycan-anchored protease. While keeping the PrtS N-terminal signal sequence for secretion and C-terminal sequence for sortase mediated peptidoglycan anchoring, the sequence of PrtS was replaced by the sequence of alpha-amylase (Amy) from Bacillus licheniformis fused or not to the sequence encoding BIL from Psy Fha. The construct corresponds to the nucleic acid sequence SEQ ID NO: 284 encoding a peptide of sequence SEQ ID NO: 285.This construct comprises a sequence encoding the signal peptide from PrtS, a gene alphaamylase, a Linker encoding Gly-Gly and a gene encoding the sortase anchoring motif for attaching the enzyme in the peptidoglycan of the bacteria.SEQ ID NO: 284 ATGAAAAAGAAAGAAACTTTCTCACTTCGGAAGTATAAAATTGGAACTGTGT CTGTTCTTTTGGGTGCAGTTTTTTTGTTTGCAGGTGCACCATCGGTAGCTGCA GATGAAGCAAATCTTAATGGGACGCTGATGCAGTATTTTGAATGGTACATGCC CAATGACGGCCAACATTGGAAGCGTTTGCAAAACGACTCGGCATATTTGGCT GAACACGGTATTACTGCCGTCTGGATTCCCCCGGCATATAAGGGAACGAGCC AAGCGGATGTGGGCTACGGTGCTTACGACCTTTATGATTTAGGGGAGTTTCAT CAAAAAGGGACGGTTCGGACAAAGTACGGCACAAAAGGAGAGCTGCAATCT GCGATCAAAAGTCTTCATTCCCGCGACATTAACGTTTACGGGGATGTGGTCAT CAACCACAAAGGCGGCGCTGATGCGACCGAAGATGTAACCGCGGTTGAAGT CGATCCCGCTGACCGCAACCGCGTAATTTCAGGAGAACACCTAATTAAAGCC TGGACACATTTTCATTTTCCGGGGCGCGGCAGCACATACAGCGATTTTAAATG GCATTGGTACCATTTTGACGGAACCGATTGGGACGAGTCCCGAAAGCTGAAC CGCATCTATAAGTTTCAAGGAAAGGCTTGGGATTGGGAAGTTTCCAATGAAA ACGGCAACTATGATTATTTGATGTATGCCGACATCGATTATGACCATCCTGATG TCGCAGCAGAAATTAAGAGATGGGGCACTTGGTATGCCAATGAACTGCAATT GGACGGTTTCCGTCTTGATGCTGTCAAACACATTAAATTTTCTTTTTTGCGGG ATTGGGTTAATCATGTCAGGGAAAAAACGGGGAAGGAAATGTTTACGGTAGC TGAATATTGGCAGAATGACTTGGGCGCGCTGGAAAACTATTTGAACAAAACA AATTTTAATCATTCAGTGTTTGACGTGCCGCTTCATTATCAGTTCCATGCTGCA TCGACACAGGGAGGCGGCTATGATATGAGGAAATTGCTGAACGGTACGGTCG TTTCCAAGCATCCGTTGAAATCGGTTACATTTGTCGATAACCATGATACACAG CCGGGGCAATCGCTTGAGTCGACTGTCCAAACATGGTTTAAGCCGCTTGCTT ACGCTTTTATTCTCACAAGGGAATCTGGATACCCTCAGGTTTTCTACGGGGAT ATGTACGGGACGAAAGGAGACTCCCAGCGCGAAATTCCTGCCTTGAAACAC AAAATTGAACCGATCTTAAAAGCGAGAAAACAGTATGCGTACGGAGCACAG CATGATTATTTCGACCACCATGACATTGTCGGCTGGACAAGGGAAGGCGACA GCTCGGTTGCAAATTCAGGTTTGGCGGCATTAATAACAGACGGACCCGGTGG GGCAAAGCGAATGTATGTCGGCCGGCAAAACGCCGGTGAGACATGGCATGA CATTACCGGAAACCGTTCGGAGCCGGTTGTCATCAATTCGGAAGGCTGGGGA GAGTTTCACGTAAACGGCGGGTCGGTTTCAATTTATGTTCAAAGAGGTGGTA AACAGGTGACTCAACTACCAAATACTGGAGAAAATGATACGAAATACTATCT TGTTCCTGGTGTCATTATTGGGCTAGGGACTCTGTTGGTAAGCATACGACGTC ACAAGGAAGAAGTATAA SEQ ID NO : 285 MKKKETFSLRKYKIGTVSVLLGAVFLFAGAPSVAADEANLNGTLMQYFEWYMP NDGQHWKREQNDSAYEAEHGITAVWIPPAYKGTSQADVGYGAYDEYDEGEFH QKGTVRTKYGTKGEEQSAIKSEHSRDINVYGDVVINHKGGADATEDVTAVEVDPADRNRVISGEHLIKAWTHFHFPGRGSTYSDFKWHWYHFDGTDWDESRKLNRI YKFQGKAWDWEVSNENGNYDYLMYADIDYDHPDVAAEIKRWGTWYANELQL DGFRLDAVKHIKFSFLRDWVNHVREKTGKEMFTVAEYWQNDLGALENYLNKT NFNHSVFDVPLHYQFHAASTQGGGYDMRKLLNGTVVSKHPLKSVTFVDNHDT QPGQSLESTVQTWFKPLAYAFILTRESGYPQVFYGDMYGTKGDSQREIPALKHKI EPILKARKQYAYGAQHDYFDHHDIVGWTREGDSSVANSGLAALITDGPGGAKR MYVGRQNAGETWHDITGNRSEPVVINSEGWGEFHVNGGSVSIYVQRGGKQVT QLPNTGENDTKYYLVPGVIIGLGTLLVSIRRHKEEV

[0165] The cultures were prepared for the amylase assay with the Sigma-Aldrich© Kit MAK009 as follows. Overnight cultures were reinoculated and incubated at 37 °C in anaerobic conditions. Once the cultures reached an OD600nm of ~0.7, the culture was centrifugated at 16000g for 5min.a) Preparation of the supernatant sample

[0166] 500pl of the clear supernatant was transferred into an Amicon Ultra-0.5 Centrifugal Filter Devices with 30 KDa cutoff. The Amicon was centrifugated for 15 minutes at 14,000 g. The residual volume in the Amicon was washed with 500pl of fresh PBS and the resulting concentrate was used as a sample for the assay (-20 pl) b) Preparation of the cells sample

[0167] The rest of the supernatant was discarded and the cells were washed twice with 500pl of PBS. The cells were finally resuspended in 200pl of PBS and, 20pl of the resuspension was used as a sample for the assay.c) DTT treatment

[0168] For the assay with the DTT treatment, the supernatant and the cells sample were prepared as follows. First, as before, 500pl of the supernatant was transferred into an Amicon Ultra-0.5 Centrifugal Filter Devices with 30 KDa cutoff. The Amicon was centrifugated for 15 minutes at 14000g. The cells were resuspended in 1ml of PBS supplemented with DTT (lOmM) and incubated at room temperature for 20 minutes.

[0169] Afterwards, the cells were pelleted by centrifugation and 500pl of the supernatant was also transferred into the Amicon. The Amicon was again centrifugated for 15 minutes at 14000g. The other 500pl of the supernatant was discarded. The pelleted cells wereagain washed in 500pl of PBS and re-peletted by centrifugation at 16000g for 5min.250p I of the supernatant mixed with 250pl of fresh PBS was used to wash the concentrate in the Amicon and the residual volume was used as the supernatant sample for the DTT treatment (~20pl). The rest of the supernatant was discarded and the pelleted cells were resuspended in 200pl of fresh PBS. 20pl of the resuspended cells was used for the assay. d) Amylase assay

[0170] Using the samples prepared with the protocol above, the assay was performed as indicated by the manufacturer, with the exception that we added 3pl of spectinomycin in the assay to avoid any potential growth or contamination.Results

[0171] In each sample, the a-amylase activity was calculated as illustrated in Figure 7.The a-amylase activity was calculated as indicated by the manufacturer's protocol by calculating the rate of paranitrophenol production. The quantity of paranitrophenol produced was evaluated by measuring the OD (405nm) of the sample each 5 min. Only the datapoints within the linear portion of the sampling curve, occurring strictly inside the first 3 hours, were considered, in a way that guaranteed R2 > 0.99 for each sample.

[0172] As demonstrated in Figure 7, after DTT treatment, the amylase activity in the cell decreases, whereas the amylase activity increases in the supernatant. These results confirm that the disulfide bridges are cleaved after DTT treatment, and that the anchored chimeric polyprotein is cleaved to form a first soluble protein that is released from its cell- wall anchor, and a second protein entity that remains peptidoglycan-anchored.

[0173] Thus, this assay illustrates the potential application of the polynucleotide of the present invention.3: Assessment of the self-: of five BIL-recombinant constructs derived from Pseudomonas

[0174] MBP-BIL-MycHisTag constructs were expressed in E. coli and purified by IMAC (ion metal affinity chromatography). Purified proteins were eluted from IMAC column with imidazole solutions (250 mM) or after incubation with reducing agent (DTT 2 mM, during 100 minutes). Samples obtained from imidazole elution were analyzed by SDS-PAGE and western blot (anti-His antibody) before and after treatment with reducing agent (DTT). Samples obtained by DTT elution were analyzed by SDS-PAGE only.

[0175] The five tested constructs include BIL domains comprising sequences SEQ ID NO: 301, 302, 86, 101 and 85. Respectively, the tested BIL sequences further comprise each N-ter and C-ter polypeptide sequences consisting of sequences SEQ ID NO: 303, 304, 186, 305 and 306:SEQ ID Origin Amino acid sequenceNO>WP_052483384.1 ANAAPEIGPCCFAAGTKVSTPDGDRAIETLKIGDIV hemagglutinin repeatWSKPEKGGKPFAATITATHVRNDQPIYRLTLKGSD containing protein LNGKTTRETLLVTPGHPFYVPAQKDFVPVIDLKLG 303[Pseudomonas sp. DRLQSLADGATENTSSEVESLELYAPVGTTYNLTV StFLB209] DVGHTFYVGDLKTWVHNTGPCDPAARG GRLLGEKGPCCFAAGTKVSTPNGDRTIESLKVGDV>WP_187519014.1 two- VWSKPEKGGKPFAAKILATHQRSDQPIYRLKLKSVpartner secretion domainRADGTAAGETLLVTPSHPFYVPAKHDFIPVIDLKPG304 containing proteinDLLQSLADGDSDNTSSEVESLELYLPEGKTYNLTV[Pseudomonas triticifolii]DVGHTFYVGELKTWVHNTGPCDLPDGY>WP_122316857.1:4155- GRKLGSEGPCCFAAGTMVSTPDGDRAIDTLKVGDI 4323 filamentous VWSKPERGGKPFAAAILATHVRTDQPIYRLKLKSV 186 hemagglutinin N-terminal RQDGQAEEETLLVTPGHPFYVPAQRDFVPVIDLKP domain-containing protein GDRLQSLEDGASENTSSEVESLELYLPEGKTYNLT [Pseudomonas cichorii] VDVGHTFYVGKLKTWVHNTGPCVLPEGYF >MDG6399431.1:3-438 LRKLEPLGPCCFAAGTMVSTPDGDRAIDTLKIGDIV polymorphic toxin-type WSKPEKGGKPFAAAILATHVRTDQPIYRLKLKSIR HINT domain-containing QDGNAESEVLMVTPSHPFYVPARRDFVPVIDLKAG 305protein [Pseudomonas DQLQSLADGAGEGASSVVESLELYLPVGKTYNLT quasicaspiana] VDVGHTFYVGKLKTWVHNTGPCPIKGEPGRKFGEAGPCCFAAGTMVSTPDGERAIDTLKVGDI >WP_198697465.1 two- VWSKPEGGGEPFAAAILATHIRTDQPIYRLKLKGR partner secretion domainQEDGKAEDETLLVTPGHPFYVPARHGFIPVIDLKPG306 containing proteinDRLQSLADGASENTSSEVESLELYLPVGKTYNLTV[Pseudomonas viridiflava]DIGHTFYVGKLKTWVHNVGPCELPEGY

[0176] For each construct, the two adjacent N-terminal cysteines, and the C-terminal cysteine, respectively defined as the N-junction and C-junction are reported in bold.

[0177] It can thus be confirmed in vitro that such recombinant constructs are also endowed with self-splicing activity, with the formation of a disulfide bridge as previously reported in Example 1.

Claims

CLAIMS1. A polynucleotide encoding a polypeptide comprising or consisting of SEQ ID NO:300.

2. The polynucleotide according to claim 1, said polypeptide comprising or consisting of SEQ ID NO: 1, preferably wherein the polypeptide comprises or consists of a sequence SEQ ID NO: 2.

3. The polynucleotide according to claim 1 or 2, wherein the polypeptide comprises or consists of a sequence selected from the group consisting of SEQ ID NO: 3 to 7, 12, 16 to 23, 26 to 47 or 49 to 102, and sequences with at least 70% identity with SEQ ID NO: 3 to 7, 12, 16 to 23, 26 to 47 or 49 to 102.

4. The polynucleotide according to claim 1 or claim 2 or claim 3, wherein the polypeptide further comprises in Nter a sequence selected from the group consisting of SEQ ID NO: 203 to 209, preferably a sequence selected from the group consisting of SEQ ID NO: 210 to 216, and more preferably a sequence selected from the group consisting of SEQ ID NO: 217 to 222 or 239.

5. The polynucleotide according to any one of claims 1 to 4, wherein the polypeptide further comprises in Cter a sequence selected from the group consisting of SEQ ID NO: 223 to 227, preferably a sequence selected from the group consisting of SEQ ID NO: 228 to 232.

6. The polynucleotide according to claim 5, comprising or consisting of a sequence of SEQ ID NO: 233.

7. A polynucleotide for expressing a chimeric polyprotein, wherein said polynucleotide comprises or consists of, from 5’ to 3’:optionally a sequence encoding a signal peptide,a sequence encoding a first protein entity,a polynucleotide according to any one of claims 1 to 6, anda sequence encoding a second protein entity,wherein the first and second protein entities may be the same or different.

8. An expression cassette comprising a polynucleotide according to any one of claims I to 7.

9. An expression vector comprising a polynucleotide according to any one of claims 1 to 7 or an expression cassette according to claim 8.

10. A host cell comprising a polynucleotide according to any one of claims 1 to 7, an expression cassette according to claim 8, or an expression vector according to claim 9.

11. A method for producing a chimeric polyprotein, wherein said polyprotein comprises a first protein entity linked to a second protein entity by a cleavable sequence comprising a disulfide bridge, wherein said method comprises expressing a polynucleotide according to claim 7.

12. The method according claim 11, wherein the cleavable sequence comprises or consists of two peptides linked by a disulfide bridge, wherein the first peptide comprises or consist of a sequence PC, and wherein the second peptide comprises or consist of a sequence XevGPC (SEQ ID NO: 238, wherein X67 is T, I or V).

13. The method according to claim 11 or claim 12, comprising a step of culturing a host cell comprising a polynucleotide according to claim 7, preferably wherein the host cell is a bacterium.

14. A chimeric polyprotein obtained or susceptible to be obtained by the method according to any one of claims 11 to 13, wherein said chimeric polyprotein comprises a first protein entity and a second protein entity, wherein said first and second protein entities are linked by a disulfide bridge under oxidative conditions and may be dissociated under reductive conditions.

15. The chimeric polyprotein of claim 14, wherein the first and second protein entities are a protein of interest and a purification tag.

16. The chimeric polyprotein of claim 14, wherein the first and second protein entities are a protein of interest and an anchoring domain.

17. The chimeric polyprotein according to any one of claims 14 to 16; wherein:- the first protein entity comprises a first cysteine forming the disulfide bridge, and the second protein entity comprises a second cysteine forming the disulfide bridge;- the first protein entity comprises or consists of sequence PC as the first cysteine;- the second protein entity comprises or consists of a sequence XevGPC (SEQ ID NO: 238, wherein X67 is T, I or V) as the second cysteine.