Compositions and methods for producing complex carbohydrate polypeptides having an isopeptide bond with a second polypeptide partner, and use thereof.

By using the SpyTag/SpyCatcher system to form isopeptide bonds in bioconjugated vaccines, the problems of complexity and low efficiency in existing technologies have been solved, enabling the efficient preparation of diverse immunogenic compositions and enhancing the immune response of vaccines.

JP2026509192APending Publication Date: 2026-03-17VAXNEWMO LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-26
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing methods for preparing bioconjugated vaccines are complex and inefficient, making it difficult to generate diverse immunogenic compositions and thus failing to fully stimulate an immune response.

Method used

Using the SpyTag/SpyCatcher system, carbohydrates are linked to carrier proteins through the formation of heteropeptide bonds between the first and second polypeptide tags, generating complex carbohydrate polypeptides that form self-assembled multimeric structures such as nanocages or virus-like particles, thus achieving efficient preparation of biological conjugated vaccines.

Benefits of technology

This has enabled the efficient preparation of bioconjugated vaccines, enhanced immune responses, and improved antibody affinity and immune memory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026509192000001_ABST
    Figure 2026509192000001_ABST
Patent Text Reader

Abstract

This disclosure subsequently provides a description of compositions and methods used to produce complex carbohydrate polypeptides using an enzyme that forms glycosidic bonds that form isopeptide bonds using a second polypeptide containing a polypeptide tag, and also provides a description of the use thereof.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This PCT application claims the interests of U.S. Provisional Application No. 63 / 448,382, U.S. Provisional Application No. 63 / 448,386, and U.S. Provisional Application No. 63 / 448,408, filed on 27 February 2023, each of which is incorporated herein by reference in whole.

[0002] Reference to electronically submitted sequence listings The sequence listing, which measures 441,619 bytes (measured on MS-Windows®), contains 440 sequences, and is located in a file named "64100_234947_SL.xml" created on February 26, 2024, is provided herein by reference through the USPTO Patent Centre and is incorporated herein by reference in its entirety. [Background technology]

[0003] Glycan protein or glycoprotein conjugate vaccines are effective therapeutic agents against a variety of bacterial pathogens and are essentially covalent conjugates of bacterial surface oligosaccharides or polysaccharides to carrier proteins. These vaccines can induce a protective immune response against O antigens or capsules present on the surface of Gram-positive or Gram-negative pathogens. Unlike pure polysaccharide vaccines, glycoprotein conjugate vaccines can stimulate robust immunological memory by inducing T cell recruitment, memory cell formation, and switching from B cell IgM to IgG antibody classes, leading to the establishment of immunological memory (Rappuoli, et al. (2019) PNAS, 116(1) 14-16, Avci, F., et al. (2011) Nat Med 17, 1602-1609). Glycoprotein conjugate vaccines can be produced through a number of different methods. Historically, conjugates used in routine medical applications have been created by chemically crosslinking purified oligosaccharides or polysaccharides to amino acid side chains on purified proteins using different linker molecules and chemistry (Berti, F. and Adamo, R. (2018) Chem. Soc. Rev., 2018, 47, 9015-9025). Bioconjugates are an alternative to chemical conjugation and rely on oligosaccharide transferase (OTase) enzymes that catalyze the covalent bonding of lipid-binding oligosaccharides or polysaccharides to specific amino acid residues on substrate proteins (Harding, C. and Feldman, M. (2019) Glycobiology, Volume 29, Issue 7, July 2019, Pages 519-529, Feldman, M. (2005) PNAS, 102(8) 3016-3021). These bioconjugates are often produced in manipulated strains of E. coli, and conjugation occurs in the bacterial periplasm. The glycan substrate of OTase binds to membrane-bound lipid carriers such as undecaprenyl pyrophosphate (UNDPP).O-linked OTases catalyze the transfer reaction of UNDPP-linked oligosaccharides or polysaccharides to serine or threonine side-chain hydroxyl groups within conserved protein motifs called sequencers (Knoot, C., et al. (2021) Glycobiology, Volume 31, Issue 9, September 2021, Pages 1192-1203; Knoot, C., et al. (2023) Glycobiology, Volume 33, Issue 1, January 2023, Pages 57-74). By co-expressing a desired glycan biosynthesis gene cluster, carrier protein(s), and OTase(s) chromosomally and / or episomatically in a bacterial host, bioconjugate vaccines can be generated in a "one-pot" biological reaction and subsequently purified downstream (Harding, C. and Feldman, M. (2019) Glycobiology, Volume 29, Issue 7, July 2019, Pages 519-529).

[0004] Several bioconjugate vaccines are currently in clinical trials, each consisting of a bacterial glycan linked to a periplasmic protein, primarily Pseudomonas aeruginosa exotoxin A (EPA), Haemophilus protein D, or CRM197 (modified diphtheria toxin) (Sorieul, C., et al. (2023) Expert Review of Vaccines, 22:1, 1055-1078). Alternatives to such carrier proteins are protein nanoparticles (NPs) or virus-like particles (VLPs): symmetrical self-assembling protein "cages" that are either derivatives of natural viral capsids or rationally engineered protein assemblies (Nguyen, B. and Tolia, N. (2021) npj Vaccines 6, 70; Bruun, T., et al. (2018) ACS Nano 2018, 12, 9, 8855-8866; Cohen, et al. (2021) PLoS ONE 16(3):e0247963). Expression of NP / VLP monomer proteins in bacterial cytoplasm results in the spontaneous assembly of megadalton-sized particles that can be purified from cellular biomass (Cohen, et al. (2021) PLoS ONE 16(3):e0247963). NP / VLP-based therapeutics have been shown to partially improve the immune response by increasing antibody avidity resulting from the larger particle size of the immunogen (Nguyen, B. and Tolia, N. (2021) npj Vaccines 6, 70). Two examples of VLP / NP are AP205 and mi3.AP205, which are derived from the CP3 coat protein of the RNA bacteriophage AP205 (Brune, K., et al. (2016) Sci Rep 6, 19234). AP205 VLPs are assembled into 120-mers with a diameter of approximately 20 nm (Cohen, et al. (2021) PLoS ONE 16(3):e0247963).mi3 is a porous dodecahedral 60-mer derived from computationally designed NPs with a diameter of 20–30 nm (Bruun, T., et al. (2018) ACS Nano 2018, 12, 9, 8855-8866) (Hsia, Y., et al. (2016) Nature volume 535, pages 136-139).

[0005] The SpyTag / SpyCatcher system originates from the immunoglobulin-like collagen adhesin domain (CnaB2) of the fibronectin-binding protein FbaB2 from Streptococcus pyogenes (Zakeri, B. et al. (2012)). The CnaB2 domain spontaneously forms an intraprotein isopeptide bond between lysine at position 31 and aspartic acid at position 117. Specifically, the aprotonated amine at Lys31 acts as a nucleophile that attacks the carbonyl carbon of Asp117, catalyzed by glutamic acid at position 77 (Zakeri, B. et al. (2012)). This isopeptide reaction occurs spontaneously and appears to be characteristic of several members of bacterial proteins in the prealbumin-like fold domain to which CnaB2 belongs. To adapt this as a tool for creating covalent bonds between two distinct polypeptides, the CnaB2 domain was split, and CnaB2 was isolated into (1) a peptide containing a C-terminal β-chain with reactive Asp117, defined as SpyTag, and (2) a protein-binding partner derived from the remaining CnaB2 polypeptide, defined as SpyCatcher (Zakeri, B. et al. (2012)). This system was named SpyTag and SpyCatcher to indicate the bacterial source (S. pyogenes) of the CnaB2 fragment. Later forms of the SpyTag and SpyCatcher system, SpyTag003 and SpyCatcher003, were created by phage display and subsequent rational engineering approaches, yielding a reaction rate of 5.5 × 10⁵ M-1 s-1. SpyTag003 / SpyCatcher003 reacts approximately 400 times faster than the original SpyTag / SpyCatcher system (Keeble, A.Het al. (2019)). SpyTag / SpyCatcher systems are widely applied to enable covalent bonding of two separate polypeptides, a SpyTag-containing polypeptide and a SpyCatcher-containing polypeptide, through the formation of an isopeptide bond.The SpyTag / SpyCatcher system has been applied to a variety of biological applications, aiming to covalently bond a target polypeptide containing a SpyTag to a different protein or material containing a SpyCatcher. These applications include, but are not limited to, anchoring different solid organic and inorganic materials to surfaces, attaching polypeptides to different polymerization architectures such as nanoparticles, virus-like particles, or adenovirus vectors, and directly attaching polypeptides to the surface of intact cells (Keeble, AH & Howarth, M. (2020), Brune, KDet al. (2016), Bruun, TUJ, Andersson, AC, Draper, SJ & Howarth, M. (2018)).

[0006] There remains a need to create novel and more effective bioconjugate vaccines, including flexible methods for generating diverse immunogenic compositions.

[0007] Summary of the Invention A fusion protein comprising (i) a glycosylated fragment and (ii) a first polypeptide tag is provided herein, wherein the first polypeptide tag can spontaneously form an isopeptide bond with a second polypeptide tag binding partner. In certain embodiments, the fusion protein is a complex carbohydrate containing a sugar covalently bonded to the fusion protein via the glycosylated fragment, and is optionally immunogenic. Representative examples of the first polypeptide tag include SpyTag (SEQ ID NO: 416), SpyTag002 (SEQ ID NO: 417), SpyTag003 (SEQ ID NO: 418), or DogTag (SEQ ID NO: 419).

[0008] In certain embodiments, the glycosylated fragment is a ComP glycosylated fragment or a variant thereof as described herein.

[0009] In certain embodiments, the glycosylated fragment is a TfpM-related pilling glycosylated fragment or a variant thereof as described herein.

[0010] In certain embodiments, the glycosylated fragment is a PilE glycosylated fragment or a variant thereof as described herein.

[0011] In certain embodiments, the glycosylated fragment is a PglB glycosylated fragment or a variant thereof as described herein.

[0012] In certain embodiments, the glycosylated fragment is a PilA glycosylated fragment or a variant thereof as described herein.

[0013] In certain embodiments, the glycosylated fragment is an STT3 glycosylated fragment or a variant thereof as described herein.

[0014] In certain embodiments, the glycosylated fragment is an N-linked glycosyltransferase glycosylated fragment.

[0015] In certain embodiments, the glycosylated fragment is an O-linked glycosyltransferase glycosylated fragment.

[0016] In certain embodiments, the glycosylated fragment is a PilA_Pa5196-related pyring glycosylated fragment or a variant thereof as described herein.

[0017] In certain embodiments, the fusion protein comprises a carrier protein, which is optionally selected from the group consisting of Escherichia coli maltose-binding protein, Pseudomonas aeruginosa exotoxin A (EPA), Pseudomonas aeruginosa PcrV, CRM197, Haemophilus influenzae protein D, cholera toxin B subunit, or tetanus toxin, and any fragment thereof.

[0018] A composition comprising a polypeptide pair comprising a first polypeptide and a second polypeptide is also provided herein, wherein the first polypeptide is the fusion protein of the Disclosure, and the second polypeptide comprises a second polypeptide tag binding partner to a first polypeptide tag of the first polypeptide, wherein the first polypeptide is bound to the second polypeptide via an isopeptide bond between the first polypeptide tag and the second polypeptide tag. In certain embodiments, the second polypeptide comprises a monomer polypeptide that can spontaneously multimerize / self-assemble into a higher-order multimer structure, which optionally is an icosahedral or dodecahedral particle (e.g., similar to a nanocage), a virus-like particle, or an adenovirus vector.

[0019] In a particular embodiment, the second polypeptide tag is SpyCatcher (SEQ ID NO: 420), SpyCatcher002 (SEQ ID NO: 421), SpyCatcher003 (SEQ ID NO: 422), or DogCatcher (SEQ ID NO: 423).

[0020] In certain embodiments, the first polypeptide is a bioconjugate comprising a sugar covalently bonded to a glycosylated fragment of the first polypeptide, and optionally, the composition is immunogenic.

[0021] Complexes comprising two or more polypeptide pairs of the present disclosure are also provided. In certain embodiments, the complex is a self-assembled multimeric higher-order structure. In certain embodiments, the self-assembled multimeric higher-order structure is an icosahedral or dodecahedral particle (e.g., similar to a nanocage), a virus-like particle, or an adenovirus vector.

[0022] A method for producing polypeptide pairs of the present disclosure is also provided, which comprises contacting a first polypeptide and a second polypeptide under conditions that allow the first polypeptide tag to spontaneously form an isopeptide bond with a second polypeptide tag binding partner.

[0023] Also provided is a method for producing a complex of the present disclosure, comprising (i) forming a self-assembled polymeric higher-order structure of the second polypeptide, and then contacting the first polypeptide and the second polypeptide under conditions that allow the first polypeptide tag to spontaneously form an isopeptide bond with the second polypeptide tag, or (ii) contacting the first polypeptide and the second polypeptide under conditions that allow the first polypeptide tag to spontaneously form an isopeptide bond with the second polypeptide tag, and then forming a self-assembled polymeric higher-order structure of the second polypeptide. In certain embodiments, the ComP glycosylated fragment is glycosylated by PglS OTase, the TfpM-related pilling glycosylated fragment is glycosylated by TfpM OTase, and optionally, the ComP glycosylated fragment is glycosylated by PglS OTase, the TfpM-related pilling glycosylated fragment is glycosylated by TfpM OTase, the PilE glycosylated fragment is glycosylated by PglL OTase, the PglB glycosylated fragment is glycosylated by PglB OTase, and the PilA glycosylated fragment is glycosylated by TfpO or PilO The fragments are glycosylated by OTase, the STT3 glycosylated fragments by the STT3 catalytic subunit, the PilA_Pa5196-associated pyring glycosylated fragments by TfpW glycosyltransferase, the N-linked glycosyltransferase fragments by N-linked glycosyltransferases from Actinobacillus pleuropneumoniae, Haemophilus influenzae, or Yersinia enterocolitica, and / or the O-linked glycosyltransferase fragments by GtfA / GtfB glycosyltransferases.

[0024] Methods for inducing an immune response in a subject by administering to the subject an effective amount of any of the compositions, complexes, and / or conjugate vaccines of the Disclosure, or a composition, complex, and / or conjugate vaccine of the Disclosure for use in inducing an immune response in the subject, are also provided herein. [Brief explanation of the drawing]

[0025] [Figure 1] A schematic diagram of glycoprotein polypeptide production using an enzyme that forms glycosidic bonds, and subsequent isopeptide bond formation using a second protein containing a polypeptide tag. The first polypeptide (protein 1) contains both (i) a SpyTag that spontaneously forms an isopeptide bond with the polypeptide tag (SpyCatcher) of the second partner polypeptide (protein 2), and (ii) a glycosylated fragment (sequon) that is recognized by a specific enzyme that covalently converts sugar into a glycosylated fragment (sequon) to form a glycosidic bond. In this example, protein 2 containing the polypeptide tag (SpyCatcher) can also self-assemble into higher-order structures such as icosahedral or dodecahedral nanoparticles similar to a nanocage, virus-like particles (VLPs), or adenovirus vectors. [Figure 2]Coomassie-stained SDS-PAGE denatured gel of purified MBP-SpyTag-v1 E. coli O16 O-antigen bioconjugate produced using the TfpM oligosaccharide transferase system, mi3-SpyCatcher, and isopeptide-bound MBP-SpyTag-v1 E. coli O16:mi3-SpyCatcher. The box diagram shows the schematic structures of the SpyTag protein (corresponding to protein 1 in Figure 1) and SpyCatcher protein (corresponding to protein 2 in Figure 1) designed to fit the enzymes that form glycosidic bonds. Abbreviations: MBP, maltose-binding protein; MBPsp, E. coli maltose-binding protein secretion (Sec) signal peptide; Pil20, 20-amino acid pyrin sequence glycosylated by TfpM oligosaccharide transferase; mi3, mi3 nanoparticle monomer; 6xHis, hexahistidine tag. For SDS-PAGE analysis, SpyCatcher and SpyTag proteins were reacted individually or in 1:1 or 2:1 ratios (based on protein concentration) in Tris-buffered saline at room temperature for 2 hours. The formation of isopeptide bond reactions was stopped by adding Laemmli buffer and then heating the samples at 100°C for 10 minutes before gel packing. Protein masses related to proteins and glycoproteins are shown above each lane. Lane A) Protein ladder with standard masses of kDa marked on the left. Lane B) Purified MBP-SpyTag-v1-O16 bioconjugate. Lane C) Purified mi3-Spycatcher. Lane D) 1:1 reaction mixture of mi3-SpyCatcher and MBP-SpyTag-v1-O16. Lane E) 2:1 reaction mixture of mi3-SpyCatcher and MBP-SpyTag-v1-O16. [Figure 3]Western blot of purified MBP-SpyTag-v1 E.coli O16 O-antigen bioconjugates generated using the TfpM oligosaccharide transferase system, mi3-SpyCatcher, and isopeptide-bonded MBP-SpyTag-v1 E.coli O16:mi3-SpyCatcher. Prior to Western blot analysis, SpyTag and SpyCatcher proteins were reacted in Tris-buffered saline for 2 hours, either alone or in 1:1 or 2:1 ratios (based on protein concentration). Protein masses related to proteins and glycoproteins are shown above each lane. Western blots were probed with anti-His tag (α-His, top panel) and E.coli O16 O-antigen antiserum (α-O16, middle panel). The combined image is shown in the bottom panel. Lane A) Protein ladder with standard mass of kDa marked on the left. Lane B) Purified MBP-SpyCatcher-v1-O16 bioconjugate. Lane C) Purified mi3-Spycatcher. Lane D) 1:1 reaction mixture of mi3-SpyTag and MBP-SpyTag-v1-O16. Lane E) 2:1 reaction mixture of mi3-SpyCatcher and MBP-SpyTag-v1-O16. [Figure 4] Size exclusion chromatography of mi3-SpyCatcher and isopeptide-bound MBP-SpyTag-v1-O16:mi3-SpyCatcher separated using a Sephacryl S-400 HR 16 / 600 column. UV absorbance traces of 250 μg of mi3-SpyCatcher (solid line) and 250 μg of mi3-SpyCatcher after reaction with 250 μg of MBP-SpyTag-v1-O16 (dashed line) in Tris-buffered saline at room temperature for 2 hours are shown. [Figure 5]Coomassie-stained SDS-PAGE denatured gel of purified MBP-SpyTag-v2 E. coli O16 O-antigen bioconjugate produced using the PglS oligosaccharide transferase system, mi3-SpyCatcher, and isopeptide-bound MBP-SpyTag-v2 E. coli O16:mi3-SpyCatcher. The box diagram shows the schematic structures of the SpyTag and SpyCatcher proteins in this experiment. Abbreviations: MBP, maltose-binding protein; MBPsp, E. coli maltose-binding protein secretion (Sec) signal peptide; ComP sequencer, sequencer derived from 23 amino acid ComP glycosylated by PglS oligosaccharide transferase; mi3, mi3 nanoparticle monomer; 6xHis, hexahistidine tag. For SDS-PAGE analysis, SpyTag and SpyCatcher proteins were reacted individually or in 1:1 or 2:1 ratios (based on protein concentration) in Tris-buffered saline at room temperature for 2 hours. The formation of isopeptide bonds was stopped by adding Laemmli buffer and then heating the sample at 100°C for 10 minutes before gel packing. Samples were taken from each reaction for SDS-PAGE analysis. Protein masses related to proteins and glycoproteins are shown above each lane. Lane A) Protein ladder with standard masses of kDa marked on the left. Lane B) Purified MBP-SpyTag-v2-O16 bioconjugate. Lane C) Purified mi3-SpyCatcher. Lane D) 1:1 reaction mixture of mi3-SpyCatcher and MBP-SpyTag-v2-O16. Lane E) A 2:1 reaction mixture of mi3-SpyCatcher and MBP-SpyTag-v2-O16. [Figure 6]Western blot of purified MBP-SpyTag-v2 E.coli O16 O-antigen bioconjugates generated using the PglS oligosaccharide transferase system, mi3-SpyCatcher, and isopeptide-bound MBP-SpyTag-v2 E.coli O16:mi3-SpyCatcher. Prior to Western blot analysis, SpyTag and SpyCatcher fusion proteins were reacted in Tris-buffered saline for 2 hours, either alone or in a 1:1 or 2:1 ratio (based on protein concentration). Western blots were probed with anti-His tag (α-His, top panel) and E.coli O16 O-antigen antiserum (α-O16, middle panel). The combined image is shown in the bottom panel. Lane A) Protein ladder with standard mass of kDa marked on the left. Lane B) Purified MBP-SpyTag-v2-O16 bioconjugate. Lane C) Purified mi3-SpyCatcher bioconjugate. Lane D) 1:1 reaction mixture of mi3-SpyCatcher and MBP-SpyTag-v2-O16. Lane E) 2:1 reaction mixture of mi3-SpyCatcher and MBP-SpyTag-v2-O16. [Figure 7]Coomassie-stained SDS-PAGE denatured gel of purified MBP-SpyTag-v1 E. coli O16 O-antigen bioconjugate produced using the TfpM oligosaccharide transferase system, AP205-SpyCatcher, and isopeptide-bound MBP-SpyTag-v1 E. coli O16:AP205-SpyCatcher. The box diagram shows the schematic structures of the SpyTag and SpyCatcher proteins in this experiment. Abbreviations: MBP, maltose-binding protein; MBPsp, E. coli maltose-binding protein secretion (Sec) signal peptide; Pil20, 20 amino acid pyrin sequence glycosylated by TfpM oligosaccharide transferase; AP205, AP205 virus-like protein coat protein monomer; 6xHis, hexahistidine tag. For SDS-PAGE analysis, SpyCatcher and SpyTag proteins were reacted individually or in 1:1 or 2:1 ratios (based on protein concentration) in Tris-buffered saline at room temperature for 2 hours before gel loading. Protein masses related to proteins and glycoproteins are shown above each lane. Lane A) Protein ladder with standard masses of kDa marked on the left. Lane B) Purified MBP-SpyTag-v1-O16 bioconjugate. Lane C) Purified AP205-SpyCatcher. Lane D) 1:1 reaction mixture of AP205-SpyCatcher and MBP-SpyTag-v1-O16. Lane E) 2:1 reaction mixture of AP205-SpyCatcher and MBP-SpyTag-v1-O16. [Figure 8]Coomassie-stained SDS-PAGE denatured gel of purified MBP-SpyTag-v2 E. coli O16 O-antigen bioconjugate produced using the PglS oligosaccharide transferase system, AP205-SpyCatcher, and isopeptide-bound MBP-SpyTag-v2 E. coli O16:AP205-SpyCatcher. The box diagram shows the schematic structures of the SpytTag and SpyCatcher proteins in this experiment. Abbreviations: MBP, maltose-binding protein; MBPsp, E. coli maltose-binding protein secretion (Sec) signal peptide; ComP sequon, sequon derived from 23 amino acid ComP glycosylated by PglS OTase; AP205, AP205 virus-like protein-coated protein monomer; 6xHis, hexahistidine tag. For SDS-PAGE analysis, SpyTag and SpyCatcher proteins were reacted individually or in 1:1 or 2:1 ratios (based on protein concentration) in Tris-buffered saline for 2 hours in a room. The formation of isopeptide bonds was stopped by adding Laemmli buffer and then heating the sample at 100°C for 10 minutes before gel packing. Samples were taken from each reaction for SDS-PAGE analysis. Protein masses related to proteins and glycoproteins are shown above each lane. Lane A) Protein ladder with standard masses of kDa marked on the left. Lane B) Purified MBP-Spycatcher-v2-O16 bioconjugate. Lane C) Purified AP205-Spycatcher. Lane D) 1:1 reaction mixture of AP205-Spytag and MBP-Spytag-v2-O16. Lane E) A 2:1 reaction mixture of AP205-Spycatcher and MBP-Spytag-v2-O16. [Figure 9]Coomassie-stained SDS-PAGE denatured gel of three purified EPA-SpyTag E. coli O16 O-antigen bioconjugates produced using the PglS oligosaccharide transferase system, mi3-SpyCatcher, and isopeptide-linked EPA-Spytag E. coli O16:mi3-SpyCatcher. The box diagram shows the schematic structure of the EPA SpyTag protein designed to fit the enzyme that forms the glycosidic bond. Abbreviations: EPA, Pseudomonas aeruginosa exotoxin A; MBPsp, E. coli maltose-binding protein secretion (Sec) signal peptide; ComP sequence, 23-amino acid pyrin sequence glycosylated by PglS oligosaccharide transferase; mi3, mi3 nanoparticle monomer; 6xHis, hexahistidine tag. The isopeptide bond formation reaction was carried out using purified EPA-Spycatcher-O16 bioconjugate and purified mi3-Spycatcher. Lane A) Protein ladder with standard mass of kDa marked on the left. Lane B) Purified mi3-Spycatcher. Lane C) Purified EPA-Spytag-v1-O16. Lane D) 1:1 reaction mixture of mi3-Spycatcher and EPA-SpyTag-v1-O16. Lane E) 1:2 reaction mixture of mi3-Spycatcher and EPA-SpyTag-v1-O16. Lane F) Purified EPA-Spytag-v2-O16. Lane G) 1:1 reaction mixture of mi3-Spycatcher and EPA-SpyTag-v2-O16. Lane H) 1:2 reaction mixture of mi3-Spycatcher and EPA-SpyTag-v2-O16. Lane I) Purified EPA-Spytag-v3-O16. Lane J) 1:1 reaction mixture of mi3-Spycatcher and EPA-SpyTag-v3-O16. Lane K) 1:2 reaction mixture of mi3-Spycatcher and EPA-SpyTag-v3-O16. No isopeptide-binding proteins were observed using EPA-Spytag-v2-O16 or EPA-Spytag-v3-O16 and mi3-Spycatcher. [Figure 10]Western blots of purified non-glycosylated EPA-SpyTag-v1 protein linker variant, mi3-SpyCatcher, and isopeptide-linked EPA-SpyTag E. coli:mi3-SpyCatcher. Each protein variant has a different amino acid linker between Spytag003 and EPA. The isopeptide bond formation reaction was performed using EPA-Spytag protein from E. coli periplasm extract and purified mi3-Spycatcher. Western blots were probed with anti-His tag antibody. Lane A) Protein ladder with standard mass of kDa marked on the left. Lane B) Purified EPA-Spytag-v1 with linker L1 (SGG). Lane C) 1:1 reaction mixture of mi3-Spycatcher and EPA-Spytag-v1 with linker L1. Lane D) Purified EPA-Spytag-v1 with linker L2 (SEQ ID NO: 430). Lane E) A 1:1 reaction mixture of mi3-Spycatcher, EPA-Spytag-v1, and linker L2. Lane F) Purified EPA-Spytag-v1 having linker L3 (SEQ ID NO: 431). Lane G) A 1:1 reaction mixture of mi3-Spycatcher, EPA-Spytag-v1, and linker L3. Lane H) Purified EPA-Spytag-v1 having linker L4 (SEQ ID NO: 432). Lane I) A 1:1 reaction mixture of mi3-Spycatcher, EPA-Spytag-v1, and linker L4. Lane J) Purified EPA-Spytag-v1 having linker L5 (SEQ ID NO: 433). Lane K) A 1:1 reaction mixture of mi3-Spycatcher, EPA-Spytag-v1, and linker L5. Lane L) Purified EPA-Spytag-v1 having linker L6 (SEQ ID NO: 434). Lane M) A 1:1 reaction mixture of mi3-Spycatcher and EPA-Spytag-v1 with linker L6. Lane N) Purified EPA-Spytag-v1 with linker L7 (SEQ ID NO: 435). Lane O) A 1:1 reaction mixture of mi3-Spycatcher and EPA-Spytag-v1 with linker L7.Lane P) Purified EPA-Spytag-v1 having linker L8 (SEQ ID NO: 436). Lane Q) 1:1 reaction mixture of mi3-Spycatcher and EPA-Spytag-v1 with linker L8. Lane R) Purified EPA-Spytag-v1 having linker L9 (SEQ ID NO: 437). Lane S) 1:1 reaction mixture of mi3-Spycatcher and EPA-Spytag-v1 with linker L9. Lane T) Purified EPA-Spytag-v1 having linker L10 (SEQ ID NO: 438). Lane U) 1:1 reaction mixture of mi3-Spycatcher and EPA-Spytag-v1 with linker L10. Some variants exhibit low isopeptide bond expression with mi3-Spycatcher and / or are unable to form such a bond. [Figure 11] Western blot of non-glycosylated EPA-Spytag-v1 over time with linker L7 (EAAAKEAAAK; SEQ ID NO: 435) and mi3-Spycatcher isopeptide formation reaction. The isopeptide bond formation reaction was performed using EPA-Spytag-v1 with linker L7 from E. coli periplasm extract and purified mi3-Spycatcher in a 1:1 ratio. The left panel shows the Western blot probed with anti-His tag antibody. The center panel shows the Western blot probed with anti-EPA antibody. The right panel is a merged image of the two channels. Lane A) Protein ladder with standard mass of kDa marked on the left. Lane B) mi3-Spycatcher only. Lane C) EPA-Spytag-v1 with linker L7 only. Lane D) Isopeptide bond formation reaction between mi3-Spycatcher and EPA-Spytag-v1 with linker L7 after 0.5 hours. Lane E) Reaction after 1 hour. Lane F) Response after 2 hours. Lane G) Response after 5 hours. Lane H) Response after 24 hours. [Figure 12]Figures A-E show a schematic diagram of the EPA-ComP110264 fusion protein, in which the ComP glycosylated fragment is fused at the C-terminus of the fusion protein. "ssDsbA" corresponds to the DsbA Sec secretion signal. GGGS (SEQ ID NO: 382) is a flexible linker between EPA and the ComP110264 fragment. Figure B shows different amino acid sequences of the ComP glycosylated fragment fused at the C-terminus of the EPA fusion protein. The bold underlined serine residues in each sequence correspond to the conserved serine 82 of ComP110264 and are glycosylation sites. The bold underlined cysteine ​​residues corresponding to Cys71 and Cys93 are also highlighted. (C2, SEQ ID NO: 383; D2, SEQ ID NO: 384; E2, SEQ ID NO: 385; F2, SEQ ID NO: 386; G2, SEQ ID NO: 387; H2, SEQ ID NO: 388; A3, SEQ ID NO: 389; B3, SEQ ID NO: 390; C3, SEQ ID NO: 391; D3, SEQ ID NO: 392; E3, SEQ ID NO: 393; F3, SEQ ID NO: 394; and C1, SEQ ID NO: 395). Figures 12C, 12D, and 12E show Western blot analysis of periplasmic extracts from E. coli SDB1 expressing PglS, CPS8 glycan, and the EPA-ComP110264 variant. Each lane in the Western blot panel corresponds to an SDB1 strain expressing a different EPA-ComP variant with a ComP glycosylated fragment corresponding to the sequence shown in Figure 12B. C shows the protein that reacts with anti-EPA antiserum. D shows the protein that reacts with anti-His antiserum. Panel E shows the merged Western blot images of Figures 12C and 12D. Equivalent amounts of periplasmic extract based on OD600 were loaded per lane. On the right side of the panels in Figures 12C–E, g0 represents unglycosylated EPA-ComP110264, and gn represents glycosylated EPA-ComP110264 with a different number of CPS8 repeat units. Protein mass markers (kDa) are shown on the left side of the panels in Figures 12C–E. [Figure 13]Figures A-D show a schematic diagram of the CRM197-ComPC1 fusion protein. "ssFlgI" corresponds to the FlgI SRP secretion signal. GGGS (SEQ ID NO: 382) is a flexible linker between CRM197 and ComPC1. Figures 13B, 13C, and 13D show Western blot analysis of the purified CRM197-ComPC1-CPS8 complex carbohydrate. B shows the protein reacting with anti-CPS8 antiserum. C shows the protein reacting with anti-CRM197 antiserum. D shows a combined Western blot image of Figures 13B and 13C. The loss of CRM197 and CPS8 signals in proteinase K (PK) treated samples indicates that the pneumococcal serotype 8 signal is bound to CRM197 and is not a result of contamination from free polysaccharides or lipid-bound polysaccharide precursors. The protein mass marker (kDa) is shown on the left side of the panels in Figures 13B–D. [Figure 14] A and B show schematic diagrams of the C-terminal and N-terminal CRM197 variants containing the C1 ComP glycosylated fragment. B shows Western blot analysis of periplasmic extracts and CPS8 glycans of E. coli SDB1 expressing CRM197-ComPC1 or ComPC1-CRM197 in the presence (+) or absence (-) of PglS. Equivalent amounts of periplasmic extract based on OD600 were loaded per lane. Protein mass markers (kDa) are shown to the left of GGGS (SEQ ID NO: 382). [Figure 15]Figures A-E show the following: A is a schematic diagram of an EPA fusion protein containing a ComP glycosylated fragment incorporated within the EPA amino acid sequence. B shows the amino acid sequences of two iGT ComP glycosylated fragments inserted between EPA residues Ala489 and Arg489. These have either two terminal cysteines ("iGCC"; SEQ ID NO: 230) or serine ("iGSS"; SEQ ID NO: 231). Figures 15C and 15D show Western blots of periplasm extracts of E. coli SDB1 expressing CPS8 glycan, EPAiGTcc or EPAiGTss, with or without PglS (+). C shows the protein reacting with anti-EPA antiserum. D shows the protein reacting with anti-His antiserum. E shows a combined Western blot image of Figures 15C and 15D. Equal volumes of periplasm extract based on OD600 were loaded per lane. The protein mass marker (kDa) is shown on the left side of the panel. [Figure 16]A-D shows schematic diagrams of the EPA constructs containing the ComP glycosylated fragments used in these experiments (from top to bottom, SEQ ID NOs: 6-28). 22-25 amino acid cleavage variants of the iGTCC ComP glycosylated fragment were inserted into the EPA coding sequence between Ala489 and Arg489. B shows the amino acid sequences of the 22 cleaved iGT ComP glycosylated fragments, with names assigned to the left. Underlined and bolded serine is the glycosylation site (iGTcc Sequence ID 230; Δ0-1 Sequence ID 232; Δ1-0 Sequence ID 243; Δ1-2 Sequence ID 245; Δ2-1 Sequence ID 256; Δ2-3 Sequence ID 258; Δ3-2 Sequence ID 269; Δ3-4 Sequence ID 271; Δ4-3 Sequence ID 282; Δ4-5 Sequence ID 284; Δ5-4 Sequence ID 295; Δ5-6 Sequence ID 297; Δ6-5 Sequence ID 308; Δ6-6 Sequence ID 309; Δ6-7 Sequence ID 310; Δ7-6 Sequence ID 321; Δ7-7 Sequence ID 322; Δ7-8 Sequence ID 323; Δ8-7 Sequence ID 334; Δ8-8 Sequence ID 335; Δ8-9 Sequence ID 336; Δ9-8 Sequence ID 346; Δ9-9 Sequence ID 347). C shows a Western blot analysis of periplasmic extracts of E. coli SDB1 expressing an EPAiGT fusion protein containing PglS, CPS8, and cleaved ComP glycosylated fragments. Each lane in the Western blot panel corresponds to a different SDB1 strain expressing an EPAiGT fusion protein containing cleaved ComP glycosylated fragments, with the ComP glycosylated fragments corresponding to the sequences shown in Figure 16B. C shows the protein reacting with an anti-EPA antiserum probe using an anti-EPA antibody. EPAiGTcc is shown for comparison. The "EPA" lane corresponds to EPA lacking any ComP-derived sequence and serves as a negative control. Equal volumes of periplasmic extract based on OD600 were loaded per lane. D shows the same Western blot as above with increased brightness of the anti-EPA signal to indicate low levels of glycosylation for the smallest ComP glycosylated fragments. [Figure 17]A, B, and C show Western blot analysis of EPA fusion proteins purified by Ni affinity chromatography, containing an iGTΔ6-6 ComP glycosylated fragment integrated between EPA residues Ala489 and Arg490. The fusion proteins were purified from SDB1 cells expressing CPS8 glycan in the presence (+) or absence (-) of PglS. A shows the protein reacting with anti-His antiserum. B shows the protein reacting with anti-CPS8 antiserum. C shows a combination of Figures 17A and 17B. Protein mass markers (kDa) are shown on the left side of the panels in Figures 17A–C. [Figure 18] Figures A and B show a schematic diagram of the EPA fusion protein containing an iGTΔ3-4 ComP glycosylated fragment incorporated between EPA residues Glu548 and Gly549. The iGTΔ3-4 amino acid sequence is listed below the schematic diagram (SEQ ID NO: 271). Figure B shows a Western blot analysis of a periplasmic extract of E. coli SDB1 expressing the EPA fusion protein containing PglS, CPS8, and the iGTΔ3-4 ComP glycosylated fragment incorporated between residues Glu548 and Gly549. The protein reacts with an anti-EPA antiserum probe using an anti-EPA antibody. [Figure 19] A, B, and C show Western blot analysis of EPA fusion proteins purified by Ni affinity chromatography, containing an iGTΔ3-4 ComP glycosylated fragment integrated between EPA residues Glu548 and Gly549. The fusion proteins were purified from SDB cells expressing CPS8 glycan in the presence (+) or absence (-) of PglS. A shows the protein reacting with anti-His antiserum. B shows the protein reacting with anti-CPS8 antiserum. C shows a combination of Figures 19A and 19B. Protein mass markers (kDa) are shown on the left side of the panels in Figures 19A–C. [Figure 20] List the ComP orthologous amino acid sequences. Predicted glycosylation sites are shown in bold. [Figure 21]The following lists the ComP Δ28 ortholog amino acid sequences in which the amino acid corresponding to the 28th N-terminal amino acid of ComPADP1:AAC45886.1 has been removed. Predicted glycosylation sites are shown in bold. [Figure 22] This shows the alignment of the ComP sequence region containing the serine (S) residue (outlined) corresponding to the serine residue at position 82 of ComP110264 (SEQ ID NO: 201), and also corresponds to the serine residue at position 84 of ComPADP1 (SEQ ID NO: 202). [Figure 23] A-D show characterizations of 13 TfpM orthologues from species of the Moraxellaceae family. Figure 23A) Phylogenetic trees of 20 TfpM orthologues, including typical TfpO, PglL, and PglS from Pseudomonas, Neisseria, and Acinetobacter, respectively. Branch confidence is indicated in red. OTase / pyrin pairs marked with an asterisk were cloned and tested in bioconjugation experiments. Figure 23B) Figure showing the design of the EPA-pyrin fusion protein and TfpM construct. Colored arrows indicate genes. Gene expression was driven by an IPTG-inducible tac promoter with a lacO operator (tac1O). The T2 terminator is marked with a black hairpin structure. Figures 23C) and 23D) Anti-EPA Western blots of whole-cell E. coli extracts expressing different EPA-pyrin carrier open reading frames and tfpM genes. Panel D is the same image as Panel C, but with increased exposure. Whole cell extracts loaded into each lane were normalized by OD600. H286A represents the M. osloensis TfpM site-directed OTase active site mutant, "g0" represents unglycosylated EPA-pyrine, and "gn" represents the CPS8-glycosylated EPA-pyrine protein. The mass of the reference protein is marked in kDa on the left side of the Western blot. [Figure 24]A phylogenetic tree showing the relative distances of TfpM, PilO, PglL, and PglS orthologues is presented. The phylogenetic tree was generated using the phylogeny.fr server (located at phylogeny.fr / on the World Wide Web), with MUSCLE, PhyML, and TreeDyn used for sequence alignment, tree calculation, and image generation, respectively. [Figure 25] This phylogenetic tree shows the relative distances of TfpM-related pyrin-like proteins, selected PilA proteins from Neisseria and Pseudomonas, and ComP from A. soli CIP 110264. Red numbers indicate branch confidence. The phylogenetic tree was generated using the phylogeny.fr server (located at phylogeny.fr / on the World Wide Web), with MUSCLE, PhyML, and TreeDyn used for sequence alignment, tree calculation, and image generation, respectively. [Figure 26] Multiple sequence alignments of selected bacterial O-linked oligosaccharide transferases are shown. The alignments were generated using Clustal Omega with default settings, available at ebi.ac.uk / Tools / msa / clustalo / on the World Wide Web. N_menigitidis_MC58_PglL (SEQ ID NO: 105), A_baylyi_ADP1_PglS (SEQ ID NO: 106), P_aeruginosa_1244_TfPO (SEQ ID NO: 107), M_osloensis_1202_TfpM (SEQ ID NO: 56), A_nosocomialis_M2_TfpO (SEQ ID NO: 108). [Figure 27] This image shows anti-EPA whole-cell Western blots examining the glycosylation status of EPA-PilMoΔ28 fusions and EPA-PilMoΔ28 C-terminal Thr167 mutants. All lanes were normalized to the same OD600. The mass of the reference protein is marked in kDa next to the Western blot. [Figure 28]Figures A and B show targeted MS / MS analysis of the HexHexA-modified C-terminal EPA-PilMoΔ28 peptide 762FLPANCRGT770 (SEQ ID NO: 61). Figure 28A) Fragmentation of EThrcD made it possible to localize the HexHexA glycosylation event to the terminal residue Thr770. Figure 28B) Fragmentation of HCD allowed for the observation of multiple y ions bound only to the Hex residue, enabling confirmation of the peptide sequence and confirmation of the binding of the disaccharide HexHexA via the Hex monosaccharide. [Figure 29] A-G demonstrate that TfpMMo can convert various bacterial glycans into EPA-PilMoΔ28 fusion proteins. Figure 29A) shows the repeating unit structures of five bacterial glycans examined with TfpMMo. Bonds between sugar monomers are indicated in parentheses. Abbreviations for glycans used: CPS8, S. pneumoniae capsular polysaccharide type 8; GBSIII, Group B Streptococcus capsular polysaccharide type III; LT2, Salmonella enterica Group B serotype LT2 O-antigen; O16, E. coli serotype O16 O-antigen; O2a, Klebsiella pneumoniae serotype O2a O-; Unless otherwise noted, all sugars are pyranose type. Abbreviations used: Glc, glucose; Gal, galactose; Galf, galacofuranose; Rha, rhamnose; GlcNAc, N-acetylglucosamine; Abe, abecose; NeuNAc, N-acetylneuraminic acid, sialic acid. Figures 29B)-29F) Antiglycan Western blots with partially purified TfpMMo-derived bioconjugates. Figure 29B) Anti-CPS8. Figure 29C) Anti-O16. Figure 29D) Anti-LT2. Figure 29E) Anti-O2a. Figure 29F) Anti-GBSIII. Figure 29G) Anti-EPA. The + / - labels in panels B)-G) indicate whether the sample was incubated with proteinase K (+) or without proteinase K (-) before SDS-PAGE separation. The mass of the reference protein is marked in kDa next to the Western blot. [Figure 30]A and B show EPA-fusion pyrin variants cleaved by the TfpMMo glycosylated, which are as small as three amino acids. Figure 30A) Sequence of EPA-fusion PilMo fragments examined for bioconjugation with TfpMMo. Blue letters indicate the C-terminal residue of EPA (i.e., EDLK; SEQ ID NO: 132). Underlined residues indicate glycine linkers positioned between the EPA and pyrin sequences. Figure 30B) Anti-EPA Western blot of whole cell extracts expressing the cleaved pyrin variant, CPS8, and TfpMMo. The calculated EPA-PilMoΔ28 mass is 80.3 kDa, and the mass of the cleaved variant ranges from 67.1 to 69.0 kDa. All lanes were normalized to the same OD600. The mass of the reference protein is marked in kDA on the left side of the Western blot. "g0" indicates unglycosylated cleaved EPA-pyrine, and "gn" indicates glycosylated EPA-pyrine protein. Unglycosylated EPA-PilMoΔ28 operates around 75 kDa. Pil20 (SEQ ID NO: 60). Pil15 (SEQ ID NO: 109). Pil13 (SEQ ID NO: 110). GGGG+Pil10 is Pil10L (SEQ ID NO: 111). Pil10 (SEQ ID NO: 112). Pil7 (SEQ ID NO: 113). Pil6 (SEQ ID NO: 114). Pil5 (SEQ ID NO: 115). Pil4 (SEQ ID NO: 116). Pil3 (SEQ ID NO: 117). EDLK+Pil2 (SEQ ID NO: 118). EDLKGGGG+Pil20 (SEQ ID NO: 122). EDLK+Pil15 (SEQ ID NO: 123). EDLK+Pil13 (SEQ ID NO: 124). EDLK+Pil10L (SEQ ID NO: 125). EDLK+Pil10 (SEQ ID NO: 126). EDLK+Pil7 (SEQ ID NO: 127). EDLK+Pil6 (SEQ ID NO: 128). EDLK+Pil5 (SEQ ID NO: 129). EDLK+Pil4 (SEQ ID NO: 130). EDLK+Pil3 (SEQ ID NO: 131). [Figure 31]This document shows multiple sequence alignments of selected pyrin proteins. The access numbers for these proteins are listed in the text. The alignments were generated using Clustal Omega with default settings, available at ebi.ac.uk / Tools / msa / clustalo / on the World Wide Web. P_aeruginosa_1244_PilA (SEQ ID NO: 119). N_menigitidis_M2_PilA (SEQ ID NO: 120). A_junii_65_pyrin (SEQ ID NO: 97). A_CIP102143_pyrin (SEQ ID NO: 88). A_CIP102637_pyrin (SEQ ID NO: 100). A_YZS-X1-1_pyrin (SEQ ID NO: 98). A_soli_110264_ComP (SEQ ID NO: 121). A_YH01026_pyrin (SEQ ID NO: 87). M_osloensis_1202_pyrin (SEQ ID NO: 57). A_junii_TUM15069_pyrine (sequence number 84). [Figure 32] Figures A-F show that purified TfpMMo-derived GBSIII bioconjugates induce a robust IgG immune response in mice. Figure 32A) Western blot of purified GBSIII-291 bioconjugate, anti-EPA channel. Figure 32B) Anti-GBSIII. Figure 32C) Combined image of A and B. Figure 32D) Coomassie stain of purified GBSIII-291 bioconjugate. Figure 32E) MS1 spectrum of intact purified GBSIII-291 bioconjugate. Protein 291 (EPA-Pil20) has a theoretical mass of 69,582.19 Da. The GBSIII-291 bioconjugate is observed in multiple mass-increasing states separated by approximately 980 Da, corresponding to the calculated mass of the GBSIII glycan repeating unit. Figure 32F) GBSIII-specific IgG dynamics during the immunization process, measured by ELISA and converted to ng / mL IgG using a standard IgG curve. **P<0.01. [Figure 33]Figures A and B show glycosylation of EPA constructs containing sequencens from different O-linked oligosaccharide transferase systems. Figure 33A) Diagram of plasmid-based operons expressing EPA with PglS or TfpM-specific sequencens. The internal glycotag ("iGT") is a 23-amino acid fragment derived from ComP110264 inserted between EPA Ala489-Arg490 and / or Glu548-Gly549. Figure 33B) Anti-EPA Western blot of SDB1 periplasm extract expressing one of four constructs and the E. coli O16 O antigen. Load per lane normalized to OD600. "g0" indicates unglycosylated EPA carrier protein, and single or double glycosylated EPA protein. The mass of the reference protein is marked in kDa on the left side of the Western blot. [Modes for carrying out the invention]

[0026] This disclosure provides a description of compositions and methods used to produce complex carbohydrate polypeptides using a glycosidic bond-forming enzyme, wherein the polypeptide can spontaneously form an isopeptide bond using a second partner polypeptide containing a polypeptide tag, and also provides a description of the use thereof.

[0027] definition It should be noted that the terms “a” or “an” refer to one or more entities; for example, “a polysaccharide” is understood to mean one or more polysaccharides. Therefore, in this specification, the terms “a” (or “an”), “one or more”, and “at least one” may be used interchangeably.

[0028] Furthermore, as used herein, “and / or” is construed to mean that each of the specified features or components is specifically disclosed, with or without the other. Accordingly, as used herein in phrases such as “A and / or B,” the term “and / or” is intended to include “A and B,” “A or B,” “A” (alone), and “B” (alone). Similarly, as used in phrases such as “A, B, and / or C,” the term “and / or” is intended to include each of the following embodiments: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).

[0029] Whenever an aspect is described using the words “comprising” or “comprises,” it is understood that other similar aspects described using terms such as “consisting of,” “consists of,” “consisting essentially of,” and / or “consists essentially of” are also provided.

[0030] Unless otherwise defined, technical and scientific terms used herein have the same meanings as those generally understood by those skilled in the art in which this disclosure relates.

[0031] A numerical range includes the numerical values ​​that define the range. Unless otherwise specified, if a list of values ​​such as 1, 2, 3, or 4 is provided, even if it is not explicitly identified as "and any range between them," the disclosure shall specifically include any range between the values ​​such as 1-3, 1-4, 2-4, etc.

[0032] The headings provided herein are for the sole purpose of facilitating reference and do not limit the various aspects or aspects of disclosure that can be obtained by referring to the entire specification.

[0033] As used herein, the term “polypeptide” is intended to encompass both the singular and plural forms “polypeptide” and refers to a molecule composed of monomers (amino acids) linearly linked by amide bonds (also called peptide bonds). The term “polypeptide” refers to any chain or more of two or more amino acids and does not refer to a specific length of product. Thus, peptides, dipeptides, tripeptides, oligopeptides, “proteins,” “amino acid chains,” or other terms used to refer to chains of two or more amino acids are included in the definition of “polypeptide,” and the term “polypeptide” can be used in place of or interchangeably with these terms. The term “polypeptide” is also intended to refer to the product of post-expression modifications of a polypeptide, which include, but are not limited to, glycosylation, acetylation, phosphorylation, amidation, derivatization with known protecting / blocking groups, cleavage by proteolysis, or modification with non-standard amino acids. Polypeptides may originate from natural biological sources or be produced by recombinant technology, but are not necessarily translated from a specified nucleic acid sequence. They may be produced by any method, including chemical synthesis.

[0034] As used herein, "protein" may refer to a single polypeptide, i.e., a single amino acid chain as defined above, but may also refer to two or more polypeptides linked together, for example, by disulfide bonds, hydrogen bonds, or hydrophobic interactions, to form a multimeric protein.

[0035] "Isolated" polypeptides or their fragments, variants, or derivatives mean polypeptides that do not exist in the natural environment and do not require a special level of purification. For example, isolated polypeptides can be removed from their original or natural environment. Recombinant-generated polypeptides and proteins expressed in host cells are considered isolated as disclosed herein, as are recombinant polypeptides that have been separated, fractionated, or partially or substantially purified by any appropriate technique.

[0036] A “vector” (as used herein interchangeably with “plasmid”) is a nucleic acid molecule that is introduced into a host cell to produce a transformed host cell. A vector may contain a nucleic acid sequence, such as an origin of replication, that enables its replication in the host cell. A vector can encode and express proteins. A vector may also contain one or more selectable marker genes and other genetic elements known in the art.

[0037] A “transformed” cell, or “host” cell, is a cell into which nucleic acid molecules have been introduced by molecular biology techniques. As used herein, the term transformation encompasses techniques that can introduce nucleic acid molecules into such cells, including transfection with viral vectors, transformation with plasmid vectors, electroporation, lipofection, and introduction of naked DNA by particle gun acceleration. A transformed cell or host cell may be a bacterial cell or a eukaryotic cell.

[0038] As used herein, the term “expression” refers to the process by which a gene produces a biochemical substance, such as a polypeptide. This process includes gene knockdown, as well as any expression of a functional presence of a gene within a cell, including but not limited to both transient and stable expression. This includes, but is not limited to, the transcription of a gene into messenger RNA (mRNA) and the translation of such mRNA into polypeptides. If the final product of interest is a biochemical substance, expression includes the production of that biochemical substance and its precursors. Gene expression produces a “gene product.” As used herein, a gene product is either a nucleic acid, such as messenger RNA, produced by the transcription of a gene, or a polypeptide translated from the transcript. Gene products as described herein further include nucleic acids that have undergone post-transcriptional modifications, e.g., polyadenylation, or polypeptides that have undergone post-translational modifications, e.g., methylation, glycosylation, lipid addition, binding to other protein subunits, or cleavage by proteolysis.

[0039] As used herein, the terms “to treat,” “to cure,” or “to treat” (for example, “to treat a subject”) mean reducing the likelihood of a disease condition, reducing the occurrence of symptoms of a disease, for example, to the extent that the subject’s survival rate is increased or discomfort is reduced. For example, treatment can mean the ability of a treatment administered to a subject to alleviate the symptoms, signs, or causes of a disease. Treatment can also mean alleviating or reducing at least one clinical symptom and / or inhibiting or delaying the progression of a condition and / or preventing or delaying the onset of a disease or illness.

[0040] "Subject," "individual," "animal," "patient," or "mammal" means any subject, especially mammalian subjects, for which diagnosis, prognosis, or treatment is desired. Mammalian subjects include humans, livestock, farm animals, sport animals, and zoo animals, such as humans, non-human primates, dogs, cats, guinea pigs, rabbits, rats, mice, horses, cattle, llamas, bears, etc.

[0041] The terms "pharmaceutical composition" or "therapeutic composition" refer to a formulation in which the biological activity of the active ingredient is effective and which does not contain additional ingredients that are unacceptably toxic to the subject to which the composition is administered. Such a composition may be sterile.

[0042] As used herein, “sugar” is a general term used to refer to carbohydrate molecules of any size, including, but not limited to, monosaccharides, disaccharides, trisaccharides, tetrasaccharides, pentasaccharides, hexasaccharides, heptasaccharides, oligosaccharides, or polysaccharides.

[0043] As used herein, “glycosidic bond” is a covalent bond between a sugar and another organic molecule, including but not limited to another sugar, protein, lipid, or nucleic acid.

[0044] As used herein, “glycosylated fragment” or “sequon” is a sequence of amino acids in a protein that functions as a sugar recognition and binding site, which is covalently converted to the protein by a glycosyltransferase or oligosaccharide transferase.

[0045] As used herein, the term “translationally fused” may mean direct binding to (e.g., a carrier protein, N-terminal leader sequence, C-terminal tag, etc. of the fusion protein) or indirect binding via an amino acid linker.

[0046] Complex carbohydrate polypeptide having an isopeptide bond with a second polypeptide partner A polypeptide pair comprising a polypeptide tag and a binding partner (e.g., another polypeptide tag) is provided herein, the polypeptide tag and the binding partner can bind to each other via the spontaneous formation of an isopeptide bond between one reactive residue contained in the binding partner and another reactive residue contained in the polypeptide tag. As used herein, when one component of a polypeptide pair is referred to, the other may be referred to as its partner.

[0047] Fusion protein Generally, one component of a polypeptide pair of the present disclosure, including what is referred to as a first polypeptide tag, may be a fusion protein. Thus, certain embodiments of the present disclosure provide a fusion protein comprising (i) a glycosylated fragment and (ii) a first polypeptide tag, the first polypeptide tag being capable of spontaneously forming an isopeptide bond with a second polypeptide tag binding partner. To avoid misunderstanding, a fusion protein comprising a glycosylated fragment may comprise a fragment comprising a glycosylation site (also referred to herein as a “sequon”) of a glycosylated protein, and in other embodiments, it may comprise a glycosylated full-length protein, the glycosylated fragment then comprising the glycosylated fragment. Numerous representative examples of glycosylated proteins and sequons, as well as related enzymes, that can be used in the compositions and methods of the present disclosure are provided in detail elsewhere herein. In certain embodiments, the fusion protein comprises a carrier protein. For example, in certain embodiments, the carrier protein may be Escherichia coli maltose-binding protein (MPB), Pseudomonas aeruginosa exotoxin A (EPA), Pseudomonas aeruginosa PcrV, CRM197, Haemophilus influenzae protein D, cholera toxin B subunit, tetanus toxin, or any fragment thereof. In certain embodiments, the glycosylated fragment may be as short as 3 amino acids. For example, TfpM OTase can recognize a 3-amino acid sequence. In certain embodiments, the glycosylated fragment may be longer, containing a full-length or nearly full-length protein. Thus, in certain embodiments, the glycosylated fragment may be at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, or 40 amino acids long.In certain embodiments, the glycosylation is not the full-length glycosylated protein, but a shorter fragment thereof, and therefore less than or equal to 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, 40, 50, 60, 80, or 100 amino acids in length. Thus, in certain embodiments, the glycosylated fragment is between any of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 24 amino acid lengths and any of 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, or 25 amino acid lengths. In certain embodiments, the glycosylated fragment is any of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, or 40 amino acid lengths, and any of 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, 40, or 50 amino acid lengths. In certain embodiments, the glycosylated fragment is any of the following amino acid lengths: 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, 40, 50, 60, or 80 amino acids.

[0048] In certain embodiments, the fusion protein is a complex carbohydrate containing sugars covalently bonded to the fusion protein via glycosylation sites / residues of a glycosylated fragment (sequon). In certain embodiments, the complex carbohydrate is immunogenic. Those skilled in the art will recognize that the sugars can be covalently bonded to the glycosylated fragment, for example, via N-bonds, O-bonds, or C-bonds.

[0049] In certain embodiments, the first polypeptide tag is translationally fused at the N-terminus of the fusion protein. In certain embodiments, the first polypeptide tag is translationally fused at the C-terminus of the fusion protein. In certain embodiments, the first polypeptide tag is translationally fused internally within the fusion protein. In certain embodiments, the first polypeptide tag is translationally fused internally within the sequence of the carrier protein.

[0050] In certain embodiments, the glycosylated fragment is translationally fused at the N-terminus of the fusion protein. In certain embodiments, the glycosylated fragment is translationally fused at the C-terminus of the fusion protein. In certain embodiments, the glycosylated fragment is translationally fused internally within the fusion protein. In certain embodiments, the glycosylated fragment is translationally fused internally within the sequence of the carrier protein.

[0051] As can be understood from the entire disclosure, by fusing internally within the fusion protein, glycosylated fragments, polypeptide tags, etc., are not located at the C-terminus or N-terminus of the fusion protein and do not contain any signal / reader sequences, purified tags (e.g., His-Tag), etc. For example, as follows: N-terminus Glycosylated fragment-carrier protein Not internal, but at the N-terminus Leader sequence-glycosylated fragment-carrier protein C-terminus Carrier protein-glycosylated fragment Not the internal terminal, but the C-terminal terminal. Carrier protein-glycosylated fragment-His-Tag internal Leader sequence - Carrier protein 1 - Glycosylated fragment - Carrier protein 2 - His-Tag internal Leader sequence-carrier protein 1-glycosylated fragment-carrier protein 1-His-Tag

[0052] As shown above, in certain embodiments of internal arrangement, glycosylated fragments, polypeptide tags, etc., may be positioned (translationally fused) between separate carrier proteins (even if they are of the same type). In certain embodiments of internal arrangement, glycosylated fragments, polypeptide tags, etc., may be positioned (translationally fused) internally within the sequence of a single carrier protein.

[0053] While not limited to any specific sequence, representative examples of a pair of polypeptide tag partners (generally referred to herein as the first polypeptide tag) include SpyTag (SEQ ID NO: 416), SpyTag002 (SEQ ID NO: 417), SpyTag003 (SEQ ID NO: 418), or DogTag (SEQ ID NO: 419). In certain embodiments, SpyTag, Spytag002, or Spytag003 are translationally fused at the N-terminus of the fusion protein (e.g., Figure 2). In certain embodiments, SpyTag, Spytag002, or Spytag003 are translationally fused at the C-terminus of the fusion protein (e.g., Figure 5). In certain embodiments, DogTag is translationally fused internally within the fusion protein.

[0054] Representative examples of the fusion proteins of this disclosure are shown in Figures 2, 5, 7, 8, and 9. For example, EPA-Spytag-v1 (SEQ ID NO: 427), EPA-Spytag-v2 (SEQ ID NO: 428), and EPA-Spytag-v3 (SEQ ID NO: 429). In certain embodiments, an amino acid linker sequence is translationally inserted between the components of the fusion protein (e.g., a signal peptide, a polypeptide tag sequence, a glycosylated fragment, a carrier protein, a histidine tag, etc.). In certain embodiments, the amino acid linker sequence is GGS, GGGGGG (SEQ ID NO: 430), GGGGGGGG (SEQ ID NO: 431), GGGGS (SEQ ID NO: 432), EAAAK (SEQ ID NO: 433), PAPAPPAPAP (SEQ ID NO: 434), EAAAKEAAAK (SEQ ID NO: 435), GGGGSPAPAP (SEQ ID NO: 436), GGGGSGGGGS (SEQ ID NO: 437), or EAAAKGGGGS (SEQ ID NO: 438). Accordingly, certain embodiments include EPA-Spytag-v1 of SEQ ID NO: 427, EPA-Spytag-v2 of SEQ ID NO: 428, or EPA-Spytag-v3 of SEQ ID NO: 429, which translateably insert one or more amino acid linkers as described above between the components of a fusion protein. For example, EPA-Spytag-v1 of SEQ ID NO: 427, in which the amino acid linker SSG is translatedably inserted after the polypeptide tag. For example, EPA-Spytag-v1 of SEQ ID NO: 440, which has the amino acid linker EAAAKEAAAK (SEQ ID NO: 435) translatedly inserted after the polypeptide tag. In certain embodiments, other amino acid linkers, such as but not limited to those disclosed herein, may also be arranged.

[0055] CompGlycosylated Fragments In certain embodiments, the glycosylated fragment is a ComP glycosylated fragment, although it is not limited to any specific glycosylated sequence. The sequence containing the glycosylated fragment of ComP may be a full-length ComP protein. In certain embodiments, the ComP glycosylated fragment contains or consists of the amino acid sequence CTGVTQIASGASAATTNVASAQC (SEQ ID NO: 412), or a fragment of SEQ ID NO: 412 that contains the amino acid ASA at positions 11-13. In certain embodiments, the ComP glycosylated fragment contains or consists of the amino acid sequence CTGVTQIASGASAATTNVASAQC (SEQ ID NO: 412), or a fragment of SEQ ID NO: 412 that is at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23 amino acids long and contains the amino acid ASA at positions 11-13 of SEQ ID NO: 412. In certain embodiments, the ComP glycosylated fragment comprises or consists of the amino acid sequence CTGVTQIASGASAATTNVASAQC (SEQ ID NO: 412), or a fragment having at least 10 amino acids in length, with the amino acid ASA at positions 11-13 of SEQ ID NO: 412. In certain embodiments, the ComP glycosylated fragment comprises or consists of the amino acid sequence CTGVTQIASGASAATTNVASAQC (SEQ ID NO: 412), or a fragment having at least 11 amino acids in length, with the amino acid ASA at positions 11-13 of SEQ ID NO: 412.

[0056] In certain embodiments, the ComP glycosylated fragment comprises or consists of a variant of SEQ ID NO: 412 having an amino acid ASA at positions 11-13 of SEQ ID NO: 412 and having 1, 2, 3, 4, 5, or 6 amino acid substitutions, additions, and / or deletions. Those skilled in the art will understand that in any of the embodiments of the glycosylated fragments of this disclosure, for example, if additions occur at specific positions and deletions occur at different positions, the net effect on the sequence length is zero. Furthermore, the additions or deletions may be consecutive and / or discontinuous. Also, in certain embodiments, the substitutions may be conservative amino acid substitutions. In certain embodiments, the ComP glycosylated fragment comprises or consists of a variant of SEQ ID NO: 412 having an amino acid ASA at positions 11-13 of SEQ ID NO: 412 and having a cumulative total of 1, 2, 3, 4, 5, or 6 amino acid substitutions, additions, and / or deletions.

[0057] In certain embodiments, the ComP glycosylated fragment comprises or consists of a variant of SEQ ID NO: 412 having an amino acid ASA at positions 11-13 of SEQ ID NO: 412 and having 1, 2, 3, 4, 5, or 6 amino acid substitutions and / or additions.

[0058] In certain embodiments, the ComP glycosylated fragment comprises or consists of a variant of SEQ ID NO: 412 having an amino acid ASA at positions 11-13 of SEQ ID NO: 412 and having 1, 2, 3, 4, 5, or 6 amino acid substitutions and / or deletions.

[0059] In certain embodiments, the ComP glycosylated fragment comprises or consists of a variant of SEQ ID NO: 412 having an amino acid ASA at positions 11-13 of SEQ ID NO: 412 and having 1, 2, 3, 4, 5, or 6 amino acid additions and / or deletions.

[0060] In certain embodiments, the ComP glycosylated fragment comprises or consists of a variant of SEQ ID NO: 412 having an amino acid ASA at positions 11-13 of SEQ ID NO: 412 and having 1, 2, 3, 4, 5, or 6 amino acid substitutions. In certain embodiments, the ComP glycosylated fragment comprises or consists of a variant of SEQ ID NO: 412 having an amino acid ASA at positions 11-13 of SEQ ID NO: 412 and having 1, 2, 3, 4, 5, or 6 amino acid additions. In certain embodiments, the ComP glycosylated fragment comprises or consists of a variant of SEQ ID NO: 412 having an amino acid ASA at positions 11-13 of SEQ ID NO: 412 and having 1, 2, 3, 4, 5, or 6 amino acid deletions.

[0061] In certain embodiments, the ComP glycosylated fragment comprises or consists of any of the further ComP glycosylated fragment sequences described elsewhere in this specification.

[0062] TfpM-related pyring glycosylated fragments In certain embodiments, the glycosylated fragment is a TfpM-related pyrin glycosylated fragment, although it is not limited to any specific glycosylated sequence. The sequence containing the TfpM-related pyrin glycosylated fragment may be a full-length TfpM-related pyrin protein. In certain embodiments, the TfpM-related pyrin glycosylated fragment contains or consists of a PilMo pyrin disulfide loop region (SEQ ID NO: 413), or a fragment containing at least the last three amino acids (i.e., RGT) from the C-terminus of TfpM-related pyrin. In certain embodiments, the TfpM-related pyring glycosylation fragment comprises or consists of a PilMo pyring disulfide loop region (SEQ ID NO: 413), or a fragment having a length of at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23 amino acids, and containing at least the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-related pyring.

[0063] In certain embodiments, the TfpM-related pyring glycosylation fragment comprises or consists of a variant of the PilMo pyring disulfide loop region (SEQ ID NO: 413) comprising the last three amino acids from the C-terminus of TfpM-related pyring (i.e., RGT) and having 1, 2, 3, 4, 5, or 6 amino acid substitutions, additions, and / or deletions.

[0064] In certain embodiments, the TfpM-related pyring glycosylation fragment comprises or consists of a variant of the PilMo pyring disulfide loop region (SEQ ID NO: 413) having 1, 2, 3, 4, 5, or 6 amino acid substitutions and / or additions, comprising the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-related pyring.

[0065] In certain embodiments, the TfpM-related pyring glycosylation fragment comprises or consists of a variant of the PilMo pyring disulfide loop region (SEQ ID NO: 413) having 1, 2, 3, 4, 5, or 6 amino acid substitutions and / or deletions, comprising the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-related pyring.

[0066] In certain embodiments, the TfpM-related pyring glycosylation fragment comprises or consists of a variant of the PilMo pyring disulfide loop region (SEQ ID NO: 413) having the last three amino acids (i.e., RGT) from the C-terminus of TfpM-related pyring and having the addition and / or deletion of 1, 2, 3, 4, 5, or 6 amino acids.

[0067] In certain embodiments, the TfpM-related pyring glycosylation fragment comprises or consists of a variant of the PilMo pyring disulfide loop region (SEQ ID NO: 413) having 1, 2, 3, 4, 5, or 6 amino acid substitutions, and including the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-related pyring. In certain embodiments, the TfpM-related pyring glycosylated fragment comprises or consists of a variant of the PilMo pyring disulfide loop region (SEQ ID NO: 413) having 1, 2, 3, 4, 5, or 6 amino acid deletions, and containing the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-related pyring.

[0068] In certain embodiments, the TfpM-related pilling glycosylation fragment comprises or consists of one of the further TfpM-related pilling glycosylation fragment sequences described elsewhere in this specification.

[0069] PilE glycosylated fragment In certain embodiments, the glycosylated fragment is a PilE glycosylated fragment, although it is not limited to any specific glycosylated sequence. The sequence containing the PilE glycosylated fragment may be a full-length PilE protein. In certain embodiments, the PilE glycosylated fragment comprises or consists of the amino acid SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO: 414), or the fragment containing at least the amino acid WPGNNTSAGV (SEQ ID NO: 439) at positions 13-22 of SEQ ID NO: 414. In a particular embodiment, the PilE glycosylated fragment comprises or consists of the amino acid SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO: 414), or a fragment having at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, or 29 amino acids in length, and containing at least the amino acid WPGNNTSAGV (SEQ ID NO: 439) at positions 13-22 of SEQ ID NO: 414.

[0070] In certain embodiments, the PilE glycosylated fragment comprises or consists of a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO: 414) having at least the amino acid WPGNNTSAGV (SEQ ID NO: 439) at positions 13-22 of SEQ ID NO: 414, and having 1, 2, 3, 4, 5, or 6 amino acid substitutions, additions, and / or deletions.

[0071] In certain embodiments, the PilE glycosylated fragment comprises or consists of a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO: 414) having at least the amino acid WPGNNTSAGV (SEQ ID NO: 439) at positions 13-22 of SEQ ID NO: 414 and having 1, 2, 3, 4, 5, or 6 amino acid substitutions and / or additions.

[0072] In certain embodiments, the PilE glycosylated fragment comprises or consists of a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO: 414) having at least the amino acid WPGNNTSAGV (SEQ ID NO: 439) at positions 13-22 of SEQ ID NO: 414 and having 1, 2, 3, 4, 5, or 6 amino acid substitutions and / or deletions.

[0073] In certain embodiments, the PilE glycosylated fragment comprises or consists of a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO: 414) having at least the amino acid WPGNNTSAGV (SEQ ID NO: 439) at positions 13-22 of SEQ ID NO: 414, and having 1, 2, 3, 4, 5, or 6 amino acid additions and / or deletions.

[0074] In certain embodiments, the PilE glycosylated fragment includes or comprises a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK(SEQ ID NO: 414) having at least the amino acid WPGNNTSAGV(SEQ ID NO: 439) at positions 13-22 of SEQ ID NO: 414 and having 1, 2, 3, 4, 5, or 6 amino acid substitutions. In certain embodiments, the PilE glycosylated fragment comprises or consists of a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO: 414) having at least one amino acid WPGNNTSAGV (SEQ ID NO: 439) at positions 13-22 of SEQ ID NO: 414 and having one, two, three, four, five, or six amino acid deletions.

[0075] PglB glycosylated fragment In certain embodiments, the glycosylated fragment is a PglB glycosylated fragment, although it is not limited to any specific glycosylated sequence. A sequence containing a PglB glycosylated fragment may be a full-length PglB protein. In certain embodiments, the PglB glycosylated fragment contains or comprises a consensus motif amino acid sequence X1X2N X3X4, where X1 is D or E, X2 is any amino acid except proline, X3 is any amino acid except proline, and X4 is S or T.

[0076] PilA glycosylated fragment In certain embodiments, the glycosylated fragment is a PilA glycosylated fragment, although it is not limited to any specific glycosylated sequence. The sequence containing the PilA glycosylated fragment may be a full-length PilA protein. In certain embodiments, the PilA glycosylated fragment contains or consists of the PilA pyring disulfide loop region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO: 415), or a fragment of PilA containing at least the last three amino acids (i.e., PKS) from the C-terminus. In certain embodiments, the PilA glycosylated fragment comprises or consists of the PilA pyring disulfide loop region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO: 415), or the fragment having at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 21 amino acid lengths and containing at least the last three amino acids (i.e., PKS) from the C-terminus of PilA.

[0077] In certain embodiments, the PilA glycosylated fragment comprises or consists of a variant of the PilA pyringisulfide loop region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO: 415) having at least the last three amino acids (i.e., PKS) from the C-terminus of PilA and having 1, 2, 3, 4, 5, or 6 amino acid substitutions, additions, and / or deletions.

[0078] In certain embodiments, the PilA glycosylated fragment comprises or consists of a variant of the PilA pyringisulfide loop region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO: 415) comprising at least the last three amino acids (i.e., PKS) from the C-terminus of PilA and having 1, 2, 3, 4, 5, or 6 amino acid substitutions and / or additions.

[0079] In certain embodiments, the PilA glycosylated fragment comprises or consists of a variant of the PilA pyringisulfide loop region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO: 415) having at least the last three amino acids (i.e., PKS) from the C-terminus of PilA and having 1, 2, 3, 4, 5, or 6 amino acid substitutions and / or deletions.

[0080] In certain embodiments, the PilA glycosylated fragment comprises or consists of a variant of the PilA pyringisulfide loop region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO: 415) having at least the last three amino acids (i.e., PKS) from the C-terminus of PilA and having the addition and / or deletion of 1, 2, 3, 4, 5, or 6 amino acids.

[0081] In certain embodiments, the PilA glycosylated fragment comprises or consists of a variant of the PilA pyring disulfide loop region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO: 415) having at least the last three amino acids (i.e., PKS) from the C-terminus of PilA and having 1, 2, 3, 4, 5, or 6 amino acid substitutions. In certain embodiments, the PilA glycosylated fragment comprises or consists of a variant of the PilA pyring disulfide loop region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO: 415) having at least the last three amino acids (i.e., PKS) from the C-terminus of PilA and having 1, 2, 3, 4, 5, or 6 amino acid deletions.

[0082] PilA Pa5196 glycosylated fragment In certain embodiments, the glycosylated fragment is a PilA_Pa5196-associated pyrin glycosylated fragment, although it is not limited to any specific glycosylated sequence. The sequence containing the PilA_Pa5196-associated pyrin glycosylated fragment may be a full-length PilA_Pa5196-associated pyrin protein. In certain embodiments, the PilA_Pa5196-associated pyrin glycosylated fragment contains or consists of chains 1 and 2 of the antiparallel beta-sheet domain of PilA_Pa5196 GKYSSVDSTIASGYPNGQITVTMTQG (SEQ ID NO: 426), or fragments thereof. In certain embodiments, the PilA_Pa5196-related pilling glycosylation fragment comprises or consists of chains 1 and 2 of the antiparallel beta-sheet domain of PilA_Pa5196 GKYSSVDSTIASGYPNGQITVTMTQG (SEQ ID NO: 426), or fragments thereof having at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acid lengths.

[0083] In certain embodiments, the glycosylated fragment is a variant of the PilA_Pa5196-related pyring glycosylated fragment that is consistent with variants of other glycosylated fragments disclosed herein, having 1, 2, 3, 4, 5, or 6 amino acid substitutions, additions, and / or deletions.

[0084] STT3 glycosylated fragment In certain embodiments, the glycosylated fragment is an STT3 glycosylated fragment, although it is not limited to any specific glycosylated sequence. A sequence containing an STT3 glycosylated fragment may be a full-length STT3 protein. In certain embodiments, the STT3 glycosylated fragment contains or consists of a consensus motif amino acid sequence NXS / T, where X is any amino acid except proline, and S / T is serine (S) or threonine (T).

[0085] N-linked glycosyltransferase glycosylated fragment In certain embodiments, the glycosylated fragment is an N-linked glycosyltransferase glycosylated fragment, although it is not limited to any specific glycosylated sequence. In certain embodiments, the N-linked glycosyltransferase glycosylated fragment comprises or consists of a consensus motif amino acid sequence NXS / T, where X is any amino acid except proline, and S / T is serine (S) or threonine (T).

[0086] O-linked glycosyltransferase glycosylated fragment In certain embodiments, the glycosylated fragment is an O-linked glycosylase glycosylated fragment, although it is not limited to any specific glycosylated sequence. In certain embodiments, the O-linked glycosylase glycosylated fragment comprises or consists of a serine or threonine-rich repeat fragment from a serine-rich repeat (SRR) adhesin of streptococci or staphylococci bacteria. In certain embodiments, the O-linked glycosylase glycosylated fragment comprises or consists of a serine (S) or threonine (T)-rich repeat from adhesin GspB of Streptococcus gordonii.

[0087] Multiple glycosylated fragments Certain embodiments of this disclosure are shown in a fusion protein comprising two or more glycosylated fragments. For example, the fusion protein comprises at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 glycosylated fragments. In a particular embodiment, the fusion protein comprises any of 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 glycosylated fragments. In certain embodiments, at least one glycosylated fragment is located at the N-terminus of the fusion protein, and at least one glycosylated fragment is located internally within the fusion protein. In certain embodiments, at least one glycosylated fragment is located at the C-terminus of the fusion protein, and at least one glycosylated fragment is located internally within the fusion protein. In certain embodiments, at least two glycosylated fragments are located internally within the fusion protein. Furthermore, in certain embodiments, one glycosylated fragment is located at the N-terminus of the fusion protein, and one glycosylated fragment is located at the C-terminus of the fusion protein. In certain embodiments, two or more glycosylated fragments are identical. For example, a fusion protein having multiple ComP glycosylated fragments. In certain embodiments, at least one of the two or more glycosylated fragments is different, or each of the glycosylated fragments is different. For example, a fusion protein in which one glycosylated fragment is a ComP glycosylated fragment and one glycosylated fragment is a TfpM-associated pilling glycosylated fragment.

[0088] In certain embodiments, the fusion protein is a complex carbohydrate comprising two or more sugars covalently bonded to the fusion protein via two or more glycosylated fragments. In certain embodiments, the fusion protein comprises at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 covalently bonded sugars. In certain embodiments, the fusion protein comprises any of 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 covalently bonded sugars. In a particular embodiment, two or more sugars are the same. In a particular embodiment, at least one of the two or more sugars is different, or each of the sugars is different.

[0089] As can be understood, the ability to combine different fusion proteins within the same complex, as well as to create fusion proteins that have multiple glycosylated fragments, can be recognized by even more different glycosylation enzymes, and / or bind different sugars, enables the mixing, matching, and amplification of various immunogenic components.

[0090] In certain embodiments, the fusion protein includes a carrier protein. Various carrier proteins have been used in complex carbohydrate vaccines, all of which are intended herein. For example, in certain embodiments, the carrier protein is selected from the group consisting of Escherichia coli maltose-binding protein, Pseudomonas aeruginosa exotoxin A (EPA), Pseudomonas aeruginosa PcrV, CRM197, Haemophilus influenzae protein D, cholera toxin B subunit, or tetanus toxin, and any fragment thereof.

[0091] polypeptide vs Certain embodiments of this disclosure are shown in a composition comprising a polypeptide pair comprising a first polypeptide and a second polypeptide. The first polypeptide comprises a first polypeptide tag which is a binding partner of the second polypeptide to a second polypeptide tag. In certain embodiments, the first polypeptide is a fusion protein of this disclosure comprising a glycosylated fragment, as described in detail elsewhere in this specification. The second polypeptide comprises a second polypeptide tag binding partner of the first polypeptide to the first polypeptide tag. The first polypeptide can be bound to the second polypeptide via an isopeptide bond between the first polypeptide tag and the second polypeptide tag. In certain embodiments, the second polypeptide comprises a monomer polypeptide that can spontaneously multimerize / self-assemble into a higher-order multimeric structure, e.g., icosahedral or dodecahedral particles (e.g., similar to nanocages), virus-like particles (VLPs), or adenovirus vectors. In certain embodiments, the second polypeptide comprises an adenovirus capsid structure protein. In certain embodiments, the second polypeptide contains the coat protein of bacteriophage AP205. In certain embodiments, the second polypeptide contains a fragment of 2-keto-3-deoxy-phosphogluconate aldolase (i301). In certain embodiments, the second polypeptide contains a fragment of mutated 2-keto-3-deoxy-phosphogluconate aldolase (mi3). In certain embodiments, the polypeptide tag of the second polypeptide (second polypeptide tag) is SpyCatcher (SEQ ID NO: 420). In certain embodiments, the second polypeptide tag is SpyCatcher002 (SEQ ID NO: 421). In certain embodiments, the second polypeptide tag is SpyCatcher003 (SEQ ID NO: 422). In certain embodiments, the second polypeptide tag is DogCatcher (SEQ ID NO: 423).

[0092] The polypeptide tag of the second polypeptide may be located terminally (N-terminus or C-terminus) or internally on the second polypeptide. In certain embodiments, the second polypeptide tag is translationally fused at the N-terminus of the second polypeptide (e.g., Figure 2). In certain embodiments, the second polypeptide tag is translationally fused at the C-terminus of the second polypeptide. In certain embodiments, the second polypeptide tag is translationally fused internally within the second polypeptide.

[0093] In certain embodiments, the first polypeptide (e.g., the fusion protein of this disclosure) is a bioconjugate comprising a sugar covalently bonded to a glycosylated fragment of the first polypeptide.

[0094] In certain embodiments, the polypeptide pair composition is immunogenic, and for example, the first polypeptide is a bioconjugate containing a sugar covalently bonded to a glycosylated fragment of the first polypeptide. In certain embodiments, the polypeptide pair composition further comprises an adjuvant and / or excipient. Exemplary adjuvants may include, but are not limited to, alum (aluminum hydroxide gel or aluminum phosphate gel), squalene emulsion (e.g., MF59, AddaS03, or AddaVax), lipid A derivatives such as monophosphoryl lipid A (MPLA), or saponins (e.g., Quil-A). In certain embodiments, the polypeptide pair composition is a pharmaceutical and / or therapeutic composition. In certain embodiments, the polypeptide pair composition is a conjugate vaccine.

[0095] Certain embodiments provide a method for producing polypeptide pairs of the present disclosure. This may be done by contacting a first polypeptide (e.g., a fusion protein of the present disclosure) and a second polypeptide under conditions that allow the first polypeptide tag to spontaneously form an isopeptide bond with its corresponding second polypeptide tag-binding partner. In certain embodiments, the method further comprises glycosylation of the first polypeptide with a sugar before contact with the second polypeptide and isopeptide bond formation (e.g., Figure 1). In certain embodiments, the first polypeptide is glycosylated in vivo (e.g., in a host cell, e.g., in bacteria) before contact with the second polypeptide and isopeptide bond formation. In certain embodiments, the method comprises isolating / purifying the first polypeptide glycosylated in vivo before contact with the second polypeptide and isopeptide bond formation. Alternatively, in certain embodiments, the first polypeptide is glycosylated after contact with the second polypeptide and isopeptide bond formation.

[0096] complex A complex comprising two or more polypeptide pairs disclosed herein is provided in this disclosure. In certain embodiments, each complex comprises 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 75, 100, 125, 150, 175, 200, 225, 250 or more complexed polypeptide pairs of this disclosure. In a particular embodiment, each complex is any of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 75, 100, 125, 150, 175, 200, 225, or 250 complexing polypeptide pairs of the present disclosure. The complex comprises up to 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 75, 100, 125, 150, 175, 200, 225, 250, 300, 400, or 500 complex polypeptide pairs of the present disclosure. In certain embodiments, such complexes are self-assembled multimeric higher-order structures. In certain embodiments, such self-assembled multimeric higher-order structures are icosahedral or dodecahedral particles (e.g., similar to nanocages), virus-like particles, or adenovirus vectors.

[0097] Since the partner polypeptide is determined by the polypeptide tag of the first polypeptide, and not, for example, the glycosylated fragment or carrier protein of the fusion protein, the second polypeptide is not limited to simply partnering with one type of first polypeptide (and vice versa). In certain embodiments, all of the first polypeptides of the complex contain the same fusion protein. However, in certain embodiments, the first polypeptides may contain different fusion proteins. In certain embodiments, at least two, three, four, five, or more of the first polypeptides of the complex contain different fusion proteins. In certain embodiments, at least two of the first polypeptides of the complex contain different fusion proteins. In certain embodiments, two, three, four, five, or six of the first polypeptides of the complex contain different fusion proteins. In certain embodiments, all of the first polypeptides of the complex are different fusion proteins.

[0098] In certain embodiments, at least one first polypeptide of the complex is a bioconjugate containing a sugar covalently bonded to a glycosylated fragment of the first polypeptide. In certain embodiments, at least about 5%, 10%, 25%, 50%, 75%, 80%, 90%, 95%, 97%, 98%, or 99% of the first polypeptide of the complex is a bioconjugate. In certain embodiments, from any of about 5%, 10%, 25%, 50%, 75%, 80%, 90%, 95%, 97%, or 98% of the first polypeptide of the complex, any of about 10%, 25%, 50%, 75%, 80%, 90%, 95%, 97%, 98%, or 99% of the first polypeptide of the complex is a bioconjugate. In certain embodiments, about 100% or 100% of the first polypeptide of the complex is a bioconjugate.

[0099] In certain embodiments, two or more of the first polypeptides in the complex are bioconjugates containing covalently bonded sugars. The number of covalently bonded sugars can be numerous, depending on the number of first polypeptides in the complex and the number of first polypeptide / second polypeptide pairs in the complex and the number of sugars bonded to each first polypeptide. For example, AP205 VLP contains approximately 180 first polypeptide binding partners (e.g., SpyCatcher) per VLP. Theoretically, if 100% of the first polypeptides are isopeptide-bonded, this allows for 180 bioconjugates per VLP. And, as described elsewhere in this specification, each bioconjugate may be covalently bonded to multiple sugars. Mi3 has a low number, approximately 60 first polypeptide binding partners per NP. In a particular embodiment, the complex contains at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1,000, 1,500, 2,000, 2,500, or 5,000 covalently bonded sugars. In a particular embodiment, the complex is one of 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1,000, 1,500, 2,000, or 2,500 covalently bonded sugars. It contains any of the following covalent sugars: from any of the following: 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1,000, 1,500, 2,000, 2,500, or 5,000 covalent sugars.In a particular embodiment, the complex contains at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, or 750 covalently bonded sugars. In a particular embodiment, the complex comprises any of 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, or 500 covalently bonded sugars, up to any of 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, or 750 covalently bonded sugars. In certain embodiments, all the sugars bound to the complex are the same. In certain embodiments, at least two, three, four, five, or more sugars bound to the complex are different. In certain embodiments, at least two of the sugars bound to the complex are different. In certain embodiments, two, three, four, five, or six sugars bound to the complex are different from one another. In certain embodiments, each of the sugars bound to the complex is different.

[0100] In certain embodiments of the complex, for example, if the complex is bound to a sugar, the complex is immunogenic.

[0101] Certain embodiments provide a pharmaceutical and / or therapeutic composition comprising the complex of the present invention and an adjuvant and / or excipient. In certain embodiments, the complex is a conjugate vaccine.

[0102] Certain embodiments provide a method for constructing the complex of the present invention. This can be done by forming a self-assembled multimeric higher-order structure of the second polypeptide of the present disclosure (e.g., icosahedral or dodecahedral particles (e.g., similar to nanocages), virus-like particles (VLPs), or adenovirus vectors), and then contacting the first polypeptide of the present disclosure (e.g., a fusion protein containing a glycosylated fragment) and the second polypeptide under conditions that allow the first polypeptide tag to spontaneously form an isopeptide bond with the second polypeptide tag (e.g., Figure 1). This can also be done by contacting the first polypeptide and the second polypeptide under conditions that allow the first polypeptide tag to spontaneously form an isopeptide bond with the second polypeptide tag, and then forming a self-assembled multimeric higher-order structure of the second polypeptide. Such methods encompass a number of possibilities and combinations for glycosylation of the first polypeptide of the complex with sugars, all of which will be understood to be contemplated herein. For example, the first polypeptide can be glycosylated before an isopeptide bond is formed between the first polypeptide and the second polypeptide (e.g., Figure 1). The first polypeptide can be glycosylated after an isopeptide bond is formed between the first polypeptide and the second polypeptide. The first polypeptide can be glycosylated before it is incorporated into the higher-order structure of the polymer (e.g., Figure 1). And the first polypeptide can be glycosylated after it has been incorporated into the higher-order structure of the polymer.

[0103] Glycosylation The glycosylated residues, glycosylated sites, glycosylated fragments, sequencers, first polypeptide fusion proteins, complexes, etc., of this disclosure may be covalently bonded to a sugar by any of a number of glycosylation methods, including but not limited to the following exemplary examples. In certain embodiments, the sugar is converted to a fusion protein containing a glycosylated fragment by the action of N-linked oligosaccharide transferase (N-OTase), O-linked oligosaccharide transferase (O-OTase), N-linked glycosyltransferase (NGT), O-linked glycosyltransferase (OGT), and / or C-mannosyltransferase (CMT). In certain embodiments, the sugar is converted to a fusion protein containing a glycosyllatin fragment by the action of PglS OTase, TfpM OTase, PglL OTase, PglB OTase, TfpO / PilO OTase, STT3 OTase, TfpW glycosyltransferase, and / or AlgB OTase. For example, in a particular embodiment, the glycosylated fragment is a ComP glycosylated fragment glycosylated by PglS OTase. In a particular embodiment, the sugar is covalently bonded to the oxygen atom in the glycosylated fragment using PglS OTase. In a particular embodiment, the glycosylated fragment is a TfpM-associated pyring glycosylated fragment glycosylated by TfpM OTase. In a particular embodiment, the sugar is covalently bonded to the oxygen atom in the glycosylated fragment using TfpM OTase. In a particular embodiment, the glycosylated fragment is a PilE glycosylated fragment glycosylated by PglL OTase. In a particular embodiment, the sugar is covalently bonded to the oxygen atom in the glycosylated fragment using PglL OTase. In a particular embodiment, the glycosylated fragment is a PglB glycosylated fragment glycosylated by PglB OTase. In certain embodiments, the sugar is covalently bonded to a nitrogen atom in the glycosylated fragment using PglB OTase. In certain embodiments, the glycosylated fragment is a PilA glycosylated fragment glycosylated by TfpO or PilO OTase.In certain embodiments, the sugar is covalently bonded to the oxygen atom in the glycosylated fragment using TfpO or PilO OTase. In certain embodiments, the glycosylated fragment is an STT3 glycosylated fragment glycosylated by an STT3 catalytic subunit. In certain embodiments, the sugar is covalently bonded to the nitrogen atom in the glycosylated fragment using STT3 OTase. In certain embodiments, the glycosylated fragment is a PilA_Pa5196-associated pyring glycosylated fragment glycosylated by a TfpW glycosyltransferase. In certain embodiments, the sugar is covalently bonded to the oxygen atom in the glycosylated fragment using a TfpW glycosyltransferase. In certain embodiments, the glycosylated fragment is an Archaeal AlgB glycosylated fragment glycosylated by an AlgB OTase. In certain embodiments, the sugar is covalently bonded to the nitrogen atom in the Archaeal AlgB glycosylated fragment glycosylated fragment using an AlgB OTase. In certain embodiments, the glycosylated fragment is an N-linked glycosyltransferase-glycosylated fragment glycosylated by an N-linked glycosyltransferase, for example, from Actinobacillus pleuropneumoniae, Haemophilus influenzae, or Yersinia enterocolitica. In certain embodiments, the sugar is covalently bonded to the nitrogen atom in the glycosylated fragment using an N-linked glycosyltransferase. In certain embodiments, the glycosylated fragment is an O-linked glycosyltransferase-glycosylated fragment glycosylated by an O-linked glycosyltransferase, for example, GtfA / GtfB glycosyltransferase. In certain embodiments, the sugar is covalently bonded to the oxygen atom in the glycosylated fragment using an O-linked glycosyltransferase. Furthermore, in certain embodiments, the sugar is covalently bonded to the carbon atom in the glycosylated fragment using a C-mannosyltransferase.

[0104] Furthermore, exemplary, in certain embodiments, the sugar is covalently bonded to an oxygen atom in a ComP glycosylated fragment (e.g., SEQ ID NO: 412 or its variants) using PglS OTase (e.g., SEQ ID NO: 400). In certain embodiments, the sugar is covalently bonded to an oxygen atom in a TfpM glycosylated fragment (e.g., SEQ ID NO: 413 or its variants) using TfpM OTase (e.g., SEQ ID NO: 402). In certain embodiments, the sugar is covalently bonded to an oxygen atom in a PilE glycosylated fragment (e.g., SEQ ID NO: 414 or its variants, e.g., SEQ ID NO: 439) using PglL OTase (e.g., SEQ ID NO: 404). In certain embodiments, the sugar is covalently bonded to an oxygen atom in a PilE glycosylated fragment (e.g., SEQ ID NO: 414 or its variants) using PglL OTase (e.g., SEQ ID NO: 404). In certain embodiments, the sugar is covalently bonded to a nitrogen atom in a PglB glycosylated fragment using PglB otase (e.g., SEQ ID NO: 405). In certain embodiments, the sugar is covalently bonded to an oxygen atom in a PilA glycosylated fragment (e.g., SEQ ID NO: 415 or its variants) using TfpO / PilO otase (e.g., SEQ ID NO: 407). In certain embodiments, the sugar is covalently bonded to a nitrogen atom in an STT3 glycosylated fragment using STT3 otase (e.g., SEQ ID NO: 408). In certain embodiments, the sugar is covalently bonded to a nitrogen atom in an Archaeal AlgB glycosylated fragment using AlgB otase (e.g., SEQ ID NO: 409). In certain embodiments, the sugar is covalently bonded to an oxygen atom in a PilA_Pa5196-related pyring glycosylated fragment (e.g., SEQ ID NO: 426 or its variants) using TfpW glycosyltransferase (e.g., SEQ ID NO: 424). In certain embodiments, the sugar is covalently bonded to the nitrogen atom in the N-linked glycosyltransferase sequence using an N-linked glycosyltransferase (e.g., SEQ ID NO: 410). In certain embodiments, the sugar is covalently bonded to the oxygen atom in the O-linked glycosyltransferase sequence using an O-linked glycosyltransferase (e.g., SEQ ID NO: 411).In certain embodiments, the sugar is covalently bonded to a carbon atom within a C-mannosyltransferase glycosylated fragment using C-mannosyltransferase.

[0105] In certain embodiments, the method is a method for producing a conjugate vaccine. This may include adding adjuvants and / or excipients to the partner pair and / or complex of the present disclosure.

[0106] Further embodiments provide systems comprising a first polypeptide and a second polypeptide of the compositions of the present disclosure. In certain embodiments, the first polypeptide is a glycosylated bioconjugate. In certain embodiments, the system comprises a polymeric higher-order structure assembled from the second polypeptide. In certain embodiments, the system comprises sugars and N-linked oligosaccharide transferases (N-OTase), O-linked oligosaccharide transferases (O-OTase), N-linked glycosyltransferases (NGT), O-linked glycosyltransferases (OGT), and / or C-mannosyltransferases (CMT), as disclosed herein.

[0107] Another embodiment provides isolated nucleic acids encoding a first polypeptide and / or a second polypeptide of the composition and / or complex of the present disclosure. Certain embodiments relate to a vector containing isolated nucleic acids. Certain embodiments relate to a host cell containing a vector.

[0108] Another embodiment provides a kit comprising two or more components, including a fusion protein, a first polypeptide, a second polypeptide, sugar, N-linked oligosaccharide transferase (N-OTase), O-linked oligosaccharide transferase (O-OTase), N-linked glycosyltransferase (NGT), O-linked glycosyltransferase (OGT), and / or C-mannosyltransferase (CMT), a bioconjugate, a multimeric higher-order structure assembled from the second polypeptide, isolated nucleic acid, a vector, and a host cell.

[0109] Another aspect provides a method for inducing an immune response in a subject by administering an effective amount of any composition, complex, and / or conjugate vaccine of the present disclosure to the subject. Compositions, complexes, and / or conjugate vaccines of the present disclosure for use in inducing an immune response in a subject are further provided.

[0110] In certain embodiments, the compositions or complexes disclosed herein are conjugate vaccines that can be administered to a subject for the prevention and / or treatment of infectious diseases and / or diseases. In certain embodiments, the conjugate vaccine is a prophylactic measure that can be used, for example, to confer immunity to infectious diseases and / or diseases to a subject. In certain embodiments, the complex carbohydrate is associated with and / or administered with an adjuvant (e.g., in a therapeutic composition). Certain embodiments provide compositions (e.g., therapeutic compositions) comprising the conjugate vaccine and adjuvant described herein. In certain embodiments, administration of the conjugate vaccine to a subject induces an immune response. In certain embodiments, the immune response induces long-term memory (memory B cells and T cells). In certain embodiments, the immunity is an antibody response. In certain embodiments, the antibody response is a serotype-specific antibody response. In certain embodiments, the antibody response is an IgG or IgM response. In certain embodiments where the antibody response is an IgG response, the IgG response is an IgG1 response. Furthermore, in certain embodiments, the conjugate vaccine generates immunological memory in the subject to which the vaccine is administered.

[0111] Certain embodiments also provide the production of vaccines against infections and / or diseases. In certain embodiments, the method comprises isolating a complex carbohydrate or fusion protein (conjugate vaccine) disclosed herein and combining the conjugate vaccine with an adjuvant. In certain embodiments, the infection is a local or systemic infection of the skin, soft tissue, blood, or organs, or is autoimmune. In certain embodiments, the vaccine is a conjugate vaccine against pneumococcal infection. In certain embodiments, the disease is pneumonia. In certain embodiments, the infection is a systemic infection and / or blood infection. In certain embodiments, the subject is a mammal. For example, in certain embodiments, it is a pig or a human.

[0112] Importantly, the embodiments disclosed herein, while not limited to pneumococcal polysaccharides, have broad applicability for generating bioconjugate vaccines for many important human and animal pathogens that are incompatible with PglB and PglL. Notable examples include the human pathogens Klebsiella pneumoniae and group B Streptococcus, as well as the porcine pathogen S. suis, all of which are highly relevant pathogens for which no licensed vaccines are available.

[0113] Methods for inducing a host immune response to a pathogen are provided herein. In certain embodiments, the pathogen is a bacterial pathogen. In certain embodiments, the host is immunized to the pathogen. In certain embodiments, the method comprises administering an effective amount of a ComP conjugate vaccine, a glycosylated fusion protein, or any other therapeutic / immunogenic composition disclosed herein to a subject requiring an immune response. Certain embodiments provide conjugate vaccines, glycosylated fusion proteins, or other therapeutic / immunogenic compositions disclosed herein for use in inducing a host immune response to a bacterial pathogen and in immunization to a bacterial pathogen. Examples of immune responses include, but are not limited to, innate responses, adaptive responses, humoral responses, antibody responses, cellular responses, B-cell responses, T-cell responses, upregulation or downregulation of cytokines, immune system crosstalk, and combinations of two or more such immune responses. In certain embodiments, the immune response is an antibody response. In certain embodiments, the immune response is an innate response, a humoral response, an antibody response, a T cell response, or a combination of two or more of these immune responses.

[0114] Methods for preventing or treating bacterial diseases and / or infections in a subject, comprising administering a conjugate vaccine, fusion protein, or composition disclosed herein to a subject in need thereof, are also provided herein. In certain embodiments, the infection is a local or systemic infection of the skin, soft tissue, blood, or organs, or is autoimmune. In certain embodiments, the disease is pneumonia. In certain embodiments, the infection is a systemic infection and / or blood infection. In certain embodiments disclosed herein, the subject is a vertebrate. In certain embodiments, the subject is a mammal such as a dog, cat, cattle, horse, pig, mouse, rat, rabbit, sheep, goat, guinea pig, monkey, ape, or llama. And, for example, in certain embodiments, the mammal is a human.

[0115] In any embodiment of the administration disclosed herein, the composition is administered by intramuscular injection, intradermal injection, intraperitoneal injection, subcutaneous injection, intravenous injection, oral administration, mucosal administration, intranasal administration, or pulmonary administration.

[0116] In a particular embodiment, a complex carbohydrate, glycosylated fusion protein, or conjugate vaccine according to any of the above claims for use in inducing a host immune response to a bacterial pathogen and / or in preventing or treating bacterial diseases and / or infections in a subject.

[0117] A minimum amount of cyanoacrylate sufficient for O-bond glycosylation. Traditional chemical conjugate vaccine synthesis is considered complex, costly, and laborious (Frasch, CEVaccine 27, 6468-6470 (2009)), but in vivo conjugates are progressing rapidly as a viable biosynthetic alternative (Huttner, A. et al. Lancet Infect Dis 17, 528-537 (2017)). These advances are best highlighted by the success of GlycoVaxyn (now an independent company directly affiliated with GlaxoSmithKline, LimmaTech Biologics AG), a clinical-stage biopharmaceutical company with multiple bioconjugate vaccines in various stages of clinical trials, one of which (Flexyn2a) has just completed a Phase 2b challenge trial. GlycoVaxyn is at the forefront of the in vivo conjugation revolution, but its ability to glycosylate carrier / receptor proteins with polysaccharides containing glucose (Glc) as a reducing end sugar has, predictably, hindered the development of pneumococcal bioconjugate vaccines.

[0118] Oligosaccharide transferase PglS (PMID23658772) and PglL, previously called PglL by Schulz et al. ComP(Harding et al. 2015) (PMID 26727908) was recently characterized as a functional OTase (Schulz, B. Let al. PLoS One 8, e62768 (2013)). Subsequent mass spectrometry studies on total glycopeptides demonstrated that PglS does not function as a general PglL-like OTase that glycosylates multiple periplasmic and outer membrane proteins (Harding, C. et al. Mol Microbiol 96, 1023-1041 (2015)). In fact, the A. baylyi ADP1 genome encodes two OTases: a PglL-like orthologue (UniProtKB / Swiss-Prot:Q6FFS6.1) that functions as a general OTase and PglS, and ComP (UniProtKB / Swiss-Prot:Q6F7F9.1) that glycosylates a single protein (Harding, CM et al. Mol Microbiol 96, 1023-1041 (2015)).

[0119] ComP is orthologized with type IV pyrin proteins such as PilA from Pseudomonas aeruginosa and PilE from Neisseria meningiditis, both of which are glycosylated by OTases TfpO (Castric, P. Microbiology 141 (Pt 5), 1247-1254 (1995)) and PglL (Power, PM et al. Mol Microbiol 49, 833-847 (2003)), respectively. TfpO and PglL also glycosylate their congener pyrin proteins at serine residues, but the site of glycosylation differs between the two systems. TfpO glycosylates its congener pyrine at a C-terminal serine residue (Comer, JE, Marshall, MA, Blanch, VJ, Deal, CD & Castric, P. Infect Immun 70, 2837-2845 (2002)), which is absent in ComP. PglL glycosylates PilE at an internal serine located at position 63 (Stimson, E. et al. Mol Microbiol 17, 1201-1214 (1995)). ComP also contains a serine residue near position 63, and the surrounding residues show moderate conservation with respect to PilE from N. meningiditis. However, comprehensive glycopeptide analysis revealed that this serine and surrounding residues are not the site of glycosylation in ComP. PglS is present in ComP 110264 : Corresponds to serine stored at position 82 of ENV58402.1 (SEQ ID NO: 201) (ComP ADP1 ComP is glycosylated at a single serine residue located at the position (which also corresponds to the conserved serine at position 84 of AAC4588631 (SEQ ID NO: 202)), which is a novel glycosylation site not previously found within the type IV pyrine superfamily. The ability of PglS to translocate polysaccharides containing glucose as a reducing terminal sugar, coupled with the identification of a novel glycosylation site within the pyrine superfamily, demonstrates that PglS is a functionally distinct OTase from PglL and TfpO.

[0120] Bioinformatics characteristics of ComP pyrin ortholog ComP was first described as a factor necessary for spontaneous transformation in Acinetobacter baylyi ADP1 (Porstendorfer, D., Drotschmann, U. & Averhoff, B. Appl Environ Microbiol 63, 4150-4157 (1997)). Subsequent studies have shown that ComP from A. baylyi ADP1 (in this specification, ComP ADP1 It has been demonstrated that ComP (also known as ComP) is glycosylated not by the common OTase PglL located elsewhere on the chromosome, but by a novel OTase, PglS, located immediately downstream of ComP (Harding, CM et al. Mol Microbiol 96, 1023-1041 (2015)). ADP1 The protein (NCBI identifier AAC45886.1) belongs to a family of proteins called type IV pyrin. Specifically, ComP shares homology with type IVa major pyrin (Giltner, CL, Nguyen, Y. & Burrows, LL Microbiol Mol Biol Rev 76, 740-772 (2012)). Type IVa pyrin shares high sequence homology at the N-terminus, which encodes a highly conserved leader sequence and an N-terminal α-helix, but the C-terminus shows significant differences across genera and even within species (Giltner, CL, Nguyen, Y. & Burrows, LL Microbiol Mol Biol Rev 76, 740-772 (2012)). To help distinguish ComP orthologues from other type IVa pyrin proteins such as PilA from A. baumannii, P. aeruginosa, and Haemophilus influenzae, as well as PilE from Neisseria species (Pelicic, V. Mol Microbiol 68, 827-837 (2008)), ComP ADP1 A BLASTp analysis was performed to compare the primary amino acid sequence of ComP with all proteins from bacteria of the genus Acinetobacter. As expected, ComP ADP1Many type IVa pilin orthologs of Acinetobacter, including those, share high homology at their N-terminals, but very few proteins show high sequence conservation across the entire amino acid sequence of ComP. At least six ComP orthologs (Figure 20) were identified based on the presence of a conserved serine at position 84 relative to ComP, and the presence of a conserved disulfide bond adjacent to a predicted glycosylation site that connects a predicted αβ loop to a β-strand region (Giltner, C.L., Nguyen, Y. & Burrows, L.L. Microbiol Mol Biol Rev 76, 740-772 (2012)). Additionally, all six ComP orthologs have both a pglS homolog immediately downstream of the comp gene and a pglL homolog located elsewhere on the chromosome. At the same time, the presence of at least the conserved serine at position 84, a disulfide loop adjacent to the glycosylation site, the presence of the pglS gene immediately downstream of comP, and the presence of a pglL homolog located elsewhere on the chromosome distinguish ComP pilin variants from other type IVa pilin variants. ADP1 Based on the presence of a conserved serine at position 84 relative to ComP, and the presence of a conserved disulfide bond adjacent to a predicted glycosylation site that connects a predicted αβ loop to a β-strand region (Giltner, C.L., Nguyen, Y.& Burrows, L.L.Microbiol Mol Biol Rev 76,740-772(2012)). Additionally, all six ComP orthologs have both a pglS homolog immediately downstream of the comp gene and a pglL homolog located elsewhere on the chromosome. At the same time, the presence of at least the conserved serine at position ADP1 84, a disulfide loop adjacent to the glycosylation site, the presence of the pglS gene immediately downstream of comP, and the presence of a pglL homolog located elsewhere on the chromosome distinguish ComP pilin variants from other type IVa pilin variants.

[0121] Thus, features common to the ComP protein that identify ComP orthologs in different Acinetobacter species are disclosed herein. The ComP protein can be distinguished from other pilins by the presence of a conserved glycosylated serine at position 84 relative to the ADP1 ComP protein, and the presence of a disulfide loop adjacent to the glycosylation site. Additionally, the presence of a pglS homolog immediately downstream of ComP is an indicator of ComP. Furthermore, for classification as a PglS OTase protein rather than a PglL OTase protein, the OTase downstream of ComP must show higher sequence conservation as PglS (ACIAD3337) compared to PglL (ACIAD0103) in A. baylyi ADP1. Also, in any embodiment disclosed herein, the ComP protein has the sequence number 201 (ComP 110264It will be apparent to those skilled in the art that it can be glycosylated onto the serine residue corresponding to the conserved serine residue at position 82 of :ENV58402.1).

[0122] CompP protein glycosylated fragments In certain embodiments, the ComP glycosylated fragment may be, or be derived from, any of the ComP proteins disclosed herein. Furthermore, the PglS OTase may be any of the following:

[0123] Previously, it was demonstrated that the PglS ortholog from Acinetobacter baylyi strain ADP1 glycosylated the ComP ortholog from A. soli strain CIP 110264 at a single serine residue located at position 82 (Harding, CM et al., 2019; WO / 2019 / 241672, the whole report is incorporated herein by reference). PglS was engineered to functionally glycosylate heterologous proteins by translationally fusing a large fragment (117 amino acids) of ComP to the C-terminus of a known carrier protein. Specifically, the 117 amino acid ComP 110264 The fragment was fused at the C-terminus of genetically inactivated exotoxin A from Pseudomonas aeruginosa (EPA) between flexible GGGS linkers (SEQ ID NO: 382). This chimeric carrier protein also possessed an N-terminal DsbA signal sequence (ssDsbA) for translocation to the periplasm via the Sec pathway, as well as a C-terminal hexahistidine tag for detection.

[0124] Even shorter ComP glycosylation fragments sufficient for glycosylation by PglS have been identified (WO / 2020 / 131236, the whole document is incorporated herein by reference). ComP fused to the C-terminus of the EPA carrier protein. 110264 The glycosylated fragment can also be glycosylated by PglS, but the ComP glycosylated fragment can glycosylate both cysteine ​​residues corresponding to Cys71 and Cys93.110264 It has been shown that glycosylation occurs only when it is present in comparison to [another form of] 110264 The fragments are shown. The glycosylated ComP fragments were PCR amplified, cloned to the C-terminus of EPA, and tested for bioconjugation with PglS. For these and all experiments described below, serotype 8 pneumococcal capsular polysaccharide (CPS8) expressed from the pB-8 plasmid was used as the glycan source (Kay, EJ, et al., 2016). CPS8 glycans were selected because they contain glucose as a reducing end sugar, and their efficient conversion to ComP by PglS has been demonstrated to date (Harding, CM et al., 2019). In addition, for these and all experiments described below, bioconjugation was performed in E. coli strain SDB1. SDB1 has deletions of WecA, which initiates biosynthesis of general enterobacteriaceae antigens and O-antigen polysaccharides, and WaaL, which transfers the undecaprenyl-pyrophosphate-linked glycan precursor to the outer core of lipid-A (Garcia-Quintanilla, F., et al., 2014). In summary, these mutations promote the accumulation of heterologously expressed lipid-binding glycan precursors, such as the CPS8 polysaccharide-lipid-linked precursor, due to exclusive use by PglS. CPS8 glycan, PglS, and fused EPA-ComP from IPTG-inducible vectors. 110264 SDB1 strains expressing the construct were cultured in LB broth, induced with midlog, and grown overnight. EPA-ComP 110264To evaluate fusion protein expression and protein glycosylation, samples were collected approximately 20 hours after induction for Western blot analysis of periplasmic extracts. Western blots were probed using antibodies against EPA (anti-EPA) and hexahistidine tag (anti-His). Probing with both antibodies allowed confirmation of whether the EPA protein and / or C-terminal ComP fragment remained intact.

[0125] Figures 12C, 12D, and 12E are from ComP. 110264 The presence of Cys71 and Cys93 residues adjacent to Ser82 in EPA-ComP occurs when the ComP glycosylated fragment is fused at the C-terminus. 110264 This reaffirms the essential role of glycosylation. As seen in Figures 12C, 12D, and 12E, fusion proteins containing ComP glycosylated fragments lacking either Cys71 or Cys93 were not glycosylated. CPS8 glycan migration was observed only in fusion proteins containing ComP glycosylated fragments with both cysteine ​​residues. The glycosylation efficiency and average number of CPS8 repeat units converted by PglS were similar for all fusion proteins containing ComP glycosylated fragments with both Cys71 and Cys93. A detailed examination of Western blot revealed chimeric EPA-ComP 110264 The variants (listed as C2, D2, E3, and F3 in Figures 12C, 12D, and 12E) were observed to react very little with the anti-His antibody compared to the anti-EPA signal (Figure 12D). Furthermore, the anti-EPA channel was found to be non-glycosylated EPA-ComP containing both Cys71 and Cys93. 110264We revealed that it migrated at a slightly lower molecular weight compared to the variant (Figure 12C). In summary, these observations suggest that the ComP fragment lacking both cysteine ​​residues is unstable and prone to C-terminal degradation, thereby likely preventing glycosylation by PglS. Without being constrained by theory, Cys71 and Cys93 form a covalent disulfide bridge to ComP 110264 It is thought that this can stabilize it.

[0126] Various proteins from different organisms, typically inactivated bacterial toxins, have been used as carriers for conjugate and bioconjugate vaccines. Cross-reactive substance 197 (CRM) 197 ) is a genetically inactivated form of diphtheria toxin, widely used as a carrier protein in multiple conjugate vaccines against Streptococcus pneumoniae, Neisseria meningitidis, and Haemophilus influenzae type b (Berti, F. & Adamo, R., 2018). CRM in Conjugate Vaccine Formulations 197 Considering the frequent use of CRM, the PglS bioconjugation system is designed for CRM 197 It has been extended to function in conjunction with the previously identified 25 amino acid "C1" ComP glycosylated fragment (ComP C1 ) is concatenated by the GGGS sequence (sequence number 382) in the CRM. 197 It was translationally fused to the C-terminus. The SRP-dependent FlgI secretion sequence (ssFlgI) was used for export to the periplasm via CRM. 197 The hexahistidine tag was added to the N-terminus (Goffin, P., et al., 2017). Finally, a C-terminal hexahistidine tag was added to aid in purification (Figure 13A). PglS and CRM 197 -ComP C1 E. coli SDB1 cells (expected size 61.8 kDa) expressing CPS8 glycan along with a carrier were cultured in a shaking flask and harvested after 24 hours. 197 -ComP C1-CPS8 complex carbohydrates were purified by three consecutive chromatography cycles. First, nickel affinity chromatography was used because the complex carbohydrates contain a C-terminal hexahistidine tag. The fractions containing the complex carbohydrates were pooled and concentrated for the glycosylated complex carbohydrates using a MonoQ column, and eluted with a linear salt gradient. A final polishing step to remove large aggregates was performed on a Superdex 200 Increase column. Anti-CRM 197 Western blotting on purified samples using pneumococcal CPS8 antiserum is performed by CRM. 197 -ComP C1 We demonstrated that the carrier was glycosylated with CPS8. By digesting the purified complex carbohydrate with proteinase K before separation on SDS-PAGE, CRM 197 And polysaccharide-specific signals are completely lost, which means that CPS8 glycan is CRM 197 -ComP C1 This indicates that the protein is covalently bonded to it.

[0127] Next, CompP C1 Glico Tags in CRM 197 We tested whether it could be moved to a different part of the fusion. Therefore, CRM 197 CompP at the N end of the code region C1 A new structure was designed by arranging the elements (Figure 14A). The FlgI secretion signal was converted to ComP C1 Placed immediately at the N-terminus of the glycosylated fragment, CRM 197 The C-terminus was tagged with hexahistidine. PglS and ComP C1 -CRM 197 E. coli SDB1 cells expressing CPS8 glycan with a carrier were cultured in a shaking flask and harvested after 24 hours. Western blot analysis of periplasm extract probed with anti-His antibody, as shown in Figure 14B, revealed ComP C1 -CRM 197It was also shown that it was glycosylated by PglS. The average number of CPS8 repeating units and glycosylation efficiency of both fusions were equivalent, which indicates that ComP C1 This demonstrates that the glycotag can be positioned at the N-terminus or C-terminus of a carrier protein.

[0128] ComP contains 11 amino acids sufficient for PglS glycosylation. 110263 Identifying the sequence Previous reports have shown that Cys71 and Cys93 are translationally fused at the C-terminus of EPA in ComP 110264 Although it has been shown that these two cysteine ​​residues and the putative disulfide crosslinks formed between them are necessary for glycosylation of fusion proteins containing glycosylated fragments (e.g., Figures 12C, 12D, and 12E), these data do not confirm whether these two cysteine ​​residues and the putative disulfide crosslinks formed between them are absolutely necessary for glycosylation by PglS under all circumstances. The N-linked sequence recognized by PglB is manipulated at multiple sites on the surface loop of EPA and used as an "internal" glycotag (Ihssen, J. et al., 2010). 110264 To determine whether Cys71 and Cys93 are necessary for PglS glycosylation, this specification uses iGT of internal glycotags. CC The cysteine-cysteine-component is the 23 amino acid compound called ComP, ranging from Cys71 to Cys93. 110264 The glycosylated fragment was incorporated into the EPA amino acid sequence. ComP 110264 iGT CC This was inserted between the residues Ala489 and Arg490 of EPA, resulting in a β-turn structure on the surface of the catalytic domain (Figure 15A). As a control, an iGT called iGTss ("serine-serine") was created, which contains serine residues instead of cysteine ​​residues at positions 71 and 93 of ComP. CCA variant of the ComP glycosylated fragment was also incorporated. This iGTSS ComP glycosylated fragment was also incorporated between the EPA residues Ala489 and Arg490. Although the serine residue is assumed to contribute a similar stereobulk as the cysteine ​​residue, it cannot be oxidized to form a disulfide bond (Figure 15B). CPS8 was incorporated into EPA. iGTcc or EPA iGTss The ability of PglS to transfer to was evaluated using the three plasmid systems described above. As can be seen in Figures 15C and 15D, EPA iGT Both the cysteine-cysteine ​​variant and the serine-serine variant are glycosylated, indicating that Cys71 and Cys93 (and the putative disulfide bond formed between them) are not required for glycosylation by PglS when the ComP fragment is introduced into the EPA protein.

[0129] Since cysteine ​​residues are not required for PglS-dependent glycosylation only when the ComP glycosylation fragment is incorporated into the fusion protein, it was intended that a shorter ComP glycosylation fragment representing the minimum O-bonded ComP sequence could be found within a 23-amino acid ComP glycosylation fragment extending from Cys71 to Cys93. To investigate this, an iGT incorporated between EPA residues Ala489 and Arg490 was introduced. CC Shorter variants of the ComP glycosylated fragment were generated to identify which ComP residues are required for glycosylation. Alternative single amino acids were used in 23 amino acid iGTs. CC By deleting from either side, 22 cleavage variants were generated, each containing Ser82, the PglS glycosylation site (Figures 16A and 16B). These variants were then used in iGT CC Named after the number of residues deleted from either side, for example, Δ3-4 is iGT CC This corresponds to the deletion of three amino acids from the N-terminus and four amino acids from the C-terminus. The shortest variant generated was 5 amino acids long. These cleavages of EPA-iGT CCThe variants were tested for bioconjugation with CPS8 and PglS in a shaking flask under the same conditions as in the aforementioned experiment. As a negative control, we included a construct expressing only the EPA coding sequence along with DsbA secretion and a hexahistidine tag.

[0130] Figure 16C shows that robust glycosylation was observed in all EPA fusion proteins containing ComP glycosylated fragments of at least 11 amino acids in length. The glycosylation ratio was 23 amino acid iGT. CC The results were equivalent to the ComP glycosylated fragment, suggesting that moderate cleavage on both sides of Ser82 does not significantly affect the glycosylation efficiency by PglS. These fusion proteins were glycosylated, but a slight decrease in glycosylation efficiency was observed as the amino acid sequence of the iGT ComP glycosylated fragment was shortened. The shortest internal ComP glycosylated fragment that was efficiently glycosylated was iGTΔ6-6 with the sequence IASGASAATTN (SEQ ID NO: 309, Figure 16C). Removal of either the N-terminal isoleucine residue (iGTΔ7-6, SEQ ID NO: 321) or the C-terminal asparagine residue (iGTΔ6-7, SEQ ID NO: 310) dramatically reduced the glycosylation efficiency of the carrier protein, suggesting that these residues play a crucial role in PglS glycosylation. Most variants smaller than iGTΔ6-6 show minimal glycosylation, with the best of these being iGTΔ7-6 with the sequence ASGASAATTN (SEQ ID NO: 321). Interestingly, even in fusion proteins containing the smallest ComP glycosylated fragments, iGTΔ9-8 (SEQ ID NO: 346) and iGTΔ9-9 (SEQ ID NO: 347) (Figure 16D), small amounts of higher molecular weight ladders were observed (6 and 5 amino acid variants, respectively), suggesting that these six and five amino acid variants were glycosylated by PglS at very low levels. This indicates that ComP recognized by PglS is not... 110264 This suggests that glycosylated sequencers can be as small as five amino acids.

[0131] Next, CPS8 glycosylated EPA fusion proteins containing the iGTΔ6-6 ComP glycosylated fragment located between residues Ala489 and Arg490 were purified from whole cell lysates using Ni affinity chromatography, and Western blot analysis was performed on the eluate using antiserum specific to either the EPA protein or the CPS8 glycan. The results of these experiments clearly show that the EPA fusion protein containing the iGTΔ6-6 ComP glycosylated fragment located between residues Ala489 and Arg490 was glycosylated at CPS8 by PglS (Figures 17A, 17B, and 17C). Overall, these experiments demonstrate that glycosylation is maintained while ComP 110264 The glycosylated fragment is divided into 117 amino acids, ComP 110264 This shows that the sequence can be shortened to a length of about 11 amino acids or even shorter. These results indicate that ComP, which has been previously shown to be necessary when fused at the C-terminus, can be shortened. 110264 The cysteine ​​residues corresponding to Cys71 and Cys93 unexpectedly indicate that they are not required for PglS-dependent glycosylation when the ComP glycosylated fragment is incorporated into the fusion protein.

[0132] The aforementioned iGT cleavage series was tested at one internal site on the EPA between residues Ala489 and Arg490. Next, a second site between EPA residues Glu548 and Gly549, incorporating the iGTΔ3-4 ComP glycosylated fragment (SEQ ID NO: 271), was tested. Similar to the first site, the second site is located on a surface-exposed loop within the catalytic domain of the EPA. This alternately tagged variant for bioconjugation with CPS8 and PglS was tested under the same conditions as the other cleavages. It was observed that this construct was glycosylated with CPS8 with similar efficiency to when iGTΔ3-4 was placed at the first site on the EPA. Next, the CPS8 glycosylated EPA fusion protein containing the iGTΔ3-4 ComP glycosylated fragment located between residues Glu548 and Gly549 was purified from whole cell lysates using Ni affinity chromatography, and Western blot analysis was performed on the eluate using antiserum specific to either the EPA protein or the CPS8 glycan. In this case as well, the results of these experiments indicate that the EPA fusion protein containing the iGTΔ3-4 ComP glycosylated fragment located between residues Glu548 and Gly549 was glycosylated at CPS8 by PglS. Overall, these experiments show that glycosylation is maintained while ComP 110264 The glycosylated fragment is divided into 117 amino acids, ComP 110264 This shows that the sequence can be shortened to about 11 amino acids or even shorter. These results indicate that ComP 110264 The cysteine ​​residues corresponding to Cys71 and Cys93 unexpectedly indicate that they are not required for PglS-dependent glycosylation when the ComP glycosylated fragment is incorporated into the fusion protein.

[0133] A complex carbohydrate comprising an oligosaccharide or polysaccharide bound to a fusion protein is provided herein. In certain embodiments, the oligosaccharide or polysaccharide is covalently bound to the fusion protein. The fusion protein comprises a glycosylated fragment of a ComP protein (as described in detail elsewhere herein). In certain embodiments of the complex carbohydrate of this disclosure, the oligosaccharide or polysaccharide comprises glucose at its reducing end.

[0134] ComP is glycosylated on a serine (S) residue. This serine residue is the same as in Sequence ID No. 201 (ComP 110264 This corresponds to position 82 of (ENV58402.1). This serine residue is conserved in the ComP protein, for example, in sequence number 202 (ComP ADP1 This corresponds to position 84 of :AAC45886.1). Therefore, in a certain embodiment, the fusion protein (and by extension the complex carbohydrate) is sequence number 202 (ComP ADP1 The serine residue corresponding to the serine residue at position 84 of :AAC45886.1) or Sequence ID No. 201 (ComP 110264 The serine residue corresponding to the serine residue at position 82 of (ENV58402.1) is glycosylated by an oligosaccharide or polysaccharide on the ComP glycosylated fragment. Figure 22 shows Sequence ID No. 201 (ComP 110264 This shows the alignment of the region of the ComP sequence containing the serine (S) residue (outlined) corresponding to the serine residue at position 82 of :ENV58402.1), which is conserved across the ComP sequence.

[0135] Those skilled in the art will recognize that by aligning the ComP sequence with SEQ ID NO: 201 (for example, either the complete or partial sequence), the conserved serine residue of the non-SEQ ID NO: 201 ComP protein corresponding to the serine residue at position 82 of SEQ ID NO: 201 can be identified. Furthermore, those skilled in the art will recognize that by aligning the ComP sequence with SEQ ID NO: 201, other residues, regions, and / or features corresponding to the residues, regions, and / or features of SEQ ID NO: 201 referred herein can be identified in the non-SEQ ID NO: 201 ComP sequence and referred to in relation to SEQ ID NO: 201. And while SEQ ID NO: 201 is generally referred to herein, by analogy, any residue, region, feature, etc. of any ComP sequence disclosed herein can be similarly referred to, for example, SEQ ID NO: 202.

[0136] A ComP protein is a protein identified as a ComP protein consistent with the description provided herein. For example, representative examples of ComP proteins include AAC45886.1 ComP [Acinetobacter sp. ADP1], ENV58402.1 Hypothesis Protein F951_00736 [Acinetobacter soli CIP 110264], APV36638.1 Competence Protein [Acinetobacter soli GFJ-2], PKD82822.1 Competence Protein [Acinetobacter radioresistens 50v1], SNX44537.1 Type IV Ciliated Aggregation Protein PilA [Acinetobacter puyangensis ANC 4466], OAL75955.1 Competence Protein [Acinetobacter sp. SFC], and ComP P5312 , and Comp ANT_H59 Examples include: In a particular embodiment, the ComP protein is sequence number 201 (ComP ADP1 ) or Sequence ID 201 (ComP 110264) and contains an amino acid sequence that is at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical, and contains a serine residue corresponding to the conserved serine residue at position 84 of SEQ ID NO: 202 or position 82 of SEQ ID NO: 201. SEQ ID NO: 202 contains a 28 - amino - acid leader sequence. In certain embodiments, the ComP protein does not contain an amino - acid leader sequence but contains a serine residue corresponding to the conserved serine residue at position 82 of SEQ ID NO: 201 (ComP 110264 : AAC45886.1), and is SEQ ID NO: 210 (ComPΔ28 ADP1 ), SEQ ID NO: 209 (ComPΔ28 110264 ), SEQ ID NO: 211 (ComPΔ28 GFJ-2 ), SEQ ID NO: 212 (ComPΔ28 P50v1 ), SEQ ID NO: 213 (ComPΔ28 4466 ), SEQ ID NO: 214 (ComPΔ28 SFC ), SEQ ID NO: 215 (ComPΔ28 P5312 ), or SEQ ID NO: 216 (ComPΔ29 ANT_H59 ) and contains an amino - acid sequence that is at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical. In certain embodiments, the ComP protein does not contain a 28 - amino - acid leader sequence but contains a serine residue corresponding to the conserved serine residue at position 82 of SEQ ID NO: 201 (ComP 110264 ) and is SEQ ID NO: 209 (ComPΔ28 110264 )" and contains an amino - acid sequence that is at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical. In certain embodiments, the ComP protein is SEQ ID NO: 210 (ComPΔ28 ADP1 ), SEQ ID NO: 209 (ComPΔ28 110264 ), SEQ ID NO: 211 (ComPΔ28 GFJ-2 ), SEQ ID NO: 212 (ComPΔ28 P50v1 ), SEQ ID NO: 213 (ComPΔ28 4466 ), SEQ ID NO: 214 (ComPΔ28 SFC ), SEQ ID NO: 215 (ComPΔ28 P5312 )), or sequence number 216 (ComPΔ29 ANT_H59 ) contains. In certain embodiments, the ComP protein is sequence number 202 (ComP ADP1 :AAC45886.1), Sequence ID 201 (ComP 110264 :ENV58402.1), Sequence ID 203 (ComP GFJ-2 :APV36638.1), Sequence ID 204 (ComP 50v1 :PKD82822.1), Sequence ID 205 (ComP 4466 :SNX44537.1), Sequence ID 206 (ComP SFC :OAL75955.1), Sequence ID 207 (ComP P5312 ), or sequence number 208 (ComP ANT_H59 )

[0137] A complex carbohydrate comprising an oligosaccharide or polysaccharide covalently bound to a fusion protein, wherein the fusion protein comprises a ComP protein (ComP) glycosylated fragment, is provided herein. In certain embodiments, the ComP glycosylated fragment is ComP 110264 It does not contain the cysteine(C) residue corresponding to the conserved cysteine(C) residue at position 71 of (SEQ ID NO: 201). In certain embodiments, the ComP glycosylated fragment is ComP 110264 It does not contain the cysteine(C) residue corresponding to the conserved cysteine(C) residue at position 93 of (SEQ ID NO: 201). As described in more detail herein, the fusion protein is ComP 110264 The ComP glycosylated fragment is glycosylated with an oligosaccharide or polysaccharide at the serine residue corresponding to the conserved serine residue at position 82 of (SEQ ID NO: 201). In certain embodiments, the ComP glycosylated fragment is located internally within the fusion protein. Furthermore, in certain embodiments, the ComP glycosylated fragment portion of the fusion protein is exposed to the solvent (or surface) and / or C of the fusion protein. 10 It is incorporated into β-turns, β-twists, β-loops, U-turns, reverse turns, chain reverses, or hairpin loops.

[0138] Since it has been discovered that when a ComP glycosylated fragment is located internally within a fusion protein, the ComP glycosylated fragment does not require an adjacent cysteine ​​residue for glycosylation, the ComP glycosylated fragments disclosed herein can be shorter than previously thought. In certain embodiments, the ComP glycosylated fragment is such that it is ComP 110264 As long as it contains the serine residue corresponding to the conserved serine residue at position 82 of (SEQ ID NO: 201), it can be shorter than 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, or 6 amino acids. In certain embodiments, the ComP glycosylated fragment has a length from one of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 to one of 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, or 22 amino acid lengths. In a particular embodiment, the fragment is a ComP protein having at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 amino acid residues at the N-terminus of a serine residue corresponding to the conserved serine residue at position 82 of SEQ ID NO: 201, e.g., X n The fragment has S[Y], where n is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 amino acid residues of the ComP protein. In certain embodiments, the fragment is a ComP protein, e.g., [X]SY, with at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 amino acid residues at the C-terminus of a serine residue corresponding to the conserved serine residue at position 82 of SEQ ID NO: 201. n The amino acid sequence of the ComP glycosylated fragment is ComP 110264 (SEQ ID NO: 201) does not extend beyond the amino acid residue corresponding to position 72 towards the N-terminus, and / or ComP 110264 It does not extend beyond the amino acid residue corresponding to position 92 of (SEQ ID NO: 201) to the C-terminus.

[0139] In accordance with the ComP protein of this disclosure, in certain embodiments, the ComP protein derived from the ComP glycosylated fragment is sequence number 209 (ComPΔ28 110264 ), Sequence ID 210 (ComPΔ28 ADP1 ), Sequence ID 211 (ComPΔ28 GFJ-2 ), Sequence ID 212 (ComPΔ28 P50v1 ), Sequence ID 213 (ComPΔ28 4466 ), Sequence ID 214 (ComPΔ28 SFC ), Sequence ID 215 (ComPΔ28 P5312 ), or sequence number 216 (ComPΔ29 ANT_H59 ) contains at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical amino acid sequences. In certain embodiments, the ComP protein derived from the ComP glycosylated fragment is sequence number 209 (ComPΔ28 110264 ), Sequence ID 210 (ComPΔ28 ADP1 ), Sequence ID 211 (ComPΔ28 GFJ-2 ), Sequence ID 212 (ComPΔ28 P50v1 ), Sequence ID 213 (ComPΔ28 4466 ), Sequence ID 214 (ComPΔ28 SFC ), Sequence ID 215 (ComPΔ28 P5312 ), or sequence number 216 (ComPΔ29 ANT_H59 ) includes.

[0140] In certain embodiments of the complex carbohydrates of this disclosure, the ComP glycosylated fragment has the following amino acid consensus sequence: [Table 1] Or it comprises or consists of a fragment of at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids in length, containing a serine (S) residue corresponding to position 11 of SEQ ID NO: 217. In certain embodiments, the fragment has at least 1, 2, 3, 4, 5, 6, 7, or 8 amino acid residues at the N-terminus of the serine (S) residue corresponding to position 11 of SEQ ID NO: 217. In certain embodiments, the fragment has at least 1, 2, 3, 4, 5, 6, 7, or 8 amino acid residues at the C-terminus of the serine (S) residue corresponding to position 11 of SEQ ID NO: 217. However, the ComP glycosylated fragment is ComP 110264 The ComP glycosylated fragment does not contain the cysteine(C) residue corresponding to the conserved cysteine(C) residue at position 71 of (SEQ ID NO: 201), and / or ComP is ComP 110264 It does not contain the cysteine(C) residue corresponding to the conserved cysteine(C) residue at position 93 of (SEQ ID NO: 201).

[0141] Certain embodiments provide a ComP glycosylated fragment which is a variant of the amino acid consensus sequence or fragment thereof of SEQ ID NO: 217, SEQ ID NO: 396, or SEQ ID NO: 397 having 1, 2, 3, 4, 5, 6, or 7 amino acid substitutions, additions, and / or deletions, wherein this variant maintains the serine (S) residue corresponding to position 11 of SEQ ID NO: 217, and this variant is ComP 110264 This variant does not contain the cysteine(C) residue corresponding to the conserved cysteine(C) residue at position 71 of (SEQ ID NO: 201), and / or this variant is ComP 110264 It does not contain the cysteine(C) residue corresponding to the conserved cysteine(C) residue at position 93 of (SEQ ID NO: 201). Those skilled in the art will understand that the number of amino acid substitutions, additions, and / or deletions that can be tolerated within a sequence without invalidating its function (e.g., its ability to function as a sequenceon) may depend on the length of the sequence. For example, a 6-amino acid sequence will tolerate fewer changes than a 21-amino acid sequence.

[0142] Whether a ComP glycosylated fragment (including fragments and variants disclosed herein and collectively referred to as ComP glycosylated fragments) can be glycosylated, and the efficiency of glycosylation, can be determined by methods described herein, among others. In certain embodiments, a ComP glycosylated fragment can be glycosylated if it is located within a fusion protein and / or within a carrier protein sequence, as described elsewhere herein. Furthermore, in certain embodiments, if a ComP glycosylated fragment or variant is located at the N-terminus and / or C-terminus of a fusion protein, it is not glycosylated, or if it is located at the N-terminus and / or C-terminus of a fusion protein, it is glycosylated by at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% less than if it were located at the N-terminus and / or C-terminus of a fusion protein.

[0143] In a particular embodiment, the fusion protein is Pseudomonas aeruginosa exotoxin A (EPA), CRM 197 The carrier protein is selected from the group consisting of cholera toxin B subunit, tetanus toxin C fragment, Haemophilus influenzae protein D, and any or more fragments thereof. For example, in a particular embodiment, the Pseudomonas aeruginosa exotoxin A (EPA) carrier protein includes the amino acid sequence of SEQ ID NO: 218, or a fragment or more fragments thereof. For example, in a particular embodiment, CRM 197 The carrier protein contains the amino acid sequence of SEQ ID NO: 224, or a fragment or multiple fragments thereof.

[0144] As can be understood from the entire disclosure, "inside the fusion protein" means that the ComP fusion protein is not located at the C-terminus or N-terminus of the fusion protein, and does not contain any C-terminal leader sequence or N-terminal tag (e.g., His-Tag).

[0145] In certain embodiments, the ComP glycosylated fragment can be bound to a carrier protein sequence via an amino acid linker.

[0146] Furthermore, in certain embodiments, the ComP glycosylated fragment can be inserted into the sequence of carrier proteins rather than between them. For example, in certain embodiments, (i) The ComP glycosylated fragment is inserted between Ala489 and Arg490 into the PDB entity 1IKQ (SEQ ID NO: 219) of Pseudomonas aeruginosa exotoxin A (EPA), (ii) The ComP glycosylated fragment is inserted between Glu548 and Gly549 into PDB entity 1IKQ (SEQ ID NO: 220) of Pseudomonas aeruginosa exotoxin A (EPA), (iii) The ComP glycosylated fragment is inserted between Ala122 and Gly123 into PDB entity 1IKQ (SEQ ID NO: 221) of Pseudomonas aeruginosa exotoxin A (EPA), (iv) The ComP glycosylated fragment is inserted between Thr355 and Gly356 in the PDB entity 1IKQ (SEQ ID NO: 222) of Pseudomonas aeruginosa exotoxin A (EPA), or (v) The ComP glycosylated fragment is inserted between Lys20 and Asp21 into the PDB entity 1IKQ (SEQ ID NO: 223) of Pseudomonas aeruginosa exotoxin A (EPA).

[0147] Furthermore, in certain embodiments, the ComP glycosylated fragment can be inserted into the sequence of carrier proteins rather than between them. For example, in certain embodiments, (i) ComP glycosylated fragments are CRM 197 For PDB entity 4AE0 (sequence number 225), it is inserted between Asn481 and Gly482, (ii) ComP glycosylated fragments, CRM 197 For PDB entity 4AE0 (sequence number 226), it is inserted between Asp392 and Gly393, (iii) ComP glycosylated fragments, CRM 197 For PDB entity 4AE0 (sequence number 227), it is inserted between Glu142 and Gly143, (iv) ComP glycosylated fragments, CRM 197 For PDB entity 4AE0 (sequence number 228), it is inserted between Asp129 and Gly130, or (v) ComP glycosylated fragments, CRM 197 This is inserted between Asn69 and Glu70 into PDB entity 4AE0 (sequence number 229).

[0148] In certain embodiments, the ComP glycosylated fragment can be located between carrier proteins, or it can be inserted into the sequence of carrier proteins(s) within a single fusion protein. In certain embodiments, the ComP glycosylated fragment can be located internally, or one or more ComP glycosylated fragments can be located at the C-terminus and / or N-terminus, sufficient for glycosylation at such a location.

[0149] A particular aspect of this disclosure is that a fusion protein can be designed to contain multiple ComP glycosylated fragments that increase the immunogenicity of the glycosylated fusion protein / complex carbohydrate. In certain embodiments, the fusion protein contains two or more, three or more, four or more, five or more, six or more, eight or more, ten or more, fifteen or more, or twenty or more ComP glycosylated fragments. In certain embodiments, the fusion protein does not contain more than three, five or more, ten or more, fifteen or more, twenty or more ComP glycosylated fragments. The identity of the ComP glycosylated fragments can also be controlled. For example, in certain embodiments, multiple ComP glycosylated fragments of the fusion protein are identical. In certain embodiments, the ComP glycosylated fragments of the fusion protein are different from each other. For example, in certain embodiments, at least three, at least four, or at least five of the ComP glycosylated fragments of the fusion protein are all different from each other. For example, in a particular embodiment, the ComP glycosylated fragments of the fusion protein are not all the same.

[0150] In certain embodiments, the oligosaccharide or polysaccharide is derived from a sugar produced by a bacterium of the genus Streptococcus. For example, in certain embodiments, the sugar is a S. pneumoniae, S. agalactiae, or S. suis capsular polysaccharide; in certain embodiments, the sugar is a serotype 8 capsular polysaccharide derived from S. pneumoniae; and in certain embodiments, the sugar is a type Ia, Ib, II, III, IV, V, VI, VII, VIII, or X capsular polysaccharide derived from S. agalactiae.

[0151] In certain embodiments, the oligosaccharide or polysaccharide is derived from a sugar produced by a bacterium of the genus Klebsiella. For example, in certain embodiments, the sugar is a capsular polysaccharide of K. pneumoniae, K. varricola, K. michinganenis, or K. oxytoca, and in certain embodiments, the sugar is an O-antigen polysaccharide of K. pneumoniae, K. varricola, K. michinganenis, or K. oxytoca.

[0152] In certain embodiments, the complex carbohydrate is produced in vivo, for example, within bacterial cells, within Escherichia coli, within bacteria of the genus Klebsiella, and / or the bacterial species is K. pneumoniae, K. varricola, K. michinganenis, or K. oxytoca.

[0153] The above complex carbohydrates (for example, CompP glycosylated fragments are CompP 110264 The ComP glycosylated fragment does not contain the cysteine(C) residue corresponding to the conserved cysteine(C) residue at position 71 of (SEQ ID NO: 201), and / or ComP is ComP 110264 A complex carbohydrate is disclosed herein, which does not contain the cysteine ​​(C) residue corresponding to the conserved cysteine ​​(C) residue at position 93 of (SEQ ID NO: 201), and in which the ComP glycosylated fragment contains or comprises the amino acid sequence of SEQ ID NOs: 232-363 or 364. The above complex carbohydrate (for example, the ComP glycosylated fragment is ComP 110264 The ComP glycosylated fragment does not contain the cysteine(C) residue corresponding to the conserved cysteine(C) residue at position 71 of (SEQ ID NO: 201), and / or ComP is ComP 110264 A complex carbohydrate is disclosed herein, which does not contain the cysteine(C) residue corresponding to the conserved cysteine(C) residue at position 93 of (SEQ ID NO: 201), and in which the ComP glycosylated fragment contains or consists of the following amino acid sequence: [Table 2]

[0154] A ComP glycosylated fragment that is a variant of any of the disclosed ComP glycosylated fragments having 1, 2, 3, 4, 5, 6, or 7 amino acid substitutions, additions, and / or deletions, wherein the variant maintains the serine residue corresponding to the conserved serine residue at position 82 of SEQ ID NO: 201, and the variant is ComP 110264 (SEQ ID NO: 201) does not contain the cysteine(C) residue corresponding to the conserved cysteine(C) residue at position 71, and / or this variant is ComP 110264 The ComP glycosylated fragments that do not contain the cysteine ​​(C) residue corresponding to the conserved cysteine ​​(C) residue at position 93 of (SEQ ID NO: 201) are also provided herein.

[0155] Whether a ComP glycosylated fragment (including fragments and variants of fragments disclosed herein and collectively referred to as ComP glycosylated fragments) can be glycosylated, and the efficiency of glycosylation, can be determined by methods described herein, among others. In certain embodiments, a ComP glycosylated fragment can be glycosylated if it is located within a fusion protein and / or within a carrier protein sequence, as described elsewhere herein. Furthermore, in certain embodiments, if a ComP glycosylated fragment is located at the N-terminus and / or C-terminus of a fusion protein, it is not glycosylated, or if it is located at the N-terminus and / or C-terminus of a fusion protein, it is glycosylated by at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% less than if it were located at the N-terminus and / or C-terminus of a fusion protein.

[0156] In certain embodiments, the complex carbohydrate is a conjugate vaccine. Therefore, in certain embodiments, this disclosure targets and provides a conjugate vaccine. In certain embodiments, the conjugate vaccine is a vaccine against Streptococcus pneumoniae serotype 8. In certain embodiments, the conjugate vaccine induces an immune response when administered to a subject. In certain embodiments, the immune response is an antibody response that induces long-term memory (memory B cells and T cells), and is optionally a serotype-specific antibody response. In certain embodiments, the antibody response is an IgG or IgM response. In certain embodiments, the antibody response is an IgG response, and optionally an IgG1 response. Also, in certain embodiments, the conjugate vaccine generates immunological memory in the subject to which the vaccine is administered.

[0157] The above describes a complex carbohydrate comprising a ComP glycosylated fragment containing an isolated fragment of the ComP protein, but the present disclosure is consistent with any description of a ComP glycosylated fragment provided anywhere in this specification, including the following appended claims, wherein the ComP glycosylated fragment is, for example, ComP 110264 It does not contain the cysteine ​​residue corresponding to the conserved cysteine ​​residue at position 71 of (SEQ ID NO: 201), and / or ComP 110264 The ComP glycosylated fragment does not contain the cysteine ​​residue corresponding to the conserved cysteine ​​residue at position 93 of (SEQ ID NO: 201), and ComP 110264 It is understood that we also explicitly provide a ComP glycosylated fragment containing a serine residue corresponding to the conserved serine residue at position 82 of (SEQ ID NO: 201).

[0158] A fusion protein comprising the ComP glycosylated fragment of the present disclosure is provided herein. In certain embodiments, the fusion protein is sequence number 201 (ComP 110264The fusion protein is glycosylated by an oligosaccharide or polysaccharide at the serine residue on the glycosylated fragment corresponding to the serine ComP glycosylated fragment residue at position 82 of the fusion protein. Furthermore, although the above describes a complex carbohydrate comprising a ComP glycosylated fragment comprising a fusion protein, it is understood that this disclosure also expressly provides a fusion protein that is consistent with any description of a fusion protein provided anywhere in this specification, including the appended claims below. In certain embodiments, the fusion protein is Pseudomonas aeruginosa exotoxin A (EPA), CRM 197 The carrier protein comprises a cholera toxin B subunit, a tetanus toxin C fragment, Haemophilus influenzae protein D, and any or more fragments thereof.

[0159] Methods for in vivo conjugation of oligosaccharides or polysaccharides to receptor polypeptides are also provided herein. In certain embodiments, the method comprises culturing a host cell containing components necessary for conjugation to an oligosaccharide or polysaccharide polypeptide. Generally, these components are an oligosaccharide transferase, a receptor polypeptide to be glycosylated, and an oligosaccharide or polysaccharide. The method comprises covalently conjugating the oligosaccharide or polysaccharide to the receptor polypeptide (the fusion protein of this disclosure) with a PglS oligosaccharide transferase (OTase), wherein the receptor polypeptide comprises a ComP glycosylated fragment as described herein. In certain embodiments, the PglS OTase is PglS 110264 (Sequence ID 365), PglS ADP1 (Sequence ID 366), PglS GFJ-2 (Sequence ID 367), PglS 50v1 (Sequence ID 368), PglS 4466 (Sequence ID 369), PglS SFC (Sequence ID 370), Pgl SP5312 (Sequence ID 371), or PglS ANT_H59 (Sequence ID 372). In certain embodiments, the oligosaccharide or polysaccharide is Sequence ID 201 (ComP 110264The ComP glycosylated fragment is bound at the serine (S) residue corresponding to the serine residue at position 82 of ). In certain embodiments, in vivo conjugation takes place in a host cell. In certain embodiments, the complex carbohydrate is produced in bacterial cells, fungal cells, yeast cells, avian cells, algal cells, insect cells, or mammalian cells. In certain embodiments, the host cell is, for example, a bacterial cell in Escherichia coli, a bacterial cell of the genus Klebsiella, and the bacterial species is K. pneumoniae, K. varricola, K. michinganenis, or K. oxytoca. Certain embodiments include culturing a host cell containing (a) a gene cluster encoding a protein required for the synthesis of oligosaccharides or polysaccharides, (b) PglS OTase, and (3) a receptor polypeptide. In certain embodiments, the production of oligosaccharides or polysaccharides is enhanced by K. pneumoniae transcription activator rmpA (K. pneumoniae NTUH K-2044) or a homolog of K. pneumoniae transcription activator rmpA (K. pneumoniae NTUH K-2044). In certain embodiments, the method further comprises expressing and / or providing such transcription activators in host cells together with other components.

[0160] In certain embodiments, complex carbohydrates are produced in cell-free systems. Examples of cell-free systems using OTases other than PglS can be found in WO2013 / 067523A1, which is incorporated herein by reference.

[0161] Also provided are host cells comprising (a) a gene cluster encoding a protein required for the synthesis of oligosaccharides or polysaccharides, (b) PglS OTase, and (3) a receptor polypeptide comprising the ComP glycosylated fragment of the present disclosure. In certain embodiments, the receptor polypeptide is a fusion protein. In certain embodiments, the host cell comprises a nucleic acid encoding PglS OTase. In certain embodiments, the host cell comprises a nucleic acid encoding the receptor polypeptide.

[0162] Isolated nucleic acids encoding the ComP glycosylated fragments and / or fusion proteins of this disclosure are also provided herein. In certain embodiments, the nucleic acid is a vector. In certain embodiments, the host cell contains the isolated nucleic acid.

[0163] The complex carbohydrates of the present invention may have one of a number of uses, including but not limited to use as a conjugate vaccine. Thus, a conjugate vaccine is produced in a particular manner. In a particular embodiment, a composition comprising the conjugate vaccine or fusion protein of the present disclosure and an adjuvant. For example, in a particular embodiment, the conjugate vaccine may contain Streptococcus pneumoniae serotype 8, Streptococcus pneumoniae serotype 1, Streptococcus pneumoniae serotype 2, Streptococcus pneumoniae serotype 4, Streptococcus pneumoniae serotype 5, Streptococcus pneumoniae serotype 6A, Streptococcus pneumoniae serotype 6B, Streptococcus pneumoniae serotype 7F, Streptococcus pneumoniae serotype 9N, Streptococcus pneumoniae serotype 9V, Streptococcus pneumoniae serotype 10A, Streptocococcus pneumoniae serotype 11A, Streptocococcus pneumoniae serotype 12F, Streptococcus pneumoniae serotype 14, Streptococcus pneumoniae serotype 15B, and Streptococcus pneumoniae serotype 17F, Streptococcus pneumoniae serotype 18C, Streptococcus pneumoniae serotype 19F, Streptococcus pneumoniae serotype 19A, Streptococcus pneumoniae serotype 20, Streptococcus pneumoniae serotype 22F, Streptococcus pneumoniae serotype 23F, Streptococcus pneumoniae serotype 33F, Klebsiella pneumoniae serotype K1, Klebsiella pneumoniae serotype K2, Klebsiella pneumoniae serotype K5, Klebsiella pneumoniae serotype K16, Klebsiellapneumoniae serotype K20, Klebsiella pneumoniae serotype K54, Klebsiella pneumoniae serotype K57, Streptococcus agalactiae serotype Ia, Streptococcus agalactiae serotype Ib, Streptococcus agalactiae serotype II, Streptococcus agalactiae serotype III, Streptococcus agalactiae serotype IV, Streptococcus agalactiae serotype V, Streptococcus agalactiae serotype VI, Streptococcus agalactiae serotype VII, Streptococcus agalactiae serotype VIII, Streptococcus agalactiae serotype IX, Streptococcus pyogenes group A carbohydrate, Enterococcus faecalis serotype A, Enterococcus faecalis serotype B, Enterococcus faecalis serotype C, Enterococcus This vaccine is against serotype D of *Enterococcus faecium*, capsular polysaccharide and lipoteichoic acid, lipooligosaccharide A of *Moraxella catarrhalis*, lipooligosaccharide B of *Moraxella catarrhalis*, lipooligosaccharide C of *Moraxella catarrhalis*, and lipoteichoic acid of *Staphylococcus aureus*. In certain embodiments, the conjugate vaccine is useful because it induces an immune response when administered to a subject. In certain embodiments, the immune response is an antibody response that induces long-term memory (memory B cells and T cells), and is optionally a serotype-specific antibody response. In certain embodiments, the antibody response is an IgG or IgM response. For example, in certain embodiments, the antibody response may be an IgG response, and in certain embodiments, it may be an IgG1 response. In certain embodiments, the conjugate vaccine generates immunological memory in the subject to which the vaccine is administered.

[0164] Disclosed herein is a pneumococcal complex carbohydrate vaccine containing a conventional vaccine carrier, which can be produced by isolating a complex carbohydrate or glycosylated fusion protein of the Disclosure, including the ComP glycosylated fragment of the Disclosure, and combining the isolated complex carbohydrate or isolated glycosylated fusion protein with an adjuvant. In certain embodiments, the ComP glycosylated fragment can be attached to a conventional carrier protein, Pseudomonas aeruginosa exotoxin A (EPA). In certain embodiments, it has been demonstrated that the glycosylated fragment / carrier fusion protein can be paired with the CPS8 polysaccharide, and the use of PglS can produce a carrier protein-CPS8 bioconjugate, which is the first pneumococcal bioconjugate vaccine of this species. For example, in certain embodiments, the EPA fusion can be paired with the CPS8 polysaccharide, and the use of PglS can produce an EPA-CPS8 bioconjugate. The EPA-CPS8 bioconjugate vaccine has been demonstrated to induce a serotype 8-specific high IgG titer, which was determined to be protective by bactericidal toxicity. Importantly, the administration of as little as 100 ng of polysaccharide in the EPA-CPS8 bioconjugate was able to provide protection. Therefore, certain embodiments provide a CPS8 pneumococcal bioconjugate vaccine.

[0165] It is intended that conjugate vaccines (such as EPA vaccine constructs) may contain additional / multiple glycosylation sites to increase the glycan-to-protein ratio and the number of serotypes, in order to develop comprehensive pneumococcal bioconjugate vaccines.

[0166] Moraxellaceae O-linked oligosaccharide transferase Furthermore, a portion of this disclosure is a novel family of bacterial O-linked oligosaccharide transferases, referred to as TfpM, from Moraxellaceae bacteria. Certain embodiments of this disclosure include either a subsequent TfpM-associated pyrin glycosylation fragment and / or a TfpM OTase. The TfpM protein is similar in size and sequence to the TfpO enzyme but can transfer long-chain polysaccharides to receptor proteins. Phylogenetic analysis has demonstrated that the TfpM protein clusters into a different clad than known bacterial oligosaccharide transferases. Using a representative TfpM enzyme from Moraxella osloensis, it was determined that TfpM glycosylates the C-terminal threonine of its congener pyrin-like protein, and the minimum sequence required for glycosylation was identified. TfpM has been demonstrated to have broad substrate tolerance and can transfer a variety of polysaccharides, including those containing reducing-terminus glucose, galactose, or 2-N-acetyl sugars. Furthermore, bioconjugates derived from TfpM were shown to induce an immune response, specifically a serotype-specific polysaccharide IgG response in mice. Therefore, the diversity of TfpM's glycan substrates and the identification of minimal TfpM sequences make this enzyme a valuable additional tool for expanding the glycotechnology toolbox.

[0167] Bioinformatics-based identification of a new class of OTases found in Moraxellaceae bacteria. To identify the gene encoding the O-linked oligosaccharide transferase, the inventors first used the basic local alignment search tool (BLAST) and PglS. ADP1 The NCBI genome and whole-genome shotgun contig sequence databases were searched using the amino acid sequence (SEQ ID NO: 1) as a query. Sequence ID 1_PglS ADP1 amino acid sequence MNSIFKKIKNYTIVSGVFFLGSAFIIPNTSNLSSTLYKELIAVLGLLILLTVKSFDYKKILIPKNFYWFLFVIFIIFIQLIVGEIYFFQDFFFSISFLVILFLSFLLGFNERLNGDDLIVKKIAWIFIIVVQISFLI AINQKIEIVQNFFLFSSSYNGRSTANLGQPNQFSTLILITLFLLCYLREKNSLNNMVFNILSFCLIFANVMTQSRSAWISVILISLLYLLKFQKKIELRRVIFFNIVFWTLVYCVPLLFNLIFFQKNSYSTFDRLTM GSSRFEIWPQLLKAVFHKPFIGYGWGQTGVAQLETINKSSTKGEWFTYSHNLFLDLMLWNGFFIGLIISILILCFLIELYSSIKNKSDLFLFFCVVAFFVHCLLEYPFAYTYFLIPVGFLCGYISTQNIKNSISYFN LSKRKLTLFLGCCWLGYVAFWVEVLDISKKNEIYARQFLFSNHVKFYNIENYILDGFSKQLDFQYLDYCELKDKYQLLDFKKVAYRYPNASIVYKYYSISAEMKMDQKSANQIIRAYSVIKNQKIIKPKLKFCSIEY

[0168] Shorten the list of hits, PglS ADP1 To reduce the possibility of identifying very similar orthologues, the search is performed using PglS. ADP1 The candidates were further narrowed down to those with less than 50% amino acid sequence identity. Some of the top hits from this narrowed list were proteins much more similar in size to the TfpO protein, but the upstream congener pyrin proteins included both a sequence flanked by ComP disulfide and a C-terminal PilA-like sequence. The first pyrin-oligosaccharide transferase pairs identified were encoded by two Acinetobacter species: A. parvus DSM 16617 and A. townerii ZZC-3 (Table 1). [Table 3]

[0169] The two oligosaccharide transferases in these strains were closely related and exhibited >96% sequence identity. Intrigued by these findings, the inventors conducted further investigations and identified other strains of Moraxellaceae that possess genes encoding similar putative oligosaccharide transferase / pyrin pairs. Numerous genes were found, such as A. parvus DSM 16617 and A. townerii ZZC-3, while most of the associated pyrin proteins encoding upstream of the oligosaccharide transferases lacked sequences like ComP. The acceptance numbers and protein sizes of 20 of these putative oligosaccharide transferase-pyrin-like protein pairs are shown in Table 1. Since the inventors did not observe any homologs in species outside of Moraxellaceae, they designated these distinct oligosaccharide transferases as TfpM proteins ("M" for Moraxellaceae) to distinguish them from other known enzymes. Because the size of the TfpM protein is similar to that of the known TfpO protein, it was initially hypothesized that these genes might encode variants of the TfpO-PilA pair, similar to those found in Acinetobacter and Pseudomonas (Harding, CM, et al. (2015) Molecular Microbiology 96, 1023-1041). However, multiple sequence alignment of 20 TfpM proteins with known PglS, PglL, and TfpO proteins demonstrated that the former has less than 26% sequence identity with typical oligosaccharide transferases. Analysis of the phylogenetic tree generated by the multiple alignment showed that the TfpM protein clusters in a different clad from the TfpO, PglS, and PglL proteins (Figures 23A and 24). In contrast, the pyrin gene located immediately upstream of tfpM was not clustered into discrete clads (Figure 25) and showed high overall identity with the PilA protein, i.e., 37%–60%.Based on their coding sequences, most of the relevant pyrin proteins belong to the type IV major pyrin family, with the exception of pyrin proteins from Acinetobacter sp. CIP102143 and Acinetobacter sp. CIP102637, which lack a characteristic type III signaling sequence at their N-terminus (Giltner Carmen, L., et al. (2012) Microbiology and Molecular Biology Reviews 76, 740-772).

[0170] TfpM orthologs use the capsular polysaccharide of Streptococcus pneumoniae serotype 8 to glycosylate the manipulated pyrin fusion protein. Although similar in size to the TfpO protein, the TfpM protein is sufficiently different in terms of its amino acid sequence to warrant further investigation. The inventors were particularly interested in investigating whether the TfpM protein could transfer only short oligosaccharides to receptor proteins like the TfpO protein. Of the 20 TfpM oligosaccharide transferases listed in Table 1, the inventors selected 13 representative examples from different classes and tested their glycosylation activity in glycosylated E. coli strains (Harding, CM, and Feldman, MF (2019) Glycobiology 29, 519-529; Feldman, MF, et al. (2005) Proceedings of the National Academy of Sciences of the United States of America 102, 3016). Previously, the inventors developed a chimeric receptor protein strategy consisting of exotoxin A protein (EPA) from Pseudomonas aeruginosa fused with soluble fragments of different sizes of ComP (the natural substrate of PglS) (Knoot, CJ, et al. (2021) Glycobiology 31, 1192-1203). All type IV pyrin-like proteins contain a conserved N-terminal pyrin signaling sequence and membrane anchoring domain, which are not required for glycosylation but are essential for pyrin stability. The fusion protein approach allows for the removal of the conserved N-terminal pyrin signaling sequence and membrane anchoring domain, thus enabling the fusion of PglS ADP1This approach was used to determine the smallest sequence that could be recognized and still efficiently glycosylated (Harding, CM, et al. (2019) Nature Communications 10, 891; Knoot, CJ, et al. (2021) Glycobiology 31, 1192-1203). The inventors adopted this approach and designed 13 synthetic double-stranded DNA blocks encoding the N-terminal cleavage fragments of the upstream pyrin gene and the downstream tfpM gene. In most strains with tfpM, the protein-coding region of the upstream pyrin gene and the start codon of tfpM overlap by one nucleotide. This gene structure was left intact in the expression construct. The synthetic DNA blocks were designed so that, when cloned into an EPA expression vector using Gibson assembly, the pyrin-coding region would be positioned in-frame with the C-terminus of EPA, creating a gene fusion that would translate to a single protein with the immediately downstream tfpM gene (Figure 23B). The size of the cleaved pyrin fragments ranged from 113 to 140 amino acids. Previous studies using TfpO from P. aeruginosa 1244 reported that adding further C-terminal residues after the serine residue prevented glycosylation; therefore, a purification tag was not added to the C-terminus of the EPA-pyrin fusion (Horzempa, J., et al. (2006) Journal of Biological Chemistry 281, 1128-1136). Expression of EPA-pyrin and TfpM proteins was driven by an IPTG-inducible tac promoter on the pEXT20 plasmid (Dykxhoorn, DM, et al. (1996) Gene 177, 133-136). The fusion protein was secreted into the periplasm using the DsbA signaling sequence at the N-terminus of EPA. The oligonucleotides and primers used for assembly are listed in Table 2. [Table 4-1] [Table 4-2]

[0171] Using this design, the inventors evaluated the ability of 13 TfpM proteins to transfer Streptococcus pneumoniae capsular polysaccharide 8 (CPS8) glycan to the congeneral pyrin domain of an EPA-pyrin fusion. The CPS8 repeat unit is a tetrasaccharide with glucose at its reducing end. In particular, PglS is to date the only known oligosaccharide transferase capable of spontaneously transferring this glycan to a receptor protein (Harding, CM, et al. (2019) Nature Communications 10, 891). The 13 EPA-pyrin fusion / TfpM expression vectors were individually transformed into E. coli SDB1 strain expressing CPS8 glycan (Feldman, MF, et al. (2005) Proceedings of the National Academy of Sciences of the United States of America 102, 3016), and protein glycosylation was evaluated. The predicted mass of the non-glycosylated fusion protein was in the range of 78.3–80.5 kDa. Some TfpM proteins glycosylate their congeneral EPA-pyrin fusions, and glycosylation creates a high molecular weight ladder (g0) above the non-glycosylated band (g0). n It was found to appear as (Figure 23C). Each high-weight band represents the binding of a glycan with one additional CPS8 repeat unit to the EPA pyrin protein. Glycosylation was readily observed in the following seven TfpM orthologues examined: Acinetobacter sp. YZSX-1-1, Acinetobacter sp. CIP102637, Acinetobacter sp. YH01026, A. junii 65, Acinetobacter sp. CIP102143, Acinetobacter sp. TUM15069, and M. osloensis 1202 (Figure 23C). Some of these extracts required higher exposure of Western blot to observe the glycosylation pattern (Figure 23D). As a control, the inventors used conserved histidine (His), which has been previously shown to be necessary for the activity of wzy_C domain-containing enzymes.286 We generated mutants of the M.osloensis TfpM protein gene with a single residue change (Figure 26). The wzy_C family pfam04932 is an "O antigen ligase" domain found in membrane-binding enzymes that catalyze the transfer and covalent bonding of lipid-bound oligosaccharides (liposaccharides) to lipid A or protein substrates. (Ruan, X., et al. (2012) Glycobiology 22, 288-299, Musumeci, MA, et al. (2014) Glycobiology 24, 39-50).

[0172] In SDB1 cell extracts expressing this H286A mutant together with EPA-pyrin fusion protein and CPS8 glycan, glycosylation was not observed (Figure 23C), which suggests that glycosylation is due to the activity of the tfpM gene product, and His 286 This indicates that this is necessary for the catalytic activity and / or stability of this oligosaccharide transferase.

[0173] TfpMMo is an O-linked oligosaccharide transferase that glycosylates the threonine at the C-terminus of a pyrine substrate. Of the 13 TfpM-pyrine pairs examined in the previous experiment, M.osloensis 1202 and Acinetobacter sp.YH01026 showed the most efficient transfer of glycans of various sizes. Because the apparent stability of the pyrine from M.osloensis FDAARGOS_1202 (hereinafter referred to as 1202) was slightly higher, the inventors selected the oligosaccharide transferase from this organism as representative for further characterization, and the enzyme was identified as TfpM Mo (Sequence ID 56) was used to refer to the intact natural M. osloensis 1202 pyrin protein throughout the text. Mo As (SEQ ID NO: 57), and the N-terminally cleaved fusion domain is Pil Mo Refer to Δ28 (sequence number 58). Next, TfpM Mo Using Pil MoThe glycosylation site of Δ28 was identified to determine whether the enzyme acts like the TfpO protein to glycosylate the C-terminal amino acid of the congener pyrin receptor, or like the PglL or PglS proteins to glycosylate an internal residue. TfpO proteins typically transfer a short oligosaccharide containing 3–6 sugars to the side chain of the C-terminal serine residue of the corresponding pyrin. All but one of the congener pyrin proteins located immediately upstream of the tfpM open reading frame terminate at a C-terminal threonine residue; specifically, the one from Psychrobacter sp.72-Oc terminates at serine. Based on this observation and the size similarity of the TfpM and TfpO proteins, the inventors hypothesized that the TfpM enzyme also transfers a glycan to the C-terminal residue and designed a point mutant of pyrin to test this. C-terminal pyrin threonine (Thr 167 Two mutants were generated, and these residues were converted to serine or alanine, and glycosylation by CPS8 was examined. Mo Whole cell E. coli extracts from strains expressing the Δ28 mutant were probed using an anti-EPA antibody by Western blotting. As shown in Figure 27, TfpM Mo EPA-Pil Mo Δ28 T167S was also able to undergo glycosylation, but glycosylation was not observed in the alanine mutant. These results indicate that Thr 167 Or Ser 167 The hydroxyl group on the side chain of the residue is TfpM Mo This suggests that it is highly likely to be a glycan binding site.

[0174] TfpM Mo To confirm the site of pyring glycosylation, CPS8-glycosylated EPA-Pil Mo Δ28 was partially purified and separated via SDS-PAGE analysis, and the separated glycoprotein was stained with Coomassie stain. EPA-Pil glycosylated with 1-3 CPS8 repeat units. MoGel slices corresponding to Δ28 were excised, digested with LysC, and glycopeptides were analyzed. LysC-derived EPA-Pil Mo Open search-based analysis of Δ28 peptide (Chick, JM, et al. (2015) Nature Biotechnology 33, 743-749, Polasky, DA, et al. (2020) Nature Methods 17, 1125-1132) found that the hexose (Hex)-hexuronic acid (HexA) modification matches the incomplete monomer (HexHexAHex2) of CPS8 glycan. 762 FLPANCRGT 770 The peptide was identified. High-energy C-trap dissociation (HCD) analysis confirmed that this disaccharide is bound via a hexose residue, and targeted electron transfer / high-energy collision dissociation (EThcD) analysis confirmed the binding of HexHexA to the C-terminal threonine residue (Figures 28A and 28B). This MS analysis did not identify the CPS8 tetrasaccharide polymer, but this is not surprising, as it is extremely difficult to detect elongated complex carbohydrates using peptide-centered LC-MS without the use of special chemical additives such as superchargers (Lin, C.-w., et al. (2016) Analytical Chemistry 88, 8484-8494). Nevertheless, the identification of a disaccharide consistent with the partially completed CPS8 tetrasaccharide still supports the idea that it is a CPS8 tetrasaccharide polymerized from a high molecular weight ladder.

[0175] TfpM Mo It transfers a polysaccharide containing glucose, galactose, or a 2-N-acetyl monosaccharide to the reducing end. Next, TfpM in the background of glycosylated E. coli MoThe aim was to investigate the range of polysaccharide substrates to which TfpM can transfer. The inventors selected polymers consisting of different reducing end sugars, polysaccharides containing different disaccharide bonds near the reducing end, and / or linear or branched repeating units. In addition to Streptococcus pneumoniae CPS8, four polysaccharides—E. coli O16 antigen, Salmonella enterica LT2 O antigen, Klebsiella pneumoniae O2a antigen, and Streptococcus capsular polysaccharide of type III B—were selected. Mo It was tested as a glycan substrate. The structures of all five tested repeating units are shown in Figure 29A (Liu, B., et al. (2020) FEMS Microbiology Reviews 44, 655-683, Curd, H., et al. (1998) Journal of Bacteriology 180, 1002-1007, Whitfield, C., et al. (1992) Journal of Bacteriology 174, 4913-4919, Pinto, V., and Berti, F. (2014) Journal of Pharmaceutical and Biomedical Analysis 98, 9-15, Geno, KA, et al. (2015) Clinical Microbiology Reviews 28, 871-899). Each of the five polysaccharides was found to be EPA-Pil in E. coli SDB1 cells. Mo Δ28 and TfpM MoEach polysaccharide was co-expressed individually, induced, and then grown for subsequent glycoprotein purification. Periplasm extracts of SDB1 cells were partially purified using anion exchange chromatography to remove any contamination with undecaprenol pyrophosphate-bound polysaccharides that could complicate Western blot interpretation. To demonstrate that the glycan-specific antibody signals observed in Western blot were actually derived from glycosylated proteins and not from lipid-bound polysaccharides contaminating the whole cell lysate, the purified glycoproteins were split into two equal fractions, and half of each was digested with proteinase K before SDS-PAGE separation and Western blotting. Since all antibodies used in this experiment were rabbit-derived, Western blots were examined using antiserum specific to each polysaccharide, and separately, with an anti-EPA antibody. As shown in Figure 7, TfpM Mo It uses all five different polysaccharides in EPA-Pil Mo It was found that the Δ28 protein was efficiently transferred. Digestion with protease K removed both the anti-glycan signal (Figures 29B, 29C, 29D, 29E, and 29F) and the anti-EPA signal (Figure 29G) from the Western blot, confirming that the anti-glycan signal originated from protein-bound polysaccharides, not from contaminating lipid-bound polysaccharides.

[0176] TfpM Mo is a severed Pil Mo The Δ28 variant can be glycosylated. All previous experiments have used PilA, an amino acid with a length of 139. Mo The N-terminal cleavage variant was used. TfpM Mo To gain insight into the minimum features required for C-terminal pyring glycosylation by Pil Mo A series of further cleaved variants of Δ28 were generated and tested to see if they could be glycosylated. The inventors first fused Pil to the C-terminus of EPA via a flexible four-residue glycine linker. Mo It generates 20 amino acid fragments, Pil 20This was named (Figure 30A). These 20 amino acid fragments were selected because they contain a disulfide loop ("DSL") region that is conserved in many type IV pyrines (Figure 31) (Horzempa, J., et al. (2006) Journal of Biological Chemistry 281, 1128-1136, Harvey, H., et al. (2009) Journal of Bacteriology 191, 6513-6524). It is noteworthy that this DSL corresponds to a different motif from the disulfide-sandwiched sequence present in the ComP protein (Knoot, CJ, et al. (2021) Glycobiology 31, 1192-1203). Based on sequence alignment with P. aeruginosa 1244 PilA, the DSL of M. osloensis pyrine is the residue Cys 148 and Cys 164 It is formed by the following. Therefore, the inventors believe that the downstream arrangement of the glycine linker is Cys 148 It was designed to start with EPA-Pil. 20 The plasmid containing the construct encoding TfpM was named pVNM297. 20 Glycosylation experiments using Pil Mo TfpM with CPS8 at a similar level to Δ28 Mo It was revealed that glycosylation is possible by (Figure 29B). To test whether the DSL region is necessary for glycosylation, several short variants lacking part of this feature were generated. In these smaller constructs, the inventors also added Cys to prevent the formation of unnatural disulfide bonds during oxidation at the periplasm. 164 The 15-, 13-, and 10-amino acid Pil was mutated to alanine (Harvey, H., et al. (2009) Journal of Bacteriology 191, 6513-6524). Mo The variants were tested, and one variant of the 10-amino acid version contained an amino acid linker (e.g., a GGGG linker), while the other variant did not. These constructs were then analyzed by Pil15 Pil 13 Pil 10L , and Pil 10 It was called (Figure 30A). TfpM Mo Pil 20 and Pil Mo Although at a level lower than Δ28, all four of these variants could be glycosylated (Figure 27B). 10 and Pil 10L Both are similarly relatively glycosylated, and the presence of an upstream glycine linker is TfpM Mo This shows that it does not significantly affect the activity of [protein name]. This linker was omitted from all subsequent constructs. Of the four proteins, Pil 13 It exhibited significantly worse glycosylation than other proteins.

[0177] From sequence alignment, it was noted that each pyrin protein glycosylated by TfpM OTases possesses a conserved "PAN / ECRG" motif near the C-terminus, just upstream of the second-to-last threonine residue (Figure 31). Since this feature is present in all glycosylated pyrin proteins, this motif is considered to be a TfpM motif. Mo It was questioned whether this was necessary for glycosylation. The inventors found a similar motif (cysteine ​​mutated to alanine - the residue in bold) [Table 5] Pil (modified to) Mo A 7-amino acid variant called Pil7 was fused to EPA, and its glycosylation was evaluated (Figure 30A). The inventors also constructed stepwise single-amino acid cleavage from 7 amino acids to 2 amino acids of this "PANARGT" sequence and evaluated the ability of TfpM to glycosylate these fragments with CPS8. The results showed that all variants except Pil2 were glycosylated. 10 TfpM at a similar level Mo This showed that glycosylation occurred by (Figure 30B). Overall glycosylation was due to Pil 20and Pil Mo Δ28 was even lower. Glycosylated Pil2 was hardly detected, but increasing the exposure of the Western blot revealed several trace amounts of ladder. This is TfpM Mo This shows that this variant can be glycosylated, but the level is significantly lower than that of the construct with one additional amino acid. From these results, TfpM Mo When it fuses to the C-terminus of a heterologous EPA protein, it forms a 3-amino acid Pil Mo It was concluded that the fragments can be recognized and glycosylated.

[0178] TfpM Mo Immunogenicity of GBSIII-derived bioconjugates EPA-Pil 20 Considering that the construct was efficiently glycosylated by TfpM, the inventors then proceeded to glycosylate EPA-Pil with type III capsular polysaccharide (GBSIII) from group B Streptococcus. 20 The immunogenicity of the protein in a mouse vaccination model was evaluated. EPA-Pil was used to assist in protein purification for these experiments. 20 A plasmid expressing the N-terminal 6x-His tag variant of the carrier protein (pVNM291) was constructed. The His tag was added immediately downstream of the predicted N-terminal DsbA signaling sequence cleavage site of EPA. pVNM291 was introduced into SDB1 cells expressing GBSIII glycan, and the resulting bioconjugate was purified by nickel-fixed metal affinity chromatography (IMAC), followed by anion exchange chromatography and size exclusion chromatography on FPLC. Western blotting and Coomassie staining of the GBSIII-291 bioconjugate, which was degraded by SDS-PAGE, revealed EPA-Pil 20 High molecular weight glycosylation of the protein and GBSIII glycan was confirmed (Figures 32A, 32B, 32C, and 32D). Purified EPA-Pil 20-Intact protein MS of the GBSIII ("GBSIII-291") conjugate supported a 20% glycan:protein ratio (Figure 32E). Each dose was formulated to contain 1 μg of GBSIII polysaccharide. As a control for these experiments, a carrier protein ("291") derived from non-glycosylated pVNM291 was purified from SDB1 cells without glycan plasmids and administered at the same protein concentration as the GBSIII bioconjugate.

[0179] TfpM Mo To test whether the GBSIII-291 bioconjugate produced was immunogenic, 5-week-old female CD-1 mice were immunized. Mice were administered either placebo (non-glycosylated 291 carrier protein) or the GBSIII-291 bioconjugate at 2-week intervals, starting with a priming dose, followed by two booster inoculations. All vaccines were formulated using Alhydrogel® 2% in a 1:9 ratio as an adjuvant. Serum was collected before each immunization and 2 weeks after the last booster inoculation. To determine the level of induced GBSIII-specific antibodies, the inventors used enzyme-linked immunosorbent assay (ELISA). All mice immunized with the GBSIII-291 bioconjugate showed high levels of anti-GBSIII IgG antibody expression, although one mouse with a low anti-GBSIII IgG response was able to boost during the immunization process (Figure 32F). As expected, GBSIII-specific IgG titers were increased in GBSIII-conjugate-vaccinated mice compared to mock-vaccinated mice (only 291, Figure 32F). In summary, these data support TfpM Mo However, this suggests that it is possible to generate bioconjugates that can induce polysaccharide-specific IgG responses.

[0180] TfpM Mo and PglS ADP1 (PglL ComP ) glycosylates a single protein that has been engineered to contain a sequence specific to each oligosaccharide transferase. Finally, the inventors wanted to determine whether a protein designed to contain sequenceons from two different OTase systems would be glycosylated by both OTases at each site. Therefore, an EPA fusion protein containing a sequenceon associated with TfpM and a sequenceon associated with PglS was constructed. For this purpose, the inventors used the residue Ala, as previously described (Knoot, CJ, et al. (2021) Glycobiology 31, 1192-1203). 489 and Arg 490 PglS Sequence [Table 6] Furthermore, Pil at the C-terminus 20 Seek On [Table 7] An EPA fusion protein containing was engineered (Figure 33A). As previously described, this construct was designed so that the open reading frame of the gene encoding the EPA fusion and the start codon of tfpM overlap by one nucleotide. The open reading frame encoding pglS from A. baylyi ADP1 was cloned 100 bp downstream of the stop codon of the tfpM open reading frame. This vector (pVNM337) was introduced into E. coli SDB1 expressing the E. coli O16 antigen, and glycosylation was examined by Western blotting. To compare with a protein containing only a single sequence, we introduced (i) TfpM-associated Pil into E. coli SDB1 expressing the O16 antigen. 20 EPA (pVNM297) containing only sequons, or (ii) residue Ala 489 and Arg 490 A construct of EPA(pVNM167) containing a PglS sequence integrated between was individually introduced. To compare with a protein having two sequencens and capable of dicly sylated, the inventors also introduced a residue Ala 489 and Arg490 Between and, as well as the residue Glu 548 and Gly 549 We also introduced an EPA construct (pVNM245) (EPA_PglS-Sequon2X) containing a sequon integrated between the two. Diagrams of these structures are shown in Figure 33A. As shown in Figure 33B, ComP or Pil Mo Western blot analysis of EPA constructs containing only one of the sequenceons showed a glycan profile of approximately 100 kDa indicating monoglycosylation. EPA constructs containing two PglS sequenceons showed a predominantly monoglycosylation profile around 100 kDa, but also exhibited a diglycosylation population shifting to around 150 kDa. TfpM Mo and PglS ADP1 Western blot analysis of the EPA fusion containing the sequence from revealed both monoglycosylation and diglycosylation populations, similar to those observed in construct pVNM245. These results led to the conclusion that the receptor protein can be glycosylated by two different OTase classes within a single expression system.

[0181] Complex carbohydrates This disclosure provides a complex carbohydrate comprising an oligosaccharide or polysaccharide covalently bound to a receptor protein. In certain embodiments, the receptor protein comprises or consists of a TfpM-related pyrine-like protein or a glycosylated fragment thereof. In certain embodiments, the oligosaccharide or polysaccharide is covalently bound to the pyrine-like protein or a glycosylated fragment thereof. In other embodiments, the TfpM-related pyrine-like protein or a glycosylated fragment thereof comprises a C-terminal serine or threonine residue, and the oligosaccharide or polysaccharide is covalently bound to the C-terminal serine or threonine. Furthermore, in certain embodiments, the receptor protein is a fusion protein comprising a TfpM-related pyrine-like protein or a glycosylated fragment thereof translatedly fused / bound to a heterogeneous amino acid sequence (e.g., a carrier protein), the TfpM-related pyrine-like protein or its glycosylated fragment being the most C-terminal sequence of the receptor protein, and consequently the receptor protein containing a C-terminal serine or threonine residue, and the oligosaccharide or polysaccharide covalently bound to the C-terminal serine or threonine. Specific examples of carrier proteins include, but are not limited to, Pseudomonas aeruginosa exotoxin A (EPA), CRM197, cholera toxin B subunit, tetanus toxin C fragment, or any fragment thereof. In certain embodiments, the TfpM-related pyrine-like protein or its glycosylated fragment is translatedly fused / bound to the heterogeneous amino acid sequence / carrier protein via an amino acid linker. In certain embodiments, the oligosaccharide or polysaccharide contains glucose at its reducing end. In certain embodiments, the complex carbohydrate is immunogenic.

[0182] In certain embodiments of the complex carbohydrates of this disclosure, the receptor protein comprises or consists of a full-length TfpM-related pyrin-like protein. In certain embodiments, the receptor protein comprises or consists of a glycosylated fragment of a TfpM-related pyrin-like protein that is shorter than the full-length TfpM-related pyrin-like protein. In certain embodiments, the length of the glycosylated fragment of the pyrin-like protein is 3 to 138 amino acids, 10 to 138 amino acids, 20 to 138 amino acids, 50 to 138 amino acids, 100 to 138 amino acids, or 116 to 138 amino acids. In certain embodiments, the length of the glycosylated fragment of the pyrin-like protein is 3 to 139 amino acids, 10 to 139 amino acids, 20 to 139 amino acids, 50 to 139 amino acids, 100 to 139 amino acids, or 116 to 139 amino acids. In certain embodiments, the length of the glycosylated fragment is 3 to 140 amino acid lengths, 10 to 140 amino acid lengths, 20 to 140 amino acid lengths, 50 to 140 amino acid lengths, 100 to 140 amino acid lengths, or 116 to 140 amino acid lengths. In certain embodiments, the length of the glycosylated fragment is 3 to 22 amino acid lengths, 10 to 22 amino acid lengths, 11 to 22 amino acid lengths, 3 to 21 amino acid lengths, 5 to 21 amino acid lengths, 10 to 21 amino acid lengths, or 11 to 21 amino acid lengths. In certain embodiments, the length of the glycosylated fragment has amino acid lengths ranging from any of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 to any of 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25.

[0183] In certain embodiments of the complex carbohydrates of this disclosure, TfpM-related pyrin-like protein or its glycosylated fragment is Pil Mo (Sequence ID 57), or Pil lacking amino acids corresponding to residues 1-28. Mo (Pil Mo Polypeptides that are or contain at least Δ28 (SEQ ID NO: 58) have at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity with respect to SEQ ID NO: 57 or SEQ ID NO: 58, for example, the C-terminal threonine is substituted with serine. In certain embodiments of the complex carbohydrates of this disclosure, the TfpM-related pyrine-like protein is 、 Pil DSM16617 (Sequence ID 82), Pil ZZC3-9 (Sequence No. 83) 、 Pil TUM15069 (Sequence ID 84), Pil AI7 (Sequence ID 85), Pil VE-C3 (Sequence number 86), Pil YH01026 (Sequence ID 87), Pil CIP102143 (Sequence number 88), Pil AI40 (Sequence ID 89), Pil F78 (Sequence ID 90), Pil S71 (Sequence ID 91), Pil ANC4282 (Sequence ID 92), Pil 72-O-c (Sequence ID 93), Pil BI730 (Sequence ID 94), Pil A3K91 (Sequence ID 95), Pil CIP102159 (Sequence ID 96), Pil junii-65 (Sequence ID 97), Pil YZS-X (Sequence ID 98), Pil T-3-2 (Sequence ID 99), and Pil CIP102637 Selected from the group consisting of (Sequence ID 100). In a particular embodiment, the TfpM-related pyrin-like protein or pyrin-like protein glycosylated fragment is Pil DSM16617 (Sequence ID 82), Pil ZZC3-9 (Sequence ID 83), Pil TUM15069 (Sequence ID 84), Pil AI7 (Sequence ID 85), Pil VE-C3(Sequence number 86), Pil YH01026 (Sequence ID 87), Pil CIP102143 (Sequence number 88), Pil AI40 (Sequence ID 89), Pil F78 (Sequence ID 90), Pil S71 (Sequence ID 91), Pil ANC4282 (Sequence ID 92), Pil 72-O-c (Sequence ID 93), Pil BI730 (Sequence ID 94), Pil A3K91 (Sequence ID 95), Pil CIP102159 (Sequence ID 96), Pil junii-65 (Sequence ID 97), Pil YZS-X (Sequence ID 98), Pil T-3-2 (Sequence ID 99), Pil CIP102637 The amino acid sequences selected from the group consisting of (SEQ ID NO: 100), and any fragment thereof (e.g., a C-terminal fragment), and / or variants in which the C-terminal threonine is substituted with serine, are included or consist of the above. In a particular embodiment, the pyrin-like protein glycosylated fragment is Pil Mo Pilling disulfide loop region (Pil Mo _DSL, Pil 20 Also known as (SEQ ID NO: 60), or its cleavage derivatives containing at least the last three amino acids from the C-terminus of pyrin, or variants in which the C-terminal threonine is substituted with serine (SEQ ID NO: 148), comprising or consisting thereof. Furthermore, in certain embodiments, the pyrin-like protein glycosylated fragment is Pil 20 (Sequence ID 60), Pil 19 (Sequence ID 133), Pil 18 (Sequence ID 134), Pil 17 (Sequence ID 135), Pil 16 (Sequence ID 136), Pil 15 (Sequence ID 109), Pil 14 (Sequence ID 137), Pil 13 (Sequence ID 110), Pil 12 (Sequence ID 138), Pil 11 (Sequence ID 139), Pil 10Pil (SEQ ID NO: 112), Pil9 (SEQ ID NO: 140), Pil8 (SEQ ID NO: 141), Pil7 (SEQ ID NO: 113), Pil6 (SEQ ID NO: 114), Pil5 (SEQ ID NO: 115), Pil4 (SEQ ID NO: 116), or Pil3 (SEQ ID NO: 117), or variants thereof having one, two, three, four, or five amino acid substitutions and maintaining the C-terminal threonine. Furthermore, in certain embodiments, the pyrin-like protein glycosylated fragment is Pil 20S (Sequence ID 148), Pil 19S (Sequence ID 149), Pil 18S (Sequence ID 150), Pil 17S (Sequence ID 151), Pil 16S (Sequence ID 152), Pil 15S (Sequence ID 153), Pil 14S (Sequence ID 154), Pil 13S (Sequence ID 155), Pil 12S (Sequence ID 156), Pil 11S (Sequence ID 157), Pil 10S (Sequence ID 158), Pil 9S (Sequence ID 159), Pil 8S (Sequence ID 160), Pil 7S (Sequence ID 161), Pil 6S (Sequence ID 162), Pil 5S (Sequence ID 163), Pil 4S (Sequence ID 164), or Pil 3S (SEQ ID NO: 165), or variants thereof having one, two, three, four, or five amino acid substitutions while maintaining the C-terminal serine.

[0184] In certain embodiments of the complex carbohydrates of this disclosure, the receptor protein may be glycosylated at two or more different positions. In certain embodiments, the receptor protein may be glycosylated by at least two different OTase classes within a single expression system. For example, in certain embodiments, the receptor protein is a fusion protein, which further comprises an additional glycosylated sequence (e.g., a glycosylated fragment) of an OTase other than TfpM oligosaccharide transferase (OTase), in addition to the glycosylated fragment of a TfpM-related pyrin-like protein located at its C-terminus. For example, the other OTase may be PglB, PglL, or PglS. In certain embodiments, the additional glycosylated sequence is an internal sequence of the fusion protein (i.e., not the majority of the C-terminus or N-terminus). In certain embodiments, the additional glycosylated sequence is an internal sequence within the sequence of the carrier protein (e.g., Figure 33A). In certain embodiments, the additional glycosylated sequence is also covalently bonded to an oligosaccharide or polysaccharide. In certain embodiments, the fusion protein contains two or more, three or more, four or more, five or more, six or more, eight or more, ten or more, fifteen or more, or twenty or more additional glycosylation sequences. In certain embodiments, the fusion protein does not contain more than two, three or five, ten or fifteen or twenty or more additional glycosylation sequences. In certain embodiments, the additional glycosylation sequences are the same. In certain embodiments, at least one additional glycosylation sequence is different from the others. In certain embodiments, at least three, four or five additional glycosylation sequences are all different from each other. Also, in certain embodiments, none of the additional glycosylation sequences are the same. In certain embodiments, the receptor protein is a fusion protein, and the fusion protein further contains an internal glycosylation fragment of ComP in addition to the glycosylation fragment of the TfpM-related pyrin-like protein located at its C-terminus. In certain embodiments, the ComP glycosylated fragment is located within the sequence of the carrier protein. In certain embodiments, the ComP glycosylated fragment is also covalently bonded to an oligosaccharide or polysaccharide.Furthermore, in certain embodiments, the ComP glycosylated fragment is... [Table 8] Or it comprises or consists of a fragment having at least one amino acid ASA at positions 11-13. In certain embodiments, the fusion protein contains two or more, three or more, four or more, five or more, six or more, eight or more, ten or more, fifteen or more, or twenty or more ComP glycosylated fragments. In certain embodiments, the fusion protein does not contain more than two, three or more, five or more, ten or more, fifteen or more, twenty or more ComP glycosylated fragments. In certain embodiments, the ComP glycosylated fragments are the same. In certain embodiments, the ComP glycosylated fragments are different from each other. Also, in certain embodiments, at least three, at least four, or at least five ComP glycosylated fragments are all different from each other. Also, in certain embodiments, none of the ComP glycosylated fragments are the same.

[0185] In certain embodiments of the complex carbohydrates of this disclosure, the oligosaccharide or polysaccharide covalently bonded to the pyrin-like protein or its glycosylated fragment has the size of at least three repeating units of the oligosaccharide or polysaccharide structure. In certain embodiments of the complex carbohydrates of this disclosure, the oligosaccharide or polysaccharide covalently bonded to the pyrin-like protein or its glycosylated fragment has the size of at least ten monosaccharides.

[0186] In certain embodiments of the complex carbohydrates of this disclosure, the oligosaccharide or polysaccharide is produced by a bacterium of the genus Streptococcus (e.g., S. pneumoniae or S. agalactiae), and the polysaccharide is a capsular polysaccharide such as Ia, Ib, II, III, IV, V, VI, VII, VIII, or IX.

[0187] In certain embodiments of the complex carbohydrates of this disclosure, the oligosaccharide or polysaccharide is produced by Klebsiella (e.g., K. pneumoniae), and the polysaccharide is a capsular polysaccharide or O antigen polysaccharide.

[0188] In certain embodiments of the complex carbohydrates of this disclosure, the oligosaccharide or polysaccharide is produced by a bacterium of the genus Salmonella, and the polysaccharide is an O antigen polysaccharide. In certain embodiments, the bacterium is S. enterica, and the S. enterica polysaccharide is an O antigen of group B.

[0189] In certain embodiments of the complex carbohydrates of this disclosure, the complex carbohydrates are produced in vivo, such as in bacterial cells. In certain embodiments, the bacteria are Escherichia coli. In certain embodiments, the bacteria are from the genus Klebsiella. In certain embodiments, the bacterial species are K. pneumoniae, K. varricola, K. michinganenis, or K. oxytoca. In certain embodiments, the complex carbohydrates are produced in a cell-free system.

[0190] In certain embodiments of the complex carbohydrates of this disclosure, the bioconjugate is a conjugate vaccine that, upon administration to a subject, induces an immune response. In certain embodiments, the immune response is an antibody response that induces long-term memory (memory B cells and T cells), and is optionally a serotype-specific antibody response. In certain embodiments, the antibody response is an IgG or IgM response. In certain embodiments, the antibody response is an IgG response, e.g., an IgG1 response. Also in certain embodiments, the conjugate vaccine generates immunomemory in the subject to which the vaccine is administered.

[0191] Glycosylated fragment This disclosure provides a pyrin-like protein glycosylated fragment comprising or consisting of an isolated fragment of the TfpM-related pyrin-like protein of this disclosure. In certain embodiments, the TfpM-related pyrin-like protein or pyrin-like protein glycosylated fragment is Pil Mo (Sequence ID 57), or Pil lacking amino acids corresponding to residues 1-28. Mo (Pil MoPolypeptides that are or contain at least Δ28 (SEQ ID NO: 58) have at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity to SEQ ID NO: 57 or SEQ ID NO: 58, for example, the C-terminal threonine is substituted with serine. In certain embodiments, the TfpM-related pyrine-like protein is Pil DSM16617 (Sequence ID 82), Pil ZZC3-9 (Sequence No. 83) 、 Pil TUM15069 (Sequence ID 84), Pil AI7 (Sequence ID 85), Pil VE-C3 (Sequence number 86), Pil YH01026 (Sequence ID 87), Pil CIP102143 (Sequence number 88), Pil AI40 (Sequence ID 89), Pil F78 (Sequence ID 90), Pil S71 (Sequence ID 91), Pil ANC4282 (Sequence ID 92), Pil 72-O-c (Sequence ID 93), Pil BI730 (Sequence ID 94), Pil A3K91 (Sequence ID 95), Pil CIP102159 (Sequence ID 96), Pil junii-65 (Sequence ID 97), Pil YZS-X (Sequence ID 98), Pil T-3-2 (Sequence ID 99), and Pil CIP102637 Selected from the group consisting of (Sequence ID 100). In a particular embodiment, the TfpM-related pyrin-like protein or pyrin-like protein glycosylated fragment is Pil DSM16617 (Sequence ID 82), Pil ZZC3-9 (Sequence ID 83), Pil TUM15069 (Sequence ID 84), Pil AI7 (Sequence ID 85), Pil VE-C3 (Sequence number 86), Pil YH01026 (Sequence ID 87), Pil CIP102143 (Sequence number 88), Pil AI40 (Sequence ID 89), Pil F78 (Sequence ID 90), Pil S71 (Sequence ID 91), Pil ANC4282 (Sequence ID 92), Pil72-O-c (Sequence ID 93), Pil BI730 (Sequence ID 94), Pil A3K91 (Sequence ID 95), Pil CIP102159 (Sequence ID 96), Pil junii-65 (Sequence ID 97), Pil YZS-X (Sequence ID 98), Pil T-3-2 (Sequence ID 99), Pil CIP102637 The present invention comprises or comprises amino acid sequences selected from the group consisting of (SEQ ID NO: 100), and any fragment thereof (e.g., a C-terminal fragment), and / or variants in which the C-terminal threonine is substituted with serine. In a particular embodiment, the pyrin-like protein glycosylation fragment comprises a PilMo pyrin disulfide loop region (Pil Mo _DSL, Pil 20 Also known as (SEQ ID NO: 60), or its cleavage derivatives containing at least the last three amino acids from the C-terminus of pyrin, or variants in which the C-terminal threonine is substituted with serine (SEQ ID NO: 148), comprising or consisting thereof. Furthermore, in certain embodiments, the pyrin-like protein glycosylated fragment is Pil 20 (Sequence ID 60), Pil 19 (Sequence ID 133), Pil 18 (Sequence ID 134), Pil 17 (Sequence ID 135), Pil 16 (Sequence ID 136), Pil 15 (Sequence ID 109), Pil 14 (Sequence ID 137), Pil 13 (Sequence ID 110), Pil 12 (Sequence ID 138), Pil 11 (Sequence ID 139), Pil 10Pil (SEQ ID NO: 112), Pil9 (SEQ ID NO: 140), Pil8 (SEQ ID NO: 141), Pil7 (SEQ ID NO: 113), Pil6 (SEQ ID NO: 114), Pil5 (SEQ ID NO: 115), Pil4 (SEQ ID NO: 116), or Pil3 (SEQ ID NO: 117), or variants thereof having one, two, three, four, or five amino acid substitutions and maintaining the C-terminal threonine. Furthermore, in certain embodiments, the pyrin-like protein glycosylated fragment is Pil 20S (Sequence ID 148), Pil 19S (Sequence ID 149), Pil 18S (Sequence ID 150), Pil 17S (Sequence ID 151), Pil 16S (Sequence ID 152), Pil 15S (Sequence ID 153), Pil 14S (Sequence ID 154), Pil 13S (Sequence ID 155), Pil 12S (Sequence ID 156), Pil 11S (Sequence ID 157), Pil 10S (Sequence ID 158), Pil 9S (Sequence ID 159), Pil 8S (Sequence ID 160), Pil 7S (Sequence ID 161), Pil 6S (Sequence ID 162), Pil 5S (Sequence ID 163), Pil 4S (Sequence ID 164), or Pil 3S (SEQ ID NO: 165), or variants thereof having one, two, three, four, or five amino acid substitutions while maintaining the C-terminal serine.

[0192] In certain embodiments of the glycosylated fragments of the present invention, the length of the isolated fragment of the TfpM-related pyrin-like protein of this disclosure is 3 to 138 amino acid lengths, 10 to 138 amino acid lengths, 20 to 138 amino acid lengths, 50 to 138 amino acid lengths, 100 to 138 amino acid lengths, or 116 to 138 amino acid lengths. In certain embodiments, the length of the glycosylated fragment is 3 to 139 amino acid lengths, 10 to 139 amino acid lengths, 20 to 139 amino acid lengths, 50 to 139 amino acid lengths, 100 to 139 amino acid lengths, or 116 to 139 amino acid lengths. In certain embodiments, the length of the glycosylated fragment is 3 to 140 amino acid lengths, 10 to 140 amino acid lengths, 20 to 140 amino acid lengths, 50 to 140 amino acid lengths, 100 to 140 amino acid lengths, or 116 to 140 amino acid lengths. In certain embodiments, the length of the glycosylated fragment is 3 to 22 amino acid lengths, 10 to 22 amino acid lengths, 11 to 22 amino acid lengths, 5 to 21 amino acid lengths, 10 to 21 amino acid lengths, or 11 to 21 amino acid lengths. In certain embodiments, the length of the glycosylated fragment is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 amino acid lengths, up to 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acid lengths. In certain embodiments, the length of the glycosylated fragment has the lengths of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids.

[0193] Fusion protein Fusion proteins comprising a TfpM-related pyrin-like protein or a glycosylated fragment thereof, translatedly fused / bound to a heterologous carrier protein, such as, but not limited to, Pseudomonas aeruginosa exotoxin A (EPA), CRM197, cholera toxin B subunit, tetanus toxin C fragment, or any fragment thereof, are provided herein. In certain embodiments, the TfpM-related pyrin-like protein or its glycosylated fragment is translatedly fused / bound to the heterologous carrier protein via an amino acid linker. In certain embodiments, the pyrin-like protein or glycosylated fragment contains a C-terminal serine or threonine residue. In certain embodiments, the pyrin-like protein or glycosylated fragment is the most C-terminal sequence of the fusion protein. Also in certain embodiments, the fusion protein contains a C-terminal serine or threonine residue. In certain embodiments, the fusion protein is glycosylated by an oligosaccharide or polysaccharide covalently bonded to the C-terminal serine or threonine. Furthermore, in certain embodiments, the fusion protein is glycosylated by an oligosaccharide or polysaccharide containing glucose at its reducing end, covalently bonded to a serine or threonine at its C-terminus. In certain embodiments, the glycosylated fusion protein is immunogenic. In certain embodiments, the glycosylated fusion protein is a conjugate vaccine.

[0194] In certain embodiments of the fusion protein of this disclosure, the fusion protein comprises a full-length TfpM-related pyrin-like protein. In certain embodiments, the fusion protein comprises or comprises a glycosylated fragment of a TfpM-related pyrin-like protein that is shorter than the full-length TfpM-related pyrin-like protein. In certain embodiments, the length of the glycosylated fragment of the pyrin-like protein is 3 to 138 amino acids, 10 to 138 amino acids, 20 to 138 amino acids, 50 to 138 amino acids, 100 to 138 amino acids, or 116 to 138 amino acids. In certain embodiments, the length of the glycosylated fragment is 3 to 139 amino acids, 10 to 139 amino acids, 20 to 139 amino acids, 50 to 139 amino acids, 100 to 139 amino acids, or 116 to 139 amino acids. In certain embodiments, the length of the glycosylated fragment is 3 to 140 amino acid lengths, 10 to 140 amino acid lengths, 20 to 140 amino acid lengths, 50 to 140 amino acid lengths, 100 to 140 amino acid lengths, or 116 to 140 amino acid lengths. In certain embodiments, the length of the glycosylated fragment is 3 to 22 amino acid lengths, 10 to 22 amino acid lengths, 11 to 22 amino acid lengths, 5 to 21 amino acid lengths, 10 to 21 amino acid lengths, or 11 to 21 amino acid lengths. In certain embodiments, the length of the glycosylated fragment has amino acid lengths ranging from any of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 to any of 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25.

[0195] In certain embodiments of the fusion protein disclosed herein, the pyrin-like protein glycosylated fragment is Pil Mo(Sequence ID 57), or Pil lacking amino acids corresponding to residues 1-28. Mo (Pil Mo Polypeptides that are or contain at least Δ28 (SEQ ID NO: 58) have at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity to SEQ ID NO: 57 or SEQ ID NO: 58, for example, the C-terminal threonine is substituted with serine. In certain embodiments, the TfpM-related pyrine-like protein is Pil DSM16617 (Sequence ID 82), Pil ZZC3-9 (Sequence No. 83) 、 Pil TUM15069 (Sequence ID 84), Pil AI7 (Sequence ID 85), Pil VE-C3 (Sequence number 86), Pil YH01026 (Sequence ID 87), Pil CIP102143 (Sequence number 88), Pil AI40 (Sequence ID 89), Pil F78 (Sequence ID 90), Pil S71 (Sequence ID 91), Pil ANC4282 (Sequence ID 92), Pil 72-O-c (Sequence ID 93), Pil BI730 (Sequence ID 94), Pil A3K91 (Sequence ID 95), Pil CIP102159 (Sequence ID 96), Pil junii-65 (Sequence ID 97), Pil YZS-X (Sequence ID 98), Pil T-3-2 (Sequence ID 99), and Pil CIP102637 Selected from the group consisting of (Sequence ID 100). In a particular embodiment, the TfpM-related pyrin-like protein or pyrin-like protein glycosylated fragment is Pil DSM16617 (Sequence ID 82), Pil ZZC3-9 (Sequence ID 83), Pil TUM15069 (Sequence ID 84), Pil AI7 (Sequence ID 85), Pil VE-C3 (Sequence number 86), Pil YH01026 (Sequence ID 87), Pil CIP102143 (Sequence number 88), Pil AI40 (Sequence ID 89), Pil F78(Sequence ID 90), Pil S71 (Sequence ID 91), Pil ANC4282 (Sequence ID 92), Pil 72-O-c (Sequence ID 93), Pil BI730 (Sequence ID 94), Pil A3K91 (Sequence ID 95), Pil CIP102159 (Sequence ID 96), Pil junii-65 (Sequence ID 97), Pil YZS-X (Sequence ID 98), Pil T-3-2 (Sequence ID 99), Pil CIP102637 The amino acid sequence selected from the group consisting of (SEQ ID NO: 100), and any fragment thereof (e.g., a C-terminal fragment), and / or variants in which the C-terminal threonine is substituted with serine, comprises or consists of the same. In a particular embodiment, the pyrin-like protein glycosylation fragment comprises the PilMo pyrin disulfide loop region (Pil Mo _DSL, Pil 20 Also known as (SEQ ID NO: 60), or its cleavage derivatives containing at least the last three amino acids from the C-terminus of pyrin, or variants in which the C-terminal threonine is substituted with serine (SEQ ID NO: 148), comprising or consisting thereof. Furthermore, in certain embodiments, the pyrin-like protein glycosylated fragment is Pil 20 (Sequence ID 60), Pil 19 (Sequence ID 133), Pil 18 (Sequence ID 134), Pil 17 (Sequence ID 135), Pil 16 (Sequence ID 136), Pil 15 (Sequence ID 109), Pil 14 (Sequence ID 137), Pil 13 (Sequence ID 110), Pil 12 (Sequence ID 138), Pil 11 (Sequence ID 139), Pil 10Pil (SEQ ID NO: 112), Pil9 (SEQ ID NO: 140), Pil8 (SEQ ID NO: 141), Pil7 (SEQ ID NO: 113), Pil6 (SEQ ID NO: 114), Pil5 (SEQ ID NO: 115), Pil4 (SEQ ID NO: 116), or Pil3 (SEQ ID NO: 117), or variants thereof having one, two, three, four, or five amino acid substitutions and maintaining the C-terminal threonine. Furthermore, in certain embodiments, the pyrin-like protein glycosylated fragment is Pil 20S (Sequence ID 148), Pil 19S (Sequence ID 149), Pil 18S (Sequence ID 150), Pil 17S (Sequence ID 151), Pil 16S (Sequence ID 152), Pil 15S (Sequence ID 153), Pil 14S (Sequence ID 154), Pil 13S (Sequence ID 155), Pil 12S (Sequence ID 156), Pil 11S (Sequence ID 157), Pil 10S (Sequence ID 158), Pil 9S (Sequence ID 159), Pil 8S (Sequence ID 160), Pil 7S (Sequence ID 161), Pil 6S (Sequence ID 162), Pil 5S (Sequence ID 163), Pil 4S (Sequence ID 164), or Pil 3S (SEQ ID NO: 165), or variants thereof having and maintaining one, two, three, four, or five amino acid substitutions in the C-terminal serine.

[0196] In certain embodiments of the fusion protein of this disclosure, the receptor protein may be glycosylated at two or more different locations. In certain embodiments, the fusion protein may be glycosylated by at least two different OTase classes within a single expression system. For example, in certain embodiments, the fusion protein further comprises a glycosylated sequence (e.g., a glycosylated fragment) of an OTase other than TfpM oligosaccharide transferase (OTase), in addition to a glycosylated fragment of a TfpM-associated pyrin-like protein located at the C-terminus. For example, the other OTase may be PglB, PglL, or PglS. In certain embodiments, the additional glycosylated sequence is an internal sequence of the fusion protein (i.e., not the majority of the C-terminus or N-terminus). In certain embodiments, the additional glycosylated sequence is an internal sequence within the sequence of the carrier protein (e.g., Figure 11A). In certain embodiments, the additional glycosylated sequence is also covalently bound to an oligosaccharide or polysaccharide. In certain embodiments, the fusion protein contains two or more, three or more, four or more, five or more, six or more, eight or more, ten or more, fifteen or more, or twenty or more additional glycosylation sequences. In certain embodiments, the fusion protein does not contain more than two, three or five, ten or fifteen or twenty or more additional glycosylation sequences. In certain embodiments, the additional glycosylation sequences are the same. In certain embodiments, at least one additional glycosylation sequence is different from the others. In certain embodiments, at least three, four or five additional glycosylation sequences are all different from each other. Also, in certain embodiments, none of the additional glycosylation sequences are the same. In certain embodiments, the fusion protein further contains an internal glycosylation fragment of ComP in addition to the glycosylation fragment of the TfpM-related pyrin-like protein located at its C-terminus. In certain embodiments, the ComP glycosylated fragment is also covalently bonded to an oligosaccharide or polysaccharide. Furthermore, in certain embodiments, the ComP glycosylated fragment is [Table 9] Or it comprises or consists of a fragment having at least one amino acid ASA at positions 11-13. In certain embodiments, the fusion protein contains two or more, three or more, four or more, five or more, six or more, eight or more, ten or more, fifteen or more, or twenty or more ComP glycosylated fragments. In certain embodiments, the fusion protein does not contain more than two, three or more, five or more, ten or more, fifteen or more, twenty or more ComP glycosylated fragments. In certain embodiments, the ComP glycosylated fragments are the same. In certain embodiments, the ComP glycosylated fragments are different from one another. In certain embodiments, at least three, at least four, or at least five ComP glycosylated fragments are all different from one another. Also, in certain embodiments, none of the ComP glycosylated fragments are the same.

[0197] In certain embodiments of the fusion proteins of this disclosure, the oligosaccharide or polysaccharide covalently bound to the pyrin-like protein or its glycosylated fragment has the size of at least three repeating units of the oligosaccharide or polysaccharide structure. In certain embodiments of the fusion proteins of this disclosure, the oligosaccharide or polysaccharide covalently bound to the pyrin-like protein or its glycosylated fragment has the size of at least ten monosaccharides.

[0198] In certain embodiments of the fusion protein of this disclosure, the oligosaccharide or polysaccharide is produced by a bacterium of the genus Streptococcus (e.g., S. pneumoniae or S. agalactiae), and the polysaccharide is a capsular polysaccharide such as Ia, Ib, II, III, IV, V, VI, VII, VIII, or IX.

[0199] In certain embodiments of the fusion proteins of this disclosure, the oligosaccharide or polysaccharide is produced by Klebsiella (e.g., K. pneumoniae), and the polysaccharide is a capsular polysaccharide or O antigen polysaccharide.

[0200] In certain embodiments of the fusion protein of this disclosure, the oligosaccharide or polysaccharide is produced by a bacterium of the genus Salmonella, and the polysaccharide is an O antigen polysaccharide. In certain embodiments, the bacterium is S. enterica, and the S. enterica polysaccharide is an O antigen of group B.

[0201] In certain embodiments of the fusion proteins of this disclosure, the glycosylated fusion protein is produced in vivo, such as in bacterial cells. In certain embodiments, the bacterium is Escherichia coli. In certain embodiments, the bacterium is from the genus Klebsiella. In certain embodiments, the bacterial species is K. pneumoniae, K. varricola, K. michinganenis, or K. oxytoca.

[0202] In certain embodiments of the fusion protein of this disclosure, the fusion protein is a vaccine that, when administered to a subject, induces an immune response. In certain embodiments, the immune response is an antibody response that induces long-term memory (memory B cells and T cells), and is optionally a serotype-specific antibody response. In certain embodiments, the antibody response is an IgG or IgM response. In certain embodiments, the antibody response is an IgG response, e.g., an IgG1 response. Also in certain embodiments, the fusion protein generates immunomemory in the subject to which the fusion protein has been administered.

[0203] Methods for generating complex carbohydrates This specification provides a method for producing complex carbohydrates. In certain embodiments, the method is carried out in vivo. In certain embodiments, the complex carbohydrates are produced in a cell-free system. An example of the use of a cell-free system utilizing an OTase other than TfpM is described in WO2013 / 067523A1, which is incorporated herein by reference. In certain embodiments, the method involves using the TfpM oligosaccharide transferase (OTase) of this disclosure to covalently conjugate an oligosaccharide or polysaccharide to a receptor protein or a glycosylated fragment thereof, which contains or consists of a TfpM-related pyrine-like protein. In certain embodiments, the pyrine-like protein or glycosylated fragment contains a C-terminal serine or threonine residue, the receptor protein contains a C-terminal serine or threonine residue, and the oligosaccharide or polysaccharide is covalently bound to the C-terminal serine or threonine residue of the receptor protein. In certain embodiments, the oligosaccharide or polysaccharide contains glucose at its reducing end. In certain embodiments, the receptor protein is a fusion protein of the Disclosure, as described in detail elsewhere in this Specification. Also, in certain embodiments, the complex carbohydrate is immunogenic.

[0204] In certain embodiments of the method for producing the complex carbohydrates of this disclosure, or any other composition or method disclosed herein, TfpM OTase contains a wzy_C superfamily domain, O antigen ligase domain, as defined by the conserved protein domain family cl04850 of the United States National Library of Science (NCBI), and / or TfpM OTase contains a wzy_C domain, O antigen ligase domain, as defined by the conserved protein domain family motif pfam04932 of the protein family (pfam) of the European Bioinformatics Laboratory (EBI, EMBL-EBI) of the European Molecular Biology Laboratory (EMBL), where pfam04932 is a protein domain family within the cl04850 superfamily protein domain. In certain embodiments, TfpM OTase contains TfpM Mo (Sequence ID 56), TfpMDSM16617 (Sequence ID 63), TfpM ZZC3 (Sequence ID 64), TfpM TUM15069 (Sequence ID 65), TfpM AI7 (Sequence ID 66), TfpM VE-C3 (Sequence ID 67), TfpM YH01026 (Sequence ID 68), TfpM CIP102143 (Sequence ID 69), TfpM AI40 (Sequence ID 70), TfpM F78 (Sequence ID 71), TfpM S71 (Sequence ID 72), TfpM ANC4282 (Sequence ID 73), TfpM CIP102159 (Sequence ID 74), TfpM junii-65 (Sequence ID 75), TfpM YZS-X (Sequence ID 76), TfpM CIP102637 (Sequence ID 77), TfpM T-3-2 (Sequence ID 78), TfpM BI730 (Sequence ID 79), TfpM A3K91 (Sequence ID 80), and / or TfpM 72-O-c (Sequence ID 81) contains at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity. In certain embodiments, TfpM OTase contains TfpM Mo (Sequence ID 56), TfpM DSM16617 (Sequence ID 63), TfpM ZZC3 (Sequence ID 64), TfpM TUM15069 (Sequence ID 65), TfpM AI7 (Sequence ID 66), TfpM VE-C3 (Sequence ID 67), TfpM YH01026 (Sequence ID 68), TfpM CIP102143 (Sequence ID 69), TfpM AI40 (Sequence ID 70), TfpM F78 (Sequence ID 71), TfpM S71 (Sequence ID 72), TfpM ANC4282 (Sequence ID 73), TfpM CIP102159 (Sequence ID 74), TfpM junii-65 (Sequence ID 75), TfpM YZS-X (Sequence ID 76), TfpM CIP102637 (Sequence ID 77), TfpMT-3-2 (Sequence ID 78), TfpM BI730 (Sequence ID 79), TfpM A3K91 (Sequence ID 80), or TfpM 72-O-c (Sequence ID 81). In a particular embodiment, TfpM OTase is TfpM Mo This is (sequence number 56).

[0205] In certain embodiments of the methods for generating the complex carbohydrates of the present disclosure, the receptor protein comprises or consists of a full-length TfpM-related pyrin-like protein. In certain embodiments, the receptor protein comprises or consists of a glycosylated fragment of a TfpM-related pyrin-like protein that is shorter than the full-length TfpM-related pyrin-like protein. In certain embodiments, the length of the glycosylated fragment of the pyrin-like protein is 3 to 138 amino acids, 10 to 138 amino acids, 20 to 138 amino acids, 50 to 138 amino acids, 100 to 138 amino acids, or 116 to 138 amino acids. In certain embodiments, the length of the glycosylated fragment of the pyrin-like protein is 3 to 139 amino acids, 10 to 139 amino acids, 20 to 139 amino acids, 50 to 139 amino acids, 100 to 139 amino acids, or 116 to 139 amino acids. In certain embodiments, the length of the glycosylated fragment is 3 to 140 amino acid lengths, 10 to 140 amino acid lengths, 20 to 140 amino acid lengths, 50 to 140 amino acid lengths, 100 to 140 amino acid lengths, or 116 to 140 amino acid lengths. In certain embodiments, the length of the glycosylated fragment is 3 to 22 amino acid lengths, 10 to 22 amino acid lengths, 11 to 22 amino acid lengths, 5 to 21 amino acid lengths, 10 to 21 amino acid lengths, or 11 to 21 amino acid lengths. In certain embodiments, the length of the glycosylated fragment has amino acid lengths ranging from any of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 to any of 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25.

[0206] In certain embodiments of the method for generating complex carbohydrates in this disclosure, the pyrin-like protein glycosylated fragment is Pil Mo (Sequence ID 57), or Pil lacking amino acids corresponding to residues 1-28. Mo (Pil Mo Polypeptides that are or contain at least Δ28 (SEQ ID NO: 58) have at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity to SEQ ID NO: 57 or SEQ ID NO: 58, for example, the C-terminal threonine is substituted with serine. In certain embodiments, the TfpM-related pyrine-like protein is Pil DSM16617 (Sequence ID 82), Pil ZZC3-9 (Sequence ID 83), Pil TUM15069 (Sequence ID 84), Pil AI7 (Sequence ID 85), Pil VE-C3 (Sequence number 86), Pil YH01026 (Sequence ID 87), Pil CIP102143 (Sequence number 88), Pil AI40 (Sequence ID 89), Pil F78 (Sequence ID 90), Pil S71 (Sequence ID 91), Pil ANC4282 (Sequence ID 92), Pil 72-O-c (Sequence ID 93), Pil BI730 (Sequence ID 94), Pil A3K91 (Sequence ID 95), Pil CIP102159 (Sequence ID 96), Pil junii-65 (Sequence ID 97), Pil YZS-X (Sequence ID 98), Pil T-3-2 (Sequence ID 99), and Pil CIP102637 Selected from the group consisting of (Sequence ID 100). In a particular embodiment, the TfpM-related pyrin-like protein or pyrin-like protein glycosylated fragment is Pil DSM16617 (Sequence ID 82), Pil ZZC3-9 (Sequence ID 83), Pil TUM15069 (Sequence ID 84), Pil AI7 (Sequence ID 85), Pil VE-C3 (Sequence number 86), Pil YH01026 (Sequence ID 87), Pil CIP102143(Sequence number 88), Pil AI40 (Sequence ID 89), Pil F78 (Sequence ID 90), Pil S71 (Sequence ID 91), Pil ANC4282 (Sequence ID 92), Pil 72-O-c (Sequence ID 93), Pil BI730 (Sequence ID 94), Pil A3K91 (Sequence ID 95), Pil CIP102159 (Sequence ID 96), Pil junii-65 (Sequence ID 97), Pil YZS-X (Sequence ID 98), Pil T-3-2 (Sequence ID 99), Pil CIP102637 The amino acid sequence selected from the group consisting of (SEQ ID NO: 100), and any fragment thereof (e.g., a C-terminal fragment), and / or variants in which the C-terminal threonine is substituted with serine, comprises or consists of the same. In a particular embodiment, the pyrin-like protein glycosylation fragment comprises the PilMo pyrin disulfide loop region (Pil Mo _DSL, Pil 20 Also known as (SEQ ID NO: 60), or its cleavage derivatives containing at least the last three amino acids from the C-terminus of pyrin, or variants in which the C-terminal threonine is substituted with serine (SEQ ID NO: 148), comprising or consisting thereof. Furthermore, in certain embodiments, the pyrin-like protein glycosylated fragment is Pil 20 (Sequence ID 60), Pil 19 (Sequence ID 133), Pil 18 (Sequence ID 134), Pil 17 (Sequence ID 135), Pil 16 (Sequence ID 136), Pil 15 (Sequence ID 109), Pil 14 (Sequence ID 137), Pil 13 (Sequence ID 110), Pil 12 (Sequence ID 138), Pil 11 (Sequence ID 139), Pil 10Pil (SEQ ID NO: 112), Pil9 (SEQ ID NO: 140), Pil8 (SEQ ID NO: 141), Pil7 (SEQ ID NO: 113), Pil6 (SEQ ID NO: 114), Pil5 (SEQ ID NO: 115), Pil4 (SEQ ID NO: 116), or Pil3 (SEQ ID NO: 117), or variants thereof having one, two, three, four, or five amino acid substitutions and maintaining the C-terminal threonine. Furthermore, in certain embodiments, the pyrin-like protein glycosylated fragment is Pil 20S (Sequence ID 148), Pil 19S (Sequence ID 149), Pil 18S (Sequence ID 150), Pil 17S (Sequence ID 151), Pil 16S (Sequence ID 152), Pil 15S (Sequence ID 153), Pil 14S (Sequence ID 154), Pil 13S (Sequence ID 155), Pil 12S (Sequence ID 156), Pil 11S (Sequence ID 157), Pil 10S (Sequence ID 158), Pil 9S (Sequence ID 159), Pil 8S (Sequence ID 160), Pil 7S (Sequence ID 161), Pil 6S (Sequence ID 162), Pil 5S (Sequence ID 163), Pil 4S (Sequence ID 164), or Pil 3S (SEQ ID NO: 165), or variants thereof having and maintaining one, two, three, four, or five amino acid substitutions in the C-terminal serine.

[0207] In certain embodiments of the method for generating the complex carbohydrate of the present disclosure, the receptor protein is a fusion protein, and the carrier protein is selected from the group consisting of Pseudomonas aeruginosa exotoxin A (EPA), CRM197, cholera toxin B subunit, tetanus toxin C fragment, and any fragment thereof. In certain embodiments, a TfpM-associated pyrin-like protein or a glycosylated fragment thereof is translationally fused / bound to a heterologous carrier protein via an amino acid linker.

[0208] In certain embodiments of the methods for generating complex carbohydrates of this disclosure, the receptor protein is a fusion protein, and the method comprises glycosylation of the receptor protein at two or more different sites. In certain embodiments, the method comprises glycosylation of the receptor protein using at least two different OTase classes within a single expression system. In certain embodiments, the fusion protein comprises at least two different OTases and two or more associated glycosylation sequences (e.g., glycosylation fragments). Representative examples of OTases that can be used in combination include PglB, PglL, PglS, TfpO, and TfpM. Those skilled in the art will recognize that OTases cannot be used together if both OTases require glycosylation sequences (sequences) at the same site, for example, both at the N-terminus or both at the C-terminus. For example, if both TfpO and TfpM typically require a sequence at the C-terminus, they cannot be used together. For example, in certain non-limiting exemplary embodiments, the receptor protein includes a glycosylated fragment of a TfpM-associated pyrin-like protein located at its C-terminus, in addition to an additional glycosylated sequence of an OTase other than TfpM oligosaccharide transferase (OTase). In certain embodiments, the other OTases are PglB, PglL, and / or PglS. In certain embodiments, one or more glycosylated sequences are internal sequences of the fusion protein (i.e., not the majority of the C-terminus or N-terminus sequence). In certain embodiments, one or more glycosylated sequences are internal sequences within the sequence of the carrier protein (e.g., Figure 33A). In certain embodiments, additional glycosylated sequences are internal sequences of the fusion protein (i.e., not the majority of the C-terminus or N-terminus sequence). In certain embodiments, additional glycosylated sequences are internal sequences within the sequence of the carrier protein (e.g., Figure 33A). In certain embodiments, at least two different glycosylated sequences of two different OTase systems are covalently bonded to an oligosaccharide or polysaccharide.In certain embodiments, the glycosylated fragment and additional glycosylated sequences of the TfpM-related pyrin-like protein located at the C-terminus of the fusion protein are covalently bonded to an oligosaccharide or polysaccharide. In certain embodiments, the fusion protein contains two or more, three or more, four or more, five or more, six or more, eight or more, ten or more, fifteen or more, or twenty or more additional glycosylated sequences. In certain embodiments, the fusion protein does not contain more than two, three or more, five or more, ten or more, fifteen or more, twenty or more additional glycosylated sequences. In certain embodiments, the additional glycosylated sequences are the same. In certain embodiments, at least one additional glycosylated sequence is different from the others. In certain embodiments, at least three, four or five additional glycosylated sequences are all different from each other. In certain embodiments, none of the additional glycosylated sequences are the same. For example, in a particular embodiment, this method includes glycosylation of the glycosylated fragment of the C-terminal TfpM-related pyrin-like protein, in addition to further glycosylation of the internal glycosylated fragment of ComP using PglS OTase. In a particular embodiment, the ComP glycosylated fragment is... [Table 10] Or it comprises or consists of a fragment having at least one amino acid ASA at positions 11-13. In certain embodiments, the fusion protein contains two or more, three or more, four or more, five or more, six or more, eight or more, ten or more, fifteen or more, or twenty or more ComP glycosylated fragments. In certain embodiments, the fusion protein does not contain more than two, three or more, five or more, ten or more, fifteen or more, twenty or more ComP glycosylated fragments. In certain embodiments, the ComP glycosylated fragments are the same. In certain embodiments, the ComP glycosylated fragments are different from one another. In certain embodiments, at least three, at least four, or at least five ComP glycosylated fragments are all different from one another. Also, in certain embodiments, none of the ComP glycosylated fragments are the same.

[0209] In certain embodiments of the method for producing the complex carbohydrates of this disclosure, (e.g., binding) occurs in vivo within a host cell. In certain embodiments, the host cell is a bacterial cell. In certain embodiments, binding occurs within Escherichia coli. In certain embodiments, binding occurs within a bacterium from the genus Klebsiella. In certain embodiments, the bacterial species is K. pneumoniae, K. varricola, K. michinganenis, or K. oxytoca.

[0210] In certain embodiments of the method for producing complex carbohydrates of the present disclosure, the method comprises culturing a host cell comprising (a) a gene cluster encoding a protein necessary for synthesizing an oligosaccharide or polysaccharide, (b) TfpM OTase, and (3) a receptor protein.

[0211] In certain embodiments of the method for generating complex carbohydrates of the present disclosure, the method generates a conjugate vaccine.

[0212] Additional Embodiments Provided herein are host cells comprising (a) a gene cluster encoding a protein necessary for the synthesis of oligosaccharides or polysaccharides, (b) the TfpM OTase of the present disclosure, and (3) a receptor protein comprising the TfpM-associated pyrin-like protein of the present disclosure or a glycosylated fragment thereof. In certain embodiments, the receptor protein is a fusion protein. In certain embodiments, the host cell comprises a nucleic acid encoding the TfpM OTase. In certain embodiments, the host cell comprises a nucleic acid encoding the receptor protein. Also, in certain embodiments, the TfpM OTase and the receptor protein are encoded by the same nucleic acid.

[0213] This specification provides isolated nucleic acids encoding glycosylated fragments and / or fusion proteins of pyrin-like proteins of the present disclosure. In certain embodiments, the nucleic acid is a vector. Also provided are host cells containing such isolated nucleic acids of the present disclosure. In certain embodiments, the host cell is a bacterial cell. In certain embodiments, the host cell is Escherichia coli. In certain embodiments, the host cell is from the genus Klebsiella. Also, in certain embodiments, the host cell is K. pneumoniae, K. varricola, K. michinganenis, or K. oxytoca.

[0214] Compositions comprising a conjugate vaccine or fusion protein of the present disclosure and an adjuvant and / or carrier are provided herein. In certain embodiments, the composition is a pharmaceutical or therapeutic composition suitable for administration to a subject / patient.

[0215] Methods for inducing a host immune response to a bacterial pathogen are provided herein, the methods comprising administering an effective amount of the conjugate vaccine, fusion protein, or composition comprising the conjugate vaccine or fusion protein and an adjuvant and / or carrier to a subject requiring an immune response. Treatment with a pharmaceutical composition comprising an immunogenic composition may be carried out separately or in combination with other treatments, as necessary. An amount sufficient to achieve this is defined as an “effective amount,” “effective dose,” or “unit dose.” An effective amount for this application will vary, for example, depending on the composition of the complex carbohydrate, the method of administration, the stage and severity of the disease being treated, the patient’s weight and overall health, and the judgment of the prescribing physician. In some embodiments, boost doses are administered over a period of time after the initial antigen administration. In certain embodiments, the immune response is an antibody response. In certain embodiments, the immune response is selected from the group consisting of innate responses, adaptive responses, humoral responses, antibody responses, cellular responses, B-cell responses, T-cell responses, cytokine upregulation or downregulation, immune system crosstalk, and two or more combinations of such immune responses. In a particular embodiment, the immune response is selected from the group consisting of innate responses, humoral responses, antibody responses, T cell responses, and two or more combinations of such immune responses.

[0216] Methods for preventing or treating bacterial diseases and / or infections in subjects are provided herein, the methods comprising administering to a subject in need of such prevention or treatment in an effective amount of a conjugate vaccine, fusion protein, or composition comprising a conjugate vaccine or fusion protein and an adjuvant and / or carrier. In certain embodiments, the subject is a mammal. In certain embodiments, the subject is a human. In certain embodiments, the subject is a pet animal. In certain embodiments, the subject is livestock. In certain embodiments, the infection is a local or systemic infection of the skin, soft tissue, blood, or organs, or is autoimmune. In certain embodiments, the disease is pneumonia. In certain embodiments, the infection is a systemic infection and / or blood infection. In certain embodiments, the conjugate vaccine, fusion protein, or composition is administered by intramuscular injection, intradermal injection, intraperitoneal injection, subcutaneous injection, intravenous injection, oral administration, mucosal administration, intranasal administration, or pulmonary administration.

[0217] A method for producing a pneumococcal conjugate vaccine against pneumococcal infection is provided herein, comprising (a) isolating a complex carbohydrate or glycosylated fusion protein of the present disclosure, and (b) combining the isolated complex carbohydrate or isolated glycosylated fusion protein with an adjuvant and / or carrier.

[0218] The following disclosure provides for use in inducing a host immune response to bacterial pathogens and / or in preventing or treating bacterial diseases and / or infections in a subject, including complex carbohydrates, glycosylated fusion proteins, or conjugate vaccines, or any composition thereof.

[0219] Recombinant nucleic acid constructs comprising a nucleotide sequence encoding a TfpM oligosaccharide transferase (OTase) operably bound to at least one heterologous transcription regulatory sequence are provided herein. In certain embodiments, the TfpM OTase is TfpM Mo(Sequence ID 56), TfpM DSM16617 (Sequence ID 63), TfpM ZZC3 (Sequence ID 64), TfpM TUM15069 (Sequence ID 65), TfpM AI7 (Sequence ID 66), TfpM VE-C3 (Sequence ID 67), TfpM YH01026 (Sequence ID 68), TfpM CIP102143 (Sequence ID 69), TfpM AI40 (Sequence ID 70), TfpM F78 (Sequence ID 71), TfpM S71 (Sequence ID 72), TfpM ANC4282 (Sequence ID 73), TfpM CIP102159 (Sequence ID 74), TfpM junii-65 (Sequence ID 75), TfpM YZS-X (Sequence ID 76), TfpM CIP102637 (Sequence ID 77), TfpM T-3-2 (Sequence ID 78), TfpM BI730 (Sequence ID 79), TfpM A3K91 (Sequence ID 80), and / or TfpM 72-O-c (Sequence ID 81) contains at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity. In certain embodiments, TfpM OTase contains TfpM Mo (Sequence ID 56), TfpM DSM16617 (Sequence ID 63), TfpM ZZC3 (Sequence ID 64), TfpM TUM15069 (Sequence ID 65), TfpM AI7 (Sequence ID 66), TfpM VE-C3 (Sequence ID 67), TfpM YH01026 (Sequence ID 68), TfpM CIP102143 (Sequence ID 69), TfpM AI40 (Sequence ID 70), TfpM F78 (Sequence ID 71), TfpM S71 (Sequence ID 72), TfpM ANC4282 (Sequence ID 73), TfpM CIP102159 (Sequence ID 74), TfpM junii-65 (Sequence ID 75), TfpM YZS-X (Sequence ID 76), TfpM CIP102637(Sequence ID 77), TfpM T-3-2 (Sequence ID 78), TfpM BI730 (Sequence ID 79), TfpM A3K91 (Sequence ID 80), and / or TfpM 72-O-c (Sequence ID 81). In a particular embodiment, TfpM OTase is TfpM Mo(Sequence ID 56). In certain embodiments, the heterologous transcription regulatory sequence is a promoter sequence. In certain embodiments, the recombinant nucleic acid construct further comprises a fusion protein of the Disclosure comprising a nucleotide sequence encoding a TfpM-related pyrin-like protein or its glycosylated fragment, or a nucleotide sequence encoding a TfpM-related pyrin-like protein or its glycosylated fragment operably bound to a nucleotide sequence encoding a TfpM OTase. In certain embodiments, the recombinant nucleic acid construct further comprises a fusion protein of the Disclosure comprising a nucleotide sequence encoding a TfpM-related pyrin-like protein or its glycosylated fragment, or a nucleotide sequence encoding a TfpM OTase, located at the 5' end of the nucleotide sequence encoding a TfpM-related pyrin-like protein or its glycosylated fragment, and operably bound to the nucleotide sequence. In certain embodiments, the fusion protein of this construct also comprises glycosylated sequences of OTases other than TfpM, e.g., PglB, PglL, PglS (e.g., ComP or its glycosylated fragment). In certain embodiments, the coding sequence of a TfpM-related pyrin-like protein or its glycosylated fragment, or a fusion protein containing a TfpM-related pyrin-like protein or its glycosylated fragment, is within 2, 5, 10, 20, 30, 40, or 50 nucleotides of the sequence encoding TfpM OTase. In certain embodiments, the coding sequence of a TfpM-related pyrin-like protein or its glycosylated fragment, or a fusion protein containing a TfpM-related pyrin-like protein or its glycosylated fragment, overlaps with an operablely bound nucleotide sequence encoding TfpM OTase. In certain embodiments, the TfpM-related pyrin-like protein comprises or consists of a full-length TfpM-related pyrin-like protein. In certain embodiments, the TfpM-related pyrin-like protein comprises or consists of a glycosylated fragment of a TfpM-related pyrin-like protein that is shorter than the full-length TfpM-related pyrin-like protein.In certain embodiments, the length of the glycosylated fragment of the pyrin-like protein is 3 to 138 amino acid lengths, 10 to 138 amino acid lengths, 20 to 138 amino acid lengths, 50 to 138 amino acid lengths, 100 to 138 amino acid lengths, or 116 to 138 amino acid lengths. In certain embodiments, the length of the glycosylated fragment of the pyrin-like protein is 3 to 139 amino acid lengths, 10 to 139 amino acid lengths, 20 to 139 amino acid lengths, 50 to 139 amino acid lengths, 100 to 139 amino acid lengths, or 116 to 139 amino acid lengths. In certain embodiments, the length of the glycosylated fragment is 3 to 140 amino acid lengths, 10 to 140 amino acid lengths, 20 to 140 amino acid lengths, 50 to 140 amino acid lengths, 100 to 140 amino acid lengths, or 116 to 140 amino acid lengths. In certain embodiments, the length of the glycosylated fragment is 3 to 22 amino acid lengths, 10 to 22 amino acid lengths, 11 to 22 amino acid lengths, 5 to 21 amino acid lengths, 10 to 21 amino acid lengths, or 11 to 21 amino acid lengths. In certain embodiments, the length of the glycosylated fragment is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 amino acid lengths, up to 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acid lengths. In certain embodiments, the length of the glycosylated fragment has an amino acid length of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25. In certain embodiments, the pyrin-like protein glycosylated fragment is Pil. Mo (Sequence ID 57), or Pil lacking amino acids corresponding to residues 1-28. Mo (Pil MoPolypeptides that are or contain at least Δ28 (SEQ ID NO: 58) have at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity to SEQ ID NO: 57 or SEQ ID NO: 58, for example, the C-terminal threonine is substituted with serine. In certain embodiments, the TfpM-related pyrine-like protein is Pil DSM16617 (Sequence ID 82), Pil ZZC3-9 (Sequence No. 83) 、 Pil TUM15069 (Sequence ID 84), Pil AI7 (Sequence ID 85), Pil VE-C3 (Sequence number 86), Pil YH01026 (Sequence ID 87), Pil CIP102143 (Sequence number 88), Pil AI40 (Sequence ID 89), Pil F78 (Sequence ID 90), Pil S71 (Sequence ID 91), Pil ANC4282 (Sequence ID 92), Pil 72-O-c (Sequence ID 93), Pil BI730 (Sequence ID 94), Pil A3K91 (Sequence ID 95), Pil CIP102159 (Sequence ID 96), Pil junii-65 (Sequence ID 97), Pil YZS-X (Sequence ID 98), Pil T-3-2 (Sequence ID 99), and Pil CIP102637 Selected from the group consisting of (Sequence ID 100). In a particular embodiment, the TfpM-related pyrin-like protein or pyrin-like protein glycosylated fragment is Pil DSM16617 (Sequence ID 82), Pil ZZC3-9 (Sequence ID 83), Pil TUM15069 (Sequence ID 84), Pil AI7 (Sequence ID 85), Pil VE-C3 (Sequence number 86), Pil YH01026 (Sequence ID 87), Pil CIP102143 (Sequence number 88), Pil AI40 (Sequence ID 89), Pil F78 (Sequence ID 90), Pil S71 (Sequence ID 91), Pil ANC4282 (Sequence ID 92), Pil72-O-c (Sequence ID 93), Pil BI730 (Sequence ID 94), Pil A3K91 (Sequence ID 95), Pil CIP102159 (Sequence ID 96), Pil junii-65 (Sequence ID 97), Pil YZS-X (Sequence ID 98), Pil T-3-2 (Sequence ID 99), Pil CIP102637 The present invention comprises or comprises amino acid sequences selected from the group consisting of (SEQ ID NO: 100), and any fragment thereof (e.g., a C-terminal fragment), and / or variants in which the C-terminal threonine is substituted with serine. In a particular embodiment, the pyrin-like protein glycosylation fragment comprises a PilMo pyrin disulfide loop region (Pil Mo _DSL, Pil 20 Also known as (SEQ ID NO: 60), or its cleavage derivatives containing at least the last three amino acids from the C-terminus of pyrin, or variants in which the C-terminal threonine is substituted with serine (SEQ ID NO: 148), comprising or consisting thereof. Furthermore, in certain embodiments, the pyrin-like protein glycosylated fragment is Pil 20 (Sequence ID 60), Pil 19 (Sequence ID 133), Pil 18 (Sequence ID 134), Pil 17 (Sequence ID 135), Pil 16 (Sequence ID 136), Pil 15 (Sequence ID 109), Pil 14 (Sequence ID 137), Pil 13 (Sequence ID 110), Pil 12 (Sequence ID 138), Pil 11 (Sequence ID 139), Pil 10 (SEQ ID NO: 112), Pil9 (SEQ ID NO: 140), Pil8 (SEQ ID NO: 141), Pil7 (SEQ ID NO: 113), Pil6 (SEQ ID NO: 114), Pil5 (SEQ ID NO: 115), Pil4 (SEQ ID NO: 116), or Pil3 (SEQ ID NO: 117), These consist of variants having one, two, three, four, or five amino acid substitutions and maintaining the C-terminal threonine. Furthermore, in certain embodiments, the pyrin-like protein glycosylated fragment is Pil 20S (Sequence ID 148), Pil 19S (Sequence ID 149), Pil 18S (Sequence ID 150), Pil 17S (Sequence ID 151), Pil 16S (Sequence ID 152), Pil 15S (Sequence ID 153), Pil 14S (Sequence ID 154), Pil 13S (Sequence ID 155), Pil 12S (Sequence ID 156), Pil 11S (Sequence ID 157), Pil 10S (Sequence ID 158), Pil 9S (Sequence ID 159), Pil 8S (Sequence ID 160), Pil 7S (Sequence ID 161), Pil 6S (Sequence ID 162), Pil 5S (Sequence ID 163), Pil 4S (Sequence ID 164), or Pil 3SThe sequence consists of (SEQ ID NO: 165), or variants thereof having one, two, three, four, or five amino acid substitutions and maintaining the C-terminal serine. In certain embodiments, the fusion protein is the fusion protein of this disclosure. In certain embodiments, the recombinant construct further comprises a nucleotide sequence encoding an additional OTase operably bound to TpfM OTase, as described elsewhere in this specification. In certain embodiments, the recombinant construct further comprises a nucleotide sequence encoding an additional OTase 3' of TpfM OTase and operably bound to TpfM OTase. In certain embodiments, the recombinant construct further comprises a nucleotide sequence encoding an additional OTase 5' of TpfM OTase and operably bound to TpfM OTase. In certain embodiments, the coding sequence for the additional OTase is within 10, 20, 30, 40, 50, 75, or 100 nucleotides of the sequence encoding TffM OTase. In certain embodiments, the recombinant construct further comprises a nucleotide sequence encoding PglS OTase 3' operably bound to TpfM OTase. In certain embodiments, the recombinant construct further comprises a nucleotide sequence encoding PglS OTase 3' of TpfM OTase and operably bound to TpfM OTase. In certain embodiments, the recombinant construct further comprises a nucleotide sequence encoding PglS OTase 5' of TpfM OTase and operably bound to TpfM OTase. In certain embodiments, the coding sequence for PglS OTase is within 10, 20, 30, 40, 50, 75, or 100 nucleotides of the sequence encoding TffM OTase. Vectors containing recombinant nucleic acid constructs are further provided herein. Host cells containing recombinant nucleic acid constructs or vectors are also further provided herein. In certain embodiments, the host cell is a bacterial cell. In certain embodiments, the host cell is Escherichia coli. In certain embodiments, the host cells are from the genus Klebsiella.In certain embodiments, the host cells are K. pneumoniae, K. varricola, K. michinganenis, or K. oxytoca. A method for generating TfpM OTase is provided herein, comprising culturing the host cells, wherein the vector is an expression vector, and the TfpM OTase is recovered. [Examples]

[0220] Crosslinking of bioconjugates to NP / VLP monomers provides repeated representation of desired glycan and protein epitopes on a scale larger than that of a single protein molecule (Liu, Y., et al. (2023) Microb Cell Fact 22, 95). The following exemplary examples describe the design and demonstration of bioconjugate-mi3 and bioconjugate-AP205 assemblies. These bioconjugates were generated using two different bacterial O-linked OTases: Acinetobacter baylyi PglS and Moraxella osloensis TfpM. PglS and TfpM belong to a recently characterized family of O-linked OTases with the broadest sugar substrate range of known OTases (Harding, CM, et al. (2019) Nat Commun 10, 891; Knoot, CJ, et al. (2023) Glycobiology, Volume 33, Pages 57-74). In particular, both PglS and TfpM can transport glycans with glucose at their reducing end, and these enzymes can be used to generate bioconjugate vaccines against various pathogens in which natural polysaccharides have this sugar at their reducing end (Harding, CM, et al. (2019) Nat Commun 10, 891, Feldman, et al. (2019) PNAS, 116(37) 18655-18663). For this application, E. coli maltose-binding protein (MBP) and P. aeruginosa EPA were engineered to contain PglS or TfpM-specific sequencers (Knoot, CJ, et al. (2021) Glycobiology, Volume 31, Pages 1192-1203; Knoot, CJ, et al. (2023) Glycobiology, Volume 33, Pages 57-74), and subsequently glycosylated with the O16O antigen from E. coli. The resulting PglS or TfpM-derived MBP or EPA bioconjugates were covalently bound to NP monomers using Spytag / Spycatcher technology.

[0221] Example 1. SpyTag-tagged MBP-O16 bioconjugate. SpyTag-treated MBP-O16 bioconjugate ("first polypeptide," e.g., Figure 1: "Protein 1") was generated in CLM24 glycosylated E. coli strains. The MBP-SpyTag fusion protein was expressed separately from the pEXT20 expression plasmid, and WbbL was expressed from the pEXT21 derivative plasmid pMF19. WbbL expression restores O16 O-antigen production in most laboratory strains of E. coli. mi3-SpyCatcher and AP205-SpyCatcher fusion proteins ("second polypeptide," e.g., Figure 1: "Protein 2") were generated separately in C41(DE3)E. coli. After culturing E. coli expression strains in TB medium, cell pellets were frozen for downstream lysis and protein purification.

[0222] SpyTag-tagged O16 bioconjugates were purified from periplasmic cell extracts using immobilized metal affinity (Ni) chromatography (IMAC). The IMAC eluate was concentrated, buffer-exchanged, and loaded onto an Akta FPLC instrument for further purification using anion exchange chromatography. Anion exchange was used to separate SpyTag-tagged non-glycosylated MBP from SpyTag-tagged glycosylated MBP-O16. Fractions containing SpyTag-tagged glycosylated MBP-O16 were pooled, concentrated, and quantified using a BCA assay kit. The purified conjugates were stored in Tris-buffered saline (TBS) at -80°C until use in isopeptide bond formation reactions.

[0223] Mi3 and AP205 nanoparticles (NPs) were purified by whole cell lysis via sonication. The lysates were centrifuged at 18,000 xG and then loaded onto IMAC resin as described above. The eluates were concentrated and loaded onto an FLPC size exclusion column to separate fully aggregated VLPs and NPs from unaggregated free monomers. Fractions containing intact aggregated VLPs or NPs were determined based on the mass of known standards run on the same column. These fractions were pooled, concentrated, quantified, and stored in TBS at 4°C until used in isopeptide bond formation reactions.

[0224] For the isopeptide formation reaction, two purified proteins (SpyTag-tagged O16-bioconjugate and VLP / NP) were mixed in TBS buffer at a molar ratio of approximately 1:1 or 2:1 mi3 / AP205:SpyTag-tagged O16-bioconjugate. These reactants were incubated at 22°C for 1.5 hours (mi3) or 3 hours (AP205). The reactants were then analyzed using SDS-PAGE Coomassie staining, Western blotting, and size exclusion chromatography using a Sephacryl S-400 HR column.

[0225] Example 2. Spytagged EPA-O16 bioconjugate. Spytagged EPA-O16 bioconjugates were generated and purified in the same manner as the MBP-O16 bioconjugates disclosed elsewhere in this specification. All versions of the Spytagged EPA were generated from the pEXT20 expression plasmid.

[0226] Crosslinking between Spytagged O16 bioconjugates and purified mi3-Spycatcher or AP205-Spycatcher generated higher molecular weight covalent crosslinked proteins, as determined using Coomassie protein staining, Western blotting, and size exclusion chromatography. Using E. coli O16 antiserum and anti-protein antibodies, it was shown that the higher molecular weight species consisted of O16 glycan-binding proteins (plural) (e.g., Figures 3 and 6). Generally, based on the intensity of crosslinked protein bands in Coomassie stained gels or Western blotting, and the observation that fewer NP / VLP monomer bands or bioconjugate bands remained after the isopeptide bonding reaction, Spytagged bioconjugates derived from TfpM reacted more completely with Spycatcher-mi3 or AP205 to form isopeptide bonded assemblies.

[0227] Size exclusion chromatography showed that the mass of the bioconjugate-NP / VLP assembly was greater than the mass before the isopeptide reaction (Figure 3).

[0228] Not all spytagged EPA bioconjugates were able to form cross-linked isopeptide species in detectable amounts using mi3-Spycatcher or AP205-Spycatcher. Of the three spytagged EPA variants tested, only EPA-Spytag-v1 generated isopeptide species (Figure 9). Furthermore, only a subset of the protein linker variant of EPA-Spytag-v1 was able to form isopeptide species using mi3-Spycatcher (Figure 10).

[0229] Based on the relative Western blot protein band intensities, the isopeptide bond formation time course indicated that the reaction between EPA-Spytag-v1 and mi3-Spycatcher was substantially over 80% complete within 1 hour of the start of the reaction (Figure 11).

[0230] Example 3. Immunization with glycosylated ComP bioconjugates induces an immune response. The T cell-dependent immune response to conjugate vaccines is characterized by the secretion of high-affinity IgG1 antibodies (Avci, FY, Li, X., Tsuji, M. & Kasper, DLNat Med 17, 1602-1609 (2011)). The immunogenicity of the CPS14-ComP bioconjugate was evaluated in a mouse vaccination model (the entire report is incorporated herein by reference, WO / 2020 / 131236). Serum recovered from mice vaccinated with the CPS14-ComP bioconjugate showed a significant increase in CPS14-specific IgG titers, but not a significant increase in IgM titers. Furthermore, secondary HRP-tagged anti-IgG subtype antibodies were used to determine which IgG subtypes showed elevated titers. IgG1 titers appeared to be higher than those of other subtypes.

[0231] Next, a second vaccination trial was conducted, comparing the immunogenicity of trivalent CPS8-, CPS9V-, and CPS14-ComP bioconjugates with that of PREVNAR 13®, the current standard of care. Serotypes 9V and 14 are included in PREVNAR 13®, and elevated IgG titers can be observed in mice immunized with PREVNAR 13® against these two serotypes. Monovalent immunization against serotype 14 also showed a significant induction of serotype-specific IgG titers, which was similar to that of preliminary immunization. All mice that received the trivalent bioconjugates showed elevated serotype-specific IgG titers compared to controls, as expected, and serum at day 49 showed significantly higher IgG titers for serotypes 8 and 14 compared to serotype 9V. Nevertheless, IgG titers against 9V were still significantly higher than those of placebo.

[0232] Example 4. Cloning and plasmid assembly All primers and oligonucleotides used in this study are listed in Table 2. The commonly used antibiotic concentrations for liquid culture and LB agar plates were ampicillin (Amp), 100 μg / mL, kanamycin (Kan), 20 μg / mL, tetracycline (Tet), 10 μg / mL, and spectinomycin (Sp), 50 μg / mL. To clone the tfpM pyrin OTase gene, HiFi gblock (Integrated DNA Technologies, IDT), designed with 25 base pair duplication at the terminals for Gibson assembly using a PCR linearized plasmid, was ordered. The plasmid backbone of these fragments was amplified from the pEXT20 plasmid encoding the P. aeruginosa EPA gene under the control of the tac promoter (pVNM57) (Dykxhoorn, DM, et al. (1996) Gene 177, 133-136) (Knoot, CJ, et al. (2021) Glycobiology 31, 1192-1203). The EPA gene had a deletion at residue E553, resulting in inactivation of the toxin. The linearized plasmids were mixed separately with each of the synthesized tfpM gBlocks and assembled using the NEBuilder HiFi DNA Assembly Kit (New England Biolabs, NEB). After assembly, the plasmids were transformed into E. coli Stellar cells (Takara Bio) by heat shock, grown at 37°C for 1 hour, and plated on LB agar supplemented with amp. Individual colonies were collected and grown in LB medium with appropriate antibiotics, and plasmids were isolated using the GeneJet Plasmid Miniprep Kit (Thermo Fisher). All plasmids were sequenced by Sanger sequencing (Genewiz). M.osloensis 1202 EPA-PilΔ28 fusion and TfpM Mo The plasmid expressing pVNM227 was named pVNM227. MoTo generate site-specific mutants, the inventors designed duplicate PCR primers that introduced the necessary codon changes in the pyrin gene and amplified each fragment from the pVNM227 plasmid. The resulting PCR product was digested with DpnI (NEB) at 37°C for 30 minutes and gel-purified from agarose gel using the Pure-Link Gel Extraction Kit (Thermo Fisher). To insert the truncated pyrin gene region, a complementary oligonucleotide with a 25 bp terminal duplication homologous to the pVNM227 PCR product was ordered. The oligonucleotides were resuspended in purified water, mixed, and annealed together in a thermocycler by heating to 98°C for 5 minutes, followed by slow cooling to 4°C at 0.1°C / min. The annealed oligonucleotides were diluted 1:5 in water and assembled with PCR-linearized pVNM227 using the NEBuilder HiFi DNA Assembly Kit (NEB). The resulting DNA was transformed into Stellar cells, and the plasmid was isolated and validated as described above. EPA-Pil 20 The plasmid containing the construct encoding TfpM was named pVNM297. N-terminal His tagged EPA-Pil 20 The variant was constructed by linearizing pVNM297 using PCR and using this fragment in a Gibson assembly with a complementary annealed oligo containing a 6xHis coding region and terminal homologous region, yielding pVNM291. pVNM167 is the previously described EPA iGTccThe plasmid (Knoot, CJ, et al. (2021) Glycobiology 31, 1192-1203) was generated by digesting with SalI. The purified SalI fragment was Gibson assembled using the pglS gene, which has a native 100 bp 5' UTR, amplified from A. baylyi ADP1 gDNA. pVNM245 was generated from the pVNM167 template by a separate PCR reaction to amplify a product with an overhang for Gibson assembly: (i) an EPA with a vector backbone containing PglS and one iGT, (ii) a second iGT for incorporation between E548 and G549, and (iii) the C-terminus of the EPA downstream of the iGT. Plasmid pVNM337 was created by amplifying tfpM from pVNM291 using primers EPA 3' F1 and pglS-tfpM R1, and then cloning the product to PCR-linearized pVNM167 amplified with pglS 5' F1 and EPA 3' R1. Phylogenetic trees of TfpM and pyrin proteins were generated using the phylogeny.fr server (worldwide web phylogeny.fr / ) with MUSCLE, PhyML, and TreeDyn for sequence alignment, tree calculation, and image generation, respectively.

[0233] Example 5. Expression of glycans and cloning of the K. pneumoniae O2a glycan gene S. pneumonia CPS8 glycan is used in plasmid pB8(Tet R (Kay, EJ, et al. (2016) Open Biology 6, 150243) was expressed from plasmid pPR1347 (Kan R (Neal, BL, et al. (1993) Journal of Bacteriology 175, 7115-7118) The E. coli O16 wbbL gene was expressed from plasmid pMF19 (Sp RGBSIII glycan was expressed from the pBBR1MCS2 derivative (Feldman, MF, et al. (2005) Proceedings of the National Academy of Sciences of the United States of America 102, 3016), and GBSIII glycan was expressed from the pBBR1MCS2 derivative (Duke, JA, et al. (2021) ACS Infectious Diseases 7, 3111-3123). Bioconjugation with the K. pneumoniae O2a O-antigen has not been reported to date. To clone the gene encoding the mechanism necessary for O2a glycan synthesis, PCR was used to amplify the wzm, wzt, wbbM, glf, wbbN, and wbbO genes (Clarke, BR, et al. (2018) Journal of Biological Chemistry 293, 4666-4679) from the K. pneumoniae strain NTUH K2044 genomic DNA. K. pneumoniae was cultured overnight in LB medium until saturated, and genomic DNA was isolated using the Wizard Genomic DNA Purification Kit (Promega). The plasmid backbone of the O2a cluster was plasmid pBBR1MCS2(Kan R The plasmids were amplified from (Kovach, ME, et al. (1995) Gene 166, 175-176). The primers for these reactions are listed in Table 2. The PCR products from these reactions were assembled using Gibson assembly with the NEBuilder HiFi DNA Assembly Kit (NEB). Stellar cells were transformed, and the plasmids were isolated and validated as described in the previous section.

[0234] Example 6. Bioconjugation and Western blot The E. coli strains used in the bioconjugation experiments were either SDB1 or CLM24 (Feldman, MF, et al. (2005) Proceedings of the National Academy of Sciences of the United States of America 102, 3016). SDB1 is a W3110 E. coli derivative with mutations in the genes encoding WecA, a glycosyltransferase that initiates the synthesis of the endogenous E. coli O16 antigen, and WaaL, an enzyme that transfers Und-PP-linked glycans to lipid A-core sugars to produce LPS. CLM24 is a W3110 derivative with only a deletion of waaL. By eliminating these genes, crosstalk between the heterologous bioconjugation system and the endogenous E. coli glycosylation pathway is prevented. To prepare E. coli strains for bioconjugation, the inventors electroporated plasmids using competent cells prepared as previously described (Knoot, CJ, et al. (2021) Glycobiology 31, 1192-1203), followed by growth in SOB at 37°C. The cells were plated onto LB agar with appropriate antibiotics. The following day, 8-10 colonies were harvested, inoculated with antibiotics into LB or TB, and grown overnight with shaking at 30°C. The next morning, the starter cultures were inoculated into either 30 mL of medium in a 125 mL Erlenmeyer flask or 1 L of medium in a 2 L flask, with an initial optical density of 600 nm (OD 600 The OD became 0.05. 600 The culture was grown with shaking at 175 RPM until the OD reached 0.4-0.6, at which point 1 mM IPTG was introduced into the culture. Unless otherwise specified, all bioconjugation experiments were performed at 30°C. After overnight introduction, total growth reached 20-24 hours, and OD 600 The ophthalmos was measured, and cells in 0.5 OD units were pelleted for analysis.

[0235] The cell pellet was suspended in 100 μl of 1X Laemmli Buffer (Biorad) and boiled at 100°C for 10 minutes. The boiled sample was briefly centrifuged at 10,000 rcf, and equivalent volumes were obtained from each lane with the same OD. 600 The proteins were normalized and loaded onto 7.5% Mini-Protean TGX gel (Biorad) for SDS-PAGE separation. The proteins were transferred to a nitrocellulose membrane using a semi-dry electrode system and blocked with Intercept Blocking Buffer (Li-Cor) for 1 hour. The membrane was incubated with primary antibody for 45 minutes using 1:1 blocking and TBST. Commercially available rabbit anti-EPA antibody and mouse anti-6xHis antibody (Millipore-Sigma) were used for protein detection. Rabbit glycan antibodies for CPS8, GBSIII, and O16 were purchased from SSI Diagnostica. The K. pneumoniae rabbit O2a antibody was generously provided by Professor Chris Whitfield (University of Guelph) (Clarke, BR, et al. (2018) Journal of Biological Chemistry 293, 4666-4679). The Salmonella B group rabbit antibody was purchased from BD. After primary incubation, the membranes were washed three times with TBST buffer for a total of 15 minutes. The membranes were then incubated for 30 minutes with the secondary antibody IRDye 680RD goat anti-mouse antibody and / or IRDye 800CW goat anti-rabbit antibody (Li-Cor) in 1:1 blocking buffer and TBST. After a final 15-minute TBST wash, the membranes were imaged using Li-Cor Odyssey CLx.

[0236] Example 7. Recombinant M. osloensis Pil Mo Δ28 Lys-C digestion Intragellation was achieved following a slightly modified protocol of Shevchenko et al. (Shevchenko, A., et al. (2006) Nat Protoc 1, 2856-2860). Gel-isolated glycosylated EPA-PilΔ28 was excised and destained twice at room temperature for 10 minutes each time with a destaining solution (50 mM NH4HCO3, 50% ethanol) while shaking at 750 RPM. The destained bands were then dehydrated with 100% ethanol for 10 minutes, dried by vacuum centrifugation for 10 minutes, and rehydrated in a 50 mM NH4HCO3 solution of 10 mM DTT. Reduction was carried out at 56°C for 60 minutes, after which the gel bands were dehydrated twice with 100% ethanol for 10 minutes each time to remove any remaining reduction buffer. Next, the reduced sample was sequentially alkylated in the dark at room temperature for 45 minutes with a 50 mM NH4HCO3 solution of 55 mM iodoacetamide. The alkylated sample was then washed four times for 10 minutes each with 50 mM NH4HCO3, followed by 100% ethanol, followed by 50 mM NH4HCO3, followed by 100% ethanol, and then dried by vacuum centrifugation. The dried alkylated sample was then rehydrated at 4°C for 1 hour with a 40 mM NH4HCO3 solution of 20 ng / μl Lys-C endoprotease (Wako Chemicals). Excess Lys-C was removed, gel sections were covered with 40 mM NH4HCO3 and incubated overnight at 37°C. The peptide was concentrated and C 18 The samples were desalted using stage tips (Ishihama, Y., et al. (2006) J Proteome Res 5, 988-994, Rappsilber, J., et al. (2007) Nat Protoc 2, 1896-1906), then eluted with buffer B (0.5% acetic acid, 80% acetonitrile (ACN)), dried, and stored at -20°C prior to LC-MS analysis.

[0237] Example 8. Recombinant M. osloensis Pil using reverse-phase LC-MS Mo Analysis of Δ28 C 18The concentrated digest is resuspended in buffer A* (0.1% TFA, 2% ACN) and PepMap100 C 18 20mm x 75μm trap and PepMap C 18 The samples were separated using a two-column chromatography system equipped with a 500 mm × 75 μm analytical column (Thermo Fisher Scientific). The samples were concentrated on a trap column at 5 μl / min for 5 minutes using 0.1% formic acid (FA), and injected at 300 nl / min into an Orbitrap Fusion® Lumos® Tribrid® mass spectrometer with a FAIMS Pro interface (Thermo Fisher Scientific) via an analytical column using a Dionex Ultimate 3000 UPLC (Thermo Fisher Scientific) by varying the concentrations of buffer A (2% DMSO, 0.1% FA) and buffer B (78% ACN, 2% DMSO, 0.1% FA). 140-minute analytical units were used for potential glycopeptide identification, and 60-minute units were used for targeted analysis. Within a specific analyte, the buffer composition was changed from 3% buffer B to 28% buffer B over 120 minutes, from 28% buffer B to 40% buffer B over 9 minutes, from 40% buffer B to 100% buffer B over 3 minutes, then held at 100% buffer B for 2 minutes, then reduced to 3% buffer B over 2 minutes, and further held at 3% buffer B for 8 minutes. The Lumos™ mass spectrometer was operated in a stepwise FAIMS data-dependent mode with three different FAIMS CVs, -25, -45, and -65, as previously described (Ahmad Izaham, AR, et al. (2021) J Proteome Res 20, 599-612), acquiring a single Orbitrap MS scan (resolution 60k) every 1.5 seconds, followed by Orbitrap HCD scans (maximum fill time 120 ms, AGC 2×10) at each of the three FAIMS CVs. 5 The Orbitrap MS-MS scan resolution was switched between 30k and NCE 25, 30, and 45. MoFor the targeted characterization of glycopeptides, analytes were performed in which the buffer composition was changed from 3% buffer B to 15% buffer B over 30 minutes, from 15% buffer B to 30% buffer B over 10 minutes, from 30% buffer B to 80% buffer B over 5 minutes, then the composition was held in 100% buffer B for 5 minutes, then reduced to 3% buffer B over 1 minute, and held in 3% buffer B for a further 9 minutes. (HCD (maximum filling time 250 ms, AGC 2.5 × 10⁻¹) 5 Orbitrap MS-MS scan resolution 30k, NCE 15, 30, 35) and EThcD (maximum loading time 250ms, AGC 2.5×10) 5 Orbitrap MS-MS scan resolution 30k, HexHexA modified glycopeptide 762 FLPANCRGT 770 Parallel reaction monitoring was performed using a calibrated charge-dependent ETD parameter (ETD reaction time controlled using Rose, CM, et al. (2015) J Am Soc Mass Spectrom 26, 1848-1857) for a +2 charge state at (687.2972 m / z) via FAIMS CV at -45.

[0238] Example 9. Pil Mo Open search for Δ28 and annotation of HexHexA-modified C-terminal peptides Pil Mo Identification of glycosylation events was achieved using open database searches, as previously described (Lewis, JM, et al. (2021) J Vis Exp). Briefly, the data files were from Moraxella osloensis Pil MoThe sequence (NCBI accession: WP_156627541.1) was searched using FragPipe (version 17.1) MSfragger 3.4 (Polasky, DA, et al. (2020) Nat Methods 17, 1125-1132, Kong, AT, et al. (2017) Nat Methods 14, 513-520). The search was performed using "Lys-C" enzyme specificity, with cysteine ​​carbamide methylation as the fixed modification and methionine oxidation as the variable modification, and a maximum of two missing cleavages was tolerated. A mass tolerance range of 0–2000 Da, called the delta mass, was allowed to enable the identification of potential glycosylation events. C-terminal peptide 762 FLPANCRGT 770 The delta mass observed on (SEQ ID NO: 61) was manually examined to identify potential glycosylation events. HexHexA-modified glycopeptide 762 FLPANCRGT 770The parallel reaction monitoring results corresponding to (SEQ ID NO: 62) were manually extracted using Freestyle Viewer (1.7SP1, Thermo Fisher Scientific), and the MS / MS data were annotated using Interactive Peptide Spectral Annotator (Brademan, DR, et al. (2019) Mol Cell Proteomics 18, S193-S201) (interactivepeptidespectralannotator.com / PeptideAnnotator.html). Spectral annotation made it possible to modify the terminal T residue not only with HexHexA (338.0849 Da) but also with Hex (162.0528 Da). The resulting MS data and search results are deposited in the PRIDE ProteomeXchange Consortium repository (Perez-Riverol, Y., et al. (2019) Nucleic Acids Res 47, D442-D450, Perez-Riverol, Y., et al. (2015) Proteomics 15, 930-949) and can be accessed using identifier PXD033468.

[0239] Example 10. Purification of bioconjugate proteins Cells for protein purification were grown in 1 L of TB medium, and the bioconjugates were isolated using an osmotic shock protocol. After overnight growth and transposition, the cells were pelletized by centrifugation and washed with 0.9% NaCl. The washed cell pellet was suspended in 200 mM Tris-HCl pH 8.5, 100 mM EDTA, and 25% sucrose and incubated at 4°C for 30 minutes with rolling. Cells were pelletized by centrifugation at 4,700 rcf for 30 minutes, and the resulting pellet was suspended in 20 mM Tris-HCl pH 8.5 and incubated at 4°C for 45 minutes with rolling. The suspension was centrifuged at 18,000 rcf for 30 minutes. The supernatant containing the periplasm fraction was concentrated and either loaded directly onto an FPLC anion exchange column, or, in the case of the His-tagged EPA-PilΔ28 bioconjugate, purified using a nickel IMAC as previously described (Knoot, CJ, et al. (2021) Glycobiology 31, 1192-1203). The periplasm extract or IMAC eluate was concentrated, buffered to 20 mM Tris-HCl pH 8.0, filtered through a 0.2 μm PES filter, and then loaded onto an Aekta pure FPLC instrument (Cytiva) equipped with a SOURCE 15Q 4.6 / 100 PE anion exchange column (Cytiva). The bioconjugate was eluted at 2 mL / min using a stepwise gradient with buffer A (20 mM Tris pH 8) and buffer B (20 mM Tris pH 8, 1 M NaCl) (incrementing by 5% from 0% B to 25% in 10 column volumes for each concentration). The bioconjugate for immunization was further purified using a Superdex 200 Increase 10 / 300 GL column. The concentrated bioconjugate pooled from the anion exchange column was loaded onto a pre-equalized Superdex 200 column in PBS buffer and eluted at a flow rate of 0.75 mL / min. The fractions containing the purified bioconjugate were pooled, concentrated, and frozen at -80°C for storage.Protein concentrations for immunization and Western blotting were measured using the Pierce BCA Protein Assay kit (Thermo Fisher). The ratio of polysaccharides to protein for calculating vaccine dosage was determined using the method described by Duke et al. (Duke, JA, et al. (2021) ACS Infectious Diseases 7, 3111-3123).

[0240] Example 11. Immunization of mice Immunization of all mice was carried out in accordance with ethical guidelines for animal experimentation and research. The experiment was conducted at the Washington University School of Medicine in St. Louis, following institutional guidelines and approved by the Institutional Animal Care and Use Committee of Washington University in St. Louis. Five-week-old female CD-1 inbred mice (Charles River Laboratories) were subcutaneously injected with 100 μL of vaccine on days 0, 14, and 28. The vaccinated groups were 291 only (5 μg protein) and GBSIII-291 (5 μg protein, 1 μg polysaccharide). Serum from the mice was collected on days 0, 14, 28, and 42. All vaccines were formulated with Alhydrogel® 2% aluminum hydroxide gel (InvivoGen) in a 1:9 ratio (50 μL of vaccine and 5.5 μL of alum dissolved in 44.5 μL of 1x sterile phosphate-buffered saline).

[0241] Example 12. Enzyme-linked immunosorbent assay (ELISA) IgG kinetic titer was measured using enzyme-linked immunosorbent assay (ELISA). Briefly, a 96-well plate (TRP Immunomaxi plate) was used to measure glycosylated E. coli expressing GBSIII capsule polysaccharide in sodium carbonate buffer (approximately 10%). 6The plates were triple-coated overnight with CFU / 100μL. The coated E. coli strains were grown as described above, cultured overnight to induce GBSIII expression, then washed, diluted, and coated onto the plates. The wells were blocked with 1% BSA in PBS, washed with 0.05% PBS-Tween (PBST), and all subsequent washes were the same. Mouse serum was diluted 1:100, added to the wells, left at room temperature for 1 hour, and then washed. Total IgG titer was detected by HRP-conjugated anti-mouse IgG (GE Lifesciences, 1:5000 dilution) added to the wells at room temperature for 1 hour. After washing, the plates were developed using a 3,3′,5,5′ tetramethylbenzidine (TMB) substrate (Biolegend) and stopped at 2N H2SO4. Optical density was measured at 450 nm using a microplate reader (Bio-Tek). To generate standard curves for data fitting, total IgG products were determined using IgG standards. Standard wells were coated with IgG in sodium carbonate buffer and then treated in the same way as the sample wells. All wells were normalized to blank wells treated in the same way as all sample wells, except for primary mouse serum. Significance was determined using the Mann-Whitney nonparametric test, with P<0.05.

[0242] The scope and breadth of this disclosure should not be limited by any of the exemplary embodiments described above, but should be defined solely in accordance with the following claims and their equivalents. *****

[0243] Certain embodiments of this disclosure may be defined in any of the following numbered paragraphs.

[0244] 1. A fusion protein comprising (i) a glycosylated fragment and (ii) a first polypeptide tag, wherein the first polypeptide tag is capable of spontaneously forming an isopeptide bond with a second polypeptide tag binding partner. Optionally, the glycosylated fragment is at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, or 40 amino acid lengths. Optionally, the glycosylated fragment is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, 40, 50, 60, 80, or 100 amino acids or less in length. Optionally, the glycosylated fragment is between 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, or 30 amino acid lengths, and between 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, or 40 amino acid lengths. Optionally, the fusion protein includes a carrier protein. 2. The fusion protein is a complex carbohydrate containing sugars covalently bonded to the fusion protein via glycosylated fragments. Optionally, the sugar is covalently bonded to the glycosylated fragment via an N-bond, O-bond, or C-bond. The fusion protein described in paragraph 1, wherein the complex carbohydrate is optionally immunogenic. 3. The fusion protein described in paragraph 1 or 2, wherein the first polypeptide tag is translationally fused at the N-terminus or C-terminus of the fusion protein. 4. The first polypeptide tag is translationally fused internally within the fusion protein. Optionally, the fusion protein described in paragraph 1 or 2, wherein the first polypeptide tag is translationally fused internally within the sequence of the carrier protein. 5. A fusion protein as described in any one of paragraphs 1-4, wherein the glycosylated fragment is translationally fused at the N-terminus or C-terminus of the fusion protein. 6. The glycosylated fragment is translationally fused internally within the fusion protein. A fusion protein as described in any one of paragraphs 1-5, wherein, optionally, a glycosylated fragment is translationally fused internally within the sequence of a carrier protein. 7. The first polypeptide tag is SpyTag (sequence number 416), SpyTag002 (sequence number 417), SpyTag003 (sequence number 418), or DogTag (sequence number 419), Optionally, SpyTag, Spytag002, or Spytag003 are translationally fused at the N-terminus or C-terminus of the fusion protein. Optionally, a fusion protein described in any one of paragraphs 1-6, wherein the DogTag is translationally fused internally within the fusion protein. 8. The glycosylated fragment is a ComP glycosylated fragment. Optionally, the ComP glycosylated fragment is either the amino acid sequence CTGVTQIASGASAATTNVASAQC (SEQ ID NO: 412) or a fragment containing the amino acid ASA. Alternatively, it may contain or consist of a variant of Sequence ID No. 412 that has amino acid ASA at positions 11-13 and has 1, 2, 3, 4, 5, or 6 amino acid substitutions, additions, and / or deletions. Optionally, the ComP glycosylated fragment has the following amino acid sequence: [Table 11] Alternatively, a fusion protein as described in any one of paragraphs 1 to 7, comprising or consisting of a variant thereof having 1, 2, 3, 4, 5, or 6 amino acid substitutions, additions, and / or deletions, which includes the amino acid ASA corresponding to positions 11-13 of Sequence ID No. 412. 9. The glycosylated fragment is a TfpM-related pyring glycosylated fragment. Optionally, the TfpM-related pyring glycosylation fragment is either the PilMo pyring disulfide loop region (SEQ ID NO: 413) or a fragment containing at least the last three amino acids from the C-terminus of the TfpM-related pyring. Alternatively, a fusion protein as described in any one of paragraphs 1 to 7, comprising or consisting of a variant thereof having one, two, three, four, five, or six amino acid substitutions, additions, and / or deletions, which includes the last three amino acids from the C-terminus of TfpM-associated pyrin. 10. The glycosylated fragment is a PilE glycosylated fragment. Optionally, the PilE glycosylated fragment is either SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO: 414), or a fragment of SEQ ID NO: 414 that contains at least one amino acid WPGNNTSAGV (SEQ ID NO: 439) at positions 13-22. Alternatively, a fusion protein as described in any one of paragraphs 1 to 7, comprising or consisting of a variant thereof having at least the amino acid WPGNNTSAGV (sequence number 439) at positions 13 to 22 of sequence number 414 and having 1, 2, 3, 4, 5, or 6 amino acid substitutions, additions, and / or deletions. 11. The glycosylated fragment is the PglB glycosylated fragment. Optionally, the PglB glycosylated fragment comprises or consists of a consensus motif amino acid sequence X1X2N X3X4, wherein X1 is D or E, X2 is any amino acid except proline, X3 is any amino acid except proline, and X4 is S or T, as described in any one of paragraphs 1 to 7. 12. The glycosylated fragment is the PilA glycosylated fragment. Optionally, the PilA glycosylated fragment is either the PilA pyring disulfide loop region (SEQ ID NO: 415) or a fragment containing at least the last three amino acids from the C-terminus of PilA. Alternatively, a fusion protein as described in any one of paragraphs 1 to 7, comprising or consisting of a variant thereof having at least the last three amino acids from the PilA terminus and having substitutions, additions, and / or deletions of 1, 2, 3, 4, 5, or 6 amino acids. 13. The glycosylated fragment is the STT3 glycosylated fragment. Optionally, the fusion protein described in any one of paragraphs 1 to 7, wherein the STT3 glycosylated fragment contains or comprises the consensus motif amino acid sequence N X1X2, where X1 is any amino acid except proline and X2 is S or T. 14. The glycosylated fragment is an N-linked glycosyltransferase glycosylated fragment. Optionally, the N-linked glycosylated fragment contains or comprises a consensus motif amino acid sequence N X1X2, where X1 is any amino acid and X2 is S or T, as described in any one of paragraphs 1 to 7. 15. The glycosylated fragment is an O-linked glycosyltransferase glycosylated fragment. Optionally, the O-linked glycosylated fragment contains or comprises a serine or threonine-rich repeat fragment from a serine-rich repeat (SRR) adhesin of streptococci or staphylococci bacteria. Optionally, the fusion protein described in any one of paragraphs 1 to 7 comprises an O-linked glycosylated fragment containing or consisting of a serine or threonine-rich repeat from adhesin GspB of Streptococcus gordonii. 16. The fusion protein contains two or more glycosylated fragments, Optionally, the fusion protein contains at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 glycosylated fragments. Optionally, the fusion protein contains any of 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23 glycosylated fragments, up to any of 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 glycosylated fragments. Optionally, at least one glycosylated fragment is located at the N-terminus or C-terminus of the fusion protein, and at least one glycosylated fragment is located internally within the fusion protein. Optionally, at least two glycosylated fragments are located internally within the fusion protein. and / or A fusion protein as described in any one of paragraphs 1 to 15, wherein, optionally, one glycosylated fragment is located at the N-terminus of the fusion protein and another glycosylated fragment is located at the C-terminus of the fusion protein. 17. Two or more glycosylated fragments are the same, At least one of the two or more glycosylated fragments is different, or Each of the glycosylated fragments is different, Optionally, one glycosylated fragment is a ComP glycosylated fragment, and one glycosylated fragment is a TfpM-related pilling glycosylated fragment. Optionally, the fusion protein is a complex carbohydrate containing two or more sugars covalently bonded to the fusion protein via two or more glycosylated fragments. Optionally, the fusion protein contains at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 covalently linked sugars. Optionally, the fusion protein contains any of 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23 covalently linked sugars, up to any of 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 covalently linked sugars. Optional, Two or more sugars are the same, At least one of the two or more sugars is different, or Each of the sugars is a different fusion protein as described in paragraph 16. 18. A fusion protein as described in any one of paragraphs 1 to 17, wherein the fusion protein is selected from the group consisting of Escherichia coli maltose-binding protein, Pseudomonas aeruginosa exotoxin A (EPA), Pseudomonas aeruginosa PcrV, CRM197, Haemophilus influenzae protein D, cholera toxin B subunit, or tetanus toxin, and any fragment thereof. 19. A composition comprising a polypeptide pair comprising a first polypeptide and a second polypeptide, The first polypeptide is a fusion protein described in any one of paragraphs 1 to 18. The second polypeptide includes a second polypeptide tag binding partner to the first polypeptide tag of the first polypeptide, The first polypeptide is bound to the second polypeptide via an isopeptide bond between the first polypeptide tag and the second polypeptide tag. Optionally, the second polypeptide comprises a monomer polypeptide that can spontaneously multimerize / self-assemble into a higher-order multimeric structure, and / or Furthermore, optionally, the higher-order multimer structure is an icosahedral or dodecahedral particle (for example, similar to a nanocage), a virus-like particle, or an adenovirus vector. 20. The second polypeptide contains adenovirus capsid structural protein, The second polypeptide contains the coat protein of bacteriophage AP205. The second polypeptide contains a fragment of 2-keto-3-deoxy-phosphogluconate aldolase (i301), or The polypeptide versus composition according to paragraph 19, wherein the second polypeptide comprises a fragment of mutated 2-keto-3-deoxy-phosphogluconate aldolase (mi3). 21. The second polypeptide tag is SpyCatcher (sequence number 420), SpyCatcher002 (sequence number 421), SpyCatcher003 (sequence number 422), or DogCatcher (sequence number 423), Optional, The first polypeptide tag is SpyTag, and the second polypeptide tag is SpyCatcher. The first polypeptide tag is SpyTag002, and the second polypeptide tag is SpyCatcher002. The first polypeptide tag is SpyTag003 and the second polypeptide tag is SpyCatcher003, or A polypeptide pair composition according to paragraph 19 or 20, wherein the first polypeptide tag is DogTag and the second polypeptide tag is DogCatcher. 22. The second polypeptide tag is translationally fused to the N-terminus or C-terminus of the second polypeptide, or A polypeptide-versus-composition according to any one of paragraphs 19-21, wherein a second polypeptide tag is internally translatably fused within the second polypeptide. 23. The first polypeptide is a bioconjugate containing a sugar covalently bonded to a glycosylated fragment of the first polypeptide. A polypeptide pair composition according to any one of paragraphs 19 to 22, wherein the composition is optionally immunogenic. 24. A polypeptide versus composition according to any one of paragraphs 19 to 23, further comprising an adjuvant and / or excipient. 25. A polypeptide pair composition according to any one of paragraphs 19 to 24, which is a pharmaceutical / therapeutic composition. 26. A polypeptide versus composition according to any one of paragraphs 19 to 25, wherein the composition is a conjugate vaccine. 27. A complex comprising two or more polypeptide pairs described in any one of paragraphs 19 to 26, Optionally, the complex comprises 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 75, 100, 125, 150, 175, 200, 225, 250 or more complexing polypeptide pairs as described in any one of paragraphs 19-26. Optionally, the complex is made from any of the 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 75, 100, 125, 150, 175, 200, 225, 250, 300, or 400 complexing polypeptide pairs listed in any one of paragraphs 19-26. Includes up to any of the 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 75, 100, 125, 150, 175, 200, 225, 250, 300, 400, or 500 complexed polypeptide pairs described in any one of paragraphs 19-26, and / or A complex in which, by arbitrary selection, the complex is a self-organized multimeric higher-order structure. 28. The complex according to paragraph 27, wherein the self-assembled multimer higher-order structure is an icosahedral or dodecahedral particle (e.g., similar to a nanocage), a virus-like particle, or an adenovirus vector. 29. The complex described in paragraph 27 or 28, wherein all of the first polypeptides of the complex contain the same fusion protein. 30. The complex according to paragraph 27 or 28, wherein at least two of the first polypeptides of the complex contain different fusion proteins. 31. The complex is a bioconjugate in which at least one first polypeptide contains a sugar covalently bonded to a glycosylated fragment of the first polypeptide. Optionally, at least about 5%, 10%, 25%, 50%, 75%, 80%, 90%, 95%, 97%, 98%, or 99% of the first polypeptide of the complex is a bioconjugate. Optionally, the bioconjugate comprises approximately 5%, 10%, 25%, 50%, 75%, 80%, 90%, 95%, 97%, or 98% of the first polypeptide of the complex, and up to approximately 10%, 25%, 50%, 75%, 80%, 90%, 95%, 97%, 98%, or 99% of the first polypeptide of the complex. Optionally, 100% of the first polypeptide in the complex is a bioconjugate. Optionally, a complex described in any one of paragraphs 27-30, wherein the complex is immunogenic. 32. Two or more of the first polypeptides of the complex are bioconjugates containing covalently bonded sugars. Optionally, the complex contains at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1,000, 1,500, 2,000, 2,500, or 5,000 covalently bonded sugars. Optionally, the complex may consist of 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1,000, 1,500, 2,000, or 2,500 covalently bonded sugars. It contains up to 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1,000, 1,500, 2,000, 2,500, or 5,000 covalently bonded sugars. If all sugars bound to the complex are of the same type, Optionally, at least one of the two or more sugars bound to the complex is different. Optionally, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, or 33 different sugars are present bound to the complex. The complex described in any one of paragraphs 27-31, wherein each of the sugars bound to the complex is different. 33. A pharmaceutical / therapeutic composition comprising the complex described in any one of paragraphs 27 to 32, and an adjuvant and / or excipient. 34. The complex or composition described in any one of paragraphs 27 to 32 or the pharmaceutical / therapeutic composition described in paragraph 33, wherein the complex or composition is a conjugate vaccine. 35. A method for producing a polypeptide pair according to any one of paragraphs 19 to 26, comprising contacting a first polypeptide and a second polypeptide under conditions that enable the first polypeptide tag to spontaneously form an isopeptide bond with a second polypeptide tag binding partner. 36. The method further comprises glycosylation of the first polypeptide with a sugar prior to contact with the second polypeptide and isopeptide bond formation. Optionally, the first polypeptide is glycosylated in vivo before contact with the second polypeptide and isopeptide bond formation. The method according to paragraph 35, optionally comprising isolating / purifying a first polypeptide glycosylated in vivo prior to contact with a second polypeptide and isopeptide bond formation. 37. The method according to paragraph 35, wherein the first polypeptide is glycosylated after contact with the second polypeptide and formation of an isopeptide bond. 38. (i) Contacting the first polypeptide and the second polypeptide under conditions that allow the second polypeptide to form a self-assembled multimeric higher-order structure, and then allow the first polypeptide tag to spontaneously form an isopeptide bond with the second polypeptide tag, or (ii) A method for producing the complex according to any one of paragraphs 27 to 32, comprising: (ii) contacting the first polypeptide and the second polypeptide under conditions that allow the first polypeptide tag to spontaneously form an isopeptide bond with the second polypeptide tag, and then form a self-assembled multimeric higher-order structure of the second polypeptide. 39. Further comprising glycosylation of the first polypeptide with sugar, optionally, The first polypeptide is glycosylated before an isopeptide bond is formed between the first polypeptide and the second polypeptide. The first polypeptide is glycosylated after an isopeptide bond is formed between the first polypeptide and the second polypeptide. The first polypeptide is glycosylated before being incorporated into the higher-order structure of the polymer, and / or The method according to paragraph 38, wherein the first polypeptide is incorporated into the higher-order structure of the polymer and then glycosylated. 40. The method according to any one of paragraphs 36, 37, or 39, wherein the sugar is converted to a fusion protein by the action of an N-linked oligosaccharide transferase (N-OTase), an O-linked oligosaccharide transferase (O-OTase), an N-linked glycosyltransferase (NGT), an O-linked glycosyltransferase (OGT), and / or a C-mannosyltransferase (CMT). 41. The ComP glycosylated fragment is glycosylated by PglS OTase, The TfpM-related pyring glycosylated fragment is glycosylated by TfpM OTase, and optionally, the ComP glycosylated fragment is glycosylated by PglS OTase, and the TfpM-related pyring glycosylated fragment is glycosylated by TfpM OTase. The PilE glycosylated fragment is glycosylated by PglL OTase, The PglB glycosylated fragment is glycosylated by PglB OTase. The PilA glycosylated fragment is glycosylated by TfpO or PilO OTase. The STT3 glycosylated fragment is glycosylated by the STT3 catalytic subunit. The PilA_Pa5196-related pyring glycosylated fragment is glycosylated by TfpW glycosyltransferase, N-linked glycosyltransferase glycosylated fragments are glycosylated by N-linked glycosyltransferases from Actinobacillus pleuropneumoniae, Haemophilus influenzae, or Yersinia enterocolitica, and / or The method according to any one of paragraphs 36, 37, 39, or 40, wherein an O-linked glycosylated fragment is glycosylated by a GtfA / GtfB glycosyltransferase. 42. The sugar covalently bonds to the oxygen atom in the glycosylated fragment using PglS OTase. A sugar is covalently bonded to the oxygen atom in the glycosylated fragment using TfpM OTase, and optionally, another sugar is covalently bonded to the oxygen atom in the glycosylated fragment using PglS OTase, and yet another sugar is covalently bonded to the oxygen atom in the glycosylated fragment using TfpM OTase. The sugar covalently bonds to the oxygen atom in the glycosylated fragment using PglL OTase. The sugar covalently bonds to the nitrogen atom in the glycosylated fragment using PglB OTase. The sugar covalently bonds to the oxygen atom in the glycosylated fragment using TfpO or PilO OTase. The sugar covalently b...

Claims

1. A fusion protein comprising (i) a glycosylated fragment and (ii) a first polypeptide tag, wherein the first polypeptide tag is capable of spontaneously forming an isopeptide bond with a second polypeptide tag binding partner, Optionally, the glycosylated fragment has a length of at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, or 40 amino acids. Optionally, the glycosylated fragment is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, 40, 50, 60, 80, or 100 amino acids or less in length. Optionally, the glycosylated fragment is any of the following amino acid lengths: 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, or 30 amino acids, and any of the following amino acid lengths: 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, or 40 amino acids. Optionally, the fusion protein includes a carrier protein.

2. The fusion protein is a complex carbohydrate containing a sugar covalently bonded to the fusion protein via the glycosylated fragment, Optionally, the sugar is covalently bonded to the glycosylated fragment via an N-bond, O-bond, or C-bond. The fusion protein according to claim 1, wherein the complex carbohydrate is optionally immunogenic.

3. The fusion protein according to claim 1, wherein the first polypeptide tag is translationally fused to the N-terminus or C-terminus of the fusion protein.

4. The first polypeptide tag is translationally fused internally within the fusion protein. The fusion protein according to claim 1, wherein the first polypeptide tag is optionally translationally fused internally within the sequence of a carrier protein.

5. The fusion protein according to claim 1, wherein the glycosylated fragment is translationally fused to the N-terminus or C-terminus of the fusion protein.

6. The glycosylated fragment is translationally fused internally within the fusion protein. The fusion protein according to claim 1, wherein the glycosylated fragment is optionally translationally fused internally within the sequence of the carrier protein.

7. The first polypeptide tag is SpyTag (sequence number 416), SpyTag002 (sequence number 417), SpyTag003 (sequence number 418), or DogTag (sequence number 419), Optionally, SpyTag, SpyTag002, or SpyTag003 is translationally fused to the N-terminus or C-terminus of the fusion protein. The fusion protein according to claim 1, wherein the DogTag is optionally translationally fused internally within the fusion protein.

8. The glycosylated fragment is a ComP glycosylated fragment. Optionally, the ComP glycosylated fragment is an amino acid sequence Table 1 Or a fragment containing the amino acid ASA, Alternatively, it may contain, or consist of, a variant of sequence number 412 having the amino acid ASA at positions 11-13 and having 1, 2, 3, 4, 5, or 6 amino acid substitutions, additions, and / or deletions. Optionally, the ComP glycosylated fragment has the following amino acid sequence: Table 2 Alternatively, the fusion protein according to claim 1, comprising or consisting of a variant thereof having 1, 2, 3, 4, 5, or 6 amino acid substitutions, additions, and / or deletions, the amino acid ASA corresponding to positions 11-13 of SEQ ID NO:

412.

9. The glycosylated fragment is a TfpM-related pyring glycosylated fragment, Optionally, the TfpM-related pyring glycosylated fragment is a PilMo pyring disulfide loop region (SEQ ID NO: 413), or a fragment containing at least the last three amino acids from the C-terminus of the TfpM-related pyring. Alternatively, the fusion protein according to claim 1, comprising or consisting of a variant thereof having one, two, three, four, five, or six amino acid substitutions, additions, and / or deletions, comprising the last three amino acids from the C-terminus of the TfpM-related pyrin.

10. The glycosylated fragment is a PilE glycosylated fragment, Optionally, the PilE glycosylated fragment is an amino acid Table 3 Alternatively, the fragment containing at least the amino acid WPGNNTSAGV (SEQ ID NO: 439) at positions 13-22 of SEQ ID NO: 414, Alternatively, the fusion protein according to claim 1, comprising or consisting of a variant thereof having at least the amino acid WPGNNTSAGV (Sequence ID 439) at positions 13 to 22 of Sequence ID No. 414, and having 1, 2, 3, 4, 5, or 6 amino acid substitutions, additions, and / or deletions.

11. The glycosylated fragment is a PglB glycosylated fragment. Optionally, the PglB glycosylated fragment is the consensus motif amino acid sequence X 1 X 2 N X 3 X 4 It includes or consists of, in the formula, X 1 However, it is D or E, and X 2 However, it is any amino acid except proline, X 3 However, it is any amino acid except proline, X 4 The fusion protein according to claim 1, wherein the protein is S or T.

12. The glycosylated fragment is a PilA glycosylated fragment, Optionally, the PilA glycosylated fragment is either the PilA pyring disulfide loop region (SEQ ID NO: 415) or the fragment containing at least the last three amino acids from the C-terminus of PilA. Alternatively, the fusion protein according to claim 1, comprising or consisting of a variant thereof that includes at least the last three amino acids from the PilA terminus and has substitutions, additions, and / or deletions of 1, 2, 3, 4, 5, or 6 amino acids.

13. The glycosylated fragment is an STT3 glycosylated fragment. Optionally, the STT3 glycosylation fragment comprises or consists of the consensus motif amino acid sequence N X 1 X 2 wherein X 1 is any amino acid except proline, and X 2 is S or T, the fusion protein according to claim 1.

14. The glycosylated fragment is an N-linked glycosyltransferase glycosylated fragment, Optionally, the N-linked glycosylated fragment may be a consensus motif amino acid sequence N X 1 X 2 It includes or consists of, in the formula, X 1 However, it is any amino acid, X 2 The fusion protein according to claim 1, wherein the protein is S or T.

15. The glycosylated fragment is an O-linked glycosyltransferase glycosylated fragment, Optionally, the O-linked glycosylated fragment comprises or consists of a serine or threonine-rich repeat fragment from a serine-rich repeat (SRR) adhesin of streptococci or staphylococci bacteria. The fusion protein according to claim 1, wherein the O-linked glycosylated fragment optionally comprises or consists of a serine or threonine-rich repeat from adhesin GspB from Streptococcus gordonii.

16. The fusion protein comprises two or more glycosylated fragments, Optionally, the fusion protein comprises at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 glycosylated fragments. Optionally, the fusion protein comprises any of 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23 glycosylated fragments, up to any of 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 glycosylated fragments. Optionally, at least one glycosylated fragment is located at the N-terminus or C-terminus of the fusion protein, and at least one glycosylated fragment is located internally within the fusion protein. Optionally, at least two glycosylated fragments are located internally within the fusion protein. and / or The fusion protein according to claim 1, wherein, optionally, one glycosylated fragment is located at the N-terminus of the fusion protein, and another glycosylated fragment is located at the C-terminus of the fusion protein.

17. The two or more glycosylated fragments are the same, At least one of the two or more glycosylated fragments is different, or Each of the glycosylated fragments is different, Optionally, one glycosylated fragment is a ComP glycosylated fragment, and one glycosylated fragment is a TfpM-related pyring glycosylated fragment. Optionally, the fusion protein is a complex carbohydrate containing two or more sugars covalently bonded to the fusion protein via two or more glycosylated fragments. Optionally, the fusion protein contains at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 covalently bonded sugars. Optionally, the fusion protein contains any of 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23 covalently bonded sugars, up to any of 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 covalently bonded sugars. Optional, The two or more sugars mentioned above are the same, At least one of the two or more sugars is different, or The fusion protein according to claim 16, wherein each of the aforementioned sugars is different.

18. The aforementioned fusion protein includes a carrier protein, The fusion protein according to claim 1, wherein the carrier protein is optionally selected from the group consisting of Escherichia coli maltose-binding protein, Pseudomonas aeruginosa exotoxin A (EPA), Pseudomonas aeruginosa PcrV, CRM197, Haemophilus influenzae protein D, cholera toxin B subunit, or tetanus toxin, and any fragment thereof.

19. A composition comprising a polypeptide pair comprising a first polypeptide and a second polypeptide, The first polypeptide is the fusion protein described in any one of claims 1 to 18. The second polypeptide includes a second polypeptide tag binding partner to the first polypeptide tag of the first polypeptide, The first polypeptide is bound to the second polypeptide via an isopeptide bond between the first polypeptide tag and the second polypeptide tag, Optionally, the second polypeptide comprises a monomer polypeptide that can spontaneously multimerize / self-assemble into a higher-order multimeric structure, and / or Furthermore, optionally, the higher-order multimer structure is an icosahedral or dodecahedral particle (for example, similar to a nanocage), a virus-like particle, or an adenovirus vector.

20. The second polypeptide comprises an adenovirus capsid structural protein, The second polypeptide comprises the coat protein of bacteriophage AP205, The second polypeptide comprises a fragment of 2-keto-3-deoxy-phosphogluconate aldolase (i301), or The polypeptide pair composition according to claim 19, wherein the second polypeptide comprises a fragment of mutated 2-keto-3-deoxy-phosphogluconate aldolase (mi3).

21. The second polypeptide tag is SpyCatcher (sequence number 420), SpyCatcher002 (sequence number 421), SpyCatcher003 (sequence number 422), or DogCatcher (sequence number 423), Optional, The first polypeptide tag is SpyTag, and the second polypeptide tag is SpyCatcher. The first polypeptide tag is SpyTag002, and the second polypeptide tag is SpyCatcher002. The first polypeptide tag is SpyTag003, and the second polypeptide tag is SpyCatcher003, or The polypeptide pair composition according to claim 19, wherein the first polypeptide tag is DogTag and the second polypeptide tag is DogCatcher.

22. The second polypeptide tag is translationally fused to the N-terminus or C-terminus of the second polypeptide, or The polypeptide pair composition according to claim 19, wherein the second polypeptide tag is translationally fused internally within the second polypeptide.

23. The first polypeptide is a bioconjugate containing a sugar covalently bonded to a glycosylated fragment of the first polypeptide, The polypeptide pair composition according to claim 19, wherein the composition is optionally immunogenic.

24. The polypeptide pair composition according to claim 19, further comprising an adjuvant and / or excipient.

25. The polypeptide pair composition according to claim 19, wherein the composition is a pharmaceutical / therapeutic composition.

26. The polypeptide pair composition according to claim 19, wherein the composition is a conjugate vaccine.

27. A complex comprising two or more polypeptide pairs according to claim 19, Optionally, the complex comprises 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 75, 100, 125, 150, 175, 200, 225, 250 or more complexed polypeptide pairs as described in any one of claims 19 to 26. Optionally, the complex is made from any of the 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 75, 100, 125, 150, 175, 200, 225, 250, 300, or 400 complexed polypeptide pairs described in any one of claims 19 to 26. A complexed polypeptide pair comprising up to any of 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 75, 100, 125, 150, 175, 200, 225, 250, 300, 400, or 500 complexed polypeptide pairs as described in any one of claims 19 to 26, and / or A complex in which, optionally, the complex is a self-organized multimeric higher-order structure.

28. The complex according to claim 27, wherein the self-assembled multimer higher-order structure is an icosahedral or dodecahedral particle (for example, similar to a nanocage), a virus-like particle, or an adenovirus vector.

29. The complex according to claim 27, wherein all of the first polypeptides of the complex contain the same fusion protein.

30. The complex according to claim 27, wherein at least two of the first polypeptides of the complex comprise different fusion proteins.

31. The complex comprises a bioconjugate in which at least one first polypeptide contains a sugar covalently bonded to a glycosylated fragment of the first polypeptide. Optionally, at least about 5%, 10%, 25%, 50%, 75%, 80%, 90%, 95%, 97%, 98%, or 99% of the first polypeptide of the complex is a bioconjugate. Optionally, the bioconjugate comprises approximately 5%, 10%, 25%, 50%, 75%, 80%, 90%, 95%, 97%, or 98% of the first polypeptide of the complex, up to approximately 10%, 25%, 50%, 75%, 80%, 90%, 95%, 97%, 98%, or 99% of the first polypeptide of the complex. Optionally, 100% of the first polypeptide of the complex is a bioconjugate. The complex according to claim 27, wherein the complex is optionally immunogenic.

32. Two or more of the first polypeptides of the complex are bioconjugates containing covalently bonded sugars. Optionally, the complex contains at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1,000, 1,500, 2,000, 2,500, or 5,000 covalently bonded sugars. Optionally, the complex may consist of any of the following covalent sugars: 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1,000, 1,500, 2,000, or 2,500. From to any of the following covalent sugars: Optionally, all sugars bound to the complex are the same. Optionally, at least one of the two or more sugars bound to the complex is different. Optionally, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, or 33 different sugars are present bound to the complex. The complex according to claim 27, wherein each of the sugars bound to the complex is optionally different.

33. A pharmaceutical / therapeutic composition comprising the complex according to claim 27 and an adjuvant and / or excipient.

34. The complex or composition according to claim 27 or the pharmaceutical / therapeutic composition according to claim 33, wherein the complex or composition is a conjugate vaccine.

35. A method for producing a polypeptide pair according to claim 19, comprising contacting the first polypeptide and the second polypeptide under conditions that enable the first polypeptide tag to spontaneously form an isopeptide bond with the second polypeptide tag binding partner.

36. The method further comprises glycosylation of the first polypeptide with a sugar prior to contact with the second polypeptide and isopeptide bond formation, Optionally, the first polypeptide is glycosylated in vivo before contact with the second polypeptide and isopeptide bond formation. The method according to claim 35, wherein the method optionally includes isolating / purifying the first polypeptide that has been glycosylated in vivo before contact with the second polypeptide and isopeptide bond formation.

37. The method according to claim 35, wherein the first polypeptide is glycosylated after contact with the second polypeptide and formation of an isopeptide bond.

38. (i) Contacting the first polypeptide and the second polypeptide under conditions that allow the second polypeptide to form a self-assembled polymeric higher-order structure, and then allow the first polypeptide tag to spontaneously form an isopeptide bond with the second polypeptide tag, or (ii) A method for producing the complex according to claim 27, comprising: (ii) contacting the first polypeptide and the second polypeptide under conditions that enable the first polypeptide tag to spontaneously form an isopeptide bond with the second polypeptide tag, and then to form a self-assembled polymeric higher-order structure of the second polypeptide.

39. The first polypeptide is further glycosylated with a sugar, optionally, The first polypeptide is glycosylated before the isopeptide bond is formed between the first polypeptide and the second polypeptide. The first polypeptide is glycosylated after the isopeptide bond is formed between the first polypeptide and the second polypeptide. The first polypeptide is glycosylated before being incorporated into the higher-order structure of the polymer, and / or The method according to claim 38, wherein the first polypeptide is incorporated into the higher-order structure of a polymer and then glycosylated.

40. The method according to claim 36, wherein the sugar is converted to the fusion protein by the action of N-linked oligosaccharide transferase (N-OTase), O-linked oligosaccharide transferase (O-OTase), N-linked glycosyltransferase (NGT), O-linked glycosyltransferase (OGT), and / or C-mannosyltransferase (CMT).

41. The ComP glycosylated fragment is glycosylated by PglS OTase, The TfpM-related pilling glycosylated fragment is glycosylated by TfpM OTase, and optionally, the ComP glycosylated fragment is glycosylated by PglS OTase, and the TfpM-related pilling glycosylated fragment is glycosylated by TfpM OTase. The PilE glycosylated fragment is glycosylated by PglL OTase, The PglB glycosylated fragment is glycosylated by PglB OTase, The PilA glycosylated fragment is glycosylated by TfpO or PilO OTase. The STT3 glycosylated fragment is glycosylated by the STT3 catalyst subunit. The PilA_Pa5196-related pyring glycosylated fragment is glycosylated by TfpW glycosyltransferase, N-linked glycosyltransferase fragments are glycosylated by N-linked glycosyltransferases from Actinobacillus pleuropneumoniae, Haemophilus influenzae, or Yersinia enterocolitica, and / or The method according to claim 36, wherein an O-linked glycosylated fragment of glycosyltransferase is glycosylated by a GtfA / GtfB glycosyltransferase.

42. The aforementioned sugar is covalently bonded to the oxygen atom in the glycosylated fragment using PglS OTase, The sugar is covalently bonded to the oxygen atom in the glycosylated fragment using TfpM OTase, optionally, the sugar is covalently bonded to the oxygen atom in the glycosylated fragment using PglS OTase, and another sugar is covalently bonded to the oxygen atom in the glycosylated fragment using TfpM OTase. The aforementioned sugar is covalently bonded to the oxygen atom in the glycosylated fragment using PglL OTase, The aforementioned sugar is covalently bonded to the nitrogen atom in the glycosylated fragment using PglB OTase, The aforementioned sugar is covalently bonded to the oxygen atom in the glycosylated fragment using TfpO or Pilo OTase. The aforementioned sugar is covalently bonded to the nitrogen atom in the glycosylated fragment using STT3 OTase. The aforementioned sugar is covalently bonded to the nitrogen atom in the glycosylated fragment using AlgB OTase, The aforementioned sugar is covalently bonded to the oxygen atom in the glycosylated fragment using TfpW glycosyltransferase, The aforementioned sugar is covalently bonded to the nitrogen atom in the glycosylated fragment using an N-linked glycosyltransferase, The aforementioned sugar is covalently bonded to the oxygen atom in the glycosylated fragment using an O-linked glycosyltransferase, The method according to claim 36, wherein the sugar is covalently bonded to a carbon atom in the glycosylated fragment using a C-mannosyltransferase.

43. The sugar is covalently bonded to an oxygen atom in a ComP glycosylated fragment (e.g., SEQ ID NO: 412, or a variant thereof) using PglS OTase (e.g., SEQ ID NO: 400), The sugar is covalently bonded to an oxygen atom in a TfpM glycosylated fragment (e.g., SEQ ID NO: 413, or a variant thereof) using TfpM OTase (e.g., SEQ ID NO: 402), The sugar is covalently bonded to an oxygen atom in a PilE glycosylated fragment (e.g., SEQ ID NO: 414, or its variant) using PglL OTase (e.g., SEQ ID NO: 404), [0472] The sugar is used to form a PglB glycosylated fragment (e.g., X) using PglB OTase (e.g., SEQ ID NO: 405). 1 X 2 N X 3 X 4 , in the formula, X 1 However, it is D or E, and X 2 However, it is any amino acid except proline, X 3 However, it is any amino acid except proline, X 4 However, it is covalently bonded to the nitrogen atom in (S or T), The sugar is covalently bonded to an oxygen atom in a PilA glycosylated fragment (e.g., SEQ ID NO: 415, or a variant thereof) using TfpO / PilO OTase (e.g., SEQ ID NO: 407), The aforementioned sugar uses STT3 OTase (e.g., SEQ ID NO: 408) to form an STT3 glycosylated fragment (e.g., NX 1 X 2 , in the formula, X 1 However, it is any amino acid except proline, X 2 However, it is covalently bonded to the nitrogen atom in (S or T), The aforementioned sugar is used with AlgB OTase (e.g., SEQ ID NO: 409) to create an Archaeal AlgB glycosylated fragment (e.g., NX). 1 X 2 , in the formula, X 1 However, it is any amino acid except proline, X 2 However, it is covalently bonded to the nitrogen atom in (S or T), The sugar is covalently bonded to an oxygen atom in a PilA_Pa5196-related pyring glycosylated fragment (e.g., SEQ ID NO: 426, or its variant) using a TfpW glycosyltransferase (e.g., SEQ ID NO: 424), The aforementioned sugar is processed using an N-linked glycosyltransferase (e.g., SEQ ID NO: 410) to obtain an N-linked glycosyltransferase sequence (e.g., N X 1 X 2 , in the formula, X 1 However, it is any amino acid, X 2 However, it is covalently bonded to the nitrogen atom in (S or T), The sugar is covalently bonded to the oxygen atom in the O-linked glycosyltransferase sequence using an O-linked glycosyltransferase (e.g., SEQ ID NO: 411), The method according to claim 42, wherein the sugar is covalently bonded to a carbon atom in a C-mannosyltransferase glycosylated fragment using C-mannosyltransferase.

44. The method according to claim 35, wherein the method is a method for producing a conjugate vaccine.

45. A system comprising the first polypeptide and the second polypeptide of the composition according to claim 1, Optionally, the first polypeptide is a glycosylated bioconjugate. Optionally, the system includes a polymeric higher-order structure assembled from the second polypeptide. Optionally, the system comprises a sugar and an N-linked oligosaccharide transferase (N-OTase), an O-linked oligosaccharide transferase (O-OTase), an N-linked glycosyltransferase (NGT), an O-linked glycosyltransferase (OGT), and / or a C-mannosyltransferase (CMT).

46. An isolated nucleic acid encoding the first polypeptide and / or the second polypeptide of the composition according to claim 26.

47. A vector comprising the isolated nucleic acid described in claim 46.

48. A host cell comprising the vector according to claim 47.

49. A kit comprising two or more components, each comprising the fusion protein, the first polypeptide, the second polypeptide, sugar, N-linked oligosaccharide transferase (N-Otase), O-linked oligosaccharide transferase (O-Otase), N-linked glycosyltransferase (NGT), O-linked glycosyltransferase (OGT), and / or C-mannosyltransferase (CMT), the bioconjugate, the multimeric higher-order structure assembled from the second polypeptide, the isolated nucleic acid, the vector, and the host cell.

50. A method for inducing an immune response in a subject by administering to the subject an effective amount of any composition, complex, and / or conjugate vaccine described in any of the prior claims, or a composition, complex, and / or conjugate vaccine described in any of the prior claims for use in inducing an immune response in the subject.

51. The glycosylated fragment is a PilA_Pa5196-related pyring glycosylated fragment, Optionally, the PilA_Pa5196-related pilling glycosylation fragment is a chain 1 and 2 of the antiparallel beta-sheet domain of PilA_Pa5196 (SEQ ID NO: 426). The fusion protein according to claim 1, or comprising, or consisting of, a variant having one, two, three, four, five, or six amino acid substitutions, additions, and / or deletions.

52. The fusion protein according to claim 1, wherein the glycosylated fragment comprises means for having a sugar bound to the glycosylated fragment by PglS OTase, TfpM OTase, PglL OTase, PglB OTase, TfpO / PilO OTase, STT3 OTase, AlgB OTase, N-linked glycosyltransferase, and / or O-linked glycosyltransferase.

53. Contains amino acid linkers, The amino acid linker is optionally selected from the group consisting of SGG, SEQ ID NO: 430, SEQ ID NO: 431, SEQ ID NO: 432, SEQ ID NO: 433, SEQ ID NO: 434, SEQ ID NO: 435, SEQ ID NO: 436, SEQ ID NO: 437, and SEQ ID NO:

438. Optionally, an amino acid linker translated immediately after a leader sequence, polypeptide tag, glycosylated fragment, carrier protein, and / or polyhistidine tag, and / or Optionally, a fusion protein according to any one of the prior claims, comprising a leader sequence, a polypeptide tag, a glycosylated fragment, a carrier protein, and / or an amino acid linker translated immediately before a polyhistidine tag.