Compositions and methods for producing glycoconjugate polypeptides having isopeptide bond to second polypeptide partner and uses of these glycoconjugate polypeptides

By using fusion proteins and oligosaccharide transferase systems in conjunction with the SpyTag/SpyCatcher system to generate glycoconjugate vaccines, the issues of flexibility and effectiveness in existing vaccine generation methods have been resolved, enabling the generation of diverse immunogenic compositions and enhanced immune responses.

CN120897997APending Publication Date: 2025-11-04VAXNEWMO LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202480018766.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-27
Filing Date
2024-02-26
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing methods for generating biological conjugate vaccines are relatively simple, lacking flexibility and effectiveness, and are difficult to meet the needs of diverse immunogenic compositions.

Method used

The fusion protein, containing glycosylated fragments and peptide tags, is used to form heteropeptide bonds using the SpyTag/SpyCatcher system. It is then combined with an oligosaccharide transferase system to generate a glycoconjugate vaccine, which self-assembles into higher-order polymeric structures such as nanoparticles or virus-like particles.

Benefits of technology

It enables the generation of diverse immunogenic compositions, improves the immune response, and enhances the immune memory and antibody response of vaccines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120897997A_ABST
    Figure CN120897997A_ABST
Patent Text Reader

Abstract

The present disclosure provides descriptions of compositions and methods for the production of glycoconjugate polypeptides using a glycosidic bond forming enzyme that then forms an isopeptide bond with a second polypeptide containing a polypeptide tag, as well as the use of the glycoconjugate polypeptides.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross Reference to Related Applications

[0002] This PCT application claims the benefit of U.S. Provisional Application No. 63 / 448,382, filed February 27, 2023, U.S. Provisional Application No. 63 / 448,386, filed February 27, 2023, and U.S. Provisional Application No. 63 / 448,408, filed February 27, 2023, each of which is incorporated by reference herein in its entirety.

[0003] Reference to Sequence Listing on File

[0004] The sequence listing in a file named “64100_234947_SL.xml” is provided with this document via the Patent Center of the USPTO and is incorporated by reference herein in its entirety, the file being 441,619 bytes in size (measured in ASCII text format), containing 440 sequences, and created on February 26, 2024. BACKGROUND

[0005] Glycan-protein or glycoprotein conjugate vaccines are effective therapeutic agents against a variety of bacterial pathogens and are essentially covalently linked to carrier proteins on the surface of bacteria, either oligosaccharides or polysaccharides. These vaccines can induce a protective immune response against the O antigen or capsule present on the surface of Gram-positive or Gram-negative pathogens. Unlike pure polysaccharide vaccines, glycoprotein conjugate vaccines can stimulate robust immune memory by inducing T cell recruitment, memory cell formation, and B cell IgM to IgG antibody class switching, thereby inducing the establishment of immune memory (Rappuoli et al. (2019) PNAS, 116(1) 14-16; Avci, F. et al. (2011) Nat Med 17, 1602–1609). Glycoprotein conjugate vaccines can be generated via a variety of different methods. Historically, vaccines for routine medical use have been synthetic conjugates made by chemically crosslinking purified oligosaccharides or polysaccharides with amino acid side chains on purified proteins using different linker molecules and chemical reactions (Berti, F. and Adamo, R. (2018) Chem. Soc. Rev., 2018, 47, 9015-9025). Bioconjugation is an alternative to chemical conjugation and relies on oligosaccharide transferases (OTases), which catalyze the covalent linking of lipid-linked oligosaccharides or polysaccharides with specific amino acid residues on the substrate protein (Harding, C. and Feldman, M. (2019) Glycobiology, Vol. 29, No. 7, July 2019, pp. 519–529; Feldman, M. (2005) PNAS, 102(8) 3016-3021). These bioconjugates are typically produced in engineered strains of *E. coli*, and the conjugation occurs in the bacterial periplasm. The glycan substrate of the OTase is linked to a membrane-bound lipid carrier, such as undecylisoprene pyrophosphate (UNDPP). O-linked OTases catalyze the transfer of oligosaccharides or polysaccharides linked to UNDPP to serine or threonine side-chain hydroxyl groups in conserved protein motifs called sequons (Knoot, C. et al. (2021) *Glycobiology*, Vol. 31, No. 9, September 2021, pp. 1192–1203; Knoot, C. et al. (2023) *Glycobiology*, Vol. 33, No. 1, January 2023, pp. 57–74). By co-expressing the desired glycan biosynthesis gene clusters, carrier proteins, and OTase in a chromosomal and / or episodic manner in a bacterial host, bioconjugate vaccines can be generated in a "one-pot" bioreactor, followed by downstream purification (Harding, C. and Feldman, M. (2019) Glycobiology, Vol. 29, No. 7, July 2019, pp. 519–529).

[0006] Several bioconjugate vaccines are currently undergoing clinical trials, and each vaccine consists of bacterial glycans linked to periplasmic proteins, primarily Pseudomonas aeruginosa exotoxin A (EPA), Haemophilus protein D, or CRM197 (a modified diphtheria toxin) (Sorieul, C. et al. (2023) Expert Review of Vaccines, 22:1, 1055-1078). Alternatives to using such carrier proteins are protein nanoparticles (NPs) or virus-like particles (VLPs): symmetrical, self-assembling protein “cages” that are derivatives of natural viral capsids or rationally engineered protein assemblies (Nguyen, B. and Tolia, N. (2021) npj Vaccines 6, 70; Bruun, T. et al. (2018) ACS Nano 2018, 12, 9, 8855–8866; Cohen et al. (2021) PLoS ONE 16(3):e0247963). Expression of NP / VLP monomeric proteins in the bacterial cytoplasm induces spontaneous assembly into mega-dalton-sized particles that can be purified from cellular biomass (Cohen et al. (2021) PLoS ONE 16(3):e0247963). NP / VLP-based therapeutics have been shown to improve immune responses, partly through increased antibody affinity caused by the larger immunogen particle size (Nguyen, B. and Tolia, N. (2021) npj Vaccines 6, 70). Two examples of VLP / NP are AP205 and mi3. AP205 is derived from the CP3 coat protein of the RNA phage AP205 (Brune, K. et al. (2016) Sci Rep 6, 19234). AP205 VLPs assemble into 120-mers with a diameter of approximately 20 nm (Cohen et al. (2021) PLoS ONE 16(3):e0247963). mi3 is a porous dodecahedral 60 polymer derived from computer-designed NPs with a diameter of 20–30 nm (Bruun, T. et al. (2018) ACS Nano 2018, 12, 9, 8855–8866)(Hsia, Y. et al. (2016) Nature Vol. 535, pp. 136–139).

[0007] The SpyTag / SpyCatcher system originates from the immunoglobulin-like collagen-adhesive domain (CnaB2), which is derived from the fibronectin-binding protein FbaB2 of *Streptococcus pyogenes* (Zakeri, B. et al. (2012)). The CnaB2 domain spontaneously forms an intraprotein heteropeptide bond between a lysine residue at position 31 and an aspartic acid residue at position 117; specifically, the aprotonated amine at Lys31 acts as a nucleophile attacking the carbonyl carbon of Asp117, an attack catalyzed by glutamate at position 77 (Zakeri, B. et al. (2012)). This heteropeptide reaction occurs spontaneously and appears to be characteristic of some members of bacterial proteins belonging to the prealbumin-like folding domain to which CnaB2 belongs. To adapt it into a tool for creating covalent fusions between two separate peptides, the CnaB2 domain was split, dividing CnaB2 into: (1) a peptide containing a C-terminal β-chain with reactive Asp117, defined as SpyTag; and (2) a protein-binding partner derived from the remaining CnaB2 peptide, defined as SpyCatcher (Zakeri, B. et al. (2012)). The system was named SpyTag and SpyCatcher to indicate the bacterial origin of the CnaB2 fragment (Streptococcus pyogenes). Subsequent forms of the SpyTag and SpyCatcher system, named SpyTag003 and SpyCatcher003, were created using phage display technology and subsequent rational engineering methods, achieving a reaction rate of 5.5 × 10⁵ M⁻¹ s⁻¹. The SpyTag003 / SpyCatcher003 reaction rate is approximately 400 times faster than the original SpyTag / SpyCatcher system (Keeble, AH et al. (2019)). The SpyTag / SpyCatcher system has been widely used to achieve covalent linkage of two separate peptides by forming isopeptide bonds, one containing a SpyTag and the other containing a SpyCatcher. The SpyTag / SpyCatcher system has been applied to a range of biological applications, all of which aim to covalently link SpyTag-containing target peptides to different proteins or materials containing SpyCatchers. These applications include, but are not limited to, anchoring peptides to the surfaces of different solid organic and inorganic materials, linking peptides to different polymerization architectures such as nanoparticles, virus-like particles or adenovirus vectors, and directly linking peptides to the surface of intact cells (Keeble, AH and Howarth, M. (2020); Brune, KD et al. (2016); Bruun, TUJ, Andersson, AC, Draper, SJ and Howarth, M. (2018)).

[0008] There is still a need to create new and more effective bioconjugate vaccines, including flexible methods for generating diverse immunogenic compositions. Summary of the Invention

[0009] This document provides a fusion protein comprising: (i) a glycosylated fragment, and (ii) a first polypeptide tag, wherein the first polypeptide tag is spontaneously capable of forming an isopeptide bond with a second polypeptide tag binding partner. In some embodiments, the fusion protein is a glycoconjugate comprising a sugar covalently linked to the fusion protein via the glycosylated fragment, and optionally possessing immunogenicity. Representative examples of the first polypeptide tag include SpyTag (SEQ ID NO:416), SpyTag002 (SEQ ID NO:417), SpyTag003 (SEQ ID NO:418), or DogTag (SEQ ID NO:419).

[0010] In some embodiments, the glycosylation fragment is a ComP glycosylation fragment or a variant thereof as described herein.

[0011] In some embodiments, the glycosylated fragment is a TfpM-associated fimbriae glycosylated fragment or a variant thereof as described herein.

[0012] In one embodiment, the glycosylation fragment is the PilE glycosylation fragment or a variant thereof as described herein.

[0013] In some embodiments, the glycosylation fragment is the PglB glycosylation fragment as described herein or a variant thereof.

[0014] In some embodiments, the glycosylation fragment is the PilA glycosylation fragment as described herein or a variant thereof.

[0015] In some embodiments, the glycosylation fragment is the STT3 glycosylation fragment or a variant thereof as described herein.

[0016] In some embodiments, the glycosylated fragment is an N-linked glycosyltransferase glycosylation fragment.

[0017] In some embodiments, the glycosylated fragment is an O-linked glycosyltransferase glycosylation fragment.

[0018] In some embodiments, the glycosylated fragment is a glycosylated fragment of the PilA_Pa5196-associated fimbriae protein as described herein, or a variant thereof.

[0019] In some embodiments, the fusion protein comprises a carrier protein, optionally selected from the group consisting of: Escherichia coli maltose-binding protein, Pseudomonas aeruginosa exotoxin A (EPA), Pseudomonas aeruginosa PcrV, CRM197, Haemophilus influenzae protein D, cholera toxin B subunit or tetanus toxin, and fragments of any of these.

[0020] This document also provides a composition comprising a peptide pair comprising a first peptide and a second peptide, wherein the first peptide is a fusion protein of this disclosure, wherein the second peptide comprises a second peptide tag binding spouse relative to a first peptide tag of the first peptide, and wherein the first peptide is linked to the second peptide via an isopeptide bond between the first peptide tag and the second peptide tag. In some embodiments, the second peptide comprises a monomeric peptide capable of spontaneously polymerizing / self-assembling into a higher-order multimeric structure; and optionally, the higher-order multimeric structure is an icosahedral or dodecahedral particle (e.g., similar to a nanocage), a virus-like particle, or an adenovirus vector.

[0021] In some embodiments, the second polypeptide tag is SpyCatcher (SEQ ID NO:420), SpyCatcher002 (SEQ ID NO:421), SpyCatcher003 (SEQ ID NO:422), or DogCatcher (SEQ ID NO:423).

[0022] In some embodiments, the first polypeptide is a bioconjugate comprising a sugar covalently linked to a glycosylated fragment of the first polypeptide; optionally, the composition is immunogenic.

[0023] A complex is also provided comprising two or more of the polypeptide pairs disclosed herein. In some embodiments, the complex is a self-assembled multimeric higher-order structure. In some embodiments, the self-assembled multimeric higher-order structure is an icosahedral or dodecahedral particle (e.g., similar to a nanocage), a virus-like particle, or an adenovirus vector.

[0024] A method for preparing the polypeptide pair disclosed herein is also provided, the method comprising contacting the first polypeptide and the second polypeptide under conditions that allow the first polypeptide tag to spontaneously form an isopeptide bond with the binding coupler of the second polypeptide tag.

[0025] A method for preparing the complex disclosed herein is also provided, the method comprising: (i) forming a self-assembled multimeric higher-order structure of a second polypeptide, and then contacting the first polypeptide and the second polypeptide under conditions that allow the first polypeptide tag to spontaneously form isopeptide bonds with the second polypeptide tag; or (ii) contacting the first polypeptide and the second polypeptide under conditions that allow the first polypeptide tag to spontaneously form isopeptide bonds with the second polypeptide tag, and then forming a self-assembled multimeric higher-order structure of the second polypeptide. In some embodiments, the ComP glycosylation fragment is glycosylated by PglS OTase; the TfpM-associated fimbriae glycosylation fragment is glycosylated by TfpM OTase, optionally wherein the ComP glycosylation fragment is glycosylated by PglS OTase and the TfpM-associated fimbriae glycosylation fragment is glycosylated by TfpM OTase; the PilE glycosylation fragment is glycosylated by PglL OTase; the PglB glycosylation fragment is glycosylated by PglB OTase; the PilA glycosylation fragment is glycosylated by TfpO or PilO OTase; the STT3 glycosylation fragment is glycosylated by STT3 catalytic subunit; the PilA_Pa5196-associated fimbriae glycosylation fragment is glycosylated by TfpW glycosyltransferase; the N-linked glycosyltransferase glycosylation fragment is derived from *Actinobacillus pleuropneumoniae*; or from *Haemophilus influenzae*. (influenzae); or N-linked glycosyltransferase glycosylation from Yersinia enterocolitica; and / or O-linked glycosyltransferase glycosylation fragments are glycosylated by GtfA / GtfB glycosyltransferases.

[0026] This document also provides a method for inducing an immune response in a subject by administering to the subject an effective amount of any composition, complex, and / or conjugate vaccine of any of the present disclosure or a composition, complex, and / or conjugate vaccine of the present disclosure for inducing an immune response in a subject. Attached Figure Description

[0027] Figure 1A schematic overview of the production of glycoprotein polypeptides using glycosidic bond-forming enzymes and the subsequent formation of isopeptide bonds with a second protein containing a polypeptide tag. The first polypeptide (protein 1) contains: (i) a SpyTag, which will spontaneously form an isopeptide bond with the polypeptide tag (SpyCatcher) of the second partner polypeptide (protein 2), and (ii) a glycosylated fragment (sequence), which is recognized by a specific enzyme that forms a glycosidic bond by covalently transferring sugar to the glycosylated fragment (sequence). In this example, protein 2 containing the polypeptide tag (SpyCatcher) can also self-assemble into higher-order structures, such as icosahedral or dodecahedral nanoparticles resembling nanocages, virus-like particles (VLPs), or adenovirus vectors.

[0028] Figure 2 The purified MBP-SpyTag-v1 E. coli O16 O-antigen bioconjugate was generated using a Coomassie-stained SDS-PAGE denaturing gel via a TfpM oligosaccharide transferase system, mi3-SpyCatcher, and isopeptide-bonded MBP-SpyTag-v1 E. coli O16:mi3-SpyCatcher. The block diagram illustrates the SpyTag protein (corresponding to...) Figure 1 Protein 1) and SpyCatcher protein (corresponding to Figure 1 Schematic structures of proteins 2) in the diagram, designed to be compatible with enzymes that form glycosidic bonds. Abbreviations: MBP, maltose-binding protein; MBPsp, E. coli maltose-binding protein secretion (Sec) signal peptide; Pil20, a 20-amino acid fimbriae sequence glycosylated by TfpM oligosaccharide transferase; mi3, mi3 nanoparticle monomer; 6xHis, hexahistidine tag. For SDS-PAGE analysis, SpyCatcher and SpyTag proteins were reacted alone or in a 1:1 or 2:1 ratio (based on protein concentration) in Tris-buffered saline at room temperature for 2 hours. The formation of isopeptide bonds was terminated by adding Laemmli buffer and subsequently heating the sample at 100°C for 10 min before gel loading. Protein masses associated with proteins and glycoproteins are indicated above each channel. Channel A) Protein gradient, with standard masses (in kDa) indicated on the left. Channel B) Purified MBP-SpyTag-v1-O16 bioconjugate. Channel C) Purified mi3-Spycatcher. Channel D) 1:1 reaction mixture of mi3-SpyCatcher and MBP-SpyTag-v1-O16. Channel E) 2:1 reaction mixture of mi3-SpyCatcher and MBP-SpyTag-v1-O16.

[0029] Figure 3Western blot analysis was performed using a TfpM oligosaccharide transferase system, mi3-SpyCatcher, and isopeptide-bonded MBP-SpyTag-v1 E. coli O16:mi3-SpyCatcher to generate purified MBP-SpyTag-v1 E. coli O16 O-antigen bioconjugates. SpyTag and SpyCatcher proteins were reacted alone or in a 1:1 or 2:1 ratio (based on protein concentration) in Tris-buffered saline for 2 hours prior to Western blot analysis. Protein masses associated with proteins and glycoproteins are labeled above each channel. Western blots were probed with an anti-His tag (α-His, top panel) and E. coli O16 O-antigen antiserum (α-O16, middle panel). Combined images are shown in the lower panel. Channel A) Protein gradient, with standard masses (in kDa) labeled on the left. Channel B) Purified MBP-SpyCatcher-v1-O16 bioconjugate. Channel C) Purified mi3-Spycatcher. Channel D) 1:1 reaction mixture of mi3-SpyTag and MBP-SpyTag-v1-O16. Channel E) 2:1 reaction mixture of mi3-SpyCatcher and MBP-SpyTag-v1-O16.

[0030] Figure 4 Size exclusion chromatogram of mi3-SpyCatcher and isopeptide-bonded MBP-SpyTag-v1-O16:mi3-SpyCatcher (separated using a Sephacryl S-400HR 16 / 600 column). The figure shows the UV absorbance traces of 250 μg mi3-SpyCatcher (solid line) and 250 μg mi3-SpyCatcher after reacting with 250 μg MBP-SpyTag-v1-O16 in Tris-buffered saline at room temperature for 2 hours (dashed line).

[0031] Figure 5The purified MBP-SpyTag-v2 *E. coli* O-antigen bioconjugate, generated using the PglS oligosaccharide transferase system, mi3-SpyCatcher, and isopeptide-bonded MBP-SpyTag-v2 *E. coli* O16:mi3-SpyCatcher, was analyzed by Coomassie staining SDS-PAGE denaturing gel. A block diagram illustrates the structural schematics of the SpyTag and SpyCatcher proteins used in this experiment. Abbreviations: MBP, maltose-binding protein; MBPsp, *E. coli* maltose-binding protein secretion (Sec) signal peptide; ComP sequence, a 23-amino acid ComP-derived sequence glycosylated by PglS oligosaccharide transferase; mi3, mi3 nanoparticle monomer; 6xHis, hexahistidine tag. For SDS-PAGE analysis, SpyTag and SpyCatcher proteins were reacted alone or in a 1:1 or 2:1 ratio (based on protein concentration) in Tris-buffered saline at room temperature for 2 hours. The formation of isopeptide bonds was terminated by adding Laemmli buffer and then heating the sample at 100°C for 10 min before gel loading. Samples from each reaction were taken for SDS-PAGE analysis. Protein masses associated with proteins and glycoproteins are indicated above each channel. Channel A) Protein gradient, with standard masses (in kDa) indicated on the left. Channel B) Purified MBP-SpyTag-v2-O16 bioconjugate. Channel C) Purified mi3-SpyCatcher. Channel D) 1:1 reaction mixture of mi3-SpyCatcher and MBP-SpyTag-v2-O16. Channel E) 2:1 reaction mixture of mi3-SpyCatcher and MBP-SpyTag-v2-O16.

[0032] Figure 6Western blot analysis was performed using the PglS oligosaccharide transferase system, mi3-SpyCatcher, and isopeptide-bonded MBP-SpyTag-v2 E. coli O16:mi3-SpyCatcher to generate purified MBP-SpyTag-v2 E. coli O16 O-antigen bioconjugates. Prior to Western blot analysis, the SpyTag and SpyCatcher fusion proteins were reacted alone or in a 1:1 or 2:1 ratio (based on protein concentration) in Tris-buffered saline for 2 hours. Western blots were probed with an anti-His tag (α-His, top panel) and E. coli O16 O-antigen antiserum (α-O16, middle panel). The combined images are shown in the lower panel. Channel A) Protein gradient, with standard mass (kDa) labeled on the left. Channel B) Purified MBP-SpyTag-v2-O16 bioconjugate. Channel C) Purified mi3-SpyCatcher bioconjugate. Channel D) A 1:1 reaction mixture of mi3-SpyCatcher and MBP-SpyTag-v2-O16. Channel E) A 2:1 reaction mixture of mi3-SpyCatcher and MBP-SpyTag-v2-O16.

[0033] Figure 7The purified MBP-SpyTag-v1 E. coli O16 O-antigen bioconjugate, generated using the TfpM oligosaccharide transferase system, AP205-SpyCatcher, and isopeptide-bonded MBP-SpyTag-v1 E. coli O16:AP205-SpyCatcher, was stained with Coomassie staining SDS-PAGE denaturing gel. A block diagram illustrates the structural schematics of the SpyTag and SpyCatcher proteins used in this experiment. Abbreviations: MBP, maltose-binding protein; MBPsp, secretion (Sec) signal peptide of E. coli maltose-binding protein; Pil20, a 20-amino acid fimbriae sequence glycosylated by TfpM oligosaccharide transferase; AP205, AP205 virus-like protein capsid protein monomer; 6xHis, hexahistidine tag. For SDS-PAGE analysis, SpyCatcher and SpyTag fusion proteins were reacted individually or in a 1:1 or 2:1 ratio (based on protein concentration) in Tris-buffered saline at room temperature for 2 hours before gel loading. Protein masses associated with proteins and glycoproteins are indicated above each channel. Channel A) Protein gradient, with standard masses (in kDa) indicated on the left. Channel B) Purified MBP-SpyTag-v1-O16 bioconjugate. Channel C) Purified AP205-SpyCatcher. Channel D) 1:1 reaction mixture of AP205-SpyCatcher and MBP-SpyTag-v1-O16. Channel E) 2:1 reaction mixture of AP205-SpyCatcher and MBP-SpyTag-v1-O16.

[0034] Figure 8The purified MBP-SpyTag-v2 *E. coli* O-antigen bioconjugate, generated using the PglS oligosaccharide transferase system, AP205-SpyCatcher, and isopeptide-bonded MBP-SpyTag-v2 *E. coli* O16:AP205-SpyCatcher, was stained with Coomassie staining SDS-PAGE denaturing gel. A block diagram illustrates the structural schematics of the SpytTag and SpyCatcher proteins used in this experiment. Abbreviations: MBP, maltose-binding protein; MBPsp, *E. coli* maltose-binding protein secretion (Sec) signal peptide; ComP sequence, a 23-amino acid ComP-derived sequence glycosylated by PglS OTase; AP205, AP205 virus-like protein capsid protein monomer; 6xHis, hexahistidine tag. For SDS-PAGE analysis, SpyTag and SpyCatcher fusion proteins were reacted alone or in a 1:1 or 2:1 ratio (based on protein concentration) in Tris-buffered saline at room temperature for 2 hours. The formation of isopeptide bonds was terminated by adding Laemmli buffer and subsequently heating the sample at 100°C for 10 min before gel loading. Samples from each reaction were taken for SDS-PAGE analysis. Protein masses associated with proteins and glycoproteins are indicated above each channel. Channel A) Protein gradient, with standard masses (in kDa) indicated on the left. Channel B) Purified MBP-Spycatcher-v2-O16 bioconjugate. Channel C) Purified AP205-Spycatcher. Channel D) 1:1 reaction mixture of AP205-Spytag and MBP-Spytag-v2-O16. Channel E) 2:1 reaction mixture of AP205-Spycatcher and MBP-Spytag-v2-O16.

[0035] Figure 9Coomassie staining SDS-PAGE denaturing gels were used to generate three purified EPA-SpyTag E. coli O-antigen bioconjugates generated using the PglS oligosaccharide transferase system, mi3-SpyCatcher, and EPA-Spytag E. coli O16:mi3-SpyCatcher via isopeptide bonds. A block diagram illustrates the schematic structure of the EPA SpyTag proteins, which are designed to be compatible with enzymes that form glycosidic bonds. Abbreviations: EPA, Pseudomonas aeruginosa exotoxin A; MBPsp, E. coli maltose-binding protein secretion (Sec) signal peptide; ComP sequence, 23-amino acid fimbriae sequence glycosylated by PglS oligosaccharide transferase; mi3, mi3 nanoparticle monomer; 6xHis, hexahistidine tag. Isopeptide bond formation reactions were performed using purified EPA-Spycatcher-O16 bioconjugates and purified mi3-Spycatcher. Channel A) Protein gradient, with standard masses (in kDa) labeled on the left. Channel B) Purified mi3-Spycatcher. Channel C) Purified EPA-Spytag-v1-O16. Channel D) 1:1 reaction mixture of mi3-Spycatcher and EPA-SpyTag-v1-O16. Channel E) 1:2 reaction mixture of mi3-Spycatcher and EPA-SpyTag-v1-O16. Channel F) Purified EPA-Spytag-v2-O16. Channel G) 1:1 reaction mixture of mi3-Spycatcher and EPA-SpyTag-v2-O16. Channel H) 1:2 reaction mixture of mi3-Spycatcher and EPA-SpyTag-v2-O16. Channel I) Purified EPA-Spytag-v3-O16. Channel J) 1:1 reaction mixture of mi3-Spycatcher and EPA-SpyTag-v3-O16. A 1:2 reaction mixture of channel K)mi3-Spycatcher and EPA-SpyTag-v3-O16. No protein binding via isopeptide bonds was observed when using either EPA-Spytag-v2-O16 or EPA-Spytag-v3-O16 with mi3-Spycatcher.

[0036] Figure 10Western blot analysis of purified, unglycosylated EPA-SpyTag-v1 protein linker variants, mi3-SpyCatcher, and isopeptide-bonded EPA-SpyTag E. coli:mi3-SpyCatcher. Each protein variant has a distinct amino acid linker between Spytag003 and EPA. An isopeptide bond formation reaction was performed using EPA-Spytag protein from E. coli periplasmic extract and purified mi3-Spycatcher. Western blot analysis was performed using an anti-His-tag antibody. Channel A) Protein gradient, with standard mass (kDa) indicated on the left. Channel B) Purified EPA-Spytag-v1 with linker L1 (SGG). Channel C) 1:1 reaction mixture of mi3-Spycatcher and EPA-Spytag-v1 with linker L1. Channel D) Purified EPA-Spytag-v1 with linker L2 (SEQ ID NO:430). Channel E) A 1:1 reaction mixture of mi3-Spycatcher and EPA-Spytag-v1 with adapter L2. Channel F) Purified EPA-Spytag-v1 with adapter L3 (SEQ ID NO: 431). Channel G) A 1:1 reaction mixture of mi3-Spycatcher and EPA-Spytag-v1 with adapter L3. Channel H) Purified EPA-Spytag-v1 with adapter L4 (SEQ ID NO: 432). Channel I) A 1:1 reaction mixture of mi3-Spycatcher and EPA-Spytag-v1 with adapter L4. Channel J) Purified EPA-Spytag-v1 with adapter L5 (SEQ ID NO: 433). Channel K) A 1:1 reaction mixture of mi3-Spycatcher and EPA-Spytag-v1 with adapter L5. Channel L) Purified EPA-Spytag-v1 with adapter L6 (SEQ ID NO: 434). Channel M) A 1:1 reaction mixture of mi3-Spycatcher and EPA-Spytag-v1 with adapter L6. Channel N) Purified EPA-Spytag-v1 with adapter L7 (SEQ ID NO:435). Channel O) A 1:1 reaction mixture of mi3-Spycatcher and EPA-Spytag-v1 with adapter L7. Channel P) Purified EPA-Spytag-v1 with adapter L8 (SEQ ID NO:436). Channel Q) A 1:1 reaction mixture of mi3-Spycatcher and EPA-Spytag-v1 with adapter L8. Channel R) Purified EPA-Spytag-v1 with adapter L9 (SEQ ID NO:437).Channel S) a 1:1 reaction mixture of mi3-Spycatcher and EPA-Spytag-v1 with linker L9. Channel T) purified EPA-Spytag-v1 with linker L10 (SEQ ID NO:438). Channel U) a 1:1 reaction mixture of mi3-Spycatcher and EPA-Spytag-v1 with linker L10. Some variants were poorly expressed and / or unable to form isopeptide bonds with mi3-Spycatcher.

[0037] Figure 11 Western blot of unglycosylated EPA-Spytag-v1 (EAAAKEAAAK; SEQ ID NO:435) with L7 linker, and the time progression of mi3-Spycatcher isopeptide formation reaction. Isopeptide formation was performed using EPA-Spytag-v1 with L7 linker from *E. coli* periplasmic extract and purified mi3-Spycatcher at a 1:1 ratio. The left panel shows the Western blot detected with an anti-His-tagged antibody. The middle panel shows the Western blot detected with an anti-EPA antibody. The right panel is a merged image of the two channels. Channel A) Protein gradient, with standard mass (kDa) labeled on the left. Channel B) mi3-Spycatcher only. Channel C) EPA-Spytag-v1 with L7 linker only. Channel D) Isopeptide reaction between mi3-Spycatcher and EPA-Spytag-v1 with L7 linker after 0.5 hours. Channel E) Reaction after 1 hour. Channel F) Reaction after 2 hours. The reaction of channel G after 5 hours. The reaction of channel H after 24 hours.

[0038] Figures 12A-E Figure 12A shows EPA-ComP 110264 A schematic diagram of the fusion protein, in which the ComP glycosylated fragment is fused at the C-terminus of the fusion protein. “ssDsbA” corresponds to the DsbA Sec secretion signal. GGGS (SEQ ID NO:382) is EPA and ComP. 110264 Flexible linkers between fragments. Figure 12B shows the different amino acid sequences of the ComP glycosylated fragment fused to the C-terminus of the EPA fusion protein. The bolded underlined serine residues in each sequence are associated with ComP. 110264The conserved serine residue 82 corresponds to and is the site of glycosylation. Bold, underlined cysteine ​​residues corresponding to Cys71 and Cys93 are also highlighted. (C2, SEQ ID NO:383; D2; SEQ ID NO:384; E2, SEQ ID NO:385; F2, SEQ ID NO:386; G2, SEQ ID NO:387; H2, SEQ ID NO:388; A3, SEQ ID NO:389; B3, SEQ ID NO:390; C3, SEQ ID NO:391; D3, SEQ ID NO:392; E3, SEQ ID NO:393; F3, SEQ ID NO:394; and C1, SEQ ID NO:395). Figures 12C, 12D, and 12E show the glycosylation sites from the expression of PglS, CPS8 glycans, and EPA-ComP. 110264 Western blot analysis of periplasmic extracts of variant *E. coli* SDB1. Each channel of the Western blot map corresponds to a strain of SDB1 expressing a different EPA-ComP variant, where the ComP glycosylated fragment corresponds to the sequence shown in Figure 12B. Figure 12C shows the protein reacting with anti-EPA antiserum. Figure 12D shows the protein reacting with anti-His antiserum. Figure 12E shows a combined Western blot image of Figures 12C and 12D. An equal amount of OD-based... 600 Periplasmic extract. On the right side of Figure 12C-E, g0 represents unglycosylated EPA-ComP. 110264 And g n EPA-ComPs glycosylated with different numbers of CPS8 repeating units 110264 Protein quality markers (in kDa) are shown on the left side of Figure 12C-E.

[0039] Figures 13A-D Figure 13A shows the CRM. 197 -ComP C1 A schematic diagram of the fusion protein. “ssFlgI” corresponds to the FlugI SRP secretion signal. GGGS (SEQ ID NO:382) is a CRM. 197 With ComP C1 The flexible connector between them. Figures 13B, 13C, and 13D show the purified CRM. 197 -ComP C1 Western blot analysis of CPS8 glycoconjugates. Figure 13B shows the proteins reacting with anti-CPS8 antiserum. Figure 13C shows the proteins reacting with anti-CRM. 197 Proteins that resist serum reactions. Figure 2Figure D shows the combined Western blot images of Figures 13B and 13C. CRM in the sample treated with proteinase K (PK) 197 The loss of CPS8 signal proves that pneumococcal serotype 8 signal is related to CRM. 197 This is related to, rather than caused by, contamination from free polysaccharides or lipid-linked polysaccharide precursors. Protein quality markers (in kDa) are shown on the left side of Figure 13B-D.

[0040] Figures 14A, B Figure 14A shows the C-terminal and N-terminal CRMs containing the C1 ComP glycosylation fragment. 197 Schematic diagram of the variant. Figure 14B shows the expression of CRM with or without (+) PglS. 197 -ComP C1 or ComP C1 Western blot analysis of periplasmic extracts of *E. coli* SDB1 containing CRM197 and CPS8 glycans. Equal amounts of OD-based samples were loaded into each channel. 600 Periplasmic extract. Protein quality markers (in kDa) are shown on the left. GGGS (SEQ ID NO:382).

[0041] Figures 15A-E Figure 15A shows a schematic diagram of the EPA fusion protein containing ComP glycosylation fragments integrated within the EPA amino acid sequence. Figure 15B shows the amino acid sequences of two iGTComP glycosylation fragments inserted between EPA residues Ala489 and Arg489. These sequences have two terminal cysteine ​​residues (“iGTComP”). CC ";SEQ ID NO:230) or serine (“iG SS ";SEQ ID NO:231). Figures 15C and 15D show the effects of (+) or (-) PglS on the expression of CPS8 glycans and EPA. iGTcc or EPA iGTss Western blot of periplasmic extracts of *E. coli* SDB1. Figure 15C shows the protein reacting with anti-EPA antiserum. Figure 15D shows the protein reacting with anti-His antiserum. Figure 15E shows a combined Western blot image of Figures 15C and 15D. Equal amounts of OD-based samples were loaded into each channel. 600 Periplasmic extract. Protein quality markers (in kDa) are shown on the left side of the figure.

[0042] Figures 16A-DFigure 16A shows a schematic diagram (from top to bottom, SEQ ID NO: 6-28) of the EPA constructs containing the ComP glycosylation fragment used in these experiments. The iGT is truncated to twenty-two to five amino acids. CC ComP glycosylation fragment variants are inserted into the EPA coding sequence between Ala489 and Arg489. Figure 16B shows the amino acid sequences of 22 truncated iGT ComP glycosylation fragments, with names labeled on the left. Underlined, bold serine residues are glycosylation sites. (iGTcc SEQ ID NO:230; Δ0-1SEQ ID NO:232; Δ1-0SEQ ID NO:243; Δ1-2SEQ ID NO:245; Δ2-1SEQ ID NO:256; Δ2-3SEQ ID NO:258; Δ3-2SEQ ID NO:269; Δ3-4SEQ ID NO:271; Δ4-3SEQ ID NO:282; Δ4-5SEQ ID NO:284; Δ5-4SEQ ID NO:295; Δ5-6SEQ ID NO:297; Δ6-5SEQ ID NO:308; Δ6-6SEQ ID NO:309; Δ6-7SEQ ID NO:310; Δ7-6SEQ ID NO:321; NO:323; Δ8-7SEQ ID NO:334; Δ8-8SEQ ID NO:335; Δ8-9SEQ ID NO:336; Δ9-8SEQ ID NO:346; Δ9-9SEQ ID NO:347). Figure 16C shows the effects of EPA expressing PglS, CPS8 and containing truncated ComP glycosylation fragments. iGT Western blot analysis of periplasmic extracts of the fusion protein from *E. coli* SDB1. Each channel in the Western blot map corresponds to the expression of different EPAs. iGT The strain SDB1 contains a fusion protein containing a truncated ComP glycosylated fragment, which corresponds to the sequence shown in Figure 16B. Figure 16C shows the protein reacting with anti-EPA antiserum, detected using an anti-EPA antibody. EPA is shown. iGTcc For comparison. The “EPA” channel corresponds to EPA lacking any ComP-derived sequences and serves as a negative control. Equal amounts of OD-based samples were loaded into each channel. 600 Periplasmic extract. Figure 16D shows the same protein blot as above, but with increased anti-EPA signal brightness to show low-level glycosylation of minimal ComP glycosylated fragments.

[0043] Figures 17A, B, CFigure 17 shows Western blot analysis of the EPA fusion protein purified by Ni affinity chromatography, which contains an iGTΔ6-6ComP glycosylated fragment integrated between the residues Ala489 and Arg490 of EPA. The fusion protein was purified from SDB1 cells expressing CPS8 glycans in the presence (+) or absence (-) of PglS. Figure 17A shows the protein reacting with anti-His antiserum. Figure 17B shows the protein reacting with anti-CPS8 antiserum. Figure 17C shows a combined graph of Figures 17A and 17B. Protein quality markers (in kDa) are shown on the left side of Figures 17A-C.

[0044] Figures 18A and 18B. Figure 18A shows a schematic diagram of the EPA fusion protein containing an iGTΔ3-4ComP glycosylated fragment integrated between residues Glu548 and Gly549 of EPA. The amino acid sequence of iGTΔ3-4 is listed below the schematic diagram (SEQ ID NO: 271). Figure 18B shows Western blot analysis of periplasmic extracts from *E. coli* SDB1 expressing the PglS, CPS8, and EPA fusion protein containing an iGTΔ3-4ComP glycosylated fragment integrated between residues Glu548 and Gly549. The figure shows the protein reacting with anti-EPA antiserum, which was detected using an anti-EPA antibody.

[0045] Figures 19A, B, C Figure 19 shows Western blot analysis of the EPA fusion protein purified by Ni affinity chromatography, which contains an iGTΔ3-4ComP glycosylated fragment integrated between the residues Glu548 and Gly549 of EPA. The fusion protein was purified from SDB cells expressing CPS8 glycans in the presence (+) or absence (-) of PglS. Figure 19A shows the protein reacting with anti-His antiserum. Figure 19B shows the protein reacting with anti-CPS8 antiserum. Figure 19C shows a combined graph of Figures 19A and 19B. Protein quality markers (in kDa) are shown on the left side of Figures 19A-C.

[0046] Figure 20 . Figure 20 The amino acid sequences of ComP orthologs are listed. Predicted glycosylation sites are shown in bold.

[0047] Figure 21 . Figure 21 The amino acid sequences of ComPΔ28 orthologs are listed, among which those are related to ComP. ADP1 The amino acids corresponding to the 28 N-terminal amino acids in AAC45886.1 have been removed. The predicted glycosylation sites are shown in bold.

[0048] Figure 22 . Figure 22 The image shows an alignment of the ComP sequence containing the serine (S) residue (boxed out), which corresponds to ComP. 110264 The serine residue at position 82 of (SEQ ID NO:201) corresponds to the same residue as ComP. ADP1 This corresponds to the serine residue at position 84 of (SEQ ID NO:202).

[0049] Figures 23A-D. Figure 23 shows the characterization of 13 TfpM orthologs from species in the Moraxellaceae family. Figure 23A) shows a branching diagram of 20 TfpM orthologs from typical TfpO, PglL, and PglS of Pseudomonas, Neisseria, and Acinetobacter, respectively. Branching confidence is indicated in red. OTase / fimbrine pairs marked with an asterisk have been cloned and tested in bioconjugation experiments. Figure 23B) shows a schematic diagram of the EPA-fimbrine fusion protein and TfpM construct design. Colored arrows indicate genes. Gene expression is driven by an IPTG-inducible tac promoter with a lacO operon (tac1O). The rrnB T2 terminator is marked with a black hairpin structure. Figures 23C) and 23D) show anti-EPA protein blots of whole-cell *E. coli* extracts expressing different EPA-fibrin vector open reading frames and the tfpM gene. Image D shows the same image as image C, but with higher exposure. Whole-cell extracts loaded in each channel were analyzed by OD... 600 Normalization was performed. H286A represents an OTase activity site mutant directed at the TfpM site in *Morala osloensis*. “g0” represents unglycosylated EPA-fimbriae protein, and “g…” n "" indicates CPS8 glycosylated EPA-fimbriae protein. The reference protein mass, in kDa, is marked on the left side of the protein blot.

[0050] Figure 24 . Figure 24 Phylogenetic diagrams of orthologs of TfpM, PilO, PglL, and PglS are shown, with relative distances indicated. The phylogenetic trees were generated using the phylogeny.fr server (Website address: phylogeny.fr / ), which uses MUSCLE, PhyML, and TreeDyn for sequence alignment, tree calculation, and image generation, respectively.

[0051] Figure 25 . Figure 25Phylogenetic diagrams of TfpM-related fimbriae-like proteins, selected PilA proteins from Neisseria and Pseudomonas, and ComP from Acinetobacter soli CIP 110264, showing relative distances. Red numbers indicate branch confidence. The phylogenetic trees were generated using the phylogeny.fr server (Website: phylogeny.fr / ), which performed sequence alignment, tree computation, and image generation using MUSCLE, PhyML, and TreeDyn, respectively.

[0052] Figure 26 . Figure 26 Multiple sequence alignments of selected bacterial O-linked oligosaccharide transferases are shown. The alignments were generated using Clustal Omega with default settings, available at ebi.ac.uk / Tools / msa / clustalo / . The selected bacteria are: *Neisseria meningitidis* MC58 PglL (SEQ ID NO:105); *Acinetobacter baylyi* ADP1 PglS (SEQ ID NO:106); *Pseudomonas aeruginosa* 1244 TfPO (SEQ ID NO:107); *Moraxella osloensis* 1202 TfpM (SEQ ID NO:56); and *Acinetobacter nosocomialis* M2 TfpO (SEQ ID NO:108).

[0053] Figure 27 . Figure 27 This study showed a whole-cell protein blot of anti-EPA, which examined EPA-Pil Mo Δ28 fusion and EPA-Pil Mo Δ28C terminal Thr 167 Glycosylation state of the mutant. All channels are relative to the same OD. 600 Normalization was performed. The reference protein mass, in kDa, was labeled next to the protein blot.

[0054] Figures 28A and 28B. Figure 28 shows C-terminated EPA-Pil modified with HexHexA. Mo Δ28 peptide 762 FLPANCRGT 770 Targeted MS / MS analysis of (SEQ ID NO:61). Figure 28A) EThcD fragmentation enables the localization of HexHexA glycosylation events to the terminal residue Thr. 770(Figure 28B) HCD fragmentation enabled the confirmation of the peptide sequence and the linkage of the disaccharide HexHexA through the Hex monosaccharide (by observing that multiple y ions were linked only to Hex residues).

[0055] Figures 29A-G Figure 29 shows TfpM Mo Multiple bacterial glycans can be transferred to EPA-Pil Mo Δ28 fusion protein. (Figure 29A) Using TfpM Mo The structures of repeating units of the five bacterial glycans tested. Connections between sugar monomers are indicated by parentheses. Glycan abbreviations used: CPS8, Streptococcus pneumoniae capsular polysaccharide type 8; GBSIII, Group B Streptococcus capsular polysaccharide type III; LT2, Salmonella enterica Group B serotype LT2 O-antigen; O16, Escherichia coli serotype O16 O-antigen; O2a, Klebsiella pneumoniae serotype O2a O-antigen; Unless otherwise specified, all sugars are in pyranose form. Abbreviations used: Glc, glucose; Gal, galactose; Galf, galactofuranose; Rha, rhamnose; GlcNAc, N-acetylglucosamine; Abe, abecosyl; NeuNAc, N-acetylneuraminic acid (sialic acid). Figures 29B-29F) show anti-glycan protein blots, which use partially purified TfpM. Mo Derived bioconjugates. Figure 29B) Anti-CPS8. Figure 29C) Anti-O16. Figure 29D) Anti-LT2. Figure 29E) Anti-O2a. Figure 29F) Anti-GBSIII. Figure 29G) Anti-EPA. In images B)-G), + / - markers indicate whether the sample was incubated with proteinase K (+) or not (-) before SDS-PAGE separation. Reference protein mass in kDa is marked next to the protein blot.

[0056] Figures 30A and 30B. Figure 30 shows TfpM. Mo It can glycosylate truncated EPA-fused fimbriae protein variants (as short as three amino acids). (Figure 30A) is used for interaction with TfpM. Mo Biologically conjugated EPA-fused Pil Mo The sequence of the fragment. Blue letters mark the C-terminal residues of EPA (i.e., EDLK; SEQ ID NO: 132). Underlined residues indicate glycine linkers placed between EPA and the fimbriae protein sequence. (Figure 30B) Expressing truncated fimbriae protein variants CPS8 and TfpM. Mo Anti-EPA protein blot analysis of whole-cell extracts. Calculated EPA-Pil MoThe Δ28 mass is 80.3 kDa, and the truncated variant has a mass range of 67.1 to 69.0 kDa. All channels are relative to the same OD. 600 Normalization was performed. Reference protein mass, in kDa, was labeled on the left side of the protein blot. "g0" indicates unglycosylated truncated EPA-fimbriae protein, and "g..." n "This indicates glycosylated EPA-pilin. Unglycosylated EPA-pilin..." Mo Δ28 is operating around 75 kDa. Pil 20 (SEQ IDNO:60). Pil 15 (SEQ ID NO:109). Pil 13 (SEQ ID NO:110). GGGG plus Pil 10 It is Pil 10L (SEQ ID NO:111). Pil 10 (SEQ ID NO:112). Pil7 (SEQ ID NO:113). Pil6 (SEQ ID NO:114). Pil5 (SEQ ID NO: 115). Pil4 (SEQ ID NO:116). Pil3 (SEQ ID NO:117). EDLK plus Pil2 (SEQ ID NO:118). EDLKGGGG plus Pil 20 (SEQ ID NO:122). EDLK plus Pil 15 (SEQ ID NO:123). EDLK plus Pil 13 (SEQ IDNO:124). EDLK plus Pil 10L (SEQ ID NO:125). EDLK plus Pil 10 (SEQ ID NO:126). EDLK plus Pil7 (SEQ ID NO:127). EDLK plus Pil6 (SEQ ID NO:128). EDLK plus Pil5 (SEQ ID NO:129). EDLK plus Pil4 (SEQ ID NO:130). EDLK plus Pil3 (SEQ ID NO:131).

[0057] Figure 31 . Figure 31Multiple sequence alignments of selected fimbriae proteins are shown. The accession numbers for these proteins are given in the text. The alignments were generated using Clustal Omega with default settings, available at ebi.ac.uk / Tools / msa / clustalo / . *Pseudomonas aeruginosa* 1244 PilA (SEQ ID NO:119). *Neisseria meningitidis* M2 PilA (SEQ ID NO:120). *Acinetobacter junii* 65 fimbriae protein (SEQ ID NO:97). *A_CIP102143_fimbriae protein (SEQ ID NO:88). *A_CIP102637_fimbriae protein (SEQ ID NO:100). *A_YZS-X1-1_fimbriae protein (SEQ ID NO:98). *Acinetobacter agrobacterium* 110264_ComP (SEQ ID NO:121). A_YH01026_fimbrine (SEQ ID NO:87). Moraxella osloi_1202_fimbrine (SEQ ID NO:57). Acinetobacter juncus_TUM15069_fimbrine (SEQ ID NO:84).

[0058] Figures 32A-F Figure 32 shows the purified TfpM Mo The derived GBSIII bioconjugate elicited a robust IgG immune response in mice. Figure 32A) Western blot of the purified GBSIII-291 bioconjugate (anti-EPA channel). Figure 32B) Anti-GBSIII. Figure 32C) Combined image of A and B. Figure 32D) Coomassie staining of the purified GBSIII-291 bioconjugate. Figure 32E) MS1 spectrum of the intact purified GBSIII-291 bioconjugate. 291 protein (EPA-Pil) 20 The theoretical mass of GBSIII-291 bioconjugate is 69,582.19 Da. Multiple mass-increasing states of GBSIII-291 bioconjugate were observed, with differences between these masses approaching 980 Da, corresponding to the calculated mass of the GBSIII glycan repeating unit. (Figure 32F) The kinetics of GBSIII-specific IgG during immunization were measured by ELISA and converted to ng / mL IgG using a standard IgG curve. **P < 0.01.

[0059] Figures 33A and 33B. Figure 33 illustrates the glycosylation of EPA constructs containing sequences from different O-linked oligosaccharide transferase systems. Figure 33A) is a schematic diagram of plasmid-based operons that express EPA with PglS or TfpM-specific sequences. The internal glycosylation tag (“iGT”) is derived from the insertion into the EPA Ala...489 –Arg 490 and / or Glu 548 –Gly 549 ComP between 110264 A 23-amino acid fragment. Figure 33B) Western blot of anti-EPA protein from SDB1 protein extract expressing one of the four constructs and E. coli O16 O-antigen. Loading amount per channel relative to OD 600 Normalization was performed. "g0" indicates unglycosylated EPA carrier protein, and monoglycosylated or disaccharidated EPA proteins are labeled. The reference protein mass, in kDa, is marked on the left side of the protein blot. Detailed Implementation

[0060] This disclosure provides a description of compositions and methods for producing glycoconjugated polypeptides using enzymes that form glycosidic bonds, wherein the polypeptide can spontaneously form isopeptide bonds with a second partner polypeptide containing a polypeptide tag, and a description of the use of the glycoconjugated polypeptide.

[0061] definition

[0062] It should be noted that the term "a (or an)" refers to one or more of the entities; for example, "polysaccharide" should be understood to mean one or more polysaccharides. Thus, the terms "a (or an)," "one or more," and "at least one" are used interchangeably in this document.

[0063] Furthermore, the term “and / or” as used herein is considered to be a specific disclosure of each of the specified features or components, with or without the other. Therefore, the term “and / or” as used in phrases such as “A and / or B” herein is intended to include “A and B”, “A or B”, “A” (alone), and “B” (alone). Similarly, the term “and / or” as used in phrases such as “A, B, and / or C” is intended to cover each of the following embodiments: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).

[0064] It should be understood that when an aspect is described in the language “comprising” or “comprises” wherever it is used in this document, similar aspects described in other aspects according to “consisting of” and / or “consisting essentially of” are also provided.

[0065] Unless otherwise defined, the technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0066] Numerical ranges include numbers within a defined range. Even if not explicitly identified by expressions such as "and any range therebetween", when listing a list of numerical values ​​(e.g., 1, 2, 3, or 4), this disclosure intentionally includes any range between such numerical values, such as 1 to 3, 1 to 4, 2 to 4, etc., unless otherwise stated.

[0067] The headings provided herein are for convenience only and are not intended to limit the various aspects of this disclosure, which can be obtained by referring to the entire specification.

[0068] As used herein, the term "polypeptide" is intended to encompass both the singular and plural forms of "polypeptide" and refers to a molecule composed of monomers (amino acids) linearly linked by amide bonds (also known as peptide bonds). The term "polypeptide" refers to any one or more chains of two or more amino acids, not a specific length of the product. Therefore, peptide, dipeptide, tripeptide, oligopeptide, "protein," "amino acid chain," or any other term used to refer to one or more chains of two or more amino acids are included within the definition of "polypeptide," and the term "polypeptide" may be used in place of or interchangeably with any of these terms. The term "polypeptide" is also intended to refer to products of post-expression modifications of a polypeptide, including but not limited to glycosylation, acetylation, phosphorylation, amidation, derivatization of known protecting / blocking groups, proteolytic cleavage, or modifications by non-standard amino acids. Polypeptides may be derived from natural biological sources or produced through recombinant technologies, but are not necessarily translated from a specified nucleic acid sequence. They can be generated in any manner, including through chemical synthesis.

[0069] As used herein, “protein” can refer to a single polypeptide, i.e., a single chain of amino acids as defined above, but can also refer to two or more polypeptides that associate together, for example, through disulfide bonds, hydrogen bonds or hydrophobic interactions, to form a multimeric protein.

[0070] "Isolated" polypeptides or fragments, variants, or derivatives thereof refer to polypeptides that are not present in their native environment. No specific degree of purification is required. For example, isolated polypeptides can be removed from their native or natural environment. Recombinant polypeptides and proteins expressed in host cells, as well as recombinant polypeptides isolated, fractionated, or partially or substantially purified by any suitable technique, are considered isolated, as disclosed herein.

[0071] A "vector" (which may also be used interchangeably with "plasmid" in this text) is a nucleic acid molecule introduced into a host cell to produce a transformed host cell. A vector may contain a nucleic acid sequence that allows it to replicate within the host cell (e.g., at the origin of replication). A vector may encode and express proteins. A vector may also contain one or more optional marker genes and other genetic elements known in the art.

[0072] "Transformed" cells or "host cells" are cells into which nucleic acid molecules have been introduced through molecular biology techniques. As used herein, the term transformation encompasses techniques that can introduce nucleic acid molecules into such cells, including transfection with viral vectors, transformation with plasmid vectors, and accelerated introduction of naked DNA via electroporation, liposome transfection, and particle guns. Transformed cells or host cells can be bacterial or eukaryotic cells.

[0073] As used herein, the term "expression" refers to the process by which a gene produces a biochemical substance (e.g., a polypeptide). This process includes any manifestation of the functional presence of a gene within a cell, including but not limited to gene knockdown and both transient and stable expression. It includes, but is not limited to, transcribing a gene into messenger RNA (mRNA) and translating such mRNA into a polypeptide. If the desired final product is a biochemical substance, expression includes the creation of that biochemical substance and any precursors thereof. Gene expression produces a "gene product." As used herein, a gene product can be a nucleic acid (e.g., messenger RNA produced by transcription of a gene) or a polypeptide translated from a transcript. Gene products described herein further include nucleic acids with post-transcriptional modifications (e.g., polyadenylation) or polypeptides with post-translational modifications (e.g., methylation, glycosylation, lipid addition, association with other protein subunits, or proteolytic cleavage).

[0074] As used herein, the term "treatment" (e.g., in the phrase "treatment of a subject") refers to reducing the likelihood of a disease pathology, decreasing the occurrence of disease symptoms, for example, to the extent that a subject's survival is prolonged or discomfort is reduced. For example, treatment can refer to the ability of a therapy to reduce disease symptoms, signs, or causes when administered to a subject. Treatment also refers to relieving or reducing at least one clinical symptom, and / or inhibiting or delaying the progression of the disease, and / or preventing or delaying the onset of the disease or condition.

[0075] "Subject," "individual," "animal," "patient," or "mammal" refers to any subject requiring diagnosis, prognosis, or treatment, particularly a mammalian subject. Mammal subjects include humans, domesticated animals, farm animals, sporting animals, and zoo animals, such as humans, non-human primates, dogs, cats, guinea pigs, rabbits, rats, mice, horses, cattle, camels, bears, etc.

[0076] The terms "pharmaceutical composition" or "therapeutic composition" refer to formulations in a form in which the biological activity of the active ingredient is effective and which do not contain any additional components that would have unacceptable toxicity to the subject to whom the composition will be administered. Such compositions may be sterile.

[0077] As used in this article, "sugar" is a general term for carbohydrate molecules of any size; including but not limited to monosaccharides, disaccharides, trisaccharides, tetrasaccharides, pentasaccharides, hexasaccharides, heptasaccharides, oligosaccharides, or polysaccharides.

[0078] As used in this article, a "glycosidic bond" is a covalent bond between a sugar and another organic molecule (including but not limited to another sugar, protein, lipid, or nucleic acid).

[0079] As used herein, a “glycosylated fragment” or “sequence segment” is a sequence of consecutive amino acids in a protein that serves as a recognition and linking site for sugars that are covalently transferred to the protein by glycosyltransferases or oligosyltransferases.

[0080] As used herein, the term "translational fusion" can refer to a direct link (e.g., a carrier protein, N-terminal leader sequence, C-terminal tag, etc.) or an indirect link via an amino acid linker. Glycoconjugate peptides having an isopeptide bond with a second polypeptide partner.

[0081] This document provides a polypeptide pair comprising a polypeptide tag and a binding partner (e.g., another polypeptide tag), wherein the polypeptide tag and the binding partner can bind to each other via spontaneous formation of an isopeptide bond between a reactive residue contained in the binding partner and another reactive residue contained in the polypeptide tag. As used herein, when referring to one component of a polypeptide pair, the other component may be referred to as its partner.

[0082] Fusion protein

[0083] One component of the polypeptide pair disclosed herein (typically containing a so-called first polypeptide tag) may be a fusion protein. Thus, certain embodiments of this disclosure provide a fusion protein comprising: (i) a glycosylated fragment, and (ii) a first polypeptide tag, wherein the first polypeptide tag can spontaneously form an isopeptide bond with a second polypeptide tag binding partner. For the avoidance of ambiguity, a fusion protein containing a glycosylated fragment may comprise a segment of a protein containing a glycosylation site (also referred to herein as a “sequence element”), which is glycosylated, or, in other embodiments, may comprise a glycosylated full-length protein, which in turn comprises the glycosylated fragment. Numerous representative examples of glycosylated proteins and sequence elements, as well as associated enzymes, that can be used in the compositions and methods of this disclosure are provided in detail elsewhere herein. In some embodiments, the fusion protein comprises a carrier protein. For example, in some embodiments, the carrier protein may be Escherichia coli maltose-binding protein (MPB), Pseudomonas aeruginosa exotoxin A (EPA), Pseudomonas aeruginosa PcrV, CRM197, Haemophilus influenzae protein D, cholerae toxin B subunit, tetanus toxin, or any fragment thereof. In some embodiments, the length of the glycosylated fragment can be as short as 3 amino acids. For example, TfpM OTase can recognize a three-amino acid sequence. In some embodiments, the glycosylated fragment can be longer, including full-length or near-full-length proteins. Thus, in some embodiments, the length of the glycosylated fragment is at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, or 40 amino acids. In some embodiments, the glycosylation is not a full-length glycosylated protein, but a shorter fragment thereof, and therefore the length does not exceed 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, 40, 50, 60, 80, or 100 amino acids. Therefore, in some embodiments, the length of the glycosylated fragment is any from 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or 24 amino acids to any from 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24 or 25 amino acids. In some embodiments, the length of the glycosylated fragment is any from 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, or 40 amino acids to any from 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, 40, or 50 amino acids.In some embodiments, the length of the glycosylated fragment is any from 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, 40, 50, 60, or 80 amino acids to 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, 40, 50, 60, 80, or 100 amino acids.

[0084] In some embodiments, the fusion protein is a glycoconjugate comprising a sugar covalently linked to a glycosylation site / residue of the fusion protein via a glycosylation fragment (sequence). In some embodiments, the glycoconjugate is immunogenic. Those skilled in the art will recognize that the sugar can be covalently linked to the glycosylation fragment, for example, via N-linking, O-linking, or C-linking.

[0085] In some embodiments, the first polypeptide tag is fused to the N-terminus of the fusion protein via translation. In some embodiments, the first polypeptide tag is fused to the C-terminus of the fusion protein via translation. In some embodiments, the first polypeptide tag is internally fused to the fusion protein via translation. In some embodiments, the first polypeptide tag is internally fused to the sequence of the carrier protein via translation.

[0086] In some embodiments, the glycosylated fragment is fused to the N-terminus of the fusion protein via translation. In some embodiments, the glycosylated fragment is fused to the C-terminus of the fusion protein via translation. In some embodiments, the glycosylated fragment is internally fused to the fusion protein via translation. In some embodiments, the glycosylated fragment is internally fused to the sequence of the carrier protein via translation.

[0087] As can be understood from this disclosure as a whole, internal fusion within the fusion protein means that glycosylated fragments, peptide tags, etc., are not located at the C-terminus or N-terminus of the fusion protein, and do not include any signal / leader sequences, purification tags (e.g., His tags), etc. For example:

[0088] N-terminal

[0089]

[0090] N-terminal, non-internal

[0091]

[0092] C-terminal

[0093]

[0094] C-terminal, non-internal

[0095]

[0096] internal

[0097]

[0098] internal

[0099]

[0100] As shown above, in some embodiments of internal placement, glycosylated fragments, peptide tags, etc., can be placed (translationally fused) between individual carrier proteins (even carrier proteins of the same type). In some embodiments of internal placement, glycosylated fragments, peptide tags, etc., can be internally placed (translationally fused) within a single carrier protein sequence.

[0101] Not limited to any specific sequence, representative examples of a single polypeptide tag partner in a pair (generally referred to herein as the first polypeptide tag) include SpyTag (SEQ ID NO:416), SpyTag002 (SEQ ID NO:417), SpyTag003 (SEQ ID NO:418), or DogTag (SEQ ID NO:419). In some embodiments, SpyTag, SpyTag002, or SpyTag003 is fused translationally to the N-terminus of the fusion protein (e.g., Figure 2 In some embodiments, SpyTag, Spytag002, or Spytag003 is fused to the C-terminus of the fusion protein via translation (e.g., Figure 5 In some embodiments, DogTag is internally fused into the fusion protein via translation.

[0102] Representative examples of the fusion proteins disclosed herein include: Figure 2 , Figure 5 , Figure 7 , Figure 8 and Figure 9As shown. For example, EPA-Spytag-v1 (SEQ ID NO:427), EPA-Spytag-v2 (SEQ ID NO:428), and EPA-Spytag-v3 (SEQ ID NO:429). In some embodiments, the amino acid linker sequence is inserted translatorily between the components of the fusion protein (e.g., signal peptide, polypeptide tag sequence, glycosylated fragment, carrier protein, histidine tag, etc.). In some embodiments, the amino acid linker sequence is GGS, GGGGGG (SEQ ID NO:430), GGGGGGGG (SEQ ID NO:431), GGGGS (SEQ ID NO:432), EAAAK (SEQ ID NO:433), PAPAPPAPAP (SEQ ID NO:434), EAAAKEAAAK (SEQ ID NO:435), GGGGSPAPAP (SEQ ID NO:436), GGGGSGGGGS (SEQ ID NO:437), or EAAAKGGGGS (SEQ ID NO:438). Therefore, some embodiments include EPA-Spytag-v1 of SEQ ID NO:427, EPA-Spytag-v2 of SEQ ID NO:428, or EPA-Spytag-v3 of SEQ ID NO:429, having one or more amino acid linkers (such as those described above) that are translated into place between the components of the fusion protein. For example, EPA-Spytag-v1 of SEQ ID NO:427 has an amino acid linker SSG that is translated into place after the peptide tag. For example, EPA-Spytag-v1 of SEQ ID NO:440 has an amino acid linker EAAAKEAAAK (SEQ ID NO:435) that is translated into place after the peptide tag. It is contemplated that in some embodiments, other amino acid linkers (such as, but not limited to, those disclosed herein) may be similarly positioned.

[0103] ComP glycosylation fragment

[0104] Not limited to any specific glycosylation sequence, in some embodiments, the glycosylated fragment is a ComP glycosylation fragment. The sequence of the ComP-containing glycosylated fragment can be the full-length ComP protein. In some embodiments, the ComP glycosylated fragment comprises or consists of the following: the amino acid sequence CTGVTQIASGASAATTNVASAQC (SEQ ID NO:412) or a fragment thereof, which contains the amino acid ASA at positions 11 to 13 of SEQ ID NO:412. In some embodiments, the ComP glycosylated fragment comprises or consists of the following: the amino acid sequence CTGVTQIASGASAATTNVASAQC (SEQ ID NO:412) or a fragment of at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23 amino acids, which contains the amino acid ASA at positions 11 to 13 of SEQ ID NO:412. In some embodiments, the ComP glycosylation fragment comprises or consists of the following: the amino acid sequence CTGVTQIASGASAATTNVASAQC (SEQ ID NO:412) or a fragment of at least 10 amino acids, wherein the fragment contains the amino acid ASA at positions 11 to 13 of SEQ ID NO:412. In some embodiments, the ComP glycosylation fragment comprises or consists of the following: the amino acid sequence CTGVTQIASGASAATTNVASAQC (SEQ ID NO:412) or a fragment of at least 11 amino acids, wherein the fragment contains the amino acid ASA at positions 11 to 13 of SEQ ID NO:412.

[0105] In some embodiments, the ComP glycosylation fragment comprises or is composed of a variant of SEQ ID NO:412, which contains amino acid ASA at positions 11 to 13 of SEQ ID NO:412 and has one, two, three, four, five, or six amino acid substitutions, additions, and / or deletions. Those skilled in the art will understand that in any embodiment of the glycosylation fragment disclosed herein, for example, if an addition occurs at one position and a deletion occurs at a different position, the net effect on sequence length is zero, etc. Further, multiple additions or deletions may be sequential and / or discontinuous. And, in some embodiments, the substitution may be a conserved amino acid substitution. In some embodiments, the ComP glycosylation fragment comprises or is composed of a variant of SEQ ID NO:412, which contains amino acid ASA at positions 11 to 13 of SEQ ID NO:412 and cumulatively has one, two, three, four, five, or six amino acid substitutions, additions, and / or deletions.

[0106] In some embodiments, the ComP glycosylation fragment comprises or is composed of a variant of SEQ ID NO:412, which contains the amino acid ASA at positions 11 to 13 of SEQ ID NO:412 and has one, two, three, four, five, or six amino acid substitutions and / or additions. In some embodiments, the ComP glycosylation fragment comprises or is composed of a variant of SEQ ID NO:412, which contains the amino acid ASA at positions 11 to 13 of SEQ ID NO:412 and cumulatively has one, two, three, four, five, or six amino acid substitutions and / or additions.

[0107] In some embodiments, the ComP glycosylation fragment comprises or is composed of a variant of SEQ ID NO:412, which contains the amino acid ASA at positions 11 to 13 of SEQ ID NO:412 and has one, two, three, four, five, or six amino acid substitutions and / or deletions. In some embodiments, the ComP glycosylation fragment comprises or is composed of a variant of SEQ ID NO:412, which contains the amino acid ASA at positions 11 to 13 of SEQ ID NO:412 and cumulatively has one, two, three, four, five, or six amino acid substitutions and / or deletions.

[0108] In some embodiments, the ComP glycosylation fragment comprises or is composed of a variant of SEQ ID NO:412, which contains the amino acid ASA at positions 11 to 13 of SEQ ID NO:412 and has one, two, three, four, five, or six amino acid additions and / or deletions. In some embodiments, the ComP glycosylation fragment comprises or is composed of a variant of SEQ ID NO:412, which contains the amino acid ASA at positions 11 to 13 of SEQ ID NO:412 and cumulatively has one, two, three, four, five, or six amino acid additions and / or deletions.

[0109] In some embodiments, the ComP glycosylation fragment comprises or is composed of a variant of SEQ ID NO:412, which contains the amino acid ASA at positions 11 to 13 of SEQ ID NO:412 and has one, two, three, four, five, or six amino acid substitutions. In some embodiments, the ComP glycosylation fragment comprises or is composed of a variant of SEQ ID NO:412, which contains the amino acid ASA at positions 11 to 13 of SEQ ID NO:412 and has one, two, three, four, five, or six amino acid additions. In some embodiments, the ComP glycosylation fragment comprises or is composed of a variant of SEQ ID NO:412, which contains the amino acid ASA at positions 11 to 13 of SEQ ID NO:412 and has one, two, three, four, five, or six amino acid deletions.

[0110] In some embodiments, the ComP glycosylation fragment includes or consists of any additional ComP glycosylation fragment sequences described elsewhere herein.

[0111] TfpM-associated fimbriae protein glycosylation fragments

[0112] Not limited to any specific glycosylation sequence, in some embodiments, the glycosylated fragment is a TfpM-associated fimbriae glycosylation fragment. The sequence containing the TfpM-associated fimbriae glycosylation fragment can be the full-length TfpM-associated fimbriae. In some embodiments, the TfpM-associated fimbriae glycosylation fragment comprises or consists of the following: the PilMo fimbriae disulfide ring region (SEQ ID NO: 413) or a fragment thereof, which contains at least the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-associated fimbriae. In some embodiments, the TfpM-associated fimbriae glycosylation fragment comprises or consists of the following: the PilMo fimbriae disulfide ring region (SEQ ID NO: 413) or a fragment of at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22 or 23 amino acids, the fragment comprising at least the last three amino acids from the C-terminus of the TfpM-associated fimbriae (i.e., RGT).

[0113] In some embodiments, the TfpM-associated fimbriae glycosylation fragment comprises or is composed of a variant of the PilMo fimbriae disulfide ring region (SEQ ID NO: 413), which includes the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-associated fimbriae and has one, two, three, four, five, or six amino acid substitutions, additions, and / or deletions. In some embodiments, the TfpM-associated fimbriae glycosylation fragment comprises or is composed of a variant of the PilMo fimbriae disulfide ring region (SEQ ID NO: 413), which includes the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-associated fimbriae and has cumulatively one, two, three, four, five, or six amino acid substitutions, additions, and / or deletions.

[0114] In some embodiments, the TfpM-associated fimbriae glycosylation fragment comprises or is composed of a variant of the PilMo fimbriae disulfide ring region (SEQ ID NO: 413), which includes the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-associated fimbriae, and has one, two, three, four, five, or six amino acid substitutions and / or additions. In some embodiments, the TfpM-associated fimbriae glycosylation fragment comprises or is composed of a variant of the PilMo fimbriae disulfide ring region (SEQ ID NO: 413), which includes the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-associated fimbriae, and cumulatively has one, two, three, four, five, or six amino acid substitutions and / or additions.

[0115] In some embodiments, the TfpM-associated fimbriae glycosylation fragment comprises or is composed of a variant of the PilMo fimbriae disulfide ring region (SEQ ID NO: 413), which includes the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-associated fimbriae and has one, two, three, four, five, or six amino acid substitutions and / or deletions. In some embodiments, the TfpM-associated fimbriae glycosylation fragment comprises or is composed of a variant of the PilMo fimbriae disulfide ring region (SEQ ID NO: 413), which includes the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-associated fimbriae and cumulatively has one, two, three, four, five, or six amino acid substitutions and / or deletions.

[0116] In some embodiments, the TfpM-associated fimbriae glycosylation fragment comprises or is composed of a variant of the PilMo fimbriae disulfide ring region (SEQ ID NO: 413), which includes the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-associated fimbriae and has one, two, three, four, five, or six amino acid additions and / or deletions. In some embodiments, the TfpM-associated fimbriae glycosylation fragment comprises or is composed of a variant of the PilMo fimbriae disulfide ring region (SEQ ID NO: 413), which includes the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-associated fimbriae and has cumulatively one, two, three, four, five, or six amino acid additions and / or deletions.

[0117] In some embodiments, the TfpM-associated fimbriae glycosylation fragment comprises or is composed of a variant of the PilMo fimbriae disulfide ring region (SEQ ID NO: 413), which includes the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-associated fimbriae and has one, two, three, four, five, or six amino acid substitutions. In some embodiments, the TfpM-associated fimbriae glycosylation fragment comprises or is composed of a variant of the PilMo fimbriae disulfide ring region (SEQ ID NO: 413), which includes the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-associated fimbriae and has one, two, three, four, five, or six amino acid additions. In some embodiments, the TfpM-associated fimbriae glycosylation fragment comprises or is composed of a variant of the PilMo fimbriae disulfide ring region (SEQ ID NO:413) containing the last three amino acids (i.e., RGT) from the C-terminus of the TfpM-associated fimbriae and having one, two, three, four, five, or six amino acid deletions.

[0118] In some embodiments, the TfpM-associated fimbriae glycosylation fragment comprises or consists of any other TfpM-associated fimbriae glycosylation fragment sequences described elsewhere herein.

[0119] PilE glycosylation fragment

[0120] Not limited to any specific glycosylation sequence, in some embodiments, the glycosylated fragment is a PilE glycosylation fragment. The sequence of the glycosylated fragment containing PilE can be the full-length PilE protein. In some embodiments, the PilE glycosylated fragment contains or consists of the following: amino acid SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO:414) or a fragment thereof, which contains at least the amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414. In some embodiments, the PilE glycosylated fragment comprises or consists of the following: amino acid SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO:414) or a fragment of at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28 or 29 amino acids, wherein the fragment contains at least the amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414.

[0121] In some embodiments, the PilE glycosylated fragment comprises or is composed of a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO:414), which contains at least the amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414 and has one, two, three, four, five, or six amino acid substitutions, additions, and / or deletions. In some embodiments, the PilE glycosylated fragment comprises or is composed of a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO:414), which contains at least the amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414 and cumulatively has one, two, three, four, five, or six amino acid substitutions, additions, and / or deletions.

[0122] In some embodiments, the PilE glycosylated fragment comprises or is composed of a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO:414), which contains at least the amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414, and has one, two, three, four, five, or six amino acid substitutions and / or additions. In some embodiments, the PilE glycosylated fragment comprises or is composed of a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO:414), which contains at least the amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414, and cumulatively has one, two, three, four, five, or six amino acid substitutions and / or additions.

[0123] In some embodiments, the PilE glycosylated fragment comprises or is composed of a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO:414), which contains at least the amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414 and has one, two, three, four, five, or six amino acid substitutions and / or deletions. In some embodiments, the PilE glycosylated fragment comprises or is composed of a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO:414), which contains at least the amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414 and cumulatively has one, two, three, four, five, or six amino acid substitutions and / or deletions.

[0124] In some embodiments, the PilE glycosylated fragment comprises or is composed of a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO:414), which contains at least the amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414, and has one, two, three, four, five, or six amino acid additions and / or deletions. In some embodiments, the PilE glycosylated fragment comprises or is composed of a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO:414), which contains at least the amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414, and cumulatively has one, two, three, four, five, or six amino acid additions and / or deletions.

[0125] In some embodiments, the PilE glycosylated fragment comprises or is composed of a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO:414), which contains at least the amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414 and has one, two, three, four, five, or six amino acid substitutions. In some embodiments, the PilE glycosylated fragment comprises or is composed of a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO:414), which contains at least the amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414 and has one, two, three, four, five, or six amino acid additions. In some embodiments, the PilE glycosylated fragment comprises or is composed of a variant of SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO:414), which contains at least the amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414 and has one, two, three, four, five or six amino acid deletions.

[0126] PglB glycosylation fragment

[0127] Not limited to any specific glycosylation sequence, in some embodiments, the glycosylated fragment is a PglB glycosylated fragment. The sequence of the glycosylated fragment containing PglB can be the full-length PglB protein. In some embodiments, the PglB glycosylated fragment contains or consists of the following: a common motif amino acid sequence X1 X2 N X3 X4, where X1 is D or E, X2 is any amino acid except proline, X3 is any amino acid except proline, and X4 is S or T.

[0128] PilA glycosylation fragment

[0129] Not limited to any specific glycosylation sequence, in some embodiments, the glycosylated fragment is a PilA glycosylated fragment. The sequence of the glycosylated fragment containing PilA can be the full-length PilA protein. In some embodiments, the PilA glycosylated fragment comprises or consists of the following: the PilA fimbriae disulfide ring region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO: 415) or a fragment thereof, which contains at least the last three amino acids (i.e., PKS) from the C-terminus of PilA. In some embodiments, the PilA glycosylated fragment comprises or consists of the following: the PilA fimbriae disulfide ring region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO:415) or a fragment of at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or 21 amino acids, the fragment comprising at least the last three amino acids from the C-terminus of PilA (i.e., PKS).

[0130] In some embodiments, the PilA glycosylated fragment comprises or is composed of a variant of the PilA fimbriae disulfide ring region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO:415), which contains at least the last three amino acids (i.e., PKS) from the C-terminus of PilA and has one, two, three, four, five, or six amino acid substitutions, additions, and / or deletions. In some embodiments, the PilA glycosylated fragment comprises or is composed of a variant of the PilA fimbriae disulfide ring region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO:415), which contains at least the last three amino acids (i.e., PKS) from the C-terminus of PilA and has cumulatively one, two, three, four, five, or six amino acid substitutions, additions, and / or deletions.

[0131] In some embodiments, the PilA glycosylated fragment comprises or is composed of a variant of the PilA fimbriae disulfide ring region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO:415), the variant comprising at least the last three amino acids (i.e., PKS) from the C-terminus of PilA, and having one, two, three, four, five, or six amino acid substitutions and / or additions. In some embodiments, the PilA glycosylated fragment comprises or is composed of a variant of the PilA fimbriae disulfide ring region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO:415), the variant comprising at least the last three amino acids (i.e., PKS) from the C-terminus of PilA, and cumulatively having one, two, three, four, five, or six amino acid substitutions and / or additions.

[0132] In some embodiments, the PilA glycosylated fragment comprises or is composed of a variant of the PilA fimbriae disulfide ring region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO:415), which contains at least the last three amino acids (i.e., PKS) from the C-terminus of PilA and has one, two, three, four, five, or six amino acid substitutions and / or deletions. In some embodiments, the PilA glycosylated fragment comprises or is composed of a variant of the PilA fimbriae disulfide ring region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO:415), which contains at least the last three amino acids (i.e., PKS) from the C-terminus of PilA and cumulatively has one, two, three, four, five, or six amino acid substitutions and / or deletions.

[0133] In some embodiments, the PilA glycosylated fragment comprises or is composed of a variant of or is composed of the PilA fimbriae disulfide ring region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO:415), which contains at least the last three amino acids (i.e., PKS) from the C-terminus of PilA and has one, two, three, four, five, or six amino acid additions and / or deletions. In some embodiments, the PilA glycosylated fragment comprises or is composed of a variant of or is composed of the PilA fimbriae disulfide ring region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO:415), which contains at least the last three amino acids (i.e., PKS) from the C-terminus of PilA and has cumulatively one, two, three, four, five, or six amino acid additions and / or deletions.

[0134] In some embodiments, the PilA glycosylated fragment comprises or is composed of a variant of the PilA fimbriae disulfide ring region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO:415), which contains at least the last three amino acids (i.e., PKS) from the C-terminus of PilA and has one, two, three, four, five, or six amino acid substitutions. In some embodiments, the PilA glycosylated fragment comprises or is composed of a variant of the PilA fimbriae disulfide ring region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO:415), which contains at least the last three amino acids (i.e., PKS) from the C-terminus of PilA and has one, two, three, four, five, or six amino acid additions. In some embodiments, the PilA glycosylated fragment comprises or is composed of a variant of or is composed of the PilA fimbriae disulfide ring region CKITKTPTAWKPNYAPANCPKS (SEQ ID NO:415), which contains at least the last three amino acids (i.e., PKS) from the C-terminus of PilA and has one, two, three, four, five or six amino acid deletions.

[0135] PilA Pa5196 glycosylation fragment

[0136] Not limited to any specific glycosylation sequence, in some embodiments, the glycosylated fragment is a glycosylated fragment of PilA_Pa5196-associated fimbriae. The sequence containing the glycosylated fragment of PilA_Pa5196-associated fimbriae can be the full-length PilA_Pa5196-associated fimbriae. In some embodiments, the PilA_Pa5196-associated fimbriae glycosylated fragment comprises or consists of the following: strands 1 and 2 of the antiparallel β-sheet domain of PilA_Pa5196 GKYSSVDSTIASGYPNGQITVTMTQG (SEQ ID NO:426), or fragments thereof. In some embodiments, the PilA_Pa5196-related fimbriae protein glycosylation fragment comprises or consists of the following: chains 1 and 2 of the antiparallel β-sheet domain of PilA_Pa5196GKYSSVDSTIASGYPNGQITVTMTQG (SEQ ID NO:426), or fragments thereof having a length of at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids.

[0137] In some embodiments, the glycosylated fragment is a variant of the PilA_Pa5196-associated fimbriae protein glycosylation fragment, which is consistent with other variants of glycosylated fragments disclosed herein, which have one, two, three, four, five, or six amino acid substitutions, additions, and / or deletions.

[0138] STT3 glycosylation fragment

[0139] Not limited to any specific glycosylation sequence, in some embodiments, the glycosylated fragment is an STT3 glycosylation fragment. The sequence of the glycosylated fragment containing STT3 can be the full-length STT3 protein. In some embodiments, the STT3 glycosylation fragment contains or consists of the common motif amino acid sequence NXS / T, where X is any amino acid other than proline, and S / T is serine (S) or threonine (T).

[0140] N-linked glycosyltransferase glycosylation fragment

[0141] Not limited to any specific glycosylation sequence, in some embodiments, the glycosylation fragment is an N-linked glycosyltransferase glycosylation fragment. In some embodiments, the N-linked glycosyltransferase glycosylation fragment comprises or consists of a common motif amino acid sequence NXS / T, where X is any amino acid other than proline, and S / T is serine (S) or threonine (T).

[0142] O-linked glycosyltransferase glycosylation fragment

[0143] Not limited to any specific glycosylation sequence, in some embodiments, the glycosylated fragment is an O-linked glycosyltransferase glycosylation fragment. In some embodiments, the O-linked glycosyltransferase glycosylation fragment comprises or consists of fragments of serine- or threonine-rich repeat sequences from serine-rich (SRR) adhesins derived from Streptococcus or Staphylococcus bacteria. In some embodiments, the O-linked glycosyltransferase glycosylation fragment comprises or consists of fragments of serine (S)- or threonine (T)-rich repeat sequences from the adhesin GspB of Streptococcus gordonii.

[0144] Multiple glycosylation fragments

[0145] Certain aspects of this disclosure relate to fusion proteins comprising two or more glycosylated fragments. For example, the fusion protein comprises at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 glycosylated fragments. In some embodiments, the fusion protein comprises any of 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23 glycosylated fragments up to any of 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 glycosylated fragments. In some embodiments, at least one glycosylation fragment is located at the N-terminus of the fusion protein and at least one glycosylation fragment is internally located within the fusion protein. In some embodiments, at least one glycosylation fragment is located at the C-terminus of the fusion protein and at least one glycosylation fragment is internally located within the fusion protein. In some embodiments, at least two glycosylation fragments are internally located within the fusion protein. Further, in some embodiments, one glycosylation fragment is located at the N-terminus of the fusion protein and one glycosylation fragment is located at the C-terminus of the fusion protein. In some embodiments, two or more glycosylation fragments are identical. For example, a fusion protein having multiple ComP glycosylation fragments. In some embodiments, at least one of the two or more glycosylation fragments is different, or each of these glycosylation fragments is different. For example, a fusion protein in which one glycosylation fragment is a ComP glycosylation fragment and one glycosylation fragment is a TfpM-associated fimbriae glycosylation fragment.

[0146] In some embodiments, the fusion protein is a glycoconjugate comprising two or more sugars covalently linked to the fusion protein via two or more glycosylated fragments. In some embodiments, the fusion protein comprises at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 covalently linked sugars. In some embodiments, the fusion protein comprises any one of 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 23 covalently linked sugars up to any one of 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 covalently linked sugars. In some embodiments, the two or more sugars are identical. In some embodiments, at least one of the two or more sugars is different, or each of these sugars is different.

[0147] It is understandable that, in addition to combining different fusion proteins within the same complex, it is possible to create fusion proteins with multiple glycosylation fragments, and further, these glycosylation fragments can be recognized by different glycosylation enzymes and / or linked to different sugars, thereby allowing for the mixing, matching, and amplification of various immunogenic components.

[0148] In some embodiments, the fusion protein comprises a carrier protein. Various carrier proteins have been used in glycoconjugate vaccines, and all of these carrier proteins are contemplated herein. For example, in some embodiments, the carrier protein is selected from the group consisting of: *E. coli* maltose-binding protein, *Pseudomonas aeruginosa* exotoxin A (EPA), *Pseudomonas aeruginosa* PcrV, CRM197, *Haemophilus influenzae* protein D, cholera toxin B subunit, or tetanus toxin, and any fragments thereof.

[0149] polypeptide pairs

[0150] Certain aspects of this disclosure relate to a composition comprising a peptide pair comprising a first peptide and a second peptide. The first peptide comprises a first peptide tag, which is a binding partner of a second peptide tag relative to the second peptide. In some embodiments, the first peptide is a fusion protein of this disclosure comprising a glycosylated fragment as described in detail elsewhere herein. The second peptide comprises a second peptide tag binding partner relative to the first peptide tag of the first peptide. The first peptide may be linked to the second peptide via an isopeptide bond between the first and second peptide tags. In some embodiments, the second peptide comprises a monomeric peptide capable of spontaneously polymerizing / self-assembling into a higher-order multimeric structure. For example, icosahedral or dodecahedral particles (e.g., similar to nanocages), virus-like particles (VLPs), or adenovirus vectors. In some embodiments, the second peptide comprises an adenovirus capsid structural protein. In some embodiments, the second peptide comprises the coat protein of bacteriophage AP205. In some embodiments, the second peptide comprises a fragment of 2-keto-3-deoxy-phosphoglucuronide aldolase (i301). In some embodiments, the second polypeptide comprises a fragment of a mutated 2-keto-3-deoxy-phosphoglucuronide (mi3). In some embodiments, the polypeptide tag of the second polypeptide (the second polypeptide tag) is SpyCatcher (SEQ ID NO:420). In some embodiments, the second polypeptide tag is SpyCatcher002 (SEQ ID NO:421). In some embodiments, the second polypeptide tag is SpyCatcher003 (SEQ ID NO:422). In some embodiments, the second polypeptide tag is DogCatcher (SEQ ID NO:423).

[0151] The peptide tag of the second peptide can be located at the terminal (N-terminus or C-terminus) or internally of the second peptide. In some embodiments, the second peptide tag is fused to the N-terminus of the second peptide via translation (e.g., Figure 2 In some embodiments, the second polypeptide tag is fused to the C-terminus of the second polypeptide via translation. In some embodiments, the second polypeptide tag is internally fused to the second polypeptide via translation.

[0152] In some embodiments, the first polypeptide (e.g., the fusion protein of this disclosure) is a bioconjugate comprising a sugar covalently linked to a glycosylated fragment of the first polypeptide.

[0153] In some embodiments, the composition comprising the peptide pair is immunogenic, for example, wherein the first peptide is a bioconjugate comprising a sugar covalently linked to a glycosylated fragment of the first peptide. In some embodiments, the peptide pair composition further comprises an adjuvant and / or excipient. Examples of adjuvants may include, but are not limited to: alum (aluminum hydroxide gel or aluminum phosphate gel), squalene emulsions (e.g., MF59, AddaSO3, or AddaVax), lipid A derivatives (such as lipid monophosphate A (MPLA)), or saponins (e.g., Quil-A). In some embodiments, the peptide pair composition is a pharmaceutical and / or therapeutic composition. In some embodiments, the peptide pair composition is a conjugate vaccine.

[0154] Some aspects provide a method for preparing peptide pairs of the present disclosure. This can be achieved by contacting a first peptide (e.g., a fusion protein of the present disclosure) and a second peptide under conditions that allow a first peptide tag to spontaneously bind to its corresponding second peptide tag to form an isopeptide bond. In some embodiments, the method further includes glycosylation of the first peptide with a sugar (e.g., see [link to relevant documentation]) before contacting the second peptide and forming an isopeptide bond. Figure 1 In some embodiments, the first polypeptide is glycosylated in vivo (e.g., in host cells, for example, in bacteria) before contacting the second polypeptide and forming an isopeptide bond. In some embodiments, the method includes isolating / purifying the in vivo glycosylated first polypeptide before contacting the second polypeptide and forming an isopeptide bond. Alternatively, in some embodiments, the first polypeptide is glycosylated after contacting the second polypeptide and forming an isopeptide bond.

[0155] complex

[0156] This disclosure provides a complex comprising two or more polypeptide pairs disclosed herein. In some embodiments, a single complex comprises 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 75, 100, 125, 150, 175, 200, 225, 250 or more of the complex polypeptide pairs disclosed herein. In some embodiments, a single complex comprises any of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 75, 100, 125, 150, 175, 200, 225, or 250 of the complex peptide pairs disclosed herein. One to any one of 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 75, 100, 125, 150, 175, 200, 225, 250, 300, 400, or 500 of the complex polypeptide pairs disclosed herein. In some embodiments, such complexes are self-assembled higher-order multimeric structures. In some embodiments, such self-assembled higher-order multimeric structures are icosahedral or dodecahedral particles (e.g., similar to nanocages), virus-like particles, or adenovirus vectors.

[0157] Because the peptide tag of the first peptide, rather than a glycosylated fragment or carrier protein (e.g., a glycosylated fragment or carrier protein of a fusion protein), determines its partner peptide, the second peptide is not limited to pairing with only one type of first peptide (and vice versa). In some embodiments, all the first peptides of the complex contain the same fusion protein. However, in some embodiments, the first peptides may contain different fusion proteins. In some embodiments, at least two, three, four, five, or more of the first peptides of the complex contain different fusion proteins. In some embodiments, at least two of the first peptides of the complex contain different fusion proteins. In some embodiments, two, three, four, five, or six of the first peptides of the complex contain different fusion proteins. In some embodiments, all the first peptides of the complex are different fusion proteins.

[0158] In some embodiments, at least one first polypeptide of the complex is a bioconjugate comprising a sugar covalently linked to a glycosylated fragment of the first polypeptide. In some embodiments, at least about 5%, 10%, 25%, 50%, 75%, 80%, 90%, 95%, 97%, 98%, or 99% of the first polypeptide in the complex is a bioconjugate. In some embodiments, any percentage of the first polypeptide from about 5%, 10%, 25%, 50%, 75%, 80%, 90%, 95%, 97%, or 98% to about 10%, 25%, 50%, 75%, 80%, 90%, 95%, 97%, 98%, or 99% of the first polypeptide in the complex is a bioconjugate. In some embodiments, about 100% or 100% of the first polypeptide in the complex is a bioconjugate.

[0159] In some embodiments, two or more of the first polypeptides in the complex are bioconjugates containing covalently linked sugars. The number of covalently linked sugars can be large due to the number of first polypeptides in the complex and will depend on the number of first polypeptide / second polypeptide pairs in the complex and the number of sugars linked to each first polypeptide. For example, each AP205 VLP contains approximately 180 first polypeptide binding partners per VLP (e.g., SpyCatcher). Theoretically, 180 bioconjugates per VLP are permissible if 100% of the sugars are bound to the first polypeptide as a bioconjugate via isopeptide bonds. Furthermore, as described elsewhere herein, each bioconjugate can be covalently linked to multiple sugars. Mi3 is lower, with approximately 60 first polypeptide binding partners per NP. In some embodiments, the complex comprises at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1,000, 1,500, 2,000, 2,500, or 5,000 covalently linked sugars. In some embodiments, the complex comprises 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1,000, 1,500, 2,000, or 2,500 covalently linked sugars. Any one of the following covalently linked sugars: 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1,000, 1,500, 2,000, 2,500, or 5,000. In some embodiments, the complex comprises at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, or 750 covalently linked sugars.In some embodiments, the complex comprises any one of 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, or 500 covalently linked sugars to any one of 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, or 750 covalently linked sugars. In some embodiments, all sugars linked to the complex are identical. In some embodiments, at least two, three, four, five, or more sugars linked to the complex are different. In some embodiments, at least two sugars linked to the complex are different. In some embodiments, two, three, four, five, or six sugars linked to the complex are different from each other. In some embodiments, each sugar linked to the complex is different.

[0160] In some embodiments of the complex, such as when it is linked to a sugar, the complex is immunogenic.

[0161] Some embodiments provide a pharmaceutical and / or therapeutic composition comprising the complex of the present invention, as well as adjuvants and / or excipients. In some embodiments, the complex is a conjugate vaccine.

[0162] Some aspects provide a method for preparing the complex of the present invention. This can be achieved by forming a self-assembled multimeric higher-order structure of the second polypeptide of the present invention (e.g., icosahedral or dodecahedral particles (e.g., similar to nanocages), virus-like particles (VLPs), or adenovirus vectors), and then contacting the first polypeptide of the present invention (e.g., a fusion protein containing a glycosylated fragment) and the second polypeptide under conditions that allow the first polypeptide tag to spontaneously form heteropeptide bonds with the second polypeptide tag (e.g., ...). Figure 1 This can also be achieved by contacting the first and second polypeptides under conditions that allow the first polypeptide tag to spontaneously form isopeptide bonds with the second polypeptide tag, and then forming a self-assembled higher-order multimeric structure of the second polypeptide. It should be understood that such methods encompass a wide range of possibilities and combinations for glycosylation of the first polypeptide of the complex with sugars, all of which are envisioned herein. For example, the first polypeptide can be glycosylated prior to the formation of isopeptide bonds between the first and second polypeptides (e.g., Figure 1Glycosylation of the first polypeptide can be performed after an isopeptide bond is formed between the first and second polypeptides. Glycosylation of the first polypeptide can also be performed before incorporation into a higher-order polymeric structure (e.g., Figure 1 Furthermore, the first polypeptide can be glycosylated after it has been incorporated into the higher-order structure of the polymer.

[0163] Glycosylation

[0164] The glycosylated residues, glycosylation sites, glycosylated fragments, sequences, first polypeptide fusion proteins, complexes, etc., disclosed herein can be covalently linked to sugars through various glycosylation methods (including but not limited to the illustrative examples below). In some embodiments, sugars are transferred to the fusion protein containing the glycosylated fragments by the action of N-linked oligosaccharide transferases (N-OTase), O-linked oligosaccharide transferases (O-Otase), N-linked glycosaccharide transferases (NGT), O-linked glycosaccharide transferases (OGT), and / or C-mannosyltransferases (CMT). In some embodiments, sugars are transferred to the fusion protein containing the glycosylated fragments by the action of PglS OTase, TfpM OTase, PglLOTase, PglB OTase, TfpO / PilO OTase, STT3 OTase, TfpW glycosaccharide transferases, and / or AlgB OTase. For example, in some embodiments, the glycosylated fragment is a ComP glycosylated fragment by PglSOTase. In some embodiments, PglS OTase is used to covalently link the sugar to an oxygen atom within the glycosylated fragment. In some embodiments, the glycosylated fragment is a TfpM-associated fimbriae glycosylated fragment glycosylated with TfpM OTase. In some embodiments, TfpM OTase is used to covalently link the sugar to an oxygen atom within the glycosylated fragment. In some embodiments, the glycosylated fragment is a PilE glycosylated fragment glycosylated with PglL OTase. In some embodiments, PglL OTase is used to covalently link the sugar to an oxygen atom within the glycosylated fragment. In some embodiments, the glycosylated fragment is a PglB glycosylated fragment glycosylated with PglB OTase. In some embodiments, PglB OTase is used to covalently link the sugar to a nitrogen atom within the glycosylated fragment. In some embodiments, the glycosylated fragment is a PilA glycosylated fragment glycosylated with TfpO or PilO OTase. In some embodiments, TfpO or PilO OTase is used to covalently link the sugar to an oxygen atom within the glycosylated fragment. In some embodiments, the glycosylated fragment is an STT3 glycosylated fragment catalyzed by STT3 catalyzed subunit glycosylation. In some embodiments, STT3 OTase is used to covalently link the sugar to a nitrogen atom within the glycosylated fragment. In some embodiments, the glycosylated fragment is a PilA_Pa5196-associated fimbriae glycosylated fragment glycosylated by TfpW glycosyltransferase. In some embodiments, TfpW glycosyltransferase is used to covalently link the sugar to an oxygen atom within the glycosylated fragment. In some embodiments, the glycosylated fragment is an archaea AlgB glycosylated fragment glycosylated by AlgB OTase. In some embodiments, AlgB OTase is used to covalently link the sugar to a nitrogen atom within the archaea AlgB glycosylated fragment.In some embodiments, the glycosylated fragment is an N-linked glycosyltransferase glycosylated fragment glycosylated by an N-linked glycosyltransferase (e.g., from *Actinobacillus pleuropneumoniae*, *Haemophilus influenzae*, or *Yersinia enterocolitica*). In some embodiments, an N-linked glycosyltransferase is used to covalently link the sugar to a nitrogen atom within the glycosylated fragment. In some embodiments, the glycosylated fragment is an O-linked glycosyltransferase glycosylated fragment glycosylated by an O-linked glycosyltransferase (e.g., GtfA / GtfB glycosyltransferase). In some embodiments, an O-linked glycosyltransferase is used to covalently link the sugar to an oxygen atom within the glycosylated fragment. Further, in some embodiments, a C-mannosyltransferase is used to covalently link the sugar to a carbon atom within the glycosylated fragment.

[0165] Further, illustratively, in some embodiments, a PglS OTase (e.g., SEQ ID NO: 400) is used to covalently link the sugar to an oxygen atom within the ComP glycosylation fragment (e.g., SEQ ID NO: 412 or a variant thereof). In some embodiments, a TfpM OTase (e.g., SEQ ID NO: 402) is used to covalently link the sugar to an oxygen atom within the TfpM glycosylation fragment (e.g., SEQ ID NO: 413 or a variant thereof). In some embodiments, a PglL OTase (e.g., SEQ ID NO: 404) is used to covalently link the sugar to an oxygen atom within the PilE glycosylation fragment (e.g., SEQ ID NO: 414 or a variant thereof, e.g., SEQ ID NO: 439). In some embodiments, a PglL OTase (e.g., SEQ ID NO: 404) is used to covalently link the sugar to an oxygen atom within the PilE glycosylation fragment (e.g., SEQ ID NO: 414 or a variant thereof). In some embodiments, a PglB oxidase (e.g., SEQ ID NO: 405) is used to covalently link the sugar to a nitrogen atom within the PglB glycosylated fragment. In some embodiments, a TfpO / PilO oxidase (e.g., SEQ ID NO: 407) is used to covalently link the sugar to an oxygen atom within the PilA glycosylated fragment (e.g., SEQ ID NO: 415 or a variant thereof). In some embodiments, an STT3 oxidase (e.g., SEQ ID NO: 408) is used to covalently link the sugar to a nitrogen atom within the STT3 glycosylated fragment. In some embodiments, an AlgB oxidase (e.g., SEQ ID NO: 409) is used to covalently link the sugar to a nitrogen atom within the archaea AlgB glycosylated fragment. In some embodiments, a TfpW glycosyltransferase (e.g., SEQ ID NO: 424) is used to covalently link the sugar to an oxygen atom within the PilA_Pa5196-associated fimbriae glycosylated fragment (e.g., SEQ ID NO: 426 or a variant thereof). In some embodiments, an N-linked glycosyltransferase (e.g., SEQ ID NO: 410) is used to covalently link the sugar to a nitrogen atom within the N-linked glycosyltransferase sequence. In some embodiments, an O-linked glycosyltransferase (e.g., SEQ ID NO: 411) is used to covalently link the sugar to an oxygen atom within the O-linked glycosyltransferase sequence. In some embodiments, a C-mannosyltransferase is used to covalently link the sugar to a carbon atom within the C-mannosyltransferase glycosylation fragment.

[0166] In some embodiments, the method is a method of producing a conjugate vaccine. This may involve adding adjuvants and / or excipients to the couples and / or complexes of this disclosure.

[0167] Another aspect provides a system comprising a first polypeptide and a second polypeptide of the compositions of this disclosure. In some embodiments, the first polypeptide is a glycosylated bioconjugate. In some embodiments, the system comprises a higher-order multimeric structure assembled from the second polypeptide. In some embodiments, the system comprises sugars and N-linked oligosaccharide transferases (N-Otase), O-linked oligosaccharide transferases (O-OTase), N-linked glycosyltransferases (NGT), O-linked glycosyltransferases (OGT), and / or C-mannosyltransferases (CMT) as disclosed herein.

[0168] On the other hand, an isolated nucleic acid is provided that encodes a first polypeptide and / or a second polypeptide of the compositions and / or complexes disclosed herein. Some embodiments relate to a vector containing the isolated nucleic acid. Some embodiments relate to a host cell containing the vector.

[0169] On the other hand, a kit is provided comprising two or more components, including: the fusion protein of the present disclosure, a first polypeptide, a second polypeptide, a sugar, an N-linked oligosaccharide transferase (N-OTase), an O-linked oligosaccharide transferase (O-OTase), an N-linked glycosyltransferase (NGT), an O-linked glycosyltransferase (OGT) and / or a C-mannosyltransferase (CMT), a bioconjugate, a multimeric higher-order structure assembled from the second polypeptide, isolated nucleic acids, a vector, and a host cell.

[0170] On the other hand, a method is provided for inducing an immune response in a subject by administering an effective amount of any of the compositions, complexes, and / or conjugate vaccines of this disclosure. Further, compositions, complexes, and / or conjugate vaccines of this disclosure are provided for inducing an immune response in a subject.

[0171] In some embodiments, the compositions or complexes disclosed herein are conjugate vaccines that can be administered to a subject to prevent and / or treat infections and / or diseases. In some embodiments, the conjugate vaccine is a prophylactic measure that can be used, for example, to immunize a subject against infection and / or disease. In some embodiments, the glycoconjugate is associated with an adjuvant (such as in a therapeutic composition) and / or administered with an adjuvant. Some embodiments provide a composition (such as a therapeutic composition) comprising the conjugate vaccine described herein and an adjuvant. In some embodiments, when the conjugate vaccine is administered to a subject, it induces an immune response. In some embodiments, the immune response triggers long-term memory (memory B cells and T cells). In some embodiments, the immunity is an antibody response. In some embodiments, the antibody response is a serotype-specific antibody response. In some embodiments, the antibody response is an IgG or IgM response. In some embodiments, when the antibody response is an IgG response, the IgG response is an IgG1 response. Further, in some embodiments, the conjugate vaccine generates immune memory in a subject who has received the vaccine.

[0172] Some embodiments also provide for generating vaccines against infections and / or diseases. In some embodiments, the method includes isolating the glycoconjugate or fusion protein disclosed herein (conjugate vaccine) and combining the conjugate vaccine with an adjuvant. In some embodiments, the infection is a local or systemic infection of the skin, soft tissue, blood, or an organ, or is inherently autoimmune. In some embodiments, the vaccine is a conjugate vaccine against pneumococcal infection. In some embodiments, the disease is pneumonia. In some embodiments, the infection is a systemic infection and / or a blood infection. In some embodiments, the subject is a mammal. For example, in some embodiments, a pig or a human.

[0173] Importantly, the aspects disclosed herein are not limited to pneumococcal polysaccharides; in fact, they have broad applicability in generating bioconjugate vaccines against many important human and animal pathogens incompatible with PglB and PglL. Notable examples include human pathogens Klebsiella pneumoniae and Group B Streptococcus, as well as the swine pathogen Streptococcus suis, all of which are extremely important pathogens for which no licensed vaccines currently exist.

[0174] This document provides methods for inducing a host immune response against a pathogen. In some embodiments, the pathogen is a bacterial pathogen. In some embodiments, the host has been immunized against the pathogen. In some embodiments, the method includes administering an effective amount of the ComP conjugate vaccine, glycosylated fusion protein, or any other therapeutic / immunogenic composition disclosed herein to a subject requiring an immune response. Some embodiments provide the conjugate vaccine, glycosylated fusion protein, or other therapeutic / immunogenic composition disclosed herein for inducing a host immune response against a bacterial pathogen and for immunizing against the bacterial pathogen. Examples of immune responses include, but are not limited to, innate responses, adaptive responses, humoral responses, antibody responses, cell-mediated responses, B cell responses, T cell responses, cytokine upregulation or downregulation, immune system crosstalk, and combinations of two or more of said immune responses. In some embodiments, the immune response is an antibody response. In some embodiments, the immune response is an innate response, a humoral response, an antibody response, a T cell response, or a combination of two or more of said immune responses.

[0175] This document also provides methods for preventing or treating bacterial diseases and / or infections in subjects, including administering the conjugate vaccines, fusion proteins, or compositions disclosed herein to subjects in need. In some embodiments, the infection is a local or systemic infection of the skin, soft tissue, blood, or organs, or is inherently autoimmune. In some embodiments, the disease is pneumonia. In some embodiments, the infection is a systemic infection and / or a blood infection. In some embodiments disclosed herein, the subject is a vertebrate. In some embodiments, the subject is a mammal, such as a dog, cat, cow, horse, pig, mouse, rat, rabbit, sheep, goat, guinea pig, monkey, ape, camel, etc. And, for example, in some embodiments, the mammal is a human.

[0176] In any of the administration examples disclosed herein, the composition is administered via intramuscular injection, intradermal injection, intraperitoneal injection, subcutaneous injection, intravenous injection, oral administration, mucosal administration, intranasal administration, or pulmonary administration.

[0177] In some embodiments, the glycoconjugate, glycosylated fusion protein, or conjugate vaccine according to any one of the preceding claims is used to induce a host immune response against a bacterial pathogen and / or to prevent or treat bacterial diseases and / or infections in a subject.

[0178] The smallest sequence sufficient for O-linked glycosylation

[0179] Traditional chemical conjugate vaccine synthesis has been considered complex, expensive, and labor-intensive (Frasch, CEVaccine 27, 6468-6470 (2009)). However, in vivo conjugation has made significant progress as a viable biosynthetic alternative (Huttner, A. et al., Lancet Infect Dis 17, 528-537 (2017)). The most prominent example of these advances is the success of GlycoVaxyn (now LimmaTechBiologics AG, an independent company directly affiliated with GlaxoSmithKline), a clinical-stage biopharmaceutical company with multiple bioconjugate vaccines at various clinical trial stages, one of which (Flexyn2a) has just completed a Phase 2b challenge study. Although GlycoVaxyn has been at the forefront of the in vivo conjugation revolution, the ability to glycosylate carrier / receptor proteins with polysaccharides containing glucose (Glc) as a reducing terminal sugar has been elusive and, unsurprisingly, has hindered the development of pneumococcal bioconjugation vaccines.

[0180] The oligosaccharide transferase PglS– was previously referred to as PglL by Schulz et al. (PMID 23658772) and as PglL by Harding et al. 2015 (PMID 26727908). ComP – It was only recently characterized as a functional OTase (Schulz, BL et al., PLoS One 8, e62768 (2013)). Subsequent mass spectrometry studies of total glycopeptides demonstrated that PglS does not act as a universal PglL-like OTase, but rather glycosylates multiple periplasmic and outer membrane proteins (Harding, CM et al., Mol Microbiol 96, 1023-1041 (2015)). In fact, the Acinetobacter bengal ADP1 genome encodes two OTases: a PglL-like ortholog (UniProtKB / Swiss-Prot: Q6FFS6.1), which acts as a universal OTase; and PglS (UniProtKB / Swiss-Prot: Q6F7F9.1), which glycosylates a single protein, ComP (Harding, CM et al., Mol Microbiol 96, 1023-1041 (2015)).

[0181] ComP is orthologous to type IV fimbriae, such as PilA from *Pseudomonas aeruginosa* and PilE from *Neisseria meningitidis*, both of which are glycosylated by OTase TfpO (Castric, P. Microbiology 141(Pt 5), 1247-1254(1995)) and PglL (Power, PM et al. Mol Microbiol 49, 833-847(2003)). Although TfpO and PglL also glycosylate their homologous fimbriae at serine residues, the glycosylation sites differ between systems. TfpO glycosylates its homologous fimbriae at a C-terminal serine residue (Comer, JE, Marshall, MA, Blanch, VJ, Deal, CD, and Castric, P. Infect Immun 70, 2837-2845(2002)), a residue not present in ComP. PglL glycosylates PilE at an internal serine residue at position 63 (Stimson, E. et al., Mol Microbiol 17, 1201-1214 (1995)). ComP also contains a serine residue near position 63, and the surrounding residues show moderate conservation for PilE from Neisseria meningitidis. However, comprehensive glycopeptide analysis shows that this serine and the surrounding residues are not glycosylation sites in ComP. PglS glycosylates PilE at a position similar to ComP. 110264 The conserved serine at position 82 of ENV58402.1 (SEQ ID NO:201) corresponds to (and also to ComP) ADP1 ComP is glycosylated at a single serine residue at position 84 (corresponding to a conserved serine residue in AAC4588631 (SEQ ID NO: 202)). This single serine residue is a novel glycosylation site not previously found in the type IV fimbriae superfamily. The ability of PglS to transfer polysaccharides containing glucose as a reducing terminal sugar, coupled with the identification of a novel glycosylation site within the fimbriae superfamily, demonstrates that PglS is a functionally distinct OTase from PglL and TfpO.

[0182] Bioinformatics characteristics of ComP fimbriae orthologs

[0183] ComP was initially described as a factor required for the natural transformation of Acinetobacter bengal ADP1 (Porstendorfer, D., Drotschmann, U., and Averhoff, B. Appl Environ Microbiol 63, 4150-4157 (1997)). In subsequent studies, it was demonstrated that ComP derived from Acinetobacter bengal ADP1 (referred to as ComP in this paper) ADP1 ComP is glycosylated by a novel OTase PglS located downstream of ComP (instead of the universal OTase PglL located elsewhere on the chromosome) (Harding, CM et al., Mol Microbiol 96, 1023-1041 (2015)). ADP1 The protein (NCBI identifier AAC45886.1) belongs to a family of proteins known as type IV fimbriae. Specifically, ComP shares homology with major type IVa fimbriae (Giltner, CL, Nguyen, Y., and Burrows, LL Microbiol Mol Biol Rev 76, 740-772 (2012)). Type IVa fimbriae share high sequence homology at their N-terminus, encoding a highly conserved leader sequence and an N-terminal α-helix; however, the C-terminus shows significant differences between genera and even within species (Giltner, CL, Nguyen, Y., and Burrows, LL Microbiol Mol Biol Rev 76, 740-772 (2012)). To help distinguish ComP orthologs from other type IVa fimbriae, such as PilA from Acinetobacter baumannii, Pseudomonas aeruginosa, and Haemophilus influenzae, and PilE from Neisseria species (Pelicic, V. Mol Microbiol 68, 827-837 (2008)), BLASTp analysis was performed. This analysis will distinguish ComP from other fimbriae. ADP1 The primary amino acid sequence was compared against all proteins from the genus *Acinetobacter*. Unsurprisingly, many orthologs of *Acinetobacter* type IVa fimbriae (including ComP) were found. ADP1 They share high homology at their N-terminus; however, very few proteins exhibit high sequence conservation throughout the entire amino acid sequence of ComP. Based on the relationship with ComP... ADP1 The presence of a conserved serine residue at position 84, and conserved disulfide bonds flanking the predicted glycosylation site (which link the predicted αβ ring to the β chain region), identified at least six ComP orthologs. Figure 20(Giltner, CL, Nguyen, Y. and Burrows, LL. Microbiol, MolBiol Rev. 76, 740-772 (2012)). Furthermore, all six ComP orthologs carry a pglS homolog downstream of the comP gene and a pglL homolog located elsewhere on the chromosome. In summary, at least the following distinguish ComP fimbriae variants from other type IVa fimbriae variants: the presence of a conserved serine residue at position 84, disulfide rings flanking the glycosylation site, the presence of the pglS gene downstream of comP, and the presence of a pglL homolog located elsewhere on the chromosome.

[0184] Therefore, this document discloses the common characteristics of ComP proteins and their identification of ComP orthologs in different Acinetobacter species. ComP proteins can be distinguished from other fimbriae proteins by the following: the presence of a conserved glycosylated serine residue at position 84 relative to the ADP1 ComP protein, and the presence of disulfide rings flanking the glycosylation site. Furthermore, the presence of a pglS homolog downstream of ComP is an indicator of ComP. Further, in order to be classified as a PglS OTase protein rather than a PglL OTase protein, the OTase downstream of ComP must exhibit higher sequence conservation than PglS (ACIAD3337) when compared to PglL (ACIAD0103) in Acinetobacter benzia ADP1. It will be apparent to those skilled in the art that, in any embodiment of this disclosure, the ComP protein comprises the same homolog as SEQ ID NO:201 (ComP 110264 The conserved serine residue at position 82 of ENV58402.1 corresponds to the serine residue and can be glycosylated on the serine residue.

[0185] ComP protein glycosylation fragments

[0186] In some embodiments, the ComP glycosylation fragment can be any ComP protein disclosed herein, or can be derived from or originate from any ComP protein disclosed herein. Further, PglS OTase can be any of the following.

[0187] It has been previously demonstrated that a PglS ortholog from *Acinetobacter bengal* strain ADP1 glycosylates a ComP ortholog from *Acinetobacter tumefaciens* strain CIP 110264 at a single serine residue at position 82 (Harding, CM et al., 2019; WO / 2019 / 241672, which is incorporated herein by reference in its entirety). PglS was engineered to functionally glycosylate heterologous proteins by translatorily fusing a large fragment (117 amino acids) of ComP to the C-terminus of a known carrier protein. Specifically, the 117-amino acid fragment of ComP... 110264 The fragment is fused at the C-terminus of genetically inactivated exotoxin A (EPA) from Pseudomonas aeruginosa to a flexible GGGS linker (SEQ ID NO: 382). This chimeric vector protein also possesses an N-terminal DsbA signal sequence (ssDsbA) for translocation to the periplasm via the Sec pathway, and a C-terminal hexahistine tag for detection.

[0188] Even shorter ComP glycosylation fragments have been identified, which are sufficient for PglS glycosylation (WO / 2020 / 131236, which is incorporated herein by reference in its entirety). Studies have shown that ComP glycosylation fused to the C-terminus of EPA carrier proteins... 110264 Glycosylated fragments can also be glycosylated by PglS, but only if the ComP glycosylated fragment is relative to ComP. 110264 It also contains cysteine ​​residues corresponding to Cys71 and Cys93. These observations were confirmed in a series of experiments aimed at identifying even shorter ComP glycosylated fragments. Figures 12A and 12B show ComP... 110264Fragments were designed to shift one amino acid relative to serine 82 from the N-terminus to the C-terminus, where serine 82 is the PglS glycosylation site when the ComP glycosylated fragment is fused to the C-terminus of the EPA carrier protein. The ComP glycosylated fragment was amplified by PCR, cloned to the C-terminus of EPA, and bioconjugation was performed using PglS. For these experiments and all experiments described below, serotype 8 pneumococcal capsular polysaccharide (CPS8) expressed from the pB-8 plasmid was used as the glycan source (Kay, EJ et al., 2016). CPS8 was chosen because it contains glucose as a reducing terminal sugar and has been previously shown to be efficiently transferred to ComP via PglS (Harding, CM et al., 2019). Furthermore, for these experiments and all experiments described below, bioconjugation was performed using *E. coli* strain SDB1. SDB1 deletions include WecA (which initiates the biosynthesis of intestinal bacterial common antigens and O-antigen polysaccharides) and WaaL (which transfers undecyprene pyrophosphate-linked glycan precursors to the outer core of lipid A) (Garcia-Quintanilla, F. et al., 2014). Overall, these mutations promote the accumulation of heterologously expressed lipid-linked glycan precursors (such as the CPS8 glycan lipid-linked precursor), which are available exclusively to PglS. The expression of CPS8 glycan, PglS, and the fusion EPA-ComP from the IPTG inducible vector was compared. 110264 The SDB1 strain of the construct was cultured in LB broth, induced in mid-log phase, and grown overnight. Samples were harvested approximately 20 hours after induction for Western blot analysis of periplasmic extracts to evaluate EPA-ComP. 110264 Expression of the fusion protein and protein glycosylation. Western blot analysis was performed using an anti-EPA antibody (anti-EPA) and an anti-hexahistine tag antibody (anti-His). Probe with both antibodies allowed determination of whether the EPA protein and / or the C-terminal ComP fragment remained intact.

[0189] Figures 12C, 12D, and 12E reiterate that when the ComP glycosylation fragment is fused to the C-terminus, ComP… 110264 The presence of Cys71 and Cys93 residues flanking Ser82 is important for EPA-ComP 110264Glycosylation is crucial. As shown in Figures 12C, 12D, and 12E, fusion proteins containing ComP glycosylated fragments lacking Cys71 or Cys93 were not glycosylated. Transfer of CPS8 glycans was observed only in fusion proteins containing ComP glycosylated fragments with both cysteine ​​residues. For all fusion proteins containing ComP glycosylated fragments with both Cys71 and Cys93, the glycosylation efficiency and average number of PglS-transferred CPS8 repeat units were similar. Careful examination of the protein blot revealed that, compared to anti-EPA signaling (Figure 12D), chimeric EPA-ComP... 110264 Variants (listed as C2, D2, E3, and F3 in Figures 12C, 12D, and 12E) showed almost no reaction with anti-His antibodies. Furthermore, the anti-EPA channel showed that it reacted with unglycosylated EPA-ComP containing both Cys71 and Cys93. 110264 Compared to other variants, these variants migrated at a slightly lower molecular weight (Figure 12C). In summary, these observations suggest that the ComP fragment lacking two cysteine ​​residues is unstable and may be prone to C-terminal degradation, thereby preventing glycosylation of PglS. Unbound by theory, it is believed that Cys71 and Cys93 can stabilize ComP by forming covalent disulfide bridges. 110264 .

[0190] Various proteins from different organisms, usually inactivated bacterial toxins, have been used as carriers for conjugates and bioconjugate vaccines. Cross-reactive material 197 (CRM) 197 ) is the genetically inactivated form of diphtheria toxin, which has been widely used as a carrier protein in various conjugate vaccines against pneumococcus, meningococcus, and Haemophilus influenzae type b (Berti, F. and Adamo, R., 2018). Given CRM 197 Due to its frequent use in conjugated vaccine formulations, the PglS bioconjugation system has been extended to include CRM. 197 They work together. For these experiments, the previously identified 25 amino acid "C1" ComP glycosylation fragment (ComP) C1 Integrating into CRM through translation 197 The C-terminus is linked via a GGGS sequence (SEQ ID NO:382). The SRP-dependent FlugI secretion sequence (ssFlgI) is added to the CRM. 197 The N-terminus was tagged to facilitate its output to the periplasm (Goffin, P. et al., 2017). Finally, a C-terminal hexahistine tag was added to aid purification (Figure 13A). The expressed CPS8 glycan, along with PglS and CRM, was then processed. 197 -ComP C1E. coli SDB1 cells with the vector (expected size 61.8 kDa) were cultured in shake flasks and harvested after 24 hours. CRM was purified by three rounds of sequential chromatography. 197 -ComP C1 -CPS8 glycoconjugates. First, nickel affinity chromatography was employed because the glycoconjugates contain a C-terminal hexahistine tag. Fractions containing the glycoconjugates were combined, and the glycosylated glycoconjugates were enriched using a MonoQ column with elution using a linear salt gradient. A final purification step was performed on a Superdex 200 Increase column to remove large aggregates. Anti-CRM chromatography was used as shown in Figures 13B, 13C, and 13D. 197 Western blot analysis of purified samples with pneumococcal CPS8 antiserum demonstrated that CRM 197 -ComP C1 The vector was glycosylated with CPS8. The purified glycoconjugate was digested with proteinase K before separation on SDS-PAGE, resulting in CRM. 197 The complete loss of polysaccharide-specific signals indicates that the CPS8 polymer is covalently linked to the CRM. 197 -ComP C1 protein.

[0191] Next, ComP was tested. C1 Can glycosylation tags be moved to CRM? 197 Another site of the fusion. Therefore, we designed a new construct that incorporates ComP C1 Placed in CRM 197 The N-terminus of the coding region (Figure 14A). The FlgI secretion signal is placed in ComP... C1 The N-terminus of the glycosylated fragment was targeted with hexahistine for CRM. 197 C-terminal labeling was performed. The expression CPS8 glycan, as well as PglS and ComP, were then used. C1 -CRM 197 The *E. coli* SDB1 cells carrying the vector were cultured in shake flasks and harvested after 24 hours. As shown in Figure 14B, Western blot analysis of the periplasmic extract, detected with an anti-His antibody, indicated that ComP... C1 -CRM 197 It was also glycosylated by PglS. The average number of CPS8 repeating units and the glycosylation efficiency of the two fusions were comparable, indicating that ComP C1 Glycosylation tags can be placed at the N-terminus or C-terminus of a carrier protein.

[0192] ComP, containing 11 amino acids sufficient for PglS glycosylation. 110263 Sequence identification

[0193] Although previous reports indicated that Cys71 and Cys93 contain ComPs that are translated into and fused to the C-terminus of the EPA. 110264 Glycosylated fragments are essential for glycosylation in fusion proteins (e.g., Figures 12C, 12D, and 12E), but these data do not determine whether the two cysteine ​​residues and the hypothetical disulfide bridge formed between them are necessary for PglS glycosylation in all cases. N-linked sequencers recognized by PglB have been engineered into multiple sites on the surface loop of EPA and used as “internal” glycosylation tags (Ihssen, J. et al., 2010). To determine ComP... 110264 Whether Cys71 and Cys93 are necessary for PglS glycosylation will depend on the entire 23 amino acids of ComP spanning from Cys71 to Cys93. 110264 Glycosylated fragments - (referred to as iGT in this paper) CC This is used for internal glycosylation tags (cysteine-cysteine) to integrate into the EPA amino acid sequence. ComP 110264 iGT CC The EPA residues Ala489 and Arg490 are inserted between the EPA residues, which are located in a β-turn structure on the surface of the catalytic domain (Fig. 15A). As a control, iGT is also integrated. CC A variant of the ComP glycosylation fragment, iGTss (“serine-serine”), contains serine residues replacing cysteine ​​residues at positions 71 and 93 of ComP. This iGTSS ComP glycosylation fragment is also integrated between residues Ala489 and Arg490 of EPA. It is speculated that the serine residues have a similar spatial volume to the cysteine ​​residues but cannot be oxidized to form disulfide bonds (Figure 15B). As described above, the transfer of CPS8 to EPA by PglS was evaluated in a three-plasmid system. iGTcc or EPA iGTss The capability. As shown in Figures 15C and 15D, EPA iGT Both the cysteine-cysteine ​​and serine-serine variants were glycosylated, demonstrating that Cys71 and Cys93 (and the presumed disulfide bond formed between them) are not necessary for PglS glycosylation when the ComP fragment is introduced into the EPA protein.

[0194] Since cysteine ​​residues are not required for PglS-dependent glycosylation only when the ComP glycosylation fragment is integrated into the fusion protein, it was hypothesized that a shorter ComP glycosylation fragment, representing a minimal O-linked ComP sequence, could be found within the 23-amino acid ComP glycosylation fragment spanning Cys71 to Cys93. To investigate this, an iGT sequence integrated between EPA residues Ala489 and Arg490 was generated. CC Shorter variants of the ComP glycosylation fragment were used to identify which ComP residues are essential for glycosylation. From the 23 amino acid iGT... CC Alternating deletion of single amino acids on both sides of the glycosylation site generates 22 truncated variants, each containing Ser82, the site of PglS glycosylation (Fig. 16A and Fig. 16B). These variants are derived from iGT CC The number of residues missing on both sides is used for naming; for example, Δ3-4 corresponds to the number of residues missing from iGT. CC The N-terminus of the EPA-iGT was deleted, along with three amino acids, and the C-terminus was deleted, resulting in the shortest variant with a length of five amino acids. These truncated EPA-iGT variants were tested in shake flasks under the same conditions as described above. CC The variants were bioconjugated with CPS8 and PglS. As a negative control, we included a construct that expressed only the EPA-coding sequence and carried DsbA secretion and a six-histidine tag.

[0195] Figure 16C shows that robust glycosylation was observed in all EPA fusion proteins containing ComP glycosylation fragments of at least 11 amino acids in length. The glycosylation ratio is compared with the 23-amino acid iGT. CCThe comparable ComP glycosylation fragments suggest that moderate truncation on both sides of Ser82 does not significantly affect the glycosylation efficiency of PglS. Although these fusion proteins are glycosylated, a slight decrease in glycosylation efficiency was observed as the amino acid sequence of the iGT ComP glycosylation fragment was shortened. The shortest highly glycosylated inner ComP glycosylation fragment is iGTΔ6-6 with the sequence IASGASAATTN (SEQ ID NO:309); Figure 16C). Removal of the N-terminal isoleucine residue (iGTΔ7-6; SEQ ID NO:321) or the C-terminal asparagine residue (iGTΔ6-7; SEQ ID NO:310) significantly reduced the glycosylation efficiency of the carrier protein, indicating that these residues play an important role in PglS glycosylation. Variants smaller than iGTΔ6-6 mostly showed minimal glycosylation, with iGTΔ7-6 with the sequence ASGASAATTN (SEQ ID NO:321) being the best. Interestingly, a small number of higher molecular weight stepwise changes were also observed in the fusion proteins containing the smallest ComP glycosylation fragments iGTΔ9-8 (SEQ ID NO:346) and iGTΔ9-9 (SEQ ID NO:347) (Fig. 16D), indicating that these six- and five-amino acid variants were glycosylated by PglS at very low levels. This implies that ComP recognized by PglS... 110264 Glycosylated sequences can be as small as five amino acids.

[0196] Next, Ni affinity chromatography was used to purify the CPS8-glycosylated EPA fusion protein containing the iGTΔ6-6ComP glycosylation fragment located between residues Ala489 and Arg490 from whole-cell lysates, and Western blot analysis was performed on the eluent using antiserum specific to EPA protein or CPS8 glycan. The results of these experiments clearly demonstrate that the EPA fusion protein containing the iGTΔ6-6ComP glycosylation fragment located between residues Ala489 and Arg490 is glycosylated by PglS with CPS8 (Figs. 17A, 17B, and 17C). Overall, these experiments indicate that ComP... 110264 Glycosylation fragments can be derived from ComPs of 117 amino acids. 110264 The sequences were shortened to 11 amino acids or less while maintaining glycosylation. These results unexpectedly suggest that the previously confirmed essential glycosylation with ComP during C-terminal fusion is possible. 110264 The cysteine ​​residues corresponding to Cys71 and Cys93 are not necessary for PglS-dependent glycosylation when the ComP glycosylated fragment is integrated into the fusion protein.

[0197] The aforementioned iGT truncated series was tested at an internal site between residues Ala489 and Arg490 on EPA. Next, a second site was tested between EPA residues Glu548 and Gly549, incorporating an iGTΔ3-4ComP glycosylation fragment (SEQ ID NO: 271). Similar to the first site, the second site is located on a surface-exposed loop within the catalytic domain of EPA. Variants with alternating tags for bioconjugation with CPS8 and PglS were tested under the same conditions as the other truncated sites. The glycosylation efficiency of this construct with CPS8 was observed to be similar to that when iGTΔ3-4 was placed at the first site on EPA. CPS8 glycosylated EPA fusion proteins containing the iGTΔ3-4ComP glycosylation fragment located between residues Glu548 and Gly549 were then purified from whole-cell lysates using Ni affinity chromatography, and the eluent was analyzed by Western blot using antiserum specific for EPA protein or CPS8 glycan. These experimental results further demonstrate that the EPA fusion protein containing the iGTΔ3-4ComP glycosylation fragment located between residues Glu548 and Gly549 is glycosylated by PglS with CPS8. Overall, these experiments indicate that ComP... 110264 Glycosylation fragments can be derived from ComPs of 117 amino acids. 110264 The sequences were shortened to 11 amino acids or less while maintaining glycosylation. These results unexpectedly suggest that, with ComP... 110264 The cysteine ​​residues corresponding to Cys71 and Cys93 are not necessary for PglS-dependent glycosylation when the ComP glycosylated fragment is integrated into the fusion protein.

[0198] This document provides glycoconjugates comprising oligosaccharides or polysaccharides linked to a fusion protein. In some embodiments, the oligosaccharide or polysaccharide is covalently linked to the fusion protein. The fusion protein comprises a glycosylated fragment of a ComP protein (as described in detail elsewhere herein). In some embodiments of the glycoconjugates disclosed herein, the oligosaccharide or polysaccharide comprises glucose at its reducing end.

[0199] ComP undergoes glycosylation at a serine (S) residue. This serine residue is associated with SEQ ID NO:201 (ComP). 110264 This corresponds to position 82 of SEQ ID NO: 202 (ComP). This serine residue is conserved in ComP proteins and, for example, corresponds to position 82 of SEQ ID NO: 202 (ComP). ADP1 Position 84 of (AAC45886.1) corresponds to this. Therefore, in some respects, the fusion protein (and thus the glycoconjugate) corresponds to SEQ ID NO:202 (ComP) on the glycosylated fragment of ComP. ADP1The serine residue at position 84 of SEQ ID NO: 201 (ComP) corresponds to or is related to the serine residue at position 84 of SEQ ID NO: 201 (ComP). 110264 Glycosylation of the serine residue at position 82 of ENV58402.1 with oligosaccharides or polysaccharides. Figure 22 The image shows an alignment of the ComP sequence containing the serine (S) residue (boxed out), which corresponds to SEQ ID NO:201 (ComP). 110264 The serine residue at position 82 of ENV58402.1 corresponds to the latter, which is conserved in the ComP sequence.

[0200] Those skilled in the art will recognize that, by comparing the ComP sequence with SEQ ID NO:201 (e.g., the complete or partial sequence), conserved serine residues in non-SEQ ID NO:201 ComP proteins can be identified, corresponding to the serine residue at position 82 of SEQ ID NO:201. Furthermore, those skilled in the art will recognize that, by comparing the ComP sequence with SEQ ID NO:201, other residues, regions, and / or features corresponding to the residues, regions, and / or features of SEQ ID NO:201 mentioned herein can be identified in non-SEQ ID NO:201 ComP sequences, and reference is made to SEQ ID NO:201. And, while SEQ ID NO:201 is generally referred to herein, by analogy any residue, region, feature, etc., of any ComP sequence disclosed herein can be similarly referred to, for example, SEQ ID NO:202.

[0201] ComP proteins are proteins that have been identified as consistent with the descriptions provided herein. Representative examples of ComP proteins include, but are not limited to: AAC45886.1 ComP [Acinetobacter ADP1]; ENV58402.1 putative protein F951_00736 [Acinetobacter agronomycetes CIP 110264]; APV36638.1 competent protein [Acinetobacter agronomycetes GFJ-2]; PKD82822.1 competent protein [Acinetobacter radioresistens 50v1]; SNX44537.1 type IV fimbriae assembly protein PilA [Acinetobacter puyangensis ANC4466]; OAL75955.1 competent protein [Acinetobacter SFC]; ComP P5312 ; and ComP ANT_H59In some respects, the ComP protein contains components similar to SEQ ID NO:201 (ComP ADP1 ) or SEQ ID NO:201(ComP 110264 The amino acid sequence is at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical, and contains a serine residue corresponding to a conserved serine residue at position 84 of SEQ ID NO:202 or position 82 of SEQ ID NO:201. SEQ ID NO:202 contains a 28-amino acid leader sequence. In some respects, the ComP protein contains a sequence identical to that in SEQ ID NO:210 (ComPΔ28). ADP1 ), SEQ ID NO:209(ComPΔ28 110264 ), SEQ ID NO:211(ComPΔ28 GFJ-2 ), SEQ ID NO:212(ComPΔ28 P50v1 ), SEQ ID NO:213(ComPΔ28 4466 ), SEQ ID NO:214(ComPΔ28 SFC ), SEQ ID NO:215(ComPΔ28 P5312 ) or SEQ ID NO:216(ComPΔ29 ANT_H59 At least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical amino acid sequences, which do not contain an amino acid leader sequence but contain an amino acid sequence identical to SEQ ID NO:201(ComP) 110264 The serine residue at position 82 of SEQ ID NO: 209 (ComPΔ28) corresponds to a conserved serine residue. In some respects, the ComP protein contains a serine residue corresponding to the conserved serine residue at position 82 of SEQ ID NO: 209 (ComPΔ28). 110264 At least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical amino acid sequence, which does not contain a 28-amino acid leader sequence but contains an amino acid sequence identical to SEQ ID NO:201(ComP) 110264 The conserved serine residue at position 82 corresponds to the serine residue in SEQ ID NO:210 (ComPΔ28). In some respects, the ComP protein contains SEQ ID NO:210 (ComPΔ28). ADP1 ), SEQ ID NO:209(ComPΔ28 110264 ), SEQ ID NO:211(ComPΔ28 GFJ-2), SEQ ID NO:212(ComPΔ28 P50v1 ), SEQ ID NO:213(ComPΔ28 4466 ), SEQ ID NO:214(ComPΔ28 SFC ), SEQ ID NO:215(ComPΔ28 P5312 ) or SEQ ID NO:216(ComPΔ29 ANT_H59 In some respects, the ComP protein is SEQ ID NO:202(ComP). ADP1 :AAC45886.1), SEQ ID NO:201(ComP 110264 :ENV58402.1), SEQ ID NO:203(ComP GFJ-2 :APV36638.1), SEQ ID NO:204(ComP 50v1 :PKD82822.1), SEQ ID NO:205(ComP 4466 :SNX44537.1), SEQ ID NO:206(ComP SFC :OAL75955.1), SEQ ID NO:207(ComP P5312 ) or SEQ ID NO:208(ComP ANT_H59 ).

[0202] This document provides a glycoconjugate comprising an oligosaccharide or polysaccharide covalently linked to a fusion protein, wherein the fusion protein comprises a ComP protein (ComP) glycosylated fragment. In some embodiments, the ComP glycosylated fragment does not contain a ComP protein. 110264 (SEQ ID NO:201) at position 71, the conserved cysteine ​​(C) residue corresponds to the cysteine ​​(C) residue. In some embodiments, the ComP glycosylation fragment does not contain the cysteine ​​(C) residue corresponding to ComP. 110264 The conserved cysteine ​​(C) residue at position 93 of (SEQ ID NO:201) corresponds to the cysteine ​​(C) residue. As described in more detail herein, the fusion protein corresponds to the ComP glycosylation fragment. 110264 The conserved serine residue at position 82 of (SEQ ID NO:201) is glycosylated with an oligosaccharide or polysaccharide at the corresponding serine residue. In some embodiments, the ComP glycosylated fragment is internally located within the fusion protein. Further, in some embodiments, the ComP glycosylated fragment of the fusion protein is partially exposed to a solvent (or surface) and / or integrated into the C-terminal region of the fusion protein. 10β-turn, β-twist, β-loop, U-turn, reverse turn, chain reversal, or hairpin loop.

[0203] Since it has been found that when the ComP glycosylation fragment is located internally within a fusion protein, its glycosylation does not require flanking cysteine ​​residues, the ComP glycosylation fragments disclosed herein may be shorter than previously thought. In some embodiments, the length of the ComP glycosylation fragment may be shorter than 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, or 6 amino acids, as long as it contains amino acids related to ComP. 110264 The ComP glycosylated fragment may be any serine residue corresponding to the conserved serine residue at position 82 of (SEQ ID NO:201). In some embodiments, the length of the ComP glycosylated fragment is any from 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 amino acids to 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 or 22 amino acids. In some embodiments, the fragment is a ComP protein having at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or 11 serine residues located at the N-terminus of the conserved serine residue at position 82 of SEQ ID NO:201, for example, X. n S[Y], where n is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 amino acid residues of the ComP protein. In some embodiments, the fragment is a ComP protein having at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 serine residues located at the C-terminus corresponding to the conserved serine residue at position 82 of SEQ ID NO: 201, for example, [X]SY n Where n is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 amino acid residues of the ComP protein. Further, in some embodiments, the amino acid sequence of the ComP glycosylated fragment extends in the N-terminal direction no more than the number of amino acid residues in the ComP protein. 110264 The amino acid residue corresponding to position 72 of (SEQ ID NO:201), and / or extending beyond the C-terminus to no more than ComP 110264 The amino acid residue corresponding to position 92 of (SEQ ID NO:201).

[0204] Consistent with the ComP protein of this disclosure, in some embodiments, the ComP protein from which the ComP glycosylation fragment is derived comprises the same as SEQ ID NO:209(ComPΔ28). 110264 )SEQ ID NO:210(ComPΔ28 ADP1), SEQ ID NO:211(ComPΔ28 GFJ-2 ), SEQ ID NO:212(ComPΔ28 P50v1 ), SEQ ID NO:213(ComPΔ28 4466 ), SEQ ID NO:214(ComPΔ28 SFC ); SEQ ID NO:215(ComPΔ28 P5312 ) or SEQ ID NO:216(ComPΔ29 ANT_H59 At least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical amino acid sequences. In some embodiments, the ComP protein from which the ComP glycosylated fragment is derived comprises SEQ ID NO:209(ComPΔ28). 110264 ), SEQ ID NO:210(ComPΔ28 ADP1 ), SEQ ID NO:211(ComPΔ28 GFJ-2 ), SEQ ID NO:212(ComPΔ28 P50v1 ), SEQ ID NO:213(ComPΔ28 4466 ), SEQ ID NO:214(ComPΔ28 SFC ); SEQ ID NO:215(ComPΔ28 P5312 ) or SEQ ID NO:216(ComPΔ29 ANT_H59 ).

[0205] In some embodiments of the glycoconjugates disclosed herein, the ComP glycosylation fragment comprises or consists of the following common amino acid sequence:

[0206] or

[0207]

[0208] Where: X1 is V, T, A or I;

[0209] X4 can be Q, T, E, A, or S;

[0210] X5 is E, Q, T, or L;

[0211] X6 is either I or V;

[0212] X7 is S, N, A, or G;

[0213] X8 is S or contains no amino acids;

[0214] X9 is G, D or contains no amino acids;

[0215] X 12 For N, S, or A;

[0216] X 13 It can be A, S, or K;

[0217] X 15 For T, S, or K;

[0218] X 18 It can be A, E, Q or L;

[0219] X 19 For T, S, or K;

[0220] X 20 It is either A or S; and

[0221] X 21 For T, Q, A, or V;

[0222] Or a fragment of at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 amino acids, the fragment containing serine corresponding to position 11 of SEQ ID NO:217. Residues. In some embodiments, the fragment contains serine at position 11 corresponding to SEQ ID NO: 217. The N-terminus of the residue has at least 1, 2, 3, 4, 5, 6, 7, or 8 amino acid residues. In some embodiments, the fragment contains serine at position 11 corresponding to SEQ ID NO:217. The C-terminus of the residue has at least 1, 2, 3, 4, 5, 6, 7, or 8 amino acid residues. However, the ComP glycosylated fragment does not contain any amino acid residues related to ComP. 110264 (SEQ ID NO:201) at position 71, the conserved cysteine ​​(C) residue corresponds to the cysteine ​​(C) residue, and / or the ComP glycosylation fragment does not contain the cysteine ​​(C) residue corresponding to ComP. 110264 The conserved cysteine ​​(C) residue at position 93 of (SEQ ID NO:201) corresponds to the cysteine ​​(C) residue.

[0223] Some embodiments provide a ComP glycosylation fragment, which is a variant of the common amino acid sequence of SEQ ID NO:217, SEQ ID NO:396, or SEQ ID NO:397, or a fragment thereof, having 1, 2, 3, 4, 5, 6, or 7 amino acid substitutions, additions, and / or deletions, wherein the variant maintains the serine corresponding to position 11 of SEQ ID NO:217. residues, and wherein the variant does not contain ComP110264 (SEQ ID NO:201) at position 71, corresponding to the conserved cysteine ​​(C) residue, and / or this variant does not contain a cysteine ​​(C) residue corresponding to ComP. 110264 The conserved cysteine ​​(C) residue at position 93 of (SEQ ID NO:201) corresponds to the cysteine ​​(C) residue. Those skilled in the art will understand that the number of tolerable amino acid substitutions, additions, and / or deletions within a sequence may depend on the sequence length without affecting function (e.g., the ability to function as a sequence segment). For example, a six-amino acid-long sequence tolerates fewer variations compared to a 21-amino acid-long sequence.

[0224] Whether a ComP glycosylated fragment (including subfractions and variants of the fragments disclosed herein, collectively referred to as ComP glycosylated fragments) can be glycosylated, and the efficiency of glycosylation, can be determined, such as by the methods described herein. In some embodiments, ComP glycosylated fragments may be glycosylated when internally located within the fusion protein and / or internally located within the carrier protein sequence (as described elsewhere herein). Further, in some embodiments, ComP glycosylated fragments or variants are not glycosylated when located at the N-terminus and / or C-terminus of the fusion protein, or their degree of glycosylation is reduced by at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% when located at the N-terminus and / or C-terminus of the fusion protein compared to when internally located within the fusion protein.

[0225] In some embodiments, the fusion protein comprises a carrier protein selected from the group consisting of: Pseudomonas aeruginosa exotoxin A (EPA), CRM 197 The cholera toxin B subunit, tetanus toxin C fragment, Haemophilus influenzae protein D, and one or more fragments thereof. For example, in some embodiments, the Pseudomonas aeruginosa exotoxin A (EPA) carrier protein comprises the amino acid sequence of SEQ ID NO:218 or one or more fragments thereof. For example, in some embodiments, CRM... 197 The carrier protein contains the amino acid sequence of SEQ ID NO:224 or one or more fragments thereof.

[0226] As can be understood from this disclosure as a whole, "internally located within the fusion protein" means that the ComP fusion protein is not located at the C-terminus or N-terminus of the fusion protein, and does not include any C-terminal leader sequence or N-terminal tag (e.g., His tag).

[0227] In some embodiments, the ComP glycosylated fragment can be linked to the carrier protein sequence via an amino acid linker.

[0228] Furthermore, in some embodiments, the ComP glycosylation fragment may be inserted into the sequence of the carrier protein, rather than between the carrier proteins. For example, in some embodiments:

[0229] (i) The ComP glycosylation fragment is inserted between Ala489 and Arg490 relative to the PDB entity 1IKQ of Pseudomonas aeruginosa exotoxin A (EPA) (SEQ ID NO:219);

[0230] (ii) The ComP glycosylated fragment is inserted between Glu548 and Gly549 relative to Pseudomonas aeruginosa exotoxin A (EPA) (SEQ ID NO:220);

[0231] (iii) The ComP glycosylated fragment is inserted between Ala122 and Gly123 relative to Pseudomonas aeruginosa exotoxin A (EPA) (SEQ ID NO:221);

[0232] (iv) The ComP glycosylated fragment is inserted between Thr355 and Gly356 relative to Pseudomonas aeruginosa exotoxin A (EPA) (SEQ ID NO: 222) PDB entity 1IKQ; or

[0233] (v) The ComP glycosylated fragment is inserted between Lys20 and Asp21 relative to Pseudomonas aeruginosa exotoxin A (EPA) (SEQ ID NO:223) PDB entity 1IKQ.

[0234] Furthermore, in some embodiments, the ComP glycosylation fragment may be inserted into the sequence of the carrier protein, rather than between the carrier proteins. For example, in some embodiments:

[0235] (i) relative to CRM 197 The PDB entity 4AE0 of (SEQ ID NO:225) has a ComP glycosylation fragment inserted between Asn481 and Gly482;

[0236] (ii) relative to CRM 197 The PDB entity 4AE0 of (SEQ ID NO:226) has a ComP glycosylation fragment inserted between Asp392 and Gly393;

[0237] (iii) Compared to CRM 197 The PDB entity 4AE0 of (SEQ ID NO:227) has a ComP glycosylation fragment inserted between Glu142 and Gly143;

[0238] (iv) Relative to CRM197 The PDB entity 4AE0 of (SEQ ID NO:228) has a ComP glycosylated fragment inserted between Asp129 and Gly130; or

[0239] (v) Relative to CRM 197 The PDB entity 4AE0 of (SEQ ID NO:229) has a ComP glycosylated fragment inserted between Asn69 and Glu70.

[0240] In some embodiments, the ComP glycosylation fragment may be located between the carrier proteins, or it may be inserted into the sequence of the carrier protein within the fusion protein. In some embodiments, the ComP glycosylation fragment may be located internally, and one or more ComP glycosylation fragments may be located at C-termini and / or N-termini sufficient for glycosylation to occur at such locations.

[0241] One aspect of this disclosure is that the fusion protein can be designed to include multiple ComP glycosylation fragments to enhance the immunogenicity of the glycosylated fusion protein / glycoconjugate. In some embodiments, the fusion protein includes two or more, three or more, four or more, five or more, six or more, eight or more, ten or more, fifteen or more, or twenty or more ComP glycosylation fragments. In some embodiments, the fusion protein does not include more than three, more than five, more than ten, more than fifteen, more than twenty, or more than twenty-five ComP glycosylation fragments. The identity of the ComP glycosylation fragments can also be controlled. For example, in some embodiments, multiple ComP glycosylation fragments of the fusion protein are identical. In some embodiments, the ComP glycosylation fragments of the fusion protein are different from each other. For example, in some embodiments, at least three, at least four, or at least five of the ComP glycosylation fragments of the fusion protein are different from each other. For example, in some embodiments, all ComP glycosylation fragments of the fusion protein are different.

[0242] In some embodiments, the oligosaccharide or polysaccharide is derived from sugars produced by bacteria of the genus Streptococcus. For example, in some embodiments, the sugar is a capsular polysaccharide of Streptococcus pneumoniae, Streptococcus agalactiae (S. agalactiae), or Streptococcus suis; in some embodiments, the sugar is a capsular polysaccharide of serotype 8 Streptococcus pneumoniae; and in some embodiments, the sugar is a capsular polysaccharide of type Ia, Ib, II, III, IV, V, VI, VII, VIII, or X of Streptococcus agalactiae.

[0243] In some embodiments, the oligosaccharide or polysaccharide is derived from sugars produced by bacteria of the genus *Klebsiella*. For example, in some embodiments, the sugar is a capsular polysaccharide of *Klebsiella pneumoniae*, *K. varricola*, *K. michinganenis*, or *K. oxytoca*; and in some embodiments, the sugar is an O-antigen polysaccharide of *Klebsiella pneumoniae*, *K. varricola*, *K. michinganenis*, or *K. oxytoca*.

[0244] In some embodiments, the glycoconjugate is produced in vivo, for example: in bacterial cells; in Escherichia coli; in bacteria from the genus Klebsiella; and / or wherein the bacterial species is Klebsiella pneumoniae, Klebsiella heterotropha, Klebsiella micrantha, or Klebsiella acidogenic.

[0245] This article provides a glycoconjugate as described above (e.g., a ComP glycosylated fragment that does not contain ComP). 110264 (SEQ ID NO:201) at position 71, the conserved cysteine ​​(C) residue corresponds to the cysteine ​​(C) residue, and / or the ComP glycosylation fragment does not contain the cysteine ​​(C) residue corresponding to ComP. 110264 (SEQ ID NO:201) at position 93, a conserved cysteine ​​(C) residue corresponding to the cysteine ​​(C) residue, wherein the ComP glycosylated fragment contains or is composed of the amino acid sequence of SEQ ID NO:232-363 or 364. This document provides a glycoconjugate as described above (e.g., the ComP glycosylated fragment does not contain the cysteine ​​(C) residue corresponding to the conserved cysteine ​​(C) residue at position 93 of ComP). 110264 (SEQ ID NO:201) at position 71, the conserved cysteine ​​(C) residue corresponds to the cysteine ​​(C) residue, and / or the ComP glycosylation fragment does not contain the cysteine ​​(C) residue corresponding to ComP. 110264 (SEQ ID NO:201) at position 93, the conserved cysteine ​​(C) residue corresponds to the cysteine ​​(C) residue, wherein the ComP glycosylated fragment contains or is composed of the following amino acid sequence:

[0246]

[0247]

[0248] or

[0249]

[0250] This document also provides a ComP glycosylation fragment, which is a variant of any of the ComP glycosylation fragments disclosed above, having 1, 2, 3, 4, 5, 6, or 7 amino acid substitutions, additions, and / or deletions, wherein the variant maintains the serine (S) residue corresponding to the conserved serine residue at position 82 of SEQ ID NO:201, and wherein the variant does not contain a ComP glycosylation fragment. 110264 (SEQ ID NO:201) at position 71, corresponding to the conserved cysteine ​​(C) residue, and / or this variant does not contain a cysteine ​​(C) residue corresponding to ComP. 110264 The conserved cysteine ​​(C) residue at position 93 of (SEQ ID NO:201) corresponds to the cysteine ​​(C) residue.

[0251] Whether a ComP glycosylated fragment (including subfractions and variants of the fragments disclosed herein, collectively referred to as ComP glycosylated fragments) can be glycosylated, and the efficiency of glycosylation, can be determined, as described herein. In some embodiments, ComP glycosylated fragments may be glycosylated when internally located within a fusion protein and / or internally located within a carrier protein sequence (as described elsewhere herein). Further, in some embodiments, ComP glycosylated fragments are not glycosylated when located at the N-terminus and / or C-terminus of a fusion protein, or their degree of glycosylation is reduced by at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% when located at the N-terminus and / or C-terminus of a fusion protein compared to when internally located within the fusion protein.

[0252] In some embodiments, the glycoconjugate is a conjugate vaccine. Therefore, this disclosure relates to and provides a conjugate vaccine in some embodiments. In some embodiments, the conjugate vaccine is a vaccine against Streptococcus pneumoniae serotype 8. In some embodiments, the conjugate vaccine induces an immune response when administered to a subject. In some embodiments, the immune response triggers long-term memory (memory B cells and T cells), is an antibody response, and optionally a serotype-specific antibody response. In some embodiments, the antibody response is an IgG or IgM response. In some embodiments, the antibody response is an IgG response; optionally, an IgG1 response. Furthermore, in some embodiments, the conjugate vaccine generates immune memory in the subject who has received the vaccine.

[0253] Although a glycoconjugate comprising a fragment of ComP glycosylation containing a separated fragment of the ComP protein has been described above, it should be understood that this disclosure also explicitly provides a ComP glycosylation fragment consistent with any and all descriptions of the ComP glycosylation fragment provided anywhere herein (including in the appended claims), for example, wherein the ComP glycosylation fragment does not contain a fragment of ComP glycosylation containing a separated fragment of the ComP protein. 110264(SEQ ID NO:201) contains a cysteine ​​residue corresponding to the conserved cysteine ​​residue at position 71, and / or does not contain a cysteine ​​residue corresponding to ComP. 110264 (SEQ ID NO:201) at position 93, the conserved cysteine ​​residue corresponds to the cysteine ​​residue, and the ComP glycosylation fragment contains the cysteine ​​residue corresponding to ComP. 110264 The serine residue corresponding to the conserved serine residue at position 82 of (SEQ ID NO:201).

[0254] This document provides a fusion protein comprising the ComP glycosylated fragment of this disclosure. In some embodiments, the fusion protein is coupled to SEQ ID NO:201 (ComP) at the glycosylated fragment. 110264 The serine residue corresponding to the ComP glycosylation fragment residue at position 82 is glycosylated with an oligosaccharide or polysaccharide. Further, although a glycoconjugate comprising a ComP glycosylation fragment comprising a fusion protein has been described above, it should be understood that this disclosure also explicitly provides a fusion protein consistent with any and all descriptions of fusion proteins provided anywhere herein (including in the appended claims). In some embodiments, the fusion protein comprises a carrier protein selected from the group consisting of: Pseudomonas aeruginosa exotoxin A (EPA), CRM... 197 Cholera toxin B subunit, tetanus toxin C fragment, Haemophilus influenzae protein D and one or more fragments thereof.

[0255] This document also provides a method for in vivo conjugation of oligosaccharides or polysaccharides to receptor peptides. In some embodiments, the method includes culturing host cells containing components necessary for conjugating oligosaccharides or polysaccharides to peptides. Generally, these components are oligosaccharide transferases, receptor peptides to be glycosylated, and oligosaccharides or polysaccharides. The method includes covalently linking oligosaccharides or polysaccharides to receptor peptides (fusion proteins of this disclosure) using a PglS oligosaccharide transferase (OTase), wherein the receptor peptide contains the ComP glycosylated fragment described herein. In some embodiments, the PglS OTase is a PglS 110264 (SEQ ID NO:365), PglS ADP1 (SEQ ID NO:366), PglS GFJ-2 (SEQ ID NO:367), PglS 50v1 (SEQ ID NO:368), PglS 4466 (SEQ ID NO:369), PglS SFC (SEQ ID NO:370), Pgl SP5312 (SEQ ID NO:371) or PglS ANT_H59(SEQ ID NO:372). In some embodiments, oligosaccharides or polysaccharides are used in conjunction with SEQ ID NO:201 (ComP) 110264 The serine residue at position 82 corresponds to the serine The residues are linked to a ComP glycosylated fragment. In some embodiments, in vivo conjugation occurs in a host cell. In some aspects, the glycoconjugate is produced in bacterial cells, fungal cells, yeast cells, avian cells, algal cells, insect cells, or mammalian cells. In some embodiments, the host cell is a bacterial cell, for example: in *Escherichia coli*; in bacteria of the genus *Klebsiella*; the bacterial species is *Klebsiella pneumoniae*, *Klebsiella heterotropha*, *Klebsiella micrantha*, or *Klebsiella acidogenic*. Some embodiments include culturing a host cell containing: (a) a gene cluster encoding a protein required for the synthesis of an oligosaccharide or polysaccharide; (b) a PglS OTase; and (c) a receptor polypeptide. In some embodiments, the production of the oligosaccharide or polysaccharide is enhanced by a homolog of *Klebsiella pneumoniae* transcription activator rmpA (*Klebsiella pneumoniae* NTUH K-2044). In some embodiments, the method further includes expressing and / or providing such transcriptional activators and other components in the host cell.

[0256] In some respects, glycoconjugates are produced in cell-free systems. Examples of applications of cell-free systems utilizing OTases other than PglS can be found in WO2013 / 067523A1, which is incorporated herein by reference.

[0257] A host cell is also provided comprising: (a) a gene cluster encoding a protein required for the synthesis of an oligosaccharide or polysaccharide; (b) a PglS OTase; and (3) a receptor polypeptide comprising a ComP glycosylated fragment of the present disclosure. In some embodiments, the receptor polypeptide is a fusion protein. In some embodiments, the host cell comprises nucleic acid encoding a PglS OTase. In some embodiments, the host cell comprises nucleic acid encoding a receptor polypeptide.

[0258] This document also provides an isolated nucleic acid encoding the ComP glycosylated fragment and / or fusion protein disclosed herein. In some embodiments, the nucleic acid is a vector. In some embodiments, a host cell contains the isolated nucleic acid.

[0259] The glycoconjugates of the present invention can have one of a variety of uses, including but not limited to use as conjugate vaccines. Therefore, conjugate vaccines are produced in certain methods. In some embodiments, compositions comprising the conjugate vaccine or fusion protein of the present disclosure and an adjuvant are used. For example, in some embodiments, the conjugate vaccine is against Klebsiella pneumoniae serotypes 8, 1, 2, 4, 5, 6A, 6B, 7F, 9N, 9V, and 1. 0A, Klebsiella pneumoniae serotype 11A, Klebsiella pneumoniae serotype 12F, Klebsiella pneumoniae serotype 14, Klebsiella pneumoniae serotype 15B, Klebsiella pneumoniae serotype 17F, Klebsiella pneumoniae serotype 18C, Klebsiella pneumoniae serotype 19F, Klebsiella pneumoniae serotype 19A, Klebsiella pneumoniae serotype 20, Klebsiella pneumoniae serotype 22F, Klebsiella pneumoniae serotype 23F, Klebsiella pneumoniae serotype 33F, Klebsiella pneumoniae serotype K1, Klebsiella pneumoniae serotype K2, Klebsiella pneumoniae serotype K5, Klebsiella pneumoniae serotype K16, Klebsiella pneumoniae serotype K20, Klebsiella pneumoniae serotype K54, Klebsiella pneumoniae serotype K57, Streptococcus agalactiae serotype Ia, Streptococcus agalactiae serotype Ib, Streptococcus agalactiae serotype II, Streptococcus agalactiae serotype III, ... Vaccines containing serotypes IV, V, VI, VII, VIII, and IX of *Streptococcus agalactiae*, group A carbohydrates of *Streptococcus pyogenes*, serotypes A, B, C, and D of *Enterococcus faecalis*, capsular polysaccharides and lipoteichoic acid of *Enterococcus faecium*, oligosaccharides A, B, and C of *Moraxella catarrhalis*, and lipoteichoic acid of *Staphylococcus aureus*. In some embodiments, conjugate vaccines are useful because they induce an immune response when administered to a subject. In some embodiments, the immune response triggers long-term memory (memory B cells and T cells), which is an antibody response, and optionally a serum-type-specific antibody response. In some embodiments, the antibody response is an IgG or IgM response. For example, in some embodiments, the antibody response may be an IgG response, and in some embodiments, it may be an IgG1 response.In some embodiments, the conjugate vaccine generates immune memory in subjects who have been given the vaccine.

[0260] This document discloses a pneumococcal glycoconjugate vaccine containing a conventional vaccine vector, which can be generated by isolating a glycoconjugate or glycosylated fusion protein of the present disclosure containing a ComP glycosylated fragment, and combining the isolated glycoconjugate or isolated glycosylated fusion protein with an adjuvant. In some embodiments, the ComP glycosylated fragment can be added to a conventional carrier protein, Pseudomonas aeruginosa exotoxin A (EPA). It has been demonstrated that, in some embodiments, the glycosylated fragment / carrier fusion protein can be paired with CPS8 polysaccharide and PglS to generate a carrier protein-CPS8 bioconjugate, which is a first-in-class pneumococcal bioconjugate vaccine. For example, in some embodiments, the EPA fusion protein can be paired with CPS8 polysaccharide and PglS to generate an EPA-CPS8 bioconjugate. It has been demonstrated that the EPA-CPS8 bioconjugate vaccine induces high IgG titers specific to serotype 8, as determined by bacterial inactivation, and these titers are protective. Importantly, protection can be provided by vaccination using only 100 ng of the polysaccharide from the EPA-CPS8 bioconjugate. Therefore, some embodiments provide a CPS8 pneumococcal bioconjugate vaccine.

[0261] It is envisioned that conjugate vaccines (such as EPA vaccine constructs) could include additional / multiple glycosylation sites to increase the glycan-to-protein ratio and expand the number of serotypes, thereby developing comprehensive pneumococcal bioconjugate vaccines.

[0262] Moraxellaceae O-linked oligosaccharide transferase

[0263] This disclosure also relates to a novel family of bacterial O-linked oligosaccharide transferases, termed TfpM, derived from Moraxellaceae bacteria. Certain embodiments of this disclosure include any one of the TfpM-associated fimbriae glycosylation fragments and / or subsequent TfpM OTases. The size and sequence of TfpM proteins are similar to TfpO enzymes, but they can transfer long-chain polysaccharides to receptor proteins. Phylogenetic analysis has demonstrated that TfpM proteins cluster in different clades from known bacterial oligosaccharide transferases. Using representative TfpM enzymes from Moraxella osloensis, it has been determined that TfpM glycosylates the C-terminal threonine residue of its homologous fimbriae-like protein, and the minimum sequence required for glycosylation has been identified. Studies have demonstrated that TfpM exhibits broad substrate tolerance and can transfer a variety of polysaccharides, including reduced-terminal glucose, galactose, or 2-N-acetyl sugars. Studies have also shown that TfpM-derived bioconjugates are immunogenic and elicit a serum-type-specific polysaccharide IgG response in mice. Therefore, the heterogeneity of TfpM glycan substrates and the identification of the smallest TfpM sequence make it a valuable addition to the toolbox for expanding glycan engineering.

[0264] Bioinformatics identification of a novel class of OTases carried by bacteria from the Moraxellaceae family.

[0265] To identify the gene encoding an O-linked oligosaccharide transferase, the inventors first used the Basic Local Alignment Search (BLAST) and PglS... ADP1 The amino acid sequence (SEQ ID NO:1) was used as the query sequence to search the NCBI Genome and Whole Genome Shotgun Contiguous Sequence Database.

[0266] SEQ ID NO:1_PglS ADP1 amino acid sequence

[0267] MNSIFKKIKNYTIVSGVFFLGSAFIIPNTSNLSSTLYKELIAVLGLLILLTVKSFDYKKILIPKNFYWFLFVIFIIFIQLIVGEIYFFQDFFFSISFLVILFLSFLLGFNERLNGDDLIVKKIAWIFIIVVQISFLI AINQKIEIVQNFFLFSSSYNGRSTANLGQPNQFSTLILITLFLLCYLREKNSLNNMVFNILSFCLIFANVMTQSRSAWISVILISLLYLLKFQKKIELRRVIFFNIVFWTLVYCVPLLFNLIFFQKNSYSTFDRLTM GSSRFEIWPQLLKAVFHKPFIGYGWGQTGVAQLETINKSSTKGEWFTYSHNLFLDLMLWNGFFIGLIISILILCFLIELYSSIKNKSDLFLFFCVVAFFVHCLLEYPFAYTYFLIPVGFLCGYISTQNIKNSISYFN LSKRKLTLFLGCCWLGYVAFWVEVLDISKKNEIYARQFLFSNHVKFYNIENYILDGFSKQLDFQYLDYCELKDKYQLLDFKKVAYRYPNASIVYKYYSISAEMKMDQKSANQIIRAYSVIKNQKIIKPKLKFCSIEY

[0268] To shorten the hit list and reduce the identification of PglS ADP1 The possibility of very similar orthologs further refines the search to include PglS. ADP1 Candidates with less than 50% amino acid sequence identity. Some top hits in this refined list are proteins that are more similar in size to TfpO proteins, but whose upstream homologous fimbriae contain both ComP disulfide flanking sequences and C-terminal PilA-like sequences. The first identified fimbriae-oligosaccharide transferase pairs encode two Acinetobacter species: Acinetobacter parvus DSM16617 and Acinetobacter townerii ZZC-3 (Table 1).

[0269] Table 1. Registry numbers of organisms and TfpM enzymes and their associated fimbriae proteins

[0270]

[0271]

[0272] The two oligosaccharide transferases from these strains were closely related, sharing >96% sequence identity. Intrigued by these findings, the inventors further observed and identified other strains within the Moralesceae family carrying genes encoding similar putative oligosaccharide transferase / fimbriae pairs. While many genes similar to those in *Acinetobacter parvum* DSM 16617 and *Acinetobacter thomsonii* ZZC-3 were found, most of the associated fimbriae encoded upstream of the oligosaccharide transferases lacked ComP-like sequence groups. Table 1 lists the accession numbers and protein sizes of twenty of these putative oligosaccharide transferase-fimbriae-like protein pairs. The inventors did not observe any homologs in species outside the Moralesceae family and, to distinguish these distinct oligosaccharide transferases from other known enzymes, named them TfpM proteins (“M” stands for Moralesceae). Given the similar size of TfpM proteins to known TfpO proteins, it was initially hypothesized that these genes encode variants of the TfpO-PilA pair (such as those found in Acinetobacter and Pseudomonas) (Harding, CM et al. (2015) Molecular Microbiology 96, 1023-1041). However, multiple sequence alignment of twenty TfpM proteins with known PglS, PglL, and TfpO proteins revealed that the former share less than 26% sequence identity with the prototype oligosaccharide transferase. Analysis of the phylogenetic tree generated by the multiple alignments showed that TfpM proteins clustered with TfpO, PglS, and PglL proteins in different clades (Figure 23A and 23B). Figure 24 Conversely, the fimbriae genes located upstream of tfpM did not cluster in discrete clades. Figure 25 The proteins showed a generally high degree of identity with the PilA protein, ranging between 37% and 60%. Based on the coding sequence, most of the associated fimbriae belong to the major type IV fimbriae family, except for those from Acinetobacter CIP102143 and Acinetobacter CIP102637, as they lack the characteristic type III signal sequence at the N-terminus (Giltner Carmen, L. et al. (2012) Microbiology and Molecular Biology Reviews 76, 740-772).

[0273] TfpM orthologs were used to glycosylate engineered fimbriae protein-fusion proteins with pneumococcal type 8 capsular polysaccharide.

[0274] Although TfpM proteins are similar in size to TfpO proteins, their amino acid sequences differ significantly and warrant further investigation. The inventors are particularly interested in determining whether TfpM proteins can transfer only short oligosaccharides to the receptor protein, similar to TfpO proteins. Of the twenty TfpM oligosaccharide transferases listed in Table 1, the inventors selected 13 representative enzymes from different clades to test their glycosylation activity in glycoengineered E. coli strains (Harding, CM and Feldman, MF (2019) Glycobiology 29, 519-529; Feldman, MF et al. (2005) Proceedings of the National Academy of Sciences of the United States of America 102, 3016). Previously, the inventors developed a chimeric receptor protein strategy consisting of fusing an exotoxin A protein (EPA) from *Pseudomonas aeruginosa* with soluble ComP fragments of varying sizes (the natural substrate of PglS) (Knoot, CJ et al. (2021) Glycobiology 31, 1192-1203). All type IV pilinoid proteins contain conserved N-terminal pilinoid signaling sequences and membrane anchoring domains, which are not essential for glycosylation but are crucial for pilinoid stability. The fusion protein approach allows for the removal of the conserved N-terminal pilinoid signaling sequence and membrane anchoring domain and has been used to identify PglS. ADP1The smallest sequence that can be recognized and still efficiently glycosylated (Harding, CM et al. (2019) Nature Communications 10, 891; Knoot, CJ et al. (2021) Glycobiology 31, 1192-1203). The inventors adapted this method and designed 13 synthetic double-stranded DNA blocks encoding the upstream fimbriae gene and the N-terminal truncated fragments of the downstream tfpM gene. In most strains carrying tfpM, the protein-coding region of the upstream fimbriae gene overlaps with the start codon of tfpM by one nucleotide. This genetic architecture remains intact in the expression construct. The synthetic DNA blocks are designed such that when cloned into the EPA expression vector using Gibson assembly, it places the fimbriae coding region within the same reading frame as the C-terminus of EPA, thereby creating a gene fusion that is translated into a single protein immediately adjacent to the downstream tfpM gene (Figure 23B). The truncated fimbriae fragments were 113 to 140 amino acids in size. A purification tag was not added to the C-terminus of the EPA-fimbriae fusion because previous studies on TfpO from *Pseudomonas aeruginosa* 1244 reported that adding an additional C-terminal residue after a serine residue prevented glycosylation (Horzempa, J. et al. (2006) *Journal of Biological Chemistry* 281, 1128-1136). Expression of the EPA-fimbriae and TfpM proteins was driven by the IPTG-inducible tac promoter on the pEXT20 plasmid (Dykxhoorn, DM et al. (1996) *Gene* 177, 133-136). The fusion protein was secreted into the periplasm using the DsbA signal sequence at the N-terminus of EPA. The oligonucleotides and primers used for assembly are listed in Table 2.

[0275] Table 2. Primers and Oligonucleotides

[0276]

[0277]

[0278]

[0279] Using this design, the inventors evaluated the ability of 13 TfpM proteins to transfer the pneumococcal capsular polysaccharide 8 (CPS8) glycan to its homologous fimbriae domain on the EPA-fimbriae fusion. The CPS8 repeat unit is a tetrasaccharide with a glucose reducing end. Notably, PglS is therefore the only known oligosaccharide transferase to date capable of spontaneously transferring this glycan to the recipient protein (Harding, CM et al. (2019) Nature Communications 10, 891). The 13 EPA-fimbriae / TfpM expression vectors were transformed into E. coli SDB1 strains expressing CPS8 glycan (Feldman, MF et al. (2005) Proceedings of the National Academy of Sciences of the United States of America 102, 3016), and protein glycosylation was evaluated. The expected mass range of the unglycosylated fusion proteins was 78.3 to 80.5 kDa. Several TfpM proteins were found to undergo glycosylation with their homologous EPA-fimbriae fusion proteins. This glycosylation was characterized by a stepwise change in molecular weight above the unglycosylated band (g0) (g...). n (Fig. 23C). Each higher-weighted band represents a glycan linked to EPA-fimbriae with an additional CPS8 repeat unit. Glycosylation was readily observed in the seven TfpM orthologs tested: Acinetobacter YZSX-1-1, Acinetobacter CIP102637, Acinetobacter YH01026, Acinetobacter jonniger 65, Acinetobacter CIP102143, Acinetobacter TUM15069, and Moraxella osloensis 1202 (Fig. 23C). Some of these extracts required higher protein blot exposures to observe the glycosylation patterns (Fig. 23D). As a control, the inventors constructed a mutant of the Moraxella osloensis TfpM protein gene that was conserved at histidine (His 286 The change involves a single residue, and this change has previously been shown to be crucial for the activity of enzymes containing the wzy_C domain. Figure 26 The wzy_C family pfam04932 is an "O-antigen ligase" domain found in membrane-bound enzymes that catalyzes the transfer and covalent linkage of lipid-linked oligosaccharides (liposes) to lipid A or protein substrates. (Ruan, X. et al. (2012) Glycobiology 22, 288-299; Musumeci, MA et al. (2014) Glycobiology 24, 39-50).

[0280] No glycosylation was observed in SDB1 cell extracts expressing the H286A mutant, as well as the EPA-fimbriae fusion and CPS8 glycan (Fig. 23C), indicating that glycosylation is caused by the activity of the tfpM gene product, and His 286 This is essential for the catalytic activity and / or stability of the oligosaccharide transferase.

[0281] TfpMMo is an O-linked oligosaccharide transferase that glycosylates the C-terminal threonine residue of its fimbriae protein substrate.

[0282] Of the 13 TfpM-filum protein pairs tested in the aforementioned experiments, those from *Moraxella osloensis* 1202 and *Acinetobacter* YH01026 exhibited the most efficient transfer of glycans of different sizes. Because the filum proteins from *Moraxella osloensis* FDARGOS_1202 (hereinafter referred to as 1202) showed slightly higher apparent stability, the inventors selected oligosaccharide transferases from this organism as representatives for further characterization and named them TfpM. Mo (SEQ ID NO:56). For clarity, the complete protozoan Moraxella osloensis 1202 fimbriae protein will be referred to as Pil throughout the text. Mo (SEQ IDNO:57), and the fusion domain with the N-terminus truncated is called Pil. Mo Δ28 (SEQ ID NO:58). Next, use TfpM. Mo Identify Pil Mo The glycosylation site at Δ28 was identified, thus determining whether the enzyme's mechanism of action is similar to that of the TfpO protein (glycosylation of the C-terminal amino acid of its homologous fimbriae receptor) or more like that of the PglL or PglS proteins (glycosylation of internal residues). The TfpO protein transfers a short oligosaccharide, typically containing 3-6 sugars, to the C-terminal serine residue side chain of its homologous fimbriae. With one exception, all homologous fimbriae upstream of the tfpM open reading frame end with a C-terminal threonine residue, i.e., fimbriae from *Psychrophilus 72-Oc* end with a serine residue. Based on this observation, and the size similarity between TfpM and TfpO proteins, the inventors hypothesized that the TfpM enzyme also transfers glycans to C-terminal residues, and therefore designed a point mutant of the fimbriae to verify this. This resulted in the generation of a C-terminal fimbriae threonine residue (Thr...). 167 Two mutants of ) were used to convert the residue to serine or alanine, and glycosylation was tested using CPS8. An anti-EPA antibody was used in Western blotting to target Pil expression. Mo Whole-cell E. coli extracts from strains of the Δ28 mutant were used for detection. Figure 27 As shown, TfpM Mo It can also be used for EPA-PilMo Δ28T167S undergoes glycosylation, but glycosylation was not observed in the alanine mutant. These results indicate that Thr... 167 or Ser 167 The hydroxyl group on the side chain of the residue may be TfpM Mo Sites that connect to glycans.

[0283] To confirm TfpM Mo The sites for glycosylation of fimbriae proteins, specifically EPA-Pil glycosylated via CPS8. Mo Δ28 was partially purified and separated by SDS-PAGE analysis, and the separated glycoproteins were stained with Coomassie Brilliant Blue. EPA-Pil glycosylated with 1–3 CPS8 repeat units was excised. Mo Gel sections corresponding to Δ28 were digested with LysC and analyzed for glycopeptides. LysC-derived EPA-Pil... Mo Analysis of the Δ28 peptide using an open-search approach (Chick, JM et al. (2015) Nature Biotechnology 33, 743-749; Polasky, DA et al. (2020) Nature Methods 17, 1125-1132) led to the identification of a peptide modified with hexose-hexuronic acid (HexA). 762 FLPANCRGT 770 A peptide was identified that is identical to the incomplete monomer of the CPS8 glycan (HexHexAHex2). High-energy C-trap dissociation (HCD) analysis supported the linkage of this disaccharide through hexose residues, and targeted electron transfer / high-energy collision dissociation (EThcD) analysis confirmed the linkage of HexHexA with the C-terminal threonine residue (Figures 28A and 28B). Although this mass spectrometry analysis did not identify a polymer of the CPS8 tetrasaccharide, this is not surprising, as it is extremely difficult to detect elongated glycoconjugates using peptide-centered LC-MS methods without the use of special chemical additives such as pressurizers (Lin, C.-w. et al. (2016) Analytical Chemistry 88, 8484-8494). Nevertheless, the identification of a disaccharide identical to the partially completed CPS8 tetrasaccharide still supports the high molecular weight stepwise transformation into a polymerized CPS8 tetrasaccharide.

[0284] TfpM Mo Polysaccharides containing glucose, galactose, or 2-N-acetyl monosaccharides are transferred at the reducing end.

[0285] Next, the aim is to explore TfpM MoThe range of polysaccharide substrates transferable in the glycoengineered E. coli environment. The inventors selected polysaccharides containing different reducing terminal sugars, different disaccharide linkages near the reducing terminals, and / or polymers composed of linear or branched repeating units. In addition to Pneumococcal CPS8, four polysaccharides were tested as TfpM: ​​E. coli O16 antigen, Salmonella enterica LT2 O- antigen, Klebsiella pneumoniae O2a antigen, and group III B streptococcal capsular polysaccharide. Mo The polysaccharide substrates. The structures of the repeat units for all five tests are shown in Figure 29A (Liu, B. et al. (2020) FEMS Microbiology Reviews 44, 655-683; Curd, H. et al. (1998) Journal of Bacteriology 180, 1002-1007; Whitfield, C. et al. (1992) Journal of Bacteriology 174, 4913-4919; Pinto, V. and Berti, F. (2014) Journal of Pharmaceutical and Biomedical Analysis 98, 9-15; Geno, KA et al. (2015) Clinical Microbiology Reviews 28, 871-899). Each of the five polysaccharides was individually reacted with EPA-Pil Mo Δ28 and TfpM Mo Co-expression, induction, and growth were performed in *E. coli* SDB1 cells for subsequent glycoprotein purification. Periplasmic extracts from SDB1 cells were partially purified using anion-exchange chromatography to remove any contaminating undecylenoyl pyrophosphate-linked polysaccharides that might obscure Western blot interpretation. To demonstrate that the glycan-specific antibody signals observed in Western blots indeed originated from glycosylated proteins, rather than from residual contaminating lipid-linked polysaccharides in whole-cell lysates, the purified glycoproteins were divided into two aliquots. One aliquot was digested with proteinase K, followed by SDS-PAGE separation and Western blotting. Western blot analysis was performed using antisera specific to each polysaccharide, and also with an anti-EPA antibody, as all antibodies used in this experiment were derived from rabbits. Figure 7 As shown, TfpM Mo It was found to efficiently transfer five different polysaccharides to EPA-Pil. MoΔ28 protein. Proteinase K digestion eliminated anti-glycan (Fig. 29B, Fig. 29C, Fig. 29D, Fig. 29E and Fig. 29F) and anti-EPA (Fig. 29G) signals in the proteomic blot, thus confirming that the anti-glycan signal originated from protein-linked polysaccharides rather than contaminating lipid-linked polysaccharides.

[0286] TfpM Mo Shortened Pil Mo Δ28 variant undergoes glycosylation

[0287] Previous experiments all utilized PilA, which is 139 amino acids in length. Mo The N-terminal truncated variant. To gain a deeper understanding of TfpM... Mo The minimum features required for C-terminal fimbriae glycosylation resulted in a series of further truncated Pil proteins. Mo The inventors first generated Δ28 variants and tested whether these variants could be glycosylated. They then developed a 20-amino acid Pil... Mo The fragment, which is fused to the EPA at the C-terminus via a flexible four-residue glycine linker, is named Pil. 20 (Figure 30A). This 20-amino acid fragment was chosen because it contains a conserved disulfide ring (“DSL”) region found in many type IV fimbriae. Figure 31 (Horzempa, J. et al. (2006) Journal of Biological Chemistry 281, 1128-1136; Harvey, H. et al. (2009) Journal of Bacteriology 191, 6513-6524). It is noteworthy that this DSL differs from the motif corresponding to the disulfide-flanked sequence in the ComP protein (Knoot, CJ et al. (2021) Glycobiology 31, 1192-1203). Based on sequence alignment with Pseudomonas aeruginosa 1244PilA, the DSL in the Moraxella osloensis fimbriae protein consists of the residue Cys. 148 and Cys 164 Therefore, the inventors designed the sequence downstream of the glycine linker to form Cys. 148 Begin. Includes EPA-Pil code. 20 The plasmid for the TfpM construct is named pVNM297. 20 Glycosylation experiments showed that it can be glycosylated by TfpM Mo Use CPS8 with Pil MoGlycosylation was performed at a similar level to Δ28 (Figure 29B). To test whether the DSL region was necessary for glycosylation, several shorter variants lacking this feature were generated. In these smaller constructs, the inventors also... 164 The amino acid was mutated to alanine to prevent the formation of unnatural disulfide linkages during periplasmic oxidation (Harvey, H. et al. (2009) Journal of Bacteriology 191, 6513-6524). Pil was tested with 15, 13, and 10 amino acids. Mo Variants, in which one of the 10 amino acid versions contains an amino acid linker (e.g., a GGGG linker), and the other does not. These constructs are respectively called Pil. 15 Pil 13 Pil 10L and Pil 10 (Figure 30A). TfpM Mo Four of these variants were capable of glycosylation, although at a lower level than Pil. 20 and Pil Mo Δ28( Figure 27 B). Pil 10 and Pil 10L The degree of glycosylation in both is comparable, indicating that the presence of the upstream glycine linker is important for TfpM. Mo Activity was not significantly affected. This linker was omitted in all subsequent constructs. Of these four proteins, Pil... 13 The degree of glycosylation is significantly lower than that of other proteins.

[0288] Sequence alignment revealed that each fimbriae glycosylated by TfpM OTase possesses a conserved “PAN / ECRG” motif located near the C-terminus, immediately upstream of the penultimate threonine residue. Figure 31 Given that this feature is present in all glycosylated fimbriae, some have questioned whether this motif is a TfpM. Mo Essential for glycosylation. The inventors used Pil, a compound containing seven amino acids. Mo The variant (called Pil7) fused into the EPA, and this variant was formed by a similar motif (modified to) The cysteine ​​residue was mutated to alanine (a thickened residue), and glycosylation was evaluated (Figure 30A). The inventors also constructed stepwise single-amino acid truncated versions of this "PANARGT" sequence, reducing it from seven amino acids to two, and evaluated the ability of TfpM to glycosylate these fragments using CPS8. The results showed that, except for Pil2, all variants were glycosylated by TfpM. Mo With il 10Glycosylation proceeded at similar levels (Figure 30B). The overall glycosylation level was again lower than that of Pil. 20 and Pil Mo Δ28. Glycosylated Pil2 was almost undetectable, but some trace stepwise changes were visible at higher protein blot exposures. This suggests that TfpM Mo This variant can be glycosylated, but the level of glycosylation is significantly lower than that of the construct containing even one additional amino acid. These results lead to the conclusion that TfpM… Mo Pil can recognize three amino acids Mo The fragment is then glycosylated.

[0289] TfpM Mo Immunogenicity of derived GBSIII bioconjugates

[0290] Given EPA-Pil 20 The construct was efficiently glycosylated with TfpM, and the inventors then evaluated the EPA-Pil glycosylated with type III capsular polysaccharide (GBSIII) from group B streptococci in a mouse vaccination model. 20 Immunogenicity of the protein. To facilitate protein purification for these experiments, an EPA-Pil protein expressing an N-terminal 6x-His tag was constructed. 20 The plasmid of the vector protein variant (pVNM291) was used. A His tag was added downstream of the predicted N-terminal DsbA signal cleavage site of EPA. pVNM291 was introduced into SDB1 cells expressing GBSIII glycan, and the resulting bioconjugate was purified using nickel-immobilized metal affinity chromatography (IMAC), followed by anion exchange chromatography and size exclusion chromatography (FPLC). Western blots and Coomassie brilliant blue staining of the GBSIII-291 bioconjugate separated by SDS-PAGE confirmed the EPA-Pil... 20 High molecular weight glycosylation of protein with GBSIII glycan (Fig. 32A, Fig. 32B, Fig. 32C, and Fig. 32D). Purified EPA-Pil 20 The complete protein MS of the GBSIII (“GBSIII-291”) conjugate supported a glycan:protein ratio of 20% (Figure 32E). Each dose was formulated to contain 1 μg of GBSIII polysaccharide. As a control for these experiments, unglycosylated pVNM291-derived carrier protein (“291”) was purified from SDB1 cells without the glycan plasmid and administered at the same protein concentration as the GBSIII bioconjugate.

[0291] To test TfpM MoThe immunogenicity of the generated GBSIII-291 bioconjugate was determined by immunizing 5-week-old female CD1 mice. Mice received either a placebo (unglycosylated 291 carrier protein) or the GBSIII-291 bioconjugate (starting with an initial dose followed by two booster doses) at two-week intervals. All vaccines were administered at a 1:9 ratio. A 2% solution was prepared as an adjuvant. Serum was collected before each immunization and two weeks after the last booster immunization. To determine the induced GBSIII-specific antibody levels, the inventors used an enzyme-linked immunosorbent assay (ELISA). High levels of anti-GBSIII IgG antibody expression were observed in all mice immunized with the GBSIII bioconjugate-291, but only one mouse showed a lower anti-GBSIII IgG response, which was enhanced during immunization (Fig. 32F). As expected, the GBSIII-specific IgG titers were increased in mice immunized with the GBSIII bioconjugate compared to mice immunized with a mimic (received only 291, Fig. 32F). Overall, these data suggest that TfpM... Mo It can produce bioconjugates that can elicit a polysaccharide-specific IgG response.

[0292] TfpM Mo and PglS ADP1 (PglL ComP Glycosylation of individual proteins engineered to contain sequences specific to each oligosaccharide transferase.

[0293] Finally, the inventors wanted to determine whether proteins engineered to contain sequences from two different OTase systems could be glycosylated by both OTases at each site. Therefore, an EPA fusion protein containing sequences associated with TfpM and PglS was constructed. To this end, the inventors designed an EPA fusion protein containing, as previously described, an EPA fusion protein integrated into the residue Ala. 489 With Arg 490 The PglS sequence (CTGVTQIASGASAATTNVASAQC) (SEQ ID NO: 59) (Knoot, CJ et al. (2021) Glycobiology 31, 1192-1203) and the Pil at the C-terminus 20The sequence (CGGTGTTVAAKFLPANCRGT) (SEQ ID NO: 60; identical to Pil DSL) (Figure 33A). As described above, this construct was designed such that the open reading frame of the gene encoding the EPA fusion overlaps with the start codon of tfpM by one nucleotide. The open reading frame encoding pglS from Acinetobacter bengal ADP1 was cloned 100 bp downstream of the stop codon of the tfpM open reading frame. This vector (pVNM337) was introduced into E. coli SDB1 expressing the E. coli O16 antigen, and its glycosylation was detected by Western blotting. To compare with proteins containing only a single sequence, the inventors introduced the following constructs individually into E. coli SDB1 expressing the O16 antigen: (i) containing only the TfpM-associated Pil 20 The sequence EPA (pVNM297) or (ii) contains residues integrated into Ala. 489 With Arg 490 The EPA (pVNM167) contains the PglS sequence between the two sequences. To compare with proteins containing two sequences that can be disaccharified, the inventors also introduced an EPA construct containing the residue Ala. 489 With Arg 490 Between and residues Glu 548 With Gly 549 The sequence subsequences (pVNM245) and (EPA_PglS sequence subsequence 2X) are shown in Figure 33A. A schematic diagram of these constructs is shown in Figure 33B. As shown in Figure 33B, each construct contains only one sequence from ComP or Pil. Mo Western blot analysis of the EPA construct containing the PglS sequence showed a glycosylation profile of approximately 100 kDa, indicating monosaccharidation. The EPA construct containing two PglS sequences showed a major monosaccharidation profile of approximately 100 kDa, but also showed a disaccharidation population migrating to approximately 150 kDa. A TfpM sequence was also included. Mo and PglS ADP1 Western blot analysis of the EPA fusion sequence revealed the presence of both monoglycosylated and disaccharidated populations, similar to those observed in the pVNM245 construct. These results suggest that the receptor protein can be glycosylated by two different OTase-like enzymes within a single expression system.

[0294] Sugar conjugates

[0295] This disclosure provides a glycoconjugate comprising an oligosaccharide or polysaccharide covalently linked to a receptor protein. In some embodiments, the receptor protein comprises or consists of the following: the TfpM-associated fimbriae-like protein of this disclosure or a glycosylated fragment thereof. In some embodiments, the oligosaccharide or polysaccharide is covalently linked to the fimbriae-like protein or a glycosylated fragment thereof. Furthermore, in some embodiments, the TfpM-associated fimbriae-like protein or a glycosylated fragment thereof comprises a C-terminal serine or threonine residue, and the oligosaccharide or polysaccharide is covalently linked to a C-terminal serine or threonine residue. Further, in some embodiments, the receptor protein is a fusion protein comprising a TfpM-associated fimbriae-like protein or a glycosylated fragment thereof fused / linked to a heterologous amino acid sequence (e.g., a carrier protein) in a translational manner, and the TfpM-associated fimbriae-like protein or its glycosylated fragment is the C-terminalmost sequence of the receptor protein, such that the receptor protein comprises a C-terminal serine or threonine residue, and the oligosaccharide or polysaccharide is covalently linked to a C-terminal serine or threonine residue. Illustrative examples of carrier proteins include, but are not limited to, fragments of Pseudomonas aeruginosa exotoxin A (EPA), CRM197, cholera toxin B subunit, tetanus toxin C fragment, or any of these. In some embodiments, TfpM-associated fimbriae-like proteins or glycosylated fragments thereof are translatorily fused / linked to a heterologous amino acid sequence / carrier protein via an amino acid linker. In some embodiments, oligosaccharides or polysaccharides contain glucose at their reducing ends. In some embodiments, the glycoconjugates are immunogenic.

[0296] In some embodiments of the glycoconjugates disclosed herein, the receptor protein comprises or is composed of a full-length TfpM-associated fimbriae-like protein. In some embodiments, the receptor protein comprises or is composed of a glycosylated fragment of a TfpM-associated fimbriae-like protein smaller than the full-length TfpM-associated fimbriae-like protein. In some embodiments, the fimbriae-like protein glycosylated fragment is 3 to 138 amino acids long, 10 to 138 amino acids long, 20 to 138 amino acids long, 50 to 138 amino acids long, 100 to 138 amino acids long, or 116 to 138 amino acids long. In some embodiments, the fimbriae-like protein glycosylated fragment is 3 to 139 amino acids long, 10 to 139 amino acids long, 20 to 139 amino acids long, 50 to 139 amino acids long, 100 to 139 amino acids long, or 116 to 139 amino acids long. In some embodiments, the glycosylated fragment is 3 to 140 amino acids long, 10 to 140 amino acids long, 20 to 140 amino acids long, 50 to 140 amino acids long, 100 to 140 amino acids long, or 116 to 140 amino acids long. In some embodiments, the glycosylated fragment is 3 to 22 amino acids long, 10 to 22 amino acids long, 11 to 22 amino acids long, 3 to 21 amino acids long, 5 to 21 amino acids long, 10 to 21 amino acids long, or 11 to 21 amino acids long. In some embodiments, the length of the glycosylated fragment is any from 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 amino acids to 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids. In some embodiments, the length of the glycosylated fragment is any from 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids.

[0297] In some embodiments of the glycoconjugates disclosed herein, the TfpM-associated fimbriae-like protein or its glycosylated fragment is Pil. Mo (SEQ ID NO:57) or Pil lacking amino acids corresponding to residues 1–28 Mo (Pil MoΔ28 (SEQ ID NO:58) or a polypeptide containing at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity with SEQ ID NO:57 or SEQ ID NO:58, for example, wherein the C-terminal threonine is replaced by a serine. In some embodiments of the glycoconjugates disclosed herein, the TfpM-associated fimbriae-like protein is selected from the group consisting of: Pil DSM16617 (SEQ IDNO:82), Pil ZZC3-9 (SEQ ID NO:83), Pil TUM15069 (SEQ ID NO:84), Pil AI7 (SEQ ID NO:85), Pil VE-C3 (SEQ ID NO:86), Pil YH01026 (SEQ ID NO:87), Pil CIP102143 (SEQ ID NO:88), Pil AI40 (SEQ ID NO:89), Pil F78 (SEQ ID NO:90), Pil S71 (SEQ ID NO:91), Pil ANC4282 (SEQ ID NO:92), Pil 72-O-c (SEQ ID NO:93), Pil BI730 (SEQ ID NO:94), Pil A3K91 (SEQ ID NO:95), Pil CIP102159 (SEQ ID NO:96), Pil junii-65 (SEQ ID NO:97), Pil YZS-X (SEQ ID NO:98), Pil T-3-2 (SEQ ID NO:99) and Pil CIP102637 (SEQ ID NO:100). In some embodiments, the TfpM-associated pilonoid protein or pilonoid glycosylated fragment comprises or consists of the following: an amino acid sequence selected from the group consisting of: Pil DSM16617 (SEQ ID NO:82), Pil ZZC3-9 (SEQ ID NO:83), Pil TUM15069 (SEQ ID NO:84), Pil AI7 (SEQ ID NO:85), Pil VE-C3 (SEQ ID NO:86), Pil YH01026 (SEQ ID NO:87), PilCIP102143 (SEQ ID NO:88), Pil AI40 (SEQ ID NO:89), Pil F78 (SEQ ID NO:90), Pil S71 (SEQ ID NO:91), Pil ANC4282 (SEQ ID NO:92), Pil 72-O-c (SEQ ID NO:93), Pil BI730 (SEQ ID NO:94), Pil A3K91 (SEQ ID NO:95), Pil CIP102159 (SEQ ID NO:96), Pil junii-65 (SEQ ID NO:97), Pil YZS-X (SEQ ID NO:98), Pil T-3-2 (SEQ ID NO:99), Pil CIP102637 (SEQ ID NO:100), and fragments thereof (e.g., C-terminal fragments) and / or variants wherein the C-terminal threonine is replaced by a serine. Furthermore, in some embodiments, the fibroin-like protein glycosylated fragment comprises or consists of the following: Pil Mo fimbriae protein disulfide ring region (Pil Mo _DSL, also known as Pil 20 (SEQ ID NO: 60) or truncated derivatives thereof, which contain at least the last three amino acids from the C-terminus of the fimbriae, or variants wherein the C-terminal threonine is replaced by a serine (SEQ ID NO: 148). Further, in some embodiments, the fimbriae-like protein glycosylation fragment consists of the following: Pil 20 (SEQ ID NO:60), Pil 19 (SEQ ID NO:133), Pil 18 (SEQ ID NO:134), Pil 17 (SEQ ID NO:135), Pil 16 (SEQ ID NO:136), Pil 15 (SEQ ID NO:109), Pil 14 (SEQ ID NO:137), Pil 13 (SEQ ID NO:110), Pil 12 (SEQ IDNO:138), Pil 11 (SEQ ID NO:139), Pil 10(SEQ ID NO:112), Pil9 (SEQ ID NO:140), Pil8 (SEQ ID NO:141), Pil7 (SEQ ID NO:113), Pil6 (SEQ ID NO:114), Pil5 (SEQ ID NO:115), Pil4 (SEQ ID NO:116), or Pil3 (SEQ ID NO:117), or variants thereof having one, two, three, four, or five amino acid substitutions and maintaining a C-terminal threonine residue. Further, in some embodiments, the fimbriae-like protein glycosylation fragment consists of the following: Pil 20S (SEQ ID NO:148), Pil 19S (SEQ ID NO:149), Pil 18S (SEQ ID NO:150), Pil 17S (SEQ ID NO:151), Pil 16S (SEQ ID NO:152), Pil 15S (SEQ ID NO:153), Pil 14S (SEQ IDNO:154), Pil 13S (SEQ ID NO:155), Pil 12S (SEQ ID NO:156), Pil 11S (SEQ ID NO:157), Pil 10S (SEQ ID NO:158), Pil 9S (SEQ ID NO:159), Pil 8S (SEQ ID NO:160), Pil 7S (SEQ ID NO:161), Pil 6S (SEQ ID NO:162), Pil 5S (SEQ ID NO:163), Pil 4S (SEQ ID NO:164) or Pil 3S (SEQ ID NO:165), or a variant thereof having one, two, three, four or five amino acid substitutions and maintaining a C-terminal serine.

[0298] In some embodiments of the glycoconjugates disclosed herein, the receptor protein may be glycosylated at two or more different locations. In some embodiments, the receptor protein may be glycosylated by at least two different OTases in an expression system. For example, in some embodiments, the receptor protein is a fusion protein, and the fusion protein, in addition to the TfpM-associated fimbriae-like protein glycosylation fragment located at its C-terminus, further includes an additional glycosylation sequence (e.g., a glycosylation fragment) of an OTase other than a TfpM oligosaccharide transferase (OTase). For example, the other OTase may be PglB, PglL, or PglS. In some embodiments, the additional glycosylation sequence is an internal sequence of the fusion protein (i.e., not the C-terminal or N-terminal terminal sequence). In some embodiments, the additional glycosylation sequence is an internal sequence of the carrier protein (e.g., Figure 33A). In some embodiments, the additional glycosylation sequence is also covalently linked to an oligosaccharide or a polysaccharide. In some embodiments, the fusion protein comprises two or more, three or more, four or more, five or more, six or more, eight or more, ten or more, fifteen or more, or twenty or more additional glycosylation sequences. In some embodiments, the fusion protein does not comprise more than two, three, five, ten, fifteen, twenty, or twenty-five additional glycosylation sequences. In some embodiments, the additional glycosylation sequences are identical. In some embodiments, at least one additional glycosylation sequence is different from each other. In some embodiments, at least three, four, or five of these additional glycosylation sequences are different from each other. And, in some embodiments, all additional glycosylation sequences are different. In some embodiments, the receptor protein is a fusion protein, and the fusion protein further comprises an internal glycosylation fragment of ComP in addition to a TfpM-associated fimbriae-like protein glycosylation fragment located at its C-terminus. In some embodiments, the ComP glycosylation fragment is internally located within the carrier protein sequence. In some embodiments, the ComP glycosylation fragment is also covalently linked to an oligosaccharide or polysaccharide. Furthermore, in some embodiments, the ComP glycosylation fragment comprises or consists of the following: CTGVTQIASGASAATTNVASAQC (SEQ ID NO: 59) or a fragment thereof containing at least the amino acid ASA at positions 11-13. In some embodiments, the fusion protein comprises two or more, three or more, four or more, five or more, six or more, eight or more, ten or more, fifteen or more, or twenty or more ComP glycosylation fragments. In some embodiments, the fusion protein does not comprise more than two, three, five, ten, fifteen, twenty, or twenty-five ComP glycosylation fragments. In some embodiments, the ComP glycosylation fragments are identical.In some embodiments, the ComP glycosylation fragments are different from each other. In some embodiments, at least three, at least four, or at least five of these ComP glycosylation fragments are different from each other. And in some embodiments, all ComP glycosylation fragments are different.

[0299] In some embodiments of the glycoconjugates disclosed herein, the oligosaccharides or polysaccharides covalently linked to the fimbriae-like protein or its glycosylated fragments are at least three repeating oligosaccharide or polysaccharide structural units in size. In some embodiments of the glycoconjugates disclosed herein, the oligosaccharides or polysaccharides covalently linked to the fimbriae-like protein or its glycosylated fragments are at least ten monosaccharides in size.

[0300] In some embodiments of the glycoconjugates disclosed herein, the oligosaccharide or polysaccharide is produced by Streptococcus bacteria (e.g., Streptococcus pneumoniae or Streptococcus agalactiae), and the polysaccharide is a capsular polysaccharide, such as Ia, Ib, II, III, IV, V, VI, VII, VIII, or IX.

[0301] In some embodiments of the glycoconjugates disclosed herein, the oligosaccharide or polysaccharide is produced by Klebsiella bacteria (e.g., Klebsiella pneumoniae), and the polysaccharide is a capsular polysaccharide or an O-antigen polysaccharide.

[0302] In some embodiments of the glycoconjugates disclosed herein, the oligosaccharide or polysaccharide is produced by Salmonella bacteria, and the polysaccharide is an O-antigen polysaccharide. In some embodiments, the bacteria are enteric Salmonella, and the enteric Salmonella polysaccharide is a group B O-antigen.

[0303] In some embodiments of the glycoconjugates disclosed herein, the glycoconjugates are produced in vivo, such as in bacterial cells. In some embodiments, the bacteria are *Escherichia coli*. In some embodiments, the bacteria are from the genus *Klebsiella*. In some embodiments, the bacterial species are *Klebsiella pneumoniae*, *Klebsiella heterotropha*, *Klebsiella micrantha*, or *Klebsiella acidogenic*. In some embodiments, the glycoconjugates are produced in a cell-free system.

[0304] In some embodiments of the glycoconjugates disclosed herein, the bioconjugate is a conjugate vaccine that induces an immune response when administered to a subject. In some embodiments, the immune response triggers long-term memory (memory B cells and T cells), is an antibody response, and optionally a serotype-specific antibody response. In some embodiments, the antibody response is an IgG or IgM response. In some embodiments, the antibody response is an IgG response, for example, an IgG1 response. And, in some embodiments, the conjugate vaccine generates immune memory in the subject administered the vaccine.

[0305] Glycosylation fragments

[0306] This disclosure provides a fibroin-like protein glycosylation fragment comprising or consisting of a separated fragment of the TfpM-associated fibroin-like protein of this disclosure. In some embodiments, the TfpM-associated fibroin-like protein or the fibroin-like protein glycosylation fragment comprises or consists of the following: Pil Mo (SEQ ID NO:57) or Pil lacking amino acids corresponding to residues 1–28 Mo (Pil Mo Δ28 (SEQ ID NO:58) or a polypeptide containing at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity with SEQ ID NO:57 or SEQ ID NO:58, for example, wherein the C-terminal threonine is replaced by a serine. In some embodiments, the TfpM-associated fimbriae-like protein is selected from the group consisting of: Pil DSM16617 (SEQ ID NO:82), Pil ZZC3-9 (SEQ ID NO:83), Pil TUM15069 (SEQ ID NO:84), Pil AI7 (SEQ ID NO:85), Pil VE-C3 (SEQ ID NO:86), Pil YH01026 (SEQ ID NO:87), Pil CIP102143 (SEQ ID NO:88), Pil AI40 (SEQ ID NO:89), Pil F78 (SEQ ID NO:90), Pil S71 (SEQ ID NO:91), Pil ANC4282 (SEQ ID NO:92), Pil 72-O-c (SEQ ID NO:93), Pil BI730 (SEQ ID NO:94), Pil A3K91 (SEQ ID NO:95), Pil CIP102159 (SEQ ID NO:96), Pil junii-65 (SEQ ID NO:97), Pil YZS-X (SEQ ID NO:98), Pil T-3-2 (SEQ ID NO:99) and Pil CIP102637 (SEQ ID NO:100). In some embodiments, the TfpM-associated pilonoid protein or pilonoid glycosylated fragment comprises or consists of the following: an amino acid sequence selected from the group consisting of: Pil DSM16617(SEQ ID NO:82), Pil ZZC3-9 (SEQ ID NO:83), Pil TUM15069 (SEQ ID NO:84), Pil AI7 (SEQ ID NO:85), Pil VE-C3 (SEQ ID NO:86), Pil YH01026 (SEQ ID NO:87), Pil CIP102143 (SEQ ID NO:88), Pil AI40 (SEQ ID NO:89), Pil F78 (SEQ ID NO:90), Pil S71 (SEQ ID NO:91), Pil ANC4282 (SEQ ID NO:92), Pil 72-O-c (SEQ ID NO:93), Pil BI730 (SEQ ID NO:94), Pil A3K91 (SEQ ID NO:95), Pil CIP102159 (SEQ ID NO:96), Pil junii-65 (SEQ IDNO:97), Pil YZS-X (SEQ ID NO:98), Pil T-3-2 (SEQ ID NO:99), Pil CIP102637 (SEQ ID NO:100), and fragments thereof (e.g., C-terminal fragments) and / or variants wherein the C-terminal threonine is replaced by a serine. Furthermore, in some embodiments, the pifiloid-like protein glycosylated fragment comprises or consists of the following: PilMo pifiloid disulfide ring region (Pil... Mo _DSL, also known as Pil 20 (SEQ ID NO: 60) or truncated derivatives thereof, which contain at least the last three amino acids from the C-terminus of the fimbriae, or variants wherein the C-terminal threonine is replaced by a serine (SEQ ID NO: 148). Further, in some embodiments, the fimbriae-like protein glycosylation fragment consists of the following: Pil 20 (SEQ ID NO:60), Pil 19 (SEQ ID NO:133), Pil 18 (SEQ ID NO:134), Pil 17 (SEQ IDNO:135), Pil 16 (SEQ ID NO:136), Pil 15(SEQ ID NO:109), Pil 14 (SEQ ID NO:137), Pil 13 (SEQ ID NO:110), Pil 12 (SEQ ID NO:138), Pil 11 (SEQ ID NO:139), Pil 10 (SEQ ID NO:112), Pil9 (SEQ ID NO:140), Pil8 (SEQ ID NO:141), Pil7 (SEQ ID NO:113), Pil6 (SEQ ID NO:114), Pil5 (SEQ ID NO:115), Pil4 (SEQ ID NO:116), or Pil3 (SEQ ID NO:117), or variants thereof having one, two, three, four, or five amino acid substitutions and maintaining a C-terminal threonine residue. Further, in some embodiments, the fimbriae-like protein glycosylation fragment consists of the following: Pil 20S (SEQ ID NO:148), Pil 19S (SEQ ID NO:149), Pil 18S (SEQ ID NO:150), Pil 17S (SEQ ID NO:151), Pil 16S (SEQ ID NO:152), Pil 15S (SEQ ID NO:153), Pil 14S (SEQ ID NO:154), Pil 13S (SEQ ID NO:155), Pil 12S (SEQ ID NO:156), Pil 11S (SEQ ID NO:157), Pil 10S (SEQ ID NO:158), Pil 9S (SEQ ID NO:159), Pil 8S (SEQ IDNO:160), Pil 7S (SEQ ID NO:161), Pil 6S (SEQ ID NO:162), Pil 5S (SEQ ID NO:163), Pil 4S (SEQ ID NO:164) or Pil 3S (SEQ ID NO:165), or a variant thereof having one, two, three, four or five amino acid substitutions and maintaining a C-terminal serine.

[0307] In some embodiments of the glycosylated fragments of the present invention, the separated fragments of the TfpM-associated fimbriae-like protein disclosed herein are 3 to 138 amino acids in length, 10 to 138 amino acids in length, 20 to 138 amino acids in length, 50 to 138 amino acids in length, 100 to 138 amino acids in length, or 116 to 138 amino acids in length. In some embodiments, the glycosylated fragments are 3 to 139 amino acids in length, 10 to 139 amino acids in length, 20 to 139 amino acids in length, 50 to 139 amino acids in length, 100 to 139 amino acids in length, or 116 to 139 amino acids in length. In some embodiments, the glycosylated fragment is 3 to 140 amino acids long, 10 to 140 amino acids long, 20 to 140 amino acids long, 50 to 140 amino acids long, 100 to 140 amino acids long, or 116 to 140 amino acids long. In some embodiments, the glycosylated fragment is 3 to 22 amino acids long, 10 to 22 amino acids long, 11 to 22 amino acids long, 5 to 21 amino acids long, 10 to 21 amino acids long, or 11 to 21 amino acids long. In some embodiments, the length of the glycosylated fragment is any from 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 amino acids to 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids. In some embodiments, the length of the glycosylated fragment is any from 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids.

[0308] Fusion protein

[0309] This document provides a fusion protein comprising a TfpM-associated fimbriae-like protein or a glycosylated fragment thereof, fused translationally to / linked to a heterologous carrier protein, such as, but not limited to, Pseudomonas aeruginosa exotoxin A (EPA), CRM197, cholera toxin B subunit, tetanus toxin C fragment, or a fragment thereof. In some embodiments, the TfpM-associated fimbriae-like protein or its glycosylated fragment is fused translationally to / linked to the heterologous carrier protein via an amino acid linker. In some embodiments, the fimbriae-like protein or its glycosylated fragment comprises a C-terminal serine or threonine residue. In some embodiments, the fimbriae-like protein or its glycosylated fragment is the C-terminal terminal sequence of the fusion protein. Furthermore, in some embodiments, the fusion protein comprises a C-terminal serine or threonine residue. Further, in some embodiments, the fusion protein is glycosylated via an oligosaccharide or polysaccharide covalently linked to a C-terminal serine or threonine residue. Furthermore, in some embodiments, the fusion protein is glycosylated by covalently linking to an oligosaccharide or polysaccharide containing glucose at its reduced terminus. In some embodiments, the glycosylated fusion protein is immunogenic. In some embodiments, the glycosylated fusion protein is a conjugated vaccine.

[0310] In some embodiments of the fusion protein disclosed herein, the fusion protein comprises a full-length TfpM-associated fimbriae-like protein. In some embodiments, the fusion protein comprises or is composed of a glycosylated fragment of the TfpM-associated fimbriae-like protein that is smaller than or consists of the full-length TfpM-associated fimbriae-like protein. In some embodiments, the fimbriae-like protein glycosylated fragment is 3 to 138 amino acids long, 10 to 138 amino acids long, 20 to 138 amino acids long, 50 to 138 amino acids long, 100 to 138 amino acids long, or 116 to 138 amino acids long. In some embodiments, the glycosylated fragment is 3 to 139 amino acids long, 10 to 139 amino acids long, 20 to 139 amino acids long, 50 to 139 amino acids long, 100 to 139 amino acids long, or 116 to 139 amino acids long. In some embodiments, the glycosylated fragment is 3 to 140 amino acids long, 10 to 140 amino acids long, 20 to 140 amino acids long, 50 to 140 amino acids long, 100 to 140 amino acids long, or 116 to 140 amino acids long. In some embodiments, the glycosylated fragment is 3 to 22 amino acids long, 10 to 22 amino acids long, 11 to 22 amino acids long, 5 to 21 amino acids long, 10 to 21 amino acids long, or 11 to 21 amino acids long. In some embodiments, the length of the glycosylated fragment is any from 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 amino acids to 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids. In some embodiments, the length of the glycosylated fragment is any from 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids.

[0311] In some embodiments of the fusion protein disclosed herein, the fimbriae-like protein glycosylated fragment comprises or consists of the following: Pil Mo (SEQ ID NO:57) or Pil lacking amino acids corresponding to residues 1–28 Mo (Pil MoΔ28 (SEQ ID NO:58) or a polypeptide containing at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity with SEQ ID NO:57 or SEQ ID NO:58, for example, wherein the C-terminal threonine is replaced by a serine. In some embodiments, the TfpM-associated fimbriae-like protein is selected from the group consisting of: Pil DSM16617 (SEQ IDNO:82), Pil ZZC3-9 (SEQ ID NO:83), Pil TUM15069 (SEQ ID NO:84), Pil AI7 (SEQ ID NO:85), Pil VE-C3 (SEQ ID NO:86), Pil YH01026 (SEQ ID NO:87), Pil CIP102143 (SEQ ID NO:88), Pil AI40 (SEQ ID NO:89), Pil F78 (SEQ ID NO:90), Pil S71 (SEQ ID NO:91), Pil ANC4282 (SEQ ID NO:92), Pil 72-O-c (SEQ ID NO:93), Pil BI730 (SEQ ID NO:94), Pil A3K91 (SEQ ID NO:95), Pil CIP102159 (SEQ ID NO:96), Pil junii-65 (SEQ ID NO:97), Pil YZS-X (SEQ ID NO:98), Pil T-3-2 (SEQ ID NO:99) and Pil CIP102637 (SEQ ID NO:100). In some embodiments, the TfpM-associated pilonoid protein or pilonoid glycosylated fragment comprises or consists of the following: an amino acid sequence selected from the group consisting of: Pil DSM16617 (SEQ ID NO:82), Pil ZZC3-9 (SEQ ID NO:83), Pil TUM15069 (SEQ ID NO:84), Pil AI7 (SEQ ID NO:85), Pil VE-C3 (SEQ ID NO:86), Pil YH01026 (SEQ ID NO:87), Pil CIP102143(SEQ ID NO:88), Pil AI40 (SEQ ID NO:89), Pil F78 (SEQ ID NO:90), Pil S71 (SEQ ID NO:91), Pil ANC4282 (SEQ ID NO:92), Pil 72-O-c (SEQ ID NO:93), Pil BI730 (SEQ ID NO:94), Pil A3K91 (SEQ ID NO:95), Pil CIP102159 (SEQ ID NO:96), Pil junii-65 (SEQ ID NO:97), Pil YZS-X (SEQ ID NO:98), Pil T-3-2 (SEQ ID NO:99), Pil CIP102637 (SEQ ID NO:100), and fragments thereof (e.g., C-terminal fragments) and / or variants wherein the C-terminal threonine is replaced by a serine. Furthermore, in some embodiments, the pifiloid-like protein glycosylated fragment comprises or consists of the following: PilMo pifiloid disulfide ring region (Pil... Mo _DSL, also known as Pil 20 (SEQ ID NO: 60) or truncated derivatives thereof, which contain at least the last three amino acids from the C-terminus of the fimbriae, or variants wherein the C-terminal threonine is replaced by a serine (SEQ ID NO: 148). Further, in some embodiments, the fimbriae-like protein glycosylation fragment consists of the following: Pil 20 (SEQ ID NO:60), Pil 19 (SEQ ID NO:133), Pil 18 (SEQ ID NO:134), Pil 17 (SEQ ID NO:135), Pil 16 (SEQ ID NO:136), Pil 15 (SEQ ID NO:109), Pil 14 (SEQ ID NO:137), Pil 13 (SEQ ID NO:110), Pil 12 (SEQ IDNO:138), Pil 11 (SEQ ID NO:139), Pil 10(SEQ ID NO:112), Pil9 (SEQ ID NO:140), Pil8 (SEQ ID NO:141), Pil7 (SEQ ID NO:113), Pil6 (SEQ ID NO:114), Pil5 (SEQ ID NO:115), Pil4 (SEQ ID NO:116), or Pil3 (SEQ ID NO:117), or variants thereof having one, two, three, four, or five amino acid substitutions and maintaining a C-terminal threonine residue. Further, in some embodiments, the fimbriae-like protein glycosylation fragment consists of the following: Pil 20S (SEQ ID NO:148), Pil 19S (SEQ ID NO:149), Pil 18S (SEQ ID NO:150), Pil 17S (SEQ ID NO:151), Pil 16S (SEQ ID NO:152), Pil 15S (SEQ ID NO:153), Pil 14S (SEQ IDNO:154), Pil 13S (SEQ ID NO:155), Pil 12S (SEQ ID NO:156), Pil 11S (SEQ ID NO:157), Pil 10S (SEQ ID NO:158), Pil 9S (SEQ ID NO:159), Pil 8S (SEQ ID NO:160), Pil 7S (SEQ ID NO:161), Pil 6S (SEQ ID NO:162), Pil 5S (SEQ ID NO:163), Pil 4S (SEQ ID NO:164) or Pil 3S (SEQ ID NO:165), or a variant thereof having one, two, three, four or five amino acid substitutions and maintaining a C-terminal serine.

[0312] In some embodiments of the fusion protein disclosed herein, the fusion protein may be glycosylated at two or more different locations. In some embodiments, the fusion protein may be glycosylated by at least two different OTases in an expression system. For example, in some embodiments, the fusion protein, in addition to a TfpM-associated fimbriae-like protein glycosylation fragment located at its C-terminus, further comprises a glycosylated sequence (e.g., a glycosylated fragment) of an OTase other than a TfpM oligosaccharide transferase (OTase). For example, the other OTase may be PglB, PglL, or PglS. In some embodiments, the additional glycosylation sequence is an internal sequence of the fusion protein (i.e., not the C-terminal or N-terminal terminal sequence). In some embodiments, the additional glycosylation sequence is an internal sequence of the carrier protein (e.g., Figure 11A). In some embodiments, the additional glycosylation sequence is also covalently linked to an oligosaccharide or polysaccharide. In some embodiments, the fusion protein comprises two or more, three or more, four or more, five or more, six or more, eight or more, ten or more, fifteen or more, or twenty or more additional glycosylation sequences. In some embodiments, the fusion protein does not comprise more than two, three, five, ten, fifteen, twenty, or twenty-five additional glycosylation sequences. In some embodiments, the additional glycosylation sequences are identical. In some embodiments, at least one additional glycosylation sequence is different from each other. In some embodiments, at least three, four, or five of these additional glycosylation sequences are different from each other. And, in some embodiments, all additional glycosylation sequences are different. In some embodiments, the fusion protein further comprises an internal glycosylation fragment of ComP in addition to the TfpM-associated fimbriae-like protein glycosylation fragment located at its C-terminus. In some embodiments, the ComP glycosylation fragment is also covalently linked to an oligosaccharide or polysaccharide. Furthermore, in some embodiments, the ComP glycosylation fragment comprises or consists of the following: CTGVTQIASGASAATTNVASAQC (SEQ ID NO: 59) or a fragment thereof containing at least the amino acid ASA at positions 11-13. In some embodiments, the fusion protein comprises two or more, three or more, four or more, five or more, six or more, eight or more, ten or more, fifteen or more, or twenty or more ComP glycosylation fragments. In some embodiments, the fusion protein does not comprise more than two, three, five, ten, fifteen, twenty, or twenty-five ComP glycosylation fragments. In some embodiments, the ComP glycosylation fragments are identical. In some embodiments, the ComP glycosylation fragments are different from each other. In some embodiments, at least three, at least four, or at least five of these ComP glycosylation fragments are different from each other. Furthermore, in some embodiments, all ComP glycosylation fragments are different.

[0313] In some embodiments of the fusion protein disclosed herein, the oligosaccharide or polysaccharide covalently linked to the fimbriae-like protein or its glycosylated fragment has a size of at least three oligosaccharide or polysaccharide structural repeating units. In some embodiments of the fusion protein disclosed herein, the oligosaccharide or polysaccharide covalently linked to the fimbriae-like protein or its glycosylated fragment has a size of at least ten monosaccharides.

[0314] In some embodiments of the fusion protein disclosed herein, the oligosaccharide or polysaccharide is produced by Streptococcus bacteria (e.g., Streptococcus pneumoniae or Streptococcus agalactiae), and the polysaccharide is a capsular polysaccharide, such as Ia, Ib, II, III, IV, V, VI, VII, VIII, or IX.

[0315] In some embodiments of the fusion protein disclosed herein, the oligosaccharide or polysaccharide is produced by Klebsiella bacteria (e.g., Klebsiella pneumoniae), and the polysaccharide is a capsular polysaccharide or an O-antigen polysaccharide.

[0316] In some embodiments of the fusion protein disclosed herein, the oligosaccharide or polysaccharide is produced by Salmonella bacteria, and the polysaccharide is an O-antigen polysaccharide. In some embodiments, the bacteria are enteric Salmonella, and the enteric Salmonella polysaccharide is a group B O-antigen.

[0317] In some embodiments of the fusion protein disclosed herein, the glycosylated fusion protein is produced in vivo, such as in bacterial cells. In some embodiments, the bacteria are *Escherichia coli*. In some embodiments, the bacteria are derived from the genus *Klebsiella*. In some embodiments, the bacterial species are *Klebsiella pneumoniae*, *Klebsiella heterotrophus*, *Klebsiella micrantha*, or *Klebsiella acidogenic*.

[0318] In some embodiments of the fusion protein disclosed herein, the fusion protein is a vaccine that induces an immune response when administered to a subject. In some embodiments, the immune response triggers long-term memory (memory B cells and T cells), is an antibody response, and optionally a serotype-specific antibody response. In some embodiments, the antibody response is an IgG or IgM response. In some embodiments, the antibody response is an IgG response, for example, an IgG1 response. And, in some embodiments, the fusion protein generates immune memory in the subject administered the fusion protein.

[0319] Methods for producing glycoconjugates

[0320] This document provides a method for producing glycoconjugates. In some embodiments, the method occurs in vivo. In some aspects, the glycoconjugates are produced in a cell-free system. Examples of applications of cell-free systems utilizing OTases other than TfpM can be found in WO2013 / 067523A1, which is incorporated herein by reference. In some embodiments, the method includes covalently linking (conjugating) an oligosaccharide or polysaccharide to a receptor protein using a TfpM oligosaccharide transferase (OTase) of this disclosure, the receptor protein comprising or consisting of: a TfpM-associated fimbriae-like protein or a glycosylated fragment thereof of this disclosure. In some embodiments, the fimbriae-like protein or glycosylated fragment comprises a C-terminal serine or threonine residue, the receptor protein comprises a C-terminal serine or threonine residue, and the oligosaccharide or polysaccharide is covalently linked to a C-terminal serine or threonine residue of the receptor protein. In some embodiments, the oligosaccharide or polysaccharide comprises glucose at its reduced end. In some embodiments, the receptor protein is a fusion protein of this disclosure as described in detail elsewhere herein. Furthermore, in some embodiments, the glycoconjugate is immunogenic.

[0321] In certain embodiments of the methods for generating glycoconjugates disclosed herein, or together with any other compositions or methods disclosed herein, TfpM OTase contains a wzy_C superfamily domain (an O-antigen ligase domain) defined by the conserved protein domain family cl04850 of the National Center for Biotechnology Information (NCBI), and / or TfpM OTase contains a wzy_C domain (an O-antigen ligase domain) defined by the conserved protein domain family motif pfam04932 of the European Institute for Molecular Biology (EMBL) and the European Bioinformatics Institute (EBI, EMBL-EBI), wherein pfam04932 is a protein domain family within the cl04850 superfamily. In some embodiments, TfpM OTase is combined with TfpM Mo (SEQ ID NO:56), TfpM DSM16617 (SEQ ID NO:63), TfpM ZZC3 (SEQ ID NO:64), TfpM TUM15069 (SEQ ID NO:65), TfpM AI7 (SEQ ID NO:66), TfpM VE-C3 (SEQ ID NO:67), TfpM YH01026 (SEQ ID NO:68), TfpM CIP102143 (SEQ ID NO:69), TfpM AI40 (SEQ ID NO:70), TfpM F78 (SEQ ID NO:71), TfpM S71 (SEQ ID NO:72), TfpM ANC4282 (SEQ ID NO:73), TfpM CIP102159 (SEQ ID NO:74), TfpM junii-65 (SEQ ID NO:75), TfpM YZS-X (SEQ ID NO:76), TfpM CIP102637 (SEQ ID NO:77), TfpM T-3-2 (SEQ ID NO:78), TfpM BI730 (SEQ ID NO:79), TfpM A3K91 (SEQ ID NO:80) and / or TfpM 72-O-c(SEQ ID NO:81) contains at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity. In some embodiments, the TfpM OTase is TfpM Mo (SEQ ID NO:56), TfpM DSM16617 (SEQ ID NO:63), TfpM ZZC3 (SEQ ID NO:64), TfpM TUM15069 (SEQ ID NO:65), TfpM AI7 (SEQ ID NO:66), TfpM VE-C3 (SEQ ID NO:67), TfpM YH01026 (SEQ ID NO:68), TfpM CIP102143 (SEQ ID NO:69), TfpM AI40 (SEQ ID NO:70), TfpM F78 (SEQ ID NO:71), TfpM S71 (SEQ ID NO:72), TfpM ANC4282 (SEQID NO:73), TfpM CIP102159 (SEQ ID NO:74), TfpM junii-65 (SEQ ID NO:75), TfpM YZS-X (SEQ ID NO:76), TfpM CIP102637 (SEQ ID NO:77), TfpM T-3-2 (SEQ ID NO:78), TfpM BI730 (SEQ ID NO:79), TfpM A3K91 (SEQ ID NO:80) or TfpM 72-O-c (SEQ ID NO:81). In some embodiments, TfpM OTase is TfpM Mo (SEQ ID NO:56).

[0322] In some embodiments of the method for generating glycoconjugates disclosed herein, the receptor protein comprises or is composed of a full-length TfpM-associated fimbriae-like protein. In some embodiments, the receptor protein comprises or is composed of a glycosylated fragment of a TfpM-associated fimbriae-like protein shorter than the full-length TfpM-associated fimbriae-like protein. In some embodiments, the fimbriae-like protein glycosylated fragment is 3 to 138 amino acids long, 10 to 138 amino acids long, 20 to 138 amino acids long, 50 to 138 amino acids long, 100 to 138 amino acids long, or 116 to 138 amino acids long. In some embodiments, the fimbriae-like protein glycosylated fragment is 3 to 139 amino acids long, 10 to 139 amino acids long, 20 to 139 amino acids long, 50 to 139 amino acids long, 100 to 139 amino acids long, or 116 to 139 amino acids long. In some embodiments, the glycosylated fragment is 3 to 140 amino acids long, 10 to 140 amino acids long, 20 to 140 amino acids long, 50 to 140 amino acids long, 100 to 140 amino acids long, or 116 to 140 amino acids long. In some embodiments, the glycosylated fragment is 3 to 22 amino acids long, 10 to 22 amino acids long, 11 to 22 amino acids long, 5 to 21 amino acids long, 10 to 21 amino acids long, or 11 to 21 amino acids long. In some embodiments, the length of the glycosylated fragment is any from 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 amino acids to 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids. In some embodiments, the length of the glycosylated fragment is any from 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids.

[0323] In some embodiments of the method for generating glycoconjugates disclosed herein, the fimbriae-like protein glycosylated fragment comprises or consists of the following: Pil Mo (SEQ ID NO:57) or Pil lacking amino acids corresponding to residues 1–28 Mo (Pil MoΔ28 (SEQ ID NO:58) or a polypeptide containing at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity with SEQ ID NO:57 or SEQ ID NO:58, for example, wherein the C-terminal threonine is replaced by a serine. In some embodiments, the TfpM-associated fimbriae-like protein is selected from the group consisting of: Pil DSM16617 (SEQ ID NO:82), Pil ZZC3-9 (SEQ ID NO:83), Pil TUM15069 (SEQ ID NO:84), Pil AI7 (SEQ ID NO:85), Pil VE-C3 (SEQ ID NO:86), Pil YH01026 (SEQ ID NO:87), Pil CIP102143 (SEQ ID NO:88), Pil AI40 (SEQ ID NO:89), Pil F78 (SEQ ID NO:90), Pil S71 (SEQ ID NO:91), Pil ANC4282 (SEQ IDNO:92), Pil 72-O-c (SEQ ID NO:93), Pil BI730 (SEQ ID NO:94), Pil A3K91 (SEQ ID NO:95), Pil CIP102159 (SEQ ID NO:96), Pil junii-65 (SEQ ID NO:97), Pil YZS-X (SEQ ID NO:98), Pil T-3-2 (SEQ ID NO:99) and Pil CIP102637 (SEQ ID NO:100). In some embodiments, the TfpM-associated pilonoid protein or pilonoid glycosylated fragment comprises or consists of the following: an amino acid sequence selected from the group consisting of: Pil DSM16617 (SEQ ID NO:82), Pil ZZC3-9 (SEQ ID NO:83), Pil TUM15069 (SEQ ID NO:84), Pil AI7 (SEQ ID NO:85), Pil VE-C3 (SEQ ID NO:86), Pil YH01026 (SEQ ID NO:87), Pil CIP102143(SEQ ID NO:88), Pil AI40 (SEQ ID NO:89), Pil F78 (SEQ ID NO:90), Pil S71 (SEQ ID NO:91), Pil ANC4282 (SEQ ID NO:92), Pil 72-O-c (SEQ ID NO:93), Pil BI730 (SEQ ID NO:94), Pil A3K91 (SEQ ID NO:95), Pil CIP102159 (SEQ ID NO:96), Pil junii-65 (SEQ ID NO:97), Pil YZS-X (SEQ ID NO:98), Pil T-3-2 (SEQ ID NO:99), Pil CIP102637 (SEQ ID NO:100), and fragments thereof (e.g., C-terminal fragments) and / or variants wherein the C-terminal threonine is replaced by a serine. Furthermore, in some embodiments, the pifiloid-like protein glycosylated fragment comprises or consists of the following: PilMo pifiloid disulfide ring region (Pil... Mo _DSL, also known as Pil 20 (SEQ ID NO: 60) or truncated derivatives thereof, which contain at least the last three amino acids from the C-terminus of the fimbriae, or variants wherein the C-terminal threonine is replaced by a serine (SEQ ID NO: 148). Further, in some embodiments, the fimbriae-like protein glycosylation fragment consists of the following: Pil 20 (SEQ ID NO:60), Pil 19 (SEQ ID NO:133), Pil 18 (SEQ ID NO:134), Pil 17 (SEQ ID NO:135), Pil 16 (SEQ ID NO:136), Pil 15 (SEQ ID NO:109), Pil 14 (SEQ ID NO:137), Pil 13 (SEQ ID NO:110), Pil 12 (SEQ IDNO:138), Pil 11 (SEQ ID NO:139), Pil 10(SEQ ID NO:112), Pil9 (SEQ ID NO:140), Pil8 (SEQ ID NO:141), Pil7 (SEQ ID NO:113), Pil6 (SEQ ID NO:114), Pil5 (SEQ ID NO:115), Pil4 (SEQ ID NO:116), or Pil3 (SEQ ID NO:117), or variants thereof having one, two, three, four, or five amino acid substitutions and maintaining a C-terminal threonine residue. Further, in some embodiments, the fimbriae-like protein glycosylation fragment consists of the following: Pil 20S (SEQ ID NO:148), Pil 19S (SEQ ID NO:149), Pil 18S (SEQ ID NO:150), Pil 17S (SEQ ID NO:151), Pil 16S (SEQ ID NO:152), Pil 15S (SEQ ID NO:153), Pil 14S (SEQ IDNO:154), Pil 13S (SEQ ID NO:155), Pil 12S (SEQ ID NO:156), Pil 11S (SEQ ID NO:157), Pil 10S (SEQ ID NO:158), Pil 9S (SEQ ID NO:159), Pil 8S (SEQ ID NO:160), Pil 7S (SEQ ID NO:161), Pil 6S (SEQ ID NO:162), Pil 5S (SEQ ID NO:163), Pil 4S (SEQ ID NO:164) or Pil 3S (SEQ ID NO:165), or a variant thereof having one, two, three, four or five amino acid substitutions and maintaining a C-terminal serine.

[0324] In some embodiments of the method for generating glycoconjugates disclosed herein, the receptor protein is a fusion protein, and the carrier protein is selected from the group consisting of: Pseudomonas aeruginosa exotoxin A (EPA), CRM197, cholera toxin B subunit, tetanus toxin C fragment, and fragments of any of these. In some embodiments, a TfpM-associated fimbriae-like protein or a glycosylated fragment thereof is fused / linked to a heterologous carrier protein via an amino acid linker in a translational manner.

[0325] In some embodiments of the method for generating glycoconjugates disclosed herein, the receptor protein is a fusion protein, and the method includes glycosylation of the receptor protein at two or more different locations. In some embodiments, the method includes glycosylation of the receptor protein with at least two different classes of OTases in an expression system. In some embodiments, the fusion protein comprises two or more glycosylated sequences (e.g., glycosylated fragments) associated with at least two different OTases. Representative examples of OTases that can be used in combination include PglB, PglL, PglS, TfpO, and TfpM. Those skilled in the art will recognize that if two OTases require glycosylated sequences (sequences) at the same location, for example, both at the N-terminus or both at the C-terminus, they cannot be used simultaneously. For example, typically, if both TfpO and TfpM require sequences at the very distal C-terminus, they cannot be used simultaneously. For example, in some non-limiting illustrative embodiments, in addition to the additional glycosylation sequence of the OTase other than the TfpM oligosaccharide transferase (OTase), the receptor protein also includes a TfpM-associated fimbriae-like protein glycosylation fragment located at its C-terminus. In some embodiments, another OTase is PglB, PglL, and / or PglS. In some embodiments, one or more glycosylation sequences are sequences within the fusion protein (i.e., not the C-terminal or N-terminal terminal sequences). In some embodiments, one or more glycosylation sequences are sequences within the carrier protein sequence (e.g., Figure 33A). In some embodiments, the additional glycosylation sequence is a sequence within the fusion protein (i.e., not the C-terminal or N-terminal terminal sequences). In some embodiments, the additional glycosylation sequence is a sequence within the carrier protein sequence (e.g., Figure 33A). In some embodiments, at least two different glycosylation sequences from two different OTase systems are covalently linked to oligosaccharides or polysaccharides. In some embodiments, a TfpM-associated fimbriae-like protein glycosylation fragment located at the C-terminus of the fusion protein, along with additional glycosylation sequences, is covalently linked to an oligosaccharide or polysaccharide. In some embodiments, the fusion protein comprises two or more, three or more, four or more, five or more, six or more, eight or more, ten or more, fifteen or more, or twenty or more additional glycosylation sequences. In some embodiments, the fusion protein does not comprise more than two, three, five, ten, fifteen, twenty, or twenty-five additional glycosylation sequences. In some embodiments, the additional glycosylation sequences are identical. In some embodiments, at least one additional glycosylation sequence is different from each other. In some embodiments, at least three, four, or five of these additional glycosylation sequences are different from each other. And, in some embodiments, all additional glycosylation sequences are different.For example, in some embodiments, the method, in addition to glycosylation of the TfpM-associated fimbriae-like protein glycosylation fragment located at its C-terminus, further includes glycosylation of the internal glycosylation fragment of ComP using PglS OTase. In some embodiments, the ComP glycosylation fragment comprises or consists of the following: CTGVTQIASGASAATTNVASAQC (SEQ ID NO: 59) or a fragment thereof containing at least the amino acid ASA at positions 11-13. In some embodiments, the fusion protein comprises two or more, three or more, four or more, five or more, six or more, eight or more, ten or more, fifteen or more, or twenty or more ComP glycosylation fragments. In some embodiments, the fusion protein does not comprise more than two, three, five, ten, fifteen, twenty, or twenty-five ComP glycosylation fragments. In some embodiments, the ComP glycosylation fragments are identical. In some embodiments, the ComP glycosylation fragments are different from each other. In some embodiments, at least three, at least four, or at least five of these ComP glycosylation fragments are different from each other. Furthermore, in some embodiments, all ComP glycosylation fragments are different.

[0326] In some embodiments, the method of generating glycoconjugates (e.g., conjugation) of this disclosure occurs in vivo within a host cell. In some embodiments, the host cell is a bacterial cell. In some embodiments, the conjugation occurs in *Escherichia coli*. In some embodiments, the conjugation occurs in *Klebsiella* bacteria. In some embodiments, the bacterial species is *Klebsiella pneumoniae*, *Klebsiella heterotropha*, *Klebsiella micrantha*, or *Klebsiella acidogenic*.

[0327] In some embodiments of the method for generating glycoconjugates disclosed herein, the method includes culturing a host cell containing: (a) a gene cluster encoding a protein required for the synthesis of an oligosaccharide or polysaccharide; (b) a TfpM OTase; and (3) a receptor protein.

[0328] In some embodiments of the method for generating glycoconjugates disclosed herein, the method generates conjugate vaccines.

[0329] Other embodiments

[0330] This document provides a host cell comprising: (a) a gene cluster encoding a protein required for the synthesis of oligosaccharides or polysaccharides; (b) the TfpM OTase of this disclosure; and (3) a receptor protein comprising a TfpM-associated fimbriae-like protein or a glycosylated fragment thereof of this disclosure. In some embodiments, the receptor protein is a fusion protein. In some embodiments, the host cell contains nucleic acid encoding the TfpM OTase. In some embodiments, the host cell contains nucleic acid encoding the receptor protein. And, in some embodiments, the TfpM OTase and the receptor protein are encoded by the same nucleic acid.

[0331] This document provides an isolated nucleic acid encoding a glycosylated fragment and / or fusion protein of the fimbriae-like protein disclosed herein. In some embodiments, the nucleic acid is a vector. A host cell comprising this isolated nucleic acid of the present disclosure is also provided. In some embodiments, the host cell is a bacterial cell. In some embodiments, the host cell is *Escherichia coli*. In some embodiments, the host cell is derived from *Klebsiella*. And in some embodiments, the host cell is *Klebsiella pneumoniae*, *Klebsiella heterotropha*, *Klebsiella micrantha*, or *Klebsiella acidogenic*.

[0332] This document provides a composition comprising a conjugated vaccine or fusion protein of the present disclosure, along with an adjuvant and / or carrier. In some embodiments, the composition is a pharmaceutical or therapeutic composition suitable for administration to a subject / patient.

[0333] This document provides a method for inducing a host immune response against a bacterial pathogen, comprising administering to a subject requiring an immune response an effective amount of the disclosed conjugate vaccine, fusion protein, or composition comprising the conjugate vaccine or fusion protein and an adjuvant and / or a carrier. Treatment with a pharmaceutical composition comprising an immunogenic composition may occur alone or in combination with other treatments, as appropriate. An amount sufficient to achieve this purpose is defined as an “effective amount,” “effective dose,” or “unit dose.” The effective amount used for this purpose will depend, for example, on factors such as the glycoconjugate composition, the manner of administration, the stage and severity of the disease being treated, the patient’s weight and overall health status, and the prescribing physician’s judgment. In some aspects, an initial dose is followed by a booster dose over a period of time. In some embodiments, the immune response is an antibody response. In some embodiments, the immune response is selected from the group consisting of: innate response, adaptive response, humoral response, antibody response, cell-mediated response, B cell response, T cell response, cytokine upregulation or downregulation, immune system crosstalk, and combinations of two or more of said immune responses. In some embodiments, the immune response is selected from the group consisting of: innate response, humoral response, antibody response, T-cell response, and a combination of two or more of the immune responses.

[0334] This document provides a method for preventing or treating bacterial diseases and / or infections in a subject, the method comprising administering to a subject in need an effective amount of a conjugate vaccine, fusion protein, or composition comprising the conjugate vaccine or fusion protein and an adjuvant and / or a carrier. In some embodiments, the subject is a mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a companion animal. In some embodiments, the subject is a livestock. In some embodiments, the infection is a local or systemic infection of the skin, soft tissue, blood, or an organ, or is inherently autoimmune. In some embodiments, the disease is pneumonia. In some embodiments, the infection is a systemic infection and / or a blood infection. In some embodiments, the conjugate vaccine, fusion protein, or composition is administered via intramuscular injection, intradermal injection, intraperitoneal injection, subcutaneous injection, intravenous injection, oral administration, mucosal administration, intranasal administration, or pulmonary administration.

[0335] This article provides a method for generating a pneumococcal conjugate vaccine against pneumococcal infection, the method comprising: (a) isolating the glycoconjugate or glycosylated fusion protein of the present disclosure; and (b) combining the isolated glycoconjugate or isolated glycosylated fusion protein with an adjuvant and / or a carrier.

[0336] This document provides glycoconjugates, glycosylated fusion proteins, or conjugate vaccines, or compositions thereof, for inducing a host immune response against a bacterial pathogen and / or preventing or treating bacterial diseases and / or infections in a subject.

[0337] This article provides a recombinant nucleic acid construct comprising a nucleotide sequence encoding a TfpM oligosaccharide transferase (OTase) operably linked to at least one heterologous transcriptional regulatory sequence. In some embodiments, the TfpM OTase is coupled to TfpM... Mo (SEQ ID NO:56), TfpM DSM16617 (SEQ ID NO:63), TfpM ZZC3 (SEQ ID NO:64), TfpM TUM15069 (SEQ ID NO:65), TfpM AI7 (SEQ ID NO:66), TfpM VE-C3 (SEQ ID NO:67), TfpM YH01026 (SEQ ID NO:68), TfpM CIP102143 (SEQ ID NO:69), TfpM AI40 (SEQ ID NO:70), TfpM F78 (SEQ ID NO:71), TfpM S71(SEQ ID NO:72), TfpM ANC4282 (SEQ ID NO:73), TfpM CIP102159 (SEQ ID NO:74), TfpM junii-65 (SEQ ID NO:75), TfpM YZS-X (SEQ ID NO:76), TfpM CIP102637 (SEQ ID NO:77), TfpM T-3-2 (SEQ ID NO:78), TfpM BI730 (SEQ ID NO:79), TfpM A3K91 (SEQ ID NO:80) and / or TfpM 72-O-c (SEQ ID NO:81) contains at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity. In some embodiments, the TfpM OTase is TfpM Mo (SEQ ID NO:56), TfpM DSM16617 (SEQ ID NO:63), TfpM ZZC3 (SEQ ID NO:64), TfpM TUM15069 (SEQ ID NO:65), TfpM AI7 (SEQ ID NO:66), TfpM VE-C3 (SEQ ID NO:67), TfpM YH01026 (SEQ ID NO:68), TfpM CIP102143 (SEQ ID NO:69), TfpM AI40 (SEQ ID NO:70), TfpM F78 (SEQ ID NO:71), TfpM S71 (SEQ ID NO:72), TfpM ANC4282 (SEQID NO:73), TfpM CIP102159 (SEQ ID NO:74), TfpM junii-65 (SEQ ID NO:75), TfpM YZS-X (SEQ ID NO:76), TfpM CIP102637 (SEQ ID NO:77), TfpM T-3-2 (SEQ ID NO:78), TfpM BI730 (SEQ ID NO:79), TfpM A3K91 (SEQ ID NO:80) and / or TfpM 72-O-c(SEQ ID NO:81). In some embodiments, TfpM OTase is TfpM Mo(SEQ ID NO:56). In some embodiments, the heterologous transcription regulatory sequence is a promoter sequence. In some embodiments, the recombinant construct further comprises a nucleotide sequence encoding a TfpM-associated fimbriae-like protein or a glycosylated fragment thereof disclosed herein, or a fusion protein of the present disclosure comprising a TfpM-associated fimbriae-like protein or a glycosylated fragment thereof operably linked to a nucleotide sequence encoding a TfpM OTase. In some embodiments, the recombinant construct further comprises a nucleotide sequence encoding a TfpM-associated fimbriae-like protein or a glycosylated fragment thereof disclosed herein, or a fusion protein of the present disclosure comprising a TfpM-associated fimbriae-like protein or a glycosylated fragment thereof located at the 5' end of a nucleotide sequence encoding a TfpM OTase and operably linked to that nucleotide sequence. In some embodiments, the fusion protein of the construct also comprises a glycosylated sequence (e.g., ComP or a glycosylated fragment thereof) of an OTase other than TfpM (such as PglB, PglL, PglS). In some embodiments, the coding sequence of the TfpM-associated fimbriae-like protein or a glycosylated fragment thereof, or a fusion protein comprising the TfpM-associated fimbriae-like protein or a glycosylated fragment thereof, is located within 2, 5, 10, 20, 30, 40, or 50 nucleotides of the sequence encoding the TfpM OTase. In some embodiments, the coding sequence of the TfpM-associated fimbriae-like protein or a glycosylated fragment thereof, or a fusion protein comprising the TfpM-associated fimbriae-like protein or a glycosylated fragment thereof, overlaps with an operatively linked nucleotide sequence encoding the TfpM OTase. In some embodiments, the TfpM-associated fimbriae-like protein comprises or is composed of the full-length TfpM-associated fimbriae-like protein. In some embodiments, the TfpM-associated fimbriae-like protein comprises or is composed of a glycosylated fragment of the TfpM-associated fimbriae-like protein smaller than the full-length TfpM-associated fimbriae-like protein. In some embodiments, the length of the fibroin-like protein glycosylation fragment is 3 to 138 amino acids, 10 to 138 amino acids, 20 to 138 amino acids, 50 to 138 amino acids, 100 to 138 amino acids, or 116 to 138 amino acids. In some embodiments, the length of the fibroin-like protein glycosylation fragment is 3 to 139 amino acids, 10 to 139 amino acids, 20 to 139 amino acids, 50 to 139 amino acids, 100 to 139 amino acids, or 116 to 139 amino acids. In some embodiments, the length of the glycosylation fragment is 3 to 140 amino acids, 10 to 140 amino acids, 20 to 140 amino acids, 50 to 140 amino acids, 100 to 140 amino acids, or 116 to 140 amino acids.In some embodiments, the length of the glycosylated fragment is 3 to 22 amino acids, 10 to 22 amino acids, 11 to 22 amino acids, 5 to 21 amino acids, 10 to 21 amino acids, or 11 to 21 amino acids. In some embodiments, the length of the glycosylated fragment is any from 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 amino acids to any from 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids. In some embodiments, the glycosylated fragment is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids in length. In some embodiments, the fimbriae-like protein glycosylated fragment comprises or consists of the following: Pil. Mo (SEQ ID NO:57) or Pil lacking amino acids corresponding to residues 1–28 Mo (Pil Mo Δ28 (SEQ ID NO:58) or a polypeptide containing at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity with SEQ ID NO:57 or SEQ ID NO:58, for example, wherein the C-terminal threonine is replaced by a serine. In some embodiments, the TfpM-associated fimbriae-like protein is selected from the group consisting of: Pil DSM16617 (SEQ ID NO:82), Pil ZZC3-9 (SEQ ID NO:83), Pil TUM15069 (SEQ ID NO:84), Pil AI7 (SEQ ID NO:85), Pil VE-C3 (SEQ ID NO:86), Pil YH01026 (SEQ ID NO:87), Pil CIP102143 (SEQ ID NO:88), Pil AI40 (SEQ ID NO:89), Pil F78 (SEQ ID NO:90), Pil S71 (SEQ ID NO:91), Pil ANC4282 (SEQ IDNO:92), Pil 72-O-c (SEQ ID NO:93), Pil BI730 (SEQ ID NO:94), PilA3K91 (SEQ ID NO:95), Pil CIP102159 (SEQ ID NO:96), Pil junii-65 (SEQ ID NO:97), Pil YZS-X (SEQ ID NO:98), Pil T-3-2 (SEQ ID NO:99) and Pil CIP102637 (SEQ ID NO:100). In some embodiments, the TfpM-associated pilonoid protein or pilonoid glycosylated fragment comprises or consists of the following: an amino acid sequence selected from the group consisting of: Pil DSM16617 (SEQ ID NO:82), Pil ZZC3-9 (SEQ ID NO:83), Pil TUM15069 (SEQ ID NO:84), Pil AI7 (SEQ ID NO:85), Pil VE-C3 (SEQ ID NO:86), Pil YH01026 (SEQ ID NO:87), Pil CIP102143 (SEQ ID NO:88), Pil AI40 (SEQ ID NO:89), Pil F78 (SEQ ID NO:90), Pil S71 (SEQ ID NO:91), Pil ANC4282 (SEQ ID NO:92), Pil 72-O-c (SEQ ID NO:93), Pil BI730 (SEQ ID NO:94), Pil A3K91 (SEQ ID NO:95), Pil CIP102159 (SEQ ID NO:96), Pil junii-65 (SEQ ID NO:97), Pil YZS-X (SEQ ID NO:98), Pil T-3-2 (SEQ ID NO:99), Pil CIP102637 (SEQ ID NO:100), and fragments thereof (e.g., C-terminal fragments) and / or variants wherein the C-terminal threonine is replaced by a serine. Furthermore, in some embodiments, the pifiloid-like protein glycosylated fragment comprises or consists of the following: PilMo pifiloid disulfide ring region (Pil... Mo _DSL, also known as Pil 20(SEQ ID NO: 60) or truncated derivatives thereof, which contain at least the last three amino acids from the C-terminus of the fimbriae, or variants wherein the C-terminal threonine is replaced by a serine (SEQ ID NO: 148). Further, in some embodiments, the fimbriae-like protein glycosylation fragment consists of the following: Pil 20 (SEQ ID NO:60), Pil 19 (SEQ ID NO:133), Pil 18 (SEQ ID NO:134), Pil 17 (SEQ ID NO:135), Pil 16 (SEQ ID NO:136), Pil 15 (SEQ ID NO:109), Pil 14 (SEQ ID NO:137), Pil 13 (SEQ ID NO:110), Pil 12 (SEQ IDNO:138), Pil 11 (SEQ ID NO:139), Pil 10 (SEQ ID NO:112), Pil9 (SEQ ID NO:140), Pil8 (SEQ ID NO:141), Pil7 (SEQ ID NO:113), Pil6 (SEQ ID NO:114), Pil5 (SEQ ID NO:115), Pil4 (SEQ ID NO:116), or Pil3 (SEQ ID NO:117), or variants thereof having one, two, three, four, or five amino acid substitutions and maintaining a C-terminal threonine residue. Further, in some embodiments, the fimbriae-like protein glycosylation fragment consists of the following: Pil 20S (SEQ ID NO:148), Pil 19S (SEQ ID NO:149), Pil 18S (SEQ ID NO:150), Pil 17S (SEQ ID NO:151), Pil 16S (SEQ ID NO:152), Pil 15S (SEQ ID NO:153), Pil 14S (SEQ IDNO:154), Pil 13S (SEQ ID NO:155), Pil 12S (SEQ ID NO:156), Pil 11S (SEQ ID NO:157), Pil10S (SEQ ID NO:158), Pil 9S (SEQ ID NO:159), Pil 8S (SEQ ID NO:160), Pil 7S (SEQ ID NO:161), Pil 6S (SEQ ID NO:162), Pil 5S (SEQ ID NO:163), Pil 4S (SEQ ID NO:164) or Pil 3S(SEQ ID NO:165), or a variant thereof having one, two, three, four, or five amino acid substitutions and maintaining a C-terminal serine residue. In some embodiments, the fusion protein is a fusion protein of the present disclosure. In some embodiments, the recombinant construct further comprises a nucleotide sequence encoding an additional OTase as described elsewhere herein, which is operatively linked to the TpfMOTase. In some embodiments, the recombinant construct further comprises a nucleotide sequence encoding an additional OTase located at the 3' of the TpfM OTase and operatively linked to the TpfM OTase. In some embodiments, the recombinant construct further comprises a nucleotide sequence encoding an additional OTase located at the 5' of the TpfM OTase and operatively linked to the TpfM OTase. In some embodiments, the coding sequence of the additional OTase is within 10, 20, 30, 40, 50, 75, or 100 nucleotides of the sequence encoding the TfpM OTase. In some embodiments, the recombinant construct further comprises a nucleotide sequence encoding PglS OTase, which is operatively linked to TpfM OTase at the 3' end. In some embodiments, the recombinant construct further comprises a nucleotide sequence encoding PglS OTase, which is located at the 3' end of TpfM OTase and operatively linked to TpfM OTase. In some embodiments, the recombinant construct further comprises a nucleotide sequence encoding PglS OTase, which is located at the 5' end of TpfM OTase and operatively linked to TpfM OTase. In some embodiments, the coding sequence for PglS OTase is within 10, 20, 30, 40, 50, 75, or 100 nucleotides of the sequence encoding TfpM OTase. This document further provides a vector comprising the recombinant nucleic acid construct. And this document further provides a host cell comprising the recombinant nucleic acid construct or vector. In some embodiments, the host cell is a bacterial cell. In some embodiments, the host cell is *Escherichia coli*. In some embodiments, the host cell is derived from the genus *Klebsiella*. Furthermore, in some embodiments, the host cell is *Klebsiella pneumoniae*, *Klebsiella heterotropha*, *Klebsiella micrantha*, or *Klebsiella acidogenic*. This document provides a method for producing TfpM OTase, the method comprising culturing host cells (wherein the vector is an expression vector) and recovering TfpM OTase.

[0338] Example

[0339] Crosslinking of bioconjugates with NP / VLP monomers can provide repetitive displays of desired glycans and protein epitopes at a scale smaller than that of a single protein molecule (Liu, Y et al. (2023) Microb Cell Fact 22, 95). The following illustrative examples describe the design and demonstration of bioconjugate-mi3 and bioconjugate-AP205 assemblies. These bioconjugates were generated using two different bacterial O-linked OTases: Acinetobacter benzi PglS and Moraxella osloensis TfpM. PglS and TfpM belong to the recently characterized families of O-linked OTases, which have the broadest range of known glycosubstrates (Harding, CM et al. (2019) Nat Commun 10, 891; Knoot, CJ et al. (2023) Glycobiology, Vol. 33, pp. 57–74). Notably, both PglS and TfpM can transfer glycans with glucose at their reducing ends, allowing these enzymes to be used to generate bioconjugate vaccines against a variety of pathogens whose native polysaccharides have this sugar at the reducing end (Harding, CM et al. (2019) Nat Commun 10, 891; Feldman et al. (2019) PNAS, 116(37) 18655-18663). For this application, *E. coli* maltose-binding protein (MBP) and *Pseudomonas aeruginosa* EPA were engineered to contain PglS-specific or TfpM-specific sequences (Knoot, CJ et al. (2021) Glycobiology, Vol. 31, pp. 1192–1203; Knoot, CJ et al. (2023) Glycobiology, Vol. 33, pp. 57–74), and subsequently glycosylated with the O16O antigen from *E. coli*. The resulting PglS-derived or TfpM-derived MBP or EPA bioconjugates were covalently linked to the NP monomer using the Spytag / Spycatcher technique.

[0340] Example 1. MBP-O16 bioconjugate with SpyTag tag.

[0341] MBP-O16 bioconjugates tagged with SpyTag (“first polypeptide”; for example, Figure 1“Protein 1” was produced in a CLM24-engineered E. coli strain. The MBP-SpyTag fusion protein was expressed alone from the pEXT20 expression plasmid, and WbbL was expressed from the plasmid pMF19 (a derivative of pEXT21). Expression of WbbL restored the production of the O16 O- antigen in most laboratory E. coli strains. The mi3- and AP205-SpyCatcher fusion protein (“second polypeptide”; e.g., Figure 1 Protein 2 was produced alone in C41(DE3) Escherichia coli. After culturing the expression strain in TB medium, the cell pellet was frozen for subsequent lysis and protein purification.

[0342] SpyTag-tagged O16-bioconjugates were purified from periplasmic cell extracts using immobilized metal affinity (Ni) chromatography (IMAC). The IMAC eluent was concentrated and buffer-exchanged before being loaded onto an Akta FPLC instrument for further purification using anion-exchange chromatography. Unglycosylated SpyTag-tagged MBP was separated from glycosylated SpyTag-tagged MBP-O16 using anion-exchange chromatography. The fractions containing glycosylated SpyTag-tagged MBP-O16 were combined, concentrated, and quantified using a BCA assay kit. The purified conjugates were stored at -80°C in Tris-buffered saline (TBS) until use in isopeptide bond formation reactions.

[0343] Mi3 and AP205 nanoparticles (NPs) were purified by whole-cell lysis via sonication. The lysates were centrifuged at 18,000 x G and then loaded onto IMAC resin as described above. The eluent was concentrated and loaded onto an FLPC-sized size exclusion column to separate fully assembled VLPs and NPs from unassembled free monomers. Fractions containing fully assembled VLPs or NPs were determined based on the mass of known standards run on the same column. These fractions were combined, concentrated, quantified, and stored in TBS at 4°C until use in isopeptide bond formation reactions.

[0344] For the isopeptide formation reaction, two purified proteins (SpyTag-tagged O16-bioconjugate and VLP / NP) were mixed in TBS buffer at a molar ratio of approximately 1:1 or approximately 2:1 of mi3 / AP205:SpyTag-tagged O16-bioconjugate. These reactions were incubated at 22°C for 1.5 h (mi3) or 3 h (AP205). The reactions were then analyzed using SDS-PAGE with Coomassie Brilliant Blue staining, Western blotting, and size exclusion chromatography using a Sephacryl S-400HR column.

[0345] Example 2. EPA-O16 bioconjugate with Spytag tag.

[0346] The generation and purification of the Spytag-tagged EPA-O16 bioconjugates were carried out in the same manner as the MBP-O16 bioconjugates disclosed elsewhere in this paper. All versions of the Spytag-tagged EPA were generated from the pEXT20 expression plasmid.

[0347] Cross-linking reactions between Spytag-tagged O16 bioconjugates and purified mi3-Spycatcher or AP205-Spycatcher yielded higher molecular weight covalently cross-linked proteins, as determined by Coomassie brilliant blue staining, Western blotting, and size exclusion chromatography. Use of E. coli O16 antiserum and antiprotein antibodies indicated that higher molecular weight species consisted of O16 glycan-linked proteins (e.g., Figure 3 and Figure 6 Typically, based on the intensity of cross-linked protein bands in Coomassie Brilliant Blue staining gels or Western blotting, and the observation of reduced NP / VLP monomer or bioconjugate bands retained after isopeptide bond reactions, TfpM-derived Spytag-tagged bioconjugates react more completely with Spycatcher-mi3 or AP205 to form assemblies bound by isopeptide bonds.

[0348] Size exclusion chromatography showed that the bioconjugate-NP / VLP assembly was larger in mass than the assembly before the isopeptide bond reaction. Figure 3 ).

[0349] Not all Spytag-tagged EPA bioconjugates form detectable amounts of cross-linked, isopeptide-bound species with mi3-Spycatcher or AP205-Spycatcher. Of the three Spytag-tagged EPA variants tested, only EPA-Spytag-v1 produced isopeptide-bound species. Figure 9 Furthermore, only a subset of EPA-Spytag-v1 protein adaptor variants were able to form species with mi3-Spycatcher via isopeptide bonds. Figure 10 ).

[0350] The time progression of isopeptide bond formation indicates that, based on the relative intensity of the protein bands in the Western blot, the reaction between EPA-Spytag-v1 and mi3-Spycatcher is essentially completed (>80%) within one hour of the reaction initiation. Figure 11 ).

[0351] Example 3. Immunization with glycosylated ComP bioconjugates to elicit an immune response.

[0352] T-cell-dependent immune responses to conjugate vaccines are characterized by the secretion of high-affinity IgG1 antibodies (Avci, FY, Li, X., Tsuji, M. and Kasper, DLNat Med 17, 1602-1609 (2011)). The immunogenicity of the CPS14-ComP bioconjugate in a mouse vaccination model was evaluated (WO / 2020 / 131236, which is incorporated herein by reference in its entirety). Serum from mice vaccinated with the CPS14-ComP bioconjugate showed a significant increase in CPS14-specific IgG titers, but no increase in IgM titers. Furthermore, HRP-tagged anti-IgG subtype secondary antibodies were used to determine which IgG subtypes showed elevated titers. IgG1 titers appeared to be higher than other subtypes.

[0353] Next, a second vaccination trial was conducted, comparing trivalent CPS8-, CPS9V-, and CPS14-ComP bioconjugates with the current standard of care, PREVNAR. Immunogenicity. Serotypes 9V and 14 are included in PREVNAR. In the middle, and in accepting PREVNAR In immunized mice, elevated IgG titers were observed against both serotypes. Monovalent immunization against serotype 14 also showed significant induction of serotype-specific IgG titers, similar to the initial immunization. Compared to controls, mice receiving the trivalent bioconjugate showed elevated serotype-specific IgG titers in both groups, as expected; serum on day 49 showed significantly higher IgG titers for serotypes 8 and 14 compared to serotype 9V. Nevertheless, IgG titers against 9V remained significantly higher than in placebo.

[0354] Example 4. Cloning and Plasmid Assembly

[0355] All primers and oligonucleotides used in this study are listed in Table 2. The working concentrations of antibiotics used in liquid culture and LB agar plates were as follows: ampicillin (Amp), 100 μg / mL; kanamycin (Kan), 20 μg / mL; tetracycline (Tet), 10 μg / mL; and spectinomycin (Sp), 50 μg / mL. To clone the tfpM fimbriae-OTase gene, HiFigblock (Integrated DNA Technologies, IDT) was ordered, with 25-base-pair overlap at the ends, for Gibson assembly with the PCR-linearized plasmid. The plasmid backbones of these fragments were amplified from the pEXT20 plasmid (pVNM57) (Dykxhoorn, DM, et al. (1996) Gene 177, 133-136), which encodes the Pseudomonas aeruginosa EPA gene under the control of the tac promoter (Knoot, CJ, et al. (2021) Glycobiology 31, 1192-1203). The EPA gene was deleting residue E553, resulting in toxin inactivation. The linearized plasmids were individually mixed with each of the synthetic tfpM gBlocks and assembled using the NEBuilder HiFi DNA Assembly Kit (New England Biolabs, NEB). After assembly, the plasmids were transformed into E. coli Stellar cells (Takara Bio) by heat shock, overgrown at 37°C for one hour, and plated on LB agar supplemented with Amp. Single colonies were picked and grown in LB medium containing appropriate antibiotics, and plasmids were isolated using the GeneJet Plasmid Mini-Prep Kit (Thermo Fisher). All plasmids were sequenced and validated by Sanger sequencing (Genewiz). The expression of *Moraxella osloensis* 1202EPA-PilΔ28 fusion and TfpM was performed. Mo The plasmid was named pVNM227. To generate Pil MoFor site-directed mutants, the inventors designed overlapping PCR primers to introduce the necessary codon changes into the fimbriae gene and amplify each fragment from the pVNM227 plasmid. The resulting PCR products were digested with DpnI (NEB) at 37°C for 30 min and purified from agarose gels using the Pure-Link Gel Extraction Kit (Thermo Fisher). To insert the truncated fimbriae gene region, complementary oligonucleotides homologous to the pVNM227 PCR product with a 25 bp overlap at the ends were ordered. The oligonucleotides were resuspended in purified water, mixed, and annealed together in a thermal cycler: heated to 98°C for 5 min, then slowly cooled to 4°C at 0.1°C / min. The annealed oligonucleotides were diluted 1:5 in water and assembled with PCR-linearized pVNM227 using the NEBuilder HiFi DNA Assembly Kit (NEB). The resulting DNA was transformed into Stellar cells, and the plasmid was isolated and validated as described above. The plasmid contains the encoding EPA-Pil. 20 The plasmid for the TfpM construct is named pVNM297. EPA-Pil with an N-terminal His tag is constructed as follows: 20 Variant: pVNM297 was linearized using PCR and used in Gibson assembly with complementary annealed oligonucleotides containing a 6xHis coding region and terminal homologous regions to generate pVNM291. pVNM167 was generated by digesting the previously described EPA with SalI. iGTccThe plasmid (Knoot, CJ et al. (2021) Glycobiology 31, 1192-1203) was generated. The purified SalI fragment was Gibson assembled with the pglS gene, amplified from Acinetobacter bengal ADP1 gDNA, which has its native 100 bp 5' UTR. pVNM245 was generated from the pVNM167 template via a separate PCR reaction to amplify a product with overhangs for Gibson assembly: (i) a vector backbone containing PglS and EPA, and an iGT; (ii) a second iGT for integration between E548 and G549; and (iii) a C-terminus of the EPA downstream of the iGT. Plasmid pVNM337 was created as follows: tfpM was amplified from pVNM291 using primers EPA 3'F1 and pglS-tfpM R1, and the product was cloned into PCR-linearized pVNM167. This PCR-linearized pVNM167 was then amplified using pglS 5'F1 and EPA 3'R1. A phylogenetic tree of TfpM and fimbriae proteins was generated using the phylogeny.fr server (website: phylogeny.fr / ). This server used MUSCLE, PhyML, and TreeDyn for sequence alignment, tree calculation, and image generation, respectively.

[0356] Example 5. Expression of glycans and cloning of the O2a glycan gene in Klebsiella pneumoniae

[0357] Streptococcus pneumoniae CPS8 glycan from plasmid pB8(Tet R The LT2 glycan of *Salmonella enterica* is expressed in plasmid pPR1347 (Kay, EJ et al. (2016) Open Biology 6, 150243). R The *E. coli* O16 wbbL gene is expressed in plasmid pMF19 (Neal, BL et al. (1993) Journal of Bacteriology 175, 7115-7118). RGBSIII glycan is expressed in pBBR1MCS2 derivatives (Feldman, MF et al. (2005) Proceedings of the National Academy of Sciences of the United States of America 102, 3016) and expressed in pBBR1MCS2 derivatives (Duke, JA et al. (2021) ACS Infectious Diseases 7, 3111-3123). Bioconjugation with Klebsiella pneumoniae O2a O-antigen has not been previously reported. To clone the genes encoding the mechanisms required for O2a glycan synthesis, the wzm, wzt, wbbM, glf, wbbN, and wbbO genes from the genomic DNA of Klebsiella pneumoniae strain NTUH K2044 were amplified by PCR (Clarke, BR et al. (2018) Journal of Biological Chemistry 293, 4666-4679). Klebsiella pneumoniae was cultured overnight in LB medium until saturation, and then genomic DNA was isolated using the Wizard Genomic DNA Purification Kit (Promega). The plasmid backbone of the O2a cluster was extracted from plasmid pBBR1MCS2 (Kan R The PCR products from these reactions were amplified in (Kovach, ME et al. (1995) Gene 166, 175-176). Primers used for these reactions are listed in Table 2. Gibson assembly of the PCR products from these reactions was performed using the NEBuilder HiFi DNA Assembly Kit (NEB). Stellar cells were transformed, and plasmids were isolated and validated as described above.

[0358] Example 6. Biological conjugation and protein blotting

[0359] The *E. coli* strains used for bioconjugation experiments were SDB1 or CLM24 (Feldman, MF et al. (2005) Proceedings of the National Academy of Sciences of the United States of America 102, 3016). SDB1 is a W3110 *E. coli* derivative with mutations in the genes encoding WecA (a glycosyltransferase that initiates the synthesis of endogenous *E. coli* O16 antigen) and WaaL (an enzyme that transfers Und-PP-linked glycan lipid A-core sugar to produce LPS). CLM24 is a W3110 derivative with only the waaL deletion absent. Elimination of these genes prevents interaction between the heterologous bioconjugation system and the endogenous *E. coli* glycosylation pathway. To prepare *E. coli* strains for bioconjugation, the inventors used competent cells prepared as described previously (Knoot, CJ et al. (2021) Glycobiology 31, 1192-1203), electroporated plasmids, and then overgrown them in SOB medium at 37°C. Cells were plated on LB agar containing appropriate antibiotics. The next day, 8–10 colonies were picked and inoculated into LB or TB agar containing antibiotics and incubated overnight with shaking at 30°C. The following morning, the starting culture was inoculated into 30 mL of medium in a 125 mL Erlenmeyer flask, or into 1 L of medium in a 2 L flask, until the initial optical density (OD) at 600 nm was reached. 600 The value was 0.05. The culture was grown with shaking at 175 RPM until the OD value was reached. 600 Once the culture reaches 0.4–0.6, induce the culture with 1 mM IPTG. Unless otherwise specified, all bioconjugation experiments are performed at 30°C. After overnight induction (total growth time 20–24 hours), measure the OD. 600 Cells with 0.5 OD units were precipitated for analysis.

[0360] Resuspend the cell pellet in 100 μl of 1X Laemml buffer (Biorad) and boil at 100 °C for 10 min. Briefly centrifuge the boiled sample at 10,000 rcf and transfer equal volumes (relative to the same OD per channel) to the appropriate concentrations. 600Normalized samples were loaded onto a 7.5% Mini-Protean TGX gel (Biorad) for SDS-PAGE separation. Proteins were transferred to a nitrocellulose membrane using a semi-dry electrode system and blocked with Intercept blocking buffer (Li-Cor) for one hour. The membrane was then incubated with primary antibodies in a 1:1 blocking and TBST solution for 45 min. For protein detection, commercially available rabbit anti-EPA antibody and mouse anti-6xHis antibody (Millipore-Sigma) were used. Rabbit glycan antibodies against CPS8, GBSIII, and O16 were purchased from SSIDidiagnostica. Klebsiella pneumoniae rabbit O2a antibody was generously donated by Professor Chris Whitfield (University of Guelph) (Clarke, BR et al. (2018) Journal of Biological Chemistry 293, 4666-4679). Salmonella group B rabbit antibody was purchased from BD Biosciences. After the initial incubation, the membrane was washed three times with TBST buffer for a total of 15 min. Subsequently, the membrane was incubated with the secondary antibodies IRDye 680RD goat anti-mouse and / or IRDye 800CW goat anti-rabbit (Li-Cor) in a 1:1 blocking buffer and TBST for 30 min. After a final 15 min TBST wash, the membrane was imaged using a Li-Cor Odyssey CLx.

[0361] Example 7. Recombinant Moraxella osloensis Pil Mo Lys-C digestion of Δ28

[0362] In-gel digestion was performed according to the protocol of Shevchenko et al. (Shevchenko, A. et al. (2006) NatProtoc1, 2856-2860), with slight modifications. The glycosylated EPA-PilΔ28 separated from the gel was excised and decolorized twice with a decolorizing solution (50 mM NH4HCO3, 50% ethanol) at 750 RPM for 10 min at room temperature. The decolorized band was then dehydrated with 100% ethanol for 10 min, dried by vacuum centrifugation for 10 min, and then rehydrated in 10 mM DTT in 50 mM NH4HCO3. The reduction reaction was carried out at 56 °C for 60 min, after which the gel band was dehydrated twice with 100% ethanol for 10 min to remove residual reduction buffer. The reduced sample was then alkylated further in 55 mM iodoacetamide in 50 mM NH4HCO3 for 45 min in the dark. The alkylated sample was then washed four times with 50 mM NH4HCO3 for 10 min, followed by washing with 100% ethanol, then with 50 mM NH4HCO3, and finally with 100% ethanol. It was then dried by vacuum centrifugation. The dried alkylated sample was then rehydrated with 20 ng / μl Lys-C endopeptide (Wako Chemicals) in 40 mM NH4HCO3 at 4 °C for 1 hr. Excess Lys-C was removed, and the gel fragments were immersed in 40 mM NH4HCO3 and incubated overnight at 37 °C. Using C... 18 The peptides were concentrated and desalted at the column tip (Ishihama, Y. et al. (2006) J Proteome Res 5, 988-994; Rappsilber, J. et al. (2007) Nat Protoc 2, 1896-1906), then eluted in buffer B (0.5% acetic acid, 80% acetonitrile (ACN)), dried, and stored at -20°C for LC-MS analysis.

[0363] Example 8. Analysis of recombinant Moraxella osloensis Pil using reversed-phase LC-MS Mo Δ28

[0364] C 18 The concentrated digest was resuspended in buffer A* (0.1% TFA, 2% ACN) and separated using a dual-column chromatography system containing a PepMap 100 C column. 18 20mm×75μm trapping column and PepMap C 18A 500 mm × 75 μm analytical column (Thermo Fisher Scientific) was used. The sample was concentrated onto the trap column using 0.1% formic acid (FA) at a rate of 5 μl / min for 5 min. Then, using a Dionex Ultimate 3000UPLC (Thermo Fisher Scientific), the concentrations of buffer A (2% DMSO, 0.1% FA) and buffer B (78% ACN, 2% DMSO, and 0.1% FA) were varied and injected via the analytical column at a rate of 300 nl / min into an Orbitrap Fusion LC system equipped with a FAIMS Pro interface. TM Lumos TM Tribrid TM In a mass spectrometer (Thermo Fisher Scientific). Identification of potential glycopeptides utilized a 140-minute analytical run, while targeting analysis utilized a 60-minute run. In the identification analytical run, the buffer composition was changed from 3% Buffer B to 28% Buffer B over 120 min, from 28% Buffer B to 40% Buffer B over 9 min, from 40% Buffer B to 100% Buffer B over 3 min, then the composition was held at 100% Buffer B for 2 min, then decreased to 3% Buffer B over 2 min, and held at 3% Buffer B for another 8 min. Lumos TM The mass spectrometer operated in step-FAIMS data-dependent mode with three different FAIMS CVs (-25, -45, and -65), as previously described (Ahmad Izaham, AR et al. (2021) J Proteome Res 20, 599-612), switching between acquiring a single Orbitrap MS scan (60k resolution) every 1.5 sec, followed by an Orbitrap HCD scan (maximum fill time 120 ms, AGC 2 × 10⁻⁶) under each of the three FAIMS CVs. 5 The Orbitrap MS-MS scan resolution was 30k, and the NCE was 25, 30, and 45. (For Pil...) MoGlycopeptides were characterized using a targeted analytical run. The buffer composition was changed from 3% Buffer B to 15% Buffer B over 30 min, then to 30% Buffer B over 10 min, and finally to 80% Buffer B over 5 min. The composition was then held at 100% Buffer B for 5 min, followed by a 1 min decrease to 3% Buffer B, and finally held at 3% Buffer B for 9 min. Parallel reaction monitoring was performed using FAIMS CV at -45°C, with HCD (maximum fill time 250 ms, AGC 2.5 × 10⁻⁶). 5 The Orbitrap MS-MS scan resolution is 30k, and the NCE is 15, 30, 35) and the EThcD (maximum fill time is 250ms, AGC is 2.5×10) is also supported. 5 Orbitrap MS-MS scans at a resolution of 30k and using HexHexA-modified glycopeptides 762 FLPANCRGT 770 The calibrated charge-dependent ETD parameter of the +2 charge state (687.2972 m / z) controlled by ETD reaction time (Rose, CM et al. (2015) J Am Soc Mass Spectrom 26, 1848-1857).

[0365] Example 9. Pil Mo Open search for Δ28 and annotation of C-terminal peptides modified with HexHexA

[0366] As mentioned earlier, an open database search was used to complete the search for Pil. Mo Identification of glycosylation events (Lewis, JM et al. (2021) J Vis Exp). In short, data files were processed in FragPipe (version 17.1) using MSfragger 3.4 (Polasky, DA et al. (2020) Nat Methods 17, 1125-1132; Kong, AT et al. (2017) Nat Methods 14, 513-520) to search for *Moraxella osloensis* Pil. Mo Sequence (NCBI accession number: WP_156627541.1). A search was performed using “Lys-C” enzyme specificity, with cysteine ​​carbamoyl methylation as the fixed modification and methionine oxidation as the variable modification, allowing up to two missed cleavages. A mass tolerance of 0 to 2000 Da (referred to as δmass) was allowed to identify potential glycosylation events. Manual examination of the C-terminal peptide was performed. 762 FLPANCRGT 770δmass was observed on (SEQ ID NO:61) to identify potential glycosylation events. Glycopeptides modified with HexHexA were manually extracted using a Freestyle Viewe (1.7SP1, Thermo Fisher Scientific). 762 FLPANCRGT 770 (SEQ ID NO:62) Corresponding parallel reaction monitoring results, MS / MS data were annotated using the Interactive Peptide Spectral Annotator (Brademan, DR et al. (2019) Mol Cell Proteomics 18, S193-S201) (Website: interactivepeptidespectralannotator.com / PeptideAnnotator.html). Spectral annotation allowed modification of terminal T residues with HexHexA (338.0849 Da) and Hex (162.0528 Da). The obtained MS data and search results have been stored in the PRIDE ProteomeXchange consortium database (Perez-Riverol, Y. et al. (2019) Nucleic Acids Res47, D442-D450; Perez-Riverol, Y. et al. (2015) Proteomics 15, 930-949) and are accessible using the identifier PXD033468.

[0367] Example 10. Purification of biological conjugate proteins

[0368] Cells for protein purification were grown in 1 L TB medium, and bioconjugates were isolated using an osmotic shock protocol. After overnight growth and induction, the cells were pelleted by centrifugation and washed in 0.9% NaCl. The washed cell pellet was resuspended in 200 mM Tris-HCl pH 8.5, 100 mM EDTA, and 25% sucrose and incubated at 4 °C for 30 min. The cells were pelleted again by centrifugation at 4,700 rcf for 30 min, and the resulting pellet was then resuspended in 20 mM Tris-HCl pH 8.5 and incubated at 4 °C for 45 min. The suspension was then centrifuged at 18,000 rcf for 30 min. The supernatant containing the periplasmic fraction was concentrated and directly loaded onto an FPLC anion exchange column, or, for His-tagged EPA-PilΔ28 bioconjugates, purified using a nickel IMAC, as previously described (Knoot, CJ et al. (2021) Glycobiology 31, 1192-1203). The periplasmic extract or IMAC eluent was concentrated and buffer-exchanged to 20 mM Tris-HCl pH 8.0, filtered through a 0.2 μm PES filter, and then loaded onto a Cytiva column equipped with a SOURCE 15Q 4.6 / 100PE anion exchange column. The bioconjugate was eluted using a pure FPLC instrument (Cytiva). The bioconjugate was eluted using a stepwise gradient of buffer A (20 mM Tris pH 8) and buffer B (20 mM Tris pH 8, 1 M NaCl) at 2 mL / min, with buffer B increasing by 5% from 0% to 25% in 10 column volumes for each concentration. The bioconjugate for immunization was further purified using a Superdex 200 Increase 10 / 300GL column. The concentrated bioconjugate collected from the anion exchange column was loaded onto a pre-equilibrated Superdex 200 column in PBS buffer and eluted at a flow rate of 0.75 mL / min. The fraction containing the purified bioconjugate was collected, concentrated, and stored frozen at -80°C. Protein concentrations for immunization and Western blotting were determined using a Pierce BCA protein assay kit (Thermo Fisher). The polysaccharide-to-protein ratio used for vaccine dosing calculations was determined using the method described by Duke et al. (Duke, JA et al. (2021) ACS Infectious Diseases 7, 3111-3123).

[0369] Example 11. Immunization of mice

[0370] All mouse immunizations were conducted in accordance with ethical guidelines for animal testing and research. Experiments were performed at Washington University School of Medicine in St. Louis according to institutional guidelines and with approval from the Washington University Institutional Animal Care and Use Committee. Five-week-old female CD-1 hybrid mice (Charles River Laboratories) were subcutaneously injected with 100 μL of the vaccine formulation on days 0, 14, and 28. Vaccination groups were either the 291-only group (5 μg protein) or the GBSIII-291 group (5 μg protein, 1 μg polysaccharide). Serum was collected from mice on days 0, 14, 28, and 42. All vaccines were administered using… 2% aluminum hydroxide gel (InvivoGen) was prepared at a ratio of 1:9 (50 μL of vaccine to 5.5 μL of alum in 44.5 μL of 1x sterile phosphate buffered saline).

[0371] Example 12. Enzyme-linked immunosorbent assay (ELISA)

[0372] IgG kinetic titers were determined using enzyme-linked immunosorbent assay (ELISA). In short, the 96-well plate (TRPImmunomaxi plate) was prepared using approximately 10... 6 Triple copies of glycoengineered *E. coli* expressing GBSIII capsular polysaccharide were coated overnight in sodium carbonate buffer. The *E. coli* strains used for coating were grown as described above, and after overnight induction to induce GBSIII expression, washed and diluted to coat the plates. Wells were blocked with 1% BSA in PBS and washed with 0.05% PBS-Tween (PBST), with all subsequent washes being identical. Mouse serum diluted 1:100 was added to the wells at room temperature and incubated for 1 hour, followed by washing. Total IgG titers were detected by HRP-conjugated anti-mouse IgG (GE Lifesciences, 1:5000 dilution) added to the wells at room temperature and incubated for 1 hour. After washing, the plates were developed using 3,3′,5,5′tetramethylbenzidine (TMB) substrate (Biolegend) and terminated with 2N H2SO4. Density was determined at 450 nm using a microplate reader (Bio-Tek). Total IgG product was determined using IgG standards to generate a standard curve for data fitting. Standard wells were coated with IgG in sodium carbonate buffer and then processed in the same manner as the sample wells. All wells were normalized relative to blank wells, which were processed in the same way as all sample wells, excluding those receiving primary mouse serum. Significance was determined using the Mann-Whitney nonparametric test, P < 0.05.

[0373] The scope and extent of this disclosure should not be limited to any of the exemplary embodiments described above, but should be defined solely by the following claims and their equivalents.

[0374] *****

[0375] Certain embodiments of this disclosure may be defined in any of the following numbered paragraphs:

[0376] 1. A fusion protein comprising: (i) a glycosylated fragment, and (ii) a first polypeptide tag, wherein the first polypeptide tag is capable of spontaneously binding to a second polypeptide tag to form an isopeptide bond;

[0377] Optionally, the length of the glycosylated fragment is at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30 or 40 amino acids;

[0378] Optionally, the length of the glycosylated fragment does not exceed 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, 40, 50, 60, 80 or 100 amino acids;

[0379] Optionally, the length of the glycosylated fragment is any from 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24 or 30 amino acids to any from 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30 or 40 amino acids; and

[0380] Optionally, the fusion protein includes a carrier protein.

[0381] 2. The fusion protein according to paragraph 1, wherein the fusion protein is a glycoconjugate comprising a sugar covalently linked to the fusion protein via the glycosylation fragment;

[0382] Optionally, the sugar is covalently linked to the glycosylated fragment via N-linking, O-linking, or C-linking; and

[0383] Optionally, the glycoconjugate is immunogenic.

[0384] 3. The fusion protein according to paragraph 1 or 2, wherein the first polypeptide tag is fused to the N-terminus or C-terminus of the fusion protein by translation.

[0385] 4. The fusion protein according to paragraph 1 or 2, wherein the first polypeptide tag is internally fused to the fusion protein by translation;

[0386] Optionally, the first polypeptide tag is internally fused into the sequence of the carrier protein via translation.

[0387] 5. The fusion protein according to any one of paragraphs 1 to 4, wherein the glycosylated fragment is fused to the N-terminus or C-terminus of the fusion protein in a translational manner.

[0388] 6. The fusion protein according to any one of paragraphs 1 to 5, wherein the glycosylated fragment is internally fused to the fusion protein in a translational manner;

[0389] Optionally, the glycosylated fragment is internally fused into the sequence of the carrier protein via translation.

[0390] 7. The fusion protein according to any one of paragraphs 1 to 6, wherein the first polypeptide tag is SpyTag (SEQ ID NO:416), SpyTag002 (SEQ ID NO:417), SpyTag003 (SEQ ID NO:418) or DogTag (SEQ ID NO:419);

[0391] Optionally, the SpyTag, Spytag002, or Spytag003 is fused to the N-terminus or C-terminus of the fusion protein via translation.

[0392] Optionally, the DogTag is internally fused into the fusion protein via translation.

[0393] 8. The fusion protein according to any one of paragraphs 1 to 7, wherein the glycosylated fragment is a ComP glycosylated fragment;

[0394] Optionally, the ComP glycosylated fragment comprises or consists of the following: the amino acid sequence CTGVTQIASGASAATTNVASAQC (SEQ ID NO:412) or a fragment containing the amino acid ASA.

[0395] or variants thereof, wherein the variant contains amino acid ASA at positions 11-13 of SEQ ID NO:412 and has one, two, three, four, five or six amino acid substitutions, additions and / or deletions;

[0396] Optionally, the ComP glycosylated fragment comprises or consists of the following amino acid sequence:

[0397] iGTccΔ0-1CTGVTQIASGASAATTNVASAQ(SEQ ID NO:232);

[0398] iGTccΔ1-0TGVTQIASGASAATTNVASAQC(SEQ ID NO:243);

[0399] iGTccΔ1-1TGVTQIASGASAATTNVASAQ(SEQ ID NO:244);

[0400] iGTccΔ1-2TGVTQIASGASAATTNVASA(SEQ ID NO:245);

[0401] iGTccΔ2-1GVTQIASGASAATTNVASAQ(SEQ ID NO:256);

[0402] iGTccΔ2-2GVTQIASGASAATTNVASA(SEQ ID NO:257);

[0403] iGTccΔ2-3GVTQIASGASAATTNVAS(SEQ ID NO:258);

[0404] iGTccΔ3-2VTQIASGASAATTNVASA(SEQ ID NO:269);

[0405] iGTccΔ3-3VTQIASGASAATTNVAS (SEQ ID NO:270);

[0406] iGTccΔ3-4VTQIASGASAATTNVA (SEQ ID NO:271);

[0407] iGTccΔ4-3 TQIASGASAATTNVAS (SEQ ID NO:282);

[0408] iGTccΔ4-4 TQIASGASAATTNVA (SEQ ID NO:283);

[0409] iGTccΔ4-5 TQIASGASAATTNV (SEQ ID NO:284);

[0410] iGTccΔ5-4 QIASGASAATTNVA (SEQ ID NO:295);

[0411] iGTccΔ5-5 QIASGASAATTNV (SEQ ID NO:296);

[0412] iGTccΔ5-6 QIASGASAATTN (SEQ ID NO:297);

[0413] iGTccΔ6-5 IASGASAATTNV (SEQ ID NO:308); or

[0414] iGTccΔ6-6 IASGASAATTN (SEQ ID NO:309),

[0415] Or a variant thereof, the variant comprising the amino acid ASA corresponding to positions 11-13 of SEQ ID NO:412, and having one, two, three, four, five or six amino acid substitutions, additions and / or deletions.

[0416] 9. The fusion protein according to any one of paragraphs 1 to 7, wherein the glycosylated fragment is a TfpM-associated fimbriae protein glycosylation fragment;

[0417] Optionally, the TfpM-associated fimbriae glycosylation fragment comprises or consists of the following: the PilMo fimbriae disulfide ring region (SEQ ID NO:413) or a fragment thereof, the fragment comprising at least the last three amino acids from the C-terminus of the TfpM-associated fimbriae;

[0418] Or a variant thereof, the variant comprising the last three amino acids from the C-terminus of the TfpM-associated fimbriae, and having one, two, three, four, five, or six amino acid substitutions, additions, and / or deletions.

[0419] 10. The fusion protein according to any one of paragraphs 1 to 7, wherein the glycosylated fragment is a PilE glycosylated fragment;

[0420] Optionally, the PilE glycosylated fragment comprises or consists of the following: amino acid SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO:414) or a fragment thereof, the fragment containing at least the amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414.

[0421] or variants thereof, wherein the variant contains at least the amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414, and has one, two, three, four, five or six amino acid substitutions, additions and / or deletions.

[0422] 11. The fusion protein according to any one of paragraphs 1 to 7, wherein the glycosylated fragment is a PglB glycosylated fragment;

[0423] Optionally, the PglB glycosylated fragment comprises or consists of the following: a common motif amino acid sequence X1 X2 N X3 X4, wherein X1 is D or E, X2 is any amino acid other than proline, X3 is any amino acid other than proline, and X4 is S or T.

[0424] 12. The fusion protein according to any one of paragraphs 1 to 7, wherein the glycosylated fragment is a PilA glycosylated fragment;

[0425] Optionally, the PilA glycosylated fragment comprises or consists of the following: the PilA fimbriae disulfide ring region (SEQ ID NO:415) or a fragment thereof, the fragment comprising at least the last three amino acids from the C-terminus of PilA;

[0426] Or a variant thereof, the variant comprising at least the last three amino acids from the end of the PilA, and having one, two, three, four, five or six amino acid substitutions, additions and / or deletions.

[0427] 13. The fusion protein according to any one of paragraphs 1 to 7, wherein the glycosylated fragment is an STT3 glycosylated fragment;

[0428] Optionally, the STT3 glycosylated fragment comprises or consists of the following: a common motif amino acid sequence N X1 X2, wherein X1 is any amino acid other than proline, and X2 is S or T.

[0429] 14. The fusion protein according to any one of paragraphs 1 to 7, wherein the glycosylated fragment is an N-linked glycosyltransferase glycosylation fragment;

[0430] Optionally, the N-linked glycosyltransferase glycosylation fragment comprises or consists of the following: a common motif amino acid sequence N X1 X2, wherein X1 is any amino acid and X2 is S or T.

[0431] 15. The fusion protein according to any one of paragraphs 1 to 7, wherein the glycosylated fragment is an O-linked glycosyltransferase glycosylation fragment;

[0432] Optionally, the O-linked glycosyltransferase glycosylation fragment comprises or consists of the following: a fragment rich in serine or threonine repeat sequences from a serine-rich repeat (SRR) adhesin derived from streptococci or staphylococci.

[0433] Optionally, the O-linked glycosyltransferase glycosylation fragment comprises or consists of a serine- or threonine-rich repeat sequence of the adhesin GspB from Streptococcus gasseri.

[0434] 16. The fusion protein according to any one of paragraphs 1 to 15, wherein the fusion protein comprises two or more glycosylated fragments;

[0435] Optionally, the fusion protein comprises at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 or 24 glycosylated fragments;

[0436] Optionally, the fusion protein comprises any one to any one of 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 glycosylated fragments.

[0437] Optionally, at least one glycosylated fragment is located at the N-terminus or C-terminus of the fusion protein and at least one glycosylated fragment is internally located within the fusion protein;

[0438] Optionally, at least two glycosylated fragments are internally located within the fusion protein;

[0439] and / or

[0440] Optionally, one of the glycosylated fragments is located at the N-terminus of the fusion protein, and the glycosylated fragment is located at the C-terminus of the fusion protein.

[0441] 17. The fusion protein as described in paragraph 16,

[0442] The two or more glycosylated fragments therein are identical;

[0443] At least one of the two or more glycosylation fragments is different; or

[0444] Each glycosylation fragment in the glycosylation fragment is different;

[0445] Optionally, one of the glycosylation fragments is a ComP glycosylation fragment, and the other of the glycosylation fragments is a TfpM-associated fimbriae protein glycosylation fragment;

[0446] Optionally, the fusion protein is a glycoconjugate comprising two or more sugars covalently linked to the fusion protein via the two or more glycosylation fragments;

[0447] Optionally, the fusion protein comprises at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 or 24 covalently linked sugars;

[0448] Optionally, the fusion protein comprises any one of 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23 covalently linked sugars to any one of 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 covalently linked sugars; and

[0449] Choose any location

[0450] The two or more sugars mentioned therein are identical;

[0451] At least one of the two or more sugars is different; or

[0452] Each of the sugars mentioned is different.

[0453] 18. The fusion protein according to any one of paragraphs 1 to 17, wherein the fusion protein comprises a carrier protein selected from the group consisting of: Escherichia coli maltose-binding protein, Pseudomonas aeruginosa exotoxin A (EPA), Pseudomonas aeruginosa PcrV, CRM197, Haemophilus influenzae protein D, cholera toxin B subunit or tetanus toxin, and fragments thereof.

[0454] 19. A composition comprising a polypeptide pair, said polypeptide pair c...

Claims

1. A fusion protein comprising: (i) a glycosylated fragment, and (ii) a first polypeptide tag, wherein the first polypeptide tag is capable of spontaneously binding to a second polypeptide tag to form an isopeptide bond; Optionally, the length of the glycosylated fragment is at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30 or 40 amino acids; Optionally, the length of the glycosylated fragment does not exceed 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30, 40, 50, 60, 80 or 100 amino acids; Optionally, the length of the glycosylated fragment is any from 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24 or 30 amino acids to any from 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 24, 30 or 40 amino acids; and Optionally, the fusion protein includes a carrier protein.

2. The fusion protein of claim 1, wherein the fusion protein is a glycoconjugate comprising a sugar covalently linked to the fusion protein via the glycosylated fragment; Optionally, the sugar is covalently linked to the glycosylated fragment via N-linking, O-linking, or C-linking; and Optionally, the glycoconjugate is immunogenic.

3. The fusion protein of claim 1, wherein the first polypeptide tag is fused to the N-terminus or C-terminus of the fusion protein by translation.

4. The fusion protein of claim 1, wherein the first polypeptide tag is internally fused into the fusion protein via translation; Optionally, the first polypeptide tag is internally fused into the sequence of the carrier protein via translation.

5. The fusion protein of claim 1, wherein the glycosylated fragment is fused to the N-terminus or C-terminus of the fusion protein via translation.

6. The fusion protein of claim 1, wherein the glycosylated fragment is internally fused into the fusion protein via translation; Optionally, the glycosylated fragment is internally fused into the sequence of the carrier protein via translation.

7. The fusion protein according to claim 1, wherein the first polypeptide tag is SpyTag (SEQ ID NO:416), SpyTag002 (SEQ ID NO:417), SpyTag003 (SEQ ID NO:418) or DogTag (SEQ ID NO:419); Optionally, the SpyTag, Spytag002, or Spytag003 is fused to the N-terminus or C-terminus of the fusion protein via translation. Optionally, the DogTag is internally fused into the fusion protein via translation.

8. The fusion protein according to claim 1, wherein the glycosylated fragment is a ComP glycosylated fragment; Optionally, the ComP glycosylated fragment comprises or consists of the following: The amino acid sequence CTGVTQIASGASAATTNVASAQC (SEQ ID NO:412) or a fragment containing the amino acid ASA, or variants thereof, wherein the variant contains amino acid ASA at positions 11-13 of SEQ ID NO:412 and has one, two, three, four, five or six amino acid substitutions, additions and / or deletions; Optionally, the ComP glycosylated fragment comprises or consists of the following amino acid sequence: Or a variant thereof, the variant comprising the amino acid ASA corresponding to positions 11-13 of SEQ ID NO:412, and having one, two, three, four, five or six amino acid substitutions, additions and / or deletions.

9. The fusion protein according to claim 1, wherein the glycosylated fragment is a TfpM-associated fimbriae protein glycosylation fragment; Optionally, the TfpM-associated fimbriae protein glycosylation fragment comprises or consists of the following: PilMo fimbriae disulfide ring region (SEQ ID NO:413) or a fragment thereof, said fragment containing at least the last three amino acids from the C-terminus of TfpM-associated fimbriae; Or a variant thereof, the variant comprising the last three amino acids from the C-terminus of the TfpM-associated fimbriae, and having one, two, three, four, five, or six amino acid substitutions, additions, and / or deletions.

10. The fusion protein according to claim 1, wherein the glycosylated fragment is a PilE glycosylated fragment; Optionally, the PilE glycosylated fragment comprises or consists of the following: Amino acid SAVTEYYLNHGEWPGNNTSAGVATSSEIK (SEQ ID NO:414) or a fragment thereof, said fragment containing at least amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414; or variants thereof, wherein the variant contains at least the amino acid WPGNNTSAGV (SEQ ID NO:439) at positions 13 to 22 of SEQ ID NO:414, and has one, two, three, four, five or six amino acid substitutions, additions and / or deletions.

11. The fusion protein according to claim 1, wherein the glycosylated fragment is a PglB glycosylated fragment; Optionally, the PglB glycosylated fragment comprises or consists of the following: The motif consists of the amino acid sequence X1X2NX3X4, where X1 is D or E, X2 is any amino acid except proline, X3 is any amino acid except proline, and X4 is S or T.

12. The fusion protein according to claim 1, wherein the glycosylated fragment is a PilA glycosylated fragment; Optionally, the PilA glycosylated fragment comprises or consists of the following: The PilA fimbriae disulfide ring region (SEQ ID NO:415) or a fragment thereof, said fragment containing at least the last three amino acids from the C-terminus of PilA; Or a variant thereof, the variant comprising at least the last three amino acids from the end of the PilA, and having one, two, three, four, five or six amino acid substitutions, additions and / or deletions.

13. The fusion protein of claim 1, wherein the glycosylated fragment is an STT3 glycosylated fragment; Optionally, the STT3 glycosylated fragment comprises or consists of the following: There is a motif amino acid sequence NX1 X2, where X1 is any amino acid except proline, and X2 is S or T.

14. The fusion protein according to claim 1, wherein the glycosylated fragment is an N-linked glycosyltransferase glycosylation fragment; Optionally, the N-linked glycosyltransferase glycosylation fragment comprises or consists of the following: There is a motif amino acid sequence N X1 X2, where X1 is any amino acid and X2 is S or T.

15. The fusion protein according to claim 1, wherein the glycosylated fragment is an O-linked glycosyltransferase glycosylation fragment; Optionally, the O-linked glycosyltransferase glycosylation fragment comprises or consists of the following: Fragments of serine or threonine repeat sequences from serine-rich or serine-rich adhesins derived from streptococci or staphylococci. Optionally, the O-linked glycosyltransferase glycosylation fragment comprises or consists of the following: The adhesin GspB from Streptococcus glabra is rich in serine or threonine repeat sequences.

16. The fusion protein of claim 1, wherein the fusion protein comprises two or more glycosylated fragments; Optionally, the fusion protein comprises at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 or 24 glycosylated fragments; Optionally, the fusion protein comprises any one to any one of 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 glycosylated fragments. Optionally, at least one glycosylated fragment is located at the N-terminus or C-terminus of the fusion protein and at least one glycosylated fragment is internally located within the fusion protein; Optionally, at least two glycosylated fragments are internally located within the fusion protein; And / or Optionally, one of the glycosylated fragments is located at the N-terminus of the fusion protein, and the glycosylated fragment is located at the C-terminus of the fusion protein.

17. The fusion protein according to claim 16, The two or more glycosylated fragments therein are identical; At least one of the two or more glycosylation fragments is different; or Each glycosylation fragment in the glycosylation fragment is different; Optionally, one of the glycosylation fragments is a ComP glycosylation fragment, and the other of the glycosylation fragments is a TfpM-associated fimbriae protein glycosylation fragment; Optionally, the fusion protein is a glycoconjugate comprising two or more sugars covalently linked to the fusion protein via the two or more glycosylation fragments; Optionally, the fusion protein comprises at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 or 24 covalently linked sugars; Optionally, the fusion protein comprises any one of 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23 covalently linked sugars to any one of 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 covalently linked sugars; and Choose any location The two or more sugars mentioned therein are identical; At least one of the two or more sugars is different; or Each of the sugars mentioned is different.

18. The fusion protein of claim 1, wherein the fusion protein comprises a carrier protein. Optionally, the carrier protein is selected from the group consisting of: Escherichia coli maltose-binding protein, Pseudomonas aeruginosa exotoxin A (EPA), Pseudomonas aeruginosa PcrV, CRM197, Haemophilus influenzae protein D, cholera toxin B subunit or tetanus toxin, and any fragment thereof.

19. A composition comprising a polypeptide pair, said polypeptide pair comprising a first polypeptide and a second polypeptide, The first polypeptide is a fusion protein according to any one of claims 1 to 18. The second polypeptide includes a second polypeptide tag binding partner relative to the first polypeptide tag of the first polypeptide, and The first polypeptide is linked to the second polypeptide via a heteropeptide bond between the first polypeptide tag and the second polypeptide tag; Optionally, the second polypeptide comprises a monomeric polypeptide capable of spontaneously polymerizing / self-assembling into a higher-order multimeric structure; and / or Further optionally, the higher-order polymeric structure is an icosahedral or dodecahedral particle (e.g., similar to a nanocage), a virus-like particle, or an adenovirus vector.

20. The polypeptide pair composition according to claim 19, The second polypeptide contains adenovirus capsid structural protein; The second polypeptide contains the coat protein of bacteriophage AP205; The second polypeptide contains a fragment of 2-keto-3-deoxy-phosphoglucuronide aldolase (i301); or The second polypeptide contains a fragment of mutated 2-keto-3-deoxy-phosphoglucuronide (mi3).

21. The polypeptide pair composition according to claim 19, wherein the second polypeptide tag is SpyCatcher (SEQ ID NO:420), SpyCatcher002 (SEQ ID NO:421), SpyCatcher003 (SEQ ID NO:422), or DogCatcher (SEQ ID NO:423); Choose any location Wherein the first polypeptide tag is SpyTag and the second polypeptide tag is SpyCatcher; The first polypeptide tag is SpyTag002 and the second polypeptide tag is SpyCatcher002; Wherein the first polypeptide tag is SpyTag003 and the second polypeptide tag is SpyCatcher003; or The first polypeptide tag is DogTag and the second polypeptide tag is DogCatcher.

22. The polypeptide pair composition according to claim 19, The second polypeptide tag is fused to the N-terminus or C-terminus of the second polypeptide via translation, or The second polypeptide tag is internally fused into the second polypeptide via translation.

23. The polypeptide pair composition of claim 19, wherein the first polypeptide is a bioconjugate comprising a sugar covalently linked to the glycosylated fragment of the first polypeptide; Optionally, the composition is immunogenic.

24. The polypeptide pair composition according to claim 19, further comprising an adjuvant and / or an excipient.

25. The polypeptide pair composition according to claim 19, wherein the composition is a pharmaceutical / therapeutic composition.

26. The polypeptide pair composition according to claim 19, wherein the composition is a conjugate vaccine.

27. A complex comprising two or more of the polypeptide pairs according to claim 19; Optionally, the complex comprises 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 75, 100, 125, 150, 175, 200, 225, 250 or more complex polypeptide pairs according to any one of claims 19 to 26; Optionally, the complex comprises any of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 75, 100, 125, 150, 175, 200, 225, 250, 300, or 400 pairs of complex polypeptides according to any one of claims 19 to 26. One to 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 75, 100, 125, 150, 175, 200, 225, 250, 300, 400 or 500 of the complex polypeptide pairs according to any one of claims 19 to 26; and / or Optionally, the complex is a self-assembled polymeric higher-order structure.

28. The complex of claim 27, wherein the self-assembled polymeric higher-order structure is an icosahedral or dodecahedral particle (e.g., similar to a nanocage), a virus-like particle, or an adenovirus vector.

29. The complex of claim 27, wherein all the first polypeptides of the complex comprise the same fusion protein.

30. The complex of claim 27, wherein at least two of the first polypeptides in the first polypeptide of the complex comprise different fusion proteins.

31. The complex of claim 27, wherein at least one first polypeptide of the complex is a bioconjugate, the bioconjugate comprising a sugar covalently linked to the glycosylated segment of the first polypeptide; Optionally, at least about 5%, 10%, 25%, 50%, 75%, 80%, 90%, 95%, 97%, 98%, or 99% of the first polypeptide in the complex is a bioconjugate; Optionally, any one of about 5%, 10%, 25%, 50%, 75%, 80%, 90%, 95%, 97%, or 98% of the first polypeptide in the complex to any one of about 10%, 25%, 50%, 75%, 80%, 90%, 95%, 97%, 98%, or 99% of the first polypeptide in the complex is a bioconjugate; Optionally, 100% of the first polypeptide in the said complex is a bioconjugate; Optionally, the complex is immunogenic.

32. The complex of claim 27, wherein two or more of the first polypeptides in the complex are bioconjugates comprising covalently linked sugars; Optionally, the complex comprises at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1,000, 1,500, 2,000, 2,500 or 5,000 covalently linked sugars; Optionally, the complex comprises 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1,000, 1,500, 2,000, or 2,500 covalently linked sugars. Any one of the following covalently linked sugars: 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 50, 75, 100, 125, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1,000, 1,500, 2,000, 2,500 or 5,000; Optionally, all sugars linked to the complex are identical; Optionally, at least one of the two or more sugars linked to the complex is different; Optionally, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32 or 33 different sugars linked to the complex are present; Optionally, each of the sugars linked to the complex is different.

33. A pharmaceutical / therapeutic composition comprising the complex according to claim 27, and an adjuvant and / or excipient.

34. The complex of claim 27 or the pharmaceutical / therapeutic composition of claim 33, wherein the complex or composition is a conjugate vaccine.

35. A method for preparing a polypeptide pair according to claim 19, the method comprising contacting the first polypeptide and the second polypeptide under conditions that allow the first polypeptide tag to spontaneously form an isopeptide bond with the second polypeptide tag binding coupler.

36. The method of claim 35, further comprising glycosylation of the first polypeptide with sugar prior to contacting the second polypeptide and forming an isopeptide bond; Optionally, the first polypeptide is glycosylated in vivo before contacting the second polypeptide and forming an isopeptide bond; and Optionally, the method includes isolating / purifying the in vivo glycosylated first polypeptide before contacting it with the second polypeptide and forming an isopeptide bond.

37. The method of claim 35, wherein the first polypeptide is glycosylated after contact with the second polypeptide and formation of an isopeptide bond.

38. A method for preparing the complex according to claim 27, the method comprising: (i) forming a self-assembled multimeric higher-order structure of the second polypeptide, and then contacting the first polypeptide and the second polypeptide under conditions that allow the first polypeptide tag to spontaneously form heteropeptide bonds with the second polypeptide tag; or (ii) The first polypeptide and the second polypeptide are brought into contact under conditions that allow the first polypeptide tag to spontaneously form heteropeptide bonds with the second polypeptide tag, and then a self-assembled multimeric higher-order structure of the second polypeptide is formed.

39. The method of claim 38, further comprising glycosylation of the first polypeptide with sugar, optionally: The first polypeptide is glycosylated before the isopeptide bond is formed between the first polypeptide and the second polypeptide; The first polypeptide is glycosylated after the isopeptide bond is formed between the first polypeptide and the second polypeptide. The first polypeptide is glycosylated before being incorporated into the higher-order structure of the polymer; And / or The process involves incorporating the first polypeptide into a higher-order polymer structure and then glycosylating the first polypeptide.

40. The method of claim 36, wherein the sugar is transferred to the fusion protein by the action of an N-linked oligosaccharide transferase (N-OTase), an O-linked oligosaccharide transferase (O-OTase), an N-linked glycosyltransferase (NGT), an O-linked glycosyltransferase (OGT), and / or a C-mannosyltransferase (CMT).

41. The method according to claim 36: The ComP glycosylation fragment was glycosylated by PglS OTase; The TfpM-associated fimbriae protein glycosylation fragment is glycosylated by TfpM OTase, optionally, the ComP glycosylation fragment is glycosylated by PglS OTase and the TfpM-associated fimbriae protein glycosylation fragment is glycosylated by TfpM OTase; The PilE glycosylation fragment was glycosylated by PglL OTase; The PglB glycosylation fragment is glycosylated by PglB OTase; The PilA glycosylated fragment is glycosylated by TfpO or PilO OTase; The STT3 glycosylation fragment was catalyzed by STT3 subunit glycosylation; Among them, the glycosylated fragment of the PilA_Pa5196-related fimbriae protein was glycosylated by TfpW glycosyltransferase; The N-linked glycosyltransferase glycosylation fragment is glycosylated by N-linked glycosyltransferases derived from *Actinomyces pleuropneumoniae*, *Haemophilus influenzae*, or *Yersinia enterocolitica*; and / or The O-linked glycosyltransferase glycosylation fragment was glycosylated by GtfA / GtfB glycosyltransferases.

42. The method according to claim 36: PglS OTase is used to covalently link the sugar to the oxygen atom within the glycosylated fragment; The sugar is covalently linked to an oxygen atom within the glycosylation fragment using TfpM OTase, optionally, wherein PglS OTase is used to covalently link the sugar to an oxygen atom within the glycosylation fragment and another sugar is covalently linked to an oxygen atom within the glycosylation fragment using TfpM OTase; PglL OTase is used to covalently link the sugar to the oxygen atom within the glycosylated fragment; PglB OTase is used to covalently link the sugar to the nitrogen atom within the glycosylated fragment; The sugar is covalently linked to the oxygen atom within the glycosylated fragment using TfpO or PilO OTase; The sugar is covalently linked to the nitrogen atom within the glycosylated fragment using STT3 OTase; AlgB OTase is used to covalently link the sugar to the nitrogen atom within the glycosylated fragment; TfpW glycosyltransferase is used to covalently link the sugar to an oxygen atom within the glycosylated fragment; In this process, an N-linked glycosyltransferase is used to covalently link the sugar to the nitrogen atom within the glycosylated fragment; O-linked glycosyltransferases are used to covalently link the sugar to an oxygen atom within the glycosylated fragment; C-mannosyltransferase is used to covalently link the sugar to a carbon atom within the glycosylated fragment.

43. The method according to claim 42: The sugar is covalently linked to an oxygen atom within a ComP glycosylation fragment (e.g., SEQ ID NO:400) using a PglS OTase (e.g., SEQ ID NO:412 or a variant thereof); The sugar is covalently linked to an oxygen atom within a TfpM glycosylation fragment (e.g., SEQ ID NO:402) using a TfpM OTase (e.g., SEQ ID NO:413 or a variant thereof); The sugar is covalently linked to an oxygen atom within the PilE glycosylation fragment (e.g., SEQ ID NO:404) using a PglL OTase (e.g., SEQ ID NO:414 or a variant thereof); The sugar is covalently linked to the nitrogen atom within a PglB glycosylation fragment (e.g., X1 X2 N X3 X4, where X1 is D or E, X2 is any amino acid other than proline, X3 is any amino acid other than proline, and X4 is S or T) using a PglB OTase (e.g., SEQ ID NO:405). The sugar is covalently linked to an oxygen atom within a PilA glycosylation fragment (e.g., SEQ ID NO:407) using a TfpO / PilO OTase (e.g., SEQ ID NO:415 or a variant thereof); The sugar is covalently linked to an STT3 glycosylation fragment (e.g., N X1 X2, where X1 is any amino acid other than proline, and X...) using an STT3 OTase (e.g., SEQ ID NO:408). 2 For nitrogen atoms in S or T); The sugar is covalently linked to the nitrogen atom within an archaeal AlgB glycosylation fragment (e.g., N X1 X2, where X1 is any amino acid other than proline and X2 is S or T) using AlgB OTase (e.g., SEQ ID NO:409). The sugar is covalently linked to an oxygen atom in a glycosylated fragment of PilA_Pa5196-associated fimbriae protein (e.g., SEQ ID NO:426 or a variant thereof) using a TfpW glycosyltransferase (e.g., SEQ ID NO:424); The sugar is covalently linked to the nitrogen atom in an N-linked glycosyltransferase sequence (e.g., N X1 X2, where X1 is any amino acid and X2 is S or T) using an N-linked glycosyltransferase (e.g., SEQ ID NO:410). The sugar is covalently linked to the oxygen atom in the O-linked glycosyltransferase sequence using an O-linked glycosyltransferase (e.g., SEQ ID NO:411); The sugar is covalently linked to a carbon atom within the C-mannosyltransferase glycosylation fragment using C-mannosyltransferase.

44. The method of claim 35, wherein the method is a method for producing a conjugate vaccine.

45. A system comprising the first polypeptide and the second polypeptide of the composition according to claim 1; Optionally, the first polypeptide is a glycosylated bioconjugate; Optionally, the system comprises a multimeric higher-order structure assembled from the second polypeptide; Optionally, the system comprises sugar and N-linked oligosaccharide transferase (N-OTase), O-linked oligosaccharide transferase (O-Otase), N-linked glycosyltransferase (NGT), O-linked glycosyltransferase (OGT) and / or C-mannosyltransferase (CMT).

46. ​​An isolated nucleic acid encoding the first polypeptide and / or the second polypeptide of the composition according to claim 26.

47. A vector comprising the isolated nucleic acid according to claim 46.

48. A host cell comprising the vector according to claim 47.

49. A kit comprising two or more components, said two or more components comprising: The fusion protein, first polypeptide, second polypeptide, sugar, N-linked oligosaccharide transferase (N-Otase), O-linked oligosaccharide transferase (O-Otase), N-linked glycosyltransferase (NGT), O-linked glycosyltransferase (OGT) and / or C-mannosyltransferase (CMT) according to any one of the preceding claims, bioconjugate, multimeric higher-order structure assembled from the second polypeptide, isolated nucleic acid, vector and host cell.

50. A method for inducing an immune response in a subject, said method being carried out by administering to the subject an effective amount of any composition, complex, and / or conjugate vaccine according to any of the preceding claims or a composition, complex, and / or conjugate vaccine for inducing an immune response in a subject according to any of the preceding claims.

51. The fusion protein according to claim 1, wherein the glycosylated fragment is a glycosylated fragment of PilA_Pa5196-associated fimbriae protein; Optionally, the PilA_Pa5196-associated fimbriae protein glycosylation fragment comprises or consists of the following: Chains 1 and 2 of the antiparallel β-fold structural domain of PilA_Pa5196 (SEQ ID NO:426), Or variants thereof having one, two, three, four, five or six amino acid substitutions, additions and / or deletions.

52. The fusion protein of claim 1, wherein the glycosylated fragment comprises means for linking the sugar to the glycosylated fragment by means of PglS OTase, TfpM OTase, PglL OTase, PglB OTase, TfpO / PilO OTase, STT3 OTase, AlgB OTase, N-linked glycosyltransferase and / or O-linked glycosyltransferase.

53. The fusion protein according to any one of the preceding claims, comprising an amino acid linker, Optionally, the amino acid linker is selected from the group consisting of: SGG, SEQ ID NO:430, SEQ ID NO:431, SEQ ID NO:432, SEQ ID NO:433, SEQ ID NO:434, SEQ ID NO:435, SEQ ID NO:436, SEQ ID NO:437 and SEQ ID NO:

438. Optionally, the amino acid linker is fused translationally immediately following the leader sequence, peptide tag, glycosylated fragment, carrier protein, and / or multiple histidine tags, and / or Optionally, the amino acid linker is fused translationally immediately preceding the leader sequence, peptide tag, glycosylated fragment, carrier protein, and / or multiple histidine tags.

Citation Information

Patent Citations

  • A prokaryote-based cell-free system for the synthesis of glycoproteins

    WO2013067523A1

  • Glycosylated COMP pilin variants, methods of making and uses thereof

    WO2019241672A2

  • O-linked glycosylation recognition motifs

    WO2020131236A1