Porphyromonas gingivalis antigenic construct

Antigens derived from Kgp, RgpA, and RgpB polypeptides delivered via mRNA provide a robust antibody response against Porphyromonas gingivalis, addressing the limitations of current therapies by inhibiting gingipain activity and adhesion, thereby preventing periodontitis.

JP2026524687APending Publication Date: 2026-07-23SANOFI SA(FR)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SANOFI SA(FR)
Filing Date
2024-07-19
Publication Date
2026-07-23

Smart Images

  • Figure 2026524687000086
    Figure 2026524687000086
  • Figure 2026524687000087
    Figure 2026524687000087
  • Figure 2026524687000088
    Figure 2026524687000088
Patent Text Reader

Abstract

The present invention relates to a composition (e.g., a vaccine composition) that can be used to immunize against P. gingivalis infection. The composition comprises a P. gingivalis antigen and a combination of antigens that can be used to immunize against P. gingivalis, either in the form of a nucleic acid encoding an antigenic protein (e.g., mRNA) or a recombinant protein antigen.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the field of treatment and prevention of Porphyromonas gingivalis (P. gingivalis) infections, such as periodontitis. In particular, the present invention relates to antigens and antigen combinations that can be used to immunize against P. gingivalis, used in the form of nucleic acids (e.g., mRNA) encoding antigenic proteins or recombinant protein antigens. [Background technology]

[0002] Periodontitis is a chronic inflammatory disease of the supporting tissues of teeth (Bostanci and Belibasakis, 2012). It affects all age groups, but has a higher incidence in older adults. Its main symptoms are bleeding or swelling of the gums, pain, and sometimes bad breath. It is characterized by the formation of periodontal pockets and subgingival plaque, which support colonization by pathogenic bacteria. In its severe forms, periodontitis can lead to destruction of the periodontal ligaments and alveolar bone, and ultimately tooth loss (Kinane et al., 2017). Periodontitis is estimated to affect about 50% of the world's population and is one of the most prevalent inflammatory diseases and a leading cause of tooth loss in adults (Mei et al., 2020). In 2022, the WHO estimated that about 19% of the world's adult population suffers from severe periodontal disease, representing more than 1 billion cases worldwide (WHO Global Oral Health Status Report, 2022).

[0003] Porphyromonas gingivalis (P. gingivalis) is a significant pathogen in periodontitis (or periodontal disease). It is a Gram-negative, non-motile, anaerobic pathogen that requires vitamin K and iron in the form of heme or hemin for its growth and produces energy by fermenting amino acids (Bostanci and Belibasakis, 2012). It is a secondary colonizer in the human oral cavity that forms communities and attaches to primary colonizers to colonize dental plaque. P. gingivalis is mainly found in deep periodontal pockets characteristic of periodontitis and has been detected in 85% of subgingival plaque samples from patients with chronic periodontitis (How et al., 2016). It is thought that P. gingivalis induces the progression of periodontitis by restructuring the normal flora of the oral cavity and promoting further colonization by pathogenic bacteria, leading to an imbalance (or dysbiosis) in the microbial biofilm state (Xu et al., 2020). Aside from its important role in periodontitis, P. gingivalis is also considered a potential risk factor for the development of several systemic diseases, including atherosclerosis, cancer, Alzheimer's disease, diabetes, and rheumatoid arthritis (Mei et al., 2020).

[0004] The main pathogenic factors of P. gingivalis include lipopolysaccharides, cilia, capsular proteins, gingipain, and outer membrane vesicles (Xu et al., 2020). Gingipain belongs to the family of cysteine ​​proteinase enzymes. They account for 85% of the extracellular proteolytic activity and 99% of the "trypsin-like activity" of P. gingivalis. They are typically located on the cell surface or on outer membrane vesicles of P. gingivalis strains, with the exception of strain HG66, which also secretes soluble gingipain into the extracellular environment (Li and Collyer, 2011).

[0005] Gingipain includes arginine-specific gingipain (RgpA and RgpB) and lysine-specific gingipain (Kgp), which cleave polypeptides at the C-terminus after arginine or lysine residues, respectively. These three proteins are encoded by individual loci found in the genomes of all P. gingivalis strains (Li and Collyer, 2011).

[0006] The primary function of gingipain is hypothesized to be the digestion of proteins for nutrition. For example, it has been proposed that Kgp cleaves host heme proteins to supply P. gingivalis with heme for its proliferation. However, gingipain has recently been found to also be involved in the pathogenicity of P. gingivalis. In particular, gingipain is thought to degrade collagen and fibrin / fibrinogen, thereby contributing to the destruction of gingival tissue, inhibiting blood coagulation, and increasing bleeding in periodontal tissue. Kgp and RgpA are also thought to mediate adhesion to host tissue and promote co-aggregation of P. gingivalis with other oral pathogens and subsequent biofilm formation. Furthermore, gingipain has been suggested to modulate the host immune response, suppressing the ability of innate and adaptive immune responses to eliminate bacteria while increasing inflammation (Aleksijevic et al., 2022).

[0007] Existing therapies for P. gingivalis-induced periodontitis include debridement from teeth by scaling (removal of plaque and calculus), and surgery in more severe cases. Adjunctive therapies include prescription of antibiotics and antimicrobial agents (Kinane et al., 2017), but these drugs are thought to have reduced efficacy against P. gingivalis due to their ability to form biofilms (Aleksijevic et al., 2022). Therefore, there is a need for an effective vaccine for the treatment and / or prevention of P. gingivalis-related diseases. It is hypothesized that targeting key pathogenic factors of P. gingivalis, such as gingipain, through reserve immunization may reduce the ability of the bacteria to cause periodontitis or migrate to distant tissues and incite other inflammatory diseases (Mei et al., 2020).

[0008] One vaccine candidate for P. gingivalis is based on a modified Kgp protein containing portions of the proteinase catalytic domain and adhesin domain of Kgp (known as Kas2-A1 - O'Brien-Simpson et al., 2011 and International Publication No. 2011014947A1). However, an improved P. gingivalis vaccine is still needed. [Overview of the Initiative] [Problems that the invention aims to solve]

[0009] The object of the present invention is to provide an antigen that can induce a functional antibody response to inhibit the proteinase activity of the gingipain catalytic domain and the hemagglutination and adhesion functions mediated by the gingipain adhesion domain. [Means for solving the problem]

[0010] The inventors have found that antigens derived from Kgp, RgpA, and / or RgpB can be used to immunize against P. gingivalis, as described herein. In particular, the inventors have found that antigens derived from Kgp, RgpA, or RgpB polypeptides in a P. gingivalis domain containing a specific portion of the Kgp, RgpA, or RgpB polypeptide induced a robust B cell (i.e., antibody) response when delivered by mRNA encoding the relevant antigen.

[0011] Accordingly, the present invention provides P. gingivalis polypeptides and nucleic acids comprising nucleotide sequences encoding such polypeptides. Polypeptide antigens described herein can be delivered by nucleic acids (e.g., mRNA) comprising the nucleotide sequences encoding the polypeptides, i.e., in that form.

[0012] The present invention also provides compositions comprising (i) a Kgp-based polypeptide or nucleic acid as described herein, and (ii) a combination of RgpA-based polypeptide or nucleic acid as described herein.

[0013] Ginger pine The term "gingipain," as used herein, refers to either the lysine-specific proteinase (Kgp) or one of the arginine-specific proteinases (RgpA and RgpB) of P. gingivalis. The term "gingipain" is used to refer to Kgp, RgpA, and RgpB. The terms "Kgp-based" and "RgpA-based," as used herein, refer to polypeptides and nucleic acids encoding polypeptides containing the Kgp or RgpA element, respectively. Polypeptides containing the Kgp and RgpA elements are referred to as "Kgp and RgpA-based."

[0014] The domain structure of gingipains is highly conserved among P. gingivalis strains. Kgp and RgpA have the same basic modular structure from the N - terminus to the C - terminus of the protein: a signal peptide, an N - terminal propeptide (cleaved in the mature protein), a protease catalytic domain (Cat), and a domain of unknown function (DUF), specifically DUF2436, followed by three truncated adhesion domains, specifically the C - terminal hemagglutinin / adhesin region composed of the K1, K2, and K3 adhesin domains. These domains are interspersed with sequences containing adhesion - binding motifs (ABMs) known as ABM1, ABM2, and ABM3. In native Kgp and RgpA, the first ABM1 and the first ABM2 are located on either side of DUF2436 (i.e., the first ABM1 is on the N - terminal side of DUF2436, between DUF2436 and the Cat domain, and the first ABM2 is on the C - terminal side of DUF2436). The second ABM1 is located on the C - terminal side of the first ABM2, followed by ABM3. The second ABM1 and ABM3 are located on the N - terminal side of the K1 adhesin domain. The K1 adhesin domain is followed by the K2 adhesin domain, followed by the second ABM2. Finally, the third ABM1 and the third ABM2 are located on either side of the K3 adhesin domain, and the protein then ends with the C - terminal domain (Li and Collyer, 2011).

[0015] The arrangement of the different domains of Kgp and RgpA is shown in Figure 1. As a specific example, Figures 2A and 2B show the wild - type sequences of Kgp and RgpA from the P. gingivalis W50 strain, in which each domain is highlighted and annotated according to the residue positions within the wild - type sequences.

[0016] The Cat, DUF2436, and K3 domains of Kgp and RgpA exhibit high sequence diversity, while the adhesin domains K1, K2, ABM1, ABM2, and ABM3 are highly conserved between RgpA and Kgp. For example, the Cat domains of Kgp and RgpA share approximately 27% sequence identity, while the DUF2436 domains of Kgp and RgpA share approximately 53% sequence identity. In contrast, each of the K1 and K2 adhesin domains of Kgp and RgpA share more than approximately 99% sequence identity.

[0017] RgpB contains a signal peptide, an N-terminal propeptide, a protease catalytic domain, and a short C-terminal domain. RgpB lacks the adhesin domain DUF2436, the adhesion-binding motif, and the K1-K3 domains. The Cat domain of RgpB shares approximately 90% sequence identity with the Cat domain of RgpA but shares only 20 - 30% sequence identity with the Cat domain of Kgp (Li and Collyer, 2011).

[0018] In a first aspect, the present invention provides a nucleic acid comprising a nucleotide sequence encoding a polypeptide, the polypeptide comprising i) at least a part of the catalytic domain of Porphyromonas gingivalis Lys-specific protease (Kgp); ii) at least a part of the domain of unknown function 2436 (DUF2436) of Porphyromonas gingivalis Kgp; iii) at least a part of the K1 adhesin domain of Porphyromonas gingivalis Kgp; iv) a first Porphyromonas gingivalis Kgp moiety comprising an adhesin-binding motif 1 (ABM1) and a first Porphyromonas gingivalis Kgp moiety comprising an adhesin-binding motif 2 (ABM2); v) the second Porphyromonas gingivalis Kgp moiety containing ABM1 and the second Porphyromonas gingivalis Kgp moiety containing ABM2; and vi) Porphyromonas gingivalis Kgp portion containing ABM3 Includes.

[0019] In a second embodiment, the present invention provides a nucleic acid comprising a nucleotide sequence encoding a polypeptide, wherein the polypeptide is i) At least a portion of the Arg-specific proteinase A (RgpA) catalytic domain or the Arg-specific proteinase B (RgpB) catalytic domain of Porphyromonas gingivalis; ii) At least a portion of Porphyromonas gingivalis RgpA DUF2436; iii) At least a portion of the RgpA K1 adhesin domain of Porphyromonas gingivalis; iv) A first Porphyromonas gingivalis RgpA moiety containing ABM1, and a first Porphyromonas gingivalis RgpA moiety containing ABM2; v) a second Porphyromonas gingivalis RgpA moiety containing ABM1, and a second Porphyromonas gingivalis RgpA moiety containing ABM2; and vi) Porphyromonas gingivalis RgpA portion containing ABM3 Includes.

[0020] In a third embodiment, the present invention is i) At least a portion of the Lys-specific proteinase (KGP) catalytic domain of Porphyromonas gingivalis; ii) At least a portion of domain 2436 (DUF2436) of the unknown function of Porphyromonas gingivalis KGP; iii) At least a portion of the KGP K1 adhesin domain of Porphyromonas gingivalis; iv) The first Porphyromonas gingivalis Kgp moiety containing the adhesin-binding motif 1 (ABM1) and the first Porphyromonas gingivalis Kgp moiety containing the adhesin-binding motif 2 (ABM2); v) the second Porphyromonas gingivalis Kgp moiety containing ABM1 and the second Porphyromonas gingivalis Kgp moiety containing ABM2; and vi) Porphyromonas gingivalis Kgp portion containing ABM3 Provides polypeptides containing the following:

[0021] In a fourth embodiment, the present invention is i) At least a portion of the Arg-specific proteinase A (RgpA) catalytic domain or the Arg-specific proteinase B (RgpB) catalytic domain of Porphyromonas gingivalis; ii) At least a portion of Porphyromonas gingivalis RgpA DUF2436; iii) At least a portion of the RgpA K1 adhesin domain of Porphyromonas gingivalis; iv) A first Porphyromonas gingivalis RgpA moiety containing ABM1, and a first Porphyromonas gingivalis RgpA moiety containing ABM2; v) a second Porphyromonas gingivalis RgpA moiety containing ABM1, and a second Porphyromonas gingivalis RgpA moiety containing ABM2; and vi) Porphyromonas gingivalis RgpA portion containing ABM3 Provides polypeptides containing the following:

[0022] In another aspect, the present invention provides a composition comprising any one of the nucleic acids of the present invention, preferably an immunogenic composition.

[0023] In another aspect, the present invention provides a composition comprising a first nucleic acid and a second nucleic acid, preferably an immunogenic composition.

[0024] In another embodiment, the present invention provides a composition comprising any one of the polypeptides of the present invention, preferably an immunogenic composition.

[0025] In another embodiment, the present invention provides a composition comprising a first polypeptide and a second polypeptide of the present invention, preferably an immunogenic composition.

[0026] In another embodiment, the present invention provides a vaccine comprising any one nucleic acid, any polypeptide, or any composition of the present invention.

[0027] The modular nature of zingipain means that the nucleic acids and polypeptides of the present invention may be composed of combinations of domains derived from different zingipains. Therefore, the nucleic acids and polypeptides of the present invention may, for example, include any of the catalytic domains described herein together with any of the DUF2436 domains described herein. As another example, any portion of zingipain containing ABM1 may be combined with any of the DUF2436 domains described herein, as described herein.

[0028] P. gingivalis strain The nucleic acids and polypeptides of the present invention may be derived from any strain of P. gingivalis. The modular structure of gingipain means that one domain of the nucleic acid or polypeptide may be derived from one strain of P. gingivalis, and another domain may be derived from a different strain of P. gingivalis. In another embodiment, all domains of the nucleic acid or polypeptide are derived from the same strain of P. gingivalis.

[0029] Examples of P. gingivalis strains are shown in Table 1 along with their corresponding GenBank sequences. Those skilled in the art can identify different domains of Kgp, RgpA, or RgpB in P. gingivalis strains, or portions of Kgp, RgpA, or RgpB containing, for example, ABM, by comparing them, for example, with the sequences of the corresponding domains or portions of Kgp or RgpA in P. gingivalis W50 strain. An example of the wild-type Kgp sequence of P. gingivalis W50 strain is provided in Sequence ID No. 157, and the corresponding nucleic acid encoding this sequence is in Sequence ID No. 160. An example of the wild-type RgpA sequence of P. gingivalis W50 strain is provided in Sequence ID No. 158, and the corresponding nucleic acid encoding this sequence is in Sequence ID No. 161. An example of the RgpB sequence of the wild-type P. gingivalis W50 strain is provided in SEQ ID NO: 159, and the corresponding nucleic acid encoding this sequence is in SEQ ID NO: 162.

[0030] [Table 1]

[0031] The nucleic acids and polypeptides of the present invention may be derived from P. gingivalis and may contain several domains that correspond to domains in naturally occurring Kgp, RgpA, or RgpB sequences. However, the nucleic acids do not encode polypeptides that are naturally occurring full-length Kgp, RgpA, or RgpB polypeptides (or mature polypeptides). Similarly, the polypeptides of the present invention are not naturally occurring full-length Kgp, RgpA, or RgpB polypeptides (or mature polypeptides).

[0032] In other words, the nucleic acids of the present invention may encode polypeptides that are modified from naturally occurring full-length Kgp, RgpA, or RgpB polypeptides. Similarly, the polypeptides of the present invention are modified from naturally occurring full-length Kgp, RgpA, or RgpB polypeptides. The modified polypeptide may be a variant of a naturally occurring polypeptide having an altered amino acid sequence, for example, by amino acid substitution, deletion, or insertion. The modified polypeptide may be a truncated or fragmented form of a naturally occurring polypeptide.

[0033] In any of the embodiments disclosed herein, the polypeptide may be a modified polypeptide according to the present invention, as described elsewhere in this specification. In any of the embodiments disclosed herein, the nucleic acid may encode a modified polypeptide according to the present invention, as described elsewhere in this specification.

[0034] The polypeptide of the present invention may also be in the form of a recombinant polypeptide. Therefore, in any of the embodiments described herein, the polypeptide is a recombinant polypeptide.

[0035] Adhesin-binding motif (ABM) Adhesin-binding motifs (ABMs) are sequences found in natural gingipain that are hypothesized to contribute to the adhesion function of gingipain. Three different gingipain ABMs have been described: ABM1, ABM2, and ABM3 (Li and Collyer, 2011). ABM1 was initially described based on the identification of a conserved sequence (Slakeski et al., 1998). ABM2 and ABM3 were described according to sequences conjugated by a specific antibody (O'Brien-Simpson et al., 2005).

[0036] The nucleic acids and polypeptides of the present invention comprise a first portion of Kgp or RgpA containing ABM1, a first portion of Kgp or RgpA containing ABM2, a second portion of Kgp or RgpA containing ABM1, and a second portion of Kgp or RgpA containing ABM2.

[0037] Gingipain portion containing ABM1 In some embodiments, the first Kgp moiety containing ABM1 is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, the first Kgp moiety containing ABM1 has the sequence of SEQ ID NO: 106. Therefore, in some embodiments, the first Kgp moiety containing ABM1 contains the sequence of SEQ ID NO: 106, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0038] In some embodiments, the first RgpA moiety containing ABM1 is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, the first RgpA moiety containing ABM1 has the sequence of SEQ ID NO: 114. Therefore, in some embodiments, the first RgpA moiety containing ABM1 contains the sequence of SEQ ID NO: 114, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0039] In some embodiments, the second Kgp moiety containing ABM1 is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, the second Kgp moiety containing ABM1 has the sequence of SEQ ID NO: 110. Therefore, in some embodiments, the second Kgp moiety containing ABM1 contains the sequence of SEQ ID NO: 110, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0040] In some embodiments, the second RgpA moiety containing ABM1 is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, the second RgpA moiety containing ABM1 has the sequence of SEQ ID NO: 102. Therefore, in some embodiments, the second RgpA moiety containing ABM1 contains the sequence of SEQ ID NO: 102, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0041] Examples of portions of gingipain (e.g., Kgp or RgpA) sequences containing the ABM1 sequence are provided in Table 2, along with the consensus sequence of ABM1. The consensus sequence of ABM1 is sequence number 120.

[0042] [Table 2]

[0043] In some embodiments, the first Kgp portion including ABM1 includes the sequence of sequence number 120, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0044] In some embodiments, the first Kgp portion including ABM1 includes the sequence of sequence number 106, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0045] In some embodiments, the first Kgp portion including ABM1 includes the sequence of sequence number 107, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0046] In some embodiments, the first Kgp portion including ABM1 includes the sequence of sequence number 89, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0047] In some embodiments, the first Kgp portion including ABM1 includes the sequence of sequence number 108, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0048] In some embodiments, the first Kgp portion including ABM1 includes the sequence of sequence number 109, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0049] In some embodiments, the first Kgp portion including ABM1 includes a sequence having at least 70% identity (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 97, 98, or 99%) to a Kgp-derived sequence bounded by a Cat domain at the N-terminus and DUF2436 at the C-terminus.

[0050] In some embodiments, the second Kgp portion including ABM1 includes the sequence of sequence number 120, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0051] In some embodiments, the second Kgp portion including ABM1 includes the sequence of sequence number 110, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0052] In some embodiments, the second Kgp portion including ABM1 includes the sequence of sequence number 111, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0053] In some embodiments, the second Kgp portion including ABM1 includes the sequence of sequence number 92, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0054] In some embodiments, the second Kgp portion including ABM1 includes the sequence of sequence number 112, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0055] In some embodiments, the second Kgp portion including ABM1 includes the sequence of sequence number 113, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0056] In some embodiments, the second Kgp portion including ABM1 includes a sequence having at least 70% identity (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) with a Kgp-derived sequence bound at the N-terminus by the ABM2 sequence and at the C-terminus by the ABM3 sequence.

[0057] In some embodiments, the first RgpA portion including ABM1 includes the sequence of sequence number 120, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0058] In some embodiments, the first RgpA portion including ABM1 includes the sequence of sequence number 114, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0059] In some embodiments, the first RgpA portion including ABM1 includes the sequence of sequence number 99, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0060] In some embodiments, the first RgpA portion including ABM1 includes the sequence of sequence number 115, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0061] In some embodiments, the first RgpA portion including ABM1 includes the sequence of sequence number 116, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0062] In some embodiments, the first RgpA portion including ABM1 includes a sequence having at least 70% identity (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) to a sequence derived from RgpA that is bounded at the N-terminus by a Cat domain and at the C-terminus by DUF2436.

[0063] In some embodiments, the second RgpA portion, including ABM1, includes the sequence of sequence number 120, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0064] In some embodiments, the second RgpA portion including ABM1 includes the sequence of sequence number 102, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0065] In some embodiments, the second RgpA portion, including ABM1, includes the sequence of sequence number 117, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0066] In some embodiments, the second RgpA portion, including ABM1, includes the sequence of sequence number 118, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0067] In some embodiments, the second RgpA portion, including ABM1, includes the sequence of sequence number 119, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0068] In some embodiments, the second RgpA portion including ABM1 includes a sequence having at least 70% identity (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 89, 90, or 99%) with a sequence derived from RgpA that is bounded at the N-terminus by the ABM2 sequence and at the C-terminus by the ABM3 sequence.

[0069] In the wild-type sequence of Kgp derived from P. gingivalis W50 strain, the first Kgp moiety containing ABM1 is located between the Cat domain and DUF2436, has a length of 35 amino acids, and contains the sequence of SEQ ID NO: 106. This sequence contains ABM1 of SEQ ID NO: 120. Therefore, in some embodiments, the first Kgp moiety containing ABM1 has a length of 10 to 35 amino acids and contains the sequence of SEQ ID NO: 120.

[0070] In some embodiments, the first Kgp moiety containing ABM1 is 15 to 30 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the first Kgp moiety containing ABM1 is 20 to 25 amino acids long and includes the sequence of SEQ ID NO: 120.

[0071] In some embodiments, the first Kgp portion containing ABM1 is 10 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the first Kgp portion containing ABM1 is 15 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the first Kgp portion containing ABM1 is 20 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the first Kgp portion containing ABM1 is 25 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the first Kgp portion containing ABM1 is 30 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the first Kgp portion containing ABM1 is 35 amino acids long and includes the sequence of SEQ ID NO: 120.

[0072] In some embodiments, the first Kgp portion containing ABM1 has an amino acid length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35, and includes the sequence of SEQ ID NO: 120.

[0073] In the wild-type Kgp sequence derived from P. gingivalis W50 strain, the second Kgp moiety containing ABM1 is located between the DUF and the K1 adhesin domain. The ABM2 sequence is located on the N-terminal side of the second moiety containing ABM1, and the ABM3 sequence is located on the C-terminal side of the second moiety containing ABM1. In this context, the second Kgp moiety containing ABM1 is 26 amino acids long and contains the sequence of SEQ ID NO: 110. This sequence contains ABM1 of SEQ ID NO: 120. Therefore, in some embodiments, the second Kgp moiety containing ABM1 is 10 to 26 amino acids long and contains the sequence of SEQ ID NO: 120.

[0074] In some embodiments, the second Kgp moiety containing ABM1 is 10 to 25 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the second Kgp moiety containing ABM1 is 15 to 20 amino acids long and includes the sequence of SEQ ID NO: 120.

[0075] In some embodiments, the second Kgp moiety containing ABM1 is 10 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the second Kgp moiety containing ABM1 is 15 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the second Kgp moiety containing ABM1 is 20 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the second Kgp moiety containing ABM1 is 25 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the second Kgp moiety containing ABM1 is 26 amino acids long and includes the sequence of SEQ ID NO: 120.

[0076] In some embodiments, the second Kgp portion containing ABM1 has an amino acid length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 26, and includes the sequence of SEQ ID NO: 120.

[0077] In the wild-type RgpA sequence derived from P. gingivalis W50 strain, the first RgpA moiety containing ABM1 is positioned between the Cat domain and the DUF domain, is 32 amino acids long, and contains the sequence of SEQ ID NO: 114. This sequence contains the ABM1 of SEQ ID NO: 120. Therefore, in some embodiments, the first RgpA moiety containing ABM1 is 10 to 32 amino acids long and contains the sequence of SEQ ID NO: 120.

[0078] In some embodiments, the first RgpA moiety containing ABM1 is 15 to 30 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the first RgpA moiety containing ABM1 is 20 to 25 amino acids long and includes the sequence of SEQ ID NO: 120.

[0079] In some embodiments, the first RgpA moiety containing ABM1 is 10 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the first RgpA moiety containing ABM1 is 15 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the first RgpA moiety containing ABM1 is 20 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the first RgpA moiety containing ABM1 is 25 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the first RgpA moiety containing ABM1 is 30 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the first RgpA moiety containing ABM1 is 32 amino acids long and includes the sequence of SEQ ID NO: 120.

[0080] In some embodiments, the first RgpA portion containing ABM1 is 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, or 32 amino acids long and contains the sequence of SEQ ID NO: 120.

[0081] In the wild-type RgpA sequence derived from P. gingivalis W50 strain, the second RgpA moiety containing ABM1 is located between the DUF and the K1 adhesin domain. The ABM2 sequence is located on the N-terminal side of the second moiety containing ABM1, and the ABM3 sequence is located on the C-terminal side of the second moiety containing ABM1. In this context, the second RgpA moiety containing ABM1 is 27 amino acids long and contains the sequence of SEQ ID NO: 119. This sequence contains ABM1 of SEQ ID NO: 120. Therefore, in some embodiments, the second RgpA moiety containing ABM1 is 10 to 27 amino acids long and contains the sequence of SEQ ID NO: 120.

[0082] In some embodiments, the second RgpA moiety containing ABM1 is 10 to 25 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the second RgpA moiety containing ABM1 is 15 to 20 amino acids long and includes the sequence of SEQ ID NO: 120.

[0083] In some embodiments, the second RgpA moiety containing ABM1 is 10 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the second RgpA moiety containing ABM1 is 15 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the second RgpA moiety containing ABM1 is 20 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the second RgpA moiety containing ABM1 is 25 amino acids long and includes the sequence of SEQ ID NO: 120. In some embodiments, the second RgpA moiety containing ABM1 is 27 amino acids long and includes the sequence of SEQ ID NO: 120.

[0084] In some embodiments, the second RgpA portion containing ABM1 has an amino acid length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27, and includes the sequence of SEQ ID NO: 120.

[0085] Gingipain portion containing ABM2 In some embodiments, the first Kgp moiety containing ABM2 is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, the first Kgp moiety containing ABM2 has the sequence of SEQ ID NO: 121. Therefore, in some embodiments, the first Kgp moiety containing ABM2 contains the sequence of SEQ ID NO: 121, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0086] In some embodiments, the first RgpA moiety containing ABM2 is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, the first RgpA moiety containing ABM2 has the sequence of SEQ ID NO: 101. Therefore, in some embodiments, the first RgpA moiety containing ABM2 contains the sequence of SEQ ID NO: 101, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0087] In some embodiments, the second Kgp moiety containing ABM2 is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, the second Kgp moiety containing ABM2 has the sequence of SEQ ID NO: 124. Therefore, in some embodiments, the second Kgp moiety containing ABM2 contains the sequence of SEQ ID NO: 124, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0088] In some embodiments, the second RgpA moiety containing ABM2 is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, the second RgpA moiety containing ABM2 has the sequence of SEQ ID NO: 126. Therefore, in some embodiments, the second RgpA moiety containing ABM21 contains the sequence of SEQ ID NO: 126, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0089] Examples of portions of gingipain (e.g., Kgp or RgpA) sequences containing the ABM2 sequence are provided in Table 3, along with the consensus sequence of ABM2. The consensus sequence of ABM2 is sequence number 130.

[0090] [Table 3]

[0091] In some embodiments, the first Kgp portion including ABM2 includes the sequence of sequence number 130, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0092] In some embodiments, the first Kgp portion including ABM2 includes the sequence of sequence number 121, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0093] In some embodiments, the first Kgp portion including ABM2 includes the sequence of sequence number 122, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0094] In some embodiments, the first Kgp portion including ABM2 includes the sequence of sequence number 91, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0095] In some embodiments, the first Kgp portion including ABM2 includes the sequence of sequence number 123, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0096] In some embodiments, the first Kgp portion including ABM2 includes a sequence having at least 70% identity (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 90, or 99%) with a Kgp-derived sequence bounded at the N-terminus by DUF2436 and at the C-terminus by the ABM1 sequence.

[0097] In some embodiments, the second Kgp portion including ABM2 includes the sequence of sequence number 130, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0098] In some embodiments, the second Kgp portion including ABM2 includes the sequence of sequence number 124, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0099] In some embodiments, the second Kgp portion including ABM2 includes the sequence of sequence number 125, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0100] In some embodiments, the second Kgp portion including ABM2 includes the sequence of sequence number 95, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0101] In some embodiments, the second Kgp portion including ABM2 has at least 70% identity (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 89, 90, or 99%) with a Kgp-derived sequence that is bounded at the N-terminus by a Kgp-derived sequence that is bound at the N-terminus by a Kgp-derived sequence that is bound at the C-terminus by a Kgp-derived sequence that is bound at the N-terminus by a Kgp-derived sequence that is bound at the C-terminus by ABM1) with respect to a Kgp-derived sequence.

[0102] In some embodiments, the first RgpA portion, including ABM2, includes the sequence of sequence number 130, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0103] In some embodiments, the first RgpA portion including ABM2 includes the sequence of sequence number 101, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0104] In some embodiments, the first RgpA portion including ABM2 includes a sequence having at least 70% identity (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) with a sequence derived from RgpA that is bounded at the N-terminus by DUF2436 and at the C-terminus by the ABM1 sequence.

[0105] In some embodiments, the second RgpA portion, including ABM2, includes the sequence of sequence number 130, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0106] In some embodiments, the second RgpA portion, including ABM2, includes the sequence of sequence number 105, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0107] In some embodiments, the second RgpA portion, including ABM2, includes the sequence of sequence number 126, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0108] In some embodiments, the second RgpA portion, including ABM2, includes the sequence of sequence number 127, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0109] In some embodiments, the second RgpA portion, including ABM2, includes the sequence of sequence number 128, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0110] In some embodiments, the second RgpA portion, including ABM2, includes the sequence of sequence number 129, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0111] In some embodiments, the second RgpA portion, including ABM2, contains a sequence having at least 70% identity (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 89, 90, or 99%) with a sequence derived from RgpA that is bounded at the N-terminus by an RgpA K2 adhesin domain and at the C-terminus by an ABM1 sequence, or a sequence derived from RgpA that is bound at the N-terminus by an RgpA K2 adhesin domain and at the C-terminus by ABM1.

[0112] In the wild-type Kgp sequence derived from P. gingivalis W50 strain, the first Kgp moiety containing ABM2 is located between DUF2436 and the K1 adhesin domain. DUF2436 is located at the N-terminus of the first moiety containing ABM2. The ABM1 and ABM3 sequences are located at the C-terminus of the first moiety containing ABM2. In this context, the first Kgp moiety containing ABM2 is 65 amino acids long and contains the sequence of SEQ ID NO: 91. This sequence contains the ABM2 of SEQ ID NO: 130. Therefore, in some embodiments, the first Kgp moiety containing ABM2 is 14 to 65 amino acids long and contains the sequence of SEQ ID NO: 130.

[0113] In some embodiments, the first Kgp moiety containing ABM2 is 20 to 60 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first Kgp moiety containing ABM2 is 25 to 55 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first Kgp moiety containing ABM2 is 30 to 50 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first Kgp moiety containing ABM2 is 35 to 45 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first Kgp moiety containing ABM2 is 35 to 40 amino acids long and includes the sequence of SEQ ID NO: 130.

[0114] In some embodiments, the first Kgp portion containing ABM2 is 14 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first Kgp portion containing ABM2 is 20 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first Kgp portion containing ABM2 is 25 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first Kgp portion containing ABM2 is 30 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first Kgp portion containing ABM2 is 35 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first Kgp portion containing ABM2 is 40 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first Kgp portion containing ABM2 is 45 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first Kgp portion containing ABM2 is 50 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first Kgp portion containing ABM2 is 55 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first Kgp portion containing ABM2 is 60 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first Kgp portion containing ABM2 is 65 amino acids long and includes the sequence of SEQ ID NO: 130.

[0115] In some embodiments, the first Kgp portion containing ABM2 is 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, or 65 amino acids long and contains the sequence of SEQ ID NO: 130.

[0116] In the wild-type sequence of Kgp derived from P. gingivalis W50 strain, the second Kgp moiety containing ABM2 is positioned between the K2 and K3 adhesin domains. The second Kgp moiety containing ABM2 is 56 amino acids long and contains the sequence of SEQ ID NO: 124. This sequence contains the ABM2 of SEQ ID NO: 130. Therefore, in some embodiments, the second Kgp moiety containing ABM2 is 14 to 56 amino acids long and contains the sequence of SEQ ID NO: 130.

[0117] In some embodiments, the second Kgp moiety containing ABM2 is 20 to 50 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second Kgp moiety containing ABM2 is 25 to 45 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second Kgp moiety containing ABM2 is 30 to 40 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second Kgp moiety containing ABM2 is 35 to 40 amino acids long and includes the sequence of SEQ ID NO: 130.

[0118] In some embodiments, the second Kgp portion containing ABM2 is 14 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second Kgp portion containing ABM2 is 20 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second Kgp portion containing ABM2 is 25 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second Kgp portion containing ABM2 is 30 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second Kgp portion containing ABM2 is 35 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second Kgp portion containing ABM2 is 40 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second Kgp portion containing ABM2 is 45 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second Kgp portion containing ABM2 is 50 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second Kgp portion containing ABM2 is 56 amino acids long and includes the sequence of SEQ ID NO: 130.

[0119] In some embodiments, the second Kgp portion containing ABM2 is 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, or 56 amino acids long and contains the sequence of SEQ ID NO: 130.

[0120] In the wild-type RgpA sequence derived from P. gingivalis W50 strain, the first RgpA moiety containing ABM2 is located between DUF2436 and the K1 adhesin domain. DUF2436 is located at the N-terminus of the first moiety containing ABM2. The ABM1 and ABM3 sequences are located at the C-terminus of the first moiety containing ABM2. In this context, the first RgpA moiety containing ABM2 is 65 amino acids long and contains the sequence of SEQ ID NO: 101. This sequence contains ABM2 of SEQ ID NO: 130. Therefore, in some embodiments, the first RgpA moiety containing ABM2 is 14 to 65 amino acids long and contains the sequence of SEQ ID NO: 130.

[0121] In some embodiments, the first RgpA moiety containing ABM2 is 20 to 60 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first RgpA moiety containing ABM2 is 25 to 55 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first RgpA moiety containing ABM2 is 30 to 50 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first RgpA moiety containing ABM2 is 35 to 45 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first RgpA moiety containing ABM2 is 35 to 40 amino acids long and includes the sequence of SEQ ID NO: 130.

[0122] In some embodiments, the first RgpA moiety containing ABM2 is 14 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first RgpA moiety containing ABM2 is 20 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first RgpA moiety containing ABM2 is 25 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first RgpA moiety containing ABM2 is 30 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first RgpA moiety containing ABM2 is 35 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first RgpA moiety containing ABM2 is 40 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first RgpA moiety containing ABM2 is 45 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first RgpA moiety containing ABM2 is 50 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first RgpA moiety containing ABM2 is 55 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first RgpA moiety containing ABM2 is 60 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the first RgpA moiety containing ABM2 is 65 amino acids long and includes the sequence of SEQ ID NO: 130.

[0123] In some embodiments, the first RgpA portion containing ABM2 is 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, or 65 amino acids long and contains the sequence of SEQ ID NO: 130.

[0124] In the wild-type RgpA sequence derived from P. gingivalis W50 strain, the second RgpA moiety containing ABM2 is positioned between the K2 and K3 adhesin domains. The second RgpA moiety containing ABM2 is 56 amino acids long and contains the sequence of SEQ ID NO: 126. This sequence contains the ABM2 of SEQ ID NO: 130. Therefore, in some embodiments, the second RgpA moiety containing ABM2 is 14 to 56 amino acids long and contains the sequence of SEQ ID NO: 130.

[0125] In some embodiments, the second RgpA moiety containing ABM2 is 20 to 50 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second RgpA moiety containing ABM2 is 25 to 45 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second RgpA moiety containing ABM2 is 30 to 40 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second RgpA moiety containing ABM2 is 35 to 40 amino acids long and includes the sequence of SEQ ID NO: 130.

[0126] In some embodiments, the second RgpA moiety containing ABM2 is 14 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second RgpA moiety containing ABM2 is 20 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second RgpA moiety containing ABM2 is 25 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second RgpA moiety containing ABM2 is 30 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second RgpA moiety containing ABM2 is 35 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second RgpA moiety containing ABM2 is 40 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second RgpA moiety containing ABM2 is 45 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second RgpA moiety containing ABM2 is 50 amino acids long and includes the sequence of SEQ ID NO: 130. In some embodiments, the second RgpA moiety containing ABM2 is 56 amino acids long and includes the sequence of SEQ ID NO: 130.

[0127] In some embodiments, the second RgpA portion containing ABM2 is 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, or 56 amino acids long and contains the sequence of SEQ ID NO: 130.

[0128] Pair of sequences containing ABM1 and ABM2 The first Kgp moiety containing ABM1 is positioned at the N-terminus of DUF2436, and the first Kgp moiety containing ABM2 is positioned at the C-terminus of DUF2436. A three-dimensional structure modeled using Alphafold2 of amino acids 229-1732 of native Kgp from P. gingivalis strain W50 (shown in Figure 3) suggests that the first moiety of Kgp containing ABM1 and the first moiety of Kgp containing ABM2 can associate to form a first fibronectin type III-like domain. The fibronectin type III-like domain is a beta-sandwich structure containing a first beta-sheet of three strands and a second beta-sheet of four strands. The first Kgp moiety containing ABM1 forms the N-terminus of the first fibronectin type III-like domain and contributes two strands to the beta-sheet containing three strands. The first Kgp moiety containing ABM2 forms the C-terminal portion of the first fibronectin type III-like domain and contributes one strand of a three-stranded beta sheet and four strands of a second beta sheet. This fibronectin type III-like domain contains ABM1 and ABM2 motifs and may play a role in mediating the adhesion function of gingipain. In fact, gingipain is known to interact with fibronectin (Li and Collyer, 2011), and the fibronectin type III domain is known to mediate the interaction with fibronectin. Furthermore, fibronectin type III-like domains have been found in other bacterial species and are presumed to perform similar functions (Konkel et al., 2010). Therefore, it may be advantageous for the polypeptide to present the ABM1 and ABM2 motifs in a three-dimensional structure similar to that of the wild-type protein.

[0129] Therefore, in some embodiments, the first Kgp moiety containing ABM1 can form the N-terminal portion of the first fibronectin type III-like domain, for example, a betasheet containing two chains. In some embodiments, the first Kgp moiety containing ABM2 can form the C-terminal portion of the first fibronectin type III-like domain, for example, a betasheet having four chains and further betasheet chains. In some embodiments, the first Kgp moiety containing ABM1 can form the N-terminal portion of the first fibronectin type III-like domain, for example, a betasheet containing two chains, and the first Kgp moiety containing ABM2 can form the C-terminal portion of the first fibronectin type III-like domain, for example, a betasheet having four chains and further betasheet chains. In certain such embodiments, the first Kgp moiety containing ABM1 and the first Kgp moiety containing ABM2 can form a first fibronectin type III-like domain having a beta-sandwich structure.

[0130] In some embodiments, the first RgpA moiety containing ABM1 can form the N-terminal portion of the first fibronectin type III-like domain, for example, a betasheet containing two strands. In some embodiments, the first RgpA moiety containing ABM2 can form the C-terminal portion of the first fibronectin type III-like domain, for example, a betasheet having four strands and further betasheet chains. In some embodiments, the first RgpA moiety containing ABM1 can form the N-terminal portion of the first fibronectin type III-like domain, for example, a betasheet containing two strands, and the first RgpA moiety containing ABM2 can form the C-terminal portion of the first fibronectin type III-like domain, for example, a betasheet having four strands and further betasheet chains. In certain such embodiments, the first RgpA moiety containing ABM1 and the first RgpA moiety containing ABM2 can form a first fibronectin type III-like domain having a beta-sandwich structure.

[0131] In some embodiments, the first Kgp moiety containing ABM1 and the first Kgp moiety containing ABM2 each contain the sequence shown in Table 4, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). In certain such embodiments, the first Kgp moiety containing ABM1 and the first Kgp moiety containing ABM2 can form a first fibronectin type III-like domain, as described in the preceding paragraph.

[0132] [Table 4]

[0133] In some embodiments, the first RgpA portion containing ABM1 and the first RgpA portion containing ABM2 each contain the sequence shown in Table 5, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). In certain such embodiments, the first RgpA portion containing ABM1 and the first RgpA portion containing ABM2 can form a first fibronectin type III-like domain, as described in the preceding paragraph.

[0134] [Table 5]

[0135] Figure 3 also clearly shows that the second Kgp moiety containing ABM1 and the second Kgp moiety containing ABM2 associate to form a second fibronectin type III-like domain. Specifically, the second Kgp moiety containing ABM1 forms the N-terminal portion of the second fibronectin type III-like domain and contributes two strands to the three-stranded beta sheet. The second Kgp moiety containing ABM2 forms the C-terminal portion of the second fibronectin type III-like domain and contributes one strand to the three-stranded beta sheet and four strands to the second beta sheet. The N-terminal portion of the second fibronectin type III-like domain associates with the C-terminal portion of the second fibronectin type III-like domain to form a beta-sandwich structure.

[0136] Therefore, in some embodiments, the second Kgp moiety containing ABM1 can form the N-terminal portion of the second fibronectin type III-like domain, for example, a betasheet containing two chains. In some embodiments, the second Kgp moiety containing ABM2 can form the C-terminal portion of the second fibronectin type III-like domain, for example, a betasheet having four chains and further betasheet chains. In some embodiments, the second Kgp moiety containing ABM1 can form the N-terminal portion of the second fibronectin type III-like domain, for example, a betasheet containing two chains, and the second Kgp moiety containing ABM2 can form the C-terminal portion of the second fibronectin type III-like domain, for example, a betasheet having four chains and further betasheet chains. In certain such embodiments, the second Kgp moiety containing ABM1 and the second Kgp moiety containing ABM2 can form a second fibronectin type III-like domain having a beta-sandwich structure.

[0137] In some embodiments, the second RgpA moiety containing ABM1 can form the N-terminal portion of the second fibronectin type III-like domain, for example, a betasheet containing two strands. In some embodiments, the second RgpA moiety containing ABM2 can form the C-terminal portion of the second fibronectin type III-like domain, for example, a betasheet having four strands and further betasheet chains. In some embodiments, the second RgpA moiety containing ABM1 can form the N-terminal portion of the second fibronectin type III-like domain, for example, a betasheet containing two strands, and the second RgpA moiety containing ABM2 can form the C-terminal portion of the second fibronectin type III-like domain, for example, a betasheet having four strands and further betasheet chains. In certain such embodiments, the second RgpA moiety containing ABM1 and the second RgpA moiety containing ABM2 can form a second fibronectin type III-like domain having a beta-sandwich structure.

[0138] In some embodiments, the second Kgp moiety containing ABM1 and the second Kgp moiety containing ABM2 each contain the sequence shown in Table 6, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). In certain such embodiments, the second Kgp moiety containing ABM1 and the second Kgp moiety containing ABM2 can form a second fibronectin type III-like domain, as described in the preceding paragraph.

[0139] [Table 6]

[0140] In some embodiments, the second RgpA moiety containing ABM1 and the second RgpA moiety containing ABM2 each contain the sequence shown in Table 7, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). In certain such embodiments, the second RgpA moiety containing ABM1 and the second RgpA moiety containing ABM2 can form a second fibronectin type III-like domain, as described in the preceding paragraph.

[0141] [Table 7]

[0142] The nucleic acids and polypeptides of the present invention, comprising the combination of a first Kgp moiety containing ABM1 and a first Kgp moiety containing ABM2 as described in Table 4, may also comprise a second Kgp moiety containing ABM1 and a second Kgp moiety containing ABM2 as described in Table 6, for example, as shown in Table 8 below. Therefore, in some embodiments, the first Kgp moiety containing ABM1 and the first Kgp moiety containing ABM2 each have the sequence shown in Table 4, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identity thereto. The sequence includes the sequence, and the second Kgp portion containing ABM1 and the second Kgp portion containing ABM2 each contain the sequence shown in Table 6, or a sequence having at least 70% identity thereto (for example, at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, the first Kgp moiety containing ABM1, the first Kgp moiety containing ABM2, the second Kgp moiety containing ABM1, and the second Kgp moiety containing ABM2 each contain the sequence shown in Table 8, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). In certain such embodiments, the first Kgp moiety containing ABM1 and the first Kgp moiety containing ABM2 can form a first fibronectin type III-like domain, and the second Kgp moiety containing ABM1 and the second Kgp moiety containing ABM2 can form a second fibronectin type III-like domain, as described in the preceding paragraph.

[0143] [Table 8]

[0144] The nucleic acids and polypeptides of the present invention, comprising the combination of a first RgpA moiety containing ABM1 and a first RgpA moiety containing ABM2 as described in Table 5, may also comprise a second RgpA moiety containing ABM1 and a second RgpA moiety containing ABM2 as described in Table 7, for example, as shown in Table 9 below. Therefore, in some embodiments, the first RgpA moiety containing ABM1 and the second RgpA moiety containing ABM2 each have the sequence shown in Table 5, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identity thereto. The sequence includes the sequence, and the second RgpA portion containing ABM1 and the second RgpA portion containing ABM2 each contain the sequence shown in Table 7, or a sequence having at least 70% identity thereto (for example, at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, the first RgpA portion containing ABM1, the first RgpA portion containing ABM2, the second RgpA portion containing ABM1, and the second RgpA portion containing ABM2 each contain the sequence shown in Table 9, or a sequence having at least 70% identity thereto (for example, at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). In certain such embodiments, a first RgpA moiety comprising ABM1 and a first RgpA moiety comprising ABM2 can form a first fibronectin type III-like domain, and a second RgpA moiety comprising ABM1 and a second RgpA moiety comprising ABM2 can form a second fibronectin type III-like domain, as described in the preceding paragraph.

[0145] [Table 9]

[0146] In some embodiments, the nucleic acid or polypeptide of the present invention comprises a first Kgp moiety containing ABM1, a first Kgp moiety containing ABM2, a second Kgp moiety containing ABM1, a second Kgp moiety containing ABM2, a first RgpA moiety containing ABM1, and a first RgpA moiety containing ABM2. In a particular embodiment, the first Kgp portion including ABM1 and the first Kgp portion including ABM2 each contain the sequence shown in Table 4 (or sequences having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%)), and the second Kgp portion including ABM1 and the second Kgp portion including ABM2 each contain the sequence shown in Table 6 (or sequences having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87 The first RgpA portion containing ABM1 and the first RgpA portion containing ABM2 each contain the sequences shown in Table 5 (or sequences having at least 70% identity thereto (for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%)), and each contains the sequences shown in Table 5 (or sequences having at least 70% identity thereto (for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%)).

[0147] For example, the first Kgp portion containing ABM1, the second Kgp portion containing ABM2, the second Kgp portion containing ABM1, the second Kgp portion containing ABM2, the first RgpA portion containing ABM1, and the first RgpA portion containing ABM2 each contain the sequences shown in Table 10, or sequences that are at least 70% (for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to them. In certain such embodiments, a first Kgp moiety containing ABM1 and a first Kgp moiety containing ABM2 can form a first fibronectin type III-like domain, a second Kgp moiety containing ABM1 and a second Kgp moiety containing ABM2 can form a second fibronectin type III-like domain, and a first RgpA moiety containing ABM1 and a first RgpA moiety containing ABM2 can form a third fibronectin type III-like domain as described in the preceding paragraph.

[0148] [Table 10]

[0149] Those skilled in the art can determine whether a particular sequence of interest forms a fibronectin type III-like domain through structural modeling in the same manner that the Kgp structure was modeled by the inventors. For example, those skilled in the art can replace the first ABM1 sequence and / or the first ABM2 sequence in the first portion of wild-type Kgp with the sequence of interest. Then, those skilled in the art can use AlphaFold2 to model the structure of a Kgp protein containing the sequence of interest and evaluate whether the sequence forms a fold that is structurally homologous to a fibronectin type III-like domain. Since a fibronectin type III-like domain is known to have a conserved beta-sandwich fold containing one beta sheet with three beta chains and one beta sheet with four chains, structural homology to a fibronectin type III-like domain can be evaluated visually. Alternatively, structural homology can be evaluated by a protein structure comparison server such as DALI (ekhidna2.biocenter.helsinki.fi / dali / lsinki.fi).

[0150] Those skilled in the art can also use functional assays to determine whether a particular sequence of interest forms a fibronectin type III-like domain. Fibronectin type III-like domains are also known to mediate interactions with fibronectin. Therefore, those skilled in the art can perform enzyme-linked immunosorbent assay (ELISA) to determine whether a Kgp construct containing the particular sequence of interest supports fibronectin binding activity. In ELISA, fibronectin is immobilized on the surface of polystyrene microplate wells, and Kgp constructs containing the sequence of interest are added to the wells in serial dilutions. The wells are washed with buffer, and the bound Kgp protein is detected with a high-affinity antibody. A similar method was used to test whether the fibronectin type III-like domain of FlpA in C. jejuni mediates binding to fibronectin (Konkel et al., 2010).

[0151] Gingipain portion containing ABM3 Wild-type Kgp and RgpA contain the ABM3 motif located at the C-terminus of DUF2436 and at the N-terminus of the K1 adhesin domain, as shown in Figures 1, 2A, and 2B. The modular nature of Kgp means that the portion of Kgp containing ABM3 may be included in any of the nucleic acids or polypeptides described herein that include the other portions of Kgp. The same modular structure of RgpA means that the portion of RgpA containing ABM3 may be included in any of the nucleic acids of the polypeptides described herein that include the other portions of RgpA.

[0152] In some embodiments, the nucleic acids and polypeptides of the present invention include a Kgp portion containing ABM3. In some embodiments, the nucleic acids and polypeptides of the present invention include an RgpA portion containing ABM3.

[0153] In some embodiments, the first Kgp moiety containing ABM3 is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, the Kgp moiety containing ABM3 has the sequence of SEQ ID NO: 131. Therefore, in some embodiments, the Kgp moiety containing ABM3 contains the sequence of SEQ ID NO: 131, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0154] In some embodiments, the first RgpA moiety containing ABM3 is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, the RgpA moiety containing ABM3 has the sequence of SEQ ID NO: 135. Therefore, in some embodiments, the RgpA moiety containing ABM3 contains the sequence of SEQ ID NO: 135, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0155] Examples of Kgp or RgpA sequences containing the ABM3 sequence are provided in Table 11, along with the consensus sequence of ABM3. The consensus sequence of ABM3 is sequence number 139.

[0156] [Table 11]

[0157] In some embodiments, the portion of Kgp containing ABM3 includes the sequence of sequence number 139, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0158] In some embodiments, the portion of Kgp containing ABM3 includes the sequence of sequence number 131, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0159] In some embodiments, the portion of Kgp containing ABM3 includes the sequence of sequence number 132, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0160] In some embodiments, the portion of Kgp containing ABM3 includes the sequence of sequence number 94, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0161] In some embodiments, the portion of Kgp containing ABM3 includes the sequence of sequence number 133, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0162] In some embodiments, the portion of Kgp containing ABM3 includes the sequence of sequence number 134, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0163] In some embodiments, the portion of RgpA containing ABM3 includes the sequence of sequence number 139, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0164] In some embodiments, the portion of RgpA containing ABM3 includes the sequence of sequence number 135, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0165] In some embodiments, the portion of RgpA containing ABM3 includes the sequence of sequence number 103, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0166] In some embodiments, the portion of RgpA containing ABM3 includes the sequence of sequence number 136, or a sequence having at least 70% identity thereto (for example, at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0167] In some embodiments, the portion of RgpA containing ABM3 includes the sequence of sequence number 137, or a sequence having at least 70% identity thereto (for example, at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0168] In some embodiments, the portion of RgpA containing ABM3 includes the sequence of sequence number 138, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0169] In the wild-type sequence of Kgp derived from P. gingivalis W50 strain, the Kgp portion containing ABM3 is located C-terminally to DUF2436 and N-terminally to the K1 adhesin domain. The ABM1 sequence is located N-terminally and adjacent to the portion containing ABM3. In this context, the Kgp portion containing ABM3 is 30 amino acids long and contains the sequence of SEQ ID NO: 132. This sequence contains ABM3 of SEQ ID NO: 139. Therefore, in some embodiments, the Kgp portion containing ABM3 is 15 to 30 amino acids long and contains the sequence of SEQ ID NO: 139.

[0170] In some embodiments, the Kgp moiety containing ABM3 is 17 to 26 amino acids long and includes the sequence of SEQ ID NO: 139. In some embodiments, the Kgp moiety containing ABM3 is 20 to 25 amino acids long and includes the sequence of SEQ ID NO: 139.

[0171] In some embodiments, the Kgp moiety containing ABM3 is 15 amino acids long and includes the sequence of SEQ ID NO: 139. In some embodiments, the Kgp moiety containing ABM3 is 17 amino acids long and includes the sequence of SEQ ID NO: 139. In some embodiments, the Kgp moiety containing ABM3 is 20 amino acids long and includes the sequence of SEQ ID NO: 139. In some embodiments, the Kgp moiety containing ABM3 is 25 amino acids long and includes the sequence of SEQ ID NO: 139. In some embodiments, the Kgp moiety containing ABM3 is 26 amino acids long and includes the sequence of SEQ ID NO: 139. In some embodiments, the first Kgp moiety containing ABM3 is 30 amino acids long and includes the sequence of SEQ ID NO: 139.

[0172] In the wild-type RgpA sequence derived from P. gingivalis W50 strain, the RgpA portion containing ABM3 is located C-terminally to DUF2436 and N-terminally to the K1 adhesin domain. The ABM1 sequence is located N-terminally to the portion containing ABM3, and the K1 adhesin domain is located C-terminally and adjacent to the portion containing ABM3. In this context, the RgpA portion containing ABM3 is 30 amino acids long and contains the sequence of SEQ ID NO: 136. This sequence contains ABM3 of SEQ ID NO: 139. Therefore, in some embodiments, the RgpA portion containing ABM3 is 15 to 30 amino acids long and contains the sequence of SEQ ID NO: 139.

[0173] In some embodiments, the RgpA moiety containing ABM3 is 17 to 26 amino acids long and includes the sequence of SEQ ID NO: 139. In some embodiments, the RgpA moiety containing ABM3 is 20 to 25 amino acids long and includes the sequence of SEQ ID NO: 139.

[0174] In some embodiments, the RgpA moiety containing ABM3 is 15 amino acids long and includes the sequence of SEQ ID NO: 139. In some embodiments, the RgpA moiety containing ABM3 is 17 amino acids long and includes the sequence of SEQ ID NO: 139. In some embodiments, the RgpA moiety containing ABM3 is 20 amino acids long and includes the sequence of SEQ ID NO: 139. In some embodiments, the RgpA moiety containing ABM3 is 25 amino acids long and includes the sequence of SEQ ID NO: 139. In some embodiments, the RgpA moiety containing ABM3 is 26 amino acids long and includes the sequence of SEQ ID NO: 139. In some embodiments, the first RgpA moiety containing ABM3 is 30 amino acids long and includes the sequence of SEQ ID NO: 139.

[0175] Position of sequences including ABM relative to other gingipain domains The first Kgp or RgpA moiety containing ABM1 may be different from its adjacent domains (i.e., the Cat domain and DUF2436), or it may overlap with one or both of its adjacent domains. In certain embodiments, the first Kgp or RgpA moiety containing ABM1 is different from the Cat domain and DUF2436. In certain embodiments, the peptide linker may be positioned between the Cat domain and the first Kgp or RgpA moiety containing ABM1. In certain embodiments, the peptide linker may be positioned between the first Kgp or RgpA moiety containing ABM1 and DUF2436. In certain embodiments, the peptide linker may be positioned between the Cat domain and the first Kgp or RgpA moiety containing ABM1, and the peptide linker may be positioned between the first Kgp or RgpA moiety containing ABM1 and DUF2436.

[0176] In some embodiments, the first Kgp or RgpA portion containing ABM1 overlaps with the Cat domain. In certain such embodiments, the first Kgp or RgpA portion containing ABM1 overlaps with the Cat domain and is different from DUF2436.

[0177] The first Kgp or RgpA moiety containing ABM2 may be different from its adjacent domains (i.e., DUF2436 and the Kgp or RgpA moiety containing ABM1), or it may overlap with one or both of its adjacent domains. In certain embodiments, the first Kgp or RgpA moiety containing ABM2 is different from DUF2436 and the Kgp or RgpA moiety containing ABM1. In certain embodiments, the peptide linker may be positioned between the Cat domain and the first Kgp or RgpA moiety containing ABM1. In certain embodiments, the peptide linker may be positioned between the first Kgp or RgpA moiety containing ABM1 and DUF2436. In certain embodiments, the peptide linker may be positioned between the Cat domain and the first Kgp or RgpA moiety containing ABM1, and the peptide linker may be positioned between the first Kgp or RgpA moiety containing ABM1 and DUF2436.

[0178] The second Kgp or RgpA moiety containing ABM1 may be different from its adjacent domains (i.e., the first Kgp or RgpA moiety containing ABM2 and the Kgp or RgpA moiety containing ABM3), or it may overlap with one or both of its adjacent domains. In certain embodiments, the second Kgp or RgpA moiety containing ABM1 is different from the first Kgp or RgpA moiety containing ABM2 and the Kgp or RgpA moiety containing ABM3. In certain embodiments, the peptide linker may be positioned between the first Kgp or RgpA moiety containing ABM2 and the second Kgp or RgpA moiety containing ABM1. In certain embodiments, the peptide linker may be positioned between the second Kgp or RgpA moiety containing ABM1 and the Kgp or RgpA moiety containing ABM3. In a particular embodiment, the peptide linker may be positioned between a first Kgp or RgpA moiety containing ABM2 and a second Kgp or RgpA moiety containing ABM1, and the peptide linker may be positioned between a second Kgp or RgpA moiety containing ABM1 and a Kgp or RgpA moiety containing ABM3.

[0179] In some embodiments, a second Kgp or RgpA portion including ABM1 overlaps with a Kgp or RgpA portion including ABM3. In certain such embodiments, a Kgp or RgpA portion including ABM1 differs from a first Kgp or RgpA portion including ABM2 and overlaps with a Kgp or RgpA portion including ABM3.

[0180] The second Kgp or RgpA moiety containing ABM2 may be different from its adjacent domain (i.e., the K2 adhesin domain), or it may overlap with one or both of its adjacent domains. In certain embodiments, the second Kgp or RgpA moiety containing ABM2 is different from the K2 adhesin domain. In certain embodiments, the peptide linker may be positioned between the second Kgp or RgpA moiety containing ABM2 and the K2 adhesin domain.

[0181] The Kgp or RgpA portion containing ABM3 may be different from or overlap with its adjacent domains (i.e., a second Kgp or RgpA portion containing ABM1 and the K1 adhesin domain). In certain embodiments, the Kgp or RgpA portion containing ABM3 overlaps with a second Kgp or RgpA portion containing ABM1. In certain embodiments, the Kgp or RgpA portion containing ABM3 overlaps with the K1 adhesin domain. In certain embodiments, the Kgp or RgpA portion containing ABM3 overlaps with a second Kgp or RgpA portion containing ABM1 and the K1 adhesin domain.

[0182] Unknown Function Domain (DUF) The nucleic acids and polypeptides of the present invention include at least a portion of unknown functional domains (DUFs) as disclosed herein. The modular nature of gingipain is such that any of the DUFs disclosed herein can be combined with any of the other domains disclosed herein. For example, any of the sequences disclosed herein, including ABM, may be combined with any of the DUFs described in the following paragraphs.

[0183] Domains of unknown function (DUFs) are protein domains whose function has not been characterized. Therefore, domains initially designated as DUFs may be renamed after their function is established, or they may be grouped into existing families of already characterized domains. DUFs are classified and organized in the Pfam database (pfam.xfam.org), and each stored DUF is assigned a number (DUF1, DUF2, etc.). This means that if two DUFs show a sufficient degree of homology, a DUF found in the first protein may be assigned the same number as a DUF found in the second protein. The Pfam database is now part of the InterPro database (www.ebi.ac.uk / interpro / ), a database that classifies proteins beyond just DUFs (Paysan-Lafosse et al., 2022).

[0184] DUF has been identified in the Kgp and RgpA of P. gingivalis and classified as DUF2436 (Dashper et al., 2017). In the InterPro database, DUF2436 is assigned registration number IPR018832. The nucleic acids and polypeptides of the present invention contain at least a portion of DUF2436.

[0185] Those skilled in the art can determine whether a particular sequence is at least part of DUF2436 by comparing it with known DUF2436 sequences. For example, the InterPro database allows for the retrieval of a particular sequence, thereby enabling the identification of sequences containing at least part of DUF2436. DUF2436 is found in many different organisms and proteins, and any DUF2436 may be used in the present invention, regardless of whether its particular sequence is found in P. gingivalis. For example, using a non-gingipain protein or DUF2436 derived from a non-P. gingivalis species may allow the remaining P. gingivalis domain of the polypeptide to fold into a structure sufficiently similar to the three-dimensional structure of wild-type gingipain.

[0186] Typically, at least a portion of DUF2436 according to the present invention is derived from P. gingivalis DUF2436. In some embodiments, DUF2436 is derived from P. gingivalis Kgp. In some embodiments, DUF2436 is derived from P. gingivalis RgpA.

[0187] In some embodiments, at least a portion of Kgp DUF2436 is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, Kgp DUF2436 has the sequence of SEQ ID NO: 168. Therefore, in some embodiments, at least a portion of Kgp DUF2436 contains a sequence that is identical to at least a portion of SEQ ID NO: 168, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) thereto. For example, at least a portion of Kgp DUF2436 contains the sequence of sequence number 168, or a sequence that is at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to it.

[0188] In some embodiments, at least a portion of RgpA DUF2436 is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, RgpA DUF2436 has the sequence of SEQ ID NO: 171. Therefore, in some embodiments, at least a portion of RgpA DUF2436 includes at least a portion of SEQ ID NO: 171, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of RgpA DUF2436 contains the sequence of sequence number 171, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0189] Examples of DUF2436 sequences that can be used in accordance with the present invention are provided in Table 12 below.

[0190] [Table 12]

[0191] In some embodiments, at least a portion of Kgp DUF2436 includes a sequence that is at least a portion of sequence number 90, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to it. For example, at least a portion of Kgp DUF2436 includes a sequence that is at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to it.

[0192] In some embodiments, at least a portion of Kgp DUF2436 includes a sequence that is at least a portion of sequence number 169, or a sequence that is at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to it. For example, at least a portion of Kgp DUF2436 includes a sequence that is at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to it.

[0193] In some embodiments, at least a portion of Kgp DUF2436 includes a sequence that is at least a portion of sequence number 170, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to it. For example, at least a portion of Kgp DUF2436 includes a sequence that is at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to it.

[0194] In some embodiments, at least a portion of RgpA DUF2436 includes a sequence that is identical to at least a portion of sequence number 100, or to it by at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of RgpA DUF2436 includes a sequence that is identical to sequence number 100, or to it by at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0195] In some embodiments, at least a portion of RgpA DUF2436 includes a sequence that is identical to at least a portion of sequence number 172, or to it by at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of RgpA DUF2436 includes a sequence that is identical to sequence number 172, or to it by at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0196] In some embodiments, at least a portion of RgpA DUF2436 includes a sequence that is at least a portion of sequence number 173, or a sequence that is at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to it. For example, at least a portion of RgpA DUF2436 includes a sequence that is at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to it.

[0197] The inventors have found that nucleic acids and polypeptides containing full-length DUF2436 may be particularly advantageous in inducing an immune response. Therefore, in certain embodiments, the nucleic acid or polypeptide of the present invention contains full-length DUF2436.

[0198] The full-length DUF2436 domain refers to the unshortened DUF2436 relative to the corresponding wild-type sequence. Therefore, in some embodiments, at least a portion of the DUF2436 is full-length DUF2436, which is the same length as the corresponding wild-type DUF2436 sequence.

[0199] For example, DUF2436 present in Kgp from P. gingivalis W50 strain is 162 amino acids long. Therefore, in some embodiments, the full-length Kgp DUF2436 is at least 162 amino acids long (e.g., 162 amino acids long). DUF2436 present in RgpA from P. gingivalis W50 strain is 163 amino acids long. Therefore, in some embodiments, the full-length RgpA DUF2436 is at least 163 amino acids long (e.g., 163 amino acids long).

[0200] In some embodiments, the full-length Kgp DUF2436 is at least 160 amino acids long (e.g., 160 amino acids long). In some embodiments, the full-length Kgp DUF2436 is at least 161 amino acids long (e.g., 161 amino acids long). In some embodiments, the full-length Kgp DUF2436 is at least 162 amino acids long (e.g., 162 amino acids long). In some embodiments, the full-length Kgp DUF2436 is at least 163 amino acids long (e.g., 166 amino acids long). In some embodiments, the full-length RgpA DUF2436 is 160 amino acids long. In some embodiments, the full-length RgpA DUF2436 is at least 161 amino acids long (e.g., 161 amino acids long). In some embodiments, the full-length RgpA DUF2436 is at least 162 amino acids long (e.g., 162 amino acids long). In some embodiments, the full-length RgpA DUF2436 is at least 163 amino acids long (e.g., 163 amino acids long).

[0201] In some embodiments, at least a portion of Kgp DUF2436 is full-length DUF2436 and includes the sequence of sequence number 168, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0202] In some embodiments, at least a portion of Kgp DUF2436 is full-length DUF2436 and includes the sequence of sequence number 90, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0203] In some embodiments, at least a portion of Kgp DUF2436 is full-length DUF2436 and includes the sequence of sequence number 169, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0204] In some embodiments, at least a portion of Kgp DUF2436 is full-length DUF2436 and includes the sequence of sequence number 170, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0205] In some embodiments, at least a portion of RgpA DUF2436 is full-length DUF2436 and includes the sequence of sequence number 171, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0206] In some embodiments, at least a portion of RgpA DUF2436 is full-length DUF2436 and includes the sequence of sequence number 100, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0207] In some embodiments, at least a portion of RgpA DUF2436 is full-length DUF2436 and includes the sequence of sequence number 172, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0208] DUF2436 can also be shortened without significantly altering the properties of the resulting polypeptide. Therefore, in some embodiments, at least a portion of DUF2436 is shortened DUF2436, where 1 to 35 amino acids are shortened. In certain such embodiments, DUF2436 is shortened by 1 to 30 amino acids, 1 to 25 amino acids, 1 to 20 amino acids, 1 to 15 amino acids, 1 to 10 amino acids, or 1 to 5 amino acids. In some embodiments, DUF2436 is shortened by 30 amino acids. In some embodiments, DUF2436 is shortened by 25 amino acids. In some embodiments, DUF2436 is shortened by 20 amino acids. In some embodiments, DUF2436 is shortened by 15 amino acids. In some embodiments, DUF2436 is shortened by 10 amino acids. In some embodiments, DUF2436 is shortened by 5 amino acids.

[0209] The shortening may occur at the N-terminus or C-terminus of DUF. Therefore, in some embodiments, DUF2436 has 1 to 35 amino acids shortened at its N-terminus. Thus, in some embodiments, at least a portion of DUF2436 is a shortened DUF2436, and the shortened DUF2436 has 1 to 35 amino acids shortened at its N-terminus. In certain such embodiments, DUF2436 has 1 to 30 amino acids shortened at its N-terminus, 1 to 25 amino acids shortened at its N-terminus, 1 to 20 amino acids shortened at its N-terminus, 1 to 15 amino acids shortened at its N-terminus, 1 to 10 amino acids shortened at its N-terminus, and 1 to 5 amino acids shortened at its N-terminus. In some embodiments, DUF2436 has 30 amino acids shortened at its N-terminus. In some embodiments, DUF2436 has 25 amino acids shortened at its N-terminus. In some embodiments, DUF2436 has 20 amino acids shortened at its N-terminus. In some embodiments, DUF2436 has 15 amino acids shortened at its N-terminus. In some embodiments, DUF2436 has 10 amino acids shortened at its N-terminus. In some embodiments, DUF2436 has 5 amino acids shortened at its N-terminus.

[0210] In some embodiments, DUF2436 has 1 to 35 amino acids shortened at its C-terminus. Thus, in some embodiments, at least a portion of DUF2436 is shortened DUF2436, and shortened DUF2436 has 1 to 35 amino acids shortened at its C-terminus. In certain such embodiments, DUF2436 has 1 to 30 amino acids shortened at its C-terminus, 1 to 25 amino acids shortened at its C-terminus, 1 to 20 amino acids shortened at its C-terminus, 1 to 15 amino acids shortened at its C-terminus, 1 to 10 amino acids shortened at its C-terminus, and 1 to 5 amino acids shortened at its C-terminus. In some embodiments, DUF2436 has 30 amino acids shortened at its C-terminus. In some embodiments, DUF2436 has 25 amino acids shortened at its C-terminus. In some embodiments, DUF2436 has 20 amino acids shortened at its C-terminus. In some embodiments, DUF2436 has 15 amino acids shortened at its C-terminus. In some embodiments, DUF2436 has 10 amino acids shortened at its C-terminus. In some embodiments, DUF2436 has 5 amino acids shortened at its C-terminus.

[0211] In some embodiments, the shortened DUF2436 includes at least a portion of sequence number 173, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0212] Variants of DUF2436 may also be used in the present invention. Such variants may be used to remove glycosylated sites, as described elsewhere in this specification.

[0213] Catalytic domain (Cat domain) The nucleic acids and polypeptides of the present invention, as disclosed herein, comprise at least a portion of a catalytic domain derived from Kgp and / or at least a portion of a catalytic domain derived from RgpA or RgpB. The modular nature of gingipain is such that any of the catalytic domains disclosed herein can be combined with any of the other domains disclosed herein. For example, any of the sequences disclosed herein, including ABM and / or DUF, may be combined with any of the catalytic domains described in the following paragraphs.

[0214] Kgp, RgpA, and RgpB are lysine-specific and arginine-specific cysteine ​​proteinases belonging to the C25 peptidase family, in which proteinase activity is mediated by a catalytic domain located at the N-terminus of the active wild-type protein. The catalytic domain as defined herein includes the C25 peptidase domain and an immunoglobulin fold (C25C) at its C-terminus (Dashper et al., 2017). The nucleic acids and polypeptides of the present invention include at least a portion of the Kgp, RgpA, and / or RgpB catalytic domains for the purpose of inducing an antibody response that inhibits the proteinase function of Kgp, RgpA, and / or RgpB. The InterPro database entry for the C25 peptidase domain is IPR001769. The InterPro database entry for the C25C domain is IPR005536.

[0215] In some embodiments, at least a portion of the Kgp catalytic domain according to the present invention is derived from P. gingivalis Kgp.

[0216] In some embodiments, at least a portion of the RgpA catalytic domain according to the present invention is derived from P. gingivalis RgpA.

[0217] In some embodiments, at least a portion of the RgpB catalytic domain according to the present invention is derived from P. gingivalis RgpB.

[0218] In some embodiments, at least a portion of the Kgp catalytic domain is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, the Kgp catalytic domain has the sequence of SEQ ID NO: 174. Therefore, in some embodiments, at least a portion of the Kgp catalytic domain includes at least a portion of SEQ ID NO: 174, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the KGP catalytic domain contains the sequence of SEQ ID NO: 174, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0219] In some embodiments, at least a portion of the RgpA catalytic domain is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, the RgpA catalytic domain has the sequence of SEQ ID NO: 178. Therefore, in some embodiments, at least a portion of the RgpA catalytic domain includes at least a portion of SEQ ID NO: 178, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the RgpA catalytic domain contains the sequence of Sequence ID No. 178, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0220] In some embodiments, at least a portion of the RgpB catalytic domain is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, the RgpB catalytic domain has the sequence of SEQ ID NO: 251. Therefore, in some embodiments, at least a portion of the RgpB catalytic domain includes at least a portion of SEQ ID NO: 251, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the RgpB catalytic domain includes the sequence of SEQ ID NO: 251, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0221] Examples of Kgp, RgpA, and RgpB catalytic domain sequences that can be used in accordance with the present invention are provided in Table 13 below.

[0222] [Table 13]

[0223] [Table 14]

[0224] [Table 15]

[0225] In some embodiments, at least a portion of the Kgp catalytic domain includes a sequence that is at least a portion of sequence number 64, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to it. For example, at least a portion of the Kgp catalytic domain includes a sequence that is at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to it.

[0226] In some embodiments, at least a portion of the Kgp catalytic domain includes a sequence that is identical to at least a portion of Sequence ID No. 174, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) thereto. For example, at least a portion of the Kgp catalytic domain includes a sequence that is identical to Sequence ID No. 174, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) thereto.

[0227] In some embodiments, at least a portion of the Kgp catalytic domain includes a sequence that is identical to at least a portion of Sequence ID No. 175, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) thereto. For example, at least a portion of the Kgp catalytic domain includes a sequence that is identical to Sequence ID No. 175, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) thereto.

[0228] In some embodiments, at least a portion of the Kgp catalytic domain includes a sequence that is identical to at least a portion of sequence number 88, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) thereto. For example, at least a portion of the Kgp catalytic domain includes a sequence that is identical to sequence number 88, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) thereto.

[0229] In some embodiments, at least a portion of the Kgp catalytic domain includes at least a portion of SEQ ID NO: 61, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the Kgp catalytic domain includes the sequence of SEQ ID NO: 61, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0230] In some embodiments, at least a portion of the Kgp catalytic domain includes a sequence that is at least a portion of sequence number 163, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical thereto. For example, at least a portion of the Kgp catalytic domain includes a sequence that is at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical thereto.

[0231] In some embodiments, at least a portion of the Kgp catalytic domain includes a sequence that is identical to at least a portion of SEQ ID NO: 176, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) thereto. For example, at least a portion of the Kgp catalytic domain includes a sequence that is identical to SEQ ID NO: 176, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) thereto.

[0232] In some embodiments, at least a portion of the Kgp catalytic domain includes a sequence that is at least a portion of SEQ ID NO: 177, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to it. For example, at least a portion of the Kgp catalytic domain includes a sequence that is at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to it.

[0233] In some embodiments, at least a portion of the RgpA catalytic domain includes a sequence that is at least a portion of sequence number 98, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to it.

[0234] In some embodiments, at least a portion of the RgpA catalytic domain includes a sequence that is identical to at least a portion of SEQ ID NO: 179, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) thereto. For example, at least a portion of the RgpA catalytic domain includes a sequence that is identical to SEQ ID NO: 179, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) thereto.

[0235] In some embodiments, at least a portion of the RgpA catalytic domain includes at least a portion of sequence number 97, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the RgpA catalytic domain includes the sequence of sequence number 97, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0236] In some embodiments, at least a portion of the RgpA catalytic domain includes a sequence that is identical to at least a portion of SEQ ID NO: 180, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) thereto. For example, at least a portion of the RgpA catalytic domain includes a sequence that is identical to SEQ ID NO: 180, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) thereto.

[0237] In some embodiments, at least a portion of the RgpA catalytic domain includes at least a portion of SEQ ID NO: 66, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the RgpA catalytic domain includes the sequence of SEQ ID NO: 66, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0238] In some embodiments, at least a portion of the RgpA catalytic domain includes a sequence that is at least a portion of sequence number 181, or at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to it. For example, at least a portion of the RgpA catalytic domain includes a sequence that is at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical to it.

[0239] In some embodiments, at least a portion of the RgpB catalytic domain includes at least a portion of SEQ ID NO: 251, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the RgpB catalytic domain includes the sequence of SEQ ID NO: 251, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0240] Preferably, the Kgp, RgpA, or RgpB catalytic domain is modified to inactivate its proteinase function. This ensures that the polypeptide does not mediate any negative effects associated with Kgp, RgpA, or RgpB proteinase function. It is possible to inactivate the proteinase function of the catalytic domain in different ways. For example, one or more residues in the active site of the catalytic domain may be mutated. Alternatively, or in addition, the catalytic domain may be shortened to form an inactivated catalytic domain.

[0241] Some of the catalytic domain sequences in Table 13 have been inactivated by mutation and / or shortening. In particular, the Kgp catalytic domain of SEQ ID NO: 64 and the RgpA catalytic domain of SEQ ID NO: 98 have been inactivated by mutation. The Kgp catalytic domains of SEQ ID NOs: 88 and 163, the RgpA catalytic domains of SEQ ID NOs: 97 and 166, and the RgpB catalytic domains of SEQ ID NOs: 252 and 253 have been inactivated by shortening.

[0242] In some embodiments, at least a portion of the Kgp catalytic domain includes a mutation that inactivates proteinase activity. In certain such embodiments, the mutation that inactivates proteinase activity is a cysteine-to-serine mutation at position 477 (C477S), where the mutation site corresponds to position 477 of the wild-type Kgp sequence of SEQ ID NO: 157.

[0243] For example, in the catalytic domains of SEQ ID NOs. 64, 174, 176, and 178, the position corresponding to position 477 of the wild-type Kgp sequence in SEQ ID NO. 157 is position 249. Therefore, the mutation that inactivates proteinase activity in the context of these catalytic domains is a cysteine-to-serine mutation (C249S) at position 249.

[0244] In some embodiments, at least a portion of the RgpA catalytic domain contains a mutation that inactivates proteinase activity. In certain such embodiments, the mutation that inactivates proteinase activity is a cysteine-to-serine mutation at position 471 (C471S), the mutation site corresponding to position 471 of the wild-type RgpA sequence of SEQ ID NO: 158.

[0245] For example, in the catalytic domains of sequence numbers 98, 179, and 181, the position corresponding to position 471 of the wild-type RgpA sequence in sequence number 158 is position C248. Therefore, the mutation that inactivates proteinase activity in the context of these catalytic domains is a cysteine-to-serine mutation (C248S) at position 248.

[0246] In some embodiments, at least a portion of the RgpB catalytic domain includes a mutation that inactivates proteinase activity. In certain such embodiments, the mutation that inactivates proteinase activity is a cysteine-to-serine mutation at position 473 (C473S), the mutation site corresponding to position 473 of the wild-type RgpB sequence of SEQ ID NO: 159.

[0247] The inventors have found that nucleic acids and polypeptides containing full-length catalytic domains may be particularly advantageous in inducing an immune response. Accordingly, in some embodiments, at least a portion of the Kgp catalytic domain includes a full-length Kgp catalytic domain. In some embodiments, at least a portion of the RgpA catalytic domain includes a full-length RgpA catalytic domain. In some embodiments, at least a portion of the RgpB catalytic domain includes a full-length RgpB catalytic domain.

[0248] A full-length Kgp, RgpA, or RgpB catalytic domain refers to a Kgp, RgpA, or RgpB catalytic domain that is not shortened relative to the corresponding wild-type sequence. Therefore, in some embodiments, at least a portion of the Kgp catalytic domain is a full-length Kgp catalytic domain that is the same length as the corresponding wild-type Kgp catalytic domain sequence. In some embodiments, at least a portion of the RgpA catalytic domain is a full-length RgpA catalytic domain that is the same length as the corresponding wild-type RgpA catalytic domain sequence. In some embodiments, at least a portion of the RgpB catalytic domain is a full-length RgpB catalytic domain that is the same length as the corresponding wild-type RgpB catalytic domain sequence.

[0249] For example, the Kgp catalytic domain present in Kgp from P. gingivalis W50 strain is 452 amino acids long. Therefore, in some embodiments, the full-length Kgp catalytic domain is at least 452 amino acids long (e.g., 452 amino acids long). The catalytic domain present in RgpA from P. gingivalis W50 strain is 438 amino acids long. Therefore, in some embodiments, the full-length RgpA catalytic domain is at least 438 amino acids long (e.g., 438 amino acids long). The catalytic domain present in RgpB from P. gingivalis W50 strain is 437 amino acids long. Therefore, in some embodiments, the full-length RgpB catalytic domain is at least 437 amino acids long (e.g., 437 amino acids long). The Kgp catalytic domain present in Kgp from other P. gingivalis strains may have a different length than the Kgp catalytic domain present in Kgp from P. gingivalis strain W50. Similarly, the RgpA catalytic domain present in RgpA from other P. gingivalis strains may have a different length than the RgpA catalytic domain present in RgpA from P. gingivalis strain W50. The RgpB catalytic domain present in RgpB from other P. gingivalis strains may also have a different length than the RgpB catalytic domain present in RgpB from P. gingivalis strain W50.

[0250] Therefore, in some embodiments, the full-length Kgp catalytic domain is at least 448 amino acids long (e.g., 448 amino acids). In some embodiments, the full-length Kgp catalytic domain is at least 449 amino acids long (e.g., 449 amino acids). In some embodiments, the full-length Kgp catalytic domain is at least 450 amino acids long (e.g., 450 amino acids). In some embodiments, the full-length Kgp catalytic domain is at least 451 amino acids long (e.g., 451 amino acids). In some embodiments, the full-length Kgp catalytic domain is at least 456 amino acids long (e.g., 456 amino acids).

[0251] In some embodiments, the full-length RgpA catalytic domain is at least 434 amino acids long (e.g., 434 amino acids). In some embodiments, the full-length RgpA catalytic domain is at least 435 amino acids long (e.g., 435 amino acids). In some embodiments, the full-length RgpA catalytic domain is at least 436 amino acids long (e.g., 436 amino acids). In some embodiments, the full-length RgpA catalytic domain is at least 437 amino acids long (e.g., 437 amino acids).

[0252] In some embodiments, the full-length RgpB catalytic domain is at least 433 amino acids long (e.g., 433 amino acids). In some embodiments, the full-length RgpB catalytic domain is at least 434 amino acids long (e.g., 434 amino acids). In some embodiments, the full-length RgpB catalytic domain is at least 435 amino acids long (e.g., 435 amino acids). In some embodiments, the full-length RgpB catalytic domain is at least 436 amino acids long (e.g., 436 amino acids).

[0253] Truncation of the Kgp, RgpA or RgpB catalytic domain can also be made such that, without significantly modifying the properties of the resulting polypeptide, for example, the polypeptide can induce an antibody that can block the catalytic function of Kgp, RgpA and / or RgpB. Thus, in some embodiments, at least a portion of the Kgp catalytic domain is a truncated Kgp catalytic domain. In some embodiments, at least a portion of the RgpA catalytic domain is a truncated RgpA catalytic domain. In some embodiments, at least a portion of the RgpB catalytic domain is a truncated RgpB catalytic domain.

[0254] Therefore, in some embodiments, at least a portion of the Kgp, RgpA, or RgpB catalytic domain is a shortened catalytic domain, where 1 to 35 amino acids are shortened. In certain such embodiments, the Kgp catalytic domain is shortened by 1 to 30 amino acids, 1 to 25 amino acids, 1 to 20 amino acids, 1 to 15 amino acids, 1 to 10 amino acids, and 1 to 5 amino acids. In certain such embodiments, the RgpA catalytic domain is shortened by 1 to 30 amino acids, 1 to 25 amino acids, 1 to 20 amino acids, 1 to 15 amino acids, 1 to 10 amino acids, and 1 to 5 amino acids. In certain such embodiments, the RgpB catalytic domain is shortened by 1 to 30 amino acids, 1 to 25 amino acids, 1 to 20 amino acids, 1 to 15 amino acids, 1 to 10 amino acids, and 1 to 5 amino acids. In some embodiments, the Kgp, RgpA, or RgpB catalytic domain is shortened by 30 amino acids. In some embodiments, the Kgp, RgpA, or RgpB catalytic domain is shortened by 25 amino acids. In some embodiments, the Kgp, RgpA, or RgpB catalytic domain is shortened by 20 amino acids. In some embodiments, the Kgp, RgpA, or RgpB catalytic domain is shortened by 15 amino acids. In some embodiments, the Kgp, RgpA, or RgpB catalytic domain is shortened by 10 amino acids. In some embodiments, the Kgp, RgpA, or RgpB catalytic domain is shortened by 5 amino acids.

[0255] The truncated Kgp, RgpA or RgpB catalytic domain can maintain the amino acid sequence of the active site as in the case of the wild type, and thus may be advantageous because it can induce antibodies specific to the natural Kgp, RgpA or RgpB active site. Thus, in some embodiments, at least a portion of the Kgp catalytic domain is a truncated Kgp catalytic domain, and the truncation inactivates the protease activity. In some embodiments, at least a portion of the RgpA catalytic domain is a truncated RgpA catalytic domain, and the truncation inactivates the protease activity. In some embodiments, at least a portion of the RgpB catalytic domain is a truncated RgpB catalytic domain, and the truncation inactivates the protease activity.

[0256] An example of a truncated Kgp catalytic domain with inactivated protease activity is the Lys-dinzipine active site peptide (KAS peptide). The KAS peptide is a part of the Kgp catalytic domain that includes a part of the Kgp catalytic domain active site. Examples of different KAS peptides are listed in Table 13, particularly the Kas2 peptide and the extended Kas2 peptide.

[0257] In some embodiments, the shortened Kgp catalytic domain (e.g., KAS peptide) is at least 36 amino acids long (e.g., 36 amino acids) and includes a portion of the Kgp catalytic domain active site. In other embodiments, the shortened Kgp catalytic domain (e.g., the extended KAS2 peptide) is at least 47 amino acids long (e.g., 47 amino acids) and includes a portion of the Kgp catalytic domain active site. In some embodiments, the shortened Kgp catalytic domain comprises a KAS peptide, which comprises a sequence having at least 70% identity to the Kas2 peptide according to SEQ ID NO: 61 (e.g., at least 75, 80, 85, 90, or 95%; or, for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). In other embodiments, the KAS peptide includes the Kas2 peptide according to SEQ ID NO: 163, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0258] In some embodiments, the shortened Kgp catalytic domain comprises a KAS peptide, which comprises a sequence having at least 70% identity to the extended Kas2 peptide according to SEQ ID NO: 175 (e.g., at least 75, 80, 85, 90, or 95%; or, for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). In other embodiments, the KAS peptide comprises a sequence having at least 70% identity to the extended Kas2 peptide according to SEQ ID NO: 88 (e.g., at least 75, 80, 85, 90, or 95%; or, for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0259] An example of a shortened RgpA or RgpB catalytic domain in which proteinase activity is inactivated is the Arg-gingipain active site peptide (RAS peptide). A RAS peptide is a portion of an RgpA or RgpB catalytic domain that contains a portion of the RgpA or RgpB catalytic domain active site. Examples of different RAS peptides are listed in Table 13, in particular the Ras2 peptide and the extended Ras2 peptide.

[0260] In some embodiments, the shortened RgpA or RgpB catalytic domain (e.g., RAS peptide) is at least 36 amino acids long (e.g., 36 amino acids) and includes a portion of the RgpA or RgpB catalytic domain active site. In other embodiments, the shortened RgpA or RgpB catalytic domain (e.g., extended RAS2 peptide) is at least 47 amino acids long (e.g., 47 amino acids) and includes a portion of the RgpA or RgpB catalytic domain active site.

[0261] In some embodiments, the shortened RgpA catalytic domain comprises a RAS peptide, the RAS peptide comprising a sequence having at least 70% identity to the Ras2 peptide according to SEQ ID NO: 180 (e.g., at least 75, 80, 85, 90, or 95%; or, for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). In other embodiments, the RAS peptide comprises a sequence having at least 70% identity to the Ras2 peptide according to SEQ ID NO: 166 (e.g., at least 75, 80, 85, 90, or 95%; or, for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0262] In some embodiments, the shortened RgpA catalytic domain includes a RAS peptide, which includes a sequence having at least 70% identity (e.g., at least 75, 80, 85, 90, or 95%; or, for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) to the extended Ras2 peptide according to SEQ ID NO: 166, or86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) to the extended Ras2 peptide according to SEQ ID NO: 166, or at least 70% identity (e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) to the extended Ras2 peptide according to SEQ ID NO: 179.

[0263] In some embodiments, the shortened RgpB catalytic domain comprises a RAS peptide, the RAS peptide comprising a sequence having at least 70% identity to the Ras2 peptide according to SEQ ID NO: 253 (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0264] In some embodiments, the shortened RgpB catalytic domain comprises a RAS peptide, the RAS peptide comprising a sequence having at least 70% identity to the extended Ras2 peptide according to SEQ ID NO: 252 (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0265] Variants of the Kgp, RgpA, or RgpB catalytic domain may also be used in the present invention. Such variants may be used to remove glycosylation sites, as described elsewhere in this specification.

[0266] As described elsewhere in this specification, the catalytic domain of Kgp has low sequence conservation with the catalytic domains of RgpA and RgpB. Therefore, it may be advantageous for nucleic acids and polypeptides to contain at least a portion of the catalytic domain of Kgp and at least a portion of the catalytic domain of RgpA or RgpB, because this may induce an immune response that can inactivate the catalytic activity of both Kgp and RgpA or RgpB. Alternatively, a composition in which the first nucleic acid or polypeptide contains at least a portion of the catalytic domain of Kgp and the second nucleic acid or polypeptide contains at least a portion of the catalytic domain of RgpA or RgpB may be advantageous.

[0267] Therefore, at least one of the Kgp catalytic domains defined in the preceding paragraph may be combined with at least one of the RgpA catalytic domains defined in the preceding paragraph. In other embodiments, at least one of the Kgp catalytic domains defined in the preceding paragraph may be combined with at least one of the RgpB catalytic domains defined in the preceding paragraph.

[0268] In some embodiments, the nucleic acid or polypeptide (i) has at least a portion of the Kgp catalytic domain (at least a portion of the Kgp catalytic domain is a shortened Kgp catalytic domain containing a KAS peptide, the KAS peptide being the Kas2 peptide according to SEQ ID NO: 88, or having at least 70% (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identity thereto (ii) a sequence including

[0269] In some embodiments, the nucleic acid or polypeptide comprises at least a portion of the Kgp catalytic domain and at least a portion of the RgpA or RgpB catalytic domain. The full-length Kgp and full-length RgpA or RgpB catalytic domains may be too long to be combined in a single nucleic acid or polypeptide with other Kgp and RgpA or RgpB domains present in the nucleic acid or polypeptide (e.g., portions of Kgp or RgpA including DUF2436, ABM, and the K1 adhesin domain). The use of shortened Kgp catalytic domains and / or shortened RgpA or RgpB catalytic domains may be used to avoid any problems such as the nucleic acid or polypeptide being too long to be accurately expressed.

[0270] Accordingly, in some embodiments, at least a portion of the Kgp catalytic domain is a shortened Kgp catalytic domain, and at least a portion of the RgpA catalytic domain is a shortened RgpA catalytic domain. Any of the shortened Kgp catalytic domains disclosed herein may be used in combination with any of the shortened RgpA catalytic domains disclosed herein. For example, in some embodiments, at least a portion of the Kgp catalytic domain is a KAS peptide (e.g., an extended Kas2 peptide of SEQ ID NO: 88, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%)), and at least a portion of the RgpA catalytic domain is a RAS peptide (e.g., an extended Ras2 peptide of SEQ ID NO: 97, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%)). In other embodiments, at least a portion of the Kgp catalytic domain is a KAS peptide (for example, the extended Kas2 peptide of SEQ ID NO: 88, or a sequence having at least 70% identity thereto (for example, at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%)), and at least a portion of the RgpA catalytic domain is a shortened RgpA catalytic domain. In other embodiments, at least a portion of the Kgp catalytic domain is a shortened Kgp catalytic domain, and at least a portion of the RgpA catalytic domain is a RAS peptide (for example, the extended Ras2 peptide of SEQ ID NO: 97, or a sequence having at least 70% identity thereto (for example, at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%)).

[0271] In other embodiments, at least a portion of the Kgp catalytic domain is a full-length catalytic domain, and at least a portion of the RgpA catalytic domain is a shortened RgpA catalytic domain. For example, in certain such embodiments, at least a portion of the Kgp catalytic domain is a full-length catalytic domain, and at least a portion of the RgpA catalytic domain is a RAS peptide (e.g., the extended Ras2 peptide of SEQ ID NO: 97, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%)).

[0272] In other embodiments, at least a portion of the Kgp catalytic domain is a shortened Kgp catalytic domain, and at least a portion of the RgpA catalytic domain is a full-length RgpA catalytic domain. For example, in certain such embodiments, at least a portion of the Kgp catalytic domain is a KAS peptide (e.g., the extended Kas2 peptide of SEQ ID NO: 88, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%)), and at least a portion of the RgpA catalytic domain is a full-length RgpA catalytic domain.

[0273] In some embodiments, at least a portion of the Kgp catalytic domain is a shortened Kgp catalytic domain, and at least a portion of the RgpB catalytic domain is a shortened RgpB catalytic domain. Any of the shortened Kgp catalytic domains disclosed herein may be used in combination with any of the shortened RgpB catalytic domains disclosed herein. For example, in some embodiments, at least a portion of the Kgp catalytic domain is a KAS peptide (e.g., an extended Kas2 peptide of SEQ ID NO: 88, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%)), and at least a portion of the RgpB catalytic domain is a RAS peptide (e.g., an extended Ras2 peptide of SEQ ID NO: 252, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%)). In other embodiments, at least a portion of the Kgp catalytic domain is a KAS peptide (for example, the extended Kas2 peptide of SEQ ID NO: 88, or a sequence having at least 70% identity thereto (for example, at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%)), and at least a portion of the RgpB catalytic domain is a shortened RgpB catalytic domain. In other embodiments, at least a portion of the Kgp catalytic domain is a shortened Kgp catalytic domain, and at least a portion of the RgpB catalytic domain is a RAS peptide (for example, the extended Ras2 peptide of SEQ ID NO: 252, or a sequence having at least 70% identity thereto (for example, at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%)).

[0274] In other embodiments, at least a portion of the Kgp catalytic domain is the full-length catalytic domain and at least a portion of the RgpB catalytic domain is a truncated RgpB catalytic domain. For example, in certain such embodiments, at least a portion of the Kgp catalytic domain is the full-length catalytic domain and at least a portion of the RgpB catalytic domain is a RAS peptide (e.g., the extended Ras2 peptide of SEQ ID NO: 252, or a sequence having at least 70% (e.g., at least 75, 80, 85, 90 or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99%) identity thereto).

[0275] In other embodiments, at least a portion of the Kgp catalytic domain is a truncated Kgp catalytic domain and at least a portion of the RgpB catalytic domain is the full-length RgpB catalytic domain. For example, in certain such embodiments, at least a portion of the Kgp catalytic domain is a KAS peptide (e.g., the extended Kas2 peptide of SEQ ID NO: 88, or a sequence having at least 70% (e.g., at least 75, 80, 85, 90 or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99%) identity thereto), and at least a portion of the RgpB catalytic domain is the full-length RgpB catalytic domain.

[0276] K1 and K2 adhesin domains The nucleic acids and polypeptides of the present invention include at least a portion of the Kgp K1 adhesin domain and / or at least a portion of the RgpA K1 adhesin domain, as disclosed herein. In some embodiments, the nucleic acids and polypeptides further include at least a portion of the Kgp K2 adhesin domain and / or at least a portion of the RgpA K1 adhesin domain.

[0277] The modular nature of gingipain is such that any of the K1 adhesin domains disclosed herein can be combined with any of the other domains disclosed herein. For example, any of the sequences disclosed above, including the ABM, DUF, or Cat domain, may be combined with any of the K1 adhesin domains described in the following paragraphs. Similarly, any of the K2 adhesin domains disclosed herein may be combined with any of the other domains disclosed herein. For example, any of the sequences disclosed above, including the ABM, DUF, or Cat domain, may be combined with any of the K2 adhesin domains described in the following paragraphs. Similarly, any of the K1 adhesin domains disclosed herein may be combined with any of the other K2 adhesin domains disclosed herein.

[0278] Wild-type Kgp and RgpA contain three adhesin domains known as K1, K2, and K3. While the degree of sequence identity between the three domains is relatively low (for example, approximately 40% sequence identity exists between K1 and K2), they are structurally homologous to each other, possessing conserved sequence motifs. However, very high sequence identity exists between the Kgp K1 adhesin domain and the RgpA K1 adhesin domain, and between the Kgp K2 adhesin domain and the RgpA K2 adhesin domain.

[0279] The three domains are members of the cleaved adhesin domain family, designated IPR011628 in the InterPro database. The cleaved adhesin domains of Kgp and RgpA have several functions in P. gingivalis, including adhesion to host tissues and colonization of host tissues, and are thought to promote co-aggregation of P. gingivalis with other oral pathogens and subsequent biofilm formation (Li and Collyer, 2011; Dashper et al., 2017). For example, cleaved adhesin domains have been reported to bind to hemoglobin, human serum albumin, and fibrinogen (Li et al., 2011; Ganuelas et al., 2013). These proteins are abundant in the blood and are most likely to be targeted by P. gingivalis during initial colonization. In addition, it has been shown that cleaved adhesin domains induce hemolysis of red blood cells in vitro, thereby enabling P. gingivalis to acquire essential heme from red blood cells (Li et al., 2011; Ganuelas et al., 2013).

[0280] The nucleic acids and polypeptides of the present invention include at least a portion of the Kgp K1 adhesin domain and / or at least a portion of the RgpA K1 adhesin domain, with the aim of inducing an antibody response that inhibits the adhesion function of Kgp and RgpA.

[0281] At least a portion of the Kgp or RgpA K1 adhesin domain according to the present invention is derived from P. gingivalis Kgp or RgpA.

[0282] In some embodiments, at least a portion of the Kgp K1 adhesin domain is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, the Kgp K1 adhesin domain has the sequence of Sequence ID No. 182. Therefore, in some embodiments, at least a portion of the Kgp K1 adhesin domain contains at least a portion of Sequence ID No. 182, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the KGP K1 adhesin domain contains the sequence of sequence number 182, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0283] In some embodiments, at least a portion of the RgpA K1 adhesin domain is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, the RgpA K1 adhesin domain has the sequence of Sequence ID No. 185. Therefore, in some embodiments, at least a portion of the RgpA K1 adhesin domain includes at least a portion of Sequence ID No. 185, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the RgpA K1 adhesin domain contains the sequence of sequence number 185, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0284] Examples of K1 adhesin domain sequences that can be used in accordance with the present invention are provided in Table 14.

[0285] [Table 16]

[0286] In some embodiments, at least a portion of the Kgp K1 adhesin domain includes at least a portion of sequence number 183, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the Kgp K1 adhesin domain includes the sequence of sequence number 183, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0287] In some embodiments, at least a portion of the Kgp K1 adhesin domain includes at least a portion of sequence number 184, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the Kgp K1 adhesin domain includes the sequence of sequence number 184, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0288] In some embodiments, at least a portion of the RgpA K1 adhesin domain includes at least a portion of sequence number 186, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the Kgp K1 adhesin domain includes the sequence of sequence number 186, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0289] The inventors have found that nucleic acids and polypeptides containing a full-length K1 adhesin domain may be particularly advantageous in inducing an immune response. This may be because polypeptides containing a full-length K1 adhesin domain can fold into a three-dimensional structure similar to the three-dimensional structure of a wild-type K1 adhesin domain, thereby enabling the construct to induce the production of antibodies that recognize the conformational epitope within the K1 adhesin domain. Therefore, in some embodiments, at least a portion of the K1 adhesin domain contains a full-length K1 adhesin domain. In some embodiments, at least a portion of the RgpA K1 adhesin domain contains a full-length RgpA K1 adhesin domain.

[0290] A full-length Kgp or RgpA K1 adhesin domain refers to a Kgp or RgpA K1 adhesin domain that is not shortened relative to the corresponding wild-type K1 adhesin domain. Thus, in some embodiments, at least a portion of the Kgp K1 adhesin domain is a full-length Kgp K1 adhesin domain that is the same length as the corresponding wild-type Kgp K1 adhesin domain. In some embodiments, at least a portion of the RgpA K1 adhesin domain is a full-length RgpA K1 adhesin domain that is the same length as the corresponding wild-type RgpA K1 adhesin domain.

[0291] For example, the Kgp K1 adhesin domain present in Kgp from P. gingivalis W50 strain is 169 amino acids long. Therefore, in some embodiments, the full-length Kgp K1 adhesin domain is at least 169 amino acids long (e.g., 169 amino acids long). The K1 adhesin domain present in RgpA from P. gingivalis W50 strain is 170 amino acids long. Therefore, in some embodiments, the full-length RgpA K1 adhesin domain is at least 170 amino acids long (e.g., 170 amino acids long).

[0292] The Kgp K1 domain present in Kgp from other P. gingivalis strains may have a different length than the Kgp K1 domain present in Kgp from P. gingivalis strain W50. Similarly, the RgpA K1 domain present in Kgp from other P. gingivalis strains may have a different length than the RgpA K1 domain present in RgpA from P. gingivalis strain W50.

[0293] Therefore, in some embodiments, the full-length Kgp K1 adhesin domain is at least 165 amino acids long (e.g., 165 amino acids long). In some embodiments, the full-length Kgp K1 adhesin domain is at least 166 amino acids long (e.g., 166 amino acids long). In some embodiments, the full-length Kgp K1 adhesin domain is at least 167 amino acids long (e.g., 167 amino acids long). In some embodiments, the full-length Kgp K1 adhesin domain is at least 168 amino acids long (e.g., 168 amino acids long). In some embodiments, the full-length Kgp K1 adhesin is at least 170 amino acids long (e.g., 170 amino acids long).

[0294] In some embodiments, the full-length RgpA K1 adhesin domain is at least 166 amino acids long (e.g., 166 amino acids). In some embodiments, the full-length RgpA K1 adhesin domain is at least 167 amino acids long (e.g., 167 amino acids). In some embodiments, the full-length RgpA K1 adhesin domain is at least 168 amino acids long (e.g., 168 amino acids). In some embodiments, the full-length RgpA K1 adhesin domain is at least 169 amino acids long (e.g., 169 amino acids). Shortening of the Kgp or RgpA K1 adhesin domain may also be done without significantly altering the properties of the resulting polypeptide, for example, so that the resulting polypeptide can still induce antibodies that block Kgp and / or RgpA adhesion. Therefore, in some embodiments, at least a portion of the Kgp K1 adhesin domain is a shortened Kgp K1 adhesin domain. In some embodiments, at least a portion of the RgpA K1 adhesin domain is a shortened RgpA K1 adhesin domain.

[0295] Therefore, in some embodiments, at least a portion of the Kgp or RgpA K1 adhesin domain is a shortened Kgp or RgpA K1 adhesin domain, in which 1 to 35 amino acids are shortened. In certain such embodiments, the Kgp or RgpA K1 adhesin domain is shortened by 1 to 30 amino acids, 1 to 25 amino acids, 1 to 20 amino acids, 1 to 15 amino acids, 1 to 10 amino acids, and 1 to 5 amino acids. In certain such embodiments, the Kgp K1 adhesin domain is shortened by 1 to 30 amino acids, 1 to 25 amino acids, 1 to 20 amino acids, 1 to 15 amino acids, 1 to 10 amino acids, and 1 to 5 amino acids. In certain such embodiments, the RgpA K1 adhesin domain is shortened by 1 to 30 amino acids, 1 to 25 amino acids, 1 to 20 amino acids, 1 to 15 amino acids, 1 to 10 amino acids, or 1 to 5 amino acids. In some embodiments, the Kgp or RgpA K1 adhesin domain is shortened by 30 amino acids. In some embodiments, the Kgp or RgpA K1 adhesin domain is shortened by 25 amino acids. In some embodiments, the Kgp or RgpA K1 adhesin domain is shortened by 20 amino acids. In some embodiments, the Kgp or RgpA K1 adhesin domain is shortened by 15 amino acids. In some embodiments, the Kgp or RgpA K1 adhesin domain is shortened by 10 amino acids. In some embodiments, the Kgp or RgpA K1 adhesin domain is shortened by 5 amino acids.

[0296] In some embodiments, the shortened Kgp K1 adhesin domain includes the sequence GTTTLSESF (sequence number 191).

[0297] In some embodiments, the shortened RgpA K1 adhesin domain includes the sequence GTTTLSESF (sequence number 192).

[0298] As can be seen in Figure 2, in wild-type Kgp and RgpA sequences, the C-terminal portion of ABM3 overlaps with the N-terminal portion of the K1 adheren domain. Therefore, in some embodiments, the portion of Kgp containing ABM3 and at least a portion of the Kgp K1 adheren domain overlap (for example, the eight C-terminal residues of the portion of Kgp containing ABM3 are also the eight N-terminal residues of the Kgp K1 adheren domain). In some embodiments, the portion of RgpA containing ABM3 and at least a portion of the RgpA K1 adheren domain overlap (for example, the nine C-terminal residues of the portion of Kgp containing ABM3 are also the nine N-terminal residues of the RgpA K1 adheren domain).

[0299] Examples of sequences that can be used in accordance with the present invention, in which the Kgp or RgpA moiety containing ABM3 and the Kgp or RgpA K1 adhesin domain overlap, are provided in Table 15.

[0300] [Table 17]

[0301] Therefore, in some embodiments, both the Kgp portion containing ABM3 and the Kgp K1 adhesin domain contain the sequence of Sequence ID No. 93, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0302] In some embodiments, both the Kgp portion containing ABM3 and the Kgp K1 adhesin domain include the sequence of Sequence ID No. 94, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0303] In some embodiments, both the Kgp portion containing ABM3 and the Kgp K1 adhesin domain include the sequence of Sequence ID No. 103, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0304] Variants of the Kgp or RgpA K1 adhesin domain may also be used in the present invention. Such variants may be used to remove glycosylation sites, as described elsewhere in this specification.

[0305] The nucleic acids and polypeptides of the present invention may contain at least a portion of additional cleaved adhesin domains in addition to at least a portion of the Kgp and / or RgpA K1 adhesin domains. The inclusion of at least a portion of the K2 adhesin domains may enable the resulting polypeptide to form a structure more closely resembling the wild-type gingipain structure, thereby providing an additional three-dimensional epitope that may be useful in eliciting an immune response. Alternatively, the inclusion of at least a portion of the K2 adhesin domains may mean that the polypeptide can induce antibodies specific to the K2 domain that can inhibit the K2-specific function of Kgp and / or RgpA. Therefore, in some embodiments, the nucleic acids and polypeptides of the present invention may contain at least a portion of the Kgp K2 adhesin domains. In some embodiments, the nucleic acids and polypeptides of the present invention may contain at least a portion of the RgpA K2 adhesin domains.

[0306] In some embodiments, the Kgp or RgpA K2 adhesin domain is full length. For example, the Kgp K2 adhesin domain present in Kgp from P. gingivalis W50 strain is 172 amino acids long. Therefore, in some embodiments, the full-length Kgp K2 adhesin domain is at least 172 amino acids long (e.g., 172 amino acids long). For example, the RgpA K2 adhesin domain present in RgpA from P. gingivalis W50 strain is 172 amino acids long. Therefore, in some embodiments, the full-length RgpA K2 adhesin domain is at least 172 amino acids long (e.g., 172 amino acids long).

[0307] The Kgp K2 domain present in Kgp from other P. gingivalis strains may have a different length than the Kgp K2 domain present in Kgp from P. gingivalis strain W50. Similarly, the RgpA K2 domain present in Kgp from other P. gingivalis strains may have a different length than the RgpA K2 domain present in RgpA from P. gingivalis strain W50.

[0308] Therefore, in some embodiments, the full-length Kgp or RgpA K2 adhesin domain is at least 168 amino acids long (e.g., 168 amino acids). In some embodiments, the full-length Kgp or RgpA K2 adhesin domain is at least 169 amino acids long (e.g., 169 amino acids). In some embodiments, the full-length Kgp or RgpA K2 adhesin domain is at least 170 amino acids long (e.g., 170 amino acids). In some embodiments, the full-length Kgp or RgpA K2 adhesin domain is at least 171 amino acids long (e.g., 171 amino acids). In some embodiments, the full-length Kgp or RgpA K2 adhesin domain is at least 171 amino acids long (e.g., 171 amino acids). In some embodiments, the full-length Kgp or RgpA K2 adhesin domain is at least 176 amino acids long (e.g., 176 amino acids). In some embodiments, the full-length Kgp or RgpA K2 adhesin domain is at least 178 amino acids long (e.g., 178 amino acids).

[0309] The shortening of the Kgp or RgpA K2 adhesin domain may also be carried out without significantly altering the properties of the resulting polypeptide, for example, so that the resulting polypeptide can still induce antibodies that block Kgp and / or RgpA adhesion. Therefore, in some embodiments, at least a portion of the Kgp K2 adhesin domain is a shortened Kgp K2 adhesin domain. In some embodiments, at least a portion of the RgpA K2 adhesin domain is a shortened RgpA K2 adhesin domain.

[0310] Therefore, in some embodiments, at least a portion of the Kgp or RgpA K2 adhesin domain is a shortened Kgp or RgpA K2 adhesin domain, in which 1 to 35 amino acids are shortened. In certain such embodiments, the Kgp K2 adhesin domain is shortened by 1 to 30 amino acids, 1 to 25 amino acids, 1 to 20 amino acids, 1 to 15 amino acids, 1 to 10 amino acids, and 1 to 5 amino acids. In certain such embodiments, the RgpA K2 adhesin domain is shortened by 1 to 30 amino acids, 1 to 25 amino acids, 1 to 20 amino acids, 1 to 15 amino acids, 1 to 10 amino acids, and 1 to 5 amino acids. In some embodiments, the Kgp or RgpA K2 adhesin domain is shortened by 30 amino acids. In some embodiments, the Kgp or RgpA K2 adhesin domain is shortened by 25 amino acids. In some embodiments, the Kgp or RgpA K2 adhesin domain is shortened by 20 amino acids. In some embodiments, the Kgp or RgpA K2 adhesin domain is shortened by 15 amino acids. In some embodiments, the Kgp or RgpA K2 adhesin domain is shortened by 10 amino acids. In some embodiments, the Kgp or RgpA K2 adhesin domain is shortened by 5 amino acids.

[0311] In some embodiments, at least a portion of the Kgp or RgpA K2 adhesin domain is derived from the P. gingivalis W50 strain. In the P. gingivalis W50 strain, the Kgp K2 adhesin domain has the sequence of Sequence ID No. 187. Therefore, in some embodiments, at least a portion of the Kgp K2 adhesin domain includes at least a portion of Sequence ID No. 187, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto. For example, at least a portion of the Kgp K2 adhesin domain includes the sequence of Sequence ID No. 187, or a sequence having at least 70% (e.g., at least 75, 80, 85, 90 or 95%; or, for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99%) identity thereto. In the P. gingivalis W50 strain, the RgpA K2 adheren domain has the sequence of Sequence ID No. 189. Therefore, in some embodiments, at least a portion of the RgpA K2 adheren domain contains at least a portion of Sequence ID No. 189, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the RgpA K2 adhesin domain contains the sequence of sequence number 189, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0312] Examples of K2 adhesin domain sequences that can be used in accordance with the present invention are provided in Table 14.

[0313] [Table 18]

[0314] In some embodiments, at least a portion of the Kgp K2 adhesin domain includes at least a portion of sequence number 96, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the Kgp K2 adhesin domain includes the sequence of sequence number 96, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0315] In some embodiments, at least a portion of the Kgp K2 adhesin domain includes at least a portion of sequence number 188, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the Kgp K2 adhesin domain includes the sequence of sequence number 188, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0316] In some embodiments, at least a portion of the RgpA K2 adhesin domain includes at least a portion of sequence number 104, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the RgpA K2 adhesin domain includes the sequence of sequence number 104, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0317] In some embodiments, at least a portion of the RgpA K2 adhesin domain includes at least a portion of sequence number 190, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). For example, at least a portion of the RgpA K2 adhesin domain includes the sequence of sequence number 190, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0318] The inventors have found that nucleic acids and polypeptides containing a full-length K2 adhesin domain may be particularly advantageous in inducing an immune response. This may be because polypeptides containing a full-length K2 adhesin domain can fold into a three-dimensional structure similar to the three-dimensional structure of a wild-type K2 adhesin domain, thereby enabling the construct to induce the production of antibodies that recognize the conformational epitope within the K2 adhesin domain. Therefore, in some embodiments, at least a portion of the K2 adhesin domain contains a full-length K2 adhesin domain. In some embodiments, at least a portion of the RgpA K2 adhesin domain contains a full-length RgpA K2 adhesin domain.

[0319] A full-length Kgp or RgpA K2 adhesin domain refers to a Kgp or RgpA K2 adhesin domain that is not shortened relative to the corresponding wild-type K2 adhesin domain. Thus, in some embodiments, at least a portion of the Kgp K2 adhesin domain is a full-length Kgp K2 adhesin domain that is the same length as the corresponding wild-type Kgp K2 adhesin domain. In some embodiments, at least a portion of the RgpA K2 adhesin domain is a full-length RgpA K2 adhesin domain that is the same length as the corresponding wild-type RgpA K2 adhesin domain.

[0320] For example, the Kgp K2 adhesin domain present in Kgp from P. gingivalis W50 strain is 172 amino acids long. Therefore, in some embodiments, the full-length Kgp K2 adhesin domain is at least 172 amino acids long (e.g., 172 amino acids long). The K2 adhesin domain present in RgpA from P. gingivalis W50 strain is 172 amino acids long. Therefore, in some embodiments, the full-length RgpA K2 adhesin domain is at least 172 amino acids long (e.g., 172 amino acids long).

[0321] Variants of the Kgp or RgpA K1 adhesin domain may also be used in the present invention. Such variants may be used to remove glycosylation sites, as described elsewhere in this specification.

[0322] In some embodiments, the nucleic acid or polypeptide comprises at least a portion of the Kgp K1 adhesin domain and at least a portion of the Kgp K2 adhesin domain. In certain such embodiments, at least a portion of the Kgp K1 adhesin domain and at least a portion of the Kgp K2 adhesin domain each comprises a sequence as shown in Table 17, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0323] [Table 19]

[0324] In some embodiments, the nucleic acid or polypeptide comprises at least a portion of the Kgp K1 adhesin domain, at least a portion of the Kgp K2 adhesin domain, and at least a portion of the RgpA K1 adhesin domain. In certain such embodiments, at least a portion of the Kgp K1 adhesin domain and at least a portion of the Kgp K2 adhesin domain each comprises the sequence shown in Table 17, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%), and the RgpA K1 adhesin domain comprises the sequence of Sequence ID No. 103.

[0325] In some embodiments, the nucleic acid or polypeptide comprises at least a portion of the RgpA K1 adhesin domain and at least a portion of the RgpA K2 adhesin domain. In certain embodiments, at least a portion of the RgpA K1 adhesin domain and at least a portion of the RgpA K2 adhesin domain each comprises a sequence as shown in Table 18, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0326] [Table 20]

[0327] Positioning of gingipain domains and sequence motifs The domains described in any of the preceding sections (i.e., a first Kgp moiety containing ABM1, a first moiety containing ABM2, a second moiety containing ABM1, a second moiety containing ABM2, at least a portion of DUF2436, at least a portion of the Kgp catalytic domain, and at least a portion of the Kgp K1 adhesin domain) can be combined to produce a nucleic acid encoding the polypeptide of the present invention.

[0328] The modular nature of the Kgp structure means that the different gingipain domains of the nucleic acid and polypeptide may be arranged in any order. However, typically, the gingipain domains are arranged in the same order as they are in wild-type Kgp and RgpA proteins. Arranging the domains in the same order as wild-type Kgp and RgpA proteins enhances polypeptide folding in a manner more closely similar to wild-type Kgp and RgpA, thereby potentially allowing, for example, the preservation of conformational epitopes.

[0329] As can be seen in Figure 1, the domain order in wild-type Kgp and RgpA is, from N-terminus to C-terminus: propeptide; catalytic domain; first portion containing ABM1; DUF2436; first portion containing ABM2; second portion containing ABM1; portion containing ABM3; K1; K2; second portion containing ABM2; K3; C-terminal domain. For embodiments that include subsets of these domains, the present domains are positioned in the same N-terminus to C-terminus order, with the omission of any of the domains from the wild-type sequence.

[0330] For example, in some embodiments, the polypeptide or nucleic acid encoding the polypeptide of the present invention comprises i) at least a portion of the Kgp catalytic domain; ii) at least a portion of Kgp DUF2436; iii) at least a portion of the Kgp K1 adhesin domain; iv) a first Kgp portion containing ABM1 and a first Kgp portion containing ABM2; v) a second Kgp portion containing ABM1 and a second Kgp portion containing ABM2; and vi) a Kgp portion containing ABM3; the domains are positioned from the N-terminus to the C-terminus of the polypeptide in the following order: Kgp catalytic domain; first Kgp portion containing ABM1; Kgp DUF2436; first Kgp portion containing ABM2; second Kgp portion containing ABM1; Kgp portion containing ABM3; Kgp K1; second Kgp portion containing ABM2.

[0331] In some embodiments, the polypeptide or nucleic acid encoding the polypeptide of the present invention comprises i) at least a portion of the RgpA catalytic domain; ii) at least a portion of RgpA DUF2436; iii) at least a portion of the RgpA K1 adheren domain; iv) a first RgpA portion containing ABM1 and a first RgpA portion containing ABM2; v) a second RgpA portion containing ABM1 and a second RgpA portion containing ABM2; vi) an RgpA portion containing ABM3; vii) at least a portion of the RgpA K2 adheren domain; the domains are in the following order from the N-terminus to the C-terminus of the polypeptide: RgpA catalytic domain; first RgpA portion containing ABM1; RgpA DUF2436; first RgpA portion containing ABM2; second RgpA portion containing ABM1; RgpA portion containing ABM3; RgpA K1; RgpA It is positioned in the second RgpA portion, which includes K2;ABM2.

[0332] In some embodiments, the polypeptide or nucleic acid encoding the polypeptide of the present invention comprises i) at least a portion of the RgpA catalytic domain; ii) at least a portion of RgpA DUF2436; iii) at least a portion of the RgpA K1 adheren domain; iv) a first RgpA portion containing ABM1 and a first RgpA portion containing ABM2; v) a second RgpA portion containing ABM1 and a second RgpA portion containing ABM2; vi) an RgpA portion containing ABM3; vii) at least a portion of the RgpA K2 adheren domain; the domains are in the following order from the N-terminus to the C-terminus of the polypeptide: first RgpA portion containing ABM1; RgpA DUF2436; first RgpA portion containing ABM2; second RgpA portion containing ABM1; RgpA portion containing ABM3; RgpA K1; RgpA K2; a second RgpA moiety containing ABM2; located in the RgpA catalytic domain.

[0333] In some embodiments, the polypeptide or nucleic acid encoding the polypeptide of the present invention comprises i) at least a portion of the Kgp catalytic domain; ii) at least a portion of Kgp DUF2436; iii) at least a portion of the Kgp K1 adhesin domain; iv) a first Kgp portion containing ABM1 and a first Kgp portion containing ABM2; v) a second Kgp portion containing ABM1 and a second Kgp portion containing ABM2; vi) a Kgp portion containing ABM3; and vii) at least a portion of the RgpA catalytic domain, the domains positioned from the N-terminus to the C-terminus of the polypeptide in the following order: Kgp catalytic domain; first Kgp portion containing ABM1; Kgp DUF2436; first Kgp portion containing ABM2; second Kgp portion containing ABM1; Kgp portion containing ABM3; Kgp K1; second Kgp portion containing ABM2; and RgpA catalytic domain.

[0334] In some embodiments, the polypeptide or nucleic acid encoding the polypeptide of the present invention comprises i) at least a portion of the Kgp catalytic domain; ii) at least a portion of Kgp DUF2436; iii) at least a portion of the Kgp K1 adheren domain; iv) a first Kgp portion containing ABM1 and a first Kgp portion containing ABM2; v) a second Kgp portion containing ABM1 and a second Kgp portion containing ABM2; vi) a Kgp portion containing ABM3; vii) at least a portion of the Kgp K2 adheren domain; and viiii) at least a portion of the RgpA catalytic domain, where the domains are arranged in the following order from the N-terminus to the C-terminus of the polypeptide: Kgp catalytic domain; first Kgp portion containing ABM1; Kgp DUF2436; first Kgp portion containing ABM2; second Kgp portion containing ABM1; Kgp portion containing ABM3; Kgp K1; Kgp K2; a second Kgp moiety containing ABM2; located in the RgpA catalytic domain.

[0335] In some embodiments, the polypeptide or nucleic acid encoding the polypeptide of the present invention comprises i) at least a portion of the Kgp catalytic domain; ii) at least a portion of Kgp DUF2436; iii) at least a portion of the Kgp K1 adheren domain; iv) a first Kgp portion containing ABM1 and a first Kgp portion containing ABM2; v) a second Kgp portion containing ABM1 and a second Kgp portion containing ABM2; vi) a Kgp portion containing ABM3; vii) at least a portion of the RgpA K2 adheren domain; and viiii) at least a portion of the RgpA catalytic domain, where the domains are arranged in the following order from the N-terminus to the C-terminus of the polypeptide: Kgp catalytic domain; first Kgp portion containing ABM1; Kgp DUF2436; first Kgp portion containing ABM2; second Kgp portion containing ABM1; Kgp portion containing ABM3; Kgp K1; RgpA K2; a second Kgp moiety containing ABM2; located in the RgpA catalytic domain.

[0336] In some embodiments, the polypeptide or nucleic acid encoding the polypeptide of the present invention comprises i) at least a portion of the Kgp catalytic domain; ii) at least a portion of Kgp DUF2436; iii) at least a portion of the Kgp K1 adheren domain; iv) a first Kgp portion containing ABM1 and a first Kgp portion containing ABM2; v) a second Kgp portion containing ABM1 and a second Kgp portion containing ABM2; vi) a Kgp portion containing ABM3; vii) at least a portion of the Kgp K2 adheren domain; and viiii) at least a portion of the RgpA catalytic domain; ix) a first RgpA portion containing ABM1 and a first RgpA portion containing ABM2; and x) at least a portion of RgpA DUF2436, where the domains are arranged in the following order from the N-terminus to the C-terminus of the polypeptide: Kgp catalytic domain; first Kgp portion containing ABM1; Kgp DUF2436; First Kgp moiety containing ABM2; Second Kgp moiety containing ABM1; Kgp moiety containing ABM3; Kgp K1; Kgp K2; Second Kgp moiety containing ABM2; First RgpA moiety containing ABM1; RgpA DUF2436; First RgpA moiety containing ABM2; Positioned in the RgpA catalytic domain.

[0337] In some embodiments, the polypeptide or nucleic acid encoding the polypeptide of the present invention comprises i) at least a portion of the Kgp catalytic domain; ii) at least a portion of Kgp DUF2436; iii) at least a portion of the Kgp K1 adheren domain; iv) a first Kgp portion containing ABM1 and a first Kgp portion containing ABM2; v) a second Kgp portion containing ABM1 and a second Kgp portion containing ABM2; vi) a Kgp portion containing ABM3; vii) at least a portion of the RgpA K2 adheren domain; and viiii) at least a portion of the RgpA catalytic domain; ix) a first RgpA portion containing ABM1 and a first RgpA portion containing ABM2; and x) at least a portion of RgpA DUF2436, where the domains are arranged in the following order from the N-terminus to the C-terminus of the polypeptide: Kgp catalytic domain; first Kgp portion containing ABM1; Kgp DUF2436; First Kgp moiety containing ABM2; Second Kgp moiety containing ABM1; Kgp moiety containing ABM3; Kgp K1; RgpA K2; Second Kgp moiety containing ABM2; First RgpA moiety containing ABM1; RgpA DUF2436; First RgpA moiety containing ABM2; Positioned in the RgpA catalytic domain.

[0338] Examples of nucleic acids or polypeptides encoding the polypeptide of the present invention are provided in Table 19. Other examples of nucleic acids or polypeptides encoding the polypeptide of the present invention are sequences provided in Table 19, with the N-terminal methionine deleted. This table also provides nucleic acid sequences encoding polypeptides, which also form part of the present invention.

[0339] [Table 21]

[0340] [Table 22]

[0341] [Table 23]

[0342] [Table 24]

[0343] [Table 25]

[0344] [Table 26]

[0345] In some embodiments, the polypeptide includes the sequence of SEQ ID NO: 1, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0346] In some embodiments, the polypeptide includes the sequence according to SEQ ID NO: 6, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0347] In some embodiments, the polypeptide includes the sequence of sequence number 11, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0348] In some embodiments, the polypeptide includes the sequence of SEQ ID NO: 16, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0349] In some embodiments, the polypeptide includes the sequence of sequence number 367, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). In some embodiments, the polypeptide includes the sequence of sequence number 371, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0350] In some embodiments, the polypeptide includes the sequence of sequence number 375, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0351] In some embodiments, the polypeptide includes the sequence of sequence number 379, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0352] In some embodiments, the polypeptide includes the sequence of sequence number 383, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0353] In some embodiments, the polypeptide includes the sequence of sequence number 397, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0354] In some embodiments, the polypeptide includes the sequence of sequence number 403, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0355] In some embodiments, the polypeptide includes the sequence of sequence number 415, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0356] In some embodiments, the polypeptide includes the sequence of sequence number 409, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0357] The N-terminal methionine in any of the above embodiments may be omitted from the polypeptide. Thus, in some embodiments, the polypeptide includes a sequence that is at least 70% (e.g., SEQ ID NO: 1 without the N-terminal methionine) identical thereto (for example, at least 75, 80, 85, 90, or 95%; or for example, at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0358] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 259 (i.e., SEQ ID NO: 6 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0359] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 264 (i.e., SEQ ID NO: 11 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0360] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 269 (i.e., SEQ ID NO: 16 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0361] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 368 (i.e., SEQ ID NO: 367 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0362] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 372 (i.e., SEQ ID NO: 371 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0363] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 376 (i.e., SEQ ID NO: 375 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0364] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 380 (i.e., SEQ ID NO: 379 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0365] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 384 (i.e., SEQ ID NO: 383 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0366] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 398 (i.e., SEQ ID NO: 397 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0367] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 404 (i.e., SEQ ID NO: 403 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0368] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 416 (i.e., SEQ ID NO: 415 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0369] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 410 (i.e., SEQ ID NO: 409 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0370] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 21.

[0371] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 22.

[0372] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 31.

[0373] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 32.

[0374] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence specified by SEQ ID NO: 41.

[0375] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence specified by SEQ ID NO: 42.

[0376] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 51.

[0377] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 52.

[0378] In certain embodiments, the amino acid sequence of the polypeptide of the present invention is encoded by a codon-optimized polynucleotide sequence.

[0379] Polypeptide variants The polypeptide sequences described herein may include one or more mutations or modifications.

[0380] In some embodiments, the polypeptide described herein may contain one or more conservative amino acid substitutions.

[0381] Mutation of glycosylation sites Glycosylation can occur in eukaryotic cells but not in prokaryotic cells. As used herein, “glycosylation” refers to the addition of a sugar unit to a protein. In particular, N-linked glycosylation is the attachment of a glycan to the amide nitrogen of an asparagine (Asn;N) residue in a protein. The process of attachment results in a glycosylated protein. This glycan may be a polysaccharide. Glycosylation can occur at any asparagine residue in a protein that is accessible by a glycosylation enzyme after the translation of the protein and is recognized by the glycosylation enzyme, most commonly at an accessible asparagine that is part of the NXS / T motif, where the second amino acid residue following asparagine is serine or threonine. Non-human glycosylation patterns may confer undesirable reactogenicity to polypeptides when used to induce antibodies. In addition, glycosylation of polypeptides that are not normally glycosylated (such as the polypeptides described herein) may alter their immunogenicity. For example, glycosylation can mask important immunogenic epitopes within a protein. Therefore, to reduce or eliminate glycosylation, either an asparagine residue or a serine / threonine residue can be modified, for example, by substitution with another amino acid.

[0382] In certain embodiments, the polypeptide as described herein includes at least one mutated glycosylation site, e.g., at least one mutated N-linked glycosylation site and / or at least one O-linked glycosylation site. In some embodiments, one or more (e.g., all) N-glycosylation sites in the polypeptide as described herein are removed. Removal of N-glycosylation sites can reduce the glycosylation of the polypeptide. In some embodiments, the polypeptide as described herein has reduced glycosylation compared to the corresponding wild-type polypeptide. Reduced glycosylation compared to the corresponding wild-type polypeptide may be observed in one or all of the polypeptide's domains. For example, at least a portion of the Kgp catalytic domain may have reduced glycosylation compared to the corresponding wild-type portion of the Kgp catalytic domain. In certain embodiments, all domains of the polypeptide have reduced glycosylation compared to the corresponding wild-type domain. Removal of N-glycosylation sites can eliminate N-glycosylation of the polypeptide.

[0383] In certain embodiments, the modification involves the substitution of one or more (e.g., all) N, S, and T amino acids in the NXS / T sequence motif, where X corresponds to any amino acid. In some embodiments, the N, S, or T amino acids are substituted with conservative amino acid substitutions.

[0384] Exemplary mutated glycosylation sites within Kgp or RgpA that can be mutated are shown in Tables 20 and 21 below. Accordingly, in any of the nucleic acids or polypeptides of the present invention disclosed herein, one or more (e.g., all) mutation sites within Kgp correspond to the positions of the wild-type sequence of SEQ ID NO: 157 as specified in Table 20. In any of the nucleic acids or polypeptides of the present invention disclosed herein, one or more (e.g., all) mutation sites within RgpA correspond to one or more (e.g., all) positions of the wild-type sequence of SEQ ID NO: 158 as specified in Table 21.

[0385] [Table 27]

[0386] [Table 28]

[0387] In some embodiments, the polypeptides described herein include one or more (e.g., all) of the mutations shown in Table 20. In some embodiments, the polypeptides described herein include one or more (e.g., all) of the mutations shown in Table 21. In some embodiments, the polypeptides described herein include one or more (e.g., all) of the mutations shown in Table 20 and one or more (e.g., all) of the mutations shown in Table 21.

[0388] In some embodiments, the Kgp-based polypeptides described herein include a single amino acid substitution at one or more (e.g., all) positions corresponding to the N-glycosylation site in the natural P. gingivalis Kgp polypeptide (e.g., SEQ ID NO: 157). In some embodiments, the Kgp-based polypeptides described herein include one or more (e.g., all) amino acid substitutions at positions corresponding to 284, 442, 574, 645, 691, 950, 968, 1089, 1316, and 1390 of SEQ ID NO: 157. In some embodiments, the Kgp-based polypeptides described herein include one or more (e.g., all) amino acid substitutions at positions corresponding to 284, 442, 574, 645, 691, 950, 968, and 1089 of SEQ ID NO: 157. In some embodiments, the KGP-based polypeptides described herein include one or more (e.g., all) amino acid substitutions at positions corresponding to 284, 442, 574, 645, 691, 950, and 968 of SEQ ID NO: 157. In some embodiments, the KGP-based polypeptides described herein include one or more (e.g., all) amino acid substitutions at positions corresponding to 442, 691, 950, 968, and 1089 of SEQ ID NO: 157. In some embodiments, the KGP-based polypeptides described herein include one or more (e.g., all) amino acid substitutions at positions corresponding to 442, 691, 950, and 968 of SEQ ID NO: 157.

[0389] In some embodiments, the RgpA-based polypeptides described herein include a single amino acid substitution at one or more (e.g., all) positions corresponding to the N-glycosylation site in the natural P. gingivalis RgpA polypeptide (e.g., SEQ ID NO: 158). In some embodiments, the RgpA-based polypeptides described herein include one or more (e.g., all) amino acid substitutions at positions corresponding to 363, 434 and 436, 508, 592, 597, 623, 629, 635, 671, 691, 766, 931, 947, 1298, and 1372 in SEQ ID NO: 158. In some embodiments, the RgpA-based polypeptides described herein include one or more (e.g., all) amino acid substitutions at positions corresponding to 363, 436, 508, 592, 597, 623, 629, 635, 671, 691, 766, 931, 947, 1298, and 1372 of SEQ ID NO: 158. In some embodiments, the RgpA-based polypeptides described herein include one or more (e.g., all) amino acid substitutions at positions 434, 671, 69, 766, 931, 947, 1298, and 1372 of SEQ ID NO: 158.

[0390] In some embodiments, the Kgp and RgpA-based polypeptides described herein include a single amino acid substitution at one or more (e.g., all) positions corresponding to the N-glycosylation sites in the natural P. gingivalis (P. gingivalis) Kgp polypeptide (e.g., SEQ ID NO: 157) and the natural P. gingivalis (P. gingivalis) RgpA polypeptide (e.g., SEQ ID NO: 158). In some embodiments, the Kgp and RgpA-based polypeptides described herein include a single amino acid substitution at one or more (e.g., all) positions corresponding to 442, 691, 950, 968, 1089, 1390 in SEQ ID NO: 157 and 434 in SEQ ID NO: 158. In some embodiments, the Kgp and RgpA-based polypeptides described herein include a single amino acid substitution at one or more (e.g., all) positions corresponding to 442, 691, 950, 968, 1089, 1390 in SEQ ID NO: 157 and 434 in SEQ ID NO: 158. In some embodiments, the Kgp and RgpA-based polypeptides described herein include a single amino acid substitution at one or more (e.g., all) positions corresponding to 442, 691, 950, 968, 1089, 1316, and 1390 in SEQ ID NO: 157 and 434 in SEQ ID NO: 158. In some embodiments, the Kgp and RgpA-based polypeptides described herein include a single amino acid substitution at one or more (e.g., all) positions corresponding to 442, 691, 950, 968, 1089, 1316, 1390 in SEQ ID NO: 157 and 434, 671, 691, 766 in SEQ ID NO: 158.

[0391] Examples of polypeptides of the present invention are provided in Table 22.

[0392] [Table 29]

[0393] [Table 30]

[0394] [Table 31]

[0395] [Table 32]

[0396] [Table 33]

[0397] [Table 34]

[0398] [Table 35]

[0399] [Table 36]

[0400] In some embodiments, the polypeptide includes the sequence of sequence number 359, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0401] In some embodiments, the polypeptide includes the sequence of sequence number 361, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0402] In some embodiments, the polypeptide includes the sequence of sequence number 363, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0403] In some embodiments, the polypeptide includes the sequence of sequence number 365, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0404] In some embodiments, the polypeptide includes the sequence of sequence number 369, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0405] In some embodiments, the polypeptide includes the sequence of sequence number 373, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0406] In some embodiments, the polypeptide includes the sequence of sequence number 377, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0407] In some embodiments, the polypeptide includes the sequence of sequence number 381, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0408] In some embodiments, the polypeptide includes the sequence of sequence number 385, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0409] In some embodiments, the polypeptide includes the sequence of sequence number 387, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0410] In some embodiments, the polypeptide includes the sequence of sequence number 389, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0411] In some embodiments, the polypeptide includes the sequence of sequence number 391, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0412] In some embodiments, the polypeptide includes the sequence of sequence number 393, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0413] In some embodiments, the polypeptide includes the sequence of sequence number 395, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0414] In some embodiments, the polypeptide includes the sequence of sequence number 399, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0415] In some embodiments, the polypeptide includes the sequence of sequence number 405, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0416] In some embodiments, the polypeptide includes the sequence of sequence number 417, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0417] In some embodiments, the polypeptide includes the sequence of sequence number 411, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0418] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 360 (i.e., SEQ ID NO: 359 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0419] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 362 (i.e., SEQ ID NO: 361 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0420] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 364 (i.e., SEQ ID NO: 363 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0421] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 366 (i.e., SEQ ID NO: 365 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0422] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 370 (i.e., SEQ ID NO: 369 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0423] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 374 (i.e., SEQ ID NO: 373 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0424] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 378 (i.e., SEQ ID NO: 377 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0425] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 382 (i.e., SEQ ID NO: 381 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0426] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 386 (i.e., SEQ ID NO: 385 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0427] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 388 (i.e., SEQ ID NO: 387 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0428] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 390 (i.e., SEQ ID NO: 389 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0429] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 392 (i.e., SEQ ID NO: 391 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0430] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 394 (i.e., SEQ ID NO: 393 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0431] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 396 (i.e., SEQ ID NO: 395 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0432] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 400 (i.e., SEQ ID NO: 399 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0433] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 406 (i.e., SEQ ID NO: 405 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0434] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 418 (i.e., SEQ ID NO: 417 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0435] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 412 (i.e., SEQ ID NO: 411 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0436] Secretion signal peptide sequence The polypeptides of the present invention as described herein may include a secretion signal peptide sequence. The secretion signal peptide may be cleaved during post-translational processing of the polypeptides described herein. Therefore, the mature form of the polypeptide may not contain a secretion signal peptide sequence. However, the nucleotide sequence encoding the secretion signal peptide sequence may be present in the nucleic acid encoding the polypeptides described herein.

[0437] In some embodiments, the polypeptides of the present invention as described herein may contain a viral or eukaryotic (e.g., human) secretory signal peptide (SS) sequence. The use of viral or eukaryotic secretory signal peptide sequences conjugated to the polypeptides described herein may offer many advantages to immunogenic compositions. Particularly when expressed from mRNA in eukaryotic cells, polypeptides of the present invention containing an SS sequence may show increased extracellular expression compared to polypeptides without an SS sequence. Increased extracellular expression may promote higher immunogenicity and, by extension, may promote better vaccine efficacy.

[0438] Viral SS sequences can be found in publicly accessible databases (e.g., NCBI or UniProt databases) that contain annotated viral polypeptide sequences and identify experimentally validated SS start and end locations.

[0439] In certain embodiments, the locations of SS sequence cleavage sites in a given known input polypeptide sequence and a SS sequence sequence can be predicted using the SignalP algorithm. The SignalP algorithm (more specifically SignalP v6.0) is described in detail in Armenteros et al. (Nature Biotechnology. 37:420-423. 2019), Teufel et al. (Nature Biotechnology. 40:1023-1025. 2022), and services.healthtech.dtu.dk / services / SignalP-6.0 / , each of which is incorporated herein by reference in its entirety. The strength of the prediction is assessed based on a cumulative rank score that takes into account the likelihood of detecting standard features of the signal sequence (SS likelihood score) and the likelihood of cleavage at the cleavage site (cleavage probability score).

[0440] In certain embodiments, the SS sequence is a viral SS sequence. In certain embodiments, the viral secretion signal peptide sequence is derived from a viral sequence in a virus capable of infecting humans. The terms “influenza,” “SARS-CoV-2,” “varicella-zoster virus (VZV),” “measles,” “rubella,” “rabies,” “Ebola,” and “smallpox” preceding the phrase “secretion signal peptide sequence” indicate that the secretion signal peptide is derived from the virus corresponding to its name.

[0441] In certain embodiments, the viral secretion signal peptide is derived from a viral sequence selected from the group consisting of influenza secretion signal peptide sequences, SARS-CoV-2 secretion signal peptide sequences, varicella-zoster virus (VZV) secretion signal peptide sequences, measles secretion signal peptide sequences, rubella secretion signal peptide sequences, mumps secretion signal peptide sequences, Ebola secretion signal peptide sequences, rabies secretion signal peptide sequences, and smallpox secretion signal peptide sequences. These specific signal peptides are derived from viral sequences in viruses that have been administered to humans as vaccines (attenuated, inactivated, or mRNA) with a proven robust safety profile.

[0442] In certain embodiments, the viral secretion signal peptide is selected from the group consisting of influenza hemagglutinin (HA) secretion signal peptide sequences, SARS-CoV-2 spike secretion signal peptide sequences, VZV gB secretion signal peptide sequences, VZV gE secretion signal peptide sequences, VZV gI secretion signal peptide sequences, VZV gK secretion signal peptide sequences, measles F protein secretion signal peptide sequences, rubella E1 protein secretion signal peptide sequences, rubella E2 protein secretion signal peptide sequences, mumps F protein secretion signal peptide sequences, Ebola GP protein secretion signal peptide sequences, rabies virus glycoprotein (rabies G) secretion signal peptide sequences, and smallpox 6kDa IC protein secretion signal peptide sequences.

[0443] In certain embodiments, the viral secretion signal peptide includes an HA secretion signal peptide sequence derived from influenza A or influenza B, preferably from influenza A.

[0444] In certain embodiments, the viral secretion signal peptide includes the signal peptide described in PCT / European Patent Application Publication No. 2023 / 062066, which is incorporated herein by reference as a whole.

[0445] The amino acid sequences of the exemplary viral secretion signal peptides of this disclosure are shown in Table 23 below. The amino acid sequences of the exemplary viral secretion signal peptides derived from influenza A or influenza B of this disclosure are shown in Table 23.1 below.

[0446] [Table 37]

[0447] [Table 38]

[0448] [Table 39]

[0449] [Table 40]

[0450] [Table 41]

[0451] In a particular embodiment, the secretory signal peptide has the sequence of SEQ ID NO: 67.

[0452] The secretion signal peptide sequence may be positioned at the N-terminus or C-terminus (e.g., the N-terminus) of the polypeptide described herein.

[0453] In certain embodiments, the SS amino acid sequence is encoded by a codon-optimized polynucleotide sequence.

[0454] In certain embodiments, the viral secretion signal peptide is bound to the antigen prokaryotic polypeptide by a linker.

[0455] Examples of polypeptides of the present invention, including secretory signaling peptides, are provided in Table 24. This table also provides nucleic acid sequences encoding the polypeptides, which also form part of the present invention. Corresponding polypeptides with mutated glycosylation sites are also included in this table. Mutations in these polypeptides are examples of the glycosylation variants described above.

[0456] [Table 42]

[0457] [Table 43]

[0458] [Table 44]

[0459] [Table 45]

[0460] [Table 46]

[0461] [Table 47]

[0462] [Table 48]

[0463] [Table 49]

[0464] [Table 50]

[0465] In some embodiments, the polypeptide includes the sequence of SEQ ID NO: 2, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0466] In some embodiments, the polypeptide includes the sequence of sequence number 7, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0467] In some embodiments, the polypeptide includes the sequence of SEQ ID NO: 12, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0468] In some embodiments, the polypeptide includes the sequence of SEQ ID NO: 17, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0469] In some embodiments, the polypeptide includes the sequence according to SEQ ID NO: 3, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0470] In some embodiments, the polypeptide includes the sequence of sequence number 279, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0471] In some embodiments, the polypeptide includes the sequence of sequence number 8, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0472] In some embodiments, the polypeptide includes the sequence of SEQ ID NO: 13, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0473] In some embodiments, the polypeptide includes the sequence of sequence number 280, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0474] In some embodiments, the polypeptide includes the sequence of sequence number 297, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0475] In some embodiments, the polypeptide includes the sequence of sequence number 18, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0476] In some embodiments, the polypeptide includes the sequence of sequence number 73, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0477] In some embodiments, the polypeptide includes the sequence of sequence number 74, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0478] In some embodiments, the polypeptide includes the sequence of sequence number 75, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0479] In some embodiments, the polypeptide includes the sequence of sequence number 283, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0480] In some embodiments, the polypeptide includes the sequence of sequence number 76, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0481] In some embodiments, the polypeptide includes the sequence of sequence number 284, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0482] In some embodiments, the polypeptide includes the sequence of sequence number 401, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0483] In some embodiments, the polypeptide includes the sequence of sequence number 413, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0484] In some embodiments, the polypeptide includes the sequence of sequence number 77, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0485] In some embodiments, the polypeptide includes the sequence of sequence number 285, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0486] In some embodiments, the polypeptide includes the sequence of sequence number 407, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0487] The N-terminal methionine in any of the above embodiments may be omitted from the polypeptide. Thus, in some embodiments, the polypeptide includes a sequence that is at least 70% (e.g., SEQ ID NO: 2 without the N-terminal methionine) of the sequence of SEQ ID NO: 255, or a sequence that is at least 75, 80, 85, 90, or 95% (e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) identical thereto.

[0488] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 260 (i.e., SEQ ID NO: 7 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0489] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 265 (i.e., SEQ ID NO: 12 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0490] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 270 (i.e., SEQ ID NO: 17 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0491] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 256 (i.e., SEQ ID NO: 3 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0492] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 281 (i.e., SEQ ID NO: 279 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0493] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 261 (i.e., SEQ ID NO: 8 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0494] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 266 (i.e., SEQ ID NO: 13 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0495] In some embodiments, the polypeptide includes a sequence that is at least 70% identical to the sequence of SEQ ID NO: 299 (i.e., SEQ ID NO: 297 without the N-terminal methionine), or a sequence that is at least 75, 80, 85, 90, or 95% identical thereto (e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). In some embodiments, the polypeptide includes a sequence that is at least 70% identical to the sequence of SEQ ID NO: 282 (i.e., SEQ ID NO: 280 without the N-terminal methionine), or a sequence that is at least 70% identical thereto (e.g., at least 75, 80, 85, 90, or 95% identical thereto, or at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0496] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 271 (i.e., SEQ ID NO: 18 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0497] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 274 (i.e., SEQ ID NO: 73 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0498] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 275 (i.e., SEQ ID NO: 74 without the N-terminal methionine), or a sequence having at least 75, 80, 85, 90, or 95% identity thereto (e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0499] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 276 (i.e., SEQ ID NO: 75 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0500] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 286 (i.e., SEQ ID NO: 283 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0501] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 277 (i.e., SEQ ID NO: 76 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0502] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 287 (i.e., SEQ ID NO: 284 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0503] In some embodiments, the polypeptide includes a sequence that is at least 70% (e.g., sequence number 401 without the N-terminal methionine) identical to sequence number 402, or a sequence that is at least 75%, 80%, 85%, 90%, or 95% identical thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0504] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 414 (i.e., SEQ ID NO: 413 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0505] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 278 (i.e., SEQ ID NO: 77 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0506] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 288 (i.e., SEQ ID NO: 285 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0507] In some embodiments, the polypeptide includes the sequence of SEQ ID NO: 408 (i.e., SEQ ID NO: 407 without the N-terminal methionine), or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). In some embodiments, the nucleic acid encoding the polypeptide includes the sequence of SEQ ID NO: 23. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence of SEQ ID NO: 24.

[0508] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 33. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 34.

[0509] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 43. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 44.

[0510] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 53. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 54.

[0511] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 25. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 26.

[0512] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 289.

[0513] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 35. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 36.

[0514] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 45. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 46.

[0515] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 290.

[0516] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 298.

[0517] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 55. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 56.

[0518] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 78. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 79.

[0519] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 80. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 81.

[0520] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 82. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 83.

[0521] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 291. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 292.

[0522] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 84. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 85.

[0523] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 293. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 294.

[0524] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 86. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 87.

[0525] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 295. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 296.

[0526] In certain embodiments, the amino acid sequence of the polypeptide of the present invention is encoded by a codon-optimized polynucleotide sequence.

[0527] Transmembrane domain (TMB) Polypeptides of the present invention as described herein may include a heteromorphic transmembrane domain (TMB). Including a TMB may be advantageous because it localizes the antigen to the cell membrane. This can reduce the intracellular localization of the antigen and further promote higher immunogenicity compared to antigens without a TMB sequence. Polypeptides containing a heteromorphic transmembrane domain may also include a secretion signal peptide sequence. In some embodiments, nucleic acids described herein encoding polypeptides containing a heteromorphic transmembrane domain also include a nucleotide sequence encoding a secretion signal peptide sequence.

[0528] TMBs may originate from any known TMB in the art, including but not limited to TMBs derived from eukaryotic transmembrane proteins (e.g., mammalian transmembrane proteins such as human transmembrane proteins), TMBs derived from prokaryotic transmembrane proteins, and TMBs derived from viral transmembrane proteins. TMBs may be further identified by in silico prediction algorithms in TMHMM prediction methods, for example, Krogh et al. (J Mol Biol. 305(3):567-580. 2001) and services.healthtech.dtu.dk / services / TMHMM-2.0 / , each of which is incorporated herein by reference in whole. Some features of TMBs are described in more detail in Albers et al. (Chapter 2 - cell membrane structures and functions. Basic Neurochemistry eighth edition. Pages 26-39. 2012), which is incorporated herein by reference. TMBs typically consist mainly of nonpolar (hydrophobic) amino acid residues, though not exclusively, and may traverse the lipid bilayer once or several times. Those skilled in the art are well aware of the methods for determining the hydrophobicity of amino acids. See Simm et al. (2016), Biol Res., 49(1):31; Wimlet and White (1996), Nat Struct Biol., 3(10):842-848; blanco.biomol.uci.edu / hydrophobicity_scales.html; and www.cgl.ucsf.edu / chimera / docs / UsersGuide / midas / hydrophob.html.

[0529] In certain embodiments, the TMB comprises or consists of (a) 15 to 50 amino acid residues, preferably 15 to 30 amino acid residues, more preferably 18 to 25 amino acid residues; and / or (b) at least 50% hydrophobic amino acid residues, preferably selected from the group consisting of alanine, isoleucine, leucine, valine, phenylalanine, tryptophan, and tyrosine; and / or (c) at least one alpha-helix.

[0530] TMBs typically contain alpha helices, each helix containing 18 to 21 amino acids sufficient to span the lipid bilayer. Therefore, in certain embodiments, the transmembrane domain contains one or more alpha helices.

[0531] In certain embodiments, the transmembrane domain is derived from an endogenous membrane protein, as further defined herein and in Albers et al. "Endogenous membrane proteins" (also known as endogenous membrane proteins) are membrane proteins that are persistently bound to the lipid membrane. In certain embodiments, the transmembrane domain is derived from an endogenous polytopic protein. Endogenous polytopic proteins are proteins that span the entire membrane. In certain embodiments, the transmembrane domain is derived from a single-pass transmembrane protein, more specifically, for example, a type I or type II bitopic membrane protein. Single-pass transmembrane proteins cross the membrane only once (i.e., bitopic membrane proteins), while multi-pass transmembrane proteins weave through the membrane and cross it several times. Single-pass transmembrane proteins can be classified into type I, where their carboxyl terminus is oriented toward the cytosol, or type II, where their amino terminus is oriented toward the cytosol. In certain embodiments, the transmembrane domain is derived from an endogenous monotopic protein. Endogenous monotopic proteins are proteins that bind to only one side of a membrane and do not completely cross the lipid bilayer.

[0532] In certain embodiments, the heterologous transmembrane domain originates from a non-human sequence.

[0533] In certain embodiments, the heterologous transmembrane domain is derived from a viral sequence. The terms “influenza,” “SARS-CoV-2,” “varicella-zoster virus (VZV),” “measles,” “rubella,” “rabies,” “Ebola,” and “smallpox” preceding the “transmembrane domain sequence” indicate that the transmembrane domain sequence is derived from the virus corresponding to that name.

[0534] In certain embodiments, the heterologous transmembrane domain is derived from a viral transmembrane domain sequence selected from the group consisting of influenza transmembrane domain sequences, SARS-CoV-2 transmembrane domain sequences, varicella-zoster virus (VZV) transmembrane domain sequences, measles transmembrane domain sequences, rubella transmembrane domain sequences, mumps transmembrane domain sequences, rabies transmembrane domain sequences, and Ebola transmembrane domain sequences. These specific transmembrane domains are derived from viral sequences in viruses that have been administered to humans as vaccines (attenuated, inactivated, or mRNA) with a proven potent safety profile.

[0535] In certain embodiments, the heterologous transmembrane domain is selected from the group consisting of influenza hemagglutinin (HA) transmembrane domain sequences, SARS-CoV-2 spike transmembrane domain sequences, VZV gB transmembrane domain sequences, VZV gE transmembrane domain sequences, VZV gI transmembrane domain sequences, VZV gK transmembrane domain sequences, measles F protein transmembrane domain sequences, rubella E1 protein transmembrane domain sequences, rubella E2 protein domain sequences, mumps F protein transmembrane domain sequences, rabies virus glycoprotein (rabies G) transmembrane domain sequences, and Ebola GP protein transmembrane domain sequences.

[0536] In certain embodiments, the heterologous transmembrane domain includes an HA transmembrane domain sequence derived from influenza A or influenza B, preferably influenza A.

[0537] The amino acid sequences of the exemplary transmembrane domains of the virus described herein are shown in Table 25 below.

[0538] [Table 51]

[0539] In a particular embodiment, the TMB sequence has the sequence GGSILAIYSTVASSLVLVVSLGAISFGG (Sequence ID 70).

[0540] In certain embodiments, the heterologous TMB sequence is located at the N-terminus or C-terminus (e.g., the C-terminus) of the polypeptide described herein.

[0541] In a particular embodiment, the TMB amino acid sequence is encoded by a codon-optimized polynucleotide sequence.

[0542] In certain embodiments, the TMB is linked to a polypeptide described herein by a linker.

[0543] Examples of polypeptides of the present invention containing TMB are provided in Table 26. This table also provides nucleic acid sequences encoding the polypeptides, which also form part of the present invention. Corresponding polypeptides with mutated glycosylation sites are also included in this table. Mutations in these polypeptides are examples of the glycosylation variants described above.

[0544] [Table 52]

[0545] [Table 53]

[0546] [Table 54]

[0547] [Table 55]

[0548] In some embodiments, the polypeptide includes the sequence of sequence number 4, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0549] In some embodiments, the polypeptide includes the sequence by SEQ ID NO: 9, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0550] In some embodiments, the polypeptide includes the sequence of SEQ ID NO: 14, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0551] In some embodiments, the polypeptide includes the sequence of sequence number 19, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0552] In some embodiments, the polypeptide includes the sequence of SEQ ID NO: 5, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0553] In some embodiments, the polypeptide includes the sequence of sequence number 10, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0554] In some embodiments, the polypeptide includes the sequence of sequence number 15, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0555] In some embodiments, the polypeptide includes the sequence of sequence number 20, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0556] The N-terminal methionine in any of the above embodiments may be omitted from the polypeptide. Thus, in some embodiments, the polypeptide includes a sequence that is at least 70% (e.g., SEQ ID NO: 4 without the N-terminal methionine) identical thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0557] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 262 (i.e., SEQ ID NO: 9 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0558] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 267 (i.e., SEQ ID NO: 14 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0559] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 272 (i.e., SEQ ID NO: 19 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0560] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 258 (i.e., SEQ ID NO: 5 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0561] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 263 (i.e., SEQ ID NO: 10 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0562] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 268 (i.e., SEQ ID NO: 15 without the N-terminal methionine), or a sequence having at least 75%, 80%, 85%, 90%, or 95% identity thereto (e.g., at least 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).

[0563] In some embodiments, the polypeptide includes a sequence having at least 70% identity to the sequence of SEQ ID NO: 273 (i.e., SEQ ID NO: 20 without the N-terminal methionine), or a sequence having at least 70% identity to it (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%).

[0564] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 27. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 28.

[0565] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 37. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 38.

[0566] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 47. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 48.

[0567] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 57. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 58.

[0568] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 29. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 30.

[0569] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 39. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 40.

[0570] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 49. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 50.

[0571] In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 59. In some embodiments, the nucleic acid encoding the polypeptide includes the sequence given by SEQ ID NO: 60.

[0572] In certain embodiments, the amino acid sequence of the polypeptide of the present invention is encoded by a codon-optimized polynucleotide sequence.

[0573] Linker In certain embodiments of this disclosure, a secretory signal peptide (SS) sequence or transmembrane domain (TMB) is directly fused to the polypeptide described herein (i.e., there is no linker, such as an amino acid linker, connecting the SS sequence or TMB to the polypeptide described herein). In certain embodiments, the Kgp or RgpA domains of the polypeptide described herein are directly fused to each other (i.e., there is no linker, such as an amino acid linker, connecting the SS sequence or TMB to the polypeptide described herein).

[0574] In other embodiments, the SS sequences and TMBs of the Disclosure are optionally linked to the polypeptides described herein by linkers. In certain embodiments, the linkers are amino acid linkers. In certain embodiments, the amino acid linkers are 1 to 10 amino acid long (for example, the amino acid linkers have lengths of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids). In other embodiments, the Kgp or RgpA domains of the polypeptides described herein are linked to each other by linkers. In certain embodiments, the linkers are amino acid linkers. In certain embodiments, the amino acid linkers are 1 to 10 amino acid long (for example, the amino acid linkers have lengths of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids).

[0575] Exemplary examples of linkers include glycine polymer (Gly)n (wherein n is at least 1, 2, 3, 4, 5, 6, 7, or 8); glycine-serine polymer (GlySer)n (wherein n is at least 1, 2, 3, 4, 5, 6, 7, or 8); glycine-alanine polymer; alanine-serine polymer; and other flexible linkers known in the art.

[0576] Glycine and glycine-serine polymers are relatively unstructured and flexible, and can therefore function as neutral tethers between the SS sequence and / or TMB and the polypeptides described herein. In certain embodiments, the linker is SGS or GSG. In some embodiments, the linker is GGS and / or GG.

[0577] Other exemplary linker sequences include, but are not limited to, the following amino acid sequences: GGG;DGGGS (SEQ ID NO: 241); TGEKP (SEQ ID NO: 242) (Liu et al. Proc.Natl.Acad.Sci.94;5525-5530.1997); GGRR (SEQ ID NO: 243); (GGGGS)n (SEQ ID NO: 244) (wherein n=1,2,3,4 or 5) (Kim et al. Proc.Natl.Acad.Sci.93;1156-1160.1996); EGKSSGSGSESKVD (SEQ ID NO: 245) (Chaudhary et al. Proc.Natl.Acad.Sci.87;1066-1070.1990); KESGSVSSEQLAQFRSLD (SEQ ID NO: 246) (Bird et al. (Cooper et al. Science. 242:423-426. 1988), GGRRGGGS (SEQ ID NO: 247); LQRDGERP (SEQ ID NO: 248); LRQKDGGGSERP (SEQ ID NO: 249); and GTSTGSGKPGSGEGSTKG (SEQ ID NO: 250) (Cooper et al. Blood. 101(4); 1637-1644. 2003). Preferred linkers are shorter, for example, consisting of 2, 3, 4, or 5 amino acids.

[0578] Further examples of linkers are shown in Chen et al. (Adv Drug Deliv Rev. 65(10): 1357-1369. 2013), which are incorporated herein by reference.

[0579] composition The present invention provides compositions comprising one or more nucleic acids of the present disclosure. The present invention also provides compositions comprising one or more polypeptides of the present disclosure. The compositions of the present invention may be, for example, pharmaceutical compositions comprising a pharmaceutically acceptable carrier, excipient, or diluent. In certain embodiments, the compositions of the present invention are immunogenic compositions. “Immunogenic composition” means a composition comprising nucleic acids or proteins that, when administered to a subject, induce an immune response, for example, an antigen-specific immune response. The immune response may be a humoral (antibody) immune response or a cellular immune response. The compositions of the present invention may be vaccine compositions. An immunogenic composition (e.g., a vaccine composition) may induce immunity (e.g., an antibody response) against P. gingivalis infection. The antibody response may include antibodies that bind to the surface of P. gingivalis bacteria or their outer membrane vesicles (OMVs) and neutralize gingipain activity associated with P. gingivalis pathogenicity. Antibodies may cross-react across various P. gingivalis strains.

[0580] As used herein, “protective immunity” or “protective immune response” refers to the induction of immunity or an immune response against an infectious agent indicated by the subject (e.g., P. gingivalis) that prevents or improves an infection or reduces at least one of its symptoms. Specifically, the induction of protective immunity or a protective immune response by administration of the compositions of the present invention is evident by the elimination or reduction of the presence of one or more symptoms of a P. gingivalis infection (e.g., periodontitis). As used herein, the term “immune response” refers to both humoral and cellular immune responses. In some embodiments, treatment with the compositions of the present invention as described herein results in protective immunity against an infection caused by P. gingivalis.

[0581] Nucleic acid composition In one embodiment, the present invention provides a composition comprising a nucleic acid as described herein, comprising a nucleotide sequence encoding a gingipain-based polypeptide as described herein. For example, in one embodiment, the composition comprises a nucleic acid as described herein, comprising a nucleotide sequence encoding a Kgp-based polypeptide. In another embodiment, the composition comprises a nucleic acid as described herein, comprising a nucleotide sequence encoding an RgpA-based polypeptide. In a further embodiment, the composition comprises a nucleic acid as described herein, comprising nucleotide sequences encoding Kgp and RgpA-based polypeptides.

[0582] In further embodiments, the present invention provides compositions comprising (a) a nucleic acid as described herein comprising a nucleotide sequence encoding a Kgp-based polypeptide as described herein; and (b) a nucleic acid as described herein comprising a nucleotide sequence encoding an RgpA-based polypeptide as described herein. Exemplary combinations are shown in Table 27, and polypeptides comprising sequences having at least 70% (e.g., 90% or 95%) identity with the sequences referenced in this table can be used as combinations of Kgp-based polypeptides and RgpA-based polypeptides.

[0583] The specific Kgp-based and Rgp-based polypeptide sequences used in these combinations specified in Table 27 may be modified as described elsewhere in this specification. For example, the DUF2436 domain may be replaced with a shortened DUF2436 domain or a DUF2436 domain having at least 70% (e.g., at least 90% or 95%) identity to the DUF2436 domain.

[0584] [Table 56]

[0585] In certain embodiments, the present invention is (a)i) at least a portion of the Kgp catalytic domain;ii) at least a portion of Kgp DUF2436;iii) at least a portion of the Kgp K1 adhesin domain;iv) a first Kgp moiety containing ABM1 and a first Kgp moiety containing ABM2;v) a second Kgp moiety containing ABM1 and a second Kgp moiety containing ABM2;and vi) a Kgp moiety containing ABM3;(The domains are positioned from the N-terminus to the C-terminus of the polypeptide in the following order: Kgp catalytic domain; first Kgp moiety containing ABM1; Kgp DUF2436; first Kgp moiety containing ABM2; second Kgp moiety containing ABM1; Kgp moiety containing ABM3; Kgp K1; second Kgp moiety containing ABM2) a first nucleic acid encoding a polypeptide; (b)i) at least a portion of the RgpA catalytic domain;ii) at least a portion of RgpA DUF2436;iii) at least a portion of the RgpA K1 adhesin domain;iv) the first RgpA portion containing ABM1 and the first RgpA portion containing ABM2;v) the second RgpA portion containing ABM1 and the second RgpA portion containing ABM2;vi) the RgpA portion containing ABM3;vii) at least a portion of the RgpA K2 adhesin domain;(Domains are in the following order from the N-terminus to the C-terminus of the polypeptide: RgpA catalytic domain; first RgpA portion containing ABM1; RgpA DUF2436; first RgpA portion containing ABM2; second RgpA portion containing ABM1; RgpA portion containing ABM3; RgpA K1; RgpA A second nucleic acid encoding a polypeptide containing K2 (positioned at a second RgpA moiety including ABM2) The present invention provides a composition containing the following:

[0586] In certain embodiments, the present invention is (a)i) at least a portion of the Kgp catalytic domain (at least a portion of the Kgp catalytic domain includes the full-length Kgp catalytic domain);ii) at least a portion of Kgp DUF2436;iii) at least a portion of the Kgp K1 adhesin domain (at least a portion of the Kgp K1 adhesin includes the full-length Kgp K1 adhesin domain);iv) a first Kgp portion containing ABM1 and a first Kgp portion containing ABM2;v) a second Kgp portion containing ABM1 and a second Kgp portion containing ABM2;and vi) a Kgp portion containing ABM3;(domains are in the following order from the N-terminus to the C-terminus of the polypeptide: Kgp catalytic domain; first Kgp portion containing ABM1; Kgp DUF2436; first Kgp portion containing ABM2; second Kgp portion containing ABM1; Kgp portion containing ABM3; Kgp A first nucleic acid encoding a polypeptide containing K1 (positioned at a second Kgp moiety including ABM2); (b)i) at least a portion of the RgpA catalytic domain;ii) at least a portion of RgpA DUF2436;iii) at least a portion of the RgpA K1 adhesin domain;iv) the first RgpA portion containing ABM1 and the first RgpA portion containing ABM2;v) the second RgpA portion containing ABM1 and the second RgpA portion containing ABM2;vi) the RgpA portion containing ABM3;vii) at least a portion of the RgpA K2 adhesin domain;(Domains are in the following order from the N-terminus to the C-terminus of the polypeptide: RgpA catalytic domain; first RgpA portion containing ABM1; RgpA DUF2436; first RgpA portion containing ABM2; second RgpA portion containing ABM1; RgpA portion containing ABM3; RgpA K1; RgpA A second nucleic acid encoding a polypeptide containing K2 (positioned at a second RgpA moiety including ABM2) The present invention provides a composition containing the following:

[0587] In certain embodiments, the present invention is (a)i) at least a portion of the Kgp catalytic domain (at least a portion of the Kgp catalytic domain includes the full-length Kgp catalytic domain);ii) at least a portion of Kgp DUF2436;iii) at least a portion of the Kgp K1 adhesin domain (at least a portion of the Kgp K1 adhesin includes the full-length Kgp K1 adhesin domain);iv) a first Kgp portion containing ABM1 and a first Kgp portion containing ABM2;v) a second Kgp portion containing ABM1 and a second Kgp portion containing ABM2;and vi) a Kgp portion containing ABM3;(domains are in the following order from the N-terminus to the C-terminus of the polypeptide: Kgp catalytic domain; first Kgp portion containing ABM1; Kgp DUF2436; first Kgp portion containing ABM2; second Kgp portion containing ABM1; Kgp portion containing ABM3; Kgp A first nucleic acid encoding a polypeptide (located at a second Kgp region including K1;ABM2; and the glycosylation site corresponding to glycosylation sites N691-T693 in the WT Kgp sequence (SEQ ID NO: 157) is mutated by a T693A substitution; and the glycosylation site corresponding to glycosylation sites N950-T952 in the WT Kgp sequence (SEQ ID NO: 157) is mutated by a T952A substitution); and (b)i) at least a portion of the RgpA catalytic domain;ii) at least a portion of RgpA DUF2436;iii) at least a portion of the RgpA K1 adhesin domain;iv) the first RgpA portion containing ABM1 and the first RgpA portion containing ABM2;v) the second RgpA portion containing ABM1 and the second RgpA portion containing ABM2;vi) the RgpA portion containing ABM3;vii) at least a portion of the RgpA K2 adhesin domain;(Domains are in the following order from the N-terminus to the C-terminus of the polypeptide: RgpA catalytic domain; first RgpA portion containing ABM1; RgpA DUF2436; first RgpA portion containing ABM2; second RgpA portion containing ABM1; RgpA portion containing ABM3; RgpA K1; RgpA A second nucleic acid encoding a polypeptide containing K2 (positioned at a second RgpA moiety including ABM2) The present invention provides a composition containing the following:

[0588] In certain embodiments, the present invention is (a) A first nucleic acid encoding a KGP-based polypeptide containing the sequence of sequence number 279, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%); and (b) A second nucleic acid encoding an RgpA-based polypeptide containing the sequence of sequence number 73, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) The present invention provides a composition containing the following:

[0589] The compositions of the present invention may comprise the combination of nucleic acids as described herein (for example, they may be formulated in the same composition). The combination of nucleic acids as described herein may alternatively comprise two or more distinct compositions (for example, as a combination of compositions for simultaneous, distinct, or sequential administration (for example, in therapeutic use as described herein)).

[0590] A composition of the Disclosure comprising one or more nucleic acids of the Disclosure may also comprise one or more additional components, such as small molecule immunostimulants (e.g., TLR agonists). A composition of the Disclosure may also comprise a delivery system for nucleic acids (e.g., RNA) described herein, such as liposomes, oil-in-water emulsions, or microparticles. In some embodiments, the composition comprises lipid nanoparticles (LNPs). In certain embodiments, the composition comprises nucleic acid molecules of the Disclosure encapsulated within LNPs.

[0591] Polypeptide composition In one embodiment, the present invention provides a composition comprising a gingipain-based polypeptide as described herein.

[0592] In one embodiment, the present invention provides a composition comprising a gingipain-based polypeptide as described herein. For example, in one embodiment, the composition comprises a KGP-based polypeptide. In another embodiment, the composition comprises an RGPA-based polypeptide. In a further embodiment, the composition comprises KGP and RGPA-based polypeptides.

[0593] In further embodiments, the present invention provides compositions comprising (a) a Kgp-based polypeptide as described herein; and (b) an RgpA-based polypeptide as described herein. Exemplary combinations are shown in Table 27, and polypeptides comprising sequences having at least 70% (e.g., 90% or 95%) identity with the sequences referenced in this table can be used as combinations of Kgp-based polypeptides and RgpA-based polypeptides.

[0594] In certain embodiments, the present invention is (a)i) at least a portion of the Kgp catalytic domain;ii) at least a portion of Kgp DUF2436;iii) at least a portion of the Kgp K1 adhesin domain;iv) a first Kgp moiety containing ABM1 and a first Kgp moiety containing ABM2;v) a second Kgp moiety containing ABM1 and a second Kgp moiety containing ABM2;and vi) a Kgp moiety containing ABM3;(The domains are positioned from the N-terminus to the C-terminus of the polypeptide in the following order: Kgp catalytic domain; first Kgp moiety containing ABM1; Kgp DUF2436; first Kgp moiety containing ABM2; second Kgp moiety containing ABM1; Kgp moiety containing ABM3; Kgp K1; second Kgp moiety containing ABM2) of the first polypeptide; (b)i) at least a portion of the RgpA catalytic domain;ii) at least a portion of RgpA DUF2436;iii) at least a portion of the RgpA K1 adhesin domain;iv) the first RgpA portion containing ABM1 and the first RgpA portion containing ABM2;v) the second RgpA portion containing ABM1 and the second RgpA portion containing ABM2;vi) the RgpA portion containing ABM3;vii) at least a portion of the RgpA K2 adhesin domain;(Domains are in the following order from the N-terminus to the C-terminus of the polypeptide: RgpA catalytic domain; first RgpA portion containing ABM1; RgpA DUF2436; first RgpA portion containing ABM2; second RgpA portion containing ABM1; RgpA portion containing ABM3; RgpA K1; RgpA A second polypeptide containing K2 (located at the second RgpA moiety including ABM2) The present invention provides a composition containing the following:

[0595] In certain embodiments, the present invention is (a)i) at least a portion of the Kgp catalytic domain (at least a portion of the Kgp catalytic domain includes the full-length Kgp catalytic domain);ii) at least a portion of Kgp DUF2436;iii) at least a portion of the Kgp K1 adhesin domain (at least a portion of the Kgp K1 adhesin includes the full-length Kgp K1 adhesin domain);iv) a first Kgp portion containing ABM1 and a first Kgp portion containing ABM2;v) a second Kgp portion containing ABM1 and a second Kgp portion containing ABM2;and vi) a Kgp portion containing ABM3;(domains are in the following order from the N-terminus to the C-terminus of the polypeptide: Kgp catalytic domain; first Kgp portion containing ABM1; Kgp DUF2436; first Kgp portion containing ABM2; second Kgp portion containing ABM1; Kgp portion containing ABM3; Kgp A first polypeptide containing K1 (positioned at a second Kgp portion including ABM2); (b)i) at least a portion of the RgpA catalytic domain;ii) at least a portion of RgpA DUF2436;iii) at least a portion of the RgpA K1 adhesin domain;iv) the first RgpA portion containing ABM1 and the first RgpA portion containing ABM2;v) the second RgpA portion containing ABM1 and the second RgpA portion containing ABM2;vi) the RgpA portion containing ABM3;vii) at least a portion of the RgpA K2 adhesin domain;(Domains are in the following order from the N-terminus to the C-terminus of the polypeptide: RgpA catalytic domain; first RgpA portion containing ABM1; RgpA DUF2436; first RgpA portion containing ABM2; second RgpA portion containing ABM1; RgpA portion containing ABM3; RgpA K1; RgpA A second polypeptide containing K2 (located at the second RgpA moiety including ABM2) The present invention provides a composition containing the following:

[0596] In certain embodiments, the present invention is (a)i) at least a portion of the Kgp catalytic domain (at least a portion of the Kgp catalytic domain includes the full-length Kgp catalytic domain);ii) at least a portion of Kgp DUF2436;iii) at least a portion of the Kgp K1 adhesin domain (at least a portion of the Kgp K1 adhesin includes the full-length Kgp K1 adhesin domain);iv) a first Kgp portion containing ABM1 and a first Kgp portion containing ABM2;v) a second Kgp portion containing ABM1 and a second Kgp portion containing ABM2;and vi) a Kgp portion containing ABM3;(domains are in the following order from the N-terminus to the C-terminus of the polypeptide: Kgp catalytic domain; first Kgp portion containing ABM1; Kgp DUF2436; first Kgp portion containing ABM2; second Kgp portion containing ABM1; Kgp portion containing ABM3; Kgp A first polypeptide comprising K1; A second Kgp moiety including ABM2; and the glycosylation site corresponding to glycosylation sites N691-T693 in the WT Kgp sequence (SEQ ID NO: 157) is mutated by a T693A substitution; and the glycosylation site corresponding to glycosylation sites N950-T952 in the WT Kgp sequence (SEQ ID NO: 157) is mutated by a T952A substitution; and (b)i) at least a portion of the RgpA catalytic domain;ii) at least a portion of RgpA DUF2436;iii) at least a portion of the RgpA K1 adhesin domain;iv) the first RgpA portion containing ABM1 and the first RgpA portion containing ABM2;v) the second RgpA portion containing ABM1 and the second RgpA portion containing ABM2;vi) the RgpA portion containing ABM3;vii) at least a portion of the RgpA K2 adhesin domain;(Domains are in the following order from the N-terminus to the C-terminus of the polypeptide: RgpA catalytic domain; first RgpA portion containing ABM1; RgpA DUF2436; first RgpA portion containing ABM2; second RgpA portion containing ABM1; RgpA portion containing ABM3; RgpA K1; RgpA A second polypeptide containing K2 (located at the second RgpA moiety including ABM2) The present invention provides a composition containing the following:

[0597] In certain embodiments, the present invention is (a) A first polypeptide comprising the sequence of sequence number 279, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%); and (b) A second polypeptide comprising the sequence of Sequence ID No. 73, or a sequence having at least 70% identity thereto (e.g., at least 75, 80, 85, 90, or 95%; or e.g., at least 75, 80, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%). The present invention provides a composition containing the following:

[0598] The specific Kgp-based and Rgp-based polypeptide sequences used in these combinations specified in Table 27 may be modified as described elsewhere in this specification. For example, the DUF2436 domain may be replaced with a shortened DUF2436 domain or a DUF2436 domain having at least 70% (e.g., at least 90% or 95%) identity to the DUF2436 domain.

[0599] The compositions of the present invention may comprise combinations of polypeptides as described herein (for example, they may be formulated in the same composition). Alternatively, the combinations of polypeptides as described herein may comprise two or more distinct compositions (for example, as a combination of compositions for simultaneous, distinct, or sequential administration (for example, in therapeutic use as described herein)).

[0600] A composition of the Disclosure comprising one or more polypeptides of the Disclosure may include an adjuvant. As used herein, “adjuvant” means a substance or vehicle that enhances the immune response to an antigen. Examples of adjuvants include, but are not limited to, suspensions of minerals (e.g., alum, aluminum hydroxide, or phosphate) to which an antigen has been adsorbed; and water-in-oil or oil-in-water emulsions in which an antigen solution is emulsified in mineral oil or water (e.g., Freund’s incomplete adjuvant). Sometimes, dead mycobacteria are included to further enhance antigenicity (e.g., Freund’s complete adjuvant). Immunostimulatory oligonucleotides (e.g., CpG motifs) can also be used as adjuvants (see, for example, U.S. Patent Nos. 6,194,388; 6,207,646; 6,214,806; 6,218,371; 6,239,116; 6,339,068; 6,406,705; and 6,429,199). Adjuvants may also include biological molecules, such as Toll-like receptor (TLR) agonists (e.g., SPA14, as described in International Publication No. 2022090359), and co-stimulatory molecules. In some embodiments, the adjuvant is AF03 (oil-in-water squalene emulsion adjuvant).

[0601] In some embodiments, the adjuvant is selected from the group consisting of aluminum-based adjuvants (e.g., AlOOH), squalene-based oil-in-water emulsion adjuvants (e.g., AF03, AS03, MF59), and liposomal adjuvants containing saponins and TLR4 agonists (e.g., SPA14, AS01).

[0602] LNP In certain embodiments, the composition of the present invention (for example, a composition comprising the nucleic acid of the present invention) further comprises lipid nanoparticles (LNPs). In certain embodiments, the nucleic acid of the present invention is encapsulated in LNPs.

[0603] The LNPs of this disclosure may include lipids from the following four categories: (i) ionized lipids (e.g., cationic lipids); (ii) PEGylated lipids; (iii) cholesterol-based lipids; and (iv) helper lipids.

[0604] A. Ionized lipids Ionized lipids facilitate mRNA encapsulation and may be cationic lipids. Cationic lipids provide a positively charged environment at low pH, making it easier to efficiently encapsulate negatively charged mRNA drug substances.

[0605] In some embodiments, the cationic lipid is OF-02: [ka] OF-02 is a non-degradable structural analog of OF-Deg-Lin. OF-Deg-Lin contains a degradable ester linkage for connecting the diketopiperazine core and the double unsaturated tail, while OF-02 contains a non-degradable 1,2-amino-alcohol linkage for connecting the same diketopiperazine core and the double unsaturated tail (Fenton et al., Adv Mater. (2016) 28:2939; U.S. Patent No. 10,201,618). Lipid A, an exemplary LNP formulation as used herein, contains OF-2.

[0606] In some embodiments, the cationic lipid is cKK-E10 (Dong et al., PNAS (2014) 111(11):3955-60; U.S. Patent No. 9,512,073): [ka]

[0607] Lipid B, an exemplary LNP formulation in this specification, contains cKK-E10.

[0608] In some embodiments, the cationic lipid is GL-HEPES-E3-E10-DS-3-E18-1(2-(4-(2-((3-(bis((Z)-2-hydroxyoctadeca-9-en-1-yl)amino)propyl)disulfaneyl)ethyl)piperazine-1-yl)ethyl4-(bis(2-hydroxydecyl)amino)butanoate) (International Publication No. 2022 / 221688), which is a HEPES-based disulfide cationic lipid having a piperazine core and having formula III: [ka]

[0609] Lipid C, an exemplary LNP formulation as described herein, contains GL-HEPES-E3-E10-DS-3-E18-1. Lipid C has the same composition as Lipid A or Lipid B, but differs in its cationic lipids.

[0610] In some embodiments, the cationic lipid is GL-HEPES-E3-E12-DS-4-E10(2-(4-(2-((3-(bis(2-hydroxydecyl)amino)butyl)disulfaneyl)ethyl)piperazine-1-yl)ethyl4-(bis(2-hydroxydodecyl)amino)butanoate) (International Publication No. 2022 / 221688), which is a HEPES-based disulfide cationic lipid having a piperazine core and having formula IV: [ka]

[0611] Lipid D, an exemplary LNP formulation as described herein, contains GL-HEPES-E3-E12-DS-4-E10. Lipid D has the same composition as Lipid A or Lipid B, but differs in its cationic lipids.

[0612] In some embodiments, the cationic lipid is GL-HEPES-E3-E12-DS-3-E14(2-(4-(2-((3-(bis(2-hydroxytetradecyl)amino)propyl)disulfaneyl)ethyl)piperazine-1-yl)ethyl4-(bis(2-hydroxydodecyl)amino)butanoate) (International Publication No. 2022 / 221688), which is a HEPES-based disulfide cationic lipid having a piperazine core and having formula V: [ka]

[0613] Lipid E, an exemplary LNP formulation as described herein, contains GL-HEPES-E3-E12-DS-3-E14. Lipid E ​​has the same composition as Lipid A or Lipid B, but differs in its cationic lipids.

[0614] The cationic lipids GL-HEPES-E3-E10-DS-3-E18-1(III), GL-HEPES-E3-E12-DS-4-E10(IV), and GL-HEPES-E3-E12-DS-3-E14(V) can be synthesized according to the general procedure shown in Scheme 1: [ka]

[0615] In some embodiments, the cationic lipid is MC3, having formula VI: [ka]

[0616] In some embodiments, the cationic lipid is SM-102(9-heptadecanyl 8-{(2-hydroxyethyl)[6-oxo-6-(undecyloxy)hexyl]amino}octanoate), which has formula VII: [ka]

[0617] In some embodiments, the cationic lipid is ALC-0315[(4-hydroxybutyl)azandiyl]di(hexane-6,1-diyl)bis(2-hexyldecanoate), which has formula VIII: [ka]

[0618] In some embodiments, the cationic lipid is cOrn-EE1, which has formula IX: [ka]

[0619] In some embodiments, the cationic lipid is cKK-E10;OF-02;[(6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31-tetraen-19-yl]4-(dimethylamino)butanoate(D-Lin-MC3-DMA);2,2-dilinoleyl-4-dimethylaminoethyl-[1,3]-dioxolane(DLin-KC2-DMA);1,2-dilinoleyloxy-N,N-dimethyl-3-aminopropane(DLin-DMA);di((Z)-non-2-en-1-yl) 9-((4-(dimethylamino)butanoyl)oxy)heptadecanedioate (L319); 9-heptadecanyl 8-{(2-hydroxyethyl)[6-oxo-6-(undecyloxy)hexyl]amino}octanoate (SM-102); [(4-hydroxybutyl)azandiyl]di(hexane-6,1-diyl)bis(2-hexyldecanoate) (ALC-0315); [3-(dimethylamino)-2-[(Z)-octadeca-9-enoyl]oxypropyl](Z)-octadeca-9-enoyl (DODAP); 2,5-bis(3-aminopropylamino)-N-[2-[di(heptadecyl)amino]-2-oxoethyl]pentanamide (DOGS); [(3S,8S,9S,10R,13R,14S,17R)-10,13-dimethyl-17-[(2R)-6-methylheptan-2-yl]-2,3,4,7,8,9,11,12,14,15,16,17-dodecahydro-1H-cyclopenta[a]phenanthrene-3-yl]N-[2-(dimethylamino)ethyl]carbamate (DC-C hol); Tetrakis(8-methylnonyl)3,3',3'',3'''-(((methylazandiyl)bis(propane-3,1diyl))bis(azantriyl))tetrapropionate(306Oi10); Decyl(2-(dioctylammonio)ethyl)phosphate(9A1P9); Ethyl 5,5-di((Z)-heptadeca-8-en-1-yl)-1-(3-(pyrrolidine-1-yl)propyl)-2,5-dihydro-1H-imidazole-2-carboxylate(A2-Iso5-2DC18);Bis(2-(dodecyldisulfanyl)ethyl)3,3'-((3-methyl-9-oxo-10-oxa-13,14-dithia-3,6-diazahexacosyl)azandiyl)dipropionate (BAME-O16B); 1,1'-((2-(4-(2-((2-((bis(2-hydroxydodecyl)amino)ethyl)(2-hydroxydodecyl)amino)ethyl)piperazine-1-yl)ethyl)azandiyl)bis(dodecane-2-ol) (C12-200); 3,6-bis(4-(bis(2-hydroxydodecyl)amino)butyl)piperazine-2,5-dione (cKK-E12); hexa(octa (I-3-yl)9,9',9'',9''',9'''',9''''''-((((Benzene-1,3,5-tricarbonyl)tris(azandiyl))tris(propane-3,1-diyl))tris(azantriyl))hexanonaate(FTT5);(((3,6-dioxopiperazine-2,5-diyl)bis(butane-4,1-diyl))bis(azantriyl))tetrakis(ethane-2,1-diyl)(9Z,9'Z,9''Z,9'''Z,12Z,12'Z,12''Z,12'''Z)-tetrakis(octadeca-9,12-dienoate)(OF-Deg-Lin);TT3;N; 1 ,N 3 ,N 5 -Tris(3-(didodecylamino)propyl)benzene-1,3,5-tricarboxamide;N1-[2-((1S)-1-[(3-aminopropyl)amino]-4-[di(3-aminopropyl)amino]butylcarboxamide)ethyl]-3,4-di[oleyloxy]-benzamide(MVL5);Heptadecan-9-yl8-((2-hydroxyethyl)(8-(nonyloxy)-8-oxooctyl)amino)octanoate(lipid5);GL-HEPES-E3-E10-DS-3-E18-1;GL-HEPES-E3-E12-DS-4-E10;GL-HEPES-E3-E12-DS-3-E14;and combinations thereof may be selected from the group.

[0620] In some embodiments, the cationic lipid is IM-001 and has formula X (European Patent No. 23306049.0): [ka]

[0621] Lipid G, an exemplary LNP formulation as described herein, contains IM-001. Lipid G has the same composition as Lipid A or Lipid B, except that it has different cationic lipids.

[0622] The cationic lipid IM-001(X) can be synthesized according to the general procedure described in Scheme 2: [ka]

[0623] Scheme 2 can be implemented as described in Example 2.

[0624] In some embodiments, the cationic lipid is IS-001 and has formula XI (European Patent No. 23306049.0): [ka]

[0625] Lipid H, an exemplary LNP formulation as described herein, contains IS-001. Lipid H has the same composition as Lipid A or Lipid B, except that it has a different cationic lipid.

[0626] The cationic lipid IS-001(XI) can be synthesized according to the general procedure described in Scheme 3: [ka]

[0627] Scheme 3 can be implemented as described in Example 3.

[0628] In some embodiments, cationic lipids are biodegradable.

[0629] In some embodiments, the cationic lipid is not biodegradable.

[0630] In some embodiments, the cationic lipid is cleavable.

[0631] In some embodiments, the cationic lipid is not cleavable.

[0632] Cationic lipids are described in more detail in Dong et al. (PNAS. 111; 11: 3955 - 60. 2014); Fenton et al. (Adv Mater. 28: 2939. 2016); U.S. Patent No. 9,512,073; and U.S. Patent No. 10,201,618, each of which is incorporated herein by reference.

[0633] B. PEGylated Lipids PEGylated lipid components provide control over the particle size and stability of the nanoparticles. By adding such components, it may be possible to provide a means to prevent complex aggregation, extend the circulation lifetime, and increase the delivery of lipid - nucleic acid pharmaceutical compositions to target tissues (Klibanov et al., FEBS Letters 268(1): 235 - 71990). These components may be selected to be rapidly exchanged from the pharmaceutical composition in vivo (see, for example, U.S. Patent No. 5,885,613).

[0634] Intended PEGylated lipids include C6 - C such as derivatized ceramides (e.g., N - octanoyl - sphingosine - 1 - [succinyl (methoxypolyethylene glycol)] (C8 PEG ceramide)) 20 (e.g., C8, C 10 , C 12 , C 14 , C 16 or C 18Examples include, but are not limited to, polyethylene glycol (PEG) up to 5 kDa in length, covalently bonded to a lipid having an alkyl chain of ) length. In some embodiments, the PEGylated lipids are 1,2-dimiristoyl-rac-glycero-3-methoxypolyethylene glycol (DMG-PEG); 1,2-distearoyl-sn-glycero-3-phosphoethanolamine-polyethylene glycol (DSPE-PEG); 1,2-dilauroyl-sn-glycero-3-phosphoethanolamine-polyethylene glycol (DLPE-PEG); or 1,2-distearoyl-rac-glycero-polyethylene glycol (DSG-PEG), PEG-DAG; PEG-PE; PEG-S-DAG; PEG-S-DMG; PEG-cer; PEG-dialkyloxypropylcarbamate; 2-[(polyethylene glycol)-2000]-N,N-ditetradecylacetamide (ALC-0159); and combinations thereof.

[0635] In certain embodiments, PEG has a high molecular weight, for example, 2000 to 2400 g / mol. In certain embodiments, PEG is PEG2000 (or PEG-2K). In certain embodiments, the PEGylated lipids of this specification are DMG-PEG2000, DSPE-PEG2000, DLPE-PEG2000, DSG-PEG2000, C8 PEG2000, or ALC-0159 (2-[(polyethylene glycol)-2000]-N,N-ditetradecylacetamide). In certain embodiments, the PEGylated lipids of this specification are DMG-PEG2000.

[0636] C. Cholesterol lipids Cholesterol components provide stability to the lipid bilayer structure within the nanoparticles. In some embodiments, the LNP comprises one or more cholesterol-based lipids. Suitable cholesterol-based lipids include, for example, DC-Choi(N,N-dimethyl-N-ethylcarboxamide cholesterol) and 1,4-bis(3-N-oleylaminopropyl)piperazine (Gao et al., Biochem Biophys Res Comm. (1991) 179:280; Wolf et al.). al., BioTechniques (1997) 23:139; US Patent No. 5,744,335), imidazole cholesterol ester ("ICE"; International Publication No. 2011 / 068810 pamphlet), sitosterol (22,23-dihydrostigmasterol), β-sitosterol, sitostanol, fucosterol, stigmasterol (stigma-5,22-dien-3-ol), ergosterol; desmosterol (3β-hydroxy-5,24-cholesterol); lanosterol (8,24-lanostadien-3b-ol); 7-dehydrocholesterol (Δ5,7-cholesterol); dihydrolanosterol (24, Examples include 25-dihydrolanosterol; thymosterol (5α-cholesta-8,24-diene-3β-ol); lasosterol (5α-cholesta-7-ene-3β-ol); diosgenin ((3β,25R)-spirosto-5-ene-3-ol); campesterol (campesto-5-ene-3β-ol); campestanol (5a-campestan-3b-ol); 24-methylenecholesterol (5,24(28)-cholestadiene-24-methylene-3β-ol); cholesteryl margallate (cholesta-5-ene-3β-ylheptadecanoate); cholesteryl oleate; cholesteryl stearate and other modified forms of cholesterol. In some embodiments, the cholesterol-based lipid used in LNP is cholesterol.

[0637] D. Helper lipids Helper lipids enhance the structural stability of LNPs and assist in LNP extrusion into endosomes. This improves the uptake and release of mRNA drug payloads. In some embodiments, the helper lipids are zwitterionic lipids with fusion properties to improve the uptake and release of drug payloads. Examples of helper lipids include 1,2-dioleoyl-SN-glycero-3-phosphoethanolamine (DOPE); 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC); 1,2-dioleoyl-sn-glycero-3-phospho-L-serine (DOPS); 1,2-dieridoyl-sn-glycero-3-phosphoethanolamine (DEPE); and 1,2-dioleoyl-sn-glycero-3-phosphocholine (DPOC), dipalmitoylphosphatidylcholine (DPPC), DMPC, 1,2-dilauroyl-sn-glycero-3-phosphocholine (DLPC), 1,2-distearoylphosphatidylethanolamine (DSPE), and 1,2-dilauroyl-sn-glycero-3-phosphoethanolamine (DLPE).

[0638] Other exemplary helper lipids include dioleoylphosphatidylcholine (DOPC), dioleoylphosphatidylglycerol (DOPG), dipalmitoylphosphatidylglycerol (DPPG), palmitoyloleoylphosphatidylcholine (POPC), palmitoyloleoylphosphatidylethanolamine (POPE), dioleoylphosphatidylethanolamine 4-(N-maleimidomethyl)-cyclohexane-l-carboxylate (DOPE-mal), dipalmitoylphosphatidylethanolamine (DPPE), dimyristoylphosphoethanolamine (DMPE), phosphatidylserine, sphingolipids, sphingomyelin, ceramides, cerebrosides, gangliosides, 16-O-monomethylPE, 16-O-dimethylPE, 18-1-transPE, l-stearoyl-2-oleoylphosphatidylethanolamine (SOPE), or combinations thereof. In certain embodiments, the helper lipid is DOPE. In certain embodiments, the helper lipid is DSPC.

[0639] In various embodiments, the LNP comprises (i) a cationic lipid selected from OF-02, cKK-E10, GL-HEPES-E3-E10-DS-3-E18-1, GL-HEPES-E3-E12-DS-4-E10, GL-HEPES-E3-E12-DS-3-E14, IM-001, or IS-001; (ii) DMG-PEG2000; (iii) cholesterol; and (iv) DOPE.

[0640] In other embodiments, the LNP comprises (i) SM-102; (ii) DMG-PEG2000; (iii) cholesterol; and (iv) DSPC.

[0641] In further other embodiments, the LNP comprises (i) ALC-0315; (ii) ALC-0159; (iii) cholesterol; and (iv) DSPC.

[0642] E. Molar ratio of lipid components The molar ratios of the above components are important for the effectiveness of LNPs in mRNA delivery. The molar ratio of cationic lipids, PEGylated lipids, cholesterol lipids, and helper lipids is A:B:C:D (where A+B+C+D=100%). In some embodiments, the molar ratio of cationic lipids in the LNP to total lipids (i.e., A) is 35-55%, for example 35-50% (e.g., 38-42%, for example 40%, or 45-50%). In some embodiments, the molar ratio of PEGylated lipid components to total lipids (i.e., B) is 0.25-2.75% (e.g., 1-2%, such as 1.5%). In some embodiments, the molar ratio of cholesterol lipids to total lipids (i.e., C) is 20-50% (e.g., 27-30%, for example 28.5%, or 38-43%). In some embodiments, the molar ratio of helper lipids to total lipids (i.e., D) is 5–35% (e.g., 28–32% such as 30% or 8–12% such as 10%). In some embodiments, the (PEGylated lipids + cholesterol) component has the same molar amount as the helper lipids. In some embodiments, the molar ratio of cationic lipids to helper lipids in LNP is greater than 1.

[0643] In certain embodiments, the LNP of this disclosure is Cationic lipids in molar ratios of 35% to 55% or 40% to 50% (for example, cationic lipids in molar ratios of 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, or 55%); Polyethylene glycol (PEG) conjugate (PEG-conjugated) lipids in molar ratios of 0.25% to 2.75% or 1.00% to 2.00% (e.g., PEG-conjugated lipids in molar ratios of 0.25%, 0.50%, 0.75%, 1.00%, 1.25%, 1.50%, 1.75%, 2.00%, 2.25%, 2.50%, or 2.75%); Cholesterol lipids in molar ratios of 20% to 50%, 25% to 45%, or 28.5% to 43% (for example, cholesterol lipids in molar ratios of 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, or 50%); and It contains helper lipids in molar ratios of 5%~35%, 8%~30%, or 10%~30% (for example, helper lipids in molar ratios of 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, or 35%). All molar ratios are relative to the total lipid content of LNPs.

[0644] In a particular embodiment, the LNP comprises a cationic lipid in a molar ratio of 40%, a PEGylated lipid in a molar ratio of 1.5%, a cholesterol-based lipid in a molar ratio of 28.5%, and a helper lipid in a molar ratio of 30%.

[0645] In certain embodiments, the LNPs of this disclosure comprise a cationic lipid in a molar ratio of 45-50%, a PEGylated lipid in a molar ratio of 1.5-1.7%, a cholesterol-based lipid in a molar ratio of 38-43%, and a helper lipid in a molar ratio of 9-10%.

[0646] In certain embodiments, the PEGylated lipid is dimyristoyl-PEG2000 (DMG-PEG2000).

[0647] In various embodiments, the cholesterol-based lipid is cholesterol.

[0648] In some embodiments, the helper lipid is 1,2-dioleoyl-SN-glycero-3-phosphoethanolamine (DOPE).

[0649] In a particular embodiment, the LNP comprises OF-02 in a molar ratio of 35% to 55%; DMG-PEG2000 in a molar ratio of 0.25% to 2.75%; cholesterol in a molar ratio of 20% to 50%; and DOPE in a molar ratio of 5% to 35%.

[0650] In certain embodiments, the LNP comprises cKK-E10 in a molar ratio of 35% to 55%; DMG-PEG2000 in a molar ratio of 0.25% to 2.75%; cholesterol in a molar ratio of 20% to 50%; and DOPE in a molar ratio of 5% to 35%.

[0651] In certain embodiments, the LNP comprises GL-HEPES-E3-E10-DS-3-E18-1 in a molar ratio of 35% to 55%; DMG-PEG2000 in a molar ratio of 0.25% to 2.75%; cholesterol in a molar ratio of 20% to 50%; and DOPE in a molar ratio of 5% to 35%.

[0652] In certain embodiments, the LNP comprises GL-HEPES-E3-E12-DS-4-E10 in a molar ratio of 35% to 55%; DMG-PEG2000 in a molar ratio of 0.25% to 2.75%; cholesterol in a molar ratio of 20% to 50%; and DOPE in a molar ratio of 5% to 35%.

[0653] In certain embodiments, the LNP comprises GL-HEPES-E3-E12-DS-3-E14 in a molar ratio of 35% to 55%; DMG-PEG2000 in a molar ratio of 0.25% to 2.75%; cholesterol in a molar ratio of 20% to 50%; and DOPE in a molar ratio of 5% to 35%.

[0654] In a particular embodiment, the LNP comprises SM-102 in a molar ratio of 35% to 55%; DMG-PEG2000 in a molar ratio of 0.25% to 2.75%; cholesterol in a molar ratio of 20% to 50%; and DSPC in a molar ratio of 5% to 35%.

[0655] In a particular embodiment, the LNP comprises ALC-0315 in a molar ratio of 35% to 55%; ALC-0159 in a molar ratio of 0.25% to 2.75%; cholesterol in a molar ratio of 20% to 50%; and DSPC in a molar ratio of 5% to 35%.

[0656] In certain embodiments, the LNP comprises OF-O2 at a molar ratio of 40%; DMG-PEG2000 at a molar ratio of 1.5%; cholesterol at a molar ratio of 28.5%; and DOPE at a molar ratio of 30%. This LNP formulation is referred to herein as “Lipid A”.

[0657] In certain embodiments, the LNP comprises cKK-E10 at a molar ratio of 40%, DMG-PEG2000 at a molar ratio of 1.5%, cholesterol at a molar ratio of 28.5%, and DOPE at a molar ratio of 30%. This LNP formulation is referred to herein as “Lipid B”.

[0658] In certain embodiments, the LNP comprises 40% molar ratio of GL-HEPES-E3-E10-DS-3-E18-1; 1.5% molar ratio of DMG-PEG2000; 28.5% molar ratio of cholesterol; and 30% molar ratio of DOPE. This LNP formulation is referred to herein as “Lipid C”.

[0659] In certain embodiments, the LNP comprises GL-HEPES-E3-E12-DS-4-E10 (40% molar ratio; DMG-PEG2000 1.5% molar ratio; cholesterol 28.5% molar ratio; and DOPE 30% molar ratio. This LNP formulation is referred to herein as "Lipid D".

[0660] In certain embodiments, the LNP comprises GL-HEPES-E3-E12-DS-3-E14 at a molar ratio of 40%; DMG-PEG2000 at a molar ratio of 1.5%; cholesterol at a molar ratio of 28.5%; and DOPE at a molar ratio of 30%. This LNP formulation is referred to herein as “Lipid E”.

[0661] In certain embodiments, the LNP comprises 50% molar ratio DLin-MC3-DMA(MC3); 1.5% molar ratio DMG-PEG2000; 38.5% molar ratio cholesterol; and 10% molar ratio DSPC. This LNP formulation is referred to herein as “Lipid F”. In certain embodiments, the LNP comprises 50% molar ratio 9-heptadecanyl 8-{(2-hydroxyethyl)[6-oxo-6-(undecyloxy)hexyl]amino}octanoate (SM-102); 10% molar ratio 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC); 38.5% molar ratio cholesterol; and 1.5% molar ratio 1,2-dimiristoyl-rac-glycero-3-methoxypolyethylene glycol-2000 (DMG-PEG2000).

[0662] In certain embodiments, the LNP comprises (4-hydroxybutyl)azandiyl]di(hexane-6,1-diyl)bis(2-hexyldecanoate) (ALC-0315) in a molar ratio of 46.3%; 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC) in a molar ratio of 9.4%; cholesterol in a molar ratio of 42.7%; and 2-[(polyethylene glycol)-2000]-N,N-ditetradecylacetamide (ALC-0159) in a molar ratio of 1.6%.

[0663] In certain embodiments, the LNP comprises (4-hydroxybutyl)azandiyl]di(hexane-6,1-diyl)bis(2-hexyldecanoate) (ALC-0315) in a molar ratio of 47.4%; 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC) in a molar ratio of 10%; cholesterol in a molar ratio of 40.9%; and 2-[(polyethylene glycol)-2000]-N,N-ditetradecylacetamide (ALC-0159) in a molar ratio of 1.7%.

[0664] In certain embodiments, the LNP comprises IM-001 at a molar ratio of 40%; DMG-PEG2000 at a molar ratio of 1.5%; cholesterol at a molar ratio of 28.5%; and DOPE at a molar ratio of 30%. This LNP formulation is referred to herein as “Lipid G”.

[0665] In certain embodiments, the LNP comprises IS-001 at a molar ratio of 40%; DMG-PEG2000 at a molar ratio of 1.5%; cholesterol at a molar ratio of 28.5%; and DOPE at a molar ratio of 30%. This LNP formulation is referred to herein as “Lipid H”.

[0666] In some embodiments, the LNP formulation is defined with respect to “Lipid A,” “Lipid B,” or “Lipid D.” In some embodiments, the LNP formulation is defined with respect to “Lipid G,” or “Lipid H.”

[0667] To calculate the actual amount of each lipid contained in the LNP formulation, first, the molar amount of the cationic lipid is determined based on the desired N / P ratio (where N is the number of nitrogen atoms in the cationic lipid and P is the number of phosphate groups in the mRNA that will be transported by the LNP). Next, the molar amounts of each of the other lipids are calculated based on the molar amount of the cationic lipid and the selected molar ratio. Then, these molar amounts are converted to weight using the molecular weight of each lipid.

[0668] Nucleic acid within F.LNP The LNP compositions described herein may include the nucleic acids (e.g., mRNA) of the present invention.

[0669] LNPs may be polyvalent as desired. In some embodiments, LNPs may carry nucleic acids, such as mRNAs encoding two or more polypeptides of the present invention, e.g., 2, 3, 4, 5, 6, 7, or 8 polypeptides. For example, an LNP may carry multiple nucleic acids (e.g., mRNAs) of the present invention, each encoding a different polypeptide of the present invention; or it may carry polycistronic mRNAs that can be translated into two or more polypeptides of the present invention (e.g., each antigen-coding sequence is separated by a nucleotide linker encoding a self-cleaving peptide, such as a 2A peptide). LNPs carrying different nucleic acids (e.g., mRNAs) typically contain (encapsulate) multiple copies of each nucleic acid. For example, an LNP carrying or encapsulating two different nucleic acids typically carries multiple copies of each of the two different nucleic acids.

[0670] In some embodiments, a single LNP formulation may contain multiple types (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more) of LNPs, each type carrying a different nucleic acid (e.g., mRNA).

[0671] If the nucleic acid is mRNA, the mRNA may be unmodified (i.e., containing only native ribonucleotides A, U, C, and / or G linked by phosphodiester bonds) or chemically modified (e.g., containing nucleotide analogs, e.g., pseudouridine (e.g., N-1-methylpseudridine), 2'-fluororibonucleotide, and 2'-methoxyribonucleotide, and / or phosphorothioate bonds). The mRNA molecule may contain a 5' cap and a poly-A tail.

[0672] G. Buffer solution and other components To stabilize nucleic acids and / or LNPs (e.g., to extend the shelf life of a vaccine product), to facilitate the administration of LNP pharmaceutical compositions, and / or to improve the in vivo expression of nucleic acids, nucleic acids and / or LNPs can be formulated in combination with one or more carriers, targeted ligands, stabilizing reagents (e.g., preservatives and antioxidants), and / or other pharmaceutically acceptable excipients. Examples of such excipients include parabens, thimerosal, thiomersal, chlorobutanol, benzalkonium chloride, and chelating agents (e.g., EDTA).

[0673] The LNP compositions of this disclosure may be provided in a liquid or freeze-dried form. Various cryoprotective agents may be used, but are not limited to sucrose, trehalose, glucose, mannitol, mannose, and dextrose. The cryoprotective agent may constitute 5 to 30% (w / v) of the LNP composition. In some embodiments, the LNP composition contains trehalose, for example, 5 to 30% (e.g., 10%) (w / v). Once formulated with the cryoprotective agent, the LNP composition may be frozen (or freeze-dried and cryopreserved) at -20°C to -80°C.

[0674] The LNP composition may be provided to the patient in a buffer solution (which may be thawed if pre-frozen, or reconstituted with the buffer solution at the bedside if pre-lyophilized). The buffer solution is preferably isotonic and suitable for, for example, intramuscular or intradermal injection. In some embodiments, the buffer solution is phosphate-buffered saline (PBS).

[0675] nucleic acid The nucleic acid of the present invention may be RNA or DNA. The nucleic acid of the present invention may be single-stranded or double-stranded. In certain embodiments, the nucleic acid is RNA, for example, mRNA.

[0676] mRNA In some embodiments, the nucleic acid of the present invention is messenger RNA (mRNA). The mRNA may be modified or unmodified. The mRNA may contain one or more coding regions and non-coding regions. The coding region is also alternatively referred to as an open reading frame (ORF). Non-coding regions in the mRNA include the 5' cap, the 5' untranslated region (UTR), the 3'UTR, and the poly-A tail. The mRNA may be purified from natural sources, produced using recombinant expression systems (e.g., in vitro transcription), and optionally purified or chemically synthesized.

[0677] In certain embodiments, the mRNA includes an ORF encoding the antigen of interest. In certain embodiments, the RNA (e.g., mRNA) further includes at least one 5'UTR, 3'UTR, poly(A) tail and / or 5' cap.

[0678] 5' Cap The 5' cap of mRNA provides resistance to nucleases found in most eukaryotic cells and can enhance translational efficiency. Several types of 5' caps are known, including the 7-methylguanosine cap ("m"). 7 Also referred to as "G" or "Cap 0"), it contains guanosine linked to the first transcribed nucleotide via a 5'-5'-triphosphate bond.

[0679] The 5' cap is typically added as follows: first, one of the terminal phosphate groups is removed from the 5' nucleotide by a phosphatase at the end of the RNA, leaving two terminal phosphates; then, guanosine triphosphate (GTP) is added to the terminal phosphate by guanylyltransferase to form a 5'5'5 triphosphate bond; and then, the 7-nitrogen of guanine is methylated by methyltransferase. Examples of cap structures include, but are not limited to, m7G(5')ppp, (5'(A, G(5')ppp(5')A, and G(5')ppp(5')G. Additional cap structures are described in U.S. Patent Application Publication 2016 / 0032356 and U.S. Patent Application Publication 2018 / 0125989, which are incorporated herein by reference.

[0680] To generate a 5'-guanosine cap structure according to the manufacturer's protocol, the following chemical RNA cap analogues can be used to simultaneously complete the 5'-cap addition of polynucleotides during an in vitro transcription reaction: 3'-O-Me-m7G(5')ppp(5')G(ARCA cap); G(5')ppp(5')A; G(5')ppp(5')G; m7G(5')ppp(5')A; m7G(5')ppp(5')G; m7G(5')ppp(5')(2'OMeA)pG; m7G(5')ppp(5')(2'OMeA)pU; m7G(5')ppp(5')(2'OMeG)pG (New England BioLabs, Ipswich, MA; TriLink Biotechnologies). Using vaccinia virus capaddase, the 5'-cap addition of modified RNA can be completed post-transcriptionally to produce the cap 0 structure: m7G(5')ppp(5')G. Using both vaccinia virus capaddase and 2'-O methyl-transferase, the cap 1 structure can be produced to generate m7G(5')ppp(5')G-2'-O-methyl. The cap 2 structure may be generated from the cap 1 structure, followed by 2'-O methylation of the third-to-last 5'-nucleotide using 2'-O methyl-transferase. The cap 3 structure may be generated from the cap 2 structure, followed by 2'-O methylation of the fourth-to-last 5'-nucleotide using 2'-O methyl-transferase.

[0681] In a particular embodiment, the mRNA of the present invention includes a 5' cap selected from the group consisting of 3'-O-Me-m7G(5')ppp(5')G (ARCA cap), G(5')ppp(5')A, G(5')ppp(5')G, m7G(5')ppp(5')A, m7G(5')ppp(5')G, m7G(5')ppp(5')(2'OMeA)pG, m7G(5')ppp(5')(2'OMeA)pU, and m7G(5')ppp(5')(2'OMeG)pG.

[0682] In a particular embodiment, the mRNA of the present invention includes the following 5' cap. [ka]

[0683] Untranslated area (UTR) In some embodiments, the mRNA of the present invention includes a 5' and / or 3' untranslated region (UTR). In mRNA, the 5' UTR begins at the transcription start site and continues to the start codon, but does not contain the start codon. The 3' UTR begins immediately after the stop codon and continues to the transcription termination signal.

[0684] In some embodiments, the mRNA disclosed herein may include a 5'UTR containing one or more elements that affect mRNA stability or translation. In some embodiments, the 5'UTR may be about 10 to 5,000 nucleotides long. In some embodiments, the 5'UTR may be about 50 to 500 nucleotides long. In some embodiments, the 5'UTR may be at least about 10 nucleotides long, about 20 nucleotides long, about 30 nucleotides long, about 40 nucleotides long, about 50 nucleotides long, about 100 nucleotides long, about 150 nucleotides long, about 200 nucleotides long, about 250 nucleotides long, about 300 nucleotides long, about 350 nucleotides long, about 400 nucleotides long, about 450 nucleotides long, about 500 nucleotides long, about 550 nucleotides long, about 600 nucleotides long, about The lengths are approximately 650 nucleotides, 700 nucleotides, 750 nucleotides, 800 nucleotides, 850 nucleotides, 900 nucleotides, 950 nucleotides, 1,000 nucleotides, 1,500 nucleotides, 2,000 nucleotides, 2,500 nucleotides, 3,000 nucleotides, 3,500 nucleotides, 4,000 nucleotides, 4,500 nucleotides, or 5,000 nucleotides.

[0685] In some embodiments, the mRNA disclosed herein may include a 3'UTR comprising one or more polyadenylation signals, protein binding sites affecting the stability of mRNA site in a cell, or one or more binding sites to miRNA. In some embodiments, the 3'UTR may be 50 to 5,000 nucleotides or longer. In some embodiments, the 3'UTR may be 50 to 1,000 nucleotides or longer. In some embodiments, the 3'UTR is at least about 50 nucleotides long, about 100 nucleotides long, about 150 nucleotides long, about 200 nucleotides long, about 250 nucleotides long, about 300 nucleotides long, about 350 nucleotides long, about 400 nucleotides long, about 450 nucleotides long, about 500 nucleotides long, about 550 nucleotides long, about 600 nucleotides long, about 650 nucleotides long, about 700 nucleotides long, about 750 nucleotides long, about 800 nucleotides long, about 850 nucleotides long, about 900 nucleotides long, about 950 nucleotides long, about 1,000 nucleotides long, about 1,500 nucleotides long, about 2,000 nucleotides long, about 2,500 nucleotides long, about 3,000 nucleotides long, about 3,500 nucleotides long, about 4,000 nucleotides long, about 4,500 nucleotides long, or about 5,000 nucleotides long.

[0686] In some embodiments, the mRNA disclosed herein may include a 5' or 3' UTR derived from a gene different from the gene encoded by the mRNA transcript (i.e., the UTR is a heterologous UTR).

[0687] In certain embodiments, the 5' and / or 3'UTR sequences may be derived from stable mRNA to enhance mRNA stability (e.g., globin, actin, GAPDH, tubulin, histone, or citrate cycle enzymes). For example, the 5'UTR sequence may contain a sub-sequence or fragment thereof of the CMV earliest 1 (IE1) gene to improve nuclease resistance and / or improve mRNA half-life. It is also conceivable to include a sequence or fragment thereof encoding human growth hormone (hGH) at the 3' end or untranslated region of the mRNA. Generally, these modifications improve mRNA stability and / or pharmacokinetic properties (e.g., half-life) compared to the unmodified counterpart, such as modifications made to improve resistance of such mRNA to in vivonuclease digestion.

[0688] Exemplary 5'UTRs include sequences derived from the CMV earliest 1(IE1) gene (U.S. Patent Application Publication No. 2014 / 0206753 and No. 2015 / 0157565, each of which is incorporated herein by reference), or sequences GGGAUCCUACC (SEQ ID NO: 140) (U.S. Patent Application Publication No. 2016 / 0151409, incorporated herein by reference).

[0689] In various embodiments, the 5'UTR may be derived from the 5'UTR of a TOP gene. TOP genes are typically characterized by the presence of a 5'-terminal oligopyrimidine (TOP) tract. Furthermore, most TOP genes are characterized by growth-related translational regulation. However, TOP genes with tissue-specific translational regulation are also known. In certain embodiments, the 5'UTR derived from the 5'UTR of a TOP gene lacks the 5'TOP motif (oligopyrimidine tract) (e.g., U.S. Patent Application Publications 2017 / 0029847, 2016 / 0304883, 2016 / 0235864, and 2016 / 0166710, each incorporated herein by reference).

[0690] In a particular embodiment, the 5'UTR is derived from the ribosomal protein large 32 (L32) gene (see U.S. Patent Application Publication No. 2017 / 0029847 above).

[0691] In a particular embodiment, the 5'UTR is derived from the 5'UTR of the h...

Claims

1. A nucleic acid comprising a nucleotide sequence encoding a polypeptide, wherein the polypeptide is i) At least a portion of the Porphyromonas gingivalis Lys-specific proteinase (KGP) catalytic domain; ii) At least a portion of domain 2436 (DUF2436) of the unknown function of Porphyromonas gingivalis KGP; iii) At least a portion of the KGP K1 adhesin domain of Porphyromonas gingivalis; iv) A first Porphyromonas gingivalis Kgp moiety containing an adhesin-binding motif 1 (ABM1) and a first Porphyromonas gingivalis Kgp moiety containing an adhesin-binding motif 2 (ABM2); v) a second Porphyromonas gingivalis Kgp moiety containing ABM1 and a second Porphyromonas gingivalis Kgp moiety containing ABM2; and vi) Kgp moiety of Porphyromonas gingivalis containing adhesin-binding motif 3 (ABM3) Nucleic acids, including

2. (a) The first Kgp portion including ABM1 includes the sequence of sequence number 120, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto; (b) The first Kgp portion including ABM2 includes the sequence of sequence number 130, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto; (c) The second Kgp portion including ABM1 includes the sequence of Sequence ID No. 120, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto; and / or (d) The nucleic acid according to claim 1, wherein the second Kgp portion including ABM2 includes the sequence according to SEQ ID NO: 130, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto.

3. (a) The first Kgp portion including ABM1 comprises the sequence according to Sequence ID No. 89, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto; (b) The first Kgp portion including ABM2 includes the sequence according to Sequence ID No. 91, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto; (c) The second Kgp portion including ABM1 includes the sequence according to SEQ ID NO: 92, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto; and / or (d) The nucleic acid according to claim 1 or 2, wherein the second Kgp portion including ABM2 includes the sequence according to SEQ ID NO: 95, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto.

4. At least a portion of Kgp DUF2436 is full-length Kgp DUF2436, and optionally, the full-length Kgp DUF2436 is the sequence according to Sequence ID No. 90 or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto. The nucleic acid according to any one of claims 1 to 3, including the nucleic acid.

5. At least a portion of the KGP catalytic domain, (a) full-length KGP catalyst domain (optionally, the full-length KGP catalyst domain includes the sequence according to Sequence ID No. 64, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto); or (b) Shortened Kgp catalytic domain (Optionally, the shortened Kgp catalytic domain includes a Lys-gingipain active site peptide (KAS peptide), for example, (i) The KAS peptide comprises a sequence having at least 70% (e.g., at least 90 or 95%) identity with the Kas2 peptide according to SEQ ID NO: 61; or (ii) The KAS peptide is the extended Kas2 peptide according to SEQ ID NO: 88, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto. (including) The nucleic acid according to any one of claims 1 to 4.

6. The polypeptide further comprises at least a portion of the Porphyromonas gingivalis Arg-specific proteinase A (RgpA) or Arg-specific proteinase B (RgpB) catalytic domain, and optionally, at least a portion of the RgpA or RgpB catalytic domain is a shortened RgpA or RgpB catalytic domain, for example, the shortened RgpA or RgpB catalytic domain comprises an Arg-gingipain active site peptide (RAS peptide), for example, (i) The RAS peptide contains a sequence that is at least 70% (e.g., at least 90 or 95%) identical to the Ras2 peptide according to SEQ ID NO: 166; or (ii) The RAS peptide includes the Ras2 peptide extended according to SEQ ID NO: 97, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto. The nucleic acid according to any one of claims 1 to 5.

7. The nucleic acid according to any one of claims 1 to 6, wherein the polypeptide further comprises at least a portion of the Porphyromonas gingivalis Kgp K2 adhesin domain, and optionally, at least a portion of the Kgp K2 adhesin domain is a full-length K2 adhesin domain, for example, the full-length Kgp K2 adhesin domain comprises the sequence of SEQ ID NO: 96, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto.

8. The nucleic acid according to claim 6, wherein the polypeptide further comprises at least a portion of the Porphyromonas gingivalis RgpA K2 adhesin domain, and optionally, at least a portion of the RgpA K2 adhesin domain is a full-length K2 adhesin domain, for example, the full-length RgpA K2 adhesin domain comprises the sequence of SEQ ID NO: 104, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto.

9. The nucleic acid according to any one of claims 6 to 8, wherein the polypeptide comprises at least a portion of RgpA DUF2436, and optionally, the RgpA DUF2436 is full-length RgpA DUF2436, for example, the full-length RgpA DUF2436 comprises the sequence of SEQ ID NO: 100, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto.

10. (a) At least a portion of the Kgp K1 adhesin domain is a full-length Kgp K1 adhesin domain, and optionally, both the full-length Kgp K1 adhesin domain and the Kgp portion including ABM3 contain the sequence of Sequence ID No. 93, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto; or (b) At least a portion of the Kgp K1 adhesin domain is a shortened Kgp K1 adhesin domain, and optionally, the shortened Kgp K1 adhesin domain and the Kgp portion including ABM3 both include the sequence of Sequence ID No. 94, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto. The nucleic acid according to any one of claims 1 to 9.

11. A nucleic acid comprising a nucleotide sequence encoding a polypeptide, wherein the polypeptide is i) At least a portion of the catalytic domain of Porphyromonas gingivalis Arg-specific proteinase A (RgpA) or Arg-specific proteinase B (RgpB), for example, at least a portion of the catalytic domain of Porphyromonas gingivalis Arg-specific proteinase A (RgpA); ii) At least a portion of Porphyromonas gingivalis RgpA DUF2436; iii) At least a portion of the RgpA K1 adhesin domain of Porphyromonas gingivalis; iv) A first Porphyromonas gingivalis RgpA moiety containing ABM1, and a first Porphyromonas gingivalis RgpA moiety containing ABM2; v) a second Porphyromonas gingivalis RgpA moiety containing ABM1, and a second Porphyromonas gingivalis RgpA moiety containing ABM2; and vi) RgpA portion of Porphyromonas gingivalis containing ABM3 Nucleic acids, including

12. (a) The first RgpA portion including ABM1 includes the sequence of sequence number 120, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto; (b) The first RgpA portion including ABM2 includes the sequence of sequence number 130, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto; (c) The second RgpA portion including ABM1 includes the sequence of SEQ ID NO: 120, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto; and / or (d) The second RgpA portion including ABM2 includes the sequence of sequence number 130, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto, The nucleic acid according to claim 11.

13. (a) The first RgpA portion including ABM1 includes the sequence according to SEQ ID NO: 99, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto; (b) The first RgpA portion including ABM2 includes the sequence of sequence number 101, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto; (c) The second RgpA portion including ABM1 includes the sequence of sequence number 102, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto; and / or (d) The second RgpA portion including ABM2 includes the sequence of sequence number 105, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto, The nucleic acid according to claim 11 or claim 12.

14. The nucleic acid according to any one of claims 11 to 13, wherein at least a portion of the RgpA DUF2436 is full-length RgpA DUF2436, and optionally the full-length RgpA DUF2436 includes the sequence of sequence number 100 or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto.

15. At least a portion of the RgpA catalytic domain, (a) full-length RgpA catalytic domain (optionally, the full-length RgpA catalytic domain includes the sequence of Sequence ID No. 98, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto); or (b) Shortened Kgp catalytic domain (Optionally, the shortened RgpA catalytic domain includes an Arg-gingipain active site peptide (RAS peptide), for example, (i) The RAS peptide comprises the Ras2 peptide according to SEQ ID NO: 166, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto; or (ii) The RAS peptide includes the extended Ras2 peptide according to SEQ ID NO: 97, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto. The nucleic acid according to any one of claims 11 to 14.

16. The nucleic acid according to any one of claims 11 to 15, wherein at least a portion of the RgpA K1 adhesin domain is a shortened RgpA K1 adhesin domain, and optionally, both the shortened RgpA K1 adhesin domain and the RgpA portion including ABM3 include the sequence of SEQ ID NO: 103, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto.

17. The nucleic acid according to any one of claims 11 to 16, wherein the polypeptide comprises at least a portion of the Porphyromonas gingivalis RgpA K2 adhesin domain, and optionally, at least a portion of the RgpA K2 adhesin domain is a full-length K2 adhesin domain, for example, the full-length Kgp K2 adhesin domain comprises the sequence of SEQ ID NO: 104, or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto.

18. The nucleic acid according to any one of claims 1 to 17, wherein the polypeptide further comprises a secretion signal peptide sequence.

19. The nucleic acid according to claim 18, wherein the secreted signal peptide comprises the secreted signal peptide sequence of the HA protein of influenza A virus, and optionally, the secreted signal peptide sequence comprises the sequence according to Sequence ID No.

67.

20. The nucleic acid according to any one of claims 1 to 19, wherein the polypeptide comprises a sequence of any one of SEQ ID NOs: 1-20, 73-77, 279, 280, 283-285, 297, 359, 361, 363, 365, 367, 369, 371, 373, 375, 377, 379, 381, 383, 385, 387, 389, 391, 393, 395, 397, 399, 401, 403, 405, 407, 409, or 411; or a sequence having at least 70% (e.g., at least 90 or 95%) identity thereto.

21. The nucleic acid is messenger RNA (mRNA), and by arbitrary selection, (i) The mRNA comprises at least one 5' untranslated portion (5'UTR), at least one 3' untranslated portion (3'UTR), and / or at least one polyadenylated (poly(A)) sequence; (ii) The mRNA is unmodified or has at least one chemical modification, optionally the mRNA has at least one chemical modification, for example the chemical modification has N1-methylpseudridine; and / or (iii) The nucleic acid according to any one of claims 1 to 20, wherein the mRNA is self-replicating mRNA or non-replicating mRNA, for example, non-replicating mRNA.

22. The polypeptide according to any one of claims 1 to 20.

23. A composition comprising a nucleic acid according to any one of claims 1 to 21, or a polypeptide according to claim 22, wherein the composition is an immunogenic composition.

24. A composition, (a) i) At least a portion of the Kgp catalytic domain (optionally, at least a portion of the Kgp catalytic domain includes a full-length Kgp catalytic domain); ii) At least a portion of KGP DUF2436; iii) At least a portion of the KGP K1 adhesin domain (optionally, at least a portion of the KGP K1 adhesin includes the full-length KGP K1 adhesin domain); iv) A first KGP portion including ABM1 and a first KGP portion including ABM2; v) A second KGP portion including ABM1 and a second KGP portion including ABM2; and vi) KGP portion including ABM3; (The domains are positioned in the following order from the N-terminus to the C-terminus of the polypeptide: Kgp catalytic domain; first Kgp moiety containing ABM1; Kgp DUF2436; first Kgp moiety containing ABM2; second Kgp moiety containing ABM1; Kgp moiety containing ABM3; Kgp K1; second Kgp moiety containing ABM2; A first nucleic acid encoding the polypeptide comprising: (Optionally, the glycosylation sites corresponding to glycosylation sites N691 to T693 in the WT Kgp sequence (SEQ ID NO: 157) are mutated by T693A substitution; and the glycosylation sites corresponding to glycosylation sites N950 to T952 in the WT Kgp sequence (SEQ ID NO: 157) are mutated by T952A substitution); and (b) i) At least a portion of the RgpA catalytic domain; ii) At least a portion of RgpA DUF2436; iii) At least a portion of the RgpA K1 adhesin domain; iv) A first RgpA portion including ABM1 and a first RgpA portion including ABM2; v) A second RgpA portion including ABM1 and a second RgpA portion including ABM2; vi) RgpA portion including ABM3; vii) At least a portion of the RgpA K2 adhesin domain; A second nucleic acid encoding the polypeptide, which comprises (the domain being positioned in the following order from the N-terminus to the C-terminus of the polypeptide: RgpA catalytic domain; first RgpA moiety containing ABM1; RgpA DUF2436; first RgpA moiety containing ABM2; second RgpA moiety containing ABM1; RgpA moiety containing ABM3; RgpA K1; RgpA K2; second RgpA moiety containing ABM2) A composition containing the following:

25. The composition is i) A first nucleic acid encoding a polypeptide comprising the sequence according to Sequence ID No. 279, or a sequence having at least 70% (e.g., at least 90% or 95%) identity thereto; and ii) A second nucleic acid encoding a polypeptide containing the sequence of Sequence ID No. 73, or a sequence having at least 70% (e.g., at least 90% or 95%) identity thereto. The composition according to claim 23 or claim 24, comprising:

26. A composition, (a) i) At least a portion of the Kgp catalytic domain (optionally, at least a portion of the Kgp catalytic domain includes a full-length Kgp catalytic domain); ii) At least a portion of KGP DUF2436; iii) At least a portion of the KGP K1 adhesin domain (optionally, at least a portion of the KGP K1 adhesin includes the full-length KGP K1 adhesin domain); iv) A first KGP portion including ABM1 and a first KGP portion including ABM2; v) A second KGP portion including ABM1 and a second KGP portion including ABM2; and vi) KGP portion including ABM3; (The domains are positioned in the following order from the N-terminus to the C-terminus of the polypeptide: Kgp catalytic domain; first Kgp moiety containing ABM1; Kgp DUF2436; first Kgp moiety containing ABM2; second Kgp moiety containing ABM1; Kgp moiety containing ABM3; Kgp K1; second Kgp moiety containing ABM2; A first polypeptide comprising: (Optionally, the glycosylation sites corresponding to glycosylation sites N691 to T693 in the WT Kgp sequence (SEQ ID NO: 157) are mutated by T693A substitution; and the glycosylation sites corresponding to glycosylation sites N950 to T952 in the WT Kgp sequence (SEQ ID NO: 157) are mutated by T952A substitution); and (b) i) At least a portion of the RgpA catalytic domain; ii) At least a portion of RgpA DUF2436; iii) At least a portion of the RgpA K1 adhesin domain; iv) A first RgpA portion including ABM1 and a first RgpA portion including ABM2; v) A second RgpA portion including ABM1 and a second RgpA portion including ABM2; vi) RgpA portion including ABM3; vii) At least a portion of the RgpA K2 adhesin domain; (The domain is positioned in the following order from the N-terminus to the C-terminus of the polypeptide: RgpA catalytic domain; first RgpA moiety containing ABM1; RgpA DUF2436; first RgpA moiety containing ABM2; second RgpA moiety containing ABM1; RgpA moiety containing ABM3; RgpA K1; RgpA K2; second RgpA moiety containing ABM2) A composition containing the following:

27. The composition is i) A first polypeptide comprising the sequence according to Sequence ID No. 279, or a sequence having at least 70% (e.g., at least 90% or 95%) identity thereto; and ii) A second polypeptide comprising the sequence of Sequence ID No. 73, or a sequence having at least 70% (e.g., at least 90% or 95%) identity thereto. The composition according to claim 23 or claim 26, comprising:

28. A nucleic acid according to any one of claims 1 to 21, a polypeptide according to claim 22, or a composition according to any one of claims 23 to 27, for use as a pharmaceutical.

29. A nucleic acid according to any one of claims 1 to 21, a polypeptide according to claim 22, or a composition according to any one of claims 23 to 27 for use in treating or preventing Porphyromonas gingivalis infection, for example, periodontitis.

30. A vaccine comprising a nucleic acid according to any one of claims 1 to 21, a polypeptide according to claim 22, or a composition according to any one of claims 23 to 27.