Predicting immunogenic t cell epitopes in pathogens using antigen and peptide features

A method using sequence analysis and machine learning generates immunogenic peptide libraries for pathogens, addressing the challenge of identifying effective T cell epitopes, enabling diagnostic and therapeutic applications for Streptococcus pneumoniae.

WO2026090277A1PCT designated stage Publication Date: 2026-04-30LA JOLLA INST FOR IMMUNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
LA JOLLA INST FOR IMMUNOLOGY
Filing Date
2025-10-22
Publication Date
2026-04-30

Smart Images

  • Figure US2025052050_30042026_PF_FP_ABST
    Figure US2025052050_30042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are methods for identifying immunogenic peptides analyzing binding affinities of a set of possible immunogenic peptide epitopes by selecting sequence data from a template antigen-binding molecule from a set of possible template antigen binding molecules; using two or more machine learning algorithms to train a model and combining into an ensemble model; and generating sequence data of one or more immunogenic peptide epitope molecules. Also provided are methods of using and compositions, including epitope megapools, and methods for detecting the presence of: a Streptococcus sp., from any one of those sequences set forth in SEQ ID NOS: 1-10), or a subsequence, portion, homologue, variant or derivative thereof; a fusion protein; a pool of 2 or more peptides; or a polynucleotide that encodes one or more peptides or proteins, or a subsequence, portion, homologue, variant or derivative thereof; vaccines, diagnostics, therapies, and kits, comprising such proteins or peptides.
Need to check novelty before this filing date? Find Prior Art

Description

PREDICTING IMMUNOGENIC T CELL EPITOPES IN PATHOGENS USING ANTIGEN AND PEPTIDE FEATURES CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application Serial No. 63 / 711,992, filed October 25, 2024, the entire contents of which are incorporated herein by reference.TECHNICAL FIELD OF THE INVENTION

[0002] The present invention relates in general to the field of proteins and peptides that are T cell epitopes and / or antigens for pathogens, and more particularly, to compositions and methods for the prevention, treatment, diagnosis, kits, and uses of such T cell epitopes and antigens, including megapools, for use in detecting and characterizing pathogen-specific responses in infection and following vaccination.STATEMENT OF FEDERALLY FUNDED RESEARCH

[0003] The inventions described in the present disclosure were made with government support under Contract No. 75N93019C00067, awarded by the National Institutes of Health. The government has certain rights in the invention.REFERENCE TO ELECTRONIC SEQUENCE LISTING

[0004] The application contains a Sequence Listing which has been submitted electronically in .XML format and is hereby incorporated by reference in its entirety. Said .XML copy, created on October 22, 2025, is named “LJII2037WO.xml” and is 9,653 bytes in size. The sequence listing contained in this .XML file is part of the specification and is hereby incorporated by reference herein in its entirety.BACKGROUND OF THE INVENTION

[0005] Without limiting the scope of the invention, its background is described in connection with infectious pathogens.

[0006] T cell lymphocytes recognize linear peptides, or epitopes, presented to them via major histocompatibility complex (MHC) molecules on host cells. CD4+ T cells recognize epitopes presented by MHC class II molecules, which are predominantly expressed on professional antigen-presenting cells, such as macrophages. These cells are the primary source of T cell epitopes derived from bacteria (1).

[0007] T cell epitope predictions have been successfully used to identify candidate peptides from infectious agents, allergens, and cancer cells (2). These predictions have predominantly relied on the features of the peptides themselves, especially their predicted MHC binding affinity. When the goal is to identify peptides likely to be immunogenic in the general population, predictions of MHC class II binding can narrow down the number of peptide candidates to the top -20% (3). However, in the case of complex bacterial pathogens expressing thousands of antigens, this will still leave tens- to hundreds of thousands of possible candidate peptides.

[0008] A need remains for identifying antigens and T cell epitopes for use in diagnostics, treatments,vaccines, kits, etc., for pathogen-related diseases and conditions, including, e.g., bacteria, viruses, protozoans, and helminths. There is additionally a specific need in the art for optimized megapools for use in detecting and characterizing pathogen-specific responses in infection and following vaccination.SUMMARY OF THE INVENTION

[0009] As embodied and broadly described herein, an aspect of the present disclosure relates to a method for generating a library of immunogenic peptide epitope molecules, the method comprising: (a) analyzing binding affinities of a set of possible immunogenic peptide epitopes for an epitope of interest from a pathogen by: selecting sequence data from a template antigen-binding molecule from a set of possible template antigen binding molecules, wherein the selected template is a different pathogen than the epitope of interest and wherein the set of possible template antigen-binding molecules consists of two or more known peptides that bind MHC and are presented to T cells, wherein the selecting sequence data comprises screening the set of possible template antigen-binding molecules based on at least two of: gene expression and conservation; one or more peptide-level features; and one or more antigen-level features; (b) using two or more machine learning algorithms to train a model and combining into an ensemble model; (c) generating sequence data of one or more immunogenic peptide epitope molecules with the ensemble model; (d) synthesizing one or more immunogenic peptide epitope molecules from the sequence data of one or more immunogenic peptide epitope molecules; and (e) screening the one or more immunogenic peptide epitope molecules synthesized for immunogenicity, wherein the molecules trigger a response by T cells that recognize the one or more immunogenic peptide epitope molecules in the context of MHC Class I or Class II molecules. In one aspect, the method further comprises selecting one or more immunogenic peptide epitope molecules that trigger a CD8 or a CD4 T cell response. In another aspect, the method generates a library of one or more immunogenic peptide epitope molecules that trigger an immune response in a host. In another aspect, the one or more peptide-level features are selected from Major Histocompatibility (MHC) Class II binding predictions, antigen strain conservation scores, and best match to a human proteome. In another aspect, the one or more antigen-level features are selected from RNA expression and subcellular localization prediction scores for each of: cytoplasm, cytoplasmic membrane, outer membrane, and extracellular. In another aspect, the MHC Class II binding is determined using NetMHCIIpan, peptide conservation for 0-3 substitutions is determined using PEPMatch, gene expression is determined using GEO Dataset, intracellular localization is determined using PSORTb, and protein existence levels using UniProt. In another aspect, the two or more machine learning algorithms are selected from Random Forest, Gradient Boosting, and XGBoost. In another aspect, the method is computer implemented. In another aspect, the model is trained with at least one of: Mycobacterium tuberculosis or Bordatella pertussis antigenic epitopes. In another aspect, the possible immunogenic peptide epitopes are viral, bacterial, fungal, protozoan, or helminthic epitopes. In another aspect, the possible immunogenic peptide epitopes are Streptococcus pneumoniae immunogenic peptide epitopes. In another aspect, the immunogenicity is in a human. In another aspect, the algorithm was used to train a model using grid search hyperparameter tuning with 10-fold cross-validation.

[0010] As embodied and broadly described herein, an aspect of the present disclosure relates to a composition comprising: one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant, or derivative thereof; a fusion protein comprising one or more amino acid sequences selected from SEQ ID NOS: 1-10; or a pool of 2 or more or more peptides comprising, consisting of, or consisting essentially of amino acid sequences selected from SEQ ID NOS: 1-10; or a polynucleotide that encodes one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant, or derivative thereof. In one aspect, the one or more peptides or proteins comprises, or wherein the fusion protein comprises two or more or more amino acid sequences selected from SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant, or derivative thereof. In another aspect, the amino acid sequence is selected from a Streptococcus pneumoniae T cell epitope selected from any one of those sequences of SEQ ID NOS: 1 to 10; and in certain examples SEQ ID NO: 1, SEQ ID NO:6, or both. In another aspect, the composition comprises: one or more Streptococcus pneumoniae peptides amino acid sequences selected from SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant, or derivative thereof; a fusion protein comprising one or more amino acid sequences selected from SEQ ID NOS: 1-10; or a pool of two or more peptides selected from SEQ ID NOS: 1-10; a polynucleotide that encodes one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant, or derivative thereof; or the peptide of SEQ ID NO:1 or SEQ ID NO:6, or both. In another aspect, the peptide or protein comprises a Streptococcus pneumoniae T cell epitope. In another aspect, the one or more peptides or proteins comprises a Streptococcus pneumoniae CD8+ or CD4+ T cell epitope. In another aspect, the Streptococcus pneumoniae T cell epitope is conserved in another Streptococcus sp. In another aspect, the one or more peptides or proteins has a length from about 9-15, 15-20, 20-25, 25-30, 30-40, 40-50, 50-75 or 75-100 amino acids. In another aspect, the one or more peptides or proteins elicits, stimulates, induces, promotes, increases, or enhances a T cell response to Streptococcus pneumoniae. In another aspect, the one or more peptides or proteins that elicits, stimulates, induces, promotes, increases, or enhances the T cell response to the Streptococcus pneumoniae is a Streptococcus pneumoniae protein or peptide, or a variant, homologue, derivative or subsequence thereof. In another aspect, the composition further comprises formulating the one or more peptides or proteins into an immunogenic formulation with an adjuvant. In another aspect, the adjuvant is selected from the group consisting of adjuvant is selected from the group consisting of alum, aluminum hydroxide, aluminum phosphate, calcium phosphate hydroxide, cytosineguanosine oligonucleotide (CpG-ODN) sequence, granulocyte macrophage colony stimulating factor (GM-CSF), monophosphoryl lipid A (MPL), poly(I:C), MF59, Quil A, N-acetyl muramyl-L-alanyl-D-isoglutamine (MDP), FIA, montanide, poly (DL-lactide-coglycolide), squalene, virosome, AS03, ASO4, IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7, IL-8, IL-10, IL-12, IL-15, IL-17, IL-18, STING, CD40L, pathogen-associated molecular patterns (PAMPs), damage-associated molecular pattern molecules (DAMPs), Freund's complete adjuvant, Freund's incomplete adjuvant, transforming growth factor (TGF)-betaantibody or antagonists, A2aR antagonists, lipopolysaccharides (LPS), Fas ligand, Trail, lymphotactin, Mannan (M-FP), APG-2, Hsp70 and Hsp90, pattern recognition receptor ligands, TLR3 ligands, TLR4 ligands, TLR5 ligands, TLR7 / 8 ligands, and TLR9 ligands. In another aspect, the composition further comprises a modulator of immune response. In another aspect, the modulator of immune response is a modulator of the innate immune response. In another aspect, the modulator is Interleukin-6 (IL-6), Interferon-gamma (IFN-y), Transforming growth factor beta (TGF-P), or Interleukin- 10 (IL-10), or an agonist or antagonist thereof.

[0011] As embodied and broadly described herein, an aspect of the present disclosure relates to a composition comprising monomers or multimers of: peptides or proteins comprising, consisting of, or consisting essentially of: one or more amino acid sequences selected from at least one of SEQ ID NOS: 1- 10; concatemers, subsequences, portions, homologues, variants, or derivatives thereof; a fusion protein comprising one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10; a polynucleotide that encodes one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant, or derivative thereof; or the peptide of SEQ ID NO:1 or SEQ ID NO:6, or both.

[0012] As embodied and broadly described herein, an aspect of the present disclosure relates to a composition comprising one or more peptide-major histocompatibility complex (MHC) monomers or multimers, wherein the peptide-MHC monomer or multimer comprises a peptide comprising, consisting of, or consisting essentially of an amino acid sequence selected from at least one of SEQ ID NOS: 1-10, in a groove of the MHC monomer or multimer.

[0013] As embodied and broadly described herein, an aspect of the present disclosure relates to a composition comprising: one or more peptides or proteins comprising, consisting of, or consisting essentially of an amino acid sequence selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof; a fusion protein comprising one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10; a pool of two or more peptides selected from at least one of SEQ ID NOS: 1-10; a polynucleotide that encodes one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof; or the peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both. In one aspect, the one or more peptides or proteins comprises, or wherein the fusion protein comprises two or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof. In another aspect, the protein or peptide comprises a Streptococcus pneumoniae T cell epitope. In another aspect, the one or more peptides or proteins comprises a Streptococcus pneumoniae CD8+ or CD4+ T cell epitope. In another aspect, the Streptococcus pneumoniae T cell epitope is not conserved in another Streptococcus pneumoniae. In another aspect, the Streptococcus pneumoniae T cell epitope is conserved in another Streptococcus pneumoniae . In another aspect, the one or more peptides or proteins has a length from about 9-15, 15-20, 20-25, 25-30, 30-40, 40-50, 50-75 or 75-100 amino acids. In another aspect, the one or morepeptides or proteins elicits, stimulates, induces, promotes, increases or enhances a T cell response to Streptococcus pneumoniae. In another aspect, the one or more peptides or proteins that elicits, stimulates, induces, promotes, increases or enhances the T cell response to Streptococcus pneumoniae is a Streptococcus pneumoniae protein or peptide, or a variant, homologue, derivative or subsequence thereof. In another aspect, the composition further comprises formulating the one or more peptides or proteins into an immunogenic formulation with an adjuvant. In another aspect, the adjuvant is selected from the group consisting of adjuvant is selected from the group consisting of alum, aluminum hydroxide, aluminum phosphate, calcium phosphate hydroxide, cytosine-guanosine oligonucleotide (CpG-ODN) sequence, granulocyte macrophage colony stimulating factor (GM-CSF), monophosphoryl lipid A (MPL), poly(I:C), MF59, Quil A, N-acetyl muramyl-L-alanyl-D-isoglutamine (MDP), FIA, montanide, poly (DL-lactide-coglycolide), squalene, virosome, AS03, ASO4, IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7, IL-8, IL-10, IL-12, IL-15, IL-17, IL-18, STING, CD40L, pathogen-associated molecular patterns (PAMPs), damage-associated molecular pattern molecules (DAMPs), Freund's complete adjuvant, Freund's incomplete adjuvant, transforming growth factor (TGF)-beta antibody or antagonists, A2aR antagonists, lipopolysaccharides (LPS), Fas ligand, Trail, lymphotactin, Mannan (M-FP), APG-2, Hsp70 and Hsp90, pattern recognition receptor ligands, TLR3 ligands, TLR4 ligands, TLR5 ligands, TLR7 / 8 ligands, and TLR9 ligands. In another aspect, the composition further comprises a modulator of immune response. In another aspect, the modulator of immune response is a modulator of the innate immune response. In another aspect, the modulator is Interleukin-6 (IL-6), Interferon-gamma (IFN-g), Transforming growth factor beta (TGF-B), or Interleukin- 10 (IL- 10), or an agonist or antagonist thereof.

[0014] As embodied and broadly described herein, an aspect of the present disclosure relates to a composition comprising monomers or multimers of: one or more peptides or proteins comprising, consisting of, or consisting essentially of: one or more Streptococcus pneumoniae amino acid sequences selected from at least one of SEQ ID NOS: 1-10, concatemers, subsequences, portions, homologues, variants or derivatives thereof; a fusion protein comprising one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10); a polynucleotide that encodes one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from at least one of SEQ ID NOS: 1-10), or a subsequence, portion, homologue, variant or derivative thereof; or the peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both.

[0015] As embodied and broadly described herein, an aspect of the present disclosure relates to a composition comprising one or more peptide-major histocompatibility complex (MHC) monomers or multimers, wherein the peptide-MHC monomer or multimer comprises a peptide comprising, consisting of, or consisting essentially of an amino acid sequence selected from at least one of SEQ ID NOS: 1-10, in a groove of the (MHC) monomer or multimer.

[0016] As embodied and broadly described herein, an aspect of the present disclosure relates to a method for detecting the presence of: (i) a Streptococcus pneumoniae or (ii) an immune response relevant to Streptococcus pneumoniae infections, vaccines or therapies, including T cells responsive to one or more Streptococcus pneumoniae peptides, comprising: providing one or more proteins or peptides for detectionof an amount or a relative amount of, and / or the activity of, and / or the state of antigen-specific T-cells; contacting a biological sample suspected of having Streptococcus pneumoniae -specific T-cells to one or more proteins or peptides for detection; and detecting an amount or a relative amount of, and / or the activity of, and / or the state of antigen-specific T-cells in the biological sample, wherein the one or more proteins or peptides for detection comprise one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10, or comprise a pool of 2 or more or more amino acid sequences set forth in any one of SEQ ID NOS: 1-10). In one aspect, the amount or a relative amount of, and / or activity of antigen-specific T-cells comprises one or more steps of identification or detection of the antigen-specific T-cells and measuring the amount of the antigen-specific T-cells. In another aspect, the one or more peptides or proteins comprises 2 or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10). In another aspect, the detecting the amount or a relative amount of, and / or activity of antigen-specific T-cells comprises indirect detection and / or direct detection. In another aspect, the method of detecting an immune response relevant to the Streptococcus pneumoniae comprises the following steps: providing an MHC monomer or an MHC multimer; contacting a population T-cells to the MHC monomer or MHC multimer; and measuring the number, activity or state of T-cells specific for the MHC monomer or MHC multimer. In another aspect, the MHC monomer or MHC multimer comprises a protein or peptide of the Streptococcus pneumoniae. In another aspect, the protein or peptide comprises a CD8+ or CD4+ T cell epitope. In another aspect, the T cell epitope is not conserved in another Streptococcus pneumoniae. In another aspect, the T cell epitope is conserved in another Streptococcus pneumoniae. In another aspect, the protein or peptide has a length from about 9-15, 15-20, 20-25, 25-30, 30-40, 40-50, 50-75 or 75-100 amino acids. In another aspect, the proteins or peptides comprise 2 or more amino acid sequences selected from any one of those sequences at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof; or the peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both. In another aspect, the method further comprises detecting the presence or amount of the one or more peptides in a biological sample, or a response thereto, which is diagnostic of a latent or active Streptococcus pneumoniae infection. In another aspect, the detecting an amount or a relative amount of, and / or the activity of, and / or the state of antigen-specific T-cells in the biological sample comprises measuring one or more of a cytokine or lymphokine secretion assay, T cell proliferation, immunoprecipitation, immunoassay, ELISA, radioimmunoassay, immunofluorescence assay, Western Blot, FACS analysis, a competitive immunoassay, a noncompetitive immunoassay, a homogeneous immunoassay a heterogeneous immunoassay, a bioassay, a reporter assay, a luciferase assay, a microarray, a surface plasmon resonance detector, a florescence resonance energy transfer, immunocytochemistry, or a cell mediated assay, or a cytokine proliferation assay. In another aspect, the method further comprises administering a treatment comprising the composition described hereinabove to the subject from which the biological sample was drawn that increases the amount or relative amount of, and / or activity of the antigen-specific T-cells.

[0017] As embodied and broadly described herein, an aspect of the present disclosure relates to a method for detecting the presence of: (i) Streptococcus pneumoniae or (ii) an immune response relevant to Streptococcus pneumoniae infections, vaccines or therapies, including T cells responsive to one or moreStreptococcus pneumoniae peptides, comprising: providing one or more proteins or peptides for detection of an amount or a relative amount of, and / or the activity of, and / or the state of antigen-specific T-cells; contacting a biological sample suspected of having Streptococcus pneumoniae -specific T-cells to one or more proteins or peptides for detection; and detecting an amount or a relative amount of, and / or the activity of, and / or the state of antigen-specific T-cells in the biological sample, wherein the one or more proteins or peptides for detection comprise one or more amino acid sequences set forth in those sequences selected from at least one of SEQ ID NOS: 1-10, or comprise a pool of 2 or more amino acid sequences set forth in those sequences selected from at least one of SEQ ID NOS: 1-10. In another aspect, detecting the amount or a relative amount of, and / or activity of antigen-specific T-cells comprises one or more steps of identification or detection of the antigen-specific T-cells and measuring the amount of the antigen-specific T-cells. In another aspect, the one or more peptides or proteins comprises 2 or more amino acid sequences selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10. In another aspect, the detecting the amount or a relative amount of, and / or activity of antigen-specific T-cells comprises indirect detection and / or direct detection. In another aspect, the method of detecting an immune response relevant to Streptococcus pneumoniae comprises the following steps: providing an MHC monomer or an MHC multimer; contacting a population T-cells to the MHC monomer or MHC multimer; and measuring the number, activity or state of T-cells specific for the MHC monomer or MHC multimer. In another aspect, the MHC monomer or MHC multimer comprises a protein or peptide of Streptococcus pneumoniae o / SEQ ID NO: 1 or SEQ ID NO:6, or both. In another aspect, the protein or peptide comprises a Streptococcus pneumoniae CD8+ or CD4+ T cell epitope. In another aspect, the Streptococcus pneumoniae T cell epitope is not conserved in another Streptococcus sp. In another aspect, the Streptococcus pneumoniae T cell epitope is conserved in another Streptococcus pneumoniae. In another aspect, the protein or peptide has a length from about 9-15, 15-20, 20-25, 25-30, 30-40, 40-50, 50-75 or 75-100 amino acids. In another aspect, the proteins or peptides comprise 2 or more amino acid sequences selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof. In another aspect, the method further comprises detecting the presence or amount of the one or more peptides in a biological sample, or a response thereto, which is diagnostic of a Streptococcus pneumoniae infection. In another aspect, detecting an amount or a relative amount of, and / or the activity of, and / or the state of antigen-specific T-cells in the biological sample comprises measuring one or more of a cytokine or lymphokine secretion assay, T cell proliferation, immunoprecipitation, immunoassay, ELISA, radioimmunoassay, immunofluorescence assay, Western Blot, FACS analysis, a competitive immunoassay, a noncompetitive immunoassay, a homogeneous immunoassay a heterogeneous immunoassay, a bioassay, a reporter assay, a luciferase assay, a microarray, a surface plasmon resonance detector, a florescence resonance energy transfer, immunocytochemistry, or a cell mediated assay, or a cytokine proliferation assay. In another aspect, the method further comprises administering a treatment comprising the composition described hereinabove to the subject from which the biological sample was drawn that increases the amount or relative amount of, and / or activity of the antigen-specific T-cells.

[0018] As embodied and broadly described herein, an aspect of the present disclosure relates to a method detecting a Streptococcus pneumoniae infection or exposure in a subject, the method comprising, consisting of, or consisting essentially of: contacting a biological sample from a subject with a composition described hereinabove; and determining if the composition elicits an immune response from the contacted cells, wherein the presence of an immune response indicates that the subject has been exposed to or infected with Streptococcus pneumoniae. In one aspect, the sample comprises T cells. In another aspect, the response comprises inducing, increasing, promoting or stimulating wti-Streptococcus pneumoniae activity of T cells. In another aspect, the T cells are CD8+ or CD4+ T cells. In another aspect, the method comprises determining whether the subject has been infected by or exposed to the Streptococcus pneumoniae more than once by determining if the subject elicits a secondary T cell immune response profde that is different from a primary T cell immune response profde . In another aspect, the method further comprises diagnosing a latent or active Streptococcus pneumoniae infection or exposure in a subject, the method comprising contacting a biological sample from a subject with a composition described hereinabove, and determining if the composition elicits a T cell immune response, wherein the T cell immune response identifies that the subject has been infected with or exposed to a Streptococcus pneumoniae. In another aspect, the method is conducted three or more days following the date of suspected infection by or exposure to a Streptococcus pneumoniae.

[0019] As embodied and broadly described herein, an aspect of the present disclosure relates to a method detecting Streptococcus pneumoniae infection or exposure in a subject, the method comprising, consisting of, or consisting essentially of: contacting a biological sample from a subject with a composition described hereinabove; and determining if the composition elicits an immune response from the contacted cells, wherein the presence of an immune response indicates that the subject has been exposed to or infected with Streptococcus pneumoniae. In another aspect, the sample comprises T cells. In another aspect, the response comprises inducing, increasing, promoting or stimulating anti- Streptococcus pneumoniae activity of T cells. In another aspect, the T cells are CD8+ or CD4+ T cells. In another aspect, the method comprises determining whether the subject has been infected by or exposed to Streptococcus pneumoniae more than once by determining if the subject elicits a secondary T cell immune response profde that is different from a primary T cell immune response profde. In another aspect, the method further comprises diagnosing a latent or active Streptococcus pneumoniae infection or exposure in a subject, the method comprising contacting a biological sample from a subject with a composition described hereinabove; and determining if the composition elicits a T cell immune response, wherein the T cell immune response identifies that the subject has been infected with or exposed to Streptococcus pneumoniae. In another aspect, the method is conducted three or more days following the date of suspected infection by or exposure to a Streptococcus pneumoniae.

[0020] As embodied and broadly described herein, an aspect of the present disclosure relates to a kit for the detection of Streptococcus pneumoniae or an immune response to Streptococcus pneumoniae in a subject comprising, consisting of or consisting essentially of: one or more T cells that specifically detect the presence of: one or more amino acid sequences selected from any one of those selected from at leastone of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof; or a fusion protein comprising one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10; or a pool of 2 or more or more peptides selected from the amino acid sequences selected from at least one of SEQ ID NOS: 1-10. In another aspect, the one or more amino acid sequences are selected from a Streptococcus pneumoniae T cell epitope selected from at least one of SEQ ID NOS: 1-10. In another aspect, the composition comprises: one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof; a fusion protein comprising one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10; a pool of 2 or more peptides selected from the amino acid sequences selected from at least one of SEQ ID NOS: 1- 10; or the peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both. In another aspect, the amino acid sequence comprises a Streptococcus pneumoniae CD8+ or CD4+ T cell epitope. In another aspect, the T cell epitope is not conserved in another Streptococcus pneumoniae. In another aspect, the T cell epitope is conserved in another Streptococcus pneumoniae. In another aspect, the fusion protein has a length from about 9-15, 15-20, 20-25, 25-30, 30-40, 40-50, 50-75 or 75-100 amino acids. In another aspect, the kit includes instruction for a diagnostic method, a process, a composition, a product, a service or component part thereof for the detection of: (i) Streptococcus pneumoniae or (ii) an immune response relevant to Streptococcus pneumoniae infections, vaccines or therapies, including T cells responsive to Streptococcus pneumoniae. In another aspect, the kit includes reagents for detecting an amount or a relative amount of, and / or the activity of, and / or the state of antigen-specific T-cells in the biological sample comprises measuring one or more of a cytokine or lymphokine secretion assay, T cell proliferation, immunoprecipitation, immunoassay, ELISA, radioimmunoassay, immunofluorescence assay, Western Blot, FACS analysis, a competitive immunoassay, a noncompetitive immunoassay, a homogeneous immunoassay a heterogeneous immunoassay, a bioassay, a reporter assay, a luciferase assay, a microarray, a surface plasmon resonance detector, a florescence resonance energy transfer, immunocytochemistry, or a cell mediated assay, or a cytokine proliferation assay. In another aspect, the kit includes reagents for determining a Human Leukocyte Antigen (HLA) profile of a subject, and selecting peptides that are presented by the HLA profile of the subject for detecting an immune response to Streptococcus pneumoniae .

[0021] As embodied and broadly described herein, an aspect of the present disclosure relates to a kit for the detection of Streptococcus pneumoniae or an immune response to Streptococcus pneumoniae in a subject comprising, consisting of or consisting essentially of: one or more T cells that specifically detect the presence of: one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof; a fusion protein comprising one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10; apool of 2 or more peptides selected from the amino acid sequences selected from at least one of SEQ ID NOS: 1-10; or the peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both. In one aspect, the one or more amino acid sequences is selected from a Streptococcus pneumoniae CD4 T cell epitope selected from at least one of SEQ ID NOS: 1-10; or both. In another aspect, the amino acid sequence comprises a Streptococcus pneumoniae CD8+ or CD4+ T cell epitope. In another aspect, the Streptococcus pneumoniae T cell epitope is not conserved in anotherStreptococcus pneumoniae. In another aspect, the Streptococcus pneumoniae T cell epitope is conserved in another Streptococcus pneumoniae. In another aspect, the fusion protein has a length from about 9-15, 15-20, 20-25, 25-30, 30-40, 40-50, 50-75 or 75-100 amino acids. In another aspect, the kit includes instruction for a diagnostic method, a process, a composition, a product, a service or component part thereof for the detection of: (i) Streptococcus pneumoniae or (ii) an immune response relevant to Streptococcus pneumoniae infections, vaccines or therapies, including T cells responsive to Streptococcus pneumoniae. In another aspect, the kit includes reagents for detecting an amount or a relative amount of, and / or the activity of, and / or the state of antigen-specific T-cells in the biological sample comprises measuring one or more of a cytokine or lymphokine secretion assay, T cell proliferation, immunoprecipitation, immunoassay, ELISA, radioimmunoassay, immunofluorescence assay, Western Blot, FACS analysis, a competitive immunoassay, a noncompetitive immunoassay, a homogeneous immunoassay a heterogeneous immunoassay, a bioassay, a reporter assay, a luciferase assay, a microarray, a surface plasmon resonance detector, a florescence resonance energy transfer, immunocytochemistry, or a cell mediated assay, or a cytokine proliferation assay. In another aspect, the kit includes reagents for determining a Human Leukocyte Antigen (HLA) profile of a subject, and selecting peptides that are presented by the HLA profile of the subject for detecting an immune response to Streptococcus pneumoniae .

[0022] As embodied and broadly described herein, an aspect of the present disclosure relates to a method of stimulating, inducing, promoting, increasing, or enhancing an immune response against a Streptococcus pneumoniae in a subject, comprising: administering a composition described hereinabove, in an amount sufficient to stimulate, induce, promote, increase, or enhance an immune response against the Streptococcus pneumoniae in the subject. In one aspect, the immune response provides the subject with protection against a Streptococcus pneumoniae infection or pathology, or one or more physiological conditions, disorders, illnesses, diseases or symptoms caused by or associated with Streptococcus pneumoniae infection or pathology. In another aspect, the immune response is specific to: one or more Streptococcus pneumoniae peptides selected from the amino acid sequences selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof.

[0023] As embodied and broadly described herein, an aspect of the present disclosure relates to a method of stimulating, inducing, promoting, increasing, or enhancing an immune response against Streptococcus pneumoniae in a subject, comprising administering a composition described hereinabove, in an amount sufficient to stimulate, induce, promote, increase, or enhance an immune response against Streptococcus pneumoniae in the subject. In one aspect, the immune response provides the subject with protection against a Streptococcus pneumoniae infection or pathology, or one or more physiological conditions, disorders, illnesses, diseases or symptoms caused by or associated with Streptococcus pneumoniae infection or pathology. In another aspect, the immune response is specific to: one or more Streptococcus pneumoniae peptides selected from the amino acid sequences set forth in those sequences selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof; or the peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both.

[0024] As embodied and broadly described herein, an aspect of the present disclosure relates to a methodof stimulating, inducing, promoting, increasing, or enhancing an immune response against Streptococcus pneumoniae in a subject, comprising: administering to a subject an amount of a protein or peptide or a polynucleotide that expresses the protein or peptide comprising, consisting of or consisting essentially of an amino acid sequence of the Streptococcus pneumoniae protein or peptide, or a variant, homologue, derivative or subsequence thereof, wherein the protein or peptide comprises at least two peptides selected from the amino acid sequences selected from at least one of SEQ ID NOS: 1-10 or a subsequence, portion, homologue, variant or derivative thereof, in an amount sufficient to prevent, stimulate, induce, promote, increase, immunize against, or enhance an immune response against Streptococcus pneumoniae in the subject. In another aspect, the immune response provides the subject with protection against Streptococcus pneumoniae infection or pathology, or one or more physiological conditions, disorders, illnesses, diseases or symptoms caused by or associated with Streptococcus pneumoniae infection or pathology.

[0025] As embodied and broadly described herein, an aspect of the present disclosure relates to a method of treating, preventing, or immunizing a subject against Streptococcus pneumoniae infection, comprising administering to a subject an amount of a protein, peptide or a polynucleotide that expresses the protein or peptide comprising, consisting of, or consisting essentially of an amino acid sequence of a Streptococcus protein or peptide, or a variant, homologue, derivative or subsequence thereof, wherein the protein or peptide comprises at least two amino acid sequences selected from at least one of SEQ ID NOS: 1-10 or a subsequence, portion, homologue, variant or derivative thereof, in an amount sufficient to treat, prevent, or immunize the subject for Streptococcus pneumoniae infection, wherein the protein or peptide comprises or consists of a Streptococcus pneumoniae T cell epitope that elicits, stimulates, induces, promotes, increases, or enhances an anti- Streptococcus pneumoniae T cell immune response. In another aspect, the one or more amino acid sequences are selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof; a fusion protein comprising one or more amino acid sequences selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10; a pool of 2 or more peptides selected from the amino acid sequences set forth in those sequences selected from at least one of SEQ ID NOS: 1-10; or the peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both. In another aspect, the anti- Streptococcus pneumoniae T cell response is a CD8+, a CD4+ T cell response, or both. In another aspect, the T cell epitope is conserved across two or more clinical isolates of Streptococcus pneumoniae or two or more circulating forms of Streptococcus pneumoniae. In another aspect, the Streptococcus pneumoniae infection is an acute infection. In another aspect, the subject is a mammal or a human. In another aspect, the method reduces Streptococcus pneumoniae bacterial titer, increases or stimulates Streptococcus pneumoniae bacterial clearance, reduces or inhibits Streptococcus pneumoniae bacterial proliferation, reduces or inhibits increases in Streptococcus pneumoniae bacterial titer or Streptococcus pneumoniae bacterial proliferation, reduces the amount of a Streptococcus pneumoniae bacterial protein or the amount of a Streptococcus pneumoniae bacterial nucleic acid, or reduces or inhibits synthesis of a Streptococcus pneumoniae bacterial protein or a Streptococcus pneumoniae bacterial nucleic acid. In another aspect, the method reduces one or more adverse physiological conditions, disorders, illness, diseases, symptoms or complications caused by or associated withStreptococcus sp. infection or pathology. In another aspect, the method improves one or more adverse physiological conditions, disorders, illness, diseases, symptoms or complications caused by or associated with Streptococcus pneumoniae infection or pathology. In another aspect, the symptom is fever or chills, cough, shortness of breath or difficulty breathing, fatigue, muscle or body aches, headache, new loss of taste or smell, sore throat, congestion or runny nose, nausea or vomiting, or diarrhea. In another aspect, the method reduces or inhibits susceptibility to Streptococcus pneumoniae infection or pathology. In another aspect, the protein or peptide, or a subsequence, portion, homologue, variant or derivative thereof, is administered prior to, substantially contemporaneously with or following exposure to or infection of the subject with Streptococcus pneumoniae. In another aspect, a plurality of Streptococcus pneumoniae T cell epitopes are administered prior to, substantially contemporaneously with or following exposure to or infection of the subject with Streptococcus pneumoniae. In another aspect, the protein or peptide, or a subsequence, portion, homologue, variant or derivative thereof is administered within 2-72 hours, 2-48 hours, 4-24 hours, 4-18 hours, or 6-12 hours after a symptom of Streptococcus pneumoniae infection or exposure develops. In another aspect, the protein or peptide, or a subsequence, portion, homologue, variant or derivative thereof is administered prior to exposure to or infection of the subject with Streptococcus pneumoniae. In another aspect, the method further comprises administering a modulator of immune response prior to, substantially contemporaneously with or following the administration to the subject of an amount of a protein or peptide. In another aspect, the modulator of immune response is a modulator of the innate immune response. In another aspect, the modulator is IL-6, IFN-y, TGF-P, or IL-10, or an agonist or antagonist thereof.

[0026] As embodied and broadly described herein, an aspect of the present disclosure relates to a method of treating, preventing, or immunizing a subject against Streptococcus pneumoniae infection, comprising administering to a subject the composition described hereinabove in an amount sufficient to treat, prevent, or immunize the subject for Streptococcus pneumoniae infection. In one aspect, the Streptococcus pneumoniae infection is an acute infection. In another aspect, the method reduces Streptococcus pneumoniae bacterial titer, increases or stimulates Streptococcus pneumoniae bacterial clearance, reduces or inhibits Streptococcus pneumoniae bacterial proliferation, reduces or inhibits increases in Streptococcus pneumoniae bacterial titer or Streptococcus pneumoniae bacterial proliferation, reduces the amount of a Streptococcus pneumoniae bacterial protein or the amount of a Streptococcus pneumoniae bacterial nucleic acid, or reduces or inhibits synthesis of a Streptococcus pneumoniae bacterial protein or a Streptococcus pneumoniae bacterial nucleic acid. In another aspect, the method reduces one or more adverse physiological conditions, disorders, illness, diseases, symptoms or complications caused by or associated with Streptococcus pneumoniae infection or pathology. In another aspect, the method improves one or more adverse physiological conditions, disorders, illness, diseases, symptoms or complications caused by or associated with Streptococcus pneumoniae infection or pathology. In another aspect, the symptom is fever or chills, cough, shortness of breath or difficulty breathing, fatigue, muscle or body aches, headache, new loss of taste or smell, sore throat, congestion or runny nose, nausea, vomiting, or diarrhea. In another aspect, the method reduces or inhibits susceptibility to Streptococcus pneumoniae infection or pathology. Inanother aspect, the composition is administered prior to, substantially contemporaneously with or following exposure to or infection of the subject with Streptococcus pneumoniae. In another aspect, the composition is administered prior to, substantially contemporaneously with or following exposure to or infection of the subject with Streptococcus pneumoniae. In another aspect, the composition is administered within 2-72 hours, 2-48 hours, 4-24 hours, 4-18 hours, or 6-12 hours after a symptom of Streptococcus pneumoniae infection or exposure develops. In another aspect, the composition is administered prior to exposure to or infection of the subject with Streptococcus pneumoniae.

[0027] As embodied and broadly described herein, an aspect of the present disclosure relates to a peptide or peptides that are immunoprevalent or immunodominant in a Streptococcus sp. bacteria obtained by a method consisting of, or consisting essentially of: obtaining an amino acid sequence of the bacteria; determining one or more sets of overlapping peptides spanning one or more bacteria antigen using unbiased selection; synthesizing one or more pools of bacterial peptides comprising the one or more sets of overlapping peptides; combining the one or more pools of bacteria peptides with Class I major histocompatibility proteins (MHC), Class II MHC, or both Class I and Class II MHC to form peptide-MHC complexes; contacting the peptide-MHC complexes with T cells from subjects exposed to the bacteria; determining which pools triggered cytokine release by the T cells; and deconvoluting from the pool of peptides that elicited cytokine release by the T cells, which peptide or peptides are immunoprevalent or immunodominant in the pool. In another aspect, the bacteria is Streptococcus sp. In another aspect, the Streptococcus sp. is Streptococcus pneumoniae. In another aspect, the immunodominant peptides are selected from 1, 2 or more peptides selected from the amino acid sequences selected from at least one of SEQ ID NOS: 1-10. In another aspect, the immunodominant peptides are selected from 1, 2 or more peptides selected from the amino acid sequences set forth in those sequences selected from at least one of SEQ ID NOS: 1-10.

[0028] As embodied and broadly described herein, an aspect of the present disclosure relates to a method of selecting an immunoprevalent or immunodominant peptide or protein of a bacteria comprising, consisting of, or consisting essentially of: obtaining an amino acid sequence of the bacteria; determining one or more sets of overlapping peptides spanning one or more bacteria antigen using unbiased selection as set forth hereinabove; synthesizing one or more pools of bacteria peptides comprising the one or more sets of overlapping peptides; combining the one or more pools of bacteria peptides with Class I major histocompatibility proteins (MHC), Class II MHC, or both Class I and Class II MHC to form peptide-MHC complexes; contacting the peptide-MHC complexes with T cells from subjects exposed to the bacteria; determining which pools triggered cytokine release by the T cells; and deconvoluting from the pool of peptides that elicited cytokine release by the T cells, which peptide or peptides are immunoprevalent or immunodominant in the pool. In one aspect, the bacteria is a Streptococcus sp. In another aspect, the Streptococcus sp. is Streptococcus pneumoniae. In another aspect, the immunodominant peptides are selected from 1, 2 or more peptides selected from the amino acid sequences selected from at least one of SEQ ID NOS: 1-10. In another aspect, the immunodominant peptides are selected from 1, 2 or more peptides selected from the amino acid sequences set forth in those sequences selected from at least one ofSEQ ID NOS: 1-10; or the peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both.

[0029] As embodied and broadly described herein, an aspect of the present disclosure relates to a polynucleotide that expresses one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10), or a subsequence, portion, homologue, variant or derivative thereof; a fusion protein comprising one or more amino acid sequences selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10; a pool of 2 or more or more peptides comprising, consisting of, or consisting essentially of amino acid sequences selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10; or the peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both.

[0030] As embodied and broadly described herein, an aspect of the present disclosure relates to a vector that comprises the polynucleotide described hereinabove. In one aspect, the vector is a bacterial vector. As embodied and broadly described herein, an aspect of the present disclosure relates to a host cell that comprises the vector described hereinabove.

[0031] As embodied and broadly described herein, an aspect of the present disclosure relates to a polynucleotide that expresses: one or more peptides or proteins comprising, consisting of, or consisting essentially of an amino acid sequence selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof; a fusion protein comprising one or more amino acid sequences selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10; a pool of 2 or more peptides selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10; or the peptide of SEQ ID NO:1 or SEQ IDNO:6, or both.

[0032] As embodied and broadly described herein, an aspect of the present disclosure relates to a vector that comprises the polynucleotide described hereinabove. In one aspect, the vector is a bacterial vector. As embodied and broadly described herein, an aspect of the present disclosure relates to a host cell that comprises the vector described hereinabove.

[0033] As embodied and broadly described herein, an aspect of the present disclosure relates to a computer-implemented method for generating a library of immunogenic peptide epitope molecules, the method comprising: receiving an electronic communication containing a set of possible immunogenic peptide epitopes for an epitope of interest from a pathogen; using a processor to select sequence data from a template antigen-binding molecule from a set of possible template antigen binding molecules, wherein the selected template is a different pathogen than the epitope of interest and wherein the set of possible template antigen-binding molecules consists of two or more known peptides that bind MHC and are presented to T cells, wherein the selecting sequence data comprises screening the set of possible template antigen-binding molecules based on at least two of: gene expression and conservation; one or more peptide-level features; and one or more antigen-level features; using two or more machine learning algorithms to train a model and combining into an ensemble model; generating sequence data of one or more immunogenic peptide epitope molecules with the ensemble model; synthesizing one or more immunogenic peptide epitope molecules from the sequence data of one or more immunogenic peptide epitope molecules;and screening the one or more immunogenic peptide epitope molecules synthesized for immunogenicity, wherein the molecules trigger a response by T cells that recognize the one or more immunogenic peptide epitope molecules in the context of MHC Class I or Class II molecules. In one aspect, the method further comprises selecting one or more immunogenic peptide epitope molecules that trigger a CD8 or a CD4 T cell response. In another aspect, the method generates a library of one or more immunogenic peptide epitope molecules that trigger an immune response in a host. In another aspect, the one or more peptide-level features are selected from Major Histocompatibility (MHC) Class II binding predictions, antigen strain conservation scores, and best match to a human proteome. In another aspect, the one or more antigenlevel features are selected from RNA expression and subcellular localization prediction scores for each of: cytoplasm, cytoplasmic membrane, outer membrane, and extracellular. In another aspect, the MHC Class II binding is determined using NetMHCIIpan, peptide conservation for 0-3 substitutions is determined using PEPMatch, gene expression is determined using GEO Dataset, intracellular localization is determined using PSORTb, and protein existence levels using UniProt. In another aspect, the two or more machine learning algorithms are selected from Random Forest, Gradient Boosting, and XGBoost. In another aspect, the method is computer implemented. In another aspect, the model is trained with at least one of: Mycobacterium tuberculosis or Bordatella pertussis antigenic epitopes. In another aspect, the possible immunogenic peptide epitopes are viral, bacterial, fungal, protozoan, or helminthic epitopes. In another aspect, the possible immunogenic peptide epitopes are Streptococcus pneumoniae immunogenic peptide epitopes. In another aspect, the immunogenicity is in a human. In another aspect, the algorithm was used to train a model using grid search hyperparameter tuning with 10-fold cross-validation.

[0034] As embodied and broadly described herein, an aspect of the present disclosure relates to a non-transitory computer-readable medium for generating a library of immunogenic peptide epitope molecules, comprising instructions stored thereon, that when executed on a processor, perform the steps of: receiving an electronic communication containing a set of possible immunogenic peptide epitopes for an epitope of interest from a pathogen; using a processor to select sequence data from a template antigen-binding molecule from a set of possible template antigen binding molecules, wherein the selected template is a different pathogen than the epitope of interest and wherein the set of possible template antigen-binding molecules consists of two or more known peptides that bind MHC and are presented to T cells, wherein the selecting sequence data comprises screening the set of possible template antigen-binding molecules based on at least two of: gene expression and conservation; one or more peptide-level features; and one or more antigen-level features; using two or more machine learning algorithms to train a model and combining into an ensemble model; generating sequence data of one or more immunogenic peptide epitope molecules with the ensemble model; synthesizing one or more immunogenic peptide epitope molecules from the sequence data of one or more immunogenic peptide epitope molecules; and screening the one or more immunogenic peptide epitope molecules synthesized for immunogenicity, wherein the molecules trigger a response by T cells that recognize the one or more immunogenic peptide epitope molecules in the context of MHC Class I or Class II molecules. In another aspect, the method further comprises selecting one or more immunogenic peptide epitope molecules that trigger a CD8 or a CD4 T cell response. In anotheraspect, the method generates a library of one or more immunogenic peptide epitope molecules that trigger an immune response in a host. In another aspect, the one or more peptide-level features are selected from Major Histocompatibility (MHC) Class II binding predictions, antigen strain conservation scores, and best match to a human proteome. In another aspect, the one or more antigen-level features are selected from RNA expression and subcellular localization prediction scores for each of: cytoplasm, cytoplasmic membrane, outer membrane, and extracellular. In another aspect, the MHC Class II binding is determined using NetMHCIIpan, peptide conservation for 0-3 substitutions is determined using PEPMatch, gene expression is determined using GEO Dataset, intracellular localization is determined using PSORTb, and protein existence levels using UniProt. In another aspect, the two or more machine learning algorithms are selected from Random Forest, Gradient Boosting, and XGBoost. In another aspect, the method is computer implemented. In another aspect, the model is trained with at least one of: Mycobacterium tuberculosis or Bordatella pertussis antigenic epitopes. In another aspect, the possible immunogenic peptide epitopes are viral, bacterial, fungal, protozoan, or helminthic epitopes. In another aspect, the possible immunogenic peptide epitopes are Streptococcus pneumoniae immunogenic peptide epitopes. In another aspect, the immunogenicity is in a human. In another aspect, the algorithm was used to train a model using grid search hyperparameter tuning with 10-fold cross-validation.BRIEF DESCRIPTION OF THE DRAWINGS

[0035] For a more complete understanding of the features and advantages of the present invention, reference is now made to the detailed description of the invention along with the accompanying figures and in which:

[0036] FIG. 1 shows the receiver operating characteristic (ROC) performance of immunogenicity prediction by the trained models on the testing data.

[0037] FIGS. 2A and 2B show the feature importance by permutation for the ensemble model after training. FIG. 2A) Each feature used for the training of each machine learning model is evaluated by the loss of ROC-AUC performance when the feature's values were permuted. Higher bars indicate greater importance to the model's predictive ability. FIG. 2B) Feature categories were combined to show the overall contribution of broader groups, highlighting which biological factors had the most influence on model performance with gene expression having the greatest impact followed by conservation and subcellular localization.

[0038] FIG. 3 shows the performance of ensemble model predicting active TB epitopes at various donor thresholds.

[0039] FIG. 4 shows the performance of ensemble model predicting Bordetella pertussis epitopes at various donor thresholds.FIGS. 5A to 5E show an overview of Streptococcus pneumoniae feature characteristics among peptides predicted to be immunogenic and non-immunogenic using the ensemble model. FIG. 5A) The highest (n=130) and lowest (n=130) .S'. pneumoniae peptide probability scores were used to select ‘predicted immunogenic’ and ‘predicted non-immunogenic peptides’. As expected, peptides predicted to beimmunogenic had. FIG. 5B) lower protein existence levels (which means better evidence for existence), FIG. 5C) higher gene expression, FIG. 5D) higher conservation across strains (no amino acid substitutions), FIG. 5E) and higher predicted MHC-II binding compared to peptides with a low ensemble probability score compared to peptides predicted to be non-immunogenic based on the ensemble model.

[0040] FIGS. 6A to 6C show peptides with high predicted immunogenic potential elicit a higher magnitude IFNy response across a higher frequency of participants. FIG. 6A) Cytokine production was measured by IFNy Fluorospot following stimulation with DMSO (negative control), PHA (positive control), megapool containing all 260 peptides (MP) and 10-peptide pools (P1-P26). Each data point represents the number of SFC per well (2xl05cells) for each participant (n=20), with bars indicating means with standard deviations. Responses were considered positive if the following 3 criteria were met: 1) number of SFC above background was >20, 2) a greater than 2-fold increase above background, 3) p < 0.05 by independent one-tailed t-test and Poisson distribution test when comparing stimuli to negative control. A maximum cutoff of 500 SFC / 2xlO5was utilized. Any responses that did not meet the 3 criteria for positivity, were considered IFNy negative and plotted as 0 SFC / 2xlO5cells. FIG. 6B) The number of participants with a positive IFNy response following stimulation. Each dot represents the average SFC across triplicates. FIG. 6C) The response frequency was calculated by dividing the number of participants with a response by the total number of participants screened. The peptide pools were tested as individual 10-peptide pools and aggregated based on any response to Pl -Pl 3 or P14-P26. The megapool was tested as a single pool of 260 peptides.

[0041] FIGS. 7A to 7C show the deconvolution of P7 into its 10 peptide components reveals two T cell epitopes with similar features and sequences . FIG.7A). The number of participants with a positive IFNy response following stimulation with their respective peptides (SEQ ID NOS: 1-10), P7, and the megapool (MP). The peptide sequences corresponding to peptides 1-10 are provided, with peptide 1 (SEQ ID NO: 1) and 6 (SEQ ID NO:6) bolded to highlight their high performance (6 / 9 participant responses) and underlined to highlight high sequence similarity (73%). FIG.7B) The number of epitopes recognized by the P7-positive participants (n=9) are visualized with a violin plot. FIG.7C) Several ensemble model prediction features (MHC-II binding, conservation and gene expression) are compared for peptide 1 (SEQ ID NO: 1), 5 (SEQ ID NO:5) and 6 (SEQ ID NO:6), along with their probability scores and UniProt specifications (Start / Stop, UniprotID, Gene).

[0042] FIGS. 8A and 8B show ROC curves of combined model on test set and the feature importance by permutation. FIG. 8A) Performance evaluated as ROC-AUC for each of the three models and ensemble on the withheld set using the combined model trained on both Mtb and B. pertussis data. The ensemble model achieved the highest ROC-AUC of 0.83 outperforming each individual model. FIG. 8B) The feature importance by feature category. Gene expression in the combined model had the greatest impact similar to the original Mtb model, however, conservation and subcellular localization had more impact in this combined model.DETAILED DESCRIPTION OF THE INVENTION

[0043] While the making and using of various embodiments of the present invention are discussed in detail below, it should be appreciated that the present invention provides many applicable inventive concepts that can be embodied in a wide variety of specific contexts. The specific embodiments discussed herein are merely illustrative of specific ways to make and use the invention and do not delimit the scope of the invention. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the disclosure described herein may be employed in practicing the invention.

[0044] To facilitate the understanding of this invention, a number of terms are defined below. Terms defined herein have meanings as commonly understood by a person of ordinary skill in the areas relevant to the present invention. Terms such as “a”, “an” and “the” are not intended to refer to only a singular entity, but include the general class of which a specific example may be used for illustration. The terminology herein is used to describe specific embodiments of the invention, but their usage does not delimit the invention, except as outlined in the claims. Unless specifically stated or obvious from context, as used herein, the term “or” is understood to be inclusive.

[0045] Previously, the inventors have identified T cell epitopes which are exclusively recognized by individuals with active tuberculosis (TB) infection and not in those that are TB neg or those with latent TB infection. The present invention includes novel methods for determining antigenic epitopes from new pathogens by training a machine learning model with data from a Mycobacterium tuberculosis (Mtb) dataset and Bordetella pertussis dataset to develop an ensemble model. The ensemble model was then applied to Streptococcus pneumoniae sequences. The resulting peptide sequences were then synthesized as predicted by the ensemble model to be immunogenic or non-immunogenic. Next, ex vivo testing with PBMCs from healthy participants showed that peptides predicted to be immunogenic elicited significantly higher I FNy responses compared to non-immunogenic peptides, validating the model for the Streptococcus pneumoniae sequences.

[0046] T cell epitope predictions have been successfully used to identify candidate peptides from infectious agents, allergens, and cancer cells (2). These predictions have predominantly relied on the features of the peptides themselves, especially their predicted MHC binding affinity. When the goal is to identify peptides likely to be immunogenic in the general population, predictions of MHC class II binding can narrow down the number of peptide candidates to the top -20% (3). However, in the case of complex bacterial pathogens expressing thousands of antigens, this will still leave tens- to hundreds of thousands of possible candidate peptides.

[0047] To address this problem, the inventors investigated the development of prediction methods that integrate the properties of specific peptides, such as their binding affinity, with properties of the antigens that they are derived from. For example, the inventors and others have shown that the expression level of an antigen impacts the likelihood that peptides contained in it will be recognized by T cells (4-6). Additionally, the inventors found that certain antigen properties can impact their immunogenicity, such asproteins secreted by or contained within type 7 secretion systems (T7SS) being the main target of immune responses in individuals who are IGRA+ (7). At the peptide level, the inventors observed that beyond MHC binding, the degree of conservation of a peptide in related antigens can increase the likelihood of T cell recognition while conservation in the host can reduce it (8,9). These combined findings were then incorporated in a unified prediction method for bacterial pathogens using machine learning approaches.

[0048] In this study, the inventors describe the peptide prediction model features and report the contribution of each feature using an existing Mtb dataset for training (7). Next, the inventors assessed the versatility of the prediction model by applying it to an existing dataset from patients with active TB and a Bordetella pertussis dataset. To further demonstrate the model’s ability to successfully discriminate between immunogenic and non-immunogenic peptides, the inventors synthesized peptides with the highest and lowest predicted values and tested for immunogenicity ex vivo in PBMCs from healthy individuals. The inventors validated the model experimentally using Streptococcus pneumoniae, the most common cause of bacterial pneumonia (10), and PBMCs from healthy individuals. Considering that most of the general population has been exposed to .S', pneumoniae by adulthood (10), PBMCs from healthy controls were used to screen for immune responses against synthesized peptides which were predicted either immunogenic or non-immunogenic for comparison. Finally, the Mtb and Bordetella pertussis datasets were combined to train a combined model to be used for broader CD4+ T cell epitope prediction against bacterial pathogens.

[0049] Model Development and Cross-Validation on a Mtb dataset. A dataset of 20,216 peptides derived from Mtb was used that was systematically tested for T cell recognition in Mtb infected participants (IGRA+; see Methods and (7)). That peptide set was created to represent each Mtb protein with 2-10 peptides, and prioritized peptides in each protein that had high predicted promiscuous MHC class II binding based on methods available at the time of selection. A total of 144 peptides were recognized by at least two participants, and these were considered as positive responses in training the machine learning models, while the remainder of peptides were considered as negatives.

[0050] For each peptide in the set, the inventors calculated six features on the peptide level, namely MHC II binding predictions (7-allele score (3)), Mtb strain conservation scores (allowing 0-3 amino acid substitutions with one score for each threshold), and the best match to the human proteome. The inventors also calculated six features on the protein level, namely RNA expression as TMP values, subcellular localization prediction scores (a score for each component: cytoplasm, cytoplasmic membrane, outer membrane, and extracellular) using PSORTb (11), and protein existence levels as curated by UniProt (Table 1). These features were selected based on prior publications showing their importance in immune recognition.

[0051] Table 1. Features collected for Mtb peptides at the peptide level and antigen level.Eevel Feature Method / SourcePeptide MHC-Binding 7-allele score NetMHCIIpanPeptide Mtb strain conservation (exact match) PEPMatchPeptide Mtb strain conservation ( 1 substitution) PEPMatchPeptide Mtb strain conservation (2 substitutions) PEPMatchPeptide Mtb strain conservation (3 substitutions) PEPMatchPeptide Best human homology percentage PEPMatchProtein RNA Expression (TPM) GEO DatasetProtein Cytoplasm localization score PSORTbProtein Cytoplasmic membrane localization score PSORTbProtein Outer membrane localization score PSORTbProtein Extracellular localization score PSORTbProtein Protein Existence Level UniProt

[0052] The inventors applied three different machine learning algorithms: XGBoost, Gradient Boosting, and Random Forest. Each algorithm was used to train a model using grid search hyperparameter tuning with 10-fold cross-validation. Model performance was evaluated using the area under curve (AUC) for the Receiver Operating Characteristic (ROC). When evaluating performance on the test data, the Random Forest model achieved an AUC of 0.88 and an AUC0.1 of 0.81. The Gradient Boosting model yielded an AUC of 0.85 and an AUC0.1 of 0.79, while the XGBoost model demonstrated an AUC of 0.90 and an AUC0.1 of 0.78. The ensemble model, which combined predictions from all three methods by averaging probabilities, further improved predictive performance, achieving an AUC of 0.91 and an AUC0.1 of 0.82. The ROC curves with AUC values of each model on the test data are shown in FIG. 1.

[0053] FIG. 1 shows the receiver operating characteristic (ROC) performance of immunogenicity prediction by the trained models on the testing data.

[0054] Feature Importance Analysis. To determine how different features contribute to the predictive performance of the model, the inventors implemented a permutation approach where each feature has its values randomly shuffled one at a time, and the decrease in model performance as measured by ROC-AUC was assessed. This method was applied individually to each of the three algorithms: Random Forest, Gradient Boosting, and XGBoost, and the results for each feature were aggregated (FIG. 2A). By far the most impactful feature was the gene expression associated with each peptide antigen. Other features, such as MHC-II binding prediction and conservation (exact match across the 15-mer peptide and allowing 2 amino acid substitutions), also demonstrated a significant effect on performance. Conversely, certainfeatures like protein existence levels, subcellular localization, and best homologous match to the human host had minimal impact. The random control feature scored the lowest, indicating that all the rest of the collected features had some effect on the model performance, even if the effect was minimal.

[0055] A disadvantage of the permutation approach for the individual feature importance evaluation is that correlated features such as the conversation scores 0, 1, 2 and 3 and the different subcellular localization scores cannot be properly evaluated, as leaving one such feature out will allow the model to compensate for it with the information in others. The inventors also combined all features from the same category, and left them out in a group for the permutation to assess their overall contribution to model performance (FIG. 2B). This showed that gene expression level was the leading feature, followed by the combined conservation features, and the combined subcellular localization features.

[0056] FIGS. 2A and 2B show the feature importance by permutation for the ensemble model after training. FIG. 2A) Each feature used for the training of each machine learning model is evaluated by the loss of ROC-AUC performance when the feature's values were permuted. Higher bars indicate greater importance to the model's predictive ability. FIG. 2B) Feature categories were combined to show the overall contribution of broader groups, highlighting which biological factors had the most influence on model performance with gene expression having the greatest impact followed by conservation and subcellular localization.

[0057] Model Performance on Independent Dataset Derived from patients with Active TB. The inventors took advantage of the fact that the same set of TB peptides originally tested for recognition in healthy IGRA+ participants were recently tested in a separate set of patients with active TB (see (12) and Methods). The inventors applied the ensemble model described above that was trained using Mtb data, and tried to predict what peptides would be positive in the active TB dataset. For the positivity threshold, the inventors started by considering peptides as positives that were recognized by at least two participants, which included only 22 peptides. This lower number compared to the first dataset can be partly explained by the lower number of participants tested, but could also be due to the increased ‘noise’ when working with blood samples from individuals with active infection. The model's predictive accuracy to distinguish the 22 peptides categorized as positive from negatives achieved an AUC of 0.79 (FIG. 3), which is an outstanding performance on an independent dataset.

[0058] Given the much lower number of ‘positive’ peptides in the active TB peptide dataset, the inventors varied the threshold for what is considered ‘positive’ from 1-3 participants, and evaluated the model’s predictive performance using AUC values. The validation of the ensemble model disclosed herein against active TB data revealed a progressive increase in predictive accuracy with higher participant recognition thresholds. At the initial threshold of recognition by at least one participant, there were 138 positive peptides, and the model achieved an AUC of 0.65. For peptides recognized by three or more participants, only seven peptides reached the positive threshold, and the AUC value of the model was 0.95, indicating an extremely high predictive capability.

[0059] FIG. 3 shows the performance of ensemble model predicting active TB epitopes at various donorthresholds. As the threshold for positive recognition increased from 1 to 3 participants, the model's AUC values progressively improved, reaching a high of 0.95 at the 3-participant threshold.

[0060] Model Performance on B. pertussis Data. While the predictive performance of the model trained on the Mtb peptide dataset on the independent active TB dataset was impressive, those datasets are linked by sharing the exact same peptides from the same pathogen. Thus, the inventors wanted to test if the model trained on the Mtb peptide dataset would have predictive power for a completely independent pathogen: Bordetella pertussis, the causative agent of whooping cough. A dataset of 24,294 peptides derived from Bordetella pertussis was utilized which had been screened in 20 participants (as described in Methods and (13)).

[0061] Given the variance in performance of the model based on the positivity threshold the inventors described above, the inventors assessed model predictions using ROC-AUC using variable thresholds, ranging from one to four participants. At the threshold of recognition by at least one participant, the ensemble model was tested against 2,939 positive peptides, yielding an AUC of 0.51. As the participant recognition threshold increased to two, the model was evaluated with 412 positive peptides and achieved an AUC of 0.57. With peptides recognized by at least three participants, the model's performance was tested against 79 peptides, resulting in an AUC of 0.76. This pattern of assessment continued as the threshold increased, with the highest AUC achieved of 0.82 for the 32 peptides recognized by at least four participants. The trend observed in the model's performance on the Bordetella pertussis dataset shows the model is more adept at identifying peptides with broad recognition among the participant population as opposed to peptides only recognized at the individual level. The ROC curves and AUC values of the ensemble model prediction at each participant threshold is shown in FIG. 4.

[0062] FIG. 4 shows the performance of ensemble model predicting Bordetella pertussis epitopes at various donor thresholds. The model shows progressively improved AUC values as the participant recognition threshold increases, with the highest AUC of 0.82 achieved at the four-participant threshold. This suggests the model is better at predicting peptides with broader recognition across multiple participants.

[0063] Model Performance to Identify Immunogenic .S', pneumoniae Peptides in a Prospective Fashion. Next, the inventors wanted to test in a prospective fashion, if the methods the inventors have developed are capable of identifying immunogenic peptides from an independent pathogen when starting from the entire proteome better then when using MHC presentation predictions alone. .S', pneumoniae was selected as a target, since it is the most common cause of community acquired pneumonia (CAP) and most adults have been exposed to the pathogen through infection or colonization during their lifetime (10). The inventors wanted to specifically test the ability of the ensemble model to select peptides that are immunogenic from .S', pneumoniae, over what would be selected based on MHC binding predictions alone. The peptides with the highest and lowest ensemble probability scores were classified into one of two groups for peptide synthesis: predicted to be immunogenic (n=130) and predicted to be non-immunogenic (n=130). As expected, the predicted immunogenic peptides had lower (=better) protein existence scores, and higherconservation, gene expression, and MHC-II binding scores compared to peptides with low predicted immunogenicity (FIGS. 5 A to 5E). Despite taking MHC-II binding into account prior to prediction to maximize the impact of the other features for peptide selection, peptides with higher ensemble probability scores achieved higher MHC-II scores.

[0064] FIGS. 5A to 5E show an overview of Streptococcus pneumoniae feature characteristics among peptides predicted to be immunogenic and non-immunogenic using the ensemble model. FIG. 5 A) The highest (n=130) and lowest (n=130) .S', pneumoniae peptide probability scores were used to select ‘predicted immunogenic’ and ‘predicted non-immunogenic peptides’. As expected, peptides predicted to be immunogenic had. FIG. 5B) lower protein existence levels (which means better evidence for existence), FIG. 5C) higher gene expression, FIG. 5D) higher conservation across strains (no amino acid substitutions), FIG. 5E) and higher predicted MHC-II binding compared to peptides with a low ensemble probability score compared to peptides predicted to be non-immunogenic based on the ensemble model.

[0065] The 130 highest and 130 lowest probability scored peptides were synthesized for screening in PBMCs from healthy participants (n=20). Peptides were combined into 26 pools of 10 peptides (P1-P26) each to minimize the number of cells required for initial screening. IFNy, one of the most prevalent pro- inflammatory cytokines, was measured using Fluorospot assay following stimulation with peptide pools and all responses were recorded. The peptide pools were tested alongside a megapool (MP) containing all 260 peptides, a negative control, and a positive control (FIG. 6A). Predicted immunogenic pools (P1-P13) had a higher magnitude of IFNy responses compared to predicted non-immunogenic pools (P14-P26). The megapool and a single 10-peptide pool (P7), were capable of eliciting a strong IFNy response. Out of the 20 participants, P7 elicited a positive IFNv signal in 9 participants, compared to single participant responses in P3, P12, P14, P15 and P22 (FIG. 6B). The megapool was recognized by 10 / 20 (50%) participants, with the majority coming from the predicted immunogenic peptides, suggesting that this comparably small number of 260 peptides already gives a good response compared to e.g. the >20,000 tested for Mtb.

[0066] To further evaluate tire ability of the prediction model to discriminate between immunogenic and non-immunogenic peptides, the inventors compared the overall response frequency of P1-P13 and P14-P26 (FIG. 6C). Half of all participants responded to at least one of the predicted immunogenic pools (Pl-P13), compared to a 5% response frequency (single responder) among the predicted non-immunogenic pools. Furthermore, the inventors were able accomplish the same response frequency (50%) with half the amount of peptides (n== 130) by restricting to those with the highest ensemble probability scores. The ensemble prediction model was successful in identifying immunogenic peptides, showing high response frequencies among the top-scoring peptides, demonstrating its potential for efficient peptide selection.

[0067] FIGS. 6A to 6C show peptides with high predicted immunogenic potential elicit a higher magnitude IFN response across a higher frequency of participants. FIG. 6A) Cytokine production was measured by IFNy Fluorospot following stimulation with DMSO (negative control), PHA (positive control), megapool containing all 260 peptides (MP) and 10-peptide pools (P1-P26). Each data pointrepresents the number of SFC per well (2xl05cells) for each participant (n=20), with bars indicating means with standard deviations. Responses were considered positive if the following 3 criteria were met: 1) number of SFC above background was >20, 2) a greater than 2-fold increase above background, 3) p < 0.05 by independent one-tailed t-test and Poisson distribution test when comparing stimuli to negative control. A maximum cutoff of 500 SFC / 2xlO5was utilized. Any responses that did not meet the 3 criteria for positivity, were considered IFNy negative and plotted as 0 SFC / 2xlO5cells. FIG. 6B) The number of participants with a positive IFNy response following stimulation. Each dot represents the average SFC across triplicates. FIG. 6C) The response frequency was calculated by dividing the number of participants with a response by the total number of participants screened. The peptide pools were tested as individual 10-peptide pools and aggregated based on any response to Pl -Pl 3 or P14-P26. The megapool was tested as a single pool of 260 peptides.

[0068] Given the dominant response to tire P7 pool, the inventors wanted to determine which specific peptides are the targeted by measuring the response against each of the 10 contained peptides via IFNy Fluorospot (FIG. 7A). Three individual peptides were recognized among participants that responded to P7, labeled as (peptide) 1 (SEQ ID NO:1), 5 (SEQ ID NO:5), and 6 (SEQ ID NO:6). Peptide 1 was recognized in 8 / 9 participants, and peptide 5 and 6 were recognized in 1 / 9 and 6 / 9 participants, respectively. (FIG. 7B). Peptide 1 and 6 demonstrated high sequence homology (73%, 11 / 15), and although encoded in different genes, the proteins they are derived from are both “ABC transporter domain-containing” proteins. Given that the start and stop index of these peptides are also nearly identical, it is likely that they come from the same protein domain. Notably, the peptides shared identical ensemble probability scores, protein existence levels, and MHC-II expression scores, while peptide 6 had higher conservation (exact match) and gene expression values, respectively (FIG. 7C). This deconvolution methodology illustrates that the ensemble model can be applied translationally for bacterial epitope discover}', and thus, pathogenic epitope discovery’, such as v iruses, bacterial, fungi, protozoans, and helminths.

[0069] FIGS. 7A to 7C show the deconvolution of P7 into its 10 peptide components reveals two T cell epitopes with similar features and sequences. FIG.7A) The number of participants with a positive IFNy response following stimulation with their respective peptides (SEQ ID NOS: 1-10), P7, and the megapool (MP). The peptide sequences corresponding to peptides 1-10 are provided, with peptide 1 (SEQ ID NO: 1) and 6 (SEQ ID NO:6) bolded to highlight their high performance (6 / 9 participant responses) and underlined to highlight high sequence similarity (73%). FIG.7B) The number of epitopes recognized by the P7-positive participants (n=9) are visualized with a violin plot. FIG.7C) Several ensemble model prediction features (MHC-II binding, conservation and gene expression) are compared for peptide 1 (SEQ ID NO: 1), 5 (SEQ ID NO:5) and 6 (SEQ ID NO:6), along with their probability scores and UniProt specifications (Start / Stop, UniprotID, Gene).

[0070] Combined Mtb and B. pertussis Model. Since the inventors have demonstrated the predictive capabilities of machine learning models on individual bacterial datasets, the inventors next trained a unified model on the combined data from Mtb and B. pertussis. This is aimed at providing a generalized prediction tool capable of identifying immunogenic CD4+ T cell epitopes across different bacterial species, thusenhancing the utility of the method disclosed herein for broad-spectrum vaccine development.

[0071] The datasets from both Mtb and Bordetella pertussis combined to comprise 44,510 peptides — 20,216 from Mtb and 24,294 from B. pertussis. Each peptide was annotated with the same set of features as in the individual models, ensuring consistency across the combined dataset. Positive peptides were defined as those recognized by at least two participants in their respective cohorts, yielding a total of 556 positive peptides. The Random Forest, Gradient Boosting, and XGBoost algorithms were retrained on this combined dataset using the same hyperparameter tuning and cross-validation procedures outlined previously.

[0072] The ensemble model, created by averaging the probabilities from the three algorithms, achieved an ROC-AUC of 0.83 and AUC0.1 of 0.72 on the combined test data (FIG. 8A). Feature importance analysis revealed that gene expression still had the greatest impact on the model prediction, but conservation and subcellular localization had more impact than the original Mtb model did (FIG. 8B).

[0073] FIGS. 8A and 8B show ROC curves of combined model on test set and the feature importance by permutation. FIG. 8A) Performance evaluated as ROC-AUC for each of the three models and ensemble on the withheld set using the combined model trained on both Mtb and B. pertussis data. The ensemble model achieved the highest ROC-AUC of 0.83 outperforming each individual model. FIG. 8B) The feature importance by feature category. Gene expression in the combined model had the greatest impact similar to the original Mtb model, however, conservation and subcellular localization had more impact in this combined model.

[0074] T cell Epitope Data. A total of 20,216 peptides, all 15-mers, derived from a proteome-wide screening of Mycobacterium tuberculosis using PBMCs from healthy IGRA+ participants (7) was used for training. The peptides were derived from 21 Mtb genomes, parsing out all possible 15-mers from their protein sequences, with a total of 1,568,148 peptides. The initial peptide selection from this study was based on MHC class II binding predictions across 22 different HLA DR, DP, and DQ class II alleles using the IEDB consensus method.

[0075] These 20,216 peptides were screened for recognition using an IFNy ELISpot assay in 28 IGRA+ healthy individuals. The peptides were also screened for recognition by an additional 21 individuals with active tuberculosis who were mid-treatment (12). This active TB dataset was used for validation of the Mtb-trained models. A total of 420 peptides were recognized by at least one IGRA+ participant, and 138 were recognized by at least one patient with active TB. For stringency, during model training, the inventors considered a positive peptide to be recognized by at least two separate participants, resulting in 144 peptides and 22 peptides classified as positive for the IGRA+ and active TB datasets, respectively. Model performance on the active TB validation set was evaluated using ROC-AUC to assess how well the trained models distinguished positive peptides from negatives across one to three participant recognition thresholds.

[0076] For independent model validation on Bordetella pertussis peptides, the inventors used an independent dataset consisting of 24,294 Bordetella pertussis peptides, which were screened in PBMCsfrom 20 participants for recognition (13). The dataset included the peptide sequence and number of participant responses. All features were collected the same way as the IGRA+ and active TB dataset. To quantify the model performance on this dataset, the inventors conducted ROC-AUC analysis at varying recognition thresholds, ranging from one to four participants, similar to how the inventors evaluated the active TB dataset.

[0077] Additionally, the trained ensemble model was used to predict the most immunogenic peptides in Streptococcus pneumoniae. All possible 15-mer residues (n=558,000) were derived from the .S'. pneumoniae (strain ATCC BAA-255 / R6) proteome (UniProt proteome ID: UP000000586) and the corresponding features were collected for each peptide. A single peptide from each protein was selected based on the highest MHC-II binding prediction scores, to maximize the influence of the other features in the model while still accounting for MHC-II binding.

[0078] Feature Collection. The inventors collected multiple features that can be attributed to each Mycobacterium tuberculosis peptide for model training. This included MHC Class II binding predictions, RNA expression levels, conservation across Mtb strains and human proteins, subcellular localization, and protein existence scores. Each feature associated with the peptides was quantified, resulting in numeric values for all data points.

[0079] MHC Class II Binding Predictions. The NetMHCIIpan 4.1 EL (eluted ligand) method was used to predict MHC binding across seven alleles (HLA-DRB 1*03:01, HLA-DRB 1*07:01, HLA-DRB 1*15:01, HLA-DRB3 * 01:01, HLA-DRB3 * 02 : 02, HLA-DRB4* 01:01, HLA-DRB5 * 01:01), which have been shown to represent a large HLA coverage of the human population. A “7-allele” score was calculated for each peptide by averaging the binding scores (ranging [0, 1]) for each allele.

[0080] RNA Expression. RNA expression data for individual Mtb peptides were obtained from a GEO dataset: Series GSE229680 (14). This dataset includes high-throughput sequencing expression profdes of Mtb, generated using the Illumina NextSeq 2000 platform. The gene expression values for the H37Rv Mtb strain cultured under normal conditions, as a control, were used. For each gene in the dataset, three control values for H37Rv, measured as transcripts per million (TPM) values, were averaged to obtain a single expression value per gene. Subsequently, using the gene symbols from which each peptide was derived, these averaged TPM values were mapped to their respective peptides. For Bordetella pertussis and Streptococcus pneumoniae, RNA expression data was sourced similarly from GEO Series GSE145049 (15) and GSE77748 (16), respectively, following the same approach for mapping averaged TPM values to their respective peptides.

[0081] Conservation in Other Mtb Strains. PEPMatch (17), apeptide-matchingtool, was used to determine the conservation of each peptide across various other Mtb strains. A total of 37 separate Mtb proteomes (supplementary material) were selected from UniProt (18) fortheir significance and non-redundancy. Each proteome was preprocessed by PEPMatch using k=3. Then, each peptide was searched across these proteomes, allowing up to three residue substitutions, corresponding to 80% homology for the 15-mer peptides. Four total conservation scores were calculated for each homology threshold (100%, 93%, 87%,and 80%) by totaling the number of matches at each threshold and dividing by the number of proteomes searched.

[0082] Conservation in Human Host. Homology to the human host was also assessed using PEPMatch. The human proteome was downloaded from UniProt (ID: UP000005640), which included protein isoforms. Given the low conservation of these peptides in human proteins, each peptide was searched with a threshold of up to seven residue substitutions. The highest homology percentage for a given peptide match was then used as the score for this feature, which ranged from 100% to 53%. If no match was found for a peptide up to seven substitutions, the next homology percentage at eight substitutions (47%) was given.

[0083] Subcellular Localization. Subcellular localization of peptides was assessed using PSORTb (11), a bioinformatics tool designed for predicting the localization of proteins within different cellular compartments. The provided Docker container was used to run the tool. Each protein sequence from which the Mtb peptides are derived was passed as the input to PSORTb. In addition, PSORTb requires a Gram staining indication, and Gram-negative was chosen for every Mtb protein. The tool categorizes proteins into four main subcellular locations: cytoplasm, cytoplasmic membrane, outer membrane, and extracellular space, and assigns a score ranging from 0 to 10 for each. A higher score indicates a greater likelihood that the protein, and by extension, the peptide derived from it, is localized in that particular cellular compartment.

[0084] Protein Existence Levels. Each protein within the UniProt database is curated with a level ranging from 1 to 5 based on the evidence supporting its existence, from direct protein sequencing to bioinformatic predictions based on genomic data. Each Mtb peptide in the study was assigned the protein existence (PE) score of its corresponding protein from UniProt using its API and a custom Python script.

[0085] Model Training. Three machine learning methods, Random Forest, Gradient Boosting, and XGBoost, were used to train separate models on the Mtb peptides and their accompanying features. The Gradient Boosting and Random Forest models were trained using the scikit-leam library, and the XGBoost model was trained separately with its independent Python bindings. Standard hyperparameter tuning using grid search with 10-fold cross-validation was used for each model. The best estimator from each grid search was taken for the final model from each method. For Random Forest, hyperparameters included 'n_estimators' [100, 200, 300], 'max_depth' [None, 10, 20], 'min_samples_splif [2, 5, 10], 'min_samples_leaf [1, 2, 4], and 'max_features' ['auto', 'sqrt']. For Gradient Boosting, the grid search covered 'n_estimators' [100, 200, 300], 'leaming_rate' [0.01, 0.05, 0.1, 0.2], 'max_depth' [3, 5, 7], and 'subsample' [0.5, 0.75, 1.0], The XGBoostmodel's hyperparameters included 'n_estimators' [100, 200, 300], 'leaming_rate' [0.01, 0.1, 0.2], 'max_depth' [3, 5, 7], 'subsample' [0.8, 1], 'colsample_bytree' [0.8, 1], 'lambda' [1, 1.5, 2], 'alpha' [0, 0.5, 1], and 'scale _pos_wcighf [1, ratio of negative to positive samples]. After training was performed for all three methods, an ensemble model was created by averaging the probabilities from each model prediction.

[0086] Model Evaluation. The models were evaluated on the withheld test data using ROC-AUC values (19) as well as AUC0.1, which is the area under the ROC curve from 0-10% false positive rate. Thesemeasures were also calculated for the ensemble model, combining the probabilities of the trained models. The evaluation compared the effectiveness in predicting peptide immunogenicity using the recognition by different numbers of total participants.

[0087] Feature Importance. To assess the contribution of each feature to the model performance, importance-by-permutation provided by scikit-leam were employed. The scikit-leam “inspection” submodule includes a permutation function for scoring each method. The permutation importance was calculated for each of the Random Forest, Gradient Boosting, and XGBoost models on the withheld test data, with 10 repeats, using the AUC values as the scoring metric for the model’s decrease in performance. In addition, a control feature was added that assigned a random value between 0 and 1 to each peptide, and this was included during training to determine its ranking amongst all features.

[0088] Ex vivo validation of the Ensemble Model using the Streptococcus pneumoniae Proteome. The peptides with the highest and lowest ensemble probability scores were classified as predicted to be immunogenic (n=130) and non-immunogenic (n=130). Peptides were purchased from TC Peptide Lab (San Diego) as crude material on a 1 mg scale. Peptides were combined into 26 pools of 10 peptides each (13 pools with high probability scores, 13 pools with low probability scores), along with a “megapool” consisting of all 260 peptides. The megapool (20) was constructed by pooling the peptides, followed by lyophilization to increase the stock concentration of the peptides resuspended in DMSO. The 10-peptide pools (2 mg / ml) and megapool (1 mg / ml) were stored in aliquots at -20°C until point of use.

[0089] Participant Selection and PBMC Preparation. All participants (n=20) provided written informed consent for participation in the study and ethical approval was obtained from the institutional review boards at La Jolla Institute for Immunology (LJI IRB VD-071 / VD-101). All individuals were >18 years old, free from any acute infection, and recruited in San Diego, California. Peripheral blood mononuclear cells (PBMCs) were isolated from whole blood by density gradient centrifugation with Ficoll (VWR International) according to the manufacturer's instructions. The cells were cryopreserved in liquid nitrogen and suspended in fetal bovine serum (FBS) containing 10% dimethyl sulfoxide (DMSO).

[0090] Each PBMC cryovial was thawed at 37 °C for 2 minutes, and cells were transferred to medium (RPMI 1640 with L-glutamin and 25 mM HEPES; VWR International), supplemented with 5% human AB serum (Gemini Bio), 1% penicillin streptomycin (Gemini Bioproducts), 1% glutamax (Gibco) and 20 U / ml benzonase nuclease (Fisher Scientific). Cells were centrifuged and resuspended in RPMI medium to determine cell concentration using trypan blue.

[0091] Peptide Screening and Deconvolution. The synthesized peptides were screened for recognition with IFNy Fluorospot assays. Participant responses were determined based on the number of IFNy spotforming cells (SFC) following stimulation with peptide pools or individual peptides and controls. Peptides were tested in 10-peptide pools, along with the megapool consisting of all 260 peptides. All 10-peptide pools with significant IFNy response were deconvoluted to determine individual peptide responses. Assays were performed in triplicates, except for the DMSO negative control, which was tested in six wells. The background was calculated by averaging the six DMSO wells by participant. The SFC was calculated byaveraging the triplicate wells for each stimulation condition. Responses were considered positive if the number of SFC above background was 1) 20 or more SFC per 106PBMC, 2) a greater than 2-fold increase, 3) p < 0.05 by independent t-test and Poisson distribution test when comparing stimuli to negative control. A maximum cutoff of 500 SFC / 2xlO5was utilized. Any responses that did not meet the 3 criteria for positivity, were plotted as 0 SFC / 2xlO5cells to highlight IFNy responses that were significantly different compared to the negative control. The response frequency was calculated by dividing the number of participants with a response by the total number of participants screened.

[0092] Plates were coated overnight at 4°C with an antibody mixture containing mouse anti-human IFNy (Mabtech, 1-D1K). The coated plates were then prepped with stimuli as follows: megapool at 2 pg / ml, individual pools at 5 pg / ml, PH A at 10 pg / ml (positive control), and media containing 0.2% DMSO (negative control). The stimuli were then combined with 2xl05PBMCs per well for 20-24 hours at 37°C in a humidified CO incubator. After incubation, the cells were removed, and the wells were washed with PBS / 0.05% Tween with an automatic plate washer. An antibody mixture containing anti-IFNy (7-B6-1-FS-BAM) (Mabtech) was prepared in PBS with 0.1% BSA and incubated for 2 hours at room temperature. The plates were then washed and incubated with anti -B AM-490 (Mabtech) for 1 hour at room temperature. A final wash and incubation with a fluorescence enhancer (Mabtech) for 15 minutes was performed and the fluorescent spots were counted with the IRIS Fluorospot reader (Mabtech).

[0093] Combined Model Training on Mtb and B. pertussis Data. The inventors combined the peptide datasets from Mtb and Bordetella pertussis for model training. The combined dataset comprised 44,510 peptides, with 20,216 peptides from Mtb and 24,294 peptides from B. pertussis. Positive peptides were defined based on recognition by at least two participants in their respective datasets, resulting in 144 positive peptides from Mtb and 412 positive peptides from B. pertussis, for a total of 556 positive peptides. The inventors retrained the Random Forest, Gradient Boosting, and XGBoost algorithms on the combined dataset using the same hyperparameter tuning and 10-fold cross-validation procedures described herein. Tire ensemble model was constructed by averaging the predicted probabilities from each of the three individual models. Model performance was evaluated as described herein. Feature importance was performed as described herein.

[0094] This disclosure demonstrates the effectiveness of machine learning models in predicting immunogenic CD4+ T cell epitopes across multiple bacteria causing airway infections, with a particular focus on Mtb, Bordetella pertussis, and Streptococcus pneumoniae . The integration of multiple biological features, including MHC class II binding predictions, RNA expression levels, conservation scores, and subcellular localization, allowed for the development of robust predictive models. Among these features, MHC class II binding predictions and gene expression levels emerged as the most critical factors influencing epitope immunogenicity. These findings highlight the importance of both the ability of peptides to bind effectively to MHC molecules and the level of antigen expression within the pathogen in determining the likelihood of eliciting a T cell response. The ensemble model disclosed herein, combining the predictions from Random Forest, Gradient Boosting, and XGBoost algorithms, showed superior performance in identifying epitopes with high immunogenic potential. This approach not only streamlinesthe epitope selection process, reducing the need for extensive peptide synthesis and experimental validation but also provides insights into the important biological features of Mtb infection.

[0095] The validation of the models disclosed herein using data from active TB participants and Bordetella pertussis peptides highlights the model's generalizability across different pathogens and disease states. The progressive increase in predictive accuracy observed with higher thresholds for number of participants recognizing each epitope in the active TB and the Bordetella pertussis validations indicates that the model is particularly effective in identifying epitopes that elicit broad immune recognition. This is crucial for vaccine design, as epitope targets recognized by a larger proportion of the population are beneficial for inducing protective immunity in a population. Additionally, the ability of the model to distinguish between immunogenic and non-immunogenic peptides in Streptococcus pneumoniae, as evidenced by the ex vivo IFNy FluoroSpot assays, further underscores the practical utility of the approach disclosed herein for in guiding experimental efforts and peptide identification. This approach can be applied to large-scale epitope discovery efforts, particularly for pathogens with few or no known CD4+ T cell targets.

[0096] The features for the model disclosed herein were chosen due to their applicability across a wide range of pathogens, particularly because they are relatively easy to obtain and do not require extensive pathogen-specific data. The model effectively predicts epitopes based on the features included. The model can be extended to other pathogens and pathogen-types and has applicability across a broader range of infectious agents, such as immunogenic peptide epitopes that are viral, bacterial, fungal, protozoan, or helminthic epitopes. The model showed high accuracy in predicting immunogenicity.

[0097] The inventors also set out to combine the Mtb and Bordetella pertussis data by training a new model using the same methods. While the AUC values are lower than the original Mtb model, the inventors demonstrate that by encompassing multiple bacterial species, the model captured a broader candidate of epitopes from pathogens from limited immunological data.

[0098] Predicting CD4+ T cell epitopes in Mtb and other pneumonia-causing bacteria using machine learning models offers a promising approach to identify the most likely immunogenic peptides, reducing the need to synthesize and test large numbers of peptides, which can be both time-consuming and costly. Machine learning models can integrate many biological features to predict which epitopes are likely to elicit strong immune responses. This peptide prediction model has huge applications for vaccine discovery and diagnostic development, since it is capable of rapidly identifying the most immunogenic bacterial epitopes. This study leveraged computational approaches to predict immunogenic CD4+ T cell epitopes across an array of different respiratory diseases, highlighting the diverse applicability of this methodology.

[0099] In conclusion, this study demonstrates that machine learning models can significantly enhance the efficiency and accuracy of CD4+ T cell epitope prediction, offering a powerful tool for vaccine development and immunological research. By reducing the need for large-scale peptide synthesis and experimental screening, this approach has the potential to accelerate the identification of immunogenic epitopes and contribute to the development of more effective vaccines against a range of bacterial pathogens.

[0100] Definitions:

[0101] The term "gene" means the segment of DNA involved in producing a protein; it includes regions preceding and following the coding region (leader and trailer) as well as intervening sequences (introns) between individual coding segments (exons). The leader, the trailer as well as the introns include regulatory elements that are necessary during the transcription and the translation of a gene. Further, a "protein gene product" is a protein expressed from a particular gene.

[0102] The word “expression” or “expressed” as used herein in reference to a gene means the transcriptional and / or translational product of that gene. The level of expression of a DNA molecule in a cell may be determined on the basis of either the amount of corresponding mRNA that is present within the cell or the amount of protein encoded by that DNA produced by the cell. The level of expression of non-coding nucleic acid molecules (e.g., sgRNA) may be detected by standard PCR or Northern blot methods well known in the art. See, Sambrooketal., 1989 Molecular Cloning: A Laboratory Manual, 18.1-18.88.

[0103] The term "amino acid" refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified, e.g., hydroxyproline, y-carboxyglutamate, and O-phosphoserine. Amino acid analogs refers to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., an a carbon that is bound to a hydrogen, a carboxyl group, an amino group, and an R group, e.g., homoserine, norleucine, methionine sulfoxide, methionine methyl sulfonium. Such analogs have modified R groups (e.g. , norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid. Amino acid mimetics refers to chemical compounds that have a structure that is different from the general chemical structure of an amino acid, but that functions in a manner similar to a naturally occurring amino acid. The terms “non-naturally occurring amino acid” and “unnatural amino acid” refer to amino acid analogs, synthetic amino acids, and amino acid mimetics which are not found in nature.

[0104] Amino acids may be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides, likewise, may be referred to by their commonly accepted single-letter codes.

[0105] The terms "polypeptide," "peptide" and "protein" are used interchangeably herein to refer to a polymer of amino acid residues, wherein the polymer may, in embodiments, be conjugated to a moiety that does not consist of amino acids. The terms apply to amino acid polymers in which one or more amino acid residue is an artificial chemical mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers and non-naturally occurring amino acid polymers. A "fusion protein" refers to a chimeric protein encoding two or more separate protein sequences that are recombinantly expressed as a single moiety.

[0106] Proteins and peptides include isolated and purified forms. Proteins and peptides also include thoseimmobilized on a substrate, as well as amino acid sequences, subsequences, portions, homologues, variants, and derivatives immobilized on a substrate.

[0107] Proteins and peptides can be included in compositions, for example, a pharmaceutical composition. In particular embodiments, a pharmaceutical composition is suitable for specific or non-specific immunotherapy, or is a vaccine composition.

[0108] Isolated nucleic acid (including isolated nucleic acid) encoding the proteins and peptides are also provided. Cells expressing a protein or peptide are further provided. Such cells include eukaryotic and prokaryotic cells, such as mammalian, insect, fungal and bacterial cells.

[0109] Methods and uses and medicaments of proteins and peptides of the invention are included. Such methods, uses and medicaments include modulating immune activity of a cell against a pathogen, for example, a bacteria or bacteria.

[0110] The term “peptide mimetic” or “peptidomimetic” refers to protein-like chain designed to mimic a peptide or protein. Peptide mimetics may be generated by modifying an existing peptide or by designing a compound that mimic peptides, including peptoids and [3-peptides.

[0111] "Conservatively modified variants" applies to both amino acid and nucleic acid sequences. With respect to particular nucleic acid sequences, "conservatively modified variants" refers to those nucleic acids that encode identical or essentially identical amino acid sequences. Because of the degeneracy of the genetic code, a number of nucleic acid sequences will encode any given protein. For instance, the codons GCA, GCC, GCG and GCU all encode the amino acid alanine. Thus, at every position where an alanine is specified by a codon, the codon can be altered to any of the corresponding codons described without altering the encoded polypeptide. Such nucleic acid variations are "silent variations," which are one species of conservatively modified variations. Every nucleic acid sequence herein which encodes a polypeptide also describes every possible silent variation of the nucleic acid. One of skill will recognize that each codon in a nucleic acid (except AUG, which is ordinarily the only codon for methionine, and TGG, which is ordinarily the only codon for tryptophan) can be modified to yield a functionally identical molecule. Accordingly, each silent variation of a nucleic acid which encodes a polypeptide is implicit in each described sequence.

[0112] As to amino acid sequences, one of skill will recognize that individual substitutions, deletions or additions to a nucleic acid, peptide, polypeptide, or protein sequence which alters, adds or deletes a single amino acid or a small percentage of amino acids in the encoded sequence is a "conservatively modified variant" where the alteration results in the substitution of an amino acid with a chemically similar amino acid. Conservative substitution tables providing functionally similar amino acids are well known in the art. Such conservatively modified variants are in addition to and do not exclude polymorphic variants, interspecies homologs, and alleles of the disclosure. The following eight groups each contain amino acids that are conservative substitutions for one another: (1) Alanine (A), Glycine (G); (2) Aspartic acid (D), Glutamic acid (E); (3) Asparagine (N), Glutamine (Q); (4) Arginine (R), Lysine (K); (5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); (6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W); (7)Serine (S), Threonine (T); and (8) Cysteine (C), Methionine (M) (see, e.g., Creighton, Proteins (1984)).

[0113] A "percentage of sequence identity" is determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide or polypeptide sequence in the comparison window may comprise additions or deletions (i. e. , gaps) as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of sequence identity.

[0114] The terms "identical" or percent "identity," in the context of two or more nucleic acids or polypeptide sequences, refer to two or more sequences or subsequences that are the same or have a specified percentage of amino acid residues or nucleotides that are the same (i.e., about 60% identity, preferably 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher identity over a specified region, when compared and aligned for maximum correspondence over a comparison window or designated region) as measured using a BLAST or BLAST 2.0 sequence comparison algorithms with default parameters described below, or by manual alignment and visual inspection (see, e.g., NCBI web site ncbi.nlm.nih.gov / BLAST / or the like). Such sequences are then said to be "substantially identical." This definition also refers to, or may be applied to, the compliment of a test sequence. The definition also includes sequences that have deletions and / or additions, as well as those that have substitutions. As described below, the preferred algorithms can account for gaps and the like. Preferably, identity exists over a region that is at least about 25 amino acids or nucleotides in length, or more preferably over a region that is 50-100 amino acids or nucleotides in length.

[0115] An amino acid or nucleotide base "position" is denoted by a number that sequentially identifies each amino acid (or nucleotide base) in the reference sequence based on its position relative to the N-terminus (or 5'-end). Due to deletions, insertions, truncations, fusions, and the like that must be taken into account when determining an optimal alignment, in general the amino acid residue number in a test sequence determined by simply counting from the N-terminus will not necessarily be the same as the number of its corresponding position in the reference sequence. For example, in a case where a variant has a deletion relative to an aligned reference sequence, there will be no amino acid in the variant that corresponds to a position in the reference sequence at the site of deletion. Where there is an insertion in an aligned reference sequence, that insertion will not correspond to a numbered amino acid position in the reference sequence. In the case of truncations or fusions there can be stretches of amino acids in either the reference or aligned sequence that do not correspond to any amino acid in the corresponding sequence.

[0116] The terms "numbered with reference to" or "corresponding to," when used in the context of the numbering of a given amino acid or polynucleotide sequence, refers to the numbering of the residues of a specified reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence.

[0117] The term “multimer” refers to a complex comprising multiple monomers (e.g., a protein complex) associated by noncovalent bonds. The monomers be substantially identical monomers, or the monomers may be different. In embodiments, the multimer is a dimer, a trimer, a tetramer, or a pentamer.

[0118] As used herein, the term "Major Histocompatibility Complex" (MHC) is a generic designation meant to encompass the histocompatibility antigen systems described in different species including the human leucocyte antigens (HLA). Typically, MHC Class I or Class II multimers are well known in the art and include but are not limited to dimers, tetramers, pentamers, hexamers, heptamers and octamers.

[0119] As used herein, the term "MHC / peptide multimer" refers to a stable multimeric complex composed of MHC protein(s) subunits loaded with a peptide of the present invention. For example, an MHC / peptide multimer (also called herein MHC / peptide complex) include, but are not limited to, an MHC / peptide dimer, trimer, tetramer, pentamer or higher valency multimer. In humans there are three major different genetic loci that encode MHC class I molecules (the MHC molecules of the human are also designated human leukocyte antigens (HLA)): HLA-A, HLA-B, HLA-C, e g., HLA-A*01, HLA-A*02, and HLA-A*11 are examples of different MHC class I alleles that can be expressed from these loci. Non-classical human MHC class I molecules such as HLA-E (homolog of mice Qa-lb) and MICA / B molecules are also encompassed by the present invention. In some embodiments, the MHC / peptide multimer is an HLA / peptide multimer selected from the group consisting of HLA-A / peptide multimer, HLA-B / peptide multimer, HLA-C / peptide multimer, HLA-E / peptide multimer, MICA / peptide multimer and MICB / peptide multimer.

[0120] In humans there are three major different genetic loci that encode MHC class II molecules: HLA-DR, HLA-DP, and HLA-DQ, each formed of two polypeptides, alpha and beta chains (A and B genes). For example, HLA-DQAl*01, HLA-DRBl*01, and HLA-DRBl*03 are different MHC class II alleles that can be expressed from these loci. It should be further noted that non-classical human MHC class II molecules such as HLA-DM and HL-DOA (homolog in mice is H2-DM and H2-O) are also encompassed by the present invention. In some embodiments, the MHC / peptide multimer is an HLA / peptide multimer selected from the group consisting of HLA-DP / peptide multimer, HLA-DQ / peptide multimer, HLA-DR / peptide multimer, HLA-DM / peptide multimer and HLA-DO / peptide multimer.

[0121] An MHC / peptide multimer may be a multimer where the heavy chain of the MHC is biotinylated, which allows combination as a tetramer with streptavidin. MHC -peptide tetramers have increased avidity for the appropriate T cell receptor (TCR) on T lymphocytes. The multimers can also be attached to paramagnetic particles or magnetic beads to facilitate removal of non-specifically bound reporter and cell sorting. Multimer staining does not kill the labelled cells, thus, cell integrity is maintained for further analysis. In some embodiments, the MHC / peptide multimer of the present invention is particularly suitable for isolating and / or identifying a population of CD8+ T cells having specificity for the peptide of the present invention (in a flow cytometry assay).

[0122] The peptides or MHC class I or class II multimer as described herein is particularly suitable for detecting T cells specific for one or more peptides of the present invention. The peptide(s) and / or the MHC / multimer complex of the present invention is particularly suitable for diagnosing Streptococcuspneumoniae infection in a subject. For example, the method comprises obtaining a blood or PBMC sample obtained from the subject with an amount of a least peptide of the present invention and detecting at least one T cell displaying a specificity for the peptide. Another diagnostic method of the present invention involves the use of a peptide of the present invention that is loaded on multimers as described above, so that the isolated CD8+ or CD4+ T cells from the subject are brought into contact with the multimers, at which the binding, activation and / or expansion of the T cells is measured. For example, following the binding to antigen presenting cells, e.g., those having the MHC class I or class II multimer, the number of CD8+ and / or CD4+ cells binding specifically to the HLA-peptide multimer may be quantified by measuring the secretion of lymphokines / cytokines, division of the T cells, or standard flow cytometry methods, such as, for example, using fluorescence activated cell sorting (FACS). The multimers can also be attached to paramagnetic ferrous or magnetic beads to facilitate removal of non-specifically bound reporter and cell sorting.

[0123] The MHC class I or class II peptide multimers as described herein can also be used as therapeutic agents. The peptide and / or the MHC class I or class II peptide multimers of the present invention are suitable for treating or preventing a Streptococcus pneumoniae infection in a subject. The MHC Class I or Class II multimers can be administered in soluble form or loaded on nanoparticles.

[0124] The term "antibody" refers to a polypeptide encoded by an immunoglobulin gene or functional fragments thereof that specifically binds and recognizes an antigen. The recognized immunoglobulin genes include the kappa, lambda, alpha, gamma, delta, epsilon, and mu constant region genes, as well as the myriad immunoglobulin variable region genes. Light chains are classified as either kappa or lambda. Heavy chains are classified as gamma, mu, alpha, delta, or epsilon, which in turn define the immunoglobulin classes, IgG, IgM, IgA, IgD and IgE, respectively.

[0125] The phrase “specifically (or selectively) binds” to an antibody or “specifically (or selectively) immunoreactive with,” when referring to a protein or peptide, refers to a binding reaction that is determinative of the presence of the protein or peptide, often in a heterogeneous population of proteins and other biologies. Thus, under designated immunoassay conditions, the specified antibodies bind to a particular protein at least two times the background and more typically more than 10 to 100 times background. Specific binding to an antibody under such conditions requires an antibody that is selected for its specificity for a particular protein. For example, polyclonal antibodies can be selected to obtain only a subset of antibodies that are specifically immunoreactive with the selected antigen and not with other proteins. This selection may be achieved by subtracting out antibodies that cross-react with other molecules . A variety of immunoassay formats may be used to select antibodies specifically immunoreactive with a particular protein. For example, solid-phase ELISA immunoassays are routinely used to select antibodies specifically immunoreactive with a protein (see, e.g., Harlow & Lane, Using Antibodies, A Laboratory Manual (1998) for a description of immunoassay formats and conditions that can be used to determine specific immunoreactivity).

[0126] Antibodies are large, complex molecules (molecular weight of -150,000 or about 1320 aminoacids) with intricate internal structure. A natural antibody molecule contains two identical pairs of polypeptide chains, each pair having one light chain and one heavy chain. Each light chain and heavy chain in turn consists of two regions: a variable ("V") region involved in binding the target antigen, and a constant ("C") region that interacts with other components of the immune system. The light and heavy chain variable regions come together in 3 -dimensional space to form a variable region that binds the antigen (for example, a receptor on the surface of a cell). Within each light or heavy chain variable region, there are three short segments (averaging 10 amino acids in length) called the complementarity determining regions ("CDRs"). The six CDRs in an antibody variable domain (three from the light chain and three from the heavy chain) fold up together in 3 -dimensional space to form the actual antibody binding site which docks onto the target antigen. The position and length of the CDRs have been precisely defined by Kabat, E. et al., Sequences of Proteins of Immunological Interest, U.S. Department of Health and Human Services, 1983, 1987. The part of a variable region not contained in the CDRs is called the framework ("FR"), which forms the environment for the CDRs.

[0127] The term "antibody" is used according to its commonly known meaning in the art. Antibodies exist, e.g., as intact immunoglobulins or as a number of well-characterized fragments produced by digestion with various peptidases. Thus, for example, pepsin digests an antibody below the disulfide linkages in the hinge region to produce F(ab)'2, a dimer of Fab which itself is a light chain joined to VH-CHI by a disulfide bond. The F(ab)'2 may be reduced under mild conditions to break the disulfide linkage in the hinge region, thereby converting the F(ab)'2 dimer into a Fab' monomer. The Fab' monomer is essentially Fab with part of the hinge region (see Fundamental Immunology (Paul ed., 3d ed. 1993). While various antibody fragments are defined in terms of the digestion of an intact antibody, one of skill will appreciate that such fragments may be synthesized de novo either chemically or by using recombinant DNA methodology. Thus, the term antibody, as used herein, also includes antibody fragments either produced by the modification of whole antibodies, orthose synthesized de novo using recombinant DNA methodologies (e.g., single chain Fv) or those identified using phage display libraries (see, e.g., McCafferty et al., Nature 348:552-554 (1990)).

[0128] An exemplary immunoglobulin (antibody) structural unit comprises a tetramer. Each tetramer is composed of two identical pairs of polypeptide chains, each pair having one “light” (about 25 kD) and one “heavy” chain (about 50-70 kD). The N-terminus of each chain defines a variable region of about 100 to 110 or more amino acids primarily responsible for antigen recognition. The terms variable light chain (VL) and variable heavy chain (VH) refer to these light and heavy chains respectively. The Fc (i.e., fragment crystallizable region) is the “base” or "tail" of an immunoglobulin and is typically composed of two heavy chains that contribute two or three constant domains depending on the class of the antibody. By binding to specific proteins, the Fc region ensures that each antibody generates an appropriate immune response for a given antigen. The Fc region also binds to various cell receptors, such as Fc receptors, and other immune molecules, such as complement proteins.

[0129] As used herein, the term “antigen” and the term “epitope” refers to a molecule or substance capable of stimulating an immune response. In one example, epitopes include but are not limited to a polypeptide and a nucleic acid encoding a polypeptide, wherein expression of the nucleic acid into a polypeptide iscapable of stimulating an immune response when the polypeptide is processed and presented on a Major Histocompatibility Complex (MHC) molecule. Generally, epitopes include peptides presented on the surface of cells non-covalently bound to the binding groove of Class I or Class II MHC, such that they can interact with T cell receptors and the respective T cell accessory molecules. However, antigens and epitopes also apply when discussing the antigen binding portion of an antibody, wherein the antibody binds to a specific structure of the antigen.

[0130] Proteolytic Processing of Antigens. Epitopes that are displayed by MHC on antigen presenting cells are cleavage peptides or products of larger peptide or protein antigen precursors. For MHC I epitopes, protein antigens are often digested by proteasomes resident in the cell. Intracellular proteasomal digestion produces peptide fragments of about 3 to 23 amino acids in length that are then loaded onto the MHC protein. Additional proteolytic activities within the cell, or in the extracellular milieu, can trim and process these fragments further. Processing of MHC Class II epitopes generally occurs via intracellular proteases from the lysosomal / endosomal compartment. The present invention includes, in one embodiment, pre-processed peptides that are attached to the anti-CD40 antibody (or fragment thereof) that directs the peptides against which an enhanced immune response is sought directly to antigen presenting cells.

[0131] As used herein, the phrase “antigen strain conservation scores” refers to the antigen from a specific pool of peptides used as a template for the machine learning analysis. Non-limiting examples of pathogens include a bacterium that causes a bacterial infection, e.g., Acinetobacter species, Actinobacillus species, Actinomycetes species, an Actinomyces species, Aerococcus species w Aeromonas species, an Anaplasma species, an Alcaligenes species, a Bacillus species, a Bacteroides species, a Bartonella species, a Bifidobacterium species, a Bordetella species, a Borrelia species, a Brucella species, a Burkholderia species, a Campylobacter species, a Capnocytophaga species, a Chlamydia species, a Citrobacter species, a Coxiella species, a Corynbacterium species, a Clostridium species, an Eikenella species, an Enterobacter species, an Escherichia species, an Enterococcus species, an Ehlichia species, an Epidermophyton species, an Erysipelothrix species, a Eubacterium species, a Francisella species, a Fusobacterium species, a Gardnerella species, a Gemella species, a Haemophilus species, ^Helicobacter species, Kingella species, a Klebsiella species, a Lactobacillus species, a Lactococcus species, a Listeria species, a Leptospira species, a Legionella species, a Leptospira species, Leuconostoc species, a Mannheimia species, a Microsporum species, a Micrococcus species, a. Moraxella species, a Morganell species, aMobiluncus species, a Micrococcus species, Mycobacterium species, a Mycoplasm species, a Nocardia species, a Neisseria species, a Pasteurelaa species, a Pediococcus species, a Peptostreptococcus species, a Pityrosporum species, a Plesiomonas species, a Prevotella species, a Porphyromonas species, a Proteus species, a Providencia species, a Pseudomonas species, a Propionib acteriums species, a Rhodococcus species, a Rickettsia species, a Rhodococcus species, a Serratia species, a Stenotrophomonas species, a Salmonella species, a Serratia species, a Shigella species, a Staphylococcus species, a Streptococcus species, a Spirillum species, a Streptobacillus species, a Treponema species, a Tropheryma species, a Trichophyton species, an Ureaplasma species, a Veillonella species, a Vibrio species, a Yersinia species, a Xanthomonas species, or combination thereof. The infection may be a viral infection which may be aMyoviridae, Podoviridae, Siphoviridae, Alloherpesviridae, Herpesviridae (including human herpes virus, and Varicella Zoster virus), Malocoherpesviridae, Lipothrixviridae, Rudiviridae, Adenoviridae, Ampullaviridae, Ascoviridae, Asfarviridae (including African swine fever virus), Baculoviridae, Cicaudaviridae, Clavaviridae, Corticoviridae, Fuselloviridae, Globuloviridae, Guttaviridae, Hytrosaviridae, Iridoviridae, Maseilleviridae, Mimiviridae, Nudiviridae, Nimaviridae, Pandoraviridae, Papillomaviridae, Phycodnaviridae, Plasmaviridae, Polydnaviruses, Polyomaviridae (including Simian virus 40, JC virus, BK virus), Poxviridae (including Cowpox and smallpox), Sphaerolipoviridae, Tectiviridae, Turriviridae, Dinodnavirus, Salterprovirus, Rhizidovirus. The viral infection may be caused by a double-stranded RNA virus, a positive sense RNA virus, a negative sense RNA virus, a retrovirus, or a combination thereof. The viral infection may be caused by a Coronaviridae virus, a Picomaviridae virus, a Caliciviridae virus, a Flaviviridae virus, a Togaviridae virus, a Bomaviridae, a Filoviridae, a Paramyxoviridae, a Pneumoviridae, a Rhabdoviridae, an Arenaviridae, a Bunyaviridae, an Orthomyxoviridae, or a Deltavirus. The viral infection may be caused by Coronavirus, SARS, SARS- COV2, Poliovirus, Rhinovirus, Hepatitis A, Norwalk virus, Yellow fever virus, West Nile virus, Hepatitis C virus, Dengue fever virus, Zika virus, Rubella virus, Ross River virus, Sindbis virus, Chikungunya virus, Boma disease virus, Ebola virus, Marburg virus, Measles virus, Mumps virus, Nipah virus, Hendra virus, Newcastle disease virus, Human respiratory syncytial virus, Rabies virus, Lassa virus, Hantavirus, Crimean-Congo hemorrhagic fever virus, Influenza, or Hepatitis D virus. Examples of protozoan parasite antigens include, but are not limited to, Babesia polypeptides, Balantidium polypeptides, Besnoitia polypeptides, Cryptosporidium polypeptides, Eimeria polypeptides, Encephalitozoon polypeptides, Entamoeba polypeptides, Giardia polypeptides, Hammondia polypeptides, Hepatozoon polypeptides, Isospora polypeptides, Leishmania polypeptides, Microsporidia polypeptides, Neospora polypeptides, Nosema polypeptides, Pentatrichomonas polypeptides, Plasmodium polypeptides. Examples of helminth parasite antigens include, but are not limited to, Acanthocheilonema polypeptides, Aelurostrongylus polypeptides, Ancylostoma polypeptides, Angiostrongylus polypeptides, Ascaris polypeptides, Brugia polypeptides, Bunostomum polypeptides, Capillaria polypeptides, Chabertia polypeptides, Cooperia polypeptides, Crenosoma polypeptides, Dictyocaulus polypeptides, Dioctophyme polypeptides, Dipetalonema polypeptides, Diphyllobothrium polypeptides, Diplydium polypeptides, Dirofilaria polypeptides, Dracunculus polypeptides, Enterobius polypeptides, Filaroides polypeptides, Haemonchus polypeptides, Lagochilascaris polypeptides, Loa polypeptides, Mansonella polypeptides, Muellerius polypeptides, Nanophyetus polypeptides, Necator polypeptides, Nematodirus polypeptides, Oesophagostomum polypeptides, Onchocerca polypeptides, Opisthorchis polypeptides, Ostertagia polypeptides, Parafdaria polypeptides, Paragonimus polypeptides, Parascaris polypeptides, Physaloptera polypeptides, Protostrongylus polypeptides, Setaria polypeptides, Spirocerca polypeptides Spirometra polypeptides, Stephanofdaria polypeptides, Strongyloides polypeptides, Strongylus polypeptides, Thelazia polypeptides, Toxascaris polypeptides, Toxocara polypeptides, Trichinella polypeptides, Trichostrongylus polypeptides, Trichuris polypeptides, Uncinaria polypeptides, and Wuchereria polypeptides, (e.g., P. falciparum circumsporozoite (PfCSP)), sporozoite surface protein 2 (PfSSP2), carboxyl terminus of liverstate antigen 1 (PfLSAl c-term), and exported protein 1 (PfExp-1), Pneumocystis polypeptides, Sarcocystis polypeptides, Schistosoma polypeptides, Theileria polypeptides, Toxoplasma polypeptides, and Trypanosoma polypeptides.

[0132] The present invention includes methods for specifically identifying the epitopes within antigens most likely to lead to the immune response sought for the specific sources of antigen presenting cells and responder T cells.

[0133] As used herein, the term “T cell epitope” refers to a specific amino acid that when present in the context of a Major or Minor Histocompatibility Complex provides a reactive site for a T cell receptor. The T-cell epitopes or peptides that stimulate the cellular arm of a subject's immune system are short peptides of about 8-25 amino acids. T-cell epitopes are recognized by T cells from animals that are immune to the antigen of interest. These T-cell epitopes or peptides can be used in assays such as the stimulation of cytokine release or secretion or evaluated by constructing major histocompatibility (MHC) proteins containing or “presenting” the peptide. Such immunogenically active fragments are often identified based on their ability to stimulate lymphocyte proliferation in response to stimulation by various fragments from the antigen of interest.

[0134] As used herein, the term “immunological response” refers to an antigen or composition is the development in a subject of a humoral and / or a cellular immune response to an antigen present in the composition of interest. For purposes of the present disclosure, a “humoral immune response” refers to an immune response mediated by antibody molecules, while a “cellular immune response” is one mediated by T-lymphocytes and / or other white blood cells. One important aspect of cellular immunity involves an antigen-specific response by cytolytic T-cells (“CTL”s). CTLs have specificity for peptide antigens that are presented in association with proteins encoded by the major histocompatibility complex (MHC) and expressed on the surfaces of cells. CTLs help induce and promote the destruction of intracellular microbes, or the lysis of cells infected with such microbes. Another aspect of cellular immunity involves an antigenspecific response by helper T-cells. Helper T-cells act to help stimulate the function, and focus the activity of, nonspecific effector cells against cells displaying peptide antigens in association with MHC molecules on their surface. A “cellular immune response” also refers to the production of cytokines, chemokines and other such molecules produced by activated T-cells and / or other white blood cells, including those derived from CD4+ and CD8+ T-cells. Hence, an immunological response may include one or more of the following effects: the production of antibodies by B-cells; and / or the activation of effector and / or suppressor T-cells and / or gamma-delta T-cells directed specifically to an antigen or antigens present in the composition or vaccine of interest. These responses may serve to neutralize infectivity, and / or mediate antibody-complement, or antibody dependent cell cytotoxicity (ADCC) to provide protection to an immunized host. Such responses can be determined using standard immunoassays and neutralization assays, well known in the art.

[0135] As used herein, the term an “immunogenic composition” and “vaccine” refer to a composition that comprises an antigenic molecule where administration of the composition to a subject or patient results inthe development in the subject of a humoral and / or a cellular immune response to the antigenic molecule of interest. “Vaccine” refers to a composition that can provide active acquired immunity to and / or therapeutic effect (e.g., treatment) of a particular disease or a pathogen. A vaccine typically contains one or more agents that can induce an immune response in a subject against a pathogen or disease, i.e., a target pathogen or disease. The immunogenic agent stimulates the body’s immune system to recognize the agent as a threat or indication of the presence of the target pathogen or disease, thereby inducing immunological memory so that the immune system can more easily recognize and destroy any of the pathogen on subsequent exposure. Vaccines can be prophylactic (e.g., preventing or ameliorating the effects of a future infection by any natural or pathogen) or therapeutic (e.g., reducing symptoms or aberrant conditions associated with infection). The administration of vaccines is referred to vaccination.

[0136] In some examples, a vaccine composition can provide nucleic acid, e.g., mRNA that encodes antigenic molecules (e.g., peptides) to a subject. The nucleic acid that is delivered via the vaccine composition in the subject can be expressed into antigenic molecules and allow the subject to acquire immunity against the antigenic molecules. In the context of the vaccination against infectious disease, the vaccine composition can provide mRNA encoding antigenic molecules that are associated with a certain pathogen, e.g., one or more peptides that are known to be expressed in the pathogen (e.g., pathogenic bacterium or bacteria).

[0137] The present invention provides nucleic acid molecules, specifically polynucleotides, primary constructs and / or mRNA that encode one or more polynucleotides that express one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from any one of those sequences of SEQ ID NOS: 1 to 10; and in certain examples SEQ ID NO: 1, SEQ ID NO:6, or both, or a subsequence, portion, homologue, variant or derivative thereof for use in immune modulation. The term "nucleic acid" refers to any compound and / or substance that comprise a polymer of nucleotides, referred to herein as polynucleotides. Exemplary nucleic acids or polynucleotides of the invention include, but are not limited to, ribonucleic acids (RNAs), deoxyribonucleic acids (DNAs), threose nucleic acids (TNAs), glycol nucleic acids (GNAs), peptide nucleic acids (PNAs), locked nucleic acids (LNAs), including diastereomers of LNAs, functionalized LNAs, or hybrids thereof.

[0138] One method of immune modulation of the present invention includes direct or indirect gene transfer, i.e., local application of a preparation containing the one or more polynucleotides (DNA, RNA, mRNA, etc.) that expresses the one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from any one of those sequences of SEQ ID NOS: 1 to 10; and in certain examples SEQ ID NO:1, SEQ ID NO:6, or both, or a subsequence, portion, homologue, variant or derivative thereof. A variety of well-known vectors can be used to deliver to cells the one or more polynucleotides or the peptides or proteins expressed by the polynucleotides, including but not limited to adenobacterial vectors and adeno-associated vectors. In addition, naked DNA, liposome delivery methods, or other novel vectors developed to deliver the polynucleotides to cells can also be beneficial. Any of a variety of promoters can be used to drive peptide or protein expression, including but not limited to endogenous promoters, constitutive promoters (e.g., cytomegalobacteria, adenobacteria, or SV40),inducible promoters (e.g., a cytokine promoter such as the interleukin- 1, tumor necrosis factor-alpha, or interleukin-6 promoter), and tissue specific promoters to express the immunogenic peptides or proteins of the present invention.

[0139] The immunization may include adenobacteria, adeno-associated bacteria, herpes bacteria, vaccinia bacteria, retrobacteriaes, or other bacterial vectors with the appropriate tropism for cells likely to present the antigenic peptide(s) or protein(s) may be used as a gene transfer delivery system for a therapeutic peptide(s) or protein(s), comprising, consisting of, or consisting essentially of an amino acid sequence selected from any one of those sequences of SEQ ID NOS: 1 to 10; and in certain examples SEQ ID NO: 1, SEQ ID NO:6, or both, or a subsequence, portion, homologue, variant or derivative thereof, gene expression construct. Bacterial vectors which do not require that the target cell be actively dividing, such as adenobacterial and adeno-associated vectors, are particularly useful when the cells are accumulating, but not proliferative. Numerous vectors useful for this purpose are generally known (Miller, Human Gene Therapy 15-14, 1990; Friedman, Science 244:1275-1281, 1989; Eglitis and Anderson, BioTechniques 6:608-614, 1988; Tolstoshev and Anderson, Current Opinion in Biotechnology 1:55-61, 1990; Sharp, The Lancet 337:1277-1278, 1991; Cometta et al., Nucleic Acid Research and Molecular Biology 36:311-322, 1987; Anderson, Science 226:401-409, 1984; Moen, Blood Cells 17:407-416, 1991; and Miller and Rosman, Bio Techniques 7:980-990, 1989; Le Gal La Salle et al., Science 259:988-990, 1993; and Johnson, Chest 107:77S-83S, 1995). Retrobacterial vectors are particularly well developed and have been used in clinical settings (Rosenberg etal.,N. Engl. J. Med 323:370, 1990; Anderson et al., U.S. Pat. No. 5,399,346).

[0140] The immunization may also include inserting the one or more polynucleotides (DNA, RNA, mRNA, etc.) that express the one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from any one of those sequences set forth in SEQ ID NO: 1 to 10, and in certain examples peptides of SEQ ID NO:1 and SEQ ID NO:6, or both, or a subsequence, portion, homologue, variant or derivative thereof into the bacterial vector, along with another gene which encodes the ligand for a receptor on a specific target cell, for example, such that the vector is now target specific. Bacterial vectors can be made target specific by attaching, for example, a sugar, a glycolipid, or a protein. Targeting can also be accomplished by using an antibody to target the bacterial vector. Those of skill in the art will know of, or can readily ascertain without undue experimentation, specific polynucleotide sequences which can be inserted into the bacterial genome or attached to a bacterial envelope to allow target specific delivery of the bacterial vector containing the gene.

[0141] Since recombinant bacteria are defective, they require assistance in order to produce infectious vector particles. This assistance can be provided, for example, by using helper cell lines that contain plasmids encoding all of the structural genes of the bacteria under the control of regulatory sequences within the bacterial genome. These plasmids are missing a nucleotide sequence which enables the packaging mechanism to recognize a polynucleotide transcript for encapsidation. These cell lines produce empty virions, since no genome is packaged. If a bacterial vector is introduced into such cells in which the packaging signal is intact, but the structural genes are replaced by other genes of interest, the vector can be packaged and vector virion produced.

[0142] Bacterial or non-bacterial approaches may also be employed for the introduction of one or more therapeutic polynucleotides that express the one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from any one of those sequences of SEQ ID NOS:1 to 10; and in certain examples SEQ ID NO:1, SEQ ID NO:6, or both), or a subsequence, portion, homologue, variant or derivative thereof, into polynucleotide-encoding polynucleotide into antigen presenting cells. The polynucleotides may be DNA, RNA, mRNA that directly encode the one or more peptides or proteins of the present invention, or may be introduced as part of an expression vector.

[0143] Another example of an immunization includes colloidal dispersion systems that include macromolecule complexes, nanocapsules, microspheres, beads, and lipid-based systems including oil-in-water emulsions, micelles, mixed micelles, and liposomes and the one or more polynucleotides that express the one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from any one of those sequences of SEQ ID NOS: 1 to 10; and in certain examples SEQ ID NO: 1, SEQ ID NO:6, or both, or a subsequence, portion, homologue, variant or derivative thereof. One non-limiting example of a colloidal system for use with the present invention is a liposome. Liposomes are artificial membrane vesicles which are useful as delivery vehicles in vitro and in vivo. It has been shown that large unilamellar vesicles (LUV), which range in size from 0.2-4.0 micrometers that can encapsulate a substantial percentage of an aqueous buffer containing large macromolecules. RNA, DNA and intact virions can be encapsulated within the aqueous interior and be delivered to cells in a biologically active form (Fraley, et al., Trends Biochem. Sci., 6:77, 1981). In addition to mammalian cells, liposomes have been used for delivery of polynucleotides in plant, yeast and bacterial cells. In order for a liposome to be an efficient gene transfer vehicle, the following characteristics should be present: (Zakut and Givol, supra) encapsulation of the genes of interest at high efficiency while not compromising their biological activity; (Feamhead, et al., supra) preferential and substantial binding to a target cell in comparison to non-target cells; (Korsmeyer, S. J., supra) delivery of the aqueous contents of the vesicle to the target cell cytoplasm at high efficiency; and (Kinoshita, et al., supra) accurate and effective expression of genetic information (Mannino, et al., Bio Techniques, 6:682, 1988).

[0144] The composition for immunizing the subject or patient may, in certain embodiments comprise a combination of phospholipid, particularly high-phase-transition-temperature phospholipids, usually in combination with steroids, especially cholesterol. Other phospholipids or other lipids may also be used. The physical characteristics of liposomes depend on pH, ionic strength, and the presence of divalent cations. The targeting of liposomes can be classified based on anatomical and mechanistic factors. Anatomical classification is based on the level of selectivity, for example, organ-specific, cell-specific, and organelle-specific. Mechanistic targeting can be distinguished based upon whether it is passive or active. Passive targeting utilizes the natural tendency of liposomes to distribute to cells of the reticuloendothelial system (RES) in organs which contain sinusoidal capillaries. Active targeting, on the other hand, involves alteration of the liposome by coupling the liposome to a specific ligand such as a monoclonal antibody, sugar, glycolipid, or protein, or by changing the composition or size of the liposome in order to achieve targeting to organs and cell types other than the naturally occurring sites of localization, specifically, cellsthat can become infected with a Streptococcus pneumoniae or interact with the proteins, peptides, and / or gene products of a Streptococcus pneumoniae, e.g., immune cells.

[0145] For any of the above approaches, the immune modulating polynucleotide construct, composition, or formulation is preferably applied to a site that will enhance the immune response. For example, the immunization may be intramuscular, intraperitoneal, enteral, parenteral, intranasal, intrapulmonary, or subcutaneous. In the gene delivery constructs of the instant invention, polynucleotide expression is directed from any suitable promoter (e.g., the human cytomegalobacteria, simian bacteria 40, actin or adenobacteria constitutive promoters; or the cytokine or metalloprotease promoters for activated synoviocyte specific expression).

[0146] In one example of the immune modifying peptide(s) or protein(s) include polynucleotides, constructs and / or mRNAs that express the one or more polynucleotides that express the one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from any one of those sequences of SEQ ID NOS: 1 to 10; and in certain examples SEQ ID NO: 1, SEQ ID NO:6, or both, or a subsequence, portion, homologue, variant or derivative thereof, that are designed to improve one or more of the stability and / or clearance in tissues, uptake and / or kinetics, cellular access by the peptide(s) or protein(s), translational, mRNA half-life, translation efficiency, immune evasion, protein production capacity, accessibility to circulation, peptide(s) or protein(s) half-life and / or presentation in the context of MHC on antigen presenting cells.

[0147] The present invention contemplates immunization for use in both active and passive immunization embodiments. Immunogenic compositions, proposed to be suitable for use as a vaccine, may be prepared most readily directly from immunogenic peptides, proteins, monomers, multimers and / or peptide-MHC complexes prepared in a manner disclosed herein. The antigenic material is generally processed to remove undesired contaminants, such as, small molecular weight molecules, incomplete proteins, or when manufactured in plant cells, plant components such as cell walls, plant proteins, and the like. Often, these immunizations are lyophilized for ease of transport and / or to increase shelf-life and can then be more readily dissolved in a desired vehicle, such as saline.

[0148] The preparation of immunizations (also referred to as vaccines) that contain the immunogenic proteins of the present invention as active ingredients is generally well understood in the art, as exemplified by United States Letters Patents 4,608,251; 4,601,903; 4,599,231; 4,599,230; 4,596,792; and 4.578,770, all incorporated herein by reference. Typically, such immunizations are prepared as injectables. The immunizations can be a liquid solution or suspension but may also be provided in a solid form suitable for solution in, or suspension in, liquid prior to injection may also be prepared. The preparation may also be emulsified. The active immunogenic ingredient is often mixed with excipients that are pharmaceutically acceptable and compatible with the active ingredient. Suitable excipients are, for example, water, saline, dextrose, glycerol, ethanol, buffers, or the like and combinations thereof. In addition, if desired, the immunization may contain minor amounts of auxiliary substances such as wetting or emulsifying agents, pH buffering agents, or adjuvants which enhance the effectiveness of the vaccines.

[0149] The immunization is / are administered in a manner compatible with the dosage formulation, and in such amount as will be therapeutically effective and immunogenic. The quantity to be administered depends on the subject to be treated, including, e.g., the capacity of the individual's immune system to synthesize antibodies, and the degree of protection desired. Precise amounts of active ingredient required to be administered depend on the judgment of the practitioner. However, suitable dosage ranges are of the order of several hundred micrograms active ingredient per vaccination. Suitable regimes for initial administration and booster shots are also variable but are typified by an initial administration followed by subsequent inoculations or other administrations.

[0150] The manner of application of the immunization may be varied widely. Any of the conventional methods for administration of a vaccine are applicable. These are believed to also include oral application on a solid physiologically acceptable base or in a physiologically acceptable dispersion, parenterally, by injection or the like. The dosage of the vaccine will depend on the route of administration and will vary according to the size of the host.

[0151] Various methods of achieving adjuvant effect for the vaccine includes use of agents such as aluminum hydroxide or phosphate (alum), commonly used as 0.05 to 0.1 percent solution in phosphate buffered saline, admixture with synthetic polymers of sugars (Carbopol) used as 0.25 percent solution, aggregation of the protein in the vaccine by heat treatment with temperatures ranging between 70° to 101°C for 30 second to 2-minute periods respectively. Aggregation by reactivating with pepsin treated (Fab) antibodies to albumin, mixture with bacterial cells such as C. parvum or endotoxins or lipopolysaccharide components of gram-negative bacteria, emulsion in physiologically acceptable oil vehicles such as mannide mono-oleate (Aracel A) or emulsion with 20 percent solution of a perfluorocarbon (Fluosol-DA) used as a block substitute may also be employed.

[0152] In many instances, it will be desirable to have multiple administrations of the vaccine, usually not exceeding six to ten immunizations, more usually not exceeding four immunizations and preferably one or more, usually at least about three immunizations. The immunizations will normally be at from two to twelve-week intervals, more usually from three to five-week intervals. Periodic boosters at intervals of 1-5 years, usually three years, will be desirable to maintain protective levels of the antibodies. The course of the immunization may be followed by assays for antibodies for the supernatant antigens. The assays may be performed by labeling with conventional labels, such as radionuclides, enzymes, fluorescent agents, and the like. These techniques are well known and may be found in a wide variety of patents, such as Hudson and Cranage, Vaccine Protocols, 2003 Humana Press, relevant portions incorporated herein by reference.

[0153] Techniques and compositions for making useful dosage forms using the present invention are described in one or more of the following references: Anderson, Philip O.; Knoben, James E.; Troutman, William G, eds., Handbook of Clinical Drug Data, Tenth Edition, McGraw-Hill, 2002; Pratt and Taylor, eds., Principles of Drug Action, Third Edition, Churchill Livingston, New York, 1990; Katzung, ed., Basic and Clinical Pharmacology, Ninth Edition, McGraw Hill, 2007; Goodman and Gilman, eds., The Pharmacological Basis of Therapeutics, Tenth Edition, McGraw Hill, 2001; Remington’s PharmaceuticalSciences, 20th Ed., Lippincott Williams & Wilkins., 2000, and updates thereto; Martindale, The Extra Pharmacopoeia, Thirty-Second Edition (The Pharmaceutical Press, London, 1999); all of which are incorporated by reference, and the like, relevant portions incorporated herein by reference.

[0154] Many suitable expression systems are commercially available, including, for example, the following: baculobacteria expression (Reilly, P. R., et al., BACULOBACTERIA EXPRESSION VECTORS: A LABORATORY MANUAL (1992); Beames, et al., Biotechniques 11:378 (1991); Pharmingen; Clontech, Palo Alto, Calif.)), vaccinia expression systems (Earl, P. L., et al., “Expression of proteins in mammalian cells using vaccinia” In Current Protocols in Molecular Biology (F. M. Ausubel, et al. Eds.), Greene Publishing Associates & Wiley Interscience, New York (1991); Moss, B., et al., U.S. Pat. No. 5,135,855, issued Aug. 4, 1992), expression in bacteria (Ausubel, F. M., et al., CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, John Wiley and Sons, Inc., Media Pa.; Clontech), expression in yeast (Rosenberg, S. and Tekamp-Olson, P., U.S. Pat. No. RE35,749, issued, Mar. 17, 1998, herein incorporated by reference; Shuster, J. R., U.S. Pat. No. 5,629,203, issued May 13, 1997, herein incorporated by reference; Gellissen, G., et al., Antonie Van Leeuwenhoek, 62(l-2):79-93 (1992); Romanos, M. A., et al., Yeast 8(6):423-488 (1992); Goeddel, D. V., Methods in Enzymology 185 (1990); Guthrie, C., and G. R. Fink, Methods in Enzymology 194 (1991)), expression in mammalian cells (Clontech; Gibco-BRL, Ground Island, N.Y.; e.g., Chinese hamster ovary (CHO) cell lines (Haynes, J., et al., Nuc. Acid. Res. 11:687-706 (1983); 1983, Lau, Y. F., et al., Mol. Cell. Biol. 4:1469-1475 (1984); Kaufman, R. J., “Selection and coamplification of heterologous genes in mammalian cells,” in Methods in Enzymology, vol. 185, pp 537-566. Academic Press, Inc., San Diego Calif. (1991)), and expression in plant cells (plant cloning vectors, Clontech Laboratories, Inc., Palo-Alto, Calif., and Pharmacia LKB Biotechnology, Inc., Piscataway, N.J.; Hood, E., et al., J. Bacteriol. 168:1291-1301 (1986); Nagel, R., et al., FEMS Microbiol. Lett. 67:325 (1990); An, et al., “Binary Vectors”, and others in Plant Molecular Biology Manual A3: 1-19 (1988); Miki, B. L. A., et al., pp. 249-265, and others in Plant DNA Infectious Agents (Hohn, T., et al., eds.) Springer-Verlag, Wien, Austria, (1987); Plant Molecular Biology: Essential Techniques, P. G. Jones and J. M. Sutton, New York, J. Wiley, 1997; Miglani, Gurbachan Dictionary of Plant Genetics and Molecular Biology, New York, Food Products Press, 1998; Henry, R. J., Practical Applications of Plant Molecular Biology, New York, Chapman & Hall, 1997), relevant portion incorporated herein by reference.

[0155] As used herein, the term “effective amount” or “effective dose” refers to that amount of the peptide or protein T cell epitopes of the invention sufficient to induce immunity, to prevent and / or ameliorate an infection or to reduce at least one symptom of an infection and / or to enhance the efficacy of another dose of peptide or protein T cell epitopes. An effective dose may refer to the amount of peptide or protein T cell epitopes sufficient to delay or minimize the onset of an infection. An effective dose may also refer to the amount of peptide or protein T cell epitopes that provides a therapeutic benefit in the treatment or management of an infection. Further, an effective dose is the amount with respect to peptide or protein T cell epitopes of the invention alone, or in combination with other therapies, that provides a therapeutic benefit in the treatment or management of an infection. An effective dose may also be the amount sufficientto enhance a subject's (e.g., a human's) own immune response against a subsequent exposure to an infectious agent. Levels of immunity can be monitored, e.g., by measuring amounts of neutralizing secretory and / or serum antibodies, e.g., by plaque neutralization, complement fixation, enzyme-linked immunosorbent, or microneutralization assay. In the case of a vaccine, an “effective dose” is one that prevents disease and / or reduces the severity of symptoms. A "reduction" of a symptom or symptoms (and grammatical equivalents of this phrase) means decreasing of the severity or frequency of the symptom(s), or elimination of the symptom(s). A "prophylactically effective amount" of a drug is an amount of a drug that, when administered to a subject, will have the intended prophylactic effect, e.g., preventing or delaying the onset (or reoccurrence) of an injury, disease, pathology or condition, or reducing the likelihood of the onset (or reoccurrence) of an injury, disease, pathology, or condition, or their symptoms, in this case, an infectious disease, and more particularly, a Streptococcus pneumoniae infection. The full prophylactic effect does not necessarily occur by administration of one dose, and may occur only after administration of a series of doses. Thus, a prophylactically effective amount may be administered in one or more administrations. Guidance can be found in the literature for appropriate dosages for given classes of pharmaceutical products. For example, for the given parameter, an effective amount will show an increase or decrease of at least 5%, 10%, 15%, 20%, 25%, 40%, 50%, 60%, 75%, 80%, 90%, or at least 100%. Efficacy can also be expressed as “-fold” increase or decrease. For example, a therapeutically effective amount can have at least a 1.2-fold, 1.5-fold, 2-fold, 5-fold, or more effect over a control. The exact amounts will depend on the purpose of the treatment, and will be ascertainable by one skilled in the art using known techniques (see, e.g., Lieberman, Pharmaceutical Dosage Forms (vols. 1-3, 1992); Lloyd, The Art, Science and Technology of Pharmaceutical Compounding (1999); Pickar, Dosage Calculations (1999); and Remington: The Science and Practice of Pharmacy, 20th Edition, 2003, Gennaro, Ed., Lippincott, Williams & Wilkins), relevant portions incorporated herein by reference.

[0156] As used herein, the term “immune stimulator” refers to a compound that enhances an immune response via the body's own chemical messengers (cytokines). These molecules comprise various cytokines, lymphokines and chemokines with immunostimulatory, immunopotentiating, and pro-inflammatory activities, such as interferons, interleukins (e.g., IL-1, IL-2, IL-3, IL-4, IL-12, IL-13); growth factors (e.g., granulocyte-macrophage (GM)-colony stimulating factor (CSF)); and other immunostimulatory molecules, such as macrophage inflammatory factor, Flt3 ligand, B7.1; B7.2, etc. The immune stimulator molecules can be administered in the same formulation as peptide or protein T cell epitopes s of the invention, or can be administered separately. Either the protein or an expression vector encoding the protein can be administered to produce an immunostimulatory effect.

[0157] As used herein, in certain embodiments, the term “protective immune response” or “protective response” refers to an immune response mediated by antibodies against an infectious agent, which is exhibited by a vertebrate (e.g., a human), which prevents or ameliorates an infection or reduces at least one symptom thereof. Peptide and protein T cell epitopes of the invention can stimulate the production of antibodies that, for example, neutralize infectious agents, blocks infectious agents from entering cells, blocks replication of said infectious agents, and / or protect host cells from infection and destruction. In otherembodiments, the term can also refer to an immune response that is mediated by T-lymphocytes and / or other white blood cells against an infectious agent, exhibited by a vertebrate (e.g., a human), that prevents or ameliorates pathogenic infection or reduces at least one symptom thereof. Peptide and protein T cell epitopes of the invention can stimulate the T cell responses that, for example, neutralize infectious agents, kill bacteria infected cells, blocks infectious agents from entering cells, blocks replication of said infectious agents, and / or protect host cells from infection and destruction.

[0158] The terms “biological sample” or “sample” refer to materials obtained from or derived from a subject or patient. A biological sample includes sections of tissues such as biopsy and autopsy samples, and frozen sections taken for histological purposes. Such samples include bodily fluids such as blood and blood fractions or products (e.g., serum, plasma, platelets, red blood cells, and the like), sputum, tissue, cultured cells (e.g., primary cultures, explants, and transformed cells) stool, urine, synovial fluid, joint tissue, synovial tissue, synoviocytes, fibroblast-like synoviocytes, macrophage -like synoviocytes, immune cells, hematopoietic cells, fibroblasts, macrophages, T cells, etc. A biological sample is typically obtained from a eukaryotic organism, such as a mammal such as a primate e.g., chimpanzee or human; cow; dog; cat; a rodent, e.g., guinea pig, rat, mouse; rabbit; or a bird; reptile; or fish.

[0159] As used herein, a “cell” refers to a cell carrying out metabolic or other function sufficient to preserve or replicate its genomic DNA. A cell can be identified by well-known methods in the art including, for example, presence of an intact membrane, staining by a particular dye, ability to produce progeny or, in the case of a gamete, ability to combine with a second gamete to produce a viable offspring. Cells may include prokaryotic and eukaryotic cells. Prokaryotic cells include but are not limited to bacteria. Eukaryotic cells include but are not limited to yeast cells and cells derived from plants and animals, for example mammalian, insect (e.g., spodoptera) and human cells. Cells may be useful when they are naturally nonadherent or have been treated not to adhere to surfaces, for example by trypsinization.

[0160] As used herein, the term "contacting" is used in accordance with its plain ordinary meaning and refers to the process of allowing at least two distinct species to become sufficiently proximal to react, interact or physically touch. It should be appreciated, however, the resulting reaction product can be produced directly from a reaction between the added reagents or from an intermediate from one or more of the added reagents which can be produced in the reaction mixture. The term “contacting” may include allowing two species to react, interact, or physically touch, wherein the two species may be, for example, an amino acid sequence, protein, or peptide as provided herein and an immune cell, such as a T cell.

[0161] As used herein, a "control" sample or value refers to a sample that serves as a reference, usually a known reference, for comparison to a test sample. For example, a test sample can be taken from a test condition, e.g., in the presence of a test compound, and compared to samples from known conditions, e.g., in the absence of the test compound (negative control), or in the presence of a known compound (positive control). A control can also represent an average value gathered from a number of tests or results. One of skill in the art will recognize that controls can be designed for assessment of any number of parameters. For example, a control can be devised to compare therapeutic benefit based on pharmacological data (e.g.,half-life) or therapeutic measures (e.g., comparison of side effects). One of skill in the art will understand which controls are valuable in a given situation and be able to analyze data based on comparisons to control values. Controls are also valuable for determining the significance of data. For example, if values for a given parameter are widely variant in controls, variation in test samples will not be considered as significant.

[0162] The term “modulator” refers to a composition that increases or decreases the level of a target molecule or the function of a target molecule or the physical state of the target of the molecule relative to the absence of the modulator.

[0163] The term “modulate” is used in accordance with its plain ordinary meaning and refers to the act of changing or varying one or more properties. “Modulation” refers to the process of changing or varying one or more properties. For example, as applied to the effects of a modulator on a target protein, to modulate means to change by increasing or decreasing a property or function of the target molecule or the amount of the target molecule.

[0164] The terms “associated” or “associated with” in the context of a substance or substance activity or function associated with a disease (e.g. a protein associated disease, a cancer (e.g., cancer, inflammatory disease, autoimmune disease, or infectious disease)) means that the disease (e.g. cancer, inflammatory disease, autoimmune disease, or infectious disease) is caused by (in whole or in part), or a symptom of the disease is caused by (in whole or in part) the substance or substance activity or function. As used herein, what is described as being associated with a disease, if a causative agent, could be a target for treatment of the disease.

[0165] The term “aberrant” as used herein refers to different from normal. When used to describe enzymatic activity or protein function, aberrant refers to activity or function that is greater or less than a normal control or the average of normal non-diseased control samples. Aberrant activity may refer to an amount of activity that results in a disease, wherein returning the aberrant activity to a normal or nondisease-associated amount (e.g., by administering a compound or using a method as described herein), results in reduction of the disease or one or more disease symptoms.

[0166] The terms “subject” or "subject in need thereof refers to a living organism who is at risk of or prone to having a disease or condition, or who is suffering from a disease or condition that can be treated by administration of a composition or pharmaceutical composition as provided herein. Non-limiting examples include humans and other primates, but also includes non-human primates such as chimpanzees and other apes and monkey species; farm animals such as cattle, sheep, pigs, goats and horses; domestic mammals such as dogs and cats; laboratory animals including rodents such as mice, rats and guinea pigs; birds, including domestic, wild and game birds such as chickens, turkeys and other gallinaceous birds, ducks, geese, and the like. The term does not denote a particular age. Thus, both adult and newborn individuals are intended to be covered. The system described above is intended for use in any of the above vertebrate species, since the immune systems of all of these vertebrates operate similarly.

[0167] The terms "disease" or "condition" refer to a state of being or health status of a patient or subjectcapable of being treated with a compound, pharmaceutical composition, or method provided herein. In embodiments, a patient or subject is human. In embodiments, the disease is Streptococcus pneumoniae infection. In certain alternative embodiments, the disease is MTB infection. In still other embodiments, the disease is whooping cough.

[0168] As used herein, "treatment" or "treating," or "palliating" or "ameliorating" are used interchangeably herein. These terms refer to an approach for obtaining beneficial or desired results including but not limited to therapeutic benefit and / or a prophylactic benefit. By therapeutic benefit is meant eradication or amelioration of the underlying disorder being treated or the disorder resulting from bacterial infection. Also, a therapeutic benefit is achieved with the eradication or amelioration of one or more of the physiological symptoms associated with bacterial infection or the underlying disorder such that an improvement is observed in the patient, notwithstanding that the patient may still be afflicted with the underlying disorder or may still be infected. For prophylactic benefit, the compositions may be administered to a patient at risk of bacterial infection, of developing a particular disease, or to a patient reporting one or more of the physiological symptoms of a disease, even though a diagnosis of this disease may not have been made. Treatment includes preventing the infection or disease, that is, causing the clinical symptoms of the disease not to develop by administration of a protective composition prior to infection or the induction of the disease; suppressing the disease, that is, causing the clinical symptoms of the disease or infection not to develop by administration of a protective composition after the inductive event or infection but prior to the clinical appearance or reappearance of the disease; inhibiting the disease, that is, arresting the development of clinical symptoms by administration of a protective composition after their initial appearance; preventing re-occurring of the disease and / or relieving the disease, that is, causing the regression of clinical symptoms by administration of a protective composition after their initial appearance. “Treatment” can also refer to any of (i) the prevention of infection or reinfection, as in a traditional vaccine, (ii) the reduction or elimination of symptoms, and (iii) the substantial or complete elimination of the pathogen in question. Treatment may be affected prophylactically (prior to infection) or therapeutically (following infection).

[0169] In addition, in certain embodiments, “treatment,” “treat,” or “treating” refers to a method of reducing the effects of one or more symptoms of infection with Streptococcus pneumoniae. Thus, in the disclosed method, treatment can refer to a 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% reduction in the severity of an established infection, disease, condition, or symptom of the infection, disease or condition. For example, a method for treating a disease is considered to be a treatment if there is a 10% reduction in one or more symptoms of the disease in a subject as compared to a control. Thus, the reduction can be a 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or any percent reduction in between 10% and 100% as compared to native or control levels. It is understood that treatment does not necessarily refer to a cure or complete ablation of the disease, condition, or symptoms of the disease or condition and / or complete prevention of infection. Further, as used herein, references to decreasing, reducing, or inhibiting include a change of 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or greater as compared to a control level and such terms can include but do not necessarily include complete elimination.

[0170] As used herein the terms “diagnose” or “diagnosing” refers to recognition of an infection, disease or condition by signs and symptoms. Diagnosing can refer to determination of whether a subject has an infection or disease, including distinguishing between latent or active infection. Diagnosis may refer to determination of the type of disease or condition a subject has or the type of bacteria the subject is infected with.

[0171] Diagnostic agents provided herein include any such agent, which are well-known in the relevant art. Among imaging agents are fluorescent and luminescent substances, including, but not limited to, a variety of organic or inorganic small molecules commonly referred to as "dyes," "labels," or "indicators." Examples include fluorescein, rhodamine, acridine dyes, Alexa dyes, and cyanine dyes. Enzymes that may be used as imaging agents in accordance with the embodiments of the disclosure include, but are not limited to, horseradish peroxidase, alkaline phosphatase, acid phosphatase, glucose oxidase, [3-galactosidase, [3-glucoronidase or [3-lactamase. Such enzymes may be used in combination with a chromogen, a Anorogenic compound or a luminogenic compound to generate a detectable signal.

[0172] The peptide(s) or protein(s) of the present invention can also be used in binding assays including, but are not limited to, immunoassays such as competitive and non-competitive assay systems using techniques such as western blots, radioimmunoassays, ELISA (enzyme linked immunosorbent assay), "sandwich" immunoassays, Meso Scale Discovery (MSD, Gaithersburg, Md.), immunoprecipitation assays, ELISPOT, precipitin reactions, gel diffusion precipitin reactions, immunodiffusion assays, agglutination assays, complement-fixation assays, immunoradiometric assays, Auorescent immunoassays, and protein A immunoassays. Such assays are routine and well known in the art (see, e.g., Ausubel et al., eds, 1994, Current Protocols in Molecular Biology, Vol. 1, John Wiley & Sons, Inc., New York, relevant portions incorporated herein by reference).

[0173] Radioactive substances that may be used as imaging agents in accordance with the embodiments of the disclosure include, but are not limited to,18F,32P,33P,45Ti,47Sc,52Fe,59Fe,62Cu,64Cu,67Cu,67Ga,68Ga,77As,86Y,90Y.89Sr,89Zr,94Tc,94Tc,99mTc, "Mo,105Pd,105Rh,mAg,inIn,123I,124I,125I,131I,142Pr,143Pr,149Pm,153Sm,154'1581Gd,161Tb,166Dy,166Ho,169Er,175Lu,177Lu,186Re,188Re,189Re,194Ir,198Au,199Au,211At,211Pb,212Bi,212Pb,213Bi,223Ra and225Ac. Paramagnetic ions that may be used as additional imaging agents in accordance with the embodiments of the disclosure include, but are not limited to, ions of transition and lanthanide metals (e.g., metals having atomic numbers of 21-29, 42, 43, 44, or 57-71). These metals include ions of Cr, V, Mn, Fe, Co, Ni, Cu, La, Ce, Pr, Nd, Pm, Sm, Eu, Gd, Tb, Dy, Ho, Er, Tm, Yb and Lu.

[0174] When the imaging agent is a radioactive metal or paramagnetic ion, the agent may be reacted with another long-tailed reagent having a long tail with one or more chelating groups attached to the long tail for binding to these ions. The long tail may be a polymer such as a polylysine, polysaccharide, or other derivatized or derivatizable chain having pendant groups to which the metals or ions may be added for binding. Examples of chelating groups that may be used according to the disclosure include, but are not limited to, ethylenediaminetetraacetic acid (EDTA), diethylenetriaminepentaacetic acid (DTPA), DOTA,NOTA, NETA, TETA, porphyrins, polyamines, crown ethers, bis-thiosemicarbazones, polyoximes, and like groups.

[0175] The terms “dose” and “dosage” are used interchangeably herein. A dose refers to the amount of active ingredient given to an individual at each administration. The dose will vary depending on a number of factors, including the range of normal doses for a given therapy, frequency of administration; size and tolerance of the individual; severity of the condition; risk of side effects; and the route of administration. One of skill will recognize that the dose can be modified depending on the above factors or based on therapeutic progress. The term “dosage form” refers to the particular format of the pharmaceutical or pharmaceutical composition, and depends on the route of administration. For example, a dosage form can be in a liquid form for nebulization, e.g., for inhalants, in a tablet or liquid, e.g., for oral delivery, or a saline solution, e.g., for injection.

[0176] As used herein, the term "administering" means oral administration, administration as a suppository, topical contact, intravenous, intraperitoneal, intramuscular, intralesional, intrathecal, intranasal or subcutaneous administration, or the implantation of a slow-release device, e.g., a mini -osmotic pump, to a subject. Administration is by any route, including parenteral and transmucosal (e.g., buccal, sublingual, palatal, gingival, nasal, vaginal, rectal, ortransdermal). Parenteral administration includes, e.g., intravenous, intramuscular, intra-arteriole, intradermal, subcutaneous, intraperitoneal, intraventricular, and intracranial. Other modes of delivery include, but are not limited to, the use of liposomal formulations, intravenous infusion, transdermal patches, etc. By "co-administer" it is meant that a composition described herein is administered at the same time, just prior to, or just after the administration of one or more additional therapies, for example cancer therapies such as chemotherapy, hormonal therapy, radiotherapy, or immunotherapy. The compounds of the invention can be administered alone or can be co-administered to the patient. Co-administration is meant to include simultaneous or sequential administration of the compounds individually or in combination (more than one compound). Thus, the preparations can also be combined, when desired, with other active substances (e.g., to reduce metabolic degradation). The compositions of the present invention can be delivered by transdermally, by a topical route, formulated as applicator sticks, solutions, suspensions, emulsions, gels, creams, ointments, pastes, jellies, paints, powders, and aerosols.

[0177] Formulations suitable for oral administration can consist of (a) liquid solutions, such as an effective amount of the antibodies provided herein suspended in diluents, such as water, saline or PEG 400; (b) capsules, sachets or tablets, each containing a predetermined amount of the active ingredient, as liquids, solids, granules or gelatin; (c) suspensions in an appropriate liquid; and (d) suitable emulsions. Tablet forms can include one or more of lactose, sucrose, mannitol, sorbitol, calcium phosphates, com starch, potato starch, microcrystalline cellulose, gelatin, colloidal silicon dioxide, talc, magnesium stearate, stearic acid, and other excipients, colorants, fillers, binders, diluents, buffering agents, moistening agents, preservatives, flavoring agents, dyes, disintegrating agents, and pharmaceutically compatible carriers. Lozenge forms can comprise the active ingredient in a flavor, e.g., sucrose, as well as pastilles comprising the active ingredient in an inert base, such as gelatin and glycerin or sucrose and acacia emulsions, gels, and the like containing,in addition to the active ingredient, carriers known in the art.

[0178] Pharmaceutical compositions can also include large, slowly metabolized macromolecules such as proteins, polysaccharides such as chitosan, polylactic acids, polyglycolic acids and copolymers (such as latex functionalized sepharose (TM), agarose, cellulose, and the like), polymeric amino acids, amino acid copolymers, and lipid aggregates (such as oil droplets or liposomes). Additionally, these carriers can function as immunostimulating agents (i.e., adjuvants).

[0179] The term "adjuvant" refers to a compound that when administered in conjunction with the compositions provided herein including embodiments thereof, augments the composition’s immune response. Generally, adjuvants are non-toxic, have high-purity, are degradable, and are stable.

[0180] Adjuvants can augment an immune response by several mechanisms including lymphocyte recruitment, stimulation of B and / or T cells, and stimulation of macrophages. The adjuvant increases the titer of induced antibodies and / or the binding affinity of induced antibodies relative to the situation if the immunogen were used alone. A variety of adjuvants can be used in combination with the agents provided herein including embodiments thereof, to elicit an immune response. Preferred adjuvants augment the intrinsic response to an immunogen without causing conformational changes in the immunogen that affect the qualitative form of the response. Preferred adjuvants include aluminum hydroxide and aluminum phosphate, 3 De-O-acylated monophosphoryl lipid A (MPL™) (see GB 2220211 (RIBI ImmunoChem Research Inc., Hamilton, Montana, now part of Corixa). Stimulon™ QS-21 is a triterpene glycoside or saponin isolated from the bark of the Quillaja Saponaria Molina tree found in South America (see Kensil et al., in Vaccine Design: The Subunit and Adjuvant Approach (eds. Powell & Newman, Plenum Press, NY, 1995); US Patent No. 5,057,540), (Aquila BioPharmaceuticals, Framingham, MA). Other adjuvants are oil in water emulsions (such as squalene or peanut oil), optionally in combination with immune stimulants, such as monophosphoryl lipid A (see Stoute etal.,N. Engl. J. Med. 336, 86-91 (1997)), pluronic polymers, and killed mycobacteria. Another adjuvant is CpG (WO 98 / 40100). Adjuvants can be administered as a component of a therapeutic composition with an active agent or can be administered separately, before, concurrently with, or after administration of the therapeutic agent.

[0181] Other adjuvants contemplated forthe invention are saponin adjuvants, such as Stimulon™ (QS-21, Aquila, Framingham, MA) or particles generated therefrom such as ISCOMs (immunostimulating complexes) and ISCOMATRIX. Other adjuvants include RC-529, GM-CSF and Complete Freund's Adjuvant (CFA) and Incomplete Freund's Adjuvant (IFA). Other adjuvants include cytokines, such as interleukins (e.g., IL-1 a and P peptides, IL-2, IL-4, IL-6, IL-12, IL-13, and IL-15), macrophage colony stimulating factor (M-CSF), granulocyte -macrophage colony stimulating factor (GM-CSF), tumor necrosis factor (TNF), chemokines, such as MIPla and and RANTES. Another class of adjuvants is glycolipid analogues including N-glycosylamides, N-glycosylureas and N-glycosylcarbamates, each of which is substituted in the sugar residue by an amino acid, as immuno-modulators or adjuvants (see US Pat. No.4,855,283). Heat shock proteins, e.g., HSP70 and HSP90, may also be used as adjuvants.

[0182] Suitable formulations for rectal administration include, for example, suppositories, which consistof the packaged nucleic acid with a suppository base. Suitable suppository bases include natural or synthetic triglycerides or paraffin hydrocarbons. In addition, it is also possible to use gelatin rectal capsules which consist of a combination of the compound of choice with a base, including, for example, liquid triglycerides, polyethylene glycols, and paraffin hydrocarbons.

[0183] Formulations suitable for parenteral administration, such as, for example, by intraarticular (in the joints), intravenous, intramuscular, intratumoral, intradermal, intraperitoneal, and subcutaneous routes, include aqueous and non-aqueous, isotonic sterile injection solutions, which can contain antioxidants, buffers, bacteriostats, and solutes that render the formulation isotonic with the blood of the intended recipient, and aqueous and non-aqueous sterile suspensions that can include suspending agents, solubilizers, thickening agents, stabilizers, and preservatives. In the practice of this invention, compositions can be administered, for example, by intravenous infusion, orally, topically, intraperitoneally, intravesically or intrathecally. Parenteral administration, oral administration, and intravenous administration are the preferred methods of administration. The formulations of compounds can be presented in unit-dose or multi-dose sealed containers, such as ampules and vials.

[0184] Injection solutions and suspensions can be prepared from sterile powders, granules, and tablets of the kind previously described. Cells transduced by nucleic acids for ex vivo therapy can also be administered intravenously or parenterally as described above.

[0185] The pharmaceutical preparation is preferably in unit dosage form. In such form the preparation is subdivided into unit doses containing appropriate quantities of the active component. The unit dosage form can be a packaged preparation, the package containing discrete quantities of preparation, such as packeted tablets, capsules, and powders in vials or ampoules. Also, the unit dosage form can be a capsule, tablet, cachet, or lozenge itself, or it can be the appropriate number of any of these in packaged form. The composition can, if desired, also contain other compatible therapeutic agents.

[0186] The combined administration contemplates co-administration, using separate formulations or a single pharmaceutical formulation, and consecutive administration in either order, wherein preferably there is a time period while both (or all) active agents simultaneously exert their biological activities.

[0187] Effective doses of the compositions provided herein vary depending upon many different factors, including means of administration, target site, physiological state of the patient, whether the patient is human or an animal, other medications administered, and whether treatment is prophylactic or therapeutic. However, a person of ordinary skill in the art would immediately recognize appropriate and / or equivalent doses looking at dosages of approved compositions for treating and preventing cancer for guidance.

[0188] As used herein, the term “pharmaceutically acceptable” is used synonymously with “physiologically acceptable” and “pharmacologically acceptable”. A pharmaceutical composition will generally comprise agents for buffering and preservation in storage, and can include buffers and carriers for appropriate delivery, depending on the route of administration. As used herein, the terms “pharmaceutically acceptable” or “pharmacologically acceptable” refer to a material which is not biologically or otherwise undesirable, i.e., the material may be administered to an individual in aformulation or composition without causing any unacceptable biological effects or interacting in a deleterious manner with any of the components of the composition in which it is contained.

[0189] “Pharmaceutically acceptable excipient” and “pharmaceutically acceptable carrier” refer to a substance that aids the administration of an active agent to and absorption by a subject and can be included in the compositions of the present invention without causing a significant adverse toxicological effect on the patient. Non-limiting examples of pharmaceutically acceptable excipients include water, NaCl, normal saline solutions, lactated Ringer’s, normal sucrose, normal glucose, binders, fillers, disintegrants, lubricants, coatings, sweeteners, flavors, salt solutions (such as Ringer's solution), alcohols, oils, gelatins, carbohydrates such as lactose, amylose or starch, fatty acid esters, hydroxymethycellulose, polyvinyl pyrrolidine, and colors, and the like. Such preparations can be sterilized and, if desired, mixed with auxiliary agents such as lubricants, preservatives, stabilizers, wetting agents, emulsifiers, salts for influencing osmotic pressure, buffers, coloring, and / or aromatic substances, and the like., that do not deleteriously react with the compounds of the invention. One of skill in the art will recognize that other pharmaceutical excipients are useful in the present invention.

[0190] The term "pharmaceutically acceptable salt" refers to salts derived from a variety of organic and inorganic counter ions well known in the art and include, by way of example only, sodium, potassium, calcium, magnesium, ammonium, tetraalkylammonium, and the like; and when the molecule contains a basic functionality, salts of organic or inorganic acids, such as hydrochloride, hydrobromide, tartrate, mesylate, acetate, maleate, oxalate and the like.

[0191] The term "preparation" is intended to include the formulation of the active compound with encapsulating material as a carrier providing a capsule in which the active component with or without other carriers, is surrounded by a carrier, which is thus in association with it. Similarly, cachets and lozenges are included. Tablets, powders, capsules, pills, cachets, and lozenges can be used as solid dosage forms suitable for oral administration.

[0192] The pharmaceutical preparation is optionally in unit dosage form. In such form the preparation is subdivided into unit doses containing appropriate quantities of the active component. The unit dosage form can be a packaged preparation, the package containing discrete quantities of preparation, such as packeted tablets, capsules, and powders in vials or ampoules. Also, the unit dosage form can be a capsule, tablet, cachet, or lozenge itself, or it can be the appropriate number of any of these in packaged form. The unit dosage form can be of a frozen dispersion.

[0193] The compositions of the present invention may additionally include components to provide sustained release and / or comfort. Such components include high molecular weight, anionic mucomimetic polymers, gelling polysaccharides and finely-divided drug carrier substrates. These components are discussed in greater detail in U.S. Pat. Nos. 4,911,920; 5,403,841; 5,212,162; and 4,861,760. The entire contents of these patents are incorporated herein by reference in their entirety for all purposes. The compositions of the present invention can also be delivered as microspheres for slow release in the body. For example, microspheres can be administered via intradermal injection of drug-containing microspheres,which slowly release subcutaneously (see Rao, J. Biomater Sci. Polym. Ed. 7:623-645, 1995; as biodegradable and injectable gel formulations (see, e.g., Gao Pharm. Res. 12:857-863, 1995); or, as microspheres for oral administration (see, e.g., Eyles, J. Pharm. Pharmacol. 49:669-674, 1997). In embodiments, the formulations of the compositions of the present invention can be delivered by the use of liposomes which fuse with the cellular membrane or are endocytosed, i.e., by employing receptor ligands attached to the liposome, that bind to surface membrane protein receptors of the cell resulting in endocytosis. By using liposomes, particularly where the liposome surface carries receptor ligands specific for target cells, or are otherwise preferentially directed to a specific organ, one can focus the delivery of the compositions of the present invention into the target cells in vivo. (See, e.g., Al -Muhammed, J. Microencapsul. 13:293-306, 1996; Chonn, Curr. Opin. Biotechnol. 6:698-708, 1995; Ostro, Am. J. Hosp. Pharm. 46:1576-1587, 1989). The compositions of the present invention can also be delivered as nanoparticles.

[0194] A person of skill in the art would readily recognize that steps of various above-described methods can be performed by one or more programmed computers, each having one or more computer processors. Herein, some embodiments are also intended to cover program storage devices, e.g., digital data storage media, which are machine or computer-readable and encode machine-executable or computer-executable programs of instructions, wherein the instructions perform some or all of the steps of said above-described methods. The program storage devices may be, e.g., digital memories, magnetic storage media such as magnetic disks and magnetic tapes, hard drives, or optically readable digital data storage media. The embodiments are also intended to cover computers programmed to perform said steps of the abovedescribed methods.

[0195] A risk score of the present invention may be calculated with an algorithm using well-known statistical analysis techniques. Non-limiting examples of statistical analysis techniques that may be used to calculate the risk score include cross-correlation, Principal Components Analysis (PCA), factor rotation, Logistic Regression (LogReg), Linear Discriminant Analysis (LDA), Eigengene Linear Discriminant Analysis (ELDA), Support Vector Machines (SVM), Random Lorest (RF), Recursive Partitioning Tree (RPART), related decision tree classification techniques, Shrunken Centroids (SC), StepAIC, Kth-Nearest Neighbor, Boosting, Decision Trees, Neural Networks, Bayesian Networks, Support Vector Machines, and Hidden Markov Models, Linear Regression or classification algorithms, Nonlinear Regression or classification algorithms, analysis of variants (ANOVA), hierarchical analysis or clustering algorithms; hierarchical algorithms using decision trees; kernel based machine algorithms such as kernel partial least squares algorithms, kernel matching pursuit algorithms, kernel Fisher's discriminate analysis algorithms, or kernel principal components analysis algorithms. In preferred embodiments, the risk score may be calculated using a random forest algorithm using the concentrations of three or more sample analytes in the panel of biomarkers. In an exemplary embodiment, the risk score is calculated as described in the examples.

[0196] The functions of the various elements shown in the figures, including any functional blocks labeled as "modules", may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with the appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term "module" should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), read-only memory (ROM) for storing software, random access memory (RAM), and nonvolatile storage. Other hardware, conventional and / or custom, may also be included.

[0197] It is contemplated that any embodiment discussed in this specification can be implemented with respect to any method, kit, reagent, or composition of the invention, and vice versa. Furthermore, compositions of the invention can be used to achieve methods of the invention.

[0198] It will be understood that particular embodiments described herein are shown by way of illustration and not as limitations of the invention. The principal features of this invention can be employed in various embodiments without departing from the scope of the invention. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, numerous equivalents to the specific procedures described herein. Such equivalents are considered to be within the scope of this invention and are covered by the claims.

[0199] All publications and patent applications mentioned in the specification are indicative of the level of skill of those skilled in the art to which this invention pertains. All publications and patent applications are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference.

[0200] The use of the word “a” or “an” when used in conjunction with the term “comprising” in the claims and / or the specification may mean “one,” but it is also consistent with the meaning of “one or more,” “at least one,” and “one or more than one.” The use of the term “or” in the claims is used to mean “and / or” unless explicitly indicated to refer to alternatives only or the alternatives are mutually exclusive, although the disclosure supports a definition that refers to only alternatives and “and / or.” Throughout this application, the term “about” is used to indicate that a value includes the inherent variation of error for the device, the method being employed to determine the value, or the variation that exists among the study subjects.

[0201] As used in this specification and claim(s), the words “comprising” (and any form of comprising, such as “comprise” and “comprises”), “having” (and any form of having, such as “have” and “has”), “including” (and any form of including, such as “includes” and “include”) or “containing” (and any form of containing, such as “contains” and “contain”) are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. In embodiments of any of the compositions and methods provided herein, “comprising” may be replaced with “consisting essentially of’ or “consisting of’. As used herein,the phrase “consisting essentially of’ requires the specified integer(s) or steps as well as those that do not materially affect the character or function of the claimed invention. As used herein, the term “consisting” is used to indicate the presence of the recited integer (e.g., a feature, an element, a characteristic, a property, a method / process step or a limitation) or group of integers (e.g., feature(s), element(s), characteristic(s), propertie(s), method / process steps or limitation(s)) only.

[0202] The term “or combinations thereof’ as used herein refers to all permutations and combinations of the listed items preceding the term. For example, “A, B, C, or combinations thereof’ is intended to include at least one of: A, B, C, AB, AC, BC, or ABC, and if order is important in a particular context, also BA, CA, CB, CBA, BCA, ACB, BAC, or CAB. Continuing with this example, expressly included are combinations that contain repeats of one or more item or term, such as BB, AAA, AB, BBC, AAABCCCC, CBBAAA, CABABB, and so forth. The skilled artisan will understand that typically there is no limit on the number of items or terms in any combination, unless otherwise apparent from the context.

[0203] As used herein, words of approximation such as, without limitation, “about”, “substantial” or “substantially” refers to a condition that when so modified is understood to not necessarily be absolute or perfect but would be considered close enough to those of ordinary skill in the art to warrant designating the condition as being present. The extent to which the description may vary will depend on how great a change can be instituted and still have one of ordinary skilled in the art recognize the modified feature as still having the required characteristics and capabilities of the unmodified feature. In general, but subject to the preceding discussion, a numerical value herein that is modified by a word of approximation such as “about” may vary from the stated value by at least ±1, 2, 3, 4, 5, 6, 7, 10, 12 or 15%.

[0204] Additionally, the section headings herein are provided for consistency with the suggestions under 37 CFR 1.77 or otherwise to provide organizational cues. These headings shall not limit or characterize the invention(s) set out in any claims that may issue from this disclosure. Specifically, and by way of example, although the headings refer to a “Field of Invention,” such claims should not be limited by the language under this heading to describe the so-called technical field. Further, a description of technology in the “Background of the Invention” section is not to be construed as an admission that technology is prior art to any invention(s) in this disclosure. Neither is the “Summary” to be considered a characterization of the invention(s) set forth in issued claims. Furthermore, any reference in this disclosure to “invention” in the singular should not be used to argue that there is only a single point of novelty in this disclosure. Multiple inventions may be set forth according to the limitations of the multiple claims issuing from this disclosure, and such claims accordingly define the invention(s), and their equivalents, that are protected thereby. In all instances, the scope of such claims shall be considered on their own merits in light of this disclosure, but should not be constrained by the headings set forth herein.

[0205] All of the compositions and / or methods disclosed and claimed herein can be made and executed without undue experimentation in light of the present disclosure. While the compositions and methods of this invention have been described in terms of preferred embodiments, it will be apparent to those of skill in the art that variations may be applied to the compositions and / or methods and in the steps or in thesequence of steps of the method described herein without departing from the concept, spirit and scope of the invention. All such similar substitutes and modifications apparent to those skilled in the art are deemed to be within the spirit, scope and concept of the invention as defined by the appended claims.

[0206] To aid the Patent Office, and any readers of any patent issued on this application in interpreting the claims appended hereto, applicants wish to note that they do not intend any of the appended claims to invoke paragraph 6 of 35 U.S.C. § 112, U.S.C. § 112 paragraph (f), or equivalent, as it exists on the date of filing hereof unless the words “means for” or “step for” are explicitly used in the particular claim.

[0207] For each of the claims, each dependent claim can depend both from the independent claim and from each of the prior dependent claims for each and every claim so long as the prior claim provides a proper antecedent basis for a claim term or element.REFERENCES

[0208] 1. Unanue ER. Antigen-presenting function of the macrophage. Annu Rev Immunol.1984;2:395-428.

[0209] 2. Peters B, Nielsen M, Sette A. T Cell Epitope Predictions. Annu Rev Immunol. 2020 Apr 26;38(Volume 38, 2020): 123-45.

[0210] 3. Paul S, Lindestam Arlehamn CS, Scriba TJ, Dillon MBC, Oseroff C, Hinz D, et al. Development and validation of a broad scheme for prediction of HLA class II restricted T cell epitopes. J Immunol Methods. 2015 Jul l;422:28-34.

[0211] 4. Kosaloglu-Yalcin Z, Lee J, Greenbaum J, Schoenberger SP, Miller A, Kim YJ, et al. Combined assessment of MHC binding and antigen abundance improves T cell epitope predictions. iScience. 2022 Feb I8;25(2): 103850.

[0212] 5. Abelin JG, Keskin DB, Sarkizova S, Hartigan CR, Zhang W, Sidney J, et al. Mass Spectrometry Profiling of HLA-Associated Peptidomes in Mono-allelic Cells Enables More Accurate Epitope Prediction. Immunity. 2017 Feb 21;46(2):315-26.

[0213] 6. Bassani-Stemberg M, Pletscher-Frankild S, Jensen LJ, Mann M. Mass spectrometry of human leukocyte antigen class I peptidomes reveals strong effects of protein abundance and turnover on antigen presentation. Mol Cell Proteomics MCP. 2015 Mar;14(3):658-73.

[0214] 7. Arlehamn CSL, Gerasimova A, Mele F, Henderson R, Swann J, Greenbaum JA, et al. Memory T Cells in Latent Mycobacterium tuberculosis Infection Are Directed against Three Antigenic Islands and Largely Contained in a CXCR3+CCR6+ Thl Subset. PLOS Pathog. 2013 Jan 24;9(l):e 1003130.

[0215] 8. Bresciani A, Paul S, Schommer N, Dillon MB, Bancroft T, Greenbaum J, et al. T-cell recognition is shaped by epitope sequence conservation in the host proteome and microbiome. Immunology. 2016 May;148(l):34-9.

[0216] 9. Lewis SA, Sutherland A, Soldevila F, Westemberg L, Aoki M, Frazier A, et al. Identification of cow milk epitopes to characterize and quantify disease-specific T cells in allergic children.J Allergy Clin Immunol. 2023 Nov;152(5): 1196-209.

[0217] 10. Gadsby NJ, Musher DM. The Microbial Etiology of Community-Acquired Pneumonia in Adults: from Classical Bacteriology to Host Transcriptional Signatures. Clin Microbiol Rev. 2022 Dec 21;35(4):e0001522.

[0218] 11. Yu NY, Wagner JR, Laird MR, Meili G, Rey S, Lo R, etal. PSORTb 3.0: improved protein subcellular localization prediction with refined localization subcategories and predictive capabilities for all prokaryotes. Bioinformatics. 2010 Jul 1;26(13): 1608-15.

[0219] 12. Panda S, Morgan J, Cheng C, Saito M, Gilman RH, Ciobanu N, et al. Identification of differentially recognized T cell epitopes in the spectrum of tuberculosis infection. Nat Commun. 2024 Jan 26;15(1):765.

[0220] 13. Antunes R da S, Garrigan E, Quiambao LG, Dhanda SK, Marrama D, Westemberg L, et al. T cell reactivity to Bordetella pertussis is highly diverse regardless of childhood vaccination. Cell Host Microbe. 2023 Aug 9;31(8): 1404-1416.e4.

[0221] 14. Malaga W, Payros D, Meunier E, Frigui W, Sayes F, Pawlik A, et al. Natural mutations in the sensor kinase of the PhoPR two-component regulatory system modulate virulence of ancestor-like tuberculosis bacilli. PLoS Pathog. 2023 Jul;19(7):el011437.

[0222] 15. Rivera-Millot A, Slupek S, Chatagnon J, Roy G, Saliou JM, Billon G, et al. Streamlined copper defenses make Bordetella pertussis reliant on custom-made operon. Commun Biol. 2021 Jan 8;4:46.

[0223] 16. Ferrandiz MJ, Martin-Galiano AJ, Amanz C, Camacho-Soguero I, Tirado-Velez JM, de la Campa AG. An increase in negative supercoiling in bacteria reveals topology-reacting gene clusters and a homeostatic response mediated by the DNA topoisomerase I gene. Nucleic Acids Res. 2016 Sep 6;44(15):7292-303.

[0224] 17. PEPMatch: a tool to identify short peptide sequence matches in large sets of proteins | BMC Bioinformatics | Full Text [Internet], [cited 2024 Sep 13], Available from: bmcbioinformatics.biomedcentral.com / articles / 10.1186 / s 12859-023-05606-4.

[0225] 18. The UniProt Consortium. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res. 2023 Jan 6;51(D1):D523-31.

[0226] 19. Richardson E, Trevizani R, Greenbaum JA, Carter H, Nielsen M, Peters B. The receiver operating characteristic curve accurately assesses imbalanced datasets. Patterns. 2024 May 31 ;5(6): 100994.

[0227] 20. da Silva Antunes R, Weiskopf D, Sidney J, Rubiro P, Peters B, Lindestam Arlehamn CS, et al. The MegaPool Approach to Characterize Adaptive CD4+ and CD8+ T Cell Responses. Curr Protoc.2023;3(ll):e934.

Claims

WHAT IS CLAIMED IS:

1. A method for generating a library of immunogenic peptide epitope molecules, the method comprising:(a) analyzing binding affinities of a set of possible immunogenic peptide epitopes for an epitope of interest from a pathogen by:selecting sequence data from a template antigen-binding molecule from a set of possible template antigen binding molecules, wherein the selected template is a different pathogen than the epitope of interest and wherein the set of possible template antigen-binding molecules consists of two or more known peptides that bind MHC and are presented to T cells, wherein the selecting sequence data comprises screening the set of possible template antigen-binding molecules based on at least two of:(1) gene expression and conservation;(2) one or more peptide-level features; and(3) one or more antigen-level features;(b) using two or more machine learning algorithms to train a model and combining into an ensemble model;(c) generating sequence data of one or more immunogenic peptide epitope molecules with the ensemble model;(d) synthesizing one or more immunogenic peptide epitope molecules from the sequence data of one or more immunogenic peptide epitope molecules; and(e) screening the one or more immunogenic peptide epitope molecules synthesized for immunogenicity, wherein the molecules trigger a response by T cells that recognize the one or more immunogenic peptide epitope molecules in the context of MHC Class I or Class II molecules.

2. The method of claim 1, further comprising selecting one or more immunogenic peptide epitope molecules that trigger a CD8 or a CD4 T cell response.

3. The method of claim 1 or claim 2, wherein the method generates a library of one or more immunogenic peptide epitope molecules that trigger an immune response in a host.

4. The method of claim 1, wherein the one or more peptide-level features are selected from Major Histocompatibility (MHC) Class II binding predictions, antigen strain conservation scores, and best match to a human proteome.

5. The method of claim 1, wherein the one or more antigen-level features are selected from RNA expression and subcellular localization prediction scores for each of: cytoplasm, cytoplasmic membrane, outer membrane, and extracellular.

6. The method of any one of claims 1 to 5, wherein MHC Class II binding is determined using NetMHCIIpan, peptide conservation for 0-3 substitutions is determined using PEPMatch, gene expression is determined using GEO Dataset, intracellular localization is determined using PSORTb, and protein existence levels using UniProt.

7. The method of claim 1, wherein the two or more machine learning algorithms are selected from Random Forest, Gradient Boosting, and XGBoost.

8. The method of claim 1, wherein the method is computer implemented.

9. The method of claim 1, wherein the model is trained with at least one of: Mycobacterium tuberculosis or Bordatella pertussis antigenic epitopes.

10. The method of claim 1, wherein the possible immunogenic peptide epitopes are viral, bacterial, fungal, protozoan, or helminthic epitopes.

11. The method of claim 1, wherein the possible immunogenic peptide epitopes are Streptococcus pneumoniae immunogenic peptide epitopes.

12. The method of claim 1, wherein the immunogenicity is in a human.

13. The method of claim 1, wherein the algorithm was used to train a model using grid search hyperparameter tuning with 10-fold cross-validation.

14. A composition comprising:one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant, or derivative thereof;a fusion protein comprising one or more amino acid sequences selected from SEQ ID NOS: 1-10; a pool of 2 or more or more peptides comprising, consisting of, or consisting essentially of amino acid sequences selected from SEQ ID NOS: 1-10; ora polynucleotide that encodes one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant, or derivative thereof.

15. The composition of claim 14, wherein the one or more peptides or proteins comprises, or wherein the fusion protein comprises two or more or more amino acid sequences selected from SEQ ID NOS: 1- 10, or a subsequence, portion, homologue, variant, or derivative thereof.

16. The composition of claim 14 or claim 15, wherein the amino acid sequence is selected from a Streptococcus pneumoniae T cell epitope selected from any one of those sequences of SEQ ID NOS: 1 to 10; and in certain examples SEQ ID NO: 1, SEQ ID NO:6, or both.

17. The composition of claim 14 or claim 15, wherein the composition comprises:one or more Streptococcus pneumoniae peptides amino acid sequences selected from SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant, or derivative thereof;a fusion protein comprising one or more amino acid sequences selected from SEQ ID NOS: 1-10; ora pool of two or more peptides selected from SEQ ID NOS: 1-10;a polynucleotide that encodes one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant, or derivative thereof; orthe peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both.

18. The composition of one of claims 14 to 17, wherein the peptide or protein comprises a Streptococcus pneumoniae T cell epitope.

19. The composition of any one of claims 14 to 18, wherein the one or more peptides or proteins comprises a Streptococcus pneumoniae CD8+ or CD4+ T cell epitope.

20. The composition of any one of claims 14 to 19, wherein the Streptococcus pneumoniae T cell epitope is conserved in another Streptococcus sp.

21. The composition of any one of claims 14 to 20, wherein the one or more peptides or proteins has a length from about 9-15, 15-20, 20-25, 25-30, 30-40, 40-50, 50-75 or 75-100 amino acids.

22. The composition of any one of claims 14 to 21, wherein the one or more peptides or proteins elicits, stimulates, induces, promotes, increases, or enhances a T cell response to Streptococcus pneumoniae.

23. The composition of claim 22, wherein the one or more peptides or proteins that elicits, stimulates, induces, promotes, increases, or enhances the T cell response to the Streptococcus pneumoniae is a Streptococcus pneumoniae protein or peptide, or a variant, homologue, derivative or subsequence thereof.

24. The composition of any one of claims 14 to 23, further comprising formulating the one or more peptides or proteins into an immunogenic formulation with an adjuvant.

25. The composition of claim 24, wherein the adjuvant is selected from the group consisting of adjuvant is selected from the group consisting of alum, aluminum hydroxide, aluminum phosphate, calcium phosphate hydroxide, cytosine-guanosine oligonucleotide (CpG-ODN) sequence, granulocyte macrophage colony stimulating factor (GM-CSF), monophosphoryl lipid A (MPL), poly(I:C), MF59, Quil A, N-acetyl muramyl-L-alanyl-D-isoglutamine (MDP), FIA, montanide, poly (DL-lactide-coglycolide), squalene, virosome, AS03, ASO4, IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7, IL-8, IL-10, IL- 12, IL-15, IL-17, IL-18, STING, CD40L, pathogen-associated molecular patterns (PAMPs), damage-associated molecular pattern molecules (DAMPs), Freund's complete adjuvant, Freund's incomplete adjuvant, transforming growth factor (TGF)-beta antibody or antagonists, A2aR antagonists, lipopolysaccharides (LPS), Fas ligand, Trail, lymphotactin, Mannan (M-FP), APG-2, Hsp70 and Hsp90, pattern recognition receptor ligands, TLR3 ligands, TLR4 ligands, TLR5 ligands, TLR7 / 8 ligands, and TLR9 ligands.

26. The composition of any one of claims 14 to 25, wherein the composition further comprises a modulator of immune response.

27. The composition of claim 26, wherein the modulator of immune response is a modulator of the innate immune response.

28. The composition of claim 14 or claim 15, wherein the modulator is Interleukin-6 (IL-6), Interferon-gamma (IFN-y), Transforming growth factor beta (TGF-P), or Interleukin- 10 (IL-10), or an agonist or antagonist thereof.

29. A composition comprising monomers or multimers of:peptides or proteins comprising, consisting of, or consisting essentially of:one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10; concatemers, subsequences, portions, homologues, variants, or derivatives thereof;a fusion protein comprising one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10;a polynucleotide that encodes one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant, or derivative thereof; orthe peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both.

30. A composition comprising one or more peptide-major histocompatibility complex (MHC) monomers or multimers, wherein the peptide-MHC monomer or multimer comprises a peptide comprising, consisting of, or consisting essentially of an amino acid sequence selected from at least one of SEQ ID NOS: 1-10, in a groove of the MHC monomer or multimer.

31. A composition comprising:one or more peptides or proteins comprising, consisting of, or consisting essentially of an amino acid sequence selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof;a fusion protein comprising one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10;a pool of two or more peptides selected from at least one of SEQ ID NOS: 1-10;a polynucleotide that encodes one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof; orthe peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both.

32. The composition of claim 31, wherein the one or more peptides or proteins comprises, or wherein the fusion protein comprises two or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof.

33. The composition of claim 31 or claim 32, wherein the protein or peptide comprises a Streptococcus pneumoniae T cell epitope.

34. The composition of any one of claims 31 to 32, wherein the one or more peptides or proteins comprises a Streptococcus pneumoniae CD8+ or CD4+ T cell epitope.

35. The composition of any one of claims 31 to 32, wherein the Streptococcus pneumoniae T cell epitope is not conserved in another Streptococcus pneumoniae.

36. The composition of any one of claims 31 to 32, wherein the Streptococcus pneumoniae T cell epitope is conserved in another Streptococcus pneumoniae.

37. The composition of any one of claims 31 to 36, wherein the one or more peptides or proteins has a length from about 9-15, 15-20, 20-25, 25-30, 30-40, 40-50, 50-75 or 75-100 amino acids.

38. The composition of any one of claims 31 to 37, wherein the one or more peptides or proteins elicits, stimulates, induces, promotes, increases or enhances a T cell response to Streptococcuspneumoniae .

39. The composition of any one of claims 31 to 38, wherein the one or more peptides or proteins that elicits, stimulates, induces, promotes, increases or enhances the T cell response to Streptococcus pneumoniae is a Streptococcus pneumoniae protein or peptide, or a variant, homologue, derivative or subsequence thereof.

40. The composition of any one of claims 31 to 39, further comprising formulating the one or more peptides or proteins into an immunogenic formulation with an adjuvant.

41. The composition of claim 40, wherein the adjuvant is selected from the group consisting of adjuvant is selected from the group consisting of alum, aluminum hydroxide, aluminum phosphate, calcium phosphate hydroxide, cytosine-guanosine oligonucleotide (CpG-ODN) sequence, granulocyte macrophage colony stimulating factor (GM-CSF), monophosphoryl lipid A (MPL), poly(I:C), MF59, Quil A, N-acetyl muramyl-L-alanyl-D-isoglutamine (MDP), FIA, montanide, poly (DL-lactide-coglycolide), squalene, virosome, AS03, ASO4, IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7, IL-8, IL-10, IL- 12, IL-15, IL-17, IL-18, STING, CD40L, pathogen-associated molecular patterns (PAMPs), damage-associated molecular pattern molecules (DAMPs), Freund's complete adjuvant, Freund's incomplete adjuvant, transforming growth factor (TGF)-beta antibody or antagonists, A2aR antagonists, lipopolysaccharides (LPS), Fas ligand, Trail, lymphotactin, Mannan (M-FP), APG-2, Hsp70 and Hsp90, pattern recognition receptor ligands, TLR3 ligands, TLR4 ligands, TLR5 ligands, TLR7 / 8 ligands, and TLR9 ligands.

42. The composition of any one of claims 31 to 41, wherein the composition further comprises a modulator of immune response.

43. The composition of claim 42, wherein the modulator of immune response is a modulator of the innate immune response.

44. The composition of claim 42 or claim 43, wherein the modulator is Interleukin-6 (IL-6), Interferon-gamma (IFN-g), Transforming growth factor beta (TGF-B), or Interleukin- 10 (IL- 10), or an agonist or antagonist thereof.

45. A composition comprising monomers or multimers of:one or more peptides or proteins comprising, consisting of, or consisting essentially of: one or more Streptococcus pneumoniae amino acid sequences selected from at least one of SEQ ID NOS: 1-10, concatemers, subsequences, portions, homologues, variants or derivatives thereof;a fusion protein comprising one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10);a polynucleotide that encodes one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from at least one of SEQ ID NOS: 1-10), or a subsequence, portion, homologue, variant or derivative thereof; orthe peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both.

46. A composition comprising one or more peptide-major histocompatibility complex (MHC) monomers or multimers, wherein the peptide-MHC monomer or multimer comprises a peptidecomprising, consisting of, or consisting essentially of an amino acid sequence selected from at least one of SEQ ID NOS: 1-10, in a groove of the (MHC) monomer or multimer.

47. A method for detecting the presence of: (i) a Streptococcus pneumoniae or (ii) an immune response relevant to Streptococcus pneumoniae infections, vaccines or therapies, including T cells responsive to one or more Streptococcus pneumoniae peptides, comprising:providing one or more proteins or peptides for detection of an amount or a relative amount of, and / or the activity of, and / or the state of antigen-specific T-cells;contacting a biological sample suspected of having Streptococcus pneumoniae -specific T-cells to one or more proteins or peptides for detection; anddetecting an amount or a relative amount of, and / or the activity of, and / or the state of antigenspecific T-cells in the biological sample, wherein the one or more proteins or peptides for detection comprise one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10, or comprise a pool of 2 or more or more amino acid sequences set forth in any one of SEQ ID NOS: 1-10).

48. The method of claim 47, wherein detecting the amount or a relative amount of, and / or activity of antigen-specific T-cells comprises one or more steps of identification or detection of the antigen-specific T-cells and measuring the amount of the antigen-specific T-cells.

49. The method of claim 47 or claim 48, wherein the one or more peptides or proteins comprises 2 or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10).

50. The method of any one of claims 47 to 49, wherein the detecting the amount or a relative amount of, and / or activity of antigen-specific T-cells comprises indirect detection and / or direct detection.

51. The method of any one of claims 47 to 50, wherein the method of detecting an immune response relevant to the Streptococcus pneumoniae comprises the following steps:providing an MHC monomer or an MHC multimer;contacting a population T-cells to the MHC monomer or MHC multimer; andmeasuring the number, activity or state of T-cells specific for the MHC monomer or MHC multimer.

52. The method of claim 47, wherein the MHC monomer or MHC multimer comprises a protein or peptide of the Streptococcus pneumoniae.

53. The method of claim 47, wherein the protein or peptide comprises a CD8+ or CD4+ T cell epitope.

54. The method of claim 47, wherein the T cell epitope is not conserved in another Streptococcus pneumoniae.

55. The method of claim 47, wherein the T cell epitope is conserved in another Streptococcus pneumoniae.

56. The method of any one of claims 47 to 55 wherein the protein or peptide has a length from about 9-15, 15-20, 20-25, 25-30, 30-40, 40-50, 50-75 or 75-100 amino acids.

57. The method of any one of claims 47 to 56, wherein the proteins or peptides comprise 2 or more amino acid sequences selected from any one of those sequences at least one of SEQ ID NOS: 1-10, or asubsequence, portion, homologue, variant or derivative thereof; or the peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both.

58. The method of any one of claims 47 to 57, further comprising detecting the presence or amount of the one or more peptides in a biological sample, or a response thereto, which is diagnostic of a latent or active Streptococcus pneumoniae infection.

59. The method of any one of claims 47 to 58, wherein detecting an amount or a relative amount of, and / or the activity of, and / or the state of antigen-specific T-cells in the biological sample comprises measuring one or more of a cytokine or lymphokine secretion assay, T cell proliferation, immunoprecipitation, immunoassay, ELISA, radioimmunoassay, immunofluorescence assay, Western Blot, FACS analysis, a competitive immunoassay, a noncompetitive immunoassay, a homogeneous immunoassay a heterogeneous immunoassay, a bioassay, a reporter assay, a luciferase assay, a microarray, a surface plasmon resonance detector, a florescence resonance energy transfer, immunocytochemistry, or a cell mediated assay, or a cytokine proliferation assay.

60. The method of any one of claims 47 to 59, further comprising administering a treatment comprising the composition of any one of claims 19-46 to the subject from which the biological sample was drawn that increases the amount or relative amount of, and / or activity of the antigen-specific T-cells.

61. A method for detecting the presence of: (i) Streptococcus pneumoniae or (ii) an immune response relevant to Streptococcus pneumoniae infections, vaccines or therapies, including T cells responsive to one or more Streptococcus pneumoniae peptides, comprising:providing one or more proteins or peptides for detection of an amount or a relative amount of, and / or the activity of, and / or the state of antigen-specific T-cells;contacting a biological sample suspected of having Streptococcus pneumoniae -specific T-cells to one or more proteins or peptides for detection; anddetecting an amount or a relative amount of, and / or the activity of, and / or the state of antigenspecific T-cells in the biological sample, wherein the one or more proteins or peptides for detection comprise one or more amino acid sequences set forth in those sequences selected from at least one of SEQ ID NOS: 1-10, or comprise a pool of 2 or more amino acid sequences set forth in those sequences selected from at least one of SEQ ID NOS: 1-10.

62. The method of claim 61, wherein detecting the amount or a relative amount of, and / or activity of antigen-specific T-cells comprises one or more steps of identification or detection of the antigen-specific T-cells and measuring the amount of the antigen-specific T-cells.

63. The method of claim 61 or claim 62, wherein the one or more peptides or proteins comprises 2 or more amino acid sequences selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10.

64. The method of any one of claims 61 to 63, wherein the detecting the amount or a relative amount of, and / or activity of antigen-specific T-cells comprises indirect detection and / or direct detection.

65. The method of any one of claims 61 to 64, wherein the method of detecting an immune response relevant to Streptococcus pneumoniae comprises the following steps:providing an MHC monomer or an MHC multimer;contacting a population T-cells to the MHC monomer or MHC multimer; andmeasuring the number, activity or state of T-cells specific for the MHC monomer or MHC multimer.

66. The method of claim 61, wherein the MHC monomer or MHC multimer comprises a protein or peptide of Streptococcus pneumoniae o / SEQ ID NO: 1 or SEQ ID NO:6, or both.

67. The method of claim 61, wherein the protein or peptide comprises a Streptococcus pneumoniae CD8+ or CD4+ T cell epitope.

68. The method of claim 55, wherein the Streptococcus pneumoniae T cell epitope is not conserved in another Streptococcus sp.

69. The method of claim 55, wherein the Streptococcus pneumoniae T cell epitope is conserved in another Streptococcus pneumoniae .

70. The method of any one of claims 49 to 57, wherein the protein or peptide has a length from about 9-15, 15-20, 20-25, 25-30, 30-40, 40-50, 50-75 or 75-100 amino acids.

71. The method of any one of claims 49 to 58, wherein the proteins or peptides comprise 2 or more amino acid sequences selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof.

72. The method of any one of claims 49 to 59, further comprising detecting the presence or amount of the one or more peptides in a biological sample, or a response thereto, which is diagnostic of a Streptococcus pneumoniae infection.

73. The method of any one of claims 49 to 60, wherein detecting an amount or a relative amount of, and / or the activity of, and / or the state of antigen-specific T-cells in the biological sample comprises measuring one or more of a cytokine or lymphokine secretion assay, T cell proliferation, immunoprecipitation, immunoassay, ELISA, radioimmunoassay, immunofluorescence assay, Western Blot, FACS analysis, a competitive immunoassay, a noncompetitive immunoassay, a homogeneous immunoassay a heterogeneous immunoassay, a bioassay, a reporter assay, a luciferase assay, a microarray, a surface plasmon resonance detector, a florescence resonance energy transfer, immunocytochemistry, or a cell mediated assay, or a cytokine proliferation assay.

74. The method of any one of claims 49 to 61, further comprising administering a treatment comprising the composition of any one of claims 19-46 to the subject from which the biological sample was drawn that increases the amount or relative amount of, and / or activity of the antigen-specific T-cells.

75. A method detecting a Streptococcus pneumoniae infection or exposure in a subject, the method comprising, consisting of, or consisting essentially of:contacting a biological sample from a subject with a composition of any one of claims 1 to 36; anddetermining if the composition elicits an immune response from the contacted cells, wherein the presence of an immune response indicates that the subject has been exposed to or infected with Streptococcus pneumoniae .

76. The method of claim 75, wherein the sample comprises T cells.

77. The method of claim 75 or claim 76, wherein the response comprises inducing, increasing, promoting or stimulating anti-Streptococcus pneumoniae activity of T cells.

78. The method of claim 75 or claim 76, wherein the T cells are CD8+ or CD4+ T cells.

79. The method of any one of claims 75 to 78, wherein the method comprises determining whether the subject has been infected by or exposed to the Streptococcus pneumoniae more than once by determining if the subject elicits a secondary T cell immune response profde that is different from a primary T cell immune response profile.

80. The method of any one of claims 75 to 79, further comprising diagnosing a latent or active Streptococcus pneumoniae infection or exposure in a subject, the method comprising contacting a biological sample from a subject with a composition of any one of claims 1 to 36, and determining if the composition elicits a T cell immune response, wherein the T cell immune response identifies that the subject has been infected with or exposed to a Streptococcus pneumoniae.

81. The method of any one of claims 75 to 80, wherein the method is conducted three or more days following the date of suspected infection by or exposure to a Streptococcus pneumoniae.

82. A method detecting Streptococcus pneumoniae infection or exposure in a subject, the method comprising, consisting of, or consisting essentially of:contacting a biological sample from a subject with a composition of any one of claims 19 to 34; and determining if the composition elicits an immune response from the contacted cells, wherein the presence of an immune response indicates that the subject has been exposed to or infected with Streptococcus pneumoniae.

83. The method of claim 82, wherein the sample comprises T cells.

84. The method of claim 82 or claim 83, wherein the response comprises inducing, increasing, promoting or stimulating anti- Streptococcus pneumoniae activity of T cells.

85. The method of claim 83 or claim 84, wherein the T cells are CD8+ or CD4+ T cells.

86. The method of any one of claims 83 to 85, wherein the method comprises determining whether the subject has been infected by or exposed to Streptococcus pneumoniae more than once by determining if the subject elicits a secondary T cell immune response profde that is different from a primary T cell immune response profde.

87. The method of any one of claims 82 to 86, further comprising diagnosing a latent or active Streptococcus pneumoniae infection or exposure in a subject, the method comprising contacting a biological sample from a subject with a composition of any one of claims 19 to 46; and determining if the composition elicits a T cell immune response, wherein the T cell immune response identifies that the subject has been infected with or exposed to Streptococcus pneumoniae.

88. The method of any one of claims 82 to 87, wherein the method is conducted three or more days following the date of suspected infection by or exposure to a Streptococcus pneumoniae.

89. A kit for the detection of Streptococcus pneumoniae or an immune response to Streptococcus pneumoniae in a subject comprising, consisting of or consisting essentially of:one or more T cells that specifically detect the presence of:one or more amino acid sequences selected from any one of those selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof;a fusion protein comprising one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10; ora pool of 2 or more or more peptides selected from the amino acid sequences selected from at least one of SEQ ID NOS: 1-10.

90. The kit of claim 89, wherein the one or more amino acid sequences are selected from a Streptococcus pneumoniae T cell epitope selected from at least one of SEQ ID NOS: 1-10.

91. The kit of claim 89 or claim 90, wherein the composition comprises:one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof;a fusion protein comprising one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10;a pool of 2 or more peptides selected from the amino acid sequences selected from at least one of SEQ ID NOS: 1-10; orthe peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both.

92. The kit of any one of claims 89 to 91, wherein the amino acid sequence comprises a Streptococcus pneumoniae CD8+ or CD4+ T cell epitope.

93. The kit of claim 92, wherein the T cell epitope is not conserved in another Streptococcus pneumoniae.

94. The kit of claim 92 or claim 93, wherein the T cell epitope is conserved in another Streptococcus pneumoniae.

95. The kit of any one of claims 89 to 94, wherein the fusion protein has a length from about 9-15, 15-20, 20-25, 25-30, 30-40, 40-50, 50-75 or 75-100 amino acids.

96. The kit of any one of claims 89 to 95, wherein the kit includes instruction for a diagnostic method, a process, a composition, a product, a service or component part thereof for the detection of: (i) Streptococcus pneumoniae or (ii) an immune response relevant to Streptococcus pneumoniae infections, vaccines or therapies, including T cells responsive to Streptococcus pneumoniae.

97. The kit of any one of claims 89 to 96, wherein the kit includes reagents for detecting an amount or a relative amount of, and / or the activity of, and / or the state of antigen-specific T-cells in the biological sample comprises measuring one or more of a cytokine or lymphokine secretion assay, T cell proliferation, immunoprecipitation, immunoassay, ELISA, radioimmunoassay, immunofluorescence assay, Western Blot, FACS analysis, a competitive immunoassay, a noncompetitive immunoassay, a homogeneous immunoassay a heterogeneous immunoassay, a bioassay, a reporter assay, a luciferase assay, a microarray, a surface plasmon resonance detector, a florescence resonance energy transfer, immunocytochemistry, or a cell mediated assay, or a cytokine proliferation assay.

98. The kit of any one of claims 89 to 87, wherein the kit includes reagents for determining a HumanLeukocyte Antigen (HLA) profile of a subject, and selecting peptides that are presented by the HLA profile of the subject for detecting an immune response to Streptococcus pneumoniae.

99. A kit for the detection of Streptococcus pneumoniae or an immune response to Streptococcus pneumoniae in a subject comprising, consisting of or consisting essentially of:one or more T cells that specifically detect the presence of:one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof;a fusion protein comprising one or more amino acid sequences selected from at least one of SEQ ID NOS: 1-10;a pool of 2 or more peptides selected from the amino acid sequences selected from at least one of SEQ ID NOS: 1-10; orthe peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both.

100. The kit of claim 99, wherein the one or more amino acid sequences is selected from a Streptococcus pneumoniae CD4 T cell epitope selected from at least one of SEQ ID NOS: 1-10; or both.

101. The kit of claims 99 or 100, wherein the amino acid sequence comprises a Streptococcus pneumoniae CD8+ or CD4+ T cell epitope.

102. The kit of claim 99, wherein the Streptococcus pneumoniae T cell epitope is not conserved in another Streptococcus pneumoniae .

103. The kit of claim 99, wherein the Streptococcus pneumoniae T cell epitope is conserved in another Streptococcus pneumoniae.

104. The kit of any one of claims 99 to 103, wherein the fusion protein has a length from about 9-15, 15-20, 20-25, 25-30, 30-40, 40-50, 50-75 or 75-100 amino acids.

105. The kit of any one of claims 99 to 104, wherein the kit includes instruction for a diagnostic method, a process, a composition, a product, a service or component part thereof for the detection of: (i) Streptococcus pneumoniae or (ii) an immune response relevant to Streptococcus pneumoniae infections, vaccines or therapies, including T cells responsive to Streptococcus pneumoniae.

106. The kit of any one of claims 99 to 104, wherein the kit includes reagents for detecting an amount or a relative amount of, and / or the activity of, and / or the state of antigen-specific T-cells in the biological sample comprises measuring one or more of a cytokine or lymphokine secretion assay, T cell proliferation, immunoprecipitation, immunoassay, ELISA, radioimmunoassay, immunofluorescence assay, Western Blot, FACS analysis, a competitive immunoassay, a noncompetitive immunoassay, a homogeneous immunoassay a heterogeneous immunoassay, a bioassay, a reporter assay, a luciferase assay, a microarray, a surface plasmon resonance detector, a florescence resonance energy transfer, immunocytochemistry, or a cell mediated assay, or a cytokine proliferation assay.

107. The kit of any one of claims 99 to 106, wherein the kit includes reagents for determining a Human Leukocyte Antigen (HLA) profile of a subject, and selecting peptides that are presented by the HLA profile of the subject for detecting an immune response to Streptococcus pneumoniae.

108. A method of stimulating, inducing, promoting, increasing, or enhancing an immune responseagainst a Streptococcus pneumoniae in a subject, comprising:administering a composition of claims 19 to 46, in an amount sufficient to stimulate, induce, promote, increase, or enhance an immune response against the Streptococcus pneumoniae in the subject.

109. The method of claim 108, wherein the immune response provides the subject with protection against a Streptococcus pneumoniae infection or pathology, or one or more physiological conditions, disorders, illnesses, diseases or symptoms caused by or associated with Streptococcus pneumoniae infection or pathology.

110. The method of claim 108 or claim 109, wherein the immune response is specific to:one or more Streptococcus pneumoniae peptides selected from the amino acid sequences selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof.

111. A method of stimulating, inducing, promoting, increasing, or enhancing an immune response against Streptococcus pneumoniae in a subject, comprising administering a composition of claims to 19 to 34, in an amount sufficient to stimulate, induce, promote, increase, or enhance an immune response against Streptococcus pneumoniae in the subject.

112. The method of claim 111, wherein the immune response provides the subject with protection against a Streptococcus pneumoniae infection or pathology, or one or more physiological conditions, disorders, illnesses, diseases or symptoms caused by or associated with Streptococcus pneumoniae infection or pathology.

113. The method of claim 111 or claim 112, wherein the immune response is specific to:one or more Streptococcus pneumoniae peptides selected from the amino acid sequences set forth in those sequences selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof; or the peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both.

114. A method of stimulating, inducing, promoting, increasing, or enhancing an immune response against Streptococcus pneumoniae in a subject, comprising:administering to a subject an amount of a protein or peptide or a polynucleotide that expresses the protein or peptide comprising, consisting of or consisting essentially of an amino acid sequence of the Streptococcus pneumoniae protein or peptide, or a variant, homologue, derivative or subsequence thereof, wherein the protein or peptide comprises at least two peptides selected from the amino acid sequences selected from at least one of SEQ ID NOS: 1-10 or a subsequence, portion, homologue, variant or derivative thereof, in an amount sufficient to prevent, stimulate, induce, promote, increase, immunize against, or enhance an immune response against Streptococcus pneumoniae in the subject.

115. The method of claim 114, wherein the immune response provides the subject with protection against Streptococcus pneumoniae infection or pathology, or one or more physiological conditions, disorders, illnesses, diseases or symptoms caused by or associated with Streptococcus pneumoniae infection or pathology.

116. A method of treating, preventing, or immunizing a subject against Streptococcus pneumoniae infection, comprising administering to a subject an amount of a protein, peptide or a polynucleotide that expresses the protein or peptide comprising, consisting of, or consisting essentially of an amino acidsequence of a Streptococcus protein or peptide, or a variant, homologue, derivative or subsequence thereof, wherein the protein or peptide comprises at least two amino acid sequences selected from at least one of SEQ ID NOS: 1-10 or a subsequence, portion, homologue, variant or derivative thereof, in an amount sufficient to treat, prevent, or immunize the subject for Streptococcus pneumoniae infection, wherein the protein or peptide comprises or consists of a Streptococcus pneumoniae T cell epitope that elicits, stimulates, induces, promotes, increases, or enhances an anti- Streptococcus pneumoniae T cell immune response.

117. The method of claim 116, wherein the one or more amino acid sequences are selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10, or a subsequence, portion, homologue, variant or derivative thereof;a fusion protein comprising one or more amino acid sequences selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10;a pool of 2 or more peptides selected from the amino acid sequences set forth in those sequences selected from at least one of SEQ ID NOS: 1-10; orthe peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both.

118. The method of claim 116, wherein the anti- Streptococcus pneumoniae T cell response is a CD8+, a CD4+ T cell response, or both.

119. The method of any of claims 116 to 118, wherein the T cell epitope is conserved across two or more clinical isolates of Streptococcus pneumoniae or two or more circulating forms of Streptococcus pneumoniae.

120. The method of claim 119, wherein the Streptococcus pneumoniae infection is an acute infection.

121. The method of any one of claims 116 to 120, wherein the subject is a mammal or a human.

122. The method of any one of claims 116 to 121, wherein the method reduces Streptococcus pneumoniae bacterial titer, increases or stimulates Streptococcus pneumoniae bacterial clearance, reduces or inhibits Streptococcus pneumoniae bacterial proliferation, reduces or inhibits increases in Streptococcus pneumoniae bacterial titer or Streptococcus pneumoniae bacterial proliferation, reduces the amount of a Streptococcus pneumoniae bacterial protein or the amount of a Streptococcus pneumoniae bacterial nucleic acid, or reduces or inhibits synthesis of a Streptococcus pneumoniae bacterial protein or a Streptococcus pneumoniae bacterial nucleic acid.

123. The method of any one of claims 116 to 122, wherein the method reduces one or more adverse physiological conditions, disorders, illness, diseases, symptoms or complications caused by or associated with Streptococcus sp. infection or pathology.

124. The method of any one of claims 116 to 123, wherein the method improves one or more adverse physiological conditions, disorders, illness, diseases, symptoms or complications caused by or associated with Streptococcus pneumoniae infection or pathology.

125. The method of claim 123 or claim 124, wherein the symptom is fever or chills, cough, shortness of breath or difficulty breathing, fatigue, muscle or body aches, headache, new loss of taste or smell, sore throat, congestion or runny nose, nausea or vomiting, or diarrhea.

126. The method of any one of claims 116 to 125, wherein the method reduces or inhibits susceptibility to Streptococcus pneumoniae infection or pathology.

127. The method of any one of claims 116 to 126, wherein the protein or peptide, or a subsequence, portion, homologue, variant or derivative thereof, is administered prior to, substantially contemporaneously with or following exposure to or infection of the subject with Streptococcus pneumoniae.

128. The method of any one of claims 116 to 127, wherein a plurality of Streptococcus pneumoniae T cell epitopes are administered prior to, substantially contemporaneously with or following exposure to or infection of the subject with Streptococcus pneumoniae.

129. The method of any one of claims 116 to 128, wherein the protein or peptide, or a subsequence, portion, homologue, variant or derivative thereof is administered within 2-72 hours, 2-48 hours, 4-24 hours, 4-18 hours, or 6-12 hours after a symptom of Streptococcus pneumoniae infection or exposure develops.

130. The method of any one of claims 116 to 129, wherein the protein or peptide, or a subsequence, portion, homologue, variant or derivative thereof is administered prior to exposure to or infection of the subject with Streptococcus pneumoniae.

131. The method of any one of claims 116 to 130, wherein the method further comprises administering a modulator of immune response prior to, substantially contemporaneously with or following the administration to the subject of an amount of a protein or peptide.

132. The method of claim 131, wherein the modulator of immune response is a modulator of the innate immune response.

133. The method of claim 131 or claim 132, wherein the modulator is IL-6, IFN-y, TGF-P, or IL-10, or an agonist or antagonist thereof.

134. A method of treating, preventing, or immunizing a subject against Streptococcus pneumoniae infection, comprising administering to a subject the composition of any one of claims 19-46 in an amount sufficient to treat, prevent, or immunize the subject for Streptococcus pneumoniae infection.

135. The method of claim 134, wherein the Streptococcus pneumoniae infection is an acute infection.

136. The method of claim 134, wherein the method reduces Streptococcus pneumoniae bacterial titer, increases or stimulates Streptococcus pneumoniae bacterial clearance, reduces or inhibits Streptococcus pneumoniae bacterial proliferation, reduces or inhibits increases in Streptococcus pneumoniae bacterial titer or Streptococcus pneumoniae bacterial proliferation, reduces the amount of a Streptococcus pneumoniae bacterial protein or the amount of a Streptococcus pneumoniae bacterial nucleic acid, or reduces or inhibits synthesis of a Streptococcus pneumoniae bacterial protein or a Streptococcus pneumoniae bacterial nucleic acid.

137. The method of any one of claims 134 to 136, wherein the method reduces one or more adverse physiological conditions, disorders, illness, diseases, symptoms or complications caused by or associated with Streptococcus pneumoniae infection or pathology.

138. The method of any one of claims 134 to 137, wherein the method improves one or more adversephysiological conditions, disorders, illness, diseases, symptoms or complications caused by or associated with Streptococcus pneumoniae infection or pathology.

139. The method of claim 137 or claim 138, wherein the symptom is fever or chills, cough, shortness of breath or difficulty breathing, fatigue, muscle or body aches, headache, new loss of taste or smell, sore throat, congestion or runny nose, nausea, vomiting, or diarrhea.

140. The method of any one of claims 134 to 139, wherein the method reduces or inhibits susceptibility to Streptococcus pneumoniae infection or pathology.

141. The method of any one of claims 134 to 140, wherein the composition is administered prior to, substantially contemporaneously with or following exposure to or infection of the subject with Streptococcus pneumoniae.

142. The method of any one of claims 134 to 141, wherein the composition is administered prior to, substantially contemporaneously with or following exposure to or infection of the subject with Streptococcus pneumoniae.

143. The method of any one of claims 134 to 142, wherein the composition is administered within 2- 72 hours, 2-48 hours, 4-24 hours, 4-18 hours, or 6-12 hours after a symptom of Streptococcus pneumoniae infection or exposure develops.

144. The method of any one of claims 134 to 143, wherein the composition is administered prior to exposure to or infection of the subject with Streptococcus pneumoniae.

145. A peptide or peptides that are immunoprevalent or immunodominant in a Streptococcus sp. bacteria obtained by a method consisting of, or consisting essentially of:obtaining an amino acid sequence of the bacteria;determining one or more sets of overlapping peptides spanning one or more bacteria antigen using unbiased selection;synthesizing one or more pools of bacterial peptides comprising the one or more sets of overlapping peptides;combining the one or more pools of bacteria peptides with Class I major histocompatibility proteins (MHC), Class II MHC, or both Class I and Class II MHC to form peptide-MHC complexes; contacting the peptide-MHC complexes with T cells from subjects exposed to the bacteria; determining which pools triggered cytokine release by the T cells; anddeconvoluting from the pool of peptides that elicited cytokine release by the T cells, which peptide or peptides are immunoprevalent or immunodominant in the pool.

146. The peptide or peptides of claim 145, wherein the bacteria is a Streptococcus sp.

147. The peptide or peptides of claim 146, wherein the Streptococcus sp. is Streptococcus pneumoniae.

148. The peptide or peptides of any one of claims 145 to 147, wherein the immunodominant peptides are selected from 1, 2 or more peptides selected from the amino acid sequences selected from at least one of SEQ ID NOS: 1-10.

149. The peptide or peptides of any one of claims 145 to 148, wherein the immunodominant peptidesare selected from 1, 2 or more peptides selected from the amino acid sequences set forth in those sequences selected from at least one of SEQ ID NOS: 1-10.

150. A method of selecting an immunoprevalent or immunodominant peptide or protein of a bacteria comprising, consisting of, or consisting essentially of:obtaining an amino acid sequence of the bacteria;determining one or more sets of overlapping peptides spanning one or more bacteria antigen using unbiased selection as set forth in claim 1 ;synthesizing one or more pools of bacteria peptides comprising the one or more sets of overlapping peptides;combining the one or more pools of bacteria peptides with Class I major histocompatibility proteins (MHC), Class II MHC, or both Class I and Class II MHC to form peptide-MHC complexes; contacting the peptide-MHC complexes with T cells from subjects exposed to the bacteria; determining which pools triggered cytokine release by the T cells; anddeconvoluting from the pool of peptides that elicited cytokine release by the T cells, which peptide or peptides are immunoprevalent or immunodominant in the pool.

151. The method of claim 150, wherein the bacteria is a Streptococcus sp.

152. The method of claim 151, wherein the Streptococcus sp. is Streptococcus pneumoniae.

153. The method of any one of claim 150 to 152, wherein the immunodominant peptides are selected from 1, 2 or more peptides selected from the amino acid sequences selected from at least one of SEQ ID NOS: 1-10.

154. The method of any one of claims 150 to 153, wherein the immunodominant peptides are selected from 1, 2 or more peptides selected from the amino acid sequences set forth in those sequences selected from at least one of SEQ ID NOS: 1-10; or the peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both.

155. A polynucleotide that expresses one or more peptides or proteins, comprising, consisting of, or consisting essentially of an amino acid sequence selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10), or a subsequence, portion, homologue, variant or derivative thereof;a fusion protein comprising one or more amino acid sequences selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10;a pool of 2 or more or more peptides comprising, consisting of, or consisting essentially of amino acid sequences selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10; orthe peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both.

156. A vector that comprises the polynucleotide of claim 155.

157. The vector of claim 156, wherein the vector is a bacterial vector.

158. A host cell that comprises the vector of claim 155 or claim 156.

159. A polynucleotide that expresses:one or more peptides or proteins comprising, consisting of, or consisting essentially of an amino acid sequence selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10,or a subsequence, portion, homologue, variant or derivative thereof;a fusion protein comprising one or more amino acid sequences selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10;a pool of 2 or more peptides selected from any one of those sequences selected from at least one of SEQ ID NOS: 1-10; orthe peptide of SEQ ID NO: 1 or SEQ ID NO:6, or both.

160. A vector that comprises the polynucleotide of claim 159.

161. The vector of claim 160, wherein the vector is a bacterial vector.

162. A host cell that comprises the vector of claim 160 or claim 161.

163. A computer-implemented method for generating a library of immunogenic peptide epitope molecules, the method comprising:(a) receiving an electronic communication containing a set of possible immunogenic peptide epitopes for an epitope of interest from a pathogen;(b) using a processor to select sequence data from a template antigen-binding molecule from a set of possible template antigen binding molecules, wherein the selected template is a different pathogen than the epitope of interest and wherein the set of possible template antigen-binding molecules consists of two or more known peptides that bind MHC and are presented to T cells, wherein the selecting sequence data comprises screening the set of possible template antigen-binding molecules based on at least two of:(1) gene expression and conservation;(2) one or more peptide-level features; and(3) one or more antigen-level features;(c) using two or more machine learning algorithms to train a model and combining into an ensemble model;(d) generating sequence data of one or more immunogenic peptide epitope molecules with the ensemble model;(e) synthesizing one or more immunogenic peptide epitope molecules from the sequence data of one or more immunogenic peptide epitope molecules; and(f) screening the one or more immunogenic peptide epitope molecules synthesized for immunogenicity, wherein the molecules trigger a response by T cells that recognize the one or more immunogenic peptide epitope molecules in the context of MHC Class I or Class II molecules.

164. The method of claim 163, selecting one or more immunogenic peptide epitope molecules that trigger a CD 8 or a CD4 T cell response.

165. The method of claim 163 or claim 164, wherein the method generates a library of one or more immunogenic peptide epitope molecules that trigger an immune response in a host.

166. The method of claim 163, wherein the one or more peptide-level features are selected from Major Histocompatibility (MHC) Class II binding predictions, antigen strain conservation scores, and best match to a human proteome.

167. The method of claim 163, wherein the one or more antigen-level features are selected from RNA expression and subcellular localization prediction scores for each of: cytoplasm, cytoplasmic membrane, outer membrane, and extracellular.

168. The method of any one of claims 163 to 167, wherein the MHC Class II binding is determined using NetMHCIIpan, peptide conservation for 0-3 substitutions is determined using PEPMatch, gene expression is determined using GEO Dataset, intracellular localization is determined using PSORTb, and protein existence levels using UniProt.

169. The method of claim 163, wherein the two or more machine learning algorithms are selected from Random Forest, Gradient Boosting, and XGBoost.

170. The method of claim 163, wherein the method is computer implemented.

171. The method of claim 163, wherein the model is trained with at least one of: Mycobacterium tuberculosis or Bordatella pertussis antigenic epitopes.

172. The method of claim 163, wherein the possible immunogenic peptide epitopes are viral, bacterial, fungal, protozoan, or helminthic epitopes.

173. The method of claim 163, wherein the possible immunogenic peptide epitopes are Streptococcus pneumoniae immunogenic peptide epitopes.

174. The method of claim 163, wherein the immunogenicity is in a human.

175. The method of claim 163, wherein the algorithm was used to train a model using grid search hyperparameter tuning with 10-fold cross-validation.

176. A non-transitory computer-readable medium for generating a library of immunogenic peptide epitope molecules, comprising instructions stored thereon, that when executed on a processor, perform the steps of:(a) receiving an electronic communication containing a set of possible immunogenic peptide epitopes for an epitope of interest from a pathogen;(b) using a processor to select sequence data from a template antigen-binding molecule from a set of possible template antigen binding molecules, wherein the selected template is a different pathogen than the epitope of interest and wherein the set of possible template antigen-binding molecules consists of two or more known peptides that bind MHC and are presented to T cells, wherein the selecting sequence data comprises screening the set of possible template antigen-binding molecules based on at least two of:(1) gene expression and conservation;(2) one or more peptide-level features; and(3) one or more antigen-level features;(c) using two or more machine learning algorithms to train a model and combining into an ensemble model;(d) generating sequence data of one or more immunogenic peptide epitope molecules with the ensemble model;(e) synthesizing one or more immunogenic peptide epitope molecules from the sequence data of one or more immunogenic peptide epitope molecules; and(f) screening the one or more immunogenic peptide epitope molecules synthesized for immunogenicity, wherein the molecules trigger a response by T cells that recognize the one or more immunogenic peptide epitope molecules in the context of MHC Class I or Class II molecules.

177. The method of claim 176, selecting one or more immunogenic peptide epitope molecules that trigger a CD 8 or a CD4 T cell response.

178. The method of claim 176 or claim 177, wherein the method generates a library of one or more immunogenic peptide epitope molecules that trigger an immune response in a host.

179. The method of claim 176, wherein the one or more peptide-level features are selected from Major Histocompatibility (MHC) Class II binding predictions, antigen strain conservation scores, and best match to a human proteome.

180. The method of claim 176, wherein the one or more antigen-level features are selected from RNA expression and subcellular localization prediction scores for each of: cytoplasm, cytoplasmic membrane, outer membrane, and extracellular.

181. The method of any one of claims 176 to 180, wherein the MHC Class II binding is determined using NetMHCIIpan, peptide conservation for 0-3 substitutions is determined using PEPMatch, gene expression is determined using GEO Dataset, intracellular localization is determined using PSORTb, and protein existence levels using UniProt.

182. The method of claim 176, wherein the two or more machine learning algorithms are selected from Random Forest, Gradient Boosting, and XGBoost.

183. The method of claim 176, wherein the method is computer implemented.

184. The method of claim 176, wherein the model is trained with at least one of: Mycobacterium tuberculosis or Bordatella pertussis antigenic epitopes.

185. The method of claim 176, wherein the possible immunogenic peptide epitopes are viral, bacterial, fungal, protozoan, or helminthic epitopes.

186. The method of claim 176, wherein the possible immunogenic peptide epitopes are Streptococcus pneumoniae immunogenic peptide epitopes.

187. The method of claim 176, wherein the immunogenicity is in a human.

188. The method of claim 176, wherein the algorithm was used to train a model using grid search hyperparameter tuning with 10-fold cross-validation.