T cell targets for tuberculosis vaccines

The Rv2140c vaccine targets follicular helper T cells to induce a protective immune response against tuberculosis, addressing the limitations of current vaccines and diagnostics by enhancing resistance and improving diagnostic accuracy.

WO2025250534A1PCT designated stage Publication Date: 2025-12-04THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/031027
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-28
Filing Date
2025-05-27
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Current tuberculosis vaccines, such as BCG, do not provide effective protection against latent tuberculosis infection or reactivation of pulmonary disease, and existing diagnostic tests struggle to distinguish between vaccinated and infected individuals, particularly in 'resisters' who show no response to standard diagnostics.

Method used

Development of a vaccine using the Mtb protein Rv2140c, specifically formulated as a polynucleotide or polypeptide, to induce a protective immune response by targeting follicular helper T cells, which secrete TGFβ, and the use of peptide-MHC multimers to identify and characterize antigen-specific CD4+ T cells.

Benefits of technology

The Rv2140c vaccine induces a protective immune response in individuals, enhancing their ability to resist tuberculosis infection and allows for improved diagnostic methods to identify effective immune responses in 'resisters' and vaccinated individuals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000070_0000
    Figure 00000070_0000
  • Figure 00000072_0000
    Figure 00000072_0000
  • Figure 00000073_0000
    Figure 00000073_0000
Patent Text Reader

Abstract

The Mtb protein Rv2140c, specific peptide epitopes derived, and polynucleotide sequences encoding such polypeptides are disclosed herein as a useful TB antigen for generating a protective immune response to development of tuberculosis. In some embodiments the Rv2140c polypeptide or polynucleotide is formulated as a vaccine. In some embodiments the Rv2140c vaccine is administered to a naïve recipient. In some embodiments the Rv2140c vaccine is administered as a booster to a previously vaccinated to exposed individual. In some embodiments the previous vaccination is a BCG vaccination.
Need to check novelty before this filing date? Find Prior Art

Description

T CELL TARGETS FOR TUBERCULOSIS VACCINESCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 652,623, filed May 28, 2024, which application is incorporated herein by reference in its entirety.REFERENCE TO ELECTRONIC SEQUENCE LISTING

[0002] The contents of the electronic sequence listing (STAN-2202WO.xml; Size: 39,044,638 bytes; and Date of Creation: May 9, 2025) is herein incorporated by reference in its entirety.BACKGROUND

[0003] Tuberculosis (TB) is a chronic infectious disease caused by infection with Mycobacterium tuberculosis (Mtb) and other Mycobacterium species. It is a major disease in developing countries, as well as an increasing problem in developed areas of the world. More than 2 billion people are believed to be infected with TB bacilli, with about 9.2 million new cases of TB and 1 .7 million deaths each year. 10% of those infected with TB bacilli will develop active TB, each person with active TB infecting an average of 10 to 15 others per year.

[0004] Mycobacterium tuberculosis infects individuals through the respiratory route. Alveolar macrophages engulf the bacterium, but it can survive and proliferate intracellularly. A complex immune response involving CD4+ and CD8+ T cells ensues, ultimately resulting in the formation of a granuloma. The isolated, but not eradicated, bacterium may persist for long periods, leaving an individual vulnerable to the later development of active TB. The active disease is most commonly manifested as an acute inflammation of the lungs, resulting in tiredness, weight loss, fever and a persistent cough. If untreated, serious complications and death typically result.

[0005] Diagnosis of latent TB infection conventionally uses a tuberculin skin test, which involves intradermal exposure to tuberculin protein-purified derivative (PPD). Antigen-specific T cell responses result in measurable induration at the injection site by 48-72 hours after injection, which indicates exposure to mycobacterial antigens. Sensitivity and specificity have, however, are a problem with this test, and individuals vaccinated with BCG cannot always be easily distinguished from infected individuals.

[0006] Alternative tests are based on in vitro T cell assays, measuring interferon-gamma release and in response to proteins including ESAT-6 and CFP-10, which are not delivered with BCG immunization. This known as the Quantiferon test and is the gold standard becausethese antigens are highly specific for Mtb infection, since there is much less cross-reactivity with BCG (the vaccine given to all children in endemic areas at birth).

[0007] Exposure to Mtb results in distinct clinical outcomes, including active pulmonary tuberculosis and asymptomatic latent Mtb infection (LTBI). Several recent studies have described a group of healthy, immunocompetent individuals referred to as Mtb “resisters” (RSTR). These individuals have been highly exposed to Mtb over the years but remain consistently negative for the current standard diagnostics for Mtb infection, the QuantiFERON test and the tuberculin skin test (TST). Studies of both antibody and T cell responses in RSTR individuals have demonstrated that they have been infected with Mtb and make Mtb-specific antibodies and T cell responses that differ from QuantiFERON positive (IGRA+) individuals.

[0008] Antigen-specific CD4+ T cells play a crucial role in protective immunity against Mtb infection, as seen from the devastating effects of HIV co-infection. It was largely believed that IFNy-expressing helper type 1 T cells (Th1 ) are one of the major mediators of this protection. However, the fact that RSTR individuals barely respond to the QuantiFERON test, casts doubt on this conclusion. Data show that at least some IGRA- individuals have been infected but make T cell responses that do not involve IFNy production.

[0009] Vaccination is commonly performed by injection of Bacillus Calmette- Guerin (BCG), an avirulent strain of M. bovis which was developed over 100 years ago. However, BCG does not prevent the establishment of latent TB or reactivation of pulmonary disease in adult life. Other vaccines are in development, including, for example, subunit vaccines, and advanced live mycobacterial vaccines which aim to replace BCG with more efficient and / or safer strains.

[0010] The development of effective tuberculosis (TB) vaccines is hindered by a lack of knowledge about the mechanisms of protective immunity and the antigens that induce protective responses. An effective T cell response to an infectious pathogen depends on multiple features of the T cells, and evaluation of these may provide insight into useful immunogens and improved vaccine strategies.SUMMARY

[0011] Individuals that are exposed to Mycobacterium species, e.g. Mycobacterium tuberculosis (Mtb), or Bacillus Calmette- Guerin (BCG) may, or may not, develop a protective immune response against future infection, against development of latent infection (LTBI), and against active tuberculosis. It is shown herein that individuals having a protective response, which may be referred to as resisters, can share certain features of immune responsiveness. Features of the protective response are that a significant portion of their CD4+T cells specific for Mtb are follicular helper T (Tfh) cells, which are specialized providers of T cell help to B cells, and are essential for germinal center formation, affinity maturation, and the developmentof most high affinity antibodies and memory B cells. These Tfh cells are found to secrete TGF , a multifunctional cytokine belonging to the transforming growth factor superfamily. Strikingly, an analysis of T cell receptor usage by these immune protective cells shows a clustering of TCR sequences that recognize an epitope of the Mtb protein Rv2140c.

[0012] The Mtb protein Rv2140c, specific peptide epitopes derived, and polynucleotide sequences encoding such polypeptides are disclosed herein as useful TB antigens for generating a protective immune response to development of tuberculosis. In some embodiments the Rv2140c polypeptide or polynucleotide is formulated as a vaccine. In some embodiments the Rv2140c vaccine is administered to a naive recipient. In some embodiments the Rv2140c vaccine is administered as a booster to a previously vaccinated to exposed individual. In some embodiments the previous vaccination is a BCG vaccination.

[0013] In some embodiments an isolated polynucleotide comprising a nucleic acid sequence encoding a polypeptide which comprises: (i) an Rv2140c protein sequence; (ii) a variant of an Rv2140c protein sequence; or (iii) an immunogenic fragment of an Rv2140c protein sequence. In some embodiments the isolated polynucleotide is used in making a medicament. In some embodiments the polynucleotide is formulated for administration as a vaccine. The polynucleotide may comprise all or part of the Rv2140c coding sequence. In some embodiments the polynucleotide comprises a sequence encoding at least the epitope PDPYAALPKLPSFSL (SEQ ID NO:1 ).

[0014] In some embodiments the Rv2140c polynucleotide is a modified mRNA in a vaccine formulation, e.g. formulated in a carrier. In some embodiments the carrier is a lipid nanoparticle (LNP), a polymeric nanoparticle, a lipid carrier such as a lipidoid, a liposome, a lipoplex, a peptide carrier, a nanoparticle mimic, a nanotube, or a conjugate. In some embodiments the mRNA has a 5' terminal cap that comprises a CapO, Cap1 , ARCA, inosine, N1 -methylguanosine, 2'-fluoro-guanosine, 7-deaza-guano sine, 8-oxo-guanosine, 2-amino-guanosine, LNA-guanosine, 2-azidoguanosine, Cap2, Cap4, 5' methylG cap, or an analog thereof. In some embodiments the mRNA comprises at least one chemically modified nucleobase, sugar, backbone, or any combination thereof. In embodiments the at least one chemically modified nucleobase is selected from the group consisting of pseudouracil (i ), N1 -methylpseudouracil (ml qj), 1 -ethylpseudouracil, 2-thiouracil (s2U), 4'-thiouracil, 5-methylcytosine, 5-methyluracil, 5-methoxyuracil, and any combination thereof.

[0015] In some embodiments an isolated polypeptide is provided, which comprises: (i) an Rv2140c protein sequence; (ii) a variant of an Rv2140c protein sequence; or (iii) an immunogenic fragment of an Rv2140c protein sequence, including without limitation a peptide comprising or consisting of SEQ ID NO:1 . Also provided is a polypeptide which comprises: (i) an Rv2140c protein sequence; (ii) a variant of an Rv2140c protein sequence; or (iii) an immunogenic fragment of an Rv2140c protein sequence; for use as a medicament, or for usein a method of immunizing an individual against infection by Mtb. In some embodiments the peptide is formulated for administration as a vaccine.

[0016] In an embodiment a method is provided for the treatment or prevention of TB comprising the administration of a safe and effective amount of a polypeptide, or a polynucleotide encoding a polypeptide comprising: (i) an Rv2140c protein sequence; (ii) a variant of an Rv2140c protein sequence; or (iii) an immunogenic fragment of an Rv2140c protein sequence; to a subject in need thereof, wherein said polypeptide induces an immune response, in particular a protective immune response specific for Mtb.

[0017] In another embodiment, an individual is assessed for the presence of a protective immune response following vaccination with an Rv2140c sequence, for example to determine the presence of T follicular helper cells specific for an Rv2140c. The release of TGFp from the Tfh may be determined, where such a release is indicative of a protective response. Tfh may be characterized as CD4+CXCR5+T cells, which may also express markers such as CD25, CD69, CD95, CD57, 0X40 (CD134) and CD40L (CD154) and induce over-expression of activation-induced cytidine deaminase in B cells. A Tfh response to Rv2140c, e g. SEQ ID NO:1 polypeptide, may include release of TGFp.DESCRIPTION OF THE FIGURES

[0018] FIGS. 1 A-1 E. RSTR exhibited a lower frequency of activated CD4+ T cells responding to Mtb lysate stimulation than LTBI. (A) Experimental design and workflow. In brief, PBMCs of 17 LTBI individuals and 19 RSTR individuals from the Ugandan household cohort were stimulated with Mtb lysate for 8h. Among them, 4 of the LTBI PBMCs and 3 of the RSTR PBMCs with additional vials were stimulated with a peptide pool covering Mtb ESAT6 and CFP10 proteins, as the two antigens are specific to Mtb and are utilized in the QuantiFERON test for the standard diagnosis of Mtb infection. After stimulation, the Mtb-reactive CD4+ T cells were defined and sorted as live CD3+, TCRap+, CD4+ T cells that co-express activation markers of CD69 and CD154 or CD69 and CD137 by flow cytometry as previously described. The single-cell TCR sequencing was performed to collect the a and p TCR pairs. Taking advantage of GLIPH3, the latest version of the GLIPH TCR analysis algorithm that is optimized to cluster homologous a and TCR pairs, we identified TCR specificity groups that are uniquely enriched in RSTR individuals but not in LTBI individuals. Leveraging the optimized T cell antigen discovery system with a novel reporter T cell line (J-FM84) we developed recently, we performed whole Mtb genome antigen screening. The identified peptide antigens were complexed with the corresponding MHC molecules to create the peptide-MHC multimer reagent (spheromer), which displays 12 peptide-MHC complexes. To study the CD4+ T cell responses after Mtb exposure in greater detail, we captured the antigen-specific CD4+ T cellsfrom Mtb-exposed QuantiFERON positive or negative individuals from an independent and well-studied ACS cohort in Cape Town with these T cell reagents. The unique patterns of magnitude, phenotype, cytokine expression, and TCR repertoire were characterized comprehensively by single-cell transcriptional study and flow cytometry. (B and C) Representative flow cytometry plots depicting the gating strategy to define and isolate Mtb- reactive CD4+ T cells in RSTR and LTBI individuals based on the activation markers after the stimulation of (B) Mtb lysate and (C) ESAT / CFP10 peptide pool. The cells are gated on live CD3+, TCRap+, and CD4+ T cells. (D and E) The frequency of activated CD4+ T cells among the total circulating CD4+ T cells in RSTR and LTBI individuals after the stimulation of (D) Mtb lysate and (E) ESAT / CFP10 peptide pool. Each dot represents an individual sample. The mean ± s.d. is shown. P-values were calculated with the Mann-Whitney U test (two-tailed). *** (p value<0.001 ), ” (p value<0.01 ), * (p value<0.05), ns (p value>0.05).

[0019] FIGS. 2A-2C. Unique T cell specificity groups enriched in RSTR individuals compared to LTBI individuals. (A) The number of unique CDR3P sequences detected in sorted, Mtb- lysate-reactive CD4+ T cells identified by single-cell TCR sequencing in RSTR individuals (red) and LTBI individuals (blue) from the Ugandan household cohort. Each dot represents an individual sample (RSTR, N = 19; LTBI, N = 17). The numbers were normalized by the total counts of the sorted cells. The mean ± s.d. is shown. P-values were calculated with the Mann- Whitney U test (two-tailed). *** (p value<0.001 ), “ (p value<0.01 ), * (p value<0.05), ns (p value>0.05). (B) Clonal expansion level of the CDR3p sequences detected in sorted, Mtb- lysate-reactive CD4+ T cells identified in RSTR individuals (red) and LTBI individuals (blue) from the Ugandan household cohort. The size of the dot represents the number of times the CDR3p sequences were detected, which is also depicted on the y-axis. The color gradient represents the number of CDR3P sequences that exhibit a certain level of clonal expansion. Plots have been aligned by individual samples on the x-axis (RSTR, N = 19; LTBI, N = 16). (C) Representative a and p TCR pairs of the 24 T cell specificity groups clustered by GLIPH3 that are enriched in RSTR individuals. The TCR clone ID, a and p CDR3 sequences, and the TRAV / , TRAJ, TRBV, TRBJ genes were listed. The TCR motif in the “cluster” row depicted the conserved signature of CDR3p identified by GLIPH3 in each T cell specificity group. The GLIPH3 predicted restricted HLA alleles for each T cell specificity group were listed. The specific Mtb antigens used to stimulate the PBMCs before the single-cell TCR seq were indicated as well.

[0020] FIGS. 3A-3H. The identification of novel Mtb antigens associated with the “resist” infection of RSTR individuals. (A) The schematic overview of the optimized T cell antigen discovery system. In brief, the artificial APC was built on the K562 cell line with the stable expression of HLA-DM, CD80 molecules, and a candidate HLA allele as previously described. The exogenous Mtb antigen was endocytosed, processed, and presented by the aAPC. Anovel TCR-deficient reporter T cell line (J-FM84) that achieved optimized sensitivity was developed and applied. The constant region of the a and p TCR chains was knocked out by CRISPR / Cas9 from the wild type Jurkat T cell line (clone E6-1 , ATCC). The T cell co-receptors CD8 / CD4 molecules and the co-stimulatory CD28 molecule were introduced to the TCR- deficient Jurkat line to facilitate TCR recognition. The luciferase expression will be triggered under the NFAT response element to report the recognition between TCR and peptide-MHC complex. (B) TCR-1 and TCR-10 were tested against their restricted HLA allele, HLA- DRB1 *1 1 :01 , by pulsing the aAPC with Mtb lysate, ESAT6 / CFP10 peptide pool, or DMSO. The aAPC with mismatched HLA-DRB3*02:02 allele was tested as isotype control. (C) TCR- 1 antigen screening of the whole Mtb proteome (321 subpools displayed in 4 plates). The color scale indicates the luminescence readout after stimulation. N=3 biological replicates. (D) TCR- 1 antigen screening of each individual ORF from the positive subpool (PL18C). (E) TCR-1 antigen screening of each overlapping peptide spanning Rv2140c protein (PL18C-12). (F) Dose-dependent response of TCR-1 to the identified peptide antigen Rv2140c (5-19). (G) TCR-10 antigen screening of each overlapping peptide spanning ESAT6 and CFP10 proteins. (H) Dose-dependent response of TCR-10 to the identified peptide antigen ESAT6 (25-39). In the bar and line charts, the mean ± s.d. is shown. N=3 biological replicates. P-values were calculated with the Mann-Whitney U test (two-tailed). *** (p value<0.001 ), “ (p value<0.01 ), * (p value<0.05), ns (p value>0.05).

[0021] FIGS. 4A-4F. Distinct magnitude, phenotype, and cytokine expression of antigenspecific CD4+ T cells in IGRA- and IGRA+ subjects from the ACS cohort. (A) Representative flow cytometry plots depicting the PBMC staining of HLA-DRB1 *1 1 :1 restricted IGRA- and IGRA+ ACS subjects and healthy subjects using the Rv2140c519 / DRI 1 and ESAT62539 / DRI 1 spheromer reagents. The plots were gated on the total CD4+ T cells, and the labeled percentage indicated the frequency of the spheromer positive CD4+ T cells among total CD4+ T cells. (B) The magnitude of CD4+T cell responses to Rv2140cs 19 / DR11 and ESAT625 39 / DR1 1 in HLA-DRB1 *11 :1 restricted IGRA- ACS subjects (n=5), IGRA+ ACS subjects (n=5), and healthy subjects ( n=5) . (C and D) Representative flow cytometry plots depicting the gating strategy to identify the CD4+ T cell subsets of (C) Rv2140cs 19 / DRI 1 -specific CD4+ T cells (in red) in IGRA- ACS subjects and (D) ESAT62539 / DR11 specific CD4+ T cells (in blue) in IGRA+ ACS subjects. The total CD4+ T cells were shown in grey and merged with the antigen-specific cells. (E) The fraction of CD4+ T cell subsets of Rv2140c5-ig / DR11 -specific CD4+ T cells in IGRA- ACS subjects (red bar) and ESAT62539 / DR1 1 specific CD4+ T cells in IGRA+ ACS subjects (blue bar). (F) The cytokine expression of Rv2140c519 / DRI 1 -specific CD4+ T cells in IGRA- ACS subjects (red curve) and ESAT62539 / DR1 1 specific CD4+ T cells in IGRA+ ACS subjects (blue curve). The curves from 5 IGRA- and 5 IGRA+ ACS subjects were merged and displayed together on the plots. PBMCs were stimulated with Mtb lysate for 8h before cytokinestaining as previously described. The staining of the unstimulated PBMCs was shown in grey and merged with the stimulated PBMCs. In the bar charts, the mean ± s.d. is shown. P-values were calculated with the Mann-Whitney U test (two-tailed). *** (p value<0.001 ), “ (p value<0.01 ), * (p value<0.05), ns (p value>0.05).

[0022] FIGS. 5A-5J. Transcriptional features of Rv2140c (519 DRI I -specific CD4+ T cells in IGRA- ACS individuals. (A) Memory subsets distribution of the Mtb lysate stimulated bulk CD4+ T cells in IGRA- ACS subjects (N=10). (B) Expression distribution of the markers used to define memory subsets of the Mtb lysate stimulated bulk CD4+ T cells in IGRA- ACS subjects (N=10). (C) CD4+ T cell subsets of the Mtb lysate stimulated bulk CD4+ T cells in IGRA- ACS subjects (N=10). (D) Expression level of transcriptional factors, surface markers, and secreted cytokines in each subset of the Mtb lysate stimulated bulk CD4+ T cells in IGRA- ACS subjects (N=10). The color scale indicates the expression level of each gene. (E) Distribution of Rv2140cs i9 / DR1 1 -specific CD4+ T cells among bulk CD4+ T cells in IGRA- ACS subjects (N=10). (F) Fraction of memory subsets of Rv2140cs 19 / DR11 -specific CD4+ T cells in IGRA- ACS subjects (red bar, N=10) and ESAT61539 / DRH -specific CD4+ T cells in IGRA+ ACS subjects (blue bar, N=5). (G) Fraction of CD4+ T cell subsets of Rv2140cs 19 / DR1 1 -specific CD4+ T cells in IGRA- ACS subjects (red bar, N=10) and ESAT61539 / DR1 1 - specific CD4+ T cells in IGRA+ ACS subjects (blue bar, N=5). (H) Expression level of transcriptional factors, surface markers, and secreted cytokines in the T follicular helper subset of Rv2140c5 i9 / DR1 1 -specific CD4+ T cells in IGRA- ACS subjects (N=10). The color scale indicates the expression level of each gene. (I) Cytokine expression level of ESAT61539 / DR1 1 - specific CD4+ T cells in IGRA+ ACS subjects (upper panel, N=5) and Rv2140c5i9 / DR1 1 - specific CD4+ T cells in IGRA- ACS subjects (lower panel, N=10). The color scale indicates the expression level of each cytokine, and the dot size represents the fraction of antigenspecific CD4+ T cells that expressed the cytokine. (J) TGF|31 expression level in each CD4+ T cell subset of Rv2140cs 19 / DR1 1 -specific CD4+ T cells in IGRA- ACS subjects (red violin plots on left, N=10) and ESAT61539 / DR11 -specific CD4+ T cells in IGRA+ ACS subjects (blue violin plots on right, N=5). P-values were calculated with the Mann-Whitney U test (two-tailed). *** (p value<0.001 ), ** (p value<0.01 ), * (p value<0.05), ns (p value>0.05).

[0023] FIGS. 6A-6D. BCG vaccine-induced Rv2140c targeting, Tfh-like CD4+ T cell response. (A) Representative flow cytometry plots depicting the staining of the BCG-vaccinated HLA- DRB1 *1 1 :1 restricted infant PBMCs using Rv2140cs-ig / DR11 and ESAT62539 / DR11 spheromer reagents. The plots were gated on the total CD4+ T cells, and the labeled percentage indicated the frequency of the spheromer positive CD4+ T cells among total CD4+ T cells. The plot from samples positive to Rv2140c5-ig / DR1 1 was shown on the left, and the right was the plot from samples negative to Rv2140cs 19 / DR11 . (B) The magnitude of CD4+T cell responses to Rv2140c5ig / DR11 and ESAT62539 / DRH in BCG-vaccinated HLA-DRB1 *1 1 :1 restricted infants (n=10) and unvaccinated healthy subjects (n=5). (C) Representative flow cytometry plots depicting the gating strategy to identify the CD4+ T cell subsets of Rv2140c5 ig / DR1 1 -specific CD4+ T cells (in green) in BCG-vaccinated infants. The total CD4+ T cells were shown in grey and merged with the antigen-specific cells. (D) The fraction of CD4+ T cell subsets of Rv2140c519 / DRI 1 -specific CD4+ T cells in BCG-vaccinated infants (green bar) compared with the Rv2140c5-ig / DR11 -specific CD4+ T cells in IGRA- ACS subjects (red bar). In the bar charts, the mean ± s.d. is shown. P-values were calculated with the Mann-Whitney U test (two-tailed). *** (p value<0.001 ), ** (p value<0.01 ), * (p value<0.05), ns (p value>0.05).

[0024] FIG. 7A-FIG.7E. Memory and CD4+ T cell subsets distribution of ESAT61539 / DR1 1 - specific CD4+ T cells in IGRA+ ACS subjects. (A) Memory subsets distribution of the Mtb lysate stimulated bulk CD4+ T cells in IGRA+ ACS subjects (N=5). (B) Expression distribution of the markers used to define memory subsets of the Mtb lysate stimulated bulk CD4+ T cells in IGRA+ ACS subjects (N=5) . (C) CD4+ T cell subsets of the Mtb lysate stimulated bulk CD4+ T cells in IGRA+ ACS subjects (N=5). (D) Expression level of transcriptional factors, surface markers, and secreted cytokines in each subset of the Mtb lysate stimulated bulk CD4+ T cells in IGRA+ ACS subjects (N=5). The color scale indicates the expression level of each gene. (E) Distribution of ESAT61539 / DR11 -specific CD4+ T cells among bulk CD4+ T cells in IGRA+ ACS subjects (N=5).

[0025] FIG. 8. Workflow to perform Mtb genome-wide T cell antigen screening. Two strategies were applied to make a thorough screening (a) the candidate TCRs and HLAs were first screened by the Mtb lysate to cover all potential Mtb antigens in the genome, followed by the protein pools of each Mtb ORF. Subsequently, the downstream screening narrowed the antigen with the single protein, peptide pool covering the positive protein, and the single peptide. In this way, we identified a antigen peptide (PDPYAALPKLPSFSL) within the Rv2140c protein that is presented by HLA-DRB1 *1 1 :01 and recognized by TCR-1 ; (b) the candidate TCRs and HLAs were further screened by ESAT6 / CFP10 peptide pool specifically. Another antigen SEQ ID NO:5, (IHSLLDEGKQSLTKL) within the ESAT6 protein was captured presented by HLA-DRB1 *1 1 :01 and recognized by TCR-10.

[0026] FIG. 9. Experimental design to study the phenotype and function of CD4+ T cells recognized Rv2140c / DR11 and ESAT6 / DR1 1. The PBMC samples were collected from another independent cohort (Adolescent Cohort Study, ACS) from South Africa (Zak et al., 2016, The Lancet). The PBMC samples from QFN+ donors (N=5) and QFN- donors (N=5) were stained with Mtb spheromer with Rv2140c / DR1 1 , ESAT6 / DR1 1 and other published ESAT6 epitopes with HLA-DRB1 *04:01 and HLA-A*02:01 to validate the Mtb infection history. To determine the phenotype and cytokine expression of the Rv2140c / DR11 -specific CD4+ T cells, a panel of surface markers and cytokines were tested with the FACS cytometry assay.

[0027] FIG. 10. The magnitude of CD4+ T cell response to Rv2140c / DR11 in HLA- DRB1 *1 1 :01 QFN- ACS donors, QFN+ ACS donors, and health control donors. Representative FACS plots are listed on right.

[0028] FIG. 1 1 A-FIG. 1 1 C. 1 1 C. The magnitude of CD4+ T cell response to (a) ESAT6 / DR11 in HLA-DRB1 *11 :01 QFN- ACS donors, QFN+ ACS donors, and health control donors. The other published ESAT6 epitopes (b) ESAT6 / DR4 and (c) ESAT6 / A2 were also tested to measure the Mtb expose history according to the HLA alleles carried by the QFN- ACS donors and QFN+ ACS donors. Representative FACS plots are listed on right.

[0029] FIG. 12. The CD4+ T cell subsets distribution of Rv2140c / DR1 1 CD4+ T cells (in red) in HLA-DRB1 *1 1 :01 QFN- ACS donors.

[0030] FIG. 13. The cytokine expression of Rv2140c / DR11 CD4+ T cells in HLA-DRB1 *1 1 :01 QFN- ACS donors after the stimulation with Mtb lysate.

[0031] FIG. 14. The TGFb expression level of the Tfh, Th9 and Th1 subsets within the Rv2140 / DR11 -specific CD4+ T cells.

[0032] FIG. 15. The cytokine expression of CD4+ T cells in HLA-DRB1 *1 1 :01 negative QFN- ACS donors after the stimulation with Rv2140c peptide pool.DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] Before the present methods and compositions are described, it is to be understood that this invention is not limited to particular method or composition described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.

[0034] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limits of that range is also specifically disclosed. Each smaller range between any stated value or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included or excluded in the range, and each range where either, neither or both limits are included in the smaller ranges is also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.

[0035] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, some potential and preferredmethods and materials are now described. All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. It is understood that the present disclosure supersedes any disclosure of an incorporated publication to the extent there is a contradiction.

[0036] It must be noted that as used herein and in the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a cell" includes a plurality of such cells and reference to "the peptide" includes reference to one or more peptides and equivalents thereof, e.g. polypeptides, known to those skilled in the art, and so forth.

[0037] The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates which may need to be independently confirmed.

[0038] As used herein, compounds which are "commercially available" may be obtained from commercial sources including but not limited to Acros Organics (Pittsburgh PA), Aldrich Chemical (Milwaukee Wl, including Sigma Chemical and Fluka), Apin Chemicals Ltd. (Milton Park UK), Avocado Research (Lancashire U.K.), BDH Inc. (Toronto, Canada), Bionet (Cornwall, U.K ), Chemservice Inc. (West Chester PA), Crescent Chemical Co. (Hauppauge NY), Eastman Organic Chemicals, Eastman Kodak Company (Rochester NY), Fisher Scientific Co. (Pittsburgh PA), Fisons Chemicals (Leicestershire UK), Frontier Scientific (Logan UT), ICN Biomedicals, Inc. (Costa Mesa CA), Key Organics (Cornwall U.K.), Lancaster Synthesis (Windham NH), Maybridge Chemical Co. Ltd. (Cornwall U.K.), Parish Chemical Co. (Orem UT), Pfaltz & Bauer, Inc. (Waterbury CN), Polyorganix (Houston TX), Pierce Chemical Co. (Rockford IL), Riedel de Haen AG (Hannover, Germany), Spectrum Quality Product, Inc. (New Brunswick, NJ), TCI America (Portland OR), Trans World Chemicals, Inc. (Rockville MD), Wako Chemicals USA, Inc. (Richmond VA), Novabiochem and Argonaut Technology.

[0039] Compounds can also be made by methods known to one of ordinary skill in the art. As used herein, "methods known to one of ordinary skill in the art" may be identified through various reference books and databases. Suitable reference books and treatises that detail the synthesis of reactants useful in the preparation of compounds of the present invention, or provide references to articles that describe the preparation, include for example, "Synthetic Organic Chemistry", John Wiley & Sons, Inc., New York; S. R. Sandler et al., "Organic Functional Group Preparations," 2nd Ed., Academic Press, New York, 1983; H. O. House, "Modern Synthetic Reactions", 2nd Ed., W. A. Benjamin, Inc. Menlo Park, Calif. 1972; T. L. Gilchrist, “Heterocyclic Chemistry”, 2nd Ed., John Wiley & Sons, New York, 1992; J. March, “Advanced Organic Chemistry: Reactions, Mechanisms and Structure”, 4th Ed.,Wiley-lnterscience, New York, 1992. Specific and analogous reactants may also be identified through the indices of known chemicals prepared by the Chemical Abstract Service of the American Chemical Society, which are available in most public and university libraries, as well as through on-line databases (the American Chemical Society, Washington, D.C., may be contacted for more details). Chemicals that are known but not commercially available in catalogs may be prepared by custom chemical synthesis houses, where many of the standard chemical supply houses (e.g., those listed above) provide custom synthesis services.

[0040] The term "Mycobacterium species of the tuberculosis complex" includes those species traditionally considered as causing the disease tuberculosis, as well as Mycobacterium environmental and opportunistic species that cause tuberculosis and lung disease in immune compromised patients, such as patients with AIDS, e.g., M. tuberculosis, M. bovis, or M. africanum, BCG, M. avium, M. intracellulare, M. celatum, M. genavense, M. haemophilum, M. kansasii, M. simiae, M. vaccae, M. fortuitum, and M. scrofulaceum (see, e.g., Harrison's Principles of Internal Medicine, Chapter 150, pp. 953-966 (16th ed., Braunwald, et al., eds., 2005). The present invention is particularly directed to infection with M. tuberculosis.

[0041] The term "active infection" refers to an infection, e.g. with Mtb that manifests disease symptoms and / or lesions. The terms "inactive infection", "dormant infection" or "latent infection" (LTBI) refers to an infection with Mtb without manifested disease symptoms and / or lesions.

[0042] The term "primary tuberculosis" refers to clinical illness directly following Mtb infection. See, Harrison's Principles of Internal Medicine, Chapter 150, pp. 953-966 (16th ed., Braunwald, et al., eds., 2005). The terms "secondary tuberculosis" or "postprimary tuberculosis" refer to the reactivation of a dormant, inactive or latent infection, see, Harrison's Principles of Internal Medicine, Chapter 150, pp. 953-966 (16th ed., Braunwald, et ai, eds., 2005).

[0043] The term "tuberculosis reactivation" refers to the later manifestation of disease symptoms in an individual that tests positive for infection (e.g. in a tuberculin skin test, suitably in an in vitro T cell based assay) test but does not have apparent disease symptoms. The positive diagnostic test indicates that the individual is infected, however, the individual may or may not have previously manifested active disease symptoms that had been treated sufficiently to bring the tuberculosis into an inactive or latent state. It will be recognized that methods for the prevention, delay or treatment of tuberculosis reactivation can be initiated in an individual manifesting active symptoms of disease.

[0044] The term "drug resistant" tuberculosis refers to an infection, e.g. with M. tuberculosis, wherein the infecting strain is not held static or killed by one or more of so- called "front-line"chemotherapeutic agents effective in treating tuberculosis, e.g., isoniazid, rifampin, ethambutol, streptomycin and pyrazinamide. The term "multi-drug resistant" tuberculosis refers to an infection wherein the infecting strain is resistant to two or more of "front-line" chemotherapeutic agents effective in treating tuberculosis.

[0045] A "chemotherapeutic agent" refers to a pharmacological agent known and used in the art to treat tuberculosis (e.g. infection by M. tuberculosis). Exemplified pharmacological agents used to treat tuberculosis include, but are not limited to amikacin, aminosalicylic acid, capreomycin, cycloserine, ethambutol, ethionamide, isoniazid, kanamycin, pyrazinamide, rifamycins (i.e., rifampin, rifapentine and rifabutin), streptomycin, ofloxacin, ciprofloxacin, clarithromycin, azithromycin and fluoroquinolones. "First-line" or "Front-line" chemotherapeutic agents used to treat tuberculosis that is not drug resistant include isoniazid, rifampin, ethambutol, streptomycin and pyrazinamide. "Second-line" chemotherapeutic agents used to treat tuberculosis that has demonstrated drug resistance to one or more "first-line" drugs include ofloxacin, ciprofloxacin, ethionamide, aminosalicylic acid, cycloserine, amikacin, kanamycin and capreomycin. Such pharmacological agents are reviewed in Chapter 48 of Goodman and Gilman's The Pharmacological Basis of Therapeutics, Hardman and Limbird eds., 2001 .

[0046] The term “tuberculosis resister” has been used to describe an individual who has been exposed to Mycobacterium tuberculosis but does not develop an active infection or clinical disease. Despite being in environments with a high prevalence of TB and having repeated exposures to the pathogen, these individuals remain free of any signs of active or latent TB infection. They may have had significant exposure to TB, often in high-risk environments such as healthcare settings or communities with high TB rates. They do not show signs of active TB (which would manifest as symptoms like a persistent cough, fever, night sweats, and weight loss) or latent TB (which is indicated by a positive tuberculin skin test (TST) or interferon-gamma release assay (IGRA) but without symptoms or contagiousness).

[0047] Resisters can be negative when tested with a quantiferon assay for Mtb even after infection with Mtb. The QuantiFERON™ assay is an interferon-y release assay (IGRA) for evaluation of tuberculosis (TB) infections (latent or active), and is recommended by the CDC as an alternative to the tuberculin skin test (TST) in certain situations. QFT contains TB mycobacterial proteins (ESAT-6, CFP-10, and TB 7.7), which are not found in the BCG vaccine or PPD (tuberculin purified protein derivative) TST injection. Because of this highly specific composition, QFT overcomes many of the shortcomings of the TST, and it is not affected by previous BCG vaccinations or exposure to non-tuberculosis Mycobacteria, both with the added benefit of providing a laboratory-based, objective result. See, for example, Anwar A et al. Diagnostic Utility of QuantiFERON-TB Gold (QFT-G) in Active PulmonaryTuberculosis. J Glob Infect Dis. 2015 Jul-Sep;7(3):108-12, herein specifically incorporated by reference.

[0048] Mycobacterium tuberculosis Rv2140c is a function unknown conserved phosphatidylethanolamine-binding protein (PEBP), homologous to Raf kinase inhibitor protein (RKIP) in humans. It is highly conserved among Mycobacteria sp., and variants obtain from, for example, a Mycobacterium species of the tuberculosis complex, e.g., a species such as M. tuberculosis, M. bovis, or M. africanum, or a Mycobacterium species that is environmental or opportunistic and that causes opportunistic infections such as lung infections in immune compromised hosts, including BCG, M. avium, M. intracellular, M. celatum, M. genavense, M. haemophilum, M. kansasii, M. simiae, M. vaccae, M. fortuitum, and M. scrofulaceum also find use for the purposes of the disclosure.

[0049] The DNA coding region of Rv2140c has a sequence (SEQ ID NO:2): ATGACAACTTCACCCGACCCGTATGCCGCGCTGCCCAAGCTGCCGTCCTTCAGCCTGA CGTCAACCTCGATCACCGATGGGCAGCCGCTGGCTACACCCCAGGTCAGCGGGATCA TGGGTGCGGGCGGGGCGGATGCCAGTCCGCAGCTGAGGTGGTCGGGATTTCCCAGC GAGACCCGCAGCTTCGCGGTAACCGTCTACGACCCTGATGCCCCCACCCTGTCCGGG TTCTGGCACTGGGCGGTGGCCAACCTGCCTGCCAACGTCACCGAGTTGCCCGAGGGT GTCGGCGATGGCCGCGAACTGCCGGGCGGGGCACTGACATTGGTCAACGACGCCGG TATGCGCCGGTATGTGGGTGCGGCGCCGCCTCCCGGTCATGGGGTGCATCGCTACTA CGTCGCGGTACACGCGGTGAAGGTCGAAAAGCTCGACCTCCCCGAGGACGCGAGTCC TGCATATCTGGGATTCAACCTGTTCCAGCACGCGATTGCACGAGCGGTCATCTTCGGC ACCTACGAGCAGCGTTAG

[0050] The encoded Rv2140c protein has the amino acid sequence (SEQ ID NO:3): MTTSPDPYAALPKLPSFSLTSTSITDGQPLATPQVSGIMGAGGADASPQLRWSGFPSETRS FAVTVYDPDAPTLSGFWHWAVANLPANVTELPEGVGDGRELPGGALTLVNDAGMRRYVG AAPPPGHGVHRYYVAVHAVKVEKLDLPEDASPAYLGFNLFQHAIARAVIFGTYEQR, where an epitope of interest includes SEQ ID NO:1 PDPYAALPKLPSFSL.

[0051] An exemplary RNA coding sequence is provided as SEQ ID NO:4. One of skill in the art will appreciate that a large number of coding sequences can be used to produce the Rv2140c protein, and the sequence may be codon-optimized and modified as required for the purpose of interest.(SEQ ID NO:4) AUGACAACUUCACCCGACCCGUAUGCCGCGCUGCCCAAGCUGCCGUCCUUCAGCCU GACGUCAACCUCGAUCACCGAUGGGCAGCCGCUGGCUACACCCCAGGUCAGCGGGA UCAUGGGUGCGGGCGGGGCGGAUGCCAGUCCGCAGCUGAGGUGGUCGGGAUUUCC CAGCGAGACCCGCAGCUUCGCGGUAACCGUCUACGACCCUGAUGCCCCCACCCUGU CCGGGUUCUGGCACUGGGCGGUGGCCAACCUGCCUGCCAACGUCACCGAGUUGCC CGAGGGUGUCGGCGAUGGCCGCGAACUGCCGGGCGGGGCACUGACAUUGGUCAAC GACGCCGGUAUGCGCCGGUAUGUGGGUGCGGCGCCGCCUCCCGGUCAUGGGGUGC AUCGCUACUACGUCGCGGUACACGCGGUGAAGGUCGAAAAGCUCGACCUCCCCGAGGACGCGAGUCCUGCAUAUCUGGGAUUCAACCUGUUCCAGCACGCGAUUGCACGAGCGGUCAUCUUCGGCACCUACGAGCAGCGUUAG

[0052] Immunogenic polypeptide sequences of Rv2140c, or polynucleotides encoding such polypeptides may comprise or consist of at least the epitope of SEQ ID NO:1. Polypeptides may comprise from at least about 12, at least about 15, at least about 18, at least about 21 , at least about 24, at least about 27, at least about 30, at least about 33, at least about 36, at least about 39, at least about 42, at least about 45, at least about 48, at least about 51 , at least about 54, at least about 57, at least about 60, at least about 63, at least about 66, at least about 69, at least about 72, at least about 75, at least about 78, at least about 81 , at least about 84, at least about 87, at least about 90, at least about 93, at least about 96, at least about 99, at least about 102, at least about 105, at least about 108, at least about 11 1 , at least about 1 14, at least about 1 17, at least about 120, at least about 123, at least about 126, at least about 129, at least about 131 , at least about 134, at least about 137, at least about 140, at least about 143, at least about 147, at least about 150, at least about 152, at least about 156, at least about 159, at least about 162, at least about 165, at least about 168, at least about 171 , at least about 174, and up to the complete protein of 176 amino acids.

[0053] Polynucleotides, including without limitation modified mRNA, may comprise or consist of at least at least about 36, at least about 45, at least about 54, at least about 63, at least about 72, at least about 81 , at least about 90, at least about 99, at least about 108, at least about 117, at least about 126, at least about 135, at least about 144, at least about 153, at least about 162, at least about 171 , at least about 180, at least about 189, at least about 198, at least about 207, at least about 216, at least about 225, at least about 234, at least about 243, at least about 252, at least about 261 , at least about 270, at least about 279, at least about 288, at least about 297, at least about 306, at least about 315, at least about 324, at least about 333, at least about 342, at least about 351 , at least about 360, at least about 369, at least about 378, at least about 387, at least about 396, at least about 405, at least about 414, at least about 423, at least about 432, at least about 441 , at least about 450, at least about 459, at least about 468, at least about 477, at least about 486, at least about 495, at least about 504, at least about 513, at least about 522, up to the complete coding sequence of 528 nucleotides. Start codons, termination codons and other non-coding elements may be included.

[0054] The terms "polypeptide", "peptide" and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. The terms also apply to amino acid polymers in which one or more amino acid residue is an artificial chemical mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers and nonnaturally occurring amino acid polymer. Suitably a polypeptide according to the present invention will consist only of naturally occurring amino acid residues, especially those amino acids encoded by the genetic code.

[0055] The term "amino acid" refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified, e.g., hydroxyproline, ra- carboxyglutamate, and O-phosphoserine. Amino acid analogues refers to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., an a carbon that is bound to a hydrogen, a carboxyl group, an amino group, and an R group, e.g., homoserine, norleucine, methionine sulfoxide, methionine methyl sulfonium. Such analogs have modified R groups (e.g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid. Amino acid mimetics refers to chemical compounds that have a structure that is different from the general chemical structure of an amino acid, but that functions in a manner similar to a naturally occurring amino acid. Suitably an amino acid is a naturally occurring amino acid or an amino acid analogue, especially a naturally occurring amino acid and in particular those amino acids encoded by the genetic code.

[0056] "Nucleic acid" refers to deoxyribonucleotides or ribonucleotides and polymers thereof in either single- or double-stranded form. The term encompasses nucleic acids containing known nucleotide analogs or modified backbone residues or linkages, which are synthetic, naturally occurring, and non-naturally occurring, which have similar binding properties as the reference nucleic acid, and which are metabolized in a manner similar to the reference nucleotides. Examples of such analogs include, without limitation, phosphorothioates, phosphoramidates, methyl phosphonates, chiral-methyl phosphonates, 2-O-methyl ribonucleotides, peptide- nucleic acids (PNAs). Suitably the term "nucleic acid" refers to naturally occurring deoxyribonucleotides or ribonucleotides and polymers thereof.

[0057] Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences, as well as the sequence explicitly indicated (suitably it refers to the sequence explicitly indicated). Specifically, degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues (Batzer et al, Nucleic Acid Res. 19:5081 (1991 ); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); Rossolini et al., Mol. Cell. Probes 8:91 -98 (1994)). The term nucleic acid is used interchangeably with gene, cDNA, mRNA, oligonucleotide, and polynucleotide.

[0058] Amino acids may be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides, likewise, may be referred to by their commonly accepted single- letter codes.

[0059] "Variants" or "conservatively modified variants" applies to both amino acid and nucleic acid sequences. With respect to particular nucleic acid sequences, conservatively modified variants refer to those nucleic acids which encode identical or essentially identical amino acid sequences, or where the nucleic acid does not encode an amino acid sequence, to essentially identical sequences.

[0060] Due to the degeneracy of the genetic code, a large number of functionally identical nucleic acids encode any given protein. For instance, the codons GCA, GCC, GCG and GCU all encode the amino acid alanine. Thus, at every position where an alanine is specified by a codon, the codon can be altered to any of the corresponding codons described without altering the encoded polypeptide. Such nucleic acid variations lead to "silent" or "degenerate" variants, which are one species of conservatively modified variations. Every nucleic acid sequence herein which encodes a polypeptide also describes every possible silent variation of the nucleic acid. One of skill will recognize that each codon in a nucleic acid (except AUG, which is ordinarily the only codon for methionine, and TGG, which is ordinarily the only codon for tryptophan) can be modified to yield a functionally identical molecule. Accordingly, each silent variation of a nucleic acid that encodes a polypeptide is implicit in each described sequence.

[0061] A polynucleotide of the invention may contain a number of silent variations (for example, 1 -50, such as 1-25, in particular 1 -5, and especially 1 codon(s) may be altered) when compared to the reference sequence. A polynucleotide of the invention may contain a number of non-silent conservative variations (for example, 1 -50, such as 1 -25, in particular 1 -5, and especially 1 codon(s) may be altered) when compared to the reference sequence. Non-silent variations are those which result in a change in the encoded amino acid sequence (either though the substitution, deletion or addition of amino acid residues). Those skilled in the art will recognize that a particular polynucleotide sequence may contain both silent and non-silent conservative variations.

[0062] In respect of variants of a protein sequence, the skilled person will recognize that individual substitutions, deletions or additions to polypeptide, which alters, adds or deletes a single amino acid or a small percentage of amino acids is a "conservatively modified variant" where the alteration(s) results in the substitution of an amino acid with a functionally similar amino acid or the substitution / deletion / addition of residues which do not substantially impact the biological function of the variant.

[0063] Conservative substitution tables providing functionally similar amino acids are well known in the art. Such conservatively modified variants are in addition to and do not excludepolymorphic variants, interspecies homologs, and alleles of the invention. A polypeptide of the invention may contain a number of conservative substitutions (for example, 1 -50, such as 1 - 25, in particular 1 -10, and especially 1 amino acid residue(s) may be altered) when compared to the reference sequence. Suitably such substitutions do not occur in the region of an epitope, and do not therefore have a significant impact on the immunogenic properties of the antigen. For example a polypeptide may have 1 , 2, 3 or more amino acid substitutions relative to the provided reference sequences.

[0064] Protein variants may also include those wherein additional amino acids are inserted compared to the reference sequence, for example, such insertions may occur at 1 -10 locations (such as 1 -5 locations, suitably 1 or 2 locations, in particular 1 location) and may, for example, involve the addition of 50 or fewer amino acids at each location (such as 20 or fewer, in particular 10 or fewer, especially 5 or fewer). Suitably such insertions do not occur in the region of an epitope, and do not therefore have a significant impact on the immunogenic properties of the antigen. One example of insertions includes a short stretch of histidine residues (e.g. 2-6 residues) to aid expression and / or purification of the antigen in question.

[0065] Protein variants include those wherein amino acids have been deleted compared to the reference sequence, for example, such deletions may occur at 1 -10 locations (such as 1 - 5 locations, suitably 1 or 2 locations, in particular 1 location) and may, for example, involve the deletion of 50 or fewer amino acids at each location (such as 20 or fewer, in particular 10 or fewer, especially 5 or fewer). Suitably such deletions do not occur in the region of an epitope, and do not therefore have a significant impact on the immunogenic properties of the antigen.

[0066] The skilled person will recognize that a particular protein variant may comprise substitutions, deletions and additions (or any combination thereof).

[0067] Variants preferably exhibit at least about 70% identity, more preferably at least about 80% identity and most preferably at least about 90% identity (such as at least about 95%, at least about 98% or at least about 99%) to the associated reference sequence.

[0068] The terms "identical" or percent "identity," in the context of two or more nucleic acids or polypeptide sequences, refer to two or more sequences or sub-sequences that are the same or have a specified percentage of amino acid residues or nucleotides that are the same (i.e., 70% identity, optionally 75%, 80%, 85%, 90%, 95%, 98% or 99% identity over a specified region), when compared and aligned for maximum correspondence over a comparison window, or designated region as measured using one of the following sequence comparison algorithms or by manual alignment and visual inspection. Such sequences are then said to be "substantially identical." This definition also refers to the compliment of a test sequence. Optionally, the identity exists over a region that is at least about 25 to about 50 amino acids or nucleotides in length, or optionally over a region that is 75-100 amino acids or nucleotidesin length. Suitably, the comparison is performed over a window corresponding to the entire length of the reference sequence.

[0069] "Isolated," as used herein, means that a polynucleotide is substantially away from other coding sequences, and that the polynucleotide does not contain large portions of unrelated coding DNA, such as large chromosomal fragments or other functional genes or polypeptide coding regions. An isolated nucleic acid is separated from other open reading frames that flank the gene and encode proteins other than the gene. Of course, this refers to the DNA segment as originally isolated, and does not exclude genes or coding regions later added to the segment by the hand of man.

[0070] As will be recognized by the skilled artisan, polynucleotides may be single-stranded (coding or antisense) or double-stranded, and may be DNA (genomic, cDNA or synthetic) or RNA molecules. RNA molecules include HnRNA molecules, which contain introns and correspond to a DNA molecule in a one-to-one manner, and mRNA molecules, which do not contain introns. Additional coding or non-coding sequences may, but need not, be present within a polynucleotide of the present invention, and a polynucleotide may, but need not, be linked to other molecules and / or support materials.

[0071] Polynucleotides may comprise a native sequence (i.e., an endogenous sequence that encodes a Mycobacterium antigen or a portion thereof) or may comprise a variant, or a biological or functional equivalent of such a sequence. Polynucleotide variants may contain one or more substitutions, additions, deletions and / or insertions, as further described below, preferably such that the immunogenicity of the encoded polypeptide is not diminished, relative to the reference protein. The effect on the immunogenicity of the encoded polypeptide may generally be assessed as described herein.

[0072] Moreover, it will be appreciated by those of ordinary skill in the art that, as a result of the degeneracy of the genetic code, there are many nucleotide sequences that encode a polypeptide as described herein. Some of these polynucleotides bear relatively low identity to the nucleotide sequence of any native gene. Nonetheless, polynucleotides that vary due to differences in codon usage are specifically contemplated by the present invention, for example polynucleotides that are optimized for human and / or primate codon selection. Further, alleles of the genes comprising the polynucleotide sequences provided herein are within the scope of the present invention. Alleles are endogenous genes that are altered as a result of one or more mutations, such as deletions, additions and / or substitutions of nucleotides. The resulting mRNA and protein may, but need not, have an altered structure or function. Alleles may be identified using standard techniques (such as hybridization, amplification and / or database sequence comparison).Vaccine Formulations

[0073] An Rv2140c polynucleotide or polypeptide sequence can be provided in a vaccine formulation, to provide an immunogen. "Antigen" or "immunogen" refers to any substance that stimulates an immune response. The term includes killed, inactivated, attenuated, or modified live bacteria, viruses, or parasites. The term antigen also includes polynucleotides, polypeptides, recombinant proteins, synthetic peptides, protein extract, cells (including bacterial cells), tissues, polysaccharides, or lipids, or fragments thereof, individually or in any combination thereof. The term antigen also includes antibodies, such as anti-idiotype antibodies or fragments thereof, and to synthetic peptide mimotopes that can mimic an antigen or antigenic determinant (epitope).

[0074] "Cellular immune response" or "cell mediated immune response" is one mediated by T-lymphocytes or other white blood cells or both, and includes the production of cytokines, chemokines and similar molecules produced by activated T-cells, white blood cells, or both.

[0075] "Emulsifier" means a substance used to make an emulsion more stable.

[0076] "Emulsion" means a composition of two immiscible liquids in which small droplets of one liquid are suspended in a continuous phase of the other liquid.

[0077] "Immune response" in a subject refers to the development of a humoral immune response, a cellular immune response, or a humoral and a cellular immune response to an antigen. Immune responses can usually be determined using standard immunoassays and neutralization assays, which are known in the art.

[0078] "Immunogenic" means evoking an immune or antigenic response. Thus an immunogenic composition would be any composition that induces an immune response.

[0079] "Pharmaceutically acceptable" refers to substances, which are within the scope of sound medical judgment, suitable for use in contact with the tissues of subjects without undue toxicity, irritation, allergic response, and the like, commensurate with a reasonable benefit-to- risk ratio, and effective for their intended use.

[0080] "Reactogenicity" refers to the side effects elicited in a subject in response to the administration of an adjuvant, an immunogenic, or a vaccine composition. It can occur at the site of administration, and is usually assessed in terms of the development of a number of symptoms. These symptoms can include inflammation, redness, and abscess. It is also assessed in terms of occurrence, duration, and severity. A "low" reaction would, for example, involve swelling that is only detectable by palpitation and not by the eye, or would be of short duration. A more severe reaction would be, for example, one that is visible to the eye or is of longer duration.

[0081] "Immunostimulatory composition" refers to a composition that includes an adjuvant, as defined herein and may optionally further include an antigen, in which case it may be more conventionally referred to as a vaccine. Administration of the composition to a subject results in an increased responsive state of immune cells. The amount of a composition that istherapeutically effective may vary depending on the presence of antigen, the adjuvant, and the condition of the subject, and can be determined by one skilled in the art. A non-antigenic adjuvant composition does not comprise an antigen for the disease of interest.

[0082] Vaccines known and used in the art include, for example, inactivated pathogen vaccines; live-attenuated pathogen vaccines; messenger RNA (mRNA) vaccines; subunit, recombinant, polysaccharide, and conjugate vaccines; toxoid vaccines; and viral vector vaccines. mRNA vaccines encode pathogen proteins that trigger an immune response. Subunit, recombinant, polysaccharide, and conjugate vaccines use specific pathogen molecules. Viral vector vaccines use a modified version of a different virus as a vector to deliver sequences encoding pathogen protein. Several different viruses have been used as vectors, including influenza, vesicular stomatitis virus (VSV), measles virus, and adenovirus.

[0083] In some embodiments, a vaccine formulation of the disclosure is an mRNA vaccine encoding an Rv2140c epitope. In some aspects, the invention is a vaccine composition of the mRNA encoding an antigen formulated in a carrier. In some embodiments the carrier is a lipid nanoparticle (LNP), a polymeric nanoparticle, a lipid carrier such as a lipidoid, a liposome, a lipoplex, a peptide carrier, a nanoparticle mimic, a nanotube, or a conjugate.

[0084] In some embodiments the mRNA has a 5' terminal cap that comprises a CapO, Cap1 , ARCA, inosine, N1 -methyl-guanosine, 2'-fluoro-guanosine, 7-deaza-guano sine, 8-oxo- guanosine, 2-amino-guanosine, LNA-guanosine, 2-azidoguanosine, Cap2, Cap4, 5' methylG cap, or an analog thereof.

[0085] In some embodiments the mRNA comprises at least one chemically modified nucleobase, sugar, backbone, or any combination thereof. In embodiments the at least one chemically modified nucleobase is selected from the group consisting of pseudouracil (i ), N1 - methylpseudouracil (m1 qj), 1 -ethylpseudouracil, 2-thiouracil (s2U), 4'-thiouracil, 5- methylcytosine, 5-methyluracil, 5-methoxyuracil, and any combination thereof.

[0086] In some embodiments at least about 25%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or 100% of the guanines, adenines, uracils or thymines are chemically modified. In some embodiments the mRNA is purified.

[0087] mRNA vaccines of the disclosure may include one or more antigens. In some embodiments the mRNA vaccine is composed of 3 or more, 4 or more, 5 or more 6 or more 7 or more, 8 or more, 9 or more antigens. In other embodiments the mRNA vaccine is composed of 1000 or less, 900 or less, 500 or less, 100 or less, 75 or less, 50 or less, 40 or less, 30 or less, 20 or less or 100 or less cancer antigens. In yet other embodiments the mRNA vaccine has 3-100, 5-100, 10-100, 15-100, 20-100, 25-100, 30-100, 35-100, 40-100, 45-100, 50-100, 55-100, 60-100, 65-100, 70-100, 75-100, 80-100, 90-100, 5-50, 10-50, 15-50, 20-50, 25-50,30-50, 35-50, 40-50, 45-50, 100-150, 100-200, 100-300, 100-400, 100-500, 50-500, 50-800, 50-1 ,000, or 100-1 ,000 antigens.

[0088] The present disclosure provides for modified nucleosides and nucleotides of a polynucleotide (e.g., RNA polynucleotides, such as mRNA polynucleotides) encoding an antigen polypeptide. As used herein, the term “nucleic acid” is used in its broadest sense and encompasses any compound and / or substance that includes a polymer of nucleotides, or derivatives or analogs thereof. These polymers are often referred to as “polynucleotides”. Accordingly, as used herein the terms “nucleic acid” and “polynucleotide” are equivalent and are used interchangeably. Exemplary nucleic acids or polynucleotides of the disclosure include, but are not limited to, ribonucleic acids (RNAs), deoxyribonucleic acids (DNAs), DNA- RNA hybrids, RNAi-inducing agents, RNAi agents, siRNAs, shRNAs, mRNAs, modified mRNAs, miRNAs, antisense RNAs, ribozymes, catalytic DNA, RNAs that induce triple helix formation, threose nucleic acids (TNAs), glycol nucleic acids (GNAs), peptide nucleic acids (PNAs), locked nucleic acids (LNAs, including LNA having a p-D-ribo configuration, a-LNA having an a-L-ribo configuration (a diastereomer of LNA), 2'-amino-LNA having a 2'-amino functionalization, and 2'-amino-a-LNA having a 2'-amino functionalization) or hybrids thereof. As used herein, the term “nucleobase” (alternatively “nucleotide base” or “nitrogenous base”) refers to a purine or pyrimidine heterocyclic compound found in nucleic acids, including any derivatives or analogs of the naturally occurring purines and pyrimidines that confer improved properties (e.g., binding affinity, nuclease resistance, chemical stability) to a nucleic acid or a portion or segment thereof. Adenine, cytosine, guanine, thymine, and uracil are the nucleobases predominately found in natural nucleic acids. Other natural, non-natural, and / or synthetic nucleobases, as known in the art and / or described herein, can be incorporated into nucleic acids.

[0089] As used herein, the term “nucleoside” refers to a compound containing a sugar molecule (e.g., a ribose in RNA or a deoxyribose in DNA), or derivative or analog thereof, covalently linked to a nucleobase (e.g., a purine or pyrimidine), or a derivative or analog thereof (also referred to herein as “nucleobase”), but lacking an internucleoside linking group (e.g., a phosphate group). As used herein, the term “nucleotide” refers to a nucleoside covalently bonded to an internucleoside linking group (e.g., a phosphate group), or any derivative, analog, or modification thereof that confers improved chemical and / or functional properties (e.g., binding affinity, nuclease resistance, chemical stability) to a nucleic acid or a portion or segment thereof. Modified nucleotides can by synthesized by any useful method, such as, for example, chemically, enzymatically, or recombinantly, to include one or more modified or non-natural nucleosides. Polynucleotides can comprise a region or regions of linked nucleosides. Such regions can have variable backbone linkages. The linkages can bestandard phosphodiester linkages, in which case the polynucleotides would comprise regions of nucleotides.

[0090] The modified polynucleotides disclosed herein can comprise various distinct modifications. In some embodiments, the modified polynucleotides contain one, two, or more (optionally different) nucleoside or nucleotide modifications. In some embodiments, a modified polynucleotide, introduced to a cell can exhibit one or more desirable properties, e.g., improved protein expression, reduced immunogenicity, or reduced degradation in the cell, as compared to an unmodified polynucleotide.

[0091] In some embodiments, a polynucleotide of the present invention (e.g., a polynucleotide comprising a nucleotide sequence encoding an Rv2140c polypeptide) is structurally modified, i.e. , comprises one or more nucleic acid structure modifications. As used herein, a “structural” modification is one in which two or more linked nucleosides are inserted, deleted, duplicated, inverted or randomized in a polynucleotide without significant chemical modification to the nucleotides themselves. Further, the term “nucleic acid structure” (used interchangeably with “polynucleotide structure”) refers to the arrangement or organization of atoms, chemical constituents, elements, motifs, and / or sequence of linked nucleotides, or derivatives or analogs thereof, that comprise a nucleic acid (e.g., an mRNA). The term also refers to the two- dimensional or three-dimensional state of a nucleic acid. Accordingly, the term “RNA structure” refers to the arrangement or organization of atoms, chemical constituents, elements, motifs, and / or sequence of linked nucleotides, or derivatives or analogs thereof, comprising an RNA molecule (e.g., an mRNA) and / or refers to a two-dimensional and / or three dimensional state of an RNA molecule. Nucleic acid structure can be further demarcated into four organizational categories referred to herein as “molecular structure”, “primary structure”, “secondary structure”, and “tertiary structure” based on increasing organizational complexity. Because chemical bonds will necessarily be broken and reformed to effect a structural modification, structural modifications are of a chemical nature and hence are chemical modifications. However, structural modifications will result in a different sequence of nucleotides. For example, the polynucleotide “ATCG” can be chemically modified to “AT-5meC-G”. The same polynucleotide can be structurally modified from “ATCG” to “ATCCCG”. Here, the dinucleotide “CC” has been inserted, resulting in a structural modification to the polynucleotide.

[0092] In some embodiments, the polynucleotides of the present invention can have a uniform chemical modification of all or any of the same nucleoside type or a population of modifications produced by mere downward titration of the same starting modification in all or any of the same nucleoside type, or a measured percent of a chemical modification of all any of the same nucleoside type but with random incorporation, such as where all uridines are replaced by a uridine analog, e.g., pseudouridine or 5-methoxyuridine. In another embodiment, the polynucleotides can have a uniform chemical modification of two, three, or four of the samenucleoside type throughout the entire polynucleotide (such as all uridines and all cytosines, etc. are modified in the same way).

[0093] Modified nucleotide base pairing encompasses not only the standard adenosinethymine, adenosine-uracil, or guanosine-cytosine base pairs, but also base pairs formed between nucleotides and / or modified nucleotides comprising non-standard or modified bases, wherein the arrangement of hydrogen bond donors and hydrogen bond acceptors permits hydrogen bonding between a non-standard base and a standard base or between two complementary non-standard base structures. One example of such non-standard base pairing is the base pairing between the modified nucleotide inosine and adenine, cytosine or uracil. Any combination of base / sugar or linker can be incorporated into polynucleotides of the present disclosure.

[0094] The skilled artisan will appreciate that, except where otherwise noted, polynucleotide sequences set forth in the instant application will recite “T”s in a representative DNA sequence but where the sequence represents RNA, the “T”s would be substituted for “U”s.

[0095] Modifications of polynucleotides (e.g., RNA polynucleotides, such as mRNA polynucleotides) that are useful in the compositions, methods and synthetic processes of the present disclosure include, but are not limited to the following nucleotides, nucleosides, and nucleobases: 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine; 2-methylthio-N6- methyladenosine; 2-methylthio-N6-threonyl carbamoyladenosine; N6- glycinylcarbamoyladenosine; N6-isopentenyladenosine; N6-methyladenosine; N6-threonyl carbamoyladenosine; 1 ,2'-O-dimethyladenosine; 1 -methyladenosine; 2'-O-methyladenosine; 2'-O-ribosyladenosine (phosphate); 2-methyladenosine; 2-methylthio-N6 isopentenyladenosine; 2-methylthio-N6-hydroxynorvalyl carbamoyladenosine; 2'-O- methyladenosine; 2'-0-ribosyladenosine (phosphate); Isopentenyladenosine; N6-(cis- hydroxyisopentenyl)adenosine; N6,2'-0-dimethyladenosine; N6,2'-O-dimethyladenosine; N6,N6,2'-O-trimethyladenosine; N6,N6-dimethyladenosine; N6-acetyladenosine; N6- hydroxynorvalylcarbamoyladenosine; N6-methyl-N6-threonylcarbamoyladenosine; 2- methyladenosine; 2-methylthio-N6-isopentenyladenosine; 7-deaza-adenosine; N1 -methyladenosine; N6, N6 (dimethyl)adenine; N6-cis-hydroxy-isopentenyl-adenosine; a-thio- adenosine; 2 (amino)adenine; 2 (aminopropyl)adenine; 2 (methylthio) N6 (isopentenyl)adenine; 2-(alkyl)adenine; 2-(aminoalkyl)adenine; 2-(aminopropyl)adenine; 2- (halo)adenine; 2-(halo)adenine; 2-(propyl) adenine; 2'-Amino-2'-deoxy-ATP; 2'-Azido-2'- deoxy-ATP; 2'-Deoxy-2'-a-aminoadenosine TP; 2'-Deoxy-2'-a-azidoadenosine TP; 6 (alkyl)adenine; 6 (methyl)adenine; 6-(alkyl)adenine; 6-(methyl)adenine; 7 (deaza)adenine; 8 (alkenyl)adenine; 8 (alkynyl)adenine; 8 (amino)adenine; 8 (thioalkyl)adenine; 8- (alkenyl)adenine; 8-(alkyl)adenine; 8-(alkynyl)adenine; 8-(amino)adenine; 8-(halo)adenine; 8- (hydroxyl)adenine; 8-(thioalkyl)adenine; 8-(thiol)adenine; 8-azido-adeno sine; aza adenine;deaza adenine; N6 (methyl)adenine; N6-(isopentyl)adenine; 7-deaza-8-aza-adenosine; 7- methyladenine; 1 -Deazaadenosine TP; 2'Fluoro-N6-Bz-deoxyadenosine TP; 2'-OMe-2- Amino-ATP; 2'O-methyl-N6-Bz-deoxyadenosine TP; 2’-a-Ethynyladenosine TP; 2- aminoadenine; 2-Aminoadenosine TP; 2-Amino-ATP; 2'-a-Trifluoromethyladenosine TP; 2- Azidoadenosine TP; 2'-b-Ethynyladenosine TP; 2-Bromoadenosine TP; 2'-b- Trifluoromethyladenosine TP; 2-Chloroadenosine TP; 2'-Deoxy-2',2'-difluoroadenosine TP; 2'- Deoxy-2'-a-mercaptoadenosine TP; 2'-Deoxy-2'-a-thiomethoxyadenosine TP; 2'-Deoxy-2'-b- aminoadenosine TP; 2'-Deoxy-2'-b-azidoadenosine TP; 2'-Deoxy-2'-b-bromoadenosine TP; 2'-Deoxy-2'-b-chloroadenosine TP; 2'-Deoxy-2'-b-fluoroadenosine TP; 2'-Deoxy-2'-b- iodoadenosine TP; 2'-Deoxy-2'-b-mercaptoadenosine TP; 2'-Deoxy-2'-b- thiomethoxyadenosine TP; 2-Fluoroadenosine TP; 2-lodoadenosine TP; 2- Mercaptoadenosine TP; 2-methoxy-adenine; 2-methylthio-adenine; 2- Trifluoromethyladenosine TP; 3-Deaza-3-bromoadenosine TP; 3-Deaza-3-chloroadenosine TP; 3-Deaza-3-fluoroadenosine TP; 3-Deaza-3-iodoadenosine TP; 3-Deazaadenosine TP; 4'- Azidoadenosine TP; 4'-Carbocyclic adenosine TP; 4'-Ethynyladenosine TP; 5'-Homo- adenosine TP; 8-Aza-ATP; 8-bromo-adenosine TP; 8-Trifluoromethyladenosine TP; 9- Deazaadenosine TP; 2-aminopurine; 7-deaza-2,6-diaminopurine; 7-deaza-8-aza-2,6- diaminopurine; 7-deaza-8-aza-2-aminopurine; 2,6-diaminopurine; 7-deaza-8-aza-adenine, 7- deaza-2-aminopurine; 2-thiocytidine; 3-methylcytidine; 5-formylcytidine; 5- hydroxymethylcytidine; 5-methylcytidine; N4-acetylcytidine; 2'-O-methylcytidine; 2'-O- methylcytidine; 5,2'-O-dimethylcytidine; 5-formyl-2'-O-methylcytidine; Lysidine; N4,2'-O- dimethylcytidine; N4-acetyl-2'-O-methylcytidine; N4-methylcytidine; N4,N4-Dimethyl-2'-OMe- Cytidine TP; 4-methylcytidine; 5-aza-cytidine; Pseudo-iso-cytidine; pyrrolo-cytidine; a-thio- cytidine; 2-(thio)cytosine; 2'-Amino-2'-deoxy-CTP; 2'-Azido-2'-deoxy-CTP; 2'-Deoxy-2'-a- aminocytidine TP; 2'-Deoxy-2'-a-azidocytidine TP; 3 (deaza) 5 (aza)cytosine; 3 (methyl)cytosine; 3-(alkyl)cytosine; 3-(deaza) 5 (aza)cytosine; 3-(methyl)cytidine; 4,2'-O- dimethylcytidine; 5 (halo)cytosine; 5 (methyl)cytosine; 5 (propynyl)cytosine; 5 (trifluoromethyl)cytosine; 5-(alkyl)cytosine; 5-(alkynyl)cytosine; 5-(halo)cytosine; 5- (propynyl)cytosine; 5-(trifluoromethyl)cytosine; 5-bromo-cytidine; 5-iodo-cytidine; 5-propynyl cytosine; 6-(azo)cytosine; 6-aza-cytidine; aza cytosine; deaza cytosine; N4 (acetyl)cytosine; 1 -methyl-1 -deaza-pseudoisocytidine; 1-methyl-pseudoisocytidine; 2-methoxy-5-methyl- cytidine; 2-methoxy-cytidine; 2-thio-5-methyl-cytidine; 4-methoxy-1 -methyl- pseudoisocytidine; 4-methoxy-pseudoisocytidine; 4-th io- 1 -methyl-1 -deaza- pseudoisocytidine; 4-thio-1 -methyl-pseudoisocytidine; 4-thio-pseudoisocytidine; 5-aza- zebularine; 5-methyl-zebularine; pyrrolo-pseudoisocytidine; Zebularine; (E)-5-(2-Bromo- vinyl)cytidi ne TP; 2,2'-anhydro-cytidine TP hydrochloride; 2'Fluor-N4-Bz-cytidine TP; 2'Fluoro- N4-Acetyl-cytidine TP; 2'-O-Methyl-N4-Acetyl-cytidine TP; 2'0-methyl-N4-Bz-cytidine TP; 2'-a-Ethynylcytidine TP; 2'-a-Trifluoromethylcytidine TP; 2'-b-Ethynylcytidine TP; 2'-b- Trifluoromethylcytidine TP; 2'-Deoxy-2',2'-difluorocytidine TP; 2'-Deoxy-2'-a-mercaptocytidine TP; 2'-Deoxy-2'-a-thiomethoxycytidine TP; 2'-Deoxy-2'-b-aminocytidine TP; 2'-Deoxy-2'-b- azidocytidine TP; 2'-Deoxy-2'-b-bromocytidine TP; 2'-Deoxy-2'-b-chlorocytidine TP; 2'-Deoxy- 2'-b-fluorocytidine TP; 2'-Deoxy-2'-b-iodocytidine TP; 2'-Deoxy-2'-b-mercaptocytidine TP; 2'- Deoxy-2'-b-thiomethoxycytidine TP; 2'-O-Methyl-5-(1 -propynyl)cytidine TP; 3'-Ethynylcytidine TP; 4'-Azidocytidine TP; 4’-Carbocyclic cytidine TP; 4'-Ethynylcytidine TP; 5-(1-Propynyl)ara- cytidine TP; 5-(2-Chloro-phenyl)-2-thiocytidine TP; 5-(4-Amino-phenyl)-2-thiocytidine TP; 5- Aminoallyl-CTP; 5-Cyanocytidine TP; 5-Ethynylara-cytidine TP; 5-Ethynylcytidine TP; 5'- Homo-cytidine TP; 5-Methoxycytidine TP; 5-Trifluoromethyl-Cytidine TP; N4-Amino-cytidine TP; N4-Benzoyl-cytidine TP; Pseudoisocytidine; 7-methylguanosine; N2,2'-O- dimethylguanosine; N2-methylguanosine; Wyosine; 1 ,2'-0-dimethylguanosine; 1- methylguanosine; 2’-0-methylguanosine; 2'-0-ribosylguanosine (phosphate); 2'-O- methylguanosine; 2'-O-ribosylguanosine (phosphate); 7-aminomethyl-7-deazaguanosine; 7- cyano-7-deazaguanosine; Archaeosine; Methylwyo sine; N2,7-dimethylguanosine; N2,N2,2'- O-trimethylguanosine; N2,N2,7-trimethylguanosine; N2,N2-dimethylguanosine; N2,7,2'-O- trimethylguanosine; 6-thio-guanosine; 7-deaza-guanosine; 8-oxo-guanosine; N1 -methylguanosine; a-thio-guanosine; 2 (propyl)guanine; 2-(alkyl)guanine; 2'-Amino-2'-deoxy-GTP; 2'- Azido-2'-deoxy-GTP; 2'-Deoxy-2'-a-aminoguanosine TP; 2'-Deoxy-2'-a-azidoguanosine TP; 6 (methyl)guanine; 6-(alkyl)guanine; 6-(methyl)guanine; 6-methyl-guanosine; 7 (alkyl)guanine;7 (deaza)guanine; 7 (methyl)guanine; 7-(alkyl)guanine; 7-(deaza)guanine; 7-(methyl)guanine;8 (alkyl)guanine; 8 (alkynyl)guanine; 8 (halo)guanine; 8 (thioalkyl)guanine; 8-(alkenyl)guanine; 8-(alkyl)guanine; 8-(alkynyl)guanine; 8-(amino)guanine; 8-(halo)guanine; 8- (hydroxyl)guanine; 8-(thioalkyl)guanine; 8-(thiol)guanine; aza guanine; deaza guanine; N (methyl)guanine; N-(methyl)guanine; 1-methyl-6-thio-guanosine; 6-methoxy-guanosine; 6- thio-7-deaza-8-aza-guanosine; 6-thio-7-deaza-guanosine; 6-thio-7-methyl-guanosine; 7- deaza-8-aza-guanosine; 7-methyl-8-oxo-guanosine; N2,N2-dimethyl-6-thio-guanosine; N2- methyl-6-thio-guanosine; 1 -Me-GTP; 2'Fluoro-N2-isobutyl-guanosine TP; 2'0-methyl-N2- isobutyl-guanosine TP; 2'-a-Ethynylguanosine TP; 2'-a-Trifluoromethylguanosine TP; 2'-b- Ethynylguanosine TP; 2'-b-Trifluoromethylguanosine TP; 2'-Deoxy-2',2'-difluoroguanosine TP; 2'-Deoxy-2'-a-mercaptoguanosine TP; 2'-Deoxy-2'-a-thiomethoxyguanosine TP; 2'- Deoxy-2'-b-aminoguanosine TP; 2'-Deoxy-2'-b-azidoguanosine TP; 2'-Deoxy-2'-b- bromoguanosine TP; 2'-Deoxy-2'-b-chloroguanosine TP; 2'-Deoxy-2'-b-fluoroguanosine TP; 2'-Deoxy-2'-b-iodoguanosine TP; 2'-Deoxy-2'-b-mercaptoguanosine TP; 2'-Deoxy-2'-b- thiomethoxyguanosine TP; 4'-Azidoguanosine TP; 4'-Carbocyclic guanosine TP; 4'- Ethynylguanosine TP; 5'-Homo-guanosine TP; 8-bromo-guanosine TP; 9-Deazaguanosine TP; N2-isobutyl-guanosine TP; 1-methylinosine; Inosine; 1 ,2'-O-dimethylinosine; 2'-O-methylinosine; 7-methylinosine; 2'-O-methylinosine; Epoxyqueuosine; galactosyl-queuosine; Mannosylqueuosine; Queuosine; allyamino-thymidine; aza thymidine; deaza thymidine; deoxy-thymidine; 2'-O-methyluridine; 2-thiouridine; 3-methyluridine; 5-carboxymethyluridine; 5-hydroxyuridine; 5-methyluridine; 5-taurinomethyl-2-thiouridine; 5-taurinomethyluridine; Dihydrouridine; Pseudouridine; (3-(3-amino-3-carboxypropyl)uridine; 1 -methyl-3-(3-amino-5- carboxypropyl)pseudouridine; 1 -methylpseduouridine; 1 -ethyl-pseudouridine; 2'-O- methyluridine; 2'-O-methylpseudouridine; 2'-O-methyluridine; 2-thio-2'-O-methyluridine; 3-(3- amino-3-carboxypropyl)uridine; 3,2'-O-dimethyluridine; 3-Methyl-pseudo-Uridine TP; 4- thiouridine; 5-(carboxyhydroxymethyl)uridine; 5-(carboxyhydroxymethyl)uridine methyl ester; 5,2'-O-dimethyluridine; 5,6-dihydro-uridine; 5-aminomethyl-2-thiouridine; 5-carbamoylmethyl- 2'-O-methyluridine; 5-carbamoylmethyluridine; 5-carboxyhydroxymethyluridine; 5- carboxyhydroxymethyluridine methyl ester; 5-carboxymethylaminomethyl-2'-0-methyluridine; 5-carboxymethylaminomethyl-2-thiouridine; 5-carboxymethylaminomethyl-2-thiouridine; 5- carboxymethylaminomethyluridine; 5-carboxymethylaminomethyluridine; 5- Carbamoylmethyluridine TP; 5-methoxycarbonylmethyl-2'-0-methyluridine; 5- methoxycarbonylmethyl-2-thiouridine; 5-methoxycarbonylmethyluridine; 5-methyluridine,), 5- methoxyuridine; 5-methyl-2-thiouridine; 5-methylaminomethyl-2-selenouridine; 5- methylaminomethyl-2-thiouridine; 5-methylaminomethyluridine; 5-Methyldihydrouridine; 5- Oxyacetic acid-Uridine TP; 5-Oxyacetic acid-methyl ester-Uridine TP; N1-methyl-pseudo- uracil; N1-ethyl-pseudo-uracil; uridine 5-oxyacetic acid; uridine 5-oxyacetic acid methyl ester; 3-(3-Amino-3-carboxypropyl)-Uridine TP; 5-(iso-Pentenylaminomethyl)-2-thiouridine TP; 5- (iso-Pentenylaminomethyl)-2'-O-methyluridine TP; 5-(iso-Pentenylaminomethyl)uridine TP; 5- propynyl uracil; a-thio-uridine; 1 (aminoalkylamino-carbonylethylenyl)-2(thio)-pseudouracil; 1 (aminoalkylaminocarbonylethylenyl)-2,4-(dithio)pseudouracil; 1(aminoalkylaminocarbonylethylenyl)-4 (thio)pseudouracil; 1(aminoalkylaminocarbonylethylenyl)-pseudouracil; 1 (aminocarbonylethylenyl)-2(thio)- pseudouracil; 1 (aminocarbonylethylenyl)-2,4-(dithio)pseudouracil; 1 (aminocarbonylethylenyl)-4 (thio)pseudouracil; 1 (aminocarbonylethylenyl)-pseudouracil; 1 substituted 2(thio)-pseudouracil; 1 substituted 2,4-(dithio)pseudouracil; 1 substituted 4 (thio)pseudouracil; 1 substituted pseudouracil; 1 -(aminoalkylamino-carbonylethylenyl)-2- (thio)-pseudouracil; 1 -Methyl-3-(3-amino-3-carboxypropyl) pseudouridine TP; 1 -Methyl-3-(3- amino-3-carboxypropyl)pseudo-UTP; 1-Methyl-pseudo-UTP; 1 -Ethyl-pseudo-UTP; 2 (thio)pseudouracil; 2’ deoxy uridine; 2' fluorouridine; 2-(thio)uracil; 2,4-(dithio)psuedouracil; 2' methyl, 2'amino, 2'azido, 2'fluro-guanosine; 2'-Amino-2'-deoxy-UTP; 2'-Azido-2'-deoxy-UTP; 2'-Azido-deoxyuridine TP; 2'-0-methylpseudouridine; 2' deoxy uridine; 2' fluorouridine; 2'- Deoxy-2'-a-aminouridine TP; 2'-Deoxy-2'-a-azidouridine TP; 2-methylpseudouridine; 3 (3 amino-3 carboxypropyl)uracil; 4 (thio)pseudouracil; 4-(thio) pseudouracil; 4-(thio)uracil; 4-thiouracil; 5 (1 ,3-diazole-1 -alkyl)uracil; 5 (2-aminopropyl)uracil; 5 (aminoalkyl)uracil; 5 (dimethylaminoalkyl)uracil; 5 (guanidiniumalkyl)uracil; 5 (methoxycarbonylmethyl)-2- (thio)uracil; 5 (methoxycarbonyl-methyl)uracil; 5 (methyl) 2 (thio)uracil ; 5 (methyl) 2,4 (dithio)uracil; 5 (methyl) 4 (thio)uracil ; 5 (methylaminomethyl)-2 (thio)uracil; 5 (methylaminomethyl)-2,4 (dithio)uracil; 5 (methylaminomethyl)-4 (thio)uracil; 5 (propynyl)uracil; 5 (trifluoromethyl)uracil; 5-(2-aminopropyl)uracil; 5-(alkyl)-2- (thio)pseudouracil; 5-(alkyl)-2,4 (dithio)pseudouracil; 5-(alkyl)-4 (thio)pseudouracil; 5- (alkyl)pseudouracil; 5-(alkyl)uracil; 5-(alkynyl)uracil; 5-(allylamino)uracil; 5-(cyanoalkyl)uracil; 5-(dialkylaminoalkyl)uracil; 5-(dimethylaminoalkyl)uracil; 5-(guanidiniumalkyl)uracil; 5- (halo)uracil; 5-(1 ,3-diazole-1 -alkyl)uracil; 5-(methoxy)uracil; 5-(methoxycarbonylmethyl)-2- (thio)uracil; 5-(methoxycarbonyl-methyl)uracil; 5-(methyl) 2(th io) uracil; 5-(methyl) 2,4 (dithio)uracil; 5-(methyl) 4 (thio)uracil; 5-(methyl)-2-(thio)pseudouracil; 5-(methyl)-2,4 (dithio)pseudouracil; 5-(methyl)-4 (thio)pseudouracil; 5-(methyl)pseudouracil; 5- (methylaminomethyl)-2 (thio)uracil; 5-(methylaminomethyl)-2,4(dithio)uracil; 5- (methylaminomethyl)-4-(thio)uracil; 5-(propynyl)uracil; 5-(trifluoromethyl)uracil; 5-aminoallyl- uridine; 5-bromo-uridine; 5-iodo-uridine; 5-uracil; 6 (azo)uracil; 6-(azo)uracil; 6-aza-uridine; allyamino-uracil; aza uracil; deaza uracil; N3 (methyl)uracil; P seudo-UTP-1 -2-ethanoic acid; Pseudouracil; 4-Thio-pseudo-UTP; 1 -carboxymethyl-pseudouridine; 1 -methyl-1 -deazapseudouridine; 1-propynyl-uridine; 1-taurinomethyl-1-methyl-uridine; 1-taurinomethyl-4-thio- uridine; 1-taurinomethyl-pseudouridine; 2-methoxy-4-thio-pseudouridine; 2-thio-1 -methyl-1 - deaza-pseudouridine; 2-thio-1-methyl-pseudouridine; 2-thio-5-aza-uridine; 2-thio- dihydropseudouridine; 2-thio-dihydrouridine; 2-thio-pseudouridine; 4-methoxy-2-thio- pseudouridine; 4-methoxy-pseudouridine; 4-thio-1 -methyl-pseudouridine; 4-thio- pseudouridine; 5-aza-uridine; Dihydropseudouridine; (±)1 -(2-Hydroxypropyl)pseudouridine TP; (2R)-1 -(2-Hydroxypropyl)pseudouridine TP; (2S)-1-(2-Hydroxypropyl)pseudouridine TP; (E)-5-(2-Bromo-vinyl)ara-uridine TP; (E)-5-(2-Bromo-vinyl)uridine TP; (Z)-5-(2-Bromo- vinyl)ara-uridine TP; (Z)-5-(2-Bromo-vinyl)uridine TP; 1 -(2,2,2-Trifluoroethyl)-pseudo-UTP; 1 - (2,2,3,3,3-Pentafluoropropyl)pseudouridine TP; 1 -(2,2-Diethoxyethyl)pseudouridine TP; 1 - (2,4,6-Trimethylbenzyl)pseudouridine TP; 1 -(2,4,6-Trimethyl-benzyl)pseudo-UTP; 1 -(2,4,6- Trimethyl-phenyl)pseudo-UTP; 1 -(2-Amino-2-carboxyethyl)pseudo-UTP; 1-(2-Amino- ethyl)pseudo-UTP; 1 -(2-Hydroxyethyl)pseudouridine TP; 1-(2-Methoxyethyl)pseudouridine TP; 1-(3,4-Bis-trifluoromethoxybenzyl)pseudouridine TP; 1 -(3,4-Dimethoxybenzyl)pseudouridine TP; 1 -(3-Amino-3-carboxypropyl)pseudo-UTP; 1-(3-Amino- propyl)pseudo-UTP; 1 -(3-Cyclopropyl-prop-2-ynyl)pseudouridine TP; 1-(4-Amino-4- carboxybutyl)pseudo-UTP; 1 -(4-Amino-benzyl)pseudo-UTP; 1 -(4-Amino-butyl)pseudo-UTP; 1 -(4-Amino-phenyl)pseudo-UTP; 1-(4-Azidobenzyl)pseudouridine TP; 1-(4- Bromobenzyl)pseudouridine TP; 1 -(4-Chlorobenzyl)pseudouridine TP; 1 -(4-Fluorobenzyl)pseudouridine TP; 1 -(4-lodobenzyl)pseudouridine TP; 1 -(4- Methanesulfonylbenzyl)pseudouridine TP; 1 -(4-Methoxybenzyl)pseudouridine TP; 1 -(4- Methoxy-benzyl)pseudo-UTP; 1 -(4-Methoxy-phenyl)pseudo-UTP; 1 -(4-Methylbenzyl)pseudouridine TP; 1 -(4-Methyl-benzyl)pseudo-UTP; 1 -(4- Nitrobenzyl)pseudouridine TP; 1-(4-Nitro-benzyl)pseudo-UTP; 1 (4-Nitro-phenyl)pseudo-UTP; 1 -(4-Thiomethoxybenzyl)pseudouridine TP; 1 -(4-Trifluoromethoxybenzyl)pseudouridine TP; 1 -(4-Trifluoromethylbenzyl)pseudouridine TP; 1 -(5-Amino-pentyl)pseudo-UTP; 1 -(6-Amino- hexyl)pseudo-UTP; 1 ,6-Dimethyl-pseudo-UTP; 1 -[3-(2-{2-[2-(2-Aminoethoxy)-ethoxy]- ethoxy}-ethoxy)-propionyl]pseudouridine TP; 1 -{3-[2-(2-Aminoethoxy)-ethoxy]-propionyl} pseudouridine TP; 1 -Acetylpseudouridine TP; 1 -Alkyl-6-(1 -propynyl)-pseudo-UTP; 1 -Alkyl-6- (2-propynyl)-pseudo-UTP; 1 -Alkyl-6-allyl-pseudo-UTP; 1-Alkyl-6-ethynyl-pseudo-UTP; 1 - Alkyl-6-homoallyl-pseudo-UTP; 1 -Alkyl-6-vinyl-pseudo-UTP; 1 -Allylpseudouridine TP; 1 - Aminomethyl-pseudo-UTP; 1 -Benzoylpseudouridine TP; 1 -Benzyloxymethylpseudouridine TP; 1 -Benzyl-pseudo-UTP; 1 -Biotinyl-PEG2-pseudouridine TP; 1 -Biotinylpseudouridine TP; 1 -Butyl-pseudo-UTP; 1 -Cyanomethylpseudouridine TP; 1 -Cyclobutylmethyl-pseudo-UTP; 1 - Cyclobutyl-pseudo-UTP; 1 -Cycloheptylmethyl-pseudo-UTP; 1-Cycloheptyl-pseudo-UTP; 1 - Cyclohexylmethyl-pseudo-UTP; 1 -Cyclohexyl-pseudo-UTP; 1 -Cyclooctylmethyl-pseudo-UTP; 1 -Cyclooctyl-pseudo-UTP; 1-Cyclopentylmethyl-pseudo-UTP; 1 -Cyclopentyl-pseudo-UTP; 1 - Cyclopropylmethyl-pseudo-UTP; 1 -Cyclopropyl-pseudo-UTP; 1 -Ethyl-pseudo-UTP; 1 -Hexyl- pseudo-UTP; 1 -Homoallylpseudouridine TP; 1 -Hydroxymethylpseudouridine TP; 1 -iso-propyl- pseudo-UTP; 1 -Me-2-thio-pseudo-UTP; 1 -Me-4-thio-pseudo-UTP; 1 -Me-alpha-thio-pseudo- UTP; 1 -Methanesulfonylmethylpseudouridine TP; 1 -Methoxymethylpseudouridine TP; 1 - Methyl-6-(2,2,2-Trifluoroethyl)pseudo-UTP; 1 -Methyl-6-(4-morpholino)-pseudo-UTP; 1 - Methyl-6-(4-thiomorpholino)-pseudo-UTP; 1 -Methyl-6-(substituted phenyl)pseudo-UTP; 1 - Methyl-6-amino-pseudo-UTP; 1 -Methyl-6-azido-pseudo-UTP; 1 -Methyl-6-bromo-pseudo- UTP; 1 -Methyl-6-butyl-pseudo-UTP; 1 -Methyl-6-chloro-pseudo-UTP; 1 -Methyl-6-cyano- pseudo-UTP; 1 -Methyl-6-dimethylamino-pseudo-UTP; 1 -Methyl-6-ethoxy-pseudo-UTP; 1 - Methyl-6-ethylcarboxylate-pseudo-UTP; 1 -Methyl-6-ethyl-pseudo-UTP; 1 -Methyl-6-fluoro- pseudo-UTP; 1 -Methyl-6-formyl-pseudo-UTP; 1 -Methyl-6-hydroxyamino-pseudo-UTP; 1 - Methyl-6-hydroxy-pseudo-UTP; 1 -Methyl-6-iodo-pseudo-UTP; 1 -Methyl-6-iso-propyl-pseudo- UTP; 1 -Methyl-6-methoxy-pseudo-UTP; 1 -Methyl-6-methylamino-pseudo-UTP; 1 -Methyl-6- phenyl-pseudo-UTP; 1 -Methyl-6-propyl-pseudo-UTP; 1 -Methyl-6-tert-butyl-pseudo-UTP; 1 - Methyl-6-trifluoromethoxy-pseudo-UTP; 1 -Methyl-6-trifluoromethyl-pseudo-UTP; 1 - Morpholinomethylpseudouridine TP; 1-Pentyl-pseudo-UTP; 1 -Phenyl-pseudo-UTP; 1 - Pivaloylpseudouridine TP; 1 -Propargylpseudouridine TP; 1 -Propyl-pseudo-UTP; 1 -propynyl- pseudouridine; 1 -p-tolyl-pseudo-UTP; 1 -tert-Butyl-pseudo-UTP; 1 - Thiomethoxymethylpseudouridine TP; 1 -Thiomorpholinomethylpseudouridine TP; 1 -Trifluoroacetylpseudouridine TP; 1 -Trifluoromethyl-pseudo-UTP; 1 -Vinylpseudouridine TP; 2,2'-anhydro-uridine TP; 2'-bromo-deoxyuridine TP; 2'-F-5-Methyl-2'-deoxy-UTP; 2'-OMe-5- Me-UTP; 2'-OMe-pseudo-UTP; 2’-a-Ethynyluridine TP; 2'-a-Trifluoromethyluridine TP; 2'-b- Ethynyluridine TP; 2'-b-Trifluoromethyluridine TP; 2’-Deoxy-2',2'-difluorouridine TP; 2'-Deoxy- 2'-a-mercaptouridine TP; 2'-Deoxy-2'-a-thiomethoxyuridine TP; 2'-Deoxy-2'-b-aminouridine TP; 2'-Deoxy-2'-b-azidouridine TP; 2'-Deoxy-2'-b-bromouridine TP; 2'-Deoxy-2'-b- chlorouridine TP; 2'-Deoxy-2'-b-fluorouridine TP; 2'-Deoxy-2'-b-iodouridine TP; 2'-Deoxy-2'-b- mercaptouridine TP; 2'-Deoxy-2'-b-thiomethoxyuridine TP; 2-methoxy-4-thio-uridine; 2- methoxyuridine; 2'-O-Methyl-5-(1 -propynyl)uridine TP; 3-Alkyl-pseudo-UTP; 4'-Azidouridine TP; 4'-Carbocyclic uridine TP; 4'-Ethynyluridine TP; 5-(1 -Propynyl)ara-uridine TP; 5-(2- Furanyl)uridine TP; 5-Cyanouridine TP; 5-Dimethylaminouridine TP; 5'-Homo-uridine TP; 5- iodo-2'-fluoro-deoxyuridine TP; 5-Phenylethynyluridine TP; 5-Trideuteromethyl-6- deuterouridine TP; 5-Trifluoromethyl-Uridine TP; 5-Vinylarauridine TP; 6-(2,2,2- Trifluoroethyl)-pseudo-UTP; 6-(4-Morpholino)-pseudo-UTP; 6-(4-Thiomorpholino)-pseudo- UTP; 6-(Substituted-Phenyl)-pseudo-UTP; 6-Amino-pseudo-UTP; 6-Azido-pseudo-UTP; 6- Bromo-pseudo-UTP; 6-Butyl-pseudo-UTP; 6-Chloro-pseudo-UTP; 6-Cyano-pseudo-UTP; 6- Dimethylamino-pseudo-UTP; 6-Ethoxy-pseudo-UTP; 6-Ethylcarboxylate-pseudo-UTP; 6- Ethyl-pseudo-UTP; 6-Fluoro-pseudo-UTP; 6-Formyl-pseudo-UTP; 6-Hydroxyamino-pseudo- UTP; 6-Hydroxy-pseudo-UTP; 6-lodo-pseudo-UTP; 6-iso-Propyl-pseudo-UTP; 6-Methoxy- pseudo-UTP; 6-Methylamino-pseudo-UTP; 6-Methyl-pseudo-UTP; 6-Phenyl-pseudo-UTP; 6- Phenyl-pseudo-UTP; 6-Propyl-pseudo-UTP; 6-tert-Butyl-pseudo-UTP; 6-Trifluoromethoxy- pseudo-UTP; 6-Trifluoromethyl-pseudo-UTP; Alpha-thio-pseudo-UTP; Pseudouridine 1 -(4- methylbenzenesulfonic acid) TP; Pseudouridine 1 -(4-methylbenzoic acid) TP; Pseudouridine TP 1 -[3-(2-ethoxy)]propionic acid; Pseudouridine TP 1 -[3-{2-(2-[2-(2-ethoxy)-ethoxy]-ethoxy)- ethoxy}]propionic acid; Pseudouridine TP 1-[3-{2-(2-[2-{2(2-ethoxy)-ethoxy}-ethoxy]-ethoxy)- ethoxy}]propionic acid; Pseudouridine TP 1 -[3-{2-(2-[2-ethoxy]-ethoxy)-ethoxy}]propionic acid; Pseudouridine TP 1 -[3-{2-(2-ethoxy)-ethoxy}]propionic acid; Pseudouridine TP 1- methylphosphonic acid; Pseudouridine TP 1 -methylphosphonic acid diethyl ester; Pseudo- UTP-N1 -3-propionic acid; Pseudo-UTP-N1 -4-butanoic acid; Pseudo-UTP-N1 -5-pentanoic acid; Pseudo-UTP-N1 -6-hexanoic acid; Pseudo-UTP-N1 -7-heptanoic acid; Pseudo-UTP-N1 - methyl-p-benzoic acid; Pseudo-UTP-N1 -p-benzoic acid; Wybutosine; Hydroxywybutosine; Isowyosine; Peroxywybutosine; undermodified hydroxywybuto sine; 4-demethylwyosine; 2,6- (diamino)purine;1 -(aza)-2-(thio)-3-(aza)-phenoxazin-1 -yl: 1 ,3-(diaza)-2-(oxo)-phenthiazin-1 - yl; 1 ,3-(diaza)-2-(oxo)-phenoxazin-1 -yl; 1 ,3,5-(triaza)-2,6-(dioxa)-naphthalene; 2(amino)purine;2,4,5-(trimethyl)phenyl;2Tnethyl, 2'amino, 2'azido, 2'fluro-cytidine;2'methyl, 2'amino, 2'azido, 2'fluro-adenine;2'methyl, 2'amino, 2'azido, 2'fluro-uridine;2'-amino-2'- deoxyribose; 2-amino-6-Chloro-purine; 2-aza-inosinyl; 2'-azido-2'-deoxyribose; 2'fluoro-2'-deoxyribose; 2'-fluoro-modified bases; 2'-O-methyl-ribose; 2-oxo-7-aminopyridopyrimidin-3-yl; 2-oxo-pyridopyrimidine-3-yl; 2-pyridinone; 3 nitropyrrole; 3-(methyl)-7- (propynyl)isocarbostyrilyl; 3-(methyl)isocarbostyrilyl; 4-(fluoro)-6-(methyl)benzimidazole; 4- (methyl)benzimidazole; 4-(methyl)indolyl; 4,6-(dimethyl)indolyl; 5 nitroindole; 5 substituted pyrimidines; 5-(methyl)isocarbostyrilyl; 5-nitroindole; 6-(aza)pyrimidine; 6-(azo)thymine; 6- (methyl)-7-(aza)indolyl; 6-chloro-purine; 6-phenyl-pyrrolo-pyrimidin-2-on-3-yl; 7- (aminoalkylhydroxy)-l -(aza)-2-(thio)-3-(aza)-phenthiazin-1 -yl; 7-(aminoalkylhydroxy)-1 -(aza)-2-(thio)-3-(aza)-phenoxazin-1 -yl; 7-(aminoalkylhydroxy)-1 ,3-(diaza)-2-(oxo)-phenoxazin-1 -yl; 7-(aminoalkylhydroxy)-1 ,3-(diaza)-2-(oxo)-phenthiazin-1 -yl; 7-(aminoalkylhydroxy)-1 ,3- (diaza)-2-(oxo)-phenoxazin-1 -yl; 7-(aza)indolyl; 7-(guanidiniumalkylhydroxy)-1 -(aza)-2-(thio)-3-(aza)-phenoxazinl-yl; 7-(guanidiniumalkylhydroxy)-1 -(aza)-2-(thio)-3-(aza)-phenthiazin-1 -yl;7-(guanidiniumalkylhydroxy)-1 -(aza)-2-(thio)-3-(aza)-phenoxazin-1 -yl; 7-(guanidiniumalkylhydroxy)-l ,3-(diaza)-2-(oxo)-phenoxazin-1 -yl; 7-(guanidiniumalkyl- hydroxy)-1 ,3-(diaza)-2-(oxo)-phenthiazin-1 -yl; 7-(guanidiniumalkylhydroxy)-1 ,3-(diaza)-2- (oxo)-phenoxazin-l -yl; 7-(propynyl)isocarbostyrilyl; 7-(propynyl)isocarbostyrilyl, propynyl-7- (aza)indolyl; 7-deaza-inosinyl; 7-substituted 1-(aza)-2-(thio)-3-(aza)-phenoxazin-1 -yl; 7- substituted 1 ,3-(diaza)-2-(oxo)-phenoxazin-1 -yl; 9-(methyl)-imidizopyridinyl; Aminoindolyl; Anthracenyl; bis-ortho-(aminoalkylhydroxy)-6-phenyl-pyrrolo-pyrimidin-2-on-3-yl; bis-ortho- substituted-6-phenyl-pyrrolo-pyrimidin-2-on-3-yl; Difluorotolyl; Hypoxanthine; Imidizopyridinyl; Inosinyl; Isocarbostyrilyl; Isoguanisine; N2-substituted purines; N6-methyl-2-amino-purine; N6-substituted purines; N-alkylated derivative; Napthalenyl; Nitrobenzimidazolyl; Nitroimidazolyl; Nitroindazolyl; Nitropyrazolyl; Nubularine; 06-substituted purines; O-alkylated derivative; ortho-(aminoalkylhydroxy)-6-phenyl-pyrrolo-pyrimidin-2-on-3-yl; ortho-substituted- 6-phenyl-pyrrolo-pyrimidin-2-on-3-yl; Oxoformycin TP; para-(aminoalkylhydroxy)-6-phenyl- pyrrolo-pyrimidin-2-on-3-yl; para-substituted-6-phenyl-pyrrolo-pyrimidin-2-on-3-yl;Pentacenyl; Phenanthracenyl; Phenyl; propynyl-7-(aza)indolyl; Pyrenyl; pyridopyrimidin-3-yl; pyridopyrimidin-3-yl, 2-oxo-7-amino-pyridopyrimidin-3-yl; pyrrolo-pyrimidin-2-on-3-yl; Pyrrolopyrimidinyl; Pyrrolopyrizinyl; Stilbenzyl; substituted 1 ,2,4-triazoles; Tetracenyl; Tubercidine; Xanthine; Xanthosine-5'-TP; 2-thio-zebularine; 5-aza-2-thio-zebularine; 7-deaza- 2-amino-purine; pyridin-4-one ribonucleoside; 2-Amino-riboside-TP; Formycin A TP; Formycin B TP; Pyrrolosine TP; 2'-OH-ara-adenosine TP; 2'-OH-ara-cytidine TP; 2'-OH-ara-uridine TP; 2'-OH-ara-guanosine TP; 5-(2-carbomethoxyvinyl)uridine TP; and N6-(19-Amino- pentaoxanonadecyl)adenosine TP.

[0096] In some embodiments, the polynucleotide (e.g., RNA polynucleotide, such as mRNA polynucleotide) includes a combination of at least two (e.g., 2, 3, 4 or more) of the aforementioned modified nucleobases.

[0097] In some embodiments, the mRNA comprises at least one chemically modified nucleoside. In some embodiments, the at least one chemically modified nucleoside is selected from the group consisting of pseudouridine (i ), 2-thiouridine (s2U), 4'-thiouridine, 5- methylcytosine, 2-thio-1 -methyl-1 -deaza-pseudouridine, 2-thio-1 -methyl-pseudouridine, 2- thio-5-aza-uridine, 2-thio-dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-pseudouridine, 4-methoxy-2-thio-pseudouridine, 4-methoxy-pseudouridine, 4-th io- 1 -methyl-pseudouridine, 4-thio-pseudouridine, 5-aza-uridine, dihydropseudouridine, 5-methyluridine, 5- methoxyuridine, 2'-O-methyl uridine, 1 -methyl-pseudouridine (ml i ), 1 -ethyl-pseudouridine (el ip), 5-methoxy-uridine (mo5U), 5-methyl-cytidine (m5C), a-thio-guanosine, a-thio- adenosine, 5-cyano uridine, 4'-thio uridine 7-deaza-adenine, 1 -methyl-adenosine (m1 A), 2- methyl-adenine (m2A), N6-methyl-adenosine (m6A), and 2,6-Diaminopurine, (I), 1 -methylinosine (ml I), wyosine (imG), methylwyosine (mimG), 7-deaza-guanosine, 7-cyano-7-deaza- guanosine (preQO), 7-aminomethyl-7-deaza-guanosine (preQ1 ), 7-methyl-guanosine (m7G), 1 -methyl-guanosine (m1 G), 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 2,8- dimethyladenosine, 2-geranylthiouridine, 2-lysidine, 2-selenouridine, 3-(3-amino-3- carboxypropyl)-5,6-dihydrouridine, 3-(3-amino-3-carboxypropyl)pseudouridine, 3- methylpseudouridine, 5-(carboxyhydroxymethyl)-2'-O-methyluridine methyl ester, 5- aminomethyl-2-geranylthiouridine, 5-aminomethyl-2-selenouridine, 5-aminomethyluridine, 5- carbamoylhydroxymethyluridine, 5-carbamoylmethyl-2-thiouridine, 5-carboxymethyl-2- thiouridine, 5-carboxymethylaminomethyl-2-geranylthiouridine, 5- carboxymethylaminomethyl-2-selenouridine, 5-cyanomethyluridine, 5-hydroxycytidine, 5- methylaminomethyl-2-geranylthiouridine, 7-aminocarboxypropyl-demethylwyosine, 7- aminocarboxypropylwyosine, 7-aminocarboxypropylwyosine methyl ester, 8- methyladenosine, N4,N4-dimethylcytidine, N6-formyladenosine, N6- hydroxymethyladenosine, agmatidine, cyclic N6-threonylcarbamoyladenosine, glutamyl- queuosine, methylated undermodified hydroxywybutosine, N4,N4,2'-O-trimethylcytidine, geranylated 5-methylaminomethyl-2-thiouridine, geranylated 5-carboxymethylaminomethyl-2- thiouridine, Qbase , preQObase, preQI base, and two or more combinations thereof. In some embodiments, the at least one chemically modified nucleoside is selected from the group consisting of pseudouridine, 1 -methyl-pseudouridine, 1 -ethyl-pseudouridine, 5- methylcytosine, 5-methoxyuridine, and a combination thereof. In some embodiments, the polynucleotide (e.g., RNA polynucleotide, such as mRNA polynucleotide) includes a combination of at least two (e.g., 2, 3, 4 or more) of the aforementioned modified nucleobases.

[0098] In certain aspects, the present disclosure provides nucleic acid molecules, specifically polynucleotides that encode one or more Rv2140c antigens, or functional fragments thereof. Features, which can be considered beneficial in some embodiments of the present disclosure, can be encoded by regions of the polynucleotide and such regions can be upstream (5') ordownstream (3') to, or within, a region that encodes a polypeptide. These regions can be incorporated into the polynucleotide before and / or after sequence optimization of the protein encoding region or open reading frame (ORF). It is not required that a polynucleotide contain both a 5' and 3' flanking region. Examples of such features include, but are not limited to, untranslated regions (UTRs), Kozak sequences, an oligo(dT) sequence, and detectable tags and can include multiple cloning sites that can have Xbal recognition.

[0099] In some embodiments, a 5' UTR and / or a 3' UTR region can be provided as flanking regions. Multiple 5' or 3' UTRs can be included in the flanking regions and can be the same or of different sequences. Any portion of the flanking regions, including none, can be sequence- optimized and any can independently contain one or more different structural or chemical modifications, before and / or after sequence optimization.

[0100] Untranslated regions (UTRs) are nucleic acid sections of a polynucleotide before a start codon (5'UTR) and after a stop codon (3'UTR) that are not translated. In some embodiments, a polynucleotide (e.g., a ribonucleic acid (RNA), e.g., a messenger RNA (mRNA)) of the invention comprising an open reading frame (ORF) encoding an antigen polypeptide further comprises UTR (e.g., a 5'UTR or functional fragment thereof, a 3'UTR or functional fragment thereof, or a combination thereof).

[0101] A UTR can be homologous or heterologous to the coding region in a polynucleotide. In some embodiments, the UTR is homologous to the ORF encoding the antigen polypeptide. In some embodiments, the UTR is heterologous to the ORF encoding the antigen polypeptide. In some embodiments, the polynucleotide comprises two or more 5'UTRs or functional fragments thereof, each of which have the same or different nucleotide sequences. In some embodiments, the polynucleotide comprises two or more 3'UTRs or functional fragments thereof, each of which have the same or different nucleotide sequences.

[0102] In some embodiments, the 5'UTR or functional fragment thereof, 3' UTR or functional fragment thereof, or any combination thereof is sequence optimized. In some embodiments, the 5'UTR or functional fragment thereof, 3' UTR or functional fragment thereof, or any combination thereof comprises at least one chemically modified nucleobase, e.g., 5- methoxyuracil.

[0103] UTRs can have features that provide a regulatory role, e.g., increased or decreased stability, localization and / or translation efficiency. A polynucleotide comprising a UTR can be administered to a cell, tissue, or organism, and one or more regulatory features can be measured using routine methods. In some embodiments, a functional fragment of a 5'UTR or 3'UTR comprises one or more regulatory features of a full length 5' or 3' UTR, respectively.

[0104] By engineering the features typically found in abundantly expressed genes of specific target organs, one can enhance the stability and protein production of a polynucleotide. For example, introduction of 5'UTR of liver-expressed mRNA, such as albumin, serum amyloid A,Apolipoprotein A / B / E, transferrin, alpha fetoprotein, erythropoietin, or Factor VIII, can enhance expression of polynucleotides in hepatic cell lines or liver. Likewise, use of 5'UTR from other tissue-specific mRNA to improve expression in that tissue is possible for muscle (e.g., MyoD, Myosin, Myoglobin, Myogenin, Herculin), for endothelial cells (e.g., Tie-1 , CD36), for myeloid cells (e.g., C / EBP, AML1 , G-CSF, GM-CSF, CD11 b, MSR, Fr-1 , i-NOS), for leukocytes (e.g., CD45, CD18), for adipose tissue (e.g., CD36, GLUT4, ACRP30, adiponectin) and for lung epithelial cells (e.g., SP-A / B / C / D).

[0105] In some embodiments, UTRs are selected from a family of transcripts whose proteins share a common function, structure, feature or property. For example, an encoded polypeptide can belong to a family of proteins (i.e., that share at least one function, structure, feature, localization, origin, or expression pattern), which are expressed in a particular cell, tissue or at some time during development. The UTRs from any of the genes or mRNA can be swapped for any other UTR of the same or different family of proteins to create a new polynucleotide.

[0106] In some embodiments, the 5'UTR and the 3'UTR can be heterologous. In some embodiments, the 5'UTR can be derived from a different species than the 3'UTR. In some embodiments, the 3'UTR can be derived from a different species than the 5'UTR.

[0107] Exemplary UTRs of the application include, but are not limited to, one or more 5'UTR and / or 3'UTR derived from the nucleic acid sequence of: a globin, such as an a- or p-globin (e.g., a Xenopus, mouse, rabbit, or human globin); a strong Kozak translational initiation signal; a CYBA (e.g., human cytochrome b-245 a polypeptide); an albumin (e.g., human albumin?); a HSD17B4 (hydroxysteroid (17-P) dehydrogenase); a virus (e.g., a tobacco etch virus (TEV), a Venezuelan equine encephalitis virus (VEEV), a Dengue virus, a cytomegalovirus (CMV) (e.g., CMV immediate early 1 (IE1 )), a hepatitis virus (e.g., hepatitis B virus), a sindbis virus, or a PAV barley yellow dwarf virus); a heat shock protein (e.g., hsp70); a translation initiation factor (e.g., elF4G); a glucose transporter (e.g., hGLUTI (human glucose transporter 1 )); an actin (e.g., human a or p actin); a GAPDH; a tubulin; a histone; a citric acid cycle enzyme; a topoisomerase (e.g., a 5'UTR of a TOP gene lacking the 5' TOP motif (the oligopyrimidine tract)); a ribosomal protein Large 32 (L32); a ribosomal protein (e.g., human or mouse ribosomal protein, such as, for example, rps9); an ATP synthase (e.g., ATP5A1 or the p subunit of mitochondrial H+-ATP synthase); a growth hormone e (e.g., bovine (bGH) or human (hGH)); an elongation factor (e.g., elongation factor 1 a1 (EEF1 A1 )); a manganese superoxide dismutase (MnSOD); a myocyte enhancer factor 2A (MEF2A); a p- F1 -ATPase, a creatine kinase, a myoglobin, a granulocyte-colony stimulating factor (G-CSF); a collagen (e.g., collagen type I, alpha 2 (Col1A2), collagen type I, alpha 1 (Col1A1 ), collagen type VI, alpha 2 (Col6A2), collagen type VI, alpha 1 (C0I6AI )); a ribophorin (e.g., ribophorin I (RPNI)); a low density lipoprotein receptor-related protein (e.g., LRP1 ); a cardiotrophin-likecytokine factor (e.g., Nnt1 ); calreticulin (Calr); a procollagen-lysine, 2-oxoglutarate 5- dioxygenase 1 (Plodl ); and a nucleobindin (e.g., Nucbl ).

[0108] In some embodiments, the 5'UTR is selected from the group consisting of a p-globin 5'UTR; a 5'UTR containing a strong Kozak translational initiation signal; a cytochrome b-245 a polypeptide (CYBA) 5'UTR; a hydroxysteroid (17- ) dehydrogenase (HSD17B4) 5'UTR; a Tobacco etch virus (TEV) 5'UTR; a Venezuelen equine encephalitis virus (TEEV) 5'UTR; a 5' proximal open reading frame of rubella virus (RV) RNA encoding nonstructural proteins; a Dengue virus (DEN) 5'UTR; a heat shock protein 70 (Hsp70) 5'UTR; a elF4G 5'UTR; a GLUT1 5'UTR; functional fragments thereof and any combination thereof.

[0109] In some embodiments, the 3'UTR is selected from the group consisting of a f3-globin 3'UTR; a CYBA 3'UTR; an albumin 3'UTR; a growth hormone (GH) 3'UTR; a VEEV 3'UTR; a hepatitis B virus (HBV) 3'UTR; a-globin 3'UTR; a DEN 3'UTR; a PAV barley yellow dwarf virus (BYDV-PAV) 3'UTR; an elongation factor 1 a1 (EEF1 A1 ) 3'UTR; a manganese superoxide dismutase (MnSOD) 3'UTR; a p subunit of mitochondrial H(+)-ATP synthase (p-mRNA) 3’UTR; a GLUT1 3'UTR; a MEF2A 3'UTR; a 0-F1 -ATPase 3’UTR; functional fragments thereof and combinations thereof.[001 10] Wild-type UTRs derived from any gene or mRNA can be incorporated into the polynucleotides of the invention. In some embodiments, a UTR can be altered relative to a wild type or native UTR to produce a variant UTR, e.g., by changing the orientation or location of the UTR relative to the ORF; or by inclusion of additional nucleotides, deletion of nucleotides, swapping or transposition of nucleotides. In some embodiments, variants of 5' or 3’ UTRs can be utilized, for example, mutants of wild type UTRs, or variants wherein one or more nucleotides are added to or removed from a terminus of the UTR.[001 11] Additionally, one or more synthetic UTRs can be used in combination with one or more non-synthetic UTRs. See, e.g., Mandal and Rossi, Nat. Protoc. 2013 8(3):568-82, and sequences available at addgene.org / Derrick_Rossi / , the contents of each are incorporated herein by reference in their entirety. UTRs or portions thereof can be placed in the same orientation as in the transcript from which they were selected or can be altered in orientation or location. Hence, a 5' and / or 3' UTR can be inverted, shortened, lengthened, or combined with one or more other 5' UTRs or 3' UTRs. In some embodiments, the polynucleotide comprises multiple UTRs, e.g., a double, a triple or a quadruple 5'UTR or 3'UTR. For example, a double UTR comprises two copies of the same UTR either in series or substantially in series. For example, a double beta-globin 3'UTR can be used (see US2010 / 0129877, the contents of which are incorporated herein by reference in its entirety).

[0112] In some embodiments, the polynucleotides of the invention comprise a 5'UTR and / or a 3'UTR selected from any one of the UTRs disclosed herein. The polynucleotides of the invention can comprise combinations of features. For example, the ORF can be flanked by a5'UTR that comprises a strong Kozak translational initiation signal and / or a 3'UTR comprising an oligo(dT) sequence for templated addition of a poly-A tail. A 5'UTR can comprise a first polynucleotide fragment and a second polynucleotide fragment from the same and / or different UTRs (see, e.g., US2010 / 0293625, herein incorporated by reference in its entirety).

[0113] Other non-UTR sequences can be used as regions or subregions within the polynucleotides of the invention. For example, introns or portions of intron sequences can be incorporated into the polynucleotides of the invention. Incorporation of intronic sequences can increase protein production as well as polynucleotide expression levels. In some embodiments, the polynucleotide of the invention comprises an internal ribosome entry site (IRES) instead of or in addition to a UTR (see, e.g., Yakubov et al., Biochem. Biophys. Res. Commun. 2010 394(1 ):189-193, the contents of which are incorporated herein by reference in their entirety). In some embodiments, the polynucleotide comprises an IRES instead of a 5'UTR sequence. In some embodiments, the polynucleotide comprises an ORF and a viral capsid sequence. In some embodiments, the polynucleotide comprises a synthetic 5'UTR in combination with a non-synthetic 3'UTR.

[0114] In some embodiments, the UTR can also include at least one translation enhancer polynucleotide, translation enhancer element, or translational enhancer elements (collectively, “TEE,” which refers to nucleic acid sequences that increase the amount of polypeptide or protein produced from a polynucleotide. As a non-limiting example, the TEE can be located between the transcription promoter and the start codon. In some embodiments, the 5'UTR comprises a TEE. In one aspect, a TEE is a conserved element in a UTR that can promote translational activity of a nucleic acid such as, but not limited to, cap-dependent or capindependent translation.

[0115] In some embodiments, the polynucleotide of the invention comprises one or multiple copies of a TEE. The TEE in a translational enhancer polynucleotide can be organized in one or more sequence segments. A sequence segment can harbor one or more of the TEEs provided herein, with each TEE being present in one or more copies. When multiple sequence segments are present in a translational enhancer polynucleotide, they can be homogenous or heterogeneous. Thus, the multiple sequence segments in a translational enhancer polynucleotide can harbor identical or different types of the TEE provided herein, identical or different number of copies of each of the TEE, and / or identical or different organization of the TEE within each sequence segment. In one embodiment, the polynucleotide of the invention comprises a translational enhancer polynucleotide sequence.

[0116] In some embodiments, a 5'UTR and / or 3'UTR comprising at least one TEE described herein can be incorporated in a monocistronic sequence such as, but not limited to, a vector system or a nucleic acid vector.

[0117] In some embodiments, a 5'UTR and / or 3'UTR of a polynucleotide of the invention comprises a TEE or portion thereof described herein. In some embodiments, the TEEs in the 3'UTR can be the same and / or different from the TEE located in the 5'UTR.

[0118] In some embodiments, the spacer separating two TEE sequences can include other sequences known in the art that can regulate the translation of the polynucleotide of the invention, e.g., miR sequences described herein (e.g., miR binding sites). As a non-limiting example, each spacer used to separate two TEE sequences can include a different miR sequence (e.g., miR binding site).

[0119] In some embodiments, a polynucleotide of the invention comprises a miR and / or TEE sequence. In some embodiments, the incorporation of a miR sequence and / or a TEE sequence into a polynucleotide of the invention can change the shape of the stem loop region, which can increase and / or decrease translation. See e.g., Kedde et al., Nature Cell Biology 2010 12(10):1014-20, herein incorporated by reference in its entirety).

[0120] In some aspects an mRNA vaccine is formulated in a lipid nanoparticle (LNP). The use of LNPs enables the effective delivery of chemically modified or unmodified mRNA vaccines. In one set of embodiments, lipid nanoparticles (LNPs) are provided. In one embodiment, a lipid nanoparticle comprises lipids including an ionizable lipid (such as an ionizable cationic lipid), a structural lipid, a phospholipid, and mRNA. Each of the LNPs described herein may be used as a formulation for the mRNA described herein. In one embodiment, a lipid nanoparticle comprises an ionizable lipid, a structural lipid, a phospholipid, and mRNA. In some embodiments, the LNP comprises an ionizable lipid, a PEG-modified lipid, a phospholipid and a structural lipid. In some embodiments, the LNP has a molar ratio of about 20-60% ionizable lipid: about 5-25% phospholipid: about 25-55% structural lipid; and about 0.5-15% PEG-modified lipid. In some embodiments, the LNP comprises a molar ratio of about 50% ionizable lipid, about 1.5% PEG-modified lipid, about 38.5% structural lipid and about 10% phospholipid. In some embodiments, the LNP comprises a molar ratio of about 55% ionizable lipid, about 2.5% PEG lipid, about 32.5% structural lipid and about 10% phospholipid. In some embodiments, the ionizable lipid is an ionizable amino or cationic lipid and the phospholipid is a neutral lipid, and the structural lipid is a cholesterol. In some embodiments, the LNP has a molar ratio of 50:38.5:10:1 .5 of ionizable lipid: cholesterol:DSPC: PEG2000- DMG.

[0121] Ionizable lipids can be selected from the non-limiting group consisting of 3- (didodecylamino)-NI , N 1 ,4-tridodecyl- 1 -piperazineethanamine (KL10), N1 -[2- (didodecylamino)ethyl]-N1 ,N4,N4-tridodecyl-1 ,4-piperazinediethanamine (KL22), 14,25- ditridecyl-15,18,21 ,24-tetraaza-octatriacontane (KL25), 1 ,2-dilinoley loxy-N , N- dimethylaminopropane (DLin-DMA), 2,2-dilinoleyl-4-dimethylaminomethyl-[1 ,3]-dioxolane(DLin-K-DMA), heptatriaconta-6,9,28,31 -tetraen-19-yl 4-(dimethylamino)butanoate (DLin- MC3-DMA), 2,2-dilinoleyl-4-(2-dimethylaminoethyl)[1 ,3]-dioxolane (DLin-KC2-DMA), 1 ,2- dioleyloxy-N,N-dimethylaminopropane (DODMA), (13Z,165Z)-N,N-dimethyl-3-nonydocosa- 13-16-dien-1 -amine (L608), 2-({8-[(3|3)-cholest-5-en-3-yloxy]octyl}oxy)-N,N-dimethyl-3- [(9Z,12Z)-octadeca-9,12-dien-1 -yloxy]propan-1 -amine (Octyl-CLinDMA), (2R)-2-({8-[(3 )- cholest-5-en-3-yloxy]octyl}oxy)-N,N-dimethyl-3-[(9Z,12Z)-octadeca-9,12-dien-1 - yloxy]propan-1 -amine (Octyl-CLinDMA (2R)), and (2S)-2-({8-[(3|3)-cholest-5-en-3- yloxy]octyl}oxy)-N,N-dimethyl-3-[(9Z,12Z)-octadeca-9,12-dien-1 -yloxy]propan-1 -amine (Octyl-CLinDMA (2S)). In addition to these, an ionizable amino lipid can also be a lipid including a cyclic amine group.

[0122] The lipid composition of the pharmaceutical composition disclosed herein can comprise one or more phospholipids, for example, one or more saturated or (poly)unsaturated phospholipids or a combination thereof. In general, phospholipids comprise a phospholipid moiety and one or more fatty acid moieties.

[0123] A phospholipid moiety can be selected, for example, from the non-limiting group consisting of phosphatidyl choline, phosphatidyl ethanolamine, phosphatidyl glycerol, phosphatidyl serine, phosphatidic acid, 2-lysophosphatidyl choline, and a sphingomyelin.

[0124] A fatty acid moiety can be selected, for example, from the non-limiting group consisting of lauric acid, myristic acid, myristoleic acid, palmitic acid, palmitoleic acid, stearic acid, oleic acid, linoleic acid, alpha-linolenic acid, erucic acid, phytanoic acid, arachidic acid, arachidonic acid, eicosapentaenoic acid, behenic acid, docosapentaenoic acid, and docosahexaenoic acid.

[0125] Particular phospholipids can facilitate fusion to a membrane. For example, a cationic phospholipid can interact with one or more negatively charged phospholipids of a membrane (e.g., a cellular or intracellular membrane). Fusion of a phospholipid to a membrane can allow one or more elements (e.g., a therapeutic agent) of a lipid-containing composition (e.g., LNPs) to pass through the membrane permitting, e.g., delivery of the one or more elements to a target tissue.

[0126] Non-natural phospholipid species including natural species with modifications and substitutions including branching, oxidation, cyclization, and alkynes are also contemplated. For example, a phospholipid can be functionalized with or cross-linked to one or more alkynes (e.g., an alkenyl group in which one or more double bonds is replaced with a triple bond). Under appropriate reaction conditions, an alkyne group can undergo a copper-catalyzed cycloaddition upon exposure to an azide. Such reactions can be useful in functionalizing a lipid bilayer of a nanoparticle composition to facilitate membrane permeation or cellular recognition or in conjugating a nanoparticle composition to a useful component such as a targeting or imaging moiety (e.g., a dye).

[0127] Phospholipids include, but are not limited to, glycerophospholipids such as phosphatidylcholines, phosphatidylethanolamines, phosphatidylserines, phosphatidylinositols, phosphatidy glycerols, and phosphatidic acids. Phospholipids also include phosphosphingolipid, such as sphingomyelin.

[0128] In certain embodiments, a phospholipid useful or potentially useful in the present invention is an analog or variant of DSPC.

[0129] In certain embodiments, a phospholipid useful or potentially useful in the present invention comprises a modified phospholipid head (e.g., a modified choline group). In certain embodiments, a phospholipid with a modified head is DSPC, or analog thereof, with a modified quaternary amine.

[0130] In certain embodiments, a phospholipid useful or potentially useful in the present invention comprises a modified tail. In certain embodiments, a phospholipid useful or potentially useful in the present invention is DSPC, or analog thereof, with a modified tail. As described herein, a “modified tail” may be a tail with shorter or longer aliphatic chains, aliphatic chains with branching introduced, aliphatic chains with substituents introduced, aliphatic chains wherein one or more methylenes are replaced by cyclic or heteroatom groups, or any combination thereof.

[0131] In certain embodiments, an alternative lipid is used in place of a phospholipid of the invention.

[0132] The LNPs disclosed herein can comprise one or more structural lipids. As used herein, the term “structural lipid” refers to sterols and also to lipids containing sterol moieties.

[0133] Incorporation of structural lipids in the lipid nanoparticle may help mitigate aggregation of other lipids in the particle. Structural lipids can be selected from the group including but not limited to, cholesterol, fecosterol, sitosterol, ergosterol, campesterol, stigmasterol, brassicasterol, tomatidine, tomatine, ursolic acid, alpha-tocopherol, hopanoids, phytosterols, steroids, and mixtures thereof. In some embodiments, the structural lipid is a sterol. As defined herein, “sterols” are a subgroup of steroids consisting of steroid alcohols. In certain embodiments, the structural lipid is a steroid. In certain embodiments, the structural lipid is cholesterol. In certain embodiments, the structural lipid is an analog of cholesterol. In certain embodiments, the structural lipid is alpha-tocopherol.

[0134] In one embodiment, the amount of the structural lipid (e.g., an sterol such as cholesterol) in the lipid composition of a pharmaceutical composition disclosed herein ranges from about 20 mol % to about 60 mol %, from about 25 mol % to about 55 mol %, from about 30 mol % to about 50 mol %, or from about 35 mol % to about 45 mol %.

[0135] In one embodiment, the amount of the structural lipid (e.g., an sterol such as cholesterol) in the lipid composition disclosed herein ranges from about 25 mol % to about 30 mol %, from about 30 mol % to about 35 mol %, or from about 35 mol % to about 40 mol %.In one embodiment, the amount of the structural lipid (e.g., a sterol such as cholesterol) in the lipid composition disclosed herein is about 24 mol %, about 29 mol %, about 34 mol %, or about 39 mol %. In some embodiments, the amount of the structural lipid (e.g., an sterol such as cholesterol) in the lipid composition disclosed herein is at least about 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49, 50, 51 , 52, 53, 54, 55, 56, 57, 58, 59, or 60 mol %.

[0136] The lipid composition of a pharmaceutical composition disclosed herein can comprise one or more a polyethylene glycol (PEG) lipid. As used herein, the term “PEG-lipid” refers to polyethylene glycol (PEG)-modified lipids. Non-limiting examples of PEG-lipids include PEG- modified phosphatidylethanolamine and phosphatidic acid, PEG-ceramide conjugates (e.g., PEG-CerC14 or PEG-CerC20), PEG-modified dialkylamines and PEG-modified 1 ,2- diacyloxypropan-3-amines. Such lipids are also referred to as PEGylated lipids. For example, a PEG lipid can be PEG-c-DOMG, PEG-DMG, PEG-DLPE, PEG-DMPE, PEG-DPPC, or a PEG-DSPE lipid.

[0137] In some embodiments, the PEG-lipid includes, but not limited to 1 ,2-dimyristoyl-sn- glycerol methoxypolyethylene glycol (PEG-DMG), 1 ,2-distearoyl-sn-glycero-3- phosphoethanolamine-N-[amino(polyethylene glycol)] (PEG-DSPE), PEG-disteryl glycerol (PEG-DSG), PEG-dipalmetoleyl, PEG-dioleyl, PEG-distearyl, PEG-diacylglycamide (PEGDAG), PEG-dipalmitoyl phosphatidylethanolamine (PEG-DPPE), or PEG-1 , 2- dimyristyloxlpropyl-3-amine (PEG-c-DMA).

[0138] In one embodiment, the PEG-lipid is selected from the group consisting of a PEG- modified phosphatidylethanolamine, a PEG-modified phosphatidic acid, a PEG-modified ceramide, a PEG-modified dialkylamine, a PEG-modified di acylglycerol, a PEG-modified dialkylglycerol, and mixtures thereof. In some embodiments, the lipid moiety of the PEG-lipids includes those having lengths of from about C14 to about C22, preferably from about C14 to about C16. In some embodiments, a PEG moiety, for example an mPEG-NH2, has a size of about 1000, 2000, 5000, 10,000, 15,000 or 20,000 daltons. In one embodiment, the PEG-lipid is PEG2k-DMG. In one embodiment, the lipid nanoparticles described herein can comprise a PEG lipid which is a non-diffusible PEG. Non-limiting examples of non-diffusible PEGs include PEG-DSG and PEG-DSPE.

[0139] In one embodiment, PEG lipids useful in the present invention can be PEGylated lipids described in International Publication No. WO2012 / 099755, the contents of which is herein incorporated by reference in its entirety. Any of these exemplary PEG lipids described herein may be modified to comprise a hydroxyl group on the PEG chain. In certain embodiments, the PEG lipid is a PEG-OH lipid. As generally defined herein, a “PEG-OH lipid” (also referred to herein as “hydroxy-PEGylated lipid”) is a PEGylated lipid having one or more hydroxyl ( — OH) groups on the lipid. In certain embodiments, the PEG-OH lipid includes one or more hydroxylgroups on the PEG chain. In certain embodiments, a PEG-OH or hydroxy- PEGylated lipid comprises an — OH group at the terminus of the PEG chain. Each possibility represents a separate embodiment of the present invention.

[0140] In one embodiment, the amount of PEG-lipid in the lipid composition of a pharmaceutical composition disclosed herein ranges from about 0.1 mol % to about 5 mol %, from about 0.5 mol % to about 5 mol %, from about 1 mol % to about 5 mol %, from about 1 .5 mol % to about 5 mol %, from about 2 mol % to about 5 mol %, from about 0.1 mol % to about 4 mol %, from about 0.5 mol % to about 4 mol %, from about 1 mol % to about 4 mol %, from about 1 .5 mol % to about 4 mol %, from about 2 mol % to about 4 mol %, from about 0.1 mol % to about 3 mol %, from about 0.5 mol % to about 3 mol %, from about 1 mol % to about 3 mol %, from about 1 .5 mol % to about 3 mol %, from about 2 mol % to about 3 mol %, from about 0.1 mol % to about 2 mol %, from about 0.5 mol % to about 2 mol %, from about 1 mol % to about 2 mol %, from about 1 .5 mol % to about 2 mol %, from about 0.1 mol % to about 1 .5 mol %, from about 0.5 mol % to about 1 .5 mol %, or from about 1 mol % to about 1 .5 mol %.

[0141] In one embodiment, the amount of PEG-lipid in the lipid composition disclosed herein is about 2 mol %. In one embodiment, the amount of PEG-lipid in the lipid composition disclosed herein is about 1 .5 mol %. In one embodiment, the amount of PEG-lipid in the lipid composition disclosed herein is at least about 0.1 , 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1 , 1 .1 , 1 .2, 1 .3, 1.4, 1.5, 1 .6, 1.7, 1 .8, 1 .9, 2, 2.1 ,2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3, 3.1 , 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4, 4.1 , 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, or 5 mol %.

[0142] In some aspects, the lipid composition of the pharmaceutical compositions disclosed herein does not comprise a PEG-lipid.

[0143] The lipid composition of a pharmaceutical composition disclosed herein can include one or more components in addition to those described above. For example, the lipid composition can include one or more permeability enhancer molecules, carbohydrates, polymers, surface altering agents (e.g., surfactants), or other components. For example, a permeability enhancer molecule can be a molecule described by U.S. Patent Application Publication No. 2005 / 0222064. Carbohydrates can include simple sugars (e.g., glucose) and polysaccharides (e.g., glycogen and derivatives and analogs thereof).

[0144] A polymer can be included in and / or used to encapsulate or partially encapsulate a pharmaceutical composition disclosed herein (e.g., a pharmaceutical composition in lipid nanoparticle form). A polymer can be biodegradable and / or biocompatible. A polymer can be selected from, but is not limited to, polyamines, polyethers, polyamides, polyesters, polycarbamates, polyureas, polycarbonates, polystyrenes, polyimides, polysulfones, polyurethanes, polyacetylenes, polyethylenes, polyethyleneimines, polyisocyanates, polyacrylates, polymethacrylates, polyacrylonitriles, and polyarylates.

[0145] The ratio between the lipid composition and the polynucleotide range can be from about 10:1 to about 60:1 (wt / wt).

[0146] In some embodiments, the ratio between the lipid composition and the polynucleotide can be about 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1 , 20:1 , 21 :1 , 22:1 , 23:1 , 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34:1, 35:1, 36:1, 37:1, 38:1, 39:1,40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51:1, 52:1, 53:1, 54:1, 55:1,56:1, 57:1, 58:1, 59:1 or 60:1 (wt / wt). In some embodiments, the wt / wt ratio of the lipid composition to the polynucleotide encoding a therapeutic agent is about 20:1 or about 15:1.

[0147] In one embodiment, the lipid nanoparticles described herein can comprise polynucleotides (e.g., mRNA) in a lipid:polynucleotide weight ratio of 5:1, 10:1, 15:1, 20:1, 25:1 , 30:1 , 35:1 , 40:1 , 45:1 , 50:1 , 55:1 , 60:1 or 70:1 , or a range or any of these ratios such as, but not limited to, 5:1 to about 10:1, from about 5:1 to about 15:1, from about 5:1 to about 20:1 , from about 5:1 to about 25:1 , from about 5:1 to about 30:1 , from about 5:1 to about 35:1 , from about 5:1 to about 40:1 , from about 5:1 to about 45:1 , from about 5:1 to about 50:1 , from about 5:1 to about 55:1 , from about 5:1 to about 60:1 , from about 5:1 to about 70:1 , from about 10:1 to about 15:1, from about 10:1 to about 20:1, from about 10:1 to about 25:1, from about10:1 to about 30:1 , from about 10:1 to about 35:1, from about 10:1 to about 40:1 , from about10:1 to about 45:1, from about 10:1 to about 50:1, from about 10:1 to about 55:1, from about10:1 to about 60:1, from about 10:1 to about 70:1, from about 15:1 to about 20:1, from about15:1 to about 25:1 ,from about 15:1 to about 30:1 , from about 15:1 to about 35:1 , from about 15:1 to about 40:1, from about 15:1 to about 45:1, from about 15:1 to about 50:1, from about 15:1 to about 55:1 , from about 15:1 to about 60:1 or from about 15:1 to about 70:1.

[0148] In one embodiment, the lipid nanoparticles described herein can comprise the polynucleotide in a concentration from approximately 0.1 mg / ml to 2 mg / ml such as, but not limited to, 0.1 mg / ml, 0.2 mg / ml, 0.3 mg / ml, 0.4 mg / ml, 0.5 mg / ml, 0.6 mg / ml, 0.7 mg / ml, 0.8 mg / ml, 0.9 mg / ml, 1.0 mg / ml, 1.1 mg / ml, 1.2 mg / ml, 1.3 mg / ml, 1.4 mg / ml, 1.5 mg / ml, 1.6 mg / ml, 1.7 mg / ml, 1.8 mg / ml, 1.9 mg / ml, 2.0 mg / ml or greater than 2.0 mg / ml.

[0149] In some embodiments, the pharmaceutical compositions disclosed herein are formulated as lipid nanoparticles (LNP). Accordingly, the present disclosure also provides nanoparticle compositions comprising (i) a lipid composition comprising a delivery agent such as a compound of Formula (I) or (III) as described herein, and (ii) a polynucleotide encoding an antigen polypeptide. In such nanoparticle composition, the lipid composition disclosed herein can encapsulate the polynucleotide encoding an antigen polypeptide.

[0150] Nanoparticle compositions are typically sized on the order of micrometers or smaller and can include a lipid bilayer. Nanoparticle compositions encompass lipid nanoparticles (LNPs), liposomes (e.g., lipid vesicles), and lipoplexes. For example, a nanoparticle composition can be a liposome having a lipid bilayer with a diameter of 500 nm or less.

[0151] Nanoparticle compositions include, for example, lipid nanoparticles (LNPs), liposomes, and lipoplexes. In some embodiments, nanoparticle compositions are vesicles including one or more lipid bilayers. In certain embodiments, a nanoparticle composition includes two or more concentric bilayers separated by aqueous compartments. Lipid bilayers can be functionalized and / or crosslinked to one another. Lipid bilayers can include one or more ligands, proteins, or channels.

[0152] In one embodiment, a lipid nanoparticle comprises an ionizable lipid, a structural lipid, a phospholipid, and mRNA. In some embodiments, the LNP comprises an ionizable lipid, a PEG-modified lipid, a phospholipid and a structural lipid. In some embodiments, the LNP has a molar ratio of about 20-60% ionizable lipid: about 5-25% phospholipid: about 25-55% structural lipid; and about 0.5-15% PEG-modified lipid. In some embodiments, the LNP comprises a molar ratio of about 50% ionizable lipid, about 1 .5% PEG-modified lipid, about 38.5% structural lipid and about 10% phospholipid. In some embodiments, the LNP comprises a molar ratio of about 55% ionizable lipid, about 2.5% PEG lipid, about 32.5% structural lipid and about 10% phospholipid. In some embodiments, the ionizable lipid is an ionizable amino lipid and the phospholipid is a neutral lipid, and the structural lipid is a cholesterol. In some embodiments, the LNP has a molar ratio of 50:38.5:10:1 .5 of ionizable lipid: cholesterol: DSPC: PEG lipid.

[0153] In some embodiments, the LNP has a polydispersity value of less than 0.4. In some embodiments, the LNP has a net neutral charge at a neutral pH. In some embodiments, the LNP has a mean diameter of 50-150 nm. In some embodiments, the LNP has a mean diameter of 80-100 nm.

[0154] As generally defined herein, the term “lipid” refers to a small molecule that has hydrophobic or amphiphilic properties. Lipids may be naturally occurring or synthetic. Examples of classes of lipids include, but are not limited to, fats, waxes, sterol-containing metabolites, vitamins, fatty acids, glycerolipids, glycerophospholipids, sphingolipids, saccharolipids, and polyketides, and prenol lipids. In some instances, the amphiphilic properties of some lipids lead them to form liposomes, vesicles, or membranes in aqueous media.

[0155] In some embodiments, a lipid nanoparticle (LNP) may comprise an ionizable lipid. As used herein, the term “ionizable lipid” has its ordinary meaning in the art and may refer to a lipid comprising one or more charged moieties. In some embodiments, an ionizable lipid may be positively charged or negatively charged. An ionizable lipid may be positively charged, in which case it can be referred to as “cationic lipid”. In certain embodiments, an ionizable lipid molecule may comprise an amine group, and can be referred to as an ionizable amino lipids. As used herein, a “charged moiety” is a chemical moiety that carries a formal electronic charge, e.g., monovalent (+1 , or -1 ), divalent (+2, or -2), trivalent (+3, or -3), etc. The chargedmoiety may be anionic (i.e., negatively charged) or cationic (i.e., positively charged). Examples of positively-charged moieties include amine groups (e.g., primary, secondary, and / or tertiary amines), ammonium groups, pyridinium group, guanidine groups, and imidizolium groups. In a particular embodiment, the charged moieties comprise amine groups. Examples of negatively-charged groups or precursors thereof, include carboxylate groups, sulfonate groups, sulfate groups, phosphonate groups, phosphate groups, hydroxyl groups, and the like. The charge of the charged moiety may vary, in some cases, with the environmental conditions, for example, changes in pH may alter the charge of the moiety, and / or cause the moiety to become charged or uncharged. In general, the charge density of the molecule may be selected as desired.

[0156] It should be understood that the terms “charged” or “charged moiety” does not refer to a “partial negative charge” or “partial positive charge” on a molecule. The terms “partial negative charge” and “partial positive charge” are given their ordinary meaning in the art. A “partial negative charge” may result when a functional group comprises a bond that becomes polarized such that electron density is pulled toward one atom of the bond, creating a partial negative charge on the atom. Those of ordinary skill in the art will, in general, recognize bonds that can become polarized in this way.

[0157] In some embodiments, the ionizable lipid is an ionizable amino lipid, sometimes referred to in the art as an “ionizable cationic lipid”. In one embodiment, the ionizable amino lipid may have a positively charged hydrophilic head and a hydrophobic tail that are connected via a linker structure.

[0158] In addition to these, an ionizable lipid may also be a lipid including a cyclic amine group. In one embodiment, the ionizable lipid may be selected from, but not limited to, an ionizable lipid described in International Publication Nos. WO2013 / 086354 and WO2013 / 1 16126; the contents of each of which are herein incorporated by reference in their entirety.

[0159] In one embodiment, the lipid may be a cleavable lipid such as those described in International Publication No. WO2012 / 170889, herein incorporated by reference in its entirety. In one embodiment, the lipid may be synthesized by methods known in the art and / or as described in International Publication Nos. WO2013086354; the contents of each of which are herein incorporated by reference in their entirety.

[0160] Nanoparticle compositions can be characterized by a variety of methods. For example, microscopy (e.g., transmission electron microscopy or scanning electron microscopy) can be used to examine the morphology and size distribution of a nanoparticle composition. Dynamic light scattering or potentiometry (e.g., potentiometric titrations) can be used to measure zeta potentials. Dynamic light scattering can also be utilized to determine particle sizes. Instruments such as the Zetasizer Nano ZS (Malvern Instruments Ltd, Malvern,Worcestershire, UK) can also be used to measure multiple characteristics of a nanoparticle composition, such as particle size, polydispersity index, and zeta potential.

[0161] Nanoparticle compositions can be characterized by a variety of methods. For example, microscopy (e.g., transmission electron microscopy or scanning electron microscopy) can be used to examine the morphology and size distribution of a nanoparticle composition. Dynamic light scattering or potentiometry (e.g., potentiometric titrations) can be used to measure zeta potentials. Dynamic light scattering can also be utilized to determine particle sizes. Instruments such as the Zetasizer Nano ZS (Malvern Instruments Ltd, Malvern, Worcestershire, UK) can also be used to measure multiple characteristics of a nanoparticle composition, such as particle size, polydispersity index, and zeta potential.

[0162] The size of the nanoparticles can help counter biological reactions such as, but not limited to, inflammation, or can increase the biological effect of the polynucleotide. As used herein, “size” or “mean size” in the context of nanoparticle compositions refers to the mean diameter of a nanoparticle composition. In one embodiment, the polynucleotide encoding an antigen polypeptide are formulated in lipid nanoparticles having a diameter from about 10 to about 100 nm such as, but not limited to, about 10 to about 20 nm, about 10 to about 30 nm, about 10 to about 40 nm, about 10 to about 50 nm, about 10 to about 60 nm, about 10 to about 70 nm, about 10 to about 80 nm, about 10 to about 90 nm, about 20 to about 30 nm, about 20 to about 40 nm, about 20 to about 50 nm, about 20 to about 60 nm, about 20 to about 70 nm, about 20 to about 80 nm, about 20 to about 90 nm, about 20 to about 100 nm, about 30 to about 40 nm, about 30 to about 50 nm, about 30 to about 60 nm, about 30 to about 70 nm, about 30 to about 80 nm, about 30 to about 90 nm, about 30 to about 100 nm, about 40 to about 50 nm, about 40 to about 60 nm, about 40 to about 70 nm, about 40 to about 80 nm, about 40 to about 90 nm, about 40 to about 100 nm, about 50 to about 60 nm, about 50 to about 70 nm, about 50 to about 80 nm, about 50 to about 90 nm, about 50 to about 100 nm, about 60 to about 70 nm, about 60 to about 80 nm, about 60 to about 90 nm, about 60 to about 100 nm, about 70 to about 80 nm, about 70 to about 90 nm, about 70 to about 100 nm, about 80 to about 90 nm, about 80 to about 100 nm and / or about 90 to about 100 nm.

[0163] In one embodiment, the nanoparticles have a diameter from about 10 to 500 nm. In one embodiment, the nanoparticle has a diameter greater than 100 nm, greater than 150 nm, greater than 200 nm, greater than 250 nm, greater than 300 nm, greater than 350 nm, greater than 400 nm, greater than 450 nm, greater than 500 nm, greater than 550 nm, greater than 600 nm, greater than 650 nm, greater than 700 nm, greater than 750 nm, greater than 800 nm, greater than 850 nm, greater than 900 nm, greater than 950 nm or greater than 1000 nm.

[0164] In some embodiments, the largest dimension of a nanoparticle composition is 1 urn or shorter (e.g., 1 pm, 900 nm, 800 nm, 700 nm, 600 nm, 500 nm, 400 nm, 300 nm, 200 nm, 175 nm, 150 nm, 125 nm, 100 nm, 75 nm, 50 nm, or shorter).

[0165] A nanoparticle composition can be relatively homogenous. A polydispersity index can be used to indicate the homogeneity of a nanoparticle composition, e.g., the particle size distribution of the nanoparticle composition. A small (e.g., less than 0.3) polydispersity index generally indicates a narrow particle size distribution. A nanoparticle composition can have a polydispersity index from about 0 to about 0.25, such as 0.01 , 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.10, 0.11 , 0.12, 0.13, 0.14, 0.15, 0.16, 0.17, 0.18, 0.19, 0.20, 0.21 , 0.22, 0.23, 0.24, or 0.25. In some embodiments, the polydispersity index of a nanoparticle composition disclosed herein can be from about 0.10 to about 0.20.

[0166] In addition to LNPs, the mRNA vaccine may be formulated in other carriers. Known carriers, for instance, may be modified or further formulated to achieve the threshold values described herein. Other carriers include but are not limited to other non-LNP lipid based carriers such as liposomes, lipoids and lipoplexes, particulate or polymeric nanoparticles, peptide carriers, nanoparticle mimics, nanotubes, conjugates, or emulsion delivery systems such as cationic submicron oil-in-water emulsions.

[0167] Liposomes are amphiphilic lipids which can form bilayers in an aqueous environment to encapsulate a RNA-containing aqueous core. These lipids can have an anionic, cationic or zwitterionic hydrophilic head group. Liposomes can be formed from a single lipid or from a mixture of lipids. A mixture may comprise (i) a mixture of anionic lipids (ii) a mixture of cationic lipids (iii) a mixture of zwitterionic lipids (iv) a mixture of anionic lipids and cationic lipids (v) a mixture of anionic lipids and zwitterionic lipids (vi) a mixture of zwitterionic lipids and cationic lipids or (vii) a mixture of anionic lipids, cationic lipids and zwitterionic lipids. Similarly, a mixture may comprise both saturated and unsaturated lipids. Exemplary phospholipids include, but are not limited to, phosphatidylethanolamines, phosphatidylcholines, phosphatidylserines, and phosphatidylglycerols. Cationic lipids include, but are not limited to, dioleoyl trimethylammonium propane (DOTAP), 1 ,2-distearyloxy-N,N-dimethyl-3- aminopropane (DSDMA), 1 ,2-dioleyloxy-N,Ndimethyl-3-aminopropane (DODMA), 1 ,2- dilinoleyloxy-N,N-dimethyl-3-aminopropane (DLinDMA), 1 ,2-dili nolenyloxy-N, N-dimethyl-3- aminopropane (DLenDMA). Zwitterionic lipids include, but are not limited to, acyl zwitterionic lipids and ether zwitterionic lipids. Examples of useful zwitterionic lipids are DPPC, DOPC and dodecylphosphocholine. The lipids can be saturated or unsaturated.

[0168] Polymeric microparticles or nanoparticles can also be used to encapsulate or adsorb mRNA. The particles may be substantially non-toxic and biodegradable. The particles useful for delivering mRNA may have an optimal size and zeta potential. For instance, the microparticles may have a diameter in the range of 0.02 pm to 8 pm. In the instances when the composition has a population of micro- or nanoparticles with different diameters, at least 80%, 85%, 90%, or 95% of those particles ideally have diameters in the range of 0.03-7 pm.The particles may also have a zeta potential of between 40-100 mV, in order to provide maximal adsorption of the mRNA to the particles.

[0169] Non-toxic and biodegradable polymers include, but are not limited to, poly(ahydroxy acids), polyhydroxy butyric acids, polylactones (including polycaprolactones), polydioxanones, polyvalerolactone, polyorthoesters, polyanhydrides, polycyanoacrylates, tyrosine-derived polycarbonates, polyvinyl-pyrrolidinones or polyester-amides, and combinations thereof. In some embodiments, the particles are formed from poly(ahydroxy acids), such as a poly(lactides) (“PLA”), copolymers of lactide and glycolide such as a poly(D,L-lactide-co-glycolide) (“PLG”), and copolymers of D,L-lactide and caprolactone. Useful PLG polymers include those having a lactide / glycolide molar ratio ranging, for example, from 20:80 to 80:20 e.g. 25:75, 40:60, 45:55, 55:45, 60:40, 75:25. Useful PLG polymers include those having a molecular weight between, for example, 5,000-200,000 Da e.g. between 10,000-100,000, 20,000-70,000, 40,000-50,000 Da.

[0170] Oil-in-water emulsions may also be used for delivering mRNA to a subject. Examples of oils useful for making the emulsions include animal (e.g., fish) oil or vegetable oil (e.g. nuts, seeds and grains). The oil may be biodegradable (metabolizable) and biocompatible. Some exemplary oils include tocopherols and squalene, a shark liver oil which is a branched, unsaturated terpenoid and combinations thereof. Terpenoids are branched chain oils that are synthesized biochemically in 5-carbon isoprene units.

[0171] The aqueous component of the emulsion can be water or can be water in which additional components have been added. For instance, it may include salts to form a buffer e.g. citrate or phosphate salts, such as sodium salts. Exemplary buffers include: a phosphate buffer; a Tris buffer; a borate buffer; a succinate buffer; a histidine buffer; or a citrate buffer.

[0172] The oil-in water emulsions ideally include one or more cationic molecules. For instance, a cationic lipid can be included in the emulsion to provide a positively charged droplet surface to which negatively-charged mRNA can attach. Useful cationic lipids include, but are not limited to: 1 ,2-dioleoyloxy-3-(trimethylammonio)propane (DOTAP), 3'-[N-(N',N'- Dimethylaminoethane)-carbamoyl]Cholesterol (DC Cholesterol), dimethyldioctadecylammonium (DDA e.g. the bromide), 1 ,2-Dimyristoyl-3-Trimethyl-AmmoniumPropane (DMTAP), dipalmitoyl(C16:0)trimethyl ammonium propane (DPTAP), distearoyltrimethylammonium propane (DSTAP). Other useful cationic lipids are: benzalkonium chloride (BAK), benzethonium chloride, cetramide (which contains tetradecyltrimethylammonium bromide and possibly small amounts of dedecyltrimethylammonium bromide and hex adecyltrimethyl ammonium bromide), cetylpyridinium chloride (CPC), cetyl trimethylammonium chloride (CTAC), N,N',N'- polyoxyethylene (10)-N-tallow-1 ,3-diaminopropane, dodecyltrimethylammonium bromide, hexadecyltrimethyl-ammonium bromide, mixed alkyl-trimethyl-ammonium bromide,benzyldimethyldodecylammonium chloride, benzyldimethylhexadecyl-ammonium chloride, benzyltrimethylammonium methoxide, cetyldimethylethylammonium bromide, dimethyldioctadecyl ammonium bromide (DDAB), methylbenzethonium chloride, decamethonium chloride, methyl mixed trialkyl ammonium chloride, methyl trioctylammonium chloride), N,N-dimethyl-N-[2 (2-methyl-4-(1 ,1 ,3,3tetramethylbutyl)-phenoxy]-ethoxy)ethyl]- benzenemetha-naminium chloride (DEBDA), dialkyldimethylammonium salts, [1 -(2,3- dioleyloxy)-propyl]-N,N,N, trimethylammonium chloride, 1 ,2-diacyl-3-(trimethylammonio) propane (acyl group=dimyristoyl, dipalmitoyl, distearoyl, dioleoyl), 1 ,2-diacyl-3 (dimethylammonio)propane (acyl group=dimyristoyl, dipalmitoyl, distearoyl, dioleoyl), 1 ,2- dioleoyl-3-(4'-trimethyl-ammonio)butanoyl-sn-glycerol, 1 ,2-dioleoyl 3-succinyl-sn-glycerol choline ester, cholesteryl (4'-trimethylammonio) butanoate), N-alkyl pyridinium salts (e.g. cetylpyridinium bromide and cetylpyridinium chloride), N-alkylpiperidinium salts, dicationic bolaform electrolytes (C12Me6; C12BU6), dialkylglycetylphosphorylcholine, lysolecithin, L- . alpha. dioleoylphosphatidylethanolamine, cholesterol hemisuccinate choline ester, lipopolyamines, including but not limited to dioctadecylamidoglycylspermine (DOGS), dipalmitoyl phosphatidylethanol-amidospermine (DPPES), lipopoly-L (or D)-lysine (LPLL, LPDL), poly(L (or D)-lysine conjugated to N-glutarylphosphatidylethanolamine, didodecyl glutamate ester with pendant amino group (C GluPhCnN), ditetradecyl glutamate ester with pendant amino group (C14GluCnN+), cationic derivatives of cholesterol, including but not limited to cholesteryl-3.beta.-oxysuccinamidoethylenetrimethylammonium salt, cholesteryl- 3.beta.-oxysuccinamidoethylene-dimethylamine, cholesteryl-3.beta.- carboxyamidoethylenetrimethylammonium salt, and cholesteryl-3.beta.- carboxyamidoethylenedimethylamine.

[0173] In addition to the oil and cationic lipid, an emulsion can include a non-ionic surfactant and / or a zwitterionic surfactant. Such surfactants include, but are not limited to: the polyoxyethylene sorbitan esters surfactants, especially polysorbate 20 and polysorbate 80; copolymers of ethylene oxide, propylene oxide, and / or butylene oxide, linear block copolymers; octoxynols; (octylphenoxy)polyethoxyethanol; phospholipids such as phosphatidylcholine; polyoxyethylene fatty ethers derived from lauryl, cetyl, stearyl and oleyl alcohols; polyoxyethylene-9-lauryl ether; and sorbitan esters.

[0174] The size of the emulsion particle may vary. In some embodiments the particles are in the range 20-750 nm, or for instance, 20-250 nm, 20-200 nm, 20-150 nm.

[0175] The vaccine compositions, polypeptides, and nucleic acids of the invention can also comprise additional polypeptides from other sources. For example, the compositions and fusion proteins of the invention can include polypeptides or nucleic acids encoding polypeptides, wherein the polypeptide enhances expression of the antigen, e.g., NS1 , an influenza virus protein (see, e.g. WO99 / 40188 and WO93 / 04175), etc. The nucleic acids ofthe invention can be engineered based on codon preference in a species of choice, e.g., humans (in the case of in vivo expression) or a particular bacterium (in the case of polypeptide production).POLYPEPTIDE COMPOSITIONS

[0176] The present invention, in other aspects, provides polypeptide compositions. Generally, a polypeptide of the invention will be an isolated Rv2140c polypeptide (i.e. separated from those components with which it may usually be found in nature). For example, a naturally- occurring protein is isolated if it is separated from some or all of the coexisting materials in the natural system. Preferably, such polypeptides are at least about 90% pure, more preferably at least about 95% pure and most preferably at least about 99% pure.

[0177] Polypeptides may be prepared using any of a variety of well known techniques. Recombinant polypeptides encoded by DNA sequences as described above may be readily prepared from the DNA sequences using any of a variety of expression vectors known to those of ordinary skill in the art. Expression may be achieved in any appropriate host cell that has been transformed or transfected with an expression vector containing a DNA molecule that encodes a recombinant polypeptide. Suitable host cells include prokaryotes, yeast, and higher eukaryotic cells, such as mammalian cells and plant cells. Preferably, the host cells employed are E. coli, yeast or a mammalian cell line such as COS or CHO. Supernatants from suitable host / vector systems which secrete recombinant protein or polypeptide into culture media may be first concentrated using a commercially available filter. Following concentration, the concentrate may be applied to a suitable purification matrix such as an affinity matrix or an ion exchange resin. Finally, one or more reverse phase HPLC steps can be employed to further purify a recombinant polypeptide.

[0178] Polypeptides of the invention, immunogenic fragments thereof, and other variants having less than about 100 amino acids, and generally less than about 50 amino acids, may also be generated by synthetic means, using techniques well known to those of ordinary skill in the art. For example, such polypeptides may be synthesized using any of the commercially available solid-phase techniques, such as the Merrifield solid-phase synthesis method, where amino acids are sequentially added to a growing amino acid chain. See Merrifield, J. Am. Chem. Soc. 85:2149-2146 (1963). Equipment for automated synthesis of polypeptides is commercially available from suppliers such as Perkin Elmer / Applied BioSystems Division (Foster City, CA), and may be operated according to the manufacturer's instructions.

[0179] Within certain specific embodiments, a polypeptide may be a fusion protein that comprises multiple polypeptides as described herein, or that comprises at least one polypeptide as described herein and an unrelated sequence, examples of such proteins include tetanus, tuberculosis and hepatitis proteins (see, e.g., Stoute et al., New Engl. J. Med.336:86-91 (1997)). A fusion partner may, for example, assist in providing T helper epitopes (an immunological fusion partner), preferably T helper epitopes recognized by humans, or may assist in expressing the protein (an expression enhancer) at higher yields than the native recombinant protein. Certain preferred fusion partners are both immunological and expression enhancing fusion partners. Other fusion partners may be selected so as to increase the solubility of the protein or to enable the protein to be targeted to desired intracellular compartments. Still further fusion partners include affinity tags, which facilitate purification of the protein.

[0180] Fusion proteins may generally be prepared using standard techniques, including chemical conjugation. Preferably, a fusion protein is expressed as a recombinant protein, allowing the production of increased levels, relative to a non-fused protein, in an expression system. Briefly, DNA sequences encoding the polypeptide components may be assembled separately, and ligated into an appropriate expression vector. The 3' end of the DNA sequence encoding one polypeptide component is ligated, with or without a peptide linker, to the 5' end of a DNA sequence encoding the second polypeptide component so that the reading frames of the sequences are in phase. This permits translation into a single fusion protein that retains the biological activity of both component polypeptides.

[0181] A peptide linker sequence may be employed to separate the first and second polypeptide components by a distance sufficient to ensure that each polypeptide folds into its secondary and tertiary structures. Such a peptide linker sequence is incorporated into the fusion protein using standard techniques well known in the art. Suitable peptide linker sequences may be chosen based on the following factors: (1 ) their ability to adopt a flexible extended conformation; (2) their inability to adopt a secondary structure that could interact with functional epitopes on the first and second polypeptides; and (3) the lack of hydrophobic or charged residues that might react with the polypeptide functional epitopes. Preferred peptide linker sequences contain Gly, Asn and Ser residues. Other near neutral amino acids, such as Thr and Ala may also be used in the linker sequence. Amino acid sequences which may be usefully employed as linkers include those disclosed in Maratea et al., Gene 40:39-46 (1985); Murphy et al., Proc. Natl. Acad. Scl. USA 83:8258-8262 (1986); U.S. Patent No. 4,935,233 and U.S. Patent No. 4,751 ,180. The linker sequence may generally be from 1 to about 50 amino acids in length. Linker sequences are not required when the first and second polypeptides have non-essential N-terminal amino acid regions that can be used to separate the functional domains and prevent steric interference.PHARMACEUTICAL COMPOSITIONS

[0182] In additional embodiments, the polynucleotide, polypeptide, and vaccine compositions disclosed herein will be formulated in pharmaceutically-acceptable or physiologically-acceptable solutions for administration to a cell or an animal, either alone, or in combination with one or more other modalities of therapy.

[0183] It will also be understood that, if desired, the nucleic acid segment (e.g., RNA or DNA) that expresses a polypeptide as disclosed herein may be administered in combination with other agents as well, such as, e.g., other proteins or polypeptides or various pharmaceutically- active agents, including chemotherapeutic agents effective against a M. tuberculosis infection. In fact, there is virtually no limit to other components that may also be included, given that the additional agents do not cause a significant adverse effect upon contact with the target cells or host tissues. The compositions may thus be delivered along with various other agents as required in the particular instance. Such compositions may be purified from host cells or other biological sources, or alternatively may be chemically synthesized as described herein. Likewise, such compositions may further comprise substituted or derivatized RNA or DNA compositions.

[0184] Formulation of pharmaceutically-acceptable excipients and carrier solutions is well- known to those of skill in the art, as is the development of suitable dosing and treatment regimens for using the particular compositions described herein in a variety of treatment regimens, including e.g., oral, parenteral, intravenous, intranasal, and intramuscular administration and formulation. Other routes of administration include via the mucosal surfaces.

[0185] Typically, formulations comprising a therapeutically effective amount deliver about 0.1 ug to about 1000 ug of polypeptide per administration, more typically about 2.5 ug to about 100 ug of polypeptide per administration. In respect of polynucleotide compositions, these typically deliver about 10 ug to about 20 mg of the inventive polynucleotide per administration, more typically about 0.1 mg to about 10 mg of the inventive polynucleotide per administration.

[0186] Naturally, the amount of active compound(s) in each therapeutically useful composition may be prepared is such a way that a suitable dosage will be obtained in any given unit dose of the compound. Factors such as solubility, bioavailability, biological half-life, route of administration, product shelf life, as well as other pharmacological considerations will be contemplated by one skilled in the art of preparing such pharmaceutical formulations, and as such, a variety of dosages and treatment regimens may be desirable.

[0187] In certain circumstances it will be desirable to deliver the pharmaceutical compositions disclosed herein orally, intranasally, parenterally, intravenously, intramuscularly, intradermally, or even intraperitoneally. Solutions of the active compounds as free base or pharmacologically acceptable salts may be prepared in water suitably mixed with a surfactant, such as hydroxypropylcellulose. Dispersions may also be prepared in glycerol, liquid polyethylene glycols, and mixtures thereof and in oils. Under ordinary conditions of storage and use, these preparations contain a preservative to prevent the growth of microorganisms.

[0188] The pharmaceutical forms suitable for injectable use include sterile aqueous solutions or dispersions and sterile powders for the extemporaneous preparation of sterile injectable solutions or dispersions (U. S. Patent 5,466,468, specifically incorporated herein by reference in its entirety). In all cases the form must be sterile and must be fluid to the extent that easy syringability exists. It must be stable under the conditions of manufacture and storage and must be preserved against the contaminating action of microorganisms, such as bacteria and fungi. The carrier can be a solvent or dispersion medium containing, for example, water, ethanol, polyol (e.g., glycerol, propylene glycol, and liquid polyethylene glycol, and the like), suitable mixtures thereof, and / or vegetable oils. Proper fluidity may be maintained, for example, by the use of a coating, such as lecithin, by the maintenance of the required particle size in the case of dispersion and by the use of surfactants. The prevention of the action of microorganisms can be facilitated by various antibacterial and antifungal agents, for example, parabens, chlorobutanol, phenol, sorbic acid, thimerosal, and the like. In many cases, it will be preferable to include isotonic agents, for example, sugars or sodium chloride. Prolonged absorption of the injectable compositions can be brought about by the use in the compositions of agents delaying absorption, for example, aluminum monostearate and gelatin.

[0189] For parenteral administration in an aqueous solution, for example, the solution should be suitably buffered if necessary and the liquid diluent first rendered isotonic with sufficient saline or glucose. These particular aqueous solutions are especially suitable for intravenous, intramuscular, subcutaneous and intraperitoneal administration. In this connection, a sterile aqueous medium that can be employed will be known to those of skill in the art in light of the present disclosure. For example, one dosage may be dissolved in 1 ml of isotonic NaCI solution and either added to 1000 ml of hypodermoclysis fluid or injected at the proposed site of infusion (see, e.g., Remington's Pharmaceutical Sciences, 15th Edition, pp. 1035-1038 and 1570-1580). Some variation in dosage will necessarily occur depending on the condition of the subject being treated. The person responsible for administration will, in any event, determine the appropriate dose for the individual subject. Moreover, for human administration, preparations should meet sterility, pyrogenicity, and the general safety and purity standards as required by FDA Office of Biologies standards.

[0190] Sterile injectable solutions are prepared by incorporating the active compounds in the required amount in the appropriate solvent with various of the other ingredients enumerated above, as required, followed by filtered sterilization. Generally, dispersions are prepared by incorporating the various sterilized active ingredients into a sterile vehicle which contains the basic dispersion medium and the required other ingredients from those enumerated above. In the case of sterile powders for the preparation of sterile injectable solutions, the preferred methods of preparation are vacuum-drying and freeze-drying techniques which yield a powderof the active ingredient plus any additional desired ingredient from a previously sterile-fi Itered solution thereof.

[0191] The compositions disclosed herein may be formulated in a neutral or salt form. Pharmaceutically-acceptable salts, include the acid addition salts (formed with the free amino groups of the protein) and which are formed with inorganic acids such as, for example, hydrochloric or phosphoric acids, or such organic acids as acetic, oxalic, tartaric, mandelic, and the like. Salts formed with the free carboxyl groups can also be derived from inorganic bases such as, for example, sodium, potassium, ammonium, calcium, or ferric hydroxides, and such organic bases as isopropylamine, trimethylamine, histidine, procaine and the like. Upon formulation, solutions will be administered in a manner compatible with the dosage formulation and in such amount as is therapeutically effective. The formulations are easily administered in a variety of dosage forms such as injectable solutions, drug-release capsules, and the like.

[0192] As used herein, "carrier" includes any and all solvents, dispersion media, vehicles, coatings, diluents, antibacterial and antifungal agents, isotonic and absorption delaying agents, buffers, carrier solutions, suspensions, colloids, and the like. The use of such media and agents for pharmaceutical active substances is well known in the art. Except insofar as any conventional media or agent is incompatible with the active ingredient, its use in the therapeutic compositions is contemplated. Supplementary active ingredients can also be incorporated into the compositions.

[0193] The phrase "pharmaceutically-acceptable" refers to molecular entities and compositions that do not produce an allergic or similar untoward reaction when administered to a human. The preparation of an aqueous composition that contains a protein as an active ingredient is well understood in the art. Typically, such compositions are prepared as injectables, either as liquid solutions or suspensions; solid forms suitable for solution in, or suspension in, liquid prior to injection can also be prepared. The preparation can also be emulsified.

[0194] In some embodiments an adjuvant composition is provided. Exemplary adjuvants are oild in water emulsions, and may comprise squalene in the oil phase. For example, AS03 is an adjuvant system composed of a-tocopherol, squalene and polysorbate 80 in an oil-in-water emulsion. MF59 is another immunologic adjuvant that comprises a squalene emulsion. The dose of adjuvant administered may depend on whether an antigen is present, on the antigen with which it is used and the antigen dosage to be applied. It is also dependent on the intended species and the desired formulation. Usually the quantity is within the range conventionally used for adjuvants. For example, adjuvants typically comprises from about 1 ,ug to about 1000 pig, inclusive, of a 1 -mL dose.Methods of vaccination

[0195] A vaccine comprising an Rv2024c polypeptide or polypeptide of the disclosure is administered to an individual in a dose effective to generate a protective immune response, i.e. a response where the individual does not suffer from a latent or active Mtb infection following exposure. In some embodiments the individual has not been previously vaccinated or exposed to Mtb. In some embodiments the individual has been exposed or vaccinated. In certain embodiments, multiple therapeutically effective doses are administered according to a defined dosing regimen, or intermittently. For example, a therapeutically effective dose can be administered once, twice, three times or more, where the administrations may be separated from 1 week, two weeks, 3 weeks, 4 weeks, 5 weeks, 6 weeks, or more. By "intermittent" administration is intended the therapeutically effective dose can be administered, for example, once every two weeks, once every three weeks, once a month, and so forth. In accordance with the methods of the present invention, a subject can receive intermittent therapy for one or more weekly or monthly cycles until the desired immune response is achieved. The agents can be administered by any acceptable route of administration as noted herein below.

[0196] The Rv2140c component may also be administered with one or more chemotherapeutic agents effective against tuberculosis (e.g. M. tuberculosis infection). Examples of such chemotherapeutic agents include, but are not limited to, amikacin, aminosalicylic acid, capreomycin, cycloserine, ethambutol, ethionamide, isoniazid, kanamycin, pyrazinamide, rifamycins (i.e., rifampin, rifapentine and rifabutin), streptomycin, ofloxacin, ciprofloxacin, clarithromycin, azithromycin and fluoroquinolones. Such chemotherapy is determined by the judgment of the treating physician using preferred drug combinations. "First-line" chemotherapeutic agents used to treat tuberculosis (e.g. M. tuberculosis infection) that is not drug resistant include isoniazid, rifampin, ethambutol, streptomycin and pyrazinamide. "Second-line" chemotherapeutic agents used to treat tuberculosis (e.g. M. tuberculosis infection) that has demonstrated drug resistance to one or more "first-line" drugs include ofloxacin, ciprofloxacin, ethionamide, aminosalicylic acid, cycloserine, amikacin, kanamycin and capreomycin. Conventional chemotherapeutic agents are generally administered over a relatively long period (ca. 9 months). Combination of conventional chemotherapeutic agents with the administration of a Rv2140c component according to the present invention may enable the chemotherapeutic treatment period to be reduced (e.g. to 8 months, 7 months, 6 months, 5 months, 4 months, 3 months or less) without a decrease in efficacy.

[0197] Of particular interest is the use of an Rv2140c component in conjunction with Bacillus Calmette-Guerin (BCG). For example, in the form of a modified BCG which recombinantly expresses Rv2140c (or a variant or fragment thereof as described herein). Alternatively, theRv2140c component may be used to enhance the response of a subject to BCG vaccination, either by co-administration or by boosting a previous BCG vaccination. When used to enhance the response of a subject to BCG vaccination, the Rv2140c component may obviously be provided in the form of a polypeptide or a polynucleotide (optionally in conjunction with additional antigenic components as described above).

[0198] The skilled person will recognize that combinations of components need not be administered together and may be applied: separately or in combination; at the same time, sequentially or within a short period; though the same or through different routes. Nevertheless, for convenience it is generally desirable (where administration regimes are compatible) to administer a combination of components as a single composition.

[0199] The polypeptides, polynucleotides and compositions of the present invention will usually be administered to humans, but are effective in other mammals including domestic mammals (e.g., dogs, cats, rabbits, rats, mice, guinea pigs, hamsters, chinchillas) and agricultural mammals (e.g., cows, pigs, sheep, goats, horses).DIAGNOSTICS

[0200] In another aspect, this invention provides methods for using one or more of the polypeptides described above to diagnose tuberculosis (for example using T cell response based assays or antibody based assays of conventional format). In other embodiments an individual is assessed for the presence of a protective immune response following vaccination, for example to determine the presence of T follicular helper cells specific for an Rv2140c. The release of TGF[3 from the Tfh may be determined, where such a release is indicative of a protective response.

[0201] Tfh may be characterized as CD4+CXCR5+T cells, which may also express markers such as CD25, CD69, CD95, CD57, 0X40 (CD134) and CD40L (CD154) and induce overexpression of activation-induced cytidine deaminase in B cells. A Tfh response to Rv2140c, e.g. SEQ ID NO:1 polypeptide, may include release of TGF[3.

[0202] Diagnosis or validation may include a method comprising: (a) obtaining a sample from the individual; (b) contacting said sample with an isolated polypeptide which comprises: (i) an Rv2140c protein sequence; (ii) a variant of an Rv2140c protein sequence; or (iii) an immunogenic fragment of an Rv2140c protein sequence; (c) quantifying the sample response.

[0203] The sample may for example be whole blood or purified cells. Suitably the sample will contain peripheral blood mononucleated cells (PBMC). In one embodiment of the invention the individual will be seropositive. In a second embodiment of the invention the individual will be seronegative.

[0204] The sample response may be quantified by a range of means known to those skilled in the art, including the monitoring of lymphocyte proliferation or the production of specific cytokines or antibodies. For example, T-cell ELISPOT may be used to monitor cytokines such as interferon gamma (IFNy), interleukin 2 (IL2) and interleukin 5 (IL5) . B-cell ELISPOT may be used to monitor the stimulation of M. tuberculosis specific antigens. The cellular response may also be characterised by the use of by intra- and extra-cellular staining and analysis by a flow cy to meter.

[0205] The present invention further provides kits for use within any of the above diagnostic methods. Such kits typically comprise two or more components necessary for performing a diagnostic assay. Components may be compounds, reagents, containers and / or equipment.

[0206] For example, one container within a kit may contain a monoclonal antibody or fragment thereof that specifically binds to a protein. Such antibodies or fragments may be provided attached to a support material, as described above. One or more additional containers may enclose elements, such as reagents or buffers, to be used in the assay. Such kits may also, or alternatively, contain a detection reagent as described above that contains a reporter group suitable for direct or indirect detection of antibody binding.

[0207] Alternatively, a kit may be designed to detect the level of mRNA encoding a protein in a biological sample. Such kits generally comprise at least one oligonucleotide probe or primer, as described above, that hybridises to a polynucleotide encoding a protein. S uch an oligonucleotide may be used, for example, within a PCR or hybridization assay. Additional components that may be present within such kits include a second oligonucleotide and / or a diagnostic reagent or container to facilitate the detection of a polynucleotide encoding a protein of the invention.

[0208] Other diagnostics kits include those designed for the detection of cell mediated responses (which may, for example, be of use in the diagnostic methods of the present invention). Such kits will typically comprise: (i) apparatus for obtaining an appropriate cell sample from a subject; (ii) means for stimulating said cell sample with an Rv2140c polypeptide (or variant thereof, immunogenic fragments thereof, or DNA encoding such polypeptides); (iii) means for detecting or quantifying the cellular response to stimulation. Suitable means for quantifying the cellular response include a B-cell ELISPOT kit or alternatively a T-cell ELISPOT kit, which are known to those skilled in the art.EXAMPLES

[0209] The following examples are provided by way of illustration only and not by way of limitation. Those of skill in the art will readily recognize a variety of noncritical parameters that could be changed or modified to yield essentially similar results.Example 1

[0210] Mycobacterium tuberculosis (Mtb) has infected billions of people worldwide, and it is a major cause of mortality and sickness, especially in Africa and Asia. The standard diagnostic for Mtb infection has been the QuantiFERON test, which measures the y-IFN response to Mtb-specific antigens. But recent results with subjects exposed to Mtb (“resisters”) have shown that the QuantiFERON test is not an accurate gage of exposure. Here, we performed a TCR repertoire analysis of Mtb-reactive T cells comparing resisters, who are QuantiFERON negative (IGRA-), with QuantiFERON positive (IGRA+) subjects using the GLIPH3 algorithm and identified a number of TCR clusters that are unique to the resisters. We identified two antigens recognized by these TCRs and made peptide-MHC multimers to detect antigen-specific CD4+ T cells in circulating lymphocytes.

[0211] Analyzing IGRA- individuals from an Mtb adolescent cohort, we found that all had detectable T cell responses to both the Mtb specific reagent (ESAT6), showing that they had been infected with Mtb, and a very robust response to a conserved mycobacterial antigen, Rv2140c. Remarkably, the T cell response to this later antigen consisted largely of the follicular helper phenotype, which has correlated with protection from Mtb in mouse models, and is not characteristic of IGRA+ individuals where a Th1* phenotype is ubiquitous.

[0212] Since the Rv2140c gene is highly conserved between mycobacterial strains, including the Bacille Calmette-Guerin (BCG) vaccine, which is given at birth and has been highly effective in preventing TB disease in children, we also surveyed infants two months post BCG vaccination and found that at least 4 of 10 had detectable Rv2140c CD4+ T cells and that these were largely of the Tfh-stem phenotype. This suggests that the response to this antigen is robust and contributes to protection, but since many individuals in endemic areas convert from IGRA- to IGRA+ over time, these results suggest that repeated exposure to Mtb leads to a major change in the T cell response and attenuated protection.

[0213] In this study, we aimed to characterize the antigen-specific CD4+ T cell responses in Mtb-exposed IGRA- individuals. Leveraging a well-established longitudinal Mtb cohort study in Uganda, we identified a group of RSTR individuals who experienced high levels of Mtb exposure through close contact with household members suffering from active pulmonary tuberculosis yet remained consistently negative for both the QuantiFERON and TST tests over an average follow-up period of 9.5 years. To investigate the Mtb-specific T cell receptor (TCR) repertoire in RSTR individuals, we performed single-cell transcriptional analysis to collect paired alpha and beta TCR chains from Mtb-reactive CD4+ T cells in both RSTR and IGRA+ LTBI individuals from the Ugandan household contact cohort. We then analyzed the collected TCR sequences using GLIPH3 — the most advanced version of the GLIPH algorithm — designed to cluster homologous TCRs based on antigen specificity. This analysis revealed 24 TCR specificity clusters that were enriched in RSTR individuals but notpresent in those with LTBI, suggesting distinct antigen-specific CD4+ T cell responses in RSTR individuals. To further investigate these responses, we employed a novel T cell antigen discovery platform described in Huang et al. to screen the whole Mtb genome for antigens recognized by RSTR-specific TCRs, and two novel peptide antigens derived from the Rv2140c and ESAT6 proteins were identified, respectively. These two antigens enabled us to delve deeper into the antigen-specific CD4+ T cell responses in RSTR individuals in greater detail by developing and applying peptide-MHC multimer reagents with the Rv2140c and ESAT6 peptides.

[0214] Thus, we analyzed individuals’ PBMCs from the independent and well-studied Adolescent Cohort Study (ACS) in Cape Town and found that essentially all of the IGRA- subjects analyzed had been exposed to Mtb as judged by ESAT6 / DRB1 :11 :01 staining, which accounted for approximately 0.1% of circulating CD4+ T cells. Notably, we observed that Rv2140c-specific CD4+ T cells were present at nearly 1% of total CD4+ T cells in circulation, which is a remarkably high frequency in the absence of any obvious stimulation, suggesting the robust response to Rv2140c in IGRA- individuals in this Mtb endemic region. These T cells had a distinctive follicular helper T cell (Tfh) phenotype, which has been linked to the control of Mtb infection in several murine studies.

[0215] In contrast, ESAT6-specific CD4+ T cells in ACS IGRA+ individuals exhibited classic Th1 / 2 dominance, consistent with previous studies. Given the high conservation of the Rv2140c gene across various mycobacteria, including the BCG vaccine strain, we explored the origin of the Rv2140c-specific response. In a cohort with 8-week-old infants who had received the BCG vaccine at birth and had no known mycobacteria exposure, we detected Rv2140c-specific CD4+ T cell responses in 4 out of 10 infants — all displaying a Tfh-stem- like phenotype. This suggests that the initial T cell priming against Rv2140c is likely induced by the BCG vaccine, generating a Tfh response. Repeated Mtb exposure later in life might then diminish and alter both the phenotype and frequency of the response.Results

[0216] RSTR CD4+ T cells exhibited a lower activation level in response to Mtb lysate compared to those from LTBI individuals. \Ne first identified and isolated Mtb-reactive CD4+ T cells from RSTR and LTBI individuals within the Ugandan household contact cohort. PBMC samples from RSTR (N=19) and LTBI (N=17) individuals were stimulated with Mtb lysate for 8 hours, after which the Mtb lysate-reactive CD4+ T cells were identified and isolated based on the upregulation of activation markers CD69, CD137, and / or CD154, as previously described (Fig. 1b). Both RSTR and LTBI individuals demonstrated a remarkable activation level. However, the frequency of the Mtb lysate-reactive CD4+ T cells among total circulating CD4+ T cells in RSTR individuals averaged 1.0%, which was significantly lowerthan the 2.4% observed in LTBI individuals (Figs. 1b and 1d). Given that ESAT6 and CFP10 proteins are specific to Mtb, we further assessed and isolated the ESAT6 / CFP10-reactive CD4+ T cells from Ugandan RSTR (N=3) and LTBI (N=4) individuals using the same approach. Overall, the magnitude of activated CD4+ T cells in response to the ESAT6 / CFP10 peptide pool stimulation was lower than that observed with Mtb lysate stimulation across both groups, with mean frequency of 0.4% and 0.59% in RSTR and LTBI individuals, respectively (Figs. 1c and 1e). Although not statistically significant, the frequency of ESAT6 / CFP10-reactive CD4+ T cells was slightly higher in LTBI compared to RSTR individuals, consistent with findings from a recently published RSTR study using the same cohort (Fig. 1e).

[0217] TCR specificity clusters associated with the “resist” of infection following Mtb exposure. Single-cell TCR sequencing has proven to be a powerful tool for collecting full- length paired alpha and beta TCR sequences, which are essential for identifying TCR specificity clusters and determining their antigen specificities. Using the BD Rhapsody transcriptional profiling platform, we obtained 12,018 paired alpha and beta TCRs from isolated Mtb-reactive CD4+ T cells in RSTR and LTBI individuals at the single-cell level. Both RSTR and LTBI individuals exhibited a substantial number of unique Mtb-specific TCR sequences (Fig. 2a). After normalizing for the total number of isolated cells, we observed similarly diverse Mtb-specific TCR repertoires in both LTBI and RSTR individuals, with an average of 178 unique CDR3p sequences per individual in LTBIs, compared to with 162 in RSTRs (Fig. 2a). Each RSTR individual demonstrated significant clonal expansion, with a considerable number of CDR3P sequences observed multiple times, ranging from 2 to 31 occurrences (Fig. 2b). Similarly, all but one LTBI individuals showed repeated CDR3|3 sequences, ranging from 2 to 39 occurrences (Fig. 2b).

[0218] Previously, we reported the GLIPH TCR analysis algorithms that can process millions of TCR sequences and group them based on their likely antigen specificities, even when derived from genetically distinct individuals. GLIPH and GLIPH2 have successfully identified hundreds of TCR specificity clusters across large and noisy TCR datasets from diverse cohort studies. However, both GLIPH and GLIPH2 rely solely on signatures within the CDR3p region. Prior studies demonstrated that clustering based on either alpha or beta TCR chains alone can yield similar specificity groups but may overlook critical features due to the heterodimeric nature of TCRs, which rely on both chains for peptide-MHC binding. To further refine the algorithm to focus on complete ab TCR heterodimers, from single T cells, we recently developed GLIPH3. GLIPH3, which also identifies discontinuous motifs, improving the recognition of conserved TCR signatures. Using this algorithm, we successfully identified 876 and 852 Mtb-specific TCR similarity clusters from RSTR and LTBI individuals,respectively. From these, we further identified 24 TCR similarity clusters that are uniquely enriched in RSTR but absent in LTBI individuals (Figs. 2c). As shown by the representative TCR sequences, each identified cluster displayed a unique conserved TCR motif, which were enriched for particular HLA class II alleles, which provided clues as to the likely restricting element.

[0219] Novel antigen discovery of TCRs associated with the “resist” of infection following Mtb exposure. To investigate the peptide antigens recognized by the TCRs derived from the GLIPH3-identified specificity clusters associated with the “resist” of infection following Mtb exposure, we employed an optimized antigen discovery system. Previously, we reported a T cell antigen discovery method by co-culturing artificial antigen-presenting cells (aAPC) and TCR-deficient reporter Jurkat T cells (J76) that can stably express candidate HLA and TCR, respectively. While this system enabled the identification of several T cell antigens, including the Mtb epitopes, its sensitivity was limited, stemming in part from the J76 is not compatible with T cell co-receptor expression, especially the CD4 molecule, likely caused by the ENU mutagenesis that was used to effect the native TCR ablation.

[0220] Here, we used a CRISPR-engineered TCR-deficient reporter Jurkat T cell line we recently developed. These lines stably express either CD4 or CD8 molecules together with an introduced TCR oc[ heterodimer, and provide significantly improved sensitivity and accuracy via a robust luciferase and GFP reporter constructs (Fig. 3a). We introduced representative TCR afS constructs using our modular TCR expression constructs from the 24 RSTR-specific TCR specificity clusters, inserted into the JFG4 line to create a panel of reporter T cells. These reporter T cells were then co-cultured with aAPCs expressing the corresponding HLAs enriched within each TCR specificity cluster. Prior to co-culture, aAPCs were pulsed with Mtb lysate or the ESAT6 / CFP10 peptide pool, following previously established protocols.

[0221] Among the 24 TCR candidates tested, TCR-1 showed a 14.9-fold induction in luciferase activity following stimulation with Mtb lysate, while TCR-10 exhibited a 9.6-fold luciferase induction in response to ESAT / CFP10 peptide pool stimulation (Fig. 3b). Both TCR- 1 and TCR-10 were restricted to HLA-DRB1 *1 1 :01 allele, as predicted by GLIPH3 analysis, and demonstrated no detectable cross-reactivity with other HLA alleles. To identify the specific Mtb antigen targeted by TCR-1 / DR1 1 , we utilized a comprehensive Mtb protein library produced via a cell-free in vitro expression system covering 95% of annotated Mtb open reading frames (ORFs). Briefly, this library included 3,294 ORFs from Mtb strain H37Rv and 430 from strain CDC1551 , obtained from BEI Resources. The ORFs were grouped into pools of 12 clones, yielding 321 subpools distributed across four 96-well plates. Screening this library, we found that TCR-1 / DR1 1 responded to the Mtb protein Rv2140c (Fig. 3d). To define the precise epitope, we employed epitope prediction tools, IEDB prediction tool andNetMHCHpan, to identify, synthesize, and screen possible antigen peptides of TCR-1 / DR11 within the Rv2140c protein.

[0222] As a result, we found peptide Rv2140c (amino acid 5-10, SEQ ID NO:1 PDPYAALPKLPSFSL) to be the peptide antigen for TCR-1 / DR11 (Fig. 3e). This was further validated by the dose-dependent luciferase response (Fig. 3f). Notably, Rv2140c has not previously been reported to be target of any T cell or antibody responses previously, and is highly conserved among Mycobacterial species. Rv2140c is known to modulate host inflammatory responses through interference with key microphage signaling pathways such as ERK and NF-KB. Given that TCR-10 / / DR1 1 responded to the ESAT6 / CFP10 peptide pool, we screened them with each individual overlapping peptide spanning ESAT6 and CFP10 proteins, and ESAT6 (amino acid 25-39, SEQ ID NO:5, ‘IHSLLDEGKQSLTKL’) was identified to be the peptide antigen (Fig. 3g). Similarly, the ESAT6(2539) peptide stimulation elicited a strong T cell response, with luciferase induction reaching 56.2-fold, even at a peptide concentration as low as 1 x10-spg / mL, further confirming specific recognition of the ESAT6<25- 39 DRB1 1 by TCR-10 (Fig. 3h).

[0223] As ESAT6 is an Mtb-specific protein and is absent from the BCG vaccine, this provides additional evidence that the RSTR individuals have been infected with Mtb, consistent with previous antibody and T cell analyses. Interestingly, although the specific ESAT6(2539) peptide has not been previously reported, several published HLA-DRB1 *11 :01 restricted epitopes overlap with it, consistent with the well-established immunodominance and virulence role of ESAT6 in Mtb pathogenesis.

[0224] Evidence that IFGA- subjects are infected with Mtb and Rv2140c is one of the major Mtb antigens recognized by their CD4+ T cell responses. The identification of these two RSTR- specific peptide antigens allowed us to investigate the antigen-specific CD4+ T cell responses in greater detail. Previous RSTR studies have shown that IGRA- individuals may, in fact, be infected with Mtb, challenging the accuracy of the QuantiFERON assay as a definitive measure of Mtb exposure. To address this, we generated peptide-MHC multimer reagents capable of detecting and isolating circulating antigen-specific CD4+ T cells. Two multimer reagents were produced for Rv2140c<5 -19 DRB11 and ESAT6(2539) / DRB1 1 , respectively, using the spheromer platform, which displays 12 peptide-MHC complexes on engineered ferritin nanoparticle to enhance sensitivity. We used these reagents to analyze an independent cohort, the Adolescent Cohort Study (ACS) in Cape Town, South Africa, which is a well- characterized longitudinal study of high school students living in an Mtb endemic region. In this study, 55% were IGRA+ and 39% were IGRA-, which have historically been considered uninfected. We first stained the PBMC from both IGRA+ and IGRA- ACS subjects with the two spheromer reagents. Strikingly, Rv2140C(s 19 DR1 1 -specific CD4+ T cells were detectedat relatively high frequencies, ranging from 0.77% to 1 .03% of total circulating CD4+ T cells in all the five IGRA- ACS subjects we tested, even in the absence of any obvious stimulation (Figs 4a and 4b, upper panel). In comparison, the Rv2140C(5 ig; / DR1 1 -specific CD4+ T cells were detectable in the five IGRA+ ACS subjects but at significantly lower frequencies, averaging only 0.14% of total CD4+ T cells.

[0225] To further validate Mtb infection in IGRA- ACS subjects, we analyzed ESAT6(2539) / DR1 1 -specific CD4+ T cell responses, which were present at an average frequency of 0.1 1 % among the total circulating CD4+ T cells (Figs 4a and 4b, lower panel), whereas IGRA+ subjects exhibited more robust responses, with frequencies ranging from 0.72% to 1 .21 %. Taken together, these findings show that many Mtb-exposed IGRA- population previously thought to be uninfected have, in fact, been exposed to Mtb. Furthermore, since ~1% of CD4+ T cells in circulation for a single epitope is a very robust response, this result for Rv2140c suggests that it is a key antigen driving CD4+ T cell responses in this population. One reason for this very strong response may be the fact that Rv2140c is highly conserved among Mycobacterial species generally, and the particular peptide used here is identical in BCG.

[0226] We next used a multi-parameter flow cytometry panel to characterize the phenotype profiles of these specific CD4+ T cells in IGRA- and IGRA+ ACS individuals. Surprisingly, pattern within the Rv2140C(5-i9> / DR1 1 -specific CD4+ T cells in IGRA- ACS individuals, we observed a striking enrichment of T follicular helper (Tfh) cells, which comprised between 22.9% to 43.4% of the antigen-specific population (Figs 4c and 4e). In contrast, ESAT6<25 39) / DR1 1 -specific CD4+ T cells from IGRA+ ACS individuals exhibited Th2 and Th17 dominance, consistent with the previous studies in subjects with latent Mtb infection (Figs 4d and 4e). While not statistically significant due to sample variation, the frequency of Th1 cells was higher on average in the ESAT6(2539) / DR1 1 -specific CD4+ T cell population in IGRA+ ACS subjects (Fig 4e). Additionally, both antigen-specific populations contained a substantial fraction of cells expressing CD45RA, indicative of a nai've-like phenotype. These cells are likely representative of T memory stem cells (TSCM), a long-lasting population has been reported previously to contribute to the durable immunity against pathogens. Interestingly, Th17 cells were not very prominent among the Rv2140c<5 19 DR1 1 -specific CD4+ T cells in IGRA- ACS subjects, which may reflect limitations in surface markers-based identification of Th17 cells (Figs 4c and 4e).

[0227] To further define the functional profile of the antigen-specific CD4+ T cell populations, we stimulated the PBMCs from both IGRA- and IGRA+ ACS individuals with Mtb lysate for 8h and assessed intracellular cytokine expression. In IGRA- ACS subjects, a notable 21 .32% of Rv214OC(5-19) / DR1 1 -specific CD4+ T cells expressed TGF|3, compared to 1 1 .35%in the ESAT6(25-39) / DR11 -specific CD4+ T cells from IGRA+ ACS subjects (Fig 4f). IL17A expression was observed in both groups, with an average of 4.32% and 9.37% in IGRA- and IGRA+ subjects, respectively (Fig 4f). Furthermore, the ESAT6(2539) / DR11 -specific CD4+ T cells in IGRA+ subjects expressed a broader range of cytokines, including IFNy, TNFa, IL- 4, and IL-13, in line with their diverse Th1 , Th2, and Th17 phenotypes (Fig 4f). These phenotypic and functional differences underscore distinct immune programming in IGRA- and IGRA+ individuals and indicate that antigen-specific Tfh responses might be one of the hallmarks of the IGRA- individuals, a phenotype which correlates with protection against Mtb in mouse models.

[0228] RV2140C(519) / DR11 -specific CD4+ T cells are largely T follicular helper cells (Tfh) and also exhibited a T memory stem cell (TSCM) phenotype in IGRA- subjects. To further define the phenotypic and functional landscape of the antigen-specific CD4+ T cells in IGRA- and IGRA+ subjects, we applied the 10xGenomics single-cell multi-omic platform with the feature barcode technology, which enabled simultaneous profiling of transcriptomes and dozens of cell surface protein expression in the same cell. By detecting the surface memory markers of the bulk CD4+ T cells of the IGRA- ACS subjects (N=10), we identified a large fraction of the na’ive- like population defined by CD45RA+, CD45RO-, CCR7+ (Figs 5a and 5b). Interestingly, aside from these naive markers, a considerable number also expressed CD95, 1 L-2Rp, and CD11 a, indicative of TSCM rather than naive T cells (TNaiVe) (Figs 5a and 5b). Despite a portion of the TSCM being mixed with the TNaivedue to their similarity on transcriptional and surface protein levels, we identified an independent TSCM cluster, with further evidence of functional differentiation towards a Tfh phenotype (Figs 5a and 5c). Notably, in addition to the central memory T cells (TCM) that were defined by the CD62L expression and the effector T cells (TE ), we identified a small subset of terminally differentiated effector T cells (TTE) marked by CD57+, CD27-, CD28-, IL-7Ra- (Figs 5a and 5b).

[0229] Following the same phenotypic definitions, TNai e, TSCM, TCM, TE , and TTE populations were also identified in the bulk CD4+ T cells from IGRA+ ACS subjects (N=5) (Figs 7a and 7b). By further characterizing the expression of the transcriptional factors, surface markers, and secreted cytokines, we found that the distinct TSCM population of the bulk CD4+ T cells exhibited the features of Tfh in the IGRA- ACS subjects (Fig 5c, clusters 6 and 7), demonstrating the high expression of BCL6 or CD185 along with IL-6, IL-21 , IL-10, and IL- 12A (Fig 5d). A prominent Th17 subset was also identified, characterized by the expression of RORyt, IL-17A, or CD161 (Fig 5c, clusters 4 and 5), and co-expressing CCR4, CCR6, and IL-26 (Fig 5d). Aside from Th17, the TE and TTE compartments contained Th1 , Th2, and Th22 subsets (Fig 5c, clusters 2 and 3). A distinct regulatory T cell (Treg) population was also detected (cluster 8), expressing the typical markers including FOXP3, GITR, and CD25 (Figs5c and 5d). Similarly, the bulk CD4+ T cells of the IGRA+ ACS subjects were defined with memory subsets of Th1 , Th17, Th2, Th22, Tfh, and Treg (Figs 7c and 7d).

[0230] The antigen-specific CD4+ T cells were labeled using dual fluorescent and oligonucleotide-conjugated peptide-MHC multimer reagents and enriched with magnetic microbeads following standard protocols. The Rv2140C(5-19 DR1 1 -specific CD4+ T cells were pinpointed within the bulk CD4+ T cells in IGRA- ACS subjects at a frequency of 5.5%, while the ESAT6(25 -39) / DR11 -specific CD4+ T cells represented 3.5% among the bulk CD4+ T cells in IGRA+ ACS subjects (Fig 5e and Fig 7e). Interestingly, the Rv2140C(5DR1 1 -specific CD4+ T cells in IGRA- ACS subjects exhibited a higher portion of TSCM and TCM, whereas the ESAT6(2539) / DR1 1 -specific CD4+ T cells in IGRA+ ACS subjects had more abundant TE (Fig 5f). Consistent with prior studies, both antigen-specific populations contained a considerable fraction of naive T cells and undifferentiated TSCM, which suggested a pool of long-lived precursors contributing to durable immune memory (Fig 5f). In terms of effector subset distribution, ESAT6<2539 DRI 1 -specific CD4+ T cells in IGRA+ ACS subjects demonstrated significantly higher frequencies of Th1 , Th2 / Th22, and Treg compared to the RV2140C(5 19) / DR1 1 -specific CD4+ T cells in IGRA- ACS subjects (Fig 5f). In contrast, the Rv2140C(5-i9) / DR1 1 -specific CD4+ T cells in IGRA- ACS subjects were dominated by Tfh and Th17 subsets, at frequencies of 29.6% and 24.1 %, respectively, highlighting a distinct immune signature (Fig 5f).

[0231] Given the known function diversity of Tfh subpopulations, including Tfh1 , Tfh2, and Tfh17, we further analyzed the transcriptional profiles of the Tfh subset within Rv2140c<5- 19) / DR11 -specific CD4+ cells in IGRA- ACS subjects (Fig 5h). In addition to the canonical Tfh markers such as BCL6, CD185, IL-6, and IL-21 , there was detectable expression of IL- 2, IL-5, CCR4, CCR5, and RORyt, suggesting the presence of Tfh1 , Tfh2, and Tfh17 subpopulations (Fig 5h). Cytokine expression profile of the antigen-specific CD4+ T cells in IGRA+ and IGRA- ACS subjects largely mirrored their phenotypic distribution (Fig 5I). Robust IFNy expression was observed in ESAT6<2539 DR1 1 -specific CD4+ T cells from IGRA+ ACS subjects, consistent with the characteristic Th1 * response, whereas IFNy was almost absent in Rv2140C(5i9) / DR11 -specific CD4+ cells in IGRA- ACS subjects (Fig 5h). Notably, there was a considerable fraction of antigen-specific cells expressed TGF in both groups (Fig 5h). In IGRA+ ACS subjects, Treg cells were the primary source of TGF within the ESAT6(2539) / DR1 1 -specific CD4+ T cells, with additional contributions from Th17 and Tfh subsets (Fig 5j). In IGRA- ACS subjects, TGF|3 expression among Rv2140C(519 DR1 1 - specific CD4+ T cells distributed across Th17, Tfh and Treg subsets, indicating a multifunctional regulatory network of the TGF|3 (Fig 5j).

[0232] Rv2140cfs g specific CD4+ T cell responses likely originated from BCG vaccination with an early Tfh-skewed phenotype. Given that the Rv2140c gene is highly conserved across various mycobacterial species, including the BCG vaccine strain, we sought to investigate whether the RV2140C(5 I ^-specific CD4+ T cell responses observed in IGRA- ACS subjects could originate from BCG vaccination. To explore this, we analyzed DR1 1 :01 individuals from a Uganda infant study. All the enrolled infants received a BCG vaccination at birth in accordance with local public health policy, and the PBMC samples were collected eight weeks post vaccination. Although the infants were not screened with QuantiFERON test or TST, they were presumed not persistently exposed to Mtb or other mycobacteria at the time of sampling. We tested the infant PBMCs (N=10) with the same spheromer reagents of RV2140C<5 -19) / DR11 and ESAT6(2539) / DR1 1 to investigate the magnitude and phenotype of the antigen-specific CD4+ T cells. Remarkably, we detected Rv2140c<519> / DR1 1 -specific CD4+ T cells in four out of ten infants, with frequencies ranging from 0.1 1 % to 0.88% of total circulating CD4+ T cells, while the remaining six were negative for this response, possibly due to the generally weak immune responses in infants (Figs 6a and 6b). Importantly, none of the ten infants were positive for the ESAT6<2539 DRI 1 reagent, further indicating a lack of Mtb exposure (Figs 6a and 6b).

[0233] Next, we applied the same multi-parameter flow cytometry panel to assess the phenotype of the Rv2140c<51 ^-specific CD4+ T cells in the BCG-vaccinated infants. Notably, we observed a prominent population of CD45RA+Tfh cells among the antigen-specific population (Figs 6c and 6d), indicating that BCG vaccination can elicit early Tfh responses in infants. Aligned with the features of T cell responses in infants, nearly 35% of the antigenspecific CD4+ T cells are in a naive-like phenotype with CD45RA expression (Figs 6c and 6d). It is worth noting that we barely detect Th1 * responses within the Rv2140C(5i9) / DR1 1 - specific CD4+ T cells from BCG-vaccinated infants (Figs 6c and 6d). Taken together, these results suggest that the initial T cell response to the Rv2140c antigen is likely a recall response from the BCG vaccine, retaining a Tfh-skewed phenotype.

[0234] It has long been thought that subjects who are consistently negative for the QuantiFERON test and / or TST are uninfected by Mtb. But recent studies of individuals in Ugandan households that had repeated contact with Mtb infected active tuberculosis disease individuals have challenged this paradigm, since some were negative for both the QuantiFERON test and TST yet showed robust humoral and cellular immunity to Mtb and have been termed “resistors” (RSTR). Here we used GLIPH3 to analyze TCR alpha and beta chains from Mtb lysate stimulated CD4+ T cells from a small group of resistors or IGRA- individuals and compared them to IGRA+ individuals from the same cohort. GLIPH3 is amodification of previous versions to allow the analysis of TCR alpha beta pairs. Of the TCR sequences and GLIPH3 clusters obtained, we identified 24 TCR clusters that were unique to the RSTR individuals. Screening a large panel of proteins translated in vitro from -3700 Mtb genes, essentially the entire genome as described in Huang et al., we were able to identify both an ESAT6 antigen and an epitope of the Rv2140c gene. While ESAT6 is a major Mtb specific protein, which is utilized in the Quantiferon test, Rv2140c has not been previously recognized as a T cell target for Mtb. Furthermore, it is highly conserved across Mycobacterial species. Both epitopes correlated with HLA-DRB1 *11 :01 expression and so we made spheromer probes with the peptides bound to that class II MHC molecule and analyzed PBMCs from the ACS cohort that were either IGRA+ or IGRA- and expressed that HLA allele. We then chose five IGRA- individuals from the ACS cohort and stained their T cells with both the ESAT6(2539) / DR11 -specific reagents. All five were stained with both probes and with the ESAT6 result showing that they had all been exposed to Mtb. This means that the RSTR results in the small Uganda population are not a unique population, but that subjects in a large cohort thousands of miles away in South Africa had similar characteristics. Even more surprising was the very robust Rv2140c response of almost 1 % of the CD4+ T cells with their Tfh phenotype. This contrasts with the predominance of the Th1 * phenotype of most Mtb-specific T cells in IGRA+ individuals shown in many studies. This Tfh phenotype has been associated with the control of Mtb infection in mouse models.

[0235] Using a pool of overlapping peptides of Rv2140c, we also stimulated 10 other IGRA- individuals who did not express DRB1 *1 1 :01 and obtained robust T cell responses, showing that this antigen was a common target of CD4+ T cells in this cohort generally. In addition, analyses of the ESAT6 and RV2140c reagents in IGRA+ DRB1 *11 :01 individuals showed a reversal of the proportions seen in IGRA-‘s that is much greater ESAT6 specificity and much less Rv2140c, but perhaps even more important little of no Tfh phenotype but more of the standard Th1 type. We conclude from this that IGRA- individuals in Mtb endemic areas have been infected by Mtb, but have a unique T cell response that may be more effective at preventing TB disease than the IGRA+ phenotype.

[0236] It was also intriguing that the Rv2140c epitope is 100% conserved in other Mycobacterial species, including BCG. For this reason, we analyzed DRB1 *1 1 :01 infants in a Uganda cohort that had been vaccinated with BCG at birth 8 weeks previously. We found that of 10 infants analyzed, the four that had detectable antigen-specific were CD45RA+ Tfh, with a strong stem-like bias, even more than the adults discussed above. Also significant was the fact that there were no detectable Th1 * T cells in any of the four Rv2140C(519 DRI 1 reactive infants. Since the Cape Town subjects are also vaccinated with BCG at birth, this data indicates that the initial T cell response to the BCG vaccine is Tfh and not Th1 *. This shows that the Tfh response we see for the Rv2140c response in IGRA- teenagers in the ACS cohortis a continuation of the BCG response, perhaps augmented by exposure to other Mycobacterial species to result in nearly 1% of all CD4 T cells being specific for this epitope.

[0237] We also know that a large fraction of IGRA- individuals become IGRA+ over time, from approximately 50% at age 10 to 70-80% of adults, and this conversion is accompanied by a dramatic change in T cell response phenotype, from largely Tfh as described here to the Th1 * phenotype seen ubiquitously in IGRA+ individuals. This suggests that the original Tfh response is changed, perhaps by repeated Mtb exposure, to Th1 *. There is some precedent for a pathogen changing an immune response, such as a cohort of monozygotic twins discordant for CMV having very different immune characteristics Brodin et al 2015, but such a wholesale change in the T cell phenotype has not been seen previously.

[0238] Also worthy of note is the considerable fraction of stem cell-like memory subset, TSCM, in the antigen-specific CD4+ T cells of both IGRA- and IGRA+ ACS donors from Cape Town, and even more dramatically in response to the BCG vaccine in the Uganda infants. This result is consistent with our recently published findings in Mtb-reactive CD4+ T cells from the Ugandan household contact cohort defined by the activation markers. Pathogen-specific TSCM cells have been widely reported in human acute and chronic infections caused by viruses and bacteria, and recent studies revealed a negative correlation between the disease severity of a parasite infection and the magnitude of circulating TSCM cells. This dynamic has also been seen in the response to the yellow fever vaccine (YFV). YFV-specific TSCM cells became detectable in the early phase after vaccination when the T cell responses were dominated by the effector T cells, persisted at a stable level, and became the majority of YFV-specific T cell population in the circulation decades after the initial vaccination. This is characteristic of very robust and long-lived responses and is probably the ideal to look for in vaccine efficacy.

[0239] In summary, we find that, as suggested by previous work defining IGRA- resistor individuals, this phenotype is widespread in the Cape Town ACS cohort, in that there are clear signs of Mtb exposure, but that the T cell response is very unique compared to that of IGRA+ individuals, being predominantly Tfh versus Th1*, especially as exemplified by the highly conserved and dominant Rv2140c antigen. Our analysis of infants recently vaccinated with BCG indicates that this Tfh response to Rv2140c originates with BCG vaccination and persists until individuals convert to the IGRA+ state, which is accompanied by a reduction in the Rv2140c response and a conversion to the Th1 * phenotype. This remarkable switch of T cell phenotype is a case of how a pathogen is able to modify a hosts immune system to better suit its needs.Example 2

[0240] The experimental design and analysis approach applied to identify the Mtb antigens recognized by resister-enriched Mtb-reactive TCRs is shown in FIG. 1 A. As illustrated in FIG. 8, two strategies were applied to make a thorough screening. (A) the candidate TCRs and HLAs were first screened by the Mtb lysate to cover all potential Mtb antigens in the genome, followed by the protein pools of each Mtb ORF. Subsequently, the downstream screening narrowed the antigen with the single protein, peptide pool covering the positive protein, and the single peptide. In this way, we identified an antigen peptide (PDPYAALPKLPSFSL) within the Rv2140c protein that is presented by HLA-DRB1 *1 1 :01 and recognized by TCR-1 . B. the candidate TCRs and HLAs were further screened by ESAT6 / CFP10 peptide pool specifically. Another antigen SEQ ID NO:5 (IHSLLDEGKQSLTKL) within the ESAT6 protein was captured presented by HLA-DRB1 *1 1 :01 and recognized by TCR-10.

[0241] The experimental design to study the phenotype and function of CD4+ T cells recognized Rv2140c / DR1 1 and ESAT6 / DR1 1 is shown in FIG. 9. The PBMC samples were collected from another independent cohort (Adolescent Cohort Study, ACS) from South Africa (Zak et al., 2016, The Lancet). The PBMC samples from QFN+ donors (N=5) and EQFN- donors (N=5) were stained with Mtb spheromer with Rv2140c / DR11 , ESAT6 / DR1 1 and other published ESAT6 epitopes with HLA-DRB1 *04:01 and HLA-A*02:01 to validate the Mtb infection history. To determine the phenotype and cytokine expression of the Rv2140c / DR1 1 - specific CD4+ T cells, a panel of surface markers and cytokines were tested with the FACS cytometry assay. FIG. 10 shows the magnitude of CD4+ T cell response to Rv2140c / DR1 1 in HLA-DRB1 *11 :01 EQFN- ACS donors, QFN+ ACS donors, and health control donors. Representative FACS plots are listed on right.

[0242] Shown in FIG. 11 is the magnitude of CD4+ T cell response to (a) ESAT6 / DR1 1 in HLA-DRB1 *11 :01 EQFN- ACS donors, QFN+ ACS donors, and health control donors. The other published ESAT6 epitopes (b) ESAT6 / DR4 and (c) ESAT6 / A2 were also tested to measure the Mtb expose history according to the HLA alleles carried by the EQFN- ACS donors and QFN+ ACS donors. Representative FACS plots are listed on right.

[0243] The CD4+ T cell subsets distribution of Rv2140c / DR1 1 CD4+ T cells (in red) in HLA- DRB1 *1 1 :01 EQFN- ACS donors is shown in FIG. 12. Up to 1 % of the CD4+ cell population comprises Tfh specific for the Mtb epitope.

[0244] The cytokine expression of Rv2140c / DR11 CD4+ T cells in HLA-DRB1 *1 1 :01 EQFN- ACS donors after the stimulation with Mtb lysate, showing significant release of TGF[3. As shown in FIG 14, the TGF0 expression level of the Tfh, Th9 and Th1 subsets within the Rv2140 / DR11 -specific CD4+ T cells was heavily skewed to Tfh cells. Other HLA types, shown in FIG. 15, also responded to stimulation with Rv2140c peptide pool.

Claims

WHAT IS CLAIMED IS:1 . An immunogenic composition comprising: an effective dose of (a)(i) an Rv2140c protein sequence of SEQ ID NO:3; (ii) a variant of an Rv2140c protein sequence; or (iii) an immunogenic fragment of an Rv2140c protein sequence; or (b) a polynucleotide comprising a nucleic acid sequence encoding the polypeptide of (a); and (c) a pharmaceutically acceptable carrier or excipient.

2. The immunogenic composition of claim 1 , further comprising an adjuvant.

3. The immunogenic composition of claim 1 or claim 2, comprising a polynucleotide encoding an Rv2140c protein sequence of SEQ ID NO:3 or an immunogenic fragment thereof.

4. The immunogenic composition of any of the previous claims, comprising a polynucleotide encoding an Rv2140c protein sequence, wherein the sequence comprises the peptide epitope of SEQ ID NO:1.

5. The immunogenic composition of claim 4, wherein the polynucleotide is a modified mRNA.

6. The immunogenic composition of claim 5, wherein the polynucleotide is packed in a carrier.

7. The immunogenic composition of claim 6, wherein the carrier is a lipid nanoparticle.

8. The immunogenic composition of any of claims 1 -7, further comprising one or more addition Mycobacterium tuberculosis proteins or polynucleotide coding sequences.

9. A method for the treatment or prevention of tuberculosis, comprising: administering an effective dose of an immunogenic composition according to any of claims 1 -8, wherein the administration induces a protective immune response.

10. The method according to claim 9, wherein the subject is a mammal.1 1 . The method according to claim 9 or claim 10, wherein the subject is a human.

12. The method according to any one of claims 9-1 1 , wherein the subject has active tuberculosis.

13. The method according to any one of claims 9-12, wherein the subject has latent tuberculosis.

14. The method according to any one of claims 9-12, wherein the subject does not have tuberculosis.

15. The method according to any one of claims 9-14, wherein the subject was previously immunized with Bacillus Calmette-Guerin (BCG).

16. The method according to any one of claims 9-15, wherein the Rv2140c protein sequence is from Mycobacterium tuberculosis.

17. The method according to claim 16, wherein the Rv2140c protein has the sequence of SEQ ID NO: 1.

18. The method according to any one of claims 9-17, further comprising the administration of one or more chemotherapeutic agents effective in treating tuberculosis.

19. The method according to claim 18, wherein the one or more chemotherapeutic agents are selected from isoniazid and rifampin.

20. The method according to any one of claims 9-18, further comprising the administration of at least one additional Mycobacterium tuberculosis antigen.21 . The method according to claim 20, wherein the additional Mycobacterium tuberculosis antigen is provided in the form of a polypeptide.

22. The method according to claim 20, wherein the additional Mycobacterium tuberculosis antigen is provided in the form of a polynucleotide.

Citation Information

Patent Citations

  • Compositions And Methods For Immunodominant Antigens of Mycobacterium Tuberculosis

    US20090285847A1

  • Nucleic acid fragments and polypeptide fragments derived from m. tuberculosis

    WO1998044119A1