RNA for preventing or treating tuberculosis
Patent Information
- Application Number
- EP2024725309
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-09-29
- Filing Date
- 2024-04-12
- Publication Date
- 2026-02-18
AI Technical Summary
Current tuberculosis vaccines face challenges such as low safety, variable efficacy, and the need for adjuvants, with existing candidates failing to demonstrate significant protection against tuberculosis in clinical trials, particularly for immunocompromised individuals and those with latent infections.
Development of RNA vaccines encoding chimeric proteins comprising multiple Mycobacterium tuberculosis antigens, designed to optimize T cell responses, which are administered intramuscularly or intravenously to induce robust and specific immune responses without the risks associated with live vaccines.
The RNA vaccines elicit robust T cell responses and provide protective immunity against tuberculosis, potentially offering safer and more effective prevention and treatment options, including for immunocompromised individuals and latent infections.
Smart Images

Figure US2024024500_17102024_PF_FP_ABST
Abstract
Description
[0001] RNA FOR PREVENTING OR TREATING TUBERCULOSIS
[0002] CROSS REFERENCE TO RELATED APPLICATIONS
[0003] The present application claims priority to United States Provisional Patent Application Nos. 63 / 496,147, filed April 14, 2023, 63 / 496,149, filed April 14, 2023, 63 / 512,681, filed July 10, 2023, 63 / 512,683, filed July 10, 2023, and US 63 / 586,839, filed September 29,2023, each of which is hereby incorporated by reference in its entirety.
[0004] Technical Field
[0005] The disclosure provides agents and methods for preventing or treating tuberculosis using RNA. The RNA encoding chimeric antigens of Mycobacterium tuberculosis, immunogenic variants or fragments thereof is formulated and administered in a way that the antigens, variants or fragments are produced by cells of a subject, in particular after intramuscular or intravenous administration of the RNA.
[0006] Background
[0007] The use of RNA to deliver foreign genetic information into target cells offers an attractive alternative to DNA. The advantages of RNA include transient expression and non-transforming character. RNA does not require nucleus infiltration for expression and moreover cannot integrate into the host genome, thereby eliminating the risk of oncogenesis.
[0008] The COVID-19 pandemic has showcased the utility and advantages of RNA technology for vaccination, as out of all COVID- 19 vaccines under development, the first two to have received emergency use authorization by the FDA were RNA-based. The biotechnology response to the COVID- 19 pandemic has highlighted the speed and flexibility of mRNA vaccines, and reveals mRNA therapeutics to be a powerful tool to address epidemic outbreaks caused by newly emerging viruses. The relative simplicity of the development process and flexibility of the manufacturing platform can markedly accelerate clinical development. As such, mRNA-based vaccine technology has attracted a lot of attention during the COVID-19 pandemic.
[0009] The first authorized vaccine was developed by BioNTech in collaboration with Pfizer. The RNA of this vaccine, BNT162b2, encodes full length spike protein modified by two proline mutations to stabilize the prefusion conformation. The RNA incorporates 1-methyl-pseudouridine, which dampens innate immune sensing and increases mRNA translation in / / Vo and is formulated in lipid nanoparticles (LNP). BNT162b2 is administered to adults intramuscularly (IM) in two 30 pg doses given 21 days apart.
[0010] Tuberculosis (TB) is caused by the bacterial pathogen Mycobacterium tuberculosis (Mtb~) and, in rarer cases, by other pathogens from the Mycobacteriaceae family and is the leading cause of death from a single infectious agent. Mtb is a gram-positive, rod-shaped bacterium from the Mycobacteriaceae family. The more than 4,000 genes encoded within an approximately 4 million base pair genome render Mtb a complex pathogenic organism. This is further emphasized by the atypical composition of its cell wall, which has a high lipid content.
[0011] Despite the observed trend for reduction in TB cases and TB-related deaths for the last 20 years, 1.42 million people died from TB alone in 2019. In addition to the active form of TB, difficulties arise from latent TB infection (LTBI), when the infected patient doesn't present clinical symptoms. The estimated 2 billion latently infected individuals worldwide pose a huge and unpredictable reservoir of Mtb (WORLD HEALTH ORGANIZATION. Global tuberculosis report 2019. Geneva, WORLD HEALTH ORGANIZATION; 2019. ISBN: 978-92-4-156571-4). The high prevalence of HIV-1 infections further increases the risk for TB disease acquisition, activation of a latent TB infection, and death from HIV-TB co-infection. In 2009, 0.2 million deaths were related to HIV-TB comorbidity. The complexity of the Mtb cell wall makes the bacterium resistant to environmental impact and to therapy with certain antibiotics. The latter further complicates anti-TB treatment especially in low- and middle-income countries (WORLD HEALTH ORGANIZATION. Global tuberculosis report 2020. Geneva, WORLD HEALTH ORGANIZATION; 2020. ISBN: 978-92-4-001313-1).
[0012] An attenuated strain of Mycobacterium bovis, bacillus Calmette-Guerin (BCG), is the only licensed TB vaccine, introduced in 1921. The use of the live vaccine BCG is not recommended for immunocompromised individuals and the protective efficacy against pulmonary TB conferred by immunization with BCG is highly variable, ranging from 50-80%. Moreover, passaging of BCG over the decades further attenuated the currently used BCG strains, reducing its protective efficacy (Brosch R, et al. Proc. Natl. Acad. Sci. U.S.A., 2007; 104(13):5596-5601). Thus, there is an unmet medical need for a safer and more effective vaccine to prevent TB, especially for a vaccine that can be administered to immunocompromised individuals.
[0013] The pipeline of clinical trials for TB vaccine candidates comprises use of live, live-attenuated, and inactivated mycobacteria, and of Mtb antigens as recombinant protein (subunit vaccine) (TuBerculosis Vaccine Initiative (TBVI). Available from: https: / / www.tbvi.eu / what-we-do / pipeline-of-vaccines / ). The drawbacks from these vaccine platforms are i. their low safety, due to replication-competent live vaccines still being infectious, II. low immunogenicity of inactivated vaccines, and ill. the need for addition of adjuvants to subunit vaccines to enhance immunogenicity. To date, most vaccine candidates have failed to demonstrate better protection from TB or from the development of TB compared to placebo in clinical trials.
[0014] For all these reasons, novel agents for preventing or treating tuberculosis are required.
[0015] Summary
[0016] The present disclosure provides compositions which are useful as TB vaccines. The compositions provided herein comprise RNA for delivering Mycobacterium tuberculosis (Mtb) antigens to a subject.
[0017] Immunity to tuberculosis is primarily shaped by T cell responses in most individuals. We have developed a platform of RNA design specifically aimed at optimizing T cell responses in RNA vaccines, which is utilized here to specifically design potent T cell inducing RNA vaccines to prevent or treat tuberculosis. Antigens for these vaccines have been selected to be included either as full-length antigens or as antigen fragments. These full-length antigens or fragments thereof are combined into chimeric proteins by stringing them together interspersed by polypeptide linkers that also function to minimize the risk of creating neoepitopes. The novel RNA vaccines encoding chimeric proteins comprising three or more antigens, as full-length antigens and antigen fragments, elicit a robust and specific T cell response in subjects infected with or at risk of contracting tuberculosis. Fifteen antigens were selected as primary targets for the novel RNA vaccines, including Wbbll, PPE18, PE13, EsxA, EsxB, EsxG, EsxH, Esxl, EsxJ, EsxK, EsxL, EsxM, EsxN, EsxV and EsxW.
[0018] Furthermore, in order to improve the rate of translation of the chimeric protein encoded by the novel RNA vaccines, increasing the amount of protein produced and presented to the immune system, we have developed ways to increase coverage of the elicited immune response while reducing or maintaining the length of the chimeric proteins. For instance, the Esx proteins EsxN, Esxl, EsxV and EsxL are highly homologous members of the same protein family and are known to be highly immunogenic. Each of these proteins is transcribed and produced with a direct secretion partner - another Esx protein - and these two secretion partners together form a heterodimer. These partners are EsxJ (partner of Esxl), EsxK (partner of EsxL), EsxM (partner of EsxN), and EsxW (partner of EsxV). This second group of proteins (together referred to as the QQILSS proteins based on the residues in their C-terminal sequences, also contains highly immunogenic proteins playing a role in virulence or dissemination. Based on sequence analysis, we identified that if a full-length Esxl protein is combined with a fragment of EsxN (or a full-length EsxN), these antigens would also induce immune responses against the other antigens EsxV and EsxL. Similarly, inclusion of the N terminal 14 amino acids of Esxl together with a full-length EsxW, would not only cover these two proteins completely, but would also cover 1) EsxM, in strains that express an ancestral version of this protein associated with dissemination within the host; 2) Strains that possess a single nucleotide polymorphism in EsxW associated with increased transmissibility; and 3) the full-length EsxK protein except for alanine residue 58.
[0019] In addition to the above, the antigens forming the core functionality of the novel RNA vaccines, can be combined with other known Mtb antigens. For instance, Mtb displays differential gene expression patterns during its active and dormant (non-dividing) phases (Andersen P, et al. Cold Spring Harb Perspect Med, 2014; 4(6):a018523). To prevent development ofTB, immunity against antigens specific for various stages of Mtb infection is preferable. The TB vaccines developed here comprising the RNA components described above are designed to induce protective immune responses against antigens specific for different stages of Mtb infection. Combinations with additional known Mtb antigens can be used to elicit a stronger immune response towards individual stages of Mtb infection or to increase the number of stages of Mtb infection, against which the vaccine elicits an immune response.
[0020] Furthermore, antigen presentation and trafficking can be optimized by adding to the chimeric protein an N-terminal signal peptide designed to traffic the recombinant protein towards the exterior of transfected cells. The C-terminus of the chimeric protein can be modified by a trafficking signal and / or transmembrane domain designed to enhance endocytosis of the recombinant protein, thereby leading to optimal MHC-II presentation and ensuing CD4+T cell responses.
[0021] The RNA vaccines described herein, e.g., comprising non-modified uridine containing mRNA (uRNA) or nucleoside modified mRNA (modRNA), expressing Mtb antigens, immunogenic variants or fragments thereof, are useful for preventing or treating tuberculosis. The RNA encoding Mtb antigens, immunogenic variants or fragments thereof is formulated and administered in a way that the antigens, variants or fragments can be produced and preferably secreted by patient cells to prevent or combat tuberculosis.
[0022] Unlike the attenuated vaccine BCG, this TB vaccine candidate does not carry the risks associated with infection and may therefore be given to people who cannot be administered live organism (such as pregnant women and immunocompromised persons).
[0023] In one aspect, the disclosure provides an RNA molecule encoding a chimeric protein comprising T cell epitopes of three or more different Mycobacterium tuberculosis antigens or immunogenic variants thereof, wherein the chimeric protein comprises one or more antigen fragments, wherein at least one Mycobacterium tuberculosis antigen or immunogenic variant thereof is represented by one or more antigen fragments and the remaining Mycobacterium tuberculosis antigens or immunogenic variants thereof are represented by one or more antigen fragments and / or one or more full-length antigens, wherein each antigen fragment or full-length antigen comprises one or more T cell epitopes, and wherein each antigen fragment or full-length antigen is separated from other antigen fragments or full-length antigens in the chimeric protein by a polypeptide linker.
[0024] In some embodiments of the RNA molecule, the chimeric protein comprises T cell epitopes of four or more different Mycobacterium tuberculosis antigens. In some embodiments, the chimeric protein comprises T cell epitopes of five or more different Mycobacterium tuberculosis antigens. In some embodiments, the chimeric protein comprises T cell epitopes of six or more different Mycobacterium tuberculosis antigens. In some embodiments, the chimeric protein comprises T cell epitopes of seven or more different Mycobacterium tuberculosis antigens.
[0025] In some embodiments, the Mycobacterium tuberculosis antigens are selected from the group of Wbbll, PPE18, PE13, EsxA, EsxB, EsxG, EsxH, Esxl, EsxJ, EsxK, EsxL, EsxM, EsxN, EsxV and EsxW.
[0026] In some embodiments: a) the Wbbll antigen comprises the amino acid sequence of SEQ ID NO: 1 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 1; b) the PPE18 antigen comprises the amino acid sequence of SEQ ID NO: 2 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 2; c) the PE13 antigen comprises the amino acid sequence of SEQ ID NO: 3 or SEQ ID NO: 4 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 3 or SEQ ID NO: 4; d) the EsxA antigen comprises the amino acid sequence of SEQ ID NO: 5 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 5; e) the EsxB antigen comprises the amino acid sequence of SEQ ID NO: 6 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 6; f) the EsxG antigen comprises the amino acid sequence of SEQ ID NO: 7 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 7; g) the EsxH antigen comprises the amino acid sequence of SEQ ID NO: 8 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 8; h) the Esxl antigen comprises the amino acid sequence of SEQ ID NO: 9 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 9; i) the EsxJ antigen comprises the amino acid sequence of SEQ ID NO: 10 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 10; j) the EsxK antigen comprises the amino acid sequence of SEQ ID NO: 11 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 11; k) the EsxL antigen comprises the amino acid sequence of SEQ ID NO: 12 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 12; l) the EsxM antigen comprises the amino acid sequence of SEQ ID NO: 13 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 13; m) the EsxN antigen comprises the amino acid sequence of SEQ ID NO: 14 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 14; n) the EsxV antigen comprises the amino acid sequence of SEQ ID NO: 15 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 15; and / or o) the EsxW antigen comprises the amino acid sequence of SEQ ID NO: 16 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 16.
[0027] In some embodiments, the Mycobacterium tuberculosis antigens comprise Wbbll, PPE18 and PE13.
[0028] In some embodiments, the chimeric protein comprises: a) a first and a second antigen fragment of PE13; b) a first, a second and a third antigen fragment of PPE18; and c) a full-length antigen of WbbLl.
[0029] In some embodiments: a) the first antigen fragment of PE13 comprises the amino acid sequence of positions 1 to 27 of SEQ ID NO: 3 or an amino acid sequence having at least 96%, 92%, 88%, 84% or 80% identity to the amino acid sequence of positions 1 to 27 of SEQ ID NO: 3; b) the second antigen fragment of PE13 comprises the amino acid sequence of positions 39 to 99 of SEQ ID NO: 3 or an amino acid sequence having at least 98%, 96%, 94%, 92%, 88%, 84% or 80% identity to the amino acid sequence of positions 39 to 99 of SEQ ID NO: 3; c) the first antigen fragment of PPE18 comprises the amino acid sequence of positions 1 to 155 of SEQ ID NO: 2 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 1 to 155 of SEQ ID NO: 2; d) the second antigen fragment of PPE18 comprises the amino acid sequence of positions 186 to 296 of SEQ ID NO: 2 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 186 to 296 of SEQ ID NO: 2; and / or e) the third antigen fragment of PPE18 comprises the amino acid sequence of positions 303 to 391 of SEQ ID NO: 2 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 303 to 391 of SEQ ID NO: 2. In some embodiments, the full-length antigens and antigen fragments in the chimeric protein are arranged in the following order: first antigen fragment of PE13 - linker - third antigen fragment of PPE18 - linker - full-length antigen of Wbbll - linker - second antigen fragment of PPE18 - linker - second antigen fragment of PE13 - linker - first antigen fragment of PPE18.
[0030] In some embodiments, the Mycobacterium tuberculosis antigens comprise EsxG, EsxH, Esxl, EsxJ, EsxN, and EsxW. In some embodiments, the Mycobacterium tuberculosis antigens additionally comprise EsxB.
[0031] In some embodiments, the chimeric protein comprises: a) a full-length antigen of EsxG; b) a full-length antigen of EsxH; c) a full-length antigen of Esxl; d) an antigen fragment of EsxJ; e) a full-length antigen of EsxN; and f) a full-length antigen of EsxW.
[0032] In some embodiments, the chimeric protein additionally comprises: a) an antigen fragment of EsxB b) a full-length antigen of EsxG; c) a full-length antigen of EsxH; d) a full-length antigen of Esxl; e) an antigen fragment of EsxJ; f) a full-length antigen of EsxN; and g) a full-length antigen of EsxW.
[0033] In some embodiments: a) the antigen fragment of EsxB comprises the amino acid sequence of positions 1 to 87 of SEQ ID NO: 6 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 1 to 87 of SEQ ID NO: 6; and / or b) the antigen fragment of EsxJ comprises the amino acid sequence of positions 1 to 14 of SEQ ID NO: 10 or an amino acid sequence having at least 93%, 86% or 79% identity to the amino acid sequence of positions 1 to 14 of SEQ ID NO: 10.
[0034] In some embodiments, the full-length antigens and antigen fragments in the chimeric protein are arranged in the following order: a) full-length antigen of EsxG - linker - full-length antigen of EsxH - linker - full-length antigen of Esxl
[0035] - linker - antigen fragment of EsxJ - linker - full-length antigen of EsxN - linker - full-length antigen of EsxW; or b) full-length antigen of EsxG - linker - full-length antigen of Esxl - linker - full-length antigen of EsxH
[0036] - linker - antigen fragment of EsxJ - linker - full-length antigen of EsxN - linker - full-length antigen of EsxW.
[0037] In some embodiments, the full-length antigens and antigen fragments in the chimeric protein are arranged in the following order: a) antigen fragment of EsxB - linker - full-length antigen of EsxG - linker - full-length antigen of EsxH
[0038] - linker - full-length antigen of Esxl - linker - antigen fragment of EsxJ - linker - full-length antigen of EsxN - linker - full-length antigen of EsxW; or b) antigen fragment of EsxB - linker - full-length antigen of EsxG - linker - full-length antigen of Esxl
[0039] - linker - full-length antigen of EsxH - linker - antigen fragment of EsxJ - linker - full-length antigen of EsxN - linker - full-length antigen of EsxW.
[0040] In some embodiments, the Mycobacterium tuberculosis antigens comprise PE13 and two or more antigens selected from PPE18, EsxW, EsxJ. EsxG, Esxl, EsxN, EsxH and EsxA.
[0041] In some embodiments, the Mycobacterium tuberculosis antigens comprise a) PE13, PPE18, Esxl, EsxJ, EsxN and EsxW; b) PE13, PPE18, EsxG, EsxJ, and EsxW; c) PE13, PPE18, EsxG, Esxl, and EsxN; d) PE13, EsxG, EsxH, Esxl, EsxJ, EsxN and EsxW; e) PE13, EsxA, EsxG, EsxH, Esxl, EsxJ, EsxN and EsxW; f) PE13, Esxl, EsxJ, EsxN and EsxW; or g) PE13, PPE18, Esxl, EsxJ, EsxN and EsxW.
[0042] In some embodiments, the chimeric protein comprises: a) a first and a second antigen fragment of PE13; b) a first, a second and a third antigen fragment of PPE18; c) a full-length antigen of Esxl; d) an antigen fragment of EsxJ; e) an antigen fragment of EsxN; and f) a full-length antigen of EsxW.
[0043] In some embodiments, the chimeric protein comprises: a) a first and a second antigen fragment of PE13; b) a first, a second and a third antigen fragment of PPE18; c) a full-length antigen of EsxG; d) an antigen fragment of EsxJ; and e) a full-length antigen of EsxW.
[0044] In some embodiments, the chimeric protein comprises: a) a first and a second antigen fragment of PE13; b) a first, a second and a third antigen fragment of PPE18; c) a full-length antigen of EsxG; d) an antigen fragment of Esxl; and e) a full-length antigen of EsxN.
[0045] In some embodiments, the chimeric protein comprises: a) a first and a second antigen fragment of PE13; b) a full-length antigen of EsxG; c) a full-length antigen of EsxH; d) a full-length antigen of Esxl; e) an antigen fragment of EsxJ; f) an antigen fragment of EsxN; and g) a full-length antigen of EsxW.
[0046] In some embodiments, the chimeric protein comprises: a) a first and a second antigen fragment of PE13; b) a first and a second antigen fragment of EsxA; c) a full-length antigen of EsxG; d) a full-length antigen of EsxH: e) a full-length antigen of Esxl; f) an antigen fragment of EsxJ; g) an antigen fragment of EsxN; and h) a full-length antigen of EsxW.
[0047] In some embodiments, the chimeric protein comprises: a) a first and a second antigen fragment of PE13; b) a full-length antigen of Esxl; c) an antigen fragment of EsxJ; d) an antigen fragment of EsxN; and e) a full-length antigen of EsxW.
[0048] In some embodiments, the chimeric protein comprises: a) a first and a second antigen fragment of PE13; b) a second and a third antigen fragment of PPE18; c) a full-length antigen of Esxl; d) an antigen fragment of EsxJ; e) an antigen fragment of EsxN; and f) a full-length antigen of EsxW.
[0049] In some embodiments: a) the first antigen fragment of PE13 comprises the amino acid sequence of positions 1 to 27 of SEQ ID NO: 3 or an amino acid sequence having at least 96%, 92%, 88%, 84% or 80% identity to the amino acid sequence of positions 1 to 27 of SEQ ID NO: 3; or the first antigen fragment of PE13 comprises the amino acid sequence of positions 1 to 29 of SEQ ID NO: 4 or an amino acid sequence having at least 96%, 92%, 88%, 84% or 80% identity to the amino acid sequence of positions 1 to 29 of SEQ ID NO: 4; b) the second antigen fragment of PE13 comprises the amino acid sequence of positions 39 to 97 of SEQ ID NO: 3 or an amino acid sequence having at least 98%, 96%, 94%, 92%, 88%, 84% or 80% identity to the amino acid sequence of positions 39 to 97 of SEQ ID NO: 3; or the second antigen fragment of PE13 comprises the amino acid sequence of positions 41 to 99 of SEQ ID NO: 4 or an amino acid sequence having at least 98%, 96%, 94%, 92%, 88%, 84% or 80% identity to the amino acid sequence of positions 41 to 99 of SEQ ID NO: 4; c) the first antigen fragment of PPE18 comprises the amino acid sequence of positions 1 to 155 of SEQ ID NO: 2 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 1 to 155 of SEQ ID NO: 2; d) the second antigen fragment of PPE18 comprises the amino acid sequence of positions 186 to 296 of SEQ ID NO: 2 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 186 to 296 of SEQ ID NO: 2; e) the third antigen fragment of PPE18 comprises the amino acid sequence of positions 303 to 391 of SEQ ID NO: 2 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 303 to 391 of SEQ ID NO: 2; f) the first antigen fragment of EsxA comprises the amino acid sequence of positions 1 to 35 of SEQ ID NO: 5 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 1 to 35 of SEQ ID NO: 5; g) the second antigen fragment of EsxA comprises the amino acid sequence of positions 26 to 81 of SEQ ID NO: 5 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 26 to 81 of SEQ ID NO: 5; h) the antigen fragment of EsxJ comprises the amino acid sequence of positions 1 to 14 of SEQ ID NO: 10 or an amino acid sequence having at least 93%, 86% or 79% identity to the amino acid sequence of positions 1 to 14 of SEQ ID NO: 10 and / or i) the antigen fragment of EsxN comprises the amino acid sequence of positions 10 to 67 of SEQ ID NO: 14 or an amino acid sequence having at least 93%, 86% or 79% identity to the amino acid sequence of positions 10 to 67 of SEQ ID NO: 14.
[0050] In some embodiments, the full-length antigens and antigen fragments in the chimeric protein are arranged in the following order: first antigen fragment of PE13 - linker - third antigen fragment of PPE18 - linker - full-length antigen of EsxW
[0051] - linker - antigen fragment of EsxJ - linker - second antigen fragment of PPE18 - linker - second antigen fragment of PE13 - linker - first antigen fragment of PPE18 - linker - full-length antigen of Esxl - linker - antigen fragment of EsxN.
[0052] In some embodiments, the full-length antigens and antigen fragments in the chimeric protein are arranged in the following order: first antigen fragment of PE13 - linker - third antigen fragment of PPE18 - linker - full-length antigen of EsxW
[0053] - linker - full-length antigen of EsxG - linker - second antigen fragment of PE13 - linker - second antigen fragment of PPE18 - linker - antigen fragment of EsxJ - linker - first antigen fragment of PPE18.
[0054] In some embodiments, the full-length antigens and antigen fragments in the chimeric protein are arranged in the following order: a) second antigen fragment of PE13 - linker - third antigen fragment of PPE18 - linker - antigen fragment of EsxN - linker - full-length antigen of EsxG - linker - first antigen fragment of PE13 - linker - second antigen fragment of PPE18 - linker - full-length antigen of Esxl - linker - first antigen fragment of PPE18; or b) second antigen fragment of PE13 - linker - first antigen fragment of PPE18 - linker - full-length antigen of Esxl - linker - second antigen fragment of PPE18 - linker - first antigen fragment of PE13 - linker - full-length antigen of EsxG - linker - antigen fragment of EsxN - linker - third antigen fragment of PPE18.
[0055] In some embodiments, the full-length antigens and antigen fragments in the chimeric protein are arranged in the following order: a) first antigen fragment of PE13 - linker - full-length antigen of EsxW - linker - full-length antigen of EsxG
[0056] - linker - second antigen fragment of PE13 - linker - full-length antigen of Esxl - linker - antigen fragment of EsxN - linker - antigen fragment of EsxJ - linker - full-length antigen of EsxH; b) second antigen fragment of PE13 - linker - full-length antigen of EsxW - linker - full-length antigen of EsxG - linker - first antigen fragment of PE13 - linker - full-length antigen of Esxl - linker - antigen fragment of EsxN
[0057] - linker - antigen fragment of EsxJ - linker - full-length antigen of EsxH; or c) second antigen fragment of PE13 - linker - full-length antigen of EsxH - linker - antigen fragment of EsxJ
[0058] - linker - antigen fragment of EsxN - linker - full-length antigen of Esxl - linker - first antigen fragment of PE13 - linker - full-length antigen of EsxG - linker - full-length antigen of EsxW.
[0059] In some embodiments, the full-length antigens and antigen fragments in the chimeric protein are arranged in the following order: a) first antigen fragment of PE13 - linker - full-length antigen of EsxW - linker - full-length antigen of EsxG
[0060] - linker - second antigen fragment of EsxA - linker - second antigen fragment of PE13 - linker - full-length antigen of Esxl - linker - first antigen fragment of EsxA - linker - antigen fragment of EsxN - linker - antigen fragment of EsxJ - linker - full-length antigen of EsxH; b) second antigen fragment of PE13 - linker - full-length antigen of EsxW - linker - full-length antigen of EsxG - linker - second antigen fragment of EsxA - linker - first antigen fragment of PE13 - linker - full-length antigen of Esxl - linker - first antigen fragment of EsxA - linker - antigen fragment of EsxN - linker - antigen fragment of EsxJ - linker - full-length antigen of EsxH; or c) second antigen fragment of PE13 - linker - full-length antigen of EsxH - linker - antigen fragment of EsxJ
[0061] - linker - antigen fragment of EsxN - linker - first antigen fragment of EsxA - linker - full-length antigen of Esxl - linker - first antigen fragment of PE13 - linker - second antigen fragment of EsxA - linker - full-length antigen of EsxG
[0062] - linker - full-length antigen of EsxW.
[0063] In some embodiments, one or more of the polypeptide linkers comprises one or more glycine and / or one or more serine amino acid. In some embodiments, one or more of the polypeptide linkers is at least 1, at least 5 or at least 10 amino acids in length. In some embodiments, one or more of the polypeptide linkers has the amino acid sequence of SEQ ID NO: 56.
[0064] In some embodiments, the chimeric protein comprises a non-native signal peptide at its N-terminus. In some embodiments, the non-native signal peptide is a human, bacterial or viral signal peptide. In some embodiments, the non-native signal peptide comprises a secretory signal. In some embodiments, the non-native signal peptide is functional in mammalian cells. In some embodiments, the non-native signal peptide comprises an amino acid sequence selected from the group of SEQ ID NOs: 17 to 37 or 65, an amino acid sequence having at least 96%, 92%, 88%, 84% or 80% identity to SEQ ID NOs: 17 to 37 or 65, an amino acid sequence encoded by a nucleotide sequence selected from the group of SEQ ID NOs: 38 to 51, or an amino acid sequence having at least 96%, 92%, 88%, 84% or 80% identity to amino acid sequences encoded by a nucleotide sequence selected from the group of SEQ ID NOs: 38 to 51. In some embodiments, the non-native signal peptide is a viral signal peptide. In some embodiments, the non-native signal peptide is a HSV-1 glycoprotein D signal peptide.
[0065] In some embodiments, the chimeric protein comprises a non-native transmembrane domain at its C-terminus. In some embodiments, the non-native transmembrane domain is a human, bacterial or viral transmembrane domain. In some embodiments, the non-native transmembrane domain comprises a trafficking domain. In some embodiments, non- native trafficking domain is an MHC class I trafficking domain. In some embodiments, the MHC class I trafficking domain comprises the amino acid sequence of SEQ ID NO: 52 or an amino acid sequence having at least 98%, 96%, 90%, or 80% identity to the amino acid sequence of SEQ ID NO: 52. In some embodiments, the non-native transmembrane domain is a viral transmembrane domain. In some embodiments, the non-native transmembrane domain is derived from HSV.
[0066] In some embodiments, the RNA molecule comprises a 5' cap. In some embodiments, the 5' cap comprises a capl structure. In some embodiments, the 5'-cap comprises m273-OGppp(mi2' °)ApG.
[0067] In some embodiments, the RNA molecule comprises a 5'-UTR. In some embodiments, the 5'-UTR comprises a modified human alpha-globin 5'-UTR. In some embodiments, the 5'-UTR comprises the nucleotide sequence of SEQ ID NO: 53, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 53.
[0068] In some embodiments, the RNA comprises a 3'-UTR. In some embodiments, the 3'-UTR comprises a first sequence from the amino terminal enhancer of split (AES) messenger RNA and a second sequence from the mitochondrial encoded 12S ribosomal RNA. In some embodiments, the 3'-UTR comprises the nucleotide sequence of SEQ ID NO: 54, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 54.
[0069] In some embodiments, the RNA molecule comprises a polyA sequence. In some embodiments, the polyA sequence is an interrupted sequence of A nucleotides. In some embodiments, the polyA sequence comprises 30 adenine nucleotides followed by 70 adenine nucleotides, wherein the 30 adenine nucleotides and 70 adenine nucleotides are separated by a nucleotide linker sequence of 10 nucleotides. In some embodiments, the polyA sequence comprises the nucleotide sequence of SEQ ID NO: 55, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 55.
[0070] In some embodiments, the RNA molecule comprises a 5'-cap, a 5'-UTR, a 3'-UTR and a polyA sequence.
[0071] In some embodiments, the RNA molecule comprises modified nucleotides, nucleosides or nucleobases. In some embodiments, the RNA molecule comprises modified uridines. In some embodiments, the RNA molecule comprises modified uridines in place of all uridines. In some embodiments, the modified uridines are Nl-methyl-pseudouridine.
[0072] In some embodiments, the coding sequence of the RNA molecule is codon-optimized and / or is characterized in that its G / C content is increased compared to the parental sequence.
[0073] In one aspect, the disclosure provides a chimeric protein encoded by the RNA molecule described herein.
[0074] In one aspect, the disclosure provides a DNA molecule encoding the RNA molecule described herein.
[0075] In one aspect, the disclosure provides a pharmaceutical composition comprising one or more RNA molecules described herein. In some embodiments of the pharmaceutical composition, the one or more RNA molecules are formulated in a lipid formulation, such as in lipid nanoparticles or liposomes. In some embodiments, the lipid formulation comprises each of: a) a cationically ionizable lipid; b) a steroid; c) a neutral lipid; and d) a polymer-conjugated lipid.
[0076] In some embodiments, the cationically ionizable lipid is present in a concentration ranging from about 40 to about 60 mol percent of the total lipids. In some embodiments, the steroid is present in a concentration ranging from about 30 to about 50 mol percent of the total lipids. In some embodiments, the neutral lipid is present in a concentration ranging from about 5 to about 15 mol percent of the total lipids. In some embodiments, the polymer-conjugated lipid is present in a concentration ranging from about 1 to about 10 mol percent of the total lipids. In some embodiments, the cationically ionizable lipid is within a range of about 40 to about 60 mole percent, the steroid is within a range of about 30 to about 50 mole percent, the neutral lipid is within a range of about 5 to about 15 mole percent, and the polymer- conjugated lipid is within a range of about 1 to about 10 mole percent.
[0077] In some embodiments, the cationically ionizable lipid comprises ((4-hydroxybutyl)azanediyl)bis(hexane-6,l-diyl)bis(2- hexyldecanoate). In some embodiments, the steroid comprises cholesterol. In some embodiments, the neutral lipid comprises a phospholipid. In some embodiments, the phospholipid comprises distearoylphosphatidylcholine (DSPC).In some embodiments, the polymer-conjugated lipid comprises a polyethylene glycol (PEG)-lipid. In some embodiments, the PEG-lipid comprises 2-[(polyethylene glycol)-2000]- / Vz / V-ditetradecylacetamide.
[0078] In some embodiments, the lipid formulation comprises:
[0079] (a) ((4-hydroxybutyl)azanediyl)bis(hexane-6,l-diyl)bis(2-hexyldecanoate);
[0080] (b) cholesterol;
[0081] (c) distearoylphosphatidylcholine (DSPC); and
[0082] (d) 2-[(polyethylene glycol)-2000]- / Vz / V-ditetradecylacetamide.
[0083] In some embodiments, ((4-hydroxybutyl)azanediyl)bis(hexane-6,l-diyl)bis(2-hexyldecanoate) is within a range of about 40 to about 60 mole percent, cholesterol is within a range of about 30 to about 50 mole percent, distearoylphosphatidylcholine (DSPC) is within a range of about 5 to about 15 mole percent, and 2-[(polyethylene glycol)-2000]-N,N-ditetradecylacetamide is within a range of about 1 to about 10 mole percent.
[0084] In some embodiments, the pharmaceutical composition further comprises one or more pharmaceutically acceptable carriers, diluents and / or excipients.
[0085] In some embodiments, the one or more RNA molecules are in a liquid formulation. In some embodiments, the one or more RNA molecules are in a frozen formulation. In some embodiments, the one or more RNA molecules are in a lyophilized formulation. In some embodiments, the one or more RNA molecules are formulated for injection. In some embodiments, the one or more RNA molecules are formulated for intramuscular administration. In some embodiments, the pharmaceutical composition is formulated for administration in human.
[0086] In one aspect, the disclosure provides a kit comprising one or more pharmaceutical compositions described herein. In some embodiments of the kit, two or more pharmaceutical compositions comprising the same or different RNA molecules described herein are in separate vials. In some embodiments, the kit further comprises instructions for use of the one or more pharmaceutical composition for treating or preventing tuberculosis.
[0087] In one aspect, the disclosure provides an RNA molecule, chimeric protein, DNA molecule, pharmaceutical composition or kit described herein for use as a medicament. In some embodiments, the use comprises a therapeutic or prophylactic treatment of a disease or disorder in a subject. In some embodiments, the use comprises the use as a vaccine against a disease or disorder in a subject. In some embodiments, the subject is a human infected with the disease or disorder or in danger of contracting the disease or disorder.
[0088] In one aspect, the disclosure provides an RNA molecule, chimeric protein, DNA molecule, pharmaceutical composition or kit described herein for use in treating or preventing tuberculosis in a subject. In some embodiments, the subject is a human suffering from tuberculosis or in danger of contracting tuberculosis. In some embodiments, the use is as a vaccine for preventing tuberculosis.
[0089] In one aspect, the disclosure provides a use of the RNA molecule, the chimeric, the DNA molecule, the pharmaceutical composition or the kit described herein for the manufacture of a medicament for preventing or treating tuberculosis.
[0090] In one aspect, the disclosure provides a method of vaccinating a subject comprising administering the RNA molecule, the chimeric protein, the DNA molecule, the pharmaceutical composition or the kit described herein to the subject. In some embodiments, the vaccination is for preventing tuberculosis. In some embodiments, administration is by intramuscular administration. In some embodiments, the method comprises administering to the subject at least one dose of the RNA, chimeric protein or pharmaceutical composition. In some embodiments, the method comprises administering to the subject at least two doses of the RNA, chimeric protein or pharmaceutical composition. In some embodiments, an amount of the RNA of at least 10 pg per dose is administered. In some embodiments, the subject is a human.
[0091] In some embodiments of the RNA molecule, chimeric protein, pharmaceutical composition or kit for use, the use or the methods described herein, the tuberculosis is caused by an infection with a Mycobacterium. In some embodiments, the Mycobacterium is selected from the group of Mycobacterium tuberculosis, Mycobacterium bovis, Mycobacterium caprae, Mycobacterium orygis, Mycobacterium africanum, Mycobacterium microti, Mycobacterium canetti and Mycobacterium pinnipedii. In some embodiments, the Mycobacterium is Mycobacterium tuberculosis.
[0092] Brief description of the Figures
[0093] Figures 1 to 11: Bioinformatic data guiding the selection of antigens and design of chimeric proteins
[0094] Each antigen of interest has been analyzed by different bioinformatic analyses as further described in the accompanying text. For each antigen is depicted: Panel A) Peptides overlapping with the human proteome. Fragments were selected to exclude any of the regions above 0.0.
[0095] Panel B) Predicted class I epitopes for a wide array of human HLA alleles. A cumulative score is listed with the y-axis representing how many peptides are predicted to bind the different HLA-alleles for each residue.
[0096] Panel C) The predicted start sites of the epitopes predicted in B.
[0097] Panel D) The predicted stop sites of the epitopes predicted in B.
[0098] Panels E-G) The same analyses as described in B-D, but for HLA class II alleles.
[0099] Figures 12-18: Domain structure of chimeric proteins
[0100] In Figures 12-18, respectively, panel A shows the domain structure of an exemplary chimeric protein disclosed herein. The chimeric protein comprises sequences from several antigens either in the form of full-length antigens or antigen fragments - separated from each other by polypeptide linkers. The exemplary chimeric protein comprises a signal peptide (sec) at its N-terminus, whereas a transmembrane domain is located at its C-terminus. The lower half of panel A shows the length of the protein sequences by number of amino acids.
[0101] In panel B, the figures respectively list the sequences and positions of individual domains within the domain structure of the exemplary chimeric protein.
[0102] Figures 19 and 20: Expression of chimeric proteins as analyzed by mass-spectrometry
[0103] In Figures 19 and 20, respectively, panel A) shows the position of tryptic digest fragments analyzed by mass- spectrometry on the sequence of the chimeric protein and panel B) shows the detection of tryptic digest fragments by mass-spectrometry.
[0104] Figures 21-23: Detection of epitope binding to four different human HLA subclasses
[0105] In Figures 21-23, respectively, panel A) shows the position of detected epitopes on the sequence of the chimeric proteins and panel B) shows the sequence and human HLA subtype presenting the epitopes identified in panel A.
[0106] Figures 24 and 25: In vivo T-cell responses to chimeric proteins in C57BI6 mice
[0107] In Figures 24 and 25, respectively, panel A) shows in the left column the in vivo T-cell responses in mice immunized with RNA encoding the chimeric proteins as measured by induction of interferon y (IFN y) in the total T-cell population - separated by domain. In the right column, IFN y induction is shown for CD4+and CD8+cells, respectively. Panels B) and C) have the same structure, showing respective induction of interleukin-2 (IL-2) and tumor necrosis factor o (TNFo). Panel D) shows the responses in NaCI mock-immunized mice as a negative control.
[0108] Figure 26: In vivo T-cell responses to chimeric proteins in transgenic human HLA A2.1 / DR1 mice
[0109] Panel A) shows the in vivo T-cell responses in mice immunized with RNA encoding the chimeric protein string 1 as measured by induction of interferon y (IFN y) in the total T-cell population - separated by domain. In panel B, IFN y induction is shown for CD4+and CD8+cells of string 1, respectively. Panels C and D and Panels E and F have the same structure but show the results for string 2 and string 3, respectively. Results are displayed as bars of the mean value ± standard deviation in combination with single values of individual mice. Figures show IFNy secretion as spot forming units in peptide-restimulated full splenoyctes (Panel A, C, E), and CD4+(white bars) or CD8+T cells (grey bars; Panel B, D, F). Figure 27: In vivo T-cell responses to chimeric proteins with different TMDs in C57BI6 mice
[0110] Panel A) shows in the left column the in vivo T-cell responses in mice immunized with RNA encoding the chimeric proteins as measured by induction of interferon y (IFN y) in the total T-cell population - separated by domain. In the right column, IFN y induction is shown for CD4+and CD8+cells, respectively. Panels B) and C) have the same structure, showing respective induction of interleukin-2 (IL-2) and tumor necrosis factor a (TNFa).
[0111] Figure 28: In vivo solenocyte and T cell response to BCG
[0112] Panel A) shows the in vivo T-cell responses in mice immunized with RNA encoding the chimeric proteins as measured by induction of interferon y (IFN y) for CD4+ cells (white bars) and CD8+ cells (grey bars) - separated by domain. Panels B) and C) have the same structure, showing respective induction of interleukin-2 (IL-2) and tumor necrosis factor o (TNFo). Panel D shows IFN y, IL-2 and TNFo induction in total splenocytes.
[0113] Figure 29-39: Visualization of IEDB and lioandomics-derived epitope hits of antigens and antigen fragments
[0114] Figures 29-39 show for the antigens or antigen fragments four subplots (labeled A-D) in the top-to-bottom order: A) "IEDB nCounts Epitopes" represent the total number of times that the corresponding epitope has been listed across multiple HLA alleles and studies in IEDB, B) "IEDB HLA alleles" shows a heatmap showing the number of times any particular HLA allele shown has been listed as being immunogenic together with an epitope in IEDB, C) "Ligandomics HLA alleles" shows a heatmap that highlights the number of times that the corresponding HLA class I allele was identified as presenting any ligandomics epitope, D) "String" represents a heatmap showing which region o f the antigen or antigen fragment has been used in one of the designed strings listed.
[0115] Figures 40-45: Domain structure of chimeric proteins
[0116] In Figures 40-45, respectively, panel A shows the domain structure of an exemplary chimeric protein disclosed herein. The chimeric protein comprises sequences from several antigens either in the form of full-length antigens or antigen fragments - separated from each other by polypeptide linkers. The exemplary chimeric protein comprises a signal peptide (sec) at its N-terminus, whereas a transmembrane domain (MITD) is located at its C-terminus.
[0117] In panel B, the figures respectively list the sequences and positions of individual domains within the domain structure of the exemplary chimeric protein.
[0118] Detailed Description
[0119] Although the present disclosure is further described in more detail below, it is to be understood that this disclosure is not limited to the particular methodologies, protocols and reagents described herein as these may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present disclosure which will be limited only by the appended claims. Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art.
[0120] In the following, the elements of the present disclosure will be described in more detail. These elements are listed with specific embodiments, however, it should be understood that they may be combined in any manner and in any number to create additional embodiments. The variously described examples and preferred embodiments should not be construed to limit the present disclosure to only the explicitly described embodiments. This description should be understood to support and encompass embodiments which combine the explicitly described embodiments with any number of the disclosed and / or preferred elements. Furthermore, any permutations and combinations of all described elements in this application should be considered disclosed by the description of the present application unless the context indicates otherwise.
[0121] For example, the present disclosure describes combinations of sequence molecules which may have different levels of sequence identity to a specified sequence, e.g., (i) sequence molecule A comprising the sequence of SEQ ID NO: a, or a sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the sequence of SEQ ID NO: a, (ii) sequence molecule B comprising the sequence of SEQ ID NO: b, or a sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the sequence of SEQ ID NO: b etc. It should be understood that the sequence molecules may be combined in any of the identity levels specified. In some embodiments, the sequence molecules are combined such that the identity levels are identical; e.g., (i) sequence molecule A comprising the sequence of SEQ ID NO: a, (ii) sequence molecule B comprising the sequence of SEQ ID NO: b etc., or (i) sequence molecule A comprising a sequence having at least 90% identity to the sequence of SEQ ID NO: a, (ii) sequence molecule B comprising a sequence having at least 90% identity to the sequence of SEQ ID NO: b etc. In some embodiments, the identity levels are independently selected and are partially or entirely different from each other, i.e., the sequence molecules are combined such that the identity levels are not identical; e.g., (i) sequence molecule A comprising the sequence of SEQ ID NO: a, (ii) sequence molecule B comprising a sequence having at least 90% identity to the sequence of SEQ ID NO: b etc., or (i) sequence molecule A comprising a sequence having at least 90% identity to the sequence of SEQ ID NO: a, (ii) sequence molecule B comprising a sequence having at least 85% identity to the sequence of SEQ ID NO: b etc.
[0122] The practice of the present disclosure will employ, unless otherwise indicated, conventional chemistry, biochemistry, cell biology, immunology, and recombinant DNA techniques which are explained in the literature in the field.
[0123] Throughout this specification and the claims which follow, unless the context requires otherwise, the word "comprise", and variations such as "comprises" and "comprising", will be understood to imply the inclusion of a stated feature, element, member, integer or step or group of features, elements, members, integers or steps but not the exclusion of any other feature, element, member, integer or step or group of features, elements, members, integers or steps. The term "consisting essentially of" limits the scope of a claim or disclosure to the specified features, elements, members, integers, or steps and those that do not materially affect the basic and novel characteristic(s) of the claim or disclosure. The term "consisting of" limits the scope of a claim or disclosure to the specified features, elements, members, integers, or steps. The term "comprising" encompasses the term "consisting essentially of" which, in turn, encompasses the term "consisting of". Thus, at each occurrence in the present application, the term "comprising" may be replaced with the term "consisting essentially of" or "consisting of". Likewise, at each occurrence in the present application, the term "consisting essentially of" may be replaced with the term "consisting of".
[0124] The terms "a", "an" and "the" and similar references used in the context of describing the present disclosure (especially in the context of the claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by the context.
[0125] All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by the context.
[0126] The use of any and all examples, or exemplary language e.g, "such as"), provided herein is intended merely to better illustrate the present disclosure and does not pose a limitation on the scope of the present disclosure otherwise claimed. No language in the specification should be construed as indicating any non-claimed element essential to the practice of the present disclosure. The term "optional" or "optionally" as used herein means that the subsequently described event, circumstance or condition may or may not occur, and that the description includes instances where said event, circumstance, or condition occurs and instances in which it does not occur.
[0127] Where used herein, "and / or" is to be taken as specific disclosure of each of the two specified features or components with or without the other. For example, "X and / or Y" is to be taken as specific disclosure of each of (i) X, (ii) Y, and (iii) X and Y, just as if each is set out individually herein.
[0128] In the context of the present disclosure, the term "about" denotes an interval of accuracy that the person of ordinary skill will understand to still ensure the technical effect of the feature in question. The term typically indicates deviation from the indicated numerical value by ±10%, ±5%, ±4%, ±3%, ±2%, ±1%, ±0.9%, ±0.8%, ±0.7%, ±0.6%, ±0.5%, ±0.4%, ±0.3%, ±0.2%, ±0.1%, ±0.05%, and for example ±0.01%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±10%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±5%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±4%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±3%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±2%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±1%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±0.9%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±0.8%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±0.7%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±0.6%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±0.5%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±0.4%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±0.3%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±0.2%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±0.1%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±0.05%. In some embodiments, "about" indicates deviation from the indicated numerical value by ±0.01%. As will be appreciated by the person of ordinary skill, the specific such deviation for a numerical value for a given technical effect will depend on the nature of the technical effect. For example, a natural or biological technical effect may generally have a larger such deviation than one for a man-made or engineering technical effect.
[0129] Recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into the specification as if it were individually recited herein.
[0130] Several documents are cited throughout the text of this specification. Each of the documents cited herein (including all patents, patent applications, scientific publications, manufacturer's specifications, instructions, etc.), whether supra or infra, are hereby incorporated by reference in their entirety. Nothing herein is to be construed as an admission that the invention is not entitled to antedate such disclosure by virtue of prior invention.
[0131] It should be noted for unambiguousness that whenever a sequence is referred to as being the sequence between the nucleotide at position x and the nucleotide at position y, the resulting sequence includes both the nucleotide at position x and the nucleotide at position y. Similarly, whenever a sequence is referred to as being the sequence between the amino acid at position x and the amino acid at position y, the resulting sequence includes both the amino acid at position x and the amino acid at position y. Moreover, while the sequences described herein, in particular in the sequence listing, refer to DNA molecules, it is clear that when it is stated in the description or the claims that an RNA comprises a nucleotide sequence as described herein, in particular in the sequence listing, the nucleotide sequence referred to is actually identical to the base-sequence of the DNA molecule described herein, in particular in the sequence listing, e.g., represented in a SEQ ID NO referred to, except that thymine is replaced by uracil.
[0132] In the following, definitions and embodiments will be provided which apply to all aspects of the present disclosure. Terms which are defined in the following have the meanings as defined, unless otherwise indicated. Any undefined terms have their art recognized meanings.
[0133] Mycobacterium tuberculosis (Mtb) is a non-motile, slowly growing and rod shaped (2-4 pm in length and 0.2-0.5 pm in width) bacterium. Mtb \s gram-positive, obligate aerobe, requires a host for growth and reproduction, and does not form spores.
[0134] The term "tuberculosis" or "TB" is used to describe the infection caused by infective agents from the genus "Mycobacterium ". Tuberculosis is a potentially fatal contagious disease that can affect almost any part of the body but is most frequently an infection of the lungs. While the majority of tuberculosis infections is caused by Mycobacterium tuberculosis, there are other Mycobacterium species that can cause tuberculosis as well. These species include Mycobacterium bovis, Mycobacterium caprae, Mycobacterium orygis, Mycobacterium africanum, Mycobacterium microti, Mycobacterium canetti and Mycobacterium pinnipedii. Mycobacterium tuberculosis and some other mycobacteria are transmitted by airborne droplet nuclei produced when an individual with active disease coughs, speaks, or sneezes. When inhaled, the droplet nuclei reach the alveoli of the lung. In susceptible individuals the organisms may then multiply and spread through lymphatics to the lymph nodes, and through the bloodstream to other sites such as the lung apices, bone marrow, kidneys, and meninges. Infections with other Mycobacterium species, such as Mycobacterium bovis or Mycobacterium caprae are also associated with the consumption of un-pasteurized milk from infected animals. The development of acquired immunity in 2 to 10 weeks results in a halt to bacterial multiplication. Lesions heal and the individual remains asymptomatic. Mycobacteria can remain dormant (latent TB) in the body after infection for years, concealed in the phagocytosed cells, and never develop into the disease. Such an individual is said to have tuberculous infection without disease, and will show a positive tuberculin test. The clinical status of latent TB is traditionally associated with the transition of Mtb to a dormant state in response to non-optimal growth conditions in vivo due to activation of the host immune response. Dormancy is a specific physiological state characterized by significant cessation of metabolic activity and growth, whereas resuscitation from dormancy is a process of restoring cell activity followed by bacterial multiplication, which in case of Mtb can lead to disease progression. The risk of developing active disease with clinical symptoms diminishes with time and may never occur, but is a lifelong risk. Approximately 5% of individuals with tuberculous infection progress to active disease.
[0135] Terms such as "reduce" or "inhibit" as used herein means the ability to cause an overall decrease, for example, of about 5% or greater, about 10% or greater, about 15% or greater, about 20% or greater, about 25% or greater, about 30% or greater, about 40% or greater, about 50% or greater, or about 75% or greater, in the level. The term "inhibit" or similar phrases includes a complete or essentially complete inhibition, i.e. a reduction to zero or essentially to zero.
[0136] Terms such as "enhance" as used herein means the ability to cause an overall increase, or enhancement, for example, by at least about 5% or greater, about 10% or greater, about 15% or greater, about 20% or greater, about 25% or greater, about 30% or greater, about 40% or greater, about 50% or greater, about 75% or greater, or about 100% or greater in the level.
[0137] "Physiological pH" as used herein refers to a pH of about 7.4. In some embodiments, physiological pH is from 7.3 to 7.5. In some embodiments, physiological pH is from 7.35 to 7.45. In some embodiments, physiological pH is 7.3, 7.35, 7.4, 7.45, or 7.5. As used in the present disclosure, "% w / v" refers to weight by volume percent, which is a unit of concentration measuring the amount of solute in grams (g) expressed as a percent of the total volume of solution in milliliters (mL).
[0138] As used in the present disclosure, "% by weight" refers to weight percent, which is a unit of concentration measuring the amount of a substance in grams (g) expressed as a percent of the total weight of the total composition in grams (g).
[0139] As used in the present disclosure, "mol %" is defined as the ratio of the number of moles of one component to the total number of moles of all components, multiplied by 100.
[0140] As used in the present disclosure, "mol % of the total lipid" is defined as the ratio of the number of moles of one lipid component to the total number of moles of all lipids, multiplied by 100. In this context, in some embodiments, the term "total lipid" includes lipids and lipid-like material.
[0141] The term "ionic strength" refers to the mathematical relationship between the number of different kinds of ionic species in a particular solution and their respective charges. Thus, ionic strength I is represented mathematically by the formula:
[0142] ' l iz‘2'ciin which c is the molar concentration of a particular ionic species and z the absolute value of its charge. The sum Z is taken over all the different kinds of ions (i) in solution.
[0143] According to the disclosure, the term "ionic strength" in some embodiments relates to the presence of monovalent ions. Regarding the presence of divalent ions, in particular divalent cations, their concentration or effective concentration (presence of free ions) due to the presence of chelating agents is, in some embodiments, sufficiently low so as to prevent degradation of the nucleic acid. In some embodiments, the concentration or effective concentration of divalent ions is below the catalytic level for hydrolysis of the phosphodiester bonds between nucleotides such as RNA nucleotides. In some embodiments, the concentration of free divalent ions is 20 pM or less. In some embodiments, there are no or essentially no free divalent ions.
[0144] "Osmolality" refers to the concentration of a particular solute expressed as the number of osmoles of solute per kilogram of solvent.
[0145] The term "lyophilizing" or "lyophilization" refers to the freeze-drying of a substance by freezing it and then reducing the surrounding pressure (e.g., below 15 Pa, such as below 10 Pa, below 5 Pa, or 1 Pa or less) to allow the frozen medium in the substance to sublimate directly from the solid phase to the gas phase. Thus, the terms "lyophilizing" and "freeze-drying" are used herein interchangeably.
[0146] The term "spray-drying" refers to spray-drying a substance by mixing (heated) gas with a fluid that is atomized (sprayed) within a vessel (spray dryer), where the solvent from the formed droplets evaporates, leading to a dry powder.
[0147] The term "reconstitute" relates to adding a solvent such as water to a dried product to return it to a liquid state such as its original liquid state. The term "recombinant" in the context of the present disclosure means "made through genetic engineering". In some embodiments, a "recombinant object" in the context of the present disclosure is not occurring naturally.
[0148] The term "naturally occurring" as used herein refers to the fact that an object can be found in nature. For example, a peptide or nucleic acid that is present in an organism (including viruses) and can be isolated from a source in nature and which has not been intentionally modified by man in the laboratory is naturally occurring. The term "found in nature" means "present in nature" and includes known objects as well as objects that have not yet been discovered and / or isolated from nature, but that may be discovered and / or isolated in the future from a natural source.
[0149] As used herein, the terms "room temperature" and "ambient temperature" are used interchangeably herein and refer to temperatures from at least about 15°C, e.g., from about 15°C to about 35°C, from about 15°C to about 30°C, from about 15°C to about 25°C, or from about 17°C to about 22°C. Such temperatures will include 15°C, 16°C, 17°C, 18°C, 19°C, 20°C, 21°C and 22°C.
[0150] The term "EDTA" refers to ethylenediaminetetraacetic acid disodium salt. All concentrations are given with respect to the EDTA disodium salt.
[0151] The term "cryoprotectant" relates to a substance that is added to a formulation in order to protect the active ingredients during the freezing stages.
[0152] The term "lyoprotectant" relates to a substance that is added to a formulation in order to protect the active ingredients during the drying stages.
[0153] According to the present disclosure, the term "peptide" refers to substances which comprise about two or more, about 3 or more, about 4 or more, about 6 or more, about 8 or more, about 10 or more, about 13 or more, about 16 or more, about 20 or more, and up to about 50, about 100 or about 150, consecutive amino acids linked to one another via peptide bonds. The term "polypeptide" refers to large peptides, in particular peptides having at least about 151 amino acids. "Peptides" and "polypeptides" are both protein molecules, although the terms "protein" and "polypeptide" are used herein usually as synonyms.
[0154] The term "biological activity" means the response of a biological system to a molecule. Such biological systems may be, for example, a cell or an organism. In some embodiments, such response is therapeutically or pharmaceutically useful.
[0155] The term "portion" refers to a fraction. With respect to a particular structure such as an amino acid sequence or protein the term "portion" thereof may designate a continuous or a discontinuous fraction of said structure.
[0156] The terms "part" and "fragment" are used interchangeably herein and refer to a continuous element. For example, a part of a structure such as an amino acid sequence or protein refers to a continuous element of said structure. When used in context of a composition, the term "part" means a portion of the composition. For example, a part of a composition may be any portion from 0.1% to 99.9% (such as 0.1%, 0.5%, 1%, 5%, 10%, 50%, 90%, or 99%) of said composition.
[0157] "Fragment", with reference to an amino acid sequence (peptide or polypeptide), relates to a part of an amino acid sequence, i.e. a sequence which represents the amino acid sequence shortened at the N-terminus and / or C-terminus. A fragment shortened at the C-terminus (N-terminal fragment) is obtainable, e.g., by translation of a truncated open reading frame that lacks the 3'-end of the open reading frame. A fragment shortened at the N-terminus (C -terminal fragment) is obtainable, e.g., by translation of a truncated open reading frame that lacks the 5'-end of the open reading frame, as long as the truncated open reading frame comprises a start codon that serves to initiate translation. A fragment of an amino acid sequence comprises, e.g., at least 50 %, at least 60 %, at least 70 %, at least 80%, at least 90% of the amino acid residues from an amino acid sequence. A fragment of an amino acid sequence comprises, e.g., at least 5, at least 6, at least 7, in particular at least 8, at least 10, at least 12, at least 15, at least 20, at least 30, at least 50, or at least 100 consecutive amino acids from an amino acid sequence. A fragment of an amino acid sequence comprises, e.g., a sequence of up to 8, in particular up to 10, up to 12, up to 15, up to 20, up to 30, up to 50, up to 80, up to 100, up to 150 or up to 200 consecutive amino acids of the amino acid sequence.
[0158] A "Mycobacterium tuberculosis antigen or immunogenic variant thereof represented by one or more antigen fragments and / or one or more full length antigens" as used herein refers to the full length Mycobacterium tuberculosis antigen or an immunogenic variant of the full length Mycobacterium tuberculosis antigen, or one or more fragments of the Mycobacterium tuberculosis antigen or an immunogenic variant of the Mycobacterium tuberculosis antigen, wherein the fragments may or may not be overlapping. An immunogenic variant of a Mycobacterium tuberculosis antigen or one or more fragments of a Mycobacterium tuberculosis antigen or an immunogenic variant of a Mycobacterium tuberculosis antigen are capable of inducing an immune response against the Mycobacterium tuberculosis antigen when delivered to a subject, e.g. in the form of a protein or an RNA transcribed by a cell of the subject. In some embodiments, a fragment of a Mycobacterium tuberculosis antigen or an immunogenic variant of a Mycobacterium tuberculosis antigen comprises at least one epitope, e.g., at least one T cell epitope, of a Mycobacterium tuberculosis antigen or an immunologically equivalent variant of said at least one epitope. In some embodiments, a fragment of a Mycobacterium tuberculosis antigen or an immunogenic variant of a Mycobacterium tuberculosis antigen comprises a fragment of, e.g., at least 5, at least 6, at least 7, in particular at least 8, at least 10, at least 12, at least 15, at least 20, at least 30, at least 50, or at least 100 consecutive amino acids of said Mycobacterium tuberculosis antigen or immunogenic variant of a Mycobacterium tuberculosis antigen.
[0159] If reference is made to RNA encoding a chimeric protein or encoding at least one full-length antigen or antigen fragment representing at least one Mycobacterium tuberculosis antigen or immunogenic variant thereof, such disclosure encompasses monocistronic and polycistronic RNAs.
[0160] If only one Mycobacterium tuberculosis antigen or immunogenic variant thereof is represented, the RNA may encode one or more fragments of the Mycobacterium tuberculosis antigen or immunogenic variant thereof. If the RNA encodes more than one fragment of the Mycobacterium tuberculosis antigen or immunogenic variant thereof, fragments of the Mycobacterium tuberculosis antigen or immunogenic variant thereof may be encoded by different open reading frames located on the same or on different RNA molecules.
[0161] If more than one Mycobacterium tuberculosis antigen or immunogenic variant thereof is represented, the RNA may encode one Mycobacterium tuberculosis antigen or immunogenic variant thereof as one or more fragments and, optionally, as an additional full length antigen, whereas the remaining Mycobacterium tuberculosis antigens or immunogenic variants thereof are encoded as the full length antigen and / or as one or more antigen fragments, respectively. In some embodiments, the RNA encodes the full-length antigen of each of the more than one Mycobacterium tuberculosis antigen or immunogenic variant thereof. In some embodiments, the RNA encodes one or more fragments of each of the more than one Mycobacterium tuberculosis antigen or immunogenic variant thereof. In some embodiments, the RNA encodes the full-length antigen of some of the more than one Mycobacterium tuberculosis antigen or immunogenic variant thereof and encodes one or more fragments of some of the more than one Mycobacterium tuberculosis antigen or immunogenic variant thereof, wherein the RNA may encode the full-length antigen as well as one or more fragments of the same Mycobacterium tuberculosis antigen or immunogenic variant thereof. The full-length antigens and / or fragments discussed above may be encoded by the same or different open reading frames located on the same or on different RNA molecules. The term "chimeric protein" is used herein as a synonym for "fusion protein" and means a protein comprising two or more subunits, such as a full-length antigen, antigen fragment and / or other functional amino acid sequence. Preferably, the fusion protein is a translational fusion between the two or more subunits. The translational fusion may be generated by genetically engineering the coding nucleotide sequence for one subunit in a reading frame with the coding nucleotide sequence of a further subunit. Subunits may be interspersed by a polypeptide linker.
[0162] The term "polypeptide linker" describes either a single amino acid or a sequence of amino acids connecting the amino acid sequences of two adjacent subunits in a linear amino acid sequence of a chimeric protein. A polypeptide linker may have a length of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15 or more amino acids. Preferably, the polypeptide linker is 10 amino acids in length. Exemplary linkers include glycine-serine-polypeptide linkers, glycine-proline-polypeptide linkers, and proline-alanine polypeptide linkers. In certain embodiments, the linker is a glycine-serine-polypeptide linker (GS linker), i.e., a peptide that consists of glycine and serine residues. In some embodiments, a GS linker comprises the amino acid sequence of SEQ ID NO: 56.
[0163] By "wild type" or "WT" or "native" herein is meant an amino acid sequence that is found in nature, including allelic variations and / or naturally occurring mutations. A wild type amino acid sequence, peptide or polypeptide has an amino acid sequence that has not been intentionally modified by man.
[0164] The term "non-native" as used herein in conjunction with amino acid sequences is meant to refer to amino acid sequences not found in nature, i.e., that have been intentionally modified by man - either in sequence or in sequence context. In one embodiment, a non-native signal peptide sequence fused or operatively linked to a Mycobacterium tuberculosis antigen denotes that said signal peptide in nature does not occur fused or operatively linked to said Mycobacterium tuberculosis antigen, either because said signal peptide can naturally be found fused or operatively linked only to other Mycobacterium tuberculosis antigens or only in other organisms, such as mammals, e.g. human, other bacteria besides Mycobacterium tuberculosis or viruses. Embodiments for such non-native signal peptides are provided herein. In another embodiment, a non-native signal peptide has been mutated in a purposeful manner (e.g., by random mutagenesis and targeted selection or by guided mutagenesis techniques, including, e.g., sequence synthesis) in order to obtain certain functional properties or to eliminate certain functional properties, resulting in a signal peptide structurally and functionally distinct from a signal peptide found in nature fused or operatively linked to the Mycobacterium tuberculosis antigen in question.
[0165] "Variant," as used herein and with reference to an amino acid sequence (peptide or polypeptide), is meant an amino acid sequence that differs from a parent amino acid sequence by virtue of at least one amino acid (e.g., a different amino acid, or a modification of the same amino acid). The parent amino acid sequence may be a naturally occurring or wild type (WT) amino acid sequence, or may be a modified version of a wild type amino acid sequence. In some embodiments, the variant amino acid sequence has at least one amino acid difference as compared to the parent amino acid sequence, e.g., from 1 to about 20 amino acid differences, such as from 1 to about 10 or from 1 to about 5 amino acid differences compared to the parent.
[0166] For the purposes of the present disclosure, "variants" of an amino acid sequence (peptide or polypeptide) may comprise amino acid insertion variants, amino acid addition variants, amino acid deletion variants and / or amino acid substitution variants. The term "variant" includes all mutants, splice variants, post-translationally modified variants, conformations, isoforms, allelic variants, species variants, and species homologs, in particular those which are naturally occurring. The term "variant" includes, in particular, fragments of an amino acid sequence.
[0167] Amino acid insertion variants comprise insertions of single or two or more amino acids in a particular amino acid sequence. In the case of amino acid sequence variants having an insertion, one or more amino acid residues are inserted into a particular site in an amino acid sequence, although random insertion with appropriate screening of the resulting product is also possible. Amino acid addition variants comprise amino- and / or carboxy-terminal fusions of one or more amino acids, such as 1, 2, 3, 5, 10, 20, 30, 50, or more amino acids. Amino acid deletion variants are characterized by the removal of one or more amino acids from the sequence, such as by removal of 1, 2, 3, 5, 10, 20, 30, 50, or more amino acids. The deletions may be in any position of the protein. Amino acid deletion variants that comprise the deletion at the N-terminal and / or C-terminal end of the protein are also called N-terminal and / or C- terminal truncation variants. Amino acid substitution variants are characterized by at least one residue in the sequence being removed and another residue being inserted in its place. Preference is given to the modifications being in positions in the amino acid sequence which are not conserved between homologous peptides or polypeptides and / or to replacing amino acids with other ones having similar properties. In some embodiments, amino acid changes in peptide and polypeptide variants are conservative amino acid changes, i.e., substitutions of similarly charged or uncharged amino acids. A conservative amino acid change involves substitution of one of a family of amino acids which are related in their side chains. Naturally occurring amino acids are generally divided into four families: acidic (aspartate, glutamate), basic (lysine, arginine, histidine), non-polar (alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), and uncharged polar (glycine, asparagine, glutamine, cysteine, serine, threonine, tyrosine) amino acids. Phenylalanine, tryptophan, and tyrosine are sometimes classified jointly as aromatic amino acids. In some embodiments, conservative amino acid substitutions include substitutions within the following groups:
[0168] - glycine, alanine;
[0169] - valine, isoleucine, leucine;
[0170] - aspartic acid, glutamic acid;
[0171] - asparagine, glutamine;
[0172] - serine, threonine;
[0173] - lysine, arginine; and
[0174] - phenylalanine, tyrosine.
[0175] In some embodiments the degree of similarity, such as identity between a given amino acid sequence and an amino acid sequence which is a variant of said given amino acid sequence, will be at least about 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%. In some embodiments, the degree of similarity or identity is given for an amino acid region which is at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90% or about 100% of the entire length of the reference amino acid sequence. For example, if the reference amino acid sequence consists of 200 amino acids, the degree of similarity or identity is given, e.g., for at least about 20, at least about 40, at least about 60, at least about 80, at least about 100, at least about 120, at least about 140, at least about 160, at least about 180, or about 200 amino acids, in some embodiments, continuous amino acids. In some embodiments, the degree of similarity or identity is given for the entire length of the reference amino acid sequence. The alignment for determining sequence similarity, such as sequence identity, can be done with art known tools, such as using the best sequence alignment, for example, using Align, using standard settings, preferably EMBOSS: : needle, Matrix: Blosum62, Gap Open 10.0, Gap Extend 0.5.
[0176] "Sequence similarity" indicates the percentage of amino acids that either are identical or that represent conservative amino acid substitutions. "Sequence identity" between two amino acid sequences indicates the percentage of amino acids that are identical between the sequences. "Sequence identity" between two nucleic acid sequences indicates the percentage of nucleotides that are identical between the sequences.
[0177] The terms "% identical" and "% identity" or similar terms are intended to refer, in particular, to the percentage of nucleotides or amino acids which are identical in an optimal alignment between the sequences to be compared. Said percentage is purely statistical, and the differences between the two sequences may be but are not necessarily randomly distributed over the entire length of the sequences to be compared. Comparisons of two sequences are usually carried out by comparing the sequences, after optimal alignment, with respect to a segment or "window of comparison", in order to identify local regions of corresponding sequences. The optimal alignment for a comparison may be carried out manually or with the aid of the local homology algorithm by Smith and Waterman, 1981, Ads App. Math. 2, 482, with the aid of the local homology algorithm by Neddleman and Wunsch, 1970, J. Mol. Biol. 48, 443, with the aid of the similarity search algorithm by Pearson and Lipman, 1988, Proc. Natl Acad. Sci. USA 88, 2444, or with the aid of computer programs using said algorithms (GAP, BESTFIT, FASTA, BLAST P, BLAST N and TFASTA in Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Drive, Madison, Wis.). In some embodiments, percent identity of two sequences is determined using the BLASTN or BLASTP algorithm, as available on the United States National Center for Biotechnology Information (NCBI) website (e.g, at blast.ncbi.nlm.nih.gov / Blast.cgi?PAGE_TYPE=BlastSearch&BLAST_SPEC=blast2seq&LINK_LOC=align2seq). In some embodiments, the algorithm parameters used for BLASTN algorithm on the NCBI website include: (i) Expect Threshold set to 10; (ii) Word Size set to 28; (iii) Max matches in a query range set to 0; (iv) Match / Mismatch Scores set to 1, - 2; (v) Gap Costs set to Linear; and (vi) the filter for low complexity regions being used. In some embodiments, the algorithm parameters used for BLASTP algorithm on the NCBI website include: (i) Expect Threshold set to 10; (ii) Word Size set to 3; (iii) Max matches in a query range set to 0; (iv) Matrix set to BLOSUM62; (v) Gap Costs set to Existence: 11 Extension: 1; and (vi) conditional compositional score matrix adjustment.
[0178] Percentage identity is obtained by determining the number of identical positions at which the sequences to be compared correspond, dividing this number by the number of positions compared (e.g, the number of positions in the reference sequence) and multiplying this result by 100.
[0179] In some embodiments, the degree of similarity or identity is given for a region which is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90% or about 100% of the entire length of the reference sequence. For example, if the reference nucleic acid sequence consists of 200 nucleotides, the degree of identity is given for at least about 100, at least about 120, at least about 140, at least about 160, at least about 180, or about 200 nucleotides, in some embodiments continuous nucleotides. In some embodiments, the degree of similarity or identity is given for the entire length of the reference sequence.
[0180] Homologous amino acid sequences exhibit according to the disclosure at least 40%, in particular at least 50%, at least 60%, at least 70%, at least 80%, at least 90% and, e.g., at least 95%, at least 98 or at least 99% identity of the amino acid residues.
[0181] The amino acid sequence variants described herein may readily be prepared by the skilled person, for example, by recombinant DNA manipulation. The manipulation of DNA sequences for preparing peptides or polypeptides having substitutions, additions, insertions or deletions, is described in detail in Molecular Cloning: A Laboratory Manual, 4thEdition, M.R. Green and J. Sambrook eds., Cold Spring Harbor Laboratory Press, Cold Spring Harbor 2012, for example. Furthermore, the peptides, polypeptides and amino acid variants described herein may be readily prepared with the aid of known peptide synthesis techniques such as, for example, by solid phase synthesis and similar methods. In some embodiments, a fragment or variant of an amino acid sequence (peptide or polypeptide) is a "functional fragment" or "functional variant". The term "functional fragment" or "functional variant" of an amino acid sequence relates to any fragment or variant exhibiting one or more functional properties identical or similar to those of the amino acid sequence from which it is derived, i.e., it is functionally equivalent. With respect to antigens or antigenic sequences, one particular function is one or more immunogenic activities displayed by the amino acid sequence from which the fragment or variant is derived. The term "functional fragment" or "functional variant", as used herein, in particular refers to a variant molecule or sequence that comprises an amino acid sequence that is altered by one or more amino acids compared to the amino acid sequence of the parent molecule or sequence and that is still capable of fulfilling one or more of the functions of the parent molecule or sequence, e.g, inducing an immune response. In some embodiments, the modifications in the amino acid sequence of the parent molecule or sequence do not significantly affect or alter the characteristics of the molecule or sequence. In different embodiments, the function of the functional fragment or functional variant may be reduced but still significantly present, e.g., function of the functional fragment or functional variant may be at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% of the parent molecule or sequence. However, in other embodiments, function of the functional fragment or functional variant may be enhanced compared to the parent molecule or sequence.
[0182] An amino acid sequence (peptide or polypeptide) "derived from" a designated amino acid sequence (peptide or polypeptide) refers to the origin of the first amino acid sequence. In some embodiments, the amino acid sequence which is derived from a particular amino acid sequence has an amino acid sequence that is identical, essentially identical or homologous to that particular sequence or a fragment thereof. Amino acid sequences derived from a particular amino acid sequence may be variants of that particular sequence or a fragment thereof. For example, it will be understood by one of ordinary skill in the art that the antigens suitable for use herein may be altered such that they vary in sequence from the naturally occurring or native sequences from which they were derived, while retaining the desirable activity of the native sequences.
[0183] In some embodiments, "isolated" means removed (e.g., purified) from the natural state or from an artificial composition, such as a composition from a production process. For example, a nucleic acid, peptide or polypeptide naturally present in a living animal is not "isolated", but the same nucleic acid, peptide or polypeptide partially or completely separated from the coexisting materials of its natural state is "isolated". An isolated nucleic acid, peptide or polypeptide can exist in substantially purified form, or can exist in a non-native environment such as, for example, a host cell.
[0184] The term "transfection" relates to the introduction of nucleic acids, in particular RNA, into a cell. For purposes of the present disclosure, the term "transfection" also includes the introduction of a nucleic acid into a cell or the uptake of a nucleic acid by such cell, wherein the cell may be present in a subject, e.g., a patient, or the cell may be in vitro, e.g., outside of a patient. Thus, according to the present disclosure, a cell for transfection of a nucleic acid described herein can be present in vitro or in vivo, e.g. the cell can form part of an organ, a tissue and / or the body of a patient. According to the disclosure, transfection can be transient or stable. For some applications of transfection, it is sufficient if the transfected genetic material is only transiently expressed. RNA can be transfected into cells to transiently express its coded protein. Since the nucleic acid introduced in the transfection process is usually not integrated into the nuclear genome, the foreign nucleic acid will be diluted through mitosis or degraded. Cells allowing episomal amplification of nucleic acids greatly reduce the rate of dilution. If it is desired that the transfected nucleic acid actually remains in the genome of the cell and its daughter cells, a stable transfection must occur. Such stable transfection can be achieved by using virus-based systems or transposon-based systems for transfection, for example. Generally, nucleic acid encoding antigen is transiently transfected into cells. RNA can be transfected into cells to transiently express its coded protein.
[0185] The disclosure includes analogs of a peptide or polypeptide. According to the present disclosure, an analog of a peptide or polypeptide is a modified form of said peptide or polypeptide from which it has been derived and has at least one functional property of said peptide or polypeptide. Eg., a pharmacological active analog of a peptide or polypeptide has at least one of the pharmacological activities of the peptide or polypeptide from which the analog has been derived. Such modifications include any chemical modification and comprise single or multiple substitutions, deletions and / or additions of any molecules associated with the peptide or polypeptide, such as carbohydrates, lipids and / or peptides or polypeptides. In some embodiments, "analogs" of peptides or polypeptides include those modified forms resulting from glycosylation, acetylation, phosphorylation, amidation, palmitoylation, myristoylation, isoprenylation, lipidation, alkylation, derivatization, introduction of protective / blocking groups, proteolytic cleavage or binding to an antibody or to another cellular ligand. The term "analog" also extends to all functional chemical equivalents of said peptides and polypeptides.
[0186] As used herein, the terms "linked", "fused", or "fusion" are used interchangeably. These terms refer to the joining together of two or more elements or components or domains.
[0187] As used herein "endogenous" refers to any material from or produced inside an organism, cell, tissue or system.
[0188] As used herein, the term "exogenous" refers to any material introduced from or produced outside an organism, cell, tissue or system.
[0189] According to various embodiments of the present disclosure, a nucleic acid such as RNA encoding a peptide or polypeptide is taken up by or introduced, i.e. transfected or transduced, into a cell which cell may be present in vitro or in a subject, resulting in expression of said peptide or polypeptide. The cell may, e.g., express the encoded peptide or polypeptide intracellularly e.g. in the cytoplasm and / or in the nucleus), may secrete the encoded peptide or polypeptide, and / or may express it on the surface. In some embodiments, the cell secretes the encoded peptide or polypeptide.
[0190] According to the present disclosure, terms such as "nucleic acid expressing" and "nucleic acid encoding" or similar terms are used interchangeably herein and with respect to a particular peptide or polypeptide mean that the nucleic acid, if present in the appropriate environment, e.g. within a cell, can be expressed to produce said peptide or polypeptide.
[0191] In particular, the term "encoding" refers to the inherent property of specific sequences of nucleotides in a polynucleotide, such as a gene, a cDNA, or an RNA (in particular, mRNA), to serve as templates for synthesis of other polymers and macromolecules in biological processes having either a defined sequence of nucleotides (i.e., rRNA, tRNA and mRNA) or a defined sequence of amino acids and the biological properties resulting therefrom. Thus, a gene encodes a protein if transcription and translation of mRNA corresponding to that gene produces the protein in a cell or other biological system. Both the coding strand, the nucleotide sequence of which is identical to the mRNA sequence and is usually provided in sequence listings, and the non-coding strand, used as the template for transcription of a gene or cDNA, can be referred to as encoding the protein or other product of that gene or cDNA.
[0192] In this respect, an "open reading frame" or "ORF" is a continuous stretch of codons beginning with a start codon and ending with a stop codon.
[0193] The term "expression" as used herein includes the transcription and / or translation of a particular nucleotide sequence. In the context of the present disclosure, the term "transcription" relates to a process, wherein the genetic code in a DNA sequence is transcribed into RNA (especially mRNA). Subsequently, the RNA may be translated into peptide or polypeptide.
[0194] With respect to RNA, the term "expression" or "translation" relates to the process in the ribosomes of a cell by which a strand of mRNA directs the assembly of a sequence of amino acids to make a peptide or polypeptide.
[0195] A medical preparation, in particular kit, described herein may comprise instructional material or instructions. As used herein, "instructional material" or "instructions" includes a publication, a recording, a diagram, or any other medium of expression which can be used to communicate the usefulness of the compositions and methods of the present disclosure. The instructional material of the kit of the present disclosure may, for example, be affixed to a container which contains the compositions / formulations of the present disclosure or be shipped together with a container which contains the compositions / formulations. Alternatively, the instructional material may be shipped separately from the container with the intention that the instructional material and the compositions be used cooperatively by the recipient.
[0196] The term "set", e.g., as used herein in the context of "set of full-length antigens and antigen fragments", means more than 1, e.g., 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, or 8 or more.
[0197] The term "at least one" as used herein in the context of "at least one RNA molecule" means 1 or more, e.g., 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, or 8 or more. In some embodiments, the term "at least one" refers to 1, 2, 3, 4, 5, 6, 7, or 8. In some embodiments, "at least one RNA molecule" refers to a set of RNA molecules, e.g., 2, 3, 4 or more RNA molecules, wherein each RNA molecule encodes an amino acid sequence comprising at least one full-length Mtb antigen or antigen fragment, immunogenic variants or fragments thereof, e.g., an amino acid sequence comprising two different Mtb antigens, immunogenic variants or fragments thereof. In some embodiments, such at least one RNA molecule or set of RNA molecules comprises the RNA molecules in a mixtures, which mixture may be obtainable by transcribing in a common reaction a mixture of DNA templates encoding said RNA molecules.
[0198] Prodrugs of a particular compound described herein are those compounds that upon administration to an individual undergo chemical conversion under physiological conditions to provide the particular compound. Additionally, prodrugs can be converted to the particular compound by chemical or biochemical methods in an ex vivo environment. For example, prodrugs can be slowly converted to the particular compound when, for example, placed in a transdermal patch reservoir with a suitable enzyme or chemical reagent. Exemplary prodrugs are esters (using an alcohol or a carboxy group contained in the particular compound) or amides (using an amino or a carboxy group contained in the particular compound) which are hydrolyzable in vivo. Specifically, any amino group which is contained in the particular compound and which bears at least one hydrogen atom can be converted into a prodrug form. Typical N-prodrug forms include carbamates, Mannich bases, enamines, and enaminones.
[0199] In the present specification, a structural formula of a compound may represent a certain isomer of said compound. It is to be understood, however, that the present disclosure includes all isomers such as geometrical isomers, optical isomers based on an asymmetrical carbon, stereoisomers, tautomers and the like which occur structurally and isomer mixtures and is not limited to the description of the formula. Furthermore, in the present specification, a structural formula of a compound may represent a specific salt and / or solvate of said compound. It is to be understood, however, that the present disclosure includes all salts (e.g., pharmaceutically acceptable salts) and solvates (e.g., hydrates) and is not limited to the description of the specific salt and / or solvate.
[0200] "Isomers" are compounds having the same molecular formula but differ in structure ("structural isomers") or in the geometrical (spatial) positioning of the functional groups and / or atoms ("stereoisomers"). "Enantiomers" are a pair of stereoisomers which are non-superimposable mirror-images of each other. A "racemic mixture" or "racemate" contains a pair of enantiomers in equal amounts and is denoted by the prefix (±). "Diastereomers" are stereoisomers which are non-superimposable and which are not mirror-images of each other. "Tautomers" are structural isomers of the same chemical substance that spontaneously and reversibly interconvert into each other, even when pure, due to the migration of individual atoms or groups of atoms; i.e., the tautomers are in a dynamic chemical equilibrium with each other. An example of tautomers are the isomers of the keto-enol-tautomerism. "Conformers" are stereoisomers that can be interconverted just by rotations about formally single bonds, and include - in particular - those leading to different 3-dimentional forms of (hetero)cyclic rings, such as chair, half-chair, boat, and twist-boat forms of cyclohexane.
[0201] The term "solvate" as used herein refers to an addition complex of a dissolved material in a solvent (such as an organic solvent (e.g., an aliphatic alcohol (such as methanol, ethanol, n-propanol, isopropanol), acetone, acetonitrile, ether, and the like), water or a mixture of two or more of these liquids), wherein the addition complex exists in the form of a crystal or mixed crystal. The amount of solvent contained in the addition complex may be stoichiometric or non- stoichiometric. A "hydrate" is a solvate wherein the solvent is water.
[0202] In isotopically labeled compounds one or more atoms are replaced by a corresponding atom having the same number of protons but differing in the number of neutrons. For example, a hydrogen atom may be replaced by a deuterium or tritium atom. Exemplary isotopes which can be used in the present disclosure include deuterium, tritium,nC,13C,14C,15N,18F,32P,32S,35S,35CI, and125I.
[0203] The term "average diameter" refers to the mean hydrodynamic diameter of particles as measured by dynamic light scattering (DLS) with data analysis using the so-called cumulant algorithm, which provides as results the so-called with the dimension of a length, and the polydispersity index (PDI), which is dimensionless (Koppel, D., J. Chem. Phys. 57, 1972, pp 4814-4820, ISO 13321). Here "average diameter", "diameter" or "size" for particles is used synonymously with this value of the
[0204] In some embodiments, the "polydispersity index" is calculated based on dynamic light scattering measurements by the so-called cumulant analysis as mentioned in the definition of the "average diameter". Under certain prerequisites, it can be taken as a measure of the size distribution of an ensemble of nanoparticles.
[0205] The "radius of gyration" (abbreviated herein as Rg) of a particle about an axis of rotation is the radial distance of a point from the axis of rotation at which, if the whole mass of the particle is assumed to be concentrated, its moment of inertia about the given axis would be the same as with its actual distribution of mass. Mathematically, Rgis the root mean square distance of the particle's components from either its center of mass or a given axis. For example, for a macromolecule composed of n mass elements, of masses mi (J= 1, 2, 3, ..., ri), located at fixed distances s / from the center of mass, Rg is the square-root of the mass average of s,2over all mass elements and can be calculated as follows:
[0206] The radius of gyration can be determined or calculated experimentally, e.g., by using light scattering. In particular, for small scattering vectors q the structure function S is defined as follows: wherein N is the number of components (Guinier's law).
[0207] The "hydrodynamic radius" (which is sometimes called "Stokes radius" or "Stokes-Einstein radius") of a particle is the radius of a hypothetical hard sphere that diffuses at the same rate as said particle. The hydrodynamic radius is related to the mobility of the particle, taking into account not only size but also solvent effects. For example, a smaller charged particle with stronger hydration may have a greater hydrodynamic radius than a larger charged particle with weaker hydration. This is because the smaller particle drags a greater number of water molecules with it as it moves through the solution. Since the actual dimensions of the particle in a solvent are not directly measurable, the hydrodynamic radius may be defined by the Stokes-Einstein equation: wherein B is the Boltzmann constant; 7~is the temperature; g is the viscosity of the solvent; and D is the diffusion coefficient. The diffusion coefficient can be determined experimentally, e.g., by using dynamic light scattering (DLS). Thus, one procedure to determine the hydrodynamic radius of a particle or a population of particles (such as the hydrodynamic radius of particles contained in a sample or control composition as disclosed herein or the hydrodynamic radius of a particle peak obtained from subjecting such a sample or control composition to field-flow fractionation) is to measure the DLS signal of said particle or population of particles (such as DLS signal of particles contained in a sample or control composition as disclosed herein or the DLS signal of a particle peak obtained from subjecting such a sample or control composition to field-flow fractionation).
[0208] The expression "light scattering" as used herein refers to the physical process where light is forced to deviate from a straight trajectory by one or more paths due to localized non-uniformities in the medium through which the light passes.
[0209] The term "UV" means ultraviolet and designates a band of the electromagnetic spectrum with a wavelength from 10 nm to 400 nm, i.e., shorter than that of visible light but longer than X-rays.
[0210] The expression "multi-angle light scattering" or "MALS" as used herein relates to a technique for measuring the light scattered by a sample into a plurality of angles. "Multi-angle" means in this respect that scattered light can be detected at different discrete angles as measured, for example, by a single detector moved over a range including the specific angles selected or an array of detectors fixed at specific angular locations. In certain embodiments, the light source used in MALS is a laser source (MALLS: multi-angle laser light scattering). Based on the MALS signal of a composition comprising particles and by using an appropriate formalism e.g, Zimm plot, Berry plot, or Debye plot), it is possible to determine the radius of gyration (Rg) and, thus, the size of said particles. Preferably, the Zimm plot is a graphical presentation using the following equation: wherein cis the mass concentration of the particles in the solvent (g / mL); A2 is the second virial coefficient (mol-mL / g2); P(0) is a form factor relating to the dependence of scattered light intensity on angle; Re is the excess Rayleigh ratio (cm1); and A* is an optical constant that is equal to 4n2q0(d / 7 / dc)2Xo-4 / VA-1, where q0is the refractive index of the solvent at the incident radiation (vacuum) wavelength, Xois the incident radiation (vacuum) wavelength (nm), Ne is Avogadro's number (mol1), and d / 7 / dc is the differential refractive index increment (mL / g) (cf., e.g., Buchholz et al. (Electrophoresis 22 (2001), 4118-4128); B.H. Zimm (J. Chem. Phys. 13 (1945), 141; P. Debye (J. Appl. Phys. 15 (1944): 338; and W. Burchard (Anal. Chem. 75 (2003), 4279-4291). Preferably, the Berry plot is calculated using the following term or the reciprocal thereof: wherein c, Re and K* are as defined above. Preferably, the Debye plot is calculated using the following term or the reciprocal thereof: wherein c, / fa and A"*are as defined above.
[0211] The expression "dynamic light scattering" or "DLS" as used herein refers to a technique to determine the size and size distribution profile of particles, in particular with respect to the hydrodynamic radius of the particles. A monochromatic light source, usually a laser, is shot through a polarizer and into a sample. The scattered light then goes through a second polarizer where it is detected and the resulting image is projected onto a screen. The particles in the solution are being hit with the light and diffract the light in all directions. The diffracted light from the particles can either interfere constructively (light regions) or destructively (dark regions). This process is repeated at short time intervals and the resulting set of speckle patterns are analyzed by an autocorrelator that compares the intensity of light at each spot over time.
[0212] The expression "static light scattering" or "SLS" as used herein refers to a technique to determine the size and size distribution profile of particles, in particular with respect to the radius of gyration of the particles, and / or the molar mass of particles. A high-intensity monochromatic light, usually a laser, is launched in a solution containing the particles. One or many detectors are used to measure the scattering intensity at one or many angles. The angular dependence is needed to obtain accurate measurements of both molar mass and size for all macromolecules of radius. Hence simultaneous measurements at several angles relative to the direction of incident light, known as multi-angle light scattering (MALS) or multi-angle laser light scattering (MALLS), is generally regarded as the standard implementation of static light scattering.
[0213] Nucleic Acids
[0214] The term "nucleic acid" comprises deoxyribonucleic acid (DNA), ribonucleic acid (RNA), combinations thereof, and modified forms thereof. The term comprises genomic DNA, cDNA, mRNA, recombinantly produced and chemically synthesized molecules. In some embodiments, a nucleic acid is DNA. In some embodiments, a nucleic acid is RNA. In some embodiments, a nucleic acid is a mixture of DNA and RNA. A nucleic acid may be present as a single-stranded or double-stranded and linear or covalently circularly closed molecule. A nucleic acid can be isolated. The term "isolated nucleic acid" means, according to the present disclosure, that the nucleic acid (i) was amplified in vitro, for example via polymerase chain reaction (PCR) for DNA or in vitro transcription (using, e.g., an RNA polymerase) for RNA, (ii) was produced recombinantly by cloning, (iii) was purified, for example, by cleavage and separation by gel electrophoresis, or (iv) was synthesized, for example, by chemical synthesis.
[0215] The term "nucleoside" (abbreviated herein as "N") relates to compounds which can be thought of as nucleotides without a phosphate group. While a nucleoside is a nucleobase linked to a sugar e.g, ribose or deoxyribose), a nucleotide is composed of a nucleoside and one or more phosphate groups. Examples of nucleosides include cytidine, uridine, pseudouridine, adenosine, and guanosine. The five standard nucleosides which usually make up naturally occurring nucleic acids are uridine, adenosine, thymidine, cytidine and guanosine. The five nucleosides are commonly abbreviated to their one letter codes U, A, T, C and G, respectively. However, thymidine is more commonly written as "dT" ("d" represents "deoxy") as it contains a 2'-deoxyribofuranose moiety rather than the ribofuranose ring found in uridine. This is because thymidine is found in deoxyribonucleic acid (DNA) and not ribonucleic acid (RNA). Conversely, uridine is found in RNA and not DNA. The remaining three nucleosides may be found in both RNA and DNA. In RNA, they would be represented as A, C and G, whereas in DNA they would be represented as dA, dC and dG.
[0216] A modified purine (A or G) or pyrimidine (C, T, or U) base moiety is, in some embodiments, modified by one or more alkyl groups, e.g., one or more CH alkyl groups, e.g., one or more methyl groups. Particular examples of modified purine or pyrimidine base moieties include N7-alkyl-guanine, N5-alkyl-adenine, 5-alkyl-cytosine, 5-alkyl-uracil, and N(l)- alkyl-uracil, such as N7-CI-4 alkyl-guanine, N5-CI-4 alkyl-adenine, 5-C1-4 alkyl-cytosine, 5-C1-4 alkyl-uracil, and N(1)-CM alkyl-uracil, preferably N7-methyl-guanine, N5-methyl-adenine, 5-methyl-cytosine, 5-methyl-uracil, and N(l)-methyl- uracil.
[0217] DNA
[0218] Herein, the term "DNA" relates to a nucleic acid molecule which is entirely or at least substantially composed of deoxyribonucleotide residues. In preferred embodiments, the DNA contains all or a majority of deoxyribonucleotide residues. As used herein, "deoxyribonucleotide" refers to a nucleotide which lacks a hydroxyl group at the 2'-position of a p-D-ribofuranosyl group. DNA encompasses without limitation, double stranded DNA, single stranded DNA, isolated DNA such as partially purified DNA, essentially pure DNA, synthetic DNA, recombinantly produced DNA, as well as modified DNA that differs from naturally occurring DNA by the addition, deletion, substitution and / or alteration of one or more nucleotides. Such alterations may refer to addition of non-nucleotide material to internal DNA nucleotides or to the end(s) of DNA. It is also contemplated herein that nucleotides in DNA may be non-standard nucleotides, such as chemically synthesized nucleotides or ribonucleotides. For the present disclosure, these altered DNAs are considered analogs of naturally-occurring DNA. A molecule contains "a majority of deoxyribonucleotide residues" if the content of deoxyribonucleotide residues in the molecule is more than 50% (such as at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%), based on the total number of nucleotide residues in the molecule. The total number of nucleotide residues in a molecule is the sum of all nucleotide residues (irrespective of whether the nucleotide residues are standard (Ze., naturally occurring) nucleotide residues or analogs thereof).
[0219] DNA may be recombinant DNA and may be obtained by cloning of a nucleic acid, in particular cDNA. The cDNA may be obtained by reverse transcription of RNA.
[0220] RNA
[0221] The term "RNA" relates to a nucleic acid molecule which includes ribonucleotide residues. In preferred embodiments, the RNA contains all or a majority of ribonucleotide residues. As used herein, "ribonucleotide" refers to a nucleotide with a hydroxyl group at the 2'-position of a p-D-ribofuranosyl group. RNA encompasses without limitation, double stranded RNA, single stranded RNA, isolated RNA such as partially purified RNA, essentially pure RNA, synthetic RNA, recombinantly produced RNA, as well as modified RNA that differs from naturally occurring RNA by the addition, deletion, substitution and / or alteration of one or more nucleotides. Such alterations may refer to addition of non- nucleotide material to internal RNA nucleotides or to the end(s) of RNA. It is also contemplated herein that nucleotides in RNA may be non-standard nucleotides, such as chemically synthesized nucleotides or deoxynucleotides. For the present disclosure, these altered / modified nucleotides can be referred to as analogs of naturally occurring nucleotides, and the corresponding RNAs containing such altered / modified nucleotides (Ze., altered / modified RNAs) can be referred to as analogs of naturally occurring RNAs. A molecule contains "a majority of ribonucleotide residues" if the content of ribonucleotide residues in the molecule is more than 50% (such as at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%), based on the total number of nucleotide residues in the molecule. The total number of nucleotide residues in a molecule is the sum of all nucleotide residues (irrespective of whether the nucleotide residues are standard (Ze., naturally occurring) nucleotide residues or analogs thereof).
[0222] "RNA" includes mRNA, tRNA, ribosomal RNA (rRNA), small nuclear RNA (snRNA), self-amplifying RNA (saRNA), transamplifying RNA (taRNA), single-stranded RNA (ssRNA), dsRNA, inhibitory RNA (such as antisense ssRNA, small interfering RNA (siRNA), or microRNA (miRNA)), activating RNA (such as small activating RNA) and immunostimulatory RNA (isRNA). In some embodiments, "RNA" refers to mRNA.
[0223] The term "in vitro transcription" or "IVT" as used herein means that the transcription (Zevthe generation of RNA) is conducted in a cell-free manner. Ze., IVT does not use living / cultured cells but rather the transcription machinery extracted from cells e.g, cell lysates or the isolated components thereof, including an RNA polymerase (preferably T7, T3 or SP6 polymerase)).
[0224] According to the present disclosure, the term '"RNA" includes "mRNA". According to the present disclosure, the term "mRNA" means "messenger-RNA" and includes a "transcript" which may be generated by using a DNA template. Generally, mRNA encodes a peptide or polypeptide. mRNA is single-stranded but may contain self-complementary sequences that allow parts of the mRNA to fold and pair with itself to form double helices.
[0225] According to the present disclosure, "dsRNA" means double-stranded RNA and is RNA with two partially or completely complementary strands.
[0226] In preferred embodiments of the present disclosure, the mRNA relates to an RNA transcript which encodes a peptide or polypeptide.
[0227] In some embodiments, the mRNA which preferably encodes a peptide or polypeptide has a length of at least 45 nucleotides (such as at least 60, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1,000, at least 1,500, at least 2,000, at least 2,500, at least 3,000, at least 3,500, at least 4,000, at least 4,500, at least 5,000, at least 6,000, at least 7,000, at least 8,000, at least 9,000 nucleotides), preferably up to 15,000, such as up to 14,000, up to 13,000, up to 12,000 nucleotides, up to 11,000 nucleotides or up to 10,000 nucleotides.
[0228] As established in the art, mRNA generally contains a 5' untranslated region (5'-UTR), a peptide / polypeptide coding region and a 3' untranslated region (3'-UTR). In some embodiments, the mRNA is produced by in vitro transcription or chemical synthesis. In some embodiments, the mRNA is produced by in vitro transcription using a DNA template. The in vitro transcription methodology is known to the skilled person; cf., e.g., Molecular Cloning: A Laboratory Manual, 4thEdition, M.R. Green and J. Sambrook eds., Cold Spring Harbor Laboratory Press, Cold Spring Harbor 2012. Furthermore, a variety of in vitro transcription kits is commercially available, e.g., from Thermo Fisher Scientific (such as TranscriptAid™ T7 kit, MEGAscript® T7 kit, MAXIscript®), New England BioLabs Inc. (such as HiScribe™ T7 kit, HiScribe™ T7 ARCA mRNA kit), Promega (such as RiboMAX™, HeLaScribe®, Riboprobe® systems), Jena Bioscience (such as SP6 or T7 transcription kits), and Epicentre (such as AmpliScribe™). For providing modified mRNA, correspondingly modified nucleotides, such as modified naturally occurring nucleotides, non-naturally occurring nucleotides and / or modified non-naturally occurring nucleotides, can be incorporated during synthesis (preferably in vitro transcription), or modifications can be effected in and / or added to the mRNA after transcription.
[0229] In some embodiments, RNA is in vitro transcribed RNA (IVT-RNA) and may be obtained by in vitro transcription of an appropriate DNA template. The promoter for controlling transcription can be any promoter for any RNA polymerase. Particular examples of RNA polymerases are the T7, T3, and SP6 RNA polymerases. Preferably, the in vitro transcription is controlled by a T7 or SP6 promoter. A DNA template for in vitro transcription may be obtained by cloning of a nucleic acid, in particular cDNA, and introducing it into an appropriate vector for in vitro transcription. The cDNA may be obtained by reverse transcription of RNA.
[0230] In some embodiments of the present disclosure, the RNA is "replicon RNA" or simply a "replicon", in particular "selfreplicating RNA" or "self-amplifying RNA". In certain embodiments, the replicon or self-replicating RNA is derived from or comprises elements derived from an ssRNA virus, in particular a positive-stranded ssRNA virus such as an alphavirus. Alphaviruses are typical representatives of positive-stranded RNA viruses. Alphaviruses replicate in the cytoplasm of infected cells (for review of the alphaviral life cycle see Jose et at, Future Microbiol., 2009, vol. 4, pp. 837-856). The total genome length of many alphaviruses typically ranges between 11,000 and 12,000 nucleotides, and the genomic RNA typically has a 5'-cap, and a 3' poly(A) tail. The genome of alphaviruses encodes non-structural proteins (involved in transcription, modification and replication of viral RNA and in protein modification) and structural proteins (forming the virus particle). There are typically two open reading frames (ORFs) in the genome. The four non-structural proteins (nsPl-nsP4) are typically encoded together by a first ORF beginning near the 5' terminus of the genome, while alphavirus structural proteins are encoded together by a second ORF which is found downstream of the first ORF and extends near the 3' terminus of the genome. Typically, the first ORF is larger than the second ORF, the ratio being roughly 2: 1. In cells infected by an alphavirus, only the nucleic acid sequence encoding non-structural proteins is translated from the genomic RNA, while the genetic information encoding structural proteins is translatable from a subgenomic transcript, which is an RNA molecule that resembles eukaryotic messenger RNA (mRNA; Gould et al, 2010, Antiviral Res., vol. 87 pp. 111-124). Following infection, i.e. at early stages of the viral life cycle, the (+) stranded genomic RNA directly acts like a messenger RNA for the translation of the open reading frame encoding the non- structural poly-protein (nsP1234).
[0231] Alphavirus-derived vectors have been proposed for delivery of foreign genetic information into target cells or target organisms. In simple approaches, the open reading frame encoding alphaviral structural proteins is replaced by an open reading frame encoding a protein of interest. Alphavirus-based trans-replication (trans-amplification) systems rely on alphavirus nucleotide sequence elements on two separate nucleic acid molecules: one nucleic acid molecule encodes a viral replicase, and the other nucleic acid molecule is capable of being replicated by said replicase in trans (hence the designation trans-replication system). Trans-replication requires the presence of both these nucleic acid molecules in a given host cell. The nucleic acid molecule capable of being replicated by the replicase in trans must comprise certain alphaviral sequence elements to allow recognition and RNA synthesis by the alphaviral replicase.
[0232] In some embodiments of the present disclosure, the RNA (in particular, mRNA) described herein (e.g., contained in the compositions / formulations of the present disclosure and / or used in the methods of the present disclosure) contains one or more modifications, e.g., in order to increase its stability and / or increase translation efficiency and / or decrease immunogenicity and / or decrease cytotoxicity. For example, in order to increase expression of the RNA (in particular, mRNA), it may be modified within the coding region, i.e., the sequence encoding the expressed peptide or polypeptide, preferably without altering the sequence of the expressed peptide or polypeptide. Such modifications are described, for example, in WO 2007 / 036366 and PCT / EP2019 / 056502, and include the following: a 5'-cap structure; an extension or truncation of the naturally occurring poly(A) tail; an alteration of the 5'- and / or 3'-untranslated regions (UTR) such as introduction of a UTR which is not related to the coding region of said RNA; the replacement of one or more naturally occurring nucleotides with synthetic nucleotides; and codon optimization (e.g, to alter, preferably increase, the GC content of the RNA). A combination of the above described modifications, i.e., incorporation of a 5'-cap structure, incorporation of a poly-A sequence, unmasking of a poly-A sequence, alteration of the 5'- and / or 3'-UTR (such as incorporation of one or more 3'-UTRs), replacing one or more naturally occurring nucleotides with synthetic nucleotides (e.g, 5-methylcytidine for cytidine and / or pseudouridine (QJ) or N(l)-methylpseudouridine (mlQJ) or 5-methyluridine (m5U) for uridine), and codon optimization, has a synergistic influence on the stability of RNA (preferably mRNA) and increase in translation efficiency. Thus, in some embodiments, the RNA (in particular, mRNA) described in the present disclosure contains a combination of at least two, at least three, at least four or all five of the above-mentioned modifications, i.e., (i) incorporation of a 5'-cap structure, (ii) incorporation of a poly-A sequence, unmasking of a poly- A sequence; (iii) alteration of the 5'- and / or 3'-UTR (such as incorporation of one or more 3'-UTRs); (iv) replacing one or more naturally occurring nucleotides with synthetic nucleotides (e.g, 5-methylcytidine for cytidine and / or pseudouridine (QJ) or N(l)-methylpseudouridine (mlQJ) or 5-methyluridine (m5U) for uridine), and (v) codon optimization.
[0233] 5'-Cao
[0234] In some embodiments, the RNA (in particular, mRNA) described herein comprises a 5'-cap structure. In some embodiments, the RNA does not have uncapped 5'-triphosphates. In some embodiments, the RNA (in particular, mRNA) may comprise a conventional 5'-cap and / or a 5'-cap analog. The term "conventional 5'-cap" refers to a cap structure found on the 5'-end of an RNA molecule and generally comprises a guanosine 5'-triphosphate (Gppp) which is connected via its triphosphate moiety to the 5'-end of the next nucleotide of the RNA (Ze., the guanosine is connected via a 5' to 5' triphosphate linkage to the rest of the RNA). The guanosine may be methylated at position N7(resulting in the cap structure m7Gppp). The term "5'-cap analog" includes a 5'-cap which is based on a conventional 5'-cap but which has been modified at either the 2'- or 3'-position of the m7guanosine structure in order to avoid an integration of the 5'-cap analog in the reverse orientation (such 5'-cap analogs are also called anti-reverse cap analogs (ARCAs)). Particularly preferred 5'-cap analogs are those having one or more substitutions at the bridging and non-bridging oxygen in the phosphate bridge, such as phosphorothioate modified 5'-cap analogs at the p-phosphate (such as m27'2 OG(5')ppSp(5')G (referred to as beta-S-ARCA or -S-ARCA)), as described in PCT / EP2019 / 056502. Providing an RNA (in particular, mRNA) with a 5'-cap structure as described herein may be achieved by in vitro transcription of a DNA template in presence of a corresponding 5'-cap compound, wherein said 5'-cap structure is co-transcriptionally incorporated into the generated RNA (in particular, mRNA) strand, or the RNA (in particular, mRNA) may be generated, for example, by in vitro transcription, and the 5'-cap structure may be attached to the RNA post-transcriptionally using capping enzymes, for example, capping enzymes of vaccinia virus.
[0235] In some embodiments, the RNA (in particular, mRNA) comprises a 5'-cap structure selected from the group consisting of m27'2 OG(5')ppSp(5')G (in particular its DI diastereomer), m27'3 OG(5')ppp(5')G, and m27'3 OGppp(mi2' °)ApG. In some embodiments, RNA comprises m27'2 OG(5')ppSp(5')G (in particular its DI diastereomer) as 5'-cap structure. In some embodiments, RNA comprises m27'3'0Gppp(mi2'0)ApG as 5'-cap structure.
[0236] In some embodiments, the RNA (in particular, mRNA) comprises a capO, capl, or cap2, preferably capl or cap2. According to the present disclosure, the term "capO" means the structure "m7GpppN", wherein N is any nucleoside bearing an OH moiety at position 2'. According to the present disclosure, the term "capl" means the structure "m7GpppNm", wherein Nm is any nucleoside bearing an OCH3 moiety at position 2'. According to the present disclosure, the term "cap2" means the structure "m7GpppNmNm", wherein each Nm is independently any nucleoside bearing an OCH3 moiety at position 2'.
[0237] The 5'-cap analog beta-S-ARCA (P-S-ARCA) has the following structure:
[0238] The "DI diastereomer of beta-S-ARCA" or "beta-S-ARCA(Dl)" is the diastereomer of beta-S-ARCA which elutes first on an HPLC column compared to the D2 diastereomer of beta-S-ARCA (beta-S-ARCA(D2)) and thus exhibits a shorter retention time. The HPLC preferably is an analytical HPLC. In some embodiments, a Supelcosil LC-18-T RP column, preferably of the format: 5 pm, 4.6 x 250 mm is used for separation, whereby a flow rate of 1.3 ml / min can be applied. In some embodiments, a gradient of methanol in ammonium acetate, for example, a 0-25% linear gradient of methanol in 0.05 M ammonium acetate, pH = 5.9, within 15 min is used. UV-detection (VWD) can be performed at 260 nm and fluorescence detection (FLD) can be performed with excitation at 280 nm and detection at 337 nm. The 5'-cap analog m27'3'0Gppp(mi2, 0)ApG (also referred to as m27'30G(5')ppp(5')m2'0ApG) which is a building block of a capl has the following structure:
[0239] An exemplary capO mRNA comprising p-S-ARCA and mRNA has the following structure:
[0240] An exemplary capO mRNA comprising m27'3 OG(5')ppp(5')G and mRNA has the following structure: An exemplary capl mRNA comprising m27'3 OGppp(mi2' °)ApG and mRNA has the following structure:
[0241] In some embodiments, the RNA is unmodified RNA and comprises a 5'-cap, which is capl. In further embodiments, the RNA comprises modified nucleotides, nucleosides or nucleobases and a 5'-cap, which is capl. For instance, such modified RNA comprises one or more modified uridines in place of unmodified uridines. In certain embodiments, the modified RNA comprising the capl structure comprises modified uridines in place of all uridines, the modified uridines preferably being Nl-methyl-pseudouridine. Polv-A tail
[0242] As used herein, the term "poly-A tail" or "poly-A sequence" refers to an uninterrupted or interrupted sequence of adenylate residues which is typically located at the 3'-end of an RNA (in particular, mRNA) molecule. Poly-A tails or poly-A sequences are known to those of skill in the art and may follow the 3'-UTR in the RNAs (in particular, mRNAs) described herein. An uninterrupted poly-A tail is characterized by consecutive adenylate residues. In nature, an uninterrupted poly-A tail is typical. RNAs (in particular, mRNAs) disclosed herein can have a poly-A tail attached to the free 3'-end of the RNA by a template-independent RNA polymerase after transcription or a poly-A tail encoded by DNA and transcribed by a template-dependent RNA polymerase.
[0243] It has been demonstrated that a poly-A tail of about 120 A nucleotides has a beneficial influence on the levels of RNA in transfected eukaryotic cells, as well as on the levels of protein that is translated from an open reading frame that is present upstream (S' of the poly-A tail (Holtkamp eta / ., 2006, Blood, vol. 108, pp. 4009-4017).
[0244] The poly-A tail may be of any length. In some embodiments, a poly-A tail comprises, essentially consists of, or consists of at least 20, at least 30, at least 40, at least 80, or at least 100 and up to 500, up to 400, up to 300, up to 200, or up to 150 A nucleotides, and, in particular, about 120 A nucleotides. In this context, "essentially consists of" means that most nucleotides in the poly-A tail, typically at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% by number of nucleotides in the poly-A tail are A nucleotides, but permits that remaining nucleotides are nucleotides other than A nucleotides, such as U nucleotides (uridylate), G nucleotides (guanylate), or C nucleotides (cytidylate). In this context, "consists of" means that all nucleotides in the poly-A tail, i.e., 100% by number of nucleotides in the poly-A tail, are A nucleotides. The term "A nucleotide" or "A" refers to adenylate.
[0245] In some embodiments, a poly-A tail is attached during RNA transcription, e.g, during preparation of in vitro transcribed RNA, based on a DNA template comprising repeated dT nucleotides (deoxythymidylate) in the strand complementary to the coding strand. The DNA sequence encoding a poly-A tail (coding strand) is referred to as poly(A) cassette.
[0246] In some embodiments, the poly(A) cassette present in the coding strand of DNA essentially consists of dA nucleotides, but is interrupted by a random sequence of the four nucleotides (dA, dC, dG, and dT). Such random sequence may be 5 to 50, 10 to 30, or 10 to 20 nucleotides in length. Such a cassette is disclosed in WO 2016 / 005324 Al, hereby incorporated by reference. Any poly(A) cassette disclosed in WO 2016 / 005324 Al may be used in the present disclosure. A poly(A) cassette that essentially consists of dA nucleotides, but is interrupted by a random sequence having an equal distribution of the four nucleotides (dA, dC, dG, dT) and having a length of e.g., 5 to 50 nucleotides shows, on DNA level, constant propagation of plasmid DNA in £ co / / and is still associated, on RNA level, with the beneficial properties with respect to supporting RNA stability and translational efficiency is encompassed. Consequently, in some embodiments, the poly-A tail contained in an RNA (in particular, mRNA) molecule described herein essentially consists of A nucleotides, but is interrupted by a random sequence of the four nucleotides (A, C, G, U). Such random sequence may be 5 to 50, 10 to 30, or 10 to 20 nucleotides in length.
[0247] In some embodiments, the poly(A) tail comprises 30 adenine nucleotides followed by 70 adenine nucleotides, wherein the 30 adenine nucleotides and 70 adenine nucleotides are separated by a linker sequence of 10 nucleotides.
[0248] In some embodiments, no nucleotides other than A nucleotides flank a poly-A tail at its 3'-end, i.e., the poly-A tail is not masked or followed at its 3'-end by a nucleotide other than A.
[0249] In some embodiments, a poly-A tail may comprise at least 20, at least 30, at least 40, at least 80, or at least 100 and up to 500, up to 400, up to 300, up to 200, or up to 150 nucleotides. In some embodiments, the poly-A tail may essentially consist of at least 20, at least 30, at least 40, at least 80, or at least 100 and up to 500, up to 400, up to 300, up to 200, or up to 150 nucleotides. In some embodiments, the poly-A tail may consist of at least 20, at least 30, at least 40, at least 80, or at least 100 and up to 500, up to 400, up to 300, up to 200, or up to 150 nucleotides. In some embodiments, the poly-A tail comprises the poly-A tail shown in SEQ ID NO: 55. In some embodiments, the poly- A tail comprises at least 100 nucleotides. In some embodiments, the poly-A tail comprises about 150 nucleotides. In some embodiments, the poly-A tail comprises about 120 nucleotides.
[0250] Untranslated regions (UTR)
[0251] In some embodiments, RNA (in particular, mRNA) described in present disclosure comprises a 5'-UTR and / or a 3'-UTR. The term "untranslated region" or "UTR" relates to a region in a DNA molecule which is transcribed but is not translated into an amino acid sequence, or to the corresponding region in an RNA molecule, such as an mRNA molecule. An untranslated region (UTR) can be present 5' (upstream) of an open reading frame (5'-UTR) and / or 3' (downstream) of an open reading frame (3'-UTR). A 5'-UTR, if present, is located at the 5'-end, upstream of the start codon of a proteinencoding region. A 5'-UTR is downstream of the 5'-cap (if present), e.g., directly adjacent to the 5'-cap. A 3'-UTR, if present, is located at the 3'-end, downstream of the termination codon of a protein-encoding region, but the term "3'- UTR" does generally not include the poly-A sequence. Thus, the 3'-UTR is upstream of the poly-A sequence (if present), e.g, directly adjacent to the poly-A sequence. Incorporation of a 3'-UTR into the 3'-non translated region of an RNA (preferably mRNA) molecule can result in an enhancement in translation efficiency. A synergistic effect may be achieved by incorporating two or more of such 3'-UTRs (which are preferably arranged in a head-to-tail orientation; cf., e.g, Holtkamp et al., Blood 108, 4009-4017 (2006)). The 3'-UTRs may be autologous or heterologous to the RNA (e.g., mRNA) into which they are introduced. In certain embodiments, the 3'-UTR is derived from a globin gene or mRNA, such as a gene or mRNA of alpha2-globin, alphal-globin, or beta-globin, e.g., beta-globin, e.g., human beta-globin. For example, the RNA (e.g., mRNA) may be modified by the replacement of the existing 3'-UTR with or the insertion of one or more, e.g., two copies of a 3'-UTR derived from a globin gene, such as alpha2-globin, alphal-globin, betaglobin, e.g., beta-globin, e.g., human beta-globin.
[0252] In some embodiments, a 5'-UTR is or comprises a modified human alpha-globin 5'-UTR. A particularly preferred 5'-UTR comprises the nucleotide sequence of SEQ ID NO: 53. In some embodiments, a 3'-UTR comprises a first sequence from the amino terminal enhancer of split (AES) messenger RNA and a second sequence from the mitochondrial encoded 12S ribosomal RNA. A particularly preferred 3'-UTR comprises the nucleotide sequence of SEQ ID NO: 54.
[0253] In some embodiments, RNA comprises a 5'-UTR comprising the nucleotide sequence of SEQ ID NO: 53, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 53.
[0254] In some embodiments, RNA comprises a 3'-UTR comprising the nucleotide sequence of SEQ ID NO: 54, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 54.
[0255] Table 1: Exemplary untranslated RNA sequences
[0256] Chemical modification
[0257] The RNA (in particular, mRNA) described herein may have modified ribonucleotides in order to increase its stability and / or decrease immunogenicity and / or decrease cytotoxicity. For example, in some embodiments, uridine in the RNA (in particular, mRNA) described herein is replaced (partially or completely, preferably completely) by a modified nucleoside. In some embodiments, the modified nucleoside is a modified uridine.
[0258] In some embodiments, the modified uridine replacing uridine is selected from the group consisting of pseudouridine (ip), Nl-methyl-pseudouridine (mlip), 5-methyl-uridine (m5U), and combinations thereof.
[0259] In some embodiments, the modified nucleoside replacing (partially or completely, preferably completely) uridine in the RNA may be any one or more of 3-methyl-uridine (m3U), 5-methoxy-uridine (mo5U), 5-aza-uridine, 6-aza-uridine, 2- thio-5-aza-uridine, 2-thio-uridine (s2U), 4-thio-uridine (s4U), 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxy- uridine (ho5U), 5-aminoallyl-uridine, 5-halo-uridine (e.g, 5-iodo-uridineor 5-bromo-uridine), uridine 5-oxyacetic acid (cmo5U), uridine 5-oxyacetic acid methyl ester (mcmo5U), 5-carboxymethyl-uridine (cm5U), 1-carboxymethyl- pseudouridine, 5-carboxyhydroxymethyl-uridine (chm5U), 5-carboxyhydroxymethyl-uridine methyl ester (mchm5U), 5- methoxycarbonylmethyl-uridine (mcm5U), 5-methoxycarbonylmethyl-2-thio-uridine (mcm5s2U), 5-aminomethyl-2- thio-uridine (nm5s2U), 5-methylaminomethyl-uridine (mnm5U), 1-ethyl-pseudouridine, 5-methylaminomethyl-2-thio- uridine (mnm5s2U), 5-methylaminomethyl-2-seleno-uridine (mnm5se2U), 5-carbamoylmethyl-uridine (ncm5U), 5- carboxymethylaminomethyl-uridine (cmnm5U), 5-carboxymethylaminomethyl-2-thio-uridine (cmnm5s2U), 5-propynyl- uridine, 1-propynyl-pseudouridine, 5-taurinomethyl-uridine (Tm5U), 1-taurinomethyl-pseudouridine, 5-taurinomethyl- 2-thio-uridine(Tm5s2U), l-taurinomethyl-4-thio-pseudouridine), 5-methyl-2-thio-uridine (m5s2U), l-methyl-4-thio- pseudouridine (mls4ip), 4-thio-l-methyl-pseudouridine, 3-methyl-pseudouridine (m3ip), 2-thio-l-methyl- pseudouridine, 1-methyl-l-deaza-pseudouridine, 2-thio-l-methyl-l-deaza-pseudouridine, dihydrouridine (D), dihydropseudouridine, 5,6-dihydrouridine, 5-methyl-dihydrouridine (m5D), 2-thio-dihydrouridine, 2-thio- dihydropseudouridine, 2-methoxy-uridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, 4-methoxy-2-thio- pseudouridine, Nl-methyl-pseudouridine, 3-(3-amino-3-carboxypropyl)uridine (acp3U), l-methyl-3-(3-amino-3- carboxypropyl)pseudouridine (acp3 ip), 5-(isopentenylaminomethyl)uridine (inm5U), 5-(isopentenylaminomethyl)-2- thio-uridine (inm5s2U), a-thio-uridine, 2'-O-methyl-uridine (Um), 5,2'-O-dimethyl-uridine (m5Um), 2'-O-methyl- pseudouridine (ipm), 2-thio-2'-O-methyl-uridine (s2Um), 5-methoxycarbonylmethyl-2'-O-methyl-uridine (mcm5Um), 5- carbamoylmethyl-2'-O-methyl-uridine (ncm5Um), 5-carboxymethylaminomethyl-2'-O-methyl-uridine (cmnm5Um), 3,2'-O-dimethyl-uridine (m3Um), 5-(isopentenylaminomethyl)-2'-O-methyl-uridine (inm5Um), 1-thio-uridine, deoxythymidine, 2'-F-ara-uridine, 2'-F-uridine, 2'-OH-ara-uridine, 5-(2-carbomethoxyvinyl) uridine, 5-[3-(l-E- propenylamino)uridine, or any other modified uridine known in the art.
[0260] An RNA (preferably mRNA) which is modified by pseudouridine (replacing partially or completely, preferably completely, uridine) is referred to herein as "QJ-modified", whereas the term "m li -modified" means that the RNA (preferably mRNA) contains N(l)-methylpseudouridine (replacing partially or completely, preferably completely, uridine). Furthermore, the term "m5U-modified" means that the RNA (preferably mRNA) contains 5-methyluridine (replacing partially or completely, preferably completely, uridine). Such i - or mlQJ- or m5U-modified RNAs usually exhibit decreased immunogenicity compared to their unmodified forms and, thus, are preferred in applications where the induction of an immune response is to be avoided or minimized. In some embodiments, the RNA (preferably mRNA) contains N(l)-methylpseudouridine replacing completely uridine.
[0261] Codon optimization and GC enrichment
[0262] The codons of the RNA (in particular, mRNA) described in the present disclosure may further be optimized, e.g., to increase the GC content of the RNA and / or to replace codons which are rare in the cell (or subject) in which the peptide or polypeptide of interest is to be expressed by codons which are synonymous frequent codons in said cell (or subject). In some embodiments, the amino acid sequence encoded by the RNA (in particular, mRNA) described in the present disclosure is encoded by a coding sequence which is codon-optimized and / or the G / C content of which is increased compared to wild type coding sequence. This also includes embodiments, wherein one or more sequence regions of the coding sequence are codon-optimized and / or increased in the G / C content compared to the corresponding sequence regions of the wild type coding sequence. In some embodiments, the codon-optimization and / or the increase in the G / C content preferably does not change the sequence of the encoded amino acid sequence.
[0263] The term "codon-optimized" refers to the alteration of codons in the coding region of a nucleic acid molecule to reflect the typical codon usage of a host organism without preferably altering the amino acid sequence encoded by the nucleic acid molecule. Within the context of the present disclosure, coding regions may be codon-optimized for optimal expression in a subject to be treated using the RNA (in particular, mRNA) described herein. Codon-optimization is based on the finding that the translation efficiency is also determined by a different frequency in the occurrence of tRNAs in cells. Thus, the sequence of RNA (in particular, mRNA) may be modified such that codons for which frequently occurring tRNAs are available are inserted in place of "rare codons".
[0264] In some embodiments, the guanosine / cytosine (G / C) content of the coding region of the RNA (in particular, mRNA) described herein is increased compared to the G / C content of the corresponding coding sequence of the wild type RNA, wherein the amino acid sequence encoded by the RNA is preferably not modified compared to the amino acid sequence encoded by the wild type RNA. This modification of the RNA sequence is based on the fact that the sequence of any RNA region to be translated is important for efficient translation of that RNA. Sequences having an increased G (guanosine) / C (cytosine) content are more stable than sequences having an increased A (adenosine) / U (uracil) content. In respect to the fact that several codons code for one and the same amino acid (so-called degeneration of the genetic code), the most favorable codons for the stability can be determined (so-called alternative codon usage). Depending on the amino acid to be encoded by the RNA, there are various possibilities for modification of the RNA sequence, compared to its wild type sequence. In particular, codons which contain A and / or U nucleotides can be modified by substituting these codons by other codons, which code for the same amino acids but contain no A and / or U or contain a lower content of A and / or U nucleotides.
[0265] In various embodiments, the G / C content of the coding region of the RNA (in particular, mRNA) described herein is increased by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, or even more compared to the G / C content of the coding region of the wild type RNA. Non-immunoaenic RNA
[0266] The term "non-immunogenic RNA" (such as "non-immunogenic mRNA") as used herein refers to RNA that does not induce a response by the immune system upon administration, e.g, to a mammal, or induces a weaker response than would have been induced by the same RNA that differs only in that it has not been subjected to the modifications and treatments that render the non-immunogenic RNA non-immunogenic, i.e., than would have been induced by standard RNA (stdRNA). In certain embodiments, non-immunogenic RNA is rendered non-immunogenic by incorporating modified nucleosides suppressing RNA-mediated activation of innate immune receptors into the RNA and / or limiting the amount of double-stranded RNA (dsRNA), e.g., by limiting the formation of double-stranded RNA (dsRNA), e.g., during in vitro transcription, and / or by removing double-stranded RNA (dsRNA), e.g., following in vitro transcription. In certain embodiments, non-immunogenic RNA is rendered non-immunogenic by incorporating modified nucleosides suppressing RNA-mediated activation of innate immune receptors into the RNA and / or by removing double-stranded RNA (dsRNA), e.g., following in vitro transcription.
[0267] For rendering the non-immunogenic RNA (especially mRNA) non-immunogenic by the incorporation of modified nucleosides, any modified nucleoside may be used as long as it lowers or suppresses immunogenicity of the RNA. Particularly preferred are modified nucleosides that suppress RNA-mediated activation of innate immune receptors. In some embodiments, the modified nucleosides comprise a replacement of one or more uridines with a nucleoside comprising a modified nucleobase. In some embodiments, the modified nucleobase is a modified uracil. In some embodiments, the nucleoside comprising a modified nucleobase is selected from the group consisting of 3-methyl- uridine (m3U), 5-methoxy-uridine (mo5U), 5-aza-uridine, 6-aza-uridine, 2-thio-5-aza-uridine, 2-thio-uridine (s2U), 4- thio-uridine (s4U), 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxy-uridine (ho5U), 5-aminoallyl-uridine, 5-halo- uridine e.g, 5-iodo-uridine or 5-bromo-uridine), uridine 5-oxyacetic acid (cmo5U), uridine 5-oxyacetic acid methyl ester (mcmo5U), 5-carboxymethyl-uridine (cm5U), 1-carboxymethyl-pseudouridine, 5-carboxyhydroxymethyl-uridine (chm5U), 5-carboxyhydroxymethyl-uridine methyl ester (mchm5U), 5-methoxycarbonylmethyl-uridine (mcm5U), 5- methoxycarbonylmethyl-2-thio-uridine (mcm5s2U), 5-aminomethyl-2-thio-uridine (nm5s2U), 5-methylaminomethyl- uridine (mnm5U), 1-ethyl-pseudouridine, 5-methylaminomethyl-2-thio-uridine (mnm5s2U), 5-methylaminomethyl-2- seleno-uridine (mnm5se2U), 5-carbamoylmethyl-uridine (ncm5U), 5-carboxymethylaminomethyl-uridine (cmnm5U), 5- carboxymethylaminomethyl-2-thio-uridine (cmnm5s2U), 5-propynyl-uridine, 1-propynyl-pseudouridine, 5- taurinomethyl-uridine (Tm5U), 1-taurinomethyl-pseudouridine, 5-taurinomethyl-2-thio-uridine(Tm5s2U), 1- taurinomethyl-4-thio-pseudouridine), 5-methyl-2-thio-uridine (m5s2U), l-methyl-4-thio-pseudouridine (m^ip), 4-thio- 1-methyl-pseudouridine, 3-methyl-pseudouridine (m3ip), 2-thio-l-methyl-pseudouridine, 1-methyl-l-deaza- pseudouridine, 2-thio-l-methyl-l-deaza-pseudouridine, dihydrouridine (D), dihydropseudouridine, 5,6-dihydrouridine, 5-methyl-dihydrouridine (m5D), 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxy-uridine, 2-methoxy-4- thio-uridine, 4-methoxy-pseudouridine, 4-methoxy-2-thio-pseudouridine, Nl-methyl-pseudouridine, 3-(3-amino-3- carboxypropyl)uridine (acp3U), l-methyl-3-(3-amino-3-carboxypropyl)pseudouridine (acp3ip), 5- (isopentenylaminomethyl)uridine (inm5U), 5-(isopentenylaminomethyl)-2-thio-uridine (inm5s2U), a-thio-uridine, 2'-O- methyl-uridine (Um), 5,2'-O-dimethyl-uridine (m5Um), 2'-O-methyl-pseudouridine (ipm), 2-thio-2'-O-methyl-uridine (s2Um), 5-methoxycarbonylmethyl-2'-O-methyl-uridine (mcm5Um), 5-carbamoylmethyl-2'-O-methyl-uridine (ncm5Um), 5-carboxymethylaminomethyl-2'-O-methyl-uridine (cmnm5Um), 3,2'-O-dimethyl-uridine (m3Um), 5- (isopentenylaminomethyl)-2'-O-methyl-uridine (inm5Um), 1-thio-uridine, deoxythymidine, 2'-F-ara-uridine, 2'-F- uridine, 2'-OH-ara-uridine, 5-(2-carbomethoxyvinyl) uridine, and 5-[3-(l-E-propenylamino)uridine. In certain embodiments, the nucleoside comprising a modified nucleobase is pseudouridine (ip), Nl-methyl-pseudouridine (mlip) or 5-methyl-uridine (m5U), in particular Nl-methyl-pseudouridine. In some embodiments, the replacement of one or more uridines with a nucleoside comprising a modified nucleobase comprises a replacement of at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 25%, at least 50%, at least 75%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% of the uridines.
[0268] During synthesis of mRNA by in vitro transcription (IVT) using T7 RNA polymerase significant amounts of aberrant products, including double-stranded RNA (dsRNA) are produced due to unconventional activity of the enzyme. dsRNA induces inflammatory cytokines and activates effector enzymes leading to protein synthesis inhibition. Formation of dsRNA can be limited during synthesis of mRNA by in vitro transcription (IVT), for example, by limiting the amount of uridine triphosphate (UTP) during synthesis. Optionally, UTP may be added once or several times during synthesis of mRNA. Also, dsRNA can be removed from RNA such as IVT RNA, for example, by ion-pair reversed phase HPLC using a non-porous or porous C-18 polystyrene-divinylbenzene (PS-DVB) matrix. Alternatively, an enzymatic based method using E. co / / RNaseIII that specifically hydrolyzes dsRNA but not ssRNA, thereby eliminating dsRNA contaminants from IVT RNA preparations can be used. Furthermore, dsRNA can be separated from ssRNA by using a cellulose material. In some embodiments, an RNA preparation is contacted with a cellulose material and the ssRNA is separated from the cellulose material under conditions which allow binding of dsRNA to the cellulose material and do not allow binding of ssRNA to the cellulose material. Suitable methods for providing ssRNA are disclosed, for example, in WO 2017 / 182524.
[0269] As the term is used herein, "remove" or "removal" refers to the characteristic of a population of first substances, such as non-immunogenic RNA, being separated from the proximity of a population of second substances, such as dsRNA, wherein the population of first substances is not necessarily devoid of the second substance, and the population of second substances is not necessarily devoid of the first substance. However, a population of first substances characterized by the removal of a population of second substances has a measurably lower content of second substances as compared to the non-separated mixture of first and second substances.
[0270] In some embodiments, the amount of double-stranded RNA (dsRNA) is limited, e.g., dsRNA (especially dsmRNA) is removed from non-immunogenic RNA , such that less than 10%, less than 5%, less than 4%, less than 3%, less than 2%, less than 1%, less than 0.5%, less than 0.3%, less than 0.1%, less than 0.05%, less than 0.03%, less than 0.01%, less than 0.005%, less than 0.004%, less than 0.003%, less than 0.002%, less than 0.001%, or less than 0.0005% of the RNA in the non-immunogenic RNA composition is dsRNA. In some embodiments, the non-immunogenic RNA (especially mRNA) is free or essentially free of dsRNA. In some embodiments, the non-immunogenic RNA (especially mRNA) composition comprises a purified preparation of single-stranded nucleoside modified RNA. In some embodiments, the non-immunogenic RNA (especially mRNA) composition comprises single-stranded nucleoside modified RNA (especially mRNA) and is substantially free of double stranded RNA (dsRNA). In some embodiments, the non-immunogenic RNA (especially mRNA) composition comprises at least 90%, at least 91%, at least 92%, at least 93 %, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, at least 99.99%, at least 99.991%, at least 99.992%, , at least 99.993%,, at least 99.994%, , at least 99.995%, at least 99.996%, at least 99.997%, or at least 99.998% single stranded nucleoside modified RNA, relative to all other nucleic acid molecules (DNA, dsRNA, etc.).
[0271] Various methods can be used to determine the amount of dsRNA. For example, a sample may be contacted with dsRNA-specific antibody and the amount of antibody binding to RNA may be taken as a measure for the amount of dsRNA in the sample. A sample containing a known amount of dsRNA may be used as a reference.
[0272] For example, RNA may be spotted onto a membrane, e.g., nylon blotting membrane. The membrane may be blocked, e.g., in TBS-T buffer (20 mM TRIS pH 7.4, 137 mM NaCI, 0.1% (v / v) TWEEN-20) containing 5% (w / v) skim milk powder. For detection of dsRNA, the membrane may be incubated with dsRNA-specific antibody, e.g., dsRNA-specific mouse mAb (English & Scientific Consulting, Szirak, Hungary). After washing, e.g., with TBS-T, the membrane may be incubated with a secondary antibody, e.g., HRP-conjugated donkey anti-mouse IgG (Jackson ImmunoResearch, Cat #715-035-150), and the signal provided by the secondary antibody may be detected.
[0273] In some embodiments, the non-immunogenic RNA (especially mRNA) is translated in a cell more efficiently than standard RNA with the same sequence. In some embodiments, translation is enhanced by a factor of 2-fold relative to its unmodified counterpart. In some embodiments, translation is enhanced by a 3-fold factor. In some embodiments, translation is enhanced by a 4-fold factor. In some embodiments, translation is enhanced by a 5-fold factor. In some embodiments, translation is enhanced by a 6-fold factor. In some embodiments, translation is enhanced by a 7-fold factor. In some embodiments, translation is enhanced by an 8-fold factor. In some embodiments, translation is enhanced by a 9-fold factor. In some embodiments, translation is enhanced by a 10-fold factor. In some embodiments, translation is enhanced by a 15-fold factor. In some embodiments, translation is enhanced by a 20-fold factor. In some embodiments, translation is enhanced by a 50-fold factor. In some embodiments, translation is enhanced by a 100- fold factor. In some embodiments, translation is enhanced by a 200-fold factor. In some embodiments, translation is enhanced by a 500-fold factor. In some embodiments, translation is enhanced by a 1000-fold factor. In some embodiments, translation is enhanced by a 2000-fold factor. In some embodiments, the factor is 10-1000-fold. In some embodiments, the factor is 10-100-fold. In some embodiments, the factor is 10-200-fold. In some embodiments, the factor is 10-300-fold. In some embodiments, the factor is 10-500-fold. In some embodiments, the factor is 20-1000- fold. In some embodiments, the factor is 30-1000-fold. In some embodiments, the factor is 50-1000-fold. In some embodiments, the factor is 100-1000-fold. In some embodiments, the factor is 200-1000-fold. In some embodiments, translation is enhanced by any other significant amount or range of amounts.
[0274] In some embodiments, the non-immunogenic RNA (especially mRNA) exhibits significantly less innate immunogenicity than standard RNA with the same sequence. In some embodiments, the non-immunogenic RNA (especially mRNA) exhibits an innate immune response that is 2-fold less than its unmodified counterpart. In some embodiments, innate immunogenicity is reduced by a 3-fold factor. In some embodiments, innate immunogenicity is reduced by a 4-fold factor. In some embodiments, innate immunogenicity is reduced by a 5-fold factor. In some embodiments, innate immunogenicity is reduced by a 6-fold factor. In some embodiments, innate immunogenicity is reduced by a 7-fold factor. In some embodiments, innate immunogenicity is reduced by an 8-fold factor. In some embodiments, innate immunogenicity is reduced by a 9-fold factor. In some embodiments, innate immunogenicity is reduced by a 10-fold factor. In some embodiments, innate immunogenicity is reduced by a 15-fold factor. In some embodiments, innate immunogenicity is reduced by a 20-fold factor. In some embodiments, innate immunogenicity is reduced by a 50-fold factor. In some embodiments, innate immunogenicity is reduced by a 100-fold factor. In some embodiments, innate immunogenicity is reduced by a 200-fold factor. In some embodiments, innate immunogenicity is reduced by a 500- fold factor. In some embodiments, innate immunogenicity is reduced by a 1000-fold factor. In some embodiments, innate immunogenicity is reduced by a 2000-fold factor.
[0275] The term "exhibits significantly less innate immunogenicity" refers to a detectable decrease in innate immunogenicity. In some embodiments, the term refers to a decrease such that an effective amount of the non-immunogenic RNA (especially mRNA) can be administered without triggering a detectable innate immune response. In some embodiments, the term refers to a decrease such that the non-immunogenic RNA (especially mRNA) can be repeatedly administered without eliciting an innate immune response sufficient to detectably reduce production of the protein encoded by the non-immunogenic RNA. In some embodiments, the decrease is such that the non-immunogenic RNA (especially mRNA) can be repeatedly administered without eliciting an innate immune response sufficient to eliminate detectable production of the protein encoded by the non-immunogenic RNA.
[0276] "Immunogenicity" is the ability of a foreign substance, such as RNA, to provoke an immune response in the body of a human or other animal. The innate immune system is the component of the immune system that is relatively unspecific and immediate. It is one of two main components of the vertebrate immune system, along with the adaptive immune system.
[0277] Antigen-coding RNA and use thereof for inducing an immune response
[0278] Generally, RNA (in particular, mRNA) described in the present disclosure comprises a nucleic acid sequence encoding a peptide or polypeptide comprising one or more Mycobacterium tuberculosis antigens, immunogenic variants or fragments thereof, for inducing an immune response against Mycobacterium tuberculosis in a subject. The peptide or polypeptide for inducing an immune response is also designated herein as "vaccine antigen" or simply "antigen".
[0279] In some embodiments, the RNA (in particular, mRNA) is translated into the respective protein upon entering cells of a subject being administered the RNA, e.g., muscle cells or antigen-presenting cells (APCs).
[0280] In some embodiments, the RNA encoding the vaccine antigen is expressed in cells of the subject to provide the vaccine antigen. In some embodiments, the RNA encoding the vaccine antigen is transiently expressed in cells of the subject. In some embodiments, the vaccine antigen is presented in the context of MHC. In some embodiments, the vaccine antigen is secreted by cells of the subject.
[0281] In some embodiments, the RNA encoding the vaccine antigen is administered intramuscularly.
[0282] In some embodiments, the RNA encoding the vaccine antigen is administered systemically, e.g., intravenously. In some embodiments, after systemic administration of the RNA encoding the vaccine antigen, expression of the RNA encoding the vaccine antigen in spleen occurs. In some embodiments, after systemic administration of the RNA encoding the vaccine antigen, expression of the RNA encoding the vaccine antigen in antigen presenting cells, preferably professional antigen presenting cells occurs. In some embodiments, the antigen presenting cells are selected from the group consisting of dendritic cells, macrophages and B cells. In some embodiments, after systemic administration of the RNA encoding the vaccine antigen, no or essentially no expression of the RNA encoding the vaccine antigen in lung and / or liver occurs. In some embodiments, after systemic administration of the RNA encoding the vaccine antigen, expression of the RNA encoding the vaccine antigen in spleen is at least 5-fold the amount of expression in lung.
[0283] A vaccine antigen comprises an epitope for inducing an immune response against a disease-associated antigen, e.g., a protein of an infectious agent (e.g., Mtb antigen), in a subject. Accordingly, the vaccine antigen comprises an antigenic sequence for inducing an immune response against a disease-associated antigen in a subject. Such antigenic sequence may correspond to a target antigen or disease-associated antigen, an immunogenic variant thereof, or an immunogenic fragment of the target antigen or disease-associated antigen or the immunogenic variant thereof. Thus, the antigenic sequence may comprise at least an epitope of a target antigen or disease-associated antigen or an immunogenic variant thereof.
[0284] The antigenic sequences, e.g., epitopes, suitable for use according to the disclosure typically may be derived from a target antigen, i.e. the antigen against which an immune response is to be elicited. For example, the antigenic sequences contained within the vaccine antigen may be a target antigen or a fragment or variant of a target antigen. The antigenic sequence or a procession product thereof, e.g., a fragment thereof, may bind to an antigen receptor such as TCR carried by immune effector cells. In some embodiments, the antigenic sequence is selected from the group consisting of the antigen expressed by a target cell to which the immune effector cells are targeted or a fragment thereof, or a variant of the antigenic sequence or the fragment.
[0285] A vaccine antigen which may be provided to a subject according to the present disclosure by administering RNA encoding the vaccine antigen, preferably results in the induction of an immune response, e.g., in the stimulation, priming and / or expansion of immune effector cells, in the subject being provided the vaccine antigen. Said immune response, e.g., stimulated, primed and / or expanded immune effector cells, is preferably directed against a target antigen, in particular a target antigen expressed in diseased cells, tissues and / or organs, i.e., a disease-associated antigen. Thus, a vaccine antigen may comprise the disease-associated antigen, or a fragment or variant thereof. In some embodiments, such fragment or variant is immunologically equivalent to the disease-associated antigen.
[0286] The term "immunologically equivalent" means that the immunologically equivalent molecule such as the immunologically equivalent amino acid sequence exhibits the same or essentially the same immunological properties and / or exerts the same or essentially the same immunological effects, e.g., with respect to the type of the immunological effect. In the context of the present disclosure, the term "immunologically equivalent" is preferably used with respect to the immunological effects or properties of antigens or antigen variants used for immunization. For example, an amino acid sequence is immunologically equivalent to a reference amino acid sequence if said amino acid sequence when exposed to the immune system of a subject induces an immune reaction having a specificity of reacting with the reference amino acid sequence. Thus, in some embodiments, a molecule which is immunologically equivalent to an antigen exhibits the same or essentially the same properties and / or exerts the same or essentially the same effects regarding the stimulation, priming and / or expansion of T cells as the antigen to which the T cells are targeted.
[0287] In the context of the present disclosure, the term "fragment of an antigen" or "variant of an antigen" means an agent which results in the induction of an immune response, e.g., in the stimulation, priming and / or expansion of immune effector cells, which immune response, e.g., stimulated, primed and / or expanded immune effector cells, targets the antigen, i.e. a disease-associated antigen, in particular when presented by diseased cells, tissues and / or organs. Thus, the vaccine antigen may correspond to or may comprise the disease-associated antigen, may correspond to or may comprise a fragment of the disease-associated antigen or may correspond to or may comprise an antigen which is homologous to the disease-associated antigen or a fragment thereof. If the vaccine antigen comprises a fragment of the disease-associated antigen or an amino acid sequence which is homologous to a fragment of the disease-associated antigen said fragment or amino acid sequence may comprise an epitope of the disease-associated antigen to which the antigen receptor of the immune effector cells is targeted or a sequence which is homologous to an epitope of the disease-associated antigen. Thus, according to the disclosure, a vaccine antigen may comprise an immunogenic fragment of a disease-associated antigen or an amino acid sequence being homologous to an immunogenic fragment of a disease-associated antigen. An "immunogenic fragment of an antigen" according to the disclosure preferably relates to a fragment of an antigen which is capable of inducing an immune response against, e.g., stimulating, priming and / or expanding immune effector cells carrying an antigen receptor binding to, the antigen or cells expressing the antigen. It is preferred that the vaccine antigen (similar to the disease-associated antigen) provides the relevant epitope for binding by the antigen receptor present on the immune effector cells. In some embodiments, the vaccine antigen or a fragment thereof (similar to the disease-associated antigen) is expressed on the surface of a cell such as an antigen-presenting cell (optionally in the context of MHC) so as to provide the relevant epitope for binding by immune effector cells. The vaccine antigen may be a recombinant antigen. In some embodiments of all aspects described herein, the RNA encoding the vaccine antigen is expressed in cells of a subject to provide the antigen or a procession product thereof for binding by the antigen receptor expressed by immune effector cells, said binding resulting in stimulation, priming and / or expansion of the immune effector cells.
[0288] An "antigen" according to the present disclosure covers any substance that will elicit an immune response and / or any substance against which an immune response or an immune mechanism such as a cellular response and / or humoral response is directed. This also includes situations wherein the antigen is processed into antigen peptides and an immune response or an immune mechanism is directed against one or more antigen peptides, in particular if presented in the context of MHC molecules. In particular, an "antigen" relates to any substance, such as a peptide or polypeptide, that reacts specifically with antibodies or T-lymphocytes (T-cells). The term "antigen" may comprise a molecule that comprises at least one epitope, such as a T cell epitope. In some embodiments, an antigen is a molecule which, optionally after processing, induces an immune reaction, which may be specific for the antigen (including cells expressing the antigen). In some embodiments, an antigen is a disease-associated antigen, such as an Mtb antigen.
[0289] In some embodiments, an antigen is presented or present on the surface of cells of the immune system such as antigen presenting cells like dendritic cells or macrophages. An antigen or a procession product thereof such as a T cell epitope is in some embodiments bound by an antigen receptor. Accordingly, an antigen or a procession product thereof may react specifically with immune effector cells such as T-lymphocytes (T cells).
[0290] According to the present disclosure, an antigen or a combination of antigens described herein may induce an immune response, wherein the immune response may comprise a humoral or cellular immune response, or both. In the context of some embodiments of the present disclosure, the antigen is presented by a cell, such as by an antigen presenting cell, in the context of MHC molecules, which results in an immune response against the antigen. An antigen may be a product which corresponds to or is derived from a naturally occurring antigen. According to the present disclosure, an antigen may correspond to a naturally occurring product.
[0291] The term "disease-associated antigen" is used in its broadest sense to refer to any antigen associated with a disease. In some embodiments, a disease-associated antigen is a molecule which contains epitopes that will stimulate a host's immune system to make a cellular antigen-specific immune response and / or a humoral antibody response against the disease. Disease-associated antigens include pathogen-associated antigens, i.e., antigens which are associated with infection by microbes, typically microbial antigens (such as bacterial or viral antigens, e.g., Mtb antigens), or antigens associated with cancer, typically tumors, such as tumor antigens.
[0292] The term "bacterial antigen" refers to any bacterial component having antigenic properties, i.e. being able to provoke an immune response in an individual. The bacterial antigen may be derived from the cell wall or cytoplasm membrane of the bacterium. The term "bacterial antigen" includes Mtb antigens, e.g., Mtb antigens as described herein.
[0293] The term "epitope" refers to an antigenic determinant in a molecule such as an antigen, i.e., to a part in or fragment of the molecule that is recognized by the immune system, for example, that is recognized by antibodies, T cells or B cells, in particular when presented in the context of MHC molecules. An epitope of a protein may comprises a continuous or discontinuous portion of said protein and, e.g., may be between about 5 and about 100, between about 5 and about 50, between about 8 and about 30, or about 10 and about 25 amino acids in length, for example, the epitope may be preferably 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids in length. In some embodiments, the epitope in the context of the present disclosure is a T cell epitope.
[0294] Terms such as "epitope", "fragment of an antigen", "immunogenic peptide" and "antigen peptide" are used interchangeably herein and, e.g., may relate to an incomplete representation of an antigen which is, e.g., capable of eliciting an immune response against the antigen or a cell expressing or comprising and presenting the antigen. In some embodiments, the terms relate to an immunogenic portion of an antigen. In some embodiments, it is a portion of an antigen that is recognized (Ze., specifically bound) by a T cell receptor, in particular if presented in the context of MHC molecules. Certain preferred immunogenic portions bind to an MHC class I or class II molecule. The term "epitope" refers to a part or fragment of a molecule such as an antigen that is recognized by the immune system. For example, the epitope may be recognized by T cells, B cells or antibodies. An epitope of an antigen may include a continuous or discontinuous portion of the antigen and may be between about 5 and about 100, such as between about 5 and about 50, between about 8 and about 30, or between about 8 and about 25 amino acids in length, for example, the epitope may be 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids in length. In some embodiments, an epitope is between about 10 and about 25 amino acids in length. The term "epitope" includes T cell epitopes.
[0295] The term "T cell epitope" refers to a part or fragment of a protein that is recognized by a T cell when presented in the context of MHC molecules, including epitopes predicted by bioinformatic means. The term "major histocompatibility complex" and the abbreviation "MHC" includes MHC class I and MHC class II molecules and relates to a complex of genes which is present in all vertebrates. MHC proteins or molecules are important for signaling between lymphocytes and antigen presenting cells or diseased cells in immune reactions, wherein the MHC proteins or molecules bind peptide epitopes and present them for recognition by T cell receptors on T cells. The proteins encoded by the MHC are expressed on the surface of cells, and display both self-antigens (peptide fragments from the cell itself) and non-self- antigens e.g, fragments of invading microorganisms) to a T cell. In the case of class I MHC / peptide complexes, the binding peptides are typically about 8 to about 10 amino acids long although longer or shorter peptides may be effective. In the case of class II MHC / peptide complexes, the binding peptides are typically about 10 to about 25 amino acids long and are in particular about 13 to about 18 amino acids long, whereas longer and shorter peptides may be effective.
[0296] In some embodiments, vaccine antigen, i.e., an antigen whose inoculation into a subject induces an immune response, is recognized by an immune effector cell. In some embodiments, the vaccine antigen if recognized by an immune effector cell is able to induce in the presence of appropriate co-stimulatory signals, stimulation, priming and / or expansion of the immune effector cell carrying an antigen receptor recognizing the vaccine antigen. In the context of the embodiments of the present disclosure, the vaccine antigen may be, e.g., presented or present on the surface of a cell, such as an antigen presenting cell.
[0297] In some embodiments, an antigen is expressed in a diseased cell (such as an infected cell).
[0298] In some embodiments, an antigen is presented by a diseased cell (such as an infected cell). In some embodiments, an antigen receptor is a TCR which binds to an epitope of an antigen presented in the context of MHC. In some embodiments, binding of a TCR when expressed by T cells and / or present on T cells to an antigen presented by cells such as antigen presenting cells results in stimulation, priming and / or expansion of said T cells. In some embodiments, binding of a TCR when expressed by T cells and / or present on T cells to an antigen presented on diseased cells results in cytolysis and / or apoptosis of the diseased cells, wherein said T cells release cytotoxic factors, e.g., perforins and granzymes.
[0299] In some embodiments, an antigen receptor is an antibody or B cell receptor which binds to an epitope in an antigen. In some embodiments, an antibody or B cell receptor binds to native epitopes of an antigen.
[0300] The terms "T cell" and 'T lymphocyte" are used interchangeably herein and include T helper cells (CD4+ T cells) and cytotoxic T cells (CTLs, CD8+ T cells) which comprise cytolytic T cells. The term "antigen-specific T cell" or similar terms relate to a T cell which recognizes the antigen to which the T cell is targeted, in particular when presented on the surface of antigen presenting cells or diseased cells in the context of MHC molecules and preferably exerts effector functions of T cells. T cells are considered to be specific for antigen if the cells kill target cells expressing an antigen. T cell specificity may be evaluated using any of a variety of standard techniques, for example, within a chromium release assay or proliferation assay. Alternatively, synthesis of lymphokines (such as interferon-y) can be measured.
[0301] In some embodiments, the term "target" shall mean an agent such as a cell or tissue which is a target for an immune response such as a cellular immune response. Targets include cells that present an antigen or an antigen epitope, i.e., a peptide fragment derived from an antigen. In some embodiments, the target cell is a cell expressing an antigen and presenting said antigen with class I MHC.
[0302] "Antigen processing" refers to the degradation of an antigen into processing products which are fragments of said antigen (e.g, the degradation of a polypeptide into peptides) and the association of one or more of these fragments e.g, via binding) with MHC molecules for presentation by cells, such as antigen-presenting cells to specific T-cells. Antigen-presenting cells can be distinguished in professional antigen presenting cells and non-professional antigen presenting cells.
[0303] The term "professional antigen presenting cells" relates to antigen presenting cells which constitutively express the Major Histocompatibility Complex class II (MHC class II) molecules required for interaction with naive T cells. If a T cell interacts with the MHC class II molecule complex on the membrane of the antigen presenting cell, the antigen presenting cell produces a co-stimulatory molecule inducing activation of the T cell. Professional antigen presenting cells comprise dendritic cells and macrophages.
[0304] The term "non-professional antigen presenting cells" relates to antigen presenting cells which do not constitutively express MHC class II molecules, but upon stimulation by certain cytokines such as interferon-gamma. Exemplary, non- professional antigen presenting cells include fibroblasts, thymic epithelial cells, thyroid epithelial cells, glial cells, pancreatic beta cells or vascular endothelial cells.
[0305] The term "dendritic cell" (DC) refers to a subtype of phagocytic cells belonging to the class of antigen presenting cells. In some embodiments, dendritic cells are derived from hematopoietic bone marrow progenitor cells. These progenitor cells initially transform into immature dendritic cells. These immature cells are characterized by high phagocytic activity and low T cell activation potential. Immature dendritic cells constantly sample the surrounding environment for pathogens such as viruses and bacteria. Once they have come into contact with a presentable antigen, they become activated into mature dendritic cells and begin to migrate to the spleen or to the lymph node. Immature dendritic cells phagocytose pathogens and degrade their proteins into small pieces and upon maturation present those fragments at their cell surface using MHC molecules. Simultaneously, they upregulate cell-surface receptors that act as co-receptors in T cell activation such as CD80, CD86, and CD40 greatly enhancing their ability to activate T cells. They also upregulate CCR7, a chemotactic receptor that induces the dendritic cell to travel through the blood stream to the spleen or through the lymphatic system to a lymph node. Here they act as antigen-presenting cells and activate helper T cells and killer T cells as well as B cells by presenting them antigens, alongside non-antigen specific co-stimulatory signals. Thus, dendritic cells can actively induce a T cell- or B cell-related immune response. In some embodiments, the dendritic cells are splenic dendritic cells.
[0306] The term "macrophage" refers to a subgroup of phagocytic cells produced by the differentiation of monocytes. Macrophages which are activated by inflammation, immune cytokines or microbial products nonspecifically engulf and kill foreign pathogens within the macrophage by hydrolytic and oxidative attack resulting in degradation of the pathogen. Peptides from degraded proteins are displayed on the macrophage cell surface where they can be recognized by T cells, and they can directly interact with antibodies on the B cell surface, resulting in T and B cell activation and further stimulation of the immune response. Macrophages belong to the class of antigen presenting cells. In some embodiments, the macrophages are splenic macrophages.
[0307] By "antigen-responsive CTL" is meant a CD8+T-cell that is responsive to an antigen or a peptide derived from said antigen, which is presented with class I MHC on the surface of antigen presenting cells.
[0308] According to the disclosure, CTL responsiveness may include sustained calcium flux, cell division, production of cytokines such as IFN-y and TNF-a, up-regulation of activation markers such as CD44 and CD69, and specific cytolytic killing of tumor antigen expressing target cells. CTL responsiveness may also be determined using an artificial reporter that accurately indicates CTL responsiveness.
[0309] "Activation" or "stimulation", as used herein, refers to the state of a cell that has been sufficiently stimulated to induce detectable cellular proliferation, such as an immune effector cell such as T cell. Activation can also be associated with initiation of signaling pathways, induced cytokine production, and detectable effector functions. The term "activated immune effector cells" refers to, among other things, immune effector cells that are undergoing cell division.
[0310] The term "priming" refers to a process wherein an immune effector cell such as a T cell has its first contact with its specific antigen and causes differentiation into effector cells such as effector T cells.
[0311] The term "expansion" refers to a process wherein a specific entity is multiplied. In some embodiments, the term is used in the context of an immunological response in which immune effector cells are stimulated by an antigen, proliferate, and the specific immune effector cell recognizing said antigen is amplified. In some embodiments, expansion leads to differentiation of the immune effector cells.
[0312] The terms "immune response" and "immune reaction" are used herein interchangeably in their conventional meaning and refer to an integrated bodily response to an antigen and may refer to a cellular immune response, a humoral immune response, or both. According to the disclosure, the term "immune response to" or "immune response against" with respect to an agent such as an antigen, cell or tissue, relates to an immune response such as a cellular response directed against the agent. An immune response may comprise one or more reactions selected from the group consisting of developing antibodies against one or more antigens and expansion of antigen-specific T-lymphocytes, such as CD4+and CD8+T-lymphocytes, e.g. CD8+T-lymphocytes, which may be detected in various proliferation or cytokine production tests in vitro.
[0313] The terms "inducing an immune response" and "eliciting an immune response" and similar terms in the context of the present disclosure refer to the induction of an immune response, such as the induction of a cellular immune response, a humoral immune response, or both. The immune response may be protective / preventive / prophylactic and / or therapeutic. The immune response may be directed against any immunogen or antigen or antigen peptide, such as against a pathogen-associated antigen e.g, an antigen of Mtb). "Inducing" in this context may mean that there was no immune response against a particular antigen or pathogen before induction, but it may also mean that there was a certain level of immune response against a particular antigen or pathogen before induction and after induction said immune response is enhanced. Thus, "inducing the immune response" in this context also includes "enhancing the immune response". In some embodiments, after inducing an immune response in an individual, said individual is protected from developing a disease such as an infectious disease or the disease condition is ameliorated by inducing an immune response.
[0314] The terms "cellular immune response", "cellular response", "cell-mediated immunity" or similar terms are meant to include a cellular response directed to cells characterized by expression of an antigen and / or presentation of an antigen with class I or class II MHC. The cellular response relates to cells called T cells or T lymphocytes which act as either "helpers" or "killers". The helper T cells (also termed CD4+T cells) play a central role by regulating the immune response and the killer cells (also termed cytotoxic T cells, cytolytic T cells, CD8+T cells or CTLs) kill cells such as diseased cells.
[0315] The term "humoral immune response" refers to a process in living organisms wherein antibodies are produced in response to agents and organisms, which they ultimately neutralize and / or eliminate. The specificity of the antibody response is mediated by T and / or B cells through membrane-associated receptors that bind antigen of a single specificity. Following binding of an appropriate antigen and receipt of various other activating signals, B lymphocytes divide, which produces memory B cells as well as antibody secreting plasma cell clones, each producing antibodies that recognize the identical antigenic epitope as was recognized by its antigen receptor. Memory B lymphocytes remain dormant until they are subsequently activated by their specific antigen. These lymphocytes provide the cellular basis of memory and the resulting escalation in antibody response when re-exposed to a specific antigen.
[0316] The term "antibody" as used herein, refers to an immunoglobulin molecule, which is able to specifically bind to an epitope on an antigen. In particular, the term "antibody" refers to a glycoprotein comprising at least two heavy (H) chains and two light (L) chains inter-connected by disulfide bonds. The term "antibody" includes monoclonal antibodies, recombinant antibodies, human antibodies, humanized antibodies, chimeric antibodies and combinations of any of the foregoing. Each heavy chain is comprised of a heavy chain variable region (VH) and a heavy chain constant region (CH). Each light chain is comprised of a light chain variable region (VL) and a light chain constant region (CL). The variable regions and constant regions are also referred to herein as variable domains and constant domains, respectively. The VH and VL regions can be further subdivided into regions of hypervariability, termed complementarity determining regions (CDRs), interspersed with regions that are more conserved, termed framework regions (FRs). Each VH and VL is composed of three CDRs and four FRs, arranged from amino-terminus to carboxy-terminus in the following order: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4. The CDRs of a VH are termed HCDR1, HCDR2 and HCDR3, the CDRs of a VL are termed LCDR1, LCDR2 and LCDR3. The variable regions of the heavy and light chains contain a binding domain that interacts with an antigen. The constant regions of an antibody comprise the heavy chain constant region (CH) and the light chain constant region (CL), wherein CH can be further subdivided into constant domain CHI, a hinge region, and constant domains CH2 and CH3 (arranged from amino-terminus to carboxy-terminus in the following order: CHI, CH2, CH3). The constant regions of the antibodies may mediate the binding of the immunoglobulin to host tissues or factors, including various cells of the immune system e.g, effector cells) and the first component (Clq) of the classical complement system. Antibodies can be intact immunoglobulins derived from natural sources or from recombinant sources and can be immunoactive portions of intact immunoglobulins. Antibodies are typically tetramers of immunoglobulin molecules. Antibodies may exist in a variety of forms including, for example, polyclonal antibodies, monoclonal antibodies, Fv, Fab and F(ab)2, as well as single chain antibodies and humanized antibodies.
[0317] The term "immunoglobulin" relates to proteins of the immunoglobulin superfamily, such as to antigen receptors such as antibodies or the B cell receptor (BCR). The immunoglobulins are characterized by a structural domain, i.e., the immunoglobulin domain, having a characteristic immunoglobulin (Ig) fold. The term encompasses membrane bound immunoglobulins as well as soluble immunoglobulins. Membrane bound immunoglobulins are also termed surface immunoglobulins or membrane immunoglobulins, which are generally part of the BCR. Soluble immunoglobulins are generally termed antibodies. Immunoglobulins generally comprise several chains, typically two identical heavy chains and two identical light chains which are linked via disulfide bonds. These chains are primarily composed of immunoglobulin domains, such as the VL (variable light chain) domain, CL (constant light chain) domain, VH (variable heavy chain) domain, and the CH (constant heavy chain) domains CHI, CH2, CH3, and CH4. There are five types of mammalian immunoglobulin heavy chains, i.e., a, 8, E, y, and |j which account for the different classes of antibodies, i.e., IgA, IgD, IgE, IgG, and IgM. As opposed to the heavy chains of soluble immunoglobulins, the heavy chains of membrane or surface immunoglobulins comprise a transmembrane domain and a short cytoplasmic domain at their carboxy-terminus. In mammals there are two types of light chains, i.e., lambda and kappa. The immunoglobulin chains comprise a variable region and a constant region. The constant region is essentially conserved within the different isotypes of the immunoglobulins, wherein the variable part is highly divers and accounts for antigen recognition.
[0318] The terms "vaccination" and "immunization" describe the process of treating an individual for therapeutic or prophylactic reasons and relate to the procedure of administering one or more immunogen(s) or antigen(s) or derivatives thereof, in particular in the form of RNA (especially mRNA) coding therefor, as described herein to an individual and stimulating an immune response against said one or more immunogen(s) or antigen(s) or cells characterized by presentation of said one or more immunogen(s) or antigen(s).
[0319] By "cell characterized by presentation of an antigen" or "cell presenting an antigen" or "MHC molecules which present an antigen on the surface of an antigen presenting cell" or similar expressions is meant a cell such as a diseased cell, in particular an infected cell, or an antigen presenting cell presenting the antigen or an antigen peptide, either directly or following processing, in the context of MHC molecules, such as MHC class I and / or MHC class II molecules. In some embodiments, the MHC molecules are MHC class I molecules.
[0320] Embodiments of antigen-coding RNA
[0321] Generally, at least four formats useful for RNA pharmaceutical compositions may be used herein, namely non-modified uridine containing mRNA (uRNA), nucleoside modified mRNA (modRNA), self-amplifying RNA (saRNA), and transamplifying RNAs.
[0322] Features of modified uridine (e.g., pseudouridine) platform may include reduced adjuvant effect, blunted immune innate immune sensor activating capacity and thus augmented polypeptide {e.g., protein) expression.
[0323] Features of self-amplifying platform may include, for example, long duration of polypeptide e.g., protein) expression, good tolerability and safety, higher likelihood for efficacy with very low RNA dose.
[0324] In some embodiments, a self-amplifying platform {e.g., RNA) comprises two nucleic acid molecules, wherein one nucleic acid molecule encodes a replicase {e.g., a viral replicase) and the other nucleic acid molecule is capable of being replicated {e.g., a replicon) by said replicase in trans (f / a / zs-replication system). In some embodiments, a selfamplifying platform {e.g., RNA) comprises a plurality of nucleic acid molecules, wherein said nucleic acids encode a plurality of replicases and / or replicons.
[0325] In some embodiments, a / rans-replication system comprises the presence of both nucleic acid molecules in a single host cell.
[0326] In some such embodiments, a nucleic acid encoding a replicase {e.g., a viral replicase) is not capable of self-replication in a target cell and / or target organism. In some such embodiments, a nucleic acid encoding a replicase {e.g., a viral replicase) lacks at least one conserved sequence element important for (-) strand synthesis based on a (+) strand template and / or for (+) strand synthesis based on a (-) strand template.
[0327] In some embodiments, a self-amplifying RNA comprises a 3' untranslated region (UTR), a 5' UTR, a cap structure, a poly adenine (polyA) tail, and any combinations thereof. In some embodiments, a self-amplifying platform does not require propagation of virus particles e.g., is not associated with undesired virus-particle formation). In some embodiments, a self-amplifying platform is not capable of forming virus particles.
[0328] In some embodiments, RNA e.g., a single stranded RNA) described herein has a length of at least 500 ribonucleotides (such as, e.g., at least 600 ribonucleotides, at least 700 ribonucleotides, at least 800 ribonucleotides, at least 900 ribonucleotides, at least 1000 ribonucleotides, at least 1250 ribonucleotides, at least 1500 ribonucleotides, at least 1750 ribonucleotides, at least 2000 ribonucleotides, at least 2500 ribonucleotides, at least 3000 ribonucleotides, at least 3500 ribonucleotides, at least 4000 ribonucleotides, at least 4500 ribonucleotides, at least 5000 ribonucleotides, or longer). In some embodiments, RNA described herein is single-stranded RNA having a length of about 800 ribonucleotides to 5000 ribonucleotides.
[0329] In some embodiments, a relevant RNA includes a polypeptide-encoding portion or a plurality of polypeptide-encoding portions. In some particular embodiments, such a portion or portions encode one or more polypeptides which are not endogenous (i.e., it is foreign) to the subject treated.
[0330] In some embodiments, the RNA described herein (e.g., contained in the compositions / formulations of the present disclosure and / or used in the methods of the present disclosure) is single-stranded RNA (in particular, mRNA) that may be translated into the respective protein upon entering cells, e.g., cells of a recipient, e.g., muscle cells or antigen- presenting cells (APCs). In addition to wild-type or codon-optimized sequences encoding an antigen sequence, the RNA may contain one or more structural elements optimized for maximal efficacy of the RNA with respect to stability and translational efficiency (5' cap, 5' UTR, 3' UTR, poly(A)-tail). In some embodiments, the RNA contains all of these elements. In some embodiments, beta-S-ARCA(Dl) (m27'2' °GppSpG) or m27'3' °Gppp(mi2' °)ApG may be utilized as specific capping structure at the 5'-end of the RNA. As 5'-UTR sequence, the 5'-UTR sequence of the human alphaglobin mRNA, optionally with an optimized 'Kozak sequence' to increase translational efficiency may be used. As 3'- UTR sequence, a combination of two sequence elements (FI element) derived from the "amino terminal enhancer of split" (AES) mRNA (called F) and the mitochondrial encoded 12S ribosomal RNA (called I) placed between the coding sequence and the poly(A)-tail to assure higher maximum protein levels and prolonged persistence of the mRNA may be used (see WO 2017 / 060314, herein incorporated by reference). Furthermore, a poly(A)-tail measuring 110 nucleotides in length, consisting of a stretch of 30 adenosine residues, followed by a 10 nucleotide linker sequence (of random nucleotides) and another 70 adenosine residues may be used. In some embodiments, the 5'-UTR comprises the nucleotide sequence of SEQ ID NO: 53, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 53. In some embodiments, the 3'-UTR comprises the nucleotide sequence of SEQ ID NO: 54, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 54. In some embodiments, the poly(A) sequence comprises the nucleotide sequence of SEQ ID NO: 55, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 55.
[0331] In some embodiments, the RNA described herein is not chemically modified, i.e. it solely contains naturally occurring nucleosides, and preferably has the composition of naturally occurring RNA.
[0332] In some embodiments, the RNA described herein is modified for optimized efficacy of the RNA (e.g., increased translation efficacy, decreased immunogenicity, and / or decreased cytotoxicity) (e.g., by replacing (partially or completely, preferably completely) naturally occurring nucleosides (in particular uridine) with synthetic nucleosides (e.g., modified nucleosides, e.g., selected from the group consisting of pseudouridine (ip), Nl-methyl-pseudouridine (mlip), and 5-methyl-uridine); and / or codon-optimization). In some embodiments, the RNA comprises a modified nucleoside in place of uridine. In some embodiments, the modified nucleoside replacing (partially or completely, preferably completely) uridine is selected from the group consisting of pseudouridine (i ), Nl-methyl-pseudouridine (mli ), and 5-methyl-uridine. In some embodiments, the RNA encoding the vaccine antigen has a coding sequence (a) which is codon-optimized, (b) the G / C content of which is increased compared to the wild type coding sequence, or (c) both (a) and (b).
[0333] In some embodiments, the RNA described herein comprises a 5' cap, a 5' UTR, a 3' UTR, and a poly(A) sequence (e.g., as described above); is modified by replacing (partially or completely, preferably completely) uridine with modified nucleosides, e.g., selected from the group consisting of pseudouridine (i ), Nl-methyl-pseudouridine (mlip), and 5- methyl-uridine; and has a coding sequence which is codon-optimized, and the G / C content of which is increased compared to the wild type coding sequence.
[0334] In some embodiments, if the present disclosure provides for a mixture of different RNA molecules, a composition comprising different RNA molecules or an administration of different RNA molecules, these different RNA molecules are present in approximately the same amount. Such different RNA molecules may be formulated in individual particulate formulations, mixed particulate formulations, or combined particulate formulations as described herein.
[0335] The present disclosure provides RNA (in particular, mRNA) comprising a nucleic acid sequence encoding an Mtb antigen, an immunogenic variant thereof, or an immunogenic fragment of the Mtb antigen or the immunogenic variant thereof.
[0336] In some embodiments, RNA (in particular, mRNA) described in the present disclosure comprises a nucleic acid sequence encoding an Mtb antigen, an immunogenic variant thereof, or an immunogenic fragment of the Mtb antigen or the immunogenic variant thereof, and is capable of expressing said Mtb antigen, immunogenic variant, or immunogenic fragment, in particular if transferred into a cell or subject, preferably a human cell or subject. Thus, in some embodiments, the RNA (in particular, mRNA) described in the present disclosure contains a coding region (open reading frame (ORF)) encoding an Mtb antigen, an immunogenic variant thereof, or an immunogenic fragment of the Mtb antigen or the immunogenic variant thereof.
[0337] In some embodiments, RNA comprises a nucleic acid sequence encoding more than one Mtb antigen, immunogenic variant thereof, or immunogenic fragment of the Mtb antigen or the immunogenic variant thereof, e.g., two, three, four or more Mtb antigens, immunogenic variants thereof, or immunogenic fragments of the Mtb antigens or the immunogenic variants thereof. In some embodiments, two or more of such Mtb antigens, immunogenic variants thereof, or immunogenic fragments of the Mtb antigens or the immunogenic variants thereof are present as a fusion protein.
[0338] In some embodiments, the one or more Mtb antigens, immunogenic variants thereof, or immunogenic fragments of the Mtb antigens or the immunogenic variants thereof encoded by the RNA may comprise or consist of naturally occurring sequences, may comprise or consist of variants of naturally occurring sequences, or may comprise or consist of sequences which are not naturally occurring, e.g., recombinant sequences. In some embodiments, the peptide or polypeptide encoded by the RNA described herein may consist of the one or more Mtb antigens, immunogenic variants thereof, or immunogenic fragments of the Mtb antigens or the immunogenic variants thereof, or may comprise the one or more Mtb antigens, immunogenic variants thereof, or immunogenic fragments of the Mtb antigens or the immunogenic variants thereof and may comprise additional sequences such as secretion signals, extended-PK groups, tags and any other sequences. In some embodiments, the additional sequences are fused to the one or more Mtb antigens, immunogenic variants thereof, or immunogenic fragments of the Mtb antigens or the immunogenic variants thereof, in some embodiments, separated by a linker. In these embodiments, the one or more Mtb antigens, immunogenic variants thereof, or immunogenic fragments of the Mtb antigens or the immunogenic variants thereof may be considered the pharmaceutically active peptide or polypeptide even if additional sequences support the function or effect of the one or more Mtb antigens, immunogenic variants thereof, or immunogenic fragments of the Mtb antigens or the immunogenic variants thereof.
[0339] According to the present disclosure, the term "pharmaceutically active peptide or polypeptide" means a peptide or polypeptide that can be used in the treatment of an individual where the expression of the peptide or polypeptide would be of benefit, e.g., in ameliorating the symptoms of a disease. Preferably, a pharmaceutically active peptide or polypeptide has curative or palliative properties and may be administered to ameliorate, relieve, alleviate, reverse, delay onset of or lessen the severity of one or more symptoms of a disease. In some embodiments, a pharmaceutically active peptide or polypeptide has a positive or advantageous effect on the condition or disease state of an individual when administered to the individual in a therapeutically effective amount. A pharmaceutically active peptide or polypeptide may have prophylactic properties and may be used to delay the onset of a disease or to lessen the severity of such disease. The term "pharmaceutically active peptide or polypeptide" includes entire peptides or polypeptides, and can also refer to pharmaceutically active fragments thereof. It can also include pharmaceutically active variants and / or analogs of a peptide or polypeptide.
[0340] Specific examples of pharmaceutically active peptides and polypeptides include, but are not limited to, antigens for vaccination such as Mtb antigens, immunogenic variants thereof, or immunogenic fragments of the Mtb antigens or the immunogenic variants thereof.
[0341] Mtb antigens, immunogenic variants thereof, or immunogenic fragments of the Mtb antigens or the immunogenic variants thereof described herein can be prepared as fusion or chimeric polypeptides that include a portion which corresponds to three or more Mtb antigens, immunogenic variants thereof, or immunogenic fragments of the Mtb antigens or the immunogenic variants thereof and a heterologous polypeptide (i.e., a polypeptide that is not an Mtb antigen, immunogenic variant thereof, or immunogenic fragment of the Mtb antigen or the immunogenic variant thereof).
[0342] According to certain embodiments, a "signal peptide" (or signal sequence) is fused, either directly or through a linker, to the N-terminus of a chimeric protein described herein.
[0343] In some embodiments, an open reading frame of the RNA described herein encodes a polypeptide that includes a signal sequence, e.g., that is functional in mammalian cells.
[0344] In some embodiments, a utilized signal sequence is "intrinsic" in that it is, in nature, associated with (e.g., linked to) the full-length antigen or antigen fragment whose sequences are included in the encoded chimeric protein.
[0345] In some embodiments, a utilized signal sequence is non-native to the encoded polypeptide - e.g., is not naturally part of a full-length antigen or antigen fragment whose sequences are included in the encoded chimeric protein.
[0346] In some embodiments, signal peptides are sequences, which are typically characterized by a length of about 15 to 30 amino acids.
[0347] In many embodiments, signal peptides are positioned at the N-terminus of an encoded chimeric protein as described herein, without being limited thereto. In some embodiments, signal peptides preferably allow the transport of the polypeptide encoded by RNAs of the present disclosure with which they are associated into a defined cellular compartment, preferably the cell surface, the endoplasmic reticulum (ER) or the endosomal-lysosomal compartment.
[0348] In some embodiments, an RNA sequence encodes a peptidoglycan hydrolase, e.g., an endolysin, that may comprise or otherwise be linked to a signal sequence (e.g., secretory sequence), such as those listed in Table 2 and 3, or a sequence having 1, 2, 3, 4, or 5 amino acid differences relative thereto. In some embodiments, a signal sequence such as MRVMAPRTLILLLSGALALTETWAGS [SEQ ID NO: 17], or a sequence having 1, 2, 3, 4, or at the most 5 amino acid differences relative thereto is utilized.
[0349] In some embodiments, a signal peptide is selected from those included in the Table 2 below and / or those encoded by the sequences in Table 3 below or a sequence having 1, 2, 3, 4, or 5 amino acid differences relative thereto:
[0350] Table 2: Exemplary signal sequences
[0351] Table 3: Exemplary nucleotide sequences encoding signal sequences
[0352] According to some embodiments, an amino acid sequence enhancing antigen processing and / or presentation is fused, either directly or through a linker, to an antigenic peptide or polypeptide (antigenic sequence), e.g., one or more Mtb antigens, immunogenic variants thereof, or immunogenic fragments of the Mtb antigens or the immunogenic variants thereof. Accordingly, in some embodiments, the RNA described herein comprises at least one coding region encoding an antigenic peptide or polypeptide and an amino acid sequence enhancing antigen processing and / or presentation.
[0353] In preferred embodiments, the transmembrane domain and / or trafficking signal is positioned at the C-terminus of a chimeric protein described herein.
[0354] In some embodiments, the amino acid sequence enhancing antigen processing and / or presentation includes, without being limited thereto, sequences derived from Human herpes simplex virus 1, envelope glycoprotein D (HSV-gDl), in particular a sequence comprising the amino acid sequence of SEQ ID NO: 140. Such sequence is designated herein as HSV-TMD. In some embodiments, an amino acid sequence enhancing antigen processing and / or presentation comprises the amino acid sequence of SEQ ID NO: 140, an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 140, or a functional fragment of the amino acid sequence of SEQ ID NO: 140, or the amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 140. In some embodiments, an amino acid sequence enhancing antigen processing and / or presentation comprises the amino acid sequence of SEQ ID NO: 140.
[0355] In some embodiments, the amino acid sequence enhancing antigen processing and / or presentation as defined herein includes, without being limited thereto, sequences derived from the human MHC class I complex (HLA-B51, haplotype A2, B27 / B51, Cw2 / Cw3), in particular a sequence comprising the amino acid sequence of SEQ ID NO: 52 or a functional variant thereof. Such sequence, which includes a transmembrane domain and a trafficking signal derived from human MHC class I, is designated herein as MITD.
[0356] In some embodiments, an amino acid sequence enhancing antigen processing and / or presentation comprises the amino acid sequence of SEQ ID NO: 52, an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 52, or a functional fragment of the amino acid sequence of SEQ ID NO: 52, or the amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 52. In some embodiments, an amino acid sequence enhancing antigen processing and / or presentation comprises the amino acid sequence of SEQ ID NO: 52.
[0357] Accordingly, in some embodiments, the RNA described herein comprises at least one coding region encoding an antigenic peptide or polypeptide and an amino acid sequence enhancing antigen processing and / or presentation, said amino acid sequence enhancing antigen processing and / or presentation preferably being fused to the antigenic peptide or polypeptide, more preferably to the C-terminus of the antigenic peptide or polypeptide as described herein.
[0358] Furthermore, a secretory sequence, e.g., a sequence comprising the amino acid sequence selected from SEQ ID NOs: 17 to 37 or 65 or encoded by the nucleotide sequence selected from SEQ ID NOs: 38 to 51, may be fused to the N- terminus of the antigenic peptide or polypeptide.
[0359] Mtb antigens, immunogenic variants thereof, or immunogenic fragments of the Mtb antigens or the immunogenic variants thereof may be fused to an extended-PK group, which increases circulation half-life. Non-limiting examples of extended-PK groups are described herein. It should be understood that other PK groups that increase the circulation half-life of peptides or polypeptides such as Mtb antigens, immunogenic variants thereof, or immunogenic fragments of the Mtb antigens or the immunogenic variants thereof are also applicable to the present disclosure. In certain embodiments, the extended-PK group is a serum albumin domain (e.g., mouse serum albumin, human serum albumin, or recombinant serum albumin).
[0360] As used herein, the term "PK" is an acronym for "pharmacokinetic" and encompasses properties of a compound including, by way of example, absorption, distribution, metabolism, and elimination by a subject. As used herein, an "extended-PK group" refers to a protein, peptide, or moiety that increases the circulation half-life of a biologically active molecule when fused to or administered together with the biologically active molecule. Examples of an extended-PK group include serum albumin (e.g., HSA), Immunoglobulin Fc or Fc fragments and variants thereof, transferrin and variants thereof, and human serum albumin (HSA) binders (as disclosed in U.S. Publication Nos. 2005 / 0287153 and 2007 / 0003549). Other exemplary extended-PK groups are disclosed in Kontermann, Expert Opin Biol Ther, 2016 Jul;16(7):903-15 which is herein incorporated by reference in its entirety. As used herein, an "extended-PK" polypeptide refers to a polypeptide moiety such as an Mtb antigen, immunogenic variant thereof, or immunogenic fragment of the Mtb antigen or the immunogenic variant thereof in combination with an extended-PK group. In some embodiments, the extended-PK polypeptide is a fusion protein in which a polypeptide moiety is linked or fused to an extended-PK group. In certain embodiments, the serum half-life of an extended-PK polypeptide is increased relative to the polypeptide alone (i.e., the polypeptide not fused to an extended-PK group). In certain embodiments, the serum half-life of the extended-PK polypeptide is at least 20%, at least 40%, at least 60%, at least 80%, at least 100%, at least 120%, at least 150%, at least 180%, at least 200%, at least 400%, at least 600%, at least 800%, or at least 1000% longer relative to the serum half-life of the polypeptide alone. In certain embodiments, the serum half-life of the extended- PK polypeptide is at least 1.5-fold, 2-fold, 2.5-fold, 3-fold, 3.5-fold, 4-fold, 4.5-fold, 5-fold, 6-fold, 7-fold, 8-fold, 10- fold, 12-fold, 13-fold, 15-fold, 17-fold, 20-fold, 22-fold, 25-fold, 27-fold, 30-fold, 35-fold, 40-fold, or 50-fold greater than the serum half-life of the polypeptide alone. In certain embodiments, the serum half-life of the extended-PK polypeptide is at least 10 hours, 15 hours, 20 hours, 25 hours, 30 hours, 35 hours, 40 hours, 50 hours, 60 hours, 70 hours, 80 hours, 90 hours, 100 hours, 110 hours, 120 hours, 130 hours, 135 hours, 140 hours, 150 hours, 160 hours, or 200 hours.
[0361] As used herein, "half-life" refers to the time taken for the serum or plasma concentration of a compound such as a peptide or polypeptide to reduce by 50%, in vivo, for example due to degradation and / or clearance or sequestration by natural mechanisms. An extended-PK polypeptide suitable for use herein is stabilized in vivo and its half-life increased by, e.g., fusion to serum albumin (e.g., human serum albumin (HSA) or mouse serum albumin (MSA)), which resist degradation and / or clearance or sequestration. The half-life can be determined in any manner known per se, such as by pharmacokinetic analysis. Suitable techniques will be clear to the person skilled in the art, and may for example generally involve the steps of suitably administering a suitable dose of the amino acid sequence or compound to a subject; collecting blood samples or other samples from said subject at regular intervals; determining the level or concentration of the amino acid sequence or compound in said blood sample; and calculating, from (a plot of) the data thus obtained, the time until the level or concentration of the amino acid sequence or compound has been reduced by 50% compared to the initial level upon dosing. Further details are provided in, e.g., standard handbooks, such as Kenneth, A. et al., Chemical Stability of Pharmaceuticals: A Handbook for Pharmacists and in Peters et al., Pharmacokinetic Analysis: A Practical Approach (1996). Reference is also made to Gibaldi, M. et al., Pharmacokinetics, 2nd Rev. Edition, Marcel Dekker (1982).
[0362] In certain embodiments, the extended-PK group includes serum albumin, or fragments thereof or variants of the serum albumin or fragments thereof (all of which for the purpose of the present disclosure are comprised by the term "albumin"). Polypeptides described herein may be fused to albumin (or a fragment or variant thereof) to form albumin fusion proteins. Such albumin fusion proteins are described in U.S. Publication No. 20070048282.
[0363] As used herein, "albumin fusion protein" refers to a protein formed by the fusion of at least one molecule of albumin (or a fragment or variant thereof) to at least one molecule of a protein such as a therapeutic protein, in particular an Mtb antigen, immunogenic variant thereof, or immunogenic fragment of the Mtb antigen or the immunogenic variant thereof. The albumin fusion protein may be generated by translation of a nucleic acid in which a polynucleotide encoding a therapeutic protein is joined in-frame with a polynucleotide encoding an albumin. The therapeutic protein and albumin, once part of the albumin fusion protein, may each be referred to as a "portion", "region" or "moiety" of the albumin fusion protein (e.g., a "therapeutic protein portion" or an "albumin protein portion"). In a highly preferred embodiment, an albumin fusion protein comprises at least one molecule of a therapeutic protein (including, but not limited to a mature form of the therapeutic protein) and at least one molecule of albumin (including but not limited to a mature form of albumin). In some embodiments, an albumin fusion protein is processed by a host cell such as a cell of the target organ for administered RNA, e.g. a liver cell, and secreted into the circulation. Processing of the nascent albumin fusion protein that occurs in the secretory pathways of the host cell used for expression of the RNA may include, but is not limited to signal peptide cleavage; formation of disulfide bonds; proper folding; addition and processing of carbohydrates (such as for example, N- and O-linked glycosylation); specific proteolytic cleavages; and / or assembly into multimeric proteins. An albumin fusion protein is preferably encoded by RNA in a non-processed form which in particular has a signal peptide at its N-terminus and following secretion by a cell is preferably present in the processed form wherein in particular the signal peptide has been cleaved off. In a most preferred embodiment, the "processed form of an albumin fusion protein" refers to an albumin fusion protein product which has undergone N- terminal signal peptide cleavage, herein also referred to as a "mature albumin fusion protein".
[0364] In preferred embodiments, albumin fusion proteins comprising a therapeutic protein have a higher plasma stability compared to the plasma stability of the same therapeutic protein when not fused to albumin. Plasma stability typically refers to the time period between when the therapeutic protein is administered in / / Vo and carried into the bloodstream and when the therapeutic protein is degraded and cleared from the bloodstream, into an organ, such as the kidney or liver, that ultimately clears the therapeutic protein from the body. Plasma stability is calculated in terms of the half-life of the therapeutic protein in the bloodstream. The half-life of the therapeutic protein in the bloodstream can be readily determined by common assays known in the art.
[0365] As used herein, "albumin" refers collectively to albumin protein or amino acid sequence, or an albumin fragment or variant, having one or more functional activities (e.g., biological activities) of albumin. In particular, "albumin" refers to human albumin or fragments or variants thereof especially the mature form of human albumin, or albumin from other vertebrates or fragments thereof, or variants of these molecules. The albumin may be derived from any vertebrate, especially any mammal, for example human, cow, sheep, or pig. Non-mammalian albumins include, but are not limited to, hen and salmon. The albumin portion of the albumin fusion protein may be from a different animal than the therapeutic protein portion.
[0366] In certain embodiments, the albumin is human serum albumin (HSA), or fragments or variants thereof, such as those disclosed in US 5,876,969, WO 2011 / 124718, WO 2013 / 075066, and WO 2011 / 0514789.
[0367] The terms, human serum albumin (HSA) and human albumin (HA) are used interchangeably herein. The terms, "albumin and "serum albumin" are broader, and encompass human serum albumin (and fragments and variants thereof) as well as albumin from other species (and fragments and variants thereof).
[0368] As used herein, a fragment of albumin sufficient to prolong the therapeutic activity or plasma stability of the therapeutic protein refers to a fragment of albumin sufficient in length or structure to stabilize or prolong the therapeutic activity or plasma stability of the protein so that the plasma stability of the therapeutic protein portion of the albumin fusion protein is prolonged or extended compared to the plasma stability in the non-fusion state.
[0369] The albumin portion of the albumin fusion proteins may comprise the full length of the albumin sequence, or may include one or more fragments thereof that are capable of stabilizing or prolonging the therapeutic activity or plasma stability. Such fragments may be of 10 or more amino acids in length or may include about 15, 20, 25, 30, 50, or more contiguous amino acids from the albumin sequence or may include part or all of specific domains of albumin. For instance, one or more fragments of HSA spanning the first two immunoglobulin-like domains may be used. In a preferred embodiment, the HSA fragment is the mature form of HSA.
[0370] Generally speaking, an albumin fragment or variant will be at least 100 amino acids long, preferably at least 150 amino acids long.
[0371] According to the disclosure, albumin may be naturally occurring albumin or a fragment or variant thereof. Albumin may be human albumin and may be derived from any vertebrate, especially any mammal. Preferably, the albumin fusion protein comprises albumin as the N-terminal portion, and a therapeutic protein as the C-terminal portion. Alternatively, an albumin fusion protein comprising albumin as the C-terminal portion, and a therapeutic protein as the N-terminal portion may also be used. In other embodiments, the albumin fusion protein has a therapeutic protein fused to both the N-terminus and the C-terminus of albumin. In a preferred embodiment, the therapeutic proteins fused at the N- and C-termini are the same therapeutic proteins. In another preferred embodiment, the therapeutic proteins fused at the N- and C-termini are different therapeutic proteins. In some embodiments, the different therapeutic proteins are both Mtb antigens, immunogenic variants thereof, or immunogenic fragments of the Mtb antigens or the immunogenic variants thereof.
[0372] In some embodiments, the therapeutic protein(s) is (are) joined to the albumin through (a) peptide linker(s). A peptide linker between the fused portions may provide greater physical separation between the moieties and thus maximize the accessibility of the therapeutic protein portion, for instance, for binding to its cognate receptor. The peptide linker may consist of amino acids such that it is flexible or more rigid. The linker sequence may be cleavable by a protease or chemically.
[0373] As used herein, the term "Fc region" refers to the portion of a native immunoglobulin formed by the respective Fc domains (or Fc moieties) of its two heavy chains. As used herein, the term "Fc domain" refers to a portion or fragment of a single immunoglobulin (Ig) heavy chain wherein the Fc domain does not comprise an Fv domain. In certain embodiments, an Fc domain begins in the hinge region just upstream of the papain cleavage site and ends at the C- terminus of the antibody. Accordingly, a complete Fc domain comprises at least a hinge domain, a CH2 domain, and a CH3 domain. In certain embodiments, an Fc domain comprises at least one of: a hinge (e.g., upper, middle, and / or lower hinge region) domain, a CH2 domain, a CH3 domain, a CH4 domain, or a variant, portion, or fragment thereof. In certain embodiments, an Fc domain comprises a complete Fc domain (i.e., a hinge domain, a CH2 domain, and a CH3 domain). In certain embodiments, an Fc domain comprises a hinge domain (or portion thereof) fused to a CH3 domain (or portion thereof). In certain embodiments, an Fc domain comprises a CH2 domain (or portion thereof) fused to a CH3 domain (or portion thereof). In certain embodiments, an Fc domain consists of a CH3 domain or portion thereof. In certain embodiments, an Fc domain consists of a hinge domain (or portion thereof) and a CH3 domain (or portion thereof). In certain embodiments, an Fc domain consists of a CH2 domain (or portion thereof) and a CH3 domain. In certain embodiments, an Fc domain consists of a hinge domain (or portion thereof) and a CH2 domain (or portion thereof). In certain embodiments, an Fc domain lacks at least a portion of a CH2 domain (e.g., all or part of a CH2 domain). An Fc domain herein generally refers to a polypeptide comprising all or part of the Fc domain of an immunoglobulin heavy-chain. This includes, but is not limited to, polypeptides comprising the entire CHI, hinge, CH2, and / or CH3 domains as well as fragments of such peptides comprising only, e.g., the hinge, CH2, and CH3 domain. The Fc domain may be derived from an immunoglobulin of any species and / or any subtype, including, but not limited to, a human IgGl, IgG2, IgG3, IgG4, IgD, IgA, IgE, or IgM antibody. The Fc domain encompasses native Fc and Fc variant molecules. As set forth herein, it will be understood by one of ordinary skill in the art that any Fc domain may be modified such that it varies in amino acid sequence from the native Fc domain of a naturally occurring immunoglobulin molecule. In certain embodiments, the Fc domain has reduced effector function (e.g., FcyR binding).
[0374] The Fc domains of a polypeptide described herein may be derived from different immunoglobulin molecules. For example, an Fc domain of a polypeptide may comprise a CH2 and / or CH3 domain derived from an IgGl molecule and a hinge region derived from an IgG3 molecule. In another example, an Fc domain can comprise a chimeric hinge region derived, in part, from an IgGl molecule and, in part, from an IgG3 molecule. In another example, an Fc domain can comprise a chimeric hinge derived, in part, from an IgGl molecule and, in part, from an IgG4 molecule. In certain embodiments, an extended-PK group includes an Fc domain or fragments thereof or variants of the Fc domain or fragments thereof (all of which for the purpose of the present disclosure are comprised by the term "Fc domain"). The Fc domain does not contain a variable region that binds to antigen. Fc domains suitable for use in the present disclosure may be obtained from a number of different sources. In certain embodiments, an Fc domain is derived from a human immunoglobulin. In certain embodiments, the Fc domain is from a human IgGl constant region. It is understood, however, that the Fc domain may be derived from an immunoglobulin of another mammalian species, including for example, a rodent (e.g. a mouse, rat, rabbit, guinea pig) or non-human primate (e.g. chimpanzee, macaque) species.
[0375] Moreover, the Fc domain (or a fragment or variant thereof) may be derived from any immunoglobulin class, including IgM, IgG, IgD, IgA, and IgE, and any immunoglobulin isotype, including IgGl, IgG2, IgG3, and IgG4.
[0376] A variety of Fc domain gene sequences (e.g., mouse and human constant region gene sequences) are available in the form of publicly accessible deposits. Constant region domains comprising an Fc domain sequence can be selected lacking a particular effector function and / or with a particular modification to reduce immunogenicity. Many sequences of antibodies and antibody-encoding genes have been published and suitable Fc domain sequences (e.g. hinge, CH2, and / or CH3 sequences, or fragments or variants thereof) can be derived from these sequences using art recognized techniques.
[0377] In certain embodiments, the extended-PK group is a serum albumin binding protein such as those described in US2005 / 0287153, US2007 / 0003549, US2007 / 0178082, US2007 / 0269422, US2010 / 0113339, W02009 / 083804, and W02009 / 133208, which are herein incorporated by reference in their entirety. In certain embodiments, the extended- PK group is transferrin, as disclosed in US 7,176,278 and US 8,158,579, which are herein incorporated by reference in their entirety. In certain embodiments, the extended-PK group is a serum immunoglobulin binding protein such as those disclosed in US2007 / 0178082, US2014 / 0220017, and US2017 / 0145062, which are herein incorporated by reference in their entirety. In certain embodiments, the extended-PK group is a fibronectin (Fn)-based scaffold domain protein that binds to serum albumin, such as those disclosed in US2012 / 0094909, which is herein incorporated by reference in its entirety. Methods of making fibronectin-based scaffold domain proteins are also disclosed in US2012 / 0094909. A non-limiting example of a Fn3-based extended-PK group is Fn3(HSA), i.e., a Fn3 protein that binds to human serum albumin.
[0378] In certain aspects, the extended-PK polypeptide, suitable for use according to the disclosure, can employ one or more peptide linkers. As used herein, the term "peptide linker" refers to a peptide or polypeptide sequence which connects two or more domains (e.g., the extended-PK moiety and a polypeptide moiety, e.g., an Mtb antigen, immunogenic variant thereof, or immunogenic fragment of the Mtb antigen or the immunogenic variant thereof) in a linear amino acid sequence of a polypeptide chain. For example, peptide linkers may be used to connect an Mtb antigen, immunogenic variant thereof, or immunogenic fragment of the Mtb antigen or the immunogenic variant thereof to a HSA domain.
[0379] Linkers suitable for fusing the extended-PK group to, e.g., an Mtb antigen, immunogenic variant thereof, or immunogenic fragment of the Mtb antigen or the immunogenic variant thereof are well known in the art and described herein above.
[0380] In the following, embodiments of vaccine RNAs are described, wherein certain terms used when describing elements thereof have the following meanings: cap: 5'-cap structure, e.g., selected from the group consisting of m27'2 OG(5')ppSp(5')G (in particular its DI diastereomer), m27'3 OG(5')ppp(5')G, and m27'3 OGppp(mi2' °)ApG. hAg-Kozak: 5'-UTR sequence of the human alpha-globin mRNA with an optimized 'Kozak sequence' to increase translational efficiency. sec / TMD: Fusion-protein tags derived from the sequence encoding the human MHC class I complex (HLA-B51, haplotype A2, B27 / B51, Cw2 / Cw3), which have been shown to improve antigen processing and presentation. Sec corresponds to the 78 bp fragment coding for the secretory signal peptide, which guides translocation of the nascent polypeptide chain into the endoplasmatic reticulum. The TMD comprises an amino acid sequence enhancing antigen processing and / or presentation. In some embodiments, the TMD comprises a trafficking signal. In some embodiments, the TMD is MITD. MITD corresponds to the transmembrane and cytoplasmic domain of the MHC class I molecule, also called MHC class I trafficking domain.
[0381] Antigen: Sequences encoding the respective vaccine antigen(s) / epitope(s), i.e., one or more Mtb antigens, immunogenic variants thereof, or immunogenic fragments of the Mtb antigens or the immunogenic variants thereof.
[0382] Glycine-serine linker (GS): Sequences coding for short peptide linkers predominantly consisting of the amino acids glycine (G) and serine (S), as commonly used for fusion proteins.
[0383] FI element: The 3'-UTR is a combination of two sequence elements derived from the "amino terminal enhancer of split" (AES) mRNA (called F) and the mitochondrial encoded 12S ribosomal RNA (called I). These were identified by an ex vivo selection process for sequences that confer RNA stability and augment total protein expression.
[0384] A30L70: A poly(A)-tail measuring 110 nucleotides in length, consisting of a stretch of 30 adenosine residues, followed by a 10 nucleotide linker sequence and another 70 adenosine residues designed to enhance RNA stability and translational efficiency in dendritic cells.
[0385] In some embodiments, vaccine RNA described herein has one of the following structures: cap-hAg-Kozak-Antigen(s)-FI-A30L70 cap-hAg-Kozak-sec-Antigen(s)-FI-A30L70 cap-hAg-Kozak-sec-Antigen(s)-TMD-FI-A30L70
[0386] In some embodiments, vaccine antigen described herein has the structure: sec-Antigen sec-Antigen-TMD
[0387] In some embodiments, hAg-Kozak comprises the nucleotide sequence of SEQ ID NO: 53. In some embodiments, sec of the encoded vaccine antigen / epitope comprises an amino acid sequence selected from SEQ ID NOs: 17 to 37 or 65 or encoded by the nucleotide sequence selected from SEQ ID NOs: 38 to 51. In some embodiments, the TMD of the encoded vaccine antigen / epitope is MITD comprising the amino acid sequence of SEQ ID NO: 52. In some embodiments, the TMD of the encoded vaccine antigen / epitope is a viral signal peptide, preferably a HSV-1 glycoprotein D signal peptide. In some embodiments, FI comprises the nucleotide sequence of SEQ ID NO: 54. In some embodiments, A30L70 comprises the nucleotide sequence of SEQ ID NO: 55.
[0388] In some embodiments, the different elements (sec, Antigen, TMD) may be linked by one or more GS linkers. In some embodiments, a GS linker of the encoded vaccine antigen / epitope comprises the amino acid sequence of SEQ ID NO: 56. In some embodiments, the sequence encoding the vaccine antigen / epitope comprises a modified nucleoside replacing (partially or completely, preferably completely) uridine, wherein the modified nucleoside is selected from the group consisting of pseudouridine (i ), Nl-methyl-pseudouridine (mli ), and 5-methyl-uridine.
[0389] In some embodiments, the sequence encoding the vaccine antigen / epitope is codon-optimized.
[0390] In some embodiments, the G / C content of the sequence encoding the vaccine antigen / epitope is increased compared to the wild type coding sequence.
[0391] In some embodiments, the RNA (in particular, mRNA) described herein comprises: a 5' UTR comprising the nucleotide sequence of SEQ ID NO: 53, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 53; a 3' UTR comprising the nucleotide sequence of SEQ ID NO: 54, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 54; and a poly-A sequence comprising the nucleotide sequence of SEQ ID NO: 55.
[0392] In some embodiments, the RNA (in particular, mRNA) described herein comprises: m27'3' °Gppp(mi2'0) ApG as capping structure at the 5'-end of the mRNA; a 5' UTR comprising the nucleotide sequence of SEQ ID NO: 53, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 53; a 3' UTR comprising the nucleotide sequence of SEQ ID NO: 54, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 54; and a poly-A sequence comprising the nucleotide sequence of SEQ ID NO: 55.
[0393] In some embodiments, the RNA is unmodified. In some embodiments, the RNA is modified. In some embodiments, the RNA comprises Nl-methyl-pseudouridine (mli ) in place of at least one uridine (e.g., in place of each uridine).
[0394] In some embodiments, the RNA (in particular, mRNA) described herein comprises: m27'3' °Gppp(mi2'0) ApG as capping structure at the 5'-end of the mRNA; a 5' UTR comprising the nucleotide sequence of SEQ ID NO: 53, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 53; a 3' UTR comprising the nucleotide sequence of SEQ ID NO: 54, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 54; and a poly-A sequence comprising the nucleotide sequence of SEQ ID NO: 55; and
[0395] Nl-methyl-pseudouridine (mli ) in place of at least one uridine (e.g., in place of each uridine).
[0396] In some embodiments, a vaccine antigen or epitope described herein is derived from Mycobacterium tuberculosis.
[0397] In some embodiments, a vaccine antigen or epitope described herein is derived from a Mycobacterium tuberculosis protein, an immunogenic variant thereof, or an immunogenic fragment of the Mycobacterium tuberculosis protein or the immunogenic variant thereof. Thus, in some embodiments, the RNA, e.g., mRNA, used in the present disclosure encodes an amino acid sequence comprising an Mtb protein, an immunogenic variant thereof, or an immunogenic fragment of the Mtb protein or the immunogenic variant thereof.
[0398] In some embodiments, a vaccine antigen or epitope described herein is derived from an Mtb protein from the acute phase of the Mtb life cycle, an immunogenic variant thereof, or an immunogenic fragment of the Mtb protein from the acute phase of the Mtb life cycle or the immunogenic variant thereof.
[0399] In some embodiments, a vaccine antigen or epitope described herein is derived from an Mtb protein from the latent phase of the Mtb life cycle, an immunogenic variant thereof, or an immunogenic fragment of the Mtb protein from the latent phase of the Mtb life cycle or the immunogenic variant thereof.
[0400] In some embodiments, a vaccine antigen or epitope described herein is derived from an Mtb protein from the resuscitation phase of the Mtb life cycle, an immunogenic variant thereof, or an immunogenic fragment of the Mtb protein from the resuscitation phase of the Mtb life cycle or the immunogenic variant thereof.
[0401] In some embodiments, RNA (in particular, mRNA) described herein (e.g., contained in the compositions / formulations of the present disclosure and / or used in the methods of the present disclosure) may be presented as a product containing the vaccine RNA as active substance and other ingredients comprising: ALC-0315 ((4- hydroxybutyl)azanediyl)bis(hexane-6,l-diyl)bis(2-hexyldecanoate), ALC-0159 (2-[(polyethylene glycol)-2000]-N,N- ditetradecylacetamide), l,2-Distearoyl-sn-glycero-3-phosphocholine (DSPC), and cholesterol.
[0402] In some embodiments, the RNA (in particular, mRNA) described herein is formulated or is to be formulated as a liquid, a solid, or a combination thereof.
[0403] In some embodiments, the RNA (in particular, mRNA) described herein is formulated or is to be formulated for injection.
[0404] In some embodiments, the RNA (in particular, mRNA) described herein is formulated or is to be formulated for intramuscular administration.
[0405] In some embodiments, the RNA (in particular, mRNA) described herein is formulated or is to be formulated as a composition, e.g., a pharmaceutical composition.
[0406] In some embodiments, the composition comprises a cationically ionizable lipid.
[0407] In some embodiments, the composition comprises a cationically ionizable lipid and one or more additional lipids. In some embodiments, the one or more additional lipids are selected from polymer-conjugated lipids, neutral lipids, and combinations thereof. In some embodiments, the neutral lipids include phospholipids, steroid lipids, and combinations thereof. In some embodiments, the one or more additional lipids are a combination of a polymer-conjugated lipid, a phospholipid, and a steroid lipid.
[0408] In some embodiments, the composition comprises a cationically ionizable lipid; a polymer-conjugated lipid which is a PEG-conjugated lipid; cholesterol; and a phospholipid. In some embodiments, the phospholipid is DSPC. In some embodiments, the phospholipid is DOPE.
[0409] In some embodiments, the composition comprises a cationically ionizable lipid; a polymer-conjugated lipid which is 2- [(polyethylene glycol)-2000]-N,N-ditetradecylacetamide; cholesterol; and a phospholipid. In some embodiments, the phospholipid is DSPC. In some embodiments, the phospholipid is DOPE.
[0410] In some embodiments, the composition comprises a cationically ionizable lipid which is ((4- hydroxybutyl)azanediyl)bis(hexane-6,l-diyl)bis(2-hexyldecanoate); a polymer-conjugated lipid which is 2- [(polyethylene glycol)-2000]-N,N-ditetradecylacetamide; cholesterol; and a phospholipid. In some embodiments, the phospholipid is DSPC. In some embodiments, the phospholipid is DOPE. In some embodiments, at least a portion of (i) the RNA, (ii) the cationically ionizable lipid, and if present, (iii) the one or more additional lipids is present in particles. In some embodiments, the particles are nanoparticles, such as lipid nanoparticles (LNPs).
[0411] In some embodiments, the composition, in particular the pharmaceutical composition, is a vaccine.
[0412] In some embodiments, the composition, in particular the pharmaceutical composition, further comprises one or more pharmaceutically acceptable carriers, diluents and / or excipients.
[0413] In some embodiments, the RNA and / or the composition, in particular the pharmaceutical composition, is / are a component of a kit.
[0414] In some embodiments, the kit further comprises instructions for use of the RNA for inducing an immune response against Mycobacterium tuberculosis in a subject.
[0415] In some embodiments, the kit further comprises instructions for use of the RNA for therapeutically or prophylactically treating a Mycobacterium tuberculosis infection in a subject.
[0416] In some embodiments, the subject is a human.
[0417] In some embodiments, the RNA (in particular, mRNA), e.g., RNA encoding vaccine antigen, described in the present disclosure is non-immunogenic. RNA encoding an immunostimulant may be administered according to the present disclosure to provide an adjuvant effect. The RNA encoding an immunostimulant may be standard RNA or non- immunogenic RNA.
[0418] Embodiments of Mycobacterium tuberculosis fMtbj antigens
[0419] The present disclosure describes Mtb antigens, immunogenic variants thereof, and immunogenic fragments of the Mtb antigens or the immunogenic variants thereof (referred to as "Mtb antigens" herein) and RNA encoding these antigens. Mtb antigens primarily described herein include Wbbll, PPE18 and PE13.
[0420] Wbbll
[0421] In some embodiments, the Mtb antigen Wbbll comprises the amino acid sequence according to SEQ ID NO: 1. A full- length antigen representing the antigen Wbbll is characterized in that it comprises the full-length amino acid sequence according to SEQ ID NO: 1, whereas an antigen fragment representing the antigen Wbbll is characterized in that it comprises an amino acid sequence which is only a part of SEQ ID NO: 1 but which still is able to induce an immune reaction to Wbbll, when delivered to a subject.
[0422] An immunogenic variant of the Mtb antigen Wbbll comprises an amino acid sequence which is "immunologically equivalent" to Wbbll and thus, is able to induce an immune reaction to Wbbll, when delivered to a subject. In some embodiments, an immunogenic variant of the Mtb antigen Wbbll comprises an amino acid sequence differing from SEQ ID NO: 1 by one or more deletions, insertions or substitutions in the amino acid sequence while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 1 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). A full-length antigen representing an immunogenic variant of the antigen Wbbll is characterized in that it comprises the full-length amino acid sequence of the immunogenic variant, whereas an antigen fragment representing an immunogenic variant of the antigen Wbbll is characterized in that it comprises an amino acid sequence which is only a part of the full-length amino acid sequence of the immunogenic variant but which still is able to induce an immune reaction to Wbbll, when delivered to a subject.
[0423] In some embodiments, the Mtb antigen Wbbll is encoded by a nucleotide sequence according to SEQ ID NO: 147.
[0424] In some embodiments, an immunogenic variant of the Mtb antigen Wbbll is encoded by a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 147 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0425] PPE18
[0426] In some embodiments, the Mtb antigen PPE18 comprises the amino acid sequence according to SEQ ID NO: 2. A full- length antigen representing the antigen PPE18 is characterized in that it comprises the full-length amino acid sequence according to SEQ ID NO: 2, whereas an antigen fragment representing the antigen PPE18 is characterized in that it comprises an amino acid sequence which is only a part of SEQ ID NO: 2 but which still is able to induce an immune reaction to PPE18, when delivered to a subject.
[0427] An immunogenic variant of the Mtb antigen PPE18 comprises an amino acid sequence which is "immunologically equivalent" to PPE18 and thus, is able to induce an immune reaction to PPE18, when delivered to a subject. In some embodiments, an immunogenic variant of the Mtb antigen PPE18 comprises an amino acid sequence differing from SEQ ID NO: 2 by one or more deletions, insertions or substitutions in the amino acid sequence while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 2 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). A full-length antigen representing an immunogenic variant of the antigen PPE18 is characterized in that it comprises the full-length amino acid sequence of the immunogenic variant, whereas an antigen fragment representing an immunogenic variant of the antigen PPE18 is characterized in that it comprises an amino acid sequence which is only a part of the full-length amino acid sequence of the immunogenic variant but which still is able to induce an immune reaction to PPE18, when delivered to a subject.
[0428] In some embodiments, the Mtb antigen PPE18 is encoded by a nucleotide sequence according to SEQ ID NO: 148.
[0429] In some embodiments, an immunogenic variant of the Mtb antigen PPE18 is encoded by a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 148 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0430] PE13
[0431] In some embodiments, the Mtb antigen PE13 comprises the amino acid sequence according to SEQ ID NO: 3. In other embodiments, the Mtb antigen PE13 comprises the amino acid sequence according to SEQ ID NO: 4. A full-length antigen representing the antigen PE13 is characterized in that it comprises the full-length amino acid sequence according to SEQ ID NO: 3 or 4, whereas an antigen fragment representing the antigen PE13 is characterized in that it comprises an amino acid sequence which is only a part of SEQ ID NO: 3 or 4 but which still is able to induce an immune reaction to PE13, when delivered to a subject.
[0432] An immunogenic variant of the Mtb antigen PE13 comprises an amino acid sequence which is "immunologically equivalent" to PE13 and thus, is able to induce an immune reaction to PE13, when delivered to a subject. In some embodiments, an immunogenic variant of the Mtb antigen PE13 comprises an amino acid sequence differing from SEQ ID NO: 3 or 4 by one or more deletions, insertions or substitutions in the amino acid sequence while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 3 or 4 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). A full-length antigen representing an immunogenic variant of the antigen PE13 is characterized in that it comprises the full-length amino acid sequence of the immunogenic variant, whereas an antigen fragment representing an immunogenic variant of the antigen PE13 is characterized in that it comprises an amino acid sequence which is only a part of the full-length amino acid sequence of the immunogenic variant but which still is able to induce an immune reaction to PE13, when delivered to a subject.
[0433] In some embodiments, the Mtb antigen PE13 is encoded by a nucleotide sequence according to SEQ ID NO: 149 or SEQ ID NO: 150.
[0434] In some embodiments, an immunogenic variant of the Mtb antigen PE13 is encoded by a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 149 or SEQ ID NO: 150 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0435] EsxA
[0436] In some embodiments, the Mtb antigen EsxA comprises the amino acid sequence according to SEQ ID NO: 5. A full- length antigen representing the antigen EsxA is characterized in that it comprises the full-length amino acid sequence according to SEQ ID NO: 5, whereas an antigen fragment representing the antigen EsxA is characterized in that it comprises an amino acid sequence which is only a part of SEQ ID NO: 5 but which still is able to induce an immune reaction to EsxA, when delivered to a subject.
[0437] An immunogenic variant of the Mtb antigen EsxA comprises an amino acid sequence which is "immunologically equivalent" to EsxA and thus, is able to induce an immune reaction to EsxA, when delivered to a subject. In some embodiments, an immunogenic variant of the Mtb antigen EsxA comprises an amino acid sequence differing from SEQ ID NO: 5 by one or more deletions, insertions or substitutions in the amino acid sequence while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 5 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). A full-length antigen representing an immunogenic variant of the antigen EsxA is characterized in that it comprises the full-length amino acid sequence of the immunogenic variant, whereas an antigen fragment representing an immunogenic variant of the antigen EsxA is characterized in that it comprises an amino acid sequence which is only a part of the full-length amino acid sequence of the immunogenic variant but which still is able to induce an immune reaction to EsxA, when delivered to a subject. EsxA is also called ESAT-6 and both terms are used synonymously herein.
[0438] In some embodiments, the Mtb antigen EsxA is encoded by a nucleotide sequence according to SEQ ID NO: 151.
[0439] In some embodiments, an immunogenic variant of the Mtb antigen EsxA is encoded by a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 151 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0440] EsxB
[0441] In some embodiments, the Mtb antigen EsxB comprises the amino acid sequence according to SEQ ID NO: 6. A full- length antigen representing the antigen EsxB is characterized in that it comprises the full-length amino acid sequence according to SEQ ID NO: 6, whereas an antigen fragment representing the antigen EsxB is characterized in that it comprises an amino acid sequence which is only a part of SEQ ID NO: 6 but which still is able to induce an immune reaction to EsxB, when delivered to a subject. An immunogenic variant of the Mtb antigen EsxB comprises an amino acid sequence which is "immunologically equivalent" to EsxB and thus, is able to induce an immune reaction to EsxB, when delivered to a subject. In some embodiments, an immunogenic variant of the Mtb antigen EsxB comprises an amino acid sequence differing from SEQ ID NO: 6 by one or more deletions, insertions or substitutions in the amino acid sequence while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 6 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). A full-length antigen representing an immunogenic variant of the antigen EsxB is characterized in that it comprises the full-length amino acid sequence of the immunogenic variant, whereas an antigen fragment representing an immunogenic variant of the antigen EsxB is characterized in that it comprises an amino acid sequence which is only a part of the full-length amino acid sequence of the immunogenic variant but which still is able to induce an immune reaction to EsxB, when delivered to a subject. EsxB is also called CFP-10 and both terms are used synonymously herein.
[0442] In some embodiments, the Mtb antigen EsxB is encoded by a nucleotide sequence according to SEQ ID NO: 152.
[0443] In some embodiments, an immunogenic variant of the Mtb antigen EsxB is encoded by a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 152 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0444] EsxG
[0445] In some embodiments, the Mtb antigen EsxG comprises the amino acid sequence according to SEQ ID NO: 7. A full- length antigen representing the antigen EsxG is characterized in that it comprises the full-length amino acid sequence according to SEQ ID NO: 7, whereas an antigen fragment representing the antigen EsxG is characterized in that it comprises an amino acid sequence which is only a part of SEQ ID NO: 7 but which still is able to induce an immune reaction to EsxG, when delivered to a subject.
[0446] An immunogenic variant of the Mtb antigen EsxG comprises an amino acid sequence which is "immunologically equivalent" to EsxG and thus, is able to induce an immune reaction to EsxG, when delivered to a subject. In some embodiments, an immunogenic variant of the Mtb antigen EsxG comprises an amino acid sequence differing from SEQ ID NO: 7 by one or more deletions, insertions or substitutions in the amino acid sequence while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 7 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). A full-length antigen representing an immunogenic variant of the antigen EsxG is characterized in that it comprises the full-length amino acid sequence of the immunogenic variant, whereas an antigen fragment representing an immunogenic variant of the antigen EsxG is characterized in that it comprises an amino acid sequence which is only a part of the full-length amino acid sequence of the immunogenic variant but which still is able to induce an immune reaction to EsxG, when delivered to a subject.
[0447] In some embodiments, the Mtb antigen EsxG is encoded by a nucleotide sequence according to SEQ ID NO: 153.
[0448] In some embodiments, an immunogenic variant of the Mtb antigen EsxG is encoded by a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 153 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0449] EsxH
[0450] In some embodiments, the Mtb antigen EsxH comprises the amino acid sequence according to SEQ ID NO: 8. A full- length antigen representing the antigen EsxH is characterized in that it comprises the full-length amino acid sequence according to SEQ ID NO: 8, whereas an antigen fragment representing the antigen EsxH is characterized in that it comprises an amino acid sequence which is only a part of SEQ ID NO: 8 but which still is able to induce an immune reaction to EsxH, when delivered to a subject.
[0451] An immunogenic variant of the Mtb antigen EsxH comprises an amino acid sequence which is "immunologically equivalent" to EsxH and thus, is able to induce an immune reaction to EsxH, when delivered to a subject. In some embodiments, an immunogenic variant of the Mtb antigen EsxH comprises an amino acid sequence differing from SEQ ID NO: 8 by one or more deletions, insertions or substitutions in the amino acid sequence while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 8 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). A full-length antigen representing an immunogenic variant of the antigen EsxH is characterized in that it comprises the full-length amino acid sequence of the immunogenic variant, whereas an antigen fragment representing an immunogenic variant of the antigen EsxH is characterized in that it comprises an amino acid sequence which is only a part of the full-length amino acid sequence of the immunogenic variant but which still is able to induce an immune reaction to EsxH, when delivered to a subject.
[0452] In some embodiments, the Mtb antigen EsxH is encoded by a nucleotide sequence according to SEQ ID NO: 154.
[0453] In some embodiments, an immunogenic variant of the Mtb antigen EsxH is encoded by a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 154 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0454] Esxl
[0455] In some embodiments, the Mtb antigen Esxl comprises the amino acid sequence according to SEQ ID NO: 9. A full- length antigen representing the antigen Esxl is characterized in that it comprises the full-length amino acid sequence according to SEQ ID NO: 9, whereas an antigen fragment representing the antigen Esxl is characterized in that it comprises an amino acid sequence which is only a part of SEQ ID NO: 9 but which still is able to induce an immune reaction to Esxl, when delivered to a subject.
[0456] An immunogenic variant of the Mtb antigen Esxl comprises an amino acid sequence which is "immunologically equivalent" to Esxl and thus, is able to induce an immune reaction to Esxl, when delivered to a subject. In some embodiments, an immunogenic variant of the Mtb antigen Esxl comprises an amino acid sequence differing from SEQ ID NO: 9 by one or more deletions, insertions or substitutions in the amino acid sequence while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 9 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). A full-length antigen representing an immunogenic variant of the antigen Esxl is characterized in that it comprises the full-length amino acid sequence of the immunogenic variant, whereas an antigen fragment representing an immunogenic variant of the antigen Esxl is characterized in that it comprises an amino acid sequence which is only a part of the full-length amino acid sequence of the immunogenic variant but which still is able to induce an immune reaction to Esxl, when delivered to a subject.
[0457] In some embodiments, the Mtb antigen Esxl is encoded by a nucleotide sequence according to SEQ ID NO: 155.
[0458] In some embodiments, an immunogenic variant of the Mtb antigen Esxl is encoded by a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 155 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN). EsxJ
[0459] In some embodiments, the Mtb antigen EsxJ comprises the amino acid sequence according to SEQ ID NO: 10. A full- length antigen representing the antigen EsxJ is characterized in that it comprises the full-length amino acid sequence according to SEQ ID NO: 10, whereas an antigen fragment representing the antigen EsxJ is characterized in that it comprises an amino acid sequence which is only a part of SEQ ID NO: 10 but which still is able to induce an immune reaction to EsxJ, when delivered to a subject.
[0460] An immunogenic variant of the Mtb antigen EsxJ comprises an amino acid sequence which is "immunologically equivalent" to EsxJ and thus, is able to induce an immune reaction to EsxJ, when delivered to a subject. In some embodiments, an immunogenic variant of the Mtb antigen EsxJ comprises an amino acid sequence differing from SEQ ID NO: 10 by one or more deletions, insertions or substitutions in the amino acid sequence while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 10 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). A full-length antigen representing an immunogenic variant of the antigen EsxJ is characterized in that it comprises the full-length amino acid sequence of the immunogenic variant, whereas an antigen fragment representing an immunogenic variant of the antigen EsxJ is characterized in that it comprises an amino acid sequence which is only a part of the full-length amino acid sequence of the immunogenic variant but which still is able to induce an immune reaction to EsxJ, when delivered to a subject.
[0461] In some embodiments, the Mtb antigen EsxJ is encoded by a nucleotide sequence according to SEQ ID NO: 156.
[0462] In some embodiments, an immunogenic variant of the Mtb antigen EsxJ is encoded by a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 156 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0463] EsxK
[0464] In some embodiments, the Mtb antigen EsxK comprises the amino acid sequence according to SEQ ID NO: 11. A full- length antigen representing the antigen EsxK is characterized in that it comprises the full-length amino acid sequence according to SEQ ID NO: 11, whereas an antigen fragment representing the antigen EsxK is characterized in that it comprises an amino acid sequence which is only a part of SEQ ID NO: 11 but which still is able to induce an immune reaction to EsxK, when delivered to a subject.
[0465] An immunogenic variant of the Mtb antigen EsxK comprises an amino acid sequence which is "immunologically equivalent" to EsxK and thus, is able to induce an immune reaction to EsxK, when delivered to a subject. In some embodiments, an immunogenic variant of the Mtb antigen EsxK comprises an amino acid sequence differing from SEQ ID NO: 11 by one or more deletions, insertions or substitutions in the amino acid sequence while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 11 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). A full-length antigen representing an immunogenic variant of the antigen EsxK is characterized in that it comprises the full-length amino acid sequence of the immunogenic variant, whereas an antigen fragment representing an immunogenic variant of the antigen EsxK is characterized in that it comprises an amino acid sequence which is only a part of the full-length amino acid sequence of the immunogenic variant but which still is able to induce an immune reaction to EsxK, when delivered to a subject.
[0466] In some embodiments, the Mtb antigen EsxK is encoded by a nucleotide sequence according to SEQ ID NO: 157. In some embodiments, an immunogenic variant of the Mtb antigen EsxK is encoded by a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 157 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0467] EsxL
[0468] In some embodiments, the Mtb antigen EsxL comprises the amino acid sequence according to SEQ ID NO: 12. A full- length antigen representing the antigen EsxL is characterized in that it comprises the full-length amino acid sequence according to SEQ ID NO: 12, whereas an antigen fragment representing the antigen EsxL is characterized in that it comprises an amino acid sequence which is only a part of SEQ ID NO: 12 but which still is able to induce an immune reaction to EsxL, when delivered to a subject.
[0469] An immunogenic variant of the Mtb antigen EsxL comprises an amino acid sequence which is "immunologically equivalent" to EsxL and thus, is able to induce an immune reaction to EsxL, when delivered to a subject. In some embodiments, an immunogenic variant of the Mtb antigen EsxL comprises an amino acid sequence differing from SEQ ID NO: 12 by one or more deletions, insertions or substitutions in the amino acid sequence while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 12 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). A full-length antigen representing an immunogenic variant of the antigen EsxL is characterized in that it comprises the full-length amino acid sequence of the immunogenic variant, whereas an antigen fragment representing an immunogenic variant of the antigen EsxL is characterized in that it comprises an amino acid sequence which is only a part of the full-length amino acid sequence of the immunogenic variant but which still is able to induce an immune reaction to EsxL, when delivered to a subject.
[0470] In some embodiments, the Mtb antigen EsxL is encoded by a nucleotide sequence according to SEQ ID NO: 158.
[0471] In some embodiments, an immunogenic variant of the Mtb antigen EsxL is encoded by a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 158 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0472] EsxM
[0473] In some embodiments, the Mtb antigen EsxM comprises the amino acid sequence according to SEQ ID NO: 13. A full- length antigen representing the antigen EsxM is characterized in that it comprises the full-length amino acid sequence according to SEQ ID NO: 13, whereas an antigen fragment representing the antigen EsxM is characterized in that it comprises an amino acid sequence which is only a part of SEQ ID NO: 13 but which still is able to induce an immune reaction to EsxM, when delivered to a subject.
[0474] An immunogenic variant of the Mtb antigen EsxM comprises an amino acid sequence which is "immunologically equivalent" to EsxM and thus, is able to induce an immune reaction to EsxM, when delivered to a subject. In some embodiments, an immunogenic variant of the Mtb antigen EsxM comprises an amino acid sequence differing from SEQ ID NO: 13 by one or more deletions, insertions or substitutions in the amino acid sequence while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 13 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). A full-length antigen representing an immunogenic variant of the antigen EsxM is characterized in that it comprises the full-length amino acid sequence of the immunogenic variant, whereas an antigen fragment representing an immunogenic variant of the antigen EsxM is characterized in that it comprises an amino acid sequence which is only a part of the full-length amino acid sequence of the immunogenic variant but which still is able to induce an immune reaction to EsxM, when delivered to a subject.
[0475] In some embodiments, the Mtb antigen EsxM is encoded by a nucleotide sequence according to SEQ ID NO: 159.
[0476] In some embodiments, an immunogenic variant of the Mtb antigen EsxM is encoded by a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 159 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0477] EsxN
[0478] In some embodiments, the Mtb antigen EsxN comprises the amino acid sequence according to SEQ ID NO: 14. A full- length antigen representing the antigen EsxN is characterized in that it comprises the full-length amino acid sequence according to SEQ ID NO: 14, whereas an antigen fragment representing the antigen EsxN is characterized in that it comprises an amino acid sequence which is only a part of SEQ ID NO: 14 but which still is able to induce an immune reaction to EsxN, when delivered to a subject.
[0479] An immunogenic variant of the Mtb antigen EsxN comprises an amino acid sequence which is "immunologically equivalent" to EsxN and thus, is able to induce an immune reaction to EsxN, when delivered to a subject. In some embodiments, an immunogenic variant of the Mtb antigen EsxN comprises an amino acid sequence differing from SEQ ID NO: 14 by one or more deletions, insertions or substitutions in the amino acid sequence while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 14 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). A full-length antigen representing an immunogenic variant of the antigen EsxN is characterized in that it comprises the full-length amino acid sequence of the immunogenic variant, whereas an antigen fragment representing an immunogenic variant of the antigen EsxN is characterized in that it comprises an amino acid sequence which is only a part of the full-length amino acid sequence of the immunogenic variant but which still is able to induce an immune reaction to EsxN, when delivered to a subject.
[0480] In some embodiments, the Mtb antigen EsxN is encoded by a nucleotide sequence according to SEQ ID NO: 160.
[0481] In some embodiments, an immunogenic variant of the Mtb antigen EsxN is encoded by a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 160 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0482] EsxV
[0483] In some embodiments, the Mtb antigen EsxV comprises the amino acid sequence according to SEQ ID NO: 15. A full- length antigen representing the antigen EsxV is characterized in that it comprises the full-length amino acid sequence according to SEQ ID NO: 15, whereas an antigen fragment representing the antigen EsxV is characterized in that it comprises an amino acid sequence which is only a part of SEQ ID NO: 15 but which still is able to induce an immune reaction to EsxV, when delivered to a subject.
[0484] An immunogenic variant of the Mtb antigen EsxV comprises an amino acid sequence which is "immunologically equivalent" to EsxV and thus, is able to induce an immune reaction to EsxV, when delivered to a subject. In some embodiments, an immunogenic variant of the Mtb antigen EsxV comprises an amino acid sequence differing from SEQ ID NO: 15 by one or more deletions, insertions or substitutions in the amino acid sequence while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 15 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). A full-length antigen representing an immunogenic variant of the antigen EsxV is characterized in that it comprises the full-length amino acid sequence of the immunogenic variant, whereas an antigen fragment representing an immunogenic variant of the antigen EsxV is characterized in that it comprises an amino acid sequence which is only a part of the full-length amino acid sequence of the immunogenic variant but which still is able to induce an immune reaction to EsxV, when delivered to a subject.
[0485] In some embodiments, the Mtb antigen EsxV is encoded by a nucleotide sequence according to SEQ ID NO: 161.
[0486] In some embodiments, an immunogenic variant of the Mtb antigen EsxV is encoded by a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 161 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0487] EsxW
[0488] In some embodiments, the Mtb antigen EsxW comprises the amino acid sequence according to SEQ ID NO: 16. A full- length antigen representing the antigen EsxW is characterized in that it comprises the full-length amino acid sequence according to SEQ ID NO: 16, whereas an antigen fragment representing the antigen EsxW is characterized in that it comprises an amino acid sequence which is only a part of SEQ ID NO: 16 but which still is able to induce an immune reaction to EsxW, when delivered to a subject.
[0489] An immunogenic variant of the Mtb antigen EsxW comprises an amino acid sequence which is "immunologically equivalent" to EsxW and thus, is able to induce an immune reaction to EsxW, when delivered to a subject. In some embodiments, an immunogenic variant of the Mtb antigen EsxW comprises an amino acid sequence differing from SEQ ID NO: 16 by one or more deletions, insertions or substitutions in the amino acid sequence while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 16 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). A full-length antigen representing an immunogenic variant of the antigen EsxW is characterized in that it comprises the full-length amino acid sequence of the immunogenic variant, whereas an antigen fragment representing an immunogenic variant of the antigen EsxW is characterized in that it comprises an amino acid sequence which is only a part of the full-length amino acid sequence of the immunogenic variant but which still is able to induce an immune reaction to EsxW, when delivered to a subject.
[0490] In some embodiments, the Mtb antigen EsxW is encoded by a nucleotide sequence according to SEQ ID NO: 162.
[0491] In some embodiments, an immunogenic variant of the Mtb antigen EsxW is encoded by a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 162 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0492] In some embodiments, amino acid sequences described herein comprise a secretory signal peptide as described herein above.
[0493] Strings
[0494] In some embodiments, the chimeric protein (also referred to as "string" herein) comprises one or more Mtb antigens and antigen fragments described above.
[0495] In some embodiments, the chimeric protein (named "string 1") comprises Wbbll, PPE18 and PE13. In some embodiments, the chimeric protein comprises one or more antigenic fragments of PPE18 and PE13 and WbbLl as full- length antigen. In some embodiments, the chimeric protein comprises two antigen fragments of PE13, three antigen fragments of PPE18, and a full-length antigen of WbbLl. In some embodiments, the chimeric protein comprises an amino acid sequence with the sequence of SEQ ID NO: 57 or an amino acid sequence that differs from SEQ ID NO: 57 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 57 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). In some embodiments, the chimeric protein as shown in SEQ ID NO: 57 is modified with respect to the signal peptide, transmembrane domains and / or linkers. In some embodiments, the chimeric protein is encoded by a nucleic acid molecule comprising a nucleotide sequence with the sequence of SEQ ID NO: 163 or 164 or a nucleotide that differs from SEQ ID NO: 163 or 164 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 163 or 164 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0496] In some embodiments, the chimeric protein (named "string 2") comprises EsxG, EsxH, Esxl, EsxJ, EsxN and EsxW. In some embodiments, the chimeric protein comprises an antigen fragment of EsxJ and full length antigens of EsxG, EsxH, Esxl, EsxN and EsxW. In some embodiments, the chimeric protein comprises an amino acid sequence with the sequence of SEQ ID NO: 59 or an amino acid sequence that differs from SEQ ID NO: 59 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 59 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). In some embodiments, the chimeric protein as shown in SEQ ID NO: 59 is modified with respect to the signal peptide, transmembrane domains and / or linkers. In some embodiments, the chimeric protein is encoded by a nucleic acid molecule comprising a nucleotide sequence with the sequence of SEQ ID NO: 165 or 166 or a nucleotide that differs from SEQ ID NO: 165 or 166 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 165 or 166 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0497] In some embodiments, the chimeric protein (named "string 3") comprises EsxB, EsxG, EsxH, Esxl, EsxJ, EsxN and EsxW. In some embodiments, the chimeric protein comprises an antigen fragment of EsxB, an antigen fragment of EsxJ and full length antigens of EsxG, EsxH, Esxl, EsxN and EsxW. In some embodiments, the chimeric protein comprises an amino acid sequence with the sequence of SEQ ID NO: 58 or an amino acid sequence that differs from SEQ ID NO: 58 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 58 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). In some embodiments, the chimeric protein as shown in SEQ ID NO: 58 is modified with respect to the signal peptide, transmembrane domains and / or linkers. In some embodiments, the chimeric protein is encoded by a nucleic acid molecule comprising a nucleotide sequence with the sequence of SEQ ID NO: 167 or 168 or a nucleotide that differs from SEQ ID NO: 167 or 168 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 167 or 168 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0498] In some embodiments, the chimeric protein (named "string 4") comprises PE13, PPE18, Esxl, EsxJ, EsxN and EsxW. In some embodiments, the chimeric protein comprises one or more antigenic fragments of PPE18, PE13, EsxJ and EsxN and full length antigens of Esxl and EsxW. In some embodiments, the chimeric protein comprises two antigen fragments of PE13, three antigen fragments of PPE18, one antigen fragment of EsxJ, one antigen fragment of EsxN and full length antigens of Esxl and EsxW. In some embodiments, the chimeric protein comprises an amino acid sequence with the sequence of SEQ ID NO: 60 or 61 or an amino acid sequence that differs from SEQ ID NO: 60 or 61 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 60 or 61 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). In some embodiments, the chimeric protein as shown in SEQ ID NO: 60 or 61 is modified with respect to the signal peptide, transmembrane domains and / or linkers. In some embodiments, the chimeric protein is encoded by a nucleic acid molecule comprising a nucleotide sequence with the sequence of any one of SEQ ID NOs: 169 to 172 or a nucleotide that differs from SEQ ID NOs: 169 to 172 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to any one of SEQ ID NOs: 169 to 172 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0499] In some embodiments, the chimeric protein (named "string 5") comprises PE13, PPE18, EsxG, EsxJ, and EsxW. In some embodiments, the chimeric protein comprises one or more antigenic fragments of PPE18, PE13 and EsxJ and full length antigens of EsxG and EsxW. In some embodiments, the chimeric protein comprises two antigen fragments of PE13, three antigen fragments of PPE18, one antigen fragment of EsxJ and full length antigens of EsxG and EsxW. In some embodiments, the chimeric protein comprises an amino acid sequence with the sequence of SEQ ID NO: 62 or 66 or an amino acid sequence that differs from SEQ ID NO: 62 or 66 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 62 or 66 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). In some embodiments, the chimeric protein as shown in SEQ ID NO: 62 or 66 is modified with respect to the signal peptide, transmembrane domains and / or linkers.
[0500] In some embodiments, the chimeric protein (named "string 6") comprises PE13, EsxG, EsxH, Esxl, EsxJ, EsxN and EsxW. In some embodiments, the chimeric protein comprises one or more antigenic fragments of PE13, EsxJ and EsxN and full length antigens of EsxG, EsxH, Esxl and EsxW. In some embodiments, the chimeric protein comprises two antigen fragments of PE13, one antigen fragment of EsxJ, one antigen fragment of EsxN and full length antigens of EsxG, EsxH, Esxl and EsxW. In some embodiments, the chimeric protein comprises an amino acid sequence with the sequence of SEQ ID NO: 63 or 67 or an amino acid sequence that differs from SEQ ID NO: 63 or 67 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 63 or 67 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). In some embodiments, the chimeric protein as shown in SEQ ID NO: 63 or 67 is modified with respect to the signal peptide, transmembrane domains and / or linkers.
[0501] In some embodiments, the chimeric protein (named "string 7") comprises PE13, EsxA, EsxG, EsxH, Esxl, EsxJ, EsxN and EsxW. In some embodiments, the chimeric protein comprises one or more antigenic fragments of PE13, EsxA, EsxJ and EsxN and full length antigens of EsxG, EsxH, Esxl and EsxW. In some embodiments, the chimeric protein comprises two antigen fragments of PE13, two antigen fragments of EsxA, one antigen fragment of EsxJ, one antigen fragment of EsxN and full length antigens of EsxG, EsxH, Esxl and EsxW. In some embodiments, the chimeric protein comprises an amino acid sequence with the sequence of SEQ ID NO: 64 or 68 or an amino acid sequence that differs from SEQ ID NO: 64 or 68 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 64 or 68 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). In some embodiments, the chimeric protein as shown in SEQ ID NO: 64 or 68 is modified with respect to the signal peptide, transmembrane domains and / or linkers.
[0502] In some embodiments, the chimeric protein (named "string 8") comprises PE13, Esxl, EsxJ, EsxN and EsxW. In some embodiments, the chimeric protein comprises one or more antigenic fragments of PE13, EsxJ and EsxN and full length antigens of Esxl and EsxW. In some embodiments, the chimeric protein comprises two antigen fragments of PE13, one antigen fragment of EsxJ, one antigen fragment of EsxN and full length antigens of Esxl and EsxW. In some embodiments, the chimeric protein comprises an amino acid sequence with the sequence of SEQ ID NO: 173 or 174 or an amino acid sequence that differs from SEQ ID NO: 173 or 174 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 173 or 174 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). In some embodiments, the chimeric protein as shown in SEQ ID NO: 173 or 174 is modified with respect to the signal peptide, transmembrane domains and / or linkers. In some embodiments, the chimeric protein is encoded by a nucleic acid molecule comprising a nucleotide sequence with the sequence of SEQ ID NO: 175 or 176 or a nucleotide that differs from SEQ ID NO: 175 or 176 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 175 or 176 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0503] In some embodiments, the chimeric protein (named "string 9") comprises PE13, PPE18, Esxl, EsxJ, EsxN and EsxW. In some embodiments, the chimeric protein comprises one or more antigenic fragments of PE13, PPE18, EsxJ and EsxN and full length antigens of Esxl and EsxW. In some embodiments, the chimeric protein comprises two antigen fragments of PE13, two antigen fragments of PPE18, one antigen fragment of EsxJ, one antigen fragment of EsxN and full length antigens of Esxl and EsxW. In some embodiments, the chimeric protein comprises an amino acid sequence with the sequence of SEQ ID NO: 177 or 178 or an amino acid sequence that differs from SEQ ID NO: 177 or 178 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 177 or 178 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). In some embodiments, the chimeric protein as shown in SEQ ID NO: 177 or 178 is modified with respect to the signal peptide, transmembrane domains and / or linkers. In some embodiments, the chimeric protein is encoded by a nucleic acid molecule comprising a nucleotide sequence with the sequence of SEQ ID NO: 179 or 180 or a nucleotide that differs from SEQ ID NO: 179 or 180 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 179 or 180 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0504] In some embodiments, the chimeric protein (named "string 10" or "string 10_R") comprises PE13, PPE18, EsxG, Esxl and EsxN. In some embodiments, the chimeric protein comprises one or more antigenic fragments of PE13, PPE18 and EsxN and full length antigens of EsxG and Esxl. In some embodiments, the chimeric protein comprises two antigen fragments of PE13, three antigen fragments of PPE18, one antigen fragment of EsxN and full length antigens of EsxG and Esxl. In some embodiments, the chimeric protein comprises an amino acid sequence with the sequence of SEQ ID NO: 141 or 142 or an amino acid sequence that differs from SEQ ID NO: 141 or 142 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 141 or 142 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). In some embodiments, the chimeric protein as shown in SEQ ID NO: 141 or 142 is modified with respect to the signal peptide, transmembrane domains and / or linkers. In some embodiments, the chimeric protein is encoded by a nucleic acid molecule comprising a nucleotide sequence with the sequence of any one of SEQ ID NO: 181 to 192 or a nucleotide that differs from SEQ ID NO: 181 to 192 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 181 to 192 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0505] In some embodiments, the chimeric protein (named "string 11" or "string 11_R") comprises PE13, EsxG, EsxH, Esxl, EsxJ, EsxN and EsxW. In some embodiments, the chimeric protein comprises one or more antigenic fragments of PE13, EsxJ and EsxN and full length antigens of EsxG, EsxH, Esxl and EsxW. In some embodiments, the chimeric protein comprises two antigen fragments of PE13, one antigen fragment of EsxJ, one antigen fragment of EsxN and full length antigens of EsxG, EsxH, Esxl and EsxW. In some embodiments, the chimeric protein comprises an amino acid sequence with the sequence of SEQ ID NO: 143 or 144 or an amino acid sequence that differs from SEQ ID NO: 143 or 144 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 143 or 144 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). In some embodiments, the chimeric protein as shown in SEQ ID NO: 143 or 144 is modified with respect to the signal peptide, transmembrane domains and / or linkers. In some embodiments, the chimeric protein is encoded by a nucleic acid molecule comprising a nucleotide sequence with the sequence of any one of SEQ ID NO: 193 to 204 or a nucleotide that differs from SEQ ID NO: 193 to 204 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 193 to 204 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0506] In some embodiments, the chimeric protein (named "string 12" or "string 12_R") comprises PE13, EsxA, EsxG, EsxH, Esxl, EsxJ, EsxN and EsxW. In some embodiments, the chimeric protein comprises one or more antigenic fragments of PE13, EsxA, EsxJ and EsxN and full length antigens of EsxG, EsxH, Esxl and EsxW. In some embodiments, the chimeric protein comprises two antigen fragments of PE13, two antigen fragments of EsxA, one antigen fragment of EsxJ, one antigen fragment of EsxN and full length antigens of EsxG, EsxH, Esxl and EsxW. In some embodiments, the chimeric protein comprises an amino acid sequence with the sequence of SEQ ID NO: 145 or 146 or an amino acid sequence that differs from SEQ ID NO: 145 or 146 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 145 or 146 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTP). In some embodiments, the chimeric protein as shown in SEQ ID NO: 145 or 146 is modified with respect to the signal peptide, transmembrane domains and / or linkers. In some embodiments, the chimeric protein is encoded by a nucleic acid molecule comprising a nucleotide sequence with the sequence of any one of SEQ ID NO: 205 to 216 or a nucleotide that differs from SEQ ID NO: 205 to 216 by one or more deletions, insertions or substitutions while having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to SEQ ID NO: 205 to 216 (e.g. determined by sequence alignment using known sequence alignment algorithms such as BLASTN).
[0507] able 4: Native amino acid and nucleotide sequences of Mycobacterium tuberculosis antigens used ccgcaggccgcggtgcagcgctaccccaacgtgcggctgctgcccacaggggccaacctcgggtacggaaccgcggtgaatcggacgatcgcccagctcggtgaaatggcgggcgat gccggcgaaccctgggtcgatgactgggtgatcgtggccaacccggacgtgcaatggggcccgggcagtatcgatgcactactggacgccgcctcccgctggccccgcgcgggcgcgc tgggcccgctgattcgggaccccgacgggtcggtgtacccgtcggcgcggcagatgcccagcctgatccgcggcggcatgcacgcagtgctcgggccgttctggccgcgcaatccgtgg acgacggcctaccggcaggagcggctggagcccagtgaacggccggtgggttggttgtcggggtcttgcctactggtgcgccggtcggcgtttggccaggtcggcggattcgacgaacg ttacttcatgtacatggaggacgtcgaccttggcgaccggcttggcaaagccggttggctgtcggtgtatgtgccgtcagccgaggttctgcaccacaaggcgcattcgacgggtcgcgac ccggcaagccatctggccgcccatcacaaaagcacctatatcttcttagccgaccgacattctggttggtggcgggctccgctgcgctggaccctgcggggatcactggcgctgcgttccca cctcatggtgcgcagttcactgcgcaggtcccgcagacggaaactgaagctggtagaagggcggcactga [SEQ ID NO: 147]
[0508] MVDFGALPPEINSARMYAGPGSASLVAAAQMWDSVASDLFSAASAFQSWWGLTVGSWIGSSAGLMVAAASPYVAWMSVTAGQAELTAAQVRVA AAAYETAYGLTVPPPVIAENRAELMILIATNLLGQNTPAIAVNEAEYGEMWAQDAAAMFGYAAATATATATLLPFEEAPEMTSAGGLLEQAAAVEEA SDTAAANQLMNNVPQALQQLAQPTQGTTPSSKLGGLWKTVSPHRSPISNMVSMANNHMSMTNSGVSMTNTLSSMLKGFAPAAAAQAVQTAAQN GVRAMSSLGSSLGSSGLGGGVAANLGRAASVGSLSVPQAWAAANQAVTPAARALPLTSLTSAAERGPGQMLGGLPVGQMGARAGGGLSGVLRVPP RPYVMPHSPAAG [SEQ ID NO: 2] atggtggatttcggggcgttaccaccggagatcaactccgcgaggatgtacgccggcccgggttcggcctcgctggtggccgcggctcagatgtgggacagcgtggcgagtgacctgttt
[0509] PPE18 tcggccgcgtcggcgtttcagtcggtggtctggggtctgacggtggggtcgtggataggttcgtcggcgggtctgatggtggcggcggcctcgccgtatgtggcgtggatgagcgtcacc gcggggcaggccgagctgaccgccgcccaggtccgggttgctgcggcggcctacgagacggcgtatgggctgacggtgcccccgccggtgatcgccgagaaccgtgctgaactgatg attctgatagcgaccaacctcttggggcaaaacaccccggcgatcgcggtcaacgaggccgaatacggcgagatgtgggcccaagacgccgccgcgatgtttggctacgccgcggcga cggcgacggcgacggcgacgttgctgccgttcgaggaggcgccggagatgaccagcgcgggtgggctcctcgagcaggccgccgcggtcgaggaggcctccgacaccgccgcggcg aaccagttgatgaacaatgtgccccaggcgctgcaacagctggcccagcccacgcagggcaccacgccttcttccaagctgggtggcctgtggaagacggtctcgccgcatcggtcgccg atcagcaacatggtgtcgatggccaacaaccacatgtcgatgaccaactcgggtgtgtcgatgaccaacaccttgagctcgatgttgaagggctttgctccggcggcggccgcccaggcc gtgcaaaccgcggcgcaaaacggggtccgggcgatgagctcgctgggcagctcgctgggttcttcgggtctgggcggtggggtggccgccaacttgggtcgggcggcctcggtcggttc gttgtcggtgccgcaggcctgggccgcggccaaccaggcagtcaccccggcggcgcgggcgctgccgctgaccagcctgaccagcgccgcggaaagagggcccgggcagatgctgg gcgggctgccggtggggcagatgggcgccagggccggtggtgggctcagtggtgtgctgcgtgttccgccgcgaccctatgtgatgccgcattctccggcggccggctag [SEQ ID NO: 148]
[0510] Antigen Amino acid sequence / nucleotide sequence
[0511] MSFVMAYPEMLAAAADTLQSIGATTVASNAAAAAPTTGWPPAADEVSALTAAHFAAHAAMYQSVSARAAAIHDQFVATLASSASSYAATEVANAA AAS [SEQ ID NO: 3] or MHVSFVMAYPEMLAAAADTLQSIGATTVASNAAAAAPTTGVVPPAADEVSALTAAHFAAHAAMYQSVSARAAAIHDQFVATLASSASSYAATEVAN AAAAS [SEQ ID NO: 4]
[0512] PE13 gtgtctttcgtgatggcatacccagagatgttggcggcggcggctgacaccctgcagagcatcggtgctaccactgtggctagcaatgccgctgcggcggccccgacgactggggtggtg ccccccgctgccgatgaggtgtcggcgctgactgcggcgcacttcgccgcacatgcggcgatgtatcagtccgtgagcgctcgggctgctgcgattcatgaccagttcgtggccacccttgc cagcagcgccagctcgtatgcggccactgaagtcgccaatgcggcggcggccagctaa [SEQ ID NO: 149] or gtgcacgtgtctttcgtgatggcatacccagagatgttggcggcggcggctgacaccctgcagagcatcggtgctaccactgtggctagcaatgccgctgcggcggccccgacgactggg gtggtgccccccgctgccgatgaggtgtcggcgctgactgcggcgcacttcgccgcacatgcggcgatgtatcagtccgtgagcgctcgggctgctgcgattcatgaccagttcgtggcca cccttgccagcagcgccagctcgtatgcggccactgaagtcgccaatgcggcggcggccagctaa [SEQ ID NO: 150]
[0513] MTEQQWNFAGIEAAASAIQGNVTSIHSLLDEGKQSLTKLAAAWGGSGSEAYQGVQQKWDATATELNNALQNLARTISEAGQAMASTEGNVTGMF A [SEQ ID NO: 5] EsxA atgacagagcagcagtggaatttcgcgggtatcgaggccgcggcaagcgcaatccagggaaatgtcacgtccattcattccctccttgacgaggggaagcagtccctgaccaagctcgc agcggcctggggcggtagcggttcggaggcgtaccagggtgtccagcaaaaatgggacgccacggctaccgagctgaacaacgcgctgcagaacctggcgcggacgatcagcgaag ccggtcaggcaatggcttcgaccgaaggcaacgtcactgggatgttcgcatag [SEQ ID NO: 151]
[0514] MAEMKTDAATLAQEAGNFERISGDLKTQIDQVESTAGSLQGQWRGAAGTAAQAAWRFQEAANKQKQELDEISTNIRQAGVQYSRADEEQQQAL SSQMGF [SEQ ID NO: 6]
[0515] EsxB atggcagagatgaagaccgatgccgctaccctcgcgcaggaggcaggtaatttcgagcggatctccggcgacctgaaaacccagatcgaccaggtggagtcgacggcaggttcgttgc agggccagtggcgcggcgcggcggggacggccgcccaggccgcggtggtgcgcttccaagaagcagccaataagcagaagcaggaactcgacgagatctcgacgaatattcgtcag gccggcgtccaatactcgagggccgacgaggagcagcagcaggcgctgtcctcgcaaatgggcttctga [SEQ ID NO: 152]
[0516] MSLLDAHIPQLVASQSAFAAKAGLMRHTIGQAEQAAMSAQAFHQGESSAAFQAAHARFVAAAAKVNTLLDVAQANLGEAAGTYVAADAAAASTYT GF [SEQ ID NO: 7]
[0517] EsxG atgagccttttggatgctcatatcccacagttggtggcctcccagtcggcgtttgccgccaaggcggggctgatgcggcacacgatcggtcaggccgagcaggcggcgatgtcggctcag gcgtttcaccagggggagtcgtcggcggcgtttcaggccgcccatgcccggtttgtggcggcggccgccaaagtcaacaccttgttggatgtcgcgcaggcgaatctgggtgaggccgcc ggtacctatgtggccgccgatgctgcggccgcgtcgacctataccgggttctga [SEQ ID NO: 153]
[0518] Antigen Amino acid sequence / nucleotide sequence
[0519] MSQIMYNYPAMLGHAGDMAGYAGTLQSLGAEIAVEQAALQSAWQGDTGITYQAWQAQWNQAMEDLVRAYHAMSSTHEANTMAMMARDTAEA AKWGG [SEQ ID NO: 8]
[0520] EsxH atgtcgcaaatcatgtacaactaccccgcgatgttgggtcacgccggggatatggccggatatgccggcacgctgcagagcttgggtgccgagatcgccgtggagcaggccgcgttgca gagtgcgtggcagggcgataccgggatcacgtatcaggcgtggcaggcacagtggaaccaggccatggaagatttggtgcgggcctatcatgcgatgtccagcacccatgaagccaac accatggcgatgatggcccgcgacacggccgaagccgccaaatggggcggctag [SEQ ID NO: 154]
[0521] MTINYQFGDVDAHGAMIRAQAGSLEAEHQAIISDVLTASDFWGGAGSAACQGFITQLGRNFQVIYEQANAHGQKVQAAGNNMAQTDSAVGSSW
[0522] A [SEQ ID NO: 9]
[0523] Esxl atgaccatcaactatcaattcggggacgtcgacgctcacggcgccatgatccgcgctcaggccgggtcgctggaggccgagcatcaggccatcatttctgatgtgttgaccgcgagtgact tttggggcggcgccggttcggcggcctgccaggggttcattacccagctgggccgtaacttccaggtgatctacgagcaggccaacgcccacgggcagaaggtgcaggctgccggcaac aacatggcacaaaccgacagcgccgtcggctccagctgggcctaa [SEQ ID NO: 155] MASRFMTDPHAMRDMAGRFEVHAQTVEDEARRMWASAQNISGAGWSGMAEATSLDTMTQMNQAFRNIVNMLHGVRDGLVRDANNYEQQEQA
[0524] SQQILSS [SEQ ID NO: 10]
[0525] EsxJ atggcctcgcgttttatgacggatccgcacgcgatgcgggacatggcgggccgttttgaggtgcacgcccagacggtggaggacgaggctcgccggatgtgggcgtccgcgcaaaacat ctcgggcgcgggctggagtggcatggccgaggcgacctcgctagacaccatgacccagatgaatcaggcgtttcgcaacatcgtgaacatgctgcacggggtgcgtgacgggctggttc gcgacgccaacaactacgaacagcaagagcaggcctcccagcagatcctcagcagctga [SEQ ID NO: 156]
[0526] MASRFMTDPHAMRDMAGRFEVHAQTVEDEARRMWASAQNISGAGWSGMAEATSLDTMAQMNQAFRNIVNMLHGVRDGLVRDANNYEQQEQA SQQILSS [SEQ ID NO: 11]
[0527] EsxK atggcctcacgttttatgacggatccgcacgcgatgcgggacatggcgggccgttttgaggtgcacgcccagacggtggaggacgaggctcgccggatgtgggcgtccgcgcaaaacat ttccggtgcgggctggagtggcatggccgaggcgacctcgctagacaccatggcccagatgaatcaggcgtttcgcaacatcgtgaacatgctgcacggggtgcgtgacgggctggttc gcgacgccaacaactacgagcagcaagagcaggcctcccagcagatcctcagcagctaa [SEQ ID NO: 157]
[0528] Amino acid nucleotide
[0529] MTINYQFGDVDAHGAMIRAQAGLLEAEHQAIIRDVLTASDFWGGAGSAACQGFITQLGRNFQVIYEQANAHGQKVQAAGNNMAQTDSAVGSSW
[0530] A [SEQ ID NO: 12]
[0531] EsxL atgaccatcaactatcaattcggggatgtcgacgctcacggcgccatgatccgcgctcaggccgggttgctggaggccgagcatcaggccatcattcgtgatgtgttgaccgcgagtgact tttggggcggcgccggttcggcggcctgccaggggttcattacccagttgggccgtaacttccaggtgatctacgagcaggccaacgcccacgggcagaaggtgcaggctgccggcaac aacatggcgcaaaccgacagcgccgtcggctccagctgggcctga [SEQ ID NO: 158]
[0532] MASRFMTDPHAMRDMAGRFEVHAQTVEDEARRMWASAQNISGAGWSGMAEATSLDTMT+MNQAFRNIVNMLHGVRDGLVRDANNYEQQEQA SQQILSS [SEQ ID NO: 13]
[0533] EsxM atggcctcacgttttatgacggatccgcatgcgatgcgggacatggcgggccgttttgaggtgcacgcccagacggtggaggacgaggctcgccggatgtgggcgtccgcgcaaaacat ttccggtgcgggctggagtggcatggccgaggcgacctcgctagacaccatgacctagatgaatcaggcgtttcgcaacatcgtgaacatgctgcacggggtgcgtgacgggctggttcg cgacgccaacaactacgaacagcaagagcaggcctcccagcagatcctgagcagctag [SEQ ID NO: 159]
[0534] MTINYQFGDVDAHGAMIRAQAASLEAEHQAIVRDVLAAGDFWGGAGSVACQEFITQLGRNFQVIYEQANAHGQKVQAAGNNMAQTDSAVGSSW A [SEQ ID NO: 14]
[0535] EsxN atgacgattaattaccagttcggggacgtcgacgctcatggcgccatgatccgcgctcaggcggcgtcgcttgaggcggagcatcaggccatcgttcgtgatgtgttggccgcgggtgact tttggggcggcgccggttcggtggcttgccaggagttcattacccagttgggccgtaacttccaggtgatctacgagcaggccaacgcccacgggcagaaggtgcaggctgccggcaac aacatggcgcaaaccgacagcgccgtcggctccagctgggcctaa [SEQ ID NO: 160]
[0536] MTINYQFGDVDAHGAMIRAQAGSLEAEHQAIISDVLTASDFWGGAGSAACQGFITQLGRNFQVIYEQANAHGQKVQAAGNNMAQTDSAVGSSW A [SEQ ID NO: 15]
[0537] EsxV atgaccatcaactatcaattcggggacgtcgacgctcacggcgccatgatccgcgctcaggccgggtcgctggaggccgagcatcaggccatcatttctgatgtgttgaccgcgagtgact tttggggcggcgccggttcggcggcctgccaggggttcattacccagctgggccgtaacttccaggtgatctacgagcaggccaacgcccacgggcagaaggtgcaggctgccggcaac aacatggcacaaaccgacagcgccgtcggctccagctgggcctaa [SEQ ID NO: 161]
[0538] Amino acid nucleotide
[0539] MTSRFMTDPHAMRDMAGRFEVHAQTVEDEARRMWASAQNISGAGWSGMAEATSLDTMTQMNQAFRNIVNMLHGVRDGLVRDANNYEQQEQA SQQILSS [SEQ ID NO: 16]
[0540] EsxW atgacctcgcgttttatgacggatccgcacgcgatgcgggacatggcgggccgttttgaggtgcacgcccagacggtggaggacgaggctcgccggatgtgggcgtccgcgcaaaacat ttccggcgcgggctggagtggcatggccgaggcgacctcgctagacaccatgacccagatgaatcaggcgtttcgcaacatcgtgaacatgctgcacggggtgcgtgacgggctggttc gcgacgccaacaactacgaacagcaagagcaggcctcccagcagatcctcagcagctga [SEQ ID NO: 162] able 5: Amino acid sequences of exemplary chimeric proteins
[0541] Chimeric . . . prot .e .in Amino acid sequence
[0542] MGGAAARLGAVILFWIVGLHGVRGMHMSFVMAYPEMLAAAADTLQSIGATTVAGGSGGGGSGGGGVAANLGRAASVGSLSVPQAWAAANQAVTPAARALPLTSLTSAAERGPGQ
[0543] MLGGLPVGQMGARAGGGLSGVLRVPPRPYVMPHSPAAGGGSGGGGSGGWAVTYSPGPHLERFLASLSLATERPVSVLLADNGSTDGTPQAAVQRYPNVRLLPTGANLGYGTAVNR
[0544] TIAQLGEMAGDAGEPWVDDWVIVANPDVQWGPGSIDALLDAASRWPRAGALGPLIRDPDGSVYPSARQMPSLIRGGMHAVLGPFWPRNPWTTAYRQERLEPSERPVGWLSGSCL
[0545] String 1 LVRRSAFGQVGGFDERYFMYMEDVDLGDRLGKAGWLSVYVPSAEVLHHKAHSTGRDPASHLAAHHKSTYIFLADRHSGWWRAPLRWTLRGSLALRSHLMVRSSLRRSRRRKLKLV
[0546] EGRHGGSGGGGSGGAAVEEASDTAAANQLMNNVPQALQQLAQPTQGTTPSSKLGGLWKTVSPHRSPISNMVSMANNHMSMTNSGVSMTNTLSSMLKGFAPAAAAQAVQTAAQN
[0547] GVRAMSSLGSSLGGSGGGGSGGWPPAADEVSALTAAHFAAHAAMYQSVSARAAAIHDQFVATLASSASSYAATEVANAAAASGGSGGGGSGGMVDFGALPPEINSARMYAGPGSA SLVAAAQMWDSVASDLFSAASAFQSWWGLTVGSWIGSSAGLMVAAASPYVAWMSVTAGQAELTAAQVRVAAAAYETAYGLTVPPPVIAENRAELMILIATNLLGQNTPAIAVNEAE YGEMWAQDAAAMFGYGGSGGGGSGGIVGIVAGLAVLAVWIGAWATVMCRRKSSGGKGGSYSQAASSDSAQGSDVSLTA* [SEQ ID NO: 57]
[0548] MGGAAARLGAVILFWIVGLHGVRGMSLLDAHIPQLVASQSAFAAKAGLMRHTIGQAEQAAMSAQAFHQGESSAAFQAAHARFVAAAAKVNTLLDVAQANLGEAAGTYVAADAAAA
[0549] STYTGFGGSGGGGSGGMTINYQFGDVDAHGAMIRAQAGSLEAEHQAIISDVLTASDFWGGAGSAACQGFITQLGRNFQVIYEQANAHGQKVQAAGNNMAQTDSAVGSSWAGGSG
[0550] _ .?GGGSGGMSQIMYNYPAMLGHAGDMAGYAGTLQSLGAEIAVEQAALQSAWQGDTGITYQAWQAQWNQAMEDLVRAYHAMSSTHEANTMAMMARDTAEAAKWGGGGSGGGGS btring z GGMASRFMTDPHAMRDGGSGGGGSGGMTINYQFGDVDAHGAMIRAQAASLEAEHQAIVRDVLAAGDFWGGAGSVACQEFITQLGRNFQVIYEQANAHGQKVQAAGNNMAQTD
[0551] SAVGSSWAGGSGGGGSGGMTSRFMTDPHAMRDMAGRFEVHAQTVEDEARRMWASAQNISGAGWSGMAEATSLDTMTQMNQAFRNIVNMLHGVRDGLVRDANNYEQQEQASQ
[0552] QILSSGGSGGGGSGGIVGIVAGLAVLAWVIGAWATVMCRRKSSGGKGGSYSQAASSDSAQGSDVSLTA* [SEQ ID NO: 59]
[0553] MGGAAARLGAVILFVVIVGLHGVRGMAEMKTDAATLAQEAGNFERISGDLKTQIDQVESTAGSLQGQWRGAAGTAAQAAVVRFQEAANKQKQELDEISTNIRQAGVQYSRADGGS GGGGSGGMSLLDAHIPQLVASQSAFAAKAGLMRHTIGQAEQAAMSAQAFHQGESSAAFQAAHARFVAAAAKVNTLLDVAQANLGEAAGTYVAADAAAASTYTGFGGSGGGGSGGM TINYQFGDVDAHGAMIRAQAGSLEAEHQAIISDVLTASDFWGGAGSAACQGFITQLGRNFQVIYEQANAHGQKVQAAGNNMAQTDSAVGSSWAGGSGGGGSGGMSQIMYNYPA
[0554] String 3 MLGHAGDMAGYAGTLQSLGAEIAVEQAALQSAWQGDTGITYQAWQAQWNQAMEDLVRAYHAMSSTHEANTMAMMARDTAEAAKWGGGGSGGGGSGGMASRFMTDPHAMRD
[0555] GGSGGGGSGGMTINYQFGDVDAHGAMIRAQAASLEAEHQAIVRDVLAAGDFWGGAGSVACQEFITQLGRNFQVIYEQANAHGQKVQAAGNNMAQTDSAVGSSWAGGSGGGGSG GMTSRFMTDPHAMRDMAGRFEVHAQTVEDEARRMWASAQNISGAGWSGMAEATSLDTMTQMNQAFRNIVNMLHGVRDGLVRDANNYEQQEQASQQILSSGGSGGGGSGGIVG IVAGLAVLAVWIGAVVATVMCRRKSSGGKGGSYSQAASSDSAQGSDVSLTA* [SEQ ID NO: 58]
[0556] Chim ....
Claims
Claims1. RNA molecule encoding a chimeric protein comprising T cell epitopes of three or more different Mycobacterium tuberculosis antigens or immunogenic variants thereof, wherein the chimeric protein comprises one or more antigen fragments, wherein at least one Mycobacterium tuberculosis antigen or immunogenic variant thereof is represented by one or more antigen fragments and the remaining Mycobacterium tuberculosis antigens or immunogenic variants thereof are represented by one or more antigen fragments and / or one or more full-length antigens, wherein each antigen fragment or full-length antigen comprises one or more T cell epitopes, and wherein each antigen fragment or full-length antigen is separated from other antigen fragments or full- length antigens in the chimeric protein by a polypeptide linker.
2. The RNA molecule of claim 1, wherein the chimeric protein comprises T cell epitopes of four or more different Mycobacterium tuberculosis antigens.
3. The RNA molecule of claim 1, wherein the chimeric protein comprises T cell epitopes of five or more different Mycobacterium tuberculosis antigens.
4. The RNA molecule of claim 1, wherein the chimeric protein comprises T cell epitopes of six or more different Mycobacterium tuberculosis antigens.
5. The RNA molecule of claim 1, wherein the chimeric protein comprises T cell epitopes of seven or more different Mycobacterium tuberculosis antigens.
6. The RNA molecule of any one of claims 1 to 5, wherein the Mycobacterium tuberculosis antigens are selected from the group of Wbbll, PPE18, PE13, EsxA, EsxB, EsxG, EsxH, Esxl, Esxl, EsxK, EsxL, EsxM, EsxN, EsxV and EsxW.
7. The RNA molecule of claim 6, wherein: p) the Wbbll antigen comprises the amino acid sequence of SEQ ID NO: 1 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 1; q) the PPE18 antigen comprises the amino acid sequence of SEQ ID NO: 2 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 2;r) the PE13 antigen comprises the amino acid sequence of SEQ ID NO: 3 or SEQ ID NO: 4 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 3 or SEQ ID NO: 4; s) the EsxA antigen comprises the amino acid sequence of SEQ ID NO: 5 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 5; t) the EsxB antigen comprises the amino acid sequence of SEQ ID NO: 6 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 6; u) the EsxG antigen comprises the amino acid sequence of SEQ ID NO: 7 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 7; v) the EsxH antigen comprises the amino acid sequence of SEQ ID NO: 8 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 8; w) the Esxl antigen comprises the amino acid sequence of SEQ ID NO: 9 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 9; x) the Esxl antigen comprises the amino acid sequence of SEQ ID NO: 10 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 10; y) the EsxK antigen comprises the amino acid sequence of SEQ ID NO: 11 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 11; z) the EsxL antigen comprises the amino acid sequence of SEQ ID NO: 12 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 12; aa) the EsxM antigen comprises the amino acid sequence of SEQ ID NO: 13 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 13; bb) the EsxN antigen comprises the amino acid sequence of SEQ ID NO: 14 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 14;cc) the EsxV antigen comprises the amino acid sequence of SEQ ID NO: 15 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 15; and / or dd) the EsxW antigen comprises the amino acid sequence of SEQ ID NO: 16 and an immunogenic variant thereof has an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of SEQ ID NO: 16.
8. The RNA molecule of claim 6 or 7, wherein the Mycobacterium tuberculosis antigens comprise Wbbll, PPE18 and PE13.
9. The RNA molecule of claim 8, wherein the chimeric protein comprises: d) a first and a second antigen fragment of PE13; e) a first, a second and a third antigen fragment of PPE18; and f) a full-length antigen of WbbLl.
10. The RNA molecule of claim 9, wherein: f) the first antigen fragment of PE13 comprises the amino acid sequence of positions 1 to 27 of SEQ ID NO: 3 or an amino acid sequence having at least 96%, 92%, 88%, 84% or 80% identity to the amino acid sequence of positions 1 to 27 of SEQ ID NO: 3; g) the second antigen fragment of PE13 comprises the amino acid sequence of positions 39 to 99 of SEQ ID NO: 3 or an amino acid sequence having at least 98%, 96%, 94%, 92%, 88%, 84% or 80% identity to the amino acid sequence of positions 39 to 99 of SEQ ID NO: 3; h) the first antigen fragment of PPE18 comprises the amino acid sequence of positions 1 to 155 of SEQ ID NO: 2 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 1 to 155 of SEQ ID NO: 2; i) the second antigen fragment of PPE18 comprises the amino acid sequence of positions 186 to 296 of SEQ ID NO: 2 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 186 to 296 of SEQ ID NO: 2; and / or j) the third antigen fragment of PPE18 comprises the amino acid sequence of positions 303 to 391 of SEQ ID NO: 2 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 303 to 391 of SEQ ID NO: 2.
11. The RNA molecule of claim 9 or 10, wherein the full-length antigens and antigen fragments in the chimeric protein are arranged in the following order:first antigen fragment of PE13 - linker - third antigen fragment of PPE18 - linker - full-length antigen of Wbbll - linker - second antigen fragment of PPE18 - linker - second antigen fragment of PE13 - linker - first antigen fragment of PPE18.
12. The RNA molecule of claim 6 or 7, wherein the Mycobacterium tuberculosis antigens comprise EsxG, EsxH, Esxl, EsxJ, EsxN, and EsxW.
13. The RNA molecule of claim 12, wherein the Mycobacterium tuberculosis antigens additionally comprise EsxB.
14. The RNA molecule of claim 12, wherein the chimeric protein comprises: g) a full-length antigen of EsxG; h) a full-length antigen of EsxH; i) a full-length antigen of Esxl; j) an antigen fragment of EsxJ; k) a full-length antigen of EsxN; and l) a full-length antigen of EsxW.
15. The RNA molecule of claim 13, wherein the chimeric protein additionally comprises: h) an antigen fragment of EsxB i) a full-length antigen of EsxG; j) a full-length antigen of EsxH; k) a full-length antigen of Esxl; l) an antigen fragment of EsxJ; m) a full-length antigen of EsxN; and n) a full-length antigen of EsxW.
16. The RNA molecule of any one of claims 14 or 15, wherein: c) the antigen fragment of EsxB comprises the amino acid sequence of positions 1 to 87 of SEQ ID NO: 6 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 1 to 87 of SEQ ID NO: 6; and / ord) the antigen fragment of EsxJ comprises the amino acid sequence of positions 1 to 14 of SEQ ID NO: 10 or an amino acid sequence having at least 93%, 86% or 79% identity to the amino acid sequence of positions 1 to 14 of SEQ ID NO: 10.
17. The RNA molecule of claim 14 or 16, wherein the full-length antigens and antigen fragments in the chimeric protein are arranged in the following order: c) full-length antigen of EsxG - linker - full-length antigen of EsxH - linker - full-length antigen of Esxl - linker - antigen fragment of EsxJ - linker - full-length antigen of EsxN - linker - full-length antigen of EsxW; or d) full-length antigen of EsxG - linker - full-length antigen of Esxl - linker - full-length antigen of EsxH - linker - antigen fragment of EsxJ - linker - full-length antigen of EsxN - linker - full- length antigen of EsxW.
18. The RNA molecule of claim 15 or 16, wherein the full-length antigens and antigen fragments in the chimeric protein are arranged in the following order: c) antigen fragment of EsxB - linker - full-length antigen of EsxG - linker - full-length antigen of EsxH - linker - full-length antigen of Esxl - linker - antigen fragment of EsxJ - linker - full-length antigen of EsxN - linker - full-length antigen of EsxW; or d) antigen fragment of EsxB - linker - full-length antigen of EsxG - linker - full-length antigen of Esxl - linker - full-length antigen of EsxH - linker - antigen fragment of EsxJ - linker - full-length antigen of EsxN - linker - full-length antigen of EsxW.
19. The RNA molecule of claim 6 or 7, wherein the Mycobacterium tuberculosis antigens comprise PE13 and two or more antigens selected from PPE18, EsxW, EsxJ, EsxG, Esxl, EsxN, EsxH and EsxA.
20. The RNA molecule of claim 19, wherein the Mycobacterium tuberculosis antigens comprise h) PE13, PPE18, Esxl, EsxJ, EsxN and EsxW; i) PE13, PPE18, EsxG, EsxJ, and EsxW; j) PE13, PPE18, EsxG, Esxl, and EsxN k) PE13, EsxG, EsxH, Esxl, EsxJ, EsxN and EsxW; l) PE13, EsxA, EsxG, EsxH, Esxl, EsxJ, EsxN and EsxW; m) PE13, Esxl, EsxJ, EsxN and EsxW; or n) PE13, PPE18, Esxl, EsxJ, EsxN and EsxW.
21. The RNA molecule of claim 19 or 20, wherein the chimeric protein comprises:g) a first and a second antigen fragment of PE13; h) a first, a second and a third antigen fragment of PPE18; i) a full-length antigen of Esxl; j) an antigen fragment of EsxJ; k) an antigen fragment of EsxN; and l) a full-length antigen of EsxW.
22. The RNA molecule of claim 19 or 20, wherein the chimeric protein comprises: f) a first and a second antigen fragment of PE13; g) a first, a second and a third antigen fragment of PPE18; h) a full-length antigen of EsxG; i) an antigen fragment of EsxJ; and j) a full-length antigen of EsxW.
23. The RNA molecule of claim 19 or 20, wherein the chimeric protein comprises: a) a first and a second antigen fragment of PE13; b) a first, a second and a third antigen fragment of PPE18; c) a full-length antigen of EsxG; d) an antigen fragment of Esxl; and e) a full-length antigen of EsxN.
24. The RNA molecule of claim 19 or 20, wherein the chimeric protein comprises: h) a first and a second antigen fragment of PE13; i) a full-length antigen of EsxG; j) a full-length antigen of EsxH; k) a full-length antigen of Esxl; l) an antigen fragment of EsxJ; m) an antigen fragment of EsxN; and n) a full-length antigen of EsxW.
25. The RNA molecule of claim 19 or 20, wherein the chimeric protein comprises: i) a first and a second antigen fragment of PE13;j) a first and a second antigen fragment of EsxA; k) a full-length antigen of EsxG; l) a full-length antigen of EsxH: m) a full-length antigen of Esxl; n) an antigen fragment of EsxJ; o) an antigen fragment of EsxN; and p) a full-length antigen of EsxW.
26. The RNA molecule of claim 19 or 20, wherein the chimeric protein comprises: f) a first and a second antigen fragment of PE13; g) a full-length antigen of Esxl; h) an antigen fragment of EsxJ; i) an antigen fragment of EsxN; and j) a full-length antigen of EsxW.
27. The RNA molecule of claim 19 or 20, wherein the chimeric protein comprises: g) a first and a second antigen fragment of PE13; h) a second and a third antigen fragment of PPE18; i) a full-length antigen of Esxl; j) an antigen fragment of EsxJ; k) an antigen fragment of EsxN; and l) a full-length antigen of EsxW.
28. The RNA molecule of any one of claims 21 to 27, wherein: j) the first antigen fragment of PE13 comprises the amino acid sequence of positions 1 to 27 of SEQ ID NO: 3 or an amino acid sequence having at least 96%, 92%, 88%, 84% or 80% identity to the amino acid sequence of positions 1 to 27 of SEQ ID NO: 3; or the first antigen fragment of PE13 comprises the amino acid sequence of positions 1 to 29 of SEQ ID NO: 4 or an amino acid sequence having at least 96%, 92%, 88%, 84% or 80% identity to the amino acid sequence of positions 1 to 29 of SEQ ID NO: 4; k) the second antigen fragment of PE13 comprises the amino acid sequence of positions 39 to 97 of SEQ ID NO: 3 or an amino acid sequence having at least 98%, 96%, 94%, 92%, 88%, 84% or 80% identity to the amino acid sequence of positions 39 to 97 of SEQ ID NO: 3; or the secondantigen fragment of PE13 comprises the amino acid sequence of positions 41 to 99 of SEQ ID NO: 4 or an amino acid sequence having at least 98%, 96%, 94%, 92%, 88%, 84% or 80% identity to the amino acid sequence of positions 41 to 99 of SEQ ID NO: 4; l) the first antigen fragment of PPE18 comprises the amino acid sequence of positions 1 to 155 of SEQ ID NO: 2 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 1 to 155 of SEQ ID NO: 2; m) the second antigen fragment of PPE18 comprises the amino acid sequence of positions 186 to 296 of SEQ ID NO: 2 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 186 to 296 of SEQ ID NO: 2; n) the third antigen fragment of PPE18 comprises the amino acid sequence of positions 303 to 391 of SEQ ID NO: 2 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 303 to 391 of SEQ ID NO: 2; o) the first antigen fragment of EsxA comprises the amino acid sequence of positions 1 to 35 of SEQ ID NO: 5 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 1 to 35 of SEQ ID NO: 5; p) the second antigen fragment of EsxA comprises the amino acid sequence of positions 26 to 81 of SEQ ID NO: 5 or an amino acid sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the amino acid sequence of positions 26 to 81 of SEQ ID NO: 5; q) the antigen fragment of EsxJ comprises the amino acid sequence of positions 1 to 14 of SEQ ID NO: 10 or an amino acid sequence having at least 93%, 86% or 79% identity to the amino acid sequence of positions 1 to 14 of SEQ ID NO: 10 and / or r) the antigen fragment of EsxN comprises the amino acid sequence of positions 10 to 67 of SEQ ID NO: 14 or an amino acid sequence having at least 93%, 86% or 79% identity to the amino acid sequence of positions 10 to 67 of SEQ ID NO: 14.
29. RNA molecule of claim 21 or 28, wherein the full-length antigens and antigen fragments in the chimeric protein are arranged in the following order: first antigen fragment of PE13 - linker - third antigen fragment of PPE18 - linker - full-length antigen of EsxW - linker - antigen fragment of EsxJ - linker - second antigen fragment of PPE18 - linker - second antigen fragment of PE13 - linker - first antigen fragment of PPE18 - linker - full-length antigen of Esxl - linker - antigen fragment of EsxN.
30. RNA molecule of claim 22 or 28, wherein the full-length antigens and antigen fragments in the chimeric protein are arranged in the following order:first antigen fragment of PE13 - linker - third antigen fragment of PPE18 - linker - full-length antigen of EsxW - linker - full-length antigen of EsxG - linker - second antigen fragment of PE13 - linker - second antigen fragment of PPE18 - linker - antigen fragment of EsxJ - linker - first antigen fragment of PPE18.
31. RNA molecule of claim 23 or 28, wherein the full-length antigens and antigen fragments in the chimeric protein are arranged in the following order: a) second antigen fragment of PE13 - linker - third antigen fragment of PPE18 - linker - antigen fragment of EsxN - linker - full-length antigen of EsxG - linker - first antigen fragment of PE13 - linker - second antigen fragment of PPE18 - linker - full-length antigen of Esxl - linker - first antigen fragment of PPE18; or b) second antigen fragment of PE13 - linker - first antigen fragment of PPE18 - linker - full-length antigen of Esxl - linker - second antigen fragment of PPE18 - linker - first antigen fragment of PE13 - linker - full-length antigen of EsxG - linker - antigen fragment of EsxN - linker - third antigen fragment of PPE18.
32. RNA molecule of claim 24 or 28, wherein the full-length antigens and antigen fragments in the chimeric protein are arranged in the following order: a) first antigen fragment of PE13 - linker - full-length antigen of EsxW - linker - full-length antigen of EsxG - linker - second antigen fragment of PE13 - linker - full-length antigen of Esxl - linker - antigen fragment of EsxN - linker - antigen fragment of EsxJ - linker - full-length antigen of EsxH; b) second antigen fragment of PE13 - linker - full-length antigen of EsxW - linker - full-length antigen of EsxG - linker - first antigen fragment of PE13 - linker - full-length antigen of Esxl - linker - antigen fragment of EsxN - linker - antigen fragment of EsxJ - linker - full-length antigen of EsxH; or c) second antigen fragment of PE13 - linker - full-length antigen of EsxH - linker - antigen fragment of EsxJ - linker - antigen fragment of EsxN - linker - full-length antigen of Esxl - linker - first antigen fragment of PE13 - linker - full-length antigen of EsxG - linker - full-length antigen of EsxW.
33. RNA molecule of claim 25 or 28, wherein the full-length antigens and antigen fragments in the chimeric protein are arranged in the following order: a) first antigen fragment of PE13 - linker - full-length antigen of EsxW - linker - full-length antigen of EsxG - linker - second antigen fragment of EsxA - linker - second antigen fragment of PE13 - linker - full-length antigen of Esxl - linker - first antigen fragment of EsxA - linker - antigen fragment of EsxN - linker - antigen fragment of EsxJ - linker - full-length antigen of EsxH; b) second antigen fragment of PE13 - linker - full-length antigen of EsxW - linker - full-length antigen of EsxG - linker - second antigen fragment of EsxA - linker - first antigen fragment of PE13 - linker - full-length antigen of Esxl - linker - first antigen fragment of EsxA - linker - antigen fragment of EsxN - linker - antigen fragment of EsxJ - linker - full-length antigen of EsxH; orc) second antigen fragment of PE13 - linker - full-length antigen of EsxH - linker - antigen fragment of EsxJ - linker - antigen fragment of EsxN - linker - first antigen fragment of EsxA - linker - full-length antigen of Esxl - linker - first antigen fragment of PE13 - linker - second antigen fragment of EsxA - linker - full-length antigen of EsxG - linker - full-length antigen of EsxW.
34. The RNA molecule of any one of claims 1 to 33, wherein one or more of the polypeptide linkers comprises one or more glycine and / or one or more serine amino acid.
35. The RNA molecule of any one of claim 1 to 34, wherein one or more of the polypeptide linkers is at least 1, at least 5 or at least 10 amino acids in length.
36. The RNA molecule of any one of claims 1 to 35, wherein one or more of the polypeptide linkers has the amino acid sequence of SEQ ID NO: 56.
37. The RNA molecule of any one of claims 1 to 36, wherein the chimeric protein comprises a non-native signal peptide at its N-terminus.
38. The RNA molecule of claim 37, wherein the non-native signal peptide is a human, bacterial or viral signal peptide.
39. The RNA molecule of claim 38, wherein the non-native signal peptide comprises a secretory signal.
40. The RNA molecule of claim 38 or 39, wherein the non-native signal peptide is functional in mammalian cells.
41. The RNA molecule of any one of claims 38 to 40, wherein the non-native signal peptide comprises an amino acid sequence selected from the group of SEQ ID NOs: 17 to 37 or 65, an amino acid sequence having at least 96%, 92%, 88%, 84% or 80% identity to SEQ ID NOs: 17 to 37 or 65, an amino acid sequence encoded by a nucleotide sequence selected from the group of SEQ ID NOs: 38 to 51, or an amino acid sequence having at least 96%, 92%, 88%, 84% or 80% identity to amino acid sequences encoded by a nucleotide sequence selected from the group of SEQ ID NOs: 38 to 51.
42. The RNA molecule of any one of claims 38 to 41, wherein the non-native signal peptide is a viral signal peptide.
43. The RNA molecule of claim 42, wherein the non-native signal peptide is a HSV-1 glycoprotein D signal peptide.
44. The RNA molecule of any one of claims 1 to 43, wherein the chimeric protein comprises a non-native transmembrane domain at its C-terminus.
45. The RNA molecule of any one of claim 44, wherein the non-native transmembrane domain is a human, bacterial or viral transmembrane domain.
46. The RNA molecule of claim 44 or 45, wherein the non-native transmembrane domain comprises a trafficking domain.
47. The RNA molecule of claim 46, wherein non-native trafficking domain is an MHC class I trafficking domain.
48. The RNA molecule of claim 47, wherein the MHC class I trafficking domain comprises the amino acid sequence of SEQ ID NO: 52 or an amino acid sequence having at least 98%, 96%, 90%, or 80% identity to the amino acid sequence of SEQ ID NO: 52.
49. The RNA molecule of claim 45, wherein the non-native transmembrane domain is a viral transmembrane domain.
50. The RNA molecule of claim 49, wherein the non-native transmembrane domain is derived from HSV.
51. The RNA molecule of any one of claims 1 to 50, wherein the RNA molecule comprises a 5' cap.
52. The RNA molecule of claim 51, wherein the 5' cap comprises a capl structure.
53. The RNA molecule of claim 51, wherein the 5'-cap comprises m27'3-OGppp(mi2' °)ApG.
54. The RNA molecule of any one of claims 1 to 53, wherein the RNA molecule comprises a 5'-UTR.
55. The RNA molecule of claim 54, wherein the 5'-UTR comprises a modified human alpha-globin 5'-UTR.
56. The RNA molecule of claim 54 or 55, wherein the 5'-UTR comprises the nucleotide sequence of SEQ IDNO: 53, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 53.
57. The RNA molecule of any one of claims 1 to 56, wherein the RNA comprises a 3'-UTR.
58. The RNA molecule of claim 57, wherein the 3'-UTR comprises a first sequence from the amino terminal enhancer of split (AES) messenger RNA and a second sequence from the mitochondrial encoded 12S ribosomal RNA.
59. The RNA molecule of claim 57 or 58, wherein the 3'-UTR comprises the nucleotide sequence of SEQ IDNO: 54, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 54.
60. The RNA molecule of any one of claims 1 to 59, wherein the RNA molecule comprises a polyA sequence.
61. The RNA molecule of claim 60, wherein the polyA sequence is an interrupted sequence of A nucleotides.
62. The RNA molecule of claim 60 or 61, wherein the polyA sequence comprises 30 adenine nucleotides followed by 70 adenine nucleotides, wherein the 30 adenine nucleotides and 70 adenine nucleotides are separated by a nucleotide linker sequence of 10 nucleotides.
63. The RNA molecule of any one of claims 60 to 62, wherein the polyA sequence comprises the nucleotide sequence of SEQ ID NO: 55, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO: 55.
64. The RNA molecule of any one of claims 60 to 63, wherein the RNA molecule comprises a 5'-cap, a 5'-UTR, a 3'-UTR and a polyA sequence.
65. The RNA molecule of any one of claims 1 to 64, wherein the RNA molecule comprises modified nucleotides, nucleosides or nucleobases.
66. The RNA molecule of claim 65, wherein the RNA molecule comprises modified uridines.
67. The RNA molecule of any one of claims 66, wherein the RNA molecule comprises modified uridines in place of all uridines.
68. The RNA molecule of claim 66 or 67, wherein the modified uridines are Nl-methyl-pseudouridine.
69. The RNA molecule of any one of claims 1 to 68, wherein the coding sequence of the RNA molecule is codon-optimized and / or is characterized in that its G / C content is increased compared to the parental sequence.
70. Chimeric protein encoded by the RNA molecule of any one of claims 1 to 69.
71. DNA molecule encoding the RNA molecule of any one of claims 1 to 69.
72. Pharmaceutical composition comprising one or more RNA molecules of any one of claims 1 to 69.
73. The pharmaceutical composition of claim 72, wherein the one or more RNA molecules are formulated in a lipid formulation, such as in lipid nanoparticles or liposomes.
74. The pharmaceutical composition of claim 73, wherein the lipid formulation comprises each of: e) a cationically ionizable lipid; f) a steroid; g) a neutral lipid; and h) a polymer-conjugated lipid.
75. The pharmaceutical composition of claim 74, wherein the cationically ionizable lipid is present in a concentration ranging from about 40 to about 60 mol percent of the total lipids.
76. The pharmaceutical composition of claim 74 or 75, wherein the steroid is present in a concentration ranging from about 30 to about 50 mol percent of the total lipids.
77. The pharmaceutical composition of any one of claims 74 to 76, wherein the neutral lipid is present in a concentration ranging from about 5 to about 15 mol percent of the total lipids.
78. The pharmaceutical composition of any one of claims 74 to 77, wherein the polymer-conjugated lipid is present in a concentration ranging from about 1 to about 10 mol percent of the total lipids.
79. The pharmaceutical composition of any one of claims 74 to 78, wherein the cationically ionizable lipid is within a range of about 40 to about 60 mole percent, the steroid is within a range of about 30 to about 50 mole percent, the neutral lipid is within a range of about 5 to about 15 mole percent, and the polymer-conjugated lipid is within a range of about 1 to about 10 mole percent.
80. The pharmaceutical composition of any one of claims 74 to 79, wherein the steroid comprises cholesterol.
81. The pharmaceutical composition of any one of claims 74 to 80, wherein the neutral lipid comprises a phospholipid.
82. The pharmaceutical composition of claim 81, wherein the phospholipid comprises distearoylphosphatidylcholine (DSPC).
83. The pharmaceutical composition of any one of claims 74 to 82, wherein the polymer-conjugated lipid comprises a polyethylene glycol (PEG)-lipid.
84. The pharmaceutical composition of any one of claims 72 to 83, wherein the pharmaceutical composition further comprises one or more pharmaceutically acceptable carriers, diluents and / or excipients.
85. The pharmaceutical composition of any one of claims 72 to 84, wherein the one or more RNA molecules are in a liquid formulation.
86. The pharmaceutical composition of any one of claims 72 to 84, wherein the one or more RNA molecules are in a frozen formulation.
87. The pharmaceutical composition of any one of claims 72 to 84, wherein the one or more RNA molecules are in a lyophilized formulation.
88. The pharmaceutical composition of any one of claims 72 to 87, wherein the one or more RNA molecules are formulated for injection.
89. The pharmaceutical composition of claim 88, wherein the one or more RNA molecules are formulated for intramuscular administration.
90. The pharmaceutical composition of any one of claims 72 to 89, therein the pharmaceutical composition is formulated for administration in human.
91. Kit comprising one or more pharmaceutical compositions of any one of claims 72 to 90.
92. The kit of claim 91, wherein two or more pharmaceutical compositions comprising the same or different RNA molecules according to any one of claims 1 to 69 are in separate vials.
93. The kit of claim 91 or 92, further comprising instructions for use of the one or more pharmaceutical composition for treating or preventing tuberculosis.
94. RNA molecule of any one of claims 1 to 69, chimeric protein of claim 70, DNA molecule of claim 71, pharmaceutical composition of any one of claims 72 to 90 or kit of any one of claims 91 to 93 for use as a medicament.
95. The RNA molecule, chimeric protein, DNA molecule, pharmaceutical composition or kit for use of claim 94, wherein the use comprises a therapeutic or prophylactic treatment of a disease or disorder in a subject.
96. The RNA molecule, chimeric protein, DNA molecule, pharmaceutical composition or kit for use of claim 94 or 95, wherein the use comprises the use as a vaccine against a disease or disorder in a subject.
97. The RNA molecule, chimeric protein, pharmaceutical composition or kit for use of claim 95 or 96, wherein the subject is a human infected with the disease or disorder or in danger of contracting the disease or disorder.
98. RNA molecule of any one of claims 1 to 69, chimeric protein of claim 70, DNA molecule of claim 71, pharmaceutical composition of any one of claims 72 to 90 or kit of any one of claims 91 to 93 for use in treating or preventing tuberculosis in a subject.
99. The RNA molecule, chimeric protein, pharmaceutical composition or kit for use of claim 98, wherein the subject is a human suffering from tuberculosis or in danger of contracting tuberculosis.
100. The RNA molecule, chimeric protein, pharmaceutical composition or kit for use of claim 98 or 99, wherein the use is as a vaccine for preventing tuberculosis.
101. Use of the RNA molecule of any one of claims 1 to 69, the chimeric protein of claim 70, the DNA molecule of claim 71, the pharmaceutical composition of any one of claims 72 to 90 or the kit of any one of claims 91 to 93 for the manufacture of a medicament for preventing or treating tuberculosis.
102. Method of vaccinating a subject comprising administering the RNA molecule of any one of claims 1 to 69, the chimeric protein of claim 70, the DNA molecule of claim 71, the pharmaceutical composition of any one of claims 72 to 90 or the kit of any one of claims 91 to 93 to the subject.
103. The method of claim 102, wherein the vaccination is for preventing tuberculosis.
104. The method of claim 102 or 103, wherein administration is by intramuscular administration.
105. The method of any one of claims 102 to 104, comprising administering to the subject at least one dose of the RNA, chimeric protein or pharmaceutical composition.
106. The method of any one of claims 102 to 105, comprising administering to the subject at least two doses of the RNA, chimeric protein or pharmaceutical composition.
107. The method of any one of claims 102 to 106, wherein an amount of the RNA of at least 10 pg per dose is administered.
108. The method of any one of claims 102 to 107, wherein the subject is a human.
109. The RNA molecule, chimeric protein, pharmaceutical composition or kit for use of any one of claims 98 to 100, the use of claim 101 or the method of any one of claims 102 to 108, wherein the tuberculosis is caused by an infection with a Mycobacterium.
110. The RNA molecule, chimeric protein, pharmaceutical composition or kit for use, the use or the method of claim 109, wherein the Mycobacterium is selected from the group of Mycobacterium tuberculosis, Mycobacterium bovis, Mycobacterium caprae, Mycobacterium orygis, Mycobacterium africanum, Mycobacterium microti, Mycobacterium canetti and Mycobacterium pinnipedii.
111. The RNA molecule, chimeric protein, pharmaceutical composition or kit for use, the use or the method of claim 110, wherein the Mycobacterium is Mycobacterium tuberculosis.