Engineered gcn4 trimerization motif

EP4713007A1Pending Publication Date: 2026-03-25MEVOX LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-15
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing vaccines face challenges due to the risk of autoimmune reactions caused by molecular mimicry between viral antigens and human proteins, particularly with the GCN4 trimerization motif, which shares sequence similarity with human proteins like C-Jun, potentially leading to adverse immune responses.

Method used

A de-humanised GCN4 trimerization motif (GCN4_M22) is designed to minimize sequence identity with human proteins, enhancing safety and stability, and is fused with vaccine antigens like the Mumps virus F protein to create immunogens that elicit immune responses without triggering autoimmune reactions.

Benefits of technology

The de-humanised GCN4 trimerization motif significantly reduces the risk of autoimmune reactions and doubles the yield of properly folded protein, providing improved vaccine safety and production efficiency for diseases involving the F protein of paramyxoviridae.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2024051265_21112024_PF_FP_ABST
    Figure GB2024051265_21112024_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to engineered GCN4 trimerization motifs, and to an immunogen comprising an engineered GCN4 trimerization motif. The invention also extends to virus-like particles and immunogenic compositions and vaccines comprising the immunogen, as well as their use in therapy and prophylaxis.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]Engineered GCN4 trimerization motif The present invention relates to engineered GCN4 trimerization motifs, and particularly, although not exclusively, to an immunogen comprising an engineered GCN4 trimerization motif. The invention also extends to virus-like particles and immunogenic compositions and vaccines comprising the immunogen, as well as their use in therapy and prophylaxis. Vaccination has led to huge benefits for population health by enabling the control of many devastating diseases. Fusion (F) glycoproteins are a well-established component of existing vaccines. For example, expression of the properly folded Mumps viral trimer requires the addition of a stabilising water-soluble engineered trimeric coiled-coil that mimics the F protein transmembrane region. However, evidence indicates that exposure to foreign antigens can potentially elicit autoimmune reactions. Autoimmune reactions result in morbidity or even death, and may be caused by infection or vaccination. Important categories of autoimmune reactions are: i) molecular mimicry; ii) bystander activation of autoreactive cells; and iii) activation of polyclonal T-cells with superantigens. Molecular mimicry arises from similarity between microorganism and human proteins, resulting in pseudo self-antigens being presented to T lymphocytes. There are numerous overlaps with pentamer and hexamer human peptides in infectious genomes, while heptamer and octamer overlaps occur, but at lower frequency. When infection carries a risk of inducing autoimmune disease, this risk is typically substantially greater than any risk associated with a corresponding vaccination. However, there is evidence for vaccine-driven autoimmunity in an extremely small proportion of individuals for particular vaccines, for example in Narcolepsy and Guillain-Barre syndrome. Other potential consequences of autoimmunity include demyelinating neuropathies, rheumatic heart disease and systemic lupus erythematosus. The general control transcription factor GCN4 protein is a parallel alpha-helical coiled coil ‘Leucine zipper’ motif. S. cerevisiae GCN4 is a transcription factor important in amino acid starvation. Stabilisation of trimeric protein quaternary structure by GCN4 is previously described in the public domain. For example, in the Protein DataBank (PDB, www.rcsb.org) accession numbers 6MJZ, 1GCM_A and 4WSG include a GCN4 stabilisation region (GCN4_init, Figure 1). 1GCM_A (Figure 1B) is a ‘GCN4 leucine zipper core mutant P-LI’ and the 1GCM structure shows a trimeric (iso)leucine zipper structure. 4WSG contains the GCN4_init sequence engaged in trimerization with the parainfluenza F glycoprotein (sequence alignment in Figure 1C). Parainfluenza is part of the paramyxovirus family that also contains Mumps. The GCN4 leucine zipper has been shown to be strongly immunogenic when fused with vaccine antigens. The 4WSG_1 structure represents the ‘prefusion’ conformation of the parainfluenza F glycoprotein, which functions to brings together the viral and host cell membranes to permit entry of the viral genome into the host cell. Crystallisation of the prefusion conformation of the influenza F protein was challenging due to spontaneous adoption of the postfusion conformation by the trimer expressed without the transmembrane region; however, removal of the hydrophobic membrane anchor is important for solubility of the expressed protein. Subsequently, Yin et al. stabilised the anchorless prefusion parainfluenza F protein by appending a designed GCN4 coiled-coil region to mimic the transmembrane region while retaining solubility, resulting in a trimeric structure. Therefore, the GCN4_init sequence was a key component required for crystallisation and structure determination of the trimeric parainfluenza F protein. Accordingly, the GCN4- mediated stabilisation of viral protein trimers is a potentially fruitful strategy in engineering self-assembling trimeric antigens for vaccine development. However, the GCN4_init construct has sequence similarity to human proteins, including the transcription factor C-Jun, resulting in potential implications for the safety of incorporating GCN4 into vaccines, such as autoimmune reactions. There is, therefore, the need for a GCN4 trimerization motif with enhanced safety, stability and solubility, which can be incorporated into vaccines. Accordingly, in a first aspect of the invention there is provided a de-humanised GCN4 trimerization motif. Advantageously, as shown in the examples, the inventors have designed a novel de-humanised GCN4 trimerization motif (GCN4_M22) that has a favourable safety profile and doubled yield relative to GCN4_init, i.e. the unmodified GCN4 sequence. In other words, this de-humanised GCN4 trimerization motif is optimised for both enhanced safety and increased protein production. The novel GCN4 coiled-coil trimerization motif will therefore allow for improved vaccine safety and production yield for multiple diseases, for example involving the F protein of other paramyxoviridae. The general control transcription factor GCN4 trimerization motif is a trimerization motif from the GCN4 protein that comprises a leucine zipper amino acid sequence that naturally forms an alpha-helical coiled-coil structure. Accordingly, in a preferred embodiment, the de-humanised GCN4 trimerization motif is a de-humanised GCN4 coiled-coil trimerization motif. The unmodified GCN4 sequence (also referred to herein as GCN4_init), may comprise an amino acid sequence of SEQ ID No: 4, which is provided herein, as follows: IEDKIEEILSKIYHIENEIARIKKLIGEAP [SEQ ID No: 4] As described above, however, this unmodified sequence comprises sequence similarity to human proteins, including the transcription factor C-Jun, resulting in potential implications for the safety of incorporating GCN4 into vaccines. Therefore, the inventors have engineered the GCN4 sequence to form a de-humanised GCN4 trimerization domain. As used herein, the term “de-humanised” will be understood to mean an engineered GCN4 amino acid sequence, which has been modified to eliminate sequence similarity with the human proteome. For example, in a preferred embodiment, the de-humanised GCN4 trimerization motif comprises less than 30%, less than 25%, less than 20%, less than 15% or less than 10% sequence identity with the human proteome. More preferably, the de-humanised GCN4 trimerization motif comprises less than 9%, less than 8%, less than 7%, less than 6%, or less than 5% sequence identity with the human proteome. Even more preferably, the de-humanised GCN4 trimerization motif comprises less than 4%, less than 3%, less than 2% or less than 1% sequence identity with the human proteome. Most preferably, the de-humanised GCN4 trimerization motif is not identical to a sequence found within the human proteome. In one embodiment, the de-humanised GCN4 trimerization motif may be referred to herein as GCN4_M22. The de-humanised GCN4 trimerization motif may comprise an amino acid sequence of SEQ ID No: 1, which is provided herein, as follows: INDRIEDILSRIYHIENNIARIQKLIGNAP [SEQ ID No: 1] Accordingly, in a preferred embodiment, the de-humanised GCN4 trimerization motif comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 1, or a fragment or variant thereof. Alternatively, in another embodiment, the de-humanised GCN4 trimerization motif may be referred to herein as GCN4_M23. The de-humanised GCN4 trimerization motif may comprise an amino acid sequence of SEQ ID No: 2, which is provided herein, as follows: INDRIEDIISRIYHIENNIARIQKIIGNAP [SEQ ID No: 2] Accordingly, in a preferred embodiment, the de-humanised GCN4 trimerization motif comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 2, or a fragment or variant thereof. Alternatively, in another embodiment, the de-humanised GCN4 trimerization motif may be referred to herein as GCN4_alt. The de-humanised GCN4 trimerization motif may comprise an amino acid sequence of SEQ ID No: 3, which is provided herein, as follows: INDRIEDILSRIYHIENNIARIRKLIGNAP [SEQ ID No: 3] Accordingly, in a preferred embodiment, the de-humanised GCN4 trimerization motif comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 3, or a fragment or variant thereof. In a preferred embodiment, the de-humanised GCN4 trimerization motif comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 1, 2 or 3, or a fragment or variant thereof, having at least 70%, 71%, 72%, 73%, 74%, or 75% sequence identity to SEQ ID No: 1, 2 or 3. Preferably, the de-humanised GCN4 trimerization motif comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 1, 2 or 3, or a fragment or variant thereof, having at least 76%, 77%, 78%, 79%, or 80% sequence identity to SEQ ID No: 1, 2 or 3. More preferably, the de-humanised GCN4 trimerization motif comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 1, 2 or 3, or a fragment or variant thereof, having at least 81%, 82%, 83%, 84%, or 85% sequence identity to SEQ ID No: 1, 2 or 3. More preferably, the de-humanised GCN4 trimerization motif comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 1, 2 or 3, or a fragment or variant thereof, having at least 86%, 87%, 88%, 89%, or 90% sequence identity to SEQ ID No: 1, 2 or 3. More preferably, the de-humanised GCN4 trimerization motif comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 1, 2 or 3, or a fragment or variant thereof, having at least 91%, 92%, 93%, 94%, or 95% sequence identity to SEQ ID No: 1, 2 or 3. More preferably, the de-humanised GCN4 trimerization motif comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 1, 2 or 3, or a fragment or variant thereof, having at least 96%, 97%, 98%, or 99% sequence identity to SEQ ID No: 1, 2 or 3. In a preferred embodiment, however, the de-humanised GCN4 trimerization motif consists of an amino acid sequence substantially as set out in SEQ ID No: 1, 2 or 3. Most preferably, the de-humanised GCN4 trimerization motif consists of an amino acid sequence substantially as set out in SEQ ID No: 1. As demonstrated in the Examples, the inventors have also shown that this novel de-humanised GCN4 trimerization motif is optimised for stability and increased protein production. The inventors avoided alteration of the hydrophobic core residues of the GCN4 trimerization motif in order to help preserve the trimeric hydrophobic packing. Accordingly, in order to promote stability, the inventors sought to: avoid charge repulsion between structurally proximal side chains, introduce helix-perpendicular and between-monomer salt bridges, enhance solubility, and balance overall charge. Protease sites were also avoided, in order to prevent cleavage of the engineered protein. As illustrated in Figures 8 and 9, the inter-helix contacts for the unmodified GCN4 trimerization motif (GCN4_init) and the de-humanised GCN4 trimerization motif (GCN4_M22) involve hydrophobic Van der Waals (VdW) interactions, electrostatic interactions and a Ser-Asn hydrogen bond. The GCN4 trimerization motif forms a trimeric coiled-coil structure, i.e. a three-helix bundle. The unmodified GCN4 trimerization motif (GCN4_init) comprises hydrogen bond, Van der Waals (VdW) / hydrophobic packing and electrostatic interactions (Lys-Glu) between two pairs of helices, however, one pair of helices have only VdW / hydrophobic interactions. Additionally, the electrostatic interactions between the two pairs of helices of the unmodified GCN4 sequence (GCN4_init) are Lys-Glu. In contrast, the de-humanised GCN4 trimerization motif (GCN4_M22) of the invention comprises electrostatic interactions (Arg-Glu) between all three pairs of helices. Additionally, these electrostatic interactions are Arg-Glu interactions. Advantageously, Arg binds to negatively charged amino acids more strongly than Lys, resulting in a stronger interaction with Glu. Moreover, Arg has a larger positive surface area for salt bridging and is more frequently represented in thermophilic organisms. Advantageously, therefore, these differences result in greater thermodynamic stability for the de-humanised GCN4 trimerization motif (GCN4_M22) relative to the unmodified GCN4 motif (GCN4_init). Accordingly, in a preferred embodiment, the de-humanised GCN4 trimerization motif comprises electrostatic interactions between all three pairs of helices. Preferably, the electrostatic interactions between all three pairs of helices are Arg- Glu interactions. In particular, three key inter-helix interactions in GCN4_M22 form a pseudo-ring in the structure, helping to confer stability. Specifically, these residue pairs are Arg22- Glu101, Arg96-Glu64 and Arg52-Glu17. Accordingly, in a preferred embodiment, the electrostatic interactions between all three pairs of helices of the de-humanised GCN4 trimerization motif comprise interactions between residue pairs Arg22-Glu101, Arg96-Glu64 and / or Arg52- Glu17. Most preferably, the electrostatic interactions between all three pairs of helices of the de-humanised GCN4 trimerization motif comprise interactions between residue pairs Arg22-Glu101, Arg96-Glu64 and Arg52-Glu17. The GCN4 trimerization motif comprises a leucine zipper amino acid sequence that naturally forms a trimeric structure. Accordingly, the GCN4 trimerization motif can be fused to soluble proteins that depend on trimerization for their therapeutic activity or proper antigenic and immunogenic structure. Thus, in a second aspect, there is provided an immunogen comprising the de- humanised GCN4 trimerization motif according to the first aspect. It will be well understood by the skilled person that an immunogen is a compound, composition, or substance that can elicit an immune response in an animal, including compositions that are injected or absorbed into an animal. Administration of an immunogen to a subject can lead to protective immunity against a pathogen of interest. The de-humanised GCN4 coiled-coil trimerization motif according to the invention will allow for improved vaccine safety and production yield for multiple diseases. Accordingly, in one embodiment, the immunogen is one which elicits an immune response against a virus. Preferably, the immunogen is one which elicits an immune response against a virus of the Adenoviridae family, Anelloviridae family, Arenaviridae family, Astroviridae family, Bornaviridae family, Bunyaviridae family, Caliciviridae family, Coronaviridae family, Filoviridae family, Flaviviridae family, Hepadnaviridae family, Hepeviridae family, Herpesviridae family, Orthomyxoviridae family, Papillomaviridae family, Paramyxoviridae family, Parvoviridae family, Picobirnaviridae family, Picobirna family, Picornaviridae family, Pneumoviridae family, Polyomaviridae family, Poxviridae family, Reoviridae family, Retroviridae family, Rhaboviridae family, Togaviridae family, or Delta family. Alternatively, the de-humanised GCN4 coiled-coil trimerization motif according to the invention may be incorporated in an immunogen that acts against other microorganisms, such as bacteria, where there are candidate antigens that may be presented in trimeric form. These include native trimers or possible synthetic trimers where the sequence has been modified to allow presentation of one or more antigen(s). Accordingly, in one embodiment, the immunogen is one which elicits an immune response against a bacterial infection. Preferably, the de-humanised GCN4 coiled-coil trimerization motif according to the invention will allow for improved vaccine safety and production yield for viruses involving the F protein of paramyxoviridae. Accordingly, in one embodiment, the immunogen is one which elicits an immune response against a virus of the Paramyxoviridae family. Preferably, the immunogen elicits an immune response against a virus selected from the group consisting of: measles virus, mumps virus, parainfluenza virus, Nipah virus, Morbillivirus canine, Rinderpest morbillis virus and respiratory syncytial virus (RSV). Preferably, therefore, the immunogen comprises an F protein selected from the group consisting of: a Measles virus F protein, a Mumps virus F protein, a parainfluenza virus F protein, a Nipah virus F protein, a Morbillivirus canine F protein, a Rinderpest morbillis virus F protein and a respiratory syncytial virus (RSV) F protein. It will be understood by the skilled person that an F protein is an envelope glycoprotein that facilitates fusion of viral and cellular membranes. In a preferred embodiment, the immunogen elicits an immune response against Mumps virus (MuV). Thus, preferably, the immunogen comprises a Mumps virus F protein. Mumps virus (MuV) is a non-segmented, negative-stranded RNA virus of the family Paramyxoviridae, subfamily Paramyxovirinae, genus Rubulavirus that causes mumps disease. Mumps virus genomic RNA contains seven tandemly linked transcription units that encode open reading frames for the nucleoprotein (N), phosphoprotein (P), V protein, I protein, matrix (M) protein, fusion (F) protein, small hydrophobic (SH) protein, hemagglutinin-neuraminidase (HN) protein, and the large (L) protein. Due to RNA editing by insertion of guanine nucleotides, the P gene (also referred to as the “V / P / I gene”) results in three mRNA transcripts corresponding to the V, P and I proteins. Specifically, faithful transcription of the P gene produces the V protein, insertion of two guanine nucleotides produces an mRNA encoding the P protein, and insertion of four guanine residues results in an mRNA encoding the I protein. The SH gene is the most variable gene amongst different genotypes of MuV and is therefore generally used as the basis for genotyping. There are 12 known genotypes of MuV, designated as genotypes A, B, C, D, F, G, H, I, J, K, L and N, that are currently circulating globally. Most current MuV vaccines are based on genotype A (Jeryl Lynn), genotype B (Urabe-AM9) or undetermined genotype (Leningrad-Zagreb) viruses. The MuV fusion (F) protein is an envelope glycoprotein of MuV that facilitates fusion of viral and cellular membranes. In nature, the F protein from MuV is initially synthesized as a single polypeptide precursor approximately 538 amino acids in length, designated F0. F0 includes an N-terminal signal peptide that directs localisation to the endoplasmic reticulum, where the signal peptide is proteolytically cleaved. The remaining F0 residues oligomerize to form a trimer and may be proteolytically processed by a cellular protease to generate two disulphide-linked fragments, F1 and F2. In MuV F the cleavage site is located approximately between residues 103 / 104. The smaller of these fragments, F2, originates from the N- terminal portion of the F0 precursor (approximately residues 20-103). The larger of these fragments, F1, includes the C-terminal portion of the F0 precursor (approximately residues 104-538) including an extracellular / luminal region (approximately residues 110-483), and a transmembrane and cytosolic region (approximately residues 484-538). The extracellular portion of the MuV F protein is the MuV F ectodomain, which includes the F2 protein and the F1 ectodomain. Three MuV F protomers oligomerize in the mature F protein, which adopts a metastable prefusion conformation that is triggered to undergo a conformational change to a postfusion conformation upon contact with a target cell membrane. This conformational change exposes a hydrophobic sequence, known as the fusion peptide, which is located at the N-terminus of the F1 ectodomain, and which associates with the host cell membrane and promotes fusion of the membrane of the virus, or an infected cell, with the target cell membrane. Accordingly, in a preferred embodiment, the immunogen comprises a Mumps virus (MuV) Fusion ectodomain trimer (F protein). Preferably, the MuV F ectodomain trimer is fused N-terminally to the de-humanised GCN4 trimerization motif. In one embodiment, the C-terminus of the MuV F ectodomain trimer may comprise an unmodified amino acid sequence. As used herein, the C-terminus of the MuV F ectodomain trimer refers to the last 12 amino acids of the MuV F ectodomain trimer sequence. Accordingly, the C-terminus of the MuV F ectodomain trimer may comprise an amino acid sequence of SEQ ID No: 5, which is provided herein, as follows: YIKESNHDQLQS [SEQ ID No: 5] Accordingly, in one embodiment, the C-terminal residue of the MuV F ectodomain trimer comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 5, or a fragment or variant thereof. However, the inventors’ analysis also led to modification of the Mumps Fusion ectodomain trimer (F protein) in proximity to the GCN4 region, to eliminate any potential cross-reactive epitopes arising from a hybrid of GCN4 and the neighbouring sequence. Accordingly, in one preferred embodiment, the C-terminus of the MuV F ectodomain trimer may comprise an amino acid sequence of SEQ ID No: 6, which is provided herein, as follows: YIKNSNHDLDS [SEQ ID No: 6] Accordingly, in a preferred embodiment, the C-terminus of the MuV F ectodomain trimer comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 6, or a fragment or variant thereof. Alternatively, in another embodiment, the C-terminus of the MuV F ectodomain trimer may comprise an amino acid sequence of SEQ ID No: 7, which is provided herein, as follows: YIKNSNHDIDS [SEQ ID No: 7] Accordingly, in a preferred embodiment, the C-terminus of the MuV F ectodomain trimer comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 7, or a fragment or variant thereof. In one embodiment, the MuV F ectodomain trimer may comprise an amino acid sequence of SEQ ID No: 8, which is provided herein, as follows: MKVSLVTCLGFAVFSFSICVNINILQQIGYIKQQVRQLSYYSQSSSSYIVVKLLPNIQPTDNSCEFKSVT QYNKTLSNLLLPIAENINNIASPSPGSRRHGGGAGIAIGIAALGVATAAQVTAAVSLVQAQTNARAIAAM KNSIQATNRAVFEVKEGTQQLAIAVQAIQNHINTIMNTQLNNMSCQILDNQLATSLGLYLTELTTCFQPQ LINPALSPISIQCLRSLLGSMTPAVVQATLSTSISAAEILSAGLMEGQIVSVLLDEMQMIVKINIPTIVT QSNALVIDFYSISSFINGQESIIQLPDRILEIGNEQWSYPAKNCKLTRHNIFCQYNEAERLSLESKLCLA GNISACVFSPIAGSYMRRFVALDGTIVANCRSLTCLCKSPSYPIYQPDHHAVTTIDLTACQTLSLDGLDF SIVSLSNITYAENLTISLSQTINTQPIDISTELIKVNASLQNAVKYIKNSNHDLDS [SEQ ID No: 8] Accordingly, in a preferred embodiment, the MuV F ectodomain trimer comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 8, or a fragment or variant thereof. In another embodiment, the MuV F ectodomain trimer may comprise an amino acid sequence of SEQ ID No: 105, which is provided herein, as follows: MKVSLVTCLGFAVFSFSICVNINILQQIGYIKQQVRQLSYYSQSSSSYIVVKLLPNIQPTDNSCEFKSVT QYNKTLSNLLLPIAENINNIASPSPGSRRHGGGAGIAIGIAALGVATAAQVTAAVSLVQAQTNARAIAAM KNSIQATNRAVFEVKEGTQQLAIAVQAIQNHINTIMNTQLNNMSCQILDNQLATSLGLYLTELTTCFQPQ LINPALSPISIQCLRSLLGSMTPAVVQATLSTSISAAEILSAGLMEGQIVSVLLDEMQMIVKINIPTIVT QSNALVIDFYSISSFINGQESIIQLPDRILEIGNEQWSYPAKNCKLTRHNIFCQYNEAERLSLESKLCLA GNISACVFSPIAGSYMRRFVALDGTIVANCRSLTCLCKSPSYPIYQPDHHAVTTIDLTACQTLSLDGLDF SIVSLSNITYAENLTISLSQTINTQPIDISTELIKVNASLQNAVKYIKNSNHDIDS [SEQ ID No: 105] Accordingly, in a preferred embodiment, the MuV F ectodomain trimer comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 105, or a fragment or variant thereof. In some preferred embodiments, the immunogen further comprises a MuV HN ectodomain. Preferably, the MuV HN ectodomain is fused C-terminally to the de- humanised GCN4 trimerization motif. MuV hemagglutinin-neuraminidase (HN) protein is an MuV envelope glycoprotein that is a type II membrane protein and facilitates attachment of MuV to host cell membranes. The full-length MuV HN protein has an N-terminal cytoplasmic tail and transmembrane domain (CT and TM, approximately amino acids 1-53), and an ectodomain (approximately amino acids 54-582) including stalk (approximately amino acids 54-130) and head regions (approximately amino acids 131-582). In one embodiment, the MuV HN ectodomain may comprise an amino acid sequence of SEQ ID No: 9, which is provided herein, as follows: NIPLVNDLRFINGINKFIIEDYATHDFSIGHPLNMPSFIPTATSPNGCTRIPSFSLGKTHWCYTHNVINA NCKDHTSSNQYVSMGILVQTASGYPMFKTLKIQYLSDGLNRKSCSIATVPDGCAMYCYVSTQLETDDYAG SSPPTQKLTLLFYNDTVTERTISPSGLEGNWATLVPGVGSGIYFENKLIFPAYGGVLPNSTLGVKSAREF FRPVNPYNPCSGPQQDLDQRALRSYFPSYFSNRRIQSAFLVCAWNQILVTNCELVVPSSNQTMMGAEGRV LLINNRLLYYQRSTSWWPYELLYEISFTFTNSGPSSVNMSWIPIYSFTRPGSGNCSGENVCPTACVSGVY LDPWPLTPYSHQSGINRNFYFTGALLNSSTTRVNPTLYVSALNNLKVLAPYGTQGLFASYTTTTCFQDTG DASVYCVYIMELASNIVGEFQILPVLTRLTIT [SEQ ID No: 9] Accordingly, in a preferred embodiment, the MuV HN ectodomain comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 9, or a fragment or variant thereof. In some preferred embodiments, the immunogen further comprises a linker region. Preferably, the linker fuses together the de-humanised GCN4 trimerization motif and the MuV HN ectodomain. As such, preferably the linker is disposed in between the de-humanised GCN4 trimerization motif and the MuV HN ectodomain. In one embodiment, the linker region may comprise an unmodified amino acid sequence. Accordingly, in one embodiment, the linker region may comprise an amino acid sequence of SEQ ID No: 10, which is provided herein, as follows: GSGGGSGG [SEQ ID No: 10] Accordingly, in one embodiment, the linker region comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 10, or a fragment or variant thereof. However, the inventors’ analysis also led to modification of the linker sequence in proximity to the GCN4 region, to eliminate any potential cross-reactive epitopes arising from a hybrid of GCN4 and the neighbouring sequence. Accordingly, in one preferred embodiment, the linker region may comprise an amino acid sequence of SEQ ID No: 11, which is provided herein, as follows: GNAAGSGG [SEQ ID No: 11] Accordingly, in one preferred embodiment, the linker region comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 11, or a fragment or variant thereof. Accordingly, in a preferred embodiment, the immunogen may comprise an amino acid sequence of SEQ ID No: 12, which is provided herein, as follows: MEFWLSWVFLVAILKGVQCVNINILQQIGYIKQQVRQLSYYSQSSSSYIVVKLLPNIQPTDNSCEFKSVT QYNKTLSNLLLPIAENINNIASPSPGSRRHGGGAGIAIGIAALGVATAAQVTAAVSLVQAQTNARAIAAM KNSIQATNRAVFEVKEGTQQLAIAVQAIQNHINTIMNTQLNNMSCQILDNQLATSLGLYLTELTTCFQPQ LINPALSPISIQCLRSLLGSMTPAVVQATLSTSISAAEILSAGLMEGQIVSVLLDEMQMIVKINIPTIVT QSNALVIDFYSISSFINGQESIIQLPDRILEIGNEQWSYPAKNCKLTRHNIFCQYNEAERLSLESKLCLA GNISACVFSPIAGSYMRRFVALDGTIVANCRSLTCLCKSPSYPIYQPDHHAVTTIDLTACQTLSLDGLDF SIVSLSNITYAENLTISLSQTINTQPIDISTELIKVNASLQNAVKYIKNSNHDLDSINDRIEDILSRIYH IENNIARIQKLIGNAPGNAAGSGGNIPLVNDLRFINGINKFIIEDYATHDFSIGHPLNMPSFIPTATSPN GCTRIPSFSLGKTHWCYTHNVINANCKDHTSSNQYVSMGILVQTASGYPMFKTLKIQYLSDGLNRKSCSI ATVPDGCAMYCYVSTQLETDDYAGSSPPTQKLTLLFYNDTVTERTISPSGLEGNWATLVPGVGSGIYFEN KLIFPAYGGVLPNSTLGVKSAREFFRPVNPYNPCSGPQQDLDQRALRSYFPSYFSNRRIQSAFLVCAWNQ ILVTNCELVVPSSNQTMMGAEGRVLLINNRLLYYQRSTSWWPYELLYEISFTFTNSGPSSVNMSWIPIYS FTRPGSGNCSGENVCPTACVSGVYLDPWPLTPYSHQSGINRNFYFTGALLNSSTTRVNPTLYVSALNNLK VLAPYGTQGLFASYTTTTCFQDTGDASVYCVYIMELASNIVGEFQILPVLTRLTIT [SEQ ID No: 12] Accordingly, in one preferred embodiment, the immunogen comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 12, or a fragment or variant thereof. In one embodiment, the GCN4_M23 may comprise an amino acid sequence of SEQ ID No: MEFWLSWVFLVAILKGVQCVNINILQQIGYIKQQVRQLSYYSQSSSSYIVVKLLPNIQPTDNSCEFKSVT QYNKTLSNLLLPIAENINNIASPSPGSRRHGGGAGIAIGIAALGVATAAQVTAAVSLVQAQTNARAIAAM KNSIQATNRAVFEVKEGTQQLAIAVQAIQNHINTIMNTQLNNMSCQILDNQLATSLGLYLTELTTCFQPQ LINPALSPISIQCLRSLLGSMTPAVVQATLSTSISAAEILSAGLMEGQIVSVLLDEMQMIVKINIPTIVT QSNALVIDFYSISSFINGQESIIQLPDRILEIGNEQWSYPAKNCKLTRHNIFCQYNEAERLSLESKLCLA GNISACVFSPIAGSYMRRFVALDGTIVANCRSLTCLCKSPSYPIYQPDHHAVTTIDLTACQTLSLDGLDF SIVSLSNITYAENLTISLSQTINTQPIDISTELIKVNASLQNAVKYIKNSNHDIDSINDRIEDIISRIYH IENNIARIQKIIGNAPGNAAGSGGNIPLVNDLRFINGINKFIIEDYATHDFSIGHPLNMPSFIPTATSPN GCTRIPSFSLGKTHWCYTHNVINANCKDHTSSNQYVSMGILVQTASGYPMFKTLKIQYLSDGLNRKSCSI ATVPDGCAMYCYVSTQLETDDYAGSSPPTQKLTLLFYNDTVTERTISPSGLEGNWATLVPGVGSGIYFEN KLIFPAYGGVLPNSTLGVKSAREFFRPVNPYNPCSGPQQDLDQRALRSYFPSYFSNRRIQSAFLVCAWNQ ILVTNCELVVPSSNQTMMGAEGRVLLINNRLLYYQRSTSWWPYELLYEISFTFTNSGPSSVNMSWIPIYS FTRPGSGNCSGENVCPTACVSGVYLDPWPLTPYSHQSGINRNFYFTGALLNSSTTRVNPTLYVSALNNLK VLAPYGTQGLFASYTTTTCFQDTGDASVYCVYIMELASNIVGEFQILPVLTRLTIT [SEQ ID No: 13] In one embodiment, the GCN4_alt may comprise an amino acid sequence of SEQ ID No 14: MEFWLSWVFLVAILKGVQCVNINILQQIGYIKQQVRQLSYYSQSSSSYIVVKLLPNIQPTDNSCEFKSVT QYNKTLSNLLLPIAENINNIASPSPGSRRHGGGAGIAIGIAALGVATAAQVTAAVSLVQAQTNARAIAAM KNSIQATNRAVFEVKEGTQQLAIAVQAIQNHINTIMNTQLNNMSCQILDNQLATSLGLYLTELTTCFQPQ LINPALSPISIQCLRSLLGSMTPAVVQATLSTSISAAEILSAGLMEGQIVSVLLDEMQMIVKINIPTIVT QSNALVIDFYSISSFINGQESIIQLPDRILEIGNEQWSYPAKNCKLTRHNIFCQYNEAERLSLESKLCLA GNISACVFSPIAGSYMRRFVALDGTIVANCRSLTCLCKSPSYPIYQPDHHAVTTIDLTACQTLSLDGLDF SIVSLSNITYAENLTISLSQTINTQPIDISTELIKVNASLQNAVKYIKNSNHDLDSINDRIEDILSRIYH IENNIARIRKLIGNAPGNAAGSGGNIPLVNDLRFINGINKFIIEDYATHDFSIGHPLNMPSFIPTATSPN GCTRIPSFSLGKTHWCYTHNVINANCKDHTSSNQYVSMGILVQTASGYPMFKTLKIQYLSDGLNRKSCSI ATVPDGCAMYCYVSTQLETDDYAGSSPPTQKLTLLFYNDTVTERTISPSGLEGNWATLVPGVGSGIYFEN KLIFPAYGGVLPNSTLGVKSAREFFRPVNPYNPCSGPQQDLDQRALRSYFPSYFSNRRIQSAFLVCAWNQ ILVTNCELVVPSSNQTMMGAEGRVLLINNRLLYYQRSTSWWPYELLYEISFTFTNSGPSSVNMSWIPIYS FTRPGSGNCSGENVCPTACVSGVYLDPWPLTPYSHQSGINRNFYFTGALLNSSTTRVNPTLYVSALNNLK VLAPYGTQGLFASYTTTTCFQDTGDASVYCVYIMELASNIVGEFQILPVLTRLTIT [SEQ ID No: 14] In a third aspect, there is provided a nucleic acid encoding the de-humanised GCN4 trimerization motif according to the first aspect or the immunogen according to the second aspect. In one embodiment, the de-humanised GCN4 trimerization motif is encoded by the nucleotide sequence of SEQ ID No: 103, as follows: ATCAACGACAGAATTGAGGACATCTTGAGCCGGATCTACCACATCGAAAACAACATCGCCAGGATCCAGA AGCTGATTGGCAACGCCCCC [SEQ ID No: 103] Accordingly, preferably the de-humanised GCN4 trimerization motif is encoded by the nucleotide sequence substantially as set out in SEQ ID No: 103, or a variant or fragment thereof. In one embodiment, the immunogen is encoded by the nucleotide sequence of SEQ ID No: 104, as follows: ATGGAGTTCTGGCTGAGCTGGGTGTTCCTTGTGGCCATCCTCAAGGGGGTTCAGTGTGTGAATATCAACA TCTTGCAGCAGATCGGATACATCAAGCAGCAAGTGCGCCAGCTCTCCTACTACAGCCAATCCTCCTCTTC GTACATTGTCGTCAAGCTTCTGCCTAATATCCAGCCCACCGACAACTCCTGCGAATTTAAGAGCGTCACT CAGTACAACAAGACCCTCTCAAACTTGCTCCTGCCCATTGCCGAGAACATCAACAACATTGCCTCGCCGT CCCCGGGCTCGAGGAGACATGGAGGCGGAGCCGGGATTGCCATCGGGATTGCCGCCCTGGGAGTGGCTAC CGCCGCCCAAGTGACCGCGGCCGTGAGCCTGGTGCAAGCCCAAACTAACGCTCGGGCCATTGCTGCCATG AAGAACTCAATCCAGGCCACCAACCGGGCAGTCTTTGAGGTGAAGGAGGGAACCCAGCAACTCGCTATCG CAGTCCAGGCGATCCAGAACCACATCAACACCATCATGAACACGCAGCTGAATAACATGTCATGCCAGAT TCTGGATAACCAGCTGGCCACTTCCTTGGGCCTCTACCTGACTGAACTAACCACATGCTTCCAGCCGCAA CTGATTAACCCTGCACTGAGCCCGATCTCCATCCAATGCTTGCGCTCGCTGCTCGGATCCATGACACCCG CCGTGGTGCAGGCAACCCTGTCAACCTCGATCTCCGCCGCAGAAATTCTCTCCGCTGGACTGATGGAGGG CCAGATTGTCTCCGTCCTGTTGGACGAGATGCAGATGATCGTGAAGATCAACATCCCTACTATCGTCACT CAGTCCAACGCACTGGTGATTGATTTCTACTCAATCTCCTCGTTCATCAACGGCCAGGAGTCAATCATCC AACTGCCTGATCGCATCCTCGAGATAGGAAATGAGCAGTGGTCCTACCCCGCTAAGAACTGCAAGCTGAC TCGGCATAACATCTTCTGCCAGTATAACGAAGCCGAGAGACTGTCCCTCGAATCCAAGCTCTGCCTGGCC GGAAACATCTCCGCCTGCGTGTTTAGCCCCATCGCCGGCAGCTACATGCGCCGGTTCGTGGCGCTTGACG GCACTATCGTCGCCAACTGCCGGTCACTGACTTGTCTGTGTAAGTCCCCTTCCTATCCAATCTACCAGCC CGACCATCATGCCGTGACTACCATCGATCTCACCGCCTGCCAAACGCTGTCCCTGGATGGGCTGGACTTC AGCATTGTGAGTCTCTCAAACATCACCTACGCAGAAAACCTTACTATTTCCCTGTCCCAAACCATCAATA CTCAACCCATTGATATCTCCACCGAACTGATTAAGGTCAACGCGTCCCTCCAGAACGCCGTCAAGTATAT CAAGAACTCCAACCACGACCTCGATAGCATCAACGACAGAATTGAGGACATCTTGAGCCGGATCTACCAC ATCGAAAACAACATCGCCAGGATCCAGAAGCTGATTGGCAACGCCCCCGGAAACGCAGCTGGTTCCGGCG GCAATATCCCTTTGGTGAACGACCTGAGATTCATCAACGGGATCAACAAGTTCATCATCGAGGATTACGC CACCCACGACTTCTCCATCGGGCACCCCCTGAATATGCCGTCCTTCATCCCAACTGCCACTTCGCCGAAC GGATGCACTCGGATCCCTTCCTTCTCGCTGGGAAAGACTCACTGGTGTTACACCCACAACGTGATCAACG CCAACTGCAAGGACCACACGTCAAGCAATCAGTACGTGTCGATGGGCATTCTGGTGCAGACCGCCTCCGG TTACCCCATGTTCAAGACCCTGAAGATTCAGTACCTGTCCGATGGGCTGAACCGGAAGTCCTGTTCCATC GCAACTGTGCCAGACGGATGCGCCATGTACTGCTATGTCAGCACCCAGCTGGAAACTGACGACTACGCCG GAAGTTCACCGCCGACCCAGAAGCTCACCCTCCTGTTCTACAACGATACTGTGACGGAAAGGACCATCTC CCCCTCCGGACTGGAAGGGAACTGGGCCACCTTGGTGCCCGGAGTGGGAAGCGGTATCTACTTCGAGAAC AAGCTGATCTTCCCTGCGTATGGCGGAGTGCTGCCGAACTCAACCCTCGGCGTGAAAAGCGCCAGGGAGT TCTTTCGCCCTGTGAACCCCTACAACCCATGCAGCGGTCCGCAGCAGGACCTCGACCAGCGCGCATTGCG CTCCTACTTCCCGAGCTACTTTAGCAATCGGCGGATCCAGTCAGCTTTTCTCGTCTGTGCGTGGAACCAG ATCCTGGTCACCAACTGCGAACTTGTGGTGCCGTCATCGAACCAGACAATGATGGGAGCGGAGGGACGAG TGCTGCTGATTAACAACAGACTCCTGTACTACCAGAGGAGCACCTCCTGGTGGCCGTACGAACTGCTGTA CGAGATCTCATTCACCTTCACTAACTCCGGTCCGTCCTCCGTCAACATGAGCTGGATTCCTATCTACTCG TTCACAAGACCCGGGTCCGGAAATTGCTCCGGAGAGAACGTCTGCCCCACTGCCTGCGTGTCCGGAGTGT ACCTCGACCCCTGGCCACTCACCCCCTACTCGCACCAGTCTGGAATCAACCGGAACTTCTACTTCACCGG TGCCCTGCTGAACTCGTCCACGACCCGCGTGAACCCCACCCTTTACGTGTCGGCCCTGAACAACCTCAAA GTGCTGGCCCCGTACGGAACTCAGGGGCTCTTTGCTTCCTACACCACCACCACGTGCTTCCAAGACACTG GCGATGCATCCGTGTACTGCGTGTACATCATGGAGCTTGCATCGAATATCGTCGGAGAGTTCCAGATCCT GCCTGTGCTGACCCGCCTGACCATAACCTGATAA [SEQ ID No: 104] Accordingly, preferably the immunogen is encoded by the nucleotide sequence substantially as set out in SEQ ID No: 104, or a variant or fragment thereof. A virus-like particle (VLP) may be decorated with the immunogen according to the second aspect, and used as a vaccine. Thus, in a fourth aspect, there is a provided a virus-like particle (VLP) comprising the immunogen according to the second aspect. Virus-like particles (VLPs) are non-replicating, viral shells, derived from any of several viruses. VLPs are generally composed of one or more viral proteins, such as, but not limited to, those proteins referred to as capsid, coat, shell, surface and / or envelope proteins, or particle-forming polypeptides derived from these proteins. VLPs can form spontaneously upon recombinant expression of the protein in an appropriate expression system. Methods for producing particular VLPs are known in the art. VLPs lack the viral components that are required for virus replication and thus represent a highly attenuated, replication-incompetent form of a virus. However, the VLP can display a polypeptide (e.g., a MuV ectodomain trimer) that is analogous to that expressed on infectious virus particles and can elicit an immune response to MuV when administered to a subject. Exemplary virus like particles and methods of their production, as well as viral proteins from several viruses that are known to form VLPs, including human papillomavirus, HIV, Semliki-Forest virus, human polyomavirus, rotavirus, parvovirus, canine parvovirus, hepatitis E virus, and Newcastle disease virus. In a fifth aspect of the invention, there is provided an immunogenic composition comprising an immunogen according to the second aspect or the VLP according to the fourth aspect, and a pharmaceutically acceptable vehicle. The term “immunogenic composition” as used throughout, refers to a composition of matter (intended to be administered to a subject) that comprises at least one immunogen or induces the expression of at least one immunogen (in the case of nucleic acid immunisation), which has the capability to elicit an immunological response in the subject to which it is administered. Such an immune response can be a cellular and / or antibody-mediated immune response directed at least against the antigen of the composition. Hence, the immunogenic composition can be referred to as a vaccine. The immunogenic composition may further comprise a suitable adjuvant. Adjuvants, such as aluminum hydroxide, Freund’s adjuvant, MPL ^ (3-O-deacylated monophosphoryl lipid A) and IL-12, TLR agonists (such as TLR-9 agonists, for example cytidine-phospho-guanosine oligodeoxynucleotide (CpG-ODN)1018), among many other suitable adjuvants well known in the art, can be included in the compositions. Suitable adjuvants are, for example, toll-like receptor agonists, alum, AlPO4, alhydrogel, Lipid-A and derivatives or variants thereof, oil-emulsions, saponins, neutral liposomes, liposomes containing the vaccine and cytokines, non- ionic block copolymers, and chemokines. Further examples of adjuvants may include an aluminium salt, a synthetic form of DNA, a carbohydrate, a tablet binder, an ion exchange resin, a preservative, a polymer, an emulsion and / or a lipid. Examples of adjuvants may include monosodium glutamate, sucrose, dextrose, aluminum bovine, human serum albumin, cytosine phosphoguanine, potassium phosphate, plasdone C, anhydrous lactose, cellulose, polacrilin potassium, glycerine, asparagine, citric acid, potassium phosphate magnesium sulfate, iron ammonium citrate, 2-phenoxyethanol, aluminium, beta-propiolactone, bovine extract, DOPC, EDTA, formaldehyde, thimerosal, phenol, potassium aluminum sulfate, potassium glutamate, sodium borate, sodium metabisulphite, urea, PLGA, PVA, PLA, PVP, cyclodextrin-based stabilisers, oil in water emulsion adjuvants and / or lipid-based adjuvants. The immunogen according to the second aspect, the virus-like particle according to the fourth aspect, and the immunogenic composition according to the fifth aspect, are particularly suitable for therapy or prophylaxis (i.e. vaccination). Hence, in a sixth aspect of the invention, there is provided the immunogen according to the second aspect, the virus-like particle according to the fourth aspect, or the immunogenic composition according to the fifth aspect, for use in therapy or prophylaxis. In a seventh aspect, there is provided the immunogen according to the second aspect, the virus-like particle according to the fourth aspect, or the immunogenic composition according to the fifth aspect, for use in eliciting an immune response. In an eighth aspect, there is provided a method of eliciting an immune response in a subject, the method comprising administering, to a subject in need thereof, a therapeutically effective amount of an immunogen according to the second aspect, the virus-like particle according to the fourth aspect, or the immunogenic composition according to the fifth aspect. It will be appreciated that the uses and methods of the invention may comprise vaccination. In a preferred embodiment, the immunogen, virus-like particle or immunogenic composition for use according to the seventh aspect, or the method according to the eighth aspect, elicits an immune response against a virus of the Adenoviridae family, Anelloviridae family, Arenaviridae family, Astroviridae family, Bornaviridae family, Bunyaviridae family, Caliciviridae family, Coronaviridae family, Filoviridae family, Flaviviridae family, Hepadnaviridae family, Hepeviridae family, Herpesviridae family, Orthomyxoviridae family, Papillomaviridae family, Paramyxoviridae family, Parvoviridae family, Picobirnaviridae family, Picobirna family, Picornaviridae family, Pneumoviridae family, Polyomaviridae family, Poxviridae family, Reoviridae family, Retroviridae family, Rhaboviridae family, Togaviridae family, or Delta family. Most preferably, the immunogen, virus-like particle or immunogenic composition for use according to the seventh aspect, or the method according to the eighth aspect, elicits an immune response against a virus of the Paramyxoviridae family. Preferably, the immunogen, virus-like particle or immunogenic composition for use according to the seventh aspect, or the method according to the eighth aspect, elicits an immune response against a virus selected from the group consisting of: measles virus, mumps virus, parainfluenza virus, Nipah virus, Morbillivirus canine, Rinderpest morbillis virus, respiratory syncytial virus (RSV). Alternatively, the immunogen, virus-like particle or immunogenic composition may elicit an immune response against other microorganisms, such as bacteria, where there are candidate antigens that may be presented in trimeric form. These include native trimers or possible synthetic trimers where the sequence has been modified to allow presentation of one or more antigen(s). Accordingly, in one embodiment, the immunogen, virus-like particle or immunogenic composition for use according to the seventh aspect, or the method according to the eighth aspect, elicits an immune response against a bacterial infection. Most preferably, the immunogen, virus-like particle or immunogenic composition for use according to the seventh aspect, or the method according to the eighth aspect, elicits an immune response against a Mumps viral infection. It will be appreciated that the immunogen or immunogenic composition according to the invention, may be used in a medicament, which may be used as a monotherapy (i.e. use of the immunogenic composition alone), for vaccination against an infection. Alternatively, the immunogen or immunogenic composition according to the invention may be used as an adjunct to, or in combination with, known therapies for treating, ameliorating, or preventing an infection. The immunogenic composition of the invention may be combined in compositions having a number of different forms depending, in particular, on the manner in which the composition is to be used. Thus, for example, the composition may be in the form of a powder, tablet, capsule, liquid, ointment, cream, gel, hydrogel, aerosol, spray, micellar solution, transdermal patch, liposome suspension, polyplex, emulsion, lipid nanoparticles (e.g. with peptide, DNA or RNA on the surface or encapsulated) or any other suitable form that may be administered to a person or animal in need of vaccination. The lipid nanoparticle may comprise one or more components selected from a group consisting of: a cationic lipid (which is preferably ionisable); phosphatidylcholine; cholesterol; and polyethylene glycol (PEG)-lipid. It will be appreciated that the vehicle of medicaments according to the invention should be one which is well-tolerated by the subject to whom it is given. Medicaments comprising the immunogenic composition of the invention may be used in a number of ways. For instance, oral administration may be required, in which case the agents may be contained within a composition that may, for example, be ingested orally in the form of a tablet, capsule or liquid. Compositions comprising agents and medicaments of the invention may be administered by inhalation (e.g. intranasally). Compositions may also be formulated for topical use. For instance, creams or ointments may be applied to the skin. The immunogenic composition of the invention may also be incorporated within a slow- or delayed-release device. Such devices may, for example, be inserted on or under the skin, and the medicament may be released over weeks or even months. The device may be located at least adjacent the treatment site. Such devices may be particularly advantageous when long-term treatment with the immunogenic composition is required and which would normally require frequent administration (e.g. at least daily injection). In a preferred embodiment, however, medicaments according to the invention may be administered to a subject by injection into the blood stream, muscle, skin or directly into a site requiring treatment. Injections may be intravenous (bolus or infusion) or subcutaneous (bolus or infusion), or intradermal (bolus or infusion), or intramuscular (bolus or infusion). It will be appreciated that the amount of immunogenic composition that is required is determined by its biological activity and bioavailability, which in turn depends on the mode of administration, the physiochemical properties of the immunogenic composition and whether it is being used as a monotherapy or in a combined therapy. The frequency of administration will also be influenced by the half-life of the active agent within the subject being treated. Optimal dosages to be administered may be determined by those skilled in the art, and will vary with the particular the immunogenic composition in use, the strength of the composition, the mode of administration, and the type and advancement of the viral infection. Additional factors depending on the particular subject being treated will result in a need to adjust dosages, including subject age, weight, gender, diet, and time of administration. Generally, a daily dose of between 0.001µg / kg of body weight and 100mg / kg of body weight of the immunogenic composition of the invention may be used for the immunisation, depending upon the agent used. More preferably, the daily dose of agent is between 1^g / kg of body weight and 100mg / kg of body weight, more preferably between 10^g / kg and ^^mg / kg body weight, and most preferably between approximately 100^g / kg and 10mg / kg body weight. Daily doses may be given as a single administration (e.g. a single daily injection or inhalation of a nasal spray). Alternatively, the immunogenic composition may require administration twice or more times during a day. As an example, the immunogenic composition may be administered as an initial primer and a subsequent boost(s), or two boosts administered at between a week or monthly intervals. Preferably, the immunogenic composition may be administered as an initial primer and a subsequent boost, administered between two to six weeks apart. Preferably, subsequent boosts may then be administered yearly to susceptible patients, such as those with weakened immune systems. Known procedures, such as those conventionally employed by the pharmaceutical industry (e.g. in vivo experimentation, clinical trials, etc.), may be used to form specific formulations of the immunogenic composition according to the invention and precise therapeutic regimes (such as daily doses of the agents and the frequency of administration). A “subject” may be a vertebrate, mammal, or domestic animal. Hence, compositions and medicaments according to the invention may be used to treat any mammal, for example livestock (e.g. a horse), pets, or may be used in other veterinary applications. Most preferably, however, the subject is a human being. A “therapeutically effective amount” of the immunogenic composition is any amount which, when administered to a subject, is the amount of the aforementioned that is needed to ameliorate, prevent or treat any given disease, preferably prophylactically. For example, the immunogenic composition of the invention may be used from about 0.001 µg to about 1 mg, and preferably from about 0.001 µg to about 500 µg. It is preferred that the amount of the immunogenic composition is an amount from about 0.01 µg to about 250 µg, and most preferably from about 0.1 µg to about 100 µg. Preferably, the immunogenic composition according to the invention is administered at a dose of 1-50^g. A “pharmaceutically acceptable vehicle” as referred to herein, is any known compound or combination of known compounds that are known to those skilled in the art to be useful in formulating pharmaceutical compositions. In one embodiment, the pharmaceutically acceptable vehicle may be a solid, and the composition may be in the form of a powder or tablet. A solid pharmaceutically acceptable vehicle may include one or more substances which may also act as flavouring agents, lubricants, solubilisers, suspending agents, dyes, fillers, glidants, compression aids, inert binders, sweeteners, preservatives, dyes, coatings, or tablet-disintegrating agents. The vehicle may also be an encapsulating material. In powders, the vehicle is a finely divided solid that is in admixture with the finely divided active agents according to the invention. In tablets, the active agent (e.g. immunogenic composition according to the invention) may be mixed with a vehicle having the necessary compression properties in suitable proportions and compacted in the shape and size desired. The powders and tablets preferably contain up to 99% of the active agents. Suitable solid vehicles include, for example calcium phosphate, magnesium stearate, talc, sugars, lactose, dextrin, starch, gelatin, cellulose, polyvinylpyrrolidine, low melting waxes and ion exchange resins. In another embodiment, the pharmaceutical vehicle may be a gel and the composition may be in the form of a cream or the like. However, the pharmaceutical vehicle may be a liquid, and the pharmaceutical composition is in the form of a solution. Liquid vehicles are used in preparing solutions, suspensions, emulsions, syrups, elixirs and pressurized compositions. The immunogenic composition according to the invention may be dissolved or suspended in a pharmaceutically acceptable liquid vehicle such as water, an organic solvent, a mixture of both or pharmaceutically acceptable oils or fats. The liquid vehicle can contain other suitable pharmaceutical additives such as solubilisers, emulsifiers, buffers, preservatives, sweeteners, flavouring agents, suspending agents, thickening agents, colours, viscosity regulators, stabilizers or osmo- regulators. Suitable examples of liquid vehicles for oral and parenteral administration include water (partially containing additives as above, e.g. cellulose derivatives, preferably sodium carboxymethyl cellulose solution), alcohols (including monohydric alcohols and polyhydric alcohols, e.g. glycols) and their derivatives, and oils (e.g. fractionated coconut oil and arachis oil). For parenteral administration, the vehicle can also be an oily ester such as ethyl oleate and isopropyl myristate. Sterile liquid vehicles are useful in sterile liquid form compositions for parenteral administration. The liquid vehicle for pressurized compositions can be a halogenated hydrocarbon or other pharmaceutically acceptable propellant. Liquid pharmaceutical compositions, which are sterile solutions or suspensions, can be utilized by, for example, subcutaneous, intradermal, intrathecal, epidural, intraperitoneal, intravenous and particularly intramuscular injection. The nucleic acid sequence, or expression cassette of the invention may be prepared as a sterile solid composition that may be dissolved or suspended at the time of administration using sterile water, saline, or other appropriate sterile injectable medium. The immunogenic composition of the invention may be administered orally in the form of a sterile solution or suspension containing other solutes or suspending agents (for example, enough saline or glucose to make the solution isotonic), bile salts, acacia, gelatin, sorbitan monoleate, polysorbate 80 (oleate esters of sorbitol and its anhydrides copolymerized with ethylene oxide) and the like. The immunogenic composition according to the invention can also be administered orally either in liquid or solid composition form. Compositions suitable for oral administration include solid forms, such as pills, capsules, granules, tablets, and powders, and liquid forms, such as solutions, syrups, elixirs, and suspensions. Forms useful for parenteral administration include sterile solutions, emulsions, and suspensions. It will be appreciated that the invention extends to any nucleic acid or peptide or variant, derivative or analogue thereof, which comprises substantially the amino acid or nucleic acid sequences of any of the sequences referred to herein, including variants or fragments thereof. The terms “substantially the amino acid / nucleotide / peptide sequence”, “variant” and “fragment”, can be a sequence that has at least 40% sequence identity with the amino acid / nucleotide / peptide sequences of any one of the sequences referred to herein, for example 40% identity with any of the sequence identified herein. Amino acid / polynucleotide / polypeptide sequences with a sequence identity which is greater than 65%, more preferably greater than 70%, even more preferably greater than 75%, and still more preferably greater than 80% sequence identity to any of the sequences referred to are also envisaged. Preferably, the amino acid / polynucleotide / polypeptide sequence has at least 85% identity with any of the sequences referred to, more preferably at least 90% identity, even more preferably at least 92% identity, even more preferably at least 95% identity, even more preferably at least 97% identity, even more preferably at least 98% identity and, most preferably at least 99% identity with any of the sequences referred to herein. The skilled technician will appreciate how to calculate the percentage identity between two amino acid / polynucleotide / polypeptide sequences. In order to calculate the percentage identity between two amino acid / polynucleotide / polypeptide sequences, an alignment of the two sequences must first be prepared, followed by calculation of the sequence identity value. The percentage identity for two sequences may take different values depending on:- (i) the method used to align the sequences, for example, ClustalW, BLAST, FASTA, Smith-Waterman (implemented in different programs), or structural alignment from 3D comparison; and (ii) the parameters used by the alignment method, for example, local vs global alignment, the pair-score matrix used (e.g. BLOSUM62, PAM250, Gonnet etc.), and gap-penalty, e.g. functional form and constants. Having made the alignment, there are many different ways of calculating percentage identity between the two sequences. For example, one may divide the number of identities by: (i) the length of shortest sequence; (ii) the length of alignment; (iii) the mean length of sequence; (iv) the number of non-gap positions; or (v) the number of equivalenced positions excluding overhangs. Furthermore, it will be appreciated that percentage identity is also strongly length dependent. Therefore, the shorter a pair of sequences is, the higher the sequence identity one may expect to occur by chance. Hence, it will be appreciated that the accurate alignment of protein or DNA sequences is a complex process. The popular multiple alignment program ClustalW (Thompson et al., 1994, Nucleic Acids Research, 22, 4673-4680; Thompson et al., 1997, Nucleic Acids Research, 24, 4876-4882) is a preferred way for generating multiple alignments of proteins or DNA in accordance with the invention. Suitable parameters for ClustalW may be as follows: For DNA alignments: Gap Open Penalty = 15.0, Gap Extension Penalty = 6.66, and Matrix = Identity. For protein alignments: Gap Open Penalty = 10.0, Gap Extension Penalty = 0.2, and Matrix = Gonnet. For DNA and Protein alignments: ENDGAP = -1, and GAPDIST = 4. Those skilled in the art will be aware that it may be necessary to vary these and other parameters for optimal sequence alignment. Preferably, calculation of percentage identities between two amino acid / polynucleotide / polypeptide sequences may then be calculated from such an alignment as (N / T)*100, where N is the number of positions at which the sequences share an identical residue, and T is the total number of positions compared including gaps and either including or excluding overhangs. Preferably, overhangs are included in the calculation. Hence, a most preferred method for calculating percentage identity between two sequences comprises (i) preparing a sequence alignment using the ClustalW program using a suitable set of parameters, for example, as set out above; and (ii) inserting the values of N and T into the following formula:- Sequence Identity = (N / T)*100. Alternative methods for identifying similar sequences will be known to those skilled in the art. For example, a substantially similar nucleotide sequence will be encoded by a sequence which hybridizes to DNA sequences or their complements under stringent conditions. By stringent conditions, the inventors mean the nucleotide hybridises to filter-bound DNA or RNA in 3x sodium chloride / sodium citrate (SSC) at approximately 45ºC followed by at least one wash in 0.2x SSC / 0.1% SDS at approximately 20-65ºC. Alternatively, a substantially similar polypeptide may differ by at least 1, but less than 5, 10, 20, 50 or 100 amino acids from any of the sequences described herein. Due to the degeneracy of the genetic code, it is clear that any nucleic acid sequence described herein could be varied or changed without substantially affecting the sequence of the protein encoded thereby, to provide a functional variant thereof. Suitable nucleotide variants are those having a sequence altered by the substitution of different codons that encode the same amino acid within the sequence, thus producing a silent (synonymous) change. Other suitable variants are those having homologous nucleotide sequences but comprising all, or portions of, sequence, which are altered by the substitution of different codons that encode an amino acid with a side chain of similar biophysical properties to the amino acid it substitutes, to produce a conservative change. For example, small non-polar, hydrophobic amino acids include glycine, alanine, leucine, isoleucine, valine, proline, and methionine. Large non-polar, hydrophobic amino acids include phenylalanine, tryptophan and tyrosine. The polar neutral amino acids include serine, threonine, cysteine, asparagine and glutamine. The positively charged (basic) amino acids include lysine, arginine and histidine. The negatively charged (acidic) amino acids include aspartic acid and glutamic acid. It will therefore be appreciated which amino acids may be replaced with an amino acid having similar biophysical properties, and the skilled technician will know the nucleotide sequences encoding these amino acids. All of the features described herein (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined with any of the above aspects in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. For a better understanding of the invention, and to show how embodiments of the same may be carried into effect, reference will now be made, by way of example, to the accompanying Figures, in which:- Figure 1 shows pairwise alignment for the unmodified GCN4 stabilisation region (GCN4_init) with records from the Protein DataBank. (A) There is an exact match to 6MJZ, although the residues are disordered in the crystal structure. A match to a trimeric sequence in 5EVM_A was identified which has electron density in the GCN4 trimerisation region. (B) The final two residues (alanine, proline) in GCN4_init do not match to 1GCM_A: the alanine is typically present as an arginine in ‘native’ GCN4 sequences, and the proline is not present. (C) There is a perfect match to the 4WSG_1 parainfluenza virus sequence, across the full length of GCN4_init. Figure 2 shows selected GCN4 designs. Provenance of the sequences shown is indicated by the brackets (bottom), specifically Mumps Fusion Glyoprotein C- terminal residues (‘Mumps F’); the coiled-coil GCN4 region engineered for prefusion conformation stabilisation (‘GCN4’); and part of the linker region (‘Linker’). GCN4_init corresponds to the unmodified sequence; regions of the alignment highlighted represent alterations to GCN4_init in the new designs. Figure 3 illustrates an overview of GCN4 engineering pipeline. The input protein sequence is analysed in two integrated streams: ‘Derisk’ and ‘Stabilization’. The ‘Derisk’ stream focuses upon elimination of sequence similarity between the input sequence and the human proteome; based upon sequence similarity and including screening against potential T Cell epitopes. ‘Stabilization’ considers the amino acid properties and protein structure, including assessment of stability by molecular dynamics. The resultant GCN4 sequence designs were taken as input for the next iteration of the engineering pipeline. In practice, candidate GCN4 sequence modifications resulted in new similarities with human protein sequences that were not returned by earlier searches – all of the sequence matches were considered together in order to ultimately produce derisked designs. Figure 4 shows multiple Sequences Alignment with GCN4_init, GCN4_M22 and sequences returned from Position Specific Iterative Basic Local Alignment and Search Tool (PSI-BLAST) searching. Amino acids are coloured according to the Taylor scheme and residue conservation may be observed, for example the hydrophobic core residues (Isoleucine, Leucine) and glutamic acid at positions 17 and 18. Figure 5 shows: A) Electron micrograph (top) and 3D structural model (bottom) of the Mumps F protein, GCN4 region, linker and Mumps HN protein. The trimerization-promoting GCN4 region is shown in yellow. B) The GCN4_init (SEQ ID No: 49) and GCN4_M22 / GCN4v6 (SEQ ID No: 50) sequences are shown with annotations to identify 7-mer and 8-mer matches with the human proteome (red boxes); these matches are eliminated in the GCN_M22 sequence. The text background colour for the sequences corresponds to Mumps F (purple), GCN4 (yellow), linker (grey) and Mumps HN (magenta). Modified residues in GCN4_M22 are shown in red, the relevance of modifications for elimination of human sequence similarity and protein stability is indicated by ‘*’ and ‘+’ respectively. Figure 6 illustrates GCN4_init peptides that potentially might cross-react with ANK3 in a human population. Results for the MHC allele HLA-B*18:01 are shown. The ‘BindLevel’ column shows NetMHCpan4.1 predicted binding strength as follows: ‘SB’ (strong binder), ‘WB’ (weak binder). The %Rank_EL defines predicted strong binders (<0.5%) and weak binders (<2%). The SB sequence EEILSKIY (SEQ ID No: 52) would involve an amino acid insertion at position 3, the SB sequence EEILSKIYH (SEQ ID No: 53) does not require an insertion but involves two ANK3 amino acid changes. The WB sequence EEILSKIYK (SEQ ID No: 55) would be produced by a single nucleotide change relative to the reference ANK3 sequence. Figure 7 shows population frequencies of HLA-B*18:01 in European populations. The frequency of HLA-B*18:01 is shown, covering up to 33.3% of individuals in some populations. Allele frequency values are expected to be approximately half of the proportion of individuals with the allele due to heterozygosity. Data provided by the Allele Frequency Net Database (AFND). Figure 8 shows mapping residue contacts for the GCN4 helices and modification for enhanced stability. Helical wheel representations (left) and LigPlot results (right) are shown for the GCN4_init (I) and GCN4_M22 (II) 3D models. Individual helices are labelled (I.i, I.ii, I.iii; II.i, II.ii II.iii). The inner Tyrosine (Y) is the N-terminus for all helices. Letter colour shows key biophysical properties: positive charge (blue), negative charge (red), hydrophobic (grey), and polar / other (orange). While the helical wheel representations have rotational symmetry, the interactions between helices in the 3D structural models are not symmetrical, which is reflected in the LigPlot results. Indeed, LigPlot identified hydrogen bond, Van der Waals (VdW) / hydrophobic packing and electrostatic interactions (Lys-Glu) between GCN4_init helices I.i, I.ii and between helices I.i, I.ii – however helices I.ii, I.iii had only VdW / hydrophobic interactions. GCN4_M22 had electrostatic interactions (Arg- Glu) between all three pairs of helices, while GCN4_init only had electrostatic interactions between two pairs of helices (I.i-I.iii and I.i-I.ii); these differences are expected to result in greater thermodynamic stability for the GCN4_M22 trimer. Figure 9 shows structural models for GCN4_init and GCN4_M22, highlighting electrostatic, H-bond and hydrophobic interactions. Homology models of the trimeric GCN4_init (I, top) and GCN4_M22 (II, bottom) are shown. The template structure was 1GCM. Helix colouring indicates regions of hydrophobic interactions for the pairs i-ii (purple), i-iii(orange) and ii-iii (blue); grey shows residues not involved in hydrophobic packing. Residues participating in electrostatic and hydrogen bond (H-bond) interactions are shown with stick and ball representation, also labelled in black text. Atoms are shown in standard colours (oxygen in red, nitrogen in blue, carbon in grey). Helix identifiers (I.i, I.ii etc.) are equivalent to those shown in Figure 6. Notably, helices I.ii and I.iii (GCN4_init) only have hydrophobic packing interactions, while helices II.ii and II.iii (GCN4_M22) have an electrostatic interaction between Glu64 (II.ii) and Arg96 (II.iii). Figure 10 shows molecular Dynamics of GCN4_init, GCN4_M22 and the homology modelling template (1GCM). Top: Two-dimensional projections of trajectories are shown. Each point represents a snapshot of the protein structure, 2D spatial proximity correlates with temporal proximity. The right-hand side of the trajectories for GCN4_init and GCN4_M22 corresponds with the end of the simulation. GCN4_init (red) adopts a greater diversity of conformations than GCN4_M22 (green). Middle: RMSD summarises protein-wide deviation of amino acid alpha- carbons from their reference positions, which are the atom positions at t=0 of the simulation. GCN4_init has greater RMSD than GCN4_M22. Bottom: RMS fluctuation shows the time-averaged deviation of the atoms across the protein sequence (referenced to t=0 positions), numbering starts at the N-terminus. GCN4_init has greater RMS fluctuation than GCN4_M22 at approximately atom positions 1550- 1600 and 2300-2400, involving the residues Asp101–Tyr103 and Gly146–Ile153. Figure 11 illustrates visualisation of RMS fluctuation for GCN4_init and GCN4_M22. Colour indicates the magnitude of fluctuation: blue (0-0.25Å), yellow (0.25-0.5Å), orange (0.5-1Å), red (>1Å). Van der Waals radii are shown for amino acids residues that have inter-helix hydrogen bonding or electrostatic interactions (CPK, space-filling), which are also labelled with black text. Overall, GCN4_init (I, top) has larger structural fluctuations than GCN4_M22 (II, bottom) at the N-terminus of helix i (yellow for GCN4_init, blue for GCN4_M22). Also, GCN4_init, GCN4_M22 are respectively red, orange at the mid-region of helix iii; showing greater instability for GCN4_init. GCN4_init Glu91 on helix I.iii shows substantial fluctuations (red) which may disrupt electrostatic interaction with Lys15 (blue) on helix I.i. Large fluctuations for GCN4_init are also observed for helix I.ii Ser42 (red), impacting hydrogen bonding with Asn6 (yellow). In contrast, the inventors found low fluctuations for the GCN_M22 salt bridge residues between helices II.iii (Glu101) and II.i (Arg22); Asn6 is also more stable in GCN_M22. Accordingly, the inventors’ modelling results predict lower mobility of key residues and regions in GCN4_M22 relative to GCN4_init, providing basis for stronger binding and therefore greater stability of the GCN4_M22 three helix bundle. Figure 12 shows expression of PreF-HN with GCN4_init and GCN4_M22. Elution profiles of Mumps PreF-HN with the GCN4_init (A) and GCN4_M22 (B) are shown. The y-axis shows ultraviolet absorbance at 280 nanometers; dark blue line with a peak around 53 minutes. The x-axis indicates elution time. The peak area for GCN4_M22 is approximately double that of the GCN4_init peak. Overall yields were respectively 0.11 mg / 100ml, 0.22 mg / 100ml for GCN4_init and GCN4_M22. Hence, the GCN4_M22 design produced a doubling in yield. Examples The inventors set out to develop an approach to derisk the unmodified GCN4 trimerization motif (GCN4_init) by modifying the amino acid residues in order to eliminate sequence similarity with the human proteome, including evaluation of predicted T-cell epitopes and potential mutations. Simultaneously, the inventors set out to enhance the coiled-coil stability of the GCN4 domain for greater yield. The stability of these new GCN4 designs was evaluated using parameters derived from homology modelling and molecular dynamics simulations. The selected designs are shown in Figure 2. Materials and Methods The inventors pre-emptively derisked the GCN4_init sequence in the context of surrounding residues from the Mumps F C-terminal region and part of the linker region (Figure 2). Sequence modifications removed sequence similarity with the human proteome and simultaneously enhanced protein stability. Analysis of the Mumps Fusion Glycoprotein (F protein) and linker sequence in proximity to the GCN4 region sought to eliminate any potential cross-reactive epitopes arising from a hybrid of GCN4 and the neighbouring sequence. Searches were carried out against the Ensembl human proteome with a) Protein BLAST (BLASTP), for sequence matching b) bespoke computer code developed by Overton for exhaustive k-mer (peptide) searching and c) PSI-BLAST to explore more distant evolutionary relationships. Sequence modifications and alignment against the human proteome were carried out for assessment and elimination of similarities to human proteins that might be introduced into the modified GCN4 sequences. The iterative modification process was also designed to eliminate potential autoimmune T-cell epitopes. In concert with derisking, modifications of the GCN4 sequence were designed in order to enhance or preserve key biophysical features driving coiled-coil stability. These included: leucine zipper hydrophobic packing, leucine zipper trimeric subunit stoichiometry, and amino acid side chain interactions including salt bridges. Multiple iterations of sequence modifications and database searching were made in order to maximise desirable sequence properties and remove human sequence similarity. This semi-automated bioinformatics protocol (Figure 3) analyses a) human protein matches to GCN4_init; b) production of desirable sequence properties, enhancing stability and yield; c) potential new matches arising with the engineered sequence and the human proteome; d) predicted T cell epitopes; e) predicted protein stability from homology modelling and molecular dynamics. The resultant sequence candidates were studied in vitro in order to ascertain biophysical properties, including yield. Sequence searching with Basic Local Alignment and Search Tool (BLAST) Protein BLAST (BLASTP) searches were applied to identify matches for the GCN4 sequence designs with human proteins. BLASTP was run with expectation value (e- value) of 1000 for sensitive detection of short sequence matches, using default parameters and the Ensembl human proteome as search database. A complementary search strategy deployed PSI-BLAST, run for three iterations with e-value threshold of 10-3and a final iteration taking e-value=10. The search database was NCBI NR. Other parameters were word size=3, BLOSUM62 matrix, gap cost=11, gap extension=1. Multiple Sequence Alignment (MSA) was performed with ClustalO and visualised in Jalview. K-mer screening Genetic variation at the population level influences potential autoimmune reactions, for example driving differences in the mechanisms that regulate immune responses to self-antigens. Therefore, vaccine derisking benefits from consideration of potential variation across populations. In order to mitigate potential undesirable epitopes arising from population variation, the inventors sought to eliminate all identical or sequence-similar 7-mers from the GCN4 region. The inventors reasoned that identical 8-mers or 9-mers - which could be recognised by the MHC class I pathway - are less likely to arise through population variation if the any GCN47- mers in the reference human proteome are eliminated. The substitutions to eliminate sequence similarity were made with reference to Blocks Substitution Matrix (BLOSUM) matrices, prioritising evolutionarily favoured amino acid substitutions that had biophysical differences; for example changing charge state, size or potential for hydrogen bonding. Substitutions were also chosen to promote stability, including engineering favourable electrostatic interactions across the three GCN4 alpha-helices. The inventors developed custom software (‘Kmer_Finder’) to identify identical K-mers between a query protein sequence and the human proteome. Kmer_Finder takes the following key steps: a) All contiguous k-mer sequences are identified in the Ensembl human proteome, forming a human k-mer reference database (Human_Kmer_DB); b) the query protein sequence (e.g. GCN4) is split into contiguous k-mers (query_Kmers) and c) the query_Kmers are compared to Human_Kmer_DB. The output of Kmer_Finder reports all human proteins with exact matches to the query_Kmers. The inventors selected k-mer values of k=7, k=8 to at the minimum peptide length recognised by the MHC (i.e. MHC class I alleles typically bind peptides from 8 to 10 amino acids in length). While 7-mers are unlikely to have high affinity to the MHC class I peptide binding cleft, a 7-mer match to the reference human proteome might contribute to an 8- mer or 9-mer match for some individual(s) due to natural variation across the human population. Therefore, exclusion of 7-mers represents a conservative threshold for elimination of potentially undesirable sequence matches between the GCN4 region and proteins in human populations. Additionally, ‘mix and match’ dual binding of non-contiguous peptides, including 7-mers, may result in epitope presentation. Accordingly, exclusion of 7-mers also helps to control for potential hybrid peptide reactivity with the MHC that might potentially have adverse consequences. Visualisation and MSA of matching 7-mers with the GCN4 sequence was produced with ClustalO and JalView. Prediction of T-cell epitopes and protein sequence properties T cells recognise peptides presented at the cell surface by MHC proteins that have a peptide-binding cleft, canonically recognising peptides from 8-10 amino acids (class I, favouring 9-mers) and 13-25 amino acids (class II). Importantly, exogenous antigens, including material from vaccination, may be processed via MHC class I (‘cross-presentation’) as well as the canonical MHC class II pathway. The majority of autoimmune diseases involve MHC class II presentation. For example, multiple sclerosis, Crohn’s disease, dilated cardiomyopathy, Rheumatoid arthritis, extra- articular rheumatoid arthritis and systemic lupus erythematosus. Genetic variation in MHC class II alleles can lead to CD4 autoimmune responses. The inventors predicted T-cell epitopes with TepiTool and NetMHCpan, the GCN4-derived query sequences were analysed for matches to the MHC class I (A-G) and II (DP, DQ, and DR). All peptide lengths were included (8-14mers for class I and 15mers for class II). The default prediction method was applied (‘IEDB recommended’) and predicted peptides with percentile rank <10 were taken forwards. Protease cleavage sites were predicted using EMBOSS tools. Solvent accessibility and secondary structure were obtained from JPred, protein disorder predicted by GlobPlot and Kyte-Doolitte hydrophobicity was visualised in JalView. Inspection of MSAs revealed patterns of amino acid conservation across GCN4 and homologous human sequences, partly arising from the leucine zipper motif (Figure 4 positions 12-36). Homology modelling and structural interactions Discovery Studio Visualizer was used for structure visualisation. The GCN4-derived sequences were input for homology models built using MODELLER 10.2, the 1GCM template was obtained from the PDB. Sequences were aligned to the reference sequence before running MODELLER with bespoke python scripts. Five models were proposed per query sequence and the global free energy determined. The 3D model with lowest normalised DOPE score was taken as input into molecular dynamics simulations. Helical wheel representations were calculated with the EMBOSS ‘pepwheel’ software and residue interactions were determined by LigPlot. Molecular Dynamics The reconstructions of the protein were prepared with GROMACS 2020.4. An AMBER03 force field was applied to the proteins, which were placed into a cube and filled with water as solvent. Six cations of sodium replaced six randomly selected molecules of water to keep the system neutral. The system was energy minimized using steepest descent minimization with an initial step size 0.1 nm, a maximum of steps number of 50000 and a maximum force of 1000 kJmol-1nm-1. According to the NVT ensemble, the system temperature was equilibrated at 300K for 100 ps and then the pressure was equilibrated to 1 bar for 100 ps using Berendsen coupling. The protein position was restrained during the equilibration of temperature and pressure. Each molecular dynamics simulation was run for 50 nanoseconds. Results and Discussion Example 1 – Sequence similarity between GCN4_init and the human proteome In order to investigate matches between GCN4 and the human proteome, the inventors developed the ‘K-mer_finder’ software to identify contiguous regions of identical sequence. The inventors also carried out sensitive BLASTP searches, with an expectation value (e-value) threshold of 1000, in order to identify sequence similar regions regardless of the statistical or evolutionary significance of the match. Matches with human protein sequences were identified for GCN4_init, including to PCIF1, DMD, GOSR2, AGRN and FAM81A. An exhaustive comparison across all 7-mers with Kmer_finder revealed identical matches for GCN4_init with 73 unique Ensembl gene identifiers. The inventors also identified human homologues of the GCN4_init sequence in PSIBLAST searches, although contiguous regions of identical or very similar amino acids were not typically identified in the more evolutionarily distant matches. The PSIBLAST results did not identify any runs of identical residues long enough to bind to the MHC, the closest being the 8-mer ‘GSGGGAGG’ (SEQ ID No: 98) within the JUNB sequence, which had one mismatch when compared with GCN4_init (Figure 5). The inventors found T cell epitopes that covered the entire GCN4_init sequence, Table 1 shows the top 40 MHC class I peptides. Table 1 - high-scoring predicted T cell epitopes for GCN4_init. The forty top-ranked MHC class I epitopes are shown for analysis of GCN4_init with TepiTool. Ranking is based on predicted IC50 value. Columns showing the start and end positions correspond to the GCN4_init sequence. The percentile rank column identifies the peptide’s predicted binding affinity against that of a large set of similarly sized peptides that were randomly selected from the SWISS-PROT database; lower values indicate stronger predicted binding. Start SEQ ID positio End Lengt No Predicted Percentile Allele n position h Peptide IC50 (nM) rank HLA- 56 A*30:01 32 40 9 RIKKLIGEA 29.09 0.13 HLA- 57 A*32:01 22 30 9 KIYHIENEI 42.07 0.05 HLA- 58 B*40:01 12 20 9 IEDKIEEIL 49.81 0.1 HLA- 59 C*03:03 1 9 9 YIKESNHQL 98.45 0.2 HLA- 60 C*03:04 1 9 9 YIKESNHQL 98.45 0.2 HLA- 61 A*68:02 18 26 9 EILSKIYHI 105.94 0.52 HLA- 62 B*18:01 16 24 9 IEEILSKIY 140.42 0.1 HLA- 63 C*12:03 1 9 9 YIKESNHQL 143.71 0.28 HLA- 64 A*02:06 1 9 9 YIKESNHQL 166.93 1.3 HLA- 65 C*03:02 1 9 9 YIKESNHQL 174.73 0.44 HLA- 66 C*16:01 1 9 9 YIKESNHQL 178.13 0.48 HLA- 67 A*68:02 29 37 9 EIARIKKLI 251.91 0.9 HLA- 68 C*17:01 1 9 9 YIKESNHQL 265.88 0.15 HLA- 69 C*12:02 1 9 9 YIKESNHQL 275.4 0.21 HLA- 70 C*02:02 1 9 9 YIKESNHQL 304.75 0.12 HLA- 71 C*02:09 1 9 9 YIKESNHQL 304.75 0.12 HLA- 72 B*08:01 1 9 9 YIKESNHQL 306.6 0.33 HLA- 73 A*02:01 1 9 9 YIKESNHQL 309.14 1.7 HLA- 74 B*40:02 28 36 9 NEIARIKKL 313.75 0.51 HLA- 75 C*14:02 1 9 9 YIKESNHQL 339.51 0.73 HLA- 76 A*68:02 4 12 9 ESNHQLQSI 351.02 1.2 HLA- 77 A*02:06 22 30 9 KIYHIENEI 365.86 2.1 HLA- 78 A*02:06 15 23 9 KIEEILSKI 442.78 2.4 HLA- 79 A*02:01 22 30 9 KIYHIENEI 447.09 2.1 HLA- 80 B*44:02 28 36 9 NEIARIKKL 449.67 0.25 HLA- 81 A*68:02 1 9 9 YIKESNHQL 552.51 1.6 HLA- 82 B*44:03 28 36 9 NEIARIKKL 564.81 0.29 HLA- 83 B*44:03 16 24 9 IEEILSKIY 613.61 0.32 HLA- 84 B*18:01 17 25 9 EEILSKIYH 623.62 0.28 HLA- 85 B*44:02 16 24 9 IEEILSKIY 644.93 0.34 HLA- 86 B*15:25 1 9 9 YIKESNHQL 665.99 2.2 HLA- 18 26 9 EILSKIYHI 87 730.67 3.3 A*02:06 HLA- 88 B*08:01 18 26 9 EILSKIYHI 733.35 0.68 HLA- 89 B*40:02 12 20 9 IEDKIEEIL 779.24 0.99 HLA- 90 C*15:02 22 30 9 KIYHIENEI 907.76 0.48 HLA- 91 B*40:01 28 36 9 NEIARIKKL 1018.86 0.68 HLA- 92 B*15:01 1 9 9 YIKESNHQL 1038.08 2.1 HLA- 93 A*02:06 35 43 9 KLIGEAPGS 1046.09 4 HLA- 94 B*18:01 28 36 9 NEIARIKKL 1096.08 0.43 The existing regions of similarity between GCN4_init and the human reference proteome might potentially be extended in length, for example to produce 9-mer epitopes, as a result of sequence polymorphism. For example, the predicted epitopes ‘KIEEILSKI’ (SEQ ID No: 78) and ‘EILSKIYHI’ (SEQ ID No: 61) (respective predicted IC500.4µM, 0.1µM) might cross-react with the membrane-associated protein ANK3 sequence positions 2511-19 (KEILSKIYK [SEQ ID No: 95]; part of the UniProt canonical ANK3 sequence). Individuals with an ANK3 mutation K2511->E would have an 8-mer peptide that matches exactly to GCN4_init and a K2519->H mutation would generate an 8-mer exact peptide match to GCN4_init that falls within the above predicted epitopes. The K2511->E mutation involves just a single nucleotide change (A->G), whereas K2519->H would require two nucleotide changes (e.g. A->C at the 1stand 3rdcodon positions). Indeed, the peptide (XEILSKIYX [SEQ ID No: 96]) has favourable auxiliary anchor residues at position four (leucine) and seven (isoleucine), while the negatively charged glutamate at position one (produced by K2511->E) could be a potent primary anchor. Accordingly, the inventors identified possible MHC class I epitopes in GCN4_init with a length of 8 amino acids. Exogenous (extracellular) material, including from vaccination, may be processed by the MHC class I pathway through cross- presentation. Importantly, analysis with NetMHCpan4.1 identified the EEILSKIY (SEQ ID No: 52) peptide as a strong binder to MHC allelle HLA-B*18:01, which is present at relatively high frequency in some populations (Figures 6 and 7). Example 2 – Design of a GCN4 trimerization sequence with an enhanced safety profile and increased stability Rational amino acid modifications were investigated in order to both derisk and improve the stability of the GCN4_native region. Multiple sequence alignments of the GCN4-derived sequences and human candidate matching regions were considered with predicted solvent accessibility, secondary structure and Kyte- Doolitte hydrophobicity. The free energy, radius of gyration and 2D projection of trajectory were taken as indicators of stability. Alteration of the hydrophobic core residues was avoided in order to help preserve the trimeric hydrophobic packing. In order to promote stability, the inventors sought to: avoid charge repulsion between structurally proximal side chains, introduce helix-perpendicular and between- monomer salt bridges, enhance solubility, and balance overall charge. Protease sites were also avoided, in order to prevent cleavage of the engineered protein. Helical wheel representations and LigPlot results for GCN4_init, GCN4_M22 are shown in Figures 8 and 9. The inter-helix contacts for GCN4_init and GCN4_M22 involve hydrophobic Van der Waals (VdW) interactions, electrostatic interactions and as a Ser-Asn hydrogen bond. Electrostatic interactions for GCN4_init are Lys- Glu; however, GCN4_M22 has Arg-Glu interactions. Arg binds to negatively charged amino acids more strongly than Lys with a difference in ΔG of at least -8.6 kJ / mol greater for interaction with Glu; accordingly, GCN4_M22 gains -17.2 kJ / mol over two Arg-Glu electrostatic interactions where GCN4_init has equivalent Lys-Glu interactions. Additionally, Arg has a larger positive surface area for salt bridging and is more frequently represented in thermophilic organisms. The pairs of helices II.ii, II.iii (GCN4_M22) and I.ii, I.iii (GCN4_init) have the same number of residues participating in Van der Waals interactions for hydrophobic packing. However, there is an additional Arg-Glu electrostatic interaction for II.ii, II.iii that confers at least - 40 kJ / mol greater thermodynamic stability to the interaction between the II.ii, II.iii pair compared with GCN4_init helices I.ii, I.iii. GCN4_init has two Glu-Lys interactions between helices I.i, I.ii, while GCN4_M22 helices II.i, II.ii have a single electrostatic interaction. While the GCN4_init additional Glu-Lys interaction would increase thermodynamic stability for helices I.i, I.ii; the net effect on stability at physiological temperatures may be small compared with the single electrostatic interaction observed for GCN4_M22. Additional to having greater similarity with the ‘native’ 1GCM template model (Table 2), the presence of electrostatic interactions between every helix pair in GCN4_M22 is expected to enhance the overall stability of the three-helix bundle relative to GCN4_init. Table 2 – Normalised Z-DOPE score for the GCN4 homology models. GCN4_M22 has the lowest score, which indicates better alignment with the 1GCM template structure. Model Trimer Z-DOPE score GCN4_M22 -0.278 GCN4_init 0.53 The inventors further explored protein stability with Molecular Dynamics (MD) simulations (Figure 10). The results of energy minimisation were taken as a measure of protein complex stability. Table 3 – energy minimization for the GCN4 homology models. Both of the models have low minimization energy, which indicates the geometrical optimization of both models performed well. Model Energy minimisation (in Kj.mol-1) GCN4_M22 -8.2061488e+05 GCN4_init -8.3671081e+05 The GCN4_init model has greater diversity of conformational states than GCN4_M22, with both models reaching a cluster of states located at the right-hand side of their respective trajectories (Fig. 10A) by the end of the simulation. The 1GCM PDB structure remained compact in one cluster. The RMSD of both proteins (mutated and nonmutated) increases from 0 to 10000 ps followed by a plateau. However, GCN4 M22 is more stable than GCN4_init, with broadly lower RMSD. Three protein regions have substantial deviation from their original positions (approximate atom positions 600-1000, 1300-1750, 2200-2400). Most of the key interactions in the GCN4_M22 model occurred in the green zone which is there is no movement from the structure. Three key inter-helix interactions GCN4_M22 form a pseudo-ring in the structure, helping to confer stability; specifically, these residue pairs are Arg22-Glu101, Arg96-Glu64 and Arg52-Glu17. In contrast, GCN4_init has electrostatic / H-bond interactions between two helices and two of these residue pairs are highly unstable (Lys15-Glu101, Asp6-Ser42). Example 3 – The engineered GCN4 sequence produced a doubling in yield Production of the Mumps F protein stabilised by GCN4 with a linker to the Mumps HN protein was investigated for a total of 16 engineered GCN4 sequences. The yield of the properly folded trimeric F protein showed significant variability across the different constructs, which were produced with both removal of human similarity and increased stability in mind. Overall yield for the engineered GCN4_M22 sequence was double that obtained for GCN4_init (Figure 12). Conclusions The inventors have identified a novel de-humanised GCN4 sequence (GCN4_M22), where matches to the Ensembl human proteome were removed from the unmodified GCN4_init input sequence. Specifically, a total of 1397-mer matches and 508-mer matches were eliminated. The inventors have also demonstrated that the de-humanised GCN4 trimerization motif is more stable than the unmodified GCN4_init and the template 1GCM protein. In particular, this is because the de- humanised GCN4_M22 presents electrostatic interactions between all three chains and in more stable regions, whereas the unmodified GCN4_init does not possess electrostatic interactions between two of the helices and the majority of these interactions occur in unstable regions. Moreover, consistent with the higher stability, the inventors also demonstrated that the dehumanised GCN4 trimerization motif had double yield relative to GCN4_init. In summary, therefore, the inventors have identified a GCN4 trimerization motif with enhanced safety, stability and solubility, which can be incorporated into vaccines without the risk of inducing autoimmune reactions.

Claims

Claims 1. A de-humanised GCN4 trimerization motif.

2. The de-humanised GCN4 trimerization motif according to claim 1, wherein the trimerization motif is a coiled-coil trimerization motif.

3. The de-humanised GCN4 trimerization motif according to either claim 1 or claim 2, comprising or consisting of an amino acid sequence substantially as set out in SEQ ID No: 1, 2 or 3, or a fragment or variant thereof having at least 70%, 71%, 72%, 73%, 74%, or 75% sequence identity to SEQ ID No: 1, 2 or 3, or preferably at least 76%, 77%, 78%, 79%, or 80% sequence identity to SEQ ID No: 1, 2 or 3.

4. The de-humanised GCN4 trimerization motif according to any preceding claim, comprising or consisting of an amino acid sequence substantially as set out in SEQ ID No: 1, 2 or 3, or a fragment or variant thereof having at least 81%, 82%, 83%, 84%, or 85% sequence identity to SEQ ID No: 1, 2 or 3, or preferably at least 86%, 87%, 88%, 89%, or 90% sequence identity to SEQ ID No: 1, 2 or 3.

5. The de-humanised GCN4 trimerization motif according to any preceding claim, comprising or consisting of an amino acid sequence substantially as set out in SEQ ID No: 1, 2 or 3, or a fragment or variant thereof having at least 91%, 92%, 93%, 94%, or 95% sequence identity to SEQ ID No: 1, 2 or 3, or preferably at least 96%, 97%, 98%, or 99% sequence identity to SEQ ID No: 1, 2 or 3.

6. The de-humanised GCN4 trimerization motif according to any preceding claim, consisting of an amino acid sequence substantially as set out in SEQ ID No: 1, 2 or 3.

7. The de-humanised GCN4 trimerization motif according to any preceding claim, consisting of an amino acid sequence substantially as set out in SEQ ID No:

1.

8. The de-humanised GCN4 trimerization motif according to any preceding claim, wherein the de-humanised GCN4 trimerization motif comprises electrostatic interactions between all three pairs of helices.

9. The de-humanised GCN4 trimerization motif according to claim 8, wherein the electrostatic interactions between all three pairs of helices are Arg-Glu interactions.

10. The de-humanised GCN4 trimerization motif according to either claim 8 or claim 9, wherein the electrostatic interactions between all three pairs of helices comprise interactions between residue pairs Arg22-Glu101, Arg96-Glu64 and / or Arg52-Glu17.

11. An immunogen comprising the de-humanised GCN4 trimerization motif according to any one of claims 1 to 10.

12. The immunogen according to claim 11, wherein the immunogen elicits an immune response against a viral or a bacterial infection.

13. The immunogen according to either claim 11 or claim 12, wherein the immunogen elicits an immune response against a virus of the Adenoviridae family, Anelloviridae family, Arenaviridae family, Astroviridae family, Bornaviridae family, Bunyaviridae family, Caliciviridae family, Coronaviridae family, Filoviridae family, Flaviviridae family, Hepadnaviridae family, Hepeviridae family, Herpesviridae family, Orthomyxoviridae family, Papillomaviridae family, Paramyxoviridae family, Parvoviridae family, Picobirnaviridae family, Picobirna family, Picornaviridae family, Pneumoviridae family, Polyomaviridae family, Poxviridae family, Reoviridae family, Retroviridae family, Rhaboviridae family, Togaviridae family, or Delta family.

14. The immunogen according to any one of claims 11 to 13, wherein the immunogen elicits an immune response against a virus selected from the group consisting of: measles virus, mumps virus, parainfluenza virus, Nipah virus, Morbillivirus canine, Rinderpest morbillis virus and respiratory syncytial virus (RSV).

15. The immunogen according to any one of claims 11 to 14, wherein the immunogen comprises an F protein selected from the group consisting of: a Measles virus F protein, a Mumps virus F protein, a parainfluenza virus F protein, a Nipah virus F protein, a Morbillivirus canine F protein, a Rinderpest morbillis virus F protein and a respiratory syncytial virus (RSV) F protein.

16. The immunogen according to any one of claims 11 to 15, wherein the immunogen comprises a Mumps virus (MuV) Fusion ectodomain trimer (F protein), preferably wherein the MuV F ectodomain trimer is fused N-terminally to the de- humanised GCN4 trimerization motif.

17. The immunogen according to claim 16, wherein the C-terminal residue of the MuV F ectodomain trimer comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 5, 6 or 7, or a fragment or variant thereof.

18. The immunogen according to either claim 16 or claim 17, wherein the MuV F ectodomain trimer comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 8 or 105, or a fragment or variant thereof.

19. The immunogen according to any one of claims 16 to 18, wherein the immunogen comprises a MuV HN ectodomain, preferably wherein the MuV HN ectodomain is fused C-terminally to the de-humanised GCN4 trimerization motif.

20. The immunogen according to claim 19, wherein the MuV HN ectodomain comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 9, or a fragment or variant thereof.

21. The immunogen according to any one of claims 16 to 20, wherein the immunogen comprises a linker region, preferably wherein the linker region is disposed in between the de-humanised GCN4 trimerization motif and the MuV HN ectodomain.

22. The immunogen according to claim 21, wherein the linker region comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 10 or 11, or a fragment or variant thereof.

23. The immunogen according to any one of claims 11 to 22, wherein the immunogen comprises or consists of an amino acid sequence substantially as set out in SEQ ID No: 12, or a fragment or variant thereof.

24. A nucleic acid encoding the de-humanised GCN4 trimerization motif according to any one of claims 1 to 10, or the immunogen according to any one of claims 11 to 23.

25. The nucleic acid according to claim 24, wherein the de-humanised GCN4 trimerization motif is encoded by the nucleotide sequence substantially as set out in SEQ ID No: 103, or a variant or fragment thereof.

26. The nucleic acid according to claim 24, wherein the immunogen is encoded by the nucleotide sequence substantially as set out in SEQ ID No: 104, or a variant or fragment thereof.

27. A virus-like particle (VLP) comprising the immunogen according to any one of claims 11 to 23.

28. An immunogenic composition comprising the immunogen according to any one of claims 11 to 23 or the VLP according to claim 27, and a pharmaceutically acceptable vehicle.

29. The immunogen according to any one of claims 11 to 23, the virus-like particle according to claim 27, or the immunogenic composition according to claim 28, for use in therapy or prophylaxis.

30. The immunogen according to any one of claims 11 to 23, the virus-like particle according to claim 27, or the immunogenic composition according to claim 28, for use in eliciting an immune response.

31. The immunogen, virus-like particle or immunogenic composition for use according to claim 30, wherein the immunogen, virus-like particle or immunogenic composition elicits an immune response against a viral or bacterial infection, preferably wherein the virus is a virus of the Adenoviridae family, Anelloviridae family, Arenaviridae family, Astroviridae family, Bornaviridae family, Bunyaviridae family, Caliciviridae family, Coronaviridae family, Filoviridae family, Flaviviridae family, Hepadnaviridae family, Hepeviridae family, Herpesviridae family, Orthomyxoviridae family, Papillomaviridae family, Paramyxoviridae family, Parvoviridae family, Picobirnaviridae family, Picobirna family, Picornaviridae family, Pneumoviridae family, Polyomaviridae family, Poxviridae family, Reoviridae family, Retroviridae family, Rhaboviridae family, Togaviridae family, or Delta family.

32. The immunogen, virus-like particle or immunogenic composition for use according to either claim 30 or claim 31, wherein the immunogen, virus-like particle or immunogenic composition elicits an immune response against a virus selected from the group consisting of: measles virus, mumps virus, parainfluenza virus, Nipah virus, Morbillivirus canine, Rinderpest morbillis virus, respiratory syncytial virus (RSV).