Vaccines and methods

JP2024180560A5Active Publication Date: 2025-05-30CAMBRIDGE ENTERPRISE LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024179966
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-09-28
Filing Date
2024-10-15
Publication Date
2025-05-30
Estimated Expiration
2039-09-27

AI Technical Summary

Technical Problem

Current vaccines for RNA viruses, particularly those causing viral hemorrhagic fevers and influenza, face challenges in inducing broadly neutralizing immune responses due to high mutation rates, limited antigen selection methods, and the unpredictability of emerging strains, leading to delayed vaccine development and limited protection.

Method used

A method involving a polypeptide library screening process to identify optimized antigenic pathogen polypeptides that can induce broadly neutralizing immune responses by using antigen-binding molecules to select polypeptides capable of binding and neutralizing multiple strains within a pathogen family, followed by expression in suitable systems like mammalian cells.

Benefits of technology

This approach enables the development of vaccines that can induce robust, cross-protective immune responses against diverse strains of RNA viruses, including enveloped viruses like Ebola and influenza, by identifying polypeptides that effectively bind and neutralize multiple variants, thereby enhancing vaccine efficacy and speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000082_0000
    Figure 00000082_0000
  • Figure 00000082_0001
    Figure 00000082_0001
  • Figure 00000082_0002
    Figure 00000082_0002
Patent Text Reader

Abstract

To provide vaccines and methods.SOLUTION: Described herein are methods for identifying optimized antigenic pathogen polypeptides capable of inducing a broadly neutralizing immune response and associated T-cell responses to a pathogen, and nucleic acid sequences encoding such polypeptides. Also described are methods for determining whether a broadly neutralizing immune response is induced in a subject following immunization with an optimized antigenic pathogen polypeptide or a nucleic acid encoding the optimized pathogen polypeptide. Further described are nucleic acid molecules, polypeptides, vectors, cells, fusion proteins, pharmaceutical compositions, and their use as vaccines against pathogens, especially against emerging or re-emerging pathogens (particularly RNA viruses).SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to methods for identifying optimized antigenic pathogen polypeptides capable of inducing a broadly neutralizing immune response against pathogens, methods for identifying nucleic acid sequences encoding such optimized antigenic pathogen polypeptides, and methods for determining whether a broadly neutralizing immune response is induced in a subject after immunization with an optimized antigenic pathogen polypeptide or a nucleic acid encoding the optimized pathogen polypeptide.The present invention also relates to nucleic acid molecules, polypeptides, vectors, cells, fusion proteins, pharmaceutical compositions, and their use as vaccines against pathogens, particularly against emerging or re-emerging pathogens, particularly RNA viruses.The present invention also relates to pseudotyped virus particles.

[0002] The underlying principle of a vaccine is to prepare the immune system for an encounter with a pathogen. Vaccines trigger the immune system to produce antibodies and T-cell responses that help fight infection. Historically, once a pathogen was isolated and grown, it was produced in large quantities, killed or weakened, and used as a vaccine. Later, recombinant genes from the isolated pathogen were used to generate recombinant proteins, which were mixed with adjuvants to stimulate an immune response. More recently, pathogen genes were cloned into vector systems (attenuated bacteria or viruses) to express and deliver antigens in vivo. All of these strategies rely on pathogens isolated from past pandemics to prevent future epidemics. For pathogens that do not change significantly or change slowly, this traditional technique works. However, some pathogens tend to mutate, and antibodies do not necessarily recognize different strains of the pathogen. Emerging and re-emerging pathogens often hide or disguise their vulnerable antigens from the immune system.

[0003] Among emerging and re-emerging diseases, a disproportionate number (37%) are caused by ribonucleic acid (RNA) viruses (Heeney, Journal of Internal Medicine 2006; 260: 399-408). RNA viruses are viruses that have RNA as their genetic material. This nucleic acid is usually single-stranded RNA (ssRNA), but can also be double-stranded RNA (dsRNA). RNA viruses generally have a much higher mutation rate compared to DNA viruses because viral RNA polymerase lacks the proofreading ability of DNA polymerase. This is one reason why it is difficult to generate effective vaccines to prevent diseases caused by RNA viruses. For the most part, current vaccine candidates against RNA viruses are limited to the virus strains used as vaccine inserts, which are often selected based on the availability of wild-type strains rather than being designed with information. Technical challenges in developing vaccines for enveloped RNA viruses include: i) viral variation in wild-type field isolate glycoproteins (GPs) results in limited breadth of protection as vaccine antigens; ii) selection of vaccine antigens expressed by the vaccine insert is highly empirical, and immunogen selection is a slow trial-and-error process; and iii) in emerging or unexpected virus outbreaks, development of new vaccine candidates is time consuming and may delay vaccine deployment.

[0004] Notable human diseases caused by RNA viruses include viral hemorrhagic fevers (VHFs), a group of diseases caused by viruses from several distinct families. In general, the term "viral hemorrhagic fevers" is used to describe severe multiorgan syndromes (i.e., multiple organ systems of the body are affected). Characteristically, the vascular system throughout the body is damaged, compromising the body's ability to regulate itself. These symptoms are often accompanied by hemorrhage (blood loss), but the blood loss itself is rarely life-threatening. Some types of hemorrhagic fever viruses can cause relatively mild illness, but many of the viruses cause severe, life-threatening illness. VHFs are caused by viruses from at least five distinct families: Arenaviridae, Bunyaviridae, Filoviridae, Flaviviridae, and Paramyxoviridae. All of the viruses in these families are RNA viruses, and all are covered with a fatty (lipid) coat or have an envelope in the coat. Survival from VHFs depends on an animal or insect host (the natural reservoir). Viruses are geographically restricted to areas where their host species survive, and humans become infected when they come into contact with an infected host. For some viruses, humans can transmit the virus to each other after transmission from the host. Human cases or outbreaks of hemorrhagic fever caused by these viruses occur sporadically and irregularly. The occurrence of outbreaks cannot be easily predicted. With some exceptions, there is no cure for VHF and no established drug treatment.

[0005] VHFs caused by arenaviruses and filoviruses together span a wide geographical area from West Africa to Central Africa, threatening adjacent areas where infected animal reservoirs may migrate but where human disease has not yet been reported. Filoviruses encode their genome in the form of single-stranded, negative-sense RNA. Two commonly known members of this family are Ebola virus and Marburg virus. Ebola is an emerging and re-emerging RNA viral disease. Outbreaks are not always caused by the exact same virus, but by different relatives (types) of the same viral family, among which there are close siblings (e.g., Ebola Mayinga and Ebola Kikwit), close cousins ​​(Taï Forest and Bundibugyo), distant cousins ​​(Sudan), and distant relatives (Marburg virus). The Ebola outbreak in West Africa in 2014 was the largest since the viral disease was first recognized. Arenaviruses are classified into two groups: Old World viruses and New World viruses. The differences between these groups are geographically and genetically distinct. At least eight arenaviruses are known to cause human diseases of varying severity. Aseptic meningitis is a severe human disease that causes inflammation involving the brain and spinal cord and can result from lymphocytic choriomeningitis virus (LCMV) infection. Hemorrhagic fever syndromes result from infection with viruses such as Guanarito virus (GTOV), Junin virus (JUNV), Lassa virus (LASV), Lujo virus (LUJV), Machupo virus (MACV), Sabia virus (SABV), or Whitewater Arroyo virus (WWAV).

[0006] Lassa fever virus (LASV), Ebola (EBOV), and Marburg (MARV) viruses are the most important hemorrhagic fevers in West and Central Africa. Lassa fever is endemic in West Africa, with estimates ranging from 300,000 to 1 million infections and 5,000 deaths per year. Lassa fever virus (LASV), Ebola (EBOV), and Marburg (MARV) viruses are all containment level 4 pathogens, with high human morbidity and mortality, no established cure, and currently no licensed vaccines against infections caused by these viruses.

[0007] Influenza viruses are members of the family Orthomyxoviridae. There are three types of influenza viruses, termed influenza A, influenza B, and influenza C. Influenza A viruses infect a wide variety of birds and mammals, including humans, horses, marine mammals, pigs, ferrets, and chickens. In animals, most influenza A viruses cause mild, localized infections of the respiratory and intestinal tracts. However, highly pathogenic influenza A strains, such as H5N1, can cause systemic infections in livestock, with mortality rates approaching 100%. In 2009, H1N1 influenza was the most common cause of human influenza. A novel strain of H1N1 of swine origin emerged in 2009 and was declared a pandemic by the World Health Organization. This strain was termed "swine flu." H1N1 influenza A viruses were also responsible for the Spanish Flu pandemic of 1918, the Fort Dix pandemic of 1976, and the Russian Flu pandemic of 1977-1978. Currently, two influenza vaccine approaches are licensed in the United States: inactivated split vaccines and live attenuated virus vaccines. Inactivated vaccines can efficiently induce humoral immune responses but generally induce cell-mediated immune responses poorly. Live virus vaccines cannot be administered to immunocompromised or pregnant patients due to increased risk of infection.

[0008] There is therefore a need to provide effective vaccines that induce broadly neutralizing immune responses to protect against emerging and re-emerging diseases, particularly those caused by viruses such as RNA viruses, including VHF and influenza. [Prior art documents] [Non-patent literature]

[0009] [Non-Patent Document 1] Heeney, Journal of Internal Medicine (2006) 260:399~408 Summary of the Invention [Means for solving the problem]

[0010] In accordance with the present invention, there is provided a method for identifying lead candidate optimized antigenic pathogen polypeptides capable of inducing a broadly neutralizing immune response against a pathogen, comprising the steps of: i) providing a polypeptide library comprising a plurality of different candidate optimized antigenic pathogen polypeptides, wherein the amino acid sequence of each different candidate is optimized from a plurality of different amino acid sequences of a pathogen polypeptide, and wherein, unlike each different amino acid sequence of the pathogen polypeptide, each different amino acid sequence of the pathogen polypeptide comprises an amino acid sequence of a polypeptide of a different isolate, each different isolate being an isolate of a pathogen of the same family as a pathogen against which it is desired to induce a broadly neutralizing immune response; ii) screening the candidate optimized antigenic pathogen polypeptides of the polypeptide library for binding by one or more broadly neutralizing antigen binding molecules, each capable of binding to and / or neutralizing pathogens of the same family as the pathogen against which it is desired to induce a broadly neutralizing immune response; and iii) identifying the candidate optimized antigenic pathogen polypeptide bound by one or more of the antigen binding molecules in step (ii) as a lead candidate optimized antigenic pathogen polypeptide capable of inducing a broadly neutralizing immune response against the pathogen. A method is provided that includes:

[0011] Optionally, each of the different isolates or each of the plurality of different isolates of a pathogen is an isolate of the same subtype or type as the pathogen against which it is desired to induce a broadly neutralizing immune response.

[0012] Optionally, each of the different isolates or each of the plurality of different isolates of a pathogen is an isolate of the same species or genus as the pathogen against which it is desired to induce a broadly neutralizing immune response.

[0013] Optionally, the different isolates include isolates of different subtypes or types within the same family as the pathogen against which it is desired to induce a broadly neutralizing immune response.

[0014] Optionally, the different isolates include isolates of different species or genus within the same family as the pathogen against which it is desired to induce a broadly neutralizing immune response.

[0015] The term "pathogen" as used herein refers to anything that can cause disease, particularly an infectious agent that can cause disease, such as a virus, bacterium, fungus, or parasite.

[0016] The term "polypeptide" as used herein refers to a polymer comprising multiple amino acid residues linked together by peptide bonds to form a chain. All proteins are polypeptides. The term "polypeptide" is used interchangeably with the term "protein". The term "polypeptide" is specifically intended to cover naturally occurring proteins as well as recombinantly or synthetically produced proteins. Optionally, the polypeptide is a modified polypeptide, e.g., a polypeptide that has been co-translationally or post-translationally modified, e.g., a glycosylated polypeptide or protein ("glycoprotein"). Glycoproteins are proteins that contain oligosaccharide chains (glycans) covalently attached to amino acid side chains. Carbohydrates are attached to proteins by co-translational or post-translational glycosylation.

[0017] "Pathogen polypeptide" refers to any polypeptide forming portion of a pathogen. Optionally, the pathogen polypeptide is a structural protein (or a portion thereof) of the pathogen. Optionally, the pathogen polypeptide is a surface-exposed structural protein (or a portion thereof) of the pathogen. Optionally, the pathogen polypeptide is a viral protein (or a portion thereof). Optionally, the pathogen polypeptide is a viral envelope protein (or a portion thereof). Optionally, the pathogen polypeptide is a glycoprotein (or a portion thereof). Optionally, the pathogen polypeptide is a viral glycoprotein (or a portion thereof). Optionally, the pathogen polypeptide is a viral envelope glycoprotein (or a portion thereof). Optionally, the pathogen polypeptide is an external viral envelope glycoprotein (or a portion thereof). Optionally, the pathogen polypeptide comprises an amino acid sequence of at least 20 amino acid residues. Optionally, the pathogen polypeptide comprises an amino acid sequence of up to 1000, 900, 800, 700, or 600 amino acid residues.

[0018] A fully assembled infectious virus is known as a virion. The simplest virion consists of nucleic acid (single- or double-stranded RNA or DNA) and a capsid protein coat. The capsid is formed as a single or double protein shell and consists of only one or a few structural protein species. Enveloped viruses have an envelope that encases their protective protein capsid. The envelope is typically derived from parts of the host membrane (phospholipids and proteins) but contains virally encoded glycoproteins.

[0019] The glycoproteins on the envelope surface serve to identify and bind to receptor sites on the host membrane. The viral envelope then fuses with the host membrane, allowing the capsid and viral genome to enter and infect the host. Virus-cell membrane fusion is the means by which all enveloped viruses, including human pathogens such as filoviruses, influenza viruses, and human immunodeficiency virus (HIV), enter cells and initiate viral infection. This membrane fusion process is carried out by one or more viral envelope glycoproteins. Fusion can occur on the cell plasma membrane or on the endosomal membrane.

[0020] Glycoproteins can help viruses evade the host immune system. Enveloped viruses have great adaptability and can change in a short time to evade the host immune system. Enveloped viruses can cause persistent infection. Enveloped RNA viruses include, for example, flaviviruses, togaviruses, coronaviruses, hepatitis D, orthomyxoviruses, paramyxoviruses, rhabdoviruses, bunyaviruses, and filoviruses. Retroviruses are enveloped viruses. Enveloped DNA viruses include herpesviruses, poxviruses, and hepadnaviruses.

[0021] Most external viral envelope proteins are glycoproteins, often assembled as dimers or trimers and present as membrane-anchored spikes. Trimeric glycoprotein (GP) spikes on the filovirus envelope mediate all stages of viral entry, including binding, entry, and fusion. The recognition sites for cellular receptors are often located in the domains furthest from the viral envelope (distal ends), while the proximal domains interact with the lipid bilayer of the envelope. Oligosaccharide side chains (glycans) are linked by N-glycosidic or, more rarely, O-glycosidic bonds. Because they are synthesized by cellular glycosyltransferases, the sugar composition of these glycans is similar to that of host cell membrane glycoproteins.

[0022] Entry of filoviruses on the cell surface has been shown to be mediated by host cell binding factors, such as C-type lectins including DC-SIGN (dendritic cell-specific ICAM3-grabbing non-integrin; also known as CD209) and L-SIGN (liver and lymph node SIGN; also known as CLEC4M), as well as several cell surface proteins, such as integrins, T-cell immunoglobulin and mucin domain-containing (TIM) proteins, and tyrosine protein kinase receptor 3 (TYRO3) family members. After binding to the cell surface, filoviruses are internalized by a macropinocytosis-like process and then transported through early and late endosomes. Then, after the viral envelope fuses with the membrane of the late endosome, the viral genome enters into the cytoplasm. In the cytoplasm, the viral genome is replicated and transcribed, and new viral proteins are synthesized to assemble progeny virions, which bud from the cell surface.

[0023] The surface glycoprotein GP of Ebola virus (EBOV) is a key component of many vaccines and a target of neutralizing antibodies. EBOV GP is synthesized as a single polypeptide and then cleaved by furin-like proteases into GP1 and GP2 subunits, which are held together through intersubunit disulfide bonds and non-covalent interactions to form a trimer of GP1-GP2 heterodimers on the viral surface. However, furin cleavage is not sufficient to prime EBOV GP. After cell entry, the virus is eventually transported to late endosomes, where GP is further primed to remove some "cap" components, thereby triggering the induction of key membrane fusion events that lead to viral entry. EBOV GP priming is mediated by the cysteine ​​proteases cathepsin B and cathepsin L, which cleave GP1 within the β13-β14 loop. Cathepsin cleavage removes approximately 60% of the amino acids from GP1, including the mucin-like domain, the glycan cap, and the outermost β-strand of the proposed receptor-binding region, resulting in a primed form of GP (termed GPcl, 19 kDa GP1 plus GP2). Unlike full-length GP, primed GPcl cannot bind to the endosomal membrane protein Niemann-Pick C1 (NPC1), an essential host entry factor for EBOV infection. The crystal structures of free NPC1-C and its complex with GPcl have been determined (Wang et al., Cell, 2016, 164, 258-268). During Ebola virus infection, the major product of the GP gene is secreted GP (sGP), a soluble dimer that lacks GP2 and the mucin-like domain but shares 295 amino acids of GP1.

[0024] Influenza virions contain a segmented negative-sense RNA genome, which encodes the following proteins: hemagglutinin (HA), neuraminidase (NA), matrix (Ml), proton ion channel protein (M2), nucleoprotein (NP), polymerase basic protein 1 (PB1), polymerase basic protein 2 (PB2), polymerase acidic protein (PA), and nonstructural protein 2 (NS2). HA, NA, Ml, and M2 are membrane-associated, whereas NP, PB1, PB2, PA, and NS2 are nucleocapsid-associated proteins. The Ml protein is the most abundant protein in influenza particles. The HA and NA proteins are envelope glycoproteins responsible for virus binding and entry of virus particles into cells, and are the source of the major immunodominant epitopes for virus neutralization and protective immunity. Both HA and NA proteins are considered to be the most important components for a prophylactic influenza vaccine.

[0025] With respect to bacteria or fungi, suitable pathogen polypeptides include polypeptides that are essential for the reproduction of the bacteria or fungus, or for the ability of the bacteria or fungus to infect or cause disease in humans. Suitable examples include surface-expressed polypeptides or proteins (e.g., Hu et al., Front.Microbiol.8:82. doi: 10.3389 / fmicb.2017.00082; Santos and Levitz, Cold Spring Harb Perspect Med. 2014; 4(11): a019711).

[0026] The term "antigenic" as used herein refers to a substance capable of inducing an immune response in a host organism. The immune response can be a humoral and / or cellular immune response. A cellular immune response is the response of cells of the immune system, such as B cells, T cells, macrophages, or polymorphonuclear leukocytes, to a stimulus, such as an antigen or a vaccine. An immune response can include any cell of the body involved in a host defense response, including, for example, epithelial cells that secrete interferons or cytokines. An immune response includes, but is not limited to, an innate immune response or inflammation. As used herein, a protective immune response refers to an immune response that protects a subject from infection or disease (i.e., prevents infection or prevents the development of disease associated with infection). Methods of measuring immune responses are well known in the art and include, for example, measuring lymphocyte (e.g., B or T cell) proliferation and / or activity, cytokine or chemokine secretion, inflammation, or antibody production.

[0027] If desired, the optimized antigenic pathogen polypeptide can induce the production of antibodies and / or a T cell response in a human or non-human animal to which the polypeptide has been administered (either as a polypeptide or, for example, expressed from an administered nucleic acid expression vector).

[0028] The term "antibody" as used herein refers to an immunoglobulin molecule with a specific amino acid sequence produced by B lymphoid cells. Antibodies are induced in humans or other animals by a specific antigen (immunogen). Antibodies are characterized by specifically reacting with an antigen in some demonstrable manner, and antibodies and antigens are each defined in relation to the other. "Inducing an antibody response" refers to the ability of an antigen or other molecule to induce the production of antibodies.

[0029] "Neutralizing" antibodies or antigen-binding molecules not only bind to pathogens, such as viruses, but also bind in such a way that they inhibit (i.e., reduce) or block infection or the progression of infection. Neutralizing antibodies or antigen-binding molecules can block interaction with receptors or bind to viral capsids to inhibit genome uncoating. The term "neutralizing antibodies" or "neutralizing antigen-binding molecules" also includes antibodies or antigen-binding molecules that can prevent infection by pathogens, such as viruses, by promoting cytokine responses or by promoting uptake and removal by immune cells. In particular, the term "neutralizing antibodies" includes antibodies (or fragments or derivatives thereof) that can inhibit or block infection (or the progression of infection) of pathogens by antibody-dependent cell-mediated cytotoxicity (ADCC) or complement-dependent cytotoxicity (CDC). Only a small subset of many antibodies that bind to viruses are neutralizing.

[0030] The term "broadly neutralizing antigen binding molecule" as used herein includes an antigen binding molecule, such as an antibody or a fragment or derivative thereof, that can inhibit (i.e., reduce), neutralize, or prevent infection of at least two different subtypes or species of pathogens, such as at least two different subtypes or species of viruses, at least two different subtypes or species of bacteria, or at least two different subtypes or species of fungi. Optionally, a broadly neutralizing antigen binding molecule can inhibit (i.e., reduce), neutralize, or prevent infection of most or all different subtypes or species of pathogens, such as most or all different subtypes or species of viruses, most or all different subtypes or species of bacteria, or most or all different subtypes or species of fungi. Optionally, a broadly neutralizing antibody can inhibit (i.e., reduce), neutralize, or prevent infection of at least two different types of members of the same family of pathogens (e.g., viruses, bacteria, or fungi).

[0031] Optionally, a plurality of different broadly neutralizing antigen binding molecules are used in step (ii) of the method of the present invention. Optionally, each different broadly neutralizing antigen binding molecule binds to a different region or epitope of a candidate optimized antigen pathogen polypeptide of the polypeptide library.

[0032] The term "broad neutralizing immune response" as used herein refers to an immune response induced in a subject that is sufficient to inhibit (i.e., reduce), neutralize, or prevent infection and / or progression of at least two different subtypes or species of pathogens, such as at least two different subtypes or species of viruses, at least two different subtypes or species of bacteria, or at least two different subtypes or species of fungi.Optionally, the broad neutralizing immune response is sufficient to inhibit, neutralize, or prevent infection and / or progression of most or all different subtypes or species of pathogens, such as most or all different subtypes or species of viruses, most or all different subtypes or species of bacteria, or most or all different subtypes or species of fungi.Optionally, the broad neutralizing immune response is sufficient to inhibit, neutralize, or prevent infection and / or progression of at least two different types of members of the same family of pathogens (e.g., viruses, bacteria, or fungi). Optionally, a broadly neutralizing immune response is sufficient to inhibit, neutralize, or prevent infection and / or progression of infection with members of at least two different genera of pathogens within the same family (e.g., viruses, bacteria, or fungi).

[0033] Some broadly neutralizing antibodies against pathogens are known. For example, some antibodies have been demonstrated to be capable of neutralizing virus isolates of various subtypes in the Filoviridae family. A systematic analysis of monoclonal antibodies against Ebola virus glycoproteins is described by Saphire et al. (Cell, 2018; 174(4): 938-952). An example of a broadly neutralizing antibody against Ebola virus is the immune-induced macaque antibody CA45 described by Zhao et al., 2017 (Cell 169, 891-904). A broadly neutralizing monoclonal antibody against HIV-1 envelope protein is referenced in Bruun et al. (PLoS ONE 9(10): e109196. doi:10.1371 / journal.pone.0109196). Corti et al. (Curr Opin Virol. 2017 Jun;24:60-69) provide a review of the specificity, antiviral and immunological mechanisms of action, and clinical development of broadly reactive monoclonal antibodies against influenza A and B viruses.

[0034] Optionally, the pathogen is a virus.

[0035] Viruses are classified primarily by phenotypic characteristics, such as morphology, nucleic acid type, mode of replication, host organism, and the type of disease they cause. One scheme for classifying viruses, the Baltimore classification system, places viruses into one of seven groups according to a combination of their nucleic acid (DNA or RNA), stranding (single- or double-stranded), sense, and method of replication: I: dsDNA viruses (e.g. adenoviruses, herpes viruses, pox viruses); · II: ssDNA viral (+ strand or "sense") DNA (e.g. parvovirus); III: dsRNA viruses (e.g., reovirus); · IV: (+)ssRNA viral (+strand or sense) RNA (e.g. picornaviruses, togaviruses); · V: (-)ssRNA viral (negative-stranded or antisense) RNA (e.g., orthomyxoviruses, filoviruses, arenaviruses, rhabdoviruses); · VI: ssRNA-RT viral (positive strand or sense) RNA with a DNA intermediate in the life cycle (e.g., retroviruses); · VII: dsDNA-RT viral DNA that has an RNA intermediate in its life cycle (e.g., hepadnaviruses).

[0036] Optionally, the virus is an RNA virus. RNA viruses include: Group III: Viruses have a double-stranded RNA genome; Group IV: The virus has a positive-sense single-stranded RNA genome. Many well-known viruses are found in this group, including picornaviruses (a family of viruses that includes well-known viruses such as Hepatitis A virus, Enteroviruses, Rhinoviruses, Polioviruses, and Foot and Mouth Disease virus), SARS virus, Hepatitis C virus, Yellow Fever virus, and Rubella virus; Group V: The virus has a negative-sense single-stranded RNA genome. Ebola and Marburg viruses are well-known members of this group, along with influenza virus, Lassa virus, measles, mumps, and rabies.

[0037] The classification of different RNA virus families according to the Baltimore classification is given in the following table: [Table 2]

[0038] Optionally, the virus is an emerging or re-emerging RNA virus. Examples of emerging or re-emerging RNA viruses include Ebola virus, Marburg virus, Lassa virus, influenza virus, MERS coronavirus, Hendra virus, and Nipah virus.

[0039] Optionally, the virus is a filovirus or arenavirus. Optionally, the virus is an Ebola virus or a Marburg virus. Optionally, the virus is a Lassa virus. Optionally, the virus is an influenza virus.

[0040] Optionally, the pathogen is a DNA virus. Optionally, the pathogen is a member of the Poxviridae family, such as monkeypox virus.

[0041] DNA viruses include: Group I: The virus has double-stranded DNA. The viruses that cause chickenpox and herpes are found in this group. Group II: The viruses have single-stranded DNA.

[0042] The Baltimore classification of different DNA virus families is given in the table below: [Table 3]

[0043] Optionally, the pathogen is a reverse transcribing virus. Reverse transcribing viruses include: Group VI: Viruses have single-stranded RNA viruses that replicate through a DNA intermediate. Retroviruses are included in this group, of which HIV is a member. Group VII: Viruses have a double-stranded DNA genome and replicate using reverse transcriptase. Hepatitis B virus can be found in this group.

[0044] The term "subtype" as used herein refers to a genetic variant or strain of a pathogen (e.g., virus, bacteria, or fungus). For example, the Ebolavirus genus is a virological taxon within the Filoviridae family. Members of this genus are called Ebola viruses. There are six known Ebola virus subtypes, each named for the region in which it was first identified: Bundibugyo, Reston, Sudan, Tai Forest, Zaire, and Bombali. Influenza A viruses are classified into subtypes based on two proteins on the surface of the virus: hemagglutinin (HA) and neuraminidase (NA). There are 18 known HA subtypes and 11 known NA subtypes. Many different combinations of HA and NA proteins are possible. For example, "H7N2 virus" designates an influenza A virus subtype that has an HA7 protein and an NA2 protein. Similarly, an "H5N1" virus has an HA5 protein and an NA1 protein.

[0045] The nomenclature of natural variant viruses of the Filoviridae family is discussed in Kuhn et al. (Arch Virol. 2013 Jan; 158(1): 301-311). According to the authors, a (natural) virus strain is "a variant of a given virus that is recognizable because it possesses some unique phenotypic characteristics that remain stable under natural conditions". Such "unique phenotypic characteristics" are biological properties that are different from the reference virus to which it is compared, such as unique antigenic properties, host range, or the manifestation of the disease it causes. "A virus variant with a simple difference in genomic sequence is not given the status of a separate strain, since there is no recognizable distinct virus phenotype". Thus, a strain is a genetically stable virus variant that differs from the natural reference virus (the typical variant) in that it causes a significantly different observable phenotype of infection (different types of disease, infecting different types of hosts, being transmitted by different means, etc.). "Genetically stable" means that the genomic changes associated with the phenotypic changes are largely conserved over time through natural selection. The degree of variation in the genome sequence is irrelevant to the classification of the variant as a strain, since distinct phenotypes sometimes result from several mutations. "Observable phenotype" means that, for example, in comparable animal experiments, the researcher is able to distinguish between animals infected with a reference control virus and those infected with a supposedly novel strain, without knowing which virus was administered to the animals and without any information on the differences between the two viruses. It is the responsibility of an international group of experts to designate a virus variant as a virus strain. So far, no natural filovirus strains have been reported that comply with this definition. All the genetic variants described, for example for EBOV, cause similar hemorrhagic fevers in humans and in experimental animals and are transmitted in the same way. None of the known EBOV genetic variants can be distinguished from the others solely from a clinical standpoint. In fact, the diversity is limited to slight differences in growth kinetics and plaque formation in vitro, or slight changes in disease duration in experimental animals, and appears to ultimately derive from limited, but stable, genome sequence differences.This is also true for the different genetic variants of MARV, RAVV, BDBV, RESTV, and SUDV (currently there is only one isolate of TAFV and no isolates of LLOV).

[0046] According to Kuhn et al., a naturally occurring genetic filovirus variant is a naturally occurring filovirus whose genomic consensus sequence differs by ≤10% from the sequence of a reference filovirus (the prototypic virus of a particular filovirus species), but is not identical to the reference filovirus, and does not cause an observable different phenotype of disease (filovirus strains are genetic filovirus variants, but most genetic filovirus variants are not filovirus strains, as the strain definition goes).

[0047] Another scheme for classifying viruses is that of the International Committee on Taxonomy of Viruses (ICTV). The system shares many features with the classification systems for cellular organisms, such as the taxon structure. However, this naming system differs in several ways from other taxonomic codes. The virus classification starts at the order level and proceeds as follows, with the taxon suffixes in italics: Eyes (-virales) Family (-viridae) Subfamily (-virinae) Genus(-virus) seed

[0048] Species designations often refer to viruses, particularly those that affect higher plants and animals.

[0049] The establishment of an order is based on the presumption that the virus families it contains most likely evolved from a common ancestor. Most of the virus families remain unexplored. As of 2017, nine orders, 131 families, 46 subfamilies, 803 genera, and 4,853 species of viruses have been defined by the ICTV. The orders are: Caudovirales, Herpesvirales, Ligamenvirales, Mononegavirales, Nidovirales, Ortervirales, Picornavirales, Bunyavirales, and Tymovirales. These orders span viruses with diverse host ranges. · Caudovirales are tailed dsDNA (group I) bacteriophages. · Herpesvirales contains the large eukaryotic dsDNA viruses. · Ligamenvirales contains linear dsDNA (group I) Archaean viruses. · Mononegavirales includes non-segmented (negative) stranded ssRNA (group V) plant and animal viruses. · Nidovirales consists of (+) stranded ssRNA (group IV) viruses that have vertebrate hosts. · Ortervirales contains single-stranded RNA and DNA viruses that replicate through a DNA intermediate (groups VI and VII). · The Picornavirales contains small (+)stranded ssRNA viruses that infect a variety of plant, insect, and animal hosts. · Tymovirales contains monopartite (+)ssRNA viruses that infect plants. · Bunyavirales contains tripartite (-) ssRNA viruses (group V).

[0050] According to the ICTV, a viral species is "a monophyletic group of viruses whose properties can be distinguished from those of other species by one or more criteria."

[0051] The term "isolate" as used herein refers to a pure pathogen sample obtained from an infected individual. A virus-infected cell already contains a population of genomes after only one round of replication, and the virions derived from these genomes differ slightly from each other. Similarly, a sample taken from an infected individual contains a large number of virions, many of which differ very slightly. As a result, an "isolate" refers to a population, and the "sequence" of an "isolate" is the consensus sequence of the population of genomes present in the analyzed sample. A virus isolate can be defined as an "example of a particular virus." A natural filovirus isolate is an example of a particular natural filovirus, or an example of a particular genetic variant. Isolates can be identical in consensus sequence or individual sequence, or slightly different from each other.

[0052] Optionally, the one or more broadly neutralizing antigen binding molecules comprise antibodies obtained, or derived from antibodies obtained, from a subject exposed to a pathogen of the same family as the pathogen against which it is desired to elicit a broadly neutralizing immune response.

[0053] Optionally, the one or more broadly neutralizing antigen binding molecules comprise antibodies obtained, or derived from antibodies obtained, from a subject exposed to a pathogen of the same subtype or type as the pathogen against which it is desired to elicit a broadly neutralizing immune response.

[0054] Optionally, the one or more broadly neutralizing antigen binding molecules comprise antibodies obtained, or derived from antibodies obtained, from a subject exposed to a pathogen of the same species or genus as the pathogen against which it is desired to elicit a broadly neutralizing immune response.

[0055] Optionally, the one or more broadly neutralizing antigen-binding molecules comprise a non-antibody antigen-binding protein. For example, the one or more broadly neutralizing antigen-binding molecules may comprise a designed ankyrin repeat protein (DARPin), an aptamer, anticalin, or a T cell receptor molecule.

[0056] DARPins are genetically engineered antibody-mimicking proteins that typically exhibit highly specific and high affinity target protein binding. They are derived from natural ankyrin proteins and contain repeated structural units that form stable protein domains with large potential target interaction surfaces. Typically, DARPins contain four or five repeats, of which the first (N-capping repeat) and the last (C-capping repeat) serve to provide a hydrophilic surface. DARPins correspond to the average size of natural ankyrin repeat protein domains. Proteins with fewer than three repeats (i.e. the capping repeat and one internal repeat) do not form sufficiently stable tertiary structures. The molecular weight of a DARPin depends on the total number of repeats: [Table 4]

[0057] 10 12 Libraries of nucleic acids encoding DARPins with random potential target interaction residues with a diversity of more than 100 variants can be generated. From these libraries, DARPins that bind to the desired target of choice with picomolar affinity and specificity can be selected using ribosome display or phage display using a signal sequence that allows for cotranslational secretion. In this way, screening of a library of DARPins can identify one or more DARPins that bind to and / or neutralize more than one subtype of a pathogen. Library-based screening to identify DARPins is described, for example, in Hartmann et al. (Molecular Therapy: Methods and Clinical Development 2018 Vol. 10: 128-143).

[0058] Optionally, the one or more antigen-binding molecules described in step (ii) of the method of the invention comprise a broadly neutralizing antibody (or a fragment or derivative thereof that retains broadly neutralizing activity), such as a broadly neutralizing monoclonal antibody (BNmAb) (or a fragment or derivative thereof that retains broadly neutralizing activity).

[0059] Optionally, the one or more antigen-binding molecules described in step (ii) of the method of the present invention comprise antibodies obtained, or derived from antibodies obtained, from subjects who have survived an outbreak of a pathogen of the same subtype, type or family as the pathogen against which it is desired to elicit a broadly neutralizing immune response.

[0060] Optionally, the one or more antigen-binding molecules described in step (ii) of the method of the present invention comprise antibodies obtained, or derived from antibodies obtained, from subjects who have survived an outbreak of a pathogen of the same species, genus or family as the pathogen against which it is desired to elicit a broadly neutralizing immune response.

[0061] The term "epidemic" as used herein refers to the occurrence of a disease in a defined facility (e.g., a hospital or medical center), community, geographic area, or period of time with more cases than would normally be expected. An epidemic can occur in a localized geographic area or spread across several countries. It can last for days or weeks, or years. The number of cases that indicate the presence of an epidemic varies depending on the pathogen, the size and type of the exposed population, the lack of previous experience or exposure to the disease, and the duration and location of the outbreak. Thus, an epidemic situation correlates with the usual frequency of the disease in the same area, in the same community, and during the same season of the year. The presence of an epidemic can be established by comparing current information with historical incidence in a population or community during the same time of the year to determine whether the number of cases observed exceeds the expected number.

[0062] Where appropriate, an outbreak of a pathogen can refer to the presence of more cases of a disease caused by the pathogen than would normally be expected in a region (e.g., a continent) or country, or in a population or community over one or more seasons or years.

[0063] Optionally, an outbreak of a pathogen (eg, a virus) is the occurrence of more cases of disease caused by the pathogen than would normally be expected in a given region (eg, a continental region) over a given season.

[0064] Optionally, an outbreak of a pathogen (eg, a virus) is the occurrence of more cases of a disease caused by the pathogen than would normally be expected in a population over a season.

[0065] Examples of continental regions include the African region: North Africa: Algeria;Canary Islands;Ceuta;Egypt;Libya;Madeira;Melilla;Morocco;Sudan;Tunisia;Western Sahara; East Africa: Burundi; Comoros; Djibouti; Eritrea; Ethiopia; Kenya; Madagascar; Malawi; Mauritius; Mayotte; Mozambique; Reunion; Rwanda; Seychelles; Somalia; South Sudan; Tanzania; Uganda; Zambia; Zimbabwe; · Central Africa: Angola; Cameroon; Central African Republic; Chad; Democratic Republic of the Congo; Republic of the Congo; Equatorial Guinea; Gabon; Sao Tome and Principe; West Africa: Benin; Burkina Faso; Cape Verde; Côte d'Ivoire; Gambia; Ghana; Guinea; Guinea-Bissau; Liberia; Mali; Mauritania; Niger; Nigeria; Saint Helena; Senegal; Sierra Leone; Togo; · Southern Africa: Botswana; Lesotho; Namibia; South Africa; Swaziland.

[0066] Optionally, the subject from which the antibody is obtained or derived is a human or non-human mammalian subject.

[0067] Candidate optimized antigenic pathogen polypeptides of the polypeptide library can be expressed using any suitable expression system, suitable examples include mammalian cells, or yeast, or insect, or bacterial cells.

[0068] Optionally, the candidate optimized antigenic pathogen polypeptides of the polypeptide library are expressed on the cell surface of an expression system, where cell surface expression increases the likelihood that the candidate optimized antigenic pathogen polypeptides will be correctly folded.

[0069] Optionally, the candidate optimized antigenic pathogen polypeptides are screened for binding by one or more antigen binding molecules by flow cytometry. For example, cells expressing the candidate optimized antigenic pathogen polypeptides may be used in a flow cytometry assay.

[0070] Optionally, the candidate optimized antigenic pathogen polypeptides are screened for binding by one or more broadly neutralizing antigen binding molecules using a first assay (e.g., flow cytometry) and for binding by one or more broadly neutralizing antigen binding molecules using a second assay (e.g., a neutralization assay).

[0071] Optionally, the pathogen is a virus, the candidate optimized antigenic pathogen polypeptide is a candidate optimized antigenic viral polypeptide, and the pathogen peptide is a viral polypeptide.

[0072] Optionally, the polypeptide library is a viral pseudotype library comprising a plurality of different viral pseudotypes, each different viral pseudotype comprising a different candidate optimized antigenic pathogen polypeptide, e.g., a different candidate optimized antigenic viral polypeptide (e.g., a viral glycoprotein).

[0073] Optionally, in step (ii), the candidate optimized antigenic viral polypeptides are screened for binding by one or more broadly neutralizing antigen binding molecules by screening the viral pseudotypes for binding and / or neutralization by one or more antigen binding molecules.

[0074] Pseudotyping is the process of producing a virus or viral vector in combination with a foreign viral envelope protein. The result is a pseudotyped viral particle. Pseudotyped particles cannot transmit phenotypic changes to progeny viral particles because they do not have the genetic material to produce additional viral envelope proteins. A "pseudotype" may be defined as a hybrid viral particle that contains a protein nucleocapsid ("core") that encapsulates a nucleic acid (RNA or DNA) genome, enclosed in a lipid "envelope" membrane that is itself derived from the host cell. This envelope is acquired as the core exits the cell by budding and contains proteins from other viruses. Many of these heterologous envelope proteins are antigenic targets for the host immune system. In pseudotypes, one or more of these envelope proteins may be derived from the virus under study. Many pseudotypes also have a foreign gene, called a "transgene," engineered into their genome. In the presence of a susceptible cell, expression of the transgene ultimately occurs when the envelope protein binds to a cellular receptor that allows entry into the cell. Rhabdoviruses (e.g., varicella-zoster virus, VSV) and retroviruses (e.g., lentiviruses) are widely used as the core for pseudotyping. For retroviruses, the key feature is that they can reverse transcribe their dimeric single-stranded RNA into a double-stranded deoxyribonucleic acid (dsDNA) copy, which is then integrated into the cellular genome through the use of viral and cellular enzymes. For retrovirus pseudotypes, this usually results in the expression of a transduction / reporter gene, the latter of which is easily quantifiable. Reporter gene expression directly correlates with the efficiency of viral envelope / receptor interaction and inversely correlates with whether individual antibody responses or antiviral agents can interfere with the natural viral entry and replication process.

[0075] Binding of viral pseudotypes to broadly neutralizing antigen-binding molecules can be measured using any suitable technique known to those skilled in the art, such as hemagglutinin inhibition (HI) assays, or enzyme-linked immunosorbent assays (ELISAs). ELISA analysis of antibody binding to glycoprotein (GP) is described in Saphire et al., 2018 (Cell 174(4): 938-952) in connection with the analysis of monoclonal antibodies against Ebola virus GP.

[0076] The production of retroviral pseudotypes and their use in pseudotype neutralization assays and immunogenicity testing is reviewed in detail in Temperton et al., 2015 (Retroviral Pseudotypes - From Scientific Tools to Clinical Utility. In: eLS. John Wiley & Sons, Ltd: Chichester. DOI: 10.1002 / 9780470015902.a0021549.pub2).

[0077] Representatives of all seven genera of retroviruses have been used in pseudotyping studies, but only gammaretrovirus or lentivirus pseudotypes have been widely used to date. Lentiviruses are a genus of the Retroviridae family, and unlike gammaretroviruses, they can infect non-proliferating cells, making them amenable to gene therapy applications involving highly differentiated or quiescent cells (e.g., G0 cell cycle phase), including muscle or neurons. The most common lentiviral vector used for pseudotyping is HIV type 1 (HIV-1), but simian immunodeficiency viruses have been used as well.

[0078] Generation of retroviral pseudotypes is accomplished through the simultaneous introduction of a foreign envelope protein gene, a core retroviral gene, and a cloned form of a transgene (e.g., a reporter or therapeutic gene) into producer cells, usually a highly transfectable cell line such as human embryonic kidney (HEK) 293 clone 17T cells (American Type Culture Collection #CRL-11268) (Pear et al., 1993, PNAS USA 90: 8392-8396).

[0079] 1. Envelope plasmid. The envelope gene of the study virus is cloned into an appropriate expression plasmid. The gene is usually derived via polymerase chain reaction amplification of the viral cDNA using specific primers or from custom gene synthesis. Some expression vectors are commercially available and utilize different, usually strong, constitutive gene promoters (e.g., human cytomegalovirus (CMV) immediate early gene), which may affect the efficacy of pseudotype generation.

[0080] 2. Retroviral gag-pol plasmid. The gag and pol genes code for a polyprotein that is then cleaved to release the structural proteins found in the core (including matrix, capsid, and nucleocapsid), and proteins involved in viral replication (protease, reverse transcriptase, and integrase) that are responsible for processing the structural proteins, converting the ssRNA viral genome to dsDNA, and ensuring integration (of the transgene) into the host cell genome. In addition, in lentiviral gag-pol constructs, the rev gene is included. The Rev protein is involved in the transport of viral mRNA from the nucleus to the cytosol for translation.

[0081] 3. Transfer / reporter plasmids: These are genes that are stably integrated into the host cell DNA, from which they are expressed via various cis-acting transcriptional elements. Transfer plasmids contain a packaging signal upstream of the gene to ensure incorporation of the viral RNA containing the gene into the viral core upon pseudotype generation.

[0082] After the cellular machinery transcribes and translates the transfected gene, the RNA dimer of the transgene (the region between the long terminal repeats; LTRs) is incorporated into the pseudotype via the packaging signal. Because the transfer plasmid is the only plasmid engineered to contain the packaging signal, no other nucleic acid is incorporated into the mature pseudotype particle. A domain at the N-terminus of Gag targets the nucleocapsid to the cell plasma membrane, into which the envelope protein is inserted. Pseudotype particles budded from the cell are enclosed in the cell membrane to form the viral envelope.

[0083] Pseudotyped viruses are released into the culture medium of producer cells. This supernatant can be titrated on target cells and the concentration of functional particles can be measured. They bind to cells via envelope protein-receptor interactions followed by membrane fusion and internalization. The pseudotyped genomes carrying the transduced / reporter genes are integrated into the host cell DNA and expressed there. The reporter gene expression level correlates with the level of transduction by viable particles. Since only the transgene is present in the pseudotype, no viral proteins are produced in the target cells and no further pseudotype production and propagation occurs. This provides safety in working with pseudotypes compared to working with wild-type viruses. Pseudotypes based on green fluorescent protein (GFP) are easily titrated using fluorescence microscopy or flow cytometry, luciferase pseudotypes by luminometry and β-galactosidase (β-gal) pseudotypes by color reaction.

[0084] Many standard serological assays measure only antibody binding, not inhibition of virus infectivity (hemagglutinin inhibition (HI) and ELISA). Neutralization assays allow sensitive detection of functional antibody responses. However, for highly contained viruses (e.g., Ebola), these assays are not widely applicable due to the need for high biosafety laboratories and specially trained personnel. The use of retroviral and lentiviral particles pseudotyped with the envelope of a pathogen as a "surrogate virus" for use in neutralization assays is one way to circumvent this problem. When using a pseudotype strategy, only the viral envelope proteins are required, with no possibility of recombinant or natural virus escape. These pseudotypes are replication-deficient and cannot produce replication-competent progeny.

[0085] Pseudotypes are excellent serum reagents for virus neutralization assays because the virions can contain reporter genes and have heterologous viral envelope proteins on their surface. The introduction of these reporter genes into target cells depends on the function of the viral envelope proteins, and therefore the titer of neutralizing antibodies against the envelope can be measured by reporter gene introduction and reduction of expression. PV neutralization assays have now been developed for a wide range of RNA viruses from multiple virus families (Temperton et al., see Table 1 above).

[0086] Pseudotype-based influenza neutralization assays have been shown to be highly efficient for measuring broadly neutralizing antibodies, making them ideal serological tools to study cross-reactive responses to multiple subtypes with pandemic potential (Corti et al., 2011, Science 333 (6044): 850-856).

[0087] Production of lentiviral vectors pseudotyped with filovirus glycoproteins is described in Sinn et al., 2017 (Methods Mol Biol. 2017;1628:65-78).

[0088] An example of a suitable general method for producing viral pseudotypes is as follows:

[0089] For transfection, 5 × 10 6 24 hours after seeding HEK-293T cells, a complex containing plasmid DNA and PEI, which facilitates the delivery of DNA into the cells, is added. Retroviral gag-pol plasmids and reporter plasmids are co-transfected with the required envelope plasmids.

[0090] An example of a suitable neutralization assay is as follows:

[0091] 1 x 10 in a 96-well plate 5 Approximately 100 × TCID50 of pseudotyped virus, which results in a relative light unit (RLU) output of 1 × 10, was incubated with dilutions of serum for 1 h at 37% (5% CO2) followed by 1 × 10 4 Target cells were added. These were incubated for a further 48 hours, after which the medium was removed and replaced with a 50:50 mixture of fresh medium and luciferase reagent. Luciferase activity was detected after 2.5 minutes by reading the plate on a luminometer. For all results, background RLU (virus alone or DEnv) was estimated prior to analysis.

[0092] Saphire et al. (supra) describe three independent assays for the evaluation of mAb neutralization in connection with the analysis of monoclonal antibodies against Ebola virus GP: i) biologically contained EBOV(ΔVP30) (Halfmann et al., 2008, Proc Natl Acad Sci USA. 2008; 105:1129-1133); and ii) authentic EBOV performed under BSL-2+, BSL-3, and BSL-4 containment; and iii) Replication-competent vesicular stomatitis virus (rVSV) harboring the EBOV GP.

[0093] Neutralization of Ebola ΔVP30-RenLuc virus An Ebola virus in which the reporter gene Renilla luciferase replaces the viral transcription factor VP30 (Ebola ΔVP30-RenLuc virus) was used to complement a Vero cell line stably expressing VP30 in trans (Vero VP30), thus allowing analysis at BSL-3 (Halfmann et al., 2008). A total of 5 × 10 of Ebola ΔVP30-RenLuc virus diluted in minimal essential medium containing 2% fetal bovine serum was used to infect the cells. 3 The focus forming units are incubated with 50 μg / ml monoclonal antibody for 3 hours at 37° C. The virus / antibody mixture is inoculated onto 9×10 cells in a 96-well plate at a multiplicity of infection (MOI) of 0.001. 3 100 / well are added to Vero VP30 cells seeded the day before and incubated at 37°C and 5% CO2 for 3 days. If used, guinea pig complement (Cedarlane) is added to minimal essential medium at a final concentration of 10%. EnduRen (Promega), a live cell luciferase substrate, is then incubated with the cells for 3 hours, after which luciferase values ​​are measured as relative light units (RLU) using a Tecan M1000 plate reader (Tecan). Assays are performed in duplicate, with known neutralizing (GP 133 / 3.16) and non-neutralizing monoclonals (VP35 5 / 69.3.2) used as positive and negative controls, respectively. Antibodies that neutralize the luciferase signal by ≥95% are defined as strong neutralizers, 50%-94% inhibition of the luciferase signal is considered to be moderate neutralizers, and neutralizers with 49% or lower inhibition are classified as weak / non-neutralizers.

[0094] Neutralization of authentic EBOV Assays to assess neutralization of authentic EBOV are performed according to the method described by Holtsberg et al. (Holtsberg et al., 2015, J Virol. 2015; 90:266-278). Vero E6 cells were plated in the inner 60 wells of a black 96-well plate at 2.5 × 10 cells / well. -4 Virus infection is performed 24 hours after seeding at 1000 cells / well. Antibodies are serially diluted twice in Vero growth medium (Eagle's minimum essential medium with Earle's salts and L-glutamine, 5% fetal bovine serum (FBS) and 1% penicillin-streptomycin) to obtain the desired final concentration (50 μg / ml), mixed with an equal amount of live EBOV and incubated at 37°C for 1 hour with mixing every 15 minutes. The antibody / virus mixture is then added to the Vero cells at an MOI of 0.2, incubated at 37°C for 1 hour, washed with PBS, and growth medium alone is added to all wells and the plate is incubated at 37°C for an additional 48 hours. Cells are then fixed with 10% neutral buffered formalin and the percentage of infected cells is determined by indirect immunofluorescence assay using EBOV-specific human mAb KZ52 and goat anti-human IgG conjugated to Alexa Fluor 488 (Molecular Probes) as the secondary antibody. Images are acquired with 20 fields / well with a 20x magnification objective in an Operetta high content imaging system (Perkin-Elmer). Operetta images are analyzed by a custom algorithm built from image analysis functions available in Harmony software (Perkin-Elmer). The percentage of inhibition for each antibody is determined relative to control cells incubated with media alone. Antibodies that reduced the percentage of infected cells by >80% were classified as strong neutralizers, while antibodies that reduced infection by 50%-79% and <50% were considered to be moderate and weak / non-neutralizers, respectively.

[0095] Neutralization of rVSV-EBOV GP A recombinant vesicular stomatitis virus (VSV) expressing both eGFP and a recombinant surface GP in place of VSV G (rVSV-EBOV) has been described previously (Wec et al. (e.g., Science; 354:350-354; Wong et al., 2010, Virol.; 84:163-175). For neutralization assays, Vero cells were cultured at 6.0 × 10 4 Cells were seeded at 100 cells / well and cultured overnight at 37°C and 5% CO2 in Eagle's Minimum Essential Medium (EMEM) supplemented with 10% fetal bovine serum (FBS) and 100 I.U. / ml penicillin and 100 µg / ml streptomycin. The next day, the virus is used to infect Vero cell layers in 96-well plates after 1 h of incubation at room temperature with serial 3-fold dilutions of antibodies starting at 330 nM (approximately 50 µg / ml) in serum-free EMEM. The amount of virus used for infection is determined based on titration of virus stocks to achieve 35-50% final infection (MOI approximately 0.1 infectious units / cell) in control wells without antibody. Virus was incubated with cells in 50% v / v EMEM supplemented with 2% FBS, 100 I.U. / ml penicillin and 100 μg / ml streptomycin for 14-16 h at 37 °C and 5% CO2, after which cells were fixed and nuclei stained with Hoescht. rVSV infectivity is measured by counting EGFP-positive cells compared to the total number of cells indicated by nuclear staining using a Cellinsight CX5 automated microscope and accompanying software (Thermo Scientific). Infection levels in control wells lacking antibody are set to 100% and infections are normalized to that value for each antibody dilution tested in triplicate. Mean values ​​are determined and a full 9-point dilution curve is used to calculate the half-maximal inhibitor concentration IC 50 is determined using GraphPad Prism version 6. IC 50 Antibodies with a Cv of ≦5 nM are considered to be strong neutralizers, and Cv of ≦5 nM are considered to be strong neutralizers. <IC 50Antibodies with <50 nM and ≦50 nM are considered to be moderate neutralizers and weak / non-neutralizers, respectively. The non-neutralizing fraction, an indicator of antibody potency, is also determined by using the antibody at the highest concentration tested, 330 nM, and measuring the GFP signal compared to that of untreated control cells. Fractions that reduce the signal by ≧98%, 50-98%, and less than 50% are considered to be strong, moderate, and weak / non-neutralizers, respectively.

[0096] Methods for screening polypeptide libraries are described in Bruun et al. (PLoS ONE 9(10): e109196).

[0097] Optionally, the method of the invention further comprises the step of generating a polypeptide library.

[0098] Optionally, the polypeptide library is generated by expressing different candidate optimized antigenic pathogen polypeptides from a nucleic acid library comprising a plurality of different nucleic acids, each different nucleic acid comprising a nucleotide sequence that encodes a different candidate optimized antigenic pathogen polypeptide of the polypeptide library.

[0099] Optionally, the different candidate optimized pathogen polypeptides are expressed in or on the surface of mammalian cells. Suitable methods are well known to those of skill in the art.

[0100] Optionally, the nucleotide sequence of each different nucleic acid of the nucleic acid library is optimized for expression of the encoded polypeptide in a mammalian cell.

[0101] Optionally, each different nucleic acid of the nucleic acid library is part of an expression vector for expression of the nucleic acid in a mammalian cell.

[0102] Optionally, the pathogen is a virus, the candidate optimized antigenic pathogen polypeptide is a candidate optimized antigenic viral polypeptide, and the pathogen peptide is a viral polypeptide.

[0103] Optionally, the nucleic acid library is a viral pseudotype vector library, wherein each different nucleic acid of the library is part of an expression vector for producing a viral pseudotype comprising an encoded viral polypeptide, and the polypeptide library is a viral pseudotype library generated by producing viral pseudotypes from the expression vectors of the viral pseudotype vector library, wherein the viral pseudotype library comprises a plurality of different viral pseudotypes, each different viral pseudotype comprising a different candidate optimized viral polypeptide encoded by a different nucleic acid sequence of the viral pseudotype vector library.

[0104] Optionally, the viral pseudotype vector library is at least 2, 3, 5, 10, 20, 30, 40, 50, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , or 10 9 Contains distinct members.

[0105] Optionally, the expression vector is also a vaccine vector.

[0106] Examples of vaccine vectors include a viral vaccine vector, a bacterial vaccine vector, an RNA vaccine vector, or a DNA vaccine vector.

[0107] Viral vaccine vectors use live viruses to deliver nucleic acid (e.g., DNA or RNA) to human or non-human animal cells. The nucleic acid contained in the virus encodes one or more antigens that elicit an immune response when expressed in infected human or non-human animal cells. Both humoral and cell-mediated immune responses can be induced by viral vaccine vectors. Viral vaccine vectors combine many of the positive properties of nucleic acid vaccines with those of live attenuated vaccines. Like nucleic acid vaccines, viral vaccine vectors deliver nucleic acid to the host to produce antigenic proteins that can be tailored to stimulate a broad range of immune responses, including antibodies. Helper T cells (CD4 + T cells), and cytotoxic T lymphocytes (CTL, CD8 + T cells) mediated immunity. Unlike nucleic acid vaccines, viral vaccine vectors also have the potential to actively invade and replicate in host cells, further activating the immune system like an adjuvant, much like live attenuated vaccines. Thus, viral vaccine vectors generally contain live attenuated viruses genetically engineered to carry nucleic acids (e.g., DNA or RNA) that code for protein antigens from unrelated organisms. Although viral vaccine vectors can generally generate stronger immune responses than nucleic acid vaccines, in some diseases viral vectors are used in combination with other vaccine technologies in a strategy called heterologous prime-boost. In this system, one vaccine is given as a priming step, and then vaccinated with an alternative vaccine as a booster. The heterologous prime-boost strategy aims to provide a stronger overall immune response. Viral vaccine vectors can be used as both prime and boost vaccines as part of this strategy. Viral vaccine vectors are reviewed in Ura et al., 2014 (Vaccines 2014, 2, 624-641) and Choi and Chang, 2013 (Clinical and Experimental Vaccine Research 2013;2:97-105).

[0108] Optionally, the viral vaccine vector is based on a viral delivery vector, such as a poxvirus (e.g., Modified Vaccinia Ankara (MVA), NYVAC, AVIPOX), herpesvirus (e.g., HSV, CMV, adenovirus of any host species), morbillivirus (e.g., measles), alphavirus (e.g., SFV, Sendai), flavivirus (e.g., yellow fever), or rhabdovirus (e.g., VSV), a bacterial delivery vector (e.g., Salmonella, E. coli), an RNA expression vector, or a DNA expression vector.

[0109] Optionally, the vector is a pEVAC-based expression vector. The pEVAC expression vector is described in more detail in Example 7 below.

[0110] In other embodiments, different candidate optimized antigenic pathogen polypeptides are expressed in or on the surface of bacteria, yeast, or insect cells.

[0111] Optionally, the method of the invention further comprises generating a nucleic acid library by synthesizing a plurality of different nucleic acids, each different nucleic acid comprising a different nucleotide sequence encoding a different candidate optimized antigenic pathogen polypeptide.

[0112] Optionally, the method of the present invention further comprises the steps of: i) obtaining amino acid sequences of pathogen polypeptides of different pathogen isolates and / or nucleotide sequences encoding the pathogen polypeptides; and ii) generating a plurality of different nucleotide sequences, each different nucleotide sequence encoding a different candidate optimized antigenic pathogen polypeptide, and the encoded amino acid sequence of each different candidate optimized antigenic pathogen polypeptide is optimized from the obtained or encoded amino acid sequence of the pathogen polypeptide and is different from each of the obtained or encoded amino acid sequences.

[0113] Optionally, generating a plurality of different nucleotide sequences in step (ii) above comprises the steps of: performing a multiple sequence alignment of the amino acid or nucleotide sequences obtained in step (i) above; identifying from the multiple sequence alignment an amino acid sequence or encoded amino acid sequence that is highly conserved among the polypeptides of the different pathogen isolates; and generating a plurality of different nucleotide sequences, each different nucleotide sequence encoding a different candidate optimized antigenic pathogen polypeptide, wherein one or more of the different nucleotide sequences comprises a sequence that encodes the highly conserved amino acid sequence or encoded amino acid sequence identified from the multiple sequence alignment.

[0114] An amino acid sequence that is highly conserved among polypeptides of different pathogen isolates or an encoded amino acid sequence can be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 150, 200, 250, 300, 400, 500, 600, 700, or 800 amino acid residues in length.

[0115] Optionally, the number of amino acid sequences of pathogen polypeptides, or the number of nucleotide sequences encoding pathogen polypeptides, of different pathogen isolates may be at least 3, 4, 5, 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 10 6 , 10 9 , or 10 12 Typically, the more sequences used for the multiple sequence alignment, the better.

[0116] Optionally, the method of the present invention further comprises the steps of identifying an amino acid sequence or an encoded amino acid sequence that is an ancestral amino acid sequence from the multiple sequence alignment; and including a sequence encoding the ancestral amino acid sequence identified from the multiple sequence alignment in one or more of the different generated nucleotide sequences.

[0117] Inclusion of one or more nucleotide sequences encoding ancestral amino acid sequences may be advantageous, since ancestral amino acid sequences that are highly conserved with respect to extant amino acid sequences are expected to be structurally and / or functionally important for the survival and / or reproduction of pathogens.Similarly, pathogen isolates may be highly diverse (especially isolates of emerging or re-emerging pathogens, such as emerging or re-emerging RNA viruses), and the evolutionary distance between these two pathogen populations may be large, so a vaccine designed to work against a pathogen population in one patient will not work in a different patient.However, their most recent common ancestors are more closely related to each other than the two pathogen populations are to each other.Thus, a vaccine designed with respect to a common ancestor will be more likely to be effective against a larger population of circulating strains.

[0118] Ancestral sequence reconstruction (ASR) is discussed in Randall et al. (Nat. Commun. 7:12847 doi: 10.1038 / ncomms 12847 (2016)). The authors define ASR as "the process of analysis of modern sequences in an evolutionary / phylogenetic tree context to infer ancestral sequences at specific nodes of the tree." Ancestral sequence reconstruction (ASR) is used in the study of molecular evolution. Unlike traditional evolutionary approaches to study proteins, ASR probes vertically for statistically inferred ancestral proteins within the nodes of the tree by horizontal comparison of related protein homologs from the termini of different branches of the phylogenetic tree (see Figure 1). A phylogenetic tree is a branching diagram that shows the evolutionary relationships in various biological species or other entities based on similarities and differences in their physical or genetic characteristics. In a rooted phylogenetic tree, each node with descendants represents the inferred immediate common ancestor of those descendants. In ASR, several related homologs of a protein of interest are selected, aligned in a multiple sequence alignment (MSA), and a phylogenetic tree is constructed with statistically inferred sequences at the branch nodes. These sequences are the so-called "ancestors." The process of synthesizing the corresponding DNA, transforming it into cells, and producing the protein is the so-called "reconstruction."

[0119] Ancestral sequences are typically calculated by maximum likelihood methods, but Bayesian methods are also implemented. Since ancestry is inferred from phylogenies, the phylogenetic morphology and composition play a major role in the ASR sequence as output. ASR does not require the actual sequence of the ancestral protein / DNA to be recreated, but rather sequences that are likely to be similar to those at the nodes. Maximum likelihood (ML) methods work by generating sequences in which the residue at each position is predicted to be most likely to occupy that position by the inference method used. Typically, this is a score matrix (similar to that used in BLAST or MSA) calculated from extant sequences. Alternative methods include maximum parsimony (MP), which builds sequences based on a sequence evolution model, the idea being that those with the fewest number of nucleotide sequence changes usually represent the most efficient and most likely path for evolution to occur. MP is often considered the least reliable reconstruction method, as it oversimplifies evolution to a degree that is probably not applicable on a billion-year scale. Other methods include Bayesian methods, which include consideration of residue uncertainty. Such methods are sometimes used to complement ML, but typically produce more ambiguous sequences (i.e., sequences containing residue positions where no clear substitutions can be predicted). In such instances, they often produce several ASR sequences that encompass most of the ambiguity relative to each other.

[0120] The ASR method and algorithms were reviewed in Joy et al., 2016, PLOS The following is a more detailed description based on the description in Computational Biology 12(7): DOI:10.1371 / journal.pcbi.1004763.

[0121] If necessary, the ASR should be at least 3, 4, 5, 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 10 6 , 10 9 , or 10 12different sequences. In some instances, the more sequences used, the better.

[0122] Optionally, each of the sequences used for the multiple sequence alignment is a full-length sequence of a pathogen polypeptide of a pathogen isolate.

[0123] Any attempt at ancestry reconstruction begins with a phylogeny. In general, a phylogeny is a tree-based hypothesis about the order in which populations (called taxa) are related in descent from a common ancestor. The observed taxa are represented by the tips or terminal nodes of the tree, which are progressively joined by branches to their common ancestor, which is usually represented by a branching point in the tree called an ancestral node or endjunction. Eventually all lineages converge to the most recent common ancestor of the entire sample of taxa. In the context of ancestry reconstruction, phylogenies are often treated as if they were known quantities (Bayesian approaches being an important exception). Since there can be a vast number of phylogenies that are roughly equally valid for explaining the data, it can be a convenient and sometimes necessary simplifying assumption to reduce the subset of phylogenies supported by the data to a single representative or point estimate. An ancestry reconstruction can be thought of as the direct result of applying a hypothetical evolutionary model to a given phylogeny. When a model contains one or more free parameters, the overall goal is to estimate these parameters based on measured features in observed taxa (sequences) that are descendants of a common ancestor. Parsimony methods are an important exception to this paradigm. They are based on the empirical rule that trait-state changes are rare, without making any attempt to quantify that rarity.

[0124] maximum savings method Parsimony refers to the principle of selecting the simplest of competing hypotheses. In the context of ancestry reconstruction, parsimony strives to find a distribution of ancestral states in a given tree that minimizes the total number of trait-state changes required to explain the states observed at the tip of the tree. This maximum parsimony method is one of the earliest formulated algorithms for reconstructing ancestral states. Maximum parsimony can be implemented by one of several algorithms. One of the earliest examples is Fitch's method (Fitch WM. Toward defining the course of evolution: minimum change for a specific tree topology. Systematic Biology. 1971;20(4):406-16), which assigns ancestral trait states parsimoniously via two traversals of a rooted binary tree. The first stage is a post-order traversal that proceeds from the tip of the tree toward the root by visiting descendant (child) nodes before their parents. First, a set of possible trait states S for the i-th ancestor is determined based on the observed trait states of its descendants. iEach assignment is the intersection of the trait states of the ancestors and their descendants. If the intersection is the empty set, it is a union. The latter case means that a trait state change has occurred between the ancestor and one of its two direct descendants. Each such event adds to a cost function of the algorithm that can be used to distinguish between alternative trees based on maximum parsimony. A preorder traversal of the tree proceeding from the root towards the tips is then performed. A trait state is then assigned to each descendant based on the trait state it shares with its parent. Since the root has no parent nodes, it may be necessary to choose trait states arbitrarily, especially if more than one possible state has been reconstructed at the root. Parsimony is intuitively preferable and very efficient, so in some cases it is still used to seed ML optimization algorithms with initial phylogenies (Stamatakis A. RAxML-VI-HPC: maximum likelihood-based phylogenetic analyses with thousands of taxa and mixed models. Bioinformatics. 2006;22:2688-90. pmid:16928733). However, they suffer from several problems: 1. Variation in evolutionary rates. Fitch's method assumes that changes between all trait states are equally likely; therefore, any change incurs the same cost for a given tree. This assumption is often unrealistic and can limit the accuracy of such methods. For example, in the evolution of nucleic acids, transitions tend to occur more frequently than transversions. This assumption can be relaxed by assigning different costs to state changes of a particular trait, resulting in a weighted parsimony algorithm (Sankoff D. Minimal mutation trees of sequences. SIAM Journal on Applied Mathematics. 1975;28(1):35-42). 2. Rapid evolution. A consequence of the "minimum evolution" heuristic that underlies such methods is that they assume that variation is rare and are therefore inappropriate when variation is the rule rather than the exception (Schluter D, Price T, Mooers AO, Ludwig D. Likelihood of ancestor states in adaptive radiation. Evolution. 1997;51(6):1699-711; Felsenstein J. Maximum likelihood and minimum-steps methods for estimating evolutionary trees from data on discrete characters. Systematic Biology. 1973;22(3):240-9). 3. Variation in time among lineages. Parsimony methods implicitly assume that the same amount of evolutionary time has passed along all branches of the tree. Because of this, they do not take into account the variation in branch lengths within a tree, which is often used to quantify the passage of evolutionary or chronological time. Due to this limitation, the techniques tend to estimate, for example, that one change occurred on a very short branch, rather than multiple changes on a very long branch. This shortcoming is addressed by model-based methods (both ML and Bayesian) that estimate the stochastic process of evolution as it unfolds along each branch of the tree (Li G, Steel M, Zhang L. More taxa are not necessarily better for the reconstruction of ancestral character states. Systematic biology. 2008;57(4):647-53). 4. Statistical justification. Even if there is no statistical model underlying the method, the estimates do not have clear uncertainties.

[0125] Maximum likelihood method (ML) ML methods of ancestral state reconstruction treat trait states at internal nodes of the tree as parameters and attempt to find parameter values ​​that maximize the probability of the data (observed trait states) given a hypothesis (model of evolution and phylogeny related to the observed sequences or taxa). Some of the earliest ML approaches to ancestry reconstruction were developed in the context of gene sequence evolution (Yang Z, Kumar S, Nei M. A new method of inference of ancestral nucleotide and amino acid sequences. Genetics. 1995;141(4):1641-50; Koshi JM, Goldstein RA. Probabilistic reconstruction of ancestral protein sequences. Journal of Molecular Evolution. 1996;42(2):313-20). Similar models have been developed for the analogous case of discrete character evolution (Pagel M. The maximum likelihood approach to reconstructing ancestral character states of discrete characters on phylogenies. Systematic biology. 1999;48(3):612-22).

[0126] These approaches use the same probabilistic framework used to infer phylogenetic trees (Felsenstein J. Evolutionary trees from DNA sequences: a maximum likelihood approach. Journal of molecular evolution. 1981;17(6):368-76). Briefly, the evolution of gene sequences is modeled by a time-reversible continuous-time Markov process. In the simplest of these, all traits undergo independent state transitions (e.g., nucleotide substitutions) at a constant rate over time. This basic model is often extended to allow different rates at each branch of the tree. Indeed, mutation rates may also vary over time (e.g., due to changes in the environment); this can be modeled by allowing a rate parameter to evolve along the tree rather than increasing the number of parameters. The model defines transition probabilities from state i to j along branches of length t (in units of evolutionary time). The likelihood of the phylogeny is calculated from a nested sum of transition probabilities that correspond to the hierarchical structure of the proposed tree. At each node, the likelihoods of its descendants are summed over all ancestral trait states possible at that node:

number

[0127] The problem of ancestry reconstruction is not to compute the overall likelihood of alternative trees, but to find the combination of trait states at each ancestral node with the best marginal ML. In general, there are two approaches to this problem. First, one can work upwards from the descendants of the tree, progressively assigning to each ancestor the trait state that is most likely considering only its direct descendants. This approach is called marginal reconstruction. This is similar to a greedy algorithm that makes locally optimal choices at each stage of an optimization problem. This can be very efficient, but is not guaranteed to reach a global optimum solution to the problem. One can then instead try to find a joint combination of ancestral trait states across the trees that jointly maximizes the likelihood of the data. Thus, this approach is called joint reconstruction. Because it is not as quick as marginal reconstruction, it is less likely to be captured by a local optimum in the non-convex objective functions that modern optimization methods and heuristics are designed to avoid. In the context of ancestry reconstruction, this means that marginal reconstructions may assign trait states to direct descendants that fall outside the congruent distribution of ancestral trait states that is a local optimum but far from the global optimum. Congruent reconstructions are more computationally complex than marginal reconstructions. Nevertheless, efficient algorithms for congruent reconstructions have been developed with time complexity that is roughly linear with the number of observed taxa or sequences.

[0128] ML-based methods of ancestry reconstruction tend to provide greater accuracy than MP methods in the presence of variation in evolutionary rates among traits (or among sites in a genome). However, these methods are still unable to accommodate variation in evolutionary rates over time, also known as heterotachy. If the evolutionary rate of a particular trait accelerates on a phylogenetic branch, the amount of evolution that has occurred on that branch will be underestimated for a given length of the branch, assuming that the evolutionary rate of that trait is constant. In addition, heterotachy can be difficult to distinguish from variation among traits in evolutionary rates.

[0129] Because ML (unlike maximum parsimony) requires the researcher to specify a model of evolution, its accuracy can be affected by the use of a grossly inaccurate model (model misspecification). Furthermore, ML can only provide a single reconstruction of the trait state (often called a "point estimate"); when the likelihood surface is highly non-convex and contains multiple peaks (local optima), a single point estimate may not provide an adequate representation and a Bayesian approach may be more appropriate.

[0130] Bayesian Inference Bayesian inference uses the likelihood of observed data to update the researcher's thinking, or prior distributions, to produce a posterior distribution. In the context of ancestry reconstruction, the goal is to estimate the posterior probability of the ancestral trait state at each internal node of a given tree. Furthermore, these probabilities can be integrated over the parameters of the evolutionary model and the posterior distribution over the space of all possible trees. This is an application of Bayes' theorem:

number

[0131] One of the first implementations of a Bayesian approach to ancestral sequence reconstruction was developed by Yang and colleagues, who defined the prior distribution using ML estimates of evolutionary models and trees, respectively. Thus, their approach is an example of an empirical Bayes method for calculating posterior probabilities of ancestral trait states, which was first implemented in the software package PAML (Yang Z. PAML 4: phylogenetic analysis by maximum likelihood. Molecular biology and evolution. 2007;24(8):1586-91). With respect to the formulation of Bayes' rule above, the empirical Bayes method fixes on the empirical estimates of the models and trees obtained from the data, effectively removing them from the posterior likelihoods, and the prior term in the equation. Moreover, Yang and colleagues (Yang Z, Kumar S, Nei M. A new method (Empirical Bayes of inference of ancestral nucleotide and amino acid sequences. Genetics. 1995;141(4):1641-50) used in the denominator the empirical distribution of site patterns in that alignment of observed nucleotide sequences (i.e., the assignment of nucleotides to the tips of the tree) instead of exhaustively calculating P(D) for all possible values ​​of S given θ. Computationally, empirical Bayes methods are similar to ML reconstructions of ancestral states, except that rather than searching for ML assignments of states based on the respective probability distributions at each internal junction, the probability distributions themselves are reported directly.

[0132] Empirical Bayesian methods for ancestry reconstruction require researchers to assume that the parameters of the evolutionary model and the tree are known without error. When the size or complexity of the data makes this an unrealistic assumption, it may be more prudent to adopt a fully hierarchical Bayesian approach and estimate a joint posterior distribution for ancestral trait states, models, and trees (Huelsenbeck JP, Bollback JP. Empirical and hierarchical Bayesian estimation of ancestral states. Systematic Biology. 2001;50(3):351-66). Huelsenbeck and Bollback first proposed a hierarchical Bayesian approach to ancestry reconstruction by sampling ancestral sequences from this joint posterior distribution using Markov chain Monte Carlo (MCMC) methods. A similar approach has also been used to reconstruct the evolution of fungal species' symbiosis with algae (lichenification) (Lutzoni F, Pagel M, Reeb V. Major fungal lineages are derived from lichen symbiotic ancestors. Nature. 2001;411(6840):937-40). For example, the MCMC Metropolis-Hastings algorithm searches for a joint posterior distribution by accepting or rejecting parameter assignments based on the ratio of posterior probabilities.

[0133] Thus, the empirical Bayes approach calculates the probability of various ancestral states for a particular tree and evolutionary model. By expressing the reconstruction of ancestral states as a set of probabilities, one can directly quantify the uncertainty in assigning any particular state to an ancestry. On the other hand, the hierarchical Bayes approach averages these probabilities over all possible trees and evolutionary models in proportion to how similar these trees and models are given the observed data.

[0134] A fully Bayesian approach is limited to the analysis of a relatively small number of sequences or taxa because the space of all possible trees quickly becomes so vast that it becomes computationally infeasible for linked samples to converge in a reasonable amount of time.

[0135] Pathogens, especially emerging or re-emerging pathogens, such as emerging or re-emerging RNA viruses, evolve at a very rapid rate, and by orders of magnitude more rapid than mammals or birds. For these organisms, ancestry reconstruction can be applied on a fairly short time scale, for example reconstructing the global or local precursors of epidemics over decades rather than millions of years. It has been proposed to use such reconstructed strains as targets for vaccine design efforts in comparison with sequences isolated from today's patients (Gaschen et al., Science. 2002;296(5577):2354-60).

[0136] According to method embodiments of the present invention, any suitable method of ARS can be used to identify an amino acid sequence that is an ancestral amino acid sequence or an encoded amino acid sequence from a multiple sequence alignment.

[0137] Optionally, identifying the ancestral amino acid sequence from the multiple sequence alignment includes performing maximum parsimony ancestral sequence reconstruction (MP-ASR).

[0138] Optionally, identifying the ancestral amino acid sequence from the multiple sequence alignment comprises performing maximum likelihood ancestral sequence reconstruction (ML-ASR).

[0139] Optionally, identifying the ancestral amino acid sequence from the multiple sequence alignment comprises performing Bayesian Inference Ancestral Sequence Reconstruction (BI-ASR).

[0140] There are many software packages available that perform ancestral sequence reconstruction. The following table (Joy et al., 2016, PLOS Computational Biology 12(7): DOI:10.1371 / journal.pcbi.1004763) provides a representative sample of the wide variety of packages that implement methods of ancestry reconstruction with different strengths and peculiarities. [Table 5]

[0141] Most of these software packages are designed to analyze genetic sequence data. For example, PAML (Yang Z. PAML 4: phylogenetic analysis by maximum likelihood. Molecular biology and evolution. 2007;24(8):1586-91) is a collection of programs for phylogenetic analysis of DNA and protein sequence alignments by ML. Ancestry reconstruction can be performed using the codeml program. HyPhy, Mesquite, and MEGA are also software packages for phylogenetic analysis of sequence data, but are designed to be more modular and customizable. HyPhy (Pond SLK, Muse SV. HyPhy: hypothesis testing using phylogenies. Statistical methods in molecular biology and evolution. 2007;24(8):1586-91) is a collection of programs for phylogenetic analysis of DNA and protein sequence alignments by ML. Ancestry reconstruction can be performed using the codeml program. HyPhy, Mesquite, and MEGA are also software packages for phylogenetic analysis of sequence data, but are designed to be more modular and customizable. evolution: Springer; 2005. p. 125-81) implements a simultaneous ML method for ancestral sequence reconstruction that can be easily adapted to reconstruct more generalized ranges of discrete ancestral trait states, such as geographic locations, by specifying customized models in its batch language (Pupko T, Pe I, Shamir R, Graur D. A fast algorithm for joint reconstruction of ancestral amino acid sequences. Molecular Biology and Evolution. 2000;17(6):890-6). Mesquite(Maddison W, Maddison D. Mesquite: a modular system for evolutionary analysis. 2.75 ed20011) provides ancestral state reconstruction methods for both discrete and continuous traits using both maximum parsimony and ML methods, as well as several visualization tools for interpreting the results of ancestry reconstruction. MEGA (Tamura K, Dudley J, Nei M, Kumar S. MEGA4: molecular evolutionary genetics analysis (MEGA) software version 4.0. Molecular biology and evolution. 2007;24(8):1596-9) is also a modular system, but with a stronger emphasis on ease of use than on customization of the analysis. As of version 5, MEGA allows users to reconstruct ancestral states using maximum parsimony, ML, and empirical Bayes methods.

[0142] Bayesian analysis of gene sequences may confer greater robustness against model misspecification. MrBayes (Huelsenbeck JP, Ronquist F. MRBAYES: Bayesian inference of phylogenetic trees. Bioinformatics. 2001;17(8):754-5) allows estimation of ancestral states at ancestral nodes using a full hierarchical Bayesian approach. The PREQUEL program distributed in the PHAST package performs comparative evolutionary genomics using ancestral sequence reconstructions (Hubisz MJ, Pollard KS, Siepel A. PHAST and RPHAST: phylogenetic analysis with space / time models. Briefings in bioinformatics. 2011;12(1):41-51). SIMMAP probabilistically maps phylogenetic variation (Bollback JP. SIMMAP: stochastic character mapping of discrete traits on phylogenies. BMC bioinformatics. 2006;7(1):88). BayesTraits (Pagel M. The maximum likelihood approach to reconstructing ancestral character states of discrete characters on phylogenies. Systematic biology. 1999;48(3):612-22) analyzes discrete or continuous traits in a Bayesian framework, evaluates evolutionary models, reconstructs ancestral states, and detects evolutionary correlations between pairs of traits.

[0143] Other software packages are more targeted at the analysis of qualitative and quantitative traits (phenotypes). For example, the ape package in the statistical computing environment R (Paradis E. Analysis of phylogenetics and evolution with R. New York: Springer; 2006) provides ancestral state reconstruction methods for both discrete and continuous traits through the ace function, which includes ML. Note that ace performs reconstruction by calculating scaled conditional likelihoods instead of marginal or joint likelihoods used by other ML-based ancestry reconstruction methods, which may negatively affect the accuracy of reconstruction at nodes other than the root. In addition to ML methods for reconstructing ancestral gene sequences, Phyrex implements a maximum parsimony-based algorithm for reconstructing ancestral gene expression profiles (by wrapping around the baseml function of PAML) (Rossnes R, Eidhammer I, Liberles DA. Phylogenetic reconstruction of ancestral character states for gene expression and mRNA splicing data. BMC bioinformatics. 2005;6(1):127).

[0144] Several software packages also reconstruct phylogeographies. BEAST (Bayesian Evolutionary Analysis by Sampling Trees) (Bouckaert R, Heled J, Kuehnert D, Vaughan T, Wu C- H, Xie D, et al. BEAST 2: a software platform for Bayesian evolutionary analysis. PLoS Comput Biol. 2014;10(4):e1003537) provides tools for reconstructing ancestral geographic locations from observed sequences annotated with location data, using Bayesian MCMC sampling methods. Diversitree (FitzJohn RG. Diversitree: comparative phylogenetic The analyses of diversification in R. Methods in Ecology and Evolution. 2012;3(6):1084-92 were carried out using Mk2 (a continuous-time Markov model of binary trait evolution) (Pagel M. Detecting Correlated Evolution on Phylogenies-a General- Method for the Comparative-Analysis of Discrete Characters. Proceedings of the Royal Society of London Series B-Biological Sciences. 1994;255(1342):37-45)) and ancestral state reconstruction under the BiSSE model. Lagrange performs analysis on the evolutionary reconstruction of geographic ranges of phylogenetic trees (Ree RH, Smith SA. Maximum likelihood inference of geographic range evolution by dispersal, local extinction, and cladogenesis. Systematic Biology. 2008;57(1):4-14). Phylomapper (Lemmon AR, Lemmon EM. A likelihood framework for estimating phylogeographic history on a continuous landscape. Systematic Biology. 2008;57(4):544-61) is a statistical framework for estimating past patterns of gene flow and ancestral geographic location. RASP (Yu Y, Harris AJ, Blair C, He X. RASP (Reconstruct Ancestral State in Phylogenies): a tool for historical biogeography. Molecular Phylogenetics and Evolution. 2015;87:46-9) estimates ancestral states using statistical DIVA, Lagrange, Bayes-Lagrange, BayArea, and BBM methods. VIP(Arias JS, Szumik CA, Goloboff PA. Spatial analysis of vicariance: a method for using direct geographical information in historical biogeography. Cladistics. 2011;27(6):617-28) infers past biogeography by examining separate geographic distributions.

[0145] Genome resequencing provides valuable information for interspecies comparative genomics. ANGES (Jones BR, Rajaraman A, Tannier E, Chauve C. ANGES: reconstructing ANcestral GEnomeS maps. Bioinformatics. 2012;28(18):2388-90) compares extant related genomes through genetic marker ancestry reconstruction. BADGER (Larget B, Kadane JB, Simon DL. A Bayesian approach to the estimation of ancestral genome arrangements. Molecular phylogenetics and evolution. 2005;36(2):214-23) uses a Bayesian approach to examine the history of gene rearrangements. Count (Csuos M. Count: evolutionary analysis of phylogenetic profiles with parsimony and likelihood. Bioinformatics. 2010;26(15):1910-2) reconstructs the evolution of gene family size. EREM (Affre L, Thompson JD, Debussche M. Genetic structure of continental and island populations of the Mediterranean endemic Cyclamen balearicum (Primulaceae). American Journal of Botany. 1997;84(4):437-51) analyzes gains and losses of genetic features coded by binary traits. PARANA(Patro R, Sefer E, Malin J, Marcais G, Navlakha S, Kingsford C. Parsimonious reconstruction of network evolution. Algorithms for Molecular Biology. 2012;7(1):1) performs parsimony-based inference of ancestral biological networks that represent gene losses and duplications.

[0146] There are also several web server based applications that allow researchers to use ML methods for ancestry reconstruction of different trait types without the need to install any software. For example, Ancestors (Diallo AB, Makarenkov V, Blanchette M. Ancestors 1.0: a web server for ancestral sequence reconstruction. Bioinformatics. 2010;26(1):130-1) is a web server for reconstructing ancestral genomes by identifying and locating syntenic regions. FastML (Ashkenazy H, Penn O, Doron-Faigenboim A, Cohen O, Cannarozzi G, Zomer O, et al. FastML: a web server for probabilistic reconstruction of ancestral sequences. Nucleic acids research. 2012;40(W1):W580-W4) is a web server for probabilistic reconstruction of ancestral sequences by ML using gap phenotypic models to reconstruct indel variation. MLGO (Hu F, Lin Y, Tang J. MLGO: phylogeny reconstruction and ancestral inference from gene-order data. BMC bioinformatics. 2014;15(1):1) is a web server for ML gene sequence analysis.

[0147] The candidate optimized antigenic pathogen polypeptides of the polypeptide library may include one or more regions of the amino acid sequence identified through the ARS. Optionally, the ancestral amino acid sequence or each region of the candidate optimized antigenic pathogen polypeptide is at least 1, 2, 3, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 amino acid residues in length. Optionally, the ancestral amino acid sequence or each region of the candidate optimized antigenic pathogen polypeptide is at most 5, 10, 15, 20, 25, 30, 40, 50, 100, 150, 200, 250, 300, 350, 400, 450, or 500, 600, 700, or 800 amino acid residues in length.

[0148] Optionally, the candidate optimized antigenic pathogen polypeptides of the polypeptide library comprise an amino acid sequence that has at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid identity along its entire length with the amino acid sequence of one or more pathogen polypeptides of the distinct isolates from which the candidate optimized antigenic pathogen polypeptides are optimized.

[0149] Optionally, the method of the present invention includes optimizing the codons of the different generated nucleotide sequences for optimal expression of the encoded candidate optimized antigenic pathogen polypeptide in an expression system. Codon optimization takes advantage of the degeneracy of the genetic code, but does not change the amino acid sequence of the encoded polypeptide. Due to degeneracy, one protein can be encoded by many alternative nucleic acid sequences. Codon preferences (codon usage bias) are different in each organism, which can be a challenge in expressing recombinant proteins in heterologous expression systems, resulting in low and unreliable expression.

[0150] Any suitable expression system may be used. Some suitable examples are well known to those of skill in the art, including expression in mammalian, yeast, insect, or bacterial cells. Optionally, the expression system comprises mammalian cells. Optionally, the expression system comprises yeast, insect, or bacterial cells.

[0151] Codon optimization methods are well known to those skilled in the art. Codon optimization algorithms can be used to design codon-optimized nucleotide sequences that code for amino acids. Such algorithms aim to provide codon-optimized sequences that maximize the expression of a polypeptide or protein in a desired expression system. Examples of suitable codon optimization algorithms include the GeneOptimizer™ algorithm (ThermoFisher), the OptimumGene™ algorithm (GenScript), and GeneGPS® (ATUM).

[0152] Optionally, the method of the present invention also includes other sequence optimizations to maximize protein expression in desired expression systems.Such gene optimization takes into account codon usage bias and other sequence-related parameters related to gene expression, such as transcription, splicing, translation, and mRNA degradation.Examples of such sequence-related parameters are shown below (parameters are categorized below as affecting transcription efficiency, translation efficiency, or protein refolding, although some parameters may affect more than one of these steps): [Table 6]

[0153] Gene optimization algorithms, such as GeneOptimizer™ and OptimumGene™, take several of these parameters into account.

[0154] Genetic optimization of human protein expression in E. coli is reviewed by Maertens et al. (Protein Science 2010 Vol. 19:1312-1326).

[0155] Optionally, the method of the invention includes optimizing the distinct nucleotide sequences for antigenicity of the encoded candidate optimized antigenic pathogen polypeptide.

[0156] Antigenicity optimization may include any of the following: (a) deletion or modification of a nucleic acid sequence encoding an amino acid sequence believed to inhibit the production and / or function of an anti-pathogen polypeptide antibody (e.g., deletion or modification of a mucin-like domain - see, e.g., Reynard et al., Journal of Virology, 2009, 9596-9601); (b) region swapping to recover one or more potentially lost encoded epitopes; (c) Site-directed mutations, e.g., of N-linked glycosylation sites. Typically, site-directed mutations are designed to delete N-linked glycosylation sites, but there may be situations in which it is desirable to introduce additional sites, e.g., to mask epitopes that elicit non-neutralizing antibodies. The ability of glycosylation to sterically block antibody binding to HA and thus provide protection against the host immune response has been demonstrated for influenza viruses. Sun et al. (Journal of Virology, 2013, 87(15):8756-8766) demonstrate that antibodies induced by viruses with multiple glycosylation sites have broader neutralizing activity than antibodies induced by viruses with fewer glycosylation sites; (d) changes to enhance stability (e.g., disulfide bond formation, reducing degradation of the encoded polypeptide by serine proteases); (e) removal of glycans (improving B cell access); (f) Insertion of a nucleic acid sequence, for example to insert a nucleic acid sequence encoding a desired epitope.

[0157] Antigenic optimization of the ectodomain of HIV-1 gp120 has been described by Joyce et al. (J Virol. 2013 Feb;87(4):2294-306).

[0158] Optionally, the different pathogen isolates include different pathogen isolates from an outbreak of a pathogen of the same subtype as the pathogen against which it is desired to induce a broadly neutralizing immune response.

[0159] Optionally, the different pathogen isolates include different pathogen isolates from a pathogen outbreak of a different subtype but the same type as the pathogen against which it is desired to induce a broadly neutralizing immune response.

[0160] Optionally, the different pathogen isolates include different pathogen isolates from an outbreak of a different type but of the same family of pathogen as the pathogen against which it is desired to induce a broadly neutralizing immune response.

[0161] Optionally, the different pathogen isolates include different previous pathogen isolates of the same subtype, type, or family of pathogens as the pathogen against which it is desired to induce a broadly neutralizing immune response.

[0162] Optionally, the different pathogen isolates include different previous pathogen isolates of the same species, genus, or family of pathogens as the pathogen against which it is desired to induce a broadly neutralizing immune response.

[0163] Optionally, the method of the invention for identifying lead candidate optimized antigenic pathogen polypeptides capable of inducing a broadly neutralizing immune response against a pathogen is an in vitro method.

[0164] According to the present invention, there is provided a method for identifying a nucleic acid sequence encoding an optimized antigenic pathogen polypeptide capable of inducing a broadly neutralizing immune response against a pathogen, comprising the steps of: i) immunizing a human or a non-human animal with a nucleic acid comprising a nucleic acid sequence encoding a lead candidate optimized antigenic pathogen polypeptide identified by a method according to the invention; ii) determining whether a broadly neutralizing immune response is induced in a human or non-human animal following immunization in step (i); and iii) if it is determined from step (ii) that a broadly neutralizing immune response is induced in a human or non-human animal, identifying the nucleic acid sequence as a nucleic acid sequence encoding an optimized antigenic pathogen polypeptide capable of inducing a broadly neutralizing immune response against the pathogen. A method is also provided that includes:

[0165] Optionally, whether a broadly neutralizing immune response is induced in a human or non-human animal is determined by determining whether antibodies in serum obtained from the human or non-human animal bind to more than one pathogen subtype within the same family as the pathogen to which a broadly neutralizing immune response is desired.

[0166] Optionally, whether a broadly neutralizing immune response is induced in a human or non-human animal is determined by determining whether antibodies in serum obtained from the human or non-human animal bind to more than one pathogen type within the same family as the pathogen to which a broadly neutralizing immune response is desired.

[0167] Any suitable non-human animal may be used. Optionally, the non-human animal is a mammal. Optionally, the mammal is a guinea pig or a mouse. Optionally, the non-human animal is a bird.

[0168] According to the present invention, i) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:1, or is identical to SEQ ID NO:1; ii) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to or to SEQ ID NO:2; iii) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:4, or is identical to SEQ ID NO:4; iv) at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:5, or identical to SEQ ID NO:5; v) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:7 or is identical to SEQ ID NO:7; or vi) a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:8 or is identical to SEQ ID NO:8. Also provided is an isolated nucleic acid molecule comprising the nucleic acid sequence or a complement thereof.

[0169] According to the present invention, i) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:10, or is identical to SEQ ID NO:10; ii) at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:12; or iii) a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:14. Also provided is an isolated nucleic acid molecule comprising the nucleic acid sequence or a complement thereof.

[0170] According to the present invention, i) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:19, or is identical to SEQ ID NO:19; ii) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:21; iii) at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:23; iv) at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:25 or identical to SEQ ID NO:25; v) at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:27 or identical to SEQ ID NO:27; vi) at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to or identical to SEQ ID NO:29; or vii) a nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to or identical to SEQ ID NO:31. Also provided is an isolated nucleic acid molecule comprising the nucleic acid sequence or a complement thereof.

[0171] According to the present invention, i) is at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:1, or is identical to the amino acid sequence encoded by SEQ ID NO:1; ii) is at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:2, or is identical to the amino acid sequence encoded by SEQ ID NO:2; iii) is at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:4, or is identical to the amino acid sequence encoded by SEQ ID NO:4; iv) at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:5, or identical to the amino acid sequence encoded by SEQ ID NO:5; v) is at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:7, or is identical to the amino acid sequence encoded by SEQ ID NO:7; vi) at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:8, or identical to the amino acid sequence encoded by SEQ ID NO:8; vii) at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:10, or identical to the amino acid sequence encoded by SEQ ID NO:10; viii) is at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:12, or is identical to the amino acid sequence encoded by SEQ ID NO:12; or ix) Further provided is an isolated polypeptide comprising an amino acid sequence that is at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:14, or is identical to the amino acid sequence encoded by SEQ ID NO:14.

[0172] According to the present invention, i) at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:3, or identical to SEQ ID NO:3; ii) is at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:6, or is identical to SEQ ID NO:6; iii) is at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:9, or is identical to SEQ ID NO:9; iv) at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:11, or identical to SEQ ID NO:11; v) at least 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:13; or vi) at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:15, or identical to SEQ ID NO:15 Also provided is an isolated polypeptide comprising the amino acid sequence.

[0173] According to the present invention, i) at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:18, or identical to SEQ ID NO:18; ii) at least 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:20; iii) at least 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:22; iv) at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:24, or identical to SEQ ID NO:24; v) at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:26, or identical to SEQ ID NO:26; vi) at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:28 or identical to SEQ ID NO:28; or vii) at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:30, or identical to SEQ ID NO:30 Also provided is an isolated polypeptide comprising the amino acid sequence.

[0174] Similarity between amino acid or nucleic acid sequences is expressed in terms of the similarity between sequences, or otherwise referred to as sequence identity.Sequence identity is often measured in terms of percentage identity (or similarity or homology), and the higher the percentage, the more similar the two sequences are.Homologues or variants of a given gene or protein will have a relatively high degree of sequence identity when aligned using standard methods.Methods for aligning sequences for comparison are well known in the art. Various programs and alignment algorithms are described in Smith and Waterman, Adv. Appl. Math. 2:482, 1981; Needleman and Wunsch, J. Mol. Biol. 48:443, 1970; Pearson and Lipman, Proc. Natl. Acad. Sci. USA 85:2444, 1988; Higgins and Sharp, Gene 73:237-244, 1988; Higgins and Sharp, CABIOS 5:151-153, 1989; Corpet et al., Nucleic Acids' Research 16:10881-10890, 1988; and Pearson and Lipman, Proc. Natl. Acad. Sci. USA 85:2444, 1988. Altschul et al., Nature Genet. 6:119-129, 1994. The NCBI Basic Local Alignment Search Tool (BLAST™) (Altschul et al., J. Mol. Biol. 215:403-410, 1990) is available from the National Center for Biotechnology Information (NCBI, Bethesda, MD) and from several sources on the Internet, and is used in conjunction with the sequence analysis programs blastp, blastn, blastx, tblastn, and tblastx.

[0175] The sequence identity between nucleic acid sequences or amino acid sequences can be determined by comparing the alignment of sequences. If the equivalent position in the compared sequences is occupied by the same nucleotide or amino acid, the molecules are identical at that position. The score of the alignment as a percentage of identity is a function of the number of identical nucleotides or amino acids at the position shared by the compared sequences. When comparing sequences, optimal alignment may require the introduction of gaps in one or more of the sequences to take into account possible insertions and deletions in the sequence. The sequence comparison method may use a gap penalty, whereby for the same number of identical molecules in the compared sequences, a sequence alignment with as few gaps as possible will achieve a higher score than an alignment with many gaps, since it reflects a higher relatedness between the two compared sequences. The calculation of maximum percent identity involves producing an optimal alignment, taking into account gap penalties.

[0176] Suitable computer programs for carrying out sequence comparisons are widely available in the commercial and public sectors. Examples include MatGat (Campanella et al., 2003, BMC Bioinformatics 4: 29; program available from http: / / bitincka.com / ledion / matgat), Gap (Needleman & Wunsch, 1970, J. Mol. Biol. 48: 443-453), FASTA (Altschul et al., 1990, J. Mol. Biol. 215: 403-410; programs available at http: / / www.ebi.ac.uk / fasta), Clustal W2.0 and X2.0 (Larkin et al. al., 2007, Bioinformatics 23: 2947-2948; program available at http: / / www.ebi.ac.uk / tools / clustalw2), and the EMBOSS pairwise alignment algorithm (Needleman & Wunsch, 1970, supra; Kruskal, 1983, In: Time warps, string edits and macromolecules: the theory and practice of sequence comparison, Sankoff & Kruskal (eds), pp 1-44, Addison Wesley; programs available at http: / / www.ebi.ac.uk / tools / emboss / align). All programs can be run using default parameters.

[0177] For example, sequence comparison may be performed using the "needle" method of the EMBOSS pairwise alignment algorithm, which determines the optimal alignment (including gaps) of two sequences when considered over their entire length and provides a percentage identity score. Default parameters for amino acid sequence comparison ("protein molecule" option) are: gap extension penalty: 0.5, gap opening penalty: 10.0, matrix: Blosum62.

[0178] Sequence comparison may be performed over the entire length of the reference sequence.

[0179] Also provided in accordance with the present invention are isolated nucleic acid molecules comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:6 and a polypeptide comprising the amino acid sequence of SEQ ID NO:9.

[0180] Also provided in accordance with the present invention are isolated nucleic acid molecules comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:13 and a polypeptide comprising the amino acid sequence of SEQ ID NO:15.

[0181] The present invention also provides a composition comprising a first nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:6, and a second nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:9.

[0182] Also provided in accordance with the present invention is a composition comprising a first nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:13, and a second nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:15.

[0183] The present invention also provides a combined preparation comprising (i) a first nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:6, and (ii) a second nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:9.

[0184] The present invention also provides a combined preparation comprising (i) a first nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:13, and (ii) a second nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:15.

[0185] Also provided in accordance with the present invention is a composition comprising a first polypeptide comprising the amino acid sequence of SEQ ID NO:6 and a second polypeptide comprising the amino acid sequence of SEQ ID NO:9.

[0186] Also provided in accordance with the present invention is a composition comprising a first polypeptide comprising the amino acid sequence of SEQ ID NO:13 and a second polypeptide comprising the amino acid sequence of SEQ ID NO:15.

[0187] There is also provided in accordance with the present invention a fusion protein comprising a first polypeptide comprising the amino acid sequence of SEQ ID NO:6 and a second polypeptide comprising the amino acid sequence of SEQ ID NO:9.

[0188] There is also provided in accordance with the present invention a fusion protein comprising a first polypeptide comprising the amino acid sequence of SEQ ID NO:13 and a second polypeptide comprising the amino acid sequence of SEQ ID NO:15.

[0189] The present invention also provides a combined preparation comprising (i) a first polypeptide comprising the amino acid sequence of SEQ ID NO:6, and (ii) a second polypeptide comprising the amino acid sequence of SEQ ID NO:9.

[0190] The present invention also provides a combined preparation comprising (i) a first polypeptide comprising the amino acid sequence of SEQ ID NO:13, and (ii) a second polypeptide comprising the amino acid sequence of SEQ ID NO:15.

[0191] The term "combination preparation" as used herein refers to a "kit of parts" in the sense that the combination components (i) and (ii) defined above can be administered independently or by using different fixed combinations of distinct amounts of the combination components (i) and (ii). The components can be administered simultaneously or sequentially. When the components are administered sequentially, the interval between administrations is preferably selected so that the therapeutic effect of the combined use of the components is greater than the effect obtained by using only one of the combination components (i) and (ii).

[0192] The components of the combination preparation may be present in one combined unit dosage form, or may be present as a first unit dosage form of component (i) and a second unit dosage form of separate component (ii). The ratio of the total amount of combination component (i) to combination component (ii) administered in the combination preparation may be varied, for example, to address the needs of a patient subpopulation being treated, or the needs of a single patient, which may be due to, for example, the patient's particular disease, age, sex, or weight.

[0193] Preferably, there is at least one beneficial effect, such as an enhancement of the effect of component (i), or component (ii), or a mutual enhancement of the effects of the combination components (i) and (ii), such as a more than additive effect, an additional beneficial effect, fewer side effects, less toxicity, or a combined therapeutic effect compared to the effective dosage of one or both of the combination components (i) and (ii), and highly preferably a synergistic effect of the combination components (i) and (ii).

[0194] The combination preparation of the present invention can be provided as a pharmaceutical combination preparation for administration to a mammal, preferably a human. Component (i) can be provided together with a pharma- ceutically acceptable carrier, excipient, or diluent, if necessary, and / or component (ii) can be provided together with a pharma- ceutically acceptable carrier, excipient, or diluent, if necessary.

[0195] Further provided in accordance with the present invention is an isolated nucleic acid molecule that encodes an amino acid sequence encoded by a nucleic acid of the present invention.

[0196] Further provided in accordance with the invention is an isolated nucleic acid molecule encoding an amino acid sequence encoded by a nucleic acid of the invention that has been codon-optimized for expression in a mammalian cell.

[0197] Further provided in accordance with the present invention is an isolated nucleic acid molecule encoding an amino acid sequence encoded by a nucleic acid of the present invention that has been genetically optimized for expression in a mammalian cell.

[0198] The present invention also provides an isolated nucleic acid encoding a polypeptide of the present invention.

[0199] Also provided in accordance with the present invention is an isolated nucleic acid molecule encoding a polypeptide of the present invention, wherein the nucleic acid is codon optimized for expression in a mammalian cell.

[0200] Also provided in accordance with the present invention is an isolated nucleic acid molecule encoding a polypeptide of the present invention, wherein the nucleic acid is genetically optimized for expression in a mammalian cell.

[0201] The present invention also provides a vector comprising the nucleic acid of the present invention.

[0202] Optionally, the vector further comprises a promoter operably linked to the nucleic acid.

[0203] Optionally, the promoter is for expression of a polypeptide encoded by the nucleic acid in a mammal.

[0204] Optionally, the promoter is for expression of a polypeptide encoded by a nucleic acid in a yeast, bacterial, or insect cell.

[0205] Optionally, the vector is a vaccine vector. Optionally, the vaccine vector is a viral vaccine vector, a bacterial vaccine vector, or a nucleic acid vector (e.g., an RNA vaccine vector or a DNA vaccine vector).

[0206] The nucleic acid molecule of the present invention can comprise DNA or RNA molecules.In the embodiment in which the nucleic acid molecule comprises an RNA molecule, it is recognized that the molecule can comprise an RNA sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to or identical to any of SEQ ID NO: 1, 2, 4, 5, 7, 8, 10, 12, 14, 19, 21, 23, 25, 27, 29 or 31, or its complement, in which each "T" nucleotide is replaced with "U".

[0207] For example, where an RNA vaccine vector is provided comprising a nucleic acid of the invention, it is recognized that the nucleic acid sequence of the nucleic acid of the invention is an RNA sequence and thus may comprise, for example, an RNA nucleic acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, or identical to, any of SEQ ID NOs: 1, 2, 4, 5, 7, 8, 10, 12, 14, 19, 21, 23, 25, 27, 29, or 31, or a complement thereof, in which each "T" nucleotide is replaced with a "U".

[0208] Also provided in accordance with the present invention is an isolated cell comprising or transfected with a vector of the present invention.

[0209] The present invention also provides viral pseudotyped particles comprising a polypeptide of the invention.

[0210] The invention also provides a method for producing viral pseudotyped particles comprising transfecting a host cell with a vector comprising a nucleic acid of the invention.

[0211] The present invention also provides a fusion protein comprising a polypeptide of the present invention.

[0212] The invention further provides a pharmaceutical composition comprising a nucleic acid of the invention and a pharma- ceutically acceptable carrier, excipient, or diluent.

[0213] The present invention also provides a pharmaceutical composition comprising a vector of the present invention and a pharma- ceutically acceptable carrier, excipient, or diluent.

[0214] The present invention also provides a pharmaceutical composition comprising a polypeptide of the present invention and a pharma- ceutically acceptable carrier, excipient, or diluent.

[0215] Optionally, the pharmaceutical composition of the invention further comprises an adjuvant to enhance the immune response in the subject to the polypeptide of the composition or the polypeptide encoded by the nucleic acid.

[0216] The present invention also provides a method for inducing an immune response against a pathogen in a subject, the method comprising administering to the subject a nucleic acid of the present invention, a polypeptide of the present invention, a vector of the present invention, or a pharmaceutical composition of the present invention.

[0217] Optionally, the pathogen is a virus. Optionally, the virus is a member of the families Filoviridae, Arenaviridae, or Orthomyxoviridae.

[0218] The present invention also provides a method for inducing an immune response in a subject against a virus of the family Filoviridae or Arenaviridae, the method comprising the step of administering to the subject a nucleic acid of the present invention, a polypeptide of the present invention, a vector of the present invention, or a pharmaceutical composition of the present invention.

[0219] The present invention also provides a method for immunizing a subject against a pathogen, the method comprising the step of administering to the subject a nucleic acid of the present invention, a polypeptide of the present invention, a vector of the present invention, or a pharmaceutical composition of the present invention.

[0220] Optionally, the pathogen is a virus. Optionally, the virus is a member of the families Filoviridae, Arenaviridae, or Orthomyxoviridae.

[0221] According to the present invention there is further provided a method of immunizing a subject against a virus of the Filoviridae family, the method comprising the step of administering to the subject a nucleic acid of the present invention, a polypeptide of the present invention, a vector of the present invention, or a pharmaceutical composition of the present invention.

[0222] The present invention also provides a method for inducing an immune response in a subject against a virus of the Filoviridae family, the method comprising the step of administering to the subject a nucleic acid of the present invention, a polypeptide of the present invention, a vector of the present invention, or a pharmaceutical composition of the present invention.

[0223] Optionally, a nucleic acid, vector, or pharmaceutical composition of the invention comprises a nucleic acid comprising a sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, or identical to, any of SEQ ID NOs: 1, 2, 4, 5, 7, 8, 10, 12, or 14, or comprises a nucleic acid encoding an amino acid sequence encoded by a nucleic acid comprising a sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, or identical to any of SEQ ID NOs: 1, 2, 4, 5, 7, 8, 10, 12, or 14.

[0224] Optionally, a polypeptide, vector, or pharmaceutical composition of the invention comprises an amino acid sequence that is at least 95%, 96%, 97%, 98%, or 99% identical to, or identical to, the amino acid sequence encoded by any of SEQ ID NOs: 1, 2, 4, 5, 7, 8, 10, 12, or 14, or comprises a polypeptide that comprises an amino acid sequence that is at least 95%, 96%, 97%, 98%, or 99% identical to, or identical to, any of SEQ ID NOs: 3, 6, 9, 11, 13, or 15.

[0225] According to the present invention there is further provided a method of immunizing a subject against a virus of the Arenaviridae family, the method comprising the step of administering to the subject a nucleic acid of the present invention, a polypeptide of the present invention, a vector of the present invention, or a pharmaceutical composition of the present invention.

[0226] The present invention also provides a method for inducing an immune response in a subject against a virus of the Arenaviridae family, the method comprising the step of administering to the subject a nucleic acid of the present invention, a polypeptide of the present invention, a vector of the present invention, or a pharmaceutical composition of the present invention.

[0227] Optionally, a nucleic acid, vector, or pharmaceutical composition of the invention comprises a nucleic acid comprising a sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, or identical to, any of SEQ ID NOs: 19, 21, 23, 25, 27, 29, or 31, or comprises a nucleic acid encoding an amino acid sequence encoded by a nucleic acid comprising a sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, or identical to any of SEQ ID NOs: 19, 21, 23, 25, 27, 29, or 31.

[0228] Optionally, a polypeptide, vector, or pharmaceutical composition of the invention comprises a polypeptide comprising an amino acid sequence that is at least 95%, 96%, 97%, 98%, or 99% identical to, or identical to, the amino acid sequence encoded by any of SEQ ID NOs: 19, 21, 23, 25, 27, 29, or 31, or comprises a polypeptide comprising an amino acid sequence that is at least 95%, 96%, 97%, 98%, or 99% identical to, or identical to, any of SEQ ID NOs: 18, 20, 22, 24, 26, 28, or 30.

[0229] Any suitable route of administration may be used. Methods of administration include, but are not limited to, intradermal, intramuscular, intraperitoneal, parenteral, intravenous, subcutaneous, vaginal, rectal, intranasal, inhalation, or oral. Parenteral administration, such as subcutaneous, intravenous, or intramuscular, is generally accomplished by injection. Injectables can be prepared in conventional forms, either as liquid solutions or suspensions, solid forms suitable for dissolution or suspension in liquid prior to injection, or as emulsions. Injectable solutions and suspensions can be prepared from sterile powders, granules, and tablets of the type already described. Administration can be systemic or local.

[0230] The composition may be administered in any suitable manner, for example with a pharma- ceutically acceptable carrier. Pharmaceutically acceptable carriers are determined in part by the particular composition to be administered and the particular method used to administer the composition. Preparations for parenteral administration include sterile aqueous or non-aqueous solutions, suspensions, and emulsions. Examples of non-aqueous solvents are propylene glycol, polyethylene glycol, vegetable oils, such as olive oil, and injectable organic esters, such as ethyl oleate. Aqueous carriers include water, alcoholic / aqueous solutions, emulsions, or suspensions, including saline and buffered media. Parenteral vehicles include sodium chloride solution, Ringer's dextrose, dextrose, and sodium chloride, lactated Ringer's, or fixed oils. Intravenous vehicles include fluid and nutrient replenishers, electrolyte replenishers (e.g., those based on Ringer's dextrose), and the like. Preservatives and other additives, such as antimicrobial agents, antioxidants, chelating agents, and inert gases, may also be present.

[0231] Some of the compositions may potentially be administered as pharma- ceutically acceptable acid or base addition salts formed by reaction with inorganic acids, such as hydrochloric acid, hydrobromic acid, perchloric acid, nitric acid, thiocyanic acid, sulfuric acid, and phosphoric acid, and organic acids, such as formic acid, acetic acid, propionic acid, glycolic acid, lactic acid, pyruvic acid, oxalic acid, malonic acid, succinic acid, maleic acid, and fumaric acid, or by reaction with inorganic bases, such as sodium hydroxide, ammonium hydroxide, potassium hydroxide, and organic bases, such as mono-, di-, tri-alkyl and aryl amines, and substituted ethanol amines.

[0232] Administration can be achieved by single dose or multiple doses.The dose administered to a subject in the context of the present disclosure must be sufficient to induce beneficial therapeutic response in the subject over time, or to inhibit or prevent infection.The required dose varies from subject to subject, depending on the species, age, weight, and general condition of the subject, the severity of the infection to be treated, the specific composition used, and its mode of administration.The appropriate dose can be determined by those skilled in the art using only routine experimentation.

[0233] Pharmaceutically acceptable carriers include, but are not limited to, saline, buffered saline, dextrose, water, glycerol, ethanol, and combinations thereof. The carrier and composition can be sterile, and the formulation is compatible with the mode of administration. The composition can also contain minor amounts of wetting or emulsifying agents, or pH buffering agents. The composition can be a liquid solution, suspension, emulsion, tablet, pill, capsule, sustained release formulation, or powder. The composition can be formulated as a suppository with conventional binders and carriers, such as triglycerides. Oral formulations can contain standard carriers, such as pharmaceutical grades of mannitol, lactose, starch, magnesium stearate, sodium saccharin, cellulose, and magnesium carbonate. Any of the common pharmaceutical carriers can be used, such as sterile saline solution or sesame oil. The vehicle can also contain conventional pharmaceutical auxiliary materials, such as pharma- ceutically acceptable salts to adjust osmotic pressure, buffers, preservatives, and the like. Other vehicles that can be used with the compositions and methods provided herein are saline and sesame oil.

[0234] In some embodiments, the composition comprises a pharma- ceutically acceptable carrier and / or an adjuvant. For example, the adjuvant can be alum, Freund's complete adjuvant, a biological adjuvant, or an immunostimulatory oligonucleotide (e.g., a CpG oligonucleotide).

[0235] The pharma- ceutically acceptable carriers (vehicles) useful in this disclosure are conventional. Remington's Pharmaceutical Sciences, by E. W. Martin, Mack Publishing Co., Easton, PA, 15 th Edition (1975) describes compositions and formulations suitable for pharmaceutical delivery of one or more therapeutic compositions, such as one or more influenza vaccines, and additional agents.

[0236] Generally, the nature of the carrier will depend on the particular mode of administration used. For example, parenteral formulations usually contain injectable fluids that contain pharma- ceutically and physiologically acceptable fluids, such as water, physiological saline, buffered saline, aqueous dextrose, glycerol, etc., as vehicles. For solid compositions (e.g., powder, pill, tablet, or capsule forms), typical non-toxic solid carriers can include, for example, pharmaceutical grades of mannitol, lactose, starch, or magnesium stearate. In addition to biologically neutral carriers, the pharmaceutical compositions to be administered can contain minor amounts of non-toxic auxiliary substances, such as wetting or emulsifying agents, preservatives, and pH buffering agents, such as sodium acetate or sorbitan monolaurate.

[0237] Optionally, the compositions of the invention are administered intramuscularly.

[0238] Optionally, the composition is administered intramuscularly, intradermally, subcutaneously, by needle, or by gene gun or electroporation.

[0239] Also provided in accordance with the present invention is a nucleic acid expression vector that includes a multiple cloning site that includes KpnI and NotI endonuclease sites.

[0240] Optionally, the multiple cloning site comprises the nucleic acid sequence of SEQ ID NO:16.

[0241] Optionally, the nucleic acid expression vector is a nucleic acid expression vector and a viral pseudotype vector.

[0242] Optionally, the nucleic acid expression vector is a vaccine vector.

[0243] Optionally, the nucleic acid expression vector comprises, from 5' to 3' direction: a promoter; a splice donor site (SD); a splice acceptor site (SA); and a terminator signal, with a multiple cloning site located between the splice acceptor site and the terminator signal.

[0244] Optionally, the promoter comprises the CMV immediate early 1 enhancer / promoter (CMV-IE-E / P) and / or the terminator signal comprises the terminator signal of the bovine growth hormone gene (Tbgh) lacking a KpnI restriction endonuclease site.

[0245] Optionally, the nucleic acid expression vector further comprises an origin of replication and a nucleic acid encoding resistance to an antibiotic. Optionally, the origin of replication comprises a pUC-plasmid origin of replication and / or the nucleic acid encodes resistance to kanamycin.

[0246] Optionally, the nucleic acid expression vector comprises the nucleic acid sequence of SEQ ID NO: 17 (pEVAC).

[0247] The polypeptide of the present invention may contain one or more conservative amino acid substitutions.Conservative amino acid substitutions are those substitutions that, when made, are least likely to interfere with the properties of the original protein, i.e., the structure and especially the function of the protein are preserved and are not significantly changed by such substitutions.Examples of conservative substitutions are shown below: [Table 7]

[0248] Conservative substitutions generally maintain (a) the structure of the polypeptide backbone in the area of ​​the substitution, e.g., as a sheet or helix conformation, (b) the charge or hydrophobicity of the molecule at the target site, or (c) the bulk of the side chain.

[0249] In general, the substitutions expected to produce the greatest changes in protein properties will be non-conservative changes, such as (a) a hydrophilic residue, such as seryl or threonyl, substituting (or being substituted by) a hydrophobic residue, such as leucyl, isoleucyl, phenylalanyl, valyl, or alanyl; (b) a cysteine ​​or proline substituting (or being substituted by) any other residue; (c) a residue having an electronegative side chain, such as lysyl, arginyl, or histidyl, substituting (or being substituted by) an electronegative residue, such as glutamyl or aspartyl; or (d) a residue having a bulky side chain, such as phenylalanine, substituting (or being substituted by) a residue having no side chain, such as glycine.

[0250] In certain embodiments of the present invention, sequence alignment and ancestral sequence reconstruction (ASR) can be used to identify highly conserved immune targets that pathogens cannot change and that will inevitably be present in future pandemics of their virus family, even the most variable RNA viruses.Synthetic gene technology can be used to produce computer-generated viral genes that are highly expressed and can be easily cloned into expression vectors, such as pEVAC vectors, which have proven to be highly versatile expression vectors for generating viral pseudotypes and for direct DNA vaccination of animals and / or humans.

[0251] A large panel of genes can be generated using pEVAC vectors so that viral pseudotypes can be generated quickly. This allows a library of viral pseudotypes, each with its own unique viral protein, to be probed by a large panel of monoclonal antibodies. This process ensures that the inserts generate conformationally correct viral surface proteins to present the most accessible targets for the Achilles heel of the virus. Narrowing the selection of candidates by pseudotyping and mAb binding provides a shortlist of top candidates to test by vaccination. This can be done in guinea pigs, where a streamlined process of pEVAC-vaccine inserts is delivered. If necessary, this allows shuttle of vaccine inserts from DNA pEVAC vectors to various viral vectors based on advanced designed convenient cloning sites. Since chimpanzee adenovectors (ChAd) were widely used in Phase I to evaluate most Ebola virus vaccine candidates for the West African pandemic, we choose to compare using the same vector for direct comparison in humans. For screening in guinea pigs, we used pEVAC-vaccine insert followed by DNA priming with the ChAd-vaccine insert.

[0252] In a particular embodiment of the invention:

[0253] 1) High-throughput "deep" sequencing technologies provide viral variation data from current and past pandemics. Analysis of this data can identify structurally highly conserved regions that preserve known B and T cell epitopes, which can be used as a scaffold for designing optimal vaccine inserts.

[0254] 2) Human monoclonal antibody (mAb) technology allows the generation of anti-viral mAbs against vaccine targets, e.g., viral envelope proteins, and identifies epitope-rich regions for targeting by broadly neutralizing monoclonal antibodies (BNmAbs).

[0255] 3) Optimal gene design and synthesis incorporates digitally modeled conserved scaffolds of the genes identified in (1) to include the broadest NmAb epitopes in these scaffolds (BN epitopes may not be optimally presented on the naturally conserved GP).

[0256] 4) Downstream information of convenient cloning sites matching the requirements of vaccine and pseudotype vectors can be taken into account during the design and synthesis of RNA and codon-optimized synthetic genes as vaccine inserts, allowing rapid and highly efficient cloning and shuttling into different screening (i.e., lentiviral pseudotype; PV) and vaccine (i.e., MVA, ChAd, VSV, DNA, etc.) vectors.

[0257] 5) Viral pseudotypes (lentiviruses) generated from the digitally designed inserts are screened for functionality in vitro via transduction and infection studies. In addition to this, neutralization assays are performed using a panel of BNmAbs and patient sera to ensure that known epitopes are preserved.

[0258] 6) Narrowing the selection to the best class of vaccine inserts from several synthetic vaccine inserts will be confirmed for immunogenicity in guinea pigs using rapid DNA priming (and if necessary) adenovirus boosting, a method that produces highly reproducible titers. In vivo screening will identify which are the most immunogenic and produce the greatest breadth of neutralization.

[0259] The central role of viral glycoproteins in cell attachment, fusion, and uncoating makes them an important antigenic target for viral vaccines and monoclonal antibody therapy pioneered during the West African Ebola outbreak. Analysis of GP sequences among EBOV species showed a high degree of diversity at the nucleotide and amino acid level (only about 60-65% nt identity). Current conventional filovirus vaccine approaches require GP-targeted vaccines to be multivalent, encoding a more conserved (about 97-98% identity in GP nucleotide sequence) specific GP for each species. Vaccines using older strains of EBOV (rVSV.ZEBOV=Kikwit) suggest that they may provide cross-protection (Henao-Restrepo AM, Lancet 2015), but there are concerns that this may be of limited efficacy against future outbreaks of other diverse highly pathogenic filoviruses.

[0260] The inventors can achieve dramatic improvements in vaccine efficacy against novel viral variants based on sequence data (including, optionally, e.g., pandemic sequence data) and generate synthetic optimized vaccine inserts to provide the broadest possible vaccine protection against future pandemics of variable RNA viruses. In certain embodiments, the inventors' novel vaccine technology: (1) Pandemic pathogen sequence (2) Broadly antiviral neutralizing monoclonal antibodies (BNmAbs) derived from pandemic survivors; and (3) computational modeling methodology. (4) Synthetic gene technology and antigen display technology (5) High-throughput virus binding and neutralization screening (6) In vivo immune selection and readout of vaccine efficacy Combine the two.

[0261] The final product is a novel immunogen used to induce the broadest spectrum of protective immune responses. We provide proof of concept that a next-generation single vaccine insert indeed induces a broad neutralization profile against the Ebolavirus genus (Zaire, Sudan, Bundibugyo) and further targets the more distantly related filovirus, Marburg virus.

[0262] Embodiments of the present invention will now be described, by way of example only, in the following examples and with reference to the accompanying drawings, in which: The present invention provides, for example, the following items. (Item 1) 1. A method for identifying lead candidate optimized antigenic pathogen polypeptides capable of inducing a broadly neutralizing immune response against a pathogen, comprising: i) providing a polypeptide library comprising a plurality of different candidate optimized antigenic pathogen polypeptides, the amino acid sequence of each different candidate being optimized from a plurality of different amino acid sequences of a pathogen polypeptide, wherein each different amino acid sequence of the pathogen polypeptides is different from each different amino acid sequence of the pathogen polypeptides, each different amino acid sequence of the pathogen polypeptides comprising an amino acid sequence of a polypeptide of a different isolate, each different isolate being an isolate of a pathogen of the same family as a pathogen against which it is desired to induce a broadly neutralizing immune response; ii) screening the candidate optimized antigenic pathogen polypeptides of the polypeptide library for binding by one or more broadly neutralizing antigen binding molecules, each capable of binding to and / or neutralizing pathogens of the same family as the pathogen against which it is desired to induce a broadly neutralizing immune response; and iii) identifying the candidate optimized antigenic pathogen polypeptide bound by one or more of the antigen binding molecules in step (ii) as a lead candidate optimized antigenic pathogen polypeptide capable of inducing a broadly neutralizing immune response against said pathogen. A method comprising: (Item 2) The method of claim 1, wherein the one or more broadly neutralizing antigen-binding molecules comprise an antibody obtained or derived from an antibody obtained from a subject exposed to a pathogen of the same family as the pathogen against which it is desired to induce a broadly neutralizing immune response. (Item 3) 3. The method of claim 1 or 2, wherein the one or more broadly neutralizing antigen-binding molecules comprise a non-antibody antigen-binding protein. (Item 4) The method of claim 3, wherein the one or more broadly neutralizing antigen binding molecules comprise designed ankyrin repeat proteins (DARPins), anticalins, aptamers, or T cell receptor molecules. (Item 5) The method of any of the preceding items, wherein the candidate optimized antigenic pathogen polypeptides of the polypeptide library are expressed in or on the surface of a mammalian cell. (Item 6) 5. The method of any of items 1 to 4, wherein the candidate optimized antigenic pathogen polypeptides of the polypeptide library are expressed in or on the surface of bacterial, yeast, or insect cells. (Item 7) The method of any of the preceding items, wherein the pathogen is a virus, the candidate optimized antigenic pathogen polypeptide is a candidate optimized antigenic viral polypeptide, and the pathogen peptide is a viral polypeptide. (Item 8) 8. The method of claim 7, wherein the polypeptide library is a viral pseudotype library comprising a plurality of different viral pseudotypes, each different viral pseudotype comprising a different candidate optimized viral polypeptide. (Item 9) In step (ii), the candidate optimized antigenic viral polypeptides are screened for binding by one or more of the antigen binding molecules by screening the viral pseudotypes for binding and / or neutralization by one or more of the antigen binding molecules. The method of claim 8. (Item 10) 8. The method of any of items 1 to 7, wherein the candidate optimized antigenic pathogen polypeptides are screened for binding by the one or more antigen binding molecules by a flow cytometry assay. (Item 11) The method of any of the preceding items, further comprising the step of generating said polypeptide library. (Item 12) 12. The method of claim 11, wherein the polypeptide library is generated by expressing the different candidate optimized antigenic pathogen polypeptides from a nucleic acid library comprising a plurality of different nucleic acids, each different nucleic acid comprising a nucleotide sequence encoding a different candidate optimized antigenic pathogen polypeptide of the polypeptide library. (Item 13) 13. The method of claim 12, wherein the different candidate optimized pathogen polypeptides are expressed in or on the surface of a mammalian cell. (Item 14) 14. The method according to item 12 or 13, wherein the nucleotide sequence of each different nucleic acid of the nucleic acid library is codon-optimized and optionally gene-optimized for expression of the encoded polypeptide in a mammalian cell. (Item 15) 15. The method of any of items 12 to 14, wherein each different nucleic acid of the nucleic acid library is part of an expression vector for expression of the nucleic acid in a mammalian cell. (Item 16) 16. The method of any of items 12 to 15, wherein the pathogen is a virus, the candidate optimized antigenic pathogen polypeptide is a candidate optimized antigenic viral polypeptide, and the pathogen peptide is a viral polypeptide. (Item 17) 17. The method of claim 16, wherein the nucleic acid library is a viral pseudotype vector library, each different nucleic acid of the library is part of an expression vector for production of a viral pseudotype comprising the encoded viral polypeptide, and the polypeptide library is a viral pseudotype library generated by producing viral pseudotypes from the expression vectors of the viral pseudotype vector library, the viral pseudotype library comprising a plurality of different viral pseudotypes, each different viral pseudotype comprising a different candidate optimized viral polypeptide encoded by a different nucleic acid sequence of the viral pseudotype vector library. (Item 18) 18. The method according to any of items 15 to 17, wherein the expression vector is also a vaccine vector. (Item 19) 19. The method of claim 18, wherein the vaccine vector is a viral vaccine vector, a bacterial vaccine vector, an RNA vaccine vector, or a DNA vaccine vector. (Item 20) 20. The method of claim 18 or 19, wherein the vaccine vector is based on a viral delivery vector, such as a poxvirus (e.g., MVA, NYVAC, AVIPOX), a herpesvirus (e.g., HSV, CMV, adenovirus of any host species), a morbillivirus (e.g., measles), an alphavirus (e.g., SFV, Sendai), a flavivirus (e.g., yellow fever), or a rhabdovirus (e.g., VSV) based viral delivery vector, a bacterial delivery vector (e.g., Salmonella, E. coli), an RNA expression vector, or a DNA expression vector. (Item 21) 21. The method according to any of items 15 to 20, wherein the vector is a pEVAC-based expression vector. (Item 22) 13. The method of claim 12, wherein the different candidate optimized antigenic pathogen polypeptides are expressed in or on the surface of bacteria, yeast, or insect cells. (Item 23) 23. The method of any of items 12 to 22, further comprising generating the nucleic acid library by synthesizing a plurality of different nucleic acids, each different nucleic acid comprising a different nucleotide sequence encoding a different candidate optimized antigenic pathogen polypeptide. (Item 24) i) obtaining the amino acid sequences of said pathogen polypeptides and / or the nucleotide sequences encoding said pathogen polypeptides of said different pathogen isolates; and ii) generating a plurality of different nucleotide sequences, each different nucleotide sequence encoding a different candidate optimized antigenic pathogen polypeptide, the encoded amino acid sequence of each different candidate optimized antigenic pathogen polypeptide being optimized from the obtained or encoded amino acid sequence of the pathogen polypeptide and different from each of the obtained or encoded amino acid sequences. 24. The method of claim 23, further comprising: (Item 25) The generation of the plurality of different nucleotide sequences in step (ii) of item 24, performing a multiple sequence alignment of the amino acid or nucleotide sequences obtained in step (i) of item 24; Identifying highly conserved or encoded amino acid sequences among the polypeptides of the different pathogen isolates from the multiple sequence alignment; and 25. The method of claim 24, comprising the step of generating a plurality of different nucleotide sequences, each different nucleotide sequence encoding a different candidate optimized antigenic pathogen polypeptide, and one or more of the different nucleotide sequences comprising a sequence encoding a highly conserved amino acid sequence or encoded amino acid sequence identified from the multiple sequence alignment. (Item 26) Identifying an amino acid sequence or an encoded amino acid sequence that is an ancestral amino acid sequence from the multiple sequence alignment; and including in one or more of the different generated nucleotide sequences a sequence encoding an ancestral amino acid sequence identified from said multiple sequence alignment. 26. The method of claim 25, further comprising: (Item 27) 27. The method according to any of items 24 to 26, comprising codon optimisation, optionally gene optimising codons, of said differentially generated nucleotide sequences for optimal expression of said encoded candidate optimised antigenic pathogen polypeptide in an expression system. (Item 28) 28. The method of claim 27, wherein the expression system comprises a mammalian cell. (Item 29) 28. The method of claim 27, wherein the expression system comprises yeast, bacterial, or insect cells. (Item 30) 30. The method of any of items 24 to 29, comprising optimizing the different nucleotide sequences for antigenicity of the encoded candidate optimized antigenic pathogen polypeptide. (Item 31) The antigenic optimization comprises: Deletion or modification of a nucleic acid sequence encoding an amino acid sequence that inhibits the production and / or function of an anti-pathogen polypeptide antibody (e.g., deletion or modification of a mucin-like domain); region swapping to recover one or more potentially lost encoded epitopes; For example, site-directed mutagenesis of N-linked glycosylation sites; changes to enhance stability (e.g., disulfide bond formation, reducing degradation of the encoded polypeptide by serine proteases); Glycan removal; Insertion of a nucleic acid sequence, e.g., to insert a nucleic acid sequence encoding a desired epitope. 31. The method according to item 30, comprising any one of the steps: (Item 32) The method of any of the preceding items, wherein the one or more broadly neutralizing antigen-binding molecules described in step (ii) of item 1 comprise a broadly neutralizing antibody, preferably a broadly neutralizing monoclonal antibody (BNmAb). (Item 33) The method according to any of the preceding items, wherein the one or more antigen-binding molecules described in step (ii) of item 1 comprise antibodies obtained, or antibodies derived from antibodies obtained, from subjects who survived an outbreak of a pathogen of the same family, and optionally the same subtype or type, as the pathogen against which it is desired to induce a broadly neutralizing immune response. (Item 34) 34. The method of claim 33, wherein the subject from which the antibody is obtained or derived is a human or non-human mammalian subject. (Item 35) 35. The method of claim 33 or 34, wherein the one or more antigen-binding molecules comprise a broadly neutralizing monoclonal antibody (BNmAb). (Item 36) The method of any of the preceding items, wherein the different pathogen isolates comprise different pathogen isolates from an outbreak of a pathogen of the same subtype as the pathogen against which it is desired to induce a broadly neutralizing immune response. (Item 37) 2. The method of any of the preceding claims, wherein the different pathogen isolates comprise different pathogen isolates from a pathogen outbreak of a different subtype but the same type as the pathogen against which it is desired to induce a broadly neutralizing immune response. (Item 38) The method of any of the preceding items, wherein the different pathogen isolates comprise different pathogen isolates from an outbreak of pathogens of a different group but the same family as the pathogen against which it is desired to induce a broadly neutralizing immune response. (Item 39) The method of any of the preceding items, wherein the different pathogen isolates comprise different previous pathogen isolates of the same subtype, type, or family as the pathogen against which it is desired to induce a broadly neutralizing immune response. (Item 40) The method of any of the preceding items, wherein each candidate optimized antigenic pathogen polypeptide comprises at least 20 amino acid residues. (Item 41) Item 11. The method of any of the preceding items, wherein the pathogen is a virus. (Item 42) 42. The method of claim 41, wherein the virus is an RNA virus. (Item 43) 43. The method of item 41 or 42, wherein the virus is an emerging or re-emerging RNA virus. (Item 44) 44. The method of any of items 41 to 43, wherein the virus is a filovirus, an arenavirus, or an orthomyxovirus. (Item 45) 44. The method according to any of items 41 to 43, wherein the virus is an Ebola virus or a Marburg virus. (Item 46) 44. The method according to any of items 41 to 43, wherein the virus is Lassa virus. (Item 47) Item 11. The method of any of the preceding items, wherein the pathogen polypeptide is a viral glycoprotein. (Item 48) 2. The method according to any of the preceding claims, which is an in vitro method. (Item 49) 1. A method for identifying a nucleic acid sequence encoding an optimized antigenic pathogen polypeptide capable of inducing a broadly neutralizing immune response against a pathogen, comprising: i) immunizing a human or non-human animal with a nucleic acid comprising a nucleic acid sequence encoding a lead candidate optimized antigenic pathogen polypeptide identified by the method of any of the preceding items; ii) determining whether a broadly neutralizing immune response has been induced in said human or said non-human animal following immunization in step (i); and iii) if it is determined from step (ii) that a broadly neutralizing immune response is induced in said human or said non-human animal, identifying said nucleic acid sequence as a nucleic acid sequence encoding an optimized antigenic pathogen polypeptide capable of inducing a broadly neutralizing immune response against said pathogen. A method comprising: (Item 50) 50. The method of claim 49, comprising determining whether a broadly neutralizing immune response is induced in the human or non-human animal by determining whether antibodies in serum obtained from the human or non-human animal bind to and / or neutralize more than one pathogen subtype. (Item 51) 51. The method of item 49 or 50, wherein the non-human animal is a mammal. (Item 52) 52. The method of claim 51, wherein the mammal is a guinea pig or a mouse. (Item 53) 51. The method according to item 49 or 50, wherein the non-human animal is a bird. (Item 54) i) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:1, or is identical to SEQ ID NO:1; ii) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:2; iii) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:4, or is identical to SEQ ID NO:4; iv) at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:5, or identical to SEQ ID NO:5; v) at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:7 or identical to SEQ ID NO:7; or vi) at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:8 or identical to SEQ ID NO:8 An isolated nucleic acid molecule comprising a nucleic acid sequence, or its complement. (Item 55) i) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:10, or is identical to SEQ ID NO:10; ii) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:12; or iii) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:14. An isolated nucleic acid molecule comprising a nucleic acid sequence, or its complement. (Item 56) i) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:19, or is identical to SEQ ID NO:19; ii) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:21; iii) at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:23; iv) at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:25 or identical to SEQ ID NO:25; v) at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:27 or identical to SEQ ID NO:27; vi) at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to or identical to SEQ ID NO:29; or vii) at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to or identical to SEQ ID NO:31. An isolated nucleic acid molecule comprising a nucleic acid sequence, or its complement. (Item 57) i) is at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:1, or is identical to the amino acid sequence encoded by SEQ ID NO:1; ii) is at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:2, or is identical to the amino acid sequence encoded by SEQ ID NO:2; iii) is at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:4, or is identical to the amino acid sequence encoded by SEQ ID NO:4; iv) at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:5, or identical to the amino acid sequence encoded by SEQ ID NO:5; v) is at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:7, or is identical to the amino acid sequence encoded by SEQ ID NO:7; vi) at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:8, or identical to the amino acid sequence encoded by SEQ ID NO:8; vii) at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:10, or identical to the amino acid sequence encoded by SEQ ID NO:10; viii) is at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:12, or is identical to the amino acid sequence encoded by SEQ ID NO:12; ix) at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:14, or identical to the amino acid sequence encoded by SEQ ID NO:14 An isolated polypeptide comprising an amino acid sequence. (Item 58) i) at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:3, or identical to SEQ ID NO:3; ii) is at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:6, or is identical to SEQ ID NO:6; iii) is at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:9, or is identical to SEQ ID NO:9; iv) at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:11, or identical to SEQ ID NO:11; v) at least 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:13; or vi) at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:15, or identical to SEQ ID NO:15 An isolated polypeptide comprising an amino acid sequence. (Item 59) i) at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:18, or identical to SEQ ID NO:18; ii) at least 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:20; iii) at least 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:22; iv) at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:24, or identical to SEQ ID NO:24; v) at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:26, or identical to SEQ ID NO:26; vi) at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:28 or identical to SEQ ID NO:28; or vii) at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:30, or identical to SEQ ID NO:30 An isolated polypeptide comprising an amino acid sequence. (Item 60) 57. An isolated nucleic acid encoding an amino acid sequence encoded by the nucleic acid of items 54, 55, or 56, which has been codon-optimized, optionally gene-optimized, for expression in a mammalian cell. (Item 61) 60. An isolated nucleic acid encoding a polypeptide according to items 57, 58 or 59, which has been codon-optimized, optionally gene-optimized, for expression in a mammalian cell. (Item 62) 62. A vector comprising the nucleic acid of item 54, 55, 56, 60, or 61. (Item 63) 63. The vector of item 62, further comprising a promoter operably linked to the nucleic acid. (Item 64) 64. The vector of item 63, wherein the promoter is for expression of the polypeptide encoded by the nucleic acid in a mammalian cell. (Item 65) 64. The vector of item 63, wherein the promoter is for expression of the polypeptide encoded by the nucleic acid in yeast or insect cells. (Item 66) 66. The vector according to any of items 62 to 65, which is a vaccine vector. (Item 67) 67. The vector of item 66, which is a viral vaccine vector, a bacterial vaccine vector, an RNA vaccine vector, or a DNA vaccine vector. (Item 68) 66. An isolated cell comprising the vector according to any of items 62 to 65. (Item 69) 60. A pseudotyped viral particle comprising a polypeptide according to item 57, 58, or 59. (Item 70) 70. A method for producing the pseudotyped viral particle of item 69, comprising transfecting a host cell with the vector of any one of items 62 to 64. (Item 71) 60. A fusion protein comprising the polypeptide of items 57, 58, or 59. (Item 72) 62. A pharmaceutical composition comprising the nucleic acid of item 54, 55, 56, 60, or 61 and a pharma- ceutically acceptable carrier, excipient, or diluent. (Item 73) 68. A pharmaceutical composition comprising a vector according to any of items 62 to 64, 66 or 67, and a pharma- ceutically acceptable carrier, excipient or diluent. (Item 74) 60. A pharmaceutical composition comprising a polypeptide according to item 57, 58 or 59 and a pharma- ceutically acceptable carrier, excipient or diluent. (Item 75) 75. The pharmaceutical composition according to any of items 72 to 74, further comprising an adjuvant for enhancing an immune response in a subject against the polypeptide of the composition or against a polypeptide encoded by the nucleic acid. (Item 76) A method for inducing an immune response against a virus of the Filoviridae family in a subject, comprising administering to the subject a nucleic acid according to any of items 54, 55, 60, or 61, a polypeptide according to item 57 or 58, a vector according to any of items 62 to 64, 66, or 67, or a pharmaceutical composition according to any of items 72 to 75. (Item 77) A method for immunizing a subject against a virus of the Filoviridae family, comprising administering to the subject a nucleic acid according to any of items 54, 55, 60, or 61, a polypeptide according to item 57 or 58, a vector according to any of items 62 to 64, 66, or 67, or a pharmaceutical composition according to any of items 72 to 75. (Item 78) 1. A method for inducing an immune response against a virus of the Arenaviridae family in a subject, comprising administering to the subject a nucleic acid according to any of items 56, 60, or 61, a polypeptide according to item 59, a vector according to any of items 62 to 64, 66, or 67, or a pharmaceutical composition according to any of items 72 to 75. (Item 79) A method for immunizing a subject against a virus of the Arenaviridae family, comprising administering to the subject a nucleic acid according to any one of items 56, 60, or 61, a polypeptide according to item 59, a vector according to any one of items 62 to 64, 66, or 67, or a pharmaceutical composition according to any one of items 72 to 75. (Item 80) 80. The method of any of items 76 to 79, wherein the composition is administered intramuscularly. (Item 81) A nucleic acid expression vector comprising a multiple cloning site containing KpnI and NotI endonuclease sites. (Item 82) 82. The vector of item 81, wherein the multiple cloning site comprises the nucleic acid sequence of SEQ ID NO:16. (Item 83) 83. The vector according to item 81 or 82, which is an expression vector and a viral pseudotype vector. (Item 84) 84. The vector according to any of items 81 to 83, which is a vaccine vector. (Item 85) 85. The vector according to any of items 81 to 84, comprising, in a 5' to 3' direction: a promoter; a splice donor site; a splice acceptor site; and a terminator signal, wherein the multiple cloning site is located between the splice acceptor site and the terminator signal. (Item 86) 86. The vector of item 85, wherein the promoter comprises a CMV immediate-early 1 enhancer / promoter and / or the terminator signal comprises the terminator signal of a bovine growth hormone gene lacking a KpnI restriction endonuclease site. (Item 87) 87. The vector according to any of items 81 to 86, further comprising an origin of replication and a nucleic acid encoding resistance to an antibiotic. (Item 88) 88. The vector of item 87, wherein the origin of replication comprises a pUC-plasmid origin of replication and / or the nucleic acid encodes resistance to kanamycin. (Item 89) 89. The vector according to any of items 81 to 88, comprising the nucleic acid sequence of SEQ ID NO: 17. (Item 90) An isolated nucleic acid molecule comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:6, and a polypeptide comprising the amino acid sequence of SEQ ID NO:9. (Item 91) An isolated nucleic acid molecule comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:13, and a polypeptide comprising the amino acid sequence of SEQ ID NO:15. (Item 92) A composition comprising a first nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:6, and a second nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:9. (Item 93) A composition comprising a first nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:13, and a second nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:15. (Item 94) A combined preparation comprising (i) a first nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:6, and (ii) a second nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:9. (Item 95) A combined preparation comprising (i) a first nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:13, and (ii) a second nucleic acid comprising a nucleotide sequence encoding a polypeptide comprising the amino acid sequence of SEQ ID NO:15. (Item 96) A composition comprising a first polypeptide comprising the amino acid sequence of SEQ ID NO:6, and a second polypeptide comprising the amino acid sequence of SEQ ID NO:9. (Item 97) A composition comprising a first polypeptide comprising the amino acid sequence of SEQ ID NO:13, and a second polypeptide comprising the amino acid sequence of SEQ ID NO:15. (Item 98) A fusion protein comprising a first polypeptide comprising the amino acid sequence of SEQ ID NO:6, and a second polypeptide comprising the amino acid sequence of SEQ ID NO:9. (Item 99) A fusion protein comprising a first polypeptide comprising the amino acid sequence of SEQ ID NO:13, and a second polypeptide comprising the amino acid sequence of SEQ ID NO:15. (Item 100) A combined preparation comprising (i) a first polypeptide comprising the amino acid sequence of SEQ ID NO:6, and (ii) a second polypeptide comprising the amino acid sequence of SEQ ID NO:9. (Item 101) A combined preparation comprising (i) a first polypeptide comprising the amino acid sequence of SEQ ID NO:13, and (ii) a second polypeptide comprising the amino acid sequence of SEQ ID NO:15. (Item 102) A nucleic acid according to any of items 54, 55, 60 or 61, a polypeptide according to item 57 or 58, a vector according to any of items 62 to 64, 66 or 67, or a pharmaceutical composition according to any of items 72 to 75 for use as a medicament. (Item 103) The nucleic acid according to any of items 54, 55, 60 or 61, the polypeptide according to item 57 or 58, the vector according to any of items 62 to 64, 66 or 67, or the pharmaceutical composition according to any of items 72 to 75, for use in the treatment of a viral infection, preferably a viral infection caused by an emerging or re-emerging virus, preferably a virus of the family Filoviridae. (Item 104) Use of a nucleic acid according to any of items 54, 55, 60 or 61, a polypeptide according to item 57 or 58, a vector according to any of items 62 to 64, 66 or 67, or a pharmaceutical composition according to any of items 72 to 75, in the manufacture of a medicament for the treatment of a viral infection, preferably a viral infection caused by an emerging or re-emerging virus, preferably a virus of the family Filoviridae. (Item 105) 5. The nucleic acid according to any of items 56, 60 or 61, the polypeptide according to item 59, the vector according to any of items 62 to 64, 66 or 67, or the pharmaceutical composition according to any of items 72 to 75 for use as a medicament. (Item 106) The nucleic acid according to any of items 56, 60 or 61, the polypeptide according to item 59, the vector according to any of items 62 to 64, 66 or 67, or the pharmaceutical composition according to any of items 72 to 75, for use in the treatment of a viral infection, preferably a viral infection caused by an emerging or re-emerging virus, preferably a virus of the Arenaviridae family. (Item 107) Use of a nucleic acid according to any of items 56, 60 or 61, a polypeptide according to item 59, a vector according to any of items 62 to 64, 66 or 67, or a pharmaceutical composition according to any of items 72 to 75, in the manufacture of a medicament for the treatment of a viral infection, preferably a viral infection caused by an emerging or re-emerging virus, preferably a virus of the Arenaviridae family. (Item 108) 97, 98, 99 or 100. The nucleic acid according to item 90 or 91, the composition according to item 92, 93, 96 or 97, the combined preparation according to item 94, 95, 100 or 101, or the fusion protein according to item 98 or 99 for use as a medicament. (Item 109) The nucleic acid according to item 90 or 91, the composition according to item 92, 93, 96 or 97, the combined preparation according to item 94, 95, 100 or 101, or the fusion protein according to item 98 or 99, for use in the treatment of a viral infection, preferably a viral infection caused by an emerging or re-emerging virus, preferably a virus of the family Filoviridae. (Item 110) Use of a nucleic acid according to item 90 or 91, a composition according to item 92, 93, 96 or 97, a combined preparation according to item 94, 95, 100 or 101, or a fusion protein according to item 98 or 99 in the manufacture of a medicament for the treatment of a viral infection, preferably a viral infection caused by an emerging or re-emerging virus, preferably a virus of the Filoviridae family. [Brief description of the drawings]

[0263] [Figure 1] FIG. 1 shows an illustration of the phylogenetic tree and its relationship to the ancestral sequence reconstruction.

[0264] [Diagram 2] Figure 2 shows a phylogenetic tree comparing Ebola and Marburg viruses. Numbers indicate percent confidence of the branches.

[0265] [Diagram 3] FIG. 3 shows the plasmid map of pEVAC.

[0266] [Figure 4] Figure 4 shows the results of a challenge study in the Ebola challenge model, which was lethal to non-vaccinated guinea pigs (group 1, bottom line), whereas vaccinated guinea pigs (group 2, top line) were all protected (left) and continued to gain weight (right).

[0267] [Diagram 5] Figure 5 shows the results of a pseudotyped virus neutralization assay illustrating the strength of neutralizing antibody responses to target antigens expressed on the surface of pseudotyped viruses representing all Ebola virus species and Marburg virus. The strength of neutralization is shown in a heat map, with red (darkest shading) being very strong neutralization, decreasing from orange to yellow (lighter shading), and no neutralization / equal to the negative control value being white. T2-4 and T2-6 are nucleic acid vaccines encoding lead candidate optimized antigenic Ebola polypeptides in combination with the T2-11 Marburg candidate, in preclinical testing with serum samples taken from immunized guinea pigs.

[0268] [Figure 6] Figure 6 shows the results of a study to determine the efficacy of nucleic acid vaccines encoding different lead candidate optimized antigenic pathogenic polypeptides identified using an embodiment of the method of the present invention. Antibody binding was measured by incubating two populations of cells carrying two different group 1 influenza A glycoproteins (H1 pandemic and seasonal) on their surface with pooled mouse serum. Any bound antibodies were then detected by a secondary antibody and the results were recorded using a flow cytometer. Binding was significantly increased before and after vaccination with all constructs, but not after vaccination with PBS (control). Overall, the vaccine candidates outperformed those from COBRA in both cases (*).

[0269] [Figure 7] FIG. 7 shows the results of a study determining the binding of cells expressing two different group 1 influenza A glycoproteins (seasonal H1N1, and pandemic-origin H1N1) on their cell surface by mouse sera from animals immunized with either COBRA or DIOS HA gene antigens.

[0270] [Figure 8] Figure 8 shows the results of cross-HA group binding (left panel) and pseudotype neutralization (right) of H7N9 (A / Shanghai 2 / 2013) by sera from DIOS or COBRA DNA immunized mice. In the right panel, the top curve is for CR9114, the two curves descending from the two lowest starting points on the left of the graph are for H1N1, and the remaining two curves are for H1N1pdm. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0271] Examples of non-optimized Ebola and Marburg virus ancestral nucleic acid sequences (i.e., sequences that are neither codon-optimized nor gene-optimized) are provided below, along with gene-optimized nucleic acid sequences that encode candidate antigenic pathogen polypeptides.

[0272] methodology For a given virus species, candidate primary sequences are downloaded, e.g., from GenBank (and any other available sources, e.g., outbreak data), and filtered to exclude identical sequences, sequences that are outside the scope of the protein of interest, and sequences with many ambiguous nucleotides. A multiple sequence alignment of the filtered sequences is generated (typically using MAFFT) and manually checked to ensure that the sequences are in the correct open reading frame. A maximum likelihood phylogeny is generated using IQTREE with automated model selection and rooted using one of several methods: outgroup sequences, midpoint rooting, centre-of-the-tree, or a tree that maximizes the association between root-to-tip distance and sampling time. Ancestral sequences are generated using HyPhy assuming MG94 with an F3x4 model of codon substitution to check that known epitopes are preserved. A phylogenetic tree with both primary and ancestral sequences is generated using IQTREE to check the placement of ancestral strains. The ancestral sequence is then modified in several ways: by deletion of regions (e.g., removal of mucin-like domains); region swapping (to restore epitopes that may be lost); mutation of specific sites (e.g., in the fusion domain of filoviruses), including editing of N-linked glycosylation sites and introduction of mutations to enhance stability. EXAMPLES

[0273] Example 1 Ebola Sudan Ancestors (T2-4) Non-optimized (SEQ ID NO:1) [ka] [ka] Gene optimization (SEQ ID NO:2) [ka] Amino acid sequence encoded by the non-optimized and gene-optimized sequence (SEQ ID NO:3): [ka]

[0274] Example 2 Comprehensive ancestor of Ebola virus (T2-6) Non-optimized (SEQ ID NO:4) [ka] Gene optimization (SEQ ID NO:5) [ka] Amino acid sequence encoded by the non-optimized and gene-optimized sequence (SEQ ID NO:6): [ka]

[0275] Example 3 Marburg virus ancestor (T2-11) Non-optimized (SEQ ID NO:7) [ka] Gene Optimization (SEQ ID NO:8) [ka] Amino acid sequence encoded by the non-optimized and gene-optimized sequence (SEQ ID NO:9): [ka]

[0276] Example 4 Layer 2-4 (SUDV_anc_-MLD) Sudan Ebola virus ancestral sequence with mucin-like domain deleted (minus, "-") Nucleotide sequence (SEQ ID NO:10): [ka] Amino acid sequence (SEQ ID NO:11): [ka]

[0277] Example 5 Layer 2-6 (SUDV_EBOV-TAFV-BDBV_anc_-MLD) Ancestral sequences for four species, Sudan, Zaire, Tai Forest, and Bundibugyo Ebola viruses, lacking the mucin-like domain Nucleotide sequence (SEQ ID NO:12): [ka] [ka] Amino acid sequence (SEQ ID NO:13): [ka]

[0278] Example 6 Layer 2-11 (RAVV_MARV_anc) Ancestral sequences for strains Marburg virus and Raven virus Nucleotide sequence (SEQ ID NO:14) [ka] [ka] Amino acid sequence (SEQ ID NO:15): [ka]

[0279] Example 7 pEVAC expression vector Figure 3 shows a map of the pEVAC expression vector. The sequence of the multiple cloning site of the vector is shown below, followed by its complete nucleotide sequence. Sequence of pEVAC multiple cloning site (MCS) (SEQ ID NO: 16): [ka] The full sequence of pEVAC (SEQ ID NO:17): CMV-IE-E / P:248-989 CMV immediate early 1 enhancer / promoter KanR:3445-4098 Kanamycin resistance SD:990-1220 Splice Donor SA:1221-1343 splice acceptor Tbgh:1392-1942 Terminator signal from bovine growth hormone pUC-ori:2096-2769 pUC plasmid replication origin [ka] [ka] [ka]

[0280] Example 8 Lead candidate optimized antigenic Ebola polypeptides capable of inducing broadly neutralizing antibody responses There has been significant interest in developing a vaccine against Ebola following the 2014 West African outbreak. Programs currently in clinical development have so far taken the "classical approach" of vaccine development, using Ebola and / or Marburg virus surface glycoproteins (GPs) from one to three strains expressed in a viral vector backbone. Antigenic specificity is derived exclusively from the EBOV strain involved, e.g. Merck uses Kikwit GP; GSK uses Mayinga EBOV and Gulu SUDV strains; Crucell and Profectus Biosiences both use Marburg virus together with Zaire and Sudan Ebola strains; Novavax's approach is unique, using the 2014 Makona EBOV strain.

[0281] Table 1 below shows flow cytometry assay results illustrating the strength of antibody binding to target antigens representing all Ebola virus species (subtypes) and Marburg virus. The strength of binding is shown by a heat map with red (darkest shading when viewed in greyscale) being very strong binding, decreasing binding from orange to yellow (fading shading when viewed in greyscale) and no binding / equal to the negative control value being white. Serum samples 1-22 were taken from individuals immunized with other Ebola virus vaccine candidates. T2-4 and T2-6 are nucleic acid vaccines encoding lead candidate optimized antigenic Ebola polypeptides in combination with the T2-11 Marburg candidate, in preclinical stage testing with serum samples taken from immunized guinea pigs. Table 1 [Table 1]

[0282] Example 9 Protection achieved by trivalent Lassa, Ebola, and Marburg virus vaccine (Tri-LEMvac) in an Ebola challenge model The inventors have developed a trivalent vaccine (Tri-LEMvac) that will generate combined vaccine efficacy against future outbreaks of hemorrhagic fever Lassa, Ebola, and Marburg virus variants.

[0283] We bioinformatically designed synthetic glycoprotein sequences from the GPC open reading frames of LASV (L) and EBOV (E) and MARV (M) from all available arenavirus and filovirus databases. These conserved sequences consist of neutralizing antibody and T cell enriched epitopes of each of these viruses. To ensure that these synthetically designed LASV, EBOV, and MARV envelopes are functional and antigenic, we expressed them as pseudotypes and quality controlled for both binding and neutralization against a panel of broadly neutralizing antibodies. Herein, we select the vaccine-derived vector Modified Vaccinia Ankara (MVA) for the construction of a trimeric LEM vaccine.

[0284] The Modified Vaccinia Ankara (MVA) vaccine platform is a non-replicating strain (i.e., does not replicate in human cells) and is a third generation smallpox vaccine and one of the most advanced recombinant poxvirus vaccine vectors in human clinical trials (Cottingham & Carroll, Vaccine, 2013, 31(39):4247-51). MVA is a robust vector system capable of simultaneously expressing up to four transgenes, facilitating strong promoters and stable insertion sites (Orubu et al, Pone, 2012,7(6)e0040167). MVA was selected because: 1) it can stably express multiple independent ORFs via compatible expression cassettes with strong and time-regulated promoters for trivalent LEM vaccination in one cost-effective vaccine lot; 2) it can induce robust B and T cell immune responses in animals and humans, especially when primed or boosted by DNA or RNA vectors; and 3) it can heat-stabilize vaccine lots for storage and transportation in developing countries lacking a cold chain (Frey et al, Vaccine, 2015, 33(39):5225-34). Proof of principle of the trivalent vaccine candidate has been demonstrated by i) validation of cassettes for independent L, E, and M GPC expression and epitope presentation; and ii) preclinical efficacy with filovirus challenge. Challenge study results are shown in Figure 4. The Ebola challenge model was fatal to non-vaccinated guinea pigs (group 1, bottom line), but all vaccinated guinea pigs (group 2, top line) were protected (left) and continued to gain weight (right).

[0285] Example 10 Pseudotyped virus neutralization assay Figure 5 shows the results of a pseudotyped virus neutralization assay illustrating the strength of neutralizing antibody responses to target antigens expressed on the surface of pseudotyped viruses representing all Ebola virus species and Marburg virus. The strength of neutralization is shown by a heat map, with red (darkest shading when viewed on a grayscale) being very strong neutralization, orange to yellow (lighter shading when viewed on a grayscale) being decreasing neutralization, and white being no neutralization / equal to the negative control value.

[0286] T2-4 and T2-6 are nucleic acid vaccines each encoding a lead candidate optimized antigenic Ebola polypeptide in combination with the T2-11 Marburg candidate, currently in preclinical testing with serum samples collected from immunized guinea pigs.

[0287] The results indicate that administering a combination of the T2-6 and T2-11 vaccine inserts produced a synergistic increase in the breadth of the immune response.

[0288] Example 11 Antibody binding assays Figure 6 shows the results of the antibody binding assay. Antibody binding was measured by incubating two populations of cells carrying two different group 1 influenza A glycoproteins (H1 pandemic and seasonal) on their surface with pooled mouse serum. Any bound antibodies were then detected by a secondary antibody and the results were recorded using a flow cytometer. Binding was significantly increased before and after vaccination with all constructs, but not after vaccination with PBS (control). Overall, the DIOS vaccine candidate outperformed the performance from COBRA in both cases ( * ).

[0289] Example 12 Comparison of immune responses induced by two different computational approaches Four groups of six mice were immunized five times at two-week intervals with 25 μg of four separate pEVAC plasmids encoding HA gene antigens designed either by a method according to an embodiment of the invention (DIOS) or by conventional methods (COBRA).

[0290] Antibody-based FACS was performed on cells expressing two different group 1 influenza A glycoproteins (seasonal H1N1, and pandemic origin H1N1) on their surface. These were used to test mouse sera from animals immunized with either COBRA or DIOS HA gene antigens. The results are shown in Figure 7.

[0291] Overall, the DIOS HA gene antigen matched or significantly outperformed the COBRA HA gene antigen ( ** p<0.01, *** p<0.001).

[0292] (Example 13) Cross-HA group binding and pseudotype neutralization of H7N9 (A / Shanghai 2 / 2013) We tested whether the DIOS-H1N1pdm vaccine of Example 12 (which produces higher levels of antibody binding to pandemic H1 HA antigens than H1N1-COBRA) can elicit antibodies that recognize and bind to different group 2 viral HAs, such as the viral HA from the potentially pandemic H7N9 strain A / Shanghai / 2 / 2013.

[0293] Figure 8 shows the results of cross-HA group binding (left panel) and pseudotype neutralization (right) of H7N9 (A / Shanghai 2 / 2013) by sera from DIOS or COBRA DNA immunized mice. H7 binding data (left) confirmed by pseudotype neutralization data (right) shows that H1N1pdm vaccinated mice showed the best neutralization compared to other groups. Significantly more binding was induced by the DIOS-H1N1pdm vaccine than the other groups tested and was comparable to the positive control broadly neutralizing monoclonal antibodies F16 (Corti et al., 2011, Supra) and CR9114 (Dreyfus et al, Science, 2012; 337(6100): 1343-1348).

[0294] These results support the conclusion that the DIOS-H1N1pdm immunogen cross-neutralizes H7 and that cross-HA group immune protection is possible for vaccines produced by the methods of the invention.

[0295] Example 14 Lassa virus glycoprotein This example describes a Lassa virus glycoprotein ancestral sequence produced using methods according to an embodiment of the invention, and modifications of the ancestral sequence to improve its immunogenicity by stabilizing the structure.

[0296] Lassa fever is a hemorrhagic disease caused by an Old World (OW) arenavirus known as Lassa virus (LASV). The virus was first isolated in Nigeria in 1969 and is currently endemic in West Africa. Due to the high morbidity and mortality associated with Lassa hemorrhagic fever, LASV has been classified as a category A pathogen.

[0297] Lassa virus is an enveloped ambisense RNA virus with a bipartite genome. The virus particle is covered with a mature glycoprotein (GP) trimeric spike that mediates viral entry. Similar to other class 1 viral fusion proteins, the envelope glycoprotein precursor (GPC) is translated as a single polypeptide and proteolytically cleaved into three subunits. Processing occurs initially in the endoplasmic reticulum (ER) by cellular signal peptidases. GPC is then transported to the cis-Golgi apparatus and processed by the cellular proprotein convertase subtilisin kexin isozyme-1 / site-1 protease (SKI-1 / S1P) to produce a non-covalently linked stable signal peptide (SSP) / GP1 / GP2 heterotrimer. Unlike other class 1 fusion proteins, the relatively long signal peptide of GPC is not degraded, which serves as a chaperone-like function necessary for the correct transport and processing of GP. SSP interacts with the cytoplasmic domain of GP2 and is involved in pH sensing. GP1 is responsible for binding to cellular receptors, while GP2 mediates membrane fusion during viral entry.

[0298] The Lassa virus glycoprotein ancestral sequence for lineages III and IV (L-10) (construct 1) was generated using methods according to embodiments of the present invention. Modifications were then independently introduced into the parent ancestral sequence (construct 1) to provide (A) SOSEP (construct 2); and (B) FLEP (construct 4), as well as combinations with a glycan knockout called NtoK to stabilize other flexible heterotrimers and prevent dissociation of the glycoprotein ectodomain from the non-covalently linked transmembrane domain (providing constructs 3 and 5).

[0299] (A) Two cysteine ​​residues were introduced at positions 207 and 360 to form a disulfide bridge (SOS) between the external and transmembrane domains of GP. To facilitate complete cleavage of these two domains, the furin cleavage site was modified from RRLL to RRRR at positions 256-259. A glutamic acid to proline mutation (EP) at position 329 prevents structural rearrangements and makes the protein less flexible.

[0300] (B) The furin cleavage site (256-RRLL-259) between the C-terminus of the ectodomain and the N-terminus of the transmembrane domain was replaced by a flexible linker with the sequence 256-GGGGSGGGGS-265. In addition, the EP mutation in (A) was introduced at position 335.

[0301] Variants of both designs were generated that further contained an asparagine to lysine mutation at position 272 or 278 to inactivate the glycosylation motif for SOSEP-NtoK or FLEP-NtoK, respectively, which may block access of some neutralizing antibodies, such as 37.7H. Construct 1: Lassa virus glycoprotein ancestral sequences for lineages III and IV (L-10=LASV_III_IV_anc) Amino acid sequence (SEQ ID NO:18): [ka] DNA sequence (SEQ ID NO:19): [ka] [ka] Construct 2: SOSEP-variant of construct 1 (L-10-SOSEP) Amino acid sequence (SEQ ID NO:20): [ka] DNA sequence (SEQ ID NO:21): [ka] [ka] Construct 3: SOSEP variant of construct 1 with N to K mutation (L-10-SOSEP-NtoK) Amino acid sequence (SEQ ID NO:22): [ka] DNA sequence (SEQ ID NO:23): [ka] Construct 4: FLEP variant of construct 1 (L-10-FLEP) Amino acid sequence (SEQ ID NO:24): [ka] DNA sequence (SEQ ID NO:25): [ka] Construct 5: FLEP variant of construct 1 with N to K mutation (L-10-FLEP-NtoK) Amino acid sequence (SEQ ID NO:26): [ka] DNA sequence (SEQ ID NO:27): [ka] Example 15 Lassa virus nucleoprotein

[0302] This example describes a Lassa virus nucleoprotein ancestral sequence produced using methods according to an embodiment of the invention. Construct 6: Lassa virus nucleoprotein ancestral sequence of Nigerian Lassa isolates (L-NP-1=L-NP-CovAnc-1_N) Amino acid sequence (SEQ ID NO:28): [ka] DNA sequence (SEQ ID NO:29): [ka] [ka] Example 16 Lassa virus nucleoprotein

[0303] This example describes a Lassa virus nucleoprotein ancestral sequence produced using methods according to an embodiment of the invention. Construct 7: Lassa virus nucleoprotein ancestral sequence of Sierra Leone isolates (L-NP-1=L-NP-CovAnc-2_SL) Amino acid sequence (SEQ ID NO:30): [ka] DNA sequence (SEQ ID NO:31): [ka] [ka]

Claims

1. An isolated nucleic acid molecule comprising: i) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:2; or ii) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:8; An isolated nucleic acid molecule comprising a nucleic acid sequence or its complement.

2. An isolated polypeptide comprising: i) is at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:2, or is identical to the amino acid sequence encoded by SEQ ID NO:2; or ii) is at least 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence encoded by SEQ ID NO:8, or is identical to the amino acid sequence encoded by SEQ ID NO:8; An isolated polypeptide comprising an amino acid sequence.

3. An isolated polypeptide comprising: i) is at least 95%, 96%, 97%, 98%, or 99% identical to, or identical to, SEQ ID NO:3; or ii) is at least 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:9, or is identical to SEQ ID NO:9; An isolated polypeptide comprising an amino acid sequence.

4. An isolated nucleic acid encoding an amino acid sequence encoded by the nucleic acid of claim 1, wherein the nucleic acid has been codon-optimized and, optionally, gene-optimized for expression in a mammalian cell.

5. An isolated nucleic acid encoding a polypeptide according to claim 2 or 3, the nucleic acid being codon-optimized and, optionally, gene-optimized for expression in a mammalian cell.

6. A vector comprising the nucleic acid described in claim 1.

7. The vector described in claim 6, further comprising a promoter operably linked to the nucleic acid.

8. The vector described in claim 7, wherein the promoter is for expression of a polypeptide encoded by the nucleic acid in a mammal.

9. The vector described in claim 7, wherein the promoter is for expression of a polypeptide encoded by the nucleic acid in yeast or insect cells.

10. A vector described in any one of claims 6 to 9, which is a vaccine vector.

11. The vector of claim 10, which is a viral vaccine vector, a bacterial vaccine vector, an RNA vaccine vector, or a DNA vaccine vector.

12. An isolated cell comprising a vector described in any one of claims 6 to 9.

13. A pseudotyped viral particle comprising a polypeptide described in claim 2 or 3.

14. A method for producing a pseudotyped viral particle as described in claim 13, comprising a step of transfecting a vector as described in any one of claims 6 to 8 into a host cell.

15. A fusion protein comprising a polypeptide described in claim 2 or 3.

16. A pharmaceutical composition comprising the nucleic acid of claim 1, 4, or 5, and a pharma- ceutically acceptable carrier, excipient, or diluent.

17. A pharmaceutical composition comprising a vector according to any one of claims 6 to 8, 10, or 11, and a pharma- ceutically acceptable carrier, excipient, or diluent.

18. A pharmaceutical composition comprising the polypeptide of claim 2 or 3 and a pharma- ceutically acceptable carrier, excipient, or diluent.

19. A pharmaceutical composition described in any of claims 16 to 18, further comprising an adjuvant for enhancing an immune response in a subject to the polypeptide of the composition, or a polypeptide encoded by the nucleic acid.

20. A composition comprising a nucleic acid according to any one of claims 1, 4 or 5, a polypeptide according to claim 2 or 3, a vector according to any one of claims 6 to 8, 10 or 11, or a pharmaceutical composition according to any one of claims 16 to 19, for use as a pharmaceutical.

21. A composition comprising a nucleic acid according to any one of claims 1, 4 or 5, a polypeptide according to claim 2 or 3, a vector according to any one of claims 6 to 8, 10 or 11, or a pharmaceutical composition according to any one of claims 16 to 19, for use in treating a viral infection, preferably caused by an emerging or re-emerging virus, preferably a virus of the Filoviridae family.

22. Use of a composition comprising a nucleic acid according to any of claims 1, 4 or 5, a polypeptide according to claim 2 or 3, a vector according to any of claims 6 to 8, 10 or 11, or a pharmaceutical composition according to any of claims 16 to 19, in the manufacture of a medicament for the treatment of a viral infection, preferably caused by an emerging or re-emerging virus, preferably a virus of the Filoviridae family.

23. A composition or pharmaceutical composition for use as described in claim 21, or the use as described in claim 22, wherein the virus is Ebola virus or Marburg virus.