Virus particles

A new class of VLPs with Class II envelope glycoproteins forms large, pH-responsive structures for enhanced cargo capacity and cellular uptake, addressing size and immunogenicity limitations in current VLPs, suitable for therapeutic and diagnostic applications.

WO2026083093A1PCT designated stage Publication Date: 2026-04-23CAMBRIDGE ENTERPRISE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CAMBRIDGE ENTERPRISE LTD
Filing Date
2025-10-17
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Current VLPs are limited by size, immunogenicity, and difficulty in assembly/disassembly, which restricts their cargo capacity and cellular uptake, particularly for therapeutic and diagnostic applications.

Method used

Development of a new class of VLPs derived from endogenous retroviruses with Class II envelope glycoproteins, capable of forming large hollow icosahedral structures with pH-dependent disassembly, allowing controlled cargo release and cellular entry.

Benefits of technology

The new VLPs offer enhanced cargo capacity, reduced immunogenicity, and efficient cellular uptake, with pH-controlled cargo release, suitable for therapeutic and diagnostic applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000058_0000
    Figure 00000058_0000
  • Figure 00000059_0000
    Figure 00000059_0000
  • Figure 00000060_0000
    Figure 00000060_0000
Patent Text Reader

Abstract

The invention relates to a recombinant, synthetic, or isolated capsid protein, as well as isolated nucleic acids encoding the same. The invention also relates to virus-like particles (VLPs) comprising the capsid protein, methods of production, uses of the VLPs for delivery, and uses of the VLPs in therapy.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] VIRUS PARTICLES

[0002] Field of the Invention

[0003] The invention relates to a recombinant, synthetic, or isolated capsid protein, as well as isolated nucleic acids encoding the same. The invention also relates to virus-like particles (VLPs) comprising the capsid protein, methods of production, uses of the VLPs for delivery, and uses of the VLPs in therapy.

[0004] Background

[0005] VLPs are non-infectious self-assembled particles derived from virus capsid proteins. VLPs are desirable delivery tools because they are safer than intact viruses, relatively cheap to produce, and are formed from a capsid scaffold that is easy to engineer using widely-available laboratory techniques. VLPs mimic the structure and shape of viruses but lack their genetic material. For decades, VLPs have been widely used as a delivery vehicle in a wide variety of applications, such as gene delivery, drug delivery, antigen display, peptide conjugation, imaging, and as screening platforms. Most currently available retroviral VLPs are ~30 nm in diameter. VLPs are typically used to carry "cargo", which may be contained within the VLP; attached to the outer surface of the VLP; or both. Due to their widespread use, certain types of VLPs are associated with immunogenicity in subjects, typically human subjects.

[0006] There exists a significant and urgent need in the art for additional types of VLPs. In particular, there exists a need for VLPs which are larger than ~30nm in diameter, and therefore have a greater capacity to carry cargo. Indeed, VLP size is a key limitation in some gene delivery applications, and existing VLPs including the commonly used adeno-associated virus (AAV) and retrovirus VLPs often have a diameter which is too small for a gene to be packaged.

[0007] Simultaneously, there exists a need in the art for VLPs which are small enough to enter a cell - this is particularly important for therapeutic and diagnostic applications. There also exists a need for VLPs which have reduced or no immunogenicity in subjects, typically human subjects, including subjects who have previously been administered other existing types of VLPs. There also exists a need for VLPs which can be assembled and / or disassembled under pre-determined and physiologically relevant conditions.

[0008] Brief Description of the Figures

[0009] Figure 1. (A): BLAST search results for homologs of Rift Valley Fever Virus (RVFV) Glycoprotein C in all species excluding viruses. (B): Domain diagram of the Gag-Pol-Env polyprotein in hookworm Ancylostoma ceylanicum, showing in particular the two subdomains of the capsid (CA) domain: the N-terminal domain (NTD) and the C-terminal domain (CTD). (C): Cryo-EM image reconstruction of Atlas Gc (left) and structures of structurally similar Class II glycoproteins from other viruses and nonviral species. Right, cell-cell fusion assay showing the membrane fusion activity of Atlas Gc under low pH conditions. (D): Env-based tree of hookworm ERVs with class II Envs. (E): RT-based tree of hookworm ERVs with atypical Env glycoproteins.

[0010] Figure 2. (A): Methodology of expressing, purifying, and assembling the Atlas capsid proteins. (B) Expression profile showing both the tagged GST-CA, and the purified CA after removal of the GST tag. (C) SEC trace and a resulting gel in which an both capsid monomers and assembled particles are represented. Figure 3. (A) and (B): Cryo-EM images of assembled capsid protein nanoparticles. Arrows indicate smaller particles of 50-55 nm and larger particles of 65 nm.

[0011] Figure 4. (A): Cryo-EM image reconstructions of the smaller and larger Atlas capsids, at overall resolutions of 3.2 A and 4.2 A, respectively. (B) Atomic models built from cryo-EM data of the smaller assembled particle at a resolution of 3.1 A. A capsid monomer is shown alone and with cryo-EM density superimposed (left). Pentameric and hexameric monomers are also shown (right). (C): The differences between the size and arrangement of the Atlas capsid particles with other retroviral capsid particles.

[0012] Figure 5. Atomic model of the Atlas particle based on the cryo-EM image reconstruction. The asymmetric unit consisting of seven CA subunits is also shown with its relevant features.

[0013] Figure 6. (A): SEC trace showing co-purification of Atlas capsid proteins with nucleic acids. (B): Electrophoretic mobility shift assay showing that Atlas capsids (generated by concentration to 7 mg / ml) can interact with all four kinds of exogenous nucleic acids but preferentially package and protect single- and double-stranded RNA. (C): Detection of RNA co-purified with Atlas capsid by qRT- PCR, showing that Atlas capsids contain RNA encoding Atlas capsid, i.e. Atlas capsids package their own mRNA.

[0014] Figure 7. (A): HEK293T and iBMDMs in which the cytoplasmic membrane, nucleus, and the capsid nanoparticles are labelled. (B): Capsid particles entering iBMDMs at physiological temperature (37°C) after 90 minutes via the endocytic pathway. Capsids do not enter cells at 4°C, indicating that entry is via endocytosis. (C): Capsid particles (indicated by arrows) entering iBMDMs after 60 minutes via the endocytic pathway. (D): Capsid particles entering HeLa cells after 15 minutes via the endocytic pathway.

[0015] Figure 8. (A): Cryo-EM microscope images showing the stability of the particles under different conditions. (B): Particles releasing the packaged nucleic acids at low pH, where the released nucleic acid strands have been highlighted for better visualisation in the right-hand pane.

[0016] Figure 9. (A): BLAST results when searching for other self-assembling capsid particles, querying with the entire Atlas ERV sequence, and filtering for coverage above 70%. (B): Diagrammatic representation of highly conserved MHR sequences in retroviruses (except viruses from the Atlas family).

[0017] Figure 10. Phylogenetic tree constructed based on the core capsid domain of the 12 family members.

[0018] Figure 11. Table showing the sequences of the identified family of capsids, as well as the core capsid protein sequences. Coverage of the entire Atlas ERV, coverage over the full Atlas capsid protein, and the sequence identity and similarity scores with respect to the Atlas capsid core are shown for each result.

[0019] Figure 12. High-confidence structural predictions for each of the 12 family members using Alphafold, as well as each of these structures superimposed on experimentally validated Atlas CA core structure.

[0020] Figure 13. Table showing the RMSD and pLDDT scores for each predictive model generated using AlphaFold for each family member, as well as the percentage coverage of the superimposed models seen in Figure 12.

[0021] Figure 14. (A): The sequence of the capsid protein of KAK6016282.1 (from the nematode Ostertagia ostertagi) and a size exclusion chromatography trace showing the resulting assembled particles after expression and purification. (B): The sequence of the capsid protein of VDL73942.1 (from the nematode Nippostrongylus brasiliensis) and a size exclusion chromatography trace showing the resulting assembled particles after expression and purification.

[0022] Figure 15. Cryo-EM images of the particles formed by the capsid protein of KAK6016282.1. Arrows indicate assembled particles in the bottom pane.

[0023] Figure 16. Cryo-EM images of the particles formed by the capsid protein of VDL73942.1. Arrows indicate assembled particles in the bottom pane.

[0024] Figure 17. (A) The sequences of the N-terminal domain (NTD) and C-terminal domain (CTD) disordered tails of the capsid protein, as well as the resulting capsid protein sequence and structure) following their truncation. (B): Size exclusion chromatography traces in which the NTD and CTD have been truncated as well as associated EM images.

[0025] Figure 18. (A): Protein structural images in which key residues required for the interactions between hexamer chains are identified. (B): Table demonstrating the effects of site-directed mutagenesis within the capsid protein on the stability of the capsid protein. (C): Cryo-EM image showing that the particles containing the K162C mutation (relative to wild-type, WT) are more stable and release less nucleic acids when incubated at pH 5.0.

[0026] Figure 19. (A): Cryo-EM images of "MA-CA" particles containing the natural matrix domain fused to the N-terminus of the capsid domain, and "CA" particles which do not contain the matrix domain. (B): Table showing the results of attempted fusion of different peptides to the N-terminus of the CA and MA domains.

[0027] Figure 20. BLOSUM62 matrix used to calculate sequence similarity.

[0028] Figure 21. Defining the Atlas family of endogenous viruses from nematodes. (A): The genetic architecture of Atlas virus. MA=matrix, CA=capsid, NC=nucleocapsid, NTD=N-terminal domain, CTD=C-terminal domain. (B): The model of Atlas Gag, predicted by Alphafold 3. Only structured domains are shown. (C): The list of 11 belpaoviruses in the Atlas family, ranked by BLAST bitscore. (D) The phylogenetic tree of the Atlas family. Atlas virus (EPB78661) is indicated by an arrow. Tree was generated by iTOL v7.

[0029] Figure 22. Cryo-EM structures of four different Atlas capsid assemblies. (A): The cryo-EM micrographs of four types of in vitro assembled capsids from the three selected species. Top left, T = 1 capsids from N. brasiliensis (VDL73942); top right, T = 4 capsids from O. ostertagi (KAK6016282); bottom left, T = 7 capsids from A. ceylanicum (Atlas EPB78661); bottom right T = 12 capsids from A. ceylanicum (Atlas EPB78661). Scale bar, 50 nm. (B): The reconstructed maps of the four types of in vitro assembled capsids, shown in scale (80-360 nm). (C): The atomic model of T = 1 capsid from N. brasiliensis (VDL73942). Its asymmetric subunit (circled) consists of a single chain as 1 / 5 of a pentamer. Both the single model (left) and the model fitted in map (right) was shown. (D) The atomic model of T = 4 capsid from O. ostertagi (KAK6016282). Its asymmetric subunit (circled) consists of four chains as 1 / 5 of a pentamer + 1 / 2 of a hexamer. (E) The atomic model of T = 7 and T = 12 capsids from A. ceylanicum (EPB78661). The asymmetric subunits consist of 1 / 5 of a pentamer + 1 hexamer; and 1 / 5 of a pentamer + 1 + + 1 / 3 of a hexamer, respectively. (F): The intra-capsomere interactions of the hexamer core in T = 7, A. ceylanicum. Key residue side chains are shown. Hydrophobic stacking is shown by dashed lines. (G) The inter-capsomere interactions in T = 7, A. ceylanicum. Left: the dimer interface. Key residue side chains are shown. Hydrogen bonds are shown by dashed lines. Right: the three-fold axis. Figure 23. Nucleic acid packaging activities of Atlas capsid particles. (A): The size exclusion chromatography trace of A. ceylanicum capsid, from a Superdex 200 16 / 30 column (Cytiva). A280: protein absorption; A260: nucleic acid absorption. (B): Cryo-EM micrographs of A. ceylanicum capsids. Packed nucleic-acid-like materials are visible within the particles. (C): Electrophoretic mobility shift assay (EMSA) of A. ceylanicum capsids with or without application of extrinsic nucleic acid (ssRNA, dsRNA, and dsDNA); and with or without treatment of Benzonase nuclease. Nuclease- resistance demonstrates that nucleic acids is packaged within particles (the nuclease is too large to enter particles). (D): Quantification of nucleic acid protection, measured from band intensities of packaged nucleic acid in three biological repeats from gels in panel (C) above.

[0030] Figure 24. Live cell imaging of A. ceylanicum capsid uptake by HEK293T, HeLa, and iBMDM cells. (A): Top panel: HEK293T and HELA: fluorescently labeled capsids; plasma membrane, as indicated, and nucleus. Bottom panel: iBMDM: fluorescently labeled capsids; acid compartments; plasma membrane, as indicated, and nucleus. Scale bars 10 pm. (B): Fixed cell imaging of A. ceylanicum, O. ostertagi, and N. brasiliensis capsid uptake by iBMDM at 37 degrees and 4 degrees, with control of unconjugated free fluorescent dye. Fluorescently labelled capsids (indicated by arrows), and nucleus. Scale bars: 10 pm.

[0031] Figure 25. Structure-based engineering of new properties in Atlas virus particles. (A): K134C mutation at the dimer interface of the A. ceylanicum capsid particle designed to generate capsid particles with structurally stabilizing disulfide-crosslinks. (B): SDS-PAGE of WT and K134C Atlas capsids in reducing and non-reducing conditions. The band in non-reducing conditions shows that the K134C capsid mutant is disulfide-crosslinked as intended. (C): Cryo-EM micrographs of WT and K134C capsids at pH 7.4 or pH 4.5. Scale bar 50 nm. The K134C mutant is resistant to low pH-induced disassembly owing to the disulfide crosslinks. (D): Left: Model of an anti-CTLA-4 (Hll) nanobody-CA fusion protein predicted by AlphaFold 3. Centre: SDS-PAGE of purified Hll-CA. Right: negative stain EM images of Hll-CA particles. Scale bars 100 nm. The micrographs demonstrate that the anti-CTLA Hll nanobody is tolerated and does not interfere with capsid assembly. These particles are expected to specifically bind to CTLA-4-positive T-cells, which are enriched in various types of cancer.

[0032] Summary of the Disclosure

[0033] The present invention therefore addresses one or more of the above problems by providing newly-identified VLPs and capsid proteins, together with encoding nucleic acids, methods of production and assembly, uses of the VLPs for delivery, and uses of the VLPs in therapy.

[0034] The present invention is based on the surprising discovery of a new class of endogenous retrovirus (ERV), in nematodes, which contains a 'rigid' shell-forming Class II envelope (Env) glycoprotein.

[0035] Identification of a Class II Env glycoprotein in a retrovirus was unexpected, particularly because retroviruses have so far always been found to have a Class I Env glycoprotein, which does not form rigid outer shells. Moreover, the inventors encountered significant technical difficulties in identifying this new class of ERV, because its Class II Env glycoprotein has very low sequence identity to known (non-retroviral) Class II Env glycoproteins, such as the Class II Env glycoprotein of Rift valley fever virus.

[0036] A further unexpected discovery was that the capsid proteins of the newly-identified class of ERVs did not contain a major homology region (MHR), which is a 20-residue long amino acid sequence that is highly conserved across ERVs (see SEQ ID NO: 111). This is particularly surprising because the MHR is typically understood to be responsible for capsid assembly, and yet the newly identified class of ERV encode capsid proteins which (as noted below) produce VLPs without the presence of an MHR. The lack of an MHR is itself remarkable in a capsid-forming protein, as deletion of MHR in retroviruses typically abolishes capsid formation.

[0037] This newly-identified family of ERVs (detailed in the Examples) was found to encode capsid proteins which self-assemble to form large hollow icosahedral VLPs formed of hexamers and pentamers of monomeric capsid protein. Two forms of VLPs are produced from A. ceylanicum (Atlas EPB78661), both of which are larger than most existing retroviral VLPs (which usually have a size of ~ 30nm) - the larger form has a diameter of ~65 nm and is formed of 720 monomers of capsid protein (60 copies of a 12-monomer asymmetric unit, which form 1 / 5 of a pentamer + 1 hexamer + % of a hexamer + 1 / 3 of a hexamer); whereas the smaller form has a diameter of "'50-55 nm and is formed of 420 monomers of capsid protein (60 copies of a 7-monomer asymmetric unit, which form 1 / 5 of a pentamer + 1 hexamer). Both forms of VLPs were found to comprise 12 pentamers of capsid protein.

[0038] A third type of in vitro assembled capsid has also been produced. This capsid is derived from N. brasiliensis (VDL73942) and has an asymmetric unit consisting of a single monomer chain which forms 1 / 5 of a pentamer. This capsid has a diameter of 18-20 nm and is formed of 60 monomers.

[0039] A fourth type of in vitro assembled capsid has also been produced. This capsid is derived from O. Ostertagi (KAK6016282) and has a 4-monomer asymmetric unit which forms 1 / 5 of a pentamer and 1 / 2 of a hexamer. This capsid has a diameter of 35-40 nm and is formed of 240 monomers.

[0040] Therefore, the newly-identified family of ERVs encode capsid proteins with a wide range of sizes. The larger capsids are particularly well-suited to accommodating more cargo, whilst the smaller capsids are particularly well-suited to reaching deeper into target tissues.

[0041] Advantageously, the VLPs of the invention have a greater cargo carrying capacity than most existing retroviral VLPs, greatly enhancing their range of available cargoes, whilst retaining the ability to readily enter a cell via the endocytic pathway. Due to their icosahedral shape, VLPs of the invention are also ideally-suited to cellular uptake. These properties are highly advantageous for therapeutic and diagnostic applications.

[0042] Moreover, since VLPs of the invention belong to a newly-identified class of VLPs, they have not previously been administered to subjects, typically human subjects. This is also highly desirable because subjects, typically human subjects display no pre-existing immunity to VLPs of the invention (regardless of whether the subject has developed pre-existing immunity to other types of VLPs) and so the present invention increases the repertoire of VLPs for use in the clinic.

[0043] VLPs of the invention were also found to demonstrate pH-dependent control of their disassembly. Specifically, VLPs of the invention may be triggered to disassemble and release their cargo by lowering the pH of their surrounding environment. This is particularly advantageous because VLPs of the invention which enter an endocytic pathway are triggered to disassemble (and thereby release their cargo) when they encounter the low pH environment of a late endosome or endo-lysosome (as compared to releasing their packaged cargo too early, or not at all). Advantageously, the inventors also determined that the pH at which the VLPs assemble or disassemble may be further optimized according to the specific application, by site-directed mutagenesis of the capsid protein e.g. using their structural data as a guide. This provides further benefits as some molecules, such nucleic acids, e.g., RNA, are sensitive to lysosome degradation and require prolonged protection with a more pH- insensitive capsid. For small molecules that are less susceptible to degradation, it may be of benefit to promote an early release. Mutagenesis can be used to control the timing of lysosome exposure. Moreover, VLPs of the invention provide further significant advantages. For example, they are easy to produce, and their established structure readily allows protein engineering. They also have abundant reactive amino acid groups on the capsid surface, thereby providing significant flexibility in e.g. chemical ligation of peptides, proteins and / or small molecules on the particle surface. This is highly desirable for e.g. antibody delivery, antigen display, specific tissue targeting or drug delivery.

[0044] VLPs of the invention are also ideally-suited to use in the delivery of small molecules, drugs or nucleic acids, because they carry a positive charge under physiological pH and display electrostatic interactions with negatively charged molecules, such as nucleic acids. VLPs of the invention also display no sequence preference for nucleic acids, thereby providing additional flexibility; protection against nuclease activity, thereby improving nucleic acid availability in vivo,- and controlled release of packaged nucleic acids, thereby enhancing delivery targeting of packaged nucleic acids.

[0045] In one aspect, the invention provides a recombinant, synthetic, or isolated capsid protein, wherein the capsid protein comprises a polypeptide sequence having at least 50% sequence identity to SEQ ID NO: 1 (e.g. at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identity); wherein the polypeptide does not comprise a major homology region (M HR).

[0046] The major homology region (MHR), which is a 20-residue long amino acid sequence that is highly conserved across ERVs. The consensus sequence of the MHR, as discussed in Tanaka, et al. (2016). J Virol 90(4), 1944-1963, is represented by SEQ. ID NO: 111:

[0047] Q[G / K]X2EX4 / 5[Y / F]X2[R / G][F / L]X3H

[0048] (with X being any residue and H being a hydrophobic residue)

[0049] SEQ ID NO: 1 is the amino acid sequence of the "core" region of the EPB78661.1 Ancylostoma ceylanicum capsid protein:

[0050] ALLNFVDASILSKMELPTFDGNMLEFPEFASRFATLVGNKAELDDTTKFSLLKSCLRGRASHAIQGLSVTAENYKIAM DILNTHFNDKVTIKHVLYSKLAELPACDPEGRNLHTLYNRMFALIRQFANGNDDSKETGLGAILLNKLPLRVKSKIYDK TANSHNLSPSELLHLLTDIVRKDTTLQEM [SEQ ID NO: 1]

[0051] The skilled person can readily identify nucleic acids which encode amino acids described herein e.g. by reference to the corresponding portion of nucleic acids described herein.

[0052] Herein, the "core" region of a capsid protein refers to the region of the capsid protein that is required to achieve self-assembly into a VLP. The "core" region of a capsid lacks the disordered N- and C- terminal regions of the endogenous capsid protein.

[0053] In one aspect, the invention provides a recombinant, synthetic, or isolated capsid protein, wherein the capsid protein comprises a polypeptide having at least 75% sequence identity, optionally at least 90% sequence identity, to a sequence selected from SEQ ID NOs: 1-12; wherein the polypeptide does not comprise a major homology region (M HR).

[0054] In some embodiments, the capsid protein comprises a polypeptide having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identity to a sequence selected from SEQ ID NOs: 1-12; wherein the polypeptide does not comprise a major homology region (MHR). SEQ ID NO: 2 is the amino acid sequence of the "core" region of the KAK6016282.1 Ostertagia ostertagi capsid protein:

[0055] SLINFVDASILSKMELPTFDGNILEFPEFSSRFATLVGNKEELDNTTKFSLLKSCLRGAALNSIQGLSLTSENYKIAMDILK THYDDKVTIKHILYSRLADLPTCDPEGRNLHNLYNRMFALVRQFTDGNNDSNETGLGAILFNKLPLRVRSQIYDRTSN SHNLSPSELLNLLTDIVRKDTTLSEM [SEQ ID NO: 2]

[0056] SEQ. ID NO: 3 is the amino acid sequence of the "core" region of the VDL73942.1 Nippostrongylus brasiliensis capsid protein:

[0057] SNISFVDASILSKLDLPIFEGNLLEFPEYWARFSTLVGDKPQLDGATKFSLLKSTLRGRALQSIKGLSITAANYPIAVDILK NHFDDRVTTRHILYTRLASLPSCDKDGRNLFALYSQMFALVRQFTTYEDDSKEYALGAILLNKLPRHIRSRIYDLSGNSE NLVPTELLHILTRIVRKETTLEEM [SEQ ID NO: 3]

[0058] SEQ ID NO: 4 is the amino acid sequence of the "core" region of the KAK6031060.1 Ostertagia ostertagi capsid protein:

[0059] PLLNFVDASILTKMELPTFDGNLLEYPEFSARFATLVGNKPQLDNTTKFCLLKSCLRGRALQSIQGLSMTAENYNIAM DILRTHFDDKVTMRHILYTKLSQLPPCDPEGHHLPVLYNRMFSLVRQFCNGEDDSKETALGALLLNKLPLRVRSQIYD KTGNSHNVTPSELLHLLTDIVRKDSTLFEI [SEQ ID NO: 4]

[0060] SEQ ID NO: 5 is the amino acid sequence of the "core" region of the KAK6026028.1 Ostertagia ostertagi capsid protein:

[0061] TLLNFVDASILSKLELPTFDGNLLDYPEFWARFATLVDNKSQLDDTTKFSLLKSCLRGRALQSVQGLSLTSANYRIAVDI LKTHYDDKVTMRHILFTKLAQLPACDPEGRHLPTLYNRMFSLVRQFCNGYDDSKETALGALLLNKLPLRVRSQIYDRT SNSHNVTPSELLHLLTDIVRKDSTLFEI [SEQ ID NO: 5]

[0062] SEQ ID NO: 6 is the amino acid sequence of the "core" region of the KAK6028305.1 Ostertagia ostertagi capsid protein:

[0063] SLLNFVDASILTKLELPTFDGNLLEYPEFASRFATLVGNKTQLDNTTKLSLLKSCLRGRALQSIQGLSMTPENFAVAMDI LRTHYDDKVTMRHILYTKLAQLPDCDPEGRNLQTLYNRMFALVRQFANSNDDSSEEALGAILLNKLPARVKSRIYDM TTHSHNLSPSELLRLLTDIVRKESVLFEM [SEQ ID NO: 6]

[0064] SEQ ID NO: 7 is the amino acid sequence of the "core" region of the RCN35992. l Ancylostoma caninum capsid protein:

[0065] DTMNFVDASILSRLDLPTFDGNLLEFPEFFARFSALIGSKKQLDDTTKFSLLKSCLKGRALQSIHGLALTANNYAIALDIL KSRYDDKVTIRHILFSQLANLPPCDPEGRHLQSLYNKMYSLTRQFCVYEDDSKEVALGAILLNKLPRHIRSKIYDKTGN AHNLTPSELLQVLTSIVHKEATLQEI [SEQ ID NO: 7]

[0066] SEQ ID NO: 8 is the amino acid sequence of the "core" region of the EPB75047.1 Ancylostoma ceylanicum capsid protein:

[0067] SLCNFVDASLLSKVDLPLFSGSILEFQEFWERFSTLIGNKPHIDDATKFSLLKSSLKGRALHCIQGLPITSANYHIAVDILK THFDDRVTIRHVLFTKLASLPACDSAGKELQVLYNQMYALIRQFCTYEDDSKEYGLGAILLNKLPRHVRSRIYDKTNN QANLTPTALIQLLTDIVRKETTLREM [SEQ ID NO: 8]

[0068] SEQ ID NO: 9 is the amino acid sequence of the "core" region of the KAK6030023.1 Ostertagia ostertagi capsid protein: QMFNFVDASLLSKIDLPTFSGSILDFQEFWERFSILVGDKPNIDDATKFSLLKSSLRGKALQCIQGLSITSANYRIAVDIL

[0069] RTHFDDKVTTRHVLYTKLANLPPCDQAGKQLQPLYNQMFALIRQFCTYEDDNKEYGLGAVLLNKFPRHIRSKIYDKTS

[0070] NQTNLTPSALIRILTDIVKKESTLHEM [SEQ ID NO: 9]

[0071] SEQ. ID NO: 10 is the amino acid sequence of the "core" region of the KAK6018921.1 Ostertagia ostertagi capsid protein:

[0072] SNFNYFDASLLSRLDLPSFSGNLIEFPEFWSRFNTLVHSKSSLTGATKFSLLKSCLRGRALQCVEGLPITDQDYETAVDIL HLNYNNPSAIRHIIYTQLSALPQCDPEGKQLQDLYLKMLRLVRQYTTMTPSSPEYGLGALLYNKLPRFVKAKIYDKIGS QRNVTPNELM MLLSDIVKKETTLRQV [SEQ ID NO: 10]

[0073] SEQ ID NO: 11 is the amino acid sequence of the "core" region of the EPB74949.1 Ancylostoma ceylanicum capsid protein:

[0074] QNLNYFDASLLTRLDLPSFSGNLLEFPEFWARYHALIHCKTTLSGATKFSLLKSCLRGRALHTIDGLPVTDDNYAIAIDIL LTTYDNPSTLRHLIYTQLSSLPQCDPDGKQLQDLYLRMLRLVRQYTAITPYSPEFALGALLYNKLPRFVRARIYDMTGG QKNLTPSELITLLEEIVRKESTLRQM [SEQ ID NO: 11]

[0075] SEQ ID NO: 12 is the amino acid sequence of the "core" region of the KAK6009265.1 Ostertagia ostertagi capsid protein:

[0076] SLLNFVDASILAKLELPTFDGNLLEYPEFACRFATLVGNKTQLDDTTKLSLLKSCLRGRALQSIQGLSMTPENYRIAMDI LRTHYDDKVTVKHILYTKLAQLPDCDPEGQLTNNEEAQIALLPHPRTSCQHFANTVETTATMETPTDSCPKNEDATIV ACTTECTKNYNQQPLSHAALMCAPVRVFNPSDPSRRITATAFLDSGSSQSYITDDLAKLLNLSTL [SEQ ID NO: 12]

[0077] In one embodiment the capsid protein comprises a polypeptide having at least 75% sequence identity, optionally at least 90% sequence identity, to a sequence selected from SEQ ID NOs: 13-48, wherein the polypeptide does not comprise a major homology region (M HR).

[0078] In some embodiments, the capsid protein comprises a polypeptide having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identity to a sequence selected from SEQ ID NOs: 13-48; wherein the polypeptide does not comprise a major homology region (MHR).

[0079] SEQ ID NO: 13 corresponds to the core region and the C-terminal tail amino acid sequence of the EPB78661.1 Ancylostoma ceylanicum capsid protein:

[0080] ALLNFVDASILSKMELPTFDGNMLEFPEFASRFATLVGNKAELDDTTKFSLLKSCLRGRASHAIQGLSVTAENYKIAM DILNTHFNDKVTIKHVLYSKLAELPACDPEGRNLHTLYNRMFALIRQFANGNDDSKETGLGAILLNKLPLRVKSKIYDK TANSHNLSPSELLHLLTDIVRKDTTLQEMSHHTRSTTPQDQYLTFHASSKIRNKRAPPN [SEQ ID NO: 13]

[0081] SEQ ID NO: 14 corresponds to the core region and the C-terminal tail amino acid sequence of the KAK6016282.1 Ostertagia ostertagi capsid protein:

[0082] SLINFVDASILSKMELPTFDGNILEFPEFSSRFATLVGNKEELDNTTKFSLLKSCLRGAALNSIQGLSLTSENYKIAMDILK THYDDKVTIKHILYSRLADLPTCDPEGRNLHNLYNRMFALVRQFTDGNNDSNETGLGAILFNKLPLRVRSQIYDRTSN SHNLSPSELLNLLTDIVRKDTTLSEMSSHIRSVAEQDHYHTFHASSKTPRRKTATFGYRGTRKQPK [SEQ ID NO: 14]

[0083] SEQ ID NO: 15 corresponds to the core region and the C-terminal tail amino acid sequence of the VDL73942.1 Nippostrongylus brasiliensis capsid protein:

[0084] SNISFVDASILSKLDLPIFEGNLLEFPEYWARFSTLVGDKPQLDGATKFSLLKSTLRGRALQSIKGLSITAANYPIAVDILK NHFDDRVTTRHILYTRLASLPSCDKDGRNLFALYSQMFALVRQFTTYEDDSKEYALGAILLNKLPRHIRSRIYDLSGNSE NLVPTELLHILTRIVRKETTLEEMEDRSNYSSDIHVNAAILSNNKATTRQQPR [SEQ ID NO: 15] SEQ ID NO: 16 corresponds to the core region and the C-terminal tail amino acid sequence of the KAK6031060.1 Ostertagia ostertagi capsid protein:

[0085] PLLNFVDASILTKMELPTFDGNLLEYPEFSARFATLVGNKPQLDNTTKFCLLKSCLRGRALQSIQGLSMTAENYNIAM DILRTHFDDKVTMRHILYTKLSQLPPCDPEGHHLPVLYNRMFSLVRQFCNGEDDSKETALGALLLNKLPLRVRSQIYD KTGNSHNVTPSELLHLLTDIVRKDSTLFEIEYHSKQSPQLSNLHQSFVAKEGHQSNNSRPLPR [SEQ ID NO: 16]

[0086] SEQ. ID NO: 17 corresponds to the core region and the C-terminal tail amino acid sequence of the KAK6026028.1 Ostertagia ostertagi capsid protein:

[0087] TLLNFVDASILSKLELPTFDGNLLDYPEFWARFATLVDNKSQLDDTTKFSLLKSCLRGRALQSVQGLSLTSANYRIAVDI LKTHYDDKVTMRHILFTKLAQLPACDPEGRHLPTLYNRMFSLVRQFCNGYDDSKETALGALLLNKLPLRVRSQIYDRT SNSHNVTPSELLHLLTDIVRKDSTLFEIEYHTRRSADTKHIDYGFHTNARTLRPTSN [SEQ ID NO: 17]

[0088] SEQ ID NO: 18 corresponds to the core region and the C-terminal tail amino acid sequence of the KAK6028305.1 Ostertagia ostertagi capsid protein:

[0089] SLLNFVDASILTKLELPTFDGNLLEYPEFASRFATLVGNKTQLDNTTKLSLLKSCLRGRALQSIQGLSMTPENFAVAMDI LRTHYDDKVTMRHILYTKLAQLPDCDPEGRNLQTLYNRMFALVRQFANSNDDSSEEALGAILLNKLPARVKSRIYDM TTHSHNLSPSELLRLLTDIVRKESVLFEMDYHSKSNQTPHSQHHGFHVIANPKNQRQLQA [SEQ ID NO: 18]

[0090] SEQ ID NO: 19 corresponds to the core region and the C-terminal tail amino acid sequence of the RCN35992. l Ancylostoma caninum capsid protein:

[0091] DTMNFVDASILSRLDLPTFDGNLLEFPEFFARFSALIGSKKQLDDTTKFSLLKSCLKGRALQSIHGLALTANNYAIALDIL KSRYDDKVTIRHILFSQLANLPPCDPEGRHLQSLYNKMYSLTRQFCVYEDDSKEVALGAILLNKLPRHIRSKIYDKTGN AHNLTPSELLQVLTSIVHKEATLQEIDYHCKQTSRSYEEGYFGLHKSHNRLSSSTQFRRSPTH [SEQ ID NO: 19]

[0092] SEQ ID NO: 20 corresponds to the core region and the C-terminal tail amino acid sequence of the EPB75047. l Ancylostoma ceylanicum capsid protein:

[0093] SLCNFVDASLLSKVDLPLFSGSILEFQEFWERFSTLIGNKPHIDDATKFSLLKSSLKGRALHCIQGLPITSANYHIAVDILK THFDDRVTIRHVLFTKLASLPACDSAGKELQVLYNQMYALIRQFCTYEDDSKEYGLGAILLNKLPRHVRSRIYDKTNN QANLTPTALIQLLTDIVRKETTLREMEFQELDSSSRYTYHVIRNQPQMSRQKKITPS [SEQ ID NO: 20]

[0094] SEQ ID NO: 21 corresponds to the core region and the C-terminal tail amino acid sequence of the KAK6030023.1 Ostertagia ostertagi capsid protein:

[0095] QMFNFVDASLLSKIDLPTFSGSILDFQEFWERFSILVGDKPNIDDATKFSLLKSSLRGKALQCIQGLSITSANYRIAVDIL RTHFDDKVTTRHVLYTKLANLPPCDQAGKQLQPLYNQMFALIRQFCTYEDDNKEYGLGAVLLNKFPRHIRSKIYDKTS NQTNLTPSALIRILTDIVKKESTLHEM EPQYNEYSNNNVFHVVNDPN [SEQ ID NO: 21]

[0096] SEQ ID NO: 22 corresponds to the core region and the C-terminal tail amino acid sequence of the KAK6018921.1 Ostertagia ostertagi capsid protein:

[0097] SNFNYFDASLLSRLDLPSFSGNLIEFPEFWSRFNTLVHSKSSLTGATKFSLLKSCLRGRALQCVEGLPITDQDYETAVDIL HLNYNNPSAIRHIIYTQLSALPQCDPEGKQLQDLYLKMLRLVRQYTTMTPSSPEYGLGALLYNKLPRFVKAKIYDKIGS QRNVTPNELM MLLSDIVKKETTLRQVEQSSNSSHSTFYVSEKRGHAPPH [SEQ ID NO: 22]

[0098] SEQ ID NO: 23 corresponds to the core region and the C-terminal tail amino acid sequence of the EPB74949. l Ancylostoma ceylanicum capsid protein: QNLNYFDASLLTRLDLPSFSGNLLEFPEFWARYHALIHCKTTLSGATKFSLLKSCLRGRALHTIDGLPVTDDNYAIAIDIL LTTYDNPSTLRHLIYTQLSSLPQCDPDGKQLQDLYLRMLRLVRQYTAITPYSPEFALGALLYNKLPRFVRARIYDMTGG QKNLTPSELITLLEEIVRKESTLRQMDASTTSHQTSYHVSEKAGRHLRPIH [SEQ ID NO: 23]

[0099] SEQ. ID NO: 24 corresponds to the core region and the C-terminal tail amino acid sequence of the KAK6009265.1 Ostertagia ostertagi capsid protein:

[0100] SLLNFVDASILAKLELPTFDGNLLEYPEFACRFATLVGNKTQLDDTTKLSLLKSCLRGRALQSIQGLSMTPENYRIAMDI LRTHYDDKVTVKHILYTKLAQLPDCDPEGQLTNNEEAQIALLPHPRTSCQHFANTVETTATMETPTDSCPKNEDATIV ACTTECTKNYNQQPLSHAALMCAPVRVFNPSDPSRRITATAFLDSGSSQSYITDDLAKLLNLSTLTTEDISISTFGSTTP

[0101] LQVTSRDHLLGIVMETGEKRLKVKSLPT [SEQ ID NO: 24]

[0102] SEQ ID NO: 25 corresponds to the core region and the N-terminal tail amino acid sequence of the EPB78661.1 Ancylostoma ceylanicum capsid protein:

[0103] PLASQMPHPQRPPFSTSPNAVHAPYTNNALLNFVDASILSKMELPTFDGNMLEFPEFASRFATLVGNKAELDDTTKF SLLKSCLRGRASHAIQGLSVTAENYKIAMDILNTHFNDKVTIKHVLYSKLAELPACDPEGRNLHTLYNRMFALIRQFAN GNDDSKETGLGAILLNKLPLRVKSKIYDKTANSHNLSPSELLHLLTDIVRKDTTLQEM [SEQ ID NO: 25]

[0104] SEQ ID NO: 26 corresponds to the core region and the N-terminal tail amino acid sequence of the KAK6016282.1 Ostertagia ostertagi capsid protein:

[0105] PKASNYDNPLVDRPNPSNVTQQAHTPCNTSLINFVDASILSKMELPTFDGNILEFPEFSSRFATLVGNKEELDNTTKFS LLKSCLRGAALNSIQGLSLTSENYKIAM DILKTHYDDKVTIKHILYSRLADLPTCDPEGRNLHNLYNRMFALVRQFTDG NNDSNETGLGAILFNKLPLRVRSQIYDRTSNSHNLSPSELLNLLTDIVRKDTTLSEM [SEQ ID NO: 26]

[0106] SEQ ID NO: 27 corresponds to the core region and the N-terminal tail amino acid sequence of the VDL73942.1 Nippostrongylus brasiliensis capsid protein:

[0107] RSSNLSAHPTQPTTLPAIDTRTYIPPFQATDASNISFVDASILSKLDLPIFEGNLLEFPEYWARFSTLVGDKPQLDGATKF SLLKSTLRGRALQSIKGLSITAANYPIAVDILKNHFDDRVTTRHILYTRLASLPSCDKDGRNLFALYSQMFALVRQFTTYE DDSKEYALGAILLNKLPRHIRSRIYDLSGNSENLVPTELLHILTRIVRKETTLEEM [SEQ ID NO: 27]

[0108] SEQ ID NO: 28 corresponds to the core region and the N-terminal tail amino acid sequence of the KAK6031060.1 Ostertagia ostertagi capsid protein:

[0109] AKRATIHHEHRSPLLHQATATHLSPDSPLLNFVDASILTKMELPTFDGNLLEYPEFSARFATLVGNKPQLDNTTKFCLLK SCLRGRALQSIQGLSMTAENYNIAMDILRTHFDDKVTMRHILYTKLSQLPPCDPEGHHLPVLYNRMFSLVRQFCNGE DDSKETALGALLLNKLPLRVRSQIYDKTGNSHNVTPSELLHLLTDIVRKDSTLFEI [SEQ ID NO: 28]

[0110] SEQ ID NO: 29 corresponds to the core region and the N-terminal tail amino acid sequence of the KAK6026028.1 Ostertagia ostertagi capsid protein:

[0111] PTSASVHQSGM IPFSSALHPDPTLLNFVDASILSKLELPTFDGNLLDYPEFWARFATLVDNKSQLDDTTKFSLLKSCLR GRALQSVQGLSLTSANYRIAVDILKTHYDDKVTMRHILFTKLAQLPACDPEGRHLPTLYNRMFSLVRQFCNGYDDSK ETALGALLLNKLPLRVRSQIYDRTSNSHNVTPSELLHLLTDIVRKDSTLFEI [SEQ ID NO: 29]

[0112] SEQ ID NO: 30 corresponds to the core region and the N-terminal tail amino acid sequence of the KAK6028305.1 Ostertagia ostertagi capsid protein:

[0113] PTAAQASSSQLQGPSNSLLNFVDASILTKLELPTFDGNLLEYPEFASRFATLVGNKTQLDNTTKLSLLKSCLRGRALQSI

[0114] QGLSMTPENFAVAMDILRTHYDDKVTMRHILYTKLAQLPDCDPEGRNLQTLYNRMFALVRQFANSNDDSSEEALG

[0115] AILLNKLPARVKSRIYDMTTHSHNLSPSELLRLLTDIVRKESVLFEM [SEQ ID NO: 30] SEQ ID NO: 31 corresponds to the core region and the N-terminal tail amino acid sequence of the RCN35992.1 Ancylostoma caninum capsid protein:

[0116] DADASVQHSTIPPNKAAKTSQVSIQPMDTM NFVDASILSRLDLPTFDGNLLEFPEFFARFSALIGSKKQLDDTTKFSLL KSCLKGRALQSIHGLALTANNYAIALDILKSRYDDKVTIRHILFSQLANLPPCDPEGRHLQSLYNKMYSLTRQFCVYEDD SKEVALGAILLNKLPRHIRSKIYDKTGNAHNLTPSELLQVLTSIVHKEATLQEI [SEQ ID NO: 31]

[0117] SEQ. ID NO: 32 corresponds to the core region and the N-terminal tail amino acid sequence of the EPB75047.1 Ancylostoma ceylanicum capsid protein:

[0118] PATSEAPSLDSQNPKLGQRNDISMPVPFPNEQVTLPHHHGTTTIDQSSRSLCNFVDASLLSKVDLPLFSGSILEFQEF WERFSTLIGNKPHIDDATKFSLLKSSLKGRALHCIQGLPITSANYHIAVDILKTHFDDRVTIRHVLFTKLASLPACDSAGK ELQVLYNQMYALIRQFCTYEDDSKEYGLGAILLNKLPRHVRSRIYDKTNNQANLTPTALIQLLTDIVRKETTLREM [SEQ ID NO: 32]

[0119] SEQ ID NO: 33 corresponds to the core region and the N-terminal tail amino acid sequence of the KAK6030023.1 Ostertagia ostertagi capsid protein:

[0120] PTFLQHQHMVNVDPAQQNPHAAPYQMFNFVDASLLSKIDLPTFSGSILDFQEFWERFSILVGDKPNIDDATKFSLLK SSLRGKALQCIQGLSITSANYRIAVDILRTHFDDKVTTRHVLYTKLANLPPCDQAGKQLQPLYNQMFALIRQFCTYED DNKEYGLGAVLLNKFPRHIRSKIYDKTSNQTNLTPSALIRILTDIVKKESTLHEM [SEQ ID NO: 33]

[0121] SEQ ID NO: 34 corresponds to the core region and the N-terminal tail amino acid sequence of the KAK6018921.1 Ostertagia ostertagi capsid protein:

[0122] GYASDQQNFFTLPHAQQGFQESSNFNYFDASLLSRLDLPSFSGNLIEFPEFWSRFNTLVHSKSSLTGATKFSLLKSCLR GRALQCVEGLPITDQDYETAVDILHLNYNNPSAIRHIIYTQLSALPQCDPEGKQLQDLYLKMLRLVRQYTTMTPSSPEY GLGALLYNKLPRFVKAKIYDKIGSQRNVTPNELM MLLSDIVKKETTLRQV [SEQ ID NO: 34]

[0123] SEQ ID NO: 35 corresponds to the core region and the N-terminal tail amino acid sequence of the EPB74949.1 Ancylostoma ceylanicum capsid protein:

[0124] PTSTQPPLLPSSAYFHPTLPNSFNQPLPSTNHPLPALSTTRPNTYPSTFQQETQNLNYFDASLLTRLDLPSFSGNLLEFP EFWARYHALIHCKTTLSGATKFSLLKSCLRGRALHTIDGLPVTDDNYAIAIDILLTTYDNPSTLRHLIYTQLSSLPQCDPD GKQLQDLYLRMLRLVRQYTAITPYSPEFALGALLYNKLPRFVRARIYDMTGGQKNLTPSELITLLEEIVRKESTLRQM [SEQ ID NO: 35]

[0125] SEQ ID NO: 36 corresponds to the core region and the N-terminal tail amino acid sequence of the KAK6009265.1 Ostertagia ostertagi capsid protein:

[0126] MQNDNSLLNFVDASILAKLELPTFDGNLLEYPEFACRFATLVGNKTQLDDTTKLSLLKSCLRGRALQSIQGLSMTPEN YRIAMDILRTHYDDKVTVKHILYTKLAQLPDCDPEGQLTNNEEAQIALLPHPRTSCQHFANTVETTATMETPTDSCPK NEDATIVACTTECTKNYNQQPLSHAALMCAPVRVFNPSDPSRRITATAFLDSGSSQSYITDDLAKLLNLSTL [SEQ ID NO: 36]

[0127] SEQ ID NO: 37 corresponds to the complete amino acid sequence of the EPB78661.1 Ancylostoma ceylanicum capsid protein:

[0128] PLASQMPHPQRPPFSTSPNAVHAPYTNNALLNFVDASILSKMELPTFDGNMLEFPEFASRFATLVGNKAELDDTTKF SLLKSCLRGRASHAIQGLSVTAENYKIAMDILNTHFNDKVTIKHVLYSKLAELPACDPEGRNLHTLYNRMFALIRQFAN GNDDSKETGLGAILLNKLPLRVKSKIYDKTANSHNLSPSELLHLLTDIVRKDTTLQEMSHHTRSTTPQDQYLTFHASSKI RNKRAPPN [SEQ ID NO: 37] SEQ ID NO: 38 is the complete amino acid sequence of the KAK6016282.1 Ostertagia ostertagi capsid protein:

[0129] PKASNYDNPLVDRPNPSNVTQQAHTPCNTSLINFVDASILSKMELPTFDGNILEFPEFSSRFATLVGNKEELDNTTKFS LLKSCLRGAALNSIQGLSLTSENYKIAM DILKTHYDDKVTIKHILYSRLADLPTCDPEGRNLHNLYNRMFALVRQFTDG NNDSNETGLGAILFNKLPLRVRSQIYDRTSNSHNLSPSELLNLLTDIVRKDTTLSEMSSHIRSVAEQDHYHTFHASSKT PRRKTATFGYRGTRKQPK [SEQ ID NO: 38]

[0130] SEQ. ID NO: 39 is the complete amino acid sequence of the VDL73942.1 Nippostrongylus brasiliensis capsid protein:

[0131] RSSNLSAHPTQPTTLPAIDTRTYIPPFQATDASNISFVDASILSKLDLPIFEGNLLEFPEYWARFSTLVGDKPQLDGATKF SLLKSTLRGRALQSIKGLSITAANYPIAVDILKNHFDDRVTTRHILYTRLASLPSCDKDGRNLFALYSQMFALVRQFTTYE DDSKEYALGAILLNKLPRHIRSRIYDLSGNSENLVPTELLHILTRIVRKETTLEEM EDRSNYSSDIHVNAAILSNNKATTR QQPR [SEQ ID NO: 39]

[0132] SEQ ID NO: 40 is the complete amino acid sequence of the KAK6031060.1 Ostertagia ostertagi capsid protein:

[0133] AKRATIHHEHRSPLLHQATATHLSPDSPLLNFVDASILTKMELPTFDGNLLEYPEFSARFATLVGNKPQLDNTTKFCLLK SCLRGRALQSIQGLSMTAENYNIAMDILRTHFDDKVTMRHILYTKLSQLPPCDPEGHHLPVLYNRMFSLVRQFCNGE DDSKETALGALLLNKLPLRVRSQIYDKTGNSHNVTPSELLHLLTDIVRKDSTLFEIEYHSKQSPQLSNLHQSFVAKEGH QSNNSRPLPR [SEQ ID NO: 40]

[0134] SEQ ID NO: 41 is the complete amino acid sequence of the KAK6026028.1 Ostertagia ostertagi capsid protein:

[0135] PTSASVHQSGM IPFSSALHPDPTLLNFVDASILSKLELPTFDGNLLDYPEFWARFATLVDNKSQLDDTTKFSLLKSCLR GRALQSVQGLSLTSANYRIAVDILKTHYDDKVTMRHILFTKLAQLPACDPEGRHLPTLYNRMFSLVRQFCNGYDDSK ETALGALLLNKLPLRVRSQIYDRTSNSHNVTPSELLHLLTDIVRKDSTLFEIEYHTRRSADTKHIDYGFHTNARTLRPTSN [SEQ ID NO: 41]

[0136] SEQ ID NO: 42 is the complete amino acid sequence of the KAK6028305.1 Ostertagia ostertagi capsid protein:

[0137] PTAAQASSSQLQGPSNSLLNFVDASILTKLELPTFDGNLLEYPEFASRFATLVGNKTQLDNTTKLSLLKSCLRGRALQSI QGLSMTPENFAVAMDILRTHYDDKVTMRHILYTKLAQLPDCDPEGRNLQTLYNRMFALVRQFANSNDDSSEEALG AILLNKLPARVKSRIYDMTTHSHNLSPSELLRLLTDIVRKESVLFEMDYHSKSNQTPHSQHHGFHVIANPKNQRQLQ A [SEQ ID NO: 42]

[0138] SEQ ID NO: 43 is the complete amino acid sequence of the RCN35992.1 Ancylostoma caninum capsid protein:

[0139] DADASVQHSTIPPNKAAKTSQVSIQPMDTM NFVDASILSRLDLPTFDGNLLEFPEFFARFSALIGSKKQLDDTTKFSLL KSCLKGRALQSIHGLALTANNYAIALDILKSRYDDKVTIRHILFSQLANLPPCDPEGRHLQSLYNKMYSLTRQFCVYEDD SKEVALGAILLNKLPRHIRSKIYDKTGNAHNLTPSELLQVLTSIVHKEATLQEIDYHCKQTSRSYEEGYFGLHKSHNRLSS STQFRRSPTH [SEQ ID NO: 43]

[0140] SEQ ID NO: 44 is the complete amino acid sequence of the EPB75047.1 Ancylostoma ceylanicum capsid protein:

[0141] PATSEAPSLDSQNPKLGQRNDISMPVPFPNEQVTLPHHHGTTTIDQSSRSLCNFVDASLLSKVDLPLFSGSILEFQEF WERFSTLIGNKPHIDDATKFSLLKSSLKGRALHCIQGLPITSANYHIAVDILKTHFDDRVTIRHVLFTKLASLPACDSAGK ELQVLYNQMYALIRQFCTYEDDSKEYGLGAILLNKLPRHVRSRIYDKTNNQANLTPTALIQLLTDIVRKETTLREM EFQ

[0142] ELDSSSRYTYHVIRNQPQMSRQKKITPS [SEQ ID NO: 44]

[0143] SEQ. ID NO: 45 is the complete amino acid sequence of the KAK6030023.1 Ostertagia ostertagi capsid protein:

[0144] PTFLQHQHMVNVDPAQQNPHAAPYQMFNFVDASLLSKIDLPTFSGSILDFQEFWERFSILVGDKPNIDDATKFSLLK SSLRGKALQCIQGLSITSANYRIAVDILRTHFDDKVTTRHVLYTKLANLPPCDQAGKQLQPLYNQMFALIRQFCTYED DNKEYGLGAVLLNKFPRHIRSKIYDKTSNQTNLTPSALIRILTDIVKKESTLHEMEPQYNEYSNNNVFHVVNDPN [SEQ ID NO: 45]

[0145] SEQ ID NO: 46 is the complete amino acid sequence of the KAK6018921.1 Ostertagia ostertagi capsid protein:

[0146] GYASDQQNFFTLPHAQQGFQESSNFNYFDASLLSRLDLPSFSGNLIEFPEFWSRFNTLVHSKSSLTGATKFSLLKSCLR GRALQCVEGLPITDQDYETAVDILHLNYNNPSAIRHIIYTQLSALPQCDPEGKQLQDLYLKMLRLVRQYTTMTPSSPEY GLGALLYNKLPRFVKAKIYDKIGSQRNVTPNELM MLLSDIVKKETTLRQVEQSSNSSHSTFYVSEKRGHAPPH [SEQ ID NO: 46]

[0147] SEQ ID NO: 47 is the complete amino acid sequence of the EPB74949.1 Ancylostoma ceylanicum capsid protein:

[0148] PTSTQPPLLPSSAYFHPTLPNSFNQPLPSTNHPLPALSTTRPNTYPSTFQQETQNLNYFDASLLTRLDLPSFSGNLLEFP EFWARYHALIHCKTTLSGATKFSLLKSCLRGRALHTIDGLPVTDDNYAIAIDILLTTYDNPSTLRHLIYTQLSSLPQCDPD GKQLQDLYLRMLRLVRQYTAITPYSPEFALGALLYNKLPRFVRARIYDMTGGQKNLTPSELITLLEEIVRKESTLRQMDA STTSHQTSYHVSEKAGRHLRPIH [SEQ ID NO: 47]

[0149] SEQ ID NO: 48 is the complete amino acid sequence of the KAK6009265.1 Ostertagia ostertagi capsid protein:

[0150] MQNDNSLLNFVDASILAKLELPTFDGNLLEYPEFACRFATLVGNKTQLDDTTKLSLLKSCLRGRALQSIQGLSMTPEN YRIAMDILRTHYDDKVTVKHILYTKLAQLPDCDPEGQLTNNEEAQIALLPHPRTSCQHFANTVETTATMETPTDSCPK NEDATIVACTTECTKNYNQQPLSHAALMCAPVRVFNPSDPSRRITATAFLDSGSSQSYITDDLAKLLNLSTLTTEDISIS TFGSTTPLQVTSRDHLLGIVMETGEKRLKVKSLPT [SEQ ID NO: 48]

[0151] The skilled person can readily identify nucleic acids which encode capsid proteins of the invention. Exemplary nucleic acids are provided by SEQ ID NOs: 49-60.

[0152] SEQ ID NO: 37 may be encoded by the nucleic acid sequence of SEQ ID NO: 49: cctctcgcgtcacagatgcctcatcctcaacgtcctccgttcagcacgtctccgaacgcagtccacgctccctacaccaacaacgctcttctcaacttt gtcgacgcttcaatcctaagcaaaatggagctcccaacgttcgacggaaacatgttggaatttcccgaattcgcatcacgattcgccactctcgttgg aaacaaagcagaacttgatgatacgacaaaattctcactactcaagtcatgcctacgtggccgcgcgtcgcacgccatccaaggactatcggttac tgcagaaaactacaaaatcgctatggatatcctcaatactcacttcaacgacaaagtcaccatcaaacacgtcctgtattcgaagttagcagaact ccccgcctgcgacccagaaggacgcaaccttcacacgctctacaatcgaatgttcgctctcatcagacaatttgcgaatggcaacgatgactcaaa ggagaccggacttggcgcaatactcttgaacaagttgccgctccgagtaaaaagcaaaatatacgacaagacagcgaactcccacaacctctcg ccaagtgaactacttcatctactaacggacatcgttcgcaaggacacaactctgcaagagatgagccaccacacacggtcgacaactccgcaaga tcaatatctgacgttccatgcatcatcaaaaataaggaataagcgagcaccaccgaac [SEQ ID NO: 49]

[0153] SEQ ID NO: 38 may be encoded by the nucleic acid sequence of SEQ ID NO: 50: cctaaagccagcaattacgacaatccactcgtcgatagaccgaatccatccaatgtgactcaacaggccatcatcactccttgcaatacgtcactca tcaattttgtcgacgcatcaatactgagcaaaatggaactacctacattcgatgggaacatactcgaattccctgaattttcgtctcgcttcgcaactc tggtcggaaacaaggaagaactcgacaacacgaccaaattttcgctgttgaagtcatgcttacgaggagcggctctgaattccatacaaggactat cgctgacgtccgagaattacaagatcgctatggacatcctgaagacacattatgacgacaaggtaacaatcaagcacatactttattcaagattgg ccgatcttcccacctgcgatcccgaaggacgcaatcttcacaatctctacaatcgaatgtttgcgctcgtacgacagttcacagatgggaacaatga ctcgaacgaaactggattgggagcgatcctcttcaacaaacttcctctacgcgtaagaagccaaatttacgatagaacgtcgaactctcacaatct atcaccgagtgagttgcttaatctcctgacggatattgttcggaaggacacaacattatccgaaatgtcgtcccatatacgctcagtggccgaacag gatcactaccatacatttcatgcctcttcaaaaactccaaggaggaaaactgcaacgtttggataccgcggtaccagaaagcagccaaag [SEQ ID NO: 50]

[0154] SEQ. ID NO: 39 may be encoded by the nucleic acid sequence of SEQ ID NO: 51: cgctcgtctaacctctctgctcacccaactcagcctactacgctacccgctattgatacaaggacgtatattcctccattccaagccactgatgcttca aatatttcgttcgtcgacgcgtccatactgagtaaactcgacctgccgatattcgaaggtaaccttctcgaatttcctgaatattgggcccgtttctcta ctctcgttggcgacaaaccacagctagatggtgctacgaaattttcacttctgaaatctacactacgcggacgagccctacagtcaatcaaaggac tgtccataaccgctgccaactatccgatcgctgttgatattctgaaaaatcacttcgatgaccgcgtcactacacgacatattctctacactcgtctgg cttcacttccttcatgtgataaagacggaagaaatctattcgctctttattcgcagatgttcgcacttgttcgccaattcacaacgtatgaagacgactc caaagaatacgccctaggagctattctactaaacaaactaccacgtcatattcgaagccgcatctacgatctaagtggaaacagcgaaaatctcgt ccctactgaactccttcacatccttacgagaatcgtcagaaaagaaacaacactagaggaaatggaagatcgctccaactactcttcagatattca cgtgaatgccgccatcctgtccaacaacaaggcaacaacaagacagcaaccgaga [SEQ ID NO: 51]

[0155] SEQ ID NO: 40 may be encoded by the nucleic acid sequence of SEQ ID NO: 52: gcaaagagagcaactatccatcacgagcatcgatcgcctctactacatcaagcaacagctacgcatttgtccccggactcccctttactcaacttcg tcgacgcttccattttaacaaaaatggaattacccacattcgatggaaacttactggagtatcccgaattttcggcgcgatttgctactttagtgggaa acaaaccacaactcgacaatacgacaaaattctgcctgctcaaatcttgcctgcgcggacgagctctccagtcaatacaaggcctatccatgaccg cggagaactacaatatcgcaatggatatacttcgtacgcacttcgacgacaaggtcaccatgagacacatattgtacaccaaactatcgcaactgc cgccgtgcgatccagaaggtcatcaccttcctgtgctgtacaacaggatgttctcgctagtgagacagttttgcaatggcgaagacgactcaaagg aaacagctctaggagcactcctcctcaacaagttaccccttcgagttcgaagccaaatttacgacaaaacggggaacagccacaatgtcactccg agcgaacttttgcatctcctgaccgacatcgtccgcaaagattcgacgctcttcgaaatcgagtatcactcaaaacaatcgccacaactgagcaac ctgcaccagagttttgtcgccaaagaaggacatcagtcaaacaactcgcgcccgcttcctcga [SEQ ID NO: 52]

[0156] SEQ ID NO: 41 may be encoded by the nucleic acid sequence of SEQ ID NO: 53: cccacttctgcctctgtacaccaatctggcatgatcccattttcatctgcacttcaccccgaccccacgcttctcaatttcgtcgacgcctccattctgtc caagttggagcttcctacattcgatggcaatcttctcgactacccggaattctgggcacgttttgctaccttagtggacaacaagtcacaactagacg atactacgaaattctcgctgctaaagtcctgcttgcgcggtcgcgctcttcagtctgttcaaggattatcactgacatctgcgaactacagaatagct gtcgacatcctcaaaacccactacgatgacaaggtcactatgaggcatatacttttcaccaagctggctcagcttccagcatgcgatcctgaaggtc gtcatcttcctactctctataatcggatgttttctcttgttcgtcagttctgcaacggctacgacgactccaaagaaactgctctcggtgcacttcttctc aacaagctaccactacgcgtacgcagtcagatctacgacagaacttcgaatagccacaacgttactcccagcgagctacttcaccttctcacggac atcgtcagaaaagattccacactattcgaaattgaataccacacacgacgaagcgctgataccaagcatatcgactacggtttccacaccaatgct cgtactctgcgtccaacttcgaat [SEQ ID NO: 53]

[0157] SEQ ID NO: 42 may be encoded by the nucleic acid sequence of SEQ ID NO: 54: ccaacggcagcacaggcatcatcatcacaactacaaggtcctagcaattcgctcctcaactttgtcgacgcctctatactgacaaaactggaactg ccaaccttcgatggtaatcttctcgagtatcccgagtttgcatcacgattcgctaccctcgtcggcaacaaaacacaattggacaacactacgaaac tgtcattgctcaaatcgtgcttacgaggacgtgcactgcagtcaatccagggattgtcaatgacaccagaaaactttgcagtcgccatggacatttta cggacccactacgacgacaaggttaccatgaggcatattctctacactaaacttgcacaacttcccgactgtgatccggaaggacgaaatttacaa acactctacaatcgcatgtttgctctcgttcgacaattcgctaatagcaacgacgattcgagtgaagaggcgctaggagcaatattgctcaacaaac tacctgcacgagtgaaaagcagaatttatgacatgacgacacactcccataatctctcaccaagtgaacttctacgcctactcacggatattgttcg caaggagtcagttctcttcgaaatggattatcattccaagtcaaaccagacaccacatagccagcatcacggcttccatgtcatcgcgaatccgaaa aatcaacggcaactacaagca [SEQ ID NO: 54]

[0158] SEQ ID NO: 43 may be encoded by the nucleic acid sequence of SEQ ID NO: 55: gacgcagatgcgagcgttcaacattcaacaatacctcccaacaaggctgcgaaaaccagtcaagtaagcatacaaccgatggatacaatgaact ttgtagacgcgtcgatcttgagcaggttggatctacccaccttcgatggaaatctactagaatttcctgaattcttcgctcgtttctctgccctcatcgg aagcaaaaaacaacttgacgataccacaaagttttcgctcctaaaatcatgtctgaaaggccgagccctccagtcaattcatggattagcgcttac cgccaataactacgcaatagctctggacatactcaagtcgaggtacgatgataaagtcaccattcgtcacatactcttcagtcaactggcaaacctt cccccatgcgacccagaaggcaggcatcttcaatccctgtacaacaagatgtattcactcacacgccagttctgcgtctatgaagatgactcgaaa gaagttgcccttggtgcaattcttctcaacaaactaccacgtcatatacgcagcaaaatctacgacaagaccggcaacgcgcacaacctcactcca agtgaacttctccaagtactgaccagcatcgtccacaaagaagcgactctacaggaaatcgactaccactgtaagcagacaagtcgctcctatga agaaggatactttggactccataaatcgcacaatagactttcttcatctacccagtttcgccgttctcctactcac [SEQ ID NO: 55]

[0159] SEQ. ID NO: 44 may be encoded by the nucleic acid sequence of SEQ ID NO: 56: cctgctacatctgaagctccctctctcgactcgcagaacccaaaattgggacaacgaaacgacatctcgatgccagtgccctttcccaacgaacaa gttacgcttccacatcatcatggaactaccacgatcgaccaatcctcgcgaagtctttgcaatttcgtcgacgcgtccctcctcagcaaagtggatct cccgttattcagtggatcaattctcgagtttcaagaattctgggaacgtttctccaccttgatcgggaataaaccacatatcgatgacgcaaccaagt tttctctgctgaaatcttcgctcaagggaagagcactccactgcatccaaggtctcccgattacatccgccaattaccacatcgcagtcgatattctc aaaacgcacttcgacgatcgagtgacaatccggcatgtgctgttcaccaaattggcgagtctaccagcatgcgattcagcaggaaaagagctaca ggtactgtacaaccaaatgtacgcgctaatacgacaattttgcacttatgaggacgacagcaaggaatatggtctaggagctattcttctgaataa actaccacgccacgtgaggagccgaatctacgacaaaaccaacaaccaagccaacctcacaccaacggcactcattcaacttctcacggatatc gtcagaaaggaaaccacattgcgagaaatggaattccaagagctagactcttctagcaggtatacatatcacgtcatccgtaaccaaccgcaaat gagtcgccaaaagaagatcacgccatcc [SEQ ID NO: 56]

[0160] SEQ ID NO: 45 may be encoded by the nucleic acid sequence of SEQ ID NO: 57: ccaacgttcctgcaacatcaacacatggtcaacgtcgatcctgcgcagcaaaatccacatgctgcaccctaccaaatgttcaatttcgttgacgcat ctcttctcagcaagatcgatctaccaacattctcaggatctattctcgacttccaagaattctgggaacgtttttccatacttgttggagacaaaccaa acatcgacgatgctaccaaattctctctcctaaaatcttcgctcagaggaaaagctctccagtgcatacaaggactgtcaattacttcagccaacta ccgcatcgctgttgacatcctcagaacgcactttgacgacaaagtcacgactcgtcatgtactttacacaaaactggcaaatttaccaccatgcgac caagctggaaaacaactgcagccgctctacaaccagatgtttgcgctaatacgacaattttgtacctacgaggacgacaacaaagaatatggtct cggagccgttctgctgaacaaatttccacggcatatcagaagtaagatctacgacaagacctccaatcagaccaatcttacgccttctgctcttattc gcatactcaccgacatcgtcaaaaaggagtcaactctgcacgaaatggagccgcaatacaacgagtacagcaacaacaacgtcttccacgtcgt caacgatcccaat [SEQ ID NO: 57]

[0161] SEQ ID NO: 46 may be encoded by the nucleic acid sequence of SEQ ID NO: 58: gggtacgcctccgatcaacagaacttcttcaccttgcctcatgcacaacaaggatttcaagagagctccaatttcaactactttgatgcgtcgctact ctcacgtttggatctcccatcattttcgggaaatctcatcgaatttccagaattttggtccagattcaacactctcgttcactcgaagtcgtcgcttacag gagcaactaaattttcactgctcaagagctgtttaagaggacgtgcacttcaatgcgtagaaggtcttccgatcacagaccaggattacgaaactg cggtcgatatactacatttgaactacaacaatccgtccgctattcgacacatcatctacactcaactttcggctttaccacagtgtgatcctgaagga aaacaacttcaagacctatacctcaagatgctacgcttagttcgacaatacacgactatgacacccagctctccagaatatggattaggagctcttc tatacaacaagctaccaagattcgttaaggctaaaatttacgacaaaattggaagtcaaagaaacgtcacacccaacgaactgatgatgcttctat cggacatcgttaagaaagagactacactgcggcaggttgaacaatcttccaactccagtcactccacattctacgtctccgaaaaacgaggacatg cacctccacac [SEQ ID NO: 58]

[0162] SEQ ID NO: 47 may be encoded by the nucleic acid sequence of SEQ ID NO: 59: ccgaccagtactcaacctccgctactaccatcaagcgcgtacttccacccgactctaccgaactccttcaatcaaccactcccttcaaccaaccatc ctctgccggctttgtcgaccacccgtcccaacacctatccttccacgttccaacaggagacccaaaacctcaactacttcgacgcgtccttactcact cgcctggacctaccatcattctcaggaaaccttctagagtttcccgaattttgggctcgatatcacgccctcattcactgcaagacgacactctctgg agcgaccaaattctctctactgaaaagctgtctcagaggacgagcacttcatacgattgatggcctcccagtgaccgacgacaactacgcgatcgc catcgacatcctactcactacatacgataacccatccactctacgtcatctcatctacacccagttatcgagtctcccgcagtgcgacccagatggca aacaactccaagacctttacctgaggatgctccgactcgtacggcaatacaccgcgatcacgccctattcgcccgagttcgcgctaggagcactac tatacaacaaactcccgcgtttcgtcagagcaagaatctatgacatgactggaggacagaaaaacctaacaccttccgagttgataaccctgctcg aggaaatcgtgagaaaggaatcaacgttacgacaaatggacgcttcgaccacgtcacaccaaacgtcgtaccacgtatcagagaaagcaggac gtcacctccggcctattcac [SEQ ID NO: 59]

[0163] SEQ. ID NO: 48 may be encoded by the nucleic acid sequence of SEQ ID NO: 60: atgcagaacgacaattcgctcctcaatttcgtcgacgcatcgattttggctaaattggaattacctacattcgatggcaacttacttgagtatccagaa tttgcctgcagatttgcaactctcgttgggaacaaaacacagctcgatgacacaaccaaactatctcttctcaagtcttgcctacgaggtcgcgccct acaatccatccaaggattatcgatgacgccagaaaactatcgcatagcgatggatattctccggacacactacgacgacaaagtaactgtcaaac acatactctacaccaaactcgctcagcttcctgattgtgatccagaaggacaactcacaaacaacgaagaagcacaaatcgctctccttcctcatc cacgaacatcatgccaacatttcgctaatactgtagagactactgcaaccatggaaacgccaaccgattcctgcccaaagaacgaagacgcgac aatagttgcctgcacgactgaatgcacaaagaattacaatcagcaacctctgtcacatgccgcccttatgtgcgcgcctgtgcgagtgttcaatcca tctgatccctcacgtcgaataaccgcgacagcttttctggattcaggatccagtcagtcttatataacagatgatttggcaaagctcctcaatctttcc actctgacaacggaagatatctcaatatcaacattcggctccactacacccctccaagtcacatctcgcgatcacctgctcggcatcgtcatggaga ctggagagaaacgtctcaaagtgaaatctctcccaacc [SEQ ID NO: 60]

[0164] The proteins of the invention include protein sequences that have been removed from their naturally occurring environment (which are "isolated"), recombinantly expressed, and chemically synthesized proteins (e.g. cell-free protein synthesis or solid-phase chemical synthesis).

[0165] When applied to a protein, the term "isolated" in the context of the present invention denotes that the protein has been removed from its natural cellular milieu and is thus free of other extraneous or unwanted polypeptide sequences. Such isolated proteins are those that are separated from their natural environment.

[0166] An example of an algorithm that is suitable for determining sequence similarity is the BLAST algorithm, which is described in Altschul et al., 1990, J. Mol. Biol. 215:403-410. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information. This algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. These initial neighborhood word hits act as starting points to find longer HSPs containing them. The word hits are expanded in both directions along each of the two sequences being compared for as far as the cumulative alignment score can be increased. Extension of the word hits is stopped when: the cumulative alignment score falls off by the quantity X from a maximum achieved value; the cumulative score goes to zero or below; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLAST program uses as defaults a word length (W) of 11, the BLOSUM62 scoring matrix (see Henikoff & Henikoff, 1992, Proc. Nat'l. Acad. Sci. USA 89:10915-10919) alignments (B) of 50, expectation (E) of 10, M'5, N'-4, and a comparison of both strands.

[0167] Any of the amino acid sequences described herein can be produced together or in conjunction with at least 1, e.g., at least (or up to) 2, 3, 5, 10, or 20 heterologous amino acids flanking each of the C- and / or N-terminal ends of the specified amino acid sequence, and or deletions of at least 1, e.g., at least (or up to) 2, 3, 5, 10, or 20 amino acids from the C- and / or N-terminal ends.

[0168] Conservative substitutions can be chosen from among a group of amino acids having a similar side chain to the reference amino acid. For example, a group of amino acids having aliphatic side chains is glycine, alanine, valine, leucine, and isoleucine; a group of amino acids having aliphatic-hydroxyl side chains is serine and threonine; a group of amino acids having amide-containing side chains is asparagine and glutamine; a group of amino acids having aromatic side chains is phenylalanine, tyrosine, and tryptophan; a group of amino acids having basic side chains is lysine, arginine, and histidine; and a group of amino acids having sulphur-containing side chains is cysteine and methionine. Accordingly, exemplary conservative substitutions for each of the naturally occurring amino acids are as follows: Ala to Ser; Arg to Lys; Asn to Gin or His; Asp to Glu; Cys to Ser or Ala; Gin to Asn; Glu to Asp; Gly to Pro; His to Asn or Gin; He to Leu or Vai; Leu to He or Vai; Lys to Arg; Gin or Glu; Met to Leu or lie; Phe to Met, Leu or Tyr; Ser to Thr; Thr to Ser; Trp to Tyr; Tyr to Trp or Phe; and, Vai to lie or Leu.

[0169] Spacers

[0170] In some embodiments, the capsid protein is attached to a spacer. Typically, the spacer comprises an amino acid sequence that is not present in the full-length capsid protein sequence.

[0171] In some embodiments, a spacer is attached to the N- terminus of the capsid protein. In some embodiments, a spacer is attached to the C- terminus of the capsid protein. In some embodiments, a spacer is attached to the N- and C-termini of the capsid protein. In some embodiments, a spacer is attached to the capsid protein between the N- and C- terminus of the capsid protein.

[0172] The inventors believe that the N-terminal tail of the capsid protein is located on the outside of the VLP. Thus, attachment of the spacer to the N-terminus of the capsid protein is understood to locate the spacer substantially at the exterior of the VLP. Attachment of cargo to an N-terminal spacer typically locates the cargo at the exterior of the VLP.

[0173] The inventors believe that the C-terminal tail of the capsid protein is located on the inside of the VLP. Thus, attachment of the spacer to the C-terminus of the capsid protein is understood to locate the spacer substantially at the interior of the VLP. Attachment of cargo to a C-terminal spacer typically locates the cargo at the interior of the VLP.

[0174] In some embodiments, the spacer is attached to cargo. In some embodiments, spacer is not attached to cargo.

[0175] In some embodiments, the spacer is attached directly to the N-terminus of a core region of a capsid protein described herein. In some embodiments, the spacer is attached directly to the C-terminus of a core region of a capsid protein described herein. In some embodiments, the spacer is attached directly to the N-terminal tail of a capsid protein described herein. In some embodiments, the spacer is attached directly to the C-terminal tail of a capsid protein described herein.

[0176] In some embodiments, the spacer is a polypeptide spacer. In some embodiments, the spacer comprises at least 2 amino acids, at least 3 amino acids, at least 4 amino acids, at least 5 amino acids, at least 10 amino acids, at least 20 amino acids, at least 30 amino acids, at least 40 amino acids, at least 50 amino acids, at least 60 amino acids, at least 70 amino acids, at least 80 amino acids, at least 90 amino acids, at least 100 amino acids, at least 110 amino acids, at least 120 amino acids, at least 130 amino acids, at least 140 amino acids, at least 150 amino acids, at least 160 amino acids, at least 170 amino acids, at least 180 amino acids, at least 190 amino acids, at least 200 amino acids. Typically, the spacer is less than 350 amino acids. In some embodiments, the spacer comprises 10- 350 amino acids. In some embodiments, the spacer comprises 10-300 amino acids. In some embodiments, the spacer comprises 10-250 amino acids. In some embodiments, the spacer comprises 10-200 amino acids. In some embodiments, the spacer comprises 10-150 amino acids. In some embodiments, the spacer comprises 10-100 amino acids. In some embodiments, the spacer comprises 50-350 amino acids. In some embodiments, the spacer comprises 50-300 amino acids. In some embodiments, the spacer comprises 50-250 amino acids. In some embodiments, the spacer comprises 50-200 amino acids. In some embodiments, the spacer comprises 50-150 amino acids. In some embodiments, the spacer comprises 50-100 amino acids.

[0177] In some embodiments, the spacer comprises 100-350 amino acids. In some embodiments, the spacer comprises 100-300 amino acids. In some embodiments, the spacer comprises 100-250 amino acids. In some embodiments, the spacer comprises 100-200 amino acids. In some embodiments, the spacer comprises 100-150 amino acids.

[0178] In some embodiments, the spacer comprises the amino acid sequence of a matrix protein, or a fragment thereof. In some embodiments, the spacer comprises a matrix protein or a fragment thereof and an additional (non-capsid protein) sequence at the N-terminus of the matrix protein or fragment thereof. In some embodiments, the spacer comprises a matrix protein or a fragment thereof and an additional (non-capsid protein) sequence at the C-terminus of the matrix protein or fragment thereof. In some embodiments, the spacer comprises a matrix protein or a fragment thereof and an additional (non-capsid protein) sequence at the N- and C- termini of the matrix protein or fragment thereof.

[0179] Matrix proteins herein describe those proteins which, in nature, form links between viral nucleocapsids and viral envelopes, and assist in maintaining viral (or VLP) structure and stability. Matrix proteins in this embodiment may be native (matrix protein A as linked to VLP capsid A), or foreign (matrix protein B as linked to VLP capsid A).

[0180] Although matrix proteins typically play a role in viral assembly, VLPs of the invention do not require the presence of a matrix protein. However, the inventors discovered that attaching cargo to the capsid protein via a matrix protein spacer helps facilitate the formation of VLPs.

[0181] In some embodiments, the matrix protein spacer comprises an amino acid sequence selected from [SEQ ID NOs: 61-71].

[0182] In one embodiment the capsid protein comprises a polypeptide having at least 75% sequence identity, optionally at least 90% sequence identity, to a sequence selected from SEQ. ID NOs: 49-59.

[0183] SEQ ID NO: 61 corresponds to the amino acid sequence of the matrix protein of EPB78661.1 Ancylostoma ceylanicum:

[0184] MPQDNSSKIRRQIGFFKKLIQRSCASIPQTFTDYDIDEKHPKFDRLDEDQLESLRLELHNIRSSLLKAYSRITSLHDEWT ALQQSDPRESKQFDEYITKYGDYRTSVTQAVTQLEELDLLLNEVDNEFRGRDLSVSSDSSENPVSNGRFDVKSSQQD DQ [SEQ ID NO: 61]

[0185] SEQ ID NO: 62 corresponds to the amino acid sequence of the matrix protein of KAK6016282.1 Ostertagia ostertagi:

[0186] MSDSAKIRRQIGFLKKQIQRYCSSIAQTLTDYSVDAKNSEFTQLDNDQLESLRFELQSTKSNLLRAYNRITFLHDEWAIL QESDANEAQLFQDYVTKYGDYRISASEAVAQLENLDLLLDDLEAEFHKRNISISSGSSDNTDDQLN [SEQ ID NO: 62]

[0187] SEQ ID NO: 63 corresponds to the amino acid sequence of the matrix protein of VDL73942.1 Nippostrongylus brasiliensis:

[0188] MSGIIRQQIGITKRQLRKALQESEGEHFDPDQVQKLKDDELLAAYESHTSAHDSLFRIYARLNRLWKQWEELMKENP

[0189] EEEEILKQYVTKYGDFRLLLNNAVSALERLDQERPLIEAELRKKKLDFEPYSESDVDSQ [SEQ ID NO: 63] SEQ ID NO: 64 corresponds to the amino acid sequence of the matrix protein of KAK6031060.1 Ostertagia ostertagi:

[0190] MSSNDSSQIRRQIGFFKKQLQRHGLSVTTTLKEYHIKTEQLDFRHLENDELESLRAEIVPLRRNLLKSYQKITKLNDEW TTLQDSNTGEQEIFNEYISKYGDYRESISTSVLQLEGLDILLNAIDQEYIKRNIHVPSDLSDATSLEDYGNELTS [SEQ. ID NO: 64]

[0191] SEQ ID NO: 65 corresponds to the amino acid sequence of the matrix protein of KAK6026028.1 Ostertagia ostertagi:

[0192] MKIITVRKSLLNSYEKITKLHDEWTILQQSVATETTNFDEYIAKYGDYRESITAAVNQLEQLDYLMNALDHEYSKRNLH IPSDSSDTASHDDNERHWN [SEQ ID NO: 65]

[0193] SEQ ID NO: 66 corresponds to the amino acid sequence of the matrix protein of KAK6028305.1 Ostertagia ostertagi:

[0194] MSPNDSIKIRRQLGFYKKQIQRLCSTATLTAKEYDVNARQPNFESLDDDEIEAFRLEVTTVRHNLLKAYSQITKLHDEW VLLQETDPAEVQQFNNYISKYGDYRNTITEAVAKLEEMDLLLSATDKECHRRHLSISSASSAVTIPDENGQRHQATAAI A [SEQ ID NO: 66]

[0195] SEQ ID NO: 67 corresponds to the amino acid sequence of the matrix protein of RCN35992.1 Ancylostoma caninum:

[0196] MFNSGKTRRHIGILKKRLLQQRNNIKDILQEYGIQEKNPVFSSLNNDEIYDFKNELLEAHQQLLRSYTKIQELHMEWS AVRITDKKEEETYNEYINTYGDYSSTVQQAVDTLDNNDELFNAIDVELSQRHLPVEARKDDLSTVNPRH [SEQ ID NO: 67]

[0197] SEQ ID NO: 68 corresponds to the amino acid sequence of the matrix protein of EPB75047.1 Ancylostoma ceylanicum:

[0198] MADSQTIRRQIGKLSRQLPRHIGAVYDVLADYELTIENVQLNNIDSNDLPMLRKDCLNARSNLLITYTRLEQLHQNW ISLIAANKEEEAVFSQFIDKYGDYRAALQDAVPILEKLDTILDAIEEEFKRRQLPLPGFIPTT [SEQ ID NO: 68]

[0199] SEQ ID NO: 69 corresponds to the amino acid sequence of the matrix protein of KAK6030023.1 Ostertagia ostertagi:

[0200] MDDSAAIRRQIGKVTRQVARCTAAVYDLFTDYDLTLDNLQLDAVASDDLPNFTKDVNLVRSNLLRYYDRLERLHQQ WQAIISTNPTEEGTFSTYIDKYGDYRHNLQQAVTILEKLDSTVDQLHHEFKKRDMPLPVPTPATSDAASHDSNYDNA KATEAIST [SEQ ID NO: 69]

[0201] SEQ ID NO: 70 corresponds to the amino acid sequence of the matrix protein of KAK6018921.1 Ostertagia ostertagi:

[0202] MSAPIRRQIGALRKLLARAITHCEEDLHENLSDFVASCSNDSLLDYVEAQNAHRDAILSKLTRLEQLNDEWSNLMTQ DQDEVQNFHNFIEKYGDYREDISKAIATLEKLDSNEAILREDMEKRGIKYETEVVQESSQRHTTVHCEDRRDCRTQN NITN [SEQ ID NO: 70]

[0203] SEQ ID NO: 71 corresponds to the amino acid sequence of the matrix protein of EPB74949.1 Ancylostoma ceylanicum:

[0204] MSAQLRRKIGIARKQLTRAVNLCEEDLKDDLPGFLQTATNDDLLDYAELQDTHHESLSTKLRKLEDLNNEWIALMAK DSSEVATFHEFISKYGDYRDDIEKAIKALQRM DSSEALIQEELEKRGIAHAFNTYSTQHSQNDGNFPHEGNKEHVTSA PPHREGLPPL [SEQ ID NO: 71] The skilled person can readily identify nucleic acids which encode matrix proteins described herein. Exemplary nucleic acids are provided by SEQ ID NOs: 72-82.

[0205] SEQ ID NO: 61 may be encoded by the nucleic acid sequence of SEQ. ID NO: 72: atgccacaggacaatagctccaaaataagaagacaaatcggcttcttcaaaaaactgatccagcgaagctgcgcgtctataccacaaaccttcac ggattacgacatcgatgagaagcatccgaagttcgatcgcttagatgaagaccagctcgaatccctccgtcttgagctccacaatatcagatcatcc ctgctgaaagcatactcaaggataacctcccttcacgacgaatggaccgctcttcagcaatccgacccgcgtgaaagcaaacaattcgacgagta tatcaccaaatatggagactacagaacttccgtcactcaagctgtgactcagctagaagaactcgatctcctactcaatgaagtggacaacgagtt tcgaggacgcgatcttagcgtgtcttcagattcgtcagaaaacccagtgtcaaatggcagatttgatgtcaagagcagccaacaagatgaccaa [SEQ ID NO: 72]

[0206] SEQ ID NO: 62 may be encoded by the nucleic acid sequence of SEQ ID NO: 73: atgtcggacagtgccaaaatccgtcgacagattggattcctcaagaagcagattcagagatactgctcatcaatcgcccaaacactgacagactac agcgtcgacgccaagaattccgaattcacgcaactggacaatgatcagttggaatcgctacgatttgaacttcaaagcaccaaatccaatcttctca gagcatacaaccgcatcacgtttttacatgacgaatgggcaatactacaagaatccgacgccaacgaagctcaactattccaagactacgtcacc aaatatggagattaccgcatatcggcatcagaagcagtcgctcaattagaaaacctcgacctccttctcgacgatctggaagctgaattccataaa cgtaacatcagtatctcttctggctcttcagacaacacagatgatcaactcaat [SEQ ID NO: 73]

[0207] SEQ ID NO: 63 may be encoded by the nucleic acid sequence of SEQ ID NO: 74: atgtccggcattattcgtcaacaaataggcatcaccaaacgacagctacgaaaggcactccaagagtccgaaggtgagcacttcgatcccgatca agttcaaaaattgaaagacgatgaacttctcgcagcctatgaatcacatacatccgcacacgactctctctttcgcatttacgcccgactgaatcgtc tttggaaacaatgggaagaacttatgaaggaaaacccggaagaggaagaaatactcaaacaatacgtaaccaaatacggagattttcgactact tctcaacaatgccgtatctgcgctggaacgacttgaccaagagcgaccactcatcgaagccgaactcaggaagaaaaaactggacttcgaaccg tactccgaatctgacgtcgactctcag [SEQ ID NO: 74]

[0208] SEQ ID NO: 64 may be encoded by the nucleic acid sequence of SEQ ID NO: 75: atgtcatccaacgacagctctcaaatacggcgtcagattggattcttcaagaaacagctccaacgtcatgggttgtcagtaacaaccacactgaag gagtaccacataaaaaccgaacaactggacttccgtcatctagaaaacgacgaacttgaatcactacgagctgagatcgttccactgagaagaaa ccttctgaagtcctaccaaaagatcacaaaattaaacgacgagtggacaacacttcaggactccaacacaggcgagcaagaaattttcaacgaat atatatcaaagtatggtgactaccgagaatcaatctctacatcagtcctccagttagagggcttggatatcctgcttaatgccattgatcaggaatac atcaagaggaacatacacgtgccttccgatctatctgatgctacttctctggaggactacggaaacgaattgacgtca [SEQ ID NO: 75]

[0209] SEQ ID NO: 65 may be encoded by the nucleic acid sequence of SEQ ID NO: 76: atgaagatcatcaccgtcagaaaaagccttctgaactcctacgaaaaaattacaaagctacatgacgaatggacaatcctacagcaatctgtagc aactgagactaccaattttgacgaatatattgctaaatatggcgactaccgtgaatcaataactgctgcagtgaaccaactcgaacaactcgacta cctgatgaacgcactcgaccatgaatactccaaaagaaatctccacattccgtcagactcctccgacaccgcttcccacgacgacaacgaaagac actggaat [SEQ ID NO: 76]

[0210] SEQ ID NO: 66 may be encoded by the nucleic acid sequence of SEQ ID NO: 77: atgtcgcccaacgacagtatcaaaataagacgtcagctgggcttctataagaagcaaattcaacggctatgctcgaccgcgacgctcactgcaaa agaatatgacgtcaacgcaaggcagcccaatttcgagagcctcgatgacgatgaaatcgaagcattccgactcgaagtcactacagtgcgccaca acttactcaaagcttattcgcaaataacgaaactccatgatgaatgggttctccttcaagaaaccgatcctgctgaagttcaacaattcaacaacta catctccaaatatggtgattatcgaaacacaattaccgaagcggtcgccaaactagaagaaatggatctactcctaagcgcgactgataaagaat gccatcgccggcatttgagtatttcctcggcatcttcagcagtaaccataccagatgaaaatggtcaacgtcatcaagcaacagccgcaatcgcc [SEQ ID NO: 77]

[0211] SEQ ID NO: 67 may be encoded by the nucleic acid sequence of SEQ ID NO: 78: atgttcaatagcggaaagacccggcgccatattggcatattgaagaaacgcctccttcaacaacgcaacaatatcaaggacatcctacaagagta cggaattcaagagaagaacccagtattctcatcactcaataacgatgaaatctatgacttcaagaacgaactcctcgaggctcaccaacaactgct ccgctcatacaccaaaatccaagagctccacatggaatggtcagcagtacgaatcaccgacaaaaaagaagaagaaacctacaacgaatacat caatacctacggagactactcgtccactgtgcagcaagccgtcgatacgttggacaataacgatgaattattcaatgccatcgacgtcgaactcagt caacgtcatctaccagtggaagcccggaaagatgacctttcaactgtcaaccctcgccat [SEQ ID NO: 78]

[0212] SEQ. ID NO: 68 may be encoded by the nucleic acid sequence of SEQ ID NO: 79: atggccgactcacagaccatacgccgccaaatcgggaagctctcaagacagctgccgagacatattggagcagtttacgacgttcttgccgactac gagctgacaatcgagaatgtgcaactcaacaatatcgattctaacgacctgccaatgttgcggaaggattgcctgaacgcccgatcgaatctgctc atcacctataccagactagagcaacttcaccagaactggatttccctaattgccgcgaacaaagaggaagaagcggtattctcgcagttcatcgat aaatatggagactatcgcgccgcactacaagatgcagtcccgattctcgagaaattggacaccattcttgacgcaattgaggaagaattcaagaga cgccaacttccactaccaggtttcatcccaaccaca [SEQ ID NO: 79]

[0213] SEQ ID NO: 69 may be encoded by the nucleic acid sequence of SEQ ID NO: 80: atggacgattctgctgcgatcagacgacaaataggaaaggtcactcgacaagttgctcgatgtacagctgcggtctacgacttattcactgactatg acctaactctcgacaacttacaactcgatgctgtcgcctctgatgaccttcctaacttcacaaaggacgtcaatcttgttcgatccaacctcctacgct actacgaccgactcgaacgtcttcatcaacaatggcaagcaattatttctaccaatccgactgaagaaggcactttcagtacatatattgataaata tggagactaccgccacaacctccagcaagctgttacaatactcgagaaactggattcaactgtcgaccaacttcatcacgaattcaaaaaaaggg acatgccactccctgtacctacacctgcgacttctgatgcagcatcacatgactcgaactacgacaacgccaaagcaacagaagccatctcaacg [SEQ ID NO: 80]

[0214] SEQ ID NO: 70 may be encoded by the nucleic acid sequence of SEQ ID NO: 81: atgagcgcaccgatccgccgtcaaataggagctcttcgcaaacttctcgcacgtgccatcacccactgtgaagaagatctacacgaaaatctatcc gacttcgttgcctcctgctcaaacgactccttgctcgactacgtggaagcacagaacgcacatcgagatgcaattctgtcgaaactcacgaggctcg aacaactcaatgacgaatggagcaacttgatgactcaagaccaagacgaggttcagaattttcacaacttcatagaaaagtatggagactacagg gaagatattagcaaagccatcgctactcttgaaaaactcgactccaacgaagcaattctccgtgaagatatggagaaaagaggaatcaaatacg aaaccgaggtggtacaagaatccagtcaacgacacaccactgttcactgcgaagacaggagagactgccgcacccagaacaacatcacgaat [SEQ ID NO: 81]

[0215] SEQ ID NO: 71 may be encoded by the nucleic acid sequence of SEQ ID NO: 82: atgagcgcacaactccgcagaaaaatcggaatcgcccgtaagcaattgacacgagctgttaatctatgcgaagaagatctcaaggatgaccttcc cggcttccttcagaccgcaacaaacgacgacctcctcgactatgctgaacttcaagacacccatcacgagtcgttgtccacgaaactgagaaaact cgaggacctcaacaacgaatggattgccctaatggccaaagactccagcgaggtagccacttttcacgaattcatctccaagtatggcgactaccg cgacgacatcgaaaaagcaatcaaagcactgcaaaggatggacagcagtgaagccctcatccaagaagaattagaaaagagaggtatcgctca tgccttcaatacgtactccacccaacattcgcaaaatgacggcaacttccctcatgaaggcaacaaagaacacgtcacttcggctcctccacatcg cgaaggactcccaccctta [SEQ ID NO: 82]

[0216] Cargo

[0217] VLPs of the invention may be used to carry cargo.

[0218] VLPs of the invention are not limited to any particular type of cargo. For example, the cargo may be a protein, small molecule or nucleic acid. In some embodiments, the cargo is a nucleic acid, such as RNA or DNA. In some embodiments, the nucleic acid(s) are suitable for CRISPR, RNA interference (RNAi) or gene therapy, an antibody, a cytostatic or a cytokine. In some embodiments, the cargo is a nucleic acid such as a plasmid or a nucleic acid encoding an antibody. In some embodiments, the cargo is a protein or tag that could be used to target the VLP to specific cells, tissues, or tumours (e.g. through a protein-receptor interaction). In some embodiments the cargo is an imaging agent. In some embodiments the imaging agent is a radionuclide. In some embodiments, the cargo is a chemical group capable of binding an imaging agent, optionally wherein the chemical group is a multi-ion-chelating group (multichelator). In some embodiments the cargo is a biocatalyst.

[0219] In one embodiment, the cargo is a fused protein or peptide capable of binding a tissue-specific cellular receptor. In one embodiment the cargo is IL-2, such that the VLP will specifically bind cell expressing IL-2 receptor. In one embodiment, the cargo comprises a fused protein or peptide capable of binding a tissue-specific cellular receptor, and an imaging agent or a chemical group capable of binding an imaging agent.

[0220] In some embodiments, cargo is contained within the VLP. In some embodiments, the cargo is attached to the interior of the VLP. In some embodiments, cargo is attached to the exterior of the VLP. In some embodiments, cargo is attached to the exterior of the VLP and cargo is attached to the interior of the VLP.

[0221] In some embodiments, the capsid is linked to cargo. In some embodiments, cargo is attached to capsid protein(s) of the VLP. In some embodiments, cargo is attached to one or more spacers. Cargo may be attached via any suitable means and is dependent on the nature of the cargo. For example, wherein the cargo is a polypeptide, the cargo may be attached to the capsid protein and / or spacer via a peptide bond (e.g. by expressing the cargo, capsid protein, and optional spacer as a single polypeptide chain). Likewise, wherein the cargo is a small molecule, the cargo may be attached to the capsid protein and / or spacer via any suitable means, for example bio-orthogonal chemistry.

[0222] In some embodiments, the cargo is reversibly bound to the VLP. This can be achieved e.g. via physicochemical interaction with or attachment to any part of the VLP or by incorporation of the cargo into the VLP.

[0223] In some embodiments, the VLP incorporates the cargo. The incorporation can be complete or incomplete. In some embodiments, the major part of the total amount of the cargo is fully incorporated into the capsid. In some embodiments, the cargo is fully encapsulated in the VLP.

[0224] In some embodiments, the VLP attaches the cargo after the assembly of VLP. In some embodiments, the cargo encoded in the same sequence as the VLP.

[0225] In some embodiments, cargo is attached to the N- terminus of the capsid protein. In some embodiments, cargo is attached to the C- terminus of the capsid protein. In some embodiments, cargo is attached to the N- and C-termini of the capsid protein. In some embodiments, a cargo is attached to the capsid protein between the N- and C- terminus of the capsid protein.

[0226] The inventors believe that the N-terminal tail of the capsid protein is located on the outside of the VLP. Thus, attachment of cargo to the N-terminus of the capsid protein is understood to locate the cargo substantially at the exterior of the VLP.

[0227] The inventors believe that the C-terminal tail of the capsid protein is located on the inside of the VLP. Thus, attachment of cargo to the C-terminus of the capsid protein is understood to locate the spacer substantially at the interior of the VLP. Nucleic acids

[0228] In one aspect, the invention provides a recombinant, synthetic, or isolated nucleic acid sequence which encodes the capsid protein of any preceding claim.

[0229] As used herein, the term "nucleic acid sequence" embraces DNA (including cDNA) and RNA sequences. The nucleic acid sequences of the invention include nucleic acid sequences that have been removed from their naturally occurring environment, recombinant or cloned DNA isolates, and chemically synthesized analogues or analogues biologically synthesized by heterologous systems.

[0230] The nucleic acids of the invention may be prepared by any means known in the art. For example, large amounts of the nucleic acid may be produced by replication in a suitable host cell. The natural or synthetic DNA fragments will be incorporated into recombinant nucleic acid constructs, typically DNA constructs, capable of introduction into and replication in a prokaryotic or eukaryotic cell.

[0231] Usually the DNA constructs will be suitable for autonomous replication in a unicellular host, such as yeast or bacteria, but may also be intended for introduction to and integration within the genome of a cultured insect, mammalian, plant or other eukaryotic cell lines.

[0232] The nucleic acids of the present invention may also be produced by chemical synthesis, e.g. by the phosphoramidite method or the tri-ester method, and may be performed on commercial automated oligonucleotide synthesizers. A double-stranded fragment may be obtained from the single stranded product of chemical synthesis either by synthesizing the complementary strand and annealing the strand together under appropriate conditions or by adding the complementary strand using DNA polymerase with an appropriate primer sequence.

[0233] When applied to a nucleic acid sequence, the term "isolated" in the context of the present invention denotes that the polynucleotide sequence has been removed from its natural genetic milieu and is thus free of other extraneous or unwanted coding sequences (but may include naturally occurring 5' and 3' untranslated regions such as promoters and terminators) and is in a form suitable for use within genetically engineered protein production systems. Such isolated molecules are those that are separated from their natural environment.

[0234] In view of the degeneracy of the genetic code, considerable sequence variation is possible among the nucleic acids of the present invention. Degenerate codons encompassing all possible codons for a given amino acid are set forth below:

[0235] Amino Acid Codons Degenerate Codon

[0236] Cys TGC TGT TGY

[0237] Ser AGC AGT TCA TCC TCG TCT WSN

[0238] Thr ACA ACC ACG ACT ACN

[0239] Pro CCA CCC CCG CCT CCN

[0240] Ala GCA GCC GCG GCT GCN

[0241] Gly GGA GGC GGG GGT GGN

[0242] Asn AAC AAT AAY Asp GAC GAT GAY

[0243] Glu GAA GAG GAR

[0244] Gin CAA CAG CAR

[0245] His CAC CAT CAY

[0246] Arg AGA AGG CGA CGC CGG CGT MGN

[0247] Lys AAA AAG AAR

[0248] Met ATG ATG

[0249] He ATA ATC ATT ATH

[0250] Leu CTA CTC CTG CTT TTA TTG YTN

[0251] Vai GTA GTC GTG GTT GTN

[0252] Phe TTC TTT TTY

[0253] Tyr TAC TAT TAY

[0254] Trp TGG TGG

[0255] Ter TAA TAG TGA TRR

[0256] Asn / Asp RAY

[0257] Glu / Gin SAR

[0258] Any NNN

[0259] One of ordinary skill in the art will appreciate that flexibility exists when determining a degenerate codon, representative of all possible codons encoding each amino acid. For example, some nucleic acids encompassed by the degenerate sequence may encode variant amino acid sequences, but one of ordinary skill in the art can easily identify such variant sequences by reference to the amino acid sequences of the present invention.

[0260] One of ordinary skill in the art appreciates that different species exhibit "preferential codon usage". As used herein, the term "preferential codon usage" refers to codons that are most frequently used in cells of a certain species, thus favouring one or a few representatives of the possible codons encoding each amino acid. For example, the amino acid threonine (Thr) may be encoded by ACA, ACC, ACG, or ACT, but in mammalian host cells ACC is the most commonly used codon; in other species, different Thr codons may be preferential. Preferential codons for a particular host cell species can be introduced into the polynucleotides of the present invention by a variety of methods known in the art. Introduction of preferential codon sequences into recombinant DNA can, for example, enhance production of the protein by making protein translation more efficient within a particular cell type or species.

[0261] Thus, in one embodiment of the invention, the nucleic acid sequence is codon optimized for expression in a host cell. "Optimized" nucleic acid sequences encode an amino acid sequence using codons that are preferred in a recombinant cell. The optimized nucleic acid sequence is typically engineered to retain completely or as much as possible of the amino acid sequence originally encoded by the starting nucleic acid sequence, which is also known as the "parental" sequence. Several methods for codon optimization are known in the art. Preferably, codon optimized sequences avoid nucleotide repeats and restriction sites that are utilized in cloning the nucleic acids of the invention, by adjusting the settings in commercial software or by manually altering the sequences to substitute codons that introduce undesired sequences, for example with highly utilized codons in the heterologous organism of interest.

[0262] In some embodiments, the nucleic acid is codon optimized for expression in a prokaryotic cell. In some embodiments, the nucleic acid is codon optimized for expression in Acinetobacter, Agrobacterium, Escherichia, Cupriavidus, Clostridium, Rhodobacter, Marinobacter, Bacillus, Klebsiella, Tatumella, Pseudomonas, Ralstonia, Rhodococcus, Methylobacterium, Methylophilus, Methylococcus, Methylomicrobium, Methylomonas, Pantoea, Streptomyces, Parachlorella, Synechococcus, Synechocystis and Thermocynechococcus. In some embodiments, the nucleic acid is codon optimized for expression in E. coll.

[0263] In some embodiments, the nucleic acid is codon optimized for expression in a eukaryotic cell. In some embodiments, the nucleic acid is codon optimized for expression in a yeast cell, a fungal cell, an algal cell and a plant cell. In some embodiments, the nucleic acid is codon optimized for expression in Pichia, Saccharomyces, Kluyveromyces, Candida, Schizosaccharomyces, Schefferomyces, Rhodosporidium, Hansenula, Klockera, Schwanniomyces, Issatchenkia, Yarrowia or Rhodotorula. In some embodiments, the nucleic acid is codon optimized for expression in S. cerevisiae, C. lipolytica, R. glutinis, S. bulderi, S. barnetti, S. exiguus, S. uvarum, S. diastaticus, K. lactis, K. marxianus K. fragile, P. kudriavzevii, S. stipitis or / . orientalis.

[0264] Compositions

[0265] The invention also provides a composition comprising a capsid protein or VLP described herein. In some embodiments, the composition is a pharmaceutically acceptable composition.

[0266] In some embodiments, the composition comprises a pharmaceutically acceptable carrier. A pharmaceutically acceptable carrier includes any and all solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying agents, and the like, compatible with pharmaceutical administration. Preferred examples of such carriers or diluents include water, saline, Ringer's solutions and dextrose solution. Supplementary active compounds can also be incorporated into the compositions. Solutions and suspensions used for parenteral administration can include a sterile diluent, such as water for injection, saline solution, polyethylene glycols, glycerin, propylene glycol or other synthetic solvents; antibacterial agents such as benzyl alcohol or methyl parabens; antioxidants such as ascorbic acid or sodium bisulfite; buffers such as acetates, citrates or phosphates, and agents for the adjustment of tonicity such as sodium chloride or dextrose. The pH can be adjusted with acids or bases, such as hydrochloric acid or sodium hydroxide. The parenteral preparation can be enclosed in ampoules, disposable syringes or multiple dose vials made of glass or plastic.

[0267] In some embodiments, the composition is sterile. Vectors

[0268] In one aspect, the invention provides a vector comprising a nucleic acid described herein.

[0269] A vector of the invention comprises a nucleic acid of the invention. A vector may comprise one or more of an origin of replication, a promoter sequence operably linked to a nucleic acid of the invention and a reporter gene or selectable marker.

[0270] The promoter may be homologous or heterologous. The promoter may be constitutive or inducible.

[0271] In one embodiment, the promoter is inducible and is activated in the presence of an inducing agent. Inducing agents include, but are not limited to, sugars, metal salts, and antibiotics. In some cases, the promoter allows constitutive expression of the polypeptide(s) of the invention.

[0272] In one embodiment, promoters that are active at different stages of growth can be used.

[0273] The promoter sequence may be operable in a prokaryotic cell, for example an E. coli cell.

[0274] The promoter sequence may be operable in a eukaryotic cell, for example a yeast cell, a fungal cell, an algal cell or a plant cell.

[0275] Where the recombinant cell is a fungal cell, the promoter can be a fungal promoter (including, but not limited to, a filamentous fungal promoter), a promoter operable in plant cells, or a promoter operable in mammalian cells. Mammalian, mammalian viral, plant and plant viral promoters can drive particularly high expression when the associated 5' UTR sequence (i.e., the sequence which begins at the transcription start site and ends one nucleotide before the start codon), normally associated with the mammalian or mammalian viral promoter is replaced by a fungal 5' UTR sequence. The source of the 5' UTR can vary provided it is operable in the filamentous fungal cell. In various embodiments, the 5' UTR can be derived from a yeast gene or a filamentous fungal gene. The 5' UTR can be from the same species as the recombinant cell or from a different species.

[0276] Promoters for recombinant expression in yeast are known in the art. Suitable promoters for S. cerevisiae include, but are not limited to, the MFal promoter, galactose inducible promoters such as GALI, GAL7, and GAL10 promoters, glycolytic enzyme promoters including the TPI and PGK promoters, the TDH3 promoter, the TEF1 promoter, the TRP1 promoter, the CYCI promoter, the CUP1 promoter, the PHO5 promoter, the ADH1 promoter, and the HDP promoter. A suitable promoter in the genus Pichia sp. is the AOXI promoter.

[0277] Suitable reporter genes or selectable markers include, but are not limited to, a drug resistance gene, a metabolic enzyme, a factor required for survival of the recombinant cell, a fluorescent marker, or an enzyme that generates a detectable product. Cells transformed with the vector can be selected based on their ability to grow in the presence of inhibitors (e.g. antibiotics) or under conditions in which untransformed cells cannot grow.

[0278] The vector may be a high copy number vector, an intermediate copy number vector, or a low copy number vector.

[0279] Recombinant cells

[0280] In one aspect, the invention provides a recombinant cell transformed with a vector described herein. The invention provides recombinant cells engineered to express a polypeptide of the invention. The invention also provides recombinant cells transformed with a nucleic acid of the invention. The nucleic acid may be extrachromosomal, on a vector (typically a plasmid). In some embodiments, the recombinant cell is transformed with a vector of the invention. In some embodiments, the recombinant cell is transiently transformed. Alternatively, the recombinant cell may be stably transformed wherein the nucleic acid is integrated in one or more copies into the genome of the cell. Integration into the cell's genome may occur at random by non-homologous recombination but preferably, the nucleic acid construct may be integrated into the cell's genome by homologous recombination, as is well known in the art.

[0281] In some embodiments, the recombinant cell is a prokaryotic cell. In some embodiments, the prokaryotic cell is selected from Acinetobacter, Agrobacterium, Escherichia, Cupriavidus, Clostridium, Rhodobacter, Marinobacter, Bacillus, Klebsiella, Tatumella, Pseudomonas, Ralstonia, Rhodococcus, Methylobacterium, Methylophilus, Methylococcus, Methylomicrobium, Methylomonas, Pantoea, Streptomyces, Parachlorella, Synechococcus, Synechocystis and Thermocynechococcus. In some embodiments, the prokaryotic cell is E. coll.

[0282] In some embodiments, the recombinant cell is a eukaryotic cell. In some embodiments, the eukaryotic cell is selected from a yeast cell, a fungal cell, an algal cell and a plant cell. In some embodiments, the eukaryotic cell is selected from Pichia, Saccharomyces, Kluyveromyces, Candida, Schizosaccharomyces, Schefferomyces, Rhodosporidium, Hansenula, Klockera, Schwanniomyces, Issatchenkia, Yarrowia or Rhodotorula. In some embodiments, the eukaryotic cell is selected from S. cerevisiae, C. lipolytica, R. glutinis, S. bulderi, S. barnetti, S. exiguus, S. uvarum, S. diastaticus, K. lactis, K. marxianus K. fragile, P. kudriavzevii, S. stipitis or / . orientalis.

[0283] In some embodiments, the recombinant cell is a nematode cell. In some embodiments, the vector is expressed in a nematode.

[0284] Recombinant cells may be cultured in conventional nutrient media modified as appropriate for activating promoters (if an inducible promoter is present), selecting transformants, or amplifying the nucleic acid sequence encoding the polypeptide(s) of the invention. Culture conditions, such as temperature, pH and the like, are those previously used with the recombinant cell selected for expression, and will be apparent to those skilled in the art. Preferred culture conditions for a given recombinant cell may be found in the scientific literature and / or from the source of the recombinant cell such as the American Type Culture Collection (ATCC).

[0285] Methods of production

[0286] VLPs of the invention may be produced using any suitable method in the art. For example, the capsid protein may be expressed in a suitable recombinant cell (e.g. E. coll or a eukaryotic cell), followed by purification of the expressed protein and / or self-assembled VLP. In some embodiments, the method comprises a step of assembling the VLPs. Purification may be achieved by any suitable method, such as size-exclusion chromatography.

[0287] Methods of triggering VLP assembly are well-known in the art. For example, VLPs may be triggered to assemble by altering the salt concentration. In one embodiment, the VLPs are triggered to assemble at high salt concentration. Alternatively or in addition, VLPs may be triggered to assemble by altering protein concentration. In one embodiment, the VLPs are triggered to assemble at high protein concentration. In one embodiment, the VLPs are triggered to assemble by altering the pH.

[0288] In some embodiments, the method comprises co-expressing the capsid protein with an Envelope protein. The skilled person will appreciate that an Env protein comprises Gc and Gn. In some embodiments, the Gc of the Env protein is a Class II Gc.

[0289] In some embodiments, the method comprises co-expressing the capsid protein with a Class II Envelope protein.

[0290] In some embodiments, the method comprises fusing the capsid protein with a Class II Envelope protein.

[0291] In some embodiments, the method comprises co-expressing the capsid protein with a matrix protein or a fragment thereof.

[0292] In some embodiments, the method comprises fusing the capsid protein with a matrix protein or a fragment thereof.

[0293] SEQ ID NO: 83 corresponds to the amino acid sequence of the Env protein of EPB78661.1 Ancylostoma ceylanicum:

[0294] LNVNAITLCVSIFFAQAAFAFNTRCPNEISINKTILYATNCAKEGIAIARYSHQKEERMCWFPVSCPLGAIRSGYPYKNG STLCGKTCECPKWATSCSFSDSDRISFSELDMIPLSIRNYKPKQVCSFNASPKCDHKKQVGHFHQVQLFDQNLLLVEE LTISIKEYIDETDFLCIDRKGWKRRAQKSITGTSRFCEKHLCAPKAKLFCAYDNPLAMLVINDTESSTSIPIKAWGTITKFY FGFPTKLTEAKKDVDAINKSCSKGGISIQ.SNTTFDVAEVCIHHYCVFLKGITSQ.TVLFPNKLVrv1YDYVASIKIWNQ.GEL SYDSELSCKAHPICEILQCYLCWERFYNTQCWNYKHFTIFMLCGATILLAAPLLCIFFKFVRLTFTIIVIKIILKRLNPIRYLR LPRRQRSFRLPAYTNRRKKKQQKRFISCTILILIQLLQVEGCSEVISVTSTEEVCTIQENKETCTFNHATTITLQPLQQQT CLTLNDPEKRPMGMLTVKPDGIKFRCNKKIEFFTRDHQIVSESVHRCHRAGSCHSDECHHVKDTDALPEFSSEANSR PGYTSCSSSCGCITCDGCFFCEPSCLFHRLYAIPTTPTIYSIFYCPSWELEVDAEISLQREDETTTSTIRLLPGRTSTWNNI RFSLIGTIVPQLPILSSAFVTNGRQTSIVKPAYAGQLQSNSVGQLQCPNLEAAKQFECHFSRNLCTCTNALHKVSCTCY DGSVEDHMEALPLPQTSKNFLVFEKDRNIYAKTHVGSALQLHIVAQDLKITTVKHTSHCQVEASDLSGCYSCTSGASL TLSCKSDNGEVLANM KCNEQTHVIRCTESGFINNILLMFDTSEVAADCTAACPGGIVNFTIKGLLAFVNERIISQSYSA TDVERNIKRDFSFVNYLSEM ISLFVNKITSFFSFWKTVAFLIVIIVLLELITKLFTKPLHDKAH [SEQ ID NO: 83]

[0295] SEQ ID NO: 84 corresponds to the amino acid sequence of the Env protein of KAK6016282.1 Ostertagia ostertagi:

[0296] HQANALTTCVSIILFCQVALATNTRCPATISTTKTILYATNCAKEGIAIALHNDTDAEKFCWFPVSCPLGAIRSEFPIVIKP NTTLCGESCECPKWATSCSYSESDRTTLSELPLIPQSLRSYQPYQVCSFEKSRQCDQKRQIGNFHQIQLFDDNILLVKQ LTIQIKEYIDDRDFTCIDKKGWKRLPRRSITGTSRFCAKHKCTPNARLFCSYGNPLAMLVINDSESSTTLPIKAWGTVTK SFFGYRAIDAEEKQQRVSVHFTEQRCTVGGVFLRSDNVIEGAELCIHNYCVFLRNFTSQTVSFPNLLIMYDHLVSIKIW NNGKLYSESELTCKAHTICETLLCYICWERLYNPQCWNYKHYLMIVILFAAGLLLTIAPTVYILLRILRFIFRILWIIVILRKLL SIIRVSPNRILRFIFRILWIMLRKLLSIIRVSPNRYRQPRLPRYTNRRKRNRQRTFLACTILILFHSSTTKGCSETVAIVITTG EEVCTIKKDEETCTFNHATVITLQPLQQQTCLTLNDPQNLPMGIVILTIKPEGIKFRCNRKDEFYTRDHKIVSESVHRCH RAGSCHADQCHEIKETDKLKEFSEVANDSPGYTSCSSSCGGLSCDGCFFSTPSCLFHRLYATPTSSTIYTVFQCPSWEL EVDAEVTLQQEDGTTAATTIHLLPGRTNTWNNIHFSLIGTIVPQLPILSSTFVTNGERTSIINPAYAGQLQSHSVGQLQ CSSYHAAEQFNCYFSRNTCTCTTTVYSTTCTCSNGNIRDRMAAQPLPQISKNFIIFESNKKFYARTNVGSAIKVQIVAE NLRITSVKHISRCQVESSDLSGCYSCTTGASLTLSCTSDNDEVLAQVKCDQHSHVIRCTEAGFLNTIFLMFDSPTVTAY CTAACPGGTVNFTIKGSLAFVNERIISPDNGTSNIQRNVASDISFVNELVEKTKDKFMSVVNGITSFFTLWELLLFLIPVI LLVSLLHRFCPRIHHDKSH [SEQ ID NO: 84] SEQ ID NO: 85 corresponds to the amino acid sequence of the Env protein of VDL73942.1 Nippostrongylus brasiliensis:

[0297] LTIAVLLCQPIAAVSAMISNSRCPSTISVARKIIYADQCTEKGLAIASIIENNRRHLCWFPIRCPSGHINIPIPLTPNTGYCG

[0298] NKCNCPKWTQTCSFYNGSKKKMSQASNVPQAIFDYKPPQVCSFHKSATCSDIHSVGIFYQIELFDGTTIIVPELHISTR EFFDENDYYCFTNDGKLVPDSVSTNTPHYGSPAFCRHHKCSAPSTNSVFCTYYSPITTLDLFNSSIIIRAWGPTARKYFS YKPIAENSEKSYLIPRCYKGGVKVETTKTLDILEACSSSTCIYMTNYDNDIPILLPTSIVLFQYTVNLNGWKEGERVFSST

[0299] LTCPGQPICETIQCRICWEKLFNVQCWSVLEIAASLFLLIVIVILLHAITPLFTILSWIVRKMLHFPPAIARVVRNACRRM

[0300] TGRRQTYDVARTTYQPRKRRRKLDKRPAVQSPTTKKNAYTTKLPNYSYSLLVRKFAYY [SEQ ID NO: 85]

[0301] SEQ. ID NO: 86 corresponds to the amino acid sequence of the Env protein of KAK6031060.1 Ostertagia ostertagi:

[0302] LCNNMVLLCLMILTIQGTSSAHTRCPEEININKTILFATECTPKGIAIARYEHDNKNVFCWFPIICPFGEIRFDASKPTVS PICGKKCKCPDWTDSCSFSSSSRTTTSQIDLMPTSIKNYKPAYVCSFNFNSSCDPGKRMGVFNQIQLFDDNVLLVEKL TITIKDYIDKNDYVCIDWKGWKRRIARRSTGTSRFCQNHDCRDNAQVFCTYDNPLALLVIDDTSSIHEGAIPIKAWGS

[0303] ITKAYYEFRKDERIESSLNAVKRSIVTPAFGLTTPLISVQCIKGGVSIDNLFHNDITEICVSQYCVFARSISSKDTLFPNSLI

[0304] MYDYSVSIKTWNNGTITDESTVSCKAHPICETLRCTFCWEHIYNVQCWTNTQMIIIALSFLIIFLILPGIRFIGKIVLVLIA

[0305] PLLSILAMLNPSKIFHEGSRPHRTQLRTYTNRRRRIHNRRFLTCVISIIIHLHYSEGCSQVSSLTGNEAKCVTTGNAETCT FNEATLLTLQPLQQATCLTLRDKRDNAVGVISIKIKGINFRCHKNVEFFTRDHKIVSESVHRCYNAGSCVKDACDNTSP TDKIKEFSSTANSSPGYTFCTISCGCIFCEGCFFCQPSCLFYRLYAIPTSSTIYTVFSCPSWETIINIEVEILQEELKIANTLQL

[0306] LPGQTTSWNNLRFTLIGNIIPQLPILSSTFM ETRRNRDRETSS [SEQ ID NO: 86]

[0307] SEQ ID NO: 87 corresponds to the amino acid sequence of the Env protein of KAK6026028.1 Ostertagia ostertagi:

[0308] FLPNALLLCFVLTLHMVSSLNTRCPEEGANINKTVIYATNCVSKGIAIARYEQLNKYAICWFPVSCPMGHVRFDISAAQ

[0309] TTLCGDECKCPQWTNSCSFSRGSRTTYSELKNIPQHLRDYRPDYVCSFNLTSTCDTTKRIGIFNQIQLYDNSFLIVKDL NVRIKDYIDKNDFVCVDRKGWIRRPNRRVSGTSRFCEQHECHPNARLFCTYDNPLALLVIDESGQEDDYRSIPIKAW GTVMKPYHDYPRTPTAEPKEKITLQIRDATSTFSESEQTMTLSMKCIKGGLLLSTQEAFDVIEACVNDYCVYAKKLTKE

[0310] AIIFPNSLIMYDYTVSIKAYENDKLRHNGHVSCKAHPICETLRCTFCWKRIYNSQCWTLLETMFFITLPLLSVILIPWLC YITKIIGLLLHIIKFLVCGTVAICKLHRKSPNLRRYTSRRRKLRTRKSSFLPCVISVLSLLHLSKGCSQVVSMNAHEEICIISN DIETCTFNEAMVVTLQPLQQETCIALKDHESQPIGVISVKINGILFQCRRNVEFFTRDHQLVSESVHRCYSAGTCDRST

[0311] CENMSPTEKSKEFSFEANNNPGYTFCTPSCGCLTCDGCFLCEKSCLFYRVYAVPTTSTIYTIFTCPSWEIVVTLEATLRQ

[0312] KDSTVSTTIQLHPGQISAWNNLKFSLIGTVVPQLPILSSTFAETDYSISVIKPAHRGQLSPHSAGQLQVFNEGTRRHVN CSFAINACQCTHGLYKASCSCSSGSVADLMQPSPLPLVSKNFKIFSEMISERAELTLYCQSSGTETTANIECPSQTQMA LCTSSGYLNILKFHFDTSSVSMICNASCPGGTVSLAVKGVLLYVDDDLIRDNLQSEAKTRDLPRDTTFLSQIPNKFKEFV

[0313] GKNSRSASN [SEQ ID NO: 87]

[0314] SEQ ID NO: 88 corresponds to the amino acid sequence of the Env protein of KAK6028305.1 Ostertagia ostertagi

[0315] TRLQSASMFVTLLMCIQVVGAINTRCPDEMNTKKVILYATTCVSQGIAIARYNESNKEKFCWFPLTCPQGAIRFEGPG

[0316] KPNLALCGERCRCPSWSQSCSFTKNWRTSMSTTNSIPEHIQNYRPPYVCSFTKSSTCDSTKKIGVFNQIQLFDNSTFI

[0317] VETLTLSIQDYIDENDFICVDRKGWIRRKTRRITGTSRFCETHTCRNDAQLFCTYDSPIAIFVSNNAPSHNIPPITVKAW

[0318] GTITKTYYGYPSKNTTRWNESSVQASIHTQCSKGGLTITSNTSLQLVELCVTSYCVFLKEVTSQQVLFPNILIMYAYEVR

[0319] LKVWNYNNVIHEKHVKCEAHPICETLQCLLCWEWIYNTQCWTYRQRAYLIIFFILTVLITPPLCLLLKLLIRTSYLMIRSL

[0320] RYFIKCAWKRKKHKEANFRRYTIRRRHNSQKKLFTCTIM IVLHIQLSKECSQFTSITADEQICTITNDTRNCVFNQATIIT

[0321] LQPLQQEACLELQDYQSQHLGTIIITLLEIQYQCHKRIEFFSRDHEIASESSHRCFMTDSCSKNACDNIKTTDTLKEFSST ANNNPGFTYCTPTCGCLLCSGCFLCHPSCLFHRIYAKPTSSSIYTIFNCPTWDLTAIAKITLTQGTSTTSTQISLTPGHTIV WNDLRLTLIGSITPQLPILSSTFMETSDNIAIVKSVHKGQLIPHTAGQLQCATLQDAKDFRCLFASESCKCANGFLKVSC SCPDGNMKRLMEPSPLPQAGKNFLIVHNNGHIFAKLNIGSALQLHIVMENLKLIAQQHKSSCIFQTSDLTGCYSCIPG AIMDLACTSDEGEVTALITCEDQYQIAKCTPRTKLNKLVFHFSTSHVYTSCSASCPGGSANITIKGTLAYVDDNLISRGS SVASERRTASSDTSLYARIVTKFTTLLSDITHSVQSFFLSLVTVKNVLLLFVVLLILNLISCAYRTIFHTVILSKKIH [SEQ ID NO: 88]

[0322] SEQ. ID NO: 89 corresponds to the amino acid sequence of the Env protein of RCN35992.1 Ancylostoma caninum:

[0323] KHFLALALAFSTFNALLADTRCPAEINTPKTIVYATNCVSKGIAIAKLDEKKKLCWFPLSCPIGSIRIPLPFRQNQGMCG PECRCPPWATSCSFSSSPRQKNSKISNVPFTIASYRPEHVCSFSPSENCDKRRKIDKFNQVQLYDDTLLVVEQLHIIIKD YIDANDFLCFDVNGTQQAAITPYTGTSYFCNKNNCDDNAELFCTFGTPSAILATPEGEESLLVKAWGVTTKEYYGHLS PETSHPSISTICTKGGIHISVTSPTTAIEACINNYCQFISNPSDSNIVFPITLVAFNYKVEVSAWKDGKLIYQQGIECQASP VCELIECSLCLELIYNPNCWSRTQIVMTASFLFAAILTLCILHPILRLLSILTRISTLVFRRTIQKVFWKVIRKPKDFAVYTGN RRNLLLAIIFAITITPTQHCSEVISVTASEETCIQNKNNNTCTFNQASVVTLQPLNQEMCLLLKNNNNSYAGMITIRVK GLYFTCKRKVEFFTRDHHFYSESVHRCRFAGSCNTGTCDKIKPTDKLPELSNTANKHPGYTFCAASCGCLTCSGCFFCL TSCLFYRYYAYPESPTVYTVFNCPLWEFTVDVHITIQKNGHSEGAHLVLHPAKTTRWNNIRLSLIAAMIPQMPILSSTF VTDGNAIAITKPATSGQLIPNTIGQLQCLTYEDAKHFKCSFPSNACVCSAATNQAVCTCSHGSISEVFKKTPLPQLTKN VMIYQKDNEVIAKTKVGTAIQLHLVTEDLKMVSTRLNTTCSIVTSELSGCYDCLAGAQISIACKSDKGDVTAEVHCGN QLQIAICTPTGHINRLILNFNSSKIRESCTLSCPGGTMSFEVTGELSYFNDGVLHQQQGVANRFSDISLALPSFNRVFE

[0324] LLRKTLNNIASPIFLLKTSFILMIFIPISTLLLRLVDACQRARTH [SEQ ID NO: 89]

[0325] SEQ ID NO: 90 corresponds to the amino acid sequence of the Env protein of EPB75047.1 Ancylostoma ceylanicum:

[0326] LNLPLLISIMAILLPCAISTSSRCPEELTLDKKIIYATPCVAKGVAVAVYSQKNRKHICWFPVTCPNGEIRDSFHQSKDMR ICGPKCECPKWSQFCSHYHNSRTQHSKASNIPSEFLHFTPEEVCSFDQSAKCSQRKKIGTFNQIELYDGTLLLVPKLQV SVKDFLSSDDFICINSEGKLVPARRPYRGTSMFCEKQQCSDDGDKFCMYDTPVALYDFSNTRSSDILIPIKAWGTTTRE FYPHNSDADTRTVFILSKCSKGGVEIRSDREIDSIEACITSYCVFASHISTTSLFFPTELVVFKYTVKISAWQNGLRVHHSS MTCPSHPICEVLHCSLCVEKIYNTQCWTNLDILLIASTSFILLSFFWIFKPVIFLFKKTMKLLLRLLNFLFNFPRRNKRESK ETNSSTTQLRTYTRHPRHHRSPRHPAVFIAVTILVYGAQRGNCCSDVISVVSATKSCSTINSTEICTFNEGATLTLQPLG QDVCLSLRNQKDEHIGLISLRVKSIHHICHQKIEFFTRDHELNSESSHRCYRAGSCQPDKCDNTKSSDTIPEFSWRASN NPGFTFCTASCQCLVCDGCFFCTPSCLFYRLYARATSNTIYTIFTCPVWELTVTAEVTLQLHTGSTTHAVTLRPGRTTRV NNIRLSLLGTISPQLPILTSAFITDGTKTAMTTRVQANVLTPQTPAQLQCASKADAITFLCQFSSRACSCSTGPYKATCT CPEGKMSKYLQQNTLPLVSKNVIIEKYDDTIAARTQVGSAINVQVNMENVKVASIQNQGTCTITASTVEGCYSCLVG AKITVVCYSTEEQTTADITCNKQHQIATCTKRGKLTELTFHFNSPEVSTSCSISCPGGQSSFHLSGLLDYVNDAQLQREI

[0327] AVNSEMTIAQPSDTFTNVTHWISSFVNKFNLSNLKFLLVCCILALFSPCIVSMTSALRDNFR [SEQ ID NO: 90]

[0328] SEQ ID NO: 91 corresponds to the amino acid sequence of the Env protein of KAK6030023.1 Ostertagia ostertagi:

[0329] LLQNAPVIFMILAALVAVMDAANTRCPTELTINKRIIYATPCVTNGIAVATYDDHEKHHLCWFPVTCPNGEIRPMPSY PNDSRLCGMKCECPHWSQFCLHYNGEHTQRSQISSIDPTLLDFTPKELYDGSLLLVPSLHVITKDYLDDNDFICISSQG NRLPPTPPYKGTFLFCERHQCDFAVDKFCVYDSPITLLDFDNNVTSHNTVPIKAWGTTTKEFYPHISGTDAEIISLMP NCSKGGVRIIVDRQISVVEACIQSYCIFAFNITTTDLLFPTELVVFKYTIIINAWKNGLHLYNSSITCPSHPICEILHCNICIE KAYNTQCWTTFDTTLIILLATVTISLYWIFAPIIYVPQKTHGTATHSFQIYGTTIALQPLGQEVCLSLKNQRNEPIGLISLR MLSIRHVCQRKVEFFTRDHEITTESVHRCYRMGSCKRDKCDNTKSSDIVEEFSWKAANSPGFTFCNPSCGCLLTCGC LSCQPSCLFYRYYATPTSETIYTVFTCPSWELTVTAEITLQLQQERKTFTMLLRPGRATRTGDITLSLLSTISPQLPILASAF ITDGTRTAMTMKVQANELTPATPAQLQCASKLDASQFLCRFSSRACSCSAGSYKATCTCPKGHISPYLSRNTLPLITKN VIIEQHDDTIAAHTRVGSAIFLQLTMENVKIVSVQNKGQCSITSSSIEGCYSCLLGAKVSIICYSTEEQTTADITCEQEQQ VAICTKTGKLNEMIFHFRSPEIATKCTVSCPGGLSSFQVEGALDYVNDAQIQDEISTKNEERIDHNENALSTVTGWIYS

[0330] LPDTITGFFLDLNLFTSIKIFLTCIILIILLYCIVFVTSILSRAK [SEQ ID NO: 91]

[0331] SEQ. ID NO: 92 corresponds to the amino acid sequence of the Env protein of KAK6018921.1 Ostertagia ostertagi:

[0332] RNVLLSILGILQIFALSAVSADIISYTPPNTRCSSQQQTIRNFVYADQCTSKGIAVAVATTEDGRSSLCWIQLTCPLGHIRL PIPQTPNTGYCGNTCKCPSWTTLCSPYNGRLTKMSTFGNVPYEIAHFIPKQVCSFTPAEHCSKHKKIGVFYQIQLYDST LVLVPELHLKTAEFLHKDDYYCFNKDGKEYSHYPPMATSGTSAFCRRHSCVEADKATHFCTYSTPVTTYEFKNNSVIIH AWGTTTRQYYPYAPLPSNESELHFTIPRCTTGGISIDTTEKFNMIEVCSAHVCVYVSDYESTTTILLPTSIVIFDYTTTIQ AWKEGTRKFFHTLKCTGKPICEVIDCIMCWEHLYNYQCWTTTQIITIVFASLVTLIFCRLLSPLFKSLWWIITKIFRVLKK PIITIFNIIRCRRRKYTIPKFPSYAKRRTERYTRILPITLMAVRLCYTCSDVVSFTTTSHNCEFNGTIETCTVDHVVNLHLQ PIGQELCLLLRDSKEQPAGTISIRLHNVAYHCQQKTQYYTRDHSFHTESNHQCWNTKSCTGEPCKGTTTSSKIPELSYY ANRRPGFTYCSPSCKCFWCDWCFRCTETCIFYRHYAEPQSADIYRVFSCPSWPITVQGTITTNMM EQEKHQDFVLQ PATPYRWNQLQIALTATASPNFPILGNFFVSNGEKTAIVDPSPANQLLPYSIGQLQCDTLEDAKKFNTCRYSPTSCQCS TTVYEAHCECPKGDITSYFTNASRTLPLSSDNIIILNENNNIIAKTSAGTSISIQLSISNLQIYTRRSQNTCTVEAGSLHGC YSCYTGANIEISCTSSVTEEIATIKCENQNQIAVCKPGGHISKLTFHFQKSYINTNCSVKCPAGETFFDVHGQLQFVTEP RLQNDVSQGKHRRSEEKGIVSFTRDIGNSIIGKMATIWTWLKNWKLILILIVTSPIIVSLLRLAFRRKSQ [SEQ ID NO: 92]

[0333] SEQ ID NO: 93 corresponds to the amino acid sequence of the Env protein of EPB74949.1 Ancylostoma ceylanicum:

[0334] MTILAIVCIISFASADFAAHNPPDTKCDKYKLSPRHIIHSDQCAHEGIVIATTKPSSSAMERIFCWFHVSCPLGHIRVSLP FQPNNGYCGDICKCPEWTNTCSHYDGKLTTTSSLPNLPITISDYRPTQVCSFNPSSLCSSERRLGAFYQIQLYDDSLVIV PELHLKTAEFFSKDDFTCFNIQGHKLATTPNPSETTGTPAYCRRYPCISPQYANTFCAYSTPVTIYHHNNDTIIIKAWGT TTRTYYPPLESSQTPSLSSYIIPRCTVGGILIDTTEQFDMVETCSSFACVYISNYSSNSKILLPSSIVIFQYTVKLTAWREGK QVFQRTQTCEGKPICEILDCFICWEHIFNPHCWTTTEIVITIASSLLLIITCRLISPALLCAWWLLKKLIQTAKKCLVKMFT CGRRSGTHTLPSFPSYRRPSRRTRKVSSLILLVYITAIQPTNQCSHIVATSSTSQYCESNSTDEQCYFDQATTLQLQPIGQ DVSLLLRNSKNEPAGTISIRVQDIAYTCRQKTEYYTRDHKFYTESHHQCWNSLSCKGETCKGVTTSTNIAEFSSYAKKH PGFTYCSPSCKCLWCDWCFWCKETCLFYKVYAYPESKTTYRVFTCPAWSVTVAGTIVLETSKLREKHDFVLHPARPHT WNKIELTLTATTVPNLPILSSYFLTDGEVVSKVEQSPAGQLIPHSVGQLQCTSENSARTFSSCTFSATACTCTPRVYEATC NCPEGNIRRYIDNSQLQLPLAAKDVVILQRGADIQAKTTSGSSTSVQVSMRGLTILTKRLNNECQIDIGSLTGCFSCIQ GARLQASCTSSLAEETAEIRCGNQTQLAMCSPHGKISELVFHFTSSSINMNCSVICPAGQSTITIQGSLDYVNNALFQ APIQQVDTHRATSSGFSLDTMSHALSTILSPIANVVELVLEQLAKIQSPPTHLEMN [SEQ ID NO: 93]

[0335] SEQ ID NO: 94 corresponds to the amino acid sequence of the Env protein of KAK6009265.1 Ostertagia ostertagi:

[0336] SRARAQTSCILPFIMLQLLGSYSATTSDTRCPEEVNVNKTVIYATSCVSRGVAVARYNDSGEEKFCWFPLSCPYGAIRF ENRHSSTTLCGEPCKCPTWSTSCSFTTGWRTATSTFSSLPRRIRQYRPPQVCSFSNSPTCDSTKKIGVFNQVQLFDNR TFLVPSLAISIQEYIDENDFICVDHKGWIRRPARRSTGTPHFCAIVAKTHKCHENPQTFCVIGPNLAVLVVDNWNSSH TVPIIVKAWGTVSKIYYGYPSNHIKSRRFIPLTSPVESQCSKGGITVASKIQLQVLEACINNYCVFLKNVSTEVILFPNSLI M YDYTVRLKALM N N AIVH EAQVGCKAH PICETLRCTLCWEWIYNTQCWTLLQVI LVFTLFVAAALTI PVI LI LCKAI AA TICLLIRLLLYLVHHMPWIRRKTQFRRYVSQRRRKKPKRIVSCTLLILLQIQSSIECSQFTSLTASEHVCTISKNSTDCTLN QATMIALQPLQQETCLQIRDHKNEHLGILSVTLRDIQYRCHKRVEFFSRDHALISESHHRCYLAGSCTQKFCDNVKP MDALEEISSIANKSPGYTYCTPTCGCLTCGGCLLCEDSCIFYRVYAKPISSSIYTIFNCPSWELSVTADITLSYGNNMTSK

[0337] RLKLTPGNVNSWNKLRLTLISTITPKLPILGSTFMSDGNDIAIIKPAHRGQLVAHTAGQLQCATFQDAQKFRCLFASET CKCSKGFHKMSCLCPDGNLKRLMQPSPLPLVGKNFLLTQQAGTVFAKLNIGSSLQLQIVMKDLHLVSLHQNSKCTFI ASDLTGCYSCISGATLDLVCTSSDGETTALLSCPGETQIAKCTTRGHLNRLILHFDSENILSSCIASCPGGSTNITIKGTLSY VNDDLIKPGLYAADQKPPLTTLNSLGKVFDSVIQFASYVVESVESLFLSLKTWKMIFLASVVLCFLVILSRMFHFMFPTI

[0338] PPRPKKH [SEQ. ID NO: 94]

[0339] The skilled person can readily identify nucleic acids which encode Class II Env proteins described herein. Exemplary nucleic acids are provided by SEQ. ID NOs: 95-106.

[0340] SEQ. ID NO: 83 may be encoded by the nucleic acid sequence of SEQ ID NO: 95: ctgaacgtcaacgccataacgctatgcgtaagcatcttcttcgcacaagccgcgttcgctttcaacactcgctgcccaaatgaaatttccatcaaca agaccatcctatacgcaaccaactgcgcaaaggaagggatagccattgctcgatacagccatcaaaaggaagaacgaatgtgctggttcccagt ctcttgtcctttaggagcaatacgatctgggtacccatacaaaaacggatcgacactatgtggaaaaacgtgcgaatgcccgaaatgggccacct catgttccttcagcgacagcgatagaataagcttctctgaattggacatgattcctctcagcataagaaactacaagccgaaacaagtatgctcctt caatgcctctccaaagtgcgatcacaaaaagcaagtaggacacttccatcaagtccaactcttcgatcagaatctactgttagtagaagaactaac tatcagcatcaaggaatatattgacgagacggacttcctatgcatcgatagaaaaggatggaaaaggcgagcacaaaaatcaataactggaacg tcacgcttctgcgagaagcacttgtgcgcaccaaaggcgaaattgttttgcgcatatgacaaccccttggccatgttagtcatcaacgatactgaatc atcaacgtcaataccgatcaaagcatggggaactataacgaaattttacttcggattccctaccaaattaacggaagcgaagaaggatgttgatgc catcaataaatcgtgctcaaaaggaggaatatcaatccaaagcaacacaacatttgacgtcgcagaagtatgcatccaccattactgcgtttttcta aaaggaataacctcacaaactgtgttgttcccgaataaacttgtcatgtacgactacgttgcatcaatcaaaatctggaaccaaggagaactctctt acgactcagaactctcctgcaaggcccacccaatctgcgagattctacagtgttatctctgctgggaaagattctacaacactcaatgctggaacta caaacacttcacgatcttcatgctttgcggcgcaactatcctcctcgcagcaccattgctatgcatcttcttcaaattcgttcgcctcaccttcaccata atgaagatcattctgaagagactcaatccgattcgctacctacgcttgccgcgaagacaacgctcatttcgattacctgcatacaccaataggagaa agaagaaacaacaaaaaagattcatatcgtgcacaatactcattctcatccaactacttcaagttgaaggatgttcagaagtcatctcagtcacatc gacagaagaagtctgcacgatccaggaaaacaaagaaacctgcaccttcaatcatgcaaccactatcacgctacagccactacagcaacagac atgcctcactctcaatgaccctgaaaaacgaccaatgggaatgttaactgtgaagcctgacggaatcaagttcagatgcaacaagaaaatcgagt tttttacgcgtgatcaccaaatcgtctccgagtctgtccatcgatgtcacagagccggaagctgccacagtgacgaatgtcatcacgtcaaagacac agatgccctcccagaattttccagcgaagctaacagcaggccaggatatacgtcatgttcgtcaagttgtggttgcatcacatgcgacggatgttttt tctgtgaacctagctgcctatttcacagattgtacgcgattccaacgacgccaacgatttactccatattttactgccccagttgggaacttgaagtcg acgccgaaatttccctacaacgagaagatgaaacaaccacatccactattcgactgcttcccggtcgaacaagcacctggaacaacattcgcttct cgctgatcggaacaatagttccacaactccccattctctcatccgcgttcgtcacaaatggaagacaaacatcaatagtgaagccagcatacgcag gacagcttcagagcaactccgttggacaactgcagtgtcccaatcttgaagccgcaaaacaattcgaatgtcatttttcgcgtaacctatgcacttgc accaatgcactacacaaagtcagctgcacatgctacgatggaagtgtagaggaccatatggaagccctcccgcttccacagacatccaaaaactt ccttgtcttcgaaaaggatcgcaacatttatgccaaaactcacgtcggatctgctctacaactccacatagtagcacaggacctgaaaatcactacc gtgaagcacacaagtcattgccaagtggaagcaagcgatttgtctggctgctacagctgcactagtggagcgtcgttaacattatcgtgcaagagc gacaacggtgaagtcctcgctaacatgaagtgcaatgaacaaactcacgtaatcagatgtacagaaagcggattcatcaacaatatattactcat gttcgacacgtctgaggttgccgccgactgcactgccgcttgcccgggagggatcgtgaacttcacaatcaaaggactgctcgcgtttgtcaacga aaggatcatctcacaaagctacagtgctacggacgttgaaaggaacatcaagagagacttttcctttgtgaattacctgtcggaaatgatatccctg tttgtaaacaaaatcacgtcatttttttctttctggaagactgtagcgtttctcattgttataatagtactattggagttaatcactaagttgttcaccaaa ccgcttcatgacaaagctcat [SEQ ID NO: 95]

[0341] SEQ ID NO: 84 may be encoded by the nucleic acid sequence of SEQ ID NO: 96: catcaagccaatgcgctcactacatgcgtcagcatcatcctattctgccaagttgcccttgctacaaacacacggtgccccgcaactatatctacca ccaagacgatactctacgcaaccaactgtgcaaaggaaggcatagccatcgcactacacaacgatactgatgcagaaaaattttgttggtttcctg tatcctgtccacttggagccataagatcagaatttccaatgaaaccaaacacaactctttgtggtgaatcctgtgaatgcccaaagtgggccacttc ctgctcatatagtgaatcagacaggactaccttgtccgaattgcccttgattccccaaagcctaaggagctaccaaccatatcaagtatgctctttcg aaaagtctcgacaatgcgatcaaaaacggcaaataggaaactttcatcaaatccaactattcgacgacaacattcttctggtaaagcaattgaca attcaaatcaaagaatatatagacgacagagacttcacgtgtatcgacaaaaaaggatggaaacgactgcccaggcgatcgatcactggaacttc acgcttctgtgcgaaacacaagtgcacgccgaacgctagactgttctgctcatatggcaatccgttggcgatgctcgtcatcaacgacagtgaatca tcaacaaccttaccaatcaaagcatggggaacggtgacaaagtcgtttttcggataccgagcaatcgacgcagaagaaaagcaacaacgcgtca gtgtccatttcactgaacaaaggtgcacagttggaggagtcttcttaagaagtgacaacgtaatcgaaggcgcagagttatgcatccacaattactg cgtttttctgaggaattttacttcccagacagtatccttcccaaacttgctgataatgtacgaccacctcgtatccatcaagatatggaataatggcaa gctgtacagcgaatccgaacttacctgcaaggctcataccatctgtgagacccttctatgctacatttgttgggaaagactctacaaccctcaatgct ggaattacaagcactatctcatgatgctcttcgctgccggtctgctccttaccatagcaccaaccgtgtacattcttctccgaatactacgcttcatctt ccgcatactatggatcatgctacgaaagctcttgtccattattcgcgtttcacctaaccgaatactacgcttcatcttccgcatactatggatcatgcta cgaaagctcttgtccattattcgcgtttcacctaaccgctatcgccaacctcgcttgccgagatatactaatcgacgaaaaaggaatcgacaacgaa catttctggcctgcaccatcctcattctctttcactcaagtacgacaaaaggatgctcagaaacagtcgccatgacaacaggcgaagaagtctgca caataaaaaaggacgaggaaacttgcaccttcaaccatgcgacggtaataacgctccaacctttgcagcagcaaacttgcctcactctcaacgat ccacaaaatcttccgatgggaatgcttacaatcaaaccagaaggaatcaagtttcgatgcaacaggaaagatgaattttacactcgtgaccacaa aattgtgtcagaatcagtccaccgctgtcaccgcgccggtagctgtcacgcagaccaatgccacgaaatcaaggaaactgacaaattgaaggaat tttccgaagtcgccaacgatagtccaggatacacatcttgttcttccagctgtggtggtttgtcttgcgatggctgtttcttctctacacctagctgcctc ttccatcggctctacgctactccaacatcatcgaccatttataccgtcttccaatgcccaagttgggaacttgaagtagacgcagaagttactctaca gcaagaagatggaacaactgctgctactacaatacaccttctcccaggccgcacgaacacttggaacaatatccacttctcgttgatcggaacaat tgtcccacaacttcctatactgtcatctaccttcgtcaccaatggagagcgaacctccataatcaatccagcttacgcgggacaattacagagtcatt ctgtaggacaactgcagtgctcttcatatcacgcagcggaacaattcaactgctacttctcgcgtaatacatgcacttgcaccactactgtatatagt acgacctgtacctgcagcaatggaaatatacgtgatcgcatggctgctcaaccacttccacaaatatcgaaaaacttcatcatcttcgaatccaaca agaagttctacgcaaggacaaatgtcggctctgccataaaagttcaaatagtcgcagaaaatctccgaataacgtcagtgaaacacatcagccgt tgtcaagtggaatcaagcgacctatccggatgctacagttgcactacgggcgcctcattaacactatcgtgcaccagcgacaacgacgaagtgctt gctcaagtcaaatgtgatcaacactcacacgtcatcagatgtacagaggctggcttcctgaacacaattttcctgatgttcgactctccaacagtgac cgcatattgcaccgctgcctgccctggagggaccgtgaacttcacgatcaaaggatctctggctttcgtcaacgaaaggattatctctccggacaac ggcacatcaaacatccagagaaacgtagcaagtgatatatctttcgtgaatgagcttgttgaaaagactaaagacaaattcatgtcagttgtgaac ggtatcacctcgttcttcactctttgggaacttcttctctttcttatccctgttatactactggttagccttctccatcgattttgtcctagaatacaccatg acaaatcgcac [SEQ ID NO: 96]

[0342] SEQ. ID NO: 85 may be encoded by the nucleic acid sequence of SEQ ID NO: 97: ctcacaattgctgtactactttgtcaaccaattgctgcagtttcggccatgatctcaaacagtagatgccccagcacaatatcagtcgctagaaaaat catctacgctgaccagtgcacagaaaagggactcgcaatcgcctcaataatcgaaaacaatcgacgacacttatgttggtttccgatacgttgtcc gtcaggacacatcaatatacccattcctctgacaccaaatacaggatactgtggaaacaagtgtaactgcccaaaatggactcaaacgtgctcatt ctacaatggaagcaagaagaaaatgtcgcaagccagtaatgttcctcaagcgatcttcgactacaaaccaccgcaagtatgctcattccacaagt cagccacatgttccgacatccacagcgtaggaatattttaccaaatcgaattattcgatggaacaactataattgtacccgaactgcacatctcaac aagggaatttttcgacgagaacgactactactgcttcacgaacgacggaaagttggtacccgattcagtttcaacgaatactccacactatggttcc cccgcattttgccgccatcacaaatgctcagcaccatcgacaaattcagtcttttgtacttactattcaccaatcacgacactggaccttttcaactcat ccatcatcatacgtgcctggggtcccacagcaagaaaatatttctcctacaagccaatcgccgaaaattcagaaaagagctacctcatacctcgat gctataagggaggtgtaaaagttgaaacaacaaaaactctcgacatacttgaggcttgcagctcttctacctgtatctacatgacgaactatgaca acgacatcccaattctattaccgacctctattgtactgttccagtataccgtgaacctgaatggatggaaagaaggcgaaagagttttctcatcaact ttgacgtgcccgggacaaccgatctgcgaaacaattcaatgcaggatttgctgggaaaagctattcaacgtacaatgttggtcagtcctcgaaata gcagcatctttgtttctgcttatcgtcattgtcatccttctccacgcaattactccacttttcacaatactcagctggatcgtaagaaaaatgttgcatttc cctccagcaatagcgagagtagtcagaaatgcatgcagaagaatgacaggcagacgtcagacctatgacgtagctcggaccacataccaacca aggaagagaagaagaaaactggacaaacgtcccgctgtacaatctccgacaacgaagaaaaatgcatatacgacgaaactacccaactactcc tacagcctattggtcaggaagtttgcttactac [SEQ ID NO: 97]

[0343] SEQ ID NO: 86 may be encoded by the nucleic acid sequence of SEQ ID NO: 98: ctgtgtaacaatatggtccttctgtgcctaatgattctgacaatacaagggacatcatcagctcatactcgctgccctgaagaaatcaacataaata agaccatactttttgctaccgaatgcactccaaaaggaatcgcgatagcacggtacgagcacgacaacaaaaacgttttctgttggtttccgatcat ttgcccatttggcgagatacgattcgatgcgtcaaaaccaacggtttcaccgatatgcggcaagaagtgcaaatgtccagattggacagattcatgc tccttttcgagcagctcaaggacaacaacatctcaaatagacctcatgccaacttctatcaagaattacaaacccgcatatgtgtgctcattcaactt caattcttcatgcgatccagggaaaaggatgggagttttcaatcaaatacaacttttcgatgacaacgtactactcgtggaaaaactcaccatcacg ataaaggactacatcgacaaaaacgactatgtgtgcatagattggaaaggatggaagaggcgcatagcaagacggtcgaccggcacgtctcgtt tctgccaaaaccacgattgtcgagacaacgcacaagtgttctgcacgtacgacaaccctctcgcgctcttagtcatagatgacacatcctctataca tgaaggagccataccaataaaggcgtggggatcgatcaccaaagcgtactatgagttcaggaaggatgagcggattgaaagttcgttgaacgca gtaaaaaggagtattgtaacgccagcctttggcctgaccactcctctcatctcagtgcagtgtatcaaaggaggagtctcaattgacaaccttttcca caacgatattaccgagatatgtgttagtcaatactgcgtttttgctcgcagcatctcttccaaggacacactatttccgaactcactcataatgtatgat tactcagtctccatcaagacatggaacaacggtacgatcacagatgagtctacagtttcctgcaaagctcatccgatatgcgaaacgttacgttgca ccttctgctgggaacacatatacaacgtacagtgctggaccaacacccagatgataataattgcactttccttcctaatcattttcctcattctaccgg gaattcgcttcatcggcaagatagtcttggtgctgattgctccacttctctccatactagcaatgttgaatccctccaaaatattccacgaaggatcgc ggcctcaccggactcaacttcgcacatatactaatcgcaggagacggatccacaaccgcagatttctcacgtgcgtcatctcaatcataatccacct tcactacagcgaaggatgttcacaagtttcttcgctaactggaaatgaagcaaaatgtgtcactaccggaaacgccgagacatgtacgttcaacga agctaccctcctcactctccaacctcttcaacaggcgacgtgcttgacactgagggataagagggacaatgctgtgggcgtgatctccatcaaaat caaaggaatcaacttcagatgccacaaaaacgtggaatttttcacccgcgaccacaaaatagtctcggaatccgttcaccgctgttacaatgctgg tagttgtgtcaaagacgcgtgtgataatacctcaccaacggacaagatcaaggaattttcgagtactgcaaacagtagcccaggctacacattctg tactatcagttgcggatgcatcttctgcgagggctgcttcttttgtcaaccgagctgcctgttttatcgactttacgcaataccaacaagttcgactatc tatacggtcttcagttgcccgagctgggaaaccatcatcaacatcgaagtcgaaatccttcaagaagaactcaaaatagccaacactctgcaactc cttcctggtcaaacaacatcatggaacaatcttcgcttcactctcataggcaacattattcctcaactgccgattctatcatcaacgtttatggaaaca agacggaatcgcgatcgtgaaaccagctca [SEQ ID NO: 98]

[0344] SEQ. ID NO: 87 may be encoded by the nucleic acid sequence of SEQ ID NO: 99: ttcttaccaaatgcacttctactgtgttttgttctcacactccacatggtttcatcgctgaatacccgctgcccagaagaaggcgcaaacatcaacaa gacggtcatatacgcaacgaactgcgtttcaaagggcatcgcaatcgcaaggtacgaacaactgaataagtatgccatatgctggtttcccgtttcc tgccctatgggacacgtacgtttcgatatttcggcagcacaaacaaccttatgtggtgacgagtgtaaatgtccgcaatggacaaactcctgttcctt cagccgcggatcaagaacaacctactcggaactgaaaaatattccacaacatcttcgcgattatcgtcccgactacgtctgttcgttcaacttgacgt ctacatgtgacaccacgaaaagaatcgggattttcaaccaaatacagctttacgacaacagcttcctcatagtcaaagatctaaacgtaagaatca aggattatatagataagaacgatttcgtatgcgtagacagaaaaggatggataagaaggccaaacagacgtgtctcaggaacgtcgcgtttttgtg agcaacacgaatgccatccaaatgctcgccttttttgcacctacgacaaccccctcgcccttctagtcatcgatgaatcaggtcaagaggacgatta tcgatccataccaatcaaagcatggggtacagtaatgaaaccgtaccatgattacccaagaacgccaacagcggaaccaaaggaaaagatcac tctacaaattcgtgatgcaacctcgacgttttcagaaagtgaacaaacaatgacgctatcaatgaaatgcatcaaaggtggtctcttgctatcgact caagaagctttcgatgtaattgaagcttgcgttaacgactactgcgtgtatgctaagaagctaacgaaggaagccataatcttcccaaattctctcat catgtacgattataccgtttcaatcaaagcctacgaaaacgataaactcaggcacaacggacatgtatcctgcaaagctcacccaatttgtgagac actcagatgtactttctgttggaaacgaatctacaactcacaatgttggactctcctggaaaccatgttcttcataactttgccactactctccgtcata ctgataccatggctttgctacatcaccaagatcattggacttctgctccacatcataaaattcctagtgtgcggaactgtcgcaatttgcaaactcca ccgaaaatcgccgaacctacgaagatacacctctcgtcgacgaaaactcaggactcgcaagtccagttttttgccttgcgttatttcagtcctttcact tcttcacctatccaaaggatgctcacaagtcgtctctatgaatgctcacgaagagatctgtatcatatcgaatgatatcgaaacttgcactttcaacg aagcaatggtcgtcacgctacagcccctgcaacaggaaacctgcatcgctctgaaagatcatgaaagccaacctatcggcgtaatctcagtgaaa atcaacggtattctgttccagtgcagaagaaatgtggaatttttcacacgagatcatcaacttgtatccgaatctgtgcacagatgctactcggcagg aacatgcgatagaagtacgtgtgagaatatgtcaccaaccgagaaaagcaaagaattttcattcgaggcgaataataaccctggatacactttttg cacaccaagctgcggatgtctaacatgcgacggctgctttctgtgcgaaaaaagttgcctcttttatagagtatatgcagttccgaccacttcgacga tctatacaatatttacatgcccaagctgggaaattgttgtcacattggaagcaacgctccggcaaaaagattcgaccgtatcaactacaattcaact tcacccaggtcaaatttccgcttggaacaatctcaagttcagtctcatagggacagtcgtcccccaactaccaatattgtcgtcgacattcgcggaa accgactactccatctcagtgatcaagcctgctcatcgaggacaactttcacctcacagtgcgggacagcttcaagtgttcaacgaaggaacacgc cgacacgttaactgctcgttcgcaatcaacgcctgtcagtgcactcatggactatataaagcctcctgttcatgtagctccggcagcgtagcagactt gatgcaaccgtcgcctttgccactggtttcgaaaaactttaagatcttttcagaaatgataagtgaaagagcggaactaactctctactgtcaatcga gcggaactgaaacaaccgctaatatcgagtgcccttctcaaacacaaatggcgctttgcacttcatcgggatacctcaacattctcaaattccactt cgacacttcatctgtctccatgatttgcaatgcttcctgtcccggaggaactgtttccttggcagtgaagggagtacttctctacgttgacgacgacct catcagagacaacttgcaatcggaggcaaagacaagagatcttccaagagataccacatttttgagtcaaatacccaacaagttcaaagaatttgt cggtaagaattccagatctgcttccaat [SEQ ID NO: 99]

[0345] SEQ ID NO: 88 may be encoded by the nucleic acid sequence of SEQ ID NO: 100: acaagacttcagtcggcttcaatgtttgtcacattactaatgtgtattcaagtcgtcggagcaatcaacacccgctgccccgatgagatgaacacca aaaaagtgatactctacgccacgacatgtgtatcccaaggaatcgcgatcgccagatataacgaatccaataaagagaaattctgctggttcccac ttacctgtcctcaaggagctataagattcgaaggtccaggaaaaccaaatcttgcactctgtggagaacgatgcagatgcccatcatggtctcaat cgtgctcttttactaagaattggagaacctcgatgtcgacaaccaactcaataccggagcacattcaaaattatcggccaccatatgtgtgctcattt acaaagtcatcaacttgcgactctaccaaaaaaatcggagttttcaatcaaatccagttattcgacaacagcacatttattgtagagacactcacctt gagcatccaagattacatagatgagaacgatttcatttgtgtggaccgcaaaggatggataaggcgcaagacccgaagaataacaggaacatca cgtttttgcgaaacgcacacctgtcgaaatgatgcacaactcttttgcacgtatgacagcccaatcgccattttcgtcagcaataacgcaccgtcaca taacataccgcccataacagtgaaggcctggggaaccattactaagacctactatggatacccatcaaagaacacaacgagatggaacgagtct tcagtccaagcatcaatacacacacaatgctccaaaggtggtttaacaattacaagcaatacatcacttcaactagtggaattgtgcgtcacgagc tactgcgtgtttctgaaagaagtcacatcacagcaggttcttttccccaacattcttatcatgtatgcttatgaagttcgactgaaagtctggaactac aacaacgtcatacacgaaaaacacgtcaaatgtgaagcacacccaatatgtgagaccctacaatgtttactgtgttgggaatggatatacaacac gcagtgctggacgtaccgtcaaagggcgtacctcataatattcttcatactcaccgtgctcataactccacctttgtgtctacttctaaagctacttatc agaacttcgtacctgatgatacgttcgcttcgatacttcatcaaatgtgcatggaaaaggaaaaagcacaaggaggccaacttcagaagatatac aattcgacgacgacacaattctcaaaagaaactcttcacatgcacgatcatgattgtccttcacatccagcttagcaaggaatgttcccagtttacgt caataacagcggacgaacaaatttgcaccataactaacgataccaggaactgcgttttcaatcaagcaacaataatcacgctgcaaccactacaa caagaagcatgtctggaacttcaagattaccaaagccaacatttgggtaccataatcattaccctactcgagatacaataccaatgtcataaacgc attgaatttttctcacgagatcatgaaatcgcatcagaatcgagccatcgctgcttcatgacggatagctgctcaaaaaatgcctgcgacaatatca agacgacagatacgttgaaagagttttcttctactgcgaacaacaacccaggtttcacatattgtacccccacatgtggttgcttactctgcagcggc tgttttctatgccatccaagctgtttatttcatagaatctacgccaaaccgacttcttcatcaatatacaccattttcaattgtcccacttgggacctaac agccatcgcgaagatcaccttgactcaaggaaccagtacaacttcaacacagatatcactcactccaggacacacaattgtatggaatgatctacg gctcacactcatcggatctatcacgcctcaactaccaatacttagttcaacattcatggaaaccagtgacaacatcgcaatcgtcaaatctgttcaca agggacaacttatacctcacacagccggtcagctccaatgcgccacattacaagatgccaaagattttcgatgtttgttcgcaagcgaatcttgcaa atgcgcgaacggattcctcaaagtctcatgtagttgcccggatggaaacatgaaaagattaatggaaccttcacctctaccacaggcaggaaaaa acttccttattgttcacaataacggacacattttcgcaaaactcaacatcggctccgcccttcaactacacattgtaatggaaaatctgaaactcatt gcgcaacagcacaagagctcctgtatatttcaaacgagcgatttaaccgggtgctacagttgcatcccaggagcgataatggatctcgcctgtacg agcgatgaaggagaggtgaccgccctgatcacttgtgaggaccaatatcaaattgccaaatgcactccacgcacaaaactcaacaaactcgttttt cacttttctaccagtcacgtatacacttcctgctctgcttcgtgcccgggagggtctgccaacataaccatcaaaggaactctggcttacgtggacga taacctcatcagccgcggatcgagtgttgccagcgaaagaagaactgcttcaagtgatacgtcattgtatgcaagaatagtaacaaaattcaccac tctcttatcagacatcactcattctgtgcaatccttttttctgagtcttgtgactgtgaaaaatgttctcttattattcgtagtcttgcttattctgaatcttat ctcttgtgcctaccgtactatttttcacacggtgatcttatcgaaaaagatccac [SEQ ID NO: 100]

[0346] SEQ. ID NO: 89 may be encoded by the nucleic acid sequence of SEQ ID NO: 101: aaacacttcttggcacttgcgctagcattttcaacattcaatgctttgctcgcagataccagatgcccagcagaaatcaacacaccaaaaacaatag tttacgcaacaaattgcgtatctaagggcatcgctatcgccaagctagatgagaagaagaagctttgctggtttccactatcatgtcccattggatcc attcgaatccccttaccatttcgacagaaccaaggaatgtgtggccccgaatgcagatgtccaccatgggcaacatcctgttcattcagttcaagtcc tcgtcagaagaattcaaaaatctccaatgttccctttacaatagcatcatatcgaccggaacatgtatgctccttctctccatccgaaaactgcgaca aaaggaggaaaattgacaaattcaatcaggtacaactgtacgatgacacgctcctggttgttgaacaactccatatcatcatcaaagactacatag atgcaaatgacttcctatgtttcgacgtaaatggaactcaacaagcagcaataacaccatacacgggaacttcgtacttctgcaacaaaaacaatt gcgatgacaatgcggaactcttttgcacttttgggacgccgagtgcaattctcgctacaccagaaggagaagaatcactcctagtaaaagcatggg gagtaacaactaaggaatactatggacatctctccccagaaacaagccacccatcaatttcaacgatatgtacaaaaggagggattcatatatcc gtaacctcacctacaacagcgatagaagcatgcatcaataactactgtcaattcatctccaatccatcggactcaaatattgtctttccaataacact tgtagcattcaactacaaggtcgaagtctctgcctggaaggatggaaaactgatataccaacaaggaattgaatgtcaagcctctcccgtatgcga actcatcgaatgttcgctctgtcttgaactcatctacaacccaaactgctggtcacggacccagattgtaatgaccgcaagcttcctcttcgccgctat cctcaccttatgtatacttcacccgattctacgccttctctcaatcttgaccaggatttcgacactagtttttcgacgaacaatccagaaggtattttgga aggtcatcagaaaacctaaggattttgcagtatacacaggaaatcgccgcaatctcttactcgccatcatcttcgccatcactatcactccaactca acactgctcagaagtcatctccgttacagcatctgaggaaacatgtattcagaacaagaataacaacacttgcaccttcaatcaagcatcggttgtg acattgcaacccctcaatcaagagatgtgtcttctcctcaagaacaacaacaatagttacgctggcatgataacaatcagagtaaaaggattatac ttcacgtgcaaaagaaaggtagaatttttcaccagggatcatcatttctattccgagtcagtgcacagatgccgatttgctggaagttgcaacaccg gaacgtgcgacaaaatcaagcccacggataaacttccagaactatccaataccgccaataaacatcccggctacacattttgtgctgcaagttgcg gatgtcttacatgcagcggctgtttcttctgcttgacaagctgtctcttttatcgttactacgcatacccggaatcgccaacggtgtatacagtttttaac tgtccgctttgggagttcactgttgatgtacacatcacgattcaaaagaatggtcactctgaaggagcacacctagttttacatcccgcaaaaacaa ctcgatggaacaatatacgtctctcactaattgcggcaatgataccacagatgccaatactatcatcaactttcgttacggacggaaacgcgatagc aataacaaaacccgccacgtcaggacagcttataccgaacacgataggtcaacttcaatgtctcacctacgaggatgcgaaacactttaaatgca gttttccaagtaatgcgtgcgtctgttcagcggcaactaatcaggccgtgtgtacctgcagccacggaagtatcagtgaggtgttcaagaaaacacc acttccacaactcaccaagaacgtcatgatatatcaaaaggacaatgaagtaatcgcaaaaaccaaagtcggaactgccatccaactccaccttg tgaccgaagacttaaaaatggtctctacacgtttgaacacaacttgttcaatagtcacgtcggaactatctggatgttacgactgcctcgcgggggc acaaatatcaatagcgtgtaagtcggacaaaggagatgtcacagcagaagtccactgtggaaatcaacttcaaattgcgatttgtaccccaactg gacatatcaacagactcattctcaacttcaattcttcgaagataagagaaagctgtacactttcctgcccgggagggacgatgtcttttgaagtcacc ggagaactcagctacttcaacgacggtgtcctacatcagcagcaaggagtagcgaatcgcttcagtgatatctcgctggctttgccttctttcaatag agtttttgagctgttgagaaaaacattgaacaatattgcctctcccatctttttgctgaaaacttcgtttatcctgatgatttttataccaatctcaactct ccttctacgccttgttgatgcctgtcagagagcaagaacacat [SEQ ID NO: 101]

[0347] SEQ. ID NO: 90 may be encoded by the nucleic acid sequence of SEQ ID NO: 102: ctgaacctacccttgctcatttcaatcatggcaattctgctaccatgcgccatctcgacaagttcccgttgtcctgaagaactgacacttgataagaa aatcatatacgcaacaccgtgtgtggcaaagggagtagctgtagcagtgtactcgcagaagaataggaagcacatctgctggtttccagtcacat gccccaacggtgaaattcgtgattccttccatcaatcaaaggatatgcgaatctgcggaccaaaatgcgaatgcccaaaatggtcgcaattctgct cacattaccacaactctcgcacacagcactcgaaggcgtcgaacatcccatcagagttccttcatttcactcctgaagaagtctgctctttcgaccaa tcggccaagtgtagccaaagaaagaaaattggaacattcaatcaaatcgaactatacgacggcacattgctcctggtcccgaagctacaagtctc tgtgaaggattttctgagttctgacgacttcatatgcatcaattccgaagggaaactagttcctgcaagacgtccctatcgaggaacctccatgttctg cgaaaaacagcaatgtagtgacgatggcgacaagttctgtatgtacgacacaccagtcgccctctatgacttcagcaataccagatcctcagacat cctgattccgataaaggcatggggaaccaccacaagagaattctatccccacaactcagacgccgacactagaaccgtcttcatattgtccaagtg ctcaaaaggtggtgtggaaatacgatcagatcgagaaatcgacagtatagaagcatgcataacttcttactgcgtctttgcatcgcacatctccact acttccctgttcttcccaacggagctggtagtcttcaagtacaccgtcaaaatctcggcatggcagaatggactacgcgttcatcactcctcgatgac ctgcccttcacatccaatatgtgaagtccttcactgcagtctttgcgtagagaaaatctacaatacgcaatgctggactaacctcgacatcctgctca tcgcatcaacgagttttatcttgctatcattcttctggatcttcaagccagtcatcttcctcttcaagaagacgatgaagcttctcctccggctcctcaac ttccttttcaactttccgcgaagaaataaaagggaaagcaaagagacgaactcctccaccacccaacttcgaacctacactcgccacccacgaca ccaccgttcacctcgacaccccgctgtcttcattgctgtaacaatccttgtatacggagcacaacgaggaaactgttgttcggacgtcatttcagttgt cagtgcaacaaaatcctgttcgaccatcaactcgactgaaatatgcactttcaacgaaggagcaaccctaactctccaaccactaggtcaggatgt atgcctgtcattgagaaatcaaaaggacgagcacatcggtctcatctcccttcgcgtgaaaagcattcaccacatctgccatcagaaaatcgaattt tttacaagggatcacgaactgaactcagaatcctcgcatagatgttacagagcaggtagttgtcaaccggacaaatgcgacaacaccaaatcttcg gatacaattcccgagtttagctggagagcatcaaacaaccctggatttaccttctgcaccgcaagttgtcagtgcctcgtatgcgatggatgcttcttc tgcacaccatcatgcttattctacagactatacgcaagagcaacgtcaaacactatctacactatcttcacctgccccgtctgggaactcacagtaa cagcagaagttacgctccaacttcacacaggaagcacgacgcacgcagtaacccttcgccccggacgaactacaagagttaacaacatcaggct gtcactcttgggcacgatttcgcctcagttgcccatcctaacgtcagctttcatcaccgatggcacgaagacagcaatgacaacgagagtacaagct aatgtactgacgccacaaacaccggcacaacttcaatgtgcttctaaagccgacgcaatcacatttctctgccaattctcttcaagagcatgctcttg ctctacaggaccgtacaaagccacatgcacttgtcctgaagggaagatgtcgaaatatctgcaacaaaacaccttgcccttggtcagcaaaaacg tcatcatcgagaaatatgatgacaccatcgcagcccgcacgcaagttggatcggcaatcaacgtccaagtaaacatggaaaacgttaaagtcgcc tcaatacagaatcaaggaacatgcaccatcacggcttccactgtcgaaggatgttattcatgcctcgtcggagccaaaataaccgtcgtatgttact caacggaagaacaaacaacagcggacatcacctgtaacaagcagcaccaaattgccacctgcacaaagagaggaaaactaaccgagttgacc ttccacttcaactccccggaagtatcaacgagttgctccatctcttgccccggaggacagtcgagtttccacctaagcggcctcctcgactatgtcaa cgacgcacaacttcagagagagatagcagtcaacagcgaaatgactatcgcccaaccaagtgatacatttacaaacgtaactcattggatatcct cgtttgtaaacaaattcaatctttccaatctgaaattcctgttagtctgttgtatccttgccttgttttcgccctgtattgtttcaatgacttccgctttgcgt gacaactttcgc [SEQ ID NO: 102]

[0348] SEQ ID NO: 91 may be encoded by the nucleic acid sequence of SEQ ID NO: 103: ctgctccaaaatgcaccagtgatcttcatgatacttgcagctttagtagctgttatggatgctgctaacacacgatgcccgactgagctcacaatcaa taagagaatcatctacgcaactccatgcgtcacaaacggcattgcggtggcaacgtacgacgatcatgaaaaacatcatctatgctggttcccagt cacgtgtccaaatggagaaatacgaccaatgccctcctatccaaatgattcacgtctttgtggaatgaaatgcgagtgtcctcattggtcacagttct gcttgcactataacggagaacatacacaaagatcgcaaatatccagcatcgatccaacacttcttgacttcacaccaaaagaactgtatgatggat cactacttttagtacccagtcttcatgtcataacaaaagattacttggatgacaacgacttcatatgcattagttcgcaaggaaaccgattacctcca acaccgccatataagggaacgtttctgttctgcgaacgacatcaatgcgacttcgccgtcgacaaattctgcgtctatgactcaccaattacgctgct cgacttcgacaacaatgtgacatcacacaacacggttccgatcaaagcatggggaacgacgacaaaggagttttatcctcacatctcgggaactg atgccgaaatcatttctctcatgccaaactgctcaaaaggtggagtcagaatcatcgtagatcgacaaataagtgtagtggaagcatgcatacagt cttactgtatcttcgccttcaacatcacaactacggatctactgtttccaacggaactggttgtcttcaagtataccatcataatcaacgcttggaaga atggactgcacctctacaactcatccattacatgcccatcccatccgatttgcgaaatacttcattgcaacatctgcatagagaaagcttacaacaca caatgctggaccactttcgacaccacgctcattatattactggctaccgtgacgatatccctttattggatcttcgctccaattatctatgtccctcaaa aaactcatgggactgcaactcattcttttcaaatatatggcacaacgatagcccttcaaccactcggacaggaggtgtgcctctccttgaaaaatca gagaaacgaaccaattggccttatttcacttcggatgctcagcatacgccacgtatgtcaacgaaaggttgagttctttacaagagatcatgagatt accaccgaatccgtccacagatgctacagaatgggaagctgtaagcgagataaatgcgacaacacaaaatcgtctgacatcgtcgaggaattca gttggaaagcagcgaattcgccaggattcacattctgcaatccaagctgtggatgtttacttacatgcggatgtctctcctgccagccctcatgtttat tctacagatactatgcaacgccaacatcggaaactatttataccgtgtttacttgcccatcatgggaactaaccgtaacagctgagataacgctgca actacaacaagaaagaaagacgttcacgatgcttctgcgtcccggaagagccactagaacaggtgatataacattatcacttctgtcaaccatttc accccaacttccaatcctcgcgtcagcttttatcactgatgggacaagaactgcgatgacgatgaaggttcaagccaatgaactgacgccagcaac gccggcacaacttcagtgtgcatccaaattagatgcttcccaattcctctgccgtttttcttcgagagcatgcagctgttctgccggatcctacaaagc aacatgcacatgcccaaaaggacatatttcaccatacctgtctcggaatacccttccactaatcaccaaaaatgtcatcatcgaacaacatgatga caccatagctgcccacacccgcgtcggatcggccatctttctacaactcaccatggaaaacgtcaagatagtttctgtacagaacaaaggacaatg ttccataacgagttcgagtattgaaggatgctattcgtgtctgctaggagccaaagtttccatcatatgctactccaccgaagaacaaacgactgct gacatcacttgcgaacaagagcaacaagtcgccatctgcaccaagactggaaagctaaacgaaatgatatttcactttcgctcaccggaaatcgc aacgaagtgcaccgtctcttgccccggagggctatccagcttccaagtagaaggcgcacttgactacgtaaacgatgcacaaatccaagacgaaa tctccaccaaaaacgaggaaaggatcgaccacaatgagaatgcgctcagtacagtaaccggatggatatattcgctcccggatacaatcactgg gtttttcctcgatctgaacctatttacatcgatcaagatattcttgacttgtatcatcttaatcattttgctatactgcattgttttcgtgacatctatcctttc tagggctaag [SEQ ID NO: 103]

[0349] SEQ. ID NO: 92 may be encoded by the nucleic acid sequence of SEQ ID NO: 104: cggaacgttctgctgtccattttgggaatccttcaaatcttcgcactttcagcagtatcagccgatattatttcgtatactccaccgaacacccgttgct cctctcaacagcaaacgataagaaattttgtttatgccgatcaatgcacttcaaaaggcatagcagttgctgttgctaccacggaagacggaagat catcgctatgctggatccagctcacttgtcctctgggacacattcgacttcccattccacagacaccgaatacaggatattgcggaaacacctgcaa atgcccttcgtggaccaccttgtgcagcccatacaacggaagattaacaaaaatgtccaccttcgggaatgtaccatatgaaatcgcccactttatt ccgaaacaagtatgctccttcacaccagctgaacactgttcgaagcacaagaagattggagttttctatcagattcagctatacgacagcacgcttg ttctcgtgccggaactgcacctcaaaacggccgaattcctccacaaagacgactattactgcttcaataaagacggaaaggaatactcgcactacc cacctatggctacttcagggacatccgcgttctgccgtcggcactcctgtgttgaagctgacaaagccactcacttctgcacgtattctactccggta accacatacgagttcaaaaataattcagtcatcatccacgcttggggcacaacaacacgacaatactacccctacgcacccctcccatcaaacga gtccgaactccactttactattcctcgatgcaccacaggaggcatatcaatcgacacgactgaaaaattcaacatgatcgaagtatgcagcgccca cgtttgcgtgtatgtatccgactacgaatctacaacaacgatcctccttccaacctctatcgttattttcgactacacgacaacaatacaagcatgga aagaaggaaccagaaaattcttccacacgctcaagtgcactgggaaaccaatatgtgaggtgatagactgcatcatgtgttgggaacacctctac aattatcaatgttggaccaccacccaaattattacaatcgtcttcgcttccttggtgacgctcatcttctgcaggctattatcaccactattcaagagtc tttggtggataattaccaaaatatttcgagtgttgaagaaaccgatcataacgatcttcaacatcataagatgccgtcgacggaagtacacaattcc gaaattcccgtcatatgcaaaaagaagaacggaacggtacacgcgcatacttccaataactctcatggccgtacgactatgctacacttgctccga cgtcgtcagcttcactacaacgtcccataattgcgaattcaacggcactatcgaaacatgcacggtggatcatgtagtaaacttgcatcttcaaccc ataggacaggagttgtgtctactactacgggactctaaggaacagccagcaggaacaatatcaatacgtcttcacaacgtcgcttatcattgtcaa caaaagactcaatactacacgcgtgatcatagctttcataccgaatccaaccaccaatgctggaatacaaagagctgtacaggagaaccatgtaa aggaacgactacttcgtcaaaaattcctgaacttagttactacgccaatagaagaccaggcttcacatattgttctcctagctgtaagtgtttctggtg cgactggtgcttcagatgcacagagacgtgcatattttatcgtcattacgcagagccccaatctgctgacatataccgagtcttcagttgcccgtcat ggcccattaccgtacaaggaaccataacgaccaatatgatggagcaagaaaaacatcaggactttgttcttcaaccagccacaccatatcgatgg aatcaattgcaaatcgcattaaccgcaactgcatcaccaaatttcccaatactcggaaacttcttcgtcagcaacggggaaaagaccgccatcgtg gatccttcaccggctaaccaactcctaccctattctataggacaactgcaatgtgatacgttggaagacgcaaagaaattcaacacatgtcgatatt ctccaacatcgtgccaatgttcaacaacggtatacgaagcccattgcgaatgtccaaagggagacataaccagctacttcaccaacgcatcacgt actcttccactgtcaagcgacaacatcatcatactcaacgagaacaacaatatcatcgcaaaaacatcagctgggacaagcatttctatccagcta tccattagtaacctgcaaatctacacgagacgatcgcagaacacatgtacggttgaagcaggatccctacatggatgctactcatgttacaccgga gccaatattgaaatctcgtgtacctcctcagtcacagaagaaatagctacgataaaatgtgaaaatcaaaaccaaatcgcagtctgtaaaccggg aggtcatatttcaaagctcacctttcattttcaaaaaagctatatcaatactaattgcagcgtgaagtgtccagcaggagagacctttttcgacgtac acggtcaattgcaattcgtcaccgagccacgcctccagaatgatgtctcacaaggaaaacatcgccgatcggaagaaaaaggaatcgtatcgttc acccgcgacatcggaaactcaattattggcaaaatggccaccatttggacatggttgaaaaattggaagctcatactcatactaattgttacttcgc ctatcatcgtatctttacttcgtttagctttcagaagaaaatcacaa [SEQ ID NO: 104]

[0350] SEQ. ID NO: 93 may be encoded by the nucleic acid sequence of SEQ ID NO: 105: atgactatactcgctattgtctgcatcatctcgttcgcgtcagcagactttgcagcgcacaatcctccagacacgaaatgcgacaagtacaagctttc gcccagacacataatccactccgatcaatgcgctcacgaagggattgtcatcgcaaccacaaaaccatcgtcctccgcaatggaaagaatcttttg ctggttccacgtatcatgtccactaggtcatattcgagtatccttaccatttcaacctaacaacggatactgtggcgacatctgcaagtgccctgaat ggacaaacacttgtagtcactacgacggcaaactgacgacaacttcatctcttccaaatctccctatcaccatctccgactatcgaccaacgcaagt ctgctccttcaatccatcatcactctgctcctccgagcgaaggttaggcgcattctaccaaatccagctctatgacgattcactcgtcattgtccctga actacacctaaaaactgcagaattcttctcaaaggacgacttcacctgcttcaacatacaaggacacaaactggcaaccacacccaacccatcag agacgacaggcactccagcgtactgtcgacggtacccgtgcatatcacctcaatacgccaatacgttctgcgcttactccacgccagtcaccatata tcatcacaacaacgacaccatcatcatcaaagcatggggaaccaccaccagaacctactacccacccctggaatcctctcaaacaccaagcctct ccagctacatcataccacgatgcacagtaggagggatcctcatcgacactacggagcaattcgacatggttgaaacatgtagctcattcgcatgcg tctacatatcgaattactcatcgaattcgaagatcctactaccttcctcgattgtcatctttcaatacaccgtgaaactcaccgcgtggagggaaggc aagcaagtctttcaaaggacacaaacatgcgaaggaaaaccaatatgcgaaatattggactgtttcatctgttgggaacacatcttcaacccacat tgttggaccacaacagaaatcgtcatcacgatcgcgtcatctctcctattgatcattacctgccgccttatatctcctgcacttctgtgtgcgtggtggc tactgaagaaactcattcagacagcgaagaagtgcctcgtcaaaatgttcacgtgtggtcgccgttccggaacgcatacgctccctagtttcccctc ctatcgaagaccatcccgccggacaaggaaagtctcctcgctaatcctcctcgtctacatcaccgccatccaacctacaaaccaatgttcccacatt gtcgccaccagcagcacaagccaatactgcgaatccaactccacagacgaacagtgttacttcgaccaagcgacaacccttcagctgcaacccat aggacaagatgtaagtcttctactacgcaacagcaaaaacgagcccgccggaaccatctccatccgcgttcaggacattgcgtatacttgccgaca aaaaacggaatactatacaagagaccacaaattttataccgaatcacaccaccaatgctggaactcactgagctgcaaaggcgaaacgtgcaaa ggagtcacaacctccacgaacatcgccgaattctccagctatgccaagaagcatcccggcttcacgtactgctcacccagttgcaaatgcctctggt gtgattggtgtttctggtgcaaagaaacgtgccttttttacaaggtctacgcgtatcctgaatcaaaaacaacctatcgcgtcttcacatgcccagca tggtcagtcacagtcgcaggaacaatagtactcgaaacgtcgaaactaagagaaaaacatgacttcgtcctccaccctgctcgtccacacacgtg gaacaaaatcgagctcactctcaccgccacaaccgtcccaaacctacccatcctttcgagctacttcctaacagacggagaagtcgtctcgaaagt ggaacaatcaccagcaggacaactaattccacactccgtaggacagcttcaatgcacctcagagaactctgccaggacattttcatcttgcaccttc tcggctactgcctgcacttgcacgccacgagtatatgaagcaacgtgcaattgtccggaaggaaacatccgaaggtatatcgacaacagtcaact acaactaccactcgcagcaaaagacgtcgtgatactccaacgtggagctgacatacaagcaaaaaccacatctggctcaagtacctccgtacaa gtctccatgcgcggactcactatcttaaccaaacgcctcaacaatgaatgtcagatcgacatcggctctctcaccggatgcttctcatgcattcaagg agcacgactgcaagcgtcatgcacgtcctcacttgcagaagaaacggccgaaatacgctgcggtaaccagacacagctcgccatgtgttctccac acggcaaaatttctgaattagtttttcatttcacatcgtcttcaattaacatgaactgttcagtcatctgccccgccggccaatcaacgatcaccatcca aggctctttggactatgtgaacaatgcgctatttcaagctcccatccagcaagtggacacgcaccgcgccacgtcctccggattctcgctcgacact atgtcccacgcactctctactatcctttctcctatagcaaatgttgtggaacttgtacttgagcagctggcgaaaatacagtcgcctccaacccacct agagatgaat [SEQ ID NO: 105]

[0351] SEQ ID NO: 94 may be encoded by the nucleic acid sequence of SEQ ID NO: 106: tccagagctcgagctcagacgtcctgcattcttcctttcatcatgttacaactactaggctcctattccgctacgaccagcgacacgagatgccccga agaagtcaatgtcaacaaaacggtcatctatgccacctcgtgcgtttcacgcggagtcgcggttgcccggtacaacgactcgggtgaagagaaatt ctgctggttcccgctttcgtgtccttacggagctatacggttcgagaatcgacactcaagtactactctctgcggggaaccatgcaaatgcccaacgt ggtctacctcctgctctttcactaccggttggaggacagccacctctacgttctcctccctgccaaggagaattcgtcagtatcggcccccacaagtc tgctcattttccaactcaccaacgtgcgactccacaaagaaaattggagtttttaaccaagtacaactcttcgacaaccggacattcctcgtaccctc cctggccataagcattcaagaatatatcgatgaaaatgacttcatctgtgtagaccacaaaggatggatacgacgaccagcaagaagatcaaccg gtacaccacatttttgcgccattgtcgccaaaacacataaatgtcatgagaacccgcaaaccttctgcgtcattggtcctaacctcgcagtactcgtc gtcgacaactggaattcctcacacacggtgccaatcatagtgaaggcttggggaacggtctcaaaaatctattacggatatcccagtaatcacatt aaaagtcgtcgcttcatacctctgacatcgcccgtggaatcccaatgttccaaaggagggataacagtcgcaagcaagatacaactacaagtgct agaagcttgcatcaacaactactgtgtgttcctgaaaaatgtctcgacagaagtcattcttttcccgaattccctcattatgtacgattatactgtacg gctaaaagcactgatgaataacgccatcgtacacgaagcgcaagtgggatgtaaagctcaccccatatgcgaaactctgcgatgcacactgtgtt gggagtggatatacaatacacagtgttggactcttctgcaagtcatattagtgtttaccctttttgtagcagcggctctcactattcctgtcatactcatt ctttgcaaagcgatagcagcaacgatatgccttctcataagattgcttctctacctcgttcatcacatgccatggatacgacggaaaacgcagttcag aagatatgtcagtcaacgacgacggaaaaaaccaaagcgaattgtaagctgcacactcttgatacttctccagatacagtcaagcatcgagtgctc ccagttcacttccttaactgccagtgagcacgtctgcaccatctccaagaactcgacggattgtaccctcaatcaggcaacgatgattgcacttcaac ctttgcagcaagagacctgtttgcaaataagagaccacaaaaatgagcatttgggcatactttccgttactcttcgtgatatccaataccgctgtcac aaaagagtcgaattcttctcgcgagaccatgcgttgatatcagaatcccaccatcgctgttacctggcaggtagctgtacccagaagttttgtgaca acgtcaaacccatggatgcactggaagaaatttcgtctatcgcaaacaaaagccccggatatacctactgcactcctacctgtggatgtttgacttg tggcggttgcctattatgcgaggatagctgcattttctatcgagtgtacgccaaacccatttcttcttcaatatatactatcttcaactgccccagctgg gaactctctgtaactgcggacatcactctgagttacggaaacaatatgacgtcaaaaagattgaagctcactccggggaatgtgaactcttggaac aagttgcgacttactctgattagcaccatcacaccaaaattaccaatacttggctcaacgttcatgtcagacggaaacgacatcgccataatcaaac cagcccaccgaggacaattagttgcacacactgctggacaactgcagtgcgccaccttccaagacgctcagaaattccgctgcctcttcgctagcg aaacatgtaaatgttcaaaaggatttcacaagatgtcttgcctctgccccgatggaaacttgaagagactaatgcagccttctccactccccttagtc ggaaagaactttctcctcacacaacaggcaggaactgtttttgcaaaactcaatataggatcgtctcttcaacttcagatagtcatgaaggatctcc acttggtgtccctacaccagaacagtaaatgtacgttcatagccagcgacctcacgggctgctatagttgcatttctggagctacattggaccttgtc tgcacaagtagtgatggagagactacggcgctcctctcttgcccaggcgagactcagatcgcaaaatgtaccacacgaggacatctcaatcgattg attcttcatttcgactcggaaaacatcctctcctcttgcatcgcctcttgccctggaggaagtaccaacattacaatcaaaggaacgttatcttacgtc aacgacgacctcatcaaaccaggactctatgcagctgatcaaaagccccctctcacaaccttaaattctctgggaaaagtttttgacagtgttatac aattcgcatcgtacgtagtggaaagtgttgagtcactctttttgagtctcaaaacctggaagatgatttttttagcatcagttgtactatgcttcttagta attctctcgcgaatgtttcatttcatgttccctacaattcctcctcgtcctaaaaagcat [SEQ. ID NO: 106]

[0352] Modifications

[0353] The capsid proteins described herein may be modified to change a property associated with the capsid protein or the property of a VLP formed from that capsid protein. Methods of modifying proteins are well-known in the art and include, for example, site-directed mutagenesis.

[0354] In one embodiment, a capsid protein is modified to alter its pH stability. In one embodiment, a capsid protein is modified to alter its thermal stability.

[0355] In one embodiment, the capsid protein comprises a mutation corresponding to K162C of SEQ. ID NO: 37. As demonstrated herein, this mutation advantageously generates more stable particles that resist disassembly at pH5 (corresponding to the acidic conditions of the endosome).

[0356] Uses

[0357] In one embodiment, the VLP of the invention is for use in an in vivo method of delivering cargo, optionally to a cell.

[0358] In one embodiment, the invention is directed to the use of the VLP described herein in an in vitro or ex vivo method of delivering cargo, optionally to a cell. In one embodiment, the VLP of the invention is for use in medicine.

[0359] In one embodiment, the VLP of the invention is for use in therapy.

[0360] In one embodiment, the VLP of the invention is for use in diagnosis, optionally wherein the VLP comprises cargo comprising an imaging agent (e.g. a radionuclide).

[0361] In one embodiment, the invention is directed to the use of the VLP described herein in a method of isolating cargo from a mixture, optionally wherein the cargo is selected from a biocatalyst or a nucleic acid.

[0362] In one embodiment, the VLP of the invention is for use in transfection.

[0363] In one embodiment, the invention is directed to the use of the VLP described herein in a method of transfection.

[0364] In one embodiment, the method or use comprises administering the VLP to a subject, optionally wherein the subject is selected from the group consisting of humans, non-human primates, mice, rats, goats, sheep, pigs, cows, horses, camels, alpacas, dogs and cats.

[0365] The subject may be selected from the group consisting of humans, non-human primates, mice, rats, goats, sheep, pigs, cows, horses, camels, alpacas, dogs and cats. In one embodiment, the subject is a human. In one embodiment, the subject is an animal.

[0366] In one embodiment, the VLP of the invention is for use in a vaccine.

[0367] In one embodiment, the VLP of the invention is for use in a bioreactor.

[0368] In one embodiment, the VLP of the invention is for use in delivering a therapeutic agent. In one embodiment, the therapeutic agent is a nucleic acid. In one embodiment, the therapeutic agent is a gene therapy package, optionally wherein the therapy package comprises nucleic acids suitable for CRISPR.

[0369] In one embodiment, the VLP of the invention is for use in delivering an imaging agent, optionally wherein the imaging agent is a contrast agent. The inventors believe that the size of the VLP described herein and the ability to target the VLP described herein to specific cells, tissues or tumours make the VLP well-suited to delivering imaging agents, e.g. radionuclides.

[0370] EXAMPLES

[0371] Example 1: Identification of endogenous retroviruses (ERVs) in nematodes with Env glycoproteins genetically unrelated to other retrovirus Env glycoproteins

[0372] Some nematode ERVs have been identified to contain Env glycoproteins that originate from infectious viruses other than retroviruses. For example, C. elegans ERVs, Cer7 and Cerl3 acquired their Env glycoproteins from a phlebovirus.

[0373] Some phleboviruses are known to have a Class II Env glycoprotein, such as the Rift Valley Fever Virus RVFV (see Malik, Henikoff & Eickbusch (2000); Genome Res;10(9):1307-18). Usually, retroviruses such as HIV contain Class I Env proteins that can diffuse freely in the fluid viral lipid bilayer envelope. By contrast, a Class II Env glycoproteins form rigid outer protein shells. The inventors have previously demonstrated the crystal structure of the RVFV Env glycoprotein, demonstrating the formation of a rigid, icosahedral shell (see Dessau & Modis (2013); Proc Natl Acad Sci USA;110(5):1696-701). The inventors then looked for other ERVs that contain the same Class II Env glycoproteins as RVFV, which demonstrated similar structural features. This was challenging due to low sequence conservation between ERVs with a Class II Env glycoprotein.

[0374] The results of a BLAST search (see Figure 1A) found "Acey_s0020.gl08" - a part of the ERV encoding a complete Gag-Pol-Env polyprotein in the hookworm species Ancylostoma ceylanicum (a domain reconstruction of this ERV is shown in Figure IB and Figure 21A-21B). This ERV was designated the name "Altas virus". The Atlas virus was found to have a Class II Env glycoprotein, which was a uniquely identified feature in retroviruses. The structure of this Class II Env glycoprotein is determined by the inventors in Merchant et al. (2022); Sci Adv;(19) and shown in Figure 1C.

[0375] The inventors therefore identify Atlas virus, an intact retrovirus-like element in the human hookworm that encodes an envelope protein genetically related to GN-GC glycoproteins from phenuiviruses. Cryo-EM, biochemical and cell-based assays show that the Atlas GC and CA proteins are fully. These hookworm endogenous viruses show that retroviruses can have class II fusion proteins instead of usual class I Env. Atlas Gc has membrane fusion activity. Small Atlas capsid particles have T = 7 CA proteins in the asymmetric unit (smallest repeating unit) of the icosahedral assembly. Large Atlas capsid particles have T = 12 CA proteins in the asymmetric unit (smallest repeating unit) of the icosahedral assembly (see Figure 22E). These assemblies are structurally divergent from other retroviral capsids. Retroviruses have unexpected structural & genetic plasticity.

[0376] Figure ID demonstrates an Env-based tree of hookworm ERVs that contain Class II Env glycoproteins. Whilst the Env glycoproteins of these hookworm ERVs differ from those of other retroviruses, the rest of their genome is similar to belpaoviruses, an ancestral retrovirus family. Figure IE shows an RT- based tree of hookworm ERVs with atypical Env glycoproteins.

[0377] Following this, the inventors went on to study the capsid proteins of the identified ERVs that have Class II Env glycoproteins. Here, the inventors were successful in expressing and purifying the capsid protein of the Atlas virus and demonstrating its unique and useful properties.

[0378] Example 2: Expression and purification of the Atlas virus capsid protein

[0379] The Atlas virus was then expressed and purified, using the method detailed below. This methodology is summarised in Figure 2A.

[0380] Details of the construct

[0381] Nucleotide sequences encoding the capsid and matrix proteins of the Atlas virus were designed and synthesized:

[0382] SEQ. ID NO: 107 corresponds to the nucleic acid “ca" sequence, which encodes the capsid domain of the Atlas virus Gag (a. a. 160-403), codon optimised for expression in E.coli: ccgctggcgagccaaatgccgcacccgcagcgtccgccgtttagcaccagcccgaacgcggtgcatgcgccgtacaccaacaacgcgctgctga acttcgttgacgcgagcattctgagcaagatggagctgccgacctttgatggtaacatgctggagttcccggaatttgcgagccgtttcgcgaccctg gtgggcaacaaagcggaactggacgataccaccaagtttagcctgctgaaaagctgcctgcgtggtcgtgcgagccatgcgatccaaggcctga gcgttaccgcggagaactataagatcgcgatggacattctgaacacccacttcaacgataaggtgaccattaaacacgttctgtacagcaaactgg cggagctgccggcgtgcgacccggaaggtcgtaacctgcacaccctgtataaccgtatgttcgcgctgatccgtcagtttgcgaacggcaacgacg atagcaaggaaaccggtctgggcgcgatcctgctgaacaaactgccgctgcgtgtgaagagcaaaatttacgacaagaccgcgaacagccaca acctgagcccgagcgagctgctgcacctgctgaccgatattgttcgtaaagacaccaccctgcaagaaatgagccatcacacccgtagcaccacc ccgcaggatcaatatctgaccttccacgcgagcagcaagatccgtaacaaacgtgctccgccgaac (SEQ ID NO: 107)

[0383] SEQ. ID NO: 108 corresponds to the nucleic acid "maca", which encodes the matrix and capsid domains of the Atlas virus Gag (a. a. 2-403), codon optimised for expression in E.coli: ccgcaggacaatagcagcaaaatccgtcgtcaaatcggtttcttcaaaaaactgatccaacgtagctgcgcgagcatcccgcaaaccttcaccga ctacgatattgacgagaagcacccgaaatttgaccgtctggatgaggaccagctggaaagcctgcgtctggaactgcacaacatccgtagcagcc tgctgaaggcgtacagccgtattaccagcctgcacgatgagtggaccgcgctgcagcaaagcgatccgcgtgagagcaagcagtttgacgaatat atcaccaaatacggtgattatcgtaccagcgtgacccaggcggttacccaactggaggaactggatctgctgctgaacgaggtggacaacgaatt tcgtggtcgtgatctgagcgttagcagcgacagcagcgaaaacccggtgagcaacggccgtttcgacgttaaaagcagccagcaagacgatcaa ccgctggcgagccaaatgccgcacccgcagcgtccgccgtttagcaccagcccgaacgcggtgcatgcgccgtacaccaacaacgcgctgctga acttcgttgacgcgagcattctgagcaagatggagctgccgacctttgatggtaacatgctggagttcccggaatttgcgagccgtttcgcgaccctg gtgggcaacaaagcggaactggacgataccaccaagtttagcctgctgaaaagctgcctgcgtggtcgtgcgagccatgcgatccaaggcctga gcgttaccgcggagaactataagatcgcgatggacattctgaacacccacttcaacgataaggtgaccattaaacacgttctgtacagcaaactgg cggagctgccggcgtgcgacccggaaggtcgtaacctgcacaccctgtataaccgtatgttcgcgctgatccgtcagtttgcgaacggcaacgacg atagcaaggaaaccggtctgggcgcgatcctgctgaacaaactgccgctgcgtgtgaagagcaaaatttacgacaagaccgcgaacagccaca acctgagcccgagcgagctgctgcacctgctgaccgatattgttcgtaaagacaccaccctgcaagaaatgagccatcacacccgtagcaccacc ccgcaggatcaatatctgaccttccacgcgagcagcaagatccgtaacaaacgtgctccgccgaac (SEQ ID NO: 108)

[0384] The sequence was codon optimized for expression in Escherichia coli and cloned into vector pGEX-6P- 1 in frame with the GST tag and the HRV 3C site by blunt end Gibson assembly. A 6 a. a. linker containing triple Gly-Ala repeats was added between the HRV 3C site and the gene insert. These sequences are provided below:

[0385] SEQ ID NO: 109 corresponds to the "pGEX6Pl-Atlas-CA" nucleic acid sequence: acgttatcgactgcacggtgcaccaatgcttctggcgtcaggcagccatcggaagctgtggtatggctgtgcaggtcgtaaatcactgcataattcg tgtcgctcaaggcgcactcccgttctggataatgttttttgcgccgacatcataacggttctggcaaatattctgaaatgagctgttgacaattaatca tcggctcgtataatgtgtggaattgtgagcggataacaatttcacacaggaaacagtattcatgtcccctatactaggttattggaaaattaagggc cttgtgcaacccactcgacttcttttggaatatcttgaagaaaaatatgaagagcatttgtatgagcgcgatgaaggtgataaatggcgaaacaaa aagtttgaattgggtttggagtttcccaatcttccttattatattgatggtgatgttaaattaacacagtctatggccatcatacgttatatagctgacaa gcacaacatgttgggtggttgtccaaaagagcgtgcagagatttcaatgcttgaaggagcggttttggatattagatacggtgtttcgagaattgcat atagtaaagactttgaaactctcaaagttgattttcttagcaagctacctgaaatgctgaaaatgttcgaagatcgtttatgtcataaaacatatttaa atggtgatcatgtaacccatcctgacttcatgttgtatgacgctcttgatgttgttttatacatggacccaatgtgcctggatgcgttcccaaaattagtt tgttttaaaaaacgtattgaagctatcccacaaattgataagtacttgaaatccagcaagtatatagcatggcctttgcagggctggcaagccacgt ttggtggtggcgaccatcctccaaaatcggatctggaagttctgttccaggggcccggtgcgggtgcgggtgcgccgctggcgagccaaatgccg cacccgcagcgtccgccgtttagcaccagcccgaacgcggtgcatgcgccgtacaccaacaacgcgctgctgaacttcgttgacgcgagcattctg agcaagatggagctgccgacctttgatggtaacatgctggagttcccggaatttgcgagccgtttcgcgaccctggtgggcaacaaagcggaact ggacgataccaccaagtttagcctgctgaaaagctgcctgcgtggtcgtgcgagccatgcgatccaaggcctgagcgttaccgcggagaactata agatcgcgatggacattctgaacacccacttcaacgataaggtgaccattaaacacgttctgtacagcaaactggcggagctgccggcgtgcgac ccggaaggtcgtaacctgcacaccctgtataaccgtatgttcgcgctgatccgtcagtttgcgaacggcaacgacgatagcaaggaaaccggtctg ggcgcgatcctgctgaacaaactgccgctgcgtgtgaagagcaaaatttacgacaagaccgcgaacagccacaacctgagcccgagcgagctg ctgcacctgctgaccgatattgttcgtaaagacaccaccctgcaagaaatgagccatcacacccgtagcaccaccccgcaggatcaatatctgacc ttccacgcgagcagcaagatccgtaacaaacgtgctccgccgaactaagcggccgcatcgtgactgactgacgatctgcctcgcgcgtttcggtga tgacggtgaaaacctctgacacatgcagctcccggagacggtcacagcttgtctgtaagcggatgccgggagcagacaagcccgtcagggcgcg tcagcgggtgttggcgggtgtcggggcgcagccatgacccagtcacgtagcgatagcggagtgtataattcttgaagacgaaagggcctcgtgat acgcctatttttataggttaatgtcatgataataatggtttcttagacgtcaggtggcacttttcggggaaatgtgcgcggaacccctatttgtttatttt tctaaatacattcaaatatgtatccgctcatgagacaataaccctgataaatgcttcaataatattgaaaaaggaagagtatgagtattcaacatttc cgtgtcgcccttattcccttttttgcggcattttgccttcctgtttttgctcacccagaaacgctggtgaaagtaaaagatgctgaagatcagttgggtg cacgagtgggttacatcgaactggatctcaacagcggtaagatccttgagagttttcgccccgaagaacgttttccaatgatgagcacttttaaagtt ctgctatgtggcgcggtattatcccgtgttgacgccgggcaagagcaactcggtcgccgcatacactattctcagaatgacttggttgagtactcacc agtcacagaaaagcatcttacggatggcatgacagtaagagaattatgcagtgctgccataaccatgagtgataacactgcggccaacttacttct gacaacgatcggaggaccgaaggagctaaccgcttttttgcacaacatgggggatcatgtaactcgccttgatcgttgggaaccggagctgaatg aagccataccaaacgacgagcgtgacaccacgatgcctgcagcaatggcaacaacgttgcgcaaactattaactggcgaactacttactctagct tcccggcaacaattaatagactggatggaggcggataaagttgcaggaccacttctgcgctcggcccttccggctggctggtttattgctgataaat ctggagccggtgagcgtgggtctcgcggtatcattgcagcactggggccagatggtaagccctcccgtatcgtagttatctacacgacggggagtc aggcaactatggatgaacgaaatagacagatcgctgagataggtgcctcactgattaagcattggtaactgtcagaccaagtttactcatatatac tttagattgatttaaaacttcatttttaatttaaaaggatctaggtgaagatcctttttgataatctcatgaccaaaatcccttaacgtgagttttcgttcc actgagcgtcagaccccgtagaaaagatcaaaggatcttcttgagatcctttttttctgcgcgtaatctgctgcttgcaaacaaaaaaaccaccgct accagcggtggtttgtttgccggatcaagagctaccaactctttttccgaaggtaactggcttcagcagagcgcagataccaaatactgtccttctag tgtagccgtagttaggccaccacttcaagaactctgtagcaccgcctacatacctcgctctgctaatcctgttaccagtggctgctgccagtggcgat aagtcgtgtcttaccgggttggactcaagacgatagttaccggataaggcgcagcggtcgggctgaacggggggttcgtgcacacagcccagctt ggagcgaacgacctacaccgaactgagatacctacagcgtgagctatgagaaagcgccacgcttcccgaagggagaaaggcggacaggtatcc ggtaagcggcagggtcggaacaggagagcgcacgagggagcttccagggggaaacgcctggtatctttatagtcctgtcgggtttcgccacctct gacttgagcgtcgatttttgtgatgctcgtcaggggggcggagcctatggaaaaacgccagcaacgcggcctttttacggttcctggccttttgctgg ccttttgctcacatgttctttcctgcgttatcccctgattctgtggataaccgtattaccgcctttgagtgagctgataccgctcgccgcagccgaacga ccgagcgcagcgagtcagtgagcgaggaagcggaagagcgcctgatgcggtattttctccttacgcatctgtgcggtatttcacaccgcataaattc cgacaccatcgaatggtgcaaaacctttcgcggtatggcatgatagcgcccggaagagagtcaattcagggtggtgaatgtgaaaccagtaacgt tatacgatgtcgcagagtatgccggtgtctcttatcagaccgtttcccgcgtggtgaaccaggccagccacgtttctgcgaaaacgcgggaaaaag tggaagcggcgatggcggagctgaattacattcccaaccgcgtggcacaacaactggcgggcaaacagtcgttgctgattggcgttgccacctcc agtctggccctgcacgcgccgtcgcaaattgtcgcggcgattaaatctcgcgccgatcaactgggtgccagcgtggtggtgtcgatggtagaacga agcggcgtcgaagcctgtaaagcggcggtgcacaatcttctcgcgcaacgcgtcagtgggctgatcattaactatccgctggatgaccaggatgcc attgctgtggaagctgcctgcactaatgttccggcgttatttcttgatgtctctgaccagacacccatcaacagtattattttctcccatgaagacggta cgcgactgggcgtggagcatctggtcgcattgggtcaccagcaaatcgcgctgttagcgggcccattaagttctgtctcggcgcgtctgcgtctggc tggctggcataaatatctcactcgcaatcaaattcagccgatagcggaacgggaaggcgactggagtgccatgtccggttttcaacaaaccatgc aaatgctgaatgagggcatcgttcccactgcgatgctggttgccaacgatcagatggcgctgggcgcaatgcgcgccattaccgagtccgggctgc gcgttggtgcggatatctcggtagtgggatacgacgataccgaagacagctcatgttatatcccgccgtcaaccaccatcaaacaggattttcgcct gctggggcaaaccagcgtggaccgcttgctgcaactctctcagggccaggcggtgaagggcaatcagctgttgcccgtctcactggtgaaaagaa aaaccaccctggcgcccaatacgcaaaccgcctctccccgcgcgttggccgattcattaatgcagctggcacgacaggtttcccgactggaaagc gggcagtgagcgcaacgcaattaatgtgagttagctcactcattaggcaccccaggctttacactttatgcttccggctcgtatgttgtgtggaattgt gagcggataacaatttcacacaggaaacagctatgaccatgattacggattcactggccgtcgttttacaacgtcgtgactgggaaaaccctggcg ttacccaacttaatcgccttgcagcacatccccctttcgccagctggcgtaatagcgaagaggcccgcaccgatcgcccttcccaacagttgcgcag cctgaatggcgaatggcgctttgcctggtttccggcaccagaagcggtgccggaaagctggctggagtgcgatcttcctgaggccgatactgtcgt cgtcccctcaaactggcagatgcacggttacgatgcgcccatctacaccaacgtaacctatcccattacggtcaatccgccgtttgttcccacggag aatccgacgggttgttactcgctcacatttaatgttgatgaaagctggctacaggaaggccagacgcgaattatttttgatggcgttggaatt (SEQ ID NO: 109)

[0386] SEQ. ID NO: 110 corresponds to the "pGEX6Pl-Atlas-MACA" nucleic acid sequence: acgttatcgactgcacggtgcaccaatgcttctggcgtcaggcagccatcggaagctgtggtatggctgtgcaggtcgtaaatcactgcataattcg tgtcgctcaaggcgcactcccgttctggataatgttttttgcgccgacatcataacggttctggcaaatattctgaaatgagctgttgacaattaatca tcggctcgtataatgtgtggaattgtgagcggataacaatttcacacaggaaacagtattcatgtcccctatactaggttattggaaaattaagggc cttgtgcaacccactcgacttcttttggaatatcttgaagaaaaatatgaagagcatttgtatgagcgcgatgaaggtgataaatggcgaaacaaa aagtttgaattgggtttggagtttcccaatcttccttattatattgatggtgatgttaaattaacacagtctatggccatcatacgttatatagctgacaa gcacaacatgttgggtggttgtccaaaagagcgtgcagagatttcaatgcttgaaggagcggttttggatattagatacggtgtttcgagaattgcat atagtaaagactttgaaactctcaaagttgattttcttagcaagctacctgaaatgctgaaaatgttcgaagatcgtttatgtcataaaacatatttaa atggtgatcatgtaacccatcctgacttcatgttgtatgacgctcttgatgttgttttatacatggacccaatgtgcctggatgcgttcccaaaattagtt tgttttaaaaaacgtattgaagctatcccacaaattgataagtacttgaaatccagcaagtatatagcatggcctttgcagggctggcaagccacgt ttggtggtggcgaccatcctccaaaatcggatctggaagttctgttccaggggcccggtgcgggtgcgggtgcgccgcaggacaatagcagcaaa atccgtcgtcaaatcggtttcttcaaaaaactgatccaacgtagctgcgcgagcatcccgcaaaccttcaccgactacgatattgacgagaagcac ccgaaatttgaccgtctggatgaggaccagctggaaagcctgcgtctggaactgcacaacatccgtagcagcctgctgaaggcgtacagccgtat taccagcctgcacgatgagtggaccgcgctgcagcaaagcgatccgcgtgagagcaagcagtttgacgaatatatcaccaaatacggtgattatc gtaccagcgtgacccaggcggttacccaactggaggaactggatctgctgctgaacgaggtggacaacgaatttcgtggtcgtgatctgagcgtta gcagcgacagcagcgaaaacccggtgagcaacggccgtttcgacgttaaaagcagccagcaagacgatcaaccgctggcgagccaaatgccgc acccgcagcgtccgccgtttagcaccagcccgaacgcggtgcatgcgccgtacaccaacaacgcgctgctgaacttcgttgacgcgagcattctga gcaagatggagctgccgacctttgatggtaacatgctggagttcccggaatttgcgagccgtttcgcgaccctggtgggcaacaaagcggaactg gacgataccaccaagtttagcctgctgaaaagctgcctgcgtggtcgtgcgagccatgcgatccaaggcctgagcgttaccgcggagaactataa gatcgcgatggacattctgaacacccacttcaacgataaggtgaccattaaacacgttctgtacagcaaactggcggagctgccggcgtgcgacc cggaaggtcgtaacctgcacaccctgtataaccgtatgttcgcgctgatccgtcagtttgcgaacggcaacgacgatagcaaggaaaccggtctg ggcgcgatcctgctgaacaaactgccgctgcgtgtgaagagcaaaatttacgacaagaccgcgaacagccacaacctgagcccgagcgagctg ctgcacctgctgaccgatattgttcgtaaagacaccaccctgcaagaaatgagccatcacacccgtagcaccaccccgcaggatcaatatctgacc ttccacgcgagcagcaagatccgtaacaaacgtgctccgccgaactaagcggccgcatcgtgactgactgacgatctgcctcgcgcgtttcggtga tgacggtgaaaacctctgacacatgcagctcccggagacggtcacagcttgtctgtaagcggatgccgggagcagacaagcccgtcagggcgcg tcagcgggtgttggcgggtgtcggggcgcagccatgacccagtcacgtagcgatagcggagtgtataattcttgaagacgaaagggcctcgtgat acgcctatttttataggttaatgtcatgataataatggtttcttagacgtcaggtggcacttttcggggaaatgtgcgcggaacccctatttgtttatttt tctaaatacattcaaatatgtatccgctcatgagacaataaccctgataaatgcttcaataatattgaaaaaggaagagtatgagtattcaacatttc cgtgtcgcccttattcccttttttgcggcattttgccttcctgtttttgctcacccagaaacgctggtgaaagtaaaagatgctgaagatcagttgggtg cacgagtgggttacatcgaactggatctcaacagcggtaagatccttgagagttttcgccccgaagaacgttttccaatgatgagcacttttaaagtt ctgctatgtggcgcggtattatcccgtgttgacgccgggcaagagcaactcggtcgccgcatacactattctcagaatgacttggttgagtactcacc agtcacagaaaagcatcttacggatggcatgacagtaagagaattatgcagtgctgccataaccatgagtgataacactgcggccaacttacttct gacaacgatcggaggaccgaaggagctaaccgcttttttgcacaacatgggggatcatgtaactcgccttgatcgttgggaaccggagctgaatg aagccataccaaacgacgagcgtgacaccacgatgcctgcagcaatggcaacaacgttgcgcaaactattaactggcgaactacttactctagct tcccggcaacaattaatagactggatggaggcggataaagttgcaggaccacttctgcgctcggcccttccggctggctggtttattgctgataaat ctggagccggtgagcgtgggtctcgcggtatcattgcagcactggggccagatggtaagccctcccgtatcgtagttatctacacgacggggagtc aggcaactatggatgaacgaaatagacagatcgctgagataggtgcctcactgattaagcattggtaactgtcagaccaagtttactcatatatac tttagattgatttaaaacttcatttttaatttaaaaggatctaggtgaagatcctttttgataatctcatgaccaaaatcccttaacgtgagttttcgttcc actgagcgtcagaccccgtagaaaagatcaaaggatcttcttgagatcctttttttctgcgcgtaatctgctgcttgcaaacaaaaaaaccaccgct accagcggtggtttgtttgccggatcaagagctaccaactctttttccgaaggtaactggcttcagcagagcgcagataccaaatactgtccttctag tgtagccgtagttaggccaccacttcaagaactctgtagcaccgcctacatacctcgctctgctaatcctgttaccagtggctgctgccagtggcgat aagtcgtgtcttaccgggttggactcaagacgatagttaccggataaggcgcagcggtcgggctgaacggggggttcgtgcacacagcccagctt ggagcgaacgacctacaccgaactgagatacctacagcgtgagctatgagaaagcgccacgcttcccgaagggagaaaggcggacaggtatcc ggtaagcggcagggtcggaacaggagagcgcacgagggagcttccagggggaaacgcctggtatctttatagtcctgtcgggtttcgccacctct gacttgagcgtcgatttttgtgatgctcgtcaggggggcggagcctatggaaaaacgccagcaacgcggcctttttacggttcctggccttttgctgg ccttttgctcacatgttctttcctgcgttatcccctgattctgtggataaccgtattaccgcctttgagtgagctgataccgctcgccgcagccgaacga ccgagcgcagcgagtcagtgagcgaggaagcggaagagcgcctgatgcggtattttctccttacgcatctgtgcggtatttcacaccgcataaattc cgacaccatcgaatggtgcaaaacctttcgcggtatggcatgatagcgcccggaagagagtcaattcagggtggtgaatgtgaaaccagtaacgt tatacgatgtcgcagagtatgccggtgtctcttatcagaccgtttcccgcgtggtgaaccaggccagccacgtttctgcgaaaacgcgggaaaaag tggaagcggcgatggcggagctgaattacattcccaaccgcgtggcacaacaactggcgggcaaacagtcgttgctgattggcgttgccacctcc agtctggccctgcacgcgccgtcgcaaattgtcgcggcgattaaatctcgcgccgatcaactgggtgccagcgtggtggtgtcgatggtagaacga agcggcgtcgaagcctgtaaagcggcggtgcacaatcttctcgcgcaacgcgtcagtgggctgatcattaactatccgctggatgaccaggatgcc attgctgtggaagctgcctgcactaatgttccggcgttatttcttgatgtctctgaccagacacccatcaacagtattattttctcccatgaagacggta cgcgactgggcgtggagcatctggtcgcattgggtcaccagcaaatcgcgctgttagcgggcccattaagttctgtctcggcgcgtctgcgtctggc tggctggcataaatatctcactcgcaatcaaattcagccgatagcggaacgggaaggcgactggagtgccatgtccggttttcaacaaaccatgc aaatgctgaatgagggcatcgttcccactgcgatgctggttgccaacgatcagatggcgctgggcgcaatgcgcgccattaccgagtccgggctgc gcgttggtgcggatatctcggtagtgggatacgacgataccgaagacagctcatgttatatcccgccgtcaaccaccatcaaacaggattttcgcct gctggggcaaaccagcgtggaccgcttgctgcaactctctcagggccaggcggtgaagggcaatcagctgttgcccgtctcactggtgaaaagaa aaaccaccctggcgcccaatacgcaaaccgcctctccccgcgcgttggccgattcattaatgcagctggcacgacaggtttcccgactggaaagc gggcagtgagcgcaacgcaattaatgtgagttagctcactcattaggcaccccaggctttacactttatgcttccggctcgtatgttgtgtggaattgt gagcggataacaatttcacacaggaaacagctatgaccatgattacggattcactggccgtcgttttacaacgtcgtgactgggaaaaccctggcg ttacccaacttaatcgccttgcagcacatccccctttcgccagctggcgtaatagcgaagaggcccgcaccgatcgcccttcccaacagttgcgcag cctgaatggcgaatggcgctttgcctggtttccggcaccagaagcggtgccggaaagctggctggagtgcgatcttcctgaggccgatactgtcgt cgtcccctcaaactggcagatgcacggttacgatgcgcccatctacaccaacgtaacctatcccattacggtcaatccgccgtttgttcccacggag aatccgacgggttgttactcgctcacatttaatgttgatgaaagctggctacaggaaggccagacgcgaattatttttgatggcgttggaatt (SEQ. ID NO: 110)

[0387] Protein expression

[0388] Chemically competent BL21 E.coli cells (CAT No. C2527) were transformed with the plasmids described above and plated on LB agar containing ampicillin as selection antibiotics for overnight incubation at 37 °C. Individual colonies were picked and grown in LB supplemented with antibiotics overnight at 37 °C, 180 RPM. The starter cultures were used to inoculate expression cultures at a 1:100 ratio. Expression cultures were grown at 37 °C, 200 RPM until they reached the exponential growth phase with OD600 = 0.6-0.8. Protein expression was induced by adding 400 pM IPTG.

[0389] Protein purification

[0390] The affinity purification protocol was modified from the method described in Pastuzyn et al. (2018) Cell, 172(1), 275-288. BL21 expression cultures containing the pGEX-6P-l-CA plasmid were grown overnight at 18 °C, 180 RPM in the presence of 400 pM IPTG. Cells were harvested by centrifugation at 5000xg for 15 min at 4 °C. For every 500 ml of culture, the pellet was resuspended in 10 ml lysis buffer (500 mM NaCI, 50 mM Tris-HCI, pH 8.0, 5% glycerol) and flash frozen in liquid nitrogen. Frozen pellets were thawed and supplemented with lysozyme; salt active nuclease (Sigma-Aldrich); complete EDTA-free protease inhibitor cocktail (Sigma-Aldrich); DTT, to a concentration of 1 mM; and PMSF, to a concentration of 0.2 mM. Cells were sonicated for 8 min with 5 s pulses, 33% duty cycle, and centrifuged at 30,000xg for 45 min at 4 °C. Supernatant was passed through a 0.45 pm filter and incubated with pre-equilibrated Pierce glutathione agarose resin (Thermo Fisher Scientific) in a gravity flow column for 90 min at 4 °C on a shaker. The bound protein was washed with 50 resin bed volumes of lysis buffer and cleaved with PreScission Protease (GE Healthcare) in cleavage buffer (150 mM NaCI, 50 mM Tris-HCI pH 7.2, 1 mM EDTA, and 1 mM DTT) overnight at 4 °C. Untagged CA was directly eluted, and the bound GST-protease complex was eluted by 20 mM reduced L-glutathione, 10 mM Tris-HCI pH 7.4. Figure 2B shows an expression profile in which the GST-tagged CA protein is identified, as well as the purified CA protein after removal of the GST tag.

[0391] Purified CA was dialysed overnight at 4 °C against assembly buffer (1 M NaCI, 2xPBS, 1 mM DTT). Following buffer exchange, CA was concentrated and incubated at 37 °C for 15 min, then chilled at 4°C for 45 min before being loaded on a Superdex 200 Increase 10 / 300 GL column (Cytiva) for sizeexclusion chromatography (SEC) analysis. Figure 2C shows a SEC trace and a resulting gel in which an both capsid monomers and assembled particles are represented. From the gel, it is evident that both peaks on the SEC trace identify the same protein. The early peak represents assembled capsid proteins which were too large to pass through the column, and hence exited the column in the void fraction. The smaller peak represents the individual capsid monomers that are present.

[0392] The identified particles (eluted from the void fraction) which assembled from the capsid monomers were then imaged using electron microscopy (see Figures 3A and 3B). These particles were found to be icosahedral in shape, with a smooth surface. As seen in Figure 3A, the particles form in two different sizes - the larger particles having a diameter of approximately 65 nm, and the smaller particles having a diameter of approximately 55 nm.

[0393] Using the cryo-EM data obtained, a 3D map of the Atlas capsid particle was reconstituted. The map shown in Figure 4A depicts the smaller particle formed by the Atlas virus capsid monomers, with a resolution of about 3.1 A. This map demonstrates that the capsid particle is a large, hollow protein nanoparticle formed by pentamers and hexamers (420 monomers in total, 60 of which form 12 pentamers). The cryo-EM structures of both the smaller and the larger particle are shown in Figure 4B. These particles differ from the structure of other endogenous retroviral particles, in both size and in arrangement, as shown in Figure 4C. Example 3: Identified properties of VLPs derived from the Atlas virus capsid protein

[0394] Atlas particles can package nucleic acids

[0395] Upon further investigation of the properties of the particle formed by monomers of the Atlas virus capsid protein, it was found that the capsid protein is rich in basic amino acids, containing 14 lysine residues and 8 arginine residues (see Figure 5).

[0396] The capsid has an isoelectric point of 9.57, and is strongly positively charged at physiological pH, which makes it possible for the capsid to bind negatively charged nucleic acids. The capsids will copurify with nucleic acids. Specifically, when RNA is added prior to assembly, the Atlas capsid proteins will package single- or double-stranded RNA, into the capsid particles, Assembly is independent of RNA sequence. Assembly with nucleic acids is shown by the size-exclusion chromatography (SEC) trace in Figure 6A and Figure 23A. Figure 23B shows cryo-EM micrographs of Atlas capsids, where packed nucleic-acid-like materials are visible within the particles.

[0397] The electrophoretic mobility shift assay of Figure 6B shows that the capsid can package different types of exogenous nucleic acids. The charge-charge interaction between Atlas CA protein and nucleic acids is negligible when the protein is in low-concentration, monomer state. However, after concentration-dependent capsid assembly, Atlas capsid can package all four kinds of nucleic acids, with a preference in packaging single- and double-stranded RNA and protecting them against nuclease treatment. The RNA co-purified with the particles was then detected using qRT-PCR, using the method shown in Figure 6C. This figure shows that Atlas particles containing RNA encoding Atlas capsid, i.e. Atlas capsids package their own mRNA. The inventors determines that the nucleic acid that co-purifies with Atlas capsid was predominantly Atlas CA mRNA, not due to any sequencespecific interactions, but because Atlas CA mRNA is the most abundant RNA species in E.coli cells expressing the protein.

[0398] The electrophoretic mobility shift assay (EMSA) of Figure 23C also shows Atlas capsids with and without application of extrinsic nucleic acid (ssRNA, dsRNA, and dsDNA); and with or without treatment of Benzonase nuclease. The nuclease-resistance of the nucleic acids demonstrates that the nucleic acids are packaged within the Atlas capsid particles, because the nuclease is too large to enter the particles. Figure 23C quantifies the protection of nucleic acids, by measuring the band intensities of the packaged nucleic acid demonstrated in three biological repeats from gels in Figure 23D.

[0399] Atlas particles can enter cells via endocytic pathway

[0400] The inventors then tested these particles in three different cell types: HEK293T cells, immortalized bone marrow-derived macrophages (iBMDMs), and HeLa cells. Here, the Atlas particles were labelled, as well as the membrane and the nucleus of cells, as seen in Figure 7A and Figure 24A.

[0401] Using live cell imaging, the inventors found very rapid uptake of the particles by the cells via the endocytic pathway (see Figures 7B-7D and Figure 24B). In Figure 7B and Figure 24B, for example, when the particles were incubated with iBMDMs, the particles remained outside the cells at 4 °C, but when the temperature was increased to 37 °C, after 90 minutes, the particles were visible inside the cells, indicating that cell-entry was via endocytosis. In other cell types, cell uptake was even faster, e.g. in Hela cells where uptake was seen within 15 mins (shown in Figure 7D). In all of these cell types it was shown that upon addition to cultured mammalian cells, the capsid particles enter the cells by endocytosis and accumulate in late endosomes. Atlas particles undergo pH-dependent disassembly

[0402] The stability of the Atlas particle was also tested. In Figure 8A it is evident that capsid particles remain stable at both a salt concentration of 1.5M, and 150mM NaCI (physiological conditions).

[0403] It was further demonstrated that when the buffer pH was reduced to 5.0, the nucleic acids were released. When the particles were incubated at this low pH of 5.0 overnight, the particles broke apart entirely.

[0404] Without being bound by theory, the inventors consider that this pH-controlled disassembly may be caused by the protonation of histidine residues on the inter-subunit interface, which disrupts hydrogen bonding between histidine and asparagine residues, thus destabilising the particle at low pH. Thus, the histidine residues may be acting as a pH sensor.

[0405] Example 4: Search for Atlas-like sequences

[0406] When conducting a PSI-BLAST search of the entire Atlas virus genome against a nematode non- redundant protein sequences database, and filtering for hits with >70% coverage, 118 hits were recovered, all of which were characterised to have a Class II Env glycoprotein (this appears unique to nematode ERVs).

[0407] However, some of the 118 sequences e.g., the hits C. nigoni and C. briggsae, did not retain the property of capsid assembly. Therefore, the 118 hits were subsequently used as a database for a second pBLAST search using Atlas CA core as the query and >50% sequence identity as filter. 11 sequences were identified, along with the query Atlas, with the common features listed below:

[0408] 1. Absence of the major homology region (MHR) (see Figure 9B)

[0409] The MHR is a 20 residue-long sequence that is highly conserved across most retroviruses, but not the Atlas virus. The consensus sequence of the MHR, as discussed in Tanaka, et al. (2016). J Virol 90(4), 1944-1963, is as follows:

[0410] Q[G / K]X2EX4 / 5[Y / F]X2[R / G][F / L]X3H

[0411] (with X being any residue and H being a hydrophobic residue)

[0412] 2. Predicted structural similarity with Atlas capsid

[0413] This will be discussed in more detail in Example 6.

[0414] Example 5: Characteristics of a family of Atlas virus-like ERVs

[0415] Using the searching strategy described in Example 4, the inventors identified 11 sequences of capsid proteins, in addition to the Atlas sequence, which they expect to assemble to form particles and demonstrate similar properties to the Atlas capsid particle.

[0416] These were identified, in part, using BLAST searching, and all 11 identified sequences have >45% identity3and >50% sequence coverage15with Atlas CA core sequence. This family of 12 Atlas-like endogenous viral elements also all have: 1. A class II Env

[0417] 2. High sim ilarityc(>58%) in the CA protein

[0418] 3. Similarity over the majority (77-92%) of the viral genome

[0419] 4. Highly similar structures predicted by AlphaFold (pLDDT > 80, RMSD < 1) (described in more detail in Example 6)aSequence identity as used in this example to refer to exact amino acid matchesbSequence coverage is used hereincSequence similarity as used in this example incorporates both exact amino acid matches and includes near-matches of amino acid residues. The similarity is calculated by the Needleman-Wunsch global alignment algorithm using the BLOSUM62 matrix. The matrix is illustrated at (https: / / en.wikipedia.org / wiki / BLOSUM), as illustrated by Figure 20. The scores are log odds ratios of the observed substitution frequency to the background frequency.

[0420] One further specific trait of these 12 identified ERVs is that they all originate from nematodes, more precisely, the Nematoda phylum. In particular they fall within the species of the Rhabditida order and Chromadorea class within the Nematoda phylum.

[0421] A phylogenetic tree was also constructed based on the "core" capsid domain of the 12 family members. As seen in Figure 10 and Figure 21D, the Atlas ERV (GenBank accession number: EPB78661.1) is positioned near the root.

[0422] A table showing the sequences and sequence-related features is presented in Figure 11. A similar table of 11 belpaoviruses in the Atlas family are shown in Figure 21C ranked by BLAST bitscore. These tables provide the following information:

[0423] Core capsid sequence - This is the core capsid (CA) sequence, without the disordered N- and C- terminal tails.

[0424] Full capsid sequence - This is the full CA sequence including tails. Note that the start and end of each sequence are inferred by analogous to those of other retroviruses based on sequence alignments.

[0425] BLAST hit coverage (or Sequence coverage) with entire Atlas ERV - The first BLAST search conducted used the entire Atlas ERV (gag-pol-env) as the search term. The number in this column refers to what percentage of the entire Atlas ERV sequence is covered in the BLAST hit associated with each row. All are over 70% because that is the cutoff we used (therefore sequences with <70% coverage were rejected). This criteria was used as it was found to empirically ensure that all retained hits had class II Env genes.

[0426] BLAST hit coverage (or Sequence coverage) with full Atlas capsid - In the second BLAST search conducted, the Atlas core CA sequence was used as the search term, and the inventors searched only against the hits from the first BLAST search. Sequences with more than 50% sequence identity in the aligned region of the sequences were retained, thus resulting in the family of 12.

[0427] Amino acid sequence identity to Atlas capsid core - This is the sequence ID (exact amino acid matches) of the core capsid sequence of each retained hit from the second BLAST search relative to Atlas CA.

[0428] Amino acid sequence similarity to Atlas capsid core (structured) sequence - As above, except that near-matches are counted as a match. Near matches include l / L / V, D / E, K / R, S / T, and some other equivalencies. Equivalencies in this case were identified using the Needleman-Wunsch global alignment algorithm using the BLOSUM62 matrix.

[0429] Bitscore - A normalized, log-scaled score used in bioinformatics, most notably with the Basic Local Alignment Search Tool (BLAST), to measure the similarity between two amino acid or nucleic acid sequences. A higher bitscore indicates a greater degree of similarity. The score is standardized so that scores can be compared across different searches, unlike the raw alignment score. A score of 50 or higher is generally considered trustworthy. Mathematical basis: the bitscore is calculated using the formula S' = (lambda*S - In (K)) / ln(2), where S is the raw alignment score, and lambda and K are constants related to the scoring system.

[0430] Example 6; Protein structural modelling for family of 12 identified capsid particles

[0431] High-confidence structural predictions have been generated for each of the capsid proteins described in Example 5, using the software AlphaFold2 or AlphaFoldB. These models are presented in Figure 12.

[0432] As seen in Figures 12 and 13, the RMSD (root mean squared deviation) of main chain atoms has been calculated when each model is superimposed onto the experimentally determined structure of Atlas virus capsid, producing values <1 A. The RMSD scores are important because they show that each family member has a very similar structure to Atlas capsid (defined as RMSD <1 A). The RMSD <1 A criterion is a well-established metric for structural similarity.

[0433] The quality of the models can also be quantified using the pLDDT score, and these values demonstrate that the models are high-confidence, pLDDT > 80. (The only exception is a portion of the last family member (red loop) which has a small non-conserved insertion).

[0434] Together these two sets of numbers clearly show that all family members are very confidently predicted to have a very similar structure indeed to Atlas capsid. These results are provided in more detail in Figure 13.

[0435] In more detail, Figure 13 provides:

[0436] RMSD (Calpha) of AlphaFold model superimposed on Atlas CA core structure - Here the inventors generated an AlphaFold structural model for each CA protein and then superimposed it on the experimentally determined Atlas CA structure. The root mean square deviation (RMSD) of the distance between equivalent pairs of Calpha carbon atoms in the aligned structures was then calculated. This is a standard metric of structural similarity.

[0437] Average pLDDT of Alphafold model - This is a confidence score for the AlphaFold model. Each atom in the model is assigned a confidence score, which expresses the probability that the model is correct. It is very robust and accurate. The value shown in this column is the average score of the entire model (all atom average). A pLDDT >70 is widely considered confident (model probably correct). 80-90 is highly confident and almost certainly correct.

[0438] Percentage sequence coverage of superimposed atomic models - For the structure superpositions of the Alphafold models onto Atlas CA, only structural elements that had an equivalent element in the Atlas reference structure were used for the alignment. Thus, if the query structure has an extra loop or domain, this was not used in the structural alignment, ensuring that the RMSD only reports on equivalent CA regions. This column shows how much of the Atlas core CA sequence is covered by the part of the query sequence retained for structural alignment, i.e., how complete the query CA is relative to Atlas CA (from a structural rather than sequence perspective). Example 7: Experimental data for capsid particles within the Atlas family

[0439] The inventors generated experimental data to support the predictive modelling data suggesting that members of the Atlas-like "family" will behave similarly in terms of capsid assembly and structure.

[0440] First, the inventors tested the other members of the Atlas family for particle assembly capacity. Two sequences were selected for this, one from the closest BLAST hit, KAK6016282.1 (from the nematode Ostertagia ostertagi), and another, VDL73942.1 (from the nematode Nippostrongylus brasiliensis). Both sequences seemed to retain particle formation capacity, with the KAK6016282.1 appearing to have a higher assembly efficiency, as appeared by the higher particle peak in size-exclusion chromatogram. These results are shown in Figures 14A and 14B.

[0441] Cryo-EM images of the particles formed by the capsid protein of KAK6016282.1 are shown in Figure

[0442] 15. Cryo-EM images of the particles formed by the capsid protein of VDL73942.1 are shown in Figure

[0443] 16.

[0444] Expression of KAK6016282.1 (O. ostertagi) and VDL73942.1 (A / . brasiliensis) capsid proteins

[0445] Chemically competent BL21 E.coli cells (CAT No. C2527) were transformed with the plasmids described above and plated on LB agar containing ampicillin as selection antibiotics for overnight incubation at 37 °C. Individual colonies were picked and grown in LB supplemented with antibiotics overnight at 37 °C, 180 RPM. The starter cultures were used to inoculate expression cultures at a 1:100 ratio. Expression cultures were grown at 37 °C, 200 RPM until they reached the exponential growth phase with OD600 = 0.6-0.8. Protein expression was induced by adding 400 pM IPTG.

[0446] Purification of KAK6016282.1 (O. ostertagi) and VDL73942.1 (A / . brasiliensis) capsid proteins

[0447] BL21 expression cultures containing the pGEX-6P-l-CA (O.ostertagi or N. brasiliensis) plasmid were grown overnight at 18 °C, 180 RPM in the presence of 400 pM IPTG. Cells were harvested by centrifugation at 5000xg for 15 min at 4 °C. For every 500 ml of culture, the pellet was resuspended in 10 ml lysis buffer (500 mM NaCI, 50 mM Tris-HCI, pH 8.0, 5% glycerol) and flash frozen in liquid nitrogen. Frozen pellets were thawed and supplemented with lysozyme; salt active nuclease (Sigma- Aldrich); complete EDTA-free protease inhibitor cocktail (Sigma-Aldrich); DTT, to a concentration of 1 mM. Cells were sonicated for 8 min with 5 s pulses, 33% duty cycle, and centrifuged at 30,000xg for 45 min at 4 °C. Supernatant was passed through a 0.45 pm filter and incubated with pre-equilibrated Pierce glutathione agarose resin (Thermo Fisher Scientific) in a gravity flow column for 90 min at 4 °C on a shaker. The bound protein was washed with 50 resin bed volumes of lysis buffer and cleaved with PreScission Protease (GE Healthcare) in cleavage buffer (150 mM NaCI, 50 mM Tris-HCI pH 7.2, 1 mM EDTA, and 1 mM DTT) overnight at 4 °C. Untagged CA was directly eluted, and the bound GST- protease complex was eluted by 20 mM reduced L-glutathione, 10 mM Tris-HCI pH 7.4. Figure 2B shows an expression profile in which the GST-tagged CA protein is identified, as well as the purified CA protein after removal of the GST tag.

[0448] Purified CA was dialysed overnight at 4 °C against assembly buffer (1 M NaCI, 2xPBS, 1 mM DTT). Following buffer exchange, CA was concentrated and incubated at 37 °C for 15 min, then chilled at 4°C for 45 min before being loaded on a Superdex 200 Increase 10 / 300 GL column (Cytiva) for sizeexclusion chromatography (SEC) analysis. The void peak (8 ml elution volume) from the Superdex 200 column contained assembled capsid proteins. Void-peak fractions containing the purified particles were collected and then imaged using electron microscopy (see Figure 22). Example 8: Structural requirements for capsid formation:

[0449] The inventors then went on to determine whether the full capsid protein sequence was required for assembly of the capsid particles.

[0450] Cryo-EM structural reconstruction did not resolve the disordered regions at N- and C-terminus. The inventors therefore designed constructs to remove the disordered loops (shown in Figure 17A) and tested the effect on protein stability and particle assembly.

[0451] As demonstrated by the size exclusion chromatography traces in Figure 17B, truncation of the N- terminal loop did not affect protein solubility and particle assembly efficiency, whereas the truncation of the C-terminal loop caused the protein to be more likely to precipitate and reduced the particle assembly efficiency. In particular, truncation of the C-terminal loop, resulted in a smaller assembled peak in the size-exclusion chromatogram, and the particles seems to be more deformed in negative stain EM. Thus, CTD truncation increases the heterogeneity and reduces the yield of particles. The inventors therefore conclude that, while the core sequence of the Atlas capsid protein (a. a. 29-213) is absolutely required for particle formation, the N-terminal disordered sequence is dispensable, and the C-terminal disordered sequence facilitates capsid formation.

[0452] Thus, the inventors found that the unstructured C-terminal tail of Atlas CA (res. 214-244) promotes - but is not essential - for capsid formation. The unstructured N-terminal tail (res. 1-28) is not required for capsid formation. Therefore, the minimal construct required for capsid formation is res. 29-213 of CA (though the 29-244 construct forms more regular and abundant particles). This minimal construct was designated as the capsid "core" region.

[0453] Example 9: Engineering the Atlas capsid VLPs to improve stability

[0454] Without being bound by theory, the inventors have considered that the pH-controlled disassembly of the Atlas particles may be caused by the protonation of histidine residues on the inter-subunit interface, which disrupts hydrogen bonding between histidine and asparagine residues, thus destabilising the particle at low pH. Thus, the histidine residues, shown in Figure 18A, may be acting as a pH sensor.

[0455] In an attempt alter the stability of the particle (i.e. the ability of the particle to remain intact at low pH), the inventors attempted to manipulate the hydrogen bonds between the hexamer chains, by inducing substitutions at different residues e.g., N172, H112, and H121. However, whilst these residues were found to be important to assembly, mutagenesis at these residues only resulted in decreasing the stability of the particles, as seen in Figure 18B.

[0456] The inventors therefore attempted to introduce a disulfide bridge between the hexamers with the K162C mutation. This position only crosslinks the monomers on a two-fold interface, without interfering with other subunits. As shown in Figure 18C, the K162C mutant of Atlas CA forms more stable particles that resist disassembly under the acidic conditions of the endosome (pH 5) with no nucleic acid release is observed.

[0457] The inventors subsequently also attempted to introduce structurally stabilizing disulfide-crosslinks by introducing a K134C mutation at the dimer interface of the Atlas capsid particle, shown in Figure 25A. Figure 25B shows an SDS-PAGE gel of WT and K134C Atlas capsids in reducing and non-reducing conditions. The band in non-reducing conditions shows that the K134C capsid mutant is disulfide- crosslinked as intended.

[0458] In Figure 25C, cryo-EM micrographs of WT and K134C capsids at pH 7.4 or pH 4.5 demonstrate that the K134C mutant is resistant to low pH-induced disassembly owing to the intermolecular disulfide crosslinks involving Cysl34. Thus, the K134C mutant of Atlas CA forms more stable particles that resist disassembly under acidic conditions.

[0459] Figure 25D demonstrates a model of an anti-CTLA-4 (Hll) nanobody-CA fusion protein predicted by AlphaFold 3 (left) and an SDS-PAGE gel of purified Hll-CA (centre). The negative stain EM images of Hll-CA particles (right) demonstrate that the anti-CTLA-4 Hll nanobody is tolerated and does not interfere with capsid assembly.

[0460] It is therefore concluded that sequences encoding small targeting domains including nanobodies can be fused to the N-terminus of Atlas capsids while preserving the ability of the capsid to form particles. Our cryo-EM structures of Atlas capsids show that the N-terminus of the capsid is located on the outer surface of the particles. Therefore, N-terminally fused targeting domains including anti- CTLA-4 Hll and other nanobodies or antibody fragments, will localize to the outer surface of Atlas particles. As a result, such N-terminally labelled particles can be expected to bind or attach specifically to cells that display the corresponding epitope on their surface. For example, these particles formed from capsids with the anti-CTLA-4 Hll fused to the capsid N-terminus are expected to specifically bind to cells that have the CTLA-4 protein on their surface, most notably CTLA-4- positive T-cells. CTLA-4-positive T-cells are involved in T-cell checkpoint blockade, a mechanism of immune suppression allows some cancers to proliferate undetected by the immune system. CTLA-4- positive T-cells are enriched in various types of cancer. The anti-CTLA-4 Hll-Atlas capsids in this example are expected to specifically bind and enter CTLA-4-positive T-cells via interaction between the Hll nanobody domain and the cellular CTLA-4 cell surface protein, thereby allowing cancer celltype specific targeting of Atlas capsids.

[0461] Example 10: Fusing protein entities on particle surface for targeted delivery

[0462] The inventors have further attempted to engineer the surface of the particle to fuse cargo to the surface.

[0463] In one attempt to do this, the inventors used the natural matrix domain fused to the capsid N- terminus of the capsid domain as a spacer for binding cargo. Therefore, "MA-CA" particles were produces, as seen in Figure 19A. In the micrograph therein, a slightly "fuzzy" surface of the MA-CA particles can be seen, in comparison to the capsid-only particles.

[0464] However, when the inventors attempted to fuse a interleukin-2 domain or an angiopep-2 domain to the CA surface instead of a MA domain, the particles failed to assemble (see Figure 19B). As a result, the inventors attempt to fuse the angiopep-2 domain to the N-terminus of the matric (MA) domain on an MA-CA particle, providing some space between the CA domain and the fused peptide. This is expected to be successful as the matrix domain should be less structurally constrained when adding a peptide to the N-terminus (in comparison to the CA domain which has an intricate assembly).

[0465] This is a desirable feature as N-terminal targeting fusion proteins would allow the targeting of the particles to specific cell types, by allowing the binding of the particles to tissue specific receptors. Furthermore, retroviral MA proteins are also known to interact with lipid envelope of retroviruses. Therefore, the addition of the MA domain is further desirable as this may increase the chance that the particles will be able to interact with lipid membranes.

[0466] The inventors have also attempted to fuse the MA domain to the N-terminus of the CA domain using alternative linkers, such as a glycine-serine linker. However, this was unsuccessful. Therefore, without being bound by theory, it is expected that the linking region between the natural MA domain and CA domain may itself function as a suitable linker when binding proteins to the CA domain. Therefore, it possible that a truncated or variant form of the MA domain may provide a suitable linker for attaching cargo the CA domain.

Claims

1. CLAIMS1. A recombinant, synthetic, or isolated capsid protein, wherein the capsid protein comprises a polypeptide sequence having at least 50% sequence identity to SEQ ID NO: 1; wherein the polypeptide does not comprise a major homology region (MHR).

2. A recombinant, synthetic, or isolated capsid protein, wherein the capsid protein comprises a polypeptide having at least 75% sequence identity, optionally at least 90% sequence identity, to a sequence selected from SEQ. ID NOs: 1-12; wherein the polypeptide does not comprise a major homology region (MHR).

3. The capsid protein of claim 1 or claim 2, wherein the capsid protein comprises a polypeptide having at least 75% sequence identity, optionally at least 90% sequence identity, to a sequence selected from SEQ. ID NOs: 13-48.

4. The capsid protein of any preceding claim, wherein a spacer is attached to the N and / or C- terminus of the capsid protein.

5. The capsid protein of claim 4, wherein the spacer comprises the amino acid sequence of a matrix protein.

6. The capsid protein of claim 5, wherein the matrix protein spacer comprises a polypeptide having at least 75% sequence identity, optionally at least 90% sequence identity, to a sequence selected from SEQ ID NOs: 49-59.

7. The capsid protein of any preceding claim wherein the capsid protein is linked to cargo.

8. A recombinant, synthetic, or isolated nucleic acid sequence which encodes the capsid protein of any preceding claim.

9. The nucleic acid of claim 8, wherein the nucleic acid sequence is codon optimized for expression in a host cell.

10. A vector comprising the nucleic acid according to claim 8 or claim 9.

11. A recombinant cell comprising the vector according to claim 10.

12. A virus-like particle (VLP) comprising the capsid protein of any one of claims 1-7.

13. The VLP of claim 12, wherein the VLP comprises cargo within the interior of the VLP.

14. The VLP of claim 12 or claim 13, wherein the VLP comprises cargo on the exterior of the VLP.

15. A method of producing the VLP according to any one of claims 12-15, the method comprising:(a) expressing the capsid protein according to any one of claims 1-7 in a host cell; and(b) purifying of the expressed capsid protein.

16. The method of claim 15, wherein the method further comprises co-expressing the capsid protein with a Class II Envelope protein.

17. The method of claims 16, wherein the Class II Envelope is selected from SEQ. ID NOs: 83-94.

18. The VLP of any one of claims 12-14 for use in an in vivo method of delivering cargo, optionally to a cell.

19. Use of the VLP of any one of claims 12-14 in an in vitro or ex vivo method of delivering cargo, optionally to a cell.

20. The VLP of any one of claims 12-14 for use in medicine.

21. The VLP of any one of claims 12-14 for use in therapy.

22. The VLP of any one of claims 12-14 for use in diagnosis, optionally wherein the VLP comprises cargo comprising an imaging agent (e.g. a radionuclide).

23. Use of the VLP of any one of claims 12-14 in a method of isolating cargo from a mixture, optionally wherein the cargo is selected from a biocatalyst or a nucleic acid.

24. The VLP of any one of claims 12-14 for use in transfection.

25. Use of the VLP of any one of claims 12-14 in a method of transfection.

26. The method or use of any one of claims 20-25, wherein the method or use comprises administering the VLP to a subject, optionally wherein the subject is selected from the group consisting of humans, non-human primates, mice, rats, goats, sheep, pigs, cows, horses, camels, alpacas, dogs and cats.

Citation Information

Patent Citations

  • Compositions and methods for efficient in VIVO delivery

    WO2023102550A2

  • Cell-type specific membrane fusion proteins

    WO2023158487A1