Protein cages

The fusion protein approach facilitates the assembly of stable protein cages for encapsulating synthetic cargo by using a cleavable sequence, addressing incompatibility and assembly defects in existing methods.

WO2026085582A1PCT designated stage Publication Date: 2026-04-30THE UNIV OF SYDNEY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Existing methods are incompatible with the packaging of synthetic cargo molecules into protein cages within live cells, necessitating in vitro assembly which often results in assembly defects and aggregation.

Method used

A fusion protein comprising an encapsulin and a second polypeptide, such as lanmodulin or maltose binding protein, with a cleavable sequence, allowing for controlled in vitro assembly into stable protein cages.

Benefits of technology

Enables the formation of uniform and thermostable protein cages capable of encapsulating diverse synthetic cargo molecules, overcoming assembly defects and aggregation issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000024_0001_TABLE
    Figure IMGF000024_0001_TABLE
  • Figure IMGF000025_0001_TABLE
    Figure IMGF000025_0001_TABLE
  • Figure IMGF000026_0001_TABLE
    Figure IMGF000026_0001_TABLE
Patent Text Reader

Abstract

The invention relates to fusion proteins comprising encapsulin and a second polypeptide comprising LanM and / or MBP, encapsulins, protein cages derived from the same, and compositions and methods of use thereof.
Need to check novelty before this filing date? Find Prior Art

Description

Protein CagesField of the invention

[0001] The invention relates to fusion proteins, encapsulins, protein cages, compositions comprising the same and methods for encapsulation and delivery of cargo molecules.Related application

[0002] This application claims priority from Australian provisional application no.2024903469, the entire contents of which are incorporated herein by reference.Sequence listing

[0003] A sequence listing in ST. 26 format is filed herewith, the entire contents of which are incorporated herein by reference.Background of the invention

[0004] Cargo-filled protein cages are powerful tools in biotechnology with promising applications as catalytic nanoreactors and vehicles for targeted drug delivery. While endogenous biomolecules can be packaged into protein cages during their expression and self-assembly inside cells, synthetic cargo molecules are typically incompatible with live cells and must be packaged in vitro.

[0005] Accordingly there is a need for new and improved approaches to the design and generation of protein cages for delivery of various cargo in vivo.

[0006] Reference to any prior art in the specification is not an acknowledgment or suggestion that this prior art forms part of the common general knowledge in any jurisdiction or that this prior art could reasonably be expected to be understood, regarded as relevant, and / or combined with other pieces of prior art by a skilled person in the art.Summary of the invention

[0007] The present invention provides a stable fusion protein or encapsulin that is useful for the generation of protein cages.

[0008] Accordingly, in a first aspect there is provided a fusion protein comprising:- a first polypeptide comprising an amino acid sequence of an encapsulin, and- a second polypeptide comprising an amino acid sequence of lanmodulin (LanM) and / or maltose binding protein (MBP),wherein the fusion protein comprises a cleavable sequence for enabling cleavage of the first polypeptide from the second polypeptide.

[0009] In preferred embodiments, the encapsulin is a Family 1 encapsulin.

[0010] In more preferred embodiments, the encapsulin is an encapsulin from a bacteria selected from the group consisting of: Quasibacillus thermotolerans (QtEnc), Thermotoga maritima (TmEnc), Myxococcus xanthus (MxEnc), Dendrosporobacter quercicolus (DgEnc), and Bacillus methanolicus (BmEnc). Accordingly, in any embodiment, the amino acid sequence of encapsulin is an amino acid sequence as set forth in SEQ ID NO: 6, 25 to 28, or 58 to 62, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto.

[0011] In one embodiment, the encapsulin comprises, consists or consists essentially of an amino acid sequence of SEQ ID NO: 6, 25 to 28, or 58 to 62, with 0 to 8 amino acid insertions, deletions, substitutions or additions (or a combination thereof). In some embodiments, the relevant amino acid sequence may have from 0 to 7, preferably from 0 to 6, preferably from 0 to 5, preferably from 0 to 4, preferably from 0 to 3, preferably from 0 to 2, preferably from 0 to 1 amino acid insertions, deletions, substitutions or additions (or a combination thereof).

[0012] In one embodiment, the encapsulin is independently able to form a protein cage and / or is competent to self-assemble. Preferably, the encapsulin comprises an amino acid sequence as set forth in SEQ ID NO: 6, 25 to 28, 58, or 59, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto.

[0013] In other embodiments, the encapsulin is independently unable to form a protein cage and / or is not competent to self-assemble. Preferably, the encapsulin comprises an amino acid sequence as set forth in SEQ ID NO: 60, 61 , or 62, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto.

[0014] Accordingly, in preferred embodiments, the invention provides a fusion protein comprising:- a first polypeptide comprising an amino acid sequence of SEQ ID NO: 6, 25 to 28, or 58 to 62, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto, and- a second polypeptide comprising an amino acid sequence of lanmodulin (LanM) and / or maltose binding protein (MBP)wherein the fusion protein comprises a cleavable sequence for enabling cleavage of the amino acid sequence of SEQ ID NO: 6, 25 to 28, or 58 to 62 (encapsulin) from the second polypeptide sequence.

[0015] In further preferred embodiments, the second polypeptide is LanM. Accordingly, in any embodiment, the amino acid sequence of the second polypeptide is an amino acid sequence as set forth in SEQ ID NO: 7, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto.

[0016] In one embodiment, the LanM comprises, consists or consists essentially of an amino acid sequence of SEQ ID NO: 7 with 0 to 8 amino acid insertions, deletions, substitutions or additions (or a combination thereof). In some embodiments, the relevant amino acid sequence may have from 0 to 7, preferably from 0 to 6, preferably from 0 to 5, preferably from 0 to 4, preferably from 0 to 3, preferably from 0 to 2, preferably from 0 to 1 amino acid insertions, deletions, substitutions or additions (ora combination thereof).

[0017] Accordingly, in preferred embodiments, the invention provides a fusion protein comprising:- a first polypeptide comprising an amino acid sequence of an encapsulin, preferably an encapsulin from Q. thermotolerans, T. maritima, M. xanthus, D. quercicolus, or B. methanolicus, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto, and- a second polypeptide comprising an amino acid sequence of SEQ ID NO: 7, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity theretowherein the fusion protein comprises a cleavable sequence for enabling cleavage of the amino acid sequence of the encapsulin from the amino acid sequence of SEQ ID NO: 7.

[0018] Further still, the invention provides a fusion protein comprising:- a first polypeptide comprising an amino acid sequence of SEQ ID NO: 6, 25 to 28, or 58 to 62, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto, and- a second polypeptide comprising an amino acid sequence of SEQ ID NO: 7, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity theretowherein the fusion protein comprises a cleavable sequence for enabling cleavage of the amino acid sequence of SEQ ID NO: 6, 25 to 28, or 58 to 62, (encapsulin) from the amino acid sequence of SEQ ID NO: 7 (LanM).

[0019] In any embodiment, the second polypeptide is MBP. Accordingly, in any embodiment, the amino acid sequence of the second polypeptide is an amino acid sequence as set forth in SEQ ID NO: 15, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto.

[0020] In one embodiment, the MBP comprises, consists or consists essentially of an amino acid sequence of SEQ ID NO: 15 with 0 to 8 amino acid insertions, deletions, substitutions or additions (or a combination thereof). In some embodiments, the relevant amino acid sequence may have from 0 to 7, preferably from 0 to 6, preferably from 0 to 5, preferably from 0 to 4, preferably from 0 to 3, preferably from 0 to 2, preferably from 0 to 1 amino acid insertions, deletions, substitutions or additions (ora combination thereof).

[0021] Accordingly, in further embodiments, the invention provides a fusion protein comprising:- a first polypeptide comprising an amino acid sequence of an encapsulin, preferably an encapsulin from Q. thermotolerans, T. maritima, M. xanthus, D. quercicolus, or B. methanolicus, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto, and- a second polypeptide comprising an amino acid sequence of SEQ ID NO: 15, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity theretowherein the fusion protein comprises a cleavable sequence for enabling cleavage of the amino acid sequence of the encapsulin from the amino acid sequence of SEQ ID NO: 15.

[0022] Further still, the invention provides a fusion protein comprising:- a first polypeptide comprising an amino acid sequence of SEQ ID NO: 6, 25 to 28, or 58 to 62, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto, and- a second polypeptide comprising an amino acid sequence of SEQ ID NO: 15, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity theretowherein the fusion protein comprises a cleavable sequence for enabling cleavage of the amino acid sequence of SEQ ID NO: 6, 25 to 28, or 58 to 62, (encapsulin) from the amino acid sequence of SEQ ID NO: 15 (MBP).

[0023] It will be appreciated that the fusion protein may comprise the first and second polypeptides in any order. For example, in one embodiment, the N-terminus of the second polypeptide is joined to the C-terminus of the first polypeptide sequence. Alternatively, the N-terminus of the first polypeptide is joined to the C-terminus of the second polypeptide sequence. Preferably, the N-terminus of the first polypeptide is joined to the C-terminus of the second polypeptide sequence (eg in the configuration LanM-QtEnc).

[0024] It will be understood that the cleavable sequence is preferably located at junction of the first polypeptide and second polypeptide sequences, such that cleavage at that sequence enables removal of the second polypeptide from first polypeptide, and subsequent self-assembly of the encapsulin and formation of the protein cage. In particularly preferred embodiments, the cleavable sequence comprises the amino acid sequence that is capable of being recognised by a protease. Examples of such sequences are well known in the art, and include, for example, the sequence Glu-Asn-Leu-Tyr-Phe-GIn 1 (Gly / Ser) which is , recognised and cleaved between the Gin and Gly / Ser residues by the TEV protease, the sequence Leu-Glu-Val-Leu-Phe-GIn 1 (Gly / Pro) which is recognised and cleaved between the Gin and Gly / Pro residues by HRV-3C protease, disulphide-containing sequences (eg: GSF-S-S-Tf and IFN-a2b-HAS) sequences sensitive to thrombin cleavage, furin cleavage, cathepsin B cleavage, and matrix metal loprotease- 1 cleavage.

[0025] Preferably the first polypeptide and the second polypeptide (such as LanM or MBP) are joined via a linker. In preferred embodiments, the cleavable sequence is comprised in the linker region. The linker region may also comprise a flexible sequence. In preferred embodiments, the linker is not a rigid linker. Examples of various flexible and rigid linkers are known to the skilled person and are described for example, in Chen et al., (2013) Advanced Drug Delivery Reviews, 65: 1357-1369.

[0026] In one embodiment, the linker comprises or consists of amino acids. The linker may be any linker known in the art to the skilled person and may be a flexible linker (such as those comprising glycine and / or glutamine residues, or repeats of glycine and serine residues), a rigid linker (such as those comprising glutamic acid and lysine residues, flanking alanine repeats, or comprising (XP)n, wherein X is any amino acid) and / or a cleavable linker (such as sequences that are susceptible by protease cleavage).

[0027] The peptide linker may be any one or more repeats of Gly-Gly-Ser (GGS), Gly-Gly-Gly-Ser (GGGS) or Gly-Gly-Gly-Gly-Ser (GGGGS) or variations thereof. In one embodiment, the linker may comprise or consist of the sequence GGGGSGGGGSGGGGS (G4S)s. In one embodiment, the peptide linker can include the amino acid sequence GGGGGS (a linker of 6 amino acids in length) or even longer. The linker may a series of repeating glycine and serine residues (GS) of different lengths, i.e., (GS)n where n is any number from 1 to 15 or more. For example, the linker may be (GS)s (i.e., GSGSGS) or longer (GS)n or longer. It will be appreciated that n can be any number including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or more.

[0028] In some examples, the fusion protein comprises the amino acid sequence as set forth in SEQ ID NOs: 8, 9, 16, 17, 37 to 44, 68 to 72, or 78 to 82.

[0029] In another aspect, the present invention includes a nucleic acid comprising or consisting of a nucleotide sequence encoding a fusion protein of the invention. Optionally, the nucleic acid sequence is a sequence as set forth in any of SEQ ID NOs: 1 to 5, 13, 14, 21 to 24, 29 to 36, 53 to 57, 63 to 67, or 73 to 77, or a sequence having at least 80%, 85%, 90%, or 95% identity thereto, and encoding a functionally equivalent variant.

[0030] In another aspect, the present invention includes a polypeptide comprising or consisting of an amino acid sequence as set forth in SEQ ID NOs: 58 to 62, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identitythereto. Preferably, the polypeptide comprises or consists of an amino acid sequence as set forth in SEQ ID NOs: 58 or 59.

[0031] In another aspect, the present invention includes a nucleic acid comprising or consisting of a nucleotide sequence encoding an encapsulin of the invention. Optionally, the nucleic acid sequence is a sequence as set forth in any of SEQ ID NOs: 53 to 57, or a sequence having at least 80%, 85%, 90%, or 95% identity thereto, and encoding a functionally equivalent variant. Preferably, the nucleic acid comprises or consists of a nucleotide sequence as set forth in SEQ ID NOs: 53 or 54.

[0032] As used herein, a functionally equivalent variant may be a fragment or variant of a polypeptide of the present invention that retains substantially similar functional activity or substantially the same biological function or activity as the polypeptide, which can be determined using assays described herein.

[0033] In another aspect, the present invention also provides a vector or construct comprising a nucleic acid of the invention.

[0034] In another aspect, the present invention provides a host cell comprising a vector or construct of the invention as described herein.

[0035] It will be appreciated that different host cells may be used in the present invention, for example bacterial cells, yeast cells, insect cells, or mammalian cells.

[0036] In a further aspect, the present invention provides a method for forming a protein cage, the method comprising:- contacting a fusion protein as described herein with an agent for cleaving the cleavable sequence of the fusion protein to cleave the second polypeptide from the fusion protein,thereby forming a protein cage.

[0037] In preferred embodiments, there is provided a method for forming a protein cage, the method comprising:- providing a fusion protein comprising: a first polypeptide comprising an amino acid sequence of SEQ ID NO: 6, 25 to 28, 58, or 59, or a functionally equivalent homolog thereof having at least 80%, 85%, 90%, or 95% identity thereto, and asecond polypeptide comprising an amino acid sequence of lanmodulin (LanM) and / or maltose binding protein (MBP);wherein the fusion protein comprises a cleavable sequence for enabling cleavage of the amino acid sequence of SEQ ID NO: 6, 25 to 28, 58, or 59, (encapsulin) from the second polypeptide- contacting the fusion protein with an agent for cleaving the cleavable sequence, thereby cleaving the second polypeptide from the fusion protein,thereby forming a protein cage.

[0038] In preferred embodiments, there is provided a method for forming a protein cage, the method comprising:- providing a fusion protein comprising: an amino acid sequence of an encapsulin, preferably an encapsulin from Q. thermotolerans, T. maritima, M. xanthus, D. quercicolus, or B. methanolicus, and an amino acid sequence of SEQ ID NO: 7 or a functionally equivalent homolog thereof having at least 80%, 85%, 90%, or 95% identity thereto,wherein the fusion protein comprises a cleavable sequence for enabling cleavage of the amino acid sequence of SEQ ID NO: 7 (LanM) from the encapsulin sequence- contacting the fusion protein with an agent for cleaving the cleavable sequence, thereby cleaving the second polypeptide from the fusion protein,thereby forming a protein cage.

[0039] In preferred embodiments, there is provided a method for forming a protein cage, the method comprising:- providing a fusion protein comprising: an amino acid sequence as set forth in SEQ ID NO: 6, 25 to 28, 58, or 59, and an amino acid sequence of SEQ ID NO: 7 or functionally equivalent homologs thereof having at least 80%, 85%, 90%, or 95% identity thereto,wherein the fusion protein comprises a cleavable sequence for enabling cleavage of the amino acid sequence of SEQ ID NO: 6, 25 to 28, 58, or 59, from the sequence of SEQ ID NO: 7- contacting the fusion protein with an agent for cleaving the cleavable sequence, thereby cleaving the second polypeptide from the fusion protein,thereby forming a protein cage.

[0040] In preferred embodiments, there is provided a method for forming a protein cage, the method comprising:- providing a fusion protein comprising: an amino acid sequence of an encapsulin, preferably an encapsulin from Q. thermotolerans, T. maritima, M. xanthus, D. quercicolus, or B. methanolicus, and an amino acid sequence of SEQ ID NO: 15 or a functionally equivalent homolog thereof having at least 80%, 85%, 90%, or 95% identity thereto,wherein the fusion protein comprises a cleavable sequence for enabling cleavage of the amino acid sequence of SEQ ID NO: 15 (MBP) from the encapsulin sequence- contacting the fusion protein with an agent for cleaving the cleavable sequence, thereby cleaving the second polypeptide from the fusion protein,thereby forming a protein cage.

[0041] In preferred embodiments, there is provided a method for forming a protein cage, the method comprising:- providing a fusion protein comprising: an amino acid sequence as set forth in SEQ ID NO: 6, 25 to 28, 58, or 59, and an amino acid sequence of SEQ ID NO: 15 or functionally equivalent homologs thereof having at least 80%, 85%, 90%, or 95% identity thereto,wherein the fusion protein comprises a cleavable sequence for enabling cleavage of the amino acid sequence of SEQ ID NO: 6, 25 to 28, 58, or 59, from the sequence of SEQ ID NO: 15- contacting the fusion protein with an agent for cleaving the cleavable sequence, thereby cleaving the sequence of SEQ ID NO: 15 from the fusion protein,thereby forming a protein cage.

[0042] In some embodiments, the method comprises further providing a polypeptide comprising an encapsulin protein prior to contacting the fusion protein with an agent for cleaving the cleavable sequence, thereby forming a hybrid protein cage. The skilled person appreciates that a hybrid protein cage may comprise at least two encapsulins and / or variants of encapsulin. For example, the method may comprise contacting a fusion protein comprising an amino acid sequence as set forth in SEQ ID NO: 6 and a polypeptide comprising an amino acid sequences as set forth in SEQ ID NO: 60 with an agent for cleaving the cleavable sequence of the fusion protein, forming a hybrid protein cage comprising amino acid sequences as set forth in SEQ ID NO: 6 and SEQ ID NO: 60. Further, the skilled person understands that hybrid protein cages may comprise encapsulins from different bacteria. The skilled person is aware that encapsulin from different bacteria may share a degree of sequence similarity and / or structural complementarity may have similar folds and be capable of combination to form hybrid protein cages. For example, a hybrid protein cage may comprise an encapsulin from Q. thermotolerans (QtEnc) and B. methanolicus (BmEnc).

[0043] In some embodiments, the polypeptide comprising an encapsulin that is further provided in the method is independently able to form a protein cage and / or is competent to self-assemble. Preferably, the polypeptide comprising an encapsulin comprises an amino acid sequence as set forth in SEQ ID NO: 6, 25 to 28, 58, or 59, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto.

[0044] In other embodiments, the polypeptide comprising an encapsulin that is further provided in the method is independently unable to form a protein cage and / or is not competent to self-assemble. Preferably, the polypeptide comprising an encapsulin comprises an amino acid sequence as set forth in SEQ ID NO: 60, 61, or 62, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto.

[0045] In a further aspect, the present invention provides a method for forming a hybrid protein cage, the method comprising:- contacting a first fusion protein and a second fusion protein as described herein with at least one agent for cleaving the cleavable sequences of the fusion proteins to cleave the second polypeptides from the fusion proteins,thereby forming a hybrid protein cage.

[0046] The skilled person understands that more than one cleavable sequence and more than one agent for cleaving the cleavable sequence may be used in the method. For example, the cleavable sequence of the first fusion protein may be cleaved by TEV protease and the cleavable sequence of the second fusion protein may be cleaved by furin. The cleavable sequence of the first fusion protein and second fusion protein may alternatively be the same sequence, allowing both fusion proteins to be cleaved simultaneously by one agent.

[0047] A hybrid protein cage formed by the method of the invention comprises different encapsulins. For example, a hybrid protein cage comprising amino acid sequence as set forth in SEQ ID NOs: 6 and 60 may be formed when the first fusion protein comprises an amino acid sequence as set forth in SEQ ID NO: 9 and the second fusion protein comprises an amino acid sequence as set forth in SEQ ID NO: 70.

[0048] In some embodiments, the second fusion protein comprises an encapsulin which is independently able to form a protein cage and / or is competent to self-assemble. Preferably, the encapsulin comprises an amino acid sequence of an encapsulin, from Q. thermotolerans, T. maritima, M. xanthus, D. quercicolus, or B. methanolicus. More preferably, the second fusion protein comprises an amino acid sequence as set forth in SEQ ID NO: 6, 25 to 28, 58, or 59 or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto. Most preferably, the second fusion protein comprises an amino acid sequence as set forth in SEQ ID NO: 8, 9, 37 to 44, 68, 69, 78, 79, 91 to 98, 104, 105, 114, or 115.

[0049] In some further embodiments, the second fusion protein comprises an encapsulin which is independently unable to form a protein cage and / or is not competent to selfassemble. Preferably, the second encapsulin comprises an amino acid sequence of an encapsulin, from Q. thermotolerans, T. maritima, M. xanthus, D. quercicolus, or B. methanolicus. More preferably, the second fusion protein comprises an amino acid sequence as set forth in SEQ ID NO: 60, 61, or 62 or a functionally equivalent variantthereof having at least 80%, 85%, 90%, or 95% identity thereto. Most preferably, the second fusion protein comprises an amino acid sequence as set forth in SEQ ID NO: 70 to 72, 80 to 82, 106 to 108, or 116 to 118.

[0050] The skilled person is aware that when encapsulins form protein cages in vitro, the population of encapsulin protein cages formed may comprise a sub-population of defective and / or partially assembled protein cages.

[0051] In a further aspect, the present invention provides a method for rescuing or repairing a defective and / or partially assembled encapsulin protein cage, the method comprising:- providing a defective and / or partially assembled encapsulin protein cage;- contacting a fusion protein as described herein with (i) an agent for cleaving the cleavable sequence of the fusion protein to cleave the second polypeptide from the fusion protein and (ii) the protein cage,thereby rescuing or repairing a defective and / or partially assembled encapsulin protein cage.

[0052] In some embodiments, the fusion protein that is contacted with the defective and / or partially assembled encapsulin protein cage comprises an encapsulin that is independently able to form a protein cage and / or is competent to self-assemble. Preferably, the fusion protein comprises an amino acid sequence as set forth in SEQ ID NO: 6, 25 to 28, 58, or 59, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto.

[0053] In further embodiments, the rescued or repaired encapsulin protein cage is a hybrid protein cage that comprises at least two encapsulins and / or variants of encapsulin.

[0054] In some embodiments of any method of the invention, the method comprises contacting the fusion protein with the agent for cleaving the cleavable sequence at a temperature of between about 4 °C to about 100 °C. The skilled person is aware that different proteases have different optimal temperatures for cleavage, for example, where engineered TEV proteases may be designed to cleave a higher or lower temperatures. In one embodiment, the agent for cleaving the cleavable sequence is a non-engineered TEV protease. In this embodiment, preferably, the temperature is between about 4 °C to about34 °C. More preferably, the temperature is about 10 °C, about 20 °C, about 25 °C, about 30 °C, or about 34 °C. Most preferably, the temperature is about 34 °C.

[0055] In any embodiment, a protein cage or hybrid protein cage formed according to the invention comprises a diameter ranging from about 15 nm to about 48 nm. Preferably, the protein cage has a diameter of about 17-45 nm, more preferably about 18 nm, 24 nm, 32 nm or 42 nm.

[0056] Preferably, the melting temperature (Tm) of a protein cage or hybrid protein cage formed according to the invention is about 55°C to about 105°C, preferably about 95°C.

[0057] Further still, the present invention provides a method for the encapsulation of at least one cargo molecule within a protein cage, the method comprising:- conjugating at least one cargo molecule to at least one cargo loading moiety (CLM),- contacting a fusion protein as described herein (eg comprising a first polypeptide comprising QtEnc, TmEnc, MxEnc, DgEnc, or BmEnc, and a second polypeptide comprising LanM and / or MBP), with an agent for cleaving the second polypeptide (eg LanM) from the fusion protein,- then contacting the cleaved fusion protein with the at least one cargo molecule conjugated to the at least one CLM under suitable conditions and for a sufficient time to enable formation of in vitro complexes of encapsulin (eg QtEnc, TmEnc, MxEnc, DgEnc, or BmEnc) and the cargo molecule,thereby encapsulating the at least one cargo molecule within a protein cage.

[0058] Single or multiple cargo molecules (for example, two or more different cargo molecules) may be individually conjugated to a CLM, for example, the CLM comprising the sequence of a cargo loading peptide (CLP) for Q. thermotolerans, and contacted with the cleaved fusion protein such that in vitro complexes comprising the encapsulin and the single or multiple cargo molecules may be formed. The skilled person appreciates that a cargo molecule may be conjugated to a CLM comprising more than one CLP (for example, a CLP for Q. thermotolerans, and a CLP for T. maritima) and be capable of forming in vitro complexes with different protein cages (for example those formed by QtEnc or TmEnc).

[0059] In some embodiments, two or more conjugated cargo molecules may be in a desired stoichiometric ratio when contacted with the cleaved fusion protein, enabling encapsulation of the two or more conjugated cargo molecules in the desired stoichiometric ratio.

[0060] It will be appreciated that CLPs and their sequences are known in the art and are well-characterised.

[0061] In a further embodiment, the CLM comprises the cargo loading peptide (CLP) for Q. thermotolerans, T. maritima, M. xanthus, D. quercicolus, or B. methanolicus. Preferably, the CLP is for the bacteria that the encapsulin that forms the protein cage is from (for example, a CLP for Q. thermotolerans with QtEnc).

[0062] In a preferred embodiment, the CLM comprises the sequence of at least one of SEQ ID NOs: 18, 19, 20, or 45 to 52 or functionally equivalent homologs thereof having at least 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% identity thereto.

[0063] In a preferred embodiment, the CLM comprises the sequence of SEQ ID NO: 18, 19, 20, and / or 45 to 52.

[0064] In another preferred embodiment, the CLM consists of the sequence of SEQ ID NO: 18, 19, 20, and / or 45 to 52.

[0065] A number of methods for conjugating different types of molecules to a CLM and / or peptide such as a CLP are known in the art. Such methods include conjugating a peptide or protein to the CLM and / or CLP by expressing the conjugate as a recombinant fusion protein, or reacting an amine functional group with a carboxylic acid functional group to form a peptide bond. In another example, small molecules may be conjugated to CLMs and / or CLPs using chemical conjugation, such as by Fmoc solid-phase peptide synthesis. Synthetic drug molecules may also be chemically conjugated to CLMs and / or CLPs, for example by reacting a maleimide-functionalised prodrug form of the synthetic drug with a Cys-modified CLM and / or CLP. Carbohydrates may be conjugated to a CLM and / or CLP by formation of an N-linked or O-linked glycopeptide or glycoprotein, and polynucleotides may be conjugated to CLMs and / or CLPs by forming protein-oligonucleotide conjugates. Cargo molecules may also be conjugated to CLMs and / or CLPs by reaction with the N-terminal amine of the peptide sequence. Azide:alkyne click chemistry reactions may beused to conjugate a cargo molecule to the CLMs and / or CLP, where the CLM and / or CLP and the cargo molecule each have been functionalised with an azide or alkyne functional group. Azide:alkyne click chemistry reactions may include copper-catalyzed alkyne-azide cycloaddition (CuAAC) or strain-promoted azide-alkyne cycloaddition (SPAAC) reactions. Sulfur fluoride exchange (SuFEx) click chemistry reactions may also be used for conjugation, for example by functionalising a cargo molecule with a sulfonyl fluoride that reacts with a tyrosine residue of a CLM and / or CLP. Non-natural amino acids may also be incorporated into the CLM and / or CLP sequence to provide chemical reactivity for conjugation with the cargo molecule. It will be appreciated that other methods and strategies for conjugation of different molecules to a CLM and / or a peptide such as CLP are known in the art.

[0066] In an embodiment, the conjugation of a cargo molecule to a CLM is by conjugation of the cargo molecule to the N-terminal or the C-terminal of the CLM.

[0067] In a further embodiment, the conjugation of a cargo molecule to a CLM is by conjugation of the cargo molecule to the N-terminal of the CLM.

[0068] In another embodiment, the conjugation of a cargo molecule to a CLM is by fusion of a cargo protein with a CLM.

[0069] In another embodiment, the conjugation of a cargo molecule to a CLM is by Fmoc solid-phase peptide synthesis.

[0070] In another embodiment, the conjugation of a cargo molecule to a CLM is by conjugating a maleimide-functionalised cargo molecule to a Cys-modified CLM.

[0071] In a further embodiment, the cargo molecule conjugated to the CLM comprises a cleavable sequence for enabling cleavage of the CLM from the cargo molecule.

[0072] In an embodiment of the method, the contacting the fusion protein with an agent for cleaving the second polypeptide from the fusion protein andcontacting the cleaved fusion protein with the at least one cargo molecule conjugated to the CLM occurs simultaneously.

[0073] In another embodiment of the method, the contacting the fusion protein with an agent for cleaving the second polypeptide andcontacting the cleaved fusion protein with the at least one cargo molecule conjugated to the CLM occurs sequentially.

[0074] In another embodiment of the method, the second polypeptide is removed prior to contacting the cleaved fusion protein with the cargo molecule conjugated to the CLM.

[0075] In some embodiments of any method of the invention, the method comprises contacting the cleaved fusion protein with the at least one cargo molecule conjugated to the CLM at a temperature of between about 4 °C to about 100 °C. Preferably, the temperature is between about 4 °C to about 34 °C. More preferably, the temperature is about 10 °C, about 20 °C, about 25 °C, about 30 °C, or about 34 °C. Most preferably, the temperature is about 34 °C.

[0076] In a further embodiment, there is provided a method for the encapsulation of a cargo molecule within a protein cage, the method comprising:- conjugating at least one cargo molecule to at least one cargo loading moiety (CLM),- contacting a fusion protein comprising the amino acid sequence of QtEnc, TmEnc, MxEnc, DqEnc, or BmEnc (ie an amino acid sequence as set forth in SEQ ID NO: 6, 25 to 28, 58, or 59 or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto) and the amino acid sequence of LanM (ie an amino acid sequence as set forth in SEQ I D NO: 7 or functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto), with an agent for cleaving the LanM amino acid sequence from the fusion protein,- then contacting the cleaved fusion protein with the at least one cargo molecule conjugated to the CLM under suitable conditions and for a sufficient time to enable formation of in vitro complexes of QtEnc, TmEnc, MxEnc, DqEnc, or BmEnc and the cargo molecule,thereby encapsulating the at least one cargo molecule within a protein cage.

[0077] In some embodiments of any method of the invention, the method comprises contacting the fusion protein with an agent for cleaving the LanM amino acid sequence from the fusion protein at a temperature of between about 4 °C to about 100 °C. In one embodiment, the agent for cleaving the LanM amino acid sequence is a non-engineeredTEV protease. In this embodiment, preferably, the temperature is between about 4 °C to about 34 °C. More preferably, the temperature is about 10 °C, about 20 °C, about 25 °C, about 30 °C, or about 34 °C. Most preferably, the temperature is about 34 °C.

[0078] In a further embodiment, there is provided a method for the encapsulation of at least one cargo molecule within a protein cage, the method comprising:- conjugating at least one cargo molecule to at least one cargo loading moiety (CLM),- contacting a fusion protein comprising the amino acid sequence of QtEnc, TmEnc, MxEnc, DqEnc, or BmEnc (ie an amino acid sequence as set forth in SEQ ID NO: 6, 25 to 28, 58, or 59 or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto) and the amino acid sequence of MBP (ie an amino acid sequence as set forth in SEQ ID NO: 15 or functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto), with an agent for cleaving the MBP amino acid sequence from the fusion protein,- then contacting the cleaved fusion protein with the at least one cargo molecule conjugated to the CLM under suitable conditions and for a sufficient time to enable formation of in vitro complexes of QtEnc, TmEnc, MxEnc, DqEnc, or BmEnc and the cargo molecule,thereby encapsulating the at least one cargo molecule within a protein cage.

[0079] In some embodiments of any method of the invention, the method comprises contacting the fusion protein with an agent for cleaving the MBP amino acid sequence from the fusion protein at a temperature of between about 4 °C to about 100 °C. In one embodiment, the agent for cleaving the MBP amino acid sequence is a non-engineered TEV protease. In this embodiment, preferably, the temperature is between about 4 °C to about 34 °C. More preferably, the temperature is about 10 °C, about 20 °C, about 25 °C, about 30 °C, or about 34 °C. Most preferably, the temperature is about 34 °C.

[0080] In preferred embodiments, the encapsulation is performed in vitro.

[0081] In a further aspect of the invention, there is provided a method for the delivery of a cargo molecule to a target cell, the method comprising:- encapsulating at least one cargo molecule in a protein cage according to a method described herein, to obtain a cargo-protein cage assembly- contacting the target cell with the cargo-protein cage assembly under suitable conditions for enabling uptake or internationalisation of the protein cage assembly by the target cellthereby delivering at least one cargo molecule to a target cell.

[0082] In some embodiments, the at least one cargo molecule is encapsulated in a protein cage or hybrid protein cage formed in other embodiments of the invention.

[0083] It will be appreciated that in some embodiments, cargo molecules may be conjugated to cargo loading moiety (CLM) using a variety of methods. For example, the cargo molecule may be a peptide that is conjugated via a polypeptide bond to a CLM. In another example, the cargo molecule may be a maleimide-functionalised prodrug that is conjugated to a Cys-modified CLM. Such methods of conjugating molecules to peptides are well-known in the art.

[0084] It will be appreciated that in any embodiment, the cargo may be any suitable cargo requiring encapsulation in a protein cage, for example, for delivery in vivo. The cargo may be selected from: a small molecule, a nucleic acid or polynucleotide (such as a DNA, RNA, siRNA or shRNA molecule), a carbohydrate, a peptide, or a protein. In certain embodiments, the cargo is a synthetic molecule (eg non-biological molecule), for example a synthetic polymer. Such cargo molecules may include modifications, for example nucleotide modifications, amino acid modifications, unnatural amino acids, and / or labels, which are well-known in the art.

[0085] In particularly preferred embodiments, the methods described herein are performed in the absence of lanthanides.

[0086] Provided herein further is a composition comprising a protein cage or hybrid protein cage as defined herein, and a pharmaceutically acceptable carrier or adjuvant.

[0087] As used herein, “at least 80% sequence identity” will be understood to provide basis for at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%,at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identity to the referenced sequence.

[0088] As used herein, except where the context requires otherwise, the term "comprise" and variations of the term, such as "comprising", "comprises" and "comprised", are not intended to exclude further additives, components, integers or steps.

[0089] Further aspects of the present invention and further embodiments of the aspects described in the preceding paragraphs will become apparent from the following description, given by way of example and with reference to the accompanying drawings.Brief description of the drawings

[0090] Figure 1: Strategies for packaging cargo into protein cages, a) Cargo packaging typically occurs in cells, where cargo that can be produced inside cells (proteins, nucleic acids) is packaged during the expression and assembly process. This method is generally incompatible with synthetic non-biological cargo, b) Current methods for in vitro packaging of synthetic cargo requires cages to be disassembled then subsequently reassembled in the presence of cargo, which can lead to significant assembly defects and aggregation, c) The inventors report a fusion-based strategy for triggering in vitro assembly in encapsulin protein cages, producing highly uniform and thermostable cages that can be loaded with diverse synthetic cargo.

[0091] Figure 2: LanM-QtEnc is an encapsulin fusion that does not self-assemble into cages, a) SDS-PAGE analysis of purified wild-type QtEnc and LanM-QtEnc fusion shows the expected protein bands at 32 and 46 kDa respectively, b) Blue native PAGE analysis of LanM-QtEnc shows that the fusion exists predominantly in a low molecular weight form, while wild-type QtEnc forms the expected high molecular weight assembly, along with a faint smeared band corresponding to other potentially aberrant assemblies, c) Size-exclusion chromatography on a Superose 6 Increase 10 / 300 GL column confirms that LanM-QtEnc exists in an unassembled state, distinct from the cage assembly formed by wild-type QtEnc. d) Size-exclusion chromatography with light scattering detection on a Bio SEC-5 2000 A HPLC column provides sufficient resolution and sensitivity at the megadalton size range to provide further evidence of aggregated or misassembled species in wild-type QtEnc. e) Negative stain transmission electron microscopy of LanM-QtEnc fusion demonstrates mostly unassembled proteins. Images were taken at 100000* magnification (scale bar = 200 nm).

[0092] Figure 3. LanM-QtEnc cleavage triggers high-fidelity assembly of encapsulin cages, a) Analytical size-exclusion chromatography of the crude cleavage mixture on a Bio SEC-5 2000 A HPLC column shows that the in vitro assembled encapsulins are highly monodisperse, without any of the additional peaks observed for wild-type QtEnc assembled in E. coli. Bottom panel: size exclusion chromatography with light scattering detection on Bio SEC-5 2000 A HPLC column shows wt QtEnc (green solid line) and in vitro assembled cleaved LanM-QtEnc (blue solid line) have similar elution profiles with their molar mass calculated from differential refractive index measurement as 9.8 ± 0.2 MDa and 8.2 ± 0.2 MDa respectively. LanM-QtEnc (blue solid line) fusion elutes much later and its molar mass calculated as 28.4 kDa ± 0.8 kDa. Molar mass is represented on each elution profile as dotted line of same colour, b) Sizeexclusion chromatography purification on a Superose 6 Increase 10 / 300 GL column shows clean conversion to a high molecular weight species, c) Blue native PAGE analysis of the purified in vitro assembled encapsulins indicates clean and quantitative conversion to a high molecular weight species, d) SDS-PAGE of the crude cleavage mixture shows that cleavage proceeded to >95% completion, e) Negative stain transmission electron microscopy of in vitro assembled encapsulins shows particles with similar cage morphology and size to wild-type QtEnc. The primary images were obtained at 40000* magnification (scale bar = 250 nm), while the inset images were obtained at 200000* magnification (scale bar = 50 nm). f) Dynamic light scattering measurements show that in vitro assembled encapsulins formed the expected size (42 nm diameter), while wildtype QtEnc was slightly larger due to the presence of minor aggregates or misassembled structures, g) Differential scanning fluorimetry shows that in vitro assembled encapsulins are more thermostable than wild-type QtEnc assembled in E. coli.

[0093] Figure 4: Selective packaging of protein cargo during in vitro assembly, a) Schematic of mNeon-Qt packaging into encapsulin by triggered in vitro assembly, b) Sizeexclusion chromatography on a Superose 6 Increase 10 / 300 GL column shows that in vitro assembly is unaffected by the presence of cargo. C) Negative stain transmission electron microscopy of in vitro packaged encapsulin cages have similar morphology and size to wild-type QtEnc. The larger images were taken at 40000* magnification (scale bar = 250 nm), while the inset images were taken at 200000* magnification (scale bar = 50nm). d) SDS-PAGE analysis of LanM-QtEnc cleavage in presence of cargo before and after SEC purification on a Superose 6 Increase 10 / 300 GL column. mNeon-Qt co-elutes with QtEnc cages while mismatched mNeon-Tm is not packaged, e) Blue native PAGE analysis shows clear in-gel fluorescence for in vitro assemblies containing mNeon-Qt protein cargo, while fluorescence is not visible for mNeon-Tm.

[0094] Figure 5: In vitro packaging of diverse synthetic cargo into encapsulin cages, a) Schematic of TMR-Qt in vitro packaged into encapsulin cages, b) Sizeexclusion chromatography monitoring absorbance at 280 nm on a Superose 6 Increase 10 / 300 GL column shows that the fidelity of in vitro assembly is unaffected by the presence of TMR cargo. Also shown: negative stain transmission electron microscopy of in vitro packaged TMR-Qt into QtEnc cages have similar morphology and size to wildtype QtEnc. The larger images were taken at 40000* magnification (scale bar = 250 nm), while the inset images were taken at 200000* magnification (scale bar = 50 nm). c) Blue native PAGE analysis shows clear in-gel fluorescence for in vitro assemblies with TMR-Qt synthetic cargo selectively packaged, while fluorescence is not visible for TMR-Tm. d) Schematic of Aldox-Qt in vitro packaged into encapsulin cages, e-f) Size-exclusion chromatography monitoring absorbance at 490 nm on a Superose 6 Increase 10 / 300 GL column shows that Aldox-Qt co-elutes with encapsulin assemblies, while doxorubicin without the CLP does not co-elute with encapsulin. g) Fluorescence intensity of the encapsulin fractions from SEC purification show aldoxorubicin fluorescence (emission at 595 nm) when Aldox-Qt is packaged, while the negative control of packaging doxorubicin without the CLP only shows background levels of fluorescence. Error bars correspond to the standard error of the mean across technical measurement replicates.

[0095] Figure 6: Cellular uptake of cargo in vitro packaged inside encapsulins. a) Confocal fluorescence microscopy images of murine RAW 264.7 cells with DAPI fluorescence (nuclear stain) and fluorescence in the FITC channel (mNeon). Merged indicates an overlay of the two channels. In vitro packaged mNeon-Qt was observed inside cells while free mNeon-Qt was not observed, b) Imaged quantification of fluorescence shows that a median of 83% of cells showed uptake of packaged mNeon-Qt, while minimal uptake was observed for free mNeon-Qt and buffer-only controls. Data analysis was performed on three independent fields of view from three biological replicates, c) Representative magnified image of RAW 264.7 cells showing intracellular localisation of packaged mNeon-Qt. d) Confocal microscopy images of RAW 264.7 cellswith DAPI and fluorescence in the TRITC channel (TMR). In vitro packaged TMR-Qt was observed inside cells while free TMR-Qt was not observed, e) Imaged quantification of fluorescence shows that a median of 50% of cells showed uptake of packaged TMR-Qt, while minimal uptake was observed for free TMR-Qt and buffer-only controls. Data analysis was performed on three independent fields of view from three biological replicates, f) Representative magnified image of RAW 264.7 cells showing intracellular localisation of packaged TMR-Qt.

[0096] Figure 7: In vitro assembly of encapsulin protein cages from MBP-QtEnc fusion, a) SDS-PAGE of fractions of MBP-QtEnc fusion obtained by Ni-NTA affinity purification, b) Purification of MBP-QtEnc fusion by size exclusion chromatography, c) SDS-PAGE showing cleavage of MBP-QtEnc fusion to produce QtEnc. d) Size exclusion chromatography on encapsulin purification column shows cleaved product (QtEnc) elutes in the encapsulin protein cage fraction.

[0097] Figure 8: In vitro assembly of encapsulin protein cages from LanM-MxEnc fusion, a) SDS-PAGE of fractions of LanM-MxEnc fusion obtained by Ni-NTA affinity purification, b) Purification of LanM-MxEnc fusion by size exclusion chromatography, c) SDS-PAGE showing cleavage of LamM-MxEnc fusion to produce MxEnc. d) Size exclusion chromatography on encapsulin purification column shows cleaved product (MxEnc) elutes in the encapsulin protein cage fraction.

[0098] Figure 9: In vitro assembly of encapsulin protein cages from LanM-TmEnc fusion, a) SDS-PAGE of fractions of LanM-TmEnc fusion obtained by Ni-NTA affinity purification, b) Purification of LanM-TmEnc fusion by size exclusion chromatography, c) SDS-PAGE showing cleavage of LamM-TmEnc fusion to produce TmEnc. d) Size exclusion chromatography on encapsulin purification column shows cleaved product (TmEnc) elutes in the encapsulin protein cage fraction.

[0099] Figure 10: MBP-QtEnc fusion can be cleaved and assembled at a range of temperatures, a) SDS-PAGE shows that MBP-QtEnc fusions may be completely cleaved at a range of temperatures, including from 10 °C to 34 °C. b) Size exclusion chromatography on encapsulin purification column shows that cleaved product (QtEnc) obtained by cleavage at different temperatures all elute in the encapsulin protein cage fraction.

[0100] Figure 11: In vitro assembly of protein cages from MBP- variant QtEnc fusions, a) SDS-PAGE of fractions of MBP-variant QtEnc fusion obtained by Ni-NTA affinity purification, b) Size exclusion chromatography on encapsulin purification column shows cleaved product (variant QtEnc) elutes in the encapsulin protein cage fraction for Glass9 and Letter11 variants but not for Vegetal 0, Pigs13, or Slay13. c) SDS-PAGE shows pure variant QtEnc for Glass9 and Letterl 1 of correct size.

[0101] Figure 12: Fusions allow in vitro assembly of hybrid protein cages from WT / variant and variant / variant encapsulins. a) Size exclusion chromatography on encapsulin purification column shows that the cleaved product of mixtures of WT and variant encapsulin elute in the encapsulin protein cage fraction, demonstrating that even assembly incompetent variants can be assembled in protein cages with assembly competent variants, b) SDS-PAGE of MBP-variant QtEnc fusion obtained by Ni-NTA affinity purification, c) SDS-PAGE of encapsulin fraction shows WT / variant hybrids as two bands due to different sizes, and variant / variant hybrids as a single band due to similar size.

[0102] Figure 13: Cargo packaging and co-packaging. Detection of A) mNeon fluorescence using Alexa488 and B) mCherry fluorescence using Alexa657 in protein cages formed from cleavage of LanM-QtEnc fusions and loaded with the following cargo: mNeon 100%, mCherry 100%, mNeon 50% and mCherry 50%, mNeon 80% and mCherry 20%, and mNeon 20% and mCherry 80%.Sequence informationTable 1: Sequence detailsSEQ ID NO: Description Nucleic acid / amino acid sequence1 Codon-optimised ATGAACAAAAGCCAACTTTATCCGGATTCACCACTGAC nucleic acid sequence GGATCAGGACTTCAACCAATTAGACCAAACCGTGATTG ofwild-type QtEnc AGGCTGCTCGTCGTCAGCTGGTGGGTCGTCGCTTCAT gene TGAGTTATATGGCCCATTGGGGCGTGGCATGCAGAGT GTCTTCAACGATATCTTCATGGAGTCTCATGAAGCGAA AATGGACTTCCAGGGCAGCTTTGACACGGAGGTAGAG TCCTCCCGTCGTGTAAACTATACCATTCCGATGTTATAT AAAGACTTCGTGCTTTACTGGCGCGATCTGGAACAGAG CAAGGCACTCGATATTCCGATCGAC I l l i CAGTGGCAGCGAACGCTGCCCGCGACGTTGCGTTCCTGGAAGATCA GATGA I l l i CCATGGAAGCAAAGAATTTGATATCCCGG GTCTGATGAACGTGAAAGGTCGCCTGACCCATCTGATT GGCAATTGGTATGAGTCGGGTAACGCCTTTCAGGATAT TGTGGAGGCCCGCAATAAATTACTCGAAATGAACCACA ATGGCCCATATGCTCTCGTGCTGTCCCCGGAGCTGTA CTCACTCTTACATCGTGTGCATAAAGACACGAATGTGC TGGAGATCGAACACGTGCGCGAGTTGATTACTGCTGG GGTTTTTCAGTCGCCTGTCCTCAAAGGGAAAAGTGGTG TGATCGTAAACACCGGTCGCAACAATCTGGATTTGGCT ATCTCGGAAGA I l l i GAGACTGCATACCTGGGCGAGG AAGGTATGAACCATCCCTTTCGCGTGTACGAGACAGTT GTTCTGCGCATCAAACGCCCGGCGGCCATTTGTACTTT AATCGATCCGGAAGAATAACodon-optimised ATGCATCACCACCATCATCACGGAGGAAGTCCGACAA nucleic acid sequence CGACCACCAAAGTTGACATCGCCGCATTTGATCCAGA of LanM-QtEnc fusion TAAAGACGGAACGATCGACTTGAAAGAGGCACTGGC gene TGCTGGTTCGGCGGCATTTGATAAACTGGACCCGGAT 6xHis tag shown in AAGGACGGTACTCTTGATGCAAAAGAATTAAAAGGGC underline GTGTCTCCGAAGCTGACTTGAAGAAACTGGACCCTGA CAACGACGGCACATTGGACAAGAAGGAGTACCTTGCLanM in bold AGCTGTGGAGCAATTCAAAGCAGCGAACCCAGACAA TEV protease site in CGATGGCACGATTGACGCTCGCGAACTGGCTTCCCC italics GGCAGGTTCCGCCTTGGTGAATCTTATTCGGGGGGGC TCCGGTGGCTCGGAGAATCT7TA I l l i CAGTCTAACAA QtEnc gene in boldAA GCCAACTTTATCCGGA TTCACCACTGA CGGA TCAGand italicsGACTTCAACCAATTAGACCAAACCGTGATTGAGGCTG CTCGTCGTCAGCTGGTGGGTCGTCGCTTCATTGAGTT ATATGGCCCATTGGGGCGTGGCATGCAGAGTGTCTTC AACGATATCTTCATGGAGTCTCATGAAGCGAAAATGG ACTTCCAGGGCAGCTTTGACACGGAGGTAGAGTCCTC CCGTCGTGTAAACTA TACCA TTCCGA TGTTA TA TAAAG ACTTCGTGCTTTACTGGCGCGATCTGGAACAGAGCAA GGCACTCGATATTCCGA TCGA CTTTTCA GTGGCA GCG AACGCTGCCCGCGACGTTGCGTTCCTGGAAGATCAG ATGATTTTCCATGGAAGCAAAGAATTTGATATCCCGG GTCTGA TGAA CGTGAAAGGTCGCCTGACCCA TCTGA T TGGCAA TTGGTA TGAGTCGGGTAACGCCTTTCAGGA T ATTGTGGAGGCCCGCAATAAATTACTCGAAATGAACC ACAATGGCCCATATGCTCTCGTGCTGTCCCCGGAGCTGTA CTCACTCTTACATCGTGTGCA TAAA GA CA CGAA T GTGCTGGA GATCGAA CA CGTGCGCGA GTTGA TTACTG CTGGGGTTTTTCAGTCGCCTGTCCTCAAAGGGAAAAG TGGTGTGATCGTAAA CA CCGGTCGCAA CAA TCTGGA T TTGGCTATCTCGGAAGA TTTTGA GACTGCA TACCTGG GCGAGGAAGGTATGAACCATCCCTTTCGCGTGTACGA GACAGTTGTTCTGCGCATCAAACGCCCGGCGGCCATT TGTACTTTAA TCGA TCCGGAA GA AT A ACodon-optimised ATGCACCACCATCATCATCACGGAGGTTCAGTCTCCAA nucleic acid sequence GGGAGAGGAGGACAATATGGCTAGTCTTCCGGCCAC of mNeonGreen with TCATGAGTTACATATCTTCGGATCCATAAACGGCGTT targeting peptide GATTTCGATATGGTGGGTCAAGGCACTGGTAACCCCA mNeon-Qt for ATGATGGCTACGAAGAACTTAACTTAAAATCTACTAA encapsulation AGGCGATCTTCAGTTTTCCCCATGGATACTTGTGCCTC 6xHis tag shown in ATATTGGCTACGGGTTTCACCAATATCTGCCTTATCCG underline GATGGAATGTCCCCCTTCCAAGCCGCAATGGTAGACG GCAGTGGCTATCAAGTCCACCGTACCATGCAGTTTGAC-terminal targeting AGATGGCGCATCCCTGACAGTTAATTATCGGTATACA peptide sequence from TACGAGGGGTCGCATATTAAAGGAGAGGCGCAAGTC Q. thermotolerans in AAGGGGACAGGGTTTCCCGCCGATGGGCCAGTCATG italics ACAAACTCGTTAACTGCCGCCGACTGGTGCAGATCGA mNeonGreen gene in AGAAAACCTACCCAAACGATAAGACGATCATATCTAC bold CTTTAAATGGTCTTACACTACGGGTAACGGAAAACGC TACAGATCAACCGCGCGGACAACGTACACCTTTGCTA AGCCCATGGCAGCGAACTACTTGAAGAATCAGCCGA TGTACGTGTTTAGAAAGACCGAGCTTAAACACTCGAA AACTGAATTGAATTTTAAAGAATGGCAGAAAGCTTTTA CGGACGTAATGGGCATGGACGAACTGTATAAGTCCG GGTCGGGAGGCTCCGGAAAAAAGAAAGGCT7TACTGT CGGGTCGTTAATTCAGTAACodon-optimised ATGCACCACCATCATCATCACGGAGGTTCAGTCTCCAA nucleic acid sequence GGGAGAGGAGGACAATATGGCTAGTCTTCCGGCCAC of mNeonGreen with TCATGAGTTACATATCTTCGGATCCATAAACGGCGTT targeting peptide GATTTCGATATGGTGGGTCAAGGCACTGGTAACCCCA mNeon-Tm for ATGATGGCTACGAAGAACTTAACTTAAAATCTACTAA encapsulation AGGCGATCTTCAGTTTTCCCCATGGATACTTGTGCCTC 6xHis tag shown in ATATTGGCTACGGGTTTCACCAATATCTGCCTTATCCG underline GATGGAATGTCCCCCTTCCAAGCCGCAATGGTAGACG GCAGTGGCTATCAAGTCCACCGTACCATGCAGTTTGA AGATGGCGCATCCCTGACAGTTAATTATCGGTATACAC-terminal targeting TACGAGGGGTCGCATATTAAAGGAGAGGCGCAAGTC peptide sequence from AAGGGGACAGGGTTTCCCGCCGATGGGCCAGTCATG T. maritima in italics ACAAACTCGTTAACTGCCGCCGACTGGTGCAGATCGA AGAAAACCTACCCAAACGATAAGACGATCATATCTACmNeonGreen gene in CTTTAAATGGTCTTACACTACGGGTAACGGAAAACGC bold TACAGATCAACCGCGCGGACAACGTACACCTTTGCTA AGCCCATGGCAGCGAACTACTTGAAGAATCAGCCGA TGTACGTGTTTAGAAAGACCGAGCTTAAACACTCGAA AACTGAATTGAATTTTAAAGAATGGCAGAAAGCTTTTA CGGACGTAATGGGCATGGACGAACTGTATAAGTCAG GCGGGAACACAGGAGGCGATTTAGGCATTCGCAAGTT ATA ACodon-optimised ATGCACCACCACCACCACCACAGCGGCGC / / / IGAATT nucleic acid sequence TAAGCTGCCGGACATTGGCGAAGGCATCCACGAAGGT ofTEV protease GAAATTGTCAAA TGGTTTGTGAAACCGGGCGA TGAAGT 6xHis tag shown in GAACGAAGACGATGTATTGTGCGAAGTGCAAAATGACA underline AGGCGGTTGTCGAAATTCCCTCCCCGGTCAAAGGGAA AGTGCTTGAAATCCTCGTCCCGGAGGGAACAGTGGCAGene in bold ACGGTCGGGCAAACGCTCATCACGCTCGATGCGCCGG dihydrolipoyllysine- GTTATGAAAACATGACGAGCAGCGGCCTGGTGCCGCG binding domain for CGGATCCGGAGAAAGCTTGTTTAAGGGACCACGTGAT solubility in italics TACAACCCGATATCGAGCACCATTTGTCATTTGACGA ATGAATCTGATGGGCACACAACATCGTTGTATGGTAT TGGATTTGGTCCCTTCATCATTACAAACAAGCACTTGT TTAGAAGAAATAATGGAACACTGTTGGTCCAATCACT ACATGGTGTATTCAAGGTCAAGAACACCACGACTTTG CAACAACACCTCATTGATGGGAGGGACATGATAATTA TTCGCATGCCTAAGGATTTCCCACCATTTCCTCAAAAG CTGAAATTTAGAGAGCCACAAAGGGAAGAGCGCATA TGTCTTGTGACAACCAACTTCCAAACTAAGAGCATGT CTAGCATGGTGTCAGACACTAGTTGCACATTCCCTTC ATCTGATGGCATATTCTGGAAGCATTGGATTCAAACC AAGGATGGGCAGTGTGGCAGTCCATTAGTATCAACTA GAGATGGGTTCATTGTTGGTATACACTCAGCATCGAA TTTCACCAACACAAACAATTATTTCACAAGCGTGCCG AAAAACTTCATGGAATTGTTGACAAATCAGGAGGCGC AGCAGTGGGTTAGTGGTTGGCGATTAAATGCTGACTC AGTATTGTGGGGGGGCCATAAAGTTTTCATGGTCAAA CCTGAAGAGCCTTTTCAGCCAGTTAAGGAAGCGAATA GGGCTCATGAATGAWt QtEnc amino acid MNKSQLYPDSPLTDQDFNQLDQTVIEAARRQLVGRRFIE sequence (bold) LYGPLGRGMQSVFNDIFMESHEAKMDFQGSFDTEVESS RRVNYTIPMLYKDFVLYWRDLEQSKALDIPIDFSVAANA ARDVAFLEDQMIFHGSKEFDIPGLMNVKGRLTHLIGNWY ESGNAFQDIVEARNKLLEMNHNGPYALVLSPELYSLLHR VHKDTNVLEIEHVRELITAGVFQSPVLKGKSGVIVNTGRN NLDLAISEDFETAYLGEEGMNHPFRVYETVVLRIKRPAAI CTLIDPEE*LanM amino acid PI I I I KVDIAAFDPDKDGTIDLKEALAAGSAAFDKLDPDK sequence DGTLDAKELKGRVSEADLKKLDPDNDGTLDKKEYLAAVE QFKAANPDNDGTIDARELASPAGSALVNLIR*6xHis-LanM-TEV- MHHHHHHGGSP / / / IKVDIAAFDPDKDGTIDLKEALAAG QtEnc amino acid SAAFDKLDPDKDGTLDAKELKGRVSEADLKKLDPDNDGT sequence LDKKEYLAAVEQFKAANPDNDGTIDARELASPAGSALVNL His tag shown in bold / PGGSGGSENLYFQSNKSQLYPDSPLTDQDFNQLDQTVI EAARRQLVGRRFIELYGPLGRGMQSVFNDIFMESHEAKMLanM shown in italics DFQGSFDTEVESSRRVNYTIPMLYKDFVLYWRDLEQSKA TEV underlined (not LDIPIDFSVAANAARDVAFLEDQMIFHGSKEFDIPGLMNVK bold) GRLTHLIGNWYESGNAFQDIVEARNKLLEMNHNGPYALV LSPELYSLLHRVHKDTNVLEIEHVRELITAGVFQSPVLKGKGlycine-serine linker SGVIVNTGRNNLDLAISEDFETAYLGEEGMNHPFRVYETVsequence in bold andVLRIKRPAAICTLIDPEE*underlineLanM-TEV-QtEnc MP / / / IKVDIAAFDPDKDGTIDLKEALAAGSAAFDKLDPD amino acid sequence KDGTLDAKELKGRVSEADLKKLDPDNDGTLDKKEYLAAV LanM shown in italics EQFKAANPDNDGTIDARELASPAGSALVNLIRGGSGGSE NLYFQSNKSQLYPDSPLTDQDFNQLDQTVIEAARRQLVG TEV underlined RRFIELYGPLGRGMQSVFNDIFMESHEAKMDFQGSFDTE Glycine-serine linker VESSRRVNYTIPMLYKDFVLYWRDLEQSKALDIPIDFSVA sequence in bold and ANAARDVAFLEDQMIFHGSKEFDIPGLMNVKGRLTHLIGN underline WYESGNAFQDIVEARNKLLEMNHNGPYALVLSPELYSLL HRVHKDTNVLEIEHVRELITAGVFQSPVLKGKSGVIVNTG RNNLDLAISEDFETAYLGEEGMNHPFRVYETVVLRIKRPA AICTLIDPEE*6xHis-mNeon-Qt CLP MHHHHHHGGSVSKGEEDNMASLPATHELHIFGSINGVDF amino acid sequence DMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGY His tag in bold GFHQYLPYPDGMSPFQAAMVDGSGYQVHRTMQFEDGA SLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSLTALinker in italics ADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTT CLP in underlineYTFAKPMAANYLKNQPMYVFRKTELKHSKTELNFKEWQK AFTDVMG M DELYKSGSGGSGKKKGFTVGSLIQ* 6xHis-mNeon-Tm CLP MHHHHHHGGSVSKGEEDNMASLPATHELHIFGSINGVDF amino acid sequence DMVGQGTGNPNDGYEELNLKSTKGDLQFSPWILVPHIGY His tag in bold GFHQYLPYPDGMSPFQAAMVDGSGYQVHRTMQFEDGA SLTVNYRYTYEGSHIKGEAQVKGTGFPADGPVMTNSLTALinker in italics ADWCRSKKTYPNDKTIISTFKWSYTTGNGKRYRSTARTT CLP in underline YTFAKPMAANYLKNQPMYVFRKTELKHSKTELNFKEWQK AFTDVMGMDELYKSGGNTGGDLGIRKL*6xHis-Lipoyl domain- MHHHHHHSGAFEFKLPDIGEGIHEGEIVKWFVKPGDEVN TEV Protease EDDVLCEVQNDKAVVEIPSPVKGKVLEILVPEGTVATVGQ TLITLDAPGYENMTSSGLVPRGSGESLFKGPRDYNPISSTI CHLTNESDGHTTSLYGIGFGPFIITNKHLFRRNNGTLLVQS LHGVFKVKNTTTLQQHLIDGRDMIIIRMPKDFPPFPQKLKF REPQREERICLVTTNFQTKSMSSMVSDTSCTFPSSDGIF WKHWIQTKDGQCGSPLVSTRDGFIVGIHSASNFTNTNNY FTSVPKNFMELLTNQEAQQWVSGWRLNADSVLWGGHK VFMVKPEEPFQPVKEANRAHE*MBP Nucleic Acid ATGAAAATCGAAGAAGGTAAACTGGTAATCTGGATTAA Sequence CGGCGATAAAGGCTATAACGGTCTCGCTGAAGTCGGT AAGAAATTCGAGAAAGATACCGGAATTAAAGTCACCGT TGAGCATCCGGATAAACTGGAAGAGAAATTCCCACAG GTTGCGGCAACTGGCGATGGCCCTGACATTATCTTCTG GGCACACGACCGCTTTGGTGGCTACGCTCAATCTGGC CTGTTGGCTGAAATCACCCCGGACAAAGCGTTCCAGG ACAAGCTGTATCCGTTTACCTGGGATGCCGTACGTTAC AACGGCAAGCTGATTGCTTACCCGATCGCTGTTGAAGC GTTATCGCTGATTTATAACAAAGATCTGCTGCCGAACC CGCCAAAAACCTGGGAAGAGATCCCGGCGCTGGATAA AGAACTGAAAGCGAAAGGTAAGAGCGCGCTGATGTTC AACCTGCAAGAACCGTACTTCACCTGGCCGCTGATTGC TGCTGACGGGGGTTATGCGTTCAAGTATGAAAACGGC AAGTACGACATTAAAGACGTGGGCGTGGATAACGCTG GCGCGAAAGCGGGTCTGACCTTCCTGGTTGACCTGAT TAAAAACAAACACATGAATGCAGACACCGATTACTCCA TCGCAGAAGCTGCCTTTAATAAAGGCGAAACAGCGATG ACCATCAACGGCCCGTGGGCATGGTCCAACATCGACA CCAGCAAAGTGAATTATGGTGTAACGGTACTGCCGACC TTCAAGGGTCAACCATCCAAACCGTTCGTTGGCGTGCT GAGCGCAGGTATTAACGCCGCCAGTCCGAACAAAGAGCTGGCAAAAGAGTTCCTCGAAAACTATCTGCTGACTGA TGAAGGTCTGGAAGCGGTTAATAAAGACAAACCGCTG GGTGCCGTAGCGCTGAAGTCTTACGAGGAAGAGTTGG CGAAAGATCCACGTATTGCCGCCACTATGGAAAACGC CCAGAAAGGTGAAATCATGCCGAACATCCCGCAGATG TCCGCTTTCTGGTATGCCGTGCGTACTGCGGTGATCAA CGCCGCCAGCGGTCGTCAGACTGTCGATGAAGCCCTG AAAGACGCGCAGACTAATTCGAGCTCGAACAACAACAA CAATAACAATAACAACAACCTCGGGATCGAGHis_MBP_TEV_QtEnc atg GGTTCTTCTCACCATCACCATCACCATGG TTCTTCT Nucleic acid sequence ATGAAAATCGAAGAAGGTAAACTGGTAATCTGGATTA ACGGCGATAAAGGCTATAACGGTCTCGCTGAAGTCG GTAAGAAATTCGAGAAAGATACCGGAATTAAAGTCACLinkers in italics CGTTGAGCATCCGGATAAACTGGAAGAGAAATTCCCA His tag underlined and CAGGTTGCGGCAACTGGCGATGGCCCTGACATTATCT italics TCTGGGCACACGACCGCTTTGGTGGCTACGCTCAATC TGGCCTGTTGGCTGAAATCACCCCGGACAAAGCGTTC MBP gene in boldCAGGACAAGCTGTATCCGTTTACCTGGGATGCCGTACProtease site in italics, GTTACAACGGCAAGCTGATTGCTTACCCGATCGCTGT bold, and underline TGAAGCGTTATCGCTGATTTATAACAAAGATCTGCTG QtEnc gene in bold CCGAACCCGCCAAAAACCTGGGAAGAGATCCCGGCG and italics CTGGATAAAGAACTGAAAGCGAAAGGTAAGAGCGCG CTGATGTTCAACCTGCAAGAACCGTACTTCACCTGGC CGCTGATTGCTGCTGACGGGGGTTATGCGTTCAAGTA TGAAAACGGCAAGTACGACATTAAAGACGTGGGCGT GGATAACGCTGGCGCGAAAGCGGGTCTGACCTTCCT GGTTGACCTGATTAAAAACAAACACATGAATGCAGAC ACCGATTACTCCATCGCAGAAGCTGCCTTTAATAAAG GCGAAACAGCGATGACCATCAACGGCCCGTGGGCAT GGTCCAACATCGACACCAGCAAAGTGAATTATGGTGT AACGGTACTGCCGACCTTCAAGGGTCAACCATCCAAA CCGTTCGTTGGCGTGCTGAGCGCAGGTATTAACGCCG CCAGTCCGAACAAAGAGCTGGCAAAAGAGTTCCTCG AAAACTATCTGCTGACTGATGAAGGTCTGGAAGCGGT TAATAAAGACAAACCGCTGGGTGCCGTAGCGCTGAA GTCTTACGAGGAAGAGTTGGCGAAAGATCCACGTATT GCCGCCACTATGGAAAACGCCCAGAAAGGTGAAATC ATGCCGAACATCCCGCAGATGTCCGCTTTCTGGTATG CCGTGCGTACTGCGGTGATCAACGCCGCCAGCGGTC GTCAGACTGTCGATGAAGCCCTGAAAGACGCGCAGACTAATTCGAGCTCGAACAACAACAACAATAACAATAA CAACAACCTCGGGATCGAGGAAAACCTGTACTTCCAA TCCAA TgcaggtggtggtggtAA CAAAA GCCAA CTTTA TCCG GA TTCA CCACTGACGGA TCA GGAC TTCAA CCA A TTA G ACCAAACCGTGATTGAGGCTGCTCGTCGTCAGCTGGT GGGTCGTCGCTTCATTGAGTTATATGGCCCATTGGGG CGTGGCA TGC A GA GTGTCTTCAACGA TATCTTCA TGG AGTCTCATGAA GCGAAAA TGGACTTCCA GGGCA GCTT TGACACGGAGGTAGAGTCCTCCCGTCGTGTAAACTAT ACCA TTCCGA TGTTA TA TAAA GA CTTCGTGCTTTA CTG GCGCGATCTGGAA CA GA GCAAGGCACTCGA TATTCC GATCGACTTTTCAGTGGCAGCGAACGCTGCCCGCGA CGTTGCGTTCCTGGAAGATCAGATGA I l l i CCATGGA AGCAAAGAATTTGATATCCCGGGTCTGATGAACGTGA AAGGTCGCCTGACCCATCTGATTGGCAATTGGTATGA GTCGGGTAACGCCTTTCAGGATATTGTGGAGGCCCGC AA TAAA TTA CTCGAAATGAACCA CAA TGGCCCA TA TG CTCTCGTGCTGTCCCCGGAGCTGTACTCACTCTTACAT CGTGTGCA TAAA GACA CGAA TGTGCTGGA GA TCGAA CACGTGCGCGAGTTGATTACTGCTGGGGTTTTTCAGT CGCCTGTCCTCAAAGGGAAAA GTGGTGTGA TCGTAAA CA CCGGTCGCAACAA TCTGGA TTTGGCTA TCTCGGAA GATTTTGAGACTGCATACCTGGGCGAGGAAGGTATGA ACCATCCCTTTCGCGTGTACGAGACAGTTGTTCTGCG CATCAAACGCCCGGCGGCCATTTGTACTTTAATCGAT CCGGAAGAATAA MBP Protein MKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEH Sequence PDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLLAEI TPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKD LLPNPPKTWEEIPALDKELKAKGKSALMFNLQEPYFTWPL IAADGGYAFKYENGKYDIKDVGVDNAGAKAGLTFLVDLIK NKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSK VNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEF LENYLLTDEGLEAVNKDKPLGAVALKSYEEELAKDPRIAA TMENAQKGEIMPNIPQMSAFWYAVRTAVINAASGRQTVD EALKDAQTNSSSNNNNNNNNNNLGIE* His_MBP_TEV_QtEnc MGSSHW7W7HGSSMKIEEGKLVIWINGDKGYNGLAEVG Protein sequence KKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWA Linkers in italics HDRFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKHis tag underlined and GKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKD italics VGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNK GETAMTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKP MBP in bold FVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKD Protease site in italics, KPLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQ bold, and underline MSAFWYAVRTAVINAASGRQTVDEALKDAQTNSSSNNN NNNNNNNLGIEEM. YFQS / VAGGGGJVKSQLYPDSPLTDQ QtEnc in bold andDFNQLDQTVIEAARRQLVGRRFIELYGPLGRGMQSVFNDitalicsIFMESHEAKMDFQGSFDTEVESSRRVNYTIPMLYKDFVL YWRDLEQSKALDIPIDFSVAANAARDVAFLEDQMIFHGS KEFDIPGLMNVKGRLTHLIGNWYESGNAFQDIVEARNKL LEMNHNGPYALVLSPELYSLLHRVHKDTNVLEIEHVRELI TA GVFQSPVLKGKSGVIVNTGRNNLDLAISEDFETA YLG EEGMNHPFRVYETWLRIKRPAAICTLIDPEE* MBP-TEV-QtEnc MKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVE amino acid sequence HPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLL Linkers in italics AEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLI YNKDLLPNPPKTWEEIPALDKELKAKGKSALMFNLQEP MBP in bold YFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKAGL Protease site in italics, TFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWA bold, and underline WSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAAS PNKELAKEFLENYLLTDEGLEAVNKDKPLGAVALKSYEEQtEnc in bold andELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVIitalicsNAASGRQTVDEALKDAQTNSSSNNNNNNNNNNLGIEEN LYFQSNAGGGGNKSQLYPDSPLTDQDFNQLDQTVIEAA RRQLVGRRFIELYGPLGRGMQSVFNDIFMESHEAKMDF QGSFDTEVESSRRVNYTIPMLYKDFVLYWRDLEQSKAL DIPIDFSVAANAARDVAFLEDQMIFHGSKEFDIPGLMNVK GRLTHLIGNWYESGNAFQDIVEARNKLLEMNHNGPYAL VLSPELYSLLHRVHKDTNVLEIEHVRELITAGVFQSPVLK GKSGVIVNTGRNNLDLAISEDFETAYLGEEGMNHPFRVY ETWLRIKRPAAICTLIDPEE*QtCLP KKKGFTVGSLIQC-modified QtCLP CKKKGFTVGSLIQ(conjugation to Aidox)QtCLP short TVGSLIQNucleic acid sequence ATGGAGTTCCTTAAGCGCTC I l l i GCTCCCCTGACAGA of the encapsulin from GAAGCAATGGCAAGAAATTGACAACCGTGCGCGTGAAATCTTCAAGACTCAGTTGTATGGACGCAAGTTTGTGGAThermotoga maritima TGTTGAGGGCCCTTACGGTTGGGAATATGCCGCTCAC (TmEnc) CCTTTGGGAGAAGTCGAAGTCCTTTCAGACGAGAACG AGGTAGTTAAGTGGGGACTTCGCAAATCTCTTCCTCTT ATCGAACTTCGTGCGACCTTTACACTGGACCTTTGGGA ATTGGACAATCTGGAGCGTGGCAAGCCCAATGTGGAT CTGTCCAGTTTGGAAGAGACTGTTCGCAAGGTGGCTG AATTTGAGGACGAAGTTATTTTTCGTGGGTGTGAAAAG AGCGGGGTTAAAGGGCTGCTGTC I l l i GAGGAACGTA AAATCGAGTGCGGTTCCACCCCAAAGGATTTACTTGAA GCAATTGTACGTGCCTTATCCA I l l i CAGCAAGGATGG CATTGAAGGGCCGTACACTTTAGTGATTAACACGGATC GTTGGATCAATTTCCTGAAAGAAGAGGCCGGGCATTAT CCCTTGGAGAAACGCGTAGAGGAGTGCTTACGTGGGG GGAAAATCATCACGACTCCCCGTATCGAGGATGCTTTG GTAGTGTCAGAGCGCGGCGGTGATTTCAAACTGATCTT GGGACAAGATTTATCAATCGGCTACGAAGACCGTGAG AAAGACGCTGTACGTCTTTTTATCACAGAGACATTTACA TTTCAGGTTGTTAATCCGGAGGCTCTTA 1 1 1 IACTGAAA TTCTAANucleic acid sequence ATGCCGGATTTCTTGGGACATGCTGAGAACCCATTGCG of the encapsulin from CGAAGAAGAGTGGGCCCGTTTAAACGAGACTGTAATC Myxococcus xanthus CAGGTTGCCCGCCGCTCTTTGGTTGGGCGTCGTATCT (MxEnc) TAGACATTTATGGCCCTTTGGGAGCGGGTGTGCAAAC GGTTCCATATGATGAGTTTCAGGGCGTAAGTCCGGGG GCCGTTGATATCGTGGGTGAGCAAGAAACCGCAATGG TGTTTACCGATGCCCGCAAGTTTAAGACCATCCCGATC ATTTATAAAGATTTCTTGCTGCATTGGCGTGATATTGAG GCGGCGCGTACCCATAACATGCCCTTAGATGTTTCGG CAGCCGCTGGGGCGGCGGCCTTATGCGCTCAGCAGG AAGACGAGTTAATCTTCTACGGCGACGCTCGCTTGGG ATACGAAGGGTTAATGACAGCTAATGGCCGTCTGACG GTTCCTTTAGGGGACTGGACCTCTCCGGGGGGTGGTT TTCAGGCAATCGTGGAAGCTACCCGCAAACTGAACGA GCAGGGACA 1 1 1 1 GG 1 CCC 1 ACGCCG 1 GG 1 1 1 1 ATCGC CTCGCTTGTATTCCCAACTTCACCGCATCTACGAAAAA ACAGGGGTGCTGGAAATTGAAACGATTCGTCAACTTGC CTCTGATGGTGTGTATCAATCAAACCGTCTGCGCGGG GAATCTGGTGTTGTGGTGTCCACAGGCCGTGAAAATAT GGACTTAGCGGTAAGCATGGACATGGTCGCTGCTTATT TAGGAGCTTCTCGTATGAATCACCC I l l i CGCGTACTGGAAGCATTGCTTCTGCGCATTAAGCACCCAGATGCGAT TTGTACATTAGAAGGAGCTGGAGCTACAGAGCGCCGC TAANucleic acid sequence ATGGACTTCCTGGATCGTAATGCGGCGCCGCTGAATG of the encapsulin from AGGGCGAATGGCAGCGTATCGACGAGGCTGTGGTTTC - Dendrosporobacter GACCGCACGCCGTACCCTGGTTGCACGTCGTATTATT quercicolus (DqEnc) GATG I 1 1 IAGGTCCGTTGGGTAGTGGTGTGTATTCTAT CCCGTACTCGGTC I l l i CTGGCAAAAGTCCTACGGGCA TTGATATGGTTGGGGAGAATGAGGAATTTGTGGTGGA GGCAAGCCGGCGTGCCACGATCAATCTGCCTATCCTG TATAAAGATTTCAAGATTATGTGGCGTGATGTGGAAGC CGACCGCCATTTGGGCTTACCGATCGATGTGTCAACC GCGGCTGTGGCAGGTAACTTTGTGGCCGTGCAAGAAG ATCATCTGATCTTTAACGGTAATGTAGAGTTGGGCCAC GATGGCTTGTTTACCGTCAAAGATCGTCAGACTGTTGC AATCAGCGATTGGAACGAAACCGGCGCCGCCCTGGCG GATGTTGTTAAAGCCGTTGGCGCGCTCAGTCAGGCTG GCCATTATGGCCCGTATGCGATGGTCGTGAGCCCGGT CCTGTTTGGTCGTATGATCCGCGTGTTTGGTAACACCG GTATGTTAGAACTGGACCAGGTGAAAGCCCTGATCACT GGCGGCGTTCATTATAGTAATGTGATCGGTGGCTCCAA AGCTGTGGTGGTGGCAACCGGTTCTCAGAATCTGAAT CTGGCCCTGGGGCAGGATATGGTAACGGCCTATATGG GTCCTACCTCCATGAATCATGTCTTTCGCGTTCTGGAA ACCGCCGCCCTGTTGGTACGTCGCCCGGATGCGATCT GCACGATTGAATAANucleic acid sequence ATGGGCAAGACACAGTTGTTTCCAGATTCACCATTGAC of the encapsulin from CGATCAAGATTTCTCGCAGCTGGACCAGACAGTGATTG - Bacillus methanolicus ACACCGCCCGCCGTCAACTCATCGGGCGTCG I 1 1 IATT (BmEnc) GAACTCTACGGCCCGTTGGGCCGCGGTATTCAGTCGA TTTTTAATGATGTATTCATTGAGAACTATGAAGCTAAAA TGGATTTCCAAGGTAG I l l i GATACGGACATCGAAACG AGCAAACGCGTGAATTACACCATTCCGTTATTATATAAG GACTTTGTTCTGTATTGGCGCGACTTGGAACAGGCGAA AGTCCTTGATATCCCGATTGAC I l l i CTGCGGCAGCTA ACGCCGCCCGTGATGTTGCGATCTTAGAGGATCAGAT GATC I l l i ATGGTAGTAAAGAGTTCGATATTCCGGGGT TGATGAATGTGAAAGGCCGTTCGACACACTTGATTGGT AATTGGTATGAATCCGGTAATGCCTTCCAGGACGTGGT GGAGGCTCGGAATAAATTGCTGGAAATGAAACATAACGGCCCGTTCGCGTTGGTCCTGTCGCCAGAATTGTACTCT CTGCTGCACCGTGTGCATAAGGATACCAATGTTCTGGA AATTGAGCATGTGCGCGAACTGGTCACAGATGGCGTTT TCCAAACACCGGTGCTGAAAGGTAAAACCGGCGTGCT GGTCAACACAGGTCGCAACAATCTGGACCTGGCTGTG TCGGAAGATTTCGATACTGCATACCTGGGTGAGGAAG GGATGAATCACCCGTTCCGCGTATACGAAACGGTCGT GCTGCGTATTAAACGCCCGAGTGCAATCTGTACCTTAG AGGATGGTGGTGAGTAAAmino acid sequence MEFLKRSFAPLTEKQWQEIDNRAREIFKTQLYGRKFVDVE of the encapsulin from GPYGWEYAAHPLGEVEVLSDENEVVKWGLRKSLPLIELR Thermotoga maritima ATFTLDLWELDNLERGKPNVDLSSLEETVRKVAEFEDEVI (TmEnc) FRGCEKSGVKGLLSFEERKIECGSTPKDLLEAIVRALSIFS KDGIEGPYTLVINTDRWINFLKEEAGHYPLEKRVEECLRG GKIITTPRIEDALVVSERGGDFKLILGQDLSIGYEDREKDA VRLFITETFTFQVVNPEALILLKF*Amino acid sequence MPDFLGHAENPLREEEWARLNETVIQVARRSLVGRRILDI of the encapsulin from YGPLGAGVQTVPYDEFQGVSPGAVDIVGEQETAMVFTD Myxococcus xanthus ARKFKTIPIIYKDFLLHWRDIEAARTHNMPLDVSAAAGAAA (MxEnc) LCAQQEDELIFYGDARLGYEGLMTANGRLTVPLGDWTSP GGGFQAIVEATRKLNEQGHFGPYAVVLSPRLYSQLHRIYE KTGVLEIETIRQLASDGVYQSNRLRGESGVVVSTGRENM DLAVSMDMVAAYLGASRMNHPFRVLEALLLRIKHPDAICT LEGAGATERR*Amino acid sequence MDFLDRNAAPLNEGEWQRIDEAVVSTARRTLVARRIIDVL of the encapsulin from GPLGSGVYSIPYSVFSGKSPTGIDMVGENEEFVVEASRR Dendrosporobacter ATINLPILYKDFKIMWRDVEADRHLGLPIDVSTAAVAGNFV quercicolus (DqEnc) AVQEDHLIFNGNVELGHDGLFTVKDRQTVAISDWNETGA ALADVVKAVGALSQAGHYGPYAMVVSPVLFGRMIRVFGN TGMLELDQVKALITGGVHYSNVIGGSKAVVVATGSQNLNL ALGQDMVTAYMGPTSMNHVFRVLETAALLVRRPDAICTIE*Amino acid sequence MGKTQLFPDSPLTDQDFSQLDQTVIDTARRQLIGRRFIEL of the encapsulin from YGPLGRGIQSIFNDVFIENYEAKMDFQGSFDTDIETSKRV Bacillus methanolicus NYTIPLLYKDFVLYWRDLEQAKVLDIPIDFSAAANAARDVA (BmEnc) ILEDQMIFYGSKEFDIPGLMNVKGRSTHLIGNWYESGNAF QDVVEARNKLLEMKHNGPFALVLSPELYSLLHRVHKDTN VLEIEHVRELVTDGVFQTPVLKGKTGVLVNTGRNNLDLAVSEDFDTAYLGEEGMNHPFRVYETVVLRIKRPSAICTLEDG GE*Nucleic acid sequence ATGcatcaccaccatcatcacaaaaaaagtCCGACAACGACCACC of the encapsulin from AAAGTTGACATCGCCGCATTTGATCCAGATAAAGACG Thermotoga maritima GAACGATCGACTTGAAAGAGGCACTGGCTGCTGGTTC fused to LanM GGCGGCATTTGATAAACTGGACCCGGATAAGGACGG (LanM_TmEnc) TACTCTTGATGCAAAAGAATTAAAAGGGCGTGTCTCC 6xHis tag shown in GAAGCTGACTTGAAGAAACTGGACCCTGACAACGAC underline GGCACATTGGACAAGAAGGAGTACCTTGCAGCTGTG GAGCAATTCAAAGCAGCGAACCCAGACAACGATGGCLanM in bold ACGATTGACGCTCGCGAACTGGCTTCCCCGGCAGGTT TEV protease site in CCGCCTTGGTGAATCTTATTCGGGGGGGCTCCGGTGG italics CTCGgagaatctttattttcagtctGAGTTCCTTAAGCGCTCTTTT GCTCCCCTGA CA GA GAA GCAA TGGCAA GAAATTGA CTmEnc gene in boldAA CCGTGCGCGTGAAATCTTCAA GACTCAGTTGTA TGand italicsGACGCAAGTTTGTGGATGTTGAGGGCCCTTACGGTTG GGAATATGCCGCTCACCCTTTGGGAGAAGTCGAAGTC CTTTCAGACGAGAACGAGGTAGTTAAGTGGGGACTTC GCAAA TCTCTTCCTCTTATCGAA CTTCGTGCGA CCTTT ACACTGGACCTTTGGGAA TTGGACAA TCTGGA GCGTG GCAAGCCCAATGTGGATCTGTCCA GTTTGGAA GA GA C TGTTCGCAAGGTGGCTGAATTTGAGGACGAAGTTATT TTTCGTGGGTGTGAAAAGAGCGGGGTTAAAGGGCTG CTGTCTTTTGA GGAA CGTAAAA TCGAGTGCGGTTCCA CCCCAAAGGATTTACTTGAAGCAATTGTACGTGCCTT ATCCATTTTCAGCAAGGATGGCATTGAAGGGCCGTAC ACTTTAGTGATTAACACGGATCGTTGGATCAATTTCCT GAAAGAAGAGGCCGGGCATTATCCCTTGGAGAAACG CGTA GAGGA GTGCTTA CGTGGGGGGAAAATCA TCA C GA CTCCCCGTA TCGAGGA TGCTTTGGTAGTGTCAGAG CGCGGCGGTGA TTTCAAACTGA TCTTGGGA CAAGATT TATCAATCGGCTACGAAGACCGTGAGAAAGACGCTGT ACGTCTTTTTATCACAGAGACATTTACATTTCAGGTTG TTAATCCGGAGGCTCTTATTTTACTGAAATTCTAANucleic acid sequence ATGcatcaccaccatcatcacqqaqqaaqtCCGACAACGACCACC of the encapsulin from AAAGTTGACATCGCCGCATTTGATCCAGATAAAGACG Myxococcus xanthus GAACGATCGACTTGAAAGAGGCACTGGCTGCTGGTTC fused to LanM GGCGGCATTTGATAAACTGGACCCGGATAAGGACGG (LanM_MxEnc) TACTCTTGATGCAAAAGAATTAAAAGGGCGTGTCTCCGAAGCTGACTTGAAGAAACTGGACCCTGACAACGAC6xHis tag shown in GGCACATTGGACAAGAAGGAGTACCTTGCAGCTGTG underline GAGCAATTCAAAGCAGCGAACCCAGACAACGATGGC ACGATTGACGCTCGCGAACTGGCTTCCCCGGCAGGTTLanM in bold CCGCCTTGGTGAATCTTATTCGGGGGGGCTCCGGTGG TEV protease site in CTCGgagaatctttattttcagtctCCGGATTTCTTGGGACATGCT italics GAGAACCCATTGCGCGAAGAAGAGTGGGCCCGTTTA AACGAGACTGTAATCCAGGTTGCCCGCCGCTCTTTGGMxEnc gene in boldTTGGGCGTCGTA TCTTAGACA TTTA TGGCCCTTTGGGAand italicsGCGGGTGTGCAAACGGTTCCATATGATGAGTTTCAGG GCGTAAGTCCGGGGGCCGTTGATATCGTGGGTGAGC AAGAAACCGCAATGGTGTTTACCGATGCCCGCAAGTT TAAGACCATCCCGATCATTTATAAAGATTTCTTGCTGC ATTGGCGTGATATTGAGGCGGCGCGTACCCATAACAT GCCCTTAGATGTTTCGGCAGCCGCTGGGGCGGCGGC CTTA TGCGCTCA GCA GGAA GA CGAGTTAA TCTTCTA C GGCGACGCTCGCTTGGGATACGAAGGGTTAATGACA GCTAATGGCCGTCTGACGGTTCCTTTAGGGGACTGGA CCTCTCCGGGGGGTGGTTTTCAGGCAATCGTGGAAGC TA CCCGCAAA CTGAA CGA GCAGGGA CATTTTGGTCCC TACGCCGTGGTTTTATCGCCTCGCTTGTATTCCCAACT TCACCGCATCTACGAAAAAACAGGGGTGCTGGAAATT GAAA CGA TTCGTCAACTTGCCTCTGATGGTGTGTA TC AATCAAACCGTCTGCGCGGGGAATCTGGTGTTGTGGT GTCCACAGGCCGTGAAAATATGGACTTAGCGGTAAG CA TGGA CA TGGTCGCTGCTTA TTTAGGAGCTTCTCGT ATGAATCACCCTTTTCGCGTACTGGAAGCATTGCTTCT GCGCA TTAAGCA CCCA GA TGCGA TTTGTA CA TTAGAA GGAGCTGGAGCTACAGAGCGCCGCTAANucleic acid sequence ATGcatcaccaccatcatcacqqaqqaaqtCCGACAACGACCACC of the encapsulin from AAAGTTGACATCGCCGCATTTGATCCAGATAAAGACG Dendrosporobacter GAACGATCGACTTGAAAGAGGCACTGGCTGCTGGTTC quercicolus fused to GGCGGCATTTGATAAACTGGACCCGGATAAGGACGG LanM (LanM_DqEnc) TACTCTTGATGCAAAAGAATTAAAAGGGCGTGTCTCC 6xHis tag shown in GAAGCTGACTTGAAGAAACTGGACCCTGACAACGAC underline GGCACATTGGACAAGAAGGAGTACCTTGCAGCTGTG GAGCAATTCAAAGCAGCGAACCCAGACAACGATGGCLanM in bold ACGATTGACGCTCGCGAACTGGCTTCCCCGGCAGGTT TEV protease site in CCGCCTTGGTGAATCTTATTCGGGGGGGCTCCGGTGG italics CTCGgagaatctttattttcagtctATGGACTTCCTGGATCGTAATGCGGCGCCGCTGAATGAGGGCGAATGGCAGCGTATCGACGAGGCTGTGGTTTCGACCGCACGCCGTACCCTG DqEnc gene in bold GTTGCACGTCGTATTATTGATGTTTTAGGTCCGTTGGGand italicsTAGTGGTGTGTATTCTATCCCGTACTCGGTCTTTTCTG GCAAAAGTCCTA CGGGCATTGA TA TGGTTGGGGA GA ATGAGGAATTTGTGGTGGAGGCAAGCCGGCGTGCCA CGATCAATCTGCCTATCCTGTATAAAGATTTCAAGATT ATGTGGCGTGATGTGGAAGCCGACCGCCATTTGGGCT TACCGATCGATGTGTCAACCGCGGCTGTGGCAGGTAA CTTTGTGGCCGTGCAAGAAGATCATCTGATCTTTAAC GGTAATGTAGAGTTGGGCCACGATGGCTTGTTTACCG TCAAA GA TCGTCAGACTGTTGCAATCA GCGA TTGGAA CGAAACCGGCGCCGCCCTGGCGGATGTTGTTAAAGC CGTTGGCGCGCTCAGTCAGGCTGGCCATTATGGCCC GTATGCGATGGTCGTGAGCCCGGTCCTGTTTGGTCGT ATGA TCCGCGTGTTTGGTAA CACCGGTA TGTTA GAAC TGGACCAGGTGAAAGCCCTGATCACTGGCGGCGTTC ATTATAGTAATGTGATCGGTGGCTCCAAAGCTGTGGT GGTGGCAACCGGTTCTCAGAATCTGAATCTGGCCCTG GGGCAGGATATGGTAACGGCCTATATGGGTCCTACCT CCATGAATCATGTCTTTCGCGTTCTGGAAACCGCCGC CCTGTTGGTACGTCGCCCGGATGCGATCTGCACGATT GAATAANucleic acid sequence ATGcatcaccaccatcatcacaaaqqaaatCCGACAACGACCACC of the encapsulin from AAAGTTGACATCGCCGCATTTGATCCAGATAAAGACG Bacillus methanolicus GAACGATCGACTTGAAAGAGGCACTGGCTGCTGGTTC fused to LanM GGCGGCATTTGATAAACTGGACCCGGATAAGGACGG (LanM_BmEnc) TACTCTTGATGCAAAAGAATTAAAAGGGCGTGTCTCC GAAGCTGACTTGAAGAAACTGGACCCTGACAACGAC6xHis tag shown inGGCACATTGGACAAGAAGGAGTACCTTGCAGCTGTGunderlineGAGCAATTCAAAGCAGCGAACCCAGACAACGATGGCLanM in bold ACGATTGACGCTCGCGAACTGGCTTCCCCGGCAGGTT TEV protease site in CCGCCTTGGTGAATCTTATTCGGGGGGGCTCCGGTGG italics CTCGgagaatctttattttcagtctATGGGCAAGACACAGTTGTTT CCAGATTCACCATTGACCGATCAAGATTTCTCGCAGCBmEnc gene in bold TGGACCAGACAGTGATTGACACCGCCCGCCGTCAACand italicsTCATCGGGCGTCG 1111 ATTGAACTCTACGGCCCGTT GGGCCGCGGTA TTCA GTCGATTTTTAA TGATGTA TTCA TTGAGAA CT A TGAAGCTAAAA TGGA TTTCCAAGGTA G TTTTGA TA CGGACATCGAAACGA GCAAA CGCGTGAA T TA CA CCA TTCCGTTA TTA TA TAA GGA CTTTGTTCTGTATTGGCGCGACTTGGAA CA GGCGAAAGTCCTTGA TATC CCGATTGACTTTTCTGCGGCAGCTAACGCCGCCCGTG ATGTTGCGATCTTAGAGGATCAGATGATC 111 IATGGT AGTAAA GAGTTCGATA TTCCGGGGTTGATGAA TGTGA AAGGCCGTTCGACACACTTGATTGGTAATTGGTATGA ATCCGGTAATGCCTTCCAGGACGTGGTGGAGGCTCG GAATAAATTGCTGGAAATGAAACATAACGGCCCGTTC GCGTTGGTCCTGTCGCCAGAATTGTACTCTCTGCTGC ACCGTGTGCATAAGGATACCAATGTTCTGGAAATTGA GCATGTGCGCGAACTGGTCACAGATGGCGTTTTCCAA ACACCGGTGCTGAAAGGTAAAACCGGCGTGCTGGTC AA CA CA GGTCGCAACAA TCTGGACCTGGCTGTGTCG GAAGATTTCGATACTGCATACCTGGGTGAGGAAGGG ATGAATCACCCGTTCCGCGTATACGAAACGGTCGTGC TGCGTATTAAACGCCCGAGTGCAATCTGTACCTTAGA GGA TGGTGGTGA GTAANucleic acid sequence ATGGGTTCTTCTCACCATCACCATCACCATGGTTCTTCT of the encapsulin from ATGAAAATCGAAGAAGGTAAACTGGTAATCTGGATTAA Thermotoga maritima CGGCGATAAAGGCTATAACGGTCTCGCTGAAGTCGGT fused to MBP AAGAAATTCGAGAAAGATACCGGAATTAAAGTCACCGT (MBP_TmEnc) TGAGCATCCGGATAAACTGGAAGAGAAATTCCCACAG 6xHis tag-MBP- GTTGCGGCAACTGGCGATGGCCCTGACATTATCTTCTG TEVsite in underline GGCACACGACCGCTTTGGTGGCTACGCTCAATCTGGC CTGTTGGCTGAAATCACCCCGGACAAAGCGTTCCAGGTmEnc in bold and ACAAGCTGTATCCGTTTACCTGGGATGCCGTACGTTAC italics AACGGCAAGCTGATTGCTTACCCGATCGCTGTTGAAGC GTTATCGCTGATTTATAACAAAGATCTGCTGCCGAACC CGCCAAAAACCTGGGAAGAGATCCCGGCGCTGGATAA AGAACTGAAAGCGAAAGGTAAGAGCGCGCTGATGTTC AACCTGCAAGAACCGTACTTCACCTGGCCGCTGATTGC TGCTGACGGGGGTTATGCGTTCAAGTATGAAAACGGC AAGTACGACATTAAAGACGTGGGCGTGGATAACGCTG GCGCGAAAGCGGGTCTGACCTTCCTGGTTGACCTGAT TAAAAACAAACACATGAATGCAGACACCGATTACTCCA TCGCAGAAGCTGCCTTTAATAAAGGCGAAACAGCGATG ACCATCAACGGCCCGTGGGCATGGTCCAACATCGACA CCAGCAAAGTGAATTATGGTGTAACGGTACTGCCGACC TTCAAGGGTCAACCATCCAAACCGTTCGTTGGCGTGCT GAGCGCAGGTATTAACGCCGCCAGTCCGAACAAAGAG CTGGCAAAAGAGTTCCTCGAAAACTATCTGCTGACTGATGAAGGTCTGGAAGCGGTTAATAAAGACAAACCGCTG GGTGCCGTAGCGCTGAAGTCTTACGAGGAAGAGTTGG CGAAAGATCCACGTATTGCCGCCACTATGGAAAACGC CCAGAAAGGTGAAATCATGCCGAACATCCCGCAGATG TCCGCTTTCTGGTATGCCGTGCGTACTGCGGTGATCAA CGCCGCCAGCGGTCGTCAGACTGTCGATGAAGCCCTG AAAGACGCGCAGACTAATTCGAGCTCGAACAACAACAA CAATAACAATAACAACAACCTCGGGATCGAGGAAAACC TGTACTTCCAATCCAATGCAGGTGGTGGTGGTGAGTTC CTTAA GCGCTCTTTTGCTCCCCTGA CAGA GAA GCAA T GGCAA GAAA TTGACAACCGTGCGCGTGAAA TCTTCAA GA CTCA GTTGTA TGGACGCAAGTTTGTGGA TGTTGA G GGCCCTTACGGTTGGGAATATGCCGCTCACCCTTTGG GA GAA GTCGAA GTCCTTTCAGA CGA GAA CGAGGTAG TTAAGTGGGGA CTTCGCAAA TCTCTTCCTCTTA TCGAA CTTCGTGCGACCTTTACACTGGACCTTTGGGAATTGG ACAA TCTGGA GCGTGGCAA GCCCAA TGTGGA TCTGTC CA GTTTGGAAGA GA CTGTTCGCAA GGTGGCTGAATTT GAGGA CGAAGTTA TTTTTCGTGGGTGTGAAAAGAGCG GGGTTAAAGGGCTGCTGTCTTTTGAGGAACGTAAAAT CGAGTGCGGTTCCACCCCAAAGGATTTACTTGAAGCA ATTGTACGTGCCTTATCCATTTTCAGCAAGGATGGCAT TGAAGGGCCGTACACTTTAGTGATTAACACGGATCGT TGGATCAATTTCCTGAAAGAAGAGGCCGGGCATTATC CCTTGGAGAAACGCGTAGAGGAGTGCTTACGTGGGG GGAAAA TCATCA CGA CTCCCCGTA TCGA GGATGCTTT GGTAGTGTCA GA GCGCGGCGGTGA TTTCAAACTGA TC TTGGGACAAGATTTATCAATCGGCTACGAAGACCGTG AGAAAGACGCTGTACGTCTTTTTATCACAGAGACATTT ACATTTCAGGTTGTTAA TCCGGAGGCTCTTA TTTTACT GAAATTCTAANucleic acid sequence ATGGGTTCTTCTCACCATCACCATCACCATGGTTCTTCT of the encapsulin from ATGAAAATCGAAGAAGGTAAACTGGTAATCTGGATTAA Myxococcus xanthus CGGCGATAAAGGCTATAACGGTCTCGCTGAAGTCGGT fused to MBP AAGAAATTCGAGAAAGATACCGGAATTAAAGTCACCGT (MBP_MxEnc) TGAGCATCCGGATAAACTGGAAGAGAAATTCCCACAG 6xHis tag-MBP- GTTGCGGCAACTGGCGATGGCCCTGACATTATCTTCTG TEVsite in underline GGCACACGACCGCTTTGGTGGCTACGCTCAATCTGGC CTGTTGGCTGAAATCACCCCGGACAAAGCGTTCCAGGMxEnc in bold and ACAAGCTGTATCCGTTTACCTGGGATGCCGTACGTTAC italicsAACGGCAAGCTGATTGCTTACCCGATCGCTGTTGAAGC GTTATCGCTGATTTATAACAAAGATCTGCTGCCGAACC CGCCAAAAACCTGGGAAGAGATCCCGGCGCTGGATAA AGAACTGAAAGCGAAAGGTAAGAGCGCGCTGATGTTC AACCTGCAAGAACCGTACTTCACCTGGCCGCTGATTGC TGCTGACGGGGGTTATGCGTTCAAGTATGAAAACGGC AAGTACGACATTAAAGACGTGGGCGTGGATAACGCTG GCGCGAAAGCGGGTCTGACCTTCCTGGTTGACCTGAT TAAAAACAAACACATGAATGCAGACACCGATTACTCCA TCGCAGAAGCTGCCTTTAATAAAGGCGAAACAGCGATG ACCATCAACGGCCCGTGGGCATGGTCCAACATCGACA CCAGCAAAGTGAATTATGGTGTAACGGTACTGCCGACC TTCAAGGGTCAACCATCCAAACCGTTCGTTGGCGTGCT GAGCGCAGGTATTAACGCCGCCAGTCCGAACAAAGAG CTGGCAAAAGAGTTCCTCGAAAACTATCTGCTGACTGA TGAAGGTCTGGAAGCGGTTAATAAAGACAAACCGCTG GGTGCCGTAGCGCTGAAGTCTTACGAGGAAGAGTTGG CGAAAGATCCACGTATTGCCGCCACTATGGAAAACGC CCAGAAAGGTGAAATCATGCCGAACATCCCGCAGATG TCCGCTTTCTGGTATGCCGTGCGTACTGCGGTGATCAA CGCCGCCAGCGGTCGTCAGACTGTCGATGAAGCCCTG AAAGACGCGCAGACTAATTCGAGCTCGAACAACAACAA CAATAACAATAACAACAACCTCGGGATCGAGGAAAACC TGTACTTCCAATCCAATGCAGGTGGTGGTGGTCCGGA TTTCTTGGGACATGCTGAGAACCCATTGCGCGAAGAA GAGTGGGCCCGTTTAAACGAGACTGTAATCCAGGTTG CCCGCCGCTCTTTGGTTGGGCGTCGTATCTTAGACAT TTATGGCCCTTTGGGAGCGGGTGTGCAAACGGTTCCA TATGATGAGTTTCAGGGCGTAAGTCCGGGGGCCGTTG ATA TCGTGGGTGAGCAAGAAA CCGCAA TGGTGTTTA C CGATGCCCGCAAGTTTAA GACCA TCCCGA TCA TTTA T AAAGATTTCTTGCTGCATTGGCGTGATATTGAGGCGG CGCGTA CCCA TAACA TGCCCTTA GA TGTTTCGGCA GC CGCTGGGGCGGCGGCCTTATGCGCTCAGCAGGAAGA CGAGTTAATCTTCTACGGCGACGCTCGCTTGGGATAC GAAGGGTTAATGACAGCTAATGGCCGTCTGACGGTTC CTTTAGGGGACTGGACCTCTCCGGGGGGTGGTTTTCA GGCAATCGTGGAAGCTACCCGCAAACTGAACGAGCA GGGACA I l l i GGICCCIACGCCGIGGI 11 IATCGCCT CGCTTGTATTCCCAACTTCACCGCATCTACGAAAAAA CA GGGGTGCTGGAAA TTGAAA CGA TTCGTCAA CTTGCCTCTGA TGGTGTGTA TCAA TCAAA CCGTCTGCGCGGG GAATCTGGTGTTGTGGTGTCCACAGGCCGTGAAAATA TGGACTTAGCGGTAAGCATGGACATGGTCGCTGCTTA TTTAGGAGCTTCTCGTATGAATCACCCTTTTCGCGTAC TGGAAGCATTGCTTCTGCGCATTAAGCACCCAGATGC GATTTGTACATTAGAAGGAGCTGGAGCTACAGAGCGC CGCTAANucleic acid sequence atgGGTTCTTCTCACCATCACCATCACCATGGTTCTTCT of the encapsulin from ATGAAAATCGAAGAAGGTAAACTGGTAATCTGGATTA Dendrosporobacter ACGGCGATAAAGGCTATAACGGTCTCGCTGAAGTCG quercicolus fused to GTAAGAAATTCGAGAAAGATACCGGAATTAAAGTCAC MBP (MBP_DqEnc) CGTTGAGCATCCGGATAAACTGGAAGAGAAATTCCCA 6xHis tag shown in CAGGTTGCGGCAACTGGCGATGGCCCTGACATTATCT underline TCTGGGCACACGACCGCTTTGGTGGCTACGCTCAATC TGGCCTGTTGGCTGAAATCACCCCGGACAAAGCGTTC MBP in bold CAGGACAAGCTGTATCCGTTTACCTGGGATGCCGTAC TEV protease site in GTTACAACGGCAAGCTGATTGCTTACCCGATCGCTGT italics TGAAGCGTTATCGCTGATTTATAACAAAGATCTGCTG CCGAACCCGCCAAAAACCTGGGAAGAGATCCCGGCGDqEnc gene in bold CTGGATAAAGAACTGAAAGCGAAAGGTAAGAGCGCGand italicsCTGATGTTCAACCTGCAAGAACCGTACTTCACCTGGC CGCTGATTGCTGCTGACGGGGGTTATGCGTTCAAGTA TGAAAACGGCAAGTACGACATTAAAGACGTGGGCGT GGATAACGCTGGCGCGAAAGCGGGTCTGACCTTCCT GGTTGACCTGATTAAAAACAAACACATGAATGCAGAC ACCGATTACTCCATCGCAGAAGCTGCCTTTAATAAAG GCGAAACAGCGATGACCATCAACGGCCCGTGGGCAT GGTCCAACATCGACACCAGCAAAGTGAATTATGGTGT AACGGTACTGCCGACCTTCAAGGGTCAACCATCCAAA CCGTTCGTTGGCGTGCTGAGCGCAGGTATTAACGCCG CCAGTCCGAACAAAGAGCTGGCAAAAGAGTTCCTCG AAAACTATCTGCTGACTGATGAAGGTCTGGAAGCGGT TAATAAAGACAAACCGCTGGGTGCCGTAGCGCTGAA GTCTTACGAGGAAGAGTTGGCGAAAGATCCACGTATT GCCGCCACTATGGAAAACGCCCAGAAAGGTGAAATC ATGCCGAACATCCCGCAGATGTCCGCTTTCTGGTATG CCGTGCGTACTGCGGTGATCAACGCCGCCAGCGGTC GTCAGACTGTCGATGAAGCCCTGAAAGACGCGCAGA CTAATTCGAGCTCGAACAACAACAACAATAACAATAA CAACAACCTCGGGATCGAGGAAAACCTGTACTTCCAATCCAATgcaggtggtggtggtATGGAC7TCCTGGATCGTAAT GCGGCGCCGCTGAATGAGGGCGAATGGCAGCGTATC GACGAGGCTGTGGTTTCGACCGCACGCCGTACCCTG GTTGCACGTCGTATTATTGATGTTTTAGGTCCGTTGGG TAGTGGTGTGTATTCTATCCCGTACTCGGTCTTTTCTG GCAAAAGTCCTA CGGGCATTGA TA TGGTTGGGGA GA ATGAGGAATTTGTGGTGGAGGCAAGCCGGCGTGCCA CGATCAATCTGCCTATCCTGTATAAAGATTTCAAGATT ATGTGGCGTGATGTGGAAGCCGACCGCCATTTGGGCT TACCGATCGATGTGTCAACCGCGGCTGTGGCAGGTAA CTTTGTGGCCGTGCAAGAAGATCATCTGATCTTTAAC GGTAATGTAGAGTTGGGCCACGATGGCTTGTTTACCG TCAAAGATCGTCAGACTGTTGCAATCAGCGATTGGAA CGAAACCGGCGCCGCCCTGGCGGATGTTGTTAAAGC CGTTGGCGCGCTCAGTCAGGCTGGCCATTATGGCCC GTATGCGATGGTCGTGAGCCCGGTCCTGTTTGGTCGT ATGA TCCGCGTGTTTGGTAA CACCGGTA TGTTA GAAC TGGACCAGGTGAAAGCCCTGATCACTGGCGGCGTTC ATTATAGTAATGTGATCGGTGGCTCCAAAGCTGTGGT GGTGGCAACCGGTTCTCAGAATCTGAATCTGGCCCTG GGGCAGGATATGGTAACGGCCTATATGGGTCCTACCT CCATGAATCATGTCTTTCGCGTTCTGGAAACCGCCGC CCTGTTGGTACGTCGCCCGGATGCGATCTGCACGATT GAATAANucleic acid sequence atqGGTTCTTCTCACCATCACCATCACCATGGTTCTTCT of the encapsulin from ATGAAAATCGAAGAAGGTAAACTGGTAATCTGGATTA Bacillus methanolicus ACGGCGATAAAGGCTATAACGGTCTCGCTGAAGTCG fused to MBP GTAAGAAATTCGAGAAAGATACCGGAATTAAAGTCAC (MBP_BmEnc) CGTTGAGCATCCGGATAAACTGGAAGAGAAATTCCCA 6xHis tag shown in CAGGTTGCGGCAACTGGCGATGGCCCTGACATTATCT underline TCTGGGCACACGACCGCTTTGGTGGCTACGCTCAATC TGGCCTGTTGGCTGAAATCACCCCGGACAAAGCGTTC MBP in bold CAGGACAAGCTGTATCCGTTTACCTGGGATGCCGTAC TEV protease site in GTTACAACGGCAAGCTGATTGCTTACCCGATCGCTGT italics TGAAGCGTTATCGCTGATTTATAACAAAGATCTGCTG CCGAACCCGCCAAAAACCTGGGAAGAGATCCCGGCGBmEnc gene in bold CTGGATAAAGAACTGAAAGCGAAAGGTAAGAGCGCGand italicsCTGATGTTCAACCTGCAAGAACCGTACTTCACCTGGC CGCTGATTGCTGCTGACGGGGGTTATGCGTTCAAGTA TGAAAACGGCAAGTACGACATTAAAGACGTGGGCGTGGATAACGCTGGCGCGAAAGCGGGTCTGACCTTCCT GGTTGACCTGATTAAAAACAAACACATGAATGCAGAC ACCGATTACTCCATCGCAGAAGCTGCCTTTAATAAAG GCGAAACAGCGATGACCATCAACGGCCCGTGGGCAT GGTCCAACATCGACACCAGCAAAGTGAATTATGGTGT AACGGTACTGCCGACCTTCAAGGGTCAACCATCCAAA CCGTTCGTTGGCGTGCTGAGCGCAGGTATTAACGCCG CCAGTCCGAACAAAGAGCTGGCAAAAGAGTTCCTCG AAAACTATCTGCTGACTGATGAAGGTCTGGAAGCGGT TAATAAAGACAAACCGCTGGGTGCCGTAGCGCTGAA GTCTTACGAGGAAGAGTTGGCGAAAGATCCACGTATT GCCGCCACTATGGAAAACGCCCAGAAAGGTGAAATC ATGCCGAACATCCCGCAGATGTCCGCTTTCTGGTATG CCGTGCGTACTGCGGTGATCAACGCCGCCAGCGGTC GTCAGACTGTCGATGAAGCCCTGAAAGACGCGCAGA CTAATTCGAGCTCGAACAACAACAACAATAACAATAA CAACAACCTCGGGATCGAGGAAAACCTGTACTTCCAA TCCAATgcaggtggtggtggtATGGGCAAGACACAG7TGT7T CCAGATTCACCATTGACCGATCAAGATTTCTCGCAGC TGGACCAGACAGTGATTGACACCGCCCGCCGTCAAC TCATCGGGCGTCG 1111 ATTGAACTCTACGGCCCGTT GGGCCGCGGTA TTCA GTCGATTTTTAA TGATGTA TTCA TTGAGAA CT A TGAAGCTAAAA TGGA TTTCCAAGGTA G TTTTGA TA CGGACATCGAAACGA GCAAA CGCGTGAA T TA CA CCA TTCCGTTA TTA TA TAAGGACTTTGTTCTGTA TTGGCGCGACTTGGAA CA GGCGAAAGTCCTTGA TATC CCGATTGACTTTTCTGCGGCAGCTAACGCCGCCCGTG ATGTTGCGATCTTAGAGGATCAGATGATC 111 IATGGT AGTAAA GAGTTCGATA TTCCGGGGTTGATGAA TGTGA AAGGCCGTTCGACACACTTGATTGGTAATTGGTATGA ATCCGGTAATGCCTTCCAGGACGTGGTGGAGGCTCG GAATAAATTGCTGGAAATGAAACATAACGGCCCGTTC GCGTTGGTCCTGTCGCCAGAATTGTACTCTCTGCTGC ACCGTGTGCATAAGGATACCAATGTTCTGGAAATTGA GCATGTGCGCGAACTGGTCACAGATGGCGTTTTCCAA ACACCGGTGCTGAAAGGTAAAACCGGCGTGCTGGTC AA CA CA GGTCGCAACAA TCTGGACCTGGCTGTGTCG GAAGATTTCGATACTGCATACCTGGGTGAGGAAGGG ATGAATCACCCGTTCCGCGTATACGAAACGGTCGTGC TGCGTATTAAACGCCCGAGTGCAATCTGTACCTTAGA GGA TGGTGGTGA GTAAAmino acid sequence MHHHHHHGGSPTTTTKVDIAAFDPDKDGTIDLKEALAAG of the encapsulin from SAAFDKLDPDKDGTLDAKELKGRVSEADLKKLDPDNDG Thermotoga maritima TLDKKEYLAAVEQFKAANPDNDGTIDARELASPAGSAL fused to LanM VN LI RGGSGGSENL YFQSEFLKRSFAPLTEKQWQEIDNR (LanM_TmEnc) AREIFKTQLYGRKFVDVEGPYGWEYAAHPLGEVEVLSD 6xHis tag shown in ENEWKWGLRKSLPLIELRATFTLDLWELDNLERGKPNV underline DLSSLEETVRKVAEFEDEVIFRGCEKSGVKGLLSFEERKI ECGSTPKDLLEAIVRALSIFSKDGIEGPYTLVINTDRWINFLanM in bold LKEEAGHYPLEKRVEECLRGGKIITTPRIEDALWSERGG TEV protease site in DFKLILGQDLSIGYEDREKDAVRLFITETFTFQWNPEALI italics LLKF*TmEnc in bold anditalicsAmino acid sequence MHHHHHHGGSPTTTTKVDIAAFDPDKDGTIDLKEALAAG of the encapsulin from SAAFDKLDPDKDGTLDAKELKGRVSEADLKKLDPDNDG Myxococcus xanthus TLDKKEYLAAVEQFKAANPDNDGTIDARELASPAGSAL fused to LanM VN LI RGGSGGSENL YFQSPDFLGHAENPLREEEWARLN (LanM_MxEnc) ETVIQVARRSLVGRRILDIYGPLGAGVQTVPYDEFQGVSP 6xHis tag shown in GAVDIVGEQETAMVFTDARKFKTIPIIYKDFLLHWRDIEA underline ARTHNMPLDVSAAAGAAALCAQQEDELIFYGDARLGYE GLMTANGRLTVPLGDWTSPGGGFQAIVEATRKLNEQGHLanM in bold FGPYAWLSPRLYSQLHRIYEKTGVLEIETIRQLASDGVY TEV protease site in QSNRLRGESGVWSTGRENMDLA VSMDMVAA YLGASR italics MNHPFRVLEALLLRIKHPDAICTLEGA GA TERR* MxEnc in bold anditalicsAmino acid sequence MHHHHHHGGSPTTTTKVDIAAFDPDKDGTIDLKEALAAG of the encapsulin from SAAFDKLDPDKDGTLDAKELKGRVSEADLKKLDPDNDG Dendrosporobacter TLDKKEYLAAVEQFKAANPDNDGTIDARELASPAGSAL quercicolus fused to VN LI RGGSGGSENL YFQSMDFLDRNAAPLNEGEWQRID LanM (LanM_DqEnc) EAWSTARRTLVARRIIDVLGPLGSGVYSIPYSVFSGKSP 6xHis tag shown in TGIDMVGENEEFWEASRRA TINLPIL YKDFKIMWRD VEA underline DRHLGLPIDVSTAAVAGNFVAVQEDHLIFNGNVELGHDG LFTVKDRQTVAISDWNETGAALADWKAVGALSQAGHYLanM in bold GPYAMWSPVLFGRMIRVFGNTGMLELDQVKALITGGVH TEV protease site in YSNVIGGSKA VWA TGSQNLNLALGQDMVTA YMGPTSM italics NHVFRVLETAALLVRRPDAICTIE*DqEnc in bold anditalicsAmino acid sequence MHHHHHHGGSPTTTTKVDIAAFDPDKDGTIDLKEALAAG of the encapsulin from SAAFDKLDPDKDGTLDAKELKGRVSEADLKKLDPDNDG Bacillus methanolicus TLDKKEYLAAVEQFKAANPDNDGTIDARELASPAGSAL fused to LanM VNLIRGGSGGSENLYFQSMGKTQLFPDSPLTDQDFSQL (LanM_BmEnc) DQTVIDTARRQLIGRRFIELYGPLGRGIQSIFNDVFIENYE 6xHis tag shown in AKMDFQGSFDTDIETSKRVNYTIPLLYKDFVLYWRDLEQ underline AKVLDIPIDFSAAANAARDVAILEDQMIFYGSKEFDIPGL MNVKGRSTHLIGNWYESGNAFQDWEARNKLLEMKHNLanM in bold GPFALVLSPELYSLLHRVHKDTNVLEIEHVRELVTDGVF TEV protease site in QTPVLKGKTGVLVNTGRNNLDLA VSEDFDTA YLGEEGM italics NHPFRVYETVVLRIKRPSAICTLEDGGE” BmEnc in bold anditalicsAmino acid sequence MGSSHHHHHHGSSMKIEEGKLVIWINGDKGYNGLAEVGK of the encapsulin from KFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHD Thermotoga maritima RFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIA fused to MBP YPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSA (MBP_TmEnc) LMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDN 6xHis tag-MBP- AGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMT TEVsite in underline INGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSA GINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVALTmEnc in bold and KSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAV italics RTAVINAASGRQTVDEALKDAQTNSSSNNNNNNNNNNL G\EENLYFQSNAGGGGEFLKRSFAPLTEKQWQEIDNRA REIFKTQLYGRKFVDVEGPYGWEYAAHPLGEVEVLSDE NEWKWGLRKSLPLIELRATFTLDLWELDNLERGKPNVD LSSLEETVRKVAEFEDEVIFRGCEKSGVKGLLSFEERKIE CGSTPKDLLEAIVRALSIFSKDGIEGPYTLVINTDRWINFL KEEAGHYPLEKRVEECLRGGKIITTPRIEDALWSERGGD FKLILGQDLSIGYEDREKDAVRLFITETFTFQWNPEALIL LKF*Amino acid sequence MGSSHHHHHHGSSMKIEEGKLVIWINGDKGYNGLAEVGK of the encapsulin from KFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHD Myxococcus xanthus RFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIA fused to MBP YPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSA (MBP_MxEnc) LMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDN 6xHis tag-MBP- AGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMT TEVsite in underline INGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVALMxEnc in bold and KSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAV italics RTAVINAASGRQTVDEALKDAQTNSSSNNNNNNNNNNL GIEENLYFQSNAGGGGPDFLGHAHVPLREEEW’ARLNET VIQVARRSLVGRRILDIYGPLGAGVQTVPYDEFQGVSPG AVDIVGEQETAMVFTDARKFKTIPIIYKDFLLHWRDIEAAR THNMPLDVSAAAGAAALCAQQEDELIFYGDARLGYEGL IVTTANGRLTVPLGDWTSPGGGFQAIVEATRKLNEQGHFG PYA WLSPRL YSQLHRIYEKTGVLEIETIRQLASDGVYQS NRLRGESGVWSTGRENMDLA VSMDMVAA YLGASRMN HPFRVLEALLLRIKHPDAICTLEGA GA TERR* Amino acid sequence MGSSHHHHHHGSSMKIEEGKLVIWINGDKGYNGLAEVG of the encapsulin from KKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWA Dendrosporobacter HDRFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNG quercicolus fused to KLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAK MBP (MBP_DqEnc) GKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKD 6xHis tag shown in VGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNK underline GETAMTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKP FVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKD MBP in bold KPLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQ TEV protease site in MSAFWYAVRTAVINAASGRQTVDEALKDAQTNSSSNNN italics NNNNNNNLGIEE / VLYFQSNAGGGG / WDFLDRNAAPLNEG EWQRIDEAWSTARRTLVARRIIDVLGPLGSGVYSIPYSVDqEnc in bold andFSGKSPTGIDMVGENEEFWEASRRATINLPILYKDFKIMitalicsWRDVEADRHLGLPIDVSTAAVAGNFVAVQEDHLIFNGNV ELGHDGLFTVKDRQTVAISDWNETGAALADWKAVGAL SQAGHYGPYAMWSPVLFGRMIRVFGNTGMLELDQVKA LITGGVHYSNVIGGSKAVWATGSQNLNLALGQDMVTAY MGPTSMNHVFRVLETAALLVRRPDAICTIE* Amino acid sequence MGSSHHHHHHGSSMKIEEGKLVIWINGDKGYNGLAEVG of the encapsulin from KKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWA Bacillus methanolicus HDRFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNG fused to MBP KLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAK (MBP_BmEnc) GKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKD 6xHis tag shown in VGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNK underline GETAMTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKP FVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKD MBP in bold KPLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQ TEV protease site in MSAFWYAVRTAVINAASGRQTVDEALKDAQTNSSSNNN italics NNNNNNNLGIEE / VLYFQSNAGGGG / WGKTQLFPDSPLTDQDFSQLDQTVIDTARRQLIGRRFIELYGPLGRGIQSIFNDVBmEnc in bold and FIENYEAKMDFQGSFDTDIETSKRVNYTIPLLYKDFVLYW italics RDLEQAKVLDIPIDFSAAANAARDVAILEDQMIFYGSKEF DIPGLMNVKGRSTHLIGNWYESGNAFQDWEARNKLLE MKHNGPFALVLSPELYSLLHRVHKDTNVLEIEHVRELVT DGVFQTPVLKGKTGVLVNTGRNNLDLAVSEDFDTAYLG EEGMNHPFRVYETWLRIKRPSAICTLEDGGE* TmCLP NTGGDLGIRKLC-modified TmCLP CNTGGDLGIRKLMxCLP PEKRLTVGSLRRC-modified MxCLP CPEKRLTVGSLRRDqCLP PPPLPDSEPDREIPADDGSLGIGSLKGTRSC-modified DqCLP CPPPLPDSEPDREIPADDGSLGIGSLKGTRS BmCLP KKKGFTVGSLIQC-modified BmCLP CKKKGFTVGSLIQNucleic acid sequence ATGAACAAAAGTCAACTATATCCCGATTCACCTTTGACC of Qt-Glass9 GATCAGGATTTCAACCAATTGGACCAAACCGTAATCGA encapsulin variant GGCGGCACGCCGTCAGCTGGTTGGTCGTCG I I I IATC Bold: sequence of GAGCTTTATGGTCCGCTGGGCCGTGGCATGCAAAGCG mutated region TTTTTAATGATATCTTCATGGAAAGCCATGAAGCGAAAA TGGATTTCCAGGGCTCGTTTGACACCGAAGTTGAGTCT AGCCGTCGTGTCAATTATACCATTCCGATGCTCTACAA GGACTTCGTGTTATACTGGCGCGATCTGGAGCAGAGC AAGGCTTTGGACATCCCGATTGACTTCTCTGTGGCGGC AAATGCGGCTCGTGACGTGGCTTTCCTGGAGGACCAG ATGATTTTTCATGGCTCCAAGGAGTTCGACATCCCAGG TCTGATGAACGTGAAAGGTCGCCTGACCCATCTGATTG GCAATTGGTACGAGAGCGGTAACGCCTTTCAAGATATT GTTGAGGCCCGTAACAAACTGCTTGAGATGAATCACAA CGGTCCGTACGCGCTGGTGTTGTCCCCGGAGCTGTAC AGCCTGCTCCACAGAGGCTCGGGTGGGTCCGGTCTG ATCACCGCGGGTGTGTTTCAGAGCCCGGTGCTGAAGG GCAAAAGCGGCGTCATCGTGAACACGGGTCGCAACAA CCTGGATCTGGCCATCAGTGAAGA I l l i GAAACCGCGT ATTTGGGCGAAGAAGGTATGAACCACCCGTTCCGTGTT TATGAAACGGTTG I l l i GCGCATTAAACGTCCGGCAGC GATTTGCACTCTGATCGACCCGGAGGAATAANucleic acid sequence ATGAATAAAAGTCAACTATATCCCGATTCACCATTGACC of Qt-Letter11 GATCAAGACTTCAACCAGCTTGACCAAACCGTTATTGA encapsulin variant AGCGGCGCGTCGTCAGCTGGTTGGTCGCCG I I I IATC Bold: sequence of GAGTTGTACGGCCCATTGGGCCGCGGTATGCAGAGCG mutated region TCTTCAATGATATCTTCATGGAAAGCCATGAAGCGAAG ATGGATTTCCAAGGTTC I l l i GATACCGAAGTGGAGTC TAGCCGTCGCGTGAATTACACCATTCCGATGCTGTACA AAGATTTCGTTCTCTACTGGCGTGATTTAGAGCAGAGC AAGGCATTGGACATCCCGATTGATTTCAGCGTGGCCG CCAATGCAGCTCGTGACGTTGCCTTTCTGGAGGACCA AATGA I l l i CCACGGCTCGAAAGAATTTGACATCCCGG GTCTGATGAACGTGAAAGGCCGCCTGACGCACCTGAT CGGCAACTGGTATGAGTCAGGAAACGCGTTTCAGGAC ATCGTGGAAGCTCGCAACAAGCTGTTGGAGATGAACC ACAACGGTCCGTATGCACTGGTGCTGAGCCCGGAATT GTACAGCCTCCTGGGCTCCGGTGGTTCCGGCCTGATT ACCGCGGGTGTTTTTCAGTCCCCGGTGCTCAAGGGCA AATCGGGCGTCATCGTCAACACCGGTAGAAATAACCTG GACCTGGCAATCAGCGAAGACTTCGAAACTGCGTATCT GGGTGAAGAGGGTATGAATCATCCGTTTCGTGTTTATG AGACGGTTGTGCTGCGTATTAAACGTCCGGCGGCGAT CTGCACCCTGATTGATCCGGAGGAGTAANucleic acid sequence ATGAATAAAAGTCAACTATATCCCGATTCACCACTGACT of Qt- Veg eta 10 GATCAAGATTTCAACCAGTTGGACCAAACCGTCATTGA encapsulin variant GGCGGCCCGTCGCCAGTTGGTCGGCCGCCG I I I IATC Bold: sequence of GAGCTGTACGGCCCTCTGGGCAGAGGCATGCAGTCAG mutated region TGTTCAATGATATTTTTATGGAGTCCCATGAAGCTAAGA TGGATTTCCAGGGTTCGTTCGACACCGAAGTTGAATCT AGCCGTCGTGTCAACTATACCATTCCGATGCTGTACAA GGACTTTG I l l i GTACTGGCGTGATTTGGAGCAGTCCA AGGCGTTGGACATCCCGATTGACTTCAGCGTGGCAGC TAATGCGGCTCGCGATGTTGCGTTTCTGGAGGATCAAA TGA I l l i CCACGGCAGCAAAGAATTTGACATCCCGGGT CTGATGAATGTTAAAGGTCGCCTGACCCATCTGATCGG CAACTGGTATGAAAGCGGCAACGCCTTCCAAGATATCG TAGAGGCGCGTAATAAACTGCTCGAGATGAACCACAAT GGTCCGTATGCCCTGGTTCTGAGTCCGGAACTCTACA GCCTGTTACACCGTGTGGGCTCCGGTGGTAGCGGCA CCGCGGGTGTTTTTCAGTCTCCGGTTCTTAAGGGTAAA AGCGGGGTGATCGTGAACACCGGTCGCAACAACCTGGACCTGGCAATTAGCGAAGACTTCGAGACGGCGTACCT GGGTGAAGAGGGAATGAACCATCCGTTTCGTGTGTAT GAGACGGTGGTGCTGCGTATTAAACGTCCGGCAGCGA TCTGCACCTTGATCGACCCGGAAGAGTAANucleic acid sequence ATGAATAAAAGTCAACTATATCCCGATTCACCACTGAC of Qt-Pigs13 GGATCAAGATTTCAACCAGTTGGACCAAACCGTGATCG encapsulin variant AGGCGGCTCGTCGTCAGCTGGTGGGTAGACG I I I I AT Bold: sequence of CGAGCTGTACGGTCCGCTGGGCCGCGGTATGCAGTCA mutated region GTTTTTAATGATATCTTTATGGAAAGCCACGAAGCGAA GATGGATTTCCAGGGTTCGTTCGATACCGAGGTGGAA AGCAGCCGTCGCGTTAACTATACCATTCCGATGTTGTA CAAAGATTTCGTGTTGTACTGGCGTGATCTGGAGCAGT CCAAGGCGCTCGACATCCCGATTGACTTCAGCGTGGC CGCTAATGCGGCTCGCGACGTCGCGTTTCTGGAGGAC CAAATGA I l l i CCATGGTTCCAAAGAATTTGACATCCCG GGTCTGATGAATGTTAAGGGTCGCTTGACGCATCTCAT CGGCAACTGGTATGAATCCGGCAATGCATTTCAGGACA TCGTGGAGGCGCGTAATAAGCTGCTGGAGATGAACCA CAACGGCCCTTACGCACTGGTGTTAAGCCCGGAGCTT TACCCGTACGGCTCTTGTATTGAGTTGATCACCGCGG GTGTGTTCCAATCTCCGGTTCTGAAAGGCAAAAGCGG CGTCATCGTCAACACCGGTCGCAACAACCTGGATCTG GCAATTAGCGAAGACTTCGAGACTGCCTATCTGGGCG AAGAGGGTATGAACCACCCGTTTCGTGTTTATGAAACC GTTGTTCTGCGTATTAAACGTCCGGCAGCGATTTGCAC CTTAATTGACCCGGAAGAATAANucleic acid sequence ATGAATAAAAGTCAACTATATCCCGATTCACCACTGAC ofQt-Slay13 CGATCAGGATTTCAACCAGCTGGACCAAACCGTGATC encapsulin variant GAAGCGGCTCGCCGCCAGCTGGTTGGTAGACG I I I IA Bold: sequence of TCGAACTGTACGGCCCGTTGGGCCGTGGTATGCAGTC mutated region CGTGTTTAACGATATCTTCATGGAAAGCCATGAGGCGA AAATGGA I l l i CAGGGTAGCTTTGATACCGAGGTTGAG TCTAGCCGTCGCGTGAACTACACCATTCCTATGCTGTA CAAGGACTTCGTGTTGTACTGGCGTGATCTGGAACAGA GCAAGGCTCTGGACATTCCGATTGACTTCAGTGTTGCA GCGAATGCTGCGCGCGATGTTGCCTTTCTAGAGGACC AAATGATCTTCCACGGCAGCAAAGAGTTCGACATCCCG GGTCTGATGAATGTCAAGGGCCGTCTGACCCACTTGAT TGGCAATTGGTATGAGTCCGGCAACGCCTTCCAAGAC ATCGTGGAAGCGCGCAACAAACTGCTGGAGATGAACCACAACGGTCCGTATGCACTGGTGTTAAGCCCGGAACT CTACAGCCTGGGTTCTGGCGGTTCGGGTATTACCGCG GGTGTGTTTCAATCCCCGGTTCTTAAGGGTAAAAGCG GCGTTATTGTCAATACCGGTCGTAACAACCTGGACCTC GCGATCAGCGAAGATTTCGAGACGGCGTATCTGGGCG AAGAAGGTATGAATCATCCGTTTCGTGTCTATGAGACG GTAGTTCTGCGTATTAAGCGTCCGGCAGCCATCTGCAC TTTGATCGACCCGGAAGAGTAAAmino acid sequence MNKSQLYPDSPLTDQDFNQLDQTVIEAARRQLVGRRFIEL of Qt-Glass9 YGPLGRGMQSVFNDIFMESHEAKMDFQGSFDTEVESSR encapsulin variant RVNYTIPMLYKDFVLYWRDLEQSKALDIPIDFSVAANAAR Bold: sequence of DVAFLEDQMIFHGSKEFDIPGLMNVKGRLTHLIGNWYES mutated region GNAFQDIVEARNKLLEMNHNGPYALVLSPELYSLLHRGS GGSGLITAGVFQSPVLKGKSGVIVNTGRNNLDLAISEDFE TAYLGEEGMNHPFRVYETVVLRIKRPAAICTLIDPEE* Amino acid sequence MNKSQLYPDSPLTDQDFNQLDQTVIEAARRQLVGRRFIEL of Qt-Letter11 YGPLGRGMQSVFNDIFMESHEAKMDFQGSFDTEVESSR encapsulin variant RVNYTIPMLYKDFVLYWRDLEQSKALDIPIDFSVAANAAR Bold: sequence of DVAFLEDQMIFHGSKEFDIPGLMNVKGRLTHLIGNWYES mutated region GNAFQDIVEARNKLLEMNHNGPYALVLSPELYSLLGSGG SGLITAGVFQSPVLKGKSGVIVNTGRNNLDLAISEDFETA YLGEEGMNHPFRVYETVVLRIKRPAAICTLIDPEE* Amino acid sequence MNKSQLYPDSPLTDQDFNQLDQTVIEAARRQLVGRRFIEL of Qt- Veg eta 10 YGPLGRGMQSVFNDIFMESHEAKMDFQGSFDTEVESSR encapsulin variant RVNYTIPMLYKDFVLYWRDLEQSKALDIPIDFSVAANAAR Bold: sequence of DVAFLEDQMIFHGSKEFDIPGLMNVKGRLTHLIGNWYES mutated region GNAFQDIVEARNKLLEMNHNGPYALVLSPELYSLLHRVG SGGSGTAGVFQSPVLKGKSGVIVNTGRNNLDLAISEDFE TAYLGEEGMNHPFRVYETVVLRIKRPAAICTLIDPEE* Amino acid sequence MNKSQLYPDSPLTDQDFNQLDQTVIEAARRQLVGRRFIEL of Qt-Pigs13 YGPLGRGMQSVFNDIFMESHEAKMDFQGSFDTEVESSR encapsulin variant RVNYTIPMLYKDFVLYWRDLEQSKALDIPIDFSVAANAAR Bold: sequence of DVAFLEDQMIFHGSKEFDIPGLMNVKGRLTHLIGNWYES mutated region GNAFQDIVEARNKLLEMNHNGPYALVLSPELYPYGSCIEL ITAGVFQSPVLKGKSGVIVNTGRNNLDLAISEDFETAYLGE EGMNHPFRVYETVVLRIKRPAAICTLIDPEE* Amino acid sequence MNKSQLYPDSPLTDQDFNQLDQTVIEAARRQLVGRRFIEL ofQt-Slay13 YGPLGRGMQSVFNDIFMESHEAKMDFQGSFDTEVESSR encapsulin variant RVNYTIPMLYKDFVLYWRDLEQSKALDIPIDFSVAANAARBold: sequence of DVAFLEDQMIFHGSKEFDIPGLMNVKGRLTHLIGNWYES mutated region GNAFQDIVEARNKLLEMNHNGPYALVLSPELYSLGSGGS GITAGVFQSPVLKGKSGVIVNTGRNNLDLAISEDFETAYL GEEGMNHPFRVYETVVLRIKRPAAICTLIDPEE* Nucleic acid sequence ATGGGTTCTTCTCACCATCACCATCACCATGGTTCTTCT of Qt-Glass9 ATGAAAATCGAAGAAGGTAAACTGGTAATCTGGATTAA encapsulin variant CGGCGATAAAGGCTATAACGGTCTCGCTGAAGTCGGT fused to MBP AAGAAATTCGAGAAAGATACCGGAATTAAAGTCACCGT (MBP_Qt-Glass9) TGAGCATCCGGATAAACTGGAAGAGAAATTCCCACAG 6xHis tag-MBP-TEV GTTGCGGCAACTGGCGATGGCCCTGACATTATCTTCTG site in underline GGCACACGACCGCTTTGGTGGCTACGCTCAATCTGGC CTGTTGGCTGAAATCACCCCGGACAAAGCGTTCCAGGQt-Glass9 Italics - ACAAGCTGTATCCGTTTACCTGGGATGCCGTACGTTAC Bold: sequence of AACGGCAAGCTGATTGCTTACCCGATCGCTGTTGAAGC mutated region GTTATCGCTGATTTATAACAAAGATCTGCTGCCGAACC CGCCAAAAACCTGGGAAGAGATCCCGGCGCTGGATAA AGAACTGAAAGCGAAAGGTAAGAGCGCGCTGATGTTC AACCTGCAAGAACCGTACTTCACCTGGCCGCTGATTGC TGCTGACGGGGGTTATGCGTTCAAGTATGAAAACGGC AAGTACGACATTAAAGACGTGGGCGTGGATAACGCTG GCGCGAAAGCGGGTCTGACCTTCCTGGTTGACCTGAT TAAAAACAAACACATGAATGCAGACACCGATTACTCCA TCGCAGAAGCTGCCTTTAATAAAGGCGAAACAGCGATG ACCATCAACGGCCCGTGGGCATGGTCCAACATCGACA CCAGCAAAGTGAATTATGGTGTAACGGTACTGCCGACC TTCAAGGGTCAACCATCCAAACCGTTCGTTGGCGTGCT GAGCGCAGGTATTAACGCCGCCAGTCCGAACAAAGAG CTGGCAAAAGAGTTCCTCGAAAACTATCTGCTGACTGA TGAAGGTCTGGAAGCGGTTAATAAAGACAAACCGCTG GGTGCCGTAGCGCTGAAGTCTTACGAGGAAGAGTTGG CGAAAGATCCACGTATTGCCGCCACTATGGAAAACGC CCAGAAAGGTGAAATCATGCCGAACATCCCGCAGATG TCCGCTTTCTGGTATGCCGTGCGTACTGCGGTGATCAA CGCCGCCAGCGGTCGTCAGACTGTCGATGAAGCCCTG AAAGACGCGCAGACTAATTCGAGCTCGAACAACAACAA CAATAACAATAACAACAACCTCGGGATCGAGGAAAACC TGTACTTCCAATCCAATGCAGGTGGTGGTGGTAACAAA AGTCAACTATATCCCGATTCACCTTTGACCGATCAGGA TTTCAACCAATTGGACCAAACCGTAATCGAGGCGGCAC GCCGTCAGCTGGTTGGTCGTCG 111 IATCGAGCTTTATGGTCCGCTGGGCCGTGGCATGCAAAGCG 1 111 IAAIG A TA TC TTCA TG GAAA GCCA TGAA G C GA AAA TG GA TTTC CAGGGCTCGTTTGACACCGAAGTTGAGTCTAGCCGTC GTGTCAATTATACCATTCCGATGCTCTACAAGGACTTC GTGTTATACTGGCGCGATCTGGAGCAGAGCAAGGCTT TGGA CA TCCCGA TTGACTTC TC TGTGGCGGCAAA TGC GGCTCGTGACGTGGCTTTCCTGGAGGACCAGATGATT TTTCATGGCTCCAAGGAGTTCGACATCCCAGGTCTGAT GAACGTGAAAGGTCGCCTGACCCATCTGATTGGCAATT GGTACGAGAGCGGTAACGCCTTTCAAGATATTGTTGAG GCCCGTAACAAACTGCTTGAGATGAATCACAACGGTCC GTACGCGCTGGTGTTGTCCCCGGAGCTGTACAGCCTG CTCCACAGAGGCTCGGGTGGGTCCGGTCTGATCACC GCGGGTGTGTTTCAGAGCCCGGTGCTGAAGGGCAAAA GCGGCGTCATCGTGAACACGGGTCGCAACAACCTGGA TCTGGCCATCAGTGAAGA 11 11 GAAACCGCGTATTTGG GCGAAGAAGGTATGAACCACCCGTTCCGTGTTTATGAA ACGGTTG 11 11 GCGCATTAAACGTCCGGCAGCGATTTG CACTCTGATCGACCCGGAGGAATAANucleic acid sequence ATGGGTTCTTCTCACCATCACCATCACCATGGTTCTTCT of Qt-Letter11 ATGAAAATCGAAGAAGGTAAACTGGTAATCTGGATTAA encapsulin variant CGGCGATAAAGGCTATAACGGTCTCGCTGAAGTCGGT fused to MBP AAGAAATTCGAGAAAGATACCGGAATTAAAGTCACCGT (MBP_Qt-Letter11) TGAGCATCCGGATAAACTGGAAGAGAAATTCCCACAG 6xHis tag-MBP-TEV GTTGCGGCAACTGGCGATGGCCCTGACATTATCTTCTG site in underline GGCACACGACCGCTTTGGTGGCTACGCTCAATCTGGC CTGTTGGCTGAAATCACCCCGGACAAAGCGTTCCAGGQt-Letter11 Italics — ACAAGCTGTATCCGTTTACCTGGGATGCCGTACGTTAC Bold: sequence of AACGGCAAGCTGATTGCTTACCCGATCGCTGTTGAAGC mutated region GTTATCGCTGATTTATAACAAAGATCTGCTGCCGAACC CGCCAAAAACCTGGGAAGAGATCCCGGCGCTGGATAA AGAACTGAAAGCGAAAGGTAAGAGCGCGCTGATGTTC AACCTGCAAGAACCGTACTTCACCTGGCCGCTGATTGC TGCTGACGGGGGTTATGCGTTCAAGTATGAAAACGGC AAGTACGACATTAAAGACGTGGGCGTGGATAACGCTG GCGCGAAAGCGGGTCTGACCTTCCTGGTTGACCTGAT TAAAAACAAACACATGAATGCAGACACCGATTACTCCA TCGCAGAAGCTGCCTTTAATAAAGGCGAAACAGCGATG ACCATCAACGGCCCGTGGGCATGGTCCAACATCGACA CCAGCAAAGTGAATTATGGTGTAACGGTACTGCCGACCTTCAAGGGTCAACCATCCAAACCGTTCGTTGGCGTGCT GAGCGCAGGTATTAACGCCGCCAGTCCGAACAAAGAG CTGGCAAAAGAGTTCCTCGAAAACTATCTGCTGACTGA TGAAGGTCTGGAAGCGGTTAATAAAGACAAACCGCTG GGTGCCGTAGCGCTGAAGTCTTACGAGGAAGAGTTGG CGAAAGATCCACGTATTGCCGCCACTATGGAAAACGC CCAGAAAGGTGAAATCATGCCGAACATCCCGCAGATG TCCGCTTTCTGGTATGCCGTGCGTACTGCGGTGATCAA CGCCGCCAGCGGTCGTCAGACTGTCGATGAAGCCCTG AAAGACGCGCAGACTAATTCGAGCTCGAACAACAACAA CAATAACAATAACAACAACCTCGGGATCGAGGAAAACC TGTACTTCCAATCCAATGCAGGTGGTGGTGGTAA TA A A AGTCAACTATATCCCGATTCACCATTGACCGATCAAGA CTTCAACCAGCTTGACCAAACCGTTATTGAAGCGGCGC GTCGTCAGCTGGTTGGTCGCCG 111 IATCGAGTTGTAC GGCCCATTGGGCCGCGGTATGCAGAGCGTCTTCAATG A TA TC TTCA TG GAAA GCCA TGAA GCGAAGA TGGA TTTC CAAGGTTCI 111 GATACCGAAGTGGAGTCTAGCCGTCG CGTGAATTACACCATTCCGATGCTGTACAAAGATTTCG TTCTCTACTGGCGTGATTTAGAGCAGAGCAAGGCATTG GACATCCCGATTGATTTCAGCGTGGCCGCCAATGCAG CTCGTGACGTTGCCTTTCTGGAGGACCAAATGA 11 1 IC CACGGC TCGAAA GAA TTTGACA TCCCGGGTC TGA TGAA CGTGAAAGGCCGCCTGACGCACCTGATCGGCAACTGG TATGAGTCAGGAAACGCGTTTCAGGACATCGTGGAAG CTCGCAACAAGCTGTTGGAGATGAACCACAACGGTCC GTA TGCACTGGTGCTGA GCCCGGAA TTGTA CAGCCTC CTGGGCTCCGGTGGTTCCGGCCTGATTACCGCGGGT GTTTTTCAGTCCCCGGTGCTCAAGGGCAAATCGGGCG TCATCGTCAACACCGGTAGAAATAACCTGGACCTGGCA ATCAGCGAAGACTTCGAAACTGCGTATCTGGGTGAAGA GGGTATGAATCATCCGTTTCGTGTTTATGAGACGGTTG TGCTGCGTATTAAACGTCCGGCGGCGATCTGCACCCT GATTGATCCGGAGGAGTAANucleic acid sequence ATGGGTTCTTCTCACCATCACCATCACCATGGTTCTTCT of Qt- Veg eta 10 ATGAAAATCGAAGAAGGTAAACTGGTAATCTGGATTAA encapsulin variant CGGCGATAAAGGCTATAACGGTCTCGCTGAAGTCGGT fused to MBP AAGAAATTCGAGAAAGATACCGGAATTAAAGTCACCGT (MBP_Qt- Vegetal 0) TGAGCATCCGGATAAACTGGAAGAGAAATTCCCACAGGTTGCGGCAACTGGCGATGGCCCTGACATTATCTTCTG6xHis tag-MBP-TEV GGCACACGACCGCTTTGGTGGCTACGCTCAATCTGGC site in underline CTGTTGGCTGAAATCACCCCGGACAAAGCGTTCCAGG ACAAGCTGTATCCGTTTACCTGGGATGCCGTACGTTACQt-Vegeta10 Italics - AACGGCAAGCTGATTGCTTACCCGATCGCTGTTGAAGC Bold: sequence of GTTATCGCTGATTTATAACAAAGATCTGCTGCCGAACC mutated region CGCCAAAAACCTGGGAAGAGATCCCGGCGCTGGATAA AGAACTGAAAGCGAAAGGTAAGAGCGCGCTGATGTTC AACCTGCAAGAACCGTACTTCACCTGGCCGCTGATTGC TGCTGACGGGGGTTATGCGTTCAAGTATGAAAACGGC AAGTACGACATTAAAGACGTGGGCGTGGATAACGCTG GCGCGAAAGCGGGTCTGACCTTCCTGGTTGACCTGAT TAAAAACAAACACATGAATGCAGACACCGATTACTCCA TCGCAGAAGCTGCCTTTAATAAAGGCGAAACAGCGATG ACCATCAACGGCCCGTGGGCATGGTCCAACATCGACA CCAGCAAAGTGAATTATGGTGTAACGGTACTGCCGACC TTCAAGGGTCAACCATCCAAACCGTTCGTTGGCGTGCT GAGCGCAGGTATTAACGCCGCCAGTCCGAACAAAGAG CTGGCAAAAGAGTTCCTCGAAAACTATCTGCTGACTGA TGAAGGTCTGGAAGCGGTTAATAAAGACAAACCGCTG GGTGCCGTAGCGCTGAAGTCTTACGAGGAAGAGTTGG CGAAAGATCCACGTATTGCCGCCACTATGGAAAACGC CCAGAAAGGTGAAATCATGCCGAACATCCCGCAGATG TCCGCTTTCTGGTATGCCGTGCGTACTGCGGTGATCAA CGCCGCCAGCGGTCGTCAGACTGTCGATGAAGCCCTG AAAGACGCGCAGACTAATTCGAGCTCGAACAACAACAA CAATAACAATAACAACAACCTCGGGATCGAGGAAAACC TGTACTTCCAATCCAATGCAGGTGGTGGTGGTAA TA A A AGTCAACTATATCCCGATTCACCACTGACTGATCAAGA TTTCAACCAGTTGGACCAAACCGTCATTGAGGCGGCC CGTCGCCAGTTGGTCGGCCGCCG 111 IATCGAGCTGT ACGGCCCTCTGGGCAGAGGCATGCAGTCAGTGTTCAA TGA TA 1111 IA TGGAGTCCCA TGAAGCTAA GA TGGA TTT CCAGGGTTCGTTCGACACCGAAGTTGAATCTAGCCGT CGTGTCAACTATACCATTCCGATGCTGTACAAGGACTT TGI I I IGTACTGGCGTGATTTGGAGCAGTCCAAGGCGT TGGACATCCCGATTGACTTCAGCGTGGCAGCTAATGC GGCTCGCGA TGTTGCGTTTC TGGAGGA TCAAA TGA TTT TCCACGGCAGCAAAGAATTTGACATCCCGGGTCTGAT GAATGTTAAAGGTCGCCTGACCCATCTGATCGGCAACT GGTATGAAAGCGGCAACGCCTTCCAAGATATCGTAGA GGCGCGTAATAAACTGCTCGAGATGAACCACAATGGTCCGTATGCCCTGGTTCTGAGTCCGGAACTCTACAGCC TGTTACACCGTGTGGGCTCCGGTGGTAGCGGCACCG CGGGTGTTTTTCA GTC TCCGGTTC TTAAGGGTAAAA GC GGGGTGATCGTGAACACCGGTCGCAACAACCTGGACC TGGCAATTAGCGAAGACTTCGAGACGGCGTACCTGGG TGAAGAGGGAATGAACCATCCGTTTCGTGTGTATGAGA CGGTGGTGCTGCGTATTAAACGTCCGGCAGCGATCTG CACCTTGATCGACCCGGAAGAGTAANucleic acid sequence ATGGGTTCTTCTCACCATCACCATCACCATGGTTCTTCT of Qt-Pigs13 ATGAAAATCGAAGAAGGTAAACTGGTAATCTGGATTAA encapsulin variant CGGCGATAAAGGCTATAACGGTCTCGCTGAAGTCGGT fused to MBP AAGAAATTCGAGAAAGATACCGGAATTAAAGTCACCGT (MBP_Qt-Pigs13) TGAGCATCCGGATAAACTGGAAGAGAAATTCCCACAG 6xHis tag-MBP-TEV GTTGCGGCAACTGGCGATGGCCCTGACATTATCTTCTG site in underline GGCACACGACCGCTTTGGTGGCTACGCTCAATCTGGC CTGTTGGCTGAAATCACCCCGGACAAAGCGTTCCAGGQt-Pigs13 Italics - ACAAGCTGTATCCGTTTACCTGGGATGCCGTACGTTAC Bold: sequence of AACGGCAAGCTGATTGCTTACCCGATCGCTGTTGAAGC mutated region GTTATCGCTGATTTATAACAAAGATCTGCTGCCGAACC CGCCAAAAACCTGGGAAGAGATCCCGGCGCTGGATAA AGAACTGAAAGCGAAAGGTAAGAGCGCGCTGATGTTC AACCTGCAAGAACCGTACTTCACCTGGCCGCTGATTGC TGCTGACGGGGGTTATGCGTTCAAGTATGAAAACGGC AAGTACGACATTAAAGACGTGGGCGTGGATAACGCTG GCGCGAAAGCGGGTCTGACCTTCCTGGTTGACCTGAT TAAAAACAAACACATGAATGCAGACACCGATTACTCCA TCGCAGAAGCTGCCTTTAATAAAGGCGAAACAGCGATG ACCATCAACGGCCCGTGGGCATGGTCCAACATCGACA CCAGCAAAGTGAATTATGGTGTAACGGTACTGCCGACC TTCAAGGGTCAACCATCCAAACCGTTCGTTGGCGTGCT GAGCGCAGGTATTAACGCCGCCAGTCCGAACAAAGAG CTGGCAAAAGAGTTCCTCGAAAACTATCTGCTGACTGA TGAAGGTCTGGAAGCGGTTAATAAAGACAAACCGCTG GGTGCCGTAGCGCTGAAGTCTTACGAGGAAGAGTTGG CGAAAGATCCACGTATTGCCGCCACTATGGAAAACGC CCAGAAAGGTGAAATCATGCCGAACATCCCGCAGATG TCCGCTTTCTGGTATGCCGTGCGTACTGCGGTGATCAA CGCCGCCAGCGGTCGTCAGACTGTCGATGAAGCCCTG AAAGACGCGCAGACTAATTCGAGCTCGAACAACAACAA CAATAACAATAACAACAACCTCGGGATCGAGGAAAACCTGTACTTCCAATCCAATGCAGGTGGTGGTGGTAA TAAA AGTCAACTATATCCCGATTCACCACTGACGGATCAAGA TTTCAACCAGTTGGACCAAACCGTGATCGAGGCGGCT CGTCGTCAGCTGGTGGGTAGACG 111 IATCGAGCTGTA CGGTCCGCTGGGCCGCGGTATGCAGTCAGTTTTTAAT GATATCTTTATGGAAAGCCACGAAGCGAAGATGGATTT CCAGGGTTCGTTCGATACCGAGGTGGAAAGCAGCCGT CGCGTTAACTATACCATTCCGATGTTGTACAAAGATTTC GTGTTGTACTGGCGTGATCTGGAGCAGTCCAAGGCGC TCGACATCCCGATTGACTTCAGCGTGGCCGCTAATGC GGCTCGCGACGTCGCGTTTCTGGAGGACCAAATGATT TTCCA TGGTTCCAAAGAA TTTGACA TCCCGGGTCTGA T GAATGTTAAGGGTCGCTTGACGCATCTCATCGGCAACT GGTATGAATCCGGCAATGCATTTCAGGACATCGTGGA GGCGCGTAATAAGCTGCTGGAGATGAACCACAACGGC CCTTACGCACTGGTGTTAAGCCCGGAGCTTTACCCGT ACGGCTCTTGTATTGAGTTGATCACCGCGGGTGTGTT CCAATCTCCGGTTCTGAAAGGCAAAAGCGGCGTCATC GTCAACACCGGTCGCAACAACCTGGATCTGGCAATTA GCGAAGACTTCGAGACTGCCTATCTGGGCGAAGAGGG TATGAACCACCCGTTTCGTGTTTATGAAACCGTTGTTCT GCGTATTAAACGTCCGGCAGCGATTTGCACCTTAATTG ACCCGGAAGAATAANucleic acid sequence ATGGGTTCTTCTCACCATCACCATCACCATGGTTCTTCT ofQt-Slay13 ATGAAAATCGAAGAAGGTAAACTGGTAATCTGGATTAA encapsulin variant CGGCGATAAAGGCTATAACGGTCTCGCTGAAGTCGGT fused to MBP AAGAAATTCGAGAAAGATACCGGAATTAAAGTCACCGT (MBP_Qt-Slay13) TGAGCATCCGGATAAACTGGAAGAGAAATTCCCACAG 6xHis tag-MBP-TEV GTTGCGGCAACTGGCGATGGCCCTGACATTATCTTCTG site in underline GGCACACGACCGCTTTGGTGGCTACGCTCAATCTGGC CTGTTGGCTGAAATCACCCCGGACAAAGCGTTCCAGGQt-Slay13 Italics - ACAAGCTGTATCCGTTTACCTGGGATGCCGTACGTTAC Bold: sequence of AACGGCAAGCTGATTGCTTACCCGATCGCTGTTGAAGC mutated region GTTATCGCTGATTTATAACAAAGATCTGCTGCCGAACC CGCCAAAAACCTGGGAAGAGATCCCGGCGCTGGATAA AGAACTGAAAGCGAAAGGTAAGAGCGCGCTGATGTTC AACCTGCAAGAACCGTACTTCACCTGGCCGCTGATTGC TGCTGACGGGGGTTATGCGTTCAAGTATGAAAACGGC AAGTACGACATTAAAGACGTGGGCGTGGATAACGCTG GCGCGAAAGCGGGTCTGACCTTCCTGGTTGACCTGATTAAAAACAAACACATGAATGCAGACACCGATTACTCCA TCGCAGAAGCTGCCTTTAATAAAGGCGAAACAGCGATG ACCATCAACGGCCCGTGGGCATGGTCCAACATCGACA CCAGCAAAGTGAATTATGGTGTAACGGTACTGCCGACC TTCAAGGGTCAACCATCCAAACCGTTCGTTGGCGTGCT GAGCGCAGGTATTAACGCCGCCAGTCCGAACAAAGAG CTGGCAAAAGAGTTCCTCGAAAACTATCTGCTGACTGA TGAAGGTCTGGAAGCGGTTAATAAAGACAAACCGCTG GGTGCCGTAGCGCTGAAGTCTTACGAGGAAGAGTTGG CGAAAGATCCACGTATTGCCGCCACTATGGAAAACGC CCAGAAAGGTGAAATCATGCCGAACATCCCGCAGATG TCCGCTTTCTGGTATGCCGTGCGTACTGCGGTGATCAA CGCCGCCAGCGGTCGTCAGACTGTCGATGAAGCCCTG AAAGACGCGCAGACTAATTCGAGCTCGAACAACAACAA CAATAACAATAACAACAACCTCGGGATCGAGGAAAACC TGTACTTCCAATCCAATGCAGGTGGTGGTGGTAA TA A A AGTCAACTATATCCCGATTCACCACTGACCGATCAGGA TTTCAACCAGCTGGACCAAACCGTGATCGAAGCGGCT CGCCGCCAGCTGGTTGGTAGACG 1 11 IATCGAACTGTA CGGCCCGTTGGGCCGTGGTATGCAGTCCGTGTTTAAC GATATCTTCATGGAAAGCCATGAGGCGAAAATGGATTT TCAGGGTAGCTTTGATACCGAGGTTGAGTCTAGCCGTC GCGTGAACTACACCATTCCTATGCTGTACAAGGACTTC GTGTTGTAC TGGCGTGA TCTGGAA CA GAGCAAGGC TC TGGACATTCCGATTGACTTCAGTGTTGCAGCGAATGCT GCGCGCGA TGTTGCCTTTC TAG A GGACCAAA TGA TC TT CCACGGCAGCAAAGAGTTCGACATCCCGGGTCTGATG AATGTCAAGGGCCGTCTGACCCACTTGATTGGCAATTG GTATGAGTCCGGCAACGCCTTCCAAGACATCGTGGAA GCGCGCAACAAACTGCTGGAGATGAACCACAACGGTC CGTATGCACTGGTGTTAAGCCCGGAACTCTACAGCCT GGGTTCTGGCGGTTCGGGTATTACCGCGGGTGTGTTT CAATCCCCGGTTCTTAAGGGTAAAAGCGGCGTTATTGT CAATACCGGTCGTAACAACCTGGACCTCGCGATCAGC GAAGATTTCGAGACGGCGTATCTGGGCGAAGAAGGTA TGAATCATCCGTTTCGTGTCTATGAGACGGTAGTTCTG CGTA TTAAGCGTCCGGCAGCCA TCTGCACTTTGA TCGA CCCGGAAGAGTAAAmino acid sequence MGSSHHHHHHGSSMKIEEGKLVIWINGDKGYNGLAEVGK of Qt-Glass9 KFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDencapsulin variant RFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIA fused to MBP YPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSA (MBP_Qt-Glass9) LMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDN AGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMT INGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSA GINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVAL KSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAV RTAVINAASGRQTVDEALKDAQTNSSSNNNNNNNNNNL GIEENLYFQSNAGGGGNKSQLYPDSPLTDQDFNQLDQTV IEAARRQLVGRRFIELYGPLGRGMQSVFNDIFMESHEAK MDFQGSFDTEVESSRRVNYTIPMLYKDFVLYWRDLEQSK ALDIPIDFSVAANAARDVAFLEDQMIFHGSKEFDIPGLMNV KGRLTHLIGNWYESGNAFQDIVEARNKLLEMNHNGPYAL VLSPELYSLLHRGSGGSGLITAGVFQSPVLKGKSGVIVNT GRNNLDLAISEDFETAYLGEEGMNHPFRVYETVVLRIKRP AAICTLIDPEE*Amino acid sequence MGSSHHHHHHGSSMKIEEGKLVIWINGDKGYNGLAEVGK of Qt-Letter11 KFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHD encapsulin variant RFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIA fused to MBP YPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSA (MBP_Qt-Letter11) LMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDN AGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMT INGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSA GINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVAL KSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAV RTAVINAASGRQTVDEALKDAQTNSSSNNNNNNNNNNL GIEENLYFQSNAGGGGNKSQLYPDSPLTDQDFNQLDQTV IEAARRQLVGRRFIELYGPLGRGMQSVFNDIFMESHEAK MDFQGSFDTEVESSRRVNYTIPMLYKDFVLYWRDLEQSK ALDIPIDFSVAANAARDVAFLEDQMIFHGSKEFDIPGLMNV KGRLTHLIGNWYESGNAFQDIVEARNKLLEMNHNGPYAL VLSPELYSLLGSGGSGLITAGVFQSPVLKGKSGVIVNTGR NNLDLAISEDFETAYLGEEGMNHPFRVYETVVLRIKRPAAI CTLIDPEE*Amino acid sequence MGSSHHHHHHGSSMKIEEGKLVIWINGDKGYNGLAEVGK of Qt- Veg eta 10 KFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHD encapsulin variant RFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIA fused to MBP YPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSA (MBP_Qt-Vegeta10) LMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSA GINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVAL KSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAV RTAVINAASGRQTVDEALKDAQTNSSSNNNNNNNNNNL GIEENLYFQSNAGGGGNKSQLYPDSPLTDQDFNQLDQTV IEAARRQLVGRRFIELYGPLGRGMQSVFNDIFMESHEAK MDFQGSFDTEVESSRRVNYTIPMLYKDFVLYWRDLEQSK ALDIPIDFSVAANAARDVAFLEDQMIFHGSKEFDIPGLMNV KGRLTHLIGNWYESGNAFQDIVEARNKLLEMNHNGPYAL VLSPELYSLLHRVGSGGSGTAGVFQSPVLKGKSGVIVNT GRNNLDLAISEDFETAYLGEEGMNHPFRVYETVVLRIKRP AAICTLIDPEE*Amino acid sequence MGSSHHHHHHGSSMKIEEGKLVIWINGDKGYNGLAEVGK of Qt-Pigs13 KFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHD encapsulin variant RFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIA fused to MBP YPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSA (MBP_Qt-Pigs13) LMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDN AGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMT INGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSA GINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVAL KSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAV RTAVINAASGRQTVDEALKDAQTNSSSNNNNNNNNNNL GIEENLYFQSNAGGGGNKSQLYPDSPLTDQDFNQLDQTV IEAARRQLVGRRFIELYGPLGRGMQSVFNDIFMESHEAK MDFQGSFDTEVESSRRVNYTIPMLYKDFVLYWRDLEQSK ALDIPIDFSVAANAARDVAFLEDQMIFHGSKEFDIPGLMNV KGRLTHLIGNWYESGNAFQDIVEARNKLLEMNHNGPYAL VLSPELYPYGSCIELITAGVFQSPVLKGKSGVIVNTGRNNL DLAISEDFETAYLGEEGMNHPFRVYETVVLRIKRPAAICTL IDPEE*Amino acid sequence MGSSHHHHHHGSSMKIEEGKLVIWINGDKGYNGLAEVGK ofQt-Slay13 KFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHD encapsulin variant RFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIA fused to MBP YPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSA (MBP_Qt-Slay13) LMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDN AGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMT INGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSA GINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVAL KSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAV RTAVINAASGRQTVDEALKDAQTNSSSNNNNNNNNNNLGIEENLYFQSNAGGGGNKSQLYPDSPLTDQDFNQLDQTV IEAARRQLVGRRFIELYGPLGRGMQSVFNDIFMESHEAK MDFQGSFDTEVESSRRVNYTIPMLYKDFVLYWRDLEQSK ALDIPIDFSVAANAARDVAFLEDQMIFHGSKEFDIPGLMNV KGRLTHLIGNWYESGNAFQDIVEARNKLLEMNHNGPYAL VLSPELYSLGSGGSGITAGVFQSPVLKGKSGVIVNTGRNN LDLAISEDFETAYLGEEGMNHPFRVYETVVLRIKRPAAICT LIDPEE*Nucleic acid sequence ATGCATCACCACCATCATCACGGAGGAAGTCCGACAA of Qt-Glass9 CGACCACCAAAGTTGACATCGCCGCATTTGATCCAGAT encapsulin variant AAAGACGGAACGATCGACTTGAAAGAGGCACTGGCTG fused to LanM CTGGTTCGGCGGCATTTGATAAACTGGACCCGGATAA (LanM_Qt-Glass9) GGACGGTACTCTTGATGCAAAAGAATTAAAAGGGCGTG 6xHis tag-LanM-TEV TCTCCGAAGCTGACTTGAAGAAACTGGACCCTGACAAC site in underline GACGGCACATTGGACAAGAAGGAGTACCTTGCAGCTG TGGAGCAATTCAAAGCAGCGAACCCAGACAACGATGGQt-variant Italics - CACGATTGACGCTCGCGAACTGGCTTCCCCGGCAGGT Bold: sequence of TCCGCCTTGGTGAATCTTATTCGGGGGGGCTCCGGTG mutated region GCTCGGAGAATCTTTA 1 1 1 1 CAGTCTAACAAAAGTCAAC TATATCCCGATTCACCTTTGACCGATCAGGATTTCAAC CAATTGGACCAAACCGTAATCGAGGCGGCACGCCGTC AGCTGGTTGGTCGTCG 1 111 ATCGAGCTTTATGGTCCG CTGGGCCGTGGCATGCAAAGCGTTTTTAATGATATCTT CATGGAAAGCCATGAAGCGAAAATGGATTTCCAGGGC TCGTTTGACACCGAAGTTGAGTCTAGCCGTCGTGTCAA TTATACCATTCCGATGCTCTACAAGGACTTCGTGTTATA CTGGCGCGA TCTGGAGCA GA GCAA GGCTTTGGACA TC CCGATTGACTTCTCTGTGGCGGCAAATGCGGCTCGTG ACGTGGCTTTCCTGGAGGACCAGATGATTTTTCATGGC TCCAAGGAGTTCGACATCCCAGGTCTGATGAACGTGAA AGGTCGCCTGACCCATCTGATTGGCAATTGGTACGAG AGCGGTAACGCCTTTCAAGATATTGTTGAGGCCCGTAA CAAACTGCTTGAGATGAATCACAACGGTCCGTACGCG CTGGTGTTGTCCCCGGAGCTGTA CA GCCTGCTCCA CA GAGGCTCGGGTGGGTCCGGTCTGATCACCGCGGGTG TGTTTCAGAGCCCGGTGCTGAAGGGCAAAAGCGGCGT CATCGTGAACACGGGTCGCAACAACCTGGATCTGGCC ATCAGTGAAGA 11 11 GAAACCGCGTATTTGGGCGAAGA AGGTATGAACCACCCGTTCCGTGTTTATGAAACGGTTG / / / IGCGCATTAAACGTCCGGCAGCGATTTGCACTCTG ATCGACCCGGAGGAATAANucleic acid sequence ATGCATCACCACCATCATCACGGAGGAAGTCCGACAA of Qt-Letter11 CGACCACCAAAGTTGACATCGCCGCATTTGATCCAGAT encapsulin variant AAAGACGGAACGATCGACTTGAAAGAGGCACTGGCTG fused to LanM CTGGTTCGGCGGCATTTGATAAACTGGACCCGGATAA (LanM_Qt-Letter11) GGACGGTACTCTTGATGCAAAAGAATTAAAAGGGCGTG 6xHis tag-LanM-TEV TCTCCGAAGCTGACTTGAAGAAACTGGACCCTGACAAC site in underline GACGGCACATTGGACAAGAAGGAGTACCTTGCAGCTG TGGAGCAATTCAAAGCAGCGAACCCAGACAACGATGGQt-variant Italics - CACGATTGACGCTCGCGAACTGGCTTCCCCGGCAGGT Bold: sequence of TCCGCCTTGGTGAATCTTATTCGGGGGGGCTCCGGTG mutated region GCTCGGAGAATCTTTA I I I I CAGTCTAA TAAAAGTCAAC TATATCCCGATTCACCATTGACCGATCAAGACTTCAAC CAGCTTGACCAAACCGTTATTGAAGCGGCGCGTCGTC AGCTGGTTGGTCGCCG 111 IATCGAGTTGTACGGCCCA TTGGGCCGCGGTATGCAGAGCGTCTTCAATGATATCTT CATGGAAAGCCATGAAGCGAAGATGGATTTCCAAGGTT CH I IGATACCGAAGTGGAGTCTAGCCGTCGCGTGAAT TACACCATTCCGATGCTGTACAAAGATTTCGTTCTCTAC TGGCGTGATTTAGAGCAGAGCAAGGCATTGGACATCC CGATTGATTTCAGCGTGGCCGCCAATGCAGCTCGTGA CGTTGCCTTTCTGGAGGACCAAATGA 111 ICCACGGCT CGAAAGAATTTGACATCCCGGGTCTGATGAACGTGAAA GGCCGCCTGACGCACCTGATCGGCAACTGGTATGAGT CAGGAAACGCGTTTCAGGACATCGTGGAAGCTCGCAA CAAGCTGTTGGAGATGAACCACAACGGTCCGTATGCA CTGGTGCTGAGCCCGGAATTGTACAGCCTCCTGGGCT CCGGTGGTTCCGGCCTGATTACCGCGGGTGTTTTTCA GTCCCCGGTGCTCAAGGGCAAATCGGGCGTCATCGTC AACACCGGTAGAAATAACCTGGACCTGGCAATCAGCG AAGACTTCGAAACTGCGTATCTGGGTGAAGAGGGTAT GAATCATCCGTTTCGTGTTTATGAGACGGTTGTGCTGC GTATTAAACGTCCGGCGGCGATCTGCACCCTGATTGAT CCGGAGGAGTAANucleic acid sequence ATGCATCACCACCATCATCACGGAGGAAGTCCGACAA of Qt- Veg eta 10 CGACCACCAAAGTTGACATCGCCGCATTTGATCCAGAT encapsulin variant AAAGACGGAACGATCGACTTGAAAGAGGCACTGGCTG fused to LanM CTGGTTCGGCGGCATTTGATAAACTGGACCCGGATAA (LanM_Qt-Vegeta10) GGACGGTACTCTTGATGCAAAAGAATTAAAAGGGCGTG6xHis tag-LanM-TEV TCTCCGAAGCTGACTTGAAGAAACTGGACCCTGACAAC site in underline GACGGCACATTGGACAAGAAGGAGTACCTTGCAGCTG TGGAGCAATTCAAAGCAGCGAACCCAGACAACGATGGQt-variant Italics - CACGATTGACGCTCGCGAACTGGCTTCCCCGGCAGGT Bold: sequence of TCCGCCTTGGTGAATCTTATTCGGGGGGGCTCCGGTG mutated region GCTCGGAGAATCTTTA I I I I CAGTCTAA TAAAAGTCAAC TATATCCCGATTCACCACTGACTGATCAAGATTTCAACC AGTTGGACCAAACCGTCATTGAGGCGGCCCGTCGCCA GTTGGTCGGCCGCCG 1111 ATCGAGCTGTACGGCCCT CTGGGCAGAGGCATGCAGTCAGTGTTCAATGATA I l l i TATGGAGTCCCATGAAGCTAAGATGGATTTCCAGGGTT CGTTCGACACCGAAGTTGAATCTAGCCGTCGTGTCAAC TATACCATTCCGATGCTGTACAAGGACTTTG 1 11 IGTAC TGGCGTGATTTGGAGCAGTCCAAGGCGTTGGACATCC CGATTGACTTCAGCGTGGCAGCTAATGCGGCTCGCGA TGTTGCGTTTCTGGAGGATCAAATGA 111 ICCACGGCA GCAAAGAATTTGACATCCCGGGTCTGATGAATGTTAAA GGTCGCCTGACCCATCTGATCGGCAACTGGTATGAAA GCGGCAACGCCTTCCAAGATATCGTAGAGGCGCGTAA TAAACTGCTCGAGATGAACCACAATGGTCCGTATGCCC TGGTTCTGA GTCCGGAACTCTACAGCCTGTTA CA CCG TGTGGGCTCCGGTGGTAGCGGCACCGCGGGTGTTTTT CAGTCTCCGGTTCTTAAGGGTAAAAGCGGGGTGATCG TGAA CACCGGTCGCAA CAA CC TGGA CC TGGCAA TTA G CGAAGACTTCGAGACGGCGTACCTGGGTGAAGAGGGA ATGAACCATCCGTTTCGTGTGTATGAGACGGTGGTGCT GCGTA TTAAACGTCCGGCAGCGA TC TGCACCTTGA TC GACCCGGAAGAGTAANucleic acid sequence ATGCATCACCACCATCATCACGGAGGAAGTCCGACAA of Qt-Pigs13 CGACCACCAAAGTTGACATCGCCGCATTTGATCCAGAT encapsulin variant AAAGACGGAACGATCGACTTGAAAGAGGCACTGGCTG fused to LanM CTGGTTCGGCGGCATTTGATAAACTGGACCCGGATAA (LanM_Qt-Pigs13) GGACGGTACTCTTGATGCAAAAGAATTAAAAGGGCGTG 6xHis tag-LanM-TEV TCTCCGAAGCTGACTTGAAGAAACTGGACCCTGACAAC site in underline GACGGCACATTGGACAAGAAGGAGTACCTTGCAGCTG TGGAGCAATTCAAAGCAGCGAACCCAGACAACGATGGQt-variant Italics - CACGATTGACGCTCGCGAACTGGCTTCCCCGGCAGGT Bold: sequence of TCCGCCTTGGTGAATCTTATTCGGGGGGGCTCCGGTG mutated region GCTCGGAGAATCTTTA 1 1 1 1 CAGTCTAA TAAAAGTCAAC TA TATCCC GA TTCA CCAC TGA CGGA TCAA GA TTTCAA CCAGTTGGACCAAACCGTGATCGAGGCGGCTCGTCGTC AGCTGGTGGGTAGACG 111 IATCGAGCTGTACGGTCC GCTGGGCCGCGGTA TGCAGTCAG 1111 IAA TGA TA TC T TTATGGAAAGCCACGAAGCGAAGATGGATTTCCAGGG TTCGTTCGATACCGAGGTGGAAAGCAGCCGTCGCGTT AACTATACCATTCCGATGTTGTACAAAGATTTCGTGTTG TACTGGCGTGATCTGGAGCAGTCCAAGGCGCTCGACA TCCCGATTGACTTCAGCGTGGCCGCTAATGCGGCTCG CGACGTCGCGTTTCTGGAGGACCAAATGA 111 ICCATG GTTCCAAAGAA TTTGACA TCCCGGGTCTGA TGAATGTT AAGGGTCGCTTGACGCATCTCATCGGCAACTGGTATG AATCCGGCAATGCATTTCAGGACATCGTGGAGGCGCG TAATAAGCTGCTGGAGATGAACCACAACGGCCCTTACG CACTGGTGTTAAGCCCGGAGCTTTACCCGTACGGCTC TTGTA TTGAGTTGA TCACCGCGGGTGTGTTCCAA TC TC CGGTTCTGAAAGGCAAAAGCGGCGTCATCGTCAACAC CGGTCGCAACAACCTGGATCTGGCAATTAGCGAAGAC TTCGAGACTGCCTATCTGGGCGAAGAGGGTATGAACC ACCCGTTTCGTGTTTATGAAACCGTTGTTCTGCGTATTA AACGTCCGGCAGCGATTTGCACCTTAATTGACCCGGAA GAATAANucleic acid sequence ATGCATCACCACCATCATCACGGAGGAAGTCCGACAA ofQt-Slay13 CGACCACCAAAGTTGACATCGCCGCATTTGATCCAGAT encapsulin variant AAAGACGGAACGATCGACTTGAAAGAGGCACTGGCTG fused to LanM CTGGTTCGGCGGCATTTGATAAACTGGACCCGGATAA (LanM_Qt-Slay13) GGACGGTACTCTTGATGCAAAAGAATTAAAAGGGCGTG 6xHis tag-LanM-TEV TCTCCGAAGCTGACTTGAAGAAACTGGACCCTGACAAC site in underline GACGGCACATTGGACAAGAAGGAGTACCTTGCAGCTG TGGAGCAATTCAAAGCAGCGAACCCAGACAACGATGGQt-variant Italics - CACGATTGACGCTCGCGAACTGGCTTCCCCGGCAGGT Bold: sequence of TCCGCCTTGGTGAATCTTATTCGGGGGGGCTCCGGTG mutated region GCTCGGAGAATCTTTA 1 1 1 1 CAGTCTAA TAAAAGTCAAC TATATCCCGATTCACCACTGACCGATCAGGATTTCAAC CAGCTGGACCAAACCGTGATCGAAGCGGCTCGCCGCC AGCTGGTTGGTAGACG 111 IATCGAACTGTACGGCCCG TTGGGCCGTGGTATGCAGTCCGTGTTTAACGATATCTT CATGGAAAGCCATGAGGCGAAAATGGA 111 ICAGGGTA GCTTTGATACCGAGGTTGAGTCTAGCCGTCGCGTGAA CTACACCATTCCTATGCTGTACAAGGACTTCGTGTTGT ACTGGCGTGATCTGGAACAGAGCAAGGCTCTGGACATTCCGATTGACTTCAGTGTTGCAGCGAATGCTGCGCGC GATGTTGCCTTTCTAGAGGACCAAATGATCTTCCACGG CAGCAAAGAGTTCGACATCCCGGGTCTGATGAATGTCA AGGGCCGTCTGACCCACTTGATTGGCAATTGGTATGA GTCCGGCAACGCCTTCCAAGACATCGTGGAAGCGCGC AACAAACTGCTGGAGATGAACCACAACGGTCCGTATG CACTGGTGTTAAGCCCGGAACTCTACAGCCTGGGTTC TGGCGGTTCGGGTATTACCGCGGGTGTGTTTCAATCC CCGGTTCTTAAGGGTAAAAGCGGCGTTATTGTCAATAC CGGTCGTAACAACCTGGACCTCGCGATCAGCGAAGAT TTCGAGACGGCGTATCTGGGCGAAGAAGGTATGAATC ATCCGTTTCGTGTCTATGAGACGGTAGTTCTGCGTATT AAGCGTCCGGCAGCCATCTGCACTTTGATCGACCCGG AAGAGTAAAmino acid sequence MHHHHHHGGSP I I I I KVDIAAFDPDKDGTIDLKEALAAG of Qt-Glass9 SAAFDKLDPDKDGTLDAKELKGRVSEADLKKLDPDNDGT encapsulin variant LDKKEYLAAVEQFKAANPDNDGTIDARELASPAGSALVNL fused to LanM IRGGSGGSENLYFQSNKSQLYPDSPLTDQDFNQLDQTVI (LanM_Qt-Glass9) EAARRQLVGRRFIELYGPLGRGMQSVFNDIFMESHEAKM DFQGSFDTEVESSRRVNYTIPMLYKDFVLYWRDLEQSKA LDIPIDFSVAANAARDVAFLEDQMIFHGSKEFDIPGLMNVK GRLTHLIGNWYESGNAFQDIVEARNKLLEMNHNGPYALV LSPELYSLLHRGSGGSGLITAGVFQSPVLKGKSGVIVNTG RNNLDLAISEDFETAYLGEEGMNHPFRVYETVVLRIKRPA AICTLIDPEE*Amino acid sequence MHHHHHHGGSP I I I I KVDIAAFDPDKDGTIDLKEALAAG of Qt-Letter11 SAAFDKLDPDKDGTLDAKELKGRVSEADLKKLDPDNDGT encapsulin variant LDKKEYLAAVEQFKAANPDNDGTIDARELASPAGSALVNL fused to LanM IRGGSGGSENLYFQSNKSQLYPDSPLTDQDFNQLDQTVI (LanM_Qt-Letter11) EAARRQLVGRRFIELYGPLGRGMQSVFNDIFMESHEAKM DFQGSFDTEVESSRRVNYTIPMLYKDFVLYWRDLEQSKA LDIPIDFSVAANAARDVAFLEDQMIFHGSKEFDIPGLMNVK GRLTHLIGNWYESGNAFQDIVEARNKLLEMNHNGPYALV LSPELYSLLGSGGSGLITAGVFQSPVLKGKSGVIVNTGRN NLDLAISEDFETAYLGEEGMNHPFRVYETVVLRIKRPAAIC TLIDPEE*Amino acid sequence MHHHHHHGGSP I I I I KVDIAAFDPDKDGTIDLKEALAAG of Qt- Veg eta 10 SAAFDKLDPDKDGTLDAKELKGRVSEADLKKLDPDNDGT encapsulin variant LDKKEYLAAVEQFKAANPDNDGTIDARELASPAGSALVNLIRGGSGGSENLYFQSNKSQLYPDSPLTDQDFNQLDQTVIfused to LanM EAARRQLVGRRFIELYGPLGRGMQSVFNDIFMESHEAKM (LanM_Qt-Vegeta10) DFQGSFDTEVESSRRVNYTIPMLYKDFVLYWRDLEQSKA LDIPIDFSVAANAARDVAFLEDQMIFHGSKEFDIPGLMNVK GRLTHLIGNWYESGNAFQDIVEARNKLLEMNHNGPYALV LSPELYSLLHRVGSGGSGTAGVFQSPVLKGKSGVIVNTG RNNLDLAISEDFETAYLGEEGMNHPFRVYETVVLRIKRPA AICTLIDPEE*81 Amino acid sequence MHHHHHHGGSP I I I I KVDIAAFDPDKDGTIDLKEALAAG of Qt-Pigs13 SAAFDKLDPDKDGTLDAKELKGRVSEADLKKLDPDNDGT encapsulin variant LDKKEYLAAVEQFKAANPDNDGTIDARELASPAGSALVNL fused to LanM IRGGSGGSENLYFQSNKSQLYPDSPLTDQDFNQLDQTVI (LanM_Qt-Pigs13) EAARRQLVGRRFIELYGPLGRGMQSVFNDIFMESHEAKM DFQGSFDTEVESSRRVNYTIPMLYKDFVLYWRDLEQSKA LDIPIDFSVAANAARDVAFLEDQMIFHGSKEFDIPGLMNVK GRLTHLIGNWYESGNAFQDIVEARNKLLEMNHNGPYALV LSPELYPYGSCIELITAGVFQSPVLKGKSGVIVNTGRNNLD LAISEDFETAYLGEEGMNHPFRVYETVVLRIKRPAAICTLI DPEE*82 Amino acid sequence MHHHHHHGGSP I I I I KVDIAAFDPDKDGTIDLKEALAAG ofQt-Slay13 SAAFDKLDPDKDGTLDAKELKGRVSEADLKKLDPDNDGT encapsulin variant LDKKEYLAAVEQFKAANPDNDGTIDARELASPAGSALVNL fused to LanM IRGGSGGSENLYFQSNKSQLYPDSPLTDQDFNQLDQTVI (LanM_Qt-Slay13) EAARRQLVGRRFIELYGPLGRGMQSVFNDIFMESHEAKM DFQGSFDTEVESSRRVNYTIPMLYKDFVLYWRDLEQSKA LDIPIDFSVAANAARDVAFLEDQMIFHGSKEFDIPGLMNVK GRLTHLIGNWYESGNAFQDIVEARNKLLEMNHNGPYALV LSPELYSLGSGGSGITAGVFQSPVLKGKSGVIVNTGRNNL DLAISEDFETAYLGEEGMNHPFRVYETVVLRIKRPAAICTL IDPEE*Detailed description of the embodiments

[0103] The ability to package synthetic cargo into protein cages is a crucial limitation for many biotechnological applications. Protein cages typically assemble during their expression inside a cellular host, thus limiting cargo packaging to biomolecules that can be produced inside cells (Figure 1a). Indeed, it is known in the art that protein cages can be readily reprogrammed to package cargo proteins or nucleic acids during assembly inside cells. However, to load synthetic non-biological cargo, protein cages must bepurified from their cellular hosts, disassembled under buffer conditions that disfavour assembly (e.g. changing pH, ionic strength, temperature, or adding chemical denaturants), then reassembled in the presence of the cargo. Drawbacks of this approach include the risk of damage to cage proteins or cargo during exposure to potentially harsh disassembly conditions, as well as suboptimal cargo loading efficiency when a passive statistical encapsulation mechanism is used. Furthermore, the subsequent in vitro reassembly process can have poor fidelity, generating a significant proportion of defective or aggregation-prone cages with inferior structural uniformity and thermal stability relative to their original correctly-assembled form.

[0104] Here, the inventors describe a fusion-based in vitro assembly method for packaging diverse synthetic cargo into encapsulin protein cages, outperforming standard in cellulo assembly methods to produce cages with superior uniformity and thermal stability. The inventors demonstrate that fluorescent dyes, proteins and cytotoxic drug molecules can all be selectively packaged with high efficiency via a peptide-mediated targeting process. The exceptional fidelity and broad compatibility of the present in vitro assembly platform enables generalisable access to cargo-filled protein cages that host novel synthetic functionality for diverse biotechnological applications.Definitions

[0105] It will be understood that the invention disclosed and defined in this specification extends to all alternative combinations of two or more of the individual features mentioned or evident from the text or drawings. All of these different combinations constitute various alternative aspects of the invention.

[0106] Reference will now be made in detail to certain embodiments of the invention. While the invention will be described in conjunction with the embodiments, it will be understood that the intention is not to limit the invention to those embodiments. On the contrary, the invention is intended to coverall alternatives, modifications, and equivalents, which may be included within the scope of the present invention as defined by the claims.

[0107] One skilled in the art will recognize many methods and materials similar or equivalent to those described herein, which could be used in the practice of the present invention. The present invention is in no way limited to the methods and materials described. It will be understood that the invention disclosed and defined in thisspecification extends to all alternative combinations of two or more of the individual features mentioned or evident from the text or drawings. All of these different combinations constitute various alternative aspects of the invention.

[0108] All of the patents and publications referred to herein are incorporated by reference in their entirety.

[0109] For purposes of interpreting this specification, terms used in the singular will also include the plural and vice versa.

[0110] The general chemical terms used in the formulae herein have their usual meaning.Encapsulins

[0111] As used herein, encapsulins are members of a family of bacterial proteins, having the ability to self-assemble to form a protein cage or nano-compartment (or nanoparticle or nanocapsule or "nanocage").

[0112] Encapsulins can form simple non-viral protein cages with exceptional stability and a native peptide-mediated mechanism for loading proteinaceous cargo, serving as ideal candidates for engineering drug delivery vehicles, vaccine scaffolds, and nanoreactors. Until now, the ability to package synthetic cargo into encapsulin cages has been a significant bottleneck, as fully-assembled assembled encapsulins are not sufficiently porous to permit entry of larger molecular cargo. Previously reported protocols for in vitro packaging require exposure to extreme buffer conditions (e.g. pH 1, pH 13, or 7 M GuHCI) for cage disassembly, leading to low-yielding reassembly with substantial cage defects and aggregation.

[0113] The present invention is based on the finding by the inventors that it is possible to obtain in vitro assembly and cargo packaging by using a protein fusion strategy to circumvent the need for disassembly. The approach by the inventors is demonstrated to outperform current in cellulo assembly and in vitro disassembly-reassembly methods by producing cargo-filled encapsulin cages with superior uniformity and thermostability.

[0114] In particular, the inventors have developed a fusion protein strategy that involves the use of an encapsulin protein, such as a Family 1 encapsulin.

[0115] Encapsulins are divided into 4 families based on evolutionary relationships. Family 1 encapsulins are the best-characterised and most commonly used encapsulins. Family 1 encapsulins are well-known to the skilled person and described in the literature, for example in Andreas and Giessen (2021), Nature communications, 12:4748 (and supplementary information), which highlights the families and describes known Family 1 encapsulins.

[0116] A Family 1 encapsulin suitable for use in the invention may be from Quasibacillus thermotolerans (also known as Bacillus thermotolerans), Thermotoga maritima, Myxococcus xanthus, Dendrosporobacter quercicolus, and Bacillus methanolicus. An exemplary sequence of the Q. thermotolerans (QtEnc), T. maritima (TmEnc), M. xanthus (MxEnc), D. quercicolus (DqEnc), and B. methanolicus (BmEnc) encapsulin sequences are provided herein as SEQ ID NO: 6, 25 to 28, and 58 to 62.Lanmodulin and maltose binding protein

[0117] The fusion proteins of the invention comprise the QtEnc, TmEnc, MxEnc, DqEnc, or BmEnc sequence, fused or linked to a polypeptide comprising LanM and / or MBP. Without wishing to be bound by theory, the inventors believe that the use of a polypeptide comprising LanM and / or MBP enables stabilisation of the fusion protein, and ensures that the encapsulin component of the fusion protein does not self-assemble until it is ready to be loaded with cargo.

[0118] As used herein, lanmodulin (LanM) refers to a natural lanthanide-binding protein. An exemplary sequence of lanmodulin is provided herein in SEQ ID NO: 7.

[0119] As used herein, maltose binding protein (MBP) refers to a periplasmic protein of E. coli which is responsible for the uptake and metabolism of maltodextrins. MBP is typically used to increase the solubility of recombinant proteins expressed in E. coli. In addition, MBP can itself be used as an affinity tag for purification of recombinant proteins. The fusion protein binds to amylose columns while all other proteins flow through. The M BP-protein fusion can be purified by eluting the column with maltose. Once the fusion protein is obtained in purified form, the protein of interest is often cleaved from MBP with a specific protease and can then be separated from MBP by affinity chromatography.

[0120] An exemplary sequence of MBP is provided herein in SEQ ID NO: 15.Linkers

[0121] The herein provided fusion proteins may comprise a linker (or “spacer”). In the context of the present invention, the polypeptide comprising or consisting of the amino acid sequence of LanM and / or M BP is fused via a linker at its C-terminus to the encapsulin amino acid sequence.

[0122] A linker is usually a peptide having a length of up to 20 amino acids. The term “linked to” or “fused to” refers to a covalent bond, e.g., a peptide bond, formed between two moieties. Accordingly, in the context of the present invention the linker may have a length of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 or 22 amino acids. For example, the herein provided fusion protein may comprise a linker between the polypeptide comprising or consisting of an amino acid sequence of LanM and the amino acid sequence of an encapsulin, such as between the N-terminus of the encapsulin and the C-terminus of the LanM. Such linkers have the advantage that they can make it more likely that the different polypeptides of the fusion protein fold independently and behave as expected.

[0123] In some aspects, the fusion protein of the present invention includes a peptide linker. The skilled person will be familiar with the design and use of various peptide linkers comprised of various amino acids, and of various lengths, which would be suitable for use as linkers in accordance with the present invention. The linker may comprise various combinations of repeated amino acid sequences.

[0124] The linker may be a flexible linker (such as those comprising repeats of glycine and serine residues), a rigid linker (such as those comprising glutamic acid and lysine residues, flanking alanine repeats) and / or a cleavable linker (such as sequences that are susceptible by protease cleavage). Examples of such linkers are known to the skilled person and are described for example, in Chen et al., (2013) Advanced Drug Delivery Reviews, 65: 1357-1369.

[0125] In some aspects, the peptide linker may include the amino acids glycine and serine in various lengths and combinations. In some aspects, the peptide linker can include the sequence Gly-Gly-Ser (GGS), Gly-Gly-Gly-Ser (GGGS) or Gly-Gly-Gly-Gly-Ser (GGGGS) and variations or repeats thereof. In some aspects, the peptide linker can include the amino acid sequence GGGGGS (a linker of 6 amino acids in length) or evenlonger. The linker may a series of repeating glycine and serine residues (GS) of different lengths, i.e. , (GS)n where n is any number from 1 to 15 or more. For example, the linker may be (GS)s (i.e., GSGSGS) or longer (GS)n or longer. It will be appreciated thatncan be any number including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or more. Fusion proteins having linkers of such length are included within the scope of the present invention. Similarly, the linker may be a series of repeating glycine residues separated by serine residues. For example (GGGGS)3 (i.e., the linker may comprise the amino acid sequence GGGGSGGGGSGGGGS (G4S)s) and variations thereof.

[0126] The peptide linker may consist of a series of repeats of Thr-Pro (TP) comprising one or more additional amino acids N and C terminal to the repeat sequence. For example, the linker may comprise or consist of the sequence GTPTPTPTPTGE (also known as the TP5 linker).

[0127] In certain aspects, the linker may be flexible and cleavable. Such linkers preferably comprise one or more recognition sites for a protease to enable cleavage.Nucleic acids

[0128] An "isolated" nucleic acid molecule is a nucleic acid molecule that is identified and separated from at least one contaminant nucleic acid molecule with which it is ordinarily associated in the natural source. An isolated nucleic acid molecule is other than in the form or setting in which it is found in nature. Isolated nucleic acid molecules therefore are distinguished from the nucleic acid molecule as it exists in natural cells.

[0129] The terms “nucleic acid molecule” and “polynucleotide” are used interchangeably herein and refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. Non-limiting examples of polynucleotides include a gene, a gene fragment, messenger RNA (mRNA), cDNA, recombinant polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A polynucleotide of the invention may be provided in isolated or purified form. A nucleic acid sequence which “encodes” a selected polypeptide is a nucleic acid molecule which is transcribed (in the case of DNA) and translated (in the case of mRNA) into a polypeptide in vivo when placed under the control of appropriate regulatory sequences. The boundaries of the coding sequence are determined by a start codon at the 5' (amino) terminus and a translation stop codon atthe 3' (carboxy) terminus. For the purposes of the invention, such nucleic acid sequences can include, but are not limited to, cDNA from viral, prokaryotic or eukaryotic mRNA, genomic sequences from viral or prokaryotic DNA or RNA, and even synthetic DNA sequences. A transcription termination sequence may be located 3' to the coding sequence.

[0130] Polynucleotides of the invention can be synthesised according to methods well known in the art, as described by way of example in Sambrook et al (1989, Molecular Cloning — a laboratory manual; Cold Spring Harbor Press).

[0131] The polynucleotide molecules of the present invention may be provided in the form of an expression cassette which includes control sequences operably linked to the inserted sequence, thus allowing forexpression of the polypeptide of the invention in vivo in a targeted subject. These expression cassettes, in turn, are typically provided within vectors (e.g., plasmids or recombinant viral vectors) which are suitable for use as reagents for nucleic acid immunization. Such an expression cassette may be administered directly to a host subject. Alternatively, a vector comprising a polynucleotide of the invention may be administered to a host subject. Preferably the polynucleotide is prepared and / or administered using a genetic vector. A suitable vector may be any vector which is capable of carrying a sufficient amount of genetic information, and allowing expression of a polypeptide of the invention.

[0132] The present invention thus includes expression vectors that comprise such polynucleotide sequences.

[0133] Furthermore, it will be appreciated that the compositions and products of the invention may comprise a mixture of polypeptides and polynucleotides. Accordingly, the invention provides a composition or product as defined herein, wherein in place of any one of the polypeptide is a polynucleotide capable of expressing said polypeptide.

[0134] Expression vectors are routinely constructed in the art of molecular biology and may for example involve the use of plasmid DNA and appropriate initiators, promoters, enhancers and other elements, such as for example polyadenylation signals which may be necessary, and which are positioned in the correct orientation, in order to allow for expression of a peptide of the invention. Other suitable vectors would be apparent to persons skilled in the art.

[0135] Thus, the methods of the present invention include delivering such a vector to a cell and allowing transcription from the vector to occur. Preferably, a polynucleotide of the invention or for use in the invention in a vector is operably linked to a control sequence which is capable of providing for the expression of the coding sequence by the host cell, i.e. the vector is an expression vector.

[0136] “Operably linked” refers to an arrangement of elements wherein the components so described are configured so as to perform their usual function. Thus, a given regulatory sequence, such as a promoter, operably linked to a nucleic acid sequence is capable of effecting the expression of that sequence when the proper enzymes are present. The promoter need not be contiguous with the sequence, so long as it functions to direct the expression thereof. Thus, for example, intervening untranslated yet transcribed sequences can be present between the promoter sequence and the nucleic acid sequence and the promoter sequence can still be considered “operably linked” to the coding sequence.

[0137] A number of expression systems have been described in the art, each of which typically consists of a vector containing a gene or nucleotide sequence of interest operably linked to expression control sequences. These control sequences include transcriptional promoter sequences and transcriptional start and termination sequences. The vectors of the invention may be for example, plasmid, virus or phage vectors provided with an origin of replication, optionally a promoter for the expression of the said polynucleotide and optionally a regulator of the promoter. A “plasmid” is a vector in the form of an extra-chromosomal genetic element. The vectors may contain one or more selectable marker genes, for example an ampicillin resistance gene in the case of a bacterial plasmid or a resistance gene for a fungal vector. Vectors may be used in vitro, for example for the production of DNA or RNA or used to transfect or transform a host cell, for example, a mammalian host cell. The vectors may also be adapted to be used in vivo, for example to allow in vivo expression of the polypeptide.

[0138] A “promoter” is a nucleotide sequence which initiates and regulates transcription of a polypeptide-encoding polynucleotide. Promoters can include inducible promoters (where expression of a polynucleotide sequence operably linked to the promoter is induced by an analyte, cofactor, regulatory protein, etc.), repressible promoters (where expression of a polynucleotide sequence operably linked to the promoter is repressed by an analyte, cofactor, regulatory protein, etc.), and constitutive promoters. It is intendedthat the term “promoter” or “control element” includes full-length promoter regions and functional (e.g., controls transcription or translation) segments of these regions.

[0139] As used herein, the term “promoter” is to be taken in its broadest context and includes the transcriptional regulatory sequences of a genomic gene, including the TATA box or initiator element, which is required for accurate transcription initiation, with or without additional regulatory elements (e.g., upstream activating sequences, transcription factor binding sites, enhancers and silencers) that alter expression of a nucleic acid, e.g., in response to a developmental and / or external stimulus, or in a tissue specific manner. In the present context, the term “promoter” is also used to describe a recombinant, synthetic or fusion nucleic acid, or derivative which confers, activates or enhances the expression of a nucleic acid to which it is operably linked. Exemplary promoters can contain additional copies of one or more specific regulatory elements to further enhance expression and / or alter the spatial expression and / or temporal expression of said nucleic acid.

[0140] Exemplary promoters active in mammalian cells include cytomegalovirus immediate early promoter (CMV-IE), human elongation factor 1-a promoter (EF1), small nuclear RNA promoters (U1a and U 1 b), a-myosin heavy chain promoter, Simian virus 40 promoter (SV40), Rous sarcoma virus promoter (RSV), Adenovirus major late promoter, P-actin promoter; hybrid regulatory element comprising a CMV enhancer / p-actin promoter or an immunoglobulin promoter or active fragment thereof. Examples of useful mammalian host cell lines are monkey kidney CV1 line transformed by SV40 (COS-7, ATCC CRL 1651); human embryonic kidney line (293 or 293 cells subcloned for growth in suspension culture; baby hamster kidney cells (BHK, ATCC CCL 10); or Chinese hamster ovary cells (CHO).

[0141] A polynucleotide, expression cassette or vector according to the present invention may additionally comprise a signal peptide sequence. The signal peptide sequence is generally inserted in operable linkage with the promoter such that the signal peptide is expressed and facilitates secretion of a polypeptide encoded by coding sequence also in operable linkage with the promoter.

[0142] Typically a signal peptide sequence encodes a peptide of 10 to 30 amino acids for example 15 to 20 amino acids. Often the amino acids are predominantly hydrophobic. In a typical situation, a signal peptide targets a growing polypeptide chain bearing thesignal peptide to the endoplasmic reticulum of the expressing cell. The signal peptide is cleaved off in the endoplasmic reticulum, allowing for secretion of the polypeptide via the Golgi apparatus. Thus, a peptide of the invention may be provided to an individual by expression from cells within the individual, and secretion from those cells.

[0143] Any appropriate expression vector (e.g., as described in Pouwels et al., Cloning Vectors: A Laboratory Manual (Elsevier, N.Y.: 1985)) and corresponding suitable host can be employed for production of recombinant polypeptides. Expression hosts include, but are not limited to, bacterial species within the genera Escherichia, Bacillus, Pseudomonas, Salmonella, mammalian or insect host cell systems including baculovirus systems (e.g., as described by Luckow et al., Bio / Technology 6: 47 (1988)), and established cell lines such as the COS-7, C127, 3T3, CHO, HeLa, and BHK cell lines, and the like. The skilled person is aware that the choice of expression host has ramifications for the type of polypeptide produced. For instance, the glycosylation of polypeptides produced in yeast or mammalian cells (e.g., COS-7 cells) will differ from that of polypeptides produced in bacterial cells, such as Escherichia coli.Polypeptides

[0144] The present invention provides fusion proteins comprised of an amino acid sequence of an encapsulin (preferably an encapsulin from Quasibacillus thermotolerans, Thermotoga maritima, Myxococcus xanthus, Dendrosporobacter quercicolus, or Bacillus methanolicus) and the amino acid sequence of lanmodulin (LanM) and / or maltose binding protein (MBP).

[0145] Typically, the fusion proteins of the invention are of the structure:[LanM] - linker - [encapsulin], or [encapsulin] - linker - [LanM], such as [LanM] - linker - [QtEnc], or [QtEnc] - linker - [LanM], wherein, preferably the linker region comprises a cleavable sequence for enabling cleavage of the LanM from the encapsulin or QtEnc. Suitable linkers are described elsewhere herein.

[0146] In further embodiments, the fusion protein of the invention may be in the structure:[MBP] - linker - [encapsulin], or [encapsulin] - linker - [MBP], such as [MBP] - linker -[QtEnc], or [QtEnc] - linker - [MBP], wherein, preferably the linker region comprises acleavable sequence for enabling cleavage of the MBP from the encapsulin or QtEnc. Suitable linkers are described elsewhere herein.

[0147] It will be appreciated that the fusion proteins of the invention are typically obtained using recombinant techniques, as are well known to the skilled person.

[0148] A "fragment" or a “variant” is a portion of a polypeptide of the present invention that retains substantially similar functional activity or substantially the same biological function or activity as the polypeptide, which can be determined using assays described herein.

[0149] It will be appreciated that when polypeptides are made using recombinant techniques, the N-terminal methionine residue is often cleaved by the cellular machinery. Accordingly, the present disclosure contemplates the use of any polypeptide sequence described herein, wherein the N-terminal methionine is absent.

[0150] “Percent (%) amino acid sequence identity” or “percent (%) identical” with respect to a polypeptide sequence, i.e. a polypeptide of the invention defined herein, is defined as the percentage of amino acid residues in a candidate sequence that are identical with the amino acid residues in the specific polypeptide of the invention, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent sequence identity, and not considering any conservative substitutions as part of the sequence identity.

[0151] Those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithms (non-limiting examples described below) needed to achieve maximal alignment over the full-length of the sequences being compared. When amino acid sequences are aligned, the percent amino acid sequence identity of a given amino acid sequence A to, with, or against a given amino acid sequence B (which can alternatively be phrased as a given amino acid sequence A that has or comprises a certain percent amino acid sequence identity to, with, or against a given amino acid sequence B) can be calculated as: percent amino acid sequence identity = X / Y100, where X is the number of amino acid residues scored as identical matches by the sequence alignment program's or algorithm's alignment of A and B and Y is the total number of amino acid residues in B. If the length of amino acid sequence A is not equal to the lengthof amino acid sequence B, the percent amino acid sequence identity of A to B will not equal the percent amino acid sequence identity of B to A.

[0152] In calculating percent identity, typically exact matches are counted. The determination of percent identity between two sequences can be accomplished using a mathematical algorithm. A nonlimiting example of a mathematical algorithm utilized for the comparison of two sequences is the algorithm of Karlin and Altschul (1990) Proc. Natl. Acad. Sci. USA 87:2264, modified as in Karlin and Altschul (1993) Proc. Natl. Acad. Sci. USA 90:5873-5877. Such an algorithm is incorporated into the BLASTN and BLASTX programs of Altschul et al. (1990) J. Mol. Biol. 215:403. To obtain gapped alignments for comparison purposes, Gapped BLAST (in BLAST 2.0) can be utilized as described in Altschul et al. (1997) Nucleic Acids Res. 25:3389. Alternatively, PSI-Blast can be used to perform an iterated search that detects distant relationships between molecules. See Altschul et al. (1997) supra. When utilizing BLAST, Gapped BLAST, and PSI-Blast programs, the default parameters of the respective programs (e.g., BLASTX and BLASTN) can be used. Alignment may also be performed manually by inspection. Another non- limiting example of a mathematical algorithm utilized for the comparison of sequences is the ClustalW algorithm (Higgins et al. (1994) Nucleic Acids Res. 22:4673-4680). ClustalW compares sequences and aligns the entirety of the amino acid or DNA sequence, and thus can provide data about the sequence conservation of the entire amino acid sequence. The ClustalW algorithm is used in several commercially available DNA / amino acid analysis software packages, such as the ALIGNX module of the Vector NTI Program Suite (Invitrogen Corporation, Carlsbad, CA). After alignment of amino acid sequences with ClustalW, the percent amino acid identity can be assessed. A non-limiting examples of a software program useful for analysis of ClustalW alignments is GENEDOC™ or JalView (http: / / www.jalview.org / ). GENEDOC™ allows assessment of amino acid (or DNA) similarity and identity between multiple proteins. Another nonlimiting example of a mathematical algorithm utilized for the comparison of sequences is the algorithm of Myers and Miller (1988) CABIOS 4:11-17. Such an algorithm is incorporated into the ALIGN program (version 2.0), which is part of the GCG Wisconsin Genetics Software Package, Version 10 (available from Accelrys, Inc., 9685 Scranton Rd., San Diego, CA, USA). When utilizing the ALIGN program for comparing amino acid sequences, a PAM 120 weight residue table, a gap length penalty of 12, and a gap penalty of 4 can be used.

[0153] The polypeptide desirably comprises an amino end and a carboxyl end. The polypeptide can comprise D-amino acids, L-amino acids or a mixture of D- and L-amino acids. The D-form of the amino acids, however, is particularly preferred since a polypeptide comprised of D-amino acids is expected to have a greater retention of its biological activity in vivo.

[0154] The polypeptide can be prepared by any of a number of conventional techniques. The polypeptide can be isolated or purified from a naturally occurring source or from a recombinant source. Recombinant production is preferred. For instance, in the case of recombinant polypeptides, a DNA fragment encoding a desired peptide can be subcloned into an appropriate vector using well-known molecular genetic techniques (see, e.g., Maniatis et al., Molecular Cloning: A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory, 1982); Sambrook et al., Molecular Cloning A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory, 1989). The fragment can be transcribed and the polypeptide subsequently translated in vitro. Commercially available kits also can be employed (e.g., such as manufactured by Clontech, Palo Alto, Calif.; Amersham Pharmacia Biotech Inc., Piscataway, N.J.; InVitrogen, Carlsbad, Calif., and the like). The polymerase chain reaction optionally can be employed in the manipulation of nucleic acids.

[0155] The term "conservative substitution" as used herein, refers to the replacement of an amino acid present in the native sequence in the peptide with a naturally or non-naturally occurring amino acid or a peptidomimetic having similar steric properties. Where the side-chain of the native amino acid to be replaced is either polar or hydrophobic, the conservative substitution should be with a naturally occurring amino acid, a non- naturally occurring amino acid or with a peptidomimetic moiety which is also polar or hydrophobic (in addition to having the same steric properties as the side-chain of the replaced amino acid).

[0156] Conservative amino acid substitution tables providing functionally similar amino acids are well known to one of ordinary skill in the art. The following six groups are examples of amino acids that may be considered to be conservative substitutions for one another:1) Alanine (A), Serine (S), Threonine (T);2) Aspartic acid (D), Glutamic acid (E);3) Asparagine (N), Glutamine (Q);4) Arginine (R), Lysine (K);5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); and6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W).

[0157] As naturally occurring amino acids are typically grouped according to their properties, conservative substitutions by naturally occurring amino acids can be determined bearing in mind the fact that replacement of charged amino acids by sterically similar non-charged amino acids are considered as conservative substitutions. For producing conservative substitutions by non-naturally occurring amino acids it is also possible to use amino acid analogs (synthetic amino acids) well known in the art. A peptidomimetic of the naturally occurring amino acid is well documented in the literature known to the skilled person and non-natural or unnatural amino acids are described further below. When affecting conservative substitutions the substituting amino acid should have the same or a similar functional group in the side chain as the original amino acid.

[0158] The phrase "non-conservative substitution" or a “non-conservative residue” as used herein refers to replacement of the amino acid as present in the parent sequence by another naturally or non-naturally occurring amino acid, having different electrochemical and / or steric properties. Thus, the side chain of the substituting amino acid can be significantly larger (or smaller) than the side chain of the native amino acid being substituted and / or can have functional groups with significantly different electronic properties than the amino acid being substituted. Examples of non-conservative substitutions of this type include the substitution of phenylalanine or cycohexylmethyl glycine for alanine, isoleucine for glycine, or -NH-CH[(-CH2)5-COOH]-CO- for aspartic acid. Non-conservative substitution includes any mutation that is not considered conservative.

[0159] A non-conservative amino acid substitution can result from changes in: (a) the structure of the amino acid backbone in the area of the substitution; (b) the charge or hydrophobicity of the amino acid; or (c) the bulk of an amino acid side chain. Substitutionsgenerally expected to produce the greatest changes in protein properties are those in which: (a) a hydrophilic residue is substituted for (or by) a hydrophobic residue; (b) a proline is substituted for (or by) any other residue; (c) a residue having a bulky side chain, e.g., phenylalanine, is substituted for (or by) one not having a side chain, e.g., glycine; or (d) a residue having an electropositive side chain, e.g., lysyl, arginyl, or histadyl, is substituted for (or by) an electronegative residue, e.g., glutamyl or aspartyl.

[0160] Alterations of the native amino acid sequence to produce mutant polypeptides, such as by insertion, deletion and / or substitution, can be done by a variety of means known to those skilled in the art. For instance, site-specific mutations can be introduced by ligating into an expression vector a synthesized oligonucleotide comprising the modified site. Alternately, oligonucleotide-directed site-specific mutagenesis procedures can be used, such as disclosed in Walder et al., Gene 42: 133 (1986); Bauer et al., Gene 37: 73 (1985); Craik, Biotechniques, 12-19 (January 1995); and U.S. Pat. Nos. 4,518,584 and 4,737,462. A preferred means for introducing mutations is the QuikChange Site-Directed Mutagenesis Kit (Stratagene, LaJolla, Calif.).

[0161] The terms "N-terminal" and "C-terminal" are used herein to designate the relative position of any amino acid sequence or polypeptide domain or structure to which they are applied. The relative positioning will be apparent from the context. That is, an "N-terminal" feature will be located at least closer to the N-terminus of the polypeptide molecule than another feature discussed in the same context (the other feature possible referred to as "C-terminal" to the first feature). Similarly, the terms "5'-" and "3'-" can be used herein to designate relative positions of features of polynucleotides.

[0162] A recombinant polypeptide made in accordance with the methods of the present invention may also be modified by, conjugated or fused to another moiety to facilitate purification of the polypeptides, or for use in immunoassays using methods known in the art. For example, a polypeptide of the invention may be modified by glycosylation, acetylation, phosphorylation, amidation, derivatization by known protecting / blocking groups, proteolytic cleavage, etc.

[0163] Modifications contemplated herein include, but are not limited to, modification to side chains, incorporating of unnatural amino acids and / or their derivatives during polypeptide synthesis and the use of crosslinkers and other methods which impose conformational constraints on the polypeptides of the invention. Any modification,including post-translational modification, that reduces the capacity of the molecule to form a dimer is contemplated herein. An example includes modification incorporated by click chemistry as known in the art. Exemplary modifications include glycosylation.

[0164] Examples of side chain modifications contemplated by the present invention include modifications of amino groups such as by reductive alkylation by reaction with an aldehyde followed by reduction with NaBH4; amidination with methylacetimidate; acylation with acetic anhydride; carbamoylation of amino groups with cyanate; trinitrobenzylation of amino groups with 2, 4, 6-trinitrobenzene sulphonic acid (TNBS); acylation of amino groups with succinic anhydride and tetrahydrophthalic anhydride; and pyridoxylation of lysine with pyridoxal-5-phosphate followed by reduction with NaBH4.

[0165] The guanidine group of arginine residues may be modified by the formation of heterocyclic condensation products with reagents such as 2,3-butanedione, phenylglyoxal and glyoxal.

[0166] The carboxyl group may be modified by carbodiimide activation via O-acylisourea formation followed by subsequent derivatisation, for example, to a corresponding amide.

[0167] Sulphydryl groups may be modified by methods such as carboxymethylation with iodoacetic acid or iodoacetamide; performic acid oxidation to cysteic acid; formation of a mixed disulphides with other thiol compounds; reaction with maleimide, maleic anhydride or other substituted maleimide; formation of mercurial derivatives using 4-chloromercuribenzoate, 4-chloromercuriphenylsulphonicacid, phenylmercury chloride, 2-chloromercuri-4-nitrophenol and other mercurials; carbamoylation with cyanate at alkaline pH.

[0168] Tryptophan residues may be modified by, for example, oxidation with N-bromosuccinimide or alkylation of the indole ring with 2-hydroxy-5-nitrobenzyl bromide or sulphenyl halides. Tyrosine residues on the other hand, may be altered by nitration with tetranitromethane to form a 3-nitrotyrosine derivative.

[0169] Modification of the imidazole ring of a histidine residue may be accomplished by alkylation with iodoacetic acid derivatives or N-carboethoxylation with diethylpyrocarbonate.

[0170] Examples of incorporating unnatural amino acids and derivatives during protein synthesis include, but are not limited to, use of norleucine, 4-amino butyric acid, 4-amino-3-hydroxy-5-phenylpentanoic acid, 6-aminohexanoic acid, t-butylglycine, norvaline, phenylglycine, ornithine, sarcosine, 4-amino-3-hydroxy-6-methylheptanoic acid, 2-thienyl alanine and / or D-isomers of amino acids. A list of unnatural amino acids contemplated herein is shown in Table 2.Table 2Non-conventional Code Non-conventional Code amino acid amino acida-aminobutyric acid Abu L-N-methylalanine Nmala a-amino-a-methylbutyrate Mgabu L-N-methylarginine Nmarg aminocyclopropane- Cpro L-N-methylasparagine Nmasn carboxylate L-N-methylaspartic acid Nmasp aminoisobutyric acid Aib L-N-methylcysteine Nmcys aminonorbornyl- Norb L-N-methylglutamine Nmgln carboxylate L-N-methylglutamic acid Nmglu cyclohexylalanine Chexa L-N-methylhistidine Nmhis cyclopentylalanine Cpen L-N-methylisolleucine NmileD-alanine Dal L-N-methylleucine NmleuD-arginine Darg L-N-methyllysine NmlysD-aspartic acid Dasp L-N-methylmethionine Nmmet D-cysteine Deys L-N-methylnorleucine NmnleD-glutamine Dgln L-N-methylnorvaline Nmnva D-glutamic acid Dglu L-N-methylornithine Nmorn D-histidine Dhis L-N-methylphenylalanine Nmphe D-isoleucine Dile L-N-methylproline Nmpro D-leucine Dleu L-N-methylserine Nmser D-lysine Dlys L-N-methylthreonine Nmthr D-methionine Dmet L-N-methyltryptophan NmtrpD-ornithine Dorn L-N-methyltyrosine NmtyrD-phenylalanine Dphe L-N-methylvaline Nmval D-proline Dpro L-N-methylethylglycine NmetgD-serine Dser L-N-methyl-t-butylglycine Nmtbug D-threonine Dthr L-norleucine Nle D- tryptophan Dtrp L-norvaline Nva D-tyrosine Dtyr a-methyl-aminoisobutyrate Maib D-valine Dval a-methyl-y-aminobutyrate Mgabu D-a-methylalanine Dmala a-methylcyclohexylalanine Mchexa D-a-methylarginine Dmarg a-methylcylcopentylalanine Mcpen D-a-methylasparagine Dmasn a-methyl-a-napthylalanine Manap D-a-methylaspartate Dmasp a-methylpenicillamine Mpen D-a-methylcysteine Dmcys N-(4-aminobutyl)glycine Nglu D-a-methylglutamine Dmgln N-(2-aminoethyl)glycine Naeg D-a-methylhistidine Dmhis N-(3-aminopropyl)glycine Norn D-a-methylisoleucine Dmile N-amino-a-methylbutyrate Nmaabu D-a-methylleucine Dmleu a-napthylalanine Anap D-a-methyllysine Dmlys N-benzylglycine Nphe D-a-methylmethionine Dmmet N-(2-carbamylethyl)glycine Ngln D-a-methylornithine Dmorn N-(carbamylmethyl)glycine Nasn D-a-methylphenylalanine Dmphe N-(2-carboxyethyl)glycine Nglu D-a-methylproline Dmpro N-(carboxymethyl)glycine Nasp D-a-methylserine Dmser N-cyclobutylglycine Ncbut D-a-methylthreonine Dmthr N-cycloheptylglycine Nchep D-a-methyltryptophan Dmtrp N-cyclohexylglycine Nchex D-a-methyltyrosine Dmty N-cyclodecylglycine Ncdec D-a-methylvaline Dmval N-cylcododecylglycine Ncdod D-N-methylalanine Dnmala N-cyclooctylglycine Ncoct D-N-methylarginine Dnmarg N-cyclopropylglycine Ncpro D-N-methylasparagine Dnmasn N-cycloundecylglycine Ncund D-N-methylaspartate Dnmasp N-(2,2-diphenylethyl)glycine Nbhm D-N-methylcysteine Dnmcys N-(3,3-diphenylpropyl)glycine Nbhe D-N-methylglutamine Dnmgln N-(3-guanidinopropyl)glycine Narg D-N-methylglutamate Dnmglu N-(1-hydroxyethyl)glycine Nthr D-N-methylhistidine Dnmhis N-(hydroxyethyl))glycine Nser D-N-methylisoleucine Dnmile N-(imidazolylethyl))glycine Nhis D-N-methylleucine Dnmleu N-(3-indolylyethyl)glycine NhtrpD-N-methyllysine Dnmlys N-methyl-v-aminobutyrate Nmgabu N-methylcyclohexylalanineNmchexa D-N-methylmethionine Dnmmet D-N-methylornithine Dnmorn N-methylcyclopentylalanine Nmcpen N-methylglycine Nala D-N-methylphenylalanine Dnmphe N-methylaminoisobutyrate Nmaib D-N-methylproline Dnmpro N-(1-methylpropyl)glycine Nile D-N-methylserine Dnmser N-(2-methylpropyl)glycine Nleu D-N-methylthreonine Dnmthr D-N-methyltryptophan Dnmtrp N-(1-methylethyl)glycine Nval D-N-methyltyrosine Dnmtyr N-methyla-napthylalanine Nmanap D-N-methylvaline Dnmval N-methylpenicillamine Nmpen y-aminobutyric acid Gabu N-(p-hydroxyphenyl)glycine Nhtyr L-f-butylglycine Tbug N-(thiomethyl)glycine Ncys L-ethylglycine Etg penicillamine Pen L-homophenylalanine Hphe L-a-methylalanine Mala L-a-methylarginine Marg L-a-methylasparagine Masn L-a-methylaspartate Masp L-a-methyl-f-butylglycine Mtbug L-a-methylcysteine Mcys L-methylethylglycine Metg L-a-methylglutamine Mgln L-a-methylglutamate Mglu L-a-methylhistidine Mhis L-a-methylhomophenylalanine Mhphe L-a-methylisoleucine Mile N-(2-methylthioethyl)glycine Nmet L-a-methylleucine Mleu L-a-methyllysine Mlys L-a-methylmethionine Mmet L-a-methylnorleucine Mnle L-a-methylnorvaline Mnva L-a-methylornithine Morn L-a-methylphenylalanine Mphe L-a-methylproline Mpro L-a-methylserine Mser L-a-methylthreonine Mthr L-a-methyltryptophan Mtrp L-a-methyltyrosine Mtyr L-a-methylvaline Mval L-N-methylhomophenylalanine Nmhphe N-(N-(2,2-diphenylethyl) Nnbhm N-(N-(3,3-diphenylpropyl) Nnbhe carbamylmethyl)glycine carbamylmethyl)glycine1-carboxy-1-(2,2-diphenyl-Nmbcethylamino)cyclopropane

[0171] Crosslinkers can be used, for example, to stabilise 3D conformations, using homo-bifunctional crosslinkers such as the bifunctional imido esters having (CH2) n spacer groups with n=1 to n=6, glutaraldehyde, N-hydroxysuccinimide esters and heterobifunctional reagents which usually contain an ami no- reactive moiety such as N-hydroxysuccinimide and another group specific-reactive moiety.Cell culture

[0172] Persons skilled in the art will be familiar with standard methods for transfecting host cells, such as mammalian cells, with a nucleic acid vector and culturing the host cell in suitable conditions for expressing genes encoded by the vector. Representative methods for transfection and culturing of mammalian cells to produce recombinant protein are described, for example in Ausubel et al., (editors), Current Protocols in Molecular Biology, Greene Pub. Associates and Wiley-lnterscience (1988, including all updates until present) orSambrooketal., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press (1989).

[0173] Means for introducing the isolated nucleic acid, vector or expression construct comprising same into a cell for expression are known to those skilled in the art. The technique used for a given cell depends on the known successful techniques. Means for introducing recombinant DNA into cells include microinjection, transfection mediated by DEAE-dextran, transfection mediated by liposomes such as by using lipofectamine (Gibco, MD, USA) and / or cellfectin (Gibco, MD, USA), PEG-mediated DNA uptake, electroporation and microparticle bombardment such as by using DNA-coated tungsten or gold particles (Agracetus Inc., Wl, USA) amongst others.

[0174] The host cells used in accordance with the present invention may be cultured in a variety of media, depending on the cell type used. Commercially available media such as Ham's FI0 (Sigma), Minimal Essential Medium ((MEM), (Sigma), RPMI-1640 (Sigma), and Dulbecco's Modified Eagle's Medium ((DMEM), Sigma) are suitable for culturing mammalian cells. Media for culturing other cell types discussed herein are known in the art.

[0175] Moreover, the skilled person will be familiar with methods for purifying expressed recombinant protein from cell culture media, including using size exclusion and affinity chromatography methods, and combinations thereof.

[0176] Where a protein is secreted into culture medium, supernatants from such expression systems can be first concentrated using a commercially available protein concentration filter, for example, an Amicon or Millipore Pellicon ultrafiltration unit. A protease inhibitor such as PMSF may be included in any of the foregoing steps to inhibit proteolysis and antibiotics may be included to prevent the growth of adventitious contaminants. Alternatively, or additionally, supernatants can be filtered and / or separated from cells expressing the protein, e.g., using continuous centrifugation.

[0177] The protein prepared from the cells can be purified using, for example, ion exchange, hydroxyapatite chromatography, hydrophobic interaction chromatography, gel electrophoresis, dialysis, affinity chromatography (e.g., lysine affinity column), or any combination of the foregoing. These methods are known in the art and described, for example in WO99 / 57134 or Ed Harlow and David Lane (editors) Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory, (1988).

[0178] The skilled artisan will also be aware that a protein can be modified to include a tag to facilitate purification or detection, e.g., a poly-histidine tag, e.g., a hexa-histidine tag, or an influenza virus hemagglutinin (HA) tag, or a Simian Virus 5 (V5) tag, or a FLAG tag, or a glutathione S-transferase (GST) tag. The resulting protein is then purified using methods known in the art, such as, affinity purification. For example, a protein comprising a hexa-his tag is purified by contacting a sample comprising the protein with nickelnitrilotriacetic acid (Ni-NTA) that specifically binds a hexa-his tag immobilized on a solid or semi-solid support, washing the sample to remove unbound protein, and subsequently eluting the bound protein. Alternatively, or in addition a ligand or antibody that binds to a tag is used in an affinity purification method.Assembly formation and cargo loading of protein cages

[0179] It will be appreciated that formation of the protein cages of the present invention requires cleavage of the fusion proteins described herein, so as to remove the LanM and / or MBP portion from the encapsulin portion, and thereby enable self-assembly of the encapsulin. Hybrid protein cages may be formed by combinations of an encapsulin and variants thereof. Hybrid protein cages may include combinations of variants of encapsulin that are independently able to form a protein cage and / or are competent to self assemble with variants of encapsulin that are independently unable to form a protein cage and / or are not competent to self-assemble.

[0180] Accordingly, it will be appreciated that cleavage of the fusion protein will be dependent on the specific cleavage sequence included in the fusion protein. For example, in the instance where the cleavage sequence is comprised of the TEV protease recognition sequence (ENLYFQ / GS), cleavage will be performed using the TEV protease, according to standard methods, such as those described herein in Example 1.

[0181] Other cleavage sequences and their associated proteases / cleavage partners are described elsewhere herein.

[0182] Following cleavage of the fusion protein, the fusion protein can be contacted with the desired cargo molecules. The cargo molecules may be nucleic acid molecules (such as RNA, DNA, siRNA, shRNA), peptides, proteins, or small molecules, including synthetic small molecule compounds.

[0183] In vitro encapsulin assembly and cargo loading may be conducted by any suitable method known to the skilled person, including but not limited to the methods described herein in Example 1.

[0184] Following formation of the loaded protein cages, and depending on the intended cargo, the protein cages may be purified and optionally further formulated in a pharmaceutically acceptable excipient or other carrier solution for enabling delivery of the loaded protein cages in vivo.

[0185] As used herein, a cargo loading moiety (CLM) comprises an amino acid sequence that enables cargo loading of a protein cage or hybrid protein cage formed by an encapsulin from specific bacteria. Such amino acid sequences may include short, conserved peptide sequences at the termini of cargo proteins, often referred to as targeting peptides (TP) or cargo loading peptides (CLP) for a bacteria, or more extended N-terminal encapsulation-mediating domains. In the broadest definition, CLM as used herein includes all sequences that enable loading of a protein cage or a hybrid protein cage formed by an encapsulin, including TP / CLP and the encapsulation-mediating domains. Cargo molecules may be functionalised for cargo loading of a protein cage or hybrid protein cage by conjugation with a CLM, which may comprise the amino acid sequence of a CLP.

[0186] For example, conjugation of a cargo molecule with a CLM comprising a CLP for Q. thermotolerans (QtCLP) enables loading of a protein cage formed by encapsulin fromQ. thermotolerans (QtEnc) with the conjugated cargo molecule. Without being bound by the examples in Table 3, CLPs for various Family 1 encapsulins are known by the skilled person and may be used in the current invention.Table 3: Examples of amino acid sequences of CLPs for specific encapsulinsSpecies Encapsulin CLPQ. thermotolerans Wt QtEnc SEQ ID NO: 6 QtCLP, C- SEQ ID NOs:modified CLP, 18, 19, 20 Qt-Glass9 SEQ ID NO: 58QtCLP shortQt-Letter11 SEQ ID NO: 59Qt-Vegeta10 SEQ ID NO: 60Qt- Pigs 13 SEQ ID NO: 61Qt-Slay13 SEQ ID NO: 62T. maritima TmEnc SEQ ID NO: 25 TmCLP, C- SEQ ID NOs:modified TmCLP 45, 46 M. xanthus MxEnc SEQ ID NO: 26 MxCLP, C- SEQ ID NOs:modified MxCLP 47, 48 D. quercicolus DqEnc SEQ ID NO: 27 DqCLP, C- SEQ ID NOs:modified DqCLP 49, 50 B. methanolicus BmEnc SEQ ID NO: 28 BmCLP, C- SEQ ID NOs:modified BmCLP 51, 52

[0187] Conjugating multiple cargo molecules to a CLM with a shared CLP sequence enables loading of a protein cage or hybrid protein cage with multiple cargo molecules in the methods of the invention. A single cargo molecule may also be conjugated to a CLM comprising more than one CLP, which functionalises the cargo molecule for loading of protein cages formed by encapsulins of different bacteria. For example, a cargo molecule may be conjugated to a CLM comprising SEQ ID NO: 18 and SEQ ID NO: 47, which would functionalise the cargo molecule for loading protein cages formed by QtEnc and MxEnc.Examples

[0188] In this work, the inventors report the high-fidelity packaging of diverse synthetic cargo into encapsulin protein cages. Encapsulins are simple non-viral protein cages with exceptional stability and a native peptide-mediated mechanism for loading proteinaceous cargo, serving as ideal candidates for engineering drug delivery vehicles, vaccine scaffolds, and nanoreactors.

[0189] Until now, the ability to package synthetic cargo into encapsulin cages has been a significant bottleneck, as fully-assembled encapsulins are not sufficiently porous to permit entry of larger molecular cargo. Previously reported protocols for in vitro packaging require exposure to extreme buffer conditions (e.g. pH 1, pH 13, or 7 M GuHCI) for cage disassembly, 36 leading to low-yielding reassembly with substantial cage defects and aggregation (Figure 1b).

[0190] The inventors achieve in vitro assembly and cargo packaging by using a protein fusion strategy to circumvent the need for disassembly (Figure 1c), outperforming current in cellulo assembly and in vitro disassembly-reassembly methods by producing cargo-filled encapsulin cages with superior uniformity and thermostability.Example 1: Materials and methodsMolecular cloning

[0191] The codon-optimised gene encoding LanM-QtEnc was synthesised as a gBIock (Integrated DNA Technologies) and cloned into a pETDuet-1 vector between the Ncol and Kpnl restriction sites by Gibson assembly using NEBuilder HiFi DNA Assembly Master Mix (New England Biolabs). Constructs were transformed into DH5a E. coli cells and the gene sequence was verified by Sanger sequencing at the Australian Genome Research Facility. Full details of DNA sequences are provided in Table 1.Recombinant protein production

[0192] Plasmids were transformed into BL21(DE3) E. coli. Single colonies were grown overnight at 37 °C in lysogeny broth (LB) as 10 mL cultures. Overnights cultures were used to inoculate 400 mL LB cultures (OD600 0.05) and grown at 37 °C to OD600 0.4-0.6, followed by induction with 0.1 mM isopropyl p-D-1-thiogalactopyranoside (IPTG) and expression overnight at 18 °C. Cells were pelleted at 3900 g for 20 min (581 OR centrifuge with S-4-104 rotor, Eppendorf) and stored at -20 °C until purification.Protein purification

[0193] The following method was used for the purification of LanM-QtEnc, mNeon-Qt, and mNeon-Tm. Frozen pellets were resuspended in 25 mL SEC buffer (50 mM Tris pH 8, 200 mM NaCI, 1 mM tris(2-carboxyethyl)phosphine (TCEP)) with the addition of 10 pg / mL DNase I (bovine pancrease, Sigma-Aldrich), 100 pg / mL lysozyme (chicken egg white, Sigma-Aldrich), and 1x protease inhibitor (complete EDTA-free Protease Inhibitor Cocktail, Roche). After 30 min on ice, cells were lysed by sonication using a Sonopuls HD 4050 with TS-106 probe (Bandelin) with a program of 55% amplitude for 11 min with a pulse time of 8 s on and 10 s off. The lysate was clarified at 17000 g for 40 min, then incubated with rocking for 1 h at 4 °C with 1 mL HisPur Ni-NTA resin (ThermoFisher Scientific) pre-equilibrated in SEC buffer. The mixture added to an Econo-Column chromatography column (Bio-Rad), and the beads were washed with 3 x 2 mL wash buffer (50 mM Tris pH 8, 200 mM NaCI, 1 mM TCEP, 20 mM imidazole), and eluted with 4 x 1 mL elution buffer (50 mM Tris pH 8, 200 mM NaCI, 1 mM TCEP, 500 mM imidazole). The eluted protein was purified by size-exclusion chromatography on a HiLoad 16 / 600 Superdex 200 pg column using SEC buffer at a flow rate of 1 mL / min on an AKTA Pure (Cytiva). All fractions containing the desired protein were concentrated and buffer exchanged using Amicon Ultra-15 centrifugal filters (10 kDa MWCO, Merck) into cleavage buffer (50 mM Tris pH 8, 50 mM NaCI, 1 mM TCEP, 0.5 mM ethylenediaminetetraacetic acid (EDTA)).

[0194] Tobacco Etch Virus (TEV) protease was purified using the same lysis and Ni-NTA affinity purification protocol. The eluted protein was dialysed into TEV-SEC buffer (20 mM HEPES pH 8, 200 mM NaCI, 2 mM EDTA, 2 mM dithiothreitol (DTT)) using Snakeskin dialysis tubing (3.5 kDa MWCO, ThermoFisher Scientific). The dialysed protein was purified by size-exclusion chromatography on a HiLoad 16 / 600 Superdex 200 pg column using TEV-SEC buffer at a flow rate of 1 mL / min on an AKTA Pure (Cytiva). Fractions containing TEV protease were concentrated and buffer exchanged using Amicon Ultra-15 centrifugal filters (10 kDa MWCO, Merck) into 40 mM HEPES pH 8, 400 mM NaCI, 4 mM EDTA, 4 mM DTT, and diluted 1:1 with glycerol before being stored at -80 °C.Solid-phase peptide synthesis

[0195] Manual peptide synthesis was performed on Rink Amide MBHA resin (0.52-0.65 mmol / g loading, Merck). Couplings were carried out by adding HATLI (4 eq) and DI PEA (4 eq) to a solution of the Fmoc-protected amino acid (4 eq) in DMF. This pre-activated mixture was then added to the resin in DMF and shaken for 1 h. The side chain protecting groups used were: t-Bu for Asp, Glu, Ser, Thr; Boc for Lys; Pbf for Arg; Trt for Asn, Gin, His. Fmoc deprotection was carried out with 20% piperidine in DMF (2 * 1 min, 1 x 10 min). N-terminal capping with 5-TAMRA was conducted by adding pre-activated 5-TAMRA (4 eq) with HATLI (4 eq) and DI PEA (8 eq) in DMF to the resin swelled in DMF and shaking at rt for 3 h.

[0196] Cleavage from the resin was achieved with TFA containing 2.5% triisopropylsilane and 2.5% water for 2 h. After cleavage, the mixture was concentrated under a stream of nitrogen. The crude residue was triturated with diethyl ether (15-20 mL) before purification by reverse-phase chromatography.

[0197] Semi-preparative reverse-phase HPLC was performed on a Waters 1525 binary HPLC pump with a Waters 2707 autosampler. Peptide samples were eluted through a Waters XBridge Peptide BEH C18 OBD™ Prep Column (10 mm x 250 mm, 5 pm) using a linear gradient system of 0.1% (v / v) TFA in MilliQ water (solvent A) and 0.1% (v / v) TFA in MeCN (solvent B) at a flow rate of 4 mL / min over 30 min. Peptides were monitored by UV absorbance at 220 nm on a Waters 2489 UV Detector with semi-prep cell. Fractions were collected with Waters Fraction Collector III.

[0198] To produce Aldox-Qt, a Cys-modified Qt CLP was synthesised and purified as per the aforementioned methodology, with the addition of an N-terminal cysteine residue. The peptide was then capped by adding acetic anhydride (4 eq) and DI PEA (4 eq) in DMF for 1 h before cleavage from the resin and HPLC purification as before. The purified peptide in DMSO (1 eq), aldoxorubicin in DMF (1 eq), and DIPEA (1 eq) were shaken in 1:1 MeCN:water at 37 °C for 3 h to yield crude Aldox-Qt. Automated flash reverse-phase column chromatography was carried out on a Biotage Selekt using pre-packed Biotage Star C4 column (10 g) for the purification of Aldox-Qt.

[0199] LCMS chromatograms were obtained using a Shimadzu Nexera-I LC-2040C Plus coupled to a Shimadzu LCMS-2020 mass spectrometer ESI single quadrupole mass detector. Peptide samples were eluted through a Shimadzu Shim-Pack Velox SP-C18 column (2.1 mm x 50 mm, 2.7 pm) with a linear gradient system of 0.1% (v / v) formic acidin MilliQ water (solvent A) and 0.1% (v / v) formic acid in MeCN (solvent B) at a flow rate of 0.4 mL / min over 6.4 min. Peptide samples were monitored by UV absorbance at 220 and 254 nm. Mass spectra were obtained by electrospray ionisation in both positive and negative modes, scanning between m / z 200 and 2000.In vitro encapsulin assembly and cargo loading

[0200] Cleavage reactions were performed on 100 pL scale by incubating 2 mg / mL LanM-QtEnc and 0.4 mg / mL TEV protease in cleavage buffer overnight at 34 °C. Cargo loading experiments included a mixture of LanM-QtEnc and either mNeon or TMR cargo in a 10:1 molar ratio, or Aidox cargo in a 2:1 molar ratio. After overnight cleavage, the reactions were passed through 0.1 mL Ni-NTA resin (for empty or mNeon loading), Capto Core 700 resin (for TMR loading). The in vitro assemblies were then purified by sizeexclusion chromatography on a Superose 6 Increase 10 / 300 GL column at 0.5 mL / min flow rate in SEC buffer.

[0201] The packaging efficiency of mNeonGreen and 5-TAMRA was estimated by measuring the absorbance of purified cargo-filled encapsulins at 506 and 550 nm respectively on a Nanodrop ND-1000 spectrophotometer (ThermoFisher Scientific). The packaging of Aidox was monitored by absorbance at 490 nm during the SEC purification. Fluorescence of the SEC purified Aidox loaded shells was measured using a Perkin Elmer EnSpire multimode plate reader (excitation 470 nm, emission 595 nm). Due to low signal intensity, 100 technical replicate measurements were obtained for each sample, with the mean intensity reported with the standard error.Analytical size exclusion chromatography

[0202] Analytical SEC was conducted on a Nexera Bio HPLC (Shimadzu) using a Bio SEC-52000 A pore size column (7.8 x 300 mm, Agilent) with a Bio SEC-52000 A guard column (7.8 x 50 mm, Agilent). Multi-angle light scattering (MALS) measurements were obtained on a Dawn 8 MALS detector (Wyatt Technologies). Concentration was calculated from change in refractive index assuming a dn / dc of 0.185. The mobile phase was phosphate-buffered saline (137 mM NaCI, 2.7 mM KCI, 10 mM Na2HPO4, 1.8 mM KH2PO4, pH 7.4) at a flow rate of 1 mL / min for 20 min, analysing 20 pL samples taken directly from cleavage reactions. Astra software v7.3.0.18 (Wyatt Technologies) was used to calculate molecular weights.Polyacrylamide gel electrophoresis

[0203] SDS-PAGE was conducted using Mini-PROTEAN TGX Stain-Free gels (BioRad) with Tris-glycine-SDS running buffer. Gels were stained using Coomassie Brilliant Blue G-250 (80 mg / L in 30 mM HCI). Unstained Protein Standard, Broad Range 10-200 kDa (New England Biolabs) was used as the ladder for all SDS-PAGE gels.

[0204] Blue Native PAGE was conducted using NativePAGE 3-12% Bis-Tris 1.0 mm Mini Protein Gels (ThermoFisher Scientific). Protein samples were diluted to an A280 of 0.4 in NativePAGE sample buffer and 9 pL was loaded in each well. NativeMark Unstained Protein Standard (ThermoFisher Scientific) was used as the ladder for all BN-PAGE gels. The anode, dark blue cathode, and light blue cathode buffers were prepared as per manufacturer instructions. Gels were run in dark blue buffer for 5 minutes at 150 V before switching to light blue buffer for 75 minutes at 150 V. The anode buffer was not changed during this time. After electrophoresis, in-gel fluorescence was measured for samples loaded with TAMRA (Trans-UV excitation 302 nm, 590 / 110 nm emission filter) or mNeonGreen (Epi-blue excitation 460-490 nm, 532 / 28 nm emission filter) on a ChemiDoc MP Imaging System (Bio-Rad). Gels were then stained with Coomassie Brilliant Blue G and then destained in water prior to imaging.Transmission electron microscopy

[0205] Encapsulin samples for negative-stain TEM were diluted to 100 pg / ml in 20 mM Tris buffer (pH 8). Gold grids (200-mesh coated with Formvar-carbon film, EMS) were made hydrophilic by glow discharge at 50 mA for 30 s. The grid was floated on 20 pl of sample for 1 min, blotted with filter paper, and washed once with distilled water before staining with 20 pl of uranyl acetate for 1 min. TEM images were captured using a JEOL 1400 microscope at 120 keV at Sydney Microscopy and Microanalysis.

[0206] Dynamic light scattering and thermal unfolding analysis

[0207] Wt QtEnc and in vitro assembled QtEnc were prepared at 0.5 mg / mL in 50 mM Tris pH 8, 200 mM NaCI and loaded onto Prometheus Panta (NanoTemper) as per the manufacturer instructions. Samples were analysed using the Size Analysis function, in which isothermal DLS scans (10 acquisitions, 5s each, 100% DLS-Laser intensity) were performed at 25 °C.

[0208] For differential scanning fluorimetry (nanoDSF), wt QtEnc, LanM-QtEnc and in vitro assembled QtEnc were prepared at 0.5 mg / mL in the same buffer and loaded onto standard capillaries (NanoTemper). Capillaries were sealed at the ends (Capillary Sealing Paste; NanoTemper). Thermal unfolding was determined on Prometheus Panta (NanoTemper) with heating ramp of 1 °C from 25 °C to 110 °C and 20% sensitivity setting.Cell culture and in vitro loaded QtEnc uptake into murine cells

[0209] Murine RAW 264.7 cells were maintained in DMEM (Gibco) supplemented with 10% fetal bovine serum (FBS) and 1% penicillin / streptomycin at 37 °C with 5% CO2.

[0210] Cells were seeded (day 1) at a density of 50,000 cells / well in black 96 well plates (ThermoFisher Scientific) overnight. On day 2, culture media was removed, cells washed with PBS, and media was replaced with starvation media (culture medium with 0% FBS). On day 3, starvation media was removed, cells were washed, and normal culture media containing 0.4 pg / ml of cargo-filled QtEnc, or equivalent control, was added to the cells. Cells were incubated with cargo-filled QtEnc for 2 h at 37 °C with 5% CO2. Following incubation, cells were washed three times with PBS with agitation for 5 mins at room temperature. Cells were fixed by incubating the cells with 4% paraformaldehyde for 15 mins with agitation at room temperature. Cells were then washed again and counterstained with DAPI. Following a third wash, 100 pL of PBS was then added to the cells in preparation for imaging. The proportion of cells with internalised fluorescent cargo-filled QtEnc was visualized via confocal microscopy using an A1 confocal microscope (Nikon). Three independent fields of view were obtained for each of three biological replicates, and the resulting images were analysed in Fiji (Imaged).Example 2: LanM-QtEnc is a stable fusion protein that does not readily selfassemble into native cages

[0211] To engineer a stable pre-assembly form of encapsulin, the inventors fused the 12 kDa protein lanmodulin (LanM) to the N-terminus of the 32 kDa encapsulin protein from Quasibacillus thermotolerans (QtEnc). The inventors bridged the two proteins by a flexible glycine-serine linker and a TEV protease cleavage site, allowing for removal of the LanM fusion as a trigger for assembly. Finally, the inventors included a hexahistidine tag at the N-terminus of LanM to facilitate purification (recognising that this is not a criticalaspect of the fusion protein). The resulting 46 kDa fusion protein is henceforth referred to as LanM-QtEnc.

[0212] Upon recombinant production in E. coli, the inventors found that LanM-QtEnc is a stable fusion protein that does not self-assemble into cages. Purification of LanM-QtEnc involved an initial affinity chromatography step over Ni-NTA resin, followed by sizeexclusion chromatography (SEC) on a Superdex 200 column to yield the pure fusion protein as determined by SDS-PAGE analysis (Figure 2a). Subsequent blue native PAGE (BN-PAGE) analysis showed that LanM-QtEnc predominantly forms a single low molecular weight species, in contrast to the expected 8 MDa cage from in cellulo assembly which is henceforth referred to as wild-type QtEnc (Figure 2b).

[0213] Attempts to visualise the protein by negative-stain transmission electron microscopy (TEM) were consistent with lack of assembled cage-like structures (Figure 2e).

[0214] The inventors used analytical SEC to confirm that LanM-QtEnc exists in a predominantly unassembled form. SEC was performed on two different columns to achieve sufficient separation across a broad size range. Separation on a Superose 6 Increase column 10 / 300 GL (5-5000 kDa) showed a clear difference between wild-type QtEnc and LanM-QtEnc (Figure 2c), with the wild-type cage eluting early in the unretained peak, and the fusion eluting 9 minutes later, near the end of the chromatograph. Meanwhile, a Bio SEC-5 HPLC column with 2000 A pore size (maximum mass range >10 MDa) also showed a substantial difference in retention time (Figure 2d), using light scattering in place of UV absorbance to enhance detection sensitivity at analytical scale. Notably for wild-type QtEnc, the inventors also observed the presence of two additional peaks with earlier retention times (Figure 2d), indicating the existence of aggregated or misassembled cages that likely formed as a result of recombinant overexpression in E. coli. This polydisperse behaviour has not previously been reported for QtEnc in the literature due to limited resolving range of analytical SEC columns used in prior studies.Example 3: Cages assembled in vitro upon fusion cleavage are more uniform and stable than cages assembled in cellulo

[0215] The inventors observed the high-fidelity in vitro assembly of encapsulin cages upon cleavage of the LanM-QtEnc fusion with TEV protease. SEC analysis on the BioSEC-5 HPLC column revealed that the in vitro assembled cages were more uniform than wild-type QtEnc assembled in E. coli, forming a single species with a calculated mass of 8.2 ± 0.2 MDa from multi-angle light scattering (MALS) measurements (Figure 3a). Encapsulin cages were isolated post-cleavage by SEC on a Superose 6 Increase 10 / 300 GL column, eluting in the unretained peak as expected for a large 8 MDa assembly (Figure 3b). BN-PAGE analysis provided further confirmation of the clean and complete conversion to an assembled state (Figure 3c), while SDS-PAGE analysis showed that cleavage had proceeded to >95% completion based on gel densitometry (Figure 3d).

[0216] The successful assembly of uniform cages with the expected diameter of 42 nm was confirmed by negative stain TEM (Figure 3e) and dynamic light scattering (DLS) measurements (Figure 3f). The 50 nm diameter calculated for wild-type QtEnc by DLS provided further evidence of non-uniformity for assemblies produced in E. coli. Notably, unexpected diameters have been observed by DLS in previous attempts to disassemble and re-assemble encapsulins in vitro using denaturants, base, and heat, as well as in vitro studies on an engineered pH-switchable version of QtEnc, both of which were attributed to some degree of misassembly or aggregation.

[0217] The in vitro assembled cages showed exceptional thermal stability that exceeded that of the wild-type QtEnc assembled in E. coli. The melting temperature (Tm) of the in vitro assemblies was 93.7 °C by differential scanning fluorimetry (DSF), compared to 55.9 °C for the LanM-QtEnc fusion prior to cleavage (Figure 3g). Meanwhile, wild-type QtEnc had a Tm of 84.4 °C, consistent with equivalent data from Giessen et al. in the first reported characterisation of QtEnc (2019, eLife 8, e46070). This increased thermal stability is most likely a direct result of higher fidelity cage assembly in the in vitro setting.Example 4: Protein cargo and synthetic molecules can be selectively and quantitatively packaged during in vitro assembly

[0218] To demonstrate that the native cargo packaging system of encapsulins was functional in vitro, the inventors encapsulated fluorescent protein mNeonGreen with the specific cargo loading peptide (CLP) for Q. thermotolerans encapsulin fused to the C-terminus (mNeon-Qt) (Figure 4a). The negative control for cargo packaging was the analogous mNeonGreen construct fused to the CLP for T. maritima encapsulin (mNeon-Tm), which does not mediate cargo packaging when co-expressed with QtEnc in E. coli. mNeon-Qt or mNeon-Tm was added during TEV cleavage of LanM-QtEnc, using a 1:10molar ratio of cargo:encapsulin to avoid the possibility of defects arising from cargo overloading.

[0219] After in vitro assembly, Ni-NTA resin was used to remove TEV protease, cleaved LanM, and any potentially unencapsulated mNeon cargo. SEC purification on a Superose 6 Increase 10 / 300 GL column showed that the efficiency of in vitro assembly was unaffected by the presence of cargo (Figure 4b), while negative stain TEM imaging confirmed the presence of uniform assemblies (Figure 4c). SDS-PAGE analysis showed that only mNeon-Qt co-eluted with the in vitro assembled encapsulins (Figure 4d), and BN-PAGE analysis confirmed that mNeon-Qt was selectively packaged into cages while mNeon-Tm was not packaged (Figure 4e). No unencapsulated mNeon-Qt cargo was observed during purification (data not shown), and absorbance measurements on purified mNeon-filled cages for total protein at 280 nm and mNeon-Qt cargo at 506 nm confirmed that mNeon-Qt had been quantitatively packaged at the original 1:10 molar ratio with no loss.

[0220] Next, the inventors showed that synthetic small molecules can be packaged in vitro (Figure 5a). The synthetic fluorescent dye 5-carboxytetramethylrhodamine (TMR) was attached to the N-terminus of the CLPs for Q. thermotolerans (TMR-Qt) and T. maritima (TMR-Tm) by Fmoc solid-phase peptide synthesis. The in vitro cargo loading experiments were repeated with TMR in place of mNeonGreen and using Capto Core 700 resin in place of Ni-NTA resin after assembly to remove TEV protease, cleaved LanM, and any potential unencapsulated cargo. Once again, the presence of synthetic cargo did not affect the fidelity and efficiency of assembly according to analytical SEC (Figure 5b) and TEM (Figure 5b). In-gel fluorescence was observed after conducting BN-PAGE on purified cages in vitro loaded with TMR-Qt, confirming successful and specific encapsulation of the synthetic dye, while no fluorescence was observed for the TMR-Tm negative control (Figure 5c).

[0221] To expand the scope of cargo packaging beyond proteins and synthetic fluorophores, the inventors packaged the synthetic drug molecule aldoxorubicin (Aidox), a maleimide-functionalised prodrug form of the chemotherapeutic doxorubicin (Dox) (Figure 5d). Aidox was conjugated onto a Cys-modified Qt CLP to generate Aldox-Qt, while free Dox was used as a negative control. Both Aldox-Qt and free Dox were in vitro packaged into in a 2:1 molar ratio of encapsulin to cargo. This higher ratio of cargopackaging was used to provide sufficient photophysical signal for detection, given the reduced brightness of Aidox relative to the previously packaged fluorophores.

[0222] Analytical SEC on a Superose 6 Increase 10 / 300 GL column showed an Aldox-specific absorbance peak (490 nm) in encapsulin fraction for in vitro packaged Aldox-Qt (Figure 5e), while no Aidox absorbance was observed in vitro packaged free Dox (Figure 5f). Furthermore, fluorescence emission was only observed in encapsulin samples purified after in vitro assembly with Aldox-Qt, while the negative control with free Dox showed equivalent signal to the buffer-only control (Figure 5g), further verifying that Aldox-Qt was in vitro loaded to QtEnc. Taken together, these results show that diverse synthetic cargo can be efficiently packaged into encapsulins in vitro by taking advantage of the highly specific encapsulin-CLP interaction.Example 5: Synthetic cargo packaged within encapsulins can be delivered into murine RAW 264.7 cells

[0223] The inventors investigated whether in vitro assembled encapsulins could facilitate uptake of fluorescent mNeon-Qt and TMR-Qt cargo into live cells. Cellular uptake of packaged cargo into murine RAW 264.7 macrophages was assessed by confocal fluorescence microscopy. Intracellular fluorescence was readily observed in 83% of cells when encapsulated mNeon-Qt was added, while no fluorescence was observed when free non-encapsulated mNeon-Qt was used (Figure 6a-c). Similarly, intracellular fluorescence was evident for packaged TMR-Qt in 50% of cells, while free TMR-Qt did not result in any fluorescence (Figure 6d-f).

[0224] This result suggests that encapsulation either enhances cellular uptake of fluorescent cargo or protects the cargo against degradation by the cellular machinery prior to imaging.

[0225] In summary, these results provide a proof of concept that in vitro assembled encapsulins are viable cargo delivery vehicles.Example 6: Summary and discussion

[0226] LanM-QtEnc is a novel protein construct that enables in vitro packaging of synthetic cargo into encapsulin cages, providing access to new biotechnological opportunities that cannot be explored using standard cell-based assembly approaches.In vitro packaging is necessary for any synthetic molecular cargo that is incompatible with living cells due to cytotoxicity, metabolic instability, or membrane impermeability. In applications where multiple cargo types are packaged simultaneously, in vitro packaging also provides the opportunity for far greater control over cargo stoichiometry than state-of-the-art genetic methods for tuning expression levels in cells. Even in the case of biomolecular cargo that can be produced in cells, in vitro packaging solves any timing issues where cargo requires additional processing prior to packaging, such as any cargo that requires multiple subunits to form an active complex, or cargo that must undergo post-translational modification to be functional.

[0227] A distinctive feature of the LanM-QtEnc system is the exceptional fidelity of cage assembly and cargo packaging. Defective assembly pathways are commonly encountered during most in vitro assembly methods, as well as in some in cellulo assembly methods. Low-fidelity assembly processes result in defects or aggregates that can reduce the overall yield of desired cages, pose challenges for sample purification, and undermine the fundamental properties that make protein cages valuable in biotechnology. In nanoreactor design for example, protein cages can act as a semi-permeable barrier to control substrate flux and sequester toxic or unstable intermediates, or alternatively act as a protective layer for fragile enzymatic cargo. Assembly defects compromise the ability of cages to achieve both of these core functions. In drug delivery applications, heterogeneous formulations arising from defective cage assemblies may be less stable in vivo, and hence less effective in shielding cytotoxic payloads from nonspecific release prior to reaching the site of action. In addition, cargo packaging efficiency is a critical parameterwhen working with valuable cargo where losses must be minimised.

[0228] The LanM-QtEnc fusion stands out as a unique example of a stable encapsulin construct that does not form a fully-assembled cage. In most literature examples of encapsulin engineering, the assembled state of encapsulins is strongly favoured. Lee et al. (2008) Journal of Virological Methods, 151: 172-180, reported several N-terminal HBCM2 fusion constructs to engineered versions of the 24 nm encapsulin from T. maritima, observing cage formation along with increased heterogeneity in most cases. A notable native example of N-terminal encapsulin fusion that does not impede assembly is the encapsulin system from Pyrococcus furiosus, which has its cargo directly fused on the N-terminal interior of the cage.46 The N-terminal fusion behaviour of LanM-QtEnc also differs from the HK97 bacteriophage capsid protein (from which encapsulins derivetheir HK97-like protein fold), where the scaffold protein gp5 is initially fused to the capsid protein during the immature Prohead-I that can form a mix of cages and dissociated capsomeres, but then is cleaved to initiate maturation to the final Head II phage capsid form.

[0229] While the robustness of assembled encapsulin cages is ideal for many biotechnological applications, this robustness has also posed a significant technical barrier for previously reported disassembly-reassembly protocols, necessitating extreme pH or chemical denaturants to attempt cargo loading with poor fidelity.

[0230] Unlike all these examples, LanM-QtEnc entirely circumvents the challenge by starting from a non-assembled state. More generally, the extreme buffer conditions required for encapsulin disassembly are not shared by all protein cages. Some VLPs are inherently capable of disassembly via relatively mild changes in experimental buffer conditions. VLPs of the brome mosaic virus coat protein can assemble at low pH and ionic strength but disassemble at neutral pH and high ionic strength, while VLPs of polyomaviruses capsid proteins such as SV40 VP1 can be isolated from cells as pentameric capsomeres and assembled with the addition of DNA. As a trade-off however, these VLPs have inherently less tolerance to the wide range of non-native conditions required for different biotechnological applications. Furthermore, many of these VLPs are reliant on negatively-charged cargo (typically nucleic acids) as a scaffold to impart stability, hence are not easily able to package other diverse forms of synthetic cargo in a selective and robust manner.Example 7: MBP-Enc and LanM-Enc fusions

[0231] Maltose-binding-protein or MBP is a common bulky protein generally used as a solubility tag to help express proteins recombinantly. Here the inventors have demonstrated that MBP may be used to prevent pre-assembly of the encapsulin (Figure 7). In some instances, MBP fusions may monomerise the encapsulins better and behave more consistently than LanM fusions. MxEnc and TmEnc (Figures 8 and 9) are encapsulins of different sizes and this example demonstrates that the invention is applicable to encapsulins from different species and of different sizes.Example 8: Cleavage and assembly of Enc fusion proteins at a range of temperatures

[0232] A range of temperatures were tested for cleavage and assembly of encapsulin fusion proteins. The same protocol as previously described was used at a range of temperatures for TEV cleavage and encapsulation. The inventors show that the cleavage and assembly also occur at a range of temperatures (shown in Figures 10A and B).Example 9: Fusions of mutant Enc and assembly of hybrid protein cages

[0233] The inventors also generated and tested MBP fusions comprising encapsulin mutants Qt-Glass9, Qt-Letter11, Qt-Vegeta10, Qt-Pigs13, and Qt-Slay13 for their ability to form protein cages (Figure 11). When the MBP protein has been cleaved, a variant that assembles into a protein cage (or a hybrid) will elute very early in the size exclusion chromatography profile because it is of a larger size. Encapsulins that remain in monomer form elute later because they are of smaller size. This is marked on Figures 11 and 12. The inventors demonstrated that Qt-Glass9 and Qt-Letter11 variants formed encapsulin protein cages independently, but Qt-Vegeta10, Qt-Pigs13, and Qt-Slay13 variants did not. The inventors also demonstrated that hybrid protein cages can be formed by combining an encapsulin that is able to independently form protein cages with a second encapsulin (or variant). Figure 12 demonstrates that hybrid protein cages may be formed regardless of the ability of the second encapsulin’s ability to independently form protein cages.Example 10: Co-packaging of cargo in protein cages

[0234] The inventors also demonstrated that multiple cargo molecules (such as mNeon and mCherry) may be co-packaged into protein cages. Figure 13 visually represents cargo packaging and co-packaging by detection of mNeon and mCherry fluorescence.Table 4: Calculated concentrations of co-packaged mNeonQtTP and mCherryQtTP in assembled encapsulin protein cage sampleSample name A506 A590 [mNeonQtTP] (pM) [mCherryQtTP) (pM)mNeon 100% 0.3 0 2.6 0mCherry 100% 0 0.135 0 1.950 / 50 0.19 0.09 1.6 1.3mNe80 / mCh20 0.242 0.08 2.1 1.1mNe20 / mCh80 0.117 0.096 1 1.3

[0235] Table 4 shows the sample names for each co-packaged encapsulin sample comprising varying ratios of mNeonQtTP (mNeon conjugated to the CLP of QtEnc) and mCherryQtTP (mCherry conjugated to the CLP of QtEnc). The absorbance values at 506 nm and 590 nm were measured and recorded for each sample and these were used to calculate the concentration of each protein using the extinction coefficients for mNeon A506 (£ = 116,000 M ) and mCherry A587(e = 72,000 MW ).Table 5: Percentage of mNeonQtTP and mCherryQtTP in each co-packaged sample.Sample name %mNeonQtTP %mCherryQtTPmNeon 100% 100 0mCherry 100% 0 10050 / 50 55 45mNe80 / mCh20 66 34mNe20 / mCh80 43 57

[0236] Table 5 shows the percentage of mNeonQtTP and mCherryQtTP present in each sample according to the concentrations calculated in Table 4.

[0237] In these experiments, the inventors performed the in vitro assembly and packaging protocol as previously described. Instead of a single cargo protein, the inventors used mNeonQtTP (a green fluorescent protein) and mCherryQtTP (a pink fluorescent protein) in varying ratios to demonstrate co-packaging. The inventors purified the encapsulin samples comprising co-packaged mNeonQtTP and mCherryQtTP and used their unique absorbance properties to determine the concentration of each fluorescent protein (Table 4) and the corresponding percentage of each protein in the sample (Table 5) to assess for alignment with the desired ratio.

[0238] It will be understood that the invention disclosed and defined in this specification extends to all alternative combinations of two or more of the individual features mentionedor evident from the text or drawings. All of these different combinations constitute various alternative aspects of the invention.

Claims

CLAIMS1. A fusion protein comprising:a first polypeptide comprising an amino acid sequence of an encapsulin, and a second polypeptide comprising an amino acid sequence of lanmodulin (LanM) and / or maltose binding protein (MBP)wherein the fusion protein comprises a cleavable sequence for enabling cleavage of the encapsulin from the amino acid sequence of the second polypeptide.

2. The fusion protein of claim 1, wherein the encapsulin is a Family 1 encapsulin, preferably from Quasibacillus thermotolerans (QtEnc), Thermotoga maritima (TmEnc), Myxococcus xanthus (MxEnc), Dendrosporobacter quercicolus (DgEnc), and Bacillus methanolicus (BmEnc).

3. The fusion protein of claim 1 or 2, wherein the amino acid sequence of an encapsulin is an amino acid sequence as set forth in SEQ ID NO: 6, 25 to 28, or 58 to 62, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto.

4. The fusion protein of any one of the preceding claims, wherein the second polypeptide comprises lanmodulin (LanM).

5. The fusion protein of claim 4, wherein the amino acid sequence of the second polypeptide is an amino acid sequence as set forth in SEQ ID NO: 7, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto.

6. The fusion protein of any one of the preceding claims, wherein the fusion protein comprises:an amino acid sequence of SEQ ID NO: 6, 25 to 28, or 58 to 62, or a functionally equivalent homolog thereof having at least 80%, 85%, 90%, or 95% identity thereto; andan amino acid sequence of SEQ ID NO: 7, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto, wherein the fusion protein comprises a cleavable sequence for enabling cleavage of the amino acid sequence of the encapsulin from the amino acid sequence of SEQ ID NO: 7.

7. The fusion protein of any one of claims 1-3, wherein the second polypeptide comprises maltose binding protein (MBP).

8. The fusion protein of claim 7, wherein the amino acid sequence of MBP is an amino acid sequence as set forth in SEQ ID NO: 15, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto.

9. The fusion protein of any one of claims 1-3 or 7-8, wherein the fusion protein comprises:an amino acid sequence of SEQ ID NO: 6, 25 to 28, or 58 to 62, or a functionally equivalent homolog thereof having at least 80%, 85%, 90%, or 95% identity thereto; andan amino acid sequence of SEQ ID NO: 15, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto, wherein the fusion protein comprises a cleavable sequence for enabling cleavage of the amino acid sequence of the encapsulin from the amino acid sequence of SEQ ID NO: 15.

10. The fusion protein of any one of the preceding claims, wherein the N-terminus of the encapsulin sequence is joined or linked to the C-terminus of the second polypeptide sequence.

11. The fusion protein of any one of the preceding claims, wherein the encapsulin and second polypeptide sequences are joined via a linker that comprises the cleavable sequence.

12. The fusion protein of any one of the preceding claims wherein the fusion protein further comprises a flexible linker linking the encapsulin and second polypeptide sequences.

13. The fusion protein of any one of claim 1-6 or 10-12, wherein the fusion protein comprises the amino acid sequence as set forth in SEQ ID NO: 8, 9, 37 to 40, or 78 to 82.

14. The fusion protein of any one of claims 1-3 or 7-12, wherein the fusion protein comprises the amino acid sequence as set forth in SEQ ID NO: 16, 17, 41 to 44, or 6815. A nucleic acid comprising or consisting of a nucleotide sequence encoding the fusion protein of any one of claims 1 to 14.

16. The nucleic acid of claim 15, wherein the sequence comprises a sequence as set forth in any of SEQ ID NOs: 1, 2, 13, 14, 21 to 24, 29 to 36, 53 to 57, 63 to 67, or 73 to 77, or a sequence having at least 80%, 85%, 90%, or 95% identity thereto, and encoding a functionally equivalent variant.

17. A nucleic acid construct or vector comprising the nucleic acid of claim 15 or 16.

18. A host cell comprising the nucleic acid construct or vector of claim 17.

19. A method for forming a protein cage, the method comprising:contacting a fusion protein of any one of claims 1-14 with an agent for cleaving the cleavable sequence of the fusion protein to cleave the second polypeptide from the fusion protein,thereby forming a protein cage.

20. A method for forming a hybrid protein cage, the method comprising contacting a first fusion protein of any one of claims 1-14 and:a polypeptide comprising an encapsulin ora second fusion protein of any one of claims 1-14with at least one agent for cleaving the cleavable sequences of the fusion protein(s) to cleave the second polypeptides from the fusion protein(s),thereby forming a hybrid protein cage.

21. A method for rescuing or repairing a defective and / or partially assembled encapsulin protein cage, the method comprising:providing a defective and / or partially assembled encapsulin protein cage; contacting a fusion protein of any one of claims 1-14 with (i) an agent for cleaving the cleavable sequence of the fusion protein to cleave the second polypeptide from the fusion protein and (ii) the protein cage,thereby rescuing or repairing a defective and / or partially assembled encapsulin protein cage.

22. A method for the encapsulation of at least one cargo molecule within a protein cage, the method comprising:conjugating at least one cargo molecule to at least one cargo loading moiety (CLM),contacting a fusion protein of any one of claims 1-14 with an agent for cleaving the second polypeptide from the fusion protein,then contacting the cleaved fusion protein with the at least one cargo molecule conjugated to the CLM under suitable conditions and for a sufficient time to enable formation of in vitro complexes of encapsulin (eg QtEnc) and the cargo molecule,thereby encapsulating the at least one cargo molecule within a protein cage.

23. A method for the encapsulation of at least one cargo molecule within a hybrid protein cage, the method comprising:conjugating at least one cargo molecule to at least one cargo loading moiety (CLM),contacting a first fusion protein of any one of claims 1-14 and:o polypeptide comprising an encapsulin oro a second fusion protein of any one of claims 1-14with at least one agent for cleaving the cleavable sequences of the fusion protein(s) to cleave the second polypeptides from the fusion protein(s), then contacting the cleaved fusion protein(s) with the at least one cargo molecule conjugated to the CLM under suitable conditions and for a sufficient time to enable formation of in vitro complexes of encapsulin (eg QtEnc) and the at least one cargo molecule,thereby encapsulating the at least one cargo molecule within a hybrid protein cage.

24. The method of any one of claims 19-23, wherein the fusion protein, first fusion protein, and / or second fusion protein comprises an amino acid sequence of SEQ ID NO: 6, 25 to 28, 58 or 59, or a functionally equivalent variant thereof having at least 80%, 85%, 90%, or 95% identity thereto.

25. The method of any one of claims 19-24, wherein the method is performed in vitro.

26. A method for the delivery of at least one cargo molecule to a target cell, the method comprising:encapsulating the at least one cargo molecule in a protein cage or hybrid protein cage obtained according to the method of any one of claims 22-25, to obtain a cargo-protein cage assembly,contacting the target cell with the cargo-protein cage assembly under suitable conditions for enabling uptake or internationalisation of the protein cage assembly by the target cell,thereby delivering a cargo molecule to a target cell.

27. The method of any one of claims 22-26, wherein the at least one cargo is selected from: a small molecule, a nucleic acid or polynucleotide (such as a DNA, RNA, siRNA or shRNA molecule), a carbohydrate, a peptide, a protein or a synthetic molecule (such as a synthetic polymer).

28. The method of any one of claims 22-27, wherein the CLM comprises the cargo loading peptide (CLP) for Q. thermotolerans, Thermotoga maritima, Myxococcus xanthus, Dendrosporobacter quercicolus, and / or Bacillus methanolicus.

29. The method of any one of claims 22-28, wherein the CLM comprises an amino acid sequence of SEQ ID NO: 18 to 20 and / or 45 to 52.

30. The method of any one of claims 22-29, wherein the conjugation of a cargo molecule to a CLM is by conjugation of the cargo molecule to the N-terminal of the CLM.

31. The method of any one of claims 22-30, wherein the conjugation of a cargo molecule to a CLM is by conjugating a maleimide-functionalised cargo molecule to a Cys-modified CLM.

32. The method of any one of claims 19-31, wherein the method is performed in the absence of lanthanides.

33. A protein cage or hybrid protein cage obtained according to the method of any one of claims 19-32.

34. A composition comprising a protein cage or hybrid protein cage of claim 33 and a pharmaceutically acceptable carrier or adjuvant.