Capsids with effector-polypeptide conjugates and uses thereof
Capsids containing endogenous Gag polypeptides and effector-polypeptide conjugates address the limitations of existing nucleic acid and therapeutic agent delivery methods by enhancing delivery efficiency and specificity, offering a promising alternative for targeted and effective therapy.
Patent Information
- Application Number
- PCT/US2024/054408
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-03
- Filing Date
- 2024-11-04
- Publication Date
- 2025-05-08
AI Technical Summary
Current methods for delivering nucleic acids and therapeutic agents face challenges such as immunogenicity, limited payload capacity, poor bio-distribution, and preexisting immunity, making them inefficient and limited in therapeutic applications.
Development of capsids containing an endogenous Gag polypeptide and effector-polypeptide conjugates, which can encapsulate or be associated with heterologous cargoes like mRNA or siRNA, to enhance delivery efficiency and specificity.
The capsid-based delivery system improves the efficiency and specificity of nucleic acid and therapeutic agent delivery, overcoming limitations of existing methods by providing a non-immunogenic, high-capacity, and targeted delivery approach.
Smart Images

Figure US2024054408_08052025_PF_FP_ABST
Abstract
Description
CAPSIDS WITH EFFECTOR-POLYPEPTIDE CONJUGATESAND USES THEREOF
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 596,071, filed November 3, 2023 and U.S. Provisional Application No. 63 / 596,047, filed November 3, 2023, both of which are hereby incorporated by reference.FIELD OF THE INVENTION
[0002] The present invention is directed to a capsid containing (a) an endogenous Gag polypeptide, (b) one or more effector-polypeptide conjugates, where each effector-polypeptide conjugate independently comprises an endogenous Gag polypeptide, and (c) optionally a heterologous cargo, such as an mRNA or siRNA. The present invention also relates to an assay that measures the encapsulation efficiency of a polynucleotide in a delivery system.BACKGROUND OF THE INVENTION
[0003] Administering diagnostic or therapeutic agents to a site of interest with precision has presented an ongoing challenge. Available methods of delivering nucleic acids to cells suffer from a number of limitations. For example, AAV viral vectors often used for gene therapy are immunogenic, have a limited payload capacity, suffer from poor bio-distribution, can only be administered by direct injection, and pose a risk of disrupting host genes by integration. Some studies have suggested that a significant proportion (e g., >50%) of the population have preexisting immunity to viral vectors such as AAV, which could severely limit their effectiveness in therapeutic applications. The utility of existing non-viral methods is also restricted by a number of shortcomings. Liposomes can be primarily delivered to the liver. Extracellular vesicles can have a limited payload capacity, limited scalability, and be subject to purification difficulties. Thus, there is a need for new and improved compositions and methods for delivering therapeutic payloads. Capsids (or virus-like particles, VLPs) disclosed herein have the potential to address many of these shortcomings.
[0004] There is a continuing need for improved delivery systems for nucleic acids and therapeutic agents.
[0005] The advent of mRNA use in vaccine has create a flurry of assay developments for determining the amount of mRNA encapsulated within a delivery system. Numerous assays have been used such as HPLC and gel shift methods. While these methods are useful, they are not always efficient. Gel shift assay can lead to inaccurate account for the encapsulation efficiency and HPLC method can be time consuming and laborious. Thus, there remains a need for an assay that can rapidly and efficiently measure the amount of polynucleotide encapsulated inside a delivery system.SUMMARY
[0006] The present invention is directed to a capsid containing an endogenous Gag polypeptide and one or more effector-polypeptide conjugates, where each effector-polypeptide conjugate independently comprises an endogenous Gag polypeptide and an effector moiety. The capsid may have a loaded, or encapsulate or be associated with, a heterologous cargo, such as an mRNA or siRNA.
[0007] One embodiment is a capsid comprising (a) an endogenous Gag polypeptide, (b) one or more effector-polypeptide conjugates, where each effector-polypeptide conjugate independently comprises an endogenous Gag polypeptide and an effector moiety, and (c) a heterologous cargo.
[0008] Another embodiment is a capsid comprising (a) an endogenous Gag polypeptide, and (b) one or more effector-polypeptide conjugates, where each effector-polypeptide conjugate independently comprises an endogenous Gag polypeptide and an effector moiety, and where the capsid is loaded with or encapsulates a heterologous cargo.
[0009] Yet another embodiment is a capsid comprising (a) an endogenous Gag polypeptide, and (b) one or more effector-polypeptide conjugates, where each effector-polypeptide conjugate independently comprises an endogenous Gag polypeptide and an effector moiety, and where a heterologous cargo is associated with the capsid.
[0010] In one embodiment of any of the capsids described herein, the capsid is a protein nanoparticle.
[0011] In one embodiment, the effector moiety is conjugated, directly or indirectly, through the N-terminus of the endogenous Gag polypeptide in the effector-polypeptide conjugate. In general, conjugation at the N-terminus results in the effector moiety being exposed outside the capsid. In another embodiment, the effector moiety is conjugated, directly or indirectly, through the C-terminus of the endogenous Gag polypeptide in the effector-polypeptide conjugate. In general, conjugation at the C-terminus results in the effector moiety being internalized within the capsid.
[0012] In some embodiments, the effector moiety has, or comprises, an amino acid sequence of any one of SEQ ID NOs: 112-122.
[0013] In one embodiment, the effector-polypeptide conjugate further comprises (i) a Spytag conjugated to endogenous Gag polypeptide and (ii) a Spycatcher conjugated to the Spytag and to the effector moiety. The Spy catcher may be conjugated through a linker (e.g., a 4 to 12 amino acid length linker, or a 5 to 10 amino acid length linker, such as a 5 amino acid linker or a 10 amino acid linker) to the endogenous polypeptide.
[0014] In another embodiment, the effector-polypeptide conjugate further comprises a thiol maleimide linker conjugated to the endogenous polypeptide and to the effector moiety.
[0015] In one embodiment, the effector-polypeptide conjugate comprises a first linker between the Gag polypeptide and the effector moiety. The first linker may comprise 5xGS or EAAK.
[0016] In one embodiment, the one or more effector-polypeptide conjugates are selected from (i) a targeted conjugate comprising an endogenous Gag polypeptide and a targeting moiety, (ii) an endosomal escape conjugate comprising an endogenous Gag polypeptide and an endosomal escape moiety, or (iii) any combination of any of the foregoing. The endosomal escape moiety may be conjugated, directly or indirectly, through the N-terminus of the endogenous Gag polypeptide. In an alternative embodiment, the endosomal escape moiety is conjugated, directly or indirectly, through the C-terminus of the endogenous Gag polypeptide. In some embodiments, the endosomal escape moiety is conjugated to the remainder of the effector- polypeptide conjugate through a cleavable linker. In one embodiment, the endosomal escape moiety may be released from the remainder of the effector-polypeptide conjugate upon entry into a cell. In some embodiments the cleavable linker is cleaved after the capsid is endocytosed into the cell. In some embodiments, the cleavable linker is cleaved during endosomal escape.
[0017] In one embodiment, the targeting moiety enhances cell-specific uptake. In another embodiment, the targeting moiety enhances nuclear localization. The targeting moiety may be a targeting polypeptide.
[0018] In one embodiment, the targeting moiety comprises a binding domain specific for a target cell of interest. In some embodiments, the binding domain comprises a receptor, an antibody, or an antigen-binding fragment. In some embodiments, the antibody fragment is selected from a Fab, a Fab’, a F(ab’)2, an Fd, an Fv, a domain antibody, a complementarity determining region (CDR), a single chain variable fragment antibody (scFv), a maxibody (e.g., scFv-Fc), a minibody (e.g., VL-VH-CHs), an intrabody, a diabody, a triabody, a tetrabody, a v- NAR (variable domain of the new antigen receptor) and a bis-scFv.
[0019] In one embodiment, the targeting moiety comprises a tag, and a binding domain is conjugated to the polypeptide through the tag. In some embodiments, the tag is selected from a SNAP tag, a biotin tag, a monomeric streptavidin, a monomeric streptavidin 2, an intein, a SunTag, an Isopeptag, a SpyTag, a SpyCatcher tag, a SnoopTag, a SnoopTagJr, a SnoopCatcher tag, a DogTag, a DogCatcher tag, a Gluthatione-S-transferase tag, a CLIP tag, a Protein A tag, a Protein G tag, a Protein AG tag, a GFP tag, an HA tag, a FLAG tag and a HiBiT-tag.
[0020] In one embodiment, the targeting polypeptide comprises one or more of LDL- receptors, apolipoprotein mimetic peptide (such as an apoA-I, apoE, or apoC-II peptide), PLA2, or any combination of any of the foregoing. In another embodiment, the targeting polypeptide is an LDL receptor selected from LRKLRK, LRKRLLRD, and LKAYKS.
[0021] In one embodiment, the targeting moiety is an antibody (e.g., a monoclonal antibody), fragment antigen-binding (Fab) protein (such as a CD71 Fab or CD117 Fab), or single-chain variable fragment (scFv). In one embodiment, the targeting moiety targets CD71. In another embodiment, the targeting moiety targets CD117.
[0022] In one embodiment, the endosomal escape conjugate comprises a first linker between the Gag polypeptide and the endosomal escape moiety. The first linker may comprise 5xGS or EAAK.
[0023] In another embodiment, the targeted conjugate comprises a first linker between the Gag polypeptide and the targeting moiety. The first linker may comprise 5xGS or EAAK.
[0024] In one embodiment of any capsid described herein, the endogenous Gag polypeptide is a native endogenous Gag polypeptide, such as dArc.
[0025] In another embodiment of any capsid described herein, the endogenous Gag polypeptide is an engineered endogenous Gag polypeptide.
[0026] In one embodiment of any capsid described herein, the endogenous Gag polypeptide is an Arc polypeptide, such as dArc.
[0027] In one embodiment, the endogenous Gag polypeptide comprises a sequence modification, wherein the sequence modification comprises a cargo binding domain, nucleic acid binding domain, zinc finger domain, sub-cellular localization signal, a nuclear localization signal, or an antibody or antigen-binding fragment thereof.
[0028] In one embodiment of any capsid described herein, the weight ratio of (a) an endogenous Gag polypeptide to (b) effector-polypeptide conjugate ranges from about 90: 10 to about 10:90, or from about 90: 10 to about 50:50. In some embodiments the weight ratio of (a) an endogenous Gag polypeptide to (b) effector-polypeptide conjugate ranges from about 98:2 to about 80:20. In some embodiments the weight ratio of (a) an endogenous Gag polypeptide to (b) effector-polypeptide conjugate is about 95:5 to about 85: 15.
[0029] In one embodiment of any capsid described herein, the cargo (a) comprises a nucleic acid (such as RNA or DNA), (b) comprises or encodes a gene editing system or a component thereof (e.g., a CRISPR / Cas system or a component thereof), (c) polypeptide, (d) a therapeutic agent, (e) an antibody or antigen-binding fragment thereof, a peptidomimetic, a nucleotidomimetic, a drug, a diagnostic tool, an imaging tool, a small molecule, or a combination of any one of the foregoing. In one preferred embodiment, the heterologous cargo is a nucleic acid.
[0030] Yet another embodiment is a method of preparing a capsid as described herein, where the method comprises the step of mixing (a) an endogenous Gag polypeptide, (b) one or more effector-polypeptide conjugates, where each effector-polypeptide conjugate independently comprises an endogenous Gag polypeptide and an effector moiety, and (c) a heterologous cargo.
[0031] The present invention also provides a method that measures the encapsulation efficiency of a polynucleotide in a delivery system. The invention disclosed a method and a kit for determining the amount of polypeptide encapsulated inside a nanoparticle. In one aspect, the invention provides a method of determining the encapsulation efficiency of a polynucleotide in a delivery system comprising the steps of a) incubating a sample containing both free polynucleotide and encapsulated polynucleotide with a probe; b) measuring the freepolynucleotide with the probe; c) treating the sample in (b) with a particle disrupting agent; d) incubating the sample in (c) with the probe; f) measuring the encapsulated polynucleotide with the probe.BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Various aspects of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings below.
[0033] FIG. 1 is a diagram showing the preparation of PNPs described herein.
[0034] FIG. 2 is examples of PNP endosomal escape and targeting moi eties.
[0035] FIG. 3 shows a thiol-maleimide linking system and its attachment to an effector moiety R. An endogenous Gag polypeptide is conjugated through the maleimide group.
[0036] FIG. 4 shows the Spy Tag- Spy Catcher system for conjugation of a ligand (such as an effector moiety) to an endogenous Gag polypeptide.
[0037] FIG. 5 is different examples of PNP functionality.
[0038] FIG. 6 is a molecular beacon assay of the dArc capsid at different pH levels.
[0039] FIG. 7 is a molecular beacon of dArc conjugated to a LDLR-Ligand at different weight ratios.
[0040] FIG. 8 is a molecular beacon assay of dArc conjugated to different LDLR amino acid sequences through various linker sequences.
[0041] FIG. 9 is a qPCR analysis of RNA uptake in cells of the dArc PNP conjugated to the LRKLRK amino acid sequence 2, 4, and 6 hours after introduction.
[0042] FIG. 10 is SLEEQ analysis of endosomal escape of dArc-PNP conjugated with LDLR targeting moiety.
[0043] FIG. 11 A-F is the co-formulating tuning for the ability of dArc and PLA2 to form capsids.
[0044] FIG. 12 A-C is the engineered dArc-PNP conjugated to PLA2 targeting moiety and capsid formation.
[0045] FIG. 13 is the engineered dArc-PNP conjugated to PLA2 enzymatic activity assay
[0046] FIG. 14 is the Engineered dArc-PNP with parvovirus PLA2 enhances cellular uptake of mRNA in HepG2 cells.
[0047] FIG. 15 is ELISA analysis of dArc-PNP conjugated with nothing, to LDLR, to PLA2, or both.
[0048] FIG. 16A-C is SpyCatcher Screening of phospholipases to determine the best weight formulation for dArc-SpyCatcher conjugated with SpyTag-PLA2.
[0049] FIG. 17 is blot analysis of the dArc-SpyCatcher / Spytag-PLA2 PNP
[0050] FIG. 18 is blot analysis of the dArc-SpyCatcher / Spytag-PLA2 PNP
[0051] FIG. 19 is mRNA entrapment at different weight ratios of dArc-SpyCatcher / SpyTag-PLA2.
[0052] FIG. 20 is molecular beacon assay at different weight ratios of dArc- SpyCatcher / SpyTag PLA2.
[0053] FIG. 21A-B is the assembly of dArc-PNP / PLA2 / HiBit
[0054] FIG. 22 is the weight ratio of dArc conjugated to HiBit at a ratio of 90 / 10 for capsid formation.
[0055] FIG. 23 is the weight ratio of dArc conjugated to HiBit at a ratio of 95 / 5 for capsid formation.
[0056] FIG. 24A-B is the protection capabilities of dArc / PLA2 / HiBit at different weight ratios in molecular beacon assays.
[0057] FIG. 25 is a SLEEQ assay measuring the endosomal escape capabilities of dArc / PLA2 / HiBit PNP.
[0058] Fig. 26 shows the number of clusters after computational analysis of the amount of viral families PLA2 is a part of.
[0059] FIG. 27 is the cluster analysis showing 5 distinct clusters of viral families with potential PLA2 usage.
[0060] FIG. 28 is strategies for releasing PLA2 from PNPs using cleavable linkers.
[0061] FIG. 29 is strategies for screening cleavable linkers to release PLA2 from PNPs.
[0062] FIG. 30 is strategies for engineering PLA2 inside of PNPs.
[0063] FIG. 31 is the particle formation capabilities of the SpyCatcher system with 5 and 10 amino acid linkers conjugated to the PNP.
[0064] FIG. 32 is the entrapment percentage of mRNA within dArc-SpyCatcher PNPs with different length amino acid linkers.
[0065] FIG. 33 is a diagram of the PNP-FAB conjugation.
[0066] FIG. 34 is the particle formation of the dArc-SpyTag with CD71FAB-SpyCatcher.
[0067] FIG. 35 is the mRNA entrapment percentage with dArc-SpyTag-BAB +dArc RGG at different ratios.
[0068] FIG. 36 is the cellular uptake of the dArc-SpyTag / CD71FAB-SpyCatcher in cells that have varying levels of CD71 receptor density.
[0069] FIG. 37 is the cellular uptake of dArc-SpyTag / CD71FAB-SpyCatcher in cells that have varying levels of CD71 receptor density analyzed this through flow cytometry.
[0070] FIG. 38 is a bar graph showing the results of a flow cytometry analysis of fluorescent mRNA uptake in CD71+ cells using PNPs that have FABCD71at different weight ratios.
[0071] FIG. 39 is a diagram showing the method measuring packaging efficiency of a nanoparticle delivery system.
[0072] FIG. 40 illustrates the effect of incubation temperature, incubation time and Mg2+concentration in the buffer on the binding of molecular beacon (MB) with the target nucleotide acid.
[0073] FIG. 41 shows the amount optimization of molecular beacon.
[0074] FIG. 42 depicts the consistence of assay readout under different conditions.
[0075] FIG. 43 shows the sensitivity of the assay.
[0076] FIG. 44 depicts the different length for stretches of T nucleotides used in the assay.
[0077] FIG. 45 illustrates the different buffer conditions for the assay.
[0078] FIG. 46 illustrates the efficiency of the assay over a gel shift assay.
[0079] FIG. 47 represents the measurement of packaging efficiency of assembledAlphavirus capsid.DETAILED DESCRIPTION OF THE INVENTION
[0080] The present invention is directed to a capsid containing an endogenous Gag polypeptide and one or more effector-polypeptide conjugates, where each effector-polypeptide conjugate independently comprises an endogenous Gag polypeptide and an effector moiety. Thecapsid may have a loaded, or encapsulate or be associated with, a heterologous cargo, such as an mRNA or siRNA.
[0081] Fig. 1 shows one embodiment of the preparation of a capsid described herein. The endogenous Gag polypeptide (referred to as a capsid protein in the figure) may be mixed with a cargo and one or both of (i) a targeted conjugate (referred to as a capsid protein + targeting moiety) and (ii) an endosomal escape conjugate (referred to as a capsid protein + endosomal escape moiety). The cargo is shown as a squiggly line on the right.
[0082] Fig. 2 lists moieties useful as targeting moieties and endosomal escape moieties. In some embodiments, the targeting moiety is an antibody or small molecule that targets a receptor in skeletal muscles such as MINAT2, CACNG1, KCNA7, RYR1, STRIT1, CHRNG, SYPL2, CACNG6, LRRC38, ENSG00000250349, SCN4A, GPR85, LSMEM1, MYADML2, CLCN1, TMEM201, SHISA4, or TMEM233. In some embodiments, the targeting moiety is an antibody or small molecule that targets the receptor KCNA7, GTEx, FANTOM5, HP A, CACNG1, SYPL2, CHRNG, or CHRNA1. In some embodiments, the targeting moiety is an antibody or small molecule that targets a receptor in heart muscle such as MYLK3, SGCG, HHATL, FITM1, SLC25A4, PLN, POPDC2, CPT1B, BVES, SSPN, SLC41A1, ITGA7, ARL6IPS PKD1L1, PRKAA2, SGCD, TIMM22, or LSMEM2. In some embodiments, the targeting moiety is an antibody or small molecule that targets a receptor in endothelial cells such as CLCN5, APLNR, or AQP1. In some embodiments, the targeting moiety is an antibody or small molecule that targets a receptor in the brain, such as NRN1, TSPAN7, or SEMA6B.
[0083] Fig. 5 shows the various components which may be conjugated to the endogenous Gag polypeptides (represented by gray trapezoids) and their arrangement.
[0084] A linker can be used to conjugate the effector moiety, such as an endosomal escape moiety (e.g., PLA2), to the polypeptide in the effector-polypeptide conjugate. The linker can include a cleavable linker as shown in Figs. 28 and 30. In some instances, the cleavable linker(s) include cathepsin, MBP, or TrxA. In one preferred embodiment, the endosomal escape moiety is conjugated through a cleavable linker to the remainder of the effector-polypeptide conjugate or polypeptide.
[0085] In some instances, an affinity tag (such as a Twinstep) is conjugated to the effector moiety, such as a targeting moiety (for example, to facilitate purification as shown in Fig. 29).
[0086] The effector-polypeptide conjugate may include a cargo binding domain, such as an RNA binding domain (e.g., RGG or RBD), which permits the cargo to bind to the effector- polypeptide conjugate. In one preferred embodiment, the cargo binding domain is conjugated to the remainder of the effector-polypeptide conjugate through a cleavable linker, such as those described herein. Suitable cargo binding domains include those described herein. In one embodiment, the cargo binding domain is conjugated, directly or indirectly, through the C- terminus of the endogenous Gag polypeptide in the effector-polypeptide conjugate.
[0087] The capsids disclosed here in, in certain embodiments, are retroviral capsids. In some embodiments, endogenous Gag (endo-Gag) polypeptides of the disclosure, such as Arc, RTL and PNMA family polypeptides, assemble into capsids for delivery of a cargo of interest. Endo-Gag polypeptides of the disclosure exhibit advantages over existing and alternate capsid-forming polypeptides, such as improved efficiency in capsid assembly, capsid disassembly, and / or capsid reassembly, e.g., reassembly with a heterologous cargo. In additional embodiments, described herein are capsids, e.g., RTL family-based, RTLIO-based, PEGlO-based, PNMA family-based, PNMA2 -based, PNMA5-based, or other endo-Gag-based capsids, for delivery of a cargo of interest. Also disclosed herein are engineered endo-Gag polypeptides, such as RTL or PNMA family polypeptides, with additional modifications of particular advantage and utility.
[0088] The sequences referenced herein are provided in Table A at the end of the specification.
[0089] The disclosure of International Application No. PCT / US2023 / 070915, filed July 25, 2023, and International Application No. PCT / US2022 / 013954, filed January 26, 2022, published as International Publication No. WO 2022 / 164942, each of which is hereby incorporated by reference in its entirety, including its disclosure of capsids and components thereof as well as their preparation and uses.
[0090] Disclosed herein, in certain embodiments, are endogenous retroviral capsid polypeptides. In some embodiments, endogenous Gag (endo-Gag) polypeptides of the disclosure, such as PNMA family polypeptides and retrotransposon Gag-like (RTL) family polypeptides, assemble into capsids for delivery of a cargo of interest. Endo-Gag polypeptides of the disclosure exhibit surprising and unexpected advantages over existing and alternate capsid-forming polypeptides, such as improved efficiency in capsid assembly, capsid disassembly, and / or capsid reassembly, e.g., reassembly with a heterologous cargo. In additional embodiments, describedherein are capsids, e g., PEGlO-based, RTLIO-based, PNMA-based or endo-Gag-based capsids, for delivery of a cargo of interest. Also disclosed herein are engineered endo-Gag polypeptides, such as PEGlO-based, RTLIO-based, other endo-Gag-based, or PNMA family polypeptides, with additional modifications of particular advantage and utility.
[0091] Expressing the endogenous Gag polypeptide as a conjugate with, for example, a bulky solubility tag, can prevent premature capsid formation and increases the efficiency of forming capsids with desired cargos.
[0092] One embodiment is a method of preparing a capsid comprising (a) expressing an endogenous Gag polypeptide conjugate in a host cell or a cell-free expression system, the endogenous Gag polypeptide conjugated to a N-terminal component which prevents, inhibits, or suppresses oligomerization of the endogenous Gag polypeptide; (b) optionally, isolating the endogenous Gag polypeptide conjugate; (c) cleaving the N-terminal tag from the endogenous Gag polypeptide conjugate to form cleaved endogenous Gag polypeptide; and (d) treating the cleaved endogenous Gag polypeptide with an assembly buffer. In one embodiment, the method further comprises contacting the endogenous Gag polypeptide with a cargo during step (d), where at least a subset of the endogenous Gag capsids comprise the cargo. In another embodiment, the method further comprises treating the isolated endogenous Gag polypeptide with a disassembly buffer prior to step (d), thereby producing or maintaining the endogenous Gag polypeptide in a disassembled state. In yet another embodiment, less than 20%, less than 10%, or less than 5% of the endogenous Gag polypeptide conjugate forms a capsid prior to step (c).
[0093] Another embodiment is a method of preparing a capsid comprising (a) cleaving the N-terminal tag from an endogenous Gag polypeptide conjugate to form cleaved endogenous Gag polypeptide, where the endogenous Gag polypeptide conjugate comprises an endogenous Gag polypeptide conjugated to a N-terminal component which prevents, inhibits, or suppresses oligomerization of the endogenous Gag polypeptide; (b) contacting the cleaved endogenous Gag polypeptide with a cargo; and (c) treating the cleaved endogenous Gag polypeptide with an assembly buffer, thereby assembling the cleaved endogenous Gag polypeptide into endogenous Gag capsids, where at least a subset of the endogenous Gag capsids comprise the cargo. In one embodiment, the method further comprises treating the endogenous Gag polypeptide with a disassembly buffer prior to step (b) or (c), thereby producing or maintaining the endogenous Gagpolypeptide in a disassembled state. In another embodiment, less than 20%, less than 10%, or less than 5% of the endogenous Gag polypeptide conjugate forms a capsid prior to step (a).
[0094] In any of the methods of preparation described herein, the N-terminal component of the endogenous Gag polypeptide conjugate may comprise (i) a cleavage site directly attached to the endogenous Gag polypeptide and (ii) a bulky solubility tag which prevents, inhibits, or suppresses oligomerization of the endogenous Gag polypeptide. In one embodiment, the bulky solubility tag comprises a polypeptide. In another embodiment, the bulky solubility tag increases the solubility of the endogenous Gag polypeptide. In yet another embodiment, the bulky solubility tag sterically blocks assembly of the endogenous Gag polypeptide. In yet another embodiment, the bulky solubility tag is selected from NusA (SEQ ID NO: 93), beta-lactamase (SEQ ID NO: 94), elongation factor Ts (SEQ ID NO: 95), peptidylprolyl isomerase (SEQ ID NO: 96), SUMO (SEQ ID NO: 97), bdSUMO, maltose binding protein (SEQ ID NO: 98 or SEQ ID NO: 64), and TriggerFactor (SEQ ID NO: 99). The cleavage site may be selected from a TEV, HRV3C, thrombin, enterokinase, factor Xa, or carb oxy peptidase cleavage site (e.g., a TEV cleavage site).
[0095] Yet another embodiment is a kit comprising (a) a first container containing a composition comprising an endogenous Gag polypeptide conjugate where the endogenous Gag polypeptide conjugate comprises an endogenous Gag polypeptide conjugated to a N-terminal component which prevents, inhibits, or suppresses oligomerization of the endogenous Gag polypeptide; and (b) a second container containing a protease for cleaving the N-terminal component from the endogenous Gag polypeptide. In the kit, the N-terminal component of the endogenous Gag polypeptide conjugate may comprise (i) a cleavage site directly attached to the endogenous Gag polypeptide and (ii) a bulky solubility tag which prevents, inhibits, or suppresses oligomerization of the endogenous Gag polypeptide. In one embodiment, the bulky solubility tag comprises a polypeptide. In another embodiment, the bulky solubility tag increases the solubility of the endogenous Gag polypeptide. In yet another embodiment, the bulky solubility tag sterically blocks assembly of the endogenous Gag polypeptide. In yet another embodiment, the bulky solubility tag is selected from NusA (SEQ ID NO: 93), beta-lactamase (SEQ ID NO: 94), elongation factor Ts (SEQ ID NO: 95), peptidylprolyl isomerase (SEQ ID NO: 96), SUMO (SEQ ID NO: 97), bdSUMO, maltose binding protein (SEQ ID NO: 98), andTriggerFactor (SEQ ID NO: 99). The cleavage site may be selected from a TEV, HRV3C, thrombin, enterokinase, factor Xa, or carboxypeptidase cleavage site (e.g., a TEV cleavage site).
[0096] Disclosed herein, in some aspects, is a method of making a capsid, the method comprising: (a) expressing an endogenous Gag polypeptide in a host cell or a cell-free expression system; (b) isolating the endogenous Gag polypeptide; and (c) treating the isolated endogenous Gag polypeptide with an assembly buffer, thereby assembling the isolated endogenous Gag polypeptide into endogenous Gag capsids.
[0097] In some embodiments, the method further comprises contacting the endogenous Gag polypeptide with a cargo during step (c), wherein at least a subset of the endogenous Gag capsids comprise the cargo. In some embodiments, the method further comprises treating the isolated endogenous Gag polypeptide with a disassembly buffer prior to step (c), thereby (e.g., disassembling a capsid comprising the endogenous Gag polypeptide) producing or maintaining the endogenous Gag polypeptide in a disassembled state. In some embodiments, the isolated endogenous Gag polypeptide is at least partially present in a capsid form prior to the treating with the disassembly buffer. In some embodiments, at least about 5% of the endogenous Gag polypeptide that is present in capsid form prior to the treating with the disassembly buffer is in the disassembled state after the treating with the disassembly buffer. In some embodiments, at least about 5% of the endogenous Gag polypeptide that is present in the disassembled state is present in capsid form after step (c).
[0098] Disclosed herein, in some aspects, is a method of loading endogenous retroviral capsids with a cargo, the method comprising: treating endogenous Gag polypeptide with a disassembly buffer, thereby (e.g., disassembling a capsid comprising the endogenous Gag polypeptide) producing or maintaining the endogenous Gag polypeptide in a disassembled state; contacting the endogenous Gag polypeptide with a cargo; and treating the endogenous Gag polypeptide with an assembly buffer, thereby assembling the endogenous Gag polypeptide into endogenous Gag capsids, wherein at least a subset of the endogenous Gag capsids comprise the cargo.
[0099] In some embodiments, the endogenous Gag polypeptide is a native endogenous Gag polypeptide. In some embodiments, the endogenous Gag polypeptide is an engineered endogenous Gag polypeptide. In some embodiments, the endogenous Gag polypeptide is a PNMA family protein. In some embodiments, the endogenous Gag polypeptide is a mammalianPNMA family protein. In some embodiments, the endogenous Gag polypeptide is a human PNMA family protein. In some embodiments, the endogenous Gag polypeptide comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55 and 77-92. In some embodiments, the endogenous Gag polypeptide comprises an amino acid sequence with at least 90% sequence identity to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55 and 77-92. In some embodiments, the endogenous Gag polypeptide is PNMA5. In some embodiments, the endogenous Gag polypeptide is PNMA5 and comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of any one of SEQ ID NOs: 5, 60, and 61. In some embodiments, the endogenous Gag polypeptide is PNMA5 and comprises an amino acid sequence with at least 90% sequence identity to any one of SEQ ID NOs: 5, 60, and 61. In some embodiments, the endogenous Gag polypeptide is PNMA5 and comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of SEQ ID NO: 60. In some embodiments, the endogenous Gag polypeptide is PNMA5 and comprises an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 60. In some embodiments, the endogenous Gag polypeptide is PNMA2. In some embodiments, the endogenous Gag polypeptide is PNMA2 and comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of any one of SEQ ID NOs: 1, 7, 8, 38, 40, 43, 46, 48, 49, and 52. In some embodiments, the endogenous Gag polypeptide is PNMA2 and comprises an amino acid sequence with at least 90% sequence identity to any one of SEQ ID NOs: 1, 7, 8, 38, 40, 43, 46, 48, 49, and 52. In some embodiments, the endogenous Gag polypeptide is PNMA2 and comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of SEQ ID NO: 7. In some embodiments, the endogenous Gag polypeptide is PNMA2 and comprises an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 7. In some embodiments, the endogenous Gag polypeptide comprises a sequence modification relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55 and 77-92.
[0100] In some embodiments, the endogenous Gag polypeptide is an RTL family protein. In some embodiments, the endogenous Gag polypeptide is a mammalian RTL family protein. In some embodiments, the endogenous Gag polypeptide is a human RTL family protein. In some embodiments, the endogenous Gag polypeptide comprises an amino acid sequence with at least90% sequence identity to at least 100 consecutive amino acids of any one of SEQ ID NOs: 60-62 and 65-74. In some embodiments, the endogenous Gag polypeptide comprises an amino acid sequence with at least 90% sequence identity to any one of SEQ ID NOs: 60-62 and 65-74. In some embodiments, the endogenous Gag polypeptide is RTL10 (BOP). In some embodiments, the endogenous Gag polypeptide is RTL10 and comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of any one of SEQ ID NOs: 60, 61, or 69. In some embodiments, the endogenous Gag polypeptide is RTL10 and comprises an amino acid sequence with at least 90% sequence identity to any one of SEQ ID NOs: 60, 61, or 69. In some embodiments, the endogenous Gag polypeptide is RTL10 and comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of SEQ ID NO: 60. In some embodiments, the endogenous Gag polypeptide is RTL10 and comprises an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 60. In some embodiments, the endogenous Gag polypeptide is PEG10. In some embodiments, the endogenous Gag polypeptide is PEG10 and comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of any one of SEQ ID NOs: 65-74. In some embodiments, the endogenous Gag polypeptide is PEG10 and comprises an amino acid sequence with at least 90% sequence identity to any one of SEQ ID NOs: 65-74. In some embodiments, the endogenous Gag polypeptide is PEG10 and comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of SEQ ID NO: 65. In some embodiments, the endogenous Gag polypeptide is PEG10 and comprises an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 65. In some embodiments, the endogenous Gag polypeptide comprises a sequence modification relative to any one of SEQ ID NOs: 60-62 and 65-74.
[0101] In some embodiments, the sequence modification comprises an amino acid deletion. In some embodiments, the sequence modification comprises an amino acid insertion. In some embodiments, the sequence modification comprises an amino acid substitution. In some embodiments, the sequence modification comprises a cargo binding domain. In some embodiments, the sequence modification comprises nucleic acid binding domain. In some embodiments, the sequence modification comprises a zinc finger domain. In some embodiments, the sequence modification comprises a sub-cellular localization signal, a nuclear localization signal, or an antibody or antigen-binding fragment thereof. In some embodiments, the sequencemodification is at an N-terminus of the endogenous Gag polypeptide. In some embodiments, the sequence modification is at a C-terminus of the endogenous Gag polypeptide. In some embodiments, the sequence modification within the endogenous Gag polypeptide. In some embodiments, the sequence modification comprises addition of a cysteine residue that is not present in a native form of the endogenous Gag polypeptide or elimination of a cysteine residue that is present in native form of the endogenous Gag polypeptide.
[0102] In some embodiments, the disassembly buffer comprises a reducing agent. In some embodiments, the reducing agent comprises reduced glutathione (GSH), beta mercaptoethanol (P-ME), Dithiothreitol (DTT), or tris(2-carboxyethyl)phosphine (TCEP). In some embodiments, the disassembly buffer comprises the reducing agent at a concentration of about 1 mM to about 20 mM. In some embodiments, the disassembly buffer comprises the reducing agent at a concentration of about 5 mM to about 100 mM. In some embodiments, the disassembly buffer comprises the reducing agent at a concentration of about 5-10 mM. In some embodiments, the disassembly buffer comprises the reducing agent at a concentration of about 10-50 mM. In some embodiments, the disassembly buffer comprises a solubilizing agent. In some embodiments, the solubilizing agent comprises a non-denaturing detergent. In some embodiments, the nondenaturing detergent is CHAPS. In some embodiments, the disassembly buffer comprises about 5-10% CHAPS. In some embodiments, the solubilizing agent comprises urea. In some embodiments, the disassembly buffer comprises about 2-10M urea. In some embodiments, the disassembly buffer comprises a pH of about 7-11. In some embodiments, the disassembly buffer comprises a pH of about 8-10. In some embodiments, the disassembly buffer comprises a pH of about 8. In some embodiments, the disassembly buffer further comprises a divalent cation. In some embodiments, the disassembly buffer further comprises MgC12. In some embodiments, the disassembly buffer comprises CHAPS, TCEP, and optionally Tris. In some embodiments, the disassembly buffer comprises about 5-10% CHAPS, about 10-20mM TCEP, and optionally about 5-100 mM Tris. In some embodiments, the disassembly buffer comprises TCEP, MgCh, and optionally Tris. In some embodiments, the disassembly buffer comprises urea, DTT, optionally NaCl, optionally NaP, and optionally glycerol. In some embodiments, the disassembly buffer comprises about 4-8M urea, about 10-80 mM DTT, optionally about 50-500 mM NaCl, optionally about 5-50 mM NaP, and optionally about 5-15% glycerol. In some embodiments, the disassembly buffer comprises urea, DTT, optionally NaCl, and optionally MgC12. In someembodiments, the assembly buffer contains less than 500 mOsm / kg of salt. In some embodiments, the assembly buffer contains about 270-330 mOsm / kg of salt. In some embodiments, the assembly buffer is a physiological buffer. In some embodiments, the assembly buffer is a phosphate buffered saline. In some embodiments, the assembly buffer comprises NaCl, KC1, Na2HPO4, and KH2PO4. In some embodiments, the assembly buffer comprises about 50 mM to about 750 mM of a monovalent salt. In some embodiments, the assembly buffer comprises about 50 mM to about 750 mM of NaCl. In some embodiments, the assembly buffer comprises glycerol. In some embodiments, the assembly buffer comprises about 5-15% glycerol. In some embodiments, the assembly buffer comprises about 10% glycerol. In some embodiments, the assembly buffer comprises a pH of about 6-9. In some embodiments, the assembly buffer comprises a pH of 7.3-8.5. In some embodiments, the assembly buffer comprises Na2HPO4 / NaH2PO4. In some embodiments, the assembly buffer does not contain a chaotropic agent and / or a reducing agent. In some embodiments, the treating in step (c) comprises incubating in in a dialysis cassette with an about 3,000-10,000 Da molecular cutoff. In some embodiments, the method produces assembled capsids of at least about 50% purity as determined by SDS-PAGE. In some embodiments, the method produces assembled capsids with at least about 50% particle homogeneity as determined by multi-angle dynamic light scattering. In some embodiments, at least about 5% of the isolated endogenous Gag polypeptide assembles to form the capsid. In some embodiments, at least 20% of the isolated endogenous Gag polypeptide assembles to form the capsid. In some embodiments, the isolated endogenous Gag polypeptide that assembles to form the capsid is determined by size exclusion chromatography. In some embodiments, the cargo is a heterologous cargo that is not associated the endogenous Gag polypeptide in nature. In some embodiments, the cargo comprises a nucleic acid. In some embodiments, the cargo comprises an RNA. In some embodiments, the cargo comprises a DNA. In some embodiments, the cargo comprises or encodes a gene editing system or a component thereof. In some embodiments, the cargo comprises or encodes a CRISPR / Cas system or a component thereof. In some embodiments, the cargo comprises a polypeptide. In some embodiments, the cargo comprises a therapeutic agent. In some embodiments, the cargo comprises an antibody or antigen-binding fragment thereof, a peptidomimetic, a nucleotidomimetic, a drug, a diagnostic tool, an imaging tool, a small molecule, or a combination thereof.
[0103] In some embodiments, the disassembly buffer comprises a monovalent salt, a reducing agent, and optionally a basic buffering agent. In some embodiments, the disassembly buffer comprises NaCl, DTT, and optionally Tris. In some embodiments, the disassembly buffer comprises about 25-1800 mM NaCl, 1-20 mM DTT, and optionally 20-75 mM Tris. In some embodiments, the disassembly buffer comprises about 500 mM NaCl, about 5mM DTT, and optionally 50 mM Tris. In some embodiments, the disassembly buffer comprises about 50 mM NaCl, about lOmM DTT, and optionally 25 mM Tris. In some embodiments, the disassembly buffer comprises at least 800 mOsm / kg of salt. In some embodiments, the disassembly buffer comprises less than 150 mOsm / kg of salt. In some embodiments, the assembly buffer contains less than 500 mOsm / kg of salt. In some embodiments, the assembly buffer contains about 270- 330 mOsm / kg of salt. In some embodiments, the assembly buffer comprises a nucleic acid. In some embodiments, the assembly buffer comprises a divalent cation. In some embodiments, the assembly buffer comprises ZnCh. In some embodiments, the assembly buffer comprises MgCh. In some embodiments, the assembly buffer comprises NaCl, ZnCh, and optionally Tris. In some embodiments, the assembly buffer comprises about 150mM NaCl, about lOpM ZnCh, and optionally about 25mM Tris. In some embodiments, the assembly buffer comprises about 150mM NaCl, about ImM DTT, about lOpM ZnCh, and optionally about 50 mM Tris. In some embodiments, the assembly buffer comprises MES and MgCh. In some embodiments, the assembly buffer comprises about 50 mM MES and about 40mM MgCh. In some embodiments, the assembly buffer comprises a pH of about 6-9. In some embodiments, the assembly buffer comprises a pH of about 7.5-8. In some embodiments, the assembly buffer comprises a pH of about 5-7. In some embodiments, the assembly buffer comprises a pH of about 6. In some embodiments, the assembly buffer does not contain a chaotropic agent and / or a reducing agent. In some embodiments, the treating in step (c) comprises incubating in in a dialysis cassette with an about 3,000-10,000 Da molecular cutoff. In some embodiments, the method produces assembled capsids of at least about 50% purity as determined by SDS-PAGE. In some embodiments, the method produces assembled capsids with at least about 50% particle homogeneity as determined by multi-angle dynamic light scattering. In some embodiments, at least about 5% of the isolated endogenous Gag polypeptide assembles to form the capsid. In some embodiments, at least 20% of the isolated endogenous Gag polypeptide assembles to form the capsid. In some embodiments, the percentage of the isolated endogenous Gag polypeptidethat assembles to form the capsid is determined by size exclusion chromatography. In some embodiments, the cargo is a heterologous cargo that is not associated the endogenous Gag polypeptide in nature. In some embodiments, the cargo comprises a nucleic acid. In some embodiments, the cargo comprises an RNA. In some embodiments, the cargo comprises a DNA. In some embodiments, the cargo comprises or encodes a gene editing system or a component thereof. In some embodiments, the cargo comprises or encodes a CRISPR / Cas system or a component thereof. In some embodiments, the cargo comprises a therapeutic agent. In some embodiments, the cargo comprises a polypeptide. In some embodiments, the cargo comprises an antibody or antigen-binding fragment thereof, a peptidomimetic, a nucleotidomimetic, a drug, a diagnostic tool, an imaging tool, a small molecule, or a combination thereof.
[0104] Disclosed herein, in some aspects, is a composition comprising an endogenous Gag polypeptide that is not Arc and an assembly buffer.
[0105] In some embodiments, the assembly buffer contains less than 500 mOsm / kg of salt. In some embodiments, the assembly buffer contains about 270-330 mOsm / kg of salt. In some embodiments, the assembly buffer is a physiological buffer. In some embodiments, the assembly buffer is a phosphate buffered saline. In some embodiments, the assembly buffer comprises NaCl, KC1, Na2HPO4, and KH2PO4. In some embodiments, the assembly buffer comprises about 50 mM to about 750 mM of a monovalent salt. In some embodiments, the assembly buffer comprises about 50 mM to about 750 mM of NaCl. In some embodiments, the assembly buffer comprises glycerol. In some embodiments, the assembly buffer comprises about 5-15% glycerol. In some embodiments, the assembly buffer comprises about 10% glycerol. In some embodiments, the assembly buffer comprises Na2HPC>4 / NaH2PO4. In some embodiments, the assembly buffer comprises a pH of about 6-9. In some embodiments, the assembly buffer comprises a pH of 7.3-8.5. In some embodiments, the assembly buffer does not contain a chaotropic agent and / or a reducing agent.
[0106] In some embodiments, the assembly buffer comprises a nucleic acid. In some embodiments, the assembly buffer comprises a divalent cation. In some embodiments, the assembly buffer comprises ZnCk. In some embodiments, the assembly buffer comprises MgCh. In some embodiments, the assembly buffer comprises NaCl, ZnCh, and optionally Tris. In some embodiments, the assembly buffer comprises about 150mM NaCl, about lOpM ZnCh, and optionally about 25mM Tris. In some embodiments, the assembly buffer comprises about150mM NaCl, about ImM DTT, about lOpM ZnCh, and optionally about 50 mM Tris. In some embodiments, the assembly buffer comprises MES and MgCb. In some embodiments, the assembly buffer comprises about 50 mM MES and about 40mM MgCh. In some embodiments, the assembly buffer comprises a pH of about 6-9. In some embodiments, the assembly buffer comprises a pH of about 7.5-8. In some embodiments, the assembly buffer comprises a pH of about 5-7. In some embodiments, the assembly buffer comprises a pH of about 6. In some embodiments, the assembly buffer does not contain a chaotropic agent and / or a reducing agent.
[0107] Disclosed herein in some aspects, is a composition comprising an endogenous Gag polypeptide and a disassembly buffer.
[0108] In some embodiments, the disassembly buffer comprises a reducing agent. In some embodiments, the reducing agent comprises reduced glutathione (GSH), beta mercaptoethanol (P-ME), Dithiothreitol (DTT), or tris(2-carboxyethyl)phosphine (TCEP).
[0109] In some embodiments, the disassembly buffer comprises the reducing agent at a concentration of about 5 mM to about 100 mM. In some embodiments, wherein the disassembly buffer comprises the reducing agent at a concentration of about 10-50 mM. In some embodiments, the disassembly buffer comprises a solubilizing agent. In some embodiments, the solubilizing agent comprises a non-denaturing detergent. In some embodiments, the nondenaturing detergent is CHAPS. In some embodiments, the disassembly buffer comprises about 5-10% CHAPS. In some embodiments, the solubilizing agent comprises urea. In some embodiments, the disassembly buffer comprises about 2-10M urea. In some embodiments, the disassembly buffer comprises a pH of about 7-11. In some embodiments, the disassembly buffer comprises a pH of about 8-10. In some embodiments, the disassembly buffer comprises a pH of about 8. In some embodiments, the disassembly buffer further comprises a divalent cation. In some embodiments, the disassembly buffer further comprises MgCh. In some embodiments, the disassembly buffer comprises CHAPS, TCEP, and optionally Tris. In some embodiments, the disassembly buffer comprises about 5-10% CHAPS, about 10-20mM TCEP, and optionally about 5-100 mM Tris. In some embodiments, the disassembly buffer comprises TCEP, MgCh, and optionally Tris. In some embodiments, the disassembly buffer comprises urea, DTT, optionally NaCl, optionally NaP, and optionally glycerol. In some embodiments, the disassembly buffer comprises about 4-8M urea, about 10-80 mM DTT, optionally about 50-500 mM NaCl,optionally about 5-50 mM NaP, and optionally about 5-15% glycerol. In some embodiments, the disassembly buffer comprises urea, DTT, optionally NaCl, and optionally MgCh.
[0110] In some embodiments, the disassembly buffer comprises the reducing agent at a concentration of about 5-10 mM. In some embodiments, the disassembly buffer comprises a pH of about 7-11. In some embodiments, the disassembly buffer comprises a pH of about 8-10. In some embodiments, the disassembly buffer comprises a pH of about 8. In some embodiments, the disassembly buffer comprises a monovalent salt, a reducing agent, and optionally a basic buffering agent. In some embodiments, the disassembly buffer comprises NaCl, DTT, and optionally Tris. In some embodiments, the disassembly buffer comprises about 25-1800 mM NaCl, 1-20 mM DTT, and optionally 20-75 mM Tris. In some embodiments, the disassembly buffer comprises about 500 mM NaCl, about 5mM DTT, and optionally 50 mM Tris. In some embodiments, the disassembly buffer comprises about 50 mM NaCl, about lOmM DTT, and optionally 25 mM Tris. In some embodiments, the disassembly buffer comprises at least 800 mOsm / kg of salt. In some embodiments, the disassembly buffer comprises less than 150 mOsm / kg of salt.[0U1] Disclosed herein, in some aspects, is a composition comprising a plurality of capsids comprising an endogenous Gag polypeptide that is not Arc, wherein the plurality of capsids is at least 50% pure as determined by SDS-PAGE.
[0112] Disclosed herein, in some aspects, is a composition comprising a plurality of capsids comprising an endogenous Gag polypeptide that is not Arc, wherein the composition comprises at least 50% particle homogeneity as determined by multi-angle dynamic light scattering.
[0113] In some embodiments, the endogenous Gag polypeptide is a native endogenous Gag polypeptide. In some embodiments, the endogenous Gag polypeptide is an engineered endogenous Gag polypeptide. In some embodiments, the endogenous Gag polypeptide is a PNMA family protein. In some embodiments, the endogenous Gag polypeptide is a mammalianPNMA family protein. In some embodiments, the endogenous Gag polypeptide is a humanPNMA family protein. In some embodiments, the endogenous Gag polypeptide comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55 and 77-92. In some embodiments, the endogenous Gag polypeptide comprises an amino acid sequence with at least 90% sequence identity to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55 and77-92. In some embodiments, the endogenous Gag polypeptide is PNMA5. In some embodiments, the endogenous Gag polypeptide is PNMA5 and comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of any one of SEQ ID NOs: 5, 77, and 78. In some embodiments, the endogenous Gag polypeptide is PNMA5 and comprises an amino acid sequence with at least 90% sequence identity to any one of SEQ ID NOs: 5, 77, and 78. In some embodiments, the endogenous Gag polypeptide is PNMA5 and comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of SEQ ID NO: 77. In some embodiments, the endogenous Gag polypeptide is PNMA5 and comprises an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 77. In some embodiments, the endogenous Gag polypeptide is PNMA2. In some embodiments, the endogenous Gag polypeptide is PNMA2 and comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of any one of SEQ ID NOs: 1, 7, 8, 38, 40, 43, 46, 48, 49, and 52. In some embodiments, the endogenous Gag polypeptide is PNMA2 and comprises an amino acid sequence with at least 90% sequence identity to any one of SEQ ID NOs: 1, 7, 8, 38, 40, 43, 46, 48, 49, and 52. In some embodiments, the endogenous Gag polypeptide is PNMA2 and comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of SEQ ID NO: 7. In some embodiments, the endogenous Gag polypeptide is PNMA2 and comprises an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 7. In some embodiments, the endogenous Gag polypeptide comprises a sequence modification relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55 and 77-92.
[0114] In some embodiments, the endogenous Gag polypeptide is An RTL family protein. In some embodiments, the endogenous Gag polypeptide is a mammalian RTL family protein. In some embodiments, the endogenous Gag polypeptide is a human RTL family protein. In some embodiments, the endogenous Gag polypeptide comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of any one of SEQ ID NOs: 60-62 and 65-74. In some embodiments, the endogenous Gag polypeptide comprises an amino acid sequence with at least 90% sequence identity to any one of SEQ ID NOs: 60-62 and 65-74. In some embodiments, the endogenous Gag polypeptide is RTL10 (BOP). In some embodiments, the endogenous Gag polypeptide is RTL 10 and comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of any one of SEQ ID NOs: 60,61, or 69. In some embodiments, the endogenous Gag polypeptide is RTL10 and comprises an amino acid sequence with at least 90% sequence identity to any one of SEQ ID NOs: 60, 61, or 69. In some embodiments, the endogenous Gag polypeptide is RTL10 and comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of SEQ ID NO: 60. In some embodiments, the endogenous Gag polypeptide is RTL10 and comprises an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 60. In some embodiments, the endogenous Gag polypeptide is PEG10. In some embodiments, the endogenous Gag polypeptide is PEG10 and comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of any one of SEQ ID NOs: 65-74. In some embodiments, the endogenous Gag polypeptide is PEG10 and comprises an amino acid sequence with at least 90% sequence identity to any one of SEQ ID NOs: 65-74. In some embodiments, the endogenous Gag polypeptide is PEG10 and comprises an amino acid sequence with at least 90% sequence identity to at least 100 consecutive amino acids of SEQ ID NO: 65. In some embodiments, the endogenous Gag polypeptide is PEG10 and comprises an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 65. In some embodiments, the endogenous Gag polypeptide comprises a sequence modification relative to any one of SEQ ID NOs: 60-62 and 65-74.
[0115] In some embodiments, the sequence modification comprises an amino acid deletion. In some embodiments, the sequence modification comprises an amino acid insertion. In some embodiments, the sequence modification comprises an amino acid substitution. In some embodiments, the sequence modification comprises a cargo binding domain. In some embodiments, the sequence modification comprises nucleic acid binding domain. In some embodiments, the sequence modification comprises a zinc finger domain. In some embodiments, the sequence modification comprises a sub-cellular localization signal, a nuclear localization signal, or an antibody or antigen-binding fragment thereof. In some embodiments, the sequence modification is at an N-terminus of the endogenous Gag polypeptide. In some embodiments, the sequence modification is at a C-terminus of the endogenous Gag polypeptide. In some embodiments, the sequence modification within the endogenous Gag polypeptide. In some embodiments, the sequence modification comprises addition of a cysteine residue that is not present in a native form of the endogenous Gag polypeptide or elimination of a cysteine residue that is present in native form of the endogenous Gag polypeptide. In some embodiments, theplurality of capsids is at least 95% pure as determined by SDS-PAGE. In some embodiments, the endogenous Gag polypeptide is capable of disassembling into a non-capsid state with an efficiency of at least 1% and reassembling into a capsid state with an efficiency of at least 1%. In some embodiments, the efficiency of the disassembling and the efficiency of the reassembling are as determined by size exclusion chromatography. In some embodiments, the size exclusion chromatography comprises quantifying a capsid peak and a monomer peak. In some embodiments, the endogenous Gag polypeptide is capable of disassembling into a non-capsid state with an efficiency of at least 20%. In some embodiments, the endogenous Gag polypeptide is capable of reassembling into the capsid state with an efficiency of at least 20%. In some embodiments, the composition further comprises a heterologous cargo that is not associated the endogenous Gag polypeptide in nature. In some embodiments, the heterologous cargo comprises a nucleic acid. In some embodiments, the heterologous cargo comprises an RNA. In some embodiments, the heterologous cargo comprises a DNA. In some embodiments, the heterologous cargo comprises or encodes a gene editing system or a component thereof. In some embodiments, the heterologous cargo comprises or encodes a CRISPR / Cas system or a component thereof. In some embodiments, the heterologous cargo comprises a polypeptide. In some embodiments, the heterologous cargo comprises a therapeutic agent. In some embodiments, the heterologous cargo comprises an antibody or antigen-binding fragment thereof, a peptidomimetic, a nucleotidomimetic, a drug, a diagnostic tool, an imaging tool, a small molecule, or a combination thereof. In some embodiments, the composition further comprises a delivery component. In some embodiments, the delivery component comprises a lipid, optionally wherein the lipid is a cationic lipid. In some embodiments, the delivery component comprises a polypeptide, optionally wherein the polypeptide is a cationic polypeptide. In some embodiments, the delivery component comprises polymer, optionally wherein the polymer is a cationic polymer. In some embodiments, the delivery component comprises a cell-penetrating peptide. In some embodiments, the delivery component comprises a fusogenic protein. In some embodiments, the delivery component comprises an endogenous retroviral envelope protein. In some embodiments, the delivery component comprises a liposome.
[0116] Disclosed herein, in some aspects is a nucleic acid encoding the endogenous Gag polypeptide of any one of the preceding embodiments. Disclosed herein, in some aspects is avector comprising the nucleic acid. Disclosed herein, in some aspects is a cell comprising the nucleic acid.
[0117] Disclosed herein, in some aspects is a method of delivering a heterologous cargo to a cell, comprising contacting the cell with the composition of any one of the preceding embodiments.
[0118] In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is contacted with the capsid at a concentration of at least about 0.001 pg / mL. In some embodiments, the contacting is in vivo. In some embodiments, the contacting is in vitro or ex vivo.
[0119] In some embodiments, the endogenous Gag polypeptide is not PNMA2.I. EFFECTOR MOIETIES AND ITS CONJUGATION TO THE ENDOGENOUS GAG POLYPEPTIDES
[0120] The effector moiety is conjugated, directly or indirectly, to an endogenous Gag polypeptide. In one preferred embodiment, the endogenous Gag polypeptide in the conjugate is the same as the unconjugated endogenous Gag polypeptide in the capsid. In an alternative embodiment, the endogenous Gag polypeptide in the conjugate is different from the unconjugated endogenous Gag polypeptide in the capsid.
[0121] In one embodiment, the effector moiety is conjugated, directly or indirectly, through the N-terminus of the endogenous Gag polypeptide in the effector-polypeptide conjugate. In general, conjugation at the N-terminus results in the effector moiety being exposed outside the capsid. In another embodiment, the effector moiety is conjugated, directly or indirectly, through the C-terminus of the endogenous Gag polypeptide in the effector-polypeptide conjugate. In general, conjugation at the C-terminus results in the effector moiety being internalized within the capsid.
[0122] In one embodiment, the effector moiety is conjugated using Spy tag and Spy catcher moieties. For instance, the effector-polypeptide conjugate may comprise (i) a Spytag conjugated to the endogenous Gag polypeptide and (ii) a Spycatcher conjugated to the Spytag and to the effector moiety. The Spycatcher may be conjugated through a linker (e.g., a 4 to 12 amino acid length linker, or a 5 to 10 amino acid length linker, such as a 5 amino acid linker or a 10 amino acid linker) to the endogenous polypeptide. In some instances, the amino acid sequence of theSpyCatcher is as shown in SEQ ID NO:339 or 340. In some instances, the amino acid sequence of the SpyTag is as shown in any one of SEQ ID NOs: 341-343.
[0123] Fig. 4 shows one example of the structure of a PNP which includes an effector- polypeptide conjugate where the effector moiety (referred to as a ligand in Fig. 4) is conjugated to a Spytag, which is conjugated to a Spycatcher. The Spycatcher is conjugated to a linker, which is conjugated to an endogenous Gag polypeptide, such as dArc. In some instances, the amino acid sequence of dArc is as shown in SEQ ID: 344. Fig. 33 shows another example of a PNP, however, here the Spytag is conjugated to a Fab or antibody.
[0124] In another embodiment, the effector-polypeptide conjugate further comprises a thiol maleimide linker conjugated to the endogenous polypeptide and to the effector moiety, such as shown in Fig. 3. In one embodiment, the endogenous Gag polypeptide is conjugated through the maleimide group within the 0 to 10 positions of the N termini. In one embodiment, the endogenous Gag polypeptide is conjugated through the maleimide group at the 3-position from the N termini. An example of such an effector-polypeptide conjugate is T4C-dArc (SEQ ID NO: 345).
[0125] In one embodiment, the one or more effector-polypeptide conjugates are selected from (i) a targeted conjugate comprising an endogenous Gag polypeptide and a targeting moiety, (ii) an endosomal escape conjugate comprising an endogenous Gag polypeptide and an endosomal escape moiety, or (iii) any combination of any of the foregoing.
[0126] The endosomal escape moiety may be conjugated, directly or indirectly, through the N-terminus of the endogenous Gag polypeptide. In an alternative embodiment, the endosomal escape moiety is conjugated, directly or indirectly, through the C-terminus of the endogenous Gag polypeptide.
[0127] In one embodiment, the Gag polypeptide. In some instances, the sequence of dArc is modified to any one of SEQ ID NOs: 336-338.
[0128] In one embodiment, the targeting moiety enhances cell-specific uptake. In another embodiment, the targeting moiety enhances nuclear localization. The targeting moiety may be a targeting polypeptide.
[0129] In one embodiment, the targeting polypeptide comprises one or more of LDL- receptors, apolipoprotein mimetic peptide (such as an apoA-I, apoE, or apoC-II peptide), PLA2,or any combination of any of the foregoing. In another embodiment, the targeting polypeptide is an LDL receptor selected from LRKLRK, LRKRLLRD, and LKAYKS.
[0130] In some instances, the PLA2 domain sequence is SEQ ID NOs: 134-234. In some instances, PLA2 is part of a PLA2-containing protein sequence. In some instances, the PLA@- containing protein sequence is SEQ ID:234 - SEQ ID:334.
[0131] In one embodiment, the targeting moiety is an antibody (e.g., a monoclonal antibody), fragment antigen-binding (Fab) protein (such as a CD71 Fab), or single-chain variable fragment.
[0132] In one embodiment, the endosomal escape conjugate comprises a first linker between the Gag polypeptide and the endosomal escape moiety. The first linker may comprise 5xGS or EAAK.
[0133] In another embodiment, the targeted conjugate comprises a first linker between the Gag polypeptide and the targeting moiety. The first linker may comprise 5xGS or EAAK.II. ENDOGENOUS GAG POLYPEPTIDES AND PNMA FAMILY POLYPEPTIDES
[0134] Endogenous Gag (endo-Gag) proteins are eukaryotic proteins that have predicted and annotated similarity to viral Gag proteins. As described herein, in some embodiments an endoGag protein is capable of assembling into a capsid (VLP) state or form. Endo-Gag proteins of the disclosure can be useful, for example, for packaging and delivering cargos (e.g., heterologous cargos) to cells.
[0135] Illustrative endo-Gag polypeptides include PNMA1, PNMA2, PNMA3, PNMA4 (MO API), PNMA5, PNMA6A, PNMA6B / 6D, PNMA6E, PNMA6F, PNMA7A, PNMA7B, PNMA8A, PNMA8B, PNMA8C, CCDC8, Arc, RTL10 (BOP), LDOC1, PEG10 (RTL2), RTL3, RTL6, RTL8A, RTL8B, and ZNF18.
[0136] An endogenous Gag protein disclosed herein can be a member of the retrotransposon Gag-like family (e g. PEG10 (RTL2), RTL3, RTL6, RTL7 (LDOC), RTL8A, RTL8B, or RTL10 (BOP)).
[0137] An endogenous Gag protein disclosed herein can be PEG10 (RTL2). In some embodiments, certain versions of PEG10 can contains two overlapping open reading frames, RF1 and RF2, and can be expressed as two proteins by -1 translational frameshifting (-1 FS): (1) a shorter gag-like polypeptide from RF1 , comprising a capsid (CA) domain (e.g., including an N- terminal CA and a C-terminal CA), a nucleocapsid (NC) domain, and CCHC-type zinc fingermotif), and (2) a longer, gag / pol-like fusion protein from RF1 / RF2 that can additionally comprise a protease (PRO), and a reverse transcriptase-like (RT) domain. Additional isoforms resulting from alternatively spliced transcript variants, and use of upstream non-AUG (CUG) start codon have been reported for PEG10. PEG10 can bind mRNA (e.g., its own mRNA) in the 5'-UTR region, in the region near the boundary between the nucleocapsid (NC) and protease (PRO) coding sequences, and in the beginning of the 3'-UTR region. Illustrative examples of PEG10 amino acid sequences are provided in Table A as SEQ ID NOs: 69-74.
[0138] In some embodiments, a PEG10 polypeptide disclosed herein comprises or consists essentially of the shorter gag-like polypeptide from RF1, comprising a capsid (CA) domain (e.g., including an N-terminal CA and a C-terminal CA), a nucleocapsid (NC) domain, and / or a CCHC-type zinc finger motif. In some embodiments, a PEG10 polypeptide disclosed herein does not contain the protease (PRO) or reverse transcriptase-like (RT) domain. In some embodiments, a PEG10 polypeptide disclosed herein comprises or consists essentially of the longer, gag / pol- like fusion protein from RF1 / RF2.
[0139] In some embodiments, a PEG10 polypeptide disclosed herein comprises a capsid (CA) domain (e.g., including an N-terminal CA and a C-terminal CA), a nucleocapsid (NC) domain, and CCHC-type zinc finger motif.
[0140] In some embodiments, a PEG10 polypeptide disclosed herein comprises a capsid (CA) domain and CCHC-type zinc finger motif. In some embodiments, a PEG10 polypeptide disclosed herein comprises a capsid (CA) domain and a nucleocapsid (NC) domain. In some embodiments, a PEG10 polypeptide disclosed herein comprises a nucleocapsid (NC) domain and CCHC-type zinc finger motif.
[0141] In some embodiments, a PEG10 polypeptide disclosed herein comprises an N- terminal CA domain, a C-terminal CA domain, a nucleocapsid (NC) domain, and CCHC-type zinc finger motif.
[0142] In some embodiments, a PEG10 polypeptide disclosed herein comprises a C-terminal CA domain, a nucleocapsid (NC) domain, and CCHC-type zinc finger motif. In some embodiments, a PEG10 polypeptide disclosed herein comprises an N-terminal CA domain, a nucleocapsid (NC) domain, and CCHC-type zinc finger motif). In some embodiments, a PEG10 polypeptide disclosed herein comprises an N-terminal CA domain, a C-terminal CA domain, and CCHC-type zinc finger motif. In some embodiments, a PEG10 polypeptide disclosed hereincomprises an N-terminal CA domain, a C-terminal CA domain, and a nucleocapsid (NC) domain.
[0143] In some embodiments, a PEG10 polypeptide disclosed herein comprises an N- terminal CA domain and a C-terminal CA domain. In some embodiments, a PEG10 polypeptide disclosed herein comprises an N-terminal CA domain, and a nucleocapsid (NC) domain. In some embodiments, a PEG10 polypeptide disclosed herein comprises an N-terminal CA domain and CCHC-type zinc finger motif. In some embodiments, a PEG10 polypeptide disclosed herein comprises an a C-terminal CA domain and a nucleocapsid (NC) domain. In some embodiments, a PEG10 polypeptide disclosed herein a C-terminal CA domain and a CCHC-type zinc finger motif. In some embodiments, a PEG10 polypeptide disclosed herein comprises an a nucleocapsid (NC) domain and CCHC-type zinc finger motif).
[0144] An endogenous Gag protein disclosed herein can be RTL10 (BOP). RTL10 can comprise a Bcl-2 homology 3 (BH3) domain, a capsid (CA) domain (e.g., including an N- terminal CA and a C-terminal CA domain), and / or a nucleocapsid (NC) domain (e.g., a cryptic NC domain).
[0145] In some embodiments, an RTL10 polypeptide disclosed herein comprises a BH3 domain, an N-terminal CA domain, a C-terminal CA domain, and a nucleocapsid (NC) domain (e.g., a cryptic NC domain).
[0146] In some embodiments, an RTL10 polypeptide disclosed herein comprises an N- terminal CA domain, a C-terminal CA domain, and a nucleocapsid (NC) domain (e g., a cryptic NC domain). In some embodiments, an RTL10 polypeptide disclosed herein comprises a BH3 domain, a C-terminal CA domain, and a nucleocapsid (NC) domain (e.g., a cryptic NC domain). In some embodiments, an RTL10 polypeptide disclosed herein comprises a BH3 domain, an N- terminal CA domain, and a nucleocapsid (NC) domain (e.g., a cryptic NC domain). In some embodiments, an RTL10 polypeptide disclosed herein comprises a BH3 domain, an N-terminal CA domain, and a C-terminal CA domain.
[0147] In some embodiments, an RTL10 polypeptide disclosed herein comprises a C- terminal CA domain and a nucleocapsid (NC) domain (e.g., a cryptic NC domain). In some embodiments, an RTL10 polypeptide disclosed herein comprises a BH3 domain and a nucleocapsid (NC) domain (e.g., a cryptic NC domain). In some embodiments, an RTL10 polypeptide disclosed herein comprises a BH3 domain, and an N-terminal CA domain. In someembodiments, an RTL10 polypeptide disclosed herein comprises an N-terminal CA domain, and a C-terminal CA domain. In some embodiments, an RTL10 polypeptide disclosed herein comprises an N-terminal CA domain a nucleocapsid (NC) domain (e.g., a cryptic NC domain). In some embodiments, an RTL10 polypeptide disclosed herein comprises a BH3 domain and a C-terminal CA domain.
[0148] In some embodiments, an RTL10 polypeptide comprises a BH3 domain, e g, LAQLGDYMS (SEQ ID NO: 76). In some embodiments, an RTL10 polypeptide lacks a BH3 domain. In some embodiments, an RTL10 polypeptide lacks a functional BH3 domain, for example, contains a variant of the BH3 domain sequence that reduces or abolishes one or more functions of the BH3 domain, such as reducing or abolishing binding to an interaction partner (e.g., VDAC binding, Bcl-2 family member binding), or reducing or abolishing induction of apoptosis. Illustrative modifications to reduce function of the BH3 domain can comprise mutation (e.g., substitution) of LI 18 and / or D123 residues, e.g., LI 18A and / or D123A substitutions.
[0149] Illustrative examples of RTL10 amino acid sequences are provided in Table A as SEQ ID NOs: 69-74.
[0150] An endogenous Gag protein disclosed herein can be a member of the Paraneoplastic Ma (PNMA) family. The Paraneoplastic Ma (PNMA) family of endo-Gag proteins comprises 15 members, PNMA1, PNMA2, PNMA3, PNMA4 (MOAP1), PNMA5, PNMA6A, PNMA6B / 6D, PNMA6E, PNMA6F, PNMA7A (ZCCHC12), PNMA7B (ZCCHC18), PNMA8A, PNMA8B, PNMA8C, and CCDC8. PNMA family members share sequence homology to the Gag protein of Long Terminal Repeat (LTR) retrotransposons. PNMA proteins are discussed in Pang et al., (2018) PNMA family: Protein interaction network and cell signalling pathways implicated in cancer and apoptosis. Cellular signalling, 45, 54-62, which is incorporated herein by reference for such disclosure.
[0151] PNMA polypeptides comprise several domains. Most PNMA family members share high sequence homology at the N-terminal conserved domain (NCD), and the central conserved domain (CCD). In some family members, a unique protein sequence or domain (UPD) is situated between NCD and CCD domains. A region rich in Lysine and Arginine basic residues, designated KRs is also found in the protein sequences of most members of PNMA family. Many PNMA family members share sequence homology at a poly-glutamic acid rich region (PolyE)near the C-terminus. Beyond the polyglutamic acid rich region is the C-terminus with varying length and low sequence homology (variable C-terminal sequence, VCS) that can be identified in the protein sequences among the PNMA family members. Homologues of PNMA proteins exist in other mammalian species, for example, chimpanzee, monkey, rat, and mouse. In some family members, electrostatic interaction between the KRs and PolyE contribute to protein conformation.
[0152] PNMA2 is one member of the PNMA family. An illustrative PNMA2 is human PNMA2, which can comprise about 364 amino acids, e.g., as provided in SEQ ID NO: 1. PNMA2 comprises a central conserved domain (CCD), for example, at residues 202-204 of SEQ ID NO: 1, and a poly-glutamic acid rich region (PolyE) near the C-terminus, for example, at residues 333-338 of SEQ ID NO: 1. PNMA2 can form heterodimers with PNMA1 and PNMA4 (MOAP1), for example to modulate (e g., inhibit) apoptotic signaling. PNMA2 can also comprise cysteine residues that have the potential to for disulfide bonds, for example, at positions 10, 136, 233, and 310 of SEQ ID NO: 1.
[0153] PNMA5 is another member of the PNMA family. An illustrative PNMA5 is human PNMA5, which can comprise about 448 amino acids, e.g., as provided in SEQ ID NO: 5.
[0154] Full sequences of illustrative PNMA family members are provided in Table A as SEQ ID NOs: 1-6 and 80-82.
[0155] Additional examples of endo-Gag proteins are disclosed in Campillos et al., (2006) Computational characterization of multiple Gag-like human proteins, Trends Genet 22(11):585- 9, which is incorporated herein by reference for such disclosure. Arc (activity-regulated cytoskeleton-associated protein) is another example of an endo-Gag protein. Arc regulates the endocytic trafficking of a-amino-3-hydroxy-5-methylisoxazole-4-propionic acid (AMPA) type glutamate receptors. Arc activities have been linked to synaptic strength and neuronal plasticity. Phenotypes of loss of Arc in experimental murine model included defective formation of longterm memory and reduced neuronal activity and plasticity.
[0156] In some embodiments, a nucleic acid sequence or amino acid sequence disclosed herein comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to a nucleic acid sequence or an amino acid sequence provided herein.
[0157] In some embodiments, an endo-Gag, RTL, or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag, RTL, or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to at least 50 consecutive amino acids of any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60-62, 65-74, and 77-92
[0158] In some embodiments, an endo-Gag, RTL, or PNMA polypeptide disclosed herein (e g., a recombinant or engineered endo-Gag, RTL, or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to at least 100 consecutive amino acids of any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60-62, 65-74, and 77-92
[0159] In some embodiments, an endo-Gag, RTL, or PNMA polypeptide disclosed herein (e g., a recombinant or engineered endo-Gag, RTL, or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to at least 200 consecutive amino acids of any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60-62, 65-74, and 77-92
[0160] In some embodiments, an endo-Gag, RTL, or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag, RTL, or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to at least 300 consecutive amino acids of any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60-62, 65-74, and 77-92.
[0161] In some embodiments, an endo-Gag, RTL, or PNMA polypeptide disclosed herein (e g., a recombinant or engineered endo-Gag, RTL, or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60-62, 65-74, and 77-92
[0162] In some embodiments, an endo-Gag, RTL, or PNMA polypeptide disclosed herein (e g., a recombinant or engineered PEG10, RTL10, endo-Gag or PNMA polypeptide) comprises one or more amino acid substitutions, deletions or insertions, from the polypeptide having the amino acid sequence of any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60- 62, 65-74, and 77-92
[0163] In some embodiments, an endo-Gag, RTL, or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered PEG10, RTL10, endo-Gag or PNMA polypeptide) comprises from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more up to about 100, 90, 80, 70, 60, 50, 45, 40, 35, 30, 25, 20, 15 amino acid insertions relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60-62, 65-74, and 77-92
[0164] In some embodiments, the endo-Gag, RTL, or PNMA polypeptide comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least or at least 50 amino acid insertions relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60-62, 65-74, and 77-92
[0165] In some embodiments, the endo-Gag, RTL, or PNMA polypeptide comprises at most 1, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, at most 11, at most 12, at most 13, at most 14, at most 15, at most 16, at most 17, at most 18, at most 19, at most 20, at most 25, at most 30, at most 35, at most 40, at most 45, or at most 50amino acid insertions relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53,55, 60-62, 65-74, and 77-92
[0166] In some embodiments, the endo-Gag, RTL, or PNMA polypeptide comprises 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-15, 1-20, 1-30, 1-40, 2-3, 2-4, 2-5, 2-6, 2-7, 2-8, 2-9, 2-10, 2-15, 2-20, 2-30, 2-40, 3-4, 3-5, 3-6, 3-7, 3-8, 3-9, 3-10, 3-15, 3-20, 3-30, 3-40, 4-5, 4-6, 4-7, 4- 8, 4-9, 4-10, 4-15, 4-20, 4-30, 5-6, 5-7, 5-8, 5-9, 5-10, 5-15, 5-20, 5-30, 5-40, 10-15, 15-20, or 20-25 amino acid insertions relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60-62, 65-74, and 77-92
[0167] In some embodiments, the endo-Gag, RTL, or PNMA polypeptide comprises 1, 2, 3,4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acid insertions relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60-62, 65-74, and 77-92
[0168] The one or more amino acid insertions can be at the N-terminus, the C-terminus, within the amino acid sequence, or a combination thereof. The amino acid insertions can be contiguous, non-contiguous, or a combination thereof.
[0169] In some embodiments, an endo-Gag, RTL, or PNMA polypeptide disclosed herein (e g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more up to about 100, 90, 80, 70, 60, 50, 45, 40, 35, 30, 25, 20, 15 amino acid deletions relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60-62, 65-74, and 77-92
[0170] In some embodiments, the endo-Gag, RTL, or PNMA polypeptide comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least or at least 50 amino acid deletions relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60-62, 65-74, and 77-92
[0171] In some embodiments, the endo-Gag, RTL, or PNMA polypeptide comprises at most 1, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, at most 11, at most 12, at most 13, at most 14, at most 15, at most 16, at most 17, at most 18, at most 19, at most 20, at most 25, at most 30, at most 35, at most 40, at most 45, or at most 50 amino acid deletions relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60-62, 65-74, and 77-92
[0172] In some embodiments, the endo-Gag, RTL, or PNMA polypeptide comprises 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-15, 1-20, 1-30, 1-40, 2-3, 2-4, 2-5, 2-6, 2-7, 2-8, 2-9, 2-10, 2-15, 2-20, 2-30, 2-40, 3-4, 3-5, 3-6, 3-7, 3-8, 3-9, 3-10, 3-15, 3-20, 3-30, 3-40, 4-5, 4-6, 4-7, 4- 8, 4-9, 4-10, 4-15, 4-20, 4-30, 5-6, 5-7, 5-8, 5-9, 5-10, 5-15, 5-20, 5-30, 5-40, 10-15, 15-20, or 20-25 amino acid deletions relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60-62, 65-74, and 77-92
[0173] In some embodiments, the endo-Gag, RTL, or PNMA polypeptide comprises 1, 2, 3,4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acid deletions relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60-62, 65-74, and 77-92
[0174] The one or more amino acid deletions can be at the N-terminus, the C-terminus, within the amino acid sequence, or a combination thereof. The amino acid deletions can be contiguous, non-contiguous, or a combination thereof.
[0175] In some embodiments, an endo-Gag, RTL, or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered PEG10, RTL10, endo-Gag or PNMA polypeptide) comprises from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more up to about 100, 90, 80, 70, 60, 50, 45, 40, 35, 30, 25, 20, 15 amino acid substitutions relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49,52, 53, 55, 60-62, 65-74, and 77-92
[0176] In some embodiments, the endo-Gag, RTL, or PNMA polypeptide comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least or at least 50 amino acid substitutions relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52,53, 55, 60-62, 65-74, and 77-92
[0177] In some embodiments, the endo-Gag, RTL, or PNMA polypeptide comprises at most 1, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, at most 11, at most 12, at most 13, at most 14, at most 15, at most 16, at most 17, at most 18, at most 19, at most 20, at most 25, at most 30, at most 35, at most 40, at most 45, or at most 50 amino acid substitutions relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60-62, 65-74, and 77-92
[0178] In some embodiments, the endo-Gag, RTL, or PNMA polypeptide comprises 1-2, 1- 3, 1 -4, 1-5, 1 -6, 1-7, 1 -8, 1-9, 1 -10, 1-15, 1-20, 1-30, 1 -40, 2-3, 2-4, 2-5, 2-6, 2-7, 2-8, 2-9, 2-10,2-15, 2-20, 2-30, 2-40, 3-4, 3-5, 3-6, 3-7, 3-8, 3-9, 3-10, 3-15, 3-20, 3-30, 3-40, 4-5, 4-6, 4-7, 4- 8, 4-9, 4-10, 4-15, 4-20, 4-30, 5-6, 5-7, 5-8, 5-9, 5-10, 5-15, 5-20, 5-30, 5-40, 10-15, 15-20, or 20-25 amino acid substitutions relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46- 49, 52, 53, 55, 60-62, 65-74, and 77-92
[0179] In some embodiments, the endo-Gag, RTL, or PNMA polypeptide comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acid substitutions relative to any one of SEQ ID NOs: 1-8, 38, 40, 41, 43, 44, 46-49, 52, 53, 55, 60-62, 65-74, and 77-92
[0180] An amino acid substitution can be a conservative or a non-conservative substitution. The one or more amino acid substitutions can be at the N-terminus, the C-terminus, within the amino acid sequence, or a combination thereof. The amino acid substitutions can be contiguous, non-contiguous, or a combination thereof.
[0181] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 1.
[0182] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 2.
[0183] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 3.
[0184] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of,or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 4.
[0185] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 5.
[0186] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 6.
[0187] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 7.
[0188] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 8.
[0189] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%,at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 38.
[0190] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 40.
[0191] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 43.
[0192] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 46.
[0193] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 48.
[0194] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, atleast 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 49.
[0195] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 52.
[0196] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 55.
[0197] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 77.
[0198] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 78.
[0199] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, atleast 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 80.
[0200] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 81.
[0201] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 82.
[0202] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 83.
[0203] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 84.
[0204] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, atleast 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 85.
[0205] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 86.
[0206] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 87.
[0207] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 88.
[0208] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 89.
[0209] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, atleast 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 90.
[0210] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 91.
[0211] In some embodiments, an endo-Gag or PNMA polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or PNMA polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 92.
[0212] In some embodiments, an endo-Gag or RTL polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or RTL polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 60.
[0213] In some embodiments, an endo-Gag or RTL polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or RTL polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 61.
[0214] In some embodiments, an endo-Gag or RTL polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or RTL polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, atleast 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 65.
[0215] In some embodiments, an endo-Gag or RTL polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or RTL polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 66.
[0216] In some embodiments, an endo-Gag or RTL polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or RTL polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 67.
[0217] In some embodiments, an endo-Gag or RTL polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or RTL polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 68.
[0218] In some embodiments, an endo-Gag or RTL polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or RTL polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 69.
[0219] In some embodiments, an endo-Gag or RTL polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or RTL polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, atleast 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 70.
[0220] In some embodiments, an endo-Gag or RTL polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or RTL polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 71.
[0221] In some embodiments, an endo-Gag or RTL polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or RTL polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 72.
[0222] In some embodiments, an endo-Gag or RTL polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or RTL polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 73.
[0223] In some embodiments, an endo-Gag or RTL polypeptide disclosed herein (e.g., a recombinant or engineered endo-Gag or RTL polypeptide) comprises, consists essentially of, or consists of an amino acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or substantially 100% sequence identity or similarity to SEQ ID NO: 74.
[0224] Various methods and software programs are used to determine the homology (e.g., identity or similarity) between two or sequences. For example, the degree of sequence identity or similarity between two sequences can be determined by comparing the two sequences using computer programs commonly employed for this purpose, such as global or local alignment algorithms. Non-limiting examples include NCBI BLAST, BLASTp, BLASTn, BLASTx,tBLASTn, Clustal W, MAFFT, Clustal Omega, AlignMe, Praline, GAP, BESTFIT, Needle (EMBOSS), Stretcher (EMBOSS), GGEARCH2SEQ, Water (EMBOSS), Matcher (EMBOSS), LALIGN, SSEARCH2SEQ, or another suitable method or algorithm, e.g., using default parameters. A Needleman and Wunsch global alignment algorithm can be used to align two sequences over their entire length, maximizing the number of matches and minimizes the number of gaps.
[0225] In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not Arc or does not comprise an amino acid sequence from Arc. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PNMA6 or does not comprise an amino acid sequence from PNMA6. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PNMA6A or does not comprise an amino acid sequence from PNMA6A. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PNMA6B or does not comprise an amino acid sequence from PNMA6B. In some embodiments, an endo-Gag polypeptide disclosed herein is not a protein from the paraneoplastic Ma (PNMA) family or does not comprise an amino acid sequence of a protein from the PNMA family. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not a protein from the retrotransposon Gag-like family or does not comprise an amino acid sequence from a protein from the retrotransposon Gag-like family. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PEG10 (RTL2) or does not comprise an amino acid sequence from PEG10 (RTL2). In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not RTL1 or does not comprise an amino acid sequence from RTL1. In some embodiments, an endo- Gag polypeptide or PNMA polypeptide disclosed herein is not RTL3 or does not comprise an amino acid sequence from RTL3. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not RTL6 or does not comprise an amino acid sequence from RTL6. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not RTL8A or does not comprise an amino acid sequence from RTL8A. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not RTL8B or does not comprise an amino acid sequence from RTL8B. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not BOP (RTL10) or does not comprise an amino acid sequence from BOP (RTL10). In some embodiments, an endo-Gag polypeptide orPNMA polypeptide disclosed herein is not LD0C1 (RTL7) or does not comprise an amino acid sequence from LD0C1 (RTL7). In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not MO API (PNMA4) or does not comprise an amino acid sequence from MO API (PNMA4). In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not ZNF18 or does not comprise an amino acid sequence from ZNF18. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not Asprvl or does not comprise an amino acid sequence from Asprvl. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not CCDC8 or does not comprise an amino acid sequence from CCDC8. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PGBD1 or does not comprise an amino acid sequence from PGBD1. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PNMA1 or does not comprise an amino acid sequence from PNMA1. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PNMA3 or does not comprise an amino acid sequence from PNMA3. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PNMA5 or does not comprise an amino acid sequence from PNMA5. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PNMA6B / 6D or does not comprise an amino acid sequence from PNMA6B / 6D. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PNMA6E or does not comprise an amino acid sequence from PNMA6E. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PNMA6F or does not comprise an amino acid sequence from PNMA6F. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PNMA7A (ZCCHC12) or does not comprise an amino acid sequence from PNMA7A (ZCCHC12). In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PNMA7B (ZCCHC18) or does not comprise an amino acid sequence from PNMA7B (ZCCHC18). In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PNMA8A or does not comprise an amino acid sequence from PNMA8A. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PNMA8B or does not comprise an amino acid sequence from PNMA8B. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PNMA8C or does not comprise an amino acid sequence from PNMA8C. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not CCDC8 or does not comprise an amino acid sequence from CCDC8. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not SCAN or does not comprise an amino acid sequence from SCAN. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not ZCCHC12 or does not comprise an amino acid sequence from ZCCHC12. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not ZCCHC18 or does not comprise an amino acid sequence from ZCCHC18. In some embodiments, an endo- Gag polypeptide or PNMA polypeptide disclosed herein is not ZNF274 or does not comprise an amino acid sequence from ZNF274. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not a SUSHI family member or does not comprise an amino acid sequence from a SUSHI family member. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not a SCAN family member or does not comprise an amino acid sequence from a SCAN family member. In some embodiments, an endo-Gag polypeptide or PNMA polypeptide disclosed herein is not PNMA2 or does not comprise an amino acid sequence from PNMA2.
[0226] In some instances, the endo Gag polypeptide (e.g., RTL family polypeptide, such as RTL10 or PEG10) comprises a full-length endo Gag polypeptide. In other instances, the endo Gag polypeptide comprises a fragment of the endo Gag, such as a truncated endo Gag (e.g., RTL family, RTL 10, or PEG10) polypeptide, that participates in the formation of a capsid. In further instances, the endo Gag polypeptide is a recombinant endo Gag polypeptide.
[0227] In some instances, the PNMA family polypeptide (e.g., PNMA2 or PNMA5) comprises a full-length PNMA polypeptide. In other instances, the PNMA family polypeptide comprises a fragment of the PNMA, such as a truncated PNMA polypeptide, that participates in the formation of a capsid. In further instances, the PNMA polypeptide is a recombinant PNMA family polypeptide.
[0228] In some instances, the endo-Gag polypeptide comprises a full-length endo-Gag protein. In other instances, the endo-Gag polypeptide comprises a fragment of an endo-Gag protein, such as a truncated endo-Gag polypeptide, that can participate in the formation of a capsid. In further instances, the endo-Gag polypeptide is a recombinant endo-Gag polypeptide.
[0229] In some instances, the PNMA family or endo-Gag polypeptide comprises from N- terminus to C-terminus the following domains: NCD, UPD, CCD, KRs, PolyE, and VCS.
[0230] In some embodiments, one or more non-essential regions which are not involved in capsid formation are removed from a PNMA family or endo-Gag protein (e.g., RTL family, RTL10, or PEG10) to generate an engineered PNMA or endo-Gag polypeptide. In such instances, one or more non-essential regions, e.g., an N-terminal region (e.g., up to 10 amino acids, up to 15 amino acids, up to 20 amino acids, up to 25 amino acids, up to 30 amino acids, up to 40 amino acids, or up to 50 amino acids), a C-terminal region (e.g., up to 10 amino acids, up to 15 amino acids, up to 20 amino acids, up to 25 amino acids, up to 30 amino acids, up to 40 amino acids, or up to 50 amino acids), or a combination thereof, are deleted from a PNMA or endo-Gag protein (e.g., a native or wild type RTL family, RTL10, PEG10, PNMA or endo-Gag protein) to generate an engineered PNMA or endo-Gag polypeptide.
[0231] In some embodiments, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50 amino acids are removed from an N-terminus of an RTL family, a PNMA or endo-Gag protein (e.g., a native or wild type RTL10, PEG10, PNMA or endo-Gag protein) to generate an engineered PNMA or endo-Gag polypeptide. In some embodiments, at most 5, at most 10, at most 15, at most 20, at most 25, at most 30, at most 35, at most 40, at most 45, or at most 50 amino acids are removed from an N-terminus of a RTL family, PNMA or endo-Gag protein (e.g., a native or wild type RTL, PNMA or endo-Gag protein) to generate an engineered PNMA or endo-Gag polypeptide. In some embodiments, about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, or about 50, amino acids are removed from an N-terminus of a RTL, PNMA or endo-Gag protein (e.g., a native or wild type RTL, PNMA or endo-Gag protein) to generate an engineered RTL, PNMA or endo-Gag polypeptide. In some embodiments, about 1 to about 50, about 1 to about 40, about 1 to about 35, about 1 to about 30, about 1 to about 27, about 1 to about 25, about 1 to about 23, about 1 to about 20, about 1 to about 15, about 1 to about 10, about 1 to about 5, about 5 to 50, about 5 to 40, about 5 to 35, about 5 to 30, about 5 to 27, about 5 to 25, about 5 to 23, about 5 to 20, about 5 to 15, about 5 to 10, about 10 to 50, about 10 to 40, about 10 to 35, about 10 to 30, about 10 to 27, about 10 to 25, about 10 to 23, about 10 to 20, about 10 to 15, about 15 to 50, about 15 to 40, about 15 to 35, about 15 to 30, about 15 to 27, about 15 to 25, about 15 to 23, about 15 to 20, about 20 to 50, about 20 to 40, about 20 to 35, about 20 to 30, about 20 to 27, about 20 to 25, or about 20 to 23 amino acids are removed from an N-terminus of a RTL, PNMAor endo-Gag protein (e g., a native or wild type RTL, PNMA or endo-Gag protein) to generate an engineered RTL, PNMA or endo-Gag polypeptide.
[0232] In some embodiments, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50 amino acids are removed from a C-terminus of an RTL family, a PNMA or endo-Gag protein (e.g., a native or wild type RTL, PNMA or endo-Gag protein, such as a native or wild type PEG10 or RTL 10) to generate an engineered RTL, PNMA or endo-Gag polypeptide. In some embodiments, at most 5, at most 10, at most 15, at most 20, at most 25, at most 30, at most 35, at most 40, at most 45, or at most 50 amino acids are removed from a C-terminus of a RTL, PNMA or endo-Gag protein (e.g., a native or wild type RTL, PNMA or endo-Gag protein) to generate an engineered RTL, PNMA or endo-Gag polypeptide. In some embodiments, about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, or about 50, amino acids are removed from a C-terminus of a RTL, RTL, PNMA or endo-Gag protein (e.g., a native or wild type RTL, PNMA or endo-Gag protein) to generate an engineered RTL, PNMA or endo-Gag polypeptide. In some embodiments, about 1 to about 50, about 1 to about 40, about 1 to about 35, about 1 to about 30, about 1 to about 27, about 1 to about 25, about 1 to about 23, about 1 to about 20, about 1 to about 15, about 1 to about 10, about 1 to about 5, about 5 to 50, about 5 to 40, about 5 to 35, about 5 to 30, about 5 to 27, about 5 to 25, about 5 to 23, about 5 to 20, about 5 to 15, about 5 to 10, about 10 to 50, about 10 to 40, about 10 to 35, about 10 to 30, about 10 to 27, about 10 to 25, about 10 to 23, about 10 to 20, about 10 to 15, about 15 to 50, about 15 to 40, about 15 to 35, about 15 to 30, about 15 to 27, about 15 to 25, about 15 to 23, about 15 to 20, about 20 to 50, about 20 to 40, about 20 to 35, about 20 to 30, about 20 to 27, about 20 to 25, or about 20 to 23 amino acids are removed from a C-terminus of a RTL, PNMA or endo-Gag protein (e.g., a native or wild type RTL, PNMA or endo-Gag protein) to generate an engineered RTL, PNMA or endo-Gag polypeptide.
[0233] In some embodiments, at least 15 amino acids are removed from a native or wild type PNMA family polypeptide (e.g., PNMA2 or PNMA5) to generate an engineered PNMA polypeptide. In some embodiments, at most 25 amino acids are removed from a C-terminus of a native or wild type PNMA to generate an engineered PNMA polypeptide. In some embodiments, about 25 amino acids are removed from a native or wild type PNMA to generate an engineered PNMA polypeptide. In some embodiments, about 10 to 30 amino acids are removed from a native or wild type PNMA to generate an engineered PNMA polypeptide. In some embodiments,about 20 to 27 amino acids are removed from a native or wild type PNMA to generate an engineered PNMA polypeptide. In some embodiments, about 24 to 26 amino acids are removed from a native or wild type PNMA to generate an engineered PNMA polypeptide.
[0234] In some cases, only the essential regions or substantially only the regions involved in capsid assembly / forming and cargo binding remain in an RTL family, a PNMA or endo-Gag polypeptide. In some embodiments, only the essential regions or substantially only the regions involved in capsid assembly / forming remain in a PEG10, RTL 10, PNMA2 or PNMA5 polypeptide. In some embodiments essential regions or regions involved in capsid assembly / forming comprise a capsid (CA) domain, an N-terminal CA domain, a C-terminal CA domain, a nucleocapsid (NC) domain, or a combination thereof. In some embodiments essential regions or regions involved in cargo binding comprise a nucleocapsid (NC) domain, a CCHC type zinc finger motif, N-terminal conserved domain (NCD), central conserved domain (CCD), unique protein sequence or domain (UPD), KRs, PolyE, variable C-terminal sequence, (VCS), or a combination thereof.
[0235] In some embodiments, an RTL family, PEG10, RTL10, PNMA, PNMA2, PNMA5, or endo-Gag polypeptide comprises truncations or modifications of domains involved in capsid forming, nucleic acid binding, or delivery.
[0236] In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: UPD, CCD, KRs, PolyE, and VCS. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: NCD, CCD, KRs, PolyE, and VCS. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: NCD, UPD, KRs, PolyE, and VCS. In some instances, the PNMA or endo-Gag polypeptide comprises from N- terminus to C-terminus the following domains: NCD, UPD, CCD, PolyE, and VCS. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: NCD, UPD, CCD, KRs, and VCS. In some instances, the PNMA or endo- Gag polypeptide comprises from N-terminus to C-terminus the following domains: NCD, UPD, CCD, KRs, and PolyE.
[0237] In some instances, each of the domains is either directly or indirectly fused to the respective two flanking domains. In some instances, a domain of an RTL, PNMA, or endo Gag polypeptide (e.g., PEG10 or RTL10) is either directly or indirectly fused to two flankingdomains (e g., domains on the C-terminal and N-terminal sides). In some instances, the domains are arranged in an order that does not impede capsid assembly and cargo binding. In some embodiments a RTL family, PNMA, or endo-Gag polypeptide (e.g., PEG10 or RTL10) disclosed herein comprises a tag, for example, a maltose binding protein (MBP tag), or an epitope tag or an affinity tag, such as a polyhistidine tag. In some embodiments a RTL family, PNMA, or endoGag polypeptide disclosed herein lacks a tag, for example, does not contain an epitope tag, does not contain an affinity tag, or does not contain a polyhistidine tag. In some embodiments a tag is initially present but is removed, e.g., by enzymatic digestion. In some embodiments a RTL family, PNMA, or endo-Gag polypeptide disclosed herein comprises an N-terminal methionine. In some embodiments a RTL, PNMA, or endo-Gag polypeptide disclosed herein lacks an N- terminal methionine. In some embodiments an N-terminal methionine is removed, co- translationally and / or via enzymatic digestion.
[0238] In some instances, the RTL or PNMA family or endo-Gag polypeptide (e.g., PEG10 or RTL 10) is an engineered PNMA or endo-Gag polypeptide. As used herein, an engineered polypeptide is a recombinant polypeptide that is not identical in sequence to a full length, wildtype polypeptide. An engineered RTL, PNMA, or endo-Gag polypeptide (e.g., PEG10 or RTL 10) can be functionally distinct from a native RTL, PNMA, or endo-Gag polypeptide, for example, comprising an additional function than the native polypeptide or lacking a function that the native polypeptide has. In some instances, the RTL, PNMA, or endo-Gag polypeptide (e.g., PEG10 or RTL10) is a native RTL, PNMA or endo-Gag polypeptide, for example, comprising, consisting essentially of, or consisting of a wild type amino acid sequence.
[0239] In some instances, a PNMA or endo-Gag polypeptide (e.g., an RTL family, such as PEG10 or RTL 10, or PNMA2, or PNMA5) is a human endo-Gag or PNMA polypeptide. In some instances, an endo-Gag or PNMA is a non-human endo-Gag or PNMA polypeptide. In additional instances, the endo-Gag or PNMA polypeptide comprises one or more domains of a human endo-Gag or PNMA polypeptide, in which at least one of the domains participates in the formation of a capsid. In additional instances, the endo-Gag or PNMA polypeptide comprises one or more domains of a non-human endo-Gag or PNMA polypeptide, in which at least one of the domains participates in the formation of a capsid.
[0240] In some instances, an engineered RTL, PNMA or endo-Gag polypeptide comprises a fragment of a RTL, PNMA or endo-Gag polypeptide from a first species and at least an additional fragment from a RTL, PNMA or endo-Gag polypeptide of a second species.
[0241] In some embodiments, an illustrative mammalian RTL, PNMA or endo-Gag protein for expression as a recombinant or engineered RTL, PNMA or endo-Gag polypeptide is from the species homo sapiens. Additional illustrative species of primate RTL, PNMA or endo-Gag proteins for expression as a recombinant or engineered RTL, PNMA or endo-Gag polypeptide include: gorilla, pongo abelii, pa paniscus, macaca nemestrina, chlorocebus sabaeus, papio anubis, rhinopithecus roxellana, macaca fascicularis, nomascus leucogenys, callithrix jacchus, aotus nancymaae, cebus capucinus imitator, saimiri boliviensis boliviensis, otolemur garnettii, macaca mulatto, and macaca fascicularis.
[0242] Illustrative species of rodent RTL, PNMA or endo-Gag proteins for expression as a recombinant or engineered RTL, PNMA or endo-Gag polypeptide includes: fukomys damarensis, microcebus murinus, heterocephalus glaber, propithecus coquereli, marmota marmota marmota, galeopterus variegatus, cavia porcellus, dipodomys ordii, octodon degus, castor canadensis nannospalax galili, carlito syrichta, chinchilla lanigera, mus musculus, ictidomys tridecemlineatus, rattus norvegicus, microtus ochrogaster, otolemur garnettii, meriones unguiculatus, cricetulus griseus, rattus norvegicus, neotoma lepida, jaculus jaculus, mustela putorius furo, mesocricetus auratus, tupaia chinensis, cricetulus griseus, chrysochloris asiatica, elephantulus edwardii, erinaceus europaeus, ochotona princeps, sorex araneus, monodelphis domestica, echinops telfairi, and condylura cristata.A. Cargo binding domain
[0243] In some instances, the RTL family, PNMA or endo-Gag polypeptide (e.g., PEG10 or RTL10) comprises a cargo binding domain. In some embodiments, an RTL family, PNMA or endo-Gag polypeptide (e.g., PEG10 or RTL10) is engineered to comprise a cargo binding domain, for example, through insertion of a cargo binding domain, or modification of a cargo binding domain that is endogenous to the RTL, PNMA or endo-Gag polypeptide. In some embodiments, the cargo binding domain is endogenous to the RTL, PNMA or endo-Gag polypeptide.
[0244] A cargo binding domain can bind a cargo covalently or non-covalently, e.g., via a linker disclosed herein. A linker can be a peptide linker, for example, a peptide linker that bindsthe cargo binding domain to a peptide or polypeptide cargo. A linker can be a non-peptide linker disclosed herein that binds to the cargo via a covalent bond or a non-covalent bond as disclosed herein. A Cargo binding domain can be joined to an RTL, PNMA or endo-Gag polypeptide covalently or non-covalently, e.g., via a linker disclosed herein.
[0245] In some embodiments a cargo binding domain binds a heterologous cargo, for example, a cargo that is not native to capsids formed from the RTL, PNMA or endo-Gag polypeptide in nature.
[0246] In some cases, the cargo binding domain comprises a nucleic acid binding domain, an RNA binding domain, a DNA binding domain, a protein binding domain, a peptide binding domain, an antibody binding domain, a small molecule binding domain, or a peptidomimetic / nucleotidomimetic binding domain. Illustrative cargo binding domains include, but are not limited to, synthetic nucleic acid (e g., RNA and / or DNA) binding domains, zinc finger domains, arginine-rich domains, domains from GPCRs, antibodies or binding fragments thereof, lipoproteins, integrins, tyrosine kinases, DNA-binding proteins, RNA-binding proteins, nucleases, ligases, proteases, integrases, isomerases, phosphatases, GTPases, aromatases, esterases, adaptor proteins, G-proteins, GEFs, cytokines, interleukins, interleukin receptors, interferons, interferon receptors, caspases, transcription factors, neurotrophic factors and their receptors, growth factors and their receptors, signal recognition particle and receptor components, extracellular matrix proteins, integral components of membrane, ribosomal proteins, translation elongation factors, translation initiation factors, GPI-anchored proteins, tissue factors, dystrophin, utrophin, dystrobrevin, any fusions, combinations, subunits, derivatives, or domains thereof. In some embodiments, a cargo binding domain comprises a domain from an endogenous Gag polypeptide, RTL family polypeptide, or a PNMA family polypeptide that is heterologous with respect to the remainder of the protein or the part of the protein that induces capsid formation. For example, a PEG10, RTL10, PNMA2 or PNMA5 polypeptide can be engineered to comprise a cargo-binding domain that is from or derived from a heterologous endo-Gag polypeptide, such as a heterologous PNMA protein, e.g., PNMA3 or another endo-Gag protein. A zinc finger domain in an endo Gag polypeptide disclosed herein can be a nucleic acid binding domain with specificity of a 3’ UTR of the cognate gene.
[0247] In some embodiments, a cargo biding domain or a nucleic acid binding domain is from or derived from a viral protein, for example, TAT protein of HIV. In some embodiments, acargo biding domain is inverted, for example, compared to a wild type version of the cargo binding domain.
[0248] A Cargo binding domain can be a synthetic cargo binding domain, for example, a synthetic nucleic acid binding domain. In some embodiments, a synthetic cargo binding domain is designed such that it comprises an extension of a C-terminal alpha helix of a RTL or PNMA polypeptide or endo-Gag polypeptide. In some embodiments, a synthetic cargo binding domain is designed such that the cargo binding domain is oriented to the interior of an assembled RTL, PNMA or endo-Gag capsid.
[0249] A nucleic acid binding domain can be from or derived from an RNA binding protein. A nucleic acid binding domain can be from or derived from a DNA binding protein.
[0250] In some embodiments, a cargo binding domain is lysine-rich, for example, comprising at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% lysine residues. In some embodiments, a cargo binding domain is arginine-rich, for example, comprising at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, or at least 40% arginine residues. In some embodiments, a cargo binding domain comprises one or more Lys / Arg (RK) repeats, for example, at least 1, at least 2, at least 3, at least 4, at least 5, at most 2, at most 3, at most 4, at most 5, about 1, about 2, about 3, about 4, or about 5 Lys / Arg repeats.
[0251] Illustrative, non-limiting examples of cargo binding domains that are nucleic acidbinding domains are the amino acid sequences SEQ ID NO: 39, SEQ ID NO: 42, and SEQ ID NO: 45, and amino acid sequences that comprise at least 70%, at least 80%, at least 85%, at least 90%, or at least 95% sequence identity to any one of SEQ ID NOs: 39, 72, and 45.
[0252] In some cases, the cargo binding domain binds to an RNA, for example, an mRNA, hairpin RNA, guide RNA, shRNA, siRNA, an mRNA, a tRNA, an rRNA, a snRNA, a microRNA, or a non-coding RNA. In some cases, the cargo binding domain binds to a DNA, for example, a ssDNA, dsDNA, or oligonucleotide. In some embodiments, a cargo binding domain binds to RNA and DNA. In some embodiments, a cargo binding domain preferentially binds RNA over DNA. In some embodiments, a cargo binding domain preferentially binds DNA over RNA. In some embodiments, a cargo binding domain preferentially binds ssDNA. In some embodiments, a cargo binding domain specifically or preferentially binds a particular nucleicacid structural motif, for example, a hairpin, such as a bulge hairpin. In some embodiments, a cargo binding domain non-specifically binds nucleic acids, RNA, and / or DNA, such as ssDNA.
[0253] In some embodiments, the PNMA family polypeptide is an engineered PNMA polypeptide with at least an RNA binding domain inserted and / or modified to bind to a heterologous cargo that is not native to capsids formed from the PNMA polypeptide in nature. In some instances, the PNMA polypeptide comprises a full-length PNMA polypeptide with at least an RNA binding domain inserted and / or modified to bind to a heterologous cargo that is not native to capsids formed from the PNMA protein in nature. In additional instances, the engineered PNMA polypeptide comprises one or more domains of a PNMA polypeptide, in which at least one of the domains participates in the formation of a capsid and in which an RNA binding domain is inserted and / or modified to bind to a heterologous cargo that the native PNMA protein does not bind to.
[0254] In some embodiments, an endo-Gag disclosed herein is an engineered Endo-Gag polypeptide with at least an RNA binding domain inserted and / or modified to bind to a heterologous cargo that is not native to capsids formed from the endo-Gag protein in nature. In some instances, the endo-Gag polypeptide comprises a full-length endo-Gag polypeptide with at least its RNA binding domain modified to bind to a heterologous cargo that is not native to capsids formed from the endo-Gag protein in nature. In other instances, the endo-Gag polypeptide comprises an engineered endo-Gag fragment comprising modification(s) in at least its RNA binding domain to bind to a heterologous cargo that a native endo-Gag protein does not bind to. In additional instances, the endo-Gag polypeptide comprises one or more domains of an engineered endo-Gag polypeptide, in which at least one of the domains participates in the formation of a capsid and in which the RNA binding domain is modified to bind to a heterologous cargo that is not native to capsids formed from the endo-Gag protein in nature.
[0255] In some embodiments, a cargo binding domain is at a C-terminus of an endo-Gag polypeptide or RTL or PNMA polypeptide disclosed herein. In some embodiments, a cargo binding domain is at an N-terminus of an endo-Gag polypeptide or RTL or PNMA polypeptide disclosed herein. In some embodiments, a cargo binding domain is within the sequence of an endo-Gag polypeptide or RTL or PNMA polypeptide disclosed herein.
[0256] In some embodiments, a cargo binding domain can be inserted adjacent to a domain of an RTL, PNMA or endo-Gag polypeptide. For example, in some embodiments, a cargobinding domain can be inserted adjacent to (e.g., N-terminal or C-terminal to) one or more of an N-terminal CA domain, C terminal CA domain, NCD, CCD, UPD, KRs, PolyE, or VCS domain.
[0257] In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: cargo binding domain, and KRs. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: KRs, and cargo binding domain.
[0258] In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: cargo binding domain, KRs, and PolyE. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: KRs, cargo binding domain, and PolyE. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: KRs, PolyE, and cargo binding domain. In some instances, each of the domains is either directly or indirectly fused to the respective two flanking domains.
[0259] In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: cargo binding domain, NCD, UPD, CCD, KRs, PolyE, and VCS. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C- terminus the following domains: NCD, cargo binding domain, UPD, CCD, KRs, PolyE, and VCS. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C- terminus the following domains: NCD, UPD, cargo binding domain, CCD, KRs, PolyE, and VCS. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C- terminus the following domains: NCD, UPD, CCD, cargo binding domain, KRs, PolyE, and VCS. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C- terminus the following domains: NCD, UPD, CCD, KRs, cargo binding domain, PolyE, and VCS. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C- terminus the following domains: NCD, UPD, CCD, KRs, PolyE, cargo binding domain, and VCS. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C- terminus the following domains: NCD, UPD, CCD, KRs, PolyE, VCS, and cargo binding domain. In some instances, each of the domains is either directly or indirectly fused to the respective two flanking domains.
[0260] In certain embodiments, a domain of a RTL, PNMA or endo-Gag polypeptide is replaced, partially replaced, or acts as an insertion site for a cargo binding domain. For example,in some embodiments, one or more of an N-terminal CA domain, C terminal CA domain, NCD, CCD, UPD, KRs, PolyE, or VCS domain can be replaced, partially replaced, or act as an insertion site for a cargo binding domain.
[0261] In some embodiments, the cargo binding domain comprises a sequence that binds to an nucleic acid secondary structure. In some embodiments, the cargo binding domain comprises a sequence that binds to a hairpin loop element. A hairpin may be capable of forming more than one loop. For example, a hairpin capable of forming two intramolecular duplexes and two loops is referred to herein as a "double hairpin". In some embodiments, the cargo binding domain comprises a sequence that binds to an aptamer sequence. For example, a cargo binding domain can comprise a binding domain of MS2, PP7, QP, F2, GA, fr, JP501, M12, R17, BZ13, JP34, JP500, KU1, Mi l, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, Cb5, <^Cb8r, (^Cbl2r, (|)Cb23r, 7s, or PRRl.
[0262] In some embodiments, a cargo binding domain sequence comprises folate (e.g, to bind folate receptor or a fragment thereof on a cargo), transferrin (e.g, to bind transferrin or a fragment thereof on a cargo), antibody CC52 (e.g, to bind rat CC531 on a cargo), anti-HER2 (e g, to bind HER2 or a fragment thereof on a cargo), anti-GD2 (e g, to bind GD2 or a fragment thereof on a cargo), anti-EGFR (e.g, to bind EGFR or a fragment thereof on a cargo), anti-VEGF or anti-VEGFR (e.g, to bind VEGF, VEGFR or a fragment thereof on a cargo), anti-CD19 (e.g, to bind CD 19 or a fragment thereof on a cargo), cyclic arginine-glycine-aspartic acid-tyrosine- cysteine peptide (e.g, to bind av 3 or a fragment thereof on a cargo), PR b peptide (e.g, to bind alpha-5 beta-1 integrin or a fragment thereof on a cargo), AG86 (e.g, to bind alpha 6 beta 4 integrin or a fragment thereof on a cargo), P6.1 peptide (e.g, to bind X HER2 receptor X or a fragment thereof on a cargo), affinity peptide LN (e.g, to bind aminopeptidase N or a fragment thereof on a cargo), a synthetic somatostatin analogue (e.g, to bind SSTR2 or a fragment thereof on a cargo), anti-CD20 (e.g, to bind CD20 or a fragment thereof on a cargo), or vice versa (e.g., the cargo binding domain comprises the domain indicated as present on the cargo, and the cargo comprises the domain indicated for the cargo binding domain.
[0263] In some embodiments, a cargo binding domain sequence comprises lambdaN (e.g, to bind BoxB or a fragment thereof on a cargo), L7A peptide (e.g, to bind C / D box or a fragment thereof on a cargo), an M52 domain (e.g, to bind MS2 RNA or a fragment thereof on a cargo), TAT peptide (e.g, to bind TAR RNA or a fragment thereof on a cargo), BC2 nanobody (e.g, tobind BC2 peptide or a fragment thereof on a cargo), Sun tag antibody (e g, to bind sun tag or a fragment thereof on a cargo), a SUMO domain (e.g, to bind SIM or a fragment thereof on a cargo), 4WW (e.g, to bind PPXY or a fragment thereof on a cargo), or vice versa (e.g., the cargo binding domain comprises the domain indicated as present on the cargo, and the cargo comprises the domain indicated for the cargo binding domain.B. Exogenous polypeptide sequence
[0264] In some embodiments, an engineered RTL, PNMA or endo-Gag polypeptide (e.g., PEG10 or RTL 10) is modified to comprise an exogenous polypeptide sequence, e.g., in addition to the cargo-binding domain. The exogenous polypeptide sequence can be fused to the RTL, PNMA or endo-Gag polypeptide. The exogenous polypeptide sequence can be covalently or non-covalently attached to the RTL, PNMA or endo-Gag polypeptide, e g., via a linker disclosed herein.
[0265] In some embodiments, the exogenous polypeptide sequence can comprise a sub- cellular localization signal, for example, a nuclear localization signal (NLS), or a sequence that targets the polypeptide to a membrane (e.g., an Arginine-rich domain). In some embodiments, the exogenous polypeptide sequence comprises a domain that binds to a cell surface molecule, for example, an antigen, a polypeptide, a receptor, a lipid (e.g., phospholipid), a lipoprotein, a glycoprotein, or the like. In some embodiments, an exogenous polypeptide recognizes and binds to receptors displayed on the surface of targeted cells. Upon reaching a cell of interest, the cargo is optionally further delivered to an intracellular target. For example, a therapeutic RNA can be translated to a protein if it comes into contact with a ribosome in the cytoplasm of the cell.
[0266] In some embodiments, the cell surface molecule is a cell surface protein. In some instances, the cell surface molecule is an antigen expressed by a cancerous cell. In some instances, the cell surface molecule is a neoepitope. In some instances, the cell surface molecule comprises one or more mutations compared to a wild-type protein. Illustrative cancer antigens that can be bound by an exogenous polypeptide sequence include, but are not limited to, alpha fetoprotein, ASLG659, B7-H3, BAFF-R, Brevican, CA125 (MUC16), CA15-3, CA19-9, carcinoembryonic antigen (CEA), CA242, CRIPTO (CR, CR1, CRGF, CRIPTO, TDGF1, teratocarcinoma-derived growth factor), CTLA-4, CXCR5, E16 (LAT1, SLC7A5), FcRH2 (IFGP4, IRTA4, SPAP1 A (SH2 domain containing phosphatase anchor protein la), SPAP1B, SPAP1C), epidermal growth factor, ETBR, Fc receptor-like protein 1 (FCRH1), GEDA, HLA-DOB (Beta subunit of MHC class II molecule (la antigen), human chorionic gonadotropin, ICOS, IL-2 receptor, IL20Ra, Immunoglobulin superfamily receptor translocation associated 2 (IRTA2), L6, Lewis Y, Lewis X, MAGE-1, MAGE-2, MAGE-3, MAGE 4, MARTI, mesothelin, MDP, MPF (SMR, MSLN), MCP1 (CCL2), macrophage inhibitory factor (MIF), MPG, MSG783, mucin, MUC1-KLH, Napi3b (SLC34A2), nectin-4, Neu oncogene product, NCA, placental alkaline phosphatase, prostate specific membrane antigen (PMSA), prostatic acid phosphatase, PSCA hlg, anti-transferrin receptor, p97, Purinergic receptor P2X ligand-gated ion channel 5 (P2X5), LY64 (Lymphocyte antigen 64 (RP105), gplOO, P21, six transmembrane epithelial antigen of prostate (STEAP1), STEAP2, Serna 5b, tumor-associated glycoprotein 72 (TAG-72), TrpM4 (BR22450, FLJ20041, TRPM4, TRPM4B, transient receptor potential cation channel, subfamily M, member 4) and the like.
[0267] In some instances, the cell surface molecule bound by an exogenous polypeptide sequence comprises a cluster of differentiation (CD) cell surface marker, for example, CD1, CD2, CD3, CD4, CD5, CD6, CD7, CD8, CD9, CD10, CDl la, CDl lb, CDl lc, CDl ld, CDwl2, CD13, CD14, CD15, CD15s, CD16, CDwl7, CD18, CD19, CD20, CD21, CD22, CD23, CD24, CD25, CD26, CD27, CD28, CD29, CD30, CD31, CD32, CD33, CD34, CD35, CD36, CD37, CD38, CD39, CD40, CD41, CD42, CD43, CD44, CD45, CD45RO, CD45RA, CD45RB, CD46, CD47, CD48, CD49a, CD49b, CD49c, CD49d, CD49e, CD49f, CD50, CD51, CD52, CD53, CD54, CD55, CD56, CD57, CD58, CD59, CDw60, CD61, CD62E, CD62L (L-selectin), CD62P, CD63, CD64, CD65, CD66a, CD66b, CD66c, CD66d, CD66e, CD71, CD79 (e.g., CD79a, CD79b), CD90, CD95 (Fas), CD103, CD104, CD125 (IL5RA), CD134 (0X40), CD137 (4- 1BB), CD152 (CTLA-4), CD221, CD274, CD279 (PD-1), CD319 (SLAMF7), CD326 (EpCAM), or the like.
[0268] In some cases, the exogenous polypeptide sequence is or is derived from a protein (e.g., a human protein), an antibody or binding fragment thereof, a viral protein, a Gag-like protein (e.g., a human Gag-like protein), or a de novo engineered protein designed to bind to a target receptor of interest. In some instances, the antibody or binding fragment thereof comprises a humanized antibody or binding fragments thereof, a murine antibody or binding fragment thereof, a chimeric antibody or binding fragment thereof, a monoclonal antibody or binding fragment thereof, a multi-specific antibody or binding fragment thereof, a bispecific antibody or biding fragment thereof, a monovalent Fab’, a divalent Fab2, F(ab)’3 fragments, a single-chainvariable fragment (scFv), a bis-scFv, an (scFv)2, a diabody, a minibody, a nanobody, a triabody, a tetrabody, a disulfide stabilized Fv protein (dsFv), a single-domain antibody (sdAb), an Ig NAR, a camelid antibody or binding fragment thereof (e.g., VHH domain), or a chemically modified derivative thereof. In some instances, the exogenous polypeptide sequence guides the delivery of a capsid formed by the engineered RTL, PNMA polypeptide to a target site of interest.
[0269] An antibody or an antigen-binding fragment thereof disclosed herein (e.g., as an exogenous polypeptide sequence or a cargo) can comprise complementarity determining regions (CDRs). In some embodiments, the CDRs determine or substantially determine binding specificity and / or affinity of the antibody or antigen-binding fragment. For example, the CDRs can be grafted onto a different suitable framework, or the framework region can be altered (e.g., via amino acid substitutions, deletions, and / or insertions), and the antigen-binding fragment or domain can retain binding for the target, and the extracellular binding domain remains functional despite the alterations outside of the CDRs. CDRs can be identified by various methods, including but not limited to the Kabat method, the Chothia method, the IMGT method, the AHO method, and the Paratome method. One antigen binding site of an antibody with heavy and light chains or variable regions therefrom comprises six CDRs, three in the hypervariable regions of the light chain variable region, and three in the hypervariable regions of the heavy chain variable region. The CDRs in the light chain are designated LI, L2, and L3, while the CDRs in the heavy chain are designated Hl, H2, and H3. CDRs can also be designated LCDR1, LCDR2, LCDR3, HCDR1, HCDR2, and HCDR3, respectively. Certain antibodies or antigen-binding domains contain less than six CDRs. For example, certain antibodies lack a light chain, and can be referred to as heavy chain only antibodies (HCAbs). HCAbs have three CDRs in a variable region referred to as VHH. A single domain antibody, or nanobody, can be generated from such a VHH region of a heavy chain only antibody.
[0270] Non-liming examples of exogenous polypeptide sequences include those that encode zinc finger domains, arginine-rich domains, domains from GPCRs, antibodies or binding fragments thereof, lipoproteins, integrins, tyrosine kinases, DNA-binding proteins, RNA-binding proteins, nucleases, ligases, proteases, integrases, isomerases, phosphatases, GTPases, aromatases, esterases, adaptor proteins, G-proteins, GEFs, cytokines, interleukins, interleukin receptors, interferons, interferon receptors, caspases, transcription factors, neurotrophic factorsand their receptors, growth factors and their receptors, signal recognition particle and receptor components, extracellular matrix proteins, integral components of membrane, ribosomal proteins, translation elongation factors, translation initiation factors, GPI-anchored proteins, tissue factors, dystrophin, utrophin, dystrobrevin, cell penetrating peptides, fusogenic proteins, viral envelope proteins, endogenous retroviral envelope proteins, any fusions, combinations, subunits, derivatives, or domains thereof.
[0271] An exogenous polypeptide sequence can be or can comprise a cell penetrating peptide (CPP). Cell-penetrating peptides (CPPs) can facilitate uptake of macromolecules through cellular membranes and enhance the delivery of CPP-modified molecules to the inside of a cell. A cellpenetrating peptide can comprise or can be a cationic CPP (e.g., TAT, R8, DPV3, DPV6, Penetratin, R9-TAT), an amphipathic CPP (e.g., pVEC, ARF, MPG, MAP, transportan), or a hydrophobic CPP (e.g., Bip4, C105Y, Melttin, gH625). A CPP can be a protein-derived CPP, a synthetic CPP, or a chimeric CPP. CPPs can be comprise amphipathic helical peptides, such as transportan and MAP, where lysine residues are major contributors to the positive charge, and Arg-rich peptides, such as TATp, Antennapedia or penetratin. Other CPPs can include: the minimal protein transduction domain of Antennapedia, a Drosophilia homeoprotein, called penetratin, which is a 16-mer peptide (residues 43-58) present in the third helix of the homeodomain; a 27-amino acid-long chimeric CPP, containing the peptide sequence from the amino terminus of the neuropeptide galanin bound via the Lys residue, mastoparan, a wasp venom peptide; VP22, a major structural component of HSV-1 facilitating intracellular transport, and transportan (18-mer) amphipathic model peptide that translocates plasma membranes of mast cells and endothelial cells by both energy-dependent and - independent mechanisms. In some embodiments, a lipid moiety is modified with CPP(s), for intracellular.
[0272] In some embodiments, the exogenous polypeptide sequence comprises a non-native cysteine residue, for example, a cysteine residue that is not present in a native form of the RTL or PNMA polypeptide or endo-Gag polypeptide (e.g., RTL10 or PEG10). The exogenous cysteine residue can be used, for example, for conjugation to a binging partner via suitable chemical reactions, such as Maleimide chemistry. In some embodiments, the exogenous polypeptide sequence comprises a non-native cysteine residue and lacks one or more native cysteine residues, for example, the one or more native cysteine residues are deleted or substituted for non-cysteine residues. The non-native cysteine residue or an exogenous polypeptide sequencecomprising the non-native cysteine residue can be fused directly or via a linker to a C-terminus, N-terminus, or within the RTL or PNMA polypeptide or endo Gag polypeptide, e.g., within or between domains of the RTL or PNMA polypeptide or endo Gag polypeptide as disclosed herein.
[0273] In some instances, the exogenous polypeptide sequence is fused directly, indirectly via a linker, or chemically to one or more of: a cargo binding domain, N-terminal CA domain, C terminal CA domain, NCD, UPD, CCD, KRs, PolyE, or VCS, if present. In some instances, the exogenous polypeptide sequence is fused directly, indirectly via a linker, or chemically to a C- terminus of the RTL, PNMA or endo-Gag polypeptide. In some instances, the exogenous polypeptide sequence is fused directly, indirectly via a linker, or chemically to an N-terminus of the RTL, PNMA or endo-Gag polypeptide.
[0274] In some embodiments, an exogenous polypeptide sequence can be inserted adjacent to a domain of a PNMA or endo-Gag polypeptide. For example, in some embodiments, an exogenous polypeptide sequence can be inserted adjacent to (e.g., N-terminal or C-terminal to) one or more of an NCD, CCD, UPD, KRs, PolyE, or VCS domain.
[0275] In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: exogenous polypeptide sequence, and PolyE. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: PolyE, and exogenous polypeptide sequence.
[0276] In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: exogenous polypeptide sequence, and KRs. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: KRs, and exogenous polypeptide sequence.
[0277] In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: exogenous polypeptide sequence, KRs, and PolyE. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: KRs, exogenous polypeptide sequence, and PolyE. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: KRs, PolyE, and exogenous polypeptide sequence. In some instances, each of the domains is either directly or indirectly fused to the respective two flanking domains.
[0278] In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: exogenous polypeptide sequence, NCD, UPD, CCD, KRs, PolyE, and VCS. In some instances, the PNMA or endo-Gag polypeptide comprises from N- terminus to C-terminus the following domains: NCD, exogenous polypeptide sequence, UPD, CCD, KRs, PolyE, and VCS. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: NCD, UPD, exogenous polypeptide sequence, CCD, KRs, PolyE, and VCS. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: NCD, UPD, CCD, exogenous polypeptide sequence, KRs, PolyE, and VCS. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: NCD, UPD, CCD, KRs, exogenous polypeptide sequence, PolyE, and VCS. In some instances, the PNMA or endo- Gag polypeptide comprises from N-terminus to C-terminus the following domains: NCD, UPD, CCD, KRs, PolyE, exogenous polypeptide sequence, and VCS. In some instances, the PNMA or endo-Gag polypeptide comprises from N-terminus to C-terminus the following domains: NCD, UPD, CCD, KRs, PolyE, VCS, and exogenous polypeptide sequence. In some instances, each of the domains is either directly or indirectly fused to the respective two flanking domains.
[0279] In certain embodiments, a domain of a RTL, PNMA or endo-Gag polypeptide is replaced, partially replaced, or acts as an insertion site for an exogenous polypeptide sequence. For example, in some embodiments, one or more of an N-terminal CA domain, C terminal CA domain, NCD, CCD, UPD, KRs, PolyE, or VCS domain can be replaced, partially replaced, or act as an insertion site for an exogenous polypeptide sequence.C. Linker
[0280] In certain embodiments, a polypeptide or capsid of the disclosure comprises a linker, for example, a peptide or non-peptide linker that joins two components covalently or non- covalently.
[0281] A linker can join a first component (e.g., domain, polypeptide, cargo binding domain, cargo, delivery component, or exogenous polypeptide sequence) to a second component (e.g., domain, polypeptide, cargo binding domain, cargo, delivery component, or exogenous polypeptide sequence).
[0282] A linker can join a cargo binding domain to a cargo (e.g., covalently or non- covalently). A linker can join a first cargo binding domain to a second cargo binding domain. Alinker can join a cargo binding domain to a different domain. A linker can join a cargo binding domain to a polypeptide (e.g., a RTL or PNMA polypeptide or endo-Gag polypeptide). A linker can join a cargo binding domain to an exogenous polypeptide sequence.
[0283] A linker can join a cargo to a cargo binding domain. A linker can join a first cargo to a second cargo. A linker can join a cargo to a domain. A linker can join a cargo to a polypeptide (e.g., a RTL or PNMA polypeptide or endo-Gag polypeptide). A linker can join a cargo to an exogenous polypeptide sequence.
[0284] A linker can join an exogenous polypeptide sequence to a polypeptide (e g., a RTL or PNMA polypeptide or endo-Gag polypeptide). A linker can join a first exogenous polypeptide sequence to a second exogenous polypeptide sequence. A linker can join an exogenous polypeptide sequence to a domain. A linker can join an exogenous polypeptide sequence to a cargo binding domain. A linker can join an exogenous polypeptide sequence to a cargo.
[0285] A linker can join a first domain to a second domain (e.g., a cargo binding domain). A linker can join a domain to a polypeptide (e.g., a RTL or PNMA polypeptide or endo-Gag polypeptide). A linker can join a domain to a cargo. A linker can join a domain to an exogenous polypeptide sequence.
[0286] A linker can join a first polypeptide (e.g., a RTL or PNMA polypeptide or end-Gag polypeptide) to a second polypeptide. A linker can join a polypeptide (e.g., a RTL or PNMA polypeptide or endo-Gag polypeptide) to a domain. A linker can join a polypeptide (e.g., a RTL or PNMA polypeptide or endo-Gag polypeptide) to a cargo binding domain. A linker can join a polypeptide (e.g., a RTL or PNMA polypeptide or end-Gag polypeptide) to a cargo. A linker can join a polypeptide (e.g., a RTL or PNMA polypeptide or end-Gag polypeptide) to an exogenous polypeptide sequence.
[0287] A linker can join any two components of a polypeptide, capsid, complex, or composition disclosed herein. For example, a linker can join a RTL or PNMA polypeptide to a cargo, an endo-Gag polypeptide to a cargo, a RTL or PNMA polypeptide to a cargo-binding domain, an endo-Gag polypeptide to a cargo-binding domain, a RTL or PNMA polypeptide to an exogenous polypeptide sequence, an endo Gag polypeptide to an exogenous polypeptide sequence, a RTL or PNMA polypeptide to a delivery component, an endo-Gag polypeptide to a delivery component, or any other two domains disclosed herein. A linker can join two domains within a RTL or PNMA polypeptide, two domains within a cargo, to domains within a cargo-binding domain, two domains within an exogenous polypeptide sequence, two domains within a delivery component, and the like. A polypeptide, capsid, complex, or composition can comprise multiple linkers. For example, a cargo, cargo-binding domain, or exogenous polypeptide sequence that comprises an antigen-binding fragment can comprise two variable regions joined by a linker, such as an scFv. A polypeptide or domain that comprises one or more linkers can be joined to another polypeptide or domains via another linker that is the same or different, for example, a cargo, cargo binding domain, or exogenous polypeptide sequence that is an scFv can comprise a linker between a VH and VL domain, and the scFv can in turn be joined to a RTL, PNMA or endo-Gag polypeptide by a second linker.
[0288] In some embodiments, the linker is a peptide linker. In some instances, the linker is a rigid linker. In other instances, the linker is a flexible linker. In some cases, the linker is a non- cleavable linker. In other cases, the linker is a cleavable linker. In additional cases, the linker comprises a linear structure, or a non-linear structure (e.g., a cyclic structure).
[0289] A flexible peptide linker can have a sequence containing stretches of glycine and serine residues. The small size of the glycine and serine residues provides flexibility and allows for mobility of the connected functional domains. The incorporation of serine or threonine can maintain the stability of the linker in aqueous solutions by forming hydrogen bonds with the water molecules, thereby reducing unfavorable interactions between the linker and protein moieties. Flexible linkers can also contain additional amino acids such as threonine and alanine to maintain flexibility, as well as polar amino acids such as lysine and glutamine to improve solubility. A rigid peptide linker can have, for example, an alpha helix-structure. An alphahelical rigid linker can act as a spacer between certain protein domains. A peptide linker can be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acid residues in length. In some cases, a linker sequence can be, for example at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 15, at least about 20, at least about 30, at least about 40, or at least about 50 amino acids in length. In some cases, a linker sequence can be, for example at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 15, at most about 20, at most about 30, at most about 40, at most about 50, at most about 60, at most about 70, atmost about 80, or at most about 100 amino acids in length. In some cases, a linker is 5-20 amino acids in length. In some cases, a linker is 10-20 amino acids in length. In some cases, a linker is 4-8 amino acids in length.
[0290] A linker can be a non-cleavable linker. A non-cleavable linker of the present disclosure can include a chemical linker that is stable. Examples of non-cleavable linkers that can be used in proteins of the present disclosure to link domains and / or polypeptides can include a thioether linker, an alkyl linker, a polymeric linker. A linker may be an SMCC linker or a PEG linker. In some embodiments, the linker may be a PEG linker. A non-cleavable linker can also include a non-proteolytically cleavable peptide linker. A non-proteolytically cleavable peptide can be inert to proteases present in a given sample, tissue, or organism. For example, a peptide can be inert or substantially inert to all or most human protease cleavage sequences, and thereby can comprise a high degree of stability within humans and human samples. Such a peptide can also comprise a secondary structure which renders a protease cleavage site inert or inaccessible to a protease. A non-cleavable linker of the present disclosure can comprise a half-life for cleavage of at least 1 hour, at least 2 hours, at least 4 hours, at least 8 hours, at least 12 hours, at least 16 hours, at least 1 day, at least 2 days, at least 3 days, at least 1 week, at least 2 weeks, or at least 1 month in the presence of human proteases at 25 °C in pH 7 buffer.
[0291] In certain embodiments, non-cleavable linkers comprise short peptides of varying lengths. Exemplary non-cleavable linkers include (EAAAK)n(SEQ ID NO: 9), or (EAAAR)n(SEQ ID NO: 10), where n is from 1 to 5, and up to 30 residues of glutamic acid-proline or lysine-proline repeats. In some embodiments, the non-cleavable linker comprises (GS)n (SEQ ID NO: 100), (SG)n (SEQ ID NO: 101), (GGGGS)n (SEQ ID NO: 11) or (GGGS)n (SEQ ID NO: 12), wherein n is 1 to 10 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10); KESGSVSSEQLAQFRSLD (SEQ ID NO: 13); or EGKSSGSGSESKST (SEQ ID NO: 14). In some embodiments, the non- cleavable linker comprises a poly-Gly / Ala polymer.
[0292] In certain embodiments, the linker is a cleavable linker, e.g., an extracellular cleavable linker or an intracellular cleavable linker. In some instances, the linker is designed for cleavage in the presence of particular conditions or in a particular environment (e.g., under physiological conditions). For example, the design of a linker for cleavage by specific conditions, such as by a specific enzyme, allows the targeting of cellular uptake to a specific location.
[0293] In some embodiments, the linker is a pH-sensitive linker. In one instance, the linker is cleaved under basic pH conditions. In other instance, the linker is cleaved under acidic pH conditions.
[0294] In some embodiments, the linker is cleaved in vivo by endogenous enzymes (e.g., proteases) such as serine proteases including but not limited to thrombin, metalloproteases, furin, cathepsin B, necrotic enzymes (e.g., calpains), and the like. For example, a cleavable linker can be used to covalently link a cargo to a cargo binding domain, and linker can be cleaved in vivo to release the cargo. Exemplary cleavable linkers include, but are not limited to, GGAANLVRGG (SEQ ID NO: 15); SGRIGFLRTA (SEQ ID NO: 16); SGRSA (SEQ ID NO: 17); GFLG (SEQ ID NO: 18); ALAL (SEQ ID NO: 19); FK; PIC(Et)F-F (SEQ ID NO: 20), where C(Et) indicates S-ethylcysteine; PR(S / T)(L / I)(S / T); DEVD (SEQ ID NO: 22); GWEHDG (SEQ ID NO: 23); RPLALWRS (SEQ ID NO: 24); or a combination thereof.
[0295] In some embodiments, the cleavable linker is conjugated to the endosomal escape moiety and to the polypeptide. In one embodiment, the endosomal escape moiety may be released from the remainder of the effector-polypeptide conjugate upon entry into a cell. In some embodiments the cleavable linker is cleaved after the capsid is endocytosed into the cell. In some embodiments, the cleavable linker is cleaved during endosomal escape.
[0296] In certain embodiments, a cleavable linker comprises a 2A linker, which can be processed into separate polypeptides co-translationally or after translation. Inclusion of a 2A linker can increase the likelihood that an appropriate ratio of components are produced (e.g., a 1 : 1, 1 :2, 1 :3, 1 :4, or 1 :5 ratio of two components). In some cases, inclusion of a 2A linker can increase the likelihood that equal or close to equal levels of two components are produced (e.g., a capsid subunit and a cargo).
[0297] A linker can be a non-peptide linker. A linker can be a chemical linker. A linker can be a chemical bond, for example, a covalent bond or a non-covalent bond. A linker of the disclosure can include a chemical linker. For example, two a first component (e.g., domain, polypeptide, cargo binding domain, cargo, or exogenous polypeptide sequence) can be joined to a second component (e.g., domain, polypeptide, cargo binding domain, cargo, or exogenous polypeptide sequence) by a chemical linker. Each chemical linker of the disclosure can be alkylene, alkenylene, alkynylene, heteroalkylene, cycloalkylene, heterocycloalkylene, arylene, or heteroarylene, any of which is optionally substituted. In some embodiments, a chemical linker ofthe disclosure can be an ester, ether, amide, thioether, or polyethyleneglycol (PEG). In some embodiments, a linker can reverse the order of the amino acids sequence in a compound where two amino acid sequences are linked, for example, so that the amino acid sequences linked by the linked are head-to-head, rather than head-to-tail. Non-limiting examples of such linkers include diesters of dicarboxylic acids, such as oxalyl diester, malonyl diester, succinyl diester, glutaryl diester, adipyl diester, pimetyl diester, fumaryl diester, maleyl diester, phthalyl diester, isophthalyl diester, and terephthalyl diester. Non-limiting examples of such linkers include diamides of dicarboxylic acids, such as oxalyl diamide, malonyl diamide, succinyl diamide, glutaryl diamide, adipyl diamide, pimetyl diamide, fumaryl diamide, maleyl diamide, phthalyl diamide, isophthalyl diamide, and terephthalyl diamide. Non-limiting examples of such linkers include diamides of diamino linkers, such as ethylene diamine, l,2-di(methylamino)ethane, 1,3- diaminopropane, l,3-di(methylamino)propane, l,4-di(methylamino)butane, 1,5- di(methylamino)pentane, l,6-di(methylamino)hexane, and pipyrizine.
[0298] Non-limiting examples of optional substituents include hydroxyl groups, sulfhydryl groups, halogens, amino groups, nitro groups, nitroso groups, cyano groups, azido groups, sulfoxide groups, sulfone groups, sulfonamide groups, carboxyl groups, carboxaldehyde groups, imine groups, alkyl groups, halo-alkyl groups, alkenyl groups, halo-alkenyl groups, alkynyl groups, halo-alkynyl groups, alkoxy groups, aryl groups, aryloxy groups, aralkyl groups, arylalkoxy groups, heterocyclyl groups, acyl groups, acyloxy groups, carbamate groups, amide groups, ureido groups, epoxy groups, and ester groups.
[0299] In some embodiments, a non-peptide linker or a chemical linker is a cleavable linker, such as a self-immolative linker.
[0300] In some embodiments, a cleavable linker comprises a chemical trigger that controls cleavage (and, e.g., release of a cargo). A chemical trigger that can be used in a linker of the disclosure can be a dipeptide trigger, cathepsin-cleavable trigger, acid-cleavable trigger, GSH- cleavable trigger, Fe(II)-cleavable trigger, enzyme-cleavable trigger, photo-responsive-cleavable trigger, bioorthogonal cleavable trigger, glycosidase cleavable triggers, phosphatase cleavable trigger, Sulfatase cleavable trigger, Hydrazone trigger, Carbonate trigger, Silyl ether trigger, Disulfide trigger, 1, 2, 4-Tri oxolane trigger, Dipeptide trigger, Triglycyl (CX) trigger, cBu-Cit trigger, P-Glucuronide trigger, -Galactoside trigger, Pyrophosphate trigger, Arylsulfate trigger,Heptamethine cyanine fluorophore trigger, O-Nitrobenzyl trigger, PC4AP trigger, or a dsProc trigger.
[0301] In some embodiments, a linker comprises a Maleimide attachment, Bis(vinylsulfonyl)piperazine attachment, N-methyl-N-phenylvinylsulfonamide attachment, Pt(II)-based attachment, carbamate attachment, carbonate attachment, quaternary ammonium attachment.
[0302] In some embodiments, a linker comprises a bond generated by carbonyl condensation, a Staudinger reaction, modified Staudinger ligation, traceless Staudinger ligation, Inverse-electron demand Diels-Alter cycloadditions, Copper-catalyzed azide-alkyne cycloaddition (CuAAC), strain-promoted azide-alkyne cycloaddition (SPAAC), Huisgen cycloaddition, click chemistry, or the like.
[0303] In some embodiments, a linker comprises a bond generated between cysteine and a cysteine reactive group, such as malemide or iodoacetamide. For example, cysteine-based conjugation can be used to link two components, such as a cargo binding domain to a cargo. One or more reactive cysteine residue(s) can be introduced at selected positions, for example, via insertion or substitution at a C-terminus of a RTL, PNMA or endo-Gag polypeptide, an N- terminus of a RTL, PNMA or endo-Gag polypeptide, or within the RTL, PNMA or endo-Gag polypeptide (e.g., in a domain disclosed herein or between to domains). The cysteine residue(s) introduced can be joined to another part of a polypeptide via a linker or spacer disclosed herein. The cysteine residue can be introduced such that it does not interfere with structure or function of the RTL, PNMA or endo-Gag polypeptide, e.g., the ability to assemble into a capsid state or form, disassemble into a non-capsid state or form, and / or re-assemble into a capsid state or form. Cysteine residues can optionally be deleted or substituted for non-reactive or less-reactive residues elsewhere to reduce or eliminate reactivity or conjugation at undesirable sites.
[0304] In some embodiments, a linker comprises a covalent attachment of two components. In some embodiments, a linker comprises a non-covalent attachment of two components. Two components present in a polypeptide, complex, or capsid can be non-covalently coupled, for example, by ionic bonds, hydrogen bonds, interactions mediated by oligomerization or dimerization domains, etc.III. CAPSIDS
[0305] In some embodiments, disclosed herein is a capsid. In some instances, the capsid comprises an endo-Gag polypeptide, for example, a RTL or PNMA family polypeptide, such as a PEG10, RTL 10 (BOP), PNMA2 or PNMA5 polypeptide. Illustrative endo-Gag polypeptides include PNMA1, PNMA2, PNMA3, PNMA4 (M0AP1), PNMA5, PNMA6A, PNMA6B / 6D, PNMA6E, PNMA6F, PNMA7A, PNMA7B, PNMA8A, PNMA8B, PNMA8C, CCDC8, Arc, BOP, LDOC1, PEG10, RTL3, RTL6, RTL8A, RTL8B, and ZNF18. In some instances, the polypeptide is a functional fragment, e.g., that is capable of forming a subunit of a capsid. RTL, PNMA or end-Gag polypeptides of the disclosure can self-assemble to form capsids, also referred to as Virus-like particles (VLPs).
[0306] In some embodiments, the capsid comprises a PNMA5-based capsid. In some instances, the PNMA5-based capsid comprises a plurality of recombinant PNMA5 polypeptides of the disclosure. In some instances, the PNMA5-based capsid comprises a plurality of engineered PNMA5 polypeptides of the disclosure. In some embodiments, the PNMA5 polypeptides are recombinant and engineered.
[0307] In some embodiments, the capsid comprises a PNMA2 -based capsid. In some instances, the PNMA2 -based capsid comprises a plurality of recombinant PNMA2 polypeptides of the disclosure. In some instances, the PNMA2-based capsid comprises a plurality of engineered PNMA2 polypeptides of the disclosure. In some embodiments, the PNMA2 polypeptides are recombinant and engineered.
[0308] In some embodiments, the capsid comprises an RTL 10-based capsid. In some instances, the RTLIO-based capsid comprises a plurality of recombinant RTL10 polypeptides of the disclosure. In some instances, the RTLIO-based capsid comprises a plurality of engineered RTL10 polypeptides of the disclosure. In some embodiments, the RTL10 polypeptides are recombinant and engineered.
[0309] In some embodiments, the capsid comprises a PEG10-based capsid. In some instances, the PEG10-based capsid comprises a plurality of recombinant PEG10 polypeptides of the disclosure. In some instances, the PEG10-based capsid comprises a plurality of engineered PEG10 polypeptides of the disclosure. In some embodiments, the PEG10 polypeptides are recombinant and engineered.
[0310] In some embodiments, the capsid comprises an endo-Gag-based capsid. In some instances, the endo-Gag capsid comprises a plurality of recombinant endo-Gag polypeptides of the disclosure. In some instances, the endo-Gag capsid comprises a plurality of engineered endo- Gag polypeptides of the disclosure. In some embodiments, the endo-Gag polypeptides are recombinant and engineered.
[0311] In certain embodiments, the assembly of RTL, PNMA and / or endo-Gag-based capsids occurs ex vivo or in vitro. In some instances, the RTL, PNMA and / or endo-Gag-based capsid is assembled in vivo.
[0312] In some embodiments, the capsid comprises a disulfide bond. The RTL, PNMA or endo-Gag polypeptides that assemble to form the capsids can comprise one or more cysteine residues that form one or more disulfide bonds. Disulfide bonds can contribute to, for example, assembly, re-assembly, and / or stability of capsids disclosed herein. In some embodiments, disulfide bonds can be reduced in a target site, such as an intracellular (e.g., cytoplasmic) environment, facilitating capsid disassembly and cargo delivery.
[0313] In some embodiments, a first RTL, PNMA or endo-Gag polypeptide subunit of a capsid forms an intermolecular disulfide bond with a second RTL, PNMA or endo-Gag polypeptide. In some embodiments, a RTL, PNMA or endo-Gag polypeptide subunit of a capsid forms an intramolecular disulfide bond. In some embodiments, a RTL, PNMA or endo-Gag polypeptide subunit of a capsid forms an intramolecular disulfide bond and an intermolecular disulfide bond. In some embodiments, disulfide bonds contribute to the formation of RTL, PNMA or endo-Gag polypeptide dimers, trimers, tetramers, and / or or other higher-order multimers.
[0314] In some embodiments, a RTL, PNMA polypeptide or endo-Gag polypeptide comprises at least one, at least two, at least three, or at least four cysteines that form disulfide bonds, e.g., upon assembly to form a capsid. In some embodiments, a RTL, PNMA polypeptide or endo-Gag polypeptide comprises one cysteine that forms a disulfide bond. In some embodiments, a RTL, PNMA polypeptide or endo-Gag polypeptide comprises two cysteines that form disulfide bonds. In some embodiments, a RTL, PNMA polypeptide or endo-Gag polypeptide comprises three cysteines that form disulfide bonds. In some embodiments, a RTL, PNMA polypeptide or endo-Gag polypeptide comprises four cysteines that form disulfide bonds.
[0315] In some embodiments, a PNMA polypeptide or endo-Gag polypeptide comprises a cysteine residue, e.g., a disulfide-forming cysteine residue at a position corresponding to, for example, any one or more of positions 10, 136, 233, and 310 of SEQ ID NO: 1. In some embodiments, a PNMA (e.g., PNMA2 or PNMA5) polypeptide or endo-Gag polypeptide comprises a cysteine residue (e g., a disulfide-forming cysteine residue) at a position corresponding to residue 10 of SEQ ID NO: 1. In some embodiments, a PNMA (e.g., PNMA2 or PNMA5) polypeptide or endo-Gag polypeptide comprises a cysteine residue (e.g., a disulfide- forming cysteine residue) at a position corresponding to residue 136 of SEQ ID NO: 1. In some embodiments, a PNMA (e.g., PNMA2 or PNMA5) polypeptide or endo-Gag polypeptide comprises a cysteine residue (e.g., a disulfide-forming cysteine residue) at a position corresponding to residue 233 of SEQ ID NO: 1. In some embodiments, a PNMA (e.g., PNMA2 or PNMA5) polypeptide or endo-Gag polypeptide comprises a cysteine residue (e.g., a disulfide- forming cysteine residue) at a position corresponding to residue 310 of SEQ ID NO: 1.
[0316] In some embodiments, the capsid has an average diameter of at least 1 nm, or more. In some instances, the capsid has an average diameter of at least 2mn, at least 3nm, at least 4nm, at least 5nm, at least lOnm, at least 11 nm, at least 12 nm, at least 13 nm, at least 14 nm, at least 15nm, at least 16 nm, at least 17 nm, at least 18 nm, at least 19 nm, at least 20nm, at least 25nm, at least 30nm, at least 40nm, at least 50nm, at least 60nm, at least 70nm, at least 80nm, at least 90nm, at least lOOnm, at least 150nm, at least 200nm, at least 300nm, at least 400nm, at least 500nm, at least 600nm, or more. In some instances, the capsid has an average diameter of at least 5nm, or more. In some instances, the capsid has an average diameter of at least 7nm, or more. In some cases, the capsid has an average diameter of at least lOnm, or more. In some instances, the capsid has an average diameter of at least 12nm, or more. In some instances, the capsid has an average diameter of at least 15nm, or more. In some instances, the capsid has an average diameter of at least 17nm, or more. In some instances, the capsid has an average diameter of at least 20nm, or more. In some cases, the capsid has an average diameter of at least 30nm, or more. In some cases, the capsid has an average diameter of at least 40nm, or more. In some cases, the capsid has an average diameter of at least 50nm, or more. In some cases, the capsid has an average diameter of at least 80nm, or more. In some cases, the capsid has an average diameter of at least lOOnm, or more. In some cases, the capsid has an average diameter of at least 200nm, or more. In some cases, the capsid has an average diameter of at least 300nm, or more. In somecases, the capsid has an average diameter of at least 400nm, or more. In some cases, the capsid has an average diameter of at least 500nm, or more. In some cases, the capsid has an average diameter of at least 600nm, or more.
[0317] In some embodiments, the capsid has an average diameter of at most 100 nm, or less. In some instances, the capsid has an average diameter of at most Inm, at most 5nm, at most lOnm, at most 20nm, at most 30nm, at most 40nm, at most 50nm, at most 60nm, at most 70nm, at most 80nm, at most 90nm, at most 95nm, at most lOOnm, at most 105nm, at most HOnm, at most 115nm, at most 120nm, at most 125nm, at most 130nm, at most 140nm, at most 150nm, at most 200nm, at most 300nm, at most 400nm, at most 500nm, at most 600nm, or less. In some cases, the capsid has an average diameter of at most 50nm, or less. In some cases, the capsid has an average diameter of at most 80nm, or less.
[0318] In some cases, the capsid has an average diameter of at most 90nm, or less. In some cases, the capsid has an average diameter of at most 95nm, or less. In some cases, the capsid has an average diameter of at most lOOnm, or less. In some cases, the capsid has an average diameter of at most 1 lOnm, or less. In some cases, the capsid has an average diameter of at most HOnm, or less. In some cases, the capsid has an average diameter of at most 150nm, or less. In some cases, the capsid has an average diameter of at most 200nm, or less. In some cases, the capsid has an average diameter of at most 300nm, or less. In some cases, the capsid has an average diameter of at most 400nm, or less. In some cases, the capsid has an average diameter of at most 500nm, or less. In some cases, the capsid has an average diameter of at most 600nm, or less.
[0319] In some embodiments, the capsid has an average diameter of about Inm, about 5nm, about lOnm, about 15nm, about 20nm, about 25nm, about 30nm, about 35nm, about 40nm, about 45nm, about 50nm, about 55nm, about 60nm, about 65nm, about 70nm, about 75nm, about 80nm, about 85nm, about 90nm, about 95nm, about lOOnm, about 105nm, about HOnm, about 120nm, about 150nm, about 200nm, about 300nm, about 400nm, about 500nm, or about 600nm. In some instances, the capsid has an average diameter of about 20nm. In some cases, the capsid has an average diameter of about 30nm. In some cases, the capsid has an average diameter of about 40nm. In some cases, the capsid has an average diameter of about 50nm. In some cases, the capsid has an average diameter of about 60nm. In some cases, the capsid has an average diameter of about 70nm. In some cases, the capsid has an average diameter of about 80nm. Insome cases, the capsid has an average diameter of about lOOnm. In some cases, the capsid has an average diameter of about 200nm.
[0320] In some embodiments, the capsid has an average diameter of from about Inm to about 600nm. In some instances, the capsid has an average diameter of from about 5nm to about 500nm, from about 5nm to about 400nm, from about 5nm to about 300nm, from about 5nm to about 200nm, from about 5nm to about lOOnm, from about 5nm to about 50nm, from about 5nm to about 30nm, lOnm to about 500nm, from about lOnm to about 400nm, from about lOnm to about 300nm, from about lOnm to about 200nm, from about lOnm to about 150nm, from about lOnm to about 120nm, from about lOnm to about lOOnm, from about lOnm to about 90nm, from about lOnm to about 50nm, from about lOnm to about 30nm, 15nm to about 500nm, from about 15nm to about 400nm, from about 15nm to about 300nm, from about 15nm to about 200nm, from about 15nm to about 150nm, from about 15nm to about 120nm, from about 15nm to about lOOnm, from about 15nm to about 90nm, from about 15nm to about 50nm, from about 20nm to about 500nm, from about 20nm to about 400nm, from about 20nm to about 300nm, from about 20nm to about 200nm, from about 20nm to about 150nm, from about 20nm to about 120nm, from about 20nm to about lOOnm, from about 20nm to about 90nm, from about 20nm to about 50nm, from about 30nm to about 500nm, from about 30nm to about 400nm, from about 30nm to about 300nm, from about 30nm to about 200nm, from about 30nm to about lOOnm, from about 30nm to about 50nm, from about 50nm to about 300nm, from about 50nm to about 200nm, or from about 50nm to about lOOnm.
[0321] Capsids in a composition of the disclosure can be, for example, at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least 92%, at least 94%, at least 96%, at least 98%, or at least 99% pure. Purity can be determined, for example, by SDS-PAGE.
[0322] Capsids in a composition of the disclosure can exhibit, for example, at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% particle homogeneity. Particle homogeneity can be determined, for example, by multi-angle dynamic light scattering (MADLS), e.g., using a particle diameter size range disclosed herein.
[0323] In some instances, the RTL (e g., PEG10 or RTL10), PNMA and / or endo-Gag-based capsid is stable at room temperature. In some cases, the RTL (e.g., PEG10 or RTL10), PNMA and / or endo-Gag-based capsid is empty. In other cases, the RTL (e.g., PEG10 or RTL10), PNMA and / or endo-Gag-based capsid is loaded (for example, loaded with a heterologous cargo disclosed herein). In some embodiments, the heterologous cargo is in an interior of at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of a plurality of capsids. In some embodiments, the heterologous cargo is in an interior of at least 50% of a plurality of the capsids. In some embodiments, the heterologous cargo is on an exterior of at least 1%, at least 3%, at least 5%, at least 7%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of a plurality of capsids.
[0324] In some embodiments, at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the heterologous cargo is in an interior of a plurality of capsids. In some embodiments, at most 1%, at most 3%, at most 5%, at most 7%, at most 10%, at most 20%, at most 30%, at most 40%, at most 50%, at most 60%, at most 70%, at most 80%, at most 90%, at most 91%, at most 92%, at most 93%, at most 94%, at most 95%, at most 96%, at most 97%, at most 98%, at most 99%, or at most 99.5% of the heterologous cargo is on an exterior of a plurality of capsids. In some embodiments, at least 1%, at least 3%, at least 5%, at least 7%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the heterologous cargo is on an exterior of a plurality of capsids.
[0325] In some instances, the RTL (e.g., PEG10 or RTL10), PNMA and / or endo-Gag-based capsid is stable at a temperature from about 2°C to about 37°C. In some instances, the RTL (e.g., PEG10 or RTL10), PNMA and / or endo-Gag-based capsid is stable at a temperature from about 2°C to about 40°C, 2°C to about 37°C, 2°C to about 25°C, 2°C to about 20°C, 4°C to about40°C, 4°C to about 37°C, 4°C to about 25°C, 4°C to about 20°C, 2°C to about 8°C, about 2°C to about 4°C, about 20°C to about 37°C, about 25°C to about 37°C, about 20°C to about 30°C, about 25°C to about 30°C, or about 30°C to about 37°C. In some cases, the RTL (e.g., PEG10 or RTL10), PNMA and / or endo-Gag-based capsid is empty. In other cases, the RTL (e.g., PEG10 or RTL 10), PNMA and / or endo-Gag-based capsid is loaded (for example, loaded with a heterologous cargo and / or a therapeutic agent, e.g., a DNA or an RNA).
[0326] In some instances, the RTL (e.g., PEG10 or RTL 10), PNMA and / or endo-Gag-based capsid is stable for at least about 1 day, at least about 2 days, at least about 4 days, at least about 5 days, at least about 7 days, at least about 14 days, at least about 28 days, at least about 30 days, at least about 60 days, at least about 2 months, at least about 3 months, at least about 4 months, at least about 5 months, at least about 6 months, at least about 12 months, at least about 18 months, at least about 24 months, at least about 3 years, at least about 5 years, or longer. In some case, the RTL (e.g., PEG10 or RTL10), PNMA and / or endo-Gag-based capsid exhibits low or minimal degradation, e.g., less than 10%, less than 5%, less than 4%, less than 3%, less than 2%, less than 1%, or less than 0.5% based on the total population of the RTL (e.g., PEG10 or RTL10), PNMA and / or endo-Gag-based capsids that are degraded. In some cases, the RTL (e.g., PEG10 or RTL10), PNMA and / or endo-Gag-based capsid is empty. In other cases, the RTL (e.g., PEG10 or RTL 10), PNMA and / or endo-Gag-based capsid is loaded (for example, loaded with a therapeutic agent, e.g., a DNA or an RNA).
[0327] In some embodiments, a capsid comprises a first endogenous retroviral capsid polypeptide and a second endogenous retroviral capsid polypeptide; wherein the amino acid sequence of the first endogenous retroviral capsid polypeptide is not identical to the amino acid sequence of the second endogenous retroviral capsid polypeptide.
[0328] In some embodiments, the first endogenous retroviral capsid polypeptide is an endo Gag polypeptide or comprises an amino acid sequence of an endo Gag polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is a native endo Gag polypeptide or comprises an amino acid sequence of a native endo Gag polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is an engineered endo Gag polypeptide or comprises an amino acid sequence of an engineered endo Gag polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is an engineered endo Gag polypeptide that comprises an amino acid deletion relative to a corresponding native endo Gagpolypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is an engineered endo Gag polypeptide that comprises an amino acid substitution relative to a corresponding native endo Gag polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is an engineered endo Gag polypeptide that comprises an amino acid insertion relative to a corresponding native endo Gag polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is an engineered endo Gag polypeptide that comprises a non-native cysteine that is not present in a corresponding native endo Gag polypeptide.
[0329] In some embodiments, the first endogenous retroviral capsid polypeptide is a RTL or PNMA polypeptide (e.g., PEG10, RTL 10, PNMA2 or PNMA5) or comprises an amino acid sequence of a RTL or PNMA polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is a native RTL or PNMA polypeptide or comprises an amino acid sequence of a native RTL or PNMA polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is an engineered RTL or PNMA polypeptide or comprises an amino acid sequence of an engineered RTL or PNMA polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is an engineered RTL or PNMA polypeptide that comprises an amino acid deletion relative to a corresponding native RTL or PNMA polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is an engineered RTL or PNMA polypeptide that comprises an amino acid substitution relative to a corresponding native RTL or PNMA polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is an engineered RTL or PNMA polypeptide that comprises an amino acid insertion relative to a corresponding native RTL or PNMA polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is an engineered RTL or PNMA polypeptide that comprises a non-native cysteine that is not present in a corresponding native RTL or PNMA polypeptide.
[0330] In some embodiments, the first endogenous retroviral capsid polypeptide is not a RTL or PNMA polypeptide or does not contain an amino acid sequence of a RTL or PNMA polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is not a native RTL or PNMA polypeptide or does not contain an amino acid sequence of a native RTL or PNMA polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide-n-is not an engineered RTL or PNMA polypeptide or does not contain an amino acid sequence of an engineered RTL or PNMA polypeptide.
[0331] In some embodiments, the first endogenous retroviral capsid polypeptide is not a PNMA2 polypeptide or does not contain an amino acid sequence of a PNMA2 polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is not a native PNMA2 polypeptide or does not contain an amino acid sequence of a native PNMA2 polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is not an engineered PNMA2 polypeptide or does not contain an amino acid sequence of an engineered PNMA2 polypeptide.
[0332] In some embodiments, the first endogenous retroviral capsid polypeptide is not a PNMA5 polypeptide or does not contain an amino acid sequence of a PNMA5polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is not a native PNMA5 polypeptide or does not contain an amino acid sequence of a native PNMA5 polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is not an engineered PNMA5 polypeptide or does not contain an amino acid sequence of an engineered PNMA5 polypeptide.
[0333] In some embodiments, the first endogenous retroviral capsid polypeptide is not an RTL 10 polypeptide or does not contain an amino acid sequence of an RTL 10 polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is not a native RTL 10 polypeptide or does not contain an amino acid sequence of a native RTL 10 polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is not an engineered RTL10 polypeptide or does not contain an amino acid sequence of an engineered RTL 10 polypeptide.
[0334] In some embodiments, the first endogenous retroviral capsid polypeptide is not a PEG10 polypeptide or does not contain an amino acid sequence of a PEG10 polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is not a native PEG10 polypeptide or does not contain an amino acid sequence of a native PEG10 polypeptide. In some embodiments, the first endogenous retroviral capsid polypeptide is not an engineered PEG10 polypeptide or does not contain an amino acid sequence of an engineered PEG10 polypeptide.
[0335] In some embodiments, the first endogenous retroviral capsid polypeptide comprises an exogenous polypeptide sequence disclosed herein. In some embodiments, the first endogenous retroviral capsid polypeptide comprises a cargo binding domain disclosed herein(e.g., an RNA, DNA, or protein-binding domain). In some embodiments, the first endogenous retroviral capsid polypeptide comprises a domain that binds to a cell surface molecule. In some embodiments, the first endogenous retroviral capsid polypeptide comprises an antibody or antigen-binding fragment thereof
[0336] In some embodiments, the second endogenous retroviral capsid polypeptide is an endo Gag polypeptide or comprises an amino acid sequence of an endo Gag polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is a native endo Gag polypeptide or comprises an amino acid sequence of a native endo Gag polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is an engineered endo Gag polypeptide or comprises an amino acid sequence of an engineered endo Gag polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is an engineered endo Gag polypeptide that comprises an amino acid deletion relative to a corresponding native endo Gag polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is an engineered endo Gag polypeptide that comprises an amino acid substitution relative to a corresponding native endo Gag polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is an engineered endo Gag polypeptide that comprises an amino acid insertion relative to a corresponding native endo Gag polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is an engineered endo Gag polypeptide that comprises a non-native cysteine that is not present in a corresponding native endo Gag polypeptide.
[0337] In some embodiments, the second endogenous retroviral capsid polypeptide is a PNMA (e.g., PNMA2 or PNMA5) polypeptide or comprises an amino acid sequence of a PNMA (e g., PNMA2 or PNMA5) polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is a native PNMA (e.g., PNMA2 or PNMA5) polypeptide or comprises an amino acid sequence of a native PNMA polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is an engineered PNMA polypeptide or comprises an amino acid sequence of an engineered PNMA polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is an engineered PNMA polypeptide that comprises an amino acid deletion relative to a corresponding native PNMA polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is an engineered PNMA polypeptide that comprises an amino acid substitution relative to a corresponding native PNMApolypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is an engineered PNMA polypeptide that comprises an amino acid insertion relative to a corresponding native PNMA2 polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is an engineered PNMA polypeptide that comprises a non-native cysteine that is not present in a corresponding native PNMA polypeptide.
[0338] In some embodiments, the second endogenous retroviral capsid polypeptide is an RTL (e.g., RTL10 or PEG10) polypeptide or comprises an amino acid sequence of an RTL polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is a native RTL (e g., RTL 10 or PEG10) polypeptide or comprises an amino acid sequence of a native RTL polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is an engineered RTL (e.g., RTL10 or PEG10) polypeptide or comprises an amino acid sequence of an engineered RTL polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is an engineered RTL polypeptide that comprises an amino acid deletion relative to a corresponding native RTL polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is an engineered RTL polypeptide that comprises an amino acid substitution relative to a corresponding native RTL polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is an engineered RTL polypeptide that comprises an amino acid insertion relative to a corresponding native RTL polypeptide. In some embodiments, the second endogenous retroviral capsid polypeptide is an engineered RTL polypeptide that comprises a non-native cysteine that is not present in a corresponding native RTL polypeptide.
[0339] In some embodiments, the second endogenous retroviral capsid polypeptide comprises an exogenous polypeptide sequence disclosed herein. In some embodiments, the second endogenous retroviral capsid polypeptide comprises a cargo binding domain disclosed herein (e.g., an RNA, DNA, or protein-binding domain). In some embodiments, the second endogenous retroviral capsid polypeptide comprises a domain that binds to a cell surface molecule. In some embodiments, the second endogenous retroviral capsid polypeptide comprises an antibody or antigen-binding fragment thereof.
[0340] In some embodiments, the first endogenous retroviral capsid polypeptide and the second endogenous retroviral capsid polypeptide are present in the capsid at a ratio of about 1 : 1, 2: 1, 3: 1 , 4: 1, 5:1 , 6: 1, 7:1, 8: 1, 9: 1, 10: 1, 20:1 , 50: 1 , 100:1 , 1 :2, 1 :3, 1 :4, 1 :5, 1 :6, 1 :7, 1 :8, 1 :9,1 : 10, 1 :20, or 1 :50. In some instances, the ratio is the comparison in molar concentration. In some instances, the ratio is the comparison in the number of capsid forming subunits. In some instances, the ratio based on the mass (e.g., number of ng) of the capsid forming subunits present.
[0341] In some embodiments, the PNMA-based capsid or endo-Gag-based capsid comprises a plurality of recombinant or engineered PNMA polypeptides and a plurality of non-PNMA proteins. In some instances, the ratio of the plurality of recombinant or engineered PNMA polypeptides to the plurality of non-PNMA proteins is 1 : 1, 2: 1, 3: 1, 4: 1, 5: 1, 6:1, 7:1, 8: 1, 9: 1, 10: 1, 20: 1, 50: 1, 100: 1, 1 :2, 1 :3, 1 :4, 1 :5, 1 :6, 1 :7, 1 :8, 1 :9, 1 : 10, 1 :20, or 1 :50. In some instances, the ratio is the comparison in molar concentration. In some instances, the ratio is the comparison in the number of capsid forming subunits.
[0342] In some embodiments, the PNMA2 -based capsid or endo-Gag-based capsid comprises a plurality of recombinant or engineered PNMA2 polypeptides and a plurality of non- PNMA2 proteins. Exemplary species of non-PNMA2 proteins include but are not limited to, PNMA1, PNMA3, PNMA4 (MOAP1), PNMA5, PNMA6A, PNMA6B / 6D, PNMA6E, PNMA6F, PNMA7A, PNMA7B, PNMA8A, PNMA8B, PNMA8C, CCDC8, Arc, RTL10 (BOP), LDOC1, PEG10, RTL3, RTL6, RTL8A, RTL8B, ZNF18, Copia, ASPRV1, a protein or a combination of proteins chosen from the SCAN domain family, and a protein or a combination of proteins chosen from the retrotransposon Gag-like family.
[0343] In some instances, the ratio of the plurality of recombinant or engineered PNMA2 polypeptides to the plurality of non-PNMA2 proteins is 1 : 1, 2:1, 3: 1, 4:1, 5: 1, 6: 1, 7: 1, 8: 1, 9:1, 10: 1, 20: 1, 50: 1, 100: 1, 1 :2, 1 :3, 1 :4, 1:5, 1 :6, 1 :7, 1 :8, 1 :9, 1 : 10, 1 :20, or 1:50. In some instances, the ratio is the comparison in molar concentration. In some instances, the ratio is the comparison in the number of capsid forming subunits.
[0344] In some embodiments, the PNMA5-based capsid or endo-Gag-based capsid comprises a plurality of recombinant or engineered PNMA5 polypeptides and a plurality of non- PNMA5 proteins. Exemplary species of non-PNMA5 proteins include but are not limited to, PNMA1, PNMA2, PNMA3, PNMA4 (MOAP1), PNMA6A, PNMA6B / 6D, PNMA6E, PNMA6F, PNMA7A, PNMA7B, PNMA8A, PNMA8B, PNMA8C, CCDC8, Arc, RTL10 (BOP), LDOC1, PEG10, RTL3, RTL6, RTL8A, RTL8B, ZNF18, Copia, ASPRV1, a protein or acombination of proteins chosen from the SCAN domain family, and a protein or a combination of proteins chosen from the retrotransposon Gag-like family.
[0345] In some instances, the ratio of the plurality of recombinant or engineered PNMA5 polypeptides to the plurality of non-PNMA5 proteins is 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 20:1, 50:1, 100:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:20, or 1:50. In some instances, the ratio is the comparison in molar concentration. In some instances, the ratio is the comparison in the number of capsid forming subunits.
[0346] In some embodiments, the RTL 10-based capsid or endo-Gag-based capsid comprises a plurality of recombinant or engineered RTL 10 polypeptides and a plurality of non-RTLIO proteins. Exemplary species of non-RTLIO proteins include but are not limited to, PNMA1, PNMA2, PNMA3, PNMA4 (MOAP1), PNMA5, PNMA6A, PNMA6B / 6D, PNMA6E, PNMA6F, PNMA7A, PNMA7B, PNMA8A, PNMA8B, PNMA8C, CCDC8, Arc, LDOC1, PEG10, RTL3, RTL6, RTL8A, RTL8B, ZNF18, Copia, ASPRV1, a protein or a combination of proteins chosen from the SCAN domain family, and a protein or a combination of proteins chosen from the retrotransposon Gag-like family.
[0347] In some instances, the ratio of the plurality of recombinant or engineered RTL 10 polypeptides to the plurality of non-RTLIO proteins is 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 20:1, 50:1, 100:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:20, or 1:50. In some instances, the ratio is the comparison in molar concentration. In some instances, the ratio is the comparison in the number of capsid forming subunits.
[0348] In some embodiments, the PEGlO-based capsid or endo-Gag-based capsid comprises a plurality of recombinant or engineered PEG10 polypeptides and a plurality of non-PEGlO proteins. Exemplary species of non-PEGlO proteins include but are not limited to, PNMA1, PNMA2, PNMA3, PNMA4 (MOAP1), PNMA5, PNMA6A, PNMA6B / 6D, PNMA6E, PNMA6F, PNMA7A, PNMA7B, PNMA8A, PNMA8B, PNMA8C, CCDC8, Arc, RTL 10 (BOP), LDOC1, RTL3, RTL6, RTL8A, RTL8B, ZNF18, Copia, ASPRV1, a protein or a combination of proteins chosen from the SCAN domain family, and a protein or a combination of proteins chosen from the retrotransposon Gag-like family.
[0349] In some instances, the ratio of the plurality of recombinant or engineered PEG10 polypeptides to the plurality of non-PEGlO proteins is 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 20:1, 50:1, 100:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:20, or 1:50. In someinstances, the ratio is the comparison in molar concentration. In some instances, the ratio is the comparison in the number of capsid forming subunits.A. Delivery component
[0350] In some embodiments, a delivery component is combined with an RTL (e.g., PEG10 or RTL 10) based or PNMA-based capsid or endo-Gag-based capsid for a targeted delivery to a site of interest. In some instances, the delivery component comprises a carrier, e.g., an extracellular vesicle such as a micelle, a liposome, or a microvesicle; or a viral envelope.
[0351] In some instances, the delivery component serves as a primary delivery vehicle for a RTL-based or PNMA-based capsid or endo-Gag-based capsid which does not comprise its own delivery component. In such cases, the delivery component directs the RTL-based or PNMA- based capsid or endo-Gag-based capsid to a target site of interest and optionally facilitates intracellular uptake.
[0352] In some embodiments, the delivery component enhances target specificity and / or sensitivity of a RTL-based or PNMA-based capsid. In such cases, the delivery component enhances the specificity and / or affinity of the RTL-based or PNMA-based capsid or endo-Gag- based capsid to the target site. In some embodiments, the delivery components enhance the specificity and / or affinity by at least 2-fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 6-fold, at least 7-fold, at least 8-fold, at least 9-fold, at least 10-fold, at least 20-fold, at least 30- fold, at least 50-fold, at least 100-fold, at least 200-fold, at least 500-fold, or more. In additional cases, the delivery components enhance the specificity and / or affinity by about 2-fold, 3-fold, 4- fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 20-fold, 30-fold, 50-fold, 100-fold, 200-fold, 500-fold, or more. In further cases, the delivery components enhance the specificity and / or affinity by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 200%, at least 500%, or more.
[0353] In some embodiments, the capsid has a low or reduced off-target effect. An off-target effect can be, for example, an effect that occurs when a cargo is delivered to an unintended tissue or cell type, or an effect that occurs when a cargo binds to an unintended interaction partner. In some cases, the off-target effect is less than 10%, less than 5%, less than 4%, less than 3%, less than 2%, less than 1%, or less than 0.5%. In some cases, the capsid does not have an off-target effect.
[0354] In some cases, the delivery component s) enhance the specificity and / or affinity by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 200%, at least 500%, or more. In further cases, the delivery components enhance the specificity and / or affinity by about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 500%, or more. Further still, the delivery component optionally reduces an off-target effect by at least 2-fold, at least 3-fold, at least 4-fold, at least 5- fold, at least 6-fold, at least 7-fold, at least 8-fold, at least 9-fold, at least 10-fold, at least 20-fold, at least 30-fold, at least 50-fold, at least 100-fold, at least 200-fold, at least 500-fold, or more. Further still, the delivery component optionally reduces off-target effect by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 200%, at least 500%, or more.
[0355] In additional instances, the delivery component serves as a first vehicle that transports an RTL-based or PNMA-based capsid to a general target region (e.g., a tumor microenvironment) and the RTL-based or PNMA-based or endo-Gag-based capsid comprises a second delivery component (for example, an exogenous polypeptide sequence as disclosed herein) that directs the RTL-based or PNMA-based capsid or endo-Gag-based capsid to a more specific target site and optionally facilitates intracellular uptake. In some embodiments, the delivery component (e.g., delivery component that serves as the first vehicle) minimizes off- target effect by at least 2-fold, at least 3 -fold, at least 4-fold, at least 5-fold, at least 6-fold, at least 7-fold, at least 8-fold, at least 9-fold, at least 10-fold, at least 20-fold, at least 30-fold, at least 50-fold, at least 100-fold, at least 200-fold, at least 500-fold, or more. In some embodiments, the delivery component minimizes off-target effect by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 200%, at least 500%, or more.
[0356] In further instances, the delivery component serves as a first vehicle that transports an RTL-based or PNMA-based capsid to a target site of interest, and the RTL-based or PNMA- based or endo-Gag-based capsid comprises a second delivery component that facilitates intracellular uptake, for example, an exogenous polypeptide sequence as disclosed herein.
[0357] In some embodiments, the delivery component comprises an extracellular vesicle. In some instances, the extracellular vesicle comprises a microvesicle, a liposome, or a micelle. In some instances, the extracellular vesicle has an average diameter of from about 10 nm to about2000 nm, from about 10 nm to about 1000 nm, from about lOnm to about 800nm, from about 20nm to about 600nm, from about 30nm to about 500nm, from about 50nm to about 200nm, or from about 80nm to about lOOnm.
[0358] In some embodiments, the delivery component comprises a microvesicle. Also known as circulating microvesicles, microvesicles are membrane-bound vesicles that comprise phospholipids. In some instances, the microvesicle has an average diameter of from about 50nm to about lOOOnm, from about lOOnm to about 800nm, from about 200nm to about 500nm, or from about 50nm to about 400nm.
[0359] In some instances, the microvesicle is originated from cell membrane inversion, exocytosis, shedding, blebbing, or budding. In some instances, the microvesicles are generated from differentiated cells. In other instances, the microvesicles are generated from undifferentiated cells, e.g., by blast cells, progenitor cells, or stem cells.
[0360] In some embodiments, the delivery component comprises a microparticle. In some embodiments, the delivery component comprises a nanoparticle. In some embodiments, the delivery component comprises a lipid nanoparticle.
[0361] In some embodiments, the delivery component comprises a liposome. In some instances, the liposome comprises a plurality of lipopeptides, which are presented on the surface of the liposome, for targeted delivery to a site or region of interest. In some cases, the liposomes fuse with the target cell, whereby the contents of the liposome are then emptied into the target cell. In some cases, a liposome is endocytosed by cells that are phagocytic. Endocytosis is then followed by intralysosomal degradation of liposomal lipids and release of the encapsulated agents.
[0362] Illustrative liposomes suitable for incorporation include, and are not limited to, multilamellar vesicles (MLV), oligolamellar vesicles (OLV), unilamellar vesicles (UV), small unilamellar vesicles (SUV), medium-sized unilamellar vesicles (MUV), large unilamellar vesicles (LUV), giant unilamellar vesicles (GUV), multivesicular vesicles (MW), single or oligolamellar vesicles made by reverse-phase evaporation method (REV), multilamellar vesicles made by the reverse-phase evaporation method (MLV-REV), stable plurilamellar vesicles (SPLV), frozen and thawed MLV (FATMLV), vesicles prepared by extrusion methods (VET), vesicles prepared by French press (FPV), vesicles prepared by fusion (FUV), dehydration- rehydration vesicles (DRV), and bubblesomes (BSV). In some instances, a liposome comprisesAmphipol (A8-35). Techniques for preparing liposomes are described in, for example, COLLOIDAL DRUG DELIVERY SYSTEMS, vol. 66 (I. Kreuter ed., Marcel Dekker, Inc. (1994)), which is incorporated herein by reference for such disclosure.
[0363] Depending on the method of preparation, liposomes are unilamellar or multilamellar, and vary in size with diameters ranging from about 20nm to greater than about lOOOnm.
[0364] In some instances, liposomes provided herein also comprise carrier lipids. In some embodiments the carrier lipids are phospholipids. Carrier lipids capable of forming liposomes include, but are not limited to, dipalmitoylphosphatidylcholine (DPPC), phosphatidylcholine (PC; lecithin), phosphatidic acid (PA), phosphatidylglycerol (PG), phosphatidylethanolamine (PE), or phosphatidylserine (PS). Other suitable phospholipids further include distearoylphosphatidylcholine (DSPC), dimyristoylphosphatidylcholine (DMPC), dipalmitoylphosphatidyglycerol (DPPG), distearoylphosphatidyglycerol (DSPG), dimyristoylphosphatidylglycerol (DMPG), dipalmitoylphosphatidic acid (DPP A); dimyristoylphosphatidic acid (DMPA), distearoylphosphatidic acid (DSPA), dipalmitoylphosphatidylserine (DPPS), dimyristoylphosphatidylserine (DMPS), distearoylphosphatidylserine (DSPS), dipalmitoylphosphatidyethanolamine (DPPE), dimyristoylphosphatidylethanolamine (DMPE), distearoylphosphatidylethanolamine (DSPE) and the like, or combinations thereof. In some embodiments, the liposomes further comprise a sterol (e g., cholesterol) which modulates liposome formation. The carrier lipids are optionally any non-phosphate polar lipids. In some embodiments, a liposome comprises an electroneutral lipid.
[0365] In some embodiments, a liposome or a delivery component comprises a cationic lipid. Cationic lipids have a head group with permanent positive charges. Non-limiting examples of cationic lipids include l,2-di-O-octadecenyl-3-trimethylammonium-propane (DOTMA), 1,2- dioleoyl-3-trimethylammonium-propane (DOTAP), Dimethyl dioctadecylammonium bromide (DDAB), and 2,3-dioleyloxy-N-[2-(sperminecarboxamido)ethyl]-N,N-dimethyl-l- propanaminium trifluoroacetate (DOSPA), and commercially available transfection reagents.
[0366] In some embodiments, the delivery component comprises a micelle. In some instances, the micelle has an average diameter from about 2nm to about 250nm, from about 20nm to about 200nm, from about 20nm to about lOOnm, or from about 50 to about lOOnm.
[0367] In some instances, the micelle is a polymeric micelle, characterized by a core shell structure, in which the hydrophobic core is surrounded by a hydrophilic shell. In some cases, thehydrophilic shell further comprises a hydrophilic polymer or copolymer and a pH sensitive component.
[0368] Exemplary hydrophilic polymers or copolymers include, but are not limited to, poly(N-substituted acrylamides), poly(N-acryloyl pyrrolidine), poly(N-acryloyl piperidine), poly(N-acryl-L-amino acid amides), poly(ethyl oxazoline), methylcellulose, hydroxypropyl acrylate, hydroxyalkyl cellulose derivatives and poly(vinyl alcohol), poly(N- isopropylacrylamide), poly(N-vinyl-2-pyrrolidone), polyethyleneglycol derivatives, and combinations thereof.
[0369] The delivery component may comprise a pH-sensitive moiety, which can include, but is not limited to, an alkylacrylic acid such as methacrylic acid, ethylacrylic acid, propyl acrylic acid and butyl acrylic acid, or an amino acid such as glutamic acid.
[0370] In some instances, a hydrophobic moiety constitutes the core of the micelle and includes, for example, a single alkyl chain, such as octadecyl acrylate or a double chain alkyl compound such as phosphatidylethanolamine or dioctadecylamine. In some cases, the hydrophobic moiety is optionally a water insoluble polymer such as a poly(lactic acid) or a poly(e-caprolactone).
[0371] Polymeric micelles exhibiting pH-sensitive properties are also contemplated and are formed, e.g., by using pH-sensitive polymers including, but not limited to, copolymers from methacrylic acid, methacrylic acid esters and acrylic acid esters, polyvinyl acetate phthalate, hydroxypropyl methyl cellulose phthalate, cellulose acetate phthalate, or cellulose acetate trimellitate.
[0372] A Delivery component can comprise a cationic moiety, for example, a cationic lipid, a cationic peptide, or a cationic polymer. A cationic moiety (e.g., lipid, peptide, or moiety) can associate with a capsid (e.g., a negative spike portion thereof) via electrostatic interactions. The presence of a cationic lipid, polymer, or peptide on a capsid can enhance binding to the negatively charged surface of a cell, facilitating increased capsid uptake.
[0373] In some embodiments, the delivery component comprises a cationic peptide, e.g., a cationic cell-penetrating peptide disclosed herein.
[0374] In some embodiments, the delivery component comprises a cationic polymer. Nonlimiting examples of cationic polymers include cationic peptides and their derivatives (e.g., polylysine, polyornithine), linear or branched synthetic polymers (e.g., polybrene,polyethyleneimine), polysaccharide-based delivery molecules (e.g., cyclodextrin, chitosan), natural polymers (e.g., histone, collagen), and activated and non-activated dendrimers. Cationic reagents can adhere to the cell membrane through electrostatic interactions and promote cellular uptake via endocytosis.
[0375] In some embodiments, the delivery component comprises a viral envelope or a component thereof. For example, viral envelopes comprise glycoproteins, phospholipids, and additional proteins obtained from a host, any of which a delivery component can comprise. In some instances, the viral envelope or component thereof is permissive to a wide range of target cells. In other instances, the viral envelope or component thereof is non -permissive and is specific to a target cell of interest. In some cases, the viral envelope or component thereof comprises a cell-specific binding protein and optionally a fusogenic molecule that aids in the fusion of the cargo into a target cell. In some cases, the viral envelope or component thereof comprises an endogenous viral envelope. In other cases, the viral envelope is a modified envelope, comprising one or more foreign proteins.
[0376] In some instances, the viral envelope or component thereof is derived from a DNA virus. Illustrative enveloped DNA viruses include viruses from the family of Herpesviridae, Poxviridae, and Hepadnavirdae . In other instances, the viral envelope or component thereof is derived from an RNA virus. Illustrative enveloped RNA viruses include viruses from the family of Bunyaviridae, Coronaviridae, Fdoviridae, Flaviviridae, Orthomyxoviridae, Paramyxoviridae, Rhabdoviridae, and Togaviridae . In additional instances, the viral envelope or component thereof is derived from a virus from the family of Retroviridae .
[0377] In some embodiments, the viral envelope or component thereof is from an oncolytic virus, such as an oncolytic DNA virus from the family of Herpesviridae (for example, HSV1) or Poxviridae (for example, Vaccinia virus and myxoma virus); or an oncolytic RNA virus from the family of Rhabdoviridae (for example, VSV) ox Paramyxoviridae (for example MV and NDV).
[0378] A component from a virus can be used for pseudotyping a capsid disclosed herein. Env is a retroviral gene that encodes the protein that forms the viral envelope. The expression of the env gene allows retroviruses to target and attach to specific cell types, and to infiltrate the target cell membrane. In some embodiments, a delivery component comprises a protein encoded by a retroviral Env gene.
[0379] In some instances, the delivery component comprises a fusogenic protein, for example, a FAST protein or an engineered variant thereof.
[0380] In some instances, the delivery component, viral envelope, or component thereof further comprises a foreign or engineered protein that binds to an antigen or a cell surface molecule. Illustrative antigens and cell surface molecules for targeting include, but are not limited to, P-glycoprotein, Her2 / Neu, erythropoietin (EPO), epidermal growth factor receptor (EGFR), vascular endothelial growth factor receptor (VEGF-R), cadherin, carcinoembryonic antigen (CEA), CD4. CD8, CD19. CD20, CD33, CD34, CD45, CD117 (c-kit), CD133, HLA-A. HLA-B, HLA-C, chemokine receptor 5 (CCRS), stem cell marker ABCG2 transporter, ovarian cancer antigen CA125, immunoglobulins, integrins, prostate specific antigen (PSA), prostate stem cell antigen (PSCA), dendritic cell-specific intercellular adhesion molecule 3-grabbing nonintegrin (DC-SIGN), thyroglobulin, granulocyte-macrophage colony stimulating factor (GM- CSF), myogenic differentiation promoting factor-1 (MyoD-1), Leu-7 (CD57), LeuM-1, cell proliferation-associated human nuclear antigen defined by the monoclonal antibody Ki-67 (Ki- 67), viral envelope proteins, HIV gpl20, or transferrin receptor.IV. CARGOS
[0381] In some embodiments, a composition or capsid disclosed herein comprises a cargo. In some embodiments, the cargo is a heterologous cargo that is not native to the endo Gag or PNMA polypeptide (e.g., is not native to the PNMA2 or PNMA5 polypeptide). In some embodiments, the cargo is a therapeutic agent. In some embodiments, the cargo is a nucleic acid molecule, a small molecule, a protein, a peptide, an antibody or binding fragment thereof, a peptidomimetic, or a nucleotidomimetic. In some instances, the cargo is a therapeutic cargo, comprising e g., one or more drugs. In some instances, the cargo comprises a diagnostic tool or component thereof, for profiling, e.g., one or more markers (such as markers associates with one or more disease phenotypes). In additional instances, the cargo comprises an imaging tool or component thereof.
[0382] In some embodiments, the cargo is a peptidomimetic. A peptidomimetic is a small protein-like polymer designed to mimic a peptide. In some instances, the peptidomimetic comprises D-peptides. In other instances, the peptidomimetic comprises L-peptides. Exemplary peptidomimetics include peptoids and P-peptides.
[0383] In some embodiments, the cargo is a nucleotidomimetic.A. Nucleic acids
[0384] In some instances, the cargo is a nucleic acid molecule. Examples of nucleic acid molecules include DNA, RNA, and mixtures of DNA and RNA. In some embodiments, the nucleic acid molecule comprises a hybrid of DNA and RNA.
[0385] In some instances, the nucleic acid molecule is a DNA polymer. In some cases, the DNA is a single stranded DNA polymer. In other cases, the DNA is a double stranded DNA polymer. In additional cases, the DNA is a hybrid of single and double stranded DNA polymers.
[0386] In some embodiments, the nucleic acid molecule is an RNA polymer, e.g., a single stranded RNA polymer, a double stranded RNA polymer, or a hybrid of single and double stranded RNA polymers. In some instances, the RNA comprises and / or encodes an antisense oligoribonucleotide, a siRNA, an mRNA, a tRNA, an rRNA, a snRNA, a shRNA, microRNA, or a non-coding RNA.
[0387] In some embodiments, the nucleic acid molecule is an antisense oligonucleotide, optionally comprising DNA, RNA, or a hybrid of DNA and RNA. In some instances, the nucleic acid molecule comprises and / or encodes an mRNA molecule. In some embodiments, the nucleic acid molecule comprises and / or encodes an RNAi molecule. In some cases, the RNAi molecule is a microRNA (miRNA) molecule. In other cases, the RNAi molecule is a siRNA molecule. The miRNA and / or siRNA are optionally double-stranded or as a hairpin, and further optionally encapsulated as precursor molecules.
[0388] In some embodiments, the nucleic acid molecule is for use in a nucleic acid-based therapy. In some instances, the nucleic acid molecule is for regulating gene expression (e.g., modulating mRNA translation or degradation), modulating RNA splicing, or RNA interference. In some cases, the nucleic acid molecule comprises and / or encodes an antisense oligonucleotide, microRNA molecule, siRNA molecule, mRNA molecule, for use in regulation of gene expression, modulating RNA splicing, or RNA interference.
[0389] In some instances, the nucleic acid molecule is for use in gene editing. Exemplary gene editing systems include, but are not limited to, CRISPR-Cas systems, zinc finger nuclease (ZFN) systems, and transcription activator-like effector nuclease (TALEN) systems. In some cases, the nucleic acid molecule comprises and / or encodes a component involved in the CRISPR-Cas systems, the ZFN systems, or the TALEN systems.
[0390] In some cases, the nucleic acid molecule is for use in antigen production for therapeutic and / or prophylactic vaccine production. For example, the nucleic acid molecule encodes an antigen that is expressed and elicits a desirable immune response (e.g., a pro- inflammatory immune response, an anti-inflammatory immune response, a tolerogenic immune response, an B cell response, an antibody response, a T cell response, a CD4+ T cell response, a CD 8+ T cell response, a Th l immune response, a Th2 immune response, a Th 17 immune response, a Treg immune response, an Ml macrophage response, an M2 macrophage response, or a combination thereof).
[0391] In some cases, the nucleic acid molecule comprises a nucleic acid enzyme. Nucleic acid enzymes are RNA molecules (e.g., ribozymes) or DNA molecules (e.g., deoxyribozymes) that have catalytic activities. In some instances, the nucleic acid molecule is a ribozyme. In other instances, the nucleic acid molecule is a deoxyribozyme. In some cases, the nucleic acid molecule is a MNAzyme, which functions as a biosensor and / or a molecular switch (see, e.g., Mokany, et al., (2010) MNAzymes, a versatile new class of nucleic acid enzymes that can function as biosensors and molecular switches, JACS 132(2): 1051-1059).
[0392] In some instances, exemplary targets of the nucleic acid molecule include, but are not limited to, ULI 23 (human cytomegalovirus), APOB, AR (androgen receptor) gene, KRAS, PCSK9, CFTR, and SMN (e.g., SMN2).
[0393] In some embodiments, the nucleic acid molecule is at least 5 nucleotides or more in length. In some instances, the nucleic acid molecule is at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 35, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 1000, at least 1500, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, or at least 9000 nucleotides or more in length. In some instances, the nucleic acid molecule is at least 10 nucleotides or more in length. In some instances, the nucleic acid molecule is at least 15 nucleotides or more in length. In some instances, the nucleic acid molecule is at least 20 nucleotides or more in length. In some instances, the nucleic acid molecule is at least 30 nucleotides or more in length. In some instances, the nucleic acid molecule is at least 40 nucleotides or more in length. In someinstances, the nucleic acid molecule is at least 50 nucleotides or more in length. In some instances, the nucleic acid molecule is at least 100 nucleotides or more in length. In some instances, the nucleic acid molecule is at least 200 nucleotides or more in length. In some instances, the nucleic acid molecule is at least 300 nucleotides or more in length. In some instances, the nucleic acid molecule is at least 500 nucleotides or more in length. In some instances, the nucleic acid molecule is at least 1000 nucleotides or more in length. In some instances, the nucleic acid molecule is at least 2000 nucleotides or more in length. In some instances, the nucleic acid molecule is at least 3000 nucleotides or more in length. In some instances, the nucleic acid molecule is at least 4000 nucleotides or more in length. In some instances, the nucleic acid molecule is at least 5000 nucleotides or more in length. In some instances, the nucleic acid molecule is at least 8000 nucleotides or more in length.
[0394] In some embodiments, the nucleic acid molecule is at most 10,000 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 15, at most 20, at most 25, at most 30, at most 35, at most 40, at most 50, at most 60, at most 70, at most 80, at most 90, at most 100, at most 150, at most 200, at most 250, at most 300, at most 400, at most 500, at most 1000, at most 1500, at most 2000, at most 3000, at most 4000, at most 5000, at most 6000, at most 7000, at most 8000, at most 9000, or at most 10,000 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 15 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 20 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 25 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 30 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 40 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 50 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 100 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 200 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 300 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 500 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 1000 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 2000 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 3000 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 4000 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 5000nucleotides or less in length. In some instances, the nucleic acid molecule is at most 8000 nucleotides or less in length. In some instances, the nucleic acid molecule is at most 9000 nucleotides or less in length.
[0395] In some embodiments, the nucleic acid molecule is about 100 nucleotides in length. In some instances, the nucleic acid molecule is about 6, about 7, about 8, about 9, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 150, about 200, about 250, about 300, about 400, about 500, about 1000, about 1500, about 2000, about 3000, about 4000, about 5000, about 6000, about 7000, about 8000, about 9000, or about 10000 nucleotides in length. In some instances, the nucleic acid molecule is about 10 nucleotides in length. In some instances, the nucleic acid molecule is about 15 nucleotides in length. In some instances, the nucleic acid molecule is about 20 nucleotides in length. In some instances, the nucleic acid molecule is about 25 nucleotides in length. In some instances, the nucleic acid molecule is about 30 nucleotides in length. In some instances, the nucleic acid molecule is about 40 nucleotides in length. In some instances, the nucleic acid molecule is about 50 nucleotides in length. In some instances, the nucleic acid molecule is about 100 nucleotides in length. In some instances, the nucleic acid molecule is about 150 nucleotides in length. In some instances, the nucleic acid molecule is about 200 nucleotides in length. In some instances, the nucleic acid molecule is about 300 nucleotides in length. In some instances, the nucleic acid molecule is about 500 nucleotides in length. In some instances, the nucleic acid molecule is about 1000 nucleotides in length. In some instances, the nucleic acid molecule is about 2000 nucleotides in length. In some instances, the nucleic acid molecule is about 3000 nucleotides in length. In some instances, the nucleic acid molecule is about 4000 nucleotides in length. In some instances, the nucleic acid molecule is about 5000 nucleotides in length. In some instances, the nucleic acid molecule is about 8000 nucleotides in length. In some instances, the nucleic acid molecule is about 9000 nucleotides in length.
[0396] In some embodiments, the nucleic acid molecule is from about 5 to about 10,000 nucleotides in length. In some instances, the nucleic acid molecule is from about 5 to about 9000 nucleotides in length, from about 5 to about 8000 nucleotides in length, from about 5 to about 7000 nucleotides in length, from about 5 to about 6000 nucleotides in length, from about 5 to about 5000 nucleotides in length, from about 5 to about 4000 nucleotides in length, from about 5 to about 3000 nucleotides in length, from about 5 to about 2000 nucleotides in length, from about5 to about 1000 nucleotides in length, from about 5 to about 500 nucleotides in length, from about 5 to about 100 nucleotides in length, from about 5 to about 50 nucleotides in length, from about 5 to about 40 nucleotides in length, from about 5 to about 30 nucleotides in length, from about 5 to about 25 nucleotides in length, from about 5 to about 20 nucleotides in length, from about 10 to about 10,000 nucleotides in length. In some instances, the nucleic acid molecule is from about 10 to about 9000 nucleotides in length, from about 10 to about 8000 nucleotides in length, from about 10 to about 7000 nucleotides in length, from about 10 to about 6000 nucleotides in length, from about 10 to about 5000 nucleotides in length, from about 10 to about 4000 nucleotides in length, from about 10 to about 3000 nucleotides in length, from about 10 to about 2000 nucleotides in length, from about 10 to about 1000 nucleotides in length, from about 10 to about 500 nucleotides in length, from about 10 to about 100 nucleotides in length, from about 10 to about 50 nucleotides in length, from about 10 to about 40 nucleotides in length, from about 10 to about 30 nucleotides in length, from about 10 to about 25 nucleotides in length, from about 10 to about 20 nucleotides in length, from about 50 to about 10,000 nucleotides in length. In some instances, the nucleic acid molecule is from about 50 to about 9000 nucleotides in length, from about 50 to about 8000 nucleotides in length, from about 50 to about 7000 nucleotides in length, from about 50 to about 6000 nucleotides in length, from about 50 to about 5000 nucleotides in length, from about 50 to about 4000 nucleotides in length, from about 50 to about 3000 nucleotides in length, from about 50 to about 2000 nucleotides in length, from about 50 to about 1000 nucleotides in length, from about 50 to about 500 nucleotides in length, from about 50 to about 100 nucleotides in length, from about 50 to about 50 nucleotides in length, from about 50 to about 40 nucleotides in length, from about 50 to about 30 nucleotides in length, from about 50 to about 25 nucleotides in length, from about 50 to about 20 nucleotides in length, from about 100 to about 10,000 nucleotides in length. In some instances, the nucleic acid molecule is from about 100 to about 9000 nucleotides in length, from about 100 to about 8000 nucleotides in length, from about 100 to about 7000 nucleotides in length, from about 100 to about 6000 nucleotides in length, from about 100 to about 5000 nucleotides in length, from about 100 to about 4000 nucleotides in length, from about 100 to about 3000 nucleotides in length, from about 100 to about 2000 nucleotides in length, from about 100 to about 1000 nucleotides in length, from about 100 to about 500 nucleotides in length, from about 100 to about 100 nucleotides in length, from about 100 to about 50 nucleotides in length, from about 100 to about40 nucleotides in length, from about 100 to about 30 nucleotides in length, from about 100 to about 25 nucleotides in length, from about 100 to about 20 nucleotides in length, from about 500 to about 10,000 nucleotides in length. In some instances, the nucleic acid molecule is from about 500 to about 9000 nucleotides in length, from about 500 to about 8000 nucleotides in length, from about 500 to about 7000 nucleotides in length, from about 500 to about 6000 nucleotides in length, from about 500 to about 5000 nucleotides in length, from about 500 to about 4000 nucleotides in length, from about 500 to about 3000 nucleotides in length, from about 500 to about 2000 nucleotides in length, from about 500 to about 1000 nucleotides in length, from about 500 to about 500 nucleotides in length, from about 500 to about 100 nucleotides in length, from about 500 to about 50 nucleotides in length, from about 500 to about 40 nucleotides in length, from about 500 to about 30 nucleotides in length, from about 500 to about 25 nucleotides in length, or from about 500 to about 20 nucleotides in length.
[0397] In some embodiments, the nucleic acid molecule comprises natural, synthetic, or artificial nucleotide analogues or bases. In some cases, the nucleic acid molecule comprises combinations of DNA, RNA and / or nucleotide analogues. In some instances, the synthetic or artificial nucleotide analogues or bases comprise modifications at one or more of ribose moiety, phosphate moiety, nucleoside moiety, or a combination thereof.
[0398] In some embodiments, a nucleotide analogue or artificial nucleotide base described above comprises a nucleic acid with a modification at a 2’ hydroxyl group of the ribose moiety. In some instances, the modification includes an H, OR, R, halo, SH, SR, NH2, NHR, NR2, or CN, wherein R is an alkyl moiety. Exemplary alkyl moiety includes, but is not limited to, halogens, sulfurs, thiols, thioethers, thioesters, amines (primary, secondary, or tertiary), amides, ethers, esters, alcohols and oxygen. In some instances, the alkyl moiety further comprises a modification. In some instances, the modification comprises an azo group, a keto group, an aldehyde group, a carboxyl group, a nitro group, a nitroso, group, a nitrile group, a heterocycle (e.g., imidazole, hydrazino or hydroxylamino) group, an isocyanate or cyanate group, or a sulfur containing group (e.g., sulfoxide, sulfone, sulfide, or disulfide). In some instances, the alkyl moiety further comprises a hetero substitution. In some instances, the carbon of the heterocyclic group is substituted by a nitrogen, oxygen or sulfur. In some instances, the heterocyclic substitution includes but is not limited to, morpholino, imidazole, and pyrrolidine.
[0399] In some instances, the modification at the 2’ hydroxyl group is a 2’-O-methyl modification or a 2’-O-methoxyethyl (2’-O-MOE) modification. In some cases, the 2’-O-methyl modification adds a methyl group to the 2’ hydroxyl group of the ribose moiety whereas the 2’0- methoxyethyl modification adds a methoxyethyl group to the 2’ hydroxyl group of the ribose moiety.
[0400] In some instances, the modification at the 2’ hydroxyl group is a 2’-O-aminopropyl modification in which an extended amine group comprising a propyl linker binds the amine group to the 2’ oxygen. In some instances, this modification neutralizes the phosphate-derived overall negative charge of the oligonucleotide molecule by introducing one positive charge from the amine group per sugar and thereby improves cellular uptake properties due to its zwitterionic properties.
[0401] In some instances, the modification at the 2’ hydroxyl group is a locked or bridged ribose modification (e g., locked nucleic acid or LNA) in which the oxygen molecule bound at the 2’ carbon is linked to the 4’ carbon by a methylene group, thus forming a 2'-C,4'-C-oxy- methylene-linked bicyclic ribonucleotide monomer.
[0402] In some embodiments, additional modifications at the 2’ hydroxyl group include 2'- deoxy, T-deoxy-2'-fluoro, 2'-O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O- DMAOE), 2'-O-dimethylaminopropyl (2'-0-DMAP), T-O- dimethylaminoethyloxy ethyl (2'-O- DMAEOE), or 2'-O-N-methylacetamido (2'-0-NMA).
[0403] In some embodiments, a nucleotide analogue comprises a modified base such as, but not limited to, N1 -methylpseudouridine, 5-propynyluridine, 5-propynylcytidine, 6- methyladenine, 6-methylguanine, N, N, -dimethyladenine, 2-propyladenine, 2propylguanine, 2- aminoadenine, 1 -methylinosine, 3 -methyluridine, 5-methylcytidine, 5 -methyluridine and other nucleotides having a modification at the 5 position, 5- (2- amino) propyl uridine, 5-halocytidine, 5-halouridine, 4-acetylcytidine, 1- methyladenosine, 2-methyladenosine, 3-methylcytidine, 6- methyluridine, 2- methylguanosine, 7-methylguanosine, 2, 2-dimethylguanosine, 5- methylaminoethyluridine, 5-methyloxyuridine, deazanucleotides (such as 7-deaza- adenosine, 6- azouridine, 6-azocytidine, or 6-azothymidine), 5-methyl-2-thiouridine, other thio bases (such as 2-thiouridine, 4-thiouridine, and 2-thiocytidine), dihydrouridine, pseudouridine, queuosine, archaeosine, naphthyl and substituted naphthyl groups, any O-and N-alkylated purines and pyrimidines (such as N6-methyladenosine, 5-methylcarbonylmethyluridine, uridine 5-oxyaceticacid, pyridine-4-one, or pyridine-2-one), phenyl and modified phenyl groups such as aminophenol or 2,4, 6-trimethoxy benzene, modified cytosines that act as G-clamp nucleotides, 8-substituted adenines and guanines, 5-substituted uracils and thymines, azapyrimidines, carboxyhydroxyalkyl nucleotides, carboxyalkylaminoalkyi nucleotides, and alkylcarbonylalkylated nucleotides. Modified nucleotides also include those nucleotides that are modified with respect to the sugar moiety, as well as nucleotides having sugars or analogs thereof that are not ribosyl. For example, the sugar moi eties, in some cases are or are based on, mannoses, arabinoses, glucopyranoses, galactopyranoses, 4'-thioribose, and other sugars, heterocycles, or carbocycles. The term nucleotide also includes universal bases. By way of example, universal bases include but are not limited to 3 -nitropyrrole, 5-nitroindole, or nebularine.
[0404] In some embodiments, a nucleotide analogue further comprises a morpholino, a peptide nucleic acid (PNA), a methylphosphonate nucleotide, a thiolphosphonate nucleotide, a 2’-fluoro N3-P5’-phosphoramidite, or a 1’, 5’- anhydrohexitol nucleic acid (HNA). Morpholino or phosphorodiamidate morpholino oligo (PMO) comprises synthetic molecules whose structure mimics natural nucleic acid structure but deviates from the normal sugar and phosphate structures. In some instances, the five-member ribose ring is substituted with a six-member morpholino ring containing four carbons, one nitrogen, and one oxygen. In some cases, the ribose monomers are linked by a phosphordiamidate group instead of a phosphate group. In such cases, the backbone alterations remove all positive and negative charges making morpholinos neutral molecules capable of crossing cellular membranes without the aid of cellular delivery agents such as those used by charged oligonucleotides.
[0405] In some embodiments, peptide nucleic acid (PNA) does not contain sugar ring or phosphate linkage and the bases are attached and appropriately spaced by oligoglycine-like molecules, therefore, eliminating a backbone charge.
[0406] In some embodiments, one or more modifications optionally occur at the internucleotide linkage. In some instances, modified internucleotide linkage includes, but is not limited to, phosphorothioates; phosphorodithioates; methylphosphonates; 5'- alkylenephosphonates; 5'-methylphosphonate; 3'-alkylene phosphonates; borontrifluoridates; borano phosphate esters and selenophosphates of 3'-5'linkage or 2'-5'linkage; phosphotriesters; thionoalkylphosphotriesters; hydrogen phosphonate linkages; alkyl phosphonates;alkylphosphonothioates; arylphosphonothioates; phosphoroselenoates; phosphorodiselenoates; phosphinates; phosphoramidates; 3'- alkylphosphoramidates; aminoalkylphosphoramidates; thionophosphoramidates; phosphoropiperazidates; phosphoroanilothioates; phosphoroanilidates; ketones; sulfones; sulfonamides; carbonates; carbamates; methylenehydrazos; methylenedimethylhydrazos; formacetals; thioformacetals; oximes; methyleneiminos; methylenemethyliminos; thioamidates; linkages with riboacetyl groups; aminoethyl glycine; silyl or siloxane linkages; alkyl or cycloalkyl linkages with or without heteroatoms of, for example, 1 to 10 carbons that are saturated or unsaturated and / or substituted and / or contain heteroatoms; linkages with morpholino structures, amides, or polyamides wherein the bases are attached to the aza nitrogens of the backbone directly or indirectly; and combinations thereof.
[0407] In some embodiments, one or more modifications comprise a modified phosphate backbone in which the modification generates a neutral or uncharged backbone. In some instances, the phosphate backbone is modified by alkylation to generate an uncharged or neutral phosphate backbone. As used herein, alkylation includes methylation, ethylation, and propylation. In some cases, an alkyl group, as used herein in the context of alkylation, refers to a linear or branched saturated hydrocarbon group containing from 1 to 6 carbon atoms. In some instances, exemplary alkyl groups include, but are not limited to, methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, sec-butyl, tert-butyl, n- pentyl, isopentyl, neopentyl, hexyl, isohexyl, 1, 1 -dimethylbutyl, 2,2-dimethylbutyl, 3.3- dimethylbutyl, and 2-ethylbutyl groups. In some cases, a modified phosphate is a phosphate group as described in U.S. Patent No. 9481905.
[0408] In some embodiments, additional modified phosphate backbones comprise methylphosphonate, ethylphosphonate, methylthiophosphonate, or methoxyphosphonate. In some cases, the modified phosphate is methylphosphonate. In some cases, the modified phosphate is ethylphosphonate. In some cases, the modified phosphate is methylthiophosphonate. In some cases, the modified phosphate is methoxyphosphonate.
[0409] In some embodiments, one or more modifications further optionally include modifications of the ribose moiety, phosphate backbone and the nucleoside, or modifications of the nucleotide analogues at the 3’ or the 5’ terminus. For example, the 3’ terminus optionally include a 3’ cationic group, or by inverting the nucleoside at the 3 ’-terminus with a 3 ’-3’ linkage. In another alternative, the 3’-terminus is optionally conjugated with an aminoalkyl group, e.g., a 3’ C5-aminoalkyl dT. In an additional alternative, the 3’-terminus is optionally conjugated withan abasic site, e.g., with an apurinic or apyrimidinic site. In some instances, the 5’-terminus is conjugated with an aminoalkyl group, e.g., a 5’-O-alkylamino substituent. In some cases, the 5’- terminus is conjugated with an abasic site, e.g., with an apurinic or apyrimidinic site.
[0410] In some embodiments, exemplary nucleic acid cargos include, but are not limited to, Fomivirsen, Mipomersen, AZD5312 (AstraZeneca), Nusinersen, and SB010 (Sterna Biologicals).
[0411] In some embodiments, a nucleic acid cargo comprises a recruitment domain. For example, an nucleic acid sequence (e.g., RNA) can comprise a secondary structure to bind to a cargo binding domain (e.g., an endogenous or engineered cargo binding domain) of a PNMA family polypeptide disclosed herein. The recruitment domain can be or comprise, for example, a hairpin loop element. The recruitment domain can be or comprise, for example, an aptamer sequence. The recruitment domain can be or comprise, for example, two or more aptamer sequences specific to the same or different cargo binding domains. In some embodiments, the recruitment domain comprises an aptamer that binds to MS2, PP7, QP, F2, GA, fir, JP501, M12, R17, BZ13, JP34, JP500, KU1, Mi l, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, (|)Cb5, 4*Cb8r, <j)Cbl2r, 4»Cb23r, 7s, or PRRl.B. Gene editing systems
[0412] In some embodiments, the cargo comprises or encodes a gene editing system or a component thereof. Non-limiting examples of gene editing tools and techniques include CRISPR, TALEN, zinc finger nuclease (ZFN), meganuclease, Mega-TAL, and transposon-based systems.
[0413] In some embodiments, the cargo comprises or encodes a CRISPR-associated polypeptide (Cas), zinc finger nuclease (ZFN), zinc finger associate gene regulation polypeptide, transcription activator-like effector nuclease (TALEN), transcription activator-like effector associated gene regulation polypeptides, meganuclease, natural master transcription factors, epigenetic modifying enzymes, recombinase, flippase, transposase, RNA-binding proteins (RBP), an Argonaute protein, any derivative thereof, any variant thereof, or any fragment thereof.
[0414] In some embodiments, the cargo comprises or encodes a CRISPR system or a component thereof. A CRISPR system can be utilized to facilitate insertion of a recombinant nucleic acid encoding a regulatable membrane protein or a component thereof into a cellgenome. For example, a CRISPR system can introduce a double stranded break at a target site in a genome or a random site of a genome.
[0415] In some embodiments, the cargo comprises or encodes a Cas protein. Non-limiting examples of Cas proteins that can be used in the CRISPR systems include Casl, Cas IB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl or Csxl2), CaslO, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, CsxlS, Csfl, Csf2, CsO, Csf4, Cpfl, c2cl, c2c3, Cas9HiFi, homologues thereof, and modified versions thereof. An unmodified CRISPR enzyme can have DNA cleavage activity, such as Cas9. A CRISPR enzyme can direct cleavage of one or both strands at a target sequence, such as within a target sequence and / or within a complement of a target sequence. For example, a CRISPR enzyme can direct cleavage of one or both strands within or within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of a target sequence. A Cas protein can be a high fidelity Cas protein. Alternatives to S. pyogenes Cas9 may include RNA-guided endonucleases from the Cpfl family that display cleavage activity in mammalian cells.
[0416] In some embodiments, the cargo comprises or encodes a guide RNA (gRNA), e.g., that complexes with a Cas protein.
[0417] In some embodiments, the cargo comprises or encodes a nucleotide sequence to be inserted in to a genome. In some embodiments, the cargo comprises or encodes a repair template, e.g., for homology-directed repair.
[0418] In some embodiments, the cargo comprises or encodes a dual nickase CRISPR system or a component thereof. A dual nickase approach may be used to introduce a double stranded break. Cas proteins can be mutated at certain amino acids within either nuclease domains, thereby deleting activity of one nuclease domain and generating a nickase Cas protein capable of generating a single strand break. A nickase along with two distinct guide RNAs targeting opposite strands may be utilized to generate a DSB within a target site (often referred to as a “double nick” or “dual nickase” CRISPR system).
[0419] In some embodiments, the cargo comprises or encodes a transposon based system or a component thereof. A transposon-based system can be utilized for insertion of a recombinant nucleic acid encoding a regulatable membrane protein of the disclosure or a component thereofinto a genome. A transposon can comprise a recombinant nucleic acid that can be inserted into a DNA sequence. A class I transposon can be transcribed into an RNA intermediate, then reverse transcribed and inserted into a DNA sequence. A class II transposon can comprise a DNA sequence that is excised from one DNA sequence and / or inserted into another DNA sequence. A class II transposon system can comprise (i) a transposon vector that contains a sequence (e.g., comprising a transgene) flanked by inverted terminal repeats, and (ii) a source for the transposase enzyme.
[0420] In some embodiments, the cargo comprises or encodes a TALEN system or a component thereof. TALENs can refer to engineered transcription activator-like effector nucleases that generally contain a central domain of DNA-binding tandem repeats and a cleavage domain. TALENs can be produced by fusing a TAL effector DNA binding domain to a DNA cleavage domain. In some cases, a DNA-binding tandem repeat comprises 33-35 amino acids in length and contains two hypervariable amino acid residues at positions 12 and 13 that can recognize at least one specific DNA base pair. A transcription activator-like effector (TALE) protein can be fused to a nuclease such as a wild-type or mutated Fokl endonuclease or the catalytic domain of Fokl.
[0421] In some embodiments, the cargo comprises or encodes a zinc finger nuclease (ZFN) or a variant, fragment, or derivative thereof. ZFN can refer to a fusion between a cleavage domain, such as a cleavage domain of Fokl, and at least one zinc finger motif (e.g., at least 2, at least 3, at least 4, or at least 5 zinc finger motifs) which can bind polynucleotides such as DNA and RNA.
[0422] In some embodiments, the cargo comprises or encodes a meganuclease. Meganucleases generally refer to rare-cutting endonucleases or homing endonucleases that can be highly sequence specific. Meganucleases can recognize DNA target sites ranging from at least 12 base pairs in length, e.g., from 12 to 40 base pairs, 12 to 50 base pairs, or 12 to 60 base pairs in length. Meganucleases can be modular DNA-binding nucleases such as any fusion protein comprising at least one catalytic domain of an endonuclease and at least one DNA binding domain or protein specifying a nucleic acid target sequence. The DNA-binding domain can contain at least one motif that recognizes single- or double-stranded DNA. A nuclease-active meganuclease can generate a double-stranded break. The meganuclease can be monomeric or dimeric. In some embodiments, the meganuclease is naturally-occurring (found in nature) orwild-type, and in other instances, the meganuclease is non-natural, artificial, engineered, synthetic, rationally designed, or man-made. In some embodiments, the meganuclease of the present disclosure includes an I-Crel meganuclease, I-Ceul meganuclease, I-Msol meganuclease, I-Scel meganuclease, variants thereof, derivatives thereof, and fragments thereof.C. Small molecules
[0423] In some embodiments, the cargo is a small molecule. In some instances, the small molecule is an inhibitor (e.g., a pan inhibitor or a selective inhibitor). In other instances, the small molecule is an activator. In additional cases, the small molecule is an agonist, antagonist, a partial agonist, a mixed agonist / antagonist, or a competitive antagonist.
[0424] In some embodiments, the small molecule is a drug that falls under the class of analgesics, antianxiety drugs, antiarrhythmics, antibacterials, antibiotics, anticoagulants and thrombolytics, anticonvulsants, antidepressants, antidiarrheals, antiemetics, antifungals, antihistamines, antihypertensives, anti-inflammatories, antineoplastics, antipsychotics, antipyretics, antivirals, barbiturates, beta-blockers, bronchodilators, common cold treatments, corticosteroids, cough suppressants, cytotoxics, decongestants, diuretics, expectorant, hormones, hypoglycemics, immunosuppressives, laxatives, muscle relaxants, sex hormones, sleeping drugs, or tranquilizers.
[0425] In some embodiments, the small molecule is an inhibitor, e.g., an inhibitor of a kinase pathway such as the Tyrosine kinase pathway or a Serine / Threonine kinase pathway. In some cases, the small molecule is a dual protein kinase inhibitor. In some cases, the small molecule is a lipid kinase inhibitor.
[0426] In some cases, the small molecule is a neuraminidase inhibitor.
[0427] In some cases, the small molecule is a carbonic anhydrase inhibitor.
[0428] In some embodiments, exemplary targets of the small molecule include, but are not limited to, vascular endothelial growth factor receptor 1 (VEGFR1), vascular endothelial growth factor receptor 2 (VEGFR2), vascular endothelial growth factor receptor 3 (VEGFR3), fibroblast growth factor receptor 1 (FGFR1), fibroblast growth factor receptor 2 (FGFR2), fibroblast growth factor receptor 3 (FGFR3), fibroblast growth factor receptor 4 (FGFR4), cyclin- dependent kinase 4 (CDK4), cyclin-dependent kinase 6 (CDK6), a receptor tyrosine kinase, a phosphoinositide 3-kinase (PI3K) isoform (e.g., PI3K6, also known as pl 108), lanus kinase 1(JAK1), Janus kinase 3 (JAK3), a receptor from the family of platelet-derived growth factor receptors (PDFG-R), and carbonic anhydrase (e.g., carbonic anhydrase I).
[0429] In some embodiments, the small molecule targets a viral protein, e.g., a viral envelope protein. In some embodiments, the small molecule decreases viral adsorption to a host cell. In some embodiments, the small molecule decreases viral entry into a host cell. In some embodiments, the small molecule decreases viral replication in a host or a host cell. In some embodiments, the small molecule decreases viral assembly.
[0430] In some embodiments, exemplary small molecule cargos include, but are not limited to, lenvatinib, palbociclib, regorafenib, idelalisib, tofacitinib, nintedanib, zanamivir, ethoxzolamide, and artemisinin.D. Proteins and peptides
[0431] In some embodiments, the cargo is a protein. In some instances, the protein is a full- length protein. In other instances, the protein is a fragment, e.g., a functional fragment. In some cases, the protein is a naturally occurring protein. In additional cases, the protein is a de novo engineered protein. In further cases, the protein is a fusion protein. In further cases, the protein is a recombinant protein. Exemplary proteins include, but are not limited to, Fc fusion proteins, anticoagulants, blood factors, bone morphogenetic proteins, enzymes, growth factors, hormones, interferons, interleukins, and thrombolytics.
[0432] In some instances, the protein is for use in an enzyme replacement therapy.
[0433] In some cases, the protein is for use in antigen production for therapeutic and / or prophylactic vaccine production. For example, the protein comprises an antigen that elicits a desirable immune response (e.g., a pro-inflammatory immune response, an anti-inflammatory immune response, a tolerogenic immune response, an B cell response, an antibody response, a T cell response, a CD4+ T cell response, a CD8+ T cell response, a Thl immune response, a Th2 immune response, a Thl7 immune response, a Treg immune response, an Ml macrophage response, an M2 macrophage response, or a combination thereof).
[0434] In some instances, exemplary protein cargos include, but are not limited to, romiplostim, liraglutide, a human growth hormone (rHGH), human insulin (BHI), follicle- stimulating hormone (FSH), Factor VIII, erythropoietin (EPO), granulocyte colony-stimulating factor (G-CSF), alpha-galactosidase A, alpha-L-iduronidase, N-acetylgalactosamine-4-sulfatase,dornase alfa, tissue plasminogen activator (TP A), glucocerebrosidase, interferon-beta-la, insulinlike growth factor 1 (IGF-1), or rasburicase.
[0435] In some embodiments, the cargo is a peptide. In some instances, the peptide is a naturally occurring peptide. In other instances, the peptide is an artificial engineered peptide or a recombinant peptide. In some cases, the peptide targets a G-protein coupled receptor, an ion channel, a microbe, an anti-microbial target, a catalytic or other Ig-family of receptors, an intracellular target, a membrane-anchored target, or an extracellular target.
[0436] In some cases, the peptide comprises at least 2 amino acids. In some cases, the peptide comprises at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 amino acids. In some cases, the peptide comprises at least 10 amino acids. In some cases, the peptide comprises at least 15 amino acids. In some cases, the peptide comprises at least 20 amino acids. In some cases, the peptide comprises at least 30 amino acids. In some cases, the peptide comprises at least 40 amino acids. In some cases, the peptide comprises at least 50 amino acids. In some cases, the peptide comprises at least 60 amino acids. In some cases, the peptide comprises at least 70 amino acids. In some cases, the peptide comprises at least 80 amino acids. In some cases, the peptide comprises at least 90 amino acids. In some cases, the peptide comprises at least 100 amino acids.
[0437] In some cases, the peptide comprises at most 3 amino acids. In some cases, the peptide comprises at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, at most 15, at most 20, at most 25, at most 30, at most 35, at most 40, at most 45, at most 50, at most 60, at most 70, at most 80, at most 90, or at most 100 amino acids. In some cases, the peptide comprises at most 10 amino acids. In some cases, the peptide comprises at most 15 amino acids. In some cases, the peptide comprises at most 20 amino acids. In some cases, the peptide comprises at most 30 amino acids. In some cases, the peptide comprises at most 40 amino acids. In some cases, the peptide comprises at most 50 amino acids. In some cases, the peptide comprises at most 60 amino acids. In some cases, the peptide comprises at most 70 amino acids. In some cases, the peptide comprises at most 80 amino acids. In some cases, the peptide comprises at most 90 amino acids. In some cases, the peptide comprises at most 100 amino acids.
[0438] In some cases, the peptide comprises from about 1 to about 10 kDa. In some cases, the peptide comprises from about 1 to about 9 kDa, about 1 to about 6 kDa, about 1 to about 5 kDa, about 1 to about 4 kDa, about 1 to about 3 kDa, about 2 to about 8 kDa, about 2 to about 6 kDa, about 2 to about 4 kDa, about 1.2 to about 2.8 kDa, about 1.5 to about 2.5 kDa, or about 1.5 to about 2 kDa.
[0439] In some embodiments, the peptide is a cyclic peptide. In some instances, the cyclic peptide is a macrocyclic peptide. In other instances, the cyclic peptide is a constrained peptide. The cyclic peptides are assembled with varied linkages, such as for example, head-to-tail, head- to-side-chain, side-chain-to-tail, and side-chain-to-side-chain linkages. In some instances, a cyclic peptide (e.g., a macrocyclic or a constrained peptide) has a molecular weight from about 500 Dalton to about 2000 Dalton. In other instances, a cyclic peptide (e.g., a macrocyclic or a constrained peptide) ranges from about 10 amino acids to about 100 amino acids, from about 10 amino acids to about 70 amino acids, or from about 10 amino acids to about 50 amino acids.
[0440] In some cases, the peptide is for use in antigen production for therapeutic and / or prophylactic vaccine production. For example, the peptide comprises an antigen that elicits a desirable immune response (e.g., a pro-inflammatory immune response, an anti-inflammatory immune response, a tolerogenic immune response, an B cell response, an antibody response, a T cell response, a CD4+ T cell response, a CD8+ T cell response, a Thl immune response, a Th2 immune response, a Thl 7 immune response, a Treg immune response, an Ml macrophage response, an M2 macrophage response, or a combination thereof).
[0441] In some embodiments, the peptide comprises natural amino acids, unnatural amino acids, or a combination thereof. In some instances, an amino acid residue refers to a molecule containing both an amino group and a carboxyl group. Suitable amino acids include, without limitation, both the D- and L-isomers of the naturally occurring amino acids, as well as non- naturally occurring amino acids prepared by organic synthesis or other metabolic routes. The term amino acid, as used herein, includes, without limitation, a-amino acids, natural amino acids, non-natural amino acids, and amino acid analogs.
[0442] In some instances, a-amino acid refers to a molecule containing both an amino group and a carboxyl group bound to a carbon which is designated the a-carbon.
[0443] In some instances, P-amino acid refers to a molecule containing both an amino group and a carboxyl group in a configuration.
[0444] In some embodiments, an amino acid analog is a racemic mixture. In some instances, the D isomer of the amino acid analog is used. In some cases, the L isomer of the amino acid analog is used. In some instances, the amino acid analog comprises chiral centers that are in the R or S configuration.
[0445] In some embodiments, exemplary peptide cargos include, but are not limited to, peginesatide, insulin, adrenocorticotropic hormone (ACTH), calcitonin, oxytocin, vasopressin, octreolide, and leuprorelin.
[0446] In some embodiments, exemplary peptide cargos include, but are not limited to, Telavancin, Dalbavancin, Oritavancin, Anidulafungin, Lanreotide, Pasireotide, Romidepsin, Linaclotide, and Peginesatide.E. Antibodies
[0447] In some embodiments, the cargo is an antibody or a binding fragment thereof. In some instances, the antibody or binding fragment thereof comprises a humanized antibody or binding fragment thereof, murine antibody or binding fragment thereof, chimeric antibody or binding fragment thereof, monoclonal antibody or binding fragment thereof, bispecific antibody or biding fragment thereof, monovalent Fab’, divalent Fab2, F(ab)'3 fragments, single-chain variable fragment (scFv), bis-scFv, (scFv)2, diabody, minibody, nanobody, triabody, tetrabody, disulfide stabilized Fv protein (dsFv), single-domain antibody (sdAb), Ig NAR, camelid antibody or binding fragment thereof, or a chemically modified derivative thereof.
[0448] In some instances, the antibody or binding fragment thereof recognizes a cell surface protein. In some instances, the cell surface protein is an antigen expressed by a cancerous cell. In some instances, the cell surface protein is a neoepitope. In some instances, the cell surface protein comprises one or more mutations compared to a wild-type protein. Exemplary cancer antigens include, but are not limited to, alpha fetoprotein, ASLG659, B7-H3, BAFF-R, Brevican, CA125 (MUC16), CA15-3, CA19-9, carcinoembryonic antigen (CEA), CA242, CRIPTO (CR, CR1, CRGF, CRIPTO, TDGF1, teratocarcinoma-derived growth factor), CTLA-4, CXCR5, E16 (LAT1, SLC7A5), FcRH2 (IFGP4, IRTA4, SPAP1 A (SH2 domain containing phosphatase anchor protein la), SPAP1B, SPAP1C), epidermal growth factor, ETBR, Fc receptor-like protein 1 (FCRH1), GEDA, HLA-DOB (Beta subunit of MHC class II molecule (la antigen), human chorionic gonadotropin, ICOS, IL-2 receptor, IL2ORot, Immunoglobulin superfamily receptor translocation associated 2 (IRTA2), L6, Lewis Y, Lewis X, MAGE-1, MAGE-2, MAGE-3,MAGE 4, MARTI, mesothelin, MDP, MPF (SMR, MSLN), MCP1 (CCL2), macrophage inhibitory factor (MIF), MPG, MSG783, mucin, MUC1-KLH, Napi3b (SLC34A2), nectin-4, Neu oncogene product, NCA, placental alkaline phosphatase, prostate specific membrane antigen (PMSA), prostatic acid phosphatase, PSCA hlg, anti -transferrin receptor, p97, Purinergic receptor P2X ligand-gated ion channel 5 (P2X5), LY64 (Lymphocyte antigen 64 (RP105), gplOO, P21, six transmembrane epithelial antigen of prostate (STEAP1), STEAP2, Sema 5b, tumor-associated glycoprotein 72 (TAG-72), TrpM4 (BR22450, FLJ20041, TRPM4, TRPM4B, transient receptor potential cation channel, subfamily M, member 4) and the like.
[0449] In some instances, the cell surface protein comprises clusters of differentiation (CD) cell surface markers. Exemplary CD cell surface markers include, but are not limited to, CD1, CD2, CD3, CD4, CD5, CD6, CD7, CD8, CD9, CD10, CDl la, CDl lb, CDl lc, CDl ld, CDwl2, CD13, CD14, CD15, CD15s, CD16, CDwl7, CD18, CD19, CD20, CD21, CD22, CD23, CD24, CD25, CD26, CD27, CD28, CD29, CD30, CD31, CD32, CD33, CD34, CD35, CD36, CD37, CD38, CD39, CD40, CD41, CD42, CD43, CD44, CD45, CD45RO, CD45RA, CD45RB, CD46, CD47, CD48, CD49a, CD49b, CD49c, CD49d, CD49e, CD49f, CD50, CD51, CD52, CD53, CD54, CD55, CD56, CD57, CD58, CD59, CDw60, CD61, CD62E, CD62L (L-selectin), CD62P, CD63, CD64, CD65, CD66a, CD66b, CD66c, CD66d, CD66e, CD71, CD79 (e.g., CD79a, CD79b), CD90, CD95 (Fas), CD103, CD104, CD125 (IL5RA), CD134 (0X40), CD137 (4- 1BB), CD152 (CTLA-4), CD221, CD274, CD279 (PD-1), CD319 (SLAMF7), CD326 (EpCAM), and the like.
[0450] In some embodiments, exemplary antibodies or binding fragments thereof include, but are not limited to, zalutumumab (HuMax-EFGr, Genmab), abagovomab (Menarini), abituzumab (Merck), adecatumumab (MT201), alacizumab pegol, alemtuzumab (Campath®, MabCampath, or Campath- 1H; Leukosite), AlloMune (BioTransplant), amatuximab (Morphotek, Inc.), anti-VEGF (Genetech), anatumomab mafenatox, apolizumab (hulDlO), ascrinvacumab (Pfizer Inc ), atezolizumab (MPDL3280A; Genentech / Roche), B43.13 (OvaRex, AltaRex Corporation), basiliximab (Simulect®, Novartis), belimumab (Benlysta®, GlaxoSmithKline), bevacizumab (Avastin®, Genentech), blinatumomab (Blincyto, AMG103; Amgen), BEC2 (ImGlone Systems Inc.), carlumab (Janssen Biotech), catumaxomab (Removab, Trion Pharma), CEAcide (Immunomedics), Cetuximab (Erbitux®, ImClone), citatuzumab bogatox (VB6-845), cixutumumab (IMC-A12, ImClone Systems Inc.), conatumumab (AMG 655, Amgen),dacetuzumab (SGN-40, huS2C6; Seattle Genetics, Inc ), daratumumab (Darzalex®, Janssen Biotech), detumomab, drozitumab (Genentech), durvalumab (Medlmmune), dusigitumab (Medlmmune), edrecolomab (MAbl7-lA, Panorex, Glaxo Wellcome), elotuzumab (Empliciti™, Bristol-Myers Squibb), emibetuzumab (Eli Lilly), enavatuzumab (Facet Biotech Corp.), enfortumab vedotin (Seattle Genetics, Inc.), enoblituzumab (MGA271, MacroGenics, Inc ), ensituxumab (Neogenix Oncology, Inc.), epratuzumab (LymphoCide, Immunomedics, Inc.), ertumaxomab (Rexomun®, Trion Pharma), etaracizumab (Abegrin, Medlmmune), farletuzumab (MORAb-003, Morphotek, Inc), FBTA05 (Lymphomun, Trion Pharma), ficlatuzumab (AVEO Pharmaceuticals), figitumumab (CP-751871, Pfizer), flanvotumab (ImClone Systems), fresolimumab (GC1008, Aanofi-Aventis), futuximab, glaximab, ganitumab (Amgen), girentuximab (Rencarex®, Wilex AG), IMAB362 (Claudiximab, Ganymed Pharmaceuticals AG), imalumab (Baxalta), IMC-1C11 (ImClone Systems), IMC-C225 (Imclone Systems Inc.), imgatuzumab (Genentech / Roche), intetumumab (Centocor, Inc.), ipilimumab (Yervoy®, Bristol- Myers Squibb), iratumumab (Medarex, Inc.), isatuximab (SAR650984, Sanofi-Aventis), labetuzumab (CEA-CIDE, Immunomedics), lexatumumab (ETR2-ST01, Cambridge Antibody Technology), lintuzumab (SGN-33, Seattle Genetics), lucatumumab (Novartis), lumiliximab, mapatumumab (HGS-ETR1, Human Genome Sciences), matuzumab (EMD 72000, Merck), milatuzumab (hLLl, Immunomedics, Inc.), mitumomab (BEC-2, ImClone Systems), narnatumab (ImClone Systems), necitumumab (Portrazza™, Eli Lilly), nesvacumab (Regeneron Pharmaceuticals), nimotuzumab (h-R3, BIOMAb EGFR, TheraCIM, Theraloc, or CIMAher; Biotech Pharmaceutical Co.), nivolumab (Opdivo®, Bristol-Myers Squibb), obinutuzumab (Gazyva or Gazyvaro; Hoffmann-La Roche), ocaratuzumab (AME-133v, LY2469298; Mentrik Biotech, LLC), ofatumumab (Arzerra®, Genmab), onartuzumab (Genentech), Ontuxizumab (Morphotek, Inc.), oregovomab (OvaRex®, AltaRex Corp.), otlertuzumab (Emergent BioSolutions), panitumumab (ABX-EGF, Amgen), pankomab (Glycotope GMBH), parsatuzumab (Genentech), patritumab, pembrolizumab (Keytruda®, Merck), pemtumomab (Theragyn, Antisoma), pertuzumab (Perjeta, Genentech), pidilizumab (CT-011, Medivation), polatuzumab vedotin (Genentech / Roche), pritumumab, racotumomab (Vaxira®, Recombio), ramucirumab (Cyramza®, ImClone Systems Inc.), rituximab (Rituxan®, Genentech), robatumumab (Schering-Plough), Seribantumab (Sanofi / Merrimack Pharmaceuticals, Inc.), sibrotuzumab, siltuximab (Sylvant™, Janssen Biotech), Smart MI95 (Protein Design Labs, Inc.),Smart ID10 (Protein Design Labs, Inc.), tabalumab (LY2127399, Eli Lilly), taplitumomab paptox, tenatumomab, teprotumumab (Roche), tetulomab, TGN1412 (CD28-SuperMAB or TAB08), tigatuzumab (CD- 1008, Daiichi Sankyo), tositumomab, trastuzumab (Herceptin®), tremelimumab (CP-672,206; Pfizer), tucotuzumab celmoleukin (EMD Pharmaceuticals), ublituximab, urelumab (BMS-663513, Bristol-Myers Squibb), volociximab (M200, Biogen Idee), and zatuximab.
[0451] In some instances, the antibody or binding fragments thereof is an antibody-drug conjugate (ADC). In some cases, the payload of the ADC comprises, for example, but is not limited to, an auristatin derivative, maytansine, a maytansinoid, a taxane, a calicheamicin, cemadotin, a duocarmycin, a pyrrolobenzodiazepine (PDB), or a tubulysin. In some instances, the payload comprises monomethyl auristatin E (MMAE) or monomethyl auristatin F (MMAF). In some instances, the payload comprises DM2 (mertansine) or DM4. In some instances, the payload comprises a pyrrolobenzodiazepine dimer.V. VECTORS AND EXPRESSION SYSTEMS
[0452] RTL or PNMA family (e.g., RTL10 or PEG10) and endo-Gag polypeptides of the disclosure can be encoded by nucleic acids, for example, for expression in a host cell or using a cell-free expression system. In certain embodiments, the RTL or PNMA polypeptides, endo-Gag polypeptides, engineered RTL or PNMA and engineered endo-Gag polypeptides described herein are encoded by vectors, e.g., plasmid vectors. In some embodiments, vectors include any suitable vector derived from either a eukaryotic or prokaryotic source. In some cases, vectors are obtained from bacteria (e.g. E. coli), insects, yeast (e.g. Pichia pastoris), algae, or mammalian sources.
[0453] Illustrative bacterial vectors include pACYC177, pASK75, pBAD vector series, pBADM vector series, pET vector series, pETM vector series, pGEX vector series, pHAT, pHAT2, pMal-c2, pMal-p2, pQE vector series, pRSET A, pRSET B, pRSET C, pTrcHis2 series, pZA31-Luc, pZE21-MCS-l, pFLAG ATS, pFLAG CTS, pFLAG MAC, pFLAG Shift-12c, pTAC -MAT-1, pFLAG CTC, and pTAC-MAT-2.
[0454] Illustrative insect vectors include pFastBacl, pFastBac DUAL, pFastBac ET, pFastBac HTa, pFastBac HTb, pFastBac HTc, pFastBac M30a, pFastBact M30b, pFastBac, M30c, pVL1392, pVL1393, pVL1393 M10, pVL1393 Ml 1, pVL1393 Ml 2, FLAG vectors such as pPolh-FLAGl or pPolh-MAT 2, or MAT vectors such as pPolh-MATl, or pPolh-MAT2.
[0455] In some cases, yeast vectors include Gateway® pDEST™ 14 vector, Gateway® pDEST™ 15 vector, Gateway® pDEST™ 17 vector, Gateway® pDEST™ 24 vector, Gateway® pYES-DEST52 vector, pBAD-DEST49 Gateway® destination vector, pAO815 Pichia vector, pFLDl Pichi pastoris vector, pGAPZA,B, & C Pichia pastoris vector, pPIC3.5K Pichia vector, pPIC6 A, B, & C Pichia vector, pPIC9K Pichia vector, pTEFl / Zeo, pYES2 yeast vector, pYES2 / CT yeast vector, pYES2 / NT A, B, & C yeast vector, or pYES3 / CT yeast vector.
[0456] Illustrative algae vectors include pChlamy-4 vector or MCS vector.
[0457] Examples of mammalian vectors include transient expression vectors or stable expression vectors. Mammalian transient expression vectors include p3xFLAG-CMV 8, pFLAG- Myc-CMV 19, pFLAG-Myc-CMV 23, pFLAG-CMV 2, pFLAG-CMV 6a,b,c, pFLAG-CMV 5.1, pFLAG-CMV 5a,b,c, p3xFLAG-CMV 7.1, pFLAG-CMV 20, p3xFLAG-Myc-CMV 24, pCMV-FLAG-MATl, pCMV-FLAG-MAT2, pBICEP-CMV 3, or pBICEP-CMV 4. Mammalian stable expression vector include pFLAG-CMV 3, p3xFLAG-CMV 9, p3xFLAG-CMV 13, pFLAG-Myc-CMV 21, p3xFLAG-Myc-CMV 25, pFLAG-CMV 4, p3xFLAG-CMV 10, p3xFLAG-CMV 14, pFLAG-Myc-CMV 22, p3xFLAG-Myc-CMV 26, pBICEP-CMV 1, or pBICEP-CMV 2.
[0458] In certain embodiments, the RTL or PNMA polypeptides, endo-Gag polypeptides, engineered RTL or PNMA and engineered endo-Gag polypeptides described herein and / or heterologous cargos are expressed using a cell-free expression system. In some instances, a cell- free system is a mixture of cytoplasmic and / or nuclear components from a cell, or isolated transcription and / or translation machinery, and is used for in vitro RNA and / or protein expression. In some cases, a cell-free system utilizes either prokaryotic cell components or eukaryotic cell components, or transcription and / or translation machinery derived therefrom. In some embodiments, nucleic acid synthesis is obtained in a cell-free system based on for example Drosophila cell, Xenopus egg, or HeLa cells (ATCC® CCL-2™). Exemplary cell-free systems include, but are not limited to, E. coli S30 Extract system, E. coli T7 S30 system, or PURExpress®.
[0459] In certain embodiments, the RTL (e.g., PEG10 or RTL 10) or PNMA polypeptides, endo-Gag polypeptides, engineered RTL or PNMA and engineered endo-Gag polypeptides described herein, and / or heterologous cargos are expressed by a host cell. In some embodiments, a host cell includes any suitable cell such as a naturally derived cell or a genetically modifiedcell. In some instances, a host cell is a production host cell. In some instances, a host cell is a eukaryotic cell. In other instances, a host cell is a prokaryotic cell. In some cases, a eukaryotic cell includes fungi (e.g., a yeast cell), an animal cell, or a plant cell. In some cases, a prokaryotic cell is a bacterial cell. Examples of bacterial cell include gram-positive bacteria or gram-negative bacteria. In some embodiments the gram-negative bacteria are anaerobic, rod-shaped, or both.
[0460] In some instances, gram-positive bacteria include Actinobacteria, Firmicutes or Teneri cutes. In some cases, gram-negative bacteria include Aquificae, Deinococcus-Thermus, Fibrobacteres-Chlorobi / Bacteroidetes (FCB group), Fusobacteria, Gemmatimonadetes, Nitrospirae, Planctomycetes-Verrucomicrobia / Chlamydiae (PVC group), Proteobacteria, Spirochaetes or Synergistetes. In some embodiments, bacteria is Acidobacteria, Chloroflexi, Chrysiogenetes, Cyanobacteria, Deferribacteres, Dictyoglomi, Thermodesulfobacteria or Thermotogae. In some embodiments, a bacterial cell is Escherichia coli, Clostridium botulinum, or Coli bacilli.
[0461] Exemplary prokaryotic host cells include, but are not limited to, BL21, Maehl™, DH10B™, TOP 10, DH5a, DHIOBac™, OmniMax™, MegaX™, DH12S™, INV110, TOPI OF’, INVaF, TOP10 / P3, ccdB Survival, PIR1, PIR2, Stbl2™, Stbl3™, or Stbl4™.
[0462] In some instances, animal cells include a cell from a vertebrate or from an invertebrate. In some cases, an animal cell includes a cell from a marine invertebrate, fish, insects, amphibian, reptile, mammal, or human. In some cases, a fungus cell includes a yeast cell, such as brewer’ s yeast, baker’s yeast, or wine yeast.
[0463] Fungi include ascomycetes such as yeast, mold, filamentous fungi, basidiomycetes, or zygomycetes. In some instances, yeast includes Ascomycota or Basidiomycota. In some cases, Ascomycota includes Saccharomycotina (true yeasts, e.g. Saccharomyces cerevisiae (baker’s yeast)) or Taphrinomycotina (e g. Schizosaccharomycetes (fission yeasts)). In some cases, Basidiomycota includes Agaricomycotina (e.g. Tremellomycetes') or Pucciniomycotina (e.g. Microbotryomyce tes) .
[0464] Exemplary yeast or filamentous fungi include, for example, the genus: Saccharomyces, Schizosaccharomyces, Candida, Pichia, Hansenula, Kluyveromyces, Zygosaccharomyces, Yarrowia, Trichosporon, Rhodosporidi, Aspergillus, Fusarium, or Trichoderma. Exemplary yeast or filamentous fungi include, for example, the species: Saccharomyces cerevisiae, Schizosaccharomyces pombe, Candida utilis, Candida boidini,Candida albicans, Candida tropicalis, Candida stellatoidea, Candida glabrata, Candida krusei, Candida parapsilosis, Candida guilliermondii, Candida viswanathii, Candida lusitaniae, Rhodotorula mucilaginosa, Pichia metanolica, Pichia angusta, Pichia pastoris, Pichia anomala, Hansenula polymorpha, Kluyveromyces lactis, Zygosaccharomyces rouxii, Yarrowia lipolytica, Trichosporon pullulans, Rhodosporidium toru-Aspergillus niger, Aspergillus nidulans, Aspergillus awamori, Aspergillus oryzae, Trichoderma reesei, Yarrowia lipolytica, Brettanomyces bruxellensis, Candida stellata, Schizosaccharomyces pombe, Torulaspora delbrueckii, Zygosaccharomyces bailii, Cryptococcus neoformans, Cryptococcus gattii, or Saccharomyces boulardii.
[0465] Exemplary yeast host cells include, but are not limited to, Pichia pastoris yeast strains such as GS115, KM71H, SMD1168, SMD1168H, and X-33; and Saccharomyces cerevisiae yeast strain such as INVScl.
[0466] In some instances, additional animal cells include cells obtained from a mollusk, arthropod, annelid or sponge. In some cases, an additional animal cell is a mammalian cell, e.g., from a human, primate, ape, equine, bovine, porcine, canine, feline or rodent. In some cases, a rodent includes mouse, rat, hamster, gerbil, hamster, chinchilla, fancy rat, or guinea pig.
[0467] Exemplary mammalian host cells include, but are not limited to, 293A cell line, 293FT cell line, 293F cells , 293 H cells, CHO DG44 cells, CHO-S cells, CHO-K1 cells, Expi293F™ cells, Flp-In™ T-REx™ 293 cell line, Flp-In™-293 cell line, Flp-In™-3T3 cell line, Flp-In™-BHK cell line, Flp-In™-CHO cell line, Flp-In™-CV-l cell line, Flp-In™-Jurkat cell line, FreeStyle™ 293-F cells, FreeStyle™ CHO-S cells, GripTite™ 293 MSR cell line, GS- CHO cell line, HepaRG™ cells, T-REx™ Jurkat cell line, Per.C6 cells, T-REx™-293 cell line, T-REx™-CHO cell line, and T-REx™-HeLa cell line.
[0468] In some instances, a mammalian host cell is a primary cell. In some instances, a mammalian host cell is a stable cell line, or a cell line that has incorporated a genetic material of interest into its own genome and has the capability to express the product of the genetic material after many generations of cell division. In some cases, a mammalian host cell is a transient cell line, or a cell line that has not incorporated a genetic material of interest into its own genome and does not have the capability to express the product of the genetic material after many generations of cell division.
[0469] Exemplary insect host cells include, but are not limited to, Drosophila S2 cells, Sf9 cells, Sf21 cells, High Five™ cells, and expresSF+® cells. Exemplary insect cell lines include, but are not limited to, strains from Chlamydomonas reinhardtii 137c, or Synechococcus elongatus PPC 7942. In some instances, plant cells include cells from algae.VI. METHODS
[0470] Disclosed herein, in certain embodiments, are methods of preparing a capsid, e.g., a capsid which encapsulates a heterologous cargo. In some embodiments, RTL or PNMA family (e.g., RTL10, PEG10, PNMA2, PNMA5, and / or other RTL or PNMA polypeptides) and / or endo-Gag polypeptides of the disclosure exhibit favorable properties for capsid assembly, disassembly, and / or reassembly.
[0471] The presence of RTL or PNMA or endo-Gag polypeptides in a capsid form or noncapsid form can be determined using, for example, size exclusion chromatography, multi-angle dynamic light scattering (MADLS), electron microscopy (EM, e.g., scanning or transmission EM), or a combination thereof. Capsid assembly, disassembly, and re-assembly efficiency can be calculated, for example, by comparing the quantity of protein present in assembled capsids to the total quantity of protein, or to the quantity of protein present in particles of smaller size and / or larger size than the capsids. In some embodiments, when calculating reassembly, the amount of protein loss during disassembly and reassembly can be measured and accounted for in the calculations.
[0472] In some embodiments, measuring capsid assembly efficiency comprises quantifying the amount of purified protein that is present in a capsid state (e.g., as capsid particles), and quantifying the amount of purified protein that is present in non-capsid states (e.g., nonassembled and / or partially-assembled states). For example, size exclusion chromatography can be used to separate monomers and oligomers from assembled capsids, and to quantify the amount of protein that is present in a capsid state versus un-assembled or partially-assembled non-capsid states.
[0473] In some embodiments, measuring capsid disassembly efficiency comprises quantifying the amount of capsid-forming protein that remains in solution after disassembling from a capsid state to a non-capsid state (e.g., oligomers, capsomers, and / or monomers). Capsomers can comprise partially assembled capsid subunits, for example, comprising at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, atleast 12, at least 13, at least 14, at least 15, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, at most 11, at most 12, at most 13, at most 14, at most 15, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, or about 15 endo Gag polypeptide monomers.
[0474] In some embodiments, the efficiency of disassembly can be determined by quantifying the percent of total protein that is a monomer peak as measured by MADLS or SEC. In some embodiments, capsids are treated with a disassembly buffer, and optionally centrifuged to get precipitate (e.g., aggregates) and / or remaining capsids out of solution. The loss of capsid structures confirmed by electron microscopy (e.g., TEM) and / or dynamic light scattering (e.g., MADLS). The efficiency of disassembly can be determined by MADLS (e.g., via a capsid peak and a monomer peak) or size exclusion chromatography (SEC). The amount of protein that remains in solution can be quantified. The amount of monomer in solution can be quantified. Disassembly efficiency can be calculated. For example, a recovery of 90ng of protein in solution from lOOng of starting capsid material indicates a disassembly efficiency of 90%.
[0475] In some embodiments, measuring capsid reassembly efficiency comprises quantifying an amount of disassembled protein in solution that reassembles into a capsid state. In some embodiments, an amount of input protein in a non-capsid state in solution is quantified (e.g., after treating with a disassembly buffer, optionally centrifugation to remove precipitate and / or capsids, and confirming a lack of capsid structures via TEM and / or MADLS). In some embodiments, disassembled protein is treated with a reassembly buffer disclosed herein, e.g., that comprises a physiological buffer, low salt, high salt, an acidic buffer plus divalent cation, does not contain a chaotropic agent, does not contain a reducing agent, and / or another assembly or reassembly buffer disclosed herein. Aggregates can be removed by centrifugation, e g., at 15,000 x g for 5 min at +4°C. The efficiency of reassembly can be determined by comparing the amounts of endo-Gag polypeptides in the monomer and capsid peaks as assayed by MADLS or SEC. Capsid formation can be confirmed by TEM and / or MADLS. In some embodiments, size exclusion chromatography performed to quantify the amount of protein in a capsid state, and this can be compared to an amount of input protein. Reassembly efficiency is calculated. For example, 8 Ing of protein detected in a capsid state by size exclusion chromatography after starting with 90 ng of protein in a non-capsid state in solution indicates a 90% reassembly efficiency. In some embodiments, reassembly efficiency is determined by comparing the amountof endo-Gag polypeptides in reassembled capsids without further purification to the input amount of endo-Gag monomer polypeptides. In some embodiments, reassembly efficiency is determined by comparing the amount of endo-Gag polypeptides in reassembled capsids with further purification (e.g., via SEC or IEC) to the input amount of endo-Gag monomer polypeptides
[0476] In some embodiments, MADLS is used to determine particle concentration and verify whether it is within the expected range for the amount of protein in solution. For example, for PNMA2, which has a molecular weight of 41509.37, the expected number of particles can be calculated from the amount of PNMA2 polypeptide based on an assumption of 60 PNMA2 monomers per capsid: 1 g Capsids = 2.42 x 1017Capsids. In another embodiment, based on the endo Gag polypeptide molecular weight, the expected number of particles can be calculated from the amount of endo Gag polypeptide based on an assumption of x (e.g., 60) monomers per capsid.
[0477] In some embodiments, an observed concentration of a capsid disclosed herein is within about ±5%, within about ±10%, within about ±20%, within about ±30%, within about ±40%, within about ±50%, within about ±60%, within about ±70%, within about ±80%, within about ±90%, within about ±2-fold, within about ±5 -fold, or within about ±10-foldof an expected particle concentration.
[0478] In some embodiments, further purification is performed (e.g., via SEC or ionexchange chromatography) if capsid reassembly is less than 5%, less than 10%, less than 20%, less than 30%, less than 40%, less than 50%, less than 60%, less than 70%, less than less than 80%, less than 85%, less than 86%, less than 87%, less than 88%, less than 89%, less than 90%, less than 91%, less than 92%, less than 93%, less than 94%, less than 95%, less than 96%, less than 97%, less than 98%, less than 99%, or less than 99.5%efficient.
[0479] In some embodiments, a RTL or PNMA family or endo-Gag polypeptide disclosed herein exhibits favorable assembly properties. In some embodiments, upon isolation, a RTL or PNMA or endo-Gag polypeptide of the disclosure (e.g., PEG10, RTL 10, PNMA2, PNMA5, or another PNMA polypeptide or another RTL or endo-Gag polypeptide) assembles to form capsids with an efficiency of at least about 0.1%, at least about 0.5%, at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, at least about 1 1%, at leastabout 12%, at least about 13%, at least about 14%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 97%, or at least about 99%. In some embodiments, at least about 0. 1%, at least about 0.5%, at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, at least about 11%, at least about 12%, at least about 13%, at least about 14%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 97%, or at least about 99% of the isolated RTL or PNMA or endo-Gag polypeptide assembles to form the capsid.
[0480] In some embodiments, an RTL or PNMA or endo-Gag polypeptide disclosed herein exhibits favorable disassembly properties. In some embodiments, after incubation in a disassembly buffer, RTL or PNMA or endo-Gag capsids (e.g., comprising PEG10, RTL10, PNMA2, PNMA5, or another RTL, endo-Gag, or PNMA polypeptide) disassemble (e.g., to monomers, or smaller subunits than the capsids) with an efficiency of at least about 0.1%, at least about 0.5%, at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, at least about 11%, at least about 12%, at least about 13%, at least about 14%, at least about ...
Claims
CLAIMS1. A capsid comprising:(a) an endogenous Gag polypeptide,(b) one or more effector-polypeptide conjugates, where each effector-polypeptide conjugate independently comprises an endogenous Gag polypeptide and an effector moiety, and(c) a heterologous cargo.
2. A capsid comprising:(a) an endogenous Gag polypeptide, and(b) one or more effector-polypeptide conjugates, where each effector-polypeptide conjugate independently comprises an endogenous Gag polypeptide and an effector moiety, and wherein the capsid is loaded with or encapsulates a heterologous cargo.
3. A cap si d compri si ng :(a) an endogenous Gag polypeptide, and(b) one or more effector-polypeptide conjugates, where each effector-polypeptide conjugate independently comprises an endogenous Gag polypeptide and an effector moiety, and wherein a heterologous cargo is associated with the capsid.
4. The capsid of any one of the preceding claims, wherein the capsid is a protein nanoparticle.
5. The capsid of any one of the preceding claims, wherein the effector moiety is conjugated, directly or indirectly, through the N-terminus of the endogenous Gag polypeptide in the effector- polypeptide conjugate.
6. The capsid of any one of the preceding claims, wherein the effector-polypeptide conjugate further comprises (i) a Spytag conjugated to endogenous Gag polypeptide and (ii) a Spycatcher conjugated to the Spytag and to the effector moiety.
7. The capsid of any one of claims 14, 16, or 17, wherein the Spycatcher is conjugated through a 4 to 12 amino acid length linker to the endogenous polypeptide.
8. The capsid of any one of the preceding claims, wherein the effector-polypeptide conjugate further comprises a thiol maleimide linker conjugated to the endogenous polypeptide and to the effector moiety.
9. The capsid of any one of the preceding claims, wherein the one or more effector- polypeptide conjugates are selected from (i) a targeted conjugate comprising an endogenous Gag polypeptide and a targeting moiety, (ii) an endosomal escape conjugate comprising an endogenous Gag polypeptide and an endosomal escape moiety, or (iii) any combination of any of the foregoing.
10. The capsid of any one of the preceding claims, wherein the endosomal escape moiety is conjugated, directly or indirectly, through the C-terminus of the endogenous Gag polypeptide.
11. The capsid of any one of the preceding claims, wherein the endosomal escape moiety is conjugated through a cleavable linker to the remainder of the effector-polypeptide conjugate.
12. The capsid of any one of the preceding claims, wherein the targeting moiety enhances cell-specific uptake.
13. The capsid of any one of the preceding claims, wherein the targeting moiety enhances nuclear localization.
14. The capsid of any one of the preceding claims, wherein the targeting moiety is a targeting polypeptide.
15. The capsid of claim 5, wherein the targeting polypeptide comprises one or more of LDL- receptors, apolipoprotein mimetic peptide, PLA2, or any combination of any of the foregoing.
16. The capsid of claim 6, wherein the targeting polypeptide is an LDL receptor selected from LRKLRK, LRKRLLRD, and LKAYKS.
17. The capsid of any one of claims 5-12, wherein the endosomal escape conjugate comprises a first linker between the Gag polypeptide and the endosomal escape moiety.
18. The capsid of any one of the preceding claims, wherein the targeted conjugate comprises a first linker between the Gag polypeptide and the targeting moiety.
19. The capsid of claim 8, wherein the first linker comprises 5xGS or EAAK.
20. The capsid of any one of the preceding claims, wherein the targeting moiety is an antibody, fragment antigen-binding (Fab) protein, or single-chain variable fragment.
21. The capsid of claim 21, wherein the targeting moiety is a monoclonal antibody.
22. The capsid of claim 21, wherein the targeting moiety is a CD71 Fab protein.
23. The capsid of any one of the preceding claims, wherein the endogenous Gag polypeptide is a native endogenous Gag polypeptide.
24. The capsid of any one of the preceding claims, wherein the endogenous Gag polypeptide is an engineered endogenous Gag polypeptide.
25. The capsid of any one of the preceding claims, wherein the endogenous Gag polypeptide is an Arc polypeptide.
26. The capsid of any one of the preceding claims, wherein the endogenous Gag polypeptide is a Drosophila Arc (dArc) polypeptide.
27. The capsid of any one of the preceding claims, wherein the endogenous Gag polypeptide comprises a sequence modification, wherein the sequence modification comprises a cargo binding domain, nucleic acid binding domain, zinc finger domain, sub-cellular localization signal, a nuclear localization signal, or an antibody or antigen-binding fragment thereof.
28. The capsid of any one of the preceding claims, wherein the weight ratio of (a) an endogenous Gag polypeptide to (b) effector-polypeptide conjugate ranges from about 98:2 to about 50:50 (such as about 98:02 to about 80:20, or about 95:5 to about 85:15).
29. The capsid of any one of the preceding claims, wherein the cargo (a) comprises a nucleic acid (such as RNA or DNA), (b) comprises or encodes a gene editing system or a component thereof (e.g., a CRISPR / Cas system or a component thereof), (c) polypeptide, (d) a therapeutic agent, (e) an antibody or antigen-binding fragment thereof, a peptidomimetic, a nucleotidomimetic, a drug, a diagnostic tool, an imaging tool, a small molecule, or a combination thereof.
30. The capsid of any one of the preceding claims, wherein the heterologous cargo is a nucleic acid.
31. A method of preparing a capsid of any one of the preceding claims, the method comprising the step of mixing:(a) an endogenous Gag polypeptide,(b) one or more effector-polypeptide conjugates, where each effector-polypeptide conjugate independently comprises an endogenous Gag polypeptide and an effector moiety, and(c) a heterologous cargo.
32. A method of determining the encapsulation efficiency of a polynucleotide in a delivery system comprising the steps of: a) incubating a sample containing both free polynucleotide and encapsulated polynucleotide with a probe; b) measuring the free polynucleotide with the probe; c) treating the sample in (b) with a particle disrupting agent; d) incubating the sample in (c) with the probe; f) measuring the encapsulated polynucleotide with the probe; wherein the packaging efficiency is calculated as 1 -(probe signal of free polynucleotide / probe signal of total polynucleotide).
33. The method of claim 32, wherein the probe is a fluorescence-related amplificant or a luminescence-related amplificant.
34. The method of claim 33, wherein the fluorescence-related amplificant is a fluorescence resonance energy transfer (FRET)-based molecular beacon probe.
35. The method of claim 32, wherein the polynucleotide comprises a single stranded region.
36. The method of claim any one of claims 32-35, wherein the polynucleotide is an antisense oligonucleotide, microRNA, RNA activating oligonucleotide, mRNA, circular mRNA, or a gene editing system.
37. The method of any one of claims 32-36, wherein the probe comprises a stretch of at least 25, at least 30, at least 35 or at least 40 thymine (T) nucleotides.
38. The method of any one of claims 32-37, wherein the probe comprises a targeting loop, a stem region, and fluorophore-quencher pairs at the 3 'and 5' ends.
39. The method of any one of claims 32-38, wherein the probe is targeting the Poly(A) tail of the polynucleotide.
40. The method of any one of claims 37-39, wherein the binding of the stretch of T nucleotides loop to the polynucleotide-PolyA tail unwinds the molecular beacon stem and separates the fluorophore and quencher, and the fluorescence signal is measured by a plate reader.
41. The method of any one of claims 32-40, wherein the probe comprises a_plate reader with a fluorescence or luminescence detector.
42. The method of any one of claims 32-41, wherein the delivery system comprises a nanoparticle.
43. The method of any one of claims 32-42, wherein the delivery system is a protein nanoparticle (PNP), lipid nanoparticle (LNP), lipid protein nanoparticle (LPNP), or an adeno- associated virus (AAV).
44. The method of any one of claims 32-43, wherein the delivery system is a protein nanoparticle (PNP).
45. The method of any one of claims 32-44, wherein the particle disrupting agent is protease K when the delivery system is a PNP or AAV.
46. The method of any one of claims 32-44, wherein the particle disrupting agent is a detergent.
47. A kit for determining he encapsulation efficiency of a polynucleotide in a delivery system comprising: a) a delivery system, wherein the delivery system comprises free polynucleotide andencapsulated polynucleotide; b) at least one probe; and c) a particle disrupting agent.
48. The kit of claim 47, wherein the probe is a fluorescence-related amplificant or a luminescence-related amplificant.
49. The kit of claim 48, wherein the fluorescence-related amplificant is a fluorescence resonance energy transfer (FRET)-based molecular beacon probe.
50. The kit of any one of claims 47-49, wherein the delivery system comprises a nanoparticle.
51. The kit of any one of claims 47-50, wherein the delivery system is a protein nanoparticle (PNP), lipid nanoparticle (LNP), lipid protein nanoparticle (LPNP), or an adeno-associated virus (AAV).
52. The kit of any one of claims 47-51, wherein the delivery system is a protein nanoparticle (PNP).
53. The kit of any one of claims 47-52, wherein the particle disrupting agent is protease K when the delivery system is a PNP.
54. The kit of any one of claims 47-53, further comprising a negative and positive control.
55. The kit of any one of claims 47-54, further comprising written instructions for using said kit.
Citation Information
Patent Citations
Method of using neutrilized DNA (N-DNA) as surface probe for high throughput detection platform
US9481905B2
Endogenous gag-based and PNMA family capsids and uses thereof
WO2024026295A1
Compositions and methods for delivering cargo to a target cell
WO2021055855A1
PNMA2-based capsids and uses thereof
WO2022164942A1
US202363596047P