Iap LTR retrotransposon compositions and methods
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2026-03-12
AI Technical Summary
Existing methods for integrating nucleic acid sequences into a genome lack site specificity and efficiency, particularly for longer sequences, and often require multiple steps or specialized proteins like CRISPR/Cas9 or Cre/loxP.
A template RNA comprising long terminal repeats (LTRs) flanking a heterologous object sequence, combined with a structural polypeptide domain and reverse transcriptase, is introduced into a cell to generate template DNA for integration into the genome using an integrase, or provided as extrachromosomal DNA without integration.
This approach enables efficient and specific integration of therapeutic DNA into host genomes, offering transient expression and reduced genomic insertion, while avoiding complex multi-step processes.
Smart Images

Figure US2025032770_12032026_PF_FP_ABST
Abstract
Description
Attorney Docket No.: 2017469-0044 IAP LTR RETROTRANSPOSON COMPOSITIONS AND METHODS CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to United States Provisional Application No.63 / 657,004, filed June 6, 2024, the entirety of which is incorporated herein by reference. BACKGROUND
[0002] Integration of a nucleic acid of interest into a genome occurs at low frequency and with little site specificity, in the absence of a specialized protein to promote the insertion event. Some existing approaches, like CRISPR / Cas9, are more suited for small edits and are less effective at integrating longer sequences. Other existing approaches, like Cre / loxP, require a first step of inserting a loxP site into the genome and then a second step of inserting a sequence of interest into the loxP site. There is a need in the art for improved proteins for inserting sequences of interest into a genome. SUMMARY OF THE INVENTION
[0003] This disclosure relates to novel compositions, systems, and methods for altering a genome at one or more locations in a host cell, tissue, or subject, in vivo or in vitro. In particular, the invention features compositions, systems, and methods for the introduction of exogenous genetic elements into a host genome. The systems described herein typically include a template RNA comprising a pair of long terminal repeats (LTRs) flanking a heterologous object sequence (e.g., encoding a therapeutic effector), which can be introduced into a target cell with a structural polypeptide domain and a reverse transcriptase polypeptide domain, or nucleic acid molecules encoding same. Inside the cell, the template RNA and reverse transcriptase polypeptide domain can be enclosed within a proteinaceous exterior (e.g., a capsid), e.g., to form a virus-like particle (VLP). The reverse transcriptase polypeptide domain can then generate a template DNA from the template RNA. The resultant template DNA can then be integrated into the genome of the cell, e.g., by an integrase from a retrovirus or a retrotransposon, e.g., an LTR retrotransposon. Additionally described here are integration-deficient systems for providing an extrachromosomal DNA molecule to a host cell that does not undergo genomic integration. Thus, this disclosure provides systems capable of producing therapeutic DNA in a host cell, e.g., DNA encoding a therapeutic protein, by reverse transcription of an RNA template comprising LTRs, wherein the therapeutic DNA is optionally integrated into the host genome.
[0004] Features of the compositions or methods can include one or more of the following enumerated embodiments. Page 1 of 92 12806854v1Attorney Docket No.: 2017469-0044 1. A template RNA comprising: a 5’ long terminal repeat (5’ LTR) comprising a sequence according to SEQ ID NO: 5, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein the 5’ LTR comprises an A at position 1 of SEQ ID NO: 5 and a G at position 2 of SEQ ID NO: 5; a 3’ long terminal repeat (3’ LTR) comprising a sequence according to SEQ ID NO: 6, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto wherein the 3’ LTR comprises an A at position 231 of SEQ ID NO: 6 and a G at position 232 of SEQ ID NO: 6; a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and a primer binding site (PBS), wherein optionally the PBS is between the 5’ LTR and the heterologous object sequence. 2. The template RNA of embodiment 1, wherein the 5’ LTR has a length of less than 354, 300, 200, 150, or 125 nucleotides. 3. The template RNA of embodiment 1, wherein the 5’ LTR has a length of 354 nucleotides. 4. The template RNA of embodiment 1 or 2, wherein the 5’ LTR has a length of 124 nucleotides. 5. The template RNA of any one of the preceding embodiments, wherein the 3’ LTR has a length of less than 354, 340, 330, 320, 310, 300, 299, or 298 nucleotides. 6. The template RNA of any one of embodiments 1-4, wherein the 3’ LTR has a length of 354 nucleotides. 7. The template RNA of any one of embodiments 1-5, wherein the 3’ LTR has a length of 297 nucleotides. 8. The template RNA of any one of embodiments 1, 3, or 5-7, wherein the 5’ UTR has a sequence according to SEQ ID NO: 4, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. Page 2 of 92 12806854v1Attorney Docket No.: 2017469-0044 9. The template RNA of any one of embodiments 1-4 or 6, wherein the 3’ UTR has a sequence according to SEQ ID NO: 4, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. 10. A template RNA comprising: a 5’ long terminal repeat (5’ LTR) comprising a sequence according to SEQ ID NO: 5 or 2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; a 3’ long terminal repeat (3’ LTR) comprising a sequence according to SEQ ID NO: 6 or 3, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and a primer binding site (PBS), wherein optionally the PBS is between the 5’ LTR and the heterologous object sequence, wherein one or both of: the 5’ LTR has a length of less than 354, 300, 200, 150, or 125 nucleotides; or the 3’ LTR has a length of less than 354, 340, 330, 320, 310, 300, 299, or 298 nucleotides. 11. The template RNA of embodiment 10, wherein the 5’ LTR comprises an A at position 1 of SEQ ID NO: 5 and a G at position 2 of SEQ ID NO: 5. 12. The template RNA of embodiment 10 or 11, wherein the 3’ LTR comprises an A at position 231 of SEQ ID NO: 6 and a G at position 232 of SEQ ID NO: 6. 13. The template RNA of any one of embodiments 10-12, wherein the 5’ LTR has a length of 124 nucleotides. 14. The template RNA of any one of embodiments 10-13, wherein the 3’ LTR has a length of 297 nucleotides. 15. A template RNA comprising (e.g., in a 5’ to 3’ direction): a) a 5’ long terminal repeat (5’ LTR) having a sequence according to SEQ ID NO: 2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; b) a primer binding site (PBS); Page 3 of 92 12806854v1Attorney Docket No.: 2017469-0044 c) the nucleic acid sequence between position 2 and position 4, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; d) the nucleic acid sequence between position 12 and position 13, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; a polypurine tract (PPT), wherein optionally the PPT is immediately adjacent to the 3’ LTR; a 3’ long terminal repeat (3’ LTR) having a sequence according to SEQ ID NO: 3 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR. 16. The template RNA of embodiment 15, wherein (c) comprises a packaging / psi signal. 17. The template RNA of embodiment 15 or 16, wherein (c)comprises the gag splice donor site. 18. The template RNA of any one of embodiments 15-17, wherein (d) comprises a central PPT (cPPT). 19. A template RNA comprising: a 5’ long terminal repeat (5’ LTR) having a sequence according to SEQ ID NO: 2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; a 3’ long terminal repeat (3’ LTR) having a sequence according to SEQ ID NO: 3 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and a primer binding site (PBS); wherein the template RNA does not comprise one or more of (e.g., does not comprise 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or all of): the nucleic acid sequence between position 1 and position 2, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; Page 4 of 92 12806854v1Attorney Docket No.: 2017469-0044 the nucleic acid sequence between position 4 and position 5, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 5 and position 6, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 6 and position 7, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 7 and position 8, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 8 and position 9, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 9 and position 10, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 10 and position 11, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 11 and position 12, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 12 and position 13, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 13 and position 14, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and / or the nucleic acid sequence between position 14 and position 15, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. Page 5 of 92 12806854v1Attorney Docket No.: 2017469-0044 20. The template RNA of any one of embodiments 15-19, which comprises a premature stop codon in gag, pro, pol, or rig. 21. A template RNA comprising: an IAP 5’ long terminal repeat (IAP 5’ LTR); an IAP 3’ long terminal repeat (IAP 3’ LTR); a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and a primer binding site (PBS), wherein optionally the PBS is between the IAP 5’ LTR and the heterologous object sequence; wherein one or both of: the IAP 5’ LTR comprises a sequence according to SEQ ID NO: 8, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; or the IAP 3’ LTR comprises a sequence according to SEQ ID NO: 9 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. 22. A template RNA comprising: an IAP 5’ long terminal repeat (IAP 5’ LTR); an IAP 3’ long terminal repeat (IAP 3’ LTR); a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and a primer binding site (PBS), wherein optionally the PBS is between the IAP 5’ LTR and the heterologous object sequence; wherein one or both of: the IAP 5’ LTR consists of: a first fragment of a full-length IAP 5’ LTR according to SEQ ID NO: 1, wherein the first fragment consists of nucleotides 231-255 of SEQ ID NO: 1 and optionally up to 2, 5, or 10 additional nucleotides adjacent to either side of said nucleotides of SEQ ID NO: 1, or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto; and a second fragment of a full-length IAP 5’ LTR according to SEQ ID NO: 1, wherein the second fragment consists of nucleotides 315-354 of SEQ ID NO: 1 and optionally up to 2, 5, or 10 additional nucleotides adjacent to either side of said nucleotides of SEQ ID NO: 1, or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto; or the IAP 3’ LTR consists of: Page 6 of 92 12806854v1Attorney Docket No.: 2017469-0044 a first fragment of a full-length IAP 3’ LTR according to SEQ ID NO: 1, wherein the first fragment consists of nucleotides 1-40 of SEQ ID NO: 1 and optionally up to 2, 5, or 10 additional nucleotides adjacent to either side of said nucleotides of SEQ ID NO: 1, or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto; and a second fragment of a full-length IAP 3’ LTR according to SEQ ID NO: 1, wherein the second fragment consists of nucleotides 231-255 of SEQ ID NO: 1 and optionally up to 2, 5, or 10 additional nucleotides adjacent to either side of said nucleotides of SEQ ID NO: 1, or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto. 23. A template RNA comprising: an IAP 5’ long terminal repeat (IAP 5’ LTR); an IAP 3’ long terminal repeat (IAP 3’ LTR); a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and a primer binding site (PBS), wherein optionally the PBS is between the IAP 5’ LTR and the heterologous object sequence; wherein one or both of: the IAP 5’ LTR comprises a mutation that reduces promoter function of the IAP 5’ LTR compared to an IAP 5’ LTR having a sequence of SEQ ID NO: 1; or the IAP 3’ LTR comprises a mutation that reduces promoter function of the IAP 3’ LTR compared to an IAP 3’ LTR having a sequence of SEQ ID NO: 1. 24. The template RNA of embodiment 23, wherein the mutation that reduces promoter function is a deletion, e.g., a deletion of one or more of (e.g. two or all of): the TATA-box; the CCAA-box; Inr motif; or the polyadenylation signal. 25. The template RNA of any one of embodiments 21-24, which is transcribed at a lower level than a control template RNA having the same sequence as the template RNA except that its 3’ LTR has a sequence of SEQ ID NO: 1. Page 7 of 92 12806854v1Attorney Docket No.: 2017469-0044 26. The template RNA of any one of embodiments 21-24, which is transcribed at a lower level than a control template RNA having the same sequence as the template RNA except that its 5’ LTR has a sequence of SEQ ID NO: 1. 27. The template RNA of any one of embodiments 21-24, which is active for integration into a target nucleic acid in an assay according to Example 5. 28. A DNA molecule comprising a sequence according to SEQ ID NO: 4, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein nucleotide 231 of SEQ ID NO: 4 is A and nucleotide 232 of SEQ ID NO: 4 is G. 29. A DNA molecule comprising a sequence according to SEQ ID NO: 7, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. 30. A DNA molecule having a sequence that is the reverse complement of the sequence of embodiment 28 or 29. 31. The DNA molecule of any one of embodiments 28-30, which is single stranded or double stranded. 32. An IAP retrotransposase polypeptide comprising: an amino acid sequence of SEQ ID NO: 12, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein one, two, or all of: position D653 of SEQ ID NO: 12 is other than D, e.g., is valine; position D710 of SEQ ID NO: 12 is other than D, e.g., is valine; or position E746 of SEQ ID NO: 12 is other than E, e.g., is valine. 33. the IAP retrotransposase polypeptide of embodiment 32, wherein position D653 of SEQ ID NO: 12 is valine. 34. the IAP retrotransposase polypeptide of embodiment 32 or 33, wherein position D710 of SEQ ID NO: 12 is valine. Page 8 of 92 12806854v1Attorney Docket No.: 2017469-0044 35. the IAP retrotransposase polypeptide of any one of embodiments 32-34, wherein position E746 of SEQ ID NO: 12 is valine. 36. A nucleic acid molecule (e.g., a DNA molecule) encoding the IAP retrotransposase polypeptide any one of embodiments 32-35. 37. The nucleic acid molecule of embodiment 36, which comprises a T at position 1958 of SEQ ID NO: 10. 38. The IAP retrotransposase polypeptide of any one of embodiments 32-35, wherein: position D710 of SEQ ID NO: 12 is aspartate (D), and position E746 of SEQ ID NO: 12 is glutamate (E). 39. The IAP retrotransposase polypeptide of any one of embodiments 32-38, which has reduced integrase activity compared to a polypeptide of SEQ ID NO: 12, e.g., reduced by about 20%, 40%, 60%, 80%, 90%, or 95%, e.g., in an assay of Example 6. 40. The IAP retrotransposase polypeptide of any one of embodiments 32-39, which leads to has transient expression of a heterologous object sequence, e.g., wherein expression is detectable for less than 9 days, e.g., in an assay of Example 6. 41. A template RNA comprising (e.g., from 5’ to 3’): a) an IAP 5’ long terminal repeat (IAP 5’ LTR); b) a primer binding site (PBS); c) a heterologous object sequence encoding a therapeutic effector; d) a central PPT (cPPT); and e) an IAP 3’ long terminal repeat (IAP 3’ LTR); wherein, between d) and e), the template RNA lacks a canonical PPT or comprises a mutation to a canonical PPT (“mutant PPT”) that reduces initiation of second strand synthesis primed by the mutant PPT. 42. The template RNA of embodiment 41, wherein the cPPT is situated upstream of the heterologous object sequence, downstream of the heterologous object sequence, or overlapping with the heterologous object sequence. Page 9 of 92 12806854v1Attorney Docket No.: 2017469-0044 43. The template RNA of embodiment 41 or 42, wherein, upon reverse transcription, the template RNA yields a greater proportion of episomes to linear dsDNA, compared to the proportion of episomes to linear dsDNA produced using a reference template RNA which comprises a canonical PPT between its cPPT and IAP 3’ LTR. 44. The template RNA of embodiment 41 or 42, wherein, upon reverse transcription, the template RNA yields fewer insertions into a host cell genome, compared to the number of insertions into a host cell genome produced using a reference template RNA which comprises a canonical PPT between its cPPT and IAP 3’ LTR. 45. A DNA molecule encoding the template RNA of any one of embodiments 1-27 or 41-44. 46. A system comprising: a template RNA of any one of embodiments 1-27 or 41-44, or a DNA encoding the template RNA, and an IAP retrotransposase polypeptide, or a nucleic acid encoding the IAP retrotransposase polypeptide. 47. The system of embodiment 46, wherein the IAP retrotransposase polypeptide comprises an amino acid sequence of SEQ ID NO: 12, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. 48. A system comprising: a template RNA, or a DNA encoding the template RNA; and an IAP retrotransposase polypeptide according to any one of embodiments 32-35 or 38-40, or a nucleic acid encoding the IAP retrotransposase polypeptide. 49. The system of embodiment 48, wherein the template RNA is a template RNA of any one of embodiments 1-27 or 41-44. 50. The system of any one of the preceding system embodiments, wherein the system is substantially free of virus. Page 10 of 92 12806854v1Attorney Docket No.: 2017469-0044 51. The system of any one of the preceding system embodiments, wherein the system is substantially free of cells. 52. A nucleic acid encoding the IAP retrotransposase polypeptide of any one of embodiments 32-35 or 38-40. 53. The nucleic acid of embodiment 52, which is an mRNA. 54. The nucleic acid of embodiment 52 or 53, which comprises one or more non-canonical or modified ribonucleotides. 55. The system or the template RNA of any one of the preceding embodiments, wherein the template RNA comprises one or more non-canonical or modified ribonucleotides. 56. A pharmaceutical composition comprising the template RNA or the system of any one of the preceding embodiments. 57. The system or IAP retrotransposase polypeptide of any one of the preceding embodiments, wherein the system, nucleic acid molecule, polypeptide, and / or DNA encoding the same, is formulated as a lipid nanoparticle (LNP). 58. A cell comprising the system of any one of embodiments 46-51, 55, or 57. 59. A cell comprising the template RNA of any one of embodiments 1-27 or 41-44. 60. A cell comprising the IAP retrotransposase polypeptide of any one of embodiments 32-35 or 38- 40. 61. A method of delivering a heterologous object sequence to a target cell, comprising introducing into the target cell (e.g., contacting the target cell with) a system of any one of embodiments 46-51, 55, or 57, and incubating the target cell under conditions suitable for production of the template DNA. Page 11 of 92 12806854v1Attorney Docket No.: 2017469-0044 BRIEF DESCRIPTION OF THE DRAWINGS
[0005] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0006] FIG.1 is a schematic diagram of an exemplary process for production of either non- integrating circular reverse transcription products or linear integration-competent reverse transcription products from a template RNA.
[0007] FIG.2A shows the retrotransposition efficiency of trans (non-autonomous) templates deletion constructs having deletion of one or more different regions without an added driver (blue dots) or the deletion constructs in combination with a ∆PBS driver IAP (pink dots). Each pair of blue dots and pink dots residing between two vertical dash lines indicates the retrotransposition activity of a construct comprising deletion of the region between the two vertical dash lines.
[0008] FIG.2B is a schematic providing numerical positions for deletion regions of the trans (non- autonomous) templates deletion constructs having deletion of one or more different regions between two vertical dash lines.
[0009] FIG.3A shows the schematic structure of various “LTR free” IAP drivers with sequence omissions, compared to WT full-length IAP driver.
[0010] FIG.3B shows the retrotransposition efficiency of LTR-free IAP drivers shown in FIG.3A.
[0011] FIG.4A shows the schematic structure of IAP retrotransposon elements comprising truncated LTRs. Each dash line represents a position on the WT full-length IAP LTR sequence.
[0012] FIG.4B shows the retrotransposition efficiency of IAP retrotransposon elements comprising truncated LTRs shown in FIG.4A.
[0013] FIG.5 shows the retrotransposition efficiency of IAP drivers comprising a D to V and / or E to V mutation in various positions of the DDE catalytic triad. Definitions
[0014] About, approximately: “About” or “approximately” as the terms are used herein applied to one or more values of interest, refer to a value that is similar to a stated reference value. In certain embodiments, the term “approximately” or “about” refers to a range of values that fall within 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less in either direction (greater than or less than) of the stated reference value unless otherwise stated or otherwise evident from the context (except where such number would exceed 100% of a possible value).
[0015] Domain: The term “domain” as used herein refers to a structure of a biomolecule that contributes to a specified function of the biomolecule. A domain may comprise a contiguous region (e.g., Page 12 of 92 12806854v1Attorney Docket No.: 2017469-0044 a contiguous sequence) or distinct, non-contiguous regions (e.g., non-contiguous sequences) of a biomolecule. Examples of protein domains include, but are not limited to, a nuclear localization sequence, a recombinase domain, a retroviral (e.g., endogenous retroviral) structural polypeptide domain, a retroviral (e.g., endogenous retroviral) reverse transcriptase polypeptide domain, a retrotransposon structural polypeptide domain, a retrotransposon reverse transcriptase polypeptide domain, a DNA recognition domain (e.g., that binds to or is capable of binding to a recognition site, e.g. as described herein), a recombinase N-terminal domain (also called a catalytic domain), a C-terminal zinc ribbon domain. In some embodiments the zinc ribbon domain further comprises a coiled-coiled motif. In some embodiments, the recombinase domain and the zinc ribbon domain are collectively referred to as the C- terminal domain. In some embodiments the N-terminal domain is linked to the C-terminal domain by an αE linker or helix. In some embodiments the N-terminal domain is between 50 and 250 amino acids, or 100-200 amino acids, or 130 - 170 amino acids, e.g., about 150 amino acids. In some embodiments the C- terminal domain is 200-800 amino acids, or 300-500 amino acids. In some embodiments the recombinase domain is between 50 and 150 amino acids. In some embodiments the zinc ribbon domain is between 30 and 100 amino acids; an example of a domain of a nucleic acid is a regulatory domain, such as a transcription factor binding domain, a recognition sequence, an arm of a recognition sequence (e.g. a 5’ or 3’ arm), a core sequence, or an object sequence (e.g., a heterologous object sequence).
[0016] Exogenous: As used herein, the term exogenous, when used with reference to a biomolecule (such as a nucleic acid sequence or polypeptide) means that the biomolecule was introduced into a host genome, cell, or organism by the hand of man. For example, a nucleic acid that is as added into an existing genome, cell, tissue, or subject using recombinant DNA techniques or other methods is exogenous to the existing nucleic acid sequence, cell, tissue or subject.
[0017] Fragment: The term “fragment,” as used herein with respect to a nucleic acid, refers to a portion of a full-length nucleic acid, or a sequence having at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to said portion. In some embodiments, the fragment recapitulates or maintains one or more function of said portion. For clarity, a fragment of a full- length LTR is not adjacent to a sequence or sequences that together form the full length LTR. The fragment may comprise one or more mutations (e.g., substitutions) relative to the corresponding part of the full-length nucleic acid.
[0018] Heterologous: The term heterologous, when used to describe a first element in reference to a second element means that the first element and second element do not exist in nature disposed as described. For example, a heterologous polypeptide, nucleic acid molecule, construct or sequence refers to (a) a polypeptide, nucleic acid molecule or portion of a polypeptide or nucleic acid molecule sequence that is not native to a cell in which it is expressed, (b) a polypeptide or nucleic acid molecule or portion of Page 13 of 92 12806854v1Attorney Docket No.: 2017469-0044 a polypeptide or nucleic acid molecule that has been altered or mutated relative to its native state, or (c) a polypeptide or nucleic acid molecule with an altered expression as compared to the native expression levels under similar conditions. For example, a heterologous regulatory sequence (e.g., promoter, enhancer) may be used to regulate expression of a gene or a nucleic acid molecule in a way that is different than the gene or a nucleic acid molecule is normally expressed in nature. In another example, a heterologous domain of a polypeptide or nucleic acid sequence (e.g., a DNA binding domain of a polypeptide or nucleic acid encoding a DNA binding domain of a polypeptide) may be disposed relative to other domains or may be a different sequence or from a different source, relative to other domains or portions of a polypeptide or its encoding nucleic acid. In certain embodiments, a heterologous nucleic acid molecule may exist in a native host cell genome but may have an altered expression level or have a different sequence or both. In other embodiments, heterologous nucleic acid molecules may not be endogenous to a host cell or host genome but instead may have been introduced into a host cell by transformation (e.g., transfection, electroporation), wherein the added molecule may integrate into the host genome or can exist as extra-chromosomal genetic material either transiently (e.g., mRNA) or semi- stably for more than one generation (e.g., episomal viral vector, plasmid or other self-replicating vector). In some embodiments, a domain is heterologous relative to another domain, if the first domain is not naturally comprised in the same polypeptide as the other domain (e.g., a fusion between two domains of different proteins from the same organism).
[0019] Long Terminal Repeat: The term “long terminal repeat” (LTR), as used herein, refers to a nucleic acid sequence, which in a wild-type context are found in pairs (which may be identical or have sequence similarity) that flank a retrovirus or an LTR retrotransposon. The term “LTR” also encompasses variants and fragments of a wild-type LTR which are functional for integration of a region of the nucleic acid molecule comprising the LTR into a target DNA molecule in the presence of factors from the retrovirus or LTR retrotransposon. An LTR is typically located at or near one end (e.g., the 5’ end or the 3’ end) of a template DNA or RNA, e.g., as described herein. In some instances, an LTR participates in integration of a heterologous object sequence comprised in the template DNA or RNA into a target DNA molecule (e.g., a genomic DNA). In some instances, the LTR, or a fragment thereof, is integrated into the target DNA molecule. In some instances, the LTR is not integrated into the target DNA molecule. In some instances, a 5’ LTR of a template DNA or RNA (e.g., as described herein) has at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a 3’ LTR sequence of the template DNA or RNA. In some instances, an LTR of a system or composition described herein has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to an LTR sequence of a naturally occurring retrovirus (e.g., endogenous retrovirus) or LTR retrotransposon. In some instances, an LTR of a system or composition described herein has at least one modification (e.g., Page 14 of 92 12806854v1Attorney Docket No.: 2017469-0044 an addition, substitution, or deletion) relative to an LTR sequence of a naturally occurring retrovirus (e.g., endogenous retrovirus) or LTR retrotransposon. In some embodiments, an LTR has promoter and / or enhancer activity. In some embodiments the LTR has no promoter activity or reduced promoter activity.
[0020] IAP 5’ long terminal repeat: As used herein, the term “IAP 5’ long terminal repeat” (“IAP 5’ LTR”) refers to an LTR that: (1) has at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a sequence of SEQ ID NO: 1 or a fragment thereof, (2) supports reverse transcription, and (3) is situated at or near the 5’ end of an RNA.
[0021] IAP 3’ long terminal repeat: As used herein, the term “IAP 3’ long terminal repeat” (“IAP 3’ LTR”) refers to an LTR that: (1) has at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a sequence of SEQ ID NO: 1 or a fragment thereof, (2) supports reverse transcription, and (3) is situated at or near the 3’ end of an RNA.
[0022] Mutation or Mutated: The term “mutated” when applied to nucleic acid sequences means that nucleotides in a nucleic acid sequence may be inserted, deleted, or changed compared to a reference (e.g., native) nucleic acid sequence. A single alteration may be made at a locus (a point mutation) or multiple nucleotides may be inserted, deleted, or changed at a single locus. In addition, one or more alterations may be made at any number of loci within a nucleic acid sequence. A nucleic acid sequence may be mutated by any method known in the art. In some embodiments a mutation occurs naturally. In some embodiments a desired mutation can be produced by any suitable method.
[0023] Nucleic acid molecule: Nucleic acid molecule refers to both RNA and DNA molecules including, without limitation, cDNA, genomic DNA, and mRNA, and also includes synthetic nucleic acid molecules, such as those that are chemically synthesized or recombinantly produced, such as RNA templates, as described herein. The nucleic acid molecule can be double-stranded or single-stranded, circular, or linear. If single-stranded, the nucleic acid molecule can be the sense strand or the antisense strand. Unless otherwise indicated, and as an example for all sequences described herein under the general format “SEQ. ID NO:,” “nucleic acid comprising SEQ. ID NO:1” refers to a nucleic acid, at least a portion which has either (i) the sequence of SEQ. ID NO:1, or (ii) a sequence complementary to SEQ. ID NO:1. The choice between the two is dictated by the context in which SEQ. ID NO:1 is used. For instance, if the nucleic acid is used as a probe, the choice between the two is dictated by the requirement that the probe be complementary to the desired target. Nucleic acid sequences of the present disclosure may be modified chemically or biochemically or may contain non-natural or derivatized nucleotide bases, as will be readily appreciated by those of skill in the art. Such modifications include, for example, labels, methylation, substitution of one or more naturally occurring nucleotides with an analog, inter-nucleotide modifications such as uncharged linkages (for example, methyl phosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (for example, phosphorothioates, Page 15 of 92 12806854v1Attorney Docket No.: 2017469-0044 phosphorodithioates, etc.), pendant moieties, (for example, polypeptides), intercalators (for example, acridine, psoralen, etc.), chelators, alkylators, and modified linkages (for example, alpha anomeric nucleic acids, etc.). Also included are synthetic molecules that mimic polynucleotides in their ability to bind to a designated sequence via hydrogen bonding and other chemical interactions. Such molecules are known in the art and include, for example, those in which peptide linkages substitute for phosphate linkages in the backbone of a molecule. Other modifications can include, for example, analogs in which the ribose ring contains a bridging moiety or other structure such as modifications found in “locked” nucleic acids.
[0024] Gene expression unit: a gene expression unit is a nucleic acid sequence comprising at least one regulatory nucleic acid sequence operably linked to at least one effector sequence. A first nucleic acid sequence is operably linked with a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For instance, a promoter or enhancer is operably linked to a coding sequence if the promoter or enhancer affects the transcription or expression of the coding sequence. Operably linked DNA sequences may be contiguous or non- contiguous. Where necessary to join two protein-coding regions, operably linked sequences may be in the same reading frame.
[0025] Host: The terms host genome or host cell, as used herein, refer to a cell and / or its genome into which protein and / or genetic material has been introduced. It should be understood that such terms are intended to refer not only to the particular subject cell and / or genome, but to the progeny of such a cell and / or the genome of the progeny of such a cell. Because certain modifications may occur in succeeding generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term “host cell” as used herein. A host genome or host cell may be an isolated cell or cell line grown in culture, or genomic material isolated from such a cell or cell line or may be a host cell or host genome which composing living tissue or an organism. In some instances, a host cell may be an animal cell or a plant cell, e.g., as described herein. In certain instances, a host cell may be a bovine cell, horse cell, pig cell, goat cell, sheep cell, chicken cell, or turkey cell. In certain instances, a host cell may be a corn cell, soy cell, wheat cell, or rice cell.
[0026] Introducing: As used herein, the term “introducing”, in the context of introducing an agent into a call, refers to causing the agent to be comprised by the cell. For example, the cell may be contacted with the agent in a way that allows the agent to pass through the cell membrane to enter the cell. Alternatively, the agent can be introduced into the cell by causing the cell to produce the agent. For instance, an agent that is a polypeptide can be introduced into the cell by contacting the cell with a nucleic acid encoding the polypeptide, under conditions that the nucleic acid enters the cell and is translated to produce the polypeptide. Page 16 of 92 12806854v1Attorney Docket No.: 2017469-0044
[0027] Contacting: As used herein, the term “contacting”, in the context of contacting a cell with an agent, comprises placing the agent at a location that allows the agent to come into physical contact with the cell. Physical contact with the cell includes, e.g., binding to the cell surface or being internalized into the cell. In some embodiments, e.g., ex vivo, contacting a cell with an agent comprises introducing the agent into media, wherein the media is in contact with the cell. In some embodiments, e.g., in vivo, contacting a cell with an agent comprises administering the agent to a subject comprising the cell, under conditions that allow the agent to come into physical contact with the cell.
[0028] Object sequence: As used herein, the term object sequence refers to a nucleic acid segment that can be desirably inserted into a target nucleic acid molecule, e.g., by a recombinase polypeptide, e.g., as described herein. In some embodiments, a template RNA or template DNA comprises a DNA recognition sequence and an object sequence that is heterologous to the DNA recognition sequence and / or the remainder of the template RNA or template DNA, generally referred to herein as a “heterologous object sequence.” An object sequence may, in some instances, be heterologous relative to the nucleic acid molecule into which it is inserted (e.g., a target DNA molecule, e.g., as described herein). In some instances, an object sequence comprises a nucleic acid sequence encoding a gene (e.g., a eukaryotic gene, e.g., a mammalian gene, e.g., a human gene) or other cargo of interest (e.g., a sequence encoding a functional RNA, e.g., an siRNA or miRNA), e.g., as described herein. In certain instances, the gene encodes a polypeptide (e.g., a blood factor or enzyme). In some instances, an object sequence comprises one or more of a nucleic acid sequence encoding a selectable marker (e.g., an auxotrophic marker or an antibiotic marker), and / or a nucleic acid control element (e.g., a promoter, enhancer, silencer, or insulator).
[0029] Polypeptide driver: As used herein, the term “polypeptide driver” refers to a polypeptide or a plurality of polypeptides that can carry out retrotransposition of a compatible template RNA. In some embodiments, the polypeptide driver comprises a single polypeptide having multiple domains, for example, two or more domains chosen from GAG, PRO, and POL. In some embodiments, the polypeptide driver comprises a plurality of separate polypeptides, wherein each polypeptide in the plurality comprises one or more domains chosen from GAG, PRO, and POL. In some embodiments, the polypeptide driver is encoded by overlapping ORFs.
[0030] Structural polypeptide domain: As used herein, the term “structural polypeptide domain” refers to a polypeptide domain that can form part of a proteinaceous exterior (e.g., a viral capsid) encapsulating a nucleic acid (e.g., a template RNA, e.g., as described herein). Retroviral env is not a structural polypeptide domain, as the term is used herein. In some instances, a structural polypeptide domain is encoded by a viral gene (e.g., a retroviral gag gene). In some instances, a structural polypeptide domain comprises a capsid protein (e.g., a CA protein and / or an NC protein, e.g., encoded by Page 17 of 92 12806854v1Attorney Docket No.: 2017469-0044 a retroviral gag gene), or a functional fragment thereof. In some instances, a structural polypeptide domain comprises a matrix protein (e.g., a MA protein, e.g., encoded by a retroviral gag gene), or a functional fragment thereof. In some instances, a structural polypeptide domain comprises a domain encoded by a retroviral gag (e.g., an endogenous retroviral gag). In some embodiments, a structural polypeptide domain comprises one or more mutations (e.g., point mutations, additions, substitutions, or deletions) relative to the amino acid sequence of a corresponding wild-type protein (e.g., a wild-type retroviral gag, CA, NC, or MA protein). In some embodiments, a structural polypeptide domain is part of a polyprotein or a fusion protein. In some embodiments, a structural polypeptide domain is not part of a polyprotein or a fusion protein.
[0031] Reverse transcriptase domain: As used herein, the term “reverse transcriptase domain” refers to a polypeptide domain capable of producing complementary DNA from a template RNA (e.g., as described herein). In some instances, a reverse transcriptase domain comprises a viral (e.g., retroviral, e.g., endogenous retroviral) reverse transcriptase, or a functional fragment thereof. In some instances, a reverse transcriptase domain produces complementary DNA from a template RNA via a primer (e.g., a tRNA primer, e.g., a lysyl tRNA primer). In some instances, a reverse transcriptase domain produces a double stranded template DNA (e.g., as described herein) from the template RNA. In some instances, a reverse transcriptase domain is encoded by a viral (e.g., retroviral, e.g., endogenous retroviral) pol gene. In some instances, a reverse transcriptase domain is encoded by a pol gene that also encodes a viral (e.g., retroviral, e.g., endogenous retroviral) integrase (IN). In some instances, a reverse transcriptase domain is encoded by a pol gene that also encodes a viral (e.g., retroviral, e.g., endogenous retroviral) protease (PR) and / or dTUPase (DU). In some embodiments, a reverse transcriptase polypeptide domain comprises one or more mutations (e.g., point mutations, additions, substitutions, or deletions) relative to the amino acid sequence of a corresponding wild-type protein (e.g., a wild-type retroviral pol, IN, PR, or DU protein). In some embodiments, a reverse transcriptase domain is part of a polyprotein or a fusion protein. In some embodiments, a reverse transcriptase domain is not part of a polyprotein or a fusion protein. In some embodiments, the reverse transcriptase domain comprises RNaseH activity. In some embodiments, a functional reverse transcriptase comprises a single protein subunit, e.g., is monomeric. In some embodiments, a functional reverse transcriptase comprises at least two subunits, e.g., is dimeric. In some embodiments, the reverse transcriptase domain is less active (or inactive) in monomeric form compared to in dimeric form. In some embodiments, a dimeric reverse transcriptase comprises two identical subunits. In some embodiments, a dimeric reverse transcriptase comprises different subunits, e.g., a p51 and a p66 subunit. In some embodiments, a reverse transcriptase comprises at least three subunits, e.g., two p51 subunits and at least one p15 subunit. In some embodiments, a reverse transcriptase comprises an RNase Page 18 of 92 12806854v1Attorney Docket No.: 2017469-0044 H domain. In some embodiments, a reverse transcriptase comprises an inactivated RNase H domain. In some embodiments, a reverse transcriptase does not comprise an RNase H domain.
[0032] LTR retrotransposon: As used herein, the term “LTR retrotransposon” in the context of a domain (e.g., LTR retrotransposon structural polypeptide domain or LTR retrotransposon reverse transcriptase polypeptide domain) refers to a polypeptide domain having sequence similarity to a corresponding domain from a wild-type LTR retrotransposon, and at least one biological function (e.g., capsid formation or reverse transcription) in common with the corresponding domain. A wild-type LTR retrotransposon does not comprise an env gene. In some embodiments, an LTR retrotransposon may comprise a retrovirus (e.g., an endogenous retrovirus) engineered to lack a functional env gene.
[0033] Retroviral: As used herein, the term “retroviral” in the context of a domain (e.g., retroviral structural polypeptide domain or retroviral reverse transcriptase polypeptide domain) refers to a polypeptide domain having sequence similarity to a corresponding domain from a wild-type retrovirus (e.g., endogenous retrovirus) and at least one biological function (e.g., capsid formation or reverse transcription) in common with the corresponding domain. A wild-type retrovirus comprises an env gene. DETAILED DESCRIPTION
[0034] This disclosure relates to compositions, systems, and methods for targeting, editing, modifying, or manipulating a DNA sequence (e.g., inserting a heterologous object DNA sequence into a target site of a mammalian genome) at one or more locations in a DNA sequence in a cell, tissue, or subject, e.g., in vivo or in vitro. Generally, the systems and compositions include a template RNA comprising a pair of long terminal repeats (LTRs) flanking a heterologous object sequence (e.g., encoding a therapeutic effector). In some instances, the LTRs are derived from a retrovirus (e.g., an endogenous retrovirus). In some instances, the LTRs are derived from a retrotransposon (e.g., an LTR retrotransposon). The template RNA is typically introduced into a target cell with a structural polypeptide domain and a reverse transcriptase polypeptide domain, or nucleic acid molecules encoding the structural polypeptide domain and the reverse transcriptase polypeptide domain. In some instances, the structural polypeptide and / or reverse transcriptase polypeptide domain are derived from a retrovirus (e.g., an endogenous retrovirus). In some instances, the structural polypeptide and / or reverse transcriptase polypeptide domain are derived from a retrotransposon (e.g., an LTR retrotransposon). The template RNA and reverse transcriptase polypeptide domain can be enclosed within a proteinaceous exterior (e.g., a capsid) in the cell, e.g., to form a virus-like particle (VLP). Within the VLP, the reverse transcriptase polypeptide domain can generate a template DNA (e.g., a linear and / or double-stranded DNA) from the template RNA. The template DNA can then optionally be integrated into the genome of the cell, e.g., by an integrase from a retrovirus (e.g., an endogenous retrovirus) or a retrotransposon, e.g., an LTR Page 19 of 92 12806854v1Attorney Docket No.: 2017469-0044 retrotransposon. The heterologous object sequence may include, e.g., a coding sequence, a regulatory sequence, and / or a gene expression unit. LTR retrotransposon systems
[0035] Long terminal repeat (LTR) retrotransposons are a type of mobile genetic elements that are widespread in eukaryotic genomes. Naturally occurring LTR retrotransposons typically have a coding region flanked by direct (i.e., not inverted) long terminal repeats. The LTR typically includes a promoter whereby the coding region may be transcribed. The coding region typically codes for the Gag and Pol polyproteins. Gag is typically processed by protease to produce structural proteins matrix (MA), capsid (CA), and nucleocapsid (NC) proteins that form the virus-like particle (VLP), and inside of which reverse transcription of the LTR retrotransposon transcript takes place. Pol typically has protease, reverse transcriptase that copies the LTR retrotransposon transcript into cDNA, RNaseH, and integrase, which integrates the cDNA into the host genome. LTR retrotransposons also typically include a primer binding site (PBS) immediately downstream of the 5´LTR and a polypurine tract (PPT) immediately upstream of the 3´LTR.
[0036] In some embodiments, a system described herein results in integration of the heterologous object sequence into the genome of the target cell. In some embodiments, the system results in integration of the heterologous object sequence into a specific site within the genome of the target cell. In some embodiments, the system results in integration of the heterologous object sequence into a random site within the genome of the target cell. In some embodiments, the integration of the heterologous object sequence into the genome of the target cell results in one or more duplications at the integration site, e.g., duplications of 4-6 (e.g., 4, 5, or 6) nucleotides in length. LTR retrotransposon genome delivery systems
[0037] The present disclosure provides compositions, systems, and methods for integrating a heterologous object sequence (e.g., encoding a therapeutic effector) into the genome of a target cell. Generally, a template RNA is introduced into a cell (e.g., as an RNA molecule, or in the form of a DNA molecule (e.g., an episome) that is transcribed into RNA in the cell). The template RNA is then enclosed in a proteinaceous exterior (e.g., capsid) within the cell, thereby forming a virus-like particle (VLP) in the cell. The template RNA is then reverse-transcribed in the VLP to generate a template DNA, e.g., thereby forming a pre-integration complex (PIC) comprising the template DNA enclosed in the proteinaceous exterior. In some embodiments, the VLP is initially formed in the cytoplasm. In some embodiments, the VLP is initially localized to the endoplasmic reticulum. The VLP does not obtain an envelope. In some embodiments, reverse transcription of the template RNA occurs while the VLP is in the cytoplasm. In Page 20 of 92 12806854v1Attorney Docket No.: 2017469-0044 some embodiments, reverse transcription of the template RNA occurs while the VLP is in the endoplasmic reticulum or another organelle compartment. In some embodiments, reverse transcription of the template RNA occurs while the VLP is in the nucleus.
[0038] Once in the nucleus, the template DNA (or a portion thereof, e.g., the heterologous object sequence) may be integrated into the genome of the cell, e.g., by an integrase (e.g., a retrotransposon integrase or a retroviral integrase, e.g., a lentiviral integrase, e.g., an HIV integrase). In some embodiments, the template DNA is not integrated into the genome of the cell. In certain embodiments, the non-integrated template DNA is circularized, e.g., to form an episome comprising the heterologous object sequence. In some embodiments, the integrated heterologous object sequence may be flanked by one or more LTRs (e.g., the 5’ LTR and / or the 3’ LTR).
[0039] In some embodiments, provided herein is a system that comprises a template RNA as described herein, or a DNA encoding the template RNA, and an IAP retrotransposase polypeptide, or a nucleic acid encoding the IAP retrotransposase polypeptide.
[0040] In some embodiments, the IAP retrotransposase polypeptide comprises an amino acid sequence of SEQ ID NO: 12, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0041] In some embodiments, the system described herein is substantially free of virus. In some embodiments, the system described herein is substantially free of cells. In some embodiments, the system is in a cell.
[0042] In some embodiments, the system comprises a nucleic acid encoding the IAP retrotransposase polypeptide as described herein. In some embodiments, the nucleic acid is an mRNA. In some embodiments, the nucleic acid comprises one or more non-canonical or modified ribonucleotides.
[0043] In some embodiments, the system comprises a template RNA that comprises one or more non-canonical or modified ribonucleotides. Template RNA Component
[0044] In some embodiments, the template RNA comprises one or more (e.g., 1, 2, 3, 4, 5, or all 6) of the following (e.g., in order from 5’ to 3’): (i) a 5’ long terminal repeat (LTR), (ii) a primer binding site (PBS), (iii) a promoter, (iv) a heterologous object sequence (e.g., comprising an open reading frame), (v) a polypurine tract, and / or (vi) a 3’ LTR. In some embodiments, the PBS has a length of about 15, 16, 17, 18, 19, or 20 nucleotides (e.g., 18 nucleotides). In some embodiments, the PBS is complementary to a sequence comprised in a tRNA (e.g., a sequence located at the 3’ end of the tRNA) normally provided by the host cell in order to start the reverse transcription.
[0045] In some embodiments, the template RNA is single stranded. Page 21 of 92 12806854v1Attorney Docket No.: 2017469-0044
[0046] In some embodiments, the polypurine tract (PPT) comprises at least 50%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% A or G nucleotides. In some embodiments, the PPT is responsible for starting the synthesis of the proviral (+) DNA strand. In some embodiments, the PPT has a length of about 7, 8, 9, 10, 11, 12, or 13 nucleotides (e.g., 10 nucleotides).
[0047] In some embodiments, the template RNA does not comprise a sequence encoding a functional viral protein (e.g., gag, pol, or a viral reverse transcriptase and / or integrase as described herein, or functional fragments thereof). In some embodiments, the heterologous object sequence is between the 5’ LTR and the 3’ LTR, and one or more sequences encoding functional viral proteins (e.g., gag, pol, or a viral reverse transcriptase and / or integrase as described herein, or functional fragments thereof) is between the 5’ LTR and 3’ LTR (e.g., between the 5’ LTR and the heterologous object sequence).
[0048] In some embodiments, the 5’ LTR is located at, or within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides of the 5’ end of the template RNA. In some embodiments, the 3’ LTR is located at, or within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides of the 3’ end of the template RNA. In some embodiments, one or more of the LTRs has a length of about 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000, 1000-1500, or 1500-2000 nucleotides. In some embodiments, one or more of the LTRs comprises a U3 region (e.g., comprising a promoter). In some embodiments, one or more of the LTRs comprises a repeated region (R). In some embodiments, one or more of the LTRs comprises a U5 region. In some embodiments, one or more of the LTRs comprises a sequence that can be specifically bound by an integrase (e.g., a retroviral or retrotransposon integrase, e.g., as described herein). In some embodiments the 5’ LTR comprises a R and U5 region and the 3’ LTR comprises a U3 and R region. In some embodiments the 5’ LTR lacks a U3 region and the 3’ LTR lacks a U5 region. In some embodiments. In some embodiments the LTR is a self-inactivating (SIN) LTR that has a ∆U3 modification intended to remove promoter or enhancer activity.
[0049] In some embodiments, a template RNA described herein comprises a 5’ LTR and a 3’ LTR having non-identical sequences.
[0050] In some embodiments, a template RNA disclosed herein comprises: (a) a 5’ LTR comprising a sequence according to SEQ ID NO: 5, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (b) a 3’ long terminal repeat (3’ LTR) comprising a sequence according to SEQ ID NO: 6, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (c) a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and (d) a primer binding site (PBS). Page 22 of 92 12806854v1Attorney Docket No.: 2017469-0044
[0051] In some embodiments, the 5’ LTR comprises an A at position 1 of SEQ ID NO: 5 and a G at position 2 of SEQ ID NO: 5. In some embodiments, the 3’ LTR comprises an A at position 231 of SEQ ID NO: 6 and a G at position 232 of SEQ ID NO: 6.
[0052] In some embodiments, the 5’ LTR comprises an A at position 1 of SEQ ID NO: 5 and a G at position 2 of SEQ ID NO: 5 and the 3’ LTR comprises an A at position 231 of SEQ ID NO: 6 and a G at position 232 of SEQ ID NO: 6.
[0053] In some embodiments, the PBS is between the 5’ LTR and the heterologous object sequence.
[0054] the 5’ LTR has a length of less than 354, 300, 200, 150, or 125 nucleotides.
[0055] In some embodiments, the 5’ LTR has a length of 354 nucleotides. In some embodiments, the 5’ LTR has a length of 124 nucleotides.
[0056] In some embodiments, the 3’ LTR has a length of less than 354, 340, 330, 320, 310, 300, 299, or 298 nucleotides. In some embodiments, the 3’ LTR has a length of 354 nucleotides. In some embodiments, 3’ LTR has a length of 297 nucleotides.
[0057] In some embodiments, the 5’ UTR has a sequence according to SEQ ID NO: 4, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0058] In some embodiments, the 3’ UTR has a sequence according to SEQ ID NO: 4, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0059] In some embodiments, a template RNA disclosed herein comprises: (a) a 5’ LTR comprising a sequence according to SEQ ID NO: 5 or 2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (b) a 3’ LTR comprising a sequence according to SEQ ID NO: 6 or 3, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (c) a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and (d) a primer binding site (PBS).
[0060] In some embodiments, the PBS is between the 5’ LTR and the heterologous object sequence.
[0061] In some embodiments, the 5’ LTR has a length of less than 354, 300, 200, 150, or 125 nucleotides. In some embodiments, the 3’ LTR has a length of less than 354, 340, 330, 320, 310, 300, 299, or 298 nucleotides. In some embodiments, the 5’ LTR has a length of less than 354, 300, 200, 150, or 125 nucleotides and the 3’ LTR has a length of less than 354, 340, 330, 320, 310, 300, 299, or 298 nucleotides.
[0062] In some embodiments, the 5’ LTR comprises an A at position 1 of SEQ ID NO: 5 and a G at position 2 of SEQ ID NO: 5. In some embodiments, the 3’ LTR comprises an A at position 231 of SEQ ID NO: 6 and a G at position 232 of SEQ ID NO: 6.
[0063] In some embodiments, the 5’ LTR has a length of 124 nucleotides.
[0064] In some embodiments, the 3’ LTR has a length of 297 nucleotides. Page 23 of 92 12806854v1Attorney Docket No.: 2017469-0044
[0065] In some embodiments, the template RNA comprises one or more non-canonical or modified ribonucleotides.
[0066] In some embodiments, provided herein are DNA molecule intermediates containing reverse transcribed LTRs. In some embodiments, a DNA molecule comprises a SEQ ID NO: 4. In some embodiments, the DNA molecule comprises a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 4. In some embodiments, the DNA molecule has the nucleotide sequence of SEQ ID NO: 4 except that nucleotide 231 of SEQ ID NO: 4 is A and nucleotide 232 of SEQ ID NO: 4 is G.
[0067] In some embodiments, the DNA molecule has a sequence that is the reverse complement of the sequence described herein. In some embodiments, the DNA molecule is single stranded. In some embodiments, the DNA molecule is double stranded.
[0068] The template RNA of the system typically comprises an object sequence for insertion into a target DNA. The object sequence may be coding or non-coding. In some embodiments, the heterologous object sequence (e.g., of a system as described herein) is about 1-50, 50-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000, 1000-2000, 2000-3000, 3000-4000, 4000- 5000, 5000-6000, 6000-7000, 7000-8000, 8000-9000, 9000-10000, or more, nucleotides in length.
[0069] In some embodiments, the object sequence may contain a non-coding sequence. For example, the template RNA may comprise a promoter or enhancer sequence. In some embodiments, the template RNA comprises a tissue specific promoter or enhancer, each of which may be unidirectional or bidirectional. It is understood that, when a template RNA is described as comprising an open reading frame or the reverse complement thereof, in some embodiments the template RNA must be converted into double stranded DNA (e.g., through reverse transcription) before the open reading frame can be transcribed and translated.
[0070] In some embodiments the template RNA has a poly-A tail at the 3’ end. In some embodiments the template RNA does not have a poly-A tail at the 3’ end.
[0071] In some embodiments, a template RNA described herein comprises one or more key transcription regulatory elements, e.g., a CCAAT box, a TATA box, a Inr, and a polyA signal. In some embodiments, the template RNA does not comprise a naturally occurring promoter of the IAP LTR. In some embodiments, and without wishing to be bound by theory, the template RNA lacking a native promoter substantially reduces genotoxicity, if any, associated with integration of a promoter sequence. In some embodiments, a template RNA described herein comprises: (a) an IAP 5’ long terminal repeat (IAP 5’ LTR); (b) an IAP 3’ long terminal repeat (IAP 3’ LTR); (c) a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and (d) a primer binding site (PBS). In some embodiments, the PBS is between the IAP 5’ LTR and the heterologous object sequence. In some Page 24 of 92 12806854v1Attorney Docket No.: 2017469-0044 embodiments, the IAP 5’ LTR comprises a sequence according to SEQ ID NO: 8, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the IAP 3’ LTR comprises a sequence according to SEQ ID NO: 9, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0072] In some embodiments, the IAP 5’ LTR comprises a sequence according to SEQ ID NO: 8, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and the IAP 3’ LTR comprises a sequence according to SEQ ID NO: 9, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto
[0073] In some embodiments, a template RNA described herein comprises: (a) an IAP 5’ long terminal repeat (IAP 5’ LTR); (b) an IAP 3’ long terminal repeat (IAP 3’ LTR); (c) a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and (d) a primer binding site (PBS). In some embodiments, the PBS is between the IAP 5’ LTR and the heterologous object sequence. In some embodiments, the IAP 5’ LTR consists of a first fragment of a full-length IAP 5’ LTR according to SEQ ID NO: 1. In some embodiments, the first fragment consists of nucleotides 231- 255 of SEQ ID NO: 1, or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto. In some embodiments, the IAP 5’ LTR consists of a second fragment of a full-length IAP 5’ LTR according to SEQ ID NO: 1, wherein the second fragment consists of nucleotides 315-354 of SEQ ID NO: 1, or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto.
[0074] In some embodiments, the IAP 5’ LTR consists of a first fragment of a full-length IAP 5’ LTR according to SEQ ID NO: 1 and a second fragment of a full-length IAP 5’ LTR according to SEQ ID NO: 1. In some embodiments, the first fragment consists of nucleotides 231-255 of SEQ ID NO: 1, or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto and the second fragment consists of nucleotides 315-354 of SEQ ID NO: 1, or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto.
[0075] In some embodiments, the IAP 3’ LTR consists of a first fragment of a full-length IAP 3’ LTR according to SEQ ID NO: 1. In some embodiments, the first fragment consists of nucleotides 1-40 of SEQ ID NO: 1, or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto. In some embodiments, the IAP 3’ LTR consists of a second fragment of a full-length IAP 3’ LTR according to SEQ ID NO: 1. In some embodiments, the second fragment consists of nucleotides 231-255 of SEQ ID NO: 1, or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto.
[0076] In some embodiments, the IAP 3’ LTR consists of a first fragment of a full-length IAP 3’ LTR according to SEQ ID NO: 1 and a second fragment of a full-length IAP 3’ LTR according to SEQ ID NO: 1. In some embodiments, the first fragment consists of nucleotides 1-40 of SEQ ID NO: 1, or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto and the second fragment Page 25 of 92 12806854v1Attorney Docket No.: 2017469-0044 consists of nucleotides 231-255 of SEQ ID NO: 1, or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto.
[0077] In some embodiments, the first fragment of a full-length IAP 5’ LTR comprises up to 2, 5, or 10 additional nucleotides adjacent to either side of said nucleotides of SEQ ID NO: 1. In some embodiments, the second fragment of a full-length IAP 5’ LTR comprises up to 2, 5, or 10 additional nucleotides adjacent to either side of said nucleotides of SEQ ID NO: 1. In some embodiments, the first fragment of a full-length IAP 3’ LTR comprises up to 2, 5, or 10 additional nucleotides adjacent to either side of said nucleotides of SEQ ID NO: 1. In some embodiments, the second fragment of a full-length IAP 3’ LTR comprises up to 2, 5, or 10 additional nucleotides adjacent to either side of said nucleotides of SEQ ID NO: 1.
[0078] In some embodiments, a template RNA described herein comprises: (a) an IAP 5’ long terminal repeat (IAP 5’ LTR); (b) an IAP 3’ long terminal repeat (IAP 3’ LTR); (c) a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and (d) a primer binding site (PBS). In some embodiments, the PBS is between the IAP 5’ LTR and the heterologous object sequence.
[0079] In some embodiments, the IAP 5’ LTR comprises a mutation that reduces promoter function of the IAP 5’ LTR compared to an IAP 5’ LTR having a sequence of SEQ ID NO: 1. In some embodiments, the IAP 3’ LTR comprises a mutation that reduces promoter function of the IAP 3’ LTR compared to an IAP 3’ LTR having a sequence of SEQ ID NO: 1. In some embodiments, the mutation is a mutation that reduces promoter function is a deletion, e.g., a deletion of one or more of (e.g. two or all of): (i) the TATA-box; (ii) the CCAA-box; (iii) Inr motif; or (iv) the polyadenylation signal.
[0080] In some embodiments, the mutation is a deletion of the TATA box. In some embodiments, the mutation is a deletion of the CCAA box. In some embodiments, the mutation is a deletion of the Inr motif. In some embodiments, the mutation is a deletion of the polyadenylation signal.
[0081] In some embodiments, the mutation is a deletion of the TATA box and the CCAA box. In some embodiments, the mutation is a deletion of the TATA box and the Inr motif. In some embodiments, the mutation is a deletion of the TATA box and the polyadenylation signal. In some embodiments, the mutation is a deletion of the CCAA box and the Inr motif. In some embodiments, the mutation is a deletion of the CCAA box and the polyadenylation signal. In some embodiments, the mutation is a deletion of the Inr motif and the polyadenylation signal. In some embodiments, the mutation is a deletion of the TATA box, the CCAA box, and the Inr motif. In some embodiments, the mutation is a deletion of the TATA box, the CCAA box, and the polyadenylation signal. In some embodiments, the mutation is a deletion of the TATA box, the Inr motif, and the polyadenylation signal. In some embodiments, the mutation is a deletion of the CCAA box, the Inr motif, and the polyadenylation signal. In some Page 26 of 92 12806854v1Attorney Docket No.: 2017469-0044 embodiments, the mutation is a deletion of the TATA box, the CCAA box, the Inr motif, and the polyadenylation signal.
[0082] In some embodiments, the template RNA described herein is transcribed at a lower level than a control template RNA having the same sequence as the template RNA except that its 3’ LTR has a sequence of SEQ ID NO: 1. In some embodiments, the template RNA described herein is transcribed at a lower level than a control template RNA having the same sequence as the template RNA except that its 5’ LTR has a sequence of SEQ ID NO: 1. In some embodiments, the template RNA described herein is active for integration into a target nucleic acid, e.g., in an assay according to Example 5.
[0083] In some embodiments, the template RNA comprises one or more non-canonical or modified ribonucleotides.
[0084] In some embodiments, provided herein are DNA molecule intermediates containing reverse transcribed LTRs. In some embodiments, the DNA molecule comprises a sequence according to SEQ ID NO: 7. In some embodiments, the DNA molecule comprises a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 4.
[0085] In some embodiments, the DNA molecule has a sequence that is the reverse complement of the sequence described herein. In some embodiments, the DNA molecule is single stranded. In some embodiments, the DNA molecule is double stranded.
[0086] In some embodiments a system or method described herein comprises a single template RNA. In some embodiments a system or method described herein comprises a plurality of template RNAs. In some embodiments, when the system comprises a plurality of nucleic acids, one or more nucleic acid comprises a conjugating domain. In some embodiments, a conjugating domain enables association of nucleic acid molecules, e.g., by hybridization of complementary sequences.
[0087] It is understood that in referring to nucleotide distances between elements in nucleotides, unless specified otherwise, distance refers to the number of nucleotides (of a single strand) or base pairs (in a double strand) that are between the elements but not part of the elements. As an example, if a first element occupies nucleotides 1-100, and a second element occupies nucleotides 102-200 of the same nucleic acid, the distance between the first element and the second element is 1 nucleotide. Polypeptide Components
[0088] In some embodiments the polypeptide driver includes (e.g., as an individual polypeptide or a domain of a fusion protein) a structural polypeptide domain, e.g., an LTR retrotransposase structural polypeptide domain. In some embodiments the polypeptide driver includes (e.g., as an individual polypeptide or a domain of a fusion protein) a reverse transcriptase polypeptide domain, e.g., an LTR retrotransposon reverse transcriptase polypeptide domain. In some embodiments, the reverse Page 27 of 92 12806854v1Attorney Docket No.: 2017469-0044 transcriptase domain is capable of reverse transcribing the template RNA, thereby producing a template DNA.
[0089] In some embodiments, the polypeptide driver comprises one polypeptide. In some embodiments, the polypeptide driver comprises a plurality of polypeptides. In some embodiments, the polypeptide driver is part of a polyprotein or a fusion protein. In some embodiments, the polypeptide driver is not part of a polyprotein or a fusion protein.
[0090] In some embodiments, the system comprises neither an envelope polypeptide domain (e.g., a retroviral envelope polypeptide domain, e.g., a lentiviral envelope polypeptide domain) nor a nucleic acid molecule encoding the envelope polypeptide domain.
[0091] Gag. Gag is processed by protease into matrix (MA), capsid (CA), and nucleocapsid (NC) proteins. MA is necessary for membrane targeting of gag polyprotein and for capsid assembly. Matrix interacts with viral membrane. CA forms the prominent hydrophobic core of the virion. (viral capsid). The best-conserved part of the gag polyprotein is the CA-like major homology region (MHR), which usually displays a central QG-X2-E-X5-F-X2-L-X2-H motif (SEQ ID NO: 14) implicated in the transposition. NC is involved in RNA packaging through recognition of a specific region of the viral genome called Ψ (PSI genome packaging). A second similarity within gag polyproteins is found in the C-terminus of the NC as a Cys-X2-Cys-X4-His-X4-Cys (CCHC) motif (SEQ ID NO: 15), which may be absent or found one, two, or three times duplicated depending on the viral species. CCHC arrays have been found to be critical for many steps in the viral life cycle, and several studies have shown they are involved in virion assembly, RNA packaging, reverse transcription, and integration processes. Each CCHC motif coordinates a zinc atom. Gag may lack Matrix in some cases, e.g. Ty3 (onlinelibrary.wiley.com / doi / abs / 10.1128 / 9781555819217.ch42). Gag may lack NC in some cases, e.g., Ty1. Gag in LTR retrotransposons typically lacks functional sequence for myristoylation and plasma membrane targeting (Ribet al 2006). In some embodiments, the structural polypeptide domain does not comprise a retroviral matrix protein domain. In some embodiments, the structural polypeptide domain comprises a capsid protein domain. In some embodiments, the structural polypeptide domain comprises a nucleocapsid protein domain.
[0092] Pol. Pol translation can be mediated by several mechanisms. For examples, the retrotransposon may include an internal ribosome entry site (IRES) for Pol. The sequence between Gag and Pol ORFs may include a small repetitive motif (such as AAAAA) that induces slippage of the ribosome, which then allows the translation of the second ORF by frameshifting. Another possible means is the use of a specific and rare transfer RNA (tRNA), causing ribosomal stalling and slippage and allowing entry into the second ORF. Gag and Pol may also occur in a ORF along with gag. The component proteins of Pol may occur in various orders (e.g., TY1 / Copia like: PR-INT-RT-RH; Page 28 of 92 12806854v1Attorney Docket No.: 2017469-0044 TY3 / Gypsy like: PR-RT-RH-INT). They may also be frameshifted from each other, as in intracisternal A particle (IAP) elements.
[0093] Protease. Proteases (PR) play a key role in the maturation process during which several peptides involved in the life cycle of the retroelement are excised by this enzyme. LTR retroelement PRs belong to clan AA of aspartic peptidases. They dimerize in their active form and may be encoded as a part of the pol polyprotein, alone or as a part of the gag polyprotein, or in frame with a dUTPase. It is well known that the structural PR homodomain is founded in a core ~90-150 residues long wherein the catalytic DTG motif is the most prominent feature along with a glycine at the C-terminal end preceded by two hydrophobic residues. At the primary structure level, the most conserved part (core) of all clan peptidases may be divided in six amino acidic patterns constituting a template we have called "DTG / ILG". The "DTG / ILG" template is the primary structure phenotype of a structural supersecondary structure, called "Andreeva’s" template (Andreeva 1991) that was previously used to describe pepsins and retropepsin. The "Andreeva’s" template is constituted by the following structural elements: an N-terminal loop (A1), a loop containing the catalytic motif (B1), an α-helix (C1) usually not preserved in retropepsins, a β-hairpin loop (D1), a hairpin loop (A2), a wide loop (B2), an α-helix (C2) towards C- terminal, and a loop (D2), which in empirically characterized retropepsins is substituted by a strand or a helical turn (Wlodawer and Gustchina 2000; Dunn et al.2002). These elements are responsible of keep both function and three-dimensional (3D) structure in characterized retropepsins and other characterized clan AA peptidases (Wlodawer and Gustchina 2000; Dunn et al.2002). It has also recently suggested that the structure of the HIV-1 (see the figure below) and other clan AA PRs have a flexibility- assisted mechanism evolutionarily preserved to favor the reactive conformation of the enzyme (Piana, Carloni, and Rothlisberger 2002; Piana, Carloni, and Parrinello 2002; Perryman, Lin, and McCammon 2004).
[0094] Reverse transcriptase. The Reverse Transcriptase (RT) is an enzyme capable of catalyzing the synthesis of DNA from a single strain of RNA or DNA. The reverse-transcription process is common among a wide range of prokaryotic and eukaryotic mobile genetic elements and requires a primer of 12- 18 bases in length usually provided by the 3´end of a host tRNA. At the primary structure level, RTs codified by Ty3 / Gypsy and Retroviridae elements expand approximately 350 residues of the pol polyprotein, including an alignable core of approximately 180 aa wherein seven conserved regions can be distinguished. At the three-dimensional (3D) structure level the RT codified by the HIV-1 retrovirus is an asymmetrical heterodimer composed of two subunits of 66 and 51 kDa, p66 and p51 respectively. P66 can be divided into five structural subdomains consisting in the RNaseH domain and four subdomains which, due to their similarity to a human right hand, are referred to as fingers, palm, thumb, and connection (Kohlstaedt et al.1992). P51 is a p-66´ derivative after proteolytic processing and excision of the RNase Page 29 of 92 12806854v1Attorney Docket No.: 2017469-0044 H. Although evidence indicate that RTs encoded by other vertebrate retroviruses also form a heterodimer, the RT may also be functionally active as a monomer.
[0095] Ribonuclease H. Ribonuclease H (RNase H) is a hydrolytic enzyme widely distributed in both prokaryotes and eukaryotes (Johnson et al.1986; Doolittle et al.1989). In Ty3 / Gypsy and Retroviridae and other LTR retroelements this enzyme is encoded as a part of the pol polyprotein and constitutes the C-terminal end of the Reverse Transcriptase (RT). RNase H is responsible for the hydrolysis of the original RNA template that is part of the RNA / DNA hybrid generated after the retrotranscription process in the viral life cycle. The three-dimensional (3D) structure of the HIV-1 RNase H is characterized by four or five α-helices and five β-sheets that interact aligning in parallel to conform the active site (Davies et al.1991). The activity of this enzyme normally requires the presence of divalent cations like Mg2+ or Mn2+ that bind to an active site constituted by a catalytic triad (Asp-443-Glu-478- Asp-498). These three residues have been proposed to be important in RNase H-mediated catalysis by HIV-1 RT (Mizarhi et al.1990; Davies et al.1991). Mutations in any of these resides inhibit the RNase H activity but have small effects on polymerase activity of the HIV-1 retrovirus (Schatz et al.1989; Mizarhi et al.1990; Davies et al.1991; Destefano et al.1994).
[0096] Integrase. Retroelement integrases (INTs) are zinc finger nucleic acid-processing enzymes that catalyze the insertion of reverse-transcribed retroviral DNA into the host genome (Chiu and Davies 2004; Nowotny 2009). These enzymes remove two bases from the end of the LTR and are responsible for the insertion of the linear double-stranded viral DNA copy into the host cell DNA. INT amino acid architecture includes three subdomains: (a) The N-terminal subdomain, which displays a conserved Zinc finger "HHCC" binding motif (Lodi et al.1995); (b) The central subdomain, which contains a catalytic core characterized by the presence of a conserved D-D-E motif (Kan et al.1991; Polard and Chandler 1995); and (c) The C-terminal subdomain, which is less preserved than the others. INT enzyme seems to be related to unspecific DNA-binding although several studies of chimeric integrases assign this function to the central core (Katzman and Sudol 1995; Shibagaki and Chow 1997), while other authors alternatively suggest that the C-terminal subdomain might interact with a sub-terminal region of the viral DNA (Jenkins et al.1997; Heuer and Brown 1997; Esposito and Craigie 1998; Heuer and Brown 1998). The functional structure of LTR retroelement-like INTs is already under study although it seems to be, together with a proviral DNA molecule and other viral and host proteins, part of a pre-integration complex of which little is known. Several studies suggest that this enzyme could act as a multimer or at least as a dimer (for a review in this topic see Craigie 2001).
[0097] Chromodomain. LTR retrotransposons may include a Chromatin Organization Modifier Domain (chromodomain). The chromodomain is a protein domain of approximately 50 residues in length, originally identified as a motif common to the Drosophila chromatin proteins Polycomb (Pc) and Page 30 of 92 12806854v1Attorney Docket No.: 2017469-0044 the heterochromatin protein1 HP1. Chromodomains are involved in chromatin remodeling and regulation of the gene expression in eukaryotes (Koonin, Zhou and Lucchesi 1995; Cavalli and Paro 1998). Almost but not all elements belonging to a lineage of Metaviridae Ty3 / Gypsy LTR retrotransposons described in the genomes of plants, fungi, and vertebrates, are carriers of a chromodomain displayed at the C-terminal end of their integrases (Malik and Eickbush 1999).
[0098] dUTPase. dUTPases (DUTs) are cellular enzymes closely similar to Uracil-DNA glycosylases and that hydrolyze dUTP to dUMP and PPi, providing a substrate for thymidylate synthase (an enzyme that converts dUMP to TMP). The expression of cellular DUTs is regulated by the cell cycle; at high levels in dividing undifferentiated cells; and at low levels in terminally non-dividing differentiated cells (Miller et al.2000). Certain retroviral lineages such as non-primate lentiviruses, betaretroviruses, and ERV-L elements encode and package DUTs into virus particles. However, depending on the genus, the dut gene is located in different zones of the internal region. While betaretroviruses codify for this enzyme in frame and N-terminal to the protease domain, lentiviruses and ERV-L elements present the ORF of this gene between or downstream to the RNaseH and INT domains (Elder et al.1992; Turelli et al.1997; Payne and Elder 2001 and references therein). In lentiviruses, DUT facilitates viral replication in non-dividing cells and prevents accumulation of G-to-A transitions in the viral genome, the role of DUT in betaretroviruses and ERV-L elements is still unclear. DUTPase domains have been also described in the genome of some Ty3 / Gypsy LTR retrotransposons (Novikova and Blinov 2008) as well as in that of two plant paretroviruses belonging to Badnavirus genus [Dioscorea bacilliform virus (DBV) and Taro bacilliform virus (TaBV)].
[0099] In some embodiments, one or more of the gag, pol, gag-pol, reverse transcriptase polypeptide domain, and / or integrase domain are derived from an LTR retrotransposon, e.g., as described herein. In some embodiments, one or more of the gag, pol, gag-pol, reverse transcriptase polypeptide domain, and / or integrase domain are derived from a retrovirus (e.g., a an endogenous retrovirus), e.g., as described herein. In some embodiments, one or more of the gag, pol, gag-pol, reverse transcriptase polypeptide domain, and / or integrase domain are derived from an endogenous retrovirus, e.g., as described herein. In some embodiments, one or more of the gag, pol, gag-pol, reverse transcriptase polypeptide domain, and / or integrase domain are introduced into the cell as proteins. In some embodiments, one or more of the gag, pol, gag-pol, reverse transcriptase polypeptide domain, and / or integrase domain are introduced into the cell as RNA (e.g., mRNA that is translated to produce the proteins). In some embodiments, one or more of the gag, pol, gag-pol, reverse transcriptase polypeptide domain, and / or integrase domain are introduced into the cell as DNA (e.g., a plasmid or episome), e.g., wherein genes encoding the gag, pol, gag-pol, reverse transcriptase polypeptide domain, and / or integrase domain are transcribed from the DNA and the resultant mRNA subsequently translated to produce the protein. In some embodiments, one or Page 31 of 92 12806854v1Attorney Docket No.: 2017469-0044 more of the gag, pol, gag-pol, reverse transcriptase polypeptide domain, and / or integrase domain is introduced into the cell by electroporation. In some instances, one or more of the gag, pol, gag-pol, reverse transcriptase polypeptide domain, and / or integrase domain is introduced into the cell via a lipid nanoparticle (LNP).
[0100] In some embodiments, the structural polypeptide domain comprises a gag polyprotein, or a functional fragment (e.g., domain) thereof (e.g., a P24, P17, or P7 / P9 domain). In some embodiments, the structural polypeptide domain lacks a plasma membrane targeting sequence. In some embodiments, the structural polypeptide domain comprises a matrix (MA) protein (e.g., a P17 protein). In some embodiments, the matrix protein is encoded as a separate polypeptide from a further structural polypeptide domain (e.g., a capsid protein and / or a nucleocapsid). In some embodiments, the structural polypeptide domain comprises a capsid (CA) protein (e.g., a P24 protein). In some embodiments, the structural polypeptide domain comprises a nucleocapsid (NC) protein (e.g., a P7 / P9 protein). In some embodiments, the structural polypeptide domain does not comprise a matrix protein. In some embodiments, the structural polypeptide domain does not comprise a nucleocapsid protein.
[0101] In some embodiments, the reverse transcriptase polypeptide domain comprises a pol polyprotein, or a functional fragment (e.g., domain) thereof (e.g., an RT, IN, PR, or DU domain). In some embodiments, the reverse transcriptase polypeptide domain comprises a retroviral or retrotransposon reverse transcriptase (RT). In some embodiments, the reverse transcriptase polypeptide domain comprises a retroviral or retrotransposon protease (PR). In some embodiments, the reverse transcriptase polypeptide domain comprises a retroviral or retrotransposon integrase (IN). In some embodiments, the reverse transcriptase polypeptide domain comprises a retroviral or retrotransposon dUTPase (DU). In some embodiments, the reverse transcriptase polypeptide domain comprises a RNase H. In some embodiments, the reverse transcriptase polypeptide domain comprises a chromodomain. In some embodiments, the reverse transcriptase polypeptide domain does not comprise a chromodomain.
[0102] In some embodiments, the structural polypeptide domain and the reverse transcriptase polypeptide domain are part of the same polypeptide (e.g., a gag-pol). In some embodiments, the structural polypeptide domain and the reverse transcriptase polypeptide domain are different polypeptides. In some embodiments, the structural polypeptide domain and the reverse transcriptase polypeptide domain are encoded by the same nucleic acid molecule (e.g., comprising an internal ribosome entry site (IRES) between the sequences encoding the structural polypeptide domain and the reverse transcriptase polypeptide domain). Page 32 of 92 12806854v1Attorney Docket No.: 2017469-0044
[0103] In some embodiments, a template RNA described herein comprises key regions of a trans (non-autonomous) template. In some embodiments, one or more 500bp regions are omitted from the template RNA relative to a wild type IAP retrotransposase polypeptide.
[0104] In some embodiments, a template RNA described herein comprises (e.g., in a 5’ to 3’ direction): (a) a 5’ long terminal repeat (5’ LTR) having a sequence according to SEQ ID NO: 2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (b) a primer binding site (PBS); (c) the nucleic acid sequence between position 2 and position 4, as shown in Figure 2B, encoding the IAP retrotransposase polypeptide , or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (d) the nucleic acid sequence between position 12 and position 13, as shown in Figure 2B, encoding the IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (e) a polypurine tract (PPT); (f) a 3’ long terminal repeat (3’ LTR) having a sequence according to SEQ ID NO: 3 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (g) a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR.
[0105] In some embodiments, the template RNA comprises a PPT that is immediately adjacent to the 3’ LTR. In some embodiments, the nucleic acid sequence between position 2 and position 4, as shown in Figure 2B, encoding the IAP retrotransposase polypeptide comprises a packaging / psi signal. In some embodiments, the nucleic acid sequence between position 2 and position 4, as shown in Figure 2B, encoding the IAP retrotransposase polypeptide comprises a gag splice donor site. In some embodiments, the nucleic acid sequence between position 12 and position 13, as shown in Figure 2B, encoding the IAP retrotransposase polypeptide comprises a central PPT (cPPT).
[0106] In some embodiments, a template RNA described herein comprises: (a) a 5’ long terminal repeat (5’ LTR) having a sequence according to SEQ ID NO: 2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (b) a 3’ long terminal repeat (3’ LTR) having a sequence according to SEQ ID NO: 3 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (c) a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and (d) a primer binding site (PBS); wherein the template RNA does not comprise one or more of (e.g., does not comprise 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or all of) the following 12 nucleic acid sequence regions: (i) the nucleic acid sequence between position 1 and position 2, as shown in Figure 2B, encoding an IAP retrotransposase polypeptide or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (ii) the nucleic acid sequence between position 4 and position 5, as shown in Figure 2B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% Page 33 of 92 12806854v1Attorney Docket No.: 2017469-0044 identity thereto; (iii) the nucleic acid sequence between position 5 and position 6, as shown in Figure 2B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (iv) the nucleic acid sequence between position 6 and position 7, as shown in Figure 2B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (v) the nucleic acid sequence between position 7 and position 8, as shown in Figure 2B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (vi) the nucleic acid sequence between position 8 and position 9, as shown in Figure 2B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (vii) the nucleic acid sequence between position 9 and position 10, as shown in Figure 2B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (viii) the nucleic acid sequence between position 10 and position 11, as shown in Figure 2B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (ix) the nucleic acid sequence between position 11 and position 12, as shown in Figure 2B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (x) the nucleic acid sequence between position 12 and position 13, as shown in Figure 2B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (xi) the nucleic acid sequence between position 13 and position 14, as shown in Figure 2B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and / or (xii) the nucleic acid sequence between position 14 and position 15, as shown in Figure 2B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0107] In some embodiments, the template RNA does not comprise one of the 12 nucleic acid sequence regions described in the preceding paragraph. In some embodiments, the template RNA does not comprise two of the 12 nucleic acid sequence regions described in the preceding paragraph. In some embodiments, the template RNA does not comprise three of the 12 nucleic acid sequence regions described in the preceding paragraph. In some embodiments, the template RNA does not comprise four of the 12 nucleic acid sequence regions described in the preceding paragraph. In some embodiments, the template RNA does not comprise five of the 12 nucleic acid sequence regions described in the preceding paragraph. In some embodiments, the template RNA does not comprise six of the 12 nucleic acid sequence regions described in the preceding paragraph. In some embodiments, the template RNA does not comprise seven of the 12 nucleic acid sequence regions described in the preceding paragraph. In some Page 34 of 92 12806854v1Attorney Docket No.: 2017469-0044 embodiments, the template RNA does not comprise eight of the 12 nucleic acid sequence regions described in the preceding paragraph. In some embodiments, the template RNA does not comprise nine of the 12 nucleic acid sequence regions described in the preceding paragraph. In some embodiments, the template RNA does not comprise ten of the 12 nucleic acid sequence regions described in the preceding paragraph. In some embodiments, the template RNA does not comprise eleven of the 12 nucleic acid sequence regions described in the preceding paragraph. In some embodiments, the template RNA does not comprise all twelve of the 12 nucleic acid sequence regions described in the preceding paragraph.
[0108] In some embodiments, the template RNA described herein comprises a premature stop codon in the sequence encoding gag, pro, pol, or rig. In some embodiments, the template RNA comprises one or more non-canonical or modified ribonucleotides. Integration-Deficient Systems
[0109] The retrotransposon systems described herein may, in some instances, be integration- deficient. In some embodiments, the integrase of the retrotransposon is substantially unable to integrate the template DNA into a target DNA (e.g., a genomic DNA). In some embodiments, the retrotransposon system is integration-deficient independent of host cell repair machinery. In some embodiments, the retrotransposon system is integration-deficient independent of a transposase, recombinase, and / or nuclease of the host cell. In embodiments, the integrase of the retrotransposon has reduced integrase activity, e.g., to at least 50%, 40%, 30%, 20%, 10%, 5%, 2%, or 1% of that of a corresponding wild-type sequence, e.g., as measured in an assay as described in Moldt et al.2008 (BMC Biotechnol.8:60; incorporated herein by reference). In some embodiments, the integrase of the retrotransposon comprises a mutation that reduces integrase activity, e.g., to at least 50%, 40%, 30%, 20%, 10%, 5%, 2%, or 1% of a corresponding wild-type sequence (e.g., a class I mutation, e.g., a mutation in a catalytic triad residue, such as mutations corresponding to D64, D116, and E152 for HIV-1 integrase). In some embodiments, one or both of the U3 and U5 attachment (att) sites at either end of the element may be mutated or deleted to impair integrase binding. In some embodiments, the system comprises an inhibitor (e.g., a small molecule inhibitor) of the integrase of the retrotransposon. Examples of inhibitors include, for HIV-1, strand-transfer inhibitors raltegravir and elvitegravir. In some embodiments, the template RNA and / or template DNA does not comprise a DNA recognition site bound by and / or recognized by the integrase of the retrotransposon.
[0110] In some embodiments, an IAP retrotransposase polypeptide is substantially unable to integrate the template DNA into a target DNA, e.g., is an integration-deficient polypeptide. In some embodiments, the integration-deficient IAP retrotransposase polypeptide has one or more mutations Page 35 of 92 12806854v1Attorney Docket No.: 2017469-0044 relative to a wild-type IAP retrotransposase polypeptide. In some embodiments, the one or more mutations are within the catalytic triad residues (e.g., within the DDE catalytic triad residues).
[0111] In some embodiments, the integration-deficient IAP retrotransposase polypeptide has an amino acid sequence of SEQ ID NO: 12, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the integration-deficient IAP retrotransposase polypeptide has an amino acid sequence of SEQ ID NO: 12 except that it comprises one, two, or all of the following mutations: position D653 of SEQ ID NO: 12 is other than D, e.g., is valine; position D710 of SEQ ID NO: 12 is other than D, e.g., is valine; or position E746 of SEQ ID NO: 12 is other than E, e.g., is valine.
[0112] In some embodiments, position D653 of SEQ ID NO: 12 is valine. In some embodiments, position D710 of SEQ ID NO: 12 is valine. In some embodiments, position E746 of SEQ ID NO: 12 is valine.
[0113] In some embodiments, position D710 of SEQ ID NO: 12 is aspartate (D), and position E746 of SEQ ID NO: 12 is glutamate (E).
[0114] In some embodiments, the integration-deficient IAP retrotransposase polypeptide has reduced integrase activity compared to a polypeptide of SEQ ID NO: 12. In some embodiments, the activity is reduced by about 20%, 40%, 60%, 80%, 90%, or 95%, e.g., as measured by an assay of Example 6.
[0115] In some embodiments, use of an integration-deficient IAP retrotransposase polypeptide described herein results in transient expression of a heterologous object sequence. In some embodiments, expression of the heterologous object sequence is detectable for less than 9 days, e.g., as assessed by an assay of Example 6.
[0116] Also provided herein is a nucleic acid molecule encoding an integration-deficient IAP retrotransposase polypeptide. In some embodiments, the nucleic acid molecule comprises a T at position 1958 of SEQ ID NO: 10, or a nucleic acid molecule having at least 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0117] In some embodiments, the IAP retrotransposase described herein is in a cell. Episomes
[0118] In some embodiments, an LTR retrotransposon-based system or method described herein can produce an episome (e.g., an episome comprising a heterologous object sequence), a circular DNA molecule. In some embodiments, an episome produced by a system or method described herein comprises an LTR. In certain embodiments, an episome produced by a system or method described herein comprises a plurality of LTRs (e.g., exactly two LTRs). In some embodiments, an episome (e.g., Page 36 of 92 12806854v1Attorney Docket No.: 2017469-0044 an episome comprising two LTRs) is formed by non-homologous end joining (NHEJ), e.g., ligating together the 5’ and 3’ ends of a linear DNA (e.g., a vector DNA as described herein). In some embodiments, an episome (e.g., an episome comprising one LTR) is produced by homologous recombination (e.g., between viral 5’ and 3’ LTRs, e.g., via strand-invasion or single-strand annealing). In some embodiments, an episome (e.g., an episome comprising on LTR) is produced by ligation of nicks, e.g., present in intermediate products of reverse transcription.
[0119] In some embodiments, the episome replicates in the target cell. In some embodiments, the episome comprises an origin of replication, e.g., a mammalian origin of replication, e.g., a human origin of replication. In some embodiments, the episome does not replicate in the target cell.
[0120] In some embodiments, a method described herein results in integration of the 5’ LTR and the 3’ LTR into the genome of the target cell. In some embodiments, the 5’ LTR and the 3’ LTR flank the heterologous object sequence after integration into the genome of the target cell.
[0121] In some embodiments, the template RNA is not part of the same nucleic acid molecule as the nucleic acid molecule that encodes the polypeptide driver. In some embodiments, the template RNA is not part of the same nucleic acid molecule as the nucleic acid molecule encoding the structural polypeptide domain.
[0122] In some embodiments, the template RNA is not part of the same nucleic acid molecule as the nucleic acid molecule encoding the reverse transcriptase polypeptide domain.
[0123] In some embodiments, the template RNA is not part of the same nucleic acid molecule as either of the nucleic acid molecule encoding the structural polypeptide domain and the nucleic acid molecule encoding the reverse transcriptase polypeptide domain.
[0124] In some embodiments, the polynucleotide (e.g., RNA) encoding the polypeptide driver does not comprise an LTR or does not comprise an LTR within 500 bp, 1 kb, 1.5 kb, or 2 kb of its coding region. In some embodiments, the polynucleotide (e.g., RNA) encoding the structural polypeptide domain does not comprise an LTR or does not comprise an LTR within 500 bp, 1 kb, 1.5 kb, or 2 kb of its coding region. In some embodiments, the polynucleotide (e.g., RNA) encoding the structural polypeptide domain does not comprise two LTRs or does not comprise two LTRs within 500 bp, 1 kb, 1.5 kb, or 2 kb of its coding region. In some embodiments, the polynucleotide (e.g., RNA) encoding the reverse transcriptase polypeptide domain does not comprise an LTR or does not comprise an LTR within 500 bp, 1 kb, 1.5 kb, or 2 kb of its coding region. In some embodiments, the polynucleotide (e.g., RNA) encoding the reverse transcriptase polypeptide domain does not comprise two LTRs or does not comprise two LTRs within 500 bp, 1 kb, 1.5 kb, or 2 kb of its coding region. Page 37 of 92 12806854v1Attorney Docket No.: 2017469-0044
[0125] In some embodiments, provided herein is a nucleic acid encoding the IAP retrotransposase polypeptide as described herein. In some embodiment, the nucleic acid is an mRNA. In some embodiments, the nucleic acid comprises one or more non-canonical or modified ribonucleotides.
[0126] Without wishing to be bound by theory, when the canonical PPT is not present (e.g., when it is replaced with a mutant variant), the second strand synthesis process will be initiated by the central PPT (cPPT), resulting in an increase in the proportion of functional episomes compared to linear integration- competent molecules.
[0127] In some embodiments, provided herein is a template RNA comprising (e.g., from 5’ to 3’): (a) an IAP 5’ long terminal repeat (IAP 5’ LTR); (b) a primer binding site (PBS); (c) a heterologous object sequence encoding a therapeutic effector molecule; (d) a central PPT (cPPT); and (e) an IAP 3’ long terminal repeat (IAP 3’ LTR). In some embodiments, between (d) and (e), the template RNA lacks a canonical PPT. In some embodiments, between (d) and (e), the template RNA comprises a mutation to a canonical PPT (“mutant PPT”), e.g., that reduces initiation of second strand synthesis primed by the mutant PPT.
[0128] In some embodiments, the cPPT is situated upstream of the heterologous object sequence. In some embodiments, the cPPT is situated downstream of the heterologous object sequence. In some embodiments, the cPPT is situated in an overlapping configuration with the heterologous object sequence.
[0129] In some embodiments, reverse transcription of the template RNA results in a greater proportion of episomes to linear dsDNA, for example, compared to the proportion of episomes to linear dsDNA produced using a reference template RNA which comprises a canonical PPT between its cPPT and IAP 3’ LTR.
[0130] In some embodiments, reverse transcription of the template RNA results in fewer insertions into a host cell genome, for example, compared to the number of insertions into a host cell genome produced using a reference template RNA which comprises a canonical PPT between its cPPT and IAP 3’ LTR.
[0131] Without wishing to be bound by theory, reverse transcription products having a circular episome (e.g. functional episome) configuration are not integrated into a host cell genome, whereas reverse transcription products having a as linear molecule (e.g. linear integration-competent molecule) configuration are capable of integration into a host cell genome. Introduction of a CAR in T cells
[0132] A LTR retrotransposon-based system described herein may be used to modify immune cells. In some embodiments, a system described herein may be used to modify T cells. In some embodiments, T-cells may include any subpopulation of T-cells, e.g., CD4+, CD8+, gamma-delta, naïve T cells, stem Page 38 of 92 12806854v1Attorney Docket No.: 2017469-0044 cell memory T cells, central memory T cells, or a mixture of subpopulations. In some embodiments, a system described herein may be used to deliver or modify a T-cell receptor (TCR) in a T cell. In some embodiments, a system described herein may be used to deliver at least one chimeric antigen receptor (CAR) to T-cells. In some embodiments, a system described herein may be used to deliver at least one CAR to natural killer (NK) cells. In some embodiments, a system described herein may be used to deliver at least one CAR to natural killer T (NKT) cells. In some embodiments, a system described herein may be used to deliver at least one CAR to a progenitor cell, e.g., a progenitor cell of T, NK, or NKT cells. In some embodiments, cells modified with at least one CAR (e.g., CAR-T cells, CAR-NK cells, CAR-NKT cells), or a combination of cells modified with at least one CAR (e.g., a mixture of CAR-NK / T cells) are used to treat a condition as identified in the targetable landscape of CAR therapies in MacKay, et al. Nat Biotechnol 38, 233-244 (2020), incorporated by reference herein in its entirety. In some embodiments, the immune cells comprise a CAR specific to a tumor or a pathogen antigen selected from a group consisting of AChR (fetal acetylcholine receptor), ADGRE2, AFP (alpha fetoprotein), BAFF-R, BCMA, CAIX (carbonic anhydrase IX), CCR1, CCR4, CEA (carcinoembryonic antigen), CD3, CD5, CD8, CD7, CD10, CD13, CD14, CD15, CD19, CD20, CD22, CD30, CD33, CLLI, CD34, CD38, CD41, CD44, CD49f, CD56, CD61, CD64, CD68, CD70,CD74, CD99,CD117, CD123, CD133, CD138, CD44v6, CD267, CD269, CDS, CLEC12A, CS1, EGP-2 (epithelial glycoprotein-2), EGP-40 (epithelial glycoprotein-40), EGFR(HER1), EGFR-VIII, EpCAM (epithelial cell adhesion molecule), EphA2, ERBB2 (HER2, human epidermal growth factor receptor 2), ERBB3, ERBB4, FBP (folate-binding protein), Flt3 receptor, folate receptor-a, GD2 (ganglioside G2), GD3 (ganglioside G3), GPC3 (glypican-3), GPI00, hTERT (human telomerase reverse transcriptase), ICAM-1, integrin B7, interleukin 6 receptor, IL13Ra2 (interleukin-13 receptor 30 subunit alpha-2), kappa-light chain, KDR (kinase insert domain receptor), LeY (Lewis Y), L1CAM (LI cell adhesion molecule), LILRB2 (leukocyte immunoglobulin like receptor B2), MARTI, MAGE-A1 (melanoma associated antigen Al), MAGE- A3, MSLN (mesothelin), MUC16 (mucin 16), MUCI (mucin I), KG2D ligands, NY-ESO-1 (cancer-testis antigen), PRI (proteinase 3), TRBCI, TRBC2, TFM-3, TACI, tyrosinase, survivin, hTERT, oncofetal antigen (h5T4), p53, PSCA (prostate stem cell antigen), PSMA (prostate-specific membrane antigen), hRORl, TAG-72 (tumor- associated glycoprotein 72), VEGF-R2 (vascular endothelial growth factor R2), WT-1 (Wilms tumor protein), and antigens of HIV (human immunodeficiency virus), hepatitis B, hepatitis C, CMV (cytomegalovirus), EBV (Epstein- Barr virus), HPV (human papilloma virus).
[0133] The LNP formulation C14-4, comprising cholesterol, phospholipid, lipid-anchored PEG, and the ionizable lipid C14-4 (Figure 2C of Billingsley et al. Nano Lett 20(3):1578-1589 (2020)) can be used for delivery to T cells, such as ex vivo delivery. Page 39 of 92 12806854v1Attorney Docket No.: 2017469-0044
[0134] Additional edits can be performed on T-cells in order to improve activity of the CAR-T cells against their cognate target. In some embodiments, a second LNP formulation of C14-4 as described comprises a Cas9 / gRNA preformed RNP complex, wherein the gRNA targets the Pdcd1 exon 1 for PD-1 inactivation, which can enhance anti-tumor activity of CAR-T cells by disruption of this inhibitory checkpoint that can otherwise trigger suppression of the cells (see Rupp et al. Sci Rep 7:737 (2017)). The application of both nanoparticle formulation thus enables lymphoma targeting by providing the anti-CD19 cargo, while simultaneously boosting efficacy by knocking out the PD-1 checkpoint inhibitor. In some embodiments, cells may be treated with the nanoparticles simultaneously. In some embodiments, the cells may be treated with the nanoparticles in separate steps, e.g., first deliver the RNP for generating the PD-1 knockout, and subsequently treat cells with the nanoparticles carrying the anti-CD19 CAR. In some embodiments, the second component of the system that improves T cell efficacy may result in the knockout of PD-1, TCR, CTLA-4, HLA-I, HLA-II, CS1, CD52, B2M, MHC-I, MHC-II, CD3, FAS, PDC1, CISH, TRAC, or a combination thereof. In some embodiments, knockdown of PD-1, TCR, CTLA- 4, HLA-I, HLA-II, CS1, CD52, B2M, MHC-I, MHC-II, CD3, FAS, PDC1, CISH, or TRAC may be preferred, e.g., using siRNA targeting PD-1. In some embodiments, siRNA targeting PD-1 may be achieved using self-delivering RNAi as described by Ligtenberg et al. Mol Ther 26(6):1482-1493 (2018) and in WO2010033247, incorporated herein by reference in its entirety, in which extensive chemical modifications of siRNAs, conferring the resulting hydrophobically modified siRNA molecules the ability to penetrate all cell types ex vivo and in vivo and achieve long-lasting specific target gene knockdown without any additional delivery formulations or techniques. In some embodiments, one or more components of the system may be delivered by other methods, e.g., electroporation. In some embodiments, additional regulators are knocked in to the cells for overexpression to control T cell- and NK cell-mediated immune responses and macrophage engulfment, e.g., PD-L1, HLA-G, CD47 (Han et al. PNAS 116(21):10441-10446 (2019)). Knock-in may be accomplished through application of an additional genome editing system as described herein with a template carrying an expression cassette for one or more such factors (3) with targeting to a safe harbor locus, e.g., AAVS1, e.g., using gRNA GGGGCCACTAGGGACAGGAT (SEQ ID NO: 13) to target the polypeptide to AAVS1.
[0135] In order to achieve delivery specifically to T-cells, targeted LNPs (tLNPs) may generated that carry a conjugated mAb against CD4. See, e.g., Ramishetti et al. ACS Nano 9(7):6706-6716 (2015). Alternatively, conjugating a mAb against CD3 can be used to target both CD4+and CD8+T-cells (Smith et al. Nat Nanotechnol 12(8):813-820 (2017)). In other embodiments, the nanoparticle used to deliver to T-cells in vivo is a constrained nanoparticle that lacks a targeting ligand, as taught by Lokugamage et al. Adv Mater 31(41):e1902251 (2019). Page 40 of 92 12806854v1Attorney Docket No.: 2017469-0044 Nucleic acid molecules Circular RNAs
[0136] It is contemplated that it may be useful to employ circular and / or linear RNA states during the formulation, delivery, or writing within the target cell. Thus, in some embodiments of any of the aspects described herein, a system comprises one or more circular RNAs (circRNAs). In some embodiments of any of the aspects described herein, a system comprises one or more linear RNAs. In some embodiments, a nucleic acid as described herein (e.g., a template nucleic acid, a nucleic acid molecule encoding a polypeptide driver, or both) is a circRNA. Doggybone DNA
[0137] In some embodiments, nucleic acid (e.g., encoding a polypeptide driver) delivered to cells is covalently closed linear DNA, or so-called “doggybone” DNA. During its lifecycle, the bacteriophage N15 employs protelomerase to convert its genome from circular plasmid DNA to a linear plasmid DNA (Ravin et al. J Mol Biol 2001). This process has been adapted for the production of covalently closed linear DNA in vitro (see, for example, WO2010086626A1). Chemically modified nucleic acids and nucleic acid end features:
[0138] A nucleic acid described herein (e.g., a template RNA; or a nucleic acid (e.g., mRNA) encoding a polypeptide driver) can comprise unmodified or modified nucleobases. Naturally occurring RNAs are synthesized from four basic ribonucleotides: ATP, CTP, UTP and GTP, but may contain post- transcriptionally modified nucleotides. Further, approximately one hundred different nucleoside modifications have been identified in RNA (Rozenski, J, Crain, P, and McCloskey, J. (1999). The RNA Modification Database: 1999 update. Nucl Acids Res 27: 196-197). An RNA can also comprise wholly synthetic nucleotides that do not occur in nature.
[0139] In some embodiments, a chemical modification is one provided in PCT / US2016 / 032454, US Pat. Pub. No.20090286852, of International Application No. WO / 2012 / 019168, WO / 2012 / 045075, WO / 2012 / 135805, WO / 2012 / 158736, WO / 2013 / 039857, WO / 2013 / 039861, WO / 2013 / 052523, WO / 2013 / 090648, WO / 2013 / 096709, WO / 2013 / 101690, WO / 2013 / 106496, WO / 2013 / 130161, WO / 2013 / 151669, WO / 2013 / 151736, WO / 2013 / 151672, WO / 2013 / 151664, WO / 2013 / 151665, WO / 2013 / 151668, WO / 2013 / 151671, WO / 2013 / 151667, WO / 2013 / 151670, WO / 2013 / 151666, WO / 2013 / 151663, WO / 2014 / 028429, WO / 2014 / 081507, WO / 2014 / 093924, WO / 2014 / 093574, WO / 2014 / 113089, WO / 2014 / 144711, WO / 2014 / 144767, WO / 2014 / 144039, WO / 2014 / 152540, WO / 2014 / 152030, WO / 2014 / 152031, WO / 2014 / 152027, WO / 2014 / 152211, WO / 2014 / 158795, WO / 2014 / 159813, WO / 2014 / 164253, WO / 2015 / 006747, WO / 2015 / 034928, WO / 2015 / 034925, Page 41 of 92 12806854v1Attorney Docket No.: 2017469-0044 WO / 2015 / 038892, WO / 2015 / 048744, WO / 2015 / 051214, WO / 2015 / 051173, WO / 2015 / 051169, WO / 2015 / 058069, WO / 2015 / 085318, WO / 2015 / 089511, WO / 2015 / 105926, WO / 2015 / 164674, WO / 2015 / 196130, WO / 2015 / 196128, WO / 2015 / 196118, WO / 2016 / 011226, WO / 2016 / 011222, WO / 2016 / 011306, WO / 2016 / 014846, WO / 2016 / 022914, WO / 2016 / 036902, WO / 2016 / 077125, or WO / 2016 / 077123, each of which is herein incorporated by reference in its entirety. It is understood that incorporation of a chemically modified nucleotide into a polynucleotide can result in the modification being incorporated into a nucleobase, the backbone, or both, depending on the location of the modification in the nucleotide. In some embodiments, the backbone modification is one provided in EP 2813570, which is herein incorporated by reference in its entirety. In some embodiments, the modified cap is one provided in US Pat. Pub. No.20050287539, which is herein incorporated by reference in its entirety.
[0140] In some embodiments, the chemically modified nucleic acid (e.g., RNA, e.g., mRNA) comprises one or more of ARCA: anti-reverse cap analog (m27.3'-OGP3G), GP3G (Unmethylated Cap Analog), m7GP3G (Monomethylated Cap Analog), m32.2.7GP3G (Trimethylated Cap Analog), m5CTP (5'-methyl-cytidine triphosphate), m6ATP (N6-methyl-adenosine-5'-triphosphate), s2UTP (2-thio-uridine triphosphate), and Ѱ (pseudouridine triphosphate).
[0141] In some embodiments, the chemically modified nucleic acid comprises a 5’ cap, e.g.: a 7- methylguanosine cap (e.g., a O-Me-m7G cap); a hypermethylated cap analog; an NAD+-derived cap analog (e.g., as described in Kiledjian, Trends in Cell Biology 28, 454-464 (2018)); or a modified, e.g., biotinylated, cap analog (e.g., as described in Bednarek et al., Phil Trans R Soc B 373, 20180167 (2018)).
[0142] In some embodiments, the chemically modified nucleic acid comprises a 3’ feature selected from one or more of: a polyA tail; a 16-nucleotide long stem-loop structure flanked by unpaired 5 nucleotides (e.g., as described by Mannironi et al., Nucleic Acid Research 17, 9113-9126 (1989)); a triple- helical structure (e.g., as described by Brown et al., PNAS 109, 19202-19207 (2012)); a tRNA, Y RNA, or vault RNA structure (e.g., as described by Labno et al., Biochemica et Biophysica Acta 1863, 3125- 3147 (2016)); incorporation of one or more deoxyribonucleotide triphosphates (dNTPs), 2’O-Methylated NTPs, or phosphorothioate-NTPs; a single nucleotide chemical modification (e.g., oxidation of the 3’ terminal ribose to a reactive aldehyde followed by conjugation of the aldehyde-reactive modified nucleotide); or chemical ligation to another nucleic acid molecule.
[0143] In some embodiments, the nucleic acid (e.g., template nucleic acid) comprises one or more modified nucleotides, e.g., selected from dihydrouridine, inosine, 7-methylguanosine, 5-methylcytidine (5mC), 5′ Phosphate ribothymidine, 2′-O-methyl ribothymidine, 2′-O-ethyl ribothymidine, 2′-fluoro ribothymidine, C-5 propynyl-deoxycytidine (pdC), C-5 propynyl-deoxyuridine (pdU), C-5 propynyl- cytidine (pC), C-5 propynyl-uridine (pU), 5-methyl cytidine, 5-methyl uridine, 5-methyl deoxycytidine, Page 42 of 92 12806854v1Attorney Docket No.: 2017469-0044 5-methyl deoxyuridine methoxy, 2,6-diaminopurine, 5′-Dimethoxytrityl-N4-ethyl-2′-deoxycytidine, C-5 propynyl-f-cytidine (pfC), C-5 propynyl-f-uridine (pfU), 5-methyl f-cytidine, 5-methyl f-uridine, C-5 propynyl-m-cytidine (pmC), C-5 propynyl-f-uridine (pmU), 5-methyl m-cytidine, 5-methyl m-uridine, LNA (locked nucleic acid), MGB (minor groove binder) pseudouridine (Ψ), 1-N-methylpseudouridine (1- Me-Ψ), or 5-methoxyuridine (5-MO-U).
[0144] In some embodiments, the nucleic acid comprises a backbone modification, e.g., a modification to a sugar or phosphate group in the backbone. In some embodiments, the nucleic acid comprises a nucleobase modification.
[0145] In some embodiments, the nucleic acid comprises one or more chemically modified nucleotides of Table M1, one or more chemical backbone modifications of Table M2, one or more chemically modified caps of Table M3. For instance, in some embodiments, the nucleic acid comprises two or more (e.g., 3, 4, 5, 6, 7, 8, 9, or 10 or more) different types of chemical modifications. As an example, the nucleic acid may comprise two or more (e.g., 3, 4, 5, 6, 7, 8, 9, or 10 or more) different types of modified nucleobases, e.g., as described herein, e.g., in Table M1. Alternatively or in combination, the nucleic acid may comprise two or more (e.g., 3, 4, 5, 6, 7, 8, 9, or 10 or more) different types of backbone modifications, e.g., as described herein, e.g., in Table M2. Alternatively or in combination, the nucleic acid may comprise one or more modified cap, e.g., as described herein, e.g., in Table M3. For instance, in some embodiments, the nucleic acid comprises one or more type of modified nucleobase and one or more type of backbone modification; one or more type of modified nucleobase and one or more modified cap; one or more type of modified cap and one or more type of backbone modification; or one or more type of modified nucleobase, one or more type of backbone modification, and one or more type of modified cap.
[0146] In some embodiments, the nucleic acid comprises one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, or more) modified nucleobases. In some embodiments, all nucleobases of the nucleic acid are modified. In some embodiments, the nucleic acid is modified at one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, or more) positions in the backbone. In some embodiments, all backbone positions of the nucleic acid are modified. Table M1. Modified nucleotides 5-aza-uridine N2-methyl-6-thio-guanosinePage 43 of 92 12806854v1Attorney Docket No.: 2017469-0044 3-methyluridine 2-thio-pseudowidine 5-carboxymethyl-uridine 3-methylmidine 1-carboxmethl-seudouridine 1-ro nl-seudomidinePage 44 of 92 12806854v1Attorney Docket No.: 2017469-0044 7-deaza-8-aza-adenine 3-(3-amino-3-carboxypropyl)uridine 7-deaza-2-aminopurine 5-methoxyuridine 7-deaza-8-aza-2-aminourine uridine 5-oxacetic acidTable M2. Backbone modifications 2’-O-Methyl backbone P tid N li Aid PNA b kbPage 45 of 92 12806854v1Attorney Docket No.: 2017469-0044 riboacetyl backbone alkene containing backbone sulfamate backbonem7GpppA m7GpppCProduction of Compositions and Systems
[0147] Methods of designing and constructing nucleic acid constructs and proteins or polypeptides (such as the systems, constructs and polypeptides described herein) are known. Generally, recombinant methods may be used. See, in general, Smales & James (Eds.), Therapeutic Proteins: Methods and Protocols (Methods in Molecular Biology), Humana Press (2005); and Crommelin, Sindelar & Meibohm (Eds.), Pharmaceutical Biotechnology: Fundamentals and Applications, Springer (2013). Methods of designing, preparing, evaluating, purifying and manipulating nucleic acid compositions are described in Green and Sambrook (Eds.), Molecular Cloning: A Laboratory Manual (Fourth Edition), Cold Spring Harbor Laboratory Press (2012). Page 46 of 92 12806854v1Attorney Docket No.: 2017469-0044 Kits, Articles of Manufacture, and Pharmaceutical Compositions:
[0148] In an aspect the disclosure provides a kit comprising a system, e.g., as described herein. In some embodiments, the kit comprises a polypeptide driver (or a nucleic acid encoding the same) and a template RNA (or DNA encoding the template RNA). In some embodiments, the kit further comprises a reagent for introducing the system into a cell, e.g., transfection reagent, LNP, and the like. In some embodiments, the kit is suitable for any of the methods described herein. In some embodiments, the kit comprises one or more elements, compositions and / or systems, or a functional fragment or component thereof, e.g., disposed in an article of manufacture. In some embodiments, the kit comprises instructions for use thereof. Chemistry, Manufacturing, and Controls (CMC):
[0149] Purification of protein therapeutics is described, for example, in Franks, Protein Biotechnology: Isolation, Characterization, and Stabilization, Humana Press (2013); and in Cutler, Protein Purification Protocols (Methods in Molecular Biology), Humana Press (2010).
[0150] In some embodiments, a system or pharmaceutical composition described herein is endotoxin free.
[0151] In some embodiments, the presence, absence, and / or level of one or more of a pyrogen, virus, fungus, bacterial pathogen, and / or host cell protein is determined. In embodiments, whether the system is free or substantially free of pyrogen, virus, fungus, bacterial pathogen, and / or host cell protein contamination is determined. Regulation of system components
[0152] It is in some embodiments desirable for a system of this invention to exhibit activity in target cells, while simultaneously having reduced activity in non-target cells. Thus, regulatory control of one or more components of the system is contemplated in some embodiments. Promoters and enhancers:
[0153] In some embodiments, a nucleic acid described herein (e.g., a nucleic acid encoding a polypeptide driver, template RNA, or an open reading frame in a heterologous object sequence) comprises a promoter sequence, e.g., a tissue specific promoter sequence. In some embodiments, the tissue-specific promoter is used to increase the target-cell specificity of a system. For instance, the promoter can be chosen on the basis that it is active in a target cell type but not active in (or active at a lower level in) a non-target cell type. Page 47 of 92 12806854v1Attorney Docket No.: 2017469-0044 miRNAs, inhibitors, and miRNA binding sites:
[0154] miRNAs and other small interfering nucleic acids generally regulate gene expression via target RNA transcript cleavage / degradation or translational repression of the target messenger RNA (mRNA). miRNAs may, in some instances, be natively expressed, typically as final 19-25 non-translated RNA products. miRNAs generally exhibit their activity through sequence-specific interactions with the 3′ untranslated regions (UTR) of target mRNAs.
[0155] In some embodiments, a nucleic acid described herein (e.g., a nucleic acid encoding a polypeptide driver, a template RNA, and / or an open reading frame in a heterologous object sequence) comprises at least one microRNA binding site. In some embodiments, the microRNA binding site is used to increase the target-cell specificity of a system. For instance, the microRNA binding site can be chosen on the basis that it is recognized by a miRNA that is present in a non-target cell type, but that is not present (or is present at a reduced level relative to the non-target cell) in a target cell type. Thus, when the RNA is present in a non-target cell, it would be bound by the miRNA, and when the RNA is present in a target cell, it would not be bound by the miRNA (or bound but at reduced levels relative to the non-target cell). While not wishing to be bound by theory, binding of the miRNA to an RNA of the system may result in destabilization or degradation of the RNA molecule or interference with translation of a coding RNA. Accordingly, the heterologous object sequence would be inserted into the genome of target cells more efficiently than into the genome of non-target cells. It is contemplated that incorporation of one or more appropriate miRNA binding sites into a nucleic acid encoding the polypeptide driver or template RNA would thus reduce integration in off-target cells, while incorporation into a heterologous object sequence would reduce expression of a transgene in off-target cells. A system having a microRNA binding site in a nucleic acid would be expected to exhibit increased specificity for target cells by the addition of more miRNA binding sites on the same or on an additional nucleic acid component of the system. In some embodiments, one or more component of a system comprises one or more miRNA binding sites to reduce activity in off-target cells.
[0156] In some embodiments, a system comprising one or more tissue-specific promoter sequences may be used in combination with one or more microRNA binding sites, e.g., as described herein. Modifications to proteins of the system Subcellular localization signals
[0157] In some embodiments, a polypeptide described herein (e.g., a polypeptide driver or a polypeptide encoded by a heterologous object sequence), comprises one or more (e.g., 2, 3, 4, 5) nuclear targeting sequences, for example, a nuclear localization sequence (NLS), e.g., as described above. In Page 48 of 92 12806854v1Attorney Docket No.: 2017469-0044 some embodiments, the NLS is a bipartite NLS. In some embodiments, an NLS facilitates the import of a protein comprising an NLS into the cell nucleus. In some embodiments, the NLS is fused to the N- terminus of a polypeptide described herein. In some embodiments, the NLS is fused to the C-terminus of a polypeptide described herein. In some embodiments, the NLS is fused to the N-terminus or the C- terminus of a polypeptide or domain described herein. In some embodiments, a linker sequence is disposed between the NLS and the neighboring domain of a polypeptide described herein. Inteins
[0158] In some embodiments, the system comprises an intein. Generally, an intein comprises a polypeptide that has the capacity to join two polypeptides or polypeptide fragments together via a peptide bond. In some embodiments, the intein is a trans-splicing intein that can join two polypeptide fragments, e.g., to form the polypeptide component of a system as described herein. In some embodiments, an intein may be encoded on the same nucleic acid molecule encoding the two polypeptide fragments. In certain embodiments, the intein may be translated as part of a larger polypeptide comprising, e.g., in order, the first polypeptide fragment, the intein, and the second polypeptide fragment. In embodiments, the translated intein may be capable of excising itself from the larger polypeptide, e.g., resulting in separation of the attached polypeptide fragments. In embodiments, the excised intein may be capable of joining the two polypeptide fragments to each other directly via a peptide bond. Exemplary inteins are described in, e.g., Table X of PCT Application No. PCT / US2021 / 020943. Other sequence modifications and improvements
[0159] In some embodiments, a polypeptide for use in any of the systems described herein can be a molecular reconstruction or ancestral reconstruction based upon the aligned polypeptide sequence of multiple instances, e.g., from sequences representing multiple copies of an LTR retrotransposon in a host genome. In some embodiments, a 5’ or 3’ untranslated region for use in any of the systems described herein can be a molecular reconstruction based upon the aligned 5’ or 3’ untranslated region of multiple retrotransposons. Based on the Accession numbers provided herein, polypeptides or nucleic acid sequences can be aligned, e.g., by using routine sequence analysis tools as Basic Local Alignment Search Tool (BLAST) or CD-Search for conserved domain analysis. Molecular reconstructions can be created based upon sequence consensus, e.g. using approaches described in Ivics et al., Cell 1997, 501 – 510 ; Wagstaff et al., Molecular Biology and Evolution 2013, 88-99. In some embodiments, the retrotransposon from which the 5’ or 3’ untranslated region or polypeptide is derived is a young or a recently active mobile element, as assessed via phylogenetic methods such as those described in Boissinot et al., Molecular Biology and Evolution 2000, 915-928. Page 49 of 92 12806854v1Attorney Docket No.: 2017469-0044 System modifications of DNA Target Sites
[0160] In some embodiments, a system described herein is capable of producing an insertion into the target site of at least 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides (and optionally no more than 500, 400, 300, 200, or 100 nucleotides). In some embodiments, a system is capable of producing an insertion into the target site of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides (and optionally no more than 500, 400, 300, 200, or 100 nucleotides). In some embodiments, a system is capable of producing an insertion into the target site of at least 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5 or 10 kilobases (and optionally no more than 1, 5, 10, or 20 kilobases).
[0161] In some embodiments, an insertion as described herein increases or decreases expression (e.g. transcription or translation) of a gene. In some embodiments, an insertion increases or decreases expression (e.g. transcription or translation) of a gene by adding sequences in a promoter or enhancer, e.g. sequences that bind transcription factors. In some embodiments, an insertion alters translation of a gene (e.g. alters an amino acid sequence), inserts or disrupts a start or stop codon, or alters or fixes the translation frame of a gene. In some embodiments, an insertion results in the functional knockout of an endogenous gene by disruption of a coding or regulatory sequence. In some embodiments, an insertion alters splicing of a gene, e.g. by inserting or disrupting a splice acceptor or donor site. In some embodiments, an insertion alters transcript or protein half-life. In some embodiments, an insertion alters protein localization in the cell (e.g. from the cytoplasm to a mitochondria, from the cytoplasm into the extracellular space (e.g. adds a secretion tag)). In some embodiments, an insertion alters (e.g. improves) protein folding (e.g. to prevent accumulation of misfolded proteins). In some embodiments, an insertion alters, increases, or decreases the activity of a gene, e.g., a protein encoded by the gene.
[0162] In some embodiments, a system or method described herein results in “scarless” insertion of the heterologous object sequence, while in some embodiments, the target site can show deletions or duplications of endogenous DNA as a result of insertion of the heterologous sequence. The mechanisms of different retrotransposons could result in different patterns of duplications or deletions in the host genome occurring during retrotransposition at the target site. In some embodiments, the system results in a scarless insertion, with no duplications or deletions in the surrounding genomic DNA. In some embodiments, the system results in a deletion of less than 1, 2, 3, 4, 5, 10, 50, or 100 bp of genomic DNA upstream of the insertion. In some embodiments, the system results in a deletion of less than 1, 2, 3, 4, 5, 10, 50, or 100 bp of genomic DNA downstream of the insertion. In some embodiments, the system results in a duplication of less than 1, 2, 3, 4, 5, 10, 50, or 100 bp of genomic DNA upstream of the insertion. In some embodiments, the system results in a duplication of less than 1, 2, 3, 4, 5, 10, 50, or 100 bp of genomic DNA downstream of the insertion. Page 50 of 92 12806854v1Attorney Docket No.: 2017469-0044 Applications
[0163] By integrating coding genes into a RNA sequence template, the system can address therapeutic needs, for example, by providing expression of a therapeutic transgene in individuals with loss-of-function mutations, by replacing gain-of-function mutations with normal transgenes, by providing regulatory sequences to eliminate gain-of-function mutation expression, and / or by controlling the expression of operably linked genes, transgenes and systems thereof. In certain embodiments, the RNA sequence template encodes a promotor region specific to the therapeutic needs of the host cell, for example a tissue specific promotor or enhancer. In still other embodiments, a promotor can be operably linked to a coding sequence.
[0164] In some embodiments, the heterologous object sequence encodes an intracellular protein (e.g., a cytoplasmic protein, a nuclear protein, an organellar protein such as a mitochondrial protein or lysosomal protein, or a membrane protein). In some embodiments, the heterologous object sequence encodes a membrane protein, e.g., a membrane protein other than a CAR, and / or an endogenous human membrane protein. In some embodiments, the heterologous object sequence encodes an extracellular protein. In some embodiments, the heterologous object sequence encodes an enzyme, a structural protein, a signaling protein, a regulatory protein, a transport protein, a sensory protein, a motor protein, a defense protein, or a storage protein. Other exemplary proteins that may be encoded by a heterologous object sequence include, without limitation, a immune receptor protein, e.g. a synthetic immune receptor protein such as a chimeric antigen receptor protein (CAR), a T cell receptor, a B cell receptor, or an antibody.
[0165] Systems described herein may also be used to modify a plant or a plant part (e.g., leaves, roots, flowers, fruits, or seeds), e.g., to increase the fitness of a plant. Administration
[0166] The composition and systems described herein may be used in vitro or in vivo. In some embodiments the system or components of the system are delivered to cells (e.g., mammalian cells, e.g., human cells), e.g., in vitro or in vivo. In some embodiments, the cells are eukaryotic cells, e.g., cells of a multicellular organism, e.g., an animal, e.g., a mammal (e.g., human, swine, bovine) a bird (e.g., poultry, such as chicken, turkey, or duck), or a fish. In some embodiments, the cells are non-human animal cells (e.g., a laboratory animal, a livestock animal, or a companion animal). In some embodiments, the cell is a stem cell (e.g., a hematopoietic stem cell), a fibroblast, or a T cell. In some embodiments, the cell is a non-dividing cell, e.g., a non-dividing fibroblast or non-dividing T cell. In some embodiments, the cell is an HSC and p53 is not upregulated or is upregulated by less than 10%, 5%, 2%, or 1%, e.g., as determined according to the method described in Example 30 of PCT Application No. PCT / US2019 / 048607, incorporated herein by reference in its entirety. In some embodiment, a system is Page 51 of 92 12806854v1Attorney Docket No.: 2017469-0044 used to make an edit in primary cells. In some embodiments, the system is used to make an edit in a cell that is not immortalized. In some embodiments, the system is used to make an edit in a cell that is euploid.
[0167] The components of the system may, in some instances, be delivered in the form of polypeptide, nucleic acid (e.g., DNA, RNA), and combinations thereof.
[0168] For instance, delivery can use any of the following combinations for delivering one or more retrotransposon proteins (e.g., an integrase, structural polypeptide domain, and / or reverse transcriptase polypeptide domain, e.g., as described herein) (e.g., as DNA encoding the retrotransposon protein, as RNA encoding the integrase protein, or as the protein itself) and the template RNA (e.g., as DNA encoding the RNA, or as RNA): 1. DNA encoding polypeptide driver + DNA encoding template RNA 2. RNA encoding polypeptide driver + DNA encoding template RNA 3. DNA encoding polypeptide driver + template RNA 4. RNA encoding polypeptide driver + template RNA 5. Polypeptide driver + DNA encoding template RNA 6. Polypeptide driver + template RNA 7. virus encoding polypeptide driver + template virus 8. virus encoding polypeptide driver + DNA encoding template RNA 9. virus encoding polypeptide driver + template RNA 10. DNA encoding polypeptide driver + template virus 11. RNA encoding polypeptide driver + template virus 12. Polypeptide driver + template virus
[0169] In some embodiments, the ratio of the construct delivering the polypeptide driver and the template RNA is between 10:1 and 1:10 (e.g., between 10:5 to 10:1, 10:5 to 10:2, 10:5 to 10:1, 5:1 to 2:1, 5:1 to 1:1, 4:1 to 2:1, 4:1 to 1:1, 3:1 to 2:1, 3:1 to 1:1, 2:1 to 1:1, 1:1 to 1:2, 1:1 to 1:3, 1:2 to 1:3, 1:1 to 1:4, 1:2 to 1:4, 1:1 to 1:5, 1:2 to 1:5, 1:1 to 1:10, 1:2 to 1:10, or 1:5 to 1:10). In certain embodiments, the ratio of the construct delivering the polypeptide driver and the template RNA is 1:1.
[0170] In one embodiment, the system and / or components of the system are delivered as nucleic acid. For example, the polypeptide may be delivered in the form of a DNA or RNA encoding the polypeptide, and the template RNA may be delivered in the form of RNA or its complementary DNA to be transcribed into RNA. In some embodiments the system or components of the system are delivered on 1, 2, 3, 4, or more distinct nucleic acid molecules. In some embodiments the system or components of the system are delivered as a combination of DNA and RNA. In some embodiments the system or components of the system are delivered as a combination of DNA and protein. In some embodiments the Page 52 of 92 12806854v1Attorney Docket No.: 2017469-0044 system or components of the system are delivered as a combination of RNA and protein. In some embodiments the genome editor polypeptide is delivered as a protein.
[0171] In some embodiments the system or components of the system are delivered to cells, e.g. mammalian cells or human cells, using a vector. The vector may be, e.g., a plasmid or a virus. In some embodiments delivery is in vivo, in vitro, ex vivo, or in situ. In some embodiments the virus is an adeno associated virus (AAV), a lentivirus, an adenovirus. In some embodiments the system or components of the system are delivered to cells with a viral-like particle or a virosome. In some embodiments the delivery uses more than one virus, viral-like particle or virosome.
[0172] In one embodiment, the compositions and systems described herein can be formulated in liposomes or other similar vesicles. Liposomes are spherical vesicle structures composed of a uni- or multilamellar lipid bilayer surrounding internal aqueous compartments and a relatively impermeable outer lipophilic phospholipid bilayer. Liposomes may be anionic, neutral or cationic. Liposomes are biocompatible, nontoxic, can deliver both hydrophilic and lipophilic drug molecules, protect their cargo from degradation by plasma enzymes, and transport their load across biological membranes and the blood brain barrier (BBB) (see, e.g., Spuch and Navarro, Journal of Drug Delivery, vol.2011, Article ID 469679, 12 pages, 2011. doi:10.1155 / 2011 / 469679 for review).
[0173] Vesicles can be made from several different types of lipids; however, phospholipids are most commonly used to generate liposomes as drug carriers. Methods for preparation of multilamellar vesicle lipids are known in the art (see for example U.S. Pat. No.6,693,086, the teachings of which relating to multilamellar vesicle lipid preparation are incorporated herein by reference). Although vesicle formation can be spontaneous when a lipid film is mixed with an aqueous solution, it can also be expedited by applying force in the form of shaking by using a homogenizer, sonicator, or an extrusion apparatus (see, e.g., Spuch and Navarro, Journal of Drug Delivery, vol.2011, Article ID 469679, 12 pages, 2011. doi:10.1155 / 2011 / 469679 for review). Extruded lipids can be prepared by extruding through filters of decreasing size, as described in Templeton et al., Nature Biotech, 15:647-652, 1997, the teachings of which relating to extruded lipid preparation are incorporated herein by reference.
[0174] Lipid nanoparticles are another example of a carrier that provides a biocompatible and biodegradable delivery system for the pharmaceutical compositions described herein. Nanostructured lipid carriers (NLCs) are modified solid lipid nanoparticles (SLNs) that retain the characteristics of the SLN, improve drug stability and loading capacity, and prevent drug leakage. Polymer nanoparticles (PNPs) are an important component of drug delivery. These nanoparticles can effectively direct drug delivery to specific targets and improve drug stability and controlled drug release. Lipid–polymer nanoparticles (PLNs), a new type of carrier that combines liposomes and polymers, may also be employed. These nanoparticles possess the complementary advantages of PNPs and liposomes. A PLN is Page 53 of 92 12806854v1Attorney Docket No.: 2017469-0044 composed of a core–shell structure; the polymer core provides a stable structure, and the phospholipid shell offers good biocompatibility. As such, the two components increase the drug encapsulation efficiency rate, facilitate surface modification, and prevent leakage of water-soluble drugs. For a review, see, e.g., Li et al.2017, Nanomaterials 7, 122; doi:10.3390 / nano7060122.
[0175] Exosomes can also be used as drug delivery vehicles for the compositions and systems described herein. For a review, see Ha et al. July 2016. Acta Pharmaceutica Sinica B. Volume 6, Issue 4, Pages 287-296; doi.org / 10.1016 / j.apsb.2016.02.001.
[0176] Fusosomes interact and fuse with target cells, and thus can be used as delivery vehicles for a variety of molecules. They generally consist of a bilayer of amphipathic lipids enclosing a lumen or cavity and a fusogen that interacts with the amphipathic lipid bilayer. The fusogen component has been shown to be engineerable in order to confer target cell specificity for the fusion and payload delivery, allowing the creation of delivery vehicles with programmable cell specificity (see, for example, the relating to fusosome design, preparation, and usage in PCT Publication No. WO / 2020014209, incorporated herein by reference in its entirety).
[0177] A system can be introduced into cells, tissues and multicellular organisms. In some embodiments the system or components of the system are delivered to the cells via mechanical means or physical means.
[0178] Formulation of protein therapeutics is described in Meyer (Ed.), Therapeutic Protein Drug Products: Practical Approaches to formulation in the Laboratory, Manufacturing, and the Clinic, Woodhead Publishing Series (2012). Lipid Nanoparticles
[0179] The methods and systems provided by the invention, may employ any suitable carrier or delivery modality, including, in certain embodiments, lipid nanoparticles (LNPs). Lipid nanoparticles, in some embodiments, comprise one or more ionic lipids, such as non-cationic lipids (e.g., neutral or anionic, or zwitterionic lipids); one or more conjugated lipids (such as PEG-conjugated lipids or lipids conjugated to polymers described in Table 5 of WO2019217941; incorporated herein by reference in its entirety); one or more sterols (e.g., cholesterol); and, optionally, one or more targeting molecules (e.g., conjugated receptors, receptor ligands, antibodies); or combinations of the foregoing.
[0180] Lipids that can be used in nanoparticle formations (e.g., lipid nanoparticles) include, for example those described in Table 4 of WO2019217941, which is incorporated by reference—e.g., a lipid- containing nanoparticle can comprise one or more of the lipids in table 4 of WO2019217941. Lipid nanoparticles can include additional elements, such as polymers, such as the polymers described in table 5 of WO2019217941, incorporated by reference. Page 54 of 92 12806854v1Attorney Docket No.: 2017469-0044
[0181] In some embodiments, conjugated lipids, when present, can include one or more of PEG- diacylglycerol (DAG) (such as l-(monomethoxy-polyethyleneglycol)-2,3-dimyristoylglycerol (PEG- DMG)), PEG-dialkyloxypropyl (DAA), PEG-phospholipid, PEG-ceramide (Cer), a pegylated phosphatidylethanoloamine (PEG-PE), PEG succinate diacylglycerol (PEGS-DAG) (such as 4-0-(2',3'- di(tetradecanoyloxy)propyl-l-0-(w-methoxy(polyethoxy)ethyl) butanedioate (PEG-S-DMG)), PEG dialkoxypropylcarbam, N-(carbonyl-methoxypoly ethylene glycol 2000)-1,2-distearoyl-sn-glycero-3- phosphoethanolamine sodium salt, and those described in Table 2 of WO2019051289 (incorporated by reference), and combinations of the foregoing.
[0182] In some embodiments, sterols that can be incorporated into lipid nanoparticles include one or more of cholesterol or cholesterol derivatives, such as those in W02009 / 127060 or US2010 / 0130588, which are incorporated by reference. Additional exemplary sterols include phytosterols, including those described in Eygeris et al (2020), dx.doi.org / 10.1021 / acs.nanolett.0c01386, incorporated herein by reference.
[0183] In some embodiments, the lipid particle comprises an ionizable lipid, a non-cationic lipid, a conjugated lipid that inhibits aggregation of particles, and a sterol. The amounts of these components can be varied independently and to achieve desired properties. For example, in some embodiments, the lipid nanoparticle comprises an ionizable lipid is in an amount from about 20 mol % to about 90 mol % of the total lipids (in other embodiments it may be 20-70% (mol), 30-60% (mol) or 40-50% (mol); about 50 mol % to about 90 mol % of the total lipid present in the lipid nanoparticle), a non-cationic lipid in an amount from about 5 mol % to about 30 mol % of the total lipids, a conjugated lipid in an amount from about 0.5 mol % to about 20 mol % of the total lipids, and a sterol in an amount from about 20 mol % to about 50 mol % of the total lipids. The ratio of total lipid to nucleic acid (e.g., encoding the polypeptide or template nucleic acid) can be varied as desired. For example, the total lipid to nucleic acid (mass or weight) ratio can be from about 10: 1 to about 30: 1.
[0184] In some embodiments, an ionizable lipid may be a cationic lipid, an ionizable cationic lipid, e.g., a cationic lipid that can exist in a positively charged or neutral form depending on pH, or an amine- containing lipid that can be readily protonated. In some embodiments, the cationic lipid is a lipid capable of being positively charged, e.g., under physiological conditions. Exemplary cationic lipids include one or more amine group(s) which bear the positive charge. In some embodiments, the lipid particle comprises a cationic lipid in formulation with one or more of neutral lipids, ionizable amine-containing lipids, biodegradable alkyn lipids, steroids, phospholipids including polyunsaturated lipids, structural lipids (e.g., sterols), PEG, cholesterol and polymer conjugated lipids. In some embodiments, the cationic lipid may be an ionizable cationic lipid. An exemplary cationic lipid as disclosed herein may have an effective pKa over 6.0. In embodiments, a lipid nanoparticle may comprise a second cationic lipid having a different Page 55 of 92 12806854v1Attorney Docket No.: 2017469-0044 effective pKa (e.g., greater than the first effective pKa), than the first cationic lipid. A lipid nanoparticle may comprise between 40 and 60 mol percent of a cationic lipid, a neutral lipid, a steroid, a polymer conjugated lipid, and a therapeutic agent, e.g., a nucleic acid (e.g., RNA) described herein (e.g., a template nucleic acid or a nucleic acid encoding a system), encapsulated within or associated with the lipid nanoparticle. In some embodiments, the nucleic acid is co-formulated with the cationic lipid. The nucleic acid may be adsorbed to the surface of an LNP, e.g., an LNP comprising a cationic lipid. In some embodiments, the nucleic acid may be encapsulated in an LNP, e.g., an LNP comprising a cationic lipid. In some embodiments, the lipid nanoparticle may comprise a targeting moiety, e.g., coated with a targeting agent. In embodiments, the LNP formulation is biodegradable. In some embodiments, a lipid nanoparticle comprising one or more lipid described herein, e.g., Formula (i), (ii), (ii), (vii) and / or (ix) encapsulates at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98% or 100% of an RNA molecule, e.g., template RNA and / or a mRNA encoding the polypeptide.
[0185] In some embodiments, the lipid to nucleic acid ratio (mass / mass ratio; w / w ratio) can be in the range of from about 1 : 1 to about 25: 1, from about 10: 1 to about 14: 1, from about 3 : 1 to about 15: 1, from about 4: 1 to about 10: 1, from about 5: 1 to about 9: 1, or about 6: 1 to about 9: 1. The amounts of lipids and nucleic acid can be adjusted to provide a desired N / P ratio, for example, N / P ratio of 3, 4, 5, 6, 7, 8, 9, 10 or higher. Generally, the lipid nanoparticle formulation’s overall lipid content can range from about 5 mg / ml to about 30 mg / mL.
[0186] Exemplary ionizable lipids that can be used in lipid nanoparticle formulations include, without limitation, those listed in Table 1 of WO2019051289, incorporated herein by reference. Additional exemplary lipids include, without limitation, one or more of the following formulae: X of US2016 / 0311759; I of US20150376115 or in US2016 / 0376224; I, II or III of US20160151284; I, IA, II, or IIA of US20170210967; I-c of US20150140070; A of US2013 / 0178541; I of US2013 / 0303587 or US2013 / 0123338; I of US2015 / 0141678; II, III, IV, or V of US2015 / 0239926; I of US2017 / 0119904; I or II of WO2017 / 117528; A of US2012 / 0149894; A of US2015 / 0057373; A of WO2013 / 116126; A of US2013 / 0090372; A of US2013 / 0274523; A of US2013 / 0274504; A of US2013 / 0053572; A of W02013 / 016058; A of W02012 / 162210; I of US2008 / 042973; I, II, III, or IV of US2012 / 01287670; I or II of US2014 / 0200257; I, II, or III of US2015 / 0203446; I or III of US2015 / 0005363; I, IA, IB, IC, ID, II, IIA, IIB, IIC, IID, or III-XXIV of US2014 / 0308304; of US2013 / 0338210; I, II, III, or IV of W02009 / 132131; A of US2012 / 01011478; I or XXXV of US2012 / 0027796; XIV or XVII of US2012 / 0058144; of US2013 / 0323269; I of US2011 / 0117125; I, II, or III of US2011 / 0256175; I, II, III, IV, V, VI, VII, VIII, IX, X, XI, XII of US2012 / 0202871; I, II, III, IV, V, VI, VII, VIII, X, XII, XIII, XIV, XV, or XVI of US2011 / 0076335; I or II of US2006 / 008378; I of US2013 / 0123338; I or X-A-Y-Z of Page 56 of 92 12806854v1Attorney Docket No.: 2017469-0044 US2015 / 0064242; XVI, XVII, or XVIII of US2013 / 0022649; I, II, or III of US2013 / 0116307; I, II, or III of US2013 / 0116307; I or II of US2010 / 0062967; I-X of US2013 / 0189351; I of US2014 / 0039032; V of US2018 / 0028664; I of US2016 / 0317458; I of US2013 / 0195920; 5, 6, or 10 of US10,221,127; III-3 of WO2018 / 081480; I-5 or I-8 of WO2020 / 081938; 18 or 25 of US9,867,888; A of US2019 / 0136231; II of WO2020 / 219876; 1 of US2012 / 0027803; OF-02 of US2019 / 0240349; 23 of US10,086,013; cKK- E12 / A6 of Miao et al (2020); C12-200 of WO2010 / 053572; 7C1 of Dahlman et al (2017); 304-O13 or 503-O13 of Whitehead et al; TS-P4C2 of US9,708,628; I of WO2020 / 106946; I of WO2020 / 106946.
[0187] In some embodiments, the ionizable lipid is MC3 (6Z,9Z,28Z,3 lZ)-heptatriaconta- 6,9,28,3 l- tetraen-l9-yl-4-(dimethylamino) butanoate (DLin-MC3-DMA or MC3), e.g., as described in Example 9 of WO2019051289A9 (incorporated by reference herein in its entirety). In some embodiments, the ionizable lipid is the lipid ATX-002, e.g., as described in Example 10 of WO2019051289A9 (incorporated by reference herein in its entirety). In some embodiments, the ionizable lipid is (l3Z,l6Z)-A,A-dimethyl- 3- nonyldocosa-l3, l6-dien-l-amine (Compound 32), e.g., as described in Example 11 of WO2019051289A9 (incorporated by reference herein in its entirety). In some embodiments, the ionizable lipid is Compound 6 or Compound 22, e.g., as described in Example 12 of WO2019051289A9 (incorporated by reference herein in its entirety). In some embodiments, the ionizable lipid is heptadecan- 9-yl 8-((2-hydroxyethyl)(6-oxo-6-(undecyloxy)hexyl)amino)octanoate (SM-102); e.g., as described in Example 1 of US9,867,888(incorporated by reference herein in its entirety). In some embodiments, the ionizable lipid is 9Z,12Z)-3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3- (diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate (LP01) e.g., as synthesized in Example 13 of WO2015 / 095340(incorporated by reference herein in its entirety). In some embodiments, the ionizable lipid is Di((Z)-non-2-en-1-yl) 9-((4- dimethylamino)butanoyl)oxy)heptadecanedioate (L319), e.g., as synthesized in Example 7, 8, or 9 of US2012 / 0027803(incorporated by reference herein in its entirety). In some embodiments, the ionizable lipid is 1,1'-((2-(4-(2-((2-(Bis(2-hydroxydodecyl)amino)ethyl)(2-hydroxydodecyl) amino)ethyl)piperazin- 1-yl)ethyl)azanediyl)bis(dodecan-2-ol) (C12-200), e.g., as synthesized in Examples 14 and 16 of WO2010 / 053572(incorporated by reference herein in its entirety). In some embodiments, the ionizable lipid is; Imidazole cholesterol ester (ICE) lipid (3S, 10R, 13R, 17R)-10, 13-dimethyl-17- ((R)-6- methylheptan-2-yl)-2, 3, 4, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17-tetradecahydro-lH- cyclopenta[a]phenanthren-3-yl 3-(1H-imidazol-4-yl)propanoate, e.g., Structure (I) from WO2020 / 106946 (incorporated by reference herein in its entirety).
[0188] Some non-limiting example of lipid compounds that may be used (e.g., in combination with other lipid components) to form lipid nanoparticles for the delivery of compositions described herein, e.g., Page 57 of 92 12806854v1Attorney Docket No.: 2017469-0044 nucleic acid (e.g., RNA) described herein (e.g., a template nucleic acid or a nucleic acid encoding a system) includes, (i)
[0189] a composition(ii)
[0190] described
[0191] described herein to the liver and / or hepatocyte cells.Page 58 of 92 12806854v1Attorney Docket No.: 2017469-0044
[0192] Formula (v) is used to deliver a composition
[0193] used to deliver a composition
[0194] a composition described herein to the liver and / or hepatocyte cells. Page 59 of 92 12806854v1Attorney Docket No.: 2017469-0044
[0195] described herein to the liver and / or hepatocyte cells.X1is O, NR1, or a direct bond, X2is C2-5 alkylene, X3is C(=0) or a direct bond, R1is H or Me, R3is Ci-3 alkyl, R2is Ci-3 alkyl, or R2taken together with the nitrogen atom to which it is attached and 1-3 carbon atoms of X2form a 4-, 5-, or 6-membered ring, or X1is NR1, R1and R2taken together with the nitrogen atoms to which they are attached form a 5- or 6-membered ring, or R2taken together with R3and the nitrogen atom to which they are attached form a 5-, 6-, or 7-membered ring, Y1is C2-12 alkylene, Y2is selected from ,n or aor absent, provided that if Z1is a direct bond, Z2is absent; R5is C5-9 alkyl or C6-10 alkoxy, R6is C5-9 alkyl or C6-10 alkoxy, W is methylene or a direct bond, and R7is H or Me, or a salt thereof, provided that if R3and R2are C2 alkyls, X1is O, X2is linear C3 alkylene, X3is C(=0), Y1is linear Ce alkylene, (Y2)n-R4isPage 60 of 92 12806854v1Attorney Docket No.: 2017469-0044 , R4is linear C5 alkyl, Z1is C2 alkylene, Z2is absent, W is methylene, and R7is H, then R5and R6are not Cx alkoxy.
[0196] In some embodiments an LNP comprising Formula (xii) is used to deliver a composition described herein to the liver and / or hepatocyte cells.
[0197] Page 61 of 92 12806854v1Attorney Docket No.: 2017469-0044
[0198] In some embodiments an LNP comprises a compound of Formula (xiii) and a compound of Formula (xiv).
[0199]
[0200] an LNP comprising a formulation of Formula (xvi) is used to deliver acomposition described herein to the lung endothelial cells.Page 62 of 92 12806854v1Attorney Docket No.: 2017469-0044 (b)
[0201] In some embodiments, a lipid compound used to form lipid nanoparticles for the delivery of compositions described herein, e.g., nucleic acid (e.g., RNA) described herein (e.g., a template nucleic acid or a nucleic acid encoding a system) is made by one of the following reactions: (b)
[0202] Exemplary non-cationic lipids include, but are not limited to, distearoyl-sn-glycero- phosphoethanolamine, distearoylphosphatidylcholine (DSPC), dioleoylphosphatidylcholine (DOPC), Page 63 of 92 12806854v1Attorney Docket No.: 2017469-0044 dipalmitoylphosphatidylcholine (DPPC), dioleoylphosphatidylglycerol (DOPG), 1,2-dioleoyl-sn-glycero- 3-phosphoethanolamine (DOPE), dipalmitoylphosphatidylglycerol (DPPG), dioleoyl- phosphatidylethanolamine (DOPE), palmitoyloleoylphosphatidylcholine (POPC), palmitoyloleoylphosphatidylethanolamine (POPE), dioleoyl-phosphatidylethanolamine 4-(N- maleimidomethyl)-cyclohexane- 1 - carboxylate (DOPE-mal), dipalmitoyl phosphatidyl ethanolamine (DPPE), dimyristoylphosphoethanolamine (DMPE), distearoyl-phosphatidyl-ethanolamine (DSPE), monomethyl-phosphatidylethanolamine (such as 16-O-monomethyl PE), dimethyl- phosphatidylethanolamine (such as 16-O-dimethyl PE), l8-l-trans PE, l-stearoyl-2-oleoyl- phosphatidyethanolamine (SOPE), hydrogenated soy phosphatidylcholine (HSPC), egg phosphatidylcholine (EPC), dioleoylphosphatidylserine (DOPS), sphingomyelin (SM), dimyristoyl phosphatidylcholine (DMPC), dimyristoyl phosphatidylglycerol (DMPG), distearoylphosphatidylglycerol (DSPG), dierucoylphosphatidylcholine (DEPC), palmitoyloleyolphosphatidylglycerol (POPG), dielaidoyl-phosphatidylethanolamine (DEPE), lecithin, phosphatidylethanolamine, lysolecithin, lysophosphatidylethanolamine, phosphatidylserine, phosphatidylinositol, sphingomyelin, egg sphingomyelin (ESM), cephalin, cardiolipin, phosphatidicacid,cerebrosides, dicetylphosphate, lysophosphatidylcholine, dilinoleoylphosphatidylcholine, or mixtures thereof. It is understood that other diacylphosphatidylcholine and diacylphosphatidylethanolamine phospholipids can also be used. The acyl groups in these lipids are preferably acyl groups derived from fatty acids having C10-C24 carbon chains, e.g., lauroyl, myristoyl, paimitoyl, stearoyl, or oleoyl. Additional exemplary lipids, in certain embodiments, include, without limitation, those described in Kim et al. (2020) dx.doi.org / 10.1021 / acs.nanolett.0c01386, incorporated herein by reference. Such lipids include, in some embodiments, plant lipids found to improve liver transfection with mRNA (e.g., DGTS In some embodiments, the non-cationic lipid may have the following structure,
[0203] Other examples of non-cationic lipids suitable for use in the lipid nanopartieles include, without limitation, nonphosphorous lipids such as, e.g., stearylamine, dodeeylamine, hexadecylamine, acetyl palmitate, glycerol ricinoleate, hexadecyl stereate, isopropyl myristate, amphoteric acrylic polymers, triethanolamine-lauryl sulfate, alkyl-aryl sulfate polyethyloxylated fatty acid amides, dioctadecyl dimethyl ammonium bromide, ceramide, sphingomyelin, and the like. Other non-cationic Page 64 of 92 12806854v1Attorney Docket No.: 2017469-0044 lipids are described in WO2017 / 099823 or US patent publication US2018 / 0028664, the contents of which is incorporated herein by reference in their entirety.
[0204] In some embodiments, the non-cationic lipid is oleic acid or a compound of Formula I, II, or IV of US2018 / 0028664, incorporated herein by reference in its entirety. The non-cationic lipid can comprise, for example, 0-30% (mol) of the total lipid present in the lipid nanoparticle. In some embodiments, the non-cationic lipid content is 5-20% (mol) or 10-15% (mol) of the total lipid present in the lipid nanoparticle. In embodiments, the molar ratio of ionizable lipid to the neutral lipid ranges from about 2:1 to about 8:1 (e.g., about 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, or 8:1).
[0205] In some embodiments, the lipid nanoparticles do not comprise any phospholipids.
[0206] In some aspects, the lipid nanoparticle can further comprise a component, such as a sterol, to provide membrane integrity. One exemplary sterol that can be used in the lipid nanoparticle is cholesterol and derivatives thereof. Non-limiting examples of cholesterol derivatives include polar analogues such as 5a-choiestanol, 53-coprostanol, choiesteryl-(2,-hydroxy)-ethyl ether, choiesteryl-(4'- hydroxy)-butyl ether, and 6-ketocholestanol; non-polar analogues such as 5a-cholestane, cholestenone, 5a-cholestanone, 5p- cholestanone, and cholesteryl decanoate; and mixtures thereof. In some embodiments, the cholesterol derivative is a polar analogue, e.g., choiesteryl-(4 '-hydroxy)-buty1 ether. Exemplary cholesterol derivatives are described in PCT publication W02009 / 127060 and US patent publication US2010 / 0130588, each of which is incorporated herein by reference in its entirety.
[0207] In some embodiments, the component providing membrane integrity, such as a sterol, can comprise 0-50% (mol) (e.g., 0-10%, 10-20%, 20-30%, 30-40%, or 40-50%) of the total lipid present in the lipid nanoparticle. In some embodiments, such a component is 20-50% (mol) 30-40% (mol) of the total lipid content of the lipid nanoparticle.
[0208] In some embodiments, the lipid nanoparticle can comprise a polyethylene glycol (PEG) or a conjugated lipid molecule. Generally, these are used to inhibit aggregation of lipid nanoparticles and / or provide steric stabilization. Exemplary conjugated lipids include, but are not limited to, PEG-lipid conjugates, polyoxazoline (POZ)-lipid conjugates, polyamide-lipid conjugates (such as ATTA-lipid conjugates), cationic-polymer lipid (CPL) conjugates, and mixtures thereof. In some embodiments, the conjugated lipid molecule is a PEG-lipid conjugate, for example, a (methoxy polyethylene glycol)- conjugated lipid.
[0209] Exemplary PEG-lipid conjugates include, but are not limited to, PEG-diacylglycerol (DAG) (such as l-(monomethoxy-polyethyleneglycol)-2,3-dimyristoylglycerol (PEG-DMG)), PEG- dialkyloxypropyl (DAA), PEG-phospholipid, PEG-ceramide (Cer), a pegylated phosphatidylethanoloamine (PEG-PE), 1,2-dimyristoyl-sn-glycerol, methoxypoly ethylene glycol (DMG- PEG-2K), PEG succinate diacylglycerol (PEGS-DAG) (such as 4-0-(2',3'-di(tetradecanoyloxy)propyl-l-0- Page 65 of 92 12806854v1Attorney Docket No.: 2017469-0044 (w-methoxy(polyethoxy)ethyl) butanedioate (PEG-S-DMG)), PEG dialkoxypropylcarbam, N-(carbonyl- methoxypolyethylene glycol 2000)-l,2-distearoyl-sn-glycero-3-phosphoethanolamine sodium salt, or a mixture thereof. Additional exemplary PEG-lipid conjugates are described, for example, in US5,885,6l3, US6,287,59l,
[0210] US2003 / 0077829, US2003 / 0077829, US2005 / 0175682, US2008 / 0020058, US2011 / 0117125, US2010 / 0130588, US2016 / 0376224, US2017 / 0119904, and US / 099823, the contents of all of which are incorporated herein by reference in their entirety. In some embodiments, a PEG-lipid is a compound of Formula III, III-a-I, III-a-2, III-b-1, III-b-2, or V of US2018 / 0028664, the content of which is incorporated herein by reference in its entirety. In some embodiments, a PEG-lipid is of Formula II of US20150376115 or US2016 / 0376224, the content of both of which is incorporated herein by reference in its entirety. In some embodiments, the PEG-DAA conjugate can be, for example, PEG-dilauryloxypropyl, PEG- dimyristyloxypropyl, PEG-dipalmityloxypropyl, or PEG-distearyloxypropyl. The PEG-lipid can be one or more of PEG-DMG, PEG-dilaurylglycerol, PEG-dipalmitoylglycerol, PEG- disterylglycerol, PEG- dilaurylglycamide, PEG-dimyristylglycamide, PEG- dipalmitoylglycamide, PEG-disterylglycamide, PEG-cholesterol (l-[8'-(Cholest-5-en-3[beta]- oxy)carboxamido-3',6'-dioxaoctanyl] carbamoyl-[omega]- methyl-poly(ethylene glycol), PEG- DMB (3,4-Ditetradecoxylbenzyl- [omega]-methyl-poly(ethylene glycol) ether), and 1,2- dimyristoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000]. In some embodiments, the PEG-lipid comprises PEG-DMG, 1,2- dimyristoyl-sn-glycero- 3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000]. In some embodiments, the PEG-lipid comprises a structure selected from: , ,Page 66 of 92 12806854v1Attorney Docket No.: 2017469-0044 , .
[0211] In some embodiments, lipids conjugated with a molecule other than a PEG can also be used in place of PEG-lipid. For example, polyoxazoline (POZ)-lipid conjugates, polyamide-lipid conjugates (such as ATTA-lipid conjugates), and cationic-polymer lipid (GPL) conjugates can be used in place of or in addition to the PEG-lipid.
[0212] Exemplary conjugated lipids, i.e., PEG-lipids, (POZ)-lipid conjugates, ATTA-lipid conjugates and cationic polymer-lipids are described in the PCT and LIS patent applications listed in Table 2 of WO2019051289A9 and in WO2020106946A1, the contents of all of which are incorporated herein by reference in their entirety.
[0213] In some embodiments an LNP comprises a compound of Formula (xix), a compound of Formula (xxi) and a compound of Formula (xxv). In some embodiments a LNP comprising a formulation of Formula (xix), Formula (xxi) and Formula (xxv) is used to deliver a composition described herein to the lung or pulmonary cells.
[0214] In some embodiments, the PEG or the conjugated lipid can comprise 0-20% (mol) of the total lipid present in the lipid nanoparticle. In some embodiments, PEG or the conjugated lipid content is 0.5- 10% or 2-5% (mol) of the total lipid present in the lipid nanoparticle. Molar ratios of the ionizable lipid, non-cationic-lipid, sterol, and PEG / conjugated lipid can be varied as needed. For example, the lipid particle can comprise 30-70% ionizable lipid by mole or by total weight of the composition, 0-60% cholesterol by mole or by total weight of the composition, 0-30% non-cationic-lipid by mole or by total weight of the composition and 1-10% conjugated lipid by mole or by total weight of the composition. Preferably, the composition comprises 30-40% ionizable lipid by mole or by total weight of the composition, 40-50% cholesterol by mole or by total weight of the composition, and 10- 20% non- cationic-lipid by mole or by total weight of the composition. In some other embodiments, the composition Page 67 of 92 12806854v1Attorney Docket No.: 2017469-0044 is 50-75% ionizable lipid by mole or by total weight of the composition, 20-40% cholesterol by mole or by total weight of the composition, and 5 to 10% non-cationic-lipid, by mole or by total weight of the composition and 1-10% conjugated lipid by mole or by total weight of the composition. The composition may contain 60-70% ionizable lipid by mole or by total weight of the composition, 25-35% cholesterol by mole or by total weight of the composition, and 5-10% non-cationic-lipid by mole or by total weight of the composition. The composition may also contain up to 90% ionizable lipid by mole or by total weight of the composition and 2 to 15% non-cationic lipid by mole or by total weight of the composition. The formulation may also be a lipid nanoparticle formulation, for example comprising 8-30% ionizable lipid by mole or by total weight of the composition, 5-30% non- cationic lipid by mole or by total weight of the composition, and 0-20% cholesterol by mole or by total weight of the composition; 4-25% ionizable lipid by mole or by total weight of the composition, 4-25% non-cationic lipid by mole or by total weight of the composition, 2 to 25% cholesterol by mole or by total weight of the composition, 10 to 35% conjugate lipid by mole or by total weight of the composition, and 5% cholesterol by mole or by total weight of the composition; or 2-30% ionizable lipid by mole or by total weight of the composition, 2-30% non-cationic lipid by mole or by total weight of the composition, 1 to 15% cholesterol by mole or by total weight of the composition, 2 to 35% conjugate lipid by mole or by total weight of the composition, and 1-20% cholesterol by mole or by total weight of the composition; or even up to 90% ionizable lipid by mole or by total weight of the composition and 2-10% non-cationic lipids by mole or by total weight of the composition, or even 100% cationic lipid by mole or by total weight of the composition. In some embodiments, the lipid particle formulation comprises ionizable lipid, phospholipid, cholesterol and a PEG-ylated lipid in a molar ratio of 50: 10:38.5: 1.5. In some other embodiments, the lipid particle formulation comprises ionizable lipid, cholesterol and a PEG-ylated lipid in a molar ratio of 60:38.5: 1.5.
[0215] In some embodiments, the lipid particle comprises ionizable lipid, non-cationic lipid (e.g. phospholipid), a sterol (e.g., cholesterol) and a PEG-ylated lipid, where the molar ratio of lipids ranges from 20 to 70 mole percent for the ionizable lipid, with a target of 40-60, the mole percent of non-cationic lipid ranges from 0 to 30, with a target of 0 to 15, the mole percent of sterol ranges from 20 to 70, with a target of 30 to 50, and the mole percent of PEG-ylated lipid ranges from 1 to 6, with a target of 2 to 5.
[0216] In some embodiments, the lipid particle comprises ionizable lipid / non-cationic- lipid / sterol / conjugated lipid at a molar ratio of 50:10:38.5:1.5.
[0217] In an aspect, the disclosure provides a lipid nanoparticle formulation comprising phospholipids, lecithin, phosphatidylcholine and phosphatidylethanolamine.
[0218] In some embodiments, one or more additional compounds can also be included. Those compounds can be administered separately or the additional compounds can be included in the lipid nanoparticles of the invention. In other words, the lipid nanoparticles can contain other compounds in Page 68 of 92 12806854v1Attorney Docket No.: 2017469-0044 addition to the nucleic acid or at least a second nucleic acid, different than the first. Without limitations, other additional compounds can be selected from the group consisting of small or large organic or inorganic molecules, monosaccharides, disaccharides, trisaccharides, oligosaccharides, polysaccharides, peptides, proteins, peptide analogs and derivatives thereof, peptidomimetics, nucleic acids, nucleic acid analogs and derivatives, an extract made from biological materials, or any combinations thereof.
[0219] In some embodiments, a lipid nanoparticle (or a formulation comprising lipid nanoparticles) lacks reactive impurities (e.g., aldehydes or ketones), or comprises less than a preselected level of reactive impurities (e.g., aldehydes or ketones). While not wishing to be bound by theory, in some embodiments, a lipid reagent is used to make a lipid nanoparticle formulation, and the lipid reagent may comprise a contaminating reactive impurity (e.g., an aldehyde or ketone). A lipid regent may be selected for manufacturing based on having less than a preselected level of reactive impurities (e.g., aldehydes or ketones). Without wishing to be bound by theory, in some embodiments, aldehydes can cause modification and damage of RNA, e.g., cross-linking between bases and / or covalently conjugating lipid to RNA (e.g., forming lipid-RNA adducts). This may, in some instances, lead to failure of a reverse transcriptase reaction and / or incorporation of inappropriate bases, e.g., at the site(s) of lesion(s), e.g., a mutation in a newly synthesized target DNA.
[0220] In some embodiments, LNPs are directed to specific tissues by the addition of targeting domains. For example, biological ligands may be displayed on the surface of LNPs to enhance interaction with cells displaying cognate receptors, thus driving association with and cargo delivery to tissues wherein cells express the receptor. In some embodiments, the biological ligand may be a ligand that drives delivery to the liver, e.g., LNPs that display GalNAc result in delivery of nucleic acid cargo to hepatocytes that display asialoglycoprotein receptor (ASGPR). The work of Akinc et al. Mol Ther 18(7):1357-1364 (2010) teaches the conjugation of a trivalent GalNAc ligand to a PEG-lipid (GalNAc- PEG-DSG) to yield LNPs dependent on ASGPR for observable LNP cargo effect (see, e.g., Figure 6). Other ligand-displaying LNP formulations, e.g., incorporating folate, transferrin, or antibodies, are discussed in WO2017223135, which is incorporated herein by reference in its entirety, in addition to the references used therein, namely Kolhatkar et al., Curr Drug Discov Technol.20118:197-206; Musacchio and Torchilin, Front Biosci.201116:1388-1412; Yu et al., Mol Membr Biol.201027:286-298; Patil et al., Crit Rev Ther Drug Carrier Syst.200825:1-61 ; Benoit et al., Biomacromolecules.201112:2708-2714; Zhao et al., Expert Opin Drug Deliv.20085:309-319; Akinc et al., Mol Ther.201018:1357-1364; Srinivasan et al., Methods Mol Biol.2012820:105-116; Ben-Arie et al., Methods Mol Biol.2012 757:497-507; Peer 2010 J Control Release.20:63-68; Peer et al., Proc Natl Acad Sci U S A.2007 104:4095-4100; Kim et al., Methods Mol Biol.2011721:339-353; Subramanya et al., Mol Ther.2010 Page 69 of 92 12806854v1Attorney Docket No.: 2017469-0044 18:2028-2037; Song et al., Nat Biotechnol.200523:709-717; Peer et al., Science.2008319:627-630; and Peer and Lieberman, Gene Ther.201118:1127-1133.
[0221] In some embodiments, LNPs are selected for tissue-specific activity by the addition of a Selective ORgan Targeting (SORT) molecule to a formulation comprising traditional components, such as ionizable cationic lipids, amphipathic phospholipids, cholesterol and poly(ethylene glycol) (PEG) lipids. The teachings of Cheng et al. Nat Nanotechnol 15(4):313-320 (2020) demonstrate that the addition of a supplemental “SORT” component precisely alters the in vivo RNA delivery profile and mediates tissue- specific (e.g., lungs, liver, spleen) gene delivery and editing as a function of the percentage and biophysical property of the SORT molecule.
[0222] In some embodiments, the LNPs comprise biodegradable, ionizable lipids. In some embodiments, the LNPs comprise (9Z,12Z)-3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3- (diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate, also called 3-((4,4- bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl (9Z,12Z)- octadeca-9,12-dienoate) or another ionizable lipid. See, e.g., lipids of WO2019 / 067992, WO / 2017 / 173054, WO2015 / 095340, and WO2014 / 136086, as well as references provided therein. In some embodiments, the term cationic and ionizable in the context of LNP lipids is interchangeable, e.g., wherein ionizable lipids are cationic depending on the pH.
[0223] In some embodiments, multiple components of a system may be prepared as a single LNP formulation, e.g., an LNP formulation comprises mRNA encoding for the polypeptide and an RNA template. Ratios of nucleic acid components may be varied in order to maximize the properties of a therapeutic. In some embodiments, the ratio of RNA template to mRNA encoding a polypeptide is about 1:1 to 100:1, e.g., about 1:1 to 20:1, about 20:1 to 40:1, about 40:1 to 60:1, about 60:1 to 80:1, or about 80:1 to 100:1, by molar ratio. In other embodiments, a system of multiple nucleic acids may be prepared by separate formulations, e.g., one LNP formulation comprising a template RNA and a second LNP formulation comprising an mRNA encoding a polypeptide. In some embodiments, the system may comprise more than two nucleic acid components formulated into LNPs. In some embodiments, the system may comprise a protein, e.g., a polypeptide, and a template RNA formulated into at least one LNP formulation.
[0224] In some embodiments, the average LNP diameter of the LNP formulation may be between 10s of nm and 100s of nm, e.g., measured by dynamic light scattering (DLS). In some embodiments, the average LNP diameter of the LNP formulation may be from about 40 nm to about 150 nm, such as about 40 nm, 45 nm, 50 nm, 55 nm, 60 nm, 65 nm, 70 nm, 75 nm, 80 nm, 85 nm, 90 nm, 95 nm, 100 nm, 105 nm, 110 nm, 115 nm, 120 nm, 125 nm, 130 nm, 135 nm, 140 nm, 145 nm, or 150 nm. In some embodiments, the average LNP diameter of the LNP formulation may be from about 50 nm to about 100 Page 70 of 92 12806854v1Attorney Docket No.: 2017469-0044 nm, from about 50 nm to about 90 nm, from about 50 nm to about 80 nm, from about 50 nm to about 70 nm, from about 50 nm to about 60 nm, from about 60 nm to about 100 nm, from about 60 nm to about 90 nm, from about 60 nm to about 80 nm, from about 60 nm to about 70 nm, from about 70 nm to about 100 nm, from about 70 nm to about 90 nm, from about 70 nm to about 80 nm, from about 80 nm to about 100 nm, from about 80 nm to about 90 nm, or from about 90 nm to about 100 nm. In some embodiments, the average LNP diameter of the LNP formulation may be from about 70 nm to about 100 nm. In a particular embodiment, the average LNP diameter of the LNP formulation may be about 80 nm. In some embodiments, the average LNP diameter of the LNP formulation may be about 100 nm. In some embodiments, the average LNP diameter of the LNP formulation ranges from about l mm to about 500 mm, from about 5 mm to about 200 mm, from about 10 mm to about 100 mm, from about 20 mm to about 80 mm, from about 25 mm to about 60 mm, from about 30 mm to about 55 mm, from about 35 mm to about 50 mm, or from about 38 mm to about 42 mm.
[0225] A LNP may, in some instances, be relatively homogenous. A polydispersity index may be used to indicate the homogeneity of a LNP, e.g., the particle size distribution of the lipid nanoparticles. A small (e.g., less than 0.3) polydispersity index generally indicates a narrow particle size distribution. A LNP may have a polydispersity index from about 0 to about 0.25, such as 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.10, 0.11, 0.12, 0.13, 0.14, 0.15, 0.16, 0.17, 0.18, 0.19, 0.20, 0.21, 0.22, 0.23, 0.24, or 0.25. In some embodiments, the polydispersity index of a LNP may be from about 0.10 to about 0.20.
[0226] The zeta potential of a LNP may be used to indicate the electrokinetic potential of the composition. In some embodiments, the zeta potential may describe the surface charge of a LNP. Lipid nanoparticles with relatively low charges, positive or negative, are generally desirable, as more highly charged species may interact undesirably with cells, tissues, and other elements in the body. In some embodiments, the zeta potential of a LNP may be from about -10 mV to about +20 mV, from about -10 mV to about +15 mV, from about -10 mV to about +10 mV, from about -10 mV to about +5 mV, from about -10 mV to about 0 mV, from about -10 mV to about -5 mV, from about -5 mV to about +20 mV, from about -5 mV to about +15 mV, from about -5 mV to about +10 mV, from about -5 mV to about +5 mV, from about -5 mV to about 0 mV, from about 0 mV to about +20 mV, from about 0 mV to about +15 mV, from about 0 mV to about +10 mV, from about 0 mV to about +5 mV, from about +5 mV to about +20 mV, from about +5 mV to about +15 mV, or from about +5 mV to about +10 mV.
[0227] The efficiency of encapsulation of a protein and / or nucleic acid, e.g., polypeptide or mRNA encoding the polypeptide, describes the amount of protein and / or nucleic acid that is encapsulated or otherwise associated with a LNP after preparation, relative to the initial amount provided. The encapsulation efficiency is desirably high (e.g., close to 100%). The encapsulation efficiency may be measured, for example, by comparing the amount of protein or nucleic acid in a solution containing the Page 71 of 92 12806854v1Attorney Docket No.: 2017469-0044 lipid nanoparticle before and after breaking up the lipid nanoparticle with one or more organic solvents or detergents. An anion exchange resin may be used to measure the amount of free protein or nucleic acid (e.g., RNA) in a solution. Fluorescence may be used to measure the amount of free protein and / or nucleic acid (e.g., RNA) in a solution. For the lipid nanoparticles described herein, the encapsulation efficiency of a protein and / or nucleic acid may be at least 50%, for example 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the encapsulation efficiency may be at least 80%. In some embodiments, the encapsulation efficiency may be at least 90%. In some embodiments, the encapsulation efficiency may be at least 95%.
[0228] A LNP may optionally comprise one or more coatings. In some embodiments, a LNP may be formulated in a capsule, film, or table having a coating. A capsule, film, or tablet including a composition described herein may have any useful size, tensile strength, hardness or density.
[0229] Additional exemplary lipids, formulations, methods, and characterization of LNPs are taught by WO2020061457, which is incorporated herein by reference in its entirety.
[0230] In some embodiments, in vitro or ex vivo cell lipofections are performed using Lipofectamine MessengerMax (Thermo Fisher) or TransIT-mRNA Transfection Reagent (Mirus Bio). In certain embodiments, LNPs are formulated using the GenVoy_ILM ionizable lipid mix (Precision NanoSystems). In certain embodiments, LNPs are formulated using 2,2‐dilinoleyl‐4‐dimethylaminoethyl‐[1,3]‐dioxolane (DLin‐KC2‐DMA) or dilinoleylmethyl‐4‐dimethylaminobutyrate (DLin-MC3-DMA or MC3), the formulation and in vivo use of which are taught in Jayaraman et al. Angew Chem Int Ed Engl 51(34):8529-8533 (2012), incorporated herein by reference in its entirety.
[0231] LNP formulations optimized for the delivery of CRISPR-Cas systems, e.g., Cas9-gRNA RNP, gRNA, Cas9 mRNA, are described in WO2019067992 and WO2019067910, both incorporated by reference.
[0232] Additional specific LNP formulations useful for delivery of nucleic acids are described in US8158601 and US8168775, both incorporated by reference, which include formulations used in patisiran, sold under the name ONPATTRO.
[0233] Exemplary dosing of LNP may include about 0.1, 0.25, 0.3, 0.5, 1, 2, 3, 4, 5, 6, 8, 10, or 100 mg / kg (RNA). Exemplary dosing of AAV comprising a nucleic acid encoding one or more components of the system may include an MOI of about 1011, 1012, 1013, and 1014vg / kg. Indications
[0234] Suitable diseases and disorders that can be treated by the systems and methods provided herein include, without limitation, diseases of the central nervous system (CNS), diseases of the eye, diseases of the heart, diseases of the hematopoietic stem cells (HSC), diseases of the kidney, diseases of Page 72 of 92 12806854v1Attorney Docket No.: 2017469-0044 the liver, diseases of the lung, diseases of the skeletal muscle, and diseases of the skin. In some embodiments, the indication is a cancer.
[0235] All publications, patent applications, patents, and other publications and references (e.g., sequence database reference numbers) cited herein are incorporated by reference in their entirety. For example, all GenBank, Unigene, and Entrez sequences referred to herein, e.g., in any Table herein, are incorporated by reference. Unless otherwise specified, the sequences specified herein (e.g., by gene name in RepBase or by accession number), including in any Table herein, refer to the database entries current as of June 6, 2024. When one gene or protein references a plurality of sequence accession numbers, all of the sequence variants are encompassed. EXAMPLES
[0236] The invention is further illustrated by the following examples. The examples are provided for illustrative purposes only and are not to be construed as limiting the scope or content of the invention in any way. Example 1: Biasing LTR retrotransposon elements towards non-integrating circular reverse transcription products
[0237] This example describes a method to generate LTR retrotransposon elements that can be reverse transcribed into functional episomes but cannot be integrated into the genome. Without wishing to be bound by theory, reverse transcription products having a circular episome (e.g. functional episome) configuration are not integrated into a host cell genome, whereas reverse transcription products having a as linear molecule (e.g. linear integration-competent molecule) configuration are capable of integration into a host cell genome. In this example, the template RNA comprises a mutated version of the canonical polypurine tract (PPT) at the 3’ end of the sequence, but its central PPT (cPPT) sequence is retained. Without wishing to be bound by theory, when the canonical PPT is not present (e.g., when it is replaced with a mutant variant), the reverse transcription process will be initiated by the cPPT, resulting in an increase in the proportion of functional episomes compared to linear integration-competent molecules.
[0238] First, a cDNA of the LTR retrotransposon is produced by reverse transcription of the RNA template molecule starting from a PBS sequence downstream of the 5’ LTR (as shown in Step 1 of Figure 1). RNase H then removes the 5’ LTR from the RNA template. The short cDNA product from the previous step anneals to the 3’ LTR, which is typically identical to the 5’-UTR. The first strand cDNA product is then extended up until the PBS site (as shown in Step 2 of Figure 1).
[0239] Then, RNase H removes most of the RNA template molecule other than short purine rich sequences (PPTs), which are typically resistant to degradation by RNase H (as shown in Step 3 of Figure Page 73 of 92 12806854v1Attorney Docket No.: 2017469-0044 1). Using one of these remaining short PPT RNA molecules as a primer, the synthesis of the second strand of the cDNA is then initiated and proceeds through the end of the 5’ LTR. The additional tRNA / PBS sequence, which initiated the cDNA synthesis in step 1, is subsequently removed by RNase H (as it is now an RNA / DNA hybrid susceptible to RNase H activity), as shown in Step 4 of Figure 1.
[0240] Following step 4, a partially complete cDNA molecule with complementary ends (PBS sequences) remains that may fold over and anneal forming a circular structure, which allows priming of further synthesis. This may yield a full-length circular molecule comprising a single LTR sequence, which comprises nicks on opposite DNA strands (as shown in Step 5 of Figure 1). Without wishing bound to theory, the distance between these nicks (e.g., single strand breaks) is dependent on whether the central PPT (cPPT) or standard PPT is used to initiate second strand synthesis (indicated by shaded areas, as shown in step 5 of Figure 1). Specifically, use of the canonical PPT leads to the nicks being closer together, while use of the cPPT leads to the nicks being further apart from each other.
[0241] Typically, a full-length double stranded linear DNA molecule with LTRs on each end is required in order for integration into host genome. To generate this product, the nicked circular DNA needs to unfold and undergo end repair (as shown in Step 6, left panel, of Figure 1). Without wishing to be bound by theory, if the central PPT is used to initiate second strand synthesis rather than the standard PPT, the region that needs to unfold is much longer and therefore requires more energy to melt. Consequently, if the overhanging “sticky ends” of the nicked circle cannot be melted, the molecule is more likely to undergo repair by host DNA repair machinery, which may yield a fully intact circular DNA molecule that cannot be integrated into a host cell genome by the LTR integrase domain. Example 2: Improved transcription start site (TSS) and termination sites
[0242] In this example, the retrotransposition efficiency of a modified IAP LTR retrotransposon comprising optimized transcription start sites (TSS) and termination sites was evaluated. The sequences of the native IAP LTRs and modified LTRs used in this example are shown below. In these sequences: The CCAAT-box of each sequence is labeled with underlined text. The TATA-box of each sequence is labeled with bolded text. The Inr motif of each sequence is labeled with capitalized, italicized, and underlined text. The +1 nucleotide (TSS) of each sequence is labeled with bolded and underlined text. The cleavage and polyadenylation signal of each sequence is labeled with bolded and italicized text. The cleavage site of each sequence is labeled with bolded and capitalized text. Native IAP92L23 LTR Sequence Page 74 of 92 12806854v1Attorney Docket No.: 2017469-0044 Tgttgggagccgcgcccacattcgccgttacaagatggcgctgacagctgtgttctaagtggtaaacaaataatctgcgcatgtgccgagggtggttcttc actctatgtgctctgccttccccgtgacgtcaactcggccgatgggctgcagccaatcagggagtgacacgtcctaggcgaaggagaattctctttaata gggacggggttttgttttctctctctctTgcttctcgctcgctcttgcttcttgcactctggctcctgaagatgtaagcaataaagttttgccgcAgaagattct ggtctgtggtgttcttcctggccgggcgtgagaacgcgtctaataaca (SEQ ID NO: 1) Note that the INR motif comprises “Tg” in the above sequence and the +1 Nucleotide (TSS) comprises the “g” that is also part of the INR motif. Portion of Native 5’ LTR gcttctcgctcgctcttgcttcttgcactctggctcctgaagatgtaagcaataaagttttgccgcAgaagattctggtctgtggtgttcttcctggccgggc gtgagaacgcgtctaataaca (SEQ ID NO: 2) Portion of Native 3’ LTR Tgttgggagccgcgcccacattcgccgttacaagatggcgctgacagctgtgttctaagtggtaaacaaataatctgcgcatgtgccgagggtggttcttc actctatgtgctctgccttccccgtgacgtcaactcggccgatgggctgcagccaatcagggagtgacacgtcctaggcgaaggagaattctctttaata gggacggggttttgttttctctctctctTgcttctcgctcgctcttgcttcttgcactctggctcctgaagatgtaagcaataaagttttgccgcA (SEQ ID NO: 3) Modified 5’ LTR AGttctcgctcgctcttgcttcttgcactctggctcctgaagatgtaagcaataaagttttgccgcAgaagattctggtctgtggtgttcttcctggccggg cgtgagaacgcgtctaataaca (SEQ ID NO: 5) Modified 3’ LTR Tgttgggagccgcgcccacattcgccgttacaagatggcgctgacagctgtgttctaagtggtaaacaaataatctgcgcatgtgccgagggtggttcttc actctatgtgctctgccttccccgtgacgtcaactcggccgatgggctgcagccaatcagggagtgacacgtcctaggcgaaggagaattctctttaata gggacggggttttgttttctctctctctTAGttctcgctcgctcttgcttcttgcactctggctcctgaagatgtaagcaataaagttttgccgcA (SEQ ID NO: 6) Full Modified IAP92L23 LTR sequence following reverse transcription Tgttgggagccgcgcccacattcgccgttacaagatggcgctgacagctgtgttctaagtggtaaacaaataatctgcgcatgtgccgagggtggttcttc actctatgtgctctgccttccccgtgacgtcaactcggccgatgggctgcagccaatcagggagtgacacgtcctaggcgaaggagaattctctttaata gggacggggttttgttttctctctctctTagttctcgctcgctcttgcttcttgcactctggctcctgaagatgtaagcaataaagttttgccgcAgaagattct ggtctgtggtgttcttcctggccgggcgtgagaacgcgtctaataaca (SEQ ID NO: 4) Page 75 of 92 12806854v1Attorney Docket No.: 2017469-0044
[0243] In the native IAP92L23 LTR sequence, the 5’ LTR starts with +1 nucleotide of TSS. In the modified 5’ LTR, the first 2 nucleotides of the sequence were changed to AG for compatibility with the CMV promoter. Within the “Full LTR sequence after reverse transcription,” the 3’ LTR starts at the beginning of the modified 3’ LTR sequence (annotated above) and goes to the natural cleavage site after the polyadenylation signal. Note that the same AG that was changed in the modified 5’ LTR was also changed in the modified 3’ LTR for complete LTR homology. Experimental Procedure
[0244] Modified IAP92L23 was synthesized with the above modified 5’ and modified 3’ LTRs in a plasmid driven by the CMV promoter. An EF1a-BFP-T2A-GFPai transgene was inserted antisense to the IAP ORFs downstream of the stop codon of the POL gene. A ∆POL version was constructed similarly but lacked the POL gene, which was used as a negative control. These constructs were transfected into 293T cells using lipofection. The retrotransposition frequency of each construct, as measured by the amount of GFP+ cells, was monitored by flow cytometry every 3-4 days. In comparison to control modified IAP, which comprises partially truncated LTRs, modified IAP92L23 reproducibly produced more signal. The retrotransposition efficiency of modified IAP92L23 and control modified IAP was dramatically higher than IAP constructs with full LTRs on both 5’ and 3’ ends previously reported in the literature (0.02-0.4% efficiency in various human cell lines). Table E2: Retrotransposition Efficiency Day 4 GFP (%) Day 8 GFP (%) Day 10 GFP (%)Example 3: Deletion Tiling
[0245] In this example, a deletion tiling experiment was conducted to identify key regions of a trans (non-autonomous) template. Using the modified IAP92L23 construct, as described above, site-directed mutagenesis was conducted to create tiled deletion constructs.
[0246] Figures 2A and 2B illustrate the design of the deletion tiling experiments. Solid lines indicate regions included in each construct and dotted lines indicate regions omitted from each construct. The “minimal” diagram illustrates a construct containing a least natural LTR sequence insertion whereas the “optimal” diagram provides an example of the “minimal” + cis expression of GAG for increased efficiency. Page 76 of 92 12806854v1Attorney Docket No.: 2017469-0044
[0247] The first deletion was the PBS region, then a small region between the PBS and the start of GAG, then the putative psi signal (GAG 5’ UTR + beginning of GAG ORF up until the 5’ splice donor in GAG), then successive 500 bp regions thereafter. These constructs were transfected into 293T cells alone or in combination with a ∆PBS driver IAP (which produces all IAP ORFs but cannot be reverse transcribed itself as it lacks a primer binding site for the first strand reverse transcription reaction). Retrotransposition activity was monitored by measuring the percent of cells expressing GFP by flow cytometry every 3-4 days.
[0248] The results are shown in Figure 2A. Each pair of blue dots and pink dots residing between two vertical dash lines indicates the retrotransposition activity of a construct comprising deletion of the region between the two vertical dash lines. The blue dot represents the retrotransposition activity of the deletion construct, whereas the pink dot represents the retrotransposition activity of the same deletion construct in combination with the driver ∆PBS IAP.
[0249] As shown in Figure 2A, the key regions for a trans template include the PBS, the beginning of GAG that may comprise packaging / psi signal, a noncoding region that may comprise a central PPT in the middle of the rve integrase domain of pol, and the PPT closest to the 3’ LTR. RIG stands for Retrotransposon Indicator Gene and is the gene cargo that the trans template carries, which may be replaced by a gene cargo that is desired for a particular application.
[0250] In addition, without wishing to be bound by theory, cis expression of GAG (expression from the template rather than provided in trans) may improve efficiency. Inclusion of the RTE element (a previously identified non-coding structural RNA region that is involved in binding of retrotransposon proteins to the RNA) may also improve efficiency. Example 4: LTR free drivers
[0251] In this example, a series of “native” IAP drivers was constructed. These native IAP drivers express the IAP proteins in appropriate ratios but do not contain IAP LTRs; therefore, these drivers cannot be replicated or integrated into cells. To generate these drivers, the native IAP sequence was cloned into pCMV-driven expression vectors. These constructs were then transfected in combination with the IAP- int∆9 trans template (containing a 500 bp deletion that spans the 3’ part of the RH domain and the 5’ part of the rve domain relative to a native sequence) at various ratios (1:1; 50% or 4:1; 80%) into 293T cells. Retrotransposition, as measured by the percentage of cells expressing GFP, was assessed by flow cytometry for every 3-4 days.
[0252] The schematic structure or the drivers is shown in Figure 3A. Compared to the IAP ∆PBS driver, native drivers v1-v3 were roughly 50% as efficient at mobilizing the IAP-int∆9 template in trans (Figure 3B). These drivers start either at the very end of the 5’ LTR (v1), at the PBS sequence (v2), or Page 77 of 92 12806854v1Attorney Docket No.: 2017469-0044 immediately after the PBS sequence (v3), as compared to WT full-length IAP construct. Native drivers v4-6, which start immediately at the ATG start codon of the GAG ORF (v4), included a 5’UTR and Kozak sequence (v5), or 5’ and 3’ UTRs (v6), as compared to WT full-length IAP construct, did not produce appreciable retrotransposition of the IAP int∆9 template (Figure 3B). Example 5: LTR promoter deletion
[0253] In this experiment, to remove the naturally occurring promoter of the IAP LTR and to limit genotoxicity typically associated with integration of a promoter sequence, identification of key transcription regulatory elements (e.g., CCAAT box, TATA box, Inr, and polyA signal) was performed. A variant of IAP 92L23 was constructed in a CMV expression vector with a truncated LTR (U3-R on the 3’ LTR and R-U5 on the 5’ LTR as described below and shown in Figure 4A. R of each sequence is labeled in underlined text. U5 of each sequence is labeled in italicized text. U3 of 3’ LTR sequence is labeled in bolded text. Full IAP92L23 LTR sequence after reverse transcription Tgttgggagccgcgcccacattcgccgttacaagatggcgctgacagctgtgttctaagtggtaaacaaataatctgcgcatgtgccgagggtg gttcttcactctatgtgctctgccttccccgtgacgtcaactcggccgatgggctgcagccaatcagggagtgacacgtcctaggcgaaggagaa ttctctttaatagggacggggttttgttttctctctctcttagttctcgctcgctcttgcttcttgcactctggctcctgaagatgtaagcaataaagttttgcc gcagaagattctggtctgtggtgttcttcctggccgggcgtgagaacgcgtctaataaca U3-R-U5 5’ LTR sequence in vector agttctcgctcgctcttgcttcttggtgttcttcctggccgggcgtgagaacgcgtctaataaca (SEQ ID NO: 8) 3’ LTR sequence in vector tgttgggagccgcgcccacattcgccgttacaagatggcgagttctcgctcgctcttgcttcttg (SEQ ID NO: 9) Truncated LTR sequence after reverse transcription Tgttgggagccgcgcccacattcgccgttacaagatggcgagttctcgctcgctcttgcttcttggtgttcttcctggccgggcgtgagaa
[0254] To truncate the LTR while maintaining the ability of the LTR go through its normal life cycle of reconstruction by reverse transcription from a template with a partial sequence on both the 5’ and 3’ ends, the ends of the LTR as well as part of the internal R region for homology between 5’ and 3’ partial LTRs present on the RNA molecule were maintained. These constructs were then transfected into 293T cells using lipofection retrotransposition, as measured by the percentage of cells expressing GFP, via flow Page 78 of 92 12806854v1Attorney Docket No.: 2017469-0044 cytometry every 3-4 days. As shown in Figure 4B, a vector containing these LTRs was capable of integration at a detectable level. No retrotransposition of the ∆POL negative control vector was observed with the same LTRs. Example 6: Integration deficient IAP
[0255] In this example, an IAP construct capable of producing episomal / transient DNA was produced by mutation of the catalytic site of the rve integrase domain. First the protein sequence of the POL gene was aligned to the well-characterized HIV integrase protein to identify key catalytic residues (e.g., the DDE catalytic triad). Site-directed mutagenesis was then performed, starting from the previously described optimized IAP92L23 vector (described in Example 2, above), where either one, two, or all three of amino acids in the DDE catalytic triad was recoded to valine using minimal nucleotide changes. Codons for DDE catalytic amino acid triads are labeled with underlined text in the sequence provided below (SEQ ID NO: 10). These vectors were then transfected into 293T cells using lipofection. Evidence of reverse transcription, indicated by percentage of cells expressing GFP, was monitored using flow cytometry every 3-4 days.
[0256] As shown in Figure 5, the VDE construct (SEQ ID NO:11) produced a transient signal of GFP, indicating that this VDE construct was reverse transcribed but not integrated into the genome in contrast to the wildtype vector. The mutation present in the VDE construct, compared to the WT POL gene, is labeled in lower case, underlined, and bolded text in the sequence provided below (SEQ ID NO: 10). WT POL gene TGGAAACCAAGACAGACAGGGTCTGGGTTTTCCTTAGCGGCCATTGGGGCAGCACGACCCAT ACCATGGAAAACAGGGGACCCAGTGTGGGTTCCTCAATGGCACCTATCCTCTGAAAAACTAG AAGCTGTGATTCAACTGGTAGAGGAACAATTAAAACTAGGCCATATTGAACCCTCTACCTCAC CTTGGAATACTCCAATTTTTGTAATTAAGAAAAAGTCAGGAAAGTGGAGACTGCTCCATGACC TCAGAGCCATTAATGAGCAAATGAACTTATTTGGCCCAGTACAGAGGGGTCTCCCTGTACTTT CCGCCTTACCACGTGGCTGGAATTTAATTATTATAGATATTAAAGATTGTTTCTTTTCTATACCTT TGTGTCCAAGGGATAGGCCCAGATTTGCCTTTACCATCCCCTCTATTAATCACATGGAACCTGA TAAGAGGTATCAATGGAAGGTCTTACCACAGGGAATGTCCAATAGTCCTACAATGTGCCAACT TTATGTGCAAGAAGCTCTTTTGCCAGTGAGGGAACAATTCCCCTCTTTAATTTTGCTCCTTTAC ATGGATGACATCCTCCTGTGCCATAAAGACCTTACCATGCTACAAAAGGCATATCCTTTTCTAC TTAAAACTTTAAGTCAGTGGGGTTTACAGATAGCCACAGAAAAGGTCCAAATTTCTGATACAG GACAATTCTTGGGCTCTGTGGTGTCCCCAGATAAGATTGTGCCCCAAAAGGTAGAGATAAGAA GAGATCACCTCCATACCTTAAATGATTTTCAAAAGCTGTTGGGAGATATTAATTGGCTCAGACC TTTTTTAAAGATTCCTTCCGCTGAGTTAAGGCCTTTGTTTAGTATTTTAGAAGGAGATCCTCATA TCTCCTCCCCTAGGACTCTTACTCTAGCTGCTAACCAGGCCTTACAAAAGGTGGAAAAAGCCT TACAGAATGCACAATTACAACGTATTGAGGATTCGCAGCCTTTCAGTTTGTGTGTCTTTAAGAC AGCACAATTGCCAACTGCAGTTTTGTGGCAGAATGGGCCATTGTTGTGGATCCATCCAAACGT ATCCCCAGCTAAAATAATAGATTGGTATCCTGATGCAATTGCACAGCTTGCCCTTAAAGGTCTA Page 79 of 92 12806854v1Attorney Docket No.: 2017469-0044 AAAGCAGCAATCACCCACTTTGGGCAAAGTCCATATCTTTTAATTGTACCTTATACCGCTGCAC AGGTTCAAACCTTGGCAGCCACATCTAATGATTGGGCAGTTTTAGTTACCTCCTTTTCAGGAA AAATAGATAACCATTATCCAAAACATCCAATCTTACAGTTTGCCCAAAATCAATCTGTTGTGTT TCCACAAATAACAGTAAGAAACCCACTTAAAAATGGGATTGTGGTATATACTGATGGATCAAA AACTGGCATAGGTGCCTATGTGGCTAATGGTAAAGTGGTATCCAAACAATATAATGAAAATTCA CCTCAAGTGGTAGAATGTTTAGTGGTCTTAGAAGTTTTAAAAACCTTTTTAGAACCCCTTAATA TTGTGTCAGATTCCTGTTATGTGGTTAATGCAGTAAATCTTTTAGAAGTGGCTGGAGTGATTAA GCCTTCCAGTAGAGTTGCCAATATTTTTCAGCAGATACAATTAGTTTTGTTATCTAGAAGATTTC CTGTTTATATTACTCATGTTAGAGCCCATTCAGGCCTACCTGGCCCCATGGCTCTGGGAAATAAT TTGGCAGATAAGGCCACTAAAGTGGTGGCTGCTGCCCTATCATCCCCGGTAGAGGCTGCAAGA AATTTTCATAACAATTTTCATGTGACGGCTGAAACATTACGCAGTCGTTTCTCCTTGACAAGAA AAGAAGCCCGTGACATTGTTACTCAATGTCAAAGCTGCTGTGAGTTCTTGCCAGTTCCTCATG TGGGAATTAACCCACGCGGTATTCGACCTCTACAGGTCTGGCAAATGGATGTTACACATGTTT CTTCCTTTGGAAAACTTCAATATCTCCATGTGTCCATTGACACATGTTCTGGCATCATGTTTGCT TCTCCATTAACCGGAGAAAAAGCCTCACATGTGATTCAACATTGCCTTGAGGCATGGAGTGCT TGGGGGAAACCCAGACTCCTTAAGACTGATAATGGACCAGCTTATACGTCTCAAAAATTCCA ACAGTTCTGCCGTCAGATGGACGTGACCCACCTGACTGGACTTCCATACAACCCTCAAGGAC AGGGTATTGTTGAGCGTGCGCATCGCACCCTCAAAACCTATCTTATAAAACAGAAGAGGGGA ACTTTTGAGGAGACTGTACCCCGAGCACCAAGAGTGTCGGTGTCTATGGCACTCTTTACACTC AATTTTTTAAATATTGATGCTCATGGCCATACTGCGGCTGAACGTCATTGTACAGAGCCAGATA GGCCCAATGAGATGGTTAAATGGAAAAATGTCCTTGATAATAAATGGTATGGCCCGGATCCTAT TTTGATAAGATCCAGGGGAGCTATCTGTGTTTTCCCACAGAATGAAAACAACCCATTTTGGATA CCAGAAAGACTCACCCGAAAAATCCAGACTGACCAAGGAAATACTAATGTCCCTCGTCTTGG TGATGTCCAGGGCGTCAATAATAAAAAGAGAGCAGCGTTGGGGGATAATGTCGACATTTCCAC TCCCAATGACGGTGATGTATAA (SEQ ID NO: 10) IAP VDE POL Gene TGGAAACCAAGACAGACAGGGTCTGGGTTTTCCTTAGCGGCCATTGGGGCAGCACGACCCAT ACCATGGAAAACAGGGGACCCAGTGTGGGTTCCTCAATGGCACCTATCCTCTGAAAAACTAG AAGCTGTGATTCAACTGGTAGAGGAACAATTAAAACTAGGCCATATTGAACCCTCTACCTCAC CTTGGAATACTCCAATTTTTGTAATTAAGAAAAAGTCAGGAAAGTGGAGACTGCTCCATGACC TCAGAGCCATTAATGAGCAAATGAACTTATTTGGCCCAGTACAGAGGGGTCTCCCTGTACTTT CCGCCTTACCACGTGGCTGGAATTTAATTATTATAGATATTAAAGATTGTTTCTTTTCTATACCTT TGTGTCCAAGGGATAGGCCCAGATTTGCCTTTACCATCCCCTCTATTAATCACATGGAACCTGA TAAGAGGTATCAATGGAAGGTCTTACCACAGGGAATGTCCAATAGTCCTACAATGTGCCAACT TTATGTGCAAGAAGCTCTTTTGCCAGTGAGGGAACAATTCCCCTCTTTAATTTTGCTCCTTTAC ATGGATGACATCCTCCTGTGCCATAAAGACCTTACCATGCTACAAAAGGCATATCCTTTTCTAC TTAAAACTTTAAGTCAGTGGGGTTTACAGATAGCCACAGAAAAGGTCCAAATTTCTGATACAG GACAATTCTTGGGCTCTGTGGTGTCCCCAGATAAGATTGTGCCCCAAAAGGTAGAGATAAGAA GAGATCACCTCCATACCTTAAATGATTTTCAAAAGCTGTTGGGAGATATTAATTGGCTCAGACC TTTTTTAAAGATTCCTTCCGCTGAGTTAAGGCCTTTGTTTAGTATTTTAGAAGGAGATCCTCATA TCTCCTCCCCTAGGACTCTTACTCTAGCTGCTAACCAGGCCTTACAAAAGGTGGAAAAAGCCT TACAGAATGCACAATTACAACGTATTGAGGATTCGCAGCCTTTCAGTTTGTGTGTCTTTAAGAC AGCACAATTGCCAACTGCAGTTTTGTGGCAGAATGGGCCATTGTTGTGGATCCATCCAAACGT ATCCCCAGCTAAAATAATAGATTGGTATCCTGATGCAATTGCACAGCTTGCCCTTAAAGGTCTA AAAGCAGCAATCACCCACTTTGGGCAAAGTCCATATCTTTTAATTGTACCTTATACCGCTGCAC AGGTTCAAACCTTGGCAGCCACATCTAATGATTGGGCAGTTTTAGTTACCTCCTTTTCAGGAA AAATAGATAACCATTATCCAAAACATCCAATCTTACAGTTTGCCCAAAATCAATCTGTTGTGTT Page 80 of 92 12806854v1Attorney Docket No.: 2017469-0044 TCCACAAATAACAGTAAGAAACCCACTTAAAAATGGGATTGTGGTATATACTGATGGATCAAA AACTGGCATAGGTGCCTATGTGGCTAATGGTAAAGTGGTATCCAAACAATATAATGAAAATTCA CCTCAAGTGGTAGAATGTTTAGTGGTCTTAGAAGTTTTAAAAACCTTTTTAGAACCCCTTAATA TTGTGTCAGATTCCTGTTATGTGGTTAATGCAGTAAATCTTTTAGAAGTGGCTGGAGTGATTAA GCCTTCCAGTAGAGTTGCCAATATTTTTCAGCAGATACAATTAGTTTTGTTATCTAGAAGATTTC CTGTTTATATTACTCATGTTAGAGCCCATTCAGGCCTACCTGGCCCCATGGCTCTGGGAAATAAT TTGGCAGATAAGGCCACTAAAGTGGTGGCTGCTGCCCTATCATCCCCGGTAGAGGCTGCAAGA AATTTTCATAACAATTTTCATGTGACGGCTGAAACATTACGCAGTCGTTTCTCCTTGACAAGAA AAGAAGCCCGTGACATTGTTACTCAATGTCAAAGCTGCTGTGAGTTCTTGCCAGTTCCTCATG TGGGAATTAACCCACGCGGTATTCGACCTCTACAGGTCTGGCAAATGGtTGTTACACATGTTTC TTCCTTTGGAAAACTTCAATATCTCCATGTGTCCATTGACACATGTTCTGGCATCATGTTTGCTT CTCCATTAACCGGAGAAAAAGCCTCACATGTGATTCAACATTGCCTTGAGGCATGGAGTGCTT GGGGGAAACCCAGACTCCTTAAGACTGATAATGGACCAGCTTATACGTCTCAAAAATTCCAA CAGTTCTGCCGTCAGATGGACGTGACCCACCTGACTGGACTTCCATACAACCCTCAAGGACA GGGTATTGTTGAGCGTGCGCATCGCACCCTCAAAACCTATCTTATAAAACAGAAGAGGGGAA CTTTTGAGGAGACTGTACCCCGAGCACCAAGAGTGTCGGTGTCTATGGCACTCTTTACACTCA ATTTTTTAAATATTGATGCTCATGGCCATACTGCGGCTGAACGTCATTGTACAGAGCCAGATAG GCCCAATGAGATGGTTAAATGGAAAAATGTCCTTGATAATAAATGGTATGGCCCGGATCCTATT TTGATAAGATCCAGGGGAGCTATCTGTGTTTTCCCACAGAATGAAAACAACCCATTTTGGATA CCAGAAAGACTCACCCGAAAAATCCAGACTGACCAAGGAAATACTAATGTCCCTCGTCTTGG TGATGTCCAGGGCGTCAATAATAAAAAGAGAGCAGCGTTGGGGGATAATGTCGACATTTCCAC TCCCAATGACGGTGATGTATAA (SEQ ID NO: 11) Wt POL amino acid sequence WKPRQTGSGFSLAAIGAARPIPWKTGDPVWVPQWHLSSEKLEAVIQLVEEQLKLGHIEPSTSPWN TPIFVIKKKSGKWRLLHDLRAINEQMNLFGPVQRGLPVLSALPRGWNLIIIDIKDCFFSIPLCPRDR PRFAFTIPSINHMEPDKRYQWKVLPQGMSNSPTMCQLYVQEALLPVREQFPSLILLLYMDDILLC HKDLTMLQKAYPFLLKTLSQWGLQIATEKVQISDTGQFLGSVVSPDKIVPQKVEIRRDHLHTLND FQKLLGDINWLRPFLKIPSAELRPLFSILEGDPHISSPRTLTLAANQALQKVEKALQNAQLQRIEDS QPFSLCVFKTAQLPTAVLWQNGPLLWIHPNVSPAKIIDWYPDAIAQLALKGLKAAITHFGQSPYL LIVPYTAAQVQTLAATSNDWAVLVTSFSGKIDNHYPKHPILQFAQNQSVVFPQITVRNPLKNGIV VYTDGSKTGIGAYVANGKVVSKQYNENSPQVVECLVVLEVLKTFLEPLNIVSDSCYVVNAVNLL EVAGVIKPSSRVANIFQQIQLVLLSRRFPVYITHVRAHSGLPGPMALGNNLADKATKVVAAALSS PVEAARNFHNNFHVTAETLRSRFSLTRKEARDIVTQCQSCCEFLPVPHVGINPRGIRPLQVWQMD VTHVSSFGKLQYLHVSIDTCSGIMFASPLTGEKASHVIQHCLEAWSAWGKPRLLKTDNGPAYTS QKFQQFCRQMDVTHLTGLPYNPQGQGIVERAHRTLKTYLIKQKRGTFEETVPRAPRVSVSMALF TLNFLNIDAHGHTAAERHCTEPDRPNEMVKWKNVLDNKWYGPDPILIRSRGAICVFPQNENNPF WIPERLTRKIQTDQGNTNVPRLGDVQGVNNKKRAALGDNVDISTPNDGDV (SEQ ID NO: 12) Page 81 of 92 12806854v1
Claims
Attorney Docket No.: 2017469-0044 CLAIMS 1. A template RNA comprising: a 5’ long terminal repeat (5’ LTR) comprising a sequence according to SEQ ID NO: 5, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein the 5’ LTR comprises an A at position 1 of SEQ ID NO: 5 and a G at position 2 of SEQ ID NO: 5; a 3’ long terminal repeat (3’ LTR) comprising a sequence according to SEQ ID NO: 6, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto wherein the 3’ LTR comprises an A at position 231 of SEQ ID NO: 6 and a G at position 232 of SEQ ID NO: 6; a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and a primer binding site (PBS), wherein optionally the PBS is between the 5’ LTR and the heterologous object sequence.
2. The template RNA of claim 1, wherein the 5’ LTR has a length of less than 354, 300, 200, 150, or 125 nucleotides.
3. The template RNA of claim 1, wherein the 5’ LTR has a length of 354 nucleotides.
4. The template RNA of claim 1 or 2, wherein the 5’ LTR has a length of 124 nucleotides.
5. The template RNA of any one of the preceding claims, wherein the 3’ LTR has a length of less than 354, 340, 330, 320, 310, 300, 299, or 298 nucleotides.
6. The template RNA of any one of claims 1-4, wherein the 3’ LTR has a length of 354 nucleotides.
7. The template RNA of any one of claims 1-5, wherein the 3’ LTR has a length of 297 nucleotides.
8. The template RNA of any one of claims 1, 3, or 5-7, wherein the 5’ UTR has a sequence according to SEQ ID NO: 4, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. Page 82 of 92 12806854v1Attorney Docket No.: 2017469-0044 9. The template RNA of any one of claims 1-4 or 6, wherein the 3’ UTR has a sequence according to SEQ ID NO: 4, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
10. A template RNA comprising: a 5’ long terminal repeat (5’ LTR) comprising a sequence according to SEQ ID NO: 5 or 2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; a 3’ long terminal repeat (3’ LTR) comprising a sequence according to SEQ ID NO: 6 or 3, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and a primer binding site (PBS), wherein optionally the PBS is between the 5’ LTR and the heterologous object sequence, wherein one or both of: the 5’ LTR has a length of less than 354, 300, 200, 150, or 125 nucleotides; or the 3’ LTR has a length of less than 354, 340, 330, 320, 310, 300, 299, or 298 nucleotides.
11. The template RNA of claim 10, wherein the 5’ LTR comprises an A at position 1 of SEQ ID NO: 5 and a G at position 2 of SEQ ID NO:
5.
12. The template RNA of claim 10 or 11, wherein the 3’ LTR comprises an A at position 231 of SEQ ID NO: 6 and a G at position 232 of SEQ ID NO:
6.
13. The template RNA of any one of claims 10-12, wherein the 5’ LTR has a length of 124 nucleotides.
14. The template RNA of any one of claims 10-13, wherein the 3’ LTR has a length of 297 nucleotides.
15. A template RNA comprising (e.g., in a 5’ to 3’ direction): a) a 5’ long terminal repeat (5’ LTR) having a sequence according to SEQ ID NO: 2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; b) a primer binding site (PBS); Page 83 of 92 12806854v1Attorney Docket No.: 2017469-0044 c) the nucleic acid sequence between position 2 and position 4, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; d) the nucleic acid sequence between position 12 and position 13, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; a polypurine tract (PPT), wherein optionally the PPT is immediately adjacent to the 3’ LTR; a 3’ long terminal repeat (3’ LTR) having a sequence according to SEQ ID NO: 3 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR.
16. The template RNA of claim 15, wherein (c) comprises a packaging / psi signal.
17. The template RNA of claim 15 or 16, wherein (c) comprises the gag splice donor site.
18. The template RNA of any one of claims 15-17, wherein (d) comprises a central PPT (cPPT).
19. A template RNA comprising: a 5’ long terminal repeat (5’ LTR) having a sequence according to SEQ ID NO: 2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; a 3’ long terminal repeat (3’ LTR) having a sequence according to SEQ ID NO: 3 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and a primer binding site (PBS); wherein the template RNA does not comprise one or more of (e.g., does not comprise 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or all of): the nucleic acid sequence between position 1 and position 2, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 4 and position 5, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; Page 84 of 92 12806854v1Attorney Docket No.: 2017469-0044 the nucleic acid sequence between position 5 and position 6, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 6 and position 7, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 7 and position 8, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 8 and position 9, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 9 and position 10, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 10 and position 11, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 11 and position 12, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 12 and position 13, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; the nucleic acid sequence between position 13 and position 14, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and / or the nucleic acid sequence between position 14 and position 15, as shown in Figure 3B, encoding an IAP retrotransposase polypeptide, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
20. The template RNA of any one of claims 15-19, which comprises a premature stop codon in gag, pro, pol, or rig. Page 85 of 92 12806854v1Attorney Docket No.: 2017469-0044 21. A template RNA comprising: an IAP 5’ long terminal repeat (IAP 5’ LTR); an IAP 3’ long terminal repeat (IAP 3’ LTR); a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and a primer binding site (PBS), wherein optionally the PBS is between the IAP 5’ LTR and the heterologous object sequence; wherein one or both of: the IAP 5’ LTR comprises a sequence according to SEQ ID NO: 8, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; or the IAP 3’ LTR comprises a sequence according to SEQ ID NO: 9 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
22. A template RNA comprising: an IAP 5’ long terminal repeat (IAP 5’ LTR); an IAP 3’ long terminal repeat (IAP 3’ LTR); a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and a primer binding site (PBS), wherein optionally the PBS is between the IAP 5’ LTR and the heterologous object sequence; wherein one or both of: the IAP 5’ LTR consists of: a first fragment of a full-length IAP 5’ LTR according to SEQ ID NO: 1, wherein the first fragment consists of nucleotides 231-255 of SEQ ID NO: 1 and optionally up to 2, 5, or 10 additional nucleotides adjacent to either side of said nucleotides of SEQ ID NO: 1, or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto; and a second fragment of a full-length IAP 5’ LTR according to SEQ ID NO: 1, wherein the second fragment consists of nucleotides 315-354 of SEQ ID NO: 1 and optionally up to 2, 5, or 10 additional nucleotides adjacent to either side of said nucleotides of SEQ ID NO: 1, or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto; or the IAP 3’ LTR consists of: a first fragment of a full-length IAP 3’ LTR according to SEQ ID NO: 1, wherein the first fragment consists of nucleotides 1-40 of SEQ ID NO: 1 and optionally up to 2, 5, or 10 additional Page 86 of 92 12806854v1Attorney Docket No.: 2017469-0044 nucleotides adjacent to either side of said nucleotides of SEQ ID NO: 1, or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto; and a second fragment of a full-length IAP 3’ LTR according to SEQ ID NO: 1, wherein the second fragment consists of nucleotides 231-255 of SEQ ID NO: 1 and optionally up to 2, 5, or 10 additional nucleotides adjacent to either side of said nucleotides of SEQ ID NO: 1, or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto.
23. A template RNA comprising: an IAP 5’ long terminal repeat (IAP 5’ LTR); an IAP 3’ long terminal repeat (IAP 3’ LTR); a heterologous object sequence encoding a therapeutic effector, positioned between the 5’ LTR and the 3’ LTR; and a primer binding site (PBS), wherein optionally the PBS is between the IAP 5’ LTR and the heterologous object sequence; wherein one or both of: the IAP 5’ LTR comprises a mutation that reduces promoter function of the IAP 5’ LTR compared to an IAP 5’ LTR having a sequence of SEQ ID NO: 1; or the IAP 3’ LTR comprises a mutation that reduces promoter function of the IAP 3’ LTR compared to an IAP 3’ LTR having a sequence of SEQ ID NO:
1.
24. The template RNA of claim 23, wherein the mutation that reduces promoter function is a deletion, e.g., a deletion of one or more of (e.g., two or all of): the TATA-box; the CCAA-box; Inr motif; or the polyadenylation signal.
25. The template RNA of any one of claims 21-24, which is transcribed at a lower level than a control template RNA having the same sequence as the template RNA except that its 3’ LTR has a sequence of SEQ ID NO:
1.
26. The template RNA of any one of claims 21-24, which is transcribed at a lower level than a control template RNA having the same sequence as the template RNA except that its 5’ LTR has a sequence of SEQ ID NO:
1. Page 87 of 92 12806854v1Attorney Docket No.: 2017469-0044 27. The template RNA of any one of claims 21-24, which is active for integration into a target nucleic acid in an assay according to Example 5.
28. A DNA molecule comprising a sequence according to SEQ ID NO: 4, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein nucleotide 231 of SEQ ID NO: 4 is A and nucleotide 232 of SEQ ID NO: 4 is G.
29. A DNA molecule comprising a sequence according to SEQ ID NO: 7, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
30. A DNA molecule having a sequence that is the reverse complement of the sequence of claim 28 or 29.
31. The DNA molecule of any one of claims 28-30, which is single stranded or double stranded.
32. An IAP retrotransposase polypeptide comprising: an amino acid sequence of SEQ ID NO: 12, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein one, two, or all of: position D653 of SEQ ID NO: 12 is other than D, e.g., is valine; position D710 of SEQ ID NO: 12 is other than D, e.g., is valine; or position E746 of SEQ ID NO: 12 is other than E, e.g., is valine.
33. The IAP retrotransposase polypeptide of claim 32, wherein position D653 of SEQ ID NO: 12 is valine.
34. The IAP retrotransposase polypeptide of claim 32 or 33, wherein position D710 of SEQ ID NO: 12 is valine.
35. The IAP retrotransposase polypeptide of any one of claims 32-34, wherein position E746 of SEQ ID NO: 12 is valine.
36. A nucleic acid molecule (e.g., a DNA molecule) encoding the IAP retrotransposase polypeptide any one of claims 32-35. Page 88 of 92 12806854v1Attorney Docket No.: 2017469-0044 37. The nucleic acid molecule of claim 36, which comprises a T at position 1958 of SEQ ID NO:
10.
38. The IAP retrotransposase polypeptide of any one of claims 32-35, wherein position D710 of SEQ ID NO: 12 is aspartate (D) and position E746 of SEQ ID NO: 12 is glutamate (E).
39. The IAP retrotransposase polypeptide of any one of claims 32-35 and 38, which has reduced integrase activity compared to a polypeptide of SEQ ID NO: 12, e.g., reduced by about 20%, 40%, 60%, 80%, 90%, or 95%, e.g., in an assay of Example 6.
40. The IAP retrotransposase polypeptide of any one of claims 32-35 and 38-39, which leads to transient expression of a heterologous object sequence, e.g., wherein expression is detectable for less than 9 days, e.g., in an assay of Example 6.
41. A template RNA comprising (e.g., from 5’ to 3’): a) an IAP 5’ long terminal repeat (IAP 5’ LTR); b) a primer binding site (PBS); c) a heterologous object sequence encoding a therapeutic effector; d) a central PPT (cPPT); and e) an IAP 3’ long terminal repeat (IAP 3’ LTR); wherein, between d) and e), the template RNA lacks a canonical PPT or comprises a mutation to a canonical PPT (“mutant PPT”) that reduces initiation of second strand synthesis primed by the mutant PPT.
42. The template RNA of claim 41, wherein the cPPT is situated upstream of the heterologous object sequence, downstream of the heterologous object sequence, or overlapping with the heterologous object sequence.
43. The template RNA of claim 41 or 42, wherein, upon reverse transcription, the template RNA yields a greater proportion of episomes to linear dsDNA, compared to the proportion of episomes to linear dsDNA produced using a reference template RNA which comprises a canonical PPT between its cPPT and IAP 3’ LTR.
44. The template RNA of claim 41 or 42, wherein, upon reverse transcription, the template RNA yields fewer insertions into a host cell genome, compared to the number of insertions into a host cell Page 89 of 92 12806854v1Attorney Docket No.: 2017469-0044 genome produced using a reference template RNA which comprises a canonical PPT between its cPPT and IAP 3’ LTR.
45. A DNA molecule encoding the template RNA of any one of claims 1-27 or 41-44.
46. A system comprising: a template RNA of any one of claims 1-27 or 41-44, or a DNA encoding the template RNA, and an IAP retrotransposase polypeptide, or a nucleic acid encoding the IAP retrotransposase polypeptide.
47. The system of claim 46, wherein the IAP retrotransposase polypeptide comprises an amino acid sequence of SEQ ID NO: 12, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
48. A system comprising: a template RNA, or a DNA encoding the template RNA; and an IAP retrotransposase polypeptide according to any one of claims 32-35 or 38-40, or a nucleic acid encoding the IAP retrotransposase polypeptide.
49. The system of claim 48, wherein the template RNA is a template RNA of any one of claims 1-27 or 41-44.
50. The system of any one of claims 46-49, wherein the system is substantially free of virus.
51. The system of any one of claims 46-50, wherein the system is substantially free of cells.
52. A nucleic acid encoding the IAP retrotransposase polypeptide of any one of claims 32-35 or 38- 40.
53. The nucleic acid of claim 52, which is an mRNA.
54. The nucleic acid of claim 52 or 53, which comprises one or more non-canonical or modified ribonucleotides. Page 90 of 92 12806854v1Attorney Docket No.: 2017469-0044 55. The system or the template RNA of any one of the preceding claims, wherein the template RNA comprises one or more non-canonical or modified ribonucleotides.
56. A pharmaceutical composition comprising the template RNA or the system of any one of the preceding claims.
57. The system or IAP retrotransposase polypeptide of any preceding claims, wherein the system, nucleic acid molecule, polypeptide, and / or DNA encoding the same, is formulated as a lipid nanoparticle (LNP).
58. A cell comprising the system of any one of claims 46-51, 55, or 57.
59. A cell comprising the template RNA of any one of claims 1-27 or 41-44.
60. A cell comprising the IAP retrotransposase polypeptide of any one of claims 32-35 or 38-40.
61. A method of delivering a heterologous object sequence to a target cell, comprising introducing into the target cell (e.g., contacting the target cell with) a system of any one of claims 46-51, 55, or 57, and incubating the target cell under conditions suitable for production of the template DNA. Page 91 of 92 12806854v1