Methods and compositions for regulating the genome
A polypeptide-endonuclease system with a template RNA enhances genome integration efficiency and specificity by integrating exogenous genetic elements directly into host genomes.
Patent Information
- Application Number
- JP2021511605
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-21
- Filing Date
- 2019-08-28
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2039-08-28
AI Technical Summary
Existing methods for integrating nucleic acids into genomes suffer from low frequency and lack of site specificity, and current techniques like CRISPR/Cas9 are inefficient for long sequence insertion, while methods like Cre/loxP require multiple steps.
A system comprising a polypeptide with a reverse transcriptase and endonuclease domain, combined with a template RNA, facilitates the integration of exogenous genetic elements into host genomes with improved specificity and efficiency.
The system enables precise and efficient insertion of heterologous sequences into target sites within genomes, reducing off-target effects and minimizing DNA repair pathway interference.
Smart Images

Figure 0007680351000480 
Figure 0007680351000481 
Figure 0007680351000482
Abstract
Description
[Technical field]
[0001] This application claims priority to U.S. Patent Application No. 62 / 723,886, filed August 28, 2018, U.S. Patent Application No. 62 / 725,778, filed August 31, 2018, U.S. Patent Application No. 62 / 850,883, filed May 21, 2019, and U.S. Patent Application No. 62 / 864,924, filed June 21, 2019, the entire contents of each of which are incorporated herein by reference.
[0002] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in ASCII format, the entire contents of which are incorporated herein by reference. The ASCII copy was created on August 28, 2019, is named V2065-7000WO_SL.txt, and is 4,004,548 bytes in size. [Background technology]
[0003] The integration of a nucleic acid of interest into a genome occurs at a low frequency and with little site specificity in the absence of special proteins to facilitate the insertion event. Some existing methods, such as CRISPR / Cas9, are more suitable for small edits and are less effective for the integration of long sequences. Other existing methods, such as Cre / loxP, require a first step of inserting a loxP site into the genome, and then a second step of inserting a sequence of interest into the loxP site. There is a need in the art for improved proteins for inserting a sequence of interest into a genome. Summary of the Invention [Means for solving the problem]
[0004] The present disclosure relates to novel compositions, systems and methods for modifying the genome of one or more locations in a host cell, tissue or subject, either in vivo or in vitro. In particular, the present invention features compositions, systems and methods for the introduction of exogenous genetic elements into a host genome.
[0005] Features of the compositions or methods may include one or more of the embodiments set forth below.
[0006] 1. The following: (a) a polypeptide or a nucleic acid encoding a polypeptide, the polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain; and (b) a template RNA (or a DNA encoding the template RNA) comprising (i) a sequence that binds to the above polypeptide and (ii) a heterologous sequence of interest that encodes a therapeutic polypeptide or that encodes a mammalian (e.g., human) polypeptide, or a fragment or variant thereof. A system for modifying DNA comprising:
[0007] 2. The following: (a) a polypeptide or a nucleic acid encoding a polypeptide, the polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain; and (b) a template RNA (or a DNA encoding the template RNA) comprising (i) a sequence that binds to the above polypeptide and (ii) a heterologous sequence of interest; 1. A system for modifying DNA comprising: i. the heterologous sequence of interest encodes a protein, e.g., an enzyme (e.g., a lysosomal enzyme) or a blood factor (e.g., Factor I, II, V, VII, X, XI, XII, or XIII); ii. the heterologous sequence of interest comprises a tissue-specific promoter or enhancer; iii. the heterologous sequence of interest encodes a polypeptide of more than 250, 300, 400, 500, or 1,000 amino acids, and optionally up to 7,500 amino acids; iv. the heterologous sequence of interest encodes a fragment of a mammalian gene, but not the entire mammalian gene, e.g., encodes one or more exons, but not the full-length protein; v. the heterologous sequence of interest encodes one or more introns; vi. the heterologous sequence of interest is other than GFP, e.g., other than a fluorescent protein or other than a reporter protein; vii. The heterologous sequence of interest is other than a T cell chimeric antigen receptor; A system that is one or more of the above.
[0008] 3. The following: (a) a polypeptide or a nucleic acid encoding a polypeptide, the polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain; and (b) a template RNA (or a DNA encoding the template RNA) comprising (i) a sequence that binds to the above polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising:
[0009] 4. The following: (a) a polypeptide or a nucleic acid encoding a polypeptide, the polypeptide comprising (i) a target DNA binding domain, (ii) a reverse transcriptase domain, and (iii) an endonuclease domain; and (b) a template RNA (or a DNA encoding the template RNA) comprising (i) a sequence that binds to the above polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising:
[0010] 5. The following: (a) a polypeptide or a nucleic acid encoding a polypeptide, said polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain, wherein one or both of (i) or (ii) are derived from an avian retrotransposase, e.g., a sequence of Table 2 or 3, or at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (b) a template RNA (or a DNA encoding the template RNA) comprising (i) a sequence that binds to the above polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising:
[0011] 6. Below: (a) a polypeptide or a nucleic acid encoding a polypeptide, the polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain, wherein the polypeptide has at least 70%, 75%, 80%, 85%, 90%, or 95% of its activity at 37° C. under otherwise similar conditions at 25° C.; and (b) a template RNA (or a DNA encoding the template RNA) comprising (i) a sequence that binds to the above polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising:
[0012] 7. The system of embodiment 6, wherein the polypeptide is derived from an avian retrotransposase, such as an avian retrotransposase listed in column 8 of Table 3, or is a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0013] 8. The system of embodiment 6, wherein the bird retrotransposase is a retrotransposase from Taeniopygia guttata, Geospiza fortis, Zonotrichia albicollis, or Tinamus guttatus, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0014] 9. The system of embodiment 6, wherein the polypeptide is derived from or has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to a retrotransposase listed in column 8 of Table 3.
[0015] 10. The system of any of the preceding embodiments, wherein the template RNA comprises a sequence of Table 3 (e.g., one or both of the 5' untranslated region of Table 3, column 6 and the 3' untranslated region of Table 3, column 7), or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0016] 11. The following: (a) a polypeptide or a nucleic acid encoding a polypeptide, the polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain; and (b) a template RNA (or a DNA encoding the template RNA) comprising (i) a sequence that binds to the above polypeptide and (ii) a heterologous sequence of interest; 1. A system for modifying DNA comprising: i. the nucleic acid encoding the polypeptide and the template RNA or the nucleic acid encoding the template RNA are separate nucleic acids; ii. the template RNA does not encode an active reverse transcriptase, e.g., contains an inactivated mutant reverse transcriptase (e.g., as described in Examples 1-2) or does not contain a reverse transcriptase sequence; or iii. the template RNA does not encode an active endonuclease, e.g., contains an inactivated endonuclease or contains no endonuclease; or iv. the template RNA contains one or more chemical modifications; A system that is one or more of the above.
[0017] 12. The following: (a) a polypeptide or a nucleic acid encoding a polypeptide, said polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain; and (b) a template RNA (or a DNA encoding the template RNA) comprising: (i) a 5' untranslated sequence linked to the polypeptide; (ii) a 3' untranslated sequence linked to the polypeptide; (iii) a heterologous sequence of interest; and (iv) a promoter operably linked to the heterologous sequence of interest. A system for modifying DNA comprising: The promoter is located between the 5' untranslated sequence bound to the polypeptide and the heterologous sequence; or The promoter is positioned between the 3' untranslated sequence which binds to the polypeptide and the heterologous sequence.
[0018] 13. The following: (a) a polypeptide or a nucleic acid encoding a polypeptide, the polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain; and (b) a template RNA (or a DNA encoding the template RNA) comprising: (i) a 5' untranslated sequence linked to the polypeptide; (ii) a 3' untranslated sequence linked to the polypeptide; and (iii) a heterologous sequence of interest. A system for modifying DNA comprising: The heterologous sequence of interest comprises an open reading frame (or its reverse complement) on the template RNA in a 5' to 3' direction; or A system in which the heterologous target sequence comprises an open reading frame (or its reverse complement) in the 3' to 5' direction on the template RNA.
[0019] 14. The following: (a) a polypeptide or a nucleic acid encoding a polypeptide, said polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain, wherein at least one of (i) or (ii) is heterologous; and (b) a template RNA (or a DNA encoding the template RNA) comprising (i) a sequence that binds to the above polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising:
[0020] 15. The following: (a) a polypeptide or a nucleic acid encoding a polypeptide, said polypeptide comprising (i) a target DNA binding domain, (ii) a reverse transcriptase domain, and (iii) an endonuclease domain, wherein at least one of (i), (ii) or (iii) is heterologous; and (b) a template RNA (or a DNA encoding the template RNA) comprising (i) a sequence that binds to the above polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising:
[0021] 16. The following: (a) a polypeptide or a nucleic acid encoding a polypeptide, the polypeptide comprising (i) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to a reverse transcriptase domain of a purine / apyrimidinic endonuclease (APE)-type non-LTR retrotransposon, and (ii) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to an endonuclease domain of an APE-type non-LTR retrotransposon; and (b) a template RNA (or a DNA encoding the template RNA) comprising (i) a sequence that binds to the above polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising:
[0022] 17. The following: (a) a polypeptide or a nucleic acid encoding a polypeptide, the polypeptide comprising: (i) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to a reverse transcriptase domain of a restriction enzyme-like endonuclease (RLE)-type non-LTR retrotransposon; (ii) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to an endonuclease domain of a RLE-type non-LTR retrotransposon; and (iii) a heterologous target DNA binding domain (e.g., a heterologous zinc finger DNA binding domain); and (b) a template RNA (or a DNA encoding the template RNA) comprising (i) a sequence that binds to the above polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising:
[0023] 18. The system of any of the preceding embodiments, wherein the template RNA comprises (iii) a promoter operably linked to a heterologous sequence of interest.
[0024] 19. The system of any of the preceding embodiments, wherein the polypeptide further comprises (iii) a DNA-binding domain.
[0025] 20. The system of embodiment 17, wherein the polypeptide comprises a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the sequence of SEQ ID NO: 1016.
[0026] 21. The system of any of the previous embodiments, wherein the polypeptide comprises a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to a sequence in Table 3, column 8.
[0027] 22. The system of any of the previous embodiments, wherein the nucleic acid encoding the polypeptide and the template RNA or the nucleic acid encoding the template RNA are covalently linked, e.g., are part of a fusion protein.
[0028] 23. The system of embodiment 22, wherein the fusion nucleic acid comprises RNA.
[0029] 24. The system of embodiment 22, wherein the fusion nucleic acid comprises DNA.
[0030] 25. The system of any of the preceding embodiments, wherein (b) comprises a template RNA.
[0031] 26. The system of embodiment 25, wherein the template RNA further comprises a nuclear localization signal.
[0032] 27. The system of any of the preceding embodiments, wherein (a) comprises RNA encoding said polypeptide.
[0033] 28. The system of embodiment 27, wherein the RNA of (a) and the RNA of (b) are separate RNA molecules.
[0034] 29. The system described in embodiment 28, wherein the RNA of (a) and the RNA of (b) are present in a ratio of 10:1 to 5:1, 5:1 to 2:1, 2:1 to 1:1, 1:1 to 1:2, 1:2 to 1:5, or 1:5 to 1:10.
[0035] 30. The system of embodiment 28, wherein the RNA of (a) does not contain a nuclear localization signal.
[0036] 31. The system of any of the previous embodiments, wherein the polypeptide further comprises a nuclear localization signal and / or a nucleolar localization signal.
[0037] 32. The system of any of the preceding embodiments, wherein (a) comprises: (i) the polypeptide and (ii) RNA encoding a nuclear localization signal and / or a nucleolar localization signal.
[0038] 33. The system of any of the preceding embodiments, wherein the RNA comprises a pseudoknot sequence, e.g., 5' to the heterologous sequence of interest.
[0039] 34. The system of embodiment 33, wherein the RNA comprises a stem-loop sequence or a helix on the 5' side of the pseudoknot sequence.
[0040] 35. The system of embodiment 33 or 34, wherein the RNA comprises one or more (e.g., two, three, or more) stem-loop sequences or helices 3' to the pseudoknot sequence, e.g., 3' to the pseudoknot sequence and 5' to the heterologous sequence of interest.
[0041] 36. The system described in any one of embodiments 33 to 35, wherein the template RNA containing the pseudoknot has catalytic activity, for example, RNA cleavage activity, for example, cis-RNA cleavage activity.
[0042] 37. The system of any of the previous embodiments, wherein the RNA comprises at least one stem-loop sequence or helix, e.g., 1, 2, 3, 4, 5 or more stem-loop sequences, hairpins or helix sequences, e.g., 3' to the heterologous sequence of interest.
[0043] 38. Any of the above numbered systems, wherein the polypeptide comprises a sequence of at least 50 amino acids (e.g., at least 100, 150, 200, 300, 500 amino acids) having at least 80% identity (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity) to a polypeptide sequence listed in Table 1, or a reverse transcriptase domain or endonuclease domain thereof.
[0044] 39. Any of the above numbered systems, wherein the polypeptide comprises a sequence of at least 50 amino acids (e.g., at least 100, 150, 200, 300, 500 amino acids) having at least 80% identity (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity) to a polypeptide sequence listed in any of Tables 2-3, or a reverse transcriptase domain, endonuclease domain, or DNA binding domain thereof.
[0045] 40. Any of the above-numbered systems, wherein the polypeptide comprises a sequence of at least 50 amino acids (e.g., at least 100, 150, 200, 300, 500 amino acids) having at least 80% identity (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity) to an amino acid sequence in Table 3, column 8, or a reverse transcriptase domain, endonuclease domain, or DNA binding domain thereof.
[0046] 41. The template RNA is a sequence of Table 3 (e.g., one or both of the 5' untranslated region of Table 3, column 6 and the 3' untranslated region of Table 3, column 7), or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, any of the above numbered systems.
[0047] 42. The system of embodiment 41, wherein the template RNA comprises a sequence of about 100-125 bp from the 3' untranslated region of Table 3, column 7, for example, the sequence comprises nucleotides 1-100, 101-200, or 201-325 of the 3' untranslated region of Table 3, column 7, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0048] 43. Any of the above numbered systems, wherein (a) comprises RNA and (b) comprises RNA.
[0049] 44. Any of the above numbered systems, which contain only RNA or which contain more RNA than DNA, with an RNA:DNA ratio of at least 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1.
[0050] 45. Any of the above numbered systems that are free of DNA or contain more than 10%, 5%, 4%, 3%, 2%, or 1% DNA on a mass or molar basis.
[0051] 46. Any of the above numbered systems, which allows DNA to be modified by insertion of a heterologous sequence of interest without interfering with DNA-dependent RNA polymerization of (b).
[0052] 47. Any of the above numbered systems, in which DNA can be modified by insertion of a heterologous sequence of interest in the presence of an inhibitor of a DNA repair pathway (e.g., SCR7, PARP inhibitors) or in a cell line with a defective DNA repair pathway (e.g., a cell line with a defective nucleotide excision repair pathway or a defective homologous recombination repair pathway).
[0053] 48. Any of the above numbered systems, which do not cause the formation of detectable levels of double-stranded breaks in target cells.
[0054] 49. Any of the above numbered systems in which reverse transcriptase activity can be used to modify DNA, optionally in the absence of homologous recombination activity.
[0055] 50. Any of the above numbered systems, wherein the template RNA has been treated to reduce secondary structure, e.g., heated to a temperature that reduces secondary structure, e.g., at least 70, 75, 80, 85, 90, or 95°C.
[0056] 51. The system of embodiment 50, wherein the template RNA is subsequently cooled, for example to a temperature that allows secondary structure, for example below 30, 25, or 20°C.
[0057] 52. Below: (a) a polypeptide or a nucleic acid encoding a polypeptide, the polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain; and (b) a template RNA (or a DNA encoding the template RNA) comprising: (i) a sequence that binds to the above polypeptide; (ii) a heterologous target sequence; (iii) a first homology domain at the 5' end of the template RNA, the first homology domain having at least 10 bases of 100% identity to a target DNA strand; and (iv) a second homology domain at the 5' end of the template RNA, the second homology domain having at least 10 bases of 100% identity to a target DNA strand. A system for modifying DNA comprising:
[0058] 53. The system of any of the preceding embodiments, wherein (a) and (b) are part of the same nucleic acid.
[0059] 54. A system described in any of embodiments 1 to 52, wherein (a) and (b) are separate nucleic acids.
[0060] 55. Any of the above numbered systems, wherein the template RNA comprises at least 10 bases of 100% identity to the target DNA strand at the 5' end of the template RNA (e.g., the target DNA strand is a human DNA sequence).
[0061] 56. The system of any of the preceding embodiments, wherein the template RNA comprises at least 10 bases of 100% identity to the target DNA strand at the 3' end of the template RNA (e.g., the target DNA strand is a human DNA sequence).
[0062] 57. A host cell (e.g., a mammalian cell, e.g., a human cell) comprising any of the preceding numbered systems.
[0063] 58. A method of modifying a target DNA strand in a cell, tissue or subject, comprising administering to the cell, tissue or subject a system of any preceding number, wherein said system reverse transcribes a template RNA sequence into a target DNA strand, thereby modifying the target DNA strand.
[0064] 59. The method of embodiment 58, wherein the cell, tissue or subject is a mammalian (e.g., human) cell, tissue or subject.
[0065] 60. The method of any of the preceding embodiments, wherein said cells are fibroblasts.
[0066] 61. The method of any of the preceding embodiments, wherein said cells are primary cells.
[0067] 62. The method of any of the preceding embodiments, wherein said cells are not immortalized.
[0068] 63. A method for modifying the genome of a mammalian cell, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, said polypeptide comprising (i) a reverse transcriptase domain, (ii) an endonuclease domain, and optionally (iii) a DNA binding domain; and (b) a template RNA (or a DNA encoding the template RNA) comprising (i) a sequence that binds to the above polypeptide, and (ii) a heterologous sequence of interest. The method comprises the step of contacting the
[0069] 64. The method of embodiment 63, wherein the polypeptide does not comprise a target DNA binding domain.
[0070] 65. The method of embodiment 63, wherein the polypeptide is derived from an APE-type transposon reverse transcriptase.
[0071] 66. The method of embodiment 63, wherein (i) the reverse transcriptase domain, (ii) the endonuclease domain, or both (i) and (ii) have a sequence in Table 1 or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, 100% identity thereto.
[0072] 67. The method of embodiment 63, wherein the polypeptide further comprises a target DNA binding domain.
[0073] 68. A method for modifying the genome of a mammalian cell, comprising: (a) RNA encoding a polypeptide, said polypeptide comprising (i) a reverse transcriptase domain, (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; and (b) a template RNA comprising (i) a sequence that binds to the polypeptide and (ii) a heterologous sequence of interest; contacting the wherein the method does not include a step of contacting a mammalian cell with DNA or wherein the compositions of (a) and (b) do not contain more than 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% DNA based on mass or molar amount of nucleic acid.
[0074] 69. The method of embodiment 68, which results in the addition of at least 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of an exogenous DNA sequence to the genome of the mammalian cell.
[0075] 70. The method of embodiment 68 or 69, which results in the addition of a protein-coding sequence to the genome of the mammalian cell.
[0076] 71. A method for inserting DNA into the genome of a mammalian cell, comprising contacting the cell with an RNA composition, the RNA composition comprising: (a) a first RNA that directs the insertion of a template RNA into the genome; and (b) A template RNA containing a heterologous sequence Including, wherein the method does not include a step of contacting the mammalian cell with DNA, or wherein the compositions of (a) and (b) do not contain more than 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% DNA by mass or molar amount of nucleic acid; The method results in the addition of at least 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of a DNA (e.g., exogenous DNA) sequence to the genome of the mammalian cell.
[0077] 72. The method of embodiment 71, wherein the first RNA encodes a polypeptide (e.g., a polypeptide of any of Tables 1, 2, or 3 herein), wherein the polypeptide directs the insertion of a template RNA into a genome.
[0078] 73. The method of embodiment 72, wherein the template RNA further comprises a sequence that binds to the polypeptide.
[0079] 74. A method for adding at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, 1000 bp of exogenous DNA to the genome of a mammalian cell without delivering the DNA to said cell.
[0080] 75. A method for adding at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, 1000 bp of exogenous DNA to the genome of a mammalian cell, the method not including a step of contacting the mammalian cell with DNA or the method including a step of contacting the mammalian cell with a composition comprising less than 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% DNA by mass or molar amount of nucleic acid.
[0081] 76. A method for adding at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, 1000 bp of exogenous DNA to the genome of a mammalian cell, comprising delivering only RNA to said mammalian cell.
[0082] 77. A method for adding at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, 1000 bp of exogenous DNA to the genome of a mammalian cell, comprising the step of delivering RNA and protein to said mammalian cell.
[0083] 78. The method of any one of embodiments 68 to 77, wherein the template RNA serves as a template for insertion of the exogenous DNA.
[0084] 79. The method of any one of embodiments 68 to 78, which does not include DNA-dependent RNA polymerization of exogenous DNA.
[0085] 80. The method of any one of embodiments 58-79, which results in the addition of at least 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of DNA to the genome of the mammalian cell.
[0086] 81. The method of any one of embodiments 68 to 80, wherein the RNA of (a) and the RNA of (b) are covalently linked, e.g., are part of the same transcript.
[0087] 82. The method of any one of embodiments 68 to 80, wherein the RNA of (a) and the RNA of (b) are separate RNAs.
[0088] 83. The method of any one of embodiments 58 to 82, which does not include a step of contacting the mammalian cell with template DNA.
[0089] 84. A method for modifying the genome of a human cell, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, said polypeptide comprising (i) a reverse transcriptase domain, (ii) an endonuclease domain, and optionally (iii) a DNA binding domain; and (b) a template RNA (or a DNA encoding the template RNA) comprising (i) a sequence that binds to the above polypeptide, and (ii) a heterologous sequence of interest. contacting the wherein said method achieves insertion of said heterologous sequence of interest into the genome of a human cell, The human cells do not exhibit upregulation of any DNA repair genes and / or tumor suppressor genes, or DNA repair genes and / or tumor suppressor genes are not upregulated by more than 10%, 5%, 2%, or 1%, for example, wherein upregulation is measured by RNA sequencing, for example, as described in Example 14.
[0090] 85. A method of adding an exogenous coding region to the genome of a cell (e.g., a mammalian cell), comprising contacting the cell with RNA comprising a non-coding strand of the exogenous coding region, optionally wherein the RNA does not comprise a coding strand of the exogenous coding region, and optionally wherein delivery comprises non-viral delivery.
[0091] 86. A method for expressing a polypeptide in a cell (e.g., a mammalian cell), comprising contacting the cell with RNA, wherein the RNA comprises a non-coding strand that is the reverse complement of a sequence that will encode the polypeptide, optionally wherein the RNA does not comprise a coding strand that encodes the polypeptide, and optionally wherein delivery comprises non-viral delivery.
[0092] 87. The method of any of embodiments 58 to 86, wherein the sequence to be inserted into the mammalian genome is a sequence exogenous to the mammalian genome.
[0093] 88. The method of any of embodiments 58 to 87, which operates independently of a DNA template.
[0094] 89. The method of any of embodiments 58 to 88, wherein the cell is part of a tissue.
[0095] 90. The method of any of embodiments 58 to 89, wherein the mammalian cell is euploid, not immortalized, part of an organism, a primary cell, non-dividing, a hepatic cell, or derived from a subject with a genetic disease.
[0096] 91. The method of any of embodiments 58-90, wherein the contacting step comprises contacting the cell with a plasmid, a virus, a virus-like particle, a virosome, a liposome, a vesicle, an exosome, or a lipid nanoparticle.
[0097] 92. The method of any one of embodiments 58-91, wherein the contacting step comprises using non-viral delivery.
[0098] 93. The method of any of embodiments 58-92, comprising contacting the cell with the template RNA (or DNA encoding the template RNA), wherein the template RNA comprises a non-coding strand of the exogenous coding region, optionally wherein the template RNA does not comprise a coding strand of the exogenous coding region, and optionally wherein delivery comprises non-viral delivery, thereby adding the exogenous coding region to the genome of the cell.
[0099] 94. The method of any of embodiments 58-93, comprising contacting the cell with the template RNA (or DNA encoding the template RNA), wherein the template RNA comprises a non-coding strand that is the reverse complement of a sequence that will encode a polypeptide, and optionally, the template RNA does not comprise a coding strand that encodes the polypeptide, and optionally, delivery comprises non-viral delivery, thereby expressing the polypeptide in the cell.
[0100] 95. The method of any of embodiments 63-94, wherein the contacting step comprises administering (a) and (b) to the subject, e.g., intravenously.
[0101] 96. The method of any of embodiments 63-95, wherein the contacting step comprises administering at least two doses of (a) and (b) to the subject.
[0102] 97. The method of any of embodiments 63-96, wherein the polypeptide reverse transcribes the template RNA sequence into the target DNA strand, thereby modifying the target DNA strand.
[0103] 98. The method of any of embodiments 63-97, wherein (a) and (b) are administered separately.
[0104] 99. The method of any of embodiments 63-97, wherein (a) and (b) are administered together.
[0105] 100. The method of any of embodiments 63 to 99, wherein the nucleic acid of (a) is not integrated into the genome of the host cell.
[0106] 101. A sequence that binds to the above polypeptide, comprising the following characteristics: (a) Located at the 3' end of the template RNA; (b) located at the 5' end of the template RNA; (b) are non-coding sequences; (c) is structural RNA; or (d) forming at least one hairpin loop structure; any preceding number method with one or more of the following:
[0107] 102. The method of any preceding number, wherein the template RNA further comprises a sequence comprising at least 20 nucleotides that are at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the target DNA strand.
[0108] 103. The method of any preceding number, wherein said template RNA further comprises a sequence comprising at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150 nucleotides that are at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to a target DNA strand.
[0109] 104. The method of any preceding number, wherein a sequence comprising at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150 nucleotides, or about 2-10, 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 10-100, or 2-100 nucleotides that is at least 80% identical to the target DNA strand, is located on the 3'-end of the template RNA.
[0110] 105. The method of any preceding number, wherein the template RNA further comprises a sequence comprising at least 100 nucleotides that are at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the target DNA strand, e.g., at the 3' end of the template RNA.
[0111] 106. The method of embodiment 104 or 105, wherein the site in the target DNA strand to which the sequence comprises at least 80% identity is proximal to (e.g., within about 0-10, 10-20, 20-30, 30-50, or 50-100 nucleotides of) the target site on the target DNA strand that is recognized (e.g., bound to and / or cleaved) by a polypeptide that comprises the endonuclease.
[0112] 107. A sequence comprising at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150 nucleotides, or about 2-10, 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 10-100, or 2-100 nucleotides, which is at least 80% identical to the target DNA strand, is present at the 3' end of the template RNA; Optionally, the site within the target DNA strand to which the sequence comprises at least 80% identity is proximal to (e.g., within about 0-10, 10-20, or 20-30 nucleotides of) a target site on the target DNA strand that is recognized (e.g., bound to and / or cleaved) by a polypeptide that comprises the endonuclease.
[0113] 108. The method of embodiment 107, wherein the target site is a site in the human genome that has the greatest identity to a native target site of the polypeptide that contains the nucleotides, for example, the site in the human genome has at least about 16, 17, 18, 19, or 20 nucleotides identical to the native target site.
[0114] 109. The method of any preceding number, wherein the template RNA has at least 3, 4, 5, 6, 7, 8, 9, or 10 bases of 100% identity to the target DNA strand.
[0115] 110. The method of any preceding number, wherein at least 3, 4, 5, 6, 7, 8, 9, or 10 bases of 100% identity to the target DNA strand are located at the 3' end of the template RNA.
[0116] 111. The method of any preceding number, wherein at least 3, 4, 5, 6, 7, 8, 9, or 10 bases of 100% identity to the target DNA strand are located at the 5' end of the template RNA.
[0117] 112. The method of any preceding number, wherein the template RNA comprises at least 3, 4, 5, 6, 7, 8, 9, or 10 bases at the 5' end of the template RNA that are 100% identical to the target DNA strand, and at least 3, 4, 5, 6, 7, 8, 9, or 10 bases at the 3' end of the template RNA that are 100% identical to the target DNA strand.
[0118] 113. The method of any of the preceding numbers, wherein the heterologous target sequence is 50 to 50,000 base pairs (e.g., 50 to 40,000 bp, 500 to 30,000 bp, 500 to 20,000 bp, 100 to 15,000 bp, 500 to 10,000 bp, 50 to 10,000 bp, 50 to 5,000 bp).
[0119] 114. The method of any of the preceding numbers, wherein the heterologous target sequence is at least 10, 25, 50, 100, 150, 200, 250, 300, 400, 500, 600, or 700 bp.
[0120] 115. The method of any of the preceding numbers, wherein the heterologous target sequence is at least 715, 750, 800, 950, 1,000, 2,000, 3,000, or 4,000 bp.
[0121] 116. The method of any of the preceding numbers, wherein the heterologous target sequence is less than 5,000, 10,000, 15,000, 20,000, 30,000, or 40,000 bp.
[0122] 117. The method of any of the preceding numbers, wherein the heterologous target sequence is less than 700, 600, 500, 400, 300, 200, 150, or 100 bp.
[0123] 118. The heterologous target sequence is as follows: (a) An open reading frame, e.g., a polypeptide, e.g., an enzyme (e.g., a lysosomal enzyme), a membrane protein, a blood factor, an exon, an intracellular protein (e.g., an organellar protein such as a cytoplasmic protein, a nuclear protein, a mitochondrial protein, or a lysosomal protein), an extracellular protein, a structural protein, a signal transduction protein, a regulatory protein, a transport protein, a sensor protein, a motor protein, a defense protein, or a storage protein-encoding sequence; (b) A non-coding and / or regulatory sequence, e.g., a sequence that binds to a transcriptional modulator, e.g., a promoter, an enhancer, an insulator; (c) splice acceptor site; (d) polyA site; (e) epigenetic modification sites; (f) Gene Expression Unit Any of the preceding number methods, including.
[0124] 119. The method of any preceding number, wherein the target DNA is a genomic safe harbor (GSH) site.
[0125] 120. The method of any preceding number, wherein said target DNA is a Natural Harbor™ site in the genome.
[0126] 121. The method of any preceding number, which achieves insertion of said heterologous sequence of interest into a target site within said genome at an average copy number of at least 0.01, 0.025, 0.05, 0.075, 0.1, 0.15, 0.2, 0.25, 0.3, 0.4, 0.5, 0.75, 1, 1.25, 1.5, 1.75, 2, 2.5, 3, 4, or 5 copies per genome.
[0127] 122. The method of any preceding number, which results in about 25-100%, 50-100%, 60-100%, 70-100%, 75-95%, 80-90% integrants into the target site in the genome uncleaved, as measured by an assay described herein, e.g., the assay of Example 6.
[0128] 123. The method of any preceding number, which achieves insertion of the heterologous sequence of interest into only one target site within the genome of the cell.
[0129] 124. The method of any preceding number, achieving insertion of said heterologous sequence of interest into a target site in a cell, wherein the inserted heterologous sequence of interest contains less than 10%, 5%, 2%, 1%, 0.5%, 0.2%, or 0.1% mutations (e.g., SNPs or one or more deletions, e.g., truncations or internal deletions) compared to the heterologous sequence of interest prior to insertion, e.g., as measured by the assay of Example 12.
[0130] 125. The method of any preceding number, achieving insertion of the heterologous sequence of interest into target sites in a plurality of cells, wherein less than 10%, 5%, 2%, or 1% of the inserted copies of the heterologous sequence of interest contain a mutation (e.g., a SNP or a deletion, e.g., a truncation or internal deletion), e.g., as measured by the assay of Example 12.
[0131] 126. The method of any preceding number, achieving insertion of said heterologous sequence of interest into the genome of a target cell, wherein the target cell exhibits no upregulation of p53 or less than 10%, 5%, 2%, or 1% upregulation of p53, e.g., wherein upregulation of p53 is measured by p53 protein levels, e.g., according to the method described in Example 30, or by levels of p53 phosphorylated at Ser15 and Ser20.
[0132] 127. The method of any preceding number, achieving insertion of said heterologous sequence of interest into the genome of a target cell, wherein the target cell does not exhibit any upregulation of DNA repair genes and / or tumor suppressor genes, or wherein DNA repair genes and / or tumor suppressor genes are not upregulated by more than 10%, 5%, 2%, or 1%, for example, wherein upregulation is measured by RNA sequencing, for example, according to the method described in Example 14.
[0133] 128. The method of any preceding number, which achieves insertion of the heterologous sequence of interest at the target site in about 1-80% of the cells, e.g., about 1-10%, 10-20%, 20-30%, 30-40%, 40-50%, 50-60%, 60-70%, or 70-80% of the cells (e.g., in copy number of one or more insertions) in a population of cells contacted with the system, as measured, e.g., using single-cell ddPCR, e.g., as described in Example 17.
[0134] 129. A method of any preceding number, which achieves insertion of the heterologous sequence of interest at a target site in about 1-80% of the cells, e.g., about 1-10%, 10-20%, 20-30%, 30-40%, 40-50%, 50-60%, 60-70%, or 70-80% of the cells, at a copy number of a single insertion, in a population of cells contacted with the system, as measured, e.g., using colony isolation and ddPCR, as described in Example 18.
[0135] 130. The method of any preceding number, achieving insertion of a heterologous sequence of interest into a target site (on-target insertion) at a higher rate than insertion into a non-target site (off-target insertion) within a population of cells, wherein the ratio of on-target insertion to off-target insertion is greater than 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, 200:1, 500:1, or 1,000:1, e.g., using the assay of Example 11.
[0136] 131. The method of any preceding number, wherein insertion of the heterologous sequence of interest is achieved in the presence of an inhibitor of a DNA repair pathway (e.g., SCR7, PARP inhibitor) or in a cell line deficient in a DNA repair pathway (e.g., a cell line deficient in a nucleotide excision repair pathway or a homologous recombination repair pathway).
[0137] 132. Any of the preceding numbered systems formulated as a pharmaceutical composition.
[0138] 133. A system of any preceding number incorporated into a pharma- ceutically acceptable carrier (e.g., vesicles, liposomes, natural or synthetic lipid bilayers, lipid nanoparticles, exosomes).
[0139] 134. A method for producing a system for modifying the genome of a mammalian cell, comprising: a) providing a template RNA according to any of the preceding embodiments, e.g., wherein the template RNA comprises (i) a sequence that binds to a polypeptide comprising a reverse transcriptase domain and an endonuclease domain, and (ii) a heterologous sequence of interest; b) treating the template RNA to restrict secondary structure, e.g., by heating the template RNA to, e.g., at least, 70, 75, 80, 85, 90, or 95° C.; and c) subsequently cooling the template RNA to a temperature that allows for secondary structure, e.g., below 30, 25, or 20° C. The method includes:
[0140] 135. The method of embodiment 134, further comprising contacting the template RNA with a polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain, or a nucleic acid (e.g., RNA) encoding the polypeptide.
[0141] 136. The method of embodiment 134 or 135, further comprising contacting said template RNA with a cell.
[0142] 137. The system or method of any of the previous embodiments, wherein said heterologous sequence of interest encodes a therapeutic polypeptide.
[0143] 138. The system or method of any of the previous embodiments, wherein said heterologous sequence of interest encodes a mammalian (e.g., human) polypeptide, or a fragment or variant thereof.
[0144] 139. The system or method of any of the preceding embodiments, wherein the heterologous sequence of interest encodes a protein, such as an enzyme (e.g., a lysosomal enzyme), a blood factor (e.g., Factor I, II, V, VII, X, XI, XII or XIII), a membrane protein, an exon, an intracellular protein (e.g., an organelle protein such as a cytoplasmic protein, a nuclear protein, a mitochondrial protein or a lysosomal protein), an extracellular protein, a structural protein, a signaling protein, a regulatory protein, a transport protein, a sensor protein, a motor protein, a defense protein, or a storage protein.
[0145] 140. The system or method of any of the previous embodiments, wherein the heterologous sequence of interest comprises a tissue-specific promoter or enhancer.
[0146] 141. The system or method of any of the previous embodiments, wherein said heterologous sequence of interest encodes a polypeptide of more than 250, 300, 400, 500, or 1,000 amino acids, optionally up to 1300 amino acids.
[0147] 142. A system or method according to any of the preceding embodiments, wherein the heterologous sequence of interest encodes a fragment of a mammalian gene, but not an entire mammalian gene, e.g., encoding one or more exons, but not a full-length protein.
[0148] 143. The system or method of any of the previous embodiments, wherein said heterologous sequence of interest encodes one or more introns.
[0149] 144. A system or method according to any of the preceding embodiments, wherein the heterologous sequence of interest is other than GFP, e.g. other than a fluorescent protein or other than a reporter protein.
[0150] 145. The system or method of any of the preceding embodiments, wherein the polypeptide comprises (i) a reverse transcriptase domain and (ii) an endonuclease domain, wherein one or both of (i) or (ii) are derived from an avian retrotransposase, e.g., a sequence of Table 2 or 3, or at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0151] 146. The system or method of any of the previous embodiments, wherein the polypeptide has at least 70%, 75%, 80%, 85%, 90%, or 95% of its activity at 37°C as compared to its activity at 25°C under otherwise similar conditions.
[0152] 147. The system or method of any of the previous embodiments, wherein the nucleic acid encoding the polypeptide and the template RNA or the nucleic acid encoding the template RNA are separate nucleic acids.
[0153] 148. The system or method of any of the previous embodiments, wherein the template RNA does not encode an active reverse transcriptase, e.g., comprises an inactivated mutant reverse transcriptase, e.g., as described in Example 1 or 2, or does not contain a reverse transcriptase sequence.
[0154] 149. The system or method of any of the previous embodiments, wherein the template RNA comprises one or more chemical modifications.
[0155] 150. A system or method according to any of the previous embodiments, wherein the heterologous sequence of interest is located between the promoter and the sequence that binds to the polypeptide.
[0156] 151. The system or method of any of the previous embodiments, wherein the promoter is located between the heterologous sequence of interest and the sequence that binds to the polypeptide.
[0157] 152. The system or method of any of the previous embodiments, wherein the heterologous sequence of interest comprises an open reading frame (or its reverse complement) in a 5' to 3' direction on the template RNA.
[0158] 153. The system or method of any of the previous embodiments, wherein the heterologous sequence of interest comprises an open reading frame (or its reverse complement) in a 3' to 5' direction on the template RNA.
[0159] 154. The system or method of any of the previous embodiments, wherein the polypeptide comprises (a) a reverse transcriptase domain and (b) an endonuclease domain, wherein at least one of (a) or (b) is heterologous.
[0160] 155. The system or method of any of the preceding embodiments, wherein the polypeptide comprises (a) a target DNA binding domain, (b) a reverse transcriptase domain, and (c) an endonuclease domain, wherein at least one of (a), (b) or (c) is heterologous.
[0161] 156. A substantially pure polypeptide comprising: (a) a reverse transcriptase domain; and (b) a heterologous endonuclease domain.
[0162] 157. A substantially pure polypeptide comprising (a) a target DNA binding domain, (b) a reverse transcriptase domain, and (c) an endonuclease domain, wherein at least one of (a), (b) or (c) is heterologous.
[0163] 158. A substantially pure polypeptide comprising (a) a reverse transcriptase domain, (b) an endonuclease domain, and (c) a heterologous target DNA binding domain.
[0164] 159. A polypeptide or a nucleic acid encoding said polypeptide, wherein the polypeptide comprises (a) a reverse transcriptase domain and (b) an endonuclease domain, wherein at least one of (a) or (b) is heterologous to the other.
[0165] 160. A polypeptide or a nucleic acid encoding said polypeptide, wherein the polypeptide comprises (a) a target DNA binding domain, (b) a reverse transcriptase domain, and (c) an endonuclease domain, wherein at least one of (a), (b) or (c) is heterologous to the others.
[0166] 161. Any of the polypeptides of embodiments 156-160, wherein the reverse transcriptase domain has at least 80% identity (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity) to a reverse transcriptase domain of an APE-type or RLE-type non-LTR retrotransposon listed in any of Tables 1-3.
[0167] 162. The polypeptide of any of embodiments 156-161, wherein the endonuclease domain has at least 80% identity, such as at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity, to an endonuclease domain of an APE-type or RLE-type non-LTR retrotransposon listed in any of Tables 1-3.
[0168] 163. Any of the polypeptides of embodiments 156-162 or any preceding numbered method, wherein the DNA binding domain has at least 80% identity, e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity, to a DNA binding domain of a sequence listed in any of Tables 1, 2, or 3.
[0169] 164. A nucleic acid encoding a polypeptide according to any preceding numbered embodiment.
[0170] 165. A vector comprising the nucleic acid according to embodiment 164.
[0171] 166. A host cell comprising a nucleic acid according to embodiment 164.
[0172] 167. A host cell comprising a polypeptide according to any preceding numbered embodiment.
[0173] 168. A host cell comprising the vector according to embodiment 165.
[0174] 169. A host cell (e.g., a human cell) comprising: (i) a heterologous sequence of interest (e.g., a sequence encoding a therapeutic polypeptide) at a target site in a chromosome; and (ii) one or both of an untranslated region on one side (e.g., upstream) of the heterologous sequence of interest (e.g., a retrotransposon untranslated sequence, e.g., a sequence in Table 3, column 6) and an untranslated region on the other side (e.g., downstream) of the heterologous sequence of interest (e.g., a retrotransposon untranslated sequence, e.g., a sequence in Table 3, column 7).
[0175] 170. A host cell (e.g., a human cell) comprising: (i) a heterologous sequence of interest (e.g., a sequence encoding a therapeutic polypeptide) at a target site within a chromosome, where the target locus is a Natural Harbor™ site, e.g., a site in Table 4 herein.
[0176] 171. The host cell of embodiment 170, further comprising one or both of: (ii) a 5' untranslated region of said heterologous sequence of interest; and (ii) a 3' untranslated region of said heterologous sequence of interest.
[0177] 172. The host cell of embodiment 170, further comprising one or both of the following: (ii) an untranslated region on one side (e.g., upstream) of the heterologous sequence of interest (e.g., a retrotransposon untranslated region, e.g., a sequence in Table 3, column 6), and an untranslated region on the other side (e.g., downstream) of the heterologous sequence of interest (e.g., a retrotransposon untranslated region, e.g., a sequence in Table 3, column 7).
[0178] 173. A host cell according to any of embodiments 169 to 173, which contains a heterologous sequence of interest only at the target site.
[0179] 174. A pharmaceutical composition comprising a system, nucleic acid, polypeptide, or vector of any preceding number; and a pharma- ceutically acceptable excipient or carrier.
[0180] 175. The pharmaceutical composition according to embodiment 174, wherein the pharma- ceutically acceptable excipient or carrier is selected from a vector (e.g., a virus or a plasmid vector), a vesicle (e.g., a liposome, an exosome, a natural or synthetic lipid bilayer), a lipid nanoparticle.
[0181] 176. The polypeptide of any of the previous embodiments, wherein said polypeptide further comprises a nuclear localization sequence.
[0182] 177. A method of modifying a target DNA strand in a cell, tissue or subject, comprising the step of administering to said cell, tissue or subject a system of any preceding number, thereby modifying the target DNA strand.
[0183] 178. The embodiment of any preceding number, wherein the polypeptide comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to an amino acid sequence listed in Table 5 (e.g., any one of SEQ ID NOs:1017-1022), or a functional fragment thereof.
[0184] 179. The embodiment of any preceding number, wherein the reverse transcriptase domain comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to an amino acid sequence listed in Table 5 (e.g., any one of SEQ ID NOs: 1017-1022), or a functional fragment thereof.
[0185] 180. The embodiment of any preceding number, wherein the retrotransposase comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to an amino acid sequence listed in Table 5 (e.g., any one of SEQ ID NOs: 1017-1022), or a functional fragment thereof.
[0186] 181. The embodiment of any preceding number, wherein the polypeptide comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to the amino acid sequence SGSETPGTSESATPES (SEQ ID NO:1023) or GGGS (SEQ ID NO:1024).
[0187] 182. The embodiment of any preceding number, wherein the reverse transcriptase domain comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to the amino acid sequence SGSETPGTSESATPES (SEQ ID NO:1023) or GGGS (SEQ ID NO:1024).
[0188] 183. The embodiment of any preceding number, wherein the retrotransposase comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to the amino acid sequence SGSETPGTSESATPES (SEQ ID NO:1023) or GGGS (SEQ ID NO:1024).
[0189] 184. The embodiment of any preceding number, wherein the polypeptide, reverse transcriptase domain, or retrotransposase comprises a linker comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to the amino acid sequence SGSETPGTSESATPES (SEQ ID NO:1023) or GGGS (SEQ ID NO:1024).
[0190] 185. The embodiment of any preceding number, wherein the polypeptide comprises a DNA-binding domain covalently linked to the remainder of the polypeptide by a linker, e.g., a linker comprising at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 125, 150, 200, 300, 400, or 500 amino acids.
[0191] 186. Embodiment 185, wherein the linker is attached to the remainder of the polypeptide at a position within the DNA-binding domain, the RNA-binding domain, the reverse transcriptase domain, or the endonuclease domain (e.g., as shown in any of Figures 17A-17F).
[0192] 187. Embodiment 185 or 186, wherein the linker is attached to the rest of the polypeptide at a position N-terminal to the alpha-helical region of the polypeptide, for example a position corresponding to version v1 as described in Example 26.
[0193] 188. Embodiment number 185 or 186, in which the linker is attached to the remainder of the polypeptide at a position C-terminal to the alpha-helical region of the polypeptide, e.g., preceding an RNA-binding motif (e.g., an a-1 RNA-binding motif), e.g., a position corresponding to version v2 as described in Example 26.
[0194] 189. Embodiment 185 or 186, wherein the linker is attached to the remainder of the polypeptide at a position C-terminal to the random coil region of the polypeptide, e.g., N-terminal to a DNA-binding motif (e.g., a c-myb DNA-binding motif), e.g., a position corresponding to version v3 as described in Example 26.
[0195] 190. Any one of embodiments 185-189, wherein the linker comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 1023) or GGGS (SEQ ID NO: 1024).
[0196] 191. The embodiment of any preceding number, wherein a polynucleotide sequence comprising at least about 500, 1000, 2000, 3000, 3500, 3600, 3700, 3800, 3900, or 4000 contiguous nucleotides from the 5' end of said template RNA sequence is integrated into a target cell genome.
[0197] 192. The embodiment of any preceding number, wherein a polynucleotide sequence comprising at least about 500, 1000, 2000, 2500, 2600, 2700, 2800, 2900, or 3000 contiguous nucleotides from the 3' end of the template RNA sequence is integrated into a target cell genome.
[0198] 193. The embodiment of any preceding number, wherein the nucleic acid sequence of the template RNA, or a portion thereof (e.g., a portion comprising at least about 100, 200, 300, 400, 500, 1000, 2000, 2500, 3000, 3500, or 4000 nucleotides), is integrated into the genome of a population of target cells at a copy number of at least about 0.21, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1.0 integrants / genome.
[0199] 194. The embodiment of any preceding number, wherein the nucleic acid sequence of the template RNA, or a portion thereof (e.g., a portion comprising at least about 100, 200, 300, 400, 500, 1000, 2000, 2500, 3000, 3500, or 4000 nucleotides), is integrated into the genome of a population of target cells at a copy number of at least about 0.085, 0.09, 0.1, 0.15, or 0.2 integrants / genome.
[0200] 195. The embodiment of any preceding number, wherein the nucleic acid sequence of the template RNA, or a portion thereof (e.g., a portion comprising at least about 100, 200, 300, 400, 500, 1000, 2000, 2500, 3000, 3500, or 4000 nucleotides), is integrated into the genome of a population of target cells at a copy number of at least about 0.036, 0.04, 0.05, 0.06, 0.07, or 0.08 integrants / genome.
[0201] 196. The embodiment of any preceding number, wherein the polypeptide comprises a functional endonuclease domain (e.g., the endonuclease domain does not comprise a mutation that abolishes endonuclease activity, e.g., as described herein).
[0202] 197. The embodiment of any preceding number, wherein the polypeptide comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to an R2 polypeptide (e.g., as described herein, e.g., R2-1_GFo) from a medium ground finch, e.g., a Galapagos finch (Geospiza fortis), or a functional fragment thereof.
[0203] 198. The embodiment of any preceding number, wherein the reverse transcriptase domain comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to an R2 polypeptide (e.g., as described herein, e.g., R2-1_GFo) from a medium ground finch, e.g., a Galapagos finch (Geospiza fortis), or a functional fragment thereof.
[0204] 199. The embodiment of any preceding number, wherein the retrotransposase comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to an R2 polypeptide from a medium ground finch, e.g., a Galapagos finch (Geospiza fortis) (e.g., R2-1_GFo, e.g., as described herein), or a functional fragment thereof.
[0205] 200. Any one of embodiments 197-199, wherein the nucleic acid sequence of the template RNA, or a portion thereof (e.g., a portion comprising at least about 100, 200, 300, 400, 500, 1000, 2000, 2500, 3000, 3500, or 4000 nucleotides), is integrated into the genome of a population of target cells at a copy number of at least about 0.21 integrants / genome.
[0206] 201. The embodiment of any preceding number, wherein the polypeptide comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to an R4 polypeptide (e.g., R4_AL, as described herein) from a large roundworm, e.g., Ascaris lumbricoides, or a functional fragment thereof.
[0207] 202. The embodiment of any preceding number, wherein the reverse transcriptase domain comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to an R4 polypeptide (e.g., as described herein, e.g., R4_AL) from a large roundworm, e.g., Ascaris lumbricoides, or a functional fragment thereof.
[0208] 203. The embodiment of any preceding number, wherein the retrotransposase comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to an R4 polypeptide (e.g., as described herein, e.g., R4_AL) from a large roundworm, e.g., Ascaris lumbricoides, or a functional fragment thereof.
[0209] 204. Any one of embodiments 201-203, wherein the nucleic acid sequence of the template RNA, or a portion thereof (e.g., a portion comprising at least about 100, 200, 300, 400, 500, 1000, 2000, 2500, 3000, 3500, or 4000 nucleotides), is integrated into the genome of a population of target cells at a copy number of at least about 0.085 integrants / genome.
[0210] 205. The embodiment of any preceding number, wherein introduction of the system into a target cell does not result in alteration (e.g., upregulation) of p53 and / or p21 protein levels, H2AX phosphorylation (e.g., γH2AX), ATM phosphorylation, ATR phosphorylation, Chk1 phosphorylation, Chk2 phosphorylation, and / or p53 phosphorylation.
[0211] 206. The embodiment of any preceding number, wherein introduction of the system into a target cell results in upregulation of p53 protein levels in the target cell to a level that is less than about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the p53 protein level induced by introduction of a site-specific nuclease, e.g., Cas9, that targets the same genomic site as the system.
[0212] 207. Embodiment 205 or 206, wherein the p53 protein level is determined according to the method described in Example 30.
[0213] 208. The embodiment of any preceding number, wherein introduction of the system into a target cell results in upregulation of p53 phosphorylation levels in the target cell genome to a level that is less than about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the p53 phosphorylation level induced by introduction of a site-specific nuclease, e.g., Cas9, that targets the same genomic site as the system.
[0214] 209. The embodiment of any preceding number, wherein introduction of the system into a target cell results in upregulation of p21 protein levels in the target cell genome to a level that is less than about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the p53 protein level induced by introduction of a site-specific nuclease, e.g., Cas9, that targets the same genomic site as the system.
[0215] 210. Embodiment 205 or 209, wherein the p21 protein level is determined according to the method described in example 30.
[0216] 211. The embodiment of any preceding number, wherein introduction of the system into a target cell results in upregulation of H2AX phosphorylation levels in the target cell genome to less than about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the H2AX phosphorylation level induced by introduction of a site-specific nuclease, e.g., Cas9, that targets the same genomic site as the system.
[0217] 212. The embodiment of any preceding number, wherein introduction of the system into a target cell results in upregulation of ATM phosphorylation levels in the target cell genome to a level that is less than about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the ATM phosphorylation level induced by introduction of a site-specific nuclease, e.g., Cas9, that targets the same genomic site as the system.
[0218] 213. The embodiment of any preceding number, wherein introduction of the system into a target cell results in upregulation of ATR phosphorylation levels in the target cell genome to a level that is less than about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the ATR phosphorylation level induced by introduction of a site-specific nuclease, e.g., Cas9, that targets the same genomic site as the system.
[0219] 214. The embodiment of any preceding number, wherein introduction of the system into a target cell results in upregulation of Chk1 phosphorylation levels in the target cell genome to a level that is less than about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the Chk1 phosphorylation level induced by introduction of a site-specific nuclease, e.g., Cas9, that targets the same genomic site as the system.
[0220] 215. The embodiment of any preceding number, wherein introduction of the system into a target cell results in upregulation of Chk2 phosphorylation levels in the target cell genome to less than about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the Chk2 phosphorylation level induced by introduction of a site-specific nuclease, e.g., Cas9, that targets the same genomic site as the system.
[0221] definition Domain: As used herein, the term "domain" refers to a structure of a biomolecule that contributes to a specific function of the biomolecule. A domain can include a continuous region (e.g., a continuous sequence) or a discrete non-contiguous region (e.g., a non-contiguous sequence) of a biomolecule. Examples of protein domains include, but are not limited to, endonuclease domains, DNA binding domains, reverse transcriptase domains; an example of a nucleic acid domain is a regulatory domain, e.g., a transcription factor binding domain.
[0222] Exogenous: As used herein, the term "exogenous" when used in reference to a biomolecule (such as a nucleic acid sequence or a polypeptide) means that the biomolecule has been introduced into a host genome, cell, or organism by human hands. For example, a nucleic acid that is added to an existing genome, cell, tissue, or subject using recombinant DNA technology or other methods is exogenous to the existing nucleic acid sequence, cell, tissue, or subject.
[0223] Genomic Safe Harbor Site (GSH Site): A genomic safe harbor site is a site within a host genome that can accommodate the integration of new genetic material, such that the inserted genetic element does not cause significant alterations of the host genome that pose a risk to the host cell or organism. A GSH site generally meets 1, 2, 3, 4, 5, 6, 7, 8, or 9 of the following criteria: (i) located >300kb from a cancer-related gene; (ii) located >300kb from a miRNA / other functional small RNA; (iii) located >50kb from the 5' gene end; (iv) located >50kb from a replication origin; (v) >50kb away from an ultraconserved element; (vi) has low transcriptional activity (i.e., no mRNA + / - 25kb); (vii) is not within a variable copy number region; (viii) is within open chromatin; and / or (ix) is unique with one copy in the human genome. Examples of GSH sites in the human genome that meet some or all of these criteria include: (i) adenovirus site 1 (AAVS1), the naturally occurring integration site of the AAV virus on chromosome 19; (ii) the chemokine (CC motif) receptor 5 (CCR5) gene, a chemokine receptor gene known as the HIV-1 co-receptor; (iii) the human orthologue of the mouse Rosa26 locus; (iv) the rDNA locus. Additional GSH sites are known and are described, for example, in Pellenz et al. epub August 20, 2018 (https: / / doi.org / 10.1101 / 396390).
[0224] Heterologous: The term "heterologous", when used to describe a first element in relation to a second element, means that the first and second elements do not naturally occur in the arrangement as described. For example, a heterologous polypeptide, nucleic acid molecule, construct or sequence refers to (a) a polypeptide, nucleic acid molecule, or a portion of a polypeptide or nucleic acid molecule sequence that is not native to the cell in which it is expressed, (b) a polypeptide or nucleic acid molecule or a portion of a polypeptide or nucleic acid molecule that is modified or mutated relative to its natural state, or (c) a polypeptide or nucleic acid molecule that has an altered expression compared to the native expression level under similar conditions. For example, a heterologous regulatory sequence (e.g., promoter, enhancer) can be used to regulate the expression of a gene or nucleic acid molecule in a manner that is different from that in which the gene or nucleic acid molecule is normally expressed in nature. In another example, a heterologous domain of a polypeptide or nucleic acid sequence (e.g., a DNA-binding domain of a polypeptide or a nucleic acid encoding a DNA-binding domain of a polypeptide) may be positioned relative to other domains or may be of a different sequence or from a different source relative to other domains or portions of a polypeptide or its encoding nucleic acid. In certain embodiments, a heterologous nucleic acid molecule may be present naturally within the host cell genome, but may have an altered expression level or a different sequence, or both. In other embodiments, a heterologous nucleic acid molecule may not be endogenous to the host cell or host genome, but may instead be introduced into the host cell by transformation (e.g., transfection, electroporation), where the added molecule may be integrated into the host genome or may exist as extrachromosomal genetic material, either transiently (e.g., mRNA) or semi-stable for two or more generations (e.g., episomal viral vectors, plasmids or other self-replicating vectors).
[0225] Mutation or Mutation: The term "(mutation)" when applied to a nucleic acid sequence means that nucleotides within a nucleic acid sequence may be inserted, deleted, or changed relative to a reference (e.g., naturally occurring) nucleic acid sequence. A single alteration may be made at a locus (point mutation), or multiple nucleotides may be inserted, deleted, or changed at a single locus. In addition, one or more alterations may be made at any number of loci within a nucleic acid sequence. Nucleic acid sequences can be mutated by any method known in the art.
[0226] Nucleic acid molecule: Nucleic acid molecule refers to both RNA and DNA molecules, including but not limited to cDNA, genomic DNA and mRNA, and also synthetic nucleic acid molecules, such as those chemically synthesized or recombinantly produced, such as RNA templates, as described herein. Nucleic acid molecules may be double-stranded or single-stranded, circular or linear. If single-stranded, the nucleic acid molecule may be the sense or antisense strand. Unless otherwise stated, and as an example of all sequences described herein in the general format "SEQ ID NO:", a nucleic acid containing "SEQ ID NO:1" refers to a nucleic acid having at least a portion thereof either (i) the sequence of SEQ ID NO:1, or (ii) a sequence complementary to SEQ ID NO:1. The choice between the two is determined by the context in which SEQ ID NO:1 is used. For example, if the nucleic acid is used as a probe, the choice between the two is determined by the requirement that the probe be complementary to the desired target. The nucleic acid sequences of the present disclosure may be chemically or biochemically modified or contain non-natural or derivatized nucleotide bases, as will be readily understood by those skilled in the art. Such modifications include, for example, labels, methylation, replacement of one or more naturally occurring nucleotides with analogs, internucleotide modifications such as uncharged bonds (e.g., methylphosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged bonds (e.g., phosphorothioates, phosphorodithioates, etc.), pendant moieties (e.g., polypeptides), intercalating agents (e.g., acridines, psoralens, etc.), chelators, alkylators, and modified linkages (e.g., α-anomeric nucleic acids, etc.). Also included are synthetic molecules that mimic polynucleotides in their ability to bind to designated sequences through hydrogen bonds and other chemical interactions. Such molecules are known in the art, and include, for example, peptide bonds in place of phosphate bonds in the backbone of the molecule. Other modifications can include, for example, analogs in which the ribose ring contains bridging moieties or other structures such as modifications found in "locked" nucleic acids.
[0227] Gene Expression Unit: A gene expression unit is a nucleic acid sequence that includes at least one regulatory nucleic acid sequence operably linked to at least one effector sequence. A first nucleic acid sequence is operably linked to a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For example, a promoter or enhancer is operably linked to a coding sequence if the promoter or enhancer affects the transcription or expression of the coding sequence. Operably linked DNA sequences can be contiguous or non-contiguous. Where necessary to link two protein coding regions, operably linked sequences may be in the same reading frame.
[0228] Host: The term host genome or host cell, as used herein, refers to a cell and / or its genome into which proteins and / or genetic material have been introduced. Such terms refer not only to the particular subject cell and / or genome, but also to the progeny of such a cell and / or the genome of the progeny of such a cell. It is understood that such progeny may not in fact be identical to the parent cell, since certain modifications may occur in subsequent generations due to mutations or environmental influences, but are still included within the scope of the term "host cell" as used herein. The host genome or host cell may be an isolated cell or cell line grown in culture, or genomic material isolated from such a cell or cell line, or may be a host cell or host genome that constitutes a living tissue or organism. In some cases, the host cell may be an animal cell or a plant cell, for example, as described herein. In certain examples, the host cell may be a bovine cell, a horse cell, a porcine cell, a caprine cell, a sheep cell, a chicken cell, or a turkey cell. In certain examples, the host cell may be a corn cell, a soybean cell, a wheat cell, or a rice cell.
[0229] Pseudoknot: "Pseudoknot sequence" as used herein refers to a nucleic acid (e.g., RNA) having a sequence with suitable self-complementarity to form a pseudoknot structure, e.g., a first segment, a second segment between the first and third segments (wherein the third segment is complementary to the first segment), and a fourth segment (wherein the fourth segment is complementary to the second segment). The pseudoknot may optionally have additional secondary structures, e.g., a stem-loop located within the second segment, a stem-loop located between the second and third segments, a sequence before the first segment, or a sequence after the fourth segment. The pseudoknot may have additional sequences between the first and second segments, between the second and third segments, or between the third and fourth segments. In some embodiments, the segments are arranged 5' to 3': first, second, third, and fourth. In some embodiments, the first and third segments include five base pairs of perfect complementarity. In some embodiments, the second and fourth segments comprise 10 base pairs, optionally with one or more (e.g., two) bulges. In some embodiments, the second segment comprises one or more unpaired nucleotides, e.g., forming a loop. In some embodiments, the third segment comprises one or more unpaired nucleotides, e.g., forming a loop.
[0230] Stem-loop sequence: As used herein, "stem-loop sequence" refers to a nucleic acid sequence (e.g., an RNA sequence) with sufficient self-complementarity to form a stem-loop, e.g., having a stem containing at least two (e.g., 3, 4, 5, 6, 7, 8, 9, or 10) base pairs and a loop containing at least three (e.g., 4) base pairs. The stem may contain a mismatch or a bulge. In an embodiment of the present invention, for example, the following items are provided: (Item 1) below: (a) a polypeptide or a nucleic acid encoding a polypeptide, the polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain, wherein one or both of (i) or (ii) are derived from an avian retrotransposase; and (b) a template RNA (or a DNA encoding the template RNA) comprising (i) a sequence that binds to the polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising: (Item 2) 2. The system according to item 1, wherein the bird retrotransposase is a retrotransposase from a zebra finch (Taeniopygia guttata), a Galapagos finch (Geospiza fortis), a white-throated sparrow (Zonotrichia albicollis), or a white-throated tinamous (Tinamus guttatus). (Item 3) below: (a) a polypeptide or a nucleic acid encoding a polypeptide, said polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain, wherein at least one of (i) or (ii) is heterologous; and (b) a template RNA (or a DNA encoding the template RNA) comprising (i) a sequence that binds to the polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising: (Item 4) below: (a) a polypeptide or a nucleic acid encoding a polypeptide, the polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain; and (b) a template RNA (or a DNA encoding the template RNA) comprising: (i) a sequence that binds to the polypeptide; (ii) a heterologous sequence of interest; (iii) a first homology domain at the 5' end of the template RNA, the first homology domain having at least 100% identity with a target DNA strand; and (iv) a second homology domain at the 5' end of the template RNA, the second homology domain having at least 100% identity with a target DNA strand. A system for modifying DNA comprising: (Item 5) 5. The system according to any one of items 1 to 4, wherein (a) comprises RNA and (b) comprises RNA. (Item 6) 6. The system according to any one of items 1 to 5, wherein (a) and (b) are part of the same nucleic acid. (Item 7) 6. The system according to any one of items 1 to 5, wherein (a) and (b) are separate nucleic acids. (Item 8) 8. The system of any one of items 1 to 7, comprising only RNA or comprising more RNA than DNA, with an RNA:DNA ratio of at least 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1. (Item 9) 9. The system according to any one of items 1 to 8, wherein the heterologous sequence of interest comprises an open reading frame in the 5' to 3' direction in the template RNA. (Item 10) 9. The system according to any one of items 1 to 8, wherein the heterologous sequence of interest comprises an open reading frame in the 3' to 5' direction in the template RNA. (Item 11) 11. The system according to any one of items 1 to 10, wherein the sequence that binds to the polypeptide is a 3' untranslated sequence. (Item 12) 12. The system of claim 11, wherein the template RNA further comprises a 5' untranslated sequence. (Item 13) 13. The system according to any one of items 1 to 12, wherein the template RNA further comprises a promoter operably linked to the heterologous sequence of interest. (Item 14) 14. The system of item 13, wherein the promoter is positioned between the 5' untranslated sequence and the heterologous sequence of interest. (Item 15) 14. The system of item 13, wherein the promoter is positioned between the 3' untranslated sequence that binds the polypeptide and the heterologous sequence of interest. (Item 16) 16. The system according to any one of items 11 to 15, wherein the 5' untranslated sequence is a sequence in Table 3, column 6, or a sequence having at least 80% identity thereto. (Item 17) 17. The system according to any one of items 11 to 16, wherein the 3' untranslated sequence is a sequence in Table 3, column 7, or a sequence having at least 80% identity thereto. (Item 18) 18. The system according to any one of items 1 to 17, wherein the heterologous sequence of interest comprises an enzyme, a membrane protein, a blood factor, an intracellular protein, an extracellular protein, a structural protein, a signal transduction protein, a regulatory protein, a transport protein, a sensor protein, a motor protein, a defense protein, or a storage protein. (Item 19) 19. The system according to any one of items 1 to 18, wherein the template RNA comprises at least 10 bases at the 5' end of the template RNA that are 100% identical to a target DNA strand. (Item 20) 20. The system according to any one of items 1 to 19, wherein the template RNA comprises at least 10 bases at the 3' end of the template RNA that are 100% identical to a target DNA strand. (Item 21) 21. A method for modifying a target DNA strand in a cell, comprising the step of administering to the cell a system according to any one of items 1 to 20, thereby modifying the target DNA strand. (Item 22) 22. The method of claim 21, which achieves the addition of at least 5 base pairs of an exogenous DNA sequence to the genome of the cell. (Item 23) 22. The method of claim 21, which achieves the addition of at least 100 base pairs of an exogenous DNA sequence to the genome of the cell. (Item 24) 24. The method according to any one of items 21 to 23, achieving insertion of the heterologous sequence of interest into the target DNA at an average copy number of at least 0.01, 0.05, or 0.5 copies per genome. (Item 25) 25. The method according to any one of items 21 to 24, achieving about 50 to 100% incorporation of the heterologous sequence of interest into the target DNA that is uncleaved. (Item 26) 26. The method according to any one of items 21 to 25, wherein the nucleic acid of (a) is not integrated into the genome of the cell. (Item 27) 27. The method according to any one of items 21 to 26, wherein the template RNA comprises at least 10 bases at the 5' end of the template RNA that are 100% identical to the target DNA strand. (Item 28) 28. The method according to any one of items 21 to 27, wherein the template RNA comprises at least 10 bases at the 3' end of the template RNA that are 100% identical to the target DNA strand. [Brief description of the drawings]
[0231] [Figure 1] FIG. 1 is a schematic diagram of the Gene Writing genome editing system. [Diagram 2] FIG. 1 is a schematic diagram of the structure of the Gene Writer genome editor polypeptide. [Diagram 3] Schematic diagram of the Gene Writer genome editor polypeptide containing heterologous DNA binding domains designed to target different sites in the genome. [Figure 4] Schematic diagram of the Gene Writer genome editor template RNA structure. [Diagram 5] FIG. 1 is a schematic diagram showing the Gene Writing genome editing system for adding gene expression units to safe harbor sites in a genome. [Figure 6]FIG. 1 is a schematic diagram showing Gene Writing genome editing to add new exons to specific introns in the genome to replace downstream exons. [Figure 7] A schematic of Miseq library construction is shown. Nested PCR was performed across the R2Tg-rDNA junction using (1) an external forward primer and a tailed internal reverse primer, followed by (2) a tailed internal forward primer and a tailed reverse primer. The internal reverse primer contains a 1-4 base stagger, an 8 nucleotide randomized UMI, and a multiplexed barcode. The UMI allows counting of individual amplification events to eliminate PCR bias. [Figure 8] 8A-8B show the results of Miseq and Matlab analysis of DNA-mediated R2Tg integration into Hek293T cells. Each graph shows the analysis of the experimental R2Tg (FIG. 8A) and the 1 bp deletion negative control. The y-axis shows the number of unique sequence alignments determined by the unique UMIs found via Matlab. The x-axis shows the sequence position of the sequence coverage. The vertical grey line on the left side of the graph shows the location of the forward primer, and the vertical grey line on the right side of the graph shows the expected Tg-rDNA junction site. The bar on the right side of the graph shows an insertion without a break, and the bar on the left side of the graph shows a break. FIG. 8A reveals that most sequences show a high degree of alignment with the expected integration product. [Figure 9] Figure 1 shows ddPCR assessment of copy number changes of R2Tg-rDNA junction in human cells between transfection conditions. The forward primer and probe were predicted to bind to the 3'UTR of R2Tg, whereas the reverse primer was targeted to human rDNA. The resulting ddPCR signal was normalized to that of the reference assay RPP30 to determine copy number. A significantly higher average copy number per genome was observed for wild-type (WT, set of bars on the left) R2Tg compared to a gene control with a 1-bp deletion altering translation (frameshift mutant control, set of bars on the right). [Figure 10]Sequence alignment and coverage of TOPO cloning nested PCR products from Example 7. The grey bars on the right side of the graph indicate the expected transgene-rDNA junctions. Most sequences show a high degree of alignment with the expected integration products. [Figure 11] 1 is a schematic diagram of an exemplary template RNA, which contains a central payload domain (e.g., including a heterologous sequence of interest, e.g., a promoter and a protein coding sequence). The payload domain is flanked by 5' and 3' protein interaction domains, e.g., sequences that can bind to a Gene Writer polypeptide, e.g., the 5' and 3' UTR sequences shown in Table 3. Flanking the protein interaction domain are 5' and 3' homology domains, which have homology to the desired insertion region in the genome. [Figure 12] Graph showing retrotransposition efficiency measured by ddPCR (digital droplet PCR) using various transfection conditions. Bars A-C represent samples containing 100 ng, 250 ng, or 500 ng, respectively, transfected with 0.15 μL Lipofectamine™ RNAiMAX. Bars D-F represent samples containing 100 ng, 250 ng, or 500 ng, respectively, transfected with 0.3 μL Lipofectamine™ RNAiMAX. Bars G-I represent samples containing 100 ng, 250 ng, or 500 ng, respectively, transfected with 1 μL TransIT®-mRNA Transfection Kit. [Figure 13]Schematic diagram of the trans-transgene delivery mechanism. The schematic shows a driver plasmid (left) with a pCEP4 backbone, which encodes the reverse transcriptase R2Tg, with a promoter and Kozak sequence upstream and a polyadenylation signal downstream. The driver plasmid can drive expression of the GeneWriter protein. The transgene plasmid (right) with a pCDNA backbone contains (in order): a CMV promoter, a rDNA homology sequence, a 5'UTR, an antisense-oriented insert, a 3'UTR, a second rDNA homology sequence, a second polyadenylation signal, and a TK promoter driving an mKate2 marker. The antisense-oriented insert contains an EF1α promoter, the coding region for EGFP including an intron, and a polyadenylation signal. The CMV promoter in the transgene plasmid is used to drive expression of a template RNA that contains a rDNA homology region, a UTR, and an antisense-oriented insert. [Figure 14] Figure 1 shows ddPCR assessment of copy number changes of transgene-rDNA junctions in human cells under transfection conditions. Forward primers and probes were designed to bind to the 3'UTR of R2Tg, and reverse primers targeted human rDNA. The resulting ddPCR signal was normalized to that of the reference assay RPP30 to determine copy number. A significantly higher average copy number per genome was observed for wild type (WT) R2Tg compared to the backbone construct without R2Tg sequence. Condition 1 shows a 9:1 driver plasmid:transgene plasmid molar ratio; condition 2 shows a 4:1 molar ratio, condition 3 shows a 1:1 molar ratio, condition 4 shows a 1:4 molar ratio, and condition 5 shows a 1:9 molar ratio. [Figure 15]Figure 15A: Hybrid capture of R2Tg identified on-target integration in the human genome. The y-axis shows the read coverage aligned with the expected target integration in the R2 ribosomal site. The 5' junction between rDNA and R2Tg is shown by the vertical line on the left, and the 3' junction is shown by the vertical line on the right. Next generation sequencing identifies reads that span the expected junction. Figure 15B shows the number of reads from this experiment classified as on-target or off-target integration at the 5' and 3' ends of the integration sequence. [Figure 16] Shown is the Sanger sequencing result of 3' junction nested PCR. Nucleotides in lower case represent the designed SNP. Nucleotides in shaded capital letters represent the WT sequence. Figure 16 discloses SEQ ID NO: 1538. [Figure 17]17A-17C are schematic diagrams depicting various covalently dimerized GeneWriter configurations. The proteins depicted are as follows: Figure 17A: wild type full-length enzyme. Figure 17B: two full-length enzymes (each containing a DNA-binding domain, an RNA-binding domain, a reverse transcriptase domain, and an endonuclease domain) linked by a linker. Figure 17C: a DNA-binding domain and an RNA-binding domain linked by a linker to a full-length enzyme. Figure 17D: a DNA-binding domain and an RNA-binding domain linked by a linker to an RNA-binding domain, a reverse transcriptase domain, and an endonuclease domain. Figure 17E: a DNA-binding domain linked by a first linker to an RNA-binding domain, which is linked by a second linker to a second RNA-binding domain, a reverse transcriptase domain, and an endonuclease domain. FIG. 17F: A DNA-binding domain linked by a first linker to an RNA-binding domain, which is linked by a second linker to multiple RNA-binding domains (in this figure, the molecule contains three RNA-binding domains), which are linked by linkers to a reverse transcriptase domain and an endonuclease domain. In some embodiments, each R2 binds to a UTR in the template RNA. In some embodiments, at least one molecule contains a reverse transcriptase domain and an endonuclease domain. In some embodiments, the protein contains multiple RNA-binding domains. In some embodiments, the modular system is partitioned and is active only when it binds to DNA, where the system uses two different DNA-binding modules, for example, a first protein containing a first DNA-binding module fused to an RNA-binding module that recruits an RNA template for target-primed reverse transcription, and a second protein that binds to an integration site and contains a second DNA-binding module fused to a reverse transcription and endonuclease module. In some embodiments, the nucleic acid encoding the GeneWriter contains an intein such that the GeneWriter protein is expressed from two separate genes and fused after translation by protein splicing.In some embodiments, the GeneWriter is derived from a non-LTR protein, for example, the R2 protein. [Figure 18]18A is a schematic diagram showing the different modular elements of the GeneWriter protein. The proteins depicted are as follows: FIG. 18A: wild-type full-length enzyme. FIG. 18B: The DNA-binding domain of GeneWriter may comprise a zinc finger, Cas9, or a transcription factor, or a fragment or variant of any of them. FIG. 18C: The reverse transcriptase domain and the RNA-binding domain together may comprise a reverse transcriptase domain that is heterologous to one or more other domains of the protein (e.g., from the R2 protein) and may optionally comprise one or more additional RNA-binding domains, or a fragment or variant of any of them. FIG. 18D: The RNA-binding domain may comprise, for example, a B-box protein, an MS2 coat protein, a dCas protein, or a UTR-binding protein, or a fragment or variant of any of them. FIG. 18E: The reverse transcriptase domain may comprise, for example, a truncated reverse transcriptase domain from the R2 protein; a reverse transcriptase domain from a virus (e.g., HIV), or a reverse transcriptase domain from AMV (avian myeloblastosis virus), or a fragment or variant of any of them. FIG. 18F: The endonuclease domain may comprise, for example, a Cas9 nickase, a Cas orthologue, FokI, or a restriction enzyme, or a fragment or variant of any thereof. In some embodiments, a separate DNA binding domain can bind to a polypeptide described herein (e.g., a DNA binding domain that has a stronger affinity for a target DNA sequence than an existing or previous DNA binding domain of the polypeptide, or a DNA binding sequence that binds to a different target DNA sequence than an existing or previous DNA binding domain of the polypeptide). In some embodiments, for example, a DNA binding domain variant can be made that has increased affinity for a target DNA sequence. In some embodiments, the DNA binding domain comprises a zinc finger. In some embodiments, the DNA binding domain is linked to the polypeptide (e.g., at the N-terminus or C-terminus) via a linker, e.g., as described herein.In some embodiments, the zinc fingers are attached to the DNA-binding domain mutants (e.g., as described herein) such that the polypeptide exhibits increased binding to a target DNA sequence (e.g., as directed by the zinc fingers) without competition with rDNA. [Figure 19] 1 is a graph showing linker mutant integration into the genome of HEK293T cells, as assessed by a ddPCR assay evaluating the copy number of R2Tg integration per genome. In the v1 mutant, the insertion is located N-terminal to the α-helical region of R2Tg preceding the predicted −1 RNA binding motif; in the v2 mutant, the insertion is located C-terminal to the α-helical region of R2Tg preceding the predicted −1 RNA binding motif; in the v3 mutant, the insertion is located C-terminal to the random coil region following the predicted c-myb DNA binding motif of R2Tg. [Figure 20A] 1 is a series of graphs showing long-read sequencing confirming the fidelity of R2Tg cis integration. Unique sequence coverage is graphed through the predicted reference sequence as determined by UMI. The vertical line on the left indicates the predicted 5' junction of rDNA and R2Tg, and the vertical line on the right indicates the 3' junction. Two separate amplicons spanning from the 5' junction to the 3' junction are shown. [Figure 20B] 1 is a series of graphs showing long-read sequencing confirming the fidelity of R2Tg cis integration. Unique sequence coverage is graphed through the predicted reference sequence as determined by UMI. The vertical line on the left indicates the predicted 5' junction of rDNA and R2Tg, and the vertical line on the right indicates the 3' junction. Two separate amplicons spanning from the 5' junction to the 3' junction are shown. [Figure 21A] FIG. 1 is a series of graphs showing long-read sequencing confirming the fidelity of R2Tg cis-integration. Unique sequence deletions (>3bp) are graphed through the predicted reference sequence as determined by UMI. The vertical line on the left indicates the predicted 5' junction of rDNA and R2Tg, and the vertical line on the right indicates the 3' junction. Two separate amplicons spanning from the 5' junction to the 3' junction are shown. [Figure 21B] FIG. 1 is a series of graphs showing long-read sequencing confirming the fidelity of R2Tg cis-integration. Unique sequence deletions (>3bp) are graphed through the predicted reference sequence as determined by UMI. The vertical line on the left indicates the predicted 5' junction of rDNA and R2Tg, and the vertical line on the right indicates the 3' junction. Two separate amplicons spanning from the 5' junction to the 3' junction are shown. [Figure 22] FIG. 1 shows an exemplary plasmid map PLV033 for cis integration of R2Gfo. [Diagram 23] Graph showing cis integration of R2Gfo, R4Al, and R2Tg into HEK293T cells. Means of four replicates are shown; error bars indicate standard deviation. [Figure 24] 1 is a graph showing that R2Tg integrates in cis into human fibroblasts. Integration efficiencies of wild-type (WT) and endonuclease (EN) control R2Tg, as measured by ddPCR of the 3′ junction of R2Tg and the rDNA target, are plotted for four replicate experiments. [Diagram 25] Western blot analysis for p53, p21, actin, and vinculin. U2OS cells were tested with the indicated compounds or plasmids: GFP, R2Tg-WT (wild type), or R2Tg-EN (endonuclease domain mutant). Plasmid transfections were performed using either Lipofectamine 3000 (Lipo) or Fugene HD (Fug). Analysis was performed 24 hours after treatment or transfection. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0232] The present disclosure relates to compositions, systems and methods for targeting, editing, modifying or manipulating DNA sequences (e.g., inserting a heterologous DNA sequence of interest at a target site in a mammalian genome), e.g., at one or more locations within the DNA sequence in a cell, tissue or subject, in vivo or in vitro. The DNA sequence of interest may include, for example, a coding sequence, a regulatory sequence, a gene expression unit.
[0233] More specifically, the present disclosure provides a retrotransposon-based system for inserting a sequence of interest into a genome. The present disclosure identifies retrotransposase sequences and associated 5'UTRs and 3'UTRs from a wide variety of organisms based, in part, on bioinformatics analysis (see Table 3). Without wishing to be bound by theory, in some embodiments, retrotransposases identified in homeothermic (warm-blooded) animal species, such as birds, may have improved thermostability compared to some other enzymes that evolved at lower temperatures, and thus thermostable retrotransposases are believed to be more suitable for use in human cells. The present disclosure can also use multiple retrotransposases from different species, e.g., different animal species, and / or various species and clades of retrotransposons (e.g., as classified by reverse transcriptase phylogeny, e.g., as described in Su et al. (2019) RNA, which is incorporated by reference in its entirety), to facilitate DNA insertion into target sites in human cells (see Example 7 and Example 28).
[0234] In some embodiments, the system described herein may have several advantages over various previous systems. For example, the present disclosure describes a retrotransposase that can insert long sequences of heterologous nucleic acid (e.g., more than 3000 nucleotides) into a genome (see, e.g., FIG. 20A). In addition, the retrotransposase described herein can insert heterologous nucleic acid into endogenous sites in a genome, such as rDNA loci (e.g., Example 7). This is in contrast to the Cre / loxP system, which requires a first step of inserting an exogenous loxP site before a second step of inserting a sequence of interest into this loxP site.
[0235] Gene-writer™ Genome Editor Non-long terminal repeat (LTR) retrotransposons are a type of mobile genetic element that is widespread in eukaryotic genomes. They include two classes: apurinic / apyrimidinic endonuclease (APE) type and restriction enzyme-like endonuclease (RLE) type. APE class retrotransposons are composed of two functional domains: an endonuclease / DNA binding domain and a reverse transcriptase domain. The RLE class is composed of three functional domains: a DNA binding domain, a reverse transcriptase domain, and an endonuclease domain. The reverse transcriptase domain of non-LTR retrotransposons functions by binding to an RNA sequence template and reverse transcribing it into target DNA in the host genome. The RNA sequence template contains a 3' untranslated region that is specifically bound by the transposase and a variable 5' region that generally has an open reading frame ("ORF") that encodes the transposase protein. The RNA sequence template may also contain a 5' untranslated region that is specifically bound by the retrotransposase.
[0236] The inventors have surprisingly found that such non-LTR retrotransposon elements can be functionally modularized and / or modified to target, edit, modify or manipulate a target DNA sequence, e.g., to insert a desired (e.g., heterologous) nucleic acid sequence into a target genome, e.g., a mammalian genome, by reverse transcription. Such modularized and modified nucleic acid, polypeptide compositions and systems are described herein and are referred to as Gene Writer™ gene editors. Gene Writer™ gene editors include: (A) a polypeptide or a nucleic acid encoding a polypeptide, where the polypeptide includes either (i) a reverse transcriptase domain and (x) an endonuclease domain that includes DNA binding functionality or (y) an endonuclease domain and a separate DNA binding domain; and (B) a template RNA that includes (i) a sequence that binds to the polypeptide and (ii) a heterologous insert sequence. For example, a Gene Writer genome editor protein can include a DNA binding domain, a reverse transcriptase domain, and an endonuclease domain. In other embodiments, the Gene Writer genome editor protein may comprise a reverse transcriptase domain and an endonuclease domain. In certain embodiments, the Gene Writer™ gene editor polypeptide may be derived from a sequence of a non-LTR retrotransposon, such as an APE-type or RLE-type retrotransposon or a portion or domain thereof. In some embodiments, the RLE-type non-LTR retrotransposon is from the R2, NeSL, HERO, R4, or CRE clade. In some embodiments, the Gene Writer genome editor is derived from the R4 element X4_Line present in the human genome. In some embodiments, the APE-type non-LTR retrotransposon is from the R1, or Tx1 clade. In some embodiments, the Gene Writer genome editor is derived from the Tx1 element Mare6 present in the human genome. The RNA template element of the Gene Writer™ genome editor system is typically heterologous to the polypeptide element and provides the sequence of interest to be inserted (reverse transcribed) into the host genome.In some embodiments, the Gene Writer genome editor protein is capable of achieving target primed reverse transcription.
[0237] In some embodiments, the Gene Writer genome editor is combined with a second polypeptide. In some embodiments, the second polypeptide is obtained from an APE-type non-LTR retrotransposon. In some embodiments, the second polypeptide has a zinc knuckle-like motif. In some embodiments, the second polypeptide is a homolog of the Gag protein.
[0238] Polypeptide components of the Gene Writer genome editor system RT domain: In certain forms of the invention, the reverse transcriptase domain of the Gene Writer system is based on the reverse transcriptase domain of APE-type or RLE-type non-LTR retrotransposon. The wild-type reverse transcriptase domain of APE-type or RLE-type non-LTR retrotransposon in the Gene Writer system can be used or modified (e.g., by insertion, deletion, or substitution of one or more residues) to alter the reverse transcriptase activity on the target DNA sequence. In some embodiments, the reverse transcriptase is modified from its native sequence to obtain an altered codon usage, e.g., an improved codon usage for human cells. In some embodiments, the reverse transcriptase is a heterologous reverse transcriptase from a different retrovirus, LTR-retrotransposon, or non-LTR retrotransposon. In certain embodiments, the Gene Writer system comprises a polypeptide comprising the reverse transcriptase domain of an RLE-type non-LTR retrotransposon from the R2, NeSL, HERO, R4, or CRE clade, or the reverse transcriptase domain of an APE-type non-LTR retrotransposon from the R1, or Tx1 clade. In certain embodiments, the Gene Writer system comprises a polypeptide comprising a reverse transcriptase domain of a retrotransposon listed in Table 1, Table 2, or Table 3. In some embodiments, the amino acid sequence of the reverse transcriptase domain of the Gene Writer system is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identical to the amino acid sequence of the reverse transcriptase domain of the retrotransposon, the DNA sequence of which is listed by reference in Table 1, Table 2, or Table 3. One of skill in the art can identify the reverse transcriptase domain based on its homology to other known reverse transcriptase domains using routine tools such as Basic Local Alignment Search Tool (BLAST). In some embodiments, the reverse transcriptase domain is modified, for example, by site-directed mutagenesis. In some embodiments, the reverse transcriptase domain is engineered to bind to a heterologous template RNA.
[0239] Endonuclease domain: In certain embodiments, the Gene Writer system described herein may use or modify (e.g., by inserting, deleting, or substituting one or more residues) the endonuclease / DNA binding domain of an APE-type retrotransposon or the endonuclease domain of an RLE-type retrotransposon. In some embodiments, the endonuclease domain or the endonuclease / DNA binding domain is modified from its native sequence to obtain modified codon usage, e.g., improved codon usage for human cells. In some embodiments, the endonuclease element is a heterologous endonuclease element, e.g., Fok1 nuclease, a type II restriction-like endonuclease (RLE-type nuclease), or another RLE-type endonuclease (also known as REL). In some embodiments, the heterologous endonuclease activity has nickase activity and does not form double-stranded breaks. The amino acid sequence of the endonuclease domain of the Gene Writer system described herein is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identical to the amino acid sequence of the endonuclease domain of the retrotransposon, whose DNA sequence is listed by reference in Table 1, Table 2, or Table 3. One skilled in the art can identify endonuclease domains based on homology with other known endonuclease domains using tools such as the Basic Local Alignment Search Tool (BLAST). In certain embodiments, the heterologous endonuclease is Fok1 or a functional fragment thereof. In certain embodiments, the heterologous endonuclease is a Holliday junction resolvase or a homolog thereof, such as the Holliday junction resolvase from Sulfolobus solfataricus-Ssol Hje (Govindaraju et al., Nucleic Acids Research 44:7, 2016).In certain embodiments, the heterologous endonuclease is a large fragment endonuclease of a spliceosomal protein, e.g., Prp8 (Mahbub et al., Mobile DNA 8:16, 2017). For example, the Gene Writer polypeptides described herein may comprise a reverse transcriptase domain from an APE-type or RLE-type retrotransposon and an endonuclease domain comprising Fok1 or a functional fragment thereof. In yet another embodiment, the homologous endonuclease domain is modified, e.g., by site-directed mutagenesis, to alter the DNA binding domain endonuclease activity. In yet another embodiment, the endonuclease domain is modified to remove any potential DNA sequence specificity.
[0240] DNA binding domain: In certain aspects, the DNA binding domain of the Gene Writer polypeptide described herein is selected, designed, or constructed to bind to a desired host DNA target sequence. In certain embodiments, the DNA binding domain of the engineered RLE is a heterologous DNA binding protein or domain compared to the natural retrotransposon sequence. In some embodiments, the heterologous DNA binding element is a zinc finger element or a TAL effector element, such as a zinc finger or TAL polypeptide or a functional fragment thereof. In some embodiments, the heterologous DNA binding element is a sequence-guided DNA binding element, such as Cas9, Cpf1, or other CRISPR-associated proteins that have been modified to have no endonuclease activity. In some embodiments, the heterologous DNA binding element retains endonuclease activity. In some embodiments, the heterologous DNA binding element is used in place of the endonuclease element of the polypeptide. In specific embodiments, the heterologous DNA binding domain may be any one or more of Cas9, a TAL domain, a ZF domain, a Myb domain, a combination thereof, or a population thereof. In certain embodiments, the heterologous DNA binding domain is a DNA binding domain of a retrotransposon listed in Table 1, Table 2, or Table 3. One skilled in the art can use tools such as the Basic Local Alignment Search Tool (BLAST) to identify DNA binding domains based on homology with other known DNA binding domains. In other embodiments, the DNA binding domain is modified, for example, by site-directed mutagenesis, increasing or decreasing DNA binding domain elements (e.g., the number and / or specificity of zinc fingers) to alter DNA binding specificity and affinity. In some embodiments, the DNA binding domain is modified from its native sequence to obtain altered codon usage, for example, improved codon usage for human cells.
[0241] In certain aspects of the invention, the host DNA binding site engineered by the Gene Writer system may be within a gene, within an intron, within an exon, within an ORF, outside the coding region of any gene, within a regulatory region of a gene, or outside a regulatory region of a gene. In other aspects, the engineered RLE may bind to one or more DNA sequences.
[0242] In certain embodiments, the Gene Writer™ gene editor system RNA further comprises a subcellular localization sequence, e.g., a nuclear localization sequence. The nuclear localization sequence can be an RNA sequence that promotes the import of the RNA into the nucleus. In certain embodiments, the nuclear localization signal is located on the template RNA. In certain embodiments, the retrotransposase polypeptide is encoded on a first RNA, the template RNA is a second separate RNA, and the nuclear localization signal is located on the template RNA but not on the RNA that encodes the retrotransposase polypeptide. Without wishing to be bound by theory, in some embodiments, the RNA encoding the retrotransposase is primarily targeted to the cytoplasm to promote its translation, whereas the template RNA is primarily targeted to the nucleus to promote its retrotransposition into the genome. In some embodiments, the nuclear localization signal is at the 3' end, 5' end, or within an internal region of the template RNA. In some embodiments, the nuclear localization signal is 3' to the heterologous sequence (e.g., adjacent to the 3' side of the heterologous sequence) or 5' to the heterologous sequence (e.g., adjacent to the 5' side of the heterologous sequence). In some embodiments, the nuclear localization signal is outside the 5'UTR or outside the 3'UTR of the template RNA. In some embodiments, the nuclear localization signal is located between the 5'UTR and the 3'UTR, where the nuclear localization signal is optionally transcribed in the transgene (e.g., the nuclear localization signal is in an antisense orientation or downstream of a transcription termination signal or a polyadenylation signal). In some embodiments, the nuclear localization signal is located inside an intron. In some embodiments, multiple identical or different nuclear localization signals are present within the RNA, e.g., within the template RNA. In some embodiments, the nuclear localization signal is less than 5, 10, 25, 50, 75, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900 or 1000 bp in length.Various RNA nuclear localization sequences can be used.For example, Lubelsky and Ulitsky, Nature 555(107-111), 2018, describes RNA sequences that drive RNA localization to the nucleus.In some embodiments, the nuclear localization signal is a SINE-derived nuclear RNA localization (SIRLOIN) signal. In some embodiments, the nuclear localization signal binds to a nuclear enriched protein. In some embodiments, the nuclear localization signal binds to a HNRNPK protein. In some embodiments, the nuclear localization signal is pyrimidine-rich, e.g., a C / T-rich, C / U-rich, C-rich, T-rich, or C-rich region. In some embodiments, the nuclear localization signal is derived from a long non-coding RNA. In some embodiments, the nuclear localization signal is derived from the MALAT1 long non-coding RNA or is the 600 nucleotide M region of MALAT1 (described in Miyagawa et al., RNA 18, (738-751), 2012). In some embodiments, the nuclear localization signal is derived from the BORG long non-coding RNA or is an AGCCC motif (described in Zhang et al., Molecular and Cellular Biology 34, 2318-2329 (2014)). In some embodiments, the nuclear localization sequence is described in Shukla et al., The EMBO Journal e98452 (2018). In some embodiments, the nuclear localization signal is derived from a non-LTR retrotransposon, an LTR retrotransposon, a retrovirus, or an endogenous retrovirus.
[0243] In certain embodiments, the Gene Writer™ gene editor system polypeptide further comprises a subcellular localization sequence, e.g., a nuclear localization sequence and / or a nucleolar localization sequence. The nuclear localization sequence and / or the nucleolar localization sequence may be an amino acid sequence that facilitates import of a protein into the nucleus and / or the nucleolus, whereby the sequence can facilitate integration of a heterologous sequence into the genome. In certain embodiments, the Gene Writer gene editor system polypeptide (e.g., a retrotransposase, e.g., a polypeptide according to any of Tables 1, 2, or 3 described herein) further comprises a nucleolar localization sequence. In certain embodiments, the retrotransposase is encoded on a first RNA, the template RNA is a second, separate RNA, and the nucleolar localization signal is encoded on the RNA encoding the retrotransposase polypeptide, but not in the template RNA. In some embodiments, the nucleolar localization signal is located at the N-terminus, C-terminus, or within an internal region of the polypeptide. In some embodiments, multiple identical or different nucleolar localization signals are used. In some embodiments, the nuclear localization signal is less than 5, 10, 25, 50, 75, or 100 amino acids in length. A variety of polypeptide nucleolar localization signals are used. For example, Yang et al., Journal of Biomedical Science 22, 33 (2015) describes a nuclear localization signal that also functions as a nucleolar localization signal. In some embodiments, the nucleolar localization signal may be a nuclear localization signal. In some embodiments, the nucleolar localization signal may overlap with a nuclear localization signal. In some embodiments, the nucleolar localization signal may include a stretch of base residues. In some embodiments, the nucleolar localization signal may be rich in arginine and lysine residues. In some embodiments, the nucleolar localization signal may be derived from a protein enriched in the nucleolus. In some embodiments, the nucleolar localization signal may be derived from a protein enriched at a ribosomal RNA locus. In some embodiments, the nucleolar localization signal may be derived from a protein that binds to rRNA.In some embodiments, the nucleolar localization signal may be derived from MSP58. In some embodiments, the nucleolar localization signal may be a monopartite motif. In some embodiments, the nucleolar localization signal may be a bipartite motif. In some embodiments, the nucleolar localization signal may be composed of multiple monopartite and bipartite motifs. In some embodiments, the nucleolar localization signal may be composed of a mixture of monopartite and bipartite motifs. In some embodiments, the nucleolar localization signal may be a double bipartite motif. In some embodiments, the nucleolar localization signal may be KRASSQALGTIPKRRSSSRFIKRKK (SEQ ID NO: 1530). In some embodiments, the nucleolar localization signal may be derived from nuclear factor-κB-inducing kinase. In some embodiments, the nucleolar localization signal may be the RKKRKKK motif (SEQ ID NO: 1531) (Birbach et al., Journal of Cell Science, 117(3615-3624), 2004).
[0244] In some embodiments, the nucleic acid described herein (e.g., an RNA encoding a GeneWriter polypeptide, or a DNA encoding that RNA) comprises a microRNA binding site. In some embodiments, the microRNA binding site is used to increase the target cell specificity of the GeneWriter system. For example, a microRNA binding site can be selected based on the criterion of being recognized by a miRNA that is present in a non-target cell type but not present (or present at a lower level compared to non-target cells) in a target cell type. Thus, when an RNA encoding a GeneWriter polypeptide is present in a non-target cell, it will be bound by the miRNA, and when an RNA encoding a GeneWriter polypeptide is present in a target cell, it will not be bound by the miRNA (or will be bound at a lower level compared to non-target cells). Without wishing to be bound by theory, binding of the miRNA to an RNA encoding a GeneWriter polypeptide may reduce the production of the GeneWriter polypeptide, for example, by degrading the mRNA encoding the polypeptide or by interfering with translation. Thus, the heterologous sequence of interest will be inserted into the genome of the target cell more efficiently than the genome of the non-target cell. Also, as described herein in the section entitled "Template RNA Component of the Gene Writer™ Gene Editor System," a system having a microRNA binding site in the RNA encoding the GeneWriter polypeptide (or encoded by DNA encoding the RNA) may be used in combination with a template RNA that is regulated by a second microRNA binding site.
[0245] [Table 1]
[0246] [Table 2]
[0247]
Table 3
[0248]
Table 4
[0249]
Table 5
[0250]
Table 6
[0251]
Table 7
[0252]
Table 8
[0253]
Table 9
[0254]
Table 10
[0255]
Table 11
[0256]
Table 12
[0257]
Table 13
[0258]
Table 14
[0259]
Table 15
[0260]
Table 16
[0261]
Table 17
[0262]
Table 18
[0263]
Table 19
[0264]
Table 20
[0265]
Table 21
[0266]
Table 22
[0267]
Table 23
[0268]
Table 24
[0269]
Table 25
[0270]
Table 26
[0271]
Table 27
[0272]
Table 28
[0273]
Table 29
[0274]
Table 30
[0275]
Table 31
[0276]
Table 32
[0277]
Table 33
[0278]
Table 34
[0279]
Table 35
[0280]
Table 36
[0281]
Table 37
[0282]
Table 38
[0283]
Table 39
[0284]
Table 40
[0285]
Table 41
[0286]
Table 42
[0287]
Table 43
[0288]
Table 44
[0289]
Table 45
[0290]
Table 46
[0291]
Table 47
[0292]
Table 48
[0293]
Table 49
[0294]
Table 50
[0295]
Table 51
[0296]
Table 52
[0297]
Table 53
[0298]
Table 54
[0299]
Table 55
[0300]
Table 56
[0301]
Table 57
[0302]
Table 58
[0303]
Table 59
[0304]
Table 60
[0305]
Table 61
[0306]
Table 62
[0307]
Table 63
[0308]
Table 64
[0309]
Table 65
[0310]
Table 66
[0311]
Table 67
[0312]
Table 68
[0313]
Table 69
[0314]
Table 70
[0315]
Table 71
[0316]
Table 72
[0317]
Table 73
[0318]
Table 74
[0319]
Table 75
[0320] [Table 76]
[0321] [Table 77]
[0322] [Table 78]
[0323] [Table 79]
[0324] A person skilled in the art can determine the nucleic acid and corresponding polypeptide sequences of each retrotransposon and their domains based on the accession numbers listed in Tables 1-3 by using a conventional sequence analysis tool, such as the Basic Local Alignment Search Tool (BLAST) for conserved domain analysis or CD-Search. Other sequence analysis tools are well known and can be found, for example, at https: / / molbiol-tools.ca, e.g., https: / / molbiol-tools.ca / Motifs.htm. SEQ ID NOs: 1-112 align with the respective rows of Table 1, and SEQ ID NOs: 113-1015 align with the first 903 rows of Table 2.
[0325] Tables 1-3 described herein provide exemplary transposon sequences, sequences of 5' and 3' untranslated regions that allow retrotransposase to bind to template RNA, and full transposon nucleic acid sequences, including amino acid sequences of retrotransposases. In some embodiments, a 5'UTR of any of Tables 1-3 allows retrotransposase to bind to template RNA. In some embodiments, a 3'UTR of any of Tables 1-3 allows retrotransposase to bind to template RNA. Thus, in some embodiments, a polypeptide for use in any of the systems described herein can be a polypeptide of any of Tables 1-3 herein, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the system further comprises one or both of the 5' or 3' untranslated regions of any of Tables 1-3 herein (or sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto) derived from the same transposon as the polypeptide set forth in the preceding sentence, e.g., as displayed in the same row of the same table. In some embodiments, the system comprises one or both of the 5' or 3' untranslated regions of any of Tables 1-3 herein, e.g., a segment of the entire transposon sequence encoding an RNA capable of binding a retrotransposase, and / or a subsequence set forth in the column entitled Predicted 5' UTR or Predicted 3' UTR.
[0326] In some embodiments, the polypeptides used in any of the systems described herein may be molecular reconstructions or ancestral sequence reconstructions based on aligned polypeptide sequences of multiple retrotransposons. In some embodiments, the 5' or 3' untranslated regions used in any of the systems described herein may be molecular reconstructions based on aligned 5' or 3' untranslated regions of multiple retrotransposons. Those skilled in the art can align polypeptide or nucleic acid sequences based on the accession numbers described herein, for example, by using routine sequence analysis tools such as Basic Local Alignment Search Tool (BLAST) or CD-Search for conserved domain analysis. Molecular reconstructions can be generated based on sequence consensus, for example, using the approaches described in Ivics et al., Cell 1997, 501-510; Wagstaff et al., Molecular Biology and Evolution 2013, 88-99. In some embodiments, the retrotransposon from which the 5' or 3' untranslated region or polypeptide is derived is a young or newly active mobile element, as assessed by biogenesis methods such as those described in Boissinot et al., Molecular Biology and Evolution 2000, 915-928.
[0327] Table 3 (below) shows exemplary Gene Writer proteins and related sequences from various retrotransposases identified using data mining. Column 1 indicates the family to which the retrotransposon belongs. Column 2 lists the component name. Column 3 indicates the accession number, if present. Column 4 lists the organism in which the retrotransposase is found. Column 5 lists the DNA sequence of the retrotransposon. Column 6 lists the predicted 5' untranslated region and column 7 lists the predicted 3' untranslated region; both are segments of the sequence in column 5 that are predicted to be capable of binding a template RNA to the retrotransposase in column 8. (It is understood that columns 5-7 are DNA sequences, and that RNA sequences according to any of columns 5-7 will typically be uracil rather than thymidine). Column 8 lists the predicted retrotransposase sequence encoded by the retrotransposon in column 5.
[0328] [Table 80]
[0329] [Table 81]
[0330] [Table 82]
[0331] [Table 83]
[0332] [Table 84]
[0333] [Table 85]
[0334]
Table 86
[0335]
Table 87
[0336]
Table 88
[0337]
Table 89
[0338]
Table 90
[0339]
Table 91
[0340]
Table 92
[0341]
Table 93
[0342]
Table 94
[0343]
Table 95
[0344]
Table 96
[0345]
Table 97
[0346]
Table 98
[0347]
Table 99
[0348]
Table 100
[0349]
Table 101
[0350]
Table 102
[0351]
Table 103
[0352]
Table 104
[0353]
Table 105
[0354]
Table 106
[0355]
Table 107
[0356]
Table 108
[0357]
Table 109
[0358]
Table 110
[0359]
Table 111
[0360]
Table 112
[0361]
Table 113
[0362]
Table 114
[0363]
Table 115
[0364]
Table 116
[0365]
Table 117
[0366]
Table 118
[0367]
Table 119
[0368]
Table 120
[0369]
Table 121
[0370]
Table 122
[0371]
Table 123
[0372]
Table 124
[0373]
Table 125
[0374]
Table 126
[0375]
Table 127
[0376]
Table 128
[0377]
Table 129
[0378]
Table 130
[0379]
Table 131
[0380]
Table 132
[0381]
Table 133
[0382]
Table 134
[0383]
Table 135
[0384]
Table 136
[0385]
Table 137
[0386]
Table 138
[0387]
Table 139
[0388]
Table 140
[0389]
Table 141
[0390]
Table 142
[0391]
Table 143
[0392]
Table 144
[0393]
Table 145
[0394]
Table 146
[0395]
Table 147
[0396]
Table 148
[0397]
Table 149
[0398]
Table 150
[0399]
Table 151
[0400]
Table 152
[0401]
Table 153
[0402]
Table 154
[0403]
Table 155
[0404]
Table 156
[0405]
Table 157
[0406]
Table 158
[0407]
Table 159
[0408]
Table 160
[0409]
Table 161
[0410]
Table 162
[0411]
Table 163
[0412]
Table 164
[0413]
Table 165
[0414]
Table 166
[0415]
Table 167
[0416]
Table 168
[0417]
Table 169
[0418]
Table 170
[0419]
Table 171
[0420]
Table 172
[0421]
Table 173
[0422]
Table 174
[0423]
Table 175
[0424]
Table 176
[0425]
Table 177
[0426]
Table 178
[0427]
Table 179
[0428]
Table 180
[0429]
Table 181
[0430]
Table 182
[0431]
Table 183
[0432]
Table 184
[0433]
Table 185
[0434]
Table 186
[0435]
Table 187
[0436]
Table 188
[0437]
Table 189
[0438]
Table 190
[0439]
Table 191
[0440]
Table 192
[0441]
Table 193
[0442]
Table 194
[0443]
Table 195
[0444]
Table 196
[0445]
Table 197
[0446]
Table 198
[0447]
Table 199
[0448]
Table 200
[0449]
Table 201
[0450]
Table 202
[0451]
Table 203
[0452]
Table 204
[0453]
Table 205
[0454]
Table 206
[0455]
Table 207
[0456]
Table 208
[0457]
Table 209
[0458]
Table 210
[0459]
Table 211
[0460]
Table 212
[0461]
Table 213
[0462]
Table 214
[0463]
Table 215
[0464]
Table 216
[0465]
Table 217
[0466]
Table 218
[0467]
Table 219
[0468]
Table 220
[0469]
Table 221
[0470]
Table 222
[0471]
Table 223
[0472]
Table 224
[0473]
Table 225
[0474]
Table 226
[0475]
Table 227
[0476]
Table 228
[0477]
Table 229
[0478]
Table 230
[0479]
Table 231
[0480]
Table 232
[0481]
Table 233
[0482]
Table 234
[0483]
Table 235
[0484]
Table 236
[0485]
Table 237
[0486]
Table 238
[0487]
Table 239
[0488]
Table 240
[0489]
Table 241
[0490]
Table 242
[0491]
Table 243
[0492]
Table 244
[0493]
Table 245
[0494]
Table 246
[0495]
Table 247
[0496]
Table 248
[0497]
Table 249
[0498]
Table 250
[0499]
Table 251
[0500]
Table 252
[0501]
Table 253
[0502]
Table 254
[0503]
Table 255
[0504]
Table 256
[0505]
Table 257
[0506]
Table 258
[0507]
Table 259
[0508]
Table 260
[0509]
Table 261
[0510]
Table 262
[0511]
Table 263
[0512]
Table 264
[0513]
Table 265
[0514]
Table 266
[0515]
Table 267
[0516]
Table 268
[0517]
Table 269
[0518]
Table 270
[0519]
Table 271
[0520]
Table 272
[0521]
Table 273
[0522]
Table 274
[0523]
Table 275
[0524]
Table 276
[0525]
Table 277
[0526]
Table 278
[0527]
Table 279
[0528]
Table 280
[0529]
Table 281
[0530]
Table 282
[0531]
Table 283
[0532]
Table 284
[0533]
Table 285
[0534]
Table 286
[0535]
Table 287
[0536]
Table 288
[0537]
Table 289
[0538]
Table 290
[0539]
Table 291
[0540]
Table 292
[0541]
Table 293
[0542]
Table 294
[0543]
Table 295
[0544]
Table 296
[0545]
Table 297
[0546]
Table 298
[0547]
Table 299
[0548]
Table 300
[0549]
Table 301
[0550]
Table 302
[0551]
Table 303
[0552]
Table 304
[0553]
Table 305
[0554]
Table 306
[0555]
Table 307
[0556]
Table 308
[0557]
Table 309
[0558]
Table 310
[0559]
Table 311
[0560]
Table 312
[0561]
Table 313
[0562]
Table 314
[0563]
Table 315
[0564]
Table 316
[0565]
Table 317
[0566]
Table 318
[0567]
Table 319
[0568]
Table 320
[0569]
Table 321
[0570]
Table 322
[0571]
Table 323
[0572]
Table 324
[0573]
Table 325
[0574]
Table 326
[0575]
Table 327
[0576]
Table 328
[0577]
Table 329
[0578]
Table 330
[0579]
Table 331
[0580]
Table 332
[0581]
Table 333
[0582]
Table 334
[0583]
Table 335
[0584]
Table 336
[0585]
Table 337
[0586]
Table 338
[0587]
Table 339
[0588]
Table 340
[0589]
Table 341
[0590]
Table 342
[0591]
Table 343
[0592]
Table 344
[0593]
Table 345
[0594]
Table 346
[0595]
Table 347
[0596]
Table 348
[0597]
Table 349
[0598]
Table 350
[0599]
Table 351
[0600]
Table 352
[0601]
Table 353
[0602]
Table 354
[0603]
Table 355
[0604]
Table 356
[0605]
Table 357
[0606]
Table 358
[0607]
Table 359
[0608] Table 360
[0609]
Table 361
[0610]
Table 362
[0611]
Table 363
[0612]
Table 364
[0613]
Table 365
[0614]
Table 366
[0615]
Table 367
[0616]
Table 368
[0617]
Table 369
[0618]
Table 370
[0619]
Table 371
[0620]
Table 372
[0621]
Table 373
[0622]
Table 374
[0623]
Table 375
[0624]
Table 376
[0625]
Table 377
[0626]
Table 378
Table 379
[0627]
Table 380
[0628]
Table 381
[0629]
Table 382
[0630]
Table 383
[0631]
Table 384
[0632]
Table 385
[0633]
Table 386
[0634]
Table 387
[0635]
Table 388
[0636]
Table 389
[0637]
Table 390
[0638]
Table 391
[0639]
Table 392
[0640]
Table 393
[0641]
Table 394
[0642]
Table 395
[0643]
Table 396
[0644]
Table 397
[0645]
Table 398
[0646]
Table 399
[0647]
Table 400
[0648]
Table 401
[0649]
Table 402
[0650]
Table 403
[0651]
Table 404
[0652]
Table 405
[0653]
Table 406
[0654]
Table 407
[0655]
Table 408
[0656]
Table 409
[0657]
Table 410
[0658]
Table 411
[0659]
Table 412
[0660]
Table 413
[0661]
Table 414
[0662]
Table 415
[0663]
Table 416
[0664]
Table 417
[0665]
Table 418
[0666]
Table 419
[0667]
Table 420
[0668]
Table 421
[0669]
Table 422
[0670]
Table 423
[0671]
Table 424
[0672]
Table 425
[0673]
Table 426
[0674]
Table 427
[0675]
Table 428
[0676]
Table 429
[0677] Table 430
[0678]
Table 431
[0679]
Table 432
[0680]
Table 433
[0681]
Table 434
[0682]
Table 435
[0683]
Table 436
[0684]
Table 437
[0685]
Table 438
[0686]
Table 439
[0687]
Table 440
[0688]
Table 441
[0689]
Table 442
[0690]
Table 443
[0691]
Table 444
[0692]
Table 445
[0693]
Table 446
[0694]
Table 447
[0695]
Table 448
[0696]
Table 449
[0697] Table 450
[0698]
Table 451
[0699]
Table 452
[0700]
Table 453
[0701]
Table 454
[0702]
Table 455
[0703]
Table 456
[0704]
Table 457
[0705]
Table 458
[0706]
Table 459
[0707]
Table 460
[0708]
Table 461
[0709]
Table 462
[0710]
Table 463
[0711]
Table 464
[0712]
Table 465
[0713]
Table 466
[0714] [Table 467]
[0715] [Table 468]
[0716] [Table 469]
[0717] Gene Writer, e.g., heat-stable gene writer Without wishing to be bound by theory, in some embodiments, retrotransposases evolved in low temperature environments may not function well at human body temperature. The present application provides several thermostable Gene Writers, including proteins derived from bird retrotransposases. Exemplary tri-transposase sequences shown in Table 3 include those of Taeniopygia guttata (zebra finch; transposon name R2-1_TG), Geospiza fortis (medium ground finch; transposon name R2-1_Gfo), Zonotrichia albicollis (zonotrichia albicollis; transposon name R2-1_ZA), and Tinamus guttatus (tinamus transposon name R2-1_TGut).
[0718] Thermostability can be measured, for example, by assaying the ability of a Gene Writer to polymerize DNA in vitro at high (e.g., 37°C) and low (e.g., 25°C) temperatures. Suitable conditions for assaying DNA polymerization activity (e.g., processivity) in vitro are described, for example, in Bibillo and Eickbush, "High Processivity of the Reverse Transcriptase from a Non-long Terminal Repeat Retrotransposon" (2002) JBC 277, 34836-34845. In some embodiments, a thermostable Gene Writer polypeptide has an activity, e.g., DNA polymerization activity, at 37°C that is less than 70%, 75%, 80%, 85%, 90%, or 95% of its activity at 25°C under otherwise similar conditions.
[0719] In some embodiments, a GeneWriter polypeptide (e.g., a sequence of Tables 1, 2, or 3, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto) is stable in a subject selected from a mammal (e.g., a human) or a bird. In some embodiments, a GeneWriter polypeptide described herein is functional at 37° C. In some embodiments, a GeneWriter polypeptide described herein has higher activity at 37° C. than at lower temperatures, e.g., 30° C., 25° C., or 20° C. In some embodiments, a GeneWriter polypeptide described herein has higher activity in human cells than in zebrafish cells.
[0720] In some embodiments, the GeneWriter Polypeptides are active in human cells cultured at 37° C. using the assays of Example 6 or Example 7 described herein.
[0721] In some embodiments, the assay comprises the following steps: (1) transfecting HEK293T cells at 10,000 cells / well into one or more wells with a diameter of 6.4 mm; (2) incubating the cells at 37° C. for 24 hr; (3) transfecting the cells with 0.5 μl of FuGENE® HD Transfection Reagent and 80 ng of DNA (the DNA contains, in order: (a) a CMV promoter; (b) a 100 bp sequence homologous to the 100 bp upstream of the target site; (c) a sequence encoding a 5′ untranslated region that binds to the GeneWriter protein; (d) a sequence encoding the GeneWriter protein; and (e) a sequence encoding the GeneWriter protein that binds to the CMV promoter. (f) a 100 bp sequence homologous to 100 bp downstream of the target site, and (g) a BGH polyadenylation sequence), and 10 μl of Opti-MEM and incubating at room temperature for 15 minutes; (4) adding the transfection mixture to the cells; (5) incubating the cells for 3 days; and (6) assaying for integration of the exogenous sequence into the target locus (e.g., rDNA) in the cell genome, e.g., one or more of the preceding steps are performed as described in Example 6 herein.
[0722] In some embodiments, the GeneWriter polypeptide achieves integration of a heterologous sequence of interest (e.g., a GFP gene) into a target locus (e.g., rDNA) at an average copy number of at least 0.01, 0.025, 0.05, 0.075, 0.1, 0.15, 0.2, 0.25, 0.3, 0.4, 0.5, 0.75, 1, 1.25, 1.5, 1.75, 2, 2.5, 3, 4, or 5 copies per genome. In some embodiments, a cell described herein (e.g., a cell comprising a heterologous sequence at a targeted insertion site) comprises a heterologous sequence of interest at an average copy number of at least 0.01, 0.025, 0.05, 0.075, 0.1, 0.15, 0.2, 0.25, 0.3, 0.4, 0.5, 0.75, 1, 1.25, 1.5, 1.75, 2, 2.5, 3, 4, or 5 copies per genome.
[0723] In some embodiments, the GeneWriter results in incorporation of the sequence into the target RNA with relatively few cleavage events at the ends. For example, in some embodiments, the Gene Writer protein (e.g., SEQ ID NO: 1016) achieves about 25-100%, 50-100%, 60-100%, 70-100%, 75-95%, 80-90%, or 86.17% incorporation into the target site that is non-cleavable, as measured by the assays described herein, e.g., the assays of Example 6 and FIG. 8. In some embodiments, the Gene Writer protein (e.g., SEQ ID NO: 1016) achieves at least about 30%, 40%, 50%, 60%, 70%, 80%, or 90% incorporation into the target site that is non-cleavable, as measured by the assays described herein. In some embodiments, integrants are classified as truncated and untruncated using an assay that involves amplification with a forward primer located 565 bp from the end of the element (e.g., a wild-type transposon sequence in Taeniopygia guttata) and a reverse primer located in the genomic DNA, e.g., rDNA, of the target insertion site. In some embodiments, the number of full-length integrants at the target insertion site is greater than the number of truncated integrants by 300-565 nucleotides within the target insertion site, e.g., the number of full-length integrants is at least 1.1×, 1.2×, 1.5×, 2×, 3×, 4×, 5×, 6×, 7×, 8×, 9×, or 10× the number of truncated integrants, or the number of full-length integrants is at least 1.1×-10×, 2×-10×, 3×-10×, or 5×-10× the number of truncated integrants.
[0724] In some embodiments, the systems or methods described herein achieve insertion of a heterologous sequence of interest only at a target site in the genome of a target cell. Insertion can be measured using a threshold of greater than 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, e.g., as described in Example 8. In some embodiments, the systems or methods described herein achieve insertion of a heterologous sequence of interest where 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 10%, 20%, 30%, 40%, or 50% of the insertion is outside the target site, e.g., using an assay described herein, e.g., the assay of Example 8.
[0725] In some embodiments, the systems or methods described herein achieve "scarless" insertion of the heterologous sequence of interest, and in some embodiments, the target site may exhibit deletion or duplication of endogenous DNA as a result of the insertion of the heterologous sequence. Different retrotransposon mechanisms of action result in different patterns of duplications or deletions in the host genome, which occur during retrotransposition at the target site. In some embodiments, the system achieves scarless insertion without incurring duplications or deletions in the surrounding genomic DNA. In some embodiments, the system causes deletion of less than 1, 2, 3, 4, 5, 10, 50, or 100 bp of genomic DNA upstream of the insertion. In some embodiments, the system causes deletion of less than 1, 2, 3, 4, 5, 10, 50, or 100 bp of genomic DNA downstream of the insertion. In some embodiments, the system causes duplication of less than 1, 2, 3, 4, 5, 10, 50, or 100 bp of genomic DNA upstream of the insertion. In some embodiments, the system results in a duplication of less than 1, 2, 3, 4, 5, 10, 50, or 100 bp of genomic DNA downstream of the insertion.
[0726] In some embodiments, the GeneWriter, or its DNA binding domain, described herein specifically binds to its target site as measured using the assay of Example 21. In some embodiments, the GeneWriter, or its DNA binding domain, binds to its target site more strongly than any other binding site in the human genome. For example, in some embodiments, the target site exhibits more than 50%, 60%, 70%, 80%, 90%, or 95% of the binding events of the GeneWriter, or its DNA binding domain, to human genomic DNA in the assay of Example 21.
[0727] Genetically engineered, e.g., dimerized GeneWriter Some non-LT retrotransposons use two subunits to complete retrotransposition (Christensen et al PNAS 2006). In some embodiments, the retrotransposases described herein comprise two linked subunits as a single polypeptide. For example, two wild-type retrotransposases could be linked with a linker to form a covalently "dimerized" protein (see FIG. 17). In some embodiments, the nucleic acid encoding the retrotransposase encodes two transposase subunits to be expressed as a single polypeptide. In some embodiments, the subunits are linked by a peptide linker, as described herein in the section entitled "Linkers" and, for example, in Chen et al Adv Drug Deliv Rev 2013. In some embodiments, the two subunits in the polypeptide are linked by a rigid linker. In some embodiments, the rigid linker is a linker that is linked by the motif (EAAAK) n (SEQ ID NO: 1534). In other embodiments, the two subunits of the polypeptide are linked by a flexible linker. In some embodiments, the flexible linker is comprised of the motif (Gly) n In some embodiments, the flexible linker is composed of the motif (GGGGS) n(SEQ ID NO: 1535). In some embodiments, the rigid or flexible linker is 1, 2, 3, 4, 5, 10, 15 or more amino acids in length to allow retrotransposition. In some embodiments, the linker is composed of a combination of rigid or flexible linker motifs.
[0728] Depending on the mechanism of action, not all functions are required from both retrotransposase subunits. In some embodiments, the fusion protein may be composed of a fully functional subunit and a second subunit lacking one or more functional domains. In some embodiments, one subunit may lack reverse transcriptase functionality. In some embodiments, one subunit may lack a reverse transcriptase domain. In some embodiments, one subunit may only have endonuclease activity. In some embodiments, one subunit may only have an endonuclease domain. In some embodiments, two subunits comprising a single polypeptide may provide complementary functionality.
[0729] In some embodiments, one subunit may lack endonuclease functionality. In some embodiments, one subunit may lack endonuclease domain. In some embodiments, one subunit may only have reverse transcriptase activity. In some embodiments, one subunit may only have reverse transcriptase domain. In some embodiments, one subunit may only have DNA-dependent DNA synthesis functionality.
[0730] Linker In some embodiments, the domains of the compositions and systems described herein (e.g., the endonuclease and reverse transcriptase domains of a polypeptide or the DNA binding and reverse transcriptase domains of a polypeptide) may be linked by a linker. The compositions described herein that include a linker element have the general form S1-L-S2, where S1 and S2 may be the same or different and represent two moieties (e.g., a polypeptide or a nucleic acid domain, respectively) linked to each other by a linker. In some embodiments, the linker may link two polypeptides. In some embodiments, the linker may link two nucleic acid molecules. In some embodiments, the linker may link a polypeptide and a nucleic acid molecule. The linker may be a chemical bond, e.g., one or more covalent or non-covalent bonds. The linker may be flexible, rigid, and / or cleavable. In some embodiments, the linker is a peptide linker. Generally, the peptide linker is at least 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acids in length, for example, 2 to 50 amino acids in length, 2 to 30 amino acids in length.
[0731] The most commonly used flexible linkers have sequences consisting primarily of stretches of Gly and Ser residues ("GS" linkers). Flexible linkers are believed to be useful for linking domains that require some degree of movement or interaction and may contain small non-polar (e.g., Gly) or polar (e.g., Ser or Thr) amino acids. The incorporation of Ser or Thr may also maintain the stability of the linker in aqueous solution by forming hydrogen bonds with water molecules, thereby reducing unfavorable interactions of the linker with other moieties. Examples of such linkers include those with the structure [GGS] ≧1 or [GGGS] ≧1(SEQ ID NO: 1536). Rigid linkers are useful for maintaining a constant distance between domains while maintaining their independent functions. Rigid linkers can also be useful when spatial separation of domains is essential to preserve the stability or biological activity of one or more components in the drug. Rigid linkers can have an alpha-helical structure or a Pro-rich sequence, (XP)n, where X refers to any amino acid, preferably Ala, Lys, or Glu. Cleavable linkers can release free functional domains in vivo. In some embodiments, the linker can be cleaved under certain conditions, for example, in the presence of a reducing agent or a protease. In vivo cleavable linkers may take advantage of the reversibility of disulfide bonds. An example is a thrombin-sensitive sequence (e.g., PRS) between two Cys residues. In vitro thrombin treatment of CPRSC (SEQ ID NO: 1537) results in cleavage of the thrombin-sensitive sequence, while the reversible disulfide bond remains intact. Such linkers are well known and are described, for example, in Chen et al. 2013. Fusion Protein Linkers: Property, Design and Functionality. Adv Drug Deliv Rev. 65(10):1357-1369. In vivo cleavage of the linker in the compositions described herein can also be performed by proteases expressed in vivo in specific cells or tissues, under pathological conditions (e.g., cancer or inflammation), or confined within specific cellular compartments. The specificity of various proteases allows for slow cleavage of the linker within the confined compartment.
[0732] In some embodiments, the amino acid linker is an endogenous amino acid (or homologous thereto) that occurs between such domains in a naturally occurring polypeptide. In some embodiments, the endogenous amino acids that occur between such domains are substituted but the length is unchanged from the natural length. In some embodiments, additional amino acid residues are added to the amino acid residues that naturally occur between the domains.
[0733] In some embodiments, amino acid linkers are computationally designed or screened to maximize protein function (Anad et al., FEBS Letters, 587:19, 2013).
[0734] Template RNA Component of the Gene Writer™ Gene Editor System The Gene Writer system described herein can transcribe an RNA sequence template into a host target DNA site by target-primed reverse transcription. By editing DNA sequences via direct reverse transcription of an RNA sequence template into a host genome, the Gene Writer system can insert a sequence of interest into a target genome without the need to introduce an exogenous DNA sequence into a host cell (unlike, for example, the CRISPR system) or even eliminate the need for an exogenous DNA insertion step. Thus, the Gene Writer system provides a platform for the use of customized RNA sequence templates containing sequences of interest, for example sequences containing heterologous gene coding and / or functional information.
[0735] In some embodiments, the template RNA encodes a Gene Writer protein in cis with a heterologous sequence of interest. Various cis constructs are described, for example, in Kuroki-Kami et al (2019) Mobile DNA 10:23, incorporated herein by reference in its entirety, and can be used in combination with any of the embodiments described herein. For example, in some embodiments, the template RNA includes a sequence encoding a heterologous sequence of interest, a Gene Writer protein (e.g., a protein comprising, for example, (i) a reverse transcriptase domain and (ii) an endonuclease domain, as described herein), a 5' untranslated region, and a 3' untranslated region. These components can be contained in various orders. In some embodiments, the Gene Writer protein and the heterologous sequence of interest are encoded in different orientations (sense vs. antisense), for example, using the arrangement shown in FIG. 3A of Kuroki-Kami et al. (ibid.). In some embodiments, the Gene Writer protein and the heterologous sequence of interest are encoded in the same orientation. In some embodiments, the nucleic acid encoding the polypeptide and the template RNA or the nucleic acid encoding the template RNA are covalently linked, e.g., are part of a fusion nucleic acid, and / or are part of the same transcript. In some embodiments, the fusion nucleic acid comprises RNA or DNA.
[0736] The nucleic acid encoding the Gene Writer polypeptide may in some instances be 5' to the heterologous sequence of interest. For example, in some embodiments, the template RNA comprises, from 5' to 3', a 5' untranslated region, a sense-encoded Gene Writer polypeptide, a sense-encoded heterologous sequence of interest, and a 3' untranslated region. In some embodiments, the template RNA comprises, from 5' to 3', a 5' untranslated region, a sense-encoded Gene Writer polypeptide, an antisense-encoded heterologous sequence of interest, and a 3' untranslated region.
[0737] In some embodiments, the RNA further comprises homology to a DNA target site.
[0738] Where a template RNA is described as comprising an open reading frame or its reverse complement, it is understood that in some embodiments the template RNA must be converted to double-stranded DNA (e.g., by reverse transcription) before the open reading frame can be transcribed and translated.
[0739] In certain embodiments, customized RNA sequence templates can be identified, designed, engineered, and constructed to contain sequences that modify or specify host genome function, for example, by introducing heterologous coding regions into the genome; affecting or causing exon structure / alternative splicing; causing disruption of an exogenous gene; causing transcriptional activation of an exogenous gene; causing epigenetic regulation of an exogenous DNA; causing up- or down-regulation of an operably linked gene, and the like. In certain embodiments, customized RNA sequence templates contain sequences that code for exons and / or transgenes, and can be engineered to provide binding sites for transcription factor activators, repressors, enhancers, and the like, as well as combinations thereof. In other embodiments, the coding sequences can also be further customized with splice acceptor sites, polyA tails. In certain embodiments, the RNA sequence can contain sequences that code for an RNA sequence template that is homologous to an RLE transposase, can be engineered to contain heterologous coding sequences, or combinations thereof.
[0740] The template RNA may have some degree of homology to the target DNA. In some embodiments, the template RNA has at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 175, 200 or more bases of strict homology to the target DNA at the 3' end of the RNA. In some embodiments, the template RNA has at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 175, 180, 200 or more bases of at least 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% homology to the target DNA, e.g., at the 5' end of the template RNA. In some embodiments, the template RNA has a 3' untranslated region derived from a non-LTR retrotransposon, e.g., a non-LTR retrotransposon described herein. In some embodiments, the template RNA has a 3' region of at least 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 100, 120, 140, 160, 180, 200 or more bases that are at least 50%, 60%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% homologous to the 3' sequence of a non-LTR retrotransposon, e.g., a non-LTR retrotransposon described herein, e.g., a non-LTR retrotransposon in Tables 1, 2, or 3. In some embodiments, the template RNA comprises a 5' untranslated region derived from a non-LTR retrotransposon, e.g., a non-LTR retrotransposon described herein.In some embodiments, the template RNA has a 5' region of at least 10, 15, 20, 25, 30, 40, 50, 60, 80, 100, 120, 140, 160, 180, or 200 or more bases of at least 40%, 50%, 60%, 70%, 80%, 90%, 95% or more homology to the 5' sequence of a non-LTR retrotransposon, e.g., a non-LTR retrotransposon described herein, e.g., a non-LTR retrotransposon in Table 2 or 3.
[0741] The template RNA component of the Gene Writer genome editing system described herein is typically capable of binding to the Gene Writer genome editing protein of the system. In some embodiments, the template RNA has a 3' region that can bind to the Gene Writer genome editing protein. The binding region, e.g., the 3' region, can be, for example, a structural RNA region that has at least one, two or three hairpin loops and can bind to the Gene Writer genome editing protein of the system.
[0742] The template RNA component of the Gene Writer genome editing system described herein is typically capable of binding to the Gene Writer genome editing protein of the system. In some embodiments, the template RNA has a 5' region capable of binding to the Gene Writer genome editing protein. The binding region, e.g., the 5' region, can be a structural RNA region, e.g., having at least one, two or three hairpin loops, capable of binding to the Gene Writer genome editing protein of the system. In some embodiments, the 5' untranslated region comprises a pseudoknot, e.g., a pseudoknot capable of binding to the Gene Writer genome editing protein.
[0743] In some embodiments, the template RNA (e.g., the untranslated region of a hairpin RNA, e.g., the 5' untranslated region) comprises a stem-loop sequence. In some embodiments, the template RNA (e.g., the untranslated region of a hairpin RNA, e.g., the 5' untranslated region) comprises a hairpin. In some embodiments, the template RNA (e.g., the untranslated region of a hairpin RNA, e.g., the 5' untranslated region) comprises a helix. In some embodiments, the template RNA (e.g., the untranslated region of a hairpin RNA, e.g., the 5' untranslated region) comprises a pseudoknot. In some embodiments, the template RNA comprises a ribozyme. In some embodiments, the ribozyme is similar to a hepatitis delta virus (HDV) ribozyme, e.g., has a secondary structure similar to that of the HDV ribozyme, and / or has one or more activities of the HDV ribozyme, e.g., autocleavage activity. See, e.g., Eickbush et al., Molecular and Cellular Biology, 2010, 3142-3150.
[0744] In some embodiments, the template RNA (e.g., the untranslated region of a hairpin RNA, e.g., the 3' untranslated region) comprises one or more stem-loop sequences or helices. Exemplary structures of the R2 3'UTR are shown, for example, in Ruschak et al. "Secondary structure models of the 3' untranslated regions of diverse R2 RNAs" RNA. 2004 Jun;10(6):978-987, e.g., FIG. 3 therein, and in Eikbush and Eikbush, "R2 and R2 / R1 hybrid non-autonomous retrotransposons derived by internal deletions of full-length elements" Mobile DNA (2012) 3:10; e.g., FIG. 3 therein, which are incorporated by reference in their entireties.
[0745] In some embodiments, the template RNA described herein comprises a sequence capable of binding to a GeneWriter protein described herein. For example, in some embodiments, the template RNA comprises an MS2 RNA sequence capable of binding to an MS2 coat protein sequence in a GeneWriter protein. In some embodiments, the template RNA comprises an RNA sequence capable of binding to a B-box sequence. In some embodiments, the template RNA comprises an RNA sequence capable of binding to a dCas sequence in a GeneWriter protein (e.g., a crRNA sequence and / or a tracrRNA sequence). In some embodiments, in addition to or instead of a UTR, the template RNA is linked (e.g., covalently) to a non-RNA UTR, e.g., a protein or a small molecule.
[0746] In some embodiments, the template RNA has a poly-A tail at the 3' end. In some embodiments, the template RNA does not include a poly-A tail at the 3' end.
[0747] In some embodiments, the template RNA has a 5' region of at least 10, 15, 20, 25, 30, 40, 50, 60, 80, 100, 120, 140, 160, 180, 200 or more bases that are at least 40%, 50%, 60%, 70%, 80%, 90%, 95% or more homologous to the 5' sequence of a non-LTR retrotransposon, e.g., a non-LTR retrotransposon described herein.
[0748] The template RNA of this system typically contains a sequence of interest for insertion into the target DNA. The sequence of interest may be coding or non-coding.
[0749] In some embodiments, the systems or methods described herein include a single template RNA. In some embodiments, the systems or methods described herein include multiple template RNAs.
[0750] In some embodiments, the sequence of interest may include an open reading frame. In some embodiments, the template RNA has a Kozak sequence. In some embodiments, the template RNA has an internal ribosome entry site. In some embodiments, the template RNA has a self-cleaving peptide, such as a T2A or P2A site. In some embodiments, the template RNA has a start codon. In some embodiments, the template RNA has a splice acceptor site. In some embodiments, the template RNA has a splice donor site. In some embodiments, the template RNA has a microRNA binding site downstream of the stop codon. In some embodiments, the template RNA has a polyA tail downstream of the stop codon of the open reading frame. In some embodiments, the template RNA includes one or more exons. In some embodiments, the template RNA includes one or more introns. In some embodiments, the template RNA includes a eukaryotic transcription terminator. In some embodiments, the template RNA includes an enhanced translation element or a translation enhancing element. In some embodiments, the RNA comprises a human T-cell leukemia virus (HTLV-1) R region. In some embodiments, the RNA comprises a post-translational regulatory element that enhances nuclear export, such as a post-translational regulatory element of Hepatitis B virus (HPRE) or Woodchuck Hepatitis virus (WPRE). In some embodiments, in the template RNA, the heterologous sequence of interest encodes a polypeptide and is encoded in the antisense orientation relative to the 5' and 3' UTR. In some embodiments, in the template RNA, the heterologous sequence of interest encodes a polypeptide and is encoded in the sense orientation relative to the 5' and 3' UTR.
[0751] In some embodiments, the nucleic acid described herein (e.g., template RNA or DNA encoding the template RNA) comprises a microRNA binding site. In some embodiments, the microRNA binding site is used to increase the target cell specificity of the GeneWriter system. For example, the microRNA binding site can be selected based on the criterion that it is recognized by a miRNA present in a non-target cell type, but not present in a target cell type (or present at a lower level compared to the target cell). Thus, if the template RNA is present in a non-target cell, it will be bound by the miRNA, and if the template RNA is present in a target cell, it will not be bound by the miRNA (or will be bound at a lower level compared to the target cell). Without wishing to be bound by theory, the binding of the template RNA to the miRNA may hinder the insertion of a heterologous sequence of interest into the genome. Thus, the heterologous sequence of interest may be inserted into the genome of a target cell more efficiently than into the genome of a non-target cell. A system having a microRNA binding site in the template RNA (or the DNA encoding it) can also be used in combination with a nucleic acid encoding a GeneWriter polypeptide, where expression of the GeneWriter polypeptide is regulated by a second microRNA binding site, e.g., as described herein, e.g., in the section entitled "Polypeptide Components of the Gene Writer Gene Editor System."
[0752] In some embodiments, the sequence of interest may include a non-coding sequence. For example, the template RNA may include a promoter or enhancer sequence. In some embodiments, the template RNA includes a tissue-specific promoter or enhancer, each of which may be unidirectional or bidirectional. In some embodiments, the promoter is an RNA polymerase I promoter, an RNA polymerase II promoter, or an RNA polymerase III promoter. In some embodiments, the promoter is a TATA element. In some embodiments, the promoter includes a B recognition element. In some embodiments, the promoter has one or more binding sites for a transcription factor. In some embodiments, the non-coding sequence is transcribed in an antisense direction relative to the 5' and 3' UTRs. In some embodiments, the non-coding sequence is transcribed in a sense direction relative to the 5' and 3' UTRs.
[0753] In some embodiments, the nucleic acid described herein (e.g., the template RNA or DNA encoding the template RNA) includes a promoter sequence, e.g., a tissue-specific promoter sequence. In some embodiments, a tissue-specific promoter is used to increase the target cell specificity of the GeneWriter system. For example, a promoter can be selected based on the criteria that it is active in the target cell type but not active (or active at a low level) in non-target cell types. Thus, even if the promoter is integrated into the genome of a non-target cell, it does not drive expression of the integrated gene (or drives only low levels of expression). A system with a tissue-specific promoter sequence in the template RNA may also be used in combination with a microRNA binding site, e.g., in the template RNA or nucleic acid encoding a GeneWriter protein, as described herein. Furthermore, a system with a tissue-specific promoter sequence in the template RNA may be used in combination with a DNA encoding a GeneWriter polypeptide driven by a tissue-specific promoter to achieve higher levels of GeneWriter protein in target cells than in non-target cells.
[0754] In some embodiments, the template RNA comprises a microRNA sequence, an siRNA sequence, a guide RNA sequence, a piwiRNA sequence.
[0755] In some embodiments, the template RNA comprises a site that regulates epigenetic modification. In some embodiments, the template RNA comprises an element that inhibits, e.g., blocks, epigenetic silencing. In some embodiments, the template RNA comprises a chromatin insulator. For example, the template RNA comprises a CTCF site or a site that is targeted for DNA methylation.
[0756] To promote higher levels or more stable gene expression, the template RNA may include features that prevent or inhibit gene silencing. In some embodiments, these features prevent or inhibit DNA methylation. In some embodiments, these features promote DNA demethylation. In some embodiments, these features prevent or inhibit histone deacetylation. In some embodiments, these features prevent or inhibit histone methylation. In some embodiments, these features promote histone acetylation. In some embodiments, these features promote histone demethylation. In some embodiments, multiple features can be incorporated into the template RNA to promote one or more of these modifications. CpG dinucleotides undergo methylation by host methyltransferases. In some embodiments, the template RNA is depleted of CpG dinucleotides, e.g., does not contain CpG nucleotides or contains a low number of CpG dinucleotides compared to the corresponding unmodified sequence. In some embodiments, the promoter driving transgene expression from the integrated DNA is depleted of CpG dinucleotides.
[0757] In some embodiments, the template RNA comprises a gene expression unit consisting of at least one regulatory region operably linked to an effector sequence, which may be a sequence (e.g., a coding sequence, such as a sequence encoding a microRNA, or a non-coding sequence) that is transcribed into RNA.
[0758] In some embodiments, the sequence of interest of the template RNA is inserted into the target genome within an endogenous intron. In some embodiments, the sequence of interest of the template RNA is inserted into the target genome, thereby acting as a new exon. In some embodiments, the insertion of the sequence of interest into the target genome results in the replacement of a native exon or the skipping of a native exon.
[0759] In some embodiments, the sequence of interest of the template RNA is inserted into the target genome at a genomic safe harbor site such as AAVS1, CCR5, or ROSA26. In some embodiments, the sequence of interest of the template RNA is added to the genome in an intergenic or intragenic region. In some embodiments, the sequence of interest of the template RNA is added to the genome 5' or 3' within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb from the endogenous active gene. In some embodiments, the sequence of interest of the template RNA is added to the genome 5' or 3' within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb from the endogenous promoter or enhancer. In some embodiments, the sequence of interest of the template RNA can be, for example, 50 to 50,000 base pairs (e.g., 50 to 40,000 bp, 500 to 30,000 bp, 500 to 20,000 bp, 100 to 15,000 bp, 500 to 10,000 bp, 50 to 10,000 bp, 50 to 50,000 bp. In some embodiments, the heterologous sequence of interest is less than 1,000, 1,300, 1500, 2,000, 3,000, 4,000, 5,000, or 7,500 nucleotides in length.
[0760] In some embodiments, the genomic safe harbor site is a Natural Harbor™ site. In some embodiments, the Natural Harbor™ site is ribosomal DNA (rDNA). In some embodiments, the Natural Harbor™ site is 5SrDNA, 18SrDNA, 5.8SrDNA, or 28SrDNA. In some embodiments, the Natural Harbor™ site is a Mutsu site in 5SrDNA. In some embodiments, the Natural Harbor™ site is an R2 site, an R5 site, an R6 site, an R4 site, an R1 site, an R9 site, or an RT site in 28SrDNA. In some embodiments, the Natural Harbor™ site is an R8 site or an R7 site in 18SrDNA. In some embodiments, the Natural Harbor™ site is DNA encoding a transfer RNA (tRNA). In some embodiments, the Natural Harbor™ site is DNA encoding a tRNA-As or a tRNA-Glu. In some embodiments, the Natural Harbor™ site is DNA encoding a spliceosomal RNA. In some embodiments, the Natural Harbor™ site is DNA encoding a small nuclear RNA (snRNA), such as U2 snRNA.
[0761] Thus, in some aspects, the disclosure provides a method of inserting a heterologous sequence of interest into a Natural Harbor™ site. In some embodiments, the method includes using a GeneWriter system as described herein, e.g., using any of the polypeptides in Tables 1-3, or a polypeptide having similarity thereto, e.g., at least 80%, 85%, 90%, or 95% identity thereto. In some embodiments, the method includes inserting a heterologous sequence of interest into a Natural Harbor™ site using an enzyme, e.g., a retrotransposase. In some aspects, the disclosure provides a host human cell comprising a heterologous sequence of interest (e.g., a sequence encoding a therapeutic polypeptide) located at a Natural Harbor™ site in the genome of the cell. In some embodiments, the Natural Harbor™ site is a site set forth in Table 4 below. In some embodiments, the heterologous sequence of interest is inserted within 20, 50, 100, 150, 200, 250, 500, or 1000 base pairs of the sequence set forth in Table 4. In some embodiments, the heterologous sequence of interest is inserted within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb of a sequence shown in Table 4. In some embodiments, the heterologous sequence of interest is inserted at a site having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to a sequence shown in Table 4. In some embodiments, the heterologous sequence of interest is inserted within 20, 50, 100, 150, 200, 250, 500, or 1000 base pairs, or within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb of a site having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to a sequence shown in Table 4.In some embodiments, the heterologous sequence of interest is inserted within a gene shown in column 5 of Table 4, or within 20, 50, 100, 150, 200, 250, 500, or 1000 base pairs of that gene, or within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb.
[0762] [Table 470]
[0763] [Table 471]
[0764] [Table 472]
[0765] [Table 473]
[0766] In some embodiments, the systems or methods described herein achieve insertion of a heterologous sequence into a target site in a human genome. In some embodiments, the target site in the human genome has sequence similarity to the corresponding target site of a corresponding wild-type retrotransposase (e.g., the retrotransposase from which the GeneWriter is derived) in the genome of the organism from which it is naturally derived. For example, in some embodiments, the identity between 40 nucleotides of the human genome sequence centered at the insertion site and 40 nucleotides of the native organism genome centered at the insertion site is less than 99.5%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 60%, or 50%, or is in the range of 50-60%, 60-70%, 70-80%, 80-90%, or 90-100%. In some embodiments, the identity between 100 nucleotides of the human genomic sequence centered at the insertion site and 100 nucleotides of the native organism genomic sequence centered at the insertion site is less than 99.5%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 60%, or 50%, or is in the range of 50-60%, 60-70%, 70-80%, 80-90%, or 90-100%. In some embodiments, the identity between 500 nucleotides of the human genomic sequence centered at the insertion site and 500 nucleotides of the native organism's genome centered at the insertion site is less than 99.5%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 60%, or 50%, or is in the range of 50-60%, 60-70%, 70-80%, 80-90%, or 90-100%.
[0767] Preparation of compositions and systems As will be appreciated by those of skill in the art, methods for designing and constructing nucleic acid constructs and proteins or polypeptides (e.g., the systems, constructs and polypeptides described herein) are routine in the art. Generally, recombinant methods may be used. For overviews, see Smales & James (Eds.), Therapeutic Proteins: Methods and Protocols (Methods in Molecular Biology), Humana Press (2005); and Crommelin, Sindelar & Meibohm (Eds.), Pharmaceutical Biotechnology: Fundamentals and Applications, Springer (2013). Methods for designing, preparing, evaluating, purifying and manipulating nucleic acid compositions are described in Green and Sambrook (Eds.), Molecular Cloning: A Laboratory Manual (Fourth Edition), Cold Spring Harbor Laboratory Press (2012).
[0768] Exemplary methods for producing pharmaceutical proteins or polypeptides described herein include expression in mammalian cells, although recombinant proteins can also be produced using insect cells, yeast, bacteria, or other cells under the control of an appropriate promoter. Mammalian expression vectors can include non-transcriptional elements, such as an origin of replication, a suitable promoter, and other 5' or 3' flanking non-transcriptional sequences, and 5' or 3' non-transcriptional sequences, such as necessary ribosome binding sites, polyadenylation sites, splice donor and acceptor sites, and termination sequences. DNA sequences derived from the SV40 viral genome, such as SV40 origin, early promoter, splice, and polyadenylation sites, can be used to provide other genetic elements required for expression of heterologous DNA sequences. Cloning and expression vectors suitable for use in bacterial, fungal, yeast, and mammalian cell hosts are described in Green & Sambrook, Molecular Cloning: A Laboratory Manual (Fourth Edition), Cold Spring Harbor Laboratory Press (2012).
[0769] A variety of mammalian cell culture systems can be used to express and produce recombinant proteins. Examples of mammalian expression systems include CHO, COS, HEK293, HeLA, and BHK cell lines. Host culture methods for protein therapeutics production are described in Zhou and Kantardjieff (Eds.), Mammalian Cell Cultures for Biologics Manufacturing (Advances in Biochemical Engineering / Biotechnology), Springer (2014). The compositions described herein can include a vector encoding a recombinant protein, e.g., a viral vector, such as a lentiviral vector. In some embodiments, the vector, e.g., a viral vector, can include a nucleic acid encoding a recombinant protein.
[0770] Purification of protein therapeutics is described in Franks, Protein Biotechnology: Isolation, Characterization, and Stabilization, Humana Press (2013); and Cutler, Protein Purification Protocols (Methods in Molecular Biology), Humana Press (2010).
[0771] Applicable The Gene Writer system can address therapeutic needs by incorporating coding genes into the RNA sequence template, for example, by providing expression of therapeutic transgenes in individuals with loss-of-function mutations, by replacing gain-of-function mutations with normal transgenes, by providing regulatory sequences to eliminate gain-of-function mutation expression, and / or by controlling expression of operably linked genes, transgenes and systems thereof. In certain embodiments, the RNA sequence template encodes a promoter region specific to the therapeutic needs of the host cell, e.g., a tissue-specific promoter or enhancer. In yet another embodiment, a promoter can be operably linked to the coding sequence.
[0772] In some embodiments, the Gene Writer™ gene editor system can provide therapeutic transgenes that express, for example, substitute blood factors or substitute enzymes, such as lysosomal enzymes. For example, the compositions, systems and methods described herein are useful for expressing in a target human genome agalsidase alfa or beta for the treatment of Fabry disease; imiglucerase, taliglucerase alfa, velaglucerase alfa, or alglucerase for Gaucher disease; sebelipase alfa for lysosomal acid lipase deficiency (Wolman disease / CESD); laronidase, idursulfase, elosulfase alfa, or galsulfase for mucopolysaccharidoses; and alglucosidase for Pompe disease. For example, the compositions, systems and methods described herein are useful for expressing in a target human genome factors I, II, V, VII, X, XI, XII, or XIII for blood factor deficiencies.
[0773] In some embodiments, the heterologous sequence of interest encodes an intracellular protein (e.g., an organelle protein such as a cytoplasmic protein, a nuclear protein, a mitochondrial protein or a lysosomal protein, or a membrane protein). In some embodiments, the heterologous sequence of interest encodes a membrane protein, e.g., a membrane protein other than CAR, and / or an integral human membrane protein. In some embodiments, the heterologous sequence of interest encodes an extracellular protein. In some embodiments, the heterologous sequence of interest encodes an enzyme, a structural protein, a signaling protein, a regulatory protein, a transport protein, a sensor protein, a motor protein, a defense protein, or a storage protein.
[0774] Administration The compositions and systems described herein may be used in vitro or in vivo. In some embodiments, the system or components of the system are delivered to a cell (e.g., a mammalian cell, e.g., a human cell), for example, in vitro or in vivo. In some embodiments, the cell is a eukaryotic cell, e.g., a cell of a multicellular organism, e.g., an animal, e.g., a mammal (e.g., a human, a pig, a cow), a bird (e.g., a poultry such as a chicken, a turkey, or a duck), or a fish. In some embodiments, the cell is a non-human animal cell (e.g., a laboratory animal, a livestock animal, or a pet). In some embodiments, the cell is a stem cell (e.g., a hematopoietic stem cell), a fibroblast, or a T cell. In some embodiments, the cell is a non-dividing cell, e.g., a non-dividing fibroblast or a non-dividing T cell. In some embodiments, the cell is an HSC, and p53 is not upregulated or is upregulated by less than 10%, 5%, 2%, or 1%, for example, as measured according to the method described in Example 30. Those skilled in the art will appreciate that the components of the Gene Writer system can be delivered in the form of polypeptides, nucleic acids (eg, DNA, RNA), and combinations thereof.
[0775] For example, delivery can use any combination of the following to deliver the retrotransposase (e.g., as DNA encoding the retrotransposase protein, as RNA encoding the retrotransposase protein, or as the protein itself) and template RNA (e.g., as DNA encoding the RNA or as the RNA): 1. Retrotransposase DNA + template DNA 2. Retrotransposase RNA + template DNA 3. Retrotransposase DNA + template RNA 4. Retrotransposase RNA + template RNA 5. Retrotransposase protein + template DNA 6. Retrotransposase protein + template RNA 7. Retrotransposase virus + template virus 8. Retrotransposase virus + template DNA 9. Retrotransposase virus + template RNA 10. Retrotransposase DNA + template virus 11. Retrotransposase RNA + template virus 12. Retrotransposase protein + template virus
[0776] As described above, in some embodiments, DNA or RNA encoding a retrotransposase protein is delivered using a virus, and in some embodiments, the template RNA (or DNA encoding the template RNA) is delivered using a virus.
[0777] In some embodiments, the system and / or system components are delivered as nucleic acids. For example, the Gene Writer polypeptide may be delivered in the form of DNA or RNA encoding the polypeptide, and the template RNA may be delivered in the form of RNA or its complementary DNA to be transcribed into RNA. In some embodiments, the system or system components are delivered on one, two, three, four, or more separate nucleic acid molecules. In some embodiments, the system or system components are delivered as a combination of DNA and RNA. In some embodiments, the system or system components are delivered as a combination of DNA and protein. In some embodiments, the system or system components are delivered as a combination of RNA and protein. In some embodiments, the system or system components are delivered as a combination of RNA and protein. In some embodiments, the system or system components are delivered as a combination of RNA and protein. In some embodiments, the Gene Writer genome editor polypeptide is delivered as a protein.
[0778] In some embodiments, the system or components of the system are delivered to cells, e.g., mammalian or human cells, using a vector. The vector may be, for example, a plasmid or a virus. In some embodiments, the delivery is in vivo, in vitro, ex vivo, or in situ. In some embodiments, the virus is an adeno-associated virus (AAV), a lentivirus, or an adenovirus. In some embodiments, the system or components of the system are delivered to cells using a virus-like particle or a virosome. In some embodiments, the delivery uses a virus, a virus-like particle, or a virosome.
[0779] In one embodiment, the compositions and systems described herein can be formulated in liposomes or other similar vesicles. Liposomes are spherical vesicular structures composed of a single or multilayer lipid bilayer membrane surrounding an internal aqueous compartment and a relatively impermeable outer lipophilic phospholipid bilayer membrane. Liposomes can be anionic, neutral or cationic. Liposomes are biocompatible, non-toxic, and can deliver both hydrophilic and lipophilic drug molecules, protecting their cargo from degradation by plasma enzymes and transporting their load across biological membranes and the blood-brain barrier (BBB) (see, e.g., Spuch and Navarro, Journal of Drug Delivery, vol. 2011, Article ID 469679, 12 pages, 2011. doi:10.1155 / 2011 / 469679 for a review).
[0780] Vesicles can be made from several different types of lipids; however, phospholipids are most commonly used to make liposomes as drug carriers. Methods for preparing multilamellar vesicle lipids are known in the art (see, for example, U.S. Pat. No. 6,693,086; its teachings on preparing multilamellar vesicle lipids are incorporated herein by reference). When lipid membrane is mixed with aqueous solution, vesicle formation can occur spontaneously, but it can also be promoted by applying force in the form of shaking using a homogenizer, ultrasonicator, or extruder (see, for example, Spuch and Navarro, Journal of Drug Delivery, vol. 2011, Article ID 469679, 12 pages, 2011.doi:10.1155 / 2011 / 469679 for review). Extruded lipids can be prepared by extrusion through size-reducing filters as described in Templeton et al., Nature Biotech, 15:647-652, 1997, the teachings of which regarding extruded lipid preparation are incorporated herein by reference.
[0781] Lipid nanoparticles are another example of carriers that provide a biocompatible and biodegradable delivery system for the pharmaceutical compositions described herein. Nanostructured lipid carriers (NLCs) are modified solid lipid nanoparticles (SLNs) that retain the characteristics of SLNs, improve drug stability and loading capacity, and prevent drug leakage. Polymer nanoparticles (PNPs) are an important component of drug delivery. These nanoparticles effectively direct drug delivery to specific targets, improve drug stability and drug release control. Lipid-polymer nanoparticles (PLNs), a new type of carrier that combines liposomes and polymers, may also be used. These nanoparticles have the complementary advantages of PNPs and liposomes. PLNs are composed of a core-shell structure; the polymer core provides a stable structure, and the phospholipid shell confers good biocompatibility. Thus, the two components enhance drug encapsulation efficiency, promote surface modification, and prevent leakage of water-soluble drugs. See, e.g., Li et al. 2017, Nanomaterials 7, 122; doi:10.3390 / nano7060122 for a review.
[0782] Exosomes can also be used as drug delivery vehicles for the compositions and systems described herein. See, e.g., Ha et al. July 2016. Acta Pharmaceutica Sinica B. Volume 6, Issue 4, Pages 287-296; https: / / doi.org / 10.1016 / j.apsb.2016.02.001 for review.
[0783] The GeneWriter system can be introduced into cells, tissues, and multicellular organisms. In some embodiments, the system or components of the system are delivered to cells using mechanical or physical means.
[0784] Formulation of protein therapeutics is described in Meyer (Ed.), Therapeutic Protein Drug Products: Practical Approaches to formulation in the Laboratory, Manufacturing, and the Clinic, Woodhead Publishing Series (2012).
[0785] All publications, patent applications, patents, and other publications and references (e.g., sequence database reference numbers) cited herein are incorporated by reference in their entirety. For example, GenBank, Unigene, and Entrez referenced herein, for example in any table herein, are all incorporated by reference. Unless otherwise indicated, sequence accession numbers listed herein, including all tables herein, refer to the most recent database entries as of August 27, 2018. When a gene or protein has multiple sequence accession numbers, all of these sequence variants are encompassed. EXAMPLES
[0786] The present invention is further described by the following examples, which are provided for illustrative purposes and should not be construed as limiting the scope or content of the invention.
[0787] Example 1: Delivery of the Gene Writer™ System into Mammalian Cells This example describes the Gene Writer™ genome editing system delivered to mammalian cells for site-specific insertion of exogenous DNA into the mammalian cell genome.
[0788] In this example, the polypeptide component of the Gene Writer™ system is the R2Bm protein from the silkworm (Bombyx mori), and the template RNA component is the RNA of the R2Bm retrotransposase from the silkworm (Bombyx mori) that contains a mutation in the reverse transcriptase domain that renders the retrotransposase inactive.
[0789] HEK293T cells are transfected with the following test substances: 1. Scrambled RNA Control 2. RNA encoding the aforementioned polypeptide 3. Template RNA as described above 4. Combination of 2 and 3
[0790] After transfection, HEK293T cells are cultured for at least 4 days and assayed for site-specific genome editing. Genomic DNA is isolated from each group of HEK293 cells. PCR is performed using primers flanking the R2Bm integration site in the 28srRNA gene. The PCR products are run on an agarose gel to measure the length of the amplified DNA.
[0791] PCR products of the expected length, indicating a successful Gene Writing™ genome editing event that inserted the mutant R2Bm retrotransposase sequence into the target genome, are observed only in cells transfected with the complete Gene Writer™ System in Group 4 above.
[0792] Example 2: Site-specific delivery of the Gene Writer™ system into insect cells This example describes the Gene Writer™ genome editing system delivered to insect cells at specific target sites in the genome.
[0793] In this example, the polypeptide component of the Gene Writer™ system is derived from R2Bm of the silkworm Bombyx mori, which has been modified by replacing its DNA-binding domain at the amino terminus of the polypeptide with a heterologous zinc finger DNA-binding domain. The zinc finger DNA-binding domain has been shown to bind DNA in the BmBLOS2 locus of B. mori cells (Takasu et al., insect Biochemistry and Molecular Biology 40(10):759-765, 2010). The template RNA is the RNA of R2Bm retrotransposase from Bombyx mori, which contains a mutation in the reverse transcriptase domain that renders the retrotransposase inactive. Additionally, the template RNA is modified at the 5' end to have 180 bases of homology to the target DNA site.
[0794] The silkworm (B. mori) insect cell line is transfected with the following test substances: 1. Scrambled RNA Control 2. RNA encoding the aforementioned polypeptide component 3. Template RNA as described above 4. Combination of 2 and 3
[0795] After transfection, the cells are cultured for at least 4 days and assayed for site-specific Gene Writing™ genome editing. Genomic DNA is isolated from the cells and PCR is performed using primers flanking the target integration site in the genome. The PCR products are run on an agarose gel to measure the length of the DNA. PCR products of the expected length are observed only in cells transfected with the complete Gene Writer™ system in group 4 above, indicating the success of the Gene Writing™ genome editing event that inserts the mutant R2Bm retrotransposase sequence into the target insect cell genome.
[0796] Example 3: Site-specific delivery of the Gene Writer™ system into mammalian cells This example describes the Gene Writer™ genome editor system, which is used to insert heterologous sequences at specific sites in a mammalian genome.
[0797] In this example, the polypeptide of the system is the R2Bm protein from the silkworm Bombyx mori, and the template RNA component is an RNA encoding the GFP protein, which is flanked at its 5' end by the 5'UTR and at its 3' end by the 3'UTR of the R2Bm retrotransposase from the silkworm Bombyx mori. The GFP gene has a ribosome entry site upstream of its start codon and a polyA tail downstream of its stop codon.
[0798] HEK293 cells are transfected with the following test substances: 1. Scrambled RNA Control 2. RNA encoding the aforementioned polypeptide 3. Template RNA encoding the aforementioned GFP 4. Combination of 2 and 3
[0799] After transfection, HEK293 cells are cultured for at least 4 days and assayed for site-specific Gene Writing™ genome editing events. Genomic DNA is isolated from HEK293 cells and PCR is performed using primers flanking the R2Bm integration site in the 28s rRNA gene. The PCR products are run on an agarose gel to measure the length of the DNA. PCR products of the expected length, indicating successful Gene Writing™ genome editing events, are detected in cells transfected with the test material of group 4 (complete Gene Writer™ system). This result demonstrates that the Gene Writing genome editing system can insert new transgenes into mammalian cell genomes.
[0800] The transfected cells are further cultured for 10 days and after multiple cell culture passages, GFP expression is assayed by flow cytometry. The percentage of GFP-positive cells from each cell population is calculated. GFP-positive cells are detected in cells transfected with the test material of group 4 (complete Gene Writer™ system). This result demonstrates that the new transgene written into the mammalian cell genome is expressed.
[0801] Example 4: Targeted delivery of gene expression units into mammalian cells using the Gene Writer™ system This example describes the construction and use of the Gene Writer genome editor to insert heterologous gene expression units into a mammalian genome.
[0802] In this example, the polypeptide of the Gene Writer system is derived from the R2Bm polypeptide of the silkworm (Bombyx mori) modified by replacing its DNA-binding domain at the amino terminus of the polypeptide with a heterologous zinc finger DNA-binding domain. The zinc finger DNA-binding domain has been shown to bind to DNA in the AAVS1 locus in human cells (Hockemeyer et al., Nature Biotechnology 27(9):851-857, 2009). The template RNA comprises a gene expression unit. The gene expression unit comprises at least one regulatory sequence operably linked to at least one coding sequence. In this example, the regulatory sequence comprises a CMV promoter and enhancer, an enhanced translation element, and a WPRE. The coding sequence is a GFP open reading frame. The gene expression unit is flanked at the 5' end by 180 bases of homology to the target DNA site and at the 3' end by the 3'UTR of the R2Bm retrotransposase from the silkworm (Bombyx mori).
[0803] HEK293 cells are transfected with the following test substances: 1. Scrambled RNA Control 2. RNA encoding the aforementioned polypeptide component 3. Template RNA containing the gene expression unit (as described above) 4. Complete Gene Writer System including both (2) and (3)
[0804] After transfection, HEK293 cells are cultured for at least 4 days and assayed for site-specific Gene Writing genome editing. Genomic DNA is isolated from HEK293 cells and PCR is performed using primers flanking the target integration site in the genome. PCR products are run on agarose gel to measure DNA length. PCR products of expected length are detected in cells transfected with test material of group 4 (complete Gene Writer™ system), indicating successful Gene Writing™ genome editing events.
[0805] Transfected cells are further cultured for 10 days, and after multiple cell culture passages, GFP expression is examined by flow cytometry. The percentage of GFP-positive cells is calculated from each cell population. GFP-positive cells are detected in the population of HEK293 cells transfected with test substances of group 4, which proves that the gene expression unit added to the mammalian cell genome by Gene Writing genome editing is expressed.
[0806] Example 5: Targeted delivery of gene expression units to intron regions of mammalian cells using the Gene Writer™ system This example describes the construction and use of the Gene Writing genome editing system to add heterologous sequences to intronic regions to act as splice acceptors for upstream exons.
[0807] The target integration site is the first intron of the albumin locus. Splicing of a new exon containing a splice acceptor site at the 5' end and a polyA tail at the 3' end into the first intron results in a mature mRNA containing the first native exon of the albumin locus spliced to the new exon. Because the first exon of albumin is removed during protein processing, cells expressing the newly formed gene unit will secrete a mature protein containing only the new exon.
[0808] In this example, the Gene Writer genome editor polypeptide is derived from the R2Bm Gene Writer genome editor of the silkworm Bombyx mori, which has been modified by replacing the amino-terminal DNA binding domain of the polypeptide with a heterologous zinc finger DNA binding domain. The zinc finger DNA binding domain has been shown to bind tightly to the albumin locus within the first intron, as described in Sarma et al., Blood 126,15:1777-1784,2015. The template RNA is an RNA encoding EPO, containing a splice acceptor site 5' adjacent to the first amino acid of mature EPO (with the start codon and signal peptide removed) and a 3' polyA tail downstream of the stop codon. The EPO RNA is flanked at the 5' end by 180 bases of homology to the target DNA site and at the 3' end by the 3'UTR of the R2Bm retrotransposase from the silkworm Bombyx mori.
[0809] HEK293 cells are transfected with the following test substances: 1. Scrambled RNA Control 2. RNA encoding the aforementioned polypeptide 3. Template RNA containing the EPO splice acceptor described above 4. Complete Gene Writer System including both (2) and (3)
[0810] After transfection, HEK293 cells are cultured for at least 4 days to assay for site-specific Gene Writing genome editing and proper mRNA processing. Genomic DNA is isolated from HEK293 cells. Reverse transcription-PCR is performed to measure mature mRNA containing the first native exon of the albumin locus and the new exon. RT-PCR reactions are performed with a forward primer that binds to the first native exon of the albumin locus and a reverse primer that binds to EPO. The RT-PCR products are run on an agarose gel to measure the length of the DNA. PCR products of the expected length, indicating a successful Gene Writing genome editing event, are detected in cells transfected with the test substance of group 4. This result demonstrates that the Gene Writing genome editing system can add a heterologous sequence encoding a gene to an intron region to act as a splice acceptor for an upstream exon.
[0811] The transfected cells are further cultured for 10 days, and after multiple cell culture passages, the cell supernatant is assayed for EPO secretion. The amount of EPO in the supernatant is measured by EPO ELISA kit. EPO is detected in HEK293 cells transfected with test substances of group 4, which proves that heterologous sequence can be added to intron region by Gene Writing genome editing to act as splice acceptor of upstream exon, and this sequence is actively expressed.
[0812] Example 6: Targeted delivery of R2Tg retrotransposons into mammalian cells This example describes the targeted integration of the R2Tg retrotransposon element (see row 1 of Table 3 herein) into mammalian cells by DNA or RNA delivery.
[0813] R2Tg is an endogenous retrotransposon derived from the zebra finch (Taenopygia guttata). Because non-LTR R2 elements are absent from the human genome and are believed to be highly site-specific, the ability of R2Tg to precisely and efficiently integrate itself into the human genome may demonstrate the ability to perform genome-targeted integration and potentially enable human gene therapy.
[0814] For the DNA delivery method, a plasmid carrying R2Tg (PLV014) was designed and synthesized such that the R2Tg element was codon-optimized and flanked with its native untranslated region (UTR), with or without an additional flanking 100 bp homology to the rDNA target locus. R2Tg element expression was driven by a mammalian CMV promoter. In addition, a 1 bp deletion mutant (678*) with a frameshift within the coding sequence of the retrotransposase was constructed as an inactivation control ("frameshift mutant"). Each plasmid was introduced into HEK393T cells by FuGENE® HD transfection reagent. 24 hours prior to transfection, HEK293T cells were seeded at 10,000 cells / well in a 96-well plate. On the day of transfection, 0.5 μl of transfection reagent and 80 ng of DNA were mixed in 10 μl of Opti-MEM and incubated at room temperature for 15 minutes. The transfection mixture was then added to the culture medium of the inoculated cells. Three days after transfection, genomic DNA was extracted for retrotransposition assay.
[0815] Next, we assessed the integration of the R2Tg transposase into the human genome. Putative integration sites in human rDNA were examined based on homology to the finch genome. Integration was assessed using Advanced Miseq and ddPCR assays.
[0816] Bias during Miseq library construction was eliminated by introducing random unique molecular indices (UMIs) in the first PCR (Figure 7). First, a nested PCR was performed by amplifying the predicted 3' junction of the R2Tg and rDNA locus for 30 cycles. At this stage, one Miseq adapter, multiplexed barcode, and 8bp UMI were introduced. A second PCR was used to further enrich the predicted products and add a second Miseq adapter. Samples were sequenced for 300 cycles on Miseq. After demultiplexing, samples were analyzed by Matlab. First, UMIs on each sequence were located by searching the adjacent sequences. A database of UMIs was created and then collapsed by uniqueness. For each unique read, a search was performed for the sequence of the predicted rDNA integration site, as well as the isolated sequences of aligned human genomic DNA and exogenous DNA. The exogenous DNA was then aligned with the predicted integration sequence. The results of the Miseq analysis pipeline are shown in Figures 8A-8B. We found abundant unique integrations at the predicted integration sites in cells treated with wild-type R2Tg constructs flanking 100 bp homology to the target rDNA, but not in those treated with the frameshift mutant control. Most integration events have the complete template RNA sequence integrated 565 bp most proximal to the integration site, as evidenced by sequencing reads that perfectly align with the expected sequence. A subset of integration events involving experimental R2Tg have either a 300 bp or 450 bp truncation, based on sequencing reads that align with the expected sequence after a gap adjacent to the target site (Figure 8A). More specifically, 86.17% of the integrants observed were uncut at the 565 bp most proximal to the integration site. In contrast, Figure 8B shows no detection of integration events at all. Constructs without flanking rDNA homology showed insignificant integration signals near noise.
[0817] Next, ddPCR is performed to confirm integration and evaluate integration efficiency. A Taqman probe was designed in the 3'UTR portion of the R2Tg element. The forward primer was synthesized to bind directly upstream of the probe, and the reverse primer was synthesized to bind to the rDNA. Thus, amplification of the expected product across the integration junction results in the degradation of the probe and the generation of a fluorescent signal. ddPCR was performed on several replication experiments of the aforementioned plasmids to determine the average copy number of R2Tg integration events. The results of the ddPCR copy number analysis (relative to the reference gene RPP30) are shown in Figure 9. Among several plasmid transfection conditions, integration of an average of 5 or more copies of R2Tg per genome at the target site was observed when delivered with homology, with a significant increase over the control construct. In contrast, the average copy number per genome in the frameshift mutant negative control was typically less than 1. Insignificant signals were observed when constructs without homology were delivered to cells. These experiments suggest efficient integration of the R2Tg retrotransposon into human cells at the target site.
[0818] For the RNA delivery method, R2Tg RNA (RNA V019) was designed such that the R2Tg element was codon-optimized and flanked by its native untranslated region (UTR). More specifically, the construct contained, in order: a T7 promoter, a 5' 28S target homology region 100 nucleotides long, an R2Tg wild-type 5'UTR, an R2Tg codon-optimized coding sequence, an R2Tg wild-type 3'UTR, and a 3' 28S target homology region 100 nucleotides long. To enhance integration, a 100 bp 28S homology sequence was added outside the UTR. R2Tg RNA was synthesized, capped and polyA-tailed. R2Tg element transcription was driven by the T7 promoter. RNA was introduced into HEK393T cells using Lipofectamine™ RNAiMAX or TransIT®-mRNA transfection reagent with a range of RNA doses. HEK293T cells were seeded in 96-well plates 24 hr prior to transfection. On the day of transfection, transfection reagent and RNA were mixed in 10 μl of Opti-MEM, and the transfection mixture was added to the medium of the seeded cells. Three days after transfection, genomic DNA was extracted and retrotransposition efficiency was measured using ddPCR with the same design as for DNA delivery.
[0819] The results of the ddPCR copy number analysis (normalized to the reference gene RPP30) are shown in Figure 12. For some transfection conditions, the average integration was measured to be 0.01 R2Tg copies per genome, significantly above the detection limit. These results demonstrate the successful integration of the R2Tg retrotransposon into human cells using the RNA delivery method.
[0820] Example 7: Targeted delivery of heterologous sequences of interest into mammalian cells using the R2Tg retrotransposon This example describes the delivery of transgenes into human cells by use of the R2Tg retrotransposon system with multiple delivery mechanisms, including RNA-mediated delivery of heterologous sequences of interest into human cells by use of the R2Tg retrotransposon system.
[0821] The R2 protein recognizes their template RNA structure within the untranslated region (UTR) of each element to form ribonucleoprotein particles that serve as intermediates for downstream integration into the host genome. Thus, the R2Tg machinery was used to engineer the decoupling of the UTR from its native environment and the integration of the UTR into another exogenous sequence to deliver the genome to the desired nucleic acid.
[0822] Trans-transgene integration was tested by constructing 1) the R2Tg coding sequence and 2) a transgene cassette flanked by R2Tg UTR sequences and 100 bp homology to 28S rDNA in separate driver and transgene plasmids, respectively. Figure 13 shows the dual plasmid system. The dual plasmid was introduced into HEK293T cells with FuGENE® HD transfection reagent at multiple driver:transgene molar ratios. In addition to the WT R2Tg driver, the backbone plasmid was used as a control. HEK293T cells were seeded at 10,000 cells / well in 96-well plates 24 hr prior to transfection. On the day of transfection, the transfection reagent and plasmid were mixed in 10 μl of Opti-MEM and incubated at room temperature for 15 min before being added to the medium of the seeded cells. Three days after transfection, genomic DNA was extracted for ddPCR assay to investigate the trans-retrotransposition efficiency. FIG. 14 shows the ddPCR results for conditions with an excess of transgene relative to the driver.
[0823] Similar to trans-transgene delivery by plasmid, RNA delivery was performed by constructing an amplicon of the coding sequence of R2Tg preceded by a T7 promoter sequence. The constructed amplicons contained the experimental R2Tg elements as well as a 1 bp deletion frameshift mutant control. Independently, an amplicon was constructed containing an exogenous sequence encoding GFP and an EGF1-α reporter with sufficient adjacent regions to drive integration into the genome by R2Tg. More specifically, the construct contained a T7 promoter driving transcription of an RNA that contained, from 5' to 3', the following: (a) a 10 nt long 5' 28S homology region, (b) a 5' untranslated region, (c) an antisense TKpA polyA sequence, (d) an antisense heterologous target sequence encoding GFP, (e) an antisense Kozak sequence, (f) an antisense EF1α promoter, (g) a 3' untranslated region that binds the GeneWriter protein, and (h) a 10 nt long 3' 28S homology region. Each RNA was transcribed using the New England Biolabs HiScribe T7 ARCA kit and purified using ZymoRNA clean and concentrator.
[0824] The resulting heterologous target RNA and R2TgRNA (either the experimental R2Tg element or the frameshift mutant) were introduced into human HEK293T cells at a 1:1 molar ratio by TransIT®-mRNA Transfection Kit. 24 hr before transfection, HEK293T cells were seeded in 96-well plates at 40,000 cells / well. On the day of transfection, 1 μl of transfection reagent and 500 ng of total RNA were mixed in 10 μl of Opti-MEM and incubated at room temperature for 5 min. The transfection mixture was then added to the medium of the seeded cells. Three days after transfection, genomic DNA was extracted for PCR assay.
[0825] Nested PCR was performed with an initial 30 rounds of PCR spanning the 3' end of the predicted transgene-rDNA junction, followed by 20 additional rounds of PCR amplification using internal primer sets. One out of three replicates of nested PCR performed on genomic DNA extracted from cells treated with the wild-type transposase reaction generated a PCR product of the expected size (approximately 596 bp). In contrast, no PCR product was observed in genomic DNA extracted from cells treated with the frameshift-inactivated R2Tg mutant control or the non-transfected control. The PCR products were gel purified using the Zero Blunt® TOPO® PCR Cloning Kit, and the resulting colonies were subjected to Sanger sequencing. Each individual PCR product sequence was then aligned to the predicted integration sequence. The fraction of PCR product sequences that aligned with the predicted integrated heterologous target sequence is shown in FIG. 10. The majority of the PCR products had the predicted integrants, as shown by the sequencing alignment adjacent to the predicted integration site on the right side of the alignment diagram. This demonstrates RNA-mediated integration of exogenous sequences into human cells by the R2Tg machinery.
[0826] Example 8: Targeted delivery of R2Tg retrotransposons into mammalian cells This example describes the targeted integration of the R2Tg retrotransposon element into mammalian cells by DNA delivery.
[0827] Plasmids carrying R2Tg(PLV014) and control plasmids were designed and synthesized as described in Example 6. Each plasmid was introduced into HEK393T cells by FuGENE® HD transfection reagent. 24 hr before transfection, HEK293T cells were seeded in 96-well plates at 10,000 cells / well. On the day of transfection, 0.5 μl of transfection reagent and 80 ng of DNA were mixed in 10 μl of Opti-MEM and incubated at room temperature for 15 minutes. The transfection mixture was then added to the medium of the seeded cells. Three days after transfection, genomic DNA was extracted for retrotransposition assay or frozen and subjected to target locus modification.
[0828] Target locus amplification was performed against the hg38 reference human genome and the rDNA locus sequence hsu13369 (GenBank: U13369.1). Target locus amplification was performed using two independent primer sets. Analysis with both primer sets revealed that the 28S rDNA locus sequence was the only integration site detected above the 1% threshold. Thus, integration of the R2Tg transposon in mammalian cells is specific for this target site.
[0829] Example 9: RNA refolding or driver:template RNA ratios for trans RNA-templated insertion into mammalian cells The RNA template is designed as in the previous example. Two RNAs consisting of a driver and a transgene payload are delivered to mammalian cells. To improve folding, denaturation of the payload RNA is performed by heating at 95°C and cooling at room temperature to promote proper secondary structure formation. In some embodiments, cooling the RNA at room temperature can increase recombination efficiency.
[0830] The molar ratio of transgene:driver is also varied to assess the preferred stoichiometry of the components. Integration is analyzed by ddPCR and sequencing. In some embodiments, a higher ratio of driver:transgene is used. In some embodiments, a higher ratio of transgene:driver is used.
[0831] Previous examples involving cis transgene integration are similarly tested for driver:payload stoichiometry. Integration is analyzed by ddPCR and sequencing. In some embodiments, higher ratios of driver transcription or translation:transgene transcription result in higher integration efficiencies. In some embodiments, higher ratios of transgene transcription:driver transcription and translation result in higher integration efficiencies.
[0832] Example 10: Hybrid capture assay Hybrid capture experiments were performed to perform an unbiased validation of the specificity of retrotransposon integration into the target site. Similar to the previous example, retrotransposon experiments were performed by introducing R2Tg flanked by its native UTR and 100 bp homology to one side of the predicted R2rDNA target. The rDNA target site had two flanking sets of 100 nucleotide identity to the corresponding native target site. Retrotransposons were delivered to human 293T cells via plasmid or mRNA. After 72 hours, genomic DNA was extracted. After extraction, each genomic DNA sample was subjected to hybrid capture according to a protocol (Twist) that included a custom probe set. Biotinylated probes were designed such that the approximately 120 bp probe spanned both strands of the R2Tg coding sequence and UTR. First, next generation libraries were generated by fragmentation of genomic DNA and ligation of sequencing adapters according to the protocol from Twist (available on the world wide web at: twistbioscience.com / ngs_protocol_custompanel_hybridcap). Probes were then hybridized to the genomic DNA library to amplify the enriched sample. The final library was sequenced on Miseq using 300bp paired-end reads. Reads were analyzed using a custom Matlab script. The resulting analysis results are shown in Figures 15A and 15B for RNA delivery. Hybrid capture showed on-target integration of R2Tg at the expected locus. For RNA delivery, one possible off-target with one read was identified at the non-expected 3' junction in the data, compared to over 100 reads at the expected locus, indicating a specificity of over 100:1. For the 5' junction, all 50 reads were at the expected locus, indicating a specificity of over 50:1. This experiment shows high integration specificity.
[0833] Example 11: Long read PacBio analysis Long-range PCR amplification can be performed to measure the integration of the desired full-length sequence into the target site of the human genome and to measure whether mutations are introduced during insertion. Retrotransposon integration experiments are performed as described in the previous example. In one example, PCR amplification is used to generate amplicons by designing one primer that targets the genomic integration site and one primer that targets the integrant sequence. In this example, these primers are designed to maximize the length of the amplified genomic locus fused with the integrant sequence. The amplicons spanning both ends of the integrant are pooled and long-read next-generation sequencing is performed to evaluate the fidelity of each integration.
[0834] In another embodiment, hybrid capture is performed as described in the previous embodiment, but with a larger target library length during the initial library generation. The resulting library is then subjected to long-read next-generation sequencing.
[0835] In some embodiments, long-read next-generation sequencing will show that there are less than 10%, 5%, 2%, 1%, 0.5%, 0.2%, or 0.1% SNPs in the integrated DNA across samples. In some embodiments, long-read next-generation sequencing will show that less than 10%, 5%, 2%, or 1% of the integrated DNA has a SNP. In some embodiments, long-read next-generation sequencing will show that less than 10%, 5%, 2%, or 1% of the integrated DNA has an internal deletion. In some embodiments, long-read next-generation sequencing will show that less than 10%, 5%, 2%, 1%, 0.5%, 0.2%, or 0.1% of the total integrated DNA across the population is deleted. In some embodiments, long-read next-generation sequencing will show that less than 10%, 5%, 2%, or 1% of the integrated DNA is truncated.
[0836] Example 12: Experiments with various homology lengths and point mutations of homology In this example, experiments are designed to characterize the preferred length and start position of homology for efficient retrotransposon integration into the target site, and the homology is used to support a mechanism of reverse transcription-driven integration.
[0837] A series of SNPs were introduced within 100 bp downstream of the homology of the R2Tg plasmid by modifying the plasmid PLV014. The design of the SNPs is shown in FIG. 16. After transfection, nested PCR was applied to recover the 3' integration junction site and generate a PCR product with an expected amplicon size of about 738 bp, and the PCR product was subjected to Sanger sequencing to confirm whether any SNPs were incorporated. In this experiment, the absence of SNP genetic markers incorporated in the junction sequence indicates that integration was driven by reverse transcription. The SNP design and sequencing results are shown in FIG. 16. SNP introduction was observed in 18 designed genetic markers, which is consistent with the integration of R2Tg directed by reverse transcription.
[0838] This example also describes the evaluation of various regions of homology to the target site to identify shorter regions that promote efficient integration into the genome. Two approaches are described in this example. First, various windows of 100 bp of homology to the target site are tested, starting from bp 1-100 3' of the target site, then bp 2-101 3' of the target site, then bp 3-103 3' of the target site, and so on, to bp 30-131 3' of the target site. Second, shorter lengths of homology to the target site sufficient for DNA integration are tested, starting from bp 0-100 3' of the target site, then bp 0-95 3' of the target site, then bp 0-90 3' of the target site, and so on, to bp 0-10 3' of the target site. Retrotransposition efficiency is measured using ddPCR after transfection of each plasmid into 293T cells.
[0839] In this example, different UTR regions of various lengths are evaluated to identify shorter sequences for efficient integration into the genome. The 3'UTR is examined by dividing this 325bp sequence into three regions, namely, 1-100bp, 101-200bp, and 201-325bp. Constructs of R2Tg containing each truncated 3'UTR were generated and the integration efficiency was examined for each.
[0840] Example 13: Assessing whether p53 or other repair pathways are upregulated This example describes the evaluation of the effect of exogenous R2Tg retrotransposition on gene expression, particularly tumor suppressor and DNA repair genes. R2Tg expressing plasmid is delivered to multiple cancer cell lines, including 293T, MCF-7, and T47D. After confirming integration into each cell line, RNA-seq is performed to evaluate the effect on gene expression profiles. Gene set enrichment analysis is then applied to evaluate whether any DNA repair pathways are upregulated after retrotransposition. MCF-7 and T47D are breast cancer cell lines that contain wild-type and mutant p53, respectively, and are used to further evaluate the relationship between p53 and retrotransposition. In some embodiments, p53 is not upregulated when the retrotransposon Gene Writer is integrated into the genome. In some embodiments, no DNA repair genes are upregulated when the retrotransposon Gene Writer is integrated into the genome. In some embodiments, no tumor suppressor genes are upregulated when the retrotransposon Gene Writer is integrated into the genome.
[0841] Example 14: Retrotransposition in the presence of DNA repair inhibitors In this example, the experiment tests the effect of various DNA repair pathways on R2Tg retrotransposition by applying DNA repair pathway inhibitors or DNA repair pathway-deficient cell lines. When applying DNA repair pathway inhibitors, a PrestoBlue cell viability assay is first performed to determine the toxicity of the inhibitor and whether any normalization should be applied to the subsequent assays. SCR7 is an inhibitor of NHEJ, which is applied in a series of dilutions during R2Tg delivery. PARP protein is a nuclear enzyme that binds to both single-strand and double-strand breaks as a homodimer. Therefore, the inhibitor is used to test related DNA repair pathways, such as the homologous recombination repair pathway and the base excision repair pathway. The experimental procedure is the same as that of SCR7. A cell line containing a defective core protein of the nucleotide excision repair (NER) pathway is used to test the effect of NER on R2Tg retrotransposition. After delivery of R2Tg elements to cells, ddPCR is used to evaluate retrotransposition in relation to the inhibition of DNA repair pathways. Also, sequencing analysis is carried out to evaluate whether specific DNA repair pathway plays a role in modifying integration junction.In some embodiments, the R2Tg integration into genome is not reduced by knocking down any DNA repair pathway, suggesting that R2Tg does not depend on host cell pathway for DNA integration.
[0842] Example 15: Retrotransposition in fibroblasts and T cells In this example, the previously performed R2Tg retrotransposition analysis of 293T cells is repeated in non-dividing cells (including fibroblasts and T cells). Compared to 293T cells, non-dividing cells can be difficult to transfect with lipid reagents. Therefore, nucleofection is used to deliver R2Tg elements. Subsequent retrotransposition assays for recombination efficiency and sequencing analysis are performed as described herein for 293T cells. In some embodiments, R2Tg is integrated into the genome of fibroblasts and T cells.
[0843] Example 16: Single cell ddPCR In this example, a quantitative assay can be used to determine the frequency of targeted genomic integration at the single cell level and the information compared to the copy number of targeted genomic integration per genome quantified from genomic DNA.
[0844] Approximately 5000 transfected cells are collected and mixed with a ddPCR reaction mixture, and then distributed into approximately 20,000 droplets, with each droplet containing one or no cells. A ddPCR assay, including 5'UTR and 3'UTR assay, is performed as described above to determine the frequency of R2 or transgene integration at the single cell level. At the same time, a control experiment is performed using genomic DNA taken from the same number of cells to determine the target genome integration efficiency per genome. In some embodiments, the frequency of target genome integration at the single cell level is calculated to be 1-80%, e.g., 25%, where the indicated percentage of cells have one or more copies of the transgene integrated into the desired locus.
[0845] Example 17: Single cell analysis by colony isolation In this example, a quantitative assay is used to determine genomic integration copy numbers in cell colonies derived from a single cell.
[0846] Single cell colonies are isolated by colony picking or limiting dilution and then cultured in a 96-well format. When the cells reach >80% confluency, half of the cells are frozen for reserve and genomic DNA from the remaining half of the cells is harvested for ddPCR. Optimized ddPCR assays, including 5'UTR and 3'UTR assays, are performed as described above to determine the frequency of R2 or transgene integration. At least 96 colonies are screened for each R2 element along with appropriate controls. The total number of colonies to be screened is determined by single cell ddPCR data, if applicable, or by the first set of single cell colony screen data. In some embodiments, the frequency of targeted genomic integration at the single cell level is calculated to be 1-80%, e.g., 25%, where the indicated percentage of cells have a single copy of the transgene integrated at the desired locus. This assay can also be used to determine the percentage of colonies with two or more copies of the transgene integrated at the desired locus.
[0847] Example 18: DNA Binding Affinity and / or Retargeting The DNA targeting module of wild-type R2 is composed of a cysteine-histidine zinc finger and a c-Myb transcription factor binding motif. This N-terminal module can be replaced with various DNA binding modules, such as DNA binding proteins (e.g., transcription factors), zinc fingers (e.g., natural or engineered motifs), and / or nucleic acid-guided, catalytically inactive endonucleases (e.g., Cas9 bound to a guide RNA (e.g., sgRNA) to form a Cas9-RNP). This DNA binding module is replaced with a naturally occurring module, and in some cases is placed with a flexible linker that connects it to the RNA binding / RT module. In addition, in some configurations, this novel DNA binding module is placed in tandem with the same and / or different DNA binding modules. Furthermore, in some configurations, the GeneWriter protein can be split, where one protein molecule contains the RNA binding module and the other protein contains the RT and endonuclease modules. In some embodiments, the exchange of DNA modules increases the specificity and / or affinity for certain genomic locations and, in some cases, allows for specific targeting of new genomic locations.
[0848] Example 19: Assay to measure DNA binding affinity The DNA binding activity (and similarly the DNA binding domains) of the GeneWriters described herein can be tested, for example, as described in this Example. The DNA binding modules are purified by expressing them recombinantly in cells (e.g., E. coli) or expressed in cell-free reactions of transcription and translation (e.g., T7 RNA polymerase + wheat germ extract). The purified DNA binding modules are tested for binding affinity by measuring Kd in a binding assay (e.g., EMSA, fluorescence anisotropy, dual filter binding, FRET, SPR, or thermophoresis (intensity change with temperature). The protein (DNA binding module) is labeled and / or the DNA module is tethered to a molecule compatible with the binding assays described above (e.g., dyes, radioisotopes (e.g., proteins:35 S-Methionine, Maleimide Dye, DNA: 32 The molecules are labeled with P terminal or internal label, DNA containing a bound amine reacted with NHS-ester dye. The concentration of these molecules is varied and measured by fitting a binding curve to calculate the binding affinity. In some assays, nucleic acid sequence specificity is examined by mutation analysis of DNA sequence or mutations to DNA binding module by amino acid changes or modifications to protein-nucleic acid complex (e.g., Cas9-RNP DNA binding module). In some embodiments, increasing the Kd of DNA binding module reduces off-target insertion and in some cases increases the activity of on-target site by increasing the residence time of R2-RNA complex at a specific genomic location.
[0849] Example 20: Assay to determine global specificity de novo The DNA-binding module is expressed in cells (e.g., animal cells, e.g., human cells) as the DNA-binding module alone, relative to the full-length retrotransposon R2, or a control without retrotransposase. The module or expression of the retrotransposon is delivered to the cells using conventional methods for delivering DNA, RNA, or proteins. The complex is cross-linked (e.g., with chemicals or UV light) or not. The cells are lysed and then treated with DNaseI so that only the bound DNA is protected from degradation. After DNA extraction and preparation of an NGS library of DNA fragments, new binding sites are identified, similar to ChIP-seq or DIG-seq. In some embodiments, potential off-target sites can be identified and tracked to eliminate false positives. In other embodiments, the assay confirms an in vitro assay for the specificity of the DNA-binding module binding to its intended site and not to other sites.
[0850] An orthogonal assay for identifying DNA-binding sites in a high-throughput manner uses the method described by Boyle et al., PNAS 2017, where DNA-binding domains are tested in a cell-free setting to determine specificity along with a systematic analysis of sequence variants for new DNA-binding modules.
[0851] Example 21: Modularity of RNA modules The RNA module binds to the R2 protein through interactions found in the reverse transcriptase module (submodule called "RNA binding"). The protein recognizes specific structures in the 5' and / or 3' UTR to interact with the RNA. In some embodiments, the exchange of UTR modules increases protein interactions, changes protein specificity binding to the UTR, stabilizes against nucleases, and / or improves cellular resistance (e.g., leading to a reduced innate immune response). In other embodiments, the addition and / or exchange of the RNA binding module of the R2 protein is compatible with the use of different sequences or ligands that link the element modules with transgenes and / or RNA. In some embodiments, the combination of new ligands instead of UTRs will result in better insertion efficiency because they have better affinity with the RNA binding domain of R2. In some embodiments, changes to the sequence of the UTR or changes to base modifications of the UTR increase the stability of the secondary structure, which leads to better interactions with the RNA binding module.
[0852] Example 22: Assays to measure RNA binding affinity to novel sequences The new UTR modules are tested in binding assays. In the case of new RNAs, they are synthesized by cell-free in vitro transcription using synthetic DNA templates, or by chemical synthesis of RNA in full length or fragments that are linked together to form a single RNA module. The binding affinity of the purified UTRs is measured by binding assays (e.g., EMSA, fluorescence anisotropy, dual filter binding, FRET, SPR, or thermophoresis (intensity change with temperature)). The UTR modules and / or RNA binding modules / RT modules are detected with or without labels as described above for labeling RNA. Measurements of various concentrations of molecules are performed to determine the binding affinity. In some embodiments, modifications to, replacement of, and / or changes to the 5' and / or 3' UTR binding modules and / or RNA binding / RT modules will achieve better interaction than wild-type R2 protein or UTR. In some embodiments, the increased interaction improves the efficiency of retrotransposition and, in some cases, increases the specificity of the R2 protein interacting with RNA.
[0853] Example 23: Alternative UTRs Without wishing to be bound by theory, in some embodiments, the UTR acts as a handle for the R2 protein to interact with an RNA, which it uses as a template for RT in conjunction with to bind to a genomic location, cleave the DNA with its endonuclease module, and then use the bound RNA as a template for RT insertion at the cleavage site in the DNA. In order for the UTR to hold the template in close proximity to the RT module, the UTR module can be replaced with different ligands, which bind to specific RNA binding sites engineered into the R2 protein. Thus, in some embodiments, the alternative non-RNA UTR is either a protein, small molecule, or other chemical, which is covalently attached via protein-protein interactions, small molecule-protein interactions, or hybridization. In some embodiments, the RNA binding module specifically binds to a non-RNA ligand that is bound to the transgene module RNA, which enhances the efficiency, stability, and / or rate of retrotransposition.
[0854] Example 24: Assays to measure activity of UTR constructs Binding assays measuring the affinity of the engineered UTR with the R2 protein are performed as described above, for example, for protein-nucleic acid interactions. In the case of protein-protein or protein-small molecule interactions, the assay uses a label on the RNA transgene module to which the UTR module is bound.
[0855] Example 25: Targeted genomic integration In this embodiment, GeneWriting technology is delivered to target cells and non-target cells, and new DNA is integrated into the genome in target cells at a higher frequency than in non-target cells. As described in more detail below, this approach utilizes non-target cells that have endogenous miRNAs that are not present (or have low levels) in target cells. Endogenous miRNAs are used to reduce DNA integration into non-target cells.
[0856] The polypeptide used is the R2Tg protein and the template RNA component is an RNA encoding the GFP protein, flanked at the 5' end by the 5'UTR of the R2Tg retrotransposase and at the 3' end by its 3'UTR. The 5'UTR is flanked by 100 bp of homology to the 5' side of the R2Tg 28s rDNA target site and the 3'UTR is flanked by 100 bp of homology to the 3' side of the R2Tg 28s rDNA target site. The GFP gene is oriented in the antisense direction with respect to the 5' and 3'UTRs and has its own promoter and polyadenylation signal.
[0857] The template RNA further comprises a microRNA recognition sequence that binds to a microRNA in a non-target cell, causing inhibition (e.g., degradation) of the template RNA prior to genome integration.
[0858] In this example, the target cells are hepatocytes and the non-target cells are macrophages from the hematopoietic system. The target cells and non-target cells are cultured separately. The template RNA and retrotransposase protein can be delivered to the cells as described herein, for example, as RNA or using a viral vector (e.g., an adeno-associated viral vector), and the template RNA is transcribed from the viral vector DNA.
[0859] Three days after treatment, cells are assayed for GFP expression and genomic integration.
[0860] GFP expression is assayed by flow cytometry, in some embodiments, GFP expression will be higher in the hepatocyte population than in the macrophage population.
[0861] Genomic integration (in terms of copy number per cell normalized to a reference gene) is assayed by droplet digital PCR using the methods described herein. In some embodiments, genomic integration will be higher in hepatocyte populations than in macrophage populations.
[0862] Example 26: Assay of DNA-binding domain modularity In this example, a series of experiments was performed to test the activity of various mutant retrotransposases and to obtain structural knowledge about these proteins. In this experiment, flexible linkers of various positions and lengths were tested in order to determine whether the DNA binding domain (DBD) is modular. These experiments also provide evidence for being able to separate the DBD from the rest of R2Tg and to replace it with any DNA targeting protein sequence. Thus, this example supports the understanding that the transposases described herein can tolerate test levels of sequence divergence at multiple positions identified by structural modeling (e.g., the predicted -1 RNA binding motif, the alpha helix, and the predicted c-myb DNA binding motif, as described below) while maintaining function.
[0863] Briefly, two linkers (Linker A: SGSETPGTSESATPES (SEQ ID NO: 1023), and Linker B: GGGS (SEQ ID NO: 1024) were inserted at three positions, denoted herein as versions v1, v2, and v3. v1 was located N-terminal to the α-helical region of R2Tg preceding the predicted -1 RNA binding motif, v2 was located C-terminal to the α-helical region of R2Tg preceding the predicted -1 RNA binding motif, and v3 was located C-terminal to the random coil region following the predicted c-myb DNA binding motif of R2Tg. For each of v1, v2, and v3, one of linkers A or B was added to the DNA plasmid expressing R2Tg by PCR, resulting in the sequences v1A (v1 + linker A), v1B (v1 + linker B), v1C (v1 + linker C), v2A (v2 + linker A), v2B (v2 + linker B), and v2C (v2 + linker C), as shown in Table 5 below. The insertion of the linkers was confirmed by Sanger sequencing, and the DNA plasmids were purified for transfection.
[0864] [Table 474]
[0865] [Table 475]
[0866] [Table 476]
[0867] [Table 477]
[0868] [Table 478]
[0869] [Table 479]
[0870] HEK293T cells were plated in 96-well plates and grown overnight at 37°C, 5% CO2. HEK293T cells were transfected with R2Tg (wild type), R2 endonuclease mutants, and linker mutants. Transfections were performed using Fugene HD transfection reagent according to the manufacturer's recommendations, with each well receiving 80ng of plasmid DNA and 0.5μL of transfection reagent. All transfections were performed in duplicate and incubated for 72 hours before genomic DNA extraction.
[0871] The activity of the mutants was measured by ddPCR assay, which quantified the copy number of R2Tg integration per genome. The 5' and 3' junctions were quantified by generating two different amplicons at each end.
[0872] v3 (near the c-myb binding motif in the DBD) had reduced interaction activity with either linker A or B. v1 (N-terminal to the α-helix preceding the -1 RNA-binding motif) had comparable activity to wild-type with linker A (16AA) and the shorter linker B (4AA). This may be related to amino acid selection, length, or 3D structure. v2 (C-terminal to the α-helix preceding the -1 RNA-binding motif) did not tolerate linker A; however, linker B had comparable and slightly better activity than wild-type. Thus, v1 and v2 can be considered as favorable positions for adding linkers that can separate the DNA-binding domain of R2Tg from the rest of the protein.
[0873] Example 27: Long-read sequencing to determine integration fidelity Retrotransposon integration experiments were carried out as described in the above examples. In one example, PCR amplification was used to generate amplicons by designing one primer targeting the genomic integration site and one primer targeting the integrant sequence. In this example, these primers were designed to maximize the length of the amplified genomic locus fused with the integrant sequence. The amplicons spanning both ends of the integrant were pooled and long-read next-generation sequencing was performed to evaluate the fidelity of each integration.
[0874] The cis construct of R2Tg was integrated into 293T cells via plasmid transfection as described herein. Amplicons spanning each end of the recombination were generated with flanking randomized UMIs to control for PCR bias. These amplicons were sequenced by PacBio next-generation sequencing. The resulting sequences were collapsed to remove reads with identical UMIs. By aligning the unique reads, coverage plots were generated, as shown in Figures 20A-20B. The sequence coverage mainly showed uniform coverage across the amplicons, indicating significant fidelity of integration. The relevant reverse transcriptase-deficient mutant control did not produce a signal. Internal deletions were also analyzed in Figures 21A-21B. Internal deletions were generally low relative to the total unique read counts and had some clustering at the 5' junction of rDNA-R2Tg.
[0875] In another example, hybrid capture may be performed as described in the previous example, except that the target library length during initial library construction is increased. The resulting libraries can then be subjected to long-read next generation sequencing.
[0876] Example 28: Targeted delivery of R2Gfo and R4Al retrotransposons into mammalian cells This example describes targeted integration of R2Gfo and R4Al retrotransposon elements by DNA delivery.
[0877] In one example, we tested the entire R2 element R2-1_GFo (Repbase; Kojima et al PLoS One 11, e0163496 (2015)) ("R2GFo") from a medium ground finch, e.g., the Galapagos finch (Geospiza fortis). In another example, we tested the entire R4 element R4_AL (Repbase; Burke et al Nucleic Acids Res. 23, 4628-34 (1995)) ("R4Al") from the roundworm (Ascaris lumbricoides). Since non-LTR R2 and R4 elements are absent in the human genome and are likely to be highly site-specific, the ability of a retrotransposon to precisely and efficiently integrate itself into the human genome indicates its ability to perform genome-targeted integration.
[0878] For cis integration of R2Gfo or R4Al elements, plasmids carrying R2Gfo (PLV033) or R4Al (PLV462) were synthesized as described in the previous example. The plasmids were synthesized such that the wild-type element was flanked by its native untranslated region (UTR) and 100 bp of homology to its rDNA target (Figure 22). Element expression was driven by a mammalian CMV promoter. We introduced each plasmid into HEK393T cells using FuGENE® HD transfection reagent. HEK293T cells were seeded in 96-well plates at 10,000 cells / well 24 hours before transfection. On the day of transfection, 0.5 μl of transfection reagent and 80 ng of DNA were mixed in 10 μl of Opti-MEM and incubated at room temperature for 15 minutes. The transfection mixture was then added to the medium of the seeded cells. Three days after transfection, genomic DNA was extracted for retrotransposition assay. In parallel, R2Tg was also delivered in the same format to serve as a comparison.
[0879] To confirm the integration and evaluate the integration efficiency, ddPCR was performed. Taqman probes were designed for the 3'UTR portion of each element. A forward primer was synthesized to bind directly upstream of the probe, and a reverse primer was synthesized to bind to the rDNA. Thus, amplification of the expected product across the integration junction degrades the probe and generates a fluorescent signal. The results of the ddPCR copy number analysis (relative to the reference gene RPP30) are shown in Figure 23. R2Gfo integration achieved an average copy number of 0.21 recombinants / genome in this experiment. R4Al achieved an average copy number of 0.085 recombinants / genome.
[0880] Example 29: Integration of retrotransposons into human fibroblasts This example describes the cis integration of R2Tg into human fibroblasts. Briefly, a plasmid designed to integrate R2Tg in cis was synthesized as described in the previous example, such that R2Tg is flanked by its native UTR and homologous sequences to its rDNA target. 0.5 μg of PLV014 (wild type) and PLV072 (EN mutant) plasmids were transfected into 100,000 human dermal fibroblasts isolated from neonatal foreskin (HDFn, C0045C, ThermoFisher Scientific), respectively, using the Neon transfection system. Two programs were performed, each with two replicates. The settings for program 1 were 1700V pulse voltage, 20ms pulse width, and 1 pulse number. The settings for program 2 were 1400V pulse voltage, 20ms pulse width, and 2 pulse numbers. Both programs achieved 95% transfection efficiency, as measured using a plasmid encoding EGFP. Three days after transfection, genomic DNA was extracted for ddPCR assay. ddPCR was performed to confirm integration and evaluate integration efficiency. A Taqman probe was designed for the 3'UTR portion of the R2Tg element. A forward primer was synthesized to bind directly upstream of the probe, and a reverse primer was synthesized to bind to the rDNA. In this way, amplification of the expected product across the integration junction degrades the probe and generates a fluorescent signal. The results of the ddPCR copy number analysis (relative to the reference gene RPP30) are shown in Figure 24. With wild-type (WT) R2Tg integration, an average copy number of 0.036 recombinants / genome was achieved in this experiment, which was significantly higher than the control R2Tg plasmid containing a point mutation that abolished endonuclease activity (EN).
[0881] Example 30: Evaluation of DNA damage response upon retrotransposon transfection DNA damage (e.g., caused by DSB formation or replication fork collapse) leads to activation of p53, which, among many other transcriptional responses, causes upregulation of p21, leading to cell cycle arrest or apoptosis. Genome editing with CRISRP / Cas9 has been shown to activate p53 and p21, a potential safety and efficacy issue for CRISPR / Cas9-based therapeutics. To prove whether R2Tg delivery into cells leads to activation of p53 and p21, U2OS cells were cultured at 4 × 10 4 U2OS cells were seeded at a density of 1000 cells / well and transfected 24 hours later with either 500 ng of R2Tg-WT plasmid or 500 ng of R2Tg-EN (a mutant of R2Tg that contains a mutation in the endonuclease (EN) domain that renders R2Tg inactive) using Fugene HD and Lipofectamine reagent. To control for transfection efficiency, U2OS cells were also transfected with a plasmid expressing GFP. Finally, as a positive control for p53 and p21, U2OS cells were treated with one of the DNA damage inducers etoposide (20 μM) or bleomycin (10 μg / ml). U2OS cells were harvested 24 hours after transfection / treatment. Protein lysates were prepared in RIPA buffer and run on SDS-PAGE gels, then transferred to nitrocellulose and probed with antibodies against p53 and p21, as well as actin and vinculin. As shown in FIG. 25, no R2Tg-induced upregulation of p53 or p21 over the GFP plasmid control was detected in any transfection condition.
Claims
1. A system for modifying DNA, comprising: (a) an RNA encoding a polypeptide, the polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain, each of (i) and (ii) being derived from a retrotransposase, wherein the polypeptide comprises an amino acid sequence of any one of SEQ ID NOs: 1016-1022, or an amino acid sequence having at least 90% identity to an amino acid sequence of any one of SEQ ID NOs: 1016-1022; and (b) a template RNA comprising (i) a sequence that binds to the polypeptide, and (ii) a heterologous sequence of interest. Including, wherein said RNA encoding said polypeptide does not contain said heterologous sequence of interest and / or said template RNA does not encode said polypeptide; and (a) and (b) are not part of the same nucleic acid; system.
2. (i) the template RNA further comprises a sequence comprising at least 20 nucleotides that are at least 80% identical to a target DNA; (ii) the target DNA is a genomic safe harbor (GSH) site; (iii) the template RNA comprises (iii) at least 3 or at least 10 bases having 100% sequence identity to the target DNA at the 3' end of the template RNA, and (iv) at least 3 or at least 10 bases having 100% sequence identity to the target DNA at the 5' end of the template RNA, and / or (iv) the sequence that binds to the polypeptide includes one or both of a 5' untranslated region and a 3' untranslated region; The system of claim 1 .
3. (i) the heterologous sequence of interest comprises a sequence encoding a polypeptide; (ii) the heterologous sequence of interest comprises a non-coding sequence or a regulatory sequence; (iii) the template RNA comprises a promoter operably linked to the heterologous sequence of interest; and / or (iv) the polypeptide further comprises one or both of a nuclear localization signal and a nucleolar localization signal; 3. A system according to claim 1 or 2.
4. The system of claim 3, wherein the polypeptide encoded by the heterologous sequence of interest is an enzyme, a membrane protein, a blood factor, an intracellular protein, an extracellular protein, a structural protein, a signaling protein, a regulatory protein, a transport protein, a motor protein, or a therapeutic protein, or a fragment thereof.
5. 5. The system of any one of claims 1 to 4, comprising only RNA or comprising more RNA than DNA, with an RNA:DNA ratio of at least 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:
1.
6. The system according to any one of claims 1 to 4, comprising only RNA, wherein the compositions (a) and (b) do not contain more than 1% DNA based on the mass or molar amount of nucleic acid.
7. The system according to any one of claims 1 to 6, wherein the reverse transcriptase activity can be used to modify DNA.
8. The system of claim 7, wherein the DNA can be modified in the absence of homologous recombination activity.
9. The system according to any one of claims 1 to 8, wherein the reverse transcriptase domain and the endonuclease domain of the polypeptide are linked by a linker.
10. The system according to any one of claims 1 to 9, disposed in a pharma- ceutically acceptable carrier.
11. The system of claim 10 , wherein the carrier comprises a lipid nanoparticle.
12. A system according to any one of claims 1 to 11 for use in a method for modifying a target DNA strand in a cell, tissue or subject.
13. 13. The system for use of claim 12, which achieves insertion of the heterologous sequence of interest into a target site within a genome with an average copy number of at least 0.01, 0.025, 0.05, 0.075, 0.1, 0.15, 0.2, 0.25, 0.3, 0.4, 0.5, 0.75, 1, 1.25, 1.5, 1.75, 2, 2.5, 3, 4, or 5 copies per genome.
14. 13. A system for use according to claim 12, which results in 25-100%, 50-100%, 60-100%, 70-100%, 75-95%, or 80-90% integration into the target site in the uncleaved genome.
15. The system according to any one of claims 1 to 14, wherein the polypeptide comprises an amino acid sequence of any one of SEQ ID NOs: 1016 to 1022, or an amino acid sequence having at least 98% identity to an amino acid sequence of any one of SEQ ID NOs: 1016 to 1022.
16. The system according to any one of claims 1 to 15, wherein the template RNA comprises a nucleic acid sequence of SEQ ID NO: 1140 or a nucleic acid sequence having at least 90% identity to the nucleic acid sequence of SEQ ID NO: 1140.
17. The system according to any one of claims 1 to 16, wherein the template RNA comprises a nucleic acid sequence of SEQ ID NO: 1263 or a nucleic acid sequence having at least 90% identity to the nucleic acid sequence of SEQ ID NO: 1263.
Citation Information
Patent Citations
Retrotransposon, DNA fragment having promoter activity, and its use
JP2002291473A
Synthetic mammalian retrotransposon gene
JP2007515941A
Chimeric endonuclease and its use
JP2013511978A
Method of line retro−position
WO2003064644A1
Tol1 FACTOR TRANSPOSASE AND DNA INTRODUCTION SYSTEM USING THE SAME
WO2008072540A1