Improved methods and compositions for modulating genome
The use of polypeptides with reverse transcriptase and endonuclease domains, combined with template RNAs, addresses the inefficiencies of existing genome editing methods by enabling precise and site-specific integration and modification of nucleic acid sequences in genomic DNA, facilitating therapeutic applications.
Patent Information
- Application Number
- JP2025141096
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-06-05
- Filing Date
- 2025-08-27
- Publication Date
- 2025-11-26
AI Technical Summary
Existing methods for integrating nucleic acid sequences into a genome lack site specificity and efficiency, particularly for long sequences, and require multiple steps like CRISPR/Cas9 for small edits or Cre/loxP for sequence insertion.
Compositions and systems featuring polypeptides with reverse transcriptase and endonuclease domains, along with template RNAs, facilitate precise insertion, deletion, or substitution of nucleotides in genomic DNA sequences, utilizing systems that include specific binding domains and heterologous sequences for targeted genome modifications.
Enables efficient and site-specific integration of exogenous genetic elements, allowing for precise modifications such as insertions and deletions of nucleotides in genomic DNA, including therapeutic polypeptides and non-coding RNAs, with high accuracy and specificity.
Smart Images

Figure 2025172847001319 
Figure 2025172847001320 
Figure 2025172847001321
Abstract
Description
[Technical Field]
[0001] Related Applications This application claims priority to U.S. Patent Application No. 62 / 985,264, filed March 4, 2020, and U.S. Patent Application No. 63 / 035,674, filed June 5, 2020, the entire contents of each of which are incorporated herein by reference. [Background technology]
[0002] The integration of a nucleic acid of interest into a genome occurs at low frequency and with little site specificity in the absence of special proteins to facilitate the insertion event. Some existing methods, such as CRISPR / Cas9, are more suitable for small edits and less effective for the integration of long sequences. Other existing methods, such as Cre / loxP, require a first step of inserting a loxP site into the genome, and then a second step of inserting a sequence of interest into the loxP site. The art needs improved proteins and methods for inserting a sequence of interest into a genome. Summary of the Invention [Means for solving the problem]
[0003] The present disclosure relates to novel compositions, systems, and methods for modifying the genome of one or more locations in a host cell, tissue, or subject, either in vivo or in vitro. In particular, the present invention features compositions, systems, and methods for the introduction of exogenous genetic elements into a host genome. The present disclosure also provides systems for modifying a genomic DNA sequence of interest, for example, by inserting, deleting, or substituting one or more nucleotides to / from the sequence of interest.
[0004] Features of the compositions or methods may include one or more of the embodiments listed below.
[0005] 1. (a) a polypeptide or a nucleic acid (e.g., DNA or mRNA) encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table Z1, or Table X) and (ii) an endonuclease domain, and one or both of (i) and (ii) have an amino acid sequence encoded by the nucleic acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising:
[0006] 2. (a) a polypeptide or a nucleic acid (e.g., DNA or mRNA) encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table Z1, or Table X) and (ii) an endonuclease domain, and one or both of (i) and (ii) have an amino acid sequence encoded by a nucleic acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides; and (b) A system for modifying DNA comprising a template RNA (or DNA encoding the template RNA) that comprises (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest.
[0007] 3. The system of any of the previous embodiments, wherein the heterologous sequence of interest encodes a therapeutic polypeptide or encodes a mammalian (e.g., human) polypeptide, or a fragment or variant thereof.
[0008] 4. The system of any of the previous embodiments, wherein the heterologous sequence of interest encodes a therapeutic non-coding RNA (e.g., miRNA).
[0009] 5. The system of any of the preceding embodiments, wherein the heterologous sequence of interest comprises, for example, a regulatory sequence (e.g., a promoter, an enhancer, a binding site for an endogenous regulatory component, e.g., an miRNA binding site) that alters expression of an endogenous gene or non-coding RNA.
[0010] 6. The system of any of the previous embodiments, wherein the regulatory sequence results in upregulation of an endogenous gene or non-coding RNA.
[0011] 7. The system of any of the previous embodiments, wherein the regulatory sequence results in downregulation of an endogenous gene or non-coding RNA.
[0012] 8. A system for modifying DNA, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, the polypeptide comprising: (i) a reverse transcriptase (RT) domain; (ii) a DNA binding domain (DBD); and (iii) an endonuclease domain, e.g., a nickase domain; and (b) a template RNA (or DNA encoding the template RNA) comprising (e.g., from 5' to 3'): (i) optionally, a sequence that binds to a target site (e.g., the second strand of a site in a target genome), (ii) optionally, a sequence that binds to a polypeptide, (iii) a heterologous sequence of interest, and (iv) a 3' target homology domain. Includes; (i) the polypeptide comprises a heterologous targeting domain (e.g., within the DBD or endonuclease domain) that specifically binds to a sequence contained within the target site; and / or (ii) The system, wherein the template RNA comprises a heterologous homologous sequence that is at least 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% homologous to a sequence contained within the target site.
[0013] 9. (a) a polypeptide or a nucleic acid (e.g., DNA or mRNA) encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table Z1, or Table X) and (ii) an endonuclease domain, and one or both of (i) and (ii) have the amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising:
[0014] 10. (a) a polypeptide or a nucleic acid (e.g., DNA or mRNA) encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, or Table Z1 or Table X) and (ii) an endonuclease domain, and one or both of (i) and (ii) have the amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising:
[0015] 11. (a) a polypeptide or a nucleic acid (e.g., DNA or mRNA) encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table Z1, or Table X) and (ii) a target DNA binding domain, and one or both of (i) and (ii) have an amino acid sequence encoded by the nucleic acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising:
[0016] 12. (a) a polypeptide or a nucleic acid (e.g., DNA or mRNA) encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table Z1, or Table X) and (ii) a target DNA binding domain, and one or both of (i) and (ii) have an amino acid sequence encoded by a nucleic acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising:
[0017] 13. (a) a polypeptide or a nucleic acid (e.g., DNA or mRNA) encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain from a protein other than a retrotransposase, e.g., from a retrovirus, e.g., a reverse transcriptase domain listed in Table Z1 or Z2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and (ii) an endonuclease domain and / or a target DNA binding domain; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising:
[0018] 14. (a) a polypeptide or a nucleic acid (e.g., DNA or mRNA) encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain from a protein other than a retrotransposase, e.g., from a retrovirus, e.g., a reverse transcriptase domain listed in Table Z1 or Z2, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, and (ii) an endonuclease domain and / or a target DNA binding domain; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; A system for modifying DNA comprising:
[0019] 15. (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table Z1 or Z2), and (ii) an endonuclease domain and / or a target DNA binding domain; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; Includes; A system for modifying DNA, wherein the sequence of the template RNA that binds to the polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to the 5' UTR or 3' UTR of the sequence of an element of Table 10 or Table X.
[0020] 16. (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table Z1 or Z2), and (ii) an endonuclease domain and / or a target DNA binding domain; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; Includes; A system for modifying DNA, wherein the sequence of the template RNA that binds to the polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to (i) the nucleotides located 5' to the start codon of the sequence of an element of Table 10 or Table X (e.g., containing a retrotransposase binding region), or (ii) the nucleotides located 3' to the stop codon of the sequence of an element of Table 10 or Table X (e.g., containing a retrotransposase binding region).
[0021] 17. (a) a polypeptide or a nucleic acid (e.g., DNA or mRNA) encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table Z1 or Z2, or Table X), and (ii) an endonuclease domain and / or a target DNA-binding domain; (b) a template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; and (c) Intein A system for modifying DNA comprising:
[0022] 18. The system of any of the previous embodiments, wherein the polypeptide comprises an intein.
[0023] 19. The system of any of the preceding embodiments, wherein the intein is a split intein.
[0024] 19. A system for modifying DNA, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA binding domain (DBD), and (iii) an endonuclease domain, e.g., a nickase domain; and (b) a template RNA (or DNA encoding the template RNA) comprising (e.g., from 5' to 3'): (i) optionally, a sequence (e.g., a CRISPR spacer) that binds to a target site (e.g., the non-edited strand of the site in the target genome), (ii) optionally, a sequence that binds to a polypeptide, (iii) a heterologous sequence of interest, and (iv) a 3' target homology domain. Including, the system.
[0025] 21. (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA binding domain (DBD), and (iii) an endonuclease domain, e.g., a nickase domain; and (b) a template RNA (or DNA encoding the template RNA) comprising (e.g., from 5' to 3'): (i) an optional sequence that binds to a target site (e.g., the non-edited strand of the site in the target genome); (ii) an optional sequence that binds to a polypeptide; (iii) a heterologous sequence of interest; and (iv) a 3' homology domain. Includes; A system for modifying DNA, wherein the RT domain has an amino acid sequence of Table 3B, Table 10, Table 11 or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0026] 22. (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA binding domain (DBD), and (iii) an endonuclease domain, e.g., a nickase domain; and (b) (e.g., from 5' to 3'): (i) an optional sequence that binds to a target site (e.g., the non-edited strand of the site in the target genome), (ii) an optional sequence that binds to a polypeptide, (iii) a heterologous sequence of interest, and (iv) a template RNA (etRNA) (or DNA encoding the template RNA) that includes a 3' homology domain. Includes; A system for modifying DNA that is capable of generating insertions of at least 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides into a target site.
[0027] 23. (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA binding domain (DBD), and (iii) an endonuclease domain, e.g., a nickase domain; and (b) a template RNA (or DNA encoding the template RNA) comprising (e.g., from 5' to 3'): (i) an optional sequence that binds to a target site (e.g., the non-edited strand of the site in the target genome); (ii) an optional sequence that binds to a polypeptide; (iii) a heterologous sequence of interest; and (iv) a 3' homology domain. Includes; 9. A system for modifying DNA, wherein the heterologous sequence of interest is at least 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 120, 140, 160, 180, 200, 500, or 1,000 nt in length.
[0028] 24. The system of any of the previous embodiments, wherein the RT domain is heterologous to the DBD; one or more DBDs are heterologous to the endonuclease domain; and / or one or more RT domains are heterologous to the endonuclease domain.
[0029] 25. (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA binding domain (DBD), and (iii) an endonuclease domain, e.g., a nickase domain; and (b) a template RNA (or DNA encoding the template RNA) comprising (e.g., from 5' to 3'): (i) an optional sequence that binds to a target site (e.g., the non-edited strand of the site in the target genome); (ii) an optional sequence that binds to a polypeptide; (iii) a heterologous sequence of interest; and (iv) a 3' homology domain. Includes; A system for modifying DNA that can create deletions at target sites of at least 81, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides.
[0030] 26. (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA binding domain (DBD), and (iii) an endonuclease domain, e.g., a nickase domain; and (b) a template (or DNA encoding a template RNA) comprising (e.g., from 5' to 3'): (i) an optional sequence that binds to a target site (e.g., the non-edited strand of the site in the target genome); (ii) an optional sequence that binds to a polypeptide; (iii) a heterologous sequence of interest; and (iv) a 3' homology domain. Includes; (a)(ii) and / or (a)(iii) is a system for modifying DNA comprising a TALE molecule; a zinc finger molecule; or a CRISPR / Cas molecule selected from Table 1 or a functional variant (e.g., mutant) thereof.
[0031] 27. (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA binding domain (DBD), and (iii) an endonuclease domain, e.g., a nickase domain; and (b) a template RNA (or DNA encoding the template RNA) comprising (e.g., from 5' to 3'): (i) an optional sequence that binds to a target site (e.g., a CRISPR spacer) (e.g., the non-edited strand of the site in the target genome); (ii) an optional sequence that binds to a polypeptide; (iii) a heterologous sequence of interest; and (iv) a 3' homology domain. Includes; A system for modifying DNA, wherein the endonuclease domain, e.g., the nickase domain, cleaves both strands of the target site DNA, and the cuts are separated from each other by at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 nucleotides.
[0032] 28. (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA binding domain (DBD), and (iii) an endonuclease domain, e.g., a nickase domain; and (b) a template RNA (or DNA encoding the template RNA) comprising (e.g., from 5' to 3'): (i) optionally, a sequence that binds to a target site (e.g., the non-edited strand of the site in the target genome); (ii) a sequence that specifically binds to an RT domain; (iii) a heterologous sequence of interest; and (iv) a 3' homology domain. A system for modifying DNA comprising:
[0033] 29. The system of any of the previous embodiments, wherein the template RNA further comprises a sequence that binds to (a)(ii) and / or (a)(iii).
[0034] 30. A system for modifying DNA, comprising: (a) a first polypeptide or a nucleic acid encoding a first polypeptide, the first polypeptide comprising (i) a reverse transcriptase (RT) domain, and (ii) optionally a DNA-binding domain; (b) a second polypeptide or a nucleic acid encoding a second polypeptide, the second polypeptide comprising: (i) a DNA-binding domain (DBD); (ii) an endonuclease domain, e.g., a nickase domain; and (c) a template RNA (or DNA encoding the template RNA) comprising (e.g., from 5' to 3'): (i) optionally, a sequence that binds to a second polypeptide (e.g., that binds to (b)(i) and / or (b)(ii)), (ii) optionally, a sequence that binds to a first polypeptide (e.g., that specifically binds to an RT domain), (iii) a heterologous sequence of interest, and (iv) a 3' homology domain. A system including:
[0035] 31. A system for modifying DNA, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises: (i) a reverse transcriptase (RT) domain, and (ii) a DNA binding domain (DBD); and (iii) an endonuclease domain, e.g., a nickase domain; (b) a first template RNA (or DNA encoding an RNA) comprising (e.g., from 5' to 3') (i) a sequence that binds to a polypeptide (e.g., that binds to (a)(ii) and / or (a)(iii)), and (ii) a sequence that binds to a target site (e.g., the non-edited strand of a site in a target genome) (e.g., where the first RNA comprises a gRNA); (c) a second template RNA (or DNA encoding the RNA) that includes (e.g., from 5' to 3'): (i) optionally, a sequence that binds to a polypeptide (e.g., that specifically binds to an RT domain), (ii) a heterologous sequence of interest, and (iii) a 3' target homology domain. A system including:
[0036] 32. The system of any of the previous embodiments, wherein the second template RNA comprises (i).
[0037] 33. The system of any of the previous embodiments, wherein the first template RNA comprises a first conjugate domain and the second template RNA comprises a second conjugate domain.
[0038] 34. The system of any of the previous embodiments, wherein the first and second conjugate domains are capable of hybridizing to one another, e.g., under stringent conditions.
[0039] 35. The system of any of the previous embodiments, wherein association of the first conjugation domain and the second conjugation domain co-localizes the first template RNA and the second template RNA.
[0040] 36. The system of any of the preceding embodiments, wherein the template RNA comprises (i).
[0041] 37. The system of any of the previous embodiments, wherein the template RNA comprises (ii).
[0042] 38. The system of any of the previous embodiments, wherein the template RNA comprises (i) and (ii).
[0043] 39. A system for modifying DNA, comprising: (a) a first polypeptide or a nucleic acid encoding a polypeptide, the polypeptide comprising a reverse transcriptase (RT) domain having the sequence of a reverse transcriptase domain of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and optionally a DNA binding domain (DBD) (e.g., a first DBD); and (b) a second polypeptide or a nucleic acid encoding a polypeptide, the polypeptide comprising: (i) a DBD (e.g., a second DBD); and (ii) an endonuclease domain, e.g., a nickase domain. A system including:
[0044] 40. The system of any of the previous embodiments, wherein the nucleic acid encoding the first polypeptide and the nucleic acid encoding the second polypeptide are two separate nucleic acids.
[0045] 41. The system of any of the previous embodiments, wherein the nucleic acid encoding the first polypeptide and the nucleic acid encoding the second polypeptide are part of the same nucleic acid molecule, e.g., present on the same vector.
[0046] 42. The system of any of the previous embodiments, having one or more (e.g., 1, 2, 3, 4, 5, 6, or all) of the following features: i. the heterologous sequence of interest encodes a protein, for example, an enzyme (e.g., a lysosomal enzyme) or a blood factor (e.g., Factor I, II, V, VII, X, XI, XII, or XIII); ii. the heterologous sequence of interest comprises a tissue-specific promoter or enhancer; iii. the heterologous sequence of interest encodes a polypeptide of more than 50, 100, 150, 200, 250, 300, 400, 500, or 1,000 amino acids, optionally up to 7,500 amino acids; iv. the heterologous sequence of interest encodes a fragment of a mammalian gene, but not the entire mammalian gene, e.g., encodes one or more exons, but not the full-length protein; v. The heterologous sequence of interest encodes one or more introns; vi. the heterologous object sequence is other than GFP, e.g., other than a fluorescent protein or other than a reporter protein; and vii. The heterologous object sequence includes only non-coding sequences, e.g., regulatory elements.
[0047] 43. The system of any of the previous embodiments, wherein the polypeptide further comprises a target DNA-binding domain, e.g., having an amino acid sequence encoded by the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0048] 44. The system of any of the previous embodiments, wherein the polypeptide further comprises a target DNA-binding domain, e.g., having an amino acid sequence encoded by the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides.
[0049] 45. The system of any of the previous embodiments, wherein the polypeptide further comprises a target DNA binding domain, e.g., having an amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0050] 46. The system of any of the previous embodiments, wherein the polypeptide further comprises a target DNA binding domain, e.g., having an amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides.
[0051] 47. The system of any of the previous embodiments, wherein the polypeptide has an activity at 37°C that is 70%, 75%, 80%, 85%, 90%, or 95% or more of its activity at 25°C under otherwise identical conditions.
[0052] 48. The system of any of the previous embodiments, wherein the polypeptide is derived from a warm-blooded organism, e.g., a bird or a mammal.
[0053] 49. The polypeptide is selected from the group consisting of CRE retrotransposase, NeSL retrotransposase, R4 retrotransposase, R2 retrotransposase, Hero retrotransposase, L1 retrotransposase, RTE retrotransposase, I retrotransposase, Jockey retrotransposase, CR1 retrotransposase, Rex1 retrotransposase, RandI / Dualen retrotransposase, Penelope or Penelope-like retrotransposase, Tx1 retrotransposase, RTEX retrotransposase, Crack retrotransposase, Nimb retrotransposase, and the like.
[0023] The system of any of the preceding embodiments, wherein the retrotransposase is derived from one or more of: Lo retrotransposase, Proto1 retrotransposase, Proto2 retrotransposase, RTETP retrotransposase, L2 retrotransposase, Tad1 retrotransposase, Loa retrotransposase, Ingi retrotransposase, Outcast retrotransposase, R1 retrotransposase, Daphne retrotransposase, L2A retrotransposase, L2B retrotransposase, Ambal retrotransposase, Vingi retrotransposase and / or Kiri retrotransposase.
[0054] 50. The system of any of the previous embodiments, wherein the polypeptide comprises an endonuclease domain from a transposable element, e.g., a restriction-like endonuclease (RLE), an apurinic / apyrimidinic endonuclease-like endonuclease (APE), a GIY-YIG endonuclease.
[0055] 51. The system of any of the preceding embodiments, wherein the endonuclease domain is intact.
[0056] 52. The system of any of the previous embodiments, wherein the endonuclease domain is inactivated.
[0057] 53. The system of any of the previous embodiments, wherein the endonuclease nicks the DNA.
[0058] 54. The system of any of the previous embodiments, wherein the endonuclease makes a double-strand break.
[0059] 55. The system of any of the previous embodiments, wherein the template RNA comprises a sequence of Table 3A or 3B or 10 (e.g., one or both of the 5' untranslated region in column 6 of Table 3A or 3B and the 3' untranslated region in column 7 of Table 3A or 3B), or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0060] 56. The system of any of the previous embodiments, wherein the template RNA comprises a sequence of Table 3A or 3B or 11 (e.g., one or both of the 5' untranslated region in column 6 of Table 3A or 3B and the 3' untranslated region in column 7 of Table 3A or 3B), or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides.
[0061] 57. The system of any of the previous embodiments, having one or more (e.g., 1, 2, 3, or all) of the following features: i. the nucleic acid encoding the polypeptide and the template RNA or the nucleic acid encoding the template RNA are separate nucleic acids; ii. the template RNA does not encode an active reverse transcriptase (including, for example, an inactivated mutant reverse transcriptase, e.g., as described in Examples 1-2) or does not contain a reverse transcriptase sequence; iii. the template RNA does not encode an active endonuclease (e.g., contains an inactivated endonuclease or no endonuclease); or iv. The template RNA comprises one or more chemical modifications.
[0062] 58. The template RNA (or DNA encoding the template RNA) comprises (i) a 5'UTR sequence linked to a polypeptide, (ii) a 3'UTR sequence linked to a polypeptide, (iii) a heterologous sequence of interest, and (iv) a promoter operably linked to the heterologous sequence of interest; the promoter is positioned between the 5' untranslated sequence and the heterologous sequence that binds the polypeptide; or 10. The system of any of the preceding embodiments, wherein the promoter is positioned between the 3' untranslated sequence that binds the polypeptide and the heterologous sequence.
[0063] 59. The template RNA (or DNA encoding the template RNA) comprises (i) a 5'UTR sequence that binds to the polypeptide, (ii) a 3'UTR sequence that binds to the polypeptide, and (iii) a heterologous sequence of interest; The system of any of the preceding embodiments, wherein the heterologous sequence of interest comprises an open reading frame (or its reverse complement) in the 5' to 3' direction of the template RNA; or wherein the heterologous sequence of interest comprises an open reading frame (or its reverse complement) in the 3' to 5' direction of the template RNA:
[0064] 60. The system of any of the previous embodiments, wherein the 5'UTR sequence of the template RNA (or DNA encoding the template RNA) has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the 5'UTR sequence of the sequence of an element of Table 3B, Table 10, or Table X.
[0065] 61. The system of any of the previous embodiments, wherein the 5' UTR sequence of the template RNA (or DNA encoding the template RNA) has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the nucleotides located 5' to the start codon of the sequence of an element of Table 3B, Table 10, or Table X (e.g., comprising a retrotransposase binding region).
[0066] 62. The system of any of the previous embodiments, wherein the 5'UTR sequence of the template RNA (or DNA encoding the template RNA) has substantial structural similarity (e.g., substantial secondary structure similarity) to the 5'UTR sequence of a sequence of an element of Table 3B, Table 10, or Table X.
[0067] 63. The system of any of the previous embodiments, wherein the 5'UTR sequence of the template RNA (or DNA encoding the template RNA) has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the 5'UTR sequence in column 6 of Table 3A or 3B or 10.
[0068] 64. The system of any of the previous embodiments, wherein the 3'UTR sequence of the template RNA (or DNA encoding the template RNA) has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the 3'UTR sequence of the sequence of an element of Table 3B, Table 10, or Table X.
[0069] 65. The system of any of the previous embodiments, wherein the 3'UTR sequence of the template RNA (or DNA encoding the template RNA) has substantial structural similarity (e.g., substantial secondary structure similarity) to the 3'UTR sequence of a sequence of an element of Table 3B, Table 10, or Table X.
[0070] 66. Any system of any of the previous embodiments, wherein the 3'UTR sequence of the template RNA (or DNA encoding the template RNA) has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the 3'UTR sequence in Table 10 or column 7 of Table 3A or 3B.
[0071] 67. Any system of any of the previous embodiments, wherein the 3'UTR sequence of the template RNA (or DNA encoding the template RNA) has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the nucleotides located 3' to the stop codon of the sequence of an element of Table 3B, Table 10, or Table X (e.g., comprising a retrotransposase binding region).
[0072] 68. The system of any of the previous embodiments, wherein the 3' UTR of the template RNA (or DNA encoding the template RNA) is flanked by homology domains, e.g., as described herein, e.g., homology domains having at least 5, 10, 20, 50, or 100 bases having at least 80% (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 100%) identity to the target DNA strand (wherein, optionally, the homology domain comprises a sequence according to a 3' homology arm of Table 11, or a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto).
[0073] 69. The system of any of the preceding embodiments, wherein the 5'UTR is flanked by homology domains, e.g., as described herein, e.g., homology domains having at least 10, 20, 50, or 100 bases having at least 80% (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 100%) identity to the target DNA strand (wherein, optionally, the homology domain comprises a sequence according to a 5' homology arm of Table 11, or a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto).
[0074] 70. The system of any of the previous embodiments, wherein at least one of the reverse transcriptase domain, the endonuclease domain, and the target DNA binding domain is heterologous, e.g., to the other domains.
[0075] 71. The system of any of the previous embodiments, wherein the endonuclease domain is heterologous to the reverse transcriptase domain and / or the target DNA binding domain.
[0076] 72. The system of any of the previous embodiments, wherein the polypeptide comprises (i) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to a reverse transcriptase domain of an apurinic / apyrimidinic endonuclease (APE)-type non-LTR retrotransposon, and (ii) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to an endonuclease domain of an APE-type non-LTR retrotransposon.
[0077] 73. The system of any of the previous embodiments, wherein the polypeptide comprises (i) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to a reverse transcriptase domain of a Penelope-like element (PLE)-type retrotransposon, and (ii) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to an endonuclease domain of a PLE-type retrotransposon, e.g., the PLE-type retrotransposon comprises a GIY-YIG endonuclease.
[0078] 74. The system of any of the previous embodiments, wherein the PLE-type retrotransposon does not comprise a functional endonuclease domain.
[0079] 75. The system of any of the previous embodiments, wherein the PLE-type retrotransposase comprises a Penelope-like element that naturally lacks an endonuclease domain (e.g., an Athena element as described in Gladyshev and Arkhipova PNAS 104, 9352-9357 (2007)).
[0080] 76. The system of any of the previous embodiments, wherein the PLE-type retrotransposase lacking a functional endonuclease domain is fused to Cas9, e.g., the PLE-type retrotransposase fused to Cas9 has DBD and / or endonuclease (e.g., nickase) function.
[0081] 77. The system of any of the previous embodiments, wherein the polypeptide comprises (i) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to a reverse transcriptase domain of a restriction enzyme-like endonuclease (RLE)-type non-LTR retrotransposon, (ii) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to an endonuclease domain of an RLE-type non-LTR retrotransposon, and (iii) a target DNA binding domain that is heterologous to (i) and / or (ii) (e.g., a heterologous zinc finger DNA binding domain).
[0082] 78. The system of any of the previous embodiments, wherein the polypeptide comprises (i) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to a reverse transcriptase domain of a restriction enzyme-like endonuclease (RLE)-type non-LTR retrotransposon, and (ii) a heterologous endonuclease domain and / or a heterologous DNA-binding domain.
[0083] 79. The system of any of the previous embodiments, wherein the polypeptide comprises (i) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to a reverse transcriptase domain of an apurinic / apyrimidinic endonuclease (APE)-type non-LTR retrotransposon, (ii) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to an endonuclease domain of an APE-type non-LTR retrotransposon, and (iii) a target DNA binding domain that is heterologous to (i) and / or (ii) (e.g., a heterologous zinc finger DNA binding domain)).
[0084] 80. The system of any of the previous embodiments, wherein the polypeptide comprises (i) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to a reverse transcriptase domain of an apurinic / apyrimidinic endonuclease (APE)-type non-LTR retrotransposon, and (ii) a heterologous endonuclease domain and / or a heterologous DNA-binding domain.
[0085] 81. The system of any of the previous embodiments, wherein the polypeptide comprises (i) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to a reverse transcriptase domain of a Penelope-like element (PLE)-type retrotransposon, (ii) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to an endonuclease domain of a PLE-type retrotransposon, and (iii) a target DNA binding domain that is heterologous to (i) and / or (ii) (e.g., a heterologous zinc finger DNA binding domain).
[0086] 82. The system of any of the previous embodiments, wherein the polypeptide comprises (i) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to a reverse transcriptase domain of a Penelope-like element (PLE)-type retrotransposon, and (ii) a heterologous endonuclease domain and / or a heterologous DNA-binding domain.
[0087] 83. The system of any of the previous embodiments, wherein the template RNA comprises (iii) a promoter operably linked to the heterologous sequence of interest.
[0088] 84. The system of any of the preceding embodiments, wherein the polypeptide further comprises (iii) a DNA-binding domain.
[0089] 85. The system of any of the previous embodiments, wherein the DNA binding domain has endonuclease activity.
[0090] 86. The system of any of the previous embodiments, wherein the endonuclease domain or endonuclease activity forms a double-stranded break in DNA.
[0091] 87. The system of any of the previous embodiments, wherein the endonuclease domain or endonuclease activity nicks DNA.
[0092] 88. The system of any of the previous embodiments, wherein the polypeptide comprises a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to a sequence in column 7 of Table 3A or 3B.
[0093] 89. The system of any of the previous embodiments, wherein the polypeptide comprises a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the sequence of an element listed in Table 10, Table 11, or Table X.
[0094] 90. The system of any of the previous embodiments, wherein the nucleic acid encoding the polypeptide and the template RNA or nucleic acid encoding the template RNA are covalently linked (e.g., are part of a fusion nucleic acid).
[0095] 91. The system of any of the previous embodiments, wherein the fusion nucleic acid comprises RNA.
[0096] 92. The system of any of the previous embodiments, wherein the fusion nucleic acid comprises DNA.
[0097] 93. The system of any of the preceding embodiments, wherein (b) comprises template RNA.
[0098] 94. The system of any of the previous embodiments, wherein the template RNA further comprises a nuclear localization signal.
[0099] 95. The system of any of the previous embodiments, wherein (a) comprises RNA encoding the polypeptide.
[0100] 96. The system of any of the previous embodiments, wherein the RNA of (a) and the RNA of (b) are separate RNA molecules.
[0101] 97. The system of any of the previous embodiments, wherein the RNA of (a) and the RNA of (b) are present in a ratio of 100:1 to 10:1, 10:1 to 5:1, 5:1 to 2:1, 2:1 to 1:1, 1:1 to 1:2, 1:2 to 1:5, 1:5 to 1:10, or 1:10 to 1:100.
[0102] 98. The system of any of the previous embodiments, wherein the RNA of (a) does not comprise a nuclear localization signal.
[0103] 99. The system of any of the previous embodiments, wherein the polypeptide further comprises a nuclear localization signal and / or a nucleolar localization signal.
[0104] 100. The system of any of the previous embodiments, wherein (a) comprises RNA encoding (i) a polypeptide and (ii) a nuclear localization signal and / or a nucleolar localization signal.
[0105] 101. The system of any of the previous embodiments, wherein the RNA comprises a pseudoknot sequence, for example, 5' to the heterologous sequence of interest.
[0106] 102. The system of any of the previous embodiments, wherein the RNA comprises a stem-loop sequence or a helix 5' to the pseudoknot sequence.
[0107] 103. The system of any of the preceding embodiments, wherein the RNA comprises one or more (e.g., two, three, or more) stem-loop sequences or helices 3' to the pseudoknot sequence, e.g., 3' to the pseudoknot sequence and 5' to the heterologous sequence of interest.
[0108] 104. The system of any of the previous embodiments, wherein the template RNA comprising the pseudoknot has catalytic activity, e.g., RNA cleavage activity, e.g., cis-RNA cleavage activity.
[0109] 105. The system of any of the previous embodiments, wherein the RNA comprises at least one stem-loop sequence, e.g., 1, 2, 3, 4, 5 or more stem-loop sequences, hairpin or helix sequences, e.g., 3' to the heterologous sequence of interest.
[0110] 106. The system of any of the previous embodiments, wherein the reverse transcriptase domain has an amino acid sequence of the reverse transcriptase domain of an element listed in Figure 10, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0111] 107. The system of any of the previous embodiments, wherein the endonuclease domain has the amino acid sequence of the endonuclease domain of an element listed in Figure 10, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0112] 108. The reverse transcriptase domain may be any of R2NS-1_CSi, R4-1_PH, RTE-1_MD, L2-18_ACar, L2-2_DRe, Dong-1_MMa, CR1_AC_1, Vingi-1_Acar, SR2, R4-1_AC, DongAa, CR1-1_PH, RTE-2_LMi, CR1-10_AMi, RTE-2_OL, L1-1_Cho, BovB, and L
[0023] The system of any of the preceding embodiments, having an amino acid sequence of the reverse transcriptase domain of ine1-2_ZM, RTE-1_Aip, CR1-3_IS, CR1-54_AAe, Tad1-65B_BG, TART-1_DWi or RTE_Ele2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0113] 109. The endonuclease domain may be any of the endonuclease domains provided herein: R2NS-1_CSi, R4-1_PH, RTE-1_MD, L2-18_ACar, L2-2_DRe, Dong-1_MMa, CR1_AC_1, Vingi-1_Acar, SR2, R4-1_AC, DongAa, CR1-1_PH, RTE-2_LMi, CR1-10_AMi, RTE-2_OL, L1-1_Cho, BovB, L
[0023] The system of any of the preceding embodiments, having an amino acid sequence of the endonuclease domain of ine1-2_ZM, RTE-1_Aip, CR1-3_IS, CR1-54_AAe, Tad1-65B_BG, TART-1_DWi or RTE_Ele2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0114] 110. The polypeptide may be any of R2NS-1_CSi, R4-1_PH, RTE-1_MD, L2-18_ACar, L2-2_DRe, Dong-1_MMa, CR1_AC_1, Vingi-1_Acar, SR2, R4-1_AC, DongAa, CR1-1_PH, RTE-2_LMi, CR1-10_AMi, RTE-2_OL, L1-1_Cho, Bo provided herein. vB, Line1-2_ZM, RTE-1_Aip, CR1-3_IS, CR1-54_AAe, Tad1-65B_BG, TART-1_DWi or RTE_Ele2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0115] 111. The sequence of the template RNA that binds to the polypeptide is the 5'UTR or any of the sequences provided herein: R2NS-1_CSi, R4-1_PH, RTE-1_MD, L2-18_ACar, L2-2_DRe, Dong-1_MMa, CR1_AC_1, Vingi-1_Acar, SR2, R4-1_AC, DongAa, CR1-1_PH, RTE-2_LMi, CR1-10_AMi, and RTE-2 100%, 95%, 96%, 97%, 98%, 99% or 100% identity to TART-1_OL, L1-1_Cho, BovB, Line1-2_ZM, RTE-1_Aip, CR1-3_IS, CR1-54_AAe, Tad1-65B_BG, TART-1_DWi or RTE_Ele2.
[0116] 112. The sequence of the template RNA that binds to the polypeptide is the 3'UTR or any of the sequences provided herein: R2NS-1_CSi, R4-1_PH, RTE-1_MD, L2-18_ACar, L2-2_DRe, Dong-1_MMa, CR1_AC_1, Vingi-1_Acar, SR2, R4-1_AC, DongAa, CR1-1_PH, RTE-2_LMi, CR1-10_AMi, and RTE-2 100%, 95%, 96%, 97%, 98%, 99% or 100% identity to TART-1_OL, L1-1_Cho, BovB, Line1-2_ZM, RTE-1_Aip, CR1-3_IS, CR1-54_AAe, Tad1-65B_BG, TART-1_DWi or RTE_Ele2.
[0117] 113. The system of any of the previous embodiments, wherein the nucleic acid encoding the polypeptide comprises a coding sequence that is codon-optimized for expression in a human cell.
[0118] 114. The system of any of the previous embodiments, wherein the template RNA comprises a coding sequence that is codon-optimized for expression in human cells.
[0119] 115. The system of any of the previous embodiments, comprising one or more circular RNA molecules (circRNAs).
[0120] 116. The system of any of the previous embodiments, wherein the circRNA encodes a Gene Writer polypeptide.
[0121] 117. The system of any of the previous embodiments, wherein the circRNA comprises a template RNA.
[0122] 118. The system of any of the preceding embodiments, wherein the circRNA is delivered to a host cell.
[0123] 119. The system of any of the previous embodiments, wherein the circRNA can be linearized, e.g., within the host cell, e.g., within the nucleus of the host cell.
[0124] 120. The system of any of the preceding embodiments, wherein the circRNA comprises a cleavage site.
[0125] 121. The system of any of the previous embodiments, wherein the circRNA further comprises a second cleavage site.
[0126] 122. The system of any of the previous embodiments, wherein the cleavage site can be cleaved (e.g., by self-cleavage) by a ribozyme, e.g., a ribozyme contained within the circRNA.
[0127] 123. The system of any of the previous embodiments, wherein the circRNA comprises a ribozyme sequence.
[0128] 124. The system of any of the previous embodiments, wherein the ribozyme sequence is capable of self-cleaving, e.g., within the host cell, e.g., in the nucleus of the host cell.
[0129] 125. The system of any of the previous embodiments, wherein the ribozyme is an inducible ribozyme.
[0130] 126. The system of any of the previous embodiments, wherein the ribozyme is a protein-responsive ribozyme, e.g., a ribozyme responsive to a nuclear protein, e.g., a genome-interacting protein, e.g., an epigenetic modifier, e.g., EZH2.
[0131] 127. The system of any of the previous embodiments, wherein the ribozyme is a nucleic acid-responsive ribozyme.
[0132] 128. The system of any of the previous embodiments, wherein the catalytic activity (e.g., autocatalytic activity) of the ribozyme is activated in the presence of a target nucleic acid molecule (e.g., an RNA molecule, such as an mRNA, miRNA, ncRNA, lncRNA, tRNA, snRNA, or mtRNA).
[0133] 129. The system of any of the previous embodiments, wherein the ribozyme is responsive to a target protein (e.g., MS2 coat protein).
[0134] 130. The system of any of the previous embodiments, wherein the target protein is localized in the cytoplasm or in the nucleus (e.g., an epigenetic modifier or a transcription factor).
[0135] 131. The system of any of the previous embodiments, wherein the ribozyme comprises a ribozyme sequence of a B2 or ALU retrotransposon, or a nucleic acid sequence having at least 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
[0136] 132. The system of any of the previous embodiments, wherein the ribozyme comprises a sequence of a tobacco ringspot virus hammerhead ribozyme, or a nucleic acid sequence having at least 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
[0137] 133. The system of any of the previous embodiments, wherein the ribozyme comprises a sequence of a hepatitis delta virus (HDV) ribozyme, or a nucleic acid sequence having at least 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
[0138] 134. The system of any of the previous embodiments, wherein the ribozyme is activated by a moiety expressed in a target cell or tissue.
[0139] 135. The system of any of the previous embodiments, wherein the ribozyme is activated by a moiety expressed within a target intracellular compartment (e.g., the nucleus, nucleolus, cytoplasm, or mitochondria).
[0140] 136. The system of any of the previous embodiments, wherein the ribozyme is comprised within a circular RNA or a linear RNA.
[0141] 137. A first circular RNA encoding a polypeptide of the Gene Writing System; A second circular RNA containing the template RNA of the Gene Writing System A system including:
[0142] 138. The system of any of the previous embodiments, wherein the template RNA, e.g., the 5'UTR, comprises a ribozyme that cleaves the template RNA (e.g., in the 5'UTR).
[0143] 139. The system of any of the above embodiments, wherein the template RNA comprises a ribozyme that is heterologous to (a)(i) (the reverse transcriptase domain), (a)(ii) (the endonuclease domain), (b)(i) (the sequence of the template RNA that binds to the polypeptide), or a combination thereof.
[0144] 140. The system of any of the previous embodiments, wherein the heterologous ribozyme can cleave the RNA that comprises the ribozyme, for example, 5' to the ribozyme, 3' to the ribozyme, or internally to the ribozyme.
[0145] 141. A lipid nanoparticle (LNP) comprising the system, polypeptide (or RNA encoding same), nucleic acid molecule, or DNA encoding the system or polypeptide of any of the previous embodiments.
[0146] 142. A first lipid nanoparticle comprising a polypeptide of a Gene Writing System (or DNA or RNA encoding the same) (e.g., as described herein); a second lipid nanoparticle containing a nucleic acid molecule of a Gene Writing system (e.g., as described herein); A system including:
[0147] 143. The system or reaction mixture of any of the previous embodiments, wherein the system, nucleic acid molecule, polypeptide, and / or DNA encoding same are formulated as lipid nanoparticles (LNPs).
[0148] 144. The LNP of any of the previous embodiments, comprising a cationic lipid.
[0149] 145. Cationic lipids have the following structure: [ka]
[0023] The LNP of any of the preceding embodiments, having:
[0150] 146. The LNP of any of the previous embodiments, further comprising one or more neutral lipids, e.g., DSPC, DPPC, DMPC, DOPC, POPC, DOPE, SM, a steroid, e.g., cholesterol, and / or one or more polymer-conjugated lipids, e.g., PEGylated lipids, e.g., PEG-DAG, PEG-PE, PEG-S-DAG, PEG-cer, or PEG dialkyloxypropylcarbamate.
[0151] 147. The system, kit, or polypeptide of any of the previous embodiments, wherein the system, polypeptide, and / or DNA encoding same is formulated as a lipid nanoparticle (LNP).
[0152] 148. The system, kit, or polypeptide of any of the previous embodiments, wherein the lipid nanoparticles (or formulation comprising a plurality of lipid nanoparticles) are devoid of reactive impurities (e.g., aldehydes) or contain reactive impurities (e.g., aldehydes) below a preselected level.
[0153] 149. The system, kit, or polypeptide of any of the previous embodiments, wherein the lipid nanoparticle (or formulation comprising a plurality of lipid nanoparticles) is devoid of aldehydes or contains aldehydes below a preselected level.
[0154] 150. The system, kit, or polypeptide of any of the previous embodiments, wherein the lipid nanoparticles are included in a formulation comprising a plurality of lipid nanoparticles.
[0155] 151. The system, kit, or polypeptide of any of the previous embodiments, wherein the lipid nanoparticle formulation is produced using one or more lipid reagents that constitute a total reactive impurity (e.g., aldehyde) content of less than 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or 0.1%.
[0156] 152. The system, kit, or polypeptide of any of the previous embodiments, wherein the lipid nanoparticle formulation is produced using one or more lipid reagents that constitute less than 3% total reactive impurity (e.g., aldehyde) content.
[0157] 153. The system, kit, or polypeptide of any of the previous embodiments, wherein the lipid nanoparticle formulation is produced using one or more lipid reagents that contain less than 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or 0.1% of any single reactive impurity (e.g., aldehyde) species.
[0158] 154. The system, kit, or polypeptide of any of the previous embodiments, wherein the lipid nanoparticle formulation is produced using one or more lipid reagents that contain less than 0.3% of any single reactive impurity (e.g., aldehyde) species.
[0159] 155. The system, kit, or polypeptide of any of the previous embodiments, wherein the lipid nanoparticle formulation is produced using one or more lipid reagents that contain less than 0.1% of any single reactive impurity (e.g., aldehyde) species.
[0160] 156. The system, kit, or polypeptide of any of the previous embodiments, wherein the lipid nanoparticle formulation comprises a total reactive impurity (e.g., aldehyde) content of less than 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or 0.1%.
[0161] 157. The system, kit, or polypeptide of any of the previous embodiments, wherein the lipid nanoparticle formulation comprises a total reactive impurity (e.g., aldehyde) content of less than 3%.
[0162] 158. The system, kit, or polypeptide of any of the previous embodiments, wherein the lipid nanoparticle formulation contains less than 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or 0.1% of any single reactive impurity (e.g., aldehyde) species.
[0163] 159. The system, kit, or polypeptide of any of the previous embodiments, wherein the lipid nanoparticle formulation contains less than 0.3% of any single reactive impurity (e.g., aldehyde) species.
[0164] 160. The system, kit, or polypeptide of any of the previous embodiments, wherein the lipid nanoparticle formulation contains less than 0.1% of any single reactive impurity (e.g., aldehyde) species.
[0165] 161. The system, kit, or polypeptide of any of the previous embodiments, wherein one or more or optionally all of the lipid reagents used in the lipid nanoparticles described herein or their formulations comprise a total reactive impurity (e.g., aldehyde) content of less than 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or 0.1%.
[0166] 162. The system or polypeptide of any of the previous embodiments, wherein one or more, or optionally all, of the lipid reagents used in the lipid nanoparticles described herein or their formulations comprise a total reactive impurity (e.g., aldehyde) content of less than 3%.
[0167] 163. The system, kit, or polypeptide of any of the previous embodiments, wherein one or more or optionally all of the lipid reagents used in the lipid nanoparticles described herein or formulations thereof contain less than 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or 0.1% of any single reactive impurity (e.g., aldehyde) species.
[0168] 164. The system, kit, or polypeptide of any of the previous embodiments, wherein one or more, or optionally all, of the lipid reagents used in the lipid nanoparticles described herein or their formulations contain less than 0.3% of any single reactive impurity (e.g., aldehyde) species.
[0169] 165. The system, kit, or polypeptide of any of the previous embodiments, wherein one or more, or optionally all, of the lipid reagents used in the lipid nanoparticles described herein or their formulations contain less than 0.1% of any single reactive impurity (e.g., aldehyde) species.
[0170] 166. The system, kit, or polypeptide of any of the previous embodiments, wherein the total aldehyde content and / or amount of any single reactive impurity (e.g., aldehyde) species is determined by liquid chromatography (LC), e.g., coupled with tandem mass spectrometry (MS / MS), e.g., according to the method described in Example 26.
[0171] 167. The system, kit, or polypeptide of any of the previous embodiments, wherein the total aldehyde content and / or amount of reactive impurity (e.g., aldehyde) species is determined by detecting one or more chemical modifications of a nucleic acid molecule (e.g., as described herein) that correlate with the presence of the reactive impurity (e.g., aldehyde), e.g., in the lipid reagent.
[0172] 168. The system, kit, or polypeptide of any of the previous embodiments, wherein the total aldehyde content and / or amount of aldehyde species is determined by detecting one or more chemical modifications of a nucleotide or nucleoside (e.g., a ribonucleotide or ribonucleoside, e.g., contained in or isolated from a nucleic acid molecule, e.g., as described herein), e.g., in a lipid reagent, that are associated with the presence of a reactive impurity (e.g., an aldehyde), e.g., as described in Example 41.
[0173] 169. The system, kit, or polypeptide of any of the previous embodiments, wherein chemical modifications of nucleic acid molecules, nucleotides, or nucleosides are detected by determining the presence of one or more modified nucleotides or nucleosides using, e.g., LC-MS / MS analysis, e.g., as described in Example 41.
[0174] 170. The polypeptide comprises a sequence of at least 50 amino acids (e.g., at least 100, 150, 200, 300, 500 amino acids) having at least 80% identity (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity) to the sequence of a polypeptide encoded by the sequence of an element of Table 10, Table 11, or Table X, or the reverse transcriptase domain or endonuclease domain thereof, of any of the above numbered systems.
[0175] 171. The polypeptide comprises a sequence of at least 50 amino acids (e.g., at least 100, 150, 200, 300, 500 amino acids) that differs by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids from the sequence of a polypeptide encoded by the sequence of an element of Table 10, Table 11, or Table X, or the reverse transcriptase domain or endonuclease domain thereof, of any of the above numbered systems.
[0176] 172. The polypeptide comprises a sequence of at least 50 amino acids (e.g., at least 100, 150, 200, 300, 500 amino acids) having at least 80% identity (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity) to the sequence of a polypeptide listed in Table 3A or 3B, or the reverse transcriptase domain, endonuclease domain, or DNA binding domain thereof, of any of the above numbered systems.
[0177] 173. The polypeptide comprises a sequence of at least 50 amino acids (e.g., at least 100, 150, 200, 300, 500 amino acids) that differs by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids from the sequence of a polypeptide listed in Table 3A or 3B, or the reverse transcriptase domain, endonuclease domain, or DNA binding domain thereof, of any of the above numbered systems.
[0178] 174. The polypeptide comprises a sequence of at least 50 amino acids (e.g., at least 100, 150, 200, 300, 500 amino acids) having at least 80% identity (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity) to an amino acid sequence in column 7 of Table 3A or 3B, or the reverse transcriptase domain, endonuclease domain, or DNA binding domain thereof, of any of the above numbered systems.
[0179] 175. The polypeptide comprises a sequence of at least 50 amino acids (e.g., at least 100, 150, 200, 300, 500 amino acids) that differs by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids from the amino acid sequence in column 7 of Table 3A or 3B, or the reverse transcriptase domain, endonuclease domain, or DNA binding domain thereof, of any of the above numbered systems.
[0180] 176. The template RNA comprises a sequence of an element of Table 10 or Table X (e.g., one or both of the 5' UTR of Table X or 10 and the 3' UTR of Table X or 10), or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, of any of the above numbered systems.
[0181] 177. The template RNA comprises a sequence of an element of Table 10 or Table X (e.g., one or both of the 5' UTR of Table X or 10 and the 3' UTR of Table X or 10), or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides.
[0182] 178. The template RNA comprises a sequence of Table 3A or 3B (e.g., one or both of the 5'UTR in column 5 of Table 3A or 3B and the 3'UTR in column 6 of Table 3A or 3B), or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, of any of the above-numbered systems.
[0183] 179. The template RNA comprises a sequence of Table 3A or 3B (e.g., one or both of the 5'UTR in column 5 of Table 3A or 3B and the 3'UTR in column 6 of Table 3A or 3B), or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides.
[0184] 180. The template RNA comprises (i) a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to the nucleotides located 5' to the start codon of the sequence of an element of Table 10 or Table X (e.g., including a retrotransposase binding region), any of the above numbered systems.
[0185] 181. The template RNA comprises (i) a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to the nucleotides located 3' to the stop codon of the sequence of an element of Table 10 or Table X (e.g., including a retrotransposase binding region), any of the above numbered systems.
[0186] 182. The system of any of the previous embodiments, wherein the template RNA comprises about 100-125 bp of a sequence from a 3'UTR of Table 10 or 11 or column 6 of Table 3A or 3B, e.g., the sequence comprises nucleotides 1-100, 101-200, or 201-325 of a 3'UTR of column 6 of Table 3A or 3B, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0187] 183. The system of any of the previous embodiments, wherein the template RNA comprises a sequence of about 100 to 125 bp from a 3'UTR of Table 10 or 11 or column 6 of Table 3A or 3B, e.g., the sequence comprises 1 to 100, 101 to 200, or 201 to 325 nucleotides of the 3'UTR of column 6 of Table 3A or 3B, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides.
[0188] 184.Any numbered system above, (a) containing RNA and (b) containing RNA.
[0189] 185. The above numbered system, which contains only RNA or contains more RNA than DNA in an RNA:DNA ratio of at least 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1 or 100:1.
[0190] 186. Any of the above numbered systems that contain no DNA or more than 10%, 5%, 4%, 3%, 2%, or 1% DNA by mass or molar amount.
[0191] 187.(b) A system according to any of the above numbers that allows DNA to be modified by inserting a heterologous sequence of interest without the intervention of DNA-dependent RNA polymerization.
[0192] 188. A system according to any of the above numbers, which allows DNA to be modified by inserting a heterologous sequence of interest via target-primed reverse transcription.
[0193] 189. Any of the above systems capable of modifying DNA by inserting a heterologous sequence of interest in the presence of an inhibitor of a DNA repair pathway (e.g., an SCR7, PARP inhibitor) or in a cell line lacking a DNA repair pathway (e.g., a cell line lacking a nucleotide excision repair pathway or a homologous recombination repair pathway).
[0194] 190. A system according to any of the above numbers that does not cause the formation of detectable levels of double-strand breaks in target cells.
[0195] 191. Any of the above systems, optionally in the absence of homologous recombination activity, that can modify DNA using reverse transcriptase activity.
[0196] 192. Any of the systems above, wherein the template RNA has been treated to reduce secondary structure (e.g., heated to a temperature that reduces secondary structure, e.g., at least 70, 75, 80, 85, 90, or 95°C).
[0197] 193. The system of any of the preceding embodiments, wherein the template RNA is subsequently cooled, e.g., to a temperature that allows secondary structure, e.g., below 37, 30, 25, or 20°C.
[0198] 194. The system of any of the previous embodiments, wherein the template RNA (or DNA encoding the template RNA) comprises a first homology domain at the 5' end of the template RNA, having at least 5 bases or at least 10 bases that are 100% identical to the target DNA strand, and a second homology domain at the 3' end of the template RNA, having at least 5 bases or at least 10 bases that are 100% identical to the target DNA strand.
[0199] 195. The system of any of the previous embodiments, wherein (a) and (b) are part of the same nucleic acid.
[0200] 196. The system of any of the previous embodiments, wherein (a) and (b) are distinct nucleic acids.
[0201] 197. The system of any of the previous embodiments, wherein the template RNA comprises at least 5 bases or at least 10 bases at the 5' end of the template RNA that are 100% identical to the target DNA strand (e.g., the target DNA strand is a human DNA sequence).
[0202] 198. The system of any of the previous embodiments, wherein the template RNA comprises at least 5 bases or at least 10 bases at the 3' end of the template RNA that are 100% identical to the target DNA strand (e.g., the target DNA strand is a human DNA sequence).
[0203] 199. The system of any of the previous embodiments, wherein the polypeptide comprises an active RNase H domain.
[0204] 200. The system of any of the previous embodiments, wherein the polypeptide does not comprise an active RNase H domain.
[0205] 201. The system of any of the previous embodiments, wherein the endogenous RNase H domain of the polypeptide is inactivated.
[0206] 202. The system of any of the previous embodiments, wherein the transposase polypeptide comprises a mutation that inactivates and / or deletes the nucleolar localization signal.
[0207] 203. The system of any of the previous embodiments, wherein the polypeptide does not comprise a functional nucleolar localization signal (e.g., does not comprise a nucleolar localization signal).
[0208] 204. The system of any of the previous embodiments, wherein the activity of the nucleolar localization signal is reduced by at least 50%, 60%, 70%, 80%, 90%, 95%, or 99%.
[0209] 205. The system of any of the previous embodiments, wherein the polypeptide comprises a nuclear localization signal (NLS), e.g., an endogenous NLS or an exogenous NLS.
[0210] 206. The polypeptide comprises (i) a first target DNA binding domain, e.g., comprising a first Zn finger domain, (ii) a reverse transcriptase domain, (iii) an endonuclease domain, and (iv) a second target DNA binding domain heterologous to the first target DNA binding domain, e.g., comprising a second Zn finger domain; The system of any of the preceding embodiments, wherein (a) binds to fewer target DNA sequences in the target cell than a similar polypeptide that comprises only the first target DNA binding domain, e.g., the presence of the second target DNA binding domain in the polypeptide along with the first DNA binding domain improves the target specificity of the polypeptide compared to the polypeptide target sequence specificity of a polypeptide that comprises only the first target DNA binding domain.
[0211] 207. The system of any of the preceding embodiments, wherein (iii) includes (iv).
[0212] 208. The system of any of the previous embodiments, wherein the second target DNA binding domain binds to a genomic DNA sequence that is less than 100, 90, 80, 70, 60, 50, 40, 30, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide away from the genomic sequence to which the first target DNA binding domain binds.
[0213] 209. The second target DNA binding domain is selected from the genome sequence to which the first target DNA binding domain binds, and the selected sequence ... 50, 60, 70, 80, 90, 100, 150, 20, 30, 40, 50, 60, 60, 70, 100, 150, 20, 30, 40, 5 ...
[0214] 210. The system of any of the previous embodiments, wherein the first or second target DNA binding domain comprises a CRISPR / Cas protein, a TAL effector domain, a Zn finger domain, or a meganuclease domain.
[0215] 211. The system of any of the previous embodiments, wherein the first strand of the target DNA can be cut at least twice (e.g., twice), and optionally the cuts are separated from each other by at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, or 200 nucleotides (and optionally no more than 500, 400, 300, 200, or 100 nucleotides)
[0216] 212. The system of any of the previous embodiments, which is capable of cleaving the first and second strands of the target DNA, and wherein the distance between the cleavages is the same as the distance between the cleavages made by the reverse transcriptase domain, e.g., when located in its endogenous polypeptide.
[0217] 213. Cutting is done from each other 1-500, 1-400, 1-300, 1-200, 1-100, 1-90, 1-80, 1-70, 1-60, 1-50, 1-40, 1-30, 1-20, 1-10, 1-5, 5-500, 5-400, 5-300, 5-200, 5-100, 5-90, 5-80, 5-70, 5-60, 5-50, 5-40, 5-30, 5-20, 5-10, 10-500, 10-400, 10-300, 10-200, 10-100, 1 0-90, 10-80, 10-70, 10-60, 10-50, 10-40, 10-30, 10-20, 20-500, 20-400, 20-300, 20-200, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, 20-30, 30-500, 30-400, 30-300, 30-200, 30-100, 30-90, 30-80, 30-70, 30-60, 30-50, 30-40, 40-500 , 40~400, 40~300, 40~200, 40~100, 40~90, 40~80, 40~70, 40~60, 40~50, 50~500, 50~400, 50~300, 50~200, 50~100, 50~90, 50~80, 50~70, 50~60, 60~500, 60~400, 60~300, 60~200, 60~100, 60~90, 60~80, 60~70, 70~500, 70~400, 70~300, 70~200, 70
[0033] The system of any of the preceding embodiments, wherein the nucleotides are separated by up to 100, 70 to 90, 70 to 80, 80 to 500, 80 to 400, 80 to 300, 80 to 200, 80 to 100, 80 to 90, 90 to 500, 90 to 400, 90 to 300, 90 to 200, 90 to 100, 100 to 500, 100 to 400, 100 to 300, 100 to 200, 200 to 500, 200 to 400, 200 to 300, 300 to 500, 300 to 400, or 400 to 500 nucleotides.
[0218] 214. The system of any of the previous embodiments, wherein the distance between cleavages is the same as the distance between cleavages made by the reverse transcriptase domain, e.g., the reverse transcriptase domain when located in its endogenous polypeptide.
[0219] 215. The system of any of the preceding embodiments, wherein both cleavages are made by the same endonuclease domain (e.g., a CRISPR / Cas protein (e.g., directed by multiple gRNAs, e.g., positioned in the template RNA)).
[0220] 216. The system of any of the previous embodiments, wherein the polypeptide further comprises a second endonuclease domain.
[0221] 217. i) a first endonuclease domain (e.g., a nickase) cleaves the strand of the target DNA that is to be edited, and a second endonuclease domain (e.g., a nickase) cleaves the strand of the target DNA that is not to be edited, or ii) The system of any of the preceding embodiments, wherein the first endonuclease domain (e.g., nickase) makes one of two cuts on the strand of the target DNA to be edited, and the second endonuclease domain (e.g., nickase) makes the other cut on the strand of the target DNA to be edited.
[0222] 218. The system of any of the previous embodiments, wherein (a), (b), or (a) and (b) further comprise a 5'UTR and / or a 3'UTR operably linked to the sequence encoding the polypeptide, the heterologous sequence of interest (e.g., a coding sequence contained in the heterologous sequence of interest), or both.
[0223] 219. The system of any of the previous embodiments, wherein the 5'UTR and / or 3'UTR increases expression of the operably linked sequence by at least 10%, 20%, 30%, 40%, 50%, 70%, 70%, 80%, 90% or 100% compared to the endogenous UTR associated with the heterologous sequence of interest or other similar nucleic acid comprising a minimal 5'UTR and a minimal 3'UTR.
[0224] 220. The system of any of the previous embodiments, wherein the template RNA (or DNA encoding the template RNA) comprises (i) a sequence that binds to a polypeptide, (ii) a heterologous sequence of interest, and (iii) a ribozyme that is heterologous to (a)(i), (a)(ii), (b)(i), or a combination thereof.
[0225] 221. The system of any of the previous embodiments, wherein (a), (b), or (a) and (b) comprise an intron that increases expression of the polypeptide, the heterologous sequence of interest (e.g., a coding sequence located in the heterologous sequence of interest), or both.
[0226] 222. A host cell (e.g., a mammalian cell, e.g., a human cell) comprising any of the systems described above.
[0227] 223. A method for modifying a target DNA strand in a cell, tissue, or subject, comprising administering to the cell, tissue, or subject a system according to any of the above numbers, wherein the system reverse-transcribes a template RNA sequence into a target DNA strand, thereby modifying the target DNA strand.
[0228] 224. The method of any of the previous embodiments, wherein the cell, tissue, or subject is a mammalian (e.g., human) cell, tissue, or subject.
[0229] 225. The method of any of the previous embodiments, wherein the tissue is liver, lung, skin, blood, immune, or muscle tissue.
[0230] 226. The method of any of the previous embodiments, wherein the cells are fibroblasts.
[0231] 227. The method of any of the previous embodiments, wherein the cells are primary cells.
[0232] 228. The method of any of the previous embodiments, wherein the cells are not immortalized.
[0233] 229. A method for modifying the genome of a mammalian cell, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; and one, two, or three of (i), (ii), and / or (iii) have an amino acid sequence encoded by the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; The method comprises contacting the
[0234] 230. A method for modifying the genome of a mammalian cell, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; and one, two, or three of (i), (ii), and / or (iii) have an amino acid sequence encoded by a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; The method comprises contacting the
[0235] 231. A method for modifying the genome of a mammalian cell, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; and one, two, or three of (i), (ii), and / or (iii) have the amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; The method comprises contacting the
[0236] 232. A method for modifying the genome of a mammalian cell, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; and one, two, or three of (i), (ii), and / or (iii) have the amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; The method comprises contacting the
[0237] 233. A method for modifying the genome of a mammalian cell, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; contacting the The method, wherein the sequence of the template RNA that binds to the polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to the 5' UTR sequence or the 3' UTR sequence of the sequence of an element of Table 3B, Table 10 or Table X.
[0238] 234. A method for modifying the genome of a mammalian cell, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; contacting the the sequence of the template RNA that binds to the polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to (i) the nucleotides located 5' to the start codon of the sequence of an element of Table 3B, Table 10 or Table X (e.g., containing a retrotransposase binding region), or (ii) the nucleotides located 3' to the stop codon of the sequence of an element of Table 3B, Table 10 or Table X (e.g., containing a retrotransposase binding region).
[0239] 235. A method for modifying the genome of a mammalian cell, comprising: (a) a polypeptide or a nucleic acid (e.g., DNA or mRNA) encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain from a protein other than a retrotransposase, e.g., from a retrovirus, e.g., a reverse transcriptase domain listed in Table Z1 or Z2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and (ii) an endonuclease domain; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; The method comprises contacting the
[0240] 236. A method for modifying the genome of a mammalian cell, comprising: (a) a polypeptide or a nucleic acid (e.g., DNA or mRNA) encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain from a protein other than a retrotransposase, e.g., from a retrovirus, e.g., a reverse transcriptase domain listed in Table Z1 or Z2, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, and (ii) an endonuclease domain; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; The method comprises contacting the
[0241] 237. A method for modifying the genome of a mammalian cell, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; (b) a template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; and (c) Intein The method comprises contacting the
[0242] 238. The method of any of the previous embodiments, wherein the polypeptide comprises an intein.
[0243] 239. The system of any of the preceding embodiments, wherein the intein is a split intein.
[0244] 240. The method of any of the previous embodiments, wherein the polypeptide does not comprise a target DNA binding domain.
[0245] 241. The method of any of the previous embodiments, wherein the polypeptide is derived from an APE-type retrotransposon reverse transcriptase.
[0246] 242. The method of any of the previous embodiments, wherein the polypeptide further comprises a target DNA-binding domain, e.g., having an amino acid sequence encoded by the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0247] 243. The method of any of the previous embodiments, wherein the polypeptide further comprises a target DNA-binding domain, e.g., having an amino acid sequence encoded by the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides.
[0248] 244. The method of any of the previous embodiments, wherein the polypeptide further comprises a target DNA binding domain, e.g., having an amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0249] 245. The method of any of the previous embodiments, wherein the polypeptide further comprises a target DNA binding domain, e.g., having an amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides.
[0250] 246. A method for modifying the genome of a mammalian cell, comprising: (a) an RNA encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; and one, two, or three of (i), (ii), and / or (iii) have an amino acid sequence encoded by the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (b) contacting with a template RNA comprising (i) a sequence that binds to the polypeptide and (ii) a heterologous sequence of interest; A method that does not include contacting a mammalian cell with DNA, or wherein the compositions of (a) and (b) do not contain more than 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% DNA by mass or molar amount of nucleic acid.
[0251] 247. A method for modifying the genome of a mammalian cell, comprising: (a) the RNA encoding the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; wherein one, two, or three of (i), (ii), and / or (iii) have an amino acid sequence encoded by the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides; and (b) contacting with a template RNA comprising (i) a sequence that binds to the polypeptide and (ii) a heterologous sequence of interest; A method that does not include contacting a mammalian cell with DNA, or wherein the compositions of (a) and (b) do not contain more than 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% DNA by mass or molar amount of nucleic acid.
[0252] 248. A method for modifying the genome of a mammalian cell, comprising: (a) RNA encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; and one, two, or three of (i), (ii), and / or (iii) have the amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (b) contacting with a template RNA comprising (i) a sequence that binds to the polypeptide and (ii) a heterologous sequence of interest; A method that does not include contacting a mammalian cell with DNA, or wherein the compositions of (a) and (b) do not contain more than 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% DNA by mass or molar amount of nucleic acid.
[0253] 249. A method for modifying the genome of a mammalian cell, comprising: (a) the RNA encoding the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; wherein one, two, or three of (i), (ii), and / or (iii) have the amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides; and (b) contacting with a template RNA comprising (i) a sequence that binds to the polypeptide and (ii) a heterologous sequence of interest; A method that does not include contacting a mammalian cell with DNA, or wherein the compositions of (a) and (b) do not contain more than 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% DNA by mass or molar amount of nucleic acid.
[0254] 250. The method of any of the previous embodiments, resulting in the addition of at least 1, 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of exogenous DNA sequence to the genome of the mammalian cell.
[0255] 251. The method of any of the previous embodiments, resulting in the deletion of at least 1, 2, 3, 4, 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of DNA sequence from the genome of the mammalian cell.
[0256] 252. The method of any of the previous embodiments, resulting in the modification of at least 1, 2, 3, 4, 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of DNA sequence from the genome of the mammalian cell.
[0257] 253. The method of any of the previous embodiments, resulting in the addition of a protein-coding sequence to the genome of a mammalian cell.
[0258] 254. The method of any of the previous embodiments, wherein the method results in a deletion of a protein-coding sequence in the genome of a mammalian cell.
[0259] 255. The method of any of the previous embodiments, wherein the method results in a modification of a protein-coding sequence in the genome of a mammalian cell.
[0260] 256. The method of any of the previous embodiments, resulting in the addition of a non-coding sequence, e.g., encoding a non-coding RNA, e.g., an miRNA, to the genome of a mammalian cell.
[0261] 257. The method of any of the previous embodiments, wherein the method results in the deletion of a non-coding sequence, e.g., encoding a non-coding RNA, e.g., an miRNA, in the genome of the mammalian cell.
[0262] 258. The method of any of the previous embodiments, resulting in a modification to the genome of a mammalian cell of a non-coding sequence, e.g., encoding a non-coding RNA, e.g., an miRNA.
[0263] 259. The method of any of the previous embodiments, which results in the addition of regulatory sequences, e.g., promoters, enhancers, miRNA binding sites, to the genome of a mammalian cell.
[0264] 260. The method of any of the previous embodiments, wherein the method results in the deletion of regulatory sequences, e.g., promoters, enhancers, miRNA binding sites, in the genome of a mammalian cell.
[0265] 261. The method of any of the previous embodiments, wherein the method results in modifications of regulatory sequences, e.g., promoters, enhancers, miRNA binding sites, in the genome of a mammalian cell.
[0266] 262. The method of any of the previous embodiments, wherein the addition, deletion or modification of a regulatory sequence to the genome of the mammalian cell results in increased expression of a coding or non-coding sequence in the genome of the mammalian cell.
[0267] 263. The method of any of the previous embodiments, wherein the addition, deletion or modification of a regulatory sequence to the genome of the mammalian cell results in a decrease in expression of a coding or non-coding sequence in the genome of the mammalian cell.
[0268] 264. A method for inserting DNA into the genome of a mammalian cell, comprising contacting the cell with an RNA composition, wherein the RNA composition comprises: (a) a first RNA that directs the insertion of a template RNA into the genome; and (b) template RNA containing a heterologous sequence Including, does not involve contacting mammalian cells with DNA, or the compositions of (a) and (b) do not contain more than 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% DNA by mass or molar amount of nucleic acid; resulting in the addition of at least 1, 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of DNA (e.g., exogenous DNA) sequence to the genome of the mammalian cell; and The method, wherein the first RNA encodes a polypeptide encoded by a sequence of an element of Table 3B, Table 10, Table 11 or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto, wherein the polypeptide directs insertion of the template RNA into the genome.
[0269] 265. A method for inserting DNA into the genome of a mammalian cell, comprising contacting the cell with an RNA composition, wherein the RNA composition comprises: (a) a first RNA that directs the insertion of a template RNA into the genome; and (b) template RNA containing a heterologous sequence Including, does not include contacting mammalian cells with DNA, or the compositions of (a) and (b) contain more than 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% DNA by mass or molar amount of nucleic acid; resulting in the addition of at least 1, 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of DNA (e.g., exogenous DNA) sequence to the genome of the mammalian cell; and The method, wherein the first RNA encodes a polypeptide encoded by a sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, wherein the polypeptide directs insertion of the template RNA into the genome.
[0270] 266. A method for inserting DNA into the genome of a mammalian cell, comprising contacting the cell with an RNA composition, wherein the RNA composition comprises: (a) a first RNA that directs the insertion of a template RNA into the genome; and (b) template RNA containing a heterologous sequence Including, does not involve contacting mammalian cells with DNA, or the compositions of (a) and (b) do not contain more than 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% DNA by mass or molar amount of nucleic acid; resulting in the addition of at least 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of DNA (e.g., exogenous DNA) sequence to the genome of the mammalian cell; The method, wherein the first RNA encodes a polypeptide of an element of Table 3B, Table 10, Table 11 or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto, wherein the polypeptide directs insertion of the template RNA into the genome.
[0271] 267. A method for inserting DNA into the genome of a mammalian cell, comprising contacting the cell with an RNA composition, wherein the RNA composition comprises: (a) a first RNA that directs the insertion of a template RNA into the genome; and (b) template RNA containing a heterologous sequence Including, does not involve contacting mammalian cells with DNA, or the compositions of (a) and (b) do not contain more than 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% DNA by mass or molar amount of nucleic acid; resulting in the addition of at least 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of DNA (e.g., exogenous DNA) sequence to the genome of the mammalian cell; and The method, wherein the first RNA encodes a polypeptide of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, wherein the polypeptide directs insertion of the template RNA into the genome.
[0272] 268. A method for inserting DNA into the genome of a mammalian cell, comprising contacting the cell with an RNA composition, wherein the RNA composition comprises: (a) a first RNA that directs the insertion of a template RNA into the genome; (b) a template RNA comprising a heterologous sequence and (i) a flanking sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to the 5' UTR sequence of the sequence of an element of Table 3B, Table 10 or Table X, and / or (ii) a flanking sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to the 3' UTR sequence of the sequence of an element of Table 3B, Table 10 or Table X. Including, does not involve contacting mammalian cells with DNA, or the compositions of (a) and (b) do not contain more than 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% DNA by mass or molar amount of nucleic acid; A method that results in the addition of at least 1, 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of a DNA (e.g., exogenous DNA) sequence to the genome of a mammalian cell.
[0273] 269. A method for inserting DNA into the genome of a mammalian cell, comprising contacting the cell with an RNA composition, wherein the RNA composition comprises: (a) a first RNA that directs the insertion of a template RNA into the genome; (b) a template RNA comprising a heterologous sequence and (i) a flanking sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to the nucleotides located 5' to the start codon of the sequence of an element of Table 3B, Table 10 or Table X (e.g., containing a retrotransposase binding region), and / or (ii) a flanking sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to the nucleotides located 3' to the stop codon of the sequence of an element of Table 3B, Table 10 or Table X (e.g., containing a retrotransposase binding region). Including, does not involve contacting mammalian cells with DNA, or the compositions of (a) and (b) do not contain more than 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% DNA by mass or molar amount of nucleic acid; A method that results in the addition of at least 1, 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of a DNA (e.g., exogenous DNA) sequence to the genome of a mammalian cell.
[0274] 270. The template RNA further comprises a sequence that binds to a polypeptide, and optionally, the sequence that binds to the polypeptide is (a) has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to the 5' UTR or 3' UTR of the sequence of an element of Table 3B, Table 10 or Table X; or The method of any of the preceding embodiments, wherein (b) the sequence of the template RNA that binds to the polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to (i) the nucleotides located 5' to the start codon of the sequence of an element of Table 3B, Table 10, or Table X (e.g., containing a retrotransposase binding region), or (ii) the nucleotides located 3' to the stop codon of the sequence of an element of Table X (e.g., containing a retrotransposase binding region).
[0275] 271. The method of any of the previous embodiments, wherein at least 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, 1000 bp of exogenous DNA is added to the genome of the mammalian cell without delivering the DNA to the cell.
[0276] 272. At least 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, or 1000 bp of exogenous DNA is added to the genome of a mammalian cell; The method of any of the previous embodiments, wherein the method does not include contacting the mammalian cell with DNA, or comprises contacting the mammalian cell with a composition having a DNA content of less than 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% by mass or molar amount of nucleic acid.
[0277] 273. The method of any of the previous embodiments, wherein at least 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, 1000 bp of exogenous DNA is added to the genome of the mammalian cell, and only RNA is delivered to the mammalian cell.
[0278] 274. The method of any of the previous embodiments, wherein at least at least 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, 1000 bp of exogenous DNA is added to the genome of the mammalian cell and RNA and protein are delivered to the mammalian cell.
[0279] 275. The method of any of the previous embodiments, wherein the template RNA serves as a template for insertion of exogenous DNA.
[0280] 276. The method of any of the previous embodiments, which does not include DNA-dependent RNA polymerization of exogenous DNA.
[0281] 277. The method of any of the previous embodiments, resulting in the addition of at least 1, 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of DNA to the genome of the mammalian cell.
[0282] 278. The method of any of the previous embodiments, wherein the RNA of (a) and the RNA of (b) are covalently linked, eg, are part of the same transcript.
[0283] 279. The method of any of the previous embodiments, wherein the RNA of (a) and the RNA of (b) are distinct RNAs.
[0284] 280. The method of any of the previous embodiments, which does not include contacting a mammalian cell with template DNA.
[0285] 281. A method for modifying the genome of a human cell, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; and one, two, or three of (i), (ii), and / or (iii) have an amino acid sequence encoded by the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; contacting the resulting in the insertion of a heterologous sequence of interest into the genome of a human cell, The method, wherein the human cells do not exhibit upregulation of DNA repair genes and / or tumor suppressor genes, or the DNA repair genes and / or tumor suppressor genes are not upregulated by more than 50%, 25%, 10%, 5%, 2%, or 1% (wherein, for example, the upregulation is measured by RNA-seq).
[0286] 282. A method for modifying the genome of a human cell, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; and one, two, or three of (i), (ii), and / or (iii) have an amino acid sequence encoded by the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; contacting the resulting in the insertion of a heterologous sequence of interest into the genome of a human cell, The method, wherein the human cells do not exhibit upregulation of DNA repair genes and / or tumor suppressor genes, or DNA repair genes and / or tumor suppressor genes are not upregulated by more than 50%, 25%, 10%, 5%, 2%, or 1% (wherein, for example, upregulation is measured by RNA-seq, as described in Example 14 of PCT / US2019 / 048607, which is incorporated herein by reference in its entirety).
[0287] 283. A method for modifying the genome of a human cell, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; and one, two, or three of (i), (ii), and / or (iii) have the amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; contacting the resulting in the insertion of a heterologous sequence of interest into the genome of a human cell, The method, wherein the human cells do not exhibit upregulation of DNA repair genes and / or tumor suppressor genes, or DNA repair genes and / or tumor suppressor genes are not upregulated by more than 50%, 25%, 10%, 5%, 2%, or 1% (wherein, for example, upregulation is measured by RNA-seq).
[0288] 284. A method for modifying the genome of a human cell, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; and one, two, or three of (i), (ii), and / or (iii) have the amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides; and (b) a template RNA (or DNA encoding the template RNA) containing (i) a sequence that binds to a polypeptide and (ii) a heterologous sequence of interest; contacting the resulting in the insertion of a heterologous sequence of interest into the genome of a human cell, The method, wherein the human cells do not exhibit upregulation of DNA repair genes and / or tumor suppressor genes, or DNA repair genes and / or tumor suppressor genes are not upregulated by more than 50%, 25%, 10%, 5%, 2%, or 1% (wherein, for example, upregulation is measured by RNA-seq).
[0289] 285. A method for modifying the genome of a human cell, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; and (b) (i) a sequence that binds to a polypeptide having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the 5' UTR sequence or the 3' UTR sequence of the sequence of an element of Table 3B, Table 10, Table 11, or Table X, and (ii) a template RNA (or DNA encoding the template RNA) that contains a heterologous sequence of interest. contacting the resulting in the insertion of a heterologous sequence of interest into the genome of a human cell, The method, wherein the human cells do not exhibit upregulation of DNA repair genes and / or tumor suppressor genes, or DNA repair genes and / or tumor suppressor genes are not upregulated by more than 50%, 25%, 10%, 5%, 2%, or 1% (wherein, for example, upregulation is measured by RNA-seq).
[0290] 286. A method for modifying the genome of a human cell, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; and (b) (i) a template RNA (or DNA encoding the template RNA) comprising a polypeptide binding sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to (i) a portion of the sequence of an element of Table 3B, Table 10 or Table X consisting of nucleotides located 5' to the start codon, or (ii) a portion of the sequence of an element of Table 3B, Table 10 or Table X consisting of nucleotides located 3' to the stop codon. contacting the resulting in the insertion of a heterologous sequence of interest into the genome of a human cell, The method, wherein the human cells do not exhibit upregulation of DNA repair genes and / or tumor suppressor genes, or DNA repair genes and / or tumor suppressor genes are not upregulated by more than 50%, 25%, 10%, 5%, 2%, or 1% (wherein, for example, upregulation is measured by RNA-seq).
[0291] 287. A method of adding an exogenous coding region to the genome of a cell (e.g., a mammalian cell), comprising contacting the cell with (i) RNA comprising the non-coding strand of the exogenous coding region (wherein, optionally, the RNA does not comprise the coding strand of the exogenous coding region), and (ii) a polypeptide comprising the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and optionally, wherein delivery comprises non-viral delivery.
[0292] 288. A method of adding an exogenous coding region to the genome of a cell (e.g., a mammalian cell), comprising contacting the cell with (i) RNA comprising the non-coding strand of the exogenous coding region (wherein, optionally, the RNA does not comprise the coding strand of the exogenous coding region), and (ii) a polypeptide comprising the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides; and optionally, delivery comprises non-viral delivery.
[0293] 289. A method for expressing a polypeptide of interest in a cell (e.g., a mammalian cell), comprising contacting the cell with (i) RNA, wherein the RNA comprises a non-coding strand that is the reverse complement of a sequence that would encode the polypeptide of interest, and optionally, the RNA does not comprise a coding sequence that encodes the polypeptide of interest, and (ii) a retrotransposase polypeptide comprising an amino acid sequence encoded by a sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and optionally, the delivery comprises non-viral delivery.
[0294] 290. A method for expressing a polypeptide of interest in a cell (e.g., a mammalian cell), comprising contacting the cell with (i) RNA, wherein the RNA comprises a non-coding strand that is the reverse complement of a sequence that would encode the polypeptide of interest, and optionally, the RNA does not comprise a coding sequence that encodes the polypeptide of interest, and (ii) a retrotransposase polypeptide comprising an amino acid sequence encoded by a sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides; and optionally, wherein delivery comprises non-viral delivery.
[0295] 291. A method for expressing a polypeptide of interest in a cell (e.g., a mammalian cell), comprising contacting the cell with (i) RNA, wherein the RNA comprises a non-coding strand that is the reverse complement of a sequence that would encode the polypeptide of interest, and optionally, the RNA does not comprise a coding sequence that encodes the polypeptide of interest, and (ii) a retrotransposase polypeptide comprising an amino acid sequence encoded by a sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and optionally, the delivery comprises non-viral delivery.
[0296] 292. A method for expressing a polypeptide of interest in a cell (e.g., a mammalian cell), comprising contacting the cell with (i) RNA, wherein the RNA comprises a non-coding strand that is the reverse complement of a sequence that would encode the polypeptide of interest, and optionally, the RNA does not comprise a coding sequence that encodes the polypeptide of interest, and (ii) a retrotransposase polypeptide comprising an amino acid sequence encoded by a sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides; and optionally, wherein delivery comprises non-viral delivery.
[0297] 293. The method of any of the previous embodiments, wherein the sequence to be inserted into the mammalian genome is a sequence that is exogenous to the mammalian genome.
[0298] 294. The method of any of the previous embodiments, wherein the exogenous sequence to be inserted into the mammalian genome does not naturally occur elsewhere in the mammalian genome.
[0299] 295. The method of any of the previous embodiments, wherein the exogenous sequence to be inserted into the mammalian genome is naturally occurring elsewhere in the mammalian genome.
[0300] 296. The method of any of the previous embodiments, which operates independently of a DNA template.
[0301] 297. The method of any of the previous embodiments, wherein the cells are part of a tissue.
[0302] 298. The method of any of the previous embodiments, wherein the mammalian cell is euploid, not immortalized, part of an organism, a primary cell, a non-dividing cell, a hepatocyte, or derived from a subject with a genetic disorder.
[0303] 299. The method of any of the previous embodiments, wherein the mammalian cells are present in a disease-free subject, e.g., to replenish the subject's genome.
[0304] 300. The method of any of the previous embodiments, wherein the contacting comprises contacting the cell with a plasmid, a virus, a virus-like particle, a virosome, a liposome, a vesicle, an exosome, a fusosome, or a lipid nanoparticle.
[0305] 301. The method of any of the previous embodiments, wherein the contacting includes the use of non-viral delivery.
[0306] 302. The method of any of the previous embodiments, comprising contacting a cell with a template RNA (or DNA encoding the template RNA), wherein the template RNA comprises the non-coding strand of the exogenous coding region, optionally wherein the template RNA does not comprise the coding strand of the exogenous coding region, and optionally wherein delivery comprises non-viral delivery, thereby adding the exogenous coding region to the genome of the cell.
[0307] 303. The method of any of the previous embodiments, comprising contacting a cell with a template RNA (or DNA encoding the template RNA), wherein the template RNA comprises a non-coding strand that is the reverse complement of a sequence encoding the polypeptide, optionally wherein the template RNA does not comprise a coding strand that encodes the polypeptide, and optionally wherein the delivery comprises non-viral delivery, thereby adding an exogenous coding region to the genome of the cell and expressing the polypeptide in the cell.
[0308] 304. The method of any of the previous embodiments, wherein the contacting comprises administering (a) and (b) to the subject, e.g., intravenously.
[0309] 305. The method of any of the previous embodiments, wherein the contacting comprises administering to the subject at least two doses of (a) and (b).
[0310] 306. The method of any of the previous embodiments, wherein the polypeptide reverse-transcribes a template RNA sequence into a target DNA strand, thereby modifying the target DNA strand.
[0311] 307. The method of any of the previous embodiments, wherein (a) and (b) are administered separately.
[0312] 308. The method of any of the previous embodiments, wherein (a) and (b) are administered together.
[0313] 309. The method of any of the previous embodiments, wherein the nucleic acid of (a) is not integrated into the genome of the host cell.
[0314] 310. The method of any of the previous embodiments, wherein the tissue is liver, lung, skin, muscle tissue (e.g., skeletal muscle), eye or ocular tissue, or central nervous system.
[0315] 311. The method of any of the previous embodiments, wherein the cell is a hematopoietic stem cell (HSC), a T cell, or a natural killer (NK) cell.
[0316] 312. The method of any of the above numbers, wherein the sequence binding to the polypeptide has one or more of the following characteristics: (a) is at the 3' end of the template RNA; (b) is at the 5' end of the template RNA; (b) is a non-coding sequence; (c) is a structured RNA; (d) forms at least one hairpin loop structure; and / or (e) is a guide RNA.
[0317] 313. The method of any of the above numbers, wherein the template RNA further comprises a sequence comprising at least 20 nucleotides that are at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the target DNA strand.
[0318] 314. The method of any of the above numbers, wherein the template RNA further comprises a sequence comprising at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150 nucleotides that are at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the target DNA strand.
[0319] 315. The method of any of the above numbers, wherein a sequence comprising at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150 nucleotides, or about 2-10, 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 10-100, or 2-100 nucleotides, which is at least 80% identical to the target DNA strand, is at the 3' end of the template RNA.
[0320] 316. The method of any of the above numbers, wherein the template RNA further comprises a sequence comprising at least 100 nucleotides that is at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the target DNA strand, e.g., at the 3' end of the template RNA.
[0321] 317. The method of any of the previous embodiments, wherein the site in the target DNA strand whose sequence comprises at least 80% identity is proximal (e.g., within about 0-10, 10-20, 20-30, 30-50, or 50-100 nucleotides) of the target site on the target DNA strand that is recognized (e.g., bound to and / or cleaved) by the polypeptide that comprises the endonuclease.
[0322] 318. The method of any number above, wherein the target RNA comprises a homology domain comprising a sequence according to the 3' homology arm of Table 11, or a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0323] 319. The method of any number above, wherein the target RNA comprises a homology domain comprising a sequence according to a 5' homology arm of Table 11, or a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0324] 320. A sequence comprising at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150 nucleotides, or about 2-10, 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 10-100, or 2-100 nucleotides, with at least 80% identity to the target DNA strand, is at the 3' end of the template RNA; Optionally, the method of any number above, wherein the site in the target DNA strand whose sequence comprises at least 80% identity is proximal (e.g., within about 0-10, 10-20, or 20-30 nucleotides) to the target site on the target DNA strand that is recognized (e.g., bound to and / or cleaved) by the polypeptide comprising the endonuclease.
[0325] 321. The method of any of the previous embodiments, wherein the target site is a site in the human genome that has closest identity to a natural target site for the polypeptide comprising the endonuclease, e.g., the target site in the human genome is identical to the natural target site by at least about 16, 17, 18, 19, or 20 nucleotides.
[0326] 322. The method of any of the above numbers, wherein the template RNA has at least 3, 4, 5, 6, 7, 8, 9, or 10 bases that are 100% identical to the target DNA strand.
[0327] 323. The method of any of the above numbers, wherein at least 3, 4, 5, 6, 7, 8, 9, or 10 bases of 100% identity to the target DNA strand are at the 3' end of the template RNA.
[0328] 324. The method of any of the above numbers, wherein at least 3, 4, 5, 6, 7, 8, 9, or 10 bases of 100% identity to the target DNA strand are at the 5' end of the template RNA.
[0329] 325. The method of any of the above numbers, wherein the template RNA comprises at least 3, 4, 5, 6, 7, 8, 9, or 10 bases of 100% identity to the target DNA strand at the 5' end of the template RNA, and at least 3, 4, 5, 6, 7, 8, 9, or 10 bases of 100% identity to the target DNA strand at the 3' end of the template RNA.
[0330] 326. The method of any of the above numbers, wherein the heterologous target sequence is 50 to 50,000 base pairs (e.g., 50 to 40,000 bp, 500 to 30,000 bp, 500 to 20,000 bp, 100 to 15,000 bp, 500 to 10,000 bp, 50 to 10,000 bp, 50 to 5,000 bp).
[0331] 327. The method of any of the above numbers, wherein the heterologous sequence of interest is at least 1, 10, 25, 50, 100, 150, 200, 250, 300, 400, 500, 600, or 700 bp.
[0332] 328. The method of any of the above numbers, wherein the heterologous sequence of interest is at least 715, 750, 800, 950, 1,000, 2,000, 3,000, or 4,000 bp.
[0333] 329. The method of any of the above numbers, wherein the heterologous sequence of interest is less than 5,000, 10,000, 15,000, 20,000, 30,000, or 40,000 bp.
[0334] 330. The method of any of the above numbers, wherein the heterologous sequence of interest is less than 700, 600, 500, 400, 300, 200, 150, or 100 bp.
[0335] 331. The heterologous target sequence is (a) an open reading frame, e.g., a sequence encoding a polypeptide, e.g., an enzyme (e.g., a lysosomal enzyme), a membrane protein, a blood factor, an exon, an intracellular protein (e.g., an organelle protein such as a cytoplasmic protein, a nuclear protein, a mitochondrial protein, or a lysosomal protein), an extracellular protein, a structural protein, a signaling protein, a regulatory protein, a transport protein, a sensory protein, a motor protein, a defense protein, a storage protein, or an immune receptor protein (e.g., a chimeric antigen receptor (CAR) protein, a T cell receptor, a B cell receptor), or an antibody; (b) non-coding and / or regulatory sequences, e.g., sequences that bind to transcriptional modulators, e.g., promoters, enhancers, insulators; (c) splice acceptor site; (d) polyA site; (e) the site of an epigenetic modification; or (f) Gene expression unit. Any of the above numbered methods, including one or more of the following:
[0336] 332. The method of any number above, wherein the target DNA is a genomic safe harbor (GSH) site.
[0337] 333. The method of any number above, wherein the target DNA is a genomic Natural Harbor™ site.
[0338] 334. The method of any number above, which results in the insertion of the heterologous sequence of interest into the genome at an average copy number of at least 0.01, 0.025, 0.05, 0.075, 0.1, 0.15, 0.2, 0.25, 0.3, 0.4, 0.5, 0.75, 1, 1.25, 1.5, 1.75, 2, 2.5, 3, 4 or 5 copies per genome.
[0339] 335. The method of any number above, which results in about 25-100%, 50-100%, 60-100%, 70-100%, 75-95%, 80%-90% integrants at the target site in the untruncated genome, as measured by an assay described herein, e.g., the assay of Example 6 of PCT Application No. PCT / US2019 / 048607.
[0340] 336. A method according to any of the above numbers, which results in the insertion of a heterologous sequence of interest at only one target site in the genome of the cell.
[0341] 337. Any of the methods described above, resulting in the insertion of a heterologous sequence of interest into a target site in a cell, wherein the inserted heterologous sequence contains less than 10%, 5%, 2%, 1%, 0.5%, 0.2%, or 0.1% mutations (e.g., SNPs, or one or more deletions, e.g., truncations or internal deletions) compared to the heterologous sequence prior to insertion, as measured, for example, by the assay of Example 12 of PCT Application No. PCT / US2019 / 048607.
[0342] 338. Any of the methods of numbers above, resulting in insertion of a heterologous sequence of interest into a target site in a plurality of cells, wherein less than 10%, 5%, 2%, or 1% of the inserted copies of the heterologous sequence contain a mutation (e.g., a SNP or a deletion, e.g., a truncation or internal deletion), e.g., as measured by the assay of Example 12 of PCT Application No. PCT / US2019 / 048607.
[0343] 339. Any of the methods of any of the above numbers, resulting in the insertion of a heterologous sequence of interest into the genome of a target cell, wherein the target cell does not exhibit p53 upregulation or exhibits less than 50%, 25%, 10%, 5%, 2%, or 1% upregulation of p53 (wherein, for example, p53 upregulation is measured by p53 protein levels or by levels of p53 phosphorylated at Ser15 and Ser20, e.g., by the methods described in Example 30).
[0344] 340. Any of the methods of the preceding numbers, resulting in the insertion of a heterologous sequence of interest into the genome of a target cell, wherein the target cell does not exhibit upregulation of DNA repair genes and / or tumor suppressor genes, or wherein DNA repair genes and / or tumor suppressor genes are not upregulated by more than 50%, 25%, 10%, 5%, 2%, or 1% (wherein, for example, upregulation is measured by RNA-seq).
[0345] 341. The method of any of the above numbers, wherein insertion of a heterologous sequence of interest into a target site (e.g., at a copy number of one insertion or two or more insertions) is achieved in about 1 to 80% of cells, e.g., about 1 to 10%, 10 to 20%, 20 to 30%, 30 to 40%, 40 to 50%, 50 to 60%, 60 to 70%, or 70 to 80% of cells in a cell population contacted with the system, as measured, e.g., using single-cell ddPCR, as described in Example 17.
[0346] 342. For example, any of the methods of the above numbers, wherein insertion of the heterologous sequence of interest into the target site occurs in about 1 to 80% of the cells in a cell population contacted with the system, e.g., about 1 to 10%, 10 to 20%, 20 to 30%, 30 to 40%, 40 to 50%, 50 to 60%, 60 to 70%, or 70 to 80% of the cells (e.g., at a copy number of one insertion), as measured, e.g., using colony isolation and ddPCR, as described in Example 18.
[0347] 343. Any of the methods described above, which result in a higher ratio of insertion of a heterologous sequence of interest into a target site (on-target insertion) than insertion into a non-target site (off-target insertion) in a cell population, e.g., a ratio of on-target insertion to off-target insertion of 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, 110:1, 120:1, 130:1, 140:1, 150:1, 160:1, 170:1, 180:1, 190:1, 200:1, 210:1, 220:1, 230:1, 240:1, 250:1, 260:1, 270:1, 280:1, 290:1, 300:1, 310:1, 320:1, 330:1, 340:1, 350:1, 360:1, 370:1, 380:1, 390:1, 400:1, 410:1, 420:1, 430:1, 440:1, 450:1, 460:1, 470:1, 480:1, 490:1, 500:1, 510:1, 520:1, 530:1, 540:1, 550:1, 560:1, 570:1, 580:1, 590:1, 600:1, 610:1, 620:1, 630:1, 640 The ratio is greater than 90:1, 100:1, 200:1, 500:1 or 1,000:1.
[0348] 344. Any of the above methods resulting in insertion of a heterologous sequence of interest in the presence of an inhibitor of a DNA repair pathway (e.g., SCR7, PARP inhibitor) or in a cell line lacking a DNA repair pathway (e.g., a cell line lacking a nucleotide excision repair pathway or a homologous recombination repair pathway).
[0349] 345. The method of any of the previous embodiments, wherein the cell has reduced Rad51 repair pathway activity, reduced expression of Rad51 or a component of the Rad51 repair pathway, or does not contain a functional Rad51 repair pathway, e.g., does not contain a functional Rad51 gene, e.g., contains a mutation (e.g., a deletion) that inactivates one or both copies of the Rad51 gene or another gene in the Rad51 repair pathway.
[0350] 346. A system of any of the above numbers formulated as a pharmaceutical composition.
[0351] 347. A system of any of the above numbers disposed in a pharmaceutically acceptable carrier (e.g., vesicles, liposomes, natural or synthetic lipid bilayers, lipid nanoparticles, exosomes).
[0352] 348. A method according to any of the above numbers, which results in the insertion of multiple (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, or 10) heterologous sequences of interest into the target cell genome.
[0353] 349. The method of any of the previous embodiments, wherein the multiple insertions occur simultaneously or sequentially.
[0354] 350. Any of the above numbered methods resulting in multiple (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, or 10) deletions of heterologous sequences of interest into the target cell genome.
[0355] 351. The method of any of the previous embodiments, wherein the multiple deletions occur simultaneously or consecutively.
[0356] 352. Any of the above methods, which results in multiple (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, or 10) base changes of a heterologous sequence of interest into the target cell genome.
[0357] 353. The method of any of the previous embodiments, wherein the multiple base changes occur simultaneously or sequentially.
[0358] 354. The method of any of the above numbers, comprising contacting a cell with a plurality of distinct template RNAs, each template RNA comprising a heterologous sequence of interest.
[0359] 355. The method of any of the previous embodiments, wherein the distinct template RNAs comprise at least two (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, or 10) distinct heterologous sequences of interest.
[0360] 356. The method of any of the previous embodiments, wherein the at least two distinct heterologous sequences of interest each comprise a distinct payload.
[0361] 357. The method of any of the previous embodiments, wherein the at least two distinct heterologous sequences of interest each comprise the same payload.
[0362] 358. The embodiment of any number above, wherein the target cells of the Gene Writing system have previously been modified at one or more loci.
[0363] 359. The embodiment of any number above, wherein the previously edited cell is a T cell.
[0364] 360. The embodiment of any number above, wherein the one or more prior modifications are selected from, e.g., a genetic knockout of an endogenous TCR (e.g., TRAC, TRBC), HLA class I (B2M), PD1, CD52, CTLA-4, TIM-3, LAG-3, or DGK.
[0365] 361. The embodiment of any number above, wherein the heterologous sequence of interest comprises a TCR or a CAR.
[0366] 362. A method for producing a system for modifying the genome of a mammalian cell, comprising: a) providing a template RNA comprising: (i) a sequence that binds to a polypeptide comprising a reverse transcriptase domain and an endonuclease domain, the sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the 5' UTR sequence or the 3' UTR sequence of the sequence of an element of Table X; and (ii) a heterologous sequence of interest. A method comprising:
[0367] 363.b) treating the template RNA to reduce secondary structure, e.g., heating the template RNA to, e.g., at least 70, 75, 80, 85, 90, or 95°C; and / or c) subsequently cooling the template RNA to, e.g., a temperature below 37, 30, 25, or 20°C that allows secondary structure. 10. The method of any of the preceding embodiments, further comprising:
[0368] 364. A method of making a system for modifying DNA (e.g., as described herein), comprising: (a) providing a template nucleic acid (e.g., an RNA or DNA template) that contains a heterologous sequence that is at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% homologous to a sequence contained within a target DNA molecule; and / or (b) providing a polypeptide of the system (e.g., comprising a DNA binding domain (DBD) and / or an endonuclease domain) that specifically binds to a sequence contained within a target DNA molecule; A method comprising:
[0369] 365. (a) including the introduction into a template nucleic acid (e.g., template RNA or DNA) of a heterologous sequence that is at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% homologous to a sequence contained within the target DNA molecule; and / or (b) introducing into a polypeptide of the system (e.g., a DNA binding domain (DBD) and / or an endonuclease domain) a heterologous targeting domain that specifically binds to a sequence contained within the target DNA molecule; 10. The method of any of the previous embodiments.
[0370] 366. The method of any of the previous embodiments, wherein the introducing of (a) comprises inserting a homologous sequence into the template nucleic acid.
[0371] 367. The method of any of the previous embodiments, wherein the introducing of (a) comprises replacing a segment of the template nucleic acid with a homologous sequence.
[0372] 368. The method of any of the previous embodiments, wherein the introducing of (a) comprises mutating one or more nucleotides (e.g., at least 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, or 100 nucleotides) of the template nucleic acid to generate a segment of the template nucleic acid having a homologous sequence.
[0373] 369. The method of any of the previous embodiments, wherein the introducing of (b) comprises inserting the amino acid sequence of the targeting domain into the amino acid sequence of the polypeptide.
[0374] 370. The method of any of the previous embodiments, wherein the introducing of (b) comprises inserting a nucleic acid sequence encoding a targeting domain into the coding sequence of a polypeptide contained within the nucleic acid molecule.
[0375] 371. The method of any of the previous embodiments, wherein the introducing of (b) comprises replacing at least a portion of the polypeptide with a targeting domain.
[0376] 372. The method of any of the previous embodiments, wherein the introduction of (a) comprises mutating one or more amino acids of the polypeptide (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500, or more amino acids).
[0377] 373. A method for modifying a target site in genomic DNA in a cell, comprising: Cells, (a) a polypeptide or a nucleic acid encoding a polypeptide (the polypeptide comprises: (i) a reverse transcriptase (RT) domain; (ii) a DNA binding domain (DBD); and (iii) an endonuclease domain, e.g., a nickase domain); and (b) a template RNA (or DNA encoding the template RNA) comprising (e.g., from 5' to 3'): (i) optionally, a sequence that binds to a target site (e.g., the second strand of a site in a target genome), (ii) optionally, a sequence that binds to a polypeptide, (iii) a heterologous sequence of interest, and (iv) a 3' target homology domain. (i) the polypeptide comprises a heterologous targeting domain (e.g., within the DBD or endonuclease domain) that specifically binds to a sequence contained within or adjacent to a target site in genomic DNA; and / or (ii) the template RNA contains a heterologous sequence that is at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% homologous to a sequence contained within or adjacent to the target site in the genomic DNA. modifying a target site in genomic DNA in a cell by contacting the cell with
[0378] 374. A method for producing a system for modifying the genome of a mammalian cell, comprising: a) providing a template RNA comprising: (i) a sequence that binds to a polypeptide comprising a reverse transcriptase domain and an endonuclease domain, the sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to (i) the nucleotides located 5' to the start codon of the sequence of an element of Table 3B, Table 10 or Table X (e.g., containing a retrotransposase binding region), or (ii) the nucleotides located 3' to the stop codon of the sequence of an element of Table 3B, Table 10 or Table X (e.g., containing a retrotransposase binding region); A method including
[0379] 375. b) treating the template RNA to reduce secondary structure, e.g., heating the template RNA to, e.g., at least 70, 75, 80, 85, 90, or 95°C; and / or c) subsequently cooling the template RNA to a temperature that allows for secondary structure, e.g., below 37, 30, 25, or 20°C; 10. The method of any of the preceding embodiments, further comprising:
[0380] 376. The method of any of the previous embodiments, wherein the system is the system of any of the previous embodiments.
[0381] 377. The method of any of the previous embodiments, further comprising contacting the template RNA with a polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain, or with a nucleic acid (e.g., RNA) encoding the polypeptide.
[0382] 378. The method of any of the preceding embodiments, further comprising contacting the template RNA with the cell.
[0383] 379. The system or method of any of the previous embodiments, wherein the heterologous sequence of interest encodes a therapeutic polypeptide.
[0384] 380. The system or method of any of the previous embodiments, wherein the heterologous sequence of interest encodes a mammalian (e.g., human) polypeptide, or a fragment or variant thereof.
[0385] 381. The system or method of any of the previous embodiments, wherein the heterologous sequence of interest encodes an enzyme (e.g., a lysosomal enzyme), a blood factor (e.g., Factor I, II, V, VII, X, XI, XII, or XIII), a membrane protein, an exon, an intracellular protein (e.g., an organelle protein such as a cytoplasmic protein, a nuclear protein, a mitochondrial protein, or a lysosomal protein), an extracellular protein, a structural protein, a signaling protein, a regulatory protein, a transport protein, a sensory protein, a motor protein, a defense protein, a storage protein, an immune receptor protein (e.g., a chimeric antigen receptor (CAR) protein, a T cell receptor, a B cell receptor), or an antibody.
[0386] 382. The system or method of any of the previous embodiments, wherein the heterologous sequence of interest comprises a tissue-specific promoter or enhancer.
[0387] 383. The system or method of any of the previous embodiments, wherein the heterologous sequence of interest encodes a polypeptide of more than 250, 300, 400, 500, or 1,000 amino acids, and optionally up to 1300 amino acids.
[0388] 384. The system or method of any of the previous embodiments, wherein the heterologous sequence of interest encodes a fragment of a mammalian gene, but not the entire mammalian gene, e.g., encoding one or more exons, but not the full-length protein.
[0389] 385. The system or method of any of the previous embodiments, wherein the heterologous sequence of interest encodes one or more introns.
[0390] 386. The system or method of any of the previous embodiments, wherein the heterologous sequence of interest is other than GFP, e.g., other than a fluorescent protein or other than a reporter protein.
[0391] 387. The system or method of any of the previous embodiments, wherein the polypeptide has an activity at 37°C that is 70%, 75%, 80%, 85%, 90%, or 95% or more of its activity at 25°C under otherwise identical conditions.
[0392] 388. The system or method of any of the previous embodiments, wherein the nucleic acid encoding the polypeptide and the template RNA or nucleic acid encoding the template RNA are separate nucleic acids.
[0393] 389. The system or method of any of the previous embodiments, wherein the template RNA does not encode an active reverse transcriptase (e.g., comprises an inactivated mutant reverse transcriptase, e.g., as described in Example 1 or 2, or does not contain a reverse transcriptase sequence).
[0394] 390. The system or method of any of the previous embodiments, wherein the template RNA comprises one or more chemical modifications.
[0395] 391. The system or method of any of the previous embodiments, wherein the heterologous sequence of interest is positioned between the promoter and the sequence that binds to the polypeptide.
[0396] 392. The system or method of any of the previous embodiments, wherein the heterologous sequence of interest is positioned between the promoter and the sequence that binds to the polypeptide.
[0397] 393. The system or method of any of the previous embodiments, wherein the heterologous sequence of interest comprises an open reading frame (or its reverse complement) in the 5' to 3' direction of the template RNA.
[0398] 394. The system or method of any of the previous embodiments, wherein the heterologous sequence of interest comprises an open reading frame (or its reverse complement) in the 3' to 5' direction of the template RNA.
[0399] 395. The system or method of any of the previous embodiments, wherein the polypeptide comprises (a) a reverse transcriptase domain and (b) an endonuclease domain, and at least one of (a) and (b) is heterologous.
[0400] 396. The system or method of any of the above embodiments, wherein the polypeptide comprises (a) a target DNA binding domain, (b) a reverse transcriptase domain, and (c) an endonuclease domain, and at least one of (a), (b), and (c) is heterologous.
[0401] 397. A substantially pure polypeptide comprising (a) a reverse transcriptase domain and (b) a heterologous endonuclease domain, wherein one or both of (a) and (b) have an amino acid sequence encoded by the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0402] 398. A substantially pure polypeptide comprising (a) a reverse transcriptase domain and (b) a heterologous endonuclease domain, wherein one or both of (a) and (b) have an amino acid sequence encoded by the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides.
[0403] 399. A substantially pure polypeptide comprising (a) a reverse transcriptase domain and (b) a heterologous endonuclease domain, wherein one or both of (a) and (b) have the amino acid sequence of an element of Table 3B, Table 10, Table 11 or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0404] 400. A substantially pure polypeptide comprising (a) a reverse transcriptase domain and (b) a heterologous endonuclease domain, wherein one or both of (a) and (b) have the amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides.
[0405] 401. A substantially pure polypeptide comprising: (a) a reverse transcriptase domain encoded by a first sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (b) an endonuclease domain encoded by a second sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein the first sequence and the second sequence are selected from elements in different rows of Table 3B, Table 10, Table 11, or Table X.
[0406] 402. A substantially pure polypeptide comprising: (a) a reverse transcriptase domain encoded by a first sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides; and (b) an endonuclease domain encoded by a second sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, wherein the first sequence and the second sequence are selected from elements in different rows of Table 3B, Table 10, Table 11, or Table X.
[0407] 403. A substantially pure polypeptide comprising: (a) a reverse transcriptase domain of a first sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (b) an endonuclease domain of a second sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein the first sequence and the second sequence are selected from elements in different rows of Table 3B, Table 10, Table 11, or Table X.
[0408] 404. A substantially pure polypeptide comprising: (a) a reverse transcriptase domain of a first sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides; and (b) an endonuclease domain of a second sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, wherein the first sequence and the second sequence are selected from elements in different rows of Table 3B, Table 10, Table 11, or Table X.
[0409] 405. The substantially pure polypeptide of any of the previous embodiments, further comprising a target DNA binding domain encoded by the sequence of an element of Table 3B, Table 10, Table 11 or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0410] 406. The substantially pure polypeptide of any of the previous embodiments, further comprising a target DNA binding domain encoded by the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides.
[0411] 407. The substantially pure polypeptide of any of the previous embodiments, further comprising a target DNA binding domain of an element of Table 3B, Table 10, Table 11 or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0412] 408. The substantially pure polypeptide of any of the previous embodiments, further comprising a target DNA binding domain of an element of Table 3B, Table 10, Table 11 or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 nucleotides.
[0413] 409. A substantially pure polypeptide comprising (a) a reverse transcriptase domain and (b) a heterologous endonuclease domain, wherein (a) has an amino acid sequence listed in Table 3B, Table 10, Table 11, or Table Z1 or Table Z2, or Table X, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and optionally (b) has the amino acid sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0414] 410. A substantially pure polypeptide comprising (a) a reverse transcriptase domain and (b) a heterologous endonuclease domain, wherein (a) has an amino acid sequence listed in Table 3B, Table 10, Table 11, or Table Z1 or Table Z2, or or Table X, or a sequence that differs from it by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, and optionally (b) has the amino acid sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence that differs from it by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides.
[0415] 411. The substantially pure polypeptide of any of the previous embodiments, further comprising a target DNA binding domain comprising an amino acid sequence of an element of Table 3B, Table 10, Table 11 or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0416] 412. The substantially pure polypeptide of any of the previous embodiments, further comprising a target DNA binding domain comprising an amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides.
[0417] 413. A substantially pure polypeptide comprising: (a) a target DNA binding domain encoded by a first sequence listed in Table 3B, Table 10, Table 11, or Table X; (b) a reverse transcriptase domain encoded by a second sequence listed in Table 3B, Table 10, Table 11, or Table X; and (c) an endonuclease domain encoded by a third sequence listed in Table 3B, Table 10, Table 11, or Table X; (i) the first sequence and the second sequence are selected from different rows of Table 3B, Table 10, Table 11, or Table X; (ii) the first sequence and the third sequence are selected from different rows of Table 3B, Table 10, Table 11, or Table X; (iii) the second sequence and the third sequence are selected from different rows of Table 3B, Table 10, Table 11, or Table X; or (iv) A polypeptide, wherein the first sequence, the second sequence, and the third sequence are each selected from different rows of Table 3B, Table 10, Table 11, or Table X.
[0418] 414. A substantially pure polypeptide comprising: (a) a target DNA binding domain of a first sequence listed in Table 3B, Table 10, Table 11, or Table X; (b) a reverse transcriptase domain of a second sequence listed in Table 3B, Table 10, Table 11, or Table X; and (c) an endonuclease domain of a third sequence listed in Table 3B, Table 10, Table 11, or Table X; (i) the first sequence and the second sequence are selected from different rows of Table 3B, Table 10, Table 11, or Table X; (ii) the first sequence and the third sequence are selected from different rows of Table 3B, Table 10, Table 11, or Table X; (iii) the second sequence and the third sequence are selected from different rows of Table 3B, Table 10, Table 11, or Table X; or (iv) A polypeptide, wherein the first sequence, the second sequence, and the third sequence are each selected from different rows of Table 3B, Table 10, Table 11, or Table X.
[0419] 415. A target DNA binding domain comprising: (a) a first amino acid sequence of an element listed, for example, in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (b) a second amino acid sequence listed in Table Z1 or Z2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (c) an endonuclease domain comprising a third amino acid sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein optionally, the elements of the first amino acid sequence and the third amino acid sequence are selected from different rows of Table 3B, Table 10, Table 11, or Table X.
[0420] 416. A substantially pure polypeptide comprising: (a) a target DNA binding domain comprising a first amino acid sequence of an element listed, e.g., in Table 3B, Table 10, Table 11, or Table X, or a sequence that differs from it by at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer polypeptide residues; (b) a reverse transcriptase domain comprising a second amino acid sequence listed in Table Z1 or Z2, or a sequence that differs from it by at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer polypeptide residues; and (c) an endonuclease domain comprising a third amino acid sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence that differs from it by at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer polypeptide residues; Optionally, the elements of the first amino acid sequence and the third amino acid sequence are selected from different rows of Table 3B, Table 10, Table 11, or Table X.
[0421] 417. A substantially pure polypeptide comprising: (a) a target DNA binding domain comprising a first amino acid sequence; (b) a reverse transcriptase domain comprising a second amino acid sequence; and (c) an endonuclease domain comprising a third amino acid sequence, wherein the first, second, and third amino acid sequences are each encoded by a sequence listed in Table 3B, Table 10, Table 11, or Table X, or comprise an amino acid sequence listed in Table Z1, or the amino acid sequence of a domain listed in Table Z2.
[0422] 418. The substantially pure polypeptide of any of the previous embodiments, wherein the first amino acid sequence is encoded by a sequence listed in Table 3B, Table 10, Table 11, or Table X.
[0423] 419. The substantially pure polypeptide of any of the above embodiments, wherein the first amino acid sequence comprises an amino acid sequence listed in Table Z1 or Z2.
[0424] 420. The substantially pure polypeptide of any of the previous embodiments, wherein the second amino acid sequence is encoded by a sequence listed in Table 3B, Table 10, Table 11, or Table X.
[0425] 421. The substantially pure polypeptide of any of the above embodiments, wherein the second amino acid sequence comprises an amino acid sequence listed in Table Z1 or Z2.
[0426] 422. The substantially pure polypeptide of any of the previous embodiments, wherein the third amino acid sequence is encoded by a sequence listed in Table 3B, Table 10, Table 11, or Table X.
[0427] 423. The substantially pure polypeptide of any of the above embodiments, wherein the third amino acid sequence comprises an amino acid sequence listed in Table Z1 or Z2.
[0428] 424. A polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (a) a reverse transcriptase domain and (b) an endonuclease domain, and one or both of (a) and (b) have an amino acid sequence encoded by a sequence of an element of Table 3B, Table 10, Table 11 or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto, and wherein at least one of (a) and (b) is heterologous to the other.
[0429] 425. A polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (a) a reverse transcriptase domain and (b) an endonuclease domain, and one or both of (a) and (b) have an amino acid sequence encoded by a sequence of an element of Table 3B, Table 10, Table 11 or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 nucleotides, and at least one of (a) and (b) is heterologous to the other.
[0430] 426. A polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide is a fusion protein comprising (a) a reverse transcriptase domain encoded by a first sequence of an element listed in Table 3B, Table 10, Table 11 or Table X, and (b) an endonuclease domain encoded by a second sequence of an element listed in Table 3B, Table 10, Table 11 or Table X, wherein the first sequence and the second sequence are selected from elements in different rows of Table 3B, Table 10, Table 11 or Table X.
[0431] 427. A polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (a) a target DNA-binding domain, (b) a reverse transcriptase domain, and (c) an endonuclease domain, wherein at least one (e.g., one, two, or all) of (a), (b), and (c) comprises an amino acid sequence encoded by a sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and wherein at least one of (a), (b), and (c) is heterologous to the others.
[0432] 428. A polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (a) a target DNA binding domain, (b) a reverse transcriptase domain, and (c) an endonuclease domain, wherein at least one (e.g., one, two, or all) of (a), (b), and (c) comprises an amino acid sequence encoded by a sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, and wherein at least one of (a), (b), and (c) is heterologous to the others.
[0433] 429. A polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide is a fusion protein comprising: (a) a target DNA binding domain encoded by a first sequence listed in Table 3B, Table 10, Table 11, or Table X; (b) a reverse transcriptase domain encoded by a second sequence listed in Table 3B, Table 10, Table 11, or Table X; and (c) an endonuclease domain encoded by a third sequence listed in Table 3B, Table 10, Table 11, or Table X; (i) the first sequence and the second sequence are selected from different rows of Table 3B, Table 10, Table 11, or Table X; (ii) the first sequence and the third sequence are selected from different rows of Table 3B, Table 10, Table 11, or Table X; (iii) the second sequence and the third sequence are selected from different rows of Table 3B, Table 10, Table 11, or Table X; or (iv) A polypeptide or a nucleic acid encoding a polypeptide, wherein the first sequence, the second sequence, and the third sequence are each selected from different rows of Table 3B, Table 10, Table 11, or Table X.
[0434] 430. The polypeptide of any of the previous embodiments, wherein the reverse transcriptase domain has at least 80% identity (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity) to a reverse transcriptase domain encoded by the sequence of an element of Table 3B, Table 10, Table 11 or Table X.
[0435] 431. The polypeptide of any of the previous embodiments, wherein the endonuclease domain has at least 80% identity, e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity, to an endonuclease domain encoded by the sequence of an element of Table 3B, Table 10, Table 11 or Table X.
[0436] 432. The polypeptide or method of any of the previous embodiments, wherein the DNA-binding domain has at least 80% identity, e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity, to a DNA-binding domain encoded by the sequence of an element of Table 3B, Table 10, Table 11 or Table X.
[0437] 433. A polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (a) a reverse transcriptase domain and (b) an endonuclease domain, and one or both of (a) and (b) have the amino acid sequence of an element of Table 3B, Table 10, Table 11 or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto, and wherein at least one of (a) and (b) is heterologous to the other.
[0438] 434. A polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (a) a reverse transcriptase domain and (b) an endonuclease domain, and one or both of (a) and (b) have an amino acid sequence of an element of Table 3B, Table 10, Table 11 or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 nucleotides, and at least one of (a) and (b) is heterologous to the other.
[0439] 435. A polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide is a fusion protein comprising (a) a reverse transcriptase domain of a first sequence listed in Table 3B, Table 10, Table 11 or Table X, and (b) an endonuclease domain of a second sequence listed in Table 3B, Table 10, Table 11 or Table X, wherein the first sequence and the second sequence are selected from different rows of Table 3B, Table 10, Table 11 or Table X.
[0440] 436. A polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (a) a target DNA binding domain, (b) a reverse transcriptase domain, and (c) an endonuclease domain, wherein at least one (e.g., one, two, or all) of (a), (b), and (c) comprises the amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and wherein at least one of (a), (b), and (c) is heterologous to the others.
[0441] 437. A polypeptide or a nucleic acid encoding a polypeptide, comprising: (a) a target DNA binding domain, (b) a reverse transcriptase domain, and (c) an endonuclease domain, wherein at least one (e.g., one, two, or all) of (a), (b), and (c) comprises the amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides, and wherein at least one of (a), (b), and (c) is heterologous to the others.
[0442] 438. A polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide is a fusion protein comprising: (a) a target DNA binding domain of a first sequence listed in Table 3B, Table 10, Table 11 or Table X; (b) a reverse transcriptase domain of a second sequence listed in Table 3B, Table 10, Table 11 or Table X; and (c) an endonuclease domain of a third sequence listed in Table 3B, Table 10, Table 11 or Table X; (i) the first sequence and the second sequence are selected from different rows of Table 3B, Table 10, Table 11, or Table X; (ii) the first sequence and the third sequence are selected from different rows of Table 3B, Table 10, Table 11, or Table X; (iii) the second sequence and the third sequence are selected from different rows of Table 3B, Table 10, Table 11, or Table X; or (iv) A polypeptide or a nucleic acid encoding a polypeptide, wherein the first sequence, the second sequence, and the third sequence are each selected from different rows of Table 3B, Table 10, Table 11, or Table X.
[0443] 439. A polypeptide or a nucleic acid encoding said polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA binding domain (DBD), and (iii) an endonuclease domain, e.g., a nickase domain, wherein the RT domain has a sequence in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0444] 440. The polypeptide of any of the previous embodiments, wherein the reverse transcriptase domain has at least 80% identity (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity) to the reverse transcriptase domain of a sequence of an element of Table 3B, Table 10, Table 11 or Table X.
[0445] 441. The polypeptide of any of the previous embodiments, wherein the endonuclease domain has at least 80% identity, e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity, to the endonuclease domain of a sequence of an element of Table 3B, Table 10, Table 11 or Table X.
[0446] 442. The polypeptide or method of any of the previous embodiments, wherein the DNA binding domain has at least 80% identity, e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity, to the DNA binding domain of a sequence of an element of Table 3B, Table 10, Table 11 or Table X.
[0447] 443. A nucleic acid encoding a polypeptide of any of the previous embodiments.
[0448] 444. A vector comprising the nucleic acid of any of the preceding embodiments.
[0449] 445. A host cell comprising the nucleic acid of any of the preceding embodiments.
[0450] 446. A host cell comprising the polypeptide of any of the preceding embodiments.
[0451] 447. A host cell comprising the vector of any of the preceding embodiments.
[0452] 448. A heterologous sequence of interest (e.g., a sequence encoding a therapeutic polypeptide) at a target site in a chromosome, and (a) one or both of an untranslated region on one side (e.g., upstream) of a heterologous sequence of interest (e.g., a retrotransposon untranslated sequence, e.g., a sequence in Table 10 or a sequence in column 5 of Table 3A or 3B) and an untranslated region on the other side (e.g., downstream) of a heterologous sequence of interest (e.g., a retrotransposon untranslated sequence, e.g., a sequence in Table 10 or a sequence in column 6 of Table 3A or 3B), and / or (b) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain and (ii) an endonuclease domain, and one or both of (i) and (ii) have an amino acid sequence encoded by the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. One or both of A host cell (e.g., a human cell) comprising:
[0453] 449. A heterologous sequence of interest (e.g., a sequence encoding a therapeutic polypeptide) at a target site in a chromosome, and (a) one or both of an untranslated region on one side (e.g., upstream) of a heterologous sequence of interest (e.g., a retrotransposon untranslated sequence, e.g., a sequence in Table 10 or a sequence in column 5 of Table 3A or 3B) and an untranslated region on the other side (e.g., downstream) of a heterologous sequence of interest (e.g., a retrotransposon untranslated sequence, e.g., a sequence in Table 10 or a sequence in column 6 of Table 3A or 3B), and / or (b) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain and (ii) an endonuclease domain, and one or both of (i) and (ii) have an amino acid sequence encoded by the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides. One or both of A host cell (e.g., a human cell) comprising:
[0454] 450. A heterologous sequence of interest (e.g., a sequence encoding a therapeutic polypeptide) at a target site in a chromosome, and (a) one or both of an untranslated region on one side (e.g., upstream) of a heterologous sequence of interest (e.g., a retrotransposon untranslated sequence, e.g., a sequence in Table 10 or a sequence in column 5 of Table 3A or 3B) and an untranslated region on the other side (e.g., downstream) of a heterologous sequence of interest (e.g., a retrotransposon untranslated sequence, e.g., a sequence in Table 10 or a sequence in column 6 of Table 3A or 3B), and / or (b) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain and (ii) an endonuclease domain, and one or both of (i) and (ii) have the amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. One or both of A host cell (e.g., a human cell) comprising:
[0455] 451. A heterologous sequence of interest (e.g., a sequence encoding a therapeutic polypeptide) at a target site in a chromosome, and (a) one or both of an untranslated region on one side (e.g., upstream) of a heterologous sequence of interest (e.g., a retrotransposon untranslated sequence, e.g., a sequence in Table 10 or a sequence in column 5 of Table 3A or 3B) and an untranslated region on the other side (e.g., downstream) of a heterologous sequence of interest (e.g., a retrotransposon untranslated sequence, e.g., a sequence in Table 10 or a sequence in column 6 of Table 3A or 3B), and / or (b) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain and (ii) an endonuclease domain, and one or both of (i) and (ii) have an amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides. One or both of A host cell (e.g., a human cell) comprising:
[0456] 452. (i) The host cell of any of the preceding embodiments, comprising a heterologous sequence of interest (e.g., a sequence encoding a therapeutic polypeptide) at a target site in a chromosome, wherein the target locus is a Natural Harbor™ site, e.g., a site in Table 4 herein.
[0457] 453. (ii) The host cell of any of the preceding embodiments, further comprising one or both of an untranslated region 5' to the heterologous sequence of interest and an untranslated region 3' to the heterologous sequence of interest.
[0458] 454.(ii) The host cell of any of the preceding embodiments, further comprising an untranslated region on one side (e.g., upstream) of the heterologous sequence of interest (e.g., a retrotransposon untranslated sequence, e.g., a sequence of Table 10 or a sequence in column 5 of Table 3A or 3B) and an untranslated region on the other side (e.g., downstream) of the heterologous sequence of interest (e.g., a retrotransposon untranslated sequence, e.g., a sequence of Table 10 or a sequence in column 6 of Table 3A or 3B).
[0459] 455. The host cell of any of the preceding embodiments, comprising a heterologous sequence of interest only at the target site.
[0460] 456. A pharmaceutical composition comprising any of the systems, nucleic acids, polypeptides, or vectors described above; and a pharmaceutically acceptable excipient or carrier.
[0461] 457. The pharmaceutical composition of any of the previous embodiments, wherein the pharmaceutically acceptable excipient or carrier is selected from a vector (e.g., a viral or plasmid vector), a vesicle (e.g., a liposome, an exosome, a natural or synthetic lipid bilayer), a fusosome, a lipid nanoparticle.
[0462] 458. The polypeptide of any of the previous embodiments, further comprising a nuclear localization sequence.
[0463] 459. A template RNA (or DNA encoding the template RNA), comprising (e.g., from 5' to 3'): (i) optionally, a sequence that binds to a target site (e.g., the non-edited strand of a site in a target genome), (ii) optionally, a sequence that binds to an endonuclease and / or DNA-binding domain of a polypeptide, (iii) a heterologous sequence of interest, and (iv) a 3' homology domain.
[0464] 460. The template RNA of any of the preceding embodiments, comprising (i).
[0465] 461. The template RNA of any of the preceding embodiments, including (ii).
[0466] 462. A template RNA (or DNA encoding the template RNA) comprising (e.g., from 5' to 3') (i) a sequence that binds to a target site (e.g., the non-edited strand of a site in a target genome), (ii) a sequence that specifically binds to the RT domain of a polypeptide, (iii) a heterologous sequence of interest, and (iv) a 3' homology domain.
[0467] 463. The template RNA of any of the previous embodiments, wherein the RT domain comprises a sequence selected from Table 3B, Table 10, Table 11 or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0468] 464.(v) The template RNA of any of the preceding embodiments, further comprising a sequence that binds to an endonuclease and / or DNA binding domain of a polypeptide (eg, the same polypeptide comprising the RT domain).
[0469] 465. The template RNA of any of the preceding embodiments, wherein the sequence of (ii) specifically binds to an RT domain of Table 3B, Table 10, Table 11, or Table X, or an RT domain sequence having at least 70, 75, 80, 85, 90, 95, or 99% identity thereto.
[0470] 466. The template RNA of any of the previous embodiments, wherein the sequence that specifically binds to the RT domain is a sequence of Table 10, Table 11, or Table X, 3A or 3B, or a sequence having at least 70, 75, 80, 85, 90, 95 or 99% identity thereto.
[0471] 467. A template RNA (or DNA encoding the template RNA), comprising, from 5' to 3': (ii) a sequence that binds to the endonuclease and / or DNA binding domain of the polypeptide; (i) a sequence that binds to a target site (e.g., the non-edited strand of a site in a target genome); (iii) a heterologous sequence of interest; and (iv) a 3' homology domain.
[0472] 468. A template RNA (or DNA encoding the template RNA) comprising, from 5' to 3', (iii) a heterologous sequence of interest, (iv) a 3' homology domain, (i) a sequence that binds to a target site (e.g., the non-edited strand of a site in a target genome), and (ii) a sequence that binds to an endonuclease and / or DNA-binding domain of a polypeptide.
[0473] 469. The system or template RNA of any of the previous embodiments, wherein the template RNA, the first template RNA, or the second template RNA comprises a sequence that specifically binds to an RT domain.
[0474] 470. The system or template RNA of any of the preceding embodiments, wherein the sequence that specifically binds to the RT domain is located between (i) and (ii).
[0475] 471. The system or template RNA of any of the preceding embodiments, wherein the sequence that specifically binds to the RT domain is located between (ii) and (iii).
[0476] 472. The system or template RNA of any of the preceding embodiments, wherein the sequence that specifically binds to the RT domain is located between (iii) and (iv).
[0477] 473. The system or template RNA of any of the preceding embodiments, wherein the sequence that specifically binds to the RT domain is located between (ivi) and (i).
[0478] 474. The system or template RNA of any of the preceding embodiments, wherein the sequence that specifically binds to the RT domain is located between (i) and (iii).
[0479] 475. (a) a first template RNA comprising (i) a sequence that binds to an endonuclease domain, e.g., a nickase domain, and / or a DNA binding domain (DBD) of a polypeptide, and (ii) a sequence that binds to a target site (e.g., the non-edited strand of a site in a target genome) (where, e.g., the first RNA comprises a gRNA); (b) (i) a sequence (e.g., the polypeptide of (a)) that specifically binds to the reverse transcriptase (RT) domain of the polypeptide (e.g., the polypeptide of (a)), (ii) a target site binding sequence (TSBS), and (iii) a second template RNA (or DNA encoding the second template RNA) comprising an RT template sequence. A system for modifying DNA comprising:
[0480] 476. The system of any of the previous embodiments, wherein the nucleic acid encoding the first template RNA and the nucleic acid encoding the second template RNA are two separate nucleic acids.
[0481] 477. The system of any of the previous embodiments, wherein the nucleic acid encoding the first template RNA and the nucleic acid encoding the second template RNA are part of the same nucleic acid molecule, e.g., present on the same vector.
[0482] 478. A method for modifying a target DNA strand in a cell, tissue, or subject, comprising administering to the cell, tissue, or subject a system according to any of the above numbers, thereby modifying the target DNA strand.
[0483] 479. The embodiment of any number above, wherein the sequence of the element of Table X is selected from Vingi-1 EE, BovB, AviRTE_Brh, Penelope_SM, and Utopia_Dyak retrotransposase.
[0484] 480. The embodiment of any number above, wherein the template RNA comprises 5'UTR and 3'UTR of the same sequence of an element of Table 3B, Table 10, or Table X.
[0485] 481. The embodiment of any number above, wherein the sequence of the element of Table X comprises a Vingi-1 EE retrotransposase and the template RNA comprises the 5'UTR and 3'UTR of Vingi-1 EE.
[0486] 482. The embodiment of any of the above numbers, wherein the sequence of the element of Table X belongs to the restriction endonuclease-like (RLE) clade.
[0487] 483. The embodiment of any of the above numbers, wherein the sequence of the element of Table X belongs to the apurinic endonuclease-like (APE) clade.
[0488] 484. The embodiment of any number above, wherein the sequence of the element of Table X belongs to the Penelope-like element (PLE) clade, and optionally, the sequence of the element of Table X comprises a GIY-YIG domain (e.g., a GIY-YIG endonuclease domain).
[0489] 485. The embodiment of any number above, wherein the sequence of the element of Table X belongs to a clade selected from the CRE, NeSL, R4, R2, Hero, L1, RTE (e.g., AviRTE_Brh or BovB), I, Jockey, CR1, Rex1, Rand1 / Dualen, Penelope (e.g., Penelope_SM), Tx1, RTEX, Crack, Nimb, Proto1, Proto2, RTETP, L2, Tad1, Loa, Ingi, Outcast, R1, Daphne, L2A, L2B, Ambal, Vingi (e.g., Vingi-1_EE), and Kiri clades.
[0490] 486. The embodiment of any number above, wherein the target DNA binding domain is heterologous to one or more other domains of the polypeptide (e.g., the reverse transcriptase domain and / or the endonuclease domain).
[0491] 487. The embodiment of any number above, wherein the heterologous target DNA binding domain comprises Cas9, Cas9 nickase, dCas9, zinc finger, or TAL domain.
[0492] 488. The embodiment of any number above, wherein the heterologous target DNA binding domain comprises a Cas domain according to Table 9 or Table 37, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0493] 489. The embodiment of any number above, wherein the heterologous target DNA binding domain comprises an N-terminal dCas9 domain.
[0494] 490. The embodiment of any number above, wherein the system further comprises a guide RNA (e.g., a U6-driven gRNA) comprising at least 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides homologous to the target DNA sequence.
[0495] 491. The embodiment of any number above, wherein the template RNA further comprises a guide RNA region comprising at least 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides homologous to the target DNA sequence.
[0496] 492. The embodiment of any of the above numbers, wherein the gRNA sequence is at the 5' end of the template.
[0497] 493. The embodiment of any of the above numbers, wherein the gRNA sequence is at the 3' end of the template.
[0498] 494. The embodiment of any number above, wherein the gRNA sequence comprises a scaffold capable of recruiting Cas9.
[0499] 495. The embodiment of any number above, wherein the gRNA sequence comprises a homology domain, e.g., as described herein.
[0500] 496. The embodiment of any number above, wherein the endonuclease domain is heterologous to one or more other domains of the polypeptide (e.g., the reverse transcriptase domain and / or the target DNA binding domain).
[0501] 497. The embodiment of any number above, wherein the heterologous endonuclease domain comprises a Cas9, Cas9 nickase, or FokI domain.
[0502] 498. The embodiment of any of the above numbers, wherein the polypeptide comprises an RNase H domain.
[0503] 499. The embodiment of any one of the above numbers, wherein the polypeptide does not include an RNase H domain, or includes an inactivated RNase H domain.
[0504] 500. The embodiment of any one of the above numbers, wherein the nucleic acid encoding the polypeptide further comprises a second open reading frame.
[0505] 501. An embodiment of any number above, wherein the nucleic acid encoding the polypeptide comprises a 2A sequence, e.g., located between the ORF1 sequence and the ORF2 sequence, and optionally the 2A sequence is selected from T2A (EGRGSLLTCGDVEENPGP), P2A (ATNFSLLKQAGDVEENPGP), E2A (QCTNYALLKLAGDVESNPGP), or F2A (VKQTLNFDLLKLAGDVESNPGP).
[0506] 502. The embodiment of any number above, wherein the polypeptide comprises an intein.
[0507] 503. The embodiment of any number above, wherein the system includes an intein (e.g., included in the second polypeptide).
[0508] 504. The embodiment of any number above, wherein the polypeptide is encoded by two or more separate open reading frames, each encoding a polypeptide fragment.
[0509] 505. The embodiment of any number above, wherein an intein (e.g., a trans-splicing intein) joins two or more polypeptide fragments to form a polypeptide.
[0510] 506. Any number of embodiments above, wherein the system includes (i) a first polypeptide fragment comprising at least one of a reverse transcriptase domain, an endonuclease domain, and a target DNA binding domain, and (ii) a second polypeptide fragment comprising at least one of a reverse transcriptase domain, an endonuclease domain, and a target DNA binding domain, wherein the first polypeptide fragment does not comprise the same type of domain as the second polypeptide fragment.
[0511] 507. (a) the first polypeptide fragment comprises a reverse transcriptase domain, the second polypeptide fragment comprises an endonuclease domain, and optionally, the first polypeptide fragment further comprises a target DNA binding domain or the second polypeptide fragment further comprises a target DNA binding domain; (b) the first polypeptide fragment comprises a reverse transcriptase domain and the second polypeptide fragment comprises a target DNA binding domain, and optionally, the first polypeptide fragment further comprises an endonuclease domain or the second polypeptide fragment further comprises an endonuclease domain; or (a) Any number of embodiments above, wherein the first polypeptide fragment comprises an endonuclease domain and the second polypeptide fragment comprises a target DNA-binding domain, and optionally, the first polypeptide fragment further comprises a reverse transcriptase domain or the second polypeptide fragment further comprises a reverse transcriptase domain.
[0512] 508. The embodiment of any number above, wherein the intein joins a first polypeptide fragment to a second polypeptide to form a polypeptide.
[0513] 509. Intein (i) fusion of a reverse transcriptase domain to an endonuclease domain; (ii) fusion of a reverse transcriptase domain to a target DNA binding domain, or (iii) fusion of an endonuclease domain to a target DNA binding domain. Any number of embodiments above that induce
[0514] 510. The embodiment of any number above, wherein the intein is heterologous to one or more (e.g., one, two, or all) of the reverse transcriptase domain, the endonuclease domain, and the target DNA binding domain.
[0515] 511. The embodiment of any number above, wherein the intein is a split intein.
[0516] 512. The embodiment of any number above, wherein the DNA encoding the polypeptide comprises a plasmid, a minicircle, Doggybone DNA (dbDNA), or ceDNA.
[0517] 513. The embodiment of any number above, wherein the RNA encoding the polypeptide comprises one or more of the following: a cap region, a polyA tail, and / or a chemical modification, e.g., one or more chemically modified nucleotides.
[0518] 514. Any of the above embodiments, wherein the RNA encoding the polypeptide comprises a circRNA.
[0519] 515. Any of the preceding embodiments, wherein the nucleic acid encoding the polypeptide is contained within a virus (e.g., AAV, adenovirus, or lentivirus, e.g., an integration-defective lentivirus).
[0520] 516. Any of the above embodiments, wherein the nucleic acid encoding the polypeptide is contained within a nanoparticle (e.g., a lipid nanoparticle), a vesicle, or a fusosome.
[0521] 517. The embodiment of any number above, wherein the polypeptide comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the amino acid sequence encoded by the sequence of an element of Table X.
[0522] 518. The embodiment of any number above, wherein the reverse transcriptase domain comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to a reverse transcriptase domain of an amino acid sequence encoded by the sequence of an element of Table 3B, Table 10, Table 11 or Table X.
[0523] 519. The embodiment of any number above, wherein the retrotransposase comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to an amino acid sequence encoded by the sequence of an element of Table 3B, Table 10, Table 11 or Table X.
[0524] 520. The embodiment of any number above, wherein the polypeptide comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the amino acid sequence of the sequence of an element of Table 3B, Table 10, Table 11 or Table X.
[0525] 521. The embodiment of any number above, wherein the reverse transcriptase domain comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the reverse transcriptase domain of the amino acid sequence of the sequence of an element of Table 3B, Table 10, Table 11 or Table X.
[0526] 522. The embodiment of any number above, wherein the retrotransposase comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the amino acid sequence of the sequence of an element of Table 3B, Table 10, Table 11 or Table X.
[0527] 523. The embodiment of any number above, wherein the polypeptide comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 1023) or GGGS (SEQ ID NO: 1024).
[0528] 524. The embodiment of any number above, wherein the reverse transcriptase domain comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 1023) or GGGS (SEQ ID NO: 1024).
[0529] 525. The embodiment of any number above, wherein the retrotransposase comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 1023) or GGGS (SEQ ID NO: 1024).
[0530] 526. The embodiment of any number above, wherein the polypeptide, reverse transcriptase domain, or retrotransposase comprises a linker comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 1023) or GGGS (SEQ ID NO: 1024).
[0531] 527. The embodiment of any number above, wherein the polypeptide comprises a DNA-binding domain covalently linked to the remainder of the polypeptide by a linker, e.g., a linker comprising at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 125, 150, 200, 300, 400, or 500 amino acids.
[0532] 528. The embodiment of any number above, wherein the linker is attached to the remainder of the polypeptide at a position in the DNA-binding domain, RNA-binding domain, reverse transcriptase domain, or endonuclease domain (e.g., as shown in any of Figures 17A-17F of PCT Application No. PCT / US2019 / 048607).
[0533] 529. The embodiment of any number above, wherein the linker is attached to the remainder of the polypeptide at a position N-terminal to the alpha helical region of the polypeptide, e.g., a position corresponding to version v1, as described in Example 26 of PCT Application No. PCT / US2019 / 048607.
[0534] 530. The embodiment of any number above, wherein the linker is attached to the remainder of the polypeptide at a position C-terminal to the alpha helical region of the polypeptide, e.g., a position corresponding to version v2, e.g., preceding the RNA binding motif (e.g., the -1 RNA binding motif), as described in Example 26 of PCT Application No. PCT / US2019 / 048607.
[0535] 531. The embodiment of any number above, wherein the linker is attached to the remainder of the polypeptide C-terminal to the random coil region of the polypeptide, e.g., at a position N-terminal to the DNA binding motif (e.g., the c-myb DNA binding motif), e.g., at a position corresponding to version v3, as described in Example 26 of PCT Application No. PCT / US2019 / 048607.
[0536] 532. The embodiment of any number above, wherein the linker comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 1023) or GGGS (SEQ ID NO: 1024).
[0537] 533. The embodiment of any number above, wherein a polynucleotide sequence comprising at least about 500, 1000, 2000, 3000, 3500, 3600, 3700, 3800, 3900, or 4000 contiguous nucleotides from the 5' end of the template RNA sequence is integrated into the target cell genome.
[0538] 534. The embodiment of any number above, wherein a polynucleotide sequence comprising at least about 500, 1000, 2000, 2500, 2600, 2700, 2800, 2900, or 3000 contiguous nucleotides from the 3' end of the template RNA sequence is integrated into the target cell genome.
[0539] 535. The embodiment of any number above, wherein the nucleic acid sequence of the template RNA, or a portion thereof (e.g., a portion comprising at least about 100, 200, 300, 400, 500, 1000, 2000, 2500, 3000, 3500, or 4000 nucleotides) is integrated into the genome of a population of target cells at a copy number of at least about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1.0 integrants / genome.
[0540] 536. The embodiment of any number above, wherein the nucleic acid sequence of the template RNA, or a portion thereof (e.g., a portion comprising at least about 100, 200, 300, 400, 500, 1000, 2000, 2500, 3000, 3500, or 4000 nucleotides), is integrated into the genome of a population of target cells at a copy number of at least about 0.01, 0.02, 0.03, 0.04, 0.05, 0.75, or 0.1 integrants / genome.
[0541] 537. The embodiment of any number above, wherein the polypeptide comprises a functional endonuclease domain (wherein, e.g., the endonuclease domain does not comprise a mutation that abolishes endonuclease activity, e.g., as described herein).
[0542] 538. The embodiment of any number above, wherein introduction of the system into the target cell does not result in changes in p53 and / or p21 protein levels (e.g., upregulation), H2AX phosphorylation (e.g., gamma H2AX), ATM phosphorylation, ATR phosphorylation, Chk1 phosphorylation, Chk2 phosphorylation, and / or p53 phosphorylation.
[0543] 539. The embodiment of any number above, wherein introduction of the system into a target cell upregulates p53 protein levels in the target cell to a level that is less than about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the p53 protein level induced by introduction of a site-specific nuclease, e.g., Cas9, that targets the same genomic site as the system.
[0544] 540. The embodiment of any number above, wherein the p53 protein level is determined by the method described in Example 30.
[0545] 541. The embodiment of any number above, wherein introduction of the system into a target cell upregulates p53 phosphorylation levels in the target cell to a level that is less than about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the p53 phosphorylation level induced by introduction of a site-specific nuclease, e.g., Cas9, that targets the same genomic site as the system.
[0546] 542. The embodiment of any number above, wherein introduction of the system into a target cell upregulates p21 protein levels in the target cell to a level that is less than about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the p53 protein levels induced by introduction of a site-specific nuclease, e.g., Cas9, that targets the same genomic site as the system.
[0547] 543. The embodiment of any number above, wherein the p21 protein level is determined by the method described in Example 30.
[0548] 544. The embodiment of any number above, wherein introduction of the system into a target cell upregulates H2AX phosphorylation levels in the target cell to a level that is less than about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the H2AX phosphorylation level induced by introduction of a site-specific nuclease, e.g., Cas9, that targets the same genomic site as the system.
[0549] 545. The embodiment of any number above, wherein introduction of the system into a target cell upregulates ATM phosphorylation levels in the target cell to a level that is less than about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the ATM phosphorylation level induced by introduction of a site-specific nuclease, e.g., Cas9, that targets the same genomic site as the system.
[0550] 546. The embodiment of any number above, wherein introduction of the system into a target cell upregulates ATR phosphorylation levels in the target cell to a level that is less than about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the ATR phosphorylation level induced by introduction of a site-specific nuclease, e.g., Cas9, that targets the same genomic site as the system.
[0551] 547. The embodiment of any number above, wherein introduction of the system into a target cell upregulates Chk1 phosphorylation levels in the target cell to a level that is less than about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the Chk1 phosphorylation level induced by introduction of a site-specific nuclease, e.g., Cas9, that targets the same genomic site as the system.
[0552] 548. The embodiment of any number above, wherein introduction of the system into a target cell upregulates Chk2 phosphorylation levels in the target cell to a level that is less than about 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the Chk2 phosphorylation level induced by introduction of a site-specific nuclease, e.g., Cas9, that targets the same genomic site as the system.
[0553] 549. The embodiment of any number above, wherein the target DNA binding domain recognizes a specific target DNA sequence.
[0554] 550. The embodiment of any number above, wherein the target DNA binding domain binds to multiple (e.g., random) target DNA sequences.
[0555] 551. The embodiment of any number above, wherein the template RNA comprises a guide RNA (e.g., a U6-driven gRNA).
[0556] 552. The embodiment of any number above, wherein the target DNA-binding domain comprises a DNA-binding domain of one or more of a retrotransposase described herein (e.g., a retrotransposase of an element of Table X, 10, 11, 3A, or 3B), Cas9, nickase-Cas9, dCas9, zinc finger, TAL, meganuclease, and / or transcription factor.
[0557] 553. The embodiment of any number above, wherein the reverse transcriptase domain comprises the reverse transcriptase domain of a retrotransposase described herein (e.g., a retrotransposase of an element of Table X, 10, 11, Z1, Z2, 3A, or 3B).
[0558] 554. The embodiment of any number above, wherein the endonuclease domain comprises an endonuclease domain of a retrotransposase described herein (e.g., a retrotransposase of an element of Table X, 10, 11, 3A, or 3B), Cas9, a nickase Cas9, a type II restriction enzyme (e.g., FokI), a Holliday junction resolvase, an RLE endonuclease domain, an APE endonuclease domain, or a GIY-YIG endonuclease domain.
[0559] 555. A polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA binding domain (DBD), and (iii) an endonuclease domain, wherein the DBD and / or the endonuclease domain comprises a heterologous targeting domain that specifically binds to a sequence contained in a target DNA molecule (e.g., genomic DNA).
[0560] 556. A template RNA (or DNA encoding a template RNA) comprising a targeting domain (e.g., a heterologous targeting domain) that specifically binds to a sequence contained in a target DNA molecule (e.g., genomic DNA), a sequence that specifically binds to the RT domain of a polypeptide, and a heterologous sequence of interest.
[0561] 557. The system, method, or template RNA of any of the preceding embodiments, wherein the polypeptide comprises a heterologous targeting domain that specifically binds to a sequence contained in a target DNA molecule (e.g., genomic DNA).
[0562] 558. The system, method, or template RNA of any of the preceding embodiments, wherein the heterologous targeting domain binds to a different nucleic acid sequence than the unmodified polypeptide.
[0563] 559. The system, method, or template RNA of any of the preceding embodiments, wherein the polypeptide does not comprise a functional endogenous targeting domain (e.g., the polypeptide does not comprise an endogenous targeting domain).
[0564] 560. The system, method, or template RNA of any of the preceding embodiments, wherein the heterologous targeting domain comprises a zinc finger (e.g., a zinc finger that specifically binds to a sequence contained in the target DNA molecule).
[0565] 561. The system, method, or template RNA of any of the preceding embodiments, wherein the heterologous targeting domain comprises a Cas domain (e.g., a Cas9 domain, or a mutant or variant thereof, e.g., a Cas9 domain that specifically binds to a sequence contained in the target DNA molecule).
[0566] 562. The system, method, or template RNA of any of the preceding embodiments, wherein the Cas domain is associated with a guide RNA (gRNA).
[0567] 563. The system, method, or template RNA of any of the preceding embodiments, wherein the heterologous targeting domain comprises an endonuclease domain (e.g., a heterologous endonuclease domain).
[0568] 564. The system, method, or template RNA of any of the preceding embodiments, wherein the endonuclease domain comprises a Cas domain (e.g., Ca9, or a mutant or variant thereof).
[0569] 565. The system, method, or template RNA of any of the preceding embodiments, wherein the Cas domain is associated with a guide RNA (gRNA).
[0570] 566. The system, method, or template RNA of any of the preceding embodiments, wherein the endonuclease domain comprises a Fok1 domain.
[0571] 567. The system, method, or template RNA of any of the previous embodiments, wherein the template nucleic acid molecule comprises at least one (e.g., one or two) heterologous homologous sequence having at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% homology to a sequence contained in the target DNA molecule (e.g., genomic DNA).
[0572] 568. The system, method, or template RNA of any of the preceding embodiments, wherein one of the at least one heterologous sequence is located at or within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides of the 5' end of the template nucleic acid molecule.
[0573] 569. The system, method, or template RNA of any of the preceding embodiments, wherein one of the at least one heterologous sequence is located at or within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides of the 3' end of the template nucleic acid molecule.
[0574] 570. The system, method, or template RNA of any of the previous embodiments, wherein the heterologous homologous sequence binds within 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides of a nick site (e.g., generated by a nickase, e.g., an endonuclease domain, e.g., as described herein) in the target DNA molecule.
[0575] 571. The system, method, or template RNA of any of the previous embodiments, wherein the heterologous homologous sequence has less than 50%, 40%, 30%, 20%, 10%, 5%, 4%, 3%, 2%, or 1% sequence identity with a nucleic acid sequence complementary to an endogenous homologous sequence in an unmodified form of the template RNA.
[0576] 572. The system, method, or template RNA of any of the previous embodiments, wherein the heterologous homologous sequence has at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% homology to a sequence of the target DNA molecule that differs from the sequence bound by (e.g., replaced by) the endogenous homologous sequence.
[0577] 573. The system, method, or template RNA of any of the previous embodiments, wherein the heterologous sequence comprises a sequence (e.g., at its 3' end) having at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% homology to a sequence located 5' to the nick site of the target DNA molecule (e.g., the site nicked by a nickase, e.g., an endonuclease domain described herein).
[0578] 574. The system, method, or template RNA of any of the preceding embodiments, wherein the heterologous sequence comprises (e.g., at its 5' end) a sequence suitable for priming target primed reverse transcription (TPRT) initiation.
[0579] 575. The system, method, or template RNA of any of the preceding embodiments, wherein the heterologous homologous sequence has at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% homology to a sequence located within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides of (e.g., 3') the target insertion site in the target DNA molecule, e.g., for a heterologous sequence of interest (e.g., as described herein).
[0580] 576. The system, method, or template RNA of any of the preceding embodiments, wherein the template nucleic acid molecule comprises a guide RNA (gRNA), e.g., as described herein.
[0581] 577. The system, method, or template RNA of any of the preceding embodiments, wherein the template nucleic acid molecule comprises a gRNA spacer sequence (e.g., within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides of its 5' end).
[0582] 578. A template RNA (or DNA encoding the template RNA) comprising (e.g., from 5' to 3') (i) a sequence that binds to a target site (e.g., the second strand of a site in a target genome), (ii) a sequence that specifically binds to the RT domain of a polypeptide, (iii) a heterologous sequence of interest, and (iv) a 3' target homology domain.
[0583] 579.(v) The template RNA of any of the preceding embodiments, further comprising a sequence that binds to an endonuclease and / or DNA binding domain of a polypeptide (eg, the same polypeptide comprising the RT domain).
[0584] 580. The template RNA of any of the previous embodiments, wherein the RT domain comprises a sequence selected from Table 3B, 10, 11 or X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0585] 581. The template RNA of any of the previous embodiments, wherein the RT domain comprises a sequence selected from Table 3B, 10, 11 or X, and wherein the RT domain further comprises a number of substitutions relative to the native sequence, e.g., at least 1, 2, 3, 4, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 substitutions.
[0586] 582. The template RNA of any of the preceding embodiments, wherein the sequence of (ii) specifically binds to the RT domain.
[0587] 583. The template RNA of any of the previous embodiments, wherein the sequence that specifically binds to the RT domain is a sequence of Table 3B or 10, e.g., a UTR sequence, or a sequence having at least 70, 75, 80, 85, 90, 95, or 99% identity thereto.
[0588] 584. A template RNA (or DNA encoding the template RNA), comprising, from 5' to 3', (ii) a sequence that binds to the endonuclease and / or DNA binding domain of the polypeptide, (i) a sequence that binds to a target site (e.g., the second strand of a site in a target genome), (iii) a heterologous sequence of interest, and (iv) a 3' homology domain.
[0589] 585. A template RNA (or DNA encoding the template RNA) comprising, from 5' to 3', (iii) a heterologous sequence of interest, (iv) a 3' homology domain, (i) a sequence that binds to a target site (e.g., the second strand of a site in a target genome), and (ii) a sequence that binds to an endonuclease and / or DNA-binding domain of a polypeptide.
[0590] 586. A system, method, kit, template RNA, or reaction mixture in which the RNA of the system (e.g., template RNA, RNA encoding the polypeptide of (a), or RNA expressed from a heterologous sequence of interest incorporated into target DNA) contains, for example, a microRNA binding site in the 3'UTR.
[0591] 587. The system, method, kit, template RNA, or reaction mixture of any of the preceding embodiments, wherein the microRNA binding site is recognized by an miRNA that is present in a non-target cell type but not present in the target cell type (or is present at reduced levels compared to the non-target cells).
[0592] 588. The system, method, kit, template RNA, or reaction mixture of any of the preceding embodiments, wherein the miRNA is miR-142 and / or the non-target cells are Kupffer cells or blood cells, e.g., immune cells.
[0593] 144. The system, method, kit, template RNA, or reaction mixture of any of the preceding embodiments, wherein the miRNA is miR-182 or miR-183, and / or the non-target cells are dorsal root ganglion neurons.
[0594] 588. The system, method, kit, template RNA, or reaction mixture of any of the preceding embodiments, wherein the system includes a first miRNA binding site recognized by a first miRNA (e.g., miR-142), and the system further includes a second miRNA binding site recognized by a second miRNA (e.g., miR-182 or miR-183), and the first miRNA binding site and the second miRNA binding site are located on the same RNA or on different RNAs of the system.
[0595] 589. The system, method, kit, template RNA, or reaction mixture of any of the preceding embodiments, wherein the template RNA comprises at least two, three, or four miRNA binding sites, e.g., the miRNA binding sites are recognized by the same or different miRNAs.
[0596] 590. The system, method, kit, template RNA, or reaction mixture of any of the preceding embodiments, wherein the RNA encoding the polypeptide of (a) comprises at least two, three, or four miRNA binding sites, e.g., the miRNA binding sites are recognized by the same or different miRNAs.
[0597] 591. The system, method, kit, template RNA, or reaction mixture of any of the preceding embodiments, wherein the RNA expressed from the heterologous sequence of interest integrated into the target DNA comprises at least two, three, or four miRNA binding sites, e.g., the miRNA binding sites are recognized by the same or different miRNAs.
[0598] definition Domain: As used herein, the term "domain" refers to a structure of a biomolecule that contributes to a specific function of the biomolecule. A domain can include a continuous region (e.g., a contiguous sequence) or a discrete, non-contiguous region (e.g., a non-contiguous sequence) of a biomolecule. Examples of protein domains include, but are not limited to, endonuclease domains, DNA binding domains, and reverse transcription domains; examples of nucleic acid domains include regulatory domains, such as transcription factor binding domains.
[0599] Exogenous: As used herein, the term "exogenous," when used in reference to a biomolecule (such as a nucleic acid sequence or polypeptide), means that the biomolecule has been introduced into a host genome, cell, or organism by human intervention. For example, a nucleic acid that is added to an existing genome, cell, tissue, or subject using recombinant DNA technology or other methods is exogenous to the existing nucleic acid sequence, cell, tissue, or subject.
[0600] Genomic Safe Harbor Site (GSH Site): A genomic safe harbor site is a site within a host genome that can accommodate the integration of new genetic material, such that the inserted genetic element does not cause significant alterations to the host genome that pose a risk to the host cell or organism. GSH sites generally meet one, two, three, four, five, six, seven, eight, or nine of the following criteria: (i) located >300 kb from a cancer-related gene; (ii) located >300 kb from an miRNA / other functional small RNA; (iii) located >50 kb from the 5' gene end; (iv) located >50 kb from a replication origin; (v) located >50 kb away from an ultraconserved element; (vi) having low transcriptional activity (i.e., no mRNA + / - 25 kb); (vii) not within a variable copy number region; (viii) located within open chromatin; and / or (ix) having one copy and being unique within the human genome. Examples of GSH sites in the human genome that meet some or all of these criteria include: (i) adenovirus site 1 (AAVS1), the naturally occurring integration site of the AAV virus on chromosome 19; (ii) the chemokine (CC motif) receptor 5 (CCR5) gene, a chemokine receptor gene known as an HIV-1 co-receptor; (iii) the human orthologue of the mouse Rosa26 locus; and (iv) the rDNA locus. Additional GSH sites are known and are described, for example, in Pellenz et al. (2018) (https: / / doi.org / 10.1101 / 396390).
[0601] Heterologous: The term "heterologous," when used to refer to a first element in relation to a second element, means that the first and second elements do not naturally exist in the arrangement described. For example, a heterologous polypeptide, nucleic acid molecule, construct, or sequence refers to (a) a polypeptide, nucleic acid molecule, or portion of a polypeptide or nucleic acid molecule sequence that is not native to the cell in which it is expressed; (b) a polypeptide or nucleic acid molecule or portion of a polypeptide or nucleic acid molecule that has been modified or mutated relative to its natural state; or (c) a polypeptide or nucleic acid molecule that has altered expression compared to native expression levels under similar conditions. For example, heterologous regulatory sequences (e.g., promoters, enhancers) can be used to regulate expression of a gene or nucleic acid molecule in a manner that differs from how the gene or nucleic acid molecule is normally expressed in nature. In another example, a heterologous domain of a polypeptide or nucleic acid sequence (e.g., a DNA-binding domain of a polypeptide or a nucleic acid encoding a DNA-binding domain of a polypeptide) may be positioned relative to other domains or portions of a polypeptide or its encoding nucleic acid, and may be of a different sequence or from a different source. In certain embodiments, a heterologous nucleic acid molecule may be present naturally within the host cell genome, but may have an altered expression level or a different sequence, or both. In other embodiments, a heterologous nucleic acid molecule may not be endogenous to the host cell or host genome, but instead may be introduced into the host cell by transformation (e.g., transfection, electroporation), where the added molecule may be integrated into the host genome or may exist as extrachromosomal genetic material, either transiently (e.g., mRNA) or semi-stable for two or more generations (e.g., episomal viral vectors, plasmids, or other self-replicating vectors). In some embodiments, a domain is heterologous to another domain if the first domain is not naturally contained in the same polypeptide as the other domain (e.g., a fusion between two domains of different proteins from the same organism).
[0602] Mutation or Mutant: The term "(mutation)," when applied to a nucleic acid sequence, means that nucleotides within a nucleic acid sequence may be inserted, deleted, or changed relative to a reference (e.g., naturally occurring) nucleic acid sequence. A single alteration may be made at a single locus (point mutation), or multiple nucleotides may be inserted, deleted, or changed at a single locus. In addition, one or more alterations may be made at any number of loci within a nucleic acid sequence. Nucleic acid sequences can be mutated by any method known in the art. In some embodiments, the mutation is naturally occurring. In some embodiments, the desired mutation can be generated by the systems described herein.
[0603] Nucleic acid molecule: Nucleic acid molecule refers to both RNA and DNA molecules, including, but not limited to, cDNA, genomic DNA, and mRNA, and also includes synthetic nucleic acid molecules, such as those chemically synthesized or recombinantly produced, such as from an RNA template, as described herein. Nucleic acid molecules can be double-stranded or single-stranded, circular, or linear. If single-stranded, the nucleic acid molecule can be the sense or antisense strand. Unless otherwise noted, and as an example of all sequences described herein in the general format "SEQ ID NO:1," a nucleic acid containing "SEQ ID NO:1" refers to a nucleic acid having, at least a portion thereof, either (i) the sequence of SEQ ID NO:1 or (ii) a sequence complementary to SEQ ID NO:1. The choice between the two is determined by the context in which SEQ ID NO:1 is used. For example, if the nucleic acid is used as a probe, the choice between the two is determined by the requirement that the probe be complementary to the desired target. The nucleic acid sequences of the present disclosure may be chemically or biochemically modified or contain non-natural or derivatized nucleotide bases, as will be readily understood by those of skill in the art. Such modifications include, for example, labels, methylation, substitution of one or more naturally occurring nucleotides with analogs, internucleotide modifications such as uncharged linkages (e.g., methylphosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), pendant moieties (e.g., polypeptides), intercalating agents (e.g., acridines, psoralens, etc.), chelators, alkylating agents, and modified linkages (e.g., α-anomeric nucleic acids, etc.). Also included are synthetic molecules that mimic polynucleotides in their ability to bind to designated sequences through hydrogen bonding and other chemical interactions. Such molecules are known in the art and include, for example, those that substitute peptide linkages for phosphate linkages in the backbone of the molecule. Other modifications can include, for example, analogs in which the ribose ring contains bridging moieties or other structures, such as modifications found in "locked" nucleic acids.
[0604] Gene Expression Unit: A gene expression unit is a nucleic acid sequence comprising at least one regulatory nucleic acid sequence operably linked to at least one effector sequence. A first nucleic acid sequence is operably linked to a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For example, a promoter or enhancer is operably linked to a coding sequence if the promoter or enhancer affects the transcription or expression of the coding sequence. Operably linked DNA sequences can be contiguous or non-contiguous. Where necessary to link two protein coding regions, operably linked sequences can be in the same reading frame.
[0605] Host: The term host genome or host cell, as used herein, refers to a cell and / or its genome into which proteins and / or genetic material have been introduced. These terms refer not only to the particular subject cell and / or genome, but also to the progeny of such a cell and / or the genomes of the progeny of such a cell. Because certain modifications may occur in subsequent generations due to mutations or environmental influences, it is understood that such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term "host cell" as used herein. A host genome or host cell may be an isolated cell or cell line grown in culture, or genomic material isolated from such a cell or cell line, or it may be a host cell or host genome comprising a living tissue or organism. In some cases, the host cell may be an animal cell or a plant cell, for example, as described herein. In certain examples, the host cell may be a bovine cell, equine cell, porcine cell, caprine cell, ovine cell, chicken cell, or turkey cell. In certain examples, the host cell may be a corn cell, soybean cell, wheat cell, or rice cell.
[0606] Pseudoknot: As used herein, "pseudoknot sequence" refers to a nucleic acid (e.g., RNA) having a sequence suitable for self-complementarity to form a pseudoknot structure, e.g., a first segment, a second segment between the first and third segments (where the third segment is complementary to the first segment), and a fourth segment (where the fourth segment is complementary to the second segment). The pseudoknot may optionally have additional secondary structure, e.g., a stem-loop disposed within the second segment, a stem-loop disposed between the second and third segments, a sequence preceding the first segment, or a sequence following the fourth segment. The pseudoknot may have additional sequence between the first and second segments, between the second and third segments, or between the third and fourth segments. In some embodiments, the segments are arranged 5' to 3': first, second, third, and fourth. In some embodiments, the first and third segments comprise 5 base pairs of perfect complementarity. In some embodiments, the second and fourth segments comprise 10 base pairs, optionally with one or more (e.g., two) bulges. In some embodiments, the second segment comprises one or more unpaired nucleotides, e.g., forming a loop. In some embodiments, the third segment comprises one or more unpaired nucleotides, e.g., forming a loop.
[0607] Stem-loop sequence: As used herein, "stem-loop sequence" refers to a nucleic acid sequence (e.g., an RNA sequence) having a stem containing sufficient self-complementarity to form a stem-loop, e.g., at least 2 (e.g., 3, 4, 5, 6, 7, 8, 9, or 10) base pairs, and a loop having at least 3 (e.g., 4) base pairs. The stem may contain mismatches or bulges.
[0608] The patent or application file contains at least one drawing executed in color. Copies of any patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. In an embodiment of the present invention, for example, the following items are provided: (Item 1) 1. A system for modifying DNA, comprising: (a) a polypeptide or a nucleic acid (e.g., DNA or mRNA) encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain and (ii) an endonuclease domain, and one or both of (i) or (ii) have an amino acid sequence encoded by a nucleic acid sequence of an element of Table 3B, Table 10, or Table 11, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (b) a template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds to the polypeptide and (ii) a heterologous target sequence; A system including: (Item 2) 1. A system for modifying DNA, comprising: (a) a polypeptide or a nucleic acid (e.g., DNA or mRNA) encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain and (ii) an endonuclease domain, and one or both of (i) or (ii) have an amino acid sequence encoded by a nucleic acid sequence of an element of Table 3B, Table 10, or Table 11, or a sequence that differs therefrom by no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides; and (b) a template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds to the polypeptide and (ii) a heterologous target sequence; A system including: (Item 3) 1. A system for modifying DNA, comprising: (a) a polypeptide or a nucleic acid (e.g., DNA or mRNA) encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain and (ii) a target DNA binding domain, and one or both of (i) or (ii) have an amino acid sequence encoded by a nucleic acid sequence of an element of Table 3B, Table 10, or Table 11, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (b) a template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds to the polypeptide and (ii) a heterologous target sequence; A system including: (Item 4) 1. A system for modifying DNA, comprising: (a) a polypeptide or a nucleic acid encoding said polypeptide, said polypeptide comprising: (i) a reverse transcriptase (RT) domain; (ii) a DNA binding domain (DBD); and (iii) an endonuclease domain, e.g., a nickase domain; and (b) a template RNA (or DNA encoding the template RNA) comprising (e.g., from 5' to 3'): (i) optionally, a sequence that binds to a target site (e.g., the unedited strand of a site in a target genome); (ii) optionally, a sequence that binds to the polypeptide; (iii) a heterologous target sequence; and (iv) a 3' homology domain. Including, The system, wherein the RT domain has a sequence in Table 3B, Table 10, or Table 11, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. [Brief explanation of the drawings]
[0609] [Figure 1] FIG. 1 is a schematic diagram of the GeneWriting™ genome editing system. [Figure 2] FIG. 1 is a schematic diagram of the structure of the Gene Writer™ genome editor polypeptide. [Figure 3] FIG. 1 is a schematic diagram of the structure of an exemplary Gene Writer™ template RNA. [Figure 4A]
[0023] Figure 1 is a series of diagrams showing example configurations of Gene Writers using domains from various sources. Gene Writers described herein may or may not contain all of the domains shown. For example, a GeneWrite may optionally lack an RNA-binding domain or may have a single domain that fulfills the functions of multiple domains, such as a Cas9 domain for DNA binding and endonuclease activity. Exemplary domains that may be included within a Gene Writer polypeptide include a DNA-binding domain (e.g., including a DNA-binding domain set forth in any of Tables X, Y, Z1, Z2, 3A, or 3B; a zinc finger; a TAL domain; Cas9; dCas9; a nickase Cas9; a transcription factor; or a meganuclease), an RNA-binding domain (e.g., including an RNA-binding domain of a B-box protein, an MS2 coat protein, dCas, or an element of a sequence set forth in any of Tables X, Y, Z1, Z2, 3A, or 3B), a reverse transcriptase domain (e.g., a reverse transcriptase domain of an element of a sequence in a table herein; or other retrotransposase (e.g., as described in Table Z1); including peptides containing a reverse transcriptase domain (e.g., as described in Table Z2), and / or an endonuclease domain (including, for example, an endonuclease domain of an element described in any of Tables X, Y, Z1, Z2, 3A, or 3B; Cas9; nickase Cas9; restriction enzyme (e.g., a type II restriction enzyme, such as FokI); meganuclease; Holliday junction resolvase; RLE retrotranspase; APE retrotransposase; or GIY-YIG retrotransposase). Exemplary Gene Writer polypeptides containing exemplary combinations of such domains are shown in the bottom panel. [Figure 4B]
[0023] Figure 1 is a series of diagrams showing example configurations of Gene Writers using domains from various sources. Gene Writers described herein may or may not contain all of the domains shown. For example, a GeneWrite may optionally lack an RNA-binding domain or may have a single domain that fulfills the functions of multiple domains, such as a Cas9 domain for DNA binding and endonuclease activity. Exemplary domains that may be included within a Gene Writer polypeptide include a DNA-binding domain (e.g., including a DNA-binding domain set forth in any of Tables X, Y, Z1, Z2, 3A, or 3B; a zinc finger; a TAL domain; Cas9; dCas9; a nickase Cas9; a transcription factor; or a meganuclease), an RNA-binding domain (e.g., including an RNA-binding domain of a B-box protein, an MS2 coat protein, dCas, or an element of a sequence set forth in any of Tables X, Y, Z1, Z2, 3A, or 3B), a reverse transcriptase domain (e.g., a reverse transcriptase domain of an element of a sequence in a table herein; or other retrotransposase (e.g., as described in Table Z1); including peptides containing a reverse transcriptase domain (e.g., as described in Table Z2), and / or an endonuclease domain (including, for example, an endonuclease domain of an element described in any of Tables X, Y, Z1, Z2, 3A, or 3B; Cas9; nickase Cas9; restriction enzyme (e.g., a type II restriction enzyme, such as FokI); meganuclease; Holliday junction resolvase; RLE retrotranspase; APE retrotransposase; or GIY-YIG retrotransposase). Exemplary Gene Writer polypeptides containing exemplary combinations of such domains are shown in the bottom panel. [Figure 5]1 is a diagram showing the modules of an exemplary Gene Writer RNA template. Individual modules of the exemplary template can be combined, rearranged, and / or omitted to generate, for example, a Gene Writer template. A=5' homology arm; B=ribozyme; C=5'UTR; D=heterologous sequence of interest; E=3'UTR; F=3' homology arm. [Figure 6] 1 is a table listing the modules of an exemplary Gene Writer RNA template. Individual modules can be combined, rearranged, and / or omitted to generate, for example, a Gene Writer template. A=5' homology arm; B=ribozyme; C=5'UTR; D=heterologous sequence of interest; E=3'UTR; F=3' homology arm. [Figure 7]
[0023] Figure 1 is a diagram showing an exemplary second-strand nicking process. (A) Cas9 nickase is fused to a Gene Writer protein. The Gene Writer protein introduces a nick in the DNA strand via its EN domain (indicated as *), and the fused Cas9 nickase introduces a nick on the top or bottom DNA strand (indicated as X).
[0024] Figure 1 is a diagram showing an exemplary second-strand nicking process. (B) Gene Writer targets DNA via its DNA-binding domain and introduces a DNA nick via its EN domain (*). Cas9 nickase is then used to generate a second nick (X) on the top or bottom strand, upstream or downstream of the EN-introduced nick. [Figure 8]This figure shows the screening of construct designs for retrotransposon-mediated integration in human cells. A driver plasmid containing a retrotransposase (driver) expression cassette is cotransfected with a template plasmid containing a retrotransposon-dependent reporter cassette. Expression from the template plasmid results in a nonfunctional GFP due to the interruption of the antisense intron, whereas transcription of the template molecule from the template plasmid results in the production of RNA, which can be spliced to remove the intron and then reverse-transcribed and integrated by the system. Thus, expression of the reporter cassette occurs only from the integrated reporter cassette (integrated gDNA, bottom), not from the template plasmid. HA = homology arms (if applicable); CMV = mammalian CMV promoter; HiBit = HiBit tag for quantification of protein expression; T7 = T7 RNA polymerase promoter; UTR = untranslated sequence, e.g., natural retrotransposon UTR; pA = poly(A) signal; SD-SA is used to indicate the splice donor and acceptor sites of the antisense intron within the GFP coding sequence. HA = homology arm (if applicable) (see, e.g., Example 6); CMV = mammalian CMV promoter; HiBit = HiBit tag for quantification of protein expression; T7 = T7 RNA polymerase promoter; UTR = untranslated sequence, e.g., natural retrotransposon UTR; pA = polyA signal; SD-SA is used to indicate the splice donor and splice acceptor sites of the antisense intron in the GFP coding sequence. [Figure 9] Candidate retrotransposons identify 25 candidates that integrate trans-payloads in human cells. A total of 163 retrotransposon systems were assayed for activity in human cells as described in Example 4. Integration, as measured by ddPCR, is shown as copies per genome per retrotransposon driver / template system. The height of each bar represents the average value of two replicates. After further optimization in Examples 4 and 5, constructs with higher activity are further highlighted in Figure 10. [Figure 10]Based on the retrotransposon hits from Table 3B, the highly active Gene Writing configurations were further improved in Examples 7, 8, and 9. When multiple configurations of a given system were tested, such as alternative coding sequences for retrotransposase (Example 8) or the addition of homology arms (Example 9), only the highest performing configuration is shown. For systems improved over the initial configurations described in Table 3B and evaluated in Example 7, the improvements described in Example 5 (Figure 11) and Example 6 (Figure 12) are detailed in Table 11. [Figure 11] Retrotransposon consensus sequences can rescue or improve transintegration activity in human cells. Integration efficiency, measured by copies / genome by ddPCR as described in Example 5, is shown for each retrotransposon driver / template system. The height of each bar represents the average value of two replicates. Open circles represent each replicate, and bars represent the original sequence (light gray) or the new consensus-generated protein sequence (dark gray). [Figure 12] Figure 12 shows that consensus motif-generated homology arm sequences can rescue or improve the transintegration activity of retrotransposons in human cells. Integration efficiency, measured by copies / genome by ddPCR (Example 6), is shown for each Gene Writer driver / template system. The height of each bar represents the average value of two replicates. Black circles and open bars represent template sequences without homology arms, while open circles and hashed bars indicate designs that include homology arms. [Figure 13A]
[0049] Figure 1 shows a luciferase activity assay for primary cells. LNPs formulated according to Example 11 were analyzed for cargo delivery to primary human hepatocytes according to Example 12. The luciferase assay revealed dose-responsive luciferase activity from cell lysates. This indicates successful delivery of RNA from the mRNA cargo into cells and expression of firefly luciferase. [Figure 13B]
[0049] Figure 1 shows a luciferase activity assay for primary cells. LNPs formulated according to Example 11 were analyzed for cargo delivery to mouse hepatocytes according to Example 12. The luciferase assay revealed dose-responsive luciferase activity from cell lysates. This indicates successful delivery of RNA from mRNA cargo to cells and expression of firefly luciferase. [Figure 14] LNP-mediated delivery of RNA cargo to mouse liver is disclosed. Firefly luciferase mRNA-containing LNPs were formulated and delivered intravenously to mice. Liver samples were harvested and assayed for luciferase activity 6, 24, and 48 hours post-administration. Reporter activity in the various formulations followed the ranking LIPIDV005 > LIPIDV004 > LIPIDV003. RNA expression was transient, and enzyme levels returned to near vehicle background by 48 hours post-administration. DETAILED DESCRIPTION OF THE INVENTION
[0610] The present disclosure relates to compositions, systems, and methods for targeting, editing, modifying, or manipulating DNA sequences (e.g., inserting a heterologous DNA sequence of interest at a target site in a mammalian genome), e.g., at one or more locations within the DNA sequence in a cell, tissue, or subject, in vivo or in vitro. The DNA sequence of interest may include, for example, a coding sequence, a regulatory sequence, or a gene expression unit.
[0611] More specifically, the present disclosure provides a retrotransposon-based system for inserting a sequence of interest into a genome. The present disclosure is based, in part, on bioinformatics analysis to identify retrotransposase sequences and associated 5'UTRs and 3'UTRs from various organisms (see Tables 3B, 10, 11, and X). Further examples of retrotransposon elements are listed, for example, in Tables 1 and 2 of PCT Application No. PCT / US2019 / 048607, the entire contents of which are incorporated herein by reference.
[0612] In some embodiments, the systems described herein may have several advantages over various previous systems. For example, the present disclosure describes retrotransposases capable of inserting long sequences of heterologous nucleic acid into a genome. Furthermore, the retrotransposases described herein can insert heterologous nucleic acid into endogenous sites in a genome, such as rDNA loci. This is in contrast to the Cre / loxP system, which requires a first step of inserting an exogenous loxP site before a second step of inserting a sequence of interest into the loxP site.
[0613] Gene-writer™ Genome Editor Non-long terminal repeat (LTR) retrotransposons are a type of mobile genetic element widely found in eukaryotic genomes. They include, for example, apurinic / apyrimidinic endonuclease (APE) type, restriction enzyme-like endonuclease (RLE) type, and Penelope-like element (PLE) type. APE-class retrotransposons contain two functional domains: an endonuclease / DNA-binding domain and a reverse transcriptase domain. Examples of APE-class retrotransposons can be found, for example, in Table 1 of PCT Application No. PCT / US2019 / 048607 (incorporated herein by reference in its entirety), including, for example, the sequence listing and sequences referenced in Table 1 therein. The RLE class is composed of three functional domains: a DNA-binding domain, a reverse transcription domain, and an endonuclease domain. Examples of RLE-class retrotransposons can be found, for example, in Table 2 of PCT Application No. PCT / US2019 / 048607 (incorporated herein by reference in its entirety), including, for example, the sequence listing and sequences referenced in Table 2 therein. The reverse transcriptase domain of non-LTR retrotransposons functions by binding to an RNA sequence template and reverse transcribing it into target DNA in the host genome. The RNA sequence template has a 3' untranslated region that specifically binds to the retrotransposase and a variable 5' region that generally contains an open reading frame ("ORF") encoding the retrotransposase protein. The RNA sequence template may also contain a 5' untranslated region that specifically binds to the retrotransposase. Penelope-like elements (PLEs) are distinct from both LTR and non-LTR retrotransposons. PLEs generally differ from those of APE and RLE elements but contain a reverse transcriptase domain similar to those of telomerase and group II introns, as well as an optional GIY-YIG endonuclease domain.
[0614] Other exemplary classes of retrotransposons include, but are not limited to, CRE, NeSL, R4, R2, Hero, L1, RTE (e.g., AviRTE_Brh or BovB), I, Jockey, CR1, Rex1, Rand1 / Dualen, Penelope or Penelope-like (PLE) (e.g., Penelope_SM), Tx1, RTEX, Crack, Nimb, Proto1, Proto2, RTETP, L2, Tad1, Loa, Ingi, Outcast, R1, Daphne, L2A, L2B, Ambal, Vingi (e.g., Vingi-1_EE), and Kiri retrotransposons.
[0615] Such retrotransposon elements described herein can be functionally modularized and / or modified to target, edit, modify, or manipulate target DNA sequences, e.g., by reverse transcription, and to insert a desired (e.g., heterologous) nucleic acid sequence into a target genome, e.g., a mammalian genome. Such modularized and modified nucleic acid, polypeptide compositions, and systems are described herein and are referred to as Gene Writer™ gene editors. The Gene Writer™ gene editor system includes: (A) a polypeptide or a nucleic acid encoding a polypeptide, where the polypeptide includes (i) a reverse transcriptase domain and (x) an endonuclease domain containing a DNA-binding function, or (y) an endonuclease domain and a separate DNA-binding domain; and (B) a template RNA that includes (i) a sequence that binds to the polypeptide, and (ii) a heterologous insert sequence. For example, a Gene Writer™ genome editor protein may include a DNA-binding domain, a reverse transcriptase domain, and an endonuclease domain. In other embodiments, a Gene Writer™ genome editor protein may include a reverse transcriptase domain and an endonuclease domain. In certain embodiments, elements of the Gene Writer™ gene editor polypeptide may be derived from the sequence of a retrotransposon, such as an APE-type, RLE-type, or PLE-type retrotransposon, or a portion or domain thereof. In some embodiments, the RLE-type non-LTR retrotransposon is derived from the R2, NeSL, HERO, R4, or CRE clade. In some embodiments, the Gene Writer genome editor is derived from the R4 element X4_Line, which is found in the human genome. In some embodiments, the APE-type non-LTR retrotransposon is derived from the R1 or Tx1 clade. In some embodiments, the Gene Writer genome editor is derived from the Tx1 element Mare6, which is found in the human genome.The RNA template element of the GeneWriter™ gene editor system is typically heterologous to the polypeptide element and provides the sequence of interest to be inserted (reverse transcribed) into the host genome. In some embodiments, the GeneWriter genome editor protein is capable of target-primed reverse transcription.
[0616] In some embodiments, the GeneWriter genome editor is combined with a second polypeptide. In some embodiments, the second polypeptide is derived from an APE-type non-LTR retrotransposon. In some embodiments, the second polypeptide has a zinc knuckle-like motif. In some embodiments, the second polypeptide is a homolog of the Gag protein.
[0617] In some embodiments, the Gene Writer genome editor includes retrotransposase sequences for elements listed in Table X. Table X provides a set of nucleic acid and amino acid sequences (listed by Repbase gene name and species) for which associated GenBank accession numbers are available. The Repbase nucleic acid and amino acid sequences for elements listed in Table X are incorporated herein by reference in their entireties. The GenBank sequences for elements listed in Table X are also incorporated herein by reference in their entireties. The nucleic acid sequences of Table X, in some instances, include sequences encoding polypeptides (e.g., open reading frames) (e.g., protein-coding sequences of corresponding Repbase entries, which are incorporated herein by reference). The nucleic acid sequences of Table X, in some instances, include 5'UTR sequences (e.g., 5'UTR sequences of corresponding Repbase entries, which are incorporated herein by reference). The nucleic acid sequences of Table X, in some instances, include 3'UTR sequences (e.g., 3'UTR sequences of corresponding Repbase entries, which are incorporated herein by reference).
[0618] In some embodiments, the open reading frames (ORFs) of amino acid sequences in Repbase are annotated. In some embodiments, the ORFs of amino acid sequences in Repbase are not annotated. If an ORF is not annotated, one of skill in the art can identify it, for example, by performing one or more translations (e.g., full-frame translations) and, optionally, comparing the translations to a reverse transcriptase sequence or consensus motif. In some embodiments, for example, an amino acid sequence of Table X as used herein is the amino acid sequence listed in the corresponding Repbase entry or encoded by the nucleic acid sequence of the corresponding Repbase entry. In some embodiments, an amino acid sequence encoded by an element of Table X is the amino acid sequence encoded by the full-length sequence of the element listed in Table X, or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the full-length sequence of an element listed in Table X may comprise one or more (e.g., all) of the 5' UTR, polypeptide coding sequence, and 3' UTR of a retrotransposon described herein. In some embodiments, the amino acid sequence of Table X is the amino acid sequence encoded by the full-length sequence of an element listed in Table X, or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the 5' UTR of an element of Table X comprises the 5' UTR of the full-length sequence of an element listed in Table X, or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the 3'UTR of an element of Table X comprises the 3'UTR of the full-length sequence of the element listed in Table X, or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
[0619] Table X also provides a list of the host organisms from which the nucleic acid sequences were obtained, and the domains present in the polypeptides encoded by the open reading frames of the nucleic acid sequences. The domains listed in Table X are indicated as domain identifiers, which correspond to the InterPro domain entries listed in Table Y. Thus, a Repbase sequence listed in Table X can encode a polypeptide comprising one or more domains indicated in Table X by their domain identifiers. The particular domains or domain types associated with each domain identifier are described in more detail in Table Y. In some embodiments, a domain of interest (e.g., as listed in Table Y) can be identified in a nucleic acid sequence (e.g., as listed in Table X) by performing a full-frame translation and identifying the amino acid sequence encoding the domain of interest.
[0620] Retrotransposon discovery tools As a result of repeat mobilization over time, transposable elements in genomic DNA often exist as tandem or interspersed repeats (Jurka Curr Opin Struct Biol 8, 333-337 (1998)). Tools that can recognize such repeats can be used to identify new elements from genomic DNA and add them to databases, such as Repbase (Jurka et al. Cytogenet Genome Res 110, 462-467 (2005)). One such tool for identifying repeats that may contain transposable elements is RepeatFinder (Volfovsky et al. Genome Biol 2 (2001)), which analyzes the repeat structure of genomic sequences. Repeats can be further collected and analyzed using additional tools, such as Censor (Kohany et al. BMC Bioinformatics 7, 474 (2006)). The Censor package retrieves genomic repeats and annotates them using various BLAST approaches against known transposable elements. Full-frame translation can be used to generate ORFs for comparison.
[0621] Other exemplary methods for identifying transposable elements include RepeatModeler2, which automates the discovery and annotation of transposable elements in genome sequences (Flynn et al., bioRxiv (2019)). In addition to achieving this with available packages such as Censor, full-frame translations of a given genome or sequence can be performed and annotated with protein domain tools such as InterProScan, which uses the InterPro database to tag domains in a given amino acid sequence (Mitchell et al., Nucleic Acids Res. 47, D351-360(2019)), allowing the identification of potential proteins containing domains related to known transposable elements (e.g., domains of elements listed in Table X).
[0622] Retrotransposons can be further classified by their reverse transcriptase domain using tools such as RT class 1 (Kapitonov et al. Gene 448, 207-213 (2009)).
[0623] Polypeptide components of the GeneWriter gene editor system RT domain: In certain aspects of the invention, the reverse transcriptase domain of the Gene Writer system is based on the reverse transcriptase domain of an APE-type or RLE-type non-LTR retrotransposon, or a PLE-type retrotransposon. The wild-type reverse transcriptase domain of an APE-type, RLE-type, or PLE-type retrotransposon can be used in the Gene Writer system, or it can be modified (e.g., by inserting, deleting, or substituting one or more residues) to alter the reverse transcriptase activity of the target DNA sequence. In some embodiments, the reverse transcriptase is modified from its native sequence to alter codon usage, e.g., to improve for use in human cells. In some embodiments, the reverse transcriptase domain is a heterologous reverse transcriptase from a different retrovirus, retron, diversity-generating retroelement, retroplasmid, group II intron, LTR retrotransposon, non-LTR retrotransposon, or other source, such as those exemplified in Table Z1 or those containing a domain listed in Table Z2. In certain embodiments, the Gene Writer system comprises a polypeptide comprising the reverse transcriptase domain of an RLE-type non-LTR retrotransposon from the R2, NeSL, HERO, R4, or CRE clade, an L1, RTE, I, Jockey, CR1, Rex1, RandI / Dualen, T1, RTEX, Crack, Nimb, Proto1, Proto2, RTETP, L2, Tad1, Loa, Ingi, Outcast, R1, Daphne, L2A, L2B, Ambal, Vingi, or Kiri clade, or an APE-type non-LTR retrotransposon from the L2A, L2B, Ambal, Vingi, or Kiri clade. In certain embodiments, the Gene Writer system comprises a polypeptide comprising the reverse transcriptase domain of a retrotransposon listed in Table 10, Table 11, Table X, Table Z1, Table Z2, or Table 3A or 3B.In embodiments, the amino acid sequence of the reverse transcriptase domain of the Gene Writer system is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to the amino acid sequence of the reverse transcriptase domain of a retrotransposon whose DNA sequence is listed in Table 10, Table 11, Table X, Table Z1, Table Z2, or Table 3A or 3B. Reverse transcription domains can be identified based on homology to other known reverse transcription domains, for example, using routine tools such as the Basic Local Alignment Search Tool (BLAST). In some embodiments, the reverse transcriptase domain is modified, for example, by site-directed mutagenesis. In some embodiments, the reverse transcriptase domain is engineered to bind to a heterologous template RNA.
[0624] In some embodiments, the polypeptide (e.g., the RT domain) comprises an RNA-binding domain, e.g., that specifically binds to an RNA sequence. In some embodiments, the template RNA comprises an RNA sequence that is specifically bound by the RNA-binding domain.
[0625] In some embodiments, the RT domain exhibits increased stringency for target primed reverse transcription (TPRT) initiation, e.g., compared to the endogenous RT domain. In some embodiments, the RT domain initiates TPRT when 3 nt within the target site immediately upstream of the first-strand nick, e.g., the genomic DNA priming the RNA template, are at least 66% or 100% complementary to the 3 nt of homology in the RNA template. In some embodiments, the RT domain initiates TPRT when there is less than a 5 nt mismatch (e.g., less than a 1, 2, 3, 4, or 5 nt mismatch) between the template RNA and the target DNA primed reverse transcription. In some embodiments, the RT domain is modified to increase the stringency of mismatches in priming the TPRT reaction, e.g., the RT domain tolerates no mismatches or tolerates fewer mismatches within the priming region compared to a wild-type (e.g., unmodified) RT domain. In some embodiments, the RT domain comprises an HIV-1 RT domain. In embodiments, the HIV-1 RT domain initiates synthesis at a lower level, even with three nucleotide mismatches, compared to alternative RT domains (e.g., as described by Jamburuthugoda and Eickbush J Mol Biol 407(5):661-672 (2011), which is incorporated herein by reference in its entirety).
[0626] In some embodiments, the RT domain forms a dimer (e.g., a heterodimer or a homodimer). In some embodiments, the RT domain is monomeric. In some embodiments, the RT domain, e.g., a retroviral RT domain, naturally functions as a monomer or a dimer (e.g., a heterodimer or a homodimer). In some embodiments, the RT domain naturally functions as a monomer, e.g., is derived from a virus that functions as a monomer. Exemplary monomeric RT domains, their viral sources, and their associated RT signatures can be found in Table 30, along with the description of the domain signatures in Table 32. In some embodiments, the RT domain of the systems described herein comprises an amino acid sequence of Table 30, or a functional fragment or variant thereof, or a sequence having at least 70%, 80%, 90%, 95%, or 99% identity thereto.In some embodiments, the RT domain is selected from the group consisting of murine leukemia virus (MLV; sometimes referred to as MoMLV) (e.g., P03355), porcine endogenous retrovirus (PERV) (e.g., UniProt Q4VFZ2), mouse mammary tumor virus (MMTV) (e.g., UniProt P03365), Mason-Pfizer monkey virus (MPMV) (e.g., UniProt P07572), bovine leukemia virus (BLV) (e.g., UniProt P03361), human T-cell leukemia virus-1 (HTLV-1) (e.g., UniProt P03362), human foamy virus (HFV) (e.g., UniProt P03363), and / or porcine endogenous retrovirus (PERV) (e.g., UniProt Q4VFZ2). P14350), simian foamy virus (SFV) (e.g., UniProt P23074), or bovine foamy / syncytial virus (BFV / BSV) (e.g., UniProt O41894), or a functional fragment or variant thereof (e.g., an amino acid sequence having at least 70%, 80%, 90%, 95%, or 99% identity thereto). In some embodiments, the RT domain is dimeric in its native functionality. Exemplary dimeric RT domains, their viral sources, and their associated RT signatures can be found in Table 31, with the domain signatures listed in Table 32. In some embodiments, the RT domain of the systems described herein comprises an amino acid sequence of Table 31, or a functional fragment or variant thereof, or a sequence having at least 70%, 80%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain is derived from a virus that functions as a dimer.In embodiments, the RT domain is selected from the group consisting of avian sarcoma / leukemia virus (ASLV) (e.g., UniProt A0A142BKH1), Rous sarcoma virus (RSV) (e.g., UniProt P03354), avian myeloblastosis virus (AMV) (e.g., UniProt Q83133), human immunodeficiency virus type I (HIV-1) (e.g., UniProt P03369), human immunodeficiency virus type II (HIV-2) (e.g., UniProt P15833), simian immunodeficiency virus (SIV) (e.g., UniProt P05896), bovine immunodeficiency virus (BIV) (e.g., UniProt P19560), equine infectious anemia virus (EIAV) (e.g., UniProt P03371), or feline immunodeficiency virus (FIV) (e.g., UniProt P16088) (Herschhorn and Hizi Cell Mol Life Sci 67(16):2717-2747(2010)), or a functional fragment or variant thereof (e.g., an amino acid sequence having at least 70%, 80%, 90%, 95%, or 99% identity thereto). Naturally, heterodimeric RT domains may, in some embodiments, also be functional as homodimers. In some embodiments, the dimeric RT domains are expressed as fusion proteins, e.g., as homodimeric or heterodimeric fusion proteins. In some embodiments, the RT function of the system is fulfilled by multiple RT domains (e.g., as described herein). In further embodiments, the multiple RT domains may be fused or separate, e.g., on the same polypeptide or on different polypeptides.
[0627] In some embodiments, a Gene Writer described herein comprises an integrase domain, e.g., the integrase domain can be part of an RT domain. In some embodiments, an RT domain (e.g., as described herein) comprises an integrase domain. In some embodiments, an RT domain (e.g., as described herein) lacks an integrase domain or comprises an integrase domain that has been inactivated by mutation or deletion. In some embodiments, a Gene Writer described herein comprises an RNase H domain, e.g., the RNase H domain can be part of an RT domain. In some embodiments, an RT domain (e.g., as described herein) comprises an RNase H domain, e.g., an endogenous RNase H domain or a heterologous RNase H domain. In some embodiments, an RT domain (e.g., as described herein) lacks an RNase H domain. In some embodiments, an RT domain (e.g., as described herein) comprises an RNase H domain that has been added, deleted, mutated, or exchanged for a heterologous RNase H domain. In some embodiments, mutation of the RNase H domain produces a polypeptide that exhibits reduced RNase activity, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% less, compared to an otherwise similar domain lacking the mutation, as measured, e.g., by the method in Kotewicz et al. Nucleic Acids Res 16(1):265-277 (1988), incorporated herein by reference in its entirety. In some embodiments, RNase H activity is eliminated.
[0628] In some embodiments, the RT domain is mutated to increase fidelity compared to other similar domains that do not have the mutation. For example, in some embodiments, a YADD or YMDD motif within the RT domain (e.g., in a reverse transcriptase) is replaced with YVDD. In some embodiments, the YADD or YMDD or YVDD substitution results in greater fidelity in retroviral reverse transcriptase activity (e.g., Jamburutugoda). and Eickbush J Mol Biol 2011; which is incorporated herein by reference in its entirety).
[0629] The diversity of reverse transcriptases (e.g., containing RT domains) includes those used by prokaryotes (Zimmerly et al. Microbiol Spectr 3(2):MDNA3-0058-2014(2015); Lampson BC(2007) Prokaryotic Reverse Transcriptases. In: Polaina J., MacCabe AP(eds) Industrial Enzymes. Springer, Dordrecht), those used by viruses (Herschhorn et al. Cell Mol Life Sci 67(16):2717-2747(2010); Menendez-Arias et al. Virus Res 234:153-176(2017)), and those used by mobile elements (Eickbush et al. Virus Res 134(1-2):221-234(2008); Craig et al. Mobile DNA III 3rd Ed. DOI:10.1128 / 9781555819217(2015)), each of which is incorporated herein by reference.
[0630] [Table 1]
[0631] [Table 2]
[0632] Table 3
[0633] Table 4
[0634] Table 5
[0635] Table 6
[0636] Table 7
[0637] Table 8
[0638] Table 9
[0639] Table 10
[0640] Table 11
[0641] Table 12
[0642] Table 13
[0643] Table 14
[0644] Table 15
[0645] Table 16
[0646] Table 17
[0647] Table 18
[0648] Table 19
[0649] Table 20
[0650] Table 21
[0651] Table 22
[0652] Table 23
[0653] Table 24
[0654] Table 25
[0655] Table 26
[0656] Table 27
[0657] Table 28
[0658] Table 29
[0659] Table 30
[0660] Table 31
[0661] Table 32
[0662] Table 33
[0663] [Table 34]
[0664] Endonuclease In some embodiments, the polypeptide comprises an endonuclease domain (e.g., a heterologous endonuclease domain). In certain embodiments, the endonuclease / DNA-binding domain of an APE-type retrotransposon, the endonuclease domain of an RLE-type retrotransposon, or the endonuclease domain of a PLE-type retrotransposon can be used or modified (e.g., by insertion, deletion, or substitution of one or more residues) in the Gene Writer systems described herein. In some embodiments, the endonuclease domain or endonuclease / DNA-binding domain is modified from its native sequence to alter codon usage, e.g., to improve for human cells. In some embodiments, the endonuclease element is a heterologous endonuclease element, such as Fok1 nuclease, Cas9, Cas9 nickase, a type II restriction enzyme-like endonuclease (RLE-type nuclease), or another RLE-type endonuclease (also known as a REL). In some embodiments, the heterologous endonuclease domain cleaves both DNA strands, forming a double-stranded break. In some embodiments, the heterologous endonuclease activity has nickase activity and does not form a double-stranded break. The amino acid sequence of the reverse transcriptase domain of the Gene Writer system described herein can be at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to the amino acid sequence of the endonuclease domain of a retrotransposon whose DNA sequence is listed in Tables X, Z1, Z2, 3A, or 3B. Endonuclease domains can be identified based on homology to other known endonuclease domains, for example, using tools such as the Basic Local Alignment Search Tool (BLAST). In certain embodiments, the heterologous endonuclease is Cas9 or Cas9 nickase, or a functional fragment thereof. In certain embodiments, the heterologous endonuclease is Fok1 or a functional fragment thereof.In certain embodiments, the heterologous endonuclease is a Holliday junction resolvase or a homolog thereof, such as Holliday junction resolvase-Ssol Hje from Sulfolobus solfataricus (Govindaraju et al., Nucleic Acids Research 44:7, 2016). In certain embodiments, the heterologous endonuclease is a large fragment endonuclease of a spliceosomal protein, such as Prp8 (Mahbub et al., Mobile DNA 8:16, 2017). For example, the Gene Writer polypeptides described herein can contain a reverse transcriptase domain derived from an APE-, RLE-, or PLE-type retrotransposon and an endonuclease domain comprising Fok1 or a functional fragment thereof. In yet other embodiments, the homologous endonuclease domain is modified, for example, by site-directed mutagenesis, to alter DNA endonuclease activity. In yet other embodiments, the endonuclease domain is modified to remove any potential DNA sequence specificity.
[0665] In some embodiments, the Gene Writer polypeptide functions to cleave a DNA target site with an endonuclease domain. In some embodiments, the endonuclease domain is also a DNA-binding domain. In some embodiments, the endonuclease domain is also a template nucleic acid (e.g., template RNA)-binding domain. For example, in some embodiments, the polypeptide comprises a CRISPR-associated endonuclease domain that binds to a template RNA, including a gRNA, and binds to and cleaves a target DNA sequence (e.g., complementary to a portion of the gRNA). In certain embodiments, the endonuclease / DNA-binding domain of an APE-type retrotransposon or the endonuclease domain of an RLE-type retrotransposon is a CRISPR-associated endonuclease domain that binds to a template RNA, including a gRNA, and cleaves the target DNA sequence. ... They can be used in Writer systems or can be modified (e.g., by inserting, deleting, or substituting one or more residues). In some embodiments, the endonuclease domain or endonuclease / DNA-binding domain is modified from its native sequence to alter codon usage, e.g., to improve for use in human cells.
[0666] In some embodiments, the endonuclease element is a heterologous endonuclease element, such as Fok1 nuclease, a Type II restriction l-like endonuclease (RLE-type nuclease), or another RLE-type endonuclease (also known as REL). In some embodiments, the heterologous endonuclease activity has nickase activity and does not form double-stranded breaks. The amino acid sequence of the endonuclease domain of the Gene Writer systems described herein may be at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to the amino acid sequence of the endonuclease domain of a retrotransposon whose DNA sequence is referenced in Table 1, 2, 3A, or 3B. Endonuclease domains can be identified based on homology to other known endonuclease domains, for example, using tools such as the Basic Local Alignment Search Tool (BLAST). In certain embodiments, the heterologous endonuclease is Fok1 or a functional fragment thereof. In certain embodiments, the heterologous endonuclease is a Holliday junction resolvase or a homolog thereof, such as the Holliday junction cleavage enzyme (Ssol Hje) from Sulfolobus solfataricus (Govindaraju et al., Nucleic Acids Research 44:7, 2016). In certain embodiments, the heterologous endonuclease is a large fragment endonuclease of a spliceosomal protein, such as Prp8 (Mahbub et al., Mobile DNA 8:16, 2017). In certain embodiments, the heterologous endonuclease is derived from a CRISPR-associated protein, such as Cas9. In certain embodiments, the heterologous endonuclease is engineered to have only ssDNA cleavage activity, for example, only nickase activity, such as, for example, a Cas9 nickase.For example, a Gene Writer polypeptide described herein may comprise a reverse transcriptase domain from an APE- or RLE-type retrotransposon and an endonuclease domain comprising Fok1 or a functional fragment thereof. In yet other embodiments, the homologous endonuclease domain is modified, e.g., by site-directed mutagenesis, to alter DNA endonuclease activity. In yet other embodiments, the endonuclease domain is modified to remove any potential DNA sequence specificity.
[0667] In some embodiments, the endonuclease domain has nickase activity and does not form double-stranded breaks. In some embodiments, the endonuclease domain forms single-stranded breaks more frequently than double-stranded breaks, for example, at least 90%, 95%, 96%, 97%, 98%, or 99% of the cuts are single-stranded breaks, or less than 10%, 5%, 4%, 3%, 2%, or 1% of the cuts are double-stranded breaks. In some embodiments, the endonuclease does not substantially form double-stranded breaks. In some embodiments, the endonuclease does not form detectable levels of double-stranded breaks.
[0668] In some embodiments, the endonuclease domain has a nickase activity that nicks the target site DNA of the strand to be edited; for example, in some embodiments, the endonuclease domain cleaves the genomic DNA of the target site near the modification site on the strand to be extended by the writing domain. In some embodiments, the endonuclease domain has a nickase activity that nicks the target site DNA of the strand to be edited, but does not nick the target site DNA of the strand to be non-edited. For example, when a polypeptide includes a CRISPR-associated endonuclease domain that has nickase activity and does not form a double-strand break, in some embodiments, the CRISPR-associated endonuclease domain nicks the target site DNA strand that contains the PAM site (e.g., does not nick the target site DNA strand that does not contain the PAM site).
[0669] In some other embodiments, the endonuclease domain has nickase activity, which creates nicks in the target site DNA of the editing strand and the non-edited strand. Without intending to be bound by theory, after the writing domain (e.g., RT domain) of the polypeptide described herein polymerizes (e.g., reverse transcribes) from a heterologous target sequence of a template nucleic acid (e.g., template RNA), the cellular DNA repair machinery must repair the nicks on the editing DNA strand. The target site DNA here contains two distinct sequences relative to the editing DNA strand: one corresponding to the original genomic DNA and the second corresponding to that polymerized from the heterologous target sequence. It is believed that the two distinct sequences equilibrate with each other, hybridizing first one and then the other with the non-edited strand, whereby the cellular DNA repair machinery incorporates the nicks into the repair target site randomly. Without intending to be bound by theory, it is believed that the introduction of additional nicks into the non-edited strand may bias the cellular DNA repair machinery to use sequences based on the heterologous target sequence more frequently than the original genomic sequence. In some embodiments, the additional nick is positioned 5' or 3' of the target site modification (e.g., insertion, deletion, or substitution) or at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, or 150 nucleotides relative to the nick on the strand being edited.
[0670] Alternatively or additionally, without intending to be bound by any particular theory, it is believed that an additional nick in the unedited strand may facilitate second strand synthesis. In some embodiments, if the Gene Writer™ has inserted or replaced a portion of the edited strand, synthesis of a new sequence corresponding to the insertion / substitution in the unedited strand is required.
[0671] In some embodiments, the polypeptide comprises a single domain with endonuclease activity (e.g., a single endonuclease domain), which nicks both the edited strand and the non-edited strand. For example, in such embodiments, the endonuclease domain can be a CRISPR-associated endonuclease domain, and the template nucleic acid (e.g., template RNA) comprises a gRNA that directs nicking of the edited strand and an additional gRNA that directs nicking of the non-edited strand. In some embodiments, the polypeptide comprises multiple domains with endonuclease activity, where a first endonuclease domain nicks the edited strand and a second endonuclease domain nicks the non-edited strand (optionally, the first endonuclease domain does not (e.g., is unable to) nick the non-edited strand and the second endonuclease domain does not (e.g., is unable to) nick the edited strand).
[0672] In some embodiments, the endonuclease domain can nick the first and second strands. In some embodiments, the first and second strand nicks occur at the same position in the target site, but not on opposite strands. In some embodiments, the second strand nick occurs at a staggered position, e.g., upstream or downstream from the first nick. In some embodiments, the endonuclease domain generates a deletion of the target site when the second strand nick is upstream of the first strand nick. In some embodiments, the endonuclease domain generates a duplication of the target site when the second strand nick is downstream of the first strand nick. In some embodiments, the endonuclease domain does not generate a duplication and / or deletion when the first and second strand nicks occur at the same position in the target site (e.g., as described in Gladyshev and Arkhipova Gene 2009; incorporated herein by reference in its entirety). In some embodiments, the endonuclease domain has altered activity depending on the protein conformation or RNA binding state, for example, to promote first-strand or second-strand nicking (e.g., as described in Christensen et al. PNAS 2006; incorporated herein by reference in its entirety).
[0673] In some embodiments, the endonuclease domain comprises a meganuclease, or a functional fragment thereof. In some embodiments, the endonuclease domain comprises a homing endonuclease, or a functional fragment thereof. In some embodiments, the endonuclease domain comprises a meganuclease from the LAGLIDADG, GIY-YIG, HNH, His-Cys Box, or PD-(D / E)XK family, e.g., having a conserved amino acid motif, e.g., as indicated in the family name, or a functional fragment or variant thereof. In some embodiments, the endonuclease domain comprises a meganuclease, or a fragment thereof, selected from, for example, I-SmaMI (Uniprot F7WD42), I-SceI (Uniprot P03882), I-AniI (Uniprot P03880), I-DmoI (Uniprot P21505), I-CreI (Uniprot P05725), I-TevI (Uniprot P13299), I-OnuI (Uniprot Q4VWW5), or I-BmoI (Uniprot Q9ANR6). In some embodiments, the meganuclease is naturally a monomer, e.g., I-SceI, I-TevI, or a dimer, e.g., I-CreI, in its functional form. For example, LAGLIDADG meganucleases with a single copy of the LAGLIDADG motif generally form homodimers, while members with two copies of the LAGLIDADG motif are generally found as monomers. In some embodiments, meganucleases that normally form as dimers are expressed as fusions, e.g., two subunits are expressed as a single ORF, optionally linked by a linker, e.g., an I-CreI dimer fusion (Rodriguez-Fornes et al. Gene Therapy 2020; the entire contents of which are incorporated herein by reference).In some embodiments, meganucleases, or functional fragments thereof, are engineered to preferentially nickase activity on one strand of double-stranded DNA molecules, e.g., I-SceI (K122I and / or K223I) (Niu et al. J Mol Biol 2008), I-AniI (K227M) (McConnell Smith et al. PNAS 2009), I-DmoI (Q42A and / or K120M) (Molina et al. J Biol Chem 2015). In some embodiments, meganucleases or functional fragments thereof with this preference for single-strand cleavage are used, for example, as endonuclease domains with nickase activity. In some embodiments, the endonuclease domain comprises a meganuclease, or functional fragment thereof, that naturally targets, or has been engineered to target, a safe harbor site, e.g., an SH6 site that targets I-CreI (Rodriguez-Fornes). et al., supra). In some embodiments, the endonuclease domain comprises a meganuclease, or a functional fragment thereof, having a sequence-tolerant catalytic domain, such as I-TevI, which recognizes the minimal motif CNNNG (Kleinstiver et al. PNAS 2012). In some embodiments, the target sequence-tolerant catalytic domain is fused to a DNA-binding domain, e.g., I-TevI is induced to be active by fusing it to (i) a zinc finger to create Tev-ZFE (Kleinstiver et al. PNAS 2012), (ii) another meganuclease to create MegaTev (Wolfs et al. Nucleic Acids Res 2014), and / or (iii) Cas9 to create TevCas9 (Wolfs et al. PNAS 2016).
[0674] In some embodiments, the endonuclease domain comprises a restriction enzyme, e.g., a Type IIS or Type IIP restriction enzyme. In some embodiments, the endonuclease domain comprises a Type IIS restriction enzyme, e.g., FokI, or a fragment or variant thereof. In some embodiments, the endonuclease domain comprises a Type IIP restriction enzyme, e.g., PvuII, or a fragment or variant thereof. In some embodiments, the dimeric restriction enzyme is expressed as a fusion, e.g., a FokI dimer fusion, such that it functions as a single strand (Minczuk et al. Nucleic Acids Res 36(12):3926-3938 (2008)).
[0675] The use of additional endonuclease domains is described, for example, in Guha and Edgell Int J Mol Sci 18(22):2565 (2017), which is incorporated herein by reference in its entirety.
[0676] In some embodiments, the endonuclease domain or DNA-binding domain (e.g., as described herein) comprises Streptococcus pyogenes Cas9 (SpCas9) or a functional fragment or variant thereof. In some embodiments, the endonuclease domain or DNA-binding domain comprises a modified SpCas9. In some embodiments, the modified SpCas9 comprises a modification that alters protospacer adjacent motif (PAM) specificity. In some embodiments, the PAM has specificity for the nucleic acid sequence 5'-NGT-3'. In some embodiments, the modified SpCas9 comprises one or more amino acid substitutions, e.g., at one or more of the following positions: L1111, D1135, G1218, E1219, A1322, or R1335, e.g., selected from the following: L1111R, D1135V, G1218R, E1219F, A1322R, R1335V. In some embodiments, the modified SpCas9 comprises the amino acid substitution T1337R and one or more additional amino acid substitutions selected from the following: L1111, D1135L, S1136R, G1218S, E1219V, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, R1335Q, T1337, T1337L, T1337Q, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337H, T1337Q, and T1337M, or a corresponding amino acid substitution thereof. In some embodiments, the modified SpCas9 comprises: (i) one or more amino acid substitutions selected from the following: D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q, and T1337; and (ii) one or more additional amino acid substitutions selected from the following: L1111R, G1218R, E1219F, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, T1337L, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337R, T1337H, T1337Q, and T1337M, or a corresponding amino acid substitution thereof.
[0677] In some embodiments, the endonuclease domain or DNA-binding domain (e.g., as described herein) comprises a Cas domain, e.g., a Cas9 domain. In several embodiments, the endonuclease domain or DNA-binding domain comprises a nuclease-active Cas domain, a Cas nickase (nCas) domain, or a nuclease-inactive Cas (dCas). In several embodiments, the endonuclease domain or DNA-binding domain comprises a nuclease-active Cas9 domain, a Cas9 nickase (nCas9) domain, or a nuclease-inactive Cas9 (dCas). In some embodiments, the endonuclease domain or DNA binding domain comprises a Cas9 domain of Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA binding domain comprises Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA binding domain comprises S. pyogenes or S. thermophilus Cas9, or a functional fragment thereof. In some embodiments, the endonuclease domain or DNA binding domain comprises a Cas9 derived from a Streptococcus pyogenes (S. pyogenes) or S. thermophilus (S. thermophilus) ... Cas9, or a functional fragment thereof. In some embodiments, the endonuclease domain or DNA binding domain comprises a Cas9 derived from a Streptococcus pyogenes (S. pyogenes) or S. thermophilus Cas9, or a functional fragment thereof. In some embodiments, the endonuclease domain or DNA binding domain comprises a Cas9 derived from a Streptococcus pyogenes (S. pyogenes) or S. thermophilus Cas9, or a functional fragment thereof. In some embodiments, the endonuclease domain or DNA binding domain comprises a Cas9 derived from a Streptococcus pyogenes (S. pyogenes) or S. thermophilus Cas9, or a functional fragment thereof. In some embodiments, the 10:5,726-737 (incorporated herein by reference). In some embodiments, the endonuclease domain or DNA-binding domain comprises a Cas, e.g., the HNH nuclease subdomain and / or RuvC1 subdomain of Cas9, or a variant thereof, as described herein. In some embodiments, the endonuclease domain or DNA-binding domain comprises Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA-binding domain comprises a Cas polypeptide (e.g., an enzyme), or a functional fragment thereof.In some embodiments, the Cas polypeptide (e.g., enzyme) is selected from the following: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (e.g., Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpf l, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Cs y3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1 , Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, C The Cas9 may be selected from sa5, a type II Cas effector protein, a type V Cas effector protein, a type VI Cas effector protein, CARF, DinG, Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12b / C2c1, Cas12c / C2c3, SpCas9(K855A), eSpCas9(1.1), SpCas9-HF1, a hyper accurate Cas9 variant (HypaCas9), homologs thereof, modified or engineered versions thereof, and / or functional fragments thereof. In some embodiments, the Cas9 comprises one or more substitutions selected from the following, for example, H840A, D10A, P475A, W476A, N477A, D1125A, W1126A, and D1127A.In embodiments, the Cas9 comprises one or more mutations at a position selected from the following: D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987, for example, one or more substitutions selected from the following: D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A. In some embodiments, the endonuclease domain or DNA binding domain is selected from the group consisting of Corynebacterium ulcerans, Corynebacterium diphtheria, Spiroplasma syrphidicola, Prevotella intermedia, Spiroplasma taiwanense, Streptococcus iniae, Belliella baltica, Psychroflexus torquis, Staphylococcus thermophilus, Listeria innocua, Campylobacter jejuni, Neisseria meningitidis, and Escherichia coli. meningitidis, Streptococcus pyogenes, or Staphylococcus aureus, or functional fragments or variants thereof.
[0678] In some embodiments, an endonuclease domain or a DNA binding domain (e.g., as described herein) comprises a Cpf1 domain that includes one or more substitutions, e.g., at positions D917, E1006A, D1255, or any combination thereof, selected from the following: D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, and D917A / E1006A / D1255A.
[0679] In some embodiments, the endonuclease domain or DNA binding domain (e.g., as described herein) comprises spCas9, spCas9-VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL.
[0680] In some embodiments, an endonuclease domain or a DNA-binding domain (e.g., as described herein) comprises an amino acid sequence listed in Table 37 below, or an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the endonuclease domain or DNA binding domain comprises an amino acid sequence that has no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 differences (e.g., mutations) relative to any of the amino acid sequences described herein.
[0681] [Table 35]
[0682] [Table 36]
[0683] [Table 37]
[0684] [Table 38]
[0685] In some embodiments, the Gene Writing polypeptide has an endonuclease domain that includes a Cas9 nickase, e.g., Cas9H840A. In some embodiments, Cas9H840A has the following amino acid sequence: Cas9 Nickase (H840A): [ka]
[0686] In some embodiments, the Gene Writing polypeptide comprises an RT domain from a retroviral reverse transcriptase, such as, for example, wild-type M-MLV RT, comprising the following sequence: M-MLV(WT): [ka]
[0687] In some embodiments, the Gene Writing polypeptide comprises an RT domain from a retroviral reverse transcriptase, such as M-MLV RT, which comprises the following sequence: [ka]
[0688] In some embodiments, the Gene Writing polypeptide comprises an RT domain from a retroviral reverse transcriptase comprising the sequence of amino acids 659 to 1329 of NP_057933. In embodiments, the Gene Writing polypeptide further comprises one additional amino acid at the N-terminus of the sequence of amino acids 659 to 1329 of NP_057933, e.g., as shown below. [ka] Core RT (bold), annotations above Ribonuclease H (underlined), annotation as above
[0689] In some embodiments, the Gene Writing polypeptide further comprises one additional amino acid at the C-terminus of the sequence of amino acids 659-1329 of NP_057933. In some embodiments, the Gene Writing polypeptide comprises a ribonuclease H1 domain (e.g., amino acids 1178-1318 of NP_057933).
[0690] In some embodiments, a retroviral reverse transcriptase domain, e.g., M-MLV RT, may contain one or more mutations from the wild-type sequence that may improve characteristics of the RT, e.g., thermostability, processivity, and / or template binding. In some embodiments, the M-MLV RT domain comprises one or more mutations selected from D200N, L603W, T330P, T306K, W313F, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, L435G, N454K, H594Q, D653N, R110S, K103L relative to the M-MLV(WT) sequence above, e.g., a combination of mutations such as D200N, L603W, and T330P, optionally further comprising T306K and W313F. In some embodiments, an M-MLV RT as used herein comprises the mutations D200N, L603W, T330P, T306K, and W313F. In embodiments, the mutant M-MLV RT comprises the following amino acid sequence: M-MLV(PE2): [ka]
[0691] In some embodiments, a Gene Writer polypeptide may include a linker, e.g., a peptide linker, such as those described in Table 7. In some embodiments, a Gene Writer polypeptide includes a flexible linker between the endonuclease and RT domain, e.g., a linker comprising the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSS. In some embodiments, the RT domain of a Gene Writer polypeptide may be located C-terminal to the endonuclease domain. In some embodiments, the RT domain of a Gene Writer polypeptide may be located N-terminal to the endonuclease domain.
[0692] [Table 39]
[0693] [Table 40]
[0694] [Table 41]
[0695] [Table 42]
[0696] In some embodiments, the Gene Writer polypeptide comprises a dCas9 sequence comprising a D10A and / or H840A mutation, for example, the sequence: [ka]
[0697] In some embodiments, a template RNA molecule for use in the system comprises, from 5' to 3', (1) a gRNA spacer; (2) a gRNA scaffold; (3) a heterologous sequence of interest; and (4) a 3' homology domain. (1) A Cas9 spacer of about 18 to 22 nt, for example, 20 nt. (2) A gRNA scaffold comprising one or more hairpin loops, e.g., one, two, or three loops, for associating the template with the nickase Cas9 domain. In some embodiments, the gRNA scaffold has the following sequence from 5' to 3': GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCC. (3) In some embodiments, the heterologous sequence of interest is, for example, 7 to 74 nt in length, e.g., 10 to 20 nt, 20 to 30 nt, 30 to 40 nt, 40 to 50 nt, 50 to 60 nt, 60 to 70 nt, or 70 to 80 nt, or 80 to 90 nt. In some embodiments, the first (approximately 5') base of the sequence is not C. (4) In some embodiments, the 3' homology domain that binds to the target priming sequence after nicking is, e.g., 3 to 20 nt, e.g., 7 to 15 nt, e.g., 12 to 14 nt, In some embodiments, the 3' homology domain has a GC content of 40 to 60%.
[0698] A second gRNA associated with the system can help drive complete integration. In some embodiments, the second gRNA can target a position 0-200 nt away from the first-strand nick, e.g., 0-50 nt, 50-100 nt, or 100-200 nt away from the first-strand nick. In some embodiments, the second gRNA can only bind to its target sequence after editing has occurred, e.g., the gRNA binds to a sequence present in the heterologous sequence of interest but not in the initial target sequence.
[0699] In some embodiments, the Gene Writing System described herein is used to perform edits in HEK293, K562, U2OS, or HeLa cells. In some embodiments, the Gene Writing System is used to perform edits in primary cells, such as primary cortical neurons from E18.5 mice.
[0700] In some embodiments, the reverse transcriptase or RT domain (e.g., as described herein) comprises a MoMLV RT sequence or a variant thereof. In embodiments, the MoMLV RT sequence comprises one or more mutations selected from D200N, L603W, T330P, T306K, W313F, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, L435G, N454K, H594Q, D653N, R110S, and K103L. In embodiments, the MoMLV RT sequence comprises a combination of mutations, such as D200N, L603W, and T330P, optionally further comprising T306K and / or W313F.
[0701] In some embodiments, the endonuclease domain (e.g., as described herein) comprises nCAS9, e.g., comprising an H840A mutation.
[0702] In some embodiments, a heterologous sequence of interest (e.g., in a system described herein) is about 1-50, about 50-100, about 100-200, about 200-300, about 300-400, about 400-500, about 500-600, about 600-700, about 700-800, about 800-900, about 900-1000, or more nucleotides in length.
[0703] In some embodiments, the RT and endonuclease domains are linked by a flexible linker, for example, comprising the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSS.
[0704] In some embodiments, the endonuclease domain is N-terminal to the RT domain. In some embodiments, the endonuclease domain is C-terminal to the RT domain.
[0705] In some embodiments, the system incorporates a heterologous sequence of interest into a target site by TPRT, for example, as described herein.
[0706] In some embodiments, a system or method described herein comprises a CRISPR DNA targeting enzyme or system, or a functional fragment or variant thereof, described in U.S. Patent Application Publication No. 20200063126, U.S. Patent Application Publication No. 20190002889, or U.S. Patent Application Publication No. 20190002875 (each of which is incorporated by reference herein in its entirety). For example, in some embodiments, a Gene Writer polypeptide or Cas endonuclease described herein comprises a polypeptide sequence described in any of the applications listed in this paragraph, and in some embodiments, a template RNA or guide RNA comprises a nucleic acid sequence described in any of the applications listed in this paragraph.
[0707] Template nucleic acid binding domain: Gene Writer polypeptides typically include a region capable of binding to a Gene Writer template nucleic acid (e.g., template RNA). In some embodiments, the template nucleic acid binding domain is an RNA-binding domain. In some embodiments, the RNA-binding domain is a modular domain capable of binding to RNA molecules containing a specific signature, e.g., a structural motif, e.g., a secondary structure present in the 3'UTR of non-LTR retrotransposons. In other embodiments, the template nucleic acid binding domain (e.g., RNA-binding domain) is contained within a reverse transcription domain, e.g., a reverse transcriptase-derived component having a known signature for RNA selection, e.g., a secondary structure present in the 3'UTR of non-LTR retrotransposons. In other embodiments, the template nucleic acid binding domain (e.g., RNA-binding domain) is contained within a DNA-binding domain. For example, in some embodiments, the DNA-binding domain is a CRISPR-associated protein that recognizes the structure of a template nucleic acid (e.g., template RNA) containing a gRNA. In some embodiments, gRNAs are short synthetic RNAs composed of a scaffold sequence involved in CRISPR-associated protein binding and a user-defined targeting sequence of approximately 20 nucleotides for a genomic target. The structure of a complete gRNA was described in Nishimasu et al., Cell 156, pp. 935-949 (2014). A gRNA (also called an sgRNA for single-guide RNA) consists of a crRNA and a tracrRNA-derived sequence linked by an artificial tetraloop. The crRNA sequence can be divided into a guide (20 nt) region and a repeat (12 nt) region, while the tracrRNA sequence can be divided into an anti-repeat (14 nt) and three tracrRNA stem loops (Nishimasu et al., Cell 156, pp. 935-949 (2014)). In practice, guide RNA sequences are generally 17 to 24 nucleotides (e.g., 19, 20, or 21 nucleotides) in length and are designed to be complementary to the target nucleic acid sequence. Custom gRNA generators and algorithms are commercially available for use in designing effective guide RNAs.In some embodiments, the gRNA comprises two RNA components from the native CRISPR system, e.g., crRNA and tracrRNA. As is well known in the art, the gRNA can also comprise a chimeric single guide RNA (sgRNA) that contains sequences from both the tracrRNA (for binding to the nuclease) and at least one crRNA (for directing the nuclease to the targeted sequence for editing / binding). Chemically modified sgRNAs have also been demonstrated to be effective for use with CRISPR-associated proteins (see, e.g., Hendel et al. (2015) Nature Biotechnol., 985-991). In some embodiments, the gRNA comprises a nucleic acid sequence complementary to a DNA sequence associated with a target gene. In some embodiments, the polypeptide comprises a DNA-binding domain comprising a CRISPR-associated protein that binds to the gRNA, allowing the DNA-binding domain to bind to the target genomic DNA sequence. In some embodiments, the gRNA is contained within a template nucleic acid (e.g., template RNA), and thus the DNA-binding domain is also a template nucleic acid-binding domain. In some embodiments, the polypeptide has RNA-binding functions in multiple domains, for example, it can bind to a gRNA structure in the CRISPR-associated DNA-binding domain and a 3'UTR structure in the reverse transcription domain derived from a non-LTR retrotransposon.
[0708] In some embodiments, the template nucleic acid (e.g., template RNA) comprises a 3' target homology domain. In some embodiments, the 3' target homology domain is located 3' to the heterologous sequence of interest and is complementary to a sequence adjacent to the site to be modified by the system described herein, or contains no more than 1, 2, 3, 4, or 5 mismatches to a sequence complementary to a sequence adjacent to the site to be modified by the system / Gene Writer™. In some embodiments, the 3' target homology domain anneals to the target site and provides a binding site and a 3' hydroxyl for initiation of TPRT by the Gene Writer polypeptide. In some embodiments, the 3' target homology domain is 3 to 5, 5 to 10, 10 to 30, 10 to 25, 10 to 20, 10 to 19, 10 to 18, 10 to 17, 10 to 16, 10 to 15, 10 to 14, 10 to 13, 10 to 12, 10 to 11, 11 to 30, 11 to 25, 11 to 20, 11 to 19, 11 to 1 8, 11-17, 11-16, 11-15, 11-14, 11-13, 11-12, 12-30, 12-25, 12-20, 12-19, 12-18, 12-17, 12-16, 12-15, 12-14, 12-13, 13-30, 13-25, 13-20, 13-19, 13-18, 13-17, 13 ~16, 13~15, 13~14, 14~30, 14~25, 14~20, 14~19, 14~18, 14~17, 14~16, 14~15, 15~30, 15~25, 15~20, 15~19, 15~18, 15~17, 15~16, 16~30, 16~25, 16~20, 16~19, 16~18, 16-17, 17-30, 17-25, 17-20, 17-19, 17-18, 18-30, 18-25, 18-20, 18-19, 19-30, 19-25, 19-20, 20-30, 20-25, or 25-30 nt, for example, 10-17, 12-16, or 12-14 nt in length.
[0709] In some embodiments, the template nucleic acid, e.g., template RNA, may comprise a gRNA (e.g., pegRNA). In some embodiments, the template nucleic acid, e.g., template RNA, may bind to a Gene Writer™ polypeptide through interaction between the gRNA of the template nucleic acid and a template nucleic acid binding domain, e.g., an RNA binding domain (e.g., a heterologous RNA binding domain). In some embodiments, the heterologous RNA binding domain is a CRISPR / Cas protein, e.g., Cas9.
[0710] In some embodiments, a template nucleic acid, e.g., a region of the template RNA, comprising a guide RNA (gRNA) adopts an underwound ribbon-like structure of the gRNA bound to a target DNA (e.g., as described in Mulepati et al., Science 19 Sep 2014: Vol. 345, Issue 6203, pp. 1479-1484). Without intending to be bound by any particular theory, it is believed that this non-canonical structure is facilitated by rotations every six nucleotides from the RNA-DNA hybrid. Thus, in some embodiments, a template nucleic acid, e.g., a region of the template RNA, comprising a gRNA, can tolerate increased mismatches with the target site at some intervals, e.g., every six bases. In some embodiments, a template nucleic acid, e.g., a region of the template RNA, comprising a gRNA that contains homology to a target site may have wobble positions at regular intervals, e.g., every six bases, that do not require base pairing with the target site.
[0711] Inducible gRNA In some embodiments, the template nucleic acid, e.g., template RNA, comprises an inducibly active gRNA. Inducible activity can be achieved by a template nucleic acid, e.g., template RNA, further comprising a blocking domain (in addition to the gRNA), where the sequence of some or all of the blocking domain is at least partially complementary to some or all of the gRNA. Thus, the blocking domain is hybridizable or substantially hybridizable to some or all of the gRNA. In some embodiments, the blocking domain and the inducibly active gRNA are disposed on the template nucleic acid, e.g., template RNA, such that the gRNA can adopt a first conformation in which the blocking domain is hybridized or substantially hybridized to the gRNA, and a second conformation in which the blocking domain is not hybridized or substantially hybridized to the gRNA. In some embodiments, in the first conformation, the gRNA is unable to bind to a Gene Writer polypeptide (e.g., a template nucleic acid binding domain, a DNA binding domain, or an endonuclease domain (e.g., a CRISPR / Cas protein)) or binds with a substantially reduced affinity compared to a similar template RNA that lacks an additional blocking domain. In some embodiments, in the second conformation, the gRNA is able to bind to a Gene Writer polypeptide (e.g., a template nucleic acid binding domain, a DNA binding domain, or an endonuclease domain (e.g., a CRISPR / Cas protein)). In some embodiments, whether the gRNA is in the first or second conformation can affect whether the DNA-binding or endonuclease activity of the Gene Writer polypeptide (e.g., of the CRISPR / Cas protein that the Gene Writer polypeptide comprises) is active. In some embodiments, hybridization of the gRNA to the blocking domain can be disrupted using an opening molecule. In some embodiments, the aperture molecule comprises an agent that binds to part or all of the gRNA or the blocking domain and inhibits hybridization of the gRNA to the blocking domain.In some embodiments, the aperture molecule comprises a nucleic acid comprising a sequence that is partially or entirely complementary to, for example, the gRNA, the blocking domain, or both. By selecting or designing an appropriate aperture molecule, provision of the aperture molecule can promote a change in the conformation of the gRNA such that it can associate with the CRISPR / Cas protein and provide the relevant function of the CRISPR / Cas protein (e.g., DNA binding and / or endonuclease activity). Without intending to be bound by any particular theory, provision of the aperture molecule at a selected time and / or location can enable spatial and temporal control of the activity of the gRNA, the CRISPR / Cas protein, or the Gene Writer system comprising them. In some embodiments, the Gene Writer may comprise a Cas protein or functional fragment thereof, such as those listed in Table 9 or Table 37, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. [Table 43]
[0712] [Table 44]
[0713] [Table 45]
[0714] [Table 46]
[0715] [Table 47]
[0716] [Table 48]
[0717] [Table 49]
[0718] [Table 50]
[0719] [Table 51]
[0720] [Table 52]
[0721] [Table 53]
[0722] Table 9B defines the components necessary to design a gRNA and / or template RNA and provides parameters for applying the Cas variants listed in Table 3A for Gene Writing. The tier indicates the preferred Cas variant, if available, for use at a given locus. The cleavage site indicates the requirement for a validated or predicted protospacer adjacent motif (PAM), and the location of the validated or predicted cleavage site (relative to the most upstream base of the PAM site). A gRNA for a given enzyme can be constructed by concatenating the crRNA, tetraloop, and tracrRNA sequences and adding a 5' spacer within the minimum and maximum lengths of the spacer that matches the protospacer at the target site. Furthermore, the predicted location of the ssDNA nick at the target is important for designing the 3' region of the template RNA, which must anneal to the sequence immediately 5' of the nick to initiate target-primed reverse transcription.
[0723] [Table 54]
[0724] Table 55
[0725] Table 56
[0726] Table 57
[0727] Table 58
[0728] In some embodiments, the aperture molecule is exogenous to the cell comprising the Gene Writer polypeptide and / or template nucleic acid. In some embodiments, the aperture molecule comprises an endogenous agent (e.g., endogenous to the cell comprising the Gene Writer polypeptide and / or template nucleic acid comprising a gRNA and a blocking domain). For example, the inducible gRNA, blocking domain, and aperture molecule can be selected such that the aperture molecule is an endogenous agent expressed in the target cell or tissue, e.g., thereby confirming activity of the Gene Writer system in the target cell or tissue. As a further example, the inducible gRNA, blocking domain, and aperture molecule can be selected such that the aperture molecule is absent or substantially not expressed in one or more non-target cells or tissues, e.g., thereby confirming activity of the Gene Writer system is absent or substantially absent, or present at a reduced level relative to the target cell or tissue, in one or more non-target cells or tissues. Exemplary blocking domains, aperture molecules, and their uses are described in PCT Publication WO 2020044039A1, which is incorporated herein by reference in its entirety. In some embodiments, a template nucleic acid, e.g., a template RNA, may include one or more UTRs (e.g., from an R2-type retrotransposon) and a gRNA. In some embodiments, the UTRs facilitate interaction of the template nucleic acid (e.g., template RNA) with a writing domain, e.g., a reverse transcriptase domain, of a Gene Writer polypeptide. In some embodiments, the gRNA facilitates interaction with a template nucleic acid binding domain (e.g., an RNA binding domain) of the polypeptide. In some embodiments, the gRNA guides the polypeptide to a matching target sequence, e.g., in the target cell genome. In some embodiments, the template nucleic acid may contain only a reverse transcriptase binding motif (e.g., a 3' UTR from R2), and the gRNA may be provided as a second nucleic acid molecule (e.g., a second RNA molecule) for target site recognition.In some embodiments, the template nucleic acid containing the RT binding motif may be present on the same molecule as the gRNA, but may be processed into two RNA molecules by a cleavage activity (e.g., a ribozyme).
[0729] In some embodiments, the template RNA may be customized to correct a given mutation in the genomic DNA of a target cell (e.g., ex vivo or in vivo, e.g., within a subject, e.g., within a target tissue or organ). For example, the mutation may be a disease-associated mutation relative to a wild-type sequence. Without intending to be bound by any particular theory, an empirical parameter set can help identify optimal initial parameters in the in silico design of the template RNA or a portion thereof. By way of non-limiting example, the following design parameters may be utilized for a selected mutation: In some embodiments, the design begins by obtaining about 500 bp (e.g., up to 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, or 700 bp, and optionally at least 20, 30, 40, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, or 650 bp) of flanking sequence on either side of the mutation to serve as the target region. In some embodiments, the template nucleic acid comprises a gRNA. Methods for designing gRNAs are known to those of skill in the art. In some embodiments, the gRNA comprises a sequence that binds to the target site (e.g., a CRISPR spacer). In some embodiments, target site-binding sequences (e.g., CRISPR spacers) for use in targeting a template nucleic acid to a target region are selected by considering the use of a particular Gene Writer polypeptide (e.g., comprising an endonuclease domain or writing domain, e.g., a CRISPR / Cas domain) (e.g., in the case of Cas9, a protospacer adjacent motif (PAM) of NGG imme...
Claims
[Claim 1] The invention described in this specification.