Improved methods and compositions for regulating genomes
The use of polypeptides with reverse transcriptase and endonuclease domains, combined with template RNAs, addresses the challenge of site-specific genome editing, enabling efficient insertion and regulation of genomic sequences up to 1,000 nucleotides.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FLAGSHIP PIONEERING INNOVATIONS VI LLC
- Filing Date
- 2021-03-04
- Publication Date
- 2026-05-21
AI Technical Summary
Existing methods for inserting target nucleic acids into genomes lack site specificity and efficiency, particularly in the absence of specialized proteins, and are less effective for longer sequences.
Compositions and systems involving polypeptides with reverse transcriptase and endonuclease domains, along with template RNAs, are used to modify genomic DNA by inserting, deleting, or substituting nucleotides, utilizing specific binding sequences and heterologous target sequences for precise genome editing.
Achieves site-specific and efficient insertion or modification of genomic sequences, including therapeutic polypeptides and non-coding RNAs, with the ability to regulate gene expression and introduce heterologous sequences up to 1,000 nucleotides.
Smart Images

Figure 0007863508001319 
Figure 0007863508001320 
Figure 0007863508001321
Abstract
Description
[Technical Field]
[0001] Related applications This application claims priority to U.S. Patent Application No. 62 / 985,264, filed on 4 March 2020, and U.S. Patent Application No. 63 / 035,674, filed on 5 June 2020, the entire contents of each of those applications being incorporated herein by reference. [Background technology]
[0002] The insertion of target nucleic acids into the genome occurs infrequently and with little site specificity, especially in the absence of specialized proteins to facilitate insertion events. Some existing methods, such as CRISPR / Cas9, are better suited to small edits and less effective for inserting longer sequences. Other existing methods, such as Cre / loxP, require a first step of inserting a loxP site into the genome, followed by a second step of inserting the target sequence into the loxP site. There is a need in this art for improved methods of inserting target sequences into proteins and genomes. [Overview of the project] [Means for solving the problem]
[0003] This disclosure relates to novel compositions, systems, and methods for modifying the genome at one or more locations in a host cell, tissue, or subject in vivo or in vitro. In particular, the present invention features compositions, systems, and methods for introducing exogenous genetic factors into a host genome. This disclosure also provides systems for modifying a target genomic DNA sequence, for example, by inserting, deleting, or substituting one or more nucleotides into / from the target sequence.
[0004] The features of the composition or method may include one or more of the embodiments listed below.
[0005] 1. (a) polypeptides or nucleic acids encoding polypeptides (e.g., DNA or mRNA) (wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table Z1 or Table X) and (ii) an endonuclease domain, where one or both of (i) and (ii) have an amino acid sequence encoded by the nucleic acid sequence of the element in Table 3B, Table 10, Table 11 or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto); and (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. A system for modifying DNA, including [specific DNA components].
[0006] 2. (a) polypeptides or nucleic acids encoding polypeptides (e.g., DNA or mRNA) (wherein a polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table Z1 or Table X) and (ii) an endonuclease domain, and one or both of (i) and (ii) have an amino acid sequence encoded by a nucleic acid sequence of an element in Table 3B, Table 10, Table 11 or Table X, or a sequence in which there are 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 or fewer nucleotides different therefrom); and (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. A system for modifying DNA, including [specific DNA components].
[0007] 3. The heterologous target sequence is a system of any of the preceding embodiments that encodes a therapeutic polypeptide or a mammalian (e.g., human) polypeptide, or a fragment or variant thereof.
[0008] 4. The heterologous target sequence is a system of any of the preceding embodiments that encodes therapeutic non-coding RNA (e.g., miRNA).
[0009] 5. A system of any of the preceding embodiments in which the heterologous target sequence includes, for example, a regulatory sequence that modifies the expression of an endogenous gene or non-coding RNA (e.g., a promoter, enhancer, binding site for an endogenous regulatory component, e.g., a miRNA binding site).
[0010] 6. The regulatory sequence is a system of any of the preceding embodiments that results in the upregulation of an endogenous gene or non-coding RNA.
[0011] 7. The regulatory sequence is a system of any of the preceding embodiments that results in the downregulation of an endogenous gene or non-coding RNA.
[0012] 8. A system for modifying DNA, (a) polypeptides or nucleic acids encoding polypeptides, wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA-binding domain (DBD); and (iii) an endonuclease domain, such as a niccasse domain; and (b) (e.g., from 5' to 3') (i) optionally, a sequence that binds to a target site (e.g., the second strand of the site in the target genome), (ii) optionally, a sequence that binds to a polypeptide, (iii) a heterologous target sequence, and (iv) a template RNA (or DNA encoding the template RNA) containing a 3' target homology domain. Includes; (i) The polypeptide contains a heterologous targeting domain (e.g., within a DBD or endonuclease domain) that specifically binds to a sequence contained within the target site; and / or (ii) A system in which the template RNA contains heterologous sequences that have at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% homology to sequences contained within the target site.
[0013] 9. (a) polypeptides or nucleic acids encoding polypeptides (e.g., DNA or mRNA) (wherein a polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table Z1 or Table X) and (ii) an endonuclease domain, where one or both of (i) and (ii) have an amino acid sequence of an element in Table 3B, Table 10, Table 11 or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto); and (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. A system for modifying DNA, including [specific DNA components].
[0014] 10. (a) polypeptides or nucleic acids encoding polypeptides (e.g., DNA or mRNA) (wherein a polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table Z1, or Table X) and (ii) an endonuclease domain, wherein one or both of (i) and (ii) have a sequence in which there are 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides that are different from the amino acid sequence of an element in Table 3B, Table 10, Table 11, or Table X); and (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. A system for modifying DNA, including [specific DNA components].
[0015] 11. (a) polypeptides or nucleic acids encoding polypeptides (e.g., DNA or mRNA) (wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table Z1 or Table X) and (ii) a target DNA binding domain, where one or both of (i) and (ii) have an amino acid sequence encoded by a nucleic acid sequence of an element in Table 3B, Table 10, Table 11 or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto); and (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. A system for modifying DNA, including [specific DNA components].
[0016] 12. (a) polypeptides or nucleic acids encoding polypeptides (e.g., DNA or mRNA) (wherein a polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table Z1 or Table X) and (ii) a target DNA-binding domain, wherein one or both of (i) and (ii) have an amino acid sequence encoded by a nucleic acid sequence of an element in Table 3B, Table 10, Table 11 or Table X, or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 or fewer different nucleotides); and (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. A system for modifying DNA, including [specific DNA components].
[0017] 13. (a) polypeptides or nucleic acids encoding polypeptides (e.g., DNA or mRNA) (wherein the polypeptide includes (i) a reverse transcriptase domain derived from a protein other than a retrotransposase, for example, a reverse transcriptase domain listed in Table Z1 or Z2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and (ii) an endonuclease domain and / or a target DNA binding domain); and, (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. A system for modifying DNA, including [specific DNA components].
[0018] 14. (a) polypeptides or nucleic acids encoding polypeptides (e.g., DNA or mRNA) (wherein a polypeptide includes (i) a reverse transcriptase domain derived from a protein other than a retrotransposase, for example, a reverse transcriptase domain listed in Table Z1 or Z2, or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 or fewer nucleotides different therefrom, and (ii) an endonuclease domain and / or a target DNA binding domain); and, (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. A system for modifying DNA, including [specific DNA components].
[0019] 15. (a) polypeptides or nucleic acids encoding polypeptides (wherein the polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table Z1 or Z2), and (ii) an endonuclease domain and / or a target DNA binding domain); and, (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. Includes; A system for modifying DNA, wherein the sequence of the template RNA bound to the polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the 5'UTR or 3'UTR of the element sequence in Table 10 or Table X.
[0020] 16. (a) polypeptides or nucleic acids encoding polypeptides (wherein the polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table Z1 or Z2), and (ii) an endonuclease domain and / or a target DNA binding domain); and, (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. Includes; A DNA modification system wherein the sequence of the template RNA that binds to the polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with (i) the nucleotide located 5' to the start codon of the sequence of the element in Table 10 or Table X (e.g., including the retrotransposase binding region) or (ii) the nucleotide located 3' to the stop codon of the sequence of the element in Table 10 or Table X (e.g., including the retrotransposase binding region).
[0021] 17. (a) polypeptides or nucleic acids encoding polypeptides (e.g., DNA or mRNA) (wherein the polypeptide comprises (i) a reverse transcriptase domain (e.g., as listed in Table 3B, Table 10, Table 11, Table Z1 or Z2 or Table X), and (ii) an endonuclease domain and / or a target DNA binding domain); (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA encoding template RNA) containing a heterologous target sequence; and (c) Intein A system for modifying DNA, including [specific DNA components].
[0022] 18. A system of any of the preceding embodiments, wherein the polypeptide comprises an intein.
[0023] 19. A system of any of the preceding embodiments in which the intein is split intein.
[0024] 19. A system for modifying DNA, (a) a polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA-binding domain (DBD), and (iii) an endonuclease domain, such as a niccasse domain; and (b) (e.g., from 5' to 3') (i) optionally, a sequence that binds to a target site (e.g., an unedited strand of the site in the target genome) (e.g., a CRISPR spacer), (ii) optionally, a sequence that binds to a polypeptide, (iii) a heterologous target sequence, and (iv) a template RNA (or DNA encoding the template RNA) containing a 3' target homology domain. A system that includes this.
[0025] 21. (a) polypeptides or nucleic acids encoding polypeptides (wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA-binding domain (DBD), and (iii) an endonuclease domain, e.g., a nickas domain); and (b) (e.g., from 5' to 3') (i) an optional sequence that binds to a target site (e.g., an unedited strand of a site in the target genome), (ii) an optional sequence that binds to a polypeptide, (iii) a heterologous target sequence, and (iv) a template RNA (or DNA encoding the template RNA) containing a 3' homology domain. Includes; An RT domain is a system for modifying DNA, having an amino acid sequence in Table 3B, Table 10, Table 11, or Table X, or a sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to such sequences.
[0026] 22. (a) polypeptides or nucleic acids encoding polypeptides (wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA-binding domain (DBD), and (iii) an endonuclease domain, e.g., a nickas domain; and (b) (e.g., from 5' to 3') (i) an optional sequence that binds to a target site (e.g., an unedited strand of a site in the target genome), (ii) an optional sequence that binds to a polypeptide, (iii) a heterologous target sequence, and (iv) a template RNA (etRNA) (or DNA encoding the template RNA) containing a 3' homology domain. Includes; A system for modifying DNA that can generate insertions of at least 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides into target sites.
[0027] 23. (a) polypeptides or nucleic acids encoding polypeptides (wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA-binding domain (DBD), and (iii) an endonuclease domain, e.g., a nickas domain); and (b) (e.g., from 5' to 3') (i) an optional sequence that binds to a target site (e.g., an unedited strand of a site in the target genome), (ii) an optional sequence that binds to a polypeptide, (iii) a heterologous target sequence, and (iv) a template RNA (or DNA encoding the template RNA) containing a 3' homology domain. Includes; A system for modifying DNA, wherein the heterologous target sequence is at least 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 120, 140, 160, 180, 200, 500, or 1,000 nt in length.
[0028] 24. Any system of the preceding embodiments, wherein the RT domain is heterogeneous with respect to the DBD; one or more DBDs are heterogeneous with respect to the endonuclease domain; and / or one or more RT domains are heterogeneous with respect to the endonuclease domain.
[0029] 25. (a) polypeptides or nucleic acids encoding polypeptides (wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA-binding domain (DBD), and (iii) an endonuclease domain, e.g., a nickasase domain); and (b) (e.g., from 5' to 3') (i) an optional sequence that binds to a target site (e.g., an unedited strand of a site in the target genome), (ii) an optional sequence that binds to a polypeptide, (iii) a heterologous target sequence, and (iv) a template RNA (or DNA encoding the template RNA) containing a 3' homology domain. Includes; A system for modifying DNA that can cause deletions at target sites of at least 81, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides.
[0030] 26. (a) polypeptides or nucleic acids encoding polypeptides (wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA-binding domain (DBD), and (iii) an endonuclease domain, e.g., a nickas domain); and (b) (e.g., from 5' to 3') (i) an optional sequence that binds to a target site (e.g., an unedited strand of a site in the target genome), (ii) an optional sequence that binds to a polypeptide, (iii) a heterologous target sequence, and (iv) a template (or DNA encoding template RNA) containing a 3' homology domain. Includes; (a)(ii) and / or (a)(iii) are systems for modifying DNA, comprising a TALE molecule; a zinc finger molecule; or a CRISPR / Cas molecule or a functional variant thereof (e.g., a mutant) selected from Table 1.
[0031] 27. (a) polypeptides or nucleic acids encoding polypeptides (wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA-binding domain (DBD), and (iii) an endonuclease domain, e.g., a nickas domain); and (b) (e.g., from 5' to 3') (i) an optional sequence that binds to a target site (e.g., a CRISPR spacer) (e.g., an unedited strand of a site in the target genome), (ii) an optional sequence that binds to a polypeptide, (iii) a heterologous target sequence, and (iv) a template RNA (or DNA encoding the template RNA) containing a 3' homology domain. Includes; An endonuclease domain, such as a nickerse domain, is a system for modifying DNA in which both strands of target site DNA are cleaved, and the cleavage is separated from each other by at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 nucleotides.
[0032] 28. (a) polypeptides or nucleic acids encoding polypeptides (wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA-binding domain (DBD), and (iii) an endonuclease domain, e.g., a nickas domain); and (b) (e.g., from 5' to 3') (i) a sequence that binds to a target site by choice (e.g., an unedited strand of a site in the target genome), (ii) a sequence that specifically binds to the RT domain, (iii) a heterologous target sequence, and (iv) a template RNA (or DNA encoding the template RNA) containing a 3' homology domain. A system for modifying DNA, including [specific DNA components].
[0033] 29. A system according to any of the preceding embodiments, wherein the template RNA further comprises sequences that bind to (a)(ii) and / or (a)(iii).
[0034] 30. A system for modifying DNA, (a) A first polypeptide or nucleic acid encoding a first polypeptide, wherein the first polypeptide comprises (i) a reverse transcriptase (RT) domain and (ii) optionally a DNA-binding domain, (b) a second polypeptide or a nucleic acid encoding a second polypeptide, the second polypeptide comprising (i) a DNA-binding domain (DBD); (ii) an endonuclease domain, such as a niccasse domain; and (c) (e.g., from 5' to 3') (i) optionally a sequence that binds to a second polypeptide (e.g., (b)(i) and / or (b)(ii)), (ii) optionally a sequence that binds to a first polypeptide (e.g., specifically to the RT domain), (iii) a heterologous target sequence, and (iv) a template RNA (or DNA encoding template RNA) containing a 3' homology domain. A system that includes this.
[0035] 31. A system for modifying DNA, (a) a polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain and (ii) a DNA-binding domain (DBD); and (iii) an endonuclease domain, such as a niccasse domain; (b) A first template RNA (or RNA-coding DNA) (for example, where the first RNA includes gRNA) comprising (i) a sequence that binds to a polypeptide (e.g., from 5' to 3') (i) a sequence that binds to (a)(ii) and / or (a)(iii) and (ii) a sequence that binds to a target site (e.g., an unedited strand of the site in the target genome); (c) (e.g., from 5' to 3') (i) optionally, a sequence that binds to the polypeptide (e.g., specifically to the RT domain), (ii) a heterologous target sequence, and (iii) a second template RNA (or RNA-coding DNA) containing a 3' target homology domain. A system that includes this.
[0036] 32. A system according to any of the preceding embodiments, wherein the second template RNA comprises (i).
[0037] 33. A system according to any of the preceding embodiments, wherein the first template RNA comprises a first conjugate domain and the second template RNA comprises a second conjugate domain.
[0038] 34. A system of any of the preceding embodiments in which the first and second conjugate domains can hybridize with each other, for example, under stringent conditions.
[0039] 35. A system according to any of the preceding embodiments, wherein the association of the first conjugate domain and the second conjugate domain causes colocalization of the first template RNA and the second template RNA.
[0040] 36. The template RNA is a system of any of the preceding embodiments, comprising (i).
[0041] 37. The template RNA is a system of any of the preceding embodiments, comprising (ii).
[0042] 38. The template RNA is a system of either of the preceding embodiments, comprising (i) and (ii).
[0043] 39. A system for modifying DNA, (a) A first polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide is a reverse transcriptase (RT) domain having a sequence of the reverse transcriptase domains of Table 3B, Table 10, Table 11 or Table X or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and optionally comprising a DNA-binding domain (DBD) (e.g., a first DBD); and (b) A second polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a DBD (e.g., a second DBD); and (ii) an endonuclease domain, e.g., a niccasse domain. A system that includes this.
[0044] 40. A system of any of the previous embodiments in which the nucleic acid encoding the first polypeptide and the nucleic acid encoding the second polypeptide are two separate nucleic acids.
[0045] 41. A system according to any of the previous embodiments, in which the nucleic acid encoding the first polypeptide and the nucleic acid encoding the second polypeptide are parts of the same nucleic acid molecule and, for example, reside on the same vector.
[0046] 42. A system of any of the preceding embodiments having one or more of the following features (e.g., 1, 2, 3, 4, 5, 6, or all of them): i. The heterologous target sequence encodes a protein, for example, an enzyme (e.g., a lysosomal enzyme) or a blood factor (e.g., factor I, II, V, VII, X, XI, XII, or XIII); ii. The heterologous target sequence includes a tissue-specific promoter or enhancer; iii. The heterologous target sequence encodes a polypeptide consisting of 50, 100, 150, 200, 250, 300, 400, 500, or more than 1,000 amino acids, or optionally up to 7,500 amino acids; iv. A heterologous target sequence codes for a fragment of a mammalian gene but not for a complete mammalian gene; for example, it codes for one or more exons but not for a full-length protein; v. A heterologous target sequence encodes one or more introns; vi. The heterologous object sequence is other than GFP, for example, other than a fluorescent protein or other than a reporter protein; and vii. A heterogeneous object array contains non-code arrays, such as only adjustment elements.
[0047] 43. A system of any of the preceding embodiments, further comprising a target DNA-binding domain, wherein the polypeptide has an amino acid sequence encoded by, for example, a sequence of an element in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0048] 44. Any system of the preceding embodiments, further comprising a target DNA-binding domain, wherein the polypeptide has an amino acid sequence encoded by, for example, the sequence of elements in Table 3B, Table 10, Table 11, or Table X, or a sequence in which there are 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom.
[0049] 45. A system according to any of the preceding embodiments, further comprising a target DNA-binding domain having, for example, an amino acid sequence of an element in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0050] 46. A system of any of the preceding embodiments, further comprising a target DNA-binding domain having, for example, an amino acid sequence of an element in Table 3B, Table 10, Table 11, or Table X, or a sequence in which there are 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom.
[0051] 47. A system according to any of the previous embodiments, wherein the polypeptide exhibits activity at 37°C that is 70%, 75%, 80%, 85%, 90%, or 95% or more of the activity at 25°C that is otherwise the same as the activity under the same conditions.
[0052] 48. A polypeptide system according to any of the preceding embodiments, derived from a warm-blooded organism, such as a bird or a mammal.
[0053] 49. Polypeptides include CRE retrotransposase, NeSL retrotransposase, R4 retrotransposase, R2 retrotransposase, Hero retrotransposase, L1 retrotransposase, RTE retrotransposase, I retrotransposase, Jockey retrotransposase, CR1 retrotransposase, Rex1 retrotransposase, RandI / Dualen retrotransposase, Penelope or Penelope-like retrotransposase, Tx1 retrotransposase, RTEX retrotransposase, Crack retrotransposase, Nimb retrotransposase. A system according to any of the preceding embodiments, derived from one or more of the following retrotransposeses: Proto1 retrotransposese, Proto2 retrotransposese, RTETP retrotransposese, L2 retrotransposese, Tad1 retrotransposese, Loa retrotransposese, Ingi retrotransposese, Outcast retrotransposese, R1 retrotransposese, Daphne retrotransposese, L2A retrotransposese, L2B retrotransposese, Ambal retrotransposese, Vingi retrotransposese, and / or Kiri retrotransposese.
[0054] 50. A system according to any of the previous embodiments, wherein the polypeptide comprises an endonuclease domain derived from a transposable element, such as a restriction-like endonuclease (RLE), a depurine / depyrimidine-site endonuclease-like endonuclease (APE), or a GIY-YIG endonuclease.
[0055] 51. The endonuclease domain is intact in any of the systems of the preceding embodiments.
[0056] 52. The endonuclease domain is inactivated in any of the systems of the preceding embodiments.
[0057] 53. Endonucleases are systems that insert nicks into DNA, as described in one of the previous embodiments.
[0058] 54. A system of any of the previous embodiments in which the endonuclease performs double-strand breaks.
[0059] 55. A system according to any of the preceding embodiments, wherein the template RNA comprises one or both of the sequences in Table 3A or 3B or 10 (for example, the 5' untranslated region in column 6 of Table 3A or 3B and the 3' untranslated region in column 7 of Table 3A or 3B), or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0060] 56. A system according to any of the preceding embodiments, wherein the template RNA comprises the sequence of Table 3A or 3B or 11 (for example, one or both of the 5' untranslated region in column 6 of Table 3A or 3B and the 3' untranslated region in column 7 of Table 3A or 3B), or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom.
[0061] 57. A system of any of the preceding embodiments having one or more of the following features (e.g., 1, 2, 3, or all of them): i. The nucleic acid encoding the polypeptide and the template RNA or the nucleic acid encoding the template RNA are distinct nucleic acids; ii. The template RNA does not encode an active reverse transcriptase (for example, including an inactivating mutant reverse transcriptase, as described in Examples 1-2), or does not contain a reverse transcriptase sequence; iii. The template RNA does not encode an active endonuclease (e.g., it contains an inactivated endonuclease or does not contain an endonuclease); or iv. The template RNA contains one or more chemical modifications.
[0062] 58. The template RNA (or DNA encoding the template RNA) comprises (i) a 5'UTR sequence that binds to the polypeptide, (ii) a 3'UTR sequence that binds to the polypeptide, (iii) a heterologous target sequence, and (iv) a promoter operably ligated to the heterologous target sequence. The promoter is positioned between the 5' untranslated sequence that binds to the polypeptide and the heterologous sequence, or The promoter is positioned between the 3' untranslated sequence that binds to the polypeptide and the heterologous sequence, as in any of the systems of the previous embodiments.
[0063] 59. The template RNA (or the DNA encoding the template RNA) includes (i) a 5'UTR sequence that binds to the polypeptide, (ii) a 3'UTR sequence that binds to the polypeptide, and (iii) a heterologous target sequence. In either of the systems of the preceding embodiments, the heterologous target sequence includes an open reading frame (or its reverse complement) in the 5' to 3' direction of the template RNA, or the heterologous target sequence includes an open reading frame (or its reverse complement) in the 3' to 5' direction of the template RNA:
[0064] 60. A system according to any of the preceding embodiments, wherein the 5'UTR sequence of the template RNA (or the DNA encoding the template RNA) has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the 5'UTR sequence of the element sequence in Table 3B, Table 10, or Table X.
[0065] 61. A system according to any of the preceding embodiments, wherein the 5'UTR sequence of the template RNA (or the DNA encoding the template RNA) has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with respect to the nucleotide located 5' to the start codon of the sequence of an element (e.g., including a retrotransposase-binding region) in Table 3B, Table 10, or Table X.
[0066] 62. A system according to any of the preceding embodiments, wherein the 5'UTR sequence of the template RNA (or the DNA encoding the template RNA) has substantial structural similarity (e.g., substantial secondary structure similarity) to the 5'UTR sequence of the element sequence in Table 3B, Table 10, or Table X.
[0067] 63. A system according to any of the preceding embodiments, wherein the 5'UTR sequence of the template RNA (or the DNA encoding the template RNA) has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the 5'UTR sequence in column 6 of Table 3A, 3B, or 10.
[0068] 64. A system according to any of the preceding embodiments, wherein the 3'UTR sequence of the template RNA (or the DNA encoding the template RNA) has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the 3'UTR sequence of the element sequence in Table 3B, Table 10, or Table X.
[0069] 65. A system according to any of the preceding embodiments, wherein the 3'UTR sequence of the template RNA (or the DNA encoding the template RNA) has substantial structural similarity (e.g., substantial secondary structure similarity) to the 3'UTR sequence of the element sequence in Table 3B, Table 10, or Table X.
[0070] 66. A system according to any of the preceding embodiments, wherein the 3'UTR sequence of the template RNA (or the DNA encoding the template RNA) has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the 3'UTR sequence in column 7 of Table 10 or Table 3A or 3B.
[0071] 67. A system according to any of the preceding embodiments, wherein the 3'UTR sequence of the template RNA (or DNA encoding the template RNA) has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with respect to the nucleotide located 3' to the stop codon of the sequence of an element (e.g., including a retrotransposase-binding region) in Table 3B, Table 10, or Table X.
[0072] 68. The 3'UTR of the template RNA (or the DNA encoding the template RNA) is flanked by a homology domain, for example, as described herein, which has at least 5, 10, 20, 50 or 100 bases having at least 80% identity (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 100%) with respect to the target DNA strand (wherein optionally, the homology domain includes a sequence following the 3' homology arm in Table 11, or a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto), as in any of the systems of the preceding embodiments.
[0073] The 69.5'UTR is flanked by homology domains, for example, as described herein, for example, homology domains having at least 10, 20, 50 or 100 bases having at least 80% identity (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 100%) with respect to the target DNA strand (wherein optionally, the homology domains include sequences following the 5' homology arm in Table 11, or sequences having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto), as in any of the systems of the preceding embodiments.
[0074] 70. A system of any of the preceding embodiments, wherein at least one of the reverse transcriptase domain, endonuclease domain, and target DNA binding domain is heterogeneous with respect to the other domains, for example.
[0075] 71. A system of any of the preceding embodiments in which the endonuclease domain is heterogeneous to the reverse transcriptase domain and / or the target DNA binding domain.
[0076] 72. A system of any of the preceding embodiments, wherein the polypeptide comprises (i) a sequence that is at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the reverse transcriptase domain of an APE-type non-LTR retrotransposon, and (ii) a sequence that is at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the endonuclease domain of an APE-type non-LTR retrotransposon.
[0077] 73. A system of any of the preceding embodiments in which the polypeptide comprises (i) a sequence that is at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the reverse transcriptase domain of a penelope-like element (PLE) type retrotransposon, and (ii) a sequence that is at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the endonuclease domain of a PLE type retrotransposon, for example, the PLE type retrotransposon comprises a GIY-YIG endonuclease.
[0078] 74. A PLE-type retrotransposon is a system of any of the preceding embodiments that does not contain a functional endonuclease domain.
[0079] 75. A PLE-type retrotransposase is a system of any of the preceding embodiments that includes a Penelope-like element that naturally lacks an endonuclease domain (e.g., an Athena element as described by Gladyshev and Arkhipova PNAS 104,9352-9357 (2007)).
[0080] 76. A PLE-type retrotransposase lacking a functional endonuclease domain is fused to Cas9, and for example, the PLE-type retrotransposase fused to Cas9 has DBD and / or endonuclease (e.g., nickase) function, as in any of the systems of the preceding embodiments.
[0081] 77. A polypeptide system of any of the preceding embodiments, comprising (i) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the reverse transcriptase domain of a restriction enzyme-like endonuclease (RLE) type non-LTR retrotransposon, (ii) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the endonuclease domain of an RLE type non-LTR retrotransposon, and (iii) a target DNA-binding domain heterologous to (i) and / or (ii) (e.g., a heterologous zinc finger DNA-binding domain).
[0082] 78. A polypeptide comprising (i) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the reverse transcriptase domain of a restriction enzyme-like endonuclease (RLE) type non-LTR retrotransposon, and (ii) a heterologous endonuclease domain and / or a heterologous DNA-binding domain, according to any of the systems of the preceding embodiments.
[0083] 79. A polypeptide system of any of the preceding embodiments comprising (i) a sequence that is at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the reverse transcriptase domain of a depurine / depyrimidine site endonuclease (APE) type non-LTR retrotransposon, (ii) a sequence that is at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the endonuclease domain of an APE type non-LTR retrotransposon, and (iii) a target DNA-binding domain that is heterologous to (i) and / or (ii) (e.g., a heterologous zinc finger DNA-binding domain).
[0084] 80. A polypeptide comprising (i) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the reverse transcriptase domain of a depurine / depyrimidine site endonuclease (APE) type non-LTR retrotransposon, and (ii) a heterologous endonuclease domain and / or a heterologous DNA-binding domain, according to any of the preceding embodiments.
[0085] 81. A polypeptide system of any of the preceding embodiments, comprising (i) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the reverse transcriptase domain of a penelope-like element (PLE) type retrotransposon, (ii) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the endonuclease domain of a PLE type retrotransposon, and (iii) a target DNA-binding domain heterologous to (i) and / or (ii) (e.g., a heterologous zinc finger DNA-binding domain).
[0086] 82. A polypeptide comprising (i) a sequence at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the reverse transcriptase domain of a Penelope-like element (PLE) type retrotransposon, and (ii) a heterologous endonuclease domain and / or a heterologous DNA-binding domain, according to any of the systems of the preceding embodiments.
[0087] 83. The template RNA is a system of any of the preceding embodiments, comprising (iii) a promoter operably ligated to a heterologous target sequence.
[0088] 84. The polypeptide system of any of the preceding embodiments, further comprising (iii) a DNA-binding domain.
[0089] 85. A system of any of the previous embodiments, wherein the DNA-binding domain has endonuclease activity.
[0090] 86. A system according to any of the preceding embodiments in which the endonuclease domain or endonuclease activity forms a double-strand break in DNA.
[0091] 87. An endonuclease domain or endonuclease activity is used to introduce a nick into DNA, as in any of the previous embodiments.
[0092] 88. A system of any of the preceding embodiments, wherein the polypeptide contains sequences that are at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the sequences in column 7 of Table 3A or 3B.
[0093] 89. A system of any of the preceding embodiments, wherein the polypeptide contains sequences that are at least 80% identical (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identical) to the sequences of elements listed in Table 10, Table 11, or Table X.
[0094] 90. A system according to any of the preceding embodiments, wherein the nucleic acid encoding the polypeptide and the template RNA, or the nucleic acid encoding the template RNA, are covalently bonded (e.g., are part of a fusion nucleic acid).
[0095] 91. A system of any of the preceding embodiments in which the fusion nucleic acid contains RNA.
[0096] 92. A system of any of the preceding embodiments in which the fusion nucleic acid contains DNA.
[0097] 93.(b) is a system of any of the above embodiments, including template RNA.
[0098] 94. A system of any of the preceding embodiments, wherein the template RNA further includes a nuclear localization signal.
[0099] 95.(a) is a system of any of the preceding embodiments, comprising RNA encoding a polypeptide.
[0100] 96. A system of any of the preceding embodiments, wherein the RNA in (a) and the RNA in (b) are distinct RNA molecules.
[0101] 97. A system according to any of the above embodiments, in which the RNA of (a) and the RNA of (b) exist in the ratios of 100:1 to 10:1, 10:1 to 5:1, 5:1 to 2:1, 2:1 to 1:1, 1:1 to 1:2, 1:2 to 1:5, 1:5 to 1:10, or 1:10 to 1:100.
[0102] 98.(a) RNA is a system of any of the preceding embodiments that does not contain a nuclear localization signal.
[0103] 99. A system of any of the prior embodiments, further comprising a polypeptide, a nuclear localization signal and / or a nucleolar localization signal.
[0104] 100.(a) is any system of the preceding embodiment, comprising (i) a polypeptide and (ii) RNA encoding a nuclear localization signal and / or a nucleolar localization signal.
[0105] 101. RNA is a system of any of the preceding embodiments, which includes a pseudoknot sequence, for example, at the 5' end of a heterologous target sequence.
[0106] 102. RNA is a system of any of the previous embodiments, comprising a stem-loop sequence or a helix at the 5' end of a pseudoknot sequence.
[0107] 103. A system of any of the above embodiments in which the RNA includes one or more (e.g., two, three or more) stem-loop sequences or helices on the 3' side of the pseudoknot sequence, for example, on the 3' side of the pseudoknot sequence and on the 5' side of the heterologous target sequence.
[0108] 104. A template RNA containing a pseudoknot has catalytic activity, e.g., RNA cleavage activity, e.g., cis-RNA cleavage activity, in any of the systems of the preceding embodiments.
[0109] 105. The RNA system, for example, includes at least one stem-loop sequence, e.g., 1, 2, 3, 4, 5 or more stem-loop sequences, hairpin or helix sequences, at the 3' end of the heterologous target sequence, according to any of the preceding embodiments.
[0110] 106. A system according to any of the preceding embodiments, wherein the reverse transcriptase domain has the amino acid sequence of the reverse transcriptase domain of the elements listed in Figure 10, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0111] 107. The system of any of the preceding embodiments, wherein the endonuclease domain has the amino acid sequence of the endonuclease domain of the elements listed in Figure 10, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0112] 108. Reverse transcriptase domains provided herein include R2NS-1_CSi, R4-1_PH, RTE-1_MD, L2-18_ACar, L2-2_DRe, Dong-1_MMa, CR1_AC_1, Vingi-1_Acar, SR2, R4-1_AC, DongAa, CR1-1_PH, RTE-2_LMi, CR1-10_AMi, RTE-2_OL, L1-1_Cho, BovB, L A system according to any of the previous embodiments, having the amino acid sequence of the reverse transcriptase domain of ine1-2_ZM, RTE-1_Aip, CR1-3_IS, CR1-54_AAe, Tad1-65B_BG, TART-1_DWi, or RTE_Ele2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0113] 109. Endonuclease domains provided herein include R2NS-1_CSi, R4-1_PH, RTE-1_MD, L2-18_ACar, L2-2_DRe, Dong-1_MMa, CR1_AC_1, Vingi-1_Acar, SR2, R4-1_AC, DongAa, CR1-1_PH, RTE-2_LMi, CR1-10_AMi, RTE-2_OL, L1-1_Cho, BovB, L A system according to any of the above embodiments, having the amino acid sequence of the endonuclease domain of ine1-2_ZM, RTE-1_Aip, CR1-3_IS, CR1-54_AAe, Tad1-65B_BG, TART-1_DWi, or RTE_Ele2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0114] 110. Polypeptides provided herein include R2NS-1_CSi, R4-1_PH, RTE-1_MD, L2-18_ACar, L2-2_DRe, Dong-1_MMa, CR1_AC_1, Vingi-1_Acar, SR2, R4-1_AC, DongAa, CR1-1_PH, RTE-2_LMi, CR1-10_AMi, RTE-2_OL, L1-1_Cho, Bo A system according to any of the above embodiments, having an amino acid sequence of vB, Line1-2_ZM, RTE-1_Aip, CR1-3_IS, CR1-54_AAe, Tad1-65B_BG, TART-1_DWi, or RTE_Ele2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0115] 111. The sequence of the template RNA that binds to the polypeptide is the 5'UTR, or R2NS-1_CSi, R4-1_PH, RTE-1_MD, L2-18_ACar, L2-2_DRe, Dong-1_MMa, CR1_AC_1, Vingi-1_Acar, SR2, R4-1_AC, DongAa, CR1-1_PH, RTE-2_LMi, CR1-10_AMi, RTE-2 A system according to any of the above embodiments, having sequences that have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with respect to _OL, L1-1_Cho, BovB, Line1-2_ZM, RTE-1_Aip, CR1-3_IS, CR1-54_AAe, Tad1-65B_BG, TART-1_DWi, or RTE_Ele2.
[0116] 112. The sequence of the template RNA that binds to the polypeptide is 3'UTR, or R2NS-1_CSi, R4-1_PH, RTE-1_MD, L2-18_ACar, L2-2_DRe, Dong-1_MMa, CR1_AC_1, Vingi-1_Acar, SR2, R4-1_AC, DongAa, CR1-1_PH, RTE-2_LMi, CR1-10_AMi, RTE-2 A system according to any of the above embodiments, having sequences that have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with respect to _OL, L1-1_Cho, BovB, Line1-2_ZM, RTE-1_Aip, CR1-3_IS, CR1-54_AAe, Tad1-65B_BG, TART-1_DWi, or RTE_Ele2.
[0117] 113. A nucleic acid encoding a polypeptide, comprising a codon-optimized coding sequence for expression in human cells, according to any of the preceding embodiments.
[0118] 114. A system of any of the preceding embodiments in which the template RNA contains a coding sequence that is codon-optimized for expression in human cells.
[0119] 115. A system according to any of the preceding embodiments, comprising one or more circular RNA molecules (circRNA).
[0120] 116.circRNA is a system of any of the previous embodiments that encodes a Gene Writer polypeptide.
[0121] 117. CircRNA is a system of any of the preceding embodiments, including template RNA.
[0122] 118. CircRNA is delivered to a host cell in one of the systems described in the previous embodiments.
[0123] 119.CircRNA is a system of any of the preceding embodiments that can be linearized, for example, within a host cell, for example, within the nucleus of a host cell.
[0124] 120.circRNA is a system from any of the previous embodiments that includes a cleavage site.
[0125] 121. The system of any of the preceding embodiments further comprising a second cleavage site for circRNA.
[0126] 122. A system of any of the preceding embodiments in which the cleavage site can be cleaved by a ribozyme, for example, a ribozyme contained within circRNA (e.g., by autocleavage).
[0127] 123.circRNA is a system of any of the previous embodiments, containing a ribozyme sequence.
[0128] 124. A system of any of the preceding embodiments in which the ribozyme sequence can be self-cleaved, for example, within a host cell, for example, within the nucleus of a host cell.
[0129] 125. The ribozyme is an inducible ribozyme, as in any of the systems of the preceding embodiments.
[0130] 126. The system of any of the preceding embodiments, wherein the ribozyme is a protein-responsive ribozyme, e.g., a ribozyme responsive to a nucleoprotein, e.g., a genome-interacting protein, e.g., an epigenetic modifier, e.g., EZH2.
[0131] 127. A system of any of the preceding embodiments in which the ribozyme is a nucleic acid-responsive ribozyme.
[0132] 128. A system according to any of the preceding embodiments, in which the catalytic activity (e.g., autocatalytic activity) of a ribozyme is activated in the presence of a target nucleic acid molecule (e.g., an RNA molecule, e.g., mRNA, miRNA, ncRNA, lncRNA, tRNA, snRNA, or mtRNA).
[0133] 129. A system of any of the preceding embodiments in which the ribozyme is responsive to a target protein (e.g., MS2 coat protein).
[0134] 130. A system according to any of the previous embodiments, in which the target protein is localized in the cytoplasm or the nucleus (e.g., an epigenetic modifier or transcription factor).
[0135] 131. A system according to any of the preceding embodiments, wherein the ribozyme comprises a ribozyme sequence of a B2 or ALU retrotransposon, or a nucleic acid sequence having at least 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
[0136] 132. A system according to any of the preceding embodiments, wherein the ribozyme comprises the sequence of the tobacco ring spot virus hammerhead ribozyme, or a nucleic acid sequence having at least 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
[0137] 133. A system according to any of the preceding embodiments, wherein the ribozyme comprises the sequence of a hepatitis delta virus (HDV) ribozyme, or a nucleic acid sequence having at least 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
[0138] 134. A system according to any of the preceding embodiments, in which a ribozyme is activated by a portion expressed within a target cell or target tissue.
[0139] 135. A system according to any of the preceding embodiments, in which a ribozyme is activated by a portion expressed within a target intracellular compartment (e.g., the nucleus, nucleolus, cytoplasm, or mitochondria).
[0140] 136. A system of either of the preceding embodiments in which the ribozyme is contained within circular RNA or linear RNA.
[0141] 137. The first circular RNA encoding the polypeptide of the Gene Writing system; A second circular RNA containing the template RNA of the Gene Writing system and A system that includes this.
[0142] 138. A system according to any of the preceding embodiments, wherein the template RNA, e.g., the 5'UTR, comprises a ribozyme that cleaves the template RNA (e.g., at the 5'UTR).
[0143] 139. The system according to any of the above embodiments, wherein the template RNA comprises a ribozyme heterogeneous to (a)(i)(reverse transcriptase domain), (a)(ii)(endonuclease domain), (b)(i)(sequence of template RNA that binds to a polypeptide), or a combination thereof.
[0144] 140. A system according to any of the preceding embodiments in which a heterologous ribozyme can cleave RNA containing the ribozyme, for example, the 5' end of the ribozyme, the 3' end of the ribozyme, or the interior of the ribozyme.
[0145] 141. Lipid nanoparticles (LNPs) comprising any of the systems, polypeptides (or RNA encoding them), nucleic acid molecules, or DNA encoding the system or polypeptide, according to any of the preceding embodiments.
[0146] 142. A first lipid nanoparticle comprising a polypeptide (or DNA or RNA encoding it) of the Gene Writing system (as described herein, for example); A second lipid nanoparticle containing a nucleic acid molecule (e.g., as described herein) of the Gene Writing system and A system that includes this.
[0147] 143. A system or reaction mixture of any of the preceding embodiments, wherein the system, nucleic acid molecules, polypeptides, and / or DNA encoding them are formulated as lipid nanoparticles (LNPs).
[0148] 144. An LNP of any of the previous embodiments, comprising a cationic lipid.
[0149] 145. Cationic lipids have the following structure: [ka] An LNP having any of the above embodiments.
[0150] 146. An LNP of any of the previous embodiments, further comprising one or more neutral lipids, e.g., DSPC, DPPC, DMPC, DOPC, POPC, DOPE, SM, steroids, e.g., cholesterol, and / or one or more polymer-conjugated lipids, e.g., pegylated lipids, e.g., PEG-DAG, PEG-PE, PEG-S-DAG, PEG-cer, or PEG-dialkyloxypropyl carbamate.
[0151] 147. A system, kit, or polypeptide of any of the preceding embodiments, wherein the system, polypeptide, and / or the DNA encoding it are formulated as lipid nanoparticles (LNPs).
[0152] 148. A system, kit, or polypeptide of any of the preceding embodiments, wherein the lipid nanoparticles (or formulations comprising multiple lipid nanoparticles) are devoid of reactive impurities (e.g., aldehydes) or contain reactive impurities (e.g., aldehydes) below a pre-selected level.
[0153] 149. A system, kit, or polypeptide of any of the preceding embodiments, wherein the lipid nanoparticles (or formulations comprising multiple lipid nanoparticles) lack an aldehyde or contain an aldehyde below a pre-selected level.
[0154] 150. A system, kit, or polypeptide of any of the preceding embodiments, comprising lipid nanoparticles in a formulation containing multiple lipid nanoparticles.
[0155] 151. A lipid nanoparticle formulation is produced using one or more lipid reagents comprising a total reactive impurity (e.g., aldehyde) content of 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or less than 0.1% of a system, kit, or polypeptide according to any of the preceding embodiments.
[0156] 152. A lipid nanoparticle formulation is a system, kit, or polypeptide according to any of the preceding embodiments, produced using one or more lipid reagents comprising a total reactive impurity (e.g., aldehyde) content of less than 3%.
[0157] 153. A lipid nanoparticle formulation is produced using one or more lipid reagents containing any single reactive impurity species (e.g., aldehyde) in amounts of 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or less than 0.1%, according to any system, kit, or polypeptide of the preceding embodiments.
[0158] 154. A lipid nanoparticle formulation is produced using one or more lipid reagents containing less than 0.3% of any single reactive impurity species (e.g., aldehyde), and is a system, kit, or polypeptide of any of the preceding embodiments.
[0159] 155. A lipid nanoparticle formulation is produced using one or more lipid reagents containing less than 0.1% of any single reactive impurity species (e.g., aldehyde), and is a system, kit, or polypeptide of any of the preceding embodiments.
[0160] 156. A system, kit, or polypeptide according to any of the preceding embodiments, comprising a lipid nanoparticle formulation with a total reactive impurity (e.g., aldehyde) content of 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or less than 0.1%.
[0161] 157. A lipid nanoparticle formulation comprising a total reactive impurity (e.g., aldehyde) content of less than 3% of any of the systems, kits, or polypeptides of the preceding embodiments.
[0162] 158. A lipid nanoparticle formulation comprising any single reactive impurity species (e.g., aldehyde) in amounts of 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or less than 0.1%, as a system, kit, or polypeptide of any of the preceding embodiments.
[0163] 159. A lipid nanoparticle formulation comprising any single reactive impurity species (e.g., aldehyde) in a concentration of less than 0.3% of any system, kit, or polypeptide according to any of the preceding embodiments.
[0164] 160. A lipid nanoparticle formulation comprising any single reactive impurity species (e.g., aldehyde) in a concentration of less than 0.1% of any system, kit, or polypeptide according to any of the preceding embodiments.
[0165] 161. One or more, or optionally all, of the lipid reagents used in the lipid nanoparticles described herein, or their formulations, constitute a total reactive impurity (e.g., aldehyde) content of 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or less than 0.1% of any of the systems, kits, or polypeptides of the preceding embodiments.
[0166] 162. One or more, or optionally all, of the lipid reagents used in the lipid nanoparticles described herein, or their formulations, constitute a total reactive impurity (e.g., aldehyde) content of less than 3%, which is any of the systems or polypeptides of the preceding embodiments.
[0167] 163. One or more, or optionally all, of the lipid reagents used in the lipid nanoparticles described herein, or their formulations, contain any single reactive impurity species (e.g., aldehyde) in amounts of 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or less than 0.1%, as a system, kit, or polypeptide of any of the preceding embodiments.
[0168] 164. One or more, or optionally all, of the lipid reagents used in the lipid nanoparticles described herein, or the formulation thereof, comprises any single reactive impurity species (e.g., aldehyde) in a concentration of less than 0.3%, as a system, kit, or polypeptide of any of the preceding embodiments.
[0169] 165. One or more, or optionally all, of the lipid reagents used in the lipid nanoparticles described herein, or the formulation thereof, comprises any single reactive impurity species (e.g., aldehyde) in a concentration of less than 0.1%, as a system, kit, or polypeptide of any of the preceding embodiments.
[0170] 166. The total aldehyde content and / or amount of any single reactive impurity species (e.g., aldehyde) is determined by liquid chromatography (LC), for example, coupled with tandem mass spectrometry (MS / MS), according to the method described in Example 26, for example, a system, kit, or polypeptide of any of the preceding embodiments.
[0171] 167. The total aldehyde content and / or amount of reactive impurities (e.g., aldehydes) is determined by detecting, for example, one or more chemical modifications of nucleic acid molecules (e.g., as described herein) associated with the presence of reactive impurities (e.g., aldehydes) in a lipid reagent, according to any system, kit, or polypeptide of the preceding embodiments.
[0172] 168. The total aldehyde content and / or amount of aldehyde species is determined by detecting one or more chemical modifications of nucleotides or nucleosides (e.g., ribonucleotides or ribonucleosides, e.g., ribonucleotides or ribonucleosides, e.g., contained in or isolated from nucleic acid molecules, e.g., as described herein) associated with the presence of reactive impurities (e.g., aldehydes) in, for example, a lipid reagent, as described in Example 41, or any other system, kit, or polypeptide of any of the preceding embodiments.
[0173] 169. A system, kit, or polypeptide of any of the preceding embodiments in which chemical modifications of nucleic acid molecules, nucleotides, or nucleosides are detected by determining the presence of one or more modified nucleotides or nucleosides using, for example, LC-MS / MS analysis, as described in Example 41.
[0174] 170. A polypeptide is a sequence of at least 50 amino acids (e.g., at least 100, 150, 200, 300, 500 amino acids) having at least 80% identity (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100%) to the polypeptide sequence encoded by the sequence of an element in Table 10, Table 11, or Table X, or any of the above numbered systems including its reverse transcriptase domain or endonuclease domain.
[0175] 171. A polypeptide is a sequence of at least 50 amino acids (e.g., at least 100, 150, 200, 300, 500 amino acids) in which 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids are different from the polypeptide sequence encoded by the sequence of elements in Table 10, Table 11, or Table X, or a system of any of the above numbers including its reverse transcriptase domain or endonuclease domain.
[0176] 172. A polypeptide is a sequence of at least 50 amino acids (e.g., at least 100, 150, 200, 300, 500 amino acids) having at least 80% identity (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100%) to the polypeptide sequences listed in Table 3A or 3B, or any of the above numbered systems including a reverse transcriptase domain, an endonuclease domain, or a DNA-binding domain.
[0177] 173. A polypeptide is a sequence of at least 50 amino acids (e.g., at least 100, 150, 200, 300, or 500 amino acids) in which 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids different from the polypeptide sequences listed in Table 3A or 3B, or any of the above numbered systems including a reverse transcriptase domain, an endonuclease domain, or a DNA-binding domain.
[0178] 174. A polypeptide is a sequence of at least 50 amino acids (e.g., at least 100, 150, 200, 300, 500 amino acids) having at least 80% identity (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity) to the amino acid sequence in column 7 of Table 3A or 3B, or any of the above numbered systems including a reverse transcriptase domain, an endonuclease domain, or a DNA-binding domain.
[0179] 175. A polypeptide is a sequence of at least 50 amino acids (e.g., at least 100, 150, 200, 300, or 500 amino acids) in which 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer amino acids are different from the amino acid sequence in column 7 of Table 3A or 3B, or any of the above numbered systems including a reverse transcriptase domain, an endonuclease domain, or a DNA-binding domain.
[0180] 176. The template RNA is any numbered above, comprising the sequence of an element in Table 10 or Table X (for example, one or both of the 5'UTR and 3'UTR of Table X or Table 10), or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0181] 177. The template RNA is any numbered system above, comprising the sequence of an element from Table 10 or Table X (for example, one or both of the 5'UTR and 3'UTR of Table X or Table 10), or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom.
[0182] 178. The template RNA is any number above, comprising the sequence of Table 3A or 3B (for example, one or both of the 5'UTR in column 5 of Table 3A or 3B and the 3'UTR in column 6 of Table 3A or 3B), or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0183] 179. The template RNA is any numbered system above, comprising the sequence in Table 3A or 3B (for example, one or both of the 5'UTR in column 5 of Table 3A or 3B and the 3'UTR in column 6 of Table 3A or 3B), or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom.
[0184] 180. The template RNA is any number of the above sequences, (i) containing sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with respect to nucleotides located 5' to the start codon of the elements of Table 10 or Table X (e.g., including retrotransposase-binding regions).
[0185] 181. The template RNA is any number of the above sequences, (i) containing sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with respect to nucleotides located 3' to the stop codon of the elements of Table 10 or Table X (e.g., including retrotransposase-binding regions).
[0186] 182. The template RNA comprises a sequence of approximately 100-125 bp from the 3'UTR of column 6 of Table 10 or 11 or Table 3A or 3B, for example, the sequence comprising nucleotides 1-100, 101-200, or 201-325 of the 3'UTR of column 6 of Table 3A or 3B, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, according to any of the preceding embodiments.
[0187] 183. A system according to any of the preceding embodiments, wherein the template RNA comprises a sequence of approximately 100-125 bp from the 3'UTR of column 6 of Table 10 or 11 or Table 3A or 3B, for example, a sequence comprising nucleotides 1-100, 101-200, or 201-325 of the 3'UTR of column 6 of Table 3A or 3B, or a sequence with 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom.
[0188] 184. Any of the above numbered systems, where (a) contains RNA and (b) contains RNA.
[0189] 185. The above numbered systems, which contain only RNA, or contain more RNA than DNA in an RNA:DNA ratio of at least 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1.
[0190] 186. Any system numbered above that does not contain DNA, or does not contain more than 10%, 5%, 4%, 3%, 2%, or 1% DNA by mass or molar amount.
[0191] Any of the above numbered systems that can modify DNA by inserting a heterologous target sequence without the need for DNA-dependent RNA polymerization as described in 187.(b).
[0192] 188. Any of the above numbered systems that can modify DNA by inserting a heterologous target sequence via target-prime reverse transcription.
[0193] 189. Any of the above numbered systems that can modify DNA by inserting heterologous target sequences in the presence of inhibitors of the DNA repair pathway (e.g., SCR7, PARP inhibitors) or in cell lines lacking the DNA repair pathway (e.g., cell lines lacking the nucleotide excision repair pathway or homologous recombination repair pathway).
[0194] 190. Any of the above numbered systems that do not cause the formation of detectable levels of double-strand breaks in target cells.
[0195] 191. Any number of the above systems that can modify DNA using reverse transcriptase activity in the absence of homologous recombination activity by arbitrary selection.
[0196] 192. The template RNA is treated to reduce secondary structure (e.g., heated to a temperature that reduces secondary structure, e.g., at least 70, 75, 80, 85, 90 or 95°C, e.g., in any of the above systems).
[0197] 193. The template RNA is subsequently cooled to a temperature that allows for secondary structure formation, for example, below 37, 30, 25, or 20°C, in any of the systems of the preceding embodiments.
[0198] 194. A system according to any of the preceding embodiments, wherein the template RNA (or DNA encoding the template RNA) comprises a first homology domain at the 5' end of the template RNA having at least 5 or at least 10 bases identical to the target DNA strand, and a second homology domain at the 3' end of the template RNA having at least 5 or at least 10 bases identical to the target DNA strand.
[0199] 195. A system of either of the preceding embodiments in which (a) and (b) are part of the same nucleic acid.
[0200] 196. A system of either of the preceding embodiments, wherein (a) and (b) are distinct nucleic acids.
[0201] 197. A system according to any of the preceding embodiments, wherein the template RNA contains at least 5 or at least 10 bases at its 5' end that are 100% identical to the target DNA strand (for example, the target DNA strand is a human DNA sequence).
[0202] 198. A system according to any of the preceding embodiments, wherein the template RNA contains at least 5 or at least 10 bases at its 3' end that are 100% identical to the target DNA strand (for example, the target DNA strand is a human DNA sequence).
[0203] 199. A polypeptide comprising an active RNase H domain, the system of any of the preceding embodiments.
[0204] 200. A system of any of the previous embodiments in which the polypeptide does not contain an active RNase H domain.
[0205] 201. A system of any of the preceding embodiments in which the endogenous RNase H domain of the polypeptide is inactivated.
[0206] 202. A transposase polypeptide comprising a mutation that inactivates and / or deletes a nucleolar localization signal, as in any of the systems of the preceding embodiments.
[0207] 203. A polypeptide system of any of the preceding embodiments that does not contain a functional nucleolar localization signal (e.g., a system that does not contain a nucleolar localization signal).
[0208] 204. A system of any of the preceding embodiments in which the activity of the nucleolar localization signal is reduced by at least 50%, 60%, 70%, 80%, 90%, 95%, or 99%.
[0209] 205. A polypeptide system comprising a nuclear localization signal (NLS), such as an endogenous NLS or an exogenous NLS, as in any of the preceding embodiments.
[0210] 206. The polypeptide comprises (i) a first target DNA binding domain, for example, containing a first Zn finger domain; (ii) a reverse transcriptase domain; (iii) an endonuclease domain; and (iv) a second target DNA binding domain heterogeneous to the first target DNA binding domain, for example, containing a second Zn finger domain; (a) binds to the target DNA sequence in fewer target cells than a similar polypeptide containing only the first target DNA binding domain. For example, when the second target DNA binding domain is present in the polypeptide together with the first DNA binding domain, the target specificity of the polypeptide is improved compared to the polypeptide target sequence specificity of a polypeptide containing only the first target DNA binding domain, in any of the previous embodiments of the system.
[0211] 207. (iii) includes (iv), in any of the previous embodiments of the system.
[0212] 208. The second target DNA binding domain binds to a genomic DNA sequence that is less than 100, 90, 80, 70, 60, 50, 40, 30, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide away from the genomic sequence to which the first target DNA binding domain binds, in any of the previous embodiments of the system.
[0213] 209. The second target DNA binding domain binds to a genomic DNA sequence that is 1 - 100, 1 - 90, 1 - 80, 1 - 60, 1 - 50, 1 - 40, 1 - 30, 1 - 20, 1 - 10, 1 - 5, 5 - 100, 5 - 90, 5 - 80, 5 - 70, 5 - 60, 5 - 50, 5 - 40, 5 - 30, 5 - 20, 5 - 10, 10 - 100, 10 - 90, 10 - 80, 10 - 70, 10 - 60, 10 - 50, 10 - 40, 10 - 30, 10 - 20, 20 - 100, 20 - 90, 20 - 80, 20 - 70, 20 - 60, 20 - 50, 20 - 40, 20 - 30, 30 - 100, 30 - 90, 30 - 80, 30 - 70, 30 - 60, 30 - 50, 30 - 40, 40 - 100, 40 - 90, 40 - 80, 40 - 70, 40 - 60, 40 - 50, 50 - 100, 50 - 90, 50 - 80, 50 - 70, 50 - 60, 60 - 100, 60 - 90, 60 - 80, 60 - 70, 70 - 100, 70 - 90, 70 - 80, 80 - 100, 80 - 90 or 90 - 100 nucleotides away from the genomic sequence to which the first target DNA binding domain binds, in any of the previous embodiments of the system.
[0214] 210. The first or second target DNA binding domain is a system of any of the previous embodiments, including a CRISPR / Cas protein, a TAL effector domain, a Zn finger domain, or a meganuclease domain.
[0215] 211. The first strand of the target DNA can be cleaved at least twice (e.g., twice), and optionally, the cleavages are at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100 or 200 nucleotides apart from each other (and optionally, 500, 400, 300, 200 or 100 nucleotides or less apart from each other), in a system of any of the previous embodiments.
[0216] 212. The first and second strands of the target DNA can be cleaved, and the distance between the cleavages is the same as the distance between the cleavages made by a reverse transcriptase domain, e.g., a reverse transcriptase domain located in its endogenous polypeptide, in a system of any of the previous embodiments.
[0217] 213. Cuts are made in the following order: 1-500, 1-400, 1-300, 1-200, 1-100, 1-90, 1-80, 1-70, 1-60, 1-50, 1-40, 1-30, 1-20, 1-10, 1-5, 5-500, 5-400, 5-300, 5-200, 5-100, 5-90, 5-80, 5-70, 5-60, 5-50, 5-40, 5-30, 5-20, 5-10, 10-500, 10-400, 10-300, 10-200, 10-100, 1 0-90, 10-80, 10-70, 10-60, 10-50, 10-40, 10-30, 10-20, 20-500, 20-400, 20-300, 20-200, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, 20-30, 30-500, 30-400, 30-300, 30-200, 30-100, 30-90, 30-80, 30-70, 30-60, 30-50, 30-40, 40-500 , 40~400, 40~300, 40~200, 40~100, 40~90, 40~80, 40~70, 40~60, 40~50, 50~500, 50~400, 50~300, 50~200, 50~100, 50~90, 50~80, 50~70, 50~60, 60~500, 60~400, 60~300, 60~200, 60~100, 60~90, 60~80, 60~70, 70~500, 70~400, 70~300, 70~200, 70 A system of any of the preceding embodiments, with nucleotides separated by ~100, 70~90, 70~80, 80~500, 80~400, 80~300, 80~200, 80~100, 80~90, 90~500, 90~400, 90~300, 90~200, 90~100, 100~500, 100~400, 100~300, 100~200, 200~500, 200~400, 200~300, 300~500, 300~400, or 400~500 nucleotides.
[0218] 214. A system of any of the preceding embodiments, wherein the distance between cleavage is the same as the distance between cleavage made by a reverse transcriptase domain, for example, a reverse transcriptase domain located in its endogenous polypeptide.
[0219] 215. A system of any of the preceding embodiments in which both cleavages are performed by the same endonuclease domain (e.g., a CRISPR / Cas protein (e.g., directed by multiple gRNAs and located in, e.g., template RNA)).
[0220] 216. A polypeptide system of any of the preceding embodiments, further comprising a second endonuclease domain.
[0221] 217. i) The first endonuclease domain (e.g., nickase) cleaves the strand of target DNA to be edited, and the second endonuclease domain (e.g., nickase) cleaves the strand of target DNA that is not edited, or ii) A system according to any of the preceding embodiments, wherein a first endonuclease domain (e.g., nickase) performs one of two cleavages of the strand of target DNA to be edited, and a second endonuclease domain (e.g., nickase) performs the other cleavage of the strand of target DNA to be edited.
[0222] 218.(a), (b), or (a) and (b) is any system of the prior embodiments, further comprising a sequence encoding a polypeptide, a heterologous target sequence (e.g., a coding sequence contained in a heterologous target sequence), or both, operably concatenated to a 5'UTR and / or a 3'UTR.
[0223] A system of any of the preceding embodiments in which the 219.5'UTR and / or 3'UTR increase the expression of the operably linked sequence by at least 10%, 20%, 30%, 40%, 50%, 70%, 70%, 80%, 90%, or 100% compared to an endogenous UTR or other similar nucleic acid containing a minimum 5'UTR and a minimum 3'UTR associated with a heterologous target sequence.
[0224] 220. A system of any of the preceding embodiments, wherein the template RNA (or DNA encoding the template RNA) comprises (i) a sequence that binds to a polypeptide, (ii) a heterologous target sequence, and (iii) a ribozyme that is heterologous to (a)(i), (a)(ii), (b)(i), or a combination thereof.
[0225] 221. (a), (b), or (a) and (b) is any system of the prior embodiments comprising an intron that increases the expression of a polypeptide, a heterologous target sequence (e.g., a coding sequence located within a heterologous target sequence), or both.
[0226] 222. A host cell containing any of the above numbered systems (e.g., mammalian cell, e.g., human cell).
[0227] 223. A method for modifying a target DNA strand in a cell, tissue, or subject, comprising administering a system of any of the above numbers to the cell, tissue, or subject, wherein the system reverse transcribes a template RNA sequence into a target DNA strand, thereby modifying the target DNA strand.
[0228] 224. A method according to any of the preceding embodiments, wherein the cells, tissues, or subjects are cells, tissues, or subjects of a mammal (e.g., human).
[0229] 225. The method according to any of the preceding embodiments, wherein the tissue is liver, lung, skin, blood, immune, or muscle tissue.
[0230] 226. The cell is a fibroblast, according to any of the preceding embodiments.
[0231] 227. The cells are primary cells, according to any of the methods of the preceding embodiments.
[0232] 228. Any of the methods of the preceding embodiments, in which the cells are not immortalized.
[0233] 229. A method for modifying the genome of mammalian cells, wherein the cells (a) A polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (such as those listed in Table 3B, Table 10, Table 11, Table X or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA binding domain; one, two or three of (i), (ii) and / or (iii) has an amino acid sequence encoded by a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to the sequence of an element in Table 3B, Table 10, Table 11 or Table X; and (b) A template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds to the polypeptide and (ii) a heterologous target sequence A method comprising contacting them.
[0234] 230. A method of modifying the genome of a mammalian cell, comprising contacting the cell with (a) A polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain (such as those listed in Table 3B, Table 10, Table 11, Table X or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA binding domain; one, two or three of (i), (ii) and / or (iii) has an amino acid sequence encoded by a sequence having no more than 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 different nucleotides therefrom; and (b) A template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds to the polypeptide and (ii) a heterologous target sequence A method comprising contacting them.
[0235] 231. A method of modifying the genome of a mammalian cell, comprising contacting the cell with (a) polypeptides or nucleic acids encoding polypeptides (wherein a polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; one, two, or three of (i), (ii), and / or (iii) having an amino acid sequence of an element in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto); and (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. A method that includes bringing into contact with it.
[0236] 232. A method for modifying the genome of mammalian cells, wherein the cells (a) polypeptides or nucleic acids encoding polypeptides (wherein a polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; one, two, or three of (i), (ii), and / or (iii) having a sequence in which there are 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides that are different from the amino acid sequence of an element in Table 3B, Table 10, Table 11, or Table X); and (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. A method that includes bringing into contact with it.
[0237] 233. A method for modifying the genome of mammalian cells, wherein the cells (a) polypeptides or nucleic acids encoding polypeptides (wherein the polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table 3B, Table 10, Table 11, Table X or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain); and (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. Including contact with; A method wherein the sequence of the template RNA that binds to the polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the 5'UTR sequence or 3'UTR sequence of the element sequence in Table 3B, Table 10, or Table X.
[0238] 234. A method for modifying the genome of a mammalian cell, wherein the cell (a) polypeptides or nucleic acids encoding polypeptides (wherein the polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table 3B, Table 10, Table 11, Table X or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain); and (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. Including contact with; A method wherein the sequence of the template RNA that binds to the polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with (i) the nucleotide located 5' to the start codon of the sequence of the element in Table 3B, Table 10, or Table X (e.g., including the retrotransposase binding region) or (ii) the nucleotide located 3' to the stop codon of the sequence of the element in Table 3B, Table 10, or Table X (e.g., including the retrotransposase binding region).
[0239] 235. A method for modifying the genome of mammalian cells, wherein the cells (a) polypeptides or nucleic acids encoding polypeptides (e.g., DNA or mRNA) (wherein the polypeptide includes (i) a reverse transcriptase domain derived from a protein other than a retrotransposase, for example, a reverse transcriptase domain listed in Table Z1 or Z2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and (ii) an endonuclease domain); and, (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. A method that includes bringing into contact with it.
[0240] 236. A method for modifying the genome of a mammalian cell, wherein the cell (a) polypeptides or nucleic acids encoding polypeptides (e.g., DNA or mRNA) (wherein the polypeptide is (i) a reverse transcriptase domain derived from a protein other than a retrotransposase, for example, a reverse transcriptase domain listed in Table Z1 or Z2, or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 or fewer nucleotides different therefrom, and (ii) an endonuclease domain; and, (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. A method that includes bringing into contact with it.
[0241] 237. A method for modifying the genome of mammalian cells, wherein the cells (a) polypeptides or nucleic acids encoding polypeptides (wherein the polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain); (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA encoding template RNA) containing a heterologous target sequence; and (c) Intein A method that includes bringing into contact with it.
[0242] 238. A method according to any of the preceding embodiments, wherein the polypeptide comprises an intein.
[0243] 239. A system of any of the preceding embodiments in which the intein is split intein.
[0244] 240. Any of the methods of the previous embodiments, wherein the polypeptide does not contain a target DNA binding domain.
[0245] 241. The polypeptide is derived from an APE-type retrotransposon reverse transcriptase, according to any of the methods of the preceding embodiments.
[0246] 242. Any method of the preceding embodiment further comprising a target DNA-binding domain having an amino acid sequence encoded by, for example, a sequence of an element in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0247] 243. Any method of the preceding embodiment, further comprising a target DNA-binding domain, wherein the polypeptide has an amino acid sequence encoded by, for example, the sequences of elements in Table 3B, Table 10, Table 11, or Table X, or sequences in which there are 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom.
[0248] 244. Any method of the preceding embodiment, further comprising a target DNA-binding domain having, for example, an amino acid sequence of an element in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0249] 245. Any method of the preceding embodiment, further comprising a target DNA-binding domain, wherein the polypeptide has, for example, an amino acid sequence of an element in Table 3B, Table 10, Table 11, or Table X, or a sequence in which there are 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom.
[0250] 246. A method for modifying the genome of a mammalian cell, wherein the cell (a) RNA encoding a polypeptide (wherein the polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; one, two or three of (i), (ii), and / or (iii) having an amino acid sequence encoded by a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto); and (b)(i) Template RNA containing a polypeptide-binding sequence and (ii) a heterologous target sequence Including contact with; A method that does not involve contacting mammalian cells with DNA, or in which compositions (a) and (b) do not contain DNA in amounts exceeding 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% by mass or molar amount of nucleic acid.
[0251] 247. A method for modifying the genome of mammalian cells, wherein the cells (a) RNA encoding a polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; one, two or three of (i), (ii), and / or (iii) has an amino acid sequence encoded by a sequence of elements in Table 3B, Table 10, Table 11, or Table X, or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom); and (b)(i) Template RNA containing a polypeptide-binding sequence and (ii) a heterologous target sequence Including contact with; A method that does not involve contacting mammalian cells with DNA, or in which compositions (a) and (b) do not contain DNA in amounts exceeding 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% by mass or molar amount of nucleic acid.
[0252] 248. A method for modifying the genome of mammalian cells, wherein the cells (a) RNA encoding a polypeptide (wherein the polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; one, two, or three of (i), (ii), and / or (iii) having an amino acid sequence of an element in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto); and (b)(i) Template RNA containing a polypeptide-binding sequence and (ii) a heterologous target sequence Including contact with; A method that does not involve contacting mammalian cells with DNA, or in which compositions (a) and (b) do not contain DNA in amounts exceeding 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% by mass or molar amount of nucleic acid.
[0253] 249. A method for modifying the genome of a mammalian cell, wherein the cell (a) RNA encoding a polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; one, two, or three of (i), (ii), and / or (iii) have sequences in which there are 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides that are different from the amino acid sequences of the elements in Table 3B, Table 10, Table 11, or Table X); and (b)(i) Template RNA containing a polypeptide-binding sequence and (ii) a heterologous target sequence Including contact with; A method that does not involve contacting mammalian cells with DNA, or in which compositions (a) and (b) do not contain DNA in amounts exceeding 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% by mass or molar amount of nucleic acid.
[0254] 250. A method according to any of the preceding embodiments that results in the addition of an exogenous DNA sequence of at least 1, 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs to the genome of a mammalian cell.
[0255] 251. A method according to any of the preceding embodiments that results in the deletion of a DNA sequence of at least 1, 2, 3, 4, 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs from the genome of a mammalian cell.
[0256] 252. Any method according to the preceding embodiments that results in a modification from the genome of a mammalian cell of at least 1, 2, 3, 4, 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of DNA sequences.
[0257] 253. A method according to any of the preceding embodiments that results in the addition of a protein-coding sequence to the genome of a mammalian cell.
[0258] 254. Any method according to the preceding embodiments for causing a deletion of a protein-coding sequence in the genome of a mammalian cell.
[0259] 255. A method according to any of the preceding embodiments for causing a modification of the protein-coding sequence in the genome of a mammalian cell.
[0260] 256. Any method of the preceding embodiment that results in the addition of a non-coding sequence encoding, for example, non-coding RNA, such as miRNA, to the genome of a mammalian cell.
[0261] 257. Any method according to the preceding embodiments for causing a deletion in the genome of a mammalian cell of a non-coding sequence that encodes, for example, non-coding RNA, such as miRNA.
[0262] 258. Any method of the preceding embodiment that results in a modification of a non-coding sequence in the genome of a mammalian cell, for example, a sequence encoding non-coding RNA, for example, miRNA.
[0263] 259. A method according to any of the preceding embodiments that results in the addition of regulatory sequences, such as promoters, enhancers, or miRNA binding sites, to the genome of a mammalian cell.
[0264] 260. A method according to any of the preceding embodiments for causing deletion of regulatory sequences, such as promoters, enhancers, or miRNA binding sites, in the genome of a mammalian cell.
[0265] 261. A method according to any of the preceding embodiments for causing modification of regulatory sequences, such as promoters, enhancers, or miRNA binding sites, in the genome of a mammalian cell.
[0266] 262. Addition, deletion, or modification of regulatory sequences to the genome of a mammalian cell, resulting in increased expression of coding or non-coding sequences in the genome of a mammalian cell, as described in any of the preceding embodiments.
[0267] 263. Addition, deletion, or modification of regulatory sequences to the genome of a mammalian cell resulting in a reduction in the expression of coding or non-coding sequences in the genome of a mammalian cell, as described in any of the preceding embodiments.
[0268] 264. A method for inserting DNA into the genome of a mammalian cell, comprising contacting the cell with an RNA composition, wherein the RNA composition is (a) A first RNA that directs the insertion of the template RNA into the genome, and (b) Template RNA containing heterologous sequences Includes, Either the combination does not involve contacting mammalian cells with DNA, or the compositions of (a) and (b) do not contain DNA in amounts exceeding 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% by mass or molar amount of nucleic acid. This results in the addition of at least 1, 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of DNA (e.g., exogenous DNA) sequences to the genome of mammalian cells, and A method wherein the first RNA encodes a polypeptide encoded by a sequence of an element in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto (where the polypeptide instructs the insertion of the template RNA into the genome).
[0269] 265. A method for inserting DNA into the genome of a mammalian cell, comprising contacting the cell with an RNA composition, wherein the RNA composition is (a) A first RNA that directs the insertion of the template RNA into the genome, and (b) Template RNA containing heterologous sequences Includes, Either the combination does not involve contacting mammalian cells with DNA, or the compositions of (a) and (b) contain more than 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% of DNA by mass or molar amount of nucleic acid. This results in the addition of at least 1, 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of DNA (e.g., exogenous DNA) sequences to the genome of mammalian cells, and A method wherein the first RNA encodes a polypeptide encoded by a sequence of elements in Table 3B, Table 10, Table 11, or Table X, or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom (where the polypeptide instructs the insertion of the template RNA into the genome).
[0270] 266. A method for inserting DNA into the genome of a mammalian cell, comprising contacting the cell with an RNA composition, wherein the RNA composition is (a) A first RNA that directs the insertion of the template RNA into the genome, and (b) Template RNA containing heterologous sequences Includes, Either the combination does not involve contacting mammalian cells with DNA, or the compositions of (a) and (b) do not contain DNA in amounts exceeding 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% by mass or molar amount of nucleic acid. This results in the addition of at least 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of DNA (e.g., exogenous DNA) sequences to the genome of mammalian cells. A method wherein the first RNA encodes a polypeptide of an element from Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto (where the polypeptide directs the insertion of the template RNA into the genome).
[0271] 267. A method for inserting DNA into the genome of a mammalian cell, comprising contacting the cell with an RNA composition, wherein the RNA composition is (a) A first RNA that directs the insertion of the template RNA into the genome, and (b) Template RNA containing heterologous sequences Includes, Either the combination does not involve contacting mammalian cells with DNA, or the compositions of (a) and (b) do not contain DNA in amounts exceeding 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% by mass or molar amount of nucleic acid. This results in the addition of at least 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of DNA (e.g., exogenous DNA) sequences to the genome of mammalian cells, and The first RNA encodes a polypeptide of an element from Table 3B, Table 10, Table 11, or Table X, or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom (where the polypeptide instructs the insertion of the template RNA into the genome), and the method.
[0272] 268. A method for inserting DNA into the genome of a mammalian cell, comprising contacting the cell with an RNA composition, wherein the RNA composition is (a) A first RNA that instructs the insertion of the template RNA into the genome, (b) heterogeneous sequences, and (i) flanking sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the 5'UTR sequence of the element sequence of Table 3B, Table 10, or Table X, and / or (ii) template RNA having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the 3'UTR sequence of the element sequence of Table 3B, Table 10, or Table X, and Includes, Either the combination does not involve contacting mammalian cells with DNA, or the compositions of (a) and (b) do not contain DNA in amounts exceeding 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% by mass or molar amount of nucleic acid. A method for adding a DNA sequence (e.g., exogenous DNA) of at least 1, 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs to the genome of a mammalian cell.
[0273] 269. A method for inserting DNA into the genome of a mammalian cell, comprising contacting the cell with an RNA composition, wherein the RNA composition is (a) A first RNA that instructs the insertion of the template RNA into the genome, (b) heterogeneous sequences, and (i) flanking sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with nucleotides located 5' to the start codon of the sequences of the elements in Table 3B, Table 10, or Table X (e.g., including the retrotransposase binding region), and / or (ii) template RNA having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with nucleotides located 3' to the stop codon of the sequences of the elements in Table 3B, Table 10, or Table X (e.g., including the retrotransposase binding region), and Includes, Either the combination does not involve contacting mammalian cells with DNA, or the compositions of (a) and (b) do not contain DNA in amounts exceeding 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% by mass or molar amount of nucleic acid. A method for adding a DNA sequence (e.g., exogenous DNA) of at least 1, 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs to the genome of a mammalian cell.
[0274] 270. The template RNA further contains a sequence that binds to the polypeptide, and by optional selection, the sequence that binds to the polypeptide is: (a) The elements in Table 3B, Table 10, or Table X have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with respect to the 5'UTR or 3'UTR of the element arrangement, or (b) The sequence of the template RNA that binds to the polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with (i) the nucleotide located 5' to the start codon of the sequence of the element in Table 3B, Table 10, or Table X (e.g., including the retrotransposase binding region) or (ii) the nucleotide located 3' to the stop codon of the sequence of the element in Table X (e.g., including the retrotransposase binding region), according to any of the methods of the prior embodiments.
[0275] 271. Any method according to the preceding embodiments, wherein at least 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, or 1000 bp of exogenous DNA is added to the genome of a mammalian cell without delivering the DNA to the cell.
[0276] 272. At least 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, 1000 bp of exogenous DNA are added to the genome of mammalian cells. A method according to any of the previous embodiments, which does not involve contacting mammalian cells with DNA, or involves contacting mammalian cells with a composition having a DNA content of less than 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, or 0.01% by mass or molar amount of nucleic acid.
[0277] 273. Any method of the preceding embodiments, wherein at least 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, or 1000 bp of exogenous DNA is added to the genome of a mammalian cell, and only RNA is delivered to the mammalian cell.
[0278] 274. Any method according to the preceding embodiments, wherein at least 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, or 1000 bp of exogenous DNA is added to the genome of a mammalian cell, and RNA and proteins are delivered to the mammalian cell.
[0279] 275. A method according to any of the preceding embodiments, wherein the template RNA acts as a template for the insertion of exogenous DNA.
[0280] 276. Any method of the preceding embodiments that does not involve DNA-dependent RNA polymerization of exogenous DNA.
[0281] 277. A method according to any of the preceding embodiments that results in the addition of at least 1, 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, or 5,000 base pairs of DNA to the genome of a mammalian cell.
[0282] 278.A method according to any of the preceding embodiments, wherein the RNA of (a) and the RNA of (b) are covalently bound, for example, are part of the same transcript.
[0283] 279. The method according to any of the preceding embodiments, wherein the RNA in (a) and the RNA in (b) are distinct RNAs.
[0284] 280. Any method of the preceding embodiments, which does not involve contacting mammalian cells with template DNA.
[0285] 281. A method for modifying the genome of human cells, wherein the cells (a) polypeptides or nucleic acids encoding polypeptides (wherein a polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; one, two, or three of (i), (ii), and / or (iii) having an amino acid sequence encoded by a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto); and (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. This includes making contact with This results in the insertion of heterologous target sequences into the genome of human cells. Human cells do not show upregulation of DNA repair genes and / or tumor suppressor genes, or DNA repair genes and / or tumor suppressor genes are not upregulated by more than 50%, 25%, 10%, 5%, 2%, or 1% (where upregulation is measured, for example, by RNA-seq), method.
[0286] 282. A method for modifying the genome of human cells, wherein the cells (a) polypeptides or nucleic acids encoding polypeptides (wherein a polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; one, two or three of (i), (ii), and / or (iii) having an amino acid sequence encoded by a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom); and (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. This includes making contact with This results in the insertion of heterologous target sequences into the genome of human cells. Human cells do not exhibit upregulation of DNA repair genes and / or tumor suppressor genes, or DNA repair genes and / or tumor suppressor genes are not upregulated by more than 50%, 25%, 10%, 5%, 2%, or 1% (wherein upregulation is measured, for example, by RNA-seq as described in Example 14 of PCT / US2019 / 048607, which is incorporated herein by its entirety by the criteria).
[0287] 283. A method for modifying the genome of human cells, wherein the cells (a) polypeptides or nucleic acids encoding polypeptides (wherein a polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; one, two, or three of (i), (ii), and / or (iii) having an amino acid sequence of an element in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto); and (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. This includes making contact with This results in the insertion of heterologous target sequences into the genome of human cells. Human cells do not show upregulation of DNA repair genes and / or tumor suppressor genes, or DNA repair genes and / or tumor suppressor genes are not upregulated by more than 50%, 25%, 10%, 5%, 2%, or 1% (where upregulation is measured, for example, by RNA-seq), method.
[0288] 284. A method for modifying the genome of human cells, wherein the cells (a) polypeptides or nucleic acids encoding polypeptides (wherein a polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table 3B, Table 10, Table 11, Table X, or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain; one, two, or three of (i), (ii), and / or (iii) having a sequence in which there are 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides that are different from the amino acid sequence of an element in Table 3B, Table 10, Table 11, or Table X); and (b)(i) a sequence that binds to a polypeptide and (ii) a template RNA (or DNA that encodes template RNA) containing a heterologous target sequence. This includes making contact with This results in the insertion of heterologous target sequences into the genome of human cells. Human cells do not show upregulation of DNA repair genes and / or tumor suppressor genes, or DNA repair genes and / or tumor suppressor genes are not upregulated by more than 50%, 25%, 10%, 5%, 2%, or 1% (where upregulation is measured, for example, by RNA-seq), method.
[0289] 285. A method for modifying the genome of human cells, wherein the cells (a) polypeptides or nucleic acids encoding polypeptides (wherein the polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table 3B, Table 10, Table 11, Table X or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain); and (b)(i) A sequence that binds to a polypeptide having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the 5'UTR or 3'UTR sequence of an element in Table 3B, Table 10, Table 11, or Table X, and (ii) a template RNA (or DNA encoding a template RNA) containing a heterologous target sequence. This includes making contact with This results in the insertion of heterologous target sequences into the genome of human cells. Human cells do not show upregulation of DNA repair genes and / or tumor suppressor genes, or DNA repair genes and / or tumor suppressor genes are not upregulated by more than 50%, 25%, 10%, 5%, 2%, or 1% (where upregulation is measured, for example, by RNA-seq), method.
[0290] 286. A method for modifying the genome of human cells, wherein the cells (a) polypeptides or nucleic acids encoding polypeptides (wherein the polypeptide comprises (i) a reverse transcriptase domain (for example, as listed in Table 3B, Table 10, Table 11, Table X or Table Z1 or Z2), (ii) an endonuclease domain, and optionally (iii) a DNA-binding domain); and (b)(i)(i) A sequence of elements in Table 3B, Table 10, or Table X consisting of nucleotides located 5' to the start codon, or (ii) A sequence to which a polypeptide binds having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with a sequence of elements in Table 3B, Table 10, or Table X consisting of nucleotides located 3' to the stop codon. (i) A template RNA (or DNA encoding template RNA) This includes making contact with This results in the insertion of heterologous target sequences into the genome of human cells. Human cells do not show upregulation of DNA repair genes and / or tumor suppressor genes, or DNA repair genes and / or tumor suppressor genes are not upregulated by more than 50%, 25%, 10%, 5%, 2%, or 1% (where upregulation is measured, for example, by RNA-seq), method.
[0291] 287. A method for adding an exogenous coding region to the genome of a cell (e.g., a mammalian cell), comprising contacting the cell with (i) RNA containing the non-coding strand of the exogenous coding region (wherein optionally, the RNA does not contain the coding strand of the exogenous coding region), and (ii) a polypeptide containing the sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein optionally, the delivery includes nonviral delivery.
[0292] 288. A method for adding an exogenous coding region to the genome of a cell (e.g., a mammalian cell), comprising contacting the cell with (i) RNA containing the non-coding strand of the exogenous coding region (wherein optionally, the RNA does not contain the coding strand of the exogenous coding region), and (ii) a polypeptide containing a sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence in which the number of nucleotides different therefrom is 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or less, wherein optionally, the delivery includes nonviral delivery.
[0293] 289. A method for expressing a target polypeptide in cells (e.g., mammalian cells), comprising contacting the cells with a retrotransposase polypeptide comprising (i) RNA (wherein the RNA comprises a non-coding strand which is the reverse complement of the sequence that will encode the target polypeptide, and optionally the RNA does not comprise the coding sequence that will encode the target polypeptide), and (ii) an amino acid sequence encoded by a sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein optionally the delivery includes non-viral delivery.
[0294] 290. A method for expressing a target polypeptide in cells (e.g., mammalian cells), comprising contacting the cells with a retrotransposase polypeptide comprising (i) RNA (wherein the RNA comprises a non-coding strand which is the reverse complement of the sequence that will encode the target polypeptide, and optionally the RNA does not comprise the coding sequence that encodes the target polypeptide), and (ii) an amino acid sequence encoded by a sequence of elements of Table 3B, Table 10, Table 11, or Table X, or a sequence in which the number of nucleotides different therefrom is 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or less, wherein optionally the delivery includes nonviral delivery.
[0295] 291. A method for expressing a target polypeptide in cells (e.g., mammalian cells), comprising contacting the cells with a retrotransposase polypeptide comprising (i) RNA (wherein the RNA comprises a non-coding strand which is the reverse complement of the sequence that will encode the target polypeptide, and optionally the RNA does not comprise the coding sequence that encodes the target polypeptide), and (ii) an amino acid sequence encoded by a sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein optionally the delivery includes nonviral delivery.
[0296] 292. A method for expressing a target polypeptide in cells (e.g., mammalian cells), comprising contacting the cells with a retrotransposase polypeptide comprising (i) RNA (wherein the RNA comprises a non-coding strand which is the reverse complement of the sequence that will encode the target polypeptide, and optionally the RNA does not comprise the coding sequence that encodes the target polypeptide), and (ii) an amino acid sequence encoded by a sequence of elements of Table 3B, Table 10, Table 11, or Table X, or a sequence in which the number of nucleotides different therefrom is 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or less, wherein optionally the delivery includes nonviral delivery.
[0297] 293. Any of the above embodiments, wherein the sequence to be inserted into the mammalian genome is an exogenous sequence to the mammalian genome.
[0298] 294. The exogenous sequence inserted into the mammalian genome is not naturally present elsewhere in the mammalian genome, as described in any of the preceding embodiments.
[0299] 295. The exogenous sequence inserted into the mammalian genome is naturally occurring elsewhere in the mammalian genome, as described in any of the preceding embodiments.
[0300] 296. Any method of the preceding embodiments that operates independently of the DNA template.
[0301] 297. A cell is part of a tissue, according to any of the preceding embodiments.
[0302] 298. Any of the preceding embodiments, wherein the mammalian cells are euploid, not immortalized, part of an organism, primary cells, non-dividing, hepatocytes, or derived from a subject with a genetic disorder.
[0303] 299. Mammalian cells are present in a disease-free subject, for example, to supplement the subject's genome, as in any of the preceding embodiments.
[0304] 300. A method of any of the preceding embodiments, wherein contact involves bringing a cell into contact with a plasmid, virus, virus-like particle, virosome, liposome, vesicle, exosome, fusosome, or lipid nanoparticle.
[0305] 301. Contact is made by any of the methods of the preceding embodiments, including the use of nonviral delivery.
[0306] 302. Any method of the preceding embodiments, comprising contacting a cell with template RNA (or DNA encoding template RNA), wherein the template RNA comprises a non-coding strand of an exogenous coding region, optionally, the template RNA does not comprise a coding strand of an exogenous coding region, optionally, the delivery comprises non-viral delivery, thereby adding the exogenous coding region to the cell's genome.
[0307] 303. Any method of the preceding embodiments, comprising contacting a cell with template RNA (or DNA encoding template RNA), wherein the template RNA comprises a non-coding strand which is the reverse complement of a sequence encoding a polypeptide, optionally, the template RNA does not contain a coding strand encoding a polypeptide, optionally, the delivery comprises non-viral delivery, thereby adding an exogenous coding region to the cell's genome and causing the polypeptide to be expressed in the cell.
[0308] 304. Contact is a method of any of the preceding embodiments, which includes, for example, administering (a) and (b) intravenously to the target.
[0309] 305. A method of any of the prior embodiments, comprising administering doses (a) and (b) to a subject at least twice.
[0310] 306. A method according to any of the preceding embodiments, wherein a polypeptide reverse transcribes a template RNA sequence onto a target DNA strand, thereby modifying the target DNA strand.
[0311] 307. A method according to any of the preceding embodiments, wherein (a) and (b) are administered separately.
[0312] 308.(a) and (b) are administered together, according to either of the preceding embodiments.
[0313] The method of any of the preceding embodiments, wherein the nucleic acid of 309.(a) is not integrated into the genome of the host cell.
[0314] 310. The method according to any of the preceding embodiments, wherein the tissue is the liver, lungs, skin, muscle tissue (e.g., skeletal muscle), eye or ocular tissue, or the central nervous system.
[0315] 311. Any method according to the preceding embodiment, wherein the cells are hematopoietic stem cells (HSCs), T cells, or natural killer (NK) cells.
[0316] 312. The sequence that binds to the polypeptide has one or more of the following characteristics, in any of the above numbered methods: (a) is located at the 3' end of the template RNA; (b) is located at the 5' end of the template RNA; (b) is a non-code array; (c) is structured RNA; (d) forms at least one hairpin loop structure; and / or (e) is the guide RNA.
[0317] 313. Any number of the above methods, wherein the template RNA further comprises a sequence containing at least 20 nucleotides with at least 80% identity to the target DNA strand (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity).
[0318] 314. The template RNA further comprises any number of the above methods, wherein the template RNA comprises a sequence containing at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150 nucleotides with at least 80% identity to the target DNA strand (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100%).
[0319] 315. A sequence containing at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150 nucleotides, or approximately 2-10, 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 10-100, or 2-100 nucleotides, located at the 3' end of the template RNA, in any of the above numbered ways.
[0320] 316. Any of the above methods, wherein the template RNA further includes, for example, at the 3' end of the template RNA a sequence containing at least 100 nucleotides with at least 80% identity to the target DNA strand (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, 100% identity).
[0321] 317. A method according to any of the preceding embodiments, wherein a site in the target DNA strand having at least 80% sequence identity is located proximal to a target site on the target DNA strand (e.g., within about 0-10, 10-20, 20-30, 30-50 or 50-100 nucleotides) and is recognized (e.g., bound and / or cleaved) by a polypeptide containing an endonuclease.
[0322] 318. The target RNA is any of the above numbered methods, comprising a homology domain containing a sequence with at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to the sequence by the 3' homology arm in Table 11.
[0323] 319. The target RNA is any of the above numbered methods, comprising a sequence by the 5' homology arm in Table 11, or a homology domain having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0324] 320. Sequences containing at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150 nucleotides, or approximately 2-10, 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 10-100, or 2-100 nucleotides, are located at the 3' end of the template RNA. By any selection, any number of the above methods, wherein a site in the target DNA strand containing at least 80% sequence identity is located proximal to the target site on the target DNA strand (e.g., within approximately 0-10, 10-20, or 20-30 nucleotides) and recognized (e.g., bound and / or cleaved) by a polypeptide containing an endonuclease.
[0325] 321. The method according to any of the preceding embodiments, wherein the target site is a site in the human genome that has the closest identity to the natural target site of the polypeptide containing the endonuclease, for example, the target site in the human genome is identical to the natural target site by at least about 16, 17, 18, 19, or 20 nucleotides.
[0326] 322. The template RNA has at least 3, 4, 5, 6, 7, 8, 9, or 10 bases that are 100% identical to the target DNA strand, in any of the above numbered methods.
[0327] 323. At least 3, 4, 5, 6, 7, 8, 9, or 10 bases of 100% identity with respect to the target DNA strand are located at the 3' end of the template RNA, in any of the above numbered ways.
[0328] 324. At least 3, 4, 5, 6, 7, 8, 9, or 10 bases of 100% identity with respect to the target DNA strand are located at the 5' end of the template RNA, in any of the above numbered ways.
[0329] 325. The template RNA is any of the above numbered methods, wherein the template RNA contains at least 3, 4, 5, 6, 7, 8, 9, or 10 bases that are 100% identical to the target DNA strand at the 5' end of the template RNA, and at least 3, 4, 5, 6, 7, 8, 9, or 10 bases that are 100% identical to the target DNA strand at the 3' end of the template RNA.
[0330] 326. The heterologous target sequence is 50 to 50,000 base pairs (e.g., 50 to 40,000 bp, 500 to 30,000 bp, 500 to 20,000 bp, 100 to 15,000 bp, 500 to 10,000 bp, 50 to 10,000 bp, 50 to 5,000 bp), using any of the above numbered methods.
[0331] 327. The heterologous target sequence is at least 1, 10, 25, 50, 100, 150, 200, 250, 300, 400, 500, 600, or 700 bp, by any of the above numbered methods.
[0332] 328. The heterologous target sequence is at least 715, 750, 800, 950, 1,000, 2,000, 3,000, or 4,000 bp, by any of the above numbered methods.
[0333] 329. The heterologous target sequence is less than 5,000, 10,000, 15,000, 20,000, 30,000, or 40,000 bp, using any of the above numbered methods.
[0334] 330. The heterologous target sequence is any of the above numbered methods, having a length of less than 700, 600, 500, 400, 300, 200, 150, or 100 bp.
[0335] 331. Heterogeneous sequences are, (a) Open reading frames, for example sequences encoding polypeptides, for example enzymes (e.g., lysosomal enzymes), membrane proteins, blood factors, exons, intracellular proteins (e.g., organelle proteins such as cytoplasmic proteins, nucleoproteins, mitochondrial proteins, or lysosomal proteins), extracellular proteins, structural proteins, signaling proteins, regulatory proteins, transport proteins, sensory proteins, motor proteins, defense proteins, storage proteins, or immune receptor proteins (e.g., chimeric antigen receptor (CAR) proteins, T cell receptors, B cell receptors), or antibodies; (b) Non-coding and / or regulatory sequences, e.g., transcription modulators, e.g., sequences that bind to promoters, enhancers, insulators; (c) Splice receptor site; (d) Poly A area; (e) epigenetic modification sites; or (f) Gene expression unit. Any of the above numbered methods, including one or more of the above.
[0336] 332. Any of the above numbered methods, wherein the target DNA is a genome-safe harbor (GSH) site.
[0337] 333. The target DNA is a genomic Natural Harbor® site, as specified by any of the above numbered methods.
[0338] 334.1 Any of the above numbered methods, which result in insertion of a heterologous target sequence into the genome with an average copy number of at least 0.01, 0.025, 0.05, 0.075, 0.1, 0.15, 0.2, 0.25, 0.3, 0.4, 0.5, 0.75, 1, 1.25, 1.5, 1.75, 2, 2.5, 3, 4, or 5 copies per genome.
[0339] 335. Any of the above numbered methods, by assays described herein, for example, the assay in Example 6 of PCT Application No. PCT / US2019 / 048607, yield approximately 25-100%, 50-100%, 60-100%, 70-100%, 75-95%, and 80-90% of the implements to the target site in the untruncate genome.
[0340] 336. Any of the above numbered methods, which result in the insertion of a heterologous target sequence into only one target site in the cell's genome.
[0341] 337. A method of any of the above numbers resulting in the insertion of a heterologous target sequence into a target site in a cell, wherein the inserted heterologous sequence contains mutations (e.g., SNPs, or one or more deletions, e.g., truncations or internal deletions) of less than 10%, 5%, 2%, 1%, 0.5%, 0.2%, or 0.1% compared to the heterologous sequence before insertion, as measured by the assay of Example 12 of PCT Application No. PCT / US2019 / 048607.
[0342] 338. A method of any of the above numbers that results in the insertion of a heterologous target sequence into a target site in multiple cells, wherein less than 10%, 5%, 2%, or 1% of the inserted heterologous sequence copies, as measured by the assay of Example 12 of PCT Application No. PCT / US2019 / 048607, contains mutations (e.g., SNPs or deletions, e.g., truncations or internal deletions).
[0343] 339. Any of the above methods for resulting in the insertion of a heterologous target sequence into a target cell genome, wherein the target cells show no p53 upregulation or show p53 upregulation of less than 50%, 25%, 10%, 5%, 2%, or 1% (wherein, for example, p53 upregulation is measured by the p53 protein level or by the level of phosphorylated p53 at Ser15 and Ser20, for example, by the method described in Example 30).
[0344] 340. Any of the above-mentioned methods for resulting in the insertion of a heterologous target sequence into the genome of a target cell, wherein the target cell does not exhibit upregulation of DNA repair genes and / or tumor suppressor genes, or the DNA repair genes and / or tumor suppressor genes are not upregulated by more than 50%, 25%, 10%, 5%, 2%, or 1% (where, for example, upregulation is measured by RNA-seq) of any of the above-mentioned methods.
[0345] 341. Any of the above numbered methods, for example, a measurement using single-cell ddPCR as described in Example 17, in which approximately 1 to 80% of the cells in a population of cells in contact with the system, for example, approximately 1 to 10%, 10 to 20%, 20 to 30%, 30 to 40%, 40 to 50%, 50 to 60%, 60 to 70%, or 70 to 80% of the cells, result in the insertion of a heterologous target sequence into the target site (for example, with one insertion or two or more insertion copy numbers).
[0346] 342. For example, in measurements using colony isolation and ddPCR as described in Example 18, for example, in approximately 1-80% of cells in a cell population brought into contact with the system, for example, approximately 1-10%, 10-20%, 20-30%, 30-40%, 40-50%, 50-60%, 60-70%, or 70-80% of cells (for example, any of the above numbered methods, which result in the insertion of a heterologous target sequence into the target site with a copy number of a single insertion).
[0347] 343. Any of the above-mentioned methods that result in a higher ratio of insertion of a heterologous target sequence into a target site (on-target insertion) than insertion into a non-target site (off-target insertion) in a cell population, for example, an assay of Example 11 of PCT Application No. PCT / US2019 / 048607 in which the on-target insertion to off-target insertion ratio is 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, A method that is greater than 90:1, 100:1, 200:1, 500:1, or 1,000:1.
[0348] 344. Any of the above-mentioned methods for causing heterologous insertion of a target sequence in the presence of a DNA repair pathway inhibitor (e.g., SCR7, PARP inhibitor) or in a cell line lacking a DNA repair pathway (e.g., a cell line lacking a nucleotide excision repair pathway or homologous recombination repair pathway).
[0349] 345. Any of the methods of the preceding embodiments, wherein the cells have reduced Rad51 repair pathway activity, reduced expression of Rad51 or components of the Rad51 repair pathway, or lack the functional Rad51 repair pathway, for example, lack the functional Rad51 gene, for example, a mutation (e.g., deletion) that inactivates one or both copies of the Rad51 gene or another gene in the Rad51 repair pathway.
[0350] 346. Any of the above numbered systems formulated as a pharmaceutical composition.
[0351] 347. Any of the above numbered systems, arranged within a pharmaceutically acceptable carrier (e.g., vesicles, liposomes, natural or synthetic lipid bilayers, lipid nanoparticles, exosomes).
[0352] 348. Any of the above numbered methods resulting in the insertion of multiple (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, or 10) heterologous target sequences into the target cell genome.
[0353] 349. Multiple insertions occur simultaneously or sequentially, according to any of the preceding embodiments.
[0354] 350. Any of the above numbered methods resulting in the deletion of multiple (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, or 10) heterologous target sequences into the target cell genome.
[0355] 351. Multiple deletions occur simultaneously or sequentially, according to any of the preceding embodiments.
[0356] 352. Any of the above numbered methods for causing multiple (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, or 10) heterologous base changes in a target cell genome.
[0357] 353. A method according to any of the preceding embodiments, wherein multiple base changes occur simultaneously or sequentially.
[0358] 354. Any of the above numbered methods, comprising contacting cells with multiple distinct template RNAs, each containing a different target sequence.
[0359] 355. A method according to any of the preceding embodiments, wherein the distinct template RNA contains at least two distinct heterologous target sequences (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, or 10).
[0360] 356. A method according to any of the preceding embodiments, wherein at least two distinct heteropurpose sequences each contain a distinct payload.
[0361] 357. A method according to any of the preceding embodiments, wherein at least two distinct heteropurpose sequences each contain the same payload.
[0362] 358. The target cells of the Gene Writing system are any number of embodiments described above, which have been previously modified at one or more gene loci.
[0363] 359. The previously edited cell is a T cell, in any of the above numbered embodiments.
[0364] 360. Any number of embodiments described above, in which one or more prior modifications are selected from, for example, gene knockouts of endogenous TCRs (e.g., TRAC, TRBC), HLA class I (B2M), PD1, CD52, CTLA-4, TIM-3, LAG-3, or DGK.
[0365] 361. The heterologous target sequence is any one of the above-mentioned embodiments, including a TCR or CAR.
[0366] 362. A method for creating a system for modifying the genome of mammalian cells, a) (i) A sequence that binds to a polypeptide comprising a reverse transcriptase domain and an endonuclease domain, and which has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the 5'UTR sequence or 3'UTR sequence of the element sequence in Table X; and (ii) A template RNA comprising a heterologous target sequence. A method that includes this.
[0367] 363.b) Processing the template RNA to reduce its secondary structure, for example, by heating the template RNA to, for example, at least 70, 75, 80, 85, 90 or 95°C, and / or c) subsequently cooling the template RNA to, for example, a temperature that allows for secondary structure, for example, 37, 30, 25 or 20°C or below. A method of any of the preceding embodiments, further including the above.
[0368] 364. A method for constructing a system for modifying DNA (for example, as described herein), (a) Prepare a template nucleic acid (e.g., template RNA or DNA) containing heterologous sequences that have at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% homology to sequences contained within the target DNA molecule, and / or (b) Prepare a polypeptide system that includes a heterologous targeting domain that specifically binds to a sequence contained within the target DNA molecule (e.g., including a DNA-binding domain (DBD) and / or an endonuclease domain). A method that includes this.
[0369] 365. (a) comprising introducing a heterologous sequence into a template nucleic acid (e.g., template RNA or DNA) which has at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% homology to a sequence contained in the target DNA molecule, and / or (b) The system comprises introducing heterologous targeting domains into the polypeptide of the system (e.g., including DNA-binding domains (DBDs) and / or endonuclease domains) that specifically bind to sequences contained within a target DNA molecule. Any of the methods of the previous embodiments.
[0370] The introduction of 366.(a) is a method of any of the preceding embodiments, which includes inserting a homologous sequence into a template nucleic acid.
[0371] The introduction of 367.(a) is a method of any of the preceding embodiments, which includes substituting a segment of a template nucleic acid with a homologous sequence.
[0372] The introduction of 368.(a) is a method of any of the preceding embodiments, comprising generating a segment of a template nucleic acid having a homologous sequence by mutating one or more nucleotides of the template nucleic acid (e.g., at least 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, or 100 nucleotides).
[0373] The introduction of 369.(b) is a method of any of the preceding embodiments, comprising inserting the amino acid sequence of the targeting domain into the amino acid sequence of the polypeptide.
[0374] The introduction of 370.(b) is a method of any of the preceding embodiments, comprising inserting a nucleic acid sequence encoding a targeting domain into the coding sequence of a polypeptide contained within a nucleic acid molecule.
[0375] The introduction of 371.(b) is any of the methods of the prior embodiments, comprising substituting at least a portion of the polypeptide with a targeting domain.
[0376] The introduction of 372.(a) is a method of any of the preceding embodiments, comprising mutating one or more amino acids of a polypeptide (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500, or more amino acids).
[0377] 373. A method for modifying a target site in intracellular genomic DNA, cells, (a) polypeptides or nucleic acids encoding polypeptides (polypeptides include (i) a reverse transcriptase (RT) domain, (ii) a DNA-binding domain (DBD); and (iii) an endonuclease domain, such as a nickas domain); and (b) (e.g., from 5' to 3') (i) optionally, a sequence that binds to a target site (e.g., the second strand of the site in the target genome), (ii) optionally, a sequence that binds to a polypeptide, (iii) a heterologous target sequence, and (iv) a template RNA (or DNA encoding the template RNA) containing a 3' target homology domain. (i) The polypeptide contains a heterologous targeting domain (for example, within a DBD or endonuclease domain) that specifically binds to a sequence located within or adjacent to a target site in genomic DNA; and / or (ii) The template RNA contains heterologous sequences that have at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% homology to sequences contained within or adjacent to the target site of the genomic DNA. A method comprising modifying a target site in intracellular genomic DNA by bringing it into contact with a substance.
[0378] 374. A method for creating a system for modifying the genome of mammalian cells, a)(i) To provide a template RNA containing a sequence that binds to a polypeptide comprising a reverse transcriptase domain and an endonuclease domain, and which has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with (i) a nucleotide located 5' to the start codon of the sequence of an element in Table 3B, Table 10, or Table X (e.g., including a retrotransposase binding region) or (ii) a nucleotide located 3' to the stop codon of the sequence of an element in Table 3B, Table 10, or Table X (e.g., including a retrotransposase binding region). Methods including
[0379] 375. b) Processing the template RNA to reduce its secondary structure, for example, by heating the template RNA to, for example, at least 70, 75, 80, 85, 90 or 95°C, and / or c) Next, the template RNA is cooled to a temperature that allows for secondary structure formation, for example, 37, 30, 25, or 20°C or below. A method of any of the preceding embodiments, further including the above.
[0380] 376. A method of any of the prior embodiments, wherein the system is any of the systems of the prior embodiments.
[0381] 377. Any method of the preceding embodiment, further comprising contacting a template RNA with a polypeptide comprising (i) a reverse transcriptase domain and (ii) an endonuclease domain, or with a nucleic acid (e.g., RNA) encoding the polypeptide.
[0382] 378. Any method of the above embodiments, further comprising bringing template RNA into contact with a cell.
[0383] 379. A system or method of any of the preceding embodiments in which a heterologous target sequence encodes a therapeutic polypeptide.
[0384] 380. A system or method of any of the preceding embodiments that encodes a heterologous (e.g., human) polypeptide, or a fragment or variant thereof.
[0385] 381. A system or method of any of the preceding embodiments in which a heterologous target sequence encodes an enzyme (e.g., a lysosomal enzyme), a blood factor (e.g., factors I, II, V, VII, X, XI, XII, or XIII), a membrane protein, an exon, an intracellular protein (e.g., organelle proteins such as cytoplasmic proteins, nucleoproteins, mitochondrial proteins, or lysosomal proteins), an extracellular protein, a structural protein, a signaling protein, a regulatory protein, a transport protein, a sensory protein, a motor protein, a protective protein, a storage protein, an immune receptor protein (e.g., a chimeric antigen receptor (CAR) protein, a T cell receptor, a B cell receptor), or an antibody.
[0386] 382. A system or method of any of the preceding embodiments, comprising a heterologous target sequence, a tissue-specific promoter, or an enhancer.
[0387] 383. A system or method according to any of the previous embodiments, wherein the heterologous target sequence encodes a polypeptide consisting of 250, 300, 400, 500, or more than 1,000 amino acids, and optionally up to 1,300 amino acids.
[0388] 384. A heterologous target sequence codes for a fragment of a mammalian gene but does not code for a complete mammalian gene, for example, one or more exons but not a full-length protein, according to any system or method of the preceding embodiments.
[0389] 385. A system or method of any of the preceding embodiments in which a heterologous sequence encodes one or more introns.
[0390] 386. Any system or method of the preceding embodiments wherein the heterologous target sequence is other than GFP, for example, other than a fluorescent protein or other than a reporter protein.
[0391] 387. A system or method according to any of the previous embodiments, wherein the polypeptide exhibits activity at 37°C that is 70%, 75%, 80%, 85%, 90%, or 95% or more of the activity at 25°C under otherwise the same conditions.
[0392] 388. A system or method of any of the preceding embodiments in which the nucleic acid encoding the polypeptide and the template RNA or the nucleic acid encoding the template RNA are separate nucleic acids.
[0393] 389. Any system or method of the preceding embodiments in which the template RNA does not encode an active reverse transcriptase (for example, including, for example, an inactivated mutant reverse transcriptase or not including a reverse transcriptase sequence, as described in Example 1 or 2).
[0394] 390. A system or method of any of the preceding embodiments, wherein the template RNA includes one or more chemical modifications.
[0395] 391. A system or method of any of the preceding embodiments in which a heterologous target sequence is placed between the promoter and the polypeptide-binding sequence.
[0396] 392. A system or method of any of the preceding embodiments in which a heterologous target sequence is placed between the promoter and the polypeptide-binding sequence.
[0397] 393. A system or method of any of the preceding embodiments, wherein the heterologous target sequence includes an open reading frame (or its reverse complement) in the 5' to 3' direction of the template RNA.
[0398] 394. A system or method of any of the preceding embodiments, wherein the heterologous target sequence includes an open reading frame (or its reverse complement) in the 3' to 5' direction of the template RNA.
[0399] 395. A system or method according to any of the preceding embodiments, wherein the polypeptide comprises (a) a reverse transcriptase domain and (b) an endonuclease domain, and at least one of (a) and (b) is heterogeneous.
[0400] 396. The system or method according to any of the above embodiments, wherein the polypeptide comprises (a) a target DNA binding domain, (b) a reverse transcriptase domain, and (c) an endonuclease domain, and at least one of (a), (b), and (c) is heterogeneous.
[0401] 397. A substantially pure polypeptide comprising (a) a reverse transcriptase domain and (b) a heterologous endonuclease domain, wherein one or both of (a) and (b) have an amino acid sequence encoded by a sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0402] 398. A substantially pure polypeptide comprising (a) a reverse transcriptase domain and (b) a heterologous endonuclease domain, wherein one or both of (a) and (b) have an amino acid sequence encoded by a sequence of an element in Table 3B, Table 10, Table 11, or Table X, or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom.
[0403] 399. A substantially pure polypeptide comprising (a) a reverse transcriptase domain and (b) a heterologous endonuclease domain, wherein one or both of (a) and (b) have an amino acid sequence of an element in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0404] 400. A substantially pure polypeptide comprising (a) a reverse transcriptase domain and (b) a heterologous endonuclease domain, wherein one or both of (a) and (b) have sequences in which the amino acid sequences of the elements of Table 3B, Table 10, Table 11, or Table X, or the number of different nucleotides is 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or less.
[0405] 401. A substantially pure polypeptide comprising (a) a reverse transcriptase domain encoded by a first sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and (b) an endonuclease domain encoded by a second sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein the first and second sequences are selected from elements in different rows of Table 3B, Table 10, Table 11, or Table X.
[0406] 402. A substantially pure polypeptide comprising (a) a reverse transcriptase domain encoded by a first sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer different nucleotides, and (b) an endonuclease domain encoded by a second sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer different nucleotides, wherein the first and second sequences are selected from elements in different rows of Table 3B, Table 10, Table 11, or Table X.
[0407] 403. A substantially pure polypeptide comprising (a) a reverse transcriptase domain of a first sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and (b) an endonuclease domain of a second sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein the first and second sequences are selected from elements in different rows of Table 3B, Table 10, Table 11, or Table X.
[0408] 404. A substantially pure polypeptide comprising (a) a reverse transcriptase domain of a first sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer different nucleotides, and (b) an endonuclease domain of a second sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer different nucleotides, wherein the first and second sequences are selected from elements in different rows of Table 3B, Table 10, Table 11, or Table X.
[0409] 405. A substantially pure polypeptide of any of the prior embodiments, further comprising a target DNA-binding domain encoded by a sequence of an element in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0410] 406. A substantially pure polypeptide of any of the prior embodiments, further comprising a target DNA-binding domain encoded by a sequence of elements in Table 3B, Table 10, Table 11, or Table X, or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom.
[0411] 407. A substantially pure polypeptide of any of the prior embodiments, further comprising a target DNA-binding domain of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0412] 408. A substantially pure polypeptide of any of the prior embodiments, further comprising a target DNA-binding domain of a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides that are elements of Table 3B, Table 10, Table 11, or Table X, or different therefrom.
[0413] 409. A substantially pure polypeptide comprising (a) a reverse transcriptase domain and (b) a heterologous endonuclease domain, wherein (a) has an amino acid sequence listed in Table 3B, Table 10, Table 11, or Table Z1 or Table Z2, or Table X, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and optionally, (b) has an amino acid sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0414] 410. A substantially pure polypeptide comprising (a) a reverse transcriptase domain and (b) a heterologous endonuclease domain, wherein (a) has an amino acid sequence listed in Table 3B, Table 10, Table 11, or Table Z1 or Table Z2, or Table X, or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom, and optionally, (b) has an amino acid sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom.
[0415] 411. A substantially pure polypeptide of any of the prior embodiments, further comprising a target DNA-binding domain containing an amino acid sequence of an element in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0416] 412. A substantially pure polypeptide of any of the prior embodiments, further comprising a target DNA-binding domain, which comprises an amino acid sequence of an element in Table 3B, Table 10, Table 11, or Table X, or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom.
[0417] 413. A substantially pure polypeptide comprising (a) a target DNA-binding domain encoded by a first sequence listed in Table 3B, Table 10, Table 11, or Table X, (b) a reverse transcriptase domain encoded by a second sequence listed in Table 3B, Table 10, Table 11, or Table X, and (c) an endonuclease domain encoded by a third sequence listed in Table 3B, Table 10, Table 11, or Table X, (i) The first and second sequences are selected from different rows of Table 3B, Table 10, Table 11, or Table X; (ii) The first and third sequences are selected from different rows of Table 3B, Table 10, Table 11, or Table X; (iii) The second and third sequences are selected from different rows of Table 3B, Table 10, Table 11, or Table X; or (iv) A polypeptide in which the first sequence, the second sequence, and the third sequence are each selected from different rows of Table 3B, Table 10, Table 11, or Table X.
[0418] 414. A substantially pure polypeptide comprising (a) a target DNA binding domain of a first sequence listed in Table 3B, Table 10, Table 11, or Table X, (b) a reverse transcriptase domain of a second sequence listed in Table 3B, Table 10, Table 11, or Table X, and (c) an endonuclease domain of a third sequence listed in Table 3B, Table 10, Table 11, or Table X, (i) The first and second sequences are selected from different rows of Table 3B, Table 10, Table 11, or Table X; (ii) The first and third sequences are selected from different rows of Table 3B, Table 10, Table 11, or Table X; (iii) The second and third sequences are selected from different rows of Table 3B, Table 10, Table 11, or Table X; or (iv) A polypeptide in which the first sequence, the second sequence, and the third sequence are each selected from different rows of Table 3B, Table 10, Table 11, or Table X.
[0419] 415. (a) A target DNA-binding domain comprising, for example, a first amino acid sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; (b) A second amino acid sequence listed in Table Z1 or Z2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. A substantially pure polypeptide comprising (c) a reverse transcriptase domain and an endonuclease domain comprising a third amino acid sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein the elements of the first amino acid sequence and the third amino acid sequence are optionally selected from different rows of Table 3B, Table 10, Table 11, or Table X.
[0420] 416. A substantially pure polypeptide comprising (a) a target DNA binding domain containing, for example, a sequence in which the first amino acid sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a polypeptide different therefrom, is at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100; (b) a reverse transcriptase domain containing a sequence in which the second amino acid sequence of an element listed in Table Z1 or Z2, or a polypeptide different therefrom, is at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100; and (c) an endonuclease domain containing a sequence in which the third amino acid sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, or a polypeptide different therefrom, is at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100; A polypeptide in which, by optional selection, the elements of the first and third amino acid sequences are selected from different rows in Table 3B, Table 10, Table 11, or Table X.
[0421] 417. A substantially pure polypeptide comprising (a) a target DNA-binding domain comprising a first amino acid sequence, (b) a reverse transcriptase domain comprising a second amino acid sequence, and (c) an endonuclease domain comprising a third amino acid sequence, wherein the first, second, and third amino acid sequences are each encoded by sequences listed in Table 3B, Table 10, Table 11, or Table X, or comprise amino acid sequences listed in Table Z1 or amino acid sequences of domains listed in Table Z2.
[0422] 418. A substantially pure polypeptide of any of the preceding embodiments, the first amino acid sequence being encoded by the sequences listed in Table 3B, Table 10, Table 11, or Table X.
[0423] 419. A substantially pure polypeptide according to any of the above embodiments, wherein the first amino acid sequence comprises an amino acid sequence listed in Table Z1 or Z2.
[0424] 420. A substantially pure polypeptide of any of the previous embodiments, wherein the second amino acid sequence is encoded by the sequences listed in Table 3B, Table 10, Table 11, or Table X.
[0425] 421. A substantially pure polypeptide according to any of the above embodiments, wherein the second amino acid sequence comprises an amino acid sequence listed in Table Z1 or Z2.
[0426] 422. A substantially pure polypeptide of any of the previous embodiments, wherein the third amino acid sequence is encoded by a sequence listed in Table 3B, Table 10, Table 11, or Table X.
[0427] 423. A substantially pure polypeptide according to any of the above embodiments, wherein the third amino acid sequence comprises an amino acid sequence listed in Table Z1 or Z2.
[0428] 424. A polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide comprises (a) a reverse transcriptase domain and (b) an endonuclease domain, and one or both of (a) and (b) have an amino acid sequence encoded by a sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and at least one of (a) and (b) is heterologous to the other.
[0429] 425. A polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide comprises (a) a reverse transcriptase domain and (b) an endonuclease domain, and one or both of (a) and (b) have an amino acid sequence encoded by a sequence of elements in Table 3B, Table 10, Table 11, or Table X, or a sequence in which there are 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides other than those therein, and at least one of (a) and (b) is heterologous to the other.
[0430] 426. A polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide is a fusion protein comprising (a) a reverse transcriptase domain encoded by a first sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, and (b) an endonuclease domain encoded by a second sequence of an element listed in Table 3B, Table 10, Table 11, or Table X, wherein the first and second sequences are selected from elements in different rows of Table 3B, Table 10, Table 11, or Table X, and the polypeptide or nucleic acid encoding a polypeptide.
[0431] 427. A polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide comprises (a) a target DNA-binding domain, (b) a reverse transcriptase domain, and (c) an endonuclease domain, and at least one of (a), (b), and (c) (e.g., one, two, or all) comprises an amino acid sequence encoded by a sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and at least one of (a), (b), and (c) is heterologous to the others.
[0432] 428. A polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide comprises (a) a target DNA-binding domain, (b) a reverse transcriptase domain, and (c) an endonuclease domain, and at least one of (a), (b), and (c) (for example, one, two, or all) comprises an amino acid sequence encoded by a sequence of elements of Table 3B, Table 10, Table 11, or Table X, or a sequence having 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom, and at least one of (a), (b), and (c) is heterologous to the others.
[0433] 429. A polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide is a fusion protein comprising (a) a target DNA-binding domain encoded by a first sequence listed in Table 3B, Table 10, Table 11, or Table X, (b) a reverse transcriptase domain encoded by a second sequence listed in Table 3B, Table 10, Table 11, or Table X, and (c) an endonuclease domain encoded by a third sequence listed in Table 3B, Table 10, Table 11, or Table X. (i) The first and second sequences are selected from different rows of Table 3B, Table 10, Table 11, or Table X; (ii) The first and third sequences are selected from different rows of Table 3B, Table 10, Table 11, or Table X; (iii) The second and third sequences are selected from different rows of Table 3B, Table 10, Table 11, or Table X; or (iv) The first sequence, the second sequence, and the third sequence are each selected from different rows of Table 3B, Table 10, Table 11, or Table X, and are polypeptides or nucleic acids encoding polypeptides.
[0434] 430. A polypeptide of any of the preceding embodiments, wherein the reverse transcriptase domain has at least 80% identity (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, or 100% identity) to a reverse transcriptase domain encoded by the sequence of elements in Table 3B, Table 10, Table 11, or Table X.
[0435] 431. A polypeptide of any of the prior embodiments, wherein the endonuclease domain has at least 80% identity, e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, or 100% identity, to an endonuclease domain encoded by a sequence of elements in Table 3B, Table 10, Table 11, or Table X.
[0436] 432. A polypeptide or method of any of the prior embodiments, wherein the DNA-binding domain has at least 80% identity, e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, or 100% identity, to the DNA-binding domain encoded by the sequence of elements in Table 3B, Table 10, Table 11, or Table X.
[0437] 433. A polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide comprises (a) a reverse transcriptase domain and (b) an endonuclease domain, and one or both of (a) and (b) have an amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and at least one of (a) and (b) is heterogeneous to the other, wherein the polypeptide or nucleic acid encoding a polypeptide.
[0438] 434. A polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide comprises (a) a reverse transcriptase domain and (b) an endonuclease domain, and one or both of (a) and (b) have an amino acid sequence of an element sequence of Table 3B, Table 10, Table 11, or Table X, or a sequence in which there are 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom, and at least one of (a) and (b) is heterologous to the other, wherein the polypeptide or nucleic acid encoding a polypeptide.
[0439] 435. A polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide is a fusion protein comprising (a) a reverse transcriptase domain of a first sequence listed in Table 3B, Table 10, Table 11, or Table X, and (b) an endonuclease domain of a second sequence listed in Table 3B, Table 10, Table 11, or Table X, wherein the first and second sequences are selected from different rows of Table 3B, Table 10, Table 11, or Table X.
[0440] 436. A polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide comprises (a) a target DNA-binding domain, (b) a reverse transcriptase domain, and (c) an endonuclease domain, and at least one of (a), (b), and (c) (for example, one, two, or all) comprises an amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and at least one of (a), (b), and (c) is heterogeneous to the others, wherein the polypeptide or nucleic acid encoding a polypeptide.
[0441] 437. A polypeptide or nucleic acid encoding a polypeptide, comprising (a) a target DNA-binding domain, (b) a reverse transcriptase domain, and (c) an endonuclease domain, wherein at least one of (a), (b), and (c) (for example, one, two, or all) comprises an amino acid sequence of an element of Table 3B, Table 10, Table 11, or Table X, or a sequence in which there are 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 or fewer nucleotides different therefrom, and at least one of (a), (b), and (c) is heterologous to the others, wherein the polypeptide or nucleic acid encoding a polypeptide.
[0442] 438. A polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide is a fusion protein comprising (a) a target DNA-binding domain of a first sequence listed in Table 3B, Table 10, Table 11, or Table X, (b) a reverse transcriptase domain of a second sequence listed in Table 3B, Table 10, Table 11, or Table X, and (c) an endonuclease domain of a third sequence listed in Table 3B, Table 10, Table 11, or Table X. (i) The first and second sequences are selected from different rows of Table 3B, Table 10, Table 11, or Table X; (ii) The first and third sequences are selected from different rows of Table 3B, Table 10, Table 11, or Table X; (iii) The second and third sequences are selected from different rows of Table 3B, Table 10, Table 11, or Table X; or (iv) The first sequence, the second sequence, and the third sequence are each selected from different rows of Table 3B, Table 10, Table 11, or Table X, and are polypeptides or nucleic acids encoding polypeptides.
[0443] 439. A polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA-binding domain (DBD), and (iii) an endonuclease domain, for example, a nickasase domain, and the RT domain has a sequence of Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0444] 440. A polypeptide of any of the previous embodiments, wherein the reverse transcriptase domain has at least 80% identity (e.g., at least 85%, 90%, 95%, 97%, 98%, 99%, or 100% identity) with respect to the reverse transcriptase domain of the elements in Table 3B, Table 10, Table 11, or Table X.
[0445] 441. A polypeptide of any of the preceding embodiments, wherein the endonuclease domain has at least 80% identity to the endonuclease domain of the elements in Table 3B, Table 10, Table 11, or Table X, for example, at least 85%, 90%, 95%, 97%, 98%, 99%, or 100% identity.
[0446] 442. A polypeptide or method of any of the prior embodiments, wherein the DNA-binding domain has at least 80% identity with respect to the DNA-binding domain of the element sequence of Table 3B, Table 10, Table 11, or Table X, for example, at least 85%, 90%, 95%, 97%, 98%, 99%, or 100% identity.
[0447] 443. A nucleic acid encoding a polypeptide of any of the embodiments described above.
[0448] 444. A vector containing a nucleic acid of any of the above embodiments.
[0449] 445. A host cell containing the nucleic acid of any of the preceding embodiments.
[0450] 446. A host cell containing a polypeptide of any of the preceding embodiments.
[0451] 447. A host cell containing a vector of any of the preceding embodiments.
[0452] 448. Heterogeneous target sequences at target sites in chromosomes (e.g., sequences encoding therapeutic polypeptides), and (a) one or both of the untranslated regions on one side (e.g., upstream) of the heterotarget sequence (e.g., retrotransposon untranslated sequences, e.g., the sequences in Table 10 or the sequences in column 5 of Table 3A or 3B) and the untranslated regions on the other side (e.g., downstream) of the heterotarget sequence (e.g., retrotransposon untranslated sequences, e.g., the sequences in Table 10 or the sequences in column 6 of Table 3A or 3B), and / or (b) polypeptides or nucleic acids encoding polypeptides (wherein a polypeptide comprises (i) a reverse transcriptase domain and (ii) an endonuclease domain, and one or both of (i) and (ii) have an amino acid sequence encoded by a sequence of an element in Table 3B, Table 10, Table 11 or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto). one or both A host cell containing (e.g., a human cell).
[0453] 449. Heterogeneous target sequences at target sites in chromosomes (e.g., sequences encoding therapeutic polypeptides), and (a) one or both of the untranslated regions on one side (e.g., upstream) of the heterotarget sequence (e.g., retrotransposon untranslated sequences, e.g., the sequences in Table 10 or the sequences in column 5 of Table 3A or 3B) and the untranslated regions on the other side (e.g., downstream) of the heterotarget sequence (e.g., retrotransposon untranslated sequences, e.g., the sequences in Table 10 or the sequences in column 6 of Table 3A or 3B), and / or (b) polypeptides or nucleic acids encoding polypeptides (wherein a polypeptide comprises (i) a reverse transcriptase domain and (ii) an endonuclease domain, and one or both of (i) and (ii) have an amino acid sequence encoded by a sequence of elements in Table 3B, Table 10, Table 11 or Table X, or a sequence in which there are 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 or fewer different nucleotides) one or both A host cell containing (e.g., a human cell).
[0454] 450. Heterogeneous target sequences at target sites in chromosomes (e.g., sequences encoding therapeutic polypeptides), and (a) one or both of the untranslated regions on one side (e.g., upstream) of the heterotarget sequence (e.g., retrotransposon untranslated sequences, e.g., the sequences in Table 10 or the sequences in column 5 of Table 3A or 3B) and the untranslated regions on the other side (e.g., downstream) of the heterotarget sequence (e.g., retrotransposon untranslated sequences, e.g., the sequences in Table 10 or the sequences in column 6 of Table 3A or 3B), and / or (b) polypeptides or nucleic acids encoding polypeptides (wherein a polypeptide comprises (i) a reverse transcriptase domain and (ii) an endonuclease domain, and one or both of (i) and (ii) have an amino acid sequence of an element of Table 3B, Table 10, Table 11 or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto). one or both A host cell containing (e.g., a human cell).
[0455] 451. Heterogeneous target sequences at target sites in chromosomes (e.g., sequences encoding therapeutic polypeptides), and (a) one or both of the untranslated regions on one side (e.g., upstream) of the heterotarget sequence (e.g., retrotransposon untranslated sequences, e.g., the sequences in Table 10 or the sequences in column 5 of Table 3A or 3B) and the untranslated regions on the other side (e.g., downstream) of the heterotarget sequence (e.g., retrotransposon untranslated sequences, e.g., the sequences in Table 10 or the sequences in column 6 of Table 3A or 3B), and / or (b) polypeptides or nucleic acids encoding polypeptides (wherein a polypeptide comprises (i) a reverse transcriptase domain and (ii) an endonuclease domain, and one or both of (i) and (ii) have an amino acid sequence of an element sequence from Table 3B, Table 10, Table 11 or Table X, or a sequence in which there are 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 or fewer different nucleotides) one or both A host cell containing (e.g., a human cell).
[0456] 452.(i) A host cell of any of the preceding embodiments comprising a heterologous target sequence (e.g., a sequence encoding a therapeutic polypeptide) at a target site in the chromosome, wherein the target locus is a Natural Harbor® site, e.g., a site in Table 4 of this specification.
[0457] 453.(ii) A host cell of any of the prior embodiments, further comprising one or both of the 5' untranslated region and the 3' untranslated region of the heterologous target sequence.
[0458] 454.(ii) A host cell of any of the preceding embodiments, further comprising an untranslated region on one side (e.g., upstream) of a heterologous target sequence (e.g., a retrotransposon untranslated sequence, e.g., the sequence in Table 10 or the sequence in column 5 of Table 3A or 3B) and an untranslated region on the other side (e.g., downstream) of the heterologous target sequence (e.g., a retrotransposon untranslated sequence, e.g., the sequence in Table 10 or the sequence in column 6 of Table 3A or 3B).
[0459] 455. A host cell of any of the previous embodiments, containing a heterologous target sequence only at the target site.
[0460] 456. A pharmaceutical composition comprising any of the above-mentioned systems, nucleic acids, polypeptides, or vectors; and a pharmaceutically acceptable excipient or carrier.
[0461] 457. A pharmaceutical composition of any of the preceding embodiments, wherein the pharmaceutically acceptable excipient or carrier is selected from vectors (e.g., viral or plasmid vectors), vesicles (e.g., liposomes, exosomes, natural or synthetic lipid bilayers), fusosomes, or lipid nanoparticles.
[0462] 458. A polypeptide of any of the preceding embodiments, further comprising a nuclear localization sequence.
[0463] 459. (For example, from 5' to 3') (i) a sequence that binds to a target site (e.g., an unedited strand of a site in the target genome), (ii) a sequence that binds to the endonuclease and / or DNA-binding domain of a polypeptide, (iii) a heterologous target sequence, and (iv) a template RNA (or DNA encoding template RNA) containing a 3' homology domain.
[0464] A template RNA of any of the previous embodiments, including 460.(i).
[0465] A template RNA of any of the preceding embodiments, including 461.(ii).
[0466] 462. A template RNA (or DNA encoding template RNA) comprising (i) a sequence that binds to a target site (e.g., an unedited strand of a site in the target genome), (ii) a sequence that specifically binds to the RT domain of a polypeptide, (iii) a heterologous target sequence, and (iv) a 3' homology domain (for example, from 5' to 3').
[0467] 463. The RT domain is a template RNA of any of the previous embodiments, comprising a sequence selected from Table 3B, Table 10, Table 11, or Table X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0468] 464.(v) A template RNA of any of the previous embodiments, further comprising a sequence that binds to the endonuclease and / or DNA-binding domain of a polypeptide (e.g., the same polypeptide including the RT domain).
[0469] The sequence of 465.(ii) is a template RNA of any of the embodiments described above, which specifically binds to an RT domain of Table 3B, Table 10, Table 11, or Table X, or to an RT domain sequence having at least 70, 75, 80, 85, 90, 95, or 99% identity thereto.
[0470] 466. A template RNA of any of the above embodiments, wherein the sequence that specifically binds to the RT domain is one of the sequences in Table 10, Table 11, or Table X, 3A, or 3B, or a sequence that is at least 70, 75, 80, 85, 90, 95, or 99% identical thereto.
[0471] A template RNA (or DNA encoding template RNA) comprising (ii) a sequence that binds to the endonuclease and / or DNA-binding domain of a polypeptide, (i) a sequence that binds to a target site (e.g., an unedited strand of a site in the target genome), (iii) a heterologous target sequence, and (iv) a 3' homology domain, extending from 467.5' to 3'.
[0472] A template RNA (or DNA encoding template RNA) comprising (iii) a heterologous target sequence, (iv) a 3' homology domain, (i) a sequence that binds to a target site (e.g., an unedited strand of a site in the target genome), and (ii) a sequence that binds to the endonuclease and / or DNA-binding domain of a polypeptide, from 468.5' to 3'.
[0473] 469. The template RNA, first template RNA, or second template RNA is a system or template RNA according to any of the preceding embodiments, comprising a sequence that specifically binds to the RT domain.
[0474] 470. The sequence that specifically binds to the RT domain is the system or template RNA of any of the previous embodiments, positioned between (i) and (ii).
[0475] 471. The sequence that specifically binds to the RT domain is any of the systems or template RNAs of the previous embodiments, positioned between (ii) and (iii).
[0476] 472. The sequence that specifically binds to the RT domain is the system or template RNA of any of the previous embodiments, located between (iii) and (iv).
[0477] 473. The sequence that specifically binds to the RT domain is any of the systems or template RNAs of the previous embodiments, located between (ivi) and (i).
[0478] 474. The sequence that specifically binds to the RT domain is any of the systems or template RNAs of the previous embodiments, located between (i) and (iii).
[0479] 475. (a)(i) A sequence that binds to the endonuclease domain of a polypeptide, e.g., the nickasase domain and / or the DNA-binding domain (DBD), and (ii) a first template RNA containing a sequence that binds to a target site (e.g., an unedited strand of a site in the target genome) (where, for example, the first RNA includes a gRNA); (b) (i) a sequence that specifically binds to the reverse transcriptase (RT) domain of a polypeptide (e.g., the polypeptide of (a)), (ii) a target site binding sequence (TSBS), and (iii) a second template RNA (or DNA encoding the second template RNA) comprising the RT template sequence. A system for modifying DNA, including [specific DNA components].
[0480] 476. A system according to any of the preceding embodiments, wherein the nucleic acid encoding the first template RNA and the nucleic acid encoding the second template RNA are two distinct nucleic acids.
[0481] 477. A system according to any of the previous embodiments, wherein the nucleic acid encoding the first template RNA and the nucleic acid encoding the second template RNA are part of the same nucleic acid molecule and, for example, reside on the same vector.
[0482] 478. A method for modifying a target DNA strand in a cell, tissue, or subject, comprising administering any of the above numbered systems to a cell, tissue, or subject to thereby modify the target DNA strand.
[0483] 479. The element sequence of Table X is selected from Vingi-1 EE, BovB, AviRTE_Brh, Penelope_SM, and Utopia_Dyak retrotransposases, in any of the above-mentioned embodiments.
[0484] 480. The template RNA is any one of the above numbered embodiments, comprising the 5'UTR and 3'UTR of the same sequence of elements in Table 3B, Table 10, or Table X.
[0485] 481. The sequence of the element in Table X contains the Vingi-1 EE retrotranspose, and the template RNA contains the 5'UTR and 3'UTR of Vingi-1EE, in any of the above numbered embodiments.
[0486] 482. The element arrangements in Table X represent any of the above-mentioned embodiments belonging to a restricted endonuclease-like (RLE) clade.
[0487] 483. The element sequences in Table X represent any of the above-mentioned embodiments belonging to the depurine-site endonuclease-like (APE) clade.
[0488] 484. The element sequence of Table X belongs to the Penelope-like element (PLE) clade, and optionally, the element sequence of Table X includes a GIY-YIG domain (e.g., a GIY-YIG endonuclease domain), as described in any of the above embodiments.
[0489] 485. The array of elements in Table X is any number of embodiments above that belong to a clade selected from the following clades: CRE, NeSL, R4, R2, Hero, L1, RTE (e.g., AviRTE_Brh or BovB), I, Jockey, CR1, Rex1, RandI / Dualen, Penelope (e.g., Penelope_SM), Tx1, RTEX, Crack, Nimb, Proto1, Proto2, RTETP, L2, Tad1, Loa, Ingi, Outcast, R1, Daphne, L2A, L2B, Ambal, Vingi (e.g., Vingi-1_EE), and Kiri.
[0490] 486. Any of the above embodiments, wherein the target DNA-binding domain is heterogeneous to one or more other domains of the polypeptide (e.g., the reverse transcriptase domain and / or the endonuclease domain).
[0491] 487. The heterologous target DNA binding domain comprises Cas9, Cas9 nickase, dCas9, zinc finger, or TAL domain, as described in any of the above embodiments.
[0492] 488. The heterologous target DNA binding domain comprises a Cas domain as shown in Table 9 or Table 37, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, as described in any of the above embodiments.
[0493] 489. The heterologous target DNA binding domain is any one of the above embodiments, including the N-terminal dCas9 domain.
[0494] 490. Any one of the above embodiments further comprises a guide RNA (e.g., U6-driven gRNA) containing at least 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides homologous to the target DNA sequence.
[0495] 491. Any one number of embodiments above, wherein the template RNA further comprises a guide RNA region containing at least 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides homologous to the target DNA sequence.
[0496] 492. The gRNA sequence is located at the 5' end of the template in any of the above numbered embodiments.
[0497] 493. The gRNA sequence is located at the 3' end of the template in any of the above numbered embodiments.
[0498] 494. The gRNA sequence is a scaffold capable of recruiting Cas9, as described in any of the above embodiments.
[0499] 495. The gRNA sequence is any of the above-mentioned embodiments, for example, including homology domains as described herein.
[0500] 496. Any one of the above embodiments, wherein the endonuclease domain is heterogeneous to one or more other domains of the polypeptide (e.g., the reverse transcriptase domain and / or the target DNA binding domain).
[0501] 497. The heterologous endonuclease domain is any one of the above-mentioned embodiments, comprising a Cas9, Cas9 nickase, or FokI domain.
[0502] 498. The polypeptide comprises any one of the above embodiments, including an RNase H domain.
[0503] 499. Any of the above embodiments wherein the polypeptide does not contain an RNase H domain or contains an inactivated RNase H domain.
[0504] 500. A nucleic acid encoding a polypeptide further comprises any number of the above embodiments, including a second open reading frame.
[0505] 501. The nucleic acid encoding the polypeptide includes, for example, a 2A sequence located between the ORF1 sequence and the ORF2 sequence, and optionally, the 2A sequence is selected from T2A(EGRGSLLTCGDVEENPGP), P2A(ATNFSLLKQAGDVEENPGP), E2A(QCTNYALLKLAGDVESNPGP), or F2A(VKQTLNFDLLKLAGDVESNPGP), in any of the above any number of embodiments.
[0506] 502. The polypeptide comprises any number of the above embodiments, wherein the polypeptide contains an intein.
[0507] 503. The system comprises any number of the above embodiments, including an intein (for example, contained in the second polypeptide).
[0508] 504. Any number of embodiments of the above, wherein the polypeptide is encoded by two or more separate open reading frames, each encoding a polypeptide fragment.
[0509] 505. An intein (e.g., a trans-splicing intein) is any one of the above embodiments, wherein two or more polypeptide fragments are joined together to form a polypeptide.
[0510] 506. Any number of embodiments described above, the system comprises (i) a first polypeptide fragment comprising at least one of a reverse transcriptase domain, an endonuclease domain, and a target DNA binding domain, and (ii) a second polypeptide fragment comprising at least one of a reverse transcriptase domain, an endonuclease domain, and a target DNA binding domain, wherein the first polypeptide fragment does not contain the same type of domain as the second polypeptide fragment.
[0511] 507.(a) The first polypeptide fragment comprises a reverse transcriptase domain, the second polypeptide fragment comprises an endonuclease domain, and optionally, the first polypeptide fragment further comprises a target DNA binding domain, or the second polypeptide fragment further comprises a target DNA binding domain; (b) The first polypeptide fragment comprises a reverse transcriptase domain, and the second polypeptide fragment comprises a target DNA binding domain, and optionally, the first polypeptide fragment further comprises an endonuclease domain, or the second polypeptide fragment further comprises an endonuclease domain; or (a) Any number of embodiments described above, wherein the first polypeptide fragment comprises an endonuclease domain, the second polypeptide fragment comprises a target DNA binding domain, and optionally, the first polypeptide fragment further comprises a reverse transcriptase domain, or the second polypeptide fragment further comprises a reverse transcriptase domain.
[0512] 508. Any number of embodiments above, wherein the intent is formed by binding a first polypeptide fragment to a second polypeptide.
[0513] 509. Intein is, (i) Fusion of the reverse transcriptase domain to the endonuclease domain, (ii) Fusion of the reverse transcriptase domain to the target DNA binding domain, or (iii) Fusion of the endonuclease domain to the target DNA binding domain An embodiment of any number above that induces the above.
[0514] 510. Any one of the above embodiments, wherein the intein is heterogeneous to one or more (e.g., one, two, or all) of the reverse transcriptase domain, endonuclease domain, and target DNA binding domain.
[0515] 511. The intent is split-intane, in any of the above numbered embodiments.
[0516] 512. The DNA encoding the polypeptide includes plasmids, minicircles, Doggybone DNA (dbDNA), or ceDNA, as per any of the above-mentioned embodiments.
[0517] 513. The RNA encoding the polypeptide is any of the above any number of embodiments, comprising: a cap region, a poly(A) tail, and / or one or more chemical modifications, e.g., one or more chemically modified nucleotides.
[0518] 514. Any of the above embodiments, wherein the RNA encoding the polypeptide includes circRNA.
[0519] 515. The nucleic acid encoding the polypeptide is contained within a virus (e.g., AAV, adenovirus, or lentivirus, e.g., an embedded-deficient lentivirus), as in any of the embodiments described above.
[0520] 516. The nucleic acid encoding the polypeptide is contained within nanoparticles (e.g., lipid nanoparticles), vesicles, or fusosomes, as in any of the embodiments described above.
[0521] 517. Any one of the above embodiments, wherein the polypeptide comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence encoded by the sequence of the elements of Table X.
[0522] 518. Any of the above-mentioned embodiments, wherein the reverse transcriptase domain comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the reverse transcriptase domain of an amino acid sequence encoded by the sequence of an element in Table 3B, Table 10, Table 11, or Table X.
[0523] 519. Any of the above-mentioned embodiments, wherein the retrotransposase comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequence encoded by the sequence of an element in Table 3B, Table 10, Table 11, or Table X.
[0524] 520. The polypeptide comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequences of the elements in Table 3B, Table 10, Table 11, or Table X, as described in any of the above embodiments.
[0525] 521. The any numbered embodiment above, wherein the reverse transcriptase domain comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the reverse transcriptase domain of the amino acid sequence of the element in Table 3B, Table 10, Table 11, or Table X.
[0526] 522. The retrotransposase comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to the amino acid sequences of the elements in Table 3B, Table 10, Table 11, or Table X, as described in any of the above embodiments.
[0527] 523. The polypeptide comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 1023) or GGGS (SEQ ID NO: 1024), any of the above-mentioned embodiments.
[0528] 524. The reverse transcriptase domain comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 1023) or GGGS (SEQ ID NO: 1024), as described in any of the above-mentioned embodiments.
[0529] 525. The retrotransposase comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 1023) or GGGS (SEQ ID NO: 1024), as described in any of the above-mentioned embodiments.
[0530] 526. The polypeptide, reverse transcriptase domain, or retrotransposase comprises a linker having an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 1023) or GGGS (SEQ ID NO: 1024), as described in any of the above embodiments.
[0531] 527. Any one of the above embodiments, wherein the polypeptide comprises a DNA-binding domain covalently bonded to the remainder of the polypeptide by a linker, for example, a linker comprising at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 125, 150, 200, 300, 400, or 500 amino acids.
[0532] 528. Any one of the above embodiments, wherein the linker is attached to the remainder of the polypeptide at a position in the DNA-binding domain, RNA-binding domain, reverse transcriptase domain, or endonuclease domain (for example, as shown in any of Figures 17A–17F of PCT Application No. PCT / US2019 / 048607).
[0533] 529. The linker is bonded to the remainder of the polypeptide at the N-terminal position of the alpha-helix region of the polypeptide, for example, at the position corresponding to version v1 as described in Example 26 of PCT Application No. PCT / US2019 / 048607, in any of the above numbered embodiments.
[0534] 530. The linker is attached to the remainder of the polypeptide, for example, at a C-terminal position of the alpha-helix region of the polypeptide, at a position corresponding to version v2, as described in Example 26 of PCT Application No. PCT / US2019 / 048607, for example, preceding the RNA-binding motif (e.g., -1 RNA-binding motif), in any of the above-mentioned embodiments.
[0535] 531. The linker is attached to the remainder of the polypeptide at a position corresponding to version v3, for example, at the C-terminal side of the random coil region of the polypeptide, for example, at the N-terminal side of the DNA binding motif (e.g., the c-myb DNA binding motif), as described in Example 26 of PCT Application No. PCT / US2019 / 048607, in any of the above-mentioned embodiments.
[0536] 532. The linker comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 1023) or GGGS (SEQ ID NO: 1024), as described in any of the above-mentioned embodiments.
[0537] 533. Any one of the above embodiments, wherein a polynucleotide sequence containing at least approximately 500, 1000, 2000, 3000, 3500, 3600, 3700, 3800, 3900, or 4000 consecutive nucleotides from the 5' end of a template RNA sequence is incorporated into the target cell genome.
[0538] 534. Any of the above-mentioned embodiments, wherein a polynucleotide sequence containing at least approximately 500, 1000, 2000, 2500, 2600, 2700, 2800, 2900, or 3000 consecutive nucleotides from the 3' end of a template RNA sequence is incorporated into the target cell genome.
[0539] 535. Any of the above-mentioned embodiments, wherein the nucleic acid sequence of a template RNA or a portion thereof (for example, a portion containing at least about 100, 200, 300, 400, 500, 1000, 2000, 2500, 3000, 3500, or 4000 nucleotides) is incorporated into the genome of a population of target cells at a copy number of at least about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1.0 integrations / genome.
[0540] 536. Any of the above-mentioned embodiments, wherein the nucleic acid sequence of a template RNA or a portion thereof (for example, a portion containing at least about 100, 200, 300, 400, 500, 1000, 2000, 2500, 3000, 3500, or 4000 nucleotides) is incorporated into the genome of a population of target cells at a copy number of at least about 0.01, 0.02, 0.03, 0.04, 0.05, 0.75, or 0.1 integrations / genome.
[0541] 537. The polypeptide comprises a functional endonuclease domain (wherein the endonuclease domain, for example, does not contain mutations that cause loss of endonuclease activity, such as those described herein), as any of the above-mentioned embodiments.
[0542] 538. Any of the above embodiments in which the introduction of the system into target cells does not result in changes in p53 and / or p21 protein levels (e.g., upregulation), H2AX phosphorylation (e.g., gamma H2AX), ATM phosphorylation, ATR phosphorylation, Chk1 phosphorylation, Chk2 phosphorylation, and / or p53 phosphorylation.
[0543] 539. Any one of the above embodiments, wherein the introduction of the system into target cells upregulates the p53 protein level in the target cells to a level less than approximately 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the p53 protein level induced by the introduction of a site-specific nuclease targeting the same genomic site as the system, such as Cas9.
[0544] 540.p53 protein levels are determined by the method described in Example 30, in any of the above numbered embodiments.
[0545] 541. Any one of the above embodiments, wherein the introduction of the system into target cells upregulates the p53 phosphorylation level in the target cells to a level less than approximately 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the p53 phosphorylation level induced by the introduction of a site-specific nuclease targeting the same genomic site as the system, such as Cas9.
[0546] 542. Any one of the above embodiments, wherein the introduction of the system into target cells upregulates the p21 protein level in the target cells to a level less than approximately 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the p53 protein level induced by the introduction of a site-specific nuclease targeting the same genomic site as the system, such as Cas9.
[0547] 543. The p21 protein level is determined by the method described in Example 30, in any of the above numbered embodiments.
[0548] 544. Any one of the above embodiments, wherein the introduction of the system into target cells upregulates the H2AX phosphorylation level in the target cells to a level less than approximately 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the H2AX phosphorylation level induced by the introduction of a site-specific nuclease targeting the same genomic site as the system, such as Cas9.
[0549] 545. Any number of the above embodiments, wherein the introduction of the system into target cells upregulates the ATM phosphorylation level in the target cells to a level less than approximately 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the ATM phosphorylation level induced by the introduction of a site-specific nuclease targeting the same genomic site as the system, such as Cas9.
[0550] 546. Any one of the above embodiments, wherein the introduction of the system into target cells upregulates the ATR phosphorylation level in the target cells to a level less than approximately 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the ATR phosphorylation level induced by the introduction of a site-specific nuclease targeting the same genomic site as the system, such as Cas9.
[0551] 547. Any one of the above embodiments, wherein the introduction of the system into target cells upregulates the Chk1 phosphorylation level in the target cells to a level less than approximately 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the Chk1 phosphorylation level induced by the introduction of a site-specific nuclease targeting the same genomic site as the system, such as Cas9.
[0552] 548. Any one of the above embodiments, wherein the introduction of the system into target cells upregulates the Chk2 phosphorylation level in the target cells to a level less than approximately 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, or 90% of the Chk2 phosphorylation level induced by the introduction of a site-specific nuclease targeting the same genomic site as the system, such as Cas9.
[0553] 549. The target DNA binding domain recognizes a specific target DNA sequence, as in any of the above any number of embodiments.
[0554] 550. Any number of embodiments described above, wherein the target DNA binding domain binds to multiple (e.g., random) target DNA sequences.
[0555] 551. The template RNA is any one of the above embodiments, including a guide RNA (e.g., U6-driven gRNA).
[0556] 552. The target DNA-binding domain comprises one or more DNA-binding domains from among the retrotransposases described herein (e.g., retrotransposases of elements X, 10, 11, 3A, or 3B), Cas9, nickase Cas9, dCas9, zinc finger, TAL, meganuclease, and / or transcription factors, as described herein, in any of the above-numbered embodiments.
[0557] 553. The reverse transcriptase domain comprises the reverse transcriptase domain of a retrotransposase described herein (e.g., a retrotransposase of the elements of Table X, 10, 11, Z1, Z2, 3A, or 3B) in any of the above-mentioned embodiments.
[0558] 554. The endonuclease domain is any one of the embodiments described herein, including the endonuclease domain of a retrotransposase described herein (e.g., retrotransposase of elements of Table X, 10, 11, 3A or 3B), Cas9, nickase Cas9, type II restriction enzyme (e.g., FokI), Holliday junction resolverase, RLE endonuclease domain, APE endonuclease domain, or GIY-YIG endonuclease domain.
[0559] 555. A polypeptide or nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA-binding domain (DBD), and (iii) an endonuclease domain, and the DBD and / or endonuclease domain comprises a heterologous targeting domain that specifically binds to a sequence contained in a target DNA molecule (e.g., genomic DNA).
[0560] 556. A template RNA (or DNA encoding template RNA) containing a targeting domain (e.g., heterologous targeting domain) that specifically binds to a sequence in a target DNA molecule (e.g., genomic DNA), a sequence that specifically binds to the RT domain of a polypeptide, and a heterologous target sequence.
[0561] 557. A polypeptide comprising a heterologous targeting domain that specifically binds to a sequence contained in a target DNA molecule (e.g., genomic DNA), a system, method, or template RNA of any of the preceding embodiments.
[0562] 558. A heterologous targeting domain is a system, method, or template RNA of any of the preceding embodiments that binds to a nucleic acid sequence different from that of an unmodified polypeptide.
[0563] 559. A polypeptide that does not contain a functional endogenous targeting domain (for example, the polypeptide does not contain an endogenous targeting domain), a system, method, or template RNA of any of the preceding embodiments.
[0564] 560. A heterologous targeting domain comprising a zinc finger (for example, a zinc finger that specifically binds to a sequence contained in a target DNA molecule) in any of the systems, methods, or template RNAs of the preceding embodiments.
[0565] 561. A heterologous targeting domain comprising a Cas domain (e.g., a Cas9 domain, or a variant or variant thereof, e.g., a Cas9 domain that specifically binds to a sequence contained in a target DNA molecule) in any of the systems, methods, or template RNAs of the preceding embodiments.
[0566] 562. A Cas domain associated with a guide RNA (gRNA), a system, method, or template RNA of any of the preceding embodiments.
[0567] 563. A heterologous targeting domain comprising an endonuclease domain (e.g., a heterologous endonuclease domain) in any of the systems, methods, or template RNAs of the preceding embodiments.
[0568] 564. The endonuclease domain comprises a Cas domain (e.g., calcium 9, or a variant or variant thereof) in any of the systems, methods, or template RNAs of the preceding embodiments.
[0569] 565. The Cas domain is associated with a guide RNA (gRNA) in any of the systems, methods, or template RNAs of the preceding embodiments.
[0570] 566. A system, method, or template RNA of any of the preceding embodiments, wherein the endonuclease domain includes the Fok1 domain.
[0571] 567. A system, method, or template RNA of any of the preceding embodiments, wherein the template nucleic acid molecule includes at least one (e.g., one or two) heterologous sequences having at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% homology to a sequence contained in a target DNA molecule (e.g., genomic DNA).
[0572] 568. A system, method, or template RNA of any of the preceding embodiments, wherein at least one heterologous sequence is located in or within approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides at the 5' end of the template nucleic acid molecule.
[0573] 569. A system, method, or template RNA of any of the preceding embodiments, wherein at least one heterologous sequence is located in or within approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides at the 3' end of the template nucleic acid molecule.
[0574] 570. A heterologous sequence is bound to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides of a nicking site in a target DNA molecule (e.g., generated by a nickase, e.g., an endonuclease domain as described herein) by any of the systems, methods, or template RNAs of the preceding embodiments.
[0575] 571. A system, method, or template RNA of any of the preceding embodiments, wherein the heterologous sequence has 50%, 40%, 30%, 20%, 10%, 5%, 4%, 3%, 2%, or less than 1% sequence identity with a nucleic acid sequence complementary to the endogenous homologous sequence of the unmodified form of the template RNA.
[0576] 572. A system, method, or template RNA of any of the preceding embodiments, wherein the heterologous sequence has at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% homology to a sequence of a target DNA molecule that is different from the sequence to which the endogenous homology sequence is joined (e.g., replaced by the heterologous sequence).
[0577] 573. A system, method, or template RNA of any of the preceding embodiments, comprising a sequence (e.g., at its 3' end) having at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% homology to a sequence located at 5' of a target DNA molecule (e.g., a site nicked by a nickase, e.g., an endonuclease domain as described herein), wherein the heterologous homologous sequence includes a sequence (e.g., at its 3' end).
[0578] 574. A heterologous sequence comprising (for example, at its 5' end) a sequence suitable for priming target primed reverse transcription (TPRT) initiation, a system, method, or template RNA of any of the preceding embodiments.
[0579] 575. A heterologous sequence having at least 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% homology to a sequence located within approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides (e.g., 3' side) of a target insertion site in a target DNA molecule for a heterologous target sequence (e.g., as described herein), in any of the systems, methods, or template RNAs of the preceding embodiments.
[0580] 576. The template nucleic acid molecule is, for example, a system, method, or template RNA of any of the preceding embodiments, including a guide RNA (gRNA) as described herein.
[0581] 577. A system, method, or template RNA of any of the above embodiments, wherein the template nucleic acid molecule comprises a gRNA spacer sequence (for example, in or within the 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides at its 5' end).
[0582] 578. A template RNA (or DNA encoding template RNA) comprising (i) a sequence that binds to a target site (e.g., the second strand of the site in the target genome), (ii) a sequence that specifically binds to the RT domain of the polypeptide, (iii) a heterologous target sequence, and (iv) a 3' target homology domain (for example, from 5' to 3').
[0583] 579.(v) A template RNA of any of the previous embodiments, further comprising a sequence that binds to the endonuclease and / or DNA-binding domain of a polypeptide (e.g., the same polypeptide including the RT domain).
[0584] The 580.RT domain is a template RNA of any of the previous embodiments, comprising a sequence selected from Table 3B, 10, 11, or X, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0585] 581. The RT domain comprises a sequence selected from Table 3B, 10, 11, or X, and the RT domain further comprises several substitutions relative to the native sequence, e.g., at least 1, 2, 3, 4, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 substitutions, as a template RNA of any of the previous embodiments.
[0586] The sequence of 582.(ii) is a template RNA from any of the previous embodiments that specifically binds to the RT domain.
[0587] 583. The template RNA of any of the previous embodiments, wherein the sequence that specifically binds to the RT domain is a sequence from Table 3B or 10, for example, a UTR sequence, or a sequence having at least 70, 75, 80, 85, 90, 95, or 99% identity thereto.
[0588] 584. From 5' to 3', a template RNA (or DNA encoding the template RNA) comprising: (ii) a sequence that binds to the endonuclease and / or DNA-binding domain of the polypeptide; (i) a sequence that binds to a target site (e.g., the second strand of the site in the target genome); (iii) a heterologous target sequence; and (iv) a 3' homology domain.
[0589] A template RNA (or DNA encoding the template RNA) comprising, from 585.5' to 3', (iii) a heterologous target sequence, (iv) a 3' homology domain, (i) a sequence that binds to a target site (e.g., the second strand of the site in the target genome), and (ii) a sequence that binds to the endonuclease and / or DNA-binding domain of the polypeptide.
[0590] 586. The RNA of the system (e.g., template RNA, RNA encoding a polypeptide of (a), or RNA expressed from a heterologous target sequence incorporated into target DNA) is, for example, a system, method, kit, template RNA, or reaction mixture containing a microRNA binding site in the 3'UTR.
[0591] 587. A system, method, kit, template RNA, or reaction mixture of any of the preceding embodiments, wherein the microRNA binding site is recognized by a miRNA that is present in non-target cell types but not in target cell types (or present in lower levels compared to non-target cells).
[0592] 588. A system, method, kit, template RNA, or reaction mixture of any of the preceding embodiments, wherein the miRNA is miR-142 and / or the non-target cell is a Kupffer cell or a blood cell, such as an immune cell.
[0593] 144. A system, method, kit, template RNA, or reaction mixture of any of the preceding embodiments, wherein the miRNA is miR-182 or miR-183, and / or the non-target cell is a dorsal root ganglion neuron.
[0594] 588. A system, method, kit, template RNA, or reaction mixture of any of the preceding embodiments, comprising a first miRNA binding site recognized by a first miRNA (e.g., miR-142), and further comprising a second miRNA binding site recognized by a second miRNA (e.g., miR-182 or miR-183), wherein the first and second miRNA binding sites are located on the same RNA or on different RNAs in the system.
[0595] 589. A system, method, kit, template RNA, or reaction mixture of any of the preceding embodiments, wherein the template RNA comprises at least two, three, or four miRNA binding sites, for example, the miRNA binding sites are recognized by the same or different miRNAs.
[0596] The RNA encoding the polypeptide of 590.(a) comprises at least two, three, or four miRNA binding sites, for example, the miRNA binding sites are recognized by the same or different miRNAs, in any of the systems, methods, kits, template RNA, or reaction mixtures of the preceding embodiments.
[0597] 591. A system, method, kit, template RNA, or reaction mixture of any of the preceding embodiments, wherein the RNA expressed from a heterologous target sequence incorporated into target DNA contains at least two, three, or four miRNA binding sites, for example, the miRNA binding sites are recognized by the same or different miRNAs.
[0598] definition Domain: As used herein, the term "domain" refers to a structure of a biomolecule that contributes to a specific function of that biomolecule. A domain may include a continuous region (e.g., a continuous sequence) or a distinct discontinuous region (e.g., a discontinuous sequence) of a biomolecule. Examples of protein domains include, but are not limited to, endonuclease domains, DNA-binding domains, and reverse transcription domains; examples of nucleic acid domains include regulatory domains, such as transcription factor-binding domains.
[0599] Extrinsic: As used herein, the term “extrinsic,” when used in relation to biomolecules (e.g., nucleic acid sequences or polypeptides), means that the biomolecule has been artificially introduced into a host genome, cell, or organism. For example, nucleic acids added to an existing genome, cell, tissue, or subject using recombinant DNA technology or other methods are extrinsic to the existing nucleic acid sequence, cell, tissue, or subject.
[0600] Genome-safe harbor sites (GSH sites): Genome-safe harbor sites are locations within the host genome that can accommodate the integration of new genetic material, for example, in such a way that the inserted genetic element does not cause significant alteration of the host genome that poses a risk to the host cell or organism. GSH sites generally meet criteria 1, 2, 3, 4, 5, 6, 7, 8, or 9 below: (i) located >300kb from oncological genes; (ii) located >300kb from miRNA / other functional small RNAs; (iii) located >50kb from the 5' gene end; (iv) located >50kb from the origin of replication; (v) located >50kb away from superconserved elements; (vi) having low transcriptional activity (i.e., lacking mRNA+ / -25kb); (vii) not being in a copy number variable region; (viii) being in open chromatin; and / or (ix) having one copy in the human genome and being unique. Examples of GSH sites in the human genome that meet some or all of these criteria include: (i) adenovirus site 1 (AAVS1), a naturally occurring integration site of the AAV virus on chromosome 19; (ii) chemokine (CC motif) receptor 5 (CCR5) gene, a chemokine receptor gene known as the HIV-1 coreceptor; (iii) the human ortholog of the mouse Rosa26 locus; and (iv) rDNA loci. Further GSH sites are known and are described, for example, in Pellenz et al. epub August 20, 2018 (https: / / doi.org / 10.1101 / 396390).
[0601] Heterogeneous: When the term “heterogeneous” is used to describe a first element in relation to a second element, it means that the first and second elements do not naturally exist in the configuration described. For example, heterogeneous polypeptides, nucleic acid molecules, constructs, or sequences refer to (a) a polypeptide, nucleic acid molecule, or portion of a polypeptide or nucleic acid molecule sequence that is not native to the cell in which it is expressed, (b) a polypeptide or nucleic acid molecule or portion of a polypeptide or nucleic acid molecule that has been modified or mutated from its native state, or (c) a polypeptide or nucleic acid molecule having modified expression compared to its native expression level under similar conditions. For example, heterogeneous regulatory sequences (e.g., promoters, enhancers) can be used to regulate the expression of a gene or nucleic acid molecule in a manner different from that which is normally expressed in nature. In another example, a heterogeneous domain of a polypeptide, or a nucleic acid sequence (e.g., a DNA-binding domain of a polypeptide, or the nucleic acid encoding the DNA-binding domain of a polypeptide), may be configured in relation to other domains, or may be a different sequence or of a different source compared to other domains or portions of the polypeptide, or the coding nucleic acid thereof. In certain embodiments, heterologous nucleic acid molecules may be present in the native host cell genome, but may have modified expression levels, different sequences, or both. In other embodiments, heterologous nucleic acid molecules may not be endogenous in the host cell or host genome, but may instead be introduced into the host cell by transformation (e.g., transfection, electroporation), where the added molecule may be integrated into the host genome or exist as extrachromosomal genetic material, either transiently (e.g., mRNA) or semi-stable for two or more generations (e.g., episomal viral vector, plasmid, or other self-replicating vector). In some embodiments, a domain is heterologous to another domain if the first domain is not naturally present in the same polypeptide as the other domain (e.g., fusion between two domains of different proteins from the same organism).
[0602] Mutation or mutant: When the term “(sudden) mutant” is applied to a nucleic acid sequence, it means that nucleotides within the nucleic acid sequence may be inserted, deleted, or altered compared to a reference (e.g., natural) nucleic acid sequence. A single alteration may occur at one locus (point mutation), or multiple nucleotides may be inserted, deleted, or altered at a single locus. In addition, one or more alterations may occur at any number of loci within a nucleic acid sequence. A nucleic acid sequence can be mutated by any method known in the art. In some embodiments, the mutation occurs naturally. In some embodiments, the desired mutation can be generated by a system described herein.
[0603] Nucleic acid molecules: Nucleic acid molecules refer to, but are not limited to, both RNA and DNA molecules, including cDNA, genomic DNA and mRNA, and also include synthetic nucleic acid molecules, such as those chemically synthesized or produced by recombination, as described herein, such as RNA templates. Nucleic acid molecules may be double-stranded or single-stranded, cyclic or linear. If single-stranded, the nucleic acid molecule may be a sense strand or an antisense strand. Unless otherwise stated, and as an example of all sequences described herein in the general format “Sequence ID,” a nucleic acid including “Sequence ID 1” means a nucleic acid in which at least a portion has either (i) the sequence of Sequence ID 1, or (ii) a sequence complementary to Sequence ID 1. The choice between the two depends on the context in which Sequence ID 1 is used. For example, when a nucleic acid is used as a probe, the choice between the two depends on the requirement that the probe is complementary to the desired target. The nucleic acid sequences of this disclosure may be chemically or biochemically modified, or may contain unnatural or derivatized nucleotide bases, as will be readily apparent to those skilled in the art. Examples of such modifications include labeling, methylation, substitution of one or more spontaneously occurring nucleotides by analogs, internucleotide modifications such as uncharged bonds (e.g., methylphosphonic acid, triester phosphate, phosphoramidates, carbamates, etc.), charged bonds (e.g., phosphorothioates, phosphorodithioates, etc.), pendant moieties (e.g., polypeptides), insertants (e.g., acridine, psoralens, etc.), chelating agents, alkylating agents, and modifying bonds (e.g., α-anomeric nucleic acids, etc.). Synthetic molecules that mimic polynucleotides in their ability to bind to a specified sequence through hydrogen bonding and other chemical interactions are also included. Such molecules are known in the art and, for example, use peptide bonds instead of phosphate bonds in the molecular backbone. Other modifications include analogs that include other structures, such as modifications found in bridging moieties or "locked" nucleic acids, where the ribose ring is located.
[0604] Gene expression unit: A gene expression unit is a nucleic acid sequence comprising at least one regulatory nucleic acid sequence operably ligated to at least one effector sequence. The first nucleic acid sequence is operably ligated to the second nucleic acid sequence when the first nucleic acid sequence is positioned functionally in relation to the second nucleic acid sequence. For example, a promoter or enhancer is operably ligated to a coding sequence if the promoter or enhancer affects the transcription or expression of the coding sequence. The operably ligated DNA sequences may be continuous or discontinuous. If it is necessary to ligate two protein coding regions, the operably ligated sequences may reside within the same reading frame.
[0605] Host: The terms host genome or host cell, as used herein, refer to the cell into which the protein and / or genetic material has been introduced and / or its genome. These terms refer not only to a specific target cell and / or genome, but also to the offspring of such cells and / or the genomes of such offspring. It should be understood that such offspring may not be identical to the parent cell in fact, as certain modifications may occur in later generations due to mutation or environmental influences, but are still included in the scope of the term “host cell” as used herein. A host genome or host cell may be an isolated cell or a cell line grown in culture or genomic material isolated from such cells or cell lines, or it may be a host cell or host genome constituting a living tissue or organism. In some cases, the host cell may be an animal cell or a plant cell, as described herein, for example. In certain cases, the host cell may be a bovine cell, a horse cell, a pig cell, a goat cell, a sheep cell, a chicken cell, or a turkey cell. In certain cases, the host cell may be a maize cell, a soybean cell, a wheat cell, or a rice cell.
[0606] Pseudoknot: As used herein, “pseudoknot sequence” refers to a nucleic acid (e.g., RNA) having, for example, a first segment, a second segment between the first and third segments (the third segment being complementary to the first segment), and a fourth segment (the fourth segment being complementary to the second segment), the self-complementarity of which is suitable for forming a pseudoknot structure. The pseudoknot may optionally have additional secondary structures, such as a stem-loop located within the second segment, a stem-loop located between the second and third segments, a sequence preceding the first segment, or a sequence following the fourth segment. The pseudoknot may have additional sequences between the first and second segments, between the second and third segments, or between the third and fourth segments. In some embodiments, the segments are arranged 5' to 3' as: first, second, third, and fourth. In some embodiments, the first and third segments contain five fully complementary base pairs. In some embodiments, the second and fourth segments optionally contain ten base pairs having one or more (e.g., two) bulges. In some embodiments, the second segment contains one or more unpaired nucleotides, for example, forming a loop. In some embodiments, the third segment contains one or more unpaired nucleotides, for example, forming a loop.
[0607] Stem-loop sequence: As used herein, “stem-loop sequence” refers to a nucleic acid sequence (e.g., RNA sequence) having a stem containing, for example, at least 2 (e.g., 3, 4, 5, 6, 7, 8, 9, or 10) base pairs and a loop containing at least 3 (e.g., 4) base pairs, where self-complementarity is sufficient to form a stem-loop. The stem may contain mismatches or bulges.
[0608] A patent file or application file must include at least one color-illustrated drawing. A copy of a patent publication or patent application publication containing a color-illustrated drawing will be provided by the Office immediately upon request and payment of the required fees. In embodiments of the present invention, for example, the following items are provided. (Item 1) A system for modifying DNA, (a) polypeptides or nucleic acids encoding polypeptides (e.g., DNA or mRNA), wherein the polypeptide comprises (i) a reverse transcriptase domain and (ii) an endonuclease domain, and one or both of (i) or (ii) have an amino acid sequence encoded by the nucleic acid sequence of the elements in Table 3B, Table 10, or Table 11, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (b) A template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds to the polypeptide and (ii) a heterogeneous target sequence. A system that includes this. (Item 2) A system for modifying DNA, (a) polypeptides or nucleic acids encoding polypeptides (e.g., DNA or mRNA), wherein the polypeptide comprises (i) a reverse transcriptase domain and (ii) an endonuclease domain, and either or both of (i) or (ii) have an amino acid sequence encoded by the nucleic acid sequence of the elements in Table 3B, Table 10, or Table 11, or a sequence that differs from it by 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides; and (b) A template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds to the polypeptide and (ii) a heterogeneous target sequence. A system that includes this. (Item 3) A system for modifying DNA, (a) polypeptides or nucleic acids encoding polypeptides (e.g., DNA or mRNA), wherein the polypeptide comprises (i) a reverse transcriptase domain and (ii) a target DNA binding domain, and either or both of (i) or (ii) have an amino acid sequence encoded by the nucleic acid sequence of the elements in Table 3B, Table 10, or Table 11, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (b) A template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds to the polypeptide and (ii) a heterogeneous target sequence. A system that includes this. (Item 4) A system for modifying DNA, (a) a polypeptide or nucleic acid encoding the polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase (RT) domain, (ii) a DNA-binding domain (DBD), and (iii) an endonuclease domain, such as a niccasse domain; and (b) (e.g., from 5' to 3') a template RNA (or DNA encoding the template RNA) comprising (i) optionally a sequence that binds to a target site (e.g., an unedited strand of a site in the target genome), (ii) optionally a sequence that binds to the polypeptide, (iii) a heterologous target sequence, and (iv) a 3' homologous domain. Includes, A system in which the RT domain has a sequence of Table 3B, Table 10, or Table 11, or a sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto. [Brief explanation of the drawing]
[0609] [Figure 1] This is a schematic diagram of the Gene Writing (trademark) genome editing system. [Figure 2] This is a schematic diagram of the structure of the Gene Writer (trademark) genome editor polypeptide. [Figure 3] This is a schematic diagram of the structure of an example GeneWriter™ template RNA. [Figure 4A]This is a series of diagrams showing examples of the stereochemistry of Gene Writers using domains derived from various sources. Gene Writers described herein may or may not include all of the domains shown. For example, GeneWrite may optionally lack an RNA-binding domain, or it may have a single domain that satisfies the functions of multiple domains, such as a Cas9 domain for DNA binding and endonuclease activity. Exemplary domains that may be included in a Gene Writer polypeptide include DNA-binding domains (e.g., DNA-binding domains listed in any of Tables X, Y, Z1, Z2, 3A, or 3B; zinc fingers; TAL domains; Cas9; dCas9; nickase Cas9; transcription factors, or meganucleases), RNA-binding domains (e.g., RNA-binding domains of B-box proteins, MS2 coat proteins, dCas, or elements of sequences listed in any of Tables X, Y, Z1, Z2, 3A, or 3B), reverse transcriptase domains (e.g., reverse transcriptase domains of elements of sequences in the tables herein; other Examples include retrotransposases (e.g., those listed in Table Z1); peptides containing reverse transcriptase domains (e.g., those listed in Table Z2); and / or endonuclease domains (e.g., endonuclease domains of elements listed in Tables X, Y, Z1, Z2, 3A, or 3B; Cas9; nickase Cas9; restriction enzymes (e.g., type II restriction enzymes, e.g., FokI); meganucleases; Holliday junction resolvers; RLE retrotranspases; APE retrotransposases; or GIY-YIG retrotransposases). Exemplary Gene Writer polypeptides containing exemplary combinations of such domains are shown in the bottom panel. [Figure 4B]This is a series of diagrams showing examples of the stereochemistry of Gene Writers using domains derived from various sources. Gene Writers described herein may or may not include all of the domains shown. For example, GeneWrite may optionally lack an RNA-binding domain, or it may have a single domain that satisfies the functions of multiple domains, such as a Cas9 domain for DNA binding and endonuclease activity. Exemplary domains that may be included in a Gene Writer polypeptide include DNA-binding domains (e.g., DNA-binding domains listed in any of Tables X, Y, Z1, Z2, 3A, or 3B; zinc fingers; TAL domains; Cas9; dCas9; nickase Cas9; transcription factors, or meganucleases), RNA-binding domains (e.g., RNA-binding domains of B-box proteins, MS2 coat proteins, dCas, or elements of sequences listed in any of Tables X, Y, Z1, Z2, 3A, or 3B), reverse transcriptase domains (e.g., reverse transcriptase domains of elements of sequences in the tables herein; other Examples include retrotransposases (e.g., those listed in Table Z1); peptides containing reverse transcriptase domains (e.g., those listed in Table Z2); and / or endonuclease domains (e.g., endonuclease domains of elements listed in Tables X, Y, Z1, Z2, 3A, or 3B; Cas9; nickase Cas9; restriction enzymes (e.g., type II restriction enzymes, e.g., FokI); meganucleases; Holliday junction resolvers; RLE retrotranspases; APE retrotransposases; or GIY-YIG retrotransposases). Exemplary Gene Writer polypeptides containing exemplary combinations of such domains are shown in the bottom panel. [Figure 5]This is a diagram showing the modules of an exemplary Gene Writer RNA template. Individual modules of the exemplary template can be combined, rearranged, and / or omitted to generate, for example, a Gene Writer template. A = 5' homologous arm; B = ribozyme; C = 5'UTR; D = heterologous target sequence; E = 3'UTR; F = 3' homologous arm. [Figure 6] This is a table showing a list of modules for an exemplary Gene Writer RNA template. Individual modules can be combined, rearranged, and / or omitted to generate, for example, a Gene Writer template. A=5' homologous arm; B=ribozyme; C=5'UTR; D=heterogeneous target sequence; E=3'UTR; F=3' homologous arm. [Figure 7] This is a diagram illustrating an exemplary second strand nicking process. (A) Cas9 nickase fuses with the Gene Writer protein. The Gene Writer protein introduces a nick into the DNA strand via its EN domain (indicated as *), and the fused Cas9 nickase introduces a nick on the upper or lower DNA strand (indicated as X). This is a diagram illustrating an exemplary second strand nicking process. (B) The Gene Writer targets DNA via its DNA-binding domain and introduces a DNA nick via its EN domain (*). Subsequently, Cas9 nickase is used to generate a second nick (X) on the upper or lower strand, upstream or downstream of the EN-introduced nick. [Figure 8]This paper presents a screening of construct designs for retrotransposon-mediated integration in human cells. A driver plasmid containing a retrotransposase (driver) expression cassette is transfused together with a template plasmid containing a retrotransposon-dependent reporter cassette. Expression from the template plasmid results in a non-functional GFP due to the disruption of antisense introns, whereas transcription of the template molecule from the template plasmid results in RNA generation, which can be spliced to remove introns, then reverse-transcribed and integrated by the system. Therefore, reporter cassette expression arises only from the integrated reporter cassette (integrated gDNA, bottom), and not from the template plasmid. HA = homologous arm (where applicable); CMV = mammalian CMV promoter; HiBit = HiBit tag for protein expression quantification; T7 = T7 RNA polymerase promoter; UTR = untranslated sequence, e.g., native retrotransposon UTR; pA = poly(A) signal; SD-SA is used to indicate the splice donor and splice acceptor sites of antisense introns within the GFP coding sequence. HA = Homology arm (if applicable) (see, e.g., Example 6); CMV = Mammalian CMV promoter; HiBit = HiBit tag for quantification of protein expression; T7 = T7 RNA polymerase promoter; UTR = Untranslated sequence, e.g., native retrotransposon UTR; pA = PolyA signal; SD-SA is used to indicate splice donor and splice acceptor sites of antisense introns in the GFP coding sequence. [Figure 9] The candidate retrotransposons identified 25 candidates that incorporate the transpayload in human cells. A total of 163 retrotransposon systems were assayed for activity in human cells, as described in Example 4. Integration, measured by ddPCR, is shown as copies / genomes for each retrotransposon driver / template system. The height of each bar represents the mean of two replicates. Constructs with higher activity are further highlighted in Figure 10 after further optimization in Examples 4 and 5. [Figure 10]Based on retrotransposon hits from Table 3B, further improved high-activity Gene Writing configurations are shown in Examples 7, 8, and 9. Only the highest-performing configuration is shown when multiple configurations of a given system are tested, such as the addition of an alternative coding sequence for the retrotransposase (Example 8) or a homology arm (Example 9). For systems improved beyond the initial configuration evaluated in Example 7, as described in Table 3B, the improvements described in Examples 5 (Figure 11) and 6 (Figure 12) are detailed in Table 11. [Figure 11] Retrotransposon consensus sequences can rescue or improve trans-integration activity in human cells. Integration efficiency, measured by copy / genome by ddPCR as described in Example 5, is shown for each retrotransposon driver / template system. The height of each bar represents the average of two replicates. White circles represent each replicate, and bars represent the original sequence (light gray) or a novel consensus-generating protein sequence (dark gray). [Figure 12] Figure 12 shows that consensus motif-generating homology arm sequences can rescue or improve the trans-integration activity of retrotransposons in human cells. Integration efficiency measured by copy / genome by ddPCR (Example 6) is shown for each Gene Writer driver / template system. The height of each bar represents the average of two replicates. Black circles and white bars represent template sequences without homology arms, while white circles and hashed bars represent designs with homology arms. [Figure 13A] This section describes a luciferase activity assay for primary cells. LNPs formulated according to Example 11 were analyzed for cargo delivery to primary human hepatocytes according to Example 12. The luciferase assay revealed dose-responsive luciferase activity from the cell lysates. This indicates successful RNA delivery from mRNA cargo to cells and successful expression of firefly luciferase. [Figure 13B]The luciferase activity assay for primary cells is shown. LNPs formulated according to Example 11 were analyzed for cargo delivery to mouse hepatocytes according to Example 12. The luciferase assay revealed dose-responsive luciferase activity from the cell lysates. This indicates successful RNA delivery from mRNA cargo to cells and expression of firefly luciferase. [Figure 14] This paper discloses LNP-mediated delivery of RNA cargo to mouse liver. Firefly luciferase mRNA-containing LNPs were formulated and delivered to mice via IV. Liver samples were harvested and assayed for luciferase activity at 6, 24, and 48 hours post-administration. Reporter activity across various formulations followed the ranking LIPIDV005 > LIPIDV004 > LIPIDV003. RNA expression was transient, and enzyme levels returned to near vehicle background levels by 48 hours post-administration. [Modes for carrying out the invention]
[0610] This disclosure relates to compositions, systems, and methods for targeting, editing, modifying, or manipulating DNA sequences at one or more locations within a DNA sequence in a cell, tissue, or subject, for example, in vivo or in vitro (e.g., inserting a heterologous target DNA sequence into a target site in a mammalian genome). The target DNA sequence may include, for example, coding sequences, regulatory sequences, or gene expression units.
[0611] More specifically, this disclosure provides a retrotransposon-based system for inserting a target sequence into a genome. This disclosure is in part based on bioinformatics analysis for identifying retrotransposase sequences and associated 5'UTR and 3'UTR from various organisms (see Tables 3B, 10, 11, and X). Further examples of retrotransposon elements are listed, for example, in Tables 1 and 2 of PCT application PCT / US2019 / 048607, which is incorporated herein by reference in its entirety.
[0612] In some embodiments, the systems described herein may have several advantages compared to various prior systems. For example, this disclosure describes a retrotransposase capable of inserting long sequences of heterologous nucleic acids into a genome. Furthermore, the retrotransposase described herein can insert heterologous nucleic acids into endogenous sites in the genome, such as rDNA loci. This is in contrast to the Cre / loxP system, which requires a first step of inserting an exogenous loxP site before a second step of inserting the target sequence into the loxP site.
[0613] Gene-writer (trademark) genome editor Non-long-chain terminal repeat (LTR) retrotransposons are a type of mobile genetic element widely found in eukaryotic genomes. These include, for example, depurine / depyrimidine site endonuclease (APE) types, restriction enzyme-like endonuclease (RLE) types, and penelope-like element (PLE) types. APE-class retrotransposons contain two functional domains: an endonuclease / DNA-binding domain and a reverse transcriptase domain. Examples of APE-class retrotransposons can be found, for example, in Table 1 of PCT application PCT / US2019 / 048607 (which is incorporated herein by reference in its entirety), including, for example, the sequence listings and sequences mentioned in Table 1. RLE-class retrotransposons consist of three functional domains: a DNA-binding domain, a reverse transcription domain, and an endonuclease domain. Examples of RLE class retrotransposons can be found, for example, in Table 2 of PCT application PCT / US2019 / 048607 (which is incorporated herein by reference in its entirety), including, for example, the sequence listings and sequences mentioned in Table 2. The reverse transcriptase domain of non-LTR retrotransposons functions by binding to an RNA sequence template and reverse transcribing it to target DNA in the host genome. The RNA sequence template has a 3' untranslated region that specifically binds to a retrotransposase and a variable 5' region that generally has an open reading frame ("ORF") encoding the retrotransposase protein. The RNA sequence template may also include a 5' untranslated region that specifically binds to a retrotransposase. Penelope-like elements (PLEs) are distinct from both LTR retrotransposons and non-LTR retrotransposons. PLEs generally contain a reverse transcriptase domain similar to that of telomerases and group II introns, but different from that of APE and RLE elements, and an optional GIY-YIG endonuclease domain.
[0614] Other exemplary classes of retrotransposons include, but are not limited to, CRE, NeSL, R4, R2, Hero, L1, RTE (e.g., AviRTE_Brh or BovB), I, Jockey, CR1, Rex1, RandI / Dualen, Penelope or Penelope-like (PLE) (e.g., Penelope_SM), Tx1, RTEX, Crack, Nimb, Proto1, Proto2, RTETP, L2, Tad1, Loa, Ingi, Outcast, R1, Daphne, L2A, L2B, Ambal, Vingi (e.g., Vingi-1_EE), and Kiri retrotransposons.
[0615] The elements of such retrotransposons described herein can be functionally modularized and / or modified, for example by reverse transcription, to target, edit, modify, or manipulate a target DNA sequence, to insert a target (e.g., heterologous) nucleic acid sequence into a target genome, e.g., a mammalian genome. Such modularized and modified nucleic acids, polypeptide compositions, and systems are described herein and referred to as Gene Writer® gene editors. A Gene Writer® gene editor system comprises: (A) a polypeptide or nucleic acid encoding a polypeptide (the polypeptide comprises (i) a reverse transcriptase domain and (x) an endonuclease domain having DNA-binding function, or (y) an endonuclease domain and a separate DNA-binding domain); and (B) a template RNA comprising (i) a sequence that binds to the polypeptide and (ii) a heterologous insertion sequence. For example, a Gene Writer genome editor protein may comprise a DNA-binding domain, a reverse transcriptase domain, and an endonuclease domain. In other embodiments, a Gene Writer® genome editor protein may comprise a reverse transcriptase domain and an endonuclease domain. In certain embodiments, elements of the Gene Writer® gene editor polypeptide may be derived from retrotransposons, such as APE-type, RLE-type, or PLE-type retrotransposons, or from the sequence of a part or domain thereof. In some embodiments, the RLE-type non-LTR retrotransposon is derived from the R2, NeSL, HERO, R4, or CRE clade. In some embodiments, the Gene Writer genome editor is derived from the R4 element X4_Line, which is found in the human genome. In some embodiments, the APE-type non-LTR retrotransposon is derived from the R1 or Tx1 clade. In some embodiments, the Gene Writer genome editor is derived from the Tx1 element Mare6, which is found in the human genome.The RNA template elements of the Gene Writer® gene editor system are typically heterogeneous to the polypeptide elements, providing a target sequence to be inserted (reverse transcribed) into the host genome. In some embodiments, the Gene Writer genome editor protein is capable of target-prime reverse transcription.
[0616] In some embodiments, the Gene Writer genome editor is combined with a second polypeptide. In some embodiments, the second polypeptide is obtained from an APE-type non-LTR retrotransposon. In some embodiments, the second polypeptide has a zinc knuckle-like motif. In some embodiments, the second polypeptide is a homolog of the Gag protein.
[0617] In some embodiments, the Gene Writer genome editor includes retrotransposase sequences of elements listed in Table X. Table X provides a set of nucleic acid and amino acid sequences (listed by Repbase gene name and species) with associated GenBank accession numbers available. The Repbase nucleic acid and amino acid sequences of elements listed in Table X are incorporated herein by reference in their entirety. The GenBank sequences of elements listed in Table X are also incorporated herein by reference in their entirety. The nucleic acid sequences in Table X include, in some examples, polypeptide-coding sequences (e.g., open reading frames) (e.g., protein-coding sequences of the corresponding Repbase entries (incorporated herein by reference)). The nucleic acid sequences in Table X include, in some examples, 5'UTR sequences (e.g., 5'UTR sequences of the corresponding Repbase entries (incorporated herein by reference)). The nucleic acid sequences in Table X include, in some examples, 3'UTR sequences (e.g., 3'UTR sequences of the corresponding Repbase entries (incorporated herein by reference)).
[0618] In some embodiments, the open reading frames (ORFs) of the amino acid sequences in the Repbase are annotated. In some embodiments, the ORFs of the amino acid sequences in the Repbase are not annotated. If the ORFs are not annotated, those skilled in the art can identify them, for example, by performing one or more translations (e.g., full-frame translations) and optionally comparing the translations to a reverse transcriptase sequence or consensus motif. In some embodiments, for example, the amino acid sequences in Table X used herein are amino acid sequences listed in the corresponding Repbase entries, or amino acid sequences encoded by the nucleic acid sequences of the corresponding Repbase entries. In some embodiments, the amino acid sequences encoded by the elements in Table X are amino acid sequences encoded by the full-length sequences of the elements listed in Table X, or sequences having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the full-length sequences of the elements listed in Table X may include one or more (e.g., all) of the 5'UTR, polypeptide coding sequence, and 3'UTR of the retrotransposon described herein. In some embodiments, the amino acid sequences in Table X are amino acid sequences encoded by the full-length sequences of the elements listed in Table X, or sequences having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the 5'UTR of the elements in Table X includes the 5'UTR of the full-length sequence of the elements listed in Table X, or sequences having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the 3'UTR of an element in Table X includes the 3'UTR of the full-length sequence of an element listed in Table X, or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
[0619] Table X also lists the host organisms from which nucleic acid sequences were obtained, and the domains present in the polypeptide encoded by the open reading frames of the nucleic acid sequences. The domains listed in Table X are indicated as domain identifiers, which correspond to the InterPro domain entries listed in Table Y. Thus, the Repbase sequences listed in Table X can encode polypeptides containing one or more domains shown in Table X by their domain identifiers. The specific domains or domain types associated with each domain identifier are described in more detail in Table Y. In some embodiments, the target domain (e.g., as listed in Table Y) can be identified in the nucleic acid sequence (e.g., as listed in Table X) by performing full-frame translation and identifying the amino acid sequence encoding the target domain.
[0620] Retrotransposon detection tool As a result of repetitive mobility over time, transposable elements in genomic DNA often exist as tandem or scattered repeats (Jurka Curr Opin Struct Biol 8,333-337 (1998)). Using tools that can recognize such repeats, new elements can be identified from genomic DNA and added to databases, such as Repbase (Jurka et al Cytogenet Genome Res 110,462-467 (2005)). One such tool for identifying repeats that may contain transposable elements is RepeatFinder (Volfovsky et al Genome Biol 2 (2001)), which analyzes the repeat structure of genomic sequences. Repeats can be further collected and analyzed using additional tools, such as Censor (Kohany et al BMC Bioinformatics 7,474 (2006)). The Censor package retrieves genomic repeats and annotates them using various BLAST approaches for known transposable elements. Using full frame translation, an ORF can be generated for comparison.
[0621] Another exemplary method for identifying transposable elements is RepeatModeler2, which automates the discovery and annotation of transposable elements in genomic sequences (Flynn et al bioRxiv (2019)). In addition to achieving this with available packages such as Censor, it is possible to perform full-frame translation of a given genome or sequence and annotate it with a protein domain tool such as InterProScan (which uses the InterPro database to tag domains in a given amino acid sequence) (Mitchell et al. Nucleic Acids Res 47, D351-360 (2019)), enabling the identification of potential proteins containing domains associated with known transposable elements (e.g., domains of elements listed in Table X).
[0622] Retrotransposons can be further classified by their reverse transcriptase domain using tools such as RT class 1 (Kapitonov et al Gene 448, 207-213 (2009)).
[0623] polypeptide components of the Gene Writer gene editor system RT domain: In certain embodiments of the present invention, the reverse transcriptase domain of the Gene Writer system is based on the reverse transcriptase domain of an APE-type or RLE-type non-LTR retrotransposon, or a PLE-type retrotransposon. The wild-type reverse transcriptase domain of an APE-type, RLE-type, or PLE-type retrotransposon can be used in the Gene Writer system or modified (e.g., by inserting, deleting, or substituting one or more residues) to alter the reverse transcriptase activity of the target DNA sequence. In some embodiments, the reverse transcriptase is modified from its native sequence so that the codon usage is altered, for example, by improving it for human cells. In some embodiments, the reverse transcriptase domain is a heterologous reverse transcriptase from a different retrovirus, retroron, diversity-generating retroelement, retroplasmid, group II intron, LTR retrotransposon, non-LTR retrotransposon, or other source, e.g., those exemplified in Table Z1 or those containing the domains listed in Table Z2. In certain embodiments, the Gene Writer system includes polypeptides containing the reverse transcriptase domain of RLE-type non-LTR retrotransposons from the R2, NeSL, HERO, R4, or CRE clades; APE-type non-LTR retrotransposons from the L1, RTE, I, Jockey, CR1, Rex1, RandI / Dualen, T1, RTEX, Crack, Nimb, Proto1, Proto2, RTETP, L2, Tad1, Loa, Ingi, Outcast, R1, Daphne, L2A, L2B, Ambal, Vingi, or Kiri clades; or PLE-type non-LTR retrotransposons. In certain embodiments, the Gene Writer system includes polypeptides containing the reverse transcriptase domain of retrotransposons listed in Table 10, Table 11, Table X, Table Z1, Table Z2, or Table 3A or 3B.In some embodiments, the amino acid sequence of the reverse transcriptase domain of the Gene Writer system is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and at least about 99% identical to the amino acid sequence of the reverse transcriptase domain of retrotransposons whose DNA sequences are described in Tables 10, 11, X, Z1, Z2, or 3A or 3B. The reverse transcriptase domain can be identified based on homology with other known reverse transcriptase domains using routine tools such as the Basic Local Alignment Search Tool (BLAST). In some embodiments, the reverse transcriptase domain is modified, for example, by site-directed mutation. In some embodiments, the reverse transcriptase domain is manipulated to bind to a heterologous template RNA.
[0624] In some embodiments, the polypeptide (e.g., the RT domain) includes an RNA-binding domain that specifically binds to an RNA sequence, for example. In some embodiments, the template RNA includes an RNA sequence that is specifically bound by the RNA-binding domain.
[0625] In some embodiments, the RT domain exhibits increased stringency for target-primed reverse transcription (TPRT) initiation compared to, for example, an endogenous RT domain. In some embodiments, the RT domain initiates TPRT when 3nts within the target site immediately upstream of the first-strand nick, for example, the genomic DNA priming the RNA template, have at least 66% or 100% complementarity to 3nts of homology in the RNA template. In some embodiments, the RT domain initiates TPRT when less than 5nts of mismatch exist between the homology of the template RNA and the target DNA-priming reverse transcription (e.g., less than 1, 2, 3, 4, or 5nts of mismatch). In some embodiments, the RT domain is modified to increase stringency in priming mismatches in the TPRT reaction, for example, the RT domain either does not tolerate any mismatches within the priming region or tolerates fewer mismatches compared to a wild-type (e.g., unmodified) RT domain. In some embodiments, the RT domain includes an HIV-1 RT domain. In several embodiments, the HIV-1 RT domain initiates synthesis at a lower level, even with three nucleotide mismatches, compared to alternative RT domains (e.g., as described by Jamburthugoda and Eickbush J Mol Biol 407(5):661-672 (2011) (the entire text of which is incorporated herein by reference)).
[0626] In some embodiments, the RT domain forms a dimer (e.g., a heterodimer or a homodimer). In some embodiments, the RT domain is a monomer. In some embodiments, the RT domain, for example, a retroviral RT domain, functions naturally as a monomer or a dimer (e.g., a heterodimer or a homodimer). In some embodiments, the RT domain functions naturally as a monomer, for example, derived from a monomeric virus. Exemplary monomeric RT domains, their viral sources, and their associated RT signatures can be found in Table 30, along with descriptions of domain signatures in Table 32. In some embodiments, the RT domain of the system described herein includes the amino acid sequence in Table 30, or a functional fragment or variant thereof, or a sequence having at least 70%, 80%, 90%, 95%, or 99% identity thereto.In several embodiments, the RT domain is murine leukemia virus (MLV; sometimes referred to as MoMLV) (e.g., P03355), porcine endogenous retrovirus (PERV) (e.g., UniProt Q4VFZ2), mouse mammary tumor virus (MMTV) (e.g., UniProt P03365), Mason-Pfizer monkey virus (MPMV) (e.g., UniProt P07572), bovine leukemia virus (BLV) (e.g., UniProt P03361), human T-cell leukemia virus-1 (HTLV-1) (e.g., UniProt P03362), human foamy virus (HFV) (e.g., UniProt The RT domains are selected from the following: P14350), simian foamy virus (SFV) (e.g., UniProt P23074), bovine foamy / syncytial virus (BFV / BSV) (e.g., UniProt O41894), or functional fragments or variants thereof (e.g., amino acid sequences having at least 70%, 80%, 90%, 95%, or 99% identity thereto). In some embodiments, the RT domain is dimerized in its innate functionality. Exemplary dimeric RT domains, their viral sources, and associated RT signatures can be found in Table 31, along with descriptions of domain signatures in Table 32. In some embodiments, the RT domains of the systems described herein include the amino acid sequences in Table 31, or their functional fragments or variants, or sequences having at least 70%, 80%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain is derived from a virus that functions as a dimer.In several embodiments, the RT domain is derived from avian sarcoma / leukemia virus (ASLV) (e.g., UniProt A0A142BKH1), Rous sarcoma virus (RSV) (e.g., UniProt P03354), avian myeloblastosis virus (AMV) (e.g., UniProt Q83133), human immunodeficiency virus type I (HIV-1) (e.g., UniProt P03369), human immunodeficiency virus type II (HIV-2) (e.g., UniProt P15833), simian immunodeficiency virus (SIV) (e.g., UniProt P05896), bovine immunodeficiency virus (BIV) (e.g., UniProt The RT domain is selected from P19560), equine infectious anemia virus (EIAV) (e.g., UniProt P03371), or feline immunodeficiency virus (FIV) (e.g., UniProt P16088) (Herschhorn and Hizi Cell Mol Life Sci 67(16):2717-2747(2010)), or a functional fragment or variant thereof (e.g., an amino acid sequence having at least 70%, 80%, 90%, 95%, or 99% identity). In nature, heterodimeric RT domains may also function as homodimers in some embodiments. In some embodiments, the dimeric RT domain is expressed as a fusion protein, for example, as a homodimeric fusion protein or a heterodimeric fusion protein. In some embodiments, the RT function of the system is satisfied by multiple RT domains (e.g., as described herein).In further embodiments, the multiple RT domains may be fused or separated, and may, for example, reside on the same polypeptide or on different polypeptides.
[0627] In some embodiments, the Gene Writer described herein includes an integrase domain, for example, the integrase domain may be part of the RT domain. In some embodiments, the RT domain (for example, as described herein) includes an integrase domain. In some embodiments, the RT domain (for example, as described herein) lacks an integrase domain or includes an integrase domain that has been inactivated by mutation or deletion. In some embodiments, the Gene Writer described herein includes a ribonuclease H domain, for example, the ribonuclease H domain may be part of the RT domain. In some embodiments, the RT domain (for example, as described herein) includes a ribonuclease H domain, for example, an endogenous ribonuclease H domain or a heterologous ribonuclease H domain. In some embodiments, the RT domain (for example, as described herein) lacks a ribonuclease H domain. In some embodiments, the RT domain (for example, as described herein) includes a ribonuclease H domain that has been added to, deleted from, mutated, or replaced with a heterologous ribonuclease H domain. In some embodiments, a mutation in the ribonuclease H domain produces a polypeptide exhibiting lower ribonuclease activity, for example, by measurement by the method in Kotewicz et al. Nucleic Acids Res 16(1):265-277 (1988) (which is incorporated herein by reference in its entirety), compared to other similar domains without the mutation, by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%. In some embodiments, the ribonuclease H activity is lost.
[0628] In some embodiments, the RT domain undergoes mutation, resulting in increased fidelity compared to other similar domains that do not have mutations. For example, in some embodiments, the YADD or YMDD motif within the RT domain (e.g., in the reverse transcriptase) is substituted with YVDD. In several embodiments, the substitution of YADD, YMDD, or YVDD results in higher fidelity in retroviral reverse transcriptase activity (e.g., described in Jamburthugoda and Eickbush J Mol Biol 2011; the entire text is incorporated herein by reference).
[0629] The diversity of reverse transcriptases (including, for example, those containing the RT domain) is used by prokaryotes (Zimmerly et al. Microbiol Spectr 3(2):MDNA3-0058-2014(2015); Lampson BC(2007) Prokaryotic Reverse Transcriptases. In:Polaina J., MacCabe AP(eds) Industrial Enzymes. Springer, Dordrecht), viruses (Herschhorn et al. Cell Mol Life Sci 67(16):2717-2747(2010); Menendez-Arias et al. Virus Res 234:153-176(2017)), and mobile factors (Eickbush et al. Virus Res 134(1-2):221-234(2008); Craig et al. Mobile DNA III 3rd See, but are not limited to, the references listed in Ed.DOI:10.1128 / 9781555819217(2015) (each of these references is incorporated herein by reference).
[0630] [Table 1]
[0631] [Table 2]
[0632] Table 3
[0633] Table 4
[0634] Table 5
[0635] Table 6
[0636] Table 7
[0637] Table 8
[0638] Table 9
[0639] Table 10
[0640] Table 11
[0641] Table 12
[0642] Table 13
[0643] Table 14
[0644] Table 15
[0645] Table 16
[0646] Table 17
[0647] Table 18
[0648] Table 19
[0649] Table 20
[0650] Table 21
[0651] Table 22
[0652] Table 23
[0653] Table 24
[0654] Table 25
[0655] Table 26
[0656] Table 27
[0657] Table 28
[0658] Table 29
[0659] Table 30
[0660] Table 31
[0661] Table 32
[0662] Table 33
[0663] [Table 34]
[0664] Endonuclease In some embodiments, the polypeptide includes an endonuclease domain (e.g., a heterologous endonuclease domain). In certain embodiments, the endonuclease / DNA binding domain of an APE-type retrotransposon, the endonuclease domain of an RLE-type retrotransposon, or the endonuclease domain of an PLE-type retrotransposon can be used in the Gene Writer system described herein, or can be modified (e.g., by insertion, deletion, or substitution of one or more residues). In some embodiments, the endonuclease domain or endonuclease / DNA binding domain is modified from its native sequence so as to alter the codon usage, for example, by improving it for human cells. In some embodiments, the endonuclease element is a heterologous endonuclease element such as Fok1 nuclease, Cas9, Cas9 nickasase, type II restriction enzyme-like endonuclease (RLE-type nuclease), or another RLE-type endonuclease (also known as REL). In some embodiments, the heterologous endonuclease domain cleaves both DNA strands, forming a double-strand break. In some embodiments, the heterologous endonuclease activity has nickase activity and does not form a double-strand break. The amino acid sequences of the reverse transcriptase domains of the Gene Writer systems described herein may be at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and at least about 99% identical to the amino acid sequences of the endonuclease domains of retrotransposons whose DNA sequences are described in Tables X, Z1, Z2, 3A, or 3B. The endonuclease domains can be identified based on homology with other known endonuclease domains using tools such as the Basic Local Alignment Search Tool (BLAST). In certain embodiments, the heterologous endonuclease is Cas9 or Cas9 nickase or a functional fragment thereof. In certain embodiments, the heterologous endonuclease is Fok1 or a functional fragment thereof.In certain embodiments, the heterologous endonuclease is a Holliday junction resolverase or its homolog, such as Holliday junction resorbase-Ssol Hje (Govindaraju et al., Nucleic Acids Research 44:7, 2016) derived from Sulfolobus solfataricus. In certain embodiments, the heterologous endonuclease is an endonuclease of a large fragment of a spliceosome protein, such as Prp8 (Mahbub et al., Mobile DNA 8:16, 2017). For example, the Gene Writer polypeptide described herein may comprise a reverse transcriptase domain derived from an APE, RLE, or PLE type retrotransposon and an endonuclease domain containing Fok1 or a functional fragment thereof. In yet another embodiment, the homologous endonuclease domain is modified, for example, by site-directed mutation, to alter DNA endonuclease activity. In yet another embodiment, the endonuclease domain is modified to remove any potential latent DNA sequence specificity.
[0665] In some embodiments, the Gene Writer polypeptide has the function of cleaving a target DNA site by an endonuclease domain. In some embodiments, the endonuclease domain is also a DNA-binding domain. In some embodiments, the endonuclease domain is also a template nucleic acid (e.g., template RNA)-binding domain. For example, in some embodiments, the polypeptide includes a CRISPR-associated endonuclease domain that binds to a template RNA, including gRNA, and binds to a target DNA sequence (e.g., complementary to a portion of the gRNA) and cleaves the target DNA sequence. In certain embodiments, the endonuclease / DNA-binding domain of an APE-type retrotransposon or an RLE-type retrotransposon can be used in the Gene Writer system described herein, or can be modified (e.g., by insertion, deletion, or substitution of one or more residues). In some embodiments, the endonuclease domain or endonuclease / DNA-binding domain is modified from its native sequence so as to alter its codon usage, for example, by improving it for human cells.
[0666] In some embodiments, the endonuclease element is a heterologous endonuclease element, such as Fok1 nuclease, type II restriction I-like endonuclease (RLE-type nuclease), or another RLE-type endonuclease (also known as REL). In some embodiments, the heterologous endonuclease activity has nickase activity and does not form double-strand breaks. The amino acid sequence of the endonuclease domain of the Gene Writer system described herein may be at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to the amino acid sequence of the endonuclease domain of a retrotransposon whose DNA sequence is referenced in Tables 1, 2, 3A, or 3B. Endonuclease domains can be identified based on homology to other known endonuclease domains, for example, using tools such as the Basic Local Alignment Search Tool (BLAST). In certain embodiments, the heterologous endonuclease is Fok1 or a functional fragment thereof. In certain embodiments, the heterologous endonuclease is a Holliday junction resolverase or its homolog, e.g., Holliday junction cleavage enzyme (Ssol Hje) from Sulfolobus solfataricus (Govindaraju et al., Nucleic Acids Research 44:7, 2016). In certain embodiments, the heterologous endonuclease is an endonuclease of a large fragment of a spliceosome protein such as Prp8 (Mahbub et al., Mobile DNA 8:16, 2017). In certain embodiments, the heterologous endonuclease is derived from a CRISPR-related protein, e.g., Cas9. In certain embodiments, the heterologous endonuclease is modified to have only ssDNA cleavage activity, for example, Cas9 nickasase, or only nickas activity.For example, the Gene Writer polypeptide described herein may include a reverse transcriptase domain from an APE-type or RLE-type retrotransposon and an endonuclease domain containing Fok1 or a functional fragment thereof. In yet another embodiment, the homologous endonuclease domain is modified, for example, by site-directed mutation to alter its DNA endonuclease activity. In yet another embodiment, the endonuclease domain is modified to remove sequence specificity of any latent DNA.
[0667] In some embodiments, the endonuclease domain has nickase activity and does not form double-strand breaks. In some embodiments, the endonuclease domain forms single-strand breaks more frequently than double-strand breaks, for example, at least 90%, 95%, 96%, 97%, 98%, or 99% of the breaks are single-strand breaks, or less than 10%, 5%, 4%, 3%, 2%, or 1% of the breaks are double-strand breaks. In some embodiments, the endonuclease substantially does not form double-strand breaks. In some embodiments, the endonuclease does not form detectable levels of double-strand breaks.
[0668] In some embodiments, the endonuclease domain has nickase activity that cleaves the target site DNA on the strand being edited; for example, in some embodiments, the endonuclease domain cleaves the genomic DNA at the target site near the modification site on the strand that will be extended by the writing domain. In some embodiments, the endonuclease domain has nickase activity that cleaves the target site DNA on the strand being edited but does not cleave the target site DNA on the strand not being edited. For example, when the polypeptide contains a CRISPR-related endonuclease domain that has nickase activity and does not form double-strand breaks, in some embodiments, the CRISPR-related endonuclease domain cleaves the target site DNA strand containing the PAM site (for example, but does not cleave the target site DNA strand that does not contain the PAM site).
[0669] In some other embodiments, the endonuclease domain has nickase activity that cleaves the target site DNA in both the edited and unedited strands. While not intended to be bound by any particular theory, after the writing domain (e.g., RT domain) of the polypeptide described herein polymerizes (e.g., reverse transcribes) from a heterologous target sequence in the template nucleic acid (e.g., template RNA), the cellular DNA repair mechanism must repair the nick on the edited DNA strand. The target site DNA here comprises two distinct sequences relative to the edited DNA strand: the first corresponding to the original genomic DNA, and the second corresponding to the one polymerized from the heterologous target sequence. The two distinct sequences are thought to equilibrate with each other, and one hybridizes with the unedited strand first, then the other, where the incorporation of the cellular DNA repair mechanism into its repair target site is considered random. While not intended to be bound by any particular theory, the introduction of further nicks into the unedited strand may bias the cellular DNA repair mechanism to use the heterologous target sequence more frequently than the original genomic sequence. In some embodiments, the additional nicks are located at 5' or 3' of the target site modification (e.g., insertion, deletion, or substitution), or at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, or 150 nucleotides relative to the nicks on the strand being edited.
[0670] Alternatively or additionally, and without intending to be bound by any particular theory, it is thought that further nicks to the unedited strand may facilitate second strand synthesis. In some embodiments, if Gene Writer™ inserts or replaces a portion of the edited strand, synthesis of a new sequence corresponding to the insertion / replacement in the unedited strand is required.
[0671] In some embodiments, the polypeptide comprises a single domain having endonuclease activity (e.g., a single endonuclease domain) that cleaves both the strand to be edited and the unedited strand. For example, in such embodiments, the endonuclease domain may be a CRISPR-associated endonuclease domain, and the template nucleic acid (e.g., template RNA) comprises a gRNA that leads to nicking of the strand to be edited and a further gRNA that leads to nicking of the unedited strand. In some embodiments, the polypeptide comprises multiple domains having endonuclease activity, where a first endonuclease domain cleaves the strand to be edited and a second endonuclease domain cleaves the unedited strand (optionally, the first endonuclease domain may not cleave the unedited strand (e.g., is unable to do so), and the second endonuclease domain may not cleave the strand to be edited (e.g., is unable to do so)).
[0672] In some embodiments, the endonuclease domain can puncture the first and second chains. In some embodiments, the first and second chain nicks occur at the same location in the target site but not on the reverse chain. In some embodiments, the second chain nick occurs at a twisted location, for example, upstream or downstream of the first nick. In some embodiments, the endonuclease domain generates a deletion at the target site if the second chain nick is upstream of the first chain nick. In some embodiments, the endonuclease domain generates a duplication of the target site if the second chain nick is downstream of the first chain nick. In some embodiments, the endonuclease domain does not generate duplication and / or deletion if the first and second chain nicks occur at the same location in the target site (for example, as described in Gladyshev and Arkhipova Gene 2009; the whole of which is incorporated herein by reference). In some embodiments, the endonuclease domain has modified activity depending on the protein structure or RNA binding state, for example (as described in Christensen et al. PNAS 2006; the whole thereof is incorporated herein by reference), which promotes the nicking of the first or second strand.
[0673] In some embodiments, the endonuclease domain includes a meganuclease or a functional fragment thereof. In some embodiments, the endonuclease domain includes a homing endonuclease or a functional fragment thereof. In some embodiments, the endonuclease domain includes a meganuclease or a functional fragment or variant thereof from the LAGLIDADG, GIY-YIG, HNH, His-Cys Box, or PD-(D / E)XK family, for example, having a conserved amino acid motif as indicated by the family name. In some embodiments, the endonuclease domain includes a meganuclease or fragment thereof selected from, for example, I-SmaMI (Uniprot F7WD42), I-SceI (Uniprot P03882), I-AniI (Uniprot P03880), I-DmoI (Uniprot P21505), I-CreI (Uniprot P05725), I-TevI (Uniprot P13299), I-OnuI (Uniprot Q4VWW5), or I-BmoI (Uniprot Q9ANR6). In some embodiments, the meganuclease is naturally present in its functional form as a monomer, e.g., I-SceI, I-TevI, or a dimer, e.g., I-CreI. For example, LAGLIDADG meganucleases having a single copy of the LAGLIDADG motif generally form homodimers, while members having two copies of the LAGLIDADG motif are generally found as monomers. In some embodiments, meganucleases that normally form as dimers are expressed as fusions, for example, as an I-CreI dimer fusion, where the two subunits are optionally linked by a linker as a single ORF (Rodriguez-Fornes et al. Gene Therapy 2020; the whole is incorporated herein by reference).In some embodiments, the meganuclease or its functional fragment is modified to preferentially nickase activity on one strand of a double-stranded DNA molecule, such as I-SceI (K122I and / or K223I) (Niu et al. J Mol Biol 2008), I-AniI (K227M) (McConnell Smith et al. PNAS 2009), or I-DmoI (Q42A and / or K120M) (Molina et al. J Biol Chem 2015). In some embodiments, the meganuclease or its functional fragment having this preference for single-strand breaks is used, for example, as an endonuclease domain having nickase activity. In some embodiments, the endonuclease domain includes a meganuclease or functional fragment thereof that naturally targets, or is modified to target, a safe port site, e.g., an SH6 site targeting I-CreI (Rodriguez-Fornes et al., above). In some embodiments, the endonuclease domain includes a meganuclease or functional fragment thereof having a sequence-resistant catalytic domain, e.g., I-TevI that recognizes the minimal motif CNNNG (Kleinstiver et al. PNAS 2012). In some embodiments, the target sequence-resistant catalytic domain is fused to the DNA-binding domain to induce activity, for example, by fusing I-TevI to (i) a Zn finger for producing Tev-ZFE (Kleinstiver et al. PNAS 2012), (ii) another meganuclease for producing MegaTev (Wolfs et al. Nucleic Acids Res 2014), and / or (iii) Cas9 for producing TevCas9 (Wolfs et al. PNAS 2016).
[0674] In some embodiments, the endonuclease domain includes a restriction enzyme, e.g., type IIS or type IIP restriction enzyme. In some embodiments, the endonuclease domain includes a type IIS restriction enzyme, e.g., FokI or a fragment or variant thereof. In some embodiments, the endonuclease domain includes a type IIP restriction enzyme, e.g., PvuII or a fragment or variant thereof. In some embodiments, the dimeric restriction enzyme is expressed as a fusion, e.g., a FokI dimer fusion, so that it functions as a single chain (Minczuk et al. Nucleic Acids Res 36(12):3926-3938(2008)).
[0675] Further uses of endonuclease domains are described, for example, in Guha and Edgell Int J Mol Sci 18(22):2565(2017) (the entire text of which is incorporated herein by reference).
[0676] In some embodiments, the endonuclease domain or DNA-binding domain (as described herein, for example) comprises Streptococcus pyogenes Cas9 (SpCas9) or a functional fragment or variant thereof. In some embodiments, the endonuclease domain or DNA-binding domain comprises modified SpCas9. In several embodiments, the modified SpCas9 comprises modifications that alter the protospacer-adjacent motif (PAM) specificity. In several embodiments, the PAM has specificity for the nucleic acid sequence 5'-NGT-3'. In several embodiments, the modified SpCas9 includes, for example, one or more amino acid substitutions at one or more positions of L1111, D1135, G1218, E1219, A1322, or R1335, which are selected from, for example, L1111R, D1135V, G1218R, E1219F, A1322R, and R1335V. In several embodiments, the modified SpCas9 includes the amino acid substitution T1337R and one or more other amino acid substitutions selected from the following: L1111, D1135L, S1136R, G1218S, E1219V, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, R1335Q, T1337, T1337L, T1337Q, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337H, T1337Q, and T1337M, or the corresponding amino acid substitutions. In several embodiments, the modified SpCas9 includes: (i) one or more amino acid substitutions selected from: D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q, and T1337; and (ii) one or more other amino acid substitutions selected from: L1111R, G1218R, E1219F, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, T1337L, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337R, T1337H, T1337Q, and T1337M, or the corresponding amino acid substitutions.
[0677] In some embodiments, the endonuclease domain or DNA-binding domain (for example, as described herein) includes a Cas domain, such as a Cas9 domain. In several embodiments, the endonuclease domain or DNA-binding domain includes a nuclease-active Cas domain, a Cas nickase (nCas) domain, or a nuclease-inactive Cas (dCas). In several embodiments, the endonuclease domain or DNA-binding domain includes a nuclease-active Cas9 domain, a Cas9 nickase (nCas9) domain, or a nuclease-inactive Cas9 (dCas). In some embodiments, the endonuclease domain or DNA-binding domain includes Cas9 domains (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA-binding domain comprises *Streptococcus pyogenes* or *S. thermophilus* Cas9, or a functional fragment thereof. In some embodiments, the endonuclease domain or DNA-binding domain comprises a Cas9 sequence, for example, as described in Chylinski, Rhun, and Charpentier (2013) RNA Biology 10:5, 726-737 (incorporated herein by reference). In some embodiments, the endonuclease domain or DNA-binding domain comprises the HNH nuclease subdomain and / or RuvC1 subdomain of Cas, for example, Cas9, or a variant thereof, as described herein.In some embodiments, the endonuclease domain or DNA-binding domain includes Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA-binding domain includes a Cas polypeptide (e.g., an enzyme) or a functional fragment thereof. In several embodiments, the Cas polypeptide (e.g., an enzyme) includes: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (e.g., Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpf l, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Cs y3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1 , Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, C Selected from sa5, type II Cas effector protein, type V Cas effector protein, type VI Cas effector protein, CARF, DinG, Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12b / C2c1, Cas12c / C2c3, SpCas9(K855A), eSpCas9(1.1), SpCas9-HF1, hyper-accurate Cas9 variants (HypaCas9), their homologs, their modified or manipulated versions, and / or their functional fragments.In some embodiments, Cas9 includes one or more substitutions selected from, for example, H840A, D10A, P475A, W476A, N477A, D1125A, W1126A, and D1127A. In some embodiments, Cas9 includes one or more mutations at positions selected from, for example, D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987, including one or more substitutions selected from, for example, D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A. In some embodiments, the endonuclease domain or DNA-binding domain is: Corynebacterium ulcerans, Corynebacterium diphtheria, Spiroplasma syrphidicola, Prevotella intermedia, Spiroplasma taiwanense, Streptococcus iniae, Bellilla baltica, Psychroflexus torquis, Streptococcus thermophilus, Listeria innocua, Campylobacter jejuni, Neisseria meningitidis. It contains a Cas (e.g., Cas9) sequence derived from *Streptococcus meningitidis*, *Streptococcus pyogenes*, or *Staphylococcus aureus*, or a functional fragment or mutant thereof.
[0678] In some embodiments, the endonuclease domain or DNA-binding domain (as described herein, for example) includes a Cpf1 domain, which comprises one or more substitutions selected from, for example, D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, and D917A / E1006A / D1255A, at positions D917, E1006A, D1255A, or any combination thereof.
[0679] In some embodiments, the endonuclease domain or DNA-binding domain (as described herein, for example) includes spCas9, spCas9-VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL.
[0680] In some embodiments, the endonuclease domain or DNA-binding domain (as described herein, for example) includes an amino acid sequence listed in Table 37 below, or an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the endonuclease domain or DNA-binding domain includes an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 or fewer differences (e.g., mutations) from any of the amino acid sequences described herein.
[0681] [Table 35]
[0682] [Table 36]
[0683] [Table 37]
[0684] [Table 38]
[0685] In some embodiments, the Gene Writing polypeptide has an endonuclease domain containing Cas9 niccas, for example, Cas9H840A. In several embodiments, Cas9H840A has the following amino acid sequence. Cas9 nickas (H840A): [ka]
[0686] In some embodiments, the Gene Writing polypeptide comprises a wild-type M-MLV RT domain from a retroviral reverse transcriptase, for example, containing the following sequence: M-MLV(WT): [ka]
[0687] In some embodiments, the Gene Writing polypeptide includes an RT domain from a retroviral reverse transcriptase, for example, M-MLV RT containing the following sequence: [ka]
[0688] In some embodiments, the Gene Writing polypeptide comprises an RT domain from a retroviral reverse transcriptase containing the sequence of amino acids 659-1329 of NP_057933. In several embodiments, the Gene Writing polypeptide further comprises one additional amino acid at the N-terminus of the sequence of amino acids 659-1329 of NP_057933, for example, as shown below. [ka] Core RT (bold), above annotation Ribonuclease H (underlined), see above note
[0689] In several embodiments, the Gene Writing polypeptide further comprises one additional amino acid at the C-terminus of the sequence of amino acids 659-1329 of NP_057933. In several embodiments, the Gene Writing polypeptide comprises a ribonuclease H1 domain (e.g., amino acids 1178-1318 of NP_057933).
[0690] In some embodiments, the retroviral reverse transcriptase domain, e.g., M-MLV RT, may contain one or more mutations from the wild-type sequence that can improve the characteristics of RT, e.g., thermal stability, processing capacity, and / or template binding. In some embodiments, the M-MLV RT domain includes a combination of mutations such as D200N, L603W, and T330P, selected from, for example, D200N, L603W, T330P, T306K, W313F, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, L435G, N454K, H594Q, D653N, R110S, and K103L, for example, further optionally including T306K and W313F. In some embodiments, the M-MLV RT used herein includes the mutations D200N, L603W, T330P, T306K, and W313F. In several embodiments, the mutant M-MLV RT includes the following amino acid sequence. M-MLV(PE2): [ka]
[0691] In some embodiments, the Gene Writer polypeptide may include a linker, such as a peptide linker, as shown in Table 7. In some embodiments, the Gene Writer polypeptide includes a flexible linker between the endonuclease and the RT domain, such as a linker containing the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSS. In some embodiments, the RT domain of the Gene Writer polypeptide may be located at the C-terminus relative to the endonuclease domain. In some embodiments, the RT domain of the Gene Writer polypeptide may be located at the N-terminus relative to the endonuclease domain.
[0692] [Table 39]
[0693] [Table 40]
[0694] [Table 41]
[0695] [Table 42]
[0696] In some embodiments, the Gene Writer polypeptide includes a dCas9 sequence containing the D10A and / or H840A mutation, for example, the following sequence. [ka]
[0697] In some embodiments, the template RNA molecule for use in the system includes, from 5' to 3', (1) a gRNA spacer; (2) a gRNA scaffold; (3) a heterologous target sequence; and (4) a 3' homology domain. In some embodiments: (1) A Cas9 spacer of approximately 18-22 nt, for example, 20 nt. (2) A gRNA scaffold comprising one or more hairpin loops, e.g., one, two, or three loops, for associating the template with the nickase Cas9 domain. In some embodiments, the gRNA scaffold has the sequence GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCC from 5' to 3'. (3) In some embodiments, the heteropurpose sequence has a length of, for example, 7–74 nt, for example, 10–20 nt, 20–30 nt, 30–40 nt, 40–50 nt, 50–60 nt, 60–70 nt, or 70–80 nt or 80–90 nt. In some embodiments, the first (approximately 5') base of the sequence is not C. (4) In some embodiments, the 3' homology domain that binds to the target priming sequence after nicking is, for example, 3-20 nt, for example, 7-15 nt, for example, 12-14 nt. In some embodiments, the 3' homology domain has a GC content of 40-60%.
[0698] A second gRNA, associated with the system, can assist in driving complete integration. In some embodiments, the second gRNA may target locations 0–200 nt away from the first strand nick, for example, 0–50 nt, 50–100 nt, or 100–200 nt away from the first strand nick. In some embodiments, the second gRNA can bind only to its target sequence after editing, for example, the gRNA binds to a sequence that is present in the heterologous target sequence but not in the initial target sequence.
[0699] In some embodiments, the Gene Writing system described herein is used to perform editing on HEK293, K562, U2OS, or HeLa cells. In some embodiments, the Gene Writing system is used to perform editing on primary cells, for example, primary cortical neurons from E18.5 mice.
[0700] In some embodiments, the reverse transcriptase or RT domain (e.g., as described herein) comprises a MoMLV RT sequence or a variant thereof. In some embodiments, the MoMLV RT sequence comprises one or more mutations selected from D200N, L603W, T330P, T306K, W313F, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, L435G, N454K, H594Q, D653N, R110S, and K103L. In some embodiments, the MoMLV RT sequence optionally comprises a combination of mutations such as D200N, L603W, and T330P, further comprising T306K and / or W313F.
[0701] In some embodiments, the endonuclease domain (e.g., as described herein) includes nCAS9, for example, the H840A mutation.
[0702] In some embodiments, the heterologous target sequence (for example, in the system described herein) has a nucleotide length of about 1 to 50, about 50 to 100, about 100 to 200, about 200 to 300, about 300 to 400, about 400 to 500, about 500 to 600, about 600 to 700, about 700 to 800, about 800 to 900, about 900 to 1000, or more.
[0703] In some embodiments, the RT and endonuclease domains are linked by a flexible linker, for example, containing the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSS.
[0704] In some embodiments, the endonuclease domain is the N-terminus of the RT domain. In some embodiments, the endonuclease domain is the C-terminus of the RT domain.
[0705] In some embodiments, the system incorporates a heterogeneous target sequence into the target site by TPRT, for example, as described herein.
[0706] In some embodiments, the systems or methods described herein include CRISPR DNA targeting enzymes or systems, or functional fragments or variants thereof, as described in U.S. Patent Publication No. 20200063126, U.S. Patent Publication No. 20190002889, or U.S. Patent Publication No. 20190002875 (each of which is incorporated herein by whole reference). For example, in some embodiments, the Gene Writer polypeptide or Cas endonuclease described herein includes a polypeptide sequence described in any of the applications cited in this paragraph, and in some embodiments, the template RNA or guide RNA includes a nucleic acid sequence described in any of the applications cited in this paragraph.
[0707] Template nucleic acid binding domain: A Gene Writer polypeptide typically includes a region capable of binding to a Gene Writer template nucleic acid (e.g., template RNA). In some embodiments, the template nucleic acid binding domain is an RNA binding domain. In some embodiments, the RNA binding domain is a modular domain capable of binding to RNA molecules containing a specific signature, e.g., a structural motif, e.g., a secondary structure present in the 3'UTR of a non-LTR retrotransposon. In other embodiments, the template nucleic acid binding domain (e.g., RNA binding domain) is contained within a reverse transcription domain, for example, a reverse transcriptase-derived component having a known signature for RNA selection, e.g., a secondary structure present in the 3'UTR of a non-LTR retrotransposon. In other embodiments, the template nucleic acid binding domain (e.g., RNA binding domain) is contained within a DNA binding domain. For example, in some embodiments, the DNA binding domain is a CRISPR-associated protein that recognizes the structure of a template nucleic acid (e.g., template RNA) containing gRNA. In some embodiments, gRNA is a short synthetic RNA composed of a scaffold sequence involved in CRISPR-associated protein binding and a user-defined targeting sequence of approximately 20 nucleotides for a genomic target. The complete structure of gRNA is described in Nishimasu et al. Cell 156, pp. 935-949 (2014). gRNA (also called sgRNA in single guide RNA) consists of crRNA and tracrRNA-derived sequences linked by an artificial tetraloop. The crRNA sequence can be divided into a guide (20nt) region and a repeat (12nt) region, while the tracrRNA sequence can be divided into an anti-repeat (14nt) and three tracrRNA stem-loops (Nishimasu et al. Cell 156, pp. 935-949 (2014)). In practice, the guide RNA sequence is generally 17-24 nucleotides long (e.g., 19, 20, or 21 nucleotides) and designed to be complementary to the target nucleic acid sequence. Custom gRNA generators and algorithms are commercially available for use in designing effective guide RNAs.In some embodiments, the gRNA comprises two RNA components derived from the natural CRISPR system, for example, crRNA and tracrRNA. As is well known in the art, gRNAs may also contain a chimeric single guide RNA (sgRNA) that includes sequences from both a tracrRNA (for binding to a nuclease) and at least one crRNA (for guiding the nuclease to a targeted sequence for editing / binding). It has also been demonstrated that chemically modified sgRNAs are effective for use with CRISPR-related proteins (see, for example, Hendel et al. (2015) Nature Biotechnol., 985-991). In some embodiments, the gRNA contains a nucleic acid sequence complementary to the DNA sequence associated with the target gene. In some embodiments, the polypeptide contains a DNA-binding domain that includes a CRISPR-related protein that binds to the gRNA, enabling the DNA-binding domain to bind to the target genomic DNA sequence. In some embodiments, the gRNA is contained within a template nucleic acid (e.g., template RNA), and therefore the DNA-binding domain is also a template nucleic acid-binding domain. In some embodiments, the polypeptide has RNA-binding function in multiple domains, which can, for example, bind to a gRNA structure in a CRISPR-related DNA-binding domain and to a 3'UTR structure in a reverse transcription domain derived from a non-LTR retrotransposon.
[0708] In some embodiments, the template nucleic acid (e.g., template RNA) includes a 3' target homology domain. In some embodiments, the 3' target homology domain is located on the 3' side of a heterologous target sequence and is complementary to a sequence adjacent to the site to be modified by the System described herein, or contains 1, 2, 3, 4, or 5 or fewer mismatches to a sequence complementary to a sequence adjacent to the site to be modified by the System / Gene Writer®. In some embodiments, the 3' target homology domain anneals to the target site and provides a binding site and a 3' hydroxyl group for TPRT initiation by the Gene Writer polypeptide. In some embodiments, the 3' target homology domain has lengths of 3-5, 5-10, 10-30, 10-25, 10-20, 10-19, 10-18, 10-17, 10-16, 10-15, 10-14, 10-13, 10-12, 10-11, 11-30, 11-25, 11-20, 11-19, and 11-1 8, 11-17, 11-16, 11-15, 11-14, 11-13, 11-12, 12-30, 12-25, 12-20, 12-19, 12-18, 12-17, 12-16, 12-15, 12-14, 12-13, 13-30, 13-25, 13-20, 13-19, 13-18, 13-17, 13 ~16, 13~15, 13~14, 14~30, 14~25, 14~20, 14~19, 14~18, 14~17, 14~16, 14~15, 15~30, 15~25, 15~20, 15~19, 15~18, 15~17, 15~16, 16~30, 16~25, 16~20, 16~19, 16~18, 16-17, 17-30, 17-25, 17-20, 17-19, 17-18, 18-30, 18-25, 18-20, 18-19, 19-30, 19-25, 19-20, 20-30, 20-25, or 25-30 nt, for example, lengths of 10-17, 12-16, or 12-14 nt.
[0709] In some embodiments, the template nucleic acid, for example, template RNA, may include gRNA (e.g., pegRNA). In some embodiments, the template nucleic acid, for example, template RNA, can bind to the Gene Writer® polypeptide through interaction between the gRNA of the template nucleic acid and a template nucleic acid binding domain, for example, an RNA binding domain (e.g., a heterologous RNA binding domain). In some embodiments, the heterologous RNA binding domain is a CRISPR / Cas protein, for example, Cas9.
[0710] In some embodiments, a template nucleic acid containing guide RNA (gRNA), e.g., a region of template RNA, adopts an underwound ribbon-like structure of the gRNA bound to target DNA (as described, e.g., Mulepati et al. Science 19 Sep 2014: Vol. 345, Issue 6203, pp. 1479-1484). While not intended to be bound by any particular theory, this non-standard structure is thought to be facilitated by a rotation of every 6 nucleotides from the RNA-DNA hybrid. Therefore, in some embodiments, a template nucleic acid containing gRNA, e.g., a region of template RNA, may tolerate an increase in mismatch with the target site at some intervals, e.g., every 6 bases. In some embodiments, a template nucleic acid containing gRNA with homology to the target site, e.g., a region of template RNA, may have fluctuating positions at regular intervals, e.g., every 6 bases, where base pairing with the target site is not required.
[0711] gRNA with inducible activity In some embodiments, the template nucleic acid, e.g., template RNA, includes an inducible activity gRNA. Inducible activity may be achieved by a template nucleic acid, e.g., template RNA, further comprising a blocking domain (in addition to the gRNA), wherein the sequence of some or all of the blocking domain is at least partially complementary to some or all of the gRNA. Therefore, the blocking domain is hybridizable or substantially hybridizable to some or all of the gRNA. In some embodiments, the blocking domain and the inducible activity gRNA are positioned on the template nucleic acid, e.g., template RNA, so that the gRNA can have a first conformation in which the blocking domain is hybridized or substantially hybridized to the gRNA, and a second conformation in which the blocking domain is not hybridized or substantially hybridized to the gRNA. In some embodiments, in the first conformation, the gRNA is unable to bind to the Gene Writer polypeptide (e.g., a template nucleic acid-binding domain, a DNA-binding domain, or an endonuclease domain (e.g., a CRISPR / Cas protein)) or binds with substantially reduced affinity compared to a similar template RNA lacking other blocking domains. In some embodiments, in the second conformation, the gRNA is able to bind to the Gene Writer polypeptide (e.g., a template nucleic acid-binding domain, a DNA-binding domain, or an endonuclease domain (e.g., a CRISPR / Cas protein)). In some embodiments, whether the gRNA is in the first or second conformation may affect whether the DNA-binding or endonuclease activity of the Gene Writer polypeptide (e.g., the CRISPR / Cas protein contained in the Gene Writer polypeptide) is active. In some embodiments, hybridization of the gRNA to the blocking domain can be disrupted using an opening molecule. In some embodiments, the opening molecule includes a drug that binds to part or all of the gRNA or blocking domain and inhibits hybridization of the gRNA to the blocking domain.In some embodiments, the opening molecule comprises a nucleic acid, for example, a gRNA, a blocking domain, or a sequence that is partially or entirely complementary to both. By selecting or designing an appropriate opening molecule, the provision of an opening molecule may facilitate a change in the three-dimensional structure of the gRNA so that it can associate with the CRISPR / Cas protein and provide the relevant function of the CRISPR / Cas protein (e.g., DNA binding and / or endonuclease activity). While not intended to be bound by any particular theory, the provision of an opening molecule at a selected time and / or location may enable spatial and temporal control of the activity of the gRNA, the CRISPR / Cas protein, or the Gene Writer system containing them. In some embodiments, the Gene Writer may comprise a Cas protein or a functional fragment thereof, as listed in Table 9 or Table 37, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it. [Table 43]
[0712] [Table 44]
[0713] [Table 45]
[0714] [Table 46]
[0715] [Table 47]
[0716] [Table 48]
[0717] [Table 49]
[0718] [Table 50]
[0719] [Table 51]
[0720] [Table 52]
[0721] [Table 53]
[0722] Table 9B defines the components required to design gRNA and / or template RNA and provides parameters for applying the Cas variants listed in Table 3A for Gene Writing. Tier indicates preferred Cas variants if available for use at a given locus. The cleavage site indicates the requirements of the validated or predicted protospacer adjacent motif (PAM) and the location of the validated or predicted cleavage site (relative to the upstream base of the PAM site). A gRNA for a given enzyme can be constructed by ligating crRNA, tetraloop, and tracrRNA sequences and further adding 5' spacers of a length within the minimum and maximum spacers that match the protospacer at the target site. Furthermore, the predicted location of the ssDNA nick at the target is important for designing the 3' region of the template RNA that needs to anneal to the sequence immediately 5' to the nick in order to initiate target prime reverse transcription.
[0723] [Table 54]
[0724] Table 55
[0725] Table 56
[0726] Table 57
[0727] Table 58
[0728] In some embodiments, the opening molecule is exogenous to cells containing the Gene Writer polypeptide and / or template nucleic acid. In some embodiments, the opening molecule comprises an endogenous agent (e.g., endogenous to cells containing the Gene Writer polypeptide and / or template nucleic acid including gRNA and a blocking domain). For example, the inducible gRNA, blocking domain, and opening molecule may be selected such that the opening molecule is an endogenous agent expressed in target cells or tissues, for example, thereby confirming the activity of the Gene Writer system in the target cells or tissues. As a further example, the inducible gRNA, blocking domain, and opening molecule may be selected such that the opening molecule is absent or substantially not expressed in one or more non-target cells or tissues, for example, thereby confirming that the activity of the Gene Writer system is absent or substantially absent, or present at a reduced level compared to the target cells or tissues. Exemplary blocking domains, opening molecules, and their uses are described in the PCT application publication, International Publication No. 2020044039A1 (which is incorporated herein by reference in its entirety). In some embodiments, the template nucleic acid, e.g., template RNA, may comprise one or more UTRs (e.g., from an R2-type retrotransposon) and gRNAs. In some embodiments, the UTR facilitates the interaction between the template nucleic acid (e.g., template RNA) and the writing domain of the Gene Writer polypeptide, e.g., the reverse transcriptase domain. In some embodiments, the gRNA facilitates the interaction between the polypeptide and the template nucleic acid binding domain (e.g., the RNA binding domain). In some embodiments, the gRNA directs the polypeptide to a matching target sequence in, for example, the target cell genome. In some embodiments, the template nucleic acid may comprise only the reverse transcriptase binding motif (e.g., the 3'UTR from R2), and the gRNA may be provided as a second nucleic acid molecule (e.g., a second RNA molecule) for target site recognition.In some embodiments, the template nucleic acid containing the RT-binding motif may reside on the same molecule as the gRNA, but it may be processed into two RNA molecules by cleavage activity (e.g., a ribozyme).
[0729] In some embodiments, the template RNA may be customized to correct a given mutation in the genomic DNA of a target cell (e.g., ex vivo or in vivo, e.g., within a subject, e.g., within a target tissue or organ). For example, the mutation may be a disease-associated mutation relative to the wild-type sequence. While not intended to be bound by any particular theory, an empirical set of parameters can be helpful in determining the optimal initial parameters in the in silico design of the template RNA or a portion thereof. As a non-limiting example, the following design parameters may be used for a selected mutation. In some embodiments, the design is initiated by obtaining an adjacent sequence of approximately 500 bp (e.g., up to 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, or 700 bp and optionally at least 20, 30, 40, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, or 650 bp) on either side of the mutation to function as a target region. In some embodiments, the template nucleic acid comprises a gRNA. Methods for designing gRNAs are known to those skilled in the art. In some embodiments, the gRNA comprises a sequence that binds to a target site (e.g., a CRISPR spacer). In some embodiments, sequences that bind to a target site (e.g., CRISPR spacers) for use in targeting a target region of a template nucleic acid are selected by considering the use of a specific Gene Writer polypeptide (e.g., including an endonuclease domain or writing domain, e.g., a CRISPR / Cas domain) (e.g., in the case of Cas9, the protospacer adjacent motif (PAM) of NGG immediately 3' to the 20nt gRNA binding region). In some embodiments, CRISPR spacers are first selected by ordering whether or not the PAM will be disrupted by Gene Writing-induced editing. In some embodiments, disruption of the PAM may increase editing efficiency.In some embodiments, PAM can also be disrupted during Gene Writing by introducing silent mutations (e.g., mutations that do not alter any amino acid residues encoded by the target nucleic acid sequence) at the target site (e.g., as part of or in addition to another modification to the target site in genomic DNA). In some embodiments, CRISPR spacers are selected by ordering the sequences by proximity of genomic sites corresponding to their desired editing positions. In some embodiments, the gRNA includes a gRNA scaffold. In some embodiments, the gRNA scaffold used may be a standard scaffold (e.g., for Cas9, 5'-GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCC-3') or may include one or more nucleotide substitutions. In some embodiments, the heterologous target sequence has at least 90% identity, e.g., at least 90%, 95%, 98%, 99%, or 100% identi...
Claims
1. A system for modifying DNA, a) A polypeptide or nucleic acid encoding the polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain and (ii) an endonuclease domain, and the polypeptide comprises an amino acid sequence having at least 95% identity with the sequence of Sequence ID No. 1988; b) (i) A 5'UTR sequence bound to the polypeptide, (ii) A 3'UTR sequence bound to the polypeptide, and (ii) Heterogeneous sequencing A template RNA containing or a nucleic acid encoding the template RNA A system that includes this.
2. The system according to claim 1, wherein the 5'UTR includes the sequence of sequence number 1986 or a sequence having at least 95% identity thereto.
3. The system according to claim 1 or 2, wherein the 3'UTR includes the sequence of sequence number 1987 or a sequence having at least 95% identity thereto.
4. A system for modifying DNA, (a) a polypeptide or nucleic acid encoding the polypeptide, wherein the polypeptide comprises (i) a reverse transcriptase domain and (ii) an endonuclease domain, and the polypeptide comprises an amino acid sequence having at least 95% identity with the sequence of Sequence ID No. 1947; (b) (i) A 5'UTR sequence bound to the polypeptide, (ii) A 3'UTR sequence bound to the polypeptide, and (iii) Heterogeneous sequencing template RNA containing and A system that includes this.
5. The system according to claim 4, wherein the 5'UTR includes the sequence of sequence number 1945 or a sequence having at least 95% identity thereto.
6. The system according to claim 4, wherein the 5'UTR includes the sequence of sequence number 1945.
7. The system according to any one of claims 4 to 6, wherein the 3'UTR includes the sequence of sequence number 1946 or a sequence having at least 95% identity therewith.
8. The system according to any one of claims 4 to 6, wherein the 3'UTR includes the sequence of sequence number 1946.
9. The system according to any one of claims 1 to 8, wherein the heterogeneous target sequence encodes a therapeutic polypeptide or a human polypeptide or a fragment or variant thereof.
10. The system according to any one of claims 1 to 8, wherein the heterogeneous target sequence encodes a chimeric antigen receptor (CAR).
11. The system according to any one of claims 1 to 8, wherein the heterogeneous purpose sequence includes a control sequence.
12. The aforementioned control array, (a) The promoter (b) enhancer, (c) The binding site of the endogenous regulatory component, (d) It is a miRNA binding site, or (e) Modifying the expression of endogenous genes or non-coding RNAs, The system according to claim 11.
13. The system according to any one of claims 1 to 12, wherein the polypeptide comprises a nuclear localization signal (NLS).
14. The system according to claim 13, wherein the NLS is fused to the N-terminus of the polypeptide.
15. The system according to claim 13, wherein the NLS is fused to the C-terminus of the polypeptide.
16. The system according to claim 13, wherein the NLS sequence comprises the amino acid sequence of PKKKRKV (SEQ ID NO: 2409).
17. The system according to any one of claims 1 to 16, wherein the polypeptide comprises a linker having the amino acid sequence SGSETPGTSEATPES (SEQ ID NO: 1023).
18. The system according to any one of claims 1 to 16, wherein the polypeptide comprises a linker having the amino acid sequence GGGS (SEQ ID NO: 1024).
19. The system according to claim 17 or 18, wherein the polypeptide further comprises an NLS, and the linker is disposed between the NLS and the remainder of the polypeptide.
20. The aforementioned template RNA is (a) Poly A portion; (b) Regulatory elements of woodchuck hepatitis virus (WPRE); and / or (c) Kozak sequence The system according to any one of claims 1 to 19, including the system described in any one of claims 1 to 19.
21. (i) the nucleic acid that encodes the polypeptide and (ii) The template RNA or the nucleic acid encoding the template RNA The system according to any one of claims 1 to 20, wherein the nucleic acids are separate nucleic acids.
22. The system according to any one of claims 1 to 21, wherein (a) comprises RNA encoding the polypeptide and (b) comprises template RNA.
23. The system according to any one of claims 1 to 22, wherein the nucleic acid encoding the polypeptide comprises a coding sequence that is codon-optimized for expression in human cells.
24. The system according to any one of claims 1 to 23, which can induce the addition, deletion, or modification of protein-coding sequences to the genome of mammalian cells.
25. The system according to any one of claims 1 to 24, which can induce the addition, deletion, or modification of non-coding RNA to the genome of a mammalian cell.
26. The system according to claim 24 or 25, wherein the insertion has a length of at least 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides.
27. The aforementioned mammalian cells, (a) Human cells; (b) Primary cells; and / or (b) T cells The system according to any one of claims 24 to 26.
28. An in vitro or ex vivo method for modifying a target DNA strand in a cell or tissue, comprising administering to the cell or tissue a system according to any one of claims 1 to 27, wherein the system reverse transcribes the template RNA sequence into the target DNA strand, thereby modifying the target DNA strand, wherein human germ cells are excluded from the cells.
29. Lipid nanoparticles (LNPs) comprising the system according to any one of claims 1 to 27.
30. A system according to any one of claims 1 to 27 for use in a method for modifying a target DNA strand in a cell, tissue, or object.