Recombinase compositions and methods of use

The use of recombinase polypeptides with specific DNA recognition sequences allows for precise and efficient insertion of genetic elements into host genomes, addressing limitations in current methods.

JP2026012799APending Publication Date: 2026-01-27FLAGSHIP PIONEERING INNOVATIONS VI LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025175403
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-08-21
Filing Date
2025-10-17
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Current methods for introducing exogenous genetic elements into a host genome are limited in precision and efficiency, particularly in vivo or in vitro applications.

Method used

A system utilizing recombinase polypeptides with specific amino acid sequences and DNA recognition sequences, including parapalindromic sequences and core sequences, to facilitate targeted insertion of heterologous sequences into host genomes.

Benefits of technology

Enables precise and efficient modification of host genomes by ensuring high sequence identity and compatibility between recombinase polypeptides and DNA recognition sequences, enhancing the frequency and accuracy of genetic element introduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012799000790
    Figure 2026012799000790
  • Figure 2026012799000791
    Figure 2026012799000791
  • Figure 2026012799000792
    Figure 2026012799000792
Patent Text Reader

Abstract

To provide methods and compositions for modulating a target genome.SOLUTION: The present disclosure relates to compositions, systems, and methods for targeting, editing, modifying, or engineering a DNA sequence (e.g., inserting a heterologous DNA sequence of interest into a target site of a mammalian genome) at one or more locations within a DNA sequence in a cell, tissue, or subject, e.g., in vivo or in vitro. The DNA sequence of interest may comprise, for example, a coding sequence, a regulatory sequence, a gene expression unit.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to novel compositions, systems, and methods for modifying the genome of one or more locations in a host cell, tissue, or subject, either in vivo or in vitro. In particular, the present invention features compositions, systems, and methods for introducing exogenous genetic elements into a host genome using recombinase polypeptides (e.g., serine recombinases, as described herein). Summary of the Invention [Means for solving the problem]

[0002] 1. A system for modifying DNA, comprising: a) a recombinase polypeptide or a nucleic acid encoding a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and b) a double-stranded insert DNA, (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), comprising a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the first and second parapalindromic sequences together comprising a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and the DNA recognition sequence further comprising a core sequence of about 2 to 20 nucleotides, the core sequence being located between a first and a second parapalindromic sequence; (ii) a heterologous sequence of interest; double-stranded intercalated DNA containing A system including:

[0003] 2. A system for modifying DNA, comprising: a) a recombinase polypeptide or a nucleic acid encoding a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and b) an insert DNA, (i) a human first parapalindromic sequence and a human second parapalindromic sequence that bind to the recombinase polypeptide of (a), each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the first and second parapalindromic sequences together comprising a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and The DNA recognition sequence further comprises a core sequence of about 2 to 20 nucleotides, the core sequence comprising a first human parapalindromic sequence and a second human parapalindromic sequence located between the first and second parapalindromic sequences; (ii) optionally, a heterologous sequence of interest; Insert DNA containing A system including:

[0004] 2a. A system for modifying DNA, comprising: a) a recombinase polypeptide or a nucleic acid encoding a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and b) a double-stranded insert DNA, (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), optionally comprising about 30-70 or 40-60 nucleotides of a sequence present within the nucleotide sequence of the left-hand region or right-hand region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; (ii) a heterologous sequence of interest; double-stranded intercalated DNA containing A system including:

[0005] 3. The system of embodiment 1 or 2, wherein the recombinase polypeptide comprises an amino acid sequence having at least 70% sequence identity to an amino acid sequence of Table 3A, 3B, or 3C.

[0006] 4. The system of embodiment 1 or 2, wherein the recombinase polypeptide comprises an amino acid sequence having at least 75% sequence identity to an amino acid sequence of Table 3A, 3B, or 3C.

[0007] 5. The system of embodiment 1 or 2, wherein the recombinase polypeptide comprises an amino acid sequence having at least 80% sequence identity to an amino acid sequence of Table 3A, 3B, or 3C.

[0008] 6. The system of embodiment 1 or 2, wherein the recombinase polypeptide comprises an amino acid sequence having at least 85% sequence identity to an amino acid sequence of Table 3A, 3B, or 3C.

[0009] 7. The system of embodiment 1 or 2, wherein the recombinase polypeptide comprises an amino acid sequence having at least 90% sequence identity to an amino acid sequence in Table 3A, 3B, or 3C.

[0010] 8. The system of embodiment 1 or 2, wherein the recombinase polypeptide comprises an amino acid sequence having at least 95% sequence identity to an amino acid sequence in Table 3A, 3B, or 3C.

[0011] 9. The system of embodiment 1 or 2, wherein the recombinase polypeptide comprises an amino acid sequence having at least 96% sequence identity to an amino acid sequence in Table 3A, 3B, or 3C.

[0012] 10. The system of embodiment 1 or 2, wherein the recombinase polypeptide comprises an amino acid sequence having at least 97% sequence identity to an amino acid sequence in Table 3A, 3B, or 3C.

[0013] 11. The system of embodiment 1 or 2, wherein the recombinase polypeptide comprises an amino acid sequence having at least 98% sequence identity to an amino acid sequence of Table 3A, 3B, or 3C.

[0014] 12. The system of embodiment 1 or 2, wherein the recombinase polypeptide comprises an amino acid sequence having at least 99% sequence identity to an amino acid sequence of Table 3A, 3B, or 3C.

[0015] 13. The system of embodiment 1 or 2, wherein the recombinase polypeptide comprises an amino acid sequence having 100% sequence identity to an amino acid sequence in Table 3A, 3B, or 3C.

[0016] 14. A system described in any one of embodiments 1 to 13, wherein (a) and (b) are in separate containers.

[0017] 15. The system of any one of embodiments 1 to 13, wherein (a) and (b) are mixed.

[0018] 15a.(b) is a system according to any one of the preceding embodiments, comprising linear double-stranded DNA.

[0019] 15b. The system of any one of embodiments 1 to 15, wherein (b) comprises circular double-stranded DNA.

[0020] 15c.(b) is (iii) a second DNA recognition sequence that binds to the recombinase polypeptide of (a), the second DNA recognition sequence having a third parapalindromic sequence and a fourth parapalindromic sequence, each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the third and fourth parapalindromic sequences together comprising a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and The second DNA recognition sequence further comprises a core sequence of about 2 to 20 nucleotides, the core sequence being located between the third and fourth parapalindromic sequences. The system of embodiment 15a, comprising:

[0021] 15d-a. The system of embodiment 15c, wherein the first DNA recognition sequence has the same sequence as the second DNA recognition sequence.

[0022] 15d-b. The system of embodiment 15c, wherein the first DNA recognition sequence does not have the same sequence as the second DNA recognition sequence (e.g., the second DNA recognition sequence includes at least one substitution, deletion, or insertion relative to the first DNA recognition sequence).

[0023] 15d1. The system of embodiment 15d-b, wherein the first DNA recognition sequence has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to the second DNA recognition sequence.

[0024] 15e. The system of any of embodiments 15c-15d1, wherein the heterologous sequence of interest is located between the first DNA recognition sequence and the second DNA recognition sequence.

[0025] 15f. A first circular RNA encoding a polypeptide of the Gene Writing System; The second circular RNA containing the template nucleic acid of the Gene Writing System A system including:

[0026] 15g. A system for modifying DNA, comprising: (a) a polypeptide or a nucleic acid encoding a polypeptide, the polypeptide comprising (i) a reverse transcriptase domain, and (ii) an endonuclease domain; and (b) a template nucleic acid comprising (i) a sequence that binds to a polypeptide, (ii) a heterologous sequence of interest, and (iii) a ribozyme that is heterologous to (a)(i), (a)(ii), (b)(i), or a combination thereof; A system including:

[0027] 15h. The system of embodiment 15g, wherein the ribozyme is heterologous to (b)(i).

[0028] 15i. The system of embodiment 15g or 15h, wherein the template nucleic acid (iv) comprises a second ribozyme, e.g., internal to (a)(i), (a)(ii), (b)(i), or a combination thereof, e.g., the second ribozyme is internal to (b)(i).

[0029] 15j. The system of embodiment 15g or 15h, wherein the heterologous ribozyme replaces a ribozyme inherent in (a)(i), (a)(ii), (b)(i), or a combination thereof, e.g., a second ribozyme is inherent in (b)(i).

[0030] The system of any of embodiments 15f to 15j, further comprising an mRNA encoding a polypeptide of the 15k. Gene Writing System.

[0031] 151. A system according to any of embodiments 15f to 15k, further comprising DNA encoding a polypeptide of the Gene Writing system.

[0032] The system of any of embodiments 15f-15l, further comprising DNA comprising the insert DNA of the 15m. Gene Writing System.

[0033] 15n. The system according to any of embodiments 15f to 15m, further comprising an insert DNA of the Gene Writing System and a DNA comprising the polypeptide.

[0034] 16. A cell (e.g., a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell; or a prokaryotic cell) comprising a recombinase polypeptide or a nucleic acid encoding a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.

[0035] 16a. A cell comprising the system of any of embodiments 1 to 15e.

[0036] 17. (i) a DNA recognition sequence that binds to a recombinase polypeptide, the DNA recognition sequence having a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the first and second parapalindromic sequences together comprise a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and the DNA recognition sequence further comprising a core sequence of about 2 to 20 nucleotides, the core sequence being located between a first and a second parapalindromic sequence; (ii) optionally, a heterologous sequence of interest; 17. The cell of embodiment 16, further comprising an insert DNA comprising:

[0037] 17a.(i)(a)'s recombinase polypeptide binds to a DNA recognition sequence, optionally comprising about 30-70 or 40-60 nucleotides of a sequence present within the left-hand region or right-hand region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; (ii) optionally, a heterologous sequence of interest; 17. The cell of embodiment 16, further comprising an insert DNA comprising:

[0038] 18. A cell (e.g., a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell; or a prokaryotic cell), (i) a DNA recognition sequence having a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the first and second parapalindromic sequences together comprise a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and the DNA recognition sequence further comprising a core sequence of about 2 to 20 nucleotides, the core sequence being located between a first and a second parapalindromic sequence; (ii) a heterologous sequence of interest; a cell comprising the vector (e.g., a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell; or a prokaryotic cell).

[0039] 18a. On a chromosome, (i) a first parapalindromic sequence of about 15 to 35 or 20 to 30 nucleotides, wherein the first parapalindromic sequence is present within a nucleotide sequence of the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence that has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic sequence, or no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; (ii) a second parapalindromic sequence of about 15 to 35 or 20 to 30 nucleotides, wherein the second parapalindromic sequence is present within a nucleotide sequence of the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence that has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic sequence, or no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; (iii) a heterologous sequence of interest located between (i) and (ii); a cell comprising the vector (e.g., a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell; or a prokaryotic cell).

[0040] 19a. The cell of embodiment 18, wherein both the DNA recognition sequence and the heterologous sequence of interest are located on an extrachromosomal nucleic acid.

[0041] 19. The cell of any of embodiments 18 or 19a, wherein the DNA recognition sequence is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90 or 100 nucleotides of the heterologous sequence of interest.

[0042] 19c. Extrachromosomal nucleic acids are (iii) a second DNA recognition sequence, comprising a third parapalindromic sequence and a fourth parapalindromic sequence, each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the third and fourth parapalindromic sequences together comprising a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and The second DNA recognition sequence further comprises a core sequence of about 2 to 20 nucleotides, the core sequence being located between the third and fourth parapalindromic sequences. 20. The cell of any of embodiments 19a or 19, comprising:

[0043] 19c1. The cell of embodiment 19c, wherein the first DNA recognition sequence has the same sequence as the second DNA recognition sequence.

[0044] 19c2. The cell of embodiment 19c, wherein the first DNA recognition sequence does not have the same sequence as the second DNA recognition sequence (e.g., the second DNA recognition sequence comprises at least one substitution, deletion, or insertion relative to the first DNA recognition sequence).

[0045] 19c3. The cell of embodiment 19c2, wherein the first DNA recognition sequence has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to the second DNA recognition sequence.

[0046] 19c4. The cell of any of embodiments 19c to 19c3, wherein the extrachromosomal nucleic acid is linear.

[0047] 19c5.(iv) a third DNA recognition sequence, comprising a fifth parapalindromic sequence and a sixth parapalindromic sequence, each parapalindromic sequence being about 15-35 or 20-30 nucleotides, and the fifth and sixth parapalindromic sequences together comprising a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and the third DNA recognition sequence further comprises a core sequence of about 2 to 20 nucleotides, the core sequence being located between the fifth and sixth parapalindromic sequences; The third DNA recognition sequence is a third DNA recognition sequence located on the chromosome. The cell of any of embodiments 19c to 19c4, comprising:

[0048] 19c6. The cell of embodiment 19c5, wherein the third DNA recognition sequence does not have the same sequence as the first DNA recognition sequence, the second DNA recognition sequence, or both the first and second DNA recognition sequences (e.g., the third DNA recognition sequence comprises at least one substitution, deletion, or insertion relative to the first and / or second DNA recognition sequence).

[0049] 19c7. The cell of embodiment 19c6, wherein the third DNA recognition sequence has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to the first DNA recognition sequence.

[0050] 19c8. The cell of any of embodiments 19c6 or 19c7, wherein the third DNA recognition sequence has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to the second DNA recognition sequence.

[0051] 19c9.(v) a fourth DNA recognition sequence, comprising a seventh parapalindromic sequence and an eighth parapalindromic sequence, each parapalindromic sequence being about 15-35 or 20-30 nucleotides, and the seventh and eighth parapalindromic sequences together comprising a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) relative to said parapalindromic region; and the fourth DNA recognition sequence further comprises a core sequence of about 2 to 20 nucleotides, the core sequence being located between the seventh and eighth parapalindromic sequences; The fourth DNA recognition sequence is a fourth DNA recognition sequence located on the same chromosome as the third DNA recognition sequence. The cell of any of embodiments 19c5 to 19c8, comprising:

[0052] 19c10. The cell of embodiment 19c9, wherein the fourth DNA recognition sequence does not have the same sequence as the first DNA recognition sequence, the second DNA recognition sequence, or both the first and second DNA recognition sequences (e.g., the fourth DNA recognition sequence comprises at least one substitution, deletion, or insertion relative to the first and / or second DNA recognition sequence).

[0053] 19c11. The cell of embodiment 19c10, wherein the fourth DNA recognition sequence has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to the first DNA recognition sequence.

[0054] 19c12. The cell of any of embodiments 19c10 or 19c11, wherein the fourth DNA recognition sequence has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to the second DNA recognition sequence.

[0055] 19c13. The cell of any of embodiments 19c9 to 19c12, wherein the fourth DNA recognition sequence has the same sequence as the third DNA recognition sequence.

[0056] 19c14. The cell of embodiment 19c13, wherein the fourth DNA recognition sequence does not have the same sequence as the fourth DNA recognition sequence (e.g., the fourth DNA recognition sequence comprises at least one substitution, deletion, or insertion relative to the third DNA recognition sequence).

[0057] 19c15. The cell of embodiment 19c14, wherein the fourth DNA recognition sequence has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to the third DNA recognition sequence.

[0058] 19c16. The cell of any of embodiments 19c10 to 19c15, wherein the third DNA recognition sequence and the fourth DNA recognition sequence are within 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, or 900 bases of each other, or within 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 kilobases of each other on the chromosome.

[0059] 20. The cell according to any of embodiments 16a to 18, wherein the DNA recognition sequence is located within a chromosome and the heterologous sequence of interest is located on an extrachromosomal nucleic acid.

[0060] 21. A cell according to any one of embodiments 16 to 20, which is a eukaryotic cell.

[0061] 22. The cell of embodiment 21, which is a mammalian cell.

[0062] 23. The cell of embodiment 22, which is a human cell.

[0063] 24. A cell according to any one of embodiments 16 to 20, which is a prokaryotic cell (e.g., a bacterial cell).

[0064] 26. The isolated eukaryotic cell of embodiment 25, which is an animal cell (e.g., a mammalian cell) or a plant cell.

[0065] 27. The isolated eukaryotic cell of embodiment 26, wherein the mammalian cell is a human cell.

[0066] 28. The isolated eukaryotic cell of embodiment 26, wherein the animal cell is a bovine cell, equine cell, porcine cell, caprine cell, ovine cell, chicken cell or turkey cell.

[0067] 29. The isolated eukaryotic cell of embodiment 26, wherein the plant cell is a corn cell, a soybean cell, a wheat cell, or a rice cell.

[0068] 30. A method for modifying the genome of a eukaryotic cell (e.g., a mammalian cell, e.g., a human cell), comprising: a) a recombinase polypeptide or a nucleic acid encoding a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and b) an insert DNA, (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), comprising a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the first and second parapalindromic sequences together comprising a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and the DNA recognition sequence further comprising a core sequence of about 2 to 20 nucleotides, the core sequence being located between a first and a second parapalindromic sequence; (ii) a heterologous sequence of interest; Insert DNA containing thereby modifying the genome of a eukaryotic cell.

[0069] 30a. A method for modifying the genome of a eukaryotic cell (e.g., a mammalian cell, e.g., a human cell), comprising: a) a recombinase polypeptide or a nucleic acid encoding a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and b) an insert DNA, (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), optionally comprising a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to about 30-70 or 40-60 nucleotides of a sequence present within the nucleotide sequence of the left-hand region or right-hand region column of Table 2A, 2B, or 2C, or said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and the DNA recognition sequence further comprising a core sequence of about 2 to 20 nucleotides, the core sequence being located between a first and a second parapalindromic sequence; (ii) a heterologous sequence of interest; Insert DNA containing thereby modifying the genome of a eukaryotic cell.

[0070] 31. A method for inserting a heterologous sequence of interest into the genome of a eukaryotic cell (e.g., a mammalian cell, e.g., a human cell), comprising: a) a recombinase polypeptide or a nucleic acid encoding a polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and b) an insert DNA, (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), comprising a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the first and second parapalindromic sequences together comprising a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and the DNA recognition sequence further comprising a core sequence of about 2 to 20 nucleotides, the core sequence being located between a first and a second parapalindromic sequence; (ii) a heterologous sequence of interest; Insert DNA containing , thereby inserting a heterologous sequence of interest into the genome of the eukaryotic cells at a frequency of, e.g., at least about 0.1% (e.g., at least about 0.1%, 0.5%, 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%) of a population of eukaryotic cells, e.g., as measured in the assay of Example 5.

[0071] 31a. A method for inserting a heterologous sequence of interest into the genome of a eukaryotic cell (e.g., a mammalian cell, e.g., a human cell), comprising: a) a recombinase polypeptide or a nucleic acid encoding a polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and b) an insert DNA, (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), optionally comprising about 30-70 or 40-60 nucleotides of a sequence present within the nucleotide sequence of the left-hand region or right-hand region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; (ii) a heterologous sequence of interest; Insert DNA containing , thereby inserting a heterologous sequence of interest into the genome of the eukaryotic cells at a frequency of, e.g., at least about 0.1% (e.g., at least about 0.1%, 0.5%, 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%) of a population of eukaryotic cells, e.g., as measured in the assay of Example 5.

[0072] 32. The method of any of embodiments 30-31a, wherein (a) and (b) are administered separately or together.

[0073] 33. The method of any one of embodiments 30-31a, wherein (a) is administered before, simultaneously with, or after administration of (b).

[0074] 34. The method of any of embodiments 30-33, wherein (a) comprises a nucleic acid encoding said polypeptide.

[0075] 35. The method of embodiment 34, wherein the nucleic acid of (a) and the insert DNA of (b) are located on the same nucleic acid molecule, for example, on the same vector.

[0076] 36. The method of embodiment 34, wherein the nucleic acid of (a) and the insert DNA of (b) are located on separate nucleic acid molecules.

[0077] 37. The method of any of embodiments 30 to 36, wherein the cells have only one endogenous DNA recognition sequence that is compatible with the DNA recognition sequence of the inserted DNA.

[0078] 38. The method of any of embodiments 30 to 36, wherein the cell has two or more endogenous DNA recognition sequences that are compatible with the DNA recognition sequences of the inserted DNA.

[0079] 38a. The insert DNA of (b) contains a second DNA recognition sequence that binds to the recombinase polypeptide of (a); the second DNA recognition sequence has a third parapalindromic sequence and a fourth parapalindromic sequence, each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the third and fourth parapalindromic sequences together comprise a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and 39. The method of any one of embodiments 30 to 38, wherein the second DNA recognition sequence further comprises a core sequence of about 2 to 20 nucleotides, the core sequence being located between the third and fourth parapalindromic sequences.

[0080] 38b. The method of embodiment 38a, wherein the first DNA recognition sequence has the same sequence as the second DNA recognition sequence.

[0081] 38c. The method of embodiment 38a, wherein the first DNA recognition sequence does not have the same sequence as the second DNA recognition sequence (e.g., the second DNA recognition sequence comprises at least one substitution, deletion, or insertion relative to the first DNA recognition sequence).

[0082] 38d. The method of embodiment 38c, wherein the first DNA recognition sequence has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to the second DNA recognition sequence.

[0083] 38e. The method of any of embodiments 38a-38d, wherein the heterologous sequence of interest is located between the first DNA recognition sequence and the second DNA recognition sequence.

[0084] 38f. The method of any preceding embodiment, wherein the recombinase polypeptide comprises an integrase, e.g., listed in Table 30 or FIG. 1A.

[0085] 38g. The method of embodiment 38f, wherein the recombinase polypeptide comprises an integrase listed in Table 30, and the DNA recognition sequence comprises a recognition sequence from the corresponding row number of Table 2A, 2B, or 2C.

[0086] 38h. The method of embodiment 38f or 38g, wherein the recombinase polypeptide comprises the amino acid sequence of Int101 (e.g., the sequence of the corresponding amino acid sequence listed in Table 3A, 3B, or 3C, e.g., corresponding to line 475 or accession ASN71805.1), and optionally, the DNA recognition sequence comprises a recognition sequence from the corresponding line number of Table 2A, 2B, or 2C (e.g., listed on line 475).

[0087] 38i. The method of embodiment 38f or 38g, wherein the recombinase polypeptide comprises the amino acid sequence of Int78 (e.g., the sequence of the corresponding amino acid sequence listed in Table 3A, 3B, or 3C, e.g., corresponding to line 371 or accession ARW58518.1), and optionally, the DNA recognition sequence comprises a recognition sequence from the corresponding line number of Table 2A, 2B, or 2C (e.g., listed on line 371).

[0088] 38j. The method of embodiment 38f or 38g, wherein the recombinase polypeptide comprises the amino acid sequence of Int79 (e.g., the sequence of the corresponding amino acid sequence listed in Table 3A, 3B, or 3C, e.g., corresponding to line 360 ​​or accession ARW58461.1), and optionally, the DNA recognition sequence comprises a recognition sequence from the corresponding line number of Table 2A, 2B, or 2C (e.g., listed on line 360).

[0089] 38k. The method of embodiment 38f or 38g, wherein the recombinase polypeptide comprises the amino acid sequence of Int30 (e.g., the sequence of the corresponding amino acid sequence listed in Table 3A, 3B, or 3C, e.g., corresponding to line 436 or accession YP_009103095.1), and optionally the DNA recognition sequence comprises a recognition sequence from the corresponding line number of Table 2A, 2B, or 2C (e.g., listed on line 436).

[0090] 38l. The method of embodiment 38f or 38g, wherein the recombinase polypeptide comprises the amino acid sequence of Int3 (e.g., the sequence of the corresponding amino acid sequence listed in Table 3A, 3B, or 3C, e.g., corresponding to line 1200 or accession YP_459991.1), and optionally, the DNA recognition sequence comprises a recognition sequence from the corresponding line number of Table 2A, 2B, or 2C (e.g., listed in line 1200).

[0091] 38m. The method of embodiment 38f or 38g, wherein the recombinase polypeptide comprises the amino acid sequence of Int38 (e.g., the sequence of the corresponding amino acid sequence listed in Table 3A, 3B, or 3C, e.g., corresponding to line 408 or accession YP_009223181.1), and optionally the DNA recognition sequence comprises a recognition sequence from the corresponding line number of Table 2A, 2B, or 2C (e.g., listed on line 408).

[0092] 38n. The method of embodiment 38f or 38g, wherein the recombinase polypeptide comprises the amino acid sequence of Int95 (e.g., the sequence of the corresponding amino acid sequence listed in Table 3A, 3B, or 3C, e.g., corresponding to line 460 or accession AFV15398.1), and optionally, the DNA recognition sequence comprises a recognition sequence from the corresponding line number of Table 2A, 2B, or 2C (e.g., listed in line 460).

[0093] 38o. The method of embodiment 38f or 38g, wherein the recombinase polypeptide comprises the amino acid sequence of Int51 (e.g., the sequence of the corresponding amino acid sequence listed in Table 3A, 3B, or 3C, e.g., corresponding to line 159 or accession AOT24690.1), and optionally the DNA recognition sequence comprises a recognition sequence from the corresponding line number of Table 2A, 2B, or 2C (e.g., listed on line 159).

[0094] 38p. The method of embodiment 38f or 38g, wherein the recombinase polypeptide comprises the amino acid sequence of Int18 (e.g., the sequence of the corresponding amino acid sequence listed in Table 3A, 3B, or 3C, e.g., corresponding to line 103 or accession AGR47239.1), and optionally the DNA recognition sequence comprises a recognition sequence from the corresponding line number of Table 2A, 2B, or 2C (e.g., listed on line 103).

[0095] 39. An isolated recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.

[0096] 40. The isolated recombinase polypeptide of embodiment 39, comprising at least one insertion, deletion, or substitution relative to the recombinase sequence of Table 3A, 3B, or 3C.

[0097] 41. The isolated recombinase polypeptide of embodiment 40, which binds to a parapalindromic region present within a eukaryotic (e.g., mammalian, e.g., human) genomic locus (e.g., a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C), or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto.

[0098] 41a. The isolated recombinase polypeptide of any of embodiments 39 or 40, which binds to a parapalindromic region present within a nucleotide sequence in the Left Region or Right Region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto.

[0099] 42. The isolated recombinase polypeptide of any of embodiments 40-41a, having at least a 2-, 3-, 4-, or 5-fold increase in affinity for a genomic locus compared to the corresponding unmodified amino acid sequence of Table 3A, 3B, or 3C.

[0100] 43. An isolated nucleic acid encoding a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.

[0101] 44. The isolated nucleic acid of embodiment 43, encoding a recombinase polypeptide comprising at least one insertion, deletion, or substitution compared to a recombinase polypeptide of Table 3A, 3B, or 3C.

[0102] 45. The isolated nucleic acid sequence according to embodiment 43 or 44, wherein the codons of the amino acid sequence have been modified (e.g., optimized) for expression in mammalian cells, such as human cells.

[0103] 46. ​​The isolated nucleic acid of any of embodiments 43 to 45, further comprising a heterologous promoter (e.g., a mammalian promoter, e.g., a tissue-specific promoter), a microRNA (e.g., a tissue-specific restricted miRNA), a polyadenylation signal, or a heterologous payload.

[0104] 47. An isolated nucleic acid (e.g., DNA), (i) a DNA recognition sequence having a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the first and second parapalindromic sequences together comprise a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and the DNA recognition sequence further comprising a core sequence of about 2 to 20 nucleotides, the core sequence being located between a first and a second parapalindromic sequence; (ii) a heterologous sequence of interest; An isolated nucleic acid (e.g., DNA) comprising:

[0105] 47a. Isolated nucleic acid (e.g., DNA), (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), optionally comprising about 30-70 or 40-60 nucleotides of, or at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to, a sequence present within the nucleotide sequence of the left-hand region or right-hand region column of Table 2A, 2B, or 2C, or a nucleotide sequence having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) relative to said parapalindromic region; (ii) optionally, a heterologous sequence of interest; An isolated nucleic acid (e.g., DNA) comprising:

[0106] 48. The isolated nucleic acid of any of embodiments 47 or 47a, which binds to a recombinase polypeptide of Table 3A, 3B, or 3C.

[0107] 48a. The isolated nucleic acid of any of embodiments 47-48, wherein the DNA recognition sequence (e.g., one or more parapalindromic sequences) comprises at least one insertion, deletion, or substitution relative to a recognition sequence (or portion thereof) present in a sequence in Table 2A, 2B, or the left region or right region columns.

[0108] 48b. The isolated nucleic acid of embodiment 48a, wherein the DNA recognition sequence (e.g., parapalindromic region) has at least a 2-, 3-, 4-, or 5-fold increase in affinity for a recombinase polypeptide compared to the corresponding unmodified DNA recognition sequence (e.g., parapalindromic region).

[0109] 48c. The isolated nucleic acid of either embodiment 48a or 48b, wherein the recombinase polypeptide has at least a 2-, 3-, 4-, or 5-fold increase in recombinase activity at a DNA recognition sequence (e.g., a parapalindromic region) compared to a corresponding unmodified DNA recognition sequence (e.g., a parapalindromic region).

[0110] 49. A method for producing a recombinase polypeptide, comprising: a) providing a nucleic acid encoding a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and b) introducing the nucleic acid into a cell (e.g., a eukaryotic or prokaryotic cell, as described herein) under conditions that allow production of the recombinase polypeptide. thereby producing a recombinase polypeptide.

[0111] 50. A method for producing a recombinase polypeptide, comprising: a) providing a cell (e.g., a eukaryotic or prokaryotic cell) containing a nucleic acid encoding a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and b) incubating the cells under conditions that allow production of the recombinase polypeptide. thereby producing a recombinase polypeptide.

[0112] 51. A method for generating an insert DNA containing a DNA recognition sequence and a heterologous sequence, comprising: a) a nucleic acid, (i) a DNA recognition sequence that binds to a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, the DNA recognition sequence having a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the first and second parapalindromic sequences, together, are those of Table 2A, 2B, or comprises a parapalindromic region present within a nucleotide sequence in the left region or right region column of 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and the DNA recognition sequence further comprising a core sequence of about 2 to 20 nucleotides, the core sequence being located between a first and a second parapalindromic sequence; (ii) a heterologous sequence of interest; providing a nucleic acid comprising: b) introducing the nucleic acid into a cell (e.g., a eukaryotic or prokaryotic cell, as described herein) under conditions that allow replication of the nucleic acid; thereby generating an insert DNA.

[0113] 51a. Nucleic acids are (iii) a second DNA recognition sequence that binds to the recombinase polypeptide, the second DNA recognition sequence having a third parapalindromic sequence and a fourth parapalindromic sequence, each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the third and fourth parapalindromic sequences together comprising a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and The second DNA recognition sequence further comprises a core sequence of about 2 to 20 nucleotides, the core sequence being located between the third and fourth parapalindromic sequences. 52. The method of embodiment 51, comprising:

[0114] 51b. The method of embodiment 51a, wherein the first DNA recognition sequence has the same sequence as the second DNA recognition sequence.

[0115] 51c. The method of embodiment 51a, wherein the first DNA recognition sequence does not have the same sequence as the second DNA recognition sequence (e.g., the second DNA recognition sequence comprises at least one substitution, deletion, or insertion relative to the first DNA recognition sequence).

[0116] 51d. The method of embodiment 51c, wherein the first DNA recognition sequence has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to the second DNA recognition sequence.

[0117] 51e. The method of any of embodiments 51a-51d, wherein the heterologous sequence of interest is located between the first DNA recognition sequence and the second DNA recognition sequence.

[0118] 51f. The method of any of embodiments 51-51e, wherein the providing step comprises using cloning techniques (e.g., restriction digestion and / or ligation), using recombinant techniques, or obtaining the nucleic acid (e.g., from a third party provider).

[0119] 52. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the recombinase polypeptide comprises at least one insertion, deletion, or substitution compared to the amino acid sequence of Table 3A, 3B, or 3C.

[0120] 53. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the recombinase polypeptide comprises a truncation at the N-terminus, the C-terminus, or both the N- and C-terminus, compared to the amino acid sequence of Table 3A, 3B, or 3C.

[0121] 54. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any of the preceding embodiments, wherein the recombinase polypeptide comprises a nuclear localization sequence, e.g., an endogenous nuclear localization sequence or a heterologous nuclear localization sequence.

[0122] 55. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the heterologous sequence of interest is inserted into the genome of cells with an efficiency of at least about 0.1% (e.g., at least about 0.1%, 0.5%, 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%) of a population of cells, e.g., as measured in the assay of Example 5.

[0123] 56. The heterologous sequence of interest is a sequence present within or at least at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 100% of the nucleotide sequences in the genome of a cell in at least about 1% of insertion events (e.g., at least about 1%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 100%), as determined, for example, by the assay of Example 4. 95%, 96%, 97%, 98% or 99% identity thereto, or having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 or fewer sequence alterations (e.g., substitutions, insertions or deletions) thereto; and / or corresponding to the line number of the recombinase listed in Table 3A, 3B or 3C.

[0124] 57. In a population of cells (e.g., contacted with the system), a heterologous sequence of interest is present at 1 to 10, e.g., 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 2 to 10, 2 to 5, 2 to 4, 3 to 10, 3 to 5, or 5 to 10 sites (e.g., within the nucleotide sequence of the left region or right region column of Table 2A, 2B, or 2C) within the genome of cells in at least 1%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 100% of the cells in the population, as determined, e.g., by the assay of Example 5.

[0033] The system, cell, method, isolated recombinase polypeptide or isolated nucleic acid of any preceding embodiment is inserted into a site comprising a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and / or corresponding to the line number of the recombinase listed in Table 3A, 3B, or 3C.

[0125] 58. In a population of cells contacted with the system, the heterologous sequence of interest is present at exactly one site within the genome of the cells (e.g., a sequence present within or at least 70%, 75%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 100%) of the cells in the population, as determined, for example, by the assay of Example 4. , 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 sequence alterations (e.g., substitutions, insertions or deletions) thereto; and / or corresponding to the line number of the recombinase listed in Table 3A, 3B or 3C.

[0126] 59. The heterologous sequence of interest may be present at 1 to 10, e.g., 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 2 to 10, 2 to 5, 2 to 4, 3 to 10, 3 to 5, or 5 to 10 sites within the genome of the cell (e.g., a sequence present within the nucleotide sequence of the left region or right region column of Table 2A, 2B, or 2C, or at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto, e.g., as determined by the assay of Example 4).

[0023] The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the recombinase is inserted into a site comprising a nucleotide sequence having, or having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or fewer sequence modifications (e.g., substitutions, insertions, or deletions) thereto; and / or corresponding to a row of a recombinase listed in Table 3A, 3B, or 3C.

[0127] 60. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the recombinase polypeptide is linked to insert DNA.

[0128] 61. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any of the previous embodiments, wherein the recombinase polypeptide is provided by providing a nucleic acid encoding the recombinase polypeptide.

[0129] 62. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, resulting in an insertion frequency of a heterologous sequence of interest into the genome of at least about 0.1% (e.g., at least about 0.1%, 0.5%, 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%) of a population of cells, e.g., as measured in the assay of Example 5.

[0130] 62a. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any of the preceding embodiments, resulting in an insertion frequency of a heterologous sequence of interest into the genome of at least about 0.1% (e.g., at least about 0.1%, 0.5%, 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%) of a population of cells, e.g., as measured in the assay of Example 13.

[0131] 62b. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, resulting in an insertion frequency of a heterologous sequence of interest into the genome of at least about 0.1% (e.g., at least about 0.1%, 0.5%, 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%) of a population of cells, e.g., as measured in the assay of Example 7.

[0132] 63. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the first parapalindromic sequence comprises a first sequence of 15 to 35 or 20 to 30 nucleotides, e.g., 13, 14, 15, 16, 17, 18, 19 or 20, of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 or 35 nucleotides present in a sequence found in the Left Region or Right Region column of Table 2A, 2B, or 2C, or a sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 or fewer substitutions, insertions, or deletions thereto.

[0133] 64. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of embodiment 63, wherein the second parapalindromic sequence comprises a second sequence of 15 to 35 or 20 to 30 nucleotides, e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides present in a sequence found in the left region or right region column of Table 2A, 2B, or 2C, 13, 14, 15, 16, 17, 18, 19, or 20, or a sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or fewer substitutions, insertions, or deletions thereto.

[0134] 65. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the insert DNA further comprises a core sequence comprising about 2-20, e.g., 2-16, nucleotides located between the first and second parapalindromic sequences found in the Left Region or Right Region columns of Table 2A, 2B, or 2C, or a sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or fewer substitutions, insertions, or deletions thereto.

[0135] 66. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the first and second parapalindromic sequences comprise completely palindromic sequences.

[0136] 67. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the first and / or second parapalindromic sequence comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 non-palindromic positions.

[0137] 69. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the first and second parapalindromic sequences are the same length.

[0138] 70. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the core sequence is about 2-20 nucleotides (e.g., 2-16 nucleotides) in length.

[0139] 71. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the core sequence, e.g., core dinucleotide, is capable of hybridizing to a corresponding sequence, e.g., a dinucleotide or its reverse complement, in the human genome.

[0140] 72. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the core sequence has at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% identity to the corresponding sequence in the human genome.

[0141] 73. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the core sequence has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or fewer mismatches to the corresponding sequence in the human genome.

[0142] 74. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the core sequence (e.g., core dinucleotide), when cleaved by the recombinase, forms sticky ends that can hybridize with corresponding sequences in the human genome.

[0143] 75. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any of the preceding embodiments, wherein the heterologous sequence of interest comprises a eukaryotic gene, e.g., a mammalian gene, e.g., a human gene, such as a blood factor (e.g., genomic factor I, II, V, VII, X, XI, XII, or XIII), or an enzyme, such as a lysosomal enzyme, or a synthetic human gene (e.g., a chimeric antigen receptor).

[0144] 76. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the inserted DNA comprises a heterologous sequence of interest and a DNA recognition sequence.

[0145] 77. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the insert DNA comprises a nucleic acid sequence encoding a recombinase polypeptide.

[0146] 78. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the insert DNA and the nucleic acid encoding the recombinase polypeptide are present in separate nucleic acid molecules.

[0147] 79. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the insert DNA and the nucleic acid encoding the recombinase polypeptide are present in the same nucleic acid molecule.

[0148] 80. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any of the preceding embodiments, wherein the insert DNA further comprises: (a) an open reading frame, e.g., a sequence encoding a polypeptide, e.g., an enzyme (e.g., a lysosomal enzyme), a blood factor, or an exon; (b) non-coding and / or regulatory sequences, such as sequences that bind transcriptional modulators, e.g., promoters (e.g., heterologous promoters), enhancers, insulators; (c) splice acceptor site; (d) polyA site; (e) epigenetic modification sites; (f) Gene expression unit.

[0149] 81. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any of the preceding embodiments, wherein the inserted DNA comprises a plasmid, a viral vector (e.g., a lentiviral vector or an episomal vector), or other self-replicating vector.

[0150] 82. The system, cell, method, isolated recombinase polypeptide or isolated nucleic acid of any of the preceding embodiments, wherein the cell does not contain an endogenous human gene contained in a heterologous sequence of interest or a protein encoded by said gene.

[0151] 83. The system, cell, method, isolated recombinase polypeptide or isolated nucleic acid of any of the preceding embodiments, wherein the cell is derived from an organism that does not contain the endogenous human gene comprised by the heterologous sequence of interest or the protein encoded by said gene.

[0152] 84. The system, cell, method, isolated recombinase polypeptide or isolated nucleic acid of any preceding embodiment, wherein the cell comprises an endogenous human DNA recognition sequence.

[0153] 85. Endogenous human DNA recognition sequences meet the following criteria: (i) located >300 kb from a cancer-related gene; (ii) located >300 kb from miRNAs / other functional small RNAs; (iii) located >50 kb from the 5′ gene end; (iv) located >50 kb from the origin of replication; (v) located >50 kb from any ultraconserved element; (vi) have low transcriptional activity (i.e., no mRNA + / - 25 kb); (vii) are not within a copy number variable region; (viii) is in open chromatin; and / or (xi) It is unique, for example, having one copy in the human genome. 85. The system, cell, method, isolated recombinase polypeptide or isolated nucleic acid of embodiment 84, wherein the recombinase is operably linked to or located within a site in the human genome having at least one, two, three, four, five, six, seven, eight or nine of the following:

[0154] 85a. The system, cell, method, isolated recombinase polypeptide or isolated nucleic acid of any of embodiments 84 or 85, wherein the cell comprises a second endogenous human DNA recognition sequence.

[0155] 85b. The second endogenous human DNA recognition sequence meets the following criteria: (i) located >300 kb from a cancer-related gene; (ii) located >300 kb from miRNAs / other functional small RNAs; (iii) located >50 kb from the 5′ gene end; (iv) located >50 kb from the origin of replication; (v) located >50 kb from any ultraconserved element; (vi) have low transcriptional activity (i.e., no mRNA + / - 25 kb); (vii) are not within a copy number variable region; (viii) is in open chromatin; and / or (xi) It is unique, for example, having one copy in the human genome. 85a. The system, cell, method, isolated recombinase polypeptide or isolated nucleic acid of embodiment 85a, wherein the recombinase polypeptide or isolated nucleic acid is operably linked to, e.g., located within, a site in the human genome having at least 1, 2, 3, 4, 5, 6, 7, 8 or 9 of the following:

[0156] 86. The system, cell, method, isolated recombinase polypeptide or isolated nucleic acid of any of the preceding embodiments, wherein the cell is an animal cell, e.g., a mammalian cell, e.g., a human cell.

[0157] 87. The system, cell, method, isolated recombinase polypeptide or isolated nucleic acid of any of the preceding embodiments, wherein the cell is a plant cell.

[0158] 88. The system, cell, method, isolated recombinase polypeptide or isolated nucleic acid of any of the preceding embodiments, wherein the cell is not genetically engineered.

[0159] 89. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any of the preceding embodiments, wherein the cell does not contain an attB or attP site.

[0160] 89a. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the cell (e.g., prior to contact with the system) comprises a pseudorecognition sequence.

[0161] 89b. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the cell (e.g., prior to contact with the system) contains exactly one pseudo-recognition sequence.

[0162] 90. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the recombinase polypeptide comprises an amino acid sequence corresponding to a single amino acid sequence in Table 3A, 3B, or 3C.

[0163] 91. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the recombinase polypeptide comprises all or part of a plurality of amino acid sequences in Table 3A, 3B, or 3C.

[0164] 92. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of embodiment 91, wherein the recombinase polypeptide comprises a first amino acid sequence from a portion of a first recombinase polypeptide sequence in Table 3A, 3B, or 3C, and a second amino acid sequence from a portion of a different second recombinase polypeptide sequence in Table 3A, 3B, or 3C.

[0165] 93. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of embodiment 92, wherein the first amino acid sequence corresponds to a domain of the first recombinase polypeptide (e.g., an N-terminal catalytic domain, a recombinase domain, a zinc ribbon domain, or a C-terminal DNA-binding domain).

[0166] 94. The system, cell, method, isolated recombinase polypeptide or isolated nucleic acid of any of embodiments 92 or 93, wherein the second amino acid sequence corresponds to a domain of the second recombinase polypeptide (e.g., an N-terminal catalytic domain, a recombinase domain, a zinc ribbon domain or a C-terminal DNA-binding domain), e.g., a domain that is different from the domain of the first amino acid sequence.

[0167] 95. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein one or more of the core sequences of the inserted DNA comprises a core dinucleotide that is modified to match a core dinucleotide of a target recognition sequence in the genomic DNA (and optionally not match at least one core dinucleotide of a non-target recognition sequence in the genomic DNA).

[0168] 96. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein one or more of the core sequences of the insert DNA comprises a core dinucleotide that has been modified to match a core dinucleotide of a recognition sequence present within the nucleotide sequence of the left region or right region column of Table 2A, 2B, or 2C (and optionally not match at least one core dinucleotide of a non-target recognition sequence present within the nucleotide sequence of the left region or right region column of Table 2A, 2B, or 2C).

[0169] 100. The system or method of any of the previous embodiments, wherein the nucleic acid encoding the recombinase polypeptide is a viral vector, e.g., an AAV vector.

[0170] 101. The system or method of any of the previous embodiments, wherein the double-stranded insert DNA is in a viral vector, e.g., an AAV vector.

[0171] 102. The system or method of any preceding embodiment, wherein the nucleic acid encoding the recombinase polypeptide is mRNA, and optionally, the mRNA is in a LNP.

[0172] 103. The system or method of any preceding embodiment, wherein the double-stranded insert DNA is not a viral vector, e.g., the double-stranded insert DNA is naked DNA or DNA in a transfection reagent.

[0173] 104. The nucleic acid encoding the recombinase polypeptide is in a first viral vector, e.g., a first AAV vector, and 10. The system or method of any preceding embodiment, wherein the insert DNA is in a second viral vector, e.g., a second AAV vector.

[0174] 105. The nucleic acid encoding the recombinase polypeptide is mRNA, optionally, the mRNA is in a LNP; and 10. The system or method of any preceding embodiment, wherein the insert DNA is in a viral vector, e.g., an AAV vector.

[0175] 106. The nucleic acid encoding the recombinase polypeptide is mRNA; 10. The system or method of any preceding embodiment, wherein the double-stranded insert DNA is not in a viral vector, e.g., the double-stranded insert DNA is naked DNA or DNA in a transfection reagent.

[0176] 107. The system or method of any preceding embodiment, wherein the insert DNA has a length of at least 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb.

[0177] 108. The system or method of any of the previous embodiments, wherein the inserted DNA does not include an antibiotic resistance gene or any other bacterial gene or moiety.

[0178] R1. The system, kit, polypeptide, or reaction mixture of any of the preceding embodiments, wherein the system comprises one or more circular RNA molecules (circRNAs).

[0179] R2. The system, kit, polypeptide or reaction mixture of embodiment R1, wherein the circRNA encodes a Gene Writer polypeptide.

[0180] R3. The system, kit, polypeptide or reaction mixture of any of embodiments R1 to R2A, wherein the circRNA is delivered to a host cell.

[0181] R4. The system, kit, polypeptide or reaction mixture of any of the preceding embodiments, wherein the circRNA is, for example, linearized in a host cell, for example in the nucleus of the host cell.

[0182] 10. The system, kit, polypeptide, or reaction mixture of any preceding embodiment, wherein the R4A.circRNA comprises a cleavage site.

[0183] The system, kit, polypeptide or reaction mixture of any embodiment R4A, wherein the R4A1.circRNA further comprises a second cleavage site.

[0184] R4B. The system, kit, polypeptide or reaction mixture of embodiment R4A or R4A1, wherein the cleavage site can be cleaved (e.g., by self-cleavage) by a ribozyme, e.g., a ribozyme contained in a circRNA.

[0185] R5. The system, kit, polypeptide, or reaction mixture of any of the preceding embodiments, wherein the circRNA comprises a ribozyme sequence.

[0186] R6. The system, kit, polypeptide or reaction mixture of embodiment R5, wherein the ribozyme sequence is capable of self-cleaving, eg, in a host cell, eg, in the nucleus of a host cell.

[0187] R6A. The system, kit, polypeptide or reaction mixture of any of embodiments R5-R6, wherein the ribozyme is an inducible ribozyme.

[0188] R7. The system, kit, polypeptide or reaction mixture of any of embodiments R5 to R6A, wherein the ribozyme is a protein-responsive ribozyme, e.g., a ribozyme responsive to a nuclear protein, e.g., a genome-interacting protein, e.g., an epigenetic modifier, e.g., EZH2.

[0189] R8. The system, kit, polypeptide or reaction mixture of any of embodiments R5 to R7, wherein the ribozyme is a nucleic acid-responsive ribozyme.

[0190] R8A. The system, kit, polypeptide, or reaction mixture of embodiment R8, wherein the catalytic activity (e.g., autocatalytic activity) of the ribozyme is activated in the presence of a target nucleic acid molecule (e.g., an RNA molecule, such as an mRNA, miRNA, ncRNA, lncRNA, tRNA, snRNA, or mtRNA).

[0191] R9A. The system, kit, polypeptide or reaction mixture of any of embodiments R5-R7, wherein the ribozyme is responsive to a target protein (eg, MS2 coat protein).

[0192] R9B. The system, kit, polypeptide or reaction mixture of embodiment R8A, wherein the target protein is localized in the cytoplasm or in the nucleus (e.g., an epigenetic modifier or a transcription factor).

[0193] R9C. The system, kit, polypeptide, or reaction mixture of any of embodiments R5-R8, wherein the ribozyme comprises a ribozyme sequence of a B2 or ALU retrotransposon or a nucleic acid sequence having at least 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.

[0194] R10A. The system, kit, polypeptide or reaction mixture of any of embodiments R5 to R8, wherein the ribozyme comprises a sequence of a tobacco ringspot virus hammerhead ribozyme or a nucleic acid sequence having at least 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.

[0195] R10B. The system, kit, polypeptide, or reaction mixture of any of embodiments R5 to R8, wherein the ribozyme comprises a nucleic acid sequence having a sequence of a hepatitis delta virus (HDV) ribozyme or at least 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.

[0196] R11. The system, kit, polypeptide or reaction mixture of any of embodiments R5-X, wherein the ribozyme is activated by a moiety expressed in a target cell or tissue.

[0197] R12. The system, kit, polypeptide or reaction mixture of any of embodiments R5 to X, wherein the ribozyme is activated by a moiety expressed in a target intracellular compartment (e.g., nucleus, nucleolus, cytoplasm or mitochondria).

[0198] R4A. The system, kit, polypeptide or reaction mixture of any preceding embodiment, wherein the ribozyme is comprised in a circular RNA or a linear RNA.

[0199] M1. The system, kit, polypeptide or reaction mixture of any of the preceding embodiments, wherein the system, polypeptide and / or DNA encoding same is formulated as a lipid nanoparticle (LNP).

[0200] M2a. The system, kit, polypeptide or reaction mixture of embodiment M1, wherein the lipid nanoparticle (or a formulation comprising a plurality of lipid nanoparticles) lacks reactive impurities (e.g., aldehydes) or contains reactive impurities (e.g., aldehydes) below a preselected level.

[0201] M2. The system, kit, polypeptide or reaction mixture of embodiment M1, wherein the lipid nanoparticle (or formulation comprising a plurality of lipid nanoparticles) is devoid of aldehydes or comprises aldehydes below a preselected level.

[0202] M3. The system, kit, polypeptide or reaction mixture of embodiment M1 or M2, wherein the lipid nanoparticle is comprised in a formulation comprising a plurality of lipid nanoparticles.

[0203] M4. The system, kit, polypeptide or reaction mixture of embodiment M3, wherein the lipid nanoparticle formulation is produced using one or more lipid reagents having a total reactive impurity (e.g., aldehyde) content of less than 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or 0.1%.

[0204] M5. The system, kit, polypeptide or reaction mixture of embodiment M4, wherein the lipid nanoparticle formulation is produced using one or more lipid reagents having a total reactive impurity (e.g., aldehyde) content of less than 3%.

[0205] M6. The system, kit, polypeptide or reaction mixture of any of embodiments M3-M5, wherein the lipid nanoparticle formulation is produced using one or more lipid reagents containing less than 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or 0.1% of any single reactive impurity (e.g., aldehyde) species.

[0206] M7. The system, kit, polypeptide or reaction mixture of embodiment M6, wherein the lipid nanoparticle formulation is produced using one or more lipid reagents containing less than 0.3% of any single reactive impurity (e.g., aldehyde) species.

[0207] M8. The system, kit, polypeptide or reaction mixture of embodiment M6, wherein the lipid nanoparticle formulation is produced using one or more lipid reagents containing less than 0.1% of any single reactive impurity (e.g., aldehyde) species.

[0208] M9. The system, kit, polypeptide or reaction mixture of any of embodiments M3-M8, wherein the lipid nanoparticle formulation has a total reactive impurity (e.g., aldehyde) content of less than 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or 0.1%.

[0209] M10. The system, kit, polypeptide or reaction mixture of embodiment M9, wherein the lipid nanoparticle formulation has a total reactive impurity (e.g., aldehyde) content of less than 3%.

[0210] M11. The system, kit, polypeptide or reaction mixture of any of embodiments M3-M10, wherein the lipid nanoparticle formulation contains less than 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2% or 0.1% of any single reactive impurity (e.g., aldehyde) species.

[0211] M12. The system, kit, polypeptide or reaction mixture of embodiment M11, wherein the lipid nanoparticle formulation contains less than 0.3% of any single reactive impurity (e.g., aldehyde) species.

[0212] M13. The system, kit, polypeptide or reaction mixture of embodiment M11, wherein the lipid nanoparticle formulation contains less than 0.1% of any single reactive impurity (e.g., aldehyde) species.

[0213] M14. The system, kit, polypeptide or reaction mixture of any of embodiments M1-M13, wherein one or more, or optionally all, of the lipid reagents used in the lipid nanoparticles or formulations thereof described herein have a total reactive impurity (e.g., aldehyde) content of less than 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, or 0.1%.

[0214] M15. The system, kit, polypeptide or reaction mixture of embodiment M14, wherein one or more, or optionally all, of the lipid reagents used in the lipid nanoparticles or formulations thereof described herein have a total reactive impurity (e.g., aldehyde) content of less than 3%.

[0215] M16. The system, kit, polypeptide or reaction mixture of any of embodiments M1-M15, wherein one or more or optionally all of the lipid reagents used in the lipid nanoparticles or formulations thereof described herein contain less than 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2% or 0.1% of any single reactive impurity (e.g., aldehyde) species.

[0216] M17. The system, kit, polypeptide or reaction mixture of embodiment M16, wherein one or more, or optionally all, of the lipid reagents used in the lipid nanoparticles or formulations thereof described herein contain less than 0.3% of any single reactive impurity (e.g., aldehyde) species.

[0217] M18. The system, kit, polypeptide or reaction mixture of embodiment M16, wherein one or more, or optionally all, of the lipid reagents used in the lipid nanoparticles or formulations thereof described herein contain less than 0.1% of any single reactive impurity (e.g., aldehyde) species.

[0218] M19. The system, kit, polypeptide, or reaction mixture of any of embodiments M1-M18, wherein the total aldehyde content and / or the amount of any single reactive impurity (e.g., aldehyde) species is determined, e.g., by liquid chromatography (LC) coupled with tandem mass spectrometry (MS / MS), e.g., according to the method described in Example 26.

[0219] M20. The system, kit, polypeptide, or reaction mixture of any of embodiments M1-M18, wherein the total aldehyde content and / or the amount of reactive impurity (e.g., aldehyde) species is determined by detecting one or more chemical modifications of a nucleic acid molecule (e.g., as described herein) associated with the presence of reactive impurities (e.g., aldehydes), e.g., in a lipid reagent.

[0220] M21. The system, kit, polypeptide, or reaction mixture of any of embodiments M1-M18, wherein the total aldehyde content and / or amount of aldehyde species is determined by detecting one or more chemical modifications of a nucleotide or nucleoside (e.g., a ribonucleotide or ribonucleoside, e.g., comprised in or isolated from a nucleic acid molecule, e.g., as described herein) associated with the presence of a reactive impurity (e.g., an aldehyde), e.g., in a lipid reagent, e.g., as described in Example 27.

[0221] M22. The system, kit, polypeptide or reaction mixture of embodiment M21, wherein chemical modifications of nucleic acid molecules, nucleotides or nucleosides are detected by determining the presence of one or more modified nucleotides or nucleosides, e.g., using LC-MS / MS analysis, e.g., as described in Example 27.

[0222] T1. A lipid nanoparticle (LNP) comprising the system, polypeptide (or RNA encoding same), nucleic acid molecule, or DNA encoding the system or polypeptide of any preceding embodiment.

[0223] T2. A first lipid nanoparticle comprising a polypeptide (or DNA or RNA encoding the same) of a Gene Writing System (e.g., as described herein); a second lipid nanoparticle containing a nucleic acid molecule of a Gene Writing system (e.g., as described herein); A system including:

[0224] T3. The system, kit, polypeptide or reaction mixture of any preceding embodiment, wherein the system, nucleic acid molecule, polypeptide and / or DNA encoding same is formulated as a lipid nanoparticle (LNP).

[0225] U1. The system, kit, polypeptide or reaction mixture of any preceding embodiment, wherein the serine recombinase comprises at least one active site signature of a serine recombinase, e.g., cd00338, cd03767, cd03768, cd03769 or cd03770.

[0226] U2. Serine recombinases can be found in public databases (e.g., InterPro, UniProt, or the Conserved Domain Database (Lu)), for example, as described herein.

[0023] The system, kit, polypeptide, or reaction mixture of any preceding embodiment, further comprising a domain identified from a gene encoding a polypeptide of interest (described by Gill, J. et al. Nucleic Acids Res 48, D265-268 (2020); the entire contents of which are incorporated herein by reference).

[0227] U3. The system, kit, polypeptide, or reaction mixture of any preceding embodiment, wherein the serine recombinase comprises a domain identified by scanning the open reading frame or full-frame translation of a nucleic acid sequence for a serine recombinase domain (e.g., as described herein), e.g., using a prediction tool, e.g., InterProScan, e.g., as described herein.

[0228] V0. The system, kit, polypeptide, cell (e.g., a cell produced by the methods herein), method, or reaction mixture of any preceding embodiment, wherein the heterologous sequence of interest is located (e.g., inserted) in a target site in the genome of the cell, and optionally the target site comprises, in order: (i) a first parapalindromic sequence (e.g., an attL site), (ii) the heterologous sequence of interest, and (iii) a second parapalindromic sequence (e.g., an attR site).

[0229] V1. The system, kit, polypeptide, cell, method, or reaction mixture of embodiment V0, wherein the cell (e.g., a cell produced by a method herein) comprises an insertion or deletion between (i) the first parapalindromic sequence and (ii) the heterologous sequence of interest, or the cell comprises an insertion or deletion between (ii) the heterologous sequence of interest and (iii) the second parapalindromic sequence.

[0230] V3. The system, kit, polypeptide, cell, method or reaction mixture of embodiment V1, wherein the insertion or deletion involves fewer than 20 nucleotides or base pairs, e.g., fewer than 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2 or fewer than 1 nucleotide or base pair, of the nucleic acid sequence at the target site.

[0231] V4. The system, kit, polypeptide, cell, method, or reaction mixture of embodiment V1, wherein the insertion comprises fewer than 20 nucleotides or base pairs, e.g., fewer than 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide or base pair.

[0232] V5. The system, kit, polypeptide, cell, method or reaction mixture of embodiment V1, wherein the deletion includes less than 20 nucleotides or base pairs, such as less than 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2 or less than 1 nucleotide or base pair, of the sequence preceding the target site.

[0233] V6. The system, kit, polypeptide, cell, method, or reaction mixture of any of embodiments V0-V5, wherein the core region (e.g., central dinucleotide) of the recognition sequence of the target site (e.g., listed in Table 4X, e.g., an attB, attP, or pseudosite thereof) comprises about 95%, 96%, 97%, 98%, 99%, or 100% identity to the core region (e.g., central dinucleotide) of the recognition sequence (on the insert DNA, e.g., listed in Table 4X, e.g., an attP or attB site).

[0234] V7. The system, kit, polypeptide, cell, method, or reaction mixture of embodiment V6, wherein the number of insertions or deletions in the target site is less than the number of insertions or deletions in an otherwise similar cell with a lower percent identity.

[0235] V8. The system, kit, polypeptide, cell, method, or reaction mixture of embodiment V7, wherein the number of insertion or deletion events is at least 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 3.0, 4.0, 5.0, 10, 20, 30, 40, 50, 60, 70, 80, 90 or at least 100 times lower.

[0236] V9. The system, kit, polypeptide, cell, method or reaction mixture of any of embodiments V0-V8, wherein the target site does not comprise multiple insertions (e.g., head-to-tail or head-to-head duos).

[0237] V9a. The system, kit, polypeptide, cell, method or reaction mixture of any of embodiments V0-V9, wherein the target site comprises less than 100, 75, 50, 45, 40, 35, 30, 25, 20, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3 or 2 copies of the heterologous sequence of interest or a fragment thereof.

[0238] V10. The system, kit, polypeptide, cell, method or reaction mixture of any of embodiments V0-V9a, wherein the target site comprises a single copy of the heterologous sequence of interest or a fragment thereof.

[0239] V11. A system, kit, polypeptide, cell, method or reaction mixture according to any of embodiments V0 to V10, wherein (e.g., in a population of cells) target sites exhibiting two or more copies of the heterologous sequence of interest or fragment thereof are less than 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 9%, 8%, 7%, 6%, 4%, 4%, 3%, 2% or 1% of target sites containing at least one copy of the heterologous sequence of interest or fragment thereof.

[0240] V12. A system, kit, polypeptide, cell, method or reaction mixture described in any of embodiments V0 to V11, wherein (e.g., in a population of cells) target sites displaying 3 or more copies of the heterologous sequence of interest or fragment thereof are less than 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 9%, 8%, 7%, 6%, 4%, 4%, 3%, 2% or 1% of target sites containing at least 1 copy of the heterologous sequence of interest or fragment thereof.

[0241] V13. A system, kit, polypeptide, cell, method or reaction mixture of any of embodiments V0-V12, wherein (e.g., in a population of cells) target sites displaying 4 or more copies of the heterologous sequence of interest or fragment thereof are less than 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 9%, 8%, 7%, 6%, 4%, 4%, 3%, 2% or 1% of target sites containing at least 1 copy of the heterologous sequence of interest or fragment thereof.

[0242] V14. The system, kit, polypeptide, cell, method, or reaction mixture of any of embodiments V0-V13, wherein the target site comprises one or more ITRs (e.g., AAV ITRs), e.g., one, two, three, four or more ITRs, e.g., the one or more ITRs are located between (i) a first parapalindromic sequence and (iii) a second parapalindromic sequence.

[0243] V15. The system, kit, polypeptide, cell, method, or reaction mixture of embodiment V14, wherein (e.g., in a population of cells) target sites comprising an ITR (e.g., an AAV ITR) between (i) a first parapalindromic sequence and (iii) a second parapalindromic sequence are at least 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of target sites comprising at least one copy of a heterologous sequence of interest or a fragment thereof.

[0244] V16. The system, kit, polypeptide, cell, method or reaction mixture of embodiment V14 or V15, wherein the insertion site comprises one or more copies of the heterologous sequence of interest or fragment thereof.

[0245] V17. The system, kit, polypeptide, cell, method or reaction mixture of any of embodiments V0 to V16, wherein the target site comprises, in order, (i) a first parapalindromic sequence, and (ii) a heterologous sequence of interest.

[0246] V18. The system, kit, polypeptide, cell, method, or reaction mixture of embodiment V17, wherein the target site does not comprise (iii) a second parapalindromic sequence.

[0247] V19. The system, kit, polypeptide, cell, method, or reaction mixture of any of embodiments V0-V17, wherein the target site comprises (iii) a second parapalindromic sequence, and (ii) is located between (i) and (iii).

[0248] V20. The system, kit, polypeptide, cell, method, or reaction mixture of any of embodiments V0-V19, wherein (e.g., in a population of cells) target sites comprising both (i) the first parapalindromic sequence and (iii) the third parapalindromic sequence comprise a high percentage of fully heterologous target sequences (e.g., at least 0.1×, 0.2×, 0.3×, 0.4×, 0.5×, 0.6×, 0.7×, 0.8×, 0.9×, 1.0×, 1.5×, 2.0×, 3×, 4×, 5×, 6×, 7×, 8×, 9×, 10× or more percent fully heterologous target sequences) relative to the percentage of target sites comprising one or fewer parapalindromic sequences (e.g., attL or attP sequences).

[0249] The present disclosure contemplates any and all combinations of any one or more of the above aspects and / or embodiments, as well as combinations with any one or more of the embodiments described in the detailed description and examples.

[0250] definition About, Approximately: "About" or "approximately," as used herein when applied to one or more values ​​of interest, refers to a value similar to the stated reference value. In certain embodiments, the term "about" or "approximately" refers to a range of values ​​that fall within 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less in either direction (more or less) of the stated reference value, unless otherwise stated or otherwise clear from the context (except where such number exceeds 100% of the possible value).

[0251] Domain: As used herein, the term "domain" refers to a structure of a biomolecule that contributes to a specific function of the biomolecule. A domain can include a continuous region (e.g., a contiguous sequence) or discrete, non-contiguous regions (e.g., a non-contiguous sequence) of a biomolecule. Examples of protein domains include, but are not limited to, a nuclear localization sequence, a recombinase domain, a DNA recognition domain (e.g., that binds to or is capable of binding to a recognition site as described herein), a recombinase N-terminal domain (also referred to as a catalytic domain), a recombinase domain, a C-terminal zinc ribbon domain, and the domains listed in Table 4. In some embodiments, the zinc ribbon domain further comprises a coiled-coil motif. In some embodiments, the recombinase domain and the zinc ribbon domain are collectively referred to as the C-terminal domain. In some embodiments, the N-terminal domain is connected to the C-terminal domain by an αE linker or helix. In some embodiments, the N-terminal domain is 50 to 250 amino acids, or 100 to 200 amino acids, or 130 to 170 amino acids, e.g., about 150 amino acids. In some embodiments, the C-terminal domain is 200 to 800 amino acids or 300 to 500 amino acids. In some embodiments, the recombinase domain is 50 to 150 amino acids. In some embodiments, the zinc ribbon domain is 30 to 100 amino acids; an example of a nucleic acid domain is a regulatory domain, such as a transcription factor binding domain, a recognition sequence, an arm of a recognition sequence (e.g., a 5' or 3' arm), a core sequence, or a sequence of interest (e.g., a heterologous sequence of interest). In some embodiments, the recombinase polypeptide comprises one or more domains (e.g., a recombinase domain or a DNA recognition domain) of a polypeptide of Table 3A, 3B, or 3C, or a fragment or variant thereof.

[0252] Exogenous: As used herein, the term "exogenous," when used in reference to a biomolecule (such as a nucleic acid sequence or polypeptide), means that the biomolecule has been introduced into a host genome, cell, or organism by human intervention. For example, a nucleic acid that is added to an existing genome, cell, tissue, or subject using recombinant DNA technology or other methods is exogenous to the existing nucleic acid sequence, cell, tissue, or subject.

[0253] Genomic Safe Harbor Site (GSH Site): A genomic safe harbor site is a site within a host genome that can accommodate the integration of new genetic material, such that the inserted genetic element does not cause significant alterations to the host genome that pose a risk to the host cell or organism. GSH sites generally meet one, two, three, four, five, six, seven, eight, or nine of the following criteria: (i) located >300 kb from a cancer-related gene; (ii) located >300 kb from an miRNA / other functional small RNA; (iii) located >50 kb from the 5' end of a gene; (iv) located >50 kb from a replication origin; (v) located >50 kb away from an ultraconserved element; (vi) having low transcriptional activity (i.e., no mRNA + / - 25 kb); (vii) not within a variable copy number region; (viii) located within open chromatin; and / or (ix) having one copy and being unique within the human genome. Examples of GSH sites in the human genome that meet some or all of these criteria include: (i) adenovirus site 1 (AAVS1), the naturally occurring integration site of the AAV virus on chromosome 19; (ii) the chemokine (CC motif) receptor 5 (CCR5) gene, a chemokine receptor gene known as an HIV-1 co-receptor; (iii) the human orthologue of the mouse Rosa26 locus; and (iv) the rDNA locus. Additional GSH sites are known and are described, for example, in Pellenz et al. (2018) (https: / / doi.org / 10.1101 / 396390).

[0254] Heterologous: The term "heterologous," when used to refer to a first element in relation to a second element, means that the first and second elements do not naturally occur in the arrangement described. For example, a heterologous polypeptide, nucleic acid molecule, construct, or sequence refers to (a) a polypeptide, nucleic acid molecule, or portion of a polypeptide or nucleic acid molecule sequence that is not native to the cell in which it is expressed; (b) a polypeptide or nucleic acid molecule, or portion of a polypeptide or nucleic acid molecule, that has been modified or mutated relative to its natural state; or (c) a polypeptide or nucleic acid molecule that has altered expression compared to native expression levels under similar conditions. For example, heterologous regulatory sequences (e.g., promoters, enhancers) can be used to regulate expression of a gene or nucleic acid molecule in a manner that differs from how the gene or nucleic acid molecule is normally expressed in nature. In certain embodiments, a heterologous nucleic acid molecule can be present in the native host cell genome but can have an altered expression level, a different sequence, or both. In other embodiments, the heterologous nucleic acid molecule may not be endogenous to the host cell or host genome, but instead may be introduced into the host cell by transformation (e.g., transfection, electroporation), where the added molecule may be integrated into the host genome or may exist as extrachromosomal genetic material either transiently (e.g., mRNA) or semi-stable for two or more generations (e.g., episomal viral vectors, plasmids or other self-replicating vectors).

[0255] Mutation or Mutant: The term "mutation," when applied to a nucleic acid sequence, means that nucleotides within a nucleic acid sequence may be inserted, deleted, or changed relative to a reference (e.g., naturally occurring) nucleic acid sequence. A single alteration may be made at one locus (point mutation), or multiple nucleotides may be inserted, deleted, or changed at a single locus. In addition, one or more alterations may be made at any number of loci within a nucleic acid sequence. Nucleic acid sequences can be mutated by any method known in the art.

[0256] Nucleic acid molecule: Nucleic acid molecule refers to both RNA and DNA molecules, including, but not limited to, cDNA, genomic DNA, and mRNA, and also includes synthetic nucleic acid molecules, such as those chemically synthesized or recombinantly produced, such as from a DNA template, as described herein. Nucleic acid molecules can be double-stranded or single-stranded, circular, or linear. If single-stranded, the nucleic acid molecule can be the sense or antisense strand. Unless otherwise noted, and as an example of all sequences described herein in the general format "SEQ ID NO:1," a nucleic acid containing "SEQ ID NO:1" refers to a nucleic acid having, at least a portion thereof, either (i) the sequence of SEQ ID NO:1 or (ii) a sequence complementary to SEQ ID NO:1. The choice between the two is determined by the context in which SEQ ID NO:1 is used. For example, if the nucleic acid is used as a probe, the choice between the two is determined by the requirement that the probe be complementary to the desired target. The nucleic acid sequences of the present disclosure can be chemically or biochemically modified or contain non-natural or derivatized nucleotide bases, as will be readily understood by those of skill in the art. Such modifications include, for example, labels, methylation, substitution of one or more naturally occurring nucleotides with analogs, internucleotide modifications such as uncharged linkages (e.g., methylphosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), pendant moieties (e.g., polypeptides), intercalating agents (e.g., acridines, psoralens, etc.), chelators, alkylating agents, and modified linkages (e.g., α-anomeric nucleic acids, etc.). Also included are synthetic molecules that mimic polynucleotides in their ability to bind to designated sequences through hydrogen bonds and other chemical interactions. Such molecules are known in the art and include, for example, those that substitute peptide linkages for phosphate linkages in the backbone of the molecule. Other modifications can include, for example, analogs in which the ribose ring contains a bridging moiety or other structures, such as those found in "locked" nucleic acids.

[0257] Gene Expression Unit: A gene expression unit is a nucleic acid sequence comprising at least one regulatory nucleic acid sequence operably linked to at least one effector sequence. A first nucleic acid sequence is operably linked to a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For example, a promoter or enhancer is operably linked to a coding sequence if the promoter or enhancer affects the transcription or expression of the coding sequence. Operably linked DNA sequences can be contiguous or non-contiguous. Where necessary to link two protein coding regions, operably linked sequences can be in the same reading frame.

[0258] Host: The term host genome or host cell, as used herein, refers to a cell and / or its genome into which proteins and / or genetic material have been introduced. These terms refer not only to the particular subject cell and / or genome, but also to the progeny of such a cell and / or the genomes of the progeny of such a cell. Because certain modifications may occur in subsequent generations due to mutations or environmental influences, it is understood that such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term "host cell" as used herein. A host genome or host cell may be an isolated cell or cell line grown in culture, or genomic material isolated from such a cell or cell line, or may be a host cell or host genome comprising a living tissue or organism. In some cases, a host cell may be an animal cell or a plant cell, e.g., as described herein. In certain examples, a host cell may be a bovine cell, equine cell, porcine cell, caprine cell, ovine cell, chicken cell, or turkey cell. In certain examples, a host cell may be a corn cell, soybean cell, wheat cell, or rice cell.

[0259] Recombinase polypeptide: As used herein, a recombinase polypeptide refers to a polypeptide having the functional ability to catalyze a recombination reaction of nucleic acid molecules (e.g., DNA molecules). A recombination reaction can involve, for example, one or more nucleic acid strand breaks (e.g., double-strand breaks), followed by joining of the two nucleic acid strand ends (e.g., sticky ends). In some cases, a recombination reaction involves, for example, insertion of an insert nucleic acid into a target site, for example, in a genome or construct. In some cases, a recombination reaction involves, for example, flipping or inverting a nucleic acid in a genome or construct. In some cases, a recombination reaction involves, for example, removing a nucleic acid from a genome or construct. In some cases, a recombinase polypeptide comprises one or more structural elements of a naturally occurring recombinase (e.g., a serine recombinase, e.g., PhiC31 recombinase or Gin recombinase). In certain cases, the recombinase polypeptide comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a recombinase described herein (e.g., listed in Table 3A, 3B, or 3C). In some embodiments, the recombinase polypeptide comprises a serine recombinase, e.g., a serine integrase. In some embodiments, the serine recombinase, e.g., a serine integrase, comprises one or more (e.g., all) of a recombinase domain, a catalytic domain, or a zinc ribbon domain. In some embodiments, the serine recombinase, e.g., a serine integrase, comprises a domain listed in Table 4 (e.g., in addition to or instead of one or more of a recombinase domain, a catalytic domain, or a zinc ribbon domain). In some cases, the recombinase polypeptide has one or more functional characteristics of a naturally occurring recombinase (e.g., a serine recombinase, such as PhiC31 recombinase or Gin recombinase). In some embodiments, the recombinase polypeptide is between 350 and 900 amino acids or between 425 and 700 amino acids.In some cases, the recombinase polypeptide recognizes (e.g., binds to) a recognition sequence in a nucleic acid molecule (e.g., a recognition sequence present in a sequence in the left region and / or right region column of Table 2A, 2B, or 2C, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto). In some embodiments, the recombinase may promote recombination between a first recognition sequence (e.g., attB or pseudo-attB) and a second genomic recognition sequence (e.g., attP or pseudo-attP). In some embodiments, the recombinase polypeptide is not active as an isolated monomer. In some embodiments, the recombinase polypeptide catalyzes a recombination reaction in cooperation with one or more other recombinase polypeptides (e.g., two or four recombinase polypeptides per recombination reaction). In some embodiments, the recombinase polypeptide is active as a dimer. In some embodiments, the recombinase assembles as a dimer at the recognition sequence. In some embodiments, the recombinase polypeptide is active as a tetramer. In some embodiments, the recombinase assembles as a tetramer at the recognition sequence. In some embodiments, the recombinase polypeptide is a recombinant (e.g., non-naturally occurring) recombinase polypeptide. In some embodiments, the recombinant recombinase polypeptide comprises an amino acid sequence from multiple recombinase polypeptides (e.g., the recombinant recombinase polypeptide comprises a first domain from a first recombinase polypeptide and a second domain from a second recombinase polypeptide).

[0260] Insert nucleic acid molecule: As described herein, an insert nucleic acid molecule (e.g., insert DNA) is a nucleic acid molecule (e.g., a DNA molecule) that is or will be inserted, at least in part, into a target site within a target nucleic acid molecule (e.g., genomic DNA). An insert nucleic acid molecule can include, for example, a nucleic acid sequence that is heterologous to the target nucleic acid molecule (e.g., genomic DNA). In some cases, the insert nucleic acid molecule includes a sequence of interest (e.g., a heterologous sequence of interest). In some cases, the insert nucleic acid molecule is a cognate DNA recognition sequence, e.g., a DNA present in the target nucleic acid. In some embodiments, the insert nucleic acid molecule is circular, and in some embodiments, the insert nucleic acid molecule is linear. In some embodiments, the insert nucleic acid molecule includes two or more DNA recognition sequences (e.g., two DNA recognition sequences), each of which is, for example, a cognate to a DNA recognition sequence present in the target nucleic acid. In some embodiments, the insert nucleic acid molecule is also referred to as a template nucleic acid molecule (e.g., template DNA).

[0261] Recognition sequence: A recognition sequence (e.g., a DNA recognition sequence) generally refers to a nucleic acid (e.g., DNA) sequence that is recognized (e.g., can be bound by) a recombinase polypeptide, e.g., as described herein. In some cases, the recognition sequence includes two recognition sequences, one located in the integration site (the site where the nucleic acid is to be integrated) and the other adjacent to the nucleic acid of interest to be introduced into the integration site. The recognition sequences are generally referred to as attB and attP. The recognition sequences can be native or modified relative to the native sequence. The recognition sequences can vary in length, but typically range from about 20 to about 200 nt, about 30 to 90 nt, and more usually 30 to 70 nucleotides. The recognition sequences are typically arranged as follows: attB includes, in relative order from 5' to 3', a first DNA sequence attB5', a core region, and a second DNA sequence attB3', attB5'-core region-attB3'. attP comprises, in relative order from 5' to 3', a first DNA sequence attP5', a core region, and a second DNA sequence attP3', attP5'-core region-attP3'. In some embodiments, attP5' and attB3' are parapalindromic (e.g., one sequence is palindromic to the other sequence or has at least 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with a palindrome to the other sequence). In some embodiments, the attP5' and attP3' recognition sequences are parapalindromic (e.g., one sequence is a palindrome to the other sequence or has at least 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to a palindrome to the other sequence). In some embodiments, the attB5' and attB3' recognition sequences are parapalindromic to each other, and the attP5' and attP3' recognition sequences are parapalindromic to each other. In some embodiments, the attB5' and attB3' and attP5' and attP3' sequences are similar, but not necessarily of the same number of nucleotides.Because attB and attP are distinct sequences, recombination results in a stretch of nucleic acid (called attL or attR, relative to the left and right) that is neither an attB nor an attP sequence. While not intending to be bound by theory, the dissimilarity between attL / attR and attB / attP likely makes the attL and attR sites unrecognizable as recombination sites for the relevant recombinase enzyme, thus reducing the probability that the enzyme will catalyze a second recombination reaction that reverses the first. The recognition sequence is typically bound by the recombinase dimer. In some embodiments, the recognition sequence is contacted by one or more of the αE helix, recombinase domain, linker domain, and / or zinc ribbon domain of the recombinase polypeptide. In some cases, the recognition sequence comprises a nucleic acid sequence present within the sequences in the left region or right region columns of Table 2A, 2B, or 2C, e.g., a 20-200 nt sequence within the sequences in the left region or right region columns of Table 2A, 2B, or 2C, e.g., a 30-70 nt sequence within the sequences in the left region or right region columns of Table 2A, 2B, or 2C, or a sequence having at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the recognition site is also referred to as an attachment site. In some embodiments, the recognition sequence is present in a genome, and when describing a recognition sequence that is the site of gene writing activity, it is referred to as a target sequence or target site.

[0262] Pseudo-recognition sequence: Recognition sequences exist in the genomes of various organisms, and they do not necessarily have the same nucleotide sequence as the wild-type recognition sequence (for a given recombinase); however, such native recognition sequences are nonetheless sufficient to promote recombinase-mediated recombination. Such recognition sequences are, among other things, referred to herein as "pseudo-recognition sequences." A "pseudo-recognition sequence" is a DNA sequence containing a recognition sequence that is recognized (e.g., can be bound by) a recombinase enzyme, where the recognition sequence differs from the corresponding wild-type recombinase recognition sequence by one or more nucleotides and / or exists as an endogenous sequence in a genome that differs from the sequence of the genome in which the wild-type recognition sequence for the recombinase remains. In some embodiments, for a given recombinase, the pseudo-recognition sequence is functionally equivalent to the wild-type recombination sequence, exists in organisms other than the organism in which the recombinase is naturally found, and may have sequence variations relative to the wild-type recognition sequence. A "pseudo attP site" or "pseudo attB site" refers to a pseudo-recognition sequence that is similar to the recognition sequence for a wild-type phage (attP) or bacterial (attB) attachment site sequence, respectively, for a phage integrase enzyme, e.g., phage PhiC31. In some embodiments, the attP or pseudo attP site is present in the genome of a host cell, while the attB or pseudo attB site is present on a targeting vector in a system described herein. In some embodiments, the attB or pseudo attB site is present in the genome of a host cell, while the attP or pseudo attP site is present on a targeting vector in a system described herein. A "pseudo att site" is a more general term that can refer to either a pseudo attP site or a pseudo attB site. An att site or pseudo att site can be present on a linear or circular nucleic acid molecule. Identification of a pseudo-recognition sequence can be achieved, for example, by using sequence alignment and analysis, where the query sequence is the recognition sequence of interest (e.g., attB and / or attP in a phage / bacterial system).For example, if a genomic recognition sequence is identified using an attB query sequence, it is referred to as a pseudo-attB site; if a genomic recognition sequence is identified using an attP query sequence, it is referred to as a pseudo-attP site. In some embodiments, the pseudo-recognition sequence shares high sequence similarity with the wild-type recognition sequence recognized by (e.g., capable of binding to) the recombinase (e.g., one or more of the αE helix, recombinase domain, linker domain, and / or zinc ribbon domain described in Li H et al., 2018, J Mol Biol, 430(21):4401-4418 (incorporated by reference)). In some embodiments, the pseudo-recognition sequence is bound to or acted upon by the recombinase more strongly than the wild-type recognition sequence of the recombinase. A pseudo-recognition sequence can also be referred to as a "pseudo-site." In some embodiments, the pseudo-site can be completely mismatched to the parent sequence, as described, for example, in Thyagarajan et al. Mol Cell Biol 21(12):3926-3934 (2001). In some embodiments, a pseudo-site as used herein may be less than 70% identical to a native recognition sequence, e.g., less than 70%, 60%, 50%, 40%, or 30% identical. In some embodiments, a pseudo-site as used herein may be more than 20% identical to a native recognition sequence, e.g., more than 20%, 30%, 40%, 50%, 60%, or 70% identical.

[0263] Hybrid recognition sequence: As used herein, a "hybrid recognition sequence" refers to a recognition sequence constructed from multiple recognition sequences, e.g., portions of wild-type and / or pseudo-recognition sequences. In some embodiments, the multiple recognition sequences are all recognition sequences for the same recombinase (e.g., wild-type and pseudo-recognition sequences recognized by the same recombinase). In some embodiments, the sequence 5' of the core sequence, e.g., attB5' or attP5' of the hybrid recombination site, matches the pseudo-recognition sequence, and the sequence 3' of the core sequence, e.g., attB3' or attP3' of the hybrid recognition sequence, matches the wild-type recognition sequence. In some embodiments, the sequence 5' of the core sequence, e.g., attB5' or attP5' of the hybrid recombination site, matches the wild-type recognition sequence, and the sequence 3' of the core sequence, e.g., attB3' or attP3' of the hybrid recognition sequence, matches the pseudo-recognition sequence. In some embodiments, the sequence 5' of the core sequence, e.g., attB5' or attP5' of the hybrid recombination site, corresponds to the pseudo-recognition sequence, and the sequence 3' of the core sequence, e.g., attB3' or attP3' of the hybrid recognition sequence, corresponds to the wild-type recognition sequence. In some embodiments, a hybrid recognition sequence can be composed of a region 5' of the core sequence from the wild-type attB site and a region 3' of the core sequence from the wild-type attP recognition sequence, or vice versa. Other combinations of such hybrid recognition sequences will be apparent to those of skill in the art in light of the teachings herein. In some embodiments, the recognition sequences suitable for use herein are hybrid recognition sequences.

[0264] Core sequence: As used herein, a core sequence refers to a nucleic acid sequence located between two arms of a recognition sequence, e.g., between a pair of parapalindromic sequences. In some embodiments, the core sequence is located between attB5' and attB3' or between attP5' and attP3'. In some cases, the core sequence can be cleaved by a recombinase polypeptide (e.g., a recombinase polypeptide that recognizes a recognition sequence comprising two parapalindromic sequences) to form, e.g., sticky ends, e.g., 3' overhangs. In some embodiments, the attB and attP core sequences are identical. In some embodiments, the attB and attP core sequences are not identical, e.g., have less than 99, 95, 90, 80, 70, 60, 50, 40, 30, or 20% identity. In some embodiments, the core sequence is about 2 to 20 nucleotides, e.g., 2 to 16 nucleotides, e.g., about 4 nucleotides in length or about 2 nucleotides in length (e.g., exactly 2 nucleotides in length). In some embodiments, the core sequence comprises a core dinucleotide corresponding to two adjacent nucleotides, such that a recombinase that recognizes a nearby para-palindromic sequence can cleave DNA on either side of the core dinucleotide, e.g., forming a sticky end. In some embodiments, the core dinucleotides of the core sequence at the attB and / or attP sites are identical, e.g., cleavage of the attP and / or attB sites forms compatible sticky ends. In some embodiments, the core sequence comprises a nucleic acid sequence present within a nucleotide sequence in the left-hand region or right-hand region columns of Table 2A, 2B, or 2C. In some embodiments, the core sequence comprises a nucleic acid sequence not present within a nucleotide sequence in the left-hand region or right-hand region columns of Table 2A, 2B, or 2C.

[0265] Sequence of interest: As used herein, the term sequence of interest refers to a nucleic acid segment that can be desirably inserted into a target nucleic acid molecule, e.g., by a recombinase polypeptide, e.g., as described herein. In some embodiments, the inserted DNA includes a DNA recognition sequence and a sequence of interest heterologous to the DNA recognition sequence (commonly referred to herein as a "heterologous sequence of interest"). A sequence of interest may, in some cases, be heterologous to the nucleic acid molecule into which it is inserted. In some cases, the sequence of interest includes a nucleic acid sequence encoding a gene (e.g., a eukaryotic cell, e.g., a mammalian cell, e.g., a human gene) or other cargo of interest (e.g., a sequence encoding a functional RNA, e.g., an siRNA or miRNA), e.g., as described herein. In particular cases, the gene encodes a polypeptide (e.g., a blood factor or enzyme). In some cases, the sequence of interest includes one or more nucleic acid sequences encoding a selectable marker (e.g., an auxotrophic marker or antibiotic marker) and / or a nucleic acid control element (e.g., a promoter, enhancer, silencer, or insulator).

[0266] Parapalindrome: As used herein, the term "parapalindrome" refers to a characteristic of a pair of nucleic acid sequences in which one of the nucleic acid sequences is either a palindrome relative to the other nucleic acid sequence, or has at least 30% (e.g., at least 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%), e.g., at least 50%, sequence identity with the other nucleic acid sequence, or has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 or fewer sequence mismatches with the other nucleic acid sequence. A "parapalindromic sequence," as used herein, refers to at least one of a pair of nucleic acid sequences that is parapalindrome relative to each other. "Parapalindromic region," as used herein, refers to a nucleic acid sequence or portion thereof that contains two parapalindromic sequences. In some cases, a parapalindromic region contains two parapalindromic sequences that flank a nucleic acid sequence segment (e.g., containing a core sequence). In an embodiment of the present invention, for example, the following items are provided: (Item 1) 1. A system for modifying DNA, comprising: a) a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or a nucleic acid encoding said recombinase polypeptide; and b) a double-stranded insert DNA, (i) A DNA recognition sequence that binds to the recombinase polypeptide of (a), comprising a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the first and second parapalindromic sequences together being a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or the parapalindromic region. a nucleotide sequence that has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to, or has no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and the DNA recognition sequence further comprises a core sequence of about 2 to 20 nucleotides, the core sequence being located between the first and second parapalindromic sequences; (ii) a heterologous sequence of interest; double-stranded intercalated DNA containing A system including: (Item 2) A eukaryotic cell (e.g., a mammalian cell, e.g., a human cell) comprising a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or a nucleic acid encoding said recombinase polypeptide. (Item 3) a eukaryotic cell (e.g., a mammalian cell, e.g., a human cell), (i) a DNA recognition sequence, comprising a first parapalindromic sequence and a second parapalindromic sequence; each parapalindromic sequence is about 15 to 35 or 20 to 30 nucleotides, and the first and second parapalindromic sequences together comprise a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; the DNA recognition sequence further comprises a core sequence of about 2 to 20 nucleotides, the core sequence being located between the first and second parapalindromic sequences; (ii) a heterologous sequence of interest; Eukaryotic cells (e.g., mammalian cells, e.g., human cells) including: (Item 4) 1. A method of modifying the genome of a eukaryotic cell (e.g., a mammalian cell, e.g., a human cell), comprising: a) a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or a nucleic acid encoding said recombinase polypeptide; and b) an insert DNA, (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), comprising a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the first and second parapalindromic sequences together comprising a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; The DNA recognition sequence further comprises a core sequence of about 2 to 20 nucleotides, a DNA recognition sequence located between the first and second parapalindromic sequences; (ii) a heterologous sequence of interest; Insert DNA containing thereby modifying the genome of said eukaryotic cell. (Item 5) 1. A method for inserting a heterologous sequence of interest into the genome of a eukaryotic cell (e.g., a mammalian cell, e.g., a human cell), comprising: a) a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or a nucleic acid encoding said polypeptide; and b) an insert DNA, (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), comprising a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the first and second parapalindromic sequences together comprising a parapalindromic region present within a nucleotide sequence in the left region or right region column of Table 2A, 2B, or 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; the DNA recognition sequence further comprises a core sequence of about 2 to 20 nucleotides, the core sequence being located between the first and second parapalindromic sequences; (ii) a heterologous sequence of interest; Insert DNA containing thereby inserting the heterologous sequence of interest into the genome of the eukaryotic cells at a frequency of, e.g., at least about 0.1% (e.g., at least about 0.1%, 0.5%, 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%) of a population of the eukaryotic cells, e.g., as measured in the assay of Example 5. (Item 6) An isolated recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. (Item 7) An isolated nucleic acid encoding a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. (Item 8) An isolated nucleic acid (e.g., DNA), (i) A DNA recognition sequence, comprising a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the first and second parapalindromic sequences, taken together, have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to a parapalindromic region present within the nucleotide sequence of the left region or right region column of Table 2A, 2B, or 2C, or to the parapalindromic region, or to no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or substitutions) thereto. a nucleotide sequence having a deletion, and the DNA recognition sequence further comprises a core sequence of about 2 to 20 nucleotides, the core sequence being located between the first and second parapalindromic sequences; (ii) a heterologous sequence of interest; An isolated nucleic acid (e.g., DNA) comprising: (Item 9) 1. A method for producing a recombinase polypeptide, comprising: a) providing a nucleic acid encoding a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and b) introducing said nucleic acid into a eukaryotic cell under conditions that allow the production of said recombinase polypeptide. thereby producing said recombinase polypeptide. (Item 10) 1. A method for generating an insert DNA comprising a DNA recognition sequence and a heterologous sequence, comprising: a) a nucleic acid, (i) a DNA recognition sequence that binds to a recombinase polypeptide comprising an amino acid sequence of Table 3A, 3B, or 3C, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, the DNA recognition sequence having a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 15 to 35 or 20 to 30 nucleotides, and the first and second parapalindromic sequences, together, are those of Table 2A, 2B, or or a parapalindromic region present within a nucleotide sequence in the left or right region of 2C, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic region, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and the DNA recognition sequence further comprises a core sequence of about 2 to 20 nucleotides, the core sequence being located between the first and second parapalindromic sequences; (ii) a heterologous sequence of interest; providing a nucleic acid comprising: b) introducing the nucleic acid into a cell (e.g., a eukaryotic or prokaryotic cell, as described herein) under conditions that allow replication of the nucleic acid; thereby producing said insert DNA. [Brief explanation of the drawings]

[0267] [Figure 1A] Activity of 10 exemplary serine integrases in human cells. HEK293T cells were transfected with an integrase expression plasmid and a template plasmid carrying a 520-bp attP-containing region followed by an EGFP reporter driven by a CMV promoter. The percentage of EGFP-positive cells observed by flow cytometry 21 days after transfection is shown. [Figure 1B] Strategies for assessing integration, stability, and expression of different AAV donor formats. Single attB* or attP* donors utilize the formation of double-stranded circularized DNA after AAV transduction into the cell nucleus. This construct also contains ITR sequences after integration. Dual attB-attB* or attP-attP* donors do not require the formation of double-stranded circularized DNA after AAV transduction. Readouts for integration stability and expression use droplet digital PCR (ddPCR) and flow cytometry (FLOW). [Figure 2]Description of AAV constructs. Line 1 shows ITR, stuffer (500), attP*, PEF1a, EGFP, WPRE, hGHpA, ITR; AAV2 serotype. Line 2 shows ITR, stuffer (500), attP, PEF1a, EGFP, WPRE, hGHpA, attP*, stuffer (500), ITR; AAV2 serotype. Line 3 shows ITR, stuffer (500), attB*, PEF1a, EGFP, WPRE, hGHpA, ITR; AAV2 serotype. Line 4 shows ITR, stuffer (500), attB, PEF1a, EGFP, WPRE, hGHpA, attB*, stuffer (500), ITR; AAV2 serotype. Row 5 shows ITR, PEF1a, hcoBXB1, WPRE, hGHpA, ITR; AAV2 serotype. Row 6 shows ITR, PEF1a, mcoBXB1, WPRE, hGHpA, ITR; AAV6 serotype. [Figure 3] Dual AAV delivery of serine integrase and template DNA into mammalian cells. (A) Schematic of the experiment. BXB1 serine recombinase and template DNA are co-delivered into the BXB landing pad cell line as separate AAV viral vectors. (B) Droplet digital PCR (ddPCR) assay to assess BXB1 serine recombinase and transgene integration (%CNV / landing pad) into the attP-attP* landing pad cell line 3 and 7 days after transduction. Black dots (to the right of each pair of gray dots) represent template-only samples, denoted at 0% on the y-axis. Gray dots (to the left of each pair of black dots) represent template + BXB1 integrase samples, denoted at 1-6% on the y-axis. [Figure 4]BXB1 integrase mRNA delivery and AAV delivery of template DNA into mammalian cells. (A) Schematic of the experiment. BXB1 serine recombinase mRNA delivery and AAV delivery of template DNA into a BXB1 landing pad cell line. (B) Droplet digital PCR (ddPCR) assay to assess BXB1 serine recombinase and transgene integration (% CNV / landing pad) into an attP-attP* landing pad cell line 3 days after mRNA transfection / AAV transduction. Black dots (to the right of each pair of gray dots) indicate template-only samples and drop to 0% on the y-axis. Gray dots (to the left of each pair of black dots) indicate template + BXB1 integrase and drop to >0% on the y-axis. [Figure 5A]General Structure of Recombinase Recognition Sites and Presence of Recognition Sites in the Left-Hand and Right-Hand Region Sequences Disclosed herein. (A) General Characteristics of Recognition Sequences. Serine recombinases as defined herein generally contain a central dinucleotide, a core sequence, and flanking arms that may be parapalindromic in nature. The attP and attB recognition sequences for Bxb1 recombinase (Table 3A, line 204) are shown herein. These sequences share a central dinucleotide, shown in bold, that is important for successful recombination between the two sites. The arms of the recognition site, indicated by a black outline, may share palindromic sequences to varying degrees and are therefore referred to herein as "parapalindromic." Nucleotides that are palindromic with respect to the opposite arm are indicated by underlined letters. Additionally, the recognition sequences share a common core between the attP and attB sites, shown here by gray shading. The core sequence includes at least the central dinucleotide but may contain additional sequences. (B) The left-hand or right-hand regions of Table 2 contain the attP site for the cognate recombinase. Table 2 includes exemplary recognition sites for exemplary recombinases described herein. As an example, the attP site for a recombinase of Table 1 or Table 3, e.g., Table 1A or Table 3A, is found in the left-hand or right-hand region of Table 2, e.g., Table 2A. Here, the attP site for Bxb1 integrase (Tables 1A and 3A, line 204) is shown in the corresponding line of Table 2A (line 204). The attP site for Bxb1 is shown as underlined and bold letters in the left-hand region sequence. [Figure 5B]General Structure of Recombinase Recognition Sites and Presence of Recognition Sites in the Left-Hand and Right-Hand Region Sequences Disclosed herein. (A) General Characteristics of Recognition Sequences. Serine recombinases as defined herein generally contain a central dinucleotide, a core sequence, and flanking arms that may be parapalindromic in nature. The attP and attB recognition sequences for Bxb1 recombinase (Table 3A, line 204) are shown herein. These sequences share a central dinucleotide, shown in bold, that is important for successful recombination between the two sites. The arms of the recognition site, indicated by a black outline, may share palindromic sequences to varying degrees and are therefore referred to herein as "parapalindromic." Nucleotides that are palindromic with respect to the opposite arm are indicated by underlined letters. Additionally, the recognition sequences share a common core between the attP and attB sites, shown here by gray shading. The core sequence includes at least the central dinucleotide but may contain additional sequences. (B) The left-hand or right-hand regions of Table 2 contain the attP site for the cognate recombinase. Table 2 includes exemplary recognition sites for exemplary recombinases described herein. As an example, the attP site for a recombinase of Table 1 or Table 3, e.g., Table 1A or Table 3A, is found in the left-hand or right-hand region of Table 2, e.g., Table 2A. Here, the attP site for Bxb1 integrase (Tables 1A and 3A, line 204) is shown in the corresponding line of Table 2A (line 204). The attP site for Bxb1 is shown as underlined and bold letters in the left-hand region sequence. DETAILED DESCRIPTION OF THE INVENTION

[0268] The present disclosure relates to compositions, systems, and methods for targeting, editing, modifying, or manipulating DNA sequences (e.g., inserting a heterologous DNA sequence of interest at a target site in a mammalian genome), for example, at one or more locations within the DNA sequence in a cell, tissue, or subject, in vivo or in vitro. The DNA sequence of interest may include, for example, a coding sequence, a regulatory sequence, or a gene expression unit.

[0269] Gene-writer™ Gene Editor The present invention provides recombinase polypeptides (e.g., serine recombinase polypeptides, e.g., as listed in Tables 3A, 3B, or 3C) that can modify or manipulate DNA sequences, e.g., by recombining two DNA sequences that contain cognate recognition sequences to which the recombinase polypeptide can bind. The Gene Writer™ gene editor system, in some embodiments, may include: (A) a polypeptide or a nucleic acid encoding a polypeptide (wherein the polypeptide comprises (i) a domain having recombinase activity and (ii) a domain having DNA-binding functionality (e.g., a recognition sequence described herein, e.g., that binds to or is capable of binding to such a recognition sequence); and (B) an insert DNA comprising (i) a sequence that binds to the polypeptide (e.g., a recognition sequence described herein) and, optionally, (ii) a sequence of interest (e.g., a heterologous sequence of interest). In some embodiments, the domain having recombinase activity and the domain having DNA-binding functionality are the same domain. For example, a Gene Writer genome editor protein may comprise a DNA-binding domain and a recombinase domain. In certain embodiments, elements of a Gene Writer™ gene editor polypeptide can be derived from the sequences of recombinase polypeptides (e.g., serine recombinases), e.g., as described herein, e.g., listed in Tables 3A, 3B, or 3C. In some embodiments, a Gene Writer™ gene editor polypeptide may comprise a sequence of a recombinase polypeptide (e.g., a serine recombinase), e.g., as described herein. The Writer genome editor is combined with a second polypeptide. In some embodiments, the second polypeptide is derived from a recombinase polypeptide (e.g., a serine recombinase), e.g., as described herein, e.g., listed in Table 3A, 3B, or 3C.

[0270] Recombinase polypeptide component of the GeneWriter gene editor system An exemplary family of recombinase polypeptides that can be used in the systems, cells, and methods described herein includes serine recombinases. Generally, serine recombinases are enzymes that catalyze site-specific recombination between two recognition sequences. The two recognition sequences can be, for example, on the same nucleic acid (e.g., DNA) molecule or on two separate nucleic acid (e.g., DNA) molecules. In some embodiments, a serine recombinase polypeptide comprises a recombinase N-terminal domain (also referred to as a catalytic domain), a recombinase domain, and a C-terminal zinc ribbon domain. In some embodiments, the zinc ribbon domain further comprises a coiled-coil motif. In some embodiments, the recombinase domain and zinc ribbon domain are collectively referred to as a C-terminal domain. In some embodiments, the N-terminal domain is 50 to 250 amino acids, or 100 to 200 amino acids, or 130 to 170 amino acids. In some embodiments, the C-terminal domain is 200 to 800 amino acids or 300 to 500 amino acids. In some embodiments, the recombinase domain is 50-150 amino acids. In some embodiments, the zinc ribbon domain is 30-100 amino acids. In some embodiments, the N-terminal domain is connected to the recombinase domain via a long helix (sometimes referred to as an αE helix or linker). In some embodiments, the recombinase domain and zinc ribbon domain are linked via a short linker. Non-limiting examples of serine recombinases and recombinase polypeptides are listed in Tables 3A, 3B, or 3C.

[0271] In some embodiments, the recombinant recombinase is constructed with a swapping domain. In some embodiments, the recombinase N-terminal domain can be paired with a heterologous recombinase C-terminal domain. In some embodiments, the catalytic domain can be paired with a heterologous recombinase domain, a zinc ribbon domain, an αE helix, and / or a short linker. In some embodiments, the C-terminal domain can comprise a heterologous recombinase domain, a zinc ribbon domain, an αE helix, and / or a short linker. In some embodiments, the DNA-binding element of the recombinase polypeptide is modified or replaced with a heterologous DNA-binding element, such as a zinc finger domain, a TAL domain, or a Watson-Crick-based targeting domain, e.g., a CRISPR / Cas system.

[0272] Without intending to be bound by any particular theory, serine recombinases utilize short, specific DNA sequences (e.g., attP and attB), which are exemplary recognition sequences. During the integration reaction, the recombinase binds to attP and attB as a dimer, mediates the association of the sites to form a tetrameric synaptic complex, and catalyzes strand exchange to integrate DNA and form new recognition sequence sites, attL and attR. The new recognition sites, attL and attR, comprise, for example, attB5'-core-attP3' and attP5'-core-attB3' in 5' to 3' order. Without intending to be bound by any particular theory, the reverse reaction, which excises DNA by site-specific recombination between the attL and attR sequences, occurs at a reduced frequency or does not occur in the absence of recombination directionality factors (RDFs). This results in stable integration with little or no detectable recombinase-mediated excision, i.e., "unidirectional" recombination.

[0273] While not being bound by the description of the mechanism, strand exchange catalyzed by recombinases typically occurs in two steps: (1) cleavage and (2) religation, which involve a covalent protein-DNA intermediate formed between the recombinase enzyme and the DNA strand. Recombinases act by binding to their DNA substrates as dimers and bringing sites together through protein-protein interactions to form a tetrameric synaptic complex. Activation of the nucleophilic serine in each of the four subunits results in DNA cleavage, generating a 2-nt 3' overhang and a transient phosphoseryl bond to the recessed 5' end. DNA strand exchange occurs through subunit rotation. The recessed 5' base and the 3' dinucleotide overhang base pair with the 3' OH, attacking the phosphoseryl bond in the reverse of the cleavage reaction to join the recombination half-sites. Further details of the structure, activity, and biology of serine recombinases are provided in the following references, which are incorporated by reference: Smith MCM. 2014. Phage-encoded serine integrases and other large Serine recombinases. Microbiol Spectrum 3(4):MDNA3-0059-2014; Rutherford K and Van Duyne G D. 2014. The ins and outs of serine integrase site-specific recombination. Current Opinion in Structural Biology 24:125-131; Van Duyne GD and Rutherford K. 2013. Large serine recombinase domain structure and attachment site binding. Critical Reviews in Biochemistry and Molecular Biology 48(5):471-491.

[0274] One skilled in the art can determine the nucleic acid and corresponding polypeptide sequences of recombinase polypeptides (e.g., serine recombinases) and their domains by using routine sequence analysis tools, such as, for example, the Basic Local Alignment Search Tool (BLAST) or CD-Search for conserved domain analysis. Other sequence analysis tools are well known and can be found, for example, at https: / / molbiol-tools.ca, e.g., http: / / molbiol-tools.ca / Motifs.htm. In some embodiments, the serine recombinases described herein comprise at least one known active site signature of a serine recombinase, e.g., cd00338, cd03767, cd03768, cd03769, or cd03770. Proteins containing these domains can also be found by searching for the domains in protein databases, such as InterPro (Mitchell et al. Nucleic Acids Res 47, D351-360 (2019)), UniProt (The UniProt Consortium Nucleic Acids Res 47, D506-515 (2019)), or conserved domain databases (Lu et al. Nucleic Acids Res 48, D265-268 (2020)), or by scanning the open reading framework or full-frame translation of nucleic acid sequences for serine recombinase domains using prediction tools, such as InterProScan.

[0275] While the present disclosure provides many specific serine recombinase sequences, it is understood that the methods described herein can also be practiced with other serine recombinases. For example, a composition or method described herein can include a serine recombinase having an active site signature selected from, e.g., cd00338, cd03767, cd03768, cd03769, or cd03770. In some embodiments, the serine recombinase is greater than 400 amino acids in length (e.g., at least 400, 500, 600, 700, 800, 900, or 1000 amino acids). In some embodiments, the recombinase includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more domains listed in any of Tables 3A-3C (e.g., listed in a single row of any of Tables 3A-3C). In some embodiments, the recombinase comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more domains listed in Table 4. In some embodiments, a method of identifying a recombinase comprises determining whether a polypeptide comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more domains listed in any of Tables 3A-3C (e.g., listed in a single row of any of 3A-3C). In some embodiments, a method of identifying a recombinase comprises determining whether a polypeptide comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more domains listed in Table 4.

[0276] Exemplary Recombinase Polypeptides In some embodiments, a Gene Writer™ gene editor system includes a recombinase polypeptide (e.g., a serine recombinase polypeptide), e.g., as described herein. Generally, a recombinase polypeptide (e.g., a serine recombinase polypeptide) specifically binds to a nucleic acid recognition sequence and catalyzes a recombination reaction at a site within the recognition sequence (e.g., a core sequence within the recognition sequence). In some embodiments, the recombinase polypeptide catalyzes recombination between the recognition sequence, or a portion thereof (e.g., its core sequence), and another nucleic acid sequence (e.g., an insert DNA that includes a cognate recognition sequence and, optionally, a sequence of interest, e.g., a heterologous sequence of interest). For example, the recombinase polypeptide (e.g., a serine recombinase polypeptide) catalyzes a recombination reaction that results in the insertion of a sequence of interest, or a portion thereof, into another nucleic acid molecule (e.g., a genomic DNA molecule, e.g., a chromosomal or mitochondrial DNA).

[0277] Tables 3A, 3B, or 3C below (see the Protseq column) provide the amino acid sequences of exemplary recombinase polypeptides, such as serine recombinases (e.g., serine integrases) or fragments thereof. Tables 2A, 2B, or 2C provide flanking nucleic acid sequences of nucleic acid sequences encoding exemplary serine recombinases of an organism of origin (see the columns labeled Left Region and Right Region, respectively); one or both of these flanking nucleic acid sequences contain the native recognition sequence or a portion thereof of the corresponding recombinase (e.g., containing an attP site or a portion thereof). Tables 3A, 3B, or 3C contain amino acid sequences not previously identified as serine recombinases, and Tables 2A, 2B, or 2C contain the corresponding flanking nucleic acid sequences (and thereby DNA recognition sequences) of serine recombinases whose DNA recognition sequences were previously unknown. Listed below are descriptions of the sequence of origin (see the description column in Table 1A, 1B, or 1C), the organism of origin of the recombinase (see the organism column in Table 1A, 1B, or 1C), the length of the amino acid sequence of the recombinase (see the protein sequence length column in Table 1A, 1B, or 1C), the genome accession number of the nucleic acid sequence encoding the recombinase (see the genome accession column in Table 1A, 1B, or 1C), the protein accession number of the recombinase (see the protein accession column in Table 1A, 1B, or 1C), and the genomic location coordinates of the recombinase-encoding sequence (including flanking nucleic acid sequences as indicated) (see the G start and G end columns in Table 1A, 1B, or 1C). Additionally, domains identified as present in exemplary recombinase sequences are identified based on InterPro analysis of the amino acid sequences (see the domain column in Table 3A, 3B, or 3C). See, e.g., http: / / omictools.com / interpro-tool. A brief key to domain nomenclature is provided in Table 4. The amino acid and genomic sequences of each accession number in Tables 1A, 1B, or 1C are incorporated herein by reference in their entirety.Each of the native recognition sequences or portions thereof present in a flanking nucleic acid sequence listed in Table 2A, 2B, or 2C may include one, two, or three of: (i) a first parapalindromic sequence, (ii) a core sequence, and / or (iii) a second parapalindromic sequence, wherein the first and second parapalindromic sequences are parapalindromic with respect to each other.

[0278] In some embodiments, when selecting a pair of parapalindromic sequences, a user of the Tables disclosed herein selects each sequence based on the sequences disclosed in the rows having the same row number as each other. For example, in some embodiments, a cell comprising a DNA recognition sequence comprising a first parapalindromic sequence and a second parapalindromic sequence comprises first and second parapalindromic sequences that are related to sequences disclosed in the same row of Table 2A, 2B, or 2C. In some embodiments, when selecting a DNA recognition sequence (e.g., a parapalindromic sequence) for use with an exemplary recombinase polypeptide, the DNA recognition sequence (e.g., a parapalindromic sequence) is selected from or related to a sequence in the row having the same row number as the exemplary recombinase polypeptide.

[0279] [Table 1]

[0280] [Table 2]

[0281] [Table 3]

[0282] [Table 4]

[0283] [Table 5]

[0284] Table 6

[0285] Table 7

[0286] Table 8

[0287] Table 9

[0288] Table 10

[0289] Table 11

[0290] Table 12

[0291] Table 13

[0292] Table 14

[0293] Table 15

[0294] Table 16

[0295] Table 17

[0296] Table 18

[0297] Table 19

[0298] Table 20

[0299] Table 21

[0300] Table 22

[0301] Table 23

[0302] Table 24

[0303] Table 25

[0304] Table 26

[0305] Table 27

[0306] Table 28

[0307] Table 29

[0308] Table 30

[0309] Table 31

[0310] Table 32

[0311] Table 33

[0312] Table 34

[0313] Table 35

[0314] Table 36

[0315] Table 37

[0316] Table 38

[0317] Table 39

[0318] Table 40

[0319] Table 41

[0320] Table 42

[0321] Table 43

[0322] Table 44

[0323] Table 45

[0324] Table 46

[0325] Table 47

[0326] Table 48

[0327] Table 49

[0328] Table 50

[0329] Table 51

[0330] Table 52

[0331] Table 53

[0332] Table 54

[0333] Table 55

[0334] Table 56

[0335] Table 57

[0336] Table 58

[0337] Table 59

[0338] Table 60

[0339] Table 61

[0340] Table 62

[0341] Table 63

[0342] Table 64

[0343] Table 65

[0344] Table 66

[0345] Table 67

[0346] Table 68

[0347] Table 69

[0348] Table 70

[0349] Table 71

[0350] Table 72

[0351] Table 73

[0352] Table 74

[0353] Table 75

[0354] Table 76

[0355] Table 77

[0356] Table 78

[0357] Table 79

[0358] Table 80

[0359] Table 81

[0360] Table 82

[0361] Table 83

[0362] Table 84

[0363] Table 85

[0364] Table 86

[0365] Table 87

[0366] Table 88

[0367] Table 89

[0368] Table 90

[0369] Table 91

[0370] Table 92

[0371] Table 93

[0372] Table 94

[0373] Table 95

[0374] Table 96

[0375] Table 97

[0376] Table 98

[0377] Table 99

[0378] Table 100

[0379] Table 101

[0380] Table 102

[0381] Table 103

[0382] Table 104

[0383] Table 105

[0384] Table 106

[0385] Table 107

[0386] Table 108

[0387] Table 109

[0388] Table 110

[0389] Table 111

[0390] Table 112

[0391] Table 113

[0392] Table 114

[0393] Table 115

[0394] Table 116

[0395] Table 117

[0396] Table 118

[0397] Table 119

[0398] Table 120

[0399] Table 121

[0400] Table 122

[0401] Table 123

[0402] Table 124

[0403] Table 125

[0404] Table 126

[0405] Table 127

[0406] Table 128

[0407] Table 129

[0408] Table 130

[0409] Table 131

[0410] Table 132

[0411] Table 133

[0412] Table 134

[0413] Table 135

[0414] Table 136

[0415] Table 137

[0416] Table 138

[0417] Table 139

[0418] Table 140

[0419] Table 141

[0420] Table 142

[0421] Table 143

[0422] Table 144

[0423] Table 145

[0424] Table 146

[0425] Table 147

[0426] Table 148

[0427] Table 149

[0428] Table 150

[0429] Table 151

[0430] Table 152

[0431] Table 153

[0432] Table 154

[0433] Table 155

[0434] Table 156

[0435] Table 157

[0436] Table 158

[0437] Table 159

[0438] Table 160

[0439] Table 161

[0440] Table 162

[0441] Table 163

[0442] Table 164

[0443] Table 165

[0444] Table 166

[0445] Table 167

[0446] Table 168

[0447] Table 169

[0448] Table 170

[0449] Table 171

[0450] Table 172

[0451] Table 173

[0452] Table 174

[0453] Table 175

[0454] Table 176

[0455] Table 177

[0456] Table 178

[0457] Table 179

[0458] Table 180

[0459] Table 181

[0460] Table 182

[0461] Table 183

[0462] Table 184

[0463] Table 185

[0464] Table 186

[0465] Table 187

[0466] Table 188

[0467] Table 189

[0468] Table 190

[0469] Table 191

[0470] Table 192

[0471] Table 193

[0472] Table 194

[0473] Table 195

[0474] Table 196

[0475] Table 197

[0476] Table 198

[0477] Table 199

[0478] Table 200

[0479] Table 201

[0480] Table 202

[0481] Table 203

[0482] Table 204

[0483] Table 205

[0484] Table 206

[0485] Table 207

[0486] Table 208

[0487] Table 209

[0488] Table 210

[0489] Table 211

[0490] Table 212

[0491] Table 213

[0492] Table 214

[0493] Table 215

[0494] Table 216

[0495] Table 217

[0496] Table 218

[0497] Table 219

[0498] Table 220

[0499] Table 221

[0500] Table 222

[0501] Table 223

[0502] Table 224

[0503] Table 225

[0504] Table 226

[0505] Table 227

[0506] Table 228

[0507] Table 229

[0508] Table 230

[0509] Table 231

[0510] Table 232

[0511] Table 233

[0512] Table 234

[0513] Table 235

[0514] Table 236

[0515] Table 237

[0516] Table 238

[0517] Table 239

[0518] Table 240

[0519] Table 241

[0520] Table 242

[0521] Table 243

[0522] Table 244

[0523] Table 245

[0524] Table 246

[0525] Table 247

[0526] Table 248

[0527] Table 249

[0528] Table 250

[0529] Table 251

[0530] Table 252

[0531] Table 253

[0532] Table 254

[0533] Table 255

[0534] Table 256

[0535] Table 257

[0536] Table 258

[0537] Table 259

[0538] Table 260

[0539] Table 261

[0540] Table 262

[0541] Table 263

[0542] Table 264

[0543] Table 265

[0544] Table 266

[0545] Table 267

[0546] Table 268

[0547] Table 269

[0548] Table 270

[0549] Table 271

[0550] Table 272

[0551] Table 273

[0552] Table 274

[0553] Table 275

[0554] Table 276

[0555] Table 277

[0556] Table 278

[0557] Table 279

[0558] Table 280

[0559] Table 281

[0560] Table 282

[0561] Table 283

[0562] Table 284

[0563] Table 285

[0564] Table 286

[0565] Table 287

[0566] Table 288

[0567] Table 289

[0568] Table 290

[0569] Table 291

[0570] Table 292

[0571] Table 293

[0572] Table 294

[0573] Table 295

[0574] Table 296

[0575] Table 297

[0576] Table 298

[0577] Table 299

[0578] Table 300

[0579] Table 301

[0580] Table 302

[0581] Table 303

[0582] Table 304

[0583] Table 305

[0584] Table 306

[0585] Table 307

[0586] Table 308

[0587] Table 309

[0588] Table 310

[0589] Table 311

[0590] Table 312

[0591] Table 313

[0592] Table 314

[0593] Table 315

[0594] Table 316

[0595] Table 317

[0596] Table 318

[0597] Table 319

[0598] Table 320

[0599] Table 321

[0600] Table 322

[0601] Table 323

[0602] Table 324

[0603] Table 325

[0604] Table 326

[0605] Table 327

[0606] Table 328

[0607] Table 329

[0608] Table 330

[0609] Table 331

[0610] Table 332

[0611] Table 333

[0612] Table 334

[0613] Table 335

[0614] Table 336

[0615] Table 337

[0616] Table 338

[0617] Table 339

[0618] Table 340

[0619] Table 341

[0620] Table 342

[0621] Table 343

[0622] Table 344

[0623] Table 345

[0624] Table 346

[0625] Table 347

[0626] Table 348

[0627] Table 349

[0628] Table 350

[0629] Table 351

[0630] Table 352

[0631] Table 353

[0632] Table 354

[0633] Table 355

[0634] Table 356

[0635] Table 357

[0636] Table 358

[0637] Table 359

[0638]

Table 360

[0639] Table 361

[0640] Table 362

[0641] Table 363

[0642] Table 364

[0643] Table 365

[0644] Table 366

[0645] Table 367

[0646] Table 368

[0647] Table 369

[0648] Table 370

[0649] Table 371

[0650] Table 372

[0651] Table 373

[0652] Table 374

[0653] Table 375

[0654] Table 376

[0655] Table 377

[0656] Table 378

[0657] Table 379

[0658] Table 380

[0659] Table 381

[0660] Table 382

[0661] Table 383

[0662] Table 384

[0663] Table 385

[0664] Table 386

[0665] Table 387

[0666] Table 388

[0667] Table 389

[0668] Table 390

[0669] Table 391

[0670] Table 392

[0671] Table 393

[0672] Table 394

[0673] Table 395

[0674] Table 396

[0675] Table 397

[0676] Table 398

[0677] Table 399

[0678] Table 400

[0679] Table 401

[0680] Table 402

[0681] Table 403

[0682] Table 404

[0683] Table 405

[0684] Table 406

[0685] Table 407

[0686] Table 408

[0687] Table 409

[0688] Table 410

[0689] Table 411

[0690] Table 412

[0691] Table 413

[0692] Table 414

[0693] Table 415

[0694] Table 416

[0695] Table 417

[0696] Table 418

[0697] Table 419

[0698] Table 420

[0699] Table 421

[0700] Table 422

[0701] Table 423

[0702] Table 424

[0703] Table 425

[0704] Table 426

[0705] Table 427

[0706] Table 428

[0707] Table 429

[0708] Table 430

[0709] Table 431

[0710] Table 432

[0711] Table 433

[0712] Table 434

[0713] Table 435

[0714] Table 436

[0715] Table 437

[0716] Table 438

[0717] Table 439

[0718] Table 440

[0719] Table 441

[0720] Table 442

[0721] Table 443

[0722] Table 444

[0723] Table 445

[0724] Table 446

[0725] Table 447

[0726] Table 448

[0727] Table 449

[0728] Table 450

[0729] Table 451

[0730] Table 452

[0731] Table 453

[0732] Table 454

[0733] Table 455

[0734] Table 456

[0735] Table 457

[0736] Table 458

[0737] Table 459

[0738] Table 460

[0739] Table 461

[0740] Table 462

[0741] Table 463

[0742] Table 464

[0743] Table 465

[0744] Table 466

[0745] Table 467

[0746] Table 468

[0747] Table 469

[0748] Table 470

[0749] Table 471

[0750] Table 472

[0751] Table 473

[0752] Table 474

[0753] Table 475

[0754] Table 476

[0755] Table 477

[0756] Table 478

[0757] Table 479

[0758] Table 480

[0759] Table 481

[0760] Table 482

[0761] Table 483

[0762] Table 484

[0763] Table 485

[0764] Table 486

[0765] Table 487

[0766] Table 488

[0767] Table 489

[0768] Table 490

[0769] Table 491

[0770] Table 492

[0771] Table 493

[0772] Table 494

[0773] Table 495

[0774] Table 496

[0775] Table 497

[0776] Table 498

[0777] Table 499

[0778] Table 500

[0779] Table 501

[0780] Table 502

[0781] Table 503

[0782] Table 504

[0783] Table 505

[0784] Table 506

[0785] Table 507

[0786] Table 508

[0787] Table 509

[0788] Table 510

[0789] Table 511

[0790] Table 512

[0791] Table 513

[0792] Table 514

[0793] Table 515

[0794] Table 516

[0795] Table 517

[0796] Table 518

[0797] Table 519

[0798] Table 520

[0799] Table 521

[0800] Table 522

[0801] Table 523

[0802] Table 524

[0803] Table 525

[0804] Table 526

[0805] Table 527

[0806] Table 528

[0807] Table 529

[0808] Table 530

[0809] Table 531

[0810] Table 532

[0811] Table 533

[0812] Table 534

[0813] Table 535

[0814] Table 536

[0815] Table 537

[0816] Table 538

[0817] Table 539

[0818] Table 540

[0819] Table 541

[0820] Table 542

[0821] Table 543

[0822] Table 544

[0823] Table 545

[0824] Table 546

[0825] Table 547

[0826] Table 548

[0827] Table 549

[0828] Table 550

[0829] Table 551

[0830] Table 552

[0831] Table 553

[0832] Table 554

[0833] Table 555

[0834] Table 556

[0835] Table 557

[0836] Table 558

[0837] Table 559

[0838] Table 560

[0839] Table 561

[0840] Table 562

[0841] Table 563

[0842] Table 564

[0843] Table 565

[0844] Table 566

[0845] Table 567

[0846] Table 568

[0847] Table 569

[0848] Table 570

[0849] Table 571

[0850] Table 572

[0851] Table 573

[0852] Table 574

[0853] Table 575

[0854] Table 576

[0855] Table 577

[0856] Table 578

[0857] Table 579

[0858] Table 580

[0859] Table 581

[0860] Table 582

[0861] Table 583

[0862] Table 584

[0863] Table 585

[0864] Table 586

[0865] Table 587

[0866] Table 588

[0867] Table 589

[0868] Table 590

[0869] Table 591

[0870] Table 592

[0871] Table 593

[0872] Table 594

[0873] Table 595

[0874] Table 596

[0875] Table 597

[0876] Table 598

[0877] Table 599

[0878]

Table 600

[0879] Table 601

[0880] Table 602

[0881] Table 603

[0882] Table 604

[0883] Table 605

[0884] Table 606

[0885] Table 607

[0886] Table 608

[0887] Table 609

[0888] Table 610

[0889] Table 611

[0890] Table 612

[0891] Table 613

[0892] Table 614

[0893] Table 615

[0894] Table 616

[0895] Table 617

[0896] Table 618

[0897] Table 619

[0898] Table 620

[0899] Table 621

[0900] Table 622

[0901] Table 623

[0902] Table 624

[0903] Table 625

[0904] Table 626

[0905] Table 627

[0906] Table 628

[0907] Table 629

[0908] Table 630

[0909] Table 631

[0910] Table 632

[0911] Table 633

[0912] Table 634

[0913] Table 635

[0914] Table 636

[0915] Table 637

[0916] Table 638

[0917] Table 639

[0918] Table 640

[0919] Table 641

[0920] Table 642

[0921] Table 643

[0922] Table 644

[0923] Table 645

[0924] Table 646

[0925] Table 647

[0926] Table 648

[0927] Table 649

[0928] Table 650

[0929] Table 651

[0930] Table 652

[0931] Table 653

[0932] Table 654

[0933] Table 655

[0934] Table 656

[0935] Table 657

[0936] Table 658

[0937] Table 659

[0938] Table 660

[0939] Table 661

[0940] Table 662

[0941] Table 663

[0942] Table 664

[0943] Table 665

[0944] Table 666

[0945] Table 667

[0946] Table 668

[0947] Table 669

[0948] Table 670

[0949] Table 671

[0950] Table 672

[0951] Table 673

[0952] Table 674

[0953] Table 675

[0954] Table 676

[0955] Table 677

[0956] Table 678

[0957] Table 679

[0958] Table 680

[0959] Table 681

[0960] Table 682

[0961] Table 683

[0962] Table 684

[0963] Table 685

[0964] Table 686

[0965] Table 687

[0966] Table 688

[0967] Table 689

[0968] Table 690

[0969] Table 691

[0970] Table 692

[0971] Table 693

[0972] Table 694

[0973] Table 695

[0974] Table 696

[0975] Table 697

[0976] Table 698

[0977] Table 699

[0978] Table 700

[0979] Table 701

[0980] Table 702

[0981] Table 703

[0982] Table 704

[0983] Table 705

[0984] Table 706

[0985] Table 707

[0986] Table 708

[0987] Table 709

[0988] Table 710

[0989] Table 711

[0990] Table 712

[0991] Table 713

[0992] Table 714

[0993] Table 715

[0994] Table 716

[0995] Table 717

[0996] Table 718

[0997] Table 719

[0998] Table 720

[0999] Table 721

[1000] Table 722

[1001] Table 723

[1002] Table 724

[1003] Table 725

[1004] Table 726

[1005] [Table 727]

[1006] [Table 728]

[1007] [Table 729]

[1008] [Table 730]

[1009] [Table 731]

[1010] [Table 732]

[1011] [Table 733]

[1012] [Table 734]

[1013] [Table 735]

[1014] In some embodiments, a sequence comprising the left region nucleic acid sequence of line 329 of Table 2A (e.g., a sequence comprising the nucleic acid sequence of SEQ ID NO: 290) is the nucleic acid sequence: [ka] Includes.

[1015] In some embodiments, a sequence comprising the left region nucleic acid sequence of line 524 of Table 2A (e.g., a sequence comprising the nucleic acid sequence of SEQ ID NO: 470) is the nucleic acid sequence: [ka] Includes.

[1016] In some embodiments, a recombinase recognition site (e.g., as described herein) comprises an attB sequence. In some embodiments, a recombinase recognition site (e.g., as described herein) comprises an attP sequence. In some embodiments, a recombinase recognition site (e.g., as described herein) comprises an attB sequence and an attP sequence. In several embodiments, the attB sequence is selected from the sequences listed in Table 4X. In several embodiments, the attP sequence is selected from the sequences listed in Table 4X. In some embodiments, a recombinase recognition site (e.g., as described herein) comprises an attB sequence and an attP sequence, wherein the attB and attP sequences each comprise a sequence listed in a single row of Table 4X.

[1017] In some embodiments, the DNA recognition sequence (e.g., as described herein) comprises an attB sequence. In some embodiments, the DNA recognition sequence (e.g., as described herein) comprises an attP sequence. In some embodiments, the DNA recognition sequence (e.g., as described herein) comprises an attB sequence and an attP sequence. In several embodiments, the attB sequence is selected from the sequences listed in Table 4X. In several embodiments, the attP sequence is selected from the sequences listed in Table 4X. In some embodiments, the DNA recognition sequence (e.g., as described herein) comprises an attB sequence and an attP sequence, wherein the attB and attP sequences each comprise a sequence listed in a single row of Table 4X.

[1018] In some embodiments, the recombinase polypeptide (e.g., comprised in a system or cell described herein) comprises an amino acid sequence listed in Table 3A, 3B, or 3C, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity thereto. In some embodiments, a recombinase polypeptide (e.g., included in a system or cell described herein) or portion thereof has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the amino acid sequence of a recombinase domain, a DNA recognition domain (e.g., that binds or is capable of binding to a recognition site, e.g., as described herein), a recombinase N-terminal domain (also called a catalytic domain), a zinc ribbon domain, a coiled-coil motif of a zinc ribbon domain, or a C-terminal domain (e.g., a recombinase domain and a zinc ribbon domain) of a recombinase polypeptide listed in Table 3A, 3B or 3C. In some embodiments, the recombinase polypeptide (e.g., comprised in a system or cell described herein) has one or more of the DNA binding activity and / or recombinase activity of a recombinase polypeptide comprising an amino acid sequence listed in Table 3A, 3B, or 3C, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity thereto.

[1019] In some embodiments, the insert DNA (e.g., contained in a system or cell described herein) comprises a nucleic acid recognition sequence present within the nucleotide sequence of the left region or right region column of Table 2A, 2B, or 2C, or a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto. In some embodiments, the insert DNA (e.g., contained in a system or cell described herein) comprises one or more (e.g., both) parapalindromic sequences present within the nucleotide sequence of the left region or right region column of Table 2A, 2B, or 2C, or a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to said parapalindromic sequences, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto. In some embodiments, the insert DNA (e.g., contained in a system or cell described herein) comprises a spacer (e.g., core sequence) of a nucleic acid recognition sequence present within the nucleotide sequence of the left region or right region column of Table 2A, 2B, or 2C, or a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto. In certain embodiments, the insert DNA further comprises a heterologous sequence of interest.

[1020] In some embodiments, the insert DNA (e.g., contained in a system or cell described herein) comprises a nucleic acid recognition sequence present within the nucleotide sequence of the left region or right region column of Table 2A, 2B, or 2C, or a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto, which is cognate to a pseudo-recognition sequence (e.g., a human recognition sequence).

[1021] In some embodiments, the insert DNA or recombinase polypeptide used in the compositions or methods described herein directs the insertion of a heterologous sequence of interest into a position having a safe harbor score of at least 3, 4, 5, 6, 7, or 8.

[1022] In certain embodiments, recombination between the insert DNA and the human DNA recognition sequence results in the formation of an integrated nucleic acid molecule comprising two recognition sequences flanking the integrated sequence (e.g., a heterologous sequence of interest). Without intending to be bound by any particular theory, serine recombinase promotes recombination between the recognition sequences comprising attB and attP sites, resulting in the formation of, for example, a recognition sequence comprising attL and attR sites flanking the integrated sequence. While serine recombinase can recognize and, e.g., bind to, attL or attR sites, serine recombinase does not significantly promote (e.g., does not promote) recombination with attL or attR sites (e.g., in the absence of additional factors). The attL and attR sites comprise the recombined portions of the attP and attB sites from which they were created. In certain embodiments, one or both of the two post-recombination recognition sequences of the integrated nucleic acid molecule contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 or more mismatches compared to one or more (e.g., one, two, or all three) of: (i) the native recognition sequence; (ii) the recognition sequence on the inserted DNA; and / or (iii) the pseudo-recognition sequence (e.g., a human DNA recognition sequence). In some embodiments, one or both of the two post-recombination recognition sequences of the integrated nucleic acid molecule contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 or more mismatches compared to the native recognition sequence. In some embodiments, the mismatches are in the core sequence.In some embodiments, it is contemplated that these differences between the recognition sequence of the integrated nucleic acid molecule and the native recognition sequence, the inserted DNA recognition sequence, and / or the human DNA recognition sequence result in a reduced binding affinity between the recombinase polypeptide and the recognition sequence of the integrated nucleic acid molecule and / or a reduced (e.g., eliminated) recombinase activity of the recombinase polypeptide towards the recognition sequence of the integrated nucleic acid molecule compared to the binding and / or activity of the recombinase towards the recognition sequence, the native recognition sequence, the inserted DNA recognition sequence, and / or the human DNA recognition sequence.

[1023] In some embodiments, the pseudo-recognition sequence (e.g., a human DNA recognition sequence) is located at or near (e.g., within 1, 2, 3, 4, 5, 10, 15, 20, 30, 40, 50, 75, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, or 10,000 nucleotides of) a genomic safe harbor site. In some embodiments, the pseudo recognition sequence (e.g., a human recognition sequence) is located at a location in the genome that meets one, two, three, four, five, six, seven, eight, or nine of the following criteria: (i) located >300 kb from a cancer-associated gene; (ii) >300 kb from an miRNA / other functional small RNA; (iii) >50 kb from the 5' end of the gene; (iv) >50 kb from the origin of replication; (v) >50 kb from any ultraconserved element; (vi) has low transcriptional activity (i.e., no mRNA + / - 25 kb); (vii) is not within a variable copy number region; (viii) is in open chromatin; and / or (ix) is unique with one copy in the human genome.

[1024] In embodiments, the cells or systems described herein comprise one or more (e.g., one, two, or three) of the following: (i) a recombinase polypeptide listed in a row having row number X of Table 3A, 3B, or 3C, or 3B, where X is number 1 through the highest row number of any of Tables 3A, 3B, or 3C, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity thereto; (ii) an insert DNA comprising a DNA recognition sequence present within a nucleotide sequence in the left region or right region column of a row having row number X of Table 2A, 2B, or 2C, or an insert DNA having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. and / or (iii) a pseudo-recognition sequence present in a sequence in the left-hand region or right-hand region column of Table 2A, 2B, or 2C. A genome containing a sequence (e.g., a human recognition sequence) or a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 sequence alterations (e.g., substitutions, insertions, or deletions) thereto.

[1025] In some embodiments, recombinase recognition sites, e.g., attB, attP, attL, or attR sites, can be predicted by available software tools. In some embodiments, recognition sites can be predicted using phage prediction tools, e.g., PhiSpy (Akhter et al. Nucleic Acids Res 40(16):e126(2012)) or PHASTER (Arndt et al. Nucleic Acids Res Res 44:W16-W21 (2016) (incorporated herein by reference). In some embodiments, the flanking regions of the integrase coding sequence in a natural context, e.g., a bacteriophage genome, a plasmid, or a bacterial genome, e.g., the left or right regions of Table 2A, 2B, or 2C, comprise the native attachment site for the recombinase enzyme. In some embodiments, the minimal attachment site can be discovered experimentally by testing fragments of the integrase flanking sequences, e.g., the left or right regions of Table 2A, 2B, or 2C, until a minimal sequence sufficient for a productive recombination reaction is found. In some embodiments, the integrase flanking sequences, e.g., the left or right regions of Table 2A, 2B, or 2C, or fragments thereof, are assayed to determine the importance of each nucleotide and profiled in a library format, e.g., according to the method of Bessen et al. Nat Commun 10:1937 (2019) (incorporated herein by reference in its entirety). In some embodiments, the recombinase or recombinase recognition site is selected through an evolutionary process for modified protein-nucleic acid interaction properties, e.g., evolving the recombinase used in the Gene Writer system as described in WO2017015545, which is incorporated by reference in its entirety. In some embodiments, the recombinase and / or recombinase recognition site is discovered through prediction of elements that will integrate into the native host genome, e.g., the ends of an integrating bacteriophage or an integrating plasmid, as described in Yang et al. Nat Methods 11(12):1261-1266 (2014), which is incorporated by reference in its entirety.

[1026] In some embodiments, attL or attR sites are present in the human genome, and the template DNA includes a cognate site; for example, the template includes an attR sequence if the genome includes an attL sequence. In some embodiments, when attL / R recognition sites are used in the Gene Writing System, the system also includes recombination directionality factors (RDFs) to enable recognition and recombination of those sites. In some embodiments, the Gene Writer polypeptide and the cognate RDF are provided as a fusion polypeptide. Exemplary recombinase-RDF fusions are described in Olorunniji et al. Nucleic Acids Res 45(14):8635-8645 (2017), the entire contents of which are incorporated herein by reference.

[1027] In some embodiments, the protein components of the Gene Writing™ system described herein can be pre-associated with a template (e.g., a DNA template). For example, in some embodiments, the Gene Writer™ polypeptide can first be combined with a DNA template to form a deoxyribonucleoprotein (DNP) complex. In some embodiments, DNP can be delivered to cells via transfection, nucleofection, viruses, vesicles, LNPs, exosomes, or fusosomes. In some embodiments, the template DNA can first be associated with a DNA bending factor, such as HMGB1, to promote excision and transposition when subsequently contacted with the transposase components. Additional description of DNP delivery can be found, for example, in Guha and Calos J Mol Biol (2020), the entire contents of which are incorporated herein by reference.

[1028] In some embodiments, the polypeptides described herein comprise one or more (e.g., 2, 3, 4, 5) nuclear targeting sequences, e.g., nuclear localization sequences (NLSs). In some embodiments, the NLS is a bipartite NLS. In some embodiments, the NLS promotes import of a protein comprising the NLS into a cell nucleus. In some embodiments, the NLS is fused to the N-terminus of a Gene Writer described herein. In some embodiments, the NLS is fused to the C-terminus of a Gene Writer. In some embodiments, the NLS is fused to the N-terminus or C-terminus of a Cas domain. In some embodiments, a linker sequence is disposed between the NLS and an adjacent domain of the Gene Writer.

[1029] In some embodiments, the NLS comprises the amino acid sequence: MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 3432), PKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 3433), RKSGKIAAIWKRPRKPKKKRKV KRTADGSEFESPKKKRKV (SEQ ID NO: 3434), KKTELQTTNAENKTKKL (SEQ ID NO: 3435), or KRGINDRNFWRGENGRKTR (SEQ ID NO: 3436), KRPAATKKAGQAKKKK (SEQ ID NO: 3437), or a functional fragment or variant thereof. Exemplary NLS sequences are also described in PCT / EP 2000 / 011690, the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences.

[1030] In some embodiments, the NLS is a bipartite NLS. A bipartite NLS typically includes two basic amino acid clusters (e.g., about 10 amino acids in length) separated by a spacer sequence. A monopartite NLS typically lacks a spacer. An example of a bipartite NLS is the nucleoplasmin NLS, which has the sequence KR[PAATKKAGQA]KKKK (SEQ ID NO: 3437) (the spacer is in parentheses). Another exemplary bipartite NLS has the sequence: PKKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 3438). Exemplary NLSs are described in WO2020051561 (which is incorporated by reference in its entirety, including the disclosure regarding nuclear localization sequences).

[1031] DNA-binding domain In some embodiments, a recombinase polypeptide (eg, included in a system or cell described herein), e.g., a tyrosine recombinase, comprises a DNA-binding domain (e.g., a target-binding domain or a template-binding domain).

[1032] In some embodiments, the recombinase polypeptides described herein can be redirected to defined target sites in the human genome. In some embodiments, the recombinases described herein can be fused to a heterologous domain, such as a heterologous DNA-binding domain. In some embodiments, the recombinase can be fused to a heterologous DNA-binding domain, such as a DNA-binding domain from a zinc finger, TAL, meganuclease, transcription factor, or sequence-guided DNA-binding element. In some embodiments, the recombinase can be fused to a sequence-guided DNA-binding element, such as a DNA-binding domain from a CRISPR-associated (Cas) DNA-binding element, such as Cas9. In some embodiments, the DNA-binding element fused to the recombinase domain can contain a mutation that inactivates other catalytic functions, such as a mutation that inactivates endonuclease activity, e.g., generating an inactive meganuclease, or a mutation that partially or completely inactivates the Cas protein, e.g., generating a nickase Cas9 or an inactive Cas9 (dCas9). For example, see Standage-Beier et al., CRISPR J 2(4):209-222 (2019) describes the use of dCas9 fused to Tn3 resolvase (integrase Cas9, iCas9) using two monomeric fusion proteins appropriately spaced at the target site for coordinated targeting of sequence-specific integration of a reporter system into the genome of HEK293 cells. Additional examples of recombinase targeting by DNA-binding domains include zinc finger fusions (zinc finger recombinase, ZFR (Gaj et al. Nucleic Acids Res 41(6):3937-3946 (2013)); RecZF (Gersbach et al. Nucleic Acids Res 38(12):4198-4206(2010)), TALE fusions (TALE recombinase, TALER (Mercer et al. Nucleic Acids Res 40(21):11163-11172(2012))), and dCas9 fusions (recombinase Cas9, recCas9 (Chaikind et al. Nucleic Acids Res 44(20):9758-9770(2016)); integrase Cas9, iCas9 (Standage-Beier et al. CRISPR J 2(4):209-222(2019))), all of which are incorporated herein by reference.

[1033] In some embodiments, the DNA-binding domain comprises Streptococcus pyogenes Cas9 (SpCas9) or a functional fragment or variant thereof. In some embodiments, the DNA-binding domain comprises a modified SpCas9. In several embodiments, the modified SpCas9 comprises a modification that alters its protospacer-adjacent motif (PAM) specificity. In several embodiments, the PAM has specificity for the nucleic acid sequence 5'-NGT-3'. In several embodiments, the modified SpCas9 comprises one or more amino acid substitutions, e.g., at one or more of L1111, D1135, G1218, E1219, A1322, or R1335, e.g., selected from L1111R, D1135V, G1218R, E1219F, A1322R, and R1335V. In some embodiments, the modified SpCas9 comprises the amino acid substitution T1337R and one or more additional amino acid substitutions selected from L1111, D1135L, S1136R, G1218S, E1219V, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, R1335Q, T1337, T1337L, T1337Q, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337H, T1337Q, and T1337M, or a corresponding amino acid substitution thereof. In some embodiments, the modified SpCas9 comprises (i) one or more amino acid substitutions selected from D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q, and T1337; and (ii) one or more additional amino acid substitutions selected from L1111R, G1218R, E1219F, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, T1337L, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337R, T1337H, T1337Q, and T1337M, or a corresponding amino acid substitution.

[1034] In some embodiments, the DNA-binding domain comprises a Cas domain, e.g., a Cas9 domain. In several embodiments, the DNA-binding domain comprises a nuclease-active Cas domain, a Cas nickase (nCas) domain, or a nuclease-inactive Cas (dCas). In several embodiments, the DNA-binding domain comprises a nuclease-active Cas9 domain, a Cas9 nickase (nCas9) domain, or a nuclease-inactive Cas9 (dCas). In some embodiments, the DNA-binding domain comprises a Cas9 domain of Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the DNA-binding domain comprises Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the DNA-binding domain comprises S. pyogenes or S. thermophilus Cas9 or a functional fragment thereof. In some embodiments, the DNA-binding domain comprises a Cas9 sequence, e.g., as described in Chylinski, Rhun, and Charpentier (2013) RNA Biology 10:5, 726-737 (incorporated herein by reference). In some embodiments, the DNA-binding domain comprises a Cas, e.g., the HNH nuclease subdomain and / or RuvC1 subdomain of Cas9 or a variant thereof, as described herein. In some embodiments, the DNA-binding domain comprises Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the DNA-binding domain comprises a Cas polypeptide (e.g., an enzyme) or a functional fragment thereof.In some embodiments, the Cas polypeptide (e.g., enzyme) is Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (e.g., Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3 , Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, C sm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Cs a5, a type II Cas effector protein, a type V Cas effector protein, a type VI Cas effector protein, CARF, DinG, Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12b / C2c1, Cas12c / C2c3, SpCas9(K855A), eSpCas9(1.1), SpCas9-HF1, a hyper accurate Cas9 variant (HypaCas9), a homolog thereof, a modified or engineered version thereof, and / or a functional fragment thereof. In some embodiments, the Cas9 comprises one or more substitutions selected from, for example, H840A, D10A, P475A, W476A, N477A, ​​D1125A, W1126A, and D1127A.In embodiments, the Cas9 comprises one or more mutations at a position selected from D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987, e.g., one or more substitutions selected from D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A. In some embodiments, the DNA binding domain is selected from the group consisting of Corynebacterium ulcerans, Corynebacterium diphtheria, Spiroplasma syrphidicola, Prevotella intermedia, Spiroplasma taiwanense, Streptococcus iniae, Belliella baltica, Psychroflexus torquis, Staphylococcus thermophilus, Listeria innocua, Campylobacter jejuni, Neisseria meningitidis, and the like. meningitidis, Streptococcus pyogenes, or Staphylococcus aureus, or functional fragments or variants thereof.

[1035] In some embodiments, the DNA binding domain comprises a Cpf1 domain comprising one or more substitutions selected from, e.g., D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A and D917A / E1006A / D1255A, e.g., at positions D917, E1006A, D1255, or any combination thereof.

[1036] In some embodiments, the DNA binding domain comprises spCas9, spCas9-VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL.

[1037] In some embodiments, the DNA-binding domain comprises an amino acid sequence listed in Table 37 below, or an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the DNA-binding domain comprises an amino acid sequence that has no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 differences (e.g., mutations) relative to any of the amino acid sequences described herein.

[1038] [Table 736]

[1039] [Table 737]

[1040] [Table 738]

[1041] In some embodiments, the Cas polypeptide binds to a gRNA that directs binding of the DNA-binding domain. In some embodiments, the gRNA comprises, for example, from 5' to 3', (1) a gRNA spacer; (2) a gRNA scaffold. In some embodiments, (1) is a Cas9 spacer of approximately 18 to 22 nt, for example, 20 nt. (2) is a gRNA scaffold comprising one or more loops, e.g., one, two, or three loops, for binding the template to the nickase Cas9 domain. In some embodiments, the gRNA scaffold comprises, from 5' to 3', the sequence: GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCC (SEQ ID NO: 3444) Carries out.

[1042] In some embodiments, the Gene Writing System described herein is used to perform editing in HEK293, K562, U2OS, or HeLa cells. In some embodiments, the Gene Writing System is used to perform editing in primary cells, such as primary cortical neurons from E18.5 mice.

[1043] In some embodiments, a system or method described herein comprises a CRISPR DNA targeting enzyme or system, or a functional fragment or variant thereof, described in U.S. Patent Application Publication No. 20200063126, U.S. Patent Application Publication No. 20190002889, or U.S. Patent Application Publication No. 20190002875 (each of which is incorporated by reference herein in its entirety). For example, in some embodiments, a GeneWriter polypeptide or Cas endonuclease described herein comprises a polypeptide sequence described in any of the applications listed in this paragraph, and in some embodiments, a guide RNA comprises a nucleic acid sequence described in any of the applications listed in this paragraph.

[1044] In some embodiments, the DNA-binding domain (e.g., the target-binding domain or the template-binding domain) comprises a meganuclease domain or a functional fragment thereof. In some embodiments, the meganuclease domain has endonuclease activity, e.g., double-strand cleavage and / or nickase activity. In other embodiments, the meganuclease domain has reduced activity, e.g., lacks endonuclease activity, e.g., the meganuclease lacks catalytic activity. In some embodiments, a meganuclease lacking catalytic activity is used as the DNA-binding domain, e.g., as described in Fonfara et al. Nucleic Acids Res 40(2):847-860 (2012), the entire contents of which are incorporated herein by reference. In several embodiments, the DNA-binding domain comprises one or more modifications relative to the wild-type DNA-binding domain, e.g., modifications by directed evolution, e.g., phage-assisted continuous evolution (PACE).

[1045] Intein In some embodiments, as described in more detail below, intein-N can be fused to the N-terminal portion of a polypeptide described herein (e.g., a Gene Writer polypeptide), e.g., at the first domain. In several embodiments, intein-C can also be fused to the C-terminal portion of a polypeptide described herein (e.g., at the second domain), e.g., linking the N-terminal portion to the C-terminal portion, thereby linking the first domain and the second domain. In some embodiments, the first domain and the second domain are each independently selected from a DNA-binding domain and a catalytic domain, e.g., a recombinase domain. In some embodiments, a single domain is disrupted using an intein strategy described herein, e.g., a DNA-binding domain, e.g., a dCas9 domain.

[1046] In some embodiments, the systems or methods described herein include an intein, which is, for example, a self-splicing protein intron (e.g., a peptide) that links flanking N-terminal and C-terminal exteins (e.g., fragments to be joined). Inteins can, in some cases, comprise a fragment of a protein that can be automatically excised to join the remaining fragment (extein) with a peptide bond in a process known as protein splicing. Inteins are also referred to as "protein inons." The process of an intein being automatically excised to join the remaining portion of a protein is referred to herein as "protein splicing" or "intein-mediated protein splicing." In some embodiments, the inteins of a precursor protein (an intein-containing protein prior to intein-mediated protein splicing) are derived from two genes. Such inteins are referred to herein as split inteins (e.g., split intein-N and split intein-C). For example, in cyanobacteria, DnaE, ​​the catalytic subunit of DNA polymerase III, is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene may be referred to herein as "intein-N." The intein encoded by the dnaE-c gene may be referred to herein as "intein-C."

[1047] The use of inteins to link heterologous protein fragments is described, for example, in Wood et al., J. Biol. Chem. 289(21);14512-9(2014), the entire contents of which are incorporated herein by reference. For example, when fused to separate protein fragments, inteins IntN and IntC can recognize each other and splice from themselves and / or simultaneously link the N- and C-terminal exteins of the protein fragments to which they are fused, thereby reconstituting a full-length protein from the two protein fragments.

[1048] In some embodiments, synthetic inteins based on the dnaE intein, Cfa-N (e.g., split intein-N), and Cfa-C (e.g., split intein-C) intein pairs are used. Examples of such inteins are described, for example, in Stevens et al., J Am Chem Soc. 2016 Feb. 24;138(7):2162-5, the entire contents of which are incorporated herein by reference. Non-limiting examples of intein pairs that can be used in accordance with the present disclosure include the Cfa DnaE intein, the Ssp GyrB intein, the Ssp DnaX intein, the Ter DnaE3 intein, the Ter ThyX intein, the Rma DnaB intein, and the Cne Prp8 intein (e.g., as described in U.S. Pat. No. 8,394,604, the entire contents of which are incorporated herein by reference).

[1049] In some embodiments, intein-N and intein-C can be fused to the N-terminal portion of split Cas9 and the C-terminal portion of split Cas9, respectively, for linking the N-terminal portion of split Cas9 with the C-terminal portion of split Cas9. For example, in some embodiments, intein-N is fused to the C-terminus of the N-terminal portion of split Cas9, forming the structure N-[N-terminal portion of split Cas9]-[intein-N]~C. In some embodiments, intein-C is fused to the N-terminus of the C-terminal portion of split Cas9, forming the structure N-[intein-C]~[C-terminal portion of split Cas9]-C. The mechanism of intein-mediated protein splicing for linking intein-linked proteins (e.g., split Cas9) is described in Shah et al., Chem Sci. 2014;5(1):446-461 (incorporated herein by reference). Methods for designing and using inteins are described, for example, in WO2020051561, WO2014004336, WO2017132580, U.S. Patent Application Publication No. 20150344549, and U.S. Patent Application Publication No. 20180127780 (each of which is incorporated herein by reference in its entirety).

[1050] In some embodiments, split refers to a division into two or more fragments. In some embodiments, the split Cas9 protein or split Cas9 is a Cas9 protein that is provided as an N-terminal fragment and a C-terminal fragment encoded by two separate nucleotide sequences. Polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein can be spliced ​​to form a reconstituted Cas9 protein. In some embodiments, the Cas9 protein is split into two fragments within a denatured region of the protein, as described in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, or as described in Jiang et al. (2016) Science 351:867-871 and PDB file: 5F9R (each of which is incorporated herein by reference in its entirety). The denatured region can be determined by one or more protein structure determination techniques known in the art, including, but not limited to, X-ray crystallography, NMR spectroscopy, electron microscopy (e.g., cryoEM), and / or in silico protein modeling. In some embodiments, the protein is split into two fragments at any C, T, A, or S within the region of SpCas9 between amino acids A292-G364, F445-K483, or E565-T637, or at the corresponding position in any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or other napDNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as protein cleavage.

[1051] In some embodiments, protein fragments range in length from about 2 to 1000 amino acids (e.g., 2 to 10, 10 to 50, 50 to 100, 100 to 200, 200 to 300, 300 to 400, 400 to 500, 500 to 600, 600 to 700, 700 to 800, 800 to 900, or 900 to 1000 amino acids). In some embodiments, protein fragments range in length from about 5 to 500 amino acids (e.g., 5 to 10, 10 to 50, 50 to 100, 100 to 200, 200 to 300, 300 to 400, or 400 to 500 amino acids). In some embodiments, protein fragments range in length from about 20 to 200 amino acids (e.g., 20 to 30, 30 to 40, 40 to 50, 50 to 100, or 100 to 200 amino acids).

[1052] In some embodiments, for example, a portion or fragment of a Gene Writer polypeptide described herein is fused to an intein. A nuclease can be fused to the N-terminus or C-terminus of an intein. In some embodiments, a portion or fragment of a fusion protein is fused to an intein and also fused to an AAV capsid protein. The intein, nuclease, and capsid protein can be fused together in any configuration (e.g., nuclease-intein-capsid, intein-nuclease-capsid, capsid-intein-nuclease, etc.). In some embodiments, the N-terminus of the intein is fused to the C-terminus of the fusion protein, and the C-terminus of the intein is fused to the N-terminus of an AAV capsid protein.

[1053] In some embodiments, a Gene Writer polypeptide (e.g., comprising a nickase Cas9 domain) is fused to Intein-N, and a polypeptide comprising a polymerase domain is fused to Intein-C.

[1054] Exemplary nucleotide and amino acid sequences of inteins are set forth below: DnaE Intein-N DNA: [ka]

[1055] DnaE Intein-N Protein: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDRGEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNLPN (SEQ ID NO: 3446)

[1056] DnaE Intein-C DNA: ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGATATTGGAGTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAGCTTCTAAT (SEQ ID NO: 3447)

[1057] Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN (SEQ ID NO: 3448)

[1058] Cfa-N DNA: [ka]

[1059] Cfa-N protein: CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQMLPIDEIFERGLDLKQVDGLP (SEQ ID NO: 3450)

[1060] Cfa-C DNA: [ka]

[1061] Cfa-C protein: MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN (SEQ ID NO: 3452)

[1062] Genomic safe harbor sites In some embodiments, the Gene Writer targets a genomic safe harbor site (e.g., directs insertion of the heterologous sequence of interest to a position with a safe harbor score of at least 3, 4, 5, 6, 7, or 8). In some embodiments, the genomic safe harbor site is a Natural Harbor™ site. In some embodiments, the Natural Harbor™ site is derived from the native target of a mobile genetic element, such as a recombinase, transposon, retrotransposon, or retrovirus. The native target of a mobile element may serve as an ideal location for genomic integration, given their evolutionary selection. In some embodiments, the Natural Harbor™ site is ribosomal DNA (rDNA). In some embodiments, the Natural Harbor™ site is 5S rDNA, 18S rDNA, 5.8S rDNA, or 28S rDNA. In some embodiments, the Natural Harbor™ site is a Mutsu site in 5S rDNA. In some embodiments, the Natural Harbor™ site is an R2 site, an R5 site, an R6 site, an R4 site, an R1 site, an R9 site, or an RT site in 28S rDNA. In some embodiments, the Natural Harbor™ site is an R8 site or an R7 site in 18S rDNA. In some embodiments, the Natural Harbor™ site is DNA encoding a transfer RNA (tRNA). In some embodiments, the Natural Harbor™ site is DNA encoding a tRNA-Asp or tRNA-Glu. In some embodiments, the Natural Harbor™ site is DNA encoding a spliceosomal RNA. In some embodiments, the Natural Harbor™ site is DNA encoding a small nuclear RNA (snRNA), such as U2 snRNA.

[1063] Thus, in some aspects, the disclosure provides methods comprising inserting a heterologous sequence of interest into a Natural Harbor™ site using the Gene Writer system described herein. In some embodiments, the Natural Harbor™ site is a site set forth in Table 4A below. In some embodiments, the heterologous sequence of interest is inserted within 20, 50, 100, 150, 200, 250, 500, or 1000 base pairs of the Natural Harbor™ site. In some embodiments, the heterologous sequence of interest is inserted within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb of the Natural Harbor™ site. In some embodiments, the heterologous sequence of interest is inserted at a site having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to a sequence shown in Table 4A. In some embodiments, the heterologous sequence of interest is inserted within 20, 50, 100, 150, 200, 250, 500, or 1000 base pairs or within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb of a site having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to a sequence shown in Table 4A. In some embodiments, the heterologous sequence of interest is inserted within a gene set forth in column 5 of Table 4A or within 20, 50, 100, 150, 200, 250, 500, or 1000 base pairs or within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb of that gene.

[1064] [Table 739]

[1065] [Table 740]

[1066] [Table 741]

[1067] [Table 742]

[1068] [Table 743]

[1069] Additional Gene Writer™ Functional Features The Gene Writers described herein may, in some cases, be characterized by one or more functional measures or characteristics. In some embodiments, the DNA-binding domain (e.g., the target-binding domain) has one or more of the functional characteristics described below. In some embodiments, the template-binding domain has one or more of the functional characteristics described below. In some embodiments, the template (e.g., the template DNA) has one or more of the functional characteristics described below. In some embodiments, the target site modified by the Gene Writer has one or more of the functional characteristics described below after modification by the Gene Writer.

[1070] Gene Writer Polypeptide DNA-binding domain In some embodiments, the DNA-binding domain can bind to a target sequence (e.g., a dsDNA target sequence) with higher affinity than a standard DNA-binding domain. In some embodiments, the standard DNA-binding domain is a DNA-binding domain from phiC31 recombinase from Streptomyces bacteriophage phiC31. In some embodiments, the DNA-binding domain can bind to a target sequence (e.g., a dsDNA target sequence) with an affinity of 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM).

[1071] In some embodiments, the affinity of a DNA-binding domain for its target sequence (e.g., a dsDNA target sequence) is measured in vitro, e.g., by thermophoresis, as described, e.g., in Asmari et al. Methods 146:107-119 (2018), which is incorporated herein by reference in its entirety.

[1072] In some embodiments, the DNA-binding domain can bind to its target sequence (e.g., a dsDNA target sequence) with an affinity of, for example, 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM) in the presence of a molar excess, e.g., about a 100-fold molar excess, of scrambled-sequence competitor dsDNA.

[1073] In some embodiments, the DNA-binding domain is found to bind to its target sequence (e.g., a dsDNA target sequence) at a higher frequency than any other sequence in the genome of the target cell, e.g., a human target cell, as measured by, e.g., ChIP-seq (e.g., in HEK293T cells), e.g., as described in He and Pu (2010) Curr. Protoc Mol Biol Chapter 21, which is incorporated herein by reference in its entirety. In some embodiments, the DNA-binding domain is found to bind to its target sequence (e.g., a dsDNA target sequence) at a frequency at least 5-fold or 10-fold higher than any other sequence in the genome of the target cell, as measured by, e.g., ChIP-seq (e.g., in HEK293T cells), e.g., as described in He and Pu (2010), supra.

[1074] Template-binding domain In some embodiments, the template-binding domain can bind to the template DNA with higher affinity than a standard DNA-binding domain. In some embodiments, the standard DNA-binding domain is a DNA-binding domain from phiC31 recombinase from Streptomyces bacteriophage phiC31. In some embodiments, the template-binding domain can bind to the template DNA with an affinity of 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM). In some embodiments, the affinity of the DNA-binding domain for its template DNA can be determined using a method such as, for example, Asmari et al. al. Methods 146:107-119 (2018), which is incorporated herein by reference in its entirety. In some embodiments, the affinity of a DNA-binding domain for its template DNA is measured in a cell (e.g., by FRET or ChIP-Seq).

[1075] In some embodiments, the DNA binding domain is selected from the group consisting of the nucleotides ... interest, e.g., Yant et al. Mol. The DNA-binding domain associates with template DNA in vitro in the presence of 10 nM competitor DNA, with at least 50% template DNA bound, as described in Cell Biol 24(20):9239-9247 (2004), the entire contents of which are incorporated herein by reference. In some embodiments, the DNA-binding domain associates with template DNA in cells (e.g., in HEK293T cells) at a frequency at least about 5-fold or 10-fold higher than with scrambled DNA. In some embodiments, the frequency of association of the DNA-binding domain with template DNA or scrambled DNA is measured by ChIP-seq, for example, as described in He and Pu (2010), supra.

[1076] target site In some embodiments, after gene writing, the target site surrounding the integration sequence contains a limited number of insertions or deletions, e.g., in about 50% or less than 10% of integration events, as determined, e.g., by long-read amplicon sequencing of the target site, as described, e.g., in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903 (incorporated herein by reference in its entirety). For example, indels have been observed after integration of insert DNA into a pseudosite in the human genome by phiC31 integrase, as described, e.g., in Thyagarajan et al. Mol Cell Biol 21(12):3926-3934 (2001) (the teachings of which are incorporated herein by reference in their entirety). In some embodiments, the Gene Writing System of the present invention can result in genomic modifications (e.g., insertions or deletions) of target sites (e.g., adjacent to, e.g., the site of integration of the insert DNA) that contain less than 20 nt of DNA, e.g., less than 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or less than 1 nt of DNA. In some embodiments, the Gene Writing System of the present invention can result in insertions of target sites (e.g., adjacent to, e.g., the site of integration of the insert DNA) that contain less than 20 nucleotides or base pairs of DNA, e.g., less than 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or less than 1 nucleotide or base pair of DNA. In some embodiments, the Gene Writing System of the present invention may result in the deletion of a target site (e.g., adjacent to, e.g., the site of integration of the insert DNA) that contains less than 20 nucleotides or base pairs of genomic DNA, e.g., less than 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or less than 1 nucleotide or base pair.In some embodiments, the rate of insertion or deletion events is lower when the core region, e.g., central dinucleotide, of the recognition sequence of the target site, e.g., attB, attP, or pseudosite thereof, comprises 100% identity to the core region, e.g., central dinucleotide, of the recognition sequence on the insert DNA, e.g., attP or attB site. In some embodiments, the rate of unintended insertion or deletion events is, for example, at least 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 3.0, 4.0, 5.0, 10, 20, 30, 40, 50, 60, 70, 80, 90, or at least 100 times lower at the target genomic site when the central dinucleotide of the recognition sequence of the target site is identical to the central dinucleotide of the recognition sequence in the insert DNA.

[1077] In some embodiments, the target site does not exhibit multiple insertion events, e.g., head-to-tail or head-to-head duplications, as determined, for example, by long-read amplicon sequencing or molecular combing (Example 29) of the target site, e.g., as described in Karst et al. (2020), supra. In some embodiments, the target site exhibits fewer than 100 insertion copies at the target site, e.g., 75 insertion copies, 50 insertion copies, 45 insertion copies, 40 insertion copies, 35 insertion copies, 30 insertion copies, 25 insertion copies, 20 insertion copies, 15 insertion copies, 14 insertion copies, 13 insertion copies, 12 insertion copies, 11 insertion copies, 10 insertion copies, 9 insertion copies, 8 insertion copies, 7 insertion copies, 6 insertion copies, 5 insertion copies, 4 insertion copies, 3 insertion copies, 2 insertion copies, or a single insertion copy. In some embodiments, target sites exhibiting two or more copies of an insertion sequence are present in less than 95% of the target sites containing an insert, e.g., less than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 9%, 8%, 7%, 6%, 4%, 4%, 3%, 2%, or less than 1% of the target sites containing an insert, as determined, e.g., by long-read amplicon sequencing or molecular combing (Example 29) of the target sites, e.g., as described in Karst et al. (2020), supra. In some embodiments, target sites exhibiting three or more copies of the insertion sequence are present in less than 95% of the target sites containing the insert, e.g., less than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 9%, 8%, 7%, 6%, 4%, 4%, 3%, 2%, or less than 1% of the target sites containing the insert, as determined, e.g., by long-read amplicon sequencing or molecular combing (Example 29) of the target sites, e.g., as described in Karst et al. (2020), supra.In some embodiments, target sites exhibiting 4 or more copies of the insertion sequence are present in less than 95% of the target sites containing the insert, e.g., less than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 9%, 8%, 7%, 6%, 4%, 4%, 3%, 2%, or less than 1% of the target sites containing the insert, as determined, e.g., by long-read amplicon sequencing or molecular combing (Example 29) of the target sites, e.g., as described in Karst et al. (2020), supra. In some embodiments, the target sites exhibit at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more copies per target site. In some embodiments, target sites exhibiting multiple copies of an insertion sequence are present in 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99% or more of the target sites containing an insert, as determined, for example, by long-read amplicon sequencing or molecular combing (Example 29) of the target sites, e.g., as described in Karst et al. (2020), supra. In some embodiments, the copies are concatemerized, i.e., concatemerized. In some embodiments, the target site contains an integration sequence corresponding to the template DNA (e.g., an entire plasmid, minicircle, or viral vector genome). In some embodiments, the target site contains a fully integrated template molecule. In some embodiments, the target site contains components of vector DNA, e.g., AAV ITRs. In some embodiments, the target site contains 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more ITRs after integration. In some embodiments, at least one ITR is present in at least 1% of the target sites after integration, e.g., at least 1%, 5%, 10%, 15%, 20%, 25%, 50%, 60%, 70%, 80%, 90, 95%, 96%, 97%, 98% or at least 99% of the target sites after integration.In some embodiments, at least one ITR is present in less than 50% of target sites after integration, e.g., less than 50%, 40%, 30%, 20%, 10%, 9%, 8%, 7%, 6%, 4%, 4%, 3%, 2%, or less than 1% of target sites after integration, as determined, e.g., by long-read amplicon sequencing or molecular combing (Example 29) of the target sites, e.g., as described in Karst et al. (2020), supra. In some embodiments, the multiple copies are arranged in a head-to-head, tail-to-tail, or head-to-tail configuration, or a mixture thereof. In some embodiments, when template DNA is initially excised from a viral vector or plasmid, e.g., by a first recombination event prior to integration, the target site does not contain an insert comprising DNA exogenous to the recognition site flanking cassette, e.g., vector DNA, e.g., AAV ITRs, in more than about 50% of events, e.g., more than about 50%, 40%, 30%, 20%, 10%, 9%, 8%, 7%, 6%, 4%, 4%, 3%, 2%, or more than about 1% of events, as determined, e.g., by long-read amplicon sequencing or molecular combing (Example 29) of the target site, e.g., as described in Karst et al. (2020), supra. In some embodiments, the integrated DNA does not contain any bacterial antibiotic resistance genes.

[1078] In some embodiments, DNA integrated at a target site by a Gene Writing system described herein comprises terminal hybrid recognition sequences (e.g., as described herein, e.g., first and / or second parapalindromic sequences), e.g., attL and attR sequences formed by recombination between the recognition sites in the insert DNA, e.g., attP or attB in the insert DNA, and a recognition site in the target DNA, e.g., an attP or attB site, or a pseudosite thereof. In some embodiments, the integrated DNA comprises one or more ITRs, e.g., 1, 2, 3, 4, or more ITRs, between the terminal hybrid recognition sequences, e.g., attL or attR sequences. In some embodiments, at least 1% of target sites with integrated DNA, e.g., at least 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or at least 90% of the integrated DNA, comprise terminal hybrid recognition sequences, e.g., ITRs between the attL and attR sequences. In some embodiments, the integrated DNA comprising an ITR between terminal hybrid recognition sequences, e.g., attL or attR sequences, comprises a single copy of the inserted DNA, e.g., a monomeric insert. In some embodiments, the monomeric insert comprises a terminal hybrid recognition sequence, e.g., attL and attR sequences, and lacks any internal ITRs. In some embodiments, the monomeric insert comprises a terminal hybrid recognition sequence, e.g., attL and attR sequences, and a single internal ITR. In some embodiments, the monomeric insert comprises a terminal hybrid recognition sequence, e.g., attL and attR sequences, and multiple internal ITRs, e.g., two internal ITRs. In some embodiments, the integrated DNA comprising a terminal hybrid recognition sequence, e.g., an ITR between attL and attR sequences, comprises multiple copies of the inserted DNA, e.g., a concatemeric insert. In some embodiments, the concatemeric insert comprises a terminal hybrid recognition sequence, e.g., attL and attR sequences, and at least two copies, e.g., at least 2, 3, or 4 copies, of the inserted DNA.In some embodiments, inserts containing terminal hybrid recognition sequences, e.g., attL and attR sequences, that contain low copies of the insert DNA are more frequent than those with higher copies of the insert DNA (e.g., inserts with one copy are more frequent than inserts with two copies, inserts with two copies are more frequent than inserts with three copies, or inserts with one copy are more frequent than inserts with three copies), and exhibit a higher frequency of occurrence, e.g., 1.1, 1.2, 1.3, 1.4, 1.5, 2.0, 2.5, 3.0, 4.0, 5.0, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more times more frequent. In some embodiments, monomeric insertions occur more frequently than dimeric insertions, e.g., at least 1.1, 1.2, 1.3, 1.4, 1.5, 2.0, 2.5, 3.0, 4.0, 5.0, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more times more frequent than dimeric insertions. In some embodiments, dimeric insertions occur more frequently than trimeric insertions, e.g., at least 1.1, 1.2, 1.3, 1.4, 1.5, 2.0, 2.5, 3.0, 4.0, 5.0, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more times more frequent than trimeric insertions. In some embodiments, monomer and dimer inserts are more frequent than concatemeric inserts (three or more inserts), e.g., at least 1.1, 1.2, 1.3, 1.4, 1.5, 2.0, 2.5, 3.0, 4.0, 5.0, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more times more frequent than concatemeric inserts. In some embodiments, concatemeric inserts comprise terminal hybrid recognition sequences, e.g., attL and attR sequences, and one or more internal recombinase recognition sequences, e.g., one, two, three, four, or more internal recognition sequences, e.g., attB or attP sequences. In some embodiments, concatemeric inserts comprise terminal hybrid recognition sequences, e.g., attL and attR sequences, and one or more internal ITRs, e.g., one, two, three, four, five, six, or more internal ITRs.The copy number, recognition sequences and ITRs of the insert DNA described herein and the relative positioning of their components can be determined using molecular combing, as described in Example 29 and in Kaykov et al. Sci Rep 6:19636 (2016), the entire contents of which are incorporated herein by reference.

[1079] In some embodiments, an insertion event may occur in which the integrated DNA does not contain terminal hybrid recognition sequences, e.g., attL and attR sequences. In some embodiments, the integrated DNA may contain one terminal recognition sequence, e.g., an attL or attR sequence. In some embodiments, the integrated DNA may not have any terminal hybrid recognition sequences, e.g., attL or attR, e.g., neither end of the integrated DNA contains a hybrid recognition sequence, e.g., an attL or attR sequence. In some embodiments, the integrated DNA that does not contain terminal hybrid recognition sequences, e.g., attL or attR sequences, comprises a fragment of the inserted DNA (e.g., an incomplete inserted DNA, e.g., an inserted DNA with an incomplete promoter, gene, or heterologous sequence of interest). In some embodiments, the integrated DNA that does not contain terminal hybrid recognition sequences, e.g., attL or attR sequences, comprises an incomplete multiple inserted DNA sequence, e.g., containing less than 1, more than 1 and less than 2, more than 2 and less than 3, more than 3 and less than 4, or another incomplete multiple copy of the complete inserted DNA.

[1080] In some embodiments, after use of the Gene Writing System, newly integrated DNA comprising terminal hybrid recognition sequences, e.g., attL and attR sequences, is present at a high frequency in the cell or population of cells as measured by an assay described herein, e.g., long-read sequencing or molecular combing, and comprises, for example, more than 50%, more than 60%, more than 70%, more than 80%, more than 90%, more than 95%, more than 96%, more than 97%, more than 98%, more than 99%, more than 99.5%, or more than 99.9% of the total insertion events compared to newly integrated DNA comprising one or several terminal hybrid recognition sequences, e.g., attL or attR sequences. In some embodiments, after use of the Gene Writing System, the newly integrated DNA containing terminal hybrid recognition sequences, e.g., attL and attR sequences, contains a lower average inserted DNA copy number per insertion event, e.g., on average, at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.5, or 2.0 fewer copies per insertion event, compared to the average inserted DNA copy number of integration events containing one or several terminal hybrid recognition sequences, e.g., attL or attP sequences. In some embodiments, after use of the Gene Writing System, the newly integrated DNA containing terminal hybrid recognition sequences, e.g., attL and attR sequences, contains a higher percentage of fully inserted DNA sequences compared to the percentage of inserted DNA sequences containing one or more terminal hybrid recognition sequences, e.g., attL or attP sequences, e.g., at least 0.1×, 0.2×, 0.3×, 0.4×, 0.5×, 0.6×, 0.7×, 0.8×, 0.9×, 1.0×, 1.5×, 2.0×, 3×, 4×, 5×, 6×, 7×, 8×, 9×, 10×, or more percent fully inserted DNA sequences.

[1081] In some embodiments, the Gene Writers described herein are capable of site-specific editing of target DNA, e.g., inserting template DNA into target DNA. In some embodiments, the site-specific Gene Writer can create an edit, e.g., an insertion, that occurs at the target site at a frequency greater than any other site in the genome. In some embodiments, the site-specific Gene Writer can create an edit, e.g., an insertion, at the target site at a frequency at least 2, 3, 4, 5, 10, 50, 100, or 1000 times greater than the frequency of all other sites in the human genome. In some embodiments, the location of the integration site is determined by one-way sequencing, e.g., as in Example 18. Incorporation of unique molecular identifiers (UMIs) into adapters or primers used for library preparation allows for quantification of individual insertion events, which can be compared between on-target insertions and all other insertions to determine preference for a defined target site. In some embodiments, an inverse PCR approach is used to determine the integration site targeted by a particular Gene Writer, e.g., as in Example 30.

[1082] In some embodiments, a gene writing system is used to edit a target DNA sequence present at a single location in the human genome. In some embodiments, gene writing is used to edit a target DNA sequence present at a single location in the human genome on a single homologous chromosome, for example, this is haplotype-specific. In some embodiments, a gene writing system is used to edit a target DNA sequence present at a single location in the human genome on two homologous chromosomes. In some embodiments, a gene writing system is used to edit a target DNA sequence present at multiple locations in the genome, for example, at least 2, 3, 4, 5, 10, 20, 50, 100, 200, 500, 1000, 5000, 10000, 100000, 200000, 500000, 1000000 (e.g., Alu element) positions in the genome. In some embodiments, the gene writing system used herein performs integration at a single target sequence in the human genome, which may be present at one or more positions. In some embodiments, the Gene Writing System used herein performs integration at a plurality of sequences that occur at least once in the human genome, for example recognizing more than 1, for example 1, 2, 3, 4, 5, 10, 20, 50, or more than 100 sequences, or fewer than 100, for example 100, 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10, or 5 sequences that occur at least once in the human genome. Thus, in some embodiments, the Gene Writer described herein may result in integration of the inserted DNA at at least 1 copy per cell, for example at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or at least 10 copies per cell, or fewer than 10 copies per cell, for example fewer than 10, 9, 8, 7, 6, 5, 4, 3, or fewer than 2 copies.

[1083] In some embodiments, the Gene Writer system can edit a genome without introducing undesired mutations. In some embodiments, the Gene Writer system can edit a genome by inserting a template, e.g., template DNA, into the genome. In some embodiments, the resulting modifications in the genome include minimal mutations relative to the template DNA sequence. In some embodiments, the average error rate of genome insertion compared to the template DNA is 10 per nucleotide. -4 , 10 -5 or 10 -6 In some embodiments, the number of mutations in the template DNA introduced into the target cell is, on average, less than 1, 2, 3, 4, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides per genome. In some embodiments, the insertion error rate in the target genome is determined by comparing the template DNA sequence with long-read amplicon sequencing across known target sites, as described, for example, in Karst et al. (2020), supra. In some embodiments, the errors counted by this method include nucleotide substitutions relative to the template sequence. In some embodiments, the errors counted by this method include nucleotide deletions relative to the template sequence. In some embodiments, the errors counted by this method include nucleotide insertions relative to the template sequence. In some embodiments, the errors counted by this method include one or more combinations of nucleotide substitutions, deletions, or insertions relative to the template sequence.

[1084] The efficiency of the integration event can be used as a measure of editing of the target site or target cell by the Gene Writer system. In some embodiments, the Gene Writer systems described herein can integrate a heterologous sequence of interest into a target site or a fraction of target cells. In some embodiments, the Gene Writer system can edit at least 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% of the target locus, as measured by detection of edits after amplification of the entire target and subsequent analysis by, for example, long-read amplicon sequencing, as described in, for example, Karst et al. (2020). In some embodiments, the Gene Writer system is capable of editing cells at an average copy number of at least 0.1, e.g., at least 0.1, 0.5, 1, 2, 3, 4, 5, 10, or 100 copies per genome, normalized to a reference gene, e.g., RPP30, across a population of cells, as determined by ddPCR using a transgene-specific primer-probe set, e.g., according to the methods described in Lin et al. Hum Gene Ther Methods 27(5):197-208 (2016).

[1085] In some embodiments, copy number per cell is analyzed by single-cell ddPCR (sc-ddPCR), e.g., according to the method of Igarashi et al. Mol Ther Methods Clin Dev 6:8-16 (2017), which is incorporated herein by reference in its entirety. In some embodiments, at least 1%, e.g., at least 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% of the target cells are positive for integration as assessed by sc-ddPCR using a transgene-specific primer-probe set. In some embodiments, the average copy number is at least 0.1, e.g., at least 0.1, 0.5, 1, 2, 3, 4, 5, 10, or 100 copies per cell, as measured by sc-ddPCR using a transgene-specific primer-probe set.

[1086] In some embodiments, the target site comprises a pair of nucleic acid sequences, wherein one of the nucleic acid sequences is palindromic with respect to the other nucleic acid sequence or has at least 20% (e.g., at least 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%), e.g., at least 50%, sequence identity with a palindromic with respect to the other nucleic acid sequence, or has no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 sequence mismatches with the other nucleic acid sequence.

[1087] Insert DNA In some embodiments, the insert DNA described herein comprises a nucleic acid sequence that can be integrated into a target DNA molecule, e.g., by a recombinase polypeptide (e.g., a serine recombinase polypeptide), e.g., as described herein. The insert DNA is typically capable of binding to one or more recombinase polypeptides (e.g., multiple copies of a recombinase polypeptide) of the system. In some embodiments, the insert DNA comprises a region that can bind to a recombinase polypeptide (e.g., a recognition sequence described herein).

[1088] In some embodiments, the insert DNA comprises a sequence of interest for insertion into the target DNA. The sequence of interest can be coding or non-coding. In some embodiments, the sequence of interest can comprise an open reading frame. In some embodiments, the insert DNA comprises a Kozak sequence. In some embodiments, the insert DNA comprises an internal ribosome entry site. In some embodiments, the insert DNA comprises a self-cleaving peptide, such as a T2A or P2A site. In some embodiments, the insert DNA comprises a start codon. In some embodiments, the insert DNA comprises a splice acceptor site. In some embodiments, the insert DNA comprises a splice donor site. In some embodiments, the insert DNA comprises a microRNA binding site, e.g., downstream of the stop codon. In some embodiments, the insert DNA comprises a polyA tail, e.g., downstream of the stop codon of the open reading frame. In some embodiments, the insert DNA comprises one or more exons. In some embodiments, the insert DNA comprises one or more introns. In some embodiments, the insert DNA comprises a eukaryotic transcription terminator. In some embodiments, the insert DNA comprises an enhanced translation element or a translation-enhancing element. In some embodiments, the insert DNA comprises a microRNA sequence, an siRNA sequence, a guide RNA sequence, or a piwiRNA sequence. In some embodiments, the insert DNA comprises a gene expression unit comprised of at least one regulatory region operably linked to an effector sequence. The effector sequence can be a sequence that is transcribed into RNA (e.g., a coding sequence, such as a sequence encoding a microRNA, or a non-coding sequence). In some embodiments, the sequence of interest can comprise a non-coding sequence. For example, the insert DNA can comprise a promoter or enhancer sequence. In some embodiments, the insert DNA comprises a tissue-specific promoter or enhancer, each of which can be unidirectional or bidirectional. In some embodiments, the promoter is an RNA polymerase I promoter, an RNA polymerase II promoter, or an RNA polymerase III promoter. In some embodiments, the promoter comprises a TATA element. In some embodiments, the promoter comprises a B recognition element.In some embodiments, the promoter has one or more binding sites for a transcription factor.

[1089] In some embodiments, the sequence of interest of the insert DNA is inserted into the target genome within an endogenous intron. In some embodiments, the sequence of interest of the insert DNA is inserted into the target genome, thereby acting as a new exon. In some embodiments, the insertion of the sequence of interest into the target genome results in the replacement of a native exon or the skipping of a native exon. In some embodiments, the sequence of interest of the insert DNA is inserted into a target site in a genomic safe harbor site such as AAVS1, CCR5, or ROSA26. In some embodiments, the sequence of interest of the insert DNA is added to the genome within an intergenic or intragenic region. In some embodiments, the sequence of interest of the insert DNA is added 5' or 3' to the genome within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb. In some embodiments, the target sequence of the inserted DNA is added to the 5' or 3' region of the genome within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb from the endogenous promoter or enhancer. In some embodiments, the target sequence of the inserted DNA may be, for example, 50 to 50,000 base pairs (e.g., 50 to 40,000 bp, 500 to 30,000 bp, 500 to 20,000 bp, 100 to 15,000 bp, 500 to 10,000 bp, 50 to 10,000 bp, or 50 to 50,000 bp. In some embodiments, the target sequence of the inserted DNA may be 1 to 50 base pairs).

[1090] In certain embodiments, insert DNA can be identified, designed, engineered, and constructed to contain sequences that modify or define genome function in a target cell or organism, for example, by introducing heterologous coding regions into the genome; influencing or effecting exon structure / alternative splicing; causing disruption of endogenous genes; causing transcriptional activation of endogenous genes; effecting epigenetic regulation of endogenous DNA; or effecting up- or down-regulation of operably linked genes. In certain embodiments, insert DNA can contain sequences encoding exons and / or transgenes and can be engineered to provide binding sites for transcription factors such as activators, repressors, enhancers, and combinations thereof. In other embodiments, the coding sequences can be further customized with splice acceptor sites, poly-A tails, etc.

[1091] The insert DNA may have some degree of homology to the target DNA. In some embodiments, the insert DNA has at least 3, 4, 5, 6, 7, 8, 9, 10, or more bases of exact homology to the target DNA or a portion thereof. In some embodiments, the insert DNA has at least 10, 15, 20, 25, 30, 40, 50, 60, 80, 100, 120, 140, 160, 180, 200, or more bases of at least 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% homology to the target DNA or a portion thereof.

[1092] As an alternative to other methods of delivery described herein, in some embodiments, the nucleic acid delivered to the cell (e.g., the nucleic acid encoding the recombinase or the template nucleic acid, or both) is designed as a minicircle, in which case the plasmid backbone sequences not relevant to GeneWriting™ are removed prior to administration to the cell. Minicircles have been shown to achieve higher transfection efficiency and gene expression compared to plasmids with backbones containing bacterial portions (e.g., bacterial origins of replication, antibiotic selection cassettes), and have been used to improve transposition efficiency (Sharma et al., 2004). al. Mol Ther Nucleic Acids 2:E74(2013)). In some embodiments, a DNA vector encoding a Gene Writer™ polypeptide is delivered as a minicircle. In some embodiments, a DNA vector containing a Gene Writer™ template is delivered as a minicircle. In some embodiments of such alternative means for delivering nucleic acids, the bacterial portion is flanked by recombination sites, e.g., attP / attB, loxP, FRT sites. In some embodiments, addition of a cognate recombinase results in intramolecular recombination and excision of the bacterial portion. In some embodiments, the recombinase sites are recognized by phiC31 recombinase. In some embodiments, the recombinase sites are recognized by Cre recombinase. In some embodiments, the recombinase sites are recognized by FLP recombinase. In some embodiments, the minicircles are delivered in a bacterial production strain, e.g., E. coli, stably expressing an inducible minicircle assembling enzyme, e.g., Kay The minicircle DNA vector is produced in a production strain according to [Para 140]. et al. Nat Biotechnol 28(12):1287-1289 (2010). Methods for minicircle DNA vector preparation and production are described in U.S. Pat. No. 9,233,174, the entire contents of which are incorporated herein by reference.

[1093] Besides plasmid DNA, minicircles can be generated by excising a desired construct, such as a recombinase expression cassette or a therapeutic expression cassette, from a viral backbone, such as an AAV vector. Previously, it has been shown that excision and circularization of the donor sequence from the viral backbone can be important for transposase-mediated integration efficiency (Yant et al. Nat Biotechnol 20(10):999-1005 (2002)). In some embodiments, minicircles are first formulated and then delivered to target cells. In other embodiments, minicircles are formed intracellularly from DNA vectors (e.g., plasmid DNA, rAAV, scAAV, ceDNA, doggiebone DNA) by co-delivery of a recombinase to achieve excision and circularization of nucleic acids (e.g., nucleic acids encoding Gene Writer™ polypeptides or DNA templates, or both) flanking the recombinase recognition sites. In some embodiments, the same recombinase is used for the first excision event (e.g., intramolecular recombination) and the second integration (e.g., target site integration) event. In some embodiments, recombination sites on the excised circular DNA (e.g., after the first recombination event, e.g., intramolecular recombination) are used as template recognition sites for the second recombination (e.g., target site recombination) event.

[1094] In some embodiments, the minicircle DNA described herein is generated by a recombinase excision event, and the Gene Writer functions to insert the minicircle DNA by a recombinase integration event. In some embodiments, the excision event and the integration event are catalyzed by the same enzyme, for example, the same serine recombinase. In some embodiments, the cassette for excision from the vector is flanked by attL and attR sites, and the excision event results in the creation of an attB or attP site that is used for integration of the cognate genomic attP or attB site. In some embodiments, the excision event involving the attL and attR sites is catalyzed by the addition of a recombination directionality factor (RDF), which enables the Gene Writer recombinase polypeptide to perform the excision. In some embodiments, the Gene Writer recombinase polypeptide functions to catalyze the integration event in the absence of an RDF.

[1095] Linker In some embodiments, domains of the compositions and systems described herein (e.g., recombinase domains and / or DNA recognition domains of the recombinase polypeptides described herein) can be linked by a linker. Compositions described herein that include a linker element have the general form S1-L-S2, where S1 and S2 can be the same or different and represent two moieties (e.g., polypeptide or nucleic acid domains, respectively) linked to each other by a linker. In some embodiments, a linker can link two polypeptides. In some embodiments, a linker can link two nucleic acid molecules. In some embodiments, a linker can link a polypeptide and a nucleic acid molecule. A linker can be a chemical bond, e.g., one or more covalent or non-covalent bonds. A linker can be flexible, rigid, and / or cleavable. In some embodiments, a linker is a peptide linker. Generally, peptide linkers are at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more amino acids in length, e.g., 2 to 50 amino acids in length, 2 to 30 amino acids in length.

[1096] The most commonly used flexible linkers have sequences consisting primarily of stretches of Gly and Ser residues ("GS" linkers). Flexible linkers are considered useful for linking domains that require some degree of movement or interaction and may contain small nonpolar (e.g., Gly) or polar (e.g., Ser or Thr) amino acids. The incorporation of Ser or Thr can also maintain the stability of the linker in aqueous solution by forming hydrogen bonds with water molecules, thereby reducing unfavorable interactions between the linker and other moieties. Examples of such linkers include those with the structure [GGS] ≧1 or [GGGS] ≧1 (SEQ ID NO: 3441). Rigid linkers are useful for maintaining a constant distance between domains while preserving their independent function. Rigid linkers can also be useful when spatial separation of domains is essential to preserve the stability or biological activity of one or more components in the drug. Rigid linkers can have an α-helical structure or a Pro-rich sequence, (XP)n, where X represents any amino acid, preferably Ala, Lys, or Glu. Cleavable linkers can release free functional domains in vivo. In some embodiments, the linker can be cleaved under certain conditions, for example, in the presence of a reducing agent or protease. In vivo cleavable linkers may utilize the reversibility of disulfide bonds. One example is a thrombin-sensitive sequence (e.g., PRS) between two Cys residues. In vitro thrombin treatment of CPRSC results in cleavage of the thrombin-sensitive sequence, while leaving the reversible disulfide bond intact. Such linkers are well known and are described, for example, in Chen et al. 2013. Fusion Protein Linkers: Property, Design and Functionality. Adv Drug Deliv Rev. 65(10):1357-1369. In vivo cleavage of the linker in the compositions described herein can also be performed by proteases expressed in vivo in specific cells or tissues, under pathological conditions (e.g., cancer or inflammation), or confined within specific cellular compartments. The specificity of various proteases allows for slow cleavage of the linker within the confined compartment.

[1097] In some embodiments, the amino acid linker is an endogenous amino acid (or homologous thereto) present between such domains in a naturally occurring polypeptide. In some embodiments, the endogenous amino acids present between such domains are substituted but the length is unchanged from the natural length. In some embodiments, additional amino acid residues are added to the amino acid residues naturally present between the domains.

[1098] In some embodiments, amino acid linkers are computationally designed or screened to maximize protein function (Anad et al., FEBS Letters, 587:19, 2013).

[1099] Additional Gene Writer features In some embodiments, the Gene Writer system can achieve complete writing without the need for endogenous host factors. In some embodiments, the system can achieve complete writing without the need for DNA repair. In some embodiments, the system can achieve complete writing without inducing a DNA damage response.

[1100] In some embodiments, the system does not require DNA repair via the NHEJ pathway, the homologous recombination repair pathway, the base excision repair pathway, or any combination thereof. Participation by a DNA repair pathway can be assayed, for example, by applying a DNA repair pathway inhibitor or a DNA repair pathway-deficient cell line. For example, when applying a DNA repair pathway inhibitor, a PrestoBlue cell viability assay can first be performed to determine the toxicity of the inhibitor and whether any normalization should be applied. SCR7 is an inhibitor of NHEJ, which can be applied in serial dilutions during Gene Writer™ delivery. PARP protein is a nuclear enzyme that binds to both single-strand and double-strand breaks as a homodimer. Therefore, the inhibitor can be used to test related DNA repair pathways, including the homologous recombination repair pathway and the base excision repair pathway. The experimental procedure is the same as that for SCR7. Cell lines lacking core proteins of the nucleotide excision repair (NER) pathway can be used to assay for Gene Writer™. The effect of NER on Writing™ can be tested. After delivery of the Gene Writer™ system into cells, ddPCR can be used to evaluate the insertion of heterologous target sequences with respect to the inhibition of DNA repair pathways. Sequencing analysis can also be performed to evaluate whether DNA repair pathways play a specific role. In some embodiments, Gene Writing™ into genomes is not reduced by knocking down DNA repair pathways as described herein. In some embodiments, Gene Writing™ into genomes is not reduced by more than 50% by knocking down DNA repair pathways.

[1101] Circular RNA in the Gene Writing System It is contemplated that it may be beneficial to use circular and / or linear RNA states during formulation, delivery, or gene writing reactions within target cells. Accordingly, in some embodiments of any of the aspects described herein, the gene writing system comprises one or more circular RNAs (circRNAs). In some embodiments of any of the aspects described herein, the gene writing system comprises one or more linear RNAs. In some embodiments, the nucleic acids described herein (e.g., nucleic acid molecules encoding gene writer polypeptides or both) are circRNAs. In some embodiments, the circular RNA molecule encodes a gene writer polypeptide. In some embodiments, a circRNA molecule encoding a gene writer polypeptide is delivered to a host cell. In some embodiments, the circular RNA molecule encodes, for example, a recombinase described herein. In some embodiments, a circRNA molecule encoding a recombinase is delivered to a host cell. In some embodiments, the circRNA molecule encoding a gene writer polypeptide is linearized (e.g., in a host cell) prior to translation.

[1102] Circular RNAs (circRNAs) are found to naturally occur in cells and have diverse functions in human cells, including both non-coding and protein-coding roles. It has been shown that circRNAs can be engineered by incorporating self-splicing introns into RNA molecules (or DNA encoding RNA molecules) that result in RNA circularization, and that engineered circRNAs can have enhanced protein production and stability (Wesselhoeft et al. Nature Communications 2018). In some embodiments, a Gene Writer™ polypeptide is encoded as a circRNA. In certain embodiments, the template nucleic acid is DNA, such as dsDNA or ssDNA.

[1103] In some embodiments, the circRNA comprises one or more ribozyme sequences. In some embodiments, the ribozyme sequence is activated, for example, for self-cleavage in the host cell, for example, thereby linearizing the circRNA. In some embodiments, the ribozyme is activated, for example, when the magnesium concentration in the host cell reaches a level sufficient for cleavage. In some embodiments, the circRNA is maintained in a low-magnesium environment before delivery to the host cell. In some embodiments, the ribozyme is a protein-responsive ribozyme. In some embodiments, the ribozyme is a nucleic acid-responsive ribozyme.

[1104] In some embodiments, the circRNA is linearized in the nucleus of a target cell. In some embodiments, linearization of the circRNA in the nucleus of a cell involves, for example, components present in the nucleus of the cell to activate the cleavage event. For example, B2 and ALU retrotransposons contain self-cleaving ribozymes whose activity is enhanced by interaction with the Polycomb protein, EZH2 (Hernandez et al. PNAS 117(1):415-425(2020)). Thus, in some embodiments, a ribozyme, such as a ribozyme derived from a B2 or ALU element that is responsive to a nuclear element, such as a nuclear protein, for example, a genome-interacting protein, such as an epigenetic modifier, EZH2, is incorporated into the circRNA, for example, in a gene writing system. In some embodiments, nuclear localization of the circRNA results in increased autocatalytic activity of the ribozyme and linearization of the circRNA.

[1105] In some embodiments, inducible ribozymes (e.g., in circRNAs described herein) are synthetically generated, for example, by utilizing protein ligand-responsive aptamer design. A system for utilizing the satellite RNA of the tobacco ringspot virus hammerhead ribozyme using the MS2 coat protein aptamer has been described (Kennedy et al. Nucleic Acids Res 42(19):12306-12321 (2014) (incorporated herein by reference in its entirety)), which results in activation of ribozyme activity in the presence of MS2 coat protein. In embodiments, such systems respond to protein ligands localized in the cytoplasm or nucleus. In some embodiments, the protein ligand is not MS2. For example, methods have been described for generating RNA aptamers for targeting ligands based on systematic evolution of ligands by exponential enrichment (SELEX) (Tuerk and Gold, Science 249(4968):505-510(1990); Ellington and Szostak, Nature 346(6287):818-822(1990); each of which is incorporated herein by reference), and in some cases assisted by in silico design (Bell et al. PNAS 117(15):8486-8493; each of which is incorporated herein by reference). Thus, in some embodiments, aptamers for target ligands are generated and incorporated into synthetic ribozyme systems, e.g., to trigger ribozyme-mediated cleavage and circRNA linearization, e.g., in the presence of a protein ligand. In some embodiments, circRNA linearization is triggered in the cytoplasm, e.g., using an aptamer that associates with the ligand in the cytoplasm. In some embodiments, circRNA linearization is triggered in the nucleus, for example, using an aptamer that associates with a ligand in the nucleus. In several embodiments, the nuclear ligand comprises an epigenetic modifier or a transcription factor. In some embodiments, the ligand that triggers linearization is present at higher levels in on-target cells than in off-target cells.

[1106] It is further contemplated that nucleic acid-responsive ribozyme systems can be used for circRNA linearization. For example, biosensors that detect specific target nucleic acid molecules to trigger ribozyme activation have been described, for example, by Penchovsky (Biotechnology Advances 32(5):1015-1027(2014) (incorporated herein by reference). By these methods, the ribozyme naturally folds in an inactive state and is activated only in the presence of a specified target nucleic acid molecule (e.g., an RNA molecule). In some embodiments, the circRNA of the Gene Writing system comprises a nucleic acid-responsive ribozyme that is activated in the presence of a specified target nucleic acid, such as an RNA, for example, an mRNA, miRNA, guide RNA, gRNA, sgRNA, ncRNA, lncRNA, tRNA, snRNA, or mtRNA. In some embodiments, the nucleic acid that triggers linearization is present at higher levels in on-target cells than in off-target cells.

[1107] In some embodiments of any of the aspects herein, the gene writing system incorporates one or more ribozymes with inducible specificity for a target tissue or cell of interest, e.g., a ribozyme activated by a ligand or nucleic acid that is present at a higher level in the target tissue or cell of interest. In some embodiments, the gene writing system incorporates a ribozyme with inducible specificity for a subcellular compartment, e.g., the nucleus, nucleolus, cytoplasm, or mitochondria. In some embodiments, the ribozyme activated by a ligand or nucleic acid is present at a higher level in the target subcellular compartment. In some embodiments, the RNA component of the gene writing system is provided as, e.g., a circRNA that is activated by linearization. In some embodiments, linearization of the circRNA encoding the gene writing polypeptide activates the molecule for translation. In some embodiments, the signal that activates the circRNA component of the gene writing system is present at a higher level in the on-target cell or tissue, e.g., such that the system is specifically activated in the on-target cell.

[1108] In some embodiments, the RNA component of the gene writing system is provided as a circRNA that is inactivated by linearization. In some embodiments, the circRNA encoding the gene writer polypeptide is inactivated by cleavage and degradation. In some embodiments, the circRNA encoding the gene writing polypeptide is inactivated by cleavage that separates the translation signal from the coding sequence of the polypeptide. In some embodiments, the signal that inactivates the circRNA component of the gene writing system is present at higher levels in off-target cells or tissues, such that the system is specifically inactivated in the off-target cells.

[1109] Evolutionary variant of Gene Writer In some embodiments, the present invention provides evolved variants of Gene Writers. Evolved variants can, in some embodiments, be produced by mutagenizing a standard Gene Writer or one of the fragments or domains contained therein. In some embodiments, one or more of the domains (e.g., catalytic domains or DNA-binding domains (e.g., target-binding domains or template-binding domains), such as sequence-guided DNA-binding elements) are evolved. One or more of these evolved variant domains can, in some embodiments, evolve alone or together with other domains. One or more evolved variant domains can, in some embodiments, be combined with a non-evolved cognate component or an evolved variant of a cognate component (e.g., one that may have evolved in a parallel or sequential manner).

[1110] In some embodiments, the process of mutagenizing a standard Gene Writer or a fragment or domain thereof comprises mutagenizing the standard Gene Writer or a fragment or domain thereof. In embodiments, the mutagenesis comprises, for example, a progressive evolution method (e.g., PACE) or a non-progressive evolution method (e.g., PANCE), as described herein. In some embodiments, the evolved Gene Writer or a fragment or domain thereof (e.g., a DNA-binding domain, e.g., a target-binding domain or a template-binding domain) comprises one or more amino acid mutations introduced into its amino acid sequence compared to the standard Gene Writer or a fragment or domain thereof. In some embodiments, the amino acid sequence mutation may comprise one or more mutated residues (e.g., conservative substitutions, non-conservative substitutions, or a combination thereof) within the amino acid sequence of the standard Gene Writer, for example, as a result of a change in the nucleotide sequence encoding the gene writer resulting in a change in a codon at any particular position within the coding sequence, a deletion of one or more amino acids (e.g., a truncated protein), an insertion of one or more amino acids, or any combination thereof. An evolved variant Gene Writer may contain variants in one or more components or domains of the Gene Writer (e.g., variants introduced into the catalytic domain, the DNA-binding domain, or a combination thereof).

[1111] In some aspects, the present invention provides Gene Writers, systems, kits, and methods that use or include evolved variants of Gene Writers, e.g., evolved variants of Gene Writers or Gene Writers produced or producible by PACE or PANCE. In embodiments, the non-evolved standard Gene Writer is a Gene Writer disclosed herein.

[1112] The term "phage-assisted incremental evolution (PACE)" as used herein generally refers to incremental evolution using phages as viral vectors. Examples of PACE technology include, for example, International PCT Application No. PCT / US2009 / 056194, filed September 8, 2009, and published March 11, 2010, as WO 2010 / 028347; International PCT Application No. PCT / US2011 / 066747, filed December 22, 2011, and published June 28, 2012, as WO 2012 / 088381; U.S. Patent No. 9,023,594, issued May 5, 2015; U.S. Patent No. 9,771,574, issued September 26, 2017; U.S. Patent No. 9,771,574, issued July 19, 2016; No. 9,394,537, filed January 20, 2015, published September 11, 2015 as WO 2015 / 134121; U.S. Pat. No. 10,179,911, issued January 15, 2019; and International PCT Application PCT / US2016 / 027795, filed April 15, 2016, published October 20, 2016 as WO 2016 / 168631, each of which is incorporated herein by reference in its entirety.

[1113] The term "phage-assisted non-gradual evolution (PANCE)" as used herein generally refers to non-gradual evolution using phages as viral vectors. Examples of PANCE technology are found in, for example, Suzuki T. et al., Crystal structures reveal an elusive functional domain of pyrrolysyl-tRNA synthetase, Nat Chem Biol. 13(12):1261-1266 (2017), the entire contents of which are incorporated herein by reference. Briefly, PANCE is a technique for rapid in vivo directed evolution using serial flask transfers of evolutionary selection phage (SP) containing the gene of interest to be evolved into fresh whole host cells (e.g., E. coli cells). The gene inside the host cell can be held constant while the gene contained in the SP is progressively evolved. After phage propagation, an aliquot of infected cells can be used to transfect the next flask containing host E. coli. This process can be repeated and / or continued until the desired phenotype has evolved, for example, for as many transfers as desired.

[1114] Methods for applying PACE and PANCE to Gene Writers will be readily understood by those skilled in the art by reference, inter alia, to the above-mentioned references. For example, further exemplary methods for directing the gradual evolution of genome-modifying proteins or systems, e.g., in a population of host cells, using phage particles can be applied to generate evolved variants of Gene Writers or fragments or subdomains thereof. Non-limiting examples of such methods are described, for example, in International PCT Application PCT / US2009 / 056194, filed September 8, 2009, and published March 11, 2010, as WO 2010 / 028347; International PCT Application PCT / US2009 / 056194, filed December 22, 2011, and published June 28, 2012, as WO 2012 / 088381; , PCT / US Patent Application Publication No. 2011 / 066747; U.S. Patent No. 9,023,594 issued May 5, 2015; U.S. Patent No. 9,771,574 issued September 26, 2017; U.S. Patent No. 9,394,537 issued July 19, 2016; WO 2015 / 134121 filed January 20, 2015, and published September 11, 2015. International PCT application PCT / US Patent Application Publication No. 2015 / 012022, published on January 15, 2019; U.S. Patent No. 10,179,911, published on January 15, 2019; International application PCT / US Patent Application Publication No. 2019 / 37216, filed on June 14, 2019; and International Publication No. WO 2019 / 023680, published on January 31, 2019, 2016. No. PCT / US2016 / 027795, filed April 15, published October 20, 2016 as WO 2016 / 168631, and International Application PCT / US2019 / 47996, filed August 23, 2019, each of which is incorporated herein by reference in its entirety.

[1115] In some non-limiting exemplary embodiments, a method for evolving an evolved mutant Gene Writer or fragment or domain thereof comprises: (a) contacting a population of host cells with a population of viral vectors (the starting Gene Writer or fragment or domain thereof) comprising a gene of interest, where (1) the host cells are suitable for infection with the viral vector; (2) the host cells express viral genes necessary for the production of viral particles; (3) the expression of at least one viral gene necessary for the production of infectious viral particles is dependent on the function of the gene of interest; and / or (4) the viral vector allows for the expression of proteins in the host cells and can be replicated by the host cells and packaged into viral particles. In some embodiments, the method comprises: (b) contacting the host cells with a mutagen using host cells comprising mutations that increase the mutation rate (e.g., by delivering a mutant plasmid or some genomic modification—e.g., a damaged DNA proofreading polymerase, an SOS gene, e.g., UmuC, UmuD′, and / or RecA (such mutations can be under the control of an inducible promoter when bound to the plasmid), or a combination thereof). In some embodiments, the method includes (c) incubating the population of host cells under conditions that allow for viral replication and viral particle production, wherein host cells are removed from the host cell population and fresh, uninfected host cells are introduced into the host cell population, thus replenishing the host cell population and forming a stream of host cells. In some embodiments, the cells are incubated under conditions that allow the gene of interest to acquire a mutation. In some embodiments, the method further includes (d) isolating from the population of host cells a mutated version of the viral vector that encodes an evolved gene product (e.g., an evolved mutant Gene Writer or a fragment or domain thereof).

[1116] Those skilled in the art will appreciate various features that can be used within the above framework. For example, in some embodiments, the viral vector or phage is a filamentous phage, e.g., an M13 phage, e.g., an M13 selection phage. In certain embodiments, the gene required for the production of infectious viral particles is M13 gene III (gIII). In some embodiments, the phage may lack functional gIII but instead contain gI, gII, gIV, gV, gVI, gVII, gVIII, gIX, and gX. In some embodiments, the generation of infectious VSV particles includes the envelope protein VSV-G. In various embodiments, different retroviral vectors can be used, e.g., murine leukemia virus vectors or lentiviral vectors. In some embodiments, retroviral vectors can be efficiently packaged using, for example, VSV-G envelope proteins as a substitute for the virus's native envelope proteins.

[1117] In some embodiments, the host cells are incubated for a suitable number of viral life cycles, e.g., at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1250, at least 1500, at least 1750, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 7500, at least 10000 or more consecutive viral life cycles, where an illustrative, non-limiting example for M13 phage is 10-20 minutes per viral life cycle. Similarly, conditions can be adjusted to control the residence time of the host cells in the population of host cells (e.g., about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 70, about 80, about 90, about 100, about 120, about 150, or about 180 minutes). The host cell population can be adjusted to control the density of the host cells, or in some embodiments, the host cell density in the inflow, e.g., 10 3 cells / ml, approximately 10 4 cells / ml, approximately 10 5 cells / ml, approximately 5-10 5 cells / ml, approximately 10 6 cells / ml, approximately 5-10 6 cells / ml, approximately 10 7 cells / ml, approximately 5-10 7 cells / ml, approximately 10 8 cells / ml, approximately 5-10 8 cells / ml, approximately 10 9 cells / ml, approximately 5-10 9 cells / ml, approximately 10 10 cells / ml or approximately 5-10 10 Cells / ml can be partly controlled.

[1118] nucleic acid Promoter In some embodiments, one or more promoters or enhancers are operably linked to a nucleic acid encoding a Gene Writer polypeptide or a template nucleic acid that controls expression of, for example, a heterologous sequence of interest. In certain embodiments, the one or more promoters or enhancers comprise cell type or tissue-specific elements. In some embodiments, the promoter or enhancer is the same as or derived from the promoter or enhancer that naturally controls expression of the heterologous sequence of interest. For example, an ornithine transcarbomylase promoter and enhancer can be used to control expression of the ornithine transcarbomylase gene in a system or method provided by the present invention for correcting an ornithine transcarbomylase deficiency. In some embodiments, the promoter is a promoter of Table 4B or a functional fragment or variant thereof.

[1119] Exemplary commercially available tissue-specific promoters can be found, for example, at a URL (e.g., https: / / www.invivogen.com / tissue-specific-promoters). In some embodiments, the promoter is a native promoter or a minimal promoter (e.g., one composed of a single fragment from the 5' region of a given gene). In some embodiments, the native promoter comprises a core promoter and its natural 5' UTR. In some embodiments, the 5' UTR comprises an intron. In other embodiments, these comprise composite promoters created by combining promoters of different origins or assembling a distal enhancer with a minimal promoter of the same origin. In some embodiments, the tissue-specific expression control sequence comprises one or more of the sequences in Table 2 or Table 3 of WO2020014209, which is incorporated by reference in its entirety.

[1120] Exemplary cell- or tissue-specific promoters are listed in the table below, and exemplary nucleic acids encoding them are known in the art and readily available using a variety of resources, for example, the NCBI database, including RefSeq., as well as the Eukaryotic Promoter Database (http: / / epd.epfl.ch / / index.php).

[1121] [Table 744]

[1122] [Table 745]

[1123] [Table 746]

[1124] [Table 747]

[1125] [Table 748]

[1126] Depending on the host / vector system utilized, any of a number of suitable transcriptional and translational control elements may be used in the expression vector, including constitutive and inducible promoters, transcriptional enhancer elements, transcriptional terminators, etc. (See, e.g., Bitter et al. (1987) Methods in Enzymology, 153:516-544; incorporated herein by reference in its entirety).

[1127] In some embodiments, the nucleic acid encoding the Gene Writer or the template nucleic acid is operably linked to a control element, e.g., a transcriptional control element such as a promoter. The transcriptional control element may, in some embodiments, be functional in either eukaryotic cells, e.g., mammalian cells; or prokaryotic cells (e.g., bacterial or archaeal cells). In some embodiments, the nucleotide sequence encoding the polypeptide is operably linked to multiple control elements, e.g., that allow expression of the nucleotide sequence encoding the polypeptide in both prokaryotic and eukaryotic cells.

[1128] For purposes of illustration, examples of spatially restricted promoters include, but are not limited to, neuron-specific promoters, adipocyte-specific promoters, cardiomyocyte-specific promoters, smooth muscle-specific promoters, photoreceptor-specific promoters, and the like. Neuron-specific spatially restricted promoters include, but are not limited to, the neuron-specific enolase (NSE) promoter (see, e.g., EMBL HSENO2, X51956); aromatic amino acid decarboxylase (AADC) promoter, neurofilament promoter (see, e.g., GenBank HUMNFL, L04147); synapsin promoter (see, e.g., GenBank HUMSYNIB, M55301); thy-1 promoter (see, e.g., Chen et al. (1987) Cell 51:7-19; and Llewellyn, et al. (2010) Nat. Med. 16(10):1161-1166); serotonin receptor promoter (see, e.g., GenBank S62283); tyrosine hydroxylase promoter (TH) (see, e.g., Oh et al. (2009) Gene Ther 16:437; Sasaoka et al. (1992) Mol. Brain Res. 16:274; Boundy et al. (1998) J. Neurosci. 18:9989; and Kaneda et al. (1991) Neuron 6:583-594); GnRH promoter (see, e.g., Radovick et al. (1991) Proc. Natl. Acad. Sci. USA 88:3402-3406); L7 promoter (see, e.g., Oberdick et al. (1990) Science 248:223-226); DNMT promoter (see, e.g., Bartge et al. (1988) Proc. Natl. Acad. Sci. USA 85:3648-3652); enkephalin promoter (see, e.g., Comb et al. (1988) EMBO J. 17:3793-3805); myelin basic protein (MBP) promoter; Ca2+ calmodulin-dependent protein kinase II-α (CamKIIα) promoter (see, e.g., Mayford et al. (1996) Proc. Natl. Acad. Sci.USA 93:13250; and Casanova et al. (2001) Genesis 31:37); CMV enhancer / platelet-derived growth factor-β promoter (see, e.g., Liu et al. (2004) Gene Therapy 11:52-60); and the like.

[1129] Adipocyte-specific spatially restricted promoters include, but are not limited to, the aP2 gene promoter / enhancer, e.g., the region from -5.4 kb to +21 bp of the human aP2 gene (see, e.g., Tozzo et al. (1997) Endocrinol. 138:1604; Ross et al. (1990) Proc. Natl. Acad. Sci. USA 87:9590; and Pavjani et al. (2005) Nat. Med. 11:797); glucose transporter-4 (GLUT4) promoter (see, e.g., Knight et al. (2003) Proc. Natl. Acad. Sci. USA 100:14725); fatty acid translocase (FAT / CD36) promoter (see, e.g., Kuriki et al. (2005) Nat. Med. 11:797); al. (2002) Biol. Pharm. Bull. 25:1476; and Sato et al. (2002) J. Biol. Chem. 277:15703); the stearoyl-CoA desaturase-1 (SCD1) promoter (Tabor et al. (1999) J. Biol. Chem. 274:20603); the leptin promoter (see, e.g., Mason et al. (1998) Endocrinol. 139:1013; and Chen et al. (1999) Biochem. Biophys. Res. Comm. 262:187); azidonectin promoter (see, e.g., Kita et al. (2005) Biochem. Biophys. Res. Comm. 331:484; and Chakrabarti (2010) Endocrinol. 151:2408); adipsin promoter (see, e.g., Platt et al. (1989) Proc. Natl. Acad. Sci. USA 86:7490); resistin promoter (see, e.g., Seo et al. (2003) Molec. Endocrinol. 17:1522); and the like.

[1130] Cardiomyocyte-specific spatially restricted promoters include, but are not limited to, regulatory sequences derived from the following genes: myosin light chain-2, α-myosin heavy chain, AE3, cardiac troponin C, cardiac actin, etc. (Franz et al. (1997) Cardiovasc. Res. 35:560-566; Robbins et al. (1995) Ann. NY Acad. Sci. 752:492-505; Linn et al. (1995) Circ. Res. 76:584-591; Parmacek et al. (1994) Mol. Cell. Biol. 14:1870-1885; Hunter et al. (1993) Hypertension 22:608-617; and Sartorelli et al. (1992) Proc. Natl. Acad. Sci. USA 89:4047-4051.

[1131] Smooth muscle cell-specific spatially restricted promoters include, but are not limited to, the SM22α promoter (see, e.g., Akyuerek et al. (2000) Mol. Med. 6:983; and U.S. Pat. No. 7,169,874); the smoothelin promoter (see, e.g., WO 2001 / 018048); the α-smooth muscle actin promoter; and the like. For example, a 0.4 kb region of the SM22α promoter (within which two CArG elements are located) has been shown to mediate vascular smooth muscle cell-specific expression (see, e.g., Kim, et al. (1997) Mol. Cell. Biol. 17, 2266-2278; Li, et al. (1996) J. Cell Biol. 132 849-859; and Moessler, et al. (1996) Development 122, 2415-2425).

[1132] Photoreceptor-specific spatially restricted promoters include, but are not limited to, the rhodopsin kinase promoter (Young et al. (2003) Ophthalmol. Vis. Sci. 44:4076); beta phosphodiesterase gene promoter (Nicoud et al. (2007) J. Gene Med. 9:1015); retinitis pigmentosa gene promoter (Nicoud et al. (2007) supra); interphotoreceptor retinoid-binding protein (IRBP) gene enhancer (Nicoud et al. (2007) supra); IRBP gene promoter (Yokoyama et al. (1992) Exp Eye Res. 55:225); and the like.

[1133] Non-limiting exemplary cell-specific promoters Cell-specific promoters known in the art can be used to direct the expression of, for example, the Gene Writer proteins described herein. Non-limiting exemplary mammalian cell-specific promoters have been characterized and used in mice that express Cre recombinase in a cell-specific manner. Some non-limiting exemplary mammalian cell-specific promoters are listed in Table 1 of U.S. Patent No. 9,845,481 (incorporated herein by reference).

[1134] In some embodiments, the cell-specific promoter is a promoter active in plants.Many exemplary cell-specific promoters are known in the art.See, for example, U.S. Patent No. 5,097,025; U.S. Patent No. 5,783,393; U.S. Patent No. 5,880,330; U.S. Patent No. 5,981,727; U.S. Patent No. 7,557,264; U.S. Patent No. 6,291,666; U.S. Patent No. 7,132,526; and U.S. Patent No. 7,323,622; and U.S. Patent Application Publication No. 2010 / 0269226; U.S. Patent Application Publication No. 2007 / 0180580; U.S. Patent Application Publication No. 2005 / 0034192; and U.S. Patent Application Publication No. 2005 / 0086712 (which are incorporated herein by reference in their entirety for all purposes).

[1135] In some embodiments, the vectors described herein comprise an expression cassette. The term "expression cassette," as used herein, refers to a nucleic acid construct comprising sufficient nucleic acid elements for expression of a nucleic acid molecule of the present invention. Typically, an expression cassette comprises a nucleic acid molecule of the present invention operably linked to a promoter sequence. The term "operably linked" refers to the association of two or more nucleic acid fragments on a single nucleic acid fragment such that the function of one is affected by the other. For example, a promoter is operably linked to a coding sequence if it is capable of affecting the expression of the coding sequence (e.g., the coding sequence is under the transcriptional control of the promoter). A coding sequence can be operably linked to a regulatory sequence in a sense or antisense orientation. In certain embodiments, the promoter is a heterologous promoter. The term "heterologous promoter," as used herein, refers to a promoter that is not known to be operably linked to a given coding sequence in nature. In certain embodiments, an expression cassette may include additional elements, such as introns, enhancers, polyadenylation sites, woodchuck response elements (WREs), and / or other elements known to affect the expression level of a coding sequence. A "promoter" typically controls the expression of a coding sequence or functional RNA. In certain embodiments, a promoter sequence includes proximal and more distal upstream elements and may further include enhancer elements. An "enhancer" is typically capable of stimulating promoter activity and may be a native element of the promoter or a heterologous element inserted to enhance the level or tissue specificity of the promoter. In certain embodiments, a promoter is derived in its entirety from a native gene. In certain embodiments, a promoter is composed of various elements from different naturally occurring promoters. In certain embodiments, a promoter comprises a synthetic nucleotide sequence. Those skilled in the art will understand that different promoters may direct the expression of a gene in different tissues or cell types, at different developmental stages, or in response to different environmental conditions or the presence or absence of agents or transcriptional cofactors.Ubiquitous, cell type-specific, tissue-specific, developmental stage-specific, and conditional promoters, such as drug-responsive promoters (e.g., tetracycline-responsive promoters), are well known to those skilled in the art. Examples of promoters include, but are not limited to, the phosphoglycerate kinase (PKG) promoter, CAG (a composite of the CMV enhancer, chicken β-actin promoter (CBA), and rabbit β-globin intron), NSE (neuron-specific enolase), synapsin, or NeuN promoter, the SV40 early promoter, the mouse mammary tumor virus LTR promoter, the adenovirus major late promoter (Ad MLP), herpes simplex virus (HSV) promoter, the cytomegalovirus (CMV) promoter, such as the CMV immediate early promoter region (CMVIE), the SFFV promoter, the Rous sarcoma virus (RSV) promoter, synthetic promoters, and hybridization promoters. Other promoters may be derived from humans or other species, including mice. Common promoters include, for example, the human cytomegalovirus (CMV) immediate-early gene promoter, the SV40 early promoter, the Rous sarcoma virus long terminal repeat, [β]-actin, the rat insulin promoter, the phosphoglycerate kinase promoter, the human α-1 antitrypsin (hAAT) promoter, the transthyretin promoter, the TBG promoter and other liver-specific promoters, the desmin promoter and similar muscle-specific promoters, the EF1-α promoter, the CAG promoter and other constitutive promoters, hybrid promoters with multiple tissue specificities, and neuron-specific promoters such as the synapsin and glyceraldehyde-3-phosphate dehydrogenase promoters, all of which are well known and readily available to those skilled in the art and can be used to achieve high-level expression of the coding sequence of interest. In addition, sequences derived from non-viral genes, such as the mouse metallothionein gene, may also be useful in the present invention.Such promoter sequences are commercially available, for example, from Stratagene (San Diego, Calif.). Further exemplary promoter sequences are described, for example, in WO 2018213786 A1, which is incorporated herein by reference in its entirety.

[1136] In some embodiments, the apolipoprotein E enhancer (ApoE) or a functional fragment thereof is used, for example, to drive expression in the liver. In some embodiments, two copies of the ApoE enhancer or a functional fragment thereof are used. In some embodiments, the ApoE enhancer or a functional fragment thereof is used in combination with a promoter, for example, the human alpha-1 antitrypsin (hAAT) promoter.

[1137] In some embodiments, the regulatory sequence confers tissue-specific gene expression capability. In some cases, the tissue-specific regulatory sequence binds to tissue-specific transcription factors that induce transcription in a tissue-specific manner. Various tissue-specific regulatory sequences (e.g., promoters, enhancers, etc.) are known in the art. Exemplary tissue-specific regulatory sequences include, but are not limited to, the following tissue-specific promoters: liver-specific thyroxine-binding globulin (TBG) promoter, insulin promoter, glucagon promoter, somatostatin promoter, pancreatic polypeptide (PPY) promoter, synapsin-1 (Syn) promoter, creatine kinase (MCK) promoter, mammalian desmin (DES) promoter, α-myosin heavy chain (a-MHC) promoter, or cardiac troponin T (cTnT) promoter. Other exemplary promoters include the β-actin promoter, the hepatitis B virus core promoter, Sandig et al., Gene Ther., 3:1002-9 (1996); the α-fetoprotein (AFP) promoter, Arbuthnot et al., Hum. Gene Ther., 7:1503-14 (1996)), the bone osteocalcin promoter (Stein et al., Mol. Biol. Rep., 24:185-96 (1997)); the bone sialoprotein promoter (Chen et al., J. Bone Miner. Res., 11:654-64 (1996)), the CD2 promoter (Hansal et al., J. Bone Miner. Res., 11:654-64 (1996)), and the α-fetoprotein (AFP) promoter (Arbuthnot et al., Hum. Gene Ther., 7:1503-14 (1996)). et al., J. Immunol., 161:1063-8(1998); immunoglobulin heavy chain promoter; T cell receptor α-chain promoter, neuronal promoters such as the neuron-specific enolase (NSE) promoter (Andersen et al., Cell. Mol. Neurobiol., 13:503-15(1993)), neurofilament light chain gene promoter (Piccioli et al., Proc. Natl. Acad. Sci. USA, 88:5611-5(1991)), and neuron-specific vgf gene promoter (Piccioli et al., Neuron, 15:373-84(1995)).Other exemplary promoter sequences are described, for example, in U.S. Patent No. 10,300,146, the entire contents of which are incorporated herein by reference. In some embodiments, tissue-specific regulatory elements, e.g., tissue-specific promoters, are selected from those known to be operably linked to genes that are highly expressed in a given tissue, e.g., as determined by RNA-seq or protein expression data, or a combination thereof. Methods for analyzing tissue specificity by expression are taught in Fagerberg et al., Mol Cell Proteomics 13(2):397-406 (2014), the entire contents of which are incorporated herein by reference.

[1138] In some embodiments, the vectors described herein are multicistronic expression constructs. Examples of multicistronic expression constructs include constructs having a first expression cassette (e.g., comprising a first promoter and a first coding nucleic acid sequence) and a second expression cassette (e.g., comprising a second promoter and a second coding nucleic acid sequence). Such multicistronic expression constructs may be particularly useful in some cases for delivery of non-translated gene products, such as hairpin RNAs, along with polypeptides, e.g., gene writers and gene writer templates. In some embodiments, multicistronic expression constructs may exhibit reduced expression levels of one or more of the included transgenes, for example, due to promoter interference or the nearby presence of incompatible nucleic acid elements. When a multicistronic expression construct is part of a viral vector, the presence of self-complementary nucleic acid sequences may, in some cases, interfere with the formation of structures necessary for viral propagation or packaging.

[1139] In some embodiments, the sequence encodes an RNA having a hairpin. In some embodiments, the hairpin RNA is a guide RNA, a template RNA, an shRNA, or a microRNA. In some embodiments, the first promoter is an RNA polymerase I promoter. In some embodiments, the first promoter is an RNA polymerase II promoter. In some embodiments, the second promoter is an RNA polymerase III promoter. In some embodiments, the second promoter is a U6 or H1 promoter. In some embodiments, the nucleic acid construct comprises AAV construct B1 or B2.

[1140] Without intending to be bound by any particular theory, it is believed that multicistronic expression constructs achieve less than optimal expression levels compared to expression systems containing a single cistron. One of the suggested causes for the reduced expression levels achieved using multicistronic expression constructs containing two or more promoter elements is the phenomenon of promoter interference (see, e.g., Curtin JA, Dane AP, Swanson A, Alexander IE, Ginn S L. Bidirectional promoter interference between two widely used internal heterologous promoters in a late-generation lentiviral construct. Gene Ther. 2008 March;15(5):384-90; and Martin-Duque P, Jezzard S, Kaftansis L, Vassaux G. Direct comparison of the insulating properties of two genetic elements in an adenoviral vector containing two different expression cassettes. Hum Gene Ther. 2004 October;15(10):995-1002; both references are incorporated herein by reference for their disclosure of the phenomenon of promoter interference). In some embodiments, the problem of promoter interference can be overcome by, for example, generating a multicistronic expression construct containing a single promoter driving transcription of multiple coding nucleic acid sequences separated by internal ribosome entry sites, or by separating cistrons containing their own promoters with transcriptional insulator elements. In some embodiments, polycistronic expression driven by a single promoter can result in uneven expression levels of the cistrons. In some embodiments, promoters cannot be efficiently separated, and separating elements may not be compatible with some gene transfer vectors, such as some retroviral vectors.

[1141] microRNA MicroRNAs (miRNAs) and other small interfering nucleic acids generally regulate gene expression by target RNA transcript cleavage / degradation or translational repression of target messenger RNAs (mRNAs). In some cases, miRNAs can be naturally expressed, typically as final 19-25 untranslated RNA products. miRNAs generally exert their activity through sequence-specific interactions with the 3' untranslated region (UTR) of target mRNAs. These endogenously expressed miRNAs can form hairpin precursors, which are subsequently processed into miRNA duplexes and then mature single-stranded miRNA molecules. This mature miRNA generally guides a multiprotein complex, miRISC, which identifies the target 3' UTR region of the target mRNA based on its complementarity to the mature miRNA. Useful transgene products can include, for example, miRNAs or miRNA-binding sites that regulate the expression of linked polypeptides. A non-limiting list of miRNA genes; the products of these genes and their homologs; are useful as transgenes or as targets for small interfering nucleic acids (e.g., miRNA sponges, antisense oligonucleotides) in methods such as those listed in U.S. Pat. No. 10,300,146, 22:25-25:48, incorporated herein by reference. In some embodiments, one or more binding sites for one or more of the aforementioned miRNAs are incorporated into a transgene, e.g., a transgene delivered by an rAAV vector, to inhibit expression of the transgene in one or more tissues of an animal carrying the transgene. In some embodiments, binding sites may be selected to control transgene expression in a tissue-specific manner. For example, a binding site for liver-specific miR-122 can be incorporated into a transgene to inhibit expression of the transgene in the liver. Other exemplary miRNA sequences are described, for example, in U.S. Pat. No. 10,300,146, incorporated herein by reference in its entirety.

[1142] miR inhibitors or miRNA inhibitors are generally agents that block miRNA expression and / or processing. Examples of such agents include, but are not limited to, microRNA antagonists, microRNA-specific antisense, microRNA sponges, and microRNA oligonucleotides (double-stranded, hairpin, short oligonucleotides), which inhibit miRNA interaction with the Drosha complex. MicroRNA inhibitors, e.g., miRNA sponges, can be expressed in cells from transgenes (see, e.g., Ebert, MS, Nature Methods, Epub 2012). (As described in Aug. 12, 2007; incorporated herein by reference in its entirety). In some embodiments, microRNA sponges or other miR inhibitors are used in conjunction with AAV. MicroRNA sponges generally specifically inhibit miRNAs via complementary heptamer seed sequences. In some embodiments, a single sponge sequence can be used to silence an entire family of miRNAs. Other methods of silencing miRNA function (derepressing miRNA targets) in cells will be apparent to those skilled in the art.

[1143] In some embodiments, the miRNAs described herein comprise a sequence listed in Table 4 of WO2020014209, which is incorporated herein by reference. The list of exemplary miRNAs from WO2020014209 is also incorporated herein by reference.

[1144] In some embodiments, it is advantageous to silence components of the Gene Writing System (e.g., nucleic acids encoding Gene Writer polypeptides, nucleic acids encoding transgenes) in a subset of cells, hi some embodiments, it is advantageous to restrict expression of components of the Gene Writing System to select cell types within a tissue of interest.

[1145] For example, in a given tissue, such as the liver, macrophages and immune cells, such as Kupffer cells in the liver, are known to be involved in the uptake of delivery vehicles for one or more components of the gene writing system. In some embodiments, at least one binding site for at least one miRNA highly expressed in macrophages and immune cells, such as Kupffer cells, is included in at least one component of the gene writing system, such as a nucleic acid encoding a gene writing polypeptide or transgene. In some embodiments, the miRNA targeting one or more binding sites is listed in the tables referenced herein, for example, miR-142, such as mature miRNA hsa-miR-142-5p or hsa-miR-142-3p.

[1146] In some embodiments, it may be beneficial to reduce Gene Writer levels and / or Gene Writer activity in cells where Gene Writer expression or overexpression of a transgene may have toxic effects. For example, it has been shown that delivery of a transgene overexpression cassette to dorsal root ganglion neurons can result in gene therapy toxicity (see Hordeaux et al. Sci Transl Med 12(569):eaba9188(2020) (incorporated herein by reference in its entirety). In some embodiments, at least one miRNA binding site can be incorporated into the nucleic acid component of a Gene Writing system to reduce expression of a system component in a neuron, e.g., a dorsal root ganglion neuron. In some embodiments, the at least one miRNA binding site incorporated into the nucleic acid component of a Gene Writing system to reduce expression of a system component in a neuron is a binding site for miR-182, e.g., mature miRNA hsa-miR-182-5p or hsa-miR-182-3p. In some embodiments, Gene Writer binding sites can be incorporated into the nucleic acid component of a Gene Writing system to reduce expression of a system component in a neuron. At least one miRNA binding site incorporated into the nucleic acid component of the writing system is a binding site for miR-183, such as the mature miRNA hsa-miR-183-5p or hsa-miR-183-3p. In some embodiments, a combination of miRNA binding sites can be used to enhance the restriction of expression of one or more components of the gene writing system to a tissue or cell type of interest.

[1147] The table below provides exemplary miRNAs and corresponding expressing cells, e.g., miRNAs that, in some embodiments, can incorporate binding sites (complementary sequences) in transgenes or polypeptide nucleic acids to reduce expression in their off-target cells.

[1148] [Table 749]

[1149] 5'UTR and 3'UTR In certain embodiments, a nucleic acid comprising an open reading frame encoding a Gene Writer polypeptide (e.g., as described herein) comprises a 5' UTR and / or a 3' UTR. In some embodiments, the 5' UTR and 3' UTR for protein expression, e.g., an mRNA (or DNA encoding an RNA) for a Gene Writer polypeptide or a heterologous sequence of interest, comprise optimized expression sequences. In some embodiments, the 5'UTR comprises GGGAAAAUAAGAGAGAAAAGAAGAGUAAGAAGAAAUAUAAGAGCCACC (SEQ ID NO: 3475) and / or the 3'UTR comprises UGAUAAUAGGCUGGAGCCUCGGUGGCCAUGCUUCUUGCCCCUUGGGCCUCCCCCCAGCCCCUCCUCCCCUUCCUGCACCCGUACCCCCGUGGUCUUUGAAUAAAGUCUGA (SEQ ID NO: 3476), e.g., as described in Richner et al. Cell 168(6):P1114-1125 (2017) (the sequences of which are incorporated herein by reference).

[1150] In some embodiments, the open reading frame of the Gene Writer system, e.g., the ORF of the mRNA (or DNA encoding the mRNA) encoding the Gene Writer polypeptide or one or more ORFs of the mRNA (or DNA encoding the mRNA) of the heterologous sequence of interest, is flanked by 5' and / or 3' untranslated regions (UTRs) that enhance its expression. In some embodiments, the 5' UTR of the mRNA component (or transcript produced from the DNA component) of the system comprises the sequence 5'-GGGAAAUAAGAGAGAAAAGAAGAGUAAGAAGAAAUAUAAGAGCCACC-3' (SEQ ID NO: 3475). In some embodiments, the 3' UTR of the mRNA component (or transcript produced from the DNA component) of the system comprises the sequence 5'-UGAUAAUAGGCUGGAGCCUCGGUGGCCAUGCUUCUUGCCCCUUGGGCCUCCCCCCAGCCCCUCCUCCCCUUCCUGCACCCGUACCCCCGUGGUCUUUGAAUAAAGUCUGA-3' (SEQ ID NO: 3476). This combination of 5'UTR and 3'UTR has been shown by Richner et al. Cell 168(6):P1114-1125 (2017), the teachings and sequences of which are incorporated herein by reference, to result in the desired expression of the operably linked ORF. In some embodiments, the systems described herein include DNA encoding a transcript, where the DNA comprises the corresponding 5'UTR and 3'UTR in the sequences listed above, with T substituted for U. In some embodiments, the DNA vector used to produce the RNA component of the system includes a promoter upstream of the 5'UTR to initiate in vitro transcription, such as a T7, T3, or SP6 promoter. The 5'UTR begins with GGG, a preferred start for optimizing transcription using T7 RNA polymerase.To adjust transcription levels and modify transcription start site nucleotides to accommodate alternative 5'UTRs, Davidson et al. Pac Symp Biocomput 433-443 (2010) teaches T7 promoter variants that meet both of these properties and methods for finding them.

[1151] Viral vectors and their components In addition to being a source of the relevant enzymes or domains described herein, such as the recombinases and DNA-binding domains used in the present invention, e.g., DNA-binding domains from Cre recombinase, λ integrase, or AAVRep proteins, viruses are useful sources of delivery vehicles for the systems described herein. Some enzymes may have multiple activities. In some embodiments, viruses used as a source of the Gene Writer delivery system or its components may be selected from the group described by Baltimore Bacteriol Rev 35(3):235-241 (1971).

[1152] In some embodiments, the virus is selected from a Group I virus, e.g., a DNA virus, which packages dsDNA into virions. In some embodiments, the Group I virus is selected from, e.g., an adenovirus, a herpesvirus, or a poxvirus.

[1153] In some embodiments, the virus is selected from a group II virus, e.g., a DNA virus that packages ssDNA into virions. In some embodiments, the group II virus is selected from, e.g., a parvovirus. In some embodiments, the parvovirus is a dependoparvovirus, e.g., an adeno-associated virus (AAV).

[1154] In some embodiments, the virus is selected from a group III virus, e.g., an RNA virus, and packages dsRNA into virions. In some embodiments, the group III virus is selected from, e.g., a Reovirus. In some embodiments, one or both strands of the dsRNA contained in such virions are coding molecules that can serve directly as mRNA upon transduction of a host cell, e.g., can be directly translated into protein upon transduction of a host cell without the need for any intervening nucleic acid replication or polymerization step.

[1155] In some embodiments, the virus is selected from a Group IV virus, e.g., an RNA virus, and packages ssRNA(+) into virions. In some embodiments, the Group IV virus is sele...

Claims

[Claim 1] The invention described in this specification.