Recombinase compositions and methods of use
Recombinase polypeptides with specific DNA recognition sequences enable high-frequency, site-specific integration of nucleic acid sequences into host genomes, addressing the limitations of existing methods.
Patent Information
- Application Number
- JP2022503494
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-15
- Filing Date
- 2020-07-17
- Publication Date
- 2025-10-23
- Estimated Expiration
- 2040-07-17
AI Technical Summary
Existing methods for integrating nucleic acid sequences into genomes suffer from low frequency and lack of site specificity, and techniques like CRISPR/Cas9 are less effective for large edits, while methods like Cre/loxP require multiple steps.
The use of recombinase polypeptides with specific amino acid sequences and DNA recognition sequences, including parapalindromic sequences, to facilitate targeted insertion of heterologous sequences into host genomes.
Achieves high-frequency and site-specific integration of exogenous genetic elements into host genomes, improving the efficiency and precision of genetic modification.
Smart Images

Figure 0007759311000179 
Figure 0007759311000180 
Figure 0007759311000001
Abstract
Description
[Technical Field]
[0001] Related Applications This application claims priority to U.S. Patent Application No. 62 / 876,165, filed July 19, 2019, and U.S. Patent Application No. 63 / 039,328, filed June 15, 2020, the entire contents of each of which are incorporated herein by reference.
[0002] Sequence Listing This application contains an electronically filed Sequence Listing in ASCII format, the entire contents of which are incorporated herein by reference. The ASCII copy was created on July 16, 2020, is designated V2065-7003WO_SL.txt, and is 2,102,102 bytes in size. [Background technology]
[0003] Integration of a nucleic acid of interest into a genome occurs at low frequency and with little site specificity in the absence of specialized proteins to facilitate the insertion event. Some existing techniques, such as CRISPR / Cas9, are more suitable for small edits and are less effective at integrating long sequences. Other existing techniques, such as Cre / loxP, require a first step of inserting a loxP site into the genome, followed by a second step of inserting a sequence of interest into the loxP site. There is a need in the art for improved compositions (e.g., proteins and nucleic acids) and methods for inserting, modifying, or deleting a sequence of interest within a genome. Summary of the Invention [Means for solving the problem]
[0004] The present disclosure relates to novel compositions, systems, and methods for modifying the genome of one or more locations in a host cell, tissue, or subject, either in vivo or in vitro. In particular, the present invention features compositions, systems, and methods for introducing exogenous genetic elements into a host genome using recombinase polypeptides (e.g., tyrosine recombinases, e.g., as described herein).
[0005] Enumeration of Embodiments 1. The following: a) a recombinase polypeptide comprising an amino acid sequence of Table 1 or 2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or a nucleic acid encoding a recombinase polypeptide; and b) Below: (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), the DNA recognition sequence having a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 10 to 30, 12 to 27, or 10 to 15 nucleotides, e.g., about 13 nucleotides, and the first and second parapalindromic sequences together comprise a parapalindromic region of a nucleotide sequence in Table 1, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or having no more than 1, 2, 3, or 4 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; and a DNA recognition sequence further comprising a core sequence of about 5 to 10 nucleotides, for example, about 8 nucleotides, the core sequence being located between a first and a second parapalindromic sequence; (ii) a heterologous sequence of interest; double-stranded intercalated DNA containing A system for modifying DNA, comprising:
[0006] 2. The following: a) a recombinase polypeptide comprising an amino acid sequence of Table 1 or 2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or a nucleic acid encoding such a recombinase polypeptide; and b) Below: (i) a human first parapalindromic sequence and a human second parapalindromic sequence of Table 1 that bind to the recombinase polypeptide of (a); (ii) optionally, a heterologous sequence of interest; Insert DNA containing A system for modifying DNA, comprising:
[0007] 3. The system of embodiment 1 or 2, wherein said recombinase polypeptide comprises an amino acid sequence having at least 70% sequence identity to an amino acid sequence in Table 2.
[0008] 4. The system of embodiment 1 or 2, wherein said recombinase polypeptide comprises an amino acid sequence having at least 75% sequence identity to an amino acid sequence in Table 2.
[0009] 5. The system of embodiment 1 or 2, wherein said recombinase polypeptide comprises an amino acid sequence having at least 80% sequence identity to an amino acid sequence in Table 2.
[0010] 6. The system of embodiment 1 or 2, wherein said recombinase polypeptide comprises an amino acid sequence having at least 85% sequence identity to an amino acid sequence in Table 2.
[0011] 7. The system of embodiment 1 or 2, wherein said recombinase polypeptide comprises an amino acid sequence having at least 90% sequence identity to an amino acid sequence in Table 2.
[0012] 8. The system of embodiment 1 or 2, wherein said recombinase polypeptide comprises an amino acid sequence having at least 95% sequence identity to an amino acid sequence in Table 2.
[0013] 9. The system of embodiment 1 or 2, wherein said recombinase polypeptide comprises an amino acid sequence having at least 96% sequence identity to an amino acid sequence in Table 2.
[0014] 10. The system of embodiment 1 or 2, wherein said recombinase polypeptide comprises an amino acid sequence having at least 97% sequence identity to an amino acid sequence in Table 2.
[0015] 11. The system of embodiment 1 or 2, wherein said recombinase polypeptide comprises an amino acid sequence having at least 98% sequence identity to an amino acid sequence in Table 2.
[0016] 12. The system of embodiment 1 or 2, wherein said recombinase polypeptide comprises an amino acid sequence having at least 99% sequence identity to an amino acid sequence in Table 2.
[0017] 13. The system of embodiment 1 or 2, wherein said recombinase polypeptide comprises an amino acid sequence having 100% sequence identity to an amino acid sequence in Table 2.
[0018] 14. A system described in any one of embodiments 1 to 13, wherein (a) and (b) are in separate containers.
[0019] 15. The system of any one of embodiments 1 to 13, wherein (a) and (b) are mixed.
[0020] 16. A cell (e.g., a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell; or a prokaryotic cell) comprising a recombinase polypeptide comprising an amino acid sequence of Table 1 or 2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or a nucleic acid encoding a recombinase polypeptide.
[0021] 17. The following: (i) a DNA recognition sequence that binds to a recombinase polypeptide, the DNA recognition sequence comprising a first parapalindromic sequence and a second parapalindromic sequence; each parapalindromic sequence is about 10 to 30, 12 to 27, or 10 to 15 nucleotides, e.g., about 13 nucleotides, and the first and second parapalindromic sequences together comprise a parapalindromic region of a nucleotide sequence of Table 1, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or having no more than 1, 2, 3, or 4 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; a DNA recognition sequence further comprising a core sequence of about 5 to 10 nucleotides, for example, about 8 nucleotides, the core sequence being located between a first and a second parapalindromic sequence; (ii) optionally, a heterologous sequence of interest; 17. The cell of embodiment 16, further comprising an insert DNA comprising:
[0022] 18. The following: (i) a DNA recognition sequence, the DNA recognition sequence comprising a first parapalindromic sequence and a second parapalindromic sequence; each parapalindromic sequence is about 10 to 30, 12 to 27, or 10 to 15 nucleotides, e.g., about 13 nucleotides, and the first and second parapalindromic sequences together comprise a parapalindromic region of a nucleotide sequence of Table 1, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or having no more than 1, 2, 3, or 4 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; a DNA recognition sequence further comprising a core sequence of about 5 to 10 nucleotides, for example, about 8 nucleotides, the core sequence being located between a first and a second parapalindromic sequence; (ii) a heterologous sequence of interest; A cell (e.g., a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell; or a prokaryotic cell), comprising:
[0023] 19. The cell of embodiment 18, wherein the DNA recognition sequence is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides of the heterologous sequence of interest.
[0024] 20. The cell of embodiment 18 or 19, wherein the DNA recognition sequence and the heterologous sequence of interest are intrachromosomal or extrachromosomal.
[0025] 21. A cell according to any one of embodiments 16 to 20, wherein the cell is a eukaryotic cell.
[0026] 22. The cell of embodiment 21, wherein the cell is a mammalian cell.
[0027] 23. The cell of embodiment 22, wherein the cell is a human cell.
[0028] 24. A cell according to any one of embodiments 16 to 20, wherein the cell is a prokaryotic cell (e.g., a bacterial cell).
[0029] 25. An isolated eukaryotic cell comprising a heterologous sequence of interest stably integrated into its genome at a genomic location listed in column 2 or 3 of Table 1.
[0030] 26. The isolated eukaryotic cell of embodiment 25, wherein the cell is an animal cell (e.g., a mammalian cell) or a plant cell.
[0031] 27. The isolated eukaryotic cell of embodiment 26, wherein the mammalian cell is a human cell.
[0032] 28. The isolated eukaryotic cell of embodiment 26, wherein said animal cell is a bovine cell, equine cell, porcine cell, caprine cell, ovine cell, chicken cell, or turkey cell.
[0033] 29. The isolated eukaryotic cell of embodiment 26, wherein the plant cell is a corn cell, a soybean cell, a wheat cell, or a rice cell.
[0034] 30. A method of modifying the genome of a eukaryotic cell (e.g., a mammalian, e.g., a human cell), comprising: a) a recombinase polypeptide comprising an amino acid sequence of Table 1 or 2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or a nucleic acid encoding a recombinase polypeptide; and b) Below: (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), the DNA recognition sequence comprising a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 10 to 30, 12 to 27, or 10 to 15 nucleotides, e.g., about 13 nucleotides, and the first and second parapalindromic sequences together comprising a parapalindromic region of a nucleotide sequence in Table 1, or a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or having no more than 1, 2, 3, or 4 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; a DNA recognition sequence further comprising a core sequence of about 5 to 10 nucleotides, for example, about 8 nucleotides, the core sequence being located between a first and a second parapalindromic sequence; (ii) a heterologous sequence of interest; Insert DNA containing contacting the This provides a method for modifying the genome of the eukaryotic cell.
[0035] 31. A method for inserting a heterologous sequence of interest into the genome of a eukaryotic cell (e.g., a mammalian, e.g., a human cell), comprising: a) a recombinase polypeptide comprising an amino acid sequence of Table 1 or 2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or a nucleic acid encoding a recombinase polypeptide; and b) Below: (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), the DNA recognition sequence comprising a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 10 to 30, 12 to 27, or 10 to 15 nucleotides, e.g., about 13 nucleotides, and the first and second parapalindromic sequences together comprising a parapalindromic region of a nucleotide sequence in Table 1 or 2; a DNA recognition sequence further comprising a core sequence of about 5 to 10 nucleotides, for example, about 8 nucleotides, the core sequence being located between a first and a second parapalindromic sequence; (ii) a heterologous sequence of interest; Insert DNA containing contacting the This includes, for example, inserting the heterologous sequence of interest into the genome of the eukaryotic cells at a frequency of at least about 0.1% (e.g., at least about 0.1%, 0.5%, 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) of a population of the eukaryotic cells, as measured, for example, by the assay of Example 5.
[0036] 32. The method of embodiment 30 or 31, wherein (a) and (b) are administered separately or together.
[0037] 33. The method of embodiment 30 or 31, wherein (a) is administered before, simultaneously with, or after the administration of (b).
[0038] 34. The method of any of embodiments 30-33, wherein (a) comprises a nucleic acid encoding the polypeptide.
[0039] 35. The method of embodiment 34, wherein the nucleic acid of (a) and the insert DNA of (b) are located on the same nucleic acid molecule, for example, on the same vector.
[0040] 36. The method of embodiment 34, wherein the nucleic acid of (a) and the insert DNA of (b) are located on separate nucleic acid molecules.
[0041] 37. The method of any one of embodiments 30 to 36, wherein the cell has only one endogenous DNA recognition sequence that is compatible with the DNA recognition sequence of the inserted DNA.
[0042] 38. The method of any one of embodiments 30 to 36, wherein the cell has two or more endogenous DNA recognition sequences that are compatible with the DNA recognition sequences of the inserted DNA.
[0043] 39. An isolated recombinase polypeptide comprising an amino acid sequence of Table 1 or 2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0044] 40. The isolated recombinase polypeptide of embodiment 39, which comprises at least one insertion, deletion, or substitution relative to the amino acid sequence of Table 1 or 2.
[0045] 41. The isolated recombinase polypeptide of embodiment 40, wherein said synthetic recombinase polypeptide binds to a eukaryotic (e.g., mammalian, e.g., human) genomic locus (e.g., a sequence in Table 1).
[0046] 42. The isolated recombinase polypeptide of embodiment 40 or 41, wherein the synthetic recombinase polypeptide has at least a 2-, 3-, 4-, or 5-fold increase in affinity for the genomic locus compared to the corresponding unmodified amino acid sequence of Table 1 or 2.
[0047] 43. An isolated nucleic acid encoding a recombinase polypeptide comprising an amino acid sequence of Table 1 or 2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0048] 44. The isolated nucleic acid of embodiment 43, encoding a recombinase polypeptide comprising at least one insertion, deletion, or substitution compared to a recombinase polypeptide of Table 1 or 2.
[0049] 45. The isolated nucleic acid of embodiment 43 or 44, which is codon-optimized for mammalian cells, such as human cells.
[0050] 46. The isolated nucleic acid of any of embodiments 43 to 45, further comprising a heterologous promoter (e.g., a mammalian promoter, e.g., a tissue-specific promoter), a microRNA (e.g., a tissue-specific restricted miRNA), a polyadenylation signal, or a heterologous payload.
[0051] 47. The following: (i) a DNA recognition sequence comprising a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 10 to 30, 12 to 27, or 10 to 15 nucleotides, e.g., about 13 nucleotides, and the first and second parapalindromic sequences together comprising a parapalindromic region of a nucleotide sequence in Table 1; the DNA recognition sequence further comprises a core sequence of about 5 to 10 nucleotides, for example, about 8 nucleotides, and the core sequence is located between a first and a second parapalindromic sequence; (ii) a heterologous sequence of interest; An isolated nucleic acid (e.g., DNA) comprising:
[0052] 48. The isolated nucleic acid of embodiment 47, which binds to a recombinase polypeptide of Table 1 or 2.
[0053] 49. A method for producing a recombinase polypeptide, the method comprising: a) providing a nucleic acid encoding a recombinase polypeptide comprising an amino acid sequence of Table 1 or 2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and b) introducing the nucleic acid into a cell (e.g., a eukaryotic or prokaryotic cell, e.g., as described herein) under conditions that allow production of the recombinase polypeptide; This provides a method for producing the recombinase polypeptide.
[0054] 50. A method for producing a recombinase polypeptide, the method comprising: a) providing a cell (e.g., a eukaryotic or prokaryotic cell) containing a nucleic acid encoding a recombinase polypeptide comprising an amino acid sequence of Table 1 or 2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and b) incubating the cells under conditions that allow production of the recombinase polypeptide; This provides a method for producing the recombinase polypeptide.
[0055] 51. A method for generating an insert DNA containing a DNA recognition sequence and a heterologous sequence, comprising: a) Below: (i) a DNA recognition sequence that binds to a recombinase polypeptide comprising an amino acid sequence of Table 1 or 2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein the DNA recognition sequence comprises a first parapalindromic sequence and a second parapalindromic sequence, each parapalindromic sequence being about 10 to 30, 12 to 27, or 10 to 15 nucleotides, e.g., about 13 nucleotides, and wherein the first and second parapalindromic sequences together comprise the parapalindromic region of a nucleotide sequence of Table 1; a DNA recognition sequence further comprising a core sequence of about 5 to 10 nucleotides, for example, about 8 nucleotides, the core sequence being located between a first and a second parapalindromic sequence; (ii) a heterologous sequence of interest; providing a nucleic acid comprising: b) introducing the nucleic acid into a cell (e.g., a eukaryotic or prokaryotic cell, e.g., as described herein) under conditions that allow replication of the nucleic acid; This is a method for producing the above-mentioned insert DNA.
[0056] 52. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the recombinase polypeptide comprises at least one insertion, deletion, or substitution compared to the amino acid sequence of Table 1 or 2.
[0057] 53. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the recombinase polypeptide comprises a truncation at the N-terminus, C-terminus, or both the N- and C-terminus, compared to the amino acid sequence of Table 1 or 2.
[0058] 54. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any of the preceding embodiments, wherein the recombinase polypeptide comprises a nuclear localization sequence, e.g., an endogenous nuclear localization sequence or a heterologous nuclear localization sequence.
[0059] 55. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the heterologous sequence of interest is inserted into the genome of the cells with an efficiency of at least about 0.1% (e.g., at least about 0.1%, 0.5%, 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) of a population of the cells, e.g., as measured in the assay of Example 5.
[0060] 56. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the heterologous sequence of interest is inserted into a site within the genome of the cell (e.g., a locus listed in column 4 of Table 1, e.g., corresponding to a recombinase row listed in column 1 of Table 1) in at least about 1% (e.g., at least about 1%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 100%) of insertion events, e.g., as measured by the assay of Example 4.
[0061] 57. In the population of cells (e.g., contacted with the system), the heterologous sequence of interest is present in the genome of at least about 1%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 100% of the cells in the population, as measured, for example, by the assay of Example 4. 0, e.g., 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, 2-10, 2-5, 2-4, 3-10, 3-5, or 5-10 sites (e.g., a locus listed in column 4 of Table 1, e.g., corresponding to the row of the recombinase listed in column 1 of Table 1).
[0062] 58. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein in a population of cells contacted with the system, the heterologous sequence of interest is inserted into exactly one site within the genome of the cells (e.g., a locus listed in column 4 of Table 1, e.g., corresponding to a recombinase row listed in column 1 of Table 1) in at least about 1%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 100% of the cells in the population, e.g., as measured by the assay of Example 4.
[0063] 59. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the heterologous sequence of interest is inserted at 1 to 10, e.g., 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 2 to 10, 2 to 5, 2 to 4, 3 to 10, 3 to 5, or 5 to 10 sites within the genome of the cell (e.g., a locus listed in column 4 of Table 1, e.g., corresponding to a row of the recombinase listed in column 1 of Table 1), e.g., as determined by the assay of Example 4.
[0064] 60. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the recombinase polypeptide is linked to the insert DNA.
[0065] 61. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the recombinase polypeptide is provided by providing a nucleic acid encoding the recombinase polypeptide.
[0066] 62. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, resulting in an insertion frequency of the heterologous sequence of interest into the genome of at least about 0.1% (e.g., at least about 0.1%, 0.5%, 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) of a population of said cells, e.g., as measured in the assay of Example 5.
[0067] 63. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the first parapalindromic sequence comprises a sequence having the first 10 to 30, 12 to 27, or 10 to 15, e.g., 10, 11, 12, 13, 14, or 15, nucleotides of the nucleotide sequence in the second or third column of Table 1, or a sequence having no more than 1, 2, or 3 substitutions, insertions, or deletions thereto.
[0068] 64. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of embodiment 63, wherein the second parapalindromic sequence further comprises a second sequence having the last 10 to 30, 12 to 27, or 10 to 15, e.g., 10, 11, 12, 13, 14, or 15, nucleotides of the same nucleotide sequence in the second or third column of Table 1, or a sequence having no more than 1, 2, or 3 substitutions, insertions, or deletions thereto.
[0069] 65. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the insert DNA further comprises a core sequence having 8 nucleotides located between the parapalindromic regions in column 3 of Table 1, or a sequence having no more than 1, 2, or 3 substitutions, insertions, or deletions thereto.
[0070] 66. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the first and second parapalindromic sequences comprise completely palindromic sequences.
[0071] 67. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the parapalindromic sequence comprises 1, 2, 3, 4, 5, or 6 non-palindromic positions.
[0072] 68. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the parapalindromic region comprises a 5' region of 10 to 30, 12 to 27, or 10 to 15, e.g., about 13 nucleotides, and / or a 3' region of 10 to 30, 12 to 27, or 10 to 15, e.g., about 13 nucleotides.
[0073] 69. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the first and second parapalindromic sequences are the same length.
[0074] 70. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the core sequence is 5 to 10 nucleotides (e.g., about 8 nucleotides) in length.
[0075] 71. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the core sequence is capable of hybridizing to a corresponding sequence in the human genome, or its reverse complement.
[0076] 72. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the core sequence has at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% identity to a corresponding sequence in the human genome.
[0077] 73. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the core sequence has no more than 1, 2, 3, 4, 5, 6, 7, 8, or 9 mismatches to the corresponding sequence in the human genome.
[0078] 74. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the core sequence, when cleaved by a recombinase, forms sticky ends that can hybridize with corresponding sequences in the human genome.
[0079] 75. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the heterologous sequence of interest comprises a eukaryotic gene, e.g., a mammalian gene, e.g., a human gene, e.g., a blood factor (e.g., genomic factor I, II, V, VII, X, XI, XII, or XIII) or an enzyme, e.g., a lysosomal enzyme, or a synthetic human gene (e.g., a chimeric antigen receptor).
[0080] 76. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the inserted DNA comprises a heterologous sequence of interest and a DNA recognition sequence.
[0081] 77. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the insert DNA comprises a nucleic acid sequence encoding the recombinase polypeptide.
[0082] 78. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the insert DNA and the nucleic acid encoding the recombinase polypeptide are present in separate nucleic acid molecules.
[0083] 79. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the insert DNA and the nucleic acid encoding the recombinase polypeptide are present in the same nucleic acid molecule.
[0084] 80. The insert DNA is: 10. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, further comprising one, two, three, four, five, or all of: (a) an open reading frame, e.g., a sequence encoding a polypeptide, e.g., an enzyme (e.g., a lysosomal enzyme), a blood factor, an exon; (b) non-coding and / or regulatory sequences, e.g., sequences that bind transcriptional modulators, e.g., promoters (e.g., heterologous promoters), enhancers, insulators; (c) splice acceptor site; (d) polyA site; (e) epigenetic modification sites; (f) Gene Expression Unit
[0085] 81. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the insert DNA comprises a plasmid, a viral vector (e.g., a lentiviral vector or an episomal vector), or other self-replicating vector.
[0086] 82. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein said cell does not contain an endogenous human gene contained in a heterologous sequence of interest or a protein encoded by said gene.
[0087] 83. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any of the preceding embodiments, wherein the cell is derived from an organism that does not contain the endogenous human gene comprised by the heterologous sequence of interest, or the protein encoded by said gene.
[0088] 84. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the cell comprises an endogenous human DNA recognition sequence.
[0089] 85. The endogenous human DNA recognition sequence satisfies the following criteria: (i) located >300 kb from cancer-related genes; (ii) located >300 kb from miRNA / other functional small RNAs; (iii) located >50 kb from the 5′ gene end; (iv) located >50 kb from the replication origin; (v) located >50 kb from any ultraconserved element; (vi) have low transcriptional activity (i.e., no mRNA + / - 25 kb); (vii) are not within a copy number variable region; (viii) is in open chromatin; and / or (xi) is unique, e.g., has one copy in the human genome; 85. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of embodiment 84, wherein the recombinase polypeptide is operably linked to or located within a site in the human genome having at least one, two, three, four, five, six, seven, eight, or nine of the following:
[0090] 86. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any of the preceding embodiments, wherein the cell is an animal cell, e.g., a mammalian cell, e.g., a human cell.
[0091] 87. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein the cell is a plant cell.
[0092] 88. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein said cell is not genetically engineered.
[0093] 89. The system, cell, method, isolated recombinase polypeptide, or isolated nucleic acid of any preceding embodiment, wherein said cell does not contain loxP sites.
[0094] 90. The system or method of any of the previous embodiments, wherein the nucleic acid encoding the recombinase polypeptide is a viral vector, e.g., an AAV vector.
[0095] 91. The system or method of any preceding embodiment, wherein the double-stranded insert DNA is in a viral vector, e.g., an AAV vector.
[0096] 92. The system or method of any preceding embodiment, wherein the nucleic acid encoding the recombinase polypeptide is mRNA, and optionally, the mRNA is in a LNP.
[0097] 93. The system or method of any preceding embodiment, wherein the double-stranded insert DNA is not a viral vector, for example, the double-stranded insert DNA is naked DNA or DNA in a transfection reagent.
[0098] 94. The nucleic acid encoding the recombinase polypeptide is in a first viral vector, e.g., a first AAV vector; and
[0013] The system or method of any preceding embodiment, wherein the insert DNA is in a second viral vector, e.g., a second AAV vector.
[0099] 95. The nucleic acid encoding the recombinase polypeptide is mRNA, and optionally, the mRNA is in a LNP; and 10. The system or method of any preceding embodiment, wherein the insert DNA is in a viral vector, e.g., an AAV vector.
[0100] 96. The nucleic acid encoding the recombinase polypeptide is mRNA; 10. The system or method of any preceding embodiment, wherein the double-stranded insert DNA is not in a viral vector, e.g., the double-stranded insert DNA is naked DNA or DNA in a transfection reagent.
[0101] 97. The system or method of any preceding embodiment, wherein the insert DNA has a length of at least 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb.
[0102] 98. The system or method of any preceding embodiment, wherein the inserted DNA does not include an antibiotic resistance gene or any other bacterial gene or moiety.
[0103] 99. The recombinase polypeptide is selected from the group consisting of Rec17 (SEQ ID NO: 1231), Rec19 (SEQ ID NO: 1233), Rec20 (SEQ ID NO: 1234), Rec27 (SEQ ID NO: 1241), Rec29 (SEQ ID NO: 1243), Rec30 (SEQ ID NO: 1244), Rec31 (SEQ ID NO: 1245), Rec32 (SEQ ID NO: 1246), Rec33 (SEQ ID NO: 1247), Rec34 (SEQ ID NO: 1248), Rec35 (SEQ ID NO: 1249), Rec36 (SEQ ID NO: 1250), Rec37 (SEQ ID NO: 1251), Rec38 (SEQ ID NO: 1252), Rec39 (SEQ ID NO: 1253), Rec 338 (SEQ ID NO: 1552), or Rec589 (SEQ ID NO: 1803), or a recombinase polypeptide having an amino acid sequence at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 sequence alterations (e.g., substitutions, insertions, or deletions) thereto.
[0104] 100. The system, cell, polypeptide, nucleic acid, or method of any preceding embodiment, wherein use of the polypeptide, system, or nucleic acid in a reporter gene inversion assay, e.g., the assay of Example 13, results in reporter gene expression in at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 55, or 60% of the cells.
[0105] 101. The reporter gene inversion assay described above comprises: i) introducing the polypeptide, system, or nucleic acid into a test cell population; ii) introducing into the test cell population a nucleic acid comprising, from 5' to 3', a promoter, a first DNA recognition sequence that binds to a recombinase polypeptide, a GFP gene in an antisense orientation, and a second DNA recognition sequence that binds to the recombinase polypeptide (e.g., the first and second DNA recognition sequences each comprise one or more sequences from the same row as the corresponding recombinase polypeptide, column 3 of Table 1); iii) incubating the test cell population for a time sufficient to allow for inversion of the GFP gene, e.g., as described in Example 13, for 2 days at 37°C; and iv) determining the percentage of cells in the test population that exhibit GFP fluorescence (e.g., a GFP fluorescence threshold of at least 1.7x (1.7-fold), 1.8x, 1.9x, 2x, 2.1x, 2.2x, or 2.3x (e.g., 2x) background fluorescence, e.g., as described in Example 13); 10. The system, cell, polypeptide, nucleic acid, or method of any preceding embodiment, comprising:
[0106] 102. The system, cell, polypeptide, nucleic acid, or method of any preceding embodiment, wherein use of the polypeptide, system, or nucleic acid in a reporter gene integration assay, e.g., the assay of Example 14, results in an average reporter gene copy number per cell of at least 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 0.7, 0.8, 0.9, or 0.95.
[0107] 103. The reporter gene integration assay, comprising: i) introducing the polypeptide, system, or nucleic acid into a test cell population; ii) introducing into the test cell population a nucleic acid comprising, from 5' to 3', a first DNA recognition sequence that binds to a recombinase polypeptide, a GFP gene, and a second DNA recognition sequence that binds to the recombinase polypeptide (e.g., the first and second DNA recognition sequences each comprise one or more sequences from the same row as the corresponding recombinase polypeptide and from the third column of Table 1); iii) incubating the test cell population for a time sufficient to allow integration of the GFP gene into the genomic DNA of the test cell population, e.g., for 2-5 days at 37°C, e.g., as described in Example 14; and iv) determining the average copy number of the GFP gene per cell in the genomic DNA of the test cell population (e.g., a threshold copy number that is at least 1.7× (1.7-fold), 1.8×, 1.9×, 2×, 2.1×, 2.2×, or 2.3× (e.g., 2×) the background copy number to be detected, e.g., as described in Example 14); 10. The system, cell, polypeptide, nucleic acid, or method of any preceding embodiment, comprising:
[0108] 104. The system, cell, polypeptide, nucleic acid, or method of any preceding embodiment, wherein the nucleic acid (e.g., isolated nucleic acid), insert DNA (e.g., double-stranded insert DNA), or heterologous sequence of interest comprises an artificial chromosome, e.g., a bacterial artificial chromosome.
[0109] 105. The system, cell, polypeptide, or nucleic acid of any of the preceding embodiments for use as a laboratory or research tool, or in a laboratory or research method.
[0110] 106. The method of any of embodiments 30-38 or 52-104, wherein the method is used as or as part of a laboratory or research method.
[0111] 107. The system, cell, polypeptide, nucleic acid, or method of any of embodiments 105 or 106, wherein the laboratory or research tool or method is used to modify animal cells, such as mammalian cells (e.g., human cells), plant cells, or fungal bacteria.
[0112] 108. A system, cell, polypeptide, nucleic acid, or method according to any of embodiments 105 to 107, wherein the laboratory or research tool or laboratory or research method is used in vitro.
[0113] The present disclosure contemplates any combination of any one or more of the above aspects and / or embodiments, as well as combinations with any one or more of the embodiments described in the detailed description and examples.
[0114] definition Domain: As used herein, the term "domain" refers to a structure of a biomolecule that contributes to a specific function of the biomolecule. A domain can include a continuous region (e.g., a contiguous sequence) or discrete, non-contiguous regions (e.g., non-contiguous sequences) of a biomolecule. Examples of protein domains include, but are not limited to, a nuclear localization sequence, a recombinase domain, a DNA recognition sequence (e.g., that binds to or is capable of binding to a recognition site, as described herein), a tyrosine recombinase N-terminal domain, and a tyrosine recombinase C-terminal domain; an example of a nucleic acid domain is a regulatory domain, e.g., a transcription factor binding domain, a parapalindromic sequence, a parapalindromic region, a core sequence, or a sequence of interest (e.g., a heterologous sequence of interest). In some embodiments, the recombinase polypeptide comprises one or more domains (e.g., a recombinase domain or a DNA recognition domain) of a polypeptide of Table 1 or 2, or a fragment or variant thereof.
[0115] Exogenous: As used herein, the term "exogenous," when used in reference to a biomolecule (such as a nucleic acid sequence or polypeptide), means that the biomolecule has been introduced into a host genome, cell, or organism by human intervention. For example, a nucleic acid that is added to an existing genome, cell, tissue, or subject using recombinant DNA technology or other methods is exogenous to the existing nucleic acid sequence, cell, tissue, or subject.
[0116] Genomic Safe Harbor Site (GSH Site): A genomic safe harbor site is a site within a host genome that can accommodate the integration of new genetic material, such that the inserted genetic element does not cause significant alterations to the host genome that pose a risk to the host cell or organism. GSH sites generally meet one, two, three, four, five, six, seven, eight, or nine of the following criteria: (i) located >300 kb from a cancer-related gene; (ii) located >300 kb from an miRNA / other functional small RNA; (iii) located >50 kb from the 5' end of a gene; (iv) located >50 kb from a replication origin; (v) located >50 kb away from an ultraconserved element; (vi) having low transcriptional activity (i.e., no mRNA + / - 25 kb); (vii) not within a variable copy number region; (viii) located within open chromatin; and / or (ix) having one copy and being unique within the human genome. Examples of GSH sites in the human genome that meet some or all of these criteria include: (i) adenovirus site 1 (AAVS1), the naturally occurring integration site of the AAV virus on chromosome 19; (ii) the chemokine (CC motif) receptor 5 (CCR5) gene, a chemokine receptor gene known as an HIV-1 co-receptor; (iii) the human orthologue of the mouse Rosa26 locus; and (iv) the rDNA locus. Additional GSH sites are known and are described, for example, in Pellenz et al. (2018) (https: / / doi.org / 10.1101 / 396390).
[0117] Heterologous: The term "heterologous," when used to refer to a first element in relation to a second element, means that the first and second elements do not naturally exist in the arrangement described. For example, a heterologous polypeptide, nucleic acid molecule, construct, or sequence refers to (a) a polypeptide, nucleic acid molecule, or portion of a polypeptide or nucleic acid molecule sequence that is not native to the cell in which it is expressed; (b) a polypeptide or nucleic acid molecule or portion of a polypeptide or nucleic acid molecule that has been modified or mutated relative to its natural state; or (c) a polypeptide or nucleic acid molecule that has altered expression compared to native expression levels under similar conditions. For example, heterologous regulatory sequences (e.g., promoters, enhancers) can be used to regulate expression of a gene or nucleic acid molecule in a manner that differs from how the gene or nucleic acid molecule is normally expressed in nature. In certain embodiments, a heterologous nucleic acid molecule can be present in the native host cell genome but can have an altered expression level, a different sequence, or both. In other embodiments, the heterologous nucleic acid molecule may not be endogenous to the host cell or host genome, but instead may be introduced into the host cell by transformation (e.g., transfection, electroporation), where the added molecule may be integrated into the host genome or may exist as extrachromosomal genetic material either transiently (e.g., mRNA) or semi-stable for two or more generations (e.g., episomal viral vectors, plasmids or other self-replicating vectors).
[0118] Mutation or Mutant: The term "mutation," when applied to a nucleic acid sequence, means that nucleotides within a nucleic acid sequence may be inserted, deleted, or changed relative to a reference (e.g., naturally occurring) nucleic acid sequence. A single alteration may be made at a single locus (point mutation), or multiple nucleotides may be inserted, deleted, or changed at a single locus. In addition, one or more alterations may be made at any number of loci within a nucleic acid sequence. Nucleic acid sequences can be mutated by any method known in the art.
[0119] Nucleic acid molecule: Nucleic acid molecule refers to both RNA and DNA molecules, including, but not limited to, cDNA, genomic DNA, and mRNA, and also includes synthetic nucleic acid molecules, such as those chemically synthesized or recombinantly produced, such as from a DNA template, as described herein. Nucleic acid molecules can be double-stranded or single-stranded, circular, or linear. If single-stranded, the nucleic acid molecule can be the sense or antisense strand. Unless otherwise noted, and as an example of all sequences described herein in the general format "SEQ ID NO:1," a nucleic acid containing "SEQ ID NO:1" refers to a nucleic acid having, at least a portion thereof, either (i) the sequence of SEQ ID NO:1 or (ii) a sequence complementary to SEQ ID NO:1. The choice between the two is determined by the context in which SEQ ID NO:1 is used. For example, if the nucleic acid is used as a probe, the choice between the two is determined by the requirement that the probe be complementary to the desired target. The nucleic acid sequences of the present disclosure may be chemically or biochemically modified or contain non-natural or derivatized nucleotide bases, as will be readily understood by those of skill in the art. Such modifications include, for example, labels, methylation, substitution of one or more naturally occurring nucleotides with analogs, internucleotide modifications such as uncharged linkages (e.g., methylphosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), pendant moieties (e.g., polypeptides), intercalating agents (e.g., acridines, psoralens, etc.), chelators, alkylating agents, and modified linkages (e.g., α-anomeric nucleic acids, etc.). Also included are synthetic molecules that mimic polynucleotides in their ability to bind to designated sequences through hydrogen bonding and other chemical interactions. Such molecules are known in the art and include, for example, those that substitute peptide linkages for phosphate linkages in the backbone of the molecule. Other modifications can include, for example, analogs in which the ribose ring contains bridging moieties or other structures, such as modifications found in "locked" nucleic acids.
[0120] Gene Expression Unit: A gene expression unit is a nucleic acid sequence comprising at least one regulatory nucleic acid sequence operably linked to at least one effector sequence. A first nucleic acid sequence is operably linked to a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For example, a promoter or enhancer is operably linked to a coding sequence if the promoter or enhancer affects the transcription or expression of the coding sequence. Operably linked DNA sequences can be contiguous or non-contiguous. Where necessary to link two protein coding regions, operably linked sequences can be in the same reading frame.
[0121] Host: The term host genome or host cell, as used herein, refers to a cell and / or its genome into which proteins and / or genetic material have been introduced. These terms refer not only to the particular subject cell and / or genome, but also to the progeny of such a cell and / or the genomes of the progeny of such a cell. Because certain modifications may occur in subsequent generations due to mutations or environmental influences, it is understood that such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term "host cell" as used herein. A host genome or host cell may be an isolated cell or cell line grown in culture, or genomic material isolated from such a cell or cell line, or it may be a host cell or host genome comprising a living tissue or organism. In some cases, the host cell may be an animal cell or a plant cell, for example, as described herein. In certain examples, the host cell may be a bovine cell, equine cell, porcine cell, caprine cell, ovine cell, chicken cell, or turkey cell. In certain examples, the host cell may be a corn cell, soybean cell, wheat cell, or rice cell.
[0122] Recombinase polypeptide: As used herein, a recombinase polypeptide refers to a polypeptide having the functional ability to catalyze a recombination reaction of nucleic acid molecules (e.g., DNA molecules). A recombination reaction can involve, for example, one or more nucleic acid strand breaks (e.g., double-strand breaks), followed by joining of the two nucleic acid strand ends (e.g., sticky ends). In some cases, a recombination reaction involves insertion of an insert nucleic acid into a target site, for example, in a genome or construct. In some cases, a recombinase polypeptide comprises one or more structural elements of a naturally occurring recombinase (e.g., a tyrosine recombinase, e.g., Cre recombinase or Flp recombinase). In certain cases, a recombinase polypeptide comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a recombinase described herein (e.g., listed in Table 1 or 2). In some cases, the recombinase polypeptide has one or more functional characteristics of a naturally occurring recombinase (e.g., a tyrosine recombinase, e.g., Cre recombinase or Flp recombinase). In some cases, the recombinase polypeptide recognizes (e.g., binds to) a recognition sequence in a nucleic acid molecule (e.g., a recognition sequence listed in Table 1 or 2, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto). In some embodiments, the recombinase polypeptide is not active as an isolated monomer. In some embodiments, the recombinase polypeptide catalyzes a recombination reaction in cooperation with one or more recombinase polypeptides (e.g., four recombinase polypeptides per recombination reaction).
[0123] Insert nucleic acid molecule: As described herein, an insert nucleic acid molecule (e.g., insert DNA) is a nucleic acid molecule (e.g., a DNA molecule) that is or will be inserted, at least in part, into a target site within a target nucleic acid molecule (e.g., genomic DNA). An insert nucleic acid molecule can include, for example, a nucleic acid sequence that is heterologous to the target nucleic acid molecule (e.g., genomic DNA). In some cases, the insert nucleic acid molecule includes a sequence of interest (e.g., a heterologous sequence of interest). In some cases, the insert nucleic acid molecule is a DNA recognition sequence, e.g., a cognate to DNA present in the target nucleic acid. In some embodiments, the insert nucleic acid molecule is circular, and in some embodiments, the insert nucleic acid molecule is linear. In some embodiments, the insert nucleic acid molecule is also referred to as a template nucleic acid molecule (e.g., template DNA).
[0124] Recognition sequence: A recognition sequence (e.g., a DNA recognition sequence) generally refers to a nucleic acid (e.g., DNA) sequence that is recognized by (e.g., can be bound by) a recombinase polypeptide, e.g., as described herein. In some cases, the recognition sequence includes two parapalindromic sequences, e.g., as described herein. In particular cases, the two parapalindromic sequences together form a parapalindromic region or a portion thereof. In some cases, the recognition sequence further includes a core sequence located between the two parapalindromic sequences, e.g., as described herein. In some cases, the recognition sequence includes a nucleic acid sequence listed in Table 1, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0125] Core sequence: As used herein, a core sequence refers to a nucleic acid sequence located between two parapalindromic sequences. In some cases, the core sequence can be cleaved by a recombinase polypeptide (e.g., a recombinase polypeptide that recognizes a recognition sequence comprising two parapalindromic sequences) to form, for example, sticky ends. In some embodiments, the core sequence is about 5 to 10 nucleotides, e.g., about 8 nucleotides in length.
[0126] Sequence of interest: As used herein, the term sequence of interest refers to a nucleic acid segment that can be desirably inserted into a target nucleic acid molecule, e.g., by a recombinase polypeptide, e.g., as described herein. In some embodiments, the inserted DNA includes a DNA recognition sequence and a sequence of interest heterologous to the DNA recognition sequence (commonly referred to herein as a "heterologous sequence of interest"). A sequence of interest may, in some cases, be heterologous to the nucleic acid molecule into which it is inserted. In some cases, the sequence of interest includes a nucleic acid sequence encoding a gene (e.g., in a eukaryotic cell, e.g., a mammalian cell, e.g., a human gene) or other cargo of interest (e.g., a sequence encoding a functional RNA, e.g., an siRNA or miRNA), e.g., as described herein. In particular cases, the gene encodes a polypeptide (e.g., a blood factor or enzyme). In some cases, the sequence of interest includes one or more nucleic acid sequences encoding a selectable marker (e.g., an auxotrophic marker or an antibiotic marker) and / or a nucleic acid control element (e.g., a promoter, enhancer, silencer, or insulator).
[0127] Parapalindrome: As used herein, the term parapalindrome refers to a characteristic of a pair of nucleic acid sequences in which one of the nucleic acid sequences is either a palindrome relative to the other nucleic acid sequence, or has at least 50% (e.g., at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with the other nucleic acid sequence, or has 1, 2, 3, 4, 5, 6, 7, or 8 or fewer sequence mismatches with the other nucleic acid sequence. A "parapalindromic sequence," as used herein, refers to at least one of a pair of nucleic acid sequences that is parapalindromic with respect to each other. A "parapalindromic region," as used herein, refers to a nucleic acid sequence, or portion thereof, that contains two parapalindromic sequences. In some cases, a parapalindromic region comprises two parapalindromic sequences flanking a nucleic acid sequence segment (eg, comprising a core sequence). [Brief explanation of the drawings]
[0128] [Figure 1] A schematic diagram of an exemplary recombinase reporter plasmid is shown. An inactive reporter plasmid containing a reverse GFP gene flanked by recombinase recognition sites (e.g., loxP) in the reverse orientation can be activated by the presence of a cognate recombinase (e.g., Cre), resulting in flipping of the GFP gene into an orientation where transcription of the coding sequence is driven by an upstream promoter (e.g., CMV). [Figure 2] A schematic diagram illustrating exemplary recombinase-mediated integration into the human genome is shown. In the top diagram, a recombinase expressed from a recombinase-expressing plasmid recognizes a first target site on the insert DNA plasmid and a second target site in the human genome and catalyzes recombination between these two sites, thereby achieving integration of the insert DNA plasmid into the second target site in the human genome. The bottom diagram shows primer and probe locations for a ddPCR assay to quantify genomic integration events. DETAILED DESCRIPTION OF THE INVENTION
[0129] The present disclosure relates to compositions, systems, and methods for targeting, editing, modifying, or manipulating DNA sequences (e.g., inserting a heterologous DNA sequence of interest at a target site in a mammalian genome), e.g., at one or more locations within the DNA sequence in a cell, tissue, or subject, in vivo or in vitro. The DNA sequence of interest may include, for example, a coding sequence, a regulatory sequence, or a gene expression unit.
[0130] Gene-writer™ Gene Editor The present invention provides recombinase polypeptides (e.g., tyrosine recombinase polypeptides, e.g., as listed in Tables 1 or 2) that can modify or manipulate DNA sequences, e.g., by recombining two DNA sequences that contain cognate recognition sequences to which the recombinase polypeptide can bind. The Gene Writer™ gene editor system, in some embodiments, may include: (A) a polypeptide or a nucleic acid encoding a polypeptide (wherein the polypeptide includes: (i) a domain having recombinase activity and (ii) a domain having DNA-binding functionality (e.g., that binds to or is capable of binding to a recognition sequence, e.g., a DNA recognition sequence, e.g., as described herein); and (B) an insert DNA that includes: (i) a sequence that binds to the polypeptide (e.g., a recognition sequence described herein), and optionally (ii) a sequence of interest (e.g., a heterologous sequence of interest). In some embodiments, the domain having recombinase activity and the domain having DNA-binding functionality are the same domain. For example, a Gene Writer genome editor protein may include a DNA-binding domain and a recombinase domain. In certain embodiments, elements of a Gene Writer™ gene editor polypeptide can be derived from the sequence of a recombinase polypeptide (e.g., tyrosine recombinase), e.g., as described herein, e.g., listed in Table 1 or 2. In some embodiments, a Gene Writer™ gene editor polypeptide can include: The Writer genome editor is combined with a second polypeptide. In some embodiments, the second polypeptide is derived from a recombinase polypeptide (e.g., a tyrosine recombinase), e.g., as described herein, e.g., listed in Table 1 or 2.
[0131] Recombinase polypeptide component of the GeneWriter gene editor system An exemplary family of recombinase polypeptides that can be used in the systems, cells, and methods described herein includes tyrosine recombinases. Generally, tyrosine recombinases are enzymes that catalyze site-specific recombination between two recognition sequences. The two recognition sequences may be, for example, on the same nucleic acid (e.g., DNA) molecule or on two separate nucleic acid (e.g., DNA) molecules. In some embodiments, the tyrosine recombinase polypeptide comprises two domains: an N-terminal domain containing the DNA contact site and a C-terminal domain containing the active site.
[0132] Tyrosine recombinases generally function by simultaneously binding two recombinase polypeptide monomers to each of the recognition sequences, resulting in four monomers participating in a single recombinase reaction. For example, as described in Gaj et al. (2014; Biotechnol. Bioeng. 111(1):1-15; the entire contents of which are incorporated herein by reference), after a pair of tyrosine recombinase monomers binds to their respective recognition sequences, the DNA-binding dimers undergo DNA strand cleavage, strand exchange, and religation to form a Holliday junction intermediate, followed by another round of DNA strand cleavage and ligation to form recombinant strands. Non-limiting examples of tyrosine recombinases include Cre recombinase and Flp recombinase, as well as the recombinase polypeptides listed in Tables 1 and 2.
[0133] Those skilled in the art can determine the nucleic acid and corresponding polypeptide sequences of recombinase polypeptides (e.g., tyrosine recombinases) and their domains by using routine sequence analysis tools such as, for example, the Basic Local Alignment Search Tool (BLAST) or CD-Search for conserved domain analysis. Other sequence analysis tools are well known and can be found, for example, at https: / / molbiol-tools.ca, e.g., https: / / molbiol-tools.ca / Motifs.htm.
[0134] Exemplary Recombinase Polypeptides In some embodiments, a Gene Writer™ gene editor system includes a recombinase polypeptide (e.g., a tyrosine recombinase polypeptide), e.g., as described herein. Generally, a recombinase polypeptide (e.g., a tyrosine recombinase polypeptide) specifically binds to a nucleic acid recognition sequence and catalyzes a recombination reaction at a site within the recognition sequence (e.g., a core sequence within the recognition sequence). In some embodiments, the recombinase polypeptide catalyzes recombination between the recognition sequence, or a portion thereof (e.g., its core sequence), and another nucleic acid sequence (e.g., an insert DNA containing a cognate recognition sequence and, optionally, a sequence of interest, e.g., a heterologous sequence of interest). For example, the recombinase polypeptide (e.g., a tyrosine recombinase polypeptide) catalyzes a recombination reaction that results in the insertion of a sequence of interest, or a portion thereof, into another nucleic acid molecule (e.g., a genomic DNA molecule, e.g., a chromosomal or mitochondrial DNA).
[0135] Table 1 below provides exemplary bidirectional tyrosine recombinase polypeptide amino acid sequences (see column 1) and their corresponding DNA recognition sequences (see columns 2 and 3), which were identified by bioinformatics. Tables 1 and 2 include amino acid sequences that have not previously been identified as bidirectional tyrosine recombinases, and also encompass the corresponding DNA recognition sequences of tyrosine recombinases for which the DNA recognition sequences were previously unknown. The amino acid sequences for each accession number listed in column 1 of Table 1 are incorporated herein by reference in their entirety.
[0136] More specifically, column 2 lists the native DNA recognition sequence (e.g., from bacteria or archaea), and column 3 lists the human DNA recognition sequence corresponding to the recombinase listed in that row. Column 4 indicates the genomic location of the human DNA recognition sequence in column 3. Column 5 lists the safe harbor score of the human DNA recognition sequence, which indicates the number of safe harbor criteria the site satisfies.
[0137] The DNA recognition sequences in Table 1 have the following domains: a first parapalindromic sequence, a core sequence, and a second parapalindromic sequence. While not intending to be bound by theory, in some embodiments, tyrosine recombinases recognize DNA recognition sequences based on the parapalindromic regions (first and second parapalindromic sequences), and there are no specific sequence requirements for the core sequence. Thus, in some embodiments, tyrosine recombinases can insert DNA into a target site in the human genome, where the target site has a core sequence that may be substantially different from the native core sequence or may be entirely derived from the native core sequence. Consequently, Table 1, column 2, includes N at these positions. In some embodiments, the core overlap sequence in the inserted DNA may be selected to match, at least in part, the corresponding sequence in the human genome. In some embodiments, the recombinase has only a single human DNA recognition sequence.
[0138] [Table 1-1]
[0139] [Table 1-2]
[0140] [Table 1-3]
[0141] [Table 1-4]
[0142] [Table 1-5]
[0143] [Table 1-6]
[0144]
Table 1-7
[0145]
Table 1-8
[0146]
Table 1-9
[0147]
Table 1-10
[0148]
Table 1-11
[0149]
Table 1-12
[0150]
Table 1-13
[0151]
Table 1-14
[0152]
Table 1-15
[0153]
Table 1-16
[0154]
Table 1-17
[0155]
Table 1-18
[0156]
Table 1-19
[0157]
Table 1-20
[0158]
Table 1-21
[0159]
Table 1-22
[0160]
Table 1-23
[0161]
Table 1-24
[0162]
Table 1-25
[0163]
Table 1-26
[0164]
Table 1-27
[0165]
Table 1-28
[0166]
Table 1-29
[0167]
Table 1-30
[0168]
Table 1-31
[0169]
Table 1-32
[0170]
Table 1-33
[0171]
Table 1-34
[0172]
Table 1-35
[0173]
Table 1-36
[0174]
Table 1-37
[0175]
Table 1-38
[0176]
Table 1-39
[0177]
Table 1-40
[0178]
Table 1-41
[0179]
Table 1-42
[0180]
Table 1-43
[0181]
Table 1-44
[0182]
Table 1-45
[0183] Non-limiting examples of amino acid sequences of tyrosine recombinases are listed by accession number in Table 1, column 1. Table 1 also lists, in column 2, exemplary native non-human (e.g., bacterial, viral, or archaeal) recognition sequences to which a given exemplary tyrosine recombinase binds. Each of the native recognition sequences listed in Table 1 typically includes three segments: (i) a first parapalindromic sequence, (ii) a spacer (e.g., a core sequence) that generally does not include a defined nucleic acid sequence, and (iii) a second parapalindromic sequence, where the first and second parapalindromic sequences are parapalindromic to each other. Table 1 further lists, in column 3, each of the exemplary recognition sequences for each of the exemplary tyrosine recombinases in the human genome. Generally, the human recognition sequences listed in column 3 of Table 1 each include three segments: (i) a first parapalindromic sequence, (ii) a spacer (e.g., a core sequence) that generally does not include a defined nucleic acid sequence, and (iii) a second parapalindromic sequence, where the first and second parapalindromic sequences are parapalindromic to each other. Table 1 includes, in column 4, the genomic locations of exemplary human recognition sequences in the human genome.
[0184] [Table 2-1]
[0185] [Table 2-2]
[0186] [Table 2-3]
[0187] [Table 2-4]
[0188] [Table 2-5]
[0189]
Table 2-6
[0190]
Table 2-7
[0191]
Table 2-8
[0192]
Table 2-9
[0193]
Table 2-10
[0194]
Table 2-11
[0195]
Table 2-12
[0196]
Table 2-13
[0197]
Table 2-14
[0198]
Table 2-15
[0199]
Table 2-16
[0200]
Table 2-17
[0201]
Table 2-18
[0202]
Table 2-19
[0203]
Table 2-20
[0204]
Table 2-21
[0205]
Table 2-22
[0206]
Table 2-23
[0207]
Table 2-24
[0208]
Table 2-25
[0209]
Table 2-26
[0210]
Table 2-27
[0211]
Table 2-28
[0212]
Table 2-29
[0213]
Table 2-30
[0214]
Table 2-31
[0215]
Table 2-32
[0216]
Table 2-33
[0217]
Table 2-34
[0218]
Table 2-35
[0219]
Table 2-36
[0220]
Table 2-37
[0221]
Table 2-38
[0222]
Table 2-39
[0223]
Table 2-40
[0224]
Table 2-41
[0225]
Table 2-42
[0226]
Table 2-43
[0227]
Table 2-44
[0228]
Table 2-45
[0229]
Table 2-46
[0230]
Table 2-47
[0231]
Table 2-48
[0232]
Table 2-49
[0233]
Table 2-50
[0234]
Table 2-51
[0235]
Table 2-52
[0236]
Table 2-53
[0237]
Table 2-54
[0238]
Table 2-55
[0239]
Table 2-56
[0240]
Table 2-57
[0241]
Table 2-58
[0242]
Table 2-59
[0243]
Table 2-60
[0244]
Table 2-61
[0245]
Table 2-62
[0246]
Table 2-63
[0247]
Table 2-64
[0248]
Table 2-65
[0249]
Table 2-66
[0250]
Table 2-67
[0251]
Table 2-68
[0252]
Table 2-69
[0253]
Table 2-70
[0254]
Table 2-71
[0255]
Table 2-72
[0256]
Table 2-73
[0257]
Table 2-74
[0258]
Table 2-75
[0259]
Table 2-76
[0260]
Table 2-77
[0261]
Table 2-78
[0262]
Table 2-79
[0263]
Table 2-80
[0264]
Table 2-81
[0265]
Table 2-82
[0266]
Table 2-83
[0267]
Table 2-84
[0268]
Table 2-85
[0269]
Table 2-86
[0270]
Table 2-87
[0271]
Table 2-88
[0272]
Table 2-89
[0273]
Table 2-90
[0274]
Table 2-91
[0275] In some embodiments, the recombinase polypeptide (e.g., contained in a system or cell described herein) comprises an amino acid sequence listed in Table 2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity thereto, or having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 sequence alterations (e.g., substitutions, insertions, or deletions) thereto. In some embodiments, the recombinase polypeptide (e.g., comprised in a system or cell described herein) has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to the amino acid sequence of the DNA-binding domain, recombinase normal, N-terminal domain, and / or C-terminal domain of a recombinase polypeptide listed in Table 2, or has no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 sequence alterations (e.g., substitutions, insertions, or deletions) thereto. In some embodiments, the recombinase polypeptide (e.g., comprised in a system or cell described herein) has one or more of the DNA binding and / or recombinase activities of a recombinase polypeptide that includes an amino acid sequence listed in Table 2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity thereto, or an amino acid sequence that has no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 sequence alterations (e.g., substitutions, insertions, or deletions) thereto.
[0276] In some embodiments, the insert DNA (e.g., contained in a system or cell described herein) comprises a nucleic acid recognition sequence listed in column 2 or 3 of Table 1, or a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or having no more than 1, 2, 3, 4, 5, 6, 7, or 8 sequence alterations (e.g., substitutions, insertions, or deletions) thereto. In some embodiments, the insert DNA (e.g., contained in a system or cell described herein) comprises one or more (e.g., both) parapalindromic sequences of the nucleic acid recognition sequences listed in column 2 or 3 of Table 1, or a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or having no more than 1, 2, 3, 4, 5, 6, 7, or 8 sequence alterations (e.g., substitutions, insertions, or deletions) thereto. In some embodiments, the insert DNA (e.g., contained in a system or cell described herein) comprises a spacer (e.g., core sequence) of a nucleic acid recognition sequence listed in column 3 of Table 1, or a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or having no more than 1, 2, 3, 4, 5, 6, 7, or 8 sequence alterations (e.g., substitutions, insertions, or deletions) thereto. In certain embodiments, the insert DNA further comprises a heterologous sequence of interest.
[0277] In some embodiments, the insert DNA (e.g., contained in a system or cell described herein) has a nucleic acid recognition sequence listed in column 2 or 3 of Table 1, or at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or no more than 1, 2, 3, 4, 5, 6, 7, or 8 sequence alterations (e.g., substitutions, insertions, or deletions) thereto. In certain embodiments, the cognate recognition sequence is located in the human genome at a position listed in column 4 of Table 1 (corresponding to the human cognate recognition sequence listed in the same row of column 3).
[0278] In some embodiments, the insert DNA or recombinase polypeptide used in the compositions or methods described herein directs insertion of a heterologous sequence of interest to a location with a safe harbor score of at least 3, 4, 5, 6, 7, or 8. In some embodiments, the insert DNA or recombinase polypeptide used in the compositions or methods described herein directs insertion of a heterologous sequence of interest to a unique genomic safe harbor site that has one copy in the human genome. As an example, a unique site can exist in one copy in a haploid human genome, such that a diploid cell can contain two copies of the site located on a homologous chromosome pair. As another example, a unique site can exist in one copy in a diploid human genome, such that a diploid cell can contain one copy of the site located on only one chromosome of a homologous chromosome pair.
[0279] In some embodiments, the three base pairs in the parapalindromic sequence ("core adjacent motif") immediately adjacent to the core sequence comprise AAA, AGA, ATA, or TAA. In some embodiments, the core adjacent motif comprises at least one A (e.g., two or three A's). In some embodiments, the core adjacent motif is ANA or NAA (where N is any nucleotide). In some embodiments, the DNA recognition sites described herein comprise a first core adjacent motif in a first parapalindromic sequence and a second core adjacent motif in a second parapalindromic sequence. In some embodiments, the first core adjacent motif and the second core adjacent motif have the same nucleotide sequence; in other embodiments, the first core adjacent motif and the second core adjacent motif have different nucleotide sequences.
[0280] In some embodiments, the DNA recognition sequence on the insert DNA has 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more mismatches compared to the human DNA recognition sequence. Without intending to be bound by any particular theory, it is believed that mismatches between DNA recognition sequences may, in some embodiments, bias recombinase activity toward integration over excision, as described, for example, in Araki et al., Nucleic Acids Research, 1997, Vol. 25, No. 4, 868-872 (incorporated herein by reference in its entirety). In some embodiments, the DNA recognition sequence on the insert DNA and / or the human DNA recognition sequence each contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more mismatches compared to the native recognition sequence recognized by the recombinase polypeptide. In certain embodiments, integration of the insert DNA with the human DNA recognition sequence results in the formation of an integrated nucleic acid molecule containing two recognition sequences flanking the integrated sequence (e.g., a heterologous sequence of interest). In certain embodiments, one or both of the two recognition sequences of the integrated nucleic acid molecule contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more mismatches compared to one or more (e.g., one, two, or all three) of the following: (i) the native recognition sequence, (ii) the recognition sequence on the insert DNA, and / or (iii) the human DNA recognition sequence. In some embodiments, the mismatches are all present on the same parapalindromic sequence. In some embodiments, the mismatches are present on different parapalindromic sequences. In some embodiments, one or both of the two recognition sequences of the integrating nucleic acid molecule contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more mismatches compared to the native recognition sequence. In some embodiments, the mismatches are present in the core sequence. In some embodiments, these differences between the recognition sequences of the integrating nucleic acid molecule and the native recognition sequence, the inserted DNA recognition sequence, and / or the human DNA recognition sequence result in a reduced binding affinity between the recombinase polypeptide and the recognition sequences of the integrating nucleic acid molecule compared to the recognition sequences of the integrating nucleic acid molecule and the native recognition sequence.
[0281] In some embodiments, the human recognition sequence (e.g., a human DNA recognition sequence, e.g., as listed in column 3 of Table 1) is located at or near (e.g., within 1, 2, 3, 4, 5, 10, 15, 20, 30, 40, 50, 75, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, or 10,000 nucleotides of) a genomic safe harbor site. In some embodiments, the human DNA recognition sequence is located at a position in the genome that meets one, two, three, four, five, six, seven, eight, or nine of the following criteria: (i) located >300 kb from a cancer-associated gene; (ii) >300 kb from a miRNA / other functional small RNA; (iii) >50 kb from the 5' gene end; (iv) >50 kb from a replication origin; (v) >50 kb from any ultraconserved element; (vi) has low transcriptional activity (i.e., no mRNA + / - 25 kb); (vii) is not within a variable copy number region; (viii) is in open chromatin; and / or (ix) is unique with one copy in the human genome. In some embodiments, a genomic location listed in column 4 of Table 1 is located at or near (e.g., within 1, 2, 3, 4, 5, 10, 15, 20, 30, 40, 50, 75, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, or 10,000 nucleotides of) a genomic safe harbor site. In some embodiments, the genomic locations listed in column 4 of Table 1 are at locations in the genome that meet one, two, three, four, five, six, seven, eight, or nine of the following criteria: (i) located >300 kb from a cancer-associated gene; (ii) >300 kb from a miRNA / other functional small RNA; (iii) >50 kb from the 5' gene end; (iv) >50 kb from a replication origin; (v) >50 kb from any ultraconserved element; (vi) has low transcriptional activity (i.e., no mRNA + / - 25 kb); (vii) is not within a variable copy number region; (viii) is in open chromatin; and / or (xi) is unique with one copy in the human genome.
[0282] In embodiments, the cells or systems described herein comprise one or more (e.g., one, two, or three) of the following: (i) a recombinase polypeptide listed in one row of column 1 of Table 1 or 2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity thereto; (ii) a DNA recognition sequence listed in the second column and same row of Table 1, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or no more than one, two, three, or four sequence modifications thereto (e.g., and / or (iii) a genome comprising a human DNA recognition sequence listed in column 3 and the same line of Table 1, or a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or having no more than 1, 2, 3, or 4 sequence alterations (e.g., substitutions, insertions, or deletions) thereto; preferably, the human DNA recognition sequence is located within the genome at a location listed in column 4 and the same line of Table 1 corresponding to the listing of the human DNA recognition sequence.
[0283] In some embodiments, protein components of the Gene Writing™ system described herein may be pre-bound to a template (e.g., a DNA template). For example, in some embodiments, a Gene Writer™ polypeptide may first be combined with a DNA template to form a deoxyribonucleoprotein (DNP) complex. In some embodiments, DNP can be delivered to cells via, for example, transfection, nucleofection, viruses, vesicles, LNPs, exosomes, or fusosomes. Further description of DNP delivery can be found, for example, in Guha and Calos J Mol Biol (2020), the entire contents of which are incorporated herein by reference.
[0284] In some embodiments, the polypeptides described herein comprise one or more (e.g., 2, 3, 4, 5) nuclear targeting sequences, e.g., nuclear localization sequences (NLSs). In some embodiments, the NLS is a bipartite NLS. In some embodiments, the NLS promotes import of a protein comprising the NLS into a cell nucleus. In some embodiments, the NLS is fused to the N-terminus of a Gene Writer described herein. In some embodiments, the NLS is fused to the C-terminus of a Gene Writer. In some embodiments, the NLS is fused to the N-terminus or C-terminus of a Cas domain. In some embodiments, a linker sequence is disposed between the NLS and an adjacent domain of the Gene Writer.
[0285] In some embodiments, the NLS comprises the amino acid sequence: MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 1822), PKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 1823), RKSGKIAAIWKRPRKPKKKRKV KRTADGSEFESPKKKRKV (SEQ ID NO: 1824), KKTELQTTNAENKTKKL (SEQ ID NO: 1825), or KRGINDRNFWRGENGRKTR (SEQ ID NO: 1826), KRPAATKKAGQAKKKK (SEQ ID NO: 1827), or a functional fragment or variant thereof. Exemplary NLS sequences are also described in PCT / EP 2000 / 011690, the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences.
[0286] In some embodiments, the NLS is a bipartite NLS. A bipartite NLS typically includes two basic amino acid clusters (e.g., about 10 amino acids in length) separated by a spacer sequence. A monopartite NLS typically lacks a spacer. An example of a bipartite NLS is the nucleoplasmin NLS, which has the sequence KR[PAATKKAGQA]KKKK (SEQ ID NO: 1828) (the spacer is in parentheses). Another exemplary bipartite NLS has the sequence: PKKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 1829). Exemplary NLSs are described in WO2020051561 (which is incorporated by reference in its entirety, including the disclosure regarding nuclear localization sequences).
[0287] DNA-binding domain In some embodiments, a recombinase polypeptide (e.g., included in a system or cell described herein), e.g., a tyrosine recombinase, comprises a DNA-binding domain (e.g., a target-binding domain or a template-binding domain).
[0288] In some embodiments, the recombinase polypeptides described herein can be redirected to a defined target site in the human genome. In some embodiments, the recombinases described herein can be fused to a heterologous domain, e.g., a heterologous DNA-binding domain. In some embodiments, the recombinases can be fused to a heterologous DNA-binding domain, e.g., a DNA-binding domain from a zinc finger, TAL, meganuclease, transcription factor, or sequence-guided DNA-binding element. In some embodiments, the recombinases can be fused to a sequence-guided DNA-binding element, e.g., a DNA-binding domain from a CRISPR-associated (Cas) DNA-binding element, e.g., Cas9. In some embodiments, the DNA-binding element fused to the recombinase domain can contain a mutation that inactivates other catalytic functions, e.g., a mutation that inactivates endonuclease activity, e.g., generating an inactive meganuclease, or a mutation that partially or completely inactivates the Cas protein, e.g., generating a nickase Cas9 or an inactive Cas9 (dCas9).
[0289] In some embodiments, the DNA-binding domain comprises Streptococcus pyogenes Cas9 (SpCas9) or a functional fragment or variant thereof. In some embodiments, the DNA-binding domain comprises a modified SpCas9. In embodiments, the modified SpCas9 comprises a modification that alters its protospacer-adjacent motif (PAM) specificity. In embodiments, the PAM has specificity for the nucleic acid sequence 5'-NGT-3'. In embodiments, the modified SpCas9 comprises one or more amino acid substitutions, e.g., at one or more of the following positions: L1111, D1135, G1218, E1219, A1322, or R1335, e.g., selected from the following: L1111R, D1135V, G1218R, E1219F, A1322R, R1335V. In some embodiments, the modified SpCas9 comprises the amino acid substitution T1337R and one or more additional amino acid substitutions selected from the following: L1111, D1135L, S1136R, G1218S, E1219V, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, R1335Q, T1337, T1337L, T1337Q, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337H, T1337Q, and T1337M, or a corresponding amino acid substitution thereof. In some embodiments, the modified SpCas9 comprises: (i) one or more amino acid substitutions selected from the following: D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q, and T1337; and (ii) one or more additional amino acid substitutions selected from the following: L1111R, G1218R, E1219F, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, T1337L, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337R, T1337H, T1337Q, and T1337M, or a corresponding amino acid substitution thereof.
[0290] In some embodiments, the DNA-binding domain comprises a Cas domain, e.g., a Cas9 domain. In several embodiments, the DNA-binding domain comprises a nuclease-active Cas domain, a Cas nickase (nCas) domain, or a nuclease-inactive Cas (dCas). In several embodiments, the DNA-binding domain comprises a nuclease-active Cas9 domain, a Cas9 nickase (nCas9) domain, or a nuclease-inactive Cas9 (dCas). In some embodiments, the DNA-binding domain comprises a Cas9 domain of Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the DNA-binding domain comprises Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the DNA-binding domain comprises S. pyogenes or S. thermophilus Cas9, or a functional fragment thereof. In some embodiments, the DNA-binding domain comprises a Cas9 sequence, e.g., as described in Chylinski, Rhun, and Charpentier (2013) RNA Biology 10:5, 726-737 (incorporated herein by reference). In some embodiments, the DNA-binding domain comprises the HNH nuclease subdomain and / or RuvC1 subdomain of a Cas, e.g., Cas9, or a variant thereof, as described herein. In some embodiments, the DNA-binding domain comprises Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the DNA-binding domain comprises a Cas polypeptide (e.g., an enzyme), or a functional fragment thereof.In some embodiments, the Cas polypeptide (e.g., enzyme) is selected from the following: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (e.g., Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpf l, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Cs y3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1 , Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, C The Cas9 may be selected from sa5, a type II Cas effector protein, a type V Cas effector protein, a type VI Cas effector protein, CARF, DinG, Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12b / C2c1, Cas12c / C2c3, SpCas9(K855A), eSpCas9(1.1), SpCas9-HF1, a hyper accurate Cas9 variant (HypaCas9), homologs thereof, modified or engineered versions thereof, and / or functional fragments thereof. In some embodiments, the Cas9 comprises one or more substitutions selected from the following, for example, H840A, D10A, P475A, W476A, N477A, D1125A, W1126A, and D1127A.In embodiments, the Cas9 comprises one or more mutations at a position selected from the following: D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987, for example, one or more substitutions selected from the following: D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A. In some embodiments, the DNA binding domain is selected from the group consisting of Corynebacterium ulcerans, Corynebacterium diphtheria, Spiroplasma syrphidicola, Prevotella intermedia, Spiroplasma taiwanense, Streptococcus iniae, Belliella baltica, Psychroflexus torquis, Staphylococcus thermophilus, Listeria innocua, Campylobacter jejuni, Neisseria meningitidis, and the like. meningitidis, Streptococcus pyogenes, or Staphylococcus aureus, or functional fragments or variants thereof.
[0291] In some embodiments, the DNA binding domain comprises a Cpf1 domain that includes one or more substitutions, e.g., at positions D917, E1006A, D1255, or any combination thereof, selected from the following: D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, and D917A / E1006A / D1255A.
[0292] In some embodiments, the DNA binding domain comprises spCas9, spCas9-VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL.
[0293] In some embodiments, the DNA-binding domain comprises an amino acid sequence listed in Table 3 below, or an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the DNA-binding domain comprises an amino acid sequence that has no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 differences (e.g., mutations) relative to any of the amino acid sequences described herein.
[0294] [Table 3-1]
[0295] [Table 3-2]
[0296] [Table 3-3]
[0297] [Table 3-4]
[0298] In some embodiments, the Cas polypeptide binds to a gRNA that directs binding of the DNA-binding domain. In some embodiments, the gRNA comprises, for example, from 5' to 3', (1) a gRNA spacer; (2) a gRNA scaffold. In some embodiments, (1) is a Cas9 spacer of approximately 18 to 22 nt, for example, 20 nt. (2) is a gRNA scaffold comprising one or more loops, e.g., one, two, or three loops, for binding the template to the nickase Cas9 domain. In some embodiments, the gRNA scaffold comprises, from 5' to 3', the sequence: GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCC (SEQ ID NO: 1835) Carries out.
[0299] In some embodiments, the Gene Writing System described herein is used to perform edits in HEK293, K562, U2OS, or HeLa cells. In some embodiments, the Gene Writing System is used to perform edits in primary cells, such as primary cortical neurons from E18.5 mice.
[0300] In some embodiments, a system or method described herein comprises a CRISPR DNA targeting enzyme or system described in U.S. Patent Application Publication No. 20200063126, U.S. Patent Application Publication No. 20190002889, or U.S. Patent Application Publication No. 20190002875 (each of which is incorporated by reference herein in its entirety), or a functional fragment or variant thereof. For example, in some embodiments, a GeneWriter polypeptide or Cas endonuclease described herein comprises a polypeptide sequence described in any of the applications listed in this paragraph, and in some embodiments, a guide RNA comprises a nucleic acid sequence described in any of the applications listed in this paragraph.
[0301] In some embodiments, the DNA-binding domain (e.g., the target-binding domain or the template-binding domain) comprises a meganuclease domain, or a functional fragment thereof. In some embodiments, the meganuclease domain has endonuclease activity, e.g., double-strand cleavage and / or nickase activity. In other embodiments, the meganuclease domain has reduced activity, e.g., lacks endonuclease activity, e.g., the meganuclease lacks catalytic activity. In some embodiments, a meganuclease lacking catalytic activity is used as the DNA-binding domain, e.g., as described in Fonfara et al. Nucleic Acids Res 40(2):847-860 (2012), the entire contents of which are incorporated herein by reference. In several embodiments, the DNA-binding domain comprises one or more modifications relative to the wild-type DNA-binding domain, e.g., modifications by directed evolution, e.g., phage-assisted continuous evolution (PACE).
[0302] Intein In some embodiments, as described in more detail below, intein-N can be fused to the N-terminal portion of a polypeptide described herein (e.g., a Gene Writer polypeptide), e.g., at the first domain. In several embodiments, intein-C can also be fused to the C-terminal portion of a polypeptide described herein (e.g., at the second domain), e.g., linking the N-terminal portion to the C-terminal portion, thereby linking the first domain and the second domain. In some embodiments, the first domain and the second domain are each independently selected from a DNA-binding domain and a catalytic domain, e.g., a recombinase domain. In some embodiments, a single domain is disrupted using an intein strategy described herein, e.g., a DNA-binding domain, e.g., a dCas9 domain.
[0303] In some embodiments, the systems or methods described herein include an intein, which is, for example, a self-splicing protein intron (e.g., a peptide) that links flanking N-terminal and C-terminal exteins (e.g., fragments to be joined). Inteins can, in some cases, comprise a fragment of a protein that can be automatically excised to join the remaining fragment (extein) with a peptide bond in a process known as protein splicing. Inteins are also referred to as "protein inons." The process of an intein being automatically excised to join the remaining portion of a protein is referred to herein as "protein splicing" or "intein-mediated protein splicing." In some embodiments, the inteins of a precursor protein (an intein-containing protein prior to intein-mediated protein splicing) are derived from two genes. Such inteins are referred to herein as split inteins (e.g., split intein-N and split intein-C). For example, in cyanobacteria, DnaE, the catalytic subunit of DNA polymerase III, is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene may be referred to herein as "intein-N." The intein encoded by the dnaE-c gene may be referred to herein as "intein-C."
[0304] The use of inteins to link heterologous protein fragments is described, for example, in Wood et al., J. Biol. Chem. 289(21);14512-9(2014), the entire contents of which are incorporated herein by reference. For example, when fused to separate protein fragments, inteins IntN and IntC can recognize each other and splice from themselves and / or simultaneously link the N- and C-terminal exteins of the protein fragments to which they are fused, thereby reconstituting a full-length protein from the two protein fragments.
[0305] In some embodiments, synthetic inteins based on the dnaE intein, Cfa-N (e.g., split intein-N), and Cfa-C (e.g., split intein-C) intein pairs are used. Examples of such inteins are described, for example, in Stevens et al., J Am Chem Soc. 2016 Feb. 24;138(7):2162-5, the entire contents of which are incorporated herein by reference. Non-limiting examples of intein pairs that can be used in accordance with the present disclosure include the following: Cfa DnaE intein, Ssp GyrB intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein, and Cne Prp8 intein (e.g., as described in U.S. Pat. No. 8,394,604, the entire contents of which are incorporated herein by reference).
[0306] In some embodiments, intein-N and intein-C can be fused to the N-terminal portion of split Cas9 and the C-terminal portion of split Cas9, respectively, for linking the N-terminal portion of split Cas9 with the C-terminal portion of split Cas9. For example, in some embodiments, intein-N is fused to the C-terminus of the N-terminal portion of split Cas9, i.e., forming the structure N-[N-terminal portion of split Cas9]-[intein-N]~C. In some embodiments, intein-C is fused to the N-terminus of the C-terminal portion of split Cas9, i.e., forming the structure N-[intein-C]~[C-terminal portion of split Cas9]-C. The mechanism of intein-mediated protein splicing for linking intein-linked proteins (e.g., split Cas9) is described in Shah et al., Chem Sci. 2014;5(1):446-461 (incorporated herein by reference). Methods for designing and using inteins are described, for example, in WO2020051561, WO2014004336, WO2017132580, U.S. Patent Application Publication No. 20150344549, and U.S. Patent Application Publication No. 20180127780 (each of which is incorporated herein by reference in its entirety).
[0307] In some embodiments, split refers to a division into two or more fragments. In some embodiments, the split Cas9 protein or split Cas9 is a Cas9 protein that is provided as an N-terminal fragment and a C-terminal fragment encoded by two separate nucleotide sequences. Polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein can be spliced to form a reconstituted Cas9 protein. In some embodiments, the Cas9 protein is split into two fragments within a denatured region of the protein, as described in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, or as described in Jiang et al. (2016) Science 351:867-871 and PDB file: 5F9R (each of which is incorporated herein by reference in its entirety). The denatured region can be determined by one or more protein structure determination techniques known in the art, including, but not limited to, X-ray crystallography, NMR spectroscopy, electron microscopy (e.g., cryoEM), and / or in silico protein modeling. In some embodiments, the protein is split into two fragments at any C, T, A, or S within the region of SpCas9 between amino acids A292-G364, F445-K483, or E565-T637, or at the corresponding position in any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or other napDNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as protein cleavage.
[0308] In some embodiments, the protein fragments are in the range of about 2 to 1000 amino acids in length (e.g., 2 to 10, 10 to 50, 50 to 100, 100 to 200, 200 to 300, 300 to 400, 400 to 500, 500 to 600, 600 to 700, 700 to 800, 800 to 900, or 900 to 1000 amino acids). In some embodiments, the protein fragments are in the range of about 5 to 500 amino acids in length (e.g., 5 to 10, 10 to 50, 50 to 100, 100 to 200, 200 to 300, 300 to 400, or 400 to 500 amino acids). In some embodiments, protein fragments range in length from about 20 to 200 amino acids (eg, 20 to 30, 30 to 40, 40 to 50, 50 to 100, or 100 to 200 amino acids).
[0309] In some embodiments, for example, a portion or fragment of a Gene Writer polypeptide described herein is fused to an intein. A nuclease can be fused to the N-terminus or C-terminus of an intein. In some embodiments, a portion or fragment of a fusion protein is fused to an intein and also fused to an AAV capsid protein. The intein, nuclease, and capsid protein can be fused together in any configuration (e.g., nuclease-intein-capsid, intein-nuclease-capsid, capsid-intein-nuclease, etc.). In some embodiments, the N-terminus of the intein is fused to the C-terminus of the fusion protein, and the C-terminus of the intein is fused to the N-terminus of an AAV capsid protein.
[0310] In some embodiments, a Gene Writer polypeptide (e.g., comprising a nickase Cas9 domain) is fused to Intein-N, and a polypeptide comprising a polymerase domain is fused to Intein-C.
[0311] Exemplary nucleotide and amino acid sequences of inteins are set forth below: DnaE Intein-N DNA: [ka]
[0312] DnaE Intein-N Protein: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDRGEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNLPN (SEQ ID NO: 1837)
[0313] DnaE Intein-C DNA: ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGATATTGGAGTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAGCTTCTAAT (SEQ ID NO: 1838)
[0314] Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN (SEQ ID NO: 1839)
[0315] Cfa-N DNA: [ka]
[0316] Cfa-N protein: CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQMLPIDEIFERGLDLKQVDGLP (SEQ ID NO: 1841)
[0317] Cfa-C DNA: [ka]
[0318] Cfa-C protein: MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN (SEQ ID NO: 1843)
[0319] Insert DNA In some embodiments, the insert DNA described herein comprises a nucleic acid sequence that can be integrated into a target DNA molecule, e.g., by a recombinase polypeptide (e.g., a tyrosine recombinase polypeptide), e.g., as described herein. The insert DNA is typically capable of binding to one or more recombinase polypeptides (e.g., multiple copies of a recombinase polypeptide) of the system. In some embodiments, the insert DNA comprises a region that can bind to a recombinase polypeptide (e.g., a recognition sequence described herein).
[0320] In some embodiments, the insert DNA comprises a sequence of interest for insertion into the target DNA. The sequence of interest may be coding or non-coding. In some embodiments, the sequence of interest may comprise an open reading frame. In some embodiments, the insert DNA comprises a Kozak sequence. In some embodiments, the insert DNA comprises an internal ribosome entry site. In some embodiments, the insert DNA comprises a self-cleaving peptide, such as a T2A or P2A site. In some embodiments, the insert DNA comprises a start codon. In some embodiments, the insert DNA comprises a splice acceptor site. In some embodiments, the insert DNA comprises a splice donor site. In some embodiments, the insert DNA comprises a microRNA binding site, e.g., downstream of the stop codon. In some embodiments, the insert DNA comprises a polyA tail, e.g., downstream of the stop codon of the open reading frame. In some embodiments, the insert DNA comprises one or more exons. In some embodiments, the insert DNA comprises one or more introns. In some embodiments, the insert DNA comprises a eukaryotic transcription terminator. In some embodiments, the insert DNA comprises an enhanced translation element or a translation-enhancing element. In some embodiments, the insert DNA comprises a microRNA sequence, an siRNA sequence, a guide RNA sequence, or a piwiRNA sequence. In some embodiments, the insert DNA comprises a gene expression unit comprised of at least one regulatory region operably linked to an effector sequence. The effector sequence may be a sequence that is transcribed into RNA (e.g., a coding sequence, such as a sequence encoding a microRNA, or a non-coding sequence). In some embodiments, the sequence of interest may comprise a non-coding sequence. For example, the insert DNA may comprise a promoter or enhancer sequence. In some embodiments, the insert DNA comprises a tissue-specific promoter or enhancer, each of which may be unidirectional or bidirectional. In some embodiments, the promoter is an RNA polymerase I promoter, an RNA polymerase II promoter, or an RNA polymerase III promoter. In some embodiments, the promoter comprises a TATA element. In some embodiments, the promoter comprises a B recognition element.In some embodiments, the promoter has one or more binding sites for a transcription factor.
[0321] In some embodiments, the sequence of interest of the insert DNA is inserted into the target genome within an endogenous intron. In some embodiments, the sequence of interest of the insert DNA is inserted into the target genome, thereby acting as a new exon. In some embodiments, the insertion of the sequence of interest into the target genome results in the replacement of a native exon or the skipping of a native exon. In some embodiments, the sequence of interest of the insert DNA is inserted into a target site in a genomic safe harbor site, such as AAVS1, CCR5, or ROSA26. In some embodiments, the sequence of interest of the insert DNA is added to the genome within an intergenic or intragenic region. In some embodiments, the sequence of interest of the insert DNA is added 5' or 3' to the genome within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb from the endogenous active gene. In some embodiments, the sequence of interest of the inserted DNA is added to the genome 5' or 3' within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb from the endogenous promoter or enhancer. In some embodiments, the target sequence of the insert DNA may be, for example, 50 to 50,000 base pairs (e.g., 50 to 40,000 bp, 500 to 30,000 bp, 500 to 20,000 bp, 100 to 15,000 bp, 500 to 10,000 bp, 50 to 10,000 bp, or 50 to 50,000 bp). In some embodiments, the target sequence of the insert DNA may be 1 to 50 base pairs (e.g., 1 to 10, 10 to 20, 20 to 30, 30 to 40, or 40 to 50 base pairs).
[0322] In certain embodiments, insert DNA can be identified, designed, engineered, and constructed to contain sequences that modify or define genome function in a target cell or organism, for example, by introducing heterologous coding regions into the genome; influencing or effecting exon structure / alternative splicing; causing disruption of endogenous genes; causing transcriptional activation of endogenous genes; effecting epigenetic regulation of endogenous DNA; or effecting up- or down-regulation of operably linked genes. In certain embodiments, insert DNA can contain sequences encoding exons and / or transgenes and can be engineered to provide binding sites for transcription factors such as activators, repressors, enhancers, and combinations thereof. In other embodiments, the coding sequences can be further customized with splice acceptor sites, poly-A tails, etc.
[0323] The insert DNA may have some degree of homology to the target DNA. In some embodiments, the insert DNA has at least 3, 4, 5, 6, 7, 8, 9, 10, or more bases of exact homology to the target DNA or a portion thereof. In some embodiments, the insert DNA has at least 10, 15, 20, 25, 30, 40, 50, 60, 80, 100, 120, 140, 160, 180, 200, or more bases of at least 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% homology to the target DNA or a portion thereof.
[0324] As an alternative to other methods of delivery described herein, in some embodiments, the nucleic acid delivered to the cell (e.g., the nucleic acid encoding the recombinase, or the template nucleic acid, or both) is designed as a minicircle, in which case the plasmid backbone sequences not related to Gene Writing™ are removed prior to administration to the cell. Minicircles have been shown to achieve higher transfection efficiency and gene expression compared to plasmids with backbones containing bacterial portions (e.g., bacterial origins of replication, antibiotic selection cassettes) and have been used to improve transposition efficiency (Sharma et al. Mol Ther Nucleic Acids 2:E74 (2013)). In some embodiments, a DNA vector encoding a Gene Writer™ polypeptide is delivered as a minicircle. In some embodiments, a DNA vector containing a Gene Writer™ template is delivered as a minicircle. In some embodiments of such alternative means for delivering nucleic acids, the bacterial portions are flanked by recombination sites, e.g., attP / attB, loxP, FRT sites. In some embodiments, addition of a cognate recombinase results in intramolecular recombination and excision of the bacterial portion. In some embodiments, the recombinase sites are recognized by phiC31 recombinase. In some embodiments, the recombinase sites are recognized by Cre recombinase. In some embodiments, the recombinase sites are recognized by FLP recombinase. In some embodiments, minicircles are generated in a bacterial production strain, e.g., E. coli stably expressing an inducible minicircle assembling enzyme, e.g., a production strain according to Kay et al. Nat Biotechnol 28(12):1287-1289 (2010). Methods for minicircle DNA vector preparation and generation are described in U.S. Pat. No. 9,233,174, the entire contents of which are incorporated herein by reference.
[0325] In addition to plasmid DNA, minicircles can be generated by excising a desired construct, such as a recombinase expression cassette or a therapeutic expression cassette, from a viral backbone, such as an AAV vector. Previous studies have demonstrated that excision and circularization of inserted DNA sequences from the viral backbone can be important for transposase-mediated integration efficiency (Yant et al., Nat Biotechnol 20(10):999-1005 (2002)). In some embodiments, minicircles are first formulated and then delivered to target cells. In other embodiments, minicircles are formed intracellularly from DNA vectors (e.g., plasmid DNA, rAAV, scAAV, ceDNA, doggie DNA) by co-delivering a recombinase to achieve excision and circularization of nucleic acids flanking the recombinase recognition sites (e.g., nucleic acids encoding Gene Writer™ polypeptides, or DNA templates, or both). In some embodiments, the same recombinase is used for the first excision event (e.g., intramolecular recombination) and the second integration (e.g., target site integration) event. In some embodiments, the recombination site on the excised circular DNA (e.g., after the first recombination event, e.g., intramolecular recombination) is used as a template recognition site for the second recombination (e.g., target site recombination) event.
[0326] Linker In some embodiments, domains of the compositions and systems described herein (e.g., recombinase domains and / or DNA recognition domains of the recombinase polypeptides described herein) can be linked by a linker. Compositions described herein that include a linker element have the general form S1-L-S2, where S1 and S2 can be the same or different and represent two moieties (e.g., polypeptide or nucleic acid domains, respectively) linked to each other by a linker. In some embodiments, a linker can link two polypeptides. In some embodiments, a linker can link two nucleic acid molecules. In some embodiments, a linker can link a polypeptide and a nucleic acid molecule. A linker can be a chemical bond, e.g., one or more covalent or non-covalent bonds. A linker can be flexible, rigid, and / or cleavable. In some embodiments, a linker is a peptide linker. Generally, the peptide linker is at least 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acids in length, for example, 2 to 50 amino acids in length, 2 to 30 amino acids in length.
[0327] The most commonly used flexible linkers have sequences consisting primarily of stretches of Gly and Ser residues ("GS" linkers). Flexible linkers are considered useful for linking domains that require some degree of movement or interaction and may contain small non-polar (e.g., Gly) or polar (e.g., Ser or Thr) amino acids. The incorporation of Ser or Thr can also maintain the stability of the linker in aqueous solution by forming hydrogen bonds with water molecules, thereby reducing unfavorable interactions between the linker and other moieties. Examples of such linkers include those with the structure [GGS] ≧1 or [GGGS] ≧1(SEQ ID NO: 1844). Rigid linkers are useful for maintaining a constant distance between domains while preserving their independent functions. Rigid linkers can also be useful when spatial separation of domains is essential to preserve the stability or biological activity of one or more components in a drug. Rigid linkers can have an α-helical structure or a Pro-rich sequence, (XP)n, where X represents any amino acid, preferably Ala, Lys, or Glu. Cleavable linkers can release free functional domains in vivo. In some embodiments, the linker can be cleaved under certain conditions, for example, in the presence of a reducing agent or protease. In vivo cleavable linkers may utilize the reversibility of disulfide bonds. One example is a thrombin-sensitive sequence (e.g., PRS) between two Cys residues. In vitro thrombin treatment of CPRSC results in cleavage of the thrombin-sensitive sequence, while leaving the reversible disulfide bond intact. Such linkers are well known and are described, for example, in Chen et al. 2013. Fusion Protein Linkers: Property, Design and Functionality. Adv Drug Deliv Rev. 65(10):1357-1369. In vivo cleavage of the linker in the compositions described herein can also be performed by proteases expressed in vivo in specific cells or tissues, under pathological conditions (e.g., cancer or inflammation), or confined within specific cellular compartments. The specificity of various proteases allows for slow cleavage of the linker within the confined compartment.
[0328] In some embodiments, the amino acid linker is an endogenous amino acid (or homologous thereto) present between such domains in a naturally occurring polypeptide. In some embodiments, the endogenous amino acids present between such domains are substituted but the length is unchanged from the natural length. In some embodiments, additional amino acid residues are added to the amino acid residues naturally present between the domains.
[0329] In some embodiments, amino acid linkers are computationally designed or screened to maximize protein function (Anad et al., FEBS Letters, 587:19, 2013).
[0330] Genomic safe harbor sites In some embodiments, the Gene Writer targets a genomic safe harbor site (e.g., directs insertion of the heterologous sequence of interest to a position with a safe harbor score of at least 3, 4, 5, 6, 7, or 8). In some embodiments, the genomic safe harbor site is a Natural Harbor™ site. In some embodiments, the Natural Harbor™ site is derived from the native target of a mobile genetic element, such as a recombinase, transposon, retrotransposon, or retrovirus. The native target of a mobile element may serve as an ideal location for genomic integration, given their evolutionary selection. In some embodiments, the Natural Harbor™ site is ribosomal DNA (rDNA). In some embodiments, the Natural Harbor™ site is 5S rDNA, 18S rDNA, 5.8S rDNA, or 28S rDNA. In some embodiments, the Natural Harbor™ site is a Mutsu site within 5S rDNA. In some embodiments, the Natural Harbor™ site is an R2 site, an R5 site, an R6 site, an R4 site, an R1 site, an R9 site, or an RT site in 28S rDNA. In some embodiments, the Natural Harbor™ site is an R8 site or an R7 site in 18S rDNA. In some embodiments, the Natural Harbor™ site is DNA encoding a transfer RNA (tRNA). In some embodiments, the Natural Harbor™ site is DNA encoding a tRNA-As or a tRNA-Glu. In some embodiments, the Natural Harbor™ site is DNA encoding a spliceosomal RNA. In some embodiments, the Natural Harbor™ site is DNA encoding a small nuclear RNA (snRNA), such as U2 snRNA.
[0331] Thus, in some aspects, the disclosure provides methods for inserting a heterologous sequence of interest into a Natural Harbor™ site using the Gene Writer system described herein. In some embodiments, the Natural Harbor™ site is a site set forth in Table 4 below. In some embodiments, the heterologous sequence of interest is inserted within 20, 50, 100, 150, 200, 250, 500, or 1000 base pairs of the Natural Harbor™ site. In some embodiments, the heterologous sequence of interest is inserted within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb of the Natural Harbor™ site. In some embodiments, the heterologous sequence of interest is inserted at a site having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to a sequence shown in Table 4. In some embodiments, the heterologous sequence of interest is inserted within 20, 50, 100, 150, 200, 250, 500, or 1000 base pairs, or within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb of a site having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to a sequence shown in Table 4. In some embodiments, the heterologous sequence of interest is inserted within a gene set forth in column 5 of Table 4, or within 20, 50, 100, 150, 200, 250, 500, or 1000 base pairs, or within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb of that gene.
[0332] [Table 4-1]
[0333] [Table 4-2]
[0334] [Table 4-3]
[0335] [Table 4-4]
[0336] [Table 4-5]
[0337] [Table 4-6]
[0338] Additional Gene Writer™ Functional Features The Gene Writers described herein may, in some cases, be characterized by one or more functional measurements or characteristics. In some embodiments, the DNA-binding domain (e.g., the target-binding domain) has one or more of the functional characteristics described below. In some embodiments, the template-binding domain has one or more of the functional characteristics described below. In some embodiments, the template (e.g., template DNA) has one or more of the functional characteristics described below. In some embodiments, the template site modified by the Gene Writer has one or more of the functional characteristics described below after modification by the Gene Writer.
[0339] Gene Writer Polypeptide DNA-binding domain In some embodiments, the DNA-binding domain can bind to a target sequence (e.g., a dsDNA target sequence) with higher affinity than a standard DNA-binding domain. In some embodiments, the standard DNA-binding domain is a DNA-binding domain derived from the Cre recombinase of bacteriophage P1. In some embodiments, the DNA-binding domain can bind to a target sequence (e.g., a dsDNA target sequence) with an affinity of 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM).
[0340] In some embodiments, the affinity of a DNA binding domain for its target sequence (e.g., a dsDNA target sequence) is measured in vitro, e.g., by thermophoresis, as described, e.g., in Asmari et al. Methods 146:107-119 (2018), which is incorporated herein by reference in its entirety.
[0341] In some embodiments, the DNA-binding domain can bind to a target sequence (e.g., a dsDNA target sequence) with an affinity of, for example, 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM), e.g., in the presence of about a 100-fold molar excess of scrambled-sequence competitor dsDNA.
[0342] In some embodiments, the DNA-binding domain is found to bind to its target sequence (e.g., a dsDNA target sequence) at a higher frequency than any other sequence in the genome of the target cell, e.g., a human target cell, as measured by, e.g., ChIP-seq (e.g., in HEK293T cells), e.g., as described in He and Pu (2010) Curr. Protoc Mol Biol Chapter 21, which is incorporated herein by reference in its entirety. In some embodiments, the DNA-binding domain is found to bind to its target sequence (e.g., a dsDNA target sequence) at a frequency at least 5-fold or 10-fold higher than any other sequence in the target cell, as measured by, e.g., ChIP-seq (e.g., in HEK293T cells), e.g., as described in He and Pu (2010), supra.
[0343] Template-binding domain In some embodiments, the template-binding domain can bind to the template DNA with higher affinity than a standard DNA-binding domain. In some embodiments, the standard DNA-binding domain is a DNA-binding domain derived from the Cre recombinase of bacteriophage P1. In some embodiments, the template-binding domain can bind to the template DNA with an affinity of 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM). In some embodiments, the affinity of the DNA-binding domain for its template DNA is measured in vitro, e.g., by thermophoresis, as described, e.g., in Asmari et al. Methods 146:107-119 (2018), which is incorporated herein by reference in its entirety. In some embodiments, the affinity of the DNA-binding domain for its template DNA is measured in a cell (e.g., by FRET or ChIP-Seq).
[0344] In some embodiments, the DNA-binding domain associates with template DNA in vitro in the presence of 10 nM competitor DNA, with at least 50% template DNA bound, as described, for example, in Yant et al. Mol Cell Biol 24(20):9239-9247 (2004), the entire contents of which are incorporated herein by reference. In some embodiments, the DNA-binding domain associates with template DNA in cells (e.g., in HEK293T cells) at a frequency at least about 5-fold or 10-fold higher than with scrambled DNA. In some embodiments, the frequency of association of the DNA-binding domain with template DNA or scrambled DNA is measured by ChIP-seq, as described in He and Pu (2010), supra.
[0345] target site In some embodiments, after gene writing, the target site surrounding the integration sequence contains a limited number of insertions or deletions, e.g., in about 50% or less than 10% of integration events, as determined, e.g., by long-read amplicon sequencing of the target site, as described, e.g., in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903 (incorporated herein by reference in its entirety). In some embodiments, the target site does not exhibit multiple insertion events, e.g., head-to-tail or head-to-head duplications, as determined, e.g., by long-read amplicon sequencing of the target site, as described, e.g., in Karst et al. (2020), supra. In some embodiments, the target site contains an integration sequence corresponding to the template DNA. In some embodiments, the target site contains a fully integrated template molecule. In some embodiments, the target site contains components of vector DNA, e.g., AAV ITRs. In some embodiments, when the template DNA is initially excised from the viral vector, e.g., by a first recombination event prior to integration, the target site does not contain non-template DNA, e.g., an insert originating from endogenous or vector DNA, e.g., AAV ITRs, in more than about 1% or 10% of events, as determined, e.g., by long-read amplicon sequencing of the target site, e.g., as described in Karst et al. (2020), supra. In some embodiments, the target site contains an integration sequence corresponding to the template DNA.
[0346] In some embodiments, the Gene Writers described herein are capable of site-specific editing of target DNA, e.g., inserting template DNA into target DNA. In some embodiments, the site-specific Gene Writers can generate edits, e.g., insertions, that occur at the target site at a frequency greater than any other site in the genome. In some embodiments, the site-specific Gene Writers can generate edits, e.g., insertions, at the target site at a frequency that is at least 2, 3, 4, 5, 10, 50, 100, or 1000 times greater than the frequency of all other sites in the human genome. In some embodiments, the location of the integration site is determined by unidirectional sequencing. Incorporation of unique molecular identifiers (UMIs) into adapters or primers used for library preparation allows for quantification of individual insertion events, which can be compared between on-target insertions and all other insertions to determine preference for a defined target site.
[0347] In some embodiments, gene writing system is used to edit the target DNA sequence that exists at a single position in the human genome.In some embodiments, gene writing is used to edit the target DNA sequence that exists at a single position in the human genome on a single homologous chromosome, for example, this is haplotype specific.In some embodiments, gene writing system is used to edit the target DNA sequence that exists at a single position in the human genome on two homologous chromosomes.In some embodiments, gene writing is used to edit the target DNA sequence that exists at multiple positions in the genome, for example, at least 2, 3, 4, 5, 10, 20, 50, 100, 200, 500, 1000, 5000, 10000, 100000, 200000, 500000, 1000000 (for example, Alu element) positions in the genome.
[0348] In some embodiments, the Gene Writer system can edit a genome without introducing undesired mutations. In some embodiments, the Gene Writer system can edit a genome by inserting a template, e.g., template DNA, into the genome. In some embodiments, the resulting modifications in the genome include minimal mutations relative to the template DNA sequence. In some embodiments, the average error rate of genome insertion compared to the template DNA is 10 per nucleotide. -4 , 10 -5 , or 10 -6 In some embodiments, the number of mutations introduced into the template DNA into the target cell is, on average, 1, 2, 3, 4, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides per genome. In some embodiments, the error rate of introduction into the target genome is determined by comparing the template DNA sequence with long-read amplicon sequencing across known target sites, as described in Karst et al. (2020), supra. In some embodiments, the errors counted by this method include nucleotide substitutions relative to the template sequence. In some embodiments, the errors counted by this method include nucleotide deletions relative to the template sequence. In some embodiments, the errors counted by this method include nucleotide insertions relative to the template sequence. In some embodiments, the errors counted by this method include any one or more combinations of nucleotide substitutions, deletions, or insertions relative to the template sequence.
[0349] The efficiency of the integration event can be used as a measure of editing of the target site or target cell by the Gene Writer system. In some embodiments, the Gene Writer systems described herein can integrate a heterologous sequence of interest into a target site or a fraction of target cells. In some embodiments, the Gene Writer system can edit at least 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% of the target locus, as measured by detection of edits when analyzed by, for example, long-read amplicon sequencing after amplification of the entire target, e.g., as described in Karst et al. (2020). In some embodiments, the Gene Writer system is capable of editing cells with an average copy number of at least 0.1, e.g., at least 0.1, 0.5, 1, 2, 3, 4, 5, 10, or 100 copies per genome, normalized to a standard gene, e.g., RPP30, across a population of cells, as determined by ddPCR using a transgene-specific primer-probe set, e.g., according to the methods described in Lin et al. Hum Gene Ther Methods 27(5):197-208 (2016).
[0350] In some embodiments, the copy number per cell is analyzed by single-cell ddPCR (sc-ddPCR), e.g., according to the method of Igarashi et al. Mol Ther Methods Clin Dev 6:8-16 (2017), which is incorporated herein by reference in its entirety. In some embodiments, at least 1%, e.g., at least 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% of the target cells are positive for integration as assessed by sc-ddPCR using a transgene-specific primer-probe set. In some embodiments, the average copy number is at least 0.1, e.g., at least 0.1, 0.5, 1, 2, 3, 4, 5, 10, or 100 copies per cell, as measured by sc-ddPCR using a transgene-specific primer-probe set.
[0351] Additional Gene Writer features In some embodiments, the Gene Writer system can achieve complete writing without the need for endogenous host factors. In some embodiments, the system can achieve complete writing without the need for DNA repair. In some embodiments, the system can achieve complete writing without inducing a DNA damage response.
[0352] In some embodiments, the system does not require DNA repair via the NHEJ pathway, the homologous recombination repair pathway, the base excision repair pathway, or any combination thereof. Participation by a DNA repair pathway can be assayed, for example, by applying a DNA repair pathway inhibitor or a DNA repair pathway-deficient cell line. For example, when applying a DNA repair pathway inhibitor, a PrestoBlue cell viability assay can be first performed to determine the toxicity of the inhibitor and whether any normalization should be applied. SCR7 is an inhibitor of NHEJ, which can be applied in serial dilutions during Gene Writer™ delivery. PARP protein is a nuclear enzyme that binds to both single-strand and double-strand breaks as a homodimer. Therefore, the inhibitor can be used to test related DNA repair pathways, including the homologous recombination repair pathway and the base excision repair pathway. The experimental procedure is the same as that for SCR7. The effect of NER on Gene Writer™ can be tested using a cell line lacking a core protein in the nucleotide excision repair (NER) pathway. After the Gene Writer™ system is delivered to cells, ddPCR can be used to evaluate the insertion of heterologous target sequence in relation to the inhibition of DNA repair pathway.Sequencing analysis can also be carried out to evaluate whether DNA repair pathway plays a specific role.In some embodiments, the Gene Writing™ in genome is not reduced by the knockdown of DNA repair pathway described herein.In some embodiments, the Gene Writing™ in genome is not reduced by more than 50% by the knockdown of DNA repair pathway.
[0353] Evolutionary variant of Gene Writer In some embodiments, the present invention provides evolved variants of Gene Writers. Evolved variants can be produced, in some embodiments, by mutagenizing a standard Gene Writer or one of the fragments or domains contained therein. In some embodiments, one or more of the domains (e.g., catalytic or DNA-binding domains (e.g., target-binding domains or template-binding domains), such as sequence-guided DNA-binding elements) are evolved. One or more of these evolved variant domains can, in some embodiments, evolve alone or together with other domains. One or more evolved variant domains can, in some embodiments, be combined with a non-evolved cognate component or an evolved variant of a cognate component (e.g., one that may have evolved in a parallel or sequential manner).
[0354] In some embodiments, the process of mutagenizing a standard Gene Writer, or a fragment or domain thereof, comprises mutagenizing the standard Gene Writer, or a fragment or domain thereof. In embodiments, the mutagenesis comprises a progressive evolution method (e.g., PACE) or a non-progressive evolution method (e.g., PANCE), e.g., as described herein. In some embodiments, the evolved Gene Writer, or a fragment or domain thereof (e.g., a DNA-binding domain, e.g., a target-binding domain or a template-binding domain), comprises one or more amino acid mutations introduced into its amino acid sequence compared to the standard Gene Writer, or a fragment or domain thereof. In some embodiments, the amino acid sequence mutation may comprise one or more mutated residues (e.g., conservative substitutions, non-conservative substitutions, or a combination thereof) within the amino acid sequence of the standard Gene Writer, for example, as a result of a change in the nucleotide sequence encoding the gene writer resulting in a change in a codon at any particular position within the coding sequence, a deletion of one or more amino acids (e.g., a truncated protein), an insertion of one or more amino acids, or any combination thereof. An evolved variant Gene Writer may contain variants in one or more components or domains of the Gene Writer (e.g., variants introduced into the catalytic domain, the DNA-binding domain, or a combination thereof).
[0355] In some aspects, the present invention provides Gene Writers, systems, kits, and methods that use or include evolved variants of Gene Writers, for example, evolved variants of Gene Writers, or Gene Writers produced or producible by PACE or PANCE. In embodiments, the non-evolved standard Gene Writer is a Gene Writer disclosed herein.
[0356] The term "phage-assisted incremental evolution (PACE)" as used herein generally refers to incremental evolution using phage as a viral vector. Examples of PACE technology include, for example, the following: International PCT Application No. PCT / US2009 / 056194, filed September 8, 2009, and published March 11, 2010 as WO 2010 / 028347; International PCT Application No. PCT / US2011 / 066747, filed December 22, 2011, and published June 28, 2012 as WO 2012 / 088381; U.S. Patent No. 9,023,594, issued May 5, 2015; U.S. Patent No. 9,771,574, issued September 26, 2017; U.S. Patent No. 9,771,574, issued July 19, 2016; No. 9,394,537, filed January 20, 2015, published September 11, 2015 as WO 2015 / 134121; U.S. Pat. No. 10,179,911, issued January 15, 2019; and International PCT Application PCT / US 2016 / 027795, filed April 15, 2016, published October 20, 2016 as WO 2016 / 168631, each of which is incorporated herein by reference in its entirety.
[0357] The term "phage-assisted non-incremental evolution (PANCE)" as used herein generally refers to non-incremental evolution using phage as a viral vector. Examples of PANCE technology are described, for example, in Suzuki T. et al., Crystal structures reveal an elusive functional domain of pyrrolysyl-tRNA synthetase, Nat Chem Biol. 13(12):1261-1266 (2017), the entire contents of which are incorporated herein by reference. Briefly, PANCE is a technique for rapid in vivo directed evolution using serial flask transfer of evolutionary selection phage (SP) containing the gene of interest to be evolved into fresh whole host cells (e.g., E. coli cells). While the gene contained in the SP is progressively evolved, the gene inside the host cell can be held constant. Following phage propagation, an aliquot of the infected cells can be used to transfect the next flask containing host E. coli. This process can be repeated and / or continued until the desired phenotype has evolved, for example, for as many transfers as desired.
[0358] Methods for applying PACE and PANCE to Gene Writers will be readily understood by those skilled in the art by reference, inter alia, to the above-mentioned references. Further exemplary methods for directing the gradual evolution of genome-modifying proteins or systems, e.g., in a population of host cells, using, for example, phage particles, can be applied to generate evolved variants of Gene Writers, or fragments or subdomains thereof. Non-limiting examples of such methods are described, for example, in the following: International PCT Application PCT / US Application No. 2009 / 056194, filed September 8, 2009, and published March 11, 2010, as WO 2010 / 028347; International PCT Application PCT / US Application No. 2012 / 088381, filed December 22, 2011, and published June 28, 2012, as WO 2012 / 088381; T application, PCT / US Patent Application No. 2011 / 066747; U.S. Patent No. 9,023,594 issued May 5, 2015; U.S. Patent No. 9,771,574 issued September 26, 2017; U.S. Patent No. 9,394,537 issued July 19, 2016; WO 2015 / 134121 filed January 20, 2015, September 11, 2015 No. 10,179,911, issued January 15, 2019; International Application PCT / US Application No. 2019 / 37216, filed June 14, 2019; International Publication No. WO 2019 / 023680, published January 31, 2019; International PCT Application PCT / US Application No. 2016 / 027795, filed April 15, 2016, published October 20, 2016 as WO 2016 / 168631; and International Application PCT / US Application No. 2019 / 47996, filed August 23, 2019, each of which is incorporated herein by reference in its entirety.
[0359] In some non-limiting exemplary embodiments, a method for evolving an evolved variant Gene Writer, or fragment or domain thereof, comprises: (a) contacting a population of host cells with a population of viral vectors comprising a gene of interest (the starting Gene Writer, or fragment or domain thereof), where (1) the host cells are suitable for infection with the viral vector; (2) the host cells express viral genes necessary for the production of viral particles; (3) the expression of at least one viral gene necessary for the production of infectious viral particles is dependent on the function of the gene of interest; and / or (4) the viral vector allows for the expression of a protein in the host cells and can be replicated by the host cells and packaged into viral particles. In some embodiments, the method includes (b) contacting the host cells with a mutagen using host cells containing mutations that enhance the mutation rate (e.g., by delivering a mutant plasmid, or any genomic modification—e.g., a damaged DNA proofreading polymerase, an SOS gene, e.g., UmuC, UmuD′, and / or RecA (which mutations can be under the control of an inducible promoter when bound to the plasmid), or a combination thereof). In some embodiments, the method includes (c) incubating the population of host cells under conditions that allow viral replication and viral particle production, whereupon host cells are removed from the host cell population and fresh, uninfected host cells are introduced into the host cell population, thus replenishing the host cell population and forming a stream of host cells. In some embodiments, the cells are incubated under conditions that allow the gene of interest to acquire a mutation. In some embodiments, the method further includes (d) isolating from the population of host cells a mutated version of the viral vector encoding an evolved gene product (e.g., an evolved mutant Gene Writer, or a fragment or domain thereof).
[0360] Those skilled in the art will appreciate various features that can be used within the above framework. For example, in some embodiments, the viral vector or phage is a filamentous phage, e.g., an M13 phage, e.g., an M13 selection phage. In certain embodiments, the gene required for the production of infectious viral particles is M13 gene III (gIII). In some embodiments, the phage may lack functional gIII but instead contain gI, gII, gIV, gV, gVI, gVII, gVIII, gIX, and gX. In some embodiments, the generation of infectious VSV particles includes the envelope protein VSV-G. In various embodiments, different retroviral vectors, e.g., murine leukemia virus vectors or lentiviral vectors, can be used. In some embodiments, retroviral vectors can be efficiently packaged, e.g., using VSV-G envelope proteins as a substitute for the virus's native envelope proteins.
[0361] In some embodiments, the host cells are incubated for a suitable number of viral life cycles, e.g., at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1250, at least 1500, at least 1750, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 7500, at least 10000, or more consecutive viral life cycles, where an illustrative, non-limiting example for M13 phage is 10-20 minutes per viral life cycle. Similarly, conditions can be adjusted to control the residence time of the host cells in the population of host cells (e.g., about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 70, about 80, about 90, about 100, about 120, about 150, or about 180 minutes). The host cell population can be adjusted to control the density of host cells, or in some embodiments, the host cell density in the inflow, e.g., 10 3 cells / ml, approximately 10 4 cells / ml, approximately 10 5 cells / ml, approximately 5-10 5 cells / ml, approximately 10 6 cells / ml, approximately 5-10 6 cells / ml, approximately 10 7 cells / ml, approximately 5-10 7 cells / ml, approximately 10 8 cells / ml, approximately 5-10 8 cells / ml, approximately 10 9 cells / ml, approximately 5-10 9 cells / ml, approximately 10 10 cells / ml, or approximately 5-10 10 Cells / ml can be partly controlled.
[0362] nucleic acid Promoter In some embodiments, one or more promoters or enhancers are operably linked to a nucleic acid encoding a Gene Writer polypeptide, or to a template nucleic acid that controls expression of, for example, a heterologous sequence of interest. In certain embodiments, the one or more promoters or enhancers comprise cell type or tissue-specific elements. In some embodiments, the promoter or enhancer is the same as or is derived from the promoter or enhancer that naturally controls expression of the heterologous sequence of interest. For example, an ornithine transcarbomylase promoter and enhancer can be used to control expression of the ornithine transcarbomylase gene in a system or method provided by the present invention for correcting an ornithine transcarbomylase deficiency. In some embodiments, the promoter is a promoter in Table 33 or a functional fragment or variant thereof.
[0363] Exemplary commercially available tissue-specific promoters can be found, for example, at a URL (e.g., https: / / www.invivogen.com / tissue-specific-promoters). In some embodiments, the promoter is a native promoter or a minimal promoter (e.g., one composed of a single fragment from the 5' region of a given gene). In some embodiments, the native promoter comprises a core promoter and its natural 5' UTR. In some embodiments, the 5' UTR comprises an intron. In other embodiments, these comprise composite promoters created by combining promoters of different origins or assembling a distal enhancer with a minimal promoter of the same origin. In some embodiments, the tissue-specific expression control sequence comprises one or more of the sequences in Table 2 or Table 3 of WO2020014209, which is incorporated by reference in its entirety.
[0364] Exemplary cell or tissue specific promoters are listed in the table below, and exemplary nucleic acids encoding them are known in the art and readily available using a variety of resources, for example, the NCBI database, including RefSeq., as well as the Eukaryotic Promoter Database (http: / / epd.epfl.ch / / index.php).
[0365] [Table 5]
[0366] [Table 6-1]
[0367] [Table 6-2]
[0368] [Table 6-3]
[0369] [Table 6-4]
[0370] Depending on the host / vector system utilized, any of a number of suitable transcriptional and translational control elements may be used in the expression vector, including constitutive and inducible promoters, transcriptional enhancer elements, transcriptional terminators, etc. (See, e.g., Bitter et al. (1987) Methods in Enzymology, 153:516-544; incorporated herein by reference in its entirety).
[0371] In some embodiments, the nucleic acid encoding the Gene Writer or the template nucleic acid is operably linked to a control element, e.g., a transcriptional control element such as a promoter. The transcriptional control element may, in some embodiments, be functional in either eukaryotic cells, e.g., mammalian cells; or prokaryotic cells (e.g., bacterial or archaeal cells). In some embodiments, the nucleotide sequence encoding the polypeptide is operably linked to multiple control elements, e.g., that allow expression of the nucleotide sequence encoding the polypeptide in both prokaryotic and eukaryotic cells.
[0372] For purposes of illustration, examples of spatially restricted promoters include, but are not limited to, neuron-specific promoters, adipocyte-specific promoters, cardiomyocyte-specific promoters, smooth muscle-specific promoters, photoreceptor-specific promoters, and the like. Neuron-specific spatially restricted promoters include, but are not limited to, the neuron-specific enolase (NSE) promoter (see, e.g., EMBL HSENO2, X51956); aromatic amino acid decarboxylase (AADC) promoter, neurofilament promoter (see, e.g., GenBank HUMNFL, L04147); synapsin promoter (see, e.g., GenBank HUMSYNIB, M55301); thy-1 promoter (see, e.g., Chen et al. (1987) Cell 51:7-19; and Llewellyn, et al. (2010) Nat. Med. 16(10):1161-1166); serotonin receptor promoter (see, e.g., GenBank S62283); tyrosine hydroxylase promoter (TH) (see, e.g., Oh et al. (2009) Gene Ther 16:437; Sasaoka et al. (1992) Mol. Brain Res. 16:274; Boundy et al. (1998) J. Neurosci. 18:9989; and Kaneda et al. (1991) Neuron 6:583-594); GnRH promoter (see, e.g., Radovick et al. (1991) Proc. Natl. Acad. Sci. USA 88:3402-3406); L7 promoter (see, e.g., Oberdick et al. (1990) Science 248:223-226); DNMT promoter (see, e.g., Bartge et al. (1988) Proc. Natl. Acad. Sci. USA 85:3648-3652); enkephalin promoter (see, e.g., Comb et al. (1988) EMBO J. 17:3793-3805); myelin basic protein (MBP) promoter; Ca2+ calmodulin-dependent protein kinase II-α (CamKIIα) promoter (see, e.g., Mayford et al. (1996) Proc. Natl. Acad. Sci.USA 93:13250; and Casanova et al. (2001) Genesis 31:37); CMV enhancer / platelet-derived growth factor-β promoter (see, e.g., Liu et al. (2004) Gene Therapy 11:52-60); and the like.
[0373] Adipocyte-specific spatially restricted promoters include, but are not limited to, the aP2 gene promoter / enhancer, e.g., the region from −5.4 kb to +21 bp of the human aP2 gene (see, e.g., Tozzo et al. (1997) Endocrinol. 138:1604; Ross et al. (1990) Proc. Natl. Acad. Sci. USA 87:9590; and Pavjani et al. (2005) Nat. Med. 11:797); glucose transporter-4 (GLUT4) promoter (see, e.g., Knight et al. (2003) Proc. Natl. Acad. Sci. USA 100:14725); fatty acid translocase (FAT / CD36) promoter (see, e.g., Kuriki et al. (2005) Nat. Med. 11:797); al. (2002) Biol. Pharm. Bull. 25:1476; and Sato et al. (2002) J. Biol. Chem. 277:15703); the stearoyl-CoA desaturase-1 (SCD1) promoter (Tabor et al. (1999) J. Biol. Chem. 274:20603); the leptin promoter (see, e.g., Mason et al. (1998) Endocrinol. 139:1013; and Chen et al. (1999) Biochem. Biophys. Res. Comm. 262:187); the azidonectin promoter (see, e.g., Kita et al. (2005) Biochem. Biophys. Res. Comm. 331:484; and Chakrabarti (2010) Endocrinol. 151:2408); adipsin promoter (see, e.g., Platt et al. (1989) Proc. Natl. Acad. Sci. USA 86:7490); resistin promoter (see, e.g., Seo et al. (2003) Molec. Endocrinol. 17:1522); and the like.
[0374] Cardiomyocyte-specific spatially restricted promoters include, but are not limited to, regulatory sequences derived from the following genes: myosin light chain-2, alpha-myosin heavy chain, AE3, cardiac troponin C, cardiac actin, and the like. Franz et al.(1997)Cardiovasc.Res.35:560-566;Robbins et al.(1995)Ann.NYAcad.Sci.752:492-505;Linn et al.(1995)Circ.Res.76:584-591;Parmacek et al. al. (1994) Mol. Cell. Biol. 14:1870-1885; Hunter et al. (1993) Hypertension 22:608-617; and Sartorelli et al. (1992) Proc. Natl. Acad. Sci. USA 89:4047-4051.
[0375] Smooth muscle cell-specific spatially restricted promoters include, but are not limited to, the SM22α promoter (see, e.g., Akyuerek et al. (2000) Mol. Med. 6:983; and U.S. Pat. No. 7,169,874); the smoothelin promoter (see, e.g., WO 2001 / 018048); the α-smooth muscle actin promoter; and the like. For example, a 0.4 kb region of the SM22α promoter (within which two CArG elements are located) has been shown to mediate vascular smooth muscle cell-specific expression (see, e.g., Kim, et al. (1997) Mol. Cell. Biol. 17, 2266-2278; Li, et al. (1996) J. Cell Biol. 132 849-859; and Moessler, et al. (1996) Development 122, 2415-2425).
[0376] Photoreceptor-specific spatially restricted promoters include, but are not limited to, the rhodopsin kinase promoter (Young et al. (2003) Ophthalmol. Vis. Sci. 44:4076); beta phosphodiesterase gene promoter (Nicoud et al. (2007) J. Gene Med. 9:1015); retinitis pigmentosa gene promoter (Nicoud et al. (2007) supra); interphotoreceptor retinoid-binding protein (IRBP) gene enhancer (Nicoud et al. (2007) supra); IRBP gene promoter (Yokoyama et al. (1992) Exp Eye Res. 55:225); and the like.
[0377] Non-limiting exemplary cell-specific promoters Cell-specific promoters known in the art can be used to direct the expression of, for example, the Gene Writer proteins described herein. Non-limiting exemplary mammalian cell-specific promoters have been characterized and used in mice that express Cre recombinase in a cell-specific manner. Some non-limiting exemplary mammalian cell-specific promoters are listed in Table 1 of U.S. Pat. No. 9,845,481 (incorporated herein by reference).
[0378] In some embodiments, the cell-specific promoter is a promoter active in plants.Many exemplary cell-specific promoters are known in the art.See, for example, U.S. Patent No. 5,097,025; U.S. Patent No. 5,783,393; U.S. Patent No. 5,880,330; U.S. Patent No. 5,981,727; U.S. Patent No. 7,557,264; U.S. Patent No. 6,291,666; U.S. Patent No. 7,132,526; and U.S. Patent No. 7,323,622; and U.S. Patent Application Publication No. 2010 / 0269226; U.S. Patent Application Publication No. 2007 / 0180580; U.S. Patent Application Publication No. 2005 / 0034192; and U.S. Patent Application Publication No. 2005 / 0086712 (which are incorporated herein by reference in their entirety for all purposes).
[0379] In some embodiments, the vectors described herein comprise an expression cassette. The term "expression cassette," as used herein, refers to a nucleic acid construct comprising sufficient nucleic acid elements for expression of a nucleic acid molecule of the present invention. Typically, an expression cassette comprises a nucleic acid molecule of the present invention operably linked to a promoter sequence. The term "operably linked" refers to the association of two or more nucleic acid fragments on a single nucleic acid fragment such that the function of one is affected by the other. For example, a promoter is operably linked to a coding sequence if it is capable of affecting the expression of the coding sequence (e.g., the coding sequence is under the transcriptional control of the promoter). A coding sequence can be operably linked to a regulatory sequence in a sense or antisense orientation. In certain embodiments, the promoter is a heterologous promoter. The term "heterologous promoter," as used herein, refers to a promoter that is not known to be operably linked to a given coding sequence in nature. In certain embodiments, an expression cassette may include additional elements, such as introns, enhancers, polyadenylation sites, woodchuck response elements (WREs), and / or other elements known to affect the expression level of a coding sequence. A "promoter" typically controls the expression of a coding sequence or functional RNA. In certain embodiments, a promoter sequence includes proximal and more distal upstream elements and may further include enhancer elements. An "enhancer" is typically capable of stimulating promoter activity and may be a native element of the promoter or a heterologous element inserted to enhance the level or tissue specificity of the promoter. In certain embodiments, a promoter is derived entirely from a native gene. In certain embodiments, a promoter is composed of various elements from different naturally occurring promoters. In certain embodiments, a promoter comprises a synthetic nucleotide sequence.Those skilled in the art will appreciate that different promoters may direct the expression of a gene in different tissues or cell types, or at different stages of development, or in response to different environmental conditions or the presence or absence of drugs or transcriptional cofactors. Ubiquitous, cell type-specific, tissue-specific, developmental stage-specific, and conditional promoters, such as drug-responsive promoters (e.g., tetracycline-responsive promoters), are well known to those skilled in the art. Examples of promoters include, but are not limited to, the phosphoglycerate kinase (PKG) promoter, CAG (a composite of the CMV enhancer, chicken β-actin promoter (CBA), and rabbit β-globin intron), NSE (neuron-specific enolase), synapsin, or NeuN promoter, SV40 early promoter, mouse mammary tumor virus LTR promoter; adenovirus major late promoter (Ad MLP); herpes simplex virus (HSV) promoter, cytomegalovirus (CMV) promoter, e.g., the CMV immediate early promoter region (CMVIE), SFFV promoter, Rous sarcoma virus (RSV) promoter, synthetic promoters, hybridization promoters, etc. Other promoters may be derived from humans or other species, including mice.Common promoters include, for example, the human cytomegalovirus (CMV) immediate-early gene promoter, the SV40 early promoter, the Rous sarcoma virus long terminal repeat, [β]-actin, the rat insulin promoter, the phosphoglycerate kinase promoter, the human α-1 antitrypsin (hAAT) promoter, the transthyretin promoter, the TBG promoter and other liver-specific promoters, the desmin promoter and similar muscle-specific promoters, the EF1-α promoter, the CAG promoter and other constitutive promoters, hybrid promoters with multiple tissue specificities, and neuron-specific promoters such as the synapsin and glyceraldehyde-3-phosphate dehydrogenase promoters, all of which are well known to those skilled in the art and readily available, and can be used to achieve high-level expression of a coding sequence of interest. In addition, sequences derived from non-viral genes, such as the mouse metallothionein gene, may also be useful in the present invention. Such promoter sequences are commercially available, for example, from Stratagene (San Diego, CA). Further exemplary promoter sequences are described, for example, in WO2018213786A1, which is incorporated by reference in its entirety.
[0380] In some embodiments, the apolipoprotein E enhancer (ApoE) or a functional fragment thereof is used, for example, to drive expression in the liver. In some embodiments, two copies of the ApoE enhancer or a functional fragment thereof are used. In some embodiments, the ApoE enhancer or a functional fragment thereof is used in combination with a promoter, for example, the human alpha-1 antitrypsin (hAAT) promoter.
[0381] In some embodiments, the regulatory sequence confers tissue-specific gene expression. In some cases, the tissue-specific regulatory sequence binds to a tissue-specific transcription factor that induces transcription in a tissue-specific manner. Various tissue-specific regulatory sequences (e.g., promoters, enhancers, etc.) are known in the art. Exemplary tissue-specific regulatory sequences include, but are not limited to, the following tissue-specific promoters: a liver-specific thyroxine-binding globulin (TBG) promoter, an insulin promoter, a glucagon promoter, a somatostatin promoter, a pancreatic polypeptide (PPY) promoter, a synapsin-1 (Syn) promoter, a creatine kinase (MCK) promoter, a mammalian desmin (DES) promoter, an α-myosin heavy chain (a-MHC) promoter, or a cardiac troponin T (cTnT) promoter. Other exemplary promoters include the β-actin promoter, the hepatitis B virus core promoter, Sandig et al., Gene Ther., 3:1002-9 (1996); the α-fetoprotein (AFP) promoter, Arbuthnot et al., Hum. Gene Ther., 7:1503-14 (1996)), the bone osteocalcin promoter (Stein et al., Mol. Biol. Rep., 24:185-96 (1997)); the bone sialoprotein promoter (Chen et al., J. Bone Miner. Res., 11:654-64 (1996)), the CD2 promoter (Hansal et al., J. Bone Miner. Res., 11:654-64 (1996)), and the α-fetoprotein (AFP) promoter (Arbuthnot et al., Hum. Gene Ther., 7:1503-14 (1996)). et al., J. Immunol., 161:1063-8 (1998); immunoglobulin heavy chain promoter; T cell receptor α-chain promoter, neuronal promoters such as the neuron-specific enolase (NSE) promoter (Andersen et al., Cell. Mol. Neurobiol., 13:503-15 (1993)), neurofilament light chain gene promoter (Piccioli et al., Proc. Natl. Acad. Sci. USA, 88:5611-5 (1991)), and neuron-specific vgf gene promoter (Piccioli et al., Neuron, 15:373-84 (1995)).Other exemplary promoter sequences are described, for example, in U.S. Patent No. 10,300,146 (incorporated herein by reference in its entirety). In some embodiments, tissue-specific regulatory elements, e.g., tissue-specific promoters, are selected from those known to be operably linked to genes that are highly expressed in a given tissue, for example, as determined by RNA-seq or protein expression data, or a combination thereof. Methods for analyzing tissue specificity by expression are taught in Fagerberg et al. Mol Cell Proteomics 13(2):397-406 (2014) (incorporated herein by reference in its entirety).
[0382] In some embodiments, the vectors described herein are multicistronic expression constructs. Examples of multicistronic expression constructs include constructs having a first expression cassette (e.g., comprising a first promoter and a first coding nucleic acid sequence) and a second expression cassette (e.g., comprising a second promoter and a second coding nucleic acid sequence). Such multicistronic expression constructs may be particularly useful in some cases for delivery of non-translated gene products, such as hairpin RNAs, along with polypeptides, e.g., gene writers and gene writer templates. In some embodiments, multicistronic expression constructs may exhibit reduced expression levels of one or more of the included transgenes due, for example, to promoter interference or the nearby presence of incompatible nucleic acid elements. When a multicistronic expression construct is part of a viral vector, the presence of self-complementary nucleic acid sequences may, in some cases, interfere with the formation of structures necessary for viral propagation or packaging.
[0383] In some embodiments, the sequence encodes an RNA having a hairpin. In some embodiments, the hairpin RNA is a guide RNA, a template RNA, an shRNA, or a microRNA. In some embodiments, the first promoter is an RNA polymerase I promoter. In some embodiments, the first promoter is an RNA polymerase II promoter. In some embodiments, the second promoter is an RNA polymerase III promoter. In some embodiments, the second promoter is a U6 or H1 promoter. In some embodiments, the nucleic acid construct comprises AAV construct B1 or B2.
[0384] Without intending to be bound by any particular theory, it is believed that multicistronic expression constructs do not achieve optimal expression levels compared to expression systems containing a single cistron. One of the suggested causes for the reduced expression levels achieved using multicistronic expression constructs containing two or more promoter elements is the phenomenon of promoter interference (see, e.g., Curtin JA, Dane AP, Swanson A, Alexander IE, Ginn S L. Bidirectional promoter interference between two widely used internal heterologous promoters in a late-generation lentiviral construct. Gene Ther. 2008 March;15(5):384-90; and Martin-Duque P, Jezzard S, Kaftansis L, Vassaux G. Direct comparison of the insulating properties of two genetic elements in an adenoviral vector containing two different expression cassettes. Hum Gene Ther. 2004 October;15(10):995-1002; both references are incorporated herein by reference for their disclosure of the phenomenon of promoter interference). In some embodiments, the problem of promoter interference can be overcome by, for example, generating a multicistronic expression construct containing a single promoter driving transcription of multiple coding nucleic acid sequences separated by internal ribosome entry sites, or by separating cistrons containing their own promoters with transcriptional insulator elements. In some embodiments, polycistronic expression driven by a single promoter can result in uneven expression levels of the cistrons. In some embodiments, promoters cannot be efficiently separated, and separating elements may not be compatible with some gene transfer vectors, such as some retroviral vectors.
[0385] microRNA miRNAs and other small interfering nucleic acids generally regulate gene expression by target RNA transcript cleavage / degradation or translational repression of target messenger RNA (mRNA). miRNAs, in some cases, can be naturally expressed, typically as final 19-25 untranslated RNA products. miRNAs generally exert their activity through sequence-specific interactions with the 3' untranslated region (UTR) of target mRNAs. These endogenously expressed miRNAs can form hairpin precursors, which are subsequently processed into miRNA duplexes and then mature single-stranded miRNA molecules. This mature miRNA generally guides a multiprotein complex, miRISC, which recognizes the target 3' UTR region of the target mRNA based on its complementarity to the mature miRNA. Useful transgene products can include, for example, miRNAs or miRNA-binding sites that regulate the expression of linked polypeptides. A non-limiting list of miRNA genes; the products of these genes and their homologs are useful as transgenes or as targets for small interfering nucleic acids (e.g., miRNA sponges, antisense oligonucleotides) in methods such as those listed in U.S. Pat. No. 10,300,146, 22:25-25:48, incorporated herein by reference. In some embodiments, one or more binding sites for one or more of the aforementioned miRNAs are incorporated into a transgene, e.g., a transgene delivered by an rAAV vector, to, for example, inhibit expression of the transgene in one or more tissues of an animal carrying the transgene. In some embodiments, the binding site may be selected to control transgene expression in a tissue-specific manner. For example, a binding site for liver-specific miR-122 can be incorporated into a transgene to inhibit expression of the transgene in the liver. Other exemplary miRNA sequences are described, for example, in U.S. Pat. No. 10,300,146, incorporated herein by reference in its entirety.
[0386] miR inhibitors or miRNA inhibitors are generally agents that block miRNA expression and / or processing. Examples of such agents include, but are not limited to, microRNA antagonists, microRNA-specific antisense, microRNA sponges, and microRNA oligonucleotides (double-stranded, hairpin, short oligonucleotides), which inhibit miRNA interaction with the Drosha complex. MicroRNA inhibitors, e.g., miRNA sponges, can be expressed in cells from transgenes (e.g., as described in Ebert, MS Nature Methods, Epub Aug. 12, 2007; incorporated herein by reference in its entirety). In some embodiments, microRNA sponges or other miR inhibitors are used in conjunction with AAV. MicroRNA sponges generally specifically inhibit miRNAs via complementary heptamer seed sequences. In some embodiments, a single sponge sequence can be used to silence an entire family of miRNAs. Other methods for silencing miRNA function (derepressing miRNA targets) in cells will be apparent to those skilled in the art.
[0387] In some embodiments, the miRNAs described herein comprise a sequence listed in Table 4 of WO2020014209, which is incorporated herein by reference. The list of exemplary miRNAs from WO2020014209 is also incorporated herein by reference.
[0388] 5'UTR and 3'UTR In certain embodiments, a nucleic acid comprising an open reading frame encoding a Gene Writer polypeptide (e.g., as described herein) comprises a 5' UTR and / or a 3' UTR. In some embodiments, the 5' UTR and 3' UTR for protein expression, e.g., an mRNA (or DNA encoding an RNA) for a Gene Writer polypeptide or a heterologous sequence of interest, comprise optimized expression sequences. In some embodiments, the 5'UTR comprises GGGAAAAUAAGAGAGAAAAGAAGAGUAAGAAGAAAUAUAAGAGCCACC (SEQ ID NO: 1867) and / or the 3'UTR comprises UGAUAAUAGGCUGGAGCCUCGGUGGCCAUGCUUCUUGCCCCUUGGGCCUCCCCCCAGCCCCUCCUCCCCUUCCUGCACCCGUACCCCCGUGGUCUUUGAAUAAAGUCUGA (SEQ ID NO: 1868), e.g., as described in Richner et al. Cell 168(6):P1114-1125 (2017) (the sequences of which are incorporated herein by reference).
[0389] In some embodiments, the open reading frame of the Gene Writer system, e.g., the ORF of the mRNA (or DNA encoding the mRNA) encoding the Gene Writer polypeptide or one or more ORFs of the mRNA (or DNA encoding the mRNA) of the heterologous sequence of interest, is flanked by 5' and / or 3' untranslated regions (UTRs) that enhance its expression. In some embodiments, the 5' UTR of the mRNA component (or transcript produced from the DNA component) of the system comprises the sequence 5'-GGGAAAUAAGAGAGAAAAGAAGAGUAAGAAGAAAUAUAAGAGCCACC-3' (SEQ ID NO: 1869). In some embodiments, the 3' UTR of the mRNA component (or transcript produced from the DNA component) of the system comprises the sequence 5'-UGAUAAUAGGCUGGAGCCUCGGUGGCCAUGCUUCUUGCCCCUUGGGCCUCCCCCCAGCCCCUCCUCCCCUUCCUGCACCCGUACCCCCGUGGUCUUUGAAUAAAGUCUGA-3' (SEQ ID NO: 1870). This combination of 5'UTR and 3'UTR has been shown by Richner et al. Cell 168(6):P1114-1125 (2017), the teachings and sequences of which are incorporated herein by reference, to result in the desired expression of the operably linked ORF. In some embodiments, the systems described herein include DNA encoding a transcript, where the DNA comprises the corresponding 5'UTR and 3'UTR in the sequences listed above, with T substituted for U. In some embodiments, the DNA vector used to produce the RNA component of the system includes a promoter upstream of the 5'UTR to initiate in vitro transcription, e.g., a T7, T3, or SP6 promoter. The 5'UTR begins with GGG, a preferred start for optimizing transcription using T7 RNA polymerase.To adjust transcription levels and modify transcription start site nucleotides to accommodate alternative 5'UTRs, Davidson et al. Pac Symp Biocomput 433-443 (2010) teaches T7 promoter variants that meet both of these properties and methods for finding them.
[0390] Viral vectors and their components In addition to being a source of the relevant enzymes or domains described herein, e.g., the recombinases and DNA-binding domains used in the present invention, e.g., the DNA-binding domains from Cre recombinase, λ integrase, or AAVRep proteins, viruses are useful sources of delivery vehicles for the systems described herein. Some enzymes may have multiple activities. In some embodiments, viruses used as a source of the Gene Writer delivery system or its components may be selected from the group described by Baltimore Bacteriol Rev 35(3):235-241 (1971).
[0391] In some embodiments, the virus is selected from a Group I virus, e.g., a DNA virus, which packages dsDNA into virions. In some embodiments, the Group I virus is selected from, e.g., an adenovirus, a herpesvirus, or a poxvirus.
[0392] In some embodiments, the virus is selected from a group II virus, e.g., a DNA virus that packages ssDNA into virions. In some embodiments, the group II virus is selected from a parvovirus. In some embodiments, the parvovirus is a dependoparvovirus, e.g., an adeno-associated virus (AAV).
[0393] In some embodiments, the virus is selected from a group III virus, e.g., an RNA virus, and packages dsRNA into virions. In some embodiments, the group III virus is selected from, e.g., a Reovirus. In some embodiments, one or both strands of the dsRNA contained in such virions are coding molecules that can serve directly as mRNA upon transduction of a host cell, e.g., can be directly translated into protein upon transduction of a host cell without the need for any intervening nucleic acid replication or polymerization step.
[0394] In some embodiments, the virus is selected from a Group IV virus, e.g., an RNA virus, and packages ssRNA(+) into virions. In some embodiments, the Group IV virus is selected from, e.g., a Coronavirus, a Picornavirus, or a Togavirus. In some embodiments, the ssRNA(+) contained in such virions is a coding molecule that can directly serve as mRNA upon transduction of a host cell, e.g., can be directly translated into protein upon transduction of a host cell without the need for any intervening nucleic acid replication or polymerization step.
[0395] In some embodiments, the virus is selected from group V viruses, for example, RNA viruses, and package ssRNA(-) into virion.In some embodiments, the virus is selected from group IV viruses, for example, Orthomyxoviruses, Rhabdoviruses.In some embodiments, the RNA virus with ssRNA(-) genome also carries the enzyme (for example, RNA-dependent RNA polymerase) inside virion, which is transduced into host cell together with viral genome, and can copy ssRNA(-) into ssRNA(+) that can be directly translated by host.
[0396] In some embodiments, the virus is selected from a group VI virus, e.g., a retrovirus, and packages ssRNA(+) into virions. In some embodiments, the group VI virus is selected from, e.g., a retrovirus. In some embodiments, the retrovirus is a lentivirus, e.g., HIV-1, HIV-2, SIV, or BIV. In some embodiments, the retrovirus is a spumavirus, e.g., a foamy virus, e.g., HFV, SFV, or BFV. In some embodiments, the ssRNA(+) contained in such virions is a coding molecule that can directly serve as mRNA upon transduction of a host cell, e.g., can be directly translated into protein upon transduction of a host cell without the need for any intervening nucleic acid replication or polymerization steps. In some embodiments, the ssRNA(+) is first reverse-transcribed and copied to generate a dsDNA genome intermediate from which mRNA can be transcribed in the host. In some embodiments, RNA viruses with ssRNA(+) genomes also carry within the virion an enzyme (e.g., RNA-dependent DNA polymerase) that is transduced into the host cell along with the viral genome and can copy the ssRNA(+) into dsDNA, which can be transcribed into mRNA and translated by the host.
[0397] In some embodiments, the virus is selected from a group VII virus, e.g., a retrovirus, and packages dsRNA into virions. In some embodiments, the group VII virus is selected from, e.g., a hepadnavirus. In some embodiments, one or both strands of the dsRNA contained in such virions are coding molecules that can directly serve as mRNA upon transduction of a host cell, e.g., can be directly translated into protein upon transduction of a host cell without the need for any intervening nucleic acid replication or polymerization steps. In some embodiments, one or both strands of the dsRNA contained in such virions are first reverse transcribed and copied to generate a dsDNA genome intermediate, from which mRNA can be transcribed by the host. In some embodiments, RNA viruses with dsRNA genomes also carry within the virion an enzyme (e.g., RNA-dependent DNA polymerase) that is transduced into host cells along with the viral genome and can copy the dsRNA into dsDNA, which can be transcribed into mRNA and translated by the host.
[0398] In some embodiments, the virions used to deliver nucleic acids in the present invention may also carry enzymes involved in the gene writing process. For example, the virions may contain a recombinase domain that is delivered to the host cell together with the nucleic acid. In some embodiments, the template nucleic acid may be associated with a gene writer polypeptide within the virion, so that both are simultaneously delivered to the target cell upon transduction of the nucleic acid from the viral particle. In some embodiments, the nucleic acid within the virion may comprise DNA, such as linear ssDNA, linear dsDNA, circular ssDNA, circular dsDNA, minicircle DNA, dbDNA, or ceDNA. In some embodiments, the nucleic acid within the virion may comprise RNA, such as linear ssRNA, linear dsRNA, circular ssRNA, or circular dsRNA. In some embodiments, the viral genome may be circularized upon transduction into the host cell; for example, a linear ssRNA molecule may form a circular ssRNA via covalent linkage, or a linear dsRNA molecule may form a circular dsRNA or one or more circular ssRNAs via covalent linkage. In some embodiments, the viral genome may replicate by rolling circle replication within the host cell. In some embodiments, the viral genome may comprise a single nucleic acid molecule, e.g., a non-segmented genome. In some embodiments, the viral genome may comprise two or more nucleic acid molecules, e.g., a segmented genome. In some embodiments, the nucleic acid within the virion may be associated with one or more proteins. In some embodiments, one or more proteins within the virion can be delivered to the host cell upon transduction. In some embodiments, a native virus may be modified for delivery of a nucleic acid by the addition of a virion packaging signal to the target nucleic acid, where the host cell is used to package the target nucleic acid containing the packaging signal.
[0399] In some embodiments, virions used as delivery vehicles may comprise commensal human viruses. In some embodiments, virions used as delivery vehicles may comprise anelloviruses, the use of which is described in WO2018232017A1, which is incorporated herein by reference in its entirety.
[0400] Preparation of compositions and systems As will be appreciated by those skilled in the art, methods for designing and constructing nucleic acid constructs and proteins or polypeptides (e.g., the systems, constructs, and polypeptides described herein) are routine in the art. Generally, recombinant methods may be used. For general information, see Smales & James (Eds.), Therapeutic Proteins: Methods and Protocols (Methods in Molecular Biology), Humana Press (2005); and Crommelin, Sindelar & Meibohm (Eds.), Pharmaceutical Biotechnology: Fundamentals and Applications, Springer (2013). Methods for designing, preparing, evaluating, purifying, and manipulating nucleic acid compositions are described in Green and Sambrook (Eds.), Molecular Cloning: A Laboratory Manual (Fourth Edition), Cold Spring Harbor Laboratory Press (2012).
[0401] Exemplary methods for producing pharmaceutical proteins or polypeptides described herein include expression in mammalian cells, although recombinant proteins can also be produced using insect cells, yeast, bacteria, or other cells under the control of an appropriate promoter. Mammalian expression vectors can include non-transcriptional elements, such as an origin of replication, a suitable promoter, and other 5' or 3' flanking non-transcribed sequences, as well as 5' or 3' non-transcribed sequences, such as necessary ribosome binding sites, polyadenylation sites, splice donor and acceptor sites, and termination sequences. DNA sequences derived from the SV40 viral genome, such as the SV40 origin, early promoter, splice, and polyadenylation sites, can be used to provide other genetic elements required for expression of heterologous DNA sequences. Cloning and expression vectors suitable for use in bacterial, fungal, yeast, and mammalian cell hosts are described in Green & Sambrook, Molecular Cloning: A Laboratory Manual (Fourth Edition), Cold Spring Harbor Laboratory Press (2012).
[0402] Various mammalian cell culture systems can be used to express and produce recombinant proteins. Examples of mammalian expression systems include CHO, COS, HEK293, HeLA, and BHK cell lines. Host culture methods for protein therapeutic production are described in Zhou and Kantardjieff (Eds.), Mammalian Cell Cultures for Biologics Manufacturing (Advances in Biochemical Engineering / Biotechnology), Springer (2014). The compositions described herein can include a vector encoding the recombinant protein, e.g., a viral vector such as a lentiviral vector. In some embodiments, the vector, e.g., a viral vector, can include a nucleic acid encoding the recombinant protein.
[0403] Purification of protein therapeutics is described in Franks, Protein Biotechnology: Isolation, Characterization, and Stabilization, Humana Press (2013); and Cutler, Protein Purification Protocols (Methods in Molecular Biology), Humana Press (2010).
[0404] RNA (e.g., gRNA or mRNA, e.g., mRNA encoding a GeneWriter) can also be produced as described herein. In some embodiments, RNA segments can be produced by chemical synthesis. In some embodiments, RNA segments can be produced by in vitro transcription of a nucleic acid template, e.g., by providing an RNA polymerase that acts on a cognate promoter of a DNA template to produce an RNA transcript. In some embodiments, in vitro transcription is performed using, e.g., T7, T3, or SP6 RNA polymerase, or a derivative thereof, acting on DNA, e.g., dsDNA, ssDNA, linear DNA, plasmid DNA, linear DNA amplicon, linearized plasmid DNA, encoding the RNA segment, under the transcriptional control of a cognate promoter, e.g., a T7, T3, or SP6 promoter. In some embodiments, a combination of chemical synthesis and in vitro transcription is used to generate RNA segments for assembly. In several embodiments, gRNAs are produced by chemical synthesis, and heterologous sequence-of-interest segments are produced by in vitro transcription. Without intending to be bound by any particular theory, in vitro transcription is believed to be more suitable for producing longer RNA molecules. In some embodiments, the reaction temperature for in vitro transcription may be lowered, e.g., below 37°C (e.g., 0-10°C, 10-20°C, or 20-30°C), to obtain a higher percentage of full-length transcripts (see Krieg Nucleic Acids Res 18:6463 (1990)), which is incorporated herein by reference in its entirety). In some embodiments, improved protocols for the synthesis of long transcripts are used to synthesize long RNAs, e.g., RNAs greater than 5 kb; for example, 27 kb transcripts can be produced in vitro using T7 RiboMAX Express (Thiel et al. J Gen Virol 82(6):1273-1281 (2001)).In some embodiments, modifications to RNA molecules described herein may be incorporated during synthesis of RNA segments (e.g., by inclusion of modified nucleotides or alternative linking chemistries), followed by synthesis of RNA segments by chemical or enzymatic processing, followed by assembly of one or more RNA segments, or a combination thereof.
[0405] In some embodiments, system mRNAs (e.g., mRNAs encoding Gene Writer polypeptides) are synthesized in vitro using T7 polymerase-mediated DNA-dependent RNA transcription from a linearized DNA template, where UTP is optionally replaced with 1-methylpseudoUTP. In some embodiments, transcripts incorporate 5' and 3' UTRs, e.g., GGGAAAAUAAGAGAGAAAAGAAGAGUAAGAAGAAAUAUAAGAGCCACC (SEQ ID NO: 1871) and UGAUAAUAGGCUGGAGCCUCGGUGGCCAUGCUUCUUGCCCCUUGGGCCUCCCCCCAGCCCCUCCUCCCCUUCCUGCACCCGUACCCCCGUGGUCUUUGAAUAAAGUCUGA (SEQ ID NO: 1872), or a functional fragment or variant thereof, and optionally include a poly-A tail, which may be encoded in the DNA template or added enzymatically post-transcriptionally. In some embodiments, a donor methyl group, e.g., S-adenosylmethionine, is added to a methylated capped RNA having a cap 0 structure to obtain a cap 1 structure, which enhances mRNA translation efficiency (Richner et al. Cell 168(6):P1114-1125(2017)).
[0406] In some embodiments, transcripts from a T7 promoter begin with a GGG motif. In some embodiments, transcripts from a T7 promoter do not begin with a GGG motif. Although a GGG motif at the transcription start site provides superior yields, slippage of the transcript against three C residues from +1 to +3 in the template strand results in T7 RNAP synthesizing a ladder of poly(G) products (Imburgio et al., Biochemistry 39(34):10419-10430 (2000)). The teachings of Davidson et al., Pac Symp Biocomput 433-443 (2010) on adjusting transcription levels and modifying transcription start site nucleotides to accommodate alternative 5'UTRs describe T7 promoter variants, and methods for their discovery, that meet both of these properties.
[0407] In some embodiments, RNA segments may be covalently linked to one another. In some embodiments, two or more RNA segments may be joined to one another using an RNA ligase, e.g., T4 RNA ligase. When using a reagent such as RNA ligase, the 5' end is typically ligated to the 3' end. In some embodiments, when two segments are joined, two linear constructs may be formed (i.e., (1) 5'-Segment 1 - Segment 2-3' and (2) 5'-Segment 2 - Segment 1-3'). In some embodiments, intramolecular circularization may also occur. Either of these issues can be addressed, for example, by blocking one 5' end or one 3' end so that the RNA ligase cannot ligate that end to the other end. In embodiments, if a 5'-Segment 1 - Segment 2-3' construct is desired, placing a blocking group at the 5' end of Segment 1 or the 3' end of Segment 2 can ensure the formation of only the correct linear ligation product and / or prevent intramolecular circularization. Compositions and methods for the covalent joining of two nucleic acid (e.g., RNA) segments are disclosed, for example, in U.S. Patent Application No. 20160102322A1, which is incorporated herein by reference in its entirety, along with methods that involve the use of RNA ligase to directional join two single-stranded RNA segments together.
[0408] For example, one example of an end blocker that can be used with T4 RNA ligase is a dideoxy terminator. T4 RNA ligase typically catalyzes the ATP-dependent ligation of a phosphodiester bond between a 5'-phosphate and a 3'-hydroxyl terminus. In some embodiments, when using T4 RNA ligase, a suitable terminus must be present at the end to be ligated. One way to block T4 RNA ligase on an end involves preventing the acquisition of the correct end format. In general, ends of RNA segments containing a 5-hydroxyl or a 3'-phosphate will not act as substrates for T4 RNA ligase.
[0409] Another exemplary method that can be used to link RNA segments is via click chemistry (e.g., as described in U.S. Pat. Nos. 7,375,234 and 7,070,941, and U.S. Patent Application Publication No. 2013 / 0046084, the entire disclosures of which are incorporated herein by reference). For example, one exemplary click chemistry reaction is between an alkyne group and an azide group (see Figure 11 of U.S. Patent Application Publication No. 20160102322A1, the entire disclosures of which are incorporated herein by reference). Any click reaction can be used to link RNA segments (e.g., Cu-azide-alkyne, strain-promoted azide-alkyne, Staudinger ligation, tetrazine ligation, photoinduced tetrazole-alkene, thiol-ene, NHS ester, epoxide, isocyanate, and aldehyde-aminooxy). In some embodiments, ligation of RNA molecules using click chemistry reactions is advantageous because click chemistry reactions are fast, modular, efficient, often do not produce toxic waste products, can be performed using water as a solvent, and / or can be set up stereospecifically.
[0410] In some embodiments, RNA segments may be linked using the azide-alkyne Huisgen cycloaddition reaction, which is typically a 1,3-dipolar cycloaddition reaction between an azide and a terminal or internal alkyne, providing a 1,2,3-triazole for linking RNA segments. Without intending to be bound by theory, one advantage of this linking method may be that the reaction can be initiated by the addition of the required Cu(I) ion. Other exemplary mechanisms by which RNA segments can be linked include, but are not limited to, the use of halogen (F, Br, I) / alkyne addition reactions, carbonyl / sulfhydryl / maleimide, and carboxyl / amine linkages. For example, one RNA molecule can be modified at the 3′ end with a thiol (using a disulfide amidite and a universal support or a disulfide-modified support) and the other RNA molecule can be modified at the 5′ end with an acrydite (using an acrylic acid phosphoramidite), and then the two RNA molecules can be linked via a Michael addition reaction. This strategy can also be applied to stepwise linking of multiple RNA molecules.Also provided is a method for linking three or more (for example, 3, 4, 5, 6, etc.) RNA molecules together.Without intending to be bound by a particular theory, this can be significant when the desired RNA molecule is longer than about 40 nucleotides, and the efficiency of chemical synthesis decreases, as pointed out in, for example, US Patent Application No. 20160102322A1 (the entirety of which is incorporated herein by reference).
[0411] By way of example, tracrRNA is typically approximately 80 nucleotides in length. Such RNA molecules can be produced by processes such as in vitro transcription or chemical synthesis. In some embodiments, when chemical synthesis is used to produce such RNA molecules, they can be produced as a single synthetic product or by ligating two or more RNA segments together. In some embodiments, when three or more RNA segments are ligated together, different methods can be used to ligate the individual segments together. Additionally, the RNA segments can be ligated together all at the same time in one pot (e.g., a container, vessel, well, tube, plate, or other container), or at different times in one pot, or at different times in different pots. In a non-limiting example, to assemble RNA segments 1, 2, and 3 in numerical order, RNA segments 1 and 2 can first be ligated together 5' to 3'. The reaction product can then be purified of reaction mixture components (e.g., by chromatography) and then placed in a second pot for ligation of its 3' end to the 5' end of RNA segment 3. The final reaction product can then be ligated to the 5' end of RNA segment 3.
[0412] In another non-limiting example, RNA segment 1 (approximately 30 nucleotides) is a target locus recognition sequence consisting of the crRNA and a portion of hairpin region 1. RNA segment 2 (approximately 35 nucleotides) contains the remainder of hairpin region 1 and a portion of the linear tracrRNA between hairpin region 1 and hairpin region 2. RNA segment 3 (approximately 35 nucleotides) contains the remainder of the linear tracrRNA between hairpin region 1 and hairpin region 2 and all of hairpin region 2. In this example, RNA segments 2 and 3 are linked 5' to 3' using click chemistry. Furthermore, both the 5' and 3' ends of the reaction product are phosphorylated. Next, the reaction product is contacted with RNA segment 1 having a 3'-terminal hydroxyl group and T4 RNA ligase to generate a guide RNA molecule.
[0413] Several alternative linking chemistries may be used to join RNA segments according to the methods of the present invention, some of which are listed in Table 6 of U.S. Patent Application Publication No. 20160102322A1, which is incorporated herein by reference in its entirety.
[0414] vector The present disclosure provides, in part, a nucleic acid, e.g., a Gene Writer polypeptide described herein, a template nucleic acid described herein, or both. In some embodiments, the vector comprises a selectable marker, e.g., an antibiotic resistance marker. In some embodiments, the antibiotic resistance marker is a kanamycin resistance marker. In some embodiments, the antibiotic resistance marker does not confer resistance to a beta-lactam antibiotic. In some embodiments, the vector does not comprise an ampicillin resistance marker. In some embodiments, the vector comprises a kanamycin resistance marker and does not comprise an ampicillin resistance marker. In some embodiments, the vector encoding the Gene Writer polypeptide integrates into the target cell genome (e.g., upon administration to the target cell, tissue, organ, or subject). In some embodiments, the vector encoding the Gene Writer polypeptide does not integrate into the target cell genome (e.g., upon administration to the target cell, tissue, organ, or subject). In some embodiments, the vector comprising the template nucleic acid (e.g., template DNA) does not integrate into the target cell genome (e.g., upon administration to the target cell, tissue, organ, or subject). In some embodiments, when the vector is integrated into the target site in the target cell genome, the selectable marker is not integrated into the genome. In some embodiments, when the vector is integrated into the target site in the target cell genome, genes or sequences involved in vector maintenance (e.g., plasmid maintenance genes) are not integrated into the genome. In some embodiments, when the vector is integrated into the target site in the target cell genome, import regulatory sequences (e.g., inverted terminal sequences, e.g., from AAV) are not integrated into the genome. In some embodiments, administration of a vector (e.g., encoding a Gene Writer polypeptide described herein, a template nucleic acid described herein, or both) to a target cell, tissue, organ, or subject results in integration of a portion of the vector into one or more target sites in the genome of the target cell, tissue, organ, or subject.In some embodiments, less than 99, 95, 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, 4, 3, 2, or 1% of the target sites containing integrated material (e.g., no target sites) contain a selectable marker (e.g., an antibiotic resistance gene), an import regulatory sequence (e.g., an inverted terminal end sequence, e.g., from an AAV), or both, from the vector.
[0415] AAV vectors In some embodiments, the vector encoding the Gene Writer polypeptide described herein, the template nucleic acid described herein, or both is an adeno-associated virus (AAV) vector, e.g., comprising the AAV genome. In some embodiments, the AAV genome comprises two genes encoding four replication proteins and three capsid proteins, respectively. In some embodiments, these genes are flanked on both sides by 145-bp inverted terminal repeats (ITRs). In some embodiments, virions comprise up to three capsid proteins (Vp1, Vp2, and / or Vp3), produced in a ratio of, for example, 1:1:10. In some embodiments, the capsid proteins are produced from the same open reading frame and / or from different splicing (Vp1) and separate translation initiation sites (Vp2 and Vp3, respectively). Generally, Vp3 is the most abundant subunit in virions and participates in receptor recognition on the cell surface, determining viral tropism. In some embodiments, Vp1 comprises a phospholipase domain, eg, at the N-terminus of Vp1, that functions in viral infectivity.
[0416] In some embodiments, the packaging capacity of a viral vector limits the size of the base editor that can be packaged into the vector, for example, the packaging capacity of an AAV may be about 4.5 kb (e.g., about 3.0, 3.5, 4.0, 4.5, 5.0, 5.5, or 6.0 kb), including, for example, one or two inverted terminal repeats (ITRs), e.g., a 145-base ITR.
[0417] In some embodiments, recombinant AAV (rAAV) contains cis-acting 145-bp ITRs flanking the vector transgene cassette, providing, for example, up to 4.5 kb for packaging of foreign DNA. After infection, rAAV, in some cases, expresses proteins described herein and persists as an episome in circular head-to-tail concatemers without integrating into the host genome. rAAV can be used, for example, in vitro and in vivo. In some embodiments, AAV-mediated gene delivery requires that the length of the gene coding sequence be equal to or greater than that of the wild-type AAV genome.
[0418] AAV delivery of genes exceeding the above sizes and / or the use of large physiological regulatory elements can be achieved, for example, by splitting the protein to be delivered into two or more fragments. In some embodiments, the N-terminal fragment is fused to a split intein-N. In some embodiments, the C-terminal fragment is fused to a split intein-C. In some embodiments, the fragments are packaged into two or more AAV vectors.
[0419] In some embodiments, dual AAV vectors are generated by splitting a large transgene expression cassette into two separate halves (5's and 3's ends, or head and tail), for example, and then packaging each half of the cassette into a single AAV vector (<5 kb). Reassembly of the full-length transgene expression cassette can then be achieved in some embodiments upon co-infection of the same cell with both dual AAV vectors. In some embodiments, co-infection is followed by one or more of the following: (1) homologous recombination (HR) between the 5's and 3's genomes (dual AAV overlapping vector); (2) ITR-mediated tail-to-head concatenation of the 5's and 3's genomes (dual AAV trans-splicing vector); and / or (3) a combination of these two mechanisms (dual AAV hybrid vector). In some embodiments, full-length protein expression is achieved in vivo using dual AAV vectors. In some embodiments, the use of a dual AAV vector platform represents an efficient and viable gene transfer strategy for transgenes greater than about 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, or 5.0 kb in size. In some embodiments, AAV vectors can also be used to transduce cells with target nucleic acids, for example, in the in vitro production of nucleic acids and peptides. In some embodiments, AAV vectors can be used in in vivo and ex vivo gene therapy procedures (see, e.g., West et al., Virology 160:38-47 (1987); U.S. Pat. No. 4,797,368; WO 93 / 24641; Kotin, Human Gene Therapy 5:793-801 (1994); Muzyczka, J. Clin. Invest. 94:1351 (1994); each of which is incorporated by reference in its entirety).The construction of recombinant AAV vectors has been described in several publications, including U.S. Pat. No. 5,173,414; Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985); Tratschin, et al., Mol. Cell. Biol. 4:2072-2081 (1984); Hermonat & Muzyczka, PNAS 81:6466-6470 (1984); and Samulski et al., J. Virol. 63:03822-3828 (1989), which are incorporated herein by reference in their entireties.
[0420] In some embodiments, the Gene Writers described herein (e.g., with or without one or more guide nucleic acids) can be delivered using AAV, lentivirus, adenovirus, or other plasmid or viral vector types, among others, using formulations and dosages obtained, for example, from U.S. Pat. No. 8,454,972 (formulations, dosages for adenovirus), U.S. Pat. No. 8,404,658 (formulations, dosages for AAV), and U.S. Pat. No. 5,846,946 (formulations, dosages for DNA plasmids), as well as clinical trials involving lentivirus, AAV, and adenovirus and publications related to such clinical trials. For example, in the case of AAV, the route of administration, formulation, and dosage can be as described in U.S. Pat. No. 8,454,972 and in clinical trials involving AAV. In the case of adenovirus, the route of administration, formulation, and dosage can be as described in U.S. Pat. No. 8,404,658 and in clinical trials involving adenovirus. For plasmid delivery, the route of administration, formulation, and dosage may be as described in U.S. Pat. No. 5,846,946 and in clinical trials involving the plasmid. Dosages may be based on or extrapolated to an average 70 kg individual (e.g., an adult male) and can be adjusted for patients, subjects, or mammals of different weights and species. The frequency of administration is within the purview of a medical or veterinary professional (e.g., a physician or veterinarian), depending on the usual factors, including the patient's or subject's age, sex, general health, and other conditions, as well as the specific medical condition or symptom being addressed. In some embodiments, the viral vector can be injected into the tissue of interest. In the case of cell-type-specific gene writing, expression of the gene writer and optional guide nucleic acid can, in some embodiments, be driven by a cell-type-specific promoter.
[0421] In some embodiments, AAV allows for low toxicity due to purification methods that do not require ultracentrifugation of cellular particles that can activate an immune response. In some embodiments, AAV allows for a low likelihood of causing insertional mutagenesis, for example, because it does not substantially integrate into the host genome.
[0422] In some embodiments, the AAV has a packaging range of about 4.4, 4.5, 4.6, 4.7, or 4.75 kb. In some embodiments, the gene writer, promoter, and transcription terminator can fit into a single viral vector. SpCas9 (4.1 kb) can be difficult to package into AAV in some cases. Therefore, in some embodiments, a gene writer with a shorter length than other gene writers or base editors is used. In some embodiments, the Gene Writer is less than about 4.5kb, 4.4kb, 4.3kb, 4.2kb, 4.1kb, 4kb, 3.9kb, 3.8kb, 3.7kb, 3.6kb, 3.5kb, 3.4kb, 3.3kb, 3.2kb, 3.1kb, 3kb, 2.9kb, 2.8kb, 2.7kb, 2.6kb, 2.5kb, 2kb, or 1.5kb.
[0423] The AAV may be AAV1, AAV2, AAV5, or any combination thereof. In some embodiments, the type of AAV is selected based on the cells to be targeted; for example, AAV serotypes 1, 2, 5, or hybrid capsids AAV1, AAV2, AAV5, or any combination thereof can be selected to target brain or neuronal cells; or AAV4 can be selected to target cardiac tissue. In some embodiments, AAV8 is selected for delivery to the liver. Exemplary AAV serotypes for these cells are described, for example, in Grimm, D. et al., J. Virol. 82:5887-5911 (2008), which is incorporated herein by reference in its entirety. In some embodiments, AAV refers to all serotypes, subtypes, and naturally occurring AAVs as well as recombinant AAVs. When AAV is used, it may refer to the virus itself or its derivatives. In some embodiments, AAV includes AAV1, AAV2, AAV3, AAV3B, AAV4, AAV5, AAV6, AAV6.2, AAV7, AAVrh.64Rl, AAVhu.37, AAVrh.8, AAVrh.32.33, AAV8, AAV9, AAV-DJ, AAV2 / 8, AAVrhlO, AAVLK03, AV10, AAV11, AAV 12, rhlO, and hybrids thereof, avian AAV, bovine AAV, canine AAV, equine AAV, primate AAV, non-primate AAV, and ovine AAV. The genome sequences of various AAV serotypes, as well as the sequences of native terminal repeats (TRs), Rep proteins, and capsid subunits, are known in the art. These sequences can be found in the literature or in official databases such as GenBank. Other exemplary AAV serotypes are listed in Table 7.
[0424] [Table 7]
[0425] In some embodiments, a pharmaceutical composition (e.g., comprising an AAV described herein) has less than 10% empty capsids, less than 8% empty capsids, less than 7% empty capsids, less than 5% empty capsids, less than 3% empty capsids, or less than 1% empty capsids. In some embodiments, the pharmaceutical composition has less than about 5% empty capsids. In some embodiments, the number of empty capsids is below the limit of detection. In some embodiments, it is advantageous for a pharmaceutical composition to have a small number of empty capsids, for example, because empty capsids may cause adverse reactions (e.g., immune, inflammatory, hepatic, and / or cardiac) and, for example, provide little or no substantial therapeutic benefit.
[0426] In some embodiments, the residual host cell protein (rHCP) in the pharmaceutical composition is greater than or equal to 1×10 13 rHCPs less than 100 ng / ml per vg / ml, e.g., 1 × 10 13 rHCP less than 40ng / ml per vg / ml or 1 x 10 13 In some embodiments, the pharmaceutical composition is 1.0 x 10 13 Less than 10 ng of rHCP per vg or 1.0 x 10 13 <5ng rHCP per vg, 1.0 × 10 13 Less than 4ng rHCP per vg or 1.0 x 10 13 In some embodiments, the residual host cell DNA (hcDNA) in the pharmaceutical composition is less than 1 x 10 13 5 x 10 per vg / ml 6 pg / ml or less of hcDNA, 1 × 10 13 1.2 × 10 per vg / ml 6 pg / ml or less of hcDNA, or 1 × 10 13 1 x 10 per vg / ml 5 In some embodiments, the residual host cell DNA in the pharmaceutical composition is 1 x 10 pg / ml hcDNA. 13 5.0 x 10 per vg 5 Less than pg, 1.0 × 10 132.0 x 10 per vg 5 Less than pg, 1.0 × 10 13 1.1 x 10 per vg 5 Less than pg, 1.0 × 10 13 1.0 x 10 per vg 5 Less than pg of hcDNA, 1.0 × 10 13 0.9 x 10 per vg 5 Less than pg of hcDNA, 1.0 × 10 13 0.8 x 10 per vg 5 Sub-pg of hcDNA, or any concentration therebetween.
[0427] In some embodiments, the residual plasmid DNA in the pharmaceutical composition is 1.0 x 10 13 1.7 × 10 per vg / ml 5 pg / ml or less, 1×1.0×10 13 1 x 10 per vg / ml 5 pg / ml, or 1.0 x 10 13 1.7 × 10 per vg / ml 6 In some embodiments, the residual DNA plasmid in the pharmaceutical composition is 1.0 x 10 13 10.0 x 10 per vg 5 Less than pg, 1.0 × 10 13 8.0 x 10 per vg 5 Less than pg or 1.0 x 10 13 6.8 x 10 per vg 5 In some embodiments, the pharmaceutical composition contains less than 1.0 x 10 pg. 13 Less than 0.5 ng per vg, 1.0 × 10 13 Less than 0.3 ng per vg, 1.0 × 10 13 Less than 0.22 ng per vg or 1.0 x 10 13 In some embodiments, the benzonase in the pharmaceutical composition contains less than 1.0 x 10 bovine serum albumin (BSA) per vg, or any intermediate concentration. 13 Less than 0.2 ng per vg, 1.0 × 10 13 Less than 0.1 ng per vg, 1.0 × 10 13 Less than 0.09 ng per vg, 1.0 × 10 13In some embodiments, the poloxamer 188 in the pharmaceutical composition is about 10-150 ppm, about 15-100 ppm, or about 20-80 ppm. In some embodiments, the cesium in the pharmaceutical composition is less than 50 pg / g (ppm), less than 30 pg / g (ppm), or less than 20 pg / g (ppm), or any intermediate concentration.
[0428] In embodiments, the pharmaceutical composition contains less than 10%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, or any percentage therebetween, of total impurities, e.g., as measured by SDS-PAGE. In embodiments, the total impurities are greater than 90%, greater than 92%, greater than 93%, greater than 94%, greater than 95%, greater than 96%, greater than 97%, greater than 98%, or any percentage therebetween, e.g., as measured by SDS-PAGE. In embodiments, any single unspecified associated impurity is greater than 5%, greater than 4%, greater than 3%, or greater than 2%, or any percentage therebetween, e.g., as measured by SDS-PAGE. In some embodiments, the pharmaceutical composition comprises a percentage of filled capsids relative to total capsids (Peak 1 + Peak 2, as measured by analytical ultracentrifugation) of greater than 85%, greater than 86%, greater than 87%, greater than 88%, greater than 89%, greater than 90%, greater than 91%, greater than 91.9%, greater than 92%, greater than 93%, or any percentage therebetween. In some embodiments of the pharmaceutical composition, the percentage of filled capsids, as measured in Peak 1 by analytical ultracentrifugation, is 20-80%, 25-75%, 30-75%, 35-75%, or 37.4-70.3%. In some embodiments of the pharmaceutical composition, the percentage of filled capsids, as measured in Peak 2 by analytical ultracentrifugation, is 20-80%, 20-70%, 22-65%, 24-62%, or 24.9-60.1%.
[0429] In one embodiment, the pharmaceutical composition contains 1.0 to 5.0 x 10 13 vg / mL, 1.2–3.0 × 10 13 vg / mL, or 1.7–2.3 × 10 13In one embodiment, the pharmaceutical composition exhibits a bioburden of less than 5 CFU / mL, less than 4 CFU / mL, less than 3 CFU / mL, less than 2 CFU / mL, or less than 1 CFU / mL, or any intermediate concentration. In some embodiments ... <85> (incorporated by reference in its entirety) is less than 1.0 EU / mL, less than 0.8 EU / mL, or less than 0.75 EU / mL. In some embodiments, the amount of endotoxin according to USP, e.g., USP <785> (incorporated by reference in its entirety) is 350-450 mOsm / kg, 370-440 mOsm / kg, or 390-430 mOsm / kg. In embodiments, the pharmaceutical composition contains fewer than 1200 particles greater than 25 μm per container, fewer than 1000 particles greater than 25 μm per container, fewer than 500 particles greater than 25 μm per container, or any intermediate value. In embodiments, the pharmaceutical composition contains fewer than 10,000 particles greater than 10 μm per container, fewer than 8000 particles greater than 10 μm per container, or fewer than 600 particles greater than 10 μm per container.
[0430] In one embodiment, the pharmaceutical composition contains 0.5 to 5.0 x 10 13 vg / mL, 1.0–4.0 × 10 13 vg / mL, 1.5–3.0 × 10 13 vg / ml or 1.7 to 2.3 × 10 13 In one embodiment, the pharmaceutical compositions described herein comprise one or more of the following: 1.0 x 10 13 Less than about 0.09 ng of benzonase per vg, less than about 30 pg / g (ppm) of cesium, about 20-80 ppm of poloxamer 188, 1.0 x 10 13 Less than approximately 0.22 ng of BSA per vg, 1.0 × 10 13 Approximately 6.8 × 10 per vg 5 Less than pg of residual DNA plasmid, 1.0 × 10 13 Approximately 1.1 × 10 per vg 5 Residual hcDNA less than pg, 1.0 × 10 13Less than approximately 4 ng rHCP per vg, pH 7.7-8.3, approximately 390-430 mOsm / kg, less than approximately 600 particles >25 μm in diameter per container, less than approximately 6000 particles >10 μm in diameter per container, approximately 1.7 × 10 13 ~2.3×10 13 vg / mL genome titer, 1.0 × 10 13 Approximately 3.9 × 10 per vg 8 ~8.4×10 10 Infectious titer in IU, 1.0 × 10 13 Approximately 100–300 pg of total protein per vg, approximately 7.5 × 10 13 A7SMA mice treated with a 2000mg / kg dose of viral vector exhibit a median survival time of >24 days, a relative efficacy of about 70-130% based on an in vitro cell assay, and / or less than about 5% empty capsids. In various embodiments, the pharmaceutical compositions described herein, comprising any of the viral particles featured herein, retain a potency of ±20%, ±15%, ±10%, or ±5% of the reference standard. In some embodiments, potency is measured using a suitable in vitro cell assay or in vivo animal model.
[0431] Additional methods for preparing, characterizing, and administering AAV particles are taught in WO2019094253, the entire contents of which are incorporated herein by reference.
[0432] Additional rAAV constructs that can be used consistent with the present invention include those described in Wang et al. 2019, including Table 1, available at: http: / / doi.org / 10.1038 / s41573-019-0012-9 (incorporated by reference in its entirety).
[0433] Kits, Articles of Manufacture, and Pharmaceutical Compositions In one aspect, the disclosure provides kits including a Gene Writer or Gene Writing system, e.g., as described herein. In some embodiments, the kit includes a Gene Writer polypeptide (or a nucleic acid encoding the polypeptide) and template DNA. In some embodiments, the kit further includes reagents for introducing the system into cells, e.g., transfection reagents, LNPs, etc. In some embodiments, the kit is suitable for any of the methods described herein. In some embodiments, the kit includes one or more elements, e.g., a composition (e.g., a pharmaceutical composition), a Gene Writer, and / or a Gene Writer system, or functional fragments or components thereof, as disposed in an article of manufacture. In some embodiments, the kit includes instructions for its use.
[0434] In one aspect, the present disclosure provides an article of manufacture, such as an article of manufacture into which a kit described herein, or elements thereof, is disposed.
[0435] In one aspect, the present disclosure provides a pharmaceutical composition comprising, for example, a Gene Writer or Gene Writing system described herein. In some embodiments, the pharmaceutical composition further comprises a pharmaceutically acceptable carrier or excipient. In some embodiments, the pharmaceutical composition comprises template DNA.
[0436] Chemistry, Manufacturing, and Controls (CMC) Purification of protein therapeutics is described, for example, in Franks, Protein Biotechnology: Isolation, Characterization, and Stabilization, Humana Press (2013); and Cutler, Protein Purification Protocols (Methods in Molecular Biology), Humana Press (2010).
[0437] In some embodiments, the Gene Writer™ system, polypeptide, and / or template nucleic acid (e.g., template DNA) meets certain quality standards. In some embodiments, the Gene Writer™ system, polypeptide, and / or template nucleic acid (e.g., template DNA) produced by the methods described herein meets certain quality standards. Accordingly, the present disclosure, in some aspects, relates to methods of producing Gene Writer™ systems, polypeptides, and / or template nucleic acids that meet certain quality standards, e.g., where the quality standards are assayed. The present disclosure, in some aspects, also relates to methods of assaying the quality standards in Gene Writer™ systems, polypeptides, and / or template nucleic acids. In some embodiments, the quality standards include, but are not limited to, one or more of the following (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12): (i) the length of the template DNA or mRNA encoding the GeneWriter polypeptide, e.g., whether the DNA or mRNA exceeds a standard length or has a length within a standard length range, e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the DNA or mRNA present exceeds 100, 125, 150, 175, or 200 nucleotides in length; (ii) the presence, absence, and / or length of a poly-A tail on the mRNA, e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the mRNAs present contain a poly-A tail (e.g., a poly-A tail at least 5, 10, 20, 30, 50, 70, 100 nucleotides in length); (iii) the presence, absence, and / or type of 5' cap on the mRNA, e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the mRNAs present contain a 5' cap, e.g., whether the cap is a 7-methylguanosine cap, e.g., an O-Me-m7G cap; (iv) the presence, absence, and / or type of one or more modified nucleotides (e.g., selected from pseudouridine, dihydrouridine, inosine, 7-methylguanosine, 1-N-methylpseudouridine (1-Me-Ψ), 5-methoxyuridine (5-MO-U), 5-methylcytidine (5mC), or locked nucleotides) in the mRNA, e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the mRNA present contains one or more modified nucleotides; (v) the stability of the template DNA or mRNA (e.g., over time and / or under preselected conditions), e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the DNA or mRNA remains intact (e.g., greater than 100, 125, 150, 175, or 200 nucleotides in length) after stability testing; (vi) the efficacy of the template DNA or mRNA in the system for modifying the DNA, e.g., whether at least 1% of the target sites are modified after assaying the system containing the DNA or mRNA for efficacy; (vii) the length of the polypeptide, first polypeptide, or second polypeptide, e.g., whether the polypeptide, first polypeptide, or second polypeptide has a length that exceeds a standard length or is within a standard length range, e.g., at least 80, 85, 90, 95, 96, 97, 98, or 99% of the polypeptides, first polypeptides, or second polypeptides present are 600, 650, 700, 750, 800, whether it is more than 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, 1600, 1700, 1800, 1900, or 2000 amino acids in length (and optionally is less than or equal to 2500, 2000, 1500, 1400, 1300, 1200, 1100, 1000, 900, 800, 700, or 600 amino acids in length); (viii) the presence, absence, and / or type of post-translational modifications to the polypeptide, the first polypeptide, or the second polypeptide, e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the polypeptide, the first polypeptide, or the second polypeptide contains phosphorylation, methylation, acetylation, myristoylation, palmitoylation, isoprenylation, glipyatyon, or lipoylation, or any combination thereof; (ix) one or more artificial, synthetic, or non-standard amino acids (e.g., ornithine, β-alanine, GABA, δ-aminolevulinic acid, PABA, D-amino acids (e.g., D-alanine or D-glutamic acid), aminoisobutyric acid, dehydroalanine, cystathionine, lanthionine, djenkolic acid) in the polypeptide, first polypeptide, or second polypeptide; the presence, absence, and / or type of (a) amino acid(s), diaminopimelic acid, homoalanine, norvaline, norleucine, homonorleucine, homoserine, O-methyl-homoserine and O-ethyl-homoserine, ethionine, selenocysteine, selenohomocysteine, selenomethionine, selenoethionine, tellurocysteine, or telluromethionine), e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the polypeptide, first polypeptide, or second polypeptide contains one or more artificial, synthetic, or non-standard amino acids; (x) stability of the polypeptide, first polypeptide, or second polypeptide (e.g., over time and / or under preselected conditions), e.g., at least 80, 85, 90, 95, 96, 97, 98, or 99% of the polypeptide, first polypeptide, or second polypeptide remains intact (e.g., 600, 650, 700, 750, 800, 850, 900, 10 ... , 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, 1600, 1700, 1800, 1900, or 2000 amino acids in length (and optionally, is not more than 2500, 2000, 1500, 1400, 1300, 1200, 1100, 1000, 900, 800, 700, or 600 amino acids in length); (xi) the efficacy of the polypeptide, the first polypeptide, or the second polypeptide in the system for modifying DNA, e.g., whether at least 1% of the target sites are modified after assaying the system including the polypeptide, the first polypeptide, or the second polypeptide for efficacy; or (xii) The presence, absence, and / or level of one or more pyrogens, viruses, fungi, bacterial pathogens, or host cell proteins, e.g., whether the system is free or substantially free of pyrogen, virus, fungi, bacterial pathogen, or host cell protein contamination.
[0438] In some embodiments, the systems or pharmaceutical compositions described herein are endotoxin-free.
[0439] In some embodiments, the presence, absence, and / or level of one or more of pyrogens, viruses, fungi, bacterial pathogens, and / or host cell proteins is determined, hi several embodiments, the system determines whether the system is free or substantially free of pyrogen, virus, fungi, bacterial pathogen, and / or host cell protein contamination.
[0440] In some embodiments, the pharmaceutical compositions or systems described herein have one or more (e.g., one, two, three, or four) of the following characteristics: (a) less than 1% (e.g., 0.5%, 0.4%, 0.3%, 0.2%, or 0.1%) of DNA template relative to the RNA encoding the polypeptide, e.g., on a molar basis; (b) less than 1% (e.g., 0.5%, 0.4%, 0.3%, 0.2%, or 0.1%) of uncapped RNA relative to the RNA encoding the polypeptide, e.g., on a molar basis; (c) less than 1% (e.g., 0.5%, 0.4%, 0.3%, 0.2%, or 0.1%) of partial-length RNA, e.g., on a molar basis, relative to the RNA encoding the polypeptide; (d) substantially free of unreacted cap dinucleotide.
[0441] Exemplary Heterologous Sequences of Interest In some embodiments, the systems or methods described herein include a heterologous sequence of interest, wherein the heterologous sequence of interest or its reverse complement encodes a protein (e.g., an antibody) or a peptide. In some embodiments, the therapeutic is one that has been approved by a regulatory agency, such as the FDA.
[0442] In some embodiments, the protein or peptide is a protein or peptide from the THPdb database (Usmani et al. PLoS One 12(7):e0181748(2017) (incorporated herein by reference in its entirety). In some embodiments, the protein or peptide is a protein or peptide disclosed in Table 8. In some embodiments, a system or method disclosed herein (e.g., one including Gene Writer) may be used to incorporate an expression cassette for a protein or peptide from Table 8 into a host cell to enable expression of the protein or peptide in the host. In some embodiments, the sequence of the protein or peptide in column 1 of Table 8 can be found in the patent or application provided in column 3 of Table 8 (incorporated herein by reference in its entirety).
[0443] In some embodiments, the protein or peptide is an antibody disclosed in Table 1 of Lu et al. J Biomed Sci 27(1):1 (2020), which is incorporated by reference in its entirety. In some embodiments, the protein or peptide is an antibody disclosed in Table 9. In some embodiments, a system or method disclosed herein (e.g., one including a Gene Writer) may be used to incorporate an expression cassette for an antibody from Table 9 into a host cell to allow expression of the antibody in the host. In some embodiments, a system or method described herein is used to express an agent that binds to a target in column 2 of Table 9 (e.g., a monoclonal antibody in column 1 of Table 9) in a subject with an indication in column 3 of Table 9.
[0444] [Table 8-1]
[0445] [Table 8-2]
[0446] [Table 8-3]
[0447] [Table 8-4]
[0448] [Table 8-5]
[0449] [Table 8-6]
[0450] [Table 8-7]
[0451]
Table 8-8
[0452]
Table 8-9
[0453]
Table 8-10
[0454]
Table 9-1
[0455]
Table 9-2
[0456]
Table 9-3
[0457] Applicable The Gene Writer system can address therapeutic needs by incorporating a coding gene into a DNA sequence template, e.g., by providing expression of a therapeutic transgene in individuals with a loss-of-function mutation (e.g., contained in a sequence of interest as described herein), by replacing a gain-of-function mutation with a normal transgene, by providing regulatory sequences to eliminate gain-of-function mutation expression, and / or by controlling the expression of operably linked genes, transgenes, and systems thereof. In certain embodiments, a sequence of interest (e.g., a heterologous sequence of interest) comprises a coding sequence that encodes a functional element (e.g., a polypeptide or non-coding RNA, as described herein) specific to the therapeutic need of the host cell. In some embodiments, a sequence of interest (e.g., a heterologous sequence of interest) comprises a promoter, e.g., a tissue-specific promoter or enhancer. In certain embodiments, a promoter can be operably linked to a coding sequence.
[0458] In some embodiments, the Gene Writer™ gene editor system can provide sequences of interest, including therapeutic agents (e.g., therapeutic transgenes) that express, for example, substitute blood factors or substitute enzymes, e.g., lysosomal enzymes. For example, the compositions, systems, and methods described herein are useful for expressing in a target human genome agalsidase alfa or beta for the treatment of Fabry disease; imiglucerase, taliglucerase alfa, velaglucerase alfa, or alglucerase for Gaucher disease; sebelipase alfa for lysosomal acid lipase deficiency (Wolman disease / CESD); laronidase, idursulfase, elosulfase alfa, or galsulfase for mucopolysaccharidoses; or alglucosidase for Pompe disease. For example, the compositions, systems, and methods described herein are useful for expressing in a target human genome factors I, II, V, VII, X, XI, XII, or XIII for blood factor deficiencies.
[0459] Administration The compositions and systems described herein can be used in vitro or in vivo. In some embodiments, the system or components of the system are delivered to cells (e.g., mammalian cells, e.g., human cells), for example, in vitro or in vivo. Those skilled in the art will understand that components of the Gene Writer system can be delivered in the form of polypeptides, nucleic acids (e.g., DNA, RNA), and combinations thereof.
[0460] In some embodiments, the system and / or system components are delivered as nucleic acids. For example, a recombinase polypeptide may be delivered in the form of DNA or RNA encoding the recombinase polypeptide. In some embodiments, the system or system components (e.g., insert DNA and nucleic acid molecule encoding the recombinase polypeptide) are delivered on one, two, three, four, or more separate nucleic acid molecules. In some embodiments, the system or system components are delivered as a combination of DNA and RNA. In some embodiments, the system or system components are delivered as a combination of DNA and protein. In some embodiments, the system or system components are delivered as a combination of RNA and protein. In some embodiments, the system or system components are delivered as a combination of RNA and protein. In some embodiments, the recombinase polypeptide is delivered as a protein.
[0461] In some embodiments, the system or components of the system are delivered to cells, e.g., mammalian or human cells, using a vector. The vector may be, for example, a plasmid or a virus. In some embodiments, delivery is in vivo, in vitro, ex vivo, or in situ. In some embodiments, the virus is an adeno-associated virus (AAV), a lentivirus, or an adenovirus. In some embodiments, the system or components of the system are delivered to cells using a virus-like particle or a virosome. In some embodiments, delivery uses a virus, a virus-like particle, or a virosome.
[0462] In one embodiment, the compositions and systems described herein can be formulated in liposomes or other similar vesicles. Liposomes are spherical vesicular structures composed of a monolayer or multilayer lipid bilayer membrane surrounding an internal aqueous compartment and a relatively impermeable outer lipophilic phospholipid bilayer membrane. Liposomes can be anionic, neutral, or cationic. Liposomes are biocompatible, non-toxic, and can deliver both hydrophilic and lipophilic drug molecules, protecting their cargo from degradation by plasma enzymes and transporting their load across biological membranes and the blood-brain barrier (BBB) (see, e.g., Spuch and Navarro, Journal of Drug Delivery, vol. 2011, Article ID 469679, 12 pages, 2011. doi:10.1155 / 2011 / 469679 for a review).
[0463] Vesicles can be made from several different types of lipids; however, phospholipids are most commonly used to make liposomes as drug carriers.The method of preparing multilamellar vesicle lipids is known in the art (see, for example, U.S. Patent No. 6,693,086; its teachings about preparing multilamellar vesicle lipids are incorporated herein by reference).Vesicle formation can occur spontaneously when lipid membrane is mixed with aqueous solution, but it can also be promoted by applying force in the form of shaking using a homogenizer, ultrasonic generator or extruder (for example, for review, see Spuch and Navarro, Journal of Drug Delivery, vol.2011, Article ID 469679, 12 pages, 2011.doi:10.1155 / 2011 / 469679). Extruded lipids can be prepared by extrusion through size-reducing filters as described in Templeton et al., Nature Biotech, 15:647-652, 1997, the teachings of which regarding extruded lipid preparation are incorporated herein by reference.
[0464] Lipid nanoparticles are another example of a carrier that provides a biocompatible and biodegradable delivery system for the pharmaceutical compositions described herein. Nanostructured lipid carriers (NLCs) are modified solid lipid nanoparticles (SLNs) that retain the characteristics of SLNs, improve drug stability and loading capacity, and prevent drug leakage. Polymer nanoparticles (PNPs) are an important component of drug delivery. These nanoparticles effectively target drug delivery to specific targets and improve drug stability and controlled drug release. Lipid-polymer nanoparticles (PLNs), a new type of carrier that combines liposomes and polymers, may also be used. These nanoparticles possess the complementary advantages of PNPs and liposomes. PLNs are composed of a core-shell structure; the polymer core provides a stable structure, while the phospholipid shell confers good biocompatibility. In this way, the two components enhance drug encapsulation efficiency, promote surface modification, and prevent leakage of water-soluble drugs. See, e.g., Li et al. 2017, Nanomaterials 7, 122; doi:10.3390 / nano7060122 for a review.
[0465] Exosomes can also be used as drug delivery vehicles for the compositions and systems described herein. See, e.g., Ha et al., July 2016. Acta Pharmaceutica Sinica B. Volume 6, Issue 4, Pages 287-296; https: / / doi.org / 10.1016 / j.apsb.2016.02.001 for a review.
[0466] In some embodiments, at least one component of the systems described herein comprises a fusosome. Fusosomes interact with and fuse with target cells, allowing them to be used as delivery vehicles for various molecules. Fusosomes are generally composed of a bilayer of amphiphilic lipids that encloses a lumen or cavity, and a fusogen that interacts with the amphiphilic lipid bilayer. The fusogen component has been shown to confer target cell specificity for fusion and payload delivery, and can be manipulated to allow for the creation of delivery vehicles with programmable cell specificity (see, e.g., the section on fusosome design, preparation, and use in WO2020014209, which is incorporated herein by reference in its entirety).
[0467] The GeneWriter system can be introduced into cells, tissues, and multicellular organisms. In some embodiments, the system or components of the system are delivered to cells using mechanical or physical means.
[0468] Formulation of protein therapeutics is described in Meyer (Ed.), Therapeutic Protein Drug Products: Practical Approaches to Formulation in the Laboratory, Manufacturing, and the Clinic, Woodhead Publishing Series (2012).
[0469] In some embodiments, the Gene Writer™ systems described herein are delivered to tissues or cells from the following: cerebrum, cerebellum, adrenal glands, ovaries, pancreas, parathyroid glands, pituitary gland, testes, thyroid gland, breast, spleen, tonsils, thymus, lymph nodes, bone marrow, lung, heart muscle, esophagus, stomach, small intestine, colon, liver, salivary glands, kidneys, prostate, blood, or other cell or tissue type. In some embodiments, the Gene Writer™ systems described herein are used to treat a disease, such as cancer, an inflammatory disease, an infectious disease, a genetic abnormality, or other disease. The cancer may be of the following: cerebrum, cerebellum, adrenal glands, ovaries, pancreas, parathyroid gland, pituitary gland, testes, thyroid gland, breast, spleen, tonsils, thymus, lymph nodes, bone marrow, lung, heart muscle, esophagus, stomach, small intestine, colon, liver, salivary glands, kidneys, prostate, blood, or other cell or tissue type, and may include multiple cancers.
[0470] In some embodiments, the Gene Writer™ systems described herein are administered enterally (e.g., oral, gastrointestinal, sublingual, sublabial, or buccal administration). In some embodiments, the Gene Writer™ systems described herein are administered parenterally (e.g., intravenous, intramuscular, subcutaneous, intradermal, epidural, intracerebral, intraventricular, epicutaneous, intranasal, intra-arterial, intra-articular, intracavernosal, intraocular, intraosseous infusion, intraperitoneal, intrathecal, intrauterine, intravaginal, intravesical, perivascular, or transmucosal administration). In some embodiments, the Gene Writer™ systems described herein are administered topically (e.g., transdermal administration).
[0471] In some embodiments, the Gene Writer™ system described herein can be used to modify animal cells, plant cells, or fungal cells. In some embodiments, the Gene Writer™ system described herein can be used to modify mammalian cells (e.g., human cells). In some embodiments, the Gene Writer™ system described herein can be used to modify cells from livestock animals (e.g., cows, horses, sheep, goats, pigs, llamas, alpacas, camels, yaks, chickens, ducks, geese, or ostriches). In some embodiments, the Gene Writer™ system described herein can be used as an experimental or research tool or in experimental or research methods, for example, to modify animal cells, e.g., mammalian cells (e.g., human cells), plant cells, or fungal cells.
[0472] In some embodiments, the Gene Writer™ systems described herein can be used to express a protein, template, or heterologous sequence of interest (e.g., in animal cells, e.g., mammalian cells (e.g., human cells), plant cells, or fungal cells). In some embodiments, the Gene Writer™ systems described herein can be used to express a protein, template, or heterologous sequence of interest under the control of an inducible promoter (e.g., a small molecule-inducible promoter). In some embodiments, the Gene Writing system or its payload is designed for tunable control, e.g., through the use of an inducible promoter. For example, a promoter driving a gene of interest, e.g., Tet, can be silent upon integration, but in some cases can be activated upon exposure to a small molecule inducer, e.g., doxycycline. In some embodiments, tunable expression allows for post-treatment control of a gene (e.g., a therapeutic gene), e.g., allowing for small molecule-dependent administration effects. In embodiments, the small molecule-dependent administration effects include altering the levels of a gene product temporally and / or spatially, e.g., through local administration. In some embodiments, promoters used in the systems described herein may be inducible, e.g., responsive to endogenous molecules of the host and / or exogenous small molecules administered thereto.
[0473] Treatment of suitable indications In some embodiments, the Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein) is used to treat a disease, disorder, or condition. In some embodiments, the Gene Writer™ system described herein, or a component or portion thereof, is used to treat a disease, disorder, or condition listed in any of Tables 10-15. In some embodiments, the Gene Writer™ system described herein, or a component or portion thereof, is used to treat, for example, a hematopoietic stem cell (HSC) disease, disorder, or condition listed in Table 10. In some embodiments, the Gene Writer™ system described herein, or a component or portion thereof, is used to treat, for example, a kidney disease, disorder, or condition listed in Table 11. In some embodiments, the Gene Writer™ system described herein, or a component or portion thereof, is used to treat, for example, a liver disease, disorder, or condition listed in Table 12. In some embodiments, the Gene Writer™ system described herein, or a component or portion thereof, is used to treat, for example, a pulmonary disease, disorder, or condition listed in Table 13. In some embodiments, the Gene Writer™ system described herein, or a component or portion thereof, is used to treat, for example, a musculoskeletal disease, disorder, or condition listed in Table 14. In some embodiments, the Gene Writer™ system described herein, or a component or portion thereof, is used to treat, for example, a skin disease, disorder, or condition listed in Table 15.
[0474] Tables 10-15: Selected indications for transGene Writers used with recombinases
[0475] [Table 10]
[0476] [Table 11]
[0477]
Table 12
[0478]
Table 13
[0479]
Table 14
[0480]
Table 15
[0481] In some embodiments, the Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein), is used to treat a genetic disease, disorder, or condition. In some embodiments, the Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein), is used to treat a subject (e.g., a human patient) diagnosed with a genetic disease, disorder, or condition. In some embodiments, the genetic disease, disorder, or condition is associated with a particular genotype, e.g., a heterozygous or homozygous genotype. In some embodiments, the genetic disease, disorder, or condition is associated with a particular mutation, e.g., a substitution, deletion, or insertion, e.g., a nucleotide expansion. In some embodiments, the genetic disease, disorder, or condition is cystic fibrosis or ornithine transcarbamylase (OTC) deficiency. In some embodiments, the Gene Writer™ systems described herein used to treat a genetic disease, disorder, or condition comprise a heterologous sequence of interest that comprises a functional (e.g., wild-type) copy of a gene that is deleted (e.g., globally or in a target cell population) in a subject (e.g., a human patient). In some embodiments, the functional copy of the gene comprises a functional (e.g., wild-type) CFTR gene or OTC gene.
[0482] In some embodiments, a Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein) is used to treat a subject (e.g., a human patient) having a level of a biomarker (e.g., associated with a disease, disorder, or condition, e.g., a genetic disease, disorder, or condition) outside of a healthy range. In some embodiments, a Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein) is used to treat a subject (e.g., a human patient) diagnosed as having a level of a biomarker (e.g., associated with a disease, disorder, or condition, e.g., a genetic disease, disorder, or condition) outside of a healthy range.
[0483] In some embodiments, the presence and / or level of a biomarker and / or genotype in a subject (e.g., a human patient) is determined before treatment with a Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein). In some embodiments, the presence and / or level of a biomarker and / or genotype in a subject (e.g., a human patient) is determined after treatment with a Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein). In some embodiments, the presence and / or level of a biomarker and / or genotype in a subject (e.g., a human patient) is determined before and after treatment with a Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein).
[0484] In some embodiments, a Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein) is administered in response to a determination that a biomarker is present in a subject (e.g., a human patient) at a level outside of a normal and / or healthy range. In some embodiments, a Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein) is readministered in response to a determination that a biomarker is present in a subject (e.g., a human patient) at a level outside of a normal and / or healthy range after an initial administration of a Gene Writer™ system described herein, or a component or portion thereof. In some embodiments, a Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein) is administered in response to a determination that a subject (e.g., a human patient), for example, or a target cell population in a subject, has a genotype (e.g., associated with a disease, disorder, or condition). In some embodiments, a Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein), is readministered to a subject (e.g., a human patient) following an initial administration of a Gene Writer™ system described herein, or a component or portion thereof, in response to a determination that the subject, e.g., or a target cell population in the subject, has a genotype (e.g., associated with a disease, disorder, or condition). In some embodiments, administration of a Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein), is continued or repeated until the biomarker is present in the subject (e.g., a human patient) at levels within the normal and / or healthy range.In some embodiments, administration of a Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein) is continued or repeated until the subject (e.g., a human patient), for example, or a target cell population within the subject, no longer possesses the genotype (e.g., associated with a disease, disorder, or condition).
[0485] In some embodiments, the Gene Writer™ systems described herein, or components or portions thereof (e.g., polypeptides or nucleic acids described herein) are used to treat a disease, disorder, or condition prenatally (e.g., in a human subject in utero, e.g., an embryo or fetus). In some embodiments, the Gene Writer™ systems described herein, or components or portions thereof (e.g., polypeptides or nucleic acids described herein) are used to treat a disease, disorder, or condition postnatally, e.g., in a human infant, toddler, or child. In some embodiments, the Gene Writer™ systems described herein, or components or portions thereof (e.g., polypeptides or nucleic acids described herein) are used to treat a disease, disorder, or condition during the neonatal period.
[0486] In some embodiments, the genotype of a subject (e.g., a human patient) treated with a Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein), or of a target cell population of the subject, remains stable throughout the development of the subject. In this context, stable refers to the absence of additional alterations to the subject's genotype (e.g., or a target cell population of the subject) after treatment with a Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein), is completed. In this context, stable can also or alternatively refer to the persistence of alterations to the subject's genotype made by a Gene Writer system described herein. Without intending to be bound by any particular theory, it may be desirable to avoid, prevent, or minimize additional alterations to the subject's genotype other than those made by the Gene Writer system. Additionally or alternatively, it may be desirable for the alteration of the genotype of the subject (e.g., or of a target cell population of the subject) to persist after completion of the treatment (e.g., for at least a selected period of time, e.g., indefinitely). In some embodiments, after completion of the treatment, the genotype of the subject (e.g., or of a target cell population of the subject) is the same as the genotype of the subject (e.g., or of a target cell population of the subject) for a selected period of time after the treatment, e.g., 1, 2, 3, 4, 5, 6, or 7 days, or 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 weeks, or 3, 4, 5, 6, 7, 8, 9, 10, or 11 months, or 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 years (e.g., indefinitely).In some embodiments, the modification to the genotype of a subject, e.g., or a target cell population of a subject, made by a Gene Writer™ system described herein, or a component or portion thereof (e.g., a polypeptide or nucleic acid described herein), persists for at least a selected period of time following treatment, e.g., 1, 2, 3, 4, 5, 6, or 7 days, or 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 weeks, or 3, 4, 5, 6, 7, 8, 9, 10, or 11 months, or 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 years (e.g., indefinitely).
[0487] Plant modification method The Gene Writer systems described herein can be used to modify plants or plant parts (eg, leaves, roots, flowers, fruits, or seeds), for example, to enhance the fitness of the plant.
[0488] A. Delivery to Plants Provided herein are methods for delivering the Gene Writer systems described herein to plants. These methods include delivering the Gene Writer systems to plants by contacting the plant, or a portion thereof, with the Gene Writer system. These methods are useful for modifying plants, for example, to enhance the fitness of the plant.
[0489] More specifically, in some embodiments, a nucleic acid described herein (e.g., a nucleic acid encoding a GeneWriter) may be encoded within a vector and inserted adjacent to a plant promoter, such as the maize ubiquitin promoter (ZmUBI), in a plant vector (e.g., pHUC411). In some embodiments, a nucleic acid described herein is introduced into a plant (e.g., japonica rice) or plant part (e.g., plant callus) via agrobacteria. In some embodiments, the systems and methods described herein can be used in plants by replacing a plant gene (e.g., hygromycin phosphotransferase (HPT)) with a null allele (e.g., containing a base substitution in the start codon). Systems and methods for modifying plant genomes are described in Xu et al., "Development of plant prime-editing systems for precise genome editing," 2020, Plant Communications.
[0490] In one aspect, provided herein is a method of increasing the fitness of a plant, the method comprising delivering to a plant a Gene Writer system described herein (e.g., in an effective amount and for a duration) to increase the fitness of the plant relative to an untreated plant (e.g., a plant that has not been delivered the Gene Writer system).
[0491] The increased fitness of the plant as a result of delivery of the Gene Writer system can be maintained in several ways, for example, thereby achieving improved plant production, such as increased yield, improved plant vigor or the quality of the product harvested from the plant, improved pre- or post-harvest characteristics (e.g., taste, appearance, shelf life) that are considered desirable in agriculture or horticulture, or improved characteristics that are otherwise beneficial to humans (e.g., reduced allergen production). Improved plant yield refers to an increase in the yield of a product of the plant (e.g., measurable by plant biomass, grain, seed or fruit yield, protein content, carbohydrate or oil content, or leaf area) by a measurable amount relative to the yield of the same product of a plant produced under the same conditions but without application of the composition, or compared to the application of a conventional plant modifier. For example, yield may be increased by at least about 0.5%, about 1%, about 2%, about 3%, about 4%, about 5%, about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 100%, or more than 100%. In some cases, the method is effective to increase yield by about 2-fold, 5-fold, 10-fold, 25-fold, 50-fold, 75-fold, 100-fold, or more than 100-fold compared to untreated plants. Yield can be expressed in terms of plant weight or volume, or plant product based on some standard. Standards can be expressed in terms of time, cultivated area, weight of plant produced, or amount of raw material used. For example, such methods can increase the yield of plant tissues such as, but not limited to, seeds, fruits, grains, nuts, tubers, roots, and leaves.
[0492] Increased plant fitness as a result of delivery of the Gene Writer system can also be measured by other means, for example, an increase or improvement in vigor rating, canopy (number of plants per area unit), plant height, stem condition, stem length, number of leaves, leaf size, canopy, appearance (e.g., dark green leaf color, etc.), root rating, emergence, protein content, increased tillering, increased leaf size, increased leaves, reduced basal leaves, stronger shoots, reduced fertilizer requirements, reduced seed requirements, more productive shoots, earlier flowering, earlier grain or seed maturity, reduced plant verse (lodging), increased shoot growth, earlier germination, or any combination of these factors, by a measurable or discernible amount compared to the same factors in plant product produced under the same conditions but without administration of the composition or with the application of a conventional plant modifier.
[0493] Accordingly, provided herein are methods of modifying plants, the methods comprising delivering to a plant an effective amount of any of the Gene Writer systems provided herein, wherein the method modifies the plant, thereby introducing or increasing a beneficial trait in the plant (e.g., about 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or more than 100%) compared to an untreated plant. In particular, the method may increase plant fitness (e.g., about 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or more than 100%) compared to an untreated plant.
[0494] In some cases, the increase in plant fitness is an increase (e.g., about 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or more than 100%) in disease resistance, drought tolerance, heat tolerance, cold tolerance, salt tolerance, metal tolerance, herbicide tolerance, drug tolerance, water use efficiency, nitrogen utilization, nitrogen stress tolerance, nitrogen fixation, pest resistance, herbivore resistance, pathogen resistance, yield, yield under water-limited conditions, vigor, growth, photosynthetic capacity, nutrition, protein content, carbohydrate content, oil content, biomass, shoot length, root length, root architecture, seed weight, or amount of harvestable product.
[0495] In some cases, the increase in fitness is an increase (e.g., about 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or more than 100%) in development, growth, yield, tolerance to abiotic stressors, or tolerance to biotic stressors. Abiotic stress refers to environmental stress conditions experienced by plants or plant parts, and includes, for example, drought stress, salt stress, heat stress, cold stress, and low nutrient stress. Biotic stress refers to environmental stress conditions experienced by plants or plant parts, and includes, for example, nematode stress, herbivorous insect stress, fungal pathogen stress, bacterial pathogen stress, or viral pathogen stress. Stress can be transient, e.g., hours, days, or months, or persistent, e.g., over the lifespan of the plant.
[0496] In some cases, an increase in plant fitness is an increase (e.g., about 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or more than 100%) in the quality of the product harvested from the plant. For example, an increase in plant fitness can be an improvement in a commercially desirable characteristic (e.g., taste or appearance) of the product harvested from the plant. In other cases, an increase in plant fitness is an increase (e.g., about 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or more than 100%) in the shelf life of the product harvested from the plant.
[0497] Alternatively, the increase in fitness can be an alteration of a characteristic beneficial to human or animal health, such as a reduction in allergen production. For example, the increase in fitness can be a reduction (e.g., about 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or more than 100%) in the production of an allergen (e.g., pollen) that stimulates an immune response in an animal (e.g., a human).
[0498] Modification of a plant (e.g., increased fitness) can result from modification of one or more plant parts. For example, a plant can be modified by contacting a plant's leaves, seeds, pollen, roots, fruit, shoots, flowers, cells, protoplasts, or tissues (e.g., meristems). Accordingly, in another aspect, provided herein is a method of increasing plant fitness, comprising contacting plant pollen with an effective amount of any of the plant modification compositions herein, which method increases the plant's fitness (e.g., by about 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or more than 100%) compared to an untreated plant.
[0499] In yet another aspect, provided herein is a method of increasing the fitness of a plant, the method comprising contacting a seed of the plant with an effective amount of any of the Gene Writer systems disclosed herein, which method increases the fitness of the plant (e.g., by about 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or more than 100%) compared to an untreated plant.
[0500] In yet another aspect, provided herein is a method comprising contacting a plant protoplast with an effective amount of any of the Gene Writer systems described herein, which method increases plant fitness (e.g., by about 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or greater than 100%) relative to an untreated plant.
[0501] In yet another aspect, provided herein is a method of increasing the fitness of a plant, the method comprising contacting a plant cell of a plant with an effective amount of any of the Gene Writer systems described herein, wherein the method increases the fitness of the plant (e.g., by about 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or more than 100%) relative to an untreated plant.
[0502] In another aspect, provided herein are methods of increasing the fitness of a plant, the methods comprising contacting a meristem of the plant with an effective amount of any of the plant modifying compositions herein, which method increases the fitness of the plant (e.g., about 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or more than 100%) relative to an untreated plant.
[0503] In another aspect, provided herein are methods of increasing plant fitness, the methods comprising contacting a plant embryo with an effective amount of any of the plant modifying compositions herein, which method increases the fitness of the plant (e.g., about 1%, 2%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or greater than 100%) relative to an untreated plant.
[0504] B. Application method Plants described herein can be exposed to any of the Gene Writer system compositions described herein in any suitable manner that allows for the delivery or administration of the composition to the plant. The Gene Writer system may be delivered alone or in combination with other active (e.g., fertilizer) or inactive substances and can be applied, for example, by spraying, injection (e.g., microinjection), plant-mediated pouring, or immersion in the form of a concentrate, gel, solution, suspension, spray, powder, pellet, briquette, brick, etc., formulated to deliver an effective concentration of the plant modifying composition. The amount and location for application of the compositions described herein will generally be determined by the plant's habitat, the life cycle stage at which the plant can be targeted by the plant modifying composition, the site where application is intended to occur, and the physical and functional properties of the plant modifying composition.
[0505] In some cases, the composition is sprayed directly onto plants, e.g., crops, for example, by backpack spraying, aerial spraying, crop spraying / dusting, etc. When the Gene Writer system is delivered to a plant, the plant receiving the Gene Writer system may be at any stage of plant growth. For example, the formulated plant modification composition can be applied as a seed coating or root treatment at an early stage of plant growth, or as a whole plant treatment at a later stage of the crop cycle. In some cases, the plant modification composition may be applied as a topical agent to the plant.
[0506] Additionally, the Gene Writer system may be applied as a systemic agent (e.g., to the soil in which the plant is grown or to the water used to water the plant) that is absorbed and distributed throughout the plant tissue. In some cases, the plant or prey organism may be genetically transformed to express the Gene Writer system.
[0507] Delayed or continuous release can also be achieved by coating the composition containing the Gene Writer system or plant-modified composition with a soluble or biodegradable coating layer, such as gelatin, which dissolves or degrades in the environment of use, after which the plant-modified composition becomes available, or by dispersing the agent in a soluble or degradable matrix. Such continuous release and / or dispersion means that the device can be advantageously used to maintain a constant effective concentration of one or more of the plant-modified compositions described herein.
[0508] In some cases, the Gene Writer system is delivered to a part of a plant, such as its leaf, seed, pollen, root, fruit, shoot, or flower, or a tissue, cell, or protoplast. In some cases, the Gene Writer system is delivered to a cell of a plant. In some cases, the Gene Writer system is delivered to a protoplast of a plant. In some cases, the Gene Writer system is delivered to a tissue of a plant. For example, the composition can be delivered to a plant meristematic tissue (e.g., an apical meristem, a lateral meristem, or an interstitial meristem). In some cases, the composition is delivered to a permanent tissue of a plant (e.g., a simple tissue (e.g., a parenchyma, a sclerenchyma, or a sclerenchyma) or a complex permanent tissue (e.g., a xylem or a phloem). In some cases, the Gene Writer system is delivered to a plant embryo.
[0509] C. Plant The Gene Writer system described herein can be delivered to a variety of plants for treatment. Plants to which the Gene Writer system can be delivered (i.e., "treated") according to the present method include whole plants and plant parts, including, but not limited to, shoot vegetative organs / structures (e.g., leaves, stems, and tubers), roots, flowers and floral organs / structures (e.g., bracts, sepals, petals, stamens, carpels, anthers, and ovules), seeds (including embryos, endosperm, cotyledons, and seed coats) and fruits (mature follicles), plant tissues (e.g., vascular tissue, ground tissue, etc.) and cells (e.g., guard cells, egg cells, etc.), and their progeny. Plant parts also refer to plant parts such as shoots, roots, stems, seeds, stipules, leaves, petals, flowers, ovules, bracts, branches, petioles, internodes, bark, pubescence, shoots, rhizomes, fronds, (grass) leaves, pollen, stamens, etc.
[0510] Plant classes that can be treated with the methods disclosed herein include higher and lower plant classes, including angiosperms (monocotyledons and dicotyledons), gymnosperms, ferns, horsetails, psilophytes, microphytes, bryophytes, and algae (e.g., multicellular or unicellular algae). Plants that can be treated according to the present methods further include any vascular plant, such as a monocotyledonous or dicotyledonous plant, or gymnosperm, including, but not limited to, alfalfa, apple, Arabidopsis, banana, canola, castor, chrysanthemum, clover, cocoa, coffee, cotton, cottonseed, corn, crambe, cranberry, cucumber, dendrobium, yam, eucalyptus, fescue, flax, gladiolus, liliacea, linseed, millet, muskmelon, mustard, oat, oil palm, oilseed rape, peanut, pineapple, ornamentals, Phaseolus, potato, rapeseed, rice, rye, ryegrass, safflower, sesame, sorghum, soybean. , sugar beets, sugarcane, sunflowers, strawberries, tobacco, tomatoes, lawn grass, wheat, and vegetable crops such as lettuce, celery, broccoli, cauliflower, and cucurbits; fruit and nut trees such as apples, pears, peaches, oranges, grapefruit, lemons, limes, almonds, pecans, walnuts, and hazel; vines such as grapes (e.g., vineyards), kiwi, and hops; fruit shrubs and brambles such as raspberries, blackberries, and gooseberries; forest trees such as ash, pine, fir, maple, oak, chestnut, and poplar; and alfalfa, canola, castor, corn, cotton, crambe, flax, linseed, mustard, oil palm, rapeseed, peanuts, potatoes, rice, safflower, sesame, soybeans, sugar beets, sunflowers, tobacco, tomatoes, and wheat. Plants that can be treated according to the methods disclosed herein include any crop plant, such as forage crops, oilseed crops, grain crops, fruit crops, vegetable crops, fiber crops, spice crops, nut crops, turf crops, sugar crops, beverage crops, and forest crops. In a particular example, the crop plant treated by the method is a soybean plant.In other particular cases, the crop plant is wheat. In other particular cases, the crop plant is corn. In other particular cases, the crop plant is cotton. In particular cases, the crop plant is alfalfa. In particular cases, the crop plant is sugar beet. In particular cases, the crop plant is rice. In particular cases, the crop plant is potato. In particular cases, the crop plant is tomato.
[0511] In certain instances, the plant is a crop plant. Examples of such crop plants include, but are not limited to, monocotyledonous and dicotyledonous plants, including, but not limited to, forage or fodder legumes, ornamental plants, food crops, trees, or shrubs selected from the following: maple spp., Allium spp., Amaranthus spp., pineapple (Ananas comosus), celery (Apium graveolens), peanut spp., asparagus (Asparagus officinalis), sugar beet (Beta vulgaris), Brassica spp. (e.g., Brassica napus, Brassica rapa ssp. (canola, oilseed rape, turnip rape), and the like. rape), Camellia sinensis, Cannabis indica, Cannabis saliva, Capsicum spp., Chestnut spp., Endive (Cichorium endivia), Watermelon (Citrullus lanatus), Citrus spp., Cocos spp., Coffee spp., Coriander (Coriandrum sativum), Hazel spp., Crataegus spp., Cucurbita spp., Cucumis spp., Carrot (Daucus carota), Beech spp. spp.), figs (Ficus carica), Fragaria spp., ginkgo (Ginkgo biloba), soybean spp. (e.g., Glycine max, Soja hispida, or Soja max), upland cotton (Gossypium hirsutum), sunflower spp. (Helianthus spp.) (e.g., sunflower (Helianthus annuus)), Hibiscus spp., Hordeum spp. (e.g., barley (Hordeum vulgare)), sweet potato (Ipomoea batatas), walnut spp., lettuce (Lactuca sativa), flax (Linum usitatissimum), litchi (Litchi chinensis), Lotus spp., Luffa acutangula, lupin spp., tomato spp. (e.g., tomato (Lycopersicon esculenturn), Lycopersicon lycopersicum), lycopersicum, Lycopersicon pyriforme), Apple species (Malus spp.), Medicago sativa, Mint species (Mentha spp.), Miscanthus sinensis, Black mulberry (Morus nigra), Musa spp., Tobacco species (Nicotiana spp.), Olive species (Olea spp.), Oryza species (e.g., rice (Oryza sativa), Oryza latifolia)), Millet (Panicum miliaceum), Switchgrass (Panicum virgatum), Passion fruit (Passiflora edulis), Italian parsley (Petroselinum crispum), Phaseolus spp., Pine spp., Pistachio (Pistacia vera), Pea spp., Poa spp., Populus spp., Cherry spp., Prunus spp., Pear (Pyrus communis), Quercus spp.), radish (Raphanus sativus), rhubarb (Rheum rhabarbarum), currants (Ribes spp.), castor beans (Ricinus communis), rubus spp., sugarcane spp. (Saccharum spp.), willow spp. (Salix spp.), elderberry spp. (Sambucus spp.), rye (Secale cereale), sesame spp., Sinapis spp., nightshade spp. (e.g. potato (Solanum tuberosum), eggplant (Solanum integrifolium) or tomato (Solanum lycopersicum)), sorghum (Sorghum bicolor), corn sorghum (Sorghum halepense, Spinacia spp., Tamarindus indica, Theobroma cacao, Trifolium spp., Triticosecale rimpaui, Triticum spp. (e.g., Triticum aestivum, Triticum durum, Triticum turgidum, Triticum hybernum, Triticum macha, Triticum sativum, or Triticum vulgare), Vaccinium spp., Vicia spp.), Vigna spp., Viola odorata, Vitis spp., and Zea mays. In particular embodiments, the crop plant is rice, rapeseed, canola, soybean, corn (maize), cotton, sugarcane, alfalfa, sorghum, or wheat.
[0512] Plants or plant parts used in the present invention include plants at any stage of plant development. In certain cases, delivery can occur during germination, seedling growth, vegetative growth, and reproductive growth. In certain cases, delivery to a plant occurs during vegetative and upright growth stages. In some cases, the composition is delivered to the pollen of a plant. In some cases, the composition is delivered to the seed of a plant. In some cases, the composition is delivered to a protoplast of a plant. In some cases, the composition is delivered to a tissue of a plant. For example, the composition can be delivered to a meristematic tissue of a plant (e.g., an apical meristem, a lateral meristem, or an interstitial meristem). In some cases, the composition is delivered to a permanent tissue of a plant (e.g., a simple tissue (e.g., a parenchyma, a sclerenchyma, or a sclerenchyma) or a complex permanent tissue (e.g., a xylem or a phloem). In some cases, the composition is delivered to a plant embryo. In some cases, the composition is delivered to a plant cell. Vegetative and reproductive growth stages are also referred to herein as "adult" or "mature" plants.
[0513] In instances where the Gene Writer system is delivered to a plant part, the plant part may be modified by a plant modifier, or the Gene Writer system may be distributed to other parts of the plant (e.g., by the plant's circulatory system) and those parts then modified by the plant modifier.
[0514] lipid nanoparticles The methods and systems provided by the present invention may, in certain embodiments, include any suitable carrier or delivery method, including lipid nanoparticles (LNPs). In some embodiments, the lipid nanoparticles include one or more ionic lipids, such as non-cationic lipids (e.g., neutral or anionic, or zwitterionic lipids); one or more conjugated lipids (e.g., PEG-conjugated lipids or lipids conjugated with polymers, such as those described in Table 5 of International Publication No. 2019217941, the entire contents of which are incorporated herein by reference); one or more sterols (e.g., cholesterol); and, optionally, one or more targeting molecules (e.g., conjugated receptors, receptor ligands, antibodies); or combinations thereof.
[0515] Lipids that can be used in nanoparticle formation (e.g., lipid nanoparticles) include, for example, those listed in Table 4 of WO2019217941 (incorporated herein by reference), e.g., lipid-containing nanoparticles can include one or more of the lipids listed in Table 4 of WO2019217941. Lipid nanoparticles can include another element, such as a polymer, e.g., a polymer listed in Table 5 of WO2019217941 (incorporated herein by reference).
[0516] In some embodiments, the conjugated lipid, if present, is PEG-diacylglycerol (DAG) (e.g., 1-(monomethoxy-polyethylene glycol)-2,3-dimyristoylglycerol (PEG-DMG)), PEG-dialkyloxypropyl (DAA), PEG-phospholipid, PEG-ceramide (Cer), PEGylated phosphatidylethanolamine (PEG-PE), PEG-diacylglycerol succinate (PEGS-DAG) (e.g., 4-O-(2',3'-di( and one or more of N-(carbonyl-1-methoxypolyethylene glycol 2000)-1,2-distearoyl-sn-glycero-3-phosphoethanolamine sodium salt, and WO 2019051289 (incorporated herein by reference), and combinations thereof.
[0517] In some embodiments, sterols that can be incorporated into lipid nanoparticles include cholesterol or cholesterol derivatives thereof, such as those described in WO 2009 / 127060 or U.S. Patent Application No. 2010 / 0130588, which are incorporated herein by reference. Another exemplary sterol is a phytosterol, including those described in Eygeris et al. (2020), dx.doi.org / 10.1021 / acs.nanolett.0c01386, which are incorporated herein by reference.
[0518] In some embodiments, lipid nanoparticles comprise an ionizable lipid, a non-cationic lipid, a conjugated lipid that inhibits particle aggregation, and a sterol. The amounts of these components can be varied independently to achieve desired properties. For example, in some embodiments, the lipid nanoparticles comprise about 20 mol% to about 90 mol% of the total lipids (in other embodiments, it can be 20-70% (mol), 30-60% (mol), or 40-50% (mol); about 50 mol% to about 90 mol% of the total lipids present in the lipid nanoparticles), about 5 mol% to about 30 mol% of the total lipids, about 0.5 mol% to about 20 mol% of the total lipids, and about 20 mol% to about 50 mol% of the total lipids. The ratio of total lipid to nucleic acid (e.g., encoding a gene writer or template nucleic acid) can be varied as desired. For example, the ratio of total lipid to nucleic acid (mass or weight) can be about 10:1 to about 30:1.
[0519] In some embodiments, the lipid to nucleic acid ratio (mass / mass ratio; w / w ratio) can range from about 1:1 to about 25:1, about 10:1 to about 14:1, about 3:1 to about 15:1, about 4:1 to about 10:1, about 5:1 to about 9:1, or about 6:1 to about 9:1. The amounts of lipid and nucleic acid can be adjusted to provide a desired N / P ratio, e.g., an N / P ratio of 3, 4, 5, 6, 7, 8, 9, or 10 or greater. Generally, the total lipid content of a lipid nanoparticle formulation can range from about 5 mg / ml to about 30 mg / ml.
[0520] Exemplary ionizable lipids that can be used in lipid nanoparticle formulations include, but are not limited to, those listed in Table 1 of International Publication No. WO2019051289 (incorporated herein by reference). Additional exemplary lipids include, but are not limited to, one or more of the following formulas: X of U.S. Patent Application No. 2016 / 0311759; I of U.S. Patent Application No. 20150376115 or U.S. Patent Application No. 2016 / 0376224; I, II, or III of U.S. Patent Application No. 20160151284; I, IA, II, or IIA of U.S. Patent Application No. 20170210967; Ic of U.S. Patent Application No. 20150140070; No. 8541, paragraph A; U.S. Patent Application No. 2013 / 0303587 or U.S. Patent Application No. 2013 / 0123338, paragraph I; U.S. Patent Application No. 2015 / 0141678, paragraph I; U.S. Patent Application No. 2015 / 0239926, paragraph II, III, IV, or V; U.S. Patent Application No. 2017 / 0119904, paragraph I; WO 2017 / 117528, paragraph I or II; U.S. Patent Application No. 2012 / 0149894, paragraph A; U.S. Patent Application No. 2015 / 005737 No. 3, A; WO 2013 / 116126, A; U.S. Patent Application No. 2013 / 0090372, A; U.S. Patent Application No. 2013 / 0274523, A; U.S. Patent Application No. 2013 / 0274504, A; U.S. Patent Application No. 2013 / 0053572, A; WO 2013 / 016058, A; WO 2012 / 162210, A; U.S. Patent Application No. 2008 / 042973, I; U.S. Patent Application I, II, III or IV of U.S. Patent Application No. 2012 / 01287670; I or II of U.S. Patent Application No. 2014 / 0200257; I, II or III of U.S. Patent Application No. 2015 / 0203446; I or III of U.S. Patent Application No. 2015 / 0005363; I, IA, IB, IC, ID, II, IIA, IIB, IIC, IID, or III-XXIV of U.S. Patent Application No. 2014 / 0308304; U.S. Patent Application No. 2013 / 0338210;I, II, III, or IV of WO 2009 / 132131; A of U.S. Patent Application No. 2012 / 01011478; I or XXXV of U.S. Patent Application No. 2012 / 0027796; XIV or XVII of U.S. Patent Application No. 2012 / 0058144; U.S. Patent Application No. 2013 / 0323269; U.S. Patent Application No. 2011 / 0117125 No. I of U.S. Patent Application No. 2011 / 0256175; U.S. Patent Application No. 2012 / 0202871; U.S. Patent Application No. 2011 / 0076335; U.S. Patent ... U.S. Patent Application No. 2006 / 008378, I or II; U.S. Patent Application No. 2013 / 0123338, I; U.S. Patent Application No. 2015 / 0064242, I or XAYZ; U.S. Patent Application No. 2013 / 0022649, XVI, XVII, or XVIII; U.S. Patent Application No. 2013 / 0116307, I, II, or III; U.S. Patent Application No. 2013 / 011630 No. 7, I, II, or III; U.S. Patent Application No. 2010 / 0062967, I or II; U.S. Patent Application No. 2013 / 0189351, I-X; U.S. Patent Application No. 2014 / 0039032, I; U.S. Patent Application No. 2018 / 0028664, V; U.S. Patent Application No. 2016 / 0317458, I; U.S. Patent Application No. 2013 / 0195920, I;
[0521] In some embodiments, the ionizable lipid is MC3(6Z,9Z,28Z,3 lZ)-heptatriaconta-6,9,28,3,l-tetraen-l9-yl-4-(dimethylpentylamino)butanoate (DLin-MC3-DMA or MC3), for example, as described in Example 9 of WO2019051289A9 (incorporated herein by reference). In some embodiments, the ionizable lipid is the lipid ATX-002, for example, as described in Example 10 of WO2019051289A9 (incorporated herein by reference). In some embodiments, the ionizable lipid is (13Z,16Z)-A,A-dimethyl-3-nonyldocosa-13,16-dien-1-amine (compound 32), e.g., as described in Example 11 of WO2019051289A9 (incorporated herein by reference). In some embodiments, the ionizable lipid is compound 6 or compound 22, e.g., as described in Example 12 of WO2019051289A9 (incorporated herein by reference).
[0522] Exemplary non-cationic lipids include, but are not limited to, distearoyl-sn-glycero-phosphoethanolamine, distearoylphosphatidylcholine (DSPC), dioleoylphosphatidylcholine (DOPC), dipalmitoylphosphatidylcholine (DPPC), dioleoylphosphatidylglycerol (DOPG), dipalmitoylphosphatidylglycerol (DPPG), dioleoyl-phosphatidylethanolamine (DOPE), palmitoyloleoylphosphatidylcholine (POPC), palmitoyloleoylphosphatidylethanolamine (POPE), 4-(N-maleimidomethyl)-2-(2-methyl-2-propanol), 4-(N-maleimidomethyl)-2-propanol, ... )-Cyclohexane-1-carboxylic acid dioleoyl-phosphoethanolamine (DOPE-mal), dipalmitoylphosphatidylethanolamine (DPPE), dipalmyristoylphosphatidylethanolamine (DMPE), distearoyl-phosphatidyl-ethanolamine (DSPE), monomethyl-phosphatidylethanolamine (e.g., 16-O-monomethyl PE), dimethyl-phosphatidylethanolamine (e.g., 16-O-dimethyl PE), 18-1-trans PE, 1-stearoyl-2-oleoyl-1-phosphatidylethanolamine (SOPE), hydrogenated soy phosphatidylcholine (HSPC), egg phosphatidylcholine (EPC), dioleoylphosphatidylserine (DOPS), sphingomyelin ( SM), dimyristoyl phosphatidylcholine (DMPC), dimyristoyl phosphatidylglycerol (DMPG), distearoyl phosphatidylglycerol (DSPG), dierucoyl phosphatidylcholine (DEPC), palmitoyloleoyl phosphatidylglycerol (POPG), dielaidoyl phosphatidylethanolamine (DEPE), lecithin, phosphatidylethanolamine, lysolecithin, lysophosphatidylethanolamine, phosphatidylserine, phosphatidylinositol, sphingomyelin, egg sphingomyelin (ESM), cephalin, cardiolipin, phosphatidic acid, cerebrosides, dicetyl phosphate, lysophosphatidylcholine, dilinoleoyl phosphatidylcholine, or mixtures thereof. Additionally, it is understood that other diacetylphosphatidylcholine and diacetylphosphatidylethanolamine phospholipids can also be used. The acyl groups in these lipids are preferably derived from fatty acids having C10-24 carbon chains, such as lauroyl, myristoyl, palmitoyl, stearoyl, or oleoyl. Additional exemplary lipids in certain embodiments include, but are not limited to, those described in Kim et al. (2020) dx.doi.org / 10.1021 / acs.nanolett.0c01386 (incorporated herein by reference).Such lipids include, in some embodiments, plant lipids (eg, DGTS) that have been shown to improve hepatic transfection with mRNA.
[0523] Other examples of non-cationic lipids suitable for use in lipid nanoparticles include, but are not limited to, non-phospholipids such as stearylamine, dodeylamine, hexadecylamine, acetyl palmitate, glycerol ricinoleate, hexadecyl stearate, isopropyl myristate, amphoteric acrylic polymers, triethanolamine lauryl sulfate, alkyl-aryl sulfate polyethoxylated fatty acid amides, dioctadecyldimethylammonium bromide, ceramide, sphingomyelin, etc. Other non-cationic lipids are described in WO 2017 / 099823 or U.S. Patent Application Publication No. 2018 / 0028664, the contents of which are incorporated herein by reference in their entireties.
[0524] In some embodiments, the non-cationic lipid is oleic acid or a compound of Formula I, II, or IV of U.S. Patent Application Publication No. 2018 / 0028664 (incorporated herein by reference in its entirety). The non-cationic lipid may comprise, for example, 0-30% (mol) of the total lipid present in the lipid nanoparticle. In some embodiments, the non-cationic lipid content is, for example, 5-20% (mol) or 10-15% (mol) of the total lipid present in the lipid nanoparticle. In several embodiments, the molar ratio of ionizable lipid to neutral lipid ranges from about 2:1 to about 8:1 (e.g., about 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, or 8:1).
[0525] In some embodiments, the lipid nanoparticles do not include any phospholipids.
[0526] In some embodiments, the lipid nanoparticles may further comprise components such as sterols to confer membrane integrity. One example of a sterol that can be used in lipid nanoparticles is cholesterol and its derivatives. Non-limiting examples of cholesterol derivatives include polar analogs such as 5a-cholestanol, 53-coprostanol, cholesteryl-(2-hydroxy)-ethyl ether, cholesteryl-(4'-hydroxy)-butyl ether, and 6-ketocholestanol; non-polar analogs such as 5a-cholestane, cholestenone, 5a-cholestanone, 5p-cholestanone, and cholesteryl decanoate; and mixtures thereof. In some embodiments, the cholesterol derivative is a polar analog, for example, cholesteryl-(4'-hydroxy)-butyl ester. Exemplary cholesterol derivatives are described in International Publication No. 2009 / 127060 and U.S. Patent Application Publication No. 2010 / 0130588 (each of which is incorporated herein by reference in its entirety).
[0527] In some embodiments, membrane integrity-conferring components, such as sterols, comprise 0-50% (mol) of the total lipid present in the lipid nanoparticle (e.g., 0-10%, 10-20%, 20-30%, 30-40%, or 40-50%). In some embodiments, such components comprise 20-50% (mol), 30-40% (mol) of the total lipid in the lipid nanoparticle.
[0528] In some embodiments, the lipid nanoparticles may contain polyethylene glycol (PEG) or conjugated lipid molecules. Generally, these inhibit aggregation of lipid nanoparticles and / or provide steric stabilization. Exemplary conjugated lipids include, but are not limited to, PEG-lipid conjugates, polyoxazoline (POZ)-lipid conjugates, polyamide-lipid conjugates (e.g., ATTA-lipid conjugates), cationic polymer lipid (CPL) conjugates, and mixtures thereof. In some embodiments, the conjugated lipid molecule is a PEG-lipid conjugate, for example, a (methoxypolyethylene glycol)-conjugated lipid.
[0529] Exemplary PEG-lipid conjugates include, but are not limited to, PEG-diacylglycerol (DAG) (e.g., 1-(monomethoxy-polyethylene glycol)-2,3-dimyristoylglycerol (PEG-DMG)), PEG-dialkyloxypropyl (DAA), PEG-phospholipid, PEG-ceramide (Cer), PEGylated phosphatidylethanolamine (PEG-PE), PEG diacylglycerol succinate (PEGS-DAG) (e.g., 4-O-(2',3'-di(tetradecanoyloxy)propyl)propionyl ... Other exemplary PEG-lipid conjugates include those described in, for example, U.S. Patent Nos. 5,885,613, 6,287,591, and U.S. Patent Application Publication No. 2003 / 0077829. No. 2003 / 0077829, U.S. Patent Application Publication No. 2005 / 0175682, U.S. Patent Application Publication No. 2008 / 0020058, U.S. Patent Application Publication No. 2011 / 0117125, U.S. Patent Application Publication No. 2010 / 0130588, U.S. Patent Application Publication No. 2016 / 0376224, U.S. Patent Application Publication No. 2017 / 0119904, and U.S. Patent Application Publication No. 2017 / 0119904, all of which are incorporated herein by reference in their entireties. In some embodiments, the PEG-lipid is a compound of formula III, III-aI, III-a-2, III-b-1, III-b-2 of U.S. Patent Application Publication No. 2018 / 0028664 (the contents of which are incorporated herein by reference in their entirety). In some embodiments, the PEG-lipid is a compound of formula II of U.S. Patent Application Publication No. 20150376115 or U.S. Patent Application Publication No. 2016 / 0376224 (the contents of which are incorporated herein by reference in their entirety).In some embodiments, the PEG-DAA conjugate can be, for example, PEG-dilauryloxypropyl, PEG-dimyristyloxypropyl, PEG-dipalmityloxypropyl, or PEG-distearyloxypropyl. PEG-lipids can be any of the following: PEG-DMG, PEG-dilaurylglycerol, PEG-dipalmitoylglycerol, PEG-disterylglycerol, PEG-dilaurylglycamide, PEG-dimyristylglycamide, PEG-dipalmitoylglycamide, PEG-disterylglycamide, PEG-cholesterol (1-[8'-(cholest-5-en-3[β]-oxy)carboxamido-3',6'-dioxaotanyl]carbamoyl-[ω]-methyl-poly(ethylene glycol), PEG-DMB (3 ,4-ditetradecoxylbenzyl-[ω]-methyl-poly(ethylene glycol) ether), and 1,2-dimyristoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000]. In some embodiments, the PEG-lipid comprises PEG-DMG, 1,2-dimyristoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000]. In some embodiments, the PEG-lipid is one or more of the following: [ka] The structure includes a structure selected from:
[0530] In some embodiments, lipids conjugated with molecules other than PEG can be used instead of PEG-lipids. For example, polyoxazoline (POZ)-lipid conjugates, polyamide-lipid conjugates (e.g., ATTA-lipid conjugates), and cationic polymer lipid (GPL) conjugates can be used instead of PEG-lipids.
[0531] Exemplary conjugated lipids, i.e., PEG-lipids, (POZ)-lipid conjugates, ATTA-lipid conjugates, and cationic polymeric lipids, are described in the PCT and LIS patent applications listed in Table 2 of WO2019051289A9, the contents of all of which are incorporated herein by reference in their entireties.
[0532] In some embodiments, PEG or conjugated lipids may account for 0-20% (mol) of the total lipids present in the lipid nanoparticles. In some embodiments, the PEG or conjugated lipid content is 0.5-10% or 2-5% (mol) of the total lipids present in the lipid nanoparticles. The molar ratios of ionizable lipids, non-cationic lipids, sterols, and PEG / conjugated lipids may vary as needed. For example, lipid particles may contain 30-70% ionizable lipids per mole or total weight of the composition, 0-60% cholesterol per mole or total weight of the composition, 0-30% non-cationic lipids per mole or total weight of the composition, and 1-10% conjugated lipids per mole or total weight of the composition. Preferably, the composition contains 30-40% ionizable lipids per mole or total weight of the composition, 40-50% cholesterol per mole or total weight of the composition, and 10-20% non-cationic lipids per mole or total weight of the composition. In some embodiments, the composition contains 50-75% ionizable lipid per mole or total weight of the composition, 20-40% cholesterol per mole or total weight of the composition, 5-10% non-cationic lipid per mole or total weight of the composition, and 1-10% conjugated lipid per mole or total weight of the composition. The composition may contain 60-70% ionizable lipid per mole or total weight of the composition, 25-35% cholesterol per mole or total weight of the composition, and 5-10% non-cationic lipid per mole or total weight of the composition. The composition may also contain up to 90% ionizable lipid per mole or total weight of the composition and 2-15% non-cationic lipid per mole or total weight of the composition.Formulations may also be prepared, for example, from 8 to 30% ionizable lipid per mole or total weight of the composition, from 5 to 30% non-cationic lipid per mole or total weight of the composition, and from 0 to 20% cholesterol per mole or total weight of the composition; from 4 to 25% ionizable lipid per mole or total weight of the composition, from 4 to 25% non-cationic lipid per mole or total weight of the composition, from 2 to 25% cholesterol per mole or total weight of the composition, from 10 to 35% conjugated lipid per mole or total weight of the composition, and from 5% cholesterol per mole or total weight of the composition; or from The lipid particle formulation may comprise 2-30% ionizable lipid per mole or total weight of the composition, 2-30% non-cationic lipid per mole or total weight of the composition, 1-15% cholesterol per mole or total weight of the composition, 2-35% conjugated lipid per mole or total weight of the composition, and 1-20% cholesterol per mole or total weight of the composition; alternatively, up to 90% ionizable lipid per mole or total weight of the composition, 2-10% non-cationic lipid per mole or total weight of the composition, or 100% cationic lipid per mole or total weight of the composition. In some embodiments, the lipid particle formulation comprises an ionizable lipid, phospholipid, cholesterol, and PEGylated lipid in a molar ratio of 50:10:38.5:1.5. In some other embodiments, the lipid particle formulation comprises an ionizable lipid, cholesterol, and PEGylated lipid in a molar ratio of 60:38.5:1.5.
[0533] In some embodiments, the lipid particles comprise an ionizable lipid, a non-cationic lipid (e.g., a phospholipid), a sterol (e.g., cholesterol), and a PEGylated lipid, wherein the molar ratio of lipids ranges from 20-70 mole % for the ionizable lipid, with a target of 40-60; the molar percentage of non-cationic lipid ranges from 0-30, with a target of 0-15; the molar percentage of sterol ranges from 20-70, with a target of 30-50; and the molar percentage of PEGylated lipid ranges from 1-6, with a target of 2-5.
[0534] In some embodiments, the lipid particles comprise an ionizable lipid / non-cationic lipid / sterol / conjugated lipid molar ratio of 50:10:38.5:1.5.
[0535] In one aspect, the present disclosure provides a lipid nanoparticle formulation comprising a phospholipid, a lecithin, a phosphatidylcholine, and a phosphatidylethanolamine.
[0536] In some embodiments, one or more additional compounds can be included. Such compounds can be administered separately, or the additional compounds can be included in the lipid nanoparticles of the present invention. In other words, the lipid nanoparticles can contain other compounds in addition to the nucleic acid or at least one second nucleic acid different from the first nucleic acid. Without limitation, the other additional compounds can be selected from the following: small or large organic or inorganic molecules, monosaccharides, disaccharides, trisaccharides, oligosaccharides, polysaccharides, peptides, proteins, peptide analogs and derivatives thereof, peptidomimetics, nucleic acids, nucleic acid analogs and derivatives thereof, extracts prepared from biological materials, or any combination thereof.
[0537] In some embodiments, the addition of a targeting domain directs LNP to a specific tissue. For example, biological ligands can be displayed on the surface of LNP to enhance interaction with cells that display cognate receptors, thereby driving binding to the receptor and cargo delivery to tissues where cells express the receptor. In some embodiments, the biological ligand can be a ligand that drives delivery to the liver, for example, LNPs that display GalNAc achieve delivery of nucleic acid cargo to hepatocytes that display asialoglycoprotein receptor (ASGPR). Akinc et al. Mol Ther 18(7):1357-1364 2010) teaches the conjugation of trivalent GalNAc ligands (GalNAc-PEG-DSG) to PEG-lipids to obtain LNPs that are dependent on ASGPR for observable LNP cargo effects (see, for example, Figure 6).Other ligand-displaying LNP formulations, including, for example, folate, transferrin, or antibodies, are described in WO 2017223135 (incorporated herein by reference), as well as in the references used therein, i.e., Kolhatkar et al., Curr Drug Discov Technol. 2011 8:197-206; Musacchio and Torchilin, Front Biosci. 2011 16:1388-1412; Yu et al., Mol Membr Biol. 2010 27:286-298; Patil et al., Crit Rev Ther Drug Carrier Syst. 2008 25:1-61; Benoit et al., Biomacromolecules. 2011 12:2708-2714; Zhao et al., Expert Opin Drug Deliv. 2008 5:309-319;Akinc et al.,Mol Ther.2010 18:1357-1364;Srinivasan et al.,Methods Mol Biol.2012 820:105-116;Ben-Arie et al.,Methods Mol Biol.2012 757:497-507;Peer 2010 J Control Release.20:63-68; Peer et al.,Proc Natl Acad Sci US A.2007 104:4095-4100;Kim et al.,Methods Mol Biol.2011 721:339-353;Subramanya et al.,Mol Ther.2010 18:2028-2037;Song et al.,Nat Biotechnol. 2005 23:709-717; Peer et al., Science. 2008 319:627-630; and Peer and Lieberman, Gene Ther. 2011 18:1127-1133.
[0538] In some embodiments, LNPs are selected for tissue-specific activity by the addition of Selective ORgan Targeting (SORT) molecules to formulations containing traditional components such as ionizable lipids, amphoteric phospholipids, cholesterol, and poly(ethylene glycol) (PEG) lipids. The teachings of Cheng et al. Nat Nanotechnol 15(4):313-320 (2020) demonstrate that the addition of supplemental "SORT" components can significantly alter in vivo RNA delivery profiles and mediate tissue-specific (e.g., lung, liver, spleen) gene delivery depending on the percentage and biophysical properties of the SORT molecules.
[0539] In some embodiments, the LNPs comprise a biodegradable, ionizable lipid. In some embodiments, the LNPs comprise (9Z,l2Z)-3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate, also referred to as (9Z,l2Z)-3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate. See, e.g., the lipids described in WO 2019 / 067992, WO 2017 / 173054, WO 2015 / 095340, and WO 2014 / 136086, and the references therein. In some embodiments, with respect to LNP lipids, the terms cationic and ionizable are interchangeable, e.g., ionizable lipids are cationic at some pH.
[0540] In some embodiments, multiple components of a Gene Writer system may be prepared in a single LNP formulation; for example, an LNP formulation may include an mRNA encoding a Gene Writer polypeptide and an RNA template. The ratio of nucleic acid components may be varied to maximize therapeutic properties. In some embodiments, the molar ratio of RNA template to mRNA encoding a Gene Writer polypeptide is about 1:1 to 100:1, e.g., about 1:1 to 20:1, about 20:1 to 40:1, about 40:1 to 60:1, about 60:1 to 80:1, or about 80:1 to 100:1. In other embodiments, multiple nucleic acid systems may be prepared in separate formulations, e.g., one LNP formulation containing a template RNA and a second LNP formulation containing an mRNA encoding a Gene Writer polypeptide. In some embodiments, a system may include three or more nucleic acid components formulated into an LNP. In some embodiments, a system may include a protein, e.g., a Gene Writer system, and a template RNA, formulated into at least one LNP.
[0541] In some embodiments, the mean LNP diameter of an LNP formulation can be tens to hundreds of nanometers, as measured, for example, by dynamic light scattering (DLS). In some embodiments, the mean LNP diameter of an LNP formulation can be about 40 nm to about 150 nm, e.g., about 40 nm, 45 nm, 50 nm, 55 nm, 60 nm, 65 nm, 70 nm, 75 nm, 80 nm, 85 nm, 90 nm, 95 nm, 100 nm, 105 nm, 110 nm, 115 nm, 120 nm, 125 nm, 130 nm, 135 nm, 140 nm, 145 nm, or 150 nm. In some embodiments, the average LNP diameter of the LNP formulation may be about 50 nm to about 100 nm, about 50 nm to about 90 nm, about 50 nm to about 80 nm, about 50 nm to about 70 nm, about 50 nm to about 60 nm, about 60 nm to about 100 nm, about 60 nm to about 90 nm, about 60 nm to about 80 nm, about 60 nm to about 70 nm, about 70 nm to about 100 nm, about 70 nm to about 90 nm, about 70 nm to about 80 nm, about 80 nm to about 100 nm, about 80 nm to about 90 nm, or about 90 nm to about 100 nm. In some embodiments, the average LNP diameter of the LNP formulation may be about 70 nm to about 100 nm. In some embodiments, the average LNP diameter of the LNP formulation may be about 80 nm. In some embodiments, the average LNP diameter of the LNP formulation may be about 100 nm. In some embodiments, the mean LNP diameter of the LNP formulation is in the range of about 1 mm to about 500 mm, about 5 mm to about 200 mm, about 10 mm to about 100 mm, about 20 mm to about 80 mm, about 25 mm to about 60 mm, about 30 mm to about 55 mm, about 35 mm to about 50 mm, or about 38 mm to about 42 mm.
[0542] LNPs can be relatively homogeneous in some cases. The polydispersity index can be used to indicate the homogeneity of LNPs, e.g., the size distribution of lipid nanoparticles. A low polydispersity index (e.g., less than 0.3) generally indicates a narrow size distribution. LNPs can have a polydispersity index of about 0 to about 0.25, e.g., 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.10, 0.11, 0.12, 0.13, 0.14, 0.15, 0.16, 0.17, 0.18, 0.19, 0.20, 0.21, 0.22, 0.23, 0.24, or 0.25. In some embodiments, the polydispersity index of LNPs can be about 0.10 to about 0.20.
[0543] The zeta potential of LNP can be used to indicate the electrokinetic potential of the composition.In some embodiments, the zeta potential can represent the surface charge of LNP.Highly charged species can undesirably interact with cells, tissues, and other elements in the body, so lipid nanoparticles with relatively low positive or negative charges are generally desirable. In some embodiments, the zeta potential of the LNP may be about -10 mV to about +20 mV, about -10 mV to about +15 mV, about -10 mV to about +10 mV, about -10 mV to about +5 mV, about -10 mV to about 0 mV, about -10 mV to about -5 mV, about -5 mV to about +20 mV, about -5 mV to about +15 mV, about -5 mV to about +10 mV, about -5 mV to about +5 mV, about -5 mV to about 0 mV, about 0 mV to about +20 mV, about 0 mV to about +15 mV, about 0 mV to about +10 mV, about 0 mV to about +5 mV, about +5 mV to about 20 mV, about +5 mV to about 15 mV, or about +5 mV to about 10 mV.
[0544] The efficiency of encapsulation of a protein and / or nucleic acid, e.g., a Gene Writer polypeptide or mRNA encoding a polypeptide, represents the amount of protein and / or nucleic acid that is encapsulated or associated with the LNP after preparation, compared to the initial amount provided. A high encapsulation efficiency (e.g., close to 100%) is desirable. The encapsulation efficiency can be measured, for example, by comparing the amount of protein or nucleic acid in a solution containing lipid nanoparticles before and after disintegrating the lipid nanoparticles with one or more solvents or detergents. Anion exchange resins can be used to measure the amount of free protein or nucleic acid (e.g., RNA) in solution. Fluorescence can be used to measure the amount of free protein or nucleic acid (e.g., RNA) in solution. For the lipid nanoparticles described herein, the encapsulation efficiency of proteins and / or nucleic acids can be at least 50%, e.g., 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the encapsulation efficiency can be at least 80%. In some embodiments, the encapsulation efficiency can be at least 90%. In some embodiments, the encapsulation efficiency can be at least 95%.
[0545] The LNPs may optionally include one or more coatings. In some embodiments, the LNPs may be formulated into capsules, films, or tablets having a coating. The capsules, films, or tablets containing the compositions described herein may have any useful size, tensile strength, hardness, or density.
[0546] Further exemplary lipids, formulations, methods, and characterization of LNPs are taught in WO2020061457, which is incorporated herein by reference in its entirety.
[0547] In some embodiments, in vitro or ex vivo cell lipofection is performed using Lipofectamine MessengerMax (Thermo Fisher) or TransIT-mRNA Transfection Reagent (Mirus Bio). In certain embodiments, LNPs are formulated using GenVoy_ILM ionizable lipid mix (Precision NanoSystems). In certain embodiments, LNPs are formulated using 2,2-dilinoleyl-4-dimethylaminoethyl-[1,3]-dioxolane (DLin-KC2-DMA) or dilinoleylmethyl-4-dimethylaminobutyrate (DLin-MC3-DMA or MC3), the formulation and in vivo use of which is taught in Jayaraman et al. Angew Chem Int Ed Engl 51(34):8529-8533 (2012), the entire contents of which are incorporated herein by reference.
[0548] LNP formulations optimized for delivery of CRISPR-Cas systems, e.g., Cas9-gRNA RNP, gRNA, Cas9 mRNA, are described in WO2019067992 and WO2019067910, both of which are incorporated herein by reference.
[0549] Another specific LNP formulation useful for delivery of nucleic acids is described in U.S. Pat. No. 8,158,601 and U.S. Pat. No. 8,168,775 (both of which are incorporated herein by reference), including the formulation used for patisiran, which is sold under the name ONPATTRO.
[0550] Exemplary doses of Gene Writer LNPs can include about 0.1, 0.25, 0.3, 0.5, 1, 2, 3, 4, 5, 6, 8, 10, or 100 mg / kg (RNA). Exemplary doses of AAVs containing nucleic acids encoding one or more components of the system can include about 10 11 , 10 12, 10 13 , and 10 14 The MOI can be given in vg / kg.
[0551] All publications, patent applications, patents, and other publications and references (e.g., sequence database reference numbers) cited herein are incorporated by reference in their entirety. For example, GenBank, Unigene, and Entrez referenced herein, including in any table herein, are all incorporated by reference. Unless otherwise indicated, sequence accession numbers listed herein, including in all tables herein, refer to the most recent database entries as of July 19, 2019. When a gene or protein has multiple sequence accession numbers, all of these sequence variants are encompassed. [Example]
[0552] The present invention is further described by the following examples, which are provided for illustrative purposes and should not be construed as limiting the scope or content of the invention.
[0553] Example 1: Delivery of the Gene Writer™ System into mammalian cells This example describes the Gene Writer™ genome editing system delivered to mammalian cells for the site-specific insertion of exogenous DNA into the mammalian cell genome.
[0554] In this example, the polypeptide component of the Gene Writer™ system is a recombinase selected from Table 1, column 1, and the template DNA component is a plasmid DNA containing a target recombination site, e.g., as listed in the corresponding row of Table 1.
[0555] HEK293T cells are transfected with the following test substances: 1. Scrambled DNA Control 2. DNA encoding the aforementioned polypeptide 3. The aforementioned template DNA 4. Combination of 2 and 3
[0556] After transfection, HEK293T cells are cultured for at least 4 days and assayed for site-specific genome editing. Genomic DNA is isolated from each group of HEK293 cells. PCR is performed using primers flanking the appropriate genomic locus selected from column 4 of Table 1. The PCR products are run on an agarose gel to measure the length of the amplified DNA.
[0557] PCR products of the expected length, indicating a successful Gene Writing™ genome editing event that inserted the DNA plasmid template into the target genome, are only observed in cells transfected with the complete Gene Writer™ System in Group 4 above.
[0558] Example 2: Targeted delivery of gene expression units into mammalian cells using the Gene Writer™ system This example describes the construction and use of the Gene Writer genome editor to insert heterologous gene expression units into mammalian genomes.
[0559] In this example, the recombinase protein is selected from Table 1, column 1. The recombinase protein targets the corresponding genomic locus listed in column 4 of Table 1 for DNA integration. The template RNA component is a plasmid DNA comprising a target recombination site and a gene expression unit. The gene expression unit comprises at least one regulatory sequence operably linked to at least one coding sequence. In this example, the regulatory sequence comprises a CMV promoter and enhancer, an enhanced translation element, and a WPRE. The coding sequence is a GFP open reading frame.
[0560] HEK293 cells are transfected with the following test substances: 1. Scrambled DNA Control 2. DNA encoding the aforementioned polypeptide 3. The aforementioned template DNA 4. Combination of 2 and 3
[0561] After transfection, HEK293 cells are cultured for at least 4 days and assayed for site-specific GeneWriting genome editing. Genomic DNA is isolated from HEK293 cells, and PCR is performed using primers flanking the target integration site in the genome. The PCR product is run on an agarose gel to measure the DNA length. PCR products of the expected length are detected in cells transfected with Group 4 test material (complete GeneWriter™ system), indicating a successful GeneWriting™ genome editing event.
[0562] The transfected cells are cultured for an additional 10 days, and after multiple cell culture passages, GFP expression is assayed by flow cytometry. The percentage of GFP-positive cells is calculated from each cell population. GFP-positive cells are detected in the population of HEK293 cells transfected with the test substance of Group 4, which demonstrates that the gene expression units added to the mammalian cell genome by GeneWriting genome editing are expressed.
[0563] Example 3: Targeted delivery of splice acceptor units into mammalian cells using the Gene Writer™ system. This example describes the construction and use of the Gene Writing genome editing system to add heterologous sequences to intronic regions to act as splice acceptors for upstream exons.
[0564] Splicing of a new exon containing a splice acceptor site at the 5' end and a polyA tail at the 3' end into the first intron results in a mature mRNA containing the first native exon of the native locus spliced to the new exon.
[0565] In this example, the recombinase protein is selected from Table 1, column 1. The recombinase protein targets the corresponding genomic locus listed in Table 1, column 4 for DNA integration. The template DNA encodes GFP with a splice acceptor site immediately 5' to the first amino acid of mature GFP (start codon removed) and a 3' poly-A tail downstream of the stop codon.
[0566] HEK293 cells are transfected with the following test substances: 1. Scrambled DNA Control 2. DNA encoding the aforementioned polypeptide 3. The aforementioned template DNA 4. Combination of 2 and 3
[0567] After transfection, HEK293 cells are cultured for at least four days to assay for site-specific GeneWriting genome editing and proper mRNA processing. Genomic DNA is isolated from HEK293 cells. Reverse transcription-PCR is performed to measure mature mRNA containing the first native exon of the target locus and the new exon. RT-PCR reactions are performed using a forward primer that binds to the first native exon of the target locus and a reverse primer that binds to GFP. RT-PCR products are run on an agarose gel to measure DNA length. PCR products of the expected length, indicating a successful GeneWriting genome editing event, are detected in cells transfected with the test substance from Group 4. This result demonstrates that the GeneWriting genome editing system can add a heterologous sequence encoding a gene to an intron region to act as a splice acceptor for an upstream exon.
[0568] The transfected cells are further cultured for 10 days, and after multiple cell culture passages, GFP expression is assayed by flow cytometry. The percentage of GFP-positive cells is calculated from each cell population. GFP-positive cells are detected in the population of HEK293T cells transfected with the test substance of Group 4, which demonstrates that the gene expression units added to the mammalian cell genome via GeneWriting genome editing are expressed.
[0569] Example 4: Specificity of Gene Writing in Mammalian Cells This example describes the Gene Writer™ genome system delivered to mammalian cells for site-specific insertion of exogenous DNA into the mammalian cell genome, and measurement of the specificity of the site-specific insertion.
[0570] In this example, GeneWriting is constructed in HEK293T cells as described in any of the preceding examples. After transfection, HEK293T cells are cultured for at least 4 days before being assayed for site-specific genome editing. Linear amplification PCR is performed as described in Schmidt et al. Nature Methods 4, 1051-1057 (2007) using a forward primer specific to the template DNA that amplifies adjacent genomic DNA. The amplified PCR products are then sequenced using next-generation sequencing technology on a MiSeq instrument. The MiSeq reads are mapped to the HEK293T genome to identify the integration site within the genome.
[0571] The percentage of...
Claims
1. below: a) a recombinase polypeptide comprising the amino acid sequence of SEQ ID NO: 1241, or an amino acid sequence having at least 95%, 96%, 97%, 98%, or 99% identity thereto, or a nucleic acid encoding said recombinase polypeptide; and b) Below: (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), wherein the DNA recognition sequence has the nucleotide sequence of SEQ ID NO: 27 or 634, or a nucleotide sequence having at least 90% or 95% identity thereto, or having no more than 1, 2, or 3 sequence alterations thereto; (ii) a heterologous sequence of interest; double-stranded insert DNA comprising A system for modifying DNA, comprising:
2. A eukaryotic cell comprising: a recombinase polypeptide comprising the amino acid sequence of SEQ ID NO: 1241, or an amino acid sequence having at least 95%, 96%, 97%, 98%, or 99% identity thereto; or a nucleic acid encoding said recombinase polypeptide.
3. below: (i) a DNA recognition sequence, the DNA recognition sequence comprising a first parapalindromic sequence and a second parapalindromic sequence; the DNA recognition sequence having the nucleotide sequence of SEQ ID NO: 27 or 634, or a nucleotide sequence having at least 90% or 95% identity thereto, or having no more than 1, 2, or 3 sequence modifications thereto; (ii) a heterologous sequence of interest; Eukaryotic cells containing
4. 1. A composition for use in modifying the genome of a eukaryotic cell, comprising a recombinase polypeptide or a nucleic acid encoding said recombinase polypeptide and an insert DNA, wherein said recombinase polypeptide or said nucleic acid encoding said recombinase polypeptide and said insert DNA are contacted with said cell, said composition comprising: a) the recombinase polypeptide comprises the amino acid sequence of SEQ ID NO: 1241, or a sequence having at least 95%, 96%, 97%, 98%, or 99% identity thereto; and b) the insert DNA is: (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), the DNA recognition sequence having a first parapalindromic sequence and a second parapalindromic sequence, the DNA recognition sequence having the nucleotide sequence of SEQ ID NO: 27 or 634, or a nucleotide sequence having at least 90% or 95% identity thereto, or having no more than 1, 2, or 3 sequence modifications thereto; (ii) a heterologous sequence of interest; Including, composition.
5. Use of a recombinase polypeptide or a nucleic acid encoding said recombinase polypeptide and an inserted DNA for the manufacture of a medicament for modifying the genome of a eukaryotic cell, characterized in that said medicament is contacted with said cell, comprising: a) the recombinase polypeptide comprises the amino acid sequence of SEQ ID NO: 1241, or a sequence having at least 95%, 96%, 97%, 98%, or 99% identity thereto; and b) the insert DNA is: (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), wherein the DNA recognition sequence has the nucleotide sequence of SEQ ID NO: 27 or 634, or a nucleotide sequence having at least 90% or 95% identity thereto, or having no more than 1, 2, or 3 sequence alterations thereto; (ii) a heterologous sequence of interest; Including, use.
6. 1. A composition for use in inserting a heterologous sequence of interest into the genome of a eukaryotic cell, comprising a recombinase polypeptide or a nucleic acid encoding said recombinase polypeptide and an insert DNA, characterized in that said composition is contacted with said cell: a) the recombinase polypeptide comprises the amino acid sequence of SEQ ID NO: 1241, or a sequence having at least 95%, 96%, 97%, 98%, or 99% identity thereto; and b) the insert DNA is: (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), wherein the DNA recognition sequence has the nucleotide sequence of SEQ ID NO: 27 or 634, or a nucleotide sequence having at least 90% or 95% identity thereto, or having no more than 1, 2, or 3 sequence alterations thereto; (ii) a heterologous sequence of interest; Including, composition.
7. Use of a recombinase polypeptide or a nucleic acid encoding said recombinase polypeptide and an insert DNA for the manufacture of an agent for inserting a heterologous sequence of interest into the genome of a eukaryotic cell, characterized in that said agent is contacted with said cell, comprising: a) the recombinase polypeptide comprises the amino acid sequence of SEQ ID NO: 1241, or a sequence having at least 95%, 96%, 97%, 98%, or 99% identity thereto; and b) the insert DNA is: (i) a DNA recognition sequence that binds to the recombinase polypeptide of (a), wherein the DNA recognition sequence has the nucleotide sequence of SEQ ID NO: 27 or 634, or a nucleotide sequence having at least 90% or 95% identity thereto, or having no more than 1, 2, or 3 sequence alterations thereto; (ii) a heterologous sequence of interest; Including, use.
8. An isolated recombinase polypeptide comprising the amino acid sequence of SEQ ID NO: 1241, or a sequence having at least 95%, 96%, 97%, 98%, or 99% identity thereto.
9. An isolated nucleic acid encoding a recombinase polypeptide comprising the amino acid sequence of SEQ ID NO: 1241, or an amino acid sequence having at least 95%, 96%, 97%, 98%, or 99% identity thereto.
10. below: (i) a DNA recognition sequence, the DNA recognition sequence comprising a first parapalindromic sequence and a second parapalindromic sequence, the DNA recognition sequence having the nucleotide sequence of SEQ ID NO: 27 or 634, or a nucleotide sequence having at least 90% or 95% identity thereto, or having no more than 1, 2, or 3 sequence modifications thereto; (ii) a heterologous sequence of interest; An isolated nucleic acid comprising:
11. 1. A method for producing a recombinase polypeptide ex vivo, said method comprising: a) providing a nucleic acid encoding a recombinase polypeptide comprising the amino acid sequence of SEQ ID NO: 1241, or a sequence having at least 95%, 96%, 97%, 98%, or 99% identity thereto; and b) introducing said nucleic acid into a eukaryotic cell under conditions that allow the production of said recombinase polypeptide; This provides a method for producing the recombinase polypeptide.
12. 1. An ex vivo method for generating an insert DNA comprising a DNA recognition sequence and a heterologous sequence, comprising: a) Below: (i) a DNA recognition sequence that binds to a recombinase polypeptide comprising the amino acid sequence of SEQ ID NO: 1241, or a sequence having at least 95%, 96%, 97%, 98%, or 99% identity thereto, wherein the DNA recognition sequence has the nucleotide sequence of SEQ ID NO: 27 or 634, or a nucleotide sequence having at least 90% or 95% identity thereto, or having no more than 1, 2, or 3 sequence modifications thereto; (ii) a heterologous sequence of interest; providing a nucleic acid comprising: b) introducing the nucleic acid ex vivo into a eukaryotic cell under conditions that allow replication of the nucleic acid; This results in a method for producing the insert DNA.
13. An isolated eukaryotic cell, which is a human cell, comprising a heterologous sequence of interest stably integrated into its genome at genomic location chr5:110266294-110266326.
14. 7. A composition for use according to claim 4 or 6, wherein the use comprises inserting the heterologous sequence of interest into the genome of the eukaryotic cells at a frequency of at least 0.1% of the population of the eukaryotic cells.
15. (i) the DNA recognition sequence is within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides of the heterologous sequence of interest; and / or (ii) the DNA recognition sequence and the heterologous sequence of interest are extrachromosomal; The system of claim 1 .
16. 2. The system of claim 1, wherein (a) the recombinase polypeptide and (b) the insert DNA are in separate containers or are mixed.
17. 10. A composition for use in modifying the genome of a eukaryotic cell, the composition comprising the system of claim 1.
18. The system of claim 1 , wherein the recombinase polypeptide comprises a nuclear localization sequence.
19. The system of claim 1 , wherein the DNA recognition sequence has the sequence set forth in SEQ ID NO: 27 or SEQ ID NO:
634.
20. The system of claim 1 , wherein the nucleic acid encoding the recombinase polypeptide is mRNA.
21. The system of claim 20 , wherein the mRNA is in an LNP.
22. 18. The composition of claim 17, wherein (a) and (b) are administered separately or together.
23. The system of claim 1 , wherein the nucleic acid encoding the recombinase polypeptide of (a) and the insert DNA of (b) are located on separate nucleic acid molecules.
Citation Information
Patent Citations
Enzymes, cells and methods for site specific recombination at asymmetric sites
WO2005081632A2