PAH adjustment system and method

A recombinant system with a reverse transcriptase and Cas9 nuclease-guided template RNA corrects PAH gene mutations, addressing the inefficiencies of current methods and providing a promising treatment for PKU by enhancing PAH activity.

JP2026509460APending Publication Date: 2026-03-19FLAGSHIP PIONEERING INNOVATIONS VI LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-14
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Current methods for inserting target nucleic acids into the genome lack site specificity and efficiency, particularly for correcting mutations in the PAH gene associated with phenylketonuria (PKU), and existing treatments like CRISPR/Cas9 and Cre/loxP are inadequate for long sequence insertion or require multiple steps.

Method used

A recombinant system comprising a nucleic acid encoding a recombinant polypeptide with a reverse transcriptase domain and Cas9 nuclease activity, guided by a template RNA with a gRNA spacer, is used to target and correct mutations in the PAH gene, enabling precise insertion, deletion, or substitution of sequences.

Benefits of technology

This system allows for precise genomic modifications, potentially restoring PAH function and reducing phenylalanine levels in PKU patients, offering a more effective treatment than existing therapies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026509460000163
    Figure 2026509460000163
  • Figure 2026509460000164
    Figure 2026509460000164
  • Figure 2026509460000165
    Figure 2026509460000165
Patent Text Reader

Abstract

This disclosure provides compositions, systems, and methods for targeting, editing, modifying, or manipulating the genome of a host cell at one or more locations in a DNA sequence within a cell, tissue, or object. A genetically modified system for treating phenylketonuria (PKU) is described. In one embodiment, this disclosure relates to a pharmaceutical composition comprising the above system and a pharmaceutically acceptable excipient or carrier, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Sequence List This application is filed electronically in XML format in accordance with WIPO standard ST.26, and includes a sequence listing which is incorporated herein by reference in its entirety. The XML copy prepared on 12 March 2024 is named V2065-7048WO_SL.xml and has a size of 8,922,351 bytes.

[0002] Cross-reference of related applications This application claims the benefits of U.S. Provisional Patent Application No. 63 / 490,442, filed March 15, 2023, U.S. Provisional Patent Application No. 63 / 491,490, filed March 21, 2023, and U.S. Provisional Patent Application No. 63 / 600,472, filed November 17, 2023. The contents of the above applications are incorporated herein by reference in their entirety. [Background technology]

[0003] The insertion of target nucleic acids into the genome occurs infrequently and with little site specificity, in the absence of specialized proteins to facilitate the insertion event. Some existing methods, such as CRISPR / Cas9, are better suited to small edits dependent on host repair pathways and are less effective for inserting long sequences. Other existing methods, such as Cre / loxP, require a first step of inserting a loxP site into the genome, followed by a second step of inserting the target sequence into the loxP site. In this art, there is a need for improved compositions (e.g., proteins and nucleic acids) and methods for inserting, modifying, or deleting target sequences within the genome.

[0004] PKU is a genetic disorder characterized by an autosomal recessive congenital metabolic disorder caused by a deficiency in the hepatic enzyme PAH. PAH catalyzes the hydroxylation of phenylalanine to tyrosine, the rate-limiting step in phenylalanine metabolism. This reaction depends on tetrahydrobiopterin (BH4), molecular oxygen, and iron as cofactors. Loss-of-function mutations in one or both copies of the PAH gene result in a non-functional or low-efficiency enzyme. This ultimately leads to a severe PKU phenotype, in which case there is a decrease in plasma tyrosine levels and phenylalanine can accumulate in the blood to toxic concentrations. Furthermore, the deficiency interferes with the normal synthesis of downstream products, including dopamine, norepinephrine, and melanin.

[0005] The PAH genome sequence and its adjacent regions span approximately 171 kb and contain 13 exons. Studies of pathogenic allele polymorphisms have identified over 500 disease-causing mutations in the PAH gene (Mitchell, et al. Genet Med. 2011;13:697-707). Approximately 62% of these mutations are characterized as missense, 13% as deletions, 11% as splices, 6% as silent, 5% as nonsense, 2% as insertions, and <1% as deletions or exon duplications. The identification of several PAH mutations has been described in terms of their effects on enzyme activity using enzyme kinetics and crystallographic studies. Mutations affecting catalytic binding mode have been observed, including Y138F, S23A, and Y377F, which are associated with a reduced tendency towards tetramerization (Flydal, et al. PNAS. 2019;116(23):11229-34). Other residues that interact with BH4 in the precatalytic conformation (amino acids 245-255, 286, 322, and 325) also interact with BH4 in the catalytic conformation, and furthermore, these sites are actually involved in the severe destabilization of PAH.

[0006] Spontaneous N-terminal PAH mutations have been identified as being distributed in a non-random pattern, clustering within residues 46–48 (GAL motif) and 65–69 (IESRP motif (SEQ ID NO: 37634)), both motifs being highly conserved in pyruvate dehydrogenase (PDH) (Gjetting, et al. Am. / J. Hum. Genet. 2001;68:1353-60). Structural-functional studies have demonstrated that mutations in these regions dramatically reduce phenylalanine binding. Most missense mutations identified in PKUs to date result in phenotypes associated with misfolding of the PAH enzyme, increased protein turnover, and loss of enzyme function. Residues in exons 7–9 and interdomain regions within subunits appear to play a crucial structural role and become hotspots for destabilization. Furthermore, recombinant hPAHs showed residual activity from mutations in the BH4-reactive domain, including R408W and Y414C, but exhibited disrupted allosteric effects, suggesting altered protein conformation (Gersting, et al. Hum. Genet. 2008;83:5-17). Mutation and structure-function analyses identified robust genotype-phenotype mappings regarding the role of PAHs in PKU, but no successful treatments exist other than lifelong symptom management strategies. Dietary therapy with phenylalanine (Phe) remains central to the treatment of PKU since its introduction in 1953. In the 1970s, combination therapy with tetrahydrobiopterin (BH4) and neurotransmitter precursors (L-dopa / carbidopa and 5-hydroxytryptophan) showed promising results in modulating PKU. Since its establishment as a treatment, synthesis of sapropterin and other small molecule isomers of BH4 has been incorporated. However, this form of therapy is generally only useful in patients with a mild subset of PAH-deficient PKU. Treatment response is thought to be related to mutations in the PAH gene that result in some residual enzyme activity. At high blood concentrations, Phe in the blood will compete with other large neutral amino acids (LNAAs) for transport across the blood-brain barrier. LNAA supplementation has been shown to decrease brain Phe concentrations despite the observation of increased plasma Phe levels. Similarly, dietary supplementation with glycomacropeptides (GMP) has been observed to significantly reduce urea production and improve protein retention and Phe utilization. However, these methods have little effect on addressing increased blood concentrations of Phe or genotype-driven substances. [Prior art documents] [Non-patent literature]

[0007] [Non-Patent Document 1] Mitchell,et al.Genet Med.2011;13:697-707 [Non-Patent Document 2] Flydal,et al.PNAS.2019;116(23):11229-34 [Non-Patent Document 3] Gjetting,et al.Am. / J.Hum.Genet.2001;68:1353-60 [Non-Patent Document 4] Gersting,et al.Hum.Genet.2008;83:5-17 [Overview of the Initiative] [Problems that the invention aims to solve]

[0008] Modern non-dietary approaches include the development of PAH-based fusion protein and enzyme replacement therapy. Enzyme replacement therapy may involve administering phenylalanine ammonia-lyase (PAL) to patients. PAL is an enzyme that catalyzes the conversion of Phe to trans-cinnamic acid and small amounts of ammonia. Early studies using PAL administered to PKU patients in enteric-coated gelatin capsules showed a decrease in Phe levels, but repeated in vivo administration resulted in an increased immune response. However, these approaches are not practical from a clinical standpoint because they require several intravenous injections due to the limited half-life of the enzyme in the blood. Gene therapy has shown promising results in promoting PAH functionality, for example, using viral vectors. However, the effectiveness of this approach is hindered by very low gene transfer rates and transient transgene expression. Therefore, novel and more effective treatments are needed to target PAH in PKU. [Means for solving the problem]

[0009] This disclosure relates to novel compositions, systems, and methods for in vivo or in vitro modification of the genome at one or more locations within a host cell, tissue, or subject. This disclosure provides recombinant systems capable of modulating phenylalanine hydroxylase (PAH) activity (e.g., inserting, modifying, or deleting a sequence of interest) and methods for treating phenylketonuria (PKU) by administering one or more such systems to modify the genome sequence to correct mutations in the PAH gene on human chromosome 12q23.2 that are involved as genetic drivers in PKU.

[0010] In one embodiment, the present disclosure relates to a system for recombining DNA to correct a human PAH gene mutation that causes PKU, comprising: (a) a nucleic acid encoding a recombinant polypeptide capable of targeted priming reverse transcription, wherein the polypeptide comprises (i) a reverse transcriptase domain and (ii) Cas9 niccas, which binds to DNA and has endonuclease activity; and (b) a template RNA comprising (i) a gRNA spacer complementary to the first portion of the human PAH gene, (ii) a gRNA scaffold that binds to the polypeptide, (iii) a heterologous target sequence containing a mutation region for correcting the mutation, and (iv) a primer-binding site (PBS) sequence at the 3' end of the template RNA containing at least 3, 4, 5, 6, 7, or 8 bases of 100% homology to the target DNA strand. In some embodiments, the PAH gene may contain the R408W mutation. The template RNA sequence may include sequences described herein, such as those in Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3.

[0011] The gRNA spacer may contain at least 15 bases with 100% homology to the target DNA at the 5' end of the template RNA. The template RNA may further contain a PBS sequence containing at least 5 bases with at least 80% homology to the target DNA strand. The template RNA may contain one or more chemical modifications.

[0012] The domains of a recombinant polypeptide may be linked by a peptide linker. A polypeptide may contain one or more peptide linkers. A recombinant polypeptide may further contain nuclear localization signals. A polypeptide may contain two or more nuclear localization signals, for example, multiple adjacent nuclear localization signals, or one or more nuclear localization signals in different regions of the polypeptide, for example, one or more nuclear localization signals at the N-terminus of the polypeptide and one or more nuclear localization signals at the C-terminus of the polypeptide. The nucleic acid encoding the recombinant polypeptide may encode one or more intein domains.

[0013] The introduction of the system into target cells may result in insertions of at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, or 1000 base pairs into exogenous DNA. The introduction of the system into target cells may result in deletions of less than 2, 3, 4, 5, 10, 50, or 100 base pairs into the genomic DNA upstream or downstream of the insertion. The introduction of the system into target cells may result in substitutions, such as substitutions of 1, 2, or 3 nucleotides, such as substitutions of consecutive nucleotides.

[0014] The heterologous target sequence may consist of at least 5, 10, 25, 50, 100, 150, 200, 250, 300, 400, 500, 600, or 700 base pairs.

[0015] In one embodiment, the disclosure relates to a pharmaceutical composition comprising the above system and a pharmaceutically acceptable excipient or carrier, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles. In one embodiment, the disclosure relates to a pharmaceutical composition comprising the above system and a plurality of pharmaceutically acceptable excipients or carriers, wherein the pharmaceutically acceptable excipients or carriers are selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles, for example, the above system is delivered by two different excipients or carriers, for example, two lipid nanoparticles and two viral vectors or one lipid nanoparticle and one viral vector. The viral vector may be an adeno-associated virus (AAV).

[0016] In one embodiment, the present disclosure relates to a host cell (e.g., a mammalian cell, e.g., a human cell) including the above-described system.

[0017] In one embodiment, the disclosure relates to a method for correcting mutations in the human PAH gene within a cell, tissue, or subject, comprising administering the above system to the cell, tissue, or subject, wherein, optionally, the correction of the mutated PAH gene comprises an amino acid substitution of W408R (replacing the pathogenic substitution R408W). The system may be introduced in vivo, in vitro, ex vivo, or in situ. The nucleic acid of (a) may be incorporated into the genome of a host cell. In some embodiments, the nucleic acid of (a) is not incorporated into the genome of a host cell. In some embodiments, the heterologous target sequence is inserted into only one target site in the host cell genome. The heterologous target sequence may be inserted into two or more target sites in the host cell genome, for example, the same corresponding site on two homologous chromosomes or two different sites on the same or different chromosomes. The heterologous target sequence may encode a mammalian polypeptide or a fragment or variant thereof. The components of the system may be delivered on one, two, three, four or more different nucleic acid molecules. The system can be introduced into host cells by electroporation or by using at least one vehicle selected from plasmid vectors, viral vectors, vesicles, and lipid nanoparticles.

[0018] The composition or method may include one or more of the embodiments listed below.

[0019] Enumerated embodiments 1. A genetically modified system, (a) template RNA (tgRNA), which has a 5' to 3' position. (1) gRNA spacer, (2) gRNA scaffold, (3) heterogeneous sequences, and (4) Primer binding site (PBS) sequence A template RNA (tgRNA) containing the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, (b) Recombinant polypeptide or nucleic acid encoding a recombinant polypeptide, wherein the recombinant polypeptide is (1) Cas domain, (2) Linker, and (3) Reverse transcriptase (RT) domain The recombinant polypeptide includes the amino acid sequence of the recombinant polypeptide of SEQ ID NO: 28 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and the recombinant polypeptide or nucleic acid A genetically modified system that includes this.

[0020] 2. A genetically modified system, (a) template RNA (tgRNA), which has a 5' to 3' position. (1) gRNA spacers having the sequence of the template RNA of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having one, two or three or fewer sequence mutations (e.g., substitutions) therein. (2) gRNA scaffold sequences of template RNAs of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or gRNA scaffolds having sequences with 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 or fewer sequence mutations therein. (3) Heteropurpose sequences of template RNAs of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or sequences having one, two or three or fewer sequence mutations therein, and (4) Primer binding site (PBS) sequences having the PBS sequence of the template RNA of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or sequences having one, two, or three or fewer sequence mutations thereto. template RNA (tgRNA) containing, (b) Recombinant polypeptide or nucleic acid encoding a recombinant polypeptide, wherein the recombinant polypeptide is (1) Cas domain, (2) Linker, and (3) Reverse transcriptase (RT) domain The recombinant polypeptide includes the amino acid sequence of the recombinant polypeptide of SEQ ID NO: 28 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and the recombinant polypeptide or nucleic acid A genetically modified system that includes this.

[0021] 3. The system of Embodiment 1 or 2, wherein the template RNA includes the nucleotide sequence of the template RNA sequence listed in Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having at least 95% identity thereto.

[0022] 4. The system of Embodiment 1 or 2, wherein the template RNA includes the nucleotide sequence of the template RNA sequence listed in Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having at least 97%, 98%, or 99% identity thereto.

[0023] 5. The system of Embodiment 1 or 2, wherein the template RNA includes the nucleotide sequence of the template RNA sequence shown in Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3.

[0024] 6. A system according to any of the preceding embodiments, wherein the template RNA includes the nucleotide sequence of SEQ ID NO: 2 or a sequence having at least 95%, 97%, 98%, or 99% identity thereto.

[0025] 7. The template RNA is a system of any of the previous embodiments, comprising the nucleotide sequence of SEQ ID NO: 2.

[0026] 8. The template RNA is a system according to any of the preceding embodiments, comprising the nucleotide sequence of Sequence ID No. 37638 or a sequence having at least 95%, 97%, 98%, or 99% identity thereto.

[0027] 9. The template RNA is a system of any of the previous embodiments, comprising the nucleotide sequence of SEQ ID NO: 37638.

[0028] 10. A system according to any of the preceding embodiments, wherein the template RNA comprises the nucleotide sequence of SEQ ID NO: 37654 or a sequence having at least 95%, 97%, 98%, or 99% identity thereto.

[0029] 11. The template RNA is a system of any of the previous embodiments, comprising the nucleotide sequence of SEQ ID NO: 37654.

[0030] 12. The template RNA is a system according to any of the preceding embodiments, comprising the nucleotide sequence of SEQ ID NO: 37653 or a sequence having at least 95%, 97%, 98%, or 99% identity thereto.

[0031] 13. The template RNA is a system of any of the previous embodiments, comprising the nucleotide sequence of SEQ ID NO: 37653.

[0032] 14. A system according to any one of Embodiments 1 to 13, wherein the recombinant polypeptide comprises the amino acid sequence of the recombinant polypeptide of SEQ ID NO: 28 or a sequence having at least 95% identity thereto.

[0033] 15. A system according to any one of Embodiments 1 to 13, wherein the recombinant polypeptide comprises the amino acid sequence of the recombinant polypeptide of SEQ ID NO: 28 or a sequence having at least 99% identity thereto.

[0034] 16. A recombinant polypeptide comprising any one of Embodiments 1 to 13, wherein the recombinant polypeptide comprises the amino acid sequence of the recombinant polypeptide of SEQ ID NO: 28.

[0035] 17. A system according to any one of Embodiments 1 to 16, comprising a nucleic acid encoding a genetically modified polypeptide, wherein the nucleic acid encoding the genetically modified polypeptide has the sequence described in Sequence ID No. 106 or a sequence having at least 95%, 97%, 98%, or 99% identity thereto.

[0036] 18. A system according to any one of Embodiments 1 to 16, comprising a nucleic acid encoding a recombinant polypeptide, wherein the nucleic acid encoding the recombinant polypeptide has the sequence described in Sequence ID No. 106.

[0037] 19. Any one of embodiments 1 to 18, further comprising a second nick RNA (ngRNA) that leads a second nick to the second strand of the human PAH gene.

[0038] 20. The system of Embodiment 19, wherein the ngRNA includes the sequences of the ngRNAs in Table 2A, E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14, or E14A, or sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.

[0039] 21.ngRNA is located from 5' to 3'. (1) gRNA spacer sequences of ngRNAs of Table 2A, E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14 or E14A, or gRNA spacers having sequences with one, two or three or fewer sequence mutations (e.g., substitutions), and (2) gRNA scaffold sequences of ngRNAs of Table 2A, E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14 or E14A, or gRNA scaffolds having sequences with 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 or fewer sequence mutations. A system of embodiment 19, including the system of embodiment 19.

[0040] 22. The system of Embodiment 19 or 20, wherein the ngRNA, together with the template RNA of the genetically modified system, has a "PAM inward orientation".

[0041] 23. A system from any one of Embodiments 19 to 22, wherein the ngRNA has the sequence described in Sequence ID No. 23 or a sequence that is at least 95%, 97%, 98%, or 99% identical thereto.

[0042] 24.ngRNA is one of the systems from Embodiments 19 to 22, having the sequence described in Sequence ID No. 23.

[0043] 25. A system from any one of Embodiments 19 to 22, wherein the ngRNA has the sequence described in Sequence ID No. 84 or a sequence that is at least 95%, 97%, 98%, or 99% identical thereto.

[0044] 26.ngRNA is one of the systems from Embodiments 19 to 22, having the sequence described in Sequence ID No. 84.

[0045] 27. A system of any one of Embodiments 1 to 26, wherein the nucleic acid encoding the genetically modified polypeptide includes RNA, for example, mRNA.

[0046] 28. A system of any one of Embodiments 1 to 27, wherein the nucleic acid encoding the genetically modified polypeptide comprises one or more chemically modified nucleotides.

[0047] 29. A nucleic acid molecule encoding a genetically modified polypeptide, for example, from 5' to 3', (a) The 5'UTR of sequences 41-44 or sequences having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, (b) A region encoding the N-terminal NLS of sequence number 36 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (c) A region encoding the Cas domain of sequence number 52 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (d) A region that codes for a linker of sequence number 54 or a sequence that is at least 90%, 95%, 96%, 97%, 98%, or 99% identical thereto. (e) A region encoding the reverse transcriptase (RT) domain of sequence number 56 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (f) A region encoding the C-terminal NLS of sequence number 38 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (g) The 3'UTR of sequences with at least 90%, 95%, 96%, 97%, 98%, or 99% identity with sequence numbers 45-48, (h) optionally, an expression element of sequence number 40, 101, or 102, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and (i) Poly(A)tails of sequences having at least 90%, 95%, 96%, 97%, 98%, or 99% identity with sequence numbers 49-51. Nucleic acid molecules containing these molecules.

[0048] 30. A nucleic acid molecule of Embodiment 29 having the sequence described in Sequence ID No. 106 or a sequence having at least 95%, 97%, 98%, or 99% identity thereto.

[0049] 31. A nucleic acid molecule of Embodiment 29 having the sequence described in Sequence ID No. 106.

[0050] 32. A nucleic acid molecule encoding a genetically modified polypeptide, for example, from 5' to 3', (a) The 5'UTR of sequences 41-44 or sequences having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, (b) A region encoding the N-terminal NLS of sequence number 37 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (c) A region encoding the Cas domain of sequence number 53 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (d) A region that codes for a linker of sequence number 55 or a sequence that is at least 90%, 95%, 96%, 97%, 98%, or 99% identical thereto. (e) A region encoding the reverse transcriptase (RT) domain of sequence number 57 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (f) A region encoding the C-terminal NLS of sequence number 39 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (g) The 3'UTR of sequences with at least 90%, 95%, 96%, 97%, 98%, or 99% identity with sequence numbers 45-48, (h) optionally, an expression element of sequence number 40, 101, or 102, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and (i) Poly(A)tails of sequences having at least 90%, 95%, 96%, 97%, 98%, or 99% identity with sequence numbers 49-51. Nucleic acid molecules containing these molecules.

[0051] 33. Any one nucleic acid molecule from Embodiments 29 to 32, comprising a Cas domain that binds to a target DNA molecule and is heterogeneous to the RT domain.

[0052] 34. A nucleic acid molecule which is mRNA, one of any one of embodiments 29 to 33.

[0053] 35. A nucleic acid molecule comprising one or more chemically modified nucleotides, one of any one of embodiments 29 to 34.

[0054] 36. Template RNA, (1) gRNA spacer, (2) gRNA scaffold, (3) heterogeneous sequences, and (4) Primer binding site (PBS) sequence A template RNA containing the nucleotide sequence of the template RNA of sequence number 15, 17, 92, or 94, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.

[0055] 37. The template RNA is a nucleic acid molecule of Embodiment 36, comprising one or more chemically modified nucleotides.

[0056] 38. A template RNA containing the nucleotide sequence of sequence number 37638 or a sequence having at least 95%, 97%, 98%, or 99% identity thereto.

[0057] 39. Template RNA containing the nucleotide sequence of sequence number 37638.

[0058] 40. A template RNA containing the nucleotide sequence of sequence number 37653 or a sequence having at least 95%, 97%, 98%, or 99% identity thereto.

[0059] 41. Template RNA containing the nucleotide sequence of sequence number 37653.

[0060] 42. A genetically modified system, (a) template RNA (tgRNA), which has a 5' to 3' position. (1) gRNA spacer, (2) gRNA scaffold, (3) heterogeneous sequences, and (4) Primer binding site (PBS) sequence A template RNA (tgRNA) containing the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, (b) nucleic acids (e.g., mRNA) that encode a genetically modified polypeptide, (1) Cas domain, (2) Linker, and (3) Reverse transcriptase (RT) domain The nucleotides comprising the recombinant polypeptide include a nucleic acid (e.g., mRNA) and one of the nucleic acids of Embodiments 29-35. A genetically modified system that includes this.

[0061] 43. A genetically modified system, (a) template RNA (tgRNA), which has a 5' to 3' position. (1) gRNA spacer, (2) gRNA scaffold, (3) heterogeneous sequences, and (4) Primer binding site (PBS) sequence A template RNA (tgRNA) containing the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, (b) nucleic acids (e.g., mRNA) that encode a genetically modified polypeptide, (1) Cas domain, (2) Linker, and (3) Reverse transcriptase (RT) domain The nucleotides encoding the recombinant polypeptide include nucleic acids (e.g., mRNA) and sequences that have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with the recombinant polypeptide sequence of Table N2, E11, or E15. A genetically modified system that includes this.

[0062] 44. A genetically modified system, (a) template RNA (tgRNA), which has a 5' to 3' position. (1) gRNA spacers having the sequence of the template RNA of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having one, two or three or fewer sequence mutations (e.g., substitutions) therein. (2) gRNA scaffold sequences of template RNAs of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or gRNA scaffolds having sequences with 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 or fewer sequence mutations therein. (3) Heteropurpose sequences of template RNAs of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or sequences having one, two or three or fewer sequence mutations therein, and (4) Primer binding site (PBS) sequences having the PBS sequence of the template RNA of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or sequences having one, two, or three or fewer sequence mutations thereto. template RNA (tgRNA) containing, (b) nucleic acids (e.g., mRNA) that encode a genetically modified polypeptide, (1) Cas domain, (2) Linker, and (3) Reverse transcriptase (RT) domain The nucleotides encoding the recombinant polypeptide include nucleic acids (e.g., mRNA) and sequences that have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with the recombinant polypeptide sequence of Table N2, E11, or E15. A genetically modified system that includes this.

[0063] 45. A system from any one of Embodiments 42 to 44, wherein the template RNA includes the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having at least 95% identity thereto.

[0064] 46. ​​A system from any one of Embodiments 42 to 44, wherein the template RNA includes a nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having at least 97%, 98%, or 99% identity thereto.

[0065] 47. A system from any one of Embodiments 42 to 44, wherein the template RNA includes the nucleotide sequence of the template RNA sequence shown in Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3.

[0066] 48. A system from any one of Embodiments 42 to 47, wherein the nucleotide encoding the recombinant polypeptide includes the nucleic acid sequence of the recombinant polypeptide of Table N2, E11, or E15, or a sequence having at least 95% identity thereto.

[0067] 49. A system from any one of Embodiments 42 to 47, wherein the nucleotide encoding the recombinant polypeptide includes the nucleic acid sequence of the recombinant polypeptide of Table N2, E11, or E15, or a sequence having at least 99% identity thereto.

[0068] 50. A system of any one of Embodiments 42 to 47, wherein the nucleotide encoding the recombinant polypeptide comprises the nucleic acid sequence of the recombinant polypeptide of Table N2, E11, or E15.

[0069] 51. Any one of embodiments 42 to 50, further comprising a second nick RNA (ngRNA) that leads a second nick to the second strand of the human PAH gene.

[0070] 52. The system of Embodiment 51, wherein the ngRNA includes the sequences of the ngRNAs in Table 2A, E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14, or E14A, or sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.

[0071] 53.ngRNA is located from 5' to 3'. (1) gRNA spacer sequences of ngRNAs of Table 2A, E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14 or E14A, or gRNA spacers having sequences with one, two or three or fewer sequence mutations (e.g., substitutions), and (2) gRNA scaffold sequences of ngRNAs of Table 2A, E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14 or E14A, or gRNA scaffolds having sequences with 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 or fewer sequence mutations. A system of embodiment 51, including the system of embodiment 51.

[0072] 54. The system of Embodiment 51 or 52, wherein the ngRNA, together with the template RNA of the genetically modified system, has a "PAM inward orientation".

[0073] 55. A system in any one of embodiments 42 to 54, wherein the nucleic acid encoding the genetically modified polypeptide includes RNA, for example, mRNA.

[0074] 56. A system of any one of Embodiments 42 to 55, wherein the nucleic acid encoding the genetically modified polypeptide comprises one or more chemically modified nucleotides.

[0075] 57. The nucleic acid molecule is formulated in lipid nanoparticles (LNPs), and is a nucleic acid molecule from any one of embodiments 29 to 35 or a template RNA from any one of embodiments 36 to 41.

[0076] 58. A system in which tgRNA, a nucleic acid molecule encoding a recombinant polypeptide, and / or ngRNA are formulated in LNP, any one of Embodiments 1-28 or 42-56.

[0077] 59. A pharmaceutical composition comprising one system from any one of embodiments 1 to 28 or 42 to 56, or one or more nucleic acids encoding it, and a pharmaceutically acceptable excipient or carrier.

[0078] 60. The pharmaceutical composition of Embodiment 59, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles (LNPs).

[0079] 61. The pharmaceutical composition of Embodiment 60, wherein the viral vector is an adeno-associated virus.

[0080] 62. A host cell (e.g., a mammalian cell, e.g., a human cell) containing any one of the nucleic acid molecules, genetically modified systems, or template RNAs of the preceding embodiments.

[0081] 63. A method for producing a nucleic acid molecule or template RNA according to any one of the preceding embodiments, comprising synthesizing the nucleic acid molecule or template RNA in vitro (for example, by in vitro transcription or solid-phase synthesis) or by introducing DNA encoding the template RNA into a host cell under conditions that enable the production of the template RNA.

[0082] 64. A method for modifying a target site in a human PAH gene within a cell, comprising contacting a cell with any one of the genetic recombination systems of Embodiments 1 to 28 or 42 to 56, or DNA encoding the same, or any one of the pharmaceutical compositions of Embodiments 59 to 61, thereby modifying a target site in a human PAH gene within a cell.

[0083] 65. A method for treating a subject having a disease or condition associated with a mutation in the human PAH gene, comprising administering to the subject one of the recombinant systems of Embodiments 1-28, 42-56, or 58, or DNA encoding the same, or one of the pharmaceutical compositions of Embodiments 59-61, thereby treating the subject having a disease or condition associated with a mutation in the human PAH gene.

[0084] 66. The method of Embodiment 65, wherein the disease or symptom is phenylketonuria (PKU) or hyperphenylalaninemia (e.g., mild or severe hyperphenylalaninemia).

[0085] 67. The subject is the method of embodiment 65 or 66, which has the R408W mutation.

[0086] 68. A method for treating a subject having a PKU, comprising administering to the subject one of the recombinant systems of Embodiments 1-28, 42-56, or 58, or DNA encoding the same, or one of the pharmaceutical compositions of Embodiments 59-61, thereby treating the subject having a PKU.

[0087] 69. A genetic recombination system or method according to any one of the preceding embodiments, wherein the introduction of the system into target cells results in the modification of a pathogenic mutation in the PAH gene.

[0088] 70. A recombinant system or method of any one of the prior embodiments wherein the pathogenic mutation is an R408W mutation, and the modification involves an amino acid substitution of W408R.

[0089] 71. A genetic recombination system or method according to any one of the preceding embodiments, wherein the introduction of the system into target cells results in a mutation that causes the restoration of PAH gene function.

[0090] 72. A genetic recombination system or method according to any one of the preceding embodiments, wherein mutation correction occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70%, or more) of the target nucleic acid.

[0091] 73. A genetic recombination system or method according to any one of the preceding embodiments, wherein mutation correction occurs in at least 10% of target cells (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70%, or more).

[0092] 74. A genetic recombination system or method according to any one of the preceding embodiments, wherein the recombination system includes a gRNA that targets the second strand, and the correction of mutations in a population of target cells is increased compared to a population of target cells treated with a recombination system that includes a template RNA without a gRNA that targets the second strand.

[0093] 75. The cell is a mammalian cell, such as a human cell, according to any one of the preceding embodiments.

[0094] 76. The subject is a human, and the method is one of the embodiments described above.

[0095] 77. Contact is performed ex vivo, for example, by any one of the preceding embodiments, wherein the cell or DNA of the subject is modified ex vivo.

[0096] 78. Contact is performed in vivo, and for example, the DNA of the cell or subject is modified in vivo, according to one of the preceding embodiments.

[0097] 79. Contacting cells or objects with a system is any one of the methods of the preceding embodiments, which includes contacting cells or cells within an object with nucleic acids (e.g., DNA or RNA) encoding a recombinant polypeptide under conditions that enable the production of the recombinant polypeptide.

[0098] 80. The heterologous target sequence comprises a region of at least five consecutive nucleotides, including 2'-fluoro modifications on alternating nucleotides, as a template RNA or system of any of the preceding embodiments.

[0099] 81. The heterologous target sequence comprises 2'-fluoromodifications on alternating nucleotides, starting at position +4 of the heterologous target sequence, in the template RNA or system of Embodiment 80.

[0100] 82. The heterologous target sequence includes 2'-fluoro modifications on alternating nucleotides at positions +4 to +8, +10, +12, +14, +16, or +18 of the heterologous target sequence, wherein the template RNA or system of Embodiment 81.

[0101] 83. The heterologous target sequence is a template RNA or system of Embodiment 80, wherein the heterologous target sequence includes alternating 2'-fluoro modifications on nucleotides in regions having a length of 2-5, 5-10, 10-15, 15-20, 20-25, 25-30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, or 80-81 nucleotides.

[0102] 84. The heterologous target sequence is a template RNA or system of Embodiment 80, wherein the heterologous target sequence includes alternating 2'-fluoro modifications on nucleotides in regions having a length of 80-90, 90-100, 100-150, 150-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000, 1000-1500, 1500-2000, 2000-2500, 2500-3000, 3000-3500, 3500-4000, 4000-4500, or 4500-5000 nucleotides.

[0103] 85. The heterologous target sequence comprises 2'-fluoromodifications on alternating nucleotides, starting at position +5 of the heterologous target sequence, in the template RNA or system of Embodiment 80.

[0104] 86. The heterologous target sequence includes 2'-fluoro modifications on alternating nucleotides at positions +5 to +9, +11, +13, +15, or +17 of the heterologous target sequence, as a template RNA or system of Embodiment 81.

[0105] 87. A template RNA or system according to any of embodiments 80 to 86, comprising a 2'-fluoromodified nucleotide at the 5' end of a heterologous target sequence (for example, at position +10 or +11 of the heterologous target sequence).

[0106] 88. A template RNA or system of Embodiment 80, comprising 2'-fluoromodified nucleotides at positions +4, +6, +8 and / or +10 of a heterogeneous target sequence.

[0107] 89. A template RNA or system of Embodiment 80, comprising 2'-fluoromodified nucleotides at positions +5, +7, +9 and / or +11 of a heterogeneous target sequence.

[0108] 90. The template RNA or system of Embodiment 80, wherein the second nucleotide from the 5' end of the heterologous target sequence includes a 2'-fluoro modification (e.g., at position +9 or +10 of the heterologous target sequence).

[0109] 91. A template RNA or system of Embodiment 80 comprising 2'-fluoromodified nucleotides at positions +5, +7 and / or +9 of a heterogeneous target sequence.

[0110] 92. A template RNA or system of Embodiment 80, comprising 2'-fluoromodified nucleotides at positions +4, +6, +8 and / or +10 of a heterogeneous target sequence. [Brief explanation of the drawing]

[0111] [Figure 1]This graph shows Phe levels in the plasma of mice treated with an exemplary genetically modified system. [Figure 2A] This graph shows the rewriting level % for each tested template RNA that has a second nick corresponding to the spacer. [Figure 2B] This graph shows the indel activity % for each of the tested template RNAs that have a second nick corresponding to the spacer. [Figure 3A] This graph shows the rewriting activity in primary mouse hepatocytes administered with the indicated LNP, which contains the indicated doses of template RNA, a second nick guide, and the RNAIVT338 recombinant polypeptide encoding mRNA, as measured by Amp-SEQ. [Figure 3B] This graph shows the percentage indel activity in primary mouse hepatocytes administered with the indicated LNP containing the RNAIVT338 recombinant polypeptide encoding the template RNA, a second nick guide, and mRNA at the indicated doses. [Figure 4] (Figure 4A) This graph shows the rewriting activity in primary mouse hepatocytes nucleofected with the indicated doses of template RNA, a second nick guide, and RNAIVT338 recombinant polypeptide, as measured by Amp-SEQ. (Figure 4B) This graph shows the indel activity % in primary mouse hepatocytes nucleofected with the indicated doses of template RNA, a second nick guide, and RNAIVT338 recombinant polypeptide. [Figure 5] This is a graph of Phe levels in plasma obtained from mice treated with a recombinant system containing the indicated template RNA and a second nick guide. [Figure 6] (Figure 6A) This graph shows the rewriting level % obtained by treating with a genetically modified system containing the template RNA and a second nick guide. (Figure 6B) This graph shows the indel activity % obtained by treating hPAH mice with a genetically modified system containing the template RNA and a second nick guide. [Figure 7A] This graph shows the rewriting level % for each of the tested template RNAs in mouse liver, specifically those without a second nick guide (left) and those with a second nick guide (right). [Figure 7B] This graph shows the indel activity percentage for each of the tested template RNAs in mouse liver, specifically those without a second nick guide (left) and those with a second nick guide (right). [Figure 8] (Figure 8A) This graph shows the rewriting level % obtained using a genetically modified system containing RNACS4134 and RNACS1810 along with various mRNAs encoding genetically modified polypeptides. (Figure 8B) This graph shows the indel activity % obtained using a genetically modified system containing various mRNAs encoding genetically modified polypeptides. [Figure 9A] This is a pair of graphs showing the rewriting level % for a recombination system containing RNACS4134 template RNA with either the second nick guide of RNACS1809 or RNACS1810, and various mRNAs encoding recombinant polypeptides. [Figure 9B] This is a pair of graphs showing the indel activity % for a recombinant system containing RNACS4134 template RNA having either the second nick guide of RNACS1809 or RNACS1810, and various mRNAs encoding recombinant polypeptides. [Figure 10A] The graphs show the Phe levels in plasma obtained from treated mice when treated with a recombinant system containing RNACS4134 template RNA having either RNACS1809 or RNACS1810's second nick guide, and various mRNAs encoding recombinant polypeptides, at the indicated dosage. [Figure 10B]This is a pair of graphs showing brain Phe levels obtained from mice treated with a recombinant system containing RNACS4134 template RNA having either RNACS1809 or RNACS1810's second nick guide, and various mRNAs encoding recombinant polypeptides, at the indicated dosage. [Figure 11] The genetic recombination system described herein is shown. The figure on the left shows a recombinant polypeptide containing a Cas nickase domain (e.g., spCas9 N863A) and a reverse transcriptase domain (RT domain) linked by a linker. The figure on the right shows template RNA containing a gRNA spacer, gRNA scaffold, heterologous target sequence, and primer-binding site sequence (PBS sequence) from 5' to 3'. The heterologous target sequence may contain a mutant region with one or more sequence differences compared to the target site. The heterologous target sequence may also contain pre-edit homology and post-edit homology regions adjacent to the mutant region. While not intended to be bound by any particular theory, it is thought that the gRNA spacer of the template RNA binds to the second strand of the target site in the genome, and the gRNA scaffold of the template RNA binds to the recombinant polypeptide, for example, to localize the recombinant polypeptide to the target site in the genome. The Cas domain of the recombinant polypeptide is thought to cleave the target site (e.g., the first strand of the target site) and bind, for example, a PBS sequence to a sequence adjacent to the site to be modified on the first strand of the target site. The RT domain of the recombinant polypeptide is thought to polymerize, for example, a sequence complementary to the heterologous target sequence, using the first strand of the target site bound to a complementary sequence containing the PBS sequence of the template RNA as a primer and the heterologous target sequence of the template RNA as a template. While not intended to be bound to any particular theory, it is thought that reverse transcription can then proceed through pre-edit homology regions, then mutation regions, and then post-edit homology regions to generate a DNA strand containing mutations identified by the heterologous target sequence. [Figure 12]This is a series of figures showing heterologous target sequences and PBS (priming) sequences of a series of template RNA variants, each containing the 2'-fluoro modification shown in the table at the top of the gray box. Two of the variants contained alternating patterns of 2'-fluoro modifications, with every other nucleotide in the partial sequence of the heterologous target sequence containing the 2'-fluoro modification. These variants further contained, in 5' to 3' order, a 2'-fluoro modified nucleotide, three 2'-OMe modified nucleotides, and three nucleotides at the 3' end of the priming region, each containing 2'-OMe and phosphorothioate modifications, respectively. These template RNA variants were tested for their ability to introduce alterations to the target nucleic acid sequence, and the rewriting efficiency and percentage of insertions or deletions (indels) obtained for each variant are shown in the graphs at the bottom left and bottom right, respectively. This figure discloses sequence numbers 37647 to 37649, in order of appearance. [Figure 13] This is a series of figures showing heterologous target sequences and PBS (priming) sequences of a series of template RNA variants, each containing the 2'-fluoro modification shown in the table at the top of the gray box. Two of the variants contained alternating patterns of 2'-fluoro modifications, with every other nucleotide in the partial sequence of the heterologous target sequence containing the 2'-fluoro modification. These variants further contained, in 5' to 3' order, a 2'-fluoro modified nucleotide, three 2'-OMe modified nucleotides, and three nucleotides at the 3' end of the priming region, each containing a 2'-OMe and / or phosphorothioate modification. These template RNA variants were tested for their ability to introduce alterations to the target nucleic acid sequence, and the resulting indel rewriting efficiency and percentage for each variant are shown in the graphs at the bottom left and bottom right, respectively. This figure discloses sequence numbers 37647 to 37649, in order of appearance. [Figure 14]This is a series of figures showing heterologous target sequences and PBS (priming) sequences of template RNA variants, each containing the 2'-fluoro modification shown in the table at the top of the gray box. The RNACS6874 variant contained an alternating pattern of 2'-fluoro modifications, with every other nucleotide in the partial sequence of the heterologous target sequence containing a 2'-fluoro modification. This variant further contained, in 5' to 3' order, a 2'-fluoro modified nucleotide, three 2'-OMe modified nucleotides, and three nucleotides at the 3' end of the priming region, each containing a 2'-OMe and a phosphorothioate modification, respectively. These template RNA variants were tested for their ability to introduce alterations to the target nucleic acid sequence, and the resulting indel rewriting efficiency and percentage for each variant are shown in the graphs at the bottom left and bottom right, respectively. This figure discloses sequence numbers 37647 to 37648, in order of appearance. [Figure 15] This is a series of figures showing heterologous target sequences and PBS (priming) sequences of a series of template RNA variants, each containing the 2'-fluoro modification shown in the table at the top of the gray box. Two of the variants contained alternating patterns of 2'-fluoro modifications, with every other nucleotide in a subsequence of the heterologous target sequence containing a 2'-fluoro modification. These variants further contained three nucleotides, each containing a 2'-fluoro modified nucleotide, a 2'-OMe modified nucleotide, and a 2'-OMe and / or phosphorothioate modification at the 3' end of the priming region, in the order of 5' to 3'. These template RNA variants were tested for their ability to introduce alterations to the target nucleic acid sequence, and the resulting indel rewriting efficiency and percentage for each variant are shown in the graphs at the bottom left and bottom right, respectively. This figure discloses sequence numbers 37650 to 37652, in order of appearance. [Figure 16]This is a series of figures showing heterologous target sequences and PBS (priming) sequences of a series of template RNA variants, each containing the 2'-fluoro modification shown in the table at the top of the gray box. Two of the variants contained alternating patterns of 2'-fluoro modifications, with every other nucleotide in a subsequence of the heterologous target sequence containing a 2'-fluoro modification. These variants further contained three nucleotides, each containing a 2'-fluoro modified nucleotide, a 2'-OMe modified nucleotide, and a 2'-OMe and / or phosphorothioate modification at the 3' end of the priming region, in the order of 5' to 3'. These template RNA variants were tested for their ability to introduce alterations to the target nucleic acid sequence, and the resulting indel rewriting efficiency and percentage for each variant are shown in the graphs at the bottom left and bottom right, respectively. This figure discloses sequence numbers 37650 to 37652, in order of appearance. [Figure 17] A pair of graphs showing the rewriting level %(17A) or indel %(17B) for a recombination system containing RNACS6874 template RNA along with various mRNAs encoding recombinant polypeptides, where the mRNA may or may not contain different forms of WPRE. [Modes for carrying out the invention]

[0112] definition When used herein in relation to chemical modification, the term “alternating nucleotides” refers to a nucleotide pattern in which all odd nucleotides in a region have the same chemical modification and all even nucleotides do not have that modification, or vice versa, or all even nucleotides in a region have the same chemical modification and all odd nucleotides do not have that modification. For example, in a 5-nucleotide length region having alternating nucleotides, the 1st, 3rd, and 5th positions of the region may all contain 2'F chemical modifications, and the 2nd and 4th positions of the region may contain unmodified nucleotides or chemical modifications other than 2'F. The 2nd and 4th positions may be the same or different. Furthermore, one or more nucleotides that all contain the same chemical modification may contain a second chemical modification. As a non-limiting example, in the above region having 2'F chemical modifications at the 1st, 3rd, and 5th positions, if only one of the nucleotides at those positions further contains a skeletal modification, the region still contains alternating nucleotides with respect to 2'F chemical modifications. Alternating nucleotides may be found in larger nucleic acid regions, and larger nucleic acids may contain one or more other non-alternating regions.

[0113] As used herein, the term "expression cassette" refers to a nucleic acid construct comprising sufficient nucleic acid elements for the expression of the nucleic acid molecule of the present invention.

[0114] As used herein, "gRNA spacer" refers to a portion of a nucleic acid that is complementary to the target nucleic acid and, together with the gRNA scaffold, can target the Cas protein to the target nucleic acid.

[0115] As used herein, "gRNA scaffold" refers to a portion of a nucleic acid that can bind to a Cas protein and, together with a gRNA spacer, can target the Cas protein to a target nucleic acid. In some embodiments, the gRNA scaffold includes a crRNA sequence, a tetraloop, and a tracrRNA sequence.

[0116] When used herein, “recombinant polypeptide” refers to a polypeptide comprising a retroviral reverse transcriptase or a polypeptide comprising an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity to a retroviral reverse transcriptase capable of incorporating a nucleic acid sequence (e.g., a sequence provided on a template nucleic acid) into a target DNA molecule (e.g., within a mammalian host cell, such as a genomic DNA molecule in a host cell). In some embodiments, the recombinant polypeptide is capable of incorporating the sequence substantially independently of the host mechanism. In some embodiments, the recombinant polypeptide incorporates the sequence at a random location in the genome, and in some embodiments, the recombinant polypeptide incorporates the sequence at a specific target site. In some embodiments, the recombinant polypeptide collectively comprises one or more domains that 1) facilitate binding to the template nucleic acid, 2) facilitate binding to the target DNA molecule, and 3) facilitate the incorporation of at least a portion of the template nucleic acid into the target DNA. The recombinant polypeptide includes both the natural polypeptide and its engineered variants having, for example, one or more amino acid substitutions to the natural sequence. Recombinant polypeptides include heterologous constructs, for example, one or more of the domains listed above are heterologous to each other, whether otherwise by heterologous fusion (or other conjugate) of wild-type domains and fusion of modifying domains, for example by substitution or fusion of heterologous subdomains or other substitution domains. Exemplary recombinant polypeptides that can be used in the manner provided herein, systems comprising them, and methods of using them are described, for example, in PCT / US2021 / 020948, which is incorporated herein by reference with respect to a recombinant polypeptide comprising a retroviral reverse transcriptase domain. In some embodiments, the recombinant polypeptide incorporates the sequence into a gene. In some embodiments, the recombinant polypeptide incorporates the sequence into a sequence outside of a gene. When used herein, “recombinant system” refers to a system comprising a recombinant polypeptide and a template nucleic acid.

[0117] As used herein, the term “domain” refers to a structure of a biomolecule that contributes to a particular function of that biomolecule. A domain may include a continuous region (e.g., a continuous sequence) or a separate discontinuous region (e.g., a discontinuous sequence) of a biomolecule. Examples of protein domains include, but are not limited to, endonuclease domains, DNA-binding domains, and reverse transcription domains, while examples of nucleic acid domains include regulatory domains, such as transcription factor-binding domains. In some embodiments, a domain (e.g., a Cas domain) may include two or more smaller domains (e.g., a DNA-binding domain and an endonuclease domain).

[0118] As used herein, the term “exomorphic,” when used in relation to biomolecules (e.g., nucleic acid sequences or polypeptides), means that the biomolecule has been artificially introduced into a host genome, cell, or organism. For example, nucleic acids added to an existing genome, cell, tissue, or subject using recombinant DNA technology or other methods are exomorphic to the existing nucleic acid sequence, cell, tissue, or subject.

[0119] The terms “first strand” and “second strand,” used herein to describe individual DNA strands of target DNA, identify two DNA strands on which the reverse transcriptase domain initiates polymerization, for example, based on the location where targeted primed synthesis begins. “First strand” refers to the strand of target DNA on which the reverse transcriptase domain initiates polymerization, for example, initiating targeted primed synthesis. “Second strand” refers to the other strand of target DNA. The designations “first strand” and “second strand” do not otherwise describe the target site DNA strands. For example, in some embodiments, the first and second strands are nicked by the polypeptide described herein, but the designations “first” and “second” strands are independent of the order in which such nicks appear.

[0120] The term “heterogeneous,” when used to describe a first element in relation to a second element, means that the first and second elements do not naturally exist in the configuration described. For example, heterogeneous polypeptides, nucleic acid molecules, constructs, or sequences refer to (a) a polypeptide, nucleic acid molecule, or portion of a polypeptide or nucleic acid molecule sequence that is not native to the cell in which it is expressed, (b) a polypeptide or nucleic acid molecule, or portion of a polypeptide or nucleic acid molecule, that has been modified or mutated from its native state, or (c) a polypeptide or nucleic acid molecule having modified expression compared to its native expression level under similar conditions. For example, heterogeneous regulatory sequences (e.g., promoters, enhancers) can be used to regulate the expression of a gene or nucleic acid molecule in a manner different from that which is normally expressed in nature. In another example, a heterogeneous domain or nucleic acid sequence of a polypeptide (e.g., the DNA-binding domain of a polypeptide or the nucleic acid encoding the DNA-binding domain of a polypeptide) may be configured in relation to other domains, or may have a different sequence or origin from a different source compared to other domains or portions of the polypeptide or its encoding nucleic acid. In certain embodiments, heterologous nucleic acid molecules may be present in the native host cell genome, but may have modified expression levels, different sequences, or both. In other embodiments, heterologous nucleic acid molecules may not be endogenous in the host cell or host genome, but may instead be introduced into the host cell by transformation (e.g., transfection, electroporation), and the added molecules may be integrated into the host genome or exist as extrachromosomal genetic material, either transiently (e.g., mRNA) or semi-stable for two or more generations (e.g., episomal viral vectors, plasmids, or other self-replicating vectors).

[0121] As used herein, “insertion” of a sequence into a target site refers to the net addition of a DNA sequence at the target site, such as a new nucleotide in a heterologous target sequence that does not have a cognitive position at the non-edited target site. In some embodiments, nucleotide alignment of the PBS sequence and the heterologous target sequence with respect to the target nucleic acid sequence will result in an alignment gap in the target nucleic acid sequence.

[0122] As used herein, the “deletion” generated by the heterologous target sequence at the target site refers to a net deletion of the DNA sequence at the target site, such as a nucleotide at a non-edited target site that does not have a cognitive position in the heterologous target sequence. In some embodiments, the nucleotide alignment of the PBS sequence and the heterologous target sequence to the target nucleic acid sequence will result in an alignment gap in the molecule containing the PBS sequence and the heterologous target sequence.

[0123] As used herein, the term “inverted terminal repeat” or “ITR” refers to an AAV viral cis-element named as such due to its symmetry. This element facilitates the effective augmentation of the AAV genome. The minimum elements for ITR function are assumed to be a Rep-binding site (RBS, 5'-GCGCGCTCGCTCGCTC-3' for AAV2, SEQ ID NO: 4601) and a terminal separation site (TRS, 5'-AGTTGG-3' for AAV2), as well as a variable palindromic sequence that enables hairpin formation. According to the present invention, an ITR comprises at least three of these elements (RBS, TRS, and the sequence that enables hairpin formation). In addition, in the present invention, the term “ITR” refers to ITRs of known native AAV serotypes (e.g., ITRs of serotypes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 AAV), chimeric ITRs formed by the fusion of ITR elements from different serotypes, and their functional variants. A "functional variant" refers to a sequence that exhibits at least 80%, 85%, or 90%, preferably at least 95%, sequence identity with a known ITR, and that allows for an increase in the sequence containing the ITR in the presence of the Rep protein.

[0124] As used herein, the term "mutant region" refers to a region in template RNA that has one or more sequence differences compared to the corresponding sequence in the target nucleic acid. Sequence differences may include, for example, substitutions, insertions, frameshifts, or deletions.

[0125] The term "(sudden) mutation," when applied to nucleic acid sequences, means that nucleotides within the nucleic acid sequence are inserted, deleted, or altered compared to a reference (e.g., natural) nucleic acid sequence. A single alteration may occur at one locus (point mutation), or multiple nucleotides may be inserted, deleted, or altered at a single locus. In addition, one or more alterations may occur at any number of loci within a single nucleic acid sequence. Nucleic acid sequences can be mutated by any method known in the art.

[0126] "Nucleic acid molecules" refer to, but are not limited to, RNA and DNA molecules, including complementary DNA ("cDNA"), genomic DNA ("gDNA"), and messenger RNA ("mRNA"), and also include synthetic nucleic acid molecules, such as those chemically synthesized or produced by recombination, as described herein, like RNA templates. Nucleic acid molecules may be double-stranded or single-stranded, cyclic, or linear. In the case of single-stranded nucleic acid molecules, the nucleic acid molecule may be a sense strand or an antisense strand. Unless otherwise stated, and as an example of all sequences described herein in the general format "Sequence ID," a nucleic acid including "Sequence ID 1" means a nucleic acid in which at least a portion has either (i) the sequence of Sequence ID 1, or (ii) a sequence complementary to Sequence ID 1. The choice between the two depends on the context in which Sequence ID 1 is used. For example, when a nucleic acid is used as a probe, the choice between the two depends on the requirement that the probe is complementary to the desired target. The nucleic acid sequences of this disclosure may be chemically or biochemically modified, or may contain unnatural or derivatized nucleotide bases, as will be readily apparent to those skilled in the art. Such modifications include, for example, labeling, methylation, substitution of one or more spontaneously occurring nucleotides by analogs, internucleotide modifications such as uncharged bonds (e.g., methylphosphonic acid, triestriates, phosphoramidates, carbamates, etc.), charged bonds (e.g., phosphorothioates, phosphorodithioates, etc.), pendant moieties (e.g., polypeptides), insertants (e.g., acridine, psoralens, etc.), chelating agents, alkylating agents, and modifying bonds (e.g., α-anomeric nucleic acids, etc.). Chemically modified bases (e.g., see Table 13), backbones (e.g., see Table 14), and modified caps (e.g., see Table 15). Synthetic molecules that mimic polynucleotides in terms of their ability to bind to specified sequences by hydrogen bonding and other chemical interactions are also included. Such molecules are known in the art, and examples include those that use peptide bonds instead of phosphate bonds in the molecular backbone, such as peptide nucleic acids (PNAs).Other modifications include analogues that include other structures such as modifications found in the ribose ring, crosslinking portion, or "locked" nucleic acid (LNA). In various embodiments, nucleic acids are associated with additional genetic elements, such as tissue-specific expression-regulatory sequences (e.g., tissue-specific promoters and tissue-specific microRNA recognition sequences) and additional elements, such as inverted repeats (e.g., inverted terminal repeats, e.g., viral elements (e.g., AAV ITR)) and tandem repeats, inverted repeats / direct repeats, homologous regions (segments with different degrees of homology to target DNA), untranslated regions (UTRs) (5', 3', or both 5' and 3' UTRs) and various combinations thereof. The nucleic acid elements of the systems provided by the present invention may be provided in various topologies, including single-stranded, double-stranded, circular, linear, open-ended linear, closed-ended linear and specific versions thereof, e.g., doggybone DNA (dbDNA), closed-ended DNA (ceDNA).

[0127] As used herein, “gene expression unit” is a nucleic acid sequence comprising at least one regulatory nucleic acid sequence operably ligated to at least one effector sequence. The first nucleic acid sequence is operably ligated to the second nucleic acid sequence when the first nucleic acid sequence is positioned functionally in relation to the second nucleic acid sequence. For example, a promoter or enhancer is operably ligated to a coding sequence if the promoter or enhancer affects the transcription or expression of the coding sequence. The operably ligated DNA sequences may be continuous or discontinuous. If it is necessary to ligate two protein coding regions, the operably ligated sequences may reside within the same reading frame.

[0128] The terms “host genome” or “host cell,” as used herein, refer to the cell into which the protein and / or genetic material has been introduced and / or its genome. These terms refer not only to a specific target cell and / or genome, but also to the offspring of such cells and / or the genomes of such offspring. It should be understood that such offspring may not be identical to the parent cell in fact, as certain modifications may occur in later generations due to mutation or environmental influences, but are still included in the scope of the term “host cell” as used herein. A host genome or host cell may be an isolated cell or a cell line grown in culture or genomic material isolated from such cells or cell lines, or it may be a host cell or host genome constituting a living tissue or organism. In some cases, the host cell may be an animal cell or a plant cell, as described herein, for example. In certain examples, the host cell may be a mammalian cell, a human cell, a bird cell, a reptile cell, a bovine cell, a horse cell, a pig cell, a goat cell, a sheep cell, a chicken cell, or a turkey cell. In certain cases, the host cells may be maize cells, soybean cells, wheat cells, or rice cells.

[0129] As used herein, “action-related” describes a functional relationship between two nucleic acid sequences, for example, 1) a promoter and 2) a heterologous target sequence, where the promoter and the heterologous target sequence (e.g., the gene of interest) are oriented such that, under appropriate conditions, the promoter drives the expression of the heterologous target sequence. For example, a template nucleic acid having a promoter and a heterologous target sequence may be, for example, a single strand with a (+) or (-) orientation. The “action-related” between the promoter and the heterologous target sequence in this template means that, regardless of whether the template nucleic acid is to be transcribed in a particular state, it will be transcribed accurately if it is in an appropriate state (e.g., with a (+) orientation in the presence of the required catalyst and NTPs, etc.). Action-related applies similarly to other pairs of nucleic acids, including other tissue-specific regulatory sequences (e.g., enhancers, repressors, and microRNA recognition sequences), IR / DR, ITR, UTR, or sequences encoding homologous regions and heterologous target sequences, or retroviral RT domains.

[0130] The terms “primer binding site sequence” or “PBS sequence,” as used herein, refer to a portion of template RNA capable of binding to a region contained in a target nucleic acid sequence. In some cases, the PBS sequence is a nucleic acid sequence containing at least 3, 4, 5, 6, 7, or 8 bases that are 100% identical to a region contained in the target nucleic acid sequence. In some embodiments, the primer region contains at least 5, 6, 7, or 8 bases that are 100% identical to a region contained in the target nucleic acid sequence. While not intended to be bound by any particular theory, in some embodiments, if the template RNA contains both a PBS sequence and a heterologous target sequence, the PBS sequence binds to a region contained in the target nucleic acid sequence, enabling the reverse transcriptase domain to use that region as a primer for reverse transcription and the heterologous target sequence as a template for reverse transcription.

[0131] As used herein, “stem-loop sequence” refers to a nucleic acid sequence (e.g., RNA sequence) having a stem containing at least 2 (e.g., 3, 4, 5, 6, 7, 8, 9, or 10) base pairs and a loop containing at least 3 (e.g., 4) base pairs, where self-complementarity is sufficient to form a stem-loop. The stem may contain mismatches or bulges.

[0132] As used herein, “tissue-specific expression-regulatory sequence” means a nucleic acid element that increases or decreases the level of a transcript containing a heterologous target sequence in a tissue-specific manner within a target tissue, for example, preferentially in an on-target tissue compared to an off-target tissue. In some embodiments, the tissue-specific expression-regulatory sequence preferentially drives or represses the transcription, activity, or half-life of a transcript containing the heterologous target sequence in a tissue-specific manner within a target tissue, for example, preferentially in an on-target tissue compared to an off-target tissue. Exemplary tissue-specific expression-regulatory sequences include tissue-specific promoters, repressors, enhancers or combinations thereof, and tissue-specific microRNA recognition sequences. Tissue specificity refers to on-target (tissues in which expression or activity of the template nucleic acid is desired or acceptable) and off-target (tissues in which expression or activity of the template nucleic acid is undesirable or unacceptable). For example, a tissue-specific promoter preferentially drives expression in an on-target tissue compared to an off-target tissue. In contrast, microRNAs that bind to tissue-specific microRNA recognition sequences are preferentially expressed in off-target tissues compared to on-target tissues, thereby reducing the expression of template nucleic acids in off-target tissues. Therefore, promoters and microRNA recognition sequences specific to the same tissue, such as a target tissue, have contrasting functions with respect to the transcription, activity, or half-life of the related sequences within the tissue (matching expression levels, i.e., promoting and suppressing high levels of microRNA in off-target tissues and low levels in on-target tissues, respectively, while promoters drive high expression in on-target tissues and low expression in off-target tissues).

[0133] Unless otherwise specified, the following numbering system is used to describe the positions of nucleotides in chemically modified and / or heterologous target sequences of template RNA in PBS. Nucleotide positions in heterologous target sequences are numbered +1, +2, +3, etc., starting from the outermost 3'- end of the heterologous target sequence. Nucleotide positions in PBS sequences are numbered -1, -2, -3, etc., starting from the outermost 5'- end of the PBS sequence. Therefore, positions +1 and -1 are directly adjacent to each other.

[0134] introduction This disclosure relates to methods for treating phenylketonuria (PKU) and compositions for targeting, editing, recombining, or manipulating DNA sequences at one or more locations within a DNA sequence in a cell, tissue, or subject, for example, in vivo or in vitro (e.g., inserting a heterologous target sequence into a target site in a mammalian genome). The heterologous target DNA sequence may include, for example, substitutions.

[0135] More specifically, the Disclosure provides a method for treating PKUs using a reverse transcriptase-based system for modifying a target genomic DNA sequence, for example, by inserting, deleting, or substituting one or more nucleotides into / from a target sequence.

[0136] This disclosure provides a method for treating PKUs using a recombination system comprising a recombinant polypeptide component and a template nucleic acid (e.g., template RNA) component. In some embodiments, the recombination system can be used to introduce modifications into target sites within the genome. In some embodiments, the recombinant polypeptide component includes a writing domain (e.g., reverse transcriptase domain), a DNA-binding domain, and an endonuclease domain (e.g., nickase domain). In some embodiments, the template nucleic acid (e.g., template RNA) includes a sequence that binds to a target site within the genome (e.g., a gRNA spacer), a sequence that binds to the recombinant polypeptide component (e.g., a gRNA scaffold), a heterologous target sequence, and a PBS sequence. While not intended to be constrained by theory, it is assumed that the template nucleic acid (e.g., template RNA) binds to the second strand of the target site within the genome and then binds to the recombinant polypeptide component (e.g., localizes the polypeptide component to the target site within the genome). The endonuclease (e.g., nickase) of the recombinant polypeptide component is thought to cleave the target site (e.g., the first strand of the target site) and bind to a sequence adjacent to the site where the PBS sequence will be modified on the first strand of the target site. The writing domain (e.g., reverse transcriptase domain) of the polypeptide component is thought to polymerize a sequence complementary to the heterologous target sequence using the first strand of the target site bound to a complementary sequence containing the PBS sequence of the template nucleic acid as a primer and the heterologous target sequence of the template nucleic acid as a template. While we do not wish to be constrained by theory, the selection of an appropriate heterologous target sequence is thought to result in the substitution, deletion, and / or insertion of one or more nucleotides at the target site.

[0137] Genetic engineering systems In some embodiments, the recombinant system described herein comprises (A) a recombinant polypeptide or a nucleic acid encoding a recombinant polypeptide, the recombinant polypeptide comprising (i) a reverse transcriptase domain and an endonuclease domain including DNA-binding functionality, and (B) a template RNA. In some embodiments, the recombinant polypeptide acts as a substantially autonomous protein mechanism capable of incorporating a template nucleic acid sequence into a target DNA molecule (e.g., within a mammalian host cell, such as a genomic DNA molecule in a host cell) substantially independently of host mechanisms. For example, the recombinant protein may comprise a DNA-binding domain, a reverse transcriptase domain, and an endonuclease domain. In some embodiments, the DNA-binding functionality may comprise an RNA component that guides the protein to the DNA sequence, such as a gRNA spacer. In other embodiments, the recombinant polypeptide may comprise a reverse transcriptase domain and an endonuclease domain. The RNA template element of the recombinant system is typically heterologous to the recombinant polypeptide element and provides a target sequence to be inserted (reverse transcribed) into the host genome. In some embodiments, the recombinant polypeptide is capable of targeted priming reverse transcription. In some embodiments, the recombinant polypeptide can undergo second chain synthesis.

[0138] Functional recombinant polypeptides can consist of unrelated DNA-binding, reverse transcription, and endonuclease domains. This modular structure allows for combinations of functional domains, e.g., dCas9 (DNA binding), AVIRE reverse transcriptase (reverse transcription), and FokI (endonuclease). In some embodiments, multiple functional domains may arise from a single protein, e.g., Cas9 or Cas9 nickase (DNA binding, endonuclease).

[0139] In some embodiments, the recombinant polypeptide collectively comprises one or more domains that 1) facilitate binding to a template nucleic acid, 2) facilitate binding to a target DNA molecule, and 3) facilitate the incorporation of at least a portion of the template nucleic acid into the target DNA. In some embodiments, the recombinant polypeptide is a modified polypeptide comprising one or more amino acid substitutions relative to the corresponding native sequence. In some embodiments, the recombinant polypeptide comprises two or more domains that are heterogeneous to each other, for example, by heterologous fusion (or other conjugate) of the otherwise wild-type domain and fusion of modified domains, for example by substitution or fusion of heterologous subdomains or other substitution domains. For example, in some embodiments, one or more of the following are the case: the RT domain is heterologous to the DBD, the DBD is heterologous to the endonuclease domain, or the RT domain is heterologous to the endonuclease domain.

[0140] In some embodiments, the template RNA molecule for use in the system includes, from 5' to 3', (1) a gRNA spacer, (2) a gRNA scaffold, (3) a heterologous target sequence, and (4) a primer binding site (PBS) sequence. In some embodiments, (1) The gRNA spacer is approximately 18-22 nt long, for example, 20 nt. (2) A gRNA scaffold comprising one or more hairpin loops, e.g., one, two, or three loops for associating a template with a Cas domain, e.g., a nickase Cas9 domain. In some embodiments, the gRNA scaffold comprises the sequence GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 21) from 5' to 3'. (3) In some embodiments, the heterogeneous target array has a length of, for example, 7 to 74, for example, 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70 or 70 to 80 nt or 80 to 90 nt. (4) In some embodiments, the PBS sequence that binds to the target priming sequence after nicking is, for example, 3-20 nt, 7-15 nt, or 12-14 nt. In some embodiments, the PBS sequence has a GC content of 40-60%.

[0141] In some embodiments, a second gRNA associated with the system may assist in driving complete integration. In some embodiments, the second gRNA may target positions 0–200 nt away from the first strand nick, for example, 0–50, 50–100, or 100–200 nt away from the first strand nick. In some embodiments, the second gRNA may bind only to its target sequence after editing, for example, the gRNA binds to a sequence that is present in the heterologous target sequence but not in the initial target sequence.

[0142] In some embodiments, editing is performed on HEK293, K562, U2OS, or HeLa cells using the genetic recombination system described herein. In some embodiments, editing is performed on primary cells, such as human primary cells, using the genetic recombination system.

[0143] In some embodiments, the recombinant polypeptides described herein include a reverse transcriptase or RT domain (e.g., as described herein) containing an AVIRE RT sequence or a variant thereof. In some embodiments, the endonuclease domain (e.g., as described herein) includes a Cas9 domain (e.g., SpCas9 containing an N863A mutation (e.g., spCas9)).

[0144] In some embodiments, the heterologous target sequences (for example, in the systems described herein) have nucleotide lengths of approximately 1-50, 50-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000 or more.

[0145] In some embodiments, the RT and endonuclease domains are linked by a flexible linker containing, for example, the amino acid sequence AEAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAKEAAAKEAAAKA (SEQ ID NO: 54).

[0146] In some embodiments, the endonuclease domain is the N-terminus of the RT domain. In some embodiments, the endonuclease domain is the C-terminus of the RT domain.

[0147] In some embodiments, the system incorporates a heterogeneous target sequence into the target site by TPRT, for example, as described herein.

[0148] In some embodiments, the genetically modified system can generate insertions at target sites of at least 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides (and optionally 500 or fewer, 400, 300, 200, or 100 nucleotides). In some embodiments, the genetically modified system can generate insertions at target sites of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides (and optionally 500, 400, 300, 200, or 100 nucleotides). In some embodiments, the genetic recombination system can generate insertions at target sites of at least 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10 kilobases (and optionally 1, 5, 10, or 20 kilobases or less). In some embodiments, the genetic recombination system can generate deletions of at least 81, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides (and optionally 500, 400, 300, or 200 nucleotides or less). In some embodiments, the genetically modified system can generate deletions of at least 81, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides (and optionally 500, 400, 300, or 200 nucleotides or less). In some embodiments, the genetically modified system can generate deletions of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides (and optionally 500, 400, 300, or 200 nucleotides or less).In some embodiments, the genetic recombination system can generate deletions of at least 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10 kilobases (and optionally 1, 5, 10, or 20 kilobases or less). In some embodiments, the genetic recombination system can generate substitutions at target sites of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 or more nucleotides. In some embodiments, the genetic recombination system can generate substitutions at target sites of 1-2, 2-3, 3-4, 4-5, 5-10, 10-15, 15-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, or 90-100 nucleotides.

[0149] In some embodiments, the substitution is a transposition mutation. In some embodiments, the substitution is a transversion mutation. In some embodiments, the substitution converts adenine to thymine, adenine to guanine, adenine to cytosine, guanine to thymine, guanine to cytosine, guanine to adenine, thymine to cytosine, thymine to adenine, thymine to guanine, cytosine to adenine, cytosine to guanine, or cytosine to thymine.

[0150] In some embodiments, insertions, deletions, substitutions, or combinations thereof increase or decrease gene expression (e.g., transcription or translation). In some embodiments, insertions, deletions, substitutions, or combinations thereof increase or decrease gene expression (e.g., transcription or translation) by modifying, adding, or deleting sequences within promoters or enhancers, such as sequences that bind to transcription factors. In some embodiments, insertions, deletions, substitutions, or combinations thereof modify gene translation (e.g., by modifying amino acid sequences), insert or delete start or stop codons, or modify or repair gene translation frames. In some embodiments, insertions, deletions, substitutions, or combinations thereof modify gene splicing, for example, by inserting, deleting, or modifying splice acceptors or donor sites. In some embodiments, insertions, deletions, substitutions, or combinations thereof modify the half-life of transcripts or proteins. In some embodiments, insertions, deletions, substitutions, or combinations thereof alter protein localization in cells (e.g., from cytoplasm to mitochondria, from cytoplasm to extracellular space (e.g., adding secretory tags)). In some embodiments, insertions, deletions, substitutions, or combinations thereof alter protein folding (e.g., improve it) (e.g., to prevent the accumulation of misfolded proteins). In some embodiments, insertions, deletions, substitutions, or combinations thereof alter, increase, or decrease the activity of a gene, for example, a protein encoded by that gene.

[0151] Exemplary recombinant polypeptides, systems comprising them, and methods of using them are described in PCT / US2021 / 020948, incorporated herein by reference, for example, amino acids and nucleic acid sequences therein, with respect to retroviral RT domains.

[0152] Exemplary recombinant polypeptides and retroviral RT domain sequences are also described, for example, in International Patent Application No. PCT / US21 / 20948, filed on March 4, 2021, for example in Tables 30, 31, and 44 therein, the entire application being incorporated herein by reference, for example, with respect to the aforementioned sequences and retroviral RTs in the tables. Accordingly, the recombinant polypeptides described herein may include an amino acid sequence or its domain (e.g., a retroviral RT domain) according to any of the tables described in this paragraph, or a functional fragment or variant thereof, or an amino acid sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity thereto.

[0153] In some embodiments, polypeptides for use in any of the systems described herein may be molecular reconstitutes or ancestral reconstitutes based on aligned polypeptide sequences of multiple homologous proteins. In some embodiments, reverse transcriptase domains for use in any of the systems described herein may be molecular reconstitutes or ancestral reconstitutes, or may be modified with specific residues based on the alignment of reverse transcriptase domains from the same or different sources. Those skilled in the art can align polypeptide or nucleic acid sequences by using common sequence analysis tools, such as the Basic Local Alignment Search Tool (BLAST) or CD-Search for conserved domain analysis, based on the acceptance numbers provided herein. Molecular reconstitutes may be prepared based on sequence consensus using methods, for example, those described in Ivics et al., Cell 1997, 501-510; Wagstaff et al., Molecular Biology and Evolution 2013, 88-99.

[0154] polypeptide components of genetically modified systems In some embodiments, recombinant polypeptides have the functions of DNA target site binding, template nucleic acid (e.g., RNA) binding, DNA target site cleavage, and template nucleic acid (e.g., RNA) writing, such as reverse transcription. In some embodiments, each function is contained within a different domain. In some embodiments, a function may be attributed to two or more domains (e.g., two or more domains exhibit functionality together). In some embodiments, two or more domains may have the same or similar functions (e.g., two or more domains independently have DNA binding functionality, for example, on two different DNA sequences). In other embodiments, one or more domains may perform one or more functions; for example, the Cas9 domain may be capable of both DNA binding and target site cleavage. In some embodiments, all domains are located within a single polypeptide. In some embodiments, the first domain is present in the first polypeptide, and the second domain is present in the second polypeptide. For example, in some embodiments, the sequence may be split between a first polypeptide and a second polypeptide, for example, the first polypeptide comprising a reverse transcriptase (RT) domain and the second polypeptide comprising a DNA-binding domain and an endonuclease domain, such as a nickasase domain. In further examples, in some embodiments, each of the first and second polypeptides comprises a DNA-binding domain (e.g., a first DNA-binding domain and a second DNA-binding domain). In some embodiments, the first and second polypeptides may be joined posttranslatically via a split intein to form a single recombinant polypeptide.

[0155] In some embodiments, the recombinant polypeptides described herein (for example, the systems described herein include recombinant polypeptides comprising) 1) a Cas domain (e.g., a Cas nickas domain, e.g., a Cas9 nickas domain), 2) a reverse transcriptase (RT) domain in which the RT domain is the C-terminus of the Cas domain, and a linker positioned between the RT domain and the Cas domain.

[0156] In some embodiments, the recombinant polypeptide includes a GG amino acid sequence between the Cas domain and the linker, an AG amino acid sequence between the RT domain and the second NLS, and / or a GG amino acid sequence between the linker and the RT domain.

[0157] Writing domain (RT domain) In certain embodiments of the present invention, the writing domain of the genetically modified system has reverse transcriptase activity and is also referred to as the reverse transcriptase domain (RT domain). In some embodiments, the RT domain includes an RT catalytic moiety and an RNA-binding region (for example, a region that binds to template RNA).

[0158] In some embodiments, the nucleic acid encoding the reverse transcriptase is modified from its native sequence to have, for example, improved and modified codon usage for human cells. In some embodiments, the reverse transcriptase domain is a heterologous reverse transcriptase from a retrovirus. In some embodiments, the RT domain, comprising a recombinant polypeptide, is mutated from its original amino acid sequence, for example, having at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 substitutions. In some embodiments, the RT domain is derived from the RT of a retrovirus, for example, the RT of reticuloendotheliosis virus of birds (AVIRE) (e.g., UniProtKB accession number P03360).

[0159] In some embodiments, the retroviral reverse transcriptase (RT) domain exhibits increased stringency for target-primed reverse transcription (TPRT) initiation compared to, for example, an endogenous RT domain. In some embodiments, the RT domain initiates TPRT when 3nts within the target site immediately upstream of the first strand nick, for example, the genomic DNA priming the RNA template, have at least 66% or 100% complementarity to 3nts of homology in the RNA template. In some embodiments, the RT domain initiates TPRT when less than 5nts of mismatch exist between the homology of the template RNA and the target DNA-priming reverse transcription (e.g., less than 1, 2, 3, 4, or 5nts of mismatch). In some embodiments, the RT domain is modified to increase stringency in priming mismatches in the TPRT reaction, for example, the RT domain either does not tolerate any mismatches within the priming region or tolerates fewer mismatches compared to a wild-type (e.g., unmodified) RT domain.

[0160] In nature, heterodimeric RT domains may also function as homodimers in some embodiments. In some embodiments, dimeric RT domains are expressed as fusion proteins, such as homodimeric fusion proteins or heterodimeric fusion proteins. In some embodiments, the RT function of the system is satisfied by multiple RT domains. In further embodiments, multiple RT domains may be fused or separated and may reside, for example, on the same polypeptide or on different polypeptides.

[0161] In some embodiments, the genetically modified system described herein includes an integrase domain, for example, the integrase domain may be part of the RT domain. In some embodiments, the RT domain (for example, as described herein) includes an integrase domain. In some embodiments, the RT domain (for example, as described herein) includes an integrase domain that lacks an integrase domain or is inactivated by mutation or deletion. In some embodiments, the genetically modified system described herein includes a ribonuclease H domain, for example, the ribonuclease H domain may be part of the RT domain. In some embodiments, the RNase H domain is not part of the RT domain but is covalently linked via a flexible linker. In some embodiments, the RT domain (for example, as described herein) includes a ribonuclease H domain, for example, an endogenous ribonuclease H domain or a heterologous ribonuclease H domain. In some embodiments, the RT domain (for example, as described herein) lacks a ribonuclease H domain. In some embodiments, the RT domain (e.g., as described herein) comprises a ribonuclease H domain that has been added to, deleted from, mutated, or replaced with a heterologous ribonuclease H domain. In some embodiments, the polypeptide comprises an inactivated endogenous RNase H domain. In some embodiments, the endogenous RNase H domain from one of the other domains of the polypeptide is genetically removed so that it is not included in the polypeptide, for example, by partially or completely cleaving the endogenous RNase H domain from the polypeptide containing the domain. In some embodiments, a mutation in the ribonuclease H domain produces a polypeptide exhibiting lower ribonuclease activity, for example, by measurement by the method in Kotewicz et al. Nucleic Acids Res 16(1):265-277 (1988) (which is incorporated herein by reference in its entirety), compared to other similar domains without the mutation, by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%.In some embodiments, ribonuclease H activity is lost.

[0162] In some embodiments, the RT domain undergoes mutation, resulting in increased fidelity compared to other similar domains that do not have mutations. For example, in some embodiments, the YADD (SEQ ID NO: 37635) or YMDD (SEQ ID NO: 37636) motif within the RT domain (e.g., in reverse transcriptase) is substituted with YVDD (SEQ ID NO: 37637). In several embodiments, substitution of YADD (SEQ ID NO: 37635), YMDD (SEQ ID NO: 37636), or YVDD (SEQ ID NO: 37637) results in higher fidelity in retroviral reverse transcriptase activity (e.g., described in Jamburthugoda and Eickbush J Mol Biol 2011, which is incorporated herein by reference in its entirety).

[0163] In some embodiments, the RT domain for use in the genetically modified systems described herein includes an RT domain from AVIRE. In some embodiments, the RT domain includes one or more mutations listed in Table 2 below, compared to a wild-type or standard RT domain (e.g., from AVIRE). In some embodiments, the RT domain includes one, two, three, four, five, or six of the mutations listed in the corresponding rows of Table 2 below.

[0164] [Table 2]

[0165] In some embodiments, the recombinant polypeptides described herein include an RT domain having an amino acid sequence according to Table 6 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto. In some embodiments, the recombinant polypeptides described herein include an RT domain encoded by a nucleic acid sequence according to Table 6 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto. In some embodiments, the nucleic acid described herein encodes an RT domain having an amino acid sequence according to Table 6 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto.

[0166] [Table 6-1]

[0167] [Table 6-2]

[0168] In some embodiments, the reverse transcriptase domain is modified, for example, by site-directed mutation. In some embodiments, the reverse transcriptase domain is modified to have improved properties. In some embodiments, the reverse transcriptase domain may be modified to have a lower error rate, for example, as described in International Publication No. 2001068895, which is incorporated herein by reference. In some embodiments, the reverse transcriptase domain may be modified to have increased thermal stability. In some embodiments, the reverse transcriptase domain may be modified to have increased processability. In some embodiments, the reverse transcriptase domain may be modified to have resistance to inhibitors. In some embodiments, the reverse transcriptase domain may be modified to be faster. In some embodiments, the reverse transcriptase domain may be modified to have increased resistance to modified nucleotides in the RNA template. In some embodiments, the reverse transcriptase domain may be modified to insert modified DNA nucleotides. In some embodiments, the reverse transcriptase domain is modified to bind to template RNA.

[0169] In some embodiments, the retroviral reverse transcriptase domain may contain one or more mutations from the wild-type sequence that can improve the characteristics of RT, such as thermal stability, processing capacity, and / or template binding.

[0170] In some embodiments, the writing domain (e.g., the RT domain) includes, for example, an RNA-binding domain that specifically binds to an RNA sequence. In some embodiments, the template RNA includes an RNA sequence that is specifically bound by the RNA-binding domain of the writing domain.

[0171] In some embodiments, the reverse transcription domain simply recognizes and reverse transcribes a specific template of the system, e.g., template RNA. In some embodiments, the template includes a sequence or structure that enables recognition and reverse transcription by the reverse transcription domain. In some embodiments, the template includes a sequence or structure that enables association of the polypeptide component of the genome modification system described herein with the RNA-binding domain. In some embodiments, the genome modification system preferentially reverse transcribes a template containing an association sequence over a template lacking the association sequence.

[0172] The writing domain may also include DNA-dependent DNA polymerase activity, such as enzymatic activity capable of writing DNA to the genome from a template DNA sequence. In some embodiments, DNA-dependent DNA polymerization is used to complete the second strand synthesis of the targeted site edit. In some embodiments, DNA-dependent DNA polymerase activity is provided by the DNA polymerase domain in the polypeptide. In some embodiments, DNA-dependent DNA polymerase activity is provided by a reverse transcriptase domain capable of DNA-dependent DNA polymerization, such as the second strand synthesis. In some embodiments, DNA-dependent DNA polymerase activity is provided by the second polypeptide of the system. In some embodiments, DNA-dependent DNA polymerase activity is optionally provided by an endogenous host cell polymerase recruited to the target site by a component of the genome modification system.

[0173] In some embodiments, the reverse transcriptase domain exhibits a lower probability of insufficient termination (P) compared to the reference reverse transcriptase domain in vitro. off ) has. In some embodiments, the reference reverse transcriptase domain is a viral reverse transcriptase domain, for example, the RT domain from AVIRE.

[0174] In some embodiments, the reverse transcriptase domain, for example, measured in 1094 nt RNA, is approximately 5 × 10¹⁴ in vitro. -3 / nt, 5×10 -4 / nt or 5×10 -6A lower probability of insufficient termination (P) is less than / nt. off ) has. In some embodiments, the insufficient termination rate in vitro is determined as described in Bibillo and Eickbush (2002) J Biol Chem 277(38):34836-34845 (which is incorporated herein by reference in its entirety).

[0175] In some embodiments, the reverse transcriptase domain can complete at least about 30% or 50% of the integration in cells. The percentage of complete integration can be measured by dividing the number of substantially full-length integration events (e.g., genomic sites containing at least 98% of the expected integration sequence) by the total number of integration events (including substantially full-length and partial) in the cell population. In several embodiments, integration in cells is determined (e.g., through integration sites) using long-read amplicon sequencing, as described, for example, in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903 (which is incorporated herein by reference in its entirety).

[0176] In several embodiments, the quantification of integration in cells involves counting the percentage of integration that includes at least about 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the DNA sequence corresponding to the template RNA (e.g., template RNA having lengths of at least 0.05, 0.1, 0.5, 0.6, 0.7, 0.7, 0.8, 0.8, 0.9, 1.0-1.2, 1.2-1.4, 1.4-1.6, 1.6-1.8, 1.8-2.0, 2-3, 3-4, or 4-5 kb).

[0177] In some embodiments, the reverse transcriptase domain is capable of polymerizing dNTPs in vitro. In multiple embodiments, the reverse transcriptase domain is capable of polymerizing dNTPs in vitro at a rate of 0.1 to 50 nt / sec (e.g., 0.1 to 1, 1 to 10, or 10 to 50 nt / sec). In multiple embodiments, the polymerization of dNTPs by the reverse transcriptase domain is measured by a single molecule assay, for example, as described in Schwartz and Quake (2009) PNAS 106(48):20294-20299 (which is incorporated herein by reference in its entirety).

[0178] In some embodiments, the reverse transcriptase domain has an in vitro error rate (e.g., nucleotide misincorporation) of, for example, 1×10 -3 ~1×10 -4 or 1×10 -4 ~1×10 -5 substitutions / nt, for example, as described in Yasukawa et al. (2017) Biochem Biophys Res Commun 492(2):147-153 (which is incorporated herein by reference in its entirety). In some embodiments, the reverse transcriptase domain has an error rate (e.g., nucleotide misincorporation) of, for example, 1×10 -3 ~1×10 -4 or 1×10 -4 ~1×10 -5 substitutions / nt within cells (e.g., HEK293T cells), for example, by long-read amplicon sequencing, as described in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903 (which is incorporated herein by reference in its entirety).

[0179] In some embodiments, the reverse transcriptase domain is capable of performing reverse transcription of target RNA in vitro. In some embodiments, the reverse transcriptase requires a primer of at least 3 nucleotides to initiate reverse transcription of the template. In some embodiments, reverse transcription of target RNA is determined by detection of cDNA from the target RNA, for example, as described in Bibillo and Eickbush (2002) J Biol Chem 277(38):34836-34845 (the entire text of which is incorporated herein by reference) (for example, when an ssDNA primer is provided that anneals to a target having at least 3, 4, 5, 6, 7, 8, 9, or 10 nt at its 3' end).

[0180] In some embodiments, the reverse transcriptase domain, for example, when converting its RNA template to cDNA, carries out reverse transcription at least 5 or 10 times more efficiently (e.g., by cDNA generation) compared to an RNA template lacking a protein-binding motif (e.g., 3'UTR). In several embodiments, the efficiency of reverse transcription is measured as described in Yasukawa et al. (2017) Biochem Biophys Res Commun 492(2):147-153 (which is incorporated herein by reference in its entirety).

[0181] In some embodiments, the reverse transcriptase domain, when expressed in cells (e.g., HEK293T cells), specifically binds to a particular RNA template at a higher frequency (e.g., about 5 or 10 times higher) than any endogenous cellular RNA. In several embodiments, the frequency of specific binding between the reverse transcriptase domain and template RNA is measured by CLIP-seq, for example, as described in Lin and Miles (2019) Nucleic Acids Res 47(11):5490-5501 (which is incorporated herein by reference in its entirety).

[0182] Template nucleic acid binding domain Recombinant polypeptides typically have a region capable of associating with a template nucleic acid (e.g., template RNA). In some embodiments, the template nucleic acid-binding domain is an RNA-binding domain. In some embodiments, the RNA-binding domain is a modular domain capable of associating with RNA molecules containing a specific signature, such as a structural motif. In other embodiments, the template nucleic acid-binding domain (e.g., RNA-binding domain) is contained within a reverse transcription domain, where, for example, a component derived from reverse transcriptase has a known signature regarding RNA preference.

[0183] In other embodiments, the template nucleic acid binding domain (e.g., RNA binding domain) is contained within the target DNA binding domain. For example, in some embodiments, the DNA binding domain is a CRISPR-related protein that recognizes the structure of a template nucleic acid (e.g., template RNA) containing gRNA. In some embodiments, the recombinant polypeptide includes a DNA binding domain containing a CRISPR-related protein that associates with a gRNA scaffold, enabling the DNA binding domain to bind to a target genomic DNA sequence. In some embodiments, the gRNA scaffold and gRNA spacer are contained within the template nucleic acid (e.g., template RNA), so the DNA binding domain is also the template nucleic acid binding domain. In some embodiments, the polypeptide has RNA binding function within multiple domains, which may bind to gRNA structures within the CRISPR-related DNA binding domain and additional sequences or structures within the reverse transcriptase domain.

[0184] In some embodiments, the RNA binding domain can bind to the template RNA with a higher affinity than a standard RNA binding domain. In some embodiments, the standard RNA binding domain is the RNA binding domain derived from Cas9 of Streptococcus pyogenes (S. pyogenes). In some embodiments, the RNA binding domain can bind to the template RNA with an affinity of 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM). In some embodiments, the affinity of the RNA binding domain for its template RNA is measured in vitro, for example, by thermal shift, as described in, for example, Asmari et al. Methods 146:107-119 (2018) (which is incorporated herein by reference in its entirety). In some embodiments, the affinity of the RNA binding domain for its template RNA is measured in cells (e.g., by FRET or CLIP-Seq).

[0185] In some embodiments, the RNA binding domain associates with the template RNA at a frequency at least about 5-fold or 10-fold higher than that with scrambled RNA in vitro. In some embodiments, the frequency of association between the RNA binding domain and the template RNA or scrambled RNA is measured by CLIP-seq, as described in, for example, Lin and Miles (2019) Nucleic Acids Res 47(11):5490-5501 (which is incorporated herein by reference in its entirety). In some embodiments, the RNA binding domain associates with the template RNA in cells (e.g., in HEK293T cells) at a frequency at least about 5-fold or 10-fold higher than that with scrambled RNA. In some embodiments, the frequency of association between the RNA binding domain and the template RNA or scrambled RNA is measured by CLIP-seq, as described in, for example, Lin and Miles (2019) (supra).

[0186] Endonuclease domain and DNA binding domain In some embodiments, the recombinant polypeptide has the function of cleaving a DNA target site via an endonuclease domain. In some embodiments, the recombinant polypeptide includes, for example, a DNA binding domain for binding to a target nucleic acid. In some embodiments, the domain of the recombinant polypeptide (e.g., the Cas domain) includes two or more smaller domains, such as a DNA binding domain and an endonuclease domain. When the DNA binding domain (e.g., the Cas domain) is described as binding to a target nucleic acid sequence, it is understood that in some embodiments, the binding is mediated by a gRNA.

[0187] In some embodiments, the domain has two functions. For example, in some embodiments, the endonuclease domain is also a DNA binding domain. In some embodiments, the endonuclease domain is also a template nucleic acid (e.g., template RNA) binding domain. For example, in some embodiments, the polypeptide includes a CRISPR-related endonuclease domain that binds to a template RNA containing a gRNA, binds to a target DNA sequence (e.g., having complementarity to a portion of the gRNA), and cleaves the target DNA sequence. In some embodiments, an endonuclease domain or an endonuclease / DNA binding domain derived from a heterologous source can be used in the gene recombination system described herein or can be modified (e.g., by insertion, deletion, or substitution of one or more residues).

[0188] In some embodiments, the nucleic acid encoding the endonuclease domain or the endonuclease / DNA binding domain is modified from its native sequence, for example, to have a modified codon usage optimized for human cells. In some embodiments, the endonuclease element is a heterologous endonuclease element, such as a Cas endonuclease (e.g., Cas9).

[0189] In certain embodiments, the DNA-binding domain of the recombinant polypeptide described herein is selected, designed, or constructed for binding to a desired host DNA target sequence. In certain embodiments, the DNA-binding domain of the polypeptide is a heterogeneous DNA-binding factor. In some embodiments, the heterogeneous DNA-binding factor is a sequence-inducible DNA-binding factor such as Cas9, Cpf1, or other CRISPR-related protein, which has been modified to lack endonuclease activity. In some embodiments, the heterogeneous DNA-binding factor retains endonuclease activity. In some embodiments, the heterogeneous DNA-binding factor retains partial endonuclease activity, such as cleaving ssDNA, and has, for example, nickase activity.

[0190] In some embodiments, the DNA-binding domain is modified, for example, by site-directed mutation, to increase or decrease DNA-binding factors (e.g., the number and / or specificity of zinc fingers), thereby altering DNA-binding specificity and affinity. In some embodiments, the nucleic acid sequence encoding the DNA-binding domain is modified from its native sequence, for example, to have improved modified codon use for human cells. In several embodiments, the DNA-binding domain includes one or more modifications to the wild-type DNA-binding domain, such as modifications by directed evolution, such as phage-assisted continuous evolution (PACE).

[0191] In some embodiments, the recombinant polypeptide includes modifications to the DNA-binding domain compared to, for example, the wild-type polypeptide. In some embodiments, the DNA-binding domain includes additions, deletions, substitutions, or modifications to the amino acid sequence of the original DNA-binding domain. In some embodiments, the DNA-binding domain is modified to include a heterologous functional domain that specifically binds to a target nucleic acid (e.g., DNA) sequence of interest. In some embodiments, the functional domain is replaced with at least a portion (e.g., all) of the DNA-binding domain of the polypeptide. In some embodiments, the Cas domain includes Cas9 or its mutants or variants (e.g., as described herein). In several embodiments, the Cas domain associates with a guide RNA (gRNA), for example, as described herein. In several embodiments, the Cas domain is guided by the gRNA to a target nucleic acid (e.g., DNA) sequence of interest. In several embodiments, the Cas domain is encoded by the same nucleic acid (e.g., RNA) molecule as the gRNA. In several embodiments, the Cas domain is encoded by a different nucleic acid (e.g., RNA) molecule than the gRNA.

[0192] In some embodiments, the DNA-binding domain can bind to a target sequence (e.g., a dsDNA target sequence) with higher affinity than a standard DNA-binding domain. In some embodiments, the standard DNA-binding domain is a DNA-binding domain derived from Cas9 of Streptococcus pyogenes. In some embodiments, the DNA-binding domain can bind to a target sequence (e.g., a dsDNA target sequence) with affinity of 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM).

[0193] In some embodiments, the affinity of the DNA-binding domain for its target sequence (e.g., a dsDNA target sequence) is measured in vitro, for example, by thermophoresis, as described in Asmari et al. Methods 146:107-119 (2018) (the entire text of which is incorporated herein by reference).

[0194] In the embodiment, the DNA-binding domain can bind to its target sequence (e.g., a dsDNA target sequence) with an affinity of, for example, 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM) in the presence of a molar excess, for example, about 100-fold molar excess, of scrambled sequence-competitive dsDNA.

[0195] In some embodiments, the DNA-binding domain is found to bind to the target sequence (e.g., a dsDNA target sequence) at a higher frequency than any other sequence in the genome of the target cell, e.g., a human target cell, as measured by ChIP-seq (e.g., in HEK293T cells), as described, for example, in He and Pu (2010) Curr. Protoc Mol Biol Chapter 21 (the entirety of which is incorporated herein by reference). In some embodiments, the DNA-binding domain is found to bind to the target sequence (e.g., a dsDNA target sequence) at a frequency at least about 5 or 10 times higher than any other sequence in the genome of the target cell, as measured by ChIP-seq (e.g., in HEK293T cells), as described, for example, in He and Pu (2010) (cited above).

[0196] In some embodiments, the endonuclease domain has nickase activity and cleaves one strand of target DNA. In some embodiments, the nickase activity reduces the formation of double-strand breaks at the target site. In some embodiments, the endonuclease domain generates a twisted nick structure in the first and second strands of target DNA. In some embodiments, the twisted nick structure generates a free 3' overhang at the target site. In some embodiments, the free 3' overhang at the target site improves editing efficiency, for example, by enhancing access to and annealing of the 3' homology region of the template nucleic acid. In some embodiments, the twisted nick structure reduces the formation of double-strand breaks at the target site.

[0197] In some embodiments, the endonuclease domain cleaves both strands of target DNA, resulting in a blunt-end cleavage of the target that, for example, does not have an ssDNA overhang on either side of the cleavage site. The amino acid sequences of the endonuclease domains of the recombinant systems described herein may be at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and at least about 99% identical to the amino acid sequences of the endonuclease domains described herein.

[0198] In certain embodiments, the heterologous endonuclease is derived from a CRISPR-related protein, such as Cas9. In certain embodiments, the heterologous endonuclease is modified to have only ssDNA cleavage activity, such as only nickase activity, for example, to be a Cas9 nickase, such as SpCas9 with a D10A, H840A, or N863A mutation. In yet another embodiment, the homologous endonuclease domain is modified to alter DNA endonuclease activity, for example, by site-directed mutation. In yet another embodiment, the endonuclease domain is modified to reduce DNA sequence specificity, for example, by cleavage to remove a domain that confers DNA sequence specificity or mutation, in order to inactivate a region that confers DNA sequence specificity.

[0199] In some embodiments, the endonuclease domain has nickase activity and does not form double-strand breaks. In some embodiments, the endonuclease domain forms single-strand breaks more frequently than double-strand breaks, for example, at least 90%, 95%, 96%, 97%, 98%, or 99% of the breaks are single-strand breaks, or less than 10%, 5%, 4%, 3%, 2%, or 1% of the breaks are double-strand breaks. In some embodiments, the endonuclease substantially does not form double-strand breaks. In some embodiments, the endonuclease does not form detectable levels of double-strand breaks.

[0200] In some embodiments, the endonuclease domain has nickase activity that cleaves target site DNA on the first strand; for example, in some embodiments, the endonuclease domain cleaves genomic DNA at a target site near a modification site on the strand that will be extended by the writing domain. In some embodiments, the endonuclease domain has nickase activity that cleaves target site DNA on the first strand but not on target site DNA on the second strand. For example, when the polypeptide contains a CRISPR-related endonuclease domain having nickase activity, in some embodiments, the CRISPR-related endonuclease domain cleaves target site DNA strands containing PAM sites (and, for example, does not cleave target site DNA strands that do not contain PAM sites). As a further example, when the polypeptide contains a CRISPR-related endonuclease domain having nickase activity, in some embodiments, the CRISPR-related endonuclease domain cleaves the target site DNA strand that does not contain a PAM site (but does not cleave the target site DNA strand that does contain a PAM site, for example).

[0201] In some other embodiments, the endonuclease domain has nickase activity that cleaves the target site DNA of the first and second strands. Although not intended to be bound by any particular theory, after the writing domain (e.g., RT domain) of the polypeptide described herein is polymerized (e.g., reverse transcribed) from a heterologous target sequence of the template nucleic acid (e.g., template RNA), the cellular DNA repair mechanism must repair the nick on the first DNA strand. The target site DNA here comprises two distinct sequences relative to the first DNA strand, the first corresponding to the original genomic DNA (e.g., with a free 5' end) and the second corresponding to the one polymerized from the heterologous target sequence (e.g., with a free 3' end). It is thought that the two distinct sequences equilibrate with each other, and one hybridizes with the second strand first, then the other, and the order of their incorporation into its repair target site by the cellular DNA repair mechanism is considered to be a stochastic process. While not intended to be bound by any particular theory, it is thought that the introduction of further nicks into the second strand may bias cellular DNA repair mechanisms to use sequences based on heterologous target sequences more frequently than the original genome sequence (Anzalone et al. Nature 576:149-157 (2019)). In some embodiments, the further nicks are located at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, or 150 nucleotides relative to the 5' or 3' of the target site modification (e.g., insertion, deletion, or substitution) or the nicks on the first strand.

[0202] Alternatively or additionally, and without intending to be bound by any particular theory, it is thought that further nicks to the second strand may facilitate the synthesis of the second strand. In some embodiments, if the recombination system inserts or substitutes a portion of the first strand, the synthesis of a new sequence corresponding to the insertion / substitution in the second strand is required.

[0203] In some embodiments, the polypeptide comprises a single domain having endonuclease activity (e.g., a single endonuclease domain), the domain cleaving both a first and a second strand. For example, in such embodiments, the endonuclease domain may be a CRISPR-associated endonuclease domain, and the template nucleic acid (e.g., template RNA) comprises a gRNA spacer that leads to nicking of the first strand and a further gRNA spacer that leads to nicking of the second strand. In some embodiments, the polypeptide comprises multiple domains having endonuclease activity, the first endonuclease domain cleaving the first strand and the second endonuclease domain cleaving the second strand (optionally, the first endonuclease domain may not cleaving the second strand (e.g., is unable to do so) and the second endonuclease domain may not cleaving the first strand (e.g., is unable to do so)).

[0204] In some embodiments, the endonuclease domain can puncture the first and second chains. In some embodiments, the first and second chain nicks occur at the same location in the target site but not on the reverse chain. In some embodiments, the second chain nick occurs at a twisted location, for example, upstream or downstream of the first nick. In some embodiments, the endonuclease domain generates a deletion in the target site if the second chain nick is upstream of the first chain nick. In some embodiments, the endonuclease domain generates a duplication of the target site if the second chain nick is downstream of the first chain nick. In some embodiments, the endonuclease domain does not generate duplication and / or deletion if the first and second chain nicks occur at the same location in the target site. In some embodiments, the endonuclease domain has modified activity depending on the protein structure or RNA binding state, for example, to promote nicking of the first or second strand (as described in whole herein by reference, e.g., Christensen et al. PNAS 2006).

[0205] In some embodiments, the recombinant polypeptide includes modifications to the endonuclease domain compared to, for example, the wild-type Cas protein. In some embodiments, the endonuclease domain includes additions, deletions, substitutions, or modifications to the amino acid sequence of the wild-type Cas protein. In some embodiments, the endonuclease domain is modified to include a heterologous functional domain that specifically binds to a target nucleic acid (e.g., DNA) sequence and / or induces its endonuclease cleavage.

[0206] In some embodiments, the endonuclease domain associates with target dsDNA at a frequency at least approximately 5 or 10 times higher than that of scrambled dsDNA in vitro. In some embodiments, the endonuclease domain associates with target dsDNA at a frequency at least approximately 5 or 10 times higher than that of scrambled dsDNA in vitro, for example, within cells (e.g., HEK293T cells). In some embodiments, the frequency of association between the endonuclease domain and target DNA or scrambled DNA is measured by ChIP-seq, for example, as described in He and Pu (2010) Curr. Protoc Mol Biol Chapter 21 (the entirety of which is incorporated herein by reference).

[0207] In some embodiments, the endonuclease domain can catalyze nick formation at target sequences by at least a 5-fold or 10-fold increase compared to, for example, non-target sequences (e.g., compared to any other genomic sequence in the target cell's genome). In some embodiments, the level of nick formation is measured using nickSeq, for example, as described in Elacqua et al. (2019) bioRxiv doi.org / 10.1101 / 867937 (which is incorporated herein by reference in its entirety).

[0208] In some embodiments, the endonuclease domain is capable of nicking DNA in vitro. In multiple embodiments, the nick results in an exposed base. In multiple embodiments, the exposed base can be detected using a nuclease sensitivity assay, such as that described in Chaudhry and Weinfeld (1995) Nucleic Acids Res 23(19):3805-3809, which is incorporated herein by reference in its entirety. In multiple embodiments, the level of the exposed base (e.g., detected by a nuclease sensitivity assay) increases by at least 10%, 50% or more compared to a reference endonuclease domain. In some embodiments, the reference endonuclease domain is an endonuclease domain derived from Cas9 of Streptococcus pyogenes (S. pyogenes).

[0209] In some embodiments, the endonuclease domain is capable of nicking DNA intracellularly. In multiple embodiments, the endonuclease domain is capable of nicking DNA in HEK293T cells. In multiple embodiments, an unrepaired nick that undergoes replication in the absence of Rad51 results in an increase in the NHEJ rate at the site of the nick, which can be detected, for example, by using a Rad51 inhibition assay, such as that described in Bothmer et al. (2017) Nat Commun 8:13905, which is incorporated herein by reference in its entirety. In multiple embodiments, the NHEJ rate increases by more than 0 to 5%. In multiple embodiments, the NHEJ rate increases to 20 to 70% (e.g., 30% to 60% or 40 to 50%), for example, upon Rad51 inhibition.

[0210] In some embodiments, the endonuclease domain releases the target after cleavage. In some embodiments, the release of the target is indirectly indicated by evaluating multiple turnovers by the enzyme, for example, as described in Yourik at al. RNA 25(1):35-44(2019) (the whole of which is incorporated herein by reference) and as shown in Figure 2. In some embodiments, the k of the endonuclease domain exp This is measured by this method, and 1 × 10 -3 ~1 × 10 -5 It is min-1.

[0211] In some embodiments, the endonuclease domain is approximately 1 × 10⁶ in vitro. 8 s -1 M -1 Catalytic efficiency exceeding (k cat / K m ) has. In some embodiments, the endonuclease domain is approximately 1 × 10 in vitro. 5 , 1 x 10 6 , 1 x 10 7 or 1 × 10 8 s -1 M -1 It has a catalytic efficiency exceeding 1 × 10⁻¹⁶. In several embodiments, the catalytic efficiency is determined as described in Chen et al. (2018) Science 360(6387):436-439 (which is incorporated herein by reference in its entirety). In some embodiments, the endonuclease domain has a catalytic efficiency of approximately 1 × 10⁻¹⁶ in the cell. 8 s -1 M -1 Catalytic efficiency exceeding (k cat / K m ) has. In some embodiments, the endonuclease domain is located in the cell at approximately 1 × 10⁻¹⁶ 5 , 1 x 10 6 , 1 x 10 7 or 1 × 10 8 s -1 M -1 It has a catalytic efficiency exceeding [a certain level].

[0212] Recombinant polypeptide containing a Cas domain In some embodiments, the recombinant polypeptides described herein include a Cas domain. In some embodiments, the Cas domain can guide the recombinant polypeptide to a target site identified by a gRNA spacer, thereby modifying the target nucleic acid sequence at "cis". In some embodiments, the recombinant polypeptide is fused to the Cas domain. In some embodiments, the recombinant polypeptide includes a CRISPR / Cas domain (also referred to herein as a CRISPR-related protein). In some embodiments, the CRISPR / Cas domain includes a protein involved in clustering and regularly arranged short palindromic sequence repeats (CRISPR), such as a Cas protein, which optionally binds to a guide RNA, such as a single guide RNA (sgRNA).

[0213] The CRISPR system is an adaptive defense system first discovered in bacteria and archaea. The CRISPR system uses RNA-guided nucleases, referred to as CRISPR-associated or "Cas" endonucleases (e.g., Cas9 or Cpf1), to cleave foreign DNA. For example, in a typical CRISPR-Cas system, the endonuclease is guided to a target nucleotide sequence (e.g., a site in the genome to be sequence-edited) by a sequence-specific, non-coding "guide RNA" that targets a single-stranded or double-stranded DNA sequence. Three classes (I-III) of CRISPR systems have been identified. Class II CRISPR systems use a single Cas endonuclease (rather than multiple Cas proteins). One class II CRISPR system includes type II Cas endonucleases such as Cas9, CRISPR RNA ("crRNA"), and transactivating crRNA ("tracrRNA"). crRNA typically contains a "spacer" sequence (protospacer), which is an RNA sequence of about 20 nucleotides corresponding to the target DNA sequence. In the wild-type system and some modified systems, crRNA binds to tracrRNA and forms a partially double-stranded structure that is cleaved by ribonuclease III. cThis also includes a region that gives rise to the rRNA / tracrRNA hybrid molecule. The crRNA / tracrRNA hybrid then guides the Cas endonuclease to recognize and cleave the target DNA sequence. The target DNA sequence is generally specific to a given Cas endonuclease and is adjacent to a "protospacer-adjacent motif" ("PAM") required for cleavage activity at the target site corresponding to the crRNA spacer. CRISPR endonucleases identified from various prokaryotic species have specific PAM sequence requirements, as listed for example for the Cas enzymes in Table 3. Examples of PAM sequences include 5'-NGG (Streptococcus pyogenes), 5'-NNAGAA (Streptococcus thermophilus CRISPR1), 5'-NGGNG (Streptococcus thermophilus CRISPR3), and 5'-NNNGATT (Neisseria meningiditis). Some endonucleases, such as Cas9 endonucleases, associate with a G-rich PAM site, such as 5'-NGG, and perform blunt-end cleavage of target DNA at three nucleotides upstream (from the 5' side) of the PAM site. Another class II CRISPR system includes a smaller V-type endonuclease Cpf1 than Cas9, such as AsCpf1 (derived from Acidaminococcus sp.) and LbCpf1 (derived from Lachnospiraceae sp.). Cpf1-associated CRISPR arrays do not require tracrRNA and are treated with mature crRNA; in other words, the Cpf1 system, in some embodiments, comprises a Cpf1 nuclease and crRNA exclusively for cleaving target DNA sequences. The Cpf1 endonuclease typically associates with T-rich PAM sites, such as 5'-TTN. Cpf1 may also recognize 5'-CTA PAM motifs.Cpf1 typically cleaves target DNA by introducing offset or twisted double-strand breaks in the 5' overhangs of 4- or 5-nucleotides, for example, 18 nucleotides downstream (3' side) from the PAM site on the coding strand and 23 nucleotides downstream from the PAM site on the complementary strand. The resulting 5-nucleotide overhangs enable more precise genome editing by DNA insertion via homologous recombination rather than insertion via blunt-end DNA. See, for example, Zetsche et al. (2015) Cell, 163:759-771.

[0214] Various CRISPR-related (Cas) genes or proteins may be used in the techniques provided herein, and the selection of Cas proteins will depend on the specific conditions of the method. Specific examples of Cas proteins include the Class II system comprising Cas1, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cpf1, C2C1, or C2C3. In some embodiments, the Cas protein, e.g., the Cas9 protein, may originate from any of various prokaryotic species. In some embodiments, a specific Cas protein, e.g., a specific Cas9 protein, is selected to recognize a specific protospacer-adjacent motif (PAM) sequence. In some embodiments, the DNA-binding domain or endonuclease domain includes a sequence-targeting polypeptide such as a Cas protein, e.g., Cas9. In certain embodiments, the Cas protein, e.g., the Cas9 protein, may be obtained from bacteria or archaea, or synthesized using known methods. In certain embodiments, the Cas protein may originate from Gram-positive or Gram-negative bacteria. In certain embodiments, the Cas protein may be derived from streptococci (e.g., Streptococcus pyogenes).

[0215] In some embodiments, the recombinant polypeptide may contain a Cas domain or a functional fragment thereof listed in Table 7 or 8, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.

[0216] [Table 7]

[0217] [Table 8-1]

[0218] [Table 8-2]

[0219] [Table 8-3]

[0220] In some embodiments, the Cas protein is required to have a protospacer-adjacent motif (PAM) located within or adjacent to the target DNA sequence to which the Cas protein binds and / or functions. In some embodiments, the PAM is or comprises NGG, YG, NNGRRT, NNNRRT, NGA, TYCV, TATV, NTTN, or NNNGATT from 5' to 3', where N represents any nucleotide, Y represents C or T, R represents A or G, and V represents A or C or G. In some embodiments, the Cas protein is a protein listed in Table 7 or 8. In some embodiments, the Cas protein includes one or more mutations that modify its PAM. Exemplary advances in the modification of the Cas enzyme for recognizing modified PAM sequences are outlined in Collias et al Nature Communications 12:555 (2021), which are incorporated herein by reference in their entirety.

[0221] In some embodiments, the Cas protein is catalytically active and cleaves one or both strands of a target DNA site. In some embodiments, following the cleavage of the target DNA site, modifications, such as insertions or deletions, are formed, for example, by cell repair mechanisms.

[0222] In some embodiments, Cas proteins are modified to inactivate or partially inactivate a nuclease, such as nuclease-deficient Cas9. While wild-type Cas9 produces double-strand breaks (DSBs) at specific DNA sequences targeted by gRNA, several CRISPR endonucleases with modified functionality are available; for example, a “nickase” version of partially inactivated Cas9 produces only single-strand breaks, and catalytically inactive Cas9 ("dCas9") does not cleave target DNA. In some embodiments, binding of dCas9 to a DNA sequence may interfere with transcription at that site due to steric hindrance. In some embodiments, binding of dCas9 to an anchor sequence may interfere with (e.g., reduce or block) the formation and / or maintenance of genomic complexes (e.g., ASMCs). In some embodiments, the DNA-binding domain includes catalytically inactive Cas9, e.g., dCas9. Numerous catalytically inactive Cas9 proteins are known in the art. In some embodiments, dCas9 contains mutations, e.g., the N863A mutation, within each endonuclease domain of the Cas protein. In some embodiments, catalytically inactive or partially inactive CRISPR / Cas domains contain one or more mutations, e.g., one or more of the mutations listed in Table 7, within the Cas protein. In some embodiments, a Cas protein listed in a given row of Table 7 contains one, two, three, or all of the mutations listed in the same row of Table 7. In some embodiments, for example, a Cas protein not listed in Table 7 contains one, two, three, or all of the mutations listed in a certain row of Table 7, or the corresponding mutation at the corresponding site in that Cas protein.

[0223] In some embodiments, catalytically inactive Cas9 proteins, such as dCas9, or partially inactivated Cas9 proteins include an N863 mutation (e.g., an N863A mutation) or a similar substitution to the amino acid corresponding to the said position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9, include a D10 mutation (e.g., D10A), a D839 mutation (e.g., D839A), an H840 mutation (e.g., H840A), and an N863 mutation (e.g., N863A) or a similar substitution to the amino acid corresponding to the said position.

[0224] In some embodiments, the DNA-binding domain or endonuclease domain may include a gRNA (e.g., a template nucleic acid containing gRNA, e.g., template RNA) or a Cas molecule linked to it (e.g., covalently).

[0225] In some embodiments, the endonuclease domain or DNA-binding domain includes Streptococcus pyogenes Cas9 (SpCas9) or a functional fragment or variant thereof. In some embodiments, the endonuclease domain or DNA-binding domain includes modified SpCas9. In several embodiments, the modified SpCas9 includes modifications that alter the protospacer-adjacent motif (PAM) specificity. In several embodiments, the PAM has specificity for the nucleic acid sequence 5'-NGT-3'. In some embodiments, the endonuclease domain or DNA-binding domain includes a Cas domain, such as a Cas9 domain. In several embodiments, the endonuclease domain or DNA-binding domain includes a nuclease-active Cas domain, a Cas niccas (nCas) domain, or a nuclease-inactive Cas (dCas). In several embodiments, the endonuclease domain or DNA-binding domain includes a nuclease-active Cas9 domain, a Cas9 niccasse (nCas9) domain, or a nuclease-inactive Cas9 (dCas9).

[0226] In some embodiments, the endonuclease domain or DNA-binding domain includes a Cas9 sequence, for example, as described in Chylinski, Rhun, and Charpentier (2013) RNA Biology 10:5, 726-737 (incorporated herein by reference). In some embodiments, the endonuclease domain or DNA-binding domain includes the HNH nuclease subdomain and / or RuvC1 subdomain of Cas, for example, Cas9 or a variant thereof, as described herein. In some embodiments, the endonuclease domain or DNA-binding domain includes a Cas polypeptide (e.g., an enzyme) or a functional fragment thereof. In some embodiments, Cas9 includes one or more substitutions and / or one or more mutations.

[0227] Further components of the exemplary genetically modified polypeptide In addition to the writing / RT domain, endonuclease, and DNA-binding domain, exemplary recombinant polypeptides for use in genetically modified systems may include further components incorporated herein.

[0228] Linker In some embodiments, the recombinant polypeptide may include a linker, such as a peptide linker, such as those listed in Table 10. In some embodiments, the recombinant polypeptide includes, from the N-terminus to the C-terminus, a Cas domain (e.g., the Cas domain in Table 8), a linker in Table 10 (or a sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity thereto), and an RT domain (e.g., the RT domain in Table 6). In some embodiments, the recombinant polypeptide includes a flexible linker between the endonuclease and the RT domain, such as the amino acid sequence AEAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAKEAAAKEAAAKA (SEQ ID NO: 54) or a sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain of the recombinant polypeptide may be located at the C-terminus relative to the endonuclease domain. In some embodiments, the RT domain of the recombinant polypeptide may be located at the N-terminus relative to the endonuclease domain.

[0229] [Table 10]

[0230] In some embodiments, the recombinant polypeptide includes (i) a linker comprising a linker sequence listed in the row of Table T1 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto, and (ii) an RT domain comprising an RT domain sequence listed in the same row of Table T1 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.

[0231] [Table T1]

[0232] In some embodiments, the recombinant polypeptide includes (i) a Cas domain comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the Cas sequence listed in the row of Table T2, (ii) a linker comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the linker sequence listed in the same row of Table T2, and (iii) an RT domain comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the RT domain sequence listed in the same row of Table T2.

[0233] [Table T2-1]

[0234] [Table T2-2]

[0235] localization array In certain embodiments, the recombinant system RNA further comprises an intracellular localization sequence, such as a nuclear localization sequence (NLS). In some embodiments, the recombinant polypeptide or the nucleic acid (e.g., RNA) encoding the recombinant polypeptide comprises an NLS. The nuclear localization sequence may be an RNA sequence that facilitates the entry of RNA into the nucleus. In certain embodiments, the nuclear localization signal is located on the template RNA. In certain embodiments, the recombinant polypeptide is encoded on a first RNA, the template RNA is a second RNA, and the nuclear localization signal is located on the template RNA and not on the RNA encoding the recombinant polypeptide. While not intended to be bound by any particular theory, in some embodiments, the RNA encoding the recombinant polypeptide is targeted primarily to the cytoplasm to facilitate its translation, while the template RNA is targeted primarily to the nucleus to facilitate insertion into the genome. In some embodiments, the nuclear localization signal is located within the 3' end, 5' end, or an internal region of the template RNA. In some embodiments, the nuclear localization signal is at the 3' of a heterologous sequence (e.g., directly at the 3' of a heterologous sequence) or at the 5' of a heterologous sequence (e.g., directly at the 5' of a heterologous sequence). In some embodiments, the nuclear localization signal is located outside the 5'UTR or outside the 3'UTR of the template RNA. In some embodiments, the nuclear localization signal is located between the 5'UTR and the 3'UTR, and optionally, the nuclear localization signal is not transcribed by the transgene (e.g., the nuclear localization signal is antisense-oriented or downstream of a transcription termination signal or polyadenylation signal). In some embodiments, the nuclear localization sequence is located inside an intron. In some embodiments, multiple identical or different nuclear localization signals are present in the RNA, for example, in the template RNA. In some embodiments, the nuclear localization signal is 5 bp, 10 bp, 25 bp, 50 bp, 75 bp, 100 bp, 150 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or less than 1000 bp in length. Various RNA nuclear localization sequences are available.For example, Lubelsky and Ulitsky, Nature 555(107-111), 2018, describe RNA sequences that drive the nuclear localization of RNA. In some embodiments, the nuclear localization signal is a SINE-derived nuclear RNA localization (SIRLOIN) signal. In some embodiments, the nuclear localization signal binds to nuclear enrichment proteins. In some embodiments, the nuclear localization signal binds to HNRNPK proteins. In some embodiments, the nuclear localization signal is abundant in pyrimidines, e.g., C / T-rich, C / U-rich, C-rich, T-rich, or U-rich regions. In some embodiments, the nuclear localization signal originates from long uncoding RNA. In some embodiments, the nuclear localization signal originates from MALAT1 long uncoding RNA or the 600-nucleotide M region of MALAT1 (described in Miyagawa et al., RNA 18,(738-751), 2012). In some embodiments, the nuclear localization signal is derived from BORG long uncoding RNA or is an AGCCC motif (as described in Zhang et al., Molecular and Cellular Biology 34, 2318-2329 (2014)). In some embodiments, the nuclear localization sequence is described in Shukla et al., The EMBO Journal e98452 (2018). In some embodiments, the nuclear localization signal is derived from a retrovirus.

[0236] In some embodiments, the polypeptide described herein comprises one or more (e.g., 2, 3, 4, 5) nuclear targeting sequences, such as nuclear localization sequences (NLS). In some embodiments, the NLS is a bipartite NLS. In some embodiments, the NLS facilitates the input of the protein containing the NLS into the cell nucleus. In some embodiments, the NLS is fused to the N-terminus of the recombinant polypeptide described herein. In some embodiments, the NLS is fused to the C-terminus of the recombinant polypeptide. In some embodiments, the NLS is fused to the N-terminus or C-terminus of the Cas domain. In some embodiments, a linker sequence is positioned between the NLS and the adjacent domain of the recombinant polypeptide.

[0237] In some embodiments, the NLS includes amino acid sequences as disclosed in Table 11. To improve intracellular localization to the nucleus, the NLS may be utilized as one or more copies within the polypeptide at one or more positions within the polypeptide, e.g., within the N-terminal domain, between peptide domains, within the C-terminal domain, or in combinations of positions, with one, two, three, or more copies of the NLS. Multiple unique sequences may be used within a single polypeptide. Sequences may be naturally monopartite or bipartite, or may be used as chimeric bipartite sequences, for example, having one or two stretches of basic amino acids. Sequence references correspond to UniProt acceptance numbers, unless indicated as SeqNLS for sequences extracted using intracellular localization prediction algorithms (Lin et al BMC Bioinformat 13:157 (2012), the entire work incorporated herein by reference).

[0238] [Table 11]

[0239] In some embodiments, the recombinant polypeptides disclosed herein include an N-terminal NLS comprising the amino acid sequence PAAKRVKLDGG (SEQ ID NO: 36) or a functional fragment thereof (for example, an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto). In some embodiments, the recombinant polypeptides disclosed herein include a C-terminal NLS comprising the amino acid sequence KRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 38) or a functional fragment thereof (for example, an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto).

[0240] Further exemplary NLS sequences are also described in PCT / EP2000 / 011690, which are incorporated herein by reference with respect to the disclosure of its exemplary nuclear localization sequences.

[0241] In some embodiments, the recombinant polypeptide comprises, in order from the N-terminus to the C-terminus, one or more N-terminal methionine residues (e.g., 1, 2, 3, 4, 5, or all 6), a first nuclear localization signal (NLS), a DNA-binding domain, a linker, an RT domain, and / or a second NLS.

[0242] In some embodiments, the recombinant polypeptide includes (i) an N-terminal NLS containing the NLS sequence of PAAKRVKLDGG (SEQ ID NO: 36) or a functional fragment thereof (for example, an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto), (ii) a Cas domain containing the Cas sequence of SEQ ID NO: 52 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto, and (iii) AEAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAKEAAAKEAAAKA (SEQ ID NO: 54) (iv) a linker comprising a linker sequence or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto; an RT domain comprising the RT domain sequence of SEQ ID NO: 56 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto; and (v) the C-terminal NLS of KRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 38) or its functional fragment (e.g., an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto). In some embodiments, the recombinant polypeptide comprises the amino acid sequence of SEQ ID NO: 28 (e.g., shown in Table T3 below) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.

[0243] [Table T3]

[0244] In some embodiments, the recombinant polypeptide further comprises an N-terminal methionine residue.

[0245] In some embodiments, the recombinant polypeptide further comprises a T2A sequence and / or a puromycin sequence (e.g., C-terminus to a second NLS). In some embodiments, the nucleic acid encoding the recombinant polypeptide (e.g., as described herein) encodes a T2A sequence, for example, the T2A sequence located between the region encoding the recombinant polypeptide and a second region, the second region optionally encoding a selectable marker, e.g., puromycin.

[0246] In certain embodiments, the recombinant polypeptide further comprises a spacer sequence between the first NLS and the DNA-binding domain. In certain embodiments, the spacer sequence between the first NLS and the DNA-binding domain comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In certain embodiments, the spacer sequence between the first NLS and the DNA-binding domain comprises the amino acid sequence GG.

[0247] In certain embodiments, the recombinant polypeptide further comprises a spacer sequence between the DNA-binding domain and the linker. In certain embodiments, the spacer sequence between the DNA-binding domain and the linker comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In certain embodiments, the spacer sequence between the DNA-binding domain and the linker comprises the amino acid sequence GG.

[0248] In certain embodiments, the recombinant polypeptide further comprises a spacer sequence between the linker and the RT domain. In certain embodiments, the spacer sequence between the linker and the RT domain comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In certain embodiments, the spacer sequence between the linker and the RT domain comprises the amino acid sequence GG.

[0249] In certain embodiments, the recombinant polypeptide further comprises a spacer sequence between the RT domain and the second NLS. In certain embodiments, the spacer sequence between the RT domain and the second NLS comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In certain embodiments, the spacer sequence between the RT domain and the second NLS comprises the amino acid sequence AG.

[0250] In certain embodiments, the recombinant polypeptide further comprises a spacer sequence between the second NLS and T2A sequence and / or the puromycin sequence. In certain embodiments, the spacer sequence between the second NLS and T2A sequence and / or the puromycin sequence comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In certain embodiments, the spacer sequence between the second NLS and T2A sequence and / or the puromycin sequence comprises the amino acid sequence GSG.

[0251] Further domains Recombinant polypeptides can bind to a target DNA sequence and a template nucleic acid (e.g., template RNA), cleave the target site, write the template to the DNA (e.g., reverse transcribe), and result in modification of the target site. In some embodiments, additional domains may be added to the polypeptide to enhance the efficiency of the process. In some embodiments, the recombinant polypeptide may include an additional DNA ligation domain to ligate the reverse-transcribed DNA to the DNA at the target site. In some embodiments, the polypeptide may include a heterologous RNA binding domain. In some embodiments, the polypeptide may include a domain having exonuclease activity from 5' to 3' (e.g., exonuclease activity from 5' to 3' enhances the repair of the modification at the target site, for example, intended for modifications across the original genome sequence). In some embodiments, the polypeptide may include a domain having exonuclease activity from 3' to 5', e.g., proofreading activity. In some embodiments, a writing domain, e.g., an RT domain, has exonuclease activity from 3' to 5', e.g., proofreading activity.

[0252] Nucleic acids encoding genetically modified polypeptides This specification provides nucleic acids encoding recombinant polypeptides for use in the genetically modified systems disclosed herein.

[0253] In some embodiments, the nucleic acid encoding the recombinant polypeptide is modified from its standard sequence, for example, to have improved modified codon usage for human cells. In certain embodiments, the nucleic acid molecule encoding the recombinant polypeptide contains one or more silent mutations in the coding region (for example, in the sequence encoding the RT domain) compared to the nucleic acid molecules described herein.

[0254] In some embodiments, the nucleic acid encoding the recombinant polypeptide (e.g., RNA) includes a post-transcriptional regulatory element that enhances nuclear export. In some embodiments, the post-transcriptional regulatory element includes those of hepatitis B virus (HPRE) or woodchuck hepatitis virus (WPRE). Exemplary WPRE regulatory element nucleic acid sequences are shown in Table N1 below. In some embodiments, the WPRE regulatory element contained in the nucleic acid encoding the recombinant polypeptide includes one or more mutations compared to the wild-type WPRE regulatory element, e.g., mutations described in Zanta-Boussif, MA et al. Gene Ther. 16, 605-619 (2009).

[0255] In some embodiments, nucleic acids encoding a recombinant polypeptide are adjacent to untranslated regions (UTRs) that modify the protein expression level. Various 5' and 3' UTRs can affect protein expression. For example, in some embodiments, the coding sequence may be preceded by a 5' UTR that modifies RNA stability or protein translation. In some embodiments, the sequence may be followed by a 3' UTR that modifies RNA stability or translation. In some embodiments, the sequence may be preceded by a 5' UTR that modifies RNA stability or translation, followed by a 3' UTR. In some embodiments, the 5' and / or 3' UTRs may be selected to enhance the protein expression level. In some embodiments, the 5' and / or 3' UTRs may be selected to modify protein expression in such a way that overproduction inhibition is minimized. In some embodiments, the UTRs are located around the coding sequence, for example, outside the coding sequence, and in other embodiments, close to the coding sequence.

[0256] In some embodiments, the system described herein includes DNA encoding a transcript, the DNA including the corresponding 5'UTR and 3'UTR sequences in which U in the sequences listed above is replaced with T). In some embodiments, the DNA vector used to produce the RNA component of the system further includes a promoter, e.g., T7, T3, or SP6 promoter, upstream of the 5'UTR to initiate in vitro transcription. The above 5'UTR starts with GGG, which is a preferred start for optimizing transcription using T7 RNA polymerase. To adjust the transcription level and modify the transcription start site nucleotide to fit an alternative 5'UTR, the teachings in Davidson et al. Pac Symp Biocomput 433-443 (2010) describe T7 promoter variants that satisfy both of these characteristics and methods for finding them.

[0257] Exemplary sequences of 5'UTRs and 3'UTRs for use in nucleic acids encoding the recombinant polypeptides described herein are shown in Tables N1, E12, and E16. In some embodiments, the nucleic acid encoding the recombinant polypeptide (e.g., RNA) further comprises a poly(A) tail for enhancing polypeptide expression.

[0258] In some embodiments, the poly(A) tail is made up solely of adenosine ribonucleotides. In some embodiments, the poly(A) tail includes, for example, non-adenosine ribonucleotides scattered within / between the regions of polyadenosine.

[0259] In some embodiments, the poly(A)tail consists of a nucleic acid sequence comprising a pattern of 12 to 18 adenosine nucleotides followed by at least 2 non-adenosine nucleotides, 22 to 28 adenosine nucleotides followed by at least 2 non-adenosine nucleotides, 32 to 38 adenosine nucleotides followed by at least 2 non-adenosine nucleotides, and 42 to 48 adenosine nucleotides.

[0260] In some embodiments, the poly(A)tail consists of a nucleic acid sequence comprising a pattern of 13 to 17 adenosine nucleotides followed by 1 to 4 non-adenosine nucleotides, 23 to 27 adenosine nucleotides followed by 1 to 5 non-adenosine nucleotides, 33 to 37 adenosine nucleotides followed by 2 to 6 non-adenosine nucleotides, and 43 to 47 adenosine nucleotides.

[0261] In some embodiments, the poly(A)tail comprises a nucleic acid sequence having a pattern of 14 to 16 adenosine nucleotides followed by at least 2 (e.g., 2 to 6) non-adenosine nucleotides, 24 to 26 adenosine nucleotides followed by at least 2 (e.g., 2 to 6) non-adenosine nucleotides, 34 to 36 adenosine nucleotides followed by at least 2 (e.g., 2 to 6) non-adenosine nucleotides, and 44 to 46 adenosine nucleotides.

[0262] In some embodiments, the poly(A)tail comprises a nucleic acid sequence having a pattern of 15 adenosine nucleotides followed by at least 2 (e.g., 2-4) non-adenosine nucleotides, 25 adenosine nucleotides followed by at least 2 (e.g., 2-4) non-adenosine nucleotides, 35 adenosine nucleotides followed by at least 2 (e.g., 2-4) non-adenosine nucleotides, and 45 adenosine nucleotides.

[0263] In some embodiments, the poly(A)tail consists of a nucleic acid sequence comprising a pattern of 15 adenosine nucleotides followed by 1 to 4 non-adenosine nucleotides, 25 adenosine nucleotides followed by 1 to 5 non-adenosine nucleotides, 35 adenosine nucleotides followed by 2 to 6 non-adenosine nucleotides, and 45 adenosine nucleotides.

[0264] In some embodiments, there are 2 to 10 non-adenosine nucleotides. In some embodiments, the number of non-adenosine nucleotides in each group of non-adenosine nucleotides increases from 5' to 3'. In some embodiments, the number of non-adenosine nucleotides in each group of non-adenosine nucleotides from 5' to 3' is 2, 3, and 4, respectively. In some embodiments, the non-adenosine nucleotides are selected from cytosine nucleotides or uridine nucleotides.

[0265] Exemplary sequences of poly(A) tails for use in nucleic acids encoding the recombinant polypeptides described herein are shown in Tables N1, E12, and E16.

[0266] [Table N1-1]

[0267] [Table N1-2]

[0268] In some embodiments, the nucleic acid encoding the recombinant polypeptide includes a nucleic acid sequence listed in Table N2, E11, or E15, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the nucleic acid encoding the recombinant polypeptide includes the nucleic acid sequence of SEQ ID NO: 40, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the nucleic acid encoding the recombinant polypeptide includes the nucleic acid sequence of SEQ ID NO: 101, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the nucleic acid encoding the recombinant polypeptide includes the nucleic acid sequence of SEQ ID NO: 102, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the nucleic acid encoding the recombinant polypeptide includes the nucleic acid sequence of SEQ ID NO: 106 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.

[0269] [Table N2-1]

[0270] [Table N2-2]

[0271] [Table N2-3]

[0272] [Table N2-4]

[0273] [Table N2-5]

[0274]

Table N2-6

[0275]

Table N2-7

[0276]

Table N2-8

[0277] Table N2-9

[0278]

Table N2-10

[0279]

Table N2-11

[0280]

Table N2-12

[0281]

Table N2-13

[0282]

Table N2-14

[0283]

Table N2-15

[0284] [Table N2-16]

[0285] [Table N2-17]

[0286] template nucleic acid The genetic recombination systems described herein can modify host target DNA sites using template nucleic acid sequences. In some embodiments, the genetic recombination systems described herein transcribe the RNA sequence template to the host target DNA site by target-prime reverse transcription (TPRT). By directly recombining the DNA sequence with the host genome via reverse transcription of the RNA sequence template, the genetic recombination systems can insert the target sequence into the target genome without the need to introduce an exogenous DNA sequence into the host cell (unlike, for example, the CRISPR system), thus eliminating the exogenous DNA insertion step. The genetic recombination systems can also delete sequences from the target genome or introduce substitutions using the target sequence. Thus, the genetic recombination systems provide a platform for the use of customized RNA sequence templates, including the target sequence, for example, a sequence containing heterologous gene coding and / or functional information.

[0287] In some embodiments, the template nucleic acid includes one or more sequences (e.g., two sequences) that bind to the recombinant polypeptide.

[0288] In some embodiments, the system or method described herein comprises a single template nucleic acid (e.g., template RNA). In some embodiments, the system or method described herein comprises multiple template nucleic acids (e.g., template RNA). For example, the system described herein comprises a first RNA comprising (e.g., from 5' to 3') a sequence that binds to a recombinant polypeptide (e.g., a DNA-binding domain and / or an endonuclease domain, e.g., gRNA) and a sequence that binds to a target site (e.g., the second strand of a site in a target genome), and a second RNA (e.g., template RNA) optionally comprising (e.g., from 5' to 3') a sequence that binds to a recombinant polypeptide (e.g., specifically to the RT domain), a heterologous target sequence, and a PBS sequence. In some embodiments, if the system comprises multiple nucleic acids, each nucleic acid comprises a conjugate domain. In some embodiments, the conjugate domain enables the association of nucleic acid molecules, for example, by hybridization of complementary sequences. For example, in some embodiments, the first RNA comprises a first conjugate domain, the second RNA comprises a second conjugate domain, and the first and second conjugate domains can hybridize with each other, for example, under stringent conditions. In some embodiments, stringent conditions in hybridization include hybridization at approximately 65°C in 4x sodium chloride / sodium citrate (SSC) followed by washing at approximately 65°C in 1x SSC.

[0289] In some embodiments, the template nucleic acid includes RNA. In some embodiments, the template nucleic acid includes DNA (e.g., single-stranded or double-stranded DNA).

[0290] In some embodiments, the template nucleic acid includes one or more (e.g., two) homology domains that are homologous to the target sequence. In some embodiments, the homology domains are about 10–20, 20–50, or 50–100 nucleotides long.

[0291] In some embodiments, the template RNA may include, for example, a gRNA sequence for directing a recombinant polypeptide to a target site of interest. In some embodiments, the template RNA includes (e.g., 5' to 3') (i) optionally a gRNA spacer that binds to a target site (e.g., the second strand of a site in the target genome), (ii) optionally a gRNA scaffold that binds to a polypeptide described herein (e.g., a recombinant polypeptide or Cas polypeptide), (iii) a heterologous target sequence including a mutation region (optionally, the heterologous target sequence includes a first homology region, a mutation region and a second homology region from 5' to 3'), and (iv) a primer-binding site (PBS) sequence including a 3' target homology domain.

[0292] The template nucleic acid (e.g., template RNA) component of the genome editing systems described herein is typically bindable to the recombinant polypeptide of the system. In some embodiments, the template nucleic acid (e.g., template RNA) has a 3' region that is bindable to the recombinant polypeptide. The binding region, e.g., the 3' region, may be a structured RNA region having, for example, at least one, two, or three hairpin loops that is bindable to the recombinant polypeptide of the system. The binding region can associate the template nucleic acid (e.g., template RNA) with any of the polypeptide modules. In some embodiments, the binding region of the template nucleic acid (e.g., template RNA) may associate with an RNA-binding domain in the polypeptide. In some embodiments, the binding region of the template nucleic acid (e.g., template RNA) may associate with a reverse transcription domain of the recombinant polypeptide (e.g., specifically binding to the RT domain). In some embodiments, the template nucleic acid (e.g., template RNA) may associate with a DNA-binding domain of the polypeptide; for example, gRNA may associate with a Cas9-derived DNA-binding domain. In some embodiments, the binding region may also provide DNA target recognition; for example, the gRNA hybridizes to a target DNA sequence and binds to a polypeptide, such as the Cas9 domain. In some embodiments, a template nucleic acid (e.g., template RNA) may associate with multiple components of the polypeptide, such as a DNA-binding domain and a reverse transcription domain.

[0293] In some embodiments, the template RNA has a poly(A) tail at its 3' end. In some embodiments, the template RNA does not have a poly(A) tail at its 3' end.

[0294] In some embodiments, the template nucleic acid is template RNA. In some embodiments, the template RNA contains one or more modified nucleotides. For example, in some embodiments, the template RNA contains one or more deoxyribonucleotides. In some embodiments, a region of the template RNA is replaced with DNA nucleotides, for example, to enhance molecular stability. For example, the 3' end of the template may contain DNA nucleotides, while the rest of the template contains RNA nucleotides that can be reverse transcribed. For example, in some embodiments, the heterologous target sequence consists mainly or entirely of RNA nucleotides (e.g., at least 90%, 95%, 98%, or 99% RNA nucleotides). In some embodiments, the PBS sequence consists mainly or entirely of DNA nucleotides (e.g., at least 90%, 95%, 98%, or 99% DNA nucleotides). In other embodiments, the heterologous target sequence for writing to the genome may contain DNA nucleotides. In some embodiments, the DNA nucleotides in the template are replicated into the genome by a domain capable of DNA-dependent DNA polymerase activity. In some embodiments, DNA-dependent DNA polymerase activity is provided by a DNA polymerase domain in the polypeptide. In some embodiments, DNA-dependent DNA polymerase activity is provided by a reverse transcriptase domain that also has the ability to perform DNA-dependent DNA polymerization, such as second strand synthesis. In some embodiments, the template molecule consists solely of DNA nucleotides.

[0295] In some embodiments, the system described herein comprises two nucleic acids, both containing the sequence of the template RNA described herein. In some embodiments, the two nucleic acids associate with each other non-covalently, for example, by direct association (e.g., by base pairing), or indirectly as part of a complex comprising one or more further molecules.

[0296] The template RNA described herein may include, from 5' to 3', (1) a gRNA spacer, (2) a gRNA scaffold, (3) a heterologous target sequence, and (4) a primer-binding site (PBS) sequence. Each of these components is described in more detail here.

[0297] gRNA spacer and gRNA scaffold The template RNA described herein may include a gRNA spacer that directs the recombination system to the target nucleic acid and a gRNA scaffold that facilitates the association of the template RNA with the Cas domain of the recombinant polypeptide. The system described herein may also include gRNA that is not part of the template nucleic acid. For example, a gRNA that includes a gRNA spacer and a gRNA scaffold but does not contain a heterologous target sequence or PBS sequence may be used to induce second strand nicking, for example, as described herein in the section titled “Second Strand Nicking”.

[0298] In some embodiments, gRNA is a short synthetic RNA consisting of a scaffold sequence involved in the binding of CRISPR-related proteins and a user-defined target sequence of approximately 20 nucleotides for a genomic target. The structure of a complete gRNA is described in Nishimasu et al. Cell 156, pp. 935-949 (2014). A gRNA (also referred to as sgRNA relative to a single guide RNA) consists of crRNA-derived and tracrRNA-derived sequences linked by an artificial tetraloop. The crRNA sequence can be divided into a guide (20nt) and repeat (12nt) region, while the tracrRNA sequence can be divided into an anti-repeat (14nt) and three tracrRNA stem-loops (Nishimasu et al. Cell 156, pp. 935-949 (2014)). In practice, the guide RNA sequence is generally designed to have a length of 17–24 nucleotides (e.g., 19, 20, or 21 nucleotides) and to be complementary to the target nucleic acid sequence. Custom gRNA generators and algorithms are commercially available for use in the design of effective guide RNAs. In some embodiments, the gRNA comprises two RNA components from the natural CRISPR system, e.g., crRNA and tracrRNA. As is well known in the art, the gRNA may also include a chimeric single RNA (sgRNA) containing sequences from both tracrRNA (for nuclease binding) and at least one crRNA (for leading the nuclease to a targeted sequence for editing / binding). Chemically modified sgRNAs have also been proven effective for use with CRISPR-related proteins; see, for example, Hendel et al. (2015) Nature Biotechnol., 985-991. In some embodiments, the gRNA spacer comprises a nucleic acid sequence complementary to the DNA sequence associated with the target gene.

[0299] In some embodiments, the template nucleic acid containing gRNA, e.g., the template RNA region, adopts an underwound ribbon-like structure of the gRNA bound to the target DNA (as described, e.g., Mulepati et al. Science 19 Sep 2014: Vol.345, Issue 6203, pp.1479-1484). While not intended to be bound to any particular theory, this non-standard structure is thought to be facilitated by a rotation of every six nucleotides from the RNA-DNA hybrid. Thus, in some embodiments, the template nucleic acid containing gRNA, e.g., the template RNA region, may tolerate an increase in mismatches with the target site at some intervals, e.g., every six bases. In some embodiments, the template nucleic acid containing gRNA with homology to the target site, e.g., the template RNA region, may have fluctuating positions at regular intervals, e.g., every six bases, where base pairing with the target site is not required.

[0300] In some embodiments, the template nucleic acid (e.g., template RNA) includes, for example, a gRNA spacer sequence of appropriate length to the Cas9 domain of a recombinant polypeptide (Table 8), and has, for example, at its 5' end, at least 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 bases with at least 80%, 85%, 90%, 95%, 99%, or 100% homology to the target site.

[0301] In some embodiments, Cas9 derivatives having enhanced activity may be used in recombinant polypeptides. In some embodiments, the Cas9 derivative may include mutations that improve the activity of the HNH endonuclease domain (see, for example, Spencer and Zhang Sci Rep 7:16836 (2017), where Cas9 derivatives and the mutations contained herein are incorporated herein by reference). In some embodiments, the Cas9 derivative may include one or more types of mutations described herein, e.g., PAM modification mutations, protein stabilization mutations, activity-enhancing mutations, and / or mutations that partially or completely inactivate one or more endonuclease domains compared to the parent enzyme (e.g., one or more mutations that eliminate endonuclease activity for one or both strands of target DNA, e.g., nickase or a catalytically inactive enzyme). In some embodiments, the Cas9 enzyme used in the systems described herein may include mutations that confer nickase activity to the enzyme, in addition to mutations that improve catalytic efficiency. Table 12 defines the components for designing gRNA and / or template RNA and provides parameters for applying Cas variants listed in Table 8 for genetic recombination. The cleavage site indicates the requirements of the validated or predicted protospacer adjacent motif (PAM) and the location of the validated or predicted cleavage site (relative to the upstream base of the PAM site). A gRNA for a given enzyme can be constructed by ligating crRNA, tetraloop, and tracrRNA sequences and further adding 5' spacers of a length within the minimum and maximum spacer lengths that match the protospacer at the target site. Furthermore, the predicted location of the ssDNA nick at the target is important for designing a PBS sequence of template RNA that can anneal to the sequence immediately 5' to the nick in order to initiate targeted priming reverse transcription. In some embodiments, the gRNA scaffold described herein includes a nucleic acid sequence in the 5' to 3' direction that includes the crRNA from Table 12, a tetraloop from the same row in Table 12, and a tracrRNA from the same row in Table 12, or a sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity thereto.In some embodiments, the gRNA or template RNA containing a scaffold further comprises a gRNA spacer having a length within the minimum and maximum spacers shown in the same row of Table 12. In some embodiments, the gRNA or template RNA having a sequence according to Table 12 is included in a system further comprising a recombinant polypeptide, the recombinant polypeptide comprising a Cas domain as shown in the same row of Table 12.

[0302] [Table 12-1]

[0303] [Table 12-2]

[0304] [Table 12-3]

[0305] In this specification, when an RNA sequence (e.g., a template RNA sequence) is described as containing a specific sequence containing thymine (T) (e.g., the sequences or portions thereof in Table 12), it is understood that the RNA sequence may (and often) contain uracil (U) instead of T. For example, the RNA sequence may contain U at all positions indicated as T in the sequences in Table 12. More specifically, this disclosure provides RNA sequences that conform to all gRNA scaffold sequences in Table 12, where the RNA sequences have U instead of each T in the sequences in Table 12. Furthermore, it is understood that terminal Us and Ts may be optionally added or removed from the tracrRNA sequence, and that when provided as RNA, it may be modified or unmodified. While not intended to be bound to specific examples, alternative gRNA scaffold sequence forms to those exemplified in Table 12, such as alternative gRNA scaffold sequences with nucleotide additions, substitutions, or deletions, or sequences with added or removed stem-loop structures, may also function with different Cas9 enzymes or their derivatives exemplified in Table 8. It is conceivable herein that gRNA scaffold sequences represent components of a recombination system that can similarly be optimized for a given system, Cas-RT fusion polypeptide, directive, target mutation, template RNA, or delivery vehicle.

[0306] Heterogeneous sequencing The template RNA described herein may include heterologous target sequences that recombinant polypeptides can use as templates for reverse transcription to write a desired sequence into a target nucleic acid. In some embodiments, the heterologous target sequence includes a post-edit homology region, a mutation region, and a pre-edit homology region from 5' to 3'. While not intended to be bound by any particular theory, the reverse transcription (RT) performing reverse transcription on the template RNA first reverse transcribes the pre-edit homology region, then the mutation region, and then the post-edit homology region, thereby generating a DNA strand containing the desired mutation with homology regions on both sides.

[0307] In some embodiments, the heteromorphic target array has a length of at least 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 8 2, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 120, 140, 160, 180, 200, 500 or 1,000 nucleotides (nts) or with a length of at least 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5 or 10 kilobases. In some embodiments, the heterogeneous target arrays have lengths of 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78 , 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 120, 140, 160, 180, 200, 500, 1,000 or 2000 nucleotides (nt) or less in length of 20, 15, 10, 9, 8, 7, 6, 5, 4 or 3 kilobases or less.In some embodiments, heterogeneous sequences have lengths of 30-1000, 40-1000, 50-1000, 60-1000, 70-1000, 74-1000, 75-1000, 76-1000, 77-1000, 78-1000, 79-1000, 80-1000, 85-1000, 90-1000, 100-1000, 120-1000, 140-1000, 160-1000, 180-1000, 200-1000, 500-1000, 30-500, 40-500, 50-500, 60-500, 7 0-500, 74-500, 75-500, 76-500, 77-500, 78-500, 79-500, 80-500, 85-500, 90-500, 100-500, 120-500, 140-500, 160-500, 180-500, 200-500, 30-200, 40-200, 50-200, 60-200, 70-200, 74-200, 75-200, 76-200, 77-200, 78-200, 79-200, 80-200, 85-200, 90-200, 100-200, 120 ~200, 140~200, 160~200, 180~200, 30~100, 40~100, 50~100, 60~100, 70~100, 74~100, 75~100, 76~100, 77~100, 78~100, 79~100, 80~100, 85~100 or 90~100 nucleotides (nt) or lengths of 1~20, 1~15, 1~10, 1~9, 1~8, 1~7, 1~6, 1~5, 1~4, 1~3, 1~2, 2~20, 2~15, 2~10, 2~9, 2~8, 2~7, 2~6, 2~5, The base weights are 2-4, 2-3, 3-20, 3-15, 3-10, 3-9, 3-8, 3-7, 3-6, 3-5, 3-4, 4-20, 4-15, 4-10, 4-9, 4-8, 4-7, 4-6, 4-5, 5-20, 5-15, 5-10, 5-9, 5-8, 5-7, 5-6, 6-20, 6-15, 6-10, 6-9, 6-8, 6-7, 7-20, 7-15, 7-10, 7-9, 7-8, 8-20, 8-15, 8-10, 8-9, 9-20, 9-15, 9-10, 10-15, 10-20, or 15-20 kilograms.In some embodiments, the heterologous target sequence has a length of 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 10-40, 10-30, or 10-20 nt, for example, a length of 10-80, 10-50, or 10-20 nt, for example, a length of approximately 10-20 nt. In some embodiments, the heterologous target sequence has a length of 8-30, 9-25, 10-20, 11-16, or 12-15 nucleotides, for example, a length of 11-16 nt. While not intended to be bound by any particular theory, in some embodiments, a larger insertion size of the edit, a larger region (e.g., the distance between the first edit / substitution and the second edit / substitution in the target region), and / or a greater number of desired edits (e.g., a mismatch of the heterologous target sequence to the target genome) may result in a longer optimal heterologous target sequence.

[0308] In certain embodiments, the template nucleic acid includes a customized RNA sequence template that can be identified, designed, modified, and constructed to include sequences that modify or specify host genome function, for example, by introducing heterologous coding regions into the genome, affecting or causing exon structure / alternative splicing, resulting in, for example, exon skipping of one or more exons, causing disruption of endogenous genes, for example, generating gene knockout, causing transcriptional activation of endogenous genes, causing epigenetic regulation of endogenous DNA, causing upregulation of one or more operably linked genes, for example, resulting in gene activation or overexpression, and causing downregulation of one or more operably linked genes, for example, generating gene knockdown. In certain embodiments, the customized RNA sequence template may be modified with intention to include binding sites and combinations thereof for transcription factor activators, repressors, enhancers, etc., to include sequences encoding exons and / or transgenes. In some embodiments, the customized template may be modified to encode a nucleic acid or peptide tag expressed in an endogenous RNA transcript or endogenous protein operably ligated to a target site. In other embodiments, the coding sequence may be further customized with a splice donor site, a splice acceptor site, or a poly-A tail.

[0309] The system's template nucleic acid (e.g., template RNA) typically includes a target sequence (e.g., heterologous target sequence) for writing a desired sequence into target DNA. The target sequence may be a coding or non-coding sequence. The template nucleic acid (e.g., template RNA) may be designed to induce an insertion, mutation, or deletion at a target DNA locus. In some embodiments, the template nucleic acid (e.g., template RNA) may be designed to induce an insertion in the target DNA. For example, the template nucleic acid (e.g., template RNA) may contain a heterologous sequence, and reverse transcription would result in the insertion of the heterologous sequence into the target DNA. In other embodiments, the RNA template may be designed to introduce a deletion into the target DNA. For example, the template nucleic acid (e.g., template RNA) may match the target DNA upstream and downstream of a desired deletion, and reverse transcription would result in the replication of the upstream and downstream sequences from a template nucleic acid (e.g., template RNA) without an intervening sequence, resulting in, for example, the deletion of an intervening sequence. In other embodiments, a template nucleic acid (e.g., template RNA) may be designed to introduce edits into target DNA. For example, the template RNA may match the target DNA sequence with one or more nucleotides as an exception, and reverse transcription would result in the replication of these edits into the target DNA, leading to mutations, such as transition or transversion mutations.

[0310] In some embodiments, writing the target sequence to a target site results in nucleotide substitutions, for example, the full length of the target sequence corresponds to the matching length of the target site having one or more mismatched bases. In some embodiments, heterologous target sequences may be designed so that combinations of sequence mutations exist, such as simultaneous addition and deletion, addition and substitution, or deletion and substitution.

[0311] In some embodiments, the heterologous target sequence may contain an open reading frame or a fragment of an open reading frame. In some embodiments, the heterologous target sequence has a Kozak sequence. In some embodiments, the heterologous target sequence has an internal ribosome entry site. In some embodiments, the heterologous target sequence has a self-cleaving peptide such as a T2A or P2A site. In some embodiments, the heterologous target sequence has a start codon. In some embodiments, the template RNA has a splice acceptor site. In some embodiments, the template RNA has a splice donor site. Exemplary splice acceptor and splice donor sites are described in International Publication No. 2016044416 (which is incorporated herein by reference in its entirety). Exemplary splice acceptor site sequences are known to those skilled in the art. In some embodiments, the template RNA has a microRNA binding site downstream of a stop codon. In some embodiments, the template RNA has a poly-A tail downstream of the stop codon of the open reading frame. In some embodiments, the template RNA includes one or more exons. In some embodiments, the template RNA includes one or more introns. In some embodiments, the template RNA includes a eukaryotic transcription terminator. In some embodiments, the template RNA includes an enhanced translation element or a translation enhancement element. In some embodiments, the RNA includes a human T-cell leukemia virus (HTLV-1) R region. In some embodiments, the RNA includes a post-transcriptional regulatory sequence that enhances nuclear export, such as that of hepatitis B virus (HPRE) or woodchuck hepatitis virus (WPRE).

[0312] In some embodiments, the heterologous target sequence may contain a non-coding sequence. For example, the template nucleic acid (e.g., template RNA) may include regulatory elements, such as promoter or enhancer sequences or miRNA binding sites. In some embodiments, integration of the target sequence at the target site results in upregulation of endogenous genes. In some embodiments, integration of the target sequence at the target site results in downregulation of endogenous genes. In some embodiments, the template nucleic acid (e.g., template RNA) includes a tissue-specific promoter or enhancer, each of which may be unidirectional or bidirectional. In some embodiments, the promoter is an RNA polymerase I promoter, an RNA polymerase II promoter, or an RNA polymerase III promoter. In some embodiments, the promoter includes a TATA element. In some embodiments, the promoter includes a B recognition element. In some embodiments, the promoter has one or more binding sites for transcription factors.

[0313] In some embodiments, the template nucleic acid (e.g., template RNA) includes a site that links epigenetic modifications. In some embodiments, the template nucleic acid (e.g., template RNA) includes a chromatin insulator. For example, the template nucleic acid (e.g., template RNA) includes a CTCF site or a site targeted for DNA methylation.

[0314] In some embodiments, the template nucleic acid (e.g., template RNA) includes a gene expression unit comprising at least one regulatory region operably ligated to an effector sequence. The effector sequence may be a sequence transcribed to RNA (e.g., a coding sequence or a non-coding sequence, such as a sequence encoding microRNA).

[0315] In some embodiments, a heterologous target sequence of a template nucleic acid (e.g., template RNA) is inserted into the target genome within an endogenous intron. In some embodiments, the heterologous target sequence of a template nucleic acid (e.g., template RNA) is inserted into the target genome and thereby acts as a new exon. In some embodiments, the insertion of the heterologous target sequence into the target genome results in the substitution or skipping of a native exon.

[0316] A template nucleic acid (e.g., template RNA) may be designed to induce an insertion, mutation, or deletion at a target DNA locus. In some embodiments, the template nucleic acid (e.g., template RNA) may be designed to induce an insertion into the target DNA. For example, the template nucleic acid (e.g., template RNA) may contain a heterologous target sequence, and reverse transcription would result in the insertion of the heterologous target sequence into the target DNA. In other embodiments, the RNA template may be designed to write a deletion into the target DNA. For example, the template nucleic acid (e.g., template RNA) may match the target DNA upstream and downstream of a desired deletion, and reverse transcription would result in the replication of the upstream and downstream sequences from a template nucleic acid (e.g., template RNA) without an intervening sequence, resulting in, for example, the deletion of the intervening sequence. In other embodiments, the template nucleic acid (e.g., template RNA) may be designed to write an edit into the target DNA. For example, template RNA may match the target DNA sequence with the exception of one or more nucleotides, and reverse transcription would result in the replication of these edits to the target DNA, leading to mutations, such as transition or transversion mutations.

[0317] In some embodiments, the pre-editing homology domain includes a nucleic acid sequence having 100% sequence identity with the nucleic acid sequence contained in the target nucleic acid molecule.

[0318] In some embodiments, the post-editing homology domain includes a nucleic acid sequence that has 100% sequence identity with the nucleic acid sequence contained in the target nucleic acid molecule.

[0319] PBS sequence In some embodiments, the template nucleic acid (e.g., template RNA) includes a PBS sequence. In some embodiments, the PBS sequence is positioned 3' to the heterologous target sequence and is complementary to a sequence adjacent to the site to be modified by the system described herein, or contains 1, 2, 3, 4, or 5 or fewer mismatches to a sequence complementary to a sequence adjacent to the site to be modified by the system / recombinant polypeptide. In some embodiments, the PBS sequence binds to within 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides of a nick site in the target nucleic acid molecule. In some embodiments, the binding of the PBS sequence to the target nucleic acid molecule enables the initiation of targeted prime reverse transcription (TPRT) by, for example, a 3' homology domain acting as a primer for TPRT. In some embodiments, the PBS sequence has lengths of 3-5, 5-10, 10-30, 10-25, 10-20, 10-19, 10-18, 10-17, 10-16, 10-15, 10-14, 10-13, 10-12, 10-11, 11-30, 11-25, 11-20, 11-19, 11-18, 11- 17, 11-16, 11-15, 11-14, 11-13, 11-12, 12-30, 12-25, 12-20, 12-19, 12-18, 12-17, 12-16, 12-15, 12-14, 12-13, 13-30, 13-25, 13-20, 13-19, 13-18, 13-17, 13-16, 1 3-15, 13-14, 14-30, 14-25, 14-20, 14-19, 14-18, 14-17, 14-16, 14-15, 15-30, 15-25, 15-20, 15-19, 15-18, 15-17, 15-16, 16-30, 16-25, 16-20, 16-19, 16-18, 16-17 , 17-30, 17-25, 17-20, 17-19, 17-18, 18-30, 18-25, 18-20, 18-19, 19-30, 19-25, 19-20, 20-30, 20-25 or 25-30 nucleotides, for example, with lengths of 10-17, 12-16 or 12-14 nucleotides. In some embodiments, the PBS sequence has lengths of 5-20, 8-16, 8-14, 8-13, 9-13, 9-12 or 10-12 nucleotides, for example, with lengths of 9-12 nucleotides.

[0320] The template nucleic acid (e.g., template RNA) may have some homology to the target DNA. In some embodiments, the template nucleic acid (e.g., template RNA) is positioned to prime the reverse transcription of the template nucleic acid (e.g., template RNA) by having a PBS sequence domain that can function as an annealing region to the target DNA. In some embodiments, the template nucleic acid (e.g., template RNA) has exact homology to the target DNA at the 3' end of the RNA by at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 175, 200 or more bases. In some embodiments, the template nucleic acid (e.g., template RNA) has, for example, at the 5' end of the template nucleic acid (e.g., template RNA) at least 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% homology to the target DNA of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 140, 150, 175, 200 or more bases.

[0321] Exemplary template arrangement In some embodiments of the systems and methods described herein, the template RNA comprises a gRNA spacer containing nucleotides of the gRNA spacer sequence in Table 1A. In some embodiments, the heterologous target sequence comprises nucleotides of the RT template sequence in Table 1A corresponding to the gRNA spacer sequence. In relation to the sequence listing, the first component corresponds to the second component if both components are in the same row in the table referenced. In some embodiments, the primer binding site (PBS) sequence comprises a sequence containing nucleotides of the PBS sequence from the same row as the RT template sequence in Table 1A.

[0322] Table 1A: Exemplary template RNA (tgRNA) for correcting pathogenic R408W mutations Table 1A provides exemplary component designs for a recombination system to modify the pathogenic R408W mutation in the PAH gene to the wild-type morphology. This table details the sequences of a complete template RNA (tgRNA) including (1) a gRNA spacer (e.g., for targeting a first strand nick), (2) a gRNA scaffold, (3) an RT (heterotarget sequence) sequence, and (4) a PBS sequence (e.g., for initiating TPRT at the first strand nick). The template in this table uses the scaffold GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 21).

[0323] [Table 1A-1]

[0324] [Table 1A-2]

[0325] [Table 1A-3]

[0326] In this specification, when an RNA sequence (e.g., a template RNA sequence) is described as containing a specific sequence containing uracil (U) (e.g., the sequences or portions thereof in Table 1A), it is understood that the RNA sequence may (and often) contain thymine (T) instead of U. For example, the RNA sequence may contain T at all positions indicated as U in the sequences in Table 1A. More specifically, this disclosure provides RNA sequences that conform to all the gRNA spacer sequences shown in Table 1A, wherein the RNA sequences have T instead of U in each of the sequences in Table 1A.

[0327] In some embodiments of the systems and methods described herein, the system further includes a second-strand targeting gRNA (ngRNA) that leads the nick to the second strand of the human PAH gene. In some embodiments, the second-strand targeting gRNA includes a spacer sequence of the ngRNA from Table 2A or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.

[0328] Table 2A: Exemplary second nicked gRNA (ngRNA) Table 2A provides exemplary second nic gRNA (ngRNA) species for any use in correcting the pathogenic R408W mutation in PAH. These ngRNAs may be used in combination with the tgRNAs listed in Table 1A. The template in this table uses the scaffold GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 21).

[0329] [Table 2A]

[0330] In some embodiments, the systems and methods provided herein may include any second nick gRNA sequence listed in Table 2A, designed to be combined with a template sequence listed in Table 1A and a recombinant polypeptide for correcting mutations in the PAH gene. The template RNA sequences shown in Tables 1A and 2A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3 may be customized depending on the target cell. For example, in some embodiments, it is necessary to inactivate the PAM sequence at the time of editing (e.g., by using a "PAM-kill" modification) to reduce the possibility of further gene editing after the initial editing (e.g., by Cas retargeting). As a result, certain template RNAs described herein are designed to write mutations (e.g., substitutions) to the PAM at the target site, thereby mutating the PAM site at the time of editing so that the sequence is no longer recognized by the recombinant polypeptide. Therefore, mutated regions within the heterologous target sequence of template RNA may include PAM inactivation sequences. While we do not wish to be constrained by theory, in some embodiments, PAM inactivation sequences prevent re-engagement of the recombinant polypeptide upon completion of recombination, or reduce re-engagement compared to template RNA lacking PAM inactivation sequences. In some embodiments, PAM inactivation sequences do not alter the amino acid sequence encoded by the gene; for example, PAM inactivation sequences result in silent mutations. In other embodiments, it is required to leave the PAM sequence intact (without PAM inactivation).

[0331] Similarly, in some embodiments, it may be desirable to modify the first three nucleotides of the RT template sequence via a “seed-kill” motif to reduce the possibility of further gene editing after the initial editing (e.g., by Cas retargeting). As a result, certain template RNAs described herein are designed to write mutations (e.g., substitutions) to a portion of the target site corresponding to the first three nucleotides of the RT template sequence, thereby, upon editing, the target site will be mutated to a sequence with lower homology to the RT template sequence. Thus, the mutated region within the heterologous target sequence of the template RNA may include a seed-kill sequence. While we do not wish to be constrained by theory, in some embodiments, the seed-kill sequence prevents re-engagement of the recombinant polypeptide upon completion of recombination, or reduces re-engagement compared to other similar template RNAs lacking a seed-kill sequence. In some embodiments, the seed-kill sequence does not modify the amino acid sequence encoded by the gene; for example, the seed-kill sequence results in a silent mutation. In other embodiments, it is necessary to leave the seed region intact, and no seed inactivation sequence is used.

[0332] In further embodiments, it may be desirable to circumvent or redirect the target cell's mismatch repair or nucleotide repair pathway, or the target cell's repair pathway, to preserve the edited strand, in order to optimize or improve gene editing efficiency. In some embodiments, multiple silent mutations (e.g., silent substitutions) may be introduced into the RT template sequence to circumvent or redirect the target cell's mismatch repair or nucleotide repair pathway, or the target cell's repair pathway, to preserve the edited strand.

[0333] Table 7A provides exemplary silent mutations at various locations within the PAH gene for use with a template to correct the R408W mutation.

[0334] [Table 7A]

[0335] In some embodiments, the template RNA contains one or more silent mutations.

[0336] Please understand that the silent mutations shown in Table 7A may be used individually with the template RNA sequences described herein, or combined in any way.

[0337] gRNA with inducible activity In some embodiments, the gRNA described herein (e.g., a gRNA that is part of a template RNA or a gRNA used for second-strand nicking) has inducible activity. Inducible activity may be achieved by a template nucleic acid, e.g., template RNA, further comprising a blocking domain (in addition to the gRNA), wherein the sequence of some or all of the blocking domain is at least partially complementary to some or all of the gRNA. Thus, the blocking domain is hybridizable or substantially hybridizable to some or all of the gRNA. In some embodiments, the blocking domain and the inducible activity gRNA are positioned on a template nucleic acid, e.g., template RNA, so that the gRNA can have a first conformation in which the blocking domain is hybridized or substantially hybridized to the gRNA, and a second conformation in which the blocking domain is not hybridized or substantially hybridized to the gRNA. In some embodiments, in the first conformation, the gRNA binds to a recombinant polypeptide (e.g., a template nucleic acid-binding domain, a DNA-binding domain, or an endonuclease domain (e.g., a CRISPR / Cas protein)) with substantially reduced affinity compared to a similar template RNA, except that it is unable to bind to the polypeptide or lacks a blocking domain. In some embodiments, in the second conformation, the gRNA is able to bind to the recombinant polypeptide (e.g., a template nucleic acid-binding domain, a DNA-binding domain, or an endonuclease domain (e.g., a CRISPR / Cas protein)). In some embodiments, whether the gRNA is in the first or second conformation may affect whether the DNA-binding or endonuclease activity of the recombinant polypeptide (e.g., the CRISPR / Cas protein contained in the recombinant polypeptide) is active.

[0338] In some embodiments, the gRNA linking the second nick has inducible activity. In some embodiments, the gRNA linking the second nick is induced after the template has been reverse transcribed. In some embodiments, hybridization of the gRNA to the blocking domain can be disrupted using an opening molecule. In some embodiments, the opening molecule includes a drug that binds to part or all of the gRNA or the blocking domain and inhibits hybridization of the gRNA to the blocking domain. In some embodiments, the opening molecule includes a nucleic acid containing a sequence that is partially or entirely complementary to the gRNA, the blocking domain, or both. By selecting or designing an appropriate opening molecule, the provision of an opening molecule may facilitate a change in the conformation of the gRNA so that it can associate with the CRISPR / Cas protein and provide the relevant function of the CRISPR / Cas protein (e.g., DNA binding and / or endonuclease activity). While not intended to be bound by any particular theory, the provision of an opening molecule at a selected time and / or location may enable spatial and temporal control of the activity of the gRNA, the CRISPR / Cas protein, or the recombination system containing them. In some embodiments, the opening molecule is exogenous to cells containing the recombinant polypeptide and / or template nucleic acid. In some embodiments, the opening molecule comprises an endogenous agent (e.g., endogenous to cells containing the recombinant polypeptide and / or template nucleic acid including gRNA and a blocking domain). For example, the inducible gRNA, blocking domain and opening molecule may be selected such that the opening molecule is an endogenous agent expressed in target cells or tissues, for example, thereby confirming the activity of the recombinant system in the target cells or tissues. As a further example, the inducible gRNA, blocking domain and opening molecule may be selected such that the opening molecule is absent or substantially not expressed in one or more non-target cells or tissues, for example, thereby confirming that the activity of the recombinant system is absent or substantially absent, or present at a reduced level compared to the target cells or tissues.Exemplary blocking domains, opening molecules, and their uses are described in the PCT application publication, International Publication No. 2020044039A1 (which is incorporated herein by reference in its entirety). In some embodiments, the template nucleic acid, e.g., template RNA, may include one or more sequences or structures and gRNA for binding by one or more components of the recombinant polypeptide, e.g., by reverse transcriptase or an RNA-binding domain. In some embodiments, the gRNA facilitates the interaction of the recombinant polypeptide with the template nucleic acid-binding domain (e.g., an RNA-binding domain). In some embodiments, the gRNA directs the recombinant polypeptide to a matching target sequence, e.g., in a target cell genome.

[0339] target nucleic acid site In some embodiments, after recombination, the target sites around the edited sequence contain a limited number of insertions or deletions in about 50% or less than 10% of the editing events, as determined, for example, by long-read amplicon sequencing of the target site, as described, for example, as described in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903 (the entirety of which is incorporated herein by reference). In some embodiments, the target site does not show multiple consecutive editing events, such as head-to-tail or head-to-head duplication, as determined, for example, by long-read amplicon sequencing of the target site, as described, for example, as described in Karst et al. bioRxiv doi.org / 10.1101 / 645903 (2020) (the entirety of which is incorporated herein by reference). In some embodiments, the target site contains an integrated sequence corresponding to the template RNA. In some embodiments, the target site does not contain an insertion derived from endogenous RNA in about 1% or more of events, as determined, for example, by long-read amplicon sequencing of the target site, as described in Karst et al.bioRxiv doi.org / 10.1101 / 645903(2020) (the whole of which is incorporated herein by reference). In some embodiments, the target site includes an embedded sequence corresponding to the template RNA.

[0340] In certain embodiments of the present invention, the host DNA binding site incorporated by the recombination system may be located within a gene, within an intron, within an exon, within an ORF, outside the coding region of any gene, within the regulatory region of a gene, or outside the regulatory region of a gene. In other embodiments, the polypeptide may bind to one or more host DNA sequences.

[0341] In some embodiments, a genetic engineering system is used to edit target loci in multiple alleles. In some embodiments, the genetic engineering system is designed to edit a specific allele. For example, a recombinant polypeptide may be directed to a specific sequence present only on one allele, and may include a template RNA homologous to a gRNA or annealing domain, rather than to the target allele, such as a second cognitive allele. In some embodiments, the genetic engineering system may modify a haplotype-specific allele. In some embodiments, a genetic engineering system targeting a specific allele preferentially targets that allele, for example, with at least 2, 4, 6, 8, or 10 times priority over the target allele.

[0342] Second chain nicking In some embodiments, the genetic recombination systems described herein include nickase activity for cleaving a first strand (e.g., in the recombinant polypeptide) and nickase activity for cleaving a second strand of target DNA (e.g., in a polypeptide separate from the recombinant polypeptide). While not intended to be bound by any particular theory as described herein, it is thought that nicking the first strand of target site DNA provides a 3'OH that can be used by the RT domain to reverse transcribe the template RNA sequence, e.g., a heterologous target sequence. While not intended to be bound by any particular theory, it is thought that the introduction of further nicks into the second strand may bias the cellular DNA repair mechanism to adopt heterologous target sequence-based sequences more frequently than the original genome sequence. In some embodiments, the further nick to the second strand is made by the same endonuclease domain (e.g., the nickase domain) as the nick to the first strand. In some embodiments, the same recombinant polypeptide performs both the nick to the first strand and the nick to the second strand. In some embodiments, the recombinant polypeptide comprises a CRISPR / Cas domain, and a further nick to the second strand is led by a further nucleic acid, for example, a second gRNA that guides the CRISPR / Cas domain to create a cleavage in the second strand. In other embodiments, the further second strand nick is made by an endonuclease domain (e.g., a nickase domain) different from the nick to the first strand. In some embodiments, the different endonuclease domain is located in the further polypeptide, separately from the recombinant polypeptide (e.g., the system of the present invention further comprises a further polypeptide). In some embodiments, the further polypeptide comprises an endonuclease domain (e.g., a nickase domain) as described herein. In some embodiments, the further polypeptide comprises, for example, a DNA-binding domain as described herein.

[0343] It is intended herein that, compared to the first strand nick, the location where the second strand nick occurs may affect one or more of the ranges in which desired recombinant DNA modifications are obtained, unwanted double-strand breaks (DSBs) occur, unwanted insertions occur, or unwanted deletions occur. While not intended to be bound by any particular theory, the second strand nick can occur in two common orientations: inward nicks and outward nicks.

[0344] In some embodiments, in an inward nick orientation, the RT domain polymerizes away from the second strand nick (e.g., using template RNA (e.g., a heterologous target sequence)). In some embodiments, in an inward nick orientation, the nick locations relative to the first strand and the nick locations relative to the second strand are located between the first and second PAM sites (e.g., in situations where both nicks are produced by polypeptides containing CRISPR / Cas domains (e.g., recombinant polypeptides)). When there are two PAMs on the outside and two nicks on the inside, this inward nick orientation may also be referred to as "PAM outward." In some embodiments, in an inward nick orientation, the nick locations relative to the first strand and the nick locations relative to the second strand are between the sites where the polypeptide and further polypeptides bind to target DNA. In some embodiments, in an inward nick orientation, the nick location relative to the second strand is located between the binding sites of the polypeptide and further polypeptides, and the nick relative to the first strand may also be located between the binding sites of the polypeptide and further polypeptides. In some embodiments, in an inward nick orientation, the nicks on the first chain and the nicks on the second chain are positioned between the PAM site and the binding site of the second polypeptide, which is at a distance from the target site.

[0345] An example of a recombinant system providing inward nick orientation comprises a recombinant polypeptide containing a CRISPR / Cas domain, a template RNA containing a gRNA that leads to nicking of a target site DNA on the first strand, and a further nucleic acid containing a further gRNA that leads to nicking at a site distant from the first nick site, wherein the sites of the first and second nicks are located between the PAM site where the two gRNAs lead the recombinant polypeptide. As a further example, another recombinant system providing inward nick orientation comprises a zinc finger molecule and a recombinant polypeptide containing a first nickase domain, a further polypeptide containing a CRISPR / Cas domain, and a further nucleic acid containing a gRNA that leads the further polypeptide to make a cleavage at a site distant from the target site DNA on the second strand, wherein the zinc finger molecule binds to the target DNA in such a way that it leads the first nickase domain to make a cleavage in the first strand of the target site, and the sites of the first and second nicks are located between the PAM site and the site where the zinc finger molecule binds. As a further example, another genetically modified system providing inward nick orientation comprises a zinc finger molecule and a recombinant polypeptide containing a first nickase domain, a TAL effector molecule and a further polypeptide containing a second nickase domain, wherein the zinc finger molecule binds to target DNA in such a way that it guides the first nickase domain to cleave the first strand of the target site, and the TAL effector molecule binds to a site at an angle to the target site in such a way that it guides the further polypeptide to cleave the second strand, and the locations of the first and second nicks are located between the binding site of the TAL effector molecule and the binding site of the zinc finger molecule.

[0346] In some embodiments, in outward nick orientation, the RT domain polymerizes toward the second strand nick (e.g., using template RNA (e.g., heterologous target sequence)). In some embodiments, in outward nick orientation, if both the first and second nicks are produced by polypeptides containing CRISPR / Cas domains (e.g., recombinant polypeptides), the first and second PAM sites are positioned between the nick locations on the first strand and the nick locations on the second strand. When there are two PAMs on the inside and two nicks on the outside, this outward nick orientation may also be referred to as "PAM inward." In some embodiments, in outward nick orientation, polypeptides (e.g., recombinant polypeptides) and further polypeptides bind to sites on target DNA between the nick locations on the first strand and the nick locations on the second strand. In some embodiments, in outward nick orientation, the nick locations on the second strand are positioned opposite the binding sites of the polypeptides and further polypeptides relative to the nick locations on the first strand. In some embodiments, in the outward orientation, the binding site of the second polypeptide, which is distanced from the PAM site and the target site, is positioned between the nick position relative to the first chain and the nick position relative to the second chain.

[0347] An example of a genetically modified system that provides outward nick orientation includes a genetically modified polypeptide containing a CRISPR / Cas domain, a template RNA containing a gRNA that leads to nicking of a target site DNA on the first strand, and a further nucleic acid containing a further gRNA that leads to nicking at a site distanced from the first nick site, where the first and second nick sites are located outside the PAM site of the site where the two gRNAs lead the genetically modified polypeptide (i.e., the PAM site is located between the first and second nick sites). As a further example, another genetically modified system providing outward nick orientation comprises a genetically modified polypeptide containing a zinc finger molecule and a first nickase domain, a further polypeptide containing a CRISPR / Cas domain, and a further nucleic acid containing a gRNA that guides the further polypeptide to make a cleavage at a site on the second strand of target DNA, wherein the zinc finger molecule binds to the target DNA in such a way that it guides the first nickase domain to make a cleavage in the first strand of the target site, and the sites of the first and second nicks are located outside the PAM site and the site where the zinc finger molecule binds (i.e., the PAM site and the site where the zinc finger molecule binds are located between the site of the first and second nicks). As a further example, another genetically modified system providing outward nick orientation comprises a recombinant polypeptide containing a zinc finger molecule and a first nickase domain, a TAL effector molecule and a further polypeptide containing a second nickase domain, wherein the zinc finger molecule binds to target DNA in such a way that it guides the first nickase domain to cleave the first strand of the target site, and the TAL effector molecule binds to a site at a distance from the target site in such a way that it guides the further polypeptide to cleave the second strand, and the locations of the first and second nicks are located outside the sites where the TAL effector molecule and the zinc finger molecule bind (i.e., the sites where the TAL effector molecule and the zinc finger molecule bind are located between the locations of the first and second nicks).

[0348] While not intended to be bound by any particular theory, in the case of recombination systems that provide a second strand nick, outward nick orientation is considered preferable in some embodiments. As described herein, inward nicks may generate a greater number of double-strand breaks (DSBs) than outward nick orientation. DSBs may be recognized by DSB repair pathways in the cell nucleus, resulting in unwanted insertions and deletions. Outward nick orientation may provide a reduced risk of DSB formation and a correspondingly smaller number of unwanted insertions and deletions. In some embodiments, unwanted insertions and deletions are insertions and deletions not encoded by the heterologous target sequence, e.g., insertions or deletions generated by double-strand break repair pathways unrelated to modifications encoded by the heterologous target sequence. In some embodiments, the desired recombination involves modifications (e.g., substitutions, insertions, or deletions) to target DNA encoded by the heterologous target sequence (e.g., achieved by recombination that writes the heterologous target sequence to a target site). In some embodiments, the first and second strand nicks are outwardly oriented.

[0349] Furthermore, the distance between the first strand nick and the second strand nick may affect one or more of the ranges over which desired recombination system DNA modifications are obtained, the range over which unwanted double-strand breaks (DSBs) occur, the range over which unwanted insertions occur, or the range over which unwanted deletions occur. While not intended to be bound by any particular theory, the second strand nick is considered advantageous because it increases the likelihood of biasing DNA repair toward the integration of heterologous target sequences into target DNA as the distance between the first and second strand nicks decreases. However, the risk of DSB formation is also considered to increase as the distance between the first and second strand nicks decreases. Accordingly, the number of unwanted insertions and / or deletions may increase as the distance between the first and second strand nicks decreases. In some embodiments, the distance between the first and second strand nicks is selected to balance the advantage of biasing DNA repair toward the integration of heterologous target sequences into target DNA with the risk of DSB formation and unwanted deletions and / or insertions. In some embodiments, when the first strand nick and the second strand nick are separated by at least a threshold distance, the system has a higher level of desired recombination system modification results, a lower level of unwanted deletions, and / or unwanted insertions, compared to otherwise similar inward nick orientation systems when the first and second nicks are separated by less than a threshold distance. In some embodiments, the threshold distance is given as follows:

[0350] In some embodiments, the first nick and the second nick are separated by at least 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides. In some embodiments, the first nick and the second nick are separated by no more than 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, or 250 nucleotides. In some embodiments, the first and second nicks are 20-200, 30-200, 40-200, 50-200, 60-200, 70-200, 80-200, 90-200, 100-200, 110-200, 120-200, 130-200, 140-200, 150-200, 160-200, 170-200, 180-200, 190-200, 20-190, 30-190, 40-190 50-190, 60-190, 70-190, 80-190, 90-190, 100-190, 110-190, 120-190, 130-190, 140-190, 150-190, 160-190, 170-190, 180-190, 20-180, 30-180, 40-180, 50-180, 60-180, 70-180, 80-180, 90-180, 100-180, 110-180, 12 0-180, 130-180, 140-180, 150-180, 160-180, 170-180, 20-170, 30-170, 40-170, 50-170, 60-170, 70-170, 80-170, 90-170, 100-170, 110-170, 120-170, 130-170, 140-170, 150-170, 160-170, 20-160, 30-160, 40-160, 50- 160, 60-160, 70-160, 80-160, 90-160, 100-160, 110-160, 120-160, 130-160, 140-160, 150-160, 20-150, 30-150, 40-150, 50-150, 60-150, 70-150, 80-150, 90-150, 100-150, 110-150, 120-150, 130-150, 140-150, 20-140,30-140, 40-140, 50-140, 60-140, 70-140, 80-140, 90-140, 100-140, 110-140, 120-140, 130-140, 20-130, 30-130, 40-130, 50-130, 60-130, 70-130, 80-130, 90- 130, 100-130, 110-130, 120-130, 20-120, 30-120, 40-120, 50-120, 60-120, 70-120, 80-120, 90-120, 100-120, 110-120, 20-110, 30-110, 40-110, 50-110, 60-110 , 70~110, 80~110, 90~110, 100~110, 20~100, 30~100, 40~100, 50~100, 60~100, 70~100, 80~100, 90~100, 20~90, 30~90, 40~90, 50~90, 60~90, 70~90, 80~90, 20~80, The nucleotides are separated by 30-80, 40-80, 50-80, 60-80, 70-80, 20-70, 30-70, 40-70, 50-70, 60-70, 20-60, 30-60, 40-60, 50-60, 20-50, 30-50, 40-50, 20-40, 30-40, or 20-30 nucleotides. In some embodiments, the first and second nicks are separated by 40-100 nucleotides.

[0351] While not intended to be bound by any particular theory, in the case of a recombination system in which a second strand nick is provided and inward nick orientation is selected, it is considered that an increase in the distance between the first strand nick and the second strand nick may be preferable. As described herein, inward nick orientation may produce a greater number of DSBs than outward nick orientation and may result in a greater amount of unwanted insertions and deletions than outward nick orientation, however, an increase in the distance between nicks may mitigate such an increase in DSBs, unwanted deletions and / or unwanted insertions. In some embodiments, inward nick orientation when the first and second nicks are at least a threshold distance apart has a greater level of desired recombination system modification results, a reduced level of unwanted deletions and / or unwanted insertions compared to an otherwise similar inward nick orientation system when the first and second nicks are less than a threshold distance apart. In some embodiments, the threshold distance is given below.

[0352] In some embodiments, the first and second strand nicks are inwardly oriented. In some embodiments, the first and second strand nicks are inwardly oriented and are separated by at least 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 220, 240, 260, 280, 300, 350, 400, 450, or 500 nucleotides, for example, at least 100 nucleotides (and optionally 500, 400, 300, 200, 190, 180, 170, 160, 150, 140, 130, or 120 nucleotides or less). In some embodiments, the first and second chain nicks are oriented inward, and the first and second chain nicks are oriented in the following order: 100-200, 110-200, 120-200, 130-200, 140-200, 150-200, 160-200, 170-200, 180-200, 190-200, 100-190, 110-190, 120-190, 130-190, 140-190, 150-190, 160-190, 170-190, 180-190, 100-180, 110-180, 120-180, 130-180, 140-180, 150-180 They are separated by 160-180, 170-180, 100-170, 110-170, 120-170, 130-170, 140-170, 150-170, 160-170, 100-160, 110-160, 120-160, 130-160, 140-160, 150-160, 100-150, 110-150, 120-150, 130-150, 140-150, 100-140, 110-140, 120-140, 130-140, 100-130, 110-130, 120-130, 100-120, 110-120, or 100-110 nucleotides.

[0353] Characteristics of chemically modified nucleic acids and nucleic acid ends Nucleic acids described herein (e.g., template nucleic acids, e.g., template RNA, or nucleic acids encoding recombinant polypeptides (e.g., mRNA) or gRNA) may include unmodified or modified nucleic acid bases. Naturally occurring RNA is synthesized from four basic ribonucleotides: ATP, CTP, UTP, and GTP, but may include post-transcriptionally modified nucleotides. Furthermore, approximately 100 different nucleoside modifications have been identified in RNA (Rozenski, J, Crain, P, and McCloskey, J. (1999). The RNA Modification Database: 1999 update. Nucleoside Acids Res 27: 196-197). RNA may also include entirely synthetic nucleotides that do not exist in nature.

[0354] Ribonucleosides having unmodified sugars containing the 2'OH group shown below. [ka]

[0355] Nucleosides containing sugars with 2'-fluoro(2'F) modification are shown below. [ka]

[0356] Nucleosides containing sugars with 2'-O-methyl (2'O-Me) modification are shown below. [ka]

[0357] The following are nucleotides containing sugars with phosphorothioate modification and 2'-O-methyl (2'O-Me) modification. [ka]

[0358] Template RNA sequences containing modified nucleotides in heteropurpose sequences and / or PBS sequences are described herein. In some embodiments, the heteropurpose sequence contains one or more 2'-O-methyl (OMe) modified nucleotides. In some embodiments, the heteropurpose sequence contains one or more 2'-fluoro modified nucleotides. In certain embodiments, the heteropurpose sequence contains a region having a pattern in which 2'-fluoro modified nucleotides alternate with unmodified nucleotides (e.g., unmodified ribonucleotides). In certain embodiments, the pattern starts from the 5' end of the heteropurpose sequence (e.g., the 5' endmost nucleotide of the heteropurpose sequence contains the 2'-fluoro modification). In other embodiments, the pattern starts from a second nucleotide from the 5' end of the heteropurpose sequence (e.g., so that the 5' endmost nucleotide of the heteropurpose sequence is unmodified and the next nucleotide contains the 2'-fluoro modification). In certain embodiments, regions having an alternating pattern of 2'-fluoromodified and unmodified nucleotides contain at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 2'-fluoromodified nucleotides. In certain embodiments, regions having an alternating pattern of 2'-fluoromodified and unmodified nucleotides contain 0-5, 5-10, 10-15, 15-20, 20-25, 25-30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, 150-200, 200- The regions have lengths of 300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000, 1000-1500, 1500-2000, 2000-2500, 2500-3000, 3000-3500, 3500-4000, 4000-4500, or 4500-5000 nucleotides. In certain embodiments, regions having alternating patterns of 2'-fluoromodified and unmodified nucleotides have a length equal to the length obtained by subtracting 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides from the heterologous target sequence.

[0359] In some embodiments, the PBS sequence of the template RNA includes one or more 2'-fluoro-modified nucleotides. In some embodiments, the PBS sequence of the template RNA includes one or more 2'-OMe-modified nucleotides. In some embodiments, the PBS sequence of the template RNA includes one or more nucleotides each containing both 2'-OMe modification and phosphorothioate modification. In certain embodiments, the 3' end of the PBS sequence includes, in order from 5' to 3', a 2'-fluoro-modified nucleotide, one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) 2'-OMe-modified nucleotides and / or one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) nucleotides each containing both 2'-OMe modification and phosphorothioate modification.

[0360] In some embodiments, the nucleotides in the linkage between the heterologous target sequence and the PBS sequence (e.g., one or more nucleotides at positions +3, +2, +1, -1, -2, and / or -3) do not contain 2'-fluoro modifications. In some embodiments, the nucleotides in the linkage between the heterologous target sequence and the PBS sequence (e.g., one or more nucleotides at positions +3, +2, +1, -1, -2, and / or -3) do not contain 2'-OMe modifications. In some embodiments, the nucleotides in the linkage between the heterologous target sequence and the PBS sequence (e.g., one or more nucleotides at positions +3, +2, +1, -1, -2, and / or -3) are unmodified nucleotides.

[0361] For example, a series of further exemplary template RNA sequences containing 2'-fluoro modifications at various positions, as tested in Example 7 of this specification, are shown in Table X3 below.

[0362] Table X3. Exemplary template RNA sequences including 2'-fluoromodification patterns. Nucleotide modifications are indicated as follows: phosphorothioate bonds indicated by asterisks, 2'-O-methyl groups indicated by "m" preceding the nucleotide, and 2'-fluoro groups indicated by / i2FN / (where N is any nucleotide). Columns 1 and 2 list the names of the sequences. Column 3 lists the nucleic acid sequences, including the indication of the nucleotide modifications described above. Column 4 lists nucleic acid sequences without modifications.

[0363] [Table X3-1]

[0364] [Table X3-2]

[0365] In some embodiments, chemical modifications are described in International Publication No. 2016 / 183482, U.S. Patent Application Publication No. 20090286852, International Patent Publication No. 2012 / 019168, International Publication No. 2012 / 045075, International Publication No. 2012 / 135805, International Publication No. 2012 / 158736, International Publication No. 2013 / 039857, International Publication No. 2013 / 039861, International Publication No. 2013 / 052523, International Publication No. 20 Pamphlet No. 13 / 090648, International Publication No. 2013 / 096709, International Publication No. 2013 / 101690, International Publication No. 2013 / 106496, International Publication No. 2013 / 130161, International Publication No. 2013 / 151669, International Publication No. 2013 / 151736, International Publication No. 2013 / 151672, International Publication No. 2013 / 151664, International Publication No. 2013 / 151665, International Publication No. 2013 / 151 Pamphlet No. 668, International Publication No. 2013 / 151671, International Publication No. 2013 / 151667, International Publication No. 2013 / 151670, International Publication No. 2013 / 151666, International Publication No. 2013 / 151663, International Publication No. 2014 / 028429, International Publication No. 2014 / 081507, International Publication No. 2014 / 093924, International Publication No. 2014 / 093574, International Publication No. 2014 / 113089 Brochures, International Publication No. 2014 / 144711, International Publication No. 2014 / 144767, International Publication No. 2014 / 144039, International Publication No. 2014 / 152540, International Publication No. 2014 / 152030, International Publication No. 2014 / 152031, International Publication No. 2014 / 152027, International Publication No. 2014 / 152211, International Publication No. 2014 / 158795, International Publication No. 2014 / 159813,International Publication Brochure No. 2014 / 164253, International Publication Brochure No. 2015 / 006747, International Publication Brochure No. 2015 / 034928, International Publication Brochure No. 2015 / 034925, International Publication Brochure No. 2015 / 038892, International Publication Brochure No. 2015 / 048744, International Publication Brochure No. 2015 / 051214, International Publication Brochure No. 2015 / 051173, International Publication Brochure No. 2015 / 051169, International Publication Brochure No. 2015 / 058069, International Publication Brochure No. 2015 / 085318, International Publication Brochure No. 2015 / 089511, International Publication Brochure No. 2015 / 105926, International Publication Brochure No. 2015 These are provided in brochures No. / 164674, International Publication No. 2015 / 196130, International Publication No. 2015 / 196128, International Publication No. 2015 / 196118, International Publication No. 2016 / 011226, International Publication No. 2016 / 011222, International Publication No. 2016 / 011306, International Publication No. 2016 / 014846, International Publication No. 2016 / 022914, International Publication No. 2016 / 036902, International Publication No. 2016 / 077125, or International Publication No. 2016 / 077123 (each of these being incorporated herein by reference in whole). It is understood that the incorporation of chemically modified nucleotides into polynucleotides can result in modifications incorporated into the nucleic acid base, the backbone, or both, depending on the location of the modification within the nucleotide. In some embodiments, backbone modifications are provided in European Patent No. 2813570 (which is incorporated herein by reference in its entirety). In some embodiments, modified caps are provided in U.S. Patent Application Publication No. 20050287539 (which is incorporated herein by reference in its entirety).

[0366] In some embodiments, chemically modified nucleic acids (e.g., RNA, e.g., mRNA) are ARCA: anti-reverse cap analog (m27.3'-OGP3G), GP3G (unmethylated cap analog), m7GP3G (monomethylated cap analog), m32.2.7GP3G (trimethylated cap analog), m5CTP (5'-methylcytidine triphosphate), m6ATP (N6-methyladenosine-5'-triphosphate), s2UTP (2-thiouridine triphosphate), cap 1 ( m7 GpppN 2’Ome It contains one or more of N and Ψ (psoiduridine triphosphate).

[0367] In some embodiments, the chemically modified nucleic acids include a 5' cap, such as a 7-methylguanosine cap (e.g., O-Me-m7G cap), a hypermethylated cap analog, an NAD+-derived cap analog (e.g., as described in Kiledjian, Trends in Cell Biology 28, 454-464 (2018)), or a modified cap analog, such as a biotinylated cap analog (e.g., as described in Bednarek et al., Phil Trans R Soc B 373, 20180167 (2018)).

[0368] In some embodiments, the chemically modified nucleic acid may be a poly-A tail, a 16-nucleotide stem-loop structure flanked by an unpaired 5-nucleotide (e.g., as described in Mannironi et al., Nucleic Acid Research 17, 9113-9126 (1989)), a triple helix structure (e.g., as described in Brown et al., PNAS 109, 19202-19207 (2012)), tRNA, Y RNA, or vault RNA structure (e.g., Labno et al., Biochemica et Biophysica Acta The 3' feature comprises one or more deoxyribonucleotide triphosphates (dNTPs) (as described in 1863,3125-3147 (2016)), the incorporation of 2'O-methylated NTPs or phosphorothioate-NTPs, a single nucleotide chemical modification (e.g., oxidation of the 3' terminal ribose to a reactive aldehyde and subsequent conjugation of the aldehyde-reactive modified nucleotide), or one or more chemical ligation to another nucleic acid molecule.

[0369] In some embodiments, nucleic acids (e.g., template nucleic acids) are, for example, dihydrouridine, inosine, 7-methylguanosine, 5-methylcytidine (5mC), 5'-ribothymidine phosphate, 2'-O-methylribothymidine, 2'-O-ethylribothymidine, 2'-fluororibothymidine, C-5-propynyl-deoxycytidine (pdC), C-5-propynyl-deoxyuridine (pdU), C-5-propynyl-cytidine (pC), C-5-propynyl-uridine (pU), 5-methylcytidine, 5-methyluridine, 5-methyldeoxycytidine, 5-methyldeoxyuridine methoxy, 2,6-diaminopropyl The material contains one or more modified nucleotides selected from 5'-dimethoxytrityl-N4-ethyl-2'-deoxycytidine, C-5 propynyl-f-cytidine (pfC), C-5 propynyl-f-uridine (pfU), 5-methyl-f-cytidine, 5-methyl-f-uridine, C-5 propynyl-m-cytidine (pmC), C-5 propynyl-f-uridine (pmU), 5-methyl-m-cytidine, 5-methyl-m-uridine, LNA (locked nucleic acid), MGB (sub-groove binder), pseudouridine (Ψ), 1-N-methylpsoiduridine (1-Me-Ψ), or 5-methoxyuridine (5-MO-U).

[0370] In some embodiments, the nucleic acid includes backbone modifications, such as modifications to sugar or phosphate groups in the backbone. In some embodiments, the nucleic acid includes nucleic acid base modifications.

[0371] In some embodiments, the nucleic acid includes one or more chemically modified nucleotides from Table 13, one or more chemical backbone modifications from Table 14, and one or more chemically modified caps from Table 15. For example, in some embodiments, the nucleic acid includes two or more (e.g., 3, 4, 5, 6, 7, 8, 9, or 10 or more) different types of chemical modifications. For example, the nucleic acid may include, for example, two or more (e.g., 3, 4, 5, 6, 7, 8, 9, or 10 or more) different types of modified nucleic acid bases from, for example, Table 13 as described herein. Alternatively or in combination, the nucleic acid may include, for example, two or more (e.g., 3, 4, 5, 6, 7, 8, 9, or 10 or more) different types of backbone modifications from, for example, Table 14 as described herein. Alternatively or in combination, the nucleic acid may include, for example, one or more modified caps from, for example, Table 15 as described herein. For example, in some embodiments, the nucleic acid includes one or more types of modified nucleic acid bases and one or more types of backbone modifications, one or more types of modified nucleic acid bases and one or more types of modified caps, one or more types of modified caps and one or more types of backbone modifications, or one or more types of modified nucleic acid bases, one or more types of backbone modifications and one or more types of modified caps.

[0372] In some embodiments, the nucleic acid contains one or more modified nucleic acid bases (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000 or more). In some embodiments, all nucleic acid bases of the nucleic acid are modified. In some embodiments, the nucleic acid is modified at one or more positions in the backbone (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000 or more). In some embodiments, all backbone positions of the nucleic acid are modified.

[0373] [Table 13-1]

[0374] [Table 13-2]

[0375] [Table 14]

[0376] [Table 15]

[0377] The nucleotides containing the template for the genetically modified system may be natural or modified bases or combinations thereof. For example, the template may include pseudouridine, dihydrouridine, inosine, 7-methylguanosine, or other modified bases. In some embodiments, the template may include locked nucleic acid nucleotides. In some embodiments, the modified bases used in the template do not inhibit the reverse transcription of the template. In some embodiments, the modified bases used in the template may improve the reverse transcription, such as specificity or fidelity.

[0378] In some embodiments, the RNA component of the system (e.g., template RNA or gRNA) includes one or more nucleotide modifications. In some embodiments, the modification pattern of the gRNA may significantly affect in vivo activity compared to an unmodified or terminally modified guide (e.g., shown in Figure 1D from Finn et al. Cell Rep 22(9):2227-2235 (2018) (the entirety of which is incorporated herein by reference)). While not intended to be bound by any particular theory, this process may be at least partially attributable to the stabilization of the RNA conferred by the modifications. Non-limiting examples of such modifications may include 2'-O-methyl (2'-O-Me), 2'-O-(2-methoxyethyl) (2'-O-MOE), 2'-fluoro (2'-F), phosphorothioate (PS) bonds between nucleotides, GC substitutions, and inverted debase bonds between nucleotides and their equivalents.

[0379] In some embodiments, the template RNA (e.g., a portion thereof that binds to a target site) or guide RNA includes a 5' terminal region. In some embodiments, the template RNA or guide RNA does not include a 5' terminal region. In some embodiments, the 5' terminal region includes a gRNA spacer region, for example, as described for sgRNA in Briner AE et al, Molecular Cell 56:333-339 (2014) (the entirety of which is incorporated herein by reference and applicable herein to, for example, all guide RNAs). In some embodiments, the 5' terminal region includes 5' terminal modifications. In some embodiments, the 5' terminal region may associate with crRNA, trRNA, sgRNA and / or dgRNA with or without the spacer region. The gRNA spacer region may, in some cases, include a guide region, guide domain, or target domain.

[0380] In some embodiments, the template RNA (e.g., a portion thereof that binds to a target site) or guide RNA described herein comprises one of the sequences shown in Table 4 of International Publication No. 2018107028A1 (which is incorporated herein by reference in its entirety). In some embodiments, if the sequence indicates a guide and / or spacer region, the composition may or may not include this region. In some embodiments, the guide RNA comprises one or more modifications of any of the sequences shown in Table 4 of International Publication No. 2018107028A1, such as those identified therein by sequence numbers. In some embodiments, the nucleotides may be the same or different, and / or the indicated modification patterns may be identical or similar to the modification patterns of guide sequences as shown in Table 4 of International Publication No. 2018107028A1. In some embodiments, the modification pattern includes the relative location and identity of the gRNA modification or the gRNA region (e.g., the 5' terminal region, the downstream stem region, the bulge region, the upstream stem region, the nexus region, the hairpin 1 region, the hairpin 2 region, the 3' terminal region). In some embodiments, the modification pattern includes at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the modifications across one or more regions of any of the sequences shown in the sequence column of Table 4 of International Publication Brochure 2018107028A1. In some embodiments, the modification pattern is at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of the modification patterns of sequences shown in the sequence column of Table 4 of International Publication Brochure No. 2018107028A1.In some embodiments, the modification pattern is at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the modification pattern of the sequence over the 5' terminal region, as shown in Table 4 of International Publication Brochure No. 2018107028A1. In some embodiments, the modification pattern is at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical across the downstream stem. In some embodiments, the modification pattern is at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical across the bulge. In some embodiments, the modification pattern is at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical across the upstream stem. In some embodiments, the modification pattern is identical across the nexus by at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the modification pattern is identical across hairpin 1 by at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the modification pattern is identical across hairpin 2 by at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%.In some embodiments, the modification pattern is identical across the 3' end by at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the modification pattern differs from the modification pattern of the sequence in Table 4 of International Publication Brochure 2018107028A1 or a region of such sequence (e.g., the 5' end, the downstream stem, the bulge, the upstream stem, the nexus, hairpin 1, hairpin 2, the 3' end) by, for example, 0, 1, 2, 3, 4, 5, 6 or more nucleotides. In some embodiments, the gRNA includes modifications that differ from the modifications of the sequence in Table 4 of International Publication Brochure 2018107028A1 by, for example, 0, 1, 2, 3, 4, 5, 6 or more nucleotides. In some embodiments, the gRNA includes modifications in the regions of the sequence shown in Table 4 of International Publication No. 2018107028A1 (e.g., the 5' end, the downstream stem, the bulge, the upstream stem, the nexus, hairpin 1, hairpin 2, and the 3' end) and different modifications by, for example, 0, 1, 2, 3, 4, 5, 6 or more nucleotides.

[0381] In some embodiments, the template RNA (e.g., a portion thereof that binds to a target site) or gRNA includes a 2'-O-methyl (2'-O-Me) modified nucleotide. In some embodiments, the gRNA includes a 2'-O-(2-methoxyethyl) (2'-O-moe) modified nucleotide. In some embodiments, the gRNA includes a 2'-fluoro (2'-F) modified nucleotide. In some embodiments, the gRNA includes a phosphorothioate (PS) bond between nucleotides. In some embodiments, the gRNA includes a 5'-terminal modification, a 3'-terminal modification, or 5' and 3'-terminal modifications. In some embodiments, the 5'-terminal modification includes a phosphorothioate (PS) bond between nucleotides. In some embodiments, the 5'-terminal modification includes a 2'-O-methyl (2'-O-Me), a 2'-O-(2-methoxyethyl) (2'-O-MOE), and / or a 2'-fluoro (2'-F) modified nucleotide. In some embodiments, the 5'-terminus modification includes at least one phosphorothioate (PS) linkage and one or more 2'-O-methyl (2'-O-Me), 2'-O-(2-methoxyethyl) (2'-O-MOE), and / or 2'-fluoro (2'-F) modified nucleotides. The terminal modification may include phosphorothioate (PS), 2'-O-methyl (2'-O-Me), 2'-O-(2-methoxyethyl) (2'-O-MOE), and / or 2'-fluoro (2'-F) modifications. Equivalent terminal modifications are also encompassed by the embodiments described herein. In some embodiments, the template RNA or gRNA includes terminal modifications in combination with modifications to one or more regions of the template RNA or gRNA. Further exemplary modifications and methods for protecting RNA, e.g., gRNA and its compositional formula are described in International Publication No. 2018126176A1 (which is incorporated herein by reference in its entirety).

[0382] In some embodiments, the template RNA described herein includes three phosphorothioate bonds at the 5' end and three phosphorothioate bonds at the 3' end. In some embodiments, the template RNA described herein includes three 2'-O-methylribonucleotides at the 5' end and three 2'-O-methylribonucleotides at the ' end. In some embodiments, the three nucleotides on the 5' end of the template RNA are 2'-O-methylribonucleotides, the internucleotide bond between the three nucleotides on the 5' end of the template RNA is a phosphorothioate bond, the three nucleotides on the 3' end of the template RNA are 2'-O-methylribonucleotides, and the internucleotide bond between the three nucleotides on the 3' end of the template RNA is a phosphorothioate bond. In some embodiments, the template RNA includes alternating blocks of ribonucleotides and 2'-O-methylribonucleotides, for example, blocks of 12 to 28 nucleotides in length. In some embodiments, the central portion of the template RNA contains alternating blocks, with the 5' and 3' ends each containing three 2'-O-methylribonucleotides and three phosphorothioate bonds.

[0383] In some embodiments, modifications (e.g., 2'-OMe-RNA, 2'-F-RNA, and PS modifications) are introduced into the template RNA or guide RNA using structural guides and phylogenetic methods, for example, as described in Mir et al. Nat Commun 9:2641 (2018) (which is incorporated herein by reference in its entirety). In some embodiments, the incorporation of 2'-F-RNA increases the thermal and nuclease stability of RNA:RNA or RNA:DNA double strands while minimally interfering with, for example, C3'-endoglycan puckering. In some embodiments, 2'-F may exhibit better tolerance than 2'-OMe at positions where 2'-OH is important for RNA:DNA double strand stability. In some embodiments, the crRNA includes one or more modifications that do not reduce Cas9 activity, for example, C10, C20, or C21 (fully modified) as listed in Supplementary Table 1 of Mir et al. Nat Commun 9:2641 (2018) (which is incorporated herein by reference in its entirety). In some embodiments, the tracrRNA contains one or more modifications that do not reduce Cas9 activity, e.g., T2, T6, T7, or T8 (fully modified) as shown in Supplementary Table 1 of Mir et al. Nat Commun 9:2641 (2018). In some embodiments, a crRNA containing one or more modifications (e.g., as described herein) may be paired with a tracrRNA containing one or more modifications, e.g., C20 and T2. In some embodiments, the gRNA contains, for example, a chimera of crRNA and tracrRNA (e.g., Jinek et al. Science 337(6096):816-821 (2012)). In several embodiments, modifications of crRNA and tracrRNA are mapped onto a single guide chimera, for example, to produce a modified gRNA with enhanced stability.

[0384] In some embodiments, the gRNA molecule may be modified by the addition or subtraction of naturally occurring components, such as hairpins. In some embodiments, the gRNA may include a gRNA in which one or more 3' hairpin elements are deleted, for example, as described in International Publication No. 2018106727 (which is incorporated herein by reference in its entirety). In some embodiments, the gRNA may contain an added hairpin structure, for example, a hairpin structure added within a spacer region, which has been shown in the teachings of Kocak et al. Nat Biotechnol 37(6):657-666 (2019) to increase the specificity of the CRISPR-Cas system. Further modifications, including examples of shortened gRNAs and specific modifications that improve in vivo activity, can be found in U.S. Patent Application Publication No. 20190316121 (which is incorporated herein by reference in its entirety).

[0385] In some embodiments, modifications to the template RNA are identified using structural guides and phylogenetic methods (e.g., as described in Mir et al. Nat Commun 9:2641 (2018), which is incorporated herein by reference in whole). In several embodiments, the modifications are identified by the inclusion or exclusion of guide regions in the template RNA. In some embodiments, the structure of the polypeptide bound to the template RNA is used to determine the nucleotides in contact with non-proteins of the RNA, and then modifications that pose a lower risk of disrupting the association between the RNA and the polypeptide can be selected. Furthermore, the secondary structure in the template RNA can be predicted in silico using a software tool, e.g., the RNA structure tool available at rna.urmc.rochester.edu / RNAstructureWeb (Bellaousov et al. Nucleic Acids Res 41:W471-W474 (2013), which is incorporated herein by reference in whole), to determine secondary structures for selecting modifications, e.g., hairpins, stems, and / or bulges.

[0386] Preparation of compositions and systems As those skilled in the art will understand, methods for designing and constructing nucleic acid constructs and proteins or polypeptides (such as the systems, constructs, and polypeptides described herein) are common in the art. Recombinant methods can generally be used. For an overview, see Smales & James (Eds.), Therapeutic Proteins: Methods and Protocols (Methods in Molecular Biology), Humana Press (2005) and Crommelin, Sindelar & Meibohm (Eds.), Pharmaceutical Biotechnology: Fundamentals and Applications, Springer (2013). Methods for designing, preparing, evaluating, purifying, and manipulating nucleic acid compositions are described in Green and Sambrook (Eds.), Molecular Cloning: A Laboratory Manual (Fourth Edition), Cold Spring Harbor Laboratory Press (2012).

[0387] This disclosure provides, in part, nucleic acids encoding recombinant polypeptides, template nucleic acids, or both as described herein, such as vectors. In some embodiments, the vector includes a selection marker, such as an antibiotic resistance marker. In some embodiments, the antibiotic resistance marker is a kanamycin resistance marker. In some embodiments, the antibiotic resistance marker does not confer resistance to β-lactam antibiotics. In some embodiments, the vector does not include an ampicillin resistance marker. In some embodiments, the vector includes a kanamycin resistance marker but does not include an ampicillin resistance marker. In some embodiments, the vector encoding a recombinant polypeptide is incorporated into the target cell genome (e.g., upon administration to target cells, tissues, organs, or subjects). In some embodiments, the vector encoding a recombinant polypeptide is not incorporated into the target cell genome (e.g., upon administration to target cells, tissues, organs, or subjects). In some embodiments, the vector encoding a template nucleic acid (e.g., template RNA) is not incorporated into the target cell genome (e.g., upon administration to target cells, tissues, organs, or subjects). In some embodiments, when the vector is integrated into a target site in the target cell genome, the selection marker is not integrated into the genome. In some embodiments, when the vector is integrated into a target site in the target cell genome, the genes or sequences involved in vector maintenance (e.g., plasmid maintenance genes) are not integrated into the genome. In some embodiments, when the vector is integrated into a target site in the target cell genome, the transfer regulatory sequences (e.g., inverted terminal sequences derived from AAV) are not integrated into the genome. In some embodiments, administration of the vector (e.g., encoding a recombinant polypeptide, a template nucleic acid, or both) to a target cell, tissue, organ, or subject results in the integration of a portion of the vector into one or more target sites in the genome of the target cell, tissue, organ, or subject.In some embodiments, 99, 95, 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, 4, 3, 2, or less than 1% (e.g., no target sites) of the target sites containing the incorporated substance include vector-derived, selection markers (e.g., antibiotic resistance genes), transfer regulatory sequences (e.g., inverted terminal sequences from AAV), or both.

[0388] Exemplary methods for producing pharmaceutical proteins or polypeptides described herein include expression in mammalian cells, but recombinant proteins can also be produced using insect cells, yeast, bacteria, or other cells under the control of a suitable promoter. Mammalian expression vectors may include non-transcriptional factors, such as origins of replication, suitable promoters, and other 5' or 3' flanking non-transcriptional sequences, as well as 5' or 3' untranslated sequences, such as required ribosome binding sites, polyadenylation sites, splice donor and acceptor sites, and termination sequences. DNA sequences derived from the SV40 viral genome, such as the SV40 origin, early promoter, splice, and polyadenylation sites, can be used to provide other genetic factors necessary for the expression of heterologous DNA sequences. Cloning and expression vectors suitable for use in bacterial, fungal, yeast, and mammalian cell hosts are described in Green & Sambrook, Molecular Cloning: A Laboratory Manual (Fourth Edition), Cold Spring Harbor Laboratory Press (2012).

[0389] Recombinant proteins can be expressed and produced using various mammalian cell culture systems. Examples of mammalian expression systems include the CHO, COS, HEK293, HeLA, and BHK cell lines. A host cell culture method for the production of protein therapeutics is described in Zhou and Kantardjieff (Eds.), Mammalian Cell Cultures for Biologics Manufacturing (Advances in Biochemical Engineering / Biotechnology), Springer (2014). The compositions described herein may include vectors encoding recombinant proteins, such as viral vectors, such as lentiviral vectors. In some embodiments, the vector, such as a viral vector, may include nucleic acids encoding recombinant proteins.

[0390] The purification of protein therapeutics is described in Franks, Protein Biotechnology: Isolation, Characterization, and Stabilization, Humana Press (2013) and Cutler, Protein Purification Protocols (Methods in Molecular Biology), Humana Press (2010).

[0391] This disclosure also provides compositions and methods for producing template nucleic acid molecules (e.g., template RNA) having specificity for recombinant polypeptides and / or genomic target sites. In one embodiment, the method includes producing an RNA segment comprising an upstream homology segment, a heterologous target sequence segment, a recombinant polypeptide binding motif, and a gRNA segment.

[0392] therapeutic use In some embodiments, the genetic recombination systems described herein can be used to modify cells (e.g., animal cells, plant cells, or fungal cells). In some embodiments, the genetic recombination systems described herein can be used to modify mammalian cells (e.g., human cells). In some embodiments, the genetic recombination systems described herein can be used to modify cells derived from domestic animals (e.g., cattle, horses, sheep, goats, pigs, llamas, alpacas, camels, yaks, chickens, ducks, geese, or ostriches). In some embodiments, the genetic recombination systems described herein can be used as experimental or research tools, or in experimental or research methods, for example, to modify animal cells, mammalian cells (e.g., human cells), plant cells, or fungal cells.

[0393] By incorporating coding genes into RNA sequence templates, providing expression of therapeutic transgenes in individuals with, for example, loss-of-function mutations, replacing gain-of-function mutations with normal transgenes, providing regulatory sequences to eliminate gain-of-function mutation expression, and / or controlling the expression of operably linked genes, transgenes, and their systems, genetically modified systems can address therapeutic needs. In certain embodiments, the RNA sequence template encodes promoter regions specific to the therapeutic needs of host cells, such as tissue-specific promoters or enhancers. In yet other embodiments, the promoters can be operably linked to the coding sequences.

[0394] Accordingly, this specification provides methods for treating phenylketonuria (PKU) or hyperphenylalaninemia (e.g., mild or severe hyperphenylalaninemia) of interest. In some embodiments, the treatment results in improvement of one or more symptoms associated with PKU or hyperphenylalaninemia.

[0395] In some embodiments, the systems described herein are used to treat subjects having the R408 mutation (e.g., R408W).

[0396] In some embodiments, treatment with the system disclosed herein results in the correction of R408W mutations in approximately 5–50% of cells (e.g., approximately 5–10%, 10–20%, 20–30%, 30–40%, 40–50%, or approximately 10%). In some embodiments, treatment with the system disclosed herein results in the correction of R408W mutations in approximately 5–50% of DNA from treated cells (e.g., approximately 5–10%, 10–20%, 20–30%, 30–40%, 40–50%, or approximately 10%).

[0397] In some embodiments, treatment with the genetically modified system described herein is compared to subjects having PKUs that have not been treated with the genetically modified system described herein. (a) Increase in phenylalanine hydroxylase (PAH) activity, efficiency and / or function, (b) A decrease in the concentration of phenylalanine in the blood and / or cerebrospinal fluid. (c) Increased concentration of tyrosine in the blood, (d) Restoration of normal synthesis of dopamine, norepinephrine and / or melanin, (e) Reduced urea production, and / or (f) Improvement of protein retention and / or Phe utilization To bring about one or more of the following.

[0398] Administration and Delivery The compositions and systems described herein may be used in vitro or in vivo. In some embodiments, the system or components of the system are delivered, for example, to cells (e.g., mammalian cells, e.g., human cells) in vitro or in vivo. In some embodiments, the cells are eukaryotic cells, e.g., multicellular organisms, e.g., animals, e.g., mammals (e.g., humans, pigs, cattle), birds (e.g., poultry, e.g., chickens, turkeys, or ducks), or fish cells. In some embodiments, the cells are non-human animal cells (e.g., laboratory animals, livestock animals, or companion animals). In some embodiments, the cells are stem cells (e.g., hematopoietic stem cells), fibroblasts, or T cells. In some embodiments, the cells are immune cells, e.g., T cells (e.g., Treg, CD4, CD8, γδ, or memory T cells), B cells (e.g., memory B cells or plasma cells), or NK cells. In some embodiments, the cells are non-dividing cells, e.g., non-dividing fibroblasts or non-dividing T cells. In some embodiments, the cells are HSCs, and p53 is either not upregulated or is upregulated to less than 10%, 5%, 2%, or 1%, as determined according to the method described in Example 30 of PCT / US2019 / 048607. Those skilled in the art will understand that components of a genetically modified system can be delivered in the form of polypeptides, nucleic acids (e.g., DNA, RNA), and combinations thereof.

[0399] In one embodiment, the system and / or components of the system are delivered as nucleic acids. For example, a recombinant polypeptide may be delivered in the form of DNA or RNA encoding the polypeptide, and a template RNA may be delivered in the form of RNA or its complementary DNA transcribed to the RNA. In some embodiments, the system or components of the system are delivered on one, two, three, four or more individual nucleic acid molecules. In some embodiments, the system or components of the system are delivered as a combination of DNA and RNA. In some embodiments, the system or components of the system are delivered as a combination of DNA and protein. In some embodiments, the system or components of the system are delivered as a combination of RNA and protein. In some embodiments, the recombinant polypeptide is delivered as a protein.

[0400] In some embodiments, the system or its components are delivered to cells, such as mammalian or human cells, using a vector. The vector may be, for example, a plasmid or a virus. In some embodiments, delivery is in vivo, in vitro, ex vivo, or in situ. In some embodiments, the virus is an adeno-associated virus (AAV), lentivirus, or adenovirus. In some embodiments, the system or its components are delivered to cells using virus-like particles or viromosomes. In some embodiments, delivery uses two or more viruses, virus-like particles, or viromosomes.

[0401] In one embodiment, the compositions and systems described herein can be formulated in liposomes or other similar vesicles. Liposomes are spherical vesicle structures consisting of a monolayer or multilayer lipid bilayer surrounding an internal aqueous compartment and a relatively impermeable outer lipophilic phospholipid bilayer. Liposomes can be anionic, neutral, or cationic. Liposomes are biocompatible, non-toxic, and can deliver both hydrophilic and lipophilic drug molecules, protecting their cargo from plasma enzyme degradation and transporting their loads across biological membranes and the blood-brain barrier (BBB) ​​(see, for example, Spuch and Navarro, Journal of Drug Delivery, vol. 2011, Article ID 469679, 12 pages, 2011. doi:10.1155 / 2011 / 469679 for a review).

[0402] Vesicles can be produced from several different types of lipids, but phospholipids are most commonly used to create liposomes as drug carriers. Methods for preparing multilayer vesicle lipids are known in the art (see, for example, U.S. Patent No. 6,693,086, the teachings relating to the preparation of multilayer vesicle lipids are incorporated herein by reference). Vesicle formation can occur spontaneously when a lipid membrane is mixed with an aqueous solution, but it can also be accelerated by applying force in the form of shaking using a homogenizer, ultrasonic generator, or extruder (see, for example, Spuch and Navarro, Journal of Drug Delivery, vol. 2011, Article ID 469679, 12 pages, 2011, doi:10.1155 / 2011 / 469679 for a review). Extruded lipids can be prepared by extrusion through a size-reducing filter, as described in Templeton et al., Nature Biotech, 15:647-652, 1997 (the teachings thereon regarding the preparation of extruded lipids are incorporated herein by reference).

[0403] Various nanoparticles can be used for delivery, including liposomes, lipid nanoparticles, cationic lipid nanoparticles, ionizable lipid nanoparticles, polymer nanoparticles, gold nanoparticles, dendrimers, cyclodextrin nanoparticles, micelles, or combinations thereof.

[0404] Lipid nanoparticles are an example of carriers that provide biocompatible and biodegradable delivery systems for the pharmaceutical compositions described herein. Nanostructured lipid carriers (NLCs) are modified solid lipid nanoparticles (SLNs) that retain the characteristics of SLNs, improve drug stability and loading capacity, and prevent drug leakage. Polymer nanoparticles (PNPs) are important components of drug delivery. These nanoparticles can effectively direct drug delivery to specific targets and improve drug stability and drug release control. Lipid-polymer nanoparticles (PLNs), i.e., carriers combining liposomes and polymers, may also be used. These nanoparticles have the complementary advantages of PNPs and liposomes. PLNs consist of a core-shell structure, where the polymer core provides a stable structure and the phospholipid shell confers good biocompatibility. Thus, the two components enhance drug encapsulation efficiency, promote surface modification, and prevent leakage of water-soluble drugs. For example, see Li et al. 2017, Nanomaterials 7, 122; doi:10.3390 / nano7060122 for a review.

[0405] Exosomes may also be used as drug delivery vehicles for the compositions and systems described herein. For a review, see Ha et al. July 2016. Acta Pharmaceutica Sinica B. Volume 6, Issue 4, Pages 287-296; doi.org / 10.1016 / j.apsb.2016.02.001.

[0406] Because fusosomes interact with and fuse with target cells, they can be used as delivery vehicles for various molecules. Fusosomes generally consist of an amphiphilic lipid bilayer enclosing a lumen or cavity and a fusogen that interacts with the amphiphilic lipid bilayer. The fusogen component has been shown to be manipulable to confer target cell specificity for fusion and payload delivery, enabling the creation of delivery vehicles with programmable cell specificity (see, for example, International Publication No. 2020014209, the teachings relating to fusosome design, preparation and use which are incorporated herein by reference).

[0407] In some embodiments, the protein components of the genetically modified system may be pre-bound to a template nucleic acid (e.g., template RNA). For example, in some embodiments, a genetically modified polypeptide may be initially combined with a template nucleic acid (e.g., template RNA) to form a ribonucleoprotein (RNP) complex. In some embodiments, RNPs may be delivered to cells, for example, via transfection, nucleofection, virus, vesicle, LNP, exosome, or fusosome.

[0408] Genetic engineering systems can be introduced into cells, tissues, and multicellular organisms. In some embodiments, the system or its components are delivered to the cells by mechanical or physical means.

[0409] The formulation of protein-based therapeutic drugs is described in Meyer (Ed.), Therapeutic Protein Drug Products: Practical Approaches to Formulation in the Laboratory, Manufacturing, and the Clinic, Woodhead Publishing Series (2012).

[0410] Tissue-specific activity / administration In some embodiments, the systems described herein may utilize one or more features (e.g., promoters or microRNA binding sites) to limit activity in off-target cells or tissues.

[0411] In some embodiments, the nucleic acids described herein (e.g., template RNA or DNA encoding template RNA) include a promoter sequence, such as a tissue-specific promoter sequence. In some embodiments, a tissue-specific promoter is used to increase the target cell specificity of the recombination system. For example, a promoter may be selected based on its activity in target cell types but inactivity (or activity at a lower level) in non-target cell types. Thus, even if the promoter is integrated into the genome of non-target cells, it will not drive the expression of the integrated gene (or will only drive low levels of expression). Furthermore, systems having a tissue-specific promoter sequence in the template RNA may be used in combination with a microRNA binding site, for example, in the template RNA or nucleic acid encoding a recombinant protein, as described herein. Systems having a tissue-specific promoter sequence in the template RNA may also be used in combination with DNA encoding a recombinant polypeptide, driven by the tissue-specific promoter, to obtain higher levels of the recombinant protein in target cells than in non-target cells, for example. In some embodiments, for example, in the case of liver indications, the tissue-specific promoter is selected from Table 3 of International Publication No. 2020014209, which is incorporated herein by reference.

[0412] In some embodiments, the nucleic acids described herein (e.g., template RNA or DNA encoding template RNA) include a microRNA binding site. In some embodiments, the microRNA binding site is used to increase the target cell specificity of the recombination system. For example, the microRNA binding site may be selected based on its recognition by miRNAs that are present in non-target cell types but not in target cell types (or present at reduced levels compared to non-target cells). Thus, if the template RNA is present in a non-target cell, it will be bound by the miRNA; if the template RNA is present in a target cell, it will not be bound by the miRNA (or will be bound at reduced levels compared to non-target cells). While not intended to be bound to any particular theory, the binding of miRNAs to template RNA may interfere with its activity, for example, with the insertion of heterologous target sequences into the genome. Therefore, this system allows for more efficient editing of the target cell genome than it would be if it were editing the genome of a non-target cell. For example, a heterologous target sequence would be inserted into the target cell genome more efficiently than it would be inserted into the non-target cell genome, or the insertion or deletion would be generated more efficiently in the target cell than in the non-target cell. Furthermore, a system having a microRNA binding site in the template RNA (or the DNA encoding it) can be used in combination with a nucleic acid encoding a recombinant polypeptide, and the expression of the recombinant polypeptide is regulated by the second microRNA binding site, for example, as described herein. In some embodiments, for example, in the case of hepatic indications, the miRNA is selected from Table 4 of International Publication No. 2020014209 (incorporated herein by reference).

[0413] In some embodiments, the template RNA includes a microRNA sequence, an siRNA sequence, a guide RNA sequence, or a piwi RNA sequence.

[0414] Promoter In some embodiments, one or more promoters or enhancers are operably linked to a nucleic acid encoding a recombinant protein or, for example, a template nucleic acid that controls the expression of a heterologous target sequence. In certain embodiments, one or more promoters or enhancers include cell type or tissue-specific elements. In some embodiments, the promoter or enhancer is the same as, or derived from, a promoter or enhancer that naturally controls the expression of a heterologous target sequence. For example, ornithine transcarbomylase promoters and enhancers can be used to control the expression of the ornithine transcarbomylase gene in a system or method provided by the present invention in order to correct ornithine transcarbomylase deletion. In some embodiments, the promoter is the promoter or a functional fragment or variant thereof from Table 16 or 17.

[0415] Exemplary commercially available tissue-specific promoters can be found, for example, at the URL (e.g., invivogen.com / tissue-specific-promoters). In some embodiments, the promoter is a native promoter or a minimal promoter (e.g., consisting of a single fragment from the 5' region of a given gene). In some embodiments, the native promoter includes a core promoter and its native 5'UTR. In some embodiments, the 5'UTR includes an intron. In other embodiments, these include composite promoters created by combining promoters of different origins or by assembling a distal enhancer with a minimal promoter of the same origin.

[0416] Exemplary cell or tissue-specific promoters are listed in the table below, and the exemplary nucleic acids they encode are known in the art and readily available through various resources, such as the NCBI database including RefSeq. and the Eukaryotic Promoter Database ( / / epd.epfl.ch / / index.php).

[0417] [Table 16]

[0418] [Table 17-1]

[0419] [Table 17-2]

[0420] [Table 17-3]

[0421] Depending on the host / vector system used, several suitable transcriptional and translational regulatory elements, including constructive and inducible promoters, transcriptional enhancer elements, and transcriptional terminators, may be used in the expression vector (see, for example, Bitter et al. (1987) Methods in Enzymology, 153:516-544, which is incorporated herein by reference in its entirety).

[0422] In some embodiments, a nucleic acid encoding a recombinant protein or a template nucleic acid is operably linked to a regulatory element, such as a transcriptional regulatory element, like a promoter. In some embodiments, the transcriptional regulatory element may be functional in either eukaryotic cells, such as mammalian cells, or prokaryotic cells (e.g., bacterial or archaeal cells). In some embodiments, a nucleotide sequence encoding a polypeptide is operably linked to a plurality of regulatory elements, for example, enabling the expression of the polypeptide-encoding nucleotide sequence in both prokaryotic and eukaryotic cells.

[0423] For illustrative purposes, examples of spatially limited promoters include, but are not limited to, neuron-specific promoters, adipocyte-specific promoters, cardiomyocyte-specific promoters, smooth muscle-specific promoters, and photoreceptor-specific promoters. Neuron-specific spatially limited promoters include, but are not limited to, the following: neuron-specific enolase (NSE) promoters (e.g., see EMBL HSENO2, X51956), aromatic amino acid decarboxylase (AADC) promoters, neurofilament promoters (e.g., see GenBank HUMNFL, L04147), synapsin promoters (e.g., see GenBank HUMSYNIB, M55301), thy-1 promoters (e.g., see Chen et al. (1987) Cell 51:7-19 and Llewellyn, et al. (2010) Nat. Med. 16(10):1161-1166), serotonin receptor promoters (e.g., see GenBank S62283), and tyrosine hydroxylase (TH) promoters (e.g., Oh et al. (2009) Gene Ther 16:437; Sasaoka et al.) See al. (1992) Mol.Brain Res. 16:274; Boundy et al. (1998) J. Neurosci. 18:9989 and Kaneda et al. (1991) Neuron 6:583-594), GnRH promoter (see, for example, Radovick et al. (1991) Proc.Natl.Acad.Sci.USA 88:3402-3406), L7 promoter (see, for example, Oberdick et al. (1990) Science 248:223-226), DNMT promoter (see, for example, Bartge et al. (1988) Proc.Natl.Acad.Sci.USA 85:3648-3652), enkephalin promoter (see, for example, Comb et al. (1988) EMBO See J.17:3793-3805), myelin basic protein (MBP) promoters, Ca2+ calmodulin-dependent protein kinase II-α (CamKIIα) promoters (e.g., Mayford et al. (1996) Proc. Natl. Acad. Sci.See USA 93:13250 and Casanova et al. (2001) Genesis 31:37), CMV enhancers / platelet-derived growth factor-β promoters (see, for example, Liu et al. (2004) Gene Therapy 11:52-60).

[0424] Examples of adipocyte-specific spatially limited promoters include, but are not limited to, the following: aP2 gene promoters / enhancers, e.g., the region of the human aP2 gene from -5.4kb to +21bp (see, e.g., Tozzo et al. (1997) Endocrinol. 138:1604; Ross et al. (1990) Proc. Natl. Acad. Sci. USA 87:9590 and Pavjani et al. (2005) Nat. Med. 11:797), glucose transporter-4 (GLUT4) promoters (see, e.g., Knight et al. (2003) Proc. Natl. Acad. Sci. USA 100:14725), fatty acid translocase (FAT / CD36) promoters (see, e.g., Kuriki et al. (2002) Biol. Pharm. Bull. 25:1476 and Sato See et al. (2002) J. Biol. Chem. 277:15703), stearoyl-CoA desaturase-1 (SCD1) promoter (Tabor et al. (1999) J. Biol. Chem. 274:20603), leptin promoter (see, for example, Mason et al. (1998) Endocrinol. 139:1013 and Chen et al. (1999) Biochem. Biophys. Res. Comm. 262:187), azidonectin promoter (see, for example, Kita et al. (2005) Biochem. Biophys. Res. Comm. 331:484 and Chakrabarti (2010) Endocrinol. 151:2408), adipsin promoter (see, for example, Platt et al. See al. (1989) Proc. Natl. Acad. Sci. USA 86:7490), resistin promoters (see, for example, Seo et al. (2003) Molec. Endocrinol. 17:1522), etc.

[0425] Cardiomyocyte-specific spatially restricted promoters include, but are not limited to, regulatory sequences derived from the following genes: myosin light chain-2, α-myosin heavy chain, AE3, cardiac troponin C, cardiac actin, etc. Franz et al.(1997)Cardiovasc.Res.35:560-566;Robbins et al.(1995)Ann.NYAcad.Sci.752:492-505;Linn et al.(1995)Circ.Res.76:584-591;Parmacek et al. al. (1994) Mol. Cell. Biol. 14:1870-1885; Hunter et al. (1993) Hypertension 22:608-617 and Sartorelli et al. (1992) Proc. Natl. Acad. Sci. USA 89:4047-4051.

[0426] Examples of smooth muscle cell-specific spatially restricted promoters include, but are not limited to, the following: the SM22α promoter (see, e.g., Akyuerek et al. (2000) Mol. Med. 6:983 and U.S. Patent No. 7,169,874), the smootherin promoter (see, e.g., International Publication No. 2001 / 018048), and the α-smooth muscle actin promoter. For example, the 0.4kb region of the SM22α promoter (which contains two CArG elements) has been shown to mediate specific expression in vascular smooth muscle cells (see, for example, Kim, et al. (1997) Mol. Cell. Biol. 17, 2266-2278; Li, et al., (1996) J. Cell Biol. 132 849-859 and Moessler, et al. (1996) Development 122, 2415-2425).

[0427] Examples of photoreceptor-specific spatially limited promoters include, but are not limited to, the following: rhodopsin kinase promoter (Young et al. (2003) Ophthalmol.Vis.Sci. 44:4076), β-phosphodiesterase gene promoter (Nicoud et al. (2007) J.Gene Med. 9:1015), retinitis pigmentosa gene promoter (Nicoud et al. (2007) op. cit.), photoreceptor-retinoid-binding protein (IRBP) gene enhancer (Nicoud et al. (2007) op. cit.), IRBP gene promoter (Yokoyama et al. (1992) Exp Eye Res. 55:225), etc.

[0428] In some embodiments, a genetically modified system, such as DNA encoding a recombinant polypeptide, DNA encoding a template RNA, or DNA or RNA encoding a heterologous target sequence, is designed so that one or more elements are operably linked to a tissue-specific promoter, such as a promoter that is active in T cells. In further embodiments, the T cell-activating promoter is inactive in other cell types, such as B cells or NK cells. In some embodiments, the T cell-activating promoter is derived from a promoter for a gene encoding a component of the T cell receptor, such as TRAC, TRBC, TRGC, or TRDC. In some embodiments, the T cell-activating promoter is derived from a promoter for a gene encoding a component of a T cell-specific cluster of differentiated proteins, such as CD3, such as CD3D, CD3E, CD3G, or CD3Z. In some embodiments, the T cell-specific promoter in the genetically modified system is discovered by comparing publicly available gene expression data across cell types and selecting a promoter from genes whose expression is enhanced in T cells. In some embodiments, promoters may be selected according to the desired expression level, for example, promoters that are active only in T cells, promoters that are active only in NK cells, or promoters that are active in both T cells and NK cells.

[0429] Cell-specific promoters known in the art can be used to instruct the expression of recombinant proteins, such as those described herein. Non-limiting, exemplary mammalian cell-specific promoters have been characterized and used in mice that express Cre recombinase in a cell-specific manner. Several non-limiting, exemplary mammalian cell-specific promoters are listed in Table 1 of U.S. Patent No. 9,845,481 (incorporated herein by reference).

[0430] In some embodiments, the vectors described herein include an expression cassette. Typically, the expression cassette includes a nucleic acid molecule of the present invention operably ligated to a promoter sequence. For example, the promoter is operably ligated to a coding sequence if it can influence the expression of the coding sequence (e.g., the coding sequence is under the transcriptional control of the promoter). The coding sequence can be operably ligated to a regulatory sequence in sense or antisense orientation. In certain embodiments, the promoter is a heterologous promoter. In certain embodiments, the expression cassette may include other elements, such as introns, enhancers, polyadenylation sites, woodchuck response elements (WREs), and / or other elements known to influence the expression level of the coding sequence. The promoter typically controls the expression of a coding sequence or functional RNA. In certain embodiments, the promoter sequence may include proximal and more distal upstream elements, and may further include enhancer elements. Enhancers can typically stimulate promoter activity and may be native elements of the promoter or heterologous elements inserted to enhance the level or tissue specificity of the promoter. In certain embodiments, the promoter as a whole is derived from a native gene. In certain embodiments, the promoter is composed of various elements derived from different naturally occurring promoters. In certain embodiments, the promoter includes a synthetic nucleotide sequence. Those skilled in the art will understand that various promoters can direct gene expression in various tissues or cell types, at various developmental stages, or in response to various environmental conditions or the presence or absence of drugs or transcription cofactors. Ubiquitous, cell type-specific, tissue-specific, developmental stage-specific, and conditional promoters, such as drug-responsive promoters (e.g., tetracycline-responsive promoters), are well known to those skilled in the art.Examples of promoters include, but are not limited to, phosphoglycerate kinase (PKG) promoters, CAG (a complex of CMV enhancer, chicken β-actin promoter (CBA), and rabbit β-globin intron), NSE (neuron-specific enolase), synapsin or NeuN promoters, SV40 early promoter, mouse mammary tumor virus LTR promoter, adenovirus major late promoter (Ad MLP), herpes simplex virus (HSV) promoter, cytomegalovirus (CMV) promoter, e.g., CMV earliest promoter region (CMVIE), SFFV promoter, Rous sarcoma virus (RSV) promoter, synthetic promoters, and hybridization promoters. Other promoters may be derived from humans or other species, including mice. Common promoters include, for example, the human cytomegalovirus (CMV) early gene promoter, the SV40 early promoter, the long terminal repeat of Rous sarcoma virus, [β]-actin, the rat insulin promoter, the phosphoglycerate kinase promoter, the human α-1 antitrypsin (hAAT) promoter, the transthyretin promoter, the TBG promoter and other liver-specific promoters, the desmin promoter and similar muscle-specific promoters, the EF1-α-promoter, multi-tissue specific hybrid promoters, and neuron-specific promoters such as synapsin and glyceraldehyde-3-phosphate dehydrogenase promoters. All of these are well known to those skilled in the art, readily available, and can be used to achieve high levels of expression of the target coding sequence. In addition, sequences derived from non-viral genes, such as the mouse metallothionein gene, may also be useful in this invention. Such promoter sequences are commercially available, for example, from Stratagene (San Diego, CA). Further exemplary promoter sequences are described, for example, in International Publication No. 2018213786A1 (which is incorporated herein by reference in its entirety).

[0431] In some embodiments, an apolipoprotein E enhancer (ApoE) or a functional fragment thereof is used to drive expression, for example, in the liver. In some embodiments, two copies of the ApoE enhancer or its functional fragment are used. In some embodiments, the ApoE enhancer or its functional fragment is used in combination with a promoter, for example, a human α-1 antitrypsin (hAAT) promoter.

[0432] In some embodiments, regulatory sequences confer tissue-specific gene expression capabilities. In some cases, tissue-specific regulatory sequences bind to tissue-specific transcription factors that induce tissue-specific transcription. Various tissue-specific regulatory sequences (e.g., promoters, enhancers, etc.) are known in the art. Exemplary tissue-specific regulatory sequences include, but are not limited to, the following tissue-specific promoters: liver-specific thyroxine-binding globulin (TBG) promoter, insulin promoter, glucagon promoter, somatostatin promoter, pancreatic polypeptide (PPY) promoter, synapsin-1 (Syn) promoter, creatine kinase (MCK) promoter, mammalian desmin (DES) promoter, α-myosin heavy chain (α-MHC) promoter, or cardiac troponin T (cTnT) promoter. Other exemplary promoters include the β-actin promoter, the hepatitis B virus core promoter (Sandig et al., Gene Ther., 3:1002-9 (1996)), the α-fetoprotein (AFP) promoter (Arbuthnot et al., Hum. Gene Ther., 7:1503-14 (1996)), the bone osteocalcin promoter (Stein et al., Mol. Biol. Rep., 24:185-96 (1997)), the bone sialoprotein promoter (Chen et al., J. Bone Miner. Res., 11:654-64 (1996)), and the CD2 promoter (Hansal et al. Examples include neuronal promoters such as immunoglobulin heavy chain promoters, T cell receptor α-chain promoters, and neuron-specific enolase (NSE) promoters (Andersen et al., Cell.Mol.Neurobiol., 13:503-15 (1993)), neurofilament light chain gene promoters (Piccioli et al., Proc. Natl.Acad.Sci.USA, 88:5611-5 (1991)), and neuron-specific vgf gene promoters (Piccioli et al., Neuron, 15:373-84 (1995)).Other exemplary promoter sequences are described, for example, in U.S. Patent No. 1,0300,146 (which is incorporated herein by reference in its entirety). In some embodiments, tissue-specific regulatory elements, such as tissue-specific promoters, are selected from those known to be operably linked to genes highly expressed in a given tissue, as measured, for example, by RNA-seq, protein expression data, or a combination thereof. Methods for analyzing tissue specificity by expression are taught in Fagerberg et al. Mol Cell Proteomics 13(2):397-406 (2014) (which is incorporated herein by reference in its entirety).

[0433] In some embodiments, the vectors described herein are multicistronic expression constructs. Examples of multicistronic expression constructs include constructs having a first expression cassette (e.g., including a first promoter and a first coding nucleic acid sequence) and a second expression cassette (e.g., including a second promoter and a second coding nucleic acid sequence). Such multicistronic expression constructs may, in some cases, be particularly useful for the delivery of untranslated gene products, such as hairpin RNA, together with polypeptides, e.g., recombinant polypeptides and recombinant templates. In some embodiments, multicistronic expression constructs may exhibit reduced expression levels of one or more of the included transgenes, for example, due to promoter interference or the proximity of incompatible nucleic acid elements. When a multicistronic expression construct is part of a viral vector, the presence of self-complementary nucleic acid sequences may, in some cases, interfere with the formation of structures necessary for viral replication or packaging.

[0434] In some embodiments, the sequence encodes RNA having a hairpin. In some embodiments, the hairpin RNA is a guide RNA, template RNA, shRNA, or microRNA. In some embodiments, the first promoter is the RNA polymerase I promoter. In some embodiments, the first promoter is the RNA polymerase II promoter. In some embodiments, the second promoter is the RNA polymerase III promoter. In some embodiments, the second promoter is the U6 or H1 promoter.

[0435] While not intended to be bound by any particular theory, multi-cistronic expression constructs are thought not to achieve optimal expression levels compared to expression systems containing only one cistron. One of the reasons shown for the reduced expression levels achieved using multi-cistronic expression constructs containing two or more promoter elements is the phenomenon of promoter interference (see, for example, Curtin JA, Dane AP, Swanson A, Alexander IE, Ginn S L. Bidirectional promoter interference between two widely used internal heterologous promoters in a late-generation lentiviral construct. Gene Ther. 2008 March;15(5):384-90 and Martin-Duque P, Jezzard S, Kaftansis L, Vassaux G. Direct comparison of the insulating properties of two genetic elements in an adenoviral vector containing two different expression cassettes. Hum Gene Ther. 2004 October;15(10):995-1002; both references are incorporated herein by reference to disclose the phenomenon of promoter interference). In some embodiments, the promoter interference problem can be resolved, for example, by generating a multi-cistronic expression construct containing only one promoter that drives the transcription of multiple coding nucleic acid sequences separated by internal ribosome entry sites, or by isolating a cistron containing its own promoter equipped with a transcription insulator element. In some embodiments, multi-cistronic expression driven by a single promoter may result in heterogeneous expression levels of the cistrons. In some embodiments, promoters cannot be efficiently isolated, and the isolated elements may not be compatible with certain gene transfer vectors, such as certain retroviral vectors.

[0436] microRNA microRNAs (miRNAs) and other small interfering nucleic acids generally regulate gene expression by cleaving / degrading target RNA transcripts or repressing the translation of target messenger RNA (mRNA). In some cases, miRNAs can be expressed naturally, typically as the final 19-25 untranslated RNA product. miRNAs generally exhibit their activity through sequence-specific interactions with the 3' untranslated region (UTR) of target mRNA. These endogenously expressed miRNAs may form hairpin precursors, which are later processed into miRNA duplexes and then into mature single-stranded miRNA molecules. This mature miRNA generally induces a multiprotein complex, miRISC, which identifies the target 3'UTR region of target mRNA based on its complementarity to the mature miRNA. Useful transgene products include, for example, miRNAs or miRNA-binding sites that regulate the expression of linked polypeptides. A non-exclusive list of miRNA genes, their products, and their homologs are useful as targets for transgenes or small interfering nucleic acids (e.g., miRNA sponges, antisense oligonucleotides) in ways such as those enumerated in U.S. Patent No. 10300146, 22:25-25:48, which is incorporated herein by reference. In some embodiments, one or more binding sites for one or more of the aforementioned miRNAs are incorporated into a transgene, such as a transgene delivered by an rAAV vector, to inhibit the expression of the transgene in one or more tissues of an animal carrying the transgene. In some embodiments, binding sites may be selected to control the expression of the transgene in a tissue-specific manner. For example, a binding site for liver-specific miR-122 can be incorporated into a transgene to inhibit the expression of that transgene in the liver. Other exemplary miRNA sequences are described, for example, in U.S. Patent No. 10300146, which is incorporated herein by reference in its entirety.

[0437] miR inhibitors or miRNA inhibitors are generally agents that block miRNA expression and / or processing. Examples of such agents, but not limited to, include microRNA antagonists, microRNA-specific antisense, microRNA sponges, and microRNA oligonucleotides (double-stranded, hairpin, and short oligonucleotides), which inhibit miRNA interaction with the Drosha complex. MicroRNA inhibitors, such as miRNA sponges, can be expressed in cells from transgenes (e.g., as described in Ebert, MS Nature Methods, Epub Aug. 12, 2007, the entire text of which is incorporated herein by reference). In some embodiments, microRNA sponges or other miR inhibitors are used in conjunction with AAV. MicroRNA sponges generally specifically inhibit miRNAs via complementary heptamer seed sequences. In some embodiments, a single sponge sequence can be used to silence an entire family of miRNAs. Other methods for silencing miRNA function in cells (de-suppression of miRNA targets) will be apparent to those skilled in the art.

[0438] In some embodiments, the recombinant system, template RNA, or polypeptide described herein is administered to a target tissue, e.g., a first tissue, or is active there (e.g., more active). In some embodiments, the recombinant system, template RNA, or polypeptide is not administered to a non-target tissue or is less active in the non-target tissue (e.g., inactive). In some embodiments, the recombinant system, template RNA, or polypeptide described herein is useful for modifying DNA in a target tissue, e.g., a first tissue (and not modifying DNA in a non-target tissue, e.g.).

[0439] In some embodiments, the recombination system comprises (a) a polypeptide as described herein or a nucleic acid encoding it, (b) a template nucleic acid as described herein (e.g., template RNA), and (c) one or more primary tissue-specific regulatory sequences specific to a target tissue, wherein one or more primary tissue-specific regulatory sequences specific to a target tissue are functionally related to (a), (b), or (a) and (b), and if related to (a), (a) comprises a nucleic acid encoding a polypeptide.

[0440] In some embodiments, the nucleic acid in (b) includes RNA.

[0441] In some embodiments, the nucleic acid in (b) includes DNA.

[0442] In some embodiments, the nucleic acid of (b) is (i) single-stranded or includes a single-stranded segment, for example, single-stranded DNA, or includes a single-stranded segment and one or more double-stranded segments, (ii) has a reverse-end repeat, or (iii) both (i) and (ii).

[0443] In some embodiments, the nucleic acid in (b) is double-stranded or includes double-stranded segments.

[0444] In some embodiments, (a) comprises a nucleic acid encoding a polypeptide.

[0445] In some embodiments, the nucleic acid in (a) includes RNA.

[0446] In some embodiments, the nucleic acid in (a) includes DNA.

[0447] In some embodiments, the nucleic acid of (a) is (i) single-stranded or includes single-stranded segments, for example, single-stranded DNA, or includes single-stranded segments and one or more double-stranded segments, (ii) has a reverse-end repeat, or (iii) both (i) and (ii).

[0448] In some embodiments, the nucleic acid of (a) is double-stranded or contains double-stranded segments.

[0449] In some embodiments, the nucleic acids in (a), (b), or (a) and (b) are linear.

[0450] In some embodiments, the nucleic acids in (a), (b), or (a) and (b) are circular, for example, plasmids or minicircles.

[0451] In some embodiments, heterogeneous target arrays are functionally related to the first promoter.

[0452] In some embodiments, one or more first tissue-specific expression regulatory sequences include a tissue-specific promoter.

[0453] In some embodiments, the tissue-specific promoter includes a first promoter functionally related to: (i) a heterologous target sequence, (ii) a nucleic acid encoding a retroviral RT, or (iii) (i) and (ii).

[0454] In some embodiments, one or more first tissue-specific regulatory sequences include tissue-specific microRNA recognition sequences functionally related to: (i) heterologous target sequences, (ii) nucleic acids encoding retroviral RT domains, or (iii)(i) and (ii).

[0455] In some embodiments, the system includes a tissue-specific promoter, and the system further includes one or more tissue-specific microRNA recognition sequences, wherein (i) the tissue-specific promoter is functionally related to (I) a heterologous target sequence, (II) a nucleic acid encoding a retroviral RT domain, or (III) (i) and (ii), and / or (ii) one or more tissue-specific microRNA recognition sequences are functionally related to (I) a heterologous target sequence, (II) a nucleic acid encoding a retroviral RT, or (III) (i) and (ii).

[0456] In some embodiments, (a) comprises a nucleic acid encoding a polypeptide, the nucleic acid comprising a promoter functionally associated with the nucleic acid encoding the polypeptide.

[0457] In some embodiments, the nucleic acid encoding the polypeptide includes one or more secondary tissue-specific expression regulatory sequences that are functionally related to the polypeptide coding sequence and are specific to the target tissue.

[0458] In some embodiments, one or more second tissue-specific expression regulatory sequences include a tissue-specific promoter.

[0459] In some embodiments, the tissue-specific promoter is a promoter functionally associated with a nucleic acid encoding a polypeptide.

[0460] In some embodiments, one or more second tissue-specific expression regulatory sequences include tissue-specific microRNA recognition sequences.

[0461] In some embodiments, the promoter functionally associated with the nucleic acid encoding the polypeptide is a tissue-specific promoter, and the system further comprises one or more tissue-specific microRNA recognition sequences.

[0462] In some embodiments, the nucleic acid component of the system provided by the present invention is a sequence (e.g., encoding a polypeptide or including a heterologous target sequence) adjacent to an untranslated region (UTR) that modifies the protein expression level. Various 5' and 3' UTRs can affect protein expression. For example, in some embodiments, the coding sequence may be preceded by a 5' UTR that modifies RNA stability or protein translation. In some embodiments, the sequence may be followed by a 3' UTR that modifies RNA stability or translation. In some embodiments, the sequence may be preceded by a 5' UTR and followed by a 3' UTR to modify RNA stability or translation. In some embodiments, the 5' and / or 3' UTRs may be selected from the 5' and 3' UTRs of complement factor 3 (C3) (CACTCCTCCCCATCCTCTCCCTCTGTCCCTCTGTCCCTCTGACCCTGCACTGTCCCAGCACC, SEQ ID NO: 11,004) or orosomucoid 1 (ORM1) (CAGGACACAGCCTTGGATCAGGACAGAGACTTGGGGGCCATCCTGCCCCTCCAACCCGACATGTGTACCTCAGCTTTTTCCCTCACTTGCATCAATAAAGCTTCTGTGTTTGGAACAGCTAA, SEQ ID NO: 11,005) (Asrani et al. RNA Biology 2018). In certain embodiments, the 5' UTR is the 5' UTR from C3 and the 3' UTR is the 3' UTR from ORM1. In certain embodiments, the 5'UTR and 3'UTR for protein expression, for example, recombinant polypeptide or mRNA (or RNA-encoding DNA) for a heterologous target sequence, include optimized expression sequences.In some embodiments, for example, as described in Richner et al. Cell 168(6):P1114-1125(2017) (the sequence of which is inc...

Claims

1. A nucleic acid molecule encoding a genetically modified polypeptide, for example, from 5' to 3', (a) 5' UTR of sequences 41-44 or sequences having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (b) A region encoding the N-terminal NLS of sequence sequence number 36 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (c) A region encoding the Cas domain of sequence number 52 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (d) A region encoding a linker of sequence sequence 54 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (e) A region encoding the reverse transcriptase (RT) domain of sequence number 56 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (f) A region encoding the C-terminal NLS of sequence number 38 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (g) The 3' UTR of sequences 45-48 or sequences having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (h) optionally, an expression element of sequence number 40, 101, or 102, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and (i) Poly(A)tails of sequences having at least 90%, 95%, 96%, 97%, 98%, or 99% identity with sequence numbers 49-51. Nucleic acid molecules containing these molecules.

2. A nucleic acid molecule encoding a genetically modified polypeptide, for example, from 5' to 3', (a) 5' UTR of sequences 41-44 or sequences having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (b) A region encoding the N-terminal NLS of sequence number 37 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (c) A region encoding the Cas domain of sequence number 53 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (d) A region encoding a linker of sequence sequence 55 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (e) A region encoding the reverse transcriptase (RT) domain of sequence number 57 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (f) A region encoding the C-terminal NLS of sequence number 39 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (g) The 3' UTR of sequences 45-48 or sequences having at least 90%, 95%, 96%, 97%, 98%, or 99% identity therewith, (h) optionally, an expression element of sequence number 40, 101, or 102, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and (i) Poly(A)tails of sequences having at least 90%, 95%, 96%, 97%, 98%, or 99% identity with sequence numbers 49-51. Nucleic acid molecules containing these molecules.

3. The nucleic acid molecule according to claim 1 or 2, wherein the Cas domain includes a Cas domain that binds to a target DNA molecule and is heterogeneous to the RT domain.

4. A nucleic acid molecule according to any one of claims 1 to 3, which is mRNA.

5. It is template RNA, (1) gRNA spacer, (2) gRNA scaffold, (3) heterologous sequences and (4) Primer binding site (PBS) sequence A template RNA containing the nucleotide sequence of the template RNA of SEQ ID NO: 15, 17, 92, or 94, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.

6. A template RNA containing the nucleotide sequence of sequence number 37638 or a sequence having at least 95%, 97%, 98%, or 99% identity thereto.

7. A template RNA containing the nucleotide sequence of sequence number 37653 or a sequence having at least 95%, 97%, 98%, or 99% identity thereto.

8. It is a genetically modified system, (a) template RNA (tgRNA), which has a 5' to 3' position (1) gRNA spacer, (2) gRNA scaffold, (3) heterologous sequences and (4) Primer binding site (PBS) sequence A template RNA (tgRNA) containing the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, (b) A recombinant polypeptide or a nucleic acid encoding the recombinant polypeptide, wherein the recombinant polypeptide is (1) Cas domain, (2) Linker, and (3) Reverse transcriptase (RT) domain The recombinant polypeptide includes the amino acid sequence of the recombinant polypeptide of SEQ ID NO: 28 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and the recombinant polypeptide or nucleic acid A genetically modified system that includes this.

9. It is a genetically modified system, (a) template RNA (tgRNA), which has a 5' to 3' position (1) gRNA spacers having the sequence of the template RNA of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having one, two or three or fewer sequence mutations (e.g., substitutions) therewith, (2) gRNA scaffold sequences of the template RNAs in Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or gRNA scaffolds having sequences with 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 or fewer sequence mutations therein, (3) A heterologous target sequence of the template RNA in Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a heterologous target sequence having one, two or three or fewer sequence mutations therein, and (4) Primer binding site (PBS) sequences having the PBS sequence of the template RNA in Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or sequences having one, two or three or fewer sequence mutations therein. template RNA (tgRNA) containing, (b) A recombinant polypeptide or a nucleic acid encoding the recombinant polypeptide, wherein the recombinant polypeptide is (1) Cas domain, (2) Linker, and (3) Reverse transcriptase (RT) domain The recombinant polypeptide includes the amino acid sequence of the recombinant polypeptide of SEQ ID NO: 28 or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and the recombinant polypeptide or nucleic acid A genetically modified system that includes this.

10. The system according to claim 8 or 9, wherein the template RNA includes a nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having at least 95% identity thereto.

11. The system according to claim 8 or 9, wherein the template RNA includes a nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having at least 97%, 98%, or 99% identity thereto.

12. The system according to claim 8 or 9, wherein the template RNA includes the nucleotide sequence of the template RNA sequence shown in Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3.

13. The system according to any one of claims 8 to 12, wherein the genetically modified polypeptide comprises the amino acid sequence of the genetically modified polypeptide of SEQ ID NO: 28 or a sequence having at least 95% identity thereto.

14. The system according to any one of claims 8 to 12, wherein the genetically modified polypeptide comprises the amino acid sequence of the genetically modified polypeptide of SEQ ID NO: 28 or a sequence having at least 99% identity thereto.

15. The system according to any one of claims 8 to 12, wherein the genetically modified polypeptide comprises the amino acid sequence of the genetically modified polypeptide of SEQ ID NO:

28.

16. It is a genetically modified system, (a) template RNA (tgRNA), which has a 5' to 3' position (1) gRNA spacer, (2) gRNA scaffold, (3) heterologous sequences and (4) Primer binding site (PBS) sequence A template RNA (tgRNA) containing the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, (b) A nucleic acid (e.g., mRNA) that encodes a genetically modified polypeptide, (1) Cas domain, (2) Linker, and (3) Reverse transcriptase (RT) domain The nucleotide encoding the recombinant polypeptide comprises a nucleic acid (e.g., mRNA) and a nucleic acid according to any one of claims 1 to 4. A genetically modified system that includes this.

17. It is a genetically modified system, (a) template RNA (tgRNA), which has a 5' to 3' position (1) gRNA spacer, (2) gRNA scaffold, (3) heterologous sequences and (4) Primer binding site (PBS) sequence A template RNA (tgRNA) containing the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, (b) A nucleic acid (e.g., mRNA) that encodes a genetically modified polypeptide, (1) Cas domain, (2) Linker, and (3) Reverse transcriptase (RT) domain The nucleotides encoding the recombinant polypeptide include nucleic acids (e.g., mRNA) and sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with the recombinant polypeptide sequence of Table N2, E11, or E15. A genetically modified system that includes this.

18. It is a genetically modified system, (a) template RNA (tgRNA), which has a 5' to 3' position (1) gRNA spacers having the sequence of the template RNA of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having one, two or three or fewer sequence mutations (e.g., substitutions) therewith, (2) gRNA scaffold sequences of the template RNAs in Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or gRNA scaffolds having sequences with 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 or fewer sequence mutations therein, (3) A heterologous target sequence of the template RNA in Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a heterologous target sequence having one, two or three or fewer sequence mutations therein, and (4) Primer binding site (PBS) sequences having the PBS sequence of the template RNA in Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or sequences having one, two or three or fewer sequence mutations therein. template RNA (tgRNA) containing, (b) A nucleic acid (e.g., mRNA) that encodes a genetically modified polypeptide, (1) Cas domain, (2) Linker, and (3) Reverse transcriptase (RT) domain The nucleotides encoding the recombinant polypeptide include nucleic acids (e.g., mRNA) and sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with the recombinant polypeptide sequence of Table N2, E11, or E15. A genetically modified system that includes this.

19. The system according to any one of claims 16 to 18, wherein the template RNA includes a nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having at least 95% identity thereto.

20. The system according to any one of claims 16 to 18, wherein the template RNA includes a nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having at least 97%, 98%, or 99% identity thereto.

21. The system according to any one of claims 16 to 18, wherein the template RNA includes the nucleotide sequence of the template RNA sequence shown in Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3.

22. The system according to any one of claims 16 to 21, wherein the nucleotide encoding the genetically modified polypeptide includes the nucleic acid sequence of the genetically modified polypeptide of Table N2, E11, or E15, or a sequence having at least 95% identity thereto.

23. The system according to any one of claims 16 to 21, wherein the nucleotide encoding the genetically modified polypeptide includes the nucleic acid sequence of the genetically modified polypeptide of Table N2, E11, or E15, or a sequence having at least 99% identity thereto.

24. The system according to any one of claims 16 to 21, wherein the nucleotide encoding the genetically modified polypeptide comprises the nucleic acid sequence of the genetically modified polypeptide in Table N2, E11, or E15.

25. The system according to any one of claims 16 to 24, further comprising a second nick RNA (ngRNA) that leads a second nick to the second strand of the human PAH gene.

26. The system according to claim 25, wherein the ngRNA includes the sequences of the ngRNAs in Table 2A, E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14 or E14A, or sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.

27. The aforementioned ngRNA is divided from 5' to 3', (1) gRNA spacer sequences of ngRNAs of Table 2A, E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14 or E14A, or gRNA spacers having sequences with one, two or three or fewer sequence mutations (e.g., substitutions), and (2) gRNA scaffold sequences of the ngRNAs in Table 2A, E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14 or E14A, or gRNA scaffolds having sequences with 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 or fewer sequence mutations therein The system according to claim 25, including the system described in claim 25.

28. The system according to any one of claims 25 to 27, wherein the ngRNA, together with the template RNA of the genetically modified system, has a "PAM inward orientation".

29. The system according to any one of claims 8 to 28, wherein the nucleic acid encoding the genetically modified polypeptide includes RNA, for example, mRNA.

30. The nucleic acid molecule is formulated in lipid nanoparticles (LNPs) and is either the nucleic acid molecule according to any one of claims 1 to 4 or the template RNA according to any one of claims 5 to 7.

31. The system according to any one of claims 8 to 29, wherein the tgRNA, the nucleic acid molecule encoding the recombinant polypeptide and / or the ngRNA are formulated in LNP.

32. A pharmaceutical composition comprising a system according to any one of claims 8 to 29 or one or more nucleic acids encoding it, and a pharmaceutically acceptable excipient or carrier.

33. The pharmaceutical composition according to claim 32, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles (LNPs).

34. The pharmaceutical composition according to claim 33, wherein the viral vector is an adeno-associated virus.

35. A host cell (e.g., a mammalian cell, e.g., a human cell) comprising a nucleic acid molecule, a genetically modified system, or a template RNA according to any one of claims 1 to 31.

36. A method for producing a nucleic acid molecule or template RNA according to any one of claims 1 to 7 or 30, comprising synthesizing the nucleic acid molecule or template RNA in vitro (for example, by in vitro transcription or solid-phase synthesis) or by introducing DNA encoding the template RNA into a host cell under conditions that enable the production of the template RNA.

37. A method for modifying a target site in a human PAH gene within a cell, comprising contacting the cell with a genetic recombination system according to any one of claims 8 to 29 or 31, or DNA encoding the same, or a pharmaceutical composition according to any one of claims 32 to 34, thereby modifying the target site in the human PAH gene within the cell.

38. A method for treating a subject having a disease or symptoms associated with a mutation in the human PAH gene, comprising administering to the subject a genetically modified system according to any one of claims 8 to 29 or 31, or DNA encoding the same, or a pharmaceutical composition according to any one of claims 32 to 34, thereby treating the subject having a disease or symptoms associated with a mutation in the human PAH gene.

39. The method according to claim 38, wherein the disease or symptom is phenylketonuria (PKU) or hyperphenylalaninemia (for example, mild or severe hyperphenylalaninemia).

40. The method according to claim 38 or 39, wherein the subject has the R408W mutation.

41. A method for treating a subject having a PKU, comprising administering to the subject a genetically modified system according to any one of claims 8 to 29 or 31, or DNA encoding the same, or a pharmaceutical composition according to any one of claims 32 to 34, thereby treating the subject having a PKU.

42. The genetic recombination system or method according to any one of claims 8 to 29, 31, or 36 to 41, wherein the introduction of the system into target cells results in the modification of the pathogenic mutation of the PAH gene.

43. The genetic recombination system or method according to any one of claims 8 to 29, 31, or 36 to 42, wherein the pathogenic mutation is an R408W mutation, and the modification comprises an amino acid substitution of W408R.

44. The genetic recombination system or method according to any one of claims 8 to 29, 31, or 36 to 43, wherein the introduction of the system into target cells results in a mutation that causes the restoration of the function of the PAH gene.

45. The genetic recombination system or method according to any one of claims 8 to 29, 31, or 36 to 44, wherein the modification of the mutation occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70%, or more) of the target nucleic acid.

46. The genetic recombination system or method according to any one of claims 8 to 29, 31, or 36 to 45, wherein the modification of the mutation occurs in at least 10% of the target cells (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70%, or more).

47. The genetic recombination system or method according to any one of claims 8 to 29, 31 or 36 to 46, wherein the genetic recombination system comprises a gRNA that targets a second strand, and the modification of the mutation in a population of target cells is increased compared to a population of target cells treated with a genetic recombination system comprising a template RNA without a gRNA that targets a second strand.

48. The method according to any one of claims 36 to 47, wherein the cells are mammalian cells such as human cells.

49. The method according to any one of claims 36 to 48, wherein the subject is a human.

50. The method according to any one of claims 36 to 49, wherein the contact is performed ex vivo, and for example, the DNA of the cell or subject is modified ex vivo.

51. The method according to any one of claims 36 to 50, wherein the contact is performed in vivo, and for example, the DNA of the cell or subject is modified in vivo.

52. The method according to any one of claims 36 to 51, wherein contacting the cells or object with the system includes contacting the cells or cells within the object with nucleic acids (e.g., DNA or RNA) encoding the recombinant polypeptide under conditions that enable the production of the recombinant polypeptide.