Gene editing system comprising reverse transcriptase

By developing a fusion protein gene editing system, using a linker to connect reverse transcriptase and nickase or nuclease, and combining guide RNA, the problem of insufficient accuracy and efficiency of gene editing in the prior art is solved, and efficient and controllable polynucleotide editing is achieved.

CN120225685APending Publication Date: 2025-06-27METAGENOMI INC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202380080715.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-28
Filing Date
2023-10-18
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

There are shortcomings in the accuracy and efficiency of existing gene editing technologies, especially in polynucleotide editing and targeted editing, which are difficult to achieve efficient and controllable results.

Method used

A fusion protein gene editing system was developed that enables precise editing of target nucleic acids by using a linker to connect reverse transcriptase and nickase or nuclease, and combining guide RNA.

Benefits of technology

This system significantly improves the accuracy and efficiency of gene editing, can achieve efficient polynucleotide editing and targeted editing in cells, and reduces editing error rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120225685A_ABST
    Figure CN120225685A_ABST
Patent Text Reader

Abstract

The present disclosure relates generally to gene editing systems comprising reverse transcriptases and fusion proteins of reverse transcriptases with nickase or nuclease, methods of making such reverse transcriptases and fusion proteins, and methods of using such reverse transcriptases and fusion proteins for site-directed genome editing in cells.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 380,194, filed on October 19, 2022; U.S. Provisional Patent Application No. 63 / 386,658, filed on December 8, 2022; U.S. Provisional Patent Application No. 63 / 387,268, filed on December 13, 2022; U.S. Provisional Patent Application No. 63 / 491,269, filed on March 20, 2023; U.S. Provisional Patent Application No. 63 / 500,228, filed on May 4, 2023; U.S. Provisional Patent Application No. 63 / 500,509, filed on May 5, 2023; and U.S. Provisional Patent Application No. 63 / 510,861, filed on June 28, 2023, each of which is incorporated herein by reference in its entirety. Summary of the Invention

[0003] The present disclosure is in part based on the development of gene editing systems that include a reverse transcriptase, a nuclease or nickase, and a guide RNA or pegRNA.

[0004] Described herein are fusion proteins that include a nickase linked to a reverse transcriptase using a linker, wherein the reverse transcriptase has at least about 80% sequence identity to any one of SEQ ID NOs: 161 - 629, 767 - 1220, 1959 - 2522, and 2582 - 2585.

[0005] Described herein are fusion proteins that include a nuclease linked to a reverse transcriptase using a linker, wherein the reverse transcriptase has at least about 80% sequence identity to any one of SEQ ID NOs: 161 - 629, 767 - 1220, 1959 - 2522, and 2582 - 2585.

[0006] Described herein are fusion proteins that include a catalytically - dead nuclease linked to a reverse transcriptase using a linker, wherein the reverse transcriptase has at least about 80% sequence identity to any one of SEQ ID NOs: 161 - 629, 767 - 1220, 1959 - 2522, and 2582 - 2585.

[0007] This document describes a gene editing system that includes: a) a nickase; b) a guide nucleic acid that is configured to form a complex with the nickase and hybridize with a target nucleic acid sequence; and c) a reverse transcriptase that has at least about 80% sequence identity with any one of SEQ ID NOs: 161-629, 767-1220, 1959-2522, and 2582-2585 and is configured to form a complex with the nickase. In some embodiments, the gene editing system further includes a nucleic acid template. In some embodiments, the nickase is a modified endonuclease. In some embodiments, the modified endonuclease is a type II CRISPR endonuclease. In some embodiments, the modified endonuclease is a type V CRISPR endonuclease. In some embodiments, the type II CRISPR endonuclease or the type V CRISPR endonuclease has nickase activity. In some embodiments, the modified endonuclease is selected from the group consisting of: spCas9(H840A), spCas9(D10A), nMG3-6(D13A), nMG3-6(H586A), nMG3-6(N609A), Cas12a, and MG29-1. In some embodiments, the modified endonuclease has at least about 80% sequence identity with any one of SEQ ID NOs: 152-154. In some embodiments, the nickase and the reverse transcriptase are linked. In some embodiments, the nickase and the reverse transcriptase are linked by a linker. In some embodiments, the linker includes at least 10, 20, or 30 amino acids. In some embodiments, the linker includes about 30-35 amino acids. In some embodiments, the linker includes about 30 amino acids. In some embodiments, the linker has at least 80% sequence identity with SEQ ID NO: 103. In some embodiments, the linker has at least 80% sequence identity with any one of SEQ ID NOs: 155-160. In some embodiments, the nickase and the reverse transcriptase are not linked. In some embodiments, the guide nucleic acid includes a spacer sequence and a crRNA. In some embodiments, the guide nucleic acid further includes a reverse transcriptase template (RTT). In some embodiments, the bases in the RTT include bulky modifications selected from the group consisting of: complex sugars or complex amino groups and / or other modifications compatible with RNA. In some embodiments, the guide nucleic acid further includes a primer binding site. In some embodiments, the primer binding site is located at the 3' end of the guide nucleic acid. In some embodiments, the primer binding site includes at least 2, 4, 6, 8, 10, 13, 16, 20, 24, 28, 32, 36, 40, 45, 50, 55, 60, or 65 nucleotides.In some embodiments, the gene editing system further comprises a transposase, an integrase or a homing endonuclease. In some embodiments, the gene editing system further comprises a retrotransposon. In some embodiments, the reverse transcriptase has a processive synthesis ability that is at least about 2 times that of Moloney murine leukemia virus (MMLV) reverse transcriptase. In some embodiments, the reverse transcriptase has a processive synthesis ability that is at most about 1 / 2 that of Moloney murine leukemia virus (MMLV) reverse transcriptase. In some embodiments, the reverse transcriptase has an error rate of less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10% or 0.05%. In some embodiments, compared with Moloney murine leukemia virus (MMLV) reverse transcriptase, the reverse transcriptase has an error rate of less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10% or 0.05%.

[0008] This document describes a gene editing system that includes: a) a nuclease; b) a guide nucleic acid that is configured to form a complex with the nuclease and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase that has at least about 80% sequence identity with any of SEQ ID NOs: 161-629, 767-1220, 1959-2522, and 2582-2585 and is configured to form a complex with the nuclease. In some embodiments, the gene editing system further includes a nucleic acid template. In some embodiments, the nuclease is a double-stranded nuclease. In some embodiments, the nuclease is a type II CRISPR endonuclease. In some embodiments, the CRISPR endonuclease is Cas9. In some embodiments, the Cas9 is catalytically dead Cas9 (dCas9). In some embodiments, the nuclease and the reverse transcriptase are linked. In some embodiments, the nuclease and the reverse transcriptase are linked by a linker. In some embodiments, the linker includes at least 10, 20, or 30 amino acids. In some embodiments, the linker includes about 30-35 amino acids. In some embodiments, the linker includes about 30 amino acids. In some embodiments, the linker has at least 80% sequence identity with SEQ ID NO: 103. In some embodiments, the linker has at least about 80% sequence identity with any of SEQ ID NOs: 155-160. In some embodiments, the nuclease and the reverse transcriptase are not linked. In some embodiments, the guide nucleic acid further includes a primer binding site. In some embodiments, the primer binding site is located at the 3' end of the guide nucleic acid. In some embodiments, the primer binding site includes at least 2, 4, 6, 8, 10, 13, 16, 20, 24, 28, 32, 36, 40, 45, 50, 55, 60, or 65 nucleotides. In some embodiments, the gene editing system further includes a transposase, integrase, or homing endonuclease. In some embodiments, the gene editing system further includes a retrotransposon. In some embodiments, the reverse transcriptase has a processive synthesis ability that is at least about 2 times that of Moloney murine leukemia virus (MMLV) reverse transcriptase. In some embodiments, the reverse transcriptase has a processive synthesis ability that is at most about 1 / 2 that of Moloney murine leukemia virus (MMLV) reverse transcriptase. In some embodiments, the reverse transcriptase has an error rate of less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10%, or 0.05%.In some embodiments, compared to Moloney murine leukemia virus (MMLV) reverse transcriptase, the reverse transcriptase has an error rate of less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10%, or 0.05%.

[0009] The present disclosure describes a gene editing system that includes: a) a nickase; b) a guide nucleic acid configured to form a complex with the nickase and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase configured to form a complex with the nickase, the reverse transcriptase having an X1X2DD motif, where X1 is F or Y, and where when X1 is Y, X2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y. In some embodiments, the X2 is A or I. In some embodiments, the X1X2DD motif is YADD (SEQ ID NO: 2572) or YIDD (SEQ ID NO: 2573). In some embodiments, the X1X2DD motif is FADD (SEQ ID NO: 2574), FVDD (SEQ ID NO: 2575), FIDD (SEQ ID NO: 2576), or FLDD (SEQ ID NO: 2577). In some embodiments, the reverse transcriptase has at least about 80% sequence identity to any one of SEQ ID NOs: 161-629, 767-1220, 1959-2522, and 2582-2585.

[0010] This disclosure describes gene editing systems that include: a) a nuclease; b) a guide nucleic acid that is configured to form a complex with the nuclease and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase that is configured to form a complex with the nuclease, the reverse transcriptase having an X1X2DD motif, where X1 is F or Y, and where when X1 is Y, X2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y. In some embodiments, the X2 is A or I. In some embodiments, the X1X2DD motif is YADD (SEQ ID NO:2572) or YIDD (SEQ ID NO:2573). In some embodiments, the X1X2DD motif is FADD (SEQ ID NO:2574), FVDD (SEQ ID NO:2575), FIDD (SEQ ID NO:2576), or FLDD (SEQ ID NO:2577). In some embodiments, the reverse transcriptase has at least about 80% sequence identity to any one of SEQ ID NOs: 161-629, 767-1220, 1959-2522, and 2582-2585.

[0011] This disclosure describes an isolated reverse transcriptase that has at least about 80% sequence identity to any one of SEQ ID NOs: 161-629, 767-1220, 1959-2522, and 2582-2585.

[0012] This disclosure describes a nucleic acid that encodes the fusion protein or gene editing system described above. In some embodiments, the nucleic acid is DNA or RNA. In some embodiments, the RNA is mRNA. In some embodiments, the nucleic acid is included in a vector. In some embodiments, the nucleic acid or the vector containing the nucleic acid is included in an adeno-associated virus or a lipid nanoparticle. In some embodiments, the nucleic acid or the vector containing the nucleic acid is included in a cell. In some embodiments, the cell is a human cell.

[0013] This disclosure describes methods for modifying double-stranded and / or single-stranded nucleic acids that include contacting a cell with the fusion protein or gene editing system described above.

[0014] This disclosure describes methods for modifying double-stranded and / or single-stranded nucleic acids in cells, the methods comprising: a) providing a guide nucleic acid to the cell to bind to a target strand of the nucleic acid; b) providing a nuclease or nickase to the cell to cleave the nucleic acid at the binding position of the guide nucleic acid; c) providing a reverse transcriptase to the cell to synthesize a modification in the target strand of the nucleic acid at the position cleaved by the nickase and / or the nuclease. In some embodiments, the reverse transcriptase has at least about 80% sequence identity to any one of SEQ ID NOs: 161-629, 767-1220, 1959-2522, and 2582-2585. In some embodiments, the modification is an insertion, deletion, or mutation. In some embodiments, the method further comprises providing an RNA or DNA template to the cell. In some embodiments, the nucleic acid is genomic or a vector. In some embodiments, the method further comprises providing a transposase, integrase, or homing endonuclease to the cell. In some embodiments, the method further comprises providing a retrotransposon to the cell. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The novel features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the disclosure will be obtained from the following detailed description that illustrates illustrative embodiments (wherein the principles of the disclosure are utilized), in the drawings:

[0016] Figure 1A - 1JJ is a bar graph showing the percentage of conversion editing of G to T for untethered reverse transcriptase (RT) candidates from the MG151 family, the candidates having eight different primer binding site (PBS) nucleotides of different lengths (PBS lengths of 2, 4, 6, 8, 10, 13, 16, and 20 nucleotides) in HEK293T cells. The MG151 candidates 80-85 ( Figure 1A - 1F ), 87-100 ( Figure 1G - 1T ), and 102-117 ( Figure 1U - 1JJ ) are shown by untreated samples, no RT, wild-type MMLV1, and wild-type MMLV2 as a control.

[0017] Figure 2Bar graph showing relative fold change in editing of untethered RT candidates from the MG151 family compared to wild-type MMLV editing (normalized to 1). Seven untethered MG151 candidates (candidates 98, 100, 99, 102, 103, 104, and 105) are shown, which have eight different PBS nucleotides of different lengths (PBS lengths of 2, 4, 6, 8, 10, 13, 16, and 20 nucleotides). The bars represent the specific PBS length tested for each candidate.

[0018] Figure 3A - 3W Bar graph showing the percentage of conversion editing of G to T for untethered reverse transcriptase (RT) candidates from the MG153 family, which have eight different PBS nucleotides of different lengths (PBS lengths of 2, 4, 6, 8, 10, 13, 16, and 20 nucleotides) in HEK293T cells. The MG153 candidates 1 - 5, 7 - 13, 15, 16, and 21 ( Figure 3A - 3O ) and 14, 17 - 20, and 25 - 27 ( Figure 3P - 3W ) are shown by untreated samples and wild-type MMLV1 as a control.

[0019] Figure 4A - 4G Bar graph showing the percentage of conversion editing of G to T for untethered reverse transcriptase (RT) candidates from the MG160 family, which have eight different PBS nucleotides of different lengths (PBS lengths of 2, 4, 6, 8, 10, 13, 16, and 20 nucleotides) in HEK293T cells. The MG160 candidates 1 - 6 and 8 ( Figure 4A - 4G ) are shown by untreated samples and wild-type MMLV1 as a control.

[0020] Figure 5A - 5G Bar graph showing the percentage of conversion editing of G to T for RT candidates from the MG160 family tethered to spCas9(H840A). The conversion of G to T for MG160 candidates 1 - 5 ( Figure 5A - 5G ) tethered to spCas9(H840A) was tested in HEK293T cells. The candidates are shown by untreated samples, wild-type MMLV1 as a control, wild-type MMLV2, spCas9(H840A)-MMLV1, and spCas9(H840A)-MMLV2.

[0021] Figure 6A - 6DShows the InDel percentage blot bar graph after targeting endogenous targets AAVS1 ( Figure 6A ), B2M ( Figure 6B ), CD5 ( Figure 6C ), and CD38 ( Figure 6D ) with the nuclease MG3-6 that binds to pegRNAs containing different lengths of PBS (PBS lengths of 2, 4, 6, 8, 10, 13, 16, and 20 nucleotides) in HEK293T cells.

[0022] Figure 7 Depicts a schematic diagram of an exemplary DNA construct for a GFP-based retrotransposition assay. The construct carries a cytomegalovirus promoter (CMVp), followed by a reverse transcriptase (RT-NLS) with an N-terminal tag (Flag-HA-NLS-MCP-linker). The EF1α promoter (EF1α) in the reverse orientation drives the expression of GFP (GFP exon 2 and GFP exon 1), only after the construct has successfully retrotransposed to the target site specified by the nuclease (inverted intron). After the primer binding site (PBS) binds to the 3' overhang generated by the nuclease, target-primed reverse transcription is initiated. NLS = nuclear localization signal; MCP = MS2 coat protein; GFP = green fluorescent protein; pA = polyA sequence, MS2 loop = site for MS2 coat protein binding.

[0023] Figure 8 Depicts the mechanism of targeting integration of reverse transcription-derived ssDNA by TnpA. The reverse transcription ncRNAs (msr shown in gray and msd shown in black) contain the expected cargo flanked by structural motifs recognized by TnpA (upper left, dashed box). The excised cargo (upper right) is circularized by TnpA and finds the targeting motif on the ssDNA target, which becomes available by binding to an RNA-guided effector (lower right, gray). TnpA mediates the integration of the ssDNA donor by cleaving the target, and the host repair mechanism repairs the integrated edit (lower left, dashed box).

[0024] Figure 9A - 9R Depicts the use of untethered MG151 candidates MG151-118 to MG151-135 for editing for the conversion of G to T across 8 different PBS lengths.

[0025] Figure 10A - 10D Depicts the use of untethered MG151 candidates MG151-123 to MG151-126 for editing for the conversion of G to T at PBS lengths of 6, 8, 10, and 13 nucleotides. Two biological replicates were performed for each candidate.

[0026] Figure 11A - 11D Depicts editing with untethered MG151 family mutants for G-to-T conversion. Figure 11A : MG151-98 wild type is shown as green bars, alongside point mutants of MG151-98, combinatorial mutants of MG151-98, and trimmed mutants of MG151-98. Figure 11A A single replicate is shown in Figure 11B and additional replicates with various MG151-98 mutations are found in Figure 11C : The MG151-99 mutant and wild-type MG151-99 have G-to-T conversion, with some mutations increasing wild-type activity. Figure 11D : The MG151-99 wild type is compared to a trimmed version of MG151-99. The trimmed 152AA of MG151-99 significantly increases the activity of G-to-T conversion, while trimming 136AA inhibits the editing activity. MMLV1 wild type is shown as gold bars, and MMLV2 (five mutant) serves as a control for each experiment.

[0027] Figure 12A - 12B Depicts untethered MG151 candidates (MG151-80 to MG151-135) tested for G-to-T conversion. Editing percentages ( Figure 12A ) and fold changes ( Figure 12B ) of G-to-T conversion relative to the MMLV wild type of PBS13. Each point represents a different PBS length, ranging from 2, 4, 6, 8, 10, 13, 16, and 20 nucleotides.

[0028] Figure 13A - 13H Depicts editing of untethered MG153 candidates tested for G-to-T conversion across 8 different PBS lengths. Figure 13H Shows MG153-53 editing when fused to Cas9.

[0029] Figure 14A - 14B Depicts untethered MG153 candidates tested for G-to-T conversion. Editing percentages ( Figure 14A ) and fold changes ( Figure 14A ) of G-to-T conversion relative to the MMLV wild type of PBS13. Each point represents a different PBS length, ranging from 2, 4, 6, 8, 10, 13, 16, and 20 nucleotides.

[0030] Figure 15A - 15UDepicts the editing using spCas9(H840A) tethered to the MG160 candidate for the G-to-T conversion across 8 different PBS lengths.

[0031] Figure 16A - 16B Depicts the MG160 candidates tethered to spCas9(H840A) that were tested for the G-to-T conversion. The percentage of editing ([ Figure 16A ) and fold change ([ Figure 16B ) of the G-to-T conversion relative to MMLV wild type of PBS13. Each point represents a different PBS length ranging from 2, 4, 6, 8, 10, 13, 16, and 20 nucleotides.

[0032] Figure 17A - 17D Depicts the untethered candidates MG151-98 and MG151-99. The RT candidates MG151-98 ([ Figure 17A ) and MG151-99 ([ Figure 17B ) were tested for 24nt insertion. The RT candidates MG151-98 ([ Figure 17C ) and MG151-99 ([ Figure 17D ) were tested for 15nt deletion.

[0033] Figure 18A - 18H Depicts the MG151 candidates including MG151-123 ([ Figure 18A and 18E ), MG151-124 ([ Figure 18B and 18F ), MG151-125 ([ Figure 18C and 18G ) and MG151-126 ([ Figure 18D and 18H ) that completed 24nt insertion ([ Figure 18A - 18D ) and 15nt deletion ([ Figure 18E - 18H ) at four PBS lengths.

[0034] Figure 19A - 19D Depicts the rational engineering of MG151-98. The MG151-98 wild type is shown as a green bar, alongside point mutations, combinatorial mutations, and trimming of MG151-98. The performance of these mutations for 24nt insertion ([ Figure 19A - 19B ) and 15nt deletion ([ Figure 19C - 19D ) is presented above together with controls MMLV1 and MMLV2.

[0035] Figure 20A - 20D Depicts the rational engineering of MG151-99. The MG151-99 wild type is shown as a green bar, alongside point mutations, combinatorial mutations, and trimming of MG151-99. The 24nt insertion ([Figure 20A - 20B ) and 15 nt deletion ( Figure 20C - 20D ) The performance of these mutations, along with control MMLV1 and MMLV2, is presented above.

[0036] Figure 21A - 21H Depicts MG153 candidates tested for 24 nt flag insertion across 4 to 8 different PBS lengths.

[0037] Figure 22A - 22H Depicts MG153 candidates tested for 15 nt deletion across 4 to 8 different PBS lengths.

[0038] Figure 23A - 23H Depicts the editing efficiency of spCas9(H840A) tethered to the MG160 candidate for 24 nt insertion at 4 to 8 different PBS lengths.

[0039] Figure 24A - 24H Depicts the editing efficiency of spCas9(H840A) tethered to the MG160 candidate for 15 nt insertion at 4 to 8 different PBS lengths.

[0040] Figure 25A - 25D Depicts the G-to-T transversion using RT in combination with the MG nickase MG3-6. Untethered ( Figure 25A - 25B ) and tethered ( Figure 25C - 25D ) systems were tested. The RTs tested included MG151-98, MG151-24, MG153-53, MG160-4, and MG151-99.

[0041] Figure 26A - 26C Depicts a screen for the ability of indicated control RTs and RT candidates to retrotranspose an RNA cargo containing GFP at a target specified by Cas9 in mammalian cells. Three days, six days, and eight days after transfection of cells with an RT-containing plasmid, a Cas9-containing plasmid, and a chemically synthesized guide RNA, the percentage of GFP-positive cells in the indicated samples was measured by flow cytometry to detect successful retrotransposition. LINE-WT (WT LINE-1RT), LINE-dead (D702Y LINE-1RT, RT-dead), NT (non-targeting guide), VEGFA (VEGFA-targeting guide).

[0042] Figure 27A - 27C Depicts the prime editing ability of engineered RTs. Figure 27A Depicts the prime editing percentage (y-axis) of MG160-4 RT across different PBS lengths (x-axis). Figure 27B Depicts the prime editing percentage (y-axis) of MG151-98 across different PBS lengths (x-axis). Figure 27CDepicts the prime editing percentage of MG153-3RT across different PBS lengths (x-axis) and (y-axis).

[0043] Figure 28 Depicts the ability of RT candidates to efficiently generate full-length cDNA from large RNA templates in mammalian cells.

[0044] Figure 29A - 29DD Depicts the editing percentage of MG160 family candidates tethered to spCas9(H840A). Candidates from the MG160 family were tethered to spCas9(H840A) and transfected into HEK293T cells to determine G-to-T editing on the VEGFA target. Chemically synthesized guides had a primer binding site length range of 2 - 20 nucleotides. Candidates MG160-473( Figure 29A ), MG160-283( Figure 29G ), MG160-379( Figure 29L ), MG160-395( Figure 29O ), MG160-9( Figure 29P ) and MG160-107( Figure 29CC ) had G-to-T editing levels (across multiple PBS lengths) equivalent to or better than that of spCas9(H840A) tethered to MMLV WT. spCas9-PE1 and spCas9-PE2 were transfected with chemically synthesized pegRNAs with a 13-nucleotide PBS length.

[0045] Figure 30A - 30C Depicts the editing percentages of G-to-T transversions, insertions, and deletions for selected MG160 candidates. In the case of MG160 candidates tethered to spCas9(H840A), RT was challenged to incorporate G-to-T transversions( Figure 30A ), 24-nucleotide insertions( Figure 30B ), and 15-nucleotide deletions( Figure 30C ) into the VEGFA target. MG160-107, MG160-473, MG160-283, MG160-379, and MG160-395 showed equivalent or improved editing levels for all types of editing at various PBS lengths compared to spCas9(H840A) tethered to MMLV WT. MG160-473 showed an editing level equivalent to that of spCas9(H840A) tethered to MMLV2 (hyperactive mutant). spCas9-PE1 and spCas9-PE2 were transfected with chemically synthesized pegRNAs with a 13-nucleotide PBS length.

[0046] Figure 31A - 31KDepicts the editing percentages of unique reverse transcriptase candidates from the MG retrotransposon family that are not tethered to spCas9(H840A). Candidates from various MG retrotransposon families were transfected in untethered form with the nickase spCas9(H840A) into HEK293T cells to determine G-to-T editing on the VEGFA target. The chemically synthesized guides had a primer binding site length range of 2 - 20 nucleotides. Candidates MG173-1( Figure 31J ) and MG173-2( Figure 31K ) were active and showed G-to-T editing above background levels across multiple PBS lengths. The controls MMLV1 and MMLV2 were untethered and transfected with spCas9(H840A) and a pegRNA with a chemically synthesized PBS length of 13 nucleotides.

[0047] Figure 32A - 32D Depicts the editing percentages of reverse transcriptase candidates from the MG group II intron family that are not tethered to spCas9(H840A). Candidates from various MG group II families were transfected in untethered form with the nickase spCas9(H840A) into HEK293T cells to determine G-to-T editing on the VEGFA target. The chemically synthesized guides had a primer binding site length range of 2 - 20 nucleotides. Candidate MG169-1( Figure 32D ) was slightly above the background editing level for G-to-T editing across multiple PBS lengths. Other MG candidates MG164-5( Figure 32A ), MG166-2( Figure 32B ) and MG167-4( Figure 32C ) did not show editing levels above background. The controls MMLV1 and MMLV2 were untethered and transfected with spCas9(H840A) and a pegRNA with a chemically synthesized PBS length of 13 nucleotides.

[0048] Figure 33A - 33D Depicts the editing percentages of WT MG160-4 and engineered mutants tethered to spCas9(H840A). Figure 33A Shows the editing percentages of seventeen engineered MG160-4 constructs tethered to spCas9(H840A) tested for G-to-T transversions on the VEGFA target in HEK293T cells. Chemically synthesized guides with PBS lengths ranging from 6 to 13 nucleotides were used to test the conversions. The point mutations H230K and H230R showed neutral changes in G-to-T editing activity, while combinations of multiple mutations substantially reduced the editing efficiency. Figure 33BShows the conversion of G to T with a selected point mutation, which shows an editing level similar to WT MG160-4. Then, 24-nucleotide insertions ( Figure 33C ) and 15-nucleotide deletions ( Figure 33D ) of MG160-4(H230K) and MG160-4(H230R) were tested. MG160-4(H230R) showed a slightly better editing level than MG160-4 WT and MG160-4(H230K) under various desired edits. spCas9-PE1 and spCas9-PE2 were transfected together with chemically synthesized pegRNAs with a PBS length of 13 nucleotides.

[0049] Figure 34 Depicts the editing percentages of WT MG153-53 and engineered mutants. The conversion of G to T on the VEGFA target of six untethered engineered MG153-53 constructs transfected with spCas9(H840A) was tested in HEK293T cells. Chemically synthesized guides with PBS lengths ranging from 6 to 13 nucleotides were used to test the conversion. The point mutation V200R showed an increase in G-to-T editing activity comparable to WT MG153-53, while combining multiple mutations significantly reduced the editing efficiency. The editing levels of MG153-53 WT and engineered constructs were comparable to or higher than those of the untethered controls TGIRT, marathon, and marathon mutants, but were significantly lower than those of the untethered MMLV WT (MMLV1) and MMLV hyperactive mutants (MMLV2).

[0050] Figure 35Depicts the editing percentage of MG3-6(H586A) with the selected RT candidates. The MG3-6(H586A) nickase was combined with the selected reverse transcriptases to perform the desired correction in the AAVS1 target. The reverse transcriptases were untethered (UT) and transfected with MG3-6(H586A) or tethered to the C-terminus of the nickase I or the N-terminus of the nickase (N) of MG3-6(H586A). The PBS lengths of the pegRNAs varied from 8, 10, 13, and 20 nucleotides. Background editing was shown at less than 0.1% editing. All selected RTs, except MG153-53, showed above-background editing by MG3-6(H586A). The activity of MG160-4 was increased when tethered to MG3-6(H586A) compared to the untethered condition. The MG151-98 engineered candidates slightly preferred to be untethered or tethered to the N-terminus of MG3-6(H586A). WT MMLV (MMLV1) and the hyperactive mutant MMLV (MMLV2) had the highest editing levels when tethered to the C-terminus of MG3-6(H586A). For each selected RT, each data point represents a single biological replicate at different PBS lengths.

[0051] Figure 36A - 36J Depicts the editing percentage of untethered MG71-2(H883A) with the selected RT candidates on the AAVS1 target. Figure 36A Shows the biological triplicate data of the selected RT candidates for five nucleotide changes on the AAVS1 target with untethered MG71-2n and chemically synthesized pegRNAs with PBS lengths of 4, 6, 8, 10, 13, and 16 nucleotides. The selected RT candidates were then tested for five nucleotide changes ( Figure 36B ), five nucleotide changes in the modified scaffold of the pegRNA ( Figure 36C ), G-to-T transversions ( Figure 36D ), 24-nucleotide insertions ( Figure 36E ), and 15-nucleotide deletions ( Figure 36F ). Across PBS lengths of 4, 6, 8, 10, 13, and 16 nucleotides, the G-to-T changes ( Figure 36G ), 15-nucleotide deletions ( Figure 36H ), 24-nucleotide insertions ( Figure 36I ), and five nucleotide changes ( Figure 36J ) of additional engineered MG151-98 candidates MG151-98(166AA), MG151-98(166AA, H171N), and MG151-98(166AA, K297P) were tested. Editing levels above background were defined as greater than 0.1%. Except for the biological triplicate data representedFigure 36A Except as otherwise noted, each figure represents a single biological replicate.

[0052] Figure 37A - 37C Depicts the editing percentages of G-to-T transversions, insertions, and deletions engineered in the MG151-98 mutant. In cases where the engineered MG151-98 candidates were not tethered to spCas9(H840A), RT was challenged to incorporate G-to-T transversions ( Figure 37A ), 24-nucleotide insertions ( Figure 37B ), and 15-nucleotide deletions ( Figure 37C ) into the VEGFA target. MG151-98(Δ166AA) enhanced the editing levels for most PBS lengths under all conditions. Specifically, when the trimmed MG151-98 construct was combined with the point mutations H171N or K297P, the editing levels were further increased, reaching levels comparable to or better than that of MMLV wild type. spCas9(H840A) not tethered to MMLV WT and MMLV2 was transfected together with chemically synthesized pegRNA with a PBS length of 13 nucleotides.

[0053] Figure 38 Depicts an overview of the mechanism for programmable genome editing achieved with Cas9, retron reverse transcriptase, and the ssDNA transposase TnpA.

[0054] Figure 39A , 39B and 39C depict an overview of the design principles for the engineered ncRNA used to generate Ec96. Figure 39A Depicts an overview of three insertion sequences of three different lengths flanked by the Hp TnpA LE / RE recognition motif. Figure 39B Depicts a figure from the specified paper (Wang et al., *Nature Microbiology* (2022)) that indicates the msdDNA region unresolved in the cryo-EM structure of Ec86 complexed with its product. Figure 39C Depicts three different alternative regions of the msd stem-loop identified for Ec86 ncRNA.

[0055] Figure 40 Depicts the predicted secondary structure of the engineered Ec86 ncRNA that contains a 200-nt insertion or a 500-nt partial kanamycin gene flanked by the reverse complement (rc) LE / RE motif of Hp TnpA. The motifs, msr, and inverted repeats (IR) required for initiating reverse transcription are highlighted.

[0056] Figure 41Depicts the quantification of msdDNA production by qPCR in reactions with or without Ec86 reverse transcriptase. WT is wild-type ncRNA. LE40RE_v1 to v3, LE200RE_v1 and v3, and LE500RE v1 to v3 are engineered ncRNA designs.

[0057] Figure 42 Depicts the insertion of chimeric products generated by the TnpA / reverse transcription system as confirmed by PCR. The PCR products are indicated by arrows. Lane numbers correspond to the following: Lane 1: LE200RE_v1 ncRNA, +RT, +TnpA; Lane 2: LE200RE_v1 ncRNA, -RT, +TnpA; Lane 3: LE200RE_v3 ncRNA, +RT, +TnpA; Lane 4: LE200RE_v3 ncRNA, -RT, +TnpA; Lane 5: LE500RE_v1 ncRNA, +RT, +TnpA; Lane 6: LE500RE_v1 ncRNA, -RT, +TnpA; Lane 7: LE500RE_v2 ncRNA, +RT, +TnpA; Lane 8: LE500RE_v2 ncRNA, -RT, +TnpA; Lane 9: LE500RE_v3 ncRNA, +RT, +TnpA; Lane 10: LE500RE_v3 ncRNA, -RT, +TnpA; Lane 11: LE200RE_v1 ncRNA, +RT, -TnpA; Lane 12: LE200RE_v1 ncRNA, -RT, -TnpA; Lane 13: LE200RE_v3 ncRNA, +RT, -TnpA; Lane 14: LE200RE_v3 ncRNA, -RT, -TnpA; Lane 15: LE500RE_v1 ncRNA, +RT, -TnpA; Lane 16: LE500RE_v1 ncRNA, -RT, -TnpA; Lane 17: LE500RE_v2 ncRNA, +RT, -TnpA; Lane 18: LE500RE_v2 ncRNA, -RT, -TnpA; Lane 19: LE500RE_v3 ncRNA, +RT, -TnpA; Lane 20: LE500RE_v3 ncRNA, -RT, -TnpA.

[0058] Figure 43Depicts the Sanger sequencing results of the ssDNA insertion products generated by TnpA, where the substrate of TnpA is generated by the Ec86 retrotransposon. The regions highlighted in the Sanger sequencing chromatogram show the junctions of the chimeric products, where the 5' sequence corresponds to the right end (RE) motif of Hp TnpA integrated with the cargo, and the 3' sequence corresponds to the ssDNA target provided in the reaction mixture. Figure 43 SEQ ID NO 2578 and 2578 are disclosed in the order of appearance respectively.

[0059] Figure 44 Depicts a method for confirming the ncRNA prediction and msd insertion tolerance of retrotransposons.

[0060] Figure 45 Depicts the secondary structure prediction of retrotrans ncRNA from the MG154 family, highlighting the 5' and 3' inverted repeat elements (IR), msr, and msd stem-loops required for initiating reverse transcription. The regions of the msd stem-loop replaced by engineered sequences are indicated.

[0061] Figure 46 Depicts the secondary structure prediction of retrotrans ncRNA from the MG155 family, highlighting the 5' and 3' inverted repeat elements (IR), msr, and msd stem-loops required for initiating reverse transcription. The regions of the msd stem-loop replaced by engineered sequences are indicated.

[0062] Figure 47 Depicts the secondary structure prediction of retrotrans ncRNA from the MG156 family, highlighting the 5' and 3' inverted repeat elements (IR), msr, and msd stem-loops required for initiating reverse transcription. The regions of the msd stem-loop replaced by engineered sequences are indicated.

[0063] Figure 48 Depicts the secondary structure prediction of retrotrans ncRNA from the MG157 family, highlighting the 5' and 3' inverted repeat elements (IR), msr, and msd stem-loops required for initiating reverse transcription. The regions of the msd stem-loop replaced by engineered sequences are indicated.

[0064] Figure 49 Depicts the secondary structure prediction of retrotrans ncRNA from the MG158 family, highlighting the 5' and 3' inverted repeat elements (IR), msr, and msd stem-loops required for initiating reverse transcription. The regions of the msd stem-loop replaced by engineered sequences are indicated.

[0065] Figure 50Depicts the predicted secondary structure of a retro ncRNA from the MG159 family, highlighting the 5' and 3' inverted repeat elements (IRs) and the msr and msd stem-loops required for reverse transcription initiation. The region of the msd stem-loop replaced by the engineered sequence is indicated.

[0066] Figure 51 Depicts the predicted secondary structure of a retro ncRNA from the MG173 family, highlighting the 5' and 3' inverted repeat elements (IRs) and the msr and msd stem-loops required for reverse transcription initiation. The region of the msd stem-loop replaced by the engineered sequence is indicated.

[0067] Figure 52 Depicts the detection of msdDNA production by qPCR. Ec86 is the positive control retrotransposon RT, and the corresponding ncRNAs tested contain a ~200 nt insertion sequence at the alternative position version 1 described previously. The ncRNAs whose activities were identified using the corresponding retrotransposon RT are colored black (msdDNA production > 10-fold above the no-RT control). The ncRNAs whose activities were not identified using the corresponding retrotransposon RT are colored light gray.

[0068] Figure 53A - 53D Depicts the percentage of editing of 5 nt changes on the AAVS1 target using MG RT and MG71-2 (H883A). The RT was tested in untethered or tethered forms (the RT on the C-terminus of MG71-2 (H883A) is indicated with nickase-RT, and the RT on the N-terminus of MG71-2 (H883A) is indicated with RT-nickase). Figure 53A : MMLV2-RT was tested untethered and tethered to MG71-2 (H883A), with the highest untethered editing levels at PBS13, nickase-RT PBS16, and RT-nickase PBS13. Figure 53B : The engineered MG151-98 (K297P, Δ166AA) was tested untethered and tethered to MG71-2 (H883A), with the highest untethered editing levels of nickase-RT and RT-nickase at PBS13, and the highest editing level was observed in the RT-nickase configuration. Figure 53C : MG160-4 (H230R) was only tested in tethered form, with the highest editing levels of nickase-RT at PBS10 and RT-nickase at PBS13. The highest editing level was observed for the RT-nickase configuration. Figure 53D: MG160-473 was tested in a tethered format, where the highest editing levels were observed in the RT-nicking enzyme configuration of PBS13. The nicking enzyme-RT configuration of MG160-473 had low read counts by NGS treatment, and the editing percentage was not determined. Correct editing indicates the expected correction with no errors found in the NGS amplicons. Incorrect editing refers to the incorporation of the expected edit but also includes errors in the NGS amplicons of the pegRNA and scaffold incorporation.

[0069] Figure 54A - 54S Depicts the editing percentages of G-to-T transversions for MG retrotransposon family candidates not tethered to spCas9(H840A). Figure 54A Summarizes the editing percentages of G-to-T transversions for untethered MG retrotransposon candidates from the MG173 family and the MG192 family across eight different PBS lengths of 2, 4, 6, 8, 10, 13, 16, and 20 nucleotides. The MG173-8 candidate showed the highest editing level compared to the other nine retrotransposon candidates. Figure 54B - 54J The editing level labeled "correct editing" shown in [Figure] represents the expected editing with no errors in the NGS amplicons, and the bars labeled "incorrect editing" refer to the incorporation of the expected edit and include errors within the NGS amplicons or errors in the incorporation from RT and / or the scaffold incorporation of the pegRNA. Figure 54K - 54S The editing levels shown in [Figure] display the editing levels across eight different PBS lengths, where the bars labeled "editing" represent the expected editing with errors in the NGS amplicons (excluding pegRNA scaffold incorporation), and the bars labeled "scaffold incorporation" represent the expected editing and scaffold incorporation of the pegRNA.

[0070] Figure 55A - 55AS Depicts the editing percentages of G-to-T transversions for MG160 family candidates tethered to spCas9(H840A). Figure 55A Summarizes the editing percentages of G-to-T transversions for tethered MG160 candidates across eight different PBS lengths of 2, 4, 6, 8, 10, 13, 16, and 20 nucleotides. MG160-45, MG160-121, MG160-136, MG160-193, MG160-232, and MG160-358 showed editing levels of 5% or higher at different PBS lengths. Figure 55B - 55WThe editing levels shown in [figure] represent the editing levels across eight different PBS lengths. The bars labeled "correct editing" represent the expected editing without errors in the NGS amplicons, and the bars labeled "incorrect editing" refer to the incorporation of the expected editing and include errors within the NGS amplicons, or errors in incorporation from RT and / or incorporation of the pegRNA scaffold. Figure 55X - 55AS The editing levels shown in [figure] display the editing levels across eight different PBSs, where the bars labeled "editing" represent the expected editing with errors in the NGS amplicons (excluding pegRNA scaffold incorporation), and the bars labeled "scaffold incorporation" represent the expected editing and scaffold incorporation of the pegRNA.

[0071] Figure 56A - 56F Depicted are the editing percentages of different edits for the VEGFA target where the MG151-98 mutant is not tethered to spCas9(H840A). The MG151-98 wild type and mutants MG151-98(D166AA, H171N) and MG151-98(D166AA, K297P) were evaluated to correct for the G-to-T transversion ( Figure 56A and 56D ) of the VEGFA target for pegRNAs with different PBS lengths of 6, 8, 10, and 13 nucleotides, Figure 56B and 56E ) 24-nucleotide insertions ( Figure 56C and 56F ) and 15-nucleotide deletions ( Figure 56A - 56C : The bars labeled "correct editing" represent the expected editing without errors in the NGS amplicons, and the bars labeled "incorrect editing" refer to the incorporation of the expected editing and include errors within the NGS amplicons, or errors in incorporation from RT and / or incorporation of the pegRNA scaffold. Controls MMLV1 and MMLV2 represent untethered spCas9(H840A), the pegRNA of PBS13, and the RT plasmids encoding MMLV1 or MMLV2, respectively. Figure 56D - 56F The editing levels shown in [figure] display the editing levels across four different PBS lengths, where the bars labeled "editing" represent the expected editing with errors in the NGS amplicons (excluding pegRNA scaffold incorporation), and the bars labeled "scaffold incorporation" represent the expected editing and scaffold incorporation of the pegRNA. Controls MMLV1 and MMLV2 represent untethered spCas9(H840A), the pegRNA of PBS13, and the RT plasmids encoding MMLV1 or MMLV2, respectively.

[0072] Figure 57A - 57HDescribes the editing percentages of G-to-T transversions of MG151 family mutants and MG153 family mutants not tethered to spCas9(H840A). For MG151-123 wild type and mutants (M304R, H287F, H178R, H178N, G279R or G279N)( Figure 57A and 57E ), MG151-126 wild type and mutants (H287F, G179R, G179N, A280R, A280K or A276R)( Figure 57B and 57F ), MG153-18 wild type and mutants (G119R, P242R or double mutant G119R and P242R)( Figure 57C and 57G ) and MG153-20 wild type and mutants (N55R, P226R or double mutant N55R and P226R)( Figure 57D and 57H ), the editing percentages of G-to-T transversions of pegRNAs with different PBS lengths of 6, 8, 10 and 13 nucleotides on the VEGFA target were evaluated. Figure 57A - 57D : Bars labeled "correct editing" represent the expected editing without errors in the NGS amplicon, and bars labeled "incorrect editing" refer to the incorporation of the expected editing and include errors within the NGS amplicon, or errors due to incorrect incorporation of RT and / or incorporation of the pegRNA scaffold. Figure 57E - 57H The editing levels shown in represent the percentage of editing levels across four different PBS lengths, where bars labeled "editing" represent the expected editing with errors in the NGS amplicon (excluding pegRNA scaffold incorporation), and bars labeled "scaffold incorporation" represent the expected editing and scaffold incorporation of the pegRNA. Controls including "no RT" represent untethered spCas9(H840A) and pegRNA of PBS13, and MMLV1 and MMLV2 represent untethered spCas9(H840A), pegRNA of PBS13 and RT plasmids encoding MMLV1 or MMLV2 respectively.

[0073] Figure 58A - 58L Depicts the editing percentages of different edits of MG160-473 mutants tethered to spCas9(H840A) for the VEGFA target. The MG160-473 wild type and mutants MG160-473(F231K) and MG160-473(F231R) were evaluated to correct G-to-T transversions on the VEGFA target of pegRNAs with different PBS lengths of 6, 8, 10, 13 and 16 nucleotides( Figure 58A 、58D , 58G and 58J), 24 nucleotide insertions ( Figure 58B , 58E , 58H and 58K) and 15 nucleotide deletions ( Figure 58C , 58F , 58I and 58L). Figure 58A -C and 58G-I: Bars labeled "correct editing" represent the expected editing without errors in the NGS amplicons, and bars labeled "incorrect editing" refer to the incorporation of the expected editing and include errors within the NGS amplicons or errors incorporated from RT and / or incorporation of the pegRNA scaffold. Figure 58D - 58F The editing levels shown in 58J-L represent the percentage of editing levels across different PBS lengths, where bars labeled "editing" represent the expected editing with errors in the NGS amplicons (excluding pegRNA scaffold incorporation), and bars labeled "scaffold incorporation" represent the expected editing and scaffold incorporation of the pegRNA. Controls including "untreated" represent cells not treated during transfection, and cas-PE1 and cas-PE2 represent tethering spCas9(H840A) to MMLV1 or MMLV2 using pegRNAs with PBS13. Asterisks indicate that the reads of the NGS samples are less than 1000.

[0074] Figure 59A - 59P Depicts the editing percentages of five nucleotide changes on the AAVS1 target with tethered MG reverse transcriptase and MG71-2n. The reverse transcriptase was tested for targeting five nucleotide changes on the AAVS1 target across six different PBS lengths (6, 8, 10, 13, 16, or 20 nucleotides) without tethering to MG71-2n, tethered to the C-terminus of MG71-2n (nickase-RT), or tethered to the N-terminus of MG71-2n (RT-nickase). The reverse transcriptases tested for this correction include: MMLV1 ( Figure 59A and 59D ), MMLV2 ( Figure 59B and 59E ), MG160-4 ( Figure 59C and 59F ), MG151-98 (D166AA) ( Figure 59G and 59J ), MG151-98 (D166AA, H171N) ( Figure 59H and 59K ), MG151-98 (D166AA, K297P) ( Figure 59I and 59L ), MG160-4 (H230R) ( Figure 59M and 59O) and MG160-473( Figure 59N and 59P ). The bars labeled "correct editing" represent the expected editing without errors in the NGS amplicons, and the bars labeled "incorrect editing" refer to the incorporation of the expected editing and include errors within the NGS amplicons, or errors resulting from the incorrect incorporation of RT and / or the incorporation of the pegRNA scaffold. The bars labeled "editing" represent the expected editing with errors in the NGS amplicons (excluding the incorporation of the pegRNA scaffold), and the bars labeled "scaffold incorporation" represent the expected editing and the incorporation of the pegRNA scaffold. Low read counts indicate that the NGS samples have fewer than 1000 reads.

[0075] Figure 60A - 60H Depicts the editing percentages of different edits on the AAVS1 target in the case where the MG reverse transcriptase is tethered to the N-terminus of MG71-2n. The reverse transcriptases MMLV1, MMLV2, MG160-4 wild type, or MG160-4(H230R) are tethered to the N-terminus of MG71-2n via a 32-amino acid linker and challenged with G-to-T transversions( Figure 60A and 60E ), 24-nucleotide insertions( Figure 60B and 60F ), 15-nucleotide deletions( Figure 60C and 60G ), or five-nucleotide changes( Figure 60D and 60H ) on the AAVS1 target using pegRNAs with PBS lengths of 8, 10, 13, and 16 nucleotides. The MG160 candidates are also tested with a pegRNA with a PBS length of 13 nucleotides without being tethered (UT) to MG71-2n. Figure 60A - 60D : The bars labeled "correct editing" represent the expected editing without errors in the NGS amplicons, and the bars labeled "incorrect editing" refer to the incorporation of the expected editing and include errors within the NGS amplicons, or errors resulting from the incorrect incorporation of RT and / or the incorporation of the pegRNA scaffold. Figure 60E - 60H The editing levels shown in represent the percentage of editing levels across four different PBS lengths, where the bars labeled "editing" represent the expected editing with errors in the NGS amplicons (excluding the incorporation of the pegRNA scaffold), and the bars labeled "scaffold incorporation" represent the expected editing and the incorporation of the pegRNA scaffold. Asterisks indicate that the NGS samples have fewer than 1000 reads.

[0076] Figure 61A - 61LDescribes the percentage editing of different edits on the AAVS1 target using the MG151-98 mutant not tethered to MG71-2n. The reverse transcriptases MMLV1, MMLV2, MG151-98 (D166AA, H171N), MG151-98 (D166AA, K297P), MG151-98 (D166AA, H171N, K297P), and untethered MG71-2n were challenged on the AAVS1 target with pegRNAs having PBS lengths of 8, 10, 13, and 16 nucleotides for G-to-T transversions ( Figure 61A and 61E ), 24-nucleotide insertions ( Figure 61B and 61F ), 15-nucleotide deletions ( Figure 61C and 61G ), or five-nucleotide changes ( Figure 61D and 61H ). Figure 61A - 61D : Bars labeled "correct edit" represent the expected edit with no errors in the NGS amplicon, and bars labeled "incorrect edit" refer to the incorporation of the expected edit and include errors within the NGS amplicon or errors arising from RT incorporation and / or pegRNA scaffold incorporation. Figure 61E - 61H The editing levels represented in Figure 61I - 61K show the percentage of editing levels across four different PBS lengths, where bars labeled "edit" represent the expected edit with errors in the NGS amplicon (excluding pegRNA scaffold incorporation), and bars labeled "scaffold incorporation" represent the expected edit of the pegRNA and scaffold incorporation.

[0077] Figure 62A - 62B Depicts the modification of the MG71-2 scaffold to increase the percentage of five-nucleotide change editing on the AAVS1 target. The scaffold of MG71-2 contains 107 nucleotides, and two modified forms of the scaffold, D2 or D2C2, result in shortened scaffold lengths of 85 nucleotides and 79 nucleotides, respectively. The D2 scaffold removes the last hairpin of the MG71-2 scaffold, and the D2C2 scaffold removes the last hairpin and the small bulge of the MG71-2 scaffold. The editing levels of five-nucleotide changes on the AAVS1 target were tested on wild-type and modified scaffolds with the reverse transcriptases MMLV2 or MG160-4 (H230R) tethered to the N-terminus of MG71-2n across PBS lengths of 8, 10, 13, and 16 nucleotides. Figure 62A:Bars labeled "correct editing" represent the expected editing without errors in the NGS amplicons, and bars labeled "incorrect editing" refer to the incorporation of the expected editing and include errors within the NGS amplicons, or errors incorporated from RT and / or incorporation of the pegRNA scaffold. Figure 62B The editing levels shown in Figure 62B represent the percentage of editing levels across eight different PBS lengths, where bars labeled "editing" represent the expected editing with errors in the NGS amplicons (excluding pegRNA scaffold incorporation), and bars labeled "scaffold incorporation" represent the expected editing and scaffold incorporation of the pegRNA.

[0078] Figure 63A - 63H Depicts the optimization of the guide RNA to improve the editing level of MG71-2n. Figures 63A - 63D Shows that the reverse transcriptases MMLV1, MMLV2, MG151-98 (D166AA, H171N), MG151-98 (D166AA, K297P), MG151-98 (D166AA, H171N, K297P), and untethered MG71-2n were challenged with five nucleotide changes on the AAVS1 target. Figures 63E - 63H Shows that the reverse transcriptases MMLV1, MMLV2, MG160-4, and MG160-4 (H230R) tethered to the N-terminus of MG71-2n and MG160-4 and MG160-4 (H230R) untethered (UT) to MG71-2n were challenged with five nucleotide changes on the AAVS1 target. Different mismatches in the pegRNA across the PBS region were tested to determine if improvements in editing could be achieved. Figure 63A 、 63C 、The PBS lengths of 8, 10, 13, and 16 nucleotides in 63E and 63G have perfect complementarity to the target region. In Figure 63B 、 63D, among 63F and 63H, PBSs with lengths of 10, 13, 16, and 20 nucleotides have perfect complementarity of 8 nucleotides in the region adjacent to the reverse transcription template (RTT), and then have different mismatches (mm) to obtain PBS lengths of 10 (2 mismatches), 13 (5 mismatches), 16 (8 mismatches), and 20 (12 mismatches) nucleotides. The bars labeled "correct editing" represent the expected editing without errors in the NGS amplicons, and the bars labeled "incorrect editing" refer to the incorporation of the expected editing and include errors within the NGS amplicons, or errors incorporated from RT and / or incorporation of the pegRNA scaffold. The bars labeled "editing" represent the expected editing with errors in the NGS amplicons (excluding pegRNA scaffold incorporation), and the bars labeled "scaffold incorporation" represent the expected editing and scaffold incorporation of the pegRNA.

[0079] Figures 64A - 64E Depicts the modification of the guide RNA of MG3-6 to improve the editing level in mammalian cells. Figure 64A : The MG3-6 wild-type mRNA was used to determine the percentage of the modification (including SNPs and InDels) level of the target amplicon AAVS1 in the NGS samples. The guide RNA consists of the scaffold and spacer of the target, and the pegRNA includes the guide RNA with PBS and RTT sequences. The modifications modL1-modL4 increased the regions of the GC content in hairpins 1-3 of the scaffold (modL1-modL3), where modL4 combines the modifications of all hairpins in the scaffold. Figures 64B - 64C Depicts the editing percentages of two nucleotide changes in the AAVS1 target measured across PBS lengths of 10 and 13 nucleotides using MMLV2 tethered to the C-terminus of MG3-6 (H586A) with wild-type and modified scaffolds modL1-modL4. As a control, "untreated" represents the cells untreated during transfection, and MG3-6 (H586A) represents the nickase and pegRNA without reverse transcriptase in cell transfection. Figures 64D - 64EDescribes the editing percentages of two nucleotide changes in the AAVS1 target measured across PBS lengths of 8, 10, 13, and 16 nucleotides with perfect complementarity to the target or PBS lengths of 10 (2 mismatches), 13 (5 mismatches), 16 (8 mismatches), and 20 (12 mismatches) using untethered MMLV1, MMLV2, MG151-98 (D166AA, H171N), and MG151-98 (D166AA, K297P), and the nickase MG3-6 (H586A). Bars labeled "correct editing" represent the expected editing without errors in the NGS amplicon, and bars labeled "incorrect editing" refer to the incorporation of the expected editing and include errors within the NGS amplicon, or errors incorporated from RT and / or incorporation of the pegRNA scaffold. Bars labeled "editing" represent the expected editing with errors in the NGS amplicon (excluding pegRNA scaffold incorporation), and bars labeled "scaffold incorporation" represent the expected editing and scaffold incorporation of the pegRNA.

[0080] Figures 65A - 65B Depicts a comparison of target recognition of guide RNAs with different PBS lengths by MG3-6 and MG3-6 / 3-8. MG3-6 wild-type and MG3-6 / 3-8 mRNA were used to determine the percentage of modified (including SNPs and InDels) levels of target amplicons AAVS1 ( Figure 65A ) and B2M ( Figure 65B ) of guide RNAs or pegRNAs with PBS lengths of 2, 4, 6, 8, 10, 13, 16, and 20 nucleotides. The guide RNA consists of the scaffold and spacer of the target, and the pegRNA includes the guide RNA with PBS and RTT sequences. MG3-6 / 3-8 showed a higher level of modification (including InDels) on the target compared to MG3-6. The control "untreated" represents cells that were not treated during transfection.

[0081] Figures 66A - 66D Depicts the identification of the MG14-241 target compatible with the prime editing system. Figure 66A : Wild-type MG14-241 mRNA or plasmid was used to determine the percentage of modified (including SNPs and InDels) levels of various targets. Guide RNAs for different targets (G1, H1, B2, E2, F2, and G2) led to different percentages of modification with target E2 (region of AAVS1), resulting in the highest level of InDels (up to approximately 60%). Figure 66B: The mRNA of MG14-241 is used to determine the percentage of modification (including SNPs and InDels) levels of the target amplicon AAVS1 in NGS samples. The guide RNA consists of a scaffold and a spacer of the target, and the pegRNA includes the guide RNA with PBS and RTT sequences. As the length of PBS increases, the percentage of modification decreases. The control "untreated" indicates cells that were not treated during transfection. Figures 66C - 66D : Editing percentages of five nucleotide changes on the AAVS1 target with untethered reverse transcriptases MMLV1, MMLV2, MG151-98 (D166AA, H171N), and MG151-98 (D166AA, K297P) and the nickase MG14-241n across eight different PBS lengths (2, 4, 6, 8, 10, 13, 16, and 20 nucleotides). MG14-241n (without RT) indicates the nickase and pegRNA without reverse transcriptase in cell transfection. The bars labeled "correct editing" represent the expected editing without errors in the NGS amplicon, and the bars labeled "incorrect editing" refer to the incorporation of the expected editing and include errors within the NGS amplicon, or errors incorporated from RT and / or incorporation of the pegRNA scaffold. The bars labeled "editing" represent the expected editing with errors in the NGS amplicon (excluding pegRNA scaffold incorporation), and the bars labeled "scaffold incorporation" represent the expected editing and scaffold incorporation of the pegRNA.

[0082] Figures 67A - 67D Depicts the design of engineered cell lines, RT-Cas chimeric proteins, and RNA cargo templates for evaluating integration by TPRT. Figure 67A Depicts a schematic showing the integration of an artificial sequence into HEK293 cells by lentivirus to generate an engineered cell line with a target site for integration. Figure 67B Depicts the percentage of indels generated by five different sgRNAs targeting an engineered landing pad. Figure 67C Depicts a schematic showing four different conformations of each RT-Cas9WT / nickase fusion generated for testing. Figure 67D Depicts six cargo designs generated for testing integration by TPRT.

[0083] Figure 68 Depicts a schematic illustration of primers for left and right end PCR for detecting integration.

[0084] Figures 69A - 69C Depicts the use of Cas9 WT-MG140-3 and sg4 at the LE using Tapestation (the box shows the band of interest; Figure 69A) Sanger sequencing was used at LE PCR (sequences matching the landing pad and cargo are shown); Figure 69B ) and Sanger sequencing was used at RE PCR to detect cargo integration (sequences matching the acquisition are shown, but the insertion of another product (Cas9) is also shown; Figure 69C ).

[0085] Figures 70A - 70B Detection of cargo integration using MG140-3-Cas9 WT and sg4 is depicted. Tapestation at LE ( Figure 70A ) and Sanger sequencing at LE PCR ( Figure 70B ) showed matches to the landing pad and mCherry cargo.

[0086] Figure 71 Detection of cargo integration using Cas9 WT-MG140-8 and sg4 by Sanger sequencing at LE is depicted.

[0087] Figures 72A - 72B Detection of cargo integration using MG153-18-CAs9 WT and sg4 by Tapestation at LE ( Figure 72A ) and Sanger sequencing at LE ( Figure 72B ) is depicted.

[0088] Figures 73A - 73C Reverse transcriptase RT activity on a homologous ncRNA loaded with a 2.2 kb cargo is depicted. Figure 73A A schematic of the substrate design for testing the activity and continuous synthesis ability of the retrotransposon RT is described. A universal template is used to test the nonspecific activity of the retrotransposon and is primed by an ssDNA primer oligonucleotide annealing to the 3'-end of the RNA. The presence of the terminal 5' and 3' reverse transcription ncRNA elements promotes the reverse transcription ncRNA to be primed by the 5' and 3' inverted repeats (IR). For both substrates, the cargo sequence is flanked by the reverse complement (rc) of the LE and RE recognition motifs of MG92-4 TnpA. This sequence is flanked by approximately 100 nt of RNA sequence, which when converted to cDNA can be quantified by multiplex TaqMan qPCR to evaluate how many cDNA molecules were synthesized at the 5' (FAM) end and 3' (HEX) end by RT. For the retrotransposon ncRNA substrate, the sequence is inserted into the alternative region of the previously identified ncRNA msd. Figure 73BDepicts the amount of ssDNA detected by FAM and HEX via multiplex TaqMan qPCR. An RT-negative control was generated by not adding any RT expression template to the cell-free expression system. The dotted line is 10 times the highest background RT-negative signal. TGIRT is a GII intron control RT, MMLV is a retroviral control RT, and Ec86 is a retroviral control RT. The label "gen" indicates testing the RT with a universal template, while "ncRNA" indicates testing the RT with a cargo-loaded cognate ncRNA. Figure 73C Depicts the confirmation of 2.2 kb ssDNA generated by RT by the tapestation D5000. Lanes correspond to the following: Lane 1: ladder; Lane 2: RT-negative gen; Lane 3: TGIRT gen; Lane 4: MG154-1 nRNA; Lane 5: MG157-1 ncRNA; Lane 6: MG157-3ncRNA; Lane 7: MG157-4 ncRNA; Lane 8: MG158-1 ncRNA; Lane 9: MG159-3ncRNA; Lane 10: MG173-1 ncRNA.

[0089] Figures 74A - 74B Depicts the screening of the ability of the retrotransposon RT MG173-1 to synthesize cDNA in mammalian cells. Figure 74A Depicts a sketch describing a method for detecting cDNA synthesis in mammalian cells. The first (FAM) and last (HEX) 100 bp of a 4.1 kb RNA template were detected using Taqman-based qPCR. Figure 74B Describes the Taqman qPCR detection of the first (FAM probe) and last (HEX probe) 100 bp PCR products amplified from cDNA synthesized from a universal 4 kb template, a universal 2 kb template, and a MG173-1-specific template flanked by 5' and 3' terminal MG173-1 ncRNA elements.

[0090] Figures 75A - 75B Depicts the insertion reaction and Sanger sequencing of the PCR of TnpA 92-4 with cDNA cargo generated by a 2.2 kb retrotransposon. Figure 75A : Lane 1: PCR of the no-template control (NTC) insertion reaction with an ssDNA superpolymer target and MG173-1 generated cDNA cargo. Lane 2: PCR of the TnpA 92-4 insertion reaction with an ssDNA superpolymer target and MG173-1 generated cDNA cargo. Figure 75B : Sanger sequencing of the chimeric insertion product generated by the insertion of the cargo generated by MG173-1 mediated by TnpA 92-4 into an ssDNA superpolymer target. Figure 75BSEQ ID NO:2579 is disclosed.

[0091] Figures 76A - 76H Depicts the therapeutic target site targeted by MG71-2. Figure 76A : The WT mRNA of MG71-2 has InDels at the therapeutic relevant sites (hPDK1, G6PC1 Q347*, and PAH R408W) with various guide RNAs. The highest InDels occur in Guide 1 of the hPDK1 gene and Guide 2 of the PAH gene targeting the R408W mutation. No InDels were detected for other guides tested against G6PC1. The positive control contains guide RNA targeting AAVS1. Figure 76B : Guide RNAs and pegRNAs with different PBS lengths of 8, 10, and 13 nucleotides were used to target the HBB gene mutation E7V. Compared with the guide RNA, the InDels decreased slightly through the pegRNA. The editing levels across eight different PBS lengths are shown. Figures 76C - 76H : Then prime editing experiments were carried out with pegRNAs using the spacers from Figures 76A - 76B The prime editing system is MG160-4(H230R) tethered to the N-terminus of MG71-2n (MG160-4(H230R)-MG71-2n) and MMLV2 tethered to the N-terminus of MG71-2n (MMLV2-MG71-2n). Figures 76C - 76D : MG160-4(H230R)-MG71-2n and MMLV2-MG71-2n target and disrupt the microRNA recognition site using pegRNAs containing 3 or 5 nucleotide (nt) mismatches incorporated into the RT template (RTT) of the pegRNA. For the 3nt mismatch incorporated into the hPDK1 microRNA recognition site, the highest editing level was observed at PBS10. Figures 76E - 76F : Prime editing systems targeting PAH R408W with PBS lengths of 8, 10, and 13nt, where the RTT length varies between 29nt and 32nt, did not show detectable editing levels. Figures 76G - 76H: MG160-4(H230R)-MG71-2n and MMLV2-MG71-2n target the HBB E7V mutation across multiple PBS lengths and achieve editing levels above background. Bars labeled "correct editing" represent the expected editing without errors in the NGS amplicon, and bars labeled "incorrect editing" refer to the incorporation of the expected editing and include errors within the NGS amplicon or errors resulting from RT incorporation and / or pegRNA scaffold incorporation. Bars labeled "editing" represent the expected editing with errors in the NGS amplicon (excluding pegRNA scaffold incorporation), and bars labeled "scaffold incorporation" represent the expected editing and scaffold incorporation of the pegRNA. *Indicates that fewer than 1000 reads were obtained for the NGS sample, and error bars represent the standard deviation of two biological replicates.

[0092] Figures 77A - 77D The data depicted demonstrate that MG71-2 recognizes multiple guide RNAs across various targets, allowing for the incorporation of larger genomic changes. Figure 77A : WT mRNA of MG71-2 has InDels at two targets (TRAC and AAVS1) with various guide RNAs. Target sites D3 and D4 on AAVS1 show some of the highest editing levels and are 69 nt apart on the AAVS1 target. The spacers of D3 and D4 are oriented in the correct orientation to be compatible with the twin, paste, and template jump (Tj) prime editing methods. Figure 77B : Tape station gel image confirming the replacement of the 69 nt sequence in the AAVS1 target with the 38 nt Bxb1 sequence using Bxb1-specific primers. Lanes G3 and H3 are two replicates of MMLV2-MG71-2n using a pegRNA containing the Bxb1 sequence and a nick guide (paste method), while lanes A4 and B4 represent two replicates of MMLV2-MG71-2n using a pegRNA containing the Bxb1 sequence and no nick guide. Lanes C4 and D4 are samples from MG151-98(H171N, K297P, 166AA)-MG71-2n using a pegRNA containing the Bxb1 sequence and no nick guide, while lanes E4 and D4 use a pegRNA containing the Bxb1 sequence and a nick guide (paste method). Figures 77C - 77D : Tapestation fragment analysis of lanes G3, H3, E4, and F4 confirmed that the amplicon contains the Bxb1 sequence.

[0093] Figures 78A - 78L Optimization of the MG71-2n system with the selected reverse transcriptase is depicted. Figures 78A - 78D: MG160-4(H230R) was cloned with a 33-amino acid linker at the N-terminus or C-terminus of MG71-2n. Additionally, MG160-4(H230R) and MG71-2n were mosaicked at five different insertion sites (S311, S355, T396, I822, and V1176). The mosaic constructs had 33-amino acid linkers at the 5' and 3' ends of MG160-4(H230R) at the insertion sites. Across four different PBS lengths, the mosaic constructs were tested for 5-nt variations and 24-nt insertions at the AAVS1 target. MG160-4(H230R) on the N-terminus of MG71-2n showed the highest editing level. Figures 78E - 78H : Various linker lengths (14AA, 15AA, 26AA, and 32AA) of the fusion of MG160-4 with the N-terminus of MG71-2 were tested along with the original 33AA linker. The 32AA and 33AA linkers had similar editing levels for both 5-nt variations and 24-nt insertions at the AAVS1 target. Figures 78I - 78L : Various linker lengths (7AA, 14AA, 15AA, 16AA, 26AA, 32AA, 44AA, and 58AA) of the fusion of RT MG160-473 or MG151-98(H171N, Δ166AA) with the N-terminus of MG71-2 were tested along with the original 33AA linker and screened for incorporation of 5-nt variations and 24-nt insertions at the AAVS1 target. Bars labeled "correct edit" represent the expected edit with no errors in the NGS amplicon, and bars labeled "incorrect edit" refer to the expected edit being incorporated and including errors within the NGS amplicon, or errors from RT incorporation and / or pegRNA scaffold incorporation. Bars labeled "edit" represent the expected edit with errors in the NGS amplicon (excluding pegRNA scaffold incorporation), and bars labeled "scaffold incorporation" represent the expected edit and scaffold incorporation of the pegRNA.

[0094] Figures 79A - 79O The therapeutic sites targeted with MG3-6-3-8 and MG3-6 are depicted. Figure 79A : WT mRNA of MG3-6 / 3-8 had InDels at treatment-related sites (A1AT, PAH R408W, G6PC1 Q347*, G6PC1 R83C, and hPDK1) with various guide RNAs. Guide RNAs indicated by dark gray bars were the spacer sequences selected for designing the pegRNA. Figures 79B - 79K: The prime editing systems MG160-4(H230R) tethered to the C-terminus of MG3-6-3-8n (MG3-6-3-8n-MG160-4(H230R)) and MMLV2 tethered to the C-terminus of MG3-6-3-8n (MG3-6-3-8n-MMLV2-) were tested at treatment-related sites. No editing was detected at the sites PAH R408W, G6PC1:R83C, and hPDK1. For A1AT and G6PC1 Q347*, some detectable editing levels were observed. Figures 79L - 79O : The editing by MG160-4(H230R) tethered to the N-terminus of MG3-6n or MG3-6-3-8n was compared with that by MMLV2 tethered to the C-terminus of MG3-6n or MG3-6-3-8n. These constructs targeted four treatment sites A1A and hPDK1. The bars labeled "correct editing" represent the expected editing without errors in the NGS amplicons, and the bars labeled "incorrect editing" refer to the incorporation of the expected editing and include errors within the NGS amplicons, or errors from RT incorporation and / or pegRNA scaffold incorporation. The bars labeled "editing" represent the expected editing with errors in the NGS amplicons (excluding pegRNA scaffold incorporation), and the bars labeled "scaffold incorporation" represent the expected editing and scaffold incorporation of the pegRNA. * indicates that fewer than 1000 reads were obtained for the NGS samples.

[0095] Figures 80A - 80D Optimization of the MG3-6n system with MG160-4 and MG160-4(H230R) is depicted. Figures 80A - 80B : MG160-4 was cloned into the N-terminus of MG3-6n with a 33AA (original linker length) as well as various linker lengths of 32AA, 44AA, and 58AA. These prime editing systems were then tested to correct two stop codons in the linker between hygromycin and the BFP-engineered cell line. PegRNAs with PBS lengths of 8, 10, and 13 nucleotides were tested. The highest editing level using a pegRNA with a PBS length of 8nt was shown with a fusion construct having 58AA. As the PBS length became longer, when using prime editing systems with different linker lengths, the differences between the linker systems showed less variability in editing levels. Figures 80C - 80D:In addition, MG160-4(H230R) and MG3-6n were inserted at five different insertion sites (K115, V208, K368, D550, and L881). The chimeric constructs had 33-amino acid linkers at the 5' and 3' ends of MG160-4(H230R) at the insertion sites. The correction of two stop codons across three different PBS lengths in the linker between hygromycin and the BFP-engineered cell line was tested for the chimeric constructs. The highest editing levels were observed when MG160-4(H230R) was tethered to the N-terminus of MG3-6n. The bars labeled "correct editing" represent the expected editing without errors in the NGS amplicons, and the bars labeled "incorrect editing" refer to the incorporation of the expected editing and include errors within the NGS amplicons, or errors resulting from RT incorporation and / or pegRNA scaffold incorporation. The bars labeled "editing" represent the expected editing with errors in the NGS amplicons (excluding pegRNA scaffold incorporation), and the bars labeled "scaffold incorporation" represent the expected editing and scaffold incorporation of the pegRNA.

[0096] Figures 81A - 81C Depicted is the screening of native reverse transcriptases tethered to the N-terminus of MG71-2n targeting AAVS1. Figure 81A :Summary of the targeting of 5-nt changes in AAVS1 by MG198 candidates tethered to the N-terminus of MG71-2n using pegRNAs with different PBS lengths (8, 10, 13, and 16 nt). Higher editing levels were observed for candidates MG198-6 and MG198-7 compared to the background. Figures 81B - 81C :MG160 candidates MG160-45, MG160-121, MG160-136, and MG160-232 were tethered to the N-terminus of MG71-2n and targeted 5-nt changes in AAVS1 using pegRNAs with different PBS lengths (8, 10, 13 nt). All MG160 candidates were slightly above background levels but showed poorer activity compared to MG160-4(H230R) and MMLV2 tethered to the N-terminus of MG71-2n. The bars labeled "correct editing" represent the expected editing without errors in the NGS amplicons, and the bars labeled "incorrect editing" refer to the incorporation of the expected editing and include errors within the NGS amplicons, or errors resulting from RT incorporation and / or pegRNA scaffold incorporation. The bars labeled "editing" represent the expected editing with errors in the NGS amplicons (excluding pegRNA scaffold incorporation), and the bars labeled "scaffold incorporation" represent the expected editing and scaffold incorporation of the pegRNA. *Indicates that fewer than 1000 reads were obtained for the NGS samples.

[0097] Figures 82A - 82IDepicts the screening of MG160 ASR candidates tethered to the N-terminus of MG71-2n for multiplex editing on the AAVS1 target. Figure 82A : Summary of the targeting of 5-nt changes in AAVS1 by MG160 ASR candidates tethered to the N-terminus of MG71-2n using pegRNAs with different PBS lengths (8, 10, 13, and 16 nt). Higher editing levels than background were observed for candidates MG160-491, MG160-492, and MG160-493. Figures 82B - 82C : Then, MG160-491, MG160-492, and MG160-493 were compared with wild-type MG160-4, MG160-4(H230R), MMLV2, and EC86 for 5-nt changes on AAVS1. All candidates were comparable to MG160-4(H230R). Then, the G-to-T transversion ( Figure 82D and 82G ), 24-nt insertion ( Figure 82E and 82H ), and 15-nt deletion (82F and 82I) of MG160-491, MG160-492, and MG160-493 were tested. Bars labeled "correct editing" represent the expected editing without errors in the NGS amplicons, and bars labeled "incorrect editing" refer to the incorporation of the expected editing and include errors within the NGS amplicons, or errors resulting from RT incorporation and / or pegRNA scaffold incorporation. Bars labeled "editing" represent the expected editing with errors in the NGS amplicons (excluding pegRNA scaffold incorporation), and bars labeled "scaffold incorporation" represent the expected editing and scaffold incorporation of the pegRNA. *Indicates that fewer than 1000 reads were obtained for the NGS samples.

[0098] Figures 83A - 83D Depicts the effect of nickase guides on prime editing efficiency. Figures 83A - 83B : Summary of prime editing efficiency using a set of nickase guides in K562 cells with MG160-4H230R-MG71-2n. Bars labeled "correct editing" represent the expected editing without errors in the NGS amplicons, and bars labeled "incorrect editing" refer to the incorporation of the expected editing and include errors within the NGS amplicons, or errors resulting from RT incorporation and / or pegRNA scaffold incorporation. Bars labeled "editing" represent the expected editing with errors in the NGS amplicons (excluding pegRNA scaffold incorporation), and bars labeled "scaffold incorporation" represent the expected editing and scaffold incorporation of the pegRNA. Figures 83C - 83D: Summary of prime editing efficiency using a set of nick guides in K562 cells with MMLV2-MG71-2n. The no nick bar indicates baseline editing with a 5nt variant guide, and the no guide indicates background editing in the mRNA sample only.

[0099] Figures 84A - 84D Depicts the effect of nick guides on prime editing efficiency in K562 and HEK293T cells. Figures 84A - 84B : Summary of prime editing efficiency using nick guides A2-H2 and A6-H6 from Figure 78 in K562 cells with MG160-4 H230R-MG71-2n, MMLV2-MG71-2n, and MG151-98-DM-SL1-MG71-2n. The no nick bar indicates baseline editing with a 5nt variant guide, and the no guide indicates background editing in the mRNA sample only. Bars labeled "correct edit" represent the expected edit with no errors in the NGS amplicon, and bars labeled "incorrect edit" refer to the expected edit being incorporated and including errors within the NGS amplicon, or errors resulting from RT incorporation and / or pegRNA scaffold incorporation. Bars labeled "edit" represent the expected edit with errors in the NGS amplicon (excluding pegRNA scaffold incorporation), and bars labeled "scaffold incorporation" represent the expected edit and scaffold incorporation of the pegRNA. Figures 84C - 84D : Summary of prime editing efficiency using nick guides A2-H2 and A6-H6 from Figure 78 in HEK293T cells with MG160-4 H230R-MG71-2n, MMLV2-MG71-2n, and MG151-98-DM-SL1-MG71-2n. The no nick bar indicates baseline editing with a 5nt variant guide, and the no guide indicates background editing in the mRNA sample only.

[0100] Figures 85A - 85B Depicts the effect of nick guides on prime editing efficiency in K562 cells. Figures 85A - 85B: Summary of prime editing efficiency using prime editors from Incision Guides A2-H2, A5-H5, and A6-H6 from Figure 78 in K562 cells with MG160-4 H230R-MG71-2n, MMLV2-MG71-2n, and MG151-98-DM-SL1-MG71-2n. pegRNAs with PBS lengths of 8, 10, 13, and 16 were used in these experiments to encode the change of a single nucleotide G to T at AAVS1. The no-incision bar represents baseline editing with the pegRNA having the indicated PBS length, and the no-prime bar represents background editing in the mRNA sample only. Bars labeled "correct edit" represent the expected edit with no errors in the NGS amplicon, and bars labeled "incorrect edit" refer to the incorporation of the expected edit and include errors within the NGS amplicon, or errors of incorporation from RT and / or scaffold incorporation of the pegRNA. Bars labeled "edit" represent the expected edit with errors in the NGS amplicon (excluding pegRNA scaffold incorporation), and bars labeled "scaffold incorporation" represent the expected edit and scaffold incorporation of the pegRNA.

[0101] Figures 86A - 86B Depicts the optimization of prime editing efficiency using incision guides. Figures 86A - 86B : Summary of prime editing efficiency using prime editor from Incision Guide E6 from Figure 78 in K562 cells with MG160-4H230R-MG71-2n. The no-incision bar indicates baseline editing with a 5nt change guide, and the no-guide indicates background editing in the mRNA sample only. Different ratios of pegRNA to incision guide were tested, and the editing efficiency was evaluated. Bars labeled "correct edit" represent the expected edit with no errors in the NGS amplicon, and bars labeled "incorrect edit" refer to the incorporation of the expected edit and include errors within the NGS amplicon, or errors of incorporation from RT and / or scaffold incorporation of the pegRNA. Bars labeled "edit" represent the expected edit with errors in the NGS amplicon (excluding pegRNA scaffold incorporation), and bars labeled "scaffold incorporation" represent the expected edit and scaffold incorporation of the pegRNA.

[0102] Brief Description of the Sequence Listing

[0103] The Sequence Listing submitted herewith provides exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems according to the present disclosure. The following is an exemplary description of the sequences therein.

[0104] SEQ ID NO:1-37 show the full-length nucleic acid sequences of the untethered MG151 family reverse transcriptases suitable for the gene editing systems described herein.

[0105] SEQ ID NO: 38 - 61 shows the full - length nucleic acid sequences of untethered MG153 family reverse transcriptases suitable for the gene - editing systems described herein.

[0106] SEQ ID NO: 62 - 68 shows the full - length nucleic acid sequences of untethered MG160 family reverse transcriptases suitable for the gene - editing systems described herein.

[0107] SEQ ID NO: 69 - 75 shows the full - length nucleic acid sequences of tethered MG160 family reverse transcriptases suitable for the gene - editing systems described herein.

[0108] SEQ ID NO: 76 - 83 shows the RNA sequences of chemically modified guide RNAs with different lengths of PBS and a single - point mutation (VEGFA spacer G changed to T) suitable for the gene - editing systems described herein.

[0109] SEQ ID NO: 84 - 91 shows the RNA sequences of chemically modified guide RNAs with different lengths of PBS and a single deletion (VEGFA spacer deletion change) suitable for the gene - editing systems described herein.

[0110] SEQ ID NO: 92 - 99 shows the RNA sequences of chemically modified guide RNAs with different lengths of PBS and a single insertion (VEGFA spacer single insertion) suitable for the gene - editing systems described herein.

[0111] SEQ ID NO: 100 - 101 shows the sequences of primers suitable for site - directed editing at the VEGFA locus.

[0112] SEQ ID NO: 102 shows the nucleic acid sequence of the VEGFA target site.

[0113] SEQ ID NO: 103 shows the nucleic acid sequence of an exemplary RT - nickase adapter.

[0114] SEQ ID NO: 104 shows the nucleic acid sequence of the MG3 effector nuclease suitable for the gene - editing systems described herein.

[0115] SEQ ID NO: 105 - 108 shows the nucleic acid sequences of endogenous targets AAVS1, B2M, CD5, and CD38.

[0116] SEQ ID NO: 109 - 140 shows the RNA sequences of chemically modified guide RNAs with spacers targeting AAVS1, B2M, CD5, and CD38 and having different lengths of PBS, which are suitable for the gene editing systems described herein.

[0117] SEQ ID NO: 141 - 148 show the sequences of primers suitable for site - directed editing at the AAVS1, B2M, CD5, and CD38 loci.

[0118] SEQ ID NO: 149 shows the RNA sequence of a chemically modified guide RNA with a spacer targeting VEGFA.

[0119] SEQ ID NO: 150 - 151 and 2580 - 2581 show the sequences of two retrotransposition assay reporter genes.

[0120] SEQ ID NO: 152 - 154 show the amino acid sequences of MG3 - 6 nucleases (nMG3 - 6 D13A, nMG3 - 6H586A, and nMG3 - 6N609A).

[0121] SEQ ID NO: 155 - 160 show the amino acid sequences of exemplary RT - nickase adaptors.

[0122] SEQ ID NO: 161 - 291 show the amino acid sequences of MG140 family retrotransposon proteins suitable for the gene editing systems described herein.

[0123] SEQ ID NO: 292 - 293 show the amino acid sequences of MG146 family retrotransposon proteins suitable for the gene editing systems described herein.

[0124] SEQ ID NO: 294 - 317 show the amino acid sequences of MG148 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0125] SEQ ID NO: 318 - 330 show the amino acid sequences of MG149 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0126] SEQ ID NO: 331 - 445 show the amino acid sequences of MG151 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0127] SEQ ID NO: 446 - 499 show the amino acid sequences of MG153 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0128] SEQ ID NO: 500 - 501 show the amino acid sequences of the MG154 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0129] SEQ ID NO: 502 - 506 show the amino acid sequences of the MG155 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0130] SEQ ID NO: 507 - 508 show the amino acid sequences of the MG156 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0131] SEQ ID NO: 509 - 513 show the amino acid sequences of the MG157 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0132] SEQ ID NO: 514 show the amino acid sequences of the MG158 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0133] SEQ ID NO: 515 - 517 show the amino acid sequences of the MG159 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0134] SEQ ID NO: 518 - 566 show the amino acid sequences of the MG160 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0135] SEQ ID NO: 567 - 571 show the amino acid sequences of the MG163 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0136] SEQ ID NO: 572 - 576 show the amino acid sequences of the MG164 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0137] SEQ ID NO: 577 - 585 show the amino acid sequences of the MG165 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0138] SEQ ID NO: 586 - 590 show the amino acid sequences of the MG166 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0139] SEQ ID NO: 591 - 595 show the amino acid sequences of the MG167 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0140] SEQ ID NO: 596 - 600 show the amino acid sequences of the MG168 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0141] SEQ ID NO: 601 - 611 show the amino acid sequences of the MG169 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0142] SEQ ID NO: 612 - 621 show the amino acid sequences of the MG170 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0143] SEQ ID NO: 622 - 626 show the amino acid sequences of the MG172 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0144] SEQ ID NO: 627 - 628 show the amino acid sequences of the MG173 family reverse transcriptase proteins suitable for the gene editing systems described herein.

[0145] SEQ ID NO: 629 shows the amino acid sequence of the MG176 family retrotransposon protein suitable for the gene editing systems described herein.

[0146] SEQ ID NO: 630 - 645 show the nuclear localization signal (NLS) suitable for the gene editing systems described herein.

[0147] SEQ ID NO: 646 shows the amino acid sequence of the MG3 - 6 nuclease suitable for the gene editing systems described herein.

[0148] SEQ ID NO: 647 shows the amino acid sequence of the MG29 - 1 nuclease suitable for the gene editing systems described herein.

[0149] SEQ ID NO: 648 shows the nucleotide sequence of the RNA template for cDNA synthesis.

[0150] SEQ ID NO: 653 shows the nucleotide sequence of MG3 - 6(H586A).

[0151] SEQ ID NO: 654 - 655 show the nucleotide sequences of the cDNA encoding gene targets.

[0152] SEQ ID NO: 656 - 697 show the full - length peptide sequences of the chemically modified guide RNAs.

[0153] SEQ ID NO: 698 - 701 show the nucleotide sequences of the primers.

[0154] SEQ ID NO: 702 - 709 show the nucleotide sequences of reverse transcriptases cloned into the tethered MG3 - 6(H586A) plasmid.

[0155] SEQ ID NO: 710 - 727 show the nucleotide sequences of genes encoding the MG151 reverse transcriptase protein, which are optimized for expression in mammalian cells and cloned into untethered plasmids.

[0156] SEQ ID NO: 728 - 749 show the nucleotide sequences of genes encoding the MG160 reverse transcriptase protein, which are optimized for expression in mammalian cells and cloned into the tethered spCas9(H840A) plasmid.

[0157] SEQ ID NO: 750 - 766 show the nucleotide sequences of genes encoding the MG151 reverse transcriptase protein, which are optimized for expression in mammalian cells and cloned into untethered plasmids.

[0158] SEQ ID NO: 767 - 784 show the full - length peptide sequence of the MG151 reverse transcriptase protein.

[0159] SEQ ID NO: 786 - 1220 show the full - length peptide sequence of the MG160 reverse transcriptase protein.

[0160] SEQ ID NO: 1221 - 1226 and 1299 show the nucleotide sequences of genes encoding the MG153 reverse transcriptase protein, which are optimized for expression in mammalian cells and cloned into untethered plasmids.

[0161] SEQ ID NO: 1227 - 1243, 1250 - 1256 and 1265 - 1271 show the nucleotide sequences of genes encoding the MG160 reverse transcriptase protein, which are optimized for expression in mammalian cells and cloned into the tethered spCas9(H840A) plasmid.

[0162] SEQ ID NO: 1245 - 1246 show the nucleotide sequence of the RT linker.

[0163] SEQ ID NO: 1257 - 1264 and 1272 - 1279 show the nucleotide sequences of genes encoding the MG160 reverse transcriptase protein, which are optimized for expression in mammalian cells and cloned into untethered plasmids.

[0164] SEQ ID NO: 1280 - 1292 and 1299 show the nucleotide sequences of genes encoding the reverse transcriptase protein, which are optimized for expression in mammalian cells and cloned into untethered plasmids.

[0165] SEQ ID NOs: 1293 - 1295 and 1300 show the nucleotide sequences of genes encoding reverse transcriptase proteins that are optimized for expression in mammalian cells and cloned into tethered or untethered plasmids.

[0166] SEQ ID NOs: 1301 - 1304 and 1309 show the nucleotide sequences of genes encoding mutant reverse transcriptase proteins that are optimized for expression in mammalian cells and cloned into tethered or untethered plasmids.

[0167] SEQ ID NOs: 1336 - 1341 show the nucleotide sequences of chemically modified guide RNAs with different lengths of PBS and a single point mutation (AAVS1 spacer G to T) suitable for the gene editing system described herein.

[0168] SEQ ID NOs: 1330 - 1335 show the nucleotide sequences of chemically modified guide RNAs with different lengths of PBS and a single deletion (AAVS1 spacer deletion change) suitable for the gene editing system described herein.

[0169] SEQ ID NOs: 1324 - 1329 show the nucleotide sequences of chemically modified guide RNAs with different lengths of PBS and a single insertion (AAVS1 spacer single insertion) suitable for the gene editing system described herein.

[0170] SEQ ID NOs: 1310 - 1315 show the nucleotide sequences of chemically modified guide RNAs with different lengths of PBS and a 5 - nucleotide change (for targeting AAVS1) suitable for the gene editing system described herein.

[0171] SEQ ID NOs: 1317 - 1323 show the nucleotide sequences of chemically modified guide RNAs with different lengths of PBS and a modified backbone (for targeting AAVS1) suitable for the gene editing system described herein.

[0172] SEQ ID NOs: 1342 - 1343 show the nucleotide sequences of the MG71 - 2 AAVS1 primers.

[0173] SEQ ID NO: 1344 shows the nucleotide sequence of the cDNA encoding the gene target.

[0174] SEQ ID NO: 1247 shows the nucleotide sequence of the spCas9(H840A) untethered or tethered plasmid.

[0175] SEQ ID NO:1248 shows the nucleotide sequence of MMLV1 that has been codon-optimized for expression in mammalian cells and cloned into tethered or untethered plasmids.

[0176] SEQ ID NO:1249 shows the nucleotide sequence of MMLV2 that has been codon-optimized for expression in mammalian cells and cloned into tethered or untethered plasmids.

[0177] SEQ ID NO:1345 - 1353 shows the nucleotide sequence of ncRNA.

[0178] SEQ ID NO:1354 - 1361 shows the nucleotide sequence of primers.

[0179] SEQ ID NO:1362 - 1393 shows the nucleotide sequence of ncRNA.

[0180] SEQ ID NO:1394 - 1401 shows the nucleotide sequence of the MG173 family reverse transcriptase that has been codon-optimized for expression in mammalian cells and cloned into an untethered plasmid.

[0181] SEQ ID NO:1402 shows the nucleotide sequence of the MG192 family reverse transcriptase that has been codon-optimized for expression in mammalian cells and cloned into an untethered plasmid.

[0182] SEQ ID NO:1403 - 1424 shows the nucleotide sequence of the MG160 family reverse transcriptase that has been codon-optimized for expression in mammalian cells and cloned into a tethered plasmid.

[0183] SEQ ID NO:1426 - 1438 shows the nucleotide sequence of the MG151 family reverse transcriptase that has been codon-optimized for expression in mammalian cells and cloned into tethered or untethered plasmids.

[0184] SEQ ID NO:1439 - 1444 shows the nucleotide sequence of the MG153 family reverse transcriptase that has been codon-optimized for expression in mammalian cells and cloned into tethered or untethered plasmids.

[0185] SEQ ID NO:1445 - 1446 shows the nucleotide sequence of the MG160 family reverse transcriptase that has been codon-optimized for expression in mammalian cells and cloned into a tethered plasmid.

[0186] SEQ ID NO:1447 shows the nucleotide sequence of the MG151 family reverse transcriptase that has been codon-optimized for expression in mammalian cells and cloned into tethered or untethered plasmids.

[0187] SEQ ID NO: 1448 - 1450 shows the nucleotide sequence of the MG71 - 2 scaffold.

[0188] SEQ ID NO: 1451 - 1462 shows the nucleotide sequences of chemically modified guide RNAs with 5 - nucleotide variations and different lengths of PBS (for targeting AAVS1) suitable for the gene - editing systems described herein.

[0189] SEQ ID NO: 1463 - 1470 shows the nucleotide sequences of chemically modified guide RNAs with modified scaffolds and different lengths of PBS (for targeting AAVS1) suitable for the gene - editing systems described herein.

[0190] SEQ ID NO: 1471 - 1474 shows the nucleotide sequences of chemically modified guide RNAs with 2 - nucleotide variations and different lengths of PBS (for targeting AAVS1) suitable for the gene - editing systems described herein.

[0191] SEQ ID NO: 1475 shows the nucleotide sequence of the mRNA encoding MG3 - 6 codon - optimized for expression in mammalian cells.

[0192] SEQ ID NO: 1476 shows the nucleotide sequence of the mRNA encoding MG3 - 6 / 3 - 8 codon - optimized for expression in mammalian cells.

[0193] SEQ ID NO: 1477 shows the nucleotide sequence of the mRNA encoding MG14 - 241 codon - optimized for expression in mammalian cells.

[0194] SEQ ID NO: 1478 shows the nucleotide sequence of the mRNA encoding MG14 - 241(H596A) codon - optimized for expression in mammalian cells.

[0195] SEQ ID NO: 1479 - 1492 shows the nucleotide sequences of chemically modified guide RNAs with 5 - nucleotide variations and different lengths of PBS (for targeting AAVS1) suitable for the gene - editing systems described herein.

[0196] SEQ ID NO: 1493 - 1504 shows the nucleotide sequences of the NGS primers.

[0197] SEQ ID NO: 1505 - 1510 shows the nucleotide sequences of the cDNA of the endogenous target.

[0198] SEQ ID NO:1511 shows the nucleotide sequence of the engineered landing pad.

[0199] SEQ ID NO:1512 - 1516 show the nucleotide sequences of the Cas9 guides targeting the engineered sites.

[0200] SEQ ID NO:1518 - 1519 show the nucleotide sequences of the primers.

[0201] SEQ ID NO:1520 - 1531 show the nucleotide sequences encoding the MGRT / Cas9 fusion protein codon - optimized for expression in mammalian systems.

[0202] SEQ ID NO:1532 - 1540 show the nucleotide sequences of the RNA cargo for integration.

[0203] SEQ ID NO:1541 - 1547 show the nucleotide sequences of the primers.

[0204] SEQ ID NO:1548 - 1555 show the nucleotide sequences of the RNA templates.

[0205] SEQ ID NO:1557 - 1560 show the nucleotide sequences of the primers.

[0206] SEQ ID NO:1561 - 1562 show the nucleotide sequences of the Taqman probes.

[0207] SEQ ID NO:1563 shows the nucleotide sequence encoding the nMRA of MG71 - 2 codon - optimized for expression in mammalian systems.

[0208] SEQ ID NO:1564 shows the nucleotide sequence of the MG71 - 2 guide.

[0209] SEQ ID NO:1566 - 1567 show the nucleotide sequences of the NGS primers.

[0210] SEQ ID NO:1568 - 1573 show the nucleotide sequences of the MG71 - 2 guides.

[0211] SEQ ID NO:1574 - 1576 show the nucleotide sequences of the MG71 - 2 pegRNA.

[0212] SEQ ID NO:1577 - 1578 show the nucleotide sequences of the NGS primers.

[0213] SEQ ID NO:1579 shows the nucleotide sequence of the cDNA encoding the endogenous target.

[0214] SEQ ID NO:1580 - 1581 show the nucleotide sequences of the NGS primers.

[0215] SEQ ID NO:1582 shows the nucleotide sequence of the cDNA encoding the endogenous target.

[0216] SEQ ID NO:1583 - 1584 show the nucleotide sequences of the NGS primers.

[0217] SEQ ID NO:1585 shows the nucleotide sequence of the cDNA encoding the endogenous target.

[0218] SEQ ID NO:1586 - 1587 show the nucleotide sequences of the NGS primers.

[0219] SEQ ID NO:1588 shows the nucleotide sequence of the cDNA encoding the endogenous target.

[0220] SEQ ID NO:1589 - 1590 show the nucleotide sequences of the NGS primers.

[0221] SEQ ID NO:1591 shows the nucleotide sequence of the cDNA encoding the endogenous target.

[0222] SEQ ID NO:1592 - 1593 show the nucleotide sequences of the reverse transcriptase that has been codon - optimized for expression in mammalian cells and cloned into a tethered plasmid.

[0223] SEQ ID NO:1596 - 1597 show the nucleotide sequences of the reverse transcriptase that has been codon - optimized for expression in mammalian cells and cloned into a tethered or untethered plasmid.

[0224] SEQ ID NO:1598 - 1609 show the nucleotide sequences of the MG71 - 2 pegRNA.

[0225] SEQ ID NO:1610 - 1620 show the nucleotide sequences of the MG71 - 2 guide.

[0226] SEQ ID NO:1621 - 1622 show the nucleotide sequences of the NGS primers.

[0227] SEQ ID NO:1623 shows the nucleotide sequence of the cDNA encoding the endogenous target.

[0228] SEQ ID NO: 1624 - 1625 shows the nucleotide sequences of NGS primers.

[0229] SEQ ID NO: 1626 shows the nucleotide sequence of the cDNA encoding an endogenous target.

[0230] SEQ ID NO: 1627 - 1628 shows the nucleotide sequences of NGS primers.

[0231] SEQ ID NO: 1629 shows the nucleotide sequence of the cDNA encoding an endogenous target.

[0232] SEQ ID NO: 1630 - 1631 shows the nucleotide sequences of NGS primers.

[0233] SEQ ID NO: 1632 shows the nucleotide sequence of the cDNA encoding an endogenous target.

[0234] SEQ ID NO: 1633 - 1634 shows the nucleotide sequences of NGS primers.

[0235] SEQ ID NO: 1635 shows the nucleotide sequence of the cDNA encoding an endogenous target.

[0236] SEQ ID NO: 1636 - 1637 shows the nucleotide sequences of NGS primers.

[0237] SEQ ID NO: 1638 shows the nucleotide sequence of the cDNA encoding an endogenous target.

[0238] SEQ ID NO: 1639 - 1640 shows the nucleotide sequences of NGS primers.

[0239] SEQ ID NO: 1641 shows the nucleotide sequence of the cDNA encoding an endogenous target.

[0240] SEQ ID NO: 1642 - 1643 shows the nucleotide sequences of NGS primers.

[0241] SEQ ID NO: 1644 shows the nucleotide sequence of the cDNA encoding an endogenous target.

[0242] SEQ ID NO: 1645 - 1646 shows the nucleotide sequences of NGS primers.

[0243] SEQ ID NO: 1647 shows the nucleotide sequence of the cDNA encoding an endogenous target.

[0244] SEQ ID NO: 1648 - 1649 shows the nucleotide sequences of the NGS primers.

[0245] SEQ ID NO: 1650 shows the nucleotide sequence of the cDNA encoding the endogenous target.

[0246] SEQ ID NO: 1651 - 1652 shows the nucleotide sequences of the NGS primers.

[0247] SEQ ID NO: 1653 shows the nucleotide sequence of the cDNA encoding the endogenous target.

[0248] SEQ ID NO: 1654 shows the nucleotide sequence of the reverse transcriptase that has been codon - optimized for expression in mammalian cells and cloned into a plasmid.

[0249] SEQ ID NO: 1656 - 1681 shows the nucleotide sequences of the MG71 - 2 pegRNA.

[0250] SEQ ID NO: 1682 shows the nucleotide sequence of the primer.

[0251] SEQ ID NO: 1683 - 1690 shows the nucleotide sequences of the MG71 - 2 pegRNA.

[0252] SEQ ID NO: 1691 - 1720 shows the nucleotide sequences of the reverse transcriptase that has been codon - optimized for expression in mammalian cells and cloned into a plasmid.

[0253] SEQ ID NO: 1722 - 1749 shows the nucleotide sequences of the MG3 - 6 / 3 - 8 guide.

[0254] SEQ ID NO: 1750 - 1751 shows the nucleotide sequences of the NGS primers.

[0255] SEQ ID NO: 1752 shows the nucleotide sequence of the cDNA encoding the endogenous target.

[0256] SEQ ID NO: 1753 - 1754 shows the nucleotide sequences of the reverse transcriptase that has been codon - optimized for expression in mammalian cells and cloned into a plasmid.

[0257] SEQ ID NO: 1755 - 1774 shows the nucleotide sequences of the MG3 - 6 / 3 - 8 pegRNA.

[0258] SEQ ID NO: 1776 - 1778 shows the nucleotide sequences of the reverse transcriptase that has been codon - optimized for expression in mammalian cells and cloned into a plasmid.

[0259] SEQ ID NO:1779 shows the nucleotide sequence of a target optimized for expression in mammalian cells by codon optimization.

[0260] SEQ ID NO:1780 - 1783 show the nucleotide sequences of reverse transcriptases optimized for expression in mammalian cells and cloned into plasmids.

[0261] SEQ ID NO:1784 - 1786 show the nucleotide sequences of MG3 - 6 pegRNA.

[0262] SEQ ID NO:1787 - 1788 show the nucleotide sequences of NGS primers.

[0263] SEQ ID NO:1789 shows the nucleotide sequence of a cDNA encoding an endogenous target.

[0264] SEQ ID NO:1790 - 1847 show the nucleotide sequences of reverse transcriptases optimized for expression in mammalian cells and cloned into plasmids.

[0265] SEQ ID NO:1848 - 1855 show the nucleotide sequences of MG71 - 2 pegRNA.

[0266] SEQ ID NO:1856 - 1858 show the nucleotide sequences of reverse transcriptases optimized for expression in mammalian cells and cloned into plasmids.

[0267] SEQ ID NO:1859 - 1862 show the nucleotide sequences of plasmids encoding MG nickase optimized for expression in mammalian cells.

[0268] SEQ ID NO:1863 - 1910 show the nucleotide sequences of MG71 - 2 guide RNAs targeting AAVS1.

[0269] SEQ ID NO:1911 - 1958 show the DNA sequences of AAVS1 target sites.

[0270] SEQ ID NO:1959 - 2002 show the full - length peptide sequences of MG140 reverse transcriptase protein.

[0271] SEQ ID NO:2003 - 2084 show the full - length peptide sequences of MG153 reverse transcriptase protein.

[0272] SEQ ID NO:2085 - 2092 show the full - length peptide sequences of MG157 reverse transcriptase protein.

[0273] SEQ ID NO: 2093 - 2112 shows the full - length peptide sequence of the MG165 reverse transcriptase protein.

[0274] SEQ ID NO: 2113 - 2156 shows the full - length peptide sequence of the MG166 reverse transcriptase protein.

[0275] SEQ ID NO: 2157 - 2186 shows the full - length peptide sequence of the MG167 reverse transcriptase protein.

[0276] SEQ ID NO: 2187 - 2223 shows the full - length peptide sequence of the MG169 reverse transcriptase protein.

[0277] SEQ ID NO: 2224 shows the full - length peptide sequence of the MG176 reverse transcriptase protein.

[0278] SEQ ID NO: 2225 - 2252 shows the full - length peptide sequence of the MG198 reverse transcriptase protein.

[0279] SEQ ID NO: 2253 - 2256 shows the full - length peptide sequence of the MG173 reverse transcriptase protein.

[0280] SEQ ID NO: 2257 - 2289 shows the full - length peptide sequence of the MG140 reverse transcriptase protein.

[0281] SEQ ID NO: 2290 - 2471 and 2582 - 2585 show the full - length peptide sequence of the MG160 reverse transcriptase protein.

[0282] SEQ ID NO: 2472 - 2517 shows the full - length peptide sequence of the MG140 retrotransposase protein.

[0283] SEQ ID NO: 2518 - 2520 shows the full - length peptide sequence of the MG160 retrotransposase protein.

[0284] SEQ ID NO: 2522 shows the full - length peptide sequence of the MG153 reverse transcriptase protein.

[0285] SEQ ID NO: 2523 - 2530 shows the nucleotide sequence of the MG140 UTR.

[0286] SEQ ID NO: 2531 - 2540 shows the nucleotide sequence of the MG153 RNA.

[0287] SEQ ID NO: 2541 - 2571 shows the nucleotide sequence of the MG140 UTR. Detailed implementation manners

[0288] Site-directed gene editing systems are powerful tools for site-directed genome engineering in cells. Programmable nucleases, such as clustered regularly interspaced short palindromic repeat (CRISPR) nucleases, have recently been used in various DNA manipulation and gene editing applications. CRISPR nucleases can be used to introduce site-directed insertions and deletions (indels) or point mutations of different lengths with or without a repair template. Single nucleotide point (SNP) mutations, deletions, and insertions represent more than 80% of pathogenic mutations. However, not all of these mutations can be precisely repaired with available gene editing systems. There is a need for clinical genome editing applications with higher efficiency and system fidelity.

[0289] In addition, repairing or inserting longer DNA fragments remains challenging, and there is a lack of a safe and effective way to target the integration of large templates into the genome, such as for gene therapy or engineered cell therapy. To date, lentiviruses or adeno-associated viruses (AAVs) have been combined with CRISPR nucleases to insert large DNA fragments, such as intact genes. However, lentivirus-mediated integration lacks targeting features because integration mostly occurs randomly in open chromatin. AAV-mediated delivery has limited cargo capacity and cannot be used for all cell types. There is a need for a safe and effective targeted genome editing system that allows for the integration of large templates.

[0290] This disclosure is partly based on the development of a gene editing system that includes a reverse transcriptase, a nuclease or nickase, and a guide RNA or pegRNA. The gene editing system can be used to introduce site-directed insertions, deletions, and mutations into the genome of a cell. In addition, it is expected that the gene editing system can be used in combination with a nucleic acid template to facilitate site-directed insertion into the genome of a cell and for large template integration.

[0291] Definition

[0292] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the claimed subject matter belongs. It should be understood that the foregoing general description and the following detailed description are merely exemplary and explanatory and do not limit any claimed subject matter. The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described.

[0293] Unless otherwise indicated, the practice of some of the methods disclosed herein employs techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA. See, for example, Sambrook and Green et al., Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (edited by F. M. Ausubel et al.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (edited by M. J. MacPherson, B. D. Hames, and G. R. Taylor (1995)); Antibodies, A Laboratory Manual (edited by Harlow and Lane (1988)), and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (edited by R. I. Freshney, (2010)).

[0294] As used herein, unless the context clearly indicates otherwise, the singular forms "a / an" and "the" are also intended to include the plural forms. Additionally, when the terms "including", "include", "having", "has", "with" or variants thereof are used in the detailed description and / or claims, such terms are intended to be inclusive in a manner similar to the term "comprising".

[0295] The term "about" or "approximately" means within an acceptable error range of a particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within one or more standard deviations in accordance with the practice in the art. Alternatively, "about" can mean a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.

[0296] As used herein, the term "nucleotide" refers to a base-sugar-phosphate combination. Envisioned nucleotides include naturally occurring nucleotides and synthetic nucleotides. Nucleotides are the monomeric units of nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide includes ribonucleoside triphosphates, adenosine triphosphate (ATP), uridine triphosphate (UTP), cytidine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives include, for example, [αS]dATP, 7-deaza-dGTP, and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance on nucleic acid molecules containing them. As used herein, the term nucleotide encompasses dideoxyribonucleoside triphosphates (ddNTPs) and derivatives thereof. Illustrative examples of ddNTPs include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides can be unlabeled or detectably labeled, such as with a moiety containing an optically detectable moiety (e.g., a fluorophore) or a quantum dot. Detectable labels include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labels of nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'-dimethylaminophenylazo)benzoic acid (DABCYL), cascade blue, Oregon green, Texas red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, which are available from Perkin Elmer, Foster City, Calif; FluoroLink deoxynucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP, which are available from Amersham, Arlington Heights, IL; fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP, which are available from Boehringer Mannheim, Indianapolis, Ind; and chromosomally labeled nucleotides, BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, cascade blue-7-UTP, cascade blue-7-dUTP, fluorescein-12-UTP, fluorescein-12-dUTP, Oregon Green 488-5-dUTP, rhodamine green-5-UTP, rhodamine green-5-dUTP, tetramethylrhodamine-6-UTP, tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP, which are available from Molecular Probes, Eugene, Oreg. The term nucleotide encompasses chemically modified nucleotides. An exemplary chemically modified nucleotide is biotin-dNTP.Non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).

[0297] The terms "polynucleotide", "oligonucleotide", and "nucleic acid" are used interchangeably to refer to polymeric forms of nucleotides of any length, deoxyribonucleotides or ribonucleotides or their analogs, in single-stranded, double-stranded, or multi-stranded form. Envisioned polynucleotides include genes or fragments thereof. Exemplary polynucleotides include, but are not limited to, DNA, RNA, coding or non-coding regions of genes or gene fragments, multiple loci (locus) defined according to linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, cell-free polynucleotides comprising cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. In polynucleotides, when referring to T, T means U (uracil) in RNA and T (thymine) in DNA. Polynucleotides can be exogenous or endogenous to a cell and / or present in a cell-free environment. The term polynucleotide encompasses modified polynucleotides (e.g., altered backbone, sugar, or nucleobase). Modifications, if present, are imparted to the nucleotide structure either before or after polymer assembly. Non-limiting examples of modifications include: 5-bromouracil, peptide nucleic acid, heterologous nucleic acid, morpholino, locked nucleic acid, glycerol nucleic acid, threose nucleic acid, dideoxynucleotide, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to a sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, wybutosine, and queuosine. The sequence of nucleotides can be interrupted by non-nucleotide components.

[0298] The term "transfection" or "transfected" refers to the introduction of a polynucleotide into a cell by non-viral or virus-based methods. The polynucleotide can be a gene sequence encoding a full-length protein or a functional portion thereof. See, e.g., Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1-18.88.

[0299] The terms "peptide", "polypeptide", and "protein" are used interchangeably herein to refer to a polymer of at least two amino acid residues joined by peptide bonds. This term does not denote a specific length of the polymer and is not intended to imply or distinguish whether the peptide is produced using recombinant techniques, chemical or enzymatic synthesis, or is naturally occurring. The term applies to both naturally occurring amino acid polymers and amino acid polymers that contain at least one modified amino acid. In some cases, the polymer may be interspersed with non-amino acids. The term includes amino acid chains of any length, including full-length proteins and proteins with or without secondary or tertiary structure (e.g., domains). The term also encompasses amino acid polymers that have been modified; for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, oxidation, and any other manipulations, such as conjugation with a labeling component. As used herein, the term "amino acid / amino acids" refers to natural and non-natural amino acids, including but not limited to modified amino acids. Modified amino acids include natural amino acids that have been chemically modified to include groups or chemical moieties that are not naturally present on the amino acid. The term "amino acid" includes both D-amino acids and L-amino acids.

[0300] As used herein, "non-natural" refers to a nucleic acid or polypeptide sequence that is not naturally occurring. Non-natural refers to a nucleic acid or polypeptide sequence that contains modifications such as mutations, insertions, or deletions. The term non-natural encompasses fusion nucleic acids or polypeptides that encode or exhibit the activity of a nucleic acid or polypeptide sequence fused to a non-natural sequence (e.g., enzyme activity, methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitination activity, etc.). Non-natural nucleic acid or polypeptide sequences include those non-natural nucleic acid or polypeptide sequences that have been genetically engineered to be linked to a naturally occurring nucleic acid or polypeptide sequence (or a variant thereof) to produce a chimeric nucleic acid or a polypeptide sequence encoding a chimeric nucleic acid or polypeptide.

[0301] As used herein, the term "promoter" refers to a regulatory DNA region that controls the transcription or expression of a polynucleotide (e.g., a gene) and that may be adjacent to or overlap with the nucleotide or nucleotide region that initiates RNA transcription. A promoter may contain specific DNA sequences that bind protein factors (commonly referred to as transcription factors) that facilitate the binding of RNA polymerase to the DNA, thereby resulting in gene transcription. Eukaryotic basal promoters typically (although not necessarily) contain a TATA box and / or a CAAT box.

[0302] As used herein, the term "expression" refers to the process of transcribing a nucleic acid sequence or polynucleotide from a DNA template (such as transcription into mRNA or other RNA transcripts) and / or the subsequent translation of the transcribed mRNA into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may be collectively referred to as a "gene product". If the polynucleotide is derived from genomic DNA, the term expression includes the splicing of mRNA in eukaryotic cells.

[0303] As used herein, "operably linked", "operably connect", "operatively linked" or their grammatical equivalents refer to the arrangement of genetic elements, such as promoters, enhancers, polyadenylation sequences, etc., where the operation (such as movement or activation) of the first genetic element has some effect on the second genetic element. The effect on the second genetic element may be, but need not be, of the same type as the operation of the first genetic element. For example, if the movement of the first element results in the activation of the second element, the two genetic elements are operably linked. For example, if a regulatory element contributes to the transcription of a coding sequence, a regulatory element containing a promoter and / or enhancer sequence is operably linked to the coding region. There may be intervening residues between the regulatory element and the coding region as long as this functional relationship is maintained.

[0304] As used herein, a "vector" refers to a macromolecule or an association of macromolecules that contains a polynucleotide or is associated with a polynucleotide and mediates the delivery of the polynucleotide into a cell. Examples of vectors include nucleic acid-based vectors (such as plasmids and viral vectors) and liposomes. Exemplary nucleic acid-based vectors contain genetic elements (such as regulatory elements) that are operably linked to a gene to facilitate the expression of the gene in a target.

[0305] As used herein, "expression cassette" and "nucleic acid cassette" are used interchangeably to refer to a component of a vector that contains a combination of nucleic acid sequences or elements (such as a therapeutic gene, a promoter, and a terminator) that are expressed together or are operably linked for expression. These terms encompass expression cassettes that include a combination of regulatory elements and one or more genes that are operably linked thereto for expression.

[0306] A "functional fragment" of a DNA or protein sequence refers to a fragment that retains a biological activity (function or structure) that is substantially similar to the biological activity of the full-length DNA or protein sequence. The biological activity of a DNA sequence includes its ability to affect expression in a manner attributable to the full-length sequence.

[0307] The terms "engineered", "synthetic", and "artificial" are used interchangeably herein to refer to an object that has been modified by human intervention. For example, these terms refer to polynucleotides or polypeptides that do not occur in nature. Engineered peptides have, but do not require, low sequence identity with naturally occurring human proteins (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1% sequence identity). For example, the VPR and VP64 domains are synthetic transactivation domains. Non-limiting examples include the following: a nucleic acid is modified by changing its sequence to a sequence that does not exist in nature; a nucleic acid is modified by ligating it to a nucleic acid that does not associate with it in nature such that the ligation product has a function not present in the original nucleic acid; an engineered nucleic acid is synthesized in vitro with a sequence that does not exist in nature; a protein is modified by changing the amino acid sequence of the protein to a sequence that does not exist in nature; an engineered protein acquires a new function or property. An "engineered" system comprises at least one engineered component.

[0308] As used herein, a "guide nucleic acid" or "guide polynucleotide" refers to a nucleic acid that can hybridize to a target nucleic acid and thereby direct an associated nuclease to the target nucleic acid. Guide nucleic acids are, but are not limited to, RNA (guide RNA or gRNA), DNA, or a mixture of RNA and DNA. A guide nucleic acid can include a crRNA or a tracrRNA or a combination of both. The term guide nucleic acid encompasses engineered guide nucleic acids and programmable guide nucleic acids that specifically bind to a target nucleic acid. A portion of the target nucleic acid can be complementary to a portion of the guide nucleic acid. The strand of the double-stranded target polynucleotide that is complementary and hybridizes to the guide nucleic acid is the complementary strand. The strand of the double-stranded target polynucleotide that is complementary to the complementary strand and thus not complementary to the guide nucleic acid is referred to as the non-complementary strand. A guide nucleic acid having a polynucleotide strand is a "single guide nucleic acid". A guide nucleic acid having two polynucleotide strands is a "dual guide nucleic acid". Unless otherwise specified, the term "guide nucleic acid" is inclusive and refers to both single guide nucleic acids and dual guide nucleic acids. A guide nucleic acid can contain a segment referred to as a "nucleic acid targeting segment" or "nucleic acid targeting sequence" or "spacer". The nucleic acid targeting segment can include sub-segments that are referred to as "protein binding segments" or "protein binding sequences" or "Cas protein binding segments".

[0309] The term "tracrRNA" or "tracr sequence" means trans-activating CRISPR RNA. The tracrRNA interacts with a CRISPR (cr)RNA to form a guide nucleic acid (e.g., guide RNA or gRNA) that can hybridize to a target nucleic acid and thereby direct an associated nuclease to the target nucleic acid.

[0310] As used herein, the term "RuvC_III domain" refers to the third discontinuous segment of the RuvC endonuclease domain (the RuvC nuclease domain comprises three discontinuous segments, RuvC_I, RuvC_II, and RuvC_III). The RuvC domain or its segments can generally be identified by alignment with a recorded domain sequence, alignment of the structure of a protein with an annotated domain, or by comparison with a Hidden Markov Model (HMM) built on a recorded domain sequence (e.g., the Pfam HMM PF18541 of RuvC_III).

[0311] As used herein, the term "HNH domain" refers to an endonuclease domain having characteristic histidine and asparagine residues. The HNH domain can generally be identified by alignment with a recorded domain sequence, alignment of the structure of a protein with an annotated domain, or by comparison with a Hidden Markov Model (HMM) built on a recorded domain sequence (e.g., the Pfam HMM PF01844 of the domain HNH).

[0312] As used herein, the term "transposon" refers to a mobile element that carries "cargo DNA" with it as it moves in and out of the genome. These transposons can differ in the type of nucleic acid being transposed, the type of repeats at the ends of the transposon, the type of cargo to be carried, or the mode of transposition (i.e., self-healing or host-healing).

[0313] As used herein, the term "transposase (transposase or transposases)" refers to an enzyme that binds to the ends of a transposon and catalyzes its movement to another part of the genome. The types of movement include the cut-and-paste mechanism and the replicative transposition mechanism.

[0314] As used herein, the term "Tn7" or "Tn7-like transposase" refers to a family of transposases that comprises three main components: the heteromultimeric transposases (TnsA and / or TnsB) and the regulatory protein (TnsC). In addition to the TnsABC transposase proteins, the Tn7 element can encode specialized target site selection proteins, TnsD and TnsE. In association with TnsABC, the sequence-specific DNA-binding protein TnsD directs transposition to a conserved site known as the "Tn7 attachment site", i.e., attTn7. TnsD is a member of a large family of proteins that also includes TniQ. TniQ has been shown to target transposition into the resolution site of plasmids.

[0315] As used herein, the terms "gene editing" and "genome editing" may be used interchangeably. Gene editing or genome editing refers to altering the nucleic acid sequence of a gene or genome. Genome editing can include, for example, insertions, deletions, and mutations. Genome editing can be performed by a gene editing system (e.g., a nuclease, reverse transcriptase, recombinase, or base editor).

[0316] As used herein, the term "recombinase" refers to an enzyme that mediates recombination of DNA fragments located between recombinase recognition sequences, resulting in excision, insertion, inversion, exchange, or translocation of the DNA fragments located between the recombinase recognition sequences.

[0317] As used herein, in the context of nucleic acid modification (e.g., genome modification), the term "recombine or recombination" refers to the process of modifying two or more nucleic acid molecules or two or more regions of a single nucleic acid molecule through the action of a recombinase protein. Recombination may particularly result in insertion, inversion, excision, or translocation of nucleic acid sequences, e.g., within or between one or more nucleic acid molecules.

[0318] As used herein, the term "complex" refers to the association of at least two components. The two components may each retain the properties / activities they had prior to forming the complex or acquire properties as a result of forming the complex. The association includes, but is not limited to, covalent bonding, non-covalent bonding (i.e., hydrogen bonding, ionic interactions, van der Waals interactions, and hydrophobic bonds), use of a linker, fusion, or any other suitable method. The components envisioned for the complex include polynucleotides, polypeptides, or combinations thereof. For example, the complex comprises an endonuclease and a guide polynucleotide.

[0319] In the context of two or more nucleic acid or polypeptide sequences, the terms "sequence identity" or "percent identity" refer to sequences that, when compared and aligned within a local or global comparison window to obtain maximum correspondence, have two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) identical or a specified percentage of identical amino acid residues or nucleotides, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example: BLASTP for polypeptide sequences longer than 30 residues, using parameters with a word length (W) of 3, an expectation value (E) of 10, and a BLOSUM62 scoring matrix, setting the gap penalty to 11 for existence, 1 for extension, and using a conditional compositional scoring matrix for adjustment; BLASTP for sequences shorter than 30 residues, using parameters with a word length (W) of 2, an expectation value (E) of 1000000, and a PAM30 scoring matrix, setting the gap penalty to 9 for gap opening and 1 for gap extension (these are the default parameters for BLASTP in the BLAST suite, available at https: / / blast.ncbi.nlm.nih.gov); CLUSTALW using Smith-Waterman homology search algorithm parameters with a match of 2, a mismatch of -1, and a gap of -1; MUSCLE using default parameters; MAFFT using parameters with retree of 2 and a maximum iteration of 1000; Novafold using default parameters; HMMER hmmalign using default parameters.

[0320] In the context of two or more nucleic acid or polypeptide sequences, the term "best alignment" refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences that have been aligned with maximum correspondence of amino acid residues or nucleotides, e.g., as determined by an alignment that produces the highest or "optimal" percent identity score.

[0321] The present disclosure includes variants of any of the enzymes described herein having one or more conservative amino acid substitutions. Such conservative substitutions can be made in the amino acid sequence of the polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be accomplished by substituting amino acids having similar hydrophobicity, polarity, and R-chain length for one another. Additionally or alternatively, by comparing the aligned sequences of homologous proteins from different species, conservative substitutions can be identified by locating amino acid residues that have mutated between species (e.g., non-conservative residues) without altering the basic function of the encoded protein. Variants of such conservative substitutions include those having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any of the reverse transcriptase protein sequences described herein (e.g., the MG140, MG146, MG148, MG149, MG151, MG153, MG154, MG155, MG156, MG157, MG158, MG159, MG160, MG163, MG164, MG165, MG166, MG167, MG168, MG169, MG170, MG172, MG173, and MG176 family reverse transcriptases or retrotransposases described herein, or any other family reverse transcriptases or retrotransposases described herein). In some embodiments, variants of such conservative substitutions are functional variants. Such functional variants can encompass sequences having substitutions such that the activity of one or more key active site residues is not disrupted.

[0322] The present disclosure also includes variants of any of the enzymes described herein that substitute one or more catalytic residues to reduce or eliminate the activity of the enzyme (e.g., variants with reduced activity). In some embodiments, a variant with reduced activity of a protein described herein comprises a disruptive substitution of at least one, at least two, or all three catalytic residues (e.g., the programmable nuclease MG3 family nickase having a D13A mutation, an H586A mutation, or an N609A mutation).

[0323] Conservative substitution tables providing functionally similar amino acids are available from a variety of references (see, e.g., Creighton, Proteins: Structures and Molecular Properties (W H Freeman &(Co.); 2nd Edition (December 1993)). The following eight groups each contain amino acids that are conservative substitutions of each other:

[0324] 1) Alanine (A), Glycine (G);

[0325] 2) Aspartic acid (D), Glutamic acid (E);

[0326] 3) Asparagine (N), Glutamine (Q);

[0327] 4) Arginine (R), Lysine (K);

[0328] 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V);

[0329] 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W);

[0330] 7) Serine (S), Threonine (T); and

[0331] 8) Cysteine (C), Methionine (M)

[0332] Gene Editing System

[0333] This disclosure describes gene editing systems that include: a) a nickase; b) a guide nucleic acid (e.g., a pegRNA or other guide RNA) that is configured to form a complex with the nickase and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase that has at least about 80% sequence identity to any of SEQ ID NOs: 161 - 629, 767 - 1220, 1959 - 2522, and 2582 - 2585 and is configured to form a complex with the nickase. This disclosure further describes gene editing systems that include: a) a nuclease; b) a guide nucleic acid (e.g., a pegRNA or other guide RNA) that is configured to form a complex with the nuclease and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase that has at least about 80% sequence identity to any of SEQ ID NOs: 161 - 629, 767 - 1220, 1959 - 2522, and 2582 - 2585 and is configured to form a complex with the nuclease. This disclosure further describes gene editing systems that include: a) a nickase; b) a guide nucleic acid (e.g., a pegRNA) that is configured to form a complex with the nickase and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase that is configured to form a complex with the nickase and has an X1X2DD motif, where X1 is F or Y, and where when X1 is Y, X2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y. This disclosure further describes gene editing systems that include: a) a nuclease; b) a guide nucleic acid (e.g., a pegRNA) that is configured to form a complex with the nuclease and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase that is configured to form a complex with the nuclease and has an X1X2DD motif, where X1 is F or Y, and where when X1 is Y, X2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y.

[0334] The gene editing systems described herein, in some embodiments, include a nickase, a nuclease, a reverse transcriptase, or a combination thereof, and are capable of introducing site - specific insertions, deletions, and mutations. In some embodiments, the nickase, nuclease, reverse transcriptase, or combination thereof is capable of integrating polynucleotides of a relatively large size. In some embodiments, the integrated polynucleotide has a size of at least about 1 kilobase (kb), 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, or greater than 10 kb.

[0335] Reverse Transcriptase

[0336] Reverse transcription is the translation of an RNA template into complementary DNA. Reverse transcription is carried out by an enzyme called reverse transcriptase (RT), which is an enzyme that produces a complementary DNA (cDNA) strand from an RNA template with RNA-dependent DNA polymerase activity. Some RT enzymes also have DNA-dependent DNA polymerase activity to produce double-stranded dsDNA. Reverse transcriptases can be of viral origin (e.g., HIV, hepatitis B, Moloney murine leukemia virus (MMLV), or avian myeloblastosis virus (AMV)) or bacterial origin (e.g., group II introns, retrotransposons / retrotransposon-like RTs, diversity-generating retroelements (DGRs), Abi-like RTs, CRISPR-associated RTs, and group II-like RTs (G2Ls)). Eukaryotic-derived reverse transcriptases include telomerase reverse transcriptase, which maintains the telomeres of eukaryotic chromosomes. Reverse transcription allows site-directed insertions, deletions, and mutations to be introduced into cDNA by encoding them in the RNA template.

[0337] In some embodiments, the reverse transcriptase is a viral, prokaryotic or eukaryotic reverse transcriptase. In some embodiments, the reverse transcriptase comprises the sequences of SEQ ID NO: 161-629, 767-1220, 1959-2522 and 2582-2585, variants thereof or functional fragments thereof. In some embodiments, the reverse transcriptase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity with any one of SEQ ID NO: 161-629, 767-1220, 1959-2522 and 2582-2585, variants thereof or functional fragments thereof. In some embodiments, the reverse transcriptase comprises a sequence having at least about 70% identity with any one of SEQ ID NO: 161-629, 767-1220, 1959-2522 and 2582-2585. In some embodiments, the reverse transcriptase comprises a sequence having at least about 75% identity with any one of SEQ ID NO: 161-629, 767-1220, 1959-2522 and 2582-2585. In some embodiments, the reverse transcriptase comprises a sequence having at least about 80% identity with any one of SEQ ID NO: 161-629, 767-1220, 1959-2522 and 2582-2585. In some embodiments, the reverse transcriptase comprises a sequence having at least about 85% identity with any one of SEQ ID NO: 161-629, 767-1220, 1959-2522 and 2582-2585. In some embodiments, the reverse transcriptase comprises a sequence having at least about 90% identity with any one of SEQ ID NO: 161-629, 767-1220, 1959-2522 and 2582-2585. In some embodiments, the reverse transcriptase comprises a sequence having at least about 95% identity with any one of SEQ ID NO: 161-629, 767-1220, 1959-2522 and 2582-2585. In some embodiments, the reverse transcriptase comprises a sequence having at least about 96% identity with any one of SEQ ID NO: 161-629, 767-1220, 1959-2522 and 2582-2585. In some embodiments, the reverse transcriptase comprises a sequence having at least about 97% identity with any one of SEQ ID NO: 161-629, 767-1220, 1959-2522 and 2582-2585.In some embodiments, the reverse transcriptase comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 161-629, 767-1220, 1959-2522, and 2582-2585. In some embodiments, the reverse transcriptase comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 161-629, 767-1220, 1959-2522, and 2582-2585. In some embodiments, the reverse transcriptase comprises a sequence having 100% identity to any one of SEQ ID NOs: 161-629, 767-1220, 1959-2522, and 2582-2585.

[0338] In some embodiments, the reverse transcriptase is a reverse transcriptase of the MG151, MG153, or MG160 family. In some embodiments, the reverse transcriptase is a reverse transcriptase of the MG140, MG146, MG148, MG149, MG151, MG153, MG154, MG155, MG156, MG157, MG158, MG159, MG160, MG163, MG164, MG165, MG166, MG167, MG168, MG169, MG170, or MG176 family. In some embodiments, the reverse transcriptase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of a reverse transcriptase or retrotransposase of the MG140, MG146, MG148, MG149, MG151, MG153, MG154, MG155, MG156, MG157, MG158, MG159, MG160, MG163, MG164, MG165, MG166, MG167, MG168, MG169, MG170, MG172, MG173, or MG176 family. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of a reverse transcriptase or retrotransposase of the MG140, MG146, MG148, MG149, MG151, MG153, MG154, MG155, MG156, MG157, MG158, MG159, MG160, MG163, MG164, MG165, MG166, MG167, MG168, MG169, MG170, MG172, MG173, or MG176 family or a variant thereof.

[0339] In some embodiments, the reverse transcriptase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1-75, 702-766, 1221-1243, 1299, 1249-1295, 1300-1304, 1309, 1394-1447, 1592-1593, 1596-1597, 1654, 1691-1720, 1753-1754, 1776-1778, 1780-1783, 1790-1847, and 1856-1858. In some embodiments, the reverse transcriptase is encoded by a nucleic acid sequence having at least 85% sequence identity to any one of SEQ ID NOs: 1-75, 702-766, 1221-1243, 1299, 1249-1295, 1300-1304, 1309, 1394-1447, 1592-1593, 1596-1597, 1654, 1691-1720, 1753-1754, 1776-1778, 1780-1783, 1790-1847, and 1856-1858. In some embodiments, the reverse transcriptase is encoded by a nucleic acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-75, 702-766, 1221-1243, 1299, 1249-1295, 1300-1304, 1309, 1394-1447, 1592-1593, 1596-1597, 1654, 1691-1720, 1753-1754, 1776-1778, 1780-1783, 1790-1847, and 1856-1858. In some embodiments, the reverse transcriptase is encoded by a nucleic acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 1-75, 702-766, 1221-1243, 1299, 1249-1295, 1300-1304, 1309, 1394-1447, 1592-1593, 1596-1597, 1654, 1691-1720, 1753-1754, 1776-1778, 1780-1783, 1790-1847, and 1856-1858.In some embodiments, the reverse transcriptase is encoded by a nucleic acid sequence having at least 96% sequence identity to any one of SEQ ID NO: 1-75, 702-766, 1221-1243, 1299, 1249-1295, 1300-1304, 1309, 1394-1447, 1592-1593, 1596-1597, 1654, 1691-1720, 1753-1754, 1776-1778, 1780-1783, 1790-1847, and 1856-1858. In some embodiments, the reverse transcriptase is encoded by a nucleic acid sequence having at least 97% sequence identity to any one of SEQ ID NO: 1-75, 702-766, 1221-1243, 1299, 1249-1295, 1300-1304, 1309, 1394-1447, 1592-1593, 1596-1597, 1654, 1691-1720, 1753-1754, 1776-1778, 1780-1783, 1790-1847, and 1856-1858. In some embodiments, the reverse transcriptase is encoded by a nucleic acid sequence having at least 98% sequence identity to any one of SEQ ID NO: 1-75, 702-766, 1221-1243, 1299, 1249-1295, 1300-1304, 1309, 1394-1447, 1592-1593, 1596-1597, 1654, 1691-1720, 1753-1754, 1776-1778, 1780-1783, 1790-1847, and 1856-1858. In some embodiments, the reverse transcriptase is encoded by a nucleic acid sequence having at least 99% sequence identity to any one of SEQ ID NO: 1-75, 702-766, 1221-1243, 1299, 1249-1295, 1300-1304, 1309, 1394-1447, 1592-1593, 1596-1597, 1654, 1691-1720, 1753-1754, 1776-1778, 1780-1783, 1790-1847, and 1856-1858. In some embodiments, the reverse transcriptase is encoded by any one of the nucleic acid sequences of SEQ ID NO: 1-75, 702-766, 1221-1243, 1299, 1249-1295, 1300-1304, 1309, 1394-1447, 1592-1593, 1596-1597, 1654, 1691-1720, 1753-1754, 1776-1778, 1780-1783, 1790-1847, and 1856-1858.

[0340] Reverse transcriptases typically have an active site core tetrad motif with the amino acid sequence XXDD. In some embodiments, the reverse transcriptase has an active site tetrad motif of X1X2DD, where X1 is F or Y, and where when X1 is Y, X2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y, and in some embodiments, X2 is A or I. In some embodiments, the X1X2DD motif is YADD (SEQ ID NO:2572) or YIDD (SEQ ID NO:2573). In some embodiments, the X1X2DD motif is FADD (SEQ ID NO:2574), FVDD (SEQ ID NO:2575), FIDD (SEQ ID NO:2576), or FLDD (SEQ ID NO:2577). In some embodiments, the reverse transcriptase is isolated. In some embodiments, the reverse transcriptase is a reverse transcriptase or retrotransposase of the MG140, MG146, MG148, MG149, MG151, MG153, MG154, MG155, MG156, MG157, MG158, MG159, MG160, MG163, MG164, MG165, MG166, MG167, MG168, MG169, MG170, MG172, MG173, or MG176 family, and the X1X2DD motif is YADD (SEQ ID NO:2572) or YIDD (SEQ ID NO:2573). In some embodiments, the reverse transcriptase is isolated. In some embodiments, the reverse transcriptase is a reverse transcriptase or retrotransposase of the MG140, MG146, MG148, MG149, MG151, MG153, MG154, MG155, MG156, MG157, MG158, MG159, MG160, MG163, MG164, MG165, MG166, MG167, MG168, MG169, MG170, MG172, MG173, or MG176 family, and the X1X2DD motif is FADD (SEQ ID NO:2574), FVDD (SEQ ID NO:2575), FIDD (SEQ ID NO:2576), or FLDD (SEQ ID NO:2577).

[0341] In some embodiments, the reverse transcriptase is less than 300 amino acids. In some embodiments, the reverse transcriptase is less than 250 amino acids. In some embodiments, the reverse transcriptase comprises at least about 50, 75, 100, 125, 150, 175, 200, 225, 250, 275, 300, or more than 300 amino acids. In some embodiments, the reverse transcriptase comprises a series of about 50 to about 300, about 75 to about 300, about 100 to about 300, about 125 to about 300, about 150 to about 300, about 175 to about 300, about 200 to about 300, about 225 to about 300, about 250 to about 300, about 275 to about 300, about 100 to about 300, about 125 to about 300, about 150 to about 300, about 175 to about 300, about 200 to about 300, about 225 to about 300, about 250 to about 300, or about 275 to about 300 amino acids.

[0342] In some embodiments, the reverse transcriptase comprises a processive synthesis ability that is at least about 2-fold that of Moloney murine leukemia virus (MMLV) reverse transcriptase. In some embodiments, the reverse transcriptase comprises a processive synthesis ability that is at most about 1 / 2 that of Moloney murine leukemia virus (MMLV) reverse transcriptase. In some embodiments, the reverse transcriptase comprises an error rate of less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10%, or 0.05%. In some embodiments, compared to Moloney murine leukemia virus (MMLV) reverse transcriptase, the reverse transcriptase comprises an error rate of less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10%, or 0.05%. Methods for measuring the processive synthesis ability of reverse transcriptase are known in the art or are described herein, such as in Example 2.

[0343] In some embodiments, the reverse transcriptase is targetable. A targetable reverse transcriptase is an engineered ribonucleoprotein complex that serves as a tool for genome editing in cells and organisms. In some embodiments, a targetable reverse transcriptase is produced by fusing a reverse transcriptase and a site-directed CRISPR nuclease variant that cleaves the non-target strand of dsDNA such that a guide RNA or pegRNA containing a primer binding site (PBS) sequence can find its complementary target sequence and hybridize thereto to initiate a reverse transcriptase reaction using a reverse transcriptase template (RTT) as a template. Two DNA flaps are generated, one containing the desired change encoded in the RTT and the other containing the original sequence; after equilibration, when the DNA flap with the desired edit is repaired by the cellular host repair mechanism, the change is integrated into the genomic DNA.

[0344] In some embodiments, the gene editing system comprises a reverse transcriptase and a nickase as described herein. In some embodiments, the gene editing system comprises a reverse transcriptase and a nuclease as described herein. In some embodiments, the gene editing system comprises a reverse transcriptase and a modified nuclease as described herein. In some embodiments, the gene editing system is programmable. In some embodiments, the modified nuclease is a site-specific nickase.

[0345] In some embodiments, the reverse transcriptase and the nuclease or nickase are linked or tethered. In some embodiments, the gene editing system comprises a fusion protein of a reverse transcriptase and a nuclease or nickase. In some embodiments, the gene editing system comprises a fusion protein comprising a nickase linked to a reverse transcriptase using a linker, wherein the reverse transcriptase has at least about 80% sequence identity with any one of SEQ ID NOs: 161-629, 767-1220, 1959-2522, and 2582-2585. In some embodiments, the gene editing system comprises a fusion protein comprising a nuclease linked to a reverse transcriptase using a linker, wherein the reverse transcriptase has at least about 80% sequence identity with any one of SEQ ID NOs: 161-629, 767-1220, 1959-2522, and 2582-2585. In some embodiments, the gene editing system comprises a fusion protein comprising a catalytically dead nuclease linked to a reverse transcriptase using a linker, wherein the reverse transcriptase has at least about 80% sequence identity with any one of SEQ ID NOs: 161-629, 767-1220, 1959-2522, and 2582-2585.

[0346] In some embodiments, the reverse transcriptase and the nuclease or nickase are joined or fused using a linker. In some embodiments, the linker comprises at least 10, 20, or 30 amino acids. In some embodiments, the linker comprises about 30-35 amino acids. In some embodiments, the linker comprises about 30 amino acids.

[0347] In some embodiments, the linker has at least 80% sequence identity with SEQ ID NO:103. In some embodiments, the linker has at least 80% sequence identity with SEQ ID NO:103. In some embodiments, the linker comprises a sequence having at least about 85% identity with SEQ ID NO:103. In some embodiments, the linker comprises a sequence having at least about 90% identity with SEQ ID NO:103. In some embodiments, the linker comprises a sequence having at least about 91% identity with SEQ ID NO:103. In some embodiments, the linker comprises a sequence having at least about 92% identity with SEQ ID NO:103. In some embodiments, the linker comprises a sequence having at least about 93% identity with SEQ ID NO:103. In some embodiments, the linker comprises a sequence having at least about 94% identity with SEQ ID NO:103. In some embodiments, the linker comprises a sequence having at least about 95% identity with SEQ ID NO:103. In some embodiments, the linker comprises a sequence having at least about 96% identity with SEQ ID NO:103. In some embodiments, the linker comprises a sequence having at least about 97% identity with SEQ ID NO:103. In some embodiments, the linker comprises a sequence having at least about 98% identity with SEQ ID NO:103. In some embodiments, the linker comprises a sequence having at least about 99% identity with SEQ ID NO:103. In some embodiments, the linker comprises a sequence having 100% identity with SEQ ID NO:103.

[0348] Suitable linkers are known in the art and include, for example, any of SEQ ID NOs:155 - 160. In some embodiments, the linker has at least 80% sequence identity with any of SEQ ID NOs:155 - 160. In some embodiments, the linker that connects any enzyme or domain described herein comprises one or more copies of a sequence that is

[0349] SGGSSGGSSGSETPGTSESATPESSGGSSGGSSAC (SEQ ID NO:155), KLGGGAPAVGGGPK (SEQ ID NO:156), (GGGGS)3 (SEQ ID NO:157), (GGGGS)2EAAAK(GGGGS)2 (SEQ ID NO:158),

[0350] (GGGGS)2(EAAAK)2(GGGGS)2 (SEQ ID NO:159), or SGSETPGTSESATPES (SEQ ID NO:160), or any other linker sequence described herein has at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity. In some embodiments, the linker comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 155-160. In some embodiments, the linker comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 155-160. In some embodiments, the linker comprises a sequence having at least about 91% identity to any one of SEQ ID NOs: 155-160. In some embodiments, the linker comprises a sequence having at least about 92% identity to any one of SEQ ID NOs: 155-160. In some embodiments, the linker comprises a sequence having at least about 93% identity to any one of SEQ ID NOs: 155-160. In some embodiments, the linker comprises a sequence having at least about 94% identity to any one of SEQ ID NOs: 155-160. In some embodiments, the linker comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 155-160. In some embodiments, the linker comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 155-160. In some embodiments, the linker comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 155-160. In some embodiments, the linker comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 155-160. In some embodiments, the linker comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 155-160. In some embodiments, the linker comprises a sequence having 100% identity to any one of SEQ ID NOs: 155-160.

[0351] In some embodiments, the nickase or nuclease and the reverse transcriptase are not linked.

[0352] In some embodiments, the reverse transcriptase, nuclease, nickase, or fusion protein described herein comprises one or more nuclear localization sequences (NLSs) near the N-terminus or C-terminus of the reverse transcriptase, nuclease, nickase, or fusion protein.

[0353] In some embodiments, the NLS comprises any one of the sequences in Table 1 below or a combination thereof:

[0354] Table 1: Exemplary NLS Sequences

[0355]

[0356]

[0357] In some embodiments, the reverse transcriptase comprises a tag. In some embodiments, the nuclease comprises a tag. In some embodiments, the nickase comprises a tag. In some embodiments, the fusion protein comprises a tag. In some embodiments, the tag is an affinity tag. Exemplary affinity tags include, but are not limited to, His tag, Flag tag, Myc tag, MBP tag, and GST tag.

[0358] In some embodiments, the reverse transcriptase comprises a protease cleavage site. In some embodiments, the nuclease comprises a protease cleavage site. In some embodiments, the nickase comprises a protease cleavage site. In some embodiments, the fusion protein comprises a protease cleavage site. Exemplary protease cleavage sites include, but are not limited to, TEV site, C3 site, Factor Xa site, and enterokinase site.

[0359] In some embodiments, the gene editing system comprises: a) a nickase; b) a guide nucleic acid (e.g., pegRNA or other guide RNA); and c) a reverse transcriptase having at least about 80% sequence identity to any one of SEQ ID NOs: 161 - 629, 767 - 1220, 1959 - 2522, and 2582 - 2585.

[0360] In some embodiments, the gene editing system comprises: a) a nuclease; b) a guide nucleic acid (e.g., pegRNA or other guide RNA); and c) a reverse transcriptase having at least about 80% sequence identity to any one of SEQ ID NOs: 161 - 629, 767 - 1220, 1959 - 2522, and 2582 - 2585.

[0361] In some embodiments, the gene editing system comprises: a) a nickase; b) a guide nucleic acid (e.g., pegRNA); and c) a reverse transcriptase having an X1X2DD motif, wherein X1 is F or Y, and wherein when X1 is Y, X2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y.

[0362] In some embodiments, the gene editing system comprises: a) a nuclease; b) a guide nucleic acid (e.g., pegRNA); and c) a reverse transcriptase having an X1X2DD motif, where X1 is F or Y, and where when X1 is Y, X2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y. In some embodiments, the X2 is A or I. In some embodiments, the X1X2DD motif is YADD (SEQ ID NO: 2572) or YIDD (SEQ ID NO: 2573). In some embodiments, the X1X2DD motif is FADD (SEQ ID NO: 2574), FVDD (SEQ ID NO: 2575), FIDD (SEQ ID NO: 2576), or FLDD (SEQ ID NO: 2577). In some embodiments, the reverse transcriptase has at least about 80% sequence identity with any one of SEQ ID NOs: 161 - 629, 767 - 1220, 1959 - 2522, and 2582 - 2585.

[0363] In some embodiments, the nuclease is configured to cleave one strand of a double-stranded target deoxyribonucleic acid (nickase). In some embodiments, the nickase or nuclease is a CRISPR nuclease as described herein. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NO: 104 and 1859 - 1862 or a variant thereof. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 70% identity to any one of SEQ ID NO: 104 and 1859 - 1862. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 75% identity to any one of SEQ ID NO: 104 and 1859 - 1862. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 80% identity to any one of SEQ ID NO: 104 and 1859 - 1862. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 85% identity to any one of SEQ ID NO: 104 and 1859 - 1862. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 90% identity to any one of SEQ ID NO: 104 and 1859 - 1862. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 95% identity to any one of SEQ ID NO: 104 and 1859 - 1862. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 96% identity to any one of SEQ ID NO: 104 and 1859 - 1862. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 97% identity to any one of SEQ ID NO: 104 and 1859 - 1862. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 98% identity to any one of SEQ ID NO: 104 and 1859 - 1862. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 99% identity to any one of SEQ ID NO: 104 and 1859 - 1862. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having 100% identity to any one of SEQ ID NO: 104 and 1859 - 1862.

[0364] In some embodiments, the system further comprises a source of Mg 2+ .

[0365] In some embodiments, the nuclease is a modified endonuclease. In some embodiments, the modified endonuclease is a type II CRISPR endonuclease or a type V CRISPR endonuclease. In some embodiments, the type II or type V CRISPR endonuclease comprises double-stranded cleavage activity, nickase activity, or can be catalytically dead. In some embodiments, the CRISPR nuclease has a modification in the HNH domain or the RuvC domain.

[0366] In some embodiments, the modified endonuclease comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NOs: 152 - 154 or variants thereof. In some embodiments, the modified endonuclease comprises at least about 80% sequence identity to any one of SEQ ID NOs: 152 - 154. In some embodiments, the modified endonuclease comprises a sequence having at least about 70% identity to any one of SEQ ID NOs: 152 - 154. In some embodiments, the modified endonuclease comprises a sequence having at least about 75% identity to any one of SEQ ID NOs: 152 - 154. In some embodiments, the modified endonuclease comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 152 - 154. In some embodiments, the modified endonuclease comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 152 - 154. In some embodiments, the modified endonuclease comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 152 - 154. In some embodiments, the modified endonuclease comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 152 - 154. In some embodiments, the modified endonuclease comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 152 - 154. In some embodiments, the modified endonuclease comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 152 - 154. In some embodiments, the modified endonuclease comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 152 - 154. In some embodiments, the modified endonuclease comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 152 - 154. In some embodiments, the modified endonuclease comprises a sequence having 100% identity to any one of SEQ ID NOs: 152 - 154.

[0367] In some embodiments, the modified endonuclease is selected from the group consisting of: spCas9(H840A), spCas9(D10A), nMG3-6(D13A), nMG3-6(H586A), nMG3-6(N609A), Cas12a, and MG29-1.

[0368] In some embodiments, the gene editing system comprises a nucleic acid template. The nucleic acid template can be RNA or DNA. The nucleic acid template can be 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 bases in length. The nucleic acid template can be 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 bases in length. In some embodiments, the nucleic acid template has a homology region homologous to a site in the genome. In some embodiments, the homology region is 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 bases in length.

[0369] In some embodiments, the gene editing system further comprises a transposase, integrase, or homing endonuclease. In some embodiments, the transposase is the transposase (Tnp) Tn5, Sleeping Beauty transposase, or Tn7 transposon. In some embodiments, the gene editing system comprises an enzyme having transposase activity. Additional enzymes having transposase activity include, but are not limited to, reverse transcriptase and IS200 / IS605 transposons.

[0370] In some embodiments, the gene editing system further comprises a retrotransposon of the present disclosure. In some embodiments, the retrotransposon is a MG140, MG146, or MG176 family retrotransposon. In some embodiments, the retrotransposon comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NOs: 161-629, 767-1220, 1959-2522, and 2582-2585 or variants thereof.

[0371] CRISPR Nuclease

[0372] In some embodiments, a nickase or an endonuclease is described herein, wherein the nickase or endonuclease is a CRISPR nuclease. In some embodiments, the CRISPR nuclease is a modified nuclease.

[0373] The CRISPR system is an RNA-guided nuclease complex that has been described as acting as an adaptive immune system in microorganisms. In its natural environment, the CRISPR system occurs in the CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) operon or locus, and the system typically consists of two parts: (i) an array of short repeat sequences (30 - 40 bp) separated by similarly short spacer sequences, which encode RNA-based targeting elements; and (ii) an ORF encoding a nuclease polypeptide guided by the RNA-based targeting element and accessory proteins / enzymes. Efficient nuclease targeting of a specific target nucleic acid sequence generally requires both: (i) complementary hybridization between the first 6 - 8 nucleic acids of the target (target seed) and the crRNA guide; and (ii) the presence of a protospacer adjacent motif (PAM) sequence within a defined vicinity of the target seed (PAM is generally a sequence not commonly represented within the host genome). Depending on the exact function and organization of the system, CRISPR systems are generally classified into 2 classes, 5 types, and 16 subtypes based on shared functional characteristics and evolutionary similarities.

[0374] Class 1 CRISPR systems have large multi-subunit effector complexes and include types I, III, and IV. Class 2 CRISPR systems typically have single polypeptide multi-domain nuclease effectors and include types II, V, and VI.

[0375] Type II CRISPR systems are considered the simplest in terms of components. In Type II CRISPR systems, processing of the CRISPR array into mature crRNAs does not require the presence of a special endonuclease subunit, but rather a small trans-encoded crRNA (tracrRNA), the region of which is complementary to the array repeat sequences; the tracrRNA interacts with its corresponding effector nuclease (e.g., Cas9) and the repeat sequences to form a precursor dsRNA structure, which is cleaved by the endogenous RNase III, thereby generating a mature effector enzyme loaded with both tracrRNA and crRNA. Type II nucleases are referred to as DNA nucleases. Type II nucleases typically exhibit a structure consisting of an RuvC-like endonuclease domain that adopts an RNase H fold, into which an unrelated HNH nuclease domain is inserted within the fold of the RuvC-like nuclease domain. The RuvC-like domain is responsible for cleavage of the target (e.g., crRNA-complementary) DNA strand, while the HNH domain is responsible for cleavage of the displaced DNA strand. Exemplary CRISPR Cas9 proteins include, but are not limited to, Cas9 from Streptococcus pyogenes (UniProtKB-Q99ZW2 (CAS9 STRP1)), Streptococcus thermophilus (UniProtKB-G3ECR1 (CAS9 STRTR)), Staphylococcus aureus (UniProtKB-J7RUA5 (CAS9 STAAU)), Campylobacter jejuni (UniProtKB-Q0P897 (CAS9CAMJE)), Campylobacter lari (UniProtKB-A0A0A8HTA3 (A0A0A8HTA3 CAMLA)), and Helicobacter canadensis (UniProtKB-C5ZYI3 (C5ZYI3 9HELI)), Francisella tularensis subsp. novicida (UniProtKB-A0Q5Y3 (CAS9_FRATN)). Additional Type II nucleases are described in International Patent Application Publications WO 2021 / 226363, WO 2022 / 159758, and WO 2022 / 056324.

[0376] Type V CRISPR systems are characterized by nuclease effectors (e.g., Cas12) that have a structure similar to that of type II effectors containing an RuvC-like domain. Similar to type II, most (but not all) type V CRISPR systems use a tracrRNA to process pre-crRNA into mature crRNA; however, unlike type II systems that require RNase III to cleave pre-crRNA into multiple crRNAs, type V systems are capable of using the effector nuclease itself to cleave pre-crRNA. Like type II CRISPR systems, type V CRISPR systems are referred to as DNA nucleases. Different from type II CRISPR systems, some type V enzymes (e.g., Cas12a) appear to have robust single-stranded non-specific deoxyribonuclease activity that is activated by the first crRNA-directed cleavage of a double-stranded target sequence.

[0377] In some embodiments, the nuclease or nickase is a CRISPR nuclease. In some embodiments, the CRISPR nuclease is a class 2 type II SpCas9 or a class 2 type V-A Cas12a (formerly Cpf1). In some embodiments, the type V-A nuclease has a guide RNA of 42-44 nucleotides, in contrast to SpCas9 which has approximately 100 nt. In some embodiments, the type V-A nuclease produces staggered cleavage sites. In some embodiments, the type V-A nuclease produces staggered cleavage sites to facilitate a directed repair pathway such as microhomology-dependent targeted integration (MITI).

[0378] The most commonly used type V-A enzymes require a 5' protospacer adjacent motif (PAM) next to the selected target site: 5'-TTTV-3' for Lachnospiraceae bacterium ND2006 LbCas12a and Acidaminococcus AsCas12a; and 5'-TTV-3' for Francisella novicida FnCas12a. In some embodiments, the PAM sequence is YTV, YYN or TTN. Additional type II nucleases are described in International Patent Application Publication WO 2021 / 226363.

[0379] In some embodiments, the nickase is a modified nuclease. In some embodiments, the modified endonuclease is a type II CRISPR endonuclease. In some embodiments, the modified endonuclease is a type II CRISPR endonuclease or a type V endonuclease. In some embodiments, the type II CRISPR endonuclease or the type V endonuclease has nickase activity.

[0380] In some embodiments, the modified endonuclease is selected from the group consisting of: spCas9(H840A), spCas9(D10A), nMG3-6(D13A), nMG3-6(H586A), nMG3-6(N609A), Cas12a, and MG29-1. In some embodiments, the modified endonuclease comprises at least about 80% sequence identity with any one of SEQ ID NOs: 152-154. In some embodiments, the nuclease comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 152-154 or a variant thereof. In some embodiments, the modified endonuclease comprises a sequence having at least about 70% identity with any one of SEQ ID NOs: 152-154. In some embodiments, the modified endonuclease comprises a sequence having at least about 75% identity with any one of SEQ ID NOs: 152-154. In some embodiments, the modified endonuclease comprises a sequence having at least about 80% identity with any one of SEQ ID NOs: 152-154. In some embodiments, the modified endonuclease comprises a sequence having at least about 85% identity with any one of SEQ ID NOs: 152-154. In some embodiments, the modified endonuclease comprises a sequence having at least about 90% identity with any one of SEQ ID NOs: 152-154. In some embodiments, the modified endonuclease comprises a sequence having at least about 95% identity with any one of SEQ ID NOs: 152-154. In some embodiments, the modified endonuclease comprises a sequence having at least about 96% identity with any one of SEQ ID NOs: 152-154. In some embodiments, the modified endonuclease comprises a sequence having at least about 97% identity with any one of SEQ ID NOs: 152-154. In some embodiments, the modified endonuclease comprises a sequence having at least about 98% identity with any one of SEQ ID NOs: 152-154. In some embodiments, the modified endonuclease comprises a sequence having at least about 99% identity with any one of SEQ ID NOs: 152-154. In some embodiments, the modified endonuclease comprises a sequence having 100% identity with any one of SEQ ID NOs: 152-154.

[0381] In some embodiments, the nuclease comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity with SEQ ID NO:646 or SEQ ID NO:647 or a variant thereof. In some embodiments, the nuclease comprises a sequence having at least about 70% identity with SEQ ID NO:646 or SEQ ID NO:647. In some embodiments, the nuclease comprises a sequence having at least about 75% identity with SEQ ID NO:646 or SEQ ID NO:647. In some embodiments, the nuclease comprises a sequence having at least about 80% identity with SEQ ID NO:646 or SEQ ID NO:647. In some embodiments, the nuclease comprises a sequence having at least about 85% identity with SEQ ID NO:646 or SEQ ID NO:647. In some embodiments, the nuclease comprises a sequence having at least about 90% identity with SEQ ID NO:646 or SEQ ID NO:647. In some embodiments, the nuclease comprises a sequence having at least about 95% identity with SEQ ID NO:646 or SEQ ID NO:647. In some embodiments, the nuclease comprises a sequence having at least about 96% identity with SEQ ID NO:646 or SEQ ID NO:647. In some embodiments, the nuclease comprises a sequence having at least about 97% identity with SEQ ID NO:646 or SEQ ID NO:647. In some embodiments, the nuclease comprises a sequence having at least about 98% identity with SEQ ID NO:646 or SEQ ID NO:647. In some embodiments, the nuclease comprises a sequence having at least about 99% identity with SEQ ID NO:646 or SEQ ID NO:647. In some embodiments, the nuclease comprises a sequence having 100% identity with SEQ ID NO:646 or SEQ ID NO:647.

[0382] In some embodiments, the nuclease is encoded by a nucleic acid sequence having at least 80% sequence identity with the nucleic acid sequence of SEQ ID NO: 653. In some embodiments, the nuclease is encoded by a nucleic acid sequence having at least 85% sequence identity with the nucleic acid sequence of SEQ ID NO: 653. In some embodiments, the nuclease is encoded by a nucleic acid sequence having at least 90% sequence identity with the nucleic acid sequence of SEQ ID NO: 653. In some embodiments, the nuclease is encoded by a nucleic acid sequence having at least 95% sequence identity with the nucleic acid sequence of SEQ ID NO: 653. In some embodiments, the nuclease is encoded by a nucleic acid sequence having at least 96% sequence identity with the nucleic acid sequence of SEQ ID NO: 653. In some embodiments, the nuclease is encoded by a nucleic acid sequence having at least 97% sequence identity with the nucleic acid sequence of SEQ ID NO: 653. In some embodiments, the nuclease is encoded by a nucleic acid sequence having at least 98% sequence identity with the nucleic acid sequence of SEQ ID NO: 653. In some embodiments, the nuclease is encoded by a nucleic acid sequence having at least 99% sequence identity with the nucleic acid sequence of SEQ ID NO: 653. In some embodiments, the nuclease is encoded by the nucleic acid sequence of SEQ ID NO: 653.

[0383] In some embodiments, the RuvC domain lacks nuclease activity. In some embodiments, the HNH domain lacks nuclease activity. In some embodiments, the modified nuclease has a modification corresponding to position H840A in Streptococcus pyogenes Cas9. In some embodiments, the modified nuclease has a modification corresponding to position D10A in Streptococcus pyogenes Cas9. In some embodiments, the modified nuclease has a modification corresponding to position D13A, which is referred to as nMG3-6(D13A) (SEQ ID NO: 152) in MG3-6 (SEQ ID NO: 646). In some embodiments, the modified nuclease has a modification corresponding to position H586A, which is referred to as nMG3-6(H586A) (SEQ ID NO: 153) in MG3-6 (SEQ ID NO: 646). In some embodiments, the modified nuclease has a modification corresponding to position N609A, which is referred to as nMG3-6(N609A) (SEQ ID NO: 154) in MG3-6 (SEQ ID NO: 646). In some embodiments, the modified nuclease is configured to cleave one strand of double-stranded target deoxyribonucleic acid. In some embodiments, the ribonucleic acid sequence configured to bind to the endonuclease comprises a tracr sequence.

[0384] In some embodiments, the nickase or nuclease comprises one or more nuclear localization sequences (NLSs) near the N-terminus or C-terminus of the nickase or nuclease.

[0385] In some embodiments, the NLS comprises any sequence or combination of sequences in Table 1 above.

[0386] Guide Nucleic Acid

[0387] In some embodiments, guide nucleic acids are provided herein, such as guide RNA (gRNA) or prime editing guide RNA (pegRNA). In polynucleotides, when T is referred to, T means U (uracil) in RNA and T (thymine) in DNA.

[0388] Prime editing enables the installation of almost any combination of point mutations, small insertions, or small deletions in the genome of living cells. The prime editing guide RNA (pegRNA) guides the prime editing protein to the target locus and also encodes the desired edit.

[0389] In some embodiments, the guide RNA targets a gene in a cell. In some embodiments, the guide RNA targets a gene in a mammalian cell. In some embodiments, the target gene is TRAC, VEGFA, AAVS1, B2M, CD5, or CD38. Exemplary guide RNAs are shown in SEQ ID NO: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855, and 1863-1910.

[0390] In some embodiments, the guide RNA is encoded by any one of the nucleic acid sequences of SEQ ID NO: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855, and 1863-1910, and has at least about 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of the nucleic acid sequences of SEQ ID NO: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855, and 1863-1910 or its reverse complement. In some embodiments, the guide RNA is encoded by a sequence having at least about 80% sequence identity with any one of the nucleic acid sequences of SEQ ID NO: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855, and 1863-1910 or its reverse complement. In some embodiments, the guide RNA is encoded by a sequence having at least about 85% sequence identity with any one of the nucleic acid sequences of SEQ ID NO: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855, and 1863-1910 or its reverse complement.In some embodiments, the guide RNA is encoded by a sequence having at least about 90% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855, and 1863-1910 or the reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 95% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855, and 1863-1910 or the reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 97% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855, and 1863-1910 or the reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 98% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855, and 1863-1910 or the reverse complement thereof.In some embodiments, the guide RNA is encoded by a sequence having at least about 99% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855, and 1863-1910 or the reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence according to any one of the nucleic acid sequences of SEQ ID NOs: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855, and 1863-1910 or the reverse complement thereof.

[0391] In some embodiments, the one or more guide RNAs are encoded by a sequence having at least about 80%, 85%, 90%, 95%, 97%, 98% or 99% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855 and 1863-1910 or the reverse complement thereof. In some embodiments, the one or more guide RNAs are encoded by a sequence having at least about 80% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855 and 1863-1910 or the reverse complement thereof. In some embodiments, the one or more guide RNAs are encoded by a sequence having at least about 85% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855 and 1863-1910 or the reverse complement thereof. In some embodiments, the one or more guide RNAs are encoded by a sequence having at least about 90% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855 and 1863-1910 or the reverse complement thereof.In some embodiments, the one or more guide RNAs are encoded by a sequence having at least about 95% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855, and 1863-1910 or the reverse complement thereof. In some embodiments, the one or more guide RNAs are encoded by a sequence having at least about 97% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855, and 1863-1910 or the reverse complement thereof. In some embodiments, the one or more guide RNAs are encoded by a sequence having at least about 98% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855, and 1863-1910 or the reverse complement thereof. In some embodiments, the one or more guide RNAs are encoded by a sequence having at least about 99% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855, and 1863-1910 or the reverse complement thereof.In some embodiments, the guide RNA is encoded by a sequence of any one of the nucleic acid sequences according to SEQ ID NO: 76-99, 109-140, 149, 656-697, 1310-1315, 1317-1341, 1451-1474, 1479-1492, 1564, 1568-1576, 1598-1620, 1656-1681, 1683-1690, 1722-1749, 1755-1774, 1784-1786, 1848-1855, and 1863-1910 or its reverse complement or its reverse complement.

[0392] In some embodiments, the guide RNA or pegRNA comprises various structural elements, including but not limited to: a spacer sequence that binds to a protospacer sequence (target sequence), a crRNA, and optionally a tracrRNA. In some embodiments, the genome editing system comprises a CRISPR guide RNA. In some embodiments, the guide RNA comprises a crRNA comprising a spacer sequence. In some embodiments, the guide RNA further comprises a tracrRNA or a modified tracrRNA.

[0393] In some embodiments, the compositions and methods provided herein comprise one or more guide RNAs. In some embodiments, the guide RNA comprises a sense sequence. In some embodiments, the guide RNA comprises an antisense sequence. In some embodiments, the guide RNA comprises a nucleotide sequence other than a region complementary or substantially complementary to a region of the target sequence. For example, the guide RNA is or is considered part of a crRNA, or is comprised in a guide crRNA (e.g., a crRNA:tracrRNA chimera).

[0394] In some embodiments, the guide RNA (e.g., gRNA) comprises synthetic nucleotides or modified nucleotides. In some embodiments, the guide RNA comprises one or more internucleotide linkers modified from natural phosphodiesters. In some embodiments, all internucleotide linkers or a contiguous nucleotide sequence thereof of the guide RNA are modified. For example, in some embodiments, the internucleotide bond comprises sulfur (S), such as a phosphorothioate internucleotide bond.

[0395] In some embodiments, the guide RNA (e.g., gRNA) comprises a modification to the ribose sugar or nucleobase. In some embodiments, the guide RNA comprises one or more nucleosides that comprise a modified sugar moiety, where the modified sugar moiety is a modification of the sugar moiety as compared to the ribose sugar moieties present in deoxyribonucleic acid (DNA) and RNA. In some embodiments, the modification is within the ribose ring structure. Exemplary modifications include, but are not limited to, replacement with a hexose ring (HNA), a bicyclic ring having a biradical bridge between the C2 and C4 carbons of the ribose ring (e.g., locked nucleic acid (LNA)), or an unlinked ribose ring that generally lacks a bond between the C2 and C3 carbons (e.g., unlocked nucleic acid (UNA)). In some embodiments, the sugar-modified nucleoside comprises bicyclohexose nucleic acid or tricyclic nucleic acid. In some embodiments, the modified nucleoside comprises a nucleoside in which the sugar moiety is replaced with a non-sugar moiety, such as peptide nucleic acid (PNA) or morpholino nucleic acid.

[0396] In some embodiments, the guide RNA comprises one or more modified sugars. In some embodiments, the sugar modification comprises a modification by changing a substituent on the ribose ring to a group other than the hydrogen or 2'-OH group that is naturally present in DNA and RNA nucleosides. In some embodiments, the substituent is introduced at the 2', 3', 4', 5' position or a combination thereof. In some embodiments, the nucleoside having a modified sugar moiety comprises a 2'-modified nucleoside, such as a 2'-substituted nucleoside. In some embodiments, the 2'-sugar-modified nucleoside is a nucleoside having a substituent other than H or -OH at the substituent position (2'-substituted nucleoside) or a nucleoside that comprises a 2'-linked biradical, and comprises 2'-substituted nucleosides and LNA (2'-4' biradical-bridged) nucleosides. Examples of 2'-substituted modified nucleosides include, but are not limited to, 2'-O-alkyl-RNA, 2'-O-methyl-RNA, 2'-alkoxy-RNA, 2'-O-methoxyethyl-RNA (MOE), 2'-amino-DNA, 2'-fluoro-RNA, and 2'-F-ANA nucleosides. In some embodiments, the modification in the ribose group comprises a modification at the 2' position of the ribose group. In some embodiments, the modification at the 2' position of the ribose group is selected from the group consisting of 2'-O-methyl, 2'-fluoro, 2'-deoxy, and 2'-O-(2-methoxyethyl).

[0397] In some embodiments, the guide RNA comprises one or more modified sugars. In some embodiments, the guide RNA comprises only modified sugars. In certain embodiments, the guide RNA comprises greater than about 10%, 25%, 50%, 75%, or 90% modified sugars. In some embodiments, the modified sugar is a bicyclic sugar. In some embodiments, the modified sugar comprises 2'-O-methoxyethyl. In some embodiments, the guide RNA comprises both nucleoside linker modifications and nucleoside modifications.

[0398] In some embodiments, the guide RNA comprises from about 15 nucleotides to about 28 nucleotides. In some embodiments, the guide RNA comprises at least 15 nucleotides. In some embodiments, the guide RNA comprises at most 28 nucleotides. In some embodiments, the guide RNA comprises from about 15 nucleotides to about 16 nucleotides, from about 15 nucleotides to about 17 nucleotides, from about 15 nucleotides to about 18 nucleotides, from about 15 nucleotides to about 19 nucleotides, from about 15 nucleotides to about 20 nucleotides, from about 15 nucleotides to about 21 nucleotides, from about 15 nucleotides to about 22 nucleotides, from about 15 nucleotides to about 23 nucleotides, from about 15 nucleotides to about 24 nucleotides, from about 15 nucleotides to about 25 nucleotides, from about 15 nucleotides to about 28 nucleotides, from about 16 nucleotides to about 17 nucleotides, from about 16 nucleotides to about 18 nucleotides, from about 16 nucleotides to about 19 nucleotides, from about 16 nucleotides to about 20 nucleotides, from about 16 nucleotides to about 21 nucleotides, from about 16 nucleotides to about 22 nucleotides, from about 16 nucleotides to about 23 nucleotides, from about 16 nucleotides to about 24 nucleotides, from about 16 nucleotides to about 25 nucleotides, from about 16 nucleotides to about 28 nucleotides, from about 17 nucleotides to about 18 nucleotides, from about 17 nucleotides to about 19 nucleotides, from about 17 nucleotides to about 20 nucleotides, from about 17 nucleotides to about 21 nucleotides, from about 17 nucleotides to about 22 nucleotides, from about 17 nucleotides to about 23 nucleotides, from about 17 nucleotides to about 24 nucleotides, from about 17 nucleotides to about 25 nucleotides, from about 17 nucleotides to about 28 nucleotides, from about 18 nucleotides to about 19 nucleotides, from about 18 nucleotides to about 20 nucleotides, from about 18 nucleotides to about 21 nucleotides, from about 18 nucleotides to about 22 nucleotides, from about 18 nucleotides to about 23 nucleotides, from about 18 nucleotides to about 24 nucleotides, from about 18 nucleotides to about 25 nucleotides, from about 18 nucleotides to about 28 nucleotides, from about 19 nucleotides to about 20 nucleotides, from about 19 nucleotides to about 21 nucleotides, from about 19 nucleotides to about 22 nucleotides, from about 19 nucleotides to about 23 nucleotides, from about 19 nucleotides to about 24 nucleotides, from about 19 nucleotides to about 25 nucleotides, from about 19 nucleotides to about 28 nucleotides, from about 20 nucleotides to about 21 nucleotides, from about 20 nucleotides to about 22 nucleotides, from about 20 nucleotides to about 23 nucleotides, from about 20 nucleotides to about 24 nucleotides, from about 20 nucleotides to about 25 nucleotides, from about 20 nucleotides to about 28 nucleotides, from about 21 nucleotides to about 22 nucleotides, from about 21 nucleotides to about 23 nucleotides, from about 21 nucleotides to about 24 nucleotides, from about 21 nucleotides to about 25 nucleotides, from about 21 nucleotides to about 28 nucleotides,From about 22 nucleotides to about 23 nucleotides, from about 22 nucleotides to about 24 nucleotides, from about 22 nucleotides to about 25 nucleotides, from about 22 nucleotides to about 28 nucleotides, from about 23 nucleotides to about 24 nucleotides, from about 23 nucleotides to about 25 nucleotides, from about 23 nucleotides to about 28 nucleotides, from about 24 nucleotides to about 25 nucleotides, from about 24 nucleotides to about 28 nucleotides, or from about 25 nucleotides to about 28 nucleotides. In some embodiments, the guide RNA comprises about 15 nucleotides, about 16 nucleotides, about 17 nucleotides, about 18 nucleotides, about 19 nucleotides, about 20 nucleotides, about 21 nucleotides, about 22 nucleotides, about 23 nucleotides, about 24 nucleotides, about 25 nucleotides, or about 28 nucleotides.

[0399] In some embodiments, the guide nucleic acid further comprises a primer binding site (PBS). In some embodiments, the primer binding site is located at the 3' end of the guide nucleic acid. In some embodiments, the primer binding site comprises at least 2, 4, 6, 8, 10, 13, 16, 20, 24, 28, 32, 36, 40, 45, 50, 55, 60, or 65 nucleotides. In some embodiments, the primer binding site comprises fewer than 2, 4, 6, or 8 nucleotides.

[0400] In some embodiments, the guide nucleic acid further comprises a reverse transcriptase template (RTT). In some embodiments, the bases in the RTT comprise bulky modifications selected from the group consisting of complex sugars, complex aminos, and / or other modifications compatible with RNA. In some embodiments, the RTT is fused to the guide RNA. In some embodiments, the guide nucleic acid further comprises a homologous sequence complementary to a region in the non-edited DNA strand. In some embodiments, the guide nucleic acid comprises a nucleic acid template. In some embodiments, the length of the RTT is at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides. In some, the length of the RTT is at least about 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides. In some embodiments, the length of the RTT is at least about 1000, 2000, 3000, 4000, or 5000 nucleotides. In some embodiments, the length of the RTT is from about 10 to about 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, or more than 200 nucleotides. In some embodiments, the length of the RTT is from about 20 to about 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, or more than 200 nucleotides. In some embodiments, the length of the RTT is from about 30 to about 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, or more than 200 nucleotides. In some embodiments, the length of the RTT is from about 40 to about 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, or more than 200 nucleotides. In some embodiments, the length of the RTT is from about 50 to about 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, or more than 200 nucleotides. In some embodiments, the length of the RTT is from about 60 to about 70, 80, 90, 100, 120, 140, 160, 180, 200, or more than 200 nucleotides. In some embodiments, the length of the RTT is from about 70 to about 80, 90, 100, 120, 140, 160, 180, 200, or more than 200 nucleotides. In some embodiments, the length of the RTT is from about 80 to about 100, 120, 140, 160, 180, 200, or more than 200 nucleotides.In some embodiments, the length of the RTT is about 100 to about 120, 140, 160, 180, 200, or more than 200 nucleotides. In some embodiments, the length of the RTT is about 100 to about 4000 nucleotides. In some embodiments, the length of the RTT is about 100 to about 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, or 4000 nucleotides. In some embodiments, the length of the RTT is about 500 to about 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, or 4000 nucleotides. In some embodiments, the length of the RTT is about 1000 to about 1500, 2000, 2500, 3000, 3500, or 4000 nucleotides. In some embodiments, the length of the RTT is about 2000 to about 2500, 3000, 3500, or 4000 nucleotides. In some embodiments, the length of the RTT is about 3000 to about 3500 or 4000 nucleotides.

[0401] Methods for preparing guide nucleic acids are known in the art. For example, guide RNAs and pegRNAs, as well as modified guide RNAs and pegRNAs, can be chemically synthesized. Additionally, nucleic acid sequences encoding guide nucleic acids can be cloned into vectors and transcribed from the vectors in vitro or in vivo using RNA polymerase.

[0402] Cell

[0403] In certain embodiments, a cell comprising the gene editing system described herein is described herein.

[0404] In some embodiments, the cell is a eukaryotic cell (e.g., a plant cell, an animal cell, a protist cell, or a fungal cell), a mammalian cell (a Chinese hamster ovary (CHO) cell, a baby hamster kidney (BHK) cell, a human embryonic kidney (HEK) cell, a murine myeloma (NS0) cell, or a human retinal cell), an immortalized cell (e.g., a HeLa cell, a COS cell, a HEK-293T cell, an MDCK cell, a 3T3 cell, a PC12 cell, a Huh7 cell, a HepG2 cell, a K562 cell, an N2a cell, or a SY5Y cell), an insect cell (e.g., a Spodoptera frugiperda cell, a Trichoplusia ni cell, a Drosophila melanogaster cell, an S2 cell, or a Heliothis virescens cell), a yeast cell (e.g., a Saccharomyces cerevisiae cell, a Cryptococcus cell, or a Candida cell), a plant cell (e.g., a parenchyma cell, a collenchyma cell, or a sclerenchyma cell), a fungal cell (e.g., a Saccharomyces cerevisiae cell, a Cryptococcus cell, or a Candida cell); or a prokaryotic cell (e.g., an Escherichia coli cell, a Streptococcus cell, a Streptomyces soil bacterium cell, or an archaeal cell). In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is an immortalized cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell is a yeast cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is a fungal cell. In some embodiments, the cell is a prokaryotic cell.

[0405] In some embodiments, the cell is A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cell, HT1080, HepG2, Huh7, K562, a primary cell, or a derivative thereof.

[0406] In some embodiments, the present disclosure provides a cell comprising the vector or nucleic acid described herein. In some embodiments, the cell expresses a gene editing system or a portion thereof. In some embodiments, the cell is a human cell. In some embodiments, the genome is edited ex vivo. In some embodiments, the genome is edited in vivo.

[0407] Delivery and Vector

[0408] In some embodiments, nucleic acid sequences encoding a gene editing system, a fusion protein comprising a nickase and a reverse transcriptase, or a guide polynucleotide are disclosed herein, the gene editing system comprising a nickase, a reverse transcriptase, and a guide polynucleotide.

[0409] In some embodiments, the nucleic acid encoding the gene editing system, fusion protein or guide polynucleotide is DNA, such as linear DNA, plasmid DNA or minicircle DNA. In some embodiments, the nucleic acid encoding the gene editing system, fusion protein or guide polynucleotide is RNA, such as mRNA.

[0410] In some embodiments, the nucleic acid encoding the gene editing system, fusion protein or guide polynucleotide is delivered by a nucleic acid-based vector. In some embodiments, the nucleic acid encoding the gene editing system, fusion protein or guide polynucleotide is delivered by a plasmid (e.g., a circular DNA molecule that can replicate autonomously within a cell), cosmid (e.g., pWE or sCos vector), artificial chromosome, human artificial chromosome (HAC), yeast artificial chromosome (YAC), bacterial artificial chromosome (BAC), P1-derived artificial chromosome (PAC), phagemid, phage derivative, telomere or virus. In some embodiments, the nucleic acid is contained in a vector selected from the list consisting of:

[0411] pSF-CMV-NEO-NH2-PPT-3XFLAG, pSF-CMV-NEO-COOH-3XFLAG, pSF-CMV-PURO-NH2-GST-TEV, pSF-OXB20-IH-TEV-FLAG(R)-6His, pCEP4 pDEST27, pSF-CMV-Ub-KrYFP, pSF-CMV-FMDV-daGFP, pEF1a-mCherry-N1 vector, pEF1a-tdTomato vector, pSF-CMV-FMDV-Hygro, pSF-CMV-PGK-Puro, pMCP tag (m), pSF-CMV-PURO-NH2-CMYC, pSF-OXB20-BetaGal, pSF-OXB20-Fluc, pSF-OXB20, pSF-Tac, pRI 101-AN DNA, pCambia2301, pTYB21, pKLAC2, pAc5.1 / V5-HisA and pDEST8.

[0412] In some embodiments, the nucleic acid-based vector contains a promoter. In some embodiments, the promoter is selected from the group consisting of: mini-promoter, inducible promoter, constitutive promoter and derivatives thereof. In some embodiments, the promoter is selected from the group consisting of: CMV, CBA, EF1a, CAG, PGK, TRE, U6, UAS, T7, Sp6, lac, araBad, trp, Ptac, p5, p19, p40, synapsin, CaMKII, GRK1 and derivatives thereof. In some embodiments, the promoter is the U6 promoter. In some embodiments, the promoter is the CAG promoter.

[0413] In some embodiments, the nucleic acid-based vector is a virus. In some embodiments, the virus is an alphavirus, parvovirus, adenovirus, AAV, baculovirus, Dengue virus, lentivirus, herpesvirus, poxvirus, anellovirus, bocavirus, vaccinia virus, or retrovirus. In some embodiments, the virus is an alphavirus. In some embodiments, the virus is a parvovirus. In some embodiments, the virus is an adenovirus. In some embodiments, the virus is AAV. In some embodiments, the virus is a baculovirus. In some embodiments, the virus is Dengue virus. In some embodiments, the virus is a lentivirus. In some embodiments, the virus is a herpesvirus. In some embodiments, the virus is a poxvirus. In some embodiments, the virus is an anellovirus. In some embodiments, the virus is a bocavirus. In some embodiments, the virus is vaccinia virus. In some embodiments, the virus is a retrovirus.

[0414] In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-rh8, AAV-rh10, AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-1, AAV-hu37, AAV-Anc80, AAV-Anc80L65, AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, AAV-3B, AAV-LK03, AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV-HSC7, AAV-HSC8, AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-HSC13, AAV-HSC14, AAV-HSC15, AAV-TT, AAV-DJ / 8, AAV-Myo, AAV-NP40, AAV-NP59, AAV-NP22, AAV-NP66, AAV-HSC16, or a derivative thereof. In some embodiments, the herpesvirus is HSV type 1, HSV-2, VZV, EBV, CMV, HHV-6, HHV-7, or HHV-8.

[0415] In some embodiments, the virus is AAV1 or a derivative thereof. In some embodiments, the virus is AAV2 or a derivative thereof. In some embodiments, the virus is AAV3 or a derivative thereof. In some embodiments, the virus is AAV4 or a derivative thereof. In some embodiments, the virus is AAV5 or a derivative thereof. In some embodiments, the virus is AAV6 or a derivative thereof. In some embodiments, the virus is AAV7 or a derivative thereof. In some embodiments, the virus is AAV8 or a derivative thereof. In some embodiments, the virus is AAV9 or a derivative thereof. In some embodiments, the virus is AAV10 or a derivative thereof. In some embodiments, the virus is AAV11 or a derivative thereof. In some embodiments, the virus is AAV12 or a derivative thereof. In some embodiments, the virus is AAV13 or a derivative thereof. In some embodiments, the virus is AAV14 or a derivative thereof. In some embodiments, the virus is AAV15 or a derivative thereof. In some embodiments, the virus is AAV16 or a derivative thereof. In some embodiments, the virus is AAV-rh8 or a derivative thereof. In some embodiments, the virus is AAV-rh10 or a derivative thereof. In some embodiments, the virus is AAV-rh20 or a derivative thereof. In some embodiments, the virus is AAV-rh39 or a derivative thereof. In some embodiments, the virus is AAV-rh74 or a derivative thereof. In some embodiments, the virus is AAV-rhM4-1 or a derivative thereof. In some embodiments, the virus is AAV-hu37 or a derivative thereof. In some embodiments, the virus is AAV-Anc80 or a derivative thereof. In some embodiments, the virus is AAV-Anc80L65 or a derivative thereof. In some embodiments, the virus is AAV-7m8 or a derivative thereof. In some embodiments, the virus is AAV-PHP-B or a derivative thereof. In some embodiments, the virus is AAV-PHP-EB or a derivative thereof. In some embodiments, the virus is AAV-2.5 or a derivative thereof. In some embodiments, the virus is AAV-2tYF or a derivative thereof. In some embodiments, the virus is AAV-3B or a derivative thereof. In some embodiments, the virus is AAV-LK03 or a derivative thereof. In some embodiments, the virus is AAV-HSC1 or a derivative thereof. In some embodiments, the virus is AAV-HSC2 or a derivative thereof. In some embodiments, the virus is AAV-HSC3 or a derivative thereof. In some embodiments, the virus is AAV-HSC4 or a derivative thereof. In some embodiments, the virus is AAV-HSC5 or a derivative thereof. In some embodiments, the virus is AAV-HSC6 or a derivative thereof. In some embodiments, the virus is AAV-HSC7 or a derivative thereof.In some embodiments, the virus is AAV-HSC8 or a derivative thereof. In some embodiments, the virus is AAV-HSC9 or a derivative thereof. In some embodiments, the virus is AAV-HSC10 or a derivative thereof. In some embodiments, the virus is AAV-HSC11 or a derivative thereof. In some embodiments, the virus is AAV-HSC12 or a derivative thereof. In some embodiments, the virus is AAV-HSC13 or a derivative thereof. In some embodiments, the virus is AAV-HSC14 or a derivative thereof. In some embodiments, the virus is AAV-HSC15 or a derivative thereof. In some embodiments, the virus is AAV-TT or a derivative thereof. In some embodiments, the virus is AAV-DJ / 8 or a derivative thereof. In some embodiments, the virus is AAV-Myo or a derivative thereof. In some embodiments, the virus is AAV-NP40 or a derivative thereof. In some embodiments, the virus is AAV-NP59 or a derivative thereof. In some embodiments, the virus is AAV-NP22 or a derivative thereof. In some embodiments, the virus is AAV-NP66 or a derivative thereof. In some embodiments, the virus is AAV-HSC16 or a derivative thereof.

[0416] In some embodiments, the virus is HSV-1 or a derivative thereof. In some embodiments, the virus is HSV-2 or a derivative thereof. In some embodiments, the virus is VZV or a derivative thereof. In some embodiments, the virus is EBV or a derivative thereof. In some embodiments, the virus is CMV or a derivative thereof. In some embodiments, the virus is HHV-6 or a derivative thereof. In some embodiments, the virus is HHV-7 or a derivative thereof. In some embodiments, the virus is HHV-8 or a derivative thereof.

[0417] In some embodiments, the nucleic acid encoding the gene editing system, fusion protein, or guide polynucleotide is delivered by a non-nucleic acid-based delivery system (e.g., a non-viral delivery system). In some embodiments, the nucleic acid is contained within a liposome. In some embodiments, the nucleic acid associates with lipids. In some embodiments, the nucleic acid associated with lipids can be encapsulated within the aqueous interior of a liposome, dispersed within the lipid bilayer of a liposome, linked to a liposome via a linking molecule that associates with both the liposome and the nucleic acid, trapped within a liposome, complexed with a liposome, dispersed within a lipid-containing solution, mixed with lipids, incorporated with lipids, contained as a suspension in lipids, contain micelles or complexed with micelles, or otherwise associated with lipids. In some embodiments, the nucleic acid is contained within a lipid nanoparticle (LNP).

[0418] In some embodiments, a nucleic acid encoding a gene editing system, a fusion protein, or a guide polynucleotide is introduced into a cell stably or transiently in any suitable manner. In some embodiments, a fusion protein or a genome editing system is transfected into a cell. In some embodiments, a cell is transduced or transfected with a nucleic acid construct encoding a fusion protein or a genome editing system. For example, a cell (e.g., with a virus encoding a fusion protein or a genome editing system) is transduced, or transfected with a nucleic acid encoding a fusion protein or a genome editing system or a translated fusion protein or genome editing system (e.g., with a plasmid encoding a fusion protein or a genome editing system). In some embodiments, the transduction is stable or transient transduction. In some embodiments, for example, when the fusion protein or genome editing system comprises a CRISPR nuclease, a cell expressing the fusion protein or genome editing system or containing the fusion protein or genome editing system is transduced or transfected with one or more gRNA or pegRNA molecules. In some embodiments, a plasmid expressing a fusion protein or genome editing system is introduced into a cell by electroporation, transient (e.g., lipid infection) and stable genomic integration (e.g., piggybac), and viral transduction (e.g., lentivirus or AAV) or other methods known to those skilled in the art. In some embodiments, a gene editing system is introduced into a cell as one or more polypeptides. In some embodiments, delivery is achieved by using an RNP complex. Methods for delivering polypeptides and / or RNPs to cells are known in the art, such as by electroporation or by cell squeezing.

[0419] Exemplary delivery methods for nucleic acids include lipofection, nucleofection, electroporation, stable genomic integration (e.g., piggybac), microinjection, biolistics, virions, liposomes, immunoliposomes, polycation or lipid nucleic acid conjugates, naked DNA, artificial virus particles, and agent-enhanced DNA uptake. Lipofection is described, for example, in U.S. Patent Nos. 5,049,386; 4,946,787; and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam TM , Lipofectin TM and SF Cell Line 4D-Nucleofector X Kit TM (Lonza). Cationic and neutral lipids suitable for efficient receptor recognition of polynucleotides in lipofection include the cationic and neutral lipids in WO 91 / 17424 and WO 91 / 16024. In some embodiments, delivery is to a cell (e.g., in vitro or ex vivo administration) or a target tissue (e.g., in vivo administration). In some embodiments, the nucleic acid is contained in a liposome or nanoparticle that specifically targets the host cell.

[0420] Additional methods for delivering nucleic acids to cells are known to those of skill in the art. See, for example, US2003 / 0087817.

[0421] Method of Use

[0422] In some embodiments, methods for modifying double-stranded and / or single-stranded nucleic acids are described herein, the methods comprising: a) providing a guide nucleic acid to a cell to bind to a target strand of the double-stranded nucleic acid; b) providing a nuclease or nickase to the cell to cleave the double-stranded nucleic acid at the binding position of the guide nucleic acid; c) providing a reverse transcriptase to the cell to synthesize a modification in the target strand of the double-stranded nucleic acid at the position cleaved by the nickase and / or the double-stranded nuclease.

[0423] In some embodiments, the method is used to introduce a modification into the genome of a cell. In some embodiments, the modification is an insertion, deletion, or mutation. In some embodiments, the method is used to introduce a site-directed insertion, deletion, and / or mutation (e.g., an insertion and a mutation) into the genome of a cell. In some embodiments, the method is used in combination with a nucleic acid template to facilitate site-directed insertion into the genome of a cell. In some embodiments, the cell is a human cell. In some embodiments, the genome of the cell or a vector included in the cell is modified. In some embodiments, the genome of the cell is modified ex vivo. In some embodiments, the genome of the cell is modified in vivo.

[0424] In some embodiments, the method further comprises providing a transposase, integrase, or homing endonuclease to the cell. In some embodiments, the method further comprises providing a retrotransposon to the cell. In some embodiments, the method further comprises providing an RNA or DNA insertion template.

[0425] In some embodiments, the methods described herein further comprise detecting a genomic modification. In some embodiments, after the genome of the cell is modified, the cell is cultured for a period of time. In some embodiments, DNA or RNA is extracted and sequenced, and the modified sequence regions are mapped and compared to the unmodified sequences. In some embodiments, the cells are stained with an antibody to the protein product that is translated from the modified nucleic acid, and the resulting stained protein or polypeptide in the cells is analyzed, for example, by flow cytometry.

[0426] The methods described herein can be used, for example, for targeted SNP correction, small insertions, or small deletions. Additionally, by using a suitable RTT, the methods described herein can be used for targeted insertion of large templates into the genome of a cell.

[0427] Kit

[0428] In some embodiments, the present disclosure provides a kit comprising one or more nucleic acid constructs encoding various components of the fusion proteins or genome editing systems described herein, e.g., nucleotide sequences encoding components of a fusion protein or genome editing system capable of modifying a target DNA sequence. In some embodiments, the nucleotide sequence comprises a heterologous promoter driving the expression of the components of the RNA genome editing system.

[0429] In some embodiments, any of the targetable reverse transcriptases or genome editing systems disclosed herein are assembled into a pharmaceutical kit, a diagnostic kit, or a research kit to facilitate their use in therapeutic, diagnostic, or research applications. The kit may include one or more containers containing any of the vectors disclosed herein and instructions for use.

[0430] The kit can be designed to assist a researcher in using the methods described herein and can take many forms. Each composition of the kit, where applicable, can be provided in liquid form (e.g., as a solution) or in solid form (e.g., as a dry powder). In certain cases, some compositions can be constituted or otherwise processed (e.g., to form an active form), e.g., by adding a suitable solvent or other species (e.g., water or cell culture medium), which may or may not be provided with the kit. As used herein, "instructions" can define the composition of the instructions and / or facilitation and generally refer to written instructions on or associated with the packaging of the present disclosure. The instructions can also include any oral or electronic instructions provided in any manner such that the user will clearly recognize that the instructions will be associated with the kit, e.g., audio-visual (e.g., videotape, DVD, etc.), Internet, and / or web-based communications, etc. In some embodiments, the written instructions are in a form prescribed by a government agency that regulates the manufacture, use, or sale of pharmaceuticals or biological products, and the instructions may also reflect the agency's approval for manufacture, use, or sale for animal administration.

[0431] Examples

[0432] The following examples are given for the purpose of illustrating various embodiments of the present disclosure and are not meant to limit the present disclosure in any way. The present examples, as well as the methods described herein, are currently representative of exemplary preferred embodiments and are not intended to limit the scope of the present disclosure. Those skilled in the art will envision variations and other uses that fall within the spirit of the present disclosure as defined by the scope of the claims.

[0433] Example 1. Bioinformatics Identification of Reverse Transcriptases in Metagenomic Databases

[0434] This example describes the identification of proteins with reverse transcriptase function by bioinformatics methods.

[0435] Bioinformatics analysis of a metagenomic database driven by the extensive assembly of microbial, viral, and eukaryotic genomes was performed to search for proteins with putative reverse transcriptase function. The analysis revealed millions of proteins with predicted reverse transcriptase function. The predicted RT hits were then bioinformatics screened to obtain full-length open reading frames (ORFs) in which high-quality RT domain hits covered more than 70% of the reference RT domain and contained the expected catalytic residues. After screening, 468 RTs were selected for their potential to develop gene editing tools (SEQ ID NO: 161 - 629). For all these identified putative RTs, the predicted active site quartet motif was [Y / F]XDD, where the most common amino acid at the first position of the quartet was tyrosine (Y, 85.2%) or phenylalanine (F, 14.5%). The second position of the quartet was more diverse, with the most common residues being alanine (A, 55.5%), isoleucine (I, 9.3%), and valine (V, 19.3%). The aspartic acid dyad (DD) was the most conserved feature of RT activity.

[0436] Example 2. Reverse transcriptase (RT) for short corrections, small insertions, and deletions

[0437] This example describes targeted genome editing using an untethered reverse transcriptase in combination with pegRNA in HEK293T cells.

[0438] Testing reverse transcriptase candidates with an untethered nickase

[0439] Reverse transcriptase (RT) candidates from the MG151 (SEQ ID NO: 1 - 37), MG153 (SEQ ID NO: 38 - 61), and MG160 families (SEQ ID NO: 62 - 75) were cloned into plasmids in which the expression of the RT candidates was driven by the CMV promoter. The plasmids were isolated for transfection into HEK293T cells. A second plasmid containing the nickase spCas9 (H840A) (where expression was driven by the CMV promoter) and the plasmid containing the RT were co - transfected. The transfection was with a chemically synthesized pegRNA (SEQ ID NO: 76 - 99) that contained the desired edit in the RT template. All components (plasmids and pegRNA) were reverse - transfected into 150,000 HEK293T cells in a 24 - well plate. Seventy - two hours after transfection, the cells were lysed in 100 μL of solution. A primer containing a barcode for next - generation sequencing (NGS) (SEQ ID NO: 100 - 101) was used to amplify a ~250 bp target (SEQ ID NO: 102) with a master mix. Then PCR purification was performed, and the samples were subjected to NGS sequencing. The FASTQ files were then processed using prime editing to determine the percentage of reads with the desired change.

[0440] MG151 family

[0441] Prime editing of untethered MG151 candidates 80 - 85 (SEQ ID NO: 1 - 6), 87 - 100 (SEQ ID NO: 7 - 20), and 102 - 117 (SEQ ID NO: 22 - 37) was tested in HEK293T cells to determine the percentage of changes with the desired correction. For each pegRNA (SEQ ID NO: 76 - 83) with different PBS lengths (2, 4, 6, 8, 10, 13, 16, 20 nucleotides), the editing percentage for each RT is shown in Figures 1A - 1JJ In a single replicate, the editing of MG151 - 98 (SEQ ID NO: 18) and MG151 - 99 (SEQ ID NO: 19) was six - fold and four - fold, respectively, that of the wild - type MMLV ([ Figure 2 ). The editing levels of the MG151 candidates MG151 - 100 (SEQ ID NO: 19), MG151 - 103 (SEQ ID NO: 23), MG151 - 104 (SEQ ID NO: 24), and MG151 - 105 (SEQ ID NO: 25) were half or comparable to that of the wild - type MMLV ([ Figure 2 ).

[0442] MG153 family

[0443] Untethered MG153 candidates 1 - 5 (SEQ ID NO:38 - 42), 7 - 21 (SEQ ID NO:44 - 58), and 25 - 27 (SEQ ID NO:59 - 61) were tested for prime editing in HEK293T cells to determine the percentage of desired corrected changes. For each pegRNA (SEQ ID NO:76 - 83) with different PBS lengths (2, 4, 6, 8, 10, 13, 16, 20 nucleotides), the editing percentage for each RT is shown in Figures 3A - 3O and 3P - 3W. MG153 - 1 (SEQ ID NO:38), MG153 - 3 (SEQ ID NO:40), MG153 - 7 (SEQ ID NO:44), MG153 - 9 (SEQ ID NO:46), MG153 - 12 (SEQ ID NO:49), and MG153 - 15 (SEQ ID NO:52) have shown editing levels above background or comparable to MMLV wild - type.

[0444] MG160 family

[0445] As described above, the activities of untethered MG160 family candidates MG160 - 1 to MG160 - 8 (SEQ ID NO:62 - 68) were tested in mammalian cells. For untethered candidates MG160 - 1 (SEQ ID NO:62) and MG160 - 4 (SEQ ID NO:65), activities above background were observed. ( Figures 4A - 4G ).

[0446] Testing reverse transcriptase candidates tethered to nickase

[0447] The activities of different RT classes with CRISPR type II nucleases were evaluated. RT candidates were cloned into plasmids containing nickase spCas9 (H840A) to generate RT - nickase fusions. The CMV promoter drives the expression of the RT - nickase fusion protein, which contains a thirty - three amino acid linker (SEQ ID NO:103) between the nickase and the RT candidate. The fusion protein was then transfected into HEK293T cells and processed by NGS as described above.

[0448] The activities of tethered MG160 candidates 1 - 5 (SEQ ID NO:69 - 73) are shown in Figures 5A - 5E . Specifically, the level of candidate MG160 - 4 (SEQ ID NO:72) is comparable to that of wild - type MMLV ( Figure 5D)。All other MG160 candidates (SEQ ID NO: 69 - 72) have at least half the activity of wild - type MMLV at a specific PBS length. Additionally, when MG160 - 1 (SEQ ID NO: 69) and MG160 - 4 (SEQ ID NO: 72) were repeated with two additional biological replicates, the editing of MG160 - 1 (SEQ ID NO: 69) was comparable to that of MMLV WT, and the editing of MG160 - 4 (SEQ ID NO: 72) was 2 - fold higher ( Figures 5F - 5G )。

[0449] The above data demonstrate that several RTs from different phylogenetic families were identified, showing comparable or higher activity to MMLV WT in the prime - editing context. Having activity across a wide range of families allows for the identification of RT candidates for different types of genome modifications (i.e., SNP correction, insertion, or deletion). At least 2 RTs of approximately 250 aa in size were identified, whose performance was similar to or better than that of MMLV WT (MG160 - 1 (SEQ ID NO: 69) and MG160 - 4 (SEQ ID NO: 72)). The small size of the RT (1 / 3 of MMLV WT) allows for efficient delivery using adeno - associated virus (AAV) and lipid nanoparticles (LNP).

[0450] Example 3. RTs for short correction, small insertions, and deletions (predictive)

[0451] This example describes targeted genome editing using additional reverse transcriptases in combination with pegRNAs in HEK293T cells.

[0452] Additional RTs from the MG151 and MG153 families, including MG151 - 101 (SEQ ID NO: 21), MG153 - 6 (SEQ ID NO: 43), or additional candidates, were tested in the untethered form as described in Example 2. This allowed for small corrections, insertions, and deletions of additional RT candidates.

[0453] RTs from the MG160 family, including MG160 - 6 (SEQ ID NO: 74), MG160 - 8 (SEQ ID NO: 75), and other candidates, were tested for editing in the tethered system as described above. This allowed for the identification of additional mini - (approx. 250 aa) RT systems that can mediate small corrections, insertions, and deletions.

[0454] Example 4. Nucleases for mediating short correction, small insertions, and deletions in combination with reverse transcriptases

[0455] This example describes targeted genome editing using a combination of an RNA-guided nuclease and pegRNA in HEK293T cells.

[0456] To evaluate the requirements for pegRNAs (gRNAs with 3' extensions) designed for the nucleases of the present disclosure, a number of PBSs of different lengths were evaluated to maintain proper nuclease-gRNA interactions. To test nuclease activity in combination with pegRNA design, InDel formation in HEK293T cells was tested. MG3-6 (SEQ ID NO:104) was used as the nuclease for pegRNA combinations with different PBS lengths. Four endogenous genomic target sites (AAVS1 (SEQ ID NO:105), B2M (SEQ ID NO:106), CD5 (SEQ ID NO:107), and CD38 (SEQ ID NO:108)) known to be recognized by wild-type MG3-6 (SEQ ID NO:104) were targeted with chemically synthesized pegRNAs having the following different PBS lengths: 2, 4, 8, 10, 13, 16, and 20 nucleotides (SEQ ID NO:109-140). MG3-6 mRNA (SEQ ID NO:104) was co-transfected with a guide RNA (control) or pegRNAs (with various PBS lengths). RNA was reverse transfected into 24-well plates with 50,000 HEK293T cells. Forty-eight hours after transfection, the cells were lysed in 100 μL solution. Primers (SEQ ID NO:141-148) were used to amplify the ~700 bp target products (SEQ ID NO:105-108) with a master mix. The samples were then purified and Sanger sequenced. The Sanger sequences were then analyzed by ICE to calculate the InDel percentage for each target site.

[0457] Sanger sequencing traces using ICE analysis showed that wild-type MG3-6 preferred pegRNAs with PBS lengths equal to or less than eight nucleotides ( Figures 6A - 6D ). Compared to pegRNAs for the corresponding target genes, WT guide RNAs (without a PBS region) binding to MG3-6 mRNA (SEQ ID NO:104) gave the highest InDel percentages for all endogenous targets (55% AAVS1, 84% B2M, 58.5% CD5, and 24% CD38). As the PBS length increased from two nucleotides to twenty nucleotides, the InDel percentages for all endogenous targets (SEQ ID NO:105-108) decreased. For example, at the target site AAVS1 (SEQ ID NO:105) with a PBS length of 2 nucleotides (SEQ IDNO:109) (53%) ( Figure 6A) The InDel percentage at [[ID=]] is similar to that observed with the WT guide RNA (SEQ ID NO: 116) (55%), but in the case of a PBS length of 20 nucleotides (SEQ ID NO: 115), the InDel percentage drops to approximately 11%. Thus, the results illustrate the general rules for pegRNA design of the MG3-6 gene editing system and highlight the importance of identifying RTs with shorter PBS length requirements.

[0458] Example 5. Use of processive RT in combination with modified pegRNAs for short corrections, small insertions, and deletions (predictive)

[0459] This example describes the use of a reverse transcriptase in combination with a CRISPR nickase and a pegRNA for targeted genome editing in HEK293T cells.

[0460] Current setups for prime editing require a pegRNA consisting of a spacer, followed by crRNA, tracr, RTT, and PBS (from 5'-3'). It has been shown that MMLV WT (MMLV1) and MMLV penta-mutant (MMLV2) have a certain level of pegRNA readthrough, thus integrating part of the tracr sequence into genomic DNA (gDNA), which is an undesired property as this design generates unwanted mutations in the genomic DNA. RTs from the GII intron family were identified, which express well in mammalian cells and show high activity for cDNA synthesis. RTs from the GII intron family generally show higher processive synthesis capabilities. Higher processive synthesis capabilities translate into the RT being able to read through structured RNAs (e.g., the crRNA-tracr portion of the pegRNA) and being able to read through small / medium-sized chemical modifications in RNA. Since RTs from GII introns show good cDNA synthesis activity and good expression in mammalian cells, they are used in the prime editing context to generate small genomic corrections, small insertions, and / or deletions. To use processive RT in the prime editing context, pegRNA readthrough as described above needs to be avoided. To achieve pegRNA readthrough by the RT, bulky modifications were incorporated into the pegRNA, e.g., into the last base of the RTT (if read from 3' to 5') or the first base (if read from 5' to 3' read segment). Bulky modifications include, for example, complex sugars or complex amines, and / or other modifications compatible with RNA.

[0461] Using Lipofectamine 2000, transfect plasmids containing nickase and a progressive RT whose activity is to be tested into cells (e.g., HEK293T cells). Transfect chemically synthesized RNAs (with or without bulky modifications) into cells using Lipofectamine MessengerMAX. 72 hours after transfection, lysate the cells in 100 μL solution. Use primers (SEQ ID NO: 100 - 101) containing barcodes for next-generation sequencing (NGS) to amplify a ~250 bp target (SEQ ID NO: 102). Then perform PCR purification and subject the samples to NGS sequencing. Process the resulting FASTQ files using prime editing to determine the percentage of reads with the desired changes.

[0462] The experiments described above allow prime editing with high-performance RT in the mammalian cell context with little or no pegRNA readthrough.

[0463] Example 6. RT for Programmable Large Cargo Integration by Target-Primed Reverse Transcription (Predictive)

[0464] This example describes the use of a reverse transcriptase with reverse transcriptase activity in combination with a CRISPR nickase and a pegRNA for targeted genome editing.

[0465] Targeted integration of large cargoes into human genomic DNA in living cells is a long-sought goal of gene editing. To date, the most effective method for integrating large cargoes into the cell genome is through the use of lentiviruses. However, lentivirus-mediated integration lacks a targeting feature as integration mostly occurs randomly in the open chromatin of the cell. For large cargo integration, an RT with high processive synthesis ability and high fidelity in combination with a nuclease is advantageous. The nuclease provides targeting in gDNA, while the RT utilizing the target-primed reverse transcription mechanism can integrate large RNA cargoes into mammalian gDNA.

[0466] The potential of RT candidates to effect large integrations was tested by their ability to retrotranspose an RNA template containing a GFP cassette that can only produce GFP (and thus fluorescence) after successful retrotransposition. The target for retrotransposition was determined by a nuclease. This nuclease generates primer sites via double-strand break events. Type II nucleases (alternatively type V nucleases) were tested to identify the optimal nuclease for generating gDNA primers. The VEGFA gene was selected for targeted integration and was targeted by the nuclease along with a chemically synthesized VEGFA guide (SEQ ID NO:149). Candidate reverse transcriptases were cloned into plasmids for mammalian expression under the CMV promoter. To localize the RT to the nucleus upon expression, one or more nuclear localization signal (NLS) sequences were added to the N-terminus and C-terminus of the RT. Additionally, the MS2 coat protein (MCP) sequence and Flag-HA (FH) tag were fused to the N-terminus of the RT. The MCP is a protein derived from the MS2 bacteriophage that recognizes a 20-nucleotide RNA stem-loop (MS2 loop) with high affinity (sub-nanomolar Kd range). The MS2 loop was added to the RT template encoded within the same plasmid, ensuring that the expressed MCP-RT fusion protein finds the RNA template for reverse transcription. Additionally, a 20-nucleotide sequence complementary to the 3' overhang generated by the nuclease serves as a primer binding site (PBS) for initiating reverse transcription. To quantify the efficiency of retrotransposition, an inverted GFP cassette driven by the EF1α promoter was cloned downstream of the RT fusion. The GFP was interrupted by an intron (two different intron sequences were tested, referred to as the normal intron and the chimeric intron), oriented such that it can only be spliced out of transcripts driven by the CMV promoter and not the EF1α promoter ( Figure 7 ). Thus, only after successful reverse transcription of this spliced RNA can the cell express GFP fluorescence. The PBS and MS2 loop were cloned downstream of the EF1α promoter, followed by a polyA sequence to stabilize the RNA template. This design ensures that the GFP fluorescence exhibited by cells expressing this plasmid is correlated with the efficiency of retrotransposition and thus gives a measure of the ability of the RT candidate to reverse transcribe and integrate large DNA fragments.

[0467] The RT candidates were cloned into GFP-based retrotransposition plasmids (SEQ ID NO:150 - 151 and 2580 - 2581) and isolated for transfection into HEK293T cells. Transfection was performed using Lipofectamine2000. After 24 hours, the cells were split into media containing puromycin to select for transfected cells expressing the plasmid. Five days later, the cells were run on a cell sorter and the percentage of GFP-positive cells in the population was quantified.

[0468] To test hundreds or thousands of RTs and / or conditions (engineered systems), the above method also allows high-throughput testing. Hundreds or thousands of conditions are pooled together, and a single pooled plasmid transfection is performed. Five days after transfection, cells expressing GFP are sorted. The best-performing RTs are identified by sequencing GFP-positive cells and mapping the RTs by using a combination of random primers and primers matching the second exon of GFP. The RTs enriched by this pooling method are then individually validated.

[0469] This method allows the identification of RTs capable of mediating large cargo integration via target-primed reverse transcription mechanisms. Thus, engineered nuclease / RT constructs allow the development of RNA-mediated large cargo integration into the genomic DNA of mammalian cells.

[0470] Example 7. RT for single-stranded DNA transposase-mediated programmable large cargo integration (predictive)

[0471] This example describes the use of a reverse transcriptase with reverse transcriptase activity in combination with TnpA for targeted genome editing.

[0472] Retrotransposons are DNA elements that contain an RT enzyme encoded downstream of a conserved non-coding structural RNA. The non-coding RNA consists of two inverted regions (referred to as msr and msd). When the retrotransposon RT recognizes the folded ncRNA, it reverse transcribes the msd portion (template), thereby generating ssDNA.

[0473] The IS200 / IS605 transposon is a mobile genetic element that integrates ssDNA at specific target sites via the TnpA transposase. TnpA excises the donor by recognizing the structural motifs at each donor end and integrates it into the recognized target site, which can be used as ssDNA.

[0474] The ssDNA generated by the retrotransposon RT can be used by TnpA as a template for the programmable integration of a desired cargo into a specific target site. Specifically, the reverse transcribed msd can contain a desired cargo (e.g., an antibiotic resistance cassette or a fluorescent marker) flanked by LE and RE structural motifs recognizable by TnpA. The TnpA transposase excises and circularizes the ssDNA donor and integrates it into the target by recognizing specific motifs obtained through an R-loop formed by RNA-guided recognition and binding of engineered (nickase or dead) effectors (e.g., MG3-6) ( Figure 8 )

[0475] Example 8. RTs for short corrections, small insertions, and deletions

[0476] Testing reverse transcriptase candidates untethered to a nickase

[0477] Retroviral reverse transcriptase (RT) candidates from the MG151 and MG153 families were cloned into plasmids in which the expression of the RT candidates was driven by the CMV promoter. The plasmids were then isolated for transfection in HEK293T cells. Another plasmid driven by the CMV promoter containing the nickase spCas9 (H840A) and the plasmid containing the RT were co-transfected. The transfection contained a chemically synthesized pegRNA (SEQ ID NOs: 656 - 697) that contained the desired edits in the RT template. All components (plasmids and pegRNA) were reverse transfected into 150,000 HEK293T cells in a 24-well plate. 72 hours after transfection, the cells were lysed in 100 μL of solution. A primer containing a barcode for next-generation sequencing (NGS) (SEQ ID NOs: 649 - 650) was used to amplify a ~250 bp target (SEQ ID NOs: 654 - 655) with a master mix. PCR purification was then performed, and the samples were sent for NGS sequencing. The FASTQ files were then processed using prime editing to determine the percentage of reads with the desired changes.

[0478] Prime editing of untethered MG151 candidates MG118 - MG135 (SEQ ID NOs: 710 - 727) was tested in HEK293T cells to determine the percentage of changes with the desired correction. For each pegRNA with different PBS lengths (2, 4, 6, 8, 10, 13, 16, and 20 nucleotides), the editing percentage for each RT is shown in Figures 9A - 9R In a single replicate, compared to MMLV WT RT, MG151 - 123 to MG151 - 126 had comparable or higher editing efficiency ( Figures 9F - 9I ). These results were replicated, and the biological replicates are shown in Figures 10A - 10D where four candidates edited at levels comparable to MMLV WT when challenged with a G - T transversion. Importantly, all of these candidates performed well with a PBS length of 8 nt, enabling shortening of the pegRNA and PBS - spacer hybridization window.

[0479] In addition, two candidates from the 151 family (MG151 - 98 and MG151 - 99) were rationally engineered to install beneficial mutations observed in other RTs (Anzalone et al., 2022). Various single or combined point mutations, as well as truncations of the RNase H domain (SEQ ID NOs: 750 - 766), were evaluated. Mutations H171N, K297P, and trimming the last 166 aa of MG151 - 98 improved prime editing efficiency, with some mutations outperforming the MMLV pentamutant ( Figures 11A - 11B)。For MG151-99, I264K, R556K, trimming of the last 152 aa proved beneficial ( Figures 11C - 11D )。 Figures 12A - 12B Shows the profile of G-T transversion assessment of MG151 candidates in HEK293T cells targeting the VEGFA gene.

[0480] Prime editing of untethered MG153 candidates MG153-29, MG153-31, MG153-33, MG153-35, MG153-36, MG153-45, and MG153-53 was tested in HEK293T cells to determine the percentage of desired corrected changes. For each pegRNA with different PBS lengths (2, 4, 6, 8, 10, 13, 16, and 20 nucleotides), the editing percentage for each RT is shown in Figures 13A - 13H . The activity levels of several RTs, including MG153-33, MG153-35, MG153-45, and MG153-53, were comparable or higher compared to MMLV WT RT ( Figures 13C - 13D and Figures 13F - 13G ). Importantly, MG153-53 performed more than 2-fold better than MMLV WT ( Figure 13G ). When tested as a fusion protein with Cas9, this candidate was also active ( Figure 13H ), demonstrating its versatility. Figures 14A - 14B Shows the profile of G-T transversion assessment of MG153 candidates in HEK293T cells targeting the VEGFA gene.

[0481] Tested reverse transcriptase candidates tethered to Cas nickase

[0482] RT candidates were cloned into a plasmid containing the nickase spCas9 (H840A) to generate RT-nickase fusions. The CMV promoter drives the expression of the fusion protein, which contains a thirty-three amino acid linker (SEQ ID NO:103) between the nickase and the RT candidate. The fusion protein was then transfected into HEK293T cells and processed by NGS as described above.

[0483] Figures 15A - 15UShows the editing activities of RT candidates MG160-17, MG160-28, MG160-31, MG160-37, MG160-40, and MG160-51 to MG160-67. Several candidates showed editing levels comparable to MMLV WT, including MG160-17, MG160-28, MG160-37, MG160-54, MG160-56, MG160-57, MG160-59, MG160-64, MG160-65, and MG160-63. Figures 16A - 16B Shows the profile of G-T transversions for the MG153 candidates evaluated in HEK293T cells targeting the VEGFA gene.

[0484] The results demonstrated that several RTs from different phylogenetic families exhibited similar or higher activities than MMLV WT RT in the prime editing context. Having activities across a broad family allows for the nomination of RT candidates that may be most suitable for different types of modifications (i.e., SNP correction, insertion, or deletion). Additionally, several RTs of approximately 250 aa in size were identified, and their performance was similar to or better than MMLV WT. Their small size (about one-third the size of MMLV WT RT) makes them promising candidates for developing compact systems capable of efficient delivery using adenovirus (AAV) and lipid nanoparticles (LNP).

[0485] RTs for small insertions and deletions

[0486] RT candidates from the MG151, MG153, and MG160 families were required to perform 24-nt insertions and 15-nt deletions in the VEGFA gene to test their ability to perform small and medium-sized corrections ( Figures 17A - 24H ). Most of the candidates that performed well in the G-T transversion experiments were also able to efficiently perform insertions and deletions. For example, well-performing candidates from the MG151 family include MG151-98, MG151-99 ( Figures 17A - 17D ), MG151-23 ( Figure 18A and 18E ) and MG151-26 ( Figure 18D and 18H ), as well as engineered variants K297P, H171N, and 166 aa trimmed for MG151-98 ( Figures 19A - 19D ), and I264K, R556K, and 152 aa trimmed for MG151-99 ( Figures 20A - 20D ). MG153-53 is a well-performing candidate in the MG153 family ( Figure 21D and 22D)。Well - performing candidates from the MG160 family include MG160 - 4( Figure 23H and 24H ), MG160 - 37( Figure 23C and 24C ), MG160 - 54( Figure 23D and 24D ), and MG160 - 64( Figure 23G and 24G ).

[0487] Example 9. Nucleases for binding to reverse transcriptase to mediate short genomic correction

[0488] The targeting required for RT - installed genomic correction, insertion, or deletion can be provided by a nickase. The nickase cleaves the non - target strand, generating a primer for reverse transcription. The gRNA accompanying the nickase is a modified version (pegRNA) consisting of a 3' extension containing an RNA template (RTT) and a PBS. The PBS and the spacer can be complementary to each other, and this complementarity can cause disruption of the gRNA structure, leading to disruption of the interaction between the pegRNA and its nickase and ultimately failure to target the gene of interest. Since each nuclease interacts with its own gRNA, the pegRNA design and requirements will vary depending on the system.

[0489] To test the versatility of RT candidates in binding to nucleases, several RTs were tested with the MG3 - 6 H586A nickase, either untethered or tethered (fused)( Figures 25A - 25D ). Moderate editing levels were detected in both the tethered and untethered systems. The editing levels can be improved by optimizing the constructs and delivery methods.

[0490] Example 10. RTs for programmable large cargo integration by target - primed reverse transcription

[0491] The ability of RT candidates to produce large integrations was tested by their ability to retrotranspose an RNA template containing a GFP cassette, which can only produce GFP (and thus fluorescence) after successful retrotransposition. The target for retrotransposition was determined by a Cas nuclease.

[0492] RT candidates were cloned into a GFP-based retrotransposition plasmid and isolated for transfection into HEK293T cells. Plasmid transfection was performed using Lipofectamine 2000, while Cas9 mRNA and chemically synthesized guides were transfected using Lipofectamine messenger max. After 24 hours, the cells were split into medium containing puromycin to select for transfected cells expressing the plasmid. Three, six, and eight days later, the cells were run on a cell sorter and the percentage of GFP-positive cells in the population was quantified.

[0493] MG candidates MG153-18 and MG153-20 showed an increase in GFP fluorescence from day 3 to day 6, above the non-targeting background, indicating successful retrotransposition in the VEGFA gene( Figures 26A - 26C ). These results indicate that MG Rt is capable of long (>1 kb) targeted integration in the human genome.

[0494] Example 11. Prime editing of engineered RT

[0495] Reverse transcriptase (RT) candidates from the MG151 family, MG160, and MG153 family were cloned into a plasmid in which the expression of the RT candidate was driven by the CMV promoter. The plasmid was then isolated for transfection in HEK293T cells. Another plasmid driven by the CMV promoter containing the nickase spCas9 (H840A) and the plasmid containing the RT were co-transfected. The transfected chemically synthesized pegRNA contained the desired edit in the RT template. All components (plasmid and pegRNA) were reverse transfected into 150,000 HEK293T cells in a 24-well plate. 72 hours after transfection, the cells were lysed in 100 uL solution. Primers containing barcodes for next-generation sequencing (NGS) were used to amplify a ~250 bp target. PCR purification was then performed and the samples were sent for NGS sequencing. The FASTQ files were then processed using prime editing to determine the percentage of reads with the desired change.

[0496] Data are shown in Figures 27A - 27CAmong three RTs from different families against primer binding sites (PBS lengths) across multiple sizes, a G-to-T transversion in the VEGFA gene was shown. The ultrasmall MG160-4 candidate was superior to MMLV WT (PE1) and behaved very similarly to the gold standard MMLV pentamutant (PE2). The MG151-98 candidate behaved close to PE1 in its WT form. Across various PBS lengths, the medium-sized 153-53 candidate was superior to PE1. Additionally, MG151-98 was rationally engineered to install beneficial mutations observed in other RTs. Various point mutations, alone or in combination, and truncations of the RNase H domain were evaluated. The mutations H171N, K297P, and trimming the last 166 aa of MG151-98 improved prime editing efficiency, with some mutations superior to the MMLV pentamutant.

[0497] Example 12. Processive RTs for large RNA template integration

[0498] Their ability to generate cDNA in a mammalian context was tested by expressing candidate reverse transcriptases in mammalian cells and detecting cDNA synthesis by qPCR. The reverse transcriptases were cloned into plasmids for mammalian expression under the CMV promoter. A 4 kb RNA template was generated by in vitro transcription and hybridized to a DNA primer.

[0499] Plasmids containing MCP fused to RT candidates under the CMV promoter were cloned and isolated for transfection in HEK293T cells. Transfection was performed using lipofectamine 2000. mRNA encoding dCas9 fused to NanoLuc was prepared. To degrade any DNA template remaining in the mRNA preparation, the reaction was treated with DNase for 1.5 h and the mRNA was cleaned. The mRNA was hybridized to the complementary DNA primer at 95 °C in 10 mM Tris pH 7.5, 50 mM NaCl for 2 min and cooled to 4 °C at a rate of 0.1 °C / sec. After transfection of the plasmid containing the MCP-RT fusion, the mRNA / DNA mixture was transfected into HEK293T cells for 6 h. Eighteen hours after mRNA / DNA transfection, the cells were lysed using a solution, adding 100 μl of QuickExtract per well in a 24-well plate. The RNA template was approximately 4247 nt. Primers designed to amplify the first and last 100 bp products from the newly synthesized cDNA (4100 bp), and a TaqMan probe were used to quantify its amplification.

[0500] Data are shown in Figure 28In the detection of the activities of the control GII intron RT TGIRT, the retrovirus MMLV (WT and pentamutant), and the positive control R2 R2Tg, as shown by the early amplification of the first and last 100 bp products. As expected for RTs with low processive synthesis ability, the retroviral RTs (MMLV) showed a high amplification level for the first 100 bp (FAM signal), but a low level for their completion of cDNA synthesis (last 100 bp) (1 / 20 of the first 100 bp, as observed by the FAM / HEX ratio signal). Group II intron-derived RTs such as MG153-18, MG153-20, MG153-51, MG153-56, MG170-1 and R2 non-LTR retrotransposon RTs such as MG140-3, MG140-8 and MG140-46 showed closer FAM / HEX ratios, demonstrating their high processive synthesis ability.

[0501] Example 13. RTs for short corrections, small insertions, and deletions

[0502] Testing reverse transcriptases tethered to the spCas9 (H840A) nickase

[0503] RT candidates were cloned into a plasmid containing the nickase spCas9 (H840A) (SEQ ID NO: 1247) to generate RT-nickase fusions. The CMV promoter drives the expression of the fusion protein, which contains a 33-amino acid linker (SEQ ID NO: 103) between the nickase and the RT candidate (SEQ ID NO: 1250-1279). The fusion protein was then transfected into HEK293T cells. The transfection contained a chemically synthesized pegRNA (SEQ ID NO: 656-679) with the desired edit in the RT template. All components (plasmid and pegRNA) were reverse transfected into 50,000 HEK293T cells in a 24-well plate. 72 hours after transfection, the cells were lysed in 100 μL of extraction solution. A primer containing a barcode for next-generation sequencing (NGS; SEQ ID NO: 100-101) was used to amplify a ~250 bp target (SEQ ID NO: 102). PCR purification was then performed, and the samples were sent for NGS sequencing. The FASTQ files were then processed using the prime editing setup to determine the percentage of reads with the desired change.

[0504] Results

[0505] In HEK293T cells, the MG160 candidate tethered to spCas9(H840A) was tested for the conversion of G to T on the VEGFA target ( Figures 29A - 29DD)。The editing percentages of each RT with pegRNAs of different PBS lengths (2, 4, 6, 8, 10, 13, 16, and 20 nucleotides; SEQ ID NO: 656 - 679) are shown in Figures 29A - 29DD Each editing level of each RT candidate represents a single biological replicate. MG160 - 473, MG160 - 283, MG160 - 379, MG160 - 395, and MG160 - 107 (SEQ ID NO: 1206, 1017, 1112, 1128, and 841) showed comparable or improved editing efficiency relative to the control spCas9 (H840A) tethered to MMLV WT (respectively Figure 29A , 29G , 29L, 29O, and 29CC). Additionally, the editing level of candidate MG160 - 473 (SEQ ID NO: 1206) was comparable to that of the control spCas9 (H840A) tethered to the hyperactive mutant MMLV (MMLV2, PE2) (SEQ ID NO: 1249; Figure 29A ). Further, candidates MG160 - 46, MG160 - 9, MG160 - 21, MG160 - 419, MG160 - 99, and MG160 - 279 (SEQ ID NO: 1251, 1265, 1270, 1271, 1274, and 1279) showed above - background activity (respectively Figure 29B , 29P , 29U, 29V, 29Y, and 29DD). Then five MG160 candidates with high G - to - T conversions were repeated to confirm the G - to - T conversions ( Figure 30A ), as well as their ability to perform 24 - nucleotide insertions ( Figure 30B ) and 15 - nucleotide deletions ( Figure 30C ) using chemically synthesized pegRNAs with PBS lengths ranging from 6, 8, 10, and 13 nucleotides (SEQ ID NO: 76 - 99). For all desired edits, MG160 - 283, MG160 - 379, MG160 - 395, and MG160 - 107 (SEQ ID NO 1017, 1112, 1128, and 841) showed editing levels similar to the control MMLV WT (SEQ ID NO: 1248), while candidate MG160 - 473 (SEQ ID NO: 1206) exhibited high editing levels comparable to the hyperactive mutant MMLV2 (SEQ ID NO: 1249) for G - to - T conversions and 24 - nucleotide insertions.

[0506] Test reverse transcriptase candidates not tethered to nickase

[0507] Reverse transcriptase (RT) candidates from different retrotransposon families MG155, MG156, MG157, MG159, and MG173, as well as from the MG group II intron families MG164, MG166, MG167, and MG169 (SEQ ID NO: 1280 - 1294), were cloned into plasmids with a CMV promoter driving RT expression. The plasmids were then isolated for transfection in HEK293T cells. Another plasmid driven by the CMV promoter containing the nickase spCas9 (H840A) (SEQ ID NO: 1247) and the plasmid containing the RT were co - transfected. The transfection contained a chemically synthesized pegRNA (SEQ ID NO: 656 - 679) that contained the desired edit in the RT template. All components (plasmids and pegRNA) were reverse - transfected into 50,000 HEK293T cells in a 24 - well plate. Seventy - two hours after transfection, the cells were lysed in 100 μL of solution. Primers containing barcodes for NGS (SEQ ID NO: 100 - 101) were used to amplify a ~250 bp target (SEQ ID NO: 102). Then PCR purification was performed, and the samples were sent for NGS sequencing. The FASTQ files were then processed using the prime editing setup to determine the percentage of reads with the desired change.

[0508] Results

[0509] pegRNAs with different PBS lengths (2, 4, 6, 8, 10, 13, 16, and 20 nucleotides; SEQ ID NO: 76 - 83; Figures 31A - 31K ) were used to test the G - to - T change of untethered retrotransposon candidates from families MG155, MG156, MG157, MG159, and MG173. In a single biological replicate experiment, the retrotransposon RT candidates MG173 - 1 (SEQ ID NO: 627; Figure 31J ) and MG173 - 2 (SEQ ID NO: 628; Figure 31K ) showed editing levels reaching approximately 4.5% and 1.5% at PBS 8 and PBS13, respectively, and had above - background editing levels at various PBS lengths. For MG155 - 3 (SEQ ID NO: 504; Figure 31A ), MG155 - 5 (SEQ ID NO: 506; Figure 31C ), and MG156 - 1 (SEQ ID NO: 507; Figure 31D ), above - background editing levels were also observed.

[0510] The untethered group II intron families MG164, MG166, MG167, and MG169 were tested to haveFigures 32A - 32D The editing levels shown in these candidates. Most of these candidates did not show detectable activity, but for MG169-1 (SEQ ID NO: 601) with different PBS lengths, some editing above background was observed ( Figure 32D ).

[0511] Example 14. Short corrections, small insertions, and deletions using engineered RT

[0512] Editing using engineered MG160-4 and MG153-53 RT candidates

[0513] The selected RT candidates MG160-4 (SEQ ID NO: 521) and MG153-53 (SEQ ID NO: 496) were rationally engineered to improve editing efficiency. Various point mutations (SEQ ID NO: 1221 - 1243) were tested individually and in combination to determine which engineered candidates could improve editing activity. Using chemically synthesized pegRNAs with PBS lengths of 6, 8, 10, and 13 nucleotides, different combinations of tethered (MG160-4) or untethered (MG153-53) to spCas9 (H840A) MG160-4 and MG153-53 mutants were tested for the G-to-T conversion on the VEGFA target. Primers containing barcodes for NGS (SEQ ID NO: 100 - 101) were used to amplify a ~250 bp target (SEQ ID NO: 102). Then PCR purification was performed, and the samples were sent for NGS sequencing. The FASTQ files were then processed to determine the percentage of reads with the desired change. Individual biological replicates were tested along with untethered controls MMLV1 and MMLV2 (SEQ ID NO: 1248 and 1249) and control RTs TGIRT, Marathon, and Marathon mutants (SEQ ID NO: 1296 - 1298).

[0514] Results

[0515] Combining four different point mutations in various combinations of MG160-4 led to a sharp decrease in editing efficiency ( Figure 33A ). However, the single-point mutations H230K (MG160-4 H230K; SEQ ID NO: 1230) and H230R (MG160-4 H230R; SEQ ID NO: 1234) showed neutral changes in the editing levels of the G-to-T conversion ( Figure 33A ). The G-to-T transversions of these engineered constructs were further tested ( Figure 33B ), as well as 24-nucleotide insertions ( Figure 33C) and a 15-nucleotide deletion ( Figure 33D ). As observed in the first replicate, MG160-4-H230K and MG160-4H230R showed a neutral change in the editing level of a G-to-T transversion ( Figure 33B ), but compared to wild-type MG160-4, for the 24-nucleotide insertion ( Figure 33C ) and deletion ( Figure 33D ), the editing level of MG160-4 H230R (SEQ ID NO:1230) increased. When the editing involved incorporation of the 24-nucleotide insertion and 15-nucleotide deletion, MG160-4 H230R (SEQ ID NO:1234) showed a slightly improved editing compared to engineered MG160-4-H230K (SEQ ID NO:1230).

[0516] The construct combining all the proposed mutations of MG153-53 (SEQ ID NO:1226) eliminated the editing activity ( Figure 34 ). Compared to WT MG153-53 (SEQ ID NO:496; Figure 34 ), the single-point mutation V200R (MG153-53(V200R); SEQ IDNO:1225) slightly enhanced the G-to-T transversion. WT MG153-53 (SEQ ID NO:496) did not perform better than controls MMLV1 and MMLV2 (SEQ ID NO:1248 and 1249), but did perform better than controls RT TGIRT, Marathon, and Marathon mutants (SEQ ID NO:1296-1298).

[0517] Example 15. Nickase for binding to reverse transcriptase to mediate short correction, small insertions, and deletions

[0518] Using RT to install site-directed genomic corrections, insertions, or deletions requires the RT system to be targetable. This example describes the use of a targetable RT system comprising an RT and a Cas nickase. The Cas nickase guided by the gRNA site specifically nicks the non-target strand, thereby generating a primer for the reverse transcription reaction. The gRNA accompanying the Cas nickase is a modified version (pegRNA) containing a 3' extension containing RTT and PBS. PBS and the spacer are complementary to each other. It is expected that this complementarity can cause disruption of the gRNA structure, leading to interaction of the pegRNA with Cas, thereby inhibiting Cas from seeking the target gene. Each Cas nuclease interacts with its own gRNA, so pegRNA design and requirements vary depending on the system.

[0519] Testing selected reverse transcriptases untethered and tethered to the MG3-6(H586A) and MG71-2(H883A) nickases

[0520] Challenge the MG3-6(H586A) (SEQ ID NO:653) or MG71-2(H883A) (SEQ ID NO:1309) nickase to introduce genomic correction at the AAVS1 target site (SEQ ID NO:654 or 1344) with a reverse transcriptase. Transfect the selected MGRT candidates (SEQ ID NO:1295 and 1299 - 1304) into HEK293T cells that are untethered Figure 35 ) from the MG3-6(H586A) (SEQ ID NO:653) plasmid, or tethered to MG3-6(H586A) (SEQ ID NO:653), where the selected RT is fused to the N-terminus or C-terminus Figure 35 ). For MG3-6(H586A) (SEQ ID NO:653) genomic correction, chemically synthesized pegRNAs (SEQ ID NO:682 - 684 and 686) with PBS lengths of 8, 10, 13, and 20 nucleotides are targeted, while for MG71-2(H883A) (SEQ ID NO:1309) genomic correction, chemically synthesized pegRNAs (SEQ ID NO:1310 - 1341) with PBS lengths of 6, 8, 10, 13, 16, and 20 nucleotides are targeted. Amplify the ~250 bp MG3-6(H586A) AAVS1 target (SEQ ID 654) or MG71-2(H883A) AAVS1 target (SEQ ID NO:1344) using primers containing barcodes for NGS (SEQ ID NO:698 - 699 for MG3-6(H586A) (SEQ ID NO:653) or SEQ ID NO:1342 - 1343 for MG71-2(H883A) (SEQ ID NO:1309)). Then perform PCR purification and send the samples for NGS sequencing. Then process the FASTQ files using the prime editing setup to determine the percentage of reads with the desired change.

[0521] Results

[0522] For the selected RT candidates for the G-to-T transversion, above-background editing (>0.1%) was observed at PBS lengths of 8, 10, 13, and 20 nucleotides Figure 35)。Interestingly, when MG3-6(H586A) (SEQ ID NO:653) and MMLV WT (MMLV1) (SEQ ID NO:1248) are tethered to the C-terminus of MG3-6(H586A), the editing level is generally higher than that in the untethered method or when the RT is tethered to the N-terminus of the nickase. When MG3-6(H586A) (SEQ ID NO:653) and the hyperactive mutant MMLV (MMLV2) (SEQ ID NO:1249) are tethered to the C-terminus of MG3-6(H586A), the editing levels are similar to those in the untethered method. In contrast, the MG151-98 engineered mutants (SEQ ID NO:1302-1304) produce higher levels of editing when tethered to the N-terminus of MG3-6(H586A) (SEQ ID NO:653) or in the untethered method ( Figure 35 )。When MG160-4 (SEQ ID NO:1295) is tethered to the N-terminus or C-terminus of MG3-6(H586A) (SEQ ID NO:653), similar editing levels are obtained, but there is no editing above background in the untethered method. All three different methods of MG153-53 (SEQ ID NO:1299) with MG3-6(H586A) (SEQ ID NO:653) show no editing activity above the background level ( Figure 35 )。

[0523] The untethered MG71-2(H883A) (SEQ ID NO:1309) with the selected RT shows editing levels for various edits, including five nucleotide changes ( Figures 36A - 36C and 36J), a nucleotide transversion of a single G to T ( Figure 36D and 36G ), a 24-nucleotide insertion ( Figure 36E and 36I ), and a 15-nucleotide deletion ( Figure 36F and 36H ). Biological triplicate data for correcting five nucleotide changes in the AAVS1 target (SEQ ID NO:1344) are shown with the selected RT. The untethered MMLV1 and MMLV2 (SEQ ID NO:1248 and 1249) with MG71-2(H883A) (SEQ ID NO:1309) show high levels of editing for all corrections ( Figures 36A - 36J ). MG153-53 (SEQ ID NO:1299) shows editing above background only when attempting to correct a 15-nucleotide deletion ( Figure 36F ). When comparing Figure 36B and Figure 36CWhen, the pegRNA scaffold changes from four consecutive Ts to a modified scaffold with four consecutive Gs. There is no significant change in the editing level between the original scaffold and the modified scaffold, so the original scaffold remains unchanged when correcting other changes (insertions, deletions, and SNPs). Interestingly, the editing level for correcting five nucleotide changes ( Figure 36B ) is higher than that for a transversion of a single G to T ( Figure 36D ). Generally, when the pegRNA PBS length is between 8 and 16 nucleotides, MG71-2(H883A) (SEQ ID NO:1309) and the selected RTs (SEQ ID NO:1295 and 1299-1301) show the highest editing levels for all corrections. Then the engineered MG151-98 candidates (SEQ ID NO:1302-1304) were tested with untethered MG71-2(H883A) (SEQ ID NO:1309) to correct the AAVS1 target (SEQ ID NO:1344; Figures 36G - 36J ). All the MG151-98 engineered candidates (SEQ ID NO:1302-1304) showed editing levels comparable to MMLV1 and MMLV2 (SEQ ID NO:1248 and 1249) for all corrections.

[0524] Both MG3-6(H586A) (SEQ ID NO:653) and MG71-2(H883A) (SEQ ID NO:1309) showed to be effective nickases, which are compatible with RT reverse transcription to integrate small corrections into genomic targets.

[0525] Example 16. Retrotransposon RT for Single-Stranded DNA Transposase (TnpA)-Mediated Programmable Large Cargo Integration

[0526] A retrotransposon is a DNA retroelement containing a reverse transcriptase (RT) gene downstream of a conserved non-coding structural RNA. The non-coding RNA consists of two inverted regions (referred to as msr and msd). When the retrotransposon RT recognizes the msr folded into a specific secondary structure (specific recognition motif), it initiates the reverse transcription of the msd portion (template), thereby generating multiple copies of single-stranded DNA (ssDNA). Overall, the reverse transcriptase has RT activity, is initiated by a specific RNA recognition motif (msr), and produces covalently bound complementary ssDNA molecules. Thus, the dependence on the recognition motif in the mrs should reduce off-target initiation and provide a mechanism to localize the template RNA / DNA to specific genomic targets.

[0527] Precise genome editing, with scarless replacement of alleles or insertion of synthetic sequences, requires in vitro delivery of donor DNA. However, there are many challenges in inducing cells to use donor DNA for homology-directed repair (HDR). To this end, it is possible to use retrotransposons to generate high-copy-number intracellular DNA molecules in the human host. Early experiments showed that the msd could be variable and could encode in situ DNA with the artificial sequence of interest. Therefore, retrotransposons can serve as a source of donor DNA for genome engineering. This biological solution can enable nuclear donor generation, thus improving the scalability and multiplexing ability of genome knock-in. Recently, experiments have shown that coupling retrotransposons with Cas9 improves the efficiency of precise genome editing by HDR in HEK293T and K563, where the HDR rate is as high as about 11%. Although these findings represent the first step in retrotransposon-based gene editing in human cells, low editing efficiency remains a challenge due to the limitations of HDR in non-cycling cells. Coupling retrotransposon-Cas9-like fusions with ssDNA integrases such as the ssDNA transposase TnpA can avoid dependence on the HDR pathway and improve DNA integration. For example, the IS200 / IS605 transposon is a mobile genetic element that integrates ssDNA at specific target sites by the TnpA transposase. TnpA excises the donor by recognizing the structural motifs at each donor end, integrates it into the recognized target site, and can be used as ssDNA.

[0528] The ssDNA generated by retrotransposon RT can be used by TnpA as a template for programmable integration of the desired cargo into specific target sites ( Figure 38 ). Specifically, the retrotranscribed msd contains the desired cargo (e.g., an antibiotic resistance cassette or a fluorescent marker) flanked by the LE and RE motifs recognizable by TnpA. The TnpA transposase excises and circularizes the ssDNA donor and integrates it into the target by recognizing specific motifs obtained through the R-loop formed by RNA-guided recognition and binding of engineered (nickase or dead) effectors (e.g., MG3-6).

[0529] Example 17. Engineering ncRNA-related retrotransposon RT to include LE, RE, and the cleavage motif of TnpA

[0530] Engineering ncRNA containing the LE / RE motifs of Hp TnpA with Ec86 retrotranscription

[0531] Using the cryo-EM structure of Ec86 complexed with its product and previous work, eight engineered Ec86 ncRNA variants (SEQ ID NOs: 1346 - 1353; Figures 39 - 40) were designed. The previous work had identified replaceable regions of the Ec86 msd stem-loop to facilitate homologous directed repair. Three different lengths of insertion sequences were designed, all of which contained the reverse complements of the left end (LE) and right end (RE) recognition motifs of the Helicobacter pylori (Hp) TnpA ssDNA transposase at the 3' and 5' flanking regions. The insertion sequence named LE40RE contained a 40 nt sequence flanked by the LE and RE of Hp TnpA, resulting in a total insertion length of 174 nt. The insertion sequences named LE200RE and LE500RE contained a partial kanamycin gene of 200 nt or 500 nt flanked by the LE / RE motifs, resulting in total insertion lengths of 334 nt and 634 nt, respectively. These three different sequences were inserted at two or three different potential replaceable regions ( Figure 40 ) within the msd stem-loop. The version named Version 1 replaced the entire msd region that did not resolve in the cryo-EM structure of Ec86 and bound to its msdDNA with the engineered sequence. Versions 2 and 3 were more progressive and conservative replacement designs, where Version 2 replaced the msd region after the bubble in the msd stem-loop, and Version 3 retained most of the msd stem-loop with the terminal eight nucleotides.

[0532] The reverse complements of the LE and RE motifs of Hp TnpA were predicted to adopt different secondary structures within the engineered ncRNA ( Figure 40 ). However, the predicted RNA folding did show that the key recognition motifs required for Ec86 RT recognition (including the terminal inverted repeat and the msr hairpin) were maintained, suggesting that the LE / RE motifs of TnpA may not disrupt the folding of the key ncRNA features required for priming.

[0533] In vitro determination of reverse transcription of Ec86 on engineered ncRNAs. In a cell-free expression system supplemented with dNTP (final 0.3 mM), Ec86 RT was co-expressed with ncRNA substrates (final 100 nM). The expression constructs were codon-optimized for E. coli and contained an N-terminal single streptococcal tag. After incubation at 37 °C for 2 hours, the reaction was quenched by heat denaturation at 95 °C for 2 minutes and then treatment with RNase A at 37 °C for 30 minutes. Ec86 activity was evaluated by qPCR using primers (SEQ ID NO: 1354 - 1355) that amplified products generated from wild-type ncRNA (SEQ ID NO: 1345), or products generated from engineered 40 nt partial kanamycin genes (SEQ ID NO: 1356 - 1357) or 200 nt and 500 nt partial kanamycin genes (SEQ ID NO: 1358 - 1359). Prior to qPCR, the resulting reverse transcription products (referred to herein as msdDNA) were diluted to ensure that the msdDNA concentration was within the linear detection range. The amount of msdDNA was quantified by extrapolation from a standard curve generated with DNA templates of known concentration. Based on these results, Ec86 RT was able to generate appreciable amounts of msdDNA from all eight engineered ncRNA designs, and the levels were comparable to those of wild-type ncRNA( Figure 41 ). This data indicates that Ec86 is tolerant to insertions up to 634 nt at three different alternative regions within the msd stem-loop.

[0534] Insertion of Hp TnpA into ssDNA generated by retrotransposon RT Ec86

[0535] Determination of ssDNA insertion by the retrotransposon of Hp TnpA. Briefly, Ec86 co-expressed engineered ncRNA substrates (LE200RE_v1 / v3 or LE500RE_v1 / v2 / v3) in a cell-free expression system as described above, followed by quenching through heat denaturation and RNase A treatment, also as described above. RNase A treatment removed any RNA in the heteroduplex formed with the generated msdDNA, thereby making the product available as ssDNA for TnpA. Subsequently, the generated ssDNA containing the LE / RE motif of Hp TnpA was mixed with Hp TnpA protein, which was also generated in a cell-free expression system in a reaction buffer containing 20 mM HEPES (pH 7.5), 160 mM NaCl, 5 mM MgCl2, 5 mM TCEP, 20 μg / mL BSA, 0.5 μg / mL poly-dIdC, and 20% glycerol. The reaction also contained 50 nM of an ssDNA insertion target including the Hp TnpA targeting motif (TTAC). The TnpA insertion reaction was carried out at 37 °C for 1 hour, after which the successful insertion of TnpA was confirmed by PCR using a primer pair annealing to the chimeric product (expected amplicon size of approximately 300 bp) with a partial kanamycin gene cargo and the ssDNA target (SEQ ID NOs: 1360 - 1361). Sanger sequencing further confirmed the insertion. Based on these results, Hp TnpA can insert into ssDNA generated by Ec86 from all 5 tested engineered ncRNAs (LE200RE_v1 / v3 and LE500RE_v1 / v2 / v3) in an RT-dependent and TnpA-dependent manner ( Figures 42 - 43 ).

[0536] Tolerance of the MG154 - 159 and MG173 families to insertion within the msd of ncRNA

[0537] Based on the predicted secondary structure of the ncRNA, the msd stem-loop was identified as the first 3' hairpin adjacent to the inverted repeat. Alternative regions of one or two versions of the msd were identified and a ~200 nt sequence encoding a partial kanamycin gene was inserted ( Figures 44 - 51 ; SEQ ID NOs: 1362 - 1393). For the indicated cases, trimmed and untrimmed versions of the ncRNA were also designed and tested ( Figure 46)。To evaluate whether the retrotransposons are tolerant to insertions within the MSD stem-loop, the corresponding retrotransposon RT was co-expressed with the engineered ncRNA in a cell-free expression system supplemented with dNTPs, and then heat-denatured and treated with RNase A as described above. The resulting MSD DNA was then diluted prior to qPCR to ensure that the concentration was within the linear detection range. qPCR was performed using primers that amplified a portion of the kanamycin sequence. The amount of MSD DNA was quantified by extrapolation from a standard curve generated using DNA templates of known concentration. A retrotransposon RT was considered active if the MSD DNA yield was more than 10-fold that of the RT-free background control. Based on these results, the following retrotranscription systems were tolerant to MSD insertions( Figure 52 ):MG155-2, MG155-3, MG155-4, MG155-5, MG156-1, MG156-2, MG157-1, MG157-3, MG157-4, MG157-5, MG158-1, MG159-1, MG159-2, MG159-3, MG173-1 and MG173-2.

[0538] Example 18. RTs for short corrections, small insertions and deletions

[0539] Reverse transcriptase candidates not tethered or tethered to the MG71-2 (H883A) nickase

[0540] The RT candidates (SEQ ID NOs: 1234, 1249 - 1250, and 1304) in the tethered system were cloned into a plasmid containing the nickase MG71-2 (H883A) (SEQ ID NO: 1309) to generate RT-nickase fusions (RT on the C-terminus or N-terminus of MG71-2 (H883A)). The CMV promoter drives the expression of the fusion protein, which contains a 33-amino acid linker (SEQ ID NO: 103) between the nickase and the RT candidate. The fusion protein was then transfected into HEK293T cells using liposomes. In the untethered system, RT was cloned into a plasmid with a CMV promoter driving RT expression. Another plasmid containing the nickase MG71-2 (H883A) driven by the EF1ɑ promoter and the plasmid containing RT were co-transfected using liposomes. Liposome transfection targeting AAVS1 was used to introduce chemically synthesized pegRNAs (SEQ ID NOs: 1310 - 1315) containing the desired edits in the RT template. All components (plasmids and pegRNAs) were reverse transfected into 50,000 HEK293T cells in a 24-well plate. Seventy-two hours after transfection, the cells were lysed. A primer containing a barcode for next-generation sequencing (NGS) (SEQ ID NOs: 1342 - 1343) was used to PCR amplify a ~250-bp target (SEQ ID NO: 1344). The samples were purified and sequenced. The sequencing data were then processed to determine the percentage of reads with the desired changes.

[0541] Engineered MG151-98 (K297P, Δ166AA) (SEQ ID NO: 1304) and MMLV2 (SEQ ID NO: 1249) were tested without or tethered to MG71-2 (H883A) (RT on the C-terminus of MG71-2 (H883A) (nickase-RT) or RT on the N-terminus of MG71-2 (H883A) (RT-nickase)) ( Figures 53A - 53B ). MG160-4 (H230R) (SEQ ID NO: 1234) and MG160-473 (SEQ ID NO: 1250) were tested when tethered to MG71-2 (H883A) (RT on the C-terminus of MG71-2 (H883A) (nickase-RT) or RT on the N-terminus of MG71-2 (H883A) (RT-nickase)) ( Figures 53C - 53D)。Challenge RT to incorporate 5 nucleotide changes at the AAVS1 target (SEQ ID NO:1344). RT was transfected with pegRNAs (SEQ ID NO:1310 - 1315) of different PBS lengths, and the data shown in Figure 53 represent the highest editing levels for each RT under each RT nickase configuration. Compared to MMLV2 tethered to the C-terminus of MG71 - 2(H883A), MG71 - 2(H883A) without tethered MMLV2 and MMLV2 tethered to the N-terminus of MG71 - 2(H883A) showed the highest levels of editing( Figure 53A )。For MG151 - 98(K297P, Δ166AA) (SEQ ID NO:1304)( Figure 53B ), similar results were shown. Previous studies have shown that candidates of the MG160 family show little or no activity in the untethered system. Engineered MG160 - 4(H230R) (SEQ ID NO:1234) and MG160 - 473 (SEQ ID NO:1250) were tested when tethered to the N-terminus or C-terminus of MG71 - 2(H883A). MG160 - 4(H230R) tethered to the N-terminus of MG71 - 2(H883A) produced a significantly higher editing level than when tethered to the C-terminus of MG71 - 2(H883A)( Figure 53C )。When MG160 - 473 was tethered to the N-terminus of MG71 - 2(H883A), the highest editing level was also shown( Figure 53D )。The data shown in Figure 53 represent "correct editing", indicating that no unexpected corrections were found in the NGS amplicons. "Incorrect editing" refers to the incorporation of the expected edit but includes errors within the NGS amplicon, or errors resulting from incorrect incorporation of RT and / or scaffold incorporation of the pegRNA. The data show that MG71 - 2(H883A) has a strong preference for RT at the N-terminus. Further, MGRT has been shown to be superior to literature controls in terms of efficiency and accuracy.

[0542] Example 19. RTs for short corrections, small insertions, and deletions

[0543] Test reverse transcriptase candidates not tethered to the spCas9(H840A) nickase

[0544] The RT candidates (SEQ ID NOs: 1394 - 1402) in the untethered system were cloned into a plasmid with a CMV promoter driving RT expression. Another plasmid driven by the EF1ɑ promoter containing the nickase spCas9 (H840A) (SEQ ID NO: 1247) and the plasmid containing RT were co-transfected by lipid transfection. A chemically synthesized pegRNA (SEQ ID NOs: 76 - 83) containing the desired edit in the RT template targeting VEGFA (SEQ ID NO: 102) was transfected by high-efficiency lipid transfection. All components (plasmids and pegRNA) were reverse transfected into 50,000 HEK293T cells in a 24-well plate. 72 hours after transfection, the cells were lysed in 100 μL of DNA extraction solution. Primers containing barcodes for next-generation sequencing (NGS) (SEQ ID NOs: 100 - 101) were used to amplify the ~250 bp target (SEQ ID NO: 102) with a high-fidelity polymerase and reaction solution. Then PCR purification was performed, and the samples were sent for NGS sequencing. The FASTQ files were then processed to determine the percentage of reads with the desired change.

[0545] Using pegRNAs with PBS lengths varying from 2 to 20 nucleotides and untethered spCas9 (H840A) (SEQ ID NO: 1247), eight candidates from the MG173 family (SEQ ID NOs: 1394 - 1401) and one candidate from the MG192 family (SEQ ID NO: 1402) were tested for the G-to-T transversion on the VEGFA target (SEQ ID NO: 102) (Figure 54). Reverse transcriptase candidates with editing levels above background (>0.1%) included MG173-3 (SEQ ID NO: 1394), MG173-8 (SEQ ID NO: 1399), MG173-9 (SEQ ID NO: 1400), and MG173-10 (SEQ ID NO: 1401), while the other reverse transcriptase candidates (SEQ ID NOs: 1395 - 1398 and 1402) were inactive against the G-to-T transversion ( Figure 54A ). The editing percentages were then further broken down to determine "correct editing", "incorrect editing", "editing", and "scaffold incorporation" ( Figures 54B - 54S)。"Correct editing" refers to the expected editing without errors in the NGS amplicon, while "incorrect editing" refers to the incorporation of the expected editing and includes errors within the NGS amplicon, or incorrect incorporation stemming from RT and / or scaffold incorporation of the pegRNA. "Editing" refers to the expected editing with errors in the NGS amplicon (excluding pegRNA scaffold incorporation), while "scaffold incorporation" indicates the expected editing and scaffold incorporation of the pegRNA. MG173-8 (SEQ ID NO:1399) showed the highest editing level compared to other reverse transcription candidates ( Figure 54A , 54G and 54P), with the highest percentage level of editing for 8 to 13 nucleotides of the PBS (SEQ ID NO:79-81).

[0546] Testing reverse transcriptase candidates tethered to the spCas9(H840A) nickase

[0547] The RT candidates (SEQ ID NO:1403-1424) in the tethered system were cloned into a plasmid containing the nickase spCas9(H840A) (SEQ ID NO:1247) to generate RT-nickase fusions. The CMV promoter drives the expression of the fusion protein, which contains a 33-amino acid linker (SEQ ID NO:103) between the nickase and the RT candidate. Transfection of these constructs, along with chemically synthesized pegRNA, followed the above transfection protocol as well as NGS sample preparation and data analysis.

[0548] In the case of tethering to spCas9(H840A) (SEQ ID NO:1247), twenty-two MG160 candidates (SEQ ID NO:1403-1424) were tested for the G-to-T transversion on the VEGFA target (SEQ ID NO:102) across eight different pegRNAs with different PBS lengths (SEQ ID NO:76-83) ( Figure 55A ). Candidates that did not show activity above background (>0.1%) under the test conditions were MG160-50 (SEQ ID NO:1409) ( Figure 55O and 55AK ), MG160-114 (SEQID NO:1404) ( Figure 55E and 55AA ), MG160-210 (SEQ ID NO:1412) ( Figure 55H and 55AD ), MG160-306 (SEQID NO:1418) ( Figure 55U and 55AQ)、MG160-416 (SEQ ID NO:1422) ( Figure 55M and 55AI ) and MG160-483 (SEQ ID NO:1424) ( Figure 55W and 55AS ). MG160 candidates with high activity levels against the G-to-T transversion include MG160-45 (SEQ ID NO:1423) ( Figure 55D and 55Z ), MG160-121 (SEQ ID NO:1405) ( Figure 55F and 55AB ), MG160-136 (SEQ ID NO:1407) ( Figure 55G and 55AC ), MG160-193 (SEQ ID NO:1410) ( Figure 55R and 55AN ), MG160-232 (SEQ ID NO:1407) ( Figure 55J and 55AF ) and MG160-358 (SEQ ID NO:1419) ( Figure 55V and 55AR ). MG160-136 (SEQ ID NO:1407) achieved an editing level of over 5% for PBS lengths of 6 - 20 nucleotides (SEQ ID NO:78 - 83), with the highest editing level under PBS 8 (SEQ ID NO:79) reaching approximately 15% for the G-to-T transversion ( Figure 55G and 55AC ). The editing percentages were then further broken down to determine "correct editing", "incorrect editing", "editing", and "scaffold incorporation" (terms described in detail above) ( Figures 55B - 55AS ).

[0549] Example 20. Short corrections, small insertions, and deletions using engineered RT

[0550] Testing engineered reverse transcriptase candidates not tethered or tethered to the spCas9(H840A) nickase

[0551] The selected RT candidates were subjected to rational engineering to improve editing efficiency. Various point mutations were tested individually and in combination to determine which engineered candidates could enhance editing activity. The selected RT candidates and engineered mutants (MG151-98 (SEQ ID No: 1300 and 1302-1304), MG151-123 (SEQ ID NO: 715 and 1426-1431), MG151-126 (SEQ ID NO: 718 and 1433-1438), MG153-18 (SEQ ID No: 55 and 1439-1441), and MG153-20 (SEQ ID No: 57 and 1442-1444)) were tested without tethering to the spCas9(H840A) (SEQ ID NO: 1247) linkage, while MG160-473 (SEQ ID NO: 1250) and mutants (SEQ ID No: 1445-1446) were tested with tethering to spCas9(H840A) (SEQ ID NO: 1247). Using chemically synthesized pegRNAs with different PBS lengths and RTTs (SEQ ID NO: 78-81, 86-90, and 94-98), the engineered reverse transcriptases were challenged for multiple edits (transversions, insertions, and deletions) of the VEGFA target (SEQ ID NO: 102). The engineered reverse transcriptases were tested without tethering to or with tethering to spCas9(H840A) (SEQ ID NO: 1247) using the same transfection protocol as described in Example 19, as well as NGS preparation and data analysis.

[0552] Using pegRNAs with PBS lengths of 6, 8, 10, and 13 nucleotides (SEQ ID NO: 78-81, 86-90, and 94-98), the wild-type MG151-98 (SEQID NO: 1300) and engineered mutants MG151-98 (Δ166AA) (SEQ ID NO: 1302), MG151-98 (H171N, Δ166AA) (SEQ ID NO: 1303), and MG151-98 (K297P, Δ166AA) (SEQ ID NO: 1304) were tested without tethering to spCas9(H840A) (SEQ ID NO: 1247) for the transversion of G to T ( Figure 56A and 56D ) on the VEGFA target (SEQ ID NO: 102), 24-nucleotide insertions ( Figure 56B and 56E ) and 15-nucleotide deletions ( Figure 56C and 56F)。In three different types of editing, trimming 166 amino acids from the C-terminus of MG151-98 (MG151-98(Δ166AA)(SEQ ID NO:1302)) resulted in no significant difference in the editing level compared to the wild type (SEQ ID NO:1300) (Figure 56). Further, compared to the wild type MG151-98 (SEQ ID NO:1300), the single point mutants H171N and K297P in combination with the 166AA trimming of the C-terminus of reverse transcriptase (SEQ ID No:1303-1304) enhanced the editing and made the editing level higher than that of MMLV1 (SEQ ID NO:1248), and comparable to that of MMLV2 (SEQ IDNO:1249) for some types of editing (Figure 56). The editing percentages were further broken down to determine "correct editing", "incorrect editing", "editing", and "scaffold incorporation" (terms described in detail in Example 19). Even though the engineered mutants of MG151-98 (SEQ ID No:1302-1304) increased the editing level, there was no significant difference in "incorrect editing" and "scaffold incorporation" compared to the wild type (Figure 56).

[0553] Using pegRNAs (SEQ ID NO:78-81, 86-90, and 94-98) with PBS lengths of 6, 8, 10, and 13 nucleotides, the wild type and engineered mutants of MG151-123 (SEQ ID NO:715, 1426-1431), MG151-126 (SEQ ID NO:718 and 1433-1438), MG153-18 (SEQ ID No:55 and 1439-1441), and MG153-20 (SEQ ID No:57 and 1442-1444) were tested for the G to T transversion on the VEGFA target (SEQ ID NO:102) (Figure 57). Compared to the wild type MG151-123 (SEQ ID NO:715), the MG151-123 mutant MG151-123(H178N)(SEQ ID NO:1429) showed increased editing levels at PBS 8 and PBS10( Figure 57A and 57E ). Other point mutations M304R, H287F, H178R, G279R, and G279N of MG151-123 (SEQ ID No:1426-1428 and 1430-1431) significantly reduced or eliminated the activity against the G to T transversion( Figure 57A and 57E)。MG151-126 (SEQ ID NO:718) and the point mutations (SEQ ID No:1433-1438) showed much lower editing levels than MG151-123 (SEQ ID NO:715) and could not be compared with MMLV1 (SEQ IDNO:1248) or MMLV2 (SEQ ID NO:1249)( Figure 57B and 57F ). Further, for MG153-18 (SEQ IDNO:55) and MG153-20 (SEQ ID NO:57), when testing the G-to-T transversion, the single point mutations (SEQ ID NO:1439-1440 and 1442-1443) and double point mutations (SEQ ID NO:1441 and 1444) showed no editing levels above background( Figure 57C 、 57D 、57G and 57H).

[0554] Using pegRNAs with PBS lengths of 6, 8, 10, 13, and 16 nucleotides (SEQ ID NO:78-82, 86-90 and 94-98), the G-to-T transversion of wild-type MG160-473 (SEQ ID NO:1250) and point mutants MG160-473 (F231R) (SEQ ID NO:1445) and MG160-473 (F231K) (SEQ ID NO:1446) on the VEGFA target (SEQ ID NO:102) was tested( Figure 58A 、 58D 、58G and 58J), 24-nucleotide insertions( Figure 58B 、 58E 、58H and 58K) and 15-nucleotide deletions( Figure 58C 、 58F 、58I and 58L). Compared with wild-type MG160-473 (SEQ ID NO:1250), the single point mutations (SEQ ID NO:1445-1446) did not result in increased editing levels. MG160-473 (SEQ ID NO:1250) was superior to tethered MMLV1 (SEQ ID NO:1248) in all editing aspects and showed comparable transversion editing levels compared with MMLV2 (SEQ ID NO:1249) (Figure 58).

[0555] Example 21. Nicking Enzymes for Binding to Reverse Transcriptase to Mediate Short Corrections, Small Insertions, and Deletions

[0556] Using RT to install genomic corrections, insertions, or deletions requires the system to be targeted. The targeting of the system is achieved by using a Cas nickase. The Cas nickase cleaves the non-target strand, generating a primer for reverse transcription. The gRNA accompanying the Cas nickase is a modified version (pegRNA) consisting of a 3' extension containing RTT and PBS. Complementarity between the PBS and the spacer can lead to disruption of the gRNA structure, causing the pegRNA to interact with Cas and thus inhibiting Cas from finding the target gene. Since each Cas nuclease interacts with its own gRNA, pegRNA design and requirements vary depending on the system.

[0557] Optimized MG71-2(H883A) nickase with MG reverse transcriptase

[0558] The selected MG reverse transcriptase candidates were challenged against the MG71-2(H883A) nickase (MG71-2n) (SEQ ID NO: 1309) to introduce genomic corrections (five nucleotide changes, a G-to-T transversion, a 24-nucleotide insertion, and a 15-nucleotide deletion) at the AAVS1 target site (SEQ ID NO: 1344) (Figures 59 - 61). The reverse transcriptases were tested without tethering to MG71-2n, tethered to the C-terminus of MG71-2n, or tethered to the N-terminus of MG71-2n with a 33AA linker (SEQ ID NO: 103). A procedure similar to the transfection and preparation protocol for the NGS samples described in Example 19 was used, except with different pegRNAs having PBS lengths of 6 to 20 nucleotides (SEQ ID NO: 1310 - 1315 and 1324 - 1341) and NGS primers (SEQ ID No: 1342 - 1343) to target the AAVS1 site with MG71-2n. Optimization of the pegRNA by modifying the scaffold of the pegRNA and incorporating mismatches in the PBS sequence was tested to determine if the editing level could be increased.

[0559] Across six different PBS lengths (6, 8, 10, 13, 16, or 20 nucleotides) containing a reverse transcription template (RTT) encoding five nucleotide changes, the selected reverse transcriptases MMLV1 (SEQ ID NO: 1248; Figure 59A and 59D )、MMLV2 (SEQ ID NO: 1249; Figure 59B and 59E )、MG160-4 (SEQ ID NO: 1295; Figure 59C and 59F )、MG151-98(Δ166AA) (SEQ ID NO: 1302; Figure 59G and 59J)、MG151-98(H178N, Δ166AA) (SEQ ID NO:1303; Figure 59H and 59K )、MG151-98(K297P, Δ166AA) (SEQ ID NO:1304; Figure 59I and 59L )、MG160-4(H230R) (SEQ ID NO:1234; Figure 59M and 59O ) and MG160-473 (SEQ ID NO:1250; Figure 59N and 59P ) were tested either untethered or tethered to MG71-2n. Generally, when using the tethering method with MG71-2n, the reverse transcriptase at the N-terminus of MG71-2n showed a higher editing level compared to the reverse transcriptase at the C-terminus of MG71-2n (Figure 59). Different reverse transcriptase candidates exhibited a preference for tethered or untethered formats (Figure 59). For example, the MG160 family candidates MG160-4, MG160-4(H230R), and MG160-473 showed much higher editing levels when tethered than in the untethered format ( Figure 59C , 59F and 59M-59P). In contrast, MG151-98(Δ166AA) and MG151-98(H178N, Δ166AA) showed higher editing levels when untethered to MG71-2n ( Figures 59G - 59L ), which may be due to the use of a non-optimal linker for MG151-98 (SEQ ID NO:1300). Generally, when targeting this region of AAVS1 with MG71-2n, the MG reverse transcriptase had fewer errors and scaffold incorporations than MMLV1 and MMLV2.

[0560] Then, MG160-4 and MG160-4(H230R) tethered to the N-terminus of MG71-2n were tested using pegRNAs with PBS lengths of 8, 10, 13, and 16 nucleotides to incorporate a G-to-T transversion, a 24-nucleotide insertion, a 15-nucleotide deletion, and five nucleotide changes at the AAVS1 target site ( Figures 60A - 60H)。MG160-4(H230R) is superior to or equivalent to wild-type MG160-4, depending on the desired correction. Additionally, MG160-4 and MG160-4(H230R) have comparable or improved editing levels compared to MMLV1 or MMLV2 tethered to the N-terminus of MG71-2n. The MG160 candidates were tested for all editing only when untethered under PBS13, and in all cases, the tethered MG160 candidates had higher activity when tethered than when untethered. Interestingly, when performing deletion correction, scaffold incorporation was much higher than for other types of editing ( Figure 60G )。However, when the reverse transcriptase was tethered to MG71-2n, scaffold incorporation appeared to decrease.

[0561] Using pegRNAs with PBS lengths of 8, 10, 13, and 16 nucleotides, engineered mutants of MG151-98(H178N, Δ166AA), MG151-98(K297P, Δ166AA), and MG151-98(H178N, K297P, Δ166AA) (SEQ ID NO:1447) untethered to MG71-2n showed successful editing of a G-to-T transversion, 24-nucleotide insertion, 15-nucleotide deletion, and five-nucleotide changes at the AAVS1 target site ( Figures 61A - 61L )。When compared to MMLV1 and MMLV2 untethered to MG71-2n, the engineered MG151-98 RT showed similar editing levels for all corrections ( Figures 61A - 61L )。When looking at the average median editing percentage for each correction, the single-point mutants MG151-98(H178N, Δ166AA), MG151-98(K297P, Δ166AA), and the double mutant MG151-98(H178N, K297P, Δ166AA) all had very similar editing levels ( Figures 61I - 61L )。Compared to other corrections, the 15-nucleotide deletion continued to have the largest amount of scaffold incorporation ( Figure 61G )。

[0562] The original guide RNA of MG71-2 contained a 107-nucleotide sequence (SEQ ID NO:1448) and a 24-nucleotide spacer. Two modified versions of the scaffold were designed: D2 (SEQ ID NO:1449) and D2C2 (SEQ ID NO:1450). The modified scaffold D2 removed the last hairpin in the scaffold, resulting in a scaffold length of 85 nucleotides. The modified scaffold D2C2 removed the last hairpin and the adjacent bulge of the original scaffold design, generating a 79-nucleotide modified scaffold. The editing levels of five nucleotide changes were tested using constructs MMLV2 or MG160-4 (H230R) tethered to the N-terminus of MG71-2n and modified pegRNAs (SEQ ID NO:1451-1458) with PBS lengths of 8, 10, 13, and 16 nucleotides (Figure 62). At PBS lengths of 10 and 13 nucleotides, the significant improvement in the increased editing levels of the two tethered constructs showed that the smaller modified scaffolds had higher editing levels (Figure 62). Further, the editing percentages analyzed by "correct editing" and "incorrect editing" ( Figure 62A ) and by "editing" and "scaffold incorporation" ( Figure 62B ) showed no significant change in the modified scaffold design compared to the original scaffold.

[0563] Due to the high complementarity between the PBS sequence and the spacer sequence of the pegRNA, incorporating mismatches in the PBS sequence may help promote higher editing levels of the desired editing. Modified mismatched pegRNAs of MG71-2n (SEQ ID NO:1459-1462) were designed to have eight nucleotides near the 3' end of the RTT and to match the target precisely in the nucleotide sequence. After these eight nucleotides, mismatches were incorporated to reach the next PBS length of the pegRNA (PBS10: 2 mismatches, PBS13: 5 mismatches, PBS16: 8 mismatches, and PBS20: 12 mismatches) (SEQ ID NO:1459-1462). When the PBS of the pegRNA contained mismatches ( Figure 63B and 63D ), MG71-2n and the selected untethered RTs (MMLV1, MMLV2, MG151-98 (H178N, Δ166AA), MG151-98 (K297P, Δ166AA), and MG151-98 (H178N, K297P, Δ166AA)) had significantly lower editing levels compared to the PBS sequence with perfect complementarity ( Figure 63A and 63C)。The same is true for the selected RTs (MMLV1, MMLV2, MG160-4, and MG160-4(H230R)) tethered to the N-terminus of MG71-2n( Figures 63E - 63H )。

[0564] Optimization of MG3-6(H586A) nickase with MG reverse transcriptase

[0565] To improve editing levels with the selected MG RTs and MG3-6(H586A) (SEQ ID NO:653), the scaffold and PBS sequences of the pegRNAs were modified to have different levels of GC content in the stem-loop of the scaffold and mismatches in the PBS sequence. In addition to the different pegRNAs (SEQ ID NO:112-113, 116, and 1463-1474) and NGS primers (SEQ ID NO:698-699), procedures similar to the transfection and NGS sample preparation protocols described above were used to target the AAVS1 locus (SEQ ID NO:654) with MG3-6n (SEQ ID NO:653).

[0566] The MG3-6 pegRNAs have four versions with modified scaffolds: modL1-4 (SEQ ID NO:1463-1470), which increase the G-C content on the first, second, and third hairpins, respectively, and modL1-modL3 (SEQ ID NO:1463-1465 and 1467-1469), and modL4 (SEQ ID NO:1466 and 1470), which combines the modifications of all three hairpins. The MG3-6 wild-type mRNA (SEQ ID NO:1475) was used to determine the percentage of modification (including SNPs and InDels) of the target amplicon AAVS1 (SEQ ID NO:654) in the NGS amplicons. The guide RNA (SEQ ID NO:116) achieved a modification percentage level of approximately 75%. The pegRNAs with PBS10 (SEQ ID NO:112) and PBS13 (SEQ ID NO:113) of the original MG3-6 scaffold achieved modifications of approximately 31% and 35%, respectively( Figure 64A )。The modification percentage levels of the pegRNAs with modifications, modL1, modL3, and modL4 (SEQ ID NO:1463, 1465-1466, 1467, and 1469-1470) decreased sharply, while the modification percentage level of modL2 (SEQ ID NO:1464 and 1468) improved slightly or remained the same as the pegRNAs containing the original scaffold design (SEQ ID NO:112-113) Figure 64A)。This also translates to the percentage of editing of two nucleotide changes in the AAVS1 target (SEQ ID NO: 654), which were measured at nucleotides 10 and 13 of the PBS length, where the original scaffold tethering MMLV2 (SEQ ID NO: 1249) to the C-terminus of MG3-6n (SEQ ID NO: 653) and the modified scaffolds modL1-modL4 (SEQ ID NO: 112-113 and 1463-1470) had an improved percentage of editing level of the modified scaffold modL2 (SEQ ID NO: 1464 and 1468). Figures 64B - 64C )。These results demonstrate the importance of the first and third hairpin sequences of the MG3-6 scaffold and suggest that future structural designs should avoid disrupting the first and third hairpins of the MG3-6 scaffold. The pegRNA was then modified to determine whether mismatches in the PBS sequence of the pegRNA could improve the editing level. Similar to the results observed with MG71-2n (Figure 63), when the pegRNA contained a mismatch in the PBS sequence (SEQ ID NO: 1471-1474), MG3-6n and the selected untethered RTs (MMLV1, MMLV2, MG151-98 (H178N, Δ166AA), and MG151-98 (K297P, Δ166AA)) showed a substantial decrease in the editing level. Figure 64D and 64E )。

[0567] The chimeras of MG3-6, MG3-6 / 3-8 (SEQ ID NO: 1476) were used to discover the target amplicons AAVS1 (SEQ ID NO: 654) Figure 65A ) and B2M (SEQ ID NO: 655 and 700-701) Figure 65BWhether the percentage of modification levels (including SNPs and InDels) can be increased. MG3-6 wild type (SEQ ID NO: 1475) and MG3-6 / 3-8 mRNA (SEQ ID NO: 1476) were used to direct InDels at the target with pegRNAs with guide RNAs and PBS lengths of 2, 4, 6, 8, 10, 13, 16, and 20 nucleotides (SEQ ID NOs: 109-124). For both the AAVS1 and B2M targets (SEQ ID NOs: 654-655), MG3-6 / 3-8 (SEQ ID NO: 1476) showed a higher level of on-target modification (including InDels) than MG3-6 (SEQ ID NO: 1475) (Figure 65). Generally, as the PBS length gets longer, both MG3-6 and MG3-6 / 3-8 have a decreasing percentage of InDels. However, MG3-6 / 3-8 has a higher InDel efficiency at specific targets and more effectively recognizes the target as the PBS length increases.

[0568] Discovery of an MG nickase bound to reverse transcriptase

[0569] The MG nuclease MG14-241 (SEQ ID NO: 1477) and the MG nickase MG14-241 (H596A) (MG14-241n) (SEQ ID NO: 1478) were tested to determine compatibility with the selected RTs for prime editing. Except for different pegRNAs (SEQ ID NOs: 1479-1492) and NGS primers (SEQ ID NOs: 1493-1504), procedures similar to the transfection and NGS sample preparation protocols described above were used to target multiple AAVS1 genomic loci (SEQ ID NOs: 1505-1510) with MG14-241 (SEQ ID NOs: 1477-1478).

[0570] Wild-type MG14-241 mRNA or plasmid (SEQ ID NO: 1477) was used to determine the percentage of modified (including SNPs and InDels) levels for various targets (G1, H1, B2, E2, F2, and G2) (SEQ ID NOs: 1505-1510). For each target with target E2 (AAVS1 region) (SEQ ID NO: 1508), different levels of InDels were observed, resulting in the highest level of InDels (up to approximately 60%) ( Figure 66A)。The mRNA of MG14-241 (SEQ ID NO:1477) was used to determine the percentage of modified (including SNPs and InDels) levels of the target amplicon E2 AAVS1 (SEQ ID NO:1508) of pegRNAs with guide RNAs (SEQ ID NO:1482) and PBS lengths of 2, 4, 6, 8, 10, 13, 16, and 20 nucleotides (SEQ ID NO:1485 - 1492). Figure 66B )。Similar to other MG nucleases, when using MG14-241 (SEQ ID NO:1477), the percentage of modification decreased as the PBS length increased. MG14-241n (SEQ ID NO:1478) and the selected untethered RTs (MMLV1, MMLV2, MG151-98 (H178N, Δ166AA), and MG151-98 (K297P, Δ166AA)) were used to determine the editing percentages of five nucleotide changes on the AAVS1 target (SEQ ID NO:1509) for all eight different PBS lengths (SEQ ID NO:1485 - 1492). Figures 66C - 66D )。For all selected RTs, the editing levels of five nucleotide changes were observed at a specific PBS length, where the untethered RTs showed the highest editing levels at PBS 8 and 10 for all selected RTs. For all selected RTs, the editing levels remained low, but further optimization of MG14-241n (SEQ ID NO:1478) and the pegRNAs could improve the editing efficiency at the selected targets.

[0571] Example 22. Site-Specific Integration of Large Cargo Templates by Non-LTR Retrotransposon RT and Group II Intron RT

[0572] Group II introns and non-LTR retrotransposases are capable of integrating large cargoes into target sites via reverse transcription of RNA templates. These reverse transcriptases (RTs) integrate RNA templates through target-primed reverse transcription (TPRT), a mechanism in which cDNA synthesis is primed by the free 3'-hydroxyl at the target DNA nick. Based on the presence of the predicted RT catalytic residues [F / Y]XDD, these enzymes were predicted to be active. To evaluate the ability of these RTs to work in conjunction with nucleases / nickases to produce programmable, site-specific integration of the cargo of interest (as opposed to their endogenous cargo), several RT-nuclease / nickase fusion constructs were designed. Additionally, various RNA templates were designed and tested for all RT-Cas fusion constructs to determine the combinations that would successfully result in targetable integration of the large cargo.

[0573] Large-Scale Site-Specific Genome Integration vi...

Claims

1. A fusion protein comprising a nickase linked to a reverse transcriptase using a linker, wherein the reverse transcriptase has at least about 80% sequence identity with any one of SEQ ID NOs: 161 - 629, 767 - 1220, 1959 - 2522, and 2582 - 2585.

2. A fusion protein comprising a nuclease linked to a reverse transcriptase using a linker, wherein the reverse transcriptase has at least about 80% sequence identity with any one of SEQ ID NOs: 161 - 629, 767 - 1220, 1959 - 2522, and 2582 - 2585.

3. A fusion protein comprising a catalytically dead nuclease linked to a reverse transcriptase using a linker, wherein the reverse transcriptase has at least about 80% sequence identity with any one of SEQ ID NOs: 161 - 629, 767 - 1220, 1959 - 2522, and 2582 - 2585.

4. A gene editing system comprising: a) A nickase; b) A guide nucleic acid configured to form a complex with the nickase and hybridize to a target nucleic acid sequence; and c) A reverse transcriptase having at least about 80% sequence identity with any one of SEQ ID NOs: 161 - 629, 767 - 1220, 1959 - 2522, and 2582 - 2585 and configured to form a complex with the nickase.

5. The gene editing system according to claim 4, wherein the nickase is a modified endonuclease.

6. The gene editing system according to claim 5, wherein the modified endonuclease is a type II CRISPR endonuclease.

7. The gene editing system according to claim 5, wherein the modified endonuclease is a type V CRISPR endonuclease.

8. The gene editing system according to any one of claims 6 to 7, wherein the type II CRISPR endonuclease or the type V CRISPR endonuclease has nickase activity.

9. The gene editing system according to claim 5, wherein the modified endonuclease is selected from the group consisting of: spCas9(H840A), spCas9(D10A), nMG3 - 6(D13A), nMG3 - 6(H586A), nMG3 - 6(N609A), Cas12a, and MG29 - 1.

10. The gene editing system according to claim 5, wherein the modified endonuclease has at least about 80% sequence identity with any one of SEQ ID NOs: 152 - 154.

11. The gene editing system according to any one of claims 4 to 10, wherein the nickase and the reverse transcriptase are fused.

12. The gene editing system according to any one of claims 4 to 10, wherein the nickase and the reverse transcriptase are linked by a linker.

13. The gene editing system according to claim 12, wherein the linker comprises at least 10, 20 or 30 amino acids.

14. The gene editing system according to claim 12, wherein the linker comprises about 30 - 35 amino acids.

15. The gene editing system according to claim 12, wherein the linker comprises about 30 amino acids.

16. The gene editing system according to claim 12, wherein the linker has at least 80% sequence identity with SEQ ID NO:

103.

17. The gene editing system according to claim 12, wherein the linker has at least 80% sequence identity with any one of SEQ ID NOs: 155 - 160.

18. The gene editing system according to any one of claims 4 to 10, wherein the nickase and the reverse transcriptase are not linked.

19. The gene editing system according to any one of claims 4 to 18, wherein the guide nucleic acid comprises a spacer sequence and a crRNA.

20. The gene editing system according to any one of claims 4 to 19, wherein the guide nucleic acid further comprises a reverse transcriptase template (RTT).

21. The gene editing system according to claim 20, wherein the bases in the RTT comprise bulky modifications selected from the group consisting of complex sugars or complex amino groups and / or other modifications compatible with RNA.

22. The gene editing system according to any one of claims 4 to 21, wherein the guide nucleic acid further comprises a primer binding site.

23. The gene editing system according to claim 22, wherein the primer binding site is located at the 3'-end of the guide nucleic acid.

24. The gene editing system according to any one of claims 22 to 23, wherein the primer binding site comprises at least 2, 4, 6, 8, 10, 13, 16, 20, 24, 28, 32, 36, 40, 45, 50, 55, 60 or 65 nucleotides.

25. The gene editing system according to any one of claims 4 to 24, wherein the nuclease is non-covalently linked to the guide nucleic acid.

26. The gene editing system according to any one of claims 4 to 24, wherein the nuclease is covalently linked to the guide nucleic acid.

27. The gene editing system according to any one of claims 4 to 24, wherein the nuclease is fused with the guide nucleic acid.

28. The gene editing system according to any one of claims 4 to 24, which further comprises a transposase, an integrase or a homing endonuclease.

29. The gene editing system according to any one of claims 4 to 28, which further comprises a retrotransposon.

30. The gene editing system according to any one of claims 4 to 29, wherein the reverse transcriptase has a processive synthesis ability that is at least about 2 times that of the Moloney murine leukemia virus (MMLV) reverse transcriptase.

31. The gene editing system according to any one of claims 4 to 29, wherein the reverse transcriptase has a processive synthesis ability that is at most about 1 / 2 of the processive synthesis ability of Moloney murine leukemia virus (MMLV) reverse transcriptase.

32. The gene editing system according to any one of claims 4 to 31, wherein the error rate of the reverse transcriptase is less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10% or 0.05%.

33. The gene editing system according to any one of claims 4 to 32, wherein the error rate of the reverse transcriptase is less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10% or 0.05% compared to Moloney murine leukemia virus (MMLV) reverse transcriptase.

34. A gene editing system, comprising: a) a nuclease; b) a guide nucleic acid configured to form a complex with the nuclease and hybridize with a target nucleic acid sequence; and c) a reverse transcriptase having at least about 80% sequence identity with any one of SEQ ID NOs: 161 - 629, 767 - 1220, 1959 - 2522, and 2582 - 2585 and configured to form a complex with the nuclease.

35. The gene editing system according to claim 34, wherein the nuclease is a double-stranded nuclease.

36. The gene editing system according to any one of claims 34 to 35, wherein the nuclease is a type II CRISPR endonuclease.

37. The gene editing system according to claim 36, wherein the CRISPR endonuclease is Cas9.

38. The gene editing system according to claim 37, wherein the Cas9 is catalytically dead Cas9 (dCas9).

39. The gene editing system according to any one of claims 34 to 38, wherein the nuclease and the reverse transcriptase are fused.

40. The gene editing system according to any one of claims 34 to 38, wherein the nuclease and the reverse transcriptase are linked by a linker.

41. The gene editing system according to claim 40, wherein the linker comprises at least 10, 20 or 30 amino acids.

42. The gene editing system according to claim 40, wherein the linker comprises about 30 - 35 amino acids.

43. The gene editing system according to claim 40, wherein the linker comprises about 30 amino acids.

44. The gene editing system according to claim 40, wherein the linker has at least 80% sequence identity with SEQ ID NO:

103.

45. The gene editing system according to claim 40, wherein the linker has at least 80% sequence identity with any one of SEQ ID NOs: 155 - 160.

46. The gene editing system according to any one of claims 34 to 38, wherein the nuclease and the reverse transcriptase are not linked.

47. The gene editing system according to any one of claims 34 to 46, wherein the guide nucleic acid further comprises a primer binding site.

48. The gene editing system according to claim 47, wherein the primer binding site is located at the 3'-end of the guide nucleic acid.

49. The gene editing system according to any one of claims 47 to 48, wherein the primer binding site comprises at least 2, 4, 6, 8, 10, 13, 16, 20, 24, 28, 32, 36, 40, 45, 50, 55, 60 or 65 nucleotides.

50. The gene editing system according to any one of claims 34 to 49, wherein the nuclease is non-covalently linked to the guide nucleic acid.

51. The gene editing system according to any one of claims 34 to 49, wherein the nuclease is covalently linked to the guide nucleic acid.

52. The gene editing system according to any one of claims 34 to 49, wherein the nuclease is fused with the guide nucleic acid.

53. The gene editing system according to any one of claims 34 to 52, which further comprises a transposase, an integrase or a homing endonuclease.

54. The gene editing system according to any one of claims 34 to 53, which further comprises a retrotransposon.

55. The gene editing system according to any one of claims 34 to 54, wherein the reverse transcriptase has a processive synthesis ability that is at least about 2-fold that of the Moloney murine leukemia virus (MMLV) reverse transcriptase.

56. The gene editing system according to any one of claims 34 to 54, wherein the reverse transcriptase has a processive synthesis ability that is at most about 1 / 2 that of the Moloney murine leukemia virus (MMLV) reverse transcriptase.

57. The gene editing system according to any one of claims 34 to 56, wherein the reverse transcriptase has an error rate of less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10% or 0.05%.

58. The gene editing system according to any one of claims 34 to 56, wherein the reverse transcriptase has an error rate that is less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10% or 0.05% compared to the Moloney murine leukemia virus (MMLV) reverse transcriptase.

59. A gene editing system, comprising: a) a nickase; b) a guide nucleic acid configured to form a complex with the nickase and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase configured to form a complex with the nickase, the reverse transcriptase having an X1X2DD motif, wherein X1 is F or Y, and wherein when X1 is Y, X2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W or Y.

60. The gene editing system according to claim 59, wherein X2 is A or I.

61. The gene editing system according to claim 59, wherein the X1X2DD motif is YADD (SEQ ID NO: 2572) or YIDD (SEQ ID NO: 2573).

62. The gene editing system according to claim 59, wherein the X1X2DD motif is FADD (SEQ ID NO: 2574), FVDD (SEQ ID NO: 2575), FIDD (SEQ ID NO: 2576), or FLDD (SEQ ID NO: 2577).

63. The gene editing system according to any one of claims 59 to 62, wherein the reverse transcriptase has at least about 80% sequence identity with any one of SEQ ID NOs: 161 - 629, 767 - 1220, 1959 - 2522, and 2582 - 2585.

64. A gene editing system comprising: a) a nuclease; b) a guide nucleic acid configured to form a complex with the nuclease and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase configured to form a complex with the nuclease, the reverse transcriptase having an X1X2DD motif, wherein X1 is F or Y, and wherein when X1 is Y, X2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y.

65. The gene editing system according to claim 64, wherein X2 is A or I.

66. The gene editing system according to claim 64, wherein the X1X2DD motif is YADD (SEQ ID NO: 2572) or YIDD (SEQ ID NO: 2573).

67. The gene editing system according to claim 64, wherein the X1X2DD motif is FADD (SEQ ID NO: 2574), FVDD (SEQ ID NO: 2575), FIDD (SEQ ID NO: 2576), or FLDD (SEQ ID NO: 2577).

68. The gene editing system according to any one of claims 64 to 67, wherein the reverse transcriptase has at least about 80% sequence identity with any one of SEQ ID NOs: 161 - 629, 767 - 1220, 1959 - 2522, and 2582 - 2585.

69. An isolated reverse transcriptase having at least about 80% sequence identity with any one of SEQ ID NOs: 161 - 629, 767 - 1220, 1959 - 2522, and 2582 - 2585.

70. A nucleic acid encoding the fusion protein according to any one of claims 1 to 3 or the gene editing system according to any one of claims 4 to 68.

71. The nucleic acid according to claim 70, wherein the nucleic acid is DNA or RNA.

72. The nucleic acid according to claim 71, wherein the RNA is mRNA.

73. A vector comprising the nucleic acid according to any one of claims 70 to 72.

74. An adeno-associated virus or lipid nanoparticle comprising the nucleic acid according to any one of claims 70 to 72 or the vector according to claim 73.

75. A cell comprising the nucleic acid according to any one of claims 70 to 72 or the vector according to claim 73.

76. The cell according to claim 75, wherein the cell is a human cell.

77. The cell according to claim 75, wherein the cell is a eukaryotic cell.

78. The cell according to claim 75, wherein the cell is a mammalian cell.

79. The cell according to claim 75, wherein the cell is an immortalized cell.

80. The cell according to claim 75, wherein the cell is an insect cell.

81. The cell according to claim 75, wherein the cell is a yeast cell.

82. The cell according to claim 75, wherein the cell is a plant cell.

83. The cell according to claim 75, wherein the cell is a fungal cell.

84. The cell according to claim 75, wherein the cell is a prokaryotic cell.

85. The cell according to claim 75, wherein the cell is A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cell, HT1080, HepG2, Huh7, K562, a primary cell or a derivative thereof.

86. The cell according to claim 75, wherein the cell is an engineered cell.

87. The cell according to claim 75, wherein the cell is a stable cell.

88. A method for modifying double-stranded and / or single-stranded nucleic acid, the method comprising contacting a cell with the fusion protein according to any one of claims 1 to 3 or the gene editing system according to any one of claims 4 to 68.

89. A method for modifying double-stranded and / or single-stranded nucleic acid, the method comprising: a) providing a guide nucleic acid to the cell to bind to a target strand of the nucleic acid; b) providing a nuclease or nickase to the cell to cleave the nucleic acid at the binding position of the guide nucleic acid; c) providing a reverse transcriptase to the cell to synthesize a modification in the target strand of the nucleic acid at the position cleaved by the nickase and / or the nuclease.

90. The method according to claim 89, wherein the reverse transcriptase has at least about 80% sequence identity with any one of SEQ ID NO: 161-629, 767-1220, 1959-2522, and 2582-2585.

91. The method according to claim 89, wherein the modification is an insertion, deletion, or mutation.

92. The method according to claim 89, which further comprises providing an RNA or DNA template.

93. The method according to claim 89, wherein the nucleic acid is genomic or a vector.

94. The method according to claim 89, further comprising providing a transposase, integrase or homing endonuclease to the cell.

95. The method according to claim 89, further comprising providing a retrotransposon to the cell.

Citation Information

Patent Citations

  • Regulation of endogenous gene expression in cells using zinc finger proteins

    US20030087817A1

  • N[ omega ,( omega -1)-dialkyloxy]- and N-[ omega ,( omega -1)-dialkenyloxy]-alk-1-yl-N,N,N-tetrasubstituted ammonium lipids and uses therefor

    US4897355A

  • N-( omega ,( omega -1)-dialkyloxy)- and N-( omega ,( omega -1)-dialkenyloxy)-alk-1-yl-N,N,N-tetrasubstituted ammonium lipids and uses therefor

    US4946787A

  • N- omega ,( omega -1)-dialkyloxy)- and N-( omega ,( omega -1)-dialkenyloxy)Alk-1-YL-N,N,N-tetrasubstituted ammonium lipids and uses therefor

    US5049386A

  • Cationic lipids for intracellular delivery of biologically active molecules

    WO1991016024A1