Gene editing systems containing reverse transcriptase
The gene editing system with a reverse transcriptase and nuclease/nickase addresses inefficiencies in existing systems by enabling precise site-specific editing and large template integration, improving genome manipulation efficacy.
Patent Information
- Application Number
- JP2025522622
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-08
- Filing Date
- 2023-10-18
- Publication Date
- 2025-12-02
AI Technical Summary
Current gene editing systems struggle to efficiently introduce site-specific insertions, deletions, and mutations, particularly with larger DNA fragments, and lack safe and efficient methods for targeted genome integration, such as in gene therapy or engineered cell therapy, due to random integration or limited cargo capacity.
A gene editing system comprising a reverse transcriptase, a nuclease or nickase, and a guide RNA, which can introduce site-specific modifications and facilitate large template integration using a nucleic acid template.
Enables precise and efficient site-specific editing, including large template integration, addressing the limitations of existing systems by enhancing fidelity and efficiency in genome manipulation.
Smart Images

Figure 2025538855000001_ABST
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 380,195, filed October 19, 2022, and U.S. Provisional Patent Application No. 63 / 386,659, filed December 8, 2022, each of which is incorporated by reference herein in its entirety. Summary of the Invention
[0002] The present disclosure is based, in part, on the development of a gene editing system that includes a reverse transcriptase, a nuclease or nickase, and a guide RNA or pegRNA.
[0003] Described herein are fusion proteins comprising a nickase linked to a reverse transcriptase using a linker, wherein the reverse transcriptase comprises at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
[0004] Described herein are fusion proteins comprising a nuclease linked to a reverse transcriptase using a linker, wherein the reverse transcriptase comprises at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
[0005] Described herein are fusion proteins comprising a catalytically deficient nuclease linked to a reverse transcriptase using a linker, wherein the reverse transcriptase comprises at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
[0006]
[0010] Described herein is a gene editing system comprising: a) a nickase; b) a guide nucleic acid configured to form a complex with the nickase and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase having at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139, and configured to form a complex with the nickase. In some embodiments, the gene editing system further comprises a nucleic acid template. In some embodiments, the nickase is a modified endonuclease. In some embodiments, the modified endonuclease is a Type II CRISPR endonuclease. In some embodiments, the modified endonuclease is a Type V CRISPR endonuclease. In some embodiments, the Type II CRISPR endonuclease or the Type V CRISPR endonuclease has nickase activity. In some embodiments, the modified endonuclease is selected from the group consisting of spCas9(H840A), spCas9(D10A), nMG3-6(D13A), nMG3-6(H586A), nMG3-6(N609A), Cas12a, and MG29-1. In some embodiments, the modified endonuclease comprises at least about 80% sequence identity to any one of SEQ ID NOs:117-119.
[0007] In some embodiments, the nickase and reverse transcriptase are linked. In some embodiments, the nickase and reverse transcriptase are linked by a linker. In some embodiments, the linker comprises at least 10, 20, or 30 amino acids. In some embodiments, the linker comprises about 30-35 amino acids. In some embodiments, the linker comprises about 30 amino acids. In some embodiments, the linker comprises at least 80% sequence identity to SEQ ID NO: 33. In some embodiments, the linker comprises at least 80% sequence identity to any one of SEQ ID NOs: 82-87. In some embodiments, the nickase and reverse transcriptase are unlinked. In some embodiments, the guide nucleic acid comprises a spacer sequence and a crRNA. In some embodiments, the guide nucleic acid further comprises a reverse transcriptase template (RTT). In some embodiments, the bases in the RTT comprise a bulk modification selected from the group consisting of a complex sugar, a complex amino group, and / or other modifications compatible with RNA. In some embodiments, the guide nucleic acid further comprises a primer binding site. In some embodiments, the primer binding site is at the 3' end of the guide nucleic acid. In some embodiments, the primer binding site comprises at least 2, 4, 6, 8, 10, 13, 16, 20, 24, 28, 32, 36, 40, 45, 50, 55, 60, or 65 nucleotides. In some embodiments, the gene editing system further comprises a transposase, an integrase, or a homing endonuclease. In some embodiments, the gene editing system further comprises a retrotransposon. In some embodiments, the reverse transcriptase comprises a processivity that is at least about two-fold higher than Moloney Murine Leukemia Virus (MMLV) reverse transcriptase. In some embodiments, the reverse transcriptase comprises a processivity that is at least about two-fold lower than Moloney Murine Leukemia Virus (MMLV) reverse transcriptase. In some embodiments, the reverse transcriptase comprises an error rate of less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10%, or 0.05%.In some embodiments, the reverse transcriptase comprises an error rate of less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10%, or 0.05% compared to Moloney Murine Leukemia Virus (MMLV) reverse transcriptase.
[0008]
[0009] Described herein is a gene editing system comprising: a) a nuclease; b) a guide nucleic acid configured to form a complex with the nuclease and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase having at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139, and configured to form a complex with the nickase. In some embodiments, the gene editing system further comprises a nucleic acid template. In some embodiments, the nuclease is a double-stranded nuclease. In some embodiments, the nuclease is a Type II CRISPR endonuclease. In some embodiments, the CRISPR endonuclease is Cas9. In some embodiments, the Cas9 is catalytically deficient Cas9 (dCas9). In some embodiments, the nuclease and the reverse transcriptase are linked. In some embodiments, the nuclease and the reverse transcriptase are linked by a linker. In some embodiments, the linker comprises at least 10, 20, or 30 amino acids. In some embodiments, the linker comprises about 30-35 amino acids. In some embodiments, the linker comprises about 30 amino acids. In some embodiments, the linker comprises at least 80% sequence identity to SEQ ID NO: 33. In some embodiments, the linker comprises at least 80% sequence identity to any one of SEQ ID NOs: 82-87. In some embodiments, the nuclease and reverse transcriptase are unlinked. In some embodiments, the guide nucleic acid further comprises a primer binding site. In some embodiments, the primer binding site is at the 3' end of the guide nucleic acid. In some embodiments, the primer binding site comprises at least 2, 4, 6, 8, 10, 13, 16, 20, 24, 28, 32, 36, 40, 45, 50, 55, 60, or 65 nucleotides. In some embodiments, the gene editing system further comprises a transposase, integrase, or homing endonuclease. In some embodiments, the gene editing system further comprises comprising a retrotransposon.In some embodiments, the reverse transcriptase comprises a processivity that is at least about 2-fold higher than Moloney murine leukemia virus (MMLV) reverse transcriptase. In some embodiments, the reverse transcriptase comprises a processivity that is at least about 2-fold lower than Moloney murine leukemia virus (MMLV) reverse transcriptase. In some embodiments, the reverse transcriptase comprises an error rate that is less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10%, or 0.05%. In some embodiments, the reverse transcriptase comprises an error rate that is less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10%, or 0.05% compared to Moloney murine leukemia virus (MMLV) reverse transcriptase.
[0009]
[0010] Described herein is a gene editing system comprising: a) a nickase; b) a guide nucleic acid configured to form a complex with the nickase and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase configured to form a complex with the nickase, wherein the reverse transcriptase has an X1X2DD motif, where X1 is F or Y, and when X1 is Y, X2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y. In some embodiments, X2 is A or I. In some embodiments, the X1X2DD motif is YADD (SEQ ID NO: 140) or YIDD (SEQ ID NO: 141). In some embodiments, the X1X2DD motif is FADD (SEQ ID NO: 142), FVDD (SEQ ID NO: 143), FIDD (SEQ ID NO: 144), or FLDD (SEQ ID NO: 145). In some embodiments, the reverse transcriptase has at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
[0010]
[0010] Described herein is a gene editing system comprising: a) a nuclease; b) a guide nucleic acid configured to form a complex with the nuclease and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase configured to form a complex with the nuclease, wherein the reverse transcriptase has an X1X2DD motif, where X1 is F or Y, and when X1 is Y, X2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y. In some embodiments, X2 is A or I. In some embodiments, the X1X2DD motif is YADD (SEQ ID NO: 140) or YIDD (SEQ ID NO: 141). In some embodiments, the X1X2DD motif is FADD (SEQ ID NO: 142), FVDD (SEQ ID NO: 143), FIDD (SEQ ID NO: 144), or FLDD (SEQ ID NO: 145). In some embodiments, the reverse transcriptase has at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
[0011] Described herein is an isolated reverse transcriptase having at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
[0012] Described herein are nucleic acids encoding the above-described fusion proteins or gene editing systems. In some embodiments, the nucleic acid is DNA or RNA. In some embodiments, the RNA is mRNA. In some embodiments, the nucleic acid is contained in a vector. In some embodiments, the nucleic acid or the vector comprising the nucleic acid is contained in an adeno-associated virus or lipid nanoparticle. In some embodiments, the nucleic acid or the vector comprising the nucleic acid is contained in a cell. In some embodiments, the cell is a human cell.
[0013] Described herein are methods for modifying double-stranded and / or single-stranded nucleic acids, comprising contacting a cell with the fusion proteins or gene editing systems described above.
[0014] Described herein are methods for modifying double-stranded and / or single-stranded nucleic acids in a cell, the methods comprising: a) providing to the cell a guide nucleic acid that binds to a target strand of the nucleic acid; b) providing to the cell a nuclease or nickase to cleave the nucleic acid at the binding site of the guide nucleic acid; and c) providing to the cell a reverse transcriptase to synthesize a modification in the target strand of the nucleic acid at the cleavage site of the nickase and / or nuclease. In some embodiments, the reverse transcriptase has at least about 80% sequence identity to any one of SEQ ID NOS: 88-116 and 138-139. In some embodiments, the modification is an insertion, deletion, or mutation. In some embodiments, the method further comprises providing to the cell an RNA or DNA template. In some embodiments, the nucleic acid is a genome or a vector. In some embodiments, the method further comprises providing to the cell a transposase, integrase, or homing endonuclease. In some embodiments, the method further comprises providing to the cell a retrotransposon. [Brief explanation of the drawings]
[0015] The novel features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings.
[0016] [Figure 1A] 1 is a bar graph showing G to T transediting rates of untethered reverse transcriptase (RT) candidates from the MG153 family. Candidate MG153-23 was tested with eight different primer binding site (PBS) nucleotides of varying lengths (PBS lengths of 2, 4, 6, 8, 10, 13, 16, and 20 nucleotides) in HEK293T cells. The graph also shows untreated samples and wild-type MMLV1 as controls. [Figure 1B]1 is a bar graph showing G to T transediting rates of untethered reverse transcriptase (RT) candidates from the MG153 family. Candidate MG153-24 was tested with eight different primer binding site (PBS) nucleotides of varying lengths (PBS lengths of 2, 4, 6, 8, 10, 13, 16, and 20 nucleotides) in HEK293T cells. The graph also shows untreated samples and wild-type MMLV1 as controls. [Figure 2] 1 is a bar graph showing G to T transediting rates of untethered reverse transcriptase (RT) candidates from the MG160 family. Candidate MG160-7 was tested with eight different primer binding site (PBS) nucleotides of varying lengths (PBS lengths of 2, 4, 6, 8, 10, 13, 16, and 20 nucleotides) in HEK293T cells. The graph also shows untreated samples and wild-type MMLV1 as controls. [Figure 3] 1 is a bar graph showing G to T transediting rates of RT candidates from the MG160 family tethered to spCas9(H840A). The MG160-7 candidate is shown along with an untreated sample, and MMLV1 and MMVL2 as controls.
[0017] Brief description of the sequence listing The Sequence Listing submitted herewith provides exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems according to the present disclosure. Below are exemplary descriptions of the sequences therein.
[0018] SEQ ID NOs: 1-3 show the full-length nucleic acid sequences of untethered MG153 family reverse transcriptases suitable for the gene editing systems described herein.
[0019] SEQ ID NO: 4 shows the full-length nucleic acid sequence of an untethered MG160 family reverse transcriptase suitable for the gene editing system described herein.
[0020] SEQ ID NO: 5 shows the full-length nucleic acid sequence of a tethered MG160 family reverse transcriptase suitable for the gene editing system described herein.
[0021] SEQ ID NOs: 6-13 show the RNA sequences of chemically modified guide RNAs with a single point mutation (VEGFA spacer G to T) with PBSs of different lengths suitable for the gene editing system described herein.
[0022] SEQ ID NOs: 14-21 show the RNA sequences of chemically modified guide RNAs with a single deletion (VEGFA spacer deletion alteration) with PBSs of different lengths suitable for the gene editing system described herein.
[0023] SEQ ID NOs: 22-29 show the RNA sequences of chemically modified guide RNAs with single insertions (VEGFA spacer single insertions) with PBSs of different lengths suitable for the gene editing system described herein.
[0024] SEQ ID NOs: 30 to 31 show the sequences of primers suitable for site-specific editing at the VEGFA site.
[0025] SEQ ID NO: 32 shows the nucleic acid sequence of the VEGFA target site.
[0026] SEQ ID NO: 33 shows the nucleic acid sequence of an exemplary RT-nickase linker.
[0027] SEQ ID NO: 34 shows the nucleic acid sequence of an MG3 effector nuclease suitable for the gene editing system described herein.
[0028] SEQ ID NOs: 35 to 38 show the nucleic acid sequences of the endogenous targets AAVS1, B2M, CD5, and CD38.
[0029] SEQ ID NOs: 39-70 show the RNA sequences of chemically modified guide RNAs with spacers targeting AAVS1, B2M, CD5, and CD38 with PBSs of different lengths suitable for the gene editing system described herein.
[0030] SEQ ID NOs: 71 to 78 show the sequences of primers suitable for site-specific editing at the AAVS1, B2M, CD5, and CD38 sites.
[0031] SEQ ID NO: 79 shows the RNA sequence of a chemically modified guide RNA with a spacer that targets VEGFA.
[0032] SEQ ID NOs: 80 to 81 show the sequences of two retrotransposition assay reporters.
[0033] SEQ ID NOs: 82-87 show the amino acid sequences of exemplary RT-nickase linkers.
[0034] SEQ ID NOs: 88 to 103 show the amino acid sequences of MG140 family retrotransposition proteins suitable for the gene editing system described herein.
[0035] SEQ ID NOs: 104 to 112 show the amino acid sequences of MG148 family reverse transcriptase proteins suitable for the gene editing system described herein.
[0036] SEQ ID NOs: 113 to 115 show the amino acid sequences of MG153 family reverse transcriptase proteins suitable for the gene editing system described herein.
[0037] SEQ ID NO: 116 shows the amino acid sequence of an MG160 family reverse transcriptase protein suitable for the gene editing system described herein.
[0038] SEQ ID NOs: 117 to 119 show the amino acid sequences of MG3-6 nucleases suitable for the gene editing system described herein.
[0039] SEQ ID NOs: 120-135 show nuclear localization signals (NLS) suitable for the gene editing system described herein.
[0040] SEQ ID NO: 136 shows the amino acid sequence of an MG3-6 nuclease suitable for the gene editing system described herein.
[0041] SEQ ID NO: 137 shows the amino acid sequence of an MG29-1 nuclease suitable for the gene editing system described herein.
[0042] SEQ ID NOs: 138-139 show the amino acid sequences of MG160 family reverse transcriptase proteins suitable for the gene editing system described herein. DETAILED DESCRIPTION OF THE INVENTION
[0043] Site-specific gene editing systems are powerful tools for site-specific genome engineering in cells. Programmable nucleases, such as clustered regularly interspaced short palindromic repeats (CRISPR) nucleases, have recently been used for a variety of DNA manipulation and gene editing applications. CRISPR nucleases can be used with or without repair templates to introduce site-specific insertions and deletions (indels) or point mutations of various lengths. Single nucleotide point (SNP) mutations, deletions, and insertions represent more than 80% of disease-causing mutations. However, not all of these mutations can be accurately repaired with available gene editing systems. Clinical genome editing applications with higher efficiency and fidelity are needed.
[0044] Furthermore, repairing or inserting longer DNA fragments remains challenging, and safe and efficient methods for targeted integration of large templates into genomes, such as in gene therapy or engineered cell therapy, are lacking. To date, lentiviruses or adeno-associated viruses (AAVs) have been used in combination with CRISPR nucleases to insert large DNA fragments, such as entire genes. However, lentivirus-mediated integration lacks targeting capabilities because integration often occurs randomly in open chromatin. AAV-mediated delivery has limited cargo capacity and is not available for all cell types. A safe and efficient targeted genome editing system that enables large template integration is needed.
[0045] The present disclosure is based in part on the development of a gene editing system comprising a reverse transcriptase, a nuclease or nickase, and a guide RNA or pegRNA. The gene editing system can be used to introduce site-specific insertions, deletions, and mutations into the genome of a cell. Furthermore, it is contemplated that the gene editing system can be used in combination with a nucleic acid template to facilitate site-specific insertion into the genome of a cell, as well as for large template integration.
[0046] definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the claimed subject matter belongs. It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of any claimed subject matter. The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
[0047] The practice of some methods disclosed herein employs, unless otherwise indicated, techniques in immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (F.M.A.usubel, et al. eds.); the series Methods in Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (M.J.MacPherson, B.D.Hames, and G.R.Taylor eds. (1995)); Harlow and Lane, eds. (1988); Antibodies, A Laboratory Manual; and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (R.I. Freshney, ed. (2010)).
[0048] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent the terms "including," "includes," "having," "has," "with," or variations thereof, are used in either the detailed description and / or claims, such terms are intended to be inclusive in the same manner as the term "comprising."
[0049] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within one or more standard deviations, as is customary in the art. Alternatively, "about" can mean within a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.
[0050] As used herein, the term "nucleotide" refers to a base-sugar-phosphate combination. Contemplated nucleotides include naturally occurring and synthetic nucleotides. A nucleotide is a monomeric unit of a nucleic acid sequence (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide includes the ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives include, for example, [αS]dATP, 7-deaza-dGTP, and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used herein, the term nucleotide encompasses dideoxyribonucleoside triphosphates (ddNTPs) and derivatives thereof. Illustrative examples of ddNTP include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP.Nucleotide can be unlabeled, or can be detectably labeled, for example, by using an optically detectable moiety (e.g., fluorophore) or a moiety containing quantum dot.Detectable labels include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labels for nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, all available from Perkin Elmer (Foster City, Calif.); FluoroLink DeoxyNucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP, fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP available from Boehringer (Mannheim, Indianapolis, Ind.), and Molecular Examples of chromosomal labeled nucleotides available from Probes (Eugene, Oreg.) include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. The term nucleotide encompasses chemically modified nucleotides. An exemplary chemically modified nucleotide is biotin-dNTP.Non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).
[0051] The terms "polynucleotide," "oligonucleotide," and "nucleic acid" are used interchangeably to refer to polymeric forms of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, in either single-, double-, or multiple-stranded form. Contemplated polynucleotides include genes or fragments thereof. Exemplary polynucleotides include, but are not limited to, DNA, RNA, coding or non-coding regions of genes or gene fragments, multiple loci (single locus) defined from binding analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, cell-free polynucleotides, including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. When referring to T, in polynucleotides, T refers to U (uracil) in RNA and T (thymine) in DNA. Polynucleotides can be exogenous or endogenous to a cell and / or can exist in a cell-free environment. The term polynucleotide encompasses modified polynucleotides (e.g., modified backbones, sugars, or nucleobases). If present, modifications to the nucleotide structure are imparted before or after assembly of the polymer. Non-limiting examples of modifications include 5-bromouracil, peptide nucleic acids, heterologous nucleic acids, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein attached to the sugar), thiol-containing nucleotides, biotin-conjugated nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosine, and wyosine. The sequence of nucleotides can be interrupted by non-nucleotide components.
[0052] The term "transfection" or "transfected" refers to the introduction of a polynucleotide into a cell by non-viral or viral-based methods. The polynucleotide may be a gene sequence encoding an entire protein or a functional portion thereof. See, e.g., Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1-18.88.
[0053] The terms "peptide," "polypeptide," and "protein" are used interchangeably herein to refer to a polymer of at least two amino acid residues joined by a peptide bond. The term does not denote a specific length of the polymer, and is not intended to imply or distinguish whether the peptide is produced using recombinant technology, chemical or enzymatic synthesis, or naturally occurring. The term applies to naturally occurring amino acid polymers as well as amino acid polymers containing at least one modified amino acid. In some cases, the polymer is interrupted by non-amino acids. The term includes amino acid chains of any length, including full-length proteins and proteins with or without secondary or tertiary structure (e.g., domains). The term also encompasses amino acid polymers that have been modified by, for example, disulfide bond formation, glycosylation, lipid formation, acetylation, phosphorylation, oxidation, and any other manipulation, such as conjugation with a labeling component. As used herein, the terms "amino acid" and "amino acids" refer to natural and unnatural amino acids, including, but not limited to, modified amino acids. Modified amino acids include amino acids that have been chemically modified to include a group or chemical moiety that does not occur naturally on the amino acid. The term "amino acid" includes both D- and L-amino acids.
[0054] As used herein, "non-naturally occurring" refers to a nucleic acid or polypeptide sequence that does not occur in nature. Non-naturally occurring refers to a nucleic acid or polypeptide sequence that does not occur in nature, including modifications such as mutations, insertions, or deletions. The term non-naturally occurring encompasses fusion nucleic acids or polypeptides in which the non-naturally occurring sequence encodes or exhibits an activity (e.g., an enzymatic activity, a methyltransferase activity, an acetyltransferase activity, a kinase activity, a ubiquitination activity, etc.) of the nucleic acid or polypeptide sequence to which it is fused. Non-naturally occurring nucleic acid or polypeptide sequences include those that are linked by genetic engineering to a naturally occurring nucleic acid or polypeptide sequence (or a variant thereof) to produce a chimeric nucleic acid or polypeptide sequence that encodes the chimeric nucleic acid or polypeptide.
[0055] As used herein, the term "promoter" refers to a regulatory DNA region that controls the transcription or expression of a polynucleotide (e.g., a gene) and may be located adjacent to or overlapping the nucleotide or region of nucleotide at which RNA transcription is initiated. A promoter may contain specific DNA sequences that bind protein factors, often referred to as transcription factors, that facilitate the binding of RNA polymerase to DNA, resulting in gene transcription. Basal promoters in eukaryotes typically, but not necessarily, contain a TATA box and / or a CAAT box.
[0056] The term "expression," as used herein, refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (such as into mRNA or other RNA transcript) and / or the process by which a transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. Collectively, the transcript and the encoded polypeptide may be referred to as a "gene product." If the polynucleotide is derived from genomic DNA, the term expression includes splicing of mRNA in a eukaryotic cell.
[0057] As used herein, "operably linked," "operable linkage," "operatively linked," or their grammatical equivalents refer to the arrangement of genetic elements, e.g., promoters, enhancers, polyadenylation sequences, etc., such that operation (e.g., movement or activation) of a first genetic element has some effect on a second genetic element. The effect on the second genetic element can be, but need not be, of the same type as operation of the first genetic element. For example, two genetic elements are operably linked if movement of the first element causes activation of the second element. A regulatory element, which may include, for example, a promoter and / or enhancer sequence, is operably linked to a coding region if the regulatory element helps to initiate transcription of the coding sequence. There can be intervening residues between the regulatory element and the coding region, so long as this functional relationship is maintained.
[0058] As used herein, "vector" refers to a polymer or association of polymers that contains or is associated with a polynucleotide and mediates delivery of the polynucleotide to a cell. Examples of vectors include nucleic acid-based vectors (e.g., plasmids and viral vectors) and liposomes. Exemplary nucleic acid-based vectors generally include genetic elements, e.g., regulatory elements, operably linked to a gene to facilitate expression of the gene in the target.
[0059] As used herein, "expression cassette" and "nucleic acid cassette" are used interchangeably to refer to a component of a vector that contains a combination of nucleic acid sequences or elements (e.g., a therapeutic gene, a promoter, and a terminator) that are to be expressed together or that are operably linked for expression. The term encompasses expression cassettes that contain a combination of one or more genes with regulatory elements that are operably linked for expression.
[0060] A "functional fragment" of a DNA or protein sequence refers to a fragment that retains a biological activity (either functional or structural) substantially similar to that of the full-length DNA or protein sequence. The biological activity of a DNA sequence includes the ability to affect expression in a manner attributable to the full-length sequence.
[0061] The terms "engineered," "synthetic," and "artificial" are used interchangeably herein to refer to entities modified by human intervention. For example, the terms refer to polynucleotides or polypeptides that do not occur in nature. Engineered peptides may, but need not, have low sequence identity (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1% sequence identity) to naturally occurring human proteins. For example, the VPR domain and the VP64 domain are synthetic transactivation domains. Non-limiting examples include: nucleic acids that are modified by changing their sequence to one that does not occur in nature; nucleic acids that are modified by ligating to a nucleic acid with which they are not naturally associated such that the ligated product possesses a function not present in the original nucleic acid; engineered nucleic acids that are synthesized in vitro using a sequence that does not occur in nature; proteins that are modified by changing their amino acid sequence to a sequence that does not occur in nature; and engineered proteins that acquire new functions or properties. An "engineered" system contains at least one engineered component.
[0062] As used herein, "guide nucleic acid" or "guide polynucleotide" refers to a nucleic acid that can hybridize to a target nucleic acid, thereby directing an associated nuclease to the target nucleic acid. A guide nucleic acid can be, but is not limited to, RNA (guide RNA or gRNA), DNA, or a mixture of RNA and DNA. A guide nucleic acid can include crRNA or tracrRNA, or a combination of both. The term guide nucleic acid encompasses engineered guide nucleic acids and programmable guide nucleic acids that specifically bind to a target nucleic acid. A portion of a target nucleic acid can be complementary to a portion of a guide nucleic acid. A strand of a double-stranded target polynucleotide that is complementary to a guide nucleic acid and hybridizes with the guide nucleic acid is the complementary strand. A strand of a double-stranded target polynucleotide that is complementary to the complementary strand and therefore not complementary to the guide nucleic acid is referred to as the non-complementary strand. A guide nucleic acid having a polynucleotide strand is a "single guide nucleic acid." A guide nucleic acid having two polynucleotide strands is a "double guide nucleic acid." Unless otherwise specified, the term "guide nucleic acid" is inclusive and refers to both single and double guide nucleic acids. A guide nucleic acid may include a segment referred to as a "nucleic acid targeting segment" or "nucleic acid targeting sequence" or "spacer." A nucleic acid targeting segment may include a subsegment referred to as a "protein binding segment" or "protein binding sequence" or "Cas protein binding segment."
[0063] The term "tracrRNA" or "tracr sequence" refers to a transactivating CRISPR RNA. The tracrRNA interacts with the CRISPR (cr) RNA to form a guide nucleic acid (e.g., guide RNA or gRNA) that can hybridize to a target nucleic acid and thereby direct an associated nuclease to the target nucleic acid.
[0064] As used herein, the term "RuvC_III domain" refers to the third, non-contiguous segment of the RuvC endonuclease domain (the RuvC nuclease domain is composed of three non-contiguous segments, RuvC_I, RuvC_II, and RuvC_III). RuvC domains or segments thereof can generally be identified by alignment to documented domain sequences, structural alignment to proteins with annotated domains, or comparison to hidden Markov models (HMMs) constructed based on documented domain sequences (e.g., Pfam HMM PF18541 for RuvC_III).
[0065] As used herein, the term "HNH domain" refers to an endonuclease domain having characteristic histidine and asparagine residues. HNH domains can generally be identified by alignment to documented domain sequences, structural alignment to proteins with annotated domains, or comparison to hidden Markov models (HMMs) constructed based on documented domain sequences (e.g., Pfam HMM PF01844 for domain HNH).
[0066] As used herein, the term "transposon" refers to mobile elements that move in and out of the genome and carry "cargo DNA" with them. These transposons can differ by the type of nucleic acid they transpose, the type of repeats at the ends of the transposon, the type of cargo they transport, or the mode of transposition (i.e., self-repair or host-repair).
[0067] As used herein, the term "transposase" or "transposases" refers to an enzyme that binds to the ends of a transposon and catalyzes its movement to another part of the genome. Types of movement include cut-and-paste and replicative transposition.
[0068] As used herein, the term "Tn7" or "Tn7-like transposase" refers to a family of transposases that contain three major components: a heteromeric transposase (TnsA and / or TnsB) along with a regulatory protein (TnsC). In addition to the TnsABC transposition proteins, Tn7 elements can encode dedicated target site selection proteins, TnsD and TnsE. In conjunction with TnsABC, the sequence-specific DNA-binding protein TnsD directs transposition into a conserved site termed the "Tn7 attachment site," i.e., attTn7. TnsD is a member of a large family of proteins that also includes TniQ. TniQ has been shown to target transposition into degradation sites of plasmids.
[0069] As used herein, the terms "gene editing" and "genome editing" can be used interchangeably. Gene editing or genome editing refers to changing the nucleic acid sequence of a gene or genome. Genome editing can include, for example, insertion, deletion, and mutation. Genome editing can be performed by a gene editing system, for example, a nuclease, a reverse transcriptase, a recombinase, or a base editor.
[0070] As used herein, the term "recombinase" refers to a site-specific enzyme that mediates the recombination of DNA between recombinase recognition sequences, resulting in the excision, integration, inversion, or exchange (e.g., transposition) of the DNA fragment between the recombinase recognition sequences.
[0071] As used herein, the terms "recombining" or "recombination" in the context of nucleic acid modification (e.g., genomic modification) refer to a process in which two or more nucleic acid molecules, or two or more regions of a single nucleic acid molecule, are modified by the action of a recombinase protein. Recombination can result in, among other things, for example, an insertion, inversion, excision, or rearrangement of a nucleic acid sequence within or between one or more nucleic acid molecules.
[0072] As used herein, the term "complex" refers to the joining of at least two components. Each of the two components may retain the properties / activities it had prior to forming the complex or may acquire properties as a result of forming the complex. Joining may be by, but is not limited to, covalent bonding, non-covalent bonding (i.e., hydrogen bonding, ionic interactions, van der Waals interactions, and hydrophobic bonding), use of a linker, fusion, or any other suitable method. Contemplated components of the complex include polynucleotides, polypeptides, or combinations thereof. For example, the complex may include an endonuclease and a guide polynucleotide.
[0073] The term "sequence identity" or "percent identity" in the context of two or more nucleic acid or polypeptide sequences refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences that are identical or have a specified percentage of identical amino acid residues or nucleotides when compared and aligned for maximum correspondence over a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP using the BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and presence of 11, gap cost at extension of 1, and using a conditional composition score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP using parameters of word length (W) of 2, expectation (E) of 1,000,000, and PAM30 scoring setting gap costs at 9 for open gaps and 1 for extended gaps for sequences shorter than 30 residues (these are the default parameters for BLASTP in the BLAST suite available at https: / / blast.ncbi.nlm.nih.gov); CLUSTALW using Smith-Waterman homology search algorithm parameters of 2 matches, -1 mismatches, and -1 gaps; MUSCLE using default parameters; MAFFT using parameters of retries of 2 and maximum repeats of 1,000; Novafold using default parameters; and HMMER hmmalign using default parameters.
[0074] In the context of two or more nucleic acid or polypeptide sequences, the term "optimally aligned" refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences aligned for maximum amino acid residue or nucleotide correspondence, as determined, for example, by the alignment producing the highest or "optimized" percent identity score.
[0075] Variants of any of the enzymes described herein having one or more conservative amino acid substitutions are included in the present disclosure. Such conservative substitutions can be made in the amino acid sequence of a polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by substituting amino acids with similar hydrophobicity, polarity, and R chain length for each other. Additionally or alternatively, by comparing aligned sequences of homologous proteins from different species, conservative substitutions can be identified by finding amino acid residues (e.g., non-conserved residues) that vary between species without changing the basic function of the encoded protein. Such conservatively substituted variants include variants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of the reverse transcriptase protein sequences described herein (e.g., MG140, MG148, MG153, or MG160), the family reverse transcriptases or retrotransposases described herein, or any other family reverse transcriptases or retrotransposases described herein). In some embodiments, such conservatively substituted variants are functional variants, which can include sequences with substitutions of one or more key active site residues such that activity is not abolished.
[0076] The present disclosure also includes variants of any of the enzymes described herein (e.g., reduced activity variants) having substitutions of one or more catalytic residues to reduce or eliminate activity of the enzyme. In some embodiments, reduced activity variants of the proteins described herein include disruptive substitutions of at least one, at least two, or all three catalytic residues (e.g., a programmable nuclease MG3 family nickase having a D13A mutation, a H586A mutation, or a N609A mutation).
[0077] Conservative substitution tables providing functionally similar amino acids are available in various references (see, for example, Creighton, Proteins: Structures and Molecular Properties (WH Freeman & Co.; 2nd edition (December 1993))). The following eight groups each contain amino acids that are conservative substitutions for one another: 1) Alanine (A), Glycine (G); 2) aspartic acid (D), glutamic acid (E); 3) asparagine (N), glutamine (Q); 4) arginine (R), lysine (K); 5) isoleucine (I), leucine (L), methionine (M), valine (V); 6) phenylalanine (F), tyrosine (Y), tryptophan (W); 7) serine (S), threonine (T); and 8) Cysteine (C), methionine (M).
[0078] Gene editing Described herein is a gene editing system comprising: a) a nickase; b) a guide nucleic acid (e.g., a pegRNA or other guide RNA) configured to form a complex with the nickase and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase having at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139, and configured to form a complex with the nickase. Also described herein is a gene editing system comprising: a) a nuclease; b) a guide nucleic acid (e.g., a pegRNA or other guide RNA) configured to form a complex with the nuclease and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase having at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139, and configured to form a complex with the nuclease.
[0010] Further described herein is a gene editing system comprising: a) a nickase; b) a guide nucleic acid (e.g., a pegRNA) configured to form a complex with the nickase and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase configured to form a complex with the nickase, wherein the reverse transcriptase has an X1X2DD motif, wherein X1 is F or Y, and when X1 is Y, X2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y.
[0010] Further described herein is a gene editing system comprising: a) a nuclease; b) a guide nucleic acid (e.g., a pegRNA) configured to form a complex with the nuclease and hybridize to a target nucleic acid sequence; and c) a reverse transcriptase configured to form a complex with the nuclease, wherein the reverse transcriptase has an X1X2DD motif, wherein X1 is F or Y, and when X1 is Y, X2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y.
[0079] In some embodiments, the gene editing systems described herein, including nickases, nucleases, reverse transcriptases, or combinations thereof, are capable of introducing site-specific insertions, deletions, and mutations. In some embodiments, the nickases, nucleases, reverse transcriptases, or combinations thereof, are capable of integrating large polynucleotides. In some embodiments, the integrated polynucleotides include at least about 1 kilobase (kb), 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, or greater than 10 kb in size.
[0080] reverse transcriptase Reverse transcription is the translation of an RNA template into complementary DNA. Reverse transcription is carried out by an enzyme called reverse transcriptase (RT), which has RNA-dependent DNA polymerase activity to generate a complementary DNA (cDNA) strand from an RNA template. Some RT enzymes also have DNA-dependent DNA polymerase activity to generate double-stranded dsDNA. Reverse transcriptases can be of viral origin (e.g., HIV, hepatitis B, Moloney murine leukemia virus (MMLV), or avian myeloblastosis virus (AMV)) or bacterial origin (e.g., group II intron, retron / retron-like RT, diversity-generating retroelement (DGR), Abi-like RT, CRISPR-associated RT, and group II-like RT (G2L)). Reverse transcriptases of eukaryotic origin include telomerase reverse transcriptase, which maintains the telomeres of eukaryotic chromosomes. Reverse transcription allows site-specific insertions, deletions, and mutations to be introduced into cDNA by encoding them in an RNA template.
[0081] In some embodiments, the reverse transcriptase is a viral, prokaryotic, or eukaryotic reverse transcriptase. In some embodiments, the reverse transcriptase comprises a sequence of SEQ ID NOs: 88-116 and 138-139, a variant thereof, or a functional fragment thereof. In some embodiments, the reverse transcriptase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139, a variant thereof, or a functional fragment thereof. In some embodiments, the reverse transcriptase comprises a sequence having at least about 70% identity to any one of SEQ ID NOs: 88-116 and 138-139. In some embodiments, the reverse transcriptase comprises a sequence having at least about 75% identity to any one of SEQ ID NOs: 88-116 and 138-139. In some embodiments, the reverse transcriptase comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 88-116 and 138-139. In some embodiments, the reverse transcriptase comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 88-116 and 138-139. In some embodiments, the reverse transcriptase comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 88-116 and 138-139. In some embodiments, the reverse transcriptase comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 88-116 and 138-139. In some embodiments, the reverse transcriptase comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 88-116 and 138-139. In some embodiments, the reverse transcriptase comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 88-116 and 138-139.In some embodiments, the reverse transcriptase comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 88-116 and 138-139. In some embodiments, the reverse transcriptase comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 88-116 and 138-139. In some embodiments, the reverse transcriptase comprises a sequence having 100% identity to any one of SEQ ID NOs: 88-116 and 138-139.
[0082] In some embodiments, the reverse transcriptase is an MG140, MG148, MG153, or MG160 family reverse transcriptase. In some embodiments, the reverse transcriptase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of an MG140, MG148, MG153, or MG160 family reverse transcriptase or retrotransposase. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of the MG140, MG148, MG153, or MG160 family reverse transcriptases or retrotransposases, or a variant thereof.
[0083] In some embodiments, the reverse transcriptase is encoded by a nucleic acid sequence having at least 80% sequence identity to the nucleic acid sequence of any one of SEQ ID NOs: 1-5. In some embodiments, the reverse transcriptase is encoded by a nucleic acid sequence having at least 85% sequence identity to the nucleic acid sequence of any one of SEQ ID NOs: 1-5. In some embodiments, the reverse transcriptase is encoded by a nucleic acid sequence having at least 90% sequence identity to the nucleic acid sequence of any one of SEQ ID NOs: 1-5. In some embodiments, the reverse transcriptase is encoded by a nucleic acid sequence having at least 95% sequence identity to the nucleic acid sequence of any one of SEQ ID NOs: 1-5. In some embodiments, the reverse transcriptase is encoded by a nucleic acid sequence having at least 96% sequence identity to the nucleic acid sequence of any one of SEQ ID NOs: 1-5. In some embodiments, the reverse transcriptase is encoded by a nucleic acid sequence having at least 97% sequence identity to the nucleic acid sequence of any one of SEQ ID NOs: 1-5. In some embodiments, the reverse transcriptase is encoded by a nucleic acid sequence having at least 98% sequence identity to the nucleic acid sequence of any one of SEQ ID NOs: 1-5. In some embodiments, the reverse transcriptase is encoded by a nucleic acid sequence having at least 99% sequence identity to the nucleic acid sequence of any one of SEQ ID NOs: 1-5. In some embodiments, the reverse transcriptase is encoded by the nucleic acid sequence of any one of SEQ ID NOs: 1-5.
[0084] Reverse transcriptases typically have an active site core tetrad motif of amino acid sequence XXDD. In some embodiments, the reverse transcriptase has an active site tetrad motif of X1X2DD, where X1 is F or Y, and when X1 is Y, X2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y. In some embodiments, X2 is A or I. In some embodiments, the X1X2DD motif is YADD (SEQ ID NO: 140) or YIDD (SEQ ID NO: 141). In some embodiments, the X1X2DD motif is FADD (SEQ ID NO: 142), FVDD (SEQ ID NO: 143), FIDD (SEQ ID NO: 144), or FLDD (SEQ ID NO: 145). In some embodiments, the reverse transcriptase is isolated. In some embodiments, the reverse transcriptase is an MG140, MG148, MG153, and MG160 family reverse transcriptase or retrotransposase, and the X1X2DD motif is YADD (SEQ ID NO: 140) or YIDD (SEQ ID NO: 141). In some embodiments, the reverse transcriptase is isolated. In some embodiments, the reverse transcriptase is an MG140, MG148, MG153, or MG160 family reverse transcriptase or retrotransposase, and the X1X2DD motif is FADD (SEQ ID NO: 142), FVDD (SEQ ID NO: 143), FIDD (SEQ ID NO: 144), or FLDD (SEQ ID NO: 145).
[0085] In some embodiments, the reverse transcriptase is smaller than 300 amino acids. In some embodiments, the reverse transcriptase is smaller than 250 amino acids. In some embodiments, the reverse transcriptase comprises at least about 50, 75, 100, 125, 150, 175, 200, 225, 250, 275, 300, or more than 300 amino acids. In some embodiments, the reverse transcriptase comprises amino acids in the range of about 50 to about 300, about 75 to about 300, about 100 to about 300, about 125 to about 300, about 150 to about 300, about 175 to about 300, about 200 to about 300, about 225 to about 300, about 250 to about 300, about 275 to about 300, about 100 to about 300, about 125 to about 300, about 150 to about 300, about 175 to about 300, about 200 to about 300, about 225 to about 300, about 250 to about 300, or about 275 to about 300.
[0086] In some embodiments, the reverse transcriptase comprises a processivity that is at least about 2-fold higher than Moloney murine leukemia virus (MMLV) reverse transcriptase. In some embodiments, the reverse transcriptase comprises a processivity that is at least about 2-fold lower than Moloney murine leukemia virus (MMLV) reverse transcriptase. In some embodiments, the reverse transcriptase comprises an error rate that is less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10%, or 0.05%. In some embodiments, the reverse transcriptase comprises an error rate that is less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10%, or 0.05% compared to Moloney murine leukemia virus (MMLV) reverse transcriptase. Methods for measuring the processivity of a reverse transcriptase are known in the art or are described herein, e.g., in Example 2.
[0087] In some embodiments, the reverse transcriptase is targetable. Targetable reverse transcriptase is an engineered ribonucleoprotein complex that serves as a tool for genome editing in cells and organisms. In some embodiments, the targetable reverse transcriptase is generated by fusing a reverse transcriptase with a site-specific CRISPR nuclease variant that nicks the non-targeted strand of dsDNA, so that a guide RNA or pegRNA containing a primer binding site (PBS) sequence can find its complementary target sequence, hybridize, and prime the reverse transcriptase reaction using a reverse transcriptase template (RTT) as a template. Two DNA flaps are produced, one containing the desired change encoded in the RTT and the other containing the original sequence. After equilibration, the change is incorporated into the genomic DNA when the DNA flap with the desired edit is repaired by the cellular host repair machinery.
[0088] In some embodiments, the gene editing system comprises a reverse transcriptase and a nickase described herein. In some embodiments, the gene editing system comprises a reverse transcriptase and a nuclease described herein. In some embodiments, the gene editing system comprises a reverse transcriptase and a modified nuclease described herein. In some embodiments, the gene editing system is programmable. In some embodiments, the modified nuclease is a site-specific nickase.
[0089] In some embodiments, the reverse transcriptase and the nuclease or nickase are linked or tethered. In some embodiments, the gene editing system comprises a fusion protein of a reverse transcriptase and a nuclease or nickase. In some embodiments, the gene editing system comprises a fusion protein comprising a nickase linked to a reverse transcriptase using a linker, wherein the reverse transcriptase comprises at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139. In some embodiments, the gene editing system comprises a fusion protein comprising a nuclease linked to a reverse transcriptase using a linker, wherein the reverse transcriptase comprises at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139. In some embodiments, the gene editing system comprises a fusion protein comprising a catalytically deficient nuclease linked to a reverse transcriptase using a linker, wherein the reverse transcriptase comprises at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
[0090] In some embodiments, the reverse transcriptase and the nuclease or nickase are linked or fused using a linker. In some embodiments, the linker comprises at least 10, 20, or 30 amino acids. In some embodiments, the linker comprises about 30-35 amino acids. In some embodiments, the linker comprises about 30 amino acids.
[0091] In some embodiments, the linker comprises at least 80% sequence identity to SEQ ID NO:33. In some embodiments, the linker comprises a sequence having at least about 85% identity to SEQ ID NO:33. In some embodiments, the linker comprises a sequence having at least about 90% identity to SEQ ID NO:33. In some embodiments, the linker comprises a sequence having at least about 91% identity to SEQ ID NO:33. In some embodiments, the linker comprises a sequence having at least about 92% identity to SEQ ID NO:33. In some embodiments, the linker comprises a sequence having at least about 93% identity to SEQ ID NO:33. In some embodiments, the linker comprises a sequence having at least about 94% identity to SEQ ID NO:33. In some embodiments, the linker comprises a sequence having at least about 95% identity to SEQ ID NO:33. In some embodiments, the linker comprises a sequence having at least about 96% identity to SEQ ID NO:33. In some embodiments, the linker comprises a sequence having at least about 97% identity to SEQ ID NO:33. In some embodiments, the linker comprises a sequence having at least about 98% identity to SEQ ID NO: 33. In some embodiments, the linker comprises a sequence having at least about 99% identity to SEQ ID NO: 33. In some embodiments, the linker comprises a sequence having 100% sequence identity to SEQ ID NO: 33.
[0092] Suitable linkers are known in the art and include, for example, any one of SEQ ID NOs: 82-87. In some embodiments, the linker comprises at least 80% sequence identity to any one of SEQ ID NOs: 82-87. In some embodiments, the linker joining any of the enzymes or domains described herein comprises at least 80% sequence identity to SGGSSGGSSGSETPGTSESATPESSGGSSGGSSAC (SEQ ID NO: 82), KLGGGAPAVGGGPK (SEQ ID NO: 83), (GGGGS)3 (SEQ ID NO: 84), (GGGGS)2EAAAK(GGGGS)2 (SEQ ID NO: 85), (GGGGS)2(EAAAK)2(GGGGS)2 (SEQ ID NO: 86), or SGSETPGTSESATPES (SEQ ID NO: 87). The linker may comprise one or more copies of a sequence having 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity, or any other linker sequence described herein. In some embodiments, the linker comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 82-87. In some embodiments, the linker comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 82-87. In some embodiments, the linker comprises a sequence having at least about 91% identity to any one of SEQ ID NOs: 82-87. In some embodiments, the linker comprises a sequence having at least about 92% identity to any one of SEQ ID NOs: 82-87. In some embodiments, the linker comprises a sequence having at least about 93% identity to any one of SEQ ID NOs: 82-87. In some embodiments, the linker comprises a sequence having at least about 94% identity to any one of SEQ ID NOs: 82-87.In some embodiments, the linker comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 82-87. In some embodiments, the linker comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 82-87. In some embodiments, the linker comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 82-87. In some embodiments, the linker comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 82-87. In some embodiments, the linker comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 82-87. In some embodiments, the linker comprises a sequence having 100% identity to any one of SEQ ID NOs: 82-87.
[0093] In some embodiments, the nickase or nuclease and the reverse transcriptase are unlinked.
[0094] In some embodiments, the reverse transcriptase, nuclease, nickase, or fusion protein described herein comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the reverse transcriptase, nuclease, nickase, or fusion protein.
[0095] In some embodiments, the NLS comprises any of the sequences in Table 1 below, or a combination thereof. [Table 1]
[0096] In some embodiments, the reverse transcriptase comprises a tag. In some embodiments, the nuclease comprises a tag. In some embodiments, the nickase comprises a tag. In some embodiments, the fusion protein comprises a tag. In some embodiments, the tag is an affinity tag. Exemplary affinity tags include, but are not limited to, a His tag, a Flag tag, a Myc tag, a MBP tag, and a GST tag.
[0097] In some embodiments, the reverse transcriptase comprises a protease cleavage site. In some embodiments, the nuclease comprises a protease cleavage site. In some embodiments, the nickase comprises a protease cleavage site. In some embodiments, the fusion protein comprises a protease cleavage site. Exemplary protease cleavage sites include, but are not limited to, a TEV site, a C3 site, a factor Xa site, and an enterokinase site.
[0098] In some embodiments, the gene editing system comprises a) a nickase; b) a guide nucleic acid (e.g., a pegRNA or other guide RNA); and c) a reverse transcriptase having at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
[0099] In some embodiments, the gene editing system comprises a) a nuclease; b) a guide nucleic acid (e.g., a pegRNA or other guide RNA); and c) a reverse transcriptase having at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
[0100] In some embodiments, the gene editing system comprises a) a nickase; b) a guide nucleic acid (e.g., a pegRNA); and c) a reverse transcriptase having an X1X2DD motif, where X1 is F or Y, and when X1 is Y, X2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y.
[0101] In some embodiments, the gene editing system comprises: a) a nuclease; b) a guide nucleic acid (e.g., a pegRNA); and c) a reverse transcriptase having an X1X2DD motif, where X1 is F or Y, and when X1 is Y, X2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y. In some embodiments, X2 is A or I. In some embodiments, the X1X2DD motif is YADD (SEQ ID NO: 140) or YIDD (SEQ ID NO: 141). In some embodiments, the X1X2DD motif is FADD (SEQ ID NO: 142), FVDD (SEQ ID NO: 143), FIDD (SEQ ID NO: 144), or FLDD (SEQ ID NO: 145). In some embodiments, the reverse transcriptase has at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
[0102] In some embodiments, the nuclease is configured to cleave one strand of a double-stranded target deoxyribonucleic acid (a nickase). In some embodiments, the nickase or nuclease is a CRISPR nuclease described herein. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 34, or a variant thereof. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 70% identity to SEQ ID NO: 34. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 75% identity to SEQ ID NO: 34. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 80% identity to SEQ ID NO: 34. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 85% identity to SEQ ID NO: 34. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 90% identity to SEQ ID NO: 34. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 95% identity to SEQ ID NO: 34. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 96% identity to SEQ ID NO: 34. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 97% identity to SEQ ID NO: 34. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 98% identity to SEQ ID NO:34.In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having at least about 99% identity to SEQ ID NO: 34. In some embodiments, the nickase or nuclease is encoded by a nucleic acid sequence having 100% identity to SEQ ID NO: 34.
[0103] In some embodiments, the system comprises Mg 2+ The present invention further comprises a source of
[0104] In some embodiments, the nuclease is a modified endonuclease. In some embodiments, the modified endonuclease is a type II CRISPR endonuclease or a type V CRISPR endonuclease. In some embodiments, the type II or type V CRISPR endonuclease may comprise double-strand cleavage activity, nickase activity, or may lack catalytic activity. In some embodiments, the CRISPR nuclease has a modification in the HNH domain or in the RuvC domain.
[0105] In some embodiments, the modified endonuclease comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NOs: 117-119, or a variant thereof. In some embodiments, the modified endonuclease comprises at least about 80% sequence identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 70% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 75% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 98% identity to any one of SEQ ID NOs:117-119.In some embodiments, the modified endonuclease comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having 100% identity to any one of SEQ ID NOs: 117-119.
[0106] In some embodiments, the modified endonuclease is selected from the group consisting of spCas9(H840A), spCas9(D10A), nMG3-6(D13A), nMG3-6(H586A), nMG3-6(N609A), Cas12a, and MG29-1.
[0107] In some embodiments, the gene editing system includes a nucleic acid template. The nucleic acid template can be RNA or DNA. The nucleic acid template can be 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 bases in length. The nucleic acid template can be 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 bases in length. In some embodiments, the nucleic acid template has a homology region that is homologous to a site in a genome. In some embodiments, the region of homology is 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 bases in length.
[0108] In some embodiments, the gene editing system further comprises a transposase, an integrase, or a homing endonuclease. In some embodiments, the transposase is a transposase (Tnp), Tn5, Sleeping Beauty transposase, or Tn7 transposon. In some embodiments, the gene editing system comprises an enzyme with transposase activity. Additional enzymes with transposase activity include, but are not limited to, retrons and IS200 / IS605 transposons.
[0109] In some embodiments, the gene editing system further comprises a retrotransposon of the present disclosure. In some embodiments, the retrotransposon is an MG140 family retrotransposon. In some embodiments, the retrotransposon comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NOs: 88-103, or a variant thereof.
[0110] CRISPR nuclease In some embodiments, the nickase or endonuclease described herein is a CRISPR nuclease. In some embodiments, the CRISPR nuclease is a modified nuclease.
[0111] CRISPR systems are RNA-directed nuclease complexes that have been described to function as adaptive immune systems in microorganisms. In their natural context, CRISPR systems occur in CRISPR (clustered regularly interspaced short palindromic repeats) operons or loci, which generally contain two parts: (i) an array of short repeat sequences (30-40 bp) separated by equally short spacer sequences that encode an RNA-based targeting element; and (ii) an ORF encoding a nuclease polypeptide directed by the RNA-based targeting element along with accessory proteins / enzymes. Efficient nuclease targeting of a specific target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6-8 nucleic acids of the target (target seed) and the crRNA guide; and (ii) the presence of a protospacer adjacent motif (PAM) sequence within a defined vicinity of the target seed (the PAM is typically a sequence not commonly represented in the host genome). Depending on the exact function and organization of the system, CRISPR systems are generally organized into two classes, five types, and 16 subtypes based on shared functional characteristics and evolutionary similarities.
[0112] Class I CRISPR systems have large, multi-subunit effector complexes and include Types I, III, and IV. Class II CRISPR systems generally have single polypeptide, multi-domain nuclease effectors and include Types II, V, and VI.
[0113] Type II CRISPR systems are considered the simplest in terms of components. In type II CRISPR systems, processing the CRISPR array into mature crRNA does not require the presence of a special endonuclease subunit, but rather a small trans-encoded crRNA (tracrRNA) with a region complementary to the array repeat sequence. The tracrRNA interacts with both its corresponding effector nuclease (e.g., Cas9) and the repeat sequence to form a precursor dsRNA structure, which is cleaved by endogenous RNAse III to generate a mature effector enzyme loaded with both tracrRNA and crRNA. Type II nucleases are known as DNA nucleases. Type II effectors generally exhibit a structure consisting of a RuvC-like endonuclease domain that adopts an RNase H fold with an unrelated HNH nuclease domain inserted into the RuvC-like nuclease domain fold. The RuvC-like domain is responsible for cleavage of the target (e.g., crRNA-complementary) DNA strand, while the HNH domain is responsible for cleavage of the displaced DNA strand. Exemplary CRISPR Cas9 proteins include those from Streptococcus pyogenes (UniProtKB-Q99ZW2 (CAS9 STRP1)), Streptococcus thermophilus (UniProtKB-G3ECR1 (CAS9 STRTR)), Staphylococcus aureus (UniProtKB-J7RUA5 (CAS9 STAAU), Campylobacter jejuni (UniProtKB-Q0P897 (CAS9 CAMJE)), Campylobacter lari (UniProtKB-A0A0A8HTA3 (A0A0A8HTA3 CAMLA), and Helicobacter canadensis (UniProtKB-C5ZYI3 (C5ZYI3 9HELI)), Francisella Examples include, but are not limited to, Cas9 from Bacillus tularensis subsp. Novicida (UniProtKB-A0Q5Y3(CAS9_FRATN)).Additional type II nucleases are described in International Patent Applications Nos. 2021 / 226363, 2022 / 159758, and 2022 / 056324.
[0114] Type V CRISPR systems are characterized by a nuclease effector (e.g., Cas12) structure similar to that of type II effectors, including a RuvC-like domain. Like type II, most (if not all) type V CRISPR systems use a tracrRNA to process pre-crRNA into mature crRNA. However, unlike type II systems, which require RNAse III to cleave the pre-crRNA into multiple crRNAs, type V systems are capable of cleaving the pre-crRNA using the effector nuclease itself. Like type II CRISPR systems, type V CRISPR systems are known as DNA nucleases. Unlike type II CRISPR systems, some type V enzymes (e.g., Cas12a) appear to possess robust single-strand nonspecific deoxyribonuclease activity that is activated by the first crRNA-directed cleavage of the double-stranded target sequence.
[0115] In some embodiments, the nuclease or nickase is a CRISPR nuclease. In some embodiments, the CRISPR nuclease is Class 2 Type II SpCas9 or Class 2 Type VA Cas12a (formerly Cpf1). In some embodiments, Type VA nucleases have guide RNAs of 42-44 nucleotides, compared to approximately 100 nt for SpCas9. In some embodiments, Type VA nucleases provide staggered cleavage sites. In some embodiments, Type VA nucleases provide staggered cleavage sites, facilitating directed repair pathways such as microhomology-dependent targeted integration (MITI).
[0116] The most commonly used Type VA enzymes require a 5' protospacer adjacent motif (PAM) adjacent to the selected target site: 5'-TTTV-3' for Lachnospiraceae bacterium ND2006 LbCas12a and Acidaminococcus species AsCas12a; and 5'-TTV-3' for Francisella novicida FnCas12a. In some embodiments, the PAM sequence is YTV, YYN, or TTN. Additional Type II nucleases are described in International Patent Application Publication No. WO 2021 / 226363.
[0117] In some embodiments, the nickase is a modified nuclease. In some embodiments, the modified endonuclease is a Type II CRISPR endonuclease. In some embodiments, the modified endonuclease is a Type II CRISPR endonuclease or a Type V endonuclease. In some embodiments, the Type II CRISPR endonuclease or a Type V endonuclease has nickase activity.
[0118] In some embodiments, the modified endonuclease is selected from the group consisting of spCas9(H840A), spCas9(D10A), nMG3-6(D13A), nMG3-6(H586A), nMG3-6(N609A), Cas12a, and MG29-1. In some embodiments, the nuclease comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 117-119, or a variant thereof. In some embodiments, the modified endonuclease comprises a sequence having at least about 70% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 75% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 117-119.In some embodiments, the modified endonuclease comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 117-119. In some embodiments, the modified endonuclease comprises a sequence having 100% identity to any one of SEQ ID NOs: 117-119.
[0119] In some embodiments, the nuclease is encoded by a nucleic acid sequence having at least 80% sequence identity to the nucleic acid sequence of SEQ ID NO: 34. In some embodiments, the nuclease is encoded by a nucleic acid sequence having at least 85% sequence identity to the nucleic acid sequence of SEQ ID NO: 34. In some embodiments, the nuclease is encoded by a nucleic acid sequence having at least 90% sequence identity to the nucleic acid sequence of SEQ ID NO: 34. In some embodiments, the nuclease is encoded by a nucleic acid sequence having at least 95% sequence identity to the nucleic acid sequence of SEQ ID NO: 34. In some embodiments, the nuclease is encoded by a nucleic acid sequence having at least 96% sequence identity to the nucleic acid sequence of SEQ ID NO: 34. In some embodiments, the nuclease is encoded by a nucleic acid sequence having at least 97% sequence identity to the nucleic acid sequence of SEQ ID NO: 34. In some embodiments, the nuclease is encoded by a nucleic acid sequence having at least 98% sequence identity to the nucleic acid sequence of SEQ ID NO: 34. In some embodiments, the nuclease is encoded by a nucleic acid sequence having at least 99% sequence identity to the nucleic acid sequence of SEQ ID NO: 34. In some embodiments, the nuclease is encoded by the nucleic acid sequence of SEQ ID NO: 34.
[0120] In some embodiments, the RuvC domain lacks nuclease activity. In some embodiments, the HNH domain lacks nuclease activity. In some embodiments, the modified nuclease has a modification corresponding to position H840A of S. pyogenes Cas9. In some embodiments, the modified nuclease has a modification corresponding to position D10A of S. pyogenes Cas9. In some embodiments, the modified nuclease has a modification corresponding to position D13A of MG3-6 (SEQ ID NO: 136) designated nMG3-6(D13A) (SEQ ID NO: 117). In some embodiments, the modified nuclease has a modification corresponding to position H586A of MG3-6 (SEQ ID NO: 136) designated nMG3-6(H586A) (SEQ ID NO: 118). In some embodiments, the modified nuclease has a modification corresponding to position N609A of MG3-6 (SEQ ID NO: 136) designated nMG3-6(N609A) (SEQ ID NO: 119). In some embodiments, the modified nuclease is configured to cleave one strand of a double-stranded target deoxyribonucleic acid. In some embodiments, the ribonucleic acid sequence configured to bind to the endonuclease comprises a tracr sequence.
[0121] In some embodiments, the nickase or nuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the nickase or nuclease.
[0122] In some embodiments, the NLS comprises any of the sequences in Table 1 above, or a combination thereof.
[0123] Guide nucleic acid In some embodiments, provided herein are guide nucleic acids, such as guide RNAs (gRNAs) or prime editing guide RNAs (pegRNAs). When referring to T, in polynucleotides, T means U (uracil) in RNA and T (thymine) in DNA.
[0124] Prime editing allows for the introduction of virtually any combination of point mutations, small insertions, or small deletions in the genome of living cells. Prime editing guide RNAs (pegRNAs) target prime editing proteins to the target locus and encode the desired edits.
[0125] In some embodiments, the guide RNA targets a gene in a cell. In some embodiments, the guide RNA targets a gene in a mammalian cell. In some embodiments, the target gene is TRAC, VEGFA, AAVS1, B2M, CD5, or CD38. Exemplary guide RNAs are set forth in SEQ ID NOs: 6-29 and 39-70.
[0126] In some embodiments, the guide RNA is encoded by any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, a sequence having at least about 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 80% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 85% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 90% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 95% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 97% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 98% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 99% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence according to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or a reverse complement thereof.
[0127] In some embodiments, one or more guide RNAs are encoded by a sequence comprising at least about 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or their reverse complements. In some embodiments, one or more guide RNAs are encoded by a sequence comprising at least about 80% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or their reverse complements. In some embodiments, one or more guide RNAs are encoded by a sequence comprising at least about 85% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or their reverse complements. In some embodiments, one or more guide RNAs are encoded by a sequence comprising at least about 90% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or their reverse complements. In some embodiments, one or more guide RNAs are encoded by a sequence comprising at least about 95% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or a reverse complement thereof. In some embodiments, one or more guide RNAs are encoded by a sequence comprising at least about 97% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or a reverse complement thereof. In some embodiments, one or more guide RNAs are encoded by a sequence comprising at least about 98% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or a reverse complement thereof. In some embodiments, one or more guide RNAs are encoded by a sequence comprising at least about 99% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence according to any one of the nucleic acid sequences of SEQ ID NOs: 6-29 and 39-70, or their reverse complements, or their reverse complements.
[0128] In some embodiments, the guide RNA or pegRNA comprises various structural elements, including, but not limited to, a spacer sequence that binds to a protospacer sequence (target sequence), a crRNA, and an optional tracrRNA. In some embodiments, the genome editing system comprises a CRISPR guide RNA. In some embodiments, the guide RNA comprises a crRNA and a spacer sequence. In some embodiments, the guide RNA additionally comprises a tracrRNA or a modified tracrRNA.
[0129] In some embodiments, the compositions and methods provided herein include one or more guide RNAs. In some embodiments, the guide RNA includes a sense sequence. In some embodiments, the guide RNA includes an antisense sequence. In some embodiments, the guide RNA includes a nucleotide sequence other than a region complementary or substantially complementary to a region of a target sequence. For example, the guide RNA is, or is considered to be, part of the crRNA or is included in the crRNA, e.g., a crRNA:tracrRNA chimera.
[0130] In some embodiments, the guide RNA (e.g., gRNA) comprises synthetic or modified nucleotides. In some embodiments, the guide RNA comprises one or more internucleoside linkers modified from natural phosphodiester. In some embodiments, the internucleoside linker of the guide RNA, or all of its contiguous nucleotide sequence, is modified. For example, in some embodiments, the internucleoside linkage comprises sulfur (S), such as a phosphorothioate internucleoside linkage.
[0131] In some embodiments, a guide RNA (e.g., gRNA) comprises a modification to the ribose sugar or nucleobase. In some embodiments, a guide RNA comprises one or more nucleosides comprising a modified sugar moiety, which is a modification of the sugar moiety compared to the ribose sugar moiety found in deoxyribose nucleic acids (DNA) and RNA. In some embodiments, the modification is within the ribose ring structure. Exemplary modifications include, but are not limited to, replacement with a hexose ring (HNA), a bicyclic ring having a biradical bridge between the C2 and C4 carbons on the ribose ring (e.g., locked nucleic acids (LNA)), or an unlinked ribose ring, which typically lacks a bond between the C2 and C3 carbons (e.g., UNA). In some embodiments, a sugar-modified nucleoside comprises a bicyclohexose nucleic acid or a tricyclic nucleic acid. In some embodiments, a modified nucleoside comprises a nucleoside in which the sugar moiety is replaced with a non-sugar moiety, e.g., a peptide nucleic acid (PNA) or a morpholino nucleic acid.
[0132] In some embodiments, the guide RNA comprises one or more modified sugars. In some embodiments, sugar modifications include modifications made by altering the substituent on the ribose ring to a group other than hydrogen or to the 2'-OH group naturally found in DNA and RNA nucleosides. In some embodiments, the substituent is introduced at the 2', 3', 4', 5' position, or a combination thereof. In some embodiments, nucleosides having modified sugar moieties include 2'-modified nucleosides, e.g., 2'-substituted nucleosides. 2'-sugar-modified nucleosides, in some embodiments, are nucleosides having a substituent other than H or -OH at the 2' position (2'-substituted nucleosides) or include a 2'-linked biradical, and include 2'-substituted nucleosides and LNA (2'-4' biradical bridged) nucleosides. Examples of 2'-substituted modified nucleosides include, but are not limited to, 2'-O-alkyl-RNA, 2'-O-methyl-RNA, 2'-alkoxy-RNA, 2'-O-methoxyethyl-RNA (MOE), 2'-amino-DNA, 2'-fluoro-RNA, and 2'-F-ANA nucleosides. In some embodiments, the modification in the ribose group comprises a modification at the 2' position of the ribose group. In some embodiments, the modification at the 2' position of the ribose group is selected from the group consisting of 2'-O-methyl, 2'-fluoro, 2'-deoxy, and 2'-O-(2-methoxyethyl).
[0133] In some embodiments, the guide RNA comprises one or more modified sugars. In some embodiments, the guide RNA comprises only modified sugars. In certain embodiments, the guide RNA comprises more than about 10%, 25%, 50%, 75%, or 90% modified sugars. In some embodiments, the modified sugar is a bicyclic sugar. In some embodiments, the modified sugar comprises a 2'-O-methoxyethyl group. In some embodiments, the guide RNA comprises both an internucleoside linker modification and a nucleoside modification.
[0134] In some embodiments, the guide RNA comprises from about 15 nucleotides to about 28 nucleotides. In some embodiments, the guide RNA comprises at least about 15 nucleotides. In some embodiments, the guide RNA comprises up to about 28 nucleotides. In some embodiments, the guide RNA is from about 15 nucleotides to about 16 nucleotides, from about 15 nucleotides to about 17 nucleotides, from about 15 nucleotides to about 18 nucleotides, from about 15 nucleotides to about 19 nucleotides, from about 15 nucleotides to about 20 nucleotides, from about 15 nucleotides to about 21 nucleotides, from about 15 nucleotides to about 22 nucleotides, from about 15 nucleotides to about 23 nucleotides, from about 15 nucleotides to about 24 nucleotides, from about 15 nucleotides to about 25 nucleotides, from about 15 nucleotides to about 28 nucleotides, from about 16 nucleotides to about 17 nucleotides, from about 16 nucleotides to about 18 nucleotides, from about 16 nucleotides to about 19 nucleotides, from about 16 nucleotides to about 20 nucleotides, from about 16 nucleotides to about 21 nucleotides, from about 16 nucleotides to about 22 nucleotides, or from about 1 6 nucleotides to about 23 nucleotides, about 16 nucleotides to about 24 nucleotides, about 16 nucleotides to about 25 nucleotides, about 16 nucleotides to about 28 nucleotides, about 17 nucleotides to about 18 nucleotides, about 17 nucleotides to about 19 nucleotides, about 17 nucleotides to about 20 nucleotides, about 17 nucleotides to about 21 nucleotides, about 17 nucleotides to about 22 nucleotides, about 17 nucleotides to about 23 nucleotides, about 17 nucleotides to about 24 nucleotides, about 17 nucleotides to about 25 nucleotides, about 17 nucleotides to about 28 nucleotides, about 18 nucleotides to about 19 nucleotides, about 18 nucleotides to about 20 nucleotides, about 18 nucleotides to about 21 nucleotides, about 18 nucleotides to about 22 nucleotides, about 18 nucleotides to about 23 nucleotides,About 18 nucleotides to about 24 nucleotides, about 18 nucleotides to about 25 nucleotides, about 18 nucleotides to about 28 nucleotides, about 19 nucleotides to about 20 nucleotides, about 19 nucleotides to about 21 nucleotides, about 19 nucleotides to about 22 nucleotides, about 19 nucleotides to about 23 nucleotides, about 19 nucleotides to about 24 nucleotides, about 19 nucleotides to about 25 nucleotides, about 19 nucleotides to about 28 nucleotides, about 20 nucleotides to about 21 nucleotides, about 20 nucleotides to about 22 nucleotides, about 20 nucleotides to about 23 nucleotides, about 20 nucleotides to about 24 nucleotides, about 20 nucleotides to about 25 nucleotides, about 20 nucleotides to about 2 The amino acid sequence comprises 8 nucleotides, about 21 nucleotides to about 22 nucleotides, about 21 nucleotides to about 23 nucleotides, about 21 nucleotides to about 24 nucleotides, about 21 nucleotides to about 25 nucleotides, about 21 nucleotides to about 28 nucleotides, about 22 nucleotides to about 23 nucleotides, about 22 nucleotides to about 24 nucleotides, about 22 nucleotides to about 25 nucleotides, about 22 nucleotides to about 28 nucleotides, about 23 nucleotides to about 24 nucleotides, about 23 nucleotides to about 25 nucleotides, about 23 nucleotides to about 28 nucleotides, about 24 nucleotides to about 25 nucleotides, about 24 nucleotides to about 28 nucleotides, or about 25 nucleotides to about 28 nucleotides. In some embodiments, the guide RNA comprises about 15 nucleotides, about 16 nucleotides, about 17 nucleotides, about 18 nucleotides, about 19 nucleotides, about 20 nucleotides, about 21 nucleotides, about 22 nucleotides, about 23 nucleotides, about 24 nucleotides, about 25 nucleotides, or about 28 nucleotides.
[0135] In some embodiments, the guide nucleic acid further comprises a primer binding site (PBS). In some embodiments, the primer binding site is 3' to the guide nucleic acid. In some embodiments, the primer binding site comprises at least 2, 4, 6, 8, 10, 13, 16, 20, 24, 28, 32, 36, 40, 45, 50, 55, 60, or 65 nucleotides. In some embodiments, the primer binding site comprises less than 2, 4, 6, or 8 nucleotides.
[0136] In some embodiments, the guide nucleic acid further comprises a reverse transcriptase template (RTT). In some embodiments, the bases in the RTT comprise a bulk modification selected from the group consisting of complex sugars, complex amino groups, and / or other modifications compatible with RNA. In some embodiments, the RTT is fused to the guide RNA. In some embodiments, the guide nucleic acid further comprises a homologous sequence that is complementary to a region in the unedited DNA strand. In some embodiments, the guide nucleic acid comprises a nucleic acid template. In some embodiments, the RTT has a length of at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides. In some, the RTT has a length of at least about 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides. In some embodiments, the RTT has a length of at least about 1000, 2000, 3000, 4000, or 5000 nucleotides. In some embodiments, the RTT has a length of about 10 to about 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, or more than 200 nucleotides. In some embodiments, the RTT has a length of about 20 to about 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, or more than 200 nucleotides. In some embodiments, the RTT has a length of about 30 to about 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, or more than 200 nucleotides. In some embodiments, the RTT has a length of about 40 to about 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, or more than 200 nucleotides. In some embodiments, the RTT has a length of about 50 to about 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, or more than 200 nucleotides. In some embodiments, the RTT has a length of about 60 to about 70, 80, 90, 100, 120, 140, 160, 180, 200, or more than 200 nucleotides.In some embodiments, the RTT has a length of about 70 to about 80, 90, 100, 120, 140, 160, 180, 200, or more than 200 nucleotides. In some embodiments, the RTT has a length of about 80 to about 100, 120, 140, 160, 180, 200, or more than 200 nucleotides. In some embodiments, the RTT has a length of about 100 to about 120, 140, 160, 180, 200, or more than 200 nucleotides. In some embodiments, the RTT has a length of about 100 to about 4000 nucleotides. In some embodiments, the RTT has a length of about 100 to about 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, or 4000 nucleotides. In some embodiments, the RTT has a length of about 500 to about 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, or 4000 nucleotides. In some embodiments, the RTT has a length of about 1000 to about 1500, 2000, 2500, 3000, 3500, or 4000 nucleotides. In some embodiments, the RTT has a length of about 2000 to about 2500, 3000, 3500, or 4000 nucleotides. In some embodiments, the RTT has a length of about 3000 to about 3500, or 4000 nucleotides.
[0137] Methods for producing guide nucleic acids are known in the art. For example, guide RNAs and pegRNAs, as well as modified guide RNAs and pegRNAs, can be chemically synthesized. Furthermore, the nucleic acid sequence encoding the guide nucleic acid can be cloned into a vector and transcribed from the vector in vitro or in vivo using RNA polymerase.
[0138] Delivery and Vectors In some embodiments, disclosed herein are nucleic acid sequences encoding the reverse transcriptases, fusion proteins, or gene editing systems described herein.
[0139] In some embodiments, the nucleic acid encoding the endonuclease system or a component thereof is DNA, e.g., linear DNA, plasmid DNA, or minicircle DNA. In some embodiments, the nucleic acid encoding the reverse transcriptase, fusion protein, or gene editing system described herein is RNA, e.g., mRNA.
[0140] In some embodiments, the nucleic acid encoding the reverse transcriptase, fusion protein, or gene editing system described herein is delivered by a nucleic acid-based vector. In some embodiments, the nucleic acid-based vector is a plasmid (e.g., a circular DNA molecule that can replicate autonomously inside a cell), a cosmid (e.g., a pWE or sCos vector), an artificial chromosome, a human artificial chromosome (HAC), a yeast artificial chromosome (YAC), a bacterial artificial chromosome (BAC), a P1-derived artificial chromosome (PAC), a phagemid, a phage derivative, a bacmid, or a virus. In some embodiments, the nucleic acid-based vector is selected from the group consisting of pSF-CMV-NEO-NH2-PPT-3XFLAG, pSF-CMV-NEO-COOH-3XFLAG, pSF-CMV-PURO-NH2-GST-TEV, pSF-OXB20-COOH-TEV-FLAG(R)-6His, pCEP4 pDEST27, pSF-CMV-Ub-KrYFP, pSF-CMV-FMDV-daGFP, pEF1a-mCherry-N1 vector, pEF1a-tdTomato vector, pSF-CMV-FMDV-Hygro, pSF-CMV-PGK-Puro, pMCP-tag(m), pSF-CMV-PURO-NH2-CMYC, pSF-OXB20-BetaGal, pSF-OXB20-Fluc, pSF-OXB20, pSF-Tac, pRI 101-AN. The vector is selected from the list consisting of DNA, pCambia2301, pTYB21, pKLAC2, pAc5.1 / V5-His A, and pDEST8.
[0141] In some embodiments, the nucleic acid-based vector comprises a promoter. In some embodiments, the promoter is selected from the group consisting of a minipromoter, an inducible promoter, a constitutive promoter, and derivatives thereof. In some embodiments, the promoter is selected from the group consisting of CMV, CBA, EF1a, CAG, PGK, TRE, U6, UAS, T7, Sp6, lac, araBad, trp, Ptac, p5, p19, p40, synapsin, CaMKII, GRK1, and derivatives thereof. In some embodiments, the promoter is a U6 promoter. In some embodiments, the promoter is a CAG promoter.
[0142] In some embodiments, the nucleic acid-based vector is a virus. In some embodiments, the virus is an alphavirus, parvovirus, adenovirus, AAV, baculovirus, dengue virus, lentivirus, herpesvirus, poxvirus, anellovirus, bocavirus, vaccinia virus, or retrovirus. In some embodiments, the virus is an alphavirus. In some embodiments, the virus is a parvovirus. In some embodiments, the virus is an adenovirus. In some embodiments, the virus is an AAV. In some embodiments, the virus is a baculovirus. In some embodiments, the virus is a dengue virus. In some embodiments, the virus is a lentivirus. In some embodiments, the virus is a herpesvirus. In some embodiments, the virus is a poxvirus. In some embodiments, the virus is anellovirus. In some embodiments, the virus is a bocavirus. In some embodiments, the virus is a vaccinia virus. In some embodiments, the virus is a retrovirus.
[0143] In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-rh8, AAV-rh 10, AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-1, AAV-hu37, AAV-Anc80, AAV-Anc80L65, AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, In some embodiments, the herpesvirus is HSV-1, HSV-2, VZV, EBV, CMV, HHV-6, HHV-7, or HHV-8.
[0144] In some embodiments, the virus is AAV1 or a derivative thereof. In some embodiments, the virus is AAV2 or a derivative thereof. In some embodiments, the virus is AAV3 or a derivative thereof. In some embodiments, the virus is AAV4 or a derivative thereof. In some embodiments, the virus is AAV5 or a derivative thereof. In some embodiments, the virus is AAV6 or a derivative thereof. In some embodiments, the virus is AAV7 or a derivative thereof. In some embodiments, the virus is AAV8 or a derivative thereof. In some embodiments, the virus is AAV9 or a derivative thereof. In some embodiments, the virus is AAV10 or a derivative thereof. In some embodiments, the virus is AAV11 or a derivative thereof. In some embodiments, the virus is AAV12 or a derivative thereof. In some embodiments, the virus is AAV13 or a derivative thereof. In some embodiments, the virus is AAV14 or a derivative thereof. In some embodiments, the virus is AAV15 or a derivative thereof. In some embodiments, the virus is AAV16 or a derivative thereof. In some embodiments, the virus is AAV-rh8 or a derivative thereof. In some embodiments, the virus is AAV-rh10 or a derivative thereof. In some embodiments, the virus is AAV-rh20 or a derivative thereof. In some embodiments, the virus is AAV-rh39 or a derivative thereof. In some embodiments, the virus is AAV-rh74 or a derivative thereof. In some embodiments, the virus is AAV-rhM4-1 or a derivative thereof. In some embodiments, the virus is AAV-hu37 or a derivative thereof. In some embodiments, the virus is AAV-Anc80 or a derivative thereof. In some embodiments, the virus is AAV-Anc80L65 or a derivative thereof. In some embodiments, the virus is AAV-7m8 or a derivative thereof. In some embodiments, the virus is AAV-PHP-B or a derivative thereof. In some embodiments, the virus is AAV-PHP-EB or a derivative thereof.In some embodiments, the virus is AAV-2.5 or a derivative thereof. In some embodiments, the virus is AAV-2tYF or a derivative thereof. In some embodiments, the virus is AAV-3B or a derivative thereof. In some embodiments, the virus is AAV-LK03 or a derivative thereof. In some embodiments, the virus is AAV-HSC1 or a derivative thereof. In some embodiments, the virus is AAV-HSC2 or a derivative thereof. In some embodiments, the virus is AAV-HSC3 or a derivative thereof. In some embodiments, the virus is AAV-HSC4 or a derivative thereof. In some embodiments, the virus is AAV-HSC5 or a derivative thereof. In some embodiments, the virus is AAV-HSC6 or a derivative thereof. In some embodiments, the virus is AAV-HSC7 or a derivative thereof. In some embodiments, the virus is AAV-HSC8 or a derivative thereof. In some embodiments, the virus is AAV-HSC9 or a derivative thereof. In some embodiments, the virus is AAV-HSC10 or a derivative thereof. In some embodiments, the virus is AAV-HSC11 or a derivative thereof. In some embodiments, the virus is AAV-HSC12 or a derivative thereof. In some embodiments, the virus is AAV-HSC13 or a derivative thereof. In some embodiments, the virus is AAV-HSC14 or a derivative thereof. In some embodiments, the virus is AAV-HSC15 or a derivative thereof. In some embodiments, the virus is AAV-TT or a derivative thereof. In some embodiments, the virus is AAV-DJ / 8 or a derivative thereof. In some embodiments, the virus is AAV-Myo or a derivative thereof. In some embodiments, the virus is AAV-NP40 or a derivative thereof. In some embodiments, the virus is AAV-NP59 or a derivative thereof. In some embodiments, the virus is AAV-NP22 or a derivative thereof. In some embodiments, the virus is AAV-NP66 or a derivative thereof. In some embodiments, the virus is AAV-HSC16 or a derivative thereof.
[0145] In some embodiments, the virus is HSV-1 or a derivative thereof. In some embodiments, the virus is HSV-2 or a derivative thereof. In some embodiments, the virus is VZV or a derivative thereof. In some embodiments, the virus is EBV or a derivative thereof. In some embodiments, the virus is CMV or a derivative thereof. In some embodiments, the virus is HHV-6 or a derivative thereof. In some embodiments, the virus is HHV-7 or a derivative thereof. In some embodiments, the virus is HHV-8 or a derivative thereof.
[0146] In some embodiments, nucleic acids encoding the reverse transcriptase, fusion protein, or gene editing system described herein are delivered by a non-nucleic acid-based delivery system (e.g., a non-viral delivery system). In some embodiments, the non-viral delivery system is a liposome. In some embodiments, the nucleic acid is associated with a lipid. The lipid-associated nucleic acid is, in some embodiments, encapsulated in the aqueous interior of the liposome, interspersed within the lipid bilayer of the liposome, attached to the liposome via a linking molecule associated with both the liposome and the nucleic acid, entrapped in the liposome, complexed with the liposome, dispersed in a solution containing a lipid, mixed with a lipid, combined with a lipid, contained as a suspension in a lipid, contained in or complexed with a micelle, or otherwise associated with a lipid. In some embodiments, the nucleic acid is contained in a lipid nanoparticle (LNP).
[0147] In some embodiments, the reverse transcriptase, fusion protein, or gene editing system described herein is introduced into a cell by any suitable method, either stably or transiently. In some embodiments, the reverse transcriptase, fusion protein, or gene editing system described herein is transfected into a cell. In some embodiments, the cell is transduced or transfected with a nucleic acid construct encoding the reverse transcriptase, fusion protein, or gene editing system described herein. For example, the cell is transduced (e.g., with a virus encoding the reverse transcriptase, fusion protein, or gene editing system described herein) or transfected (e.g., with a plasmid encoding the reverse transcriptase, fusion protein, or gene editing system described herein) with a nucleic acid encoding the reverse transcriptase, fusion protein, or gene editing system described herein, or with the reverse transcriptase, fusion protein, or gene editing system described herein. In some embodiments, the transduction is stable or transient transduction. In some embodiments, cells expressing or containing the reverse transcriptase, fusion protein, or gene editing system described herein are transduced or transfected with one or more gRNA molecules, e.g., when the reverse transcriptase, fusion protein, or gene editing system described herein comprises a CRISPR nuclease. In some embodiments, a plasmid expressing the reverse transcriptase, fusion protein, or gene editing system described herein is introduced into the cell through electroporation, transient (e.g., lipofection) and stable genomic integration (e.g., piggybac), viral transduction (e.g., lentivirus or AAV), or other methods known to those skilled in the art. In some embodiments, the gene editing system is introduced into the cell as one or more polypeptides. In some embodiments, delivery is achieved through the use of RNP complexes. Methods for delivering polypeptides and / or RNPs to cells, for example, by electroporation or by cell squeezing, are known in the art.
[0148] Exemplary methods for nucleic acid delivery include lipofection, nucleofection, electroporation, stable genome integration (e.g., piggyback), microinjection, biolistics, virosomes, liposomes, immunoliposomes, polycations or lipid-nucleic acid conjugates, naked DNA, artificial virions, and drug-enhanced uptake of DNA. Lipofection is described, for example, in U.S. Pat. Nos. 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam™, Lipofectin™, and SF Cell Line 4D-Nucleofector X Kit™ (Lonza)). Cationic and neutral lipids suitable for efficient receptor-recognition lipofection of polynucleotides include the lipids of WO91 / 17424 and WO91 / 16024. In some embodiments, delivery is to a cell (e.g., in vitro or ex vivo administration) or a target tissue (e.g., in vivo administration). In some embodiments, the nucleic acid is contained in a liposome or nanoparticle that specifically targets the host cell.
[0149] Additional methods for delivery of nucleic acids into cells are known to those of skill in the art, see, e.g., US2003 / 0087817.
[0150] In some embodiments, the present disclosure provides a cell comprising a vector or nucleic acid described herein. In some embodiments, the cell expresses a gene editing system or portion thereof. In some embodiments, the cell is a human cell. In some embodiments, the cell is genome edited ex vivo. In some embodiments, the cell is genome edited in vivo.
[0151] cell In certain embodiments, described herein are cells comprising the systems or vectors described herein.
[0152] In some embodiments, the cell is a eukaryotic cell (e.g., a plant cell, an animal cell, a protist cell, or a fungal cell), a mammalian cell (Chinese hamster ovary (CHO) cell, baby hamster kidney (BHK), human embryonic kidney (HEK), mouse myeloma (NS0), or human retinal cell), an immortalized cell (e.g., a HeLa cell, a COS cell, a HEK-293T cell, an MDCK cell, a 3T3 cell, a PC12 cell, a Huh7 cell, a HepG2 cell, a K562 cell, an N2a cell, or a SY5Y cell), an insect cell (e.g., a Spodoptera frugiperda cell, a Trichoplusia ni cell, a Drosophila melanogaster cell, an S2 cell, or a Heliothis virescens cell), a yeast cell (e.g., a Saccharomyces In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is an immortalized cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell is a yeast cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is a fungal cell. In some embodiments, the cell is a prokaryotic cell.
[0153] In some embodiments, the cells are A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof.
[0154] How to use In some embodiments, described herein are methods for modifying double-stranded and / or single-stranded nucleic acids, the methods comprising: a) providing a cell with a guide nucleic acid that binds to a target strand of the double-stranded nucleic acid; b) providing the cell with a nuclease or nickase to cleave the double-stranded nucleic acid at the site of binding of the guide nucleic acid; and c) providing the cell with a reverse transcriptase to synthesize a modification in the target strand of the double-stranded nucleic acid at the site of cleavage by the nickase and / or double-stranded nuclease.
[0155] In some embodiments, the method is used to introduce a modification into the genome of a cell. In some embodiments, the modification is an insertion, deletion, or mutation. In some embodiments, the method is used to introduce site-specific insertions, deletions, and / or mutations (e.g., insertions and mutations) into the genome of a cell. In some embodiments, the method is used in combination with a nucleic acid template to facilitate site-specific insertion into the genome of a cell. In some embodiments, the cell is a human cell. In some embodiments, the genome of the cell or a vector contained in the cell is modified. In some embodiments, the genome of the cell is modified ex vivo. In some embodiments, the genome of the cell is modified in vivo.
[0156] In some embodiments, the method further comprises providing a transposase, an integrase, or a homing endonuclease to the cell. In some embodiments, the method further comprises providing a retrotransposon to the cell. In some embodiments, the method further comprises providing an RNA or DNA insertion template.
[0157] In some embodiments, the methods described herein further include detecting the genomic modification. In some embodiments, after the cell genome is modified, the cells are cultured for a certain period of time. In some embodiments, DNA or RNA is extracted and sequenced, and the modified sequence region is mapped and compared to the unmodified sequence. In some embodiments, the cells are stained with an antibody to the protein product translated from the modified nucleic acid, and the resulting stained protein or polypeptide in the cells is analyzed, for example, by flow cytometry.
[0158] The methods described herein can be used, for example, for targeted SNP correction, small insertions, or small deletions. In addition, the methods described herein can be used for targeted insertion of large templates into the genome of a cell by using a suitable RTT.
[0159] kit In some embodiments, the present disclosure provides kits that include one or more nucleic acid constructs encoding various components of the fusion proteins or genome editing systems described herein, including, for example, nucleotide sequences encoding components of a fusion protein or genome editing system capable of modifying a target DNA sequence. In some embodiments, the nucleotide sequences include a heterologous promoter that drives expression of the RNA genome editing system components.
[0160] In some embodiments, any of the targetable reverse transcriptase or genome editing systems disclosed herein are incorporated into pharmaceutical, diagnostic, or research kits to facilitate their use in therapeutic, diagnostic, or research applications. The kits may include one or more containers housing any of the vectors disclosed herein and instructions for use.
[0161] Kits can be designed to facilitate use of the methods described herein by researchers and can take many forms. Each of the components of the kit can be provided in liquid form (e.g., in solution) or solid form (e.g., dry powder), where applicable. In some embodiments, the compositions are configurable or otherwise processable (e.g., to an active form), for example, by the addition of a suitable solvent or other species (e.g., water or cell culture medium), which may or may not be provided with the kit. As used herein, "instructions" defines instructional and / or promotional components and can typically involve written instructions on or associated with the packaging of the present disclosure. Instructions can also include any oral or electronic instructions, such as audiovisual (e.g., videotape, DVD, etc.), internet, and / or web-based communications, provided in any manner such that the user clearly recognizes that the instructions are associated with the kit. The written instructions, in some embodiments, are in a form prescribed by a government agency that regulates the manufacture, use, or sale of pharmaceutical or biological products, and the instructions may also reflect approval by the agency of manufacture, use, or sale for animal administration. [Example]
[0162] The following examples are given for the purpose of illustrating various embodiments of the present disclosure and are not intended to limit the disclosure in any way. The examples, together with the methods described herein, represent currently preferred embodiments, are exemplary, and are not intended as limitations on the scope of the present disclosure. Modifications therein and other uses encompassed within the spirit of the present disclosure as defined by the scope of the claims will occur to those skilled in the art.
[0163] Example 1. Bioinformatics Identification of Reverse Transcriptases This example describes the identification of proteins with reverse transcriptase function by a bioinformatics approach.
[0164] Extensive assembly-driven metagenomic databases of microbial, viral, and eukaryotic genomes were bioinformatically analyzed for proteins with putative reverse transcriptase function. The analysis revealed millions of proteins with predicted reverse transcriptase function. The predicted RT hits were then bioinformatically filtered for complete open reading frames (ORFs) with high-quality RT domain hits, covering more than 70% of the canonical RT domain and containing predicted catalytic residues. After filtering, 29 RTs were selected for their potential to develop gene editing tools (SEQ ID NOS: 88-116 and 138-139). For all of these identified putative RTs, the predicted active site tetrad motif was [Y / F]ADD, with the most frequent amino acids at position 1 of the tetrad being tyrosine (Y, 96.6%) versus phenylalanine (F, 3.4%). The aspartate dyad (DD) is the most conserved feature of RT activity.
[0165] Example 2. Reverse Transcriptase (RT) for Short Corrections, Small Insertions, and Deletions This example describes the use of untethered reverse transcriptase in combination with pegRNA for targeted genome editing in HEK293T cells.
[0166] Testing reverse transcriptase candidates from nickases Reverse transcriptase (RT) candidates from the MG153 and MG160 families were cloned into a plasmid in which expression of the RT candidate was driven by a CMV promoter. The plasmid was isolated for transfection into HEK293T cells. A second plasmid containing the nickase spCas9 (H840A), whose expression is driven by a CMV promoter, and the RT-containing plasmid were co-transfected. Chemically synthesized PEG-RNA (SEQ ID NOs: 6-29) containing the desired edit in the RT template was transfected. All components (plasmid and PEG-RNA) were reverse transfected into 150,000 HEK293T cells in a 24-well plate. 72 hours after transfection, cells were lysed in 100 μL of DNA extraction solution. Next-generation sequencing (NGS) barcode-containing primers (SEQ ID NOs: 30-31) were used to amplify an approximately 250 bp target (SEQ ID NO: 32). PCR cleanup was then performed, and the samples were subjected to NGS sequencing. The FASTQ files were then processed to determine the percentage of reads with the desired alterations.
[0167] MG153 family The untethered MG153 candidates, MG153-23 and MG153-24 (SEQ ID NOs: 2-3), were tested for prime editing in HEK293T cells to determine the desired rate of correction. The editing rates for each RT are shown in Figure 1 for each pegRNA with various PBS lengths (2, 4, 6, 8, 10, 13, 16, and 20 nucleotides).
[0168] MG160 family An untethered MG160 family candidate, MG160-7 (SEQ ID NO: 5), was tested in mammalian cells for the above activities, and no activity was detected above background in this assay (FIG. 2).
[0169] Testing nickase-tethered reverse transcriptase candidates The activity of various RT classes was evaluated using CRISPR type II nucleases. RT candidates were cloned into a plasmid containing the nickase spCas9 (H840A) to generate RT-nickase fusions. A CMV promoter drove the expression of RT-nickase fusion proteins containing a 33-amino acid linker (SEQ ID NO: 33) between the nickase and the RT candidate. The fusion proteins were then transfected into HEK293T cells and processed for NGS as described above.
[0170] The activity of tethered MG160-7 (SEQ ID NO: 5) is shown in Figure 3. The activity detected in this assay was not above background.
[0171] The above data demonstrate the identification of candidate RTs that showed comparable or even greater activity than MMLV WT in the context of prime editing. Activity across a broad family of RTs allows for the identification of candidate RTs for different types of genome modifications (i.e., SNP correction, insertions, or deletions). The small size of the identified RTs (one-third that of MMLV WT) allows for efficient delivery using adeno-associated virus (AAV) and lipid nanoparticles (LNPs).
[0172] Example 3. RT for short corrections, small insertions, and deletions (predictive) Additional RTs from the MG153 family, including MG153-22 (SEQ ID NO: 1), will be tested in an untethered format as described in Example 2. This will allow for the identification of additional RT candidates for small corrections, insertions, and deletions.
[0173] References Anzalone AV,Gao XD,Podracky CJ,Nelson AT,Koblan LW,Raguram A,Levy JM,Mercer JAM,Liu DR.Programmable deletion,replacement,integration and inversion of large DNA sequences with twin prime editing.Nat Biotechnol.2022,40(5):731-740.doi:10.1038 / s41587-021-01133-w.Epub 2021 Dec 9.PMID:34887556;PMCID:PMC9117393.
[0174] Clement K,Rees H,Canver MC,Gehrke JM,Farouni R,Hsu JY,Cole MA,Liu DR,Joung JK,Bauer DE,Pinello L.CRISPResso2 provides accurate and rapid genome editing sequence analysis.Nat Biotechnol.2019,37(3):224-226.doi:10.1038 / s41587-019-0032-3.PMID:30809026;PMCID:PMC6533916.
[0175] Gonzalez-Delgado A,Mestre MR,Martinez-Abarca F,Toro N.Prokaryotic reverse transcriptases:from retroelements to specialized defense systems.FEMS Microbiol Rev.2021,45(6):fuab025.doi:10.1093 / femsre / fuab025.PMID:33983378;PMCID:PMC8632793.
[0176] He S,Corneloup A,Guynet C,Lavatine L,Caumont-Sarcos A,Siguier P,Marty B,Dyda F,Chandler M,Ton Hoang B.The IS200 / IS605 Family and “Peel and Paste”Single-strand Transposition Mechanism.Microbiol Spectr.2015,3(4).Doi:10.1128 / microbiolspec.MDNA3-0039-2014.PMID:26350330.
[0177] equivalent The present disclosure may be embodied in other specific forms without departing from their spirit or essential characteristics. Accordingly, the foregoing embodiments are to be considered in all respects as illustrative and not limiting of the disclosure described herein. The scope of the present disclosure is, therefore, indicated by the appended claims, rather than the foregoing description, and all changes that come within the meaning and range of equivalency of the claims are intended to be embraced therein. [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7] [Table 2-8] Table 2-9 Table 2-10 Table 2-11 Table 2-12 Table 2-13 Table 2-14 Table 2-15 Table 2-16 Table 2-17 Table 2-18 Table 2-19 Table 2-20 Table 2-21 Table 2-22 Table 2-23 Table 2-24 Table 2-25 Table 2-26 Table 2-27 Table 2-28 Table 2-29 Table 2-30 Table 2-31 Table 2-32
Claims
1. A fusion protein comprising a nickase linked to a reverse transcriptase using a linker, wherein the reverse transcriptase comprises at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
2. A fusion protein comprising a nuclease linked to a reverse transcriptase using a linker, wherein the reverse transcriptase comprises at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
3. A fusion protein comprising a catalytically deficient nuclease linked to a reverse transcriptase using a linker, wherein the reverse transcriptase comprises at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
4. A gene editing system, a) nickase; b) a guide nucleic acid configured to form a complex with the nickase and hybridize to a target nucleic acid sequence; c) a reverse transcriptase comprising at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139, and configured to form a complex with the nickase.
5. 5. The gene editing system of Claim 4, wherein the nickase is a modifying endonuclease.
6. 6. The gene editing system of claim 5, wherein the modified endonuclease is a type II CRISPR endonuclease.
7. 6. The gene editing system of claim 5, wherein the modified endonuclease is a type V CRISPR endonuclease.
8. The gene editing system of claim 6 or 7, wherein the type II CRISPR endonuclease or the type V CRISPR endonuclease has nickase activity.
9. 6. The gene editing system of claim 5, wherein the modified endonuclease is selected from the group consisting of spCas9(H840A), spCas9(D10A), nMG3-6(D13A), nMG3-6(H586A), nMG3-6(N609A), Cas12a, and MG29-1.
10. 6. The gene editing system of Claim 5, wherein the modified endonuclease comprises at least about 80% sequence identity to any one of SEQ ID NOs: 117-119.
11. The gene editing system of any one of claims 4 to 10, wherein the nickase and the reverse transcriptase are fused together.
12. The gene editing system of any one of claims 4 to 10, wherein the nickase and the reverse transcriptase are linked by a linker.
13. 12. The gene editing system of Claim 11, wherein the linker comprises at least 10, 20, or 30 amino acids.
14. 12. The gene editing system of claim 11, wherein the linker comprises about 30 to 35 amino acids.
15. 12. The gene editing system of Claim 11, wherein the linker comprises about 30 amino acids.
16. 12. The gene editing system of Claim 11, wherein the linker comprises at least 80% sequence identity to SEQ ID NO:
33.
17. 12. The gene editing system of Claim 11, wherein the linker comprises at least 80% sequence identity to any one of SEQ ID NOs: 82-87.
18. The gene editing system of any one of claims 4 to 17, wherein the guide nucleic acid comprises a spacer sequence and crRNA.
19. The gene editing system of any one of claims 4 to 18, wherein the guide nucleic acid further comprises a reverse transcriptase template (RTT).
20. 20. The gene editing system of Claim 19, wherein the bases in the RTT comprise a bulk modification selected from the group of complex sugars, or complex amino groups, and / or other modifications compatible with RNA.
21. The gene editing system of any one of claims 4 to 20, wherein the guide nucleic acid further comprises a primer binding site.
22. 22. The gene editing system of Claim 21, wherein the primer binding site is at the 3' end of the guide nucleic acid.
23. 23. The gene editing system of Claim 21 or 22, wherein the primer binding site comprises at least 2, 4, 6, 8, 10, 13, 16, 20, 24, 28, 32, 36, 40, 45, 50, 55, 60, or 65 nucleotides.
24. The gene editing system of any one of claims 4 to 23, further comprising a transposase, integrase, or homing endonuclease.
25. The gene editing system of any one of claims 4 to 23, further comprising a retrotransposon.
26. 26. The gene editing system of any one of claims 4-25, wherein the reverse transcriptase comprises at least about two-fold higher processivity than Moloney murine leukemia virus (MMLV) reverse transcriptase.
27. 26. The gene editing system of any one of claims 4-25, wherein the reverse transcriptase comprises a processivity that is at least about 2-fold lower than Moloney murine leukemia virus (MMLV) reverse transcriptase.
28. 28. The gene editing system of any one of claims 4-27, wherein the reverse transcriptase comprises an error rate of less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10%, or 0.05%.
29. 29. The gene editing system of any one of claims 4-28, wherein the reverse transcriptase comprises an error rate of less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10%, or 0.05% compared to Moloney murine leukemia virus (MMLV) reverse transcriptase.
30. A gene editing system, a) a nuclease; b) a guide nucleic acid configured to form a complex with the nuclease and hybridize to a target nucleic acid sequence; c) a reverse transcriptase having at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139, and configured to form a complex with the nuclease.
31. 31. The gene editing system of Claim 30, wherein the nuclease is a double-stranded nuclease.
32. 32. The gene editing system of claim 30 or 31, wherein the nuclease is a type II CRISPR endonuclease.
33. 33. The gene editing system of Claim 32, wherein the CRISPR endonuclease is Cas9.
34. 34. The gene editing system of Claim 33, wherein the Cas9 is catalytically deficient Cas9 (dCas9).
35. The gene editing system of any one of claims 30 to 34, wherein the nuclease and the reverse transcriptase are fused together.
36. 36. The gene editing system of Claim 35, wherein the nuclease and the reverse transcriptase are linked by a linker.
37. 37. The gene editing system of Claim 36, wherein the linker comprises at least 10, 20, or 30 amino acids.
38. 37. The gene editing system of Claim 36, wherein the linker comprises about 30 to 35 amino acids.
39. 37. The gene editing system of Claim 36, wherein the linker comprises about 30 amino acids.
40. 37. The gene editing system of Claim 36, wherein the linker comprises at least 80% sequence identity to SEQ ID NO:
33.
41. 37. The gene editing system of Claim 36, wherein the linker comprises at least 80% sequence identity to any one of SEQ ID NOs: 82-87.
42. The gene editing system of any one of claims 30 to 34, wherein the nuclease and the reverse transcriptase are unlinked.
43. The gene editing system of any one of claims 30 to 42, wherein the guide nucleic acid further comprises a primer binding site.
44. 44. The gene editing system of Claim 43, wherein the primer binding site is at the 3' end of the guide nucleic acid.
45. 45. The gene editing system of Claim 43 or 44, wherein the primer binding site comprises at least 2, 4, 6, 8, 10, 13, 16, 20, 24, 28, 32, 36, 40, 45, 50, 55, 60, or 65 nucleotides.
46. 46. The gene editing system of any one of claims 30 to 45, further comprising a transposase, integrase, or homing endonuclease.
47. The gene editing system of any one of claims 30 to 46, further comprising a retrotransposon.
48. 48. The gene editing system of any one of Claims 30-47, wherein the reverse transcriptase comprises at least about two-fold higher processivity than Moloney murine leukemia virus (MMLV) reverse transcriptase.
49. 48. The gene editing system of any one of Claims 30-47, wherein the reverse transcriptase comprises a processivity that is at least about 2-fold lower than Moloney murine leukemia virus (MMLV) reverse transcriptase.
50. 50. The gene editing system of any one of Claims 30-49, wherein the reverse transcriptase comprises an error rate of less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10%, or 0.05%.
51. 50. The gene editing system of any one of claims 30-49, wherein the reverse transcriptase comprises an error rate of less than about 2.5%, 2.0%, 1.5%, 1%, 0.5%, 0.25%, 0.10%, or 0.05% compared to Moloney murine leukemia virus (MMLV) reverse transcriptase.
52. A gene editing system, a) nickase; b) a guide nucleic acid configured to form a complex with the nickase and hybridize to a target nucleic acid sequence; c) a reverse transcriptase configured to form a complex with said nickase, said reverse transcriptase comprising: X 1 X 2 DD motif, X 1 is F or Y, and X 1 If Y, then X 2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y.
53. The X 2 is A or I.
54. The X 1 X 2 53. The gene editing system of claim 52, wherein the DD motif is YADD (SEQ ID NO: 140) or YIDD (SEQ ID NO: 141).
55. The X 1 X 2 53. The gene editing system of Claim 52, wherein the DD motif is FADD (SEQ ID NO: 142), FVDD (SEQ ID NO: 143), FIDD (SEQ ID NO: 144), or FLDD (SEQ ID NO: 145).
56. 56. The gene editing system of any one of Claims 52-55, wherein the reverse transcriptase has at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
57. A gene editing system, a) a nuclease; b) a guide nucleic acid configured to form a complex with the nuclease and hybridize to a target nucleic acid sequence; c) a reverse transcriptase configured to form a complex with said nuclease, said reverse transcriptase comprising X 1 X 2 DD motif, X 1 is F or Y, and X 1 If Y, then X 2 is A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, V, W, or Y.
58. The X 2 is A or I.
59. The X 1 X 2 58. The gene editing system of claim 57, wherein the DD motif is YADD (SEQ ID NO: 140) or YIDD (SEQ ID NO: 141).
60. The X 1 X 2 58. The gene editing system of claim 57, wherein the DD motif is FADD (SEQ ID NO: 142), FVDD (SEQ ID NO: 143), FIDD (SEQ ID NO: 144), or FLDD (SEQ ID NO: 145).
61. 61. The gene editing system of any one of Claims 57-60, wherein the reverse transcriptase has at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
62. An isolated reverse transcriptase having at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
63. A nucleic acid encoding the fusion protein of any one of claims 1 to 3 or the gene editing system of any one of claims 4 to 61.
64. 64. The nucleic acid of claim 63, wherein the nucleic acid is DNA or RNA.
65. 65. The nucleic acid of claim 64, wherein the RNA is mRNA.
66. A vector comprising the nucleic acid of any one of claims 63 to 65.
67. 67. An adeno-associated virus or lipid nanoparticle comprising a nucleic acid according to any one of claims 63 to 65 or a vector according to claim 66.
68. A cell comprising the nucleic acid of any one of claims 63 to 65 or the vector of claim 66.
69. 69. The cell of claim 68, wherein the cell is a human cell.
70. 62. A method for modifying double-stranded and / or single-stranded nucleic acid, the method comprising contacting a cell with a fusion protein of any one of claims 1 to 3 or a gene editing system of any one of claims 4 to 61.
71. 1. A method for modifying double-stranded and / or single-stranded nucleic acids, comprising: a) providing a cell with a guide nucleic acid that hybridizes to a target nucleic acid sequence; b) providing a nuclease or nickase to the cell to cleave the nucleic acid at the binding site of the guide nucleic acid; c) providing a reverse transcriptase to said cell to synthesize a modification in the target strand of said nucleic acid at the site of cleavage by said nickase and / or nuclease.
72. 72. The method of claim 71, wherein the reverse transcriptase has at least about 80% sequence identity to any one of SEQ ID NOs: 88-116 and 138-139.
73. 72. The method of claim 71, wherein the modification is an insertion, deletion, or mutation.
74. 72. The method of claim 71, further comprising providing an RNA or DNA template.
75. 72. The method of claim 71, wherein the nucleic acid is a genome or a vector.
76. 72. The method of claim 71, further comprising providing the cell with a transposase, integrase, or homing endonuclease.
77. 72. The method of claim 71, further comprising providing said cell with a retrotransposon.