Evolved adenine deaminase and RNA-guided nuclease fusion proteins with internal insertion sites and methods of use

Fusion proteins with internal insertion sites for deaminases in RNA-guided nucleases improve editing efficiency and specificity, addressing limitations in existing RGN systems for precise genetic modifications.

JP2025537710APending Publication Date: 2025-11-20LIFEEDIT THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025525634
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-02-17
Filing Date
2023-11-06
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

Existing RNA-guided nuclease (RGN) systems for genome editing, such as CRISPR-Cas proteins, face limitations in editing efficiency and specificity, particularly in modifying nucleobases outside the canonical editing window, which hampers precise genetic modifications in targeted genome editing applications.

Method used

Development of fusion proteins comprising RNA-guided nucleases (RGNs) with internal insertion sites for deaminases, such as adenine deaminases, which enhance editing activity and shift the editing window, allowing for improved targeted nucleic acid modifications.

Benefits of technology

The fusion proteins exhibit enhanced editing efficiency and specificity, enabling precise genetic modifications, including correcting genetic defects and introducing beneficial traits in mammalian and plant cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025537710000047
    Figure 2025537710000047
  • Figure 2025537710000048
    Figure 2025537710000048
  • Figure 2025537710000049
    Figure 2025537710000049
Patent Text Reader

Abstract

Compositions and methods are provided that include a deaminase for targeted editing of nucleic acids. Also provided are compositions and methods for localizing a heterologous polypeptide to a target DNA molecule and for targeted editing of nucleic acids. Fusion proteins are provided that include an RNA-guided nuclease (RGN) and at least one heterologous polypeptide inserted therein, and fusion proteins that include a DNA-binding polypeptide and a deaminase. The heterologous polypeptide can be a prime editing polypeptide or a base-editing polypeptide. The compositions also include a nucleic acid molecule encoding the deaminase or the fusion protein. Vectors and host cells are also provided that include a nucleic acid molecule encoding the deaminase or the fusion protein.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application Nos. 63 / 382,344, filed November 4, 2022, and 63 / 485,642, filed February 17, 2023, each of which is incorporated by reference in its entirety.

[0002] Reference to sequence listing submitted electronically as an XML file This application contains a Sequence Listing that has been submitted in xml format via the USPTO Patent Center and is incorporated herein by reference in its entirety. The xml copy, created on November 6, 2023, is named L103438_1330WO_Seq_List.xml and is 1.12MB in size.

[0003] The present invention relates to the fields of molecular biology and gene editing. [Background technology]

[0004] Targeted genome editing or modification is rapidly becoming an important tool for basic and applied research. Early methods involve engineering nucleases such as meganucleases, zinc finger fusion proteins, or TALENs, which require the generation of chimeric nucleases with engineered, programmable, sequence-specific DNA-binding domains specific for each particular target sequence. RNA-guided nucleases (RGNs), such as the CRISPR-associated (Cas) proteins of the clustered regularly interspaced short palindromic repeats (CRISPR)-Cas bacterial system, enable targeting of specific sequences by complexing the nuclease with a guide RNA that specifically hybridizes to the specific target sequence. Producing target-specific guide RNAs is less costly and more efficient than generating chimeric nucleases for each target sequence. Such RNA-guided nucleases can be used to edit genomes by introducing sequence-specific double-strand breaks that are repaired via error-prone non-homologous end joining (NHEJ) and introduce mutations at specific genomic locations.

[0005] Additionally, RGNs are useful for targeted DNA editing approaches. Targeted editing of nucleic acid sequences, such as targeted cleavage, to allow the introduction of specific modifications into genomic DNA allows for highly sensitive approaches to studying gene function and gene expression. RGNs can also be used to generate chimeric proteins that use the RNA-guided activity of RGNs in combination with DNA-modifying enzymes, such as deaminases, for targeted base editing. Targeted editing can be developed to target genetic diseases in humans or to introduce agronomically beneficial mutations into crop genomes. The development of genome editing tools provides new approaches to gene editing-based mammalian therapeutics and agrobiotechnology. Summary of the Invention

[0006]

[0003] Compositions and methods are provided for localizing a heterologous polynucleotide to a target DNA molecule and modifying the target DNA molecule. The provided compositions comprise a deaminase polypeptide and a fusion protein comprising at least one heterologous polypeptide, such as an RNA-guided nuclease (RGN) and a deaminase inserted therein. The heterologous polypeptide can be a prime editing polypeptide or a base-editing polypeptide. Fusion proteins comprising a base-editing polypeptide are also referred to herein as base editors. Thus, in some embodiments, the base editor comprises a nucleic acid molecule-binding polypeptide (e.g., a DNA-binding polypeptide) and a deaminase polypeptide. In some embodiments, the base editor comprises an RGN having a deaminase inserted therein. In some embodiments, a base editor comprising an RGN having a deaminase polypeptide inserted therein exhibits improved editing activity and / or a shifted editing window compared to the parent base editor. In some embodiments, the parent base editor comprises a deaminase fused to the amino terminus of an RGN. Also provided are nucleic acid molecules encoding the disclosed deaminase polypeptides and fusion proteins, vectors comprising the nucleic acid molecules, and cells comprising the deaminase polypeptides or fusion proteins, nucleic acid molecules, or vectors.

[0007] Systems comprising the disclosed fusion proteins (or nucleic acid molecules encoding them) and one or more guide RNAs (or one or more nucleic acid molecules encoding them), together with cells comprising the systems, are provided for localizing heterologous polypeptides or deaminases to target DNA molecules. Methods for localizing heterologous polypeptides or deaminases and / or modifying target DNA molecules include delivering the disclosed systems to a target DNA molecule or a cell comprising a target DNA molecule.

[0008] Pharmaceutical compositions are provided comprising the deaminases and fusion proteins of the present disclosure, nucleic acid molecules encoding same (or vectors or cells comprising same), or systems (or cells comprising same). Methods for treating a subject having or at risk of developing a disease, disorder, or condition include administering to the subject a fusion protein of the present disclosure, a nucleic acid molecule encoding same (or vectors or cells comprising same), a system (or cells comprising same), or a pharmaceutical composition. In some embodiments, the disease is associated with a causative mutation, and treating includes correcting the causative mutation. [Brief explanation of the drawings]

[0009] [Figure 1] Figures 1A and 1B depict a rational design approach to improve editing efficiency at bases outside the canonical editing window of the parent adenine base editor LPG50148-nAPG07433.1. Figure 1A provides a schematic cartoon of the linear base editor, showing the parent adenine base editor (ABE) as an end-to-end fusion of LPG50148 and nAPG07433.1, and the internal base editor in which LPG50148 is inserted within nAPG07433.1. Figure 1B shows a structural model of APG07433.1 with the insertion point (blue sphere). LPG50148 (green) was inserted after the position indicated. The insertion sites for G910, V872, and E900 are within the wedge domain, the insertion sites for T642 and R670 are within the HNH domain, and the insertion sites for D737, A802, S775, S772, K778, and R30 are within the RuvCIII domain. [Figure 2]Figures 2A and 2B show the editing efficiency and editing window of nAPG07433.1 inlaid base editors (IBEs) LPG20002 to LPG20006 and the HNH-substituted LPG20001. Figure 2A provides the total editing efficiency of the first five designed nAPG07433.1 IBEs tested on genomic targets for the total percentage of AG substitutions. Figure 2B shows the editing window of LPG20001 to LPG20006. The heat map provides the editing rate for each position. Rows represent five individual guides (from top to bottom: SGN001681, SGN001062, SGN000968, SGN000754, and SGN001064). Positions 1 to 25 represent the positions of the first to 25th targeted adenines within the protospacer, and position 1 is the first adenine 5' of the PAM. [Figure 3] A-C show the editing efficiency and editing window of nAPG07433.1 IBE LPG20007-LPG20011. The heatmap provides the editing rate for each position. Rows represent four individual guides (top to bottom: SGN001062, SGN000968, SGN000754, and SGN001064). [Figure 4] Figures A-C illustrate nAPG07433.1 IBEs 2 and 8. Figure 4A provides a structural representation of the insertion points of nAPG07433.1 IBEs 2 (S772) and 8 (T642). Figure 4B depicts the schematic domain organization of LPG20002 and LPG20008 compared to the parent ABE. Figure 4C shows the average editing efficiency of LPG20002 and LPG20008 for three target sites (SGN001947, SGN001064, and SGN000909) compared to the parent ABE. [Figure 5]Figures 5A-C show evaluation of nAPG07433.1 IBE2 with deleted linkers. Figure 5A depicts the insertion of truncated LPG50148 into nAPG07433.1 and deletion of the N-terminal linker, C-terminal linker, or both linkers to generate nAPG07433.1 IBE2 variants. Figure 5B shows the total editing efficiency (percentage of A-G substitutions) of all LPG20002 variants tested at four genomic targets. Figure 5C provides a heatmap showing the editing rate at each position. Rows represent four individual target sites (top to bottom: SGN001062, SGN001064, SGN000754, and SGN000968). [Figure 6] Figures 6A and 6B show the therapeutic potential of engineered IBE. Figure 6A provides results from disrupting a splice donor in exon 1 in the transthyretin (TTR) gene as a strategy for gene knockout. Plasmids were delivered into HEK cells via nucleofection. Figure 6B shows the use of LPG20002 variants to reduce β2-microglobulin (B2M) surface presentation. mRNA was delivered into T cells via nucleofection. [Figure 7A] Figure 1 shows the editing activity of nAPG05586 IBE. A provides the average editing rate across the editing window (adenines 5-24) of nAPG05586 IBE with four guide RNAs after delivery of a plasmid encoding the IBE and guide RNAs into HEK293T cells via nucleofection. [Figure 7B] Figure 1 shows the editing activity of nAPG05586 IBEs. A provides the editing rate at all adenine positions within the editing window for each of the nAPG05586 IBEs. [Figure 8A] Figure 1 shows the editing activity of nAPG01604 IBE. A provides the average editing rate across the editing window (adenines 5-19) of nAPG01604 IBE with four guide RNAs after delivery of a plasmid encoding the IBE and guide RNAs into HEK293T cells via nucleofection. [Figure 8B]Figure 1 shows the editing activity of nAPG01604 IBEs. A provides the editing rate at all adenine positions within the editing window for each of the nAPG01604 IBEs. [Figure 9A] Figure 1 shows the editing activity of nLPG10145 IBE. A provides the average editing rate across the editing window (adenines 4-25) of LPG10145 IBE with four guide RNAs after delivery of a plasmid encoding the IBE and guide RNA into HEK293T cells via nucleofection. [Figure 9B] Figure 1 shows the editing activity of nLPG10145 IBEs. Figure 2 provides the editing rate at all adenine positions within the editing window for each of the nLPG10145 IBEs. [Figure 10A] 1 depicts the first step of a directed evolution strategy to identify an adenine base editor variant with greater activity than the parent ABE (LPG50148 fused to dAPG07433.1). Shown is an end-to-end fusion protein of the LPG50148 adenosine deaminase variant and a nuclease-inactive version of the APG07433.1 RNA-guided nuclease, expressed in E. coli, that converts a stop codon to a sense codon in the kanamycin resistance gene, conferring resistance. [Figure 10B] This figure depicts the first step of a directed evolution strategy to identify adenine base editor variants with better activity than the parent ABE (LPG50148 fused to dAPG07433.1). It shows how evolved ABEs were enriched for those with better editing activity than the parent ABE. E. coli harboring a target plasmid was transformed with a library of mutant LPG50148 fused to dAPG07433.1. To grow in high concentrations of kanamycin, cells are required to repair the antibiotic resistance gene (KanR*) with higher editing efficiency than the parent ABE. Selection was performed separately for three sequence contexts: TAT, CAC, or TAG. After the enrichment step, variants with up to 8-fold improvement were identified. [Figure 11]We present a complete directed evolution strategy for identifying adenine base editor variants with superior activity compared to the parent ABE. ABE variants with single point mutations exhibiting up to eight times greater activity than the parent ABE were selected in bacteria and then evaluated in a first-round human cell screen in human HEK293T cells. Single point mutants identified in selections performed against the TAT or CAC sequence context were evaluated at eight endogenous genomic sites, while single point mutants identified in selections performed against the TAG sequence context were evaluated at four endogenous genomic sites. The single mutations showing the highest average fold changes were combined to generate the top hits: for TAT, 20 (LPG50148 with L35N, V81S, N156R, and L162W mutations, shown as SEQ ID NO: 319), 60 (LPG50148 with I75W and F155W mutations, shown as SEQ ID NO: 359), 67 (LPG50148 with V81S, F155W, and A160D mutations, shown as SEQ ID NO: 366), and 68 (LPG50148 with V81S, C145A, and F155W mutations, shown as SEQ ID NO: 370). 48, shown as SEQ ID NO: 367), for CAC, 106 (LPG50148 with D76G and V81S mutations, shown as SEQ ID NO: 405), and 107 (LPG50148 with V81S and M117L mutations, shown as SEQ ID NO: 406), and for TAG, 88 (LPG50148 with I40L and V105M mutations, shown as SEQ ID NO: 387), and 103 (LPG50148 with V81S, C145M, and A160W mutations, shown as SEQ ID NO: 402). [Figure 12] Higher editing efficiency of evolved ABE variants LPG50148.2.20, LPG50148.2.60, LPG50148.2.67, LPG50148.2.68, and LPG50148.2.69 compared to the parent ABE when the ABE was delivered as mRNA. [Figure 13]Figures 13A and 13B provide the results of robustness testing of evolved ABE variants LPG50148.2.20, LPG50148.2.60, LPG50148.2.67, LPG50148.2.68, LPG50148.2.103, LPG50148.2.106, and LPG50148.2.107. Figure 13A shows the percentage of guides (from 48 genomic sites) with total AT substitutions greater than 20% editing. Figure 13B shows the average AG substitutions across all genomic sites. There were three significant effects: variants LPG50148.2.20 (p≦0.001), LPG50148.2.103 (p<0.001), and LPG50148.2.67 (p=0.01). [Figure 14] Figure 1 shows the shift in the optimal editing window for three of the evolved ABE variants. The graph shows the average AT substitutions for the parent ABE and the three evolved ABE variants LPG50148.2.20, LPG50148.2.67, and LPG50148.2.103 by position, indicating the editor activity window. Numbers indicate positions within the protospacer upstream of the PAM sequence. The optimal window is ≥ 30% of maximum editing. [Figure 15] Figure 1 shows AG editing in different sequence contexts for multiple evolved ABE variants (LPG50148.2.20, LPG50148.2.60, LPG50148.2.67, LPG50148.2.68, LPG50148.2.103, LPG50148.2.106, and LPG50148.2.107) and the parent ABE. Editing rates reflect the median value for each context across 48 genomic sites. [Figure 16] 1 provides the AG conversion rate at position 11 in the target sequence of SGN008393 using LPG20047 in primary mouse hepatocytes. [Figure 17A] We show that engineering a third-generation deaminase based on a second-generation deaminase provided higher levels of AG editing when fused to the N-terminus of nAPG07433.1 against multiple targets. [Figure 17B]We show that engineering a third-generation deaminase based on a second-generation deaminase provided higher levels of AG editing when fused to the N-terminus of nAPG07433.1 against multiple targets. [Figure 17C] We show that engineering a third-generation deaminase based on a second-generation deaminase provided higher levels of AG editing when fused to the N-terminus of nAPG07433.1 against multiple targets. [Figure 18] Improved performance of the third-generation base editor LPG50324 fused to nAPG07433.1 is shown. [Figure 19] 1 shows improved editing utilizing LPG50310 deaminase relative to LPG50265 deaminase in HEK293T cells with plasmid delivery when fused to nAPG07433.1. DETAILED DESCRIPTION OF THE INVENTION

[0010] Many modifications and other embodiments of the inventions described herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions. It is to be understood, therefore, that the invention is not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.

[0011] I. Overview The present disclosure provides deaminases and fusion proteins comprising a nucleic acid molecule-binding polypeptide, such as a DNA-binding polypeptide, and a deaminase polypeptide. In certain embodiments, the DNA-binding polypeptide is a sequence-specific DNA-binding polypeptide in that the DNA-binding polypeptide binds to a target sequence more frequently than it binds to randomized background sequences. In some embodiments, the DNA-binding polypeptide is or is derived from a meganuclease, a zinc finger fusion protein, or a TALEN. In some embodiments, the fusion protein comprises an RNA-guided DNA-binding polypeptide and a deaminase polypeptide. In some embodiments, the fusion protein is an RNA-guided nuclease (RGN), such as an RNA-guided DNA-binding polypeptide, such as a CRISPR-Cas (e.g., Cas9) polypeptide, that binds to a guide RNA (also referred to as gRNA), which in turn binds to the target nucleic acid sequence via strand hybridization.

[0012] The present disclosure also provides fusion proteins (and nucleic acid molecules encoding the same) comprising RGN and at least one heterologous polypeptide inserted within RGN. The heterologous polypeptide can be inserted within a surface location of RGN (e.g., immediately after an amino acid residue on the surface of RGN). Fusion proteins of the present disclosure can comprise RGN with a base-editing polypeptide or a prime-editing polypeptide, some of which exhibit improved editing activity and / or a shifted editing window compared to end-to-end fusion proteins.

[0013] The terms "protein," "peptide," and "polypeptide" are used interchangeably herein and refer to a polymer of amino acid residues linked together by peptide (amide) bonds. These terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide is at least three amino acids long. A protein, peptide, or polypeptide may refer to an individual protein or a collection of proteins. One or more of the amino acids in a protein, peptide, or polypeptide may be modified by the addition of chemicals such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, linkers for conjugation, functionalization, or other modifications. A protein, peptide, or polypeptide may also be a single molecule or a multimolecular complex. A protein, peptide, or polypeptide may be simply a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, recombinant, synthetic, or any combination thereof.

[0014] The term "fusion protein" as used herein refers to a hybrid polypeptide comprising protein domains from at least two different proteins. A fusion protein may comprise more than one different domain, for example, RGN and a deaminase. The fusion protein of the present invention comprises a deaminase of the present disclosure and a nucleic acid molecule-binding polypeptide or a heterologous protein (e.g., a deaminase) inserted within the amino acid sequence of RGN, which in some cases can disrupt a domain within the RGN protein. In some embodiments, the fusion protein is in a complex with or associated with a nucleic acid, for example, RNA.

[0015] The heterologous polypeptide inserted into the RGN according to the present invention can be a base-editing polypeptide, such as a deaminase polypeptide or an active variant or fragment thereof, which directly chemically modifies (e.g., deaminates) nucleobases, resulting in the conversion of one nucleobase to another. Deamination of nucleobases by deaminase can result in point mutations at the respective residues, referred to herein as "nucleic acid editing" or "base editing." Thus, a fusion protein comprising an RNA-guided nuclease (RGN) polypeptide and a deaminase can be used for targeted editing of nucleic acid sequences.

[0016] The fusion protein of the present disclosure can include an RGN fused to a prime editing polypeptide. Prime editing is a versatile and precise genome editing method that uses a nucleic acid programmable DNA binding protein that functions in association with a polymerase to directly write new genetic information into a specified DNA site (e.g., as described in US11,447,770B1, WO2021072328, WO2021226558, WO2020156575, WO2021042047, and US11193123, each of which is incorporated herein by reference in its entirety). The prime editing system uses the nickase RGN, and the system is programmed with a prime editing (PE) guide RNA ("PEgRNA").

[0017] The fusion proteins of the present disclosure are useful for targeted editing of DNA in vitro, e.g., for generating genetically modified cells. These genetically modified cells can be plant or animal cells. Such fusion proteins can also be useful for introducing targeted mutations, e.g., correcting genetic defects in mammalian cells ex vivo, e.g., in cells obtained from a subject that are then reintroduced into the same or another subject, and for introducing targeted mutations, e.g., correcting genetic defects in mammalian subjects or introducing inactivating mutations into disease-associated genes. Such fusion proteins can also be useful for introducing targeted mutations into plant cells, e.g., for introducing beneficial or agronomically valuable traits or alleles.

[0018] Any of the fusion proteins provided herein can be produced by any method known in the art. For example, the proteins provided herein can be produced through recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire contents of which are incorporated herein by reference.

[0019] II. Nucleic Acid Molecule-Binding Polypeptides Some aspects of the present disclosure provide fusion proteins comprising a nucleic acid molecule-binding polypeptide and a deaminase polypeptide. While binding to and targeted editing of RNA molecules is contemplated by the present invention, in some embodiments, the nucleic acid molecule-binding polypeptide of the fusion protein is a DNA-binding polypeptide. Such fusion proteins are useful for targeted editing of DNA in vitro, ex vivo, or in vivo. These fusion proteins are active in mammalian cells and are useful for targeted editing of DNA molecules.

[0020] In some embodiments, a fusion protein of the present disclosure comprises a DNA-binding polypeptide. As used herein, the term "DNA-binding polypeptide" refers to any polypeptide capable of binding to DNA. In certain embodiments, the DNA-binding polypeptide portion of a fusion protein of the present disclosure binds to double-stranded DNA. In some embodiments, the DNA-binding polypeptide binds to DNA in a sequence-specific manner. As used herein, the terms "sequence-specific" or "sequence-specific manner" refer to selective interaction with a specific nucleotide sequence.

[0021] Two polynucleotide sequences can be considered substantially complementary when the two sequences hybridize to each other under stringent conditions. Similarly, a DNA-binding polypeptide is considered to bind to a particular target sequence in a sequence-specific manner if the DNA-binding polypeptide binds to that sequence under stringent conditions. By "stringent conditions" or "stringent hybridization conditions" is intended conditions under which two polynucleotide sequences hybridize to each other (or a polypeptide binds to its specific target sequence) to a detectably higher degree than other sequences (e.g., at least 2-fold over background). Stringent conditions are sequence-dependent and will be different in different circumstances. Typically, stringent conditions would be conditions in which the salt concentration is less than 1.5 M Na ion, typically about 0.01 to 1.0 M Na ion (or other salt), at pH 7.0 to 8.3, and the temperature is at least about 30°C for short sequences (e.g., 10 to 50 nucleotides) and at least about 60°C for long sequences (e.g., greater than 50 nucleotides). Stringent conditions can also be achieved with the addition of destabilizing agents such as formamide. Exemplary low stringency conditions include hybridization in a buffer of 30-35% formamide, 1 M NaCl, and 1% SDS (sodium dodecyl sulfate) at 37°C, followed by a wash in 1x to 2x SSC (20x SSC = 3.0 M NaCl / 0.3 M trisodium citrate) at 50-55°C. Exemplary moderate stringency conditions include hybridization in 40-45% formamide, 1.0 M NaCl, and 1% SDS at 37°C, followed by a wash in 0.5x to 1x SSC at 55-60°C. Exemplary high stringency conditions include hybridization in 50% formamide, 1 M NaCl, and 1% SDS at 37°C, followed by a wash in 0.1x SSC at 60-65°C. Optionally, the wash buffer may contain about 0.1% to about 1% SDS. The duration of hybridization is generally less than about 24 hours, usually about 4 to about 12 hours. The duration of the wash period will be at least long enough to reach equilibrium.

[0022] Tm is the temperature (under defined ionic strength and pH) at which 50% of a complementary target sequence hybridizes to a perfectly matched sequence. For DNA-DNA hybrids, Tm can be estimated from the equation of Meinkoth and Wahl (1984) Anal. Biochem. 138:267-284: Tm = 81.5 ° C + 16.6 (log M) + 0.41 (% GC) - 0.61 (% form) - 500 / L, where M is the molar concentration of monovalent cations, % GC is the percentage of guanosine and cytosine nucleotides in DNA, % form is the percentage of formamide in the hybridization solution, and L is the hybrid length in base pairs. Generally, stringent conditions are selected to be about 5 ° C lower than the thermal melting point (Tm) of a specific sequence and its complementary strand at a defined ionic strength and pH. However, strictly stringent conditions can utilize hybridization and / or wash temperatures that are 1, 2, 3, or 4° C. below the thermal melting point (Tm). Moderately stringent conditions can utilize hybridization and / or wash temperatures that are 6, 7, 8, 9, or 10° C. below the thermal melting point (Tm). Low stringency conditions can utilize hybridization and / or wash temperatures that are 11, 12, 13, 14, 15, or 20° C. below the thermal melting point (Tm). Using the equations, hybridization and wash compositions, and desired Tm, one of skill in the art will understand that variations in the stringency of hybridization and / or wash solutions are essentially described.Extensive guides to nucleic acid hybridization can be found in Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, Part I, Chapter 2 (Elsevier, New York), and Ausubel et al., eds. (1995) Current Protocols in Molecular Biology, Chapter 2 (Greene Publishing and Wiley-Interscience, New York). See Sambrook et al. (1989) Molecular Cloning: A Laboratory Manual (2d ed., Cold Spring Harbor Laboratory Press, Plainview, New York).

[0023] In certain embodiments, the sequence-specific DNA-binding polypeptide is an RNA-guided DNA-binding polypeptide (RGDBP). As used herein, the terms "RNA-guided DNA-binding polypeptide" and "RGDBP" refer to a polypeptide capable of binding to DNA through hybridization of an associated RNA molecule with a target DNA sequence.

[0024] In some embodiments, the DNA-binding polypeptide of the fusion protein is a nuclease, such as a sequence-specific nuclease. As used herein, the term "nuclease" refers to an enzyme that catalyzes the cleavage of phosphodiester bonds between nucleotides in a nucleic acid molecule. In some embodiments, the DNA-binding polypeptide is an endonuclease capable of cleaving phosphodiester bonds between nucleotides in a nucleic acid molecule, while in certain embodiments, the DNA-binding polypeptide is an exonuclease capable of cleaving nucleotides at either end (5' or 3') of a nucleic acid molecule. In some embodiments, the sequence-specific nuclease is selected from the group consisting of meganucleases, zinc finger nucleases, TAL effector DNA-binding domain-nuclease fusion proteins (TALENs), and RNA-guided nucleases (RGNs), or variants thereof with reduced or inhibited nuclease activity.

[0025] As used herein, the term "meganuclease" or "homing endonuclease" refers to an endonuclease that binds to a recognition site within double-stranded DNA that is 12 to 40 bp in length. Non-limiting examples of meganucleases are those that belong to the LAGLIDADG family, which contains the conserved amino acid motif LAGLIDADG (SEQ ID NO: 700). The term "meganuclease" can refer to dimeric or single-chain meganucleases.

[0026] As used herein, the term "zinc finger nuclease" or "ZFN" refers to a chimeric protein comprising a zinc finger DNA-binding domain and a nuclease domain.

[0027] As used herein, the term "TAL effector DNA binding domain-nuclease fusion protein" or "TALEN" refers to a chimeric protein comprising a TAL effector DNA binding domain and a nuclease domain.

[0028] In certain embodiments, the DNA-binding polypeptide is capable of generating a single-stranded region within a double-stranded DNA molecule. An example of a single-stranded region is the single-stranded loop contained within an R-loop, which is a triple-stranded nucleic acid structure containing a region of single-stranded DNA formed within a double-stranded DNA molecule resulting from hybridization of a complementary strand to a single-stranded RNA or DNA molecule. The adenine within or adjacent to the single-stranded region of an R-loop can be deaminated by an adenine deaminase that has activity on single-stranded nucleic acids (e.g., ssDNA). In some of these embodiments, the DNA-binding polypeptide capable of generating an R-loop within a double-stranded DNA molecule is an RNA-guided DNA-binding polypeptide or an RGN nuclease. As used herein, the term "RNA-guided nuclease" or "RGN" refers to an RNA-guided DNA-binding polypeptide with nuclease activity. RGNs are considered to be "RNA-guided" because the guide RNA forms a complex with the RNA-guided nuclease to guide the RNA-guided nuclease to bind to the target sequence and, in some embodiments, introduce a single- or double-stranded break in the target sequence.

[0029] Although RGN may be capable of cleaving a target sequence upon binding, the term RGN also encompasses nuclease-inactive RGN, which is capable of binding to a target sequence but not cleaving it. The term "cleaving" or "cleavage" refers to the hydrolysis of at least one phosphodiester bond in the backbone of one or both strands of a double-stranded target sequence (e.g., a target DNA sequence), which can result in either a single-stranded break or a double-stranded break within the target DNA sequence. Cleavage of a target sequence by RGN can result in a single-stranded or double-stranded break. An RGN capable of cleaving only one strand of a double-stranded target nucleic acid molecule is referred to herein as a nickase. Such an RGN has a single functional nuclease domain. An RGN nickase may be a naturally occurring nickase, or an RGN protein that naturally cleaves both strands of a double-stranded nucleic acid molecule that has been mutated in one or more nuclease domains, thereby reducing or eliminating the nuclease activity of these mutated domains to become a nickase.

[0030] RNA-guided nucleases (RGNs) enable targeted manipulation of single sites within the genome and are useful in the context of gene targeting for therapeutic and research applications. In various organisms, including mammals, RNA-guided nucleases have been used for genome manipulation by stimulating either non-homologous end joining or homologous recombination. RGNs are CRISPR-Cas proteins, which are RNA-guided nucleases guided to target sequences by guide RNAs (gRNAs) as part of the clustered regularly interspaced short palindromic repeats (CRISPR) RNA-guided nuclease system, or active variants or fragments thereof.

[0031] Some aspects of the present disclosure provide fusion proteins comprising an RNA-guided DNA-binding polypeptide and a deaminase polypeptide, such as an adenine deaminase polypeptide. In some embodiments, the RNA-guided DNA-binding polypeptide is an RNA-guided nuclease (RGN). In further embodiments, the RNA-guided nuclease is a naturally occurring CRISPR-Cas protein or an active variant or fragment thereof. CRISPR-Cas systems are classified as class 1 or class 2 systems. Class 2 systems contain a single effector nuclease and include types II, V, and VI. Class 1 and 2 systems are subdivided into types (I, II, III, IV, V, VI), and some types are further divided into subtypes (e.g., type II-A, type II-B, type II-C, type VA, type VB).

[0032] In certain embodiments, the RGN is a naturally occurring type II CRISPR-Cas protein or an active variant or fragment thereof. As used herein, the terms "type II CRISPR-Cas protein," "type II CRISPR-Cas effector protein," or "type II RNA-guided nuclease" refer to an RGN that requires a trans-activating RNA (tracrRNA) and contains two nuclease domains (i.e., RuvC and HNH), each responsible for cleaving one strand of a double-stranded DNA molecule. In some embodiments, the present invention provides a fusion protein comprising a deaminase of the present disclosure fused to a Cas9 protein, such as Streptococcus pyogenes Cas9 (SpCas9), the sequence of which is set forth as SEQ ID NO: 415, or an SpCas9 nickase, as described in U.S. Patent Nos. 10,000,772 and 8,697,359, each of which is incorporated herein by reference in its entirety. In some embodiments, the present invention provides a fusion protein comprising a deaminase of the present disclosure fused to Streptococcus thermophilus Cas9 (StCas9), the sequence of which is set forth as SEQ ID NO:701, or a StCas9 nickase, as disclosed in U.S. Patent No. 10,113,167, the entire contents of which are incorporated herein by reference. In some embodiments, the present invention provides a fusion protein comprising a deaminase of the present disclosure fused to Streptococcus aureus Cas9 (SaCas9), the sequence of which is set forth as SEQ ID NO:702, or a SaCas9 nickase, as disclosed in U.S. Patent No. 9,752,132, the entire contents of which are incorporated herein by reference.

[0033] In some embodiments, the CRISPR-Cas protein is a naturally occurring type V CRISPR-Cas protein or an active variant or fragment thereof. As used herein, the terms "type V CRISPR-Cas protein," "type V CRISPR-Cas effector protein," or "type V RNA-guided nuclease" refer to an RGN that cleaves dsDNA, contains a single RuvC nuclease domain or split RuvC nuclease domains, and lacks an HNH domain (Zetsche et al. 2015, Cell doi:10.1016 / j.cell.2015.09.038; Shmakov et al. 2017, Nat Rev Microbiol doi:10.1038 / nrmicro.2016.184; Yan et al. 2018, Science doi:10.1126 / science.aav7271; Harrington et al. 2018, Science doi:10.1126 / science.aav4294). In some embodiments, the fusion proteins of the present disclosure comprise Cas12 (e.g., Cas12a). Note that Cas12a, also referred to as Cpf1, does not require tracrRNA, whereas other type V CRISPR-Cas proteins, such as Cas12b, do. Most type V effectors can also target ssDNA (single-stranded DNA), often without a PAM requirement (Zetsche et al., 2015; Yan et al., 2018; Harrington et al., 2018). The terms "type V CRISPR-Cas protein" and "type V RGN" encompass unique RGNs containing a split RuvC nuclease domain, such as those disclosed in WO2021 / 138247, the contents of each of which are incorporated herein by reference in their entireties.In some embodiments, the present invention provides fusion proteins comprising a deaminase of the present disclosure fused to either Francisella novicida Cas12a (FnCas12a), the sequence of which is set forth as SEQ ID NO:703 and which is disclosed in U.S. Patent No. 9,790,490, which is incorporated by reference herein in its entirety, or a nuclease-inactivated mutant of FnCas12a disclosed in U.S. Patent No. 9,790,490.

[0034] In some embodiments, the CRISPR-Cas protein is a naturally occurring type VI CRISPR-Cas protein or an active variant or fragment thereof. As used herein, the terms "type VI CRISPR-Cas protein," "type VI CRISPR-Cas effector protein," or "type VI RGN" refer to a CRISPR-Cas effector protein that contains two HEPN domains that cleave RNA without requiring tracrRNA. In some embodiments, the present invention provides fusion proteins comprising a deaminase of the present disclosure fused to Cas13.

[0035] In some embodiments, the fusion protein of the present disclosure comprises RGN, or a nickase or nuclease-inactive variant thereof, as disclosed in International Application Publication Nos. WO2019 / 236566, WO2020 / 139783, WO2021 / 030344, WO2021 / 138247, WO2021 / 231437, or WO2021 / 217002, or International Application No. PCT / IB2023 / 058160 filed August 12, 2023, each of which is incorporated by reference in its entirety.

[0036] In some embodiments, the fusion proteins of the present disclosure comprise an RGN, or a nickase- or nuclease-inactive variant thereof, listed in Table 1 and / or set forth as SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699. Guide RNA sequences (crRNA repeats and tracrRNA sequences) and consensus PAM sequences that can be used with each RGN in Table 1 are also provided. In certain embodiments, the fusion protein comprises a nucleotide sequence of about 80% or more, about 81% or more, about 82% or more, about 83% or more, about 84% or more, about 85% or more, about 86% or more, about 87% or more, about 88% or more, about 89% or more, about 90% or more, about 91% or more, about 92% or more, about 93% or more, about 94% or more, about 95% or more, about 96% or more, about 97% or more, about 98% or more, about 99% or more, about 100% or more, about 101% or more, about 102% or more, about 103% or more, about 104% or more, about 105% or more, about 106% or more, about 107% or more, about 108% or more, about 109% or more, about 110% or more, about 111% or more, about 112% or more, about 113% or more, about 114% or more, about 115% or more, about 116% or more, about 117% or more, about 118% or more, about 119% or more, about 120% or more, about 121% or more, about 122% or more, about 12 and / or set forth as SEQ ID NOS: 1-4, 49-162, 435, 575, 576, 698, and 699, having 80% to 99% or more sequence identity with any one of the amino acid sequences listed in Table 1 and / or set forth as SEQ ID NOS: 1-4, 49-162, 435, 575, 576, 698, and 699, including, but not limited to, about 96% or more, about 97% or more, about 98% or more, about 99% or more, or more.In some embodiments, the fusion protein comprises an RGN having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or more sequence identity to the RGN amino acid sequence disclosed in Table 1 and / or set forth as SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699. In other embodiments, the fusion protein comprises a fragment of an RGN listed in Table 1, such as one that differs by only 1-10 amino acid residues, such as only 1-15 amino acid residues, 6-10, 5, 4, 3, 2, or 1. In specific embodiments, RGN comprises an N- or C-terminal truncation, which may include a deletion of 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, or more amino acids from either the N- or C-terminus of the polypeptide. In some embodiments, RGN comprises an internal deletion, which may include a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, or more amino acids. [Table 1-1] [Table 1-2] [Table 1-3]

[0037] The RGN component of the fusion protein of the present disclosure is selected from the group consisting of APG07433.1 (disclosed in International Patent Publication No. WO2019 / 236566, which is incorporated herein by reference in its entirety, and set forth herein as SEQ ID NO: 1), APG05586 (disclosed in International Patent Publication No. WO2021 / 217002, which is incorporated herein by reference in its entirety, and set forth herein as SEQ ID NO: 2), APG01604 (disclosed in International Patent Publication No. WO2021 / 217002, which is incorporated herein by reference in its entirety, and set forth herein as SEQ ID NO: 3), and APG01604 (disclosed in International Patent Publication No. WO2021 / 217002, which is incorporated herein by reference in its entirety, and set forth herein as SEQ ID NO: 4). The nickase may be an active fragment or variant (i.e., capable of binding to a nucleic acid molecule in an RNA-guided manner), such as LPG10145 (disclosed in International Patent Publication No. WO2023 / 139557, which is incorporated by reference in its entirety, and set forth herein as SEQ ID NO: 4), LPG10196 (disclosed in International Application No. PCT / IB2023 / 058160, filed August 12, 2023, which is incorporated by reference in its entirety, and set forth herein as SEQ ID NO: 131), or a nickase variant thereof.

[0038] RGN can have 80% to 99% or more sequence identity to any one of SEQ ID NOs: 1, 2, 3, 4, and 131 (which retain RGN activity), including but not limited to at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more. In certain embodiments, a fusion protein comprises RGN having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 1, 2, 3, 4, and 131 (which retains RGN activity). The fusion protein can comprise an active fragment of RGN set forth as any one of SEQ ID NOs: 1, 2, 3, 4, and 131, such as one that differs from any one of SEQ ID NOs: 1, 2, 3, 4, and 131 by no more than 1 to 10 amino acid residues, such as no more than 1 to 15 amino acid residues, no more than 6 to 10 amino acid residues, no more than 5 amino acid residues, no more than 4 amino acid residues, no more than 3 amino acid residues, no more than 2 amino acid residues, or no more than 1 amino acid residue. In certain embodiments, an active RGN fragment comprises an N- or C-terminal truncation that may include a deletion of 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, or more amino acids from either the N- or C-terminus of the polypeptide set forth in any one of SEQ ID NOs: 1, 2, 3, 4, and 131. In some embodiments, the active RGN variant or fragment comprises an internal deletion, which can include deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, or more amino acids of any one of SEQ ID NOs: 1, 2, 3, 4, and 131.An active fragment of RGN can comprise at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, or more consecutive amino acid residues of the amino acid sequence set forth in any one of SEQ ID NOs: 1, 2, 3, 4, and 131.

[0039] In some embodiments, the RGN of the fusion protein is an RGN nickase containing a mutation (e.g., a D10A mutation, based on the Streptococcus pyogenes Cas9 sequence, whose amino acid numbering is set forth as SEQ ID NO: 415) that renders RGN capable of cleaving only the non-base-edited target strand of a nucleic acid duplex (the strand that contains the PAM and is base-paired to the gRNA). Nickases containing the D10A mutation or equivalent mutations have an inactivated RuvC nuclease domain and cleave the target strand. D10A nickases cannot cleave the non-target strand of DNA, i.e., the strand where base editing is desired. In these embodiments, RGN nicks the target strand, while the complementary non-target strand is modified by a deaminase. Cellular DNA repair machinery can use the modified non-target strand as a template to repair the nicked target strand, thereby introducing a mutation into the DNA.

[0040] Thus, in some embodiments, the nickase comprises an inactive RuvC domain. The RuvC domain has an RNase H fold (see, e.g., Nishimasu et al. (2014) Cell 156(5):935-949, incorporated by reference in its entirety). The RuvC domain of RGN is often a split RuvC domain that includes two or more non-adjacent regions within the linear amino acid sequence. For example, the RuvC domain of Streptococcus pyogenes Cas9 includes amino acid residues 1-59, 718-769, and 909-1098 of SEQ ID NO: 415. A non-limiting example of a mutation in the RuvC domain that inactivates its nuclease activity is the D10A mutation, which mutates the first aspartic acid residue within the split RuvC nuclease domain. nAPG07433.1 (represented as SEQ ID NO:49) is a nickase variant (with an inactivated RuvC domain) of APG07433.1, represented as SEQ ID NO:1 and described in WO2019 / 236566 (incorporated herein by reference in its entirety). nAPG05586 (represented as SEQ ID NO:50) is a nickase variant (with an inactivated RuvC domain) of APG05586 (represented as SEQ ID NO:2). nAPG01604 (represented as SEQ ID NO:51) is a nickase variant (with an inactivated RuvC domain) of APG01604 (represented as SEQ ID NO:3). nLPG10145 (represented as SEQ ID NO:52) is a nickase variant (with an inactivated RuvC domain) of LPG10145 (represented as SEQ ID NO:4). nLPG10196 (shown as SEQ ID NO: 698) is a nickase variant (with an inactivated RuvC domain) of LPG10196 (shown as SEQ ID NO: 131).

[0041] In some embodiments, the RGN of the fusion protein is an RGN nickase containing a mutation (e.g., an H840A mutation, based on the Streptococcus pyogenes Cas9 sequence, whose amino acid numbering is set forth as SEQ ID NO: 415) that renders RGN capable of cleaving only the non-target strand of a nucleic acid duplex (the strand that does not contain a PAM and is not base-paired to the gRNA). In some of these embodiments, the nickase contains an inactive HNH nuclease domain. The HNH nuclease domain of RGN has a ββα-metallofold (see, e.g., Nishimasu et al. 2014). For example, the HNH nuclease domain of Streptococcus pyogenes Cas9 contains amino acid residues 775-908 of SEQ ID NO: 415. A non-limiting example of a mutation in the HNH domain that inactivates its nuclease activity is the H840A mutation, which mutates the first histidine of the HNH nuclease domain. RGN with an inactivated HNH domain acts on the non-target strand. nAPG07433.1 (shown as SEQ ID NO: 56) is a nickase variant (with an inactivated HNH domain) of APG07433.1 shown as SEQ ID NO: 1 and described in WO2019 / 236566 (incorporated herein by reference in its entirety). nAPG05586 (shown as SEQ ID NO: 53) is a nickase variant (with an inactivated HNH domain) of APG05586 (shown as SEQ ID NO: 2). nAPG01604 (shown as SEQ ID NO: 54) is a nickase variant (with an inactivated HNH domain) of APG01604 (shown as SEQ ID NO: 3). nLPG10145 (shown as SEQ ID NO: 55) is a nickase variant (with an inactivated HNH domain) of LPG10145 (shown as SEQ ID NO: 4). nLPG10196 (shown as SEQ ID NO: 699) is a nickase variant (with an inactivated HNH domain) of LPG10196 (shown as SEQ ID NO: 131).

[0042] Methods for inactivating the RuvC and / or HNH domains of RGN are known in the art and generally involve mutating the first aspartic acid in the split RuvC domain and / or the first histidine in the HNH domain. Typically, the aspartic acid or histidine residue is mutated to alanine. Other amino acid residues in the RuvC domain that can be mutated to inactivate the nuclease activity of the domain include Glu762, His983, and Asp986 (typically to alanine), with amino acid numbering based on the Streptococcus pyogenes Cas9 sequence set forth in SEQ ID NO:415. Other amino acid residues in the HNH domain that can be mutated include D839 and N863 (typically to alanine), with amino acid numbering based on the Streptococcus pyogenes Cas9 sequence set forth in SEQ ID NO:415.

[0043] Unless otherwise specified, the use of the nomenclature "n" followed by the name of a nuclease refers to a nickase in which the RuvC domain is inactivated (e.g., has a D10A mutation).

[0044] In some embodiments, the fusion protein comprises an RGN nickase that retains nickase activity, comprising an amino acid sequence having about 60% to about 99.5% identity to any one of SEQ ID NOs: 49-56, 698, and 699. In some embodiments, the RGN nickase that retains nickase activity comprises an amino acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to any one of SEQ ID NOs: 49-56, 698, and 699.

[0045] In some embodiments, the fusion protein comprises an RGN nickase comprising an amino acid sequence having 80% to 99% or more sequence identity to any one of SEQ ID NOs: 49-56, 698, and 699, including but not limited to at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or more.

[0046] In some embodiments, the RGN of the fusion protein is nuclease-inactive. As used herein, an RGN protein mutated to be nuclease-inactive or "dead" may be referred to as an RNA-guided DNA-binding polypeptide, or a nuclease-inactive RGN or a nuclease-inactive RGN. Methods for generating nuclease-inactive RGN are known in the art and generally involve mutating only one or all of the nuclease domains of RGN to render the nuclease domain(s) inactive. In those embodiments in which RGN contains only a single nuclease domain (e.g., a RuvC domain), the nuclease-inactive variant has at least one mutation in the RuvC domain that results in the inactivation of the RuvC nuclease domain. In those embodiments in which RGN contains more than one nuclease domain, such as a RuvC and an HNH domain, at least one mutation in each of the RuvC and HNH domains renders both nuclease domains inactive.

[0047] One exemplary suitable nuclease-inactive RGN is the D10A / H840A Cas9 mutant (see, e.g., Qi et al., Cell. 2013;152(5):1173-83, the entire contents of which are incorporated herein by reference). Additionally, suitable nuclease-inactive variants of other known RNA-guided nucleases (RGNs) can be determined.

[0048] Other additional exemplary suitable nuclease-inactive RGN variants include, but are not limited to, the D10A / D839A / H840A and D10A / D839A / H840A / N863A mutant domains (see, e.g., Mali et al., Nature Biotechnology. 2013;31(9):833-838, the entire contents of which are incorporated herein by reference).

[0049] Additional suitable RGN proteins mutated into nickases or inactive nucleases will be apparent to those of skill in the art based on this disclosure and knowledge within the art (e.g., RGNs disclosed in PCT Publication No. WO2019 / 236566, which is incorporated herein by reference in its entirety), and are within the scope of this disclosure.

[0050] Any method known in the art for introducing mutations into an amino acid sequence, such as PCR-mediated mutagenesis and site-directed mutagenesis, can be used to generate nickase- or nuclease-inactive RGN. See, e.g., U.S. Publication No. 2014 / 0068797 and U.S. Patent No. 9,790,490, each of which is incorporated herein by reference in its entirety.

[0051] Fusion proteins of the invention containing RGN utilize a guide RNA that binds to the RGN component of the fusion protein and guides the fusion protein to a target sequence. The term "guide RNA" refers to a nucleotide sequence that has sufficient complementarity with a target nucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of an associated RGN to the target nucleotide sequence. More specifically, when the target nucleotide sequence is double-stranded, such as in the case of DNA, the target nucleotide sequence consists of a target strand (including a PAM sequence) and a non-target strand. In these embodiments, the guide RNA has sufficient complementarity with the non-target strand of the double-stranded target sequence (e.g., a target DNA sequence) to hybridize with the non-target strand and direct sequence-specific binding of an associated RNA-guided nuclease (RGN) to the target sequence (e.g., a target DNA sequence). Thus, in some embodiments, the guide RNA includes a spacer that is identical in sequence to the target strand, except that uracil (U) replaces thymidine (T) in the guide RNA.

[0052] A guide RNA is one or more RNA molecules (typically one or two) that bind to RGN, guide RGN to bind to a specific target nucleotide sequence, and, in cases where the RGN has nickase or nuclease activity, can also cleave the target nucleotide sequence. A guide RNA comprises a CRISPR RNA (crRNA), and in some embodiments, a trans-activating CRISPR RNA (tracrRNA). In some embodiments, a portion of the guide RNA comprises DNA nucleotides. In certain embodiments, the guide RNA comprises artificial, non-naturally occurring nucleotide analogs, or one or more nucleotides are chemically modified and include modifications described in International Application No. PCT / IB2023 / 058418, filed August 25, 2023, which is incorporated by reference in its entirety.

[0053] CRISPR RNA comprises a spacer sequence and a CRISPR repeat sequence. A "spacer sequence" is a nucleotide sequence that directly hybridizes with a non-target strand of a target sequence of interest. The spacer sequence is engineered to be fully or partially complementary to the non-target strand of a target sequence of interest. In various embodiments, the spacer sequence comprises from about 8 nucleotides to about 30 nucleotides, or more. For example, the spacer sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides in length. In some embodiments, the spacer sequence is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides in length. In some embodiments, the spacer sequence is about 10 to about 26 nucleotides in length, or about 12 to about 30 nucleotides in length. In some embodiments, the spacer sequence is about 30 nucleotides in length. In some embodiments, the spacer sequence is 30 nucleotides in length.In some embodiments, the degree of complementarity between a spacer sequence and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is 50% to 99% or more, including, but not limited to, about 50% or more, about 60% or more, about 70% or more, about 75% or more, about 80% or more, about 81% or more, about 82% or more, about 83% or more, about 84% or more, about 85% or more, about 86% or more, about 87% or more, about 88% or more, about 89% or more, about 90% or more, about 91% or more, about 92% or more, about 93% or more, about 94% or more, about 95% or more, about 96% or more, about 97% or more, about 98% or more, about 99% or more, or more. In some embodiments, the degree of complementarity between a spacer sequence and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more. In some embodiments, the spacer sequence does not contain secondary structure, which can be predicted using any suitable polynucleotide folding algorithm known in the art, including, but not limited to, mFold (see, e.g., Zuker and Stiegler (1981) Nucleic Acids Res. 9:133-148) and RNAfold (see, e.g., Gruber et al. (2008) Cell 106(1):23-24).

[0054] CRISPR RNA repeats comprise a nucleotide sequence that, by itself or in concert with a hybridized tracrRNA, forms a structure recognized by an RGN molecule. In various embodiments, the CRISPR RNA repeat comprises from about 8 nucleotides to about 30 nucleotides, or more. In some embodiments, the CRISPR RNA repeat comprises from about 8 nucleotides to about 30 nucleotides, or more. For example, the CRISPR repeat can be about 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more nucleotides in length. In some embodiments, the CRISPR repeat sequences are 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides in length. In some embodiments, the degree of complementarity between a CRISPR repeat sequence and its corresponding tracrRNA sequence, when optimally aligned using a suitable alignment algorithm, is 50%-99% or more, including, but not limited to, about 50% or more, about 60% or more, about 70% or more, about 75% or more, about 80% or more, about 81% or more, about 82% or more, about 83% or more, about 84% or more, about 85% or more, about 86% or more, about 87% or more, about 88% or more, about 89% or more, about 90% or more, about 91% or more, about 92% or more, about 93% or more, about 94% or more, about 95% or more, about 96% or more, about 97% or more, about 98% or more, about 99% or more, or more.In some embodiments, the degree of complementarity between a CRISPR repeat sequence and its corresponding tracrRNA sequence, when optimally aligned using a suitable alignment algorithm, is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more.

[0055] In some embodiments, the guide RNA further comprises a tracrRNA molecule. The transactivating CRISPR RNA or tracrRNA molecule comprises a nucleotide sequence comprising a region of sufficient complementarity to hybridize to the CRISPR repeats of the crRNA, referred to herein as the anti-repeat region. In some embodiments, the tracrRNA molecule further comprises a region with a secondary structure (e.g., a stem-loop) or forms a secondary structure upon hybridization with its corresponding crRNA. In some embodiments, the region of the tracrRNA that is fully or partially complementary to the CRISPR repeats is at the 5' end of the molecule, and the 3' end of the tracrRNA comprises a secondary structure. This region of secondary structure generally comprises several hairpin structures, including a nexus hairpin found adjacent to the anti-repeat sequence. There is often a terminal hairpin at the 3' end of the tracrRNA, which can vary in structure and number, but often comprises a GC-rich Rho-independent transcription terminator hairpin followed by a string of Us at the 3' end. See, e.g., Briner et al. (2014) Molecular Cell 56:333-339, Briner and Barrangou (2016) Cold Spring Harb Protoc, doi:10.1101 / pdb.top090902, and U.S. Publication No. 2017 / 0275648, each of which is incorporated herein by reference in its entirety.

[0056] In various embodiments, the anti-repeat region of the tracrRNA that is fully or partially complementary to the CRISPR repeat sequence comprises from about 6 nucleotides to about 30 nucleotides, or more. For example, the region of base pairing between the tracrRNA anti-repeat sequence and the CRISPR repeat sequence can be about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides in length. In some embodiments, the region of base pairing between the tracrRNA anti-repeat sequence and the CRISPR repeat sequence is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more nucleotides in length. In some embodiments, the anti-repeat region of the tracrRNA that is fully or partially complementary to the CRISPR repeat sequence is about 10 nucleotides in length. In some embodiments, the anti-repeat region of the tracrRNA that is fully or partially complementary to the CRISPR repeat sequence is 10 nucleotides in length.In some embodiments, the degree of complementarity between a CRISPR repeat sequence and its corresponding tracrRNA anti-repeat sequence, when optimally aligned using a suitable alignment algorithm, is about 50% or more, about 60% or more, about 70% or more, about 75% or more, about 80% or more, about 81% or more, about 82% or more, about 83% or more, about 84% or more, about 85% or more, about 86% or more. or greater, about 87% or greater, about 88% or greater, about 89% or greater, about 90% or greater, about 91% or greater, about 92% or greater, about 93% or greater, about 94% or greater, about 95% or greater, about 96% or greater, about 97% or greater, about 98% or greater, about 99% or greater, or greater, 50% to 99% or greater. In some embodiments, the degree of complementarity between a CRISPR repeat sequence and its corresponding tracrRNA anti-repeat sequence, when optimally aligned using a suitable alignment algorithm, is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or more.

[0057] In various embodiments, the entire tracrRNA comprises from about 60 nucleotides to more than about 210 nucleotides. In some embodiments, the entire tracrRNA comprises from about 60 nucleotides to more than 210 nucleotides. For example, the tracrRNA can be about 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 150, 160, 170, 180, 190, 200, 210, or more nucleotides in length. In some embodiments, the tracrRNA is 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 150, 160, 170, 180, 190, 200, 210, or more nucleotides in length. In some embodiments, the tracrRNA is about 100 to about 200 nucleotides in length, including about 95, 96, 97, 98, 99, 100, 105, 106, 107, 108, 109, and 100 nucleotides in length. In some embodiments, the tracrRNA is 100-110 nucleotides in length, including 95, 96, 97, 98, 99, 100, 105, 106, 107, 108, 109, and 110 nucleotides in length.

[0058] The guide RNA forms a complex with an RNA-guided DNA-binding polypeptide or an RNA-guided nuclease, or a fusion protein containing the same, and guides the RNA-guided nuclease or the fusion protein containing the same to bind to a target sequence. When the guide RNA forms a complex with an RGN, the bound RGN introduces a single- or double-strand break in the target sequence. After the target sequence is cleaved, the break can be repaired so that the DNA sequence of the target sequence is modified during the repair process. Provided herein are methods for using mutant variants of RNA-guided nucleases, either nuclease-inactive or nickase-active, linked to a base-editing polypeptide (e.g., a deaminase) or a primed editing polypeptide to modify a target sequence in the DNA of a host cell. Because the polypeptide is capable of binding to a target sequence but not necessarily cleaving it, mutant variants of RNA-guided nucleases in which nuclease activity is inactivated or significantly reduced may be referred to as RNA-guided DNA-binding polypeptides (RGDBPs). RNA-guided nucleases that are capable of cleaving only one strand of a double-stranded nucleic acid molecule are referred to herein as nickases.

[0059] The target nucleotide sequence is bound by an RNA-guided DNA-binding polypeptide (e.g., RGN) and hybridizes with a guide RNA (e.g., RGN) that associates with RGDBP. If RGDBP possesses nuclease activity, including nickase activity (i.e., is RGN), it can subsequently cleave the target sequence.

[0060] The guide RNA can be a single-guide RNA or a dual-guide RNA system. A single-guide RNA comprises a crRNA and optionally a tracrRNA on a single molecule of RNA, while a dual-guide RNA system comprises a crRNA and a tracrRNA present on two distinct RNA molecules hybridized to each other through at least a portion of the CRISPR repeat sequence of the crRNA and at least a portion of the tracrRNA, which may be fully or partially complementary to the CRISPR repeat sequence of the crRNA. In some of these embodiments in which the guide RNA is a single-guide RNA, the crRNA and optionally the tracrRNA are separated by a linker nucleotide sequence.

[0061] Suitable crRNA repeat, tracrRNA, and guide RNA sequences for the RGN component of the fusion proteins of the present disclosure are disclosed in International Patent Publication Nos. WO2019 / 236566, WO2021 / 030344, WO2020 / 139783, WO2021 / 217002, WO2021 / 138247, WO2021 / 231437, WO2023 / 139557, and PCT International Application No. PCT / IB2023 / 058160 filed August 12, 2023, each of which is incorporated by reference herein in its entirety, and / or are provided in Table 1 herein.

[0062] Generally, the linker nucleotide sequence between the crRNA and tracrRNA does not contain complementary bases to avoid the formation of secondary structures within or involving the nucleotides of the linker nucleotide sequence. In some embodiments, the linker nucleotide sequence between the crRNA and tracrRNA is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, or more nucleotides in length. In some embodiments, the linker nucleotide sequence between the crRNA and tracrRNA is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more nucleotides in length. In some embodiments, the linker nucleotide sequence of a single guide RNA is at least 4 nucleotides in length. In some embodiments, the linker nucleotide sequence of a single guide RNA is 4 nucleotides in length.

[0063] In certain embodiments, the guide RNA can be introduced into a target cell, organelle, or embryo as an RNA molecule. The guide RNA can be in vitro transcribed or chemically synthesized. In some embodiments, a nucleotide sequence encoding the guide RNA is introduced into a cell, organelle, or embryo. In some embodiments, the nucleotide sequence encoding the guide RNA is operably linked to a promoter (e.g., an RNA polymerase III promoter). The promoter can be a native promoter or heterologous to the guide RNA-encoding nucleotide sequence. In some embodiments, the promoter is selected from any one of the promoters disclosed in International Application No. PCT / US2022 / 032940, filed June 10, 2022, which is incorporated herein by reference in its entirety.

[0064] In various embodiments, the guide RNA can be introduced into a target cell, organelle, or embryo as a ribonucleoprotein complex, as described herein, wherein the guide RNA binds to an RNA-guided nuclease polypeptide.

[0065] The guide RNA guides the associated RGDBP (e.g., RGN) to a specific target nucleotide sequence of interest through hybridization of the guide RNA to the target nucleotide sequence. The target nucleotide sequence can comprise DNA, RNA, or a combination of both, and can be single-stranded or double-stranded. The target nucleotide sequence can be genomic DNA (i.e., chromosomal DNA), plasmid DNA, or an RNA molecule (e.g., messenger RNA, ribosomal RNA, transfer RNA, microRNA, small interfering RNA). The target nucleotide sequence can be bound (and in some embodiments, cleaved) by an RNA-guided DNA-binding polypeptide in vitro or in a cell. The chromosomal sequence targeted by the RGDBP (e.g., RGN) can be a nuclear, plastid, or mitochondrial chromosomal sequence. In some embodiments, the target nucleotide sequence is unique in the target genome.

[0066] In some embodiments, the target nucleotide sequence is adjacent to a protospacer adjacent motif (PAM). The PAM is generally within about 1 to about 10 nucleotides of the target nucleotide sequence, including about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 nucleotides from the target nucleotide sequence. In some embodiments, the PAM is within 1 to 10 nucleotides of the target nucleotide sequence, including 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides from the target nucleotide sequence. Unless otherwise specified, the PAM is immediately adjacent to the target nucleotide sequence at either its 5' or 3' end. In some embodiments, the PAM is 3' of the target sequence. Generally, the PAM is a consensus sequence of about 2-6 nucleotides, but in some embodiments is 1, 2, 3, 4, 5, 6, 7, 8, 9, or more nucleotides in length.

[0067] Suitable PAM sequences for the RGN component of the fusion proteins of the present disclosure are disclosed in International Patent Publication Nos. WO2019 / 236566, WO2021 / 030344, WO2020 / 139783, WO2021 / 217002, WO2021 / 138247, WO2021 / 231437, WO2023 / 139557, and PCT International Application No. PCT / IB2023 / 058160 filed August 12, 2023, each of which is incorporated by reference in its entirety, and / or are provided in Table 1 herein.

[0068] A PAM limits the sequences that a given RGDBP (e.g., RGN) can target because the PAM must be proximal to the target nucleotide sequence. Upon recognizing its corresponding PAM sequence, RGN can cleave the target nucleotide sequence at a specific cleavage site. As used herein, a cleavage site consists of two specific nucleotides within the target nucleotide sequence at which the nucleotide sequence is cleaved by RGN. The cleavage site can include the first and second, second and third, third and fourth, fourth and fifth, fifth and sixth, seventh and eighth, or eighth or ninth nucleotides from the PAM in either the 5' or 3' direction. Because RGN can cleave the target nucleotide sequence to produce crossed ends, in some embodiments, the cleavage site is defined based on a distance of two nucleotides from the PAM on the positive (+) strand of the polynucleotide and a distance of two nucleotides from the PAM on the negative (-) strand of the polynucleotide.

[0069] RGDBP and RGN can be used to deliver fusion polypeptides, polynucleotides, or small molecule payloads to specific genomic locations.

[0070] In embodiments in which the fusion protein contains a meganuclease as the DNA-binding component, the target sequence can contain a pair of inverted 9-base pair "half-sites" separated by four base pairs. In the case of a single-chain meganuclease, the N-terminal domain of the protein contacts the first half-site, and the C-terminal domain of the protein contacts the second half-site. Cleavage by the meganuclease results in a four-base pair 3' overhang. In embodiments in which the DNA-binding polypeptide contains a compact TALEN, the recognition sequence contains a first CNNNGN sequence recognized by the I-TevI ​​domain, followed by a nonspecific spacer 4-16 base pairs long, followed by a second sequence 16-22 bp long (this sequence typically has a 5' T base) recognized by the TAL effector domain. In those embodiments in which the DNA-binding polypeptide component of the fusion protein comprises a zinc finger, the DNA-binding domain typically recognizes an 18 bp recognition sequence comprising a pair of 9 base pair "half-sites" separated by 2-10 base pairs, and cleavage by the nuclease generates blunt ends or 5' overhangs of variable length (frequently 4 base pairs).

[0071] III. Deaminases Some aspects of the present invention provide adenine deaminases produced through directed evolution and optimization of the previously disclosed adenine deaminase LPG50148 (represented herein as SEQ ID NO:5 without its start methionine), as disclosed in International Application No. WO 2022 / 056254, which is incorporated herein by reference in its entirety. These evolved adenine deaminases are represented as SEQ ID NOs:6, 300-414, 596-598, and 720-723 (again without the start methionine).

[0072] The term "deaminase" refers to an enzyme that catalyzes a deamination reaction. The deaminase of the present invention is a nucleobase deaminase, and the terms "deaminase" and "nucleobase deaminase" are used interchangeably herein. The deaminase can be a naturally occurring deaminase enzyme, or an active fragment or variant thereof. The deaminase can be active on single-stranded nucleic acids, such as ssDNA or ssRNA, or on double-stranded nucleic acids, such as dsDNA or dsRNA. In some embodiments, the deaminase can only deaminate ssDNA and does not act on dsDNA.

[0073] The deaminases of the present invention can be used to edit DNA or RNA molecules and are useful as deaminases alone or as components of fusion proteins. In some embodiments, the deaminases can be used to edit ssDNA or ssRNA molecules. Deamination of adenine, adenosine, or deoxyadenosine generates inosine, which is treated as guanine by polymerases.

[0074] To date, there are no known naturally occurring adenine deaminases that deaminate adenine in DNA. Several methods have been employed to evolve and optimize adenine deaminases that act on tRNA (ADAT) proteins so that they are active on DNA molecules in mammalian cells (Gaudelli et al., 2017; Koblan, L. Wet et al., 2018, Nat Biotechnol 36, 843-846; Richter, M. F. et al., 2020, Nat Biotechnol, doi:10.1038 / s41587-020-0562-8, each of which is incorporated herein by reference in its entirety).

[0075] In some embodiments, the adenine deaminases of the present disclosure are used in combination with a cytosine deaminase that catalyzes the hydrolytic deamination of cytosine, cytidine, or deoxycytidine to uracil (such as those disclosed in International Application Publication Nos. WO2020 / 189783 and WO2022 / 204093, and U.S. Application Publication No. 2022 / 0145296, each of which is incorporated herein by reference).

[0076] Deaminases of the present disclosure, or active variants or fragments thereof, can be introduced into cells as part of a deaminase-DNA binding polypeptide fusion and / or can be co-expressed with the DNA binding polypeptide-deaminase fusion to increase the efficiency of introducing desired A>N (where N is C, T, or G) mutations, such as A>G mutations, into target DNA molecules.

[0077] Deaminases of the present disclosure comprise an amino acid sequence having about 50% to about 100% identity to any of SEQ ID NOs: 5, 6, 300-414, 596-598, and 720-723. In some embodiments, the deaminase has about 50% to about 100% identity to any of SEQ ID NOs: 5, 6, 300-414, 596-598, and 720-723 and at least one of the amino acid residues set forth in any one of Tables 2, 4, 6, and 19. In some embodiments, the deaminase has an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to any of SEQ ID NOs: 5, 6, 300-414, 596-598, and 720-723, and comprises at least one of the amino acid residues set forth in any one of Tables 2, 4, 6, and 19. In some embodiments, the deaminase comprises an amino acid sequence selected from SEQ ID NOs: 303, 319, 366, 375, 406, 596, 597, and 720. Non-limiting examples of such fusion proteins are described in the Examples section herein.

[0078] The deaminase can have at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or more sequence identity to SEQ ID NO:5, and the deaminase can comprise the following amino acid residues: a) C at the position corresponding to position 2 of SEQ ID NO: 5; b) F or C at the position corresponding to position 22 of SEQ ID NO: 5; c) Q at a position corresponding to position 23 of SEQ ID NO: 5; d) N at a position corresponding to position 35 of SEQ ID NO: 5; e) L at a position corresponding to position 40 of SEQ ID NO: 5; f) Q at a position corresponding to position 46 of SEQ ID NO: 5; g) A or M at the position corresponding to position 68 of SEQ ID NO: 5; h) H at a position corresponding to position 72 of SEQ ID NO: 5; i) W, A, Q, Y, or D at a position corresponding to position 75 of SEQ ID NO: 5; j) G at a position corresponding to position 76 of SEQ ID NO: 5; k) S at a position corresponding to position 81 of SEQ ID NO: 5; l) M at a position corresponding to position 105 of SEQ ID NO: 5; m) H at a position corresponding to position 108 of SEQ ID NO: 5; n) L at a position corresponding to position 109 of SEQ ID NO: 5; o) I or L at a position corresponding to position 117 of SEQ ID NO: 5; p) F at a position corresponding to position 120 of SEQ ID NO: 5; q) E, T, A, G, or V at a position corresponding to position 121 of SEQ ID NO: 5; r) I or H at a position corresponding to position 122 of SEQ ID NO: 5; s) K at a position corresponding to position 125 of SEQ ID NO: 5; t) A at a position corresponding to position 126 of SEQ ID NO: 5; u) H at a position corresponding to position 135 of SEQ ID NO: 5; v) V at a position corresponding to position 137 of SEQ ID NO: 5; w) Y at a position corresponding to position 138 of SEQ ID NO: 5; x) L at a position corresponding to position 139 of SEQ ID NO: 5; y) K or A at a position corresponding to position 142 of SEQ ID NO: 5; z) A, Q, L, or M at a position corresponding to position 145 of SEQ ID NO: 5; aa) K at a position corresponding to position 148 of SEQ ID NO: 5, bb) Q at a position corresponding to position 151 of SEQ ID NO: 5; cc) E or R at a position corresponding to position 153 of SEQ ID NO: 5; dd) W at a position corresponding to position 155 of SEQ ID NO: 5; ee) R or V at a position corresponding to position 156 of SEQ ID NO: 5; ff) F at a position corresponding to position 157 of SEQ ID NO: 5; gg) R at a position corresponding to position 158 of SEQ ID NO: 5; hh) Q at a position corresponding to position 159 of SEQ ID NO: 5; ii) D or W at a position corresponding to position 160 of SEQ ID NO: 5; jj) W, H, V, or A at a position corresponding to position 162 of SEQ ID NO: 5; kk) R at a position corresponding to position 165 of SEQ ID NO: 5, and ll) has at least one H at a position corresponding to position 166 of SEQ ID NO:5.

[0079] In some embodiments, the deaminase is a) N at a position corresponding to position 35 of SEQ ID NO: 5 and W at a position corresponding to position 162 of SEQ ID NO: 5; b) N at a position corresponding to position 35 of SEQ ID NO: 5 and Q at a position corresponding to position 46 of SEQ ID NO: 5; c) S at a position corresponding to position 81 of SEQ ID NO: 5 and R at a position corresponding to position 156 of SEQ ID NO: 5; d) Q at a position corresponding to position 46 of SEQ ID NO: 5, and R at a position corresponding to position 156 of SEQ ID NO: 5; e) R at a position corresponding to position 156 of SEQ ID NO: 5 and W at a position corresponding to position 162 of SEQ ID NO: 5; f) A at a position corresponding to position 68 of SEQ ID NO: 5, and S at a position corresponding to position 81 of SEQ ID NO: 5; g) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at a position corresponding to position 35 of SEQ ID NO:5, Q at a position corresponding to position 46 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; i) W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; j) N at a position corresponding to position 35 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; k) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; l) M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; m) Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; n) M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; o) N at a position corresponding to position 35, Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; p) S at a position corresponding to position 81 of SEQ ID NO: 5 and W at a position corresponding to position 155 of SEQ ID NO: 5; q) W at a position corresponding to position 75 of SEQ ID NO: 5 and S at a position corresponding to position 81 of SEQ ID NO: 5; r) S at a position corresponding to position 81 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; s) S at a position corresponding to position 81 of SEQ ID NO: 5 and A at a position corresponding to position 145 of SEQ ID NO: 5; t) W at a position corresponding to position 75 of SEQ ID NO: 5, and W at a position corresponding to position 155 of SEQ ID NO: 5; u) W at a position corresponding to position 155 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; v) A at a position corresponding to position 145 of SEQ ID NO:5 and W at a position corresponding to position 155 of SEQ ID NO:5; w) W at a position corresponding to position 75 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; x) W at a position corresponding to position 75 of SEQ ID NO: 5 and A at a position corresponding to position 145 of SEQ ID NO: 5; y) A at a position corresponding to position 145 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; z) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; aa) S at a position corresponding to position 81 of SEQ ID NO: 5, W at a position corresponding to position 155 of SEQ ID NO: 5, and D at a position corresponding to position 160 of SEQ ID NO: 5; bb) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; cc) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; dd) a W at a position corresponding to position 75 of SEQ ID NO:5, an S at a position corresponding to position 81 of SEQ ID NO:5, and an A at a position corresponding to position 145 of SEQ ID NO:5; ee) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ff) W at a position corresponding to position 75 of SEQ ID NO: 5, W at a position corresponding to position 155 of SEQ ID NO: 5, and D at a position corresponding to position 160 of SEQ ID NO: 5; gg) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a W at a position corresponding to position 155 of SEQ ID NO:5; hh) A at a position corresponding to position 145 of SEQ ID NO: 5, W at a position corresponding to position 155 of SEQ ID NO: 5, and D at a position corresponding to position 160 of SEQ ID NO: 5; ii) W at a position corresponding to position 75 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; jj) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; kk) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; ll) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; mm) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; nn) W at a position corresponding to position 75 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; oo) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; pp) L at a position corresponding to position 40 of SEQ ID NO: 5, and M at a position corresponding to position 105 of SEQ ID NO: 5; qq) L at a position corresponding to position 40 of SEQ ID NO: 5, and V at a position corresponding to position 121 of SEQ ID NO: 5; rr) L at a position corresponding to position 40 of SEQ ID NO: 5, and M at a position corresponding to position 145 of SEQ ID NO: 5; ss) L at a position corresponding to position 40 of SEQ ID NO: 5, and Q at a position corresponding to position 159 of SEQ ID NO: 5; tt) L at a position corresponding to position 40 of SEQ ID NO: 5 and W at a position corresponding to position 160 of SEQ ID NO: 5; uu) M at a position corresponding to position 105 of SEQ ID NO: 5 and V at a position corresponding to position 121 of SEQ ID NO: 5; vv) M at a position corresponding to position 105 of SEQ ID NO: 5, and M at a position corresponding to position 145 of SEQ ID NO: 5; ww) M at a position corresponding to position 105 of SEQ ID NO: 5, and Q at a position corresponding to position 159 of SEQ ID NO: 5; xx) M at a position corresponding to position 105 of SEQ ID NO: 5 and W at a position corresponding to position 160 of SEQ ID NO: 5; yy) V at a position corresponding to position 121 of SEQ ID NO: 5, and M at a position corresponding to position 145 of SEQ ID NO: 5; zz) V at a position corresponding to position 121 of SEQ ID NO: 5 and Q at a position corresponding to position 159 of SEQ ID NO: 5; aaa) V at a position corresponding to position 121 of SEQ ID NO: 5 and W at a position corresponding to position 160 of SEQ ID NO: 5; bbb) M at a position corresponding to position 145 of SEQ ID NO: 5, and Q at a position corresponding to position 159 of SEQ ID NO: 5; ccc) M at a position corresponding to position 145 of SEQ ID NO: 5, and W at a position corresponding to position 160 of SEQ ID NO: 5; ddd) Q at a position corresponding to position 159 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; eee) S at a position corresponding to position 81 of SEQ ID NO:5, M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; fff) S at a position corresponding to position 81 of SEQ ID NO:5, V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ggg) A at a position corresponding to position 68 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; hhh) G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; iii) S at a position corresponding to position 81 of SEQ ID NO:5 and L at a position corresponding to position 117 of SEQ ID NO:5; jjj) A at a position corresponding to position 68 of SEQ ID NO: 5 and G at a position corresponding to position 76 of SEQ ID NO: 5; kkk) A at a position corresponding to position 68 of SEQ ID NO: 5, and L at a position corresponding to position 117 of SEQ ID NO: 5; lll) G at a position corresponding to position 76 of SEQ ID NO: 5 and L at a position corresponding to position 117 of SEQ ID NO: 5; mmm) an A at a position corresponding to position 68 of SEQ ID NO:5, a G at a position corresponding to position 76 of SEQ ID NO:5, and an S at a position corresponding to position 81 of SEQ ID NO:5; nnn) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ooo) G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ppp) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; qqq) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; rrr) S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 109 of SEQ ID NO:5, E at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; sss) A at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, A at a position corresponding to position 121 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ttt) Q at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, I at a position corresponding to position 117 of SEQ ID NO:5, Q at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; uuu) Y at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, G at a position corresponding to position 121 of SEQ ID NO:5, L at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; vvv) a C at a position corresponding to position 22 of SEQ ID NO:5, an A at a position corresponding to position 68 of SEQ ID NO:5, a Y at a position corresponding to position 75 of SEQ ID NO:5, a G at a position corresponding to position 76 of SEQ ID NO:5, an S at a position corresponding to position 81 of SEQ ID NO:5, an H at a position corresponding to position 108 of SEQ ID NO:5, an L at a position corresponding to position 117 of SEQ ID NO:5, an H at a position corresponding to position 122 of SEQ ID NO:5, a Y at a position corresponding to position 138 of SEQ ID NO:5, an L at a position corresponding to position 139 of SEQ ID NO:5, an A at a position corresponding to position 142 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, an R at a position corresponding to position 153 of SEQ ID NO:5, a W at a position corresponding to position 155 of SEQ ID NO:5, an R at a position corresponding to position 156 of SEQ ID NO:5, a D at a position corresponding to position 160 of SEQ ID NO:5, and an H at a position corresponding to position 162 of SEQ ID NO:5; or www) has an S at the position corresponding to position 81 of SEQ ID NO:5.

[0080] In some embodiments, the deaminase comprises an amino acid sequence having an N at a position corresponding to position 35 of SEQ ID NO:5, an S at a position corresponding to position 81 of SEQ ID NO:5, an R at a position corresponding to position 156 of SEQ ID NO:5, and a W at a position corresponding to position 162 of SEQ ID NO:5. In some embodiments, the deaminase comprises an amino acid sequence having a W at a position corresponding to position 75 of SEQ ID NO:5, an S at a position corresponding to position 81 of SEQ ID NO:5, a W at a position corresponding to position 155 of SEQ ID NO:5, and a D at a position corresponding to position 160 of SEQ ID NO:5. In some embodiments, the deaminase comprises an amino acid sequence having an S at a position corresponding to position 81 of SEQ ID NO:5, and an L at a position corresponding to position 117 of SEQ ID NO:5. In some embodiments, the deaminase comprises an amino acid sequence having an A at a position corresponding to position 68 of SEQ ID NO:5, and an S at a position corresponding to position 81 of SEQ ID NO:5. In some embodiments, the deaminase comprises an amino acid sequence having a Y at a position corresponding to position 75 of SEQ ID NO:5, an S at a position corresponding to position 81 of SEQ ID NO:5, an L at a position corresponding to position 117 of SEQ ID NO:5, a G at a position corresponding to position 121 of SEQ ID NO:5, an L at a position corresponding to position 145 of SEQ ID NO:5, an F at a position corresponding to position 155 of SEQ ID NO:5, and a D at a position corresponding to position 160 of SEQ ID NO:5. In some embodiments, the deaminase comprises an amino acid sequence having an S at a position corresponding to position 81 of SEQ ID NO:5, a W at a position corresponding to position 155 of SEQ ID NO:5, and a D at a position corresponding to position 160 of SEQ ID NO:5. In some embodiments, the deaminase comprises an amino acid sequence having an S at a position corresponding to position 81 of SEQ ID NO:5, an L at a position corresponding to position 109 of SEQ ID NO:5, an E at a position corresponding to position 153 of SEQ ID NO:5, a W at a position corresponding to position 155 of SEQ ID NO:5, and a D at a position corresponding to position 160 of SEQ ID NO:5.

[0081] In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO: 303. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO: 319. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO: 365. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO: 375. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO: 406. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO: 596. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO: 597. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO: 720.

[0082] In some embodiments, the deaminase consists of the amino acid sequence of SEQ ID NO: 303. In some embodiments, the deaminase consists of the amino acid sequence of SEQ ID NO: 319. In some embodiments, the deaminase consists of the amino acid sequence of SEQ ID NO: 365. In some embodiments, the deaminase consists of the amino acid sequence of SEQ ID NO: 375. In some embodiments, the deaminase consists of the amino acid sequence of SEQ ID NO: 406. In some embodiments, the deaminase consists of the amino acid sequence of SEQ ID NO: 596. In some embodiments, the deaminase consists of the amino acid sequence of SEQ ID NO: 597. In some embodiments, the deaminase consists of the amino acid sequence of SEQ ID NO: 720.

[0083] The deaminases of the present invention can have improved deaminase activity compared to the parent LPG50148 deaminase. Improved deaminase activity can be measured by any method known in the art for measuring deamination of nucleobases (e.g., adenine) alone or when fused to a DNA-binding polypeptide (e.g., RGN). The deaminases disclosed herein can have improved deaminase activity, such as about 1.1-fold, about 1.5-fold, about 2-fold, about 2.5-fold, about 3-fold, about 3.5-fold, about 4-fold, about 4.5-fold, about 5-fold, about 5.5-fold, about 6-fold, about 6.5-fold, about 7-fold, about 7.5-fold, about 8-fold, about 8.5-fold, about 9-fold, about 9.5-fold, about 10-fold, about 11-fold, about 12-fold, about 13-fold, about 15-fold, about 16-fold, about 17-fold, about 18-fold, about 19-fold, about 20-fold, about 21-fold, about 22-fold, about 23-fold, about 24-fold, about 25-fold, about 26-fold, about 27-fold, about 28-fold, about 29-fold, about 30-fold, about 31-fold, about 32-fold, about 33-fold, about 34-fold, about 35-fold, about 36-fold, about 37-fold, about 38-fold, about 39-fold, about 40-fold, about 41-fold, about 42-fold, about 43-fold, about 44-fold, about 45-fold, about 46-fold, about 47-fold, about 48-fold, about 49-fold The efficiency of deaminating nucleobases can be 1.1-10, 1.1-20, 1.1-30, or more times greater than the parent LPG50148 deaminase, including, but not limited to, 9-fold, about 20-fold, about 21-fold, about 22-fold, about 23-fold, about 24-fold, about 25-fold, about 26-fold, about 27-fold, about 28-fold, about 29-fold, and about 30-fold greater.

[0084] IV. Heterologous Polypeptides In some embodiments, the fusion protein of the present disclosure comprises a heterologous polypeptide inserted within RGN. The heterologous polypeptide is any polypeptide that is not naturally bound to the RGN protein. In some embodiments, the heterologous polypeptide comprises a protein domain having biological activity. The heterologous polypeptide can comprise a detectable label, a selectable marker, or a purification tag. In some embodiments, the heterologous polypeptide is a base-editing polypeptide (e.g., a deaminase) or a prime-editing polypeptide.

[0085] A. Prime Edited Polypeptides The heterologous polypeptide of the fusion protein of the present disclosure can be a prime editing polypeptide, such as a reverse transcriptase.

[0086] Prime editing is a versatile and precise genome editing method that uses nucleic acid-programmable DNA-binding proteins that function in association with polymerases to write new genetic information directly into designated DNA sites (e.g., as described in US11,447,770B1, WO2021072328, WO2021226558, WO2020156575, WO2021042047, and US11193123, each of which is incorporated by reference in its entirety). The prime editing system uses the nickase RGN (generally with an inactivated HNH domain), and the system is programmed with a prime editing (PE) guide RNA ("PEgRNA"). The PEgRNA is a guide RNA that specifies the target sequence and provides a template for polymerization of a replacement strand containing the edit through an engineered extension on the guide RNA (e.g., at the 5' or 3' end, or in an internal portion of the guide RNA). The RGN nickase / prime editing polypeptide fusion is guided to the target sequence by the PEG RNA and nicks the target strand upstream of the sequence to be edited and upstream of the PAM, creating a 3' flap on the target strand. The PEG RNA contains a primer binding site (PBS) that is complementary to the 3' flap of the target strand. In some embodiments, the PBS is at least about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides in length. In certain embodiments, the pegRNA comprises a PBS at least 5 (e.g., at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 28, 19, or 20) nucleotides in length. In some embodiments, the pegRNA may comprise a PBS at least 8 nucleotides in length. Hybridization of the PBS and the 3' flap of the target strand allows polymerization of a replacement strand containing the edit, using the extension of the pegRNA as a template.The extension of the PEG RNA can be formed from RNA or DNA. In the case of RNA extension, the polymerase of the prime editor can be an RNA-dependent DNA polymerase (such as a reverse transcriptase). In the case of DNA extension, the polymerase of the prime editor can be a DNA-dependent DNA polymerase.

[0087] The replacement strand containing the desired edit (e.g., a single nucleobase substitution) shares the same sequence as the target strand of the target sequence to be edited (except for the desired edit). Through DNA repair and / or replication mechanisms, the target strand of the target sequence is replaced by a newly synthesized replacement strand containing the desired edit. In some cases, prime editing can be considered a "search-and-replace" genome editing technique because the prime editor not only searches for and locates the desired target sequence to be edited, but also simultaneously encodes the replacement strand containing the desired edit, which is installed in place of the corresponding target strand of the target sequence. Thus, in some embodiments, a guide RNA useful in the compositions and methods of the present disclosure includes an extension containing an editing template for prime editing. In some embodiments, a prime editing polypeptide that can be fused to an RGN comprises a DNA polymerase (e.g., an RNA-dependent DNA polymerase). In certain embodiments, the DNA polymerase is a reverse transcriptase. In certain embodiments, the RGN is a nickase, such as one in which the HNH domain is inactivated.

[0088] B. Base-edited polypeptide The heterologous polypeptide of the fusion protein of the present disclosure can be a base-editing polypeptide, such as a deaminase.

[0089] In some embodiments, the deaminase component of a fusion protein of the disclosure (in which the deaminase is inserted into an RGN) is any one of SEQ ID NOs: 5, 6, 229-414, 596, 597, and 720-723, or an active fragment or variant thereof, including the deaminases disclosed herein (SEQ ID NOs: 300-414, 596-598, and 720-723). The deaminase component can have 80% to 99% or more sequence identity to any one of SEQ ID NOs: 5, 6, 229-414, 596, 597, and 720-723 (which retain deaminase activity), including but not limited to at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or more. The fusion protein can comprise an active fragment of a deaminase set forth as any one of SEQ ID NOs: 5, 6, 229-414, 596, 597, and 720-723, such as one that differs from any one of SEQ ID NOs: 5, 6, 229-414, 596, 597, and 720-723 by as few as 1-15 amino acid residues, 6-10, as few as 5, as few as 4, as few as 3, as few as 2, or as few as 1 amino acid residue. In some embodiments, the deaminase inserted into RGN has about 50% to about 100% identity to any one of SEQ ID NOs: 5, 6, 300-414, 596-598, and 720-723 and at least one of the amino acid residues set forth in any one of Tables 2, 4, 6, and 19.In some embodiments, the deaminase has an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to any of SEQ ID NOs: 5, 6, 300-414, 596-598, and 720-723, and has at least one of the amino acid residues set forth in any one of Tables 2, 4, 6, and 19.

[0090] In certain embodiments, the active deaminase comprises an N- or C-terminal truncation that can include a deletion of 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, or more amino acids from either the N- or C-terminus of the polypeptide set forth in any one of SEQ ID NOs: 5, 6, 229-414, 596, 597, and 720-723. A fusion protein can comprise a deaminase lacking the first and last amino acid residues compared to the parent deaminase from which the deaminase is derived, and deaminase activity is retained. A fusion protein can comprise a deaminase lacking the first and last amino acid residues of any one of SEQ ID NOs: 5, 6, 229-414, 596, 597, and 720-723, and deaminase activity is retained. In some embodiments, active deaminase variants comprise an internal deletion, which can include deletion of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, or more amino acids of any one of SEQ ID NOs: 5, 6, 229-414, 596, 597, and 720-723. An active fragment of a deaminase can contain at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, or more consecutive amino acid residues of the amino acid sequence set forth in any one of SEQ ID NOs: 5, 6, 229-414, 596, 597, and 720-723.

[0091] In some embodiments, the deaminase component of a fusion protein of the disclosure (in which the deaminase is inserted into RGN) is any one of SEQ ID NOs: 303, 319, 365, 375, 406, 597, and 720, or an active variant or fragment thereof.

[0092] A "base editor" is a fusion protein comprising an RGN operably linked to a deaminase such that the fusion protein is capable of deaminating a nucleobase of a target nucleic acid sequence when bound to a guide RNA. The RGN of a base editor generally comprises an RGN nickase (typically one with an inactivated RuvC domain) or an inactive RGN that does not have nuclease activity. A base editor fusion protein in which a deaminase is inserted within an RGN is referred to herein as an "inlaid base editor."

[0093] Fusion proteins of the present disclosure can include adenine deaminase, such as those disclosed herein (e.g., 5, 6, 248-414, 596, 597, and 720-723). Base editors comprising RGN and adenine deaminase are referred to herein as "A base editors," "adenine base editors," or "ABEs" and can be used for targeted editing of nucleic acid sequences. To date, no naturally occurring adenine deaminases are known that deaminate adenine in DNA. Several methods have been employed to evolve and optimize adenine deaminase acting on tRNA (ADAT) proteins to be active on DNA molecules in mammalian cells (Gaudelli et al., 2017; Koblan, L. Wet et al., 2018, Nat Biotechnol 36, 843-846; Richter, M. F. et al., 2020, Nat Biotechnol, doi:10.1038 / s41587-020-0562-8, each incorporated by reference in its entirety).

[0094] Non-limiting examples of adenine deaminases that can be used in the present invention include, but are not limited to, those disclosed in PCT International Publication No. WO2022 / 056254 and SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723, or active variants or fragments thereof.

[0095] Adenine base editors (ABEs) comprise RGN and adenine deaminase. ABEs function through the deamination of adenine to inosine on DNA target molecules (Gaudelli, NM et al. 2017). Inosine is recognized as guanine by polymerases, allowing for the incorporation of cytosine directly opposite the inosine on the complementary DNA strand. After one round of replication following deamination, there is a resulting A:T to G:C base pair change in the genome. In some embodiments, the adenine deaminase fusion protein or an active variant or fragment thereof introduces an A>N mutation into a DNA molecule, where N is C, G, or T. In further embodiments, they introduce an A>G mutation into a DNA molecule.

[0096] In some embodiments, a fusion protein of the present disclosure comprising an RGN having a deaminase inserted therein can comprise a cytosine deaminase that catalyzes the hydrolytic deamination of cytosine, cytidine, or deoxycytidine to uracil. The cytosine deaminase can act on either DNA or RNA, typically acting on single-stranded nucleic acid molecules. In some embodiments, the cytosine deaminase is an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the cytosine deaminase is an APOBECl family deaminase. In some embodiments, the cytosine deaminase is activation-induced cytidine deaminase (AID). In some embodiments, the cytosine deaminase is an ACF1 / ASE deaminase.

[0097] Fusion proteins comprising RGN and cytosine deaminase are referred to herein as "C base editors," "cytosine base editors," or "CBEs." CBEs can convert cytosine to uracil, which can then be converted to thymine during DNA replication or repair. In some embodiments, CBEs convert cytosine to guanine or cytosine to adenine. Without being bound by any theory or mechanism of action, the conversion of cytosine to guanine or adenine by a cytosine base editor is believed to result from the deamination of cytosine to uracil and the subsequent activity of uracil DNA glycosylase during base excision repair of the uracil residue. In some embodiments, the cytosine deaminase or active variant or fragment thereof of the fusion protein introduces a C>N mutation into a DNA molecule, where N is A, G, or T. In further embodiments, they introduce a C>G or C>A mutation into a DNA molecule.

[0098] Non-limiting examples of cytosine deaminases that can be used in the present invention include, but are not limited to, those disclosed in PCT International Publication Nos. WO2020 / 139783 and WO2022 / 204093, and SEQ ID NOs: 229-247, or active variants or fragments thereof.

[0099] The mutation rate of adenines or cytosines within or adjacent to the target sequence bound by the RGN-deaminase fusion protein can be measured using any method known in the art, including polymerase chain reaction (PCR), restriction fragment length polymorphism (RFLP), or DNA sequencing.

[0100] V. Fusion Proteins The present invention provides various types of fusion proteins. Those skilled in the art will understand that whenever a protein is fused to another protein, the N-terminal methionine (Met) of the C-terminal protein may optionally be omitted from the sequence. Thus, as used herein, a reference to a given amino acid sequence fused to another amino acid sequence explicitly includes such a sequence with or optionally without an N-terminal Met, regardless of whether such sequence includes an N-terminal Met in the sequence listing.

[0101] Some embodiments of the present invention involve fusion proteins comprising a heterologous polypeptide inserted into RGN. The heterologous polypeptide can be inserted into RGN at its surface (i.e., inserted at or between surface amino acid residues within surface loops).

[0102] The heterologous polypeptide can be inserted into linker domain 2, wedge (WED) domain, RuvC domain, HNH domain, Rec-2 domain, or PAM-interacting (PI) domain. In some embodiments, the RuvC domain is the RuvCIII domain. The Rec or recognition lobe mediates nucleic acid binding through multiple Rec domains (e.g., Rec1-3) by sensing nucleic acids, regulating HNH conformational transitions, and locking the catalytic HNH domain at the cleavage site. The wedge domain is involved in recognition of the guide RNA scaffold. Non-limiting examples of domains within RGN include RuvC-I from amino acid residues 1-54, BH from amino acid residues 55-83, REC1 from amino acid residues 84-244, REC2 from amino acid residues 245-462, RuvC-II from amino acid residues 463-521, L1 from amino acid residues 522-552, HNH from amino acid residues 553-672, L2 from amino acid residues 673-685, RuvC-III from amino acid residues 686-833, WED from amino acid residues 834-938, and PI from amino acid residues 939-1071, all related to APG07433.1 set forth as SEQ ID NO:1. APG05586 (shown as SEQ ID NO:2) has the following domains: RuvC-I from amino acid residues 1-33, BH from amino acid residues 34-71, REC1 from amino acid residues 72-232, REC2 from amino acid residues 233-468, RuvC-II from amino acid residues 469-517, L1 from amino acid residues 518-552, HNH from amino acid residues 553-672, L2 from amino acid residues 673-687, RuvC-III from amino acid residues 688-837, WED from amino acid residues 838-998, and PI from amino acid residues 999-1150.LPG10145 (shown as SEQ ID NO:4) has the following domains: RuvC-I from amino acid residues 1-42, BH from amino acid residues 43-79, REC1 from amino acid residues 80-236, REC2 from amino acid residues 237-476, RuvC-II from amino acid residues 477-524, L1 from amino acid residues 525-560, HNH from amino acid residues 561-676, L2 from amino acid residues 677-690, RuvC-III from amino acid residues 691-828, WED from amino acid residues 829-976, and PI from amino acid residues 977-1130. APG01604 (shown as SEQ ID NO:3) has the following domains: RuvC-I from amino acid residues 1-40, BH from amino acid residues 41-74, REC1 from amino acid residues 75-223, REC2 from amino acid residues 224-430, RuvC-II from amino acid residues 431-483, L1 from amino acid residues 484-516, HNH from amino acid residues 517-631, L2 from amino acid residues 632-651, RuvC-III from amino acid residues 652-775, WED from amino acid residues 776-909, and PI from amino acid residues 910-1052.

[0103] The PAM-interacting domain is the domain that binds to the PAM sequence. The general domain of an RGN protein can be determined through structural comparison with RGN proteins that have defined domains.

[0104] In those embodiments in which the fusion protein comprises an RGN having at least 90% sequence identity to SEQ ID NO: 1 or 49, the heterologous polypeptide (e.g., deaminase) is located within the RGN and i) the amino acid position corresponding to position 30 of SEQ ID NO: 1; ii) an amino acid position corresponding to position 642 of SEQ ID NO: 1; iii) an amino acid position corresponding to position 670 of SEQ ID NO: 1; iv) an amino acid position corresponding to position 737 of SEQ ID NO: 1; v) an amino acid position corresponding to position 772 of SEQ ID NO: 1; vi) an amino acid position corresponding to position 775 of SEQ ID NO: 1; vii) the amino acid position corresponding to position 778 of SEQ ID NO: 1; and viii) It can be inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 802 of SEQ ID NO:1.

[0105] In those embodiments in which the fusion protein comprises an RGN having at least 90% sequence identity to SEQ ID NO: 2 or 50, the heterologous polypeptide comprises a heterologous polypeptide within the RGN: i) the amino acid position corresponding to position 678 of SEQ ID NO: 2; ii) an amino acid position corresponding to position 736 of SEQ ID NO: 2; iii) an amino acid position corresponding to position 778 of SEQ ID NO: 2; iv) an amino acid position corresponding to position 788 of SEQ ID NO: 2, and v) It can be inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 922 of SEQ ID NO:2.

[0106] In those embodiments in which the fusion protein comprises an RGN having at least 90% sequence identity to SEQ ID NO: 3 or 51, the heterologous polypeptide comprises, within the RGN: i) the amino acid position corresponding to position 725 of SEQ ID NO: 3; ii) an amino acid position corresponding to position 739 of SEQ ID NO: 3, and iii) It can be inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 744 of SEQ ID NO:3.

[0107] In those embodiments in which the fusion protein comprises an RGN having at least 90% sequence identity to SEQ ID NO: 4 or 52, the heterologous polypeptide comprises a heterologous polypeptide comprising, within the RGN: i) the amino acid position corresponding to position 347 of SEQ ID NO: 4; ii) an amino acid position corresponding to position 524 of SEQ ID NO: 4; iii) an amino acid position corresponding to position 666 of SEQ ID NO: 4; iv) an amino acid position corresponding to position 680 of SEQ ID NO: 4; v) an amino acid position corresponding to position 740 of SEQ ID NO: 4; vi) an amino acid position corresponding to position 785 of SEQ ID NO: 4; vii) an amino acid position corresponding to position 910 of SEQ ID NO: 4, and viii) It can be inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 1077 of SEQ ID NO:4.

[0108] In those embodiments in which the fusion protein comprises an RGN having at least 90% sequence identity to SEQ ID NO: 131 or 698, the heterologous polypeptide comprises a heterologous polypeptide comprising, within the RGN: i) the amino acid position corresponding to position 766 of SEQ ID NO: 131, and ii) It can be inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 806 of SEQ ID NO:131.

[0109] Fusion proteins of the invention can comprise a base editor fusion protein, wherein the heterologous polypeptide is a deaminase, and the deaminase is inserted within an RGN, such as an RGN nickase or a nuclease-inactive RGN. In some embodiments, the RGN component of the base editor fusion protein comprises an RGN nickase or a nuclease-inactive RGN having about 60% to about 99.5% identity to any one of SEQ ID NOs: 49-52 and 698. In some embodiments, the RGN nickase or nuclease-inactive RGN comprises an amino acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to any one of SEQ ID NOs: 49-52 and 698.

[0110] In some embodiments, the base editor fusion protein comprises an RGN nickase or nuclease-inactive RGN comprising an amino acid sequence having 80% to 99% or more sequence identity to any one of SEQ ID NOs: 49-52 and 698, including but not limited to at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more.

[0111] Fusion proteins of the invention can include base editor fusion proteins (wherein the heterologous polypeptide is a deaminase and the deaminase is inserted into an RGN, such as an RGN nickase or a nuclease-inactive RGN), which have improved editing activity compared to a parent base editor comprising a deaminase fused to the amino terminus of an RGN. Improved editing activity can be measured by any method known in the art. The base editor fusion proteins disclosed herein exhibit improved editing activity of about 1.1-fold, about 1.5-fold, about 2-fold, about 2.5-fold, about 3-fold, about 3.5-fold, about 4-fold, about 4.5-fold, about 5-fold, about 5.5-fold, about 6-fold, about 6.5-fold, about 7-fold, about 7.5-fold, about 8-fold, about 8.5-fold, about 9-fold, about 9.5-fold, about 10-fold, about 11-fold, about 12-fold, about 13-fold, about 15-fold, about 16-fold, about 17-fold, about 18-fold, about 19-fold, about 20-fold, about 21-fold, about 22-fold, about 23-fold, about 24-fold, about 25-fold, about 26-fold, about 27-fold, about 28-fold, about 29-fold, about 30-fold, about 31-fold, about 32-fold, about 33-fold, about 34-fold, about 35-fold, about 36-fold, about 37-fold, about 38-fold, about 39-fold, about 40-fold, about 41-fold, about 42-fold, about 43-fold, about 44-fold, about 45-fold, about 46-fold, about 47-fold, about 48-fold The editing efficiency may be 1.1-fold to 10-fold, 1.1-fold to 20-fold, 1.1-fold to 30-fold, or more, greater than the parent end-to-end fusion protein, including, but not limited to, 7-fold, about 18-fold, about 19-fold, about 20-fold, about 21-fold, about 22-fold, about 23-fold, about 24-fold, about 25-fold, about 26-fold, about 27-fold, about 28-fold, about 29-fold, and about 30-fold.

[0112] Binding of the base editor fusion protein to the target sequence results in modification of nucleotides adjacent to the target sequence. The nucleobases adjacent to the target sequence that are modified by the deaminase can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 base pairs from the 5' or 3' end of the target sequence. The full range of nucleotides that can be edited using a base editor fusion protein (e.g., an RGN fused to a deaminase), typically expressed as the distance of nucleotides from the PAM sequence (e.g., the nucleotides at positions 8 and 22 upstream, i.e., 5', from the PAM sequence), is referred to herein as the "editing window" of a particular base editor fusion protein. The editing window can be, for example, 1 to 100 base pairs 5' or 3' of the PAM sequence, including, but not limited to, 5 to 50, 5 to 25, 8 to 22, 10 to 20, or 10 to 21 base pairs 5' or 3' of the PAM sequence.

[0113] Non-limiting examples of inlaid base editor fusion proteins include: a) LPG50274 / LPG50148.2.76 (shown as SEQ ID NO: 375) or an active variant or fragment thereof, inserted within nAPG05586 (shown as SEQ ID NO: 50) or an active variant or fragment thereof after the amino acid position corresponding to position 922 of SEQ ID NO: 2, such as the sequence shown as SEQ ID NO: 599; b) LPG50274 / LPG50148.2.76 (shown as SEQ ID NO: 375) or an active variant or fragment thereof, inserted within nLPG10145 (shown as SEQ ID NO: 52) or an active variant or fragment thereof after the amino acid position corresponding to position 910 of SEQ ID NO: 4, such as the sequence shown as SEQ ID NO: 600; c) LPG50221 / LPG50148.2.20 (shown as SEQ ID NO: 319) or an active variant or fragment thereof, inserted within nAPG05586 (shown as SEQ ID NO: 50) or an active variant or fragment thereof after the amino acid position corresponding to position 922 of SEQ ID NO: 2, such as the sequence shown as SEQ ID NO: 602; d) LPG50221 / LPG50148.2.20 (shown as SEQ ID NO: 319) or an active variant or fragment thereof, inserted within nLPG10145 (shown as SEQ ID NO: 52) or an active variant or fragment thereof after the amino acid position corresponding to position 910 of SEQ ID NO: 4, such as the sequence shown as SEQ ID NO: 603; e) LPG50274 / LPG50148.2.76 (shown as SEQ ID NO: 375) or an active variant or fragment thereof, inserted within nAPG07433.1 (shown as SEQ ID NO: 49) or an active variant or fragment thereof after the amino acid position corresponding to position 772 of SEQ ID NO: 1, such as the sequence shown as SEQ ID NO: 606; f) LPG50274 / LPG50148.2.76 (shown as SEQ ID NO: 375) or an active variant or fragment thereof, inserted within nAPG05586 (shown as SEQ ID NO: 50) or an active variant or fragment thereof after the amino acid position corresponding to position 678 of SEQ ID NO: 2, such as the sequence shown as SEQ ID NO: 610; g) LPG50221 / LPG50148.2.20 (shown as SEQ ID NO: 319) or an active variant or fragment thereof, inserted within nAPG07433.1 (shown as SEQ ID NO: 49) or an active variant or fragment thereof after the amino acid position corresponding to position 772 of SEQ ID NO: 1, such as the sequence shown as SEQ ID NO: 612; h) LPG50324 (shown as SEQ ID NO: 597) or an active variant or fragment thereof, inserted within nAPG05586 (shown as SEQ ID NO: 50) or an active variant or fragment thereof after the amino acid position corresponding to position 922 of SEQ ID NO: 2, such as the sequence shown as SEQ ID NO: 614; i) LPG50221 / LPG50148.2.20 (shown as SEQ ID NO: 319) or an active variant or fragment thereof, inserted within nAPG05586 (shown as SEQ ID NO: 50) or an active variant or fragment thereof after the amino acid position corresponding to position 678 of SEQ ID NO: 2, such as the sequence shown as SEQ ID NO: 615; j) LPG50319 / LPG50148.2.107 (shown as SEQ ID NO: 406) or an active variant or fragment thereof, inserted within nAPG05586 (shown as SEQ ID NO: 50) or an active variant or fragment thereof after the amino acid position corresponding to position 678 of SEQ ID NO: 2, such as the sequence shown as SEQ ID NO: 617; k) LPG50324 (shown as SEQ ID NO: 597) or an active variant or fragment thereof, inserted within nAPG05586 (shown as SEQ ID NO: 50) or an active variant or fragment thereof after the amino acid position corresponding to position 678 of SEQ ID NO: 2, such as the sequence shown as SEQ ID NO: 619; l) LPG50320 (shown as SEQ ID NO: 596) or an active variant or fragment thereof, inserted within nAPG05586 (shown as SEQ ID NO: 50) or an active variant or fragment thereof after the amino acid position corresponding to position 678 of SEQ ID NO: 2, such as the sequence shown as SEQ ID NO: 620; m) LPG50319 / LPG50148.2.107 (shown as SEQ ID NO: 406) or an active variant or fragment thereof, inserted within nLPG10145 (shown as SEQ ID NO: 52) or an active variant or fragment thereof after the amino acid position corresponding to position 910 of SEQ ID NO: 4, such as the sequence shown as SEQ ID NO: 621; n) LPG50324 (shown as SEQ ID NO: 597) or an active variant or fragment thereof, inserted in nLPG10145 (shown as SEQ ID NO: 52) or an active variant or fragment thereof after the amino acid position corresponding to position 910 of SEQ ID NO: 4, such as the sequence shown as SEQ ID NO: 622; o) LPG50319 / LPG50148.2.107 (shown as SEQ ID NO: 406) or an active variant or fragment thereof, inserted within nAPG05586 (shown as SEQ ID NO: 50) or an active variant or fragment thereof after the amino acid position corresponding to position 922 of SEQ ID NO: 2, such as the sequence shown as SEQ ID NO: 623; p) LPG50320 (shown as SEQ ID NO: 596) or an active variant or fragment thereof, inserted in nLPG10145 (shown as SEQ ID NO: 52) or an active variant or fragment thereof after the amino acid position corresponding to position 910 of SEQ ID NO: 4, such as the sequence shown as SEQ ID NO: 624; q) LPG50320 (shown as SEQ ID NO: 596) or an active variant or fragment thereof, inserted within nAPG05586 (shown as SEQ ID NO: 50) or an active variant or fragment thereof after the amino acid position corresponding to position 922 of SEQ ID NO: 2, such as the sequence shown as SEQ ID NO: 626; r) LPG50319 / LPG50148.2.107 (shown as SEQ ID NO: 406) or an active variant or fragment thereof, inserted within nAPG07433.1 (shown as SEQ ID NO: 49) or an active variant or fragment thereof after the amino acid position corresponding to position 772 of SEQ ID NO: 1, such as the sequence shown as SEQ ID NO: 630; s) LPG50320 (shown as SEQ ID NO: 596) or an active variant or fragment thereof inserted within nAPG07433.1 (shown as SEQ ID NO: 49) or an active variant or fragment thereof after the amino acid position corresponding to position 772 of SEQ ID NO: 1, such as the sequence shown as SEQ ID NO: 631, and t) LPG50324 (shown as SEQ ID NO:597) or an active variant or fragment thereof inserted into nAPG07433.1 (shown as SEQ ID NO:49) or an active variant or fragment thereof after the amino acid position corresponding to position 772 of SEQ ID NO:1, such as the sequence shown as SEQ ID NO:632.

[0114] Fusion proteins of the invention can include base editor fusion proteins (wherein the heterologous polypeptide is a deaminase and the deaminase is inserted within an RGN, such as an RGN nickase or a nuclease-inactive RGN), which have a shifted editing window compared to a parent base editor comprising a deaminase fused to the amino terminus of an RGN. The editing window can be shifted to be narrower or wider than the parent base editor. The shifted editing window is a wider or narrower window (e.g., by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides) compared to the parent end-to-end fusion protein, and / or the editing window is moved in the 5' or 3' direction of the target molecule, e.g., by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides in the 5' or 3' direction, compared to the parent end-to-end fusion protein.

[0115] In some embodiments, the fusion protein comprises one or more peptide linkers between the heterologous polypeptide and RGN. Those skilled in the art will understand that whenever a protein is fused to another protein, a linker may optionally be added to the sequence. Therefore, as used herein, reference to a given fusion protein explicitly includes such a sequence, with or without a linker, regardless of whether such a sequence is listed as a linker. The linker between the deaminase and RGN can determine the editing window of the fusion protein. To achieve optimal length and rigidity of deaminase activity for a specific application, the following structure may be used: (GGGGS) n and (G) n Formation of highly flexible linkers (EAAAK) n and (XP) nA variety of linker lengths and flexibilities can be employed, ranging from more rigid linkers of the same or similar structure. As used herein, the term "peptide linker" refers to a peptide that links two polypeptides. In some embodiments, the linker joins an RNA-guided nuclease and a deaminase. In some embodiments, the linker joins an inactive or inactive RGN and a deaminase. In some embodiments, the linker is 3 to 100 amino acids in length, e.g., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated. In some embodiments, shorter linkers are preferred to reduce the overall size or length of the fusion protein or its coding sequence. Non-limiting examples of peptide linkers between RGN and the heterologous polypeptide inserted therein include GGS, SGG, and GSSG (shown as SEQ ID NO: 642). The linker can be at the N-terminus, C-terminus, or both of the inserted heterologous polypeptide.

[0116] In some embodiments, the RGN and the heterologous polypeptide inserted therein are fused directly to each other without a linker sequence between them.

[0117] Another aspect of the present invention involves a fusion protein comprising a DNA-binding polypeptide (e.g., a nuclease-inactive or nickase RGN) operably linked to a deaminase of the present invention (e.g., SEQ ID NOS: 6, 300-414, 596-598, and 720-723, or an active fragment or variant thereof). In some embodiments, the DNA-binding polypeptide (e.g., a nuclease-inactive RGN or nickase RGN) fused to a deaminase of the present invention can be targeted to a specific location of a nucleic acid molecule (i.e., a target nucleic acid molecule), in some embodiments a specific genomic locus, to modify expression of a desired sequence. In some embodiments, binding of the fusion protein to the target sequence results in deamination of a nucleobase, resulting in the conversion of one nucleobase to another. In some embodiments, binding of the fusion protein to the target sequence results in deamination of a nucleobase adjacent to the target sequence.

[0118] Some aspects of the present disclosure provide fusion proteins comprising (i) a DNA-binding polypeptide (e.g., a nuclease-inactive or nickase RGN polypeptide), (ii) a deaminase polypeptide, and optionally (iii) a second deaminase. The second deaminase can be the same as the first deaminase or can be a different deaminase. In some embodiments, both the first and second deaminase are adenine deaminases of the invention.

[0119] The present disclosure provides fusion proteins in various configurations, in some embodiments, the fusion protein is an end-to-end fusion in which the deaminase polypeptide is fused to the N-terminus of the DNA-binding polypeptide (e.g., an RGN polypeptide) or the deaminase polypeptide is fused to the C-terminus of the DNA-binding polypeptide (e.g., an RGN polypeptide).

[0120] Non-limiting examples of end-to-end base editor fusion proteins include those comprising a deaminase, or an active variant or fragment thereof, fused to the N-terminus of a nickase, or an active variant or fragment thereof, of any one of the end-to-end fusion proteins used in the Examples herein; a) LPG50221 / LPG50148.2.20 (shown as SEQ ID NO: 319) or an active variant or fragment thereof fused to the N-terminus of nAPG07433.1 (shown as SEQ ID NO: 49) or an active variant or fragment thereof, such as the sequence shown as SEQ ID NO: 578 or 582; b) LPG50274 / LPG50148.2.76 (shown as SEQ ID NO: 375) or an active variant or fragment thereof fused to the N-terminus of nAPG07433.1 (shown as SEQ ID NO: 49) or an active variant or fragment thereof, such as the sequence shown as SEQ ID NO: 594 or 577; c) LPG50319 / LPG50148.2.107 (shown as SEQ ID NO: 406) or an active variant or fragment thereof fused to the N-terminus of nAPG07433.1 (shown as SEQ ID NO: 49) or an active variant or fragment thereof, such as the sequence shown as SEQ ID NO: 588 or 616; d) LPG50320 (shown as SEQ ID NO: 596) or an active variant or fragment thereof fused to the N-terminus of nAPG07433.1 (shown as SEQ ID NO: 49) or an active variant or fragment thereof, such as the sequence shown as SEQ ID NO: 593 or 604, and e) LPG50324 (shown as SEQ ID NO: 597) or an active variant or fragment thereof fused to the N-terminus of nAPG07433.1 (shown as SEQ ID NO: 49) or an active variant or fragment thereof, such as the sequence shown as SEQ ID NO: 605.

[0121] In some embodiments, the end-to-end deaminase and the DNA-binding polypeptide (e.g., an RNA-guided DNA-binding polypeptide) are fused to one another via a linker. The term "linker," as used herein, refers to a chemical group or molecule that links two molecules or moieties, such as the binding domain and cleavage domain of a nuclease. In some embodiments, the linker joins an RNA-guided nuclease and a deaminase. In some embodiments, the linker joins an inactive or inactive RGN and a deaminase. In further embodiments, the linker joins two deaminases. In some embodiments, the linker joins an RNA-guided nuclease and a USP. In some embodiments, the linker joins a deaminase and a USP. In certain embodiments, the linker joins an RNA-guided nuclease-deaminase fusion with a USP. Typically, the linker is positioned between or adjacent to two groups, molecules, or other moieties and is connected to each via a covalent bond, thus connecting the two. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 3 to 100 amino acids in length, e.g., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated. In some embodiments, shorter linkers are preferred to reduce the overall size or length of the fusion protein or its coding sequence.

[0122] In some embodiments, the linker between the deaminase and the DNA-binding polypeptide in the end-to-end fusion comprises any one of SEQ ID NOs: 704, 705, and 707-712. Additional suitable linker motifs and linker configurations will be apparent to those of skill in the art. In some embodiments, suitable linker motifs and configurations include those described in Chen et al., 2013 (Adv Drug Deliv Rev. 65(10):1357-69, the contents of which are incorporated herein by reference in their entirety). Additional suitable linker sequences will be apparent to those of skill in the art. In some embodiments, the linker sequence comprises the amino acid sequence set forth as SEQ ID NO: 233 or 234.

[0123] In some embodiments, the general structure of exemplary fusion proteins provided herein comprises the structure: [NH]-[deaminase]-[DBP]-[COOH], [NH]-[DBP]-[deaminase]-[COOH], [NH]-[DBP]-[deaminase]-[deaminase]-[COOH], [NH]-[deaminase]-[DBP]-[deaminase]-[COOH], or [NH]-[deaminase]-[deaminase]-[DBP]-[COOH], where DBP is a DNA-binding polypeptide, NH is the N-terminus of the fusion protein, and COOH is the C-terminus of the fusion protein. In some embodiments, the fusion protein comprises more than two deaminase polypeptides.

[0124] In certain embodiments, exemplary fusion proteins provided herein comprise the general structure: [NH]-[deaminase]-[RGN]-[COOH], [NH]-[RGN]-[deaminase]-[COOH], [NH]-[RGN]-[deaminase]-[deaminase]-[COOH], [NH]-[deaminase]-[RGN]-[deaminase]-[COOH], or [NH]-[deaminase]-[deaminase]-[RGN]-[COOH], where NH is the N-terminus of the fusion protein and COOH is the C-terminus of the fusion protein. In some embodiments, the fusion protein comprises more than two deaminase polypeptides.

[0125] In some embodiments, the fusion protein comprises the structure: [NH2]-[deaminase]-[nuclease-inactive RGN]-[COOH], [NH2]-[deaminase]-[deaminase]-[nuclease-inactive RGN]-[COOH], [NH2]-[nuclease-inactive RGN]-[deaminase]-[COOH], [NH2]-[deaminase]-[nuclease-inactive RGN]-[deaminase]-[COOH], or [NH2]-[nuclease-inactive RGN]-[deaminase]-[deaminase]-[COOH]. It should be understood that "nuclease-inactive RGN" refers to any RGN, including any CRISPR-Cas protein that has been mutated to be nuclease-inactive. In some embodiments, the fusion protein comprises more than two deaminase polypeptides.

[0126] In some embodiments, the fusion protein comprises the structure: [NH2]-[deaminase]-[RGN nickase]-[COOH], [NH2]-[deaminase]-[deaminase]-[RGN nickase]-[COOH], [NH2]-[RGN nickase]-[deaminase]-[COOH], [NH2]-[deaminase]-[RGN nickase]-[deaminase]-[COOH], or [NH2]-[RGN nickase]-[deaminase]-[deaminase]-[COOH]. It should be understood that "RGN nickase" refers to any RGN, including any CRISPR-Cas protein that has been mutated to be active as a nickase.

[0127] In some embodiments, the fusion proteins provided herein comprise the full-length sequence of a deaminase. However, in some embodiments, the fusion proteins provided herein do not comprise the full-length sequence of a deaminase, but only a fragment thereof.

[0128] In some embodiments, a fusion protein of the invention comprises a DNA-binding polypeptide (e.g., RGN) and a deaminase, wherein the deaminase has an amino acid sequence having about 50% to about 100% identity to any of SEQ ID NOs: 6, 300-414, 596-598, and 720-723. In some embodiments, a fusion protein of the invention comprises a DNA-binding polypeptide (e.g., RGN) and a deaminase, wherein the deaminase has an amino acid sequence having about 50% to about 100% identity to any of SEQ ID NOs: 6, 300-414, 596-598, and 720-723, and comprises at least one amino acid residue set forth in any one of Tables 2, 4, 6, and 19. In some embodiments, a fusion protein of the invention comprises a DNA-binding polypeptide (e.g., RGN) and a deaminase, wherein the deaminase has an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to any of SEQ ID NOs: 6, 300-414, 596-598, and 720-723, and comprises at least one of the amino acid residues set forth in any one of Tables 2, 4, 6, and 19. Examples of such fusion proteins are described in the Examples section herein.

[0129] In some embodiments, the fusion protein comprises one deaminase polypeptide. In some embodiments, the fusion protein comprises at least two deaminase polypeptides operably linked directly or via a peptide linker. In some embodiments, the fusion protein comprises one deaminase polypeptide, and a second deaminase polypeptide is co-expressed with the fusion protein.

[0130] In some embodiments, the "-" used in the general structures above indicates the presence of an optional linker sequence. In some embodiments, the fusion proteins provided herein do not include a linker sequence. In some embodiments, at least one of the optional linker sequences is present.

[0131] Other exemplary features that may be present on the deaminases or fusion proteins of the present disclosure are localization sequences, such as a nuclear localization sequence, a cytoplasmic localization sequence, an export sequence such as a nuclear export sequence, or other localization sequences, as well as sequence tags that are useful for solubilizing, purifying, or detecting the fusion protein. Suitable localization signal sequences and protein tag sequences provided herein include, but are not limited to, a biotin carboxylase carrier protein (BCCP) tag, a myc tag, a calmodulin tag, a FLAG tag (e.g., a 3XFLAG tag), a hemagglutinin (HA) tag, a polyhistidine tag, also referred to as a histidine tag or His tag, a maltose-binding protein (MBP) tag, a nus tag, a glutathione-S-transferase (GST) tag, a green fluorescent protein (GFP) tag, a thioredoxin tag, an S tag, a Softag (e.g., Softag 1, Softag 3), a streptotag, a biotin ligase tag, a FlAsH tag, a V5 tag, and an SBP tag. Additional suitable sequences will be apparent to those skilled in the art.

[0132] The compositions and methods of the present disclosure can utilize a deaminase or fusion protein that includes at least one nuclear localization signal (NLS) to enhance transport of the deaminase or fusion protein to the nucleus of a cell. Nuclear localization signals are known in the art and generally include a stretch of basic amino acids (see, e.g., Lange et al., J. Biol. Chem. (2007) 282:5101-5105). In some embodiments, the deaminase or fusion protein includes two, three, four, five, six, or more nuclear localization signals. The nuclear localization signal(s) can be heterologous NLSs. Non-limiting examples of nuclear localization signals useful for the RGN of the present disclosure include the nuclear localization signals of SV40 large T antigen, nucleoplasmin, c-Myc (e.g., Ray et al. (2015) Bioconjug Chem 26(6):1004-7), POLD1 DNA polymerase delta 1, active adenosine deaminase acting on RNA (ADAR2), interacting with RNA polymerase II 1 (Iwr1) (Czeko E., et al., Mol Cell 2011, incorporated herein by reference in its entirety), and OpT (U.S. Publication No. 2020 / 0109382, incorporated herein by reference in its entirety). In embodiments, the RGN comprises an NLS sequence set forth as any one of SEQ ID NOs: 430, 431, and 713-718. The deaminase or fusion protein can comprise one or more NLS sequences at its N-terminus, C-terminus, or both the N-terminus and C-terminus. For example, the deaminase or fusion protein can comprise two NLS sequences in the N-terminal region and four NLS sequences in the C-terminal region. In some embodiments, the deaminase or fusion protein comprises an SV40 NLS (such as the sequence set forth as SEQ ID NO: 430) at its N-terminus and a nucleoplasmin NLS (such as the sequence set forth as SEQ ID NO: 431) at its C-terminus. In some embodiments, the deaminase or fusion protein comprises a c-Myc promoter (such as the sequence set forth as SEQ ID NO: 718) at both its N-terminus and its C-terminus.

[0133] When the NLS is attached to the N-terminus, C-terminus, or both of the deaminase or fusion protein, an NLS linker protein can be present to separate the deaminase or fusion protein from the NLS. Such an NLS linker protein can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more amino acids in length. In some embodiments, the NLS linker protein is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 amino acids in length. In some embodiments, the NLS linker protein between the NLS and the deaminase or fusion protein has the sequence set forth as SEQ ID NO: 706, 719, or 728. In some embodiments, the deaminase or fusion protein comprises a c-Myc promoter at its N-terminus (such as the sequence set forth as SEQ ID NO: 718) separated from the deaminase or fusion protein by an NLS linker protein having the sequence set forth as SEQ ID NO: 728, and a c-Myc promoter at its C-terminus (such as the sequence set forth as SEQ ID NO: 718) separated from the deaminase or fusion protein by an NLS linker protein having the sequence set forth as SEQ ID NO: 728.

[0134] In some embodiments, the compositions and methods of the present disclosure utilize a deaminase or fusion protein that includes at least one cell-penetrating domain that facilitates cellular uptake of the deaminase or fusion protein. Cell-penetrating domains are known in the art and generally comprise a stretch of positively charged amino acid residues (i.e., polycationic cell-penetrating domains), alternating polar and non-polar amino acid residues (i.e., amphipathic cell-penetrating domains), or hydrophobic amino acid residues (i.e., hydrophobic cell-penetrating domains) (see, e.g., Milletti F. (2012) Drug Discov Today 17:850-860). A non-limiting example of a cell-penetrating domain is the trans-activating transcriptional activator (TAT) from human immunodeficiency virus 1.

[0135] The nuclear localization signal and / or cell permeability domain can be located at the N-terminus, C-terminus, and / or internal position of the deaminase or fusion protein.

[0136] VI. Nucleotides encoding deaminases, fusion proteins, and / or gRNAs The present disclosure provides polynucleotides encoding the deaminase, fusion protein, or gRNA of the present disclosure. In some embodiments, the polynucleotide encodes a fusion protein comprising a deaminase and a DNA-binding polypeptide, such as a meganuclease, a zinc finger fusion protein, or a TALEN. The present disclosure further provides polynucleotides encoding fusion proteins comprising a deaminase and an RNA-guided DNA-binding polypeptide (RGDBP). Such an RNA-guided DNA-binding polypeptide can be an RGN or an RGN variant. The protein variant can be a nuclease-inactive or a nickase. The RGN can be a CRISPR-Cas protein, or an active variant or fragment thereof. Examples of CRISPR-Cas nucleases are well known in the art, and similar corresponding mutations can generate mutant variants that are also nickases or are nuclease-inactive.

[0137]

[0010] Embodiments of the present invention provide polynucleotides encoding fusion proteins comprising an RGDBP (e.g., RGN) and a deaminase as described herein. In some embodiments, a second polynucleotide encodes a guide RNA required by the RGDBP to target a nucleotide sequence of interest. In some embodiments, the guide RNA and the fusion protein are encoded by the same polynucleotide.

[0138] In some embodiments, the guide RNA and the fusion protein are encoded by the same polynucleotide, while in other embodiments, the guide RNA and the fusion protein are encoded by two separate polynucleotides.

[0139] The use of the term "polynucleotide" is not intended to limit the present disclosure to polynucleotides comprising DNA, although such DNA polynucleotides are contemplated. Those skilled in the art will recognize that polynucleotides can include ribonucleotides (RNA) (e.g., mRNA) and combinations of ribonucleotides and deoxyribonucleotides. Such deoxyribonucleotides and ribonucleotides include both naturally occurring molecules and synthetic analogues. Polynucleotides disclosed herein also encompass all forms of sequence, including, but not limited to, single-stranded forms, double-stranded forms, stem and loop structures, circular forms (including, e.g., circular RNA), and the like.

[0140] An embodiment of the present invention is a nucleic acid molecule comprising a sequence encoding a deaminase having about 50% to about 100% identity to any one of SEQ ID NOs: 5, 6, 300 to 414, 596 to 598, and 720 to 723. An embodiment of the present invention is a nucleic acid molecule comprising a sequence encoding a deaminase having about 50% to about 100% identity to any one of SEQ ID NOs: 5, 6, 300 to 414, 596 to 598, and 720 to 723, and having at least one amino acid residue set forth in any one of Tables 2, 4, 6, and 19. An embodiment of the present invention is a nucleic acid molecule comprising a sequence encoding a deaminase having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to any of SEQ ID NOs: 5, 6, 300-414, 596-598, and 720-723, and having at least one of the amino acid residues set forth in any one of Tables 2, 4, 6, and 19, wherein the nucleic acid molecule encodes a deaminase having deaminase activity. The nucleic acid molecule may further comprise a heterologous promoter or terminator. The nucleic acid molecule may encode a fusion protein, wherein the encoded deaminase is operably linked to a DNA-binding polypeptide and, optionally, a second deaminase. In some embodiments, the nucleic acid molecule encodes a fusion protein, wherein the encoded deaminase is operably linked to RGN and, optionally, a second deaminase.

[0141] In some embodiments, nucleic acid molecules comprising polynucleotides encoding deaminases or fusion proteins of the invention are codon-optimized for expression in a target organism. A "codon-optimized" coding sequence is a polynucleotide coding sequence having a frequency of codon usage designed to mimic the frequency of preferred codon usage or transcription conditions of a particular host cell. Expression in a particular host cell or organism is enhanced as a result of modifying one or more codons at the nucleic acid level such that the translated amino acid sequence is unchanged. Nucleic acid molecules can be codon-optimized either in whole or in part. Codon tables and other references providing preference information for a wide range of organisms are available in the art (e.g., see Campbell and Gowri (1990) Plant Physiol. 92:1-11 for a discussion of plant-preferred codon usage). Methods for synthesizing plant-preferred genes are available in the art. See, for example, US Pat. Nos. 5,380,831 and 5,436,391, and Murray et al. (1989) Nucleic Acids Res. 17:477-498, which are incorporated herein by reference.

[0142] In some embodiments, polynucleotides encoding the deaminases, fusion proteins, and / or gRNAs described herein are provided in expression cassettes for in vitro expression or expression in a cell, organelle, embryo, or organism of interest. The cassette may include 5' and 3' regulatory sequences operably linked to the polynucleotides encoding the deaminases, fusion proteins, and / or gRNAs provided herein, allowing for expression of the polynucleotides. The cassette may further include at least one additional gene or genetic element cotransformed into the organism. When additional genes or elements are included, the components are operably linked. The term "operably linked" is intended to mean a functional linkage between two or more elements. For example, an operable linkage between a promoter and a coding region of interest (e.g., a region encoding a deaminase, RGN, and / or gRNA) is a functional linkage that allows for expression of the coding region of interest. Operably linked elements may be contiguous or non-contiguous. When used to refer to the junction of two protein-coding regions, operably linked means that the coding regions are in the same reading frame. In some embodiments, the additional gene(s) or element(s) are provided on multiple expression cassettes. For example, a nucleotide sequence encoding a deaminase or fusion protein of the present disclosure may be present on one expression cassette, while a nucleotide sequence encoding a gRNA may be present on a separate expression cassette. Another example may have a nucleotide sequence encoding a deaminase of the present disclosure alone on a first expression cassette, a second expression cassette encoding a fusion protein comprising the deaminase, and a nucleotide sequence encoding a gRNA on a third expression cassette. Such expression cassettes are provided with multiple restriction and / or recombination sites for inserting polynucleotides under the transcriptional control of regulatory regions. An expression cassette containing a selectable marker gene may also be present.

[0143] An expression cassette can include, in the 5'-3' direction of transcription, a transcriptional (and in some embodiments, translational) initiation region (i.e., promoter), a polynucleotide encoding a deaminase of the present invention, a polynucleotide encoding a fusion protein of the present invention, and a transcriptional (and in some embodiments, translational) termination region (i.e., termination region) functional in the organism of interest. Promoters useful in the present invention are capable of directing or driving expression of a coding sequence in a host cell. Regulatory regions (e.g., promoters, transcriptional regulatory regions, and translational termination regions) can be endogenous or heterologous to the host cell or to each other. As used herein, "heterologous" with respect to a sequence refers to a sequence that originates from a foreign species or, if from the same species, has been substantially modified in composition and / or genomic locus from its native form by deliberate human intervention. As used herein, a chimeric gene comprises a coding sequence operably linked to a transcriptional initiation region that is heterologous to the coding sequence.

[0144] Convenient termination regions are available from the Ti plasmid of A. tumefaciens, such as the octopine synthase and nopaline synthase termination regions. Also, Guerineau et al. (1991) Mol.Gen.Genet.262:141-144, Proudfoot (1991) Cell 64:671-674, Sanfacon et al. (1991) Genes Dev.5:141-149, Mogen et al. (1990) Plant Cell 2:1261-1272, Munroe et al. (1990) Gene 91:151-158, Ballas et al. (1989) Nucleic Acids Res. 17:7891-7903, and Joshi et al. (1987) Nucleic Acids Res. 15:9627-9639.

[0145] Additional regulatory signals include, but are not limited to, transcription initiation start sites, operators, activators, enhancers, other regulatory elements, ribosome binding sites, start codons, termination signals, etc. See, e.g., U.S. Patent Nos. 5,039,523 and 4,853,331, EPO 0480762A2, Sambrook et al. (1992) Molecular Cloning: A Laboratory Manual, ed. Maniatis et al. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY), hereinafter "Sambrook 11," Davis et al., eds. (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, NY, and references cited therein.

[0146] In preparing expression cassettes, various DNA fragments may be manipulated to provide DNA sequences in the proper orientation, and, if necessary, in the proper reading frame. To this end, adapters or linkers may be employed to join DNA fragments, or other manipulations may be involved to provide convenient restriction sites, remove unnecessary DNA, remove restriction sites, etc. To this end, in vitro mutagenesis, primer repair, restriction, annealing, resubstitutions, e.g., transitions and transversions, may be involved.

[0147] Many promoters can be used in the practice of the present invention. The promoter can be selected based on the desired result. The nucleic acid can be combined with a constitutive, inducible, developmental stage-specific, cell type-specific, tissue-preferred, tissue-specific, or other promoter for expression in the organism of interest. See, for example, WO 99 / 43838 and the promoters described in U.S. Patent Nos. 8,575,425, 7,790,846, 8,147,856, 8,586832, 7,772,369, 7,534,939, 6,072,050, 5,659,026, 5,608,149, 5,608,144, 5,604,121, 5,569,597, 5,466,785, 5,399,680, 5,268,463, 5,608,142, and 6,177,611.

[0148] For expression in plants, constitutive promoters include the CaMV 35S promoter (Odell et al. (1985) Nature 313:810-812), rice actin (McElroy et al. (1990) Plant Cell 2:163-171), ubiquitin (Christensen et al. (1989) Plant Mol. Biol. 12:619-632 and Christensen et al. (1992) Plant Mol. Biol. 18:675-689), pEMU (Last et al. (1991) Theor. Appl. Genet. 81:581-588), and MAS (Velten et al. (1984) EMBO J. 3:2723-2730).

[0149] Examples of inducible promoters are the Adh1 promoter, which is inducible by hypoxia or cold stress, the Hsp70 promoter, which is inducible by heat stress, the PPDK promoter and the peptocarboxylase promoter, both of which are inducible by light, the safener-induced In2-2 promoter (U.S. Pat. No. 5,364,780), the oxine-induced, tapetum-specific Axig1 promoter (PCT US01 / 22169), which is also active in callus, steroid-responsive promoters (e.g., the estrogen-induced ERE promoter and glucocorticoid-inducible promoters in Schena et al. (1991) Proc. Natl. Acad. Sci. USA 88:10421-10425 and McNellis et al. (1998) Plant J. 14(2):247-257), and tetracycline-inducible and tetracycline-repressible promoters (e.g., Gatz et al. (1998) Plant J. 14(2):247-257), which are also incorporated herein by reference. Chemically inducible promoters such as those described by G. et al. (1991) Mol. Gen. Genet. 227:229-237, and US Pat. Nos. 5,814,618 and 5,789,156, are also useful.

[0150] In some embodiments, tissue-specific or tissue-preferred promoters are utilized to target expression of an expression construct in a specific tissue. In certain embodiments, tissue-specific or tissue-preferred promoters are active in plant tissues. Examples of promoters under developmental control in plants include promoters that preferentially initiate transcription in certain tissues, such as leaves, roots, fruits, seeds, or flowers. A "tissue-specific" promoter is a promoter that initiates transcription only in certain tissues. Unlike constitutive expression of a gene, tissue-specific expression is the result of several interacting levels of gene regulation. Therefore, promoters from homologous or closely related plant species may be preferred to achieve efficient and reliable expression of a transgene in a specific tissue. In some embodiments, expression involves a tissue-preferred promoter. A "tissue-preferred" promoter is a promoter that preferentially initiates transcription in a specific tissue, but not necessarily entirely or exclusively.

[0151] In some embodiments, a nucleic acid molecule encoding a deaminase or fusion protein described herein comprises a cell-type-specific promoter. A "cell-type-specific" promoter is a promoter that drives expression primarily in a particular cell type in one or more organs. Some examples of plant cells in which a functional cell-type-specific promoter may be primarily active in plants include, for example, BETL cells, root vascular cells, leaves, stem cells, and stem cells. A nucleic acid molecule can also comprise a cell-type-preferred promoter. A "cell-type-preferred" promoter is a promoter that primarily, but not necessarily entirely or exclusively, drives expression in a particular cell type in one or more organs. Some examples of plant cells in which a functional cell-type-preferred promoter may be primarily active in plants include, for example, BETL cells, root vascular cells, leaves, stem cells, and stem cells.

[0152] In some embodiments, the nucleic acid sequence encoding the deaminase, fusion protein, and / or gRNA is operably linked to a promoter sequence recognized by a phage RNA polymerase, e.g., for in vitro mRNA synthesis. In such embodiments, the in vitro transcribed RNA can be purified for use in the methods described herein. For example, the promoter sequence can be a T7, T3, or SP6 promoter sequence, or a variation of the T7, T3, or SP6 promoter sequence. In such embodiments, the expressed protein and / or RNA can be purified for use in the genome modification methods described herein.

[0153] In certain embodiments, the polynucleotide encoding the deaminase, fusion protein, and / or gRNA is linked to a polyadenylation signal (e.g., the SV40 polyA signal and other signals functional in plants) and / or at least one transcription termination sequence. In some embodiments, the sequence encoding the deaminase or fusion protein is linked to sequence(s) encoding at least one nuclear localization signal, at least one cell-penetrating domain, and / or at least one signal peptide capable of transporting the protein to a specific subcellular location, as described elsewhere herein.

[0154] In some embodiments, the polynucleotide encoding the deaminase, fusion protein, and / or gRNA is present in a vector or vectors. A "vector" refers to a polynucleotide composition for transferring, delivering, or introducing a nucleic acid into a host cell. Suitable vectors include plasmid vectors, phagemids, cosmids, artificial / minichromosomes, transposons, and viral vectors (e.g., lentiviral vectors, adeno-associated viral vectors, baculoviral vectors). In some embodiments, the vector comprises additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences), selectable marker sequences (e.g., antibiotic resistance genes), origins of replication, etc. Additional information can be found in "Current Protocols in Molecular Biology" Ausubel et al., John Wiley & Sons, New York, 2003, or "Molecular Cloning: A Laboratory Manual" Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3111. rd edition, 2001.

[0155] In some embodiments, the vector contains a selectable marker gene for the selection of transformed cells. The selectable marker gene is utilized for the selection of transformed cells or tissues. Marker genes include genes encoding antibiotic resistance, such as those encoding neomycin phosphotransferase II (NEO) and hygromycin phosphotransferase (HPT), and genes conferring resistance to herbicidal compounds, such as glufosinate ammonium, bromoxynil, imidazolinone, and 2,4-dichlorophenoxyacetate (2,4-D).

[0156] In some embodiments, the expression cassette or vector comprising the sequence encoding the fusion protein further comprises a sequence encoding a gRNA. In some embodiments, the sequence(s) encoding the gRNA are operably linked to at least one transcriptional control sequence for expression of the gRNA in an organism or host cell of interest. For example, the polynucleotide encoding the gRNA can be operably linked to a promoter sequence recognized by RNA polymerase III (Pol III). Examples of suitable Pol III promoters include, but are not limited to, mammalian U6, U3, H1, and 7SL RNA promoters, rice U6 and U3 promoters, and the promoters disclosed in PCT International Application No. PCT / US2022 / 032940, filed June 10, 2022, which is incorporated herein by reference in its entirety.

[0157] As shown, an organism of interest can be transformed using an expression construct containing a nucleotide sequence encoding a deaminase, a fusion protein, and / or a gRNA. Methods for transformation involve introducing a nucleotide construct into the organism of interest. By "introducing," we mean introducing the nucleotide construct into a host cell so that the construct gains access to the interior of the host cell. The methods of the present invention do not require a specific method for introducing the nucleotide construct into a host organism, so long as the nucleotide construct gains access to the interior of at least one cell of the host organism. In some embodiments, an mRNA encoding the deaminase or fusion protein is introduced into the host cell. In some embodiments, where the fusion protein comprises an RGN, an mRNA encoding the fusion protein is introduced into the cell and a gRNA is introduced into the cell. The host cell can be a eukaryotic or prokaryotic cell. In some embodiments, the eukaryotic host cell is a plant cell, a mammalian cell, or an insect cell. Methods for introducing nucleotide constructs into plant and other host cells are known in the art and include, but are not limited to, stable transformation methods, transient transformation methods, and virus-mediated methods.

[0158] The method results in transformed organisms, including plants, including whole plants, as well as plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, propagules, embryos, and their progeny. Plant cells can be differentiated or undifferentiated (e.g., callus, suspension culture cells, protoplasts, leaf cells, root cells, phloem cells, pollen).

[0159] A "transgenic organism" or "transformed organism" or "stably transformed" organism, cell, or tissue refers to an organism that incorporates or integrates a polynucleotide encoding the deaminase or fusion protein of the present invention. It is recognized that other exogenous or endogenous nucleic acid sequences or DNA fragments can also be incorporated into host cells. Agrobacterium- and biolistic transformation remain the two predominant approaches for transforming plant cells. However, host cell transformation can also be carried out by infection, transfection, microinjection, electroporation, microprojection, biolistic or particle bombardment, electroporation, silica / carbon fiber, ultrasound-mediated, PEG-mediated, calcium phosphate co-precipitation, polycation DMSO technology, DEAE-dextran procedures, as well as viral, liposome-mediated, and the like. Viral-mediated transfer of polynucleotides encoding deaminases, fusion proteins, and / or gRNAs includes retrovirus-, lentivirus-, adenovirus-, and adeno-associated virus-mediated transfer and expression, as well as the use of caulimoviruses (e.g., cauliflower mosaic virus), geminiviruses (e.g., bean golden yellow mosaic virus or corn streak virus), and RNA plant viruses (e.g., tobacco mosaic virus).

[0160] Transformation protocols and protocols for introducing polypeptide or polynucleotide sequences into plants can vary depending on the type of host cell (e.g., monocotyledonous or dicotyledonous plant cell) targeted for transformation. Methods for transformation are known in the art and include those described in U.S. Patent Nos. 8,575,425, 7,692,068, 8,802,934, and 7,541,517, each of which is incorporated herein by reference. Also, Rakoczy-Trojanowska, M. (2002) Cell Mol Biol Lett.7:849-858, Jones et al. (2005) Plant Methods 1:5, Rivera et al. (2012) Physics of Life Reviews 9:308-345, Bartlett et al. (2008) Plant Methods 4:1-12, Bates, GW (1999) Methods in Molecular Biology 111:359-366, Binns and Thomashow (1988) Annual Reviews in Microbiology 42:575-606, Christou, P. (1992) The Plant Journal 2:275-281, Christou, P. (1995) Euphytica 85:13-27, Tzfira et. See also Yao et al. (2004) TRENDS in Genetics 20:375-383, Yao et al. (2006) Journal of Experimental Botany 57:3737-3746, Zupan and Zambryski (1995) Plant Physiology 107:1041-1047, Jones et al. (2005) Plant Methods 1:5.

[0161] Transformation can result in stable or transient integration of a nucleic acid into a cell. "Stable transformation" is intended to mean that a nucleotide construct introduced into a host cell is integrated into the genome of the host cell and is capable of being inherited by its progeny. "Transient transformation" is intended to mean that a polynucleotide is introduced into a host cell without being integrated into the genome of the host cell.

[0162] Methods for chloroplast transformation are known in the art. See, e.g., Svab et al. (1990) Proc. Natl. Acad. Sci. USA 87:8526-8530; Svab and Maliga (1993) Proc. Natl. Acad. Sci. USA 90:913-917; Svab and Maliga (1993) EMBO J. 12:601-606. This method relies on particle gun delivery of DNA containing a selectable marker and targeting the DNA to the plastid genome through homologous recombination. Additionally, plastid transformation can be achieved by transactivation of a silent plastid-derived transgene through tissue-preferential expression of a nuclear-encoded, plastid-directed RNA polymerase. Such a system is reported in McBride et al. (1994) Proc. Natl. Acad. Sci. USA 91:7301-7305.

[0163] The transformed cells can be grown into transgenic organisms, such as plants, according to conventional methods. See, for example, McCormick et al. (1986) Plant Cell Reports 5:81-84. These plants can then be grown and pollinated with either the same transformed strain or a different strain, and the resulting hybrids carrying the deaminase or fusion protein polynucleotide can be identified. To ensure that the deaminase or fusion protein polynucleotide is stably maintained and inherited, two or more generations can be grown, and then the seeds can be harvested to ensure the presence of the deaminase or fusion protein polynucleotide. In this way, the present invention provides transformed seeds (also referred to as "transgenic seeds") carrying the nucleotide construct of the present invention, e.g., the expression cassette of the present invention, stably integrated into their genome.

[0164] In some embodiments, transformed cells are introduced into an organism. These cells may be derived from an organism, and the cells are transformed using an ex vivo approach.

[0165] The sequences provided herein can be used for the transformation of any plant species, including, but not limited to, monocotyledons or dicotyledons. Examples of plants of interest include, but are not limited to, corn (maize), sorghum, wheat, sunflower, tomato, crucifers, pepper, potato, cotton, rice, soybean, sugar beet, sugarcane, tobacco, barley and rapeseed, Brassica, alfalfa, rye, millet, safflower, peanut, sweet potato, cassava, coffee, coconut, pineapple, citrus trees, cocoa, tea, banana, avocado, fig, guava, mango, olive, papaya, cashew, macadamia, almond, oat, vegetables, ornamentals, and conifers.

[0166] Vegetables include, but are not limited to, tomatoes, lettuce, green beans, lima beans, peas, and members of the genus Curcumis, such as cucumbers, cantaloupes, and muskmelons. Ornamental plants include, but are not limited to, azaleas, hydrangeas, hibiscus, roses, tulips, daffodils, petunias, carnations, poinsettias, and chrysanthemums. Preferably, the plants of the present invention are crop plants (e.g., corn, sorghum, wheat, sunflowers, tomatoes, crucifers, peppers, potatoes, cotton, rice, soybeans, sugar beets, sugarcane, tobacco, barley, rapeseed, etc.).

[0167] As used herein, the term "plant" includes intact plant cells, plant protoplasts, plant cell tissue cultures capable of regenerating plants, plant callus, plant mass, and intact plant cells in plants or parts of plants such as embryos, pollen, ovules, seeds, leaves, flowers, branches, fruit, grains, ears, cobs, husks, stems, roots, root tips, anthers, etc. Grain is intended to mean mature seeds produced by commercial growers for purposes other than seed growth or reproduction. Progeny, variants, and mutants of regenerated plants are also included within the scope of the present invention, provided that these parts contain the introduced polynucleotide. Further provided are processed plant products or by-products carrying the sequences disclosed herein, including, for example, soybean meal.

[0168] In some embodiments, polynucleotides encoding deaminases, fusion proteins, and / or gRNAs are used to transform any eukaryotic species, including, but not limited to, animals (e.g., mammals, insects, fish, birds, and reptiles), fungi, amoebae, algae, and yeast. In some embodiments, polynucleotides encoding deaminases, fusion proteins, and / or gRNAs are used to transform any prokaryotic species, including, but not limited to, archaea and bacteria (e.g., Bacillus spp., Klebsiella spp., Streptomyces spp., Rhizobium spp., Escherichia spp., Pseudomonas spp., Salmonella spp., Shigella spp., Vibrio spp., Yersinia spp., Mycoplasma spp., Agrobacterium spp., and Lactobacillus spp.).

[0169] In some embodiments, conventional viral and non-viral gene transfer methods are used to introduce nucleic acids into mammalian cells or target tissues. Using such methods, nucleic acids encoding the deaminase or fusion protein of the present invention and, optionally, a gRNA can be administered to cells in culture or in a host organism. Non-viral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of the vectors described herein), naked nucleic acid, and nucleic acid complexed with a delivery vehicle such as a liposome. Viral vector delivery systems include DNA and RNA viruses that have either episomal or integrated genomes after delivery to cells. Non-limiting examples include vectors utilizing caulimoviruses (e.g., cauliflower mosaic virus), geminiviruses (e.g., bean golden yellow mosaic virus or corn streak virus), and RNA plant viruses (e.g., tobacco mosaic virus). For reviews of gene therapy procedures, see Anderson, Science 256:808-813 (1992); Nabel & Feigner, TIBTECH 11:211-217 (1993); Mitani & Caskey, TIBTECH 11:162-166 (1993); Dillon, TIBTECH 11:167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10):1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience 8:35-36 (1995); Kremer & Perricaudet, British Medical Bulletin 51(1):31-44 (1995); Haddada et al., in Current Topics in See Microbiology and Immunology, Doerfler and Bohm (eds) (1995), and Yu et al., Gene Therapy 1:13-26 (1994).

[0170] Non-viral methods for nucleic acid delivery include lipofection, Agrobacterium-mediated transformation, nucleofection, microinjection, gene guns, virosomes, liposomes, immunoliposomes, polycation or lipid:nucleic acid conjugates, naked DNA, artificial virions, and drug-enhanced uptake of DNA. Lipofection is described, for example, in U.S. Pat. Nos. 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam™ and Lipofectin™). Cationic and neutral lipids suitable for efficient receptor-recognition lipofection of polynucleotides include those described in Feigner, WO91 / 17424, and WO91 / 16024. Delivery can be to cells (in vitro or ex vivo administration) or target tissues (e.g., in vivo administration). The preparation of lipid:nucleic acid complexes, including targeted liposomes such as immunolipid complexes, is well known to those skilled in the art (e.g., Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Ther. 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad et al., Cancer Res. 52:4817-4820 (1992), U.S. Patent Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787).

[0171] The use of RNA or DNA virus-based systems for nucleic acid delivery takes advantage of highly evolved processes for targeting viruses to specific cells in the body and transporting the viral payload to the nucleus. Viral vectors can be administered directly to patients (in vivo) or used to treat cells in vitro, with the modified cells optionally administered to patients (ex vivo). Traditional virus-based systems can include retroviral, lentiviral, adenoviral, adeno-associated viral, and herpes simplex viral vectors for gene transfer. Integration into the host genome is possible with retroviral, lentiviral, and adeno-associated viral gene transfer methods, often resulting in long-term expression of the inserted transgene. Additionally, high transduction efficiencies have been observed in many different cell types and target tissues.

[0172] Retroviral tropism can be modified by incorporating foreign envelope proteins, expanding the potential target population of target cells. Lentiviral vectors are retroviral vectors capable of transducing or infecting non-dividing cells and typically produce high viral titers. Therefore, the choice of retroviral gene transfer system will depend on the target tissue. Retroviral vectors consist of cis-acting long terminal repeats capable of packaging up to 6-10 kb of foreign sequence. Minimal cis-acting LTRs are sufficient for vector replication and packaging, which are then used to integrate the desired gene into target cells and provide persistent transgene expression. Widely used retroviral vectors include those based on murine leukemia virus (MuLV), gibbon ape leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (see, e.g., Buchscher et al., J. Virol. 66:2731-2739 (1992); Johann et al., J. Virol. 66:1635-1640 (1992); Sommnerfelt et al., Virol. 176:58-59 (1990); Wilson et al., J. Virol. 63:2374-2378 (1989); Miller et al., J. Virol. 65:2220-2224 (1991); PCT / US94 / 05700).

[0173] For applications where transient expression is preferred, an adenovirus-based system can be used. Adenovirus-based vectors are capable of very high transduction efficiency in many cell types and do not require cell division. High titers and expression levels have been obtained using such vectors. This vector can be produced in large quantities in a relatively simple system. Adeno-associated virus ("AAV") vectors can also be used to transduce cells with target nucleic acids, for example, in the in vitro production of nucleic acids and peptides, and in in vivo and ex vivo gene therapy procedures (see, e.g., West et al., Virology 160:38-47 (1987); U.S. Pat. No. 4,797,368; WO 93 / 24641; Katin, Human Gene Therapy 5:793-801 (1994); Muzyczka, J. Clin. Invest. 94:1351 (1994). Construction of recombinant AAV vectors is described in U.S. Pat. No. 5,173,414; Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985); Tratschin, et al. The process has been described in many publications, including Hermonat & Muzyczka, PNAS 81:6466-6470 (1984), and Samulski et al., J. Virol. 63:03822-3828 (1989). Packaging cells are typically used to form viral particles capable of infecting host cells. Such cells include 293 cells, which package adenovirus, and ψJ2 or PA317 cells, which package retrovirus.

[0174] Viral vectors used in gene therapy are usually produced by generating cell lines that package nucleic acid vectors into viral particles. The vectors typically contain minimal viral sequences necessary for packaging and subsequent integration into the host; other viral sequences are replaced by an expression cassette for the polynucleotide(s) to be expressed. Missing viral functions are typically supplied in trans by the packaging cell line. For example, AAV vectors used in gene therapy typically contain only the ITR sequences from the AAV genome necessary for packaging and integration into the host genome. Viral DNA is packaged in a cell line containing a helper plasmid encoding other AAV genes, i.e., rep and cap, but lacking the ITR sequences.

[0175] The cell line can be infected with adenovirus as a helper. The helper virus promotes the replication of AAV vectors and the expression of AAV genes from the helper plasmid. The helper plasmid is not packaged in significant amounts due to the lack of ITR sequences. Contamination by adenovirus can be reduced, for example, by heat treatment, to which adenovirus is more sensitive than AAV. Additional methods for delivering nucleic acids to cells are known to those skilled in the art. See, for example, US2003 / 0087817 (incorporated herein by reference).

[0176] Ideally, the coding sequence of a fusion protein of the present invention and the corresponding guide RNA for targeting the fusion protein can all be packaged into a single AAV vector. The generally accepted size limit for AAV vectors is 4.7 kb, but larger sizes may be contemplated at the expense of reduced packing efficiency. To ensure that the expression cassettes for both the fusion protein and its corresponding guide RNA can fit into an AAV vector, an activity-deleted variant of RGN may be used. In addition to shortening the amino acid sequence, and thus the coding sequence of the RGN and / or deaminase of the fusion protein, the peptide linker connecting the RGN and deaminase may also be shortened or removed. Finally, genetic elements such as promoters, enhancers, and / or terminators may be manipulated through deletion analysis to determine the minimum size required for each to function. The present invention also teaches methods of using such fusion proteins for targeted base editing via in vivo AAV vector delivery.

[0177] In some embodiments, host cells are transiently or non-transiently transfected with one or more vectors described herein. In some embodiments, cells are transfected as they naturally occur in a subject. In some embodiments, transfected cells are obtained from a subject.

[0178] In some embodiments, the transfected cell is a eukaryotic cell. In some embodiments, the eukaryotic cell is an animal cell (e.g., mammalian, insect, fish, bird, and reptile). In some embodiments, the transfected cell is a human cell. In some embodiments, the transfected cell is a cell of hematopoietic origin, such as an immune cell (i.e., a cell of the innate or adaptive immune system), including, but not limited to, a B cell, a T cell, a natural killer (NK) cell, a pluripotent stem cell, an induced pluripotent stem cell, a chimeric antigen receptor T (CAR-T) cell, a monocyte, a macrophage, and a dendritic cell. The transfected cell can be an allogeneic cell (e.g., an allogeneic T cell) or an autologous cell (e.g., an autologous T cell).

[0179] In some embodiments, the cells are derived from cells taken from a subject, such as a cell line. In some embodiments, the cells or cell line are prokaryotic. In some embodiments, the cells or cell line are eukaryotic. In further embodiments, the cells or cell line are derived from insect, avian, plant, or fungal species. In some embodiments, the cells or cell line may be mammalian, such as, for example, human, monkey, mouse, cow, pig, goat, hamster, rat, cat, or dog. A wide variety of cell lines for tissue culture are known in the art. Examples of cell lines include C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLaS3, Huhl, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Panel, PC-3, TFl, CTLL-2, CIR, Rat6, CVI, RPTE, AlO, T24, 182, A37 5, ARH-77, Calul, SW480, SW620, SKOV3, SK-UT, CaCo2, P388Dl, SEM-K2, WEHI-231, HB56, TIB 55, lurkat, 145.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, Hep G2, HeLa B, HeLa T4.COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelium, BALB / 3T3 mouse embryonic fibroblasts, 3T3 Swiss, 3T3-Ll, 132-d5 human fetal fibroblasts, 10.1 mouse fibroblasts, 293-T, 3T3, 721, 9L, A2780, A2780ADR, A2 780cis, A172, A20, A253, A431, A-549, ALC, B16, B35, BCP-I cells, BEAS-2B, bEnd.3, BHK-21, BR 293, BxPC3, C3H-10Tl / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-Kl, CHO-K2, CHO-T, CHO Dhfr- / -, COR-L23, COR-L23 / CPR, COR-L235010, CORL23 / R23, COS-7, COV-434, CML Tl, CMT, CT26, D17, DH82, DU145, DuCaP, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA 2, HEK-293, HeLa, Hepalclc7, HL-60, HMEC, HT-29, lurkat, lY cells, K562 cells, Ku812, KCL22, KGl, KYOl, LNCap, Ma-Mel 1-48, MC-38, MCF-7, MCF-l0A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCKII, MDCKII, MOR / 0.2R, MONO-MAC 6, MTD-lA, MyEnd, NCI-H69 / CPR, NCI-H69 / LX10, NCI-H69 / LX20, NCI-H69 / LX4, NIH-3T3, NALM-1, NW-145, OPCN / OPCT cell line, Peer, PNT-lA / PNT Examples of suitable cell lines include, but are not limited to, 2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THP1 cell lines, U373, U87, U937, VcaP, Vero cells, WM39, WT-49, X63, YAC-1, YAR, and transgenic varieties thereof. Cell lines are available from a variety of sources known to those skilled in the art (see, for example, the American Type Culture Collection (ATCC), Manassas, Va.).

[0180] In some embodiments, cells transfected with one or more vectors described herein are used to establish new cell lines containing one or more vector-derived sequences. In some embodiments, cells transiently transfected with a fusion protein of the invention and optionally a gRNA, or a ribonucleoprotein complex of the invention, and modified through the activity of the fusion protein or ribonucleoprotein complex, are used to establish new cell lines, including cells containing the modification but lacking any other exogenous sequences. In some embodiments, cells transiently or non-transiently transfected with one or more vectors described herein, or cell lines derived from such cells, are used in evaluating one or more test compounds.

[0181] In some embodiments, one or more vectors described herein are used to produce non-human transgenic animals or transgenic plants. In some embodiments, the transgenic animal is an insect. In further embodiments, the insect is an insect such as a mosquito or a tick. In some embodiments, the insect is a plant pest such as the corn rootworm or fall armyworm. In some embodiments, the transgenic animal is a bird such as a chicken, turkey, goose, or duck. In some embodiments, the transgenic animal is a mammal such as a human, mouse, rat, hamster, monkey, ape, rabbit, pig, cow, horse, goat, sheep, cat, or dog.

[0182] VII. Polypeptide and Polynucleotide Variants and Fragments The present disclosure provides deaminases and fusion proteins comprising the same, as well as fusion proteins comprising RGN and a heterologous polypeptide (e.g., deaminase) active on DNA molecules. RGN can include any one of SEQ ID NOS: 1-4, 49-162, 435, 575, 576, 698, and 699, active variants or fragments thereof, and polynucleotides encoding same. Deaminases can include any one of SEQ ID NOS: 6, 229-414, 596, 597, and 720-723, active variants or fragments thereof, and polynucleotides encoding same.

[0183] The activity of the variants or fragments may be altered compared to the polynucleotide or polypeptide of interest, but the variants and fragments should retain the function of the polynucleotide or polypeptide of interest. For example, the variants or fragments may have increased activity, decreased activity, a different spectrum of activity, or any other alteration in activity compared to the polynucleotide or polypeptide of interest.

[0184] Fragments and variants of the deaminases of the invention that have adenine deaminase activity retain that activity when they are part of a fusion protein that further comprises a DNA-binding polypeptide or fragment thereof.

[0185] Fragments and variants of the fusion proteins of the invention, such as fusion protein fragments and variants comprising a fragment or variant of a deaminase or a fragment or variant of a DNA-binding polypeptide (e.g., RGN), that have base editing activity, retain that activity.

[0186] The term "fragment" refers to a portion of a polynucleotide or polypeptide sequence of the invention. A "fragment" or "biologically active portion" includes a polynucleotide comprising a sufficient number of consecutive nucleotides to retain biological activity. A "fragment" or "biologically active portion" includes a polypeptide comprising a sufficient number of consecutive amino acid residues to retain biological activity. Fragments of RGN or deaminase disclosed herein include those shorter than the full-length sequence due to the use of alternative downstream start sites. In some embodiments, a biologically active portion of a deaminase or RGN is a polypeptide comprising, for example, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, or more consecutive amino acid residues of any of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699, or a variant thereof. In some embodiments, a biologically active portion of a deaminase or RGN is a polypeptide comprising 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, or more consecutive amino acid residues of any of SEQ ID NOs: 5, 6, 229-414, 596, 597, and 720-723, or a variant thereof. Such biologically active portions can be prepared by recombinant techniques and evaluated for activity.

[0187] In general, "variant" is intended to mean a substantially similar sequence. With respect to polynucleotides, variants include deletions and / or additions of one or more nucleotides at one or more internal sites within the naturally occurring polynucleotide and / or substitutions of one or more nucleotides at one or more sites within the naturally occurring polynucleotide. As used herein, a "native" or "wild-type" polynucleotide or polypeptide includes a naturally occurring nucleotide sequence or amino acid sequence, respectively. With respect to polynucleotides, conservative variants include sequences that, due to the degeneracy of the genetic code, encode the native amino acid sequence of a gene of interest. Naturally occurring allelic variants such as these can be identified using well-known molecular biology techniques, e.g., using polymerase chain reaction (PCR) and hybridization techniques, as outlined below. Variant polynucleotides also include synthetically derived polynucleotides, such as those generated using site-directed mutagenesis, but still encoding a polypeptide or polynucleotide of interest. Variants of a particular polynucleotide disclosed herein have from about 40% to about 99% or more sequence identity with the particular polynucleotide, as determined by sequence alignment programs and parameters described elsewhere herein. Generally, variants of a particular polynucleotide disclosed herein have at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or more sequence identity with the particular polynucleotide, as determined by sequence alignment programs and parameters described elsewhere herein.

[0188] Variants of a particular polynucleotide (i.e., a reference polynucleotide) disclosed herein can also be assessed by comparing the percent sequence identity between the polypeptide encoded by the variant polynucleotide and the polypeptide encoded by the reference polynucleotide. The percent sequence identity between any two polypeptides can be calculated using sequence alignment programs and parameters described elsewhere herein. When any given pair of polynucleotides disclosed herein is assessed by comparing the percent sequence identity shared by the two polypeptides they encode, the percent sequence identity between the two encoded polypeptides will be from about 40% to about 99% or more, and in some embodiments, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or more sequence identity.

[0189] In some embodiments, a polynucleotide of the present disclosure encodes a deaminase comprising an amino acid sequence having at least about 40% to about 99% identity to the amino acid sequence of any of SEQ ID NOs: 5, 6, 300-414, 596-598, and 720-723, wherein the deaminase comprises at least one of the amino acid residues set forth in any one of Tables 2, 4, 6, and 19. In some embodiments, a polynucleotide of the disclosure encodes a deaminase comprising an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or more identity to the amino acid sequence of any of SEQ ID NOs: 5, 6, 300-414, 596-598, and 720-723, wherein the deaminase comprises at least one of the amino acid residues set forth in any one of Tables 2, 4, 6, and 19.

[0190] Biologically active variants of adenine deaminase of the invention can differ by as few as 1 to 10 amino acid residues, such as as few as 1 to 15 amino acid residues, 6 to 10, as few as 5, as few as 4, as few as 3, as few as 2, or as few as 1 amino acid residue. In specific embodiments, the polypeptides comprise N- or C-terminal truncations, which can include deletions of 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, or more amino acids from either the N- or C-terminus of the polypeptide. In some embodiments, the polypeptide comprises an internal deletion, which can include a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, or more amino acids.

[0191] Modifications can be made to the RGNs and deaminases provided herein to generate variant proteins and polynucleotides. Human-designed changes can be introduced through the application of site-directed mutagenesis techniques. In some embodiments, naturally occurring, unknown, or uncharacterized polynucleotides and / or polypeptides structurally and / or functionally related to the sequences disclosed herein can also be identified as being within the scope of the present invention. Conservative amino acid substitutions can be made in non-conserved regions that do not alter the function of the polypeptide as an adenine deaminase.

[0192] Variant polynucleotides and proteins also encompass sequences and proteins derived from mutagenic and recombinogenic procedures, such as DNA shuffling. In such procedures, one or more different RGNs or deaminases are engineered to generate new deaminases possessing desired properties. In this manner, libraries of recombinant polynucleotides are generated from a population of related sequence polynucleotides that share substantial sequence identity and contain sequence regions capable of homologous recombination in vitro or in vivo. For example, using this approach, sequence motifs encoding domains of interest can be engineered to have increased K in the case of enzymes. mThe sequences provided herein can be shuffled between other, later-identified genes to obtain new genes encoding proteins with improved properties of interest, such as: Strategies for such DNA shuffling are known in the art. See, for example, Stemmer (1994) Proc. Natl. Acad. Sci. USA 91:10747-10751, Stemmer (1994) Nature 370:389-391, Crameri et al. (1997) Nature Biotech. 15:436-438, Moore et al. (1997) J. Mol. Biol. 272:336-347, Zhang et al. (1997) Proc. Natl. Acad. Sci. USA 94:4504-4509, Crameri et al. (1998) Nature 391:288-291, and U.S. Patent Nos. 5,605,793 and 5,837,458. "Shuffled" nucleic acids are nucleic acids produced by a shuffling procedure, such as any of the shuffling procedures described herein. Shuffled nucleic acids are produced by recombining (physically or virtually) two or more nucleic acids (or character strings), e.g., artificially, and optionally in a recursive manner. Generally, one or more screening steps are used in the shuffling process to identify nucleic acids of interest, and this screening step can be performed before or after any recombination step. In some (but not all) shuffling embodiments, it is desirable to perform multiple rounds of recombination before selection to increase the diversity of the pool being screened. The overall process of recombination and selection is optionally repeated recursively. Depending on the context, shuffling can refer to the overall process of recombination and selection, or alternatively, can simply refer to the recombination portion of the overall process.

[0193] As used herein, "sequence identity" or "identity" in the context of two polynucleotide or polypeptide sequences refers to the residues of the two sequences that are the same when aligned for maximum correspondence over a specified comparison window. It is recognized that non-identical residue positions often differ by conservative amino acid substitutions, in which an amino acid residue is substituted with another amino acid residue having similar chemical properties (e.g., charge or hydrophobicity), thus not altering the functional properties of the molecule. Protein sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity." Means for measuring sequence similarity are well known to those of skill in the art. Typically, this involves scoring conservative substitutions as partial rather than complete mismatches. Thus, for example, conservative substitutions are given a score between zero and one, with identical amino acids being given a score of one and non-conservative substitutions being given a score of zero. Scoring of conservative substitutions is calculated, for example, as implemented in the program PC / GENE (Intelligenetics, Mountain View, California).

[0194] As used herein, "percentage of sequence identity" is determined by comparing two optimally aligned sequences over a comparison window, where the portion of the polynucleotide sequence within the comparison window may contain additions or deletions (i.e., gaps) compared to a reference sequence (which does not contain additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions where the same nucleic acid base or amino acid residue occurs in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity.

[0195] Unless otherwise specified, sequence identity / similarity values ​​provided herein refer to values ​​obtained using GAP version 10, or any equivalent program, using the following parameters: % identity and % similarity for nucleotide sequences using a GAP weight of 50 and a length weight of 3, and the nwsgapdna.cmp scoring matrix; % identity and % similarity for amino acid sequences using a GAP weight of 8 and a length weight of 2, and the BLOSUM62 scoring matrix. By "equivalent program" is intended any sequence comparison program that produces alignments with identical nucleotide or amino acid residue matches and identical percent sequence identity for any two sequences in question when compared to corresponding alignments produced by GAP version 10.

[0196] Two sequences are "optimally aligned" when they are aligned for similarity scoring using a defined amino acid substitution matrix (e.g., BLOSUM62), gap existence penalties, and gap extension penalties to achieve the highest possible score for the pair of sequences. Amino acid substitution matrices and their use in quantifying similarity between two sequences are well known in the art and are described, for example, in Dayhoff et al. (1978) "A model of evolutionary change in proteins." In "Atlas of Protein Sequence and Structure," Vol. 5, Suppl. 3 (ed. MO Dayhoff), pp. 345-352, Natl. Biomed. Res. Found., Washington, DC, and Henikoff et al. (1992) Proc. Natl. Acad. Sci. USA 89:10915-10919. The BLOSUM62 matrix is ​​often used as the default scoring substitution matrix in sequence alignment protocols. A gap existence penalty is imposed for introducing a single amino acid gap into one of the aligned sequences, and a gap extension penalty is imposed for each additional empty amino acid position inserted into an already opened gap. The alignment is defined by the amino acid positions in each sequence where the alignment begins and ends, and optionally by inserting one or more gaps in one or both sequences to achieve the highest possible score. While optimal alignment and scoring can be achieved manually, this process is facilitated by the use of computer-implemented alignment algorithms, such as gapped BLAST 2.0, described in Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402 and made publicly available at the National Center for Biotechnology Information Website (www.ncbi.nlm.nih.gov).Optimal alignments containing multiple alignments can be prepared using, for example, PSI-BLAST, available through www.ncbi.nlm.nih.gov and described by Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402.

[0197] For an amino acid sequence that is optimally aligned with a reference sequence, the amino acid residue "corresponds" to the position in the reference sequence with which the residue is paired in the alignment. The "position" is indicated by a number that sequentially identifies each amino acid in the reference sequence based on its position relative to the N-terminus. Due to deletions, insertions, truncations, fusions, etc., which must be considered when determining optimal alignment, the number of amino acid residues in a test sequence, determined by simple counting from the N-terminus, is generally not necessarily the same as the number of their corresponding positions in the reference sequence. For example, if there is a deletion in the aligned test sequence, there will be no amino acid corresponding to the position in the reference sequence at the deletion site. If there is an insertion in the aligned reference sequence, the insertion will not correspond to any amino acid position in the reference sequence. In the case of a truncation or fusion, there may be a stretch of amino acids in the reference or aligned sequence that does not correspond to any amino acid in the corresponding sequence.

[0198] VIII. Antibodies Also encompassed are antibodies to the deaminases, fusion proteins, or ribonucleoproteins of the present invention, including deaminases having the amino acid sequences set forth as any one of SEQ ID NOS: 6, 300-414, 596-598, and 720-723, or active variants or fragments thereof. Methods for producing antibodies are well known in the art (see, e.g., Harlow and Lane (1988) Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY, and U.S. Patent No. 4,196,265). These antibodies can be used in kits for detecting and isolating the deaminases or fusion proteins or ribonucleoproteins described herein. Accordingly, the present disclosure provides kits that include an antibody that specifically binds to a polypeptide or ribonucleoprotein described herein, including, for example, a polypeptide comprising a sequence at least 85% identical to any of SEQ ID NOs: 5, 6, 300-414, 596-598, and 720-723, wherein the deaminase comprises at least one of the amino acid residues set forth in any one of Tables 2, 4, 6, and 19.

[0199] IX. SYSTEMS AND RIBONUCLEOPROTEIN COMPLEXES FOR BINDING AND / OR MODIFYING TARGET SEQUENCES OF INTEREST AND METHODS OF MAKING SAME The present disclosure provides a system for targeting a nucleic acid sequence and modifying the target nucleic acid sequence. In some embodiments, the system targets a heterologous polypeptide to a nucleic acid sequence and can modify the target nucleic acid sequence. In those embodiments in which the heterologous polypeptide comprises a base-editing polypeptide (e.g., a deaminase) or a prime-editing polypeptide, the base-editing polypeptide (e.g., a deaminase) or prime-editing polypeptide (e.g., an RGN nickase or a nuclease-inactive RGN) fused to an RGN is involved in modifying the target nucleic acid sequence. A guide RNA hybridizes to the target sequence of interest and forms a complex with the RGN component of the fusion protein, thereby guiding the fusion protein to bind to the target sequence.

[0200] In some embodiments, an RNA-guided DNA-binding polypeptide such as an RGN (e.g., an RGN nickase or a nuclease-active RGN) and a gRNA are involved in targeting the ribonucleoprotein complex to a nucleic acid sequence of interest, and a deaminase polypeptide fused to RGDBP is involved in modifying the target nucleic acid sequence. In those embodiments in which the deaminase is an adenine deaminase, the base editor modifies A>N. In some embodiments, the adenine deaminase converts A>G. The guide RNA hybridizes to the target sequence of interest and forms a complex with the RNA-guided DNA-binding polypeptide, thereby guiding the RNA-guided DNA-binding polypeptide to bind to the target sequence. The RNA-guided DNA-binding polypeptide is part of a fusion protein that also includes a deaminase, such as those described herein. In some embodiments, the RNA-guided DNA-binding polypeptide is an RGN, such as Cas9. Other examples of RNA-guided DNA-binding polypeptides include RGNs, such as those described in International Patent Applications Nos. 2019 / 236566 and 2020 / 139783, each of which is incorporated by reference in its entirety. In some embodiments, the RNA-guided DNA-binding polypeptide is a type II CRISPR-Cas polypeptide, or an active variant or fragment thereof. In some embodiments, the RNA-guided DNA-binding polypeptide is a type V CRISPR-Cas polypeptide, or an active variant or fragment thereof. In some embodiments, the RNA-guided DNA-binding polypeptide is a type VI CRISPR-Cas polypeptide. In some embodiments, the DNA-binding polypeptide of the fusion protein does not require an RNA guide, such as a zinc finger nuclease, TALEN, or meganuclease polypeptide. In some embodiments, the nuclease activity of the DNA-binding polypeptide is partially or completely inactivated. In a further embodiment, the RNA-guided DNA-binding polypeptide comprises the amino acid sequence of an RGN, such as, for example, APG07433.1 (SEQ ID NO: 1), or an active variant or fragment thereof, such as nickase nAPG07433.1 (SEQ ID NO: 49).

[0201] In some embodiments, the systems provided herein for binding and modifying a target sequence of interest are ribonucleoprotein complexes, which are at least one molecule of RNA bound to at least one protein (i.e., a fusion protein). The ribonucleoprotein complexes provided herein comprise a fusion protein comprising at least one guide RNA as the RNA component, and an RNA-guided DNA-binding polypeptide (e.g., RGN) and a heterologous polypeptide (e.g., a deaminase) inserted therein as the protein components. In some embodiments, the ribonucleoprotein complex is purified from cells or organisms transformed with a polynucleotide encoding the fusion protein and the guide RNA, and cultured under conditions that allow expression of the fusion protein and the guide RNA.

[0202] Methods for producing a deaminase, a fusion protein, or a fusion protein-ribonucleoprotein complex are provided. Such methods include culturing cells comprising a nucleotide sequence encoding the deaminase, the fusion protein, and, in some embodiments, a nucleotide sequence encoding the guide RNA, under conditions in which the deaminase or the fusion protein (and, in some embodiments, the guide RNA) is expressed. The deaminase, the fusion protein, or the fusion ribonucleoprotein can then be purified from a lysate of the cultured cells.

[0203] Methods for purifying deaminases, fusion proteins, or fusion ribonucleoproteins from lysates of biological samples are known in the art (e.g., size exclusion and / or affinity chromatography, 2D-PAGE, HPLC, reverse-phase chromatography, immunoprecipitation). In certain methods, the deaminase or fusion protein is recombinantly produced and includes a purification tag to aid in its purification, including, but not limited to, glutathione-S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, HA, nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, 6xHis, biotin carboxyl carrier protein (BCCP), and calmodulin. Generally, the tagged deaminase, fusion protein, or fusion ribonucleoprotein complex is purified using immunoprecipitation or other similar methods known in the art.

[0204] An "isolated" or "purified" polypeptide, or biologically active portion thereof, is substantially or essentially free from components that normally accompany or interact with the polypeptide as found in its naturally occurring environment. Thus, an isolated or purified polypeptide is substantially free of other cellular material or culture medium when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized. A protein that is substantially free of cellular material includes preparations of the protein having less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% (by dry weight) of contaminating proteins. When the protein of the invention or a biologically active portion thereof is recombinantly produced, optimally, the culture medium represents less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% (by dry weight) of chemical precursors or non-target protein chemicals.

[0205] Certain methods provided herein for binding and / or cleaving a target sequence of interest involve the use of an in vitro assembled ribonucleoprotein complex. In vitro assembly of a ribonucleoprotein complex can be carried out using any method known in the art in which a fusion protein containing RGDBP (e.g., RGN) is contacted with a guide RNA under conditions that allow the RGDBP (e.g., RGN) fusion protein to bind to the guide RNA. As used herein, "contact," "contacting," or "contacted" refers to placing components of a desired reaction together under conditions suitable for carrying out the desired reaction. The RGDBP (e.g., RGN) fusion protein can be purified from a biological sample, cell lysate, or culture medium, produced via in vitro translation, or chemically synthesized. The guide RNA can be purified from a biological sample, cell lysate, or culture medium, transcribed in vitro, or chemically synthesized. The RGDBP (e.g., RGN) fusion protein and guide RNA can be contacted in solution (e.g., buffered saline) to allow in vitro assembly of the ribonucleoprotein complex.

[0206] X. Methods for localizing heterologous polypeptides to or modifying target DNA molecules The present disclosure provides methods for localizing a heterologous polypeptide (e.g., a deaminase) to a target nucleic acid molecule of interest (e.g., a target DNA molecule) or modifying the target nucleic acid molecule. The methods comprise delivering a fusion protein of the present disclosure or a polynucleotide encoding the same to the target sequence, or to a cell, organelle, or embryo containing the target sequence. In certain embodiments, the methods comprise delivering a system comprising at least one guide RNA or a polynucleotide encoding the same and at least a fusion protein of the present disclosure to the target sequence, or to a cell, organelle, or embryo containing the target sequence.

[0207] In some embodiments, the method includes contacting a DNA molecule with (a) a fusion protein and (b) a gRNA that targets the fusion protein of (a) to a target nucleotide sequence of the DNA molecule, wherein the DNA molecule is contacted with the fusion protein and the gRNA in amounts and under suitable conditions effective for binding of the RGBDP (e.g., RGN) component to the target sequence.

[0208] The target DNA molecule can contain a sequence associated with a disease or disorder, and editing of at least one nucleobase within the causative mutation can result in a sequence not associated with the disease or disorder. In some embodiments, the disease or disorder affects an animal. In further embodiments, the disease or disorder affects a mammal, such as a human, cow, horse, dog, cat, goat, sheep, pig, monkey, rat, mouse, or hamster. In some embodiments, the target DNA sequence is present in an allele of a crop plant, and a particular allele of a trait of interest results in a plant with lower agronomic value. Editing at least one nucleobase within this region results in an allele that improves the trait and increases the agronomic value of the plant.

[0209] After delivery of the polynucleotide encoding the guide RNA and / or fusion protein, the cell or embryo can then be cultured under conditions in which the guide RNA and / or fusion protein is expressed. In various embodiments, the method comprises contacting the target sequence with a ribonucleoprotein complex comprising the gRNA and the fusion protein. In certain embodiments, the method comprises introducing a ribonucleoprotein complex of the invention into a cell, organelle, or embryo containing the target sequence. The ribonucleoprotein complex of the invention can be a complex purified from a biological sample, a complex recombinantly produced and then purified, or a complex assembled in vitro, as described herein. In those embodiments in which the ribonucleoprotein complex contacted with the target sequence or organelle or embryo is assembled in vitro, the method can further comprise in vitro assembly of the complex prior to contacting with the target sequence, cell, organelle, or embryo.

[0210] The purified or in vitro assembled ribonucleoprotein complexes of the present invention can be introduced into cells, organelles, or embryos using any method known in the art, including, but not limited to, electroporation. In some embodiments, the fusion protein (or a polynucleotide encoding it) and a polynucleotide encoding or comprising a guide RNA are introduced into cells, organelles, or embryos using any method known in the art (e.g., electroporation).

[0211] Upon delivery to or contact with the target sequence, or a cell, organelle, or embryo containing the target sequence, the guide RNA directs the fusion protein to bind to the target sequence in a sequence-specific manner, which can then be modified in cases where the fusion protein comprises a base-editing polypeptide (e.g., a deaminase) or a prime-editing polypeptide.

[0212] In some embodiments, where the fusion protein is a base editor (i.e., comprises a base-editing polypeptide such as a deaminase), binding of the fusion protein to the target sequence results in modification of nucleotides adjacent to the target sequence. The nucleobases adjacent to the target sequence that are modified by the deaminase can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 base pairs from the 5' or 3' end of the target sequence. The full range of nucleotides that can be edited using a base editor fusion protein (e.g., RGN fused to a deaminase), typically expressed as the distance of nucleotides from the PAM sequence (e.g., nucleotides at positions 8 and 22 upstream, i.e., 5', from the PAM sequence), is referred to herein as the "editing window" of a particular base editor fusion protein. The editing window can be, for example, 1 to 100 base pairs 5' or 3' of the PAM sequence, including, but not limited to, 5 to 50, 5 to 25, 8 to 22, 10 to 20, or 10 to 21 base pairs 5' or 3' of the PAM sequence. Fusion proteins comprising an adenine deaminase, such as those disclosed herein, and an RNA-guided DNA-binding polypeptide (e.g., RGN), can introduce targeted A>N mutations into target DNA molecules. In some embodiments, the fusion protein introduces targeted A>G mutations into target DNA molecules.

[0213] Methods for measuring the binding of fusion proteins to target sequences are known in the art and include chromatin immunoprecipitation assays, gel mobility shift assays, DNA pull-down assays, reporter assays, and microplate capture and detection assays. Similarly, methods for measuring the cleavage or modification of target sequences are known in the art and include in vitro or in vivo cleavage assays in which cleavage is confirmed using PCR, sequencing, or gel electrophoresis, with or without an appropriate label (e.g., radioisotope, fluorescent substance) attached to the target sequence to facilitate detection of degradation products. In some embodiments, a nicking-triggered exponential amplification (NTEXPAR) assay is used (see, e.g., Zhang et al. (2016) Chem. Sci. 7:4951-4957). In vivo cleavage can be assessed using the Surveyor assay (Guschin et al. (2010) Methods Mol Biol 649:247-256).

[0214] Methods for measuring the base editing activity of a base editor fusion protein or a prime editor fusion protein are known in the art and include those described herein, as well as any form of sequence analysis of the target nucleic acid molecule after contact with the base editor or prime editor (e.g., PCR, sequencing, or gel electrophoresis, with or without the attachment of an appropriate label). Base editor fusion proteins comprising a deaminase disclosed herein can exhibit improved editing activity (i.e., improved editing efficiency or reduced RNA editing), a larger editing window, or both, when compared to a similar base editor fusion protein comprising the parent LPG50148 deaminase. The base editing efficiency of a first base editor fusion protein comprising a deaminase of the invention (e.g., an active variant or fragment thereof comprising at least one of the amino acid residues set forth in any one of SEQ ID NOs: 5, 6, 300-414, 596-598, and 720-723, or Tables 2, 4, 6, and 19) is about 1.1-fold, about 1.5-fold, about 2-fold, about 2.5-fold, about 3-fold, about 3.5-fold, about 4-fold, about 4.5-fold, about 5-fold, about 5.5-fold, about 6-fold, about 6.5-fold, about 7-fold, about 7.5-fold, about 8-fold, about 8.5-fold, about 9-fold, about 9.5-fold, about 10-fold, The first base editor fusion protein can have an editing efficiency that is 1.1-10 fold, 1.1-20 fold, 1.1-30 fold, or more, greater than a second base editor fusion protein comprising the parent LPG50148 deaminase and the same DNA-binding polypeptide (e.g., RGN) as the parent LPG50148 deaminase, including, but not limited to, about 11-fold, about 12-fold, about 13-fold, about 15-fold, about 16-fold, about 17-fold, about 18-fold, about 19-fold, about 20-fold, about 21-fold, about 22-fold, about 23-fold, about 24-fold, about 25-fold, about 26-fold, about 27-fold, about 28-fold, about 29-fold, and about 30-fold greater. Base editor fusion proteins comprising a deaminase of the invention can exhibit reduced RNA editing when compared to similar base editor fusion proteins comprising the parent LPG50148 deaminase.A first base editor fusion protein comprising a deaminase of the invention can exhibit an RNA editing rate of about 1% to about 99% of that of a second base editor fusion protein comprising the parent LPG50148 deaminase and the same DNA-binding polypeptide (e.g., RGN) as the first base editor fusion protein, including but not limited to about 1%, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, and about 99%.

[0215] The editing window of a first base editor fusion protein comprising a deaminase of the invention (e.g., SEQ ID NOS: 5, 6, 300-414, 596-598, and 720-723, or an active variant or fragment thereof comprising at least one of the amino acid residues set forth in any one of Tables 2, 4, 6, and 19) can be at least one nucleotide longer on either or both sides of the editing window of a second base editor fusion protein comprising the parent LPG50148 deaminase and the same DNA binding protein (e.g., RGN) as the first base editor fusion protein, including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides longer on either or both sides of the parent base editor editing window.

[0216] In some embodiments, the methods involve the use of a fusion protein in which an RGDBP (e.g., RGN) is complexed with more than one guide RNA. The more than one guide RNA can target different regions of a single gene or can target multiple genes. This multiple targeting allows the deaminase or prime editing polypeptide of the fusion protein to modify the nucleic acid, thereby introducing multiple mutations into a target nucleic acid molecule (e.g., a genome) of interest.

[0217] In embodiments where the method involves the use of an RNA-guided nuclease (RGN), such as the nickase RGN (i.e., capable of cleaving only one strand of a double-stranded polynucleotide, e.g., nAPG07433.1 (SEQ ID NO: 49)), the method can include introducing two different RGNs or RGN variants that target the same or overlapping target sequences and cleave different strands of the polynucleotide. For example, an RGN nickase that cleaves only the positive (+) strand of a double-stranded polynucleotide can be introduced along with a second RGN nickase that cleaves only the negative (-) strand of the double-stranded polynucleotide. In some embodiments, two different fusion proteins are provided, each containing a different RGN with a different PAM recognition sequence, so that a greater variety of nucleotide sequences can be targeted for mutation.

[0218] Those skilled in the art will understand that any of the methods disclosed herein can be used to target a single target sequence or multiple target sequences. Thus, the methods include the use of a fusion protein comprising a single RGDBP (e.g., RGN) in combination with multiple distinct guide RNAs, which can target multiple distinct sequences within a single gene and / or multiple genes. Also encompassed herein are methods in which multiple distinct guide RNAs are introduced in combination with multiple distinct RGN fusion proteins. These guide RNA and guide RNA / fusion protein systems can target multiple distinct sequences within a single gene and / or multiple genes.

[0219] In some embodiments, fusion proteins comprising an RNA-guided DNA-binding polypeptide (e.g., RGN) and a heterologous polypeptide (e.g., a deaminase) can be used to generate mutations within a target gene or target region of a gene of interest. In some embodiments, the fusion proteins of the invention can be used for saturation mutagenesis of a target gene or region of a target gene of interest, followed by high-throughput forward genetic screening to identify novel mutations and / or phenotypes. In some embodiments, the fusion proteins described herein can be used to generate mutations at target genomic locations that may or may not include coding DNA sequences. Libraries of cell lines generated by the targeted mutagenesis described above can also be useful for studying gene function or gene expression.

[0220] XI. Target Polynucleotides In one aspect, the present invention provides methods for modifying a target polynucleotide in a eukaryotic cell, which may be in vivo, ex vivo, or in vitro. In some embodiments, the method includes sampling a cell or population of cells from a human or non-human animal or plant (including microalgae) and modifying the cell or cells. Culturing can occur ex vivo at any stage. The cells can be reintroduced into a human, non-human animal, or plant (including microalgae).

[0221] Using natural variability, plant breeders combine the most useful genes for desirable qualities such as yield, quality, uniformity, durability, and resistance to pests. These desirable qualities also include growth, day-length preference, temperature requirements, the onset of flowering or reproductive development, fatty acid content, insect resistance, disease resistance, nematode resistance, fungus resistance, herbicide resistance, and tolerance to various environmental factors, including adverse soil conditions, including drought, heat, humidity, cold, wind, and high salinity. Sources of these useful genes include native or exotic varieties, heirloom varieties, wild plant relatives, and induced mutations, such as treating plant material with mutagens. Using the present invention, plant breeders are provided with a new tool for inducing mutations. Thus, those skilled in the art can employ the present invention to induce increases in useful genes with greater precision than previous mutagens, thereby accelerating and improving plant breeding programs.

[0222] The target polynucleotide of a fusion protein of the present invention can be any polynucleotide, endogenous or exogenous to a eukaryotic cell. For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. In some embodiments, the target polynucleotide is a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). In some embodiments, the target sequence of a fusion protein of the present invention is associated with a PAM (protospacer adjacent motif), i.e., a short sequence recognized by an RNA-guided DNA-binding polypeptide (e.g., RGN). While the exact sequence and length requirements of the PAM vary depending on the RNA-guided DNA-binding polypeptide used, the PAM is typically a 2-5 base pair sequence adjacent to the protospacer (i.e., the target sequence).

[0223] The target polynucleotides of the fusion proteins of the present invention can include numerous disease-associated genes and polynucleotides, as well as signal transduction biochemical pathway-related genes and polynucleotides. Examples of target polynucleotides include sequences associated with signal transduction biochemical pathways, such as signal transduction biochemical pathway-related genes or polynucleotides. Examples of target polynucleotides include disease-associated genes or polynucleotides. A "disease-associated" gene or polynucleotide refers to any gene or polynucleotide whose transcription or translation product occurs at an abnormal level or in an abnormal form in cells derived from disease-affected tissue compared to non-diseased control tissue or cells. It can be a gene that becomes expressed at an abnormally high level, or it can be a gene that becomes expressed at an abnormally low level, and altered expression correlates with the development and / or progression of the disease. A disease-associated gene also refers to a gene that harbors mutation(s) or genetic variation that is directly involved in or in linkage disequilibrium with gene(s) involved in the etiology (e.g., causative mutation) of the disease. The transcription or translation product can be known or unknown and can be at normal or abnormal levels.

[0224] Non-limiting examples of disease-associated genes that can be targeted using the methods and compositions of the present disclosure are available on the World Wide Web from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.).

[0225] In some embodiments, the method comprises contacting a DNA molecule comprising a target DNA sequence with a fusion protein of the invention, wherein the DNA molecule is contacted with the fusion protein in an amount and under suitable conditions effective to modify the target DNA molecule. In certain embodiments, the method comprises contacting a DNA molecule comprising the target DNA sequence with (a) an RGN fusion protein of the invention comprising a base-editing polypeptide, and (b) a gRNA that targets the fusion protein of (a) to a target nucleotide sequence in a DNA strand, wherein the DNA molecule is contacted with the fusion protein and the gRNA in an amount and under suitable conditions effective to deaminate at least one nucleobase. In some embodiments, the target DNA sequence comprises a sequence associated with a disease or disorder, and deamination of the nucleobase results in a sequence not associated with the disease or disorder. In some embodiments, the target DNA sequence is present in an allele of a crop, and a particular allele of a trait of interest results in a plant with lower agronomic value. Deamination of the nucleobase results in an allele that improves the trait and increases the agronomic value of the plant.

[0226] In some embodiments, the target DNA sequence contains a point mutation associated with a disease or disorder, and deamination of the mutant base results in a sequence that is not associated with the disease or disorder, hi some embodiments, deamination corrects the point mutation in the sequence associated with the disease or disorder.

[0227] In some embodiments, the sequence associated with the disease or disorder encodes a protein, and deamination or prime editing introduces a stop codon into the sequence associated with the disease or disorder, resulting in truncation of the encoded protein. In some embodiments, the contacting is performed in vivo in a subject suspected of having, having, or diagnosed with a disease or disorder. In some embodiments, the disease or disorder is a disease associated with a point mutation or single base mutation in the genome. In some embodiments, the disease is a genetic disease, cancer, metabolic disease, or lysosomal storage disease.

[0228] XII. Pharmaceutical Compositions and Methods of Treatment Provided herein is a method of treating a disease in a subject in need thereof, the method comprising administering to a subject in need thereof a fusion protein of the present disclosure or a polynucleotide encoding same, a gRNA or a polynucleotide encoding same, a fusion protein system of the present disclosure, or a cell modified by or containing any one of these compositions.

[0229] In some embodiments, the treatment involves in vivo gene editing by administering a fusion protein, gRNA, or fusion protein system of the present disclosure, or a polynucleotide(s) encoding the same, to a subject in need thereof. In some embodiments, the treatment involves ex vivo gene editing, in which cells are genetically modified ex vivo with a fusion protein, gRNA, or fusion protein system of the present disclosure, or a polynucleotide(s) encoding the same, and then the modified cells are administered to the subject. In some embodiments, the genetically modified cells are then derived from the subject receiving the modified cells, and the transplanted cells are referred to herein as autologous. In some embodiments, the genetically modified cells are derived from a different subject (i.e., donor) within the same species as the subject receiving the modified cells (i.e., recipient), and the transplanted cells are referred to herein as allogeneic. In some examples described herein, the cells can be grown in culture prior to administration to the subject in need thereof.

[0230] In some embodiments, the disease treated with the compositions of the present disclosure is a disease that can be treated with immunotherapy, such as chimeric antigen receptor (CAR) T cells. Such diseases include, but are not limited to, cancer.

[0231] In some embodiments, modification of the target sequence (e.g., deamination) results in the correction of a genetic defect or a point mutation that leads to loss of function in the gene product. In some embodiments, the genetic defect is associated with a disease or disorder, e.g., a lysosomal storage disorder, or a metabolic disease such as, e.g., type I diabetes. Thus, in some embodiments, the disease treated with the compositions of the present disclosure is associated with a sequence (i.e., the sequence is causative of the disease or disorder or is responsible for symptoms associated with the disease or disorder) that is mutated to treat the disease or disorder or alleviate symptoms associated with the disease or disorder.

[0232] In some embodiments, the disease treated with the compositions of the present disclosure is associated with a causative mutation. As used herein, "causative mutation" refers to a specific nucleotide, multiple nucleotides, or nucleotide sequence in a genome that contributes to the severity or presence of a disease or disorder in a subject. Correction of the causative mutation leads to amelioration of at least one symptom attributable to the disease or disorder. In some embodiments, correction of the causative mutation leads to amelioration of at least one symptom attributable to the disease or disorder. In some embodiments, the causative mutation is adjacent to a PAM site recognized by RGDBP (e.g., RGN) of the fusion proteins disclosed herein. The causative mutation can be corrected with the fusion protein. Non-limiting examples of disease-associated genes and mutations are available on the World Wide Web from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.).

[0233] In some embodiments, the methods provided herein are used to introduce inactivating point mutations into genes or alleles encoding gene products associated with a disease or disorder. For example, in some embodiments, methods are provided herein employing fusion proteins to introduce inactivating point mutations into cancer genes (e.g., in the treatment of proliferative disorders). Inactivating mutations, in some embodiments, can generate premature stop codons in the coding sequence, resulting in expression of truncated gene products, e.g., truncated proteins lacking the function of the full-length protein. In some embodiments, the goal of the methods provided herein is to restore the function of dysfunctional genes through genome editing. The fusion proteins provided herein can be validated for in vitro gene editing-based human therapy, for example, by correcting disease-associated mutations in human cell cultures. It will be understood by those skilled in the art that the fusion proteins provided herein, e.g., fusion proteins comprising an RNA-guided DNA-binding polypeptide and an adenine deaminase polypeptide, can be used to correct any single-point A>G mutation. Deamination of the mutant A to G leads to correction of the mutation.

[0234] As used herein, "treatment" or "treating," or "alleviating," or "ameliorating," are used interchangeably. These terms refer to an approach to obtaining beneficial or desired results, including, but not limited to, therapeutic benefit and / or prophylactic benefit. Therapeutic benefit refers to any therapeutically relevant improvement or effect in one or more diseases, conditions, or symptoms being treated. For prophylactic benefit, the composition may be administered to a subject at risk of developing a particular disease, condition, or symptom, or a subject reporting one or more physiological symptoms of a disease, even if the disease, condition, or symptom has not yet manifested.

[0235] The term "effective amount" or "therapeutically effective amount" refers to an amount of an agent sufficient to produce a beneficial or desired result. A therapeutically effective amount may vary depending on one or more of the subject and disease state being treated, the subject's weight and age, the severity of the disease state, the mode of administration, etc., and these can be readily determined by one of ordinary skill in the art. The specific dose may vary depending on one or more of the particular agent selected, the administration regimen to be followed, whether it is administered in combination with other compounds, the timing of administration, and the delivery system in which it is delivered.

[0236] The term "administering" refers to the placement of an active ingredient in a subject by a method or route that results in at least partial localization of the introduced active ingredient at a desired site, such as a site of injury or repair, so that a desired effect(s) occurs. In those embodiments in which cells are administered, the cells can be administered by any suitable route that results in delivery to a desired location in a subject where at least a portion of the transplanted cells or components of the cells remain viable. The survival period of the cells after administration to a subject can be as short as a few hours, e.g., 24 hours to a few days, to as long as several years, or even the lifetime of the patient, i.e., long-term engraftment. For example, in some aspects described herein, an effective amount of photoreceptor cells or retinal progenitor cells is administered via a systemic route of administration, e.g., intraperitoneal or intravenous.

[0237] In some embodiments, administering comprises administering by viral delivery. In some embodiments, administering comprises administering by electroporation. In some embodiments, administering comprises administering by nanoparticle delivery. In some embodiments, administering comprises administering by liposome delivery. In some embodiments, administering comprises administering by lipid nanoparticle (LNP) delivery. In some embodiments, LNPs or lipid components thereof are described in WO2022173531, WO2022150485, or U.S. Provisional Application No. 63 / 492,537, filed March 28, 2023, the contents of each of which are incorporated herein by reference in their entirety. Any effective route of administration can be used to administer an effective amount of the pharmaceutical compositions described herein. In some embodiments, administering comprises administering by a method selected from the group consisting of intravenous, subcutaneous, intramuscular, oral, rectal, aerosol, parenteral, ophthalmic, pulmonary, transdermal, vaginal, otic, nasal, and topical administration, or any combination thereof. In some embodiments, administration by injection or infusion is used for delivery of the cells.

[0238] As used herein, the term "subject" refers to any individual for whom diagnosis, treatment, or therapy is desired. In some embodiments, the subject is an animal. In some embodiments, the subject is a mammal. In some embodiments, the subject is a human. The human can be an adult, an adolescent, a child, or an infant.

[0239] The effectiveness of treatment can be determined by a skilled clinician. However, a treatment is considered "effective treatment" if any one or all of the signs or symptoms of a disease or disorder are modified in a beneficial manner (e.g., reduced by at least 10%) or other clinically acceptable symptoms or markers of the disease are improved or ameliorated. Efficacy can also be measured by the individual's failure to deteriorate (e.g., the progression of the disease is halted or at least slowed), as assessed by the need for hospitalization or medical intervention. Methods for measuring these indicators are well known to those skilled in the art. Treatment includes: (1) inhibiting the disease, e.g., preventing or slowing the progression of symptoms, or (2) relieving the disease, e.g., causing regression of symptoms, and (3) preventing or reducing the likelihood of the onset of symptoms.

[0240] Pharmaceutical compositions are provided that include a fusion protein of the present disclosure or a polynucleotide encoding same, a system of the present disclosure, or a cell containing any of the fusion protein, the polynucleotide encoding same, or the system, and a pharmaceutically acceptable carrier.

[0241] As used herein, a "pharmaceutically acceptable carrier" refers to a material that does not cause significant irritation to an organism and does not abolish the activity and properties of the active ingredient (e.g., a deaminase or fusion protein, or a nucleic acid molecule encoding same). The carrier must be of sufficiently high purity and sufficiently low toxicity to be suitable for administration to the subject being treated. The carrier may be inert or may possess pharmaceutical benefits. In some embodiments, a pharmaceutically acceptable carrier comprises one or more compatible solid or liquid fillers, diluents, or encapsulating substances that are suitable for administration to humans or other vertebrates. In some embodiments, a pharmaceutical composition comprises a pharmaceutically acceptable carrier that is not naturally occurring. In some embodiments, the pharmaceutically acceptable carrier and the active ingredient are not found together in nature and are therefore heterologous.

[0242] Pharmaceutical compositions used in the methods of the present disclosure can be formulated with suitable carriers, excipients, and other agents that provide for suitable transport, delivery, tolerability, etc. Many suitable formulations are known to those skilled in the art. See, for example, Remington, The Science and Practice of Pharmacy (21 st See, e.g., ed. 2005). Non-limiting examples include sterile diluents such as water for injection, saline, fixed oils, polyethylene glycol, glycerin, propylene glycol, or other synthetic solvents; antibacterial agents such as benzyl alcohol or methylparabens; antioxidants such as ascorbic acid or sodium bisulfite; chelating agents such as ethylenediaminetetraacetic acid; buffers such as acetates, citrates, or phosphates, and agents for adjusting tonicity such as sodium chloride or dextrose. For intravenous administration, particular carriers are saline or phosphate-buffered saline (PBS). Pharmaceutical compositions for oral or parenteral use can be prepared in a unit dosage form suitable for the dosage of the active ingredient. Such unit dosage forms include, for example, tablets, pills, capsules, injections (ampoules), suppositories, and the like. These compositions may also contain adjuvants, including preservatives, wetting agents, emulsifying agents, and dispersing agents. Prevention of the action of microorganisms can be ensured by various antibacterial and antifungal agents, for example, parabens, chlorobutanol, phenol, sorbic acid. It may also be desirable to include isotonic agents, for example, sugars, sodium chloride, etc. Prolonged absorption of the injectable pharmaceutical form can be brought about by the use of agents delaying absorption, for example, aluminum monostearate and gelatin.

[0243] In some embodiments in which cells containing or modified with the disclosed fusion proteins, systems, or polynucleotides encoding same are administered to a subject, the cells are administered as a suspension with a pharmaceutically acceptable carrier. Those skilled in the art will recognize that pharmaceutically acceptable carriers used in cell compositions do not contain amounts of buffers, compounds, cryopreservatives, preservatives, or other agents that substantially interfere with the viability of the cells delivered to a subject. Cell-containing formulations can include, for example, an osmotic buffer that allows the integrity of the cell membrane to be maintained, and optionally, nutrients to maintain cell viability or enhance engraftment upon administration. Such formulations and suspensions are known to those skilled in the art and / or can be adapted for use with the cells described herein using routine experimentation.

[0244] The cell compositions can also be emulsified or presented as liposomal compositions, provided that the emulsification procedure does not adversely affect cell viability. The cells and any other active ingredients can be mixed with excipients that are pharmaceutically acceptable, compatible with the active ingredients, and in amounts suitable for use in the therapeutic methods described herein.

[0245] Additional agents included in the cell composition can include pharmaceutically acceptable salts of the components therein. Pharmaceutically acceptable salts include acid addition salts (formed with the free amino groups of the polypeptide) formed with inorganic acids such as hydrochloric acid or phosphoric acid, or organic acids such as acetic acid, tartaric acid, mandelic acid, etc. Salts formed with free carboxyl groups can also be derived from inorganic bases such as sodium, potassium, ammonium, calcium, or ferric hydroxide, and organic bases such as isopropylamine, trimethylamine, 2-ethylaminoethanol, histidine, procaine, etc.

[0246] XIII. Cells Containing Polynucleotide Gene Modifications Provided herein are cells and organisms that contain a target nucleic acid molecule of interest that has been modified using a process mediated by a fusion protein, optionally with a gRNA, as described herein.

[0247] In some embodiments, the fusion protein comprises a deaminase polypeptide comprising the amino acid sequence of any of SEQ ID NOs: 5, 6, 300-414, 596-598, and 720-723, or an active variant or fragment thereof comprising at least one of the amino acid residues set forth in any one of Tables 2, 4, 6, and 19. In some embodiments, the fusion protein comprises a deaminase comprising an amino acid sequence having about 50% to about 99% or more identity to any of SEQ ID NOs: 5, 6, 300-414, 596-598, and 720-723. In some embodiments, the fusion protein comprises a deaminase comprising an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to any of SEQ ID NOs: 5, 6, 300-414, 596-598, and 720-723. In some embodiments, the fusion protein comprises a deaminase and a DNA-binding polypeptide (e.g., an RNA-guided DNA-binding polypeptide). In further embodiments, the fusion protein comprises a deaminase and an RGN or a variant thereof, such as, for example, APG07433.1 (SEQ ID NO: 1) or its nickase variant nAPG07433.1 (SEQ ID NO: 49). In some embodiments, the fusion protein comprises a deaminase and Cas9 or a variant thereof, such as, for example, dCas9 or nickase-Cas9. In some embodiments, the fusion protein comprises a nuclease-inactive or nickase variant of a Type II CRISPR-Cas polypeptide. In some embodiments, the fusion protein comprises a nuclease-inactive or nickase variant of a Type V CRISPR-Cas polypeptide. In some embodiments, the fusion protein comprises a nuclease-inactive or nickase variant of a Type VI CRISPR-Cas polypeptide.

[0248] The modified cells can be eukaryotic (e.g., mammalian, plant, insect, or avian cells) or prokaryotic. Also provided are organelles and embryos comprising at least one nucleotide sequence modified by a process utilizing the fusion proteins described herein. The genetically modified cells, organisms, organelles, and embryos can be heterozygous or homozygous for the modified nucleotide sequence. The mutation(s) introduced by the fusion protein can result in altered expression (up- or down-regulation), inactivation, or expression of a modified protein product or incorporated sequence. In these cases, where the mutation(s) result in either gene inactivation or expression of a non-functional protein product, the genetically modified cell, organism, organelle, or embryo is referred to as a "knockout." The knockout phenotype can be the result of a deletion mutation (i.e., deletion of at least one nucleotide), an insertion mutation (i.e., insertion of at least one nucleotide), or a nonsense mutation (i.e., substitution of at least one nucleotide such that a stop codon is introduced).

[0249] In some embodiments, the mutation(s) introduced by the fusion protein result in the production of a variant protein product. The expressed variant protein product can have at least one amino acid substitution and / or at least one amino acid addition or deletion. The variant protein product can exhibit modified properties or activities, including, but not limited to, altered enzymatic activity or substrate specificity, when compared to the wild-type protein.

[0250] In some embodiments, the mutation(s) introduced by the fusion protein result in an altered expression pattern of the protein. As a non-limiting example, a mutation(s) in a regulatory region controlling expression of a protein product may result in overexpression or downregulation of the protein product, or an altered tissue or temporal expression pattern.

[0251] The modified cells can be grown into organisms such as plants according to conventional methods. See, for example, McCormick et al. (1986) Plant Cell Reports 5:81-84. These plants can then be grown and pollinated with either the same modified strain or a different strain, and the resulting hybrids can have the genetic modification. The present invention provides genetically modified seeds. Progeny, variants, and mutants of the regenerated plants are also included within the scope of the present invention, provided that these portions contain the genetic modification. Further provided are processed plant products or by-products that carry the genetic modification, including, for example, soybean meal.

[0252] The methods provided herein can be used to modify any plant species, including, but not limited to, monocotyledonous or dicotyledonous plants. Examples of plants of interest include, but are not limited to, corn (maize), sorghum, wheat, sunflower, tomato, cruciferous plants, pepper, potato, cotton, rice, soybean, sugar beet, sugarcane, tobacco, barley and rapeseed, Brassica, alfalfa, rye, millet, safflower, peanut, sweet potato, cassava, coffee, coconut, pineapple, citrus trees, cocoa, tea, banana, avocado, fig, guava, mango, olive, papaya, cashew, macadamia, almond, oat, vegetables, ornamental plants, and conifers.

[0253] Vegetables include, but are not limited to, tomatoes, lettuce, green beans, lima beans, peas, and members of the genus Curcumis, such as cucumbers, cantaloupes, and muskmelons. Ornamental plants include, but are not limited to, azaleas, hydrangeas, hibiscus, roses, tulips, daffodils, petunias, carnations, poinsettias, and chrysanthemums. Preferably, the plants of the present invention are crop plants (e.g., corn, sorghum, wheat, sunflowers, tomatoes, crucifers, peppers, potatoes, cotton, rice, soybeans, sugar beets, sugarcane, tobacco, barley, rapeseed, etc.).

[0254] The methods provided herein can also be used to genetically modify any prokaryotic species, including, but not limited to, archaea and bacteria (e.g., Bacillus sp., Klebsiella sp., Streptomyces sp., Rhizobium sp., Escherichia sp., Pseudomonas sp., Salmonella sp., Shigella sp., Vibrio sp., Yersinia sp., Mycoplasma sp., Agrobacterium, Lactobacillus sp.).

[0255] The methods provided herein can be used to genetically modify any eukaryotic species or cells therefrom, including, but not limited to, animals (e.g., mammals, insects, fish, birds, and reptiles), fungi, amoebae, algae, and yeast. In some embodiments, cells modified by the methods of the present disclosure include cells of hematopoietic origin, such as cells of the immune system (i.e., cells of the innate or adaptive immune system), including, but not limited to, B cells, T cells, natural killer (NK) cells, pluripotent stem cells, induced pluripotent stem cells, chimeric antigen receptor T (CAR-T) cells, monocytes, macrophages, and dendritic cells.

[0256] The modified cells can be introduced into an organism. These cells can be derived from the same organism (e.g., a human) in the case of autologous cell transplantation, where the cells are modified using an ex vivo approach. In some embodiments, the cells are derived from another organism within the same species (e.g., another human) in the case of allogeneic cell transplantation.

[0257] XIV. Kit Some aspects of the present disclosure provide kits comprising the deaminase or fusion protein of the present invention. In certain embodiments, the present disclosure provides kits comprising the fusion protein and a guide RNA. In addition, in some embodiments, the kits include suitable reagents, buffers, and / or instructions for using the fusion protein, for example, for in vitro or in vivo DNA or RNA editing. In some embodiments, the kits include instructions for designing and using suitable gRNAs for targeted editing of nucleic acid sequences.

[0258] The articles "a" and "an" are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, "a polypeptide" means one or more polypeptides.

[0259] All publications and patent applications mentioned in this specification are indicative of the level of those skilled in the art to which this disclosure pertains. All publications and patent applications are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference.

[0260] Although the foregoing invention has been described in some detail by way of illustration and example, for purposes of clarity of understanding, it will be apparent that certain changes and modifications may be practiced within the scope of the appended claims.

[0261] Non-limiting embodiments include the following:

[0262] 1. A fusion protein comprising an RNA-guided nuclease (RGN) and at least one heterologous polypeptide inserted therein, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699. 2. The fusion protein of embodiment 1, wherein the RGN comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699. 3. The fusion protein of embodiment 1, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NOs: 1 to 4, 49 to 162, 435, 575, 576, 698, and 699. 4. The fusion protein of embodiment 1, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1 to 4 and 57 to 139. 5. The fusion protein of embodiment 1, wherein the RGN comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 1 to 4 and 57 to 139. 6. The fusion protein of embodiment 1, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NOs: 1 to 4 and 57 to 139. 7. The fusion protein of any one of embodiments 1 to 6, wherein the heterologous polypeptide is inserted within linker domain 2, wedge domain, RuvC domain, HNH domain, Rec-2 domain, or PAM-interacting domain of the RGN. 8. The fusion protein of embodiment 7, wherein the RuvC domain is a RuvCIII domain. 9. a) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 1, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 30 of SEQ ID NO: 1; ii) an amino acid position corresponding to position 642 of SEQ ID NO: 1; iii) an amino acid position corresponding to position 670 of SEQ ID NO: 1; iv) an amino acid position corresponding to position 737 of SEQ ID NO: 1; v) an amino acid position corresponding to position 772 of SEQ ID NO: 1; vi) an amino acid position corresponding to position 775 of SEQ ID NO: 1; vii) the amino acid position corresponding to position 778 of SEQ ID NO: 1; and viii) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 802 of SEQ ID NO: 1; b) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:2, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 678 of SEQ ID NO: 2; ii) an amino acid position corresponding to position 736 of SEQ ID NO: 2; iii) an amino acid position corresponding to position 778 of SEQ ID NO: 2; iv) an amino acid position corresponding to position 788 of SEQ ID NO: 2, and v) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 922 of SEQ ID NO:2; c) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 3, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 725 of SEQ ID NO: 3; ii) an amino acid position corresponding to position 739 of SEQ ID NO: 3, and iii) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 744 of SEQ ID NO: 3; or d) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 4, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 347 of SEQ ID NO: 4; ii) an amino acid position corresponding to position 524 of SEQ ID NO: 4; iii) an amino acid position corresponding to position 666 of SEQ ID NO: 4; iv) an amino acid position corresponding to position 680 of SEQ ID NO: 4; v) an amino acid position corresponding to position 740 of SEQ ID NO: 4; vi) an amino acid position corresponding to position 785 of SEQ ID NO: 4; vii) an amino acid position corresponding to position 910 of SEQ ID NO: 4; and viii) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 1077 of SEQ ID NO: 4; or e) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 131, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 766 of SEQ ID NO: 131, and ii) The fusion protein according to any one of claims 1 to 8, wherein the fusion protein is inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 806 of SEQ ID NO: 131. 10. The fusion protein of any one of embodiments 1 to 9, wherein the heterologous polypeptide comprises a base-editing polypeptide or a prime-editing polypeptide. 11. The fusion protein of embodiment 10, wherein the prime editing polypeptide comprises a DNA polymerase. 12. The fusion protein of embodiment 10, wherein the prime editing polypeptide is a reverse transcriptase. 13. The fusion protein of embodiment 10, wherein the base-editing polypeptide comprises a deaminase. 14. The fusion protein of embodiment 13, wherein the deaminase is a cytosine deaminase or an adenine deaminase. 15. The fusion protein of embodiment 13 or 14, wherein the fusion protein is a base editor fusion protein, and the base editor fusion protein has improved editing activity compared to a parent base editor comprising the deaminase fused to the amino terminus of the RGN. 16. The at least one heterologous polypeptide is located within an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 1; a) the amino acid position corresponding to position 642 of SEQ ID NO: 1; b) an amino acid position corresponding to position 670 of SEQ ID NO: 1; c) an amino acid position corresponding to position 737 of SEQ ID NO: 1; d) an amino acid position corresponding to position 772 of SEQ ID NO: 1; e) the amino acid position corresponding to position 775 of SEQ ID NO: 1, and f) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 778 of SEQ ID NO: 1; 16. The fusion protein of any one of embodiments 13 to 15, wherein the fusion protein is a base editor fusion protein, and the base editor fusion protein has a shifted editing window compared to a parent base editor comprising the deaminase fused to the amino terminus of the RGN. 17. The fusion protein of any one of embodiments 13 to 16, wherein the deaminase lacks the first and last amino acid residues compared to the parent deaminase from which the deaminase is derived. 18. The fusion protein of any one of embodiments 13-17, wherein the deaminase comprises an adenine deaminase and the fusion protein is an adenine base editor (ABE) fusion protein. 19. The fusion protein of embodiment 18, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723. 20. The fusion protein of embodiment 18, wherein the adenine deaminase comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723. 21. The fusion protein of embodiment 18, wherein the adenine deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723. 22. The fusion protein of embodiment 18, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 5 or 6. 23. The fusion protein of embodiment 18, wherein the adenine deaminase comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 5 or 6. 24. The fusion protein of embodiment 18, wherein the adenine deaminase comprises the amino acid sequence of SEQ ID NO: 5 or 6. 25. The fusion protein of embodiment 18, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650. 26. The fusion protein of embodiment 18, wherein the ABE fusion protein comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650. 27. The fusion protein of embodiment 18, wherein the ABE fusion protein comprises an amino acid sequence set forth in any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650. 28. The fusion protein of embodiment 18, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parent ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN. 29. The fusion protein of embodiment 18, wherein the ABE fusion protein has at least 95% sequence identity to any one of SEQ ID NOs: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parent ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN. 30. The fusion protein of embodiment 18, wherein the ABE fusion protein comprises an amino acid sequence set forth in any one of SEQ ID NOs: 8 to 10 and 14 to 16. 31. The fusion protein of any one of embodiments 1 to 30, wherein said RGN is a nickase or nuclease-inactive RGN. 32. The fusion protein of embodiment 31, wherein the RGN nickase comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 49-56, 698, and 699, and retains nickase activity. 33. The fusion protein of embodiment 31, wherein the RGN nickase comprises the amino acid sequence of any one of SEQ ID NOs: 49-56, 698, and 699. 34. A fusion protein according to one of embodiments 1 to 33, wherein said RGN and said heterologous polypeptide are directly fused to each other without a linker sequence in between. 35. The fusion protein of any of embodiments 1 to 34, wherein the fusion protein further comprises at least one nuclear localization signal (NLS). 36. A ribonucleoprotein (RNP) complex comprising a fusion protein according to any one of embodiments 1 to 35 and 232 to 243 and a guide RNA bound to the fusion protein. 37. A nucleic acid molecule comprising a polynucleotide encoding a fusion protein, wherein the fusion protein comprises an RNA-guided nuclease (RGN) and at least one heterologous polypeptide inserted therein, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699. 38. The nucleic acid molecule of embodiment 37, wherein the RGN comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699. 39. The nucleic acid molecule of embodiment 37, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699. 40. The nucleic acid molecule of embodiment 37, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-4, 57-139. 41. The nucleic acid molecule of embodiment 37, wherein the RGN comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 1-4, 57-139. 42. The nucleic acid molecule of embodiment 37, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NOs: 1 to 4, 57 to 139. 43. The nucleic acid molecule of any one of embodiments 1 to 6, wherein the heterologous polypeptide is inserted within linker domain 2, wedge domain, RuvC domain, HNH domain, Rec-2 domain, or PAM-interacting domain of the RGN. 44. The nucleic acid molecule of embodiment 7, wherein the RuvC domain is a RuvCIII domain. 45. a) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 1, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 30 of SEQ ID NO: 1; ii) an amino acid position corresponding to position 642 of SEQ ID NO: 1; iii) an amino acid position corresponding to position 670 of SEQ ID NO: 1; iv) an amino acid position corresponding to position 737 of SEQ ID NO: 1; v) an amino acid position corresponding to position 772 of SEQ ID NO: 1; vi) an amino acid position corresponding to position 775 of SEQ ID NO: 1; vii) the amino acid position corresponding to position 778 of SEQ ID NO: 1; and viii) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 802 of SEQ ID NO: 1; b) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:2, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 678 of SEQ ID NO: 2; ii) an amino acid position corresponding to position 736 of SEQ ID NO: 2; iii) an amino acid position corresponding to position 778 of SEQ ID NO: 2; iv) an amino acid position corresponding to position 788 of SEQ ID NO: 2, and v) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 922 of SEQ ID NO:2; c) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 3, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 725 of SEQ ID NO: 3; ii) an amino acid position corresponding to position 739 of SEQ ID NO: 3, and iii) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 744 of SEQ ID NO: 3; or d) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 4, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 347 of SEQ ID NO: 4; ii) an amino acid position corresponding to position 524 of SEQ ID NO: 4; iii) an amino acid position corresponding to position 666 of SEQ ID NO: 4; iv) an amino acid position corresponding to position 680 of SEQ ID NO: 4; v) an amino acid position corresponding to position 740 of SEQ ID NO: 4; vi) an amino acid position corresponding to position 785 of SEQ ID NO: 4; vii) an amino acid position corresponding to position 910 of SEQ ID NO: 4; and viii) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 1077 of SEQ ID NO: 4; or e) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 131, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 766 of SEQ ID NO: 131, and ii) the nucleic acid molecule according to any one of embodiments 37 to 44, which is inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 806 of SEQ ID NO: 131. 46. ​​The nucleic acid molecule of embodiment 45, wherein the heterologous polypeptide comprises a base-editing polypeptide or a prime-editing polypeptide. 47. The nucleic acid molecule of embodiment 46, wherein the prime editing polypeptide comprises a DNA polymerase. 48. The nucleic acid molecule of embodiment 46, wherein the prime editing polypeptide comprises a reverse transcriptase. 49. The nucleic acid molecule of embodiment 46, wherein the base-editing polypeptide comprises a deaminase. 50. The nucleic acid molecule of embodiment 49, wherein the deaminase is a cytosine deaminase or an adenine deaminase. 51. The nucleic acid molecule of embodiment 49 or 50, wherein the fusion protein is a base editor fusion protein, and the base editor fusion protein has improved editing activity compared to a parent base editor comprising the deaminase fused to the amino terminus of the RGN. 52. The at least one heterologous polypeptide is located within the RGN, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 1; a) the amino acid position corresponding to position 642 of SEQ ID NO: 1; b) an amino acid position corresponding to position 670 of SEQ ID NO: 1; c) an amino acid position corresponding to position 737 of SEQ ID NO: 1; d) an amino acid position corresponding to position 772 of SEQ ID NO: 1; e) the amino acid position corresponding to position 775 of SEQ ID NO: 1, and f) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 778 of SEQ ID NO: 1; 52. The nucleic acid molecule of any one of embodiments 49 to 51, wherein the fusion protein is a base editor fusion protein, and the base editor fusion protein has a shifted editing window compared to a parent base editor comprising the deaminase fused to the amino terminus of the RGN. 53. The nucleic acid molecule of any one of embodiments 49 to 52, wherein the deaminase lacks the first and last amino acid residues compared to the parent deaminase from which the deaminase is derived. 54. The nucleic acid molecule of any one of embodiments 49 to 53, wherein the deaminase comprises an adenine deaminase and is an adenine base editor (ABE) fusion protein. 55. The nucleic acid molecule of embodiment 54, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723. 56. The nucleic acid molecule of embodiment 54, wherein the adenine deaminase comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723. 57. The nucleic acid molecule of embodiment 54, wherein the adenine deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723. 58. The nucleic acid molecule of embodiment 54, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 5 or 6. 59. The nucleic acid molecule of embodiment 54, wherein the adenine deaminase comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 5 or 6. 60. The nucleic acid molecule of embodiment 54, wherein the adenine deaminase comprises the amino acid sequence of SEQ ID NO: 5 or 6. 61. The nucleic acid molecule of embodiment 54, wherein the fusion protein comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650. 62. The nucleic acid molecule of embodiment 54, wherein the fusion protein comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650. 63. The nucleic acid molecule of embodiment 54, wherein the ABE fusion protein comprises an amino acid sequence set forth in any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650. 64. The nucleic acid molecule of embodiment 54, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parent ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN. 65. The nucleic acid molecule of embodiment 54, wherein the ABE fusion protein comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parent ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN. 66. The nucleic acid molecule of embodiment 54, wherein the ABE fusion protein comprises an amino acid sequence set forth in any one of SEQ ID NOs: 8 to 10 and 14 to 16. 67. The nucleic acid molecule of any one of embodiments 37 to 66, wherein said RGN is a nickase or nuclease-inactive RGN. 68. The nucleic acid molecule of embodiment 67, wherein the RGN nickase comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 49-56, 698, and 699, and retains nickase activity. 69. The nucleic acid molecule of embodiment 67, wherein the RGN nickase comprises the amino acid sequence of any one of SEQ ID NOs: 49-56, 698, and 699. 70. The nucleic acid molecule of any one of embodiments 37 to 69, wherein said RGN and said heterologous polypeptide are directly fused to each other without a linker sequence between them. 71. The nucleic acid molecule of any of embodiments 37 to 70, wherein the fusion protein further comprises at least one nuclear localization signal (NLS). 72. The nucleic acid molecule of any one of embodiments 37 to 71, wherein the nucleic acid molecule is codon-optimized for expression in a eukaryotic cell. 73. The nucleic acid molecule of any one of embodiments 37 to 71, wherein the nucleic acid molecule is codon-optimized for expression in a prokaryotic cell. 74. A vector comprising a nucleic acid molecule according to any one of embodiments 37 to 73 and 244 to 255. 75. The vector of embodiment 74, further comprising at least one nucleotide sequence encoding a guide RNA capable of hybridizing to a non-target strand of the target sequence of the RGN. 76. The vector of embodiment 74, wherein the guide RNA is a single guide RNA (sgRNA). 77. The vector of embodiment 74, wherein the guide RNA is a dual guide RNA. 78. A cell comprising a fusion protein according to any one of embodiments 1 to 35 and 232 to 243, or an RNP complex according to embodiment 36. 79. A cell comprising a fusion protein according to any one of embodiments 1 to 35 and 232 to 243, wherein the cell further comprises a guide RNA. 80. A cell comprising a nucleic acid molecule according to any one of embodiments 37 to 73 and 244 to 255. 81. A cell comprising a vector according to any one of embodiments 74 to 77. 82. The cell according to any one of embodiments 78 to 81, wherein the cell is a prokaryotic cell. 83. The cell according to any one of embodiments 78 to 81, wherein the cell is a eukaryotic cell. 84. The cell of embodiment 83, wherein the eukaryotic cell is a mammalian cell. 85. The cell of embodiment 84, wherein the mammalian cell is a human cell. 86. The cell of embodiment 85, wherein the human cell is an immune cell. 87. The cell of embodiment 86, wherein the immune cell is a stem cell. 88. The cell of embodiment 87, wherein the stem cell is an induced pluripotent stem cell. 89. The cell of embodiment 83, wherein the eukaryotic cell is an insect or avian cell. 90. The cell of embodiment 83, wherein the eukaryotic cell is a fungal cell. 91. The cell of embodiment 83, wherein the eukaryotic cell is a plant cell. 92. A plant comprising a cell according to embodiment 91. 93. A seed comprising the cells of embodiment 91. 94. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a fusion protein according to any one of embodiments 1 to 35 and 232 to 243, an RNP complex according to embodiment 36, a nucleic acid molecule according to any one of embodiments 37 to 73 and 244 to 255, a vector according to any one of embodiments 74 to 77, or a cell according to any one of embodiments 84 to 88. 95. A method for producing a fusion protein, comprising culturing a cell according to any one of embodiments 78 to 91 under conditions in which the fusion protein is expressed. 96. A method for producing a fusion protein, comprising introducing a nucleic acid molecule according to any one of embodiments 37 to 73 and 237 to 243 or a vector according to any one of embodiments 74 to 77 into a cell, and culturing the cell under conditions in which the fusion protein is expressed. 97. The method of embodiment 95 or 96, further comprising purifying the fusion protein. 98. A method for producing an RGN fusion ribonucleoprotein complex, comprising: introducing into a cell a nucleic acid molecule described in any one of embodiments 37 to 73 and 237 to 243, and a nucleic acid comprising an expression cassette encoding a guide RNA, or a vector described in any one of embodiments 74 to 77; and culturing the cells under conditions in which the fusion protein and gRNA are expressed and form an RGN fusion ribonucleoprotein complex. 99. The method of embodiment 98, further comprising purifying the RGN ribonucleoprotein complex. 100. A system for localizing a heterologous polypeptide to a target DNA molecule containing a target DNA sequence, comprising: a) a fusion protein or a nucleotide sequence encoding the fusion protein, the fusion protein comprising an RNA-guided nuclease (RGN) and at least one heterologous polypeptide inserted therein, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699, or a nucleotide sequence encoding the fusion protein; b) one or more guide RNAs, or one or more nucleotide sequences encoding one or more guide RNAs, capable of hybridizing to a non-target strand of the target DNA sequence; A system wherein one or more guide RNAs are capable of forming a complex with a fusion protein to guide the fusion protein to bind to the target DNA sequence. 101. The system of embodiment 100, wherein the RGN comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699. 102. The system of embodiment 100, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699. 103. The system according to embodiment 100, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NOs: 1-4 and 57-139. 104. The system according to embodiment 100, wherein the RGN comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 1-4 and 57-139. 105. The system according to embodiment 100, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NOs: 1 to 4 and 57 to 139. 106. A system according to any one of embodiments 100 to 105, wherein the heterologous polypeptide is inserted within linker domain 2, wedge domain, RuvC domain, HNH domain, Rec-2 domain, or PAM interaction domain of the RGN. 107. The system of embodiment 106, wherein the RuvC domain is a RuvCIII domain. 108. i) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 1, and the at least one heterologous polypeptide is located within the RGN: A. The amino acid position corresponding to position 30 of SEQ ID NO: 1; B. The amino acid position corresponding to position 642 of SEQ ID NO: 1; C. The amino acid position corresponding to position 670 of SEQ ID NO: 1; D. The amino acid position corresponding to position 737 of SEQ ID NO: 1; E. an amino acid position corresponding to position 772 of SEQ ID NO: 1; F. an amino acid position corresponding to position 775 of SEQ ID NO: 1; G. the amino acid position corresponding to position 778 of SEQ ID NO: 1, and H. inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 802 of SEQ ID NO: 1; ii) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:2, and the at least one heterologous polypeptide is located within the RGN: A. The amino acid position corresponding to position 678 of SEQ ID NO:2; B. The amino acid position corresponding to position 736 of SEQ ID NO:2; C. The amino acid position corresponding to position 778 of SEQ ID NO:2; D. the amino acid position corresponding to position 788 of SEQ ID NO:2, and E. inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 922 of SEQ ID NO:2; iii) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 3, and the at least one heterologous polypeptide is located within the RGN: A. The amino acid position corresponding to position 725 of SEQ ID NO:3; B. The amino acid position corresponding to position 739 of SEQ ID NO:3, and C. inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 744 of SEQ ID NO:3; iv) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 4, and the at least one heterologous polypeptide is located within the RGN: A. The amino acid position corresponding to position 347 of SEQ ID NO:4; B. The amino acid position corresponding to position 524 of SEQ ID NO:4; C. The amino acid position corresponding to position 666 of SEQ ID NO:4; D. The amino acid position corresponding to position 680 of SEQ ID NO:4; E. an amino acid position corresponding to position 740 of SEQ ID NO:4; F. an amino acid position corresponding to position 785 of SEQ ID NO:4; G. the amino acid position corresponding to position 910 of SEQ ID NO:4, and H. inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 1077 of SEQ ID NO:4; or v) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 131, and the at least one heterologous polypeptide is located within the RGN: A. an amino acid position corresponding to position 766 of SEQ ID NO: 131, and B. The system of any one of embodiments 100 to 107, wherein the system is inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 806 of SEQ ID NO: 131. 109. The system of any one of embodiments 100 to 108, wherein the heterologous polypeptide comprises a base-editing polypeptide or a prime-editing polypeptide. 110. The system of embodiment 109, wherein the prime editing polypeptide comprises a DNA polymerase. 111. The system of embodiment 109, wherein the prime editing polypeptide comprises a reverse transcriptase. 112. The system of embodiment 109, wherein the base-editing polypeptide comprises a deaminase. 113. The system according to embodiment 112, wherein the deaminase is a cytosine deaminase or an adenine deaminase. 114. The system of embodiment 112 or 113, wherein the fusion protein is a base editor fusion protein, and the base editor fusion protein has improved editing activity compared to a parent base editor comprising the deaminase fused to the amino terminus of the RGN. 115. The at least one heterologous polypeptide is located within the RGN, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 1; a) the amino acid position corresponding to position 642 of SEQ ID NO: 1; b) an amino acid position corresponding to position 670 of SEQ ID NO: 1; c) an amino acid position corresponding to position 737 of SEQ ID NO: 1; d) an amino acid position corresponding to position 772 of SEQ ID NO: 1; e) the amino acid position corresponding to position 775 of SEQ ID NO: 1, and f) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 778 of SEQ ID NO: 1; 115. The system of any one of embodiments 112-114, wherein the fusion protein is a base editor fusion protein, and wherein the base editor fusion protein has a shifted editing window compared to a parent base editor comprising the deaminase fused to the amino terminus of the RGN. 116. The system of any one of embodiments 112-115, wherein the deaminase lacks the first and last amino acid residues compared to the parent deaminase from which the deaminase is derived. 117. The system of any one of embodiments 112-116, wherein the deaminase comprises an adenine deaminase and the fusion protein is an adenine base editor (ABE) fusion protein. 118. The system of embodiment 117, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723. 119. The system of embodiment 117, wherein the adenine deaminase comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723. 120. The system of embodiment 117, wherein the adenine deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723. 121. The system of embodiment 117, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 5 or 6. 122. The system of embodiment 117, wherein the adenine deaminase comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 5 or 6. 123. The system of embodiment 117, wherein the adenine deaminase comprises the amino acid sequence of SEQ ID NO: 5 or 6. 124. The system of embodiment 117, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650. 125. The system of embodiment 117, wherein the ABE fusion protein comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650. 126. The system of embodiment 117, wherein the ABE fusion protein comprises an amino acid sequence set forth in any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650. 127. The system of embodiment 117, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parent ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN. 128. The system of embodiment 117, wherein the ABE fusion protein comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parent ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN. 129. The system of embodiment 117, wherein the ABE fusion protein comprises an amino acid sequence set forth in any one of SEQ ID NOs: 8 to 10 and 14 to 16. 130. The system according to any one of embodiments 100 to 129, wherein the RGN is a nickase or nuclease-inactive RGN. 131. The system of embodiment 130, wherein the RGN nickase comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 49-56, 698, and 699, and retains nickase activity. 132. The system of embodiment 130, wherein the RGN nickase comprises the amino acid sequence of any one of SEQ ID NOs: 49-56, 698, and 699. 133. The system according to any one of embodiments 100 to 132, wherein said RGN and said heterologous polypeptide are directly fused to each other without a linker sequence between them. 134. The system according to any of embodiments 100 to 133, wherein the fusion protein further comprises at least one nuclear localization signal (NLS). 135. The system according to any one of embodiments 100 to 134, wherein the nucleotide sequence encoding the fusion protein is codon-optimized for expression in eukaryotic cells. 136. The system according to any one of embodiments 100 to 135, wherein at least one of the nucleotide sequences encoding one or more guide RNAs and the nucleotide sequence encoding the fusion protein is operably linked to a promoter heterologous to the nucleotide sequence. 137. A system according to any one of embodiments 100 to 136, wherein the target DNA sequence is a eukaryotic target DNA sequence. 138. A system according to any one of embodiments 100 to 137, wherein the target DNA sequence is located adjacent to a protospacer adjacent motif (PAM) recognized by RGN. 139. A system according to any one of embodiments 100 to 138, wherein the nucleotide sequence encoding one or more guide RNAs and the nucleotide sequence encoding the fusion protein are located on one vector. 140. A cell comprising a system according to any one of embodiments 100 to 139 and 256 to 267. 141. The cell of embodiment 140, wherein the cell is a eukaryotic cell. 142. The cell according to embodiment 141, wherein the eukaryotic cell is a mammalian cell. 143. The cell of embodiment 142, wherein the mammalian cell is a human cell. 144. The cell of embodiment 143, wherein the human cell is an immune cell. 145. The cell of embodiment 144, wherein the immune cell is a stem cell. 146. The cell according to embodiment 145, wherein the stem cell is an induced pluripotent stem cell. 147. The cell according to embodiment 141, wherein the eukaryotic cell is an insect cell. 148. The cell according to embodiment 141, wherein the eukaryotic cell is a plant cell. 149. A plant comprising a cell according to embodiment 148. 150. A seed comprising cells according to embodiment 148. 151. The cell according to embodiment 140, wherein the cell is a prokaryotic cell. 152. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a system according to any one of embodiments 100 to 139 and 256 to 267, or a cell according to any one of embodiments 142 to 146. 153. A method for localizing a heterologous polypeptide to a target DNA molecule comprising a target DNA sequence, the method comprising delivering a system described in any one of embodiments 100 to 139 and 256 to 267 to the target DNA molecule or to a cell comprising the target DNA molecule. 154. A method for modifying a target DNA molecule comprising a target DNA sequence, the method comprising delivering a system according to any one of embodiments 100 to 139 and 256 to 267 to the target DNA molecule or to a cell comprising the target DNA molecule. 155. The method of embodiment 154, wherein the deaminase comprises adenine deaminase, and the modified target DNA molecule comprises an A>N mutation of at least one nucleotide within the target DNA molecule, wherein N is C, G, or T. 156. The method of embodiment 155, wherein the modified target DNA molecule comprises an A>G mutation of at least one nucleotide in the target DNA molecule. 157. A method for localizing a heterologous polypeptide to a target DNA molecule comprising a target DNA sequence, comprising: a) i) one or more guide RNAs capable of hybridizing to the non-target strand of the target DNA sequence; and ii) assembling a ribonucleotide complex in vitro by combining a fusion protein comprising an RNA-guided nuclease (RGN) and at least one heterologous polypeptide inserted therein, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699; Under conditions favorable for the formation of a ribonucleotide complex, b) contacting the target DNA molecule or a cell containing the target DNA molecule with an in vitro assembled ribonucleotide complex; A method in which one or more guide RNAs hybridize to a non-target strand of a target DNA sequence, thereby directing the fusion protein to bind to the target DNA sequence. 158. The method of embodiment 157, wherein the RGN comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699. 159. The method of embodiment 157, wherein the RGN comprises an amino acid sequence having the amino acid sequence of any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699. 160. The method of embodiment 157, wherein the RGN comprises at least 90% sequence identity with any one of SEQ ID NOs: 1-4 and 57-139. 161. The method of embodiment 157, wherein the RGN comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 1-4 and 57-139. 162. The method of embodiment 157, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NOs: 1 to 4 and 57 to 139. 163. The method of any one of embodiments 157 to 162, wherein the heterologous polypeptide is inserted within linker domain 2, wedge domain, RuvC domain, HNH domain, Rec-2 domain, or PAM interaction domain of the RGN. 164. The method of embodiment 163, wherein the RuvC domain is a RuvCIII domain. 165. a) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 1, and the at least one heterologous polypeptide is located within the RGN: I) the amino acid position corresponding to position 30 of SEQ ID NO: 1; II) the amino acid position corresponding to position 642 of SEQ ID NO: 1; III) an amino acid position corresponding to position 670 of SEQ ID NO: 1; IV) an amino acid position corresponding to position 737 of SEQ ID NO: 1; V) an amino acid position corresponding to position 772 of SEQ ID NO: 1; VI) an amino acid position corresponding to position 775 of SEQ ID NO: 1; VII) the amino acid position corresponding to position 778 of SEQ ID NO: 1; and VIII) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 802 of SEQ ID NO: 1; b) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:2, and the at least one heterologous polypeptide is located within the RGN: I) the amino acid position corresponding to position 678 of SEQ ID NO: 2; II) the amino acid position corresponding to position 736 of SEQ ID NO: 2; III) an amino acid position corresponding to position 778 of SEQ ID NO: 2; IV) an amino acid position corresponding to position 788 of SEQ ID NO: 2; and V) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 922 of SEQ ID NO: 2; c) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 3, and the at least one heterologous polypeptide is located within the RGN: I) the amino acid position corresponding to position 725 of SEQ ID NO: 3; II) an amino acid position corresponding to position 739 of SEQ ID NO: 3; III) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 744 of SEQ ID NO: 3; d) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 4, and the at least one heterologous polypeptide is located within the RGN: I) the amino acid position corresponding to position 347 of SEQ ID NO: 4; II) the amino acid position corresponding to position 524 of SEQ ID NO: 4; III) an amino acid position corresponding to position 666 of SEQ ID NO: 4; IV) an amino acid position corresponding to position 680 of SEQ ID NO: 4; V) an amino acid position corresponding to position 740 of SEQ ID NO: 4; VI) an amino acid position corresponding to position 785 of SEQ ID NO: 4; VII) the amino acid position corresponding to position 910 of SEQ ID NO: 4, and VIII) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 1077 of SEQ ID NO: 4, or e) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 131, and the at least one heterologous polypeptide is located within the RGN: I) the amino acid position corresponding to position 766 of SEQ ID NO: 131, and II) the amino acid position corresponding to position 806 of SEQ ID NO: 131, immediately after the amino acid position selected from the group consisting of: the method according to any one of embodiments 157 to 164. 166. The method of any one of embodiments 157 to 165, wherein the heterologous polypeptide comprises a base-editing polypeptide or a prime-editing polypeptide, and the method further comprises modifying the target DNA molecule to produce a modified target DNA molecule. 167. The method of embodiment 166, wherein the target DNA molecule contains a mutation that causes a disease or disorder, and modification of the target DNA molecule corrects the causative mutation. 168. The method of embodiment 167, wherein correcting the causative mutation comprises correcting a nonsense mutation. 169. The method of any one of embodiments 166-168, wherein the prime editing polypeptide comprises a DNA polymerase. 170. The method of any one of embodiments 166-168, wherein the prime editing polypeptide ...

Claims

1. A fusion protein comprising an RNA-guided nuclease (RGN) and at least one heterologous polypeptide inserted therein, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699.

2. The fusion protein of claim 1, wherein the RGN comprises any one of the amino acid sequences of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699.

3. The fusion protein of claim 1 or 2, wherein the heterologous polypeptide is inserted within linker domain 2, wedge domain, RuvC domain, HNH domain, Rec-2 domain, or PAM-interacting domain of the RGN.

4. a) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 1, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 30 of SEQ ID NO: 1; ii) an amino acid position corresponding to position 642 of SEQ ID NO: 1; iii) an amino acid position corresponding to position 670 of SEQ ID NO: 1; iv) an amino acid position corresponding to position 737 of SEQ ID NO: 1; v) an amino acid position corresponding to position 772 of SEQ ID NO: 1; vi) an amino acid position corresponding to position 775 of SEQ ID NO: 1; vii) an amino acid position corresponding to position 778 of SEQ ID NO: 1; and viii) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 802 of SEQ ID NO: 1; or b) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:2, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 678 of SEQ ID NO:2; ii) an amino acid position corresponding to position 736 of SEQ ID NO:2; iii) an amino acid position corresponding to position 778 of SEQ ID NO:2; iv) an amino acid position corresponding to position 788 of SEQ ID NO:2; and v) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 922 of SEQ ID NO:2; c) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 3, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 725 of SEQ ID NO: 3; ii) an amino acid position corresponding to position 739 of SEQ ID NO: 3; and iii) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 744 of SEQ ID NO: 3; or d) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 4, and the at least one heterologous polypeptide is located within the RGN: i) an amino acid position corresponding to position 347 of SEQ ID NO:4; ii) an amino acid position corresponding to position 524 of SEQ ID NO:4; iii) an amino acid position corresponding to position 666 of SEQ ID NO:4; iv) an amino acid position corresponding to position 680 of SEQ ID NO: 4; v) an amino acid position corresponding to position 740 of SEQ ID NO:4; vi) an amino acid position corresponding to position 785 of SEQ ID NO:4; vii) an amino acid position corresponding to position 910 of SEQ ID NO:4; and viii) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 1077 of SEQ ID NO: 4; or e) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 131, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 766 of SEQ ID NO: 131, and ii) the amino acid position corresponding to position 806 of SEQ ID NO:

131.

5. The fusion protein of any one of claims 1 to 4, wherein the heterologous polypeptide comprises a base-editing polypeptide or a prime-editing polypeptide.

6. The fusion protein of Claim 5, wherein the base-editing polypeptide comprises a deaminase.

7. 7. The fusion protein of claim 6, wherein the fusion protein is a base editor fusion protein, and the base editor fusion protein has improved editing activity compared to a parent base editor comprising the deaminase fused to the amino terminus of the RGN.

8. wherein the at least one heterologous polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 1; a) the amino acid position corresponding to position 642 of SEQ ID NO: 1; b) an amino acid position corresponding to position 670 of SEQ ID NO: 1; c) an amino acid position corresponding to position 737 of SEQ ID NO: 1; d) an amino acid position corresponding to position 772 of SEQ ID NO: 1; e) an amino acid position corresponding to position 775 of SEQ ID NO: 1; and f) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 778 of SEQ ID NO: 1; 8. The fusion protein of claim 6 or 7, wherein the fusion protein is a base editor fusion protein, and the base editor fusion protein has a shifted editing window compared to a parent base editor comprising the deaminase fused to the amino terminus of the RGN.

9. 9. The fusion protein of any one of claims 6 to 8, wherein the deaminase lacks the first and last amino acid residues compared to the parent deaminase from which the deaminase is derived.

10. The fusion protein of any one of claims 6 to 9, wherein the deaminase comprises adenine deaminase and the fusion protein is an adenine base editor (ABE) fusion protein.

11. 11. The fusion protein of claim 10, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723.

12. The fusion protein of claim 10, wherein the adenine deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723.

13. The fusion protein of claim 10, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650.

14. The fusion protein of claim 10, wherein the ABE fusion protein comprises an amino acid sequence set forth in any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650.

15. 11. The fusion protein of claim 10, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parent ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN.

16. The fusion protein of any one of claims 1 to 15, wherein the RGN is a nickase or nuclease inactive RGN.

17. 17. The fusion protein of claim 16, wherein the RGN nickase comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 49-56, 698, and 699 and retains nickase activity.

18. 17. The fusion protein of claim 16, wherein the RGN nickase comprises the amino acid sequence of any one of SEQ ID NOs: 49-56, 698, and 699.

19. A ribonucleoprotein (RNP) complex comprising the fusion protein of any one of claims 1 to 18, 110, and 111, and a guide RNA bound to the fusion protein.

20. A nucleic acid molecule comprising a polynucleotide encoding a fusion protein, wherein the fusion protein comprises an RNA-guided nuclease (RGN) and at least one heterologous polypeptide inserted therein, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699.

21. 21. The nucleic acid molecule of claim 20, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699.

22. 22. The nucleic acid molecule of claim 20 or 21, wherein the heterologous polypeptide is inserted within linker domain 2, wedge domain, RuvC domain, HNH domain, Rec-2 domain, or PAM-interacting domain of the RGN.

23. a) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 1, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 30 of SEQ ID NO: 1; ii) an amino acid position corresponding to position 642 of SEQ ID NO: 1; iii) an amino acid position corresponding to position 670 of SEQ ID NO: 1; iv) an amino acid position corresponding to position 737 of SEQ ID NO: 1; v) an amino acid position corresponding to position 772 of SEQ ID NO: 1; vi) an amino acid position corresponding to position 775 of SEQ ID NO: 1; vii) an amino acid position corresponding to position 778 of SEQ ID NO: 1; and viii) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 802 of SEQ ID NO: 1; or b) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:2, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 678 of SEQ ID NO:2; ii) an amino acid position corresponding to position 736 of SEQ ID NO:2; iii) an amino acid position corresponding to position 778 of SEQ ID NO:2; iv) an amino acid position corresponding to position 788 of SEQ ID NO:2; and v) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 922 of SEQ ID NO:2; c) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 3, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 725 of SEQ ID NO: 3; ii) an amino acid position corresponding to position 739 of SEQ ID NO: 3; and iii) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 744 of SEQ ID NO: 3; or d) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 4, and the at least one heterologous polypeptide is located within the RGN: i) an amino acid position corresponding to position 347 of SEQ ID NO:4; ii) an amino acid position corresponding to position 524 of SEQ ID NO:4; iii) an amino acid position corresponding to position 666 of SEQ ID NO:4; iv) an amino acid position corresponding to position 680 of SEQ ID NO: 4; v) an amino acid position corresponding to position 740 of SEQ ID NO:4; vi) an amino acid position corresponding to position 785 of SEQ ID NO:4; vii) an amino acid position corresponding to position 910 of SEQ ID NO:4; and viii) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 1077 of SEQ ID NO: 4; or e) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 131, and the at least one heterologous polypeptide is located within the RGN: i) the amino acid position corresponding to position 766 of SEQ ID NO: 131, and ii) the nucleic acid molecule according to any one of claims 20 to 22, wherein the nucleic acid molecule is inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 806 of SEQ ID NO:

131.

24. The nucleic acid molecule of any one of claims 20 to 23, wherein the heterologous polypeptide comprises a base-editing polypeptide or a prime-editing polypeptide.

25. 25. The nucleic acid molecule of Claim 24, wherein the base-editing polypeptide comprises a deaminase.

26. 26. The nucleic acid molecule of Claim 25, wherein the fusion protein is a base editor fusion protein, and the base editor fusion protein has improved editing activity compared to a parent base editor comprising the deaminase fused to the amino terminus of the RGN.

27. Within the RGN, the at least one heterologous polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 1; a) the amino acid position corresponding to position 642 of SEQ ID NO: 1; b) an amino acid position corresponding to position 670 of SEQ ID NO: 1; c) an amino acid position corresponding to position 737 of SEQ ID NO: 1; d) an amino acid position corresponding to position 772 of SEQ ID NO: 1; e) an amino acid position corresponding to position 775 of SEQ ID NO: 1; and f) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 778 of SEQ ID NO: 1; 27. The nucleic acid molecule of Claim 25 or 26, wherein the fusion protein is a base editor fusion protein, and the base editor fusion protein has a shifted editing window compared to a parent base editor comprising the deaminase fused to the amino terminus of the RGN.

28. 28. The nucleic acid molecule of any one of claims 25 to 27, wherein the deaminase lacks the first and last amino acid residues compared to the parent deaminase from which the deaminase is derived.

29. 29. The nucleic acid molecule of any one of claims 25 to 28, wherein the deaminase comprises an adenine deaminase and is an adenine base editor (ABE) fusion protein.

30. 30. The nucleic acid molecule of claim 29, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723.

31. 30. The nucleic acid molecule of claim 29, wherein the adenine deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723.

32. 30. The nucleic acid molecule of claim 29, wherein the fusion protein comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650.

33. 30. The nucleic acid molecule of claim 29, wherein the ABE fusion protein comprises an amino acid sequence set forth in any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650.

34. 30. The nucleic acid molecule of Claim 29, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parent ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN.

35. The nucleic acid molecule of any one of claims 20 to 34, wherein the RGN is a nickase or nuclease inactive RGN.

36. 36. The nucleic acid molecule of claim 35, wherein the RGN nickase comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 49-56, 698, and 699, and retains nickase activity.

37. 36. The nucleic acid molecule of claim 35, wherein the RGN nickase comprises the amino acid sequence of any one of SEQ ID NOs: 49-56, 698, and 699.

38. A vector comprising a nucleic acid molecule according to any one of claims 20 to 37 and 120 to 126.

39. 39. The vector of claim 38, further comprising at least one nucleotide sequence encoding a guide RNA capable of hybridizing to a non-target strand of the target sequence of the RGN.

40. A cell comprising a fusion protein according to any one of claims 1 to 18 and 113 to 119, or an RNP complex according to claim 19.

41. 120. A cell comprising the fusion protein of any one of claims 1 to 18 and 113 to 119, wherein the cell further comprises a guide RNA.

42. A cell comprising a nucleic acid molecule according to any one of claims 20 to 37 and 120 to 126, or a vector according to claim 38 or 39.

43. The cell according to any one of claims 40 to 42, wherein the cell is a mammalian cell.

44. The cell according to any one of claims 40 to 42, wherein the cell is a plant cell.

45. 45. A plant or seed comprising the cell of claim 44.

46. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a fusion protein according to any one of claims 1 to 18 and 113 to 119, an RNP complex according to claim 19, a nucleic acid molecule according to any one of claims 20 to 37 and 120 to 126, a vector according to claim 38 or 39, or a cell according to claim 43.

47. 127. A method for producing an RGN-fusion ribonucleoprotein complex, comprising: introducing into a cell a nucleic acid molecule described in any one of claims 20 to 37 and 120 to 126, and a nucleic acid comprising an expression cassette encoding a guide RNA, or a vector described in claim 38 or 39; and culturing the cell under conditions in which the fusion protein and the gRNA are expressed to form an RGN-fusion ribonucleoprotein complex.

48. 1. A system for localizing a heterologous polypeptide to a target DNA molecule comprising a target DNA sequence, comprising: a) a fusion protein or a nucleotide sequence encoding said fusion protein, said fusion protein comprising an RNA-guided nuclease (RGN) and at least one heterologous polypeptide inserted therein, wherein said RGN comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699; b) one or more guide RNAs capable of hybridizing to non-target strands of the target DNA sequence, or one or more nucleotide sequences encoding the one or more guide RNAs; wherein the one or more guide RNAs are capable of forming a complex with the fusion protein to guide the fusion protein to bind to the target DNA sequence.

49. 49. The system of claim 48, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699.

50. 50. The system of claim 48 or 49, wherein the heterologous polypeptide is inserted within linker domain 2, wedge domain, RuvC domain, HNH domain, Rec-2 domain, or PAM-interacting domain of the RGN.

51. i) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 1, and the at least one heterologous polypeptide is located within the RGN: A. the amino acid position corresponding to position 30 of SEQ ID NO: 1; B. The amino acid position corresponding to position 642 of SEQ ID NO:1; C. an amino acid position corresponding to position 670 of SEQ ID NO: 1; D. The amino acid position corresponding to position 737 of SEQ ID NO:1; E. an amino acid position corresponding to position 772 of SEQ ID NO:1; F. an amino acid position corresponding to position 775 of SEQ ID NO: 1; G. an amino acid position corresponding to position 778 of SEQ ID NO: 1, and H. inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 802 of SEQ ID NO: 1; ii) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:2, and the at least one heterologous polypeptide is located within the RGN: A. the amino acid position corresponding to position 678 of SEQ ID NO:2; B. The amino acid position corresponding to position 736 of SEQ ID NO:2; C. an amino acid position corresponding to position 778 of SEQ ID NO:2; D. an amino acid position corresponding to position 788 of SEQ ID NO:2, and E. inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 922 of SEQ ID NO:2; iii) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 3, and the at least one heterologous polypeptide is located within the RGN: A. the amino acid position corresponding to position 725 of SEQ ID NO:3; B. an amino acid position corresponding to position 739 of SEQ ID NO:3, and C. inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 744 of SEQ ID NO:3; iv) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 4, and the at least one heterologous polypeptide is located within the RGN: A. an amino acid position corresponding to position 347 of SEQ ID NO:4; B. The amino acid position corresponding to position 524 of SEQ ID NO:4; C. an amino acid position corresponding to position 666 of SEQ ID NO:4; D. an amino acid position corresponding to position 680 of SEQ ID NO:4; E. an amino acid position corresponding to position 740 of SEQ ID NO:4; F. the amino acid position corresponding to position 785 of SEQ ID NO:4; G. an amino acid position corresponding to position 910 of SEQ ID NO:4, and H. inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 1077 of SEQ ID NO:4; or v) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 131, and the at least one heterologous polypeptide is located within the RGN: A. an amino acid position corresponding to position 766 of SEQ ID NO: 131, and B. The system of any one of claims 48 to 50, wherein the system is inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 806 of SEQ ID NO:

131.

52. 52. The system of any one of claims 48 to 51, wherein the heterologous polypeptide comprises a base-editing polypeptide or a prime-editing polypeptide.

53. 53. The system of Claim 52, wherein the base-editing polypeptide comprises a deaminase.

54. 54. The system of Claim 53, wherein the fusion protein is a base editor fusion protein, and the base editor fusion protein has improved editing activity compared to a parent base editor comprising the deaminase fused to the amino terminus of the RGN.

55. Within the RGN, the at least one heterologous polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 1; a) the amino acid position corresponding to position 642 of SEQ ID NO: 1; b) an amino acid position corresponding to position 670 of SEQ ID NO: 1; c) an amino acid position corresponding to position 737 of SEQ ID NO: 1; d) an amino acid position corresponding to position 772 of SEQ ID NO: 1; e) an amino acid position corresponding to position 775 of SEQ ID NO: 1; and f) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 778 of SEQ ID NO: 1; 55. The system of claim 53 or 54, wherein the fusion protein is a base editor fusion protein, and the base editor fusion protein has a shifted editing window compared to a parent base editor comprising the deaminase fused to the amino terminus of the RGN.

56. 56. The system of any one of claims 53 to 55, wherein the deaminase lacks the first and last amino acid residues compared to the parent deaminase from which the deaminase is derived.

57. 57. The system of any one of claims 53 to 56, wherein the deaminase comprises an adenine deaminase and the fusion protein is an adenine base editor (ABE) fusion protein.

58. 58. The system of claim 57, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723.

59. 58. The system of claim 57, wherein the adenine deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723.

60. 58. The system of claim 57, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650.

61. 58. The system of claim 57, wherein the ABE fusion protein comprises an amino acid sequence set forth in any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650.

62. 58. The system of claim 57, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parent ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN.

63. 63. The system of any one of claims 48 to 62, wherein the RGN is a nickase or nuclease inactive RGN.

64. 64. The system of claim 63, wherein the RGN nickase comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 49-56, 698, and 699 and retains nickase activity.

65. 64. The system of claim 63, wherein the RGN nickase comprises the amino acid sequence of any one of SEQ ID NOs: 49-56, 698, and 699.

66. 66. The system of any one of claims 48 to 65, wherein at least one of the nucleotide sequences encoding the one or more guide RNAs and the nucleotide sequence encoding the fusion protein is operably linked to a promoter heterologous to the nucleotide sequence.

67. 67. The system of any one of claims 48 to 66, wherein the target DNA sequence is a eukaryotic target DNA sequence.

68. 68. The system of any one of claims 48 to 67, wherein the target DNA sequence is located adjacent to a protospacer adjacent motif (PAM) recognized by the RGN.

69. 69. The system of any one of claims 48 to 68, wherein the nucleotide sequence encoding the one or more guide RNAs and the nucleotide sequence encoding the fusion protein are located on one vector.

70. A cell comprising the system of any one of claims 48 to 69, 117 and 118.

71. 71. The cell of claim 70, wherein the cell is a mammalian cell.

72. 71. The cell of claim 70, wherein the cell is a plant cell.

73. 73. A plant or seed comprising the cell of claim 72.

74. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a system according to any one of claims 48 to 69 and 127 to 133, or a cell according to claim 71.

75. 134. A method for localizing a heterologous polypeptide to a target DNA molecule comprising a target DNA sequence, the method comprising delivering a system according to any one of claims 48 to 69 and 127 to 133 to said target DNA molecule or to a cell comprising said target DNA molecule.

76. 134. A method for modifying a target DNA molecule comprising a target DNA sequence, the method comprising delivering a system according to any one of claims 48 to 69 and 127 to 133 to said target DNA molecule or to a cell comprising said target DNA molecule.

77. 1. A method for localizing a heterologous polypeptide to a target DNA molecule comprising a target DNA sequence, comprising: a) i) one or more guide RNAs capable of hybridizing to the non-target strand of the target DNA sequence; and ii) assembling a ribonucleotide complex in vitro by combining a fusion protein comprising an RNA-guided nuclease (RGN) and at least one heterologous polypeptide inserted therein, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699; Under conditions suitable for the formation of the ribonucleotide complex, b) contacting the target DNA molecule or a cell containing the target DNA molecule with the in vitro assembled ribonucleotide complex; wherein the one or more guide RNAs hybridize to a non-target strand of the target DNA sequence, thereby directing the fusion protein to bind to the target DNA sequence.

78. 78. The method of claim 77, wherein the RGN comprises an amino acid sequence having the amino acid sequence of any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699.

79. 79. The method of claim 77 or 78, wherein the heterologous polypeptide is inserted within linker domain 2, wedge domain, RuvC domain, HNH domain, Rec-2 domain, or PAM-interacting domain of the RGN.

80. a) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 1, and the at least one heterologous polypeptide is located within the RGN: I) an amino acid position corresponding to position 30 of SEQ ID NO: 1; II) an amino acid position corresponding to position 642 of SEQ ID NO: 1; III) an amino acid position corresponding to position 670 of SEQ ID NO: 1; IV) an amino acid position corresponding to position 737 of SEQ ID NO: 1; W) the amino acid position corresponding to position 772 of SEQ ID NO: 1; VI) an amino acid position corresponding to position 775 of SEQ ID NO: 1; VII) an amino acid position corresponding to position 778 of SEQ ID NO: 1; and VIII) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 802 of SEQ ID NO: 1; or b) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:2, and the at least one heterologous polypeptide is located within the RGN: I) an amino acid position corresponding to position 678 of SEQ ID NO:2; II) an amino acid position corresponding to position 736 of SEQ ID NO:2; III) an amino acid position corresponding to position 778 of SEQ ID NO:2; IV) an amino acid position corresponding to position 788 of SEQ ID NO:2; and V) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 922 of SEQ ID NO:2; c) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 3, and the at least one heterologous polypeptide is located within the RGN: I) an amino acid position corresponding to position 725 of SEQ ID NO:3; II) an amino acid position corresponding to position 739 of SEQ ID NO:3; III) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 744 of SEQ ID NO: 3; or d) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 4, and the at least one heterologous polypeptide is located within the RGN: I) an amino acid position corresponding to position 347 of SEQ ID NO:4; II) an amino acid position corresponding to position 524 of SEQ ID NO:4; III) an amino acid position corresponding to position 666 of SEQ ID NO:4; IV) an amino acid position corresponding to position 680 of SEQ ID NO: 4; V) an amino acid position corresponding to position 740 of SEQ ID NO: 4; VI) an amino acid position corresponding to position 785 of SEQ ID NO:4; VII) an amino acid position corresponding to position 910 of SEQ ID NO: 4; and VIII) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 1077 of SEQ ID NO: 4; or e) the RGN comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 131, and the at least one heterologous polypeptide is located within the RGN: I) an amino acid position corresponding to position 766 of SEQ ID NO: 131; and II) the amino acid position corresponding to position 806 of SEQ ID NO:

131.

81. 81. The method of any one of Claims 77-80, wherein the heterologous polypeptide comprises a base-editing polypeptide or a prime-editing polypeptide, and the method further comprises modifying the target DNA molecule to produce a modified target DNA molecule.

82. 82. The method of claim 81, wherein the target DNA molecule contains a mutation that causes a disease or disorder, and wherein modification of the target DNA molecule corrects the causative mutation.

83. 83. The method of claim 82, wherein said correction of said causative mutation comprises correcting a nonsense mutation.

84. The method of any one of Claims 81 to 83, wherein the base-editing polypeptide comprises a deaminase.

85. 85. The method of Claim 84, wherein the fusion protein is a base editor fusion protein, and wherein the base editor fusion protein has improved editing activity compared to a parent base editor comprising the deaminase fused to the amino terminus of the RGN.

86. Within the RGN, the at least one heterologous polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 1; a) the amino acid position corresponding to position 642 of SEQ ID NO: 1; b) an amino acid position corresponding to position 670 of SEQ ID NO: 1; c) an amino acid position corresponding to position 737 of SEQ ID NO: 1; d) an amino acid position corresponding to position 772 of SEQ ID NO: 1; e) an amino acid position corresponding to position 775 of SEQ ID NO: 1; and f) inserted immediately after an amino acid position selected from the group consisting of the amino acid position corresponding to position 778 of SEQ ID NO: 1; 86. The method of claim 84 or 85, wherein the fusion protein is a base editor fusion protein, and the base editor fusion protein has a shifted editing window compared to a parent base editor comprising the deaminase fused to the amino terminus of the RGN.

87. 87. The method of any one of claims 84-86, wherein the deaminase lacks the first and last amino acid residues compared to the parent deaminase from which the deaminase is derived.

88. 88. The method of any one of claims 84 to 87, wherein the deaminase comprises adenine deaminase and the fusion protein is an adenine base editor (ABE) fusion protein.

89. 89. The method of claim 88, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723.

90. 89. The method of claim 88, wherein the adenine deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 5, 6, 248-414, 596, 597, and 720-723.

91. 89. The method of claim 88, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650.

92. 89. The method of claim 88, wherein the ABE fusion protein comprises an amino acid sequence set forth in any one of SEQ ID NOs: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647, and 650.

93. 89. The method of Claim 88, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parent ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN.

94. 94. The method of any one of claims 77 to 93, wherein the RGN is a nickase or nuclease inactive RGN.

95. 95. The method of claim 94, wherein the RGN nickase comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 49-56, 698, and 699, and retains nickase activity.

96. 95. The method of claim 94, wherein the RGN nickase comprises the amino acid sequence of any one of SEQ ID NOs: 49-56, 698, and 699.

97. 97. The method of any one of claims 77 to 96, wherein at least one of the nucleotide sequences encoding one or more guide RNAs and the nucleotide sequence encoding the fusion protein is operably linked to a promoter heterologous to said nucleotide sequence.

98. 98. The method of any one of claims 77 to 97, wherein the target DNA sequence is a eukaryotic target DNA sequence.

99. 99. The method of any one of claims 77 to 98, wherein the target DNA sequence is located adjacent to a protospacer adjacent motif (PAM) recognized by the RGN.

100. 100. The method of any one of claims 77 to 99, wherein the target DNA sequence is intracellular.

101. 101. The method of claim 100, further comprising selecting cells containing the modified target DNA molecule.

102. A cell comprising a target DNA molecule modified by the method of claim 101.

103. The cell of claim 102, wherein the cell is a mammalian cell.

104. 103. The cell of claim 102, wherein the cell is a plant cell.

105. 105. A plant or seed comprising the cell of claim 104.

106. A pharmaceutical composition comprising the cells of claim 103 and a pharmaceutically acceptable carrier.

107. 1. A method for treating a subject having or at risk of developing a disease, disorder, or condition, comprising: 133。 A method comprising administering to the subject a fusion protein according to any one of claims 1 to 18 and 113 to 119, an RNP complex according to claim 19, a nucleic acid molecule according to any one of claims 20 to 37 and 120 to 126, a vector according to claim 38 or 39, a cell according to any one of claims 43, 71 and 103, a system according to any one of claims 48 to 69 and 127 to 133, or a pharmaceutical composition according to any one of claims 46, 74 and 106.

108. 108. The method of claim 107, wherein the disease is associated with a causative mutation and the treatment comprises correcting the causative mutation.

109. 133. Use of a fusion protein according to any one of claims 1 to 18 and 113 to 119, an RNP complex according to claim 19, a nucleic acid molecule according to any one of claims 20 to 37 and 120 to 126, a vector according to claim 38 or 39, a cell according to any one of claims 43, 71 and 103, or a system according to any one of claims 48 to 69 and 127 to 133 for the treatment of a disease, disorder or condition in a subject having or at risk of developing said disease, disorder or condition.

110. 110. The use of claim 109, wherein the disease is associated with a causative mutation and the treatment comprises correcting the causative mutation.

111. 133. Use of a fusion protein according to any one of claims 1 to 18 and 113 to 119, an RNP complex according to claim 19, a nucleic acid molecule according to any one of claims 20 to 37 and 120 to 126, a vector according to claim 38 or 39, a cell according to any one of claims 43, 71 and 103, or a system according to any one of claims 48 to 69 and 127 to 133 for the manufacture of a medicament useful for treating a disease, disorder or condition.

112. 112. The use of claim 111, wherein the disease is associated with a causative mutation and an effective amount of the agent corrects the causative mutation.

113. The adenine deaminase comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO:5, and comprising the following amino acid residues: a) a C at the position corresponding to position 2 of SEQ ID NO: 5; b) F or C at the position corresponding to position 22 of SEQ ID NO: 5; c) Q at a position corresponding to position 23 of SEQ ID NO:5; d) N at a position corresponding to position 35 of SEQ ID NO: 5; e) L at a position corresponding to position 40 of SEQ ID NO: 5; f) Q at a position corresponding to position 46 of SEQ ID NO:5; g) A or M at the position corresponding to position 68 of SEQ ID NO: 5; h) H at a position corresponding to position 72 of SEQ ID NO: 5; i) W, A, Q, Y, or D at a position corresponding to position 75 of SEQ ID NO:5; j) G at a position corresponding to position 76 of SEQ ID NO:5; k) S at a position corresponding to position 81 of SEQ ID NO: 5; l) M at a position corresponding to position 105 of SEQ ID NO: 5; m) H at a position corresponding to position 108 of SEQ ID NO: 5; n) L at a position corresponding to position 109 of SEQ ID NO: 5; o) I or L at a position corresponding to position 117 of SEQ ID NO: 5; p) F at a position corresponding to position 120 of SEQ ID NO: 5; q) E, T, A, G, or V at a position corresponding to position 121 of SEQ ID NO: 5; r) I or H at a position corresponding to position 122 of SEQ ID NO: 5; s) K at a position corresponding to position 125 of SEQ ID NO:5; t) A at a position corresponding to position 126 of SEQ ID NO:5; u) H at a position corresponding to position 135 of SEQ ID NO: 5; v) V at a position corresponding to position 137 of SEQ ID NO: 5; w) Y at a position corresponding to position 138 of SEQ ID NO: 5; x) L at a position corresponding to position 139 of SEQ ID NO: 5; y) K or A at a position corresponding to position 142 of SEQ ID NO: 5; z) A, Q, L, or M at a position corresponding to position 145 of SEQ ID NO: 5; aa) K at a position corresponding to position 148 of SEQ ID NO: 5; bb) Q at a position corresponding to position 151 of SEQ ID NO:5; cc) E or R at a position corresponding to position 153 of SEQ ID NO:5; dd) W at a position corresponding to position 155 of SEQ ID NO: 5; ee) R or V at a position corresponding to position 156 of SEQ ID NO:5; ff) F at a position corresponding to position 157 of SEQ ID NO: 5; gg) R at a position corresponding to position 158 of SEQ ID NO:5; hh) Q at a position corresponding to position 159 of SEQ ID NO: 5; ii) D or W at a position corresponding to position 160 of SEQ ID NO: 5; jj) W, H, V, or A at a position corresponding to position 162 of SEQ ID NO:5; kk) R at a position corresponding to position 165 of SEQ ID NO:5, and ll) at least one of H at a position corresponding to position 166 of SEQ ID NO:

5.

114. The adenine deaminase a) N at a position corresponding to position 35 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; b) N at a position corresponding to position 35 of SEQ ID NO:5, and Q at a position corresponding to position 46 of SEQ ID NO:5; c) S at a position corresponding to position 81 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; d) Q at a position corresponding to position 46 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; e) R at a position corresponding to position 156 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; f) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5; g) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at a position corresponding to position 35 of SEQ ID NO:5, Q at a position corresponding to position 46 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; i) a W at a position corresponding to position 155 of SEQ ID NO:5, an R at a position corresponding to position 156 of SEQ ID NO:5, and a W at a position corresponding to position 162 of SEQ ID NO:5; j) N at a position corresponding to position 35 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; k) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; l) M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; m) Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; n) M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; o) N at a position corresponding to position 35, Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; p) S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; q) W at a position corresponding to position 75 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; r) S at a position corresponding to position 81 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; s) S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; t) W at a position corresponding to position 75 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; u) W at a position corresponding to position 155 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; v) A at a position corresponding to position 145 of SEQ ID NO:5 and W at a position corresponding to position 155 of SEQ ID NO:5; w) W at a position corresponding to position 75 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; x) W at a position corresponding to position 75 of SEQ ID NO:5 and A at a position corresponding to position 145 of SEQ ID NO:5; y) A at a position corresponding to position 145 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; z) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; aa) S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; bb) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; cc) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; dd) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; ee) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ff) W at a position corresponding to position 75 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; gg) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a W at a position corresponding to position 155 of SEQ ID NO:5; hh) A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ii) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a D at a position corresponding to position 160 of SEQ ID NO:5; jj) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; kk) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; ll) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; mm) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; nn) W at a position corresponding to position 75 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; oo) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; pp) L at a position corresponding to position 40 of SEQ ID NO: 5, and M at a position corresponding to position 105 of SEQ ID NO: 5; qq) L at a position corresponding to position 40 of SEQ ID NO:5, and V at a position corresponding to position 121 of SEQ ID NO:5; rr) L at a position corresponding to position 40 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; ss) L at a position corresponding to position 40 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; tt) L at a position corresponding to position 40 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; uu) M at a position corresponding to position 105 of SEQ ID NO:5 and V at a position corresponding to position 121 of SEQ ID NO:5; vv) M at a position corresponding to position 105 of SEQ ID NO: 5, and M at a position corresponding to position 145 of SEQ ID NO: 5; ww) M at a position corresponding to position 105 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; xx) M at a position corresponding to position 105 of SEQ ID NO: 5 and W at a position corresponding to position 160 of SEQ ID NO: 5; yy) V at a position corresponding to position 121 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; zz) V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; aaa) V at a position corresponding to position 121 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; bbb) M at a position corresponding to position 145 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ccc) M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; ddd) Q at a position corresponding to position 159 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; eee) S at a position corresponding to position 81 of SEQ ID NO:5, M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; fff) S at a position corresponding to position 81 of SEQ ID NO:5, V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ggg) A at a position corresponding to position 68 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; hhh) G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; iii) S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; jjj) A at a position corresponding to position 68 of SEQ ID NO:5 and G at a position corresponding to position 76 of SEQ ID NO:5; kkk) A at a position corresponding to position 68 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; lll) G at a position corresponding to position 76 of SEQ ID NO:5 and L at a position corresponding to position 117 of SEQ ID NO:5; mmm) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; nnn) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ooo) G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ppp) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; qqq) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; rrr) S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 109 of SEQ ID NO:5, E at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; sss) A at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, A at a position corresponding to position 121 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ttt) Q at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, I at a position corresponding to position 117 of SEQ ID NO:5, Q at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; uuu) Y at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, G at a position corresponding to position 121 of SEQ ID NO:5, L at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; vvv) C at a position corresponding to position 22 of SEQ ID NO:5, A at a position corresponding to position 68 of SEQ ID NO:5, Y at a position corresponding to position 75 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, H at a position corresponding to position 122 of SEQ ID NO:5, Y at a position corresponding to position 138 of SEQ ID NO:5, L at a position corresponding to position 139 of SEQ ID NO:5, A at a position corresponding to position 142 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, R at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, D at a position corresponding to position 160 of SEQ ID NO:5, and H at a position corresponding to position 162 of SEQ ID NO:5; or www) S at a position corresponding to position 81 of SEQ ID NO:

5.

115. 115. The fusion protein of claim 113 or 114, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

116. 115. The fusion protein of claim 113 or 114, wherein the deaminase has an amino acid sequence set forth as any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

117. the fusion protein a) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; b) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 52 after an amino acid position corresponding to position 910 of SEQ ID NO: 4; c) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 319 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; d) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 319 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 52 after an amino acid position corresponding to position 910 of SEQ ID NO: 4; e) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; f) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; g) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 319 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; h) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 597 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; i) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 319 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; j) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 406 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; k) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 597 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; l) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 596 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; m) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 406 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; n) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 597 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; o) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 406 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; p) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 596 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; q) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 596 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; r) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 406 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; s) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 596 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; and t) the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 597 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1. The fusion protein of claim 113 or 114, wherein the deaminase is selected from the group consisting of:

118. the fusion protein a) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 375 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; b) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 375 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; c) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 319 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; d) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 319 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; e) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 375 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; f) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 375 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; g) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 319 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; h) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 597 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; i) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 319 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; j) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 406 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; k) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 597 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; l) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 596 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; m) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 406 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; n) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 597 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; o) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 406 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; p) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 596 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; q) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 596 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; r) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 406 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; s) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 596 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; and t) the deaminase has the amino acid sequence set forth as SEQ ID NO: 597 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1. The fusion protein of claim 113 or 114, wherein the deaminase has the amino acid sequence set forth as SEQ ID NO: 597 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1.

119. 119. The fusion protein of claim 118, wherein the fusion protein has a sequence set forth as any one of SEQ ID NOs: 599, 600, 602, 603, 606, 610, 612, 614, 615, 617, 619-624, 626, and 630-632.

120. The adenine deaminase comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO:5, and comprising the following amino acid residues: a) a C at the position corresponding to position 2 of SEQ ID NO: 5; b) F or C at the position corresponding to position 22 of SEQ ID NO: 5; c) Q at a position corresponding to position 23 of SEQ ID NO:5; d) N at a position corresponding to position 35 of SEQ ID NO: 5; e) L at a position corresponding to position 40 of SEQ ID NO: 5; f) Q at a position corresponding to position 46 of SEQ ID NO:5; g) A or M at the position corresponding to position 68 of SEQ ID NO: 5; h) H at a position corresponding to position 72 of SEQ ID NO: 5; i) W, A, Q, Y, or D at a position corresponding to position 75 of SEQ ID NO:5; j) G at a position corresponding to position 76 of SEQ ID NO:5; k) S at a position corresponding to position 81 of SEQ ID NO: 5; l) M at a position corresponding to position 105 of SEQ ID NO: 5; m) H at a position corresponding to position 108 of SEQ ID NO: 5; n) L at a position corresponding to position 109 of SEQ ID NO: 5; o) I or L at a position corresponding to position 117 of SEQ ID NO: 5; p) F at a position corresponding to position 120 of SEQ ID NO: 5; q) E, T, A, G, or V at a position corresponding to position 121 of SEQ ID NO: 5; r) I or H at a position corresponding to position 122 of SEQ ID NO: 5; s) K at a position corresponding to position 125 of SEQ ID NO:5; t) A at a position corresponding to position 126 of SEQ ID NO:5; u) H at a position corresponding to position 135 of SEQ ID NO: 5; v) V at a position corresponding to position 137 of SEQ ID NO: 5; w) Y at a position corresponding to position 138 of SEQ ID NO: 5; x) L at a position corresponding to position 139 of SEQ ID NO: 5; y) K or A at a position corresponding to position 142 of SEQ ID NO: 5; z) A, Q, L, or M at a position corresponding to position 145 of SEQ ID NO: 5; aa) K at a position corresponding to position 148 of SEQ ID NO: 5; bb) Q at a position corresponding to position 151 of SEQ ID NO:5; cc) E or R at a position corresponding to position 153 of SEQ ID NO:5; dd) W at a position corresponding to position 155 of SEQ ID NO: 5; ee) R or V at a position corresponding to position 156 of SEQ ID NO:5; ff) F at a position corresponding to position 157 of SEQ ID NO: 5; gg) R at a position corresponding to position 158 of SEQ ID NO:5; hh) Q at a position corresponding to position 159 of SEQ ID NO: 5; ii) D or W at a position corresponding to position 160 of SEQ ID NO: 5; jj) W, H, V, or A at a position corresponding to position 162 of SEQ ID NO:5; kk) R at a position corresponding to position 165 of SEQ ID NO:5, and 11) The nucleic acid molecule of claim 29, comprising at least one of H at a position corresponding to position 166 of SEQ ID NO:

5.

121. The adenine deaminase a) N at a position corresponding to position 35 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; b) N at a position corresponding to position 35 of SEQ ID NO:5, and Q at a position corresponding to position 46 of SEQ ID NO:5; c) S at a position corresponding to position 81 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; d) Q at a position corresponding to position 46 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; e) R at a position corresponding to position 156 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; f) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5; g) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at a position corresponding to position 35 of SEQ ID NO:5, Q at a position corresponding to position 46 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; i) a W at a position corresponding to position 155 of SEQ ID NO:5, an R at a position corresponding to position 156 of SEQ ID NO:5, and a W at a position corresponding to position 162 of SEQ ID NO:5; j) N at a position corresponding to position 35 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; k) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; l) M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; m) Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; n) M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; o) N at a position corresponding to position 35, Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; p) S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; q) W at a position corresponding to position 75 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; r) S at a position corresponding to position 81 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; s) S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; t) W at a position corresponding to position 75 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; u) W at a position corresponding to position 155 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; v) A at a position corresponding to position 145 of SEQ ID NO:5 and W at a position corresponding to position 155 of SEQ ID NO:5; w) W at a position corresponding to position 75 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; x) W at a position corresponding to position 75 of SEQ ID NO:5 and A at a position corresponding to position 145 of SEQ ID NO:5; y) A at a position corresponding to position 145 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; z) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; aa) S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; bb) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; cc) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; dd) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; ee) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ff) W at a position corresponding to position 75 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; gg) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a W at a position corresponding to position 155 of SEQ ID NO:5; hh) A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ii) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a D at a position corresponding to position 160 of SEQ ID NO:5; jj) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; kk) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; ll) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; mm) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; nn) W at a position corresponding to position 75 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; oo) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; pp) L at a position corresponding to position 40 of SEQ ID NO: 5, and M at a position corresponding to position 105 of SEQ ID NO: 5; qq) L at a position corresponding to position 40 of SEQ ID NO:5, and V at a position corresponding to position 121 of SEQ ID NO:5; rr) L at a position corresponding to position 40 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; ss) L at a position corresponding to position 40 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; tt) L at a position corresponding to position 40 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; uu) M at a position corresponding to position 105 of SEQ ID NO:5 and V at a position corresponding to position 121 of SEQ ID NO:5; vv) M at a position corresponding to position 105 of SEQ ID NO: 5, and M at a position corresponding to position 145 of SEQ ID NO: 5; ww) M at a position corresponding to position 105 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; xx) M at a position corresponding to position 105 of SEQ ID NO: 5 and W at a position corresponding to position 160 of SEQ ID NO: 5; yy) V at a position corresponding to position 121 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; zz) V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; aaa) V at a position corresponding to position 121 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; bbb) M at a position corresponding to position 145 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ccc) M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; ddd) Q at a position corresponding to position 159 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; eee) S at a position corresponding to position 81 of SEQ ID NO:5, M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; fff) S at a position corresponding to position 81 of SEQ ID NO:5, V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ggg) A at a position corresponding to position 68 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; hhh) G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; iii) S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; jjj) A at a position corresponding to position 68 of SEQ ID NO:5 and G at a position corresponding to position 76 of SEQ ID NO:5; kkk) A at a position corresponding to position 68 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; lll) G at a position corresponding to position 76 of SEQ ID NO:5 and L at a position corresponding to position 117 of SEQ ID NO:5; mmm) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; nnn) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ooo) G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ppp) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; qqq) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; rrr) S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 109 of SEQ ID NO:5, E at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; sss) A at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, A at a position corresponding to position 121 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ttt) Q at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, I at a position corresponding to position 117 of SEQ ID NO:5, Q at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; uuu) Y at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, G at a position corresponding to position 121 of SEQ ID NO:5, L at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; vvv) C at a position corresponding to position 22 of SEQ ID NO:5, A at a position corresponding to position 68 of SEQ ID NO:5, Y at a position corresponding to position 75 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, H at a position corresponding to position 122 of SEQ ID NO:5, Y at a position corresponding to position 138 of SEQ ID NO:5, L at a position corresponding to position 139 of SEQ ID NO:5, A at a position corresponding to position 142 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, R at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, D at a position corresponding to position 160 of SEQ ID NO:5, and H at a position corresponding to position 162 of SEQ ID NO:5; or www) S at a position corresponding to position 81 of SEQ ID NO:

5.

122. 122. The nucleic acid molecule of claim 120 or 121, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

123. 122. The nucleic acid molecule of claim 120 or 121, wherein the deaminase has an amino acid sequence set forth as any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

124. the fusion protein a) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; b) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 52 after an amino acid position corresponding to position 910 of SEQ ID NO: 4; c) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 319 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; d) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 319 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 52 after an amino acid position corresponding to position 910 of SEQ ID NO: 4; e) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; f) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; g) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 319 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; h) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 597 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; i) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 319 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; j) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 406 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; k) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 597 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; l) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 596 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; m) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 406 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; n) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 597 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; o) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 406 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; p) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 596 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; q) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 596 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; r) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 406 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; s) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 596 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; and t) the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 597 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1. The nucleic acid molecule of claim 120 or 121, wherein the deaminase is selected from the group consisting of: a fusion protein having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 597 and inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1.

125. the fusion protein a) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 375 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; b) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 375 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; c) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 319 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; d) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 319 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; e) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 375 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; f) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 375 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; g) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 319 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; h) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 597 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; i) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 319 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; j) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 406 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; k) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 597 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; l) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 596 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; m) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 406 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; n) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 597 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; o) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 406 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; p) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 596 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; q) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 596 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; r) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 406 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; s) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 596 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; and t) the deaminase has the amino acid sequence set forth as SEQ ID NO: 597 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1. The nucleic acid molecule of claim 120 or 121, wherein the deaminase is selected from the group consisting of: a fusion protein having the amino acid sequence set forth as SEQ ID NO: 597 and inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1.

126. 126. The nucleic acid molecule of claim 125, wherein the fusion protein has a sequence set forth as any one of SEQ ID NOs: 599, 600, 602, 603, 606, 610, 612, 614, 615, 617, 619-624, 626, and 630-632.

127. The adenine deaminase comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO:5, and comprising the following amino acid residues: a) a C at the position corresponding to position 2 of SEQ ID NO: 5; b) F or C at the position corresponding to position 22 of SEQ ID NO: 5; c) Q at a position corresponding to position 23 of SEQ ID NO:5; d) N at a position corresponding to position 35 of SEQ ID NO: 5; e) L at a position corresponding to position 40 of SEQ ID NO: 5; f) Q at a position corresponding to position 46 of SEQ ID NO:5; g) A or M at the position corresponding to position 68 of SEQ ID NO: 5; h) H at a position corresponding to position 72 of SEQ ID NO: 5; i) W, A, Q, Y, or D at a position corresponding to position 75 of SEQ ID NO:5; j) G at a position corresponding to position 76 of SEQ ID NO:5; k) S at a position corresponding to position 81 of SEQ ID NO: 5; l) M at a position corresponding to position 105 of SEQ ID NO: 5; m) H at a position corresponding to position 108 of SEQ ID NO: 5; n) L at a position corresponding to position 109 of SEQ ID NO: 5; o) I or L at a position corresponding to position 117 of SEQ ID NO: 5; p) F at a position corresponding to position 120 of SEQ ID NO: 5; q) E, T, A, G, or V at a position corresponding to position 121 of SEQ ID NO: 5; r) I or H at a position corresponding to position 122 of SEQ ID NO: 5; s) K at a position corresponding to position 125 of SEQ ID NO:5; t) A at a position corresponding to position 126 of SEQ ID NO:5; u) H at a position corresponding to position 135 of SEQ ID NO: 5; v) V at a position corresponding to position 137 of SEQ ID NO: 5; w) Y at a position corresponding to position 138 of SEQ ID NO: 5; x) L at a position corresponding to position 139 of SEQ ID NO: 5; y) K or A at a position corresponding to position 142 of SEQ ID NO: 5; z) A, Q, L, or M at a position corresponding to position 145 of SEQ ID NO:5; aa) K at a position corresponding to position 148 of SEQ ID NO: 5; bb) Q at a position corresponding to position 151 of SEQ ID NO:5; cc) E or R at a position corresponding to position 153 of SEQ ID NO: 5; dd) W at a position corresponding to position 155 of SEQ ID NO: 5; ee) R or V at a position corresponding to position 156 of SEQ ID NO:5; ff) F at a position corresponding to position 157 of SEQ ID NO: 5; gg) R at a position corresponding to position 158 of SEQ ID NO:5; hh) Q at a position corresponding to position 159 of SEQ ID NO: 5; ii) D or W at a position corresponding to position 160 of SEQ ID NO: 5; jj) W, H, V, or A at a position corresponding to position 162 of SEQ ID NO:5; kk) R at a position corresponding to position 165 of SEQ ID NO:5, and ll) at least one of H at a position corresponding to position 166 of SEQ ID NO:

5.

128. The adenine deaminase a) N at a position corresponding to position 35 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; b) N at a position corresponding to position 35 of SEQ ID NO:5, and Q at a position corresponding to position 46 of SEQ ID NO:5; c) S at a position corresponding to position 81 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; d) Q at a position corresponding to position 46 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; e) R at a position corresponding to position 156 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; f) A at a position corresponding to position 68 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; g) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at a position corresponding to position 35 of SEQ ID NO:5, Q at a position corresponding to position 46 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; i) a W at a position corresponding to position 155 of SEQ ID NO:5, an R at a position corresponding to position 156 of SEQ ID NO:5, and a W at a position corresponding to position 162 of SEQ ID NO:5; j) N at a position corresponding to position 35 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; k) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; l) M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; m) Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; n) M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; o) N at a position corresponding to position 35, Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; p) S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; q) W at a position corresponding to position 75 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; r) S at a position corresponding to position 81 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; s) S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; t) W at a position corresponding to position 75 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; u) W at a position corresponding to position 155 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; v) A at a position corresponding to position 145 of SEQ ID NO:5 and W at a position corresponding to position 155 of SEQ ID NO:5; w) W at a position corresponding to position 75 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; x) W at a position corresponding to position 75 of SEQ ID NO:5 and A at a position corresponding to position 145 of SEQ ID NO:5; y) A at a position corresponding to position 145 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; z) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; aa) S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; bb) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; cc) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; dd) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; ee) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ff) W at a position corresponding to position 75 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; gg) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a W at a position corresponding to position 155 of SEQ ID NO:5; hh) A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ii) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a D at a position corresponding to position 160 of SEQ ID NO:5; jj) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; kk) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; ll) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; mm) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; nn) W at a position corresponding to position 75 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; oo) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; pp) L at a position corresponding to position 40 of SEQ ID NO: 5, and M at a position corresponding to position 105 of SEQ ID NO: 5; qq) L at a position corresponding to position 40 of SEQ ID NO:5, and V at a position corresponding to position 121 of SEQ ID NO:5; rr) L at a position corresponding to position 40 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; ss) L at a position corresponding to position 40 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; tt) L at a position corresponding to position 40 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; uu) M at a position corresponding to position 105 of SEQ ID NO:5 and V at a position corresponding to position 121 of SEQ ID NO:5; vv) M at a position corresponding to position 105 of SEQ ID NO: 5, and M at a position corresponding to position 145 of SEQ ID NO: 5; ww) M at a position corresponding to position 105 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; xx) M at a position corresponding to position 105 of SEQ ID NO: 5 and W at a position corresponding to position 160 of SEQ ID NO: 5; yy) V at a position corresponding to position 121 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; zz) V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; aaa) V at a position corresponding to position 121 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; bbb) M at a position corresponding to position 145 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ccc) M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; ddd) Q at a position corresponding to position 159 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; eee) S at a position corresponding to position 81 of SEQ ID NO:5, M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; fff) S at a position corresponding to position 81 of SEQ ID NO:5, V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ggg) A at a position corresponding to position 68 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; hhh) G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; iii) S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; jjj) A at a position corresponding to position 68 of SEQ ID NO:5 and G at a position corresponding to position 76 of SEQ ID NO:5; kkk) A at a position corresponding to position 68 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; lll) G at a position corresponding to position 76 of SEQ ID NO:5 and L at a position corresponding to position 117 of SEQ ID NO:5; mmm) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; nnn) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ooo) G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ppp) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; qqq) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; rrr) S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 109 of SEQ ID NO:5, E at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; sss) A at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, A at a position corresponding to position 121 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ttt) Q at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, I at a position corresponding to position 117 of SEQ ID NO:5, Q at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; uuu) Y at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, G at a position corresponding to position 121 of SEQ ID NO:5, L at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; vvv) C at a position corresponding to position 22 of SEQ ID NO:5, A at a position corresponding to position 68 of SEQ ID NO:5, Y at a position corresponding to position 75 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, H at a position corresponding to position 122 of SEQ ID NO:5, Y at a position corresponding to position 138 of SEQ ID NO:5, L at a position corresponding to position 139 of SEQ ID NO:5, A at a position corresponding to position 142 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, R at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, D at a position corresponding to position 160 of SEQ ID NO:5, and H at a position corresponding to position 162 of SEQ ID NO:5; or www) S at position corresponding to position 81 of SEQ ID NO:

5.

129. 129. The system of claim 127 or 128, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

130. 129. The system of claim 127 or 128, wherein the deaminase has an amino acid sequence set forth as any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

131. the fusion protein a) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; b) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 52 after an amino acid position corresponding to position 910 of SEQ ID NO: 4; c) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 319 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; d) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 319 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 52 after an amino acid position corresponding to position 910 of SEQ ID NO: 4; e) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; f) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; g) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 319 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; h) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 597 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; i) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 319 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; j) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 406 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; k) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 597 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; l) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 596 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; m) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 406 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; n) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 597 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; o) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 406 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; p) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 596 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; q) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 596 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; r) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 406 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; s) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 596 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; and t) the deaminase is selected from the group consisting of a fusion protein having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 597 and inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1. The system of claim 127 or 128, wherein the deaminase is selected from the group consisting of: a fusion protein having at least 90% sequence identity to SEQ ID NO: 597 and inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1.

132. a) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 375 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; a) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 375 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; b) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 319 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; c) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 319 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; d) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 375 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; e) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 375 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; f) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 319 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; g) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 597 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; h) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 319 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; i) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 406 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; j) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 597 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; k) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 596 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; l) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 406 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; m) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 597 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; n) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 406 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; o) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 596 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; p) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 596 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; q) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 406 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; r) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 596 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; and s) the deaminase is a fusion protein having the amino acid sequence set forth as SEQ ID NO: 597 and inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1. The system of claim 127 or 128, wherein the deaminase is a fusion protein having the amino acid sequence set forth as SEQ ID NO: 597 and inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1.

133. 133. The system of claim 132, wherein the fusion protein has a sequence set forth as any one of SEQ ID NOs: 599, 600, 602, 603, 606, 610, 612, 614, 615, 617, 619-624, 626, and 630-632.

134. The adenine deaminase comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO:5, and comprising the following amino acid residues: a) a C at the position corresponding to position 2 of SEQ ID NO: 5; b) F or C at the position corresponding to position 22 of SEQ ID NO: 5; c) Q at a position corresponding to position 23 of SEQ ID NO:5; d) N at a position corresponding to position 35 of SEQ ID NO: 5; e) L at a position corresponding to position 40 of SEQ ID NO: 5; f) Q at a position corresponding to position 46 of SEQ ID NO:5; g) A or M at the position corresponding to position 68 of SEQ ID NO: 5; h) H at a position corresponding to position 72 of SEQ ID NO: 5; i) W, A, Q, Y, or D at a position corresponding to position 75 of SEQ ID NO:5; j) G at a position corresponding to position 76 of SEQ ID NO:5; k) S at a position corresponding to position 81 of SEQ ID NO: 5; l) M at a position corresponding to position 105 of SEQ ID NO: 5; m) H at a position corresponding to position 108 of SEQ ID NO: 5; n) L at a position corresponding to position 109 of SEQ ID NO: 5; o) I or L at a position corresponding to position 117 of SEQ ID NO: 5; p) F at a position corresponding to position 120 of SEQ ID NO: 5; q) E, T, A, G, or V at a position corresponding to position 121 of SEQ ID NO: 5; r) I or H at a position corresponding to position 122 of SEQ ID NO: 5; s) K at a position corresponding to position 125 of SEQ ID NO:5; t) A at a position corresponding to position 126 of SEQ ID NO:5; u) H at a position corresponding to position 135 of SEQ ID NO: 5; v) V at a position corresponding to position 137 of SEQ ID NO: 5; w) Y at a position corresponding to position 138 of SEQ ID NO: 5; x) L at a position corresponding to position 139 of SEQ ID NO: 5; y) K or A at a position corresponding to position 142 of SEQ ID NO: 5; z) A, Q, L, or M at a position corresponding to position 145 of SEQ ID NO: 5; aa) K at a position corresponding to position 148 of SEQ ID NO: 5; bb) Q at a position corresponding to position 151 of SEQ ID NO:5; cc) E or R at a position corresponding to position 153 of SEQ ID NO:5; dd) W at a position corresponding to position 155 of SEQ ID NO: 5; ee) R or V at a position corresponding to position 156 of SEQ ID NO:5; ff) F at a position corresponding to position 157 of SEQ ID NO: 5; gg) R at a position corresponding to position 158 of SEQ ID NO:5; hh) Q at a position corresponding to position 159 of SEQ ID NO: 5; ii) D or W at a position corresponding to position 160 of SEQ ID NO: 5; jj) W, H, V, or A at a position corresponding to position 162 of SEQ ID NO:5; kk) R at a position corresponding to position 165 of SEQ ID NO:5, and ll) at least one of H at a position corresponding to position 166 of SEQ ID NO:

5.

135. The adenine deaminase a) N at a position corresponding to position 35 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; b) N at a position corresponding to position 35 of SEQ ID NO:5, and Q at a position corresponding to position 46 of SEQ ID NO:5; c) S at a position corresponding to position 81 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; d) Q at a position corresponding to position 46 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; e) R at a position corresponding to position 156 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; f) A at a position corresponding to position 68 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; g) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at a position corresponding to position 35 of SEQ ID NO:5, Q at a position corresponding to position 46 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; i) a W at a position corresponding to position 155 of SEQ ID NO:5, an R at a position corresponding to position 156 of SEQ ID NO:5, and a W at a position corresponding to position 162 of SEQ ID NO:5; j) N at a position corresponding to position 35 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; k) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; l) M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; m) Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; n) M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; o) N at a position corresponding to position 35, Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; p) S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; q) W at a position corresponding to position 75 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; r) S at a position corresponding to position 81 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; s) S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; t) W at a position corresponding to position 75 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; u) W at a position corresponding to position 155 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; v) A at a position corresponding to position 145 of SEQ ID NO:5 and W at a position corresponding to position 155 of SEQ ID NO:5; w) W at a position corresponding to position 75 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; x) W at a position corresponding to position 75 of SEQ ID NO:5 and A at a position corresponding to position 145 of SEQ ID NO:5; y) A at a position corresponding to position 145 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; z) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; aa) S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; bb) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; cc) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; dd) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; ee) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ff) W at a position corresponding to position 75 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; gg) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a W at a position corresponding to position 155 of SEQ ID NO:5; hh) A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ii) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a D at a position corresponding to position 160 of SEQ ID NO:5; jj) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; kk) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; ll) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; mm) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; nn) W at a position corresponding to position 75 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; oo) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; pp) L at a position corresponding to position 40 of SEQ ID NO: 5, and M at a position corresponding to position 105 of SEQ ID NO: 5; qq) L at a position corresponding to position 40 of SEQ ID NO:5, and V at a position corresponding to position 121 of SEQ ID NO:5; rr) L at a position corresponding to position 40 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; ss) L at a position corresponding to position 40 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; tt) L at a position corresponding to position 40 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; uu) M at a position corresponding to position 105 of SEQ ID NO:5 and V at a position corresponding to position 121 of SEQ ID NO:5; vv) M at a position corresponding to position 105 of SEQ ID NO: 5, and M at a position corresponding to position 145 of SEQ ID NO: 5; ww) M at a position corresponding to position 105 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; xx) M at a position corresponding to position 105 of SEQ ID NO: 5 and W at a position corresponding to position 160 of SEQ ID NO: 5; yy) V at a position corresponding to position 121 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; zz) V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; aaa) V at a position corresponding to position 121 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; bbb) M at a position corresponding to position 145 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ccc) M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; ddd) Q at a position corresponding to position 159 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; eee) S at a position corresponding to position 81 of SEQ ID NO:5, M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; fff) S at a position corresponding to position 81 of SEQ ID NO:5, V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ggg) A at a position corresponding to position 68 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; hhh) G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; iii) S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; jjj) A at a position corresponding to position 68 of SEQ ID NO:5 and G at a position corresponding to position 76 of SEQ ID NO:5; kkk) A at a position corresponding to position 68 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; lll) G at a position corresponding to position 76 of SEQ ID NO:5 and L at a position corresponding to position 117 of SEQ ID NO:5; mmm) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; nnn) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ooo) G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ppp) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; qqq) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; rrr) S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 109 of SEQ ID NO:5, E at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; sss) A at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, A at a position corresponding to position 121 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ttt) Q at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, I at a position corresponding to position 117 of SEQ ID NO:5, Q at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; uuu) Y at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, G at a position corresponding to position 121 of SEQ ID NO:5, L at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; vvv) C at a position corresponding to position 22 of SEQ ID NO:5, A at a position corresponding to position 68 of SEQ ID NO:5, Y at a position corresponding to position 75 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, H at a position corresponding to position 122 of SEQ ID NO:5, Y at a position corresponding to position 138 of SEQ ID NO:5, L at a position corresponding to position 139 of SEQ ID NO:5, A at a position corresponding to position 142 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, R at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, D at a position corresponding to position 160 of SEQ ID NO:5, and H at a position corresponding to position 162 of SEQ ID NO:5; or www) S at a position corresponding to position 81 of SEQ ID NO:

5.

136. 136. The method of claim 134 or 135, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

137. 136. The method of claim 134 or 135, wherein the deaminase has an amino acid sequence set forth as any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

138. the fusion protein a) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; b) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 52 after an amino acid position corresponding to position 910 of SEQ ID NO: 4; c) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 319 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; d) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 319 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 52 after an amino acid position corresponding to position 910 of SEQ ID NO: 4; e) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; f) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; g) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 319 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; h) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 597 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; i) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 319 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; j) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 406 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; k) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 597 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; l) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 596 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; m) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 406 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; n) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 597 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; o) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 406 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; p) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 596 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; q) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 596 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; r) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 406 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; s) a fusion protein in which the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 596 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; and t) a fusion protein wherein the deaminase has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 597 and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1. The method of claim 134 or 135, wherein the deaminase is selected from the group consisting of:

139. the fusion protein a) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 375 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; b) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 375 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; c) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 319 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; d) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 319 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; e) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 375 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; f) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 375 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; g) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 319 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; h) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 597 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; i) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 319 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; j) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 406 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; k) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 597 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; l) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 596 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 50 after the amino acid position corresponding to position 678 of SEQ ID NO: 2; m) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 406 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; n) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 597 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; o) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 406 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; p) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 596 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 52 after the amino acid position corresponding to position 910 of SEQ ID NO: 4; q) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 596 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 50 after the amino acid position corresponding to position 922 of SEQ ID NO: 2; r) a fusion protein in which the deaminase has the amino acid sequence shown as SEQ ID NO: 406 and is inserted into the RGN having the amino acid sequence shown as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; s) a fusion protein in which the deaminase has the amino acid sequence set forth as SEQ ID NO: 596 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO: 1; and t) the deaminase has the amino acid sequence set forth as SEQ ID NO: 597 and is inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1. The method of claim 134 or 135, wherein the deaminase is selected from the group consisting of: a fusion protein having the amino acid sequence set forth as SEQ ID NO: 597 and inserted into the RGN having the amino acid sequence set forth as SEQ ID NO: 49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1.

140. 140. The method of claim 139, wherein the fusion protein has a sequence set forth as any one of SEQ ID NOs: 599, 600, 602, 603, 606, 610, 612, 614, 615, 617, 619-624, 626, and 630-632.

141. 1. A deaminase comprising an amino acid sequence having at least 85% sequence identity to SEQ ID NO:5, comprising the following amino acid residues: a) a C at the position corresponding to position 2 of SEQ ID NO: 5; b) F or C at the position corresponding to position 22 of SEQ ID NO: 5; c) Q at a position corresponding to position 23 of SEQ ID NO:5; d) N at a position corresponding to position 35 of SEQ ID NO: 5; e) L at a position corresponding to position 40 of SEQ ID NO: 5; f) Q at a position corresponding to position 46 of SEQ ID NO:5; g) A or M at the position corresponding to position 68 of SEQ ID NO: 5; h) H at a position corresponding to position 72 of SEQ ID NO: 5; i) W, A, Q, Y, or D at a position corresponding to position 75 of SEQ ID NO:5; j) G at a position corresponding to position 76 of SEQ ID NO:5; k) S at a position corresponding to position 81 of SEQ ID NO: 5; l) M at a position corresponding to position 105 of SEQ ID NO: 5; m) H at a position corresponding to position 108 of SEQ ID NO: 5; n) L at a position corresponding to position 109 of SEQ ID NO: 5; o) I or L at a position corresponding to position 117 of SEQ ID NO: 5; p) F at a position corresponding to position 120 of SEQ ID NO: 5; q) E, T, A, G, or V at a position corresponding to position 121 of SEQ ID NO: 5; r) I or H at a position corresponding to position 122 of SEQ ID NO: 5; s) K at a position corresponding to position 125 of SEQ ID NO:5; t) A at a position corresponding to position 126 of SEQ ID NO:5; u) H at a position corresponding to position 135 of SEQ ID NO: 5; v) V at a position corresponding to position 137 of SEQ ID NO: 5; w) Y at a position corresponding to position 138 of SEQ ID NO: 5; x) L at a position corresponding to position 139 of SEQ ID NO: 5; y) K or A at a position corresponding to position 142 of SEQ ID NO: 5; z) A, Q, L, or M at a position corresponding to position 145 of SEQ ID NO: 5; aa) K at a position corresponding to position 148 of SEQ ID NO: 5; bb) Q at a position corresponding to position 151 of SEQ ID NO:5; cc) E or R at a position corresponding to position 153 of SEQ ID NO:5; dd) W at a position corresponding to position 155 of SEQ ID NO: 5; ee) R or V at a position corresponding to position 156 of SEQ ID NO:5; ff) F at a position corresponding to position 157 of SEQ ID NO: 5; gg) R at a position corresponding to position 158 of SEQ ID NO:5; hh) Q at a position corresponding to position 159 of SEQ ID NO: 5; ii) D or W at a position corresponding to position 160 of SEQ ID NO: 5; jj) W, H, V, or A at a position corresponding to position 162 of SEQ ID NO:5; kk) R at a position corresponding to position 165 of SEQ ID NO:5, and ll) A deaminase comprising at least one H at a position corresponding to position 166 of SEQ ID NO:

5.

142. The deaminase a) N at a position corresponding to position 35 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; b) N at a position corresponding to position 35 of SEQ ID NO:5, and Q at a position corresponding to position 46 of SEQ ID NO:5; c) S at a position corresponding to position 81 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; d) Q at a position corresponding to position 46 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; e) R at a position corresponding to position 156 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; f) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5; g) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at a position corresponding to position 35 of SEQ ID NO:5, Q at a position corresponding to position 46 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; i) a W at a position corresponding to position 155 of SEQ ID NO:5, an R at a position corresponding to position 156 of SEQ ID NO:5, and a W at a position corresponding to position 162 of SEQ ID NO:5; j) N at a position corresponding to position 35 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; k) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; l) M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; m) Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; n) M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; o) N at a position corresponding to position 35, Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; p) S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; q) W at a position corresponding to position 75 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; r) S at a position corresponding to position 81 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; s) S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; t) W at a position corresponding to position 75 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; u) W at a position corresponding to position 155 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; v) A at a position corresponding to position 145 of SEQ ID NO:5 and W at a position corresponding to position 155 of SEQ ID NO:5; w) W at a position corresponding to position 75 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; x) W at a position corresponding to position 75 of SEQ ID NO:5 and A at a position corresponding to position 145 of SEQ ID NO:5; y) A at a position corresponding to position 145 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; z) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; aa) S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; bb) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; cc) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; dd) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; ee) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ff) W at a position corresponding to position 75 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; gg) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a W at a position corresponding to position 155 of SEQ ID NO:5; hh) A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ii) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a D at a position corresponding to position 160 of SEQ ID NO:5; jj) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; kk) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; ll) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; mm) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; nn) W at a position corresponding to position 75 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; oo) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; pp) L at a position corresponding to position 40 of SEQ ID NO: 5, and M at a position corresponding to position 105 of SEQ ID NO: 5; qq) L at a position corresponding to position 40 of SEQ ID NO:5, and V at a position corresponding to position 121 of SEQ ID NO:5; rr) L at a position corresponding to position 40 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; ss) L at a position corresponding to position 40 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; tt) L at a position corresponding to position 40 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; uu) M at a position corresponding to position 105 of SEQ ID NO:5 and V at a position corresponding to position 121 of SEQ ID NO:5; vv) M at a position corresponding to position 105 of SEQ ID NO: 5, and M at a position corresponding to position 145 of SEQ ID NO: 5; ww) M at a position corresponding to position 105 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; xx) M at a position corresponding to position 105 of SEQ ID NO: 5 and W at a position corresponding to position 160 of SEQ ID NO: 5; yy) V at a position corresponding to position 121 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; zz) V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; aaa) V at a position corresponding to position 121 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; bbb) M at a position corresponding to position 145 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ccc) M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; ddd) Q at a position corresponding to position 159 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; eee) S at a position corresponding to position 81 of SEQ ID NO:5, M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; fff) S at a position corresponding to position 81 of SEQ ID NO:5, V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ggg) A at a position corresponding to position 68 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; hhh) G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; iii) S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; jjj) A at a position corresponding to position 68 of SEQ ID NO:5 and G at a position corresponding to position 76 of SEQ ID NO:5; kkk) A at a position corresponding to position 68 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; lll) G at a position corresponding to position 76 of SEQ ID NO:5 and L at a position corresponding to position 117 of SEQ ID NO:5; mmm) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; nnn) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ooo) G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ppp) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; qqq) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; rrr) S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 109 of SEQ ID NO:5, E at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; sss) A at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, A at a position corresponding to position 121 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ttt) Q at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, I at a position corresponding to position 117 of SEQ ID NO:5, Q at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; uuu) Y at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, G at a position corresponding to position 121 of SEQ ID NO:5, L at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; vvv) C at a position corresponding to position 22 of SEQ ID NO:5, A at a position corresponding to position 68 of SEQ ID NO:5, Y at a position corresponding to position 75 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, H at a position corresponding to position 122 of SEQ ID NO:5, Y at a position corresponding to position 138 of SEQ ID NO:5, L at a position corresponding to position 139 of SEQ ID NO:5, A at a position corresponding to position 142 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, R at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, D at a position corresponding to position 160 of SEQ ID NO:5, and H at a position corresponding to position 162 of SEQ ID NO:5; or www) S at a position corresponding to position 81 of SEQ ID NO:

5. The deaminase of claim 141.

143. 143. The deaminase of claim 141 or 142, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

144. The deaminase of claim 141 or 142, wherein the deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

145. 145. The deaminase of any one of claims 141 to 144, wherein the deaminase has improved deaminase activity compared to SEQ ID NO:

5.

146. The deaminase of any one of claims 141 to 145, wherein the deaminase is an adenine deaminase.

147. 1. A nucleic acid molecule comprising a polynucleotide encoding a deaminase having at least 85% sequence identity to SEQ ID NO:5, wherein the deaminase comprises the following amino acid residues: a) a C at the position corresponding to position 2 of SEQ ID NO: 5; b) F or C at the position corresponding to position 22 of SEQ ID NO: 5; c) Q at a position corresponding to position 23 of SEQ ID NO:5; d) N at a position corresponding to position 35 of SEQ ID NO: 5; e) L at a position corresponding to position 40 of SEQ ID NO: 5; f) Q at a position corresponding to position 46 of SEQ ID NO:5; g) A or M at the position corresponding to position 68 of SEQ ID NO: 5; h) H at a position corresponding to position 72 of SEQ ID NO: 5; i) W, A, Q, Y, or D at a position corresponding to position 75 of SEQ ID NO:5; j) G at a position corresponding to position 76 of SEQ ID NO:5; k) S at a position corresponding to position 81 of SEQ ID NO: 5; l) M at a position corresponding to position 105 of SEQ ID NO: 5; m) H at a position corresponding to position 108 of SEQ ID NO: 5; n) L at a position corresponding to position 109 of SEQ ID NO: 5; o) I or L at a position corresponding to position 117 of SEQ ID NO: 5; p) F at a position corresponding to position 120 of SEQ ID NO: 5; q) E, T, A, G, or V at a position corresponding to position 121 of SEQ ID NO: 5; r) I or H at a position corresponding to position 122 of SEQ ID NO: 5; s) K at a position corresponding to position 125 of SEQ ID NO:5; t) A at a position corresponding to position 126 of SEQ ID NO:5; u) H at a position corresponding to position 135 of SEQ ID NO: 5; v) V at a position corresponding to position 137 of SEQ ID NO: 5; w) Y at a position corresponding to position 138 of SEQ ID NO: 5; x) L at a position corresponding to position 139 of SEQ ID NO: 5; y) K or A at a position corresponding to position 142 of SEQ ID NO: 5; z) A, Q, L, or M at a position corresponding to position 145 of SEQ ID NO: 5; aa) K at a position corresponding to position 148 of SEQ ID NO: 5; bb) Q at a position corresponding to position 151 of SEQ ID NO:5; cc) E or R at a position corresponding to position 153 of SEQ ID NO:5; dd) W at a position corresponding to position 155 of SEQ ID NO: 5; ee) R or V at a position corresponding to position 156 of SEQ ID NO:5; ff) F at a position corresponding to position 157 of SEQ ID NO: 5; gg) R at a position corresponding to position 158 of SEQ ID NO:5; hh) Q at a position corresponding to position 159 of SEQ ID NO: 5; ii) D or W at a position corresponding to position 160 of SEQ ID NO: 5; jj) W, H, V, or A at a position corresponding to position 162 of SEQ ID NO:5; kk) R at a position corresponding to position 165 of SEQ ID NO:5, and ll) A nucleic acid molecule comprising at least one H at a position corresponding to position 166 of SEQ ID NO:

5.

148. The deaminase a) N at a position corresponding to position 35 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; b) N at a position corresponding to position 35 of SEQ ID NO:5, and Q at a position corresponding to position 46 of SEQ ID NO:5; c) S at a position corresponding to position 81 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; d) Q at a position corresponding to position 46 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; e) R at a position corresponding to position 156 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; f) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5; g) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at a position corresponding to position 35 of SEQ ID NO:5, Q at a position corresponding to position 46 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; i) a W at a position corresponding to position 155 of SEQ ID NO:5, an R at a position corresponding to position 156 of SEQ ID NO:5, and a W at a position corresponding to position 162 of SEQ ID NO:5; j) N at a position corresponding to position 35 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; k) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; l) M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; m) Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; n) M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; o) N at a position corresponding to position 35, Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; p) S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; q) W at a position corresponding to position 75 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; r) S at a position corresponding to position 81 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; s) S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; t) W at a position corresponding to position 75 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; u) W at a position corresponding to position 155 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; v) A at a position corresponding to position 145 of SEQ ID NO:5 and W at a position corresponding to position 155 of SEQ ID NO:5; w) W at a position corresponding to position 75 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; x) W at a position corresponding to position 75 of SEQ ID NO:5 and A at a position corresponding to position 145 of SEQ ID NO:5; y) A at a position corresponding to position 145 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; z) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; aa) S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; bb) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; cc) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; dd) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; ee) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ff) W at a position corresponding to position 75 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; gg) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a W at a position corresponding to position 155 of SEQ ID NO:5; hh) A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ii) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a D at a position corresponding to position 160 of SEQ ID NO:5; jj) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; kk) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; ll) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; mm) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; nn) W at a position corresponding to position 75 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; oo) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; pp) L at a position corresponding to position 40 of SEQ ID NO: 5, and M at a position corresponding to position 105 of SEQ ID NO: 5; qq) L at a position corresponding to position 40 of SEQ ID NO:5, and V at a position corresponding to position 121 of SEQ ID NO:5; rr) L at a position corresponding to position 40 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; ss) L at a position corresponding to position 40 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; tt) L at a position corresponding to position 40 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; uu) M at a position corresponding to position 105 of SEQ ID NO:5 and V at a position corresponding to position 121 of SEQ ID NO:5; vv) M at a position corresponding to position 105 of SEQ ID NO: 5, and M at a position corresponding to position 145 of SEQ ID NO: 5; ww) M at a position corresponding to position 105 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; xx) M at a position corresponding to position 105 of SEQ ID NO: 5 and W at a position corresponding to position 160 of SEQ ID NO: 5; yy) V at a position corresponding to position 121 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; zz) V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; aaa) V at a position corresponding to position 121 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; bbb) M at a position corresponding to position 145 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ccc) M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; ddd) Q at a position corresponding to position 159 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; eee) S at a position corresponding to position 81 of SEQ ID NO:5, M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; fff) S at a position corresponding to position 81 of SEQ ID NO:5, V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ggg) A at a position corresponding to position 68 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; hhh) G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; iii) S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; jjj) A at a position corresponding to position 68 of SEQ ID NO:5 and G at a position corresponding to position 76 of SEQ ID NO:5; kkk) A at a position corresponding to position 68 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; lll) G at a position corresponding to position 76 of SEQ ID NO:5 and L at a position corresponding to position 117 of SEQ ID NO:5; mmm) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; nnn) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ooo) G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ppp) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; qqq) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; rrr) S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 109 of SEQ ID NO:5, E at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; sss) A at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, A at a position corresponding to position 121 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ttt) Q at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, I at a position corresponding to position 117 of SEQ ID NO:5, Q at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; uuu) Y at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, G at a position corresponding to position 121 of SEQ ID NO:5, L at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; vvv) C at a position corresponding to position 22 of SEQ ID NO:5, A at a position corresponding to position 68 of SEQ ID NO:5, Y at a position corresponding to position 75 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, H at a position corresponding to position 122 of SEQ ID NO:5, Y at a position corresponding to position 138 of SEQ ID NO:5, L at a position corresponding to position 139 of SEQ ID NO:5, A at a position corresponding to position 142 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, R at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, D at a position corresponding to position 160 of SEQ ID NO:5, and H at a position corresponding to position 162 of SEQ ID NO:5; or www) The nucleic acid molecule of claim 147, comprising an S at a position corresponding to position 81 of SEQ ID NO:

5.

149. 149. The nucleic acid molecule of claim 147 or 148, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

150. 149. The nucleic acid molecule of claim 147 or 148, wherein the deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

151. 151. The nucleic acid molecule of any one of claims 147 to 150, wherein the deaminase has improved deaminase activity compared to SEQ ID NO:

5.

152. The nucleic acid molecule of any one of claims 147 to 151, wherein the deaminase is adenine deaminase.

153. 153. The nucleic acid molecule of any one of claims 147 to 152, further comprising a heterologous promoter operably linked to said polynucleotide.

154. A vector comprising the nucleic acid molecule of any one of claims 147 to 153.

155. 155. A cell comprising the deaminase of any one of claims 141 to 146, the nucleic acid molecule of any one of claims 147 to 153, or the vector of claim 154.

156. The cell of claim 155, wherein the cell is a mammalian cell.

157. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and the deaminase of any one of claims 141 to 146, the nucleic acid molecule of any one of claims 147 to 153, the vector of claim 154, or the cell of claim 156.

158. 1. A fusion protein comprising a DNA-binding polypeptide and a deaminase having at least 85% sequence identity to SEQ ID NO:5, wherein the deaminase comprises the following amino acid residues: a) a C at the position corresponding to position 2 of SEQ ID NO: 5; b) F or C at the position corresponding to position 22 of SEQ ID NO: 5; c) Q at a position corresponding to position 23 of SEQ ID NO:5; d) N at a position corresponding to position 35 of SEQ ID NO: 5; e) L at a position corresponding to position 40 of SEQ ID NO: 5; f) Q at a position corresponding to position 46 of SEQ ID NO:5; g) A or M at the position corresponding to position 68 of SEQ ID NO: 5; h) H at a position corresponding to position 72 of SEQ ID NO: 5; i) W, A, Q, Y, or D at a position corresponding to position 75 of SEQ ID NO:5; j) G at a position corresponding to position 76 of SEQ ID NO:5; k) S at a position corresponding to position 81 of SEQ ID NO: 5; l) M at a position corresponding to position 105 of SEQ ID NO: 5; m) H at a position corresponding to position 108 of SEQ ID NO: 5; n) L at a position corresponding to position 109 of SEQ ID NO: 5; o) I or L at a position corresponding to position 117 of SEQ ID NO: 5; p) F at a position corresponding to position 120 of SEQ ID NO: 5; q) E, T, A, G, or V at a position corresponding to position 121 of SEQ ID NO: 5; r) I or H at a position corresponding to position 122 of SEQ ID NO: 5; s) K at a position corresponding to position 125 of SEQ ID NO:5; t) A at a position corresponding to position 126 of SEQ ID NO:5; u) H at a position corresponding to position 135 of SEQ ID NO: 5; v) V at a position corresponding to position 137 of SEQ ID NO: 5; w) Y at a position corresponding to position 138 of SEQ ID NO: 5; x) L at a position corresponding to position 139 of SEQ ID NO: 5; y) K or A at a position corresponding to position 142 of SEQ ID NO: 5; z) A, Q, L, or M at a position corresponding to position 145 of SEQ ID NO:5; aa) K at a position corresponding to position 148 of SEQ ID NO: 5; bb) Q at a position corresponding to position 151 of SEQ ID NO:5; cc) E or R at a position corresponding to position 153 of SEQ ID NO: 5; dd) W at a position corresponding to position 155 of SEQ ID NO: 5; ee) R or V at a position corresponding to position 156 of SEQ ID NO:5; ff) F at a position corresponding to position 157 of SEQ ID NO: 5; gg) R at a position corresponding to position 158 of SEQ ID NO:5; hh) Q at a position corresponding to position 159 of SEQ ID NO: 5; ii) D or W at a position corresponding to position 160 of SEQ ID NO: 5; jj) W, H, V, or A at a position corresponding to position 162 of SEQ ID NO:5; kk) R at a position corresponding to position 165 of SEQ ID NO:5, and ll) A fusion protein comprising at least one of H at a position corresponding to position 166 of SEQ ID NO:

5.

159. The deaminase a) N at a position corresponding to position 35 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; b) N at a position corresponding to position 35 of SEQ ID NO:5, and Q at a position corresponding to position 46 of SEQ ID NO:5; c) S at a position corresponding to position 81 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; d) Q at a position corresponding to position 46 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; e) R at a position corresponding to position 156 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; f) A at a position corresponding to position 68 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; g) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at a position corresponding to position 35 of SEQ ID NO:5, Q at a position corresponding to position 46 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; i) a W at a position corresponding to position 155 of SEQ ID NO:5, an R at a position corresponding to position 156 of SEQ ID NO:5, and a W at a position corresponding to position 162 of SEQ ID NO:5; j) N at a position corresponding to position 35 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; k) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; l) M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; m) Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; n) M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; o) N at a position corresponding to position 35, Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; p) S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; q) W at a position corresponding to position 75 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; r) S at a position corresponding to position 81 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; s) S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; t) W at a position corresponding to position 75 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; u) W at a position corresponding to position 155 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; v) A at a position corresponding to position 145 of SEQ ID NO:5 and W at a position corresponding to position 155 of SEQ ID NO:5; w) W at a position corresponding to position 75 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; x) W at a position corresponding to position 75 of SEQ ID NO:5 and A at a position corresponding to position 145 of SEQ ID NO:5; y) A at a position corresponding to position 145 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; z) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; aa) S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; bb) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; cc) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; dd) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; ee) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ff) W at a position corresponding to position 75 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; gg) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a W at a position corresponding to position 155 of SEQ ID NO:5; hh) A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ii) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a D at a position corresponding to position 160 of SEQ ID NO:5; jj) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; kk) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; ll) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; mm) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; nn) W at a position corresponding to position 75 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; oo) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; pp) L at a position corresponding to position 40 of SEQ ID NO: 5, and M at a position corresponding to position 105 of SEQ ID NO: 5; qq) L at a position corresponding to position 40 of SEQ ID NO:5, and V at a position corresponding to position 121 of SEQ ID NO:5; rr) L at a position corresponding to position 40 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; ss) L at a position corresponding to position 40 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; tt) L at a position corresponding to position 40 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; uu) M at a position corresponding to position 105 of SEQ ID NO:5 and V at a position corresponding to position 121 of SEQ ID NO:5; vv) M at a position corresponding to position 105 of SEQ ID NO: 5, and M at a position corresponding to position 145 of SEQ ID NO: 5; ww) M at a position corresponding to position 105 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; xx) M at a position corresponding to position 105 of SEQ ID NO: 5 and W at a position corresponding to position 160 of SEQ ID NO: 5; yy) V at a position corresponding to position 121 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; zz) V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; aaa) V at a position corresponding to position 121 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; bbb) M at a position corresponding to position 145 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ccc) M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; ddd) Q at a position corresponding to position 159 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; eee) S at a position corresponding to position 81 of SEQ ID NO:5, M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; fff) S at a position corresponding to position 81 of SEQ ID NO:5, V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ggg) A at a position corresponding to position 68 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; hhh) G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; iii) S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; jjj) A at a position corresponding to position 68 of SEQ ID NO:5 and G at a position corresponding to position 76 of SEQ ID NO:5; kkk) A at a position corresponding to position 68 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; lll) G at a position corresponding to position 76 of SEQ ID NO:5 and L at a position corresponding to position 117 of SEQ ID NO:5; mmm) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; nnn) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ooo) G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ppp) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; qqq) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; rrr) S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 109 of SEQ ID NO:5, E at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; sss) A at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, A at a position corresponding to position 121 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ttt) Q at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, I at a position corresponding to position 117 of SEQ ID NO:5, Q at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; uuu) Y at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, G at a position corresponding to position 121 of SEQ ID NO:5, L at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; vvv) C at a position corresponding to position 22 of SEQ ID NO:5, A at a position corresponding to position 68 of SEQ ID NO:5, Y at a position corresponding to position 75 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, H at a position corresponding to position 122 of SEQ ID NO:5, Y at a position corresponding to position 138 of SEQ ID NO:5, L at a position corresponding to position 139 of SEQ ID NO:5, A at a position corresponding to position 142 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, R at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, D at a position corresponding to position 160 of SEQ ID NO:5, and H at a position corresponding to position 162 of SEQ ID NO:5; or www) S at a position corresponding to position 81 of SEQ ID NO:

5.

160. 160. The fusion protein of claim 158 or 159, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

161. 160. The fusion protein of claim 158 or 159, wherein the deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

162. 162. The fusion protein of any one of claims 158 to 161, wherein the deaminase has improved deaminase activity compared to SEQ ID NO:

5.

163. 163. The fusion protein of any one of claims 158 to 162, wherein the deaminase is adenine deaminase.

164. 164. The fusion protein of any one of claims 158 to 163, wherein the DNA-binding polypeptide is an RNA-guided nuclease (RGN) polypeptide.

165. The fusion protein of claim 164, wherein the RGN is an RGN nickase.

166. 165. The fusion protein of claim 164, wherein the RGN has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699.

167. The fusion protein of claim 165, wherein the RGN nickase has an amino acid sequence set forth as any one of SEQ ID NOs: 49-52 and 698.

168. the fusion protein a) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 319 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; b) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; c) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 406 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; d) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 596 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; e) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 597 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; f) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 303 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; g) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 366 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; and h) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 720, fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO:

49.

169. the fusion protein a) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 319 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; b) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 375 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; c) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 406 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; d) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 596 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; e) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 597 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; f) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 303 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; g) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 366 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; and h) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 720 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO:

49.

170. 160. The fusion protein of claim 158 or 159, wherein the fusion protein has a sequence set forth as any one of SEQ ID NOs: 577, 578, 582, 584, 588, 593, 594, 604, 605, 616, and 729.

171. 171. A ribonucleoprotein (RNP) complex comprising the fusion protein of any one of claims 158 to 170 and a guide RNA bound to the fusion protein.

172. 1. A nucleic acid molecule comprising a polynucleotide encoding a fusion protein comprising a DNA-binding polypeptide and a deaminase, wherein the deaminase comprises the following amino acid residues: a) a C at the position corresponding to position 2 of SEQ ID NO: 5; b) F or C at the position corresponding to position 22 of SEQ ID NO: 5; c) Q at a position corresponding to position 23 of SEQ ID NO:5; d) N at a position corresponding to position 35 of SEQ ID NO: 5; e) L at a position corresponding to position 40 of SEQ ID NO: 5; f) Q at a position corresponding to position 46 of SEQ ID NO:5; g) A or M at the position corresponding to position 68 of SEQ ID NO: 5; h) H at a position corresponding to position 72 of SEQ ID NO: 5; i) W, A, Q, Y, or D at a position corresponding to position 75 of SEQ ID NO:5; j) G at a position corresponding to position 76 of SEQ ID NO:5; k) S at a position corresponding to position 81 of SEQ ID NO: 5; l) M at a position corresponding to position 105 of SEQ ID NO: 5; m) H at a position corresponding to position 108 of SEQ ID NO: 5; n) L at a position corresponding to position 109 of SEQ ID NO: 5; o) I or L at a position corresponding to position 117 of SEQ ID NO: 5; p) F at a position corresponding to position 120 of SEQ ID NO: 5; q) E, T, A, G, or V at a position corresponding to position 121 of SEQ ID NO: 5; r) I or H at a position corresponding to position 122 of SEQ ID NO: 5; s) K at a position corresponding to position 125 of SEQ ID NO:5; t) A at a position corresponding to position 126 of SEQ ID NO:5; u) H at a position corresponding to position 135 of SEQ ID NO: 5; v) V at a position corresponding to position 137 of SEQ ID NO: 5; w) Y at a position corresponding to position 138 of SEQ ID NO: 5; x) L at a position corresponding to position 139 of SEQ ID NO: 5; y) K or A at a position corresponding to position 142 of SEQ ID NO: 5; z) A, Q, L, or M at a position corresponding to position 145 of SEQ ID NO: 5; aa) K at a position corresponding to position 148 of SEQ ID NO: 5; bb) Q at a position corresponding to position 151 of SEQ ID NO:5; cc) E or R at a position corresponding to position 153 of SEQ ID NO:5; dd) W at a position corresponding to position 155 of SEQ ID NO: 5; ee) R or V at a position corresponding to position 156 of SEQ ID NO:5; ff) F at a position corresponding to position 157 of SEQ ID NO: 5; gg) R at a position corresponding to position 158 of SEQ ID NO:5; hh) Q at a position corresponding to position 159 of SEQ ID NO: 5; ii) D or W at a position corresponding to position 160 of SEQ ID NO: 5; jj) W, H, V, or A at a position corresponding to position 162 of SEQ ID NO:5; kk) R at a position corresponding to position 165 of SEQ ID NO:5, and ll) A nucleic acid molecule comprising at least one H at a position corresponding to position 166 of SEQ ID NO:

5.

173. The deaminase a) N at a position corresponding to position 35 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; b) N at a position corresponding to position 35 of SEQ ID NO:5, and Q at a position corresponding to position 46 of SEQ ID NO:5; c) S at a position corresponding to position 81 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; d) Q at a position corresponding to position 46 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; e) R at a position corresponding to position 156 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; f) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5; g) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at a position corresponding to position 35 of SEQ ID NO:5, Q at a position corresponding to position 46 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; i) a W at a position corresponding to position 155 of SEQ ID NO:5, an R at a position corresponding to position 156 of SEQ ID NO:5, and a W at a position corresponding to position 162 of SEQ ID NO:5; j) N at a position corresponding to position 35 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; k) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; l) M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; m) Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; n) M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; o) N at a position corresponding to position 35, Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; p) S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; q) W at a position corresponding to position 75 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; r) S at a position corresponding to position 81 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; s) S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; t) W at a position corresponding to position 75 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; u) W at a position corresponding to position 155 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; v) A at a position corresponding to position 145 of SEQ ID NO:5 and W at a position corresponding to position 155 of SEQ ID NO:5; w) W at a position corresponding to position 75 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; x) W at a position corresponding to position 75 of SEQ ID NO:5 and A at a position corresponding to position 145 of SEQ ID NO:5; y) A at a position corresponding to position 145 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; z) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; aa) S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; bb) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; cc) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; dd) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; ee) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ff) W at a position corresponding to position 75 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; gg) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a W at a position corresponding to position 155 of SEQ ID NO:5; hh) A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ii) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a D at a position corresponding to position 160 of SEQ ID NO:5; jj) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; kk) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; ll) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; mm) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; nn) W at a position corresponding to position 75 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; oo) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; pp) L at a position corresponding to position 40 of SEQ ID NO: 5, and M at a position corresponding to position 105 of SEQ ID NO: 5; qq) L at a position corresponding to position 40 of SEQ ID NO:5, and V at a position corresponding to position 121 of SEQ ID NO:5; rr) L at a position corresponding to position 40 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; ss) L at a position corresponding to position 40 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; tt) L at a position corresponding to position 40 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; uu) M at a position corresponding to position 105 of SEQ ID NO:5 and V at a position corresponding to position 121 of SEQ ID NO:5; vv) M at a position corresponding to position 105 of SEQ ID NO: 5, and M at a position corresponding to position 145 of SEQ ID NO: 5; ww) M at a position corresponding to position 105 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; xx) M at a position corresponding to position 105 of SEQ ID NO: 5 and W at a position corresponding to position 160 of SEQ ID NO: 5; yy) V at a position corresponding to position 121 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; zz) V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; aaa) V at a position corresponding to position 121 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; bbb) M at a position corresponding to position 145 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ccc) M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; ddd) Q at a position corresponding to position 159 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; eee) S at a position corresponding to position 81 of SEQ ID NO:5, M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; fff) S at a position corresponding to position 81 of SEQ ID NO:5, V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ggg) A at a position corresponding to position 68 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; hhh) G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; iii) S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; jjj) A at a position corresponding to position 68 of SEQ ID NO:5 and G at a position corresponding to position 76 of SEQ ID NO:5; kkk) A at a position corresponding to position 68 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; lll) G at a position corresponding to position 76 of SEQ ID NO:5 and L at a position corresponding to position 117 of SEQ ID NO:5; mmm) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; nnn) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ooo) G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ppp) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; qqq) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; rrr) S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 109 of SEQ ID NO:5, E at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; sss) A at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, A at a position corresponding to position 121 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ttt) Q at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, I at a position corresponding to position 117 of SEQ ID NO:5, Q at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; uuu) Y at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, G at a position corresponding to position 121 of SEQ ID NO:5, L at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; vvv) C at a position corresponding to position 22 of SEQ ID NO:5, A at a position corresponding to position 68 of SEQ ID NO:5, Y at a position corresponding to position 75 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, H at a position corresponding to position 122 of SEQ ID NO:5, Y at a position corresponding to position 138 of SEQ ID NO:5, L at a position corresponding to position 139 of SEQ ID NO:5, A at a position corresponding to position 142 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, R at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, D at a position corresponding to position 160 of SEQ ID NO:5, and H at a position corresponding to position 162 of SEQ ID NO:5; or www) The nucleic acid molecule of claim 172, comprising an S at a position corresponding to position 81 of SEQ ID NO:

5.

174. 174. The nucleic acid molecule of claim 172 or 173, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

175. 174. The nucleic acid molecule of claim 172 or 173, wherein the deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

176. 176. The nucleic acid molecule of any one of claims 172 to 175, wherein the deaminase has improved deaminase activity compared to SEQ ID NO:

5.

177. The nucleic acid molecule of any one of claims 172 to 176, wherein the deaminase is adenine deaminase.

178. 178. The nucleic acid molecule of any one of claims 172 to 177, wherein the DNA-binding polypeptide is an RNA-guided nuclease (RGN) polypeptide.

179. The nucleic acid molecule of claim 178, wherein the RGN is an RGN nickase.

180. 179. The nucleic acid molecule of claim 178, wherein the RGN has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699.

181. 180. The nucleic acid molecule of claim 179, wherein the RGN nickase has an amino acid sequence set forth as any one of SEQ ID NOs: 49-52 and 698.

182. the fusion protein a) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 319 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; b) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; c) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 406 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; d) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 596 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; e) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 597 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; f) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 303 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; g) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 366 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; and h) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 720, fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO:

49.

183. the fusion protein a) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 319 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; b) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 375 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; c) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 406 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; d) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 596 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; e) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 597 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; f) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 303 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; g) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 366 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; and h) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 720, fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO:

49.

184. 174. The nucleic acid molecule of claim 172 or 173, wherein the fusion protein has a sequence set forth as any one of SEQ ID NOs: 577, 578, 582, 584, 588, 593, 594, 604, 605, 616, and 729.

185. A vector comprising the nucleic acid molecule of any one of claims 172 to 184.

186. 185. A vector comprising the nucleic acid molecule of any one of claims 172 to 184, further comprising at least one nucleotide sequence encoding a guide RNA capable of hybridizing to a non-target strand of the target sequence.

187. A cell comprising the fusion protein of any one of claims 158 to 170 or the RNP complex of claim 171.

188. 171. A cell comprising the fusion protein of any one of claims 158 to 170, wherein the cell further comprises a guide RNA.

189. A cell comprising the nucleic acid molecule of any one of claims 172 to 184.

190. A cell comprising the vector of claims 185 and 186.

191. The cell of any one of claims 187 to 190, wherein the cell is a mammalian cell.

192. The cell of any one of claims 187 to 190, wherein the cell is a plant cell.

193. 193. A plant or seed comprising the cell of claim 192.

194. A pharmaceutical composition comprising a pharmaceutically acceptable carrier, a fusion protein according to any one of claims 158 to 170, an RNP complex according to claim 171, a nucleic acid molecule according to any one of claims 172 to 184, a vector according to claim 185 or 186, or a cell according to claim 191.

195. 186。 A method for producing an RGN fusion ribonucleoprotein complex, comprising: introducing into a cell a nucleic acid molecule described in any one of claims 172 to 184 or a vector described in claim 185, and a nucleic acid molecule comprising an expression cassette encoding a guide RNA or a vector described in claim 186; and culturing the cell under conditions in which the fusion protein and the gRNA are expressed to form an RGN fusion ribonucleoprotein complex.

196. 1. A system for modifying a target DNA molecule comprising a target DNA sequence, comprising: a) a fusion protein comprising an RNA-guided nuclease polypeptide and a deaminase, wherein the deaminase has an amino acid sequence having at least 85% sequence identity to SEQ ID NO:5, or a nucleotide sequence encoding the fusion protein, wherein the deaminase comprises the following amino acid residues: i) a C at the position corresponding to position 2 of SEQ ID NO: 5; ii) F or C at a position corresponding to position 22 of SEQ ID NO: 5; iii) Q at a position corresponding to position 23 of SEQ ID NO:5; iv) N at a position corresponding to position 35 of SEQ ID NO: 5; v) L at a position corresponding to position 40 of SEQ ID NO: 5; vi) Q at a position corresponding to position 46 of SEQ ID NO:5; vii) A or M at the position corresponding to position 68 of SEQ ID NO:5; viii) H at a position corresponding to position 72 of SEQ ID NO:5; ix) W, A, Q, Y, or D at a position corresponding to position 75 of SEQ ID NO:5; x) G at a position corresponding to position 76 of SEQ ID NO:5; xi) S at a position corresponding to position 81 of SEQ ID NO:5; xii) M at a position corresponding to position 105 of SEQ ID NO:5; xiii) H at a position corresponding to position 108 of SEQ ID NO:5; xiv) L at a position corresponding to position 109 of SEQ ID NO: 5; xv) I or L at a position corresponding to position 117 of SEQ ID NO: 5; xvi) F at a position corresponding to position 120 of SEQ ID NO:5; xvii) E, T, A, G, or V at a position corresponding to position 121 of SEQ ID NO:5; xviii) I or H at a position corresponding to position 122 of SEQ ID NO:5; xix) K at a position corresponding to position 125 of SEQ ID NO:5; xx) A at a position corresponding to position 126 of SEQ ID NO: 5; xxi) H at a position corresponding to position 135 of SEQ ID NO:5; xxii) V at a position corresponding to position 137 of SEQ ID NO:5; xxiii) Y at a position corresponding to position 138 of SEQ ID NO:5; xxiv) L at a position corresponding to position 139 of SEQ ID NO: 5; xxv) K or A at a position corresponding to position 142 of SEQ ID NO: 5; xxvi) A, Q, L, or M at a position corresponding to position 145 of SEQ ID NO:5; xxvii) K at a position corresponding to position 148 of SEQ ID NO:5; xxviii) Q at a position corresponding to position 151 of SEQ ID NO:5; xxix) E or R at a position corresponding to position 153 of SEQ ID NO: 5; xxx) W at a position corresponding to position 155 of SEQ ID NO: 5; xxxi) R or V at a position corresponding to position 156 of SEQ ID NO:5; xxxii) F at a position corresponding to position 157 of SEQ ID NO:5; xxxiii) R at a position corresponding to position 158 of SEQ ID NO:5; xxxiv) Q at a position corresponding to position 159 of SEQ ID NO:5; xxxv) D or W at a position corresponding to position 160 of SEQ ID NO: 5; xxxvi) W, H, V, or A at a position corresponding to position 162 of SEQ ID NO:5; xxxvii) R at a position corresponding to position 165 of SEQ ID NO:5, and xxxviii) a fusion protein comprising at least one of H at a position corresponding to position 166 of SEQ ID NO:5; b) one or more guide RNAs capable of hybridizing to non-target strands of the target DNA sequence, or one or more nucleotide sequences encoding the one or more guide RNAs; wherein the one or more guide RNAs are capable of forming a complex with the fusion protein to guide the fusion protein to bind to the target DNA sequence and modify the target DNA molecule.

197. The deaminase a) N at a position corresponding to position 35 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; b) N at a position corresponding to position 35 of SEQ ID NO:5, and Q at a position corresponding to position 46 of SEQ ID NO:5; c) S at a position corresponding to position 81 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; d) Q at a position corresponding to position 46 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; e) R at a position corresponding to position 156 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; f) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5; g) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at a position corresponding to position 35 of SEQ ID NO:5, Q at a position corresponding to position 46 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; i) a W at a position corresponding to position 155 of SEQ ID NO:5, an R at a position corresponding to position 156 of SEQ ID NO:5, and a W at a position corresponding to position 162 of SEQ ID NO:5; j) N at a position corresponding to position 35 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; k) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; l) M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; m) Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; n) M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; o) N at a position corresponding to position 35, Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; p) S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; q) W at a position corresponding to position 75 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; r) S at a position corresponding to position 81 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; s) S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; t) W at a position corresponding to position 75 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; u) W at a position corresponding to position 155 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; v) A at a position corresponding to position 145 of SEQ ID NO:5 and W at a position corresponding to position 155 of SEQ ID NO:5; w) W at a position corresponding to position 75 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; x) W at a position corresponding to position 75 of SEQ ID NO:5 and A at a position corresponding to position 145 of SEQ ID NO:5; y) A at a position corresponding to position 145 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; z) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; aa) S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; bb) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; cc) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; dd) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; ee) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ff) W at a position corresponding to position 75 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; gg) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a W at a position corresponding to position 155 of SEQ ID NO:5; hh) A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ii) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a D at a position corresponding to position 160 of SEQ ID NO:5; jj) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; kk) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; ll) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; mm) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; nn) W at a position corresponding to position 75 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; oo) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; pp) L at a position corresponding to position 40 of SEQ ID NO: 5, and M at a position corresponding to position 105 of SEQ ID NO: 5; qq) L at a position corresponding to position 40 of SEQ ID NO:5, and V at a position corresponding to position 121 of SEQ ID NO:5; rr) L at a position corresponding to position 40 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; ss) L at a position corresponding to position 40 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; tt) L at a position corresponding to position 40 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; uu) M at a position corresponding to position 105 of SEQ ID NO:5 and V at a position corresponding to position 121 of SEQ ID NO:5; vv) M at a position corresponding to position 105 of SEQ ID NO: 5, and M at a position corresponding to position 145 of SEQ ID NO: 5; ww) M at a position corresponding to position 105 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; xx) M at a position corresponding to position 105 of SEQ ID NO: 5 and W at a position corresponding to position 160 of SEQ ID NO: 5; yy) V at a position corresponding to position 121 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; zz) V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; aaa) V at a position corresponding to position 121 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; bbb) M at a position corresponding to position 145 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ccc) M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; ddd) Q at a position corresponding to position 159 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; eee) S at a position corresponding to position 81 of SEQ ID NO:5, M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; fff) S at a position corresponding to position 81 of SEQ ID NO:5, V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ggg) A at a position corresponding to position 68 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; hhh) G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; iii) S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; jjj) A at a position corresponding to position 68 of SEQ ID NO:5 and G at a position corresponding to position 76 of SEQ ID NO:5; kkk) A at a position corresponding to position 68 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; lll) G at a position corresponding to position 76 of SEQ ID NO:5 and L at a position corresponding to position 117 of SEQ ID NO:5; mmm) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; nnn) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ooo) G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ppp) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; qqq) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; rrr) S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 109 of SEQ ID NO:5, E at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; sss) A at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, A at a position corresponding to position 121 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ttt) Q at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, I at a position corresponding to position 117 of SEQ ID NO:5, Q at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; uuu) Y at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, G at a position corresponding to position 121 of SEQ ID NO:5, L at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; vvv) C at a position corresponding to position 22 of SEQ ID NO:5, A at a position corresponding to position 68 of SEQ ID NO:5, Y at a position corresponding to position 75 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, H at a position corresponding to position 122 of SEQ ID NO:5, Y at a position corresponding to position 138 of SEQ ID NO:5, L at a position corresponding to position 139 of SEQ ID NO:5, A at a position corresponding to position 142 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, R at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, D at a position corresponding to position 160 of SEQ ID NO:5, and H at a position corresponding to position 162 of SEQ ID NO:5; or www) S at a position corresponding to position 81 of SEQ ID NO:

5.

198. 198. The system of claim 196 or 197, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

199. 198. The system of claim 196 or 197, wherein the deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

200. 200. The system of any one of claims 196 to 199, wherein the deaminase has improved deaminase activity compared to SEQ ID NO:

5.

201. 201. The system of any one of claims 196 to 200, wherein the deaminase is adenine deaminase.

202. 202. The system of any one of Claims 196-201, wherein at least one of the nucleotide sequences encoding one or more guide RNAs and the nucleotide sequence encoding the fusion protein is operably linked to a promoter heterologous to said nucleotide sequence.

203. 203. The system of any one of claims 196 to 202, wherein the target DNA sequence is a eukaryotic target DNA sequence.

204. 204. The system of any one of claims 196 to 203, wherein the target DNA sequence is located adjacent to a protospacer adjacent motif (PAM) recognized by the RGN.

205. 205. The system of any one of claims 196 to 204, wherein the RGN of the fusion protein has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699.

206. The system of any one of claims 196 to 205, wherein the RGN of the fusion protein is an RGN nickase.

207. The system of claim 206, wherein the RGN nickase has an amino acid sequence set forth as any one of SEQ ID NOs: 49-52 and 698.

208. the fusion protein a) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 319 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; b) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; c) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 406 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; d) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 596 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; e) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 597 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; f) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 303 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; g) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 366 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; and h) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 720, fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO:

49.

209. the fusion protein a) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 319 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; b) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 375 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; c) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 406 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; d) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 596 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; e) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 597 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; f) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 303 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; g) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 366 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; and h) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 720, fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO:

49.

210. 198. The system of claim 196 or 197, wherein the fusion protein has a sequence set forth as any one of SEQ ID NOs: 577, 578, 582, 584, 588, 593, 594, 604, 605, 616, and 729.

211. A cell comprising the system described in any one of claims 196 to 210.

212. 212. The cell of claim 211, wherein the cell is a mammalian cell.

213. 212. The cell of claim 211, wherein the cell is a plant cell.

214. 214. A plant or seed comprising the cell of claim 213.

215. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and the system of any one of claims 196 to 210 or the cell of claim 212.

216. 211. A method for modifying a target DNA molecule comprising a target DNA sequence, the method comprising delivering a system according to any one of claims 196 to 210 to said target DNA molecule or to a cell comprising said target DNA molecule.

217. 1. A method for modifying a target DNA molecule comprising a target sequence, comprising: a) i) one or more guide RNAs capable of hybridizing to the non-target strand of the target DNA sequence; and ii) assembling an RGN-deaminase ribonucleotide complex in vitro by combining a fusion protein comprising an RNA-guided nuclease polypeptide (RGN) and at least one deaminase, wherein the deaminase has an amino acid sequence having at least 85% sequence identity to SEQ ID NO:5; Under conditions suitable for the formation of the RGN-deaminase ribonucleotide complex, the deaminase reacts with the following amino acid residues: A) C at the position corresponding to position 2 of SEQ ID NO: 5; B) F or C at a position corresponding to position 22 of SEQ ID NO: 5; C) Q at a position corresponding to position 23 of SEQ ID NO:5; D) N at a position corresponding to position 35 of SEQ ID NO:5; E) L at a position corresponding to position 40 of SEQ ID NO:5; F) Q at a position corresponding to position 46 of SEQ ID NO:5; G) A or M at a position corresponding to position 68 of SEQ ID NO: 5; H) H at a position corresponding to position 72 of SEQ ID NO: 5; II) W, A, Q, Y, or D at a position corresponding to position 75 of SEQ ID NO:5; J) G at a position corresponding to position 76 of SEQ ID NO:5; K) S at a position corresponding to position 81 of SEQ ID NO:5; LI) M at a position corresponding to position 105 of SEQ ID NO:5; MI) H at a position corresponding to position 108 of SEQ ID NO: 5; N) L at a position corresponding to position 109 of SEQ ID NO: 5; O) I or L at a position corresponding to position 117 of SEQ ID NO: 5; P) F at a position corresponding to position 120 of SEQ ID NO: 5; Q) E, T, A, G, or V at a position corresponding to position 121 of SEQ ID NO: 5; R) I or H at a position corresponding to position 122 of SEQ ID NO:5; S) K at a position corresponding to position 125 of SEQ ID NO:5; T) A at a position corresponding to position 126 of SEQ ID NO:5; U) H at a position corresponding to position 135 of SEQ ID NO: 5; V) V at a position corresponding to position 137 of SEQ ID NO: 5; W) Y at a position corresponding to position 138 of SEQ ID NO: 5; X) L at a position corresponding to position 139 of SEQ ID NO: 5; Y) K or A at a position corresponding to position 142 of SEQ ID NO: 5; Z) A, Q, L, or M at a position corresponding to position 145 of SEQ ID NO: 5; AA) K at a position corresponding to position 148 of SEQ ID NO:5; DD) Q at a position corresponding to position 151 of SEQ ID NO: 5; EE) E or R at a position corresponding to position 153 of SEQ ID NO: 5; DD) W at a position corresponding to position 155 of SEQ ID NO: 5; EE) R or V at a position corresponding to position 156 of SEQ ID NO:5; GG) F at a position corresponding to position 157 of SEQ ID NO:5; GG) R at a position corresponding to position 158 of SEQ ID NO:5; HH) Q at a position corresponding to position 159 of SEQ ID NO:5; II) D or W at a position corresponding to position 160 of SEQ ID NO:5; KK) W, H, V, or A at a position corresponding to position 162 of SEQ ID NO:5; KK) R at a position corresponding to position 165 of SEQ ID NO:5, and LL) having at least one H at a position corresponding to position 166 of SEQ ID NO:5; b) contacting the target DNA molecule or a cell containing the target DNA molecule with the in vitro assembled RGN-ribonucleotide complex; wherein the one or more guide RNAs hybridize to a non-target strand of the target DNA sequence, thereby directing the fusion protein to bind to the target DNA sequence, resulting in modification of the target DNA sequence.

218. The deaminase a) N at a position corresponding to position 35 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; b) N at a position corresponding to position 35 of SEQ ID NO:5, and Q at a position corresponding to position 46 of SEQ ID NO:5; c) S at a position corresponding to position 81 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; d) Q at a position corresponding to position 46 of SEQ ID NO:5, and R at a position corresponding to position 156 of SEQ ID NO:5; e) R at a position corresponding to position 156 of SEQ ID NO:5 and W at a position corresponding to position 162 of SEQ ID NO:5; f) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5; g) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at a position corresponding to position 35 of SEQ ID NO:5, Q at a position corresponding to position 46 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; i) a W at a position corresponding to position 155 of SEQ ID NO:5, an R at a position corresponding to position 156 of SEQ ID NO:5, and a W at a position corresponding to position 162 of SEQ ID NO:5; j) N at a position corresponding to position 35 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; k) N at a position corresponding to position 35 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; l) M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; m) Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; n) M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; o) N at a position corresponding to position 35, Q at a position corresponding to position 46 of SEQ ID NO:5, M at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, and W at a position corresponding to position 162 of SEQ ID NO:5; p) S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; q) W at a position corresponding to position 75 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; r) S at a position corresponding to position 81 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; s) S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; t) W at a position corresponding to position 75 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; u) W at a position corresponding to position 155 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; v) A at a position corresponding to position 145 of SEQ ID NO:5 and W at a position corresponding to position 155 of SEQ ID NO:5; w) W at a position corresponding to position 75 of SEQ ID NO: 5 and D at a position corresponding to position 160 of SEQ ID NO: 5; x) W at a position corresponding to position 75 of SEQ ID NO:5 and A at a position corresponding to position 145 of SEQ ID NO:5; y) A at a position corresponding to position 145 of SEQ ID NO:5 and D at a position corresponding to position 160 of SEQ ID NO:5; z) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; aa) S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; bb) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; cc) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; dd) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and A at a position corresponding to position 145 of SEQ ID NO:5; ee) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ff) W at a position corresponding to position 75 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; gg) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a W at a position corresponding to position 155 of SEQ ID NO:5; hh) A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ii) a W at a position corresponding to position 75 of SEQ ID NO:5, an A at a position corresponding to position 145 of SEQ ID NO:5, and a D at a position corresponding to position 160 of SEQ ID NO:5; jj) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; kk) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 155 of SEQ ID NO:5; ll) S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; mm) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; nn) W at a position corresponding to position 75 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; oo) W at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; pp) L at a position corresponding to position 40 of SEQ ID NO: 5, and M at a position corresponding to position 105 of SEQ ID NO: 5; qq) L at a position corresponding to position 40 of SEQ ID NO:5, and V at a position corresponding to position 121 of SEQ ID NO:5; rr) L at a position corresponding to position 40 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; ss) L at a position corresponding to position 40 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; tt) L at a position corresponding to position 40 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; uu) M at a position corresponding to position 105 of SEQ ID NO:5 and V at a position corresponding to position 121 of SEQ ID NO:5; vv) M at a position corresponding to position 105 of SEQ ID NO: 5, and M at a position corresponding to position 145 of SEQ ID NO: 5; ww) M at a position corresponding to position 105 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; xx) M at a position corresponding to position 105 of SEQ ID NO: 5 and W at a position corresponding to position 160 of SEQ ID NO: 5; yy) V at a position corresponding to position 121 of SEQ ID NO:5, and M at a position corresponding to position 145 of SEQ ID NO:5; zz) V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; aaa) V at a position corresponding to position 121 of SEQ ID NO:5 and W at a position corresponding to position 160 of SEQ ID NO:5; bbb) M at a position corresponding to position 145 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ccc) M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; ddd) Q at a position corresponding to position 159 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; eee) S at a position corresponding to position 81 of SEQ ID NO:5, M at a position corresponding to position 145 of SEQ ID NO:5, and W at a position corresponding to position 160 of SEQ ID NO:5; fff) S at a position corresponding to position 81 of SEQ ID NO:5, V at a position corresponding to position 121 of SEQ ID NO:5, and Q at a position corresponding to position 159 of SEQ ID NO:5; ggg) A at a position corresponding to position 68 of SEQ ID NO:5 and S at a position corresponding to position 81 of SEQ ID NO:5; hhh) G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; iii) S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; jjj) A at a position corresponding to position 68 of SEQ ID NO:5 and G at a position corresponding to position 76 of SEQ ID NO:5; kkk) A at a position corresponding to position 68 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; lll) G at a position corresponding to position 76 of SEQ ID NO:5 and L at a position corresponding to position 117 of SEQ ID NO:5; mmm) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and S at a position corresponding to position 81 of SEQ ID NO:5; nnn) A at a position corresponding to position 68 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ooo) G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; ppp) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; qqq) A at a position corresponding to position 68 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, and L at a position corresponding to position 117 of SEQ ID NO:5; rrr) S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 109 of SEQ ID NO:5, E at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; sss) A at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, A at a position corresponding to position 121 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; ttt) Q at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, I at a position corresponding to position 117 of SEQ ID NO:5, Q at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; uuu) Y at a position corresponding to position 75 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, G at a position corresponding to position 121 of SEQ ID NO:5, L at a position corresponding to position 145 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, and D at a position corresponding to position 160 of SEQ ID NO:5; vvv) C at a position corresponding to position 22 of SEQ ID NO:5, A at a position corresponding to position 68 of SEQ ID NO:5, Y at a position corresponding to position 75 of SEQ ID NO:5, G at a position corresponding to position 76 of SEQ ID NO:5, S at a position corresponding to position 81 of SEQ ID NO:5, H at a position corresponding to position 108 of SEQ ID NO:5, L at a position corresponding to position 117 of SEQ ID NO:5, H at a position corresponding to position 122 of SEQ ID NO:5, Y at a position corresponding to position 138 of SEQ ID NO:5, L at a position corresponding to position 139 of SEQ ID NO:5, A at a position corresponding to position 142 of SEQ ID NO:5, A at a position corresponding to position 145 of SEQ ID NO:5, R at a position corresponding to position 153 of SEQ ID NO:5, W at a position corresponding to position 155 of SEQ ID NO:5, R at a position corresponding to position 156 of SEQ ID NO:5, D at a position corresponding to position 160 of SEQ ID NO:5, and H at a position corresponding to position 162 of SEQ ID NO:5; or www) S at a position corresponding to position 81 of SEQ ID NO:

5.

219. 219. The method of claim 217 or 218, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

220. 219. The method of claim 217 or 218, wherein the deaminase has an amino acid sequence of any one of SEQ ID NOs: 300-414, 596-598, and 720-723.

221. 221. The method of any one of claims 217 to 220, wherein the deaminase has improved deaminase activity compared to SEQ ID NO:

5.

222. 222. The method of any one of claims 217 to 221, wherein the deaminase is adenine deaminase.

223. 223. The method of any one of claims 217 to 222, wherein the RGN of the fusion protein has an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 1-4, 49-162, 435, 575, 576, 698, and 699.

224. The method of any one of claims 217 to 222, wherein the RGN of the fusion protein is an RGN nickase.

225. 225. The method of claim 224, wherein the RGN nickase has an amino acid sequence set forth as any one of SEQ ID NOs: 49-52 and 698.

226. the fusion protein a) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 319 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; b) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 375 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; c) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 406 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; d) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 596 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; e) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 597 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; f) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 303 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; g) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 366 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 49; and h) a fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 720, fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO:

49.

227. the fusion protein a) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 319 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as SEQ ID NO: 49; b) a fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence set forth as SEQ ID NO: 375 and is fused to the N-terminus of an RGN comprising the amino acid sequence set forth as S...