Evolved adenine deaminase and RNA-guided nuclease fusion proteins with internal insertion sites and methods of use

By inserting deaminase peptides into RNA-guided nucleases to form fusion proteins, the problems of low editing efficiency and limited editing window of existing RNA-guided nucleases have been solved, achieving more efficient genome editing that is applicable to gene function research, treatment of hereditary diseases, and agricultural biotechnology.

CN121127583APending Publication Date: 2025-12-12LIFEEDIT THERAPEUTICS INC

Patent Information

Application Number
CN202380090553.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-17
Filing Date
2023-11-06
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing RNA-guided nuclease systems suffer from low editing efficiency and limited editing windows in genome editing, making it difficult to efficiently modify specific genomic locations.

Method used

By designing a fusion protein containing RNA-guided nuclease and deaminase peptides, and inserting heterologous peptides into the RNA-guided nuclease, a base editor is formed, improving editing activity and editing window.

Benefits of technology

It enables more efficient genome editing, expands the editing window, and improves the ability to modify specific genomic locations, making it suitable for gene function research, treatment of hereditary diseases, and agricultural biotechnology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005482960170000121
    Figure BDA0005482960170000121
  • Figure BDA0005482960170000131
    Figure BDA0005482960170000131
  • Figure BDA0005482960170000141
    Figure BDA0005482960170000141
Patent Text Reader

Abstract

Compositions and methods comprising a deaminase for targeted editing of nucleic acids are provided. Also provided are compositions and methods for localizing a heterologous polypeptide to a target DNA molecule, and compositions and methods for targeted editing of nucleic acids. Fusion proteins comprising an RNA-guided nuclease (RGN) and at least one heterologous polypeptide inserted therein are provided, as well as fusion proteins comprising a DNA binding polypeptide and a deaminase. The heterologous polypeptide may be a pilot editing polypeptide or a base editing polypeptide. Compositions also include nucleic acid molecules encoding a deaminase or fusion protein. Vectors and host cells comprising the nucleic acid molecules encoding the deaminase or fusion protein are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 382,344, filed November 4, 2022, and U.S. Provisional Patent Application No. 63 / 485,642, filed February 17, 2023, each of which is incorporated herein by reference in its entirety.

[0003] The sequence list is submitted electronically in the form of an XML file.

[0004] This application contains a sequence list that has been filed in XML format with the USPTO Patent Center and is incorporated herein by reference in its entirety. The XML copy created on November 6, 2023, is named L103438_1330WO_Seq_List.xml and has a size of 1.12 MB. Technical Field

[0005] This invention relates to the fields of molecular biology and gene editing. Background of the Invention

[0007] Targeted genome editing or modification is rapidly becoming an important tool in both basic and applied research. Early approaches involved engineered nucleases, such as meganucleases, zinc finger fusion proteins, or TALENs, requiring the creation of chimeric nucleases with engineered, programmable, sequence-specific DNA-binding domains that are specific to each particular target sequence. RNA-guided nucleases (RGNs), such as CRISPR-Cas bacterial system-associated clustered regularly spaced short palindromic repeats (CRISPR) proteins, allow targeting specific sequences by complexing the nuclease with guide RNA that specifically hybridizes to the target sequence. Generating target-specific guide RNA is less costly and more efficient than generating chimeric nucleases for each target sequence. Such RNA-guided nucleases can be used to edit the genome by introducing sequence-specific double-strand breaks to introduce mutations at specific genomic locations, which are then repaired via error-prone non-homologous end joining (NHEJ).

[0008] Furthermore, RGNs can be used in targeted DNA editing protocols. Targeted editing of nucleic acid sequences that allows for the introduction of specific modifications into genomic DNA (e.g., targeted cleavage) enables highly detailed protocols for studying gene function and expression. RGNs can also be used to generate chimeric proteins that combine the RNA-guided activity of RGNs with DNA-modifying enzymes (such as deaminases) for targeted base editing. Targeted editing can be deployed to target genetic diseases in humans or to introduce agronomically beneficial mutations into the genomes of crop plants. The development of genome editing tools has provided new avenues for gene-editing-based mammalian therapies and agricultural biotechnology. Summary of the Invention

[0009] Compositions and methods for directing heteropolynucleotides to and modifying target DNA molecules are provided. The provided compositions comprise a deaminase polypeptide and a fusion protein comprising an RNA-guided nuclease (RGN) and at least one heteropolypeptide, such as a deaminase, inserted therein. The heteropolypeptide may be a lead editing polypeptide or a base editing polypeptide. Fusion proteins comprising base editing polypeptides are also referred to herein as base editors. Thus, in some embodiments, a base editor comprises a nucleic acid molecule-binding polypeptide (e.g., a DNA-binding polypeptide) and a deaminase polypeptide. In some embodiments, a base editor comprises an RGN having an inserted deaminase. In some embodiments, a base editor comprising an RGN having an inserted deaminase polypeptide exhibits improved editing activity and / or a shifted editing window compared to a parental base editor. In some embodiments, a parental base editor comprises a deaminase fused to the amino terminus of an RGN. Nucleic acid molecules encoding the disclosed deaminase polypeptide and fusion protein, vectors comprising nucleic acid molecules, and cells comprising the deaminase polypeptide or fusion protein, the nucleic acid molecule, or the vector are also provided.

[0010] A system comprising a disclosed fusion protein (or a nucleic acid molecule encoding the fusion protein) and one or more guide RNAs (or one or more nucleic acid molecules encoding the one or more guide RNAs) is provided, as well as a cell comprising the system, the system being used to target a heterologous peptide or deaminase to a target DNA molecule. Methods for targeting a heterologous peptide or deaminase and / or modifying a target DNA molecule include delivering the disclosed system to the target DNA molecule or a cell containing the target DNA molecule.

[0011] Pharmaceutical compositions comprising the deaminase and fusion protein of this disclosure, a nucleic acid molecule encoding the deaminase and fusion protein (or a vector or cell comprising the deaminase and fusion protein), or a system (or a cell comprising the system) are provided. Methods of treating a subject suffering from a disease, condition, or illness, or at risk of developing such a disease, condition, or illness, include administering to the subject the fusion protein of this disclosure, a nucleic acid molecule encoding the fusion protein (or a vector or cell comprising the fusion protein), a system (or a cell comprising the system), or a pharmaceutical composition. In some embodiments, the disease, condition, or illness is associated with a causal mutation, and the treatment includes correcting the causal mutation. Attached Figure Description

[0012] Figure 1 A reasonable design scheme is described to improve the editing efficiency of bases outside the standard editing window of the parental adenine base editor LPG50148-nAPG07433.1. Figure 1 A provides a schematic diagram of a linear base editor, which shows a parental adenine base editor (ABE) as an end-to-end fusion between LPG50148 and nAPG07433.1, and an internal base editor in which LPG50148 is inserted within nAPG07433.1. Figure 1 B shows the structural model of APG07433.1 with insertion points (blue spheres). LPG50148 (green) is inserted after the position shown. Insertion sites G910, V872, and E900 are within the wedge-shaped structural domain; insertion sites T642 and R670 are within the HNH structural domain; and insertion sites D737, A802, S775, S772, K778, and R30 are within the RuvCIII structural domain.

[0013] Figure 2 The editing efficiency and editing window of the nAPG07433.1 Inlaid Base Editor (IBE) LPG20002-LPG20006 and HNH substitute LPG20001 are shown. Figure 2 A provides the total editing efficiency of the first five nAPG07433.1 IBE designs, with the total percentage of A to G substitutions in these IBEs tested against genomic targets. Figure 2 B shows the editing windows for LPG20001-LPG20006. The heatmap provides the percentage of editing at each position. The rows represent five individual guides (from top to bottom: SGN001681, SGN001062, SGN000968, SGN000754, and SGN001064). Positions 1-25 indicate the positions of the first to 25th target adenines within the anterior septal sequence, where position 1 is the first adenine 5' of the PAM.

[0014] Figure 3 The image shows the editing efficiency and editing window of nAPG07433.1 IBE LPG20007-LPG20011. The heatmap provides the editing percentage for each location. The rows represent four individual guides (from top to bottom: SGN001062, SGN000968, SGN000754, and SGN001064).

[0015] Figure 4 nAPG07433.1 IBE 2 and 8 are described. Figure 4 A provides structural representations of the insertion points for nAPG07433.1 IBE 2(S772) and 8(T642). Figure 4 B depicts the schematic domain organization of LPG20002 and LPG20008 compared to their parent ABE. Figure 4 C shows the average editing efficiency of LPG20002 and LPG20008 on three target sites (SGN001947, SGN001064 and SGN000909) compared with the parental ABE.

[0016] Figure 5 An evaluation of nAPG07433.1 IBE2 with a missing connector is shown. Figure 5 A describes inserting a truncated LPG50148 into nAPG07433.1, with the N-end connector, C-end connector, or both connectors missing, to produce the nAPG07433.1 IBE2 variant. Figure 5 B shows the total editing efficiency (percentage of A to G substitutions) of all LPG20002 variants tested with four genomic targets. Figure 5 C provides a heatmap showing the percentage of edits for each location. The rows represent four individual target sites (from top to bottom: SGN001062, SGN001064, SGN000754, and SGN000968).

[0017] Figure 6 This demonstrates the therapeutic potential of engineered IBEs. Figure 6 A provides the results of disrupting the splicing donor in exon 1 of the transthyretin (TTR) gene as a gene knockout strategy. The plasmid was delivered to HEK cells via nuclear transfection. Figure 6 B illustrates the use of the LPG20002 variant to reduce the surface presentation of β2 microglobulin (B2M). The mRNA was delivered to T cells via nuclear transfection.

[0018] Figure 7 shows the editing activity of nAPG05586 IBE. Figure 7AThe average editing rate is provided on the editing windows (adenine 5-24) of nAPG05586 IBE with four guide RNAs after delivery of plasmids encoding IBE and guide RNA into HEK293T cells via nuclear transfection. Figure 7B The edit rate at each adenine position within the edit window of each of the nAPG05586 IBE is provided.

[0019] Figure 8 shows the editing activity of nAPG01604 IBE. Figure 8A The average editing rate is provided on the editing windows (adenine 5-19) of nAPG01604 IBE with four guide RNAs after delivery of plasmids encoding IBE and guide RNA into HEK293T cells via nuclear transfection. Figure 8B The edit rate at each adenine position within the edit window of each of the nAPG01604 IBE is provided.

[0020] Figure 9 shows the editing activity of nLPG10145 IBE. Figure 9A The average editing rate is provided on the editing windows (adenine 4–25) of LPG10145 IBE with four guide RNAs after delivery of plasmids encoding IBE and guide RNA into HEK293T cells via nuclear transfection. Figure 9B The edit rate at each adenine position within the edit window of each of the nLPG10145 IBEs is provided.

[0021] Figure 10A and Figure 10B The first step of a directed evolution strategy was described, which was used to identify an adenine base editor variant (LPG50148 fused with dAPG07433.1) that was more active than the parental ABE. Figure 10A An end-to-end fusion protein of an LPG50148 adenosine deaminase variant and an inactive version of the APG07433.1 RNA-guided nuclease is shown, which is expressed in E. coli and converts the stop codon in the kanamycin resistance gene to a sense codon to confer resistance. Figure 10B This study demonstrates how evolved ABEs with greater editing activity than their parental counterparts can be enriched. *E. coli* carrying the target plasmid were transformed with a library of the mutant LPG50148, fused with dAPG07433.1. To grow in high concentrations of kanamycin, cells were instructed to repair antibiotic resistance genes (KanR*) with higher editing efficiency than the parental ABEs. Selection was made individually for three sequence contexts: TAT, CAC, or TAG. Following the enrichment step, variants with up to eight-fold improvements were identified.

[0022] Figure 11 A complete directed evolution strategy for identifying adenine base editor variants with greater activity than the parental ABE is illustrated. ABE variants with single-point mutations exhibiting up to 8-fold greater activity compared to the parental ABE were selected in bacteria and evaluated in the first round of human cell screening in human HEK293T cells. Single-point mutants identified in selections targeting TAT or CAC sequence contexts were evaluated at eight endogenous genomic loci, while single-point mutants identified in selections targeting TAG sequence contexts were evaluated at four endogenous genomic loci. The single mutations showing the highest average fold change were combined to produce the highest hit counts for TAT: 20 (LPG50148 with L35N, V81S, N156R, and L162W mutations as shown in SEQ ID NO:319), 60 (LPG50148 with I75W and F155W mutations as shown in SEQ ID NO:359), 67 (LPG50148 with V81S, F155W, and A160D mutations as shown in SEQ ID NO:366), and 68 (LPG50148 with V81S, C145A, and F155W mutations as shown in SEQ ID NO:367); the highest hit counts for CAC: 106 (LPG50148 with D76G and V81S mutations as shown in SEQ ID NO:405) and 107 (LPG50148 with I75W and F155W mutations as shown in SEQ ID NO:319). LPG50148 with the V81S and M117L mutations shown in NO:406; and the highest hit counts of TAG: 88 (LPG50148 with the I40L and V105M mutations shown in SEQ ID NO:387) and 103 (LPG50148 with the V81S, C145M and A160W mutations shown in SEQ ID NO:402).

[0023] Figure 12 The study demonstrates that when ABE is delivered in mRNA form, the evolved ABE variants LPG50148.2.20, LPG50148.2.60, LPG50148.2.67, LPG50148.2.68, and LPG50148.2.69 exhibit higher editing efficiency compared to the parental ABE.

[0024] Figure 13 Results of robustness tests for the evolved ABE variants LPG50148.2.20, LPG50148.2.60, LPG50148.2.67, LPG50148.2.68, LPG50148.2.103, LPG50148.2.106, and LPG50148.2.107 are provided. Figure 13A shows the percentage of guides with a total AT substitution greater than 20% edited (from 48 genomic sites). Figure 13 B shows the average A-to-G substitutions at all genomic loci. There are three significant effects: variants LPG50148.2.20 (p≤0.001), LPG50148.2.103 (p<0.001), and LPG50148.2.67 (p=0.01).

[0025] Figure 14 The shifts in the optimal edit window for the three evolutionary ABE variants are shown. These figures illustrate the average AT substitution of the parental ABE and the three evolutionary ABE variants LPG50148.2.20, LPG50148.2.67, and LPG50148.2.103 by showing the position of the active editor window. The numbers indicate the position within the pre-spacer sequence upstream of the PAM sequence. The optimal window represents the maximum edit of ≥30%.

[0026] Figure 15 A-to-G edits within different sequence contexts are shown for multiple evolutionary ABE variants (LPG50148.2.20, LPG50148.2.60, LPG50148.2.67, LPG50148.2.68, LPG50148.2.103, LPG50148.2.106, and LPG50148.2.107) and the parental ABE. The percentage of edits reflects the median for each context at forty-eight genomic loci.

[0027] Figure 16 The percentage of A-to-G conversion at position 11 in the target sequence of SGN008393 was obtained using LPG20047 from primary mouse hepatocytes.

[0028] Figure 17 illustrates how engineering a third-generation deaminase based on a second-generation deaminase provides a higher level of A-to-G editing for multiple targets when fused to the N-terminus of nAPG07433.1.

[0029] Figure 18 Improved performance of the third-generation base editor LPG50324 fused with nAPG07433.1 is demonstrated.

[0030] Figure 19 This demonstrates an improved editing of LPG50310 deaminase relative to LPG50265 deaminase in HEK293T cells with plasmid delivery when fused with nAPG07433.1. Detailed Implementation

[0031] Benefiting from the teachings presented in the foregoing description, those skilled in the art will conceive of numerous modifications and other embodiments of the invention shown herein. Therefore, it should be understood that the invention is not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. Consequently, although specific terminology is used herein, it is used only in a general and descriptive sense and not for limiting purposes.

[0032] I. Overview

[0033] This disclosure provides deaminases and fusion proteins comprising a nucleic acid molecule-binding polypeptide (such as a DNA-binding polypeptide) and a deaminase polypeptide. In some embodiments, the DNA-binding polypeptide is a sequence-specific DNA-binding polypeptide because it binds to the target sequence at a greater frequency than it binds to a randomized background sequence. In some embodiments, the DNA-binding polypeptide is or is derived from a meganuclease, zinc finger fusion protein, or TALEN. In some embodiments, the fusion protein comprises an RNA-guided DNA-binding polypeptide and a deaminase polypeptide. In some embodiments, the RNA-guided DNA-binding polypeptide is an RNA-guided nuclease (RGN), such as a CRISPR-Cas (e.g., Cas9) polypeptide, which binds to a guide RNA (also referred to as gRNA), which then binds to the target nucleic acid sequence via strand hybridization.

[0034] This disclosure also provides a fusion protein comprising an RGN and at least one heterologous polypeptide inserted within the RGN (and a nucleic acid molecule encoding the fusion protein). The heterologous polypeptide may be inserted into a surface location of the RGN (e.g., immediately following an amino acid residue located on the surface of the RGN). Compared to end-to-end fusion proteins, the fusion proteins of this disclosure may comprise an RGN having a base-editing polypeptide or a lead-editing polypeptide, some of which exhibit improved editing activity and / or a transposed editing window.

[0035] The terms “protein,” “peptide,” and “polypeptide” are used interchangeably herein and refer to polymers of amino acid residues linked together by peptide (amide) bonds. The term refers to a protein, peptide, or polypeptide of any size, structure, or function. Typically, a protein, peptide, or polypeptide is at least three amino acids long. A protein, peptide, or polypeptide can refer to a single protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide can be modified, for example, by adding chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, linkers for conjugation, functionalization, or other modifications. Proteins, peptides, or polypeptides can also be single molecules or multi-molecule complexes. Proteins, peptides, or polypeptides can simply be fragments of naturally occurring proteins or peptides. Proteins, peptides, or polypeptides can be naturally occurring, recombinant, or synthetic, or any combination thereof.

[0036] As used herein, the term "fusion protein" refers to a hybrid polypeptide comprising protein domains from at least two different proteins. A fusion protein may comprise more than one different domain, such as an RGN and a deaminase. The fusion proteins of the present invention comprise a deaminase and a nucleic acid molecule-binding polypeptide of the present disclosure or a heterologous protein (e.g., a deaminase) inserted into the amino acid sequence of an RGN, which in some instances may interrupt domains within the RGN protein. In some embodiments, the fusion protein forms a complex with or is associated with a nucleic acid (e.g., RNA).

[0037] The heterologous polypeptide inserted into the RGN according to the present invention can be a base-editing polypeptide (e.g., a deaminase polypeptide or its active variant or fragment) that directly chemically modifies nucleosides (e.g., deaminates nucleosides), resulting in the conversion from one nucleobase to another. Deamination of nucleosides by deaminases can lead to point mutations at the corresponding residues, which is referred to herein as “nucleic acid editing” or “base editing.” Therefore, fusion proteins comprising RNA-guided nuclease (RGN) polypeptides and deaminases can be used for targeted editing of nucleic acid sequences.

[0038] The fusion protein disclosed herein may comprise an RGN fused to a leader editing peptide. Leader editing is a general and precise genome editing method that uses a polymerase-associated nucleic acid programmable DNA-binding protein to directly write new genetic information into a designated DNA site (described, for example, in the following references: US11,447,770B1; WO2021072328; WO2021226558; WO2020156575; WO2021042047; US11193123; each of which is incorporated herein by reference in its entirety). The leader editing system uses an RGN as a nickase, and the system is programmed with a leader editing (PE) guide RNA (“PEgRNA”).

[0039] The fusion proteins disclosed herein can be used for targeted editing of DNA in vitro, for example, for the generation of genetically modified cells. These genetically modified cells can be plant cells or animal cells. Such fusion proteins can also be used to introduce targeted mutations, for example, to correct genetic defects in mammalian cells in vitro (e.g., genetic defects in cells obtained from a subject and subsequently reintroduced into the same or another subject); and can be used to introduce targeted mutations, for example, to correct genetic defects in mammalian subjects or to introduce deactivation mutations in disease-related genes. Such fusion proteins can also be used to introduce targeted mutations in plant cells, for example, to introduce beneficial or agronomically valuable traits or alleles.

[0040] Any fusion protein provided herein can be produced by any method known in the art. For example, the proteins provided herein can be produced via recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire contents of which are incorporated herein by reference.

[0041] II. Nucleic acid molecules bind to polypeptides

[0042] Some aspects of this disclosure provide fusion proteins comprising a nucleic acid molecule-binding polypeptide and a deaminase polypeptide. While the invention contemplates binding to and targeted editing of RNA molecules, in some embodiments, the nucleic acid molecule-binding polypeptide of the fusion protein is a DNA-binding polypeptide. Such fusion proteins can be used for targeted editing of DNA in vitro, ex vivo, or in vivo. These fusion proteins are active in mammalian cells and can be used for targeted editing of DNA molecules.

[0043] In some embodiments, the fusion protein of this disclosure comprises a DNA-binding polypeptide. As used herein, the term "DNA-binding polypeptide" refers to any polypeptide capable of binding to DNA. In some embodiments, the DNA-binding polypeptide portion of the fusion protein of this disclosure binds to double-stranded DNA. In some embodiments, the DNA-binding polypeptide binds to DNA in a sequence-specific manner. As used herein, the terms "sequence-specific" or "sequence-specific manner" refer to selective interaction with a specific nucleotide sequence.

[0044] When two polynucleotide sequences hybridize to each other under stringent conditions, the two sequences can be considered substantially complementary. Similarly, if a DNA-binding polypeptide binds to its sequence under stringent conditions, the DNA-binding polypeptide is considered to bind to the specific target sequence in a sequence-specific manner. "Stringent conditions" or "stringent hybridization conditions" are expected conditions under which two polynucleotide sequences (or polypeptides binding to their specific target sequences) will bind to each other to a greater degree of detectability than other sequences (e.g., at least twice the background). Stringent conditions depend on the sequence and will vary under different conditions. Typically, stringent conditions are those where the salt concentration is below 1.5 M Na ions, typically around 0.01 to 1.0 M Na ion concentration (or other salt) at pH 7.0 to 8.3, and the temperature is at least 30°C for short sequences (e.g., 10 to 50 nucleotides) and at least 60°C for long sequences (e.g., greater than 50 nucleotides). Stringent conditions can also be achieved by adding a destabilizing agent (such as formamide). Exemplary low-toughness conditions include hybridization with a buffer of 30 to 35% formamide, 1M NaCl, and 1% SDS (sodium dodecyl sulfate) at 37°C, followed by washing in 1X to 2X SSC (20X SSC = 3.0M NaCl / 0.3M trisodium citrate) at 50 to 55°C. Exemplary medium-toughness conditions include hybridization with 40 to 45% formamide, 1.0M NaCl, and 1% SDS at 37°C, followed by washing in 0.5X to 1X SSC at 55 to 60°C. Exemplary high-toughness conditions include hybridization with 50% formamide, 1M NaCl, and 1% SDS at 37°C, followed by washing in 0.1X SSC at 60 to 65°C. Optionally, the wash buffer may contain about 0.1% to about 1% SDS. The duration of hybridization is typically less than about 24 hours, typically about 4 to about 12 hours. The duration of washing will be at least sufficient to reach equilibrium.

[0045] Tm is the temperature (at defined ionic strength and pH) at which 50% of the complementary target sequence hybridizes with a perfectly matched sequence. For DNA-DNA hybrids, Tm can be estimated from the following equation: Meinkoth and Wahl (1984) Anal. Biochem. 138:267-284: Tm = 81.5 °C + 16.6 (log M) + 0.41 (% GC) - 0.61 (% formamide) - 500 / L; where M is the molar concentration of the monovalent cation, % GC is the percentage of guanosine and cytosine nucleotides in the DNA, % formamide is the percentage of formamide in the hybridization solution, and L is the length of the hybrid in the base pair. Generally, stringent conditions are chosen to be approximately 5 °C lower than the thermal melting point (Tm) of the specific sequence and its complement at defined ionic strength and pH. However, extremely stringent conditions can be achieved by hybridization and / or washing at temperatures 1, 2, 3, or 4 °C below the specific melting point (Tm); moderately stringent conditions can be achieved by hybridization and / or washing at temperatures 6, 7, 8, 9, or 10 °C below the specific melting point (Tm); and low-stringent conditions can be achieved by hybridization and / or washing at temperatures 11, 12, 13, 14, 15, or 20 °C below the specific melting point (Tm). Using this equation, the hybridization and washing compositions, and the desired Tm, those skilled in the art will understand that a variation in the stringency of the hybridization and / or washing solutions is essentially described. Extensive guidelines for nucleic acid hybridization can be found in the following literature: Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, Part I, Chapter 2 (Elsevier, New York); and Ausubel et al., edited edition, (1995) Current Protocols in Molecular Biology, Chapter 2 (Greene Publishing and Wiley-Interscience, New York). See also Sambrook et al. (1989) Molecular Cloning: A Laboratory Manual (2nd edition, Cold Spring Harbor Laboratory Press, Plainview, New York).

[0046] In some embodiments, the sequence-specific DNA-binding polypeptide is an RNA-guided DNA-binding polypeptide (RGDBP). As used herein, the terms "RNA-guided DNA-binding polypeptide" and "RGDBP" refer to polypeptides capable of binding to DNA through hybridization of a related RNA molecule with a target DNA sequence.

[0047] In some embodiments, the DNA-binding polypeptide of the fusion protein is a nuclease, such as a sequence-specific nuclease. As used herein, the term "nuclease" refers to an enzyme that catalyzes the cleavage of phosphodiester bonds between nucleotides in a nucleic acid molecule. In some embodiments, the DNA-binding polypeptide is an endonuclease capable of cleaving phosphodiester bonds between nucleotides within a nucleic acid molecule, while in some embodiments, the DNA-binding polypeptide is an exonuclease capable of cleaving nucleotides at either end (5' or 3') of a nucleic acid molecule. In some embodiments, the sequence-specific nuclease is selected from the group consisting of megnucleases, zinc finger nucleases, TAL effector DNA-binding domain nuclease fusion proteins (TALENs), and RNA-guided nucleases (RGNs) or variants thereof, wherein the nuclease activity has been reduced or inhibited.

[0048] As used herein, the term "meganuclease" or "homing endonuclease" refers to an endonuclease that binds to a recognition site within double-stranded DNA of 12 to 40 bp in length. Non-limiting examples of meganucleases are those belonging to the LAGLIDADG family that contain the conserved amino acid motif LAGLIDADG (SEQ ID NO: 700). The term "meganuclease" can refer to a dimer or a single-stranded meganuclease.

[0049] As used herein, the term "zinc finger nuclease" or "ZFN" refers to a chimeric protein containing a zinc finger DNA-binding domain and a nuclease domain.

[0050] As used herein, the term "TAL effector DNA-binding domain nuclease fusion protein" or "TALEN" refers to a chimeric protein containing a TAL effector DNA-binding domain and a nuclease domain.

[0051] In some embodiments, the DNA-binding polypeptide is a DNA-binding polypeptide capable of generating a single-stranded region within a double-stranded DNA molecule. An example of a single-stranded region is a single-stranded loop contained within an R-loop, which is a triple-stranded nucleic acid structure containing a single-stranded DNA region formed within a double-stranded DNA molecule produced by hybridization of the complementary strand with a single-stranded RNA or DNA molecule. Adenine in or near the single-stranded region of the R-loop can be deaminated by an adenine deaminase active against single-stranded nucleic acids (e.g., ssDNA). In some of these embodiments, the DNA-binding polypeptide capable of generating an R-loop within a double-stranded DNA molecule is an RNA-guided DNA-binding polypeptide or an RGN nuclease. As used herein, the terms "RNA-guided nuclease" or "RGN" refer to an RNA-guided DNA-binding polypeptide with nuclease activity. RGN is considered "RNA-guided" because the guiding RNA forms a complex with the RNA-guided nuclease to direct the RNA-guided nuclease to bind to a target sequence and, in some embodiments, introduce single-stranded or double-stranded breaks at the target sequence.

[0052] Although RGNs can cleave target sequences upon binding, the term RGN also encompasses nuclease-dead RGNs that can bind to the target sequence but do not cleave it. The term "cleave" or "cleavage" refers to the hydrolysis of at least one phosphodiester bond within the backbone of one or both strands of a double-stranded target sequence (e.g., a target DNA sequence), which can result in single-strand or double-strand breaks within the target DNA sequence. RGN cleavage of the target sequence can result in single-strand or double-strand breaks. RGNs capable only of cleaving the single strand of a double-stranded target nucleic acid molecule are referred to herein as cleaving enzymes. Such RGNs possess a single functional nuclease domain. RGN cleaving enzymes can be naturally occurring cleaving enzymes or can be RGN proteins that naturally cleave both strands of a double-stranded nucleic acid molecule but have been mutated in one or more nuclease domains, resulting in reduced or eliminated nuclease activity in these mutated domains to become cleaving enzymes.

[0053] RNA-guided nucleases (RGNs) allow for targeted manipulation of single sites within the genome and are useful in the context of gene targeting for both therapeutic and research applications. In a variety of organisms, including mammals, RNA-guided nucleases have been used for genome engineering by stimulating non-homologous end joining or homologous recombination. RGNs include CRISPR-Cas proteins, which are RNA-guided nucleases directed to target sequences by a guide RNA (gRNA) or its active variant or fragment that is part of a clustered regularly spaced short palindromic repeat (CRISPR) RNA-guided nuclease system.

[0054] Some aspects of this disclosure provide fusion proteins comprising an RNA-guided DNA-binding polypeptide and a deaminase polypeptide (such as an adenine deaminase polypeptide). In some embodiments, the RNA-guided DNA-binding polypeptide is an RNA-guided nuclease (RGN). In further embodiments, the RNA-guided nuclease is a naturally occurring CRISPR-Cas protein or its active variant or fragment. CRISPR-Cas systems are classified as Category 1 or Category 2 systems. Category 2 systems comprise single-effect nucleases and include types II, V, and VI. Category 1 and Category 2 systems are further subdivided into types (Type I, Type II, Type III, Type IV, Type V, Type VI), some of which are further subdivided into subtypes (e.g., Type II-A, Type II-B, Type II-C, Type VA, Type VB).

[0055] In some embodiments, the RGN is a naturally occurring type II CRISPR-Cas protein or its active variant or fragment. As used herein, the terms “type II CRISPR-Cas protein,” “type II CRISPR-Cas effector protein,” or “type II RNA-guided nuclease” refer to an RGN that requires trans-activated RNA (tracrRNA) and contains two nuclease domains (i.e., RuvC and HNH), each of which is responsible for cleaving single strands of a double-stranded DNA molecule. In some embodiments, the present invention provides a fusion protein comprising a deaminase of the present disclosure fused to a Cas9 protein, such as Streptococcus pyogenes Cas9 (SpCas9), the sequence of which is shown in SEQ ID NO:415, or the SpCas9 nickase, described in U.S. Patent Nos. 10,000,772 and 8,697,359, each of which is incorporated herein by reference in its entirety. In some embodiments, the present invention provides a fusion protein comprising a deaminase of the present disclosure fused to Streptococcus thermophilus Cas9 (StCas9), the sequence of which is shown in SEQ ID NO:701, or a StCas9 nickase described in U.S. Patent No. 10,113,167, which is incorporated herein by reference in its entirety. In some embodiments, the present invention provides a fusion protein comprising a deaminase of the present disclosure fused to Streptococcus aureus Cas9 (SaCas9), as shown in SEQ ID NO:702, or a SaCas9 nickase described in U.S. Patent No. 9,752,132, which is incorporated herein by reference in its entirety.

[0056] In some embodiments, the CRISPR-Cas protein is a naturally occurring type V CRISPR-Cas protein or its active variant or fragment. As used herein, the terms “type V CRISPR-Cas protein,” “type V CRISPR-Cas effector protein,” or “type V RNA-guided nuclease” refer to an RGN that cleaves dsDNA and contains a single RuvC nuclease domain or a split RuvC nuclease domain lacking an HNH domain (Zetsche et al. 2015, Cell doi:10.1016 / j.cell.2015.09.038; Shmakov et al. 2017, Nat Rev Microbioldoi:10.1038 / nrmicro.2016.184; Yan et al. 2018, Science doi:10.1126 / science.aav7271; Harrington et al. 2018, Science doi:10.1126 / science.aav4294). In some embodiments, the fusion proteins disclosed herein comprise Cas12 (e.g., Cas12a). Notably, Cas12a is also referred to as Cpf1 and does not require tracrRNA, although other type V CRISPR-Cas proteins (such as Cas12b) do require tracrRNA. Most type V effectors can also target ssDNA (single-stranded DNA) and generally do not require PAM (Zetsche et al., 2015; Yan et al., 2018; Harrington et al., 2018). The terms “type V CRISPR-Cas protein” and “type V RGN” encompass unique RGNs containing a split RuvC nuclease domain, such as those disclosed in WO 2021 / 138247, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the present invention provides a fusion protein comprising the deaminase disclosed herein, the deaminase being fused to *Francisella novicida* Cas12a (FnCas12a), the sequence of which is shown in SEQ ID NO:703 and disclosed in U.S. Patent No. 9,790,490, which is incorporated herein by reference in its entirety, or any of the nuclease-inactivating mutants of FnCas12a disclosed in U.S. Patent No. 9,790,490.

[0057] In some embodiments, the CRISPR-Cas protein is a naturally occurring type VI CRISPR-Cas protein or its active variant or fragment. As used herein, the terms "type VI CRISPR-Cas protein," "type VI CRISPR-Cas effector protein," or "type VI RGN" refer to a CRISPR-Cas effector protein that does not require tracrRNA and contains a HEPN domain of two cleavage RNAs. In some embodiments, the present invention provides a fusion protein comprising a deaminase of the present disclosure fused to Cas13.

[0058] In some embodiments, the fusion protein of this disclosure comprises RGN, or its nickase or nuclease-dead variants are those disclosed in the following documents: International Application Publication Nos. WO 2019 / 236566, WO 2020 / 139783, WO 2021 / 030344, WO 2021 / 138247, WO 2021 / 231437 or WO 2021 / 217002, or International Application No. PCT / IB2023 / 058160, filed on August 12, 2023, each of which is incorporated herein by reference in its entirety.

[0059] In some embodiments, the fusion proteins of this disclosure comprise RGNs listed in Table 1 and / or variants thereof, such as SEQ ID NO:1-4, 49-162, 435, 575, 576, 698, and 699, representing nickase or nuclease death. Guide RNA sequences (crRNA repeats and tracrRNA sequences) that can be used with each RGN in Table 1, as well as a common PAM sequence, are also provided. In some embodiments, the fusion protein comprises an active variant of the RGN listed in Table 1 and / or such as SEQ ID NO:1-4, 49-162, 435, 575, 576, 698, and 699 (an active variant capable of binding to a nucleic acid molecule in an RNA-guided manner), which has a sequence identity between 80% and 99% or higher with any of the amino acid sequences listed in Table 1 and / or such as SEQ ID NO:1-4, 49-162, 435, 575, 576, 698, and 699, including but not limited to about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or higher. In some embodiments, the fusion protein comprises an RGN having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or higher sequence identity with the RGN amino acid sequences disclosed in Table 1 and / or such as SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698, and 699. In other embodiments, the fusion protein comprises fragments of the RGN listed in Table 1, such as fragments differing by as few as 1-15 amino acid residues, as few as 1-10 (e.g., 6-10), as few as 5, as few as 4, as few as 3, as few as 2, or as few as 1 amino acid residue. In certain embodiments, the RGN comprises an N-terminal or C-terminal truncation, which may include at least the deletion of 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60 or more amino acids from the N or C terminus of the polypeptide. In some embodiments, the RGN comprises an internal deletion, which may include at least the deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60 or more amino acids.

[0060] Table 1. Non-restrictive examples of RNA-guided nucleases

[0061]

[0062]

[0063]

[0064]

[0065] *ND: The RGN component of the fusion protein of this disclosure that is not specified may be APG07433.1 (disclosed in International Patent Publication No. WO2019 / 236566, which is incorporated herein by reference in its entirety; and is shown herein as SEQ ID NO:1), APG05586 (disclosed in International Patent Publication No. WO 2021 / 217002, which is incorporated herein by reference in its entirety; and is shown herein as SEQ ID NO:2), APG01604 (disclosed in International Patent Publication No. WO2021 / 217002, which is incorporated herein by reference in its entirety; and is shown herein as SEQ ID NO:3), LPG10145 (disclosed in International Patent Publication No. WO 2023 / 139557 ...1), LPG05586 (disclosed in International Patent Publication No. WO 2021 / 217002, which is incorporated herein by reference in its entirety; and is shown herein as SEQ ID NO:3), LPG01604 (disclosed in International Patent Publication No. WO2021 / 217002, which is incorporated herein by reference in its entirety; and is shown herein as SEQ ID NO:3), LPG10145 (disclosed in International Patent Publication No. WO 2023 / 139557, which is incorporated herein by reference in its entirety; and is shown herein as SEQ ID NO:1), LPG05586 (disclosed in International Patent Publication No. WO 2021 / 217002, which is incorporated herein by reference in its entirety (as shown in NO:4), LPG10196 (disclosed in International Application No. PCT / IB2023 / 058160, filed on August 12, 2023, which is incorporated herein by reference in its entirety; and is shown herein as SEQ ID NO:131); or its active fragments or variants (i.e., active fragments or variants capable of binding to nucleic acid molecules in an RNA-guided manner), such as its nickase variants.

[0066] The RGN may have sequence identity between 80% and 99% or higher with any of the RGNs in SEQ ID NO: 1, 2, 3, 4 and 131 (which retain RGN activity), including but not limited to at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or higher. In some embodiments, the fusion protein comprises an RGN having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher sequence identity with any of SEQ ID NO:1, 2, 3, 4 and 131 (retaining RGN activity). The fusion protein may comprise an active fragment of the RGN as shown in any of SEQ ID NO:1, 2, 3, 4 and 131, such as an active fragment differing from any of SEQ ID NO:1, 2, 3, 4 or 131 by as few as 1-15 amino acid residues, as few as 1-10 (e.g., 6-10), as few as 5, as few as 4, as few as 3, as few as 2 or as few as 1 amino acid residue. In some embodiments, the active RGN fragment comprises an N-terminal or C-terminal truncation, which may contain at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60 or more amino acids deleted from the N- or C-terminus of the polypeptides shown in any one of SEQ ID NO: 1, 2, 3, 4, and 131. In some embodiments, the active RGN variant or fragment comprises an internal deletion, which may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60 or more amino acids deleted from the N- or C-terminus of the polypeptides shown in any one of SEQ ID NO: 1, 2, 3, 4, and 131. The active fragment of RGN may contain at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more consecutive amino acid residues of the amino acid sequence shown in any of SEQ ID NO: 1, 2, 3, 4 and 131.

[0067] In some embodiments, the RGN of the fusion protein is an RGN nickase containing a mutation (e.g., a D10A mutation, where the amino acid numbering is based on the *Streptococcus pyogenes* Cas9 sequence as shown in SEQ ID NO:415) that enables the RGN to cleave only the non-base-editing target strand of the nucleic acid duplex (the strand containing PAM and paired with gRNA bases). Nickases containing the D10A mutation or an equivalent mutation have an inactivated RuvC nuclease domain and cleave the target strand. The D10A nickase cannot cleave the non-target strand of DNA, i.e., the strand intended for base editing. In these embodiments, the RGN introduces a nick into the target strand, while the complementary non-target strand is modified with a deaminase. Cellular DNA repair mechanisms can use the modified non-target strand as a template to repair the nicked target strand, thereby introducing a mutation into the DNA.

[0068] Therefore, in some embodiments, the nickase comprises an inactive RuvC domain. The RuvC domain has an RNase H fold structure (see, for example, Nishimasu et al. (2014) Cell 156(5):935-949, which is incorporated herein by reference in its entirety). The RuvC domain of RGN is typically a split RuvC domain containing two or more non-neighboring regions within a linear amino acid sequence. For example, the RuvC domain of Streptococcus pyogenes Cas9 contains amino acid residues 1-59, 718-769, and 909-1098 of SEQ ID NO:415. A non-limiting example of a mutation within the RuvC domain that inactivates its nuclease activity is the D10A mutation, which mutates the 1-episode ester residue in the split RuvC nuclease domain. nAPG07433.1 (as shown in SEQ ID NO:49) is a nickase variant of APG07433.1 (with an inactivated RuvC domain) as shown in SEQ ID NO:1, and is described in WO 2019 / 236566 (which is incorporated herein by reference in its entirety). nAPG05586 (as shown in SEQ ID NO:50) is a nickase variant of APG05586 (as shown in SEQ ID NO:2) (with an inactivated RuvC domain). nAPG01604 (as shown in SEQ ID NO:51) is a nickase variant of APG01604 (as shown in SEQ ID NO:3) (with an inactivated RuvC domain). nLPG10145 (as shown in SEQ ID NO:52) is a nickase variant of LPG10145 (as shown in SEQ ID NO:4) (with an inactivated RuvC domain). nLPG10196 (as shown in SEQ ID NO:698) is a nickase variant of LPG10196 (as shown in SEQ ID NO:131) (with an inactivated RuvC domain).

[0069] In some embodiments, the RGN of the fusion protein is an RGN nickase containing a mutation (e.g., the H840A mutation, where the amino acid numbering is based on the *Streptococcus pyogenes* Cas9 sequence as shown in SEQ ID NO:415) that enables the RGN to cleave only the non-target strand of the nucleic acid duplex (the strand that does not contain PAM and is not paired with gRNA bases). In some of these embodiments, the nickase contains an inactive HNH nuclease domain. The HNH nuclease domain of the RGN has a ββα-metallic fold (see, for example, Nishimasu et al. 2014). The HNH nuclease domain of *Streptococcus pyogenes* Cas9 contains, for example, amino acid residues 775-908 of SEQ ID NO:415. A non-limiting example of a mutation that inactivates its nuclease activity within the HNH domain is the H840A mutation, which mutates the first histidine of the HNH nuclease domain. The RGN with the inactivated HNH domain acts on the non-target strand. nAPG07433.1 (as shown in SEQ ID NO:56) is a nicking enzyme variant of APG07433.1 (with an inactivated HNH domain) as shown in SEQ ID NO:1, and is described in WO 2019 / 236566 (which is incorporated herein by reference in its entirety). nAPG05586 (as shown in SEQ ID NO:53) is a nicking enzyme variant of APG05586 (as shown in SEQ ID NO:2) (with an inactivated HNH domain). nAPG01604 (as shown in SEQ ID NO:54) is a nicking enzyme variant of APG01604 (as shown in SEQ ID NO:3) (with an inactivated HNH domain). nLPG10145 (as shown in SEQ ID NO:55) is a nicking enzyme variant of LPG10145 (as shown in SEQ ID NO:4) (with an inactivated HNH domain). nLPG10196 (as shown in SEQ ID NO:699) is a nickase variant of LPG10196 (as shown in SEQ ID NO:131) (with an inactivated HNH domain).

[0070] Methods for inactivating the RuvC and / or HNH domains of RGN are known in the art and typically involve mutating the first aspartic acid residue and / or the first histidine residue of the HNH domain within the split RuvC domain. Typically, the aspartic acid or histidine residue is mutated to alanine. Other amino acid residues within the RuvC domain that can be mutated to give the domain inactive nuclease activity include Glu762, His983, and Asp986 (typically alanine), wherein the amino acid numbering is based on the *Streptococcus pyogenes* Cas9 sequence as shown in SEQ ID NO:415. Other amino acid residues within the HNH domain that can be mutated include D839 and N863 (typically alanine), wherein the amino acid numbering is based on the *Streptococcus pyogenes* Cas9 sequence as shown in SEQ ID NO:415.

[0071] Unless otherwise stated, when the nomenclature “n” is followed by the nuclease name, it refers to a nickase in which the RuvC domain has been inactivated (e.g., with a D10A mutation).

[0072] In some embodiments, the fusion protein comprises an RGN nickase that retains nickase activity, comprising an amino acid sequence having about 60% to about 99.5% identity with any one of SEQ ID NO:49-56, 698, and 699. In some embodiments, the RGN nickase that retains RGN nickase activity comprises an amino acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with any one of SEQ ID NO:49-56, 698, and 699.

[0073] In some embodiments, the fusion protein comprises an RGN nickase comprising an amino acid sequence having sequence identity between 80% and 99% or higher with any one of SEQ ID NO:49-56, 698 and 699, including but not limited to at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or higher.

[0074] In some embodiments, the RGN of the fusion protein is nuclease-dead. As used herein, an RGN protein mutated to become nuclease-inactive or “dead” can be referred to as an RNA-guided DNA-binding polypeptide, a nuclease-inactive RGN, or a nuclease-dead RGN. Methods for generating nuclease-inactive RGNs are known in the art and generally involve mutating a unique nuclease domain or all nuclease domains within the nuclease domain of the RGN to render the nuclease domain inactive. In those embodiments where the RGN contains only a single nuclease domain (e.g., a RuvC domain), the nuclease-inactive variant will have at least one mutation within the RuvC domain that results in the inactivation of the RuvC nuclease domain. In those embodiments where the RGN contains more than one nuclease domain (such as RuvC and HNH domains), at least one mutation within each of the RuvC or HNH domains renders both nuclease domains inactive.

[0075] An exemplary suitable nuclease-inactive RGN is the D10A / H840A-Cas9 mutant (see, for example, Qi et al. Cell. 2013; 152(5):1173-83, the entire contents of which are incorporated herein by reference). Furthermore, suitable nuclease-inactive variants of other known RNA-guided nucleases (RGNs) can be identified.

[0076] Other additional exemplary suitable nuclease inactive RGN variants include, but are not limited to, the D10A / D839A / H840A and D10A / D839A / H840A / N863A mutant domains (see, for example, Mali et al., Nature Biotechnology. 2013; 31(9):833-838, the full contents of which are incorporated herein by reference).

[0077] Based on this disclosure and knowledge in the art, additional suitable RGN proteins mutated into nickases or inactive nucleases will be apparent to those skilled in the art (such as the example RGN disclosed, for example, in PCT Publication No. WO 2019 / 236566, which is incorporated herein by reference in its entirety), and are within the scope of this disclosure.

[0078] Any method known in the art for introducing mutations into an amino acid sequence, such as PCR-mediated mutagenesis and site-directed mutagenesis, can be used to generate RGNs that result in nickase or nuclease death. See, for example, U.S. Publication No. 2014 / 0068797 and U.S. Patent No. 9,790,490; each of these patents is incorporated herein by reference in its entirety.

[0079] The fusion protein of the present invention, comprising an RGN, utilizes a guide RNA that binds to the RGN component of the fusion protein and guides the fusion protein to a target sequence. The term "guide RNA" refers to a nucleotide sequence that is sufficiently complementary to the target nucleotide sequence to hybridize with the target sequence and guide the associated RGN to bind sequence-specifically to the target nucleotide sequence. More specifically, when the target nucleotide sequence is double-stranded, as in the case of DNA, the target nucleotide sequence consists of a target strand (which contains a PAM sequence) and a non-target strand. In these embodiments, the guide RNA is sufficiently complementary to the non-target strand of the double-stranded target sequence (e.g., the target DNA sequence) such that the guide RNA hybridizes with the non-target strand and guides the associated RNA-guided nuclease (RGN) to bind sequence-specifically to the target sequence (e.g., the target DNA sequence). Thus, in some embodiments, the guide RNA includes a spacer with the same sequence as the target strand, except that uracil (U) replaces thymidine (T) in the guide RNA.

[0080] Guide RNA is one or more RNA molecules (usually one or two) that can bind to RGN and guide RGN to bind to a specific target nucleotide sequence, and in those examples where RGN has nicking or nuclease activity, can also cleave the target nucleotide sequence. Guide RNA comprises CRISPR RNA (crRNA), and in some embodiments, comprises trans-activating CRISPR RNA (tracrRNA). In some embodiments, a portion of the guide RNA comprises DNA nucleotides. In some embodiments, the guide RNA comprises artificial, non-naturally occurring nucleotide analogs, or one or more nucleotides are chemically modified, including modifications described in International Application No. PCT / IB2023 / 058418, filed August 25, 2023, which is incorporated herein by reference in its entirety.

[0081] CRISPR RNA comprises a spacer sequence and a CRISPR repeat sequence. The "spacer sequence" is a nucleotide sequence that directly hybridizes to the non-target strand of the target sequence of interest. The spacer sequence is engineered to be fully or partially complementary to the non-target strand of the target sequence of interest. In various embodiments, the spacer sequence comprises about 8 nucleotides to about 30 nucleotides or more. For example, the length of the spacer sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30 or more nucleotides. In some embodiments, the spacer sequence is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides in length. In some embodiments, the spacer sequence is about 10 to about 26 nucleotides or about 12 to about 30 nucleotides in length. In some embodiments, the spacer sequence is about 30 nucleotides in length. In some embodiments, the spacer sequence is 30 nucleotides in length. In some embodiments, when optimal alignment is performed using a suitable alignment algorithm, the complementarity between the spacer subsequence and its corresponding target sequence is between 50% and 99% or higher, including but not limited to about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% or higher. In some embodiments, when optimal alignment is performed using a suitable alignment algorithm, the complementarity between the spacer sequence and its corresponding target sequence is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher. In some embodiments, the spacer sequence is free of secondary structures, which can be predicted using any suitable polynucleotide folding algorithm known in the art, including but not limited to mFold (see, e.g., Zuker and Stiegler (1981) Nucleic Acids Res. 9:133-148) and RNAfold (see, e.g., Gruber et al. (2008) Cell 106(1):23-24).

[0082] CRISPR RNA repeat sequences comprise nucleotide sequences recognized by RGN molecules, which independently or collaboratively form structures with hybridized tracrRNA. In various embodiments, the CRISPR RNA repeat sequence comprises about 8 nucleotides to about 30 nucleotides or more. In some embodiments, the CRISPR RNA repeat sequence comprises 8 to 30 nucleotides or more. For example, the length of the CRISPR repeat sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30 or more nucleotides. In some embodiments, the length of the CRISPR repeat sequence is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides. In some embodiments, when optimal alignment is performed using a suitable alignment algorithm, the complementarity between the CRISPR repeat sequence and its corresponding tracrRNA sequence is between 50% and 99% or higher, including but not limited to about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% or higher. In some embodiments, when optimal alignment is performed using a suitable alignment algorithm, the complementarity between the CRISPR repeat sequence and its corresponding tracrRNA sequence is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher.

[0083] In some embodiments, the guide RNA further comprises a tracrRNA molecule. The trans-activating CRISPR RNA or tracrRNA molecule comprises a nucleotide sequence containing a region of sufficient complementarity to hybridize with the CRISPR repeat sequence of the crRNA; this region is referred to herein as the anti-repetition region. In some embodiments, the tracrRNA molecule further comprises a region having secondary structures (e.g., stem-loop), or forming secondary structures upon hybridization with its corresponding crRNA. In some embodiments, the tracrRNA region, which is fully or partially complementary to the CRISPR repeat sequence, is located at the 5' end of the molecule, and the 3' end of the tracrRNA contains secondary structures. This secondary structure region typically contains several hairpin structures, including linker hairpins found adjacent to the anti-repetition sequence. A terminal hairpin is typically present at the 3' end of the tracrRNA; its structure and number may vary, but it typically contains a GC-rich Rho-independent transcription terminator hairpin, followed by a string of Us at the 3' end. See, for example, Briner et al. (2014) Molecular Cell 56:333-339, Briner and Barrangou (2016) Cold Spring Harb Protoc; doi:10.1101 / pdb.top090902, and U.S. Publication No. 2017 / 0275648, each of which is incorporated herein by reference in its entirety.

[0084] In various embodiments, the anti-repetition region of the tracrRNA, which is fully or partially complementary to the CRISPR repeat sequence, comprises about 6 nucleotides to about 30 nucleotides or more. For example, the length of the base-pairing region between the tracrRNA anti-repetition sequence and the CRISPR repeat sequence can be about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30 or more nucleotides. In some embodiments, the length of the base-pairing region between the anti-repetitive sequence of the tracrRNA and the CRISPR repeat sequence can be 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides. In some embodiments, the anti-repetitive region of the tracrRNA that is completely or partially complementary to the CRISPR repeat sequence is about 10 nucleotides long. In some embodiments, the anti-repetitive region of the tracrRNA that is completely or partially complementary to the CRISPR repeat sequence is 10 nucleotides long. In some embodiments, when optimal alignment is performed using a suitable alignment algorithm, the complementarity between the CRISPR repeat sequence and its corresponding tracrRNA anti-repetitive sequence is between 50% and 99% or higher, including but not limited to about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% or higher. In some embodiments, when optimal alignment is performed using a suitable alignment algorithm, the complementarity between the CRISPR repeat sequence and its corresponding tracrRNA anti-repetitive sequence is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or higher.

[0085] In various embodiments, the entire tracrRNA contains from about 60 nucleotides to more than about 210 nucleotides. In some embodiments, the entire tracrRNA contains from 60 nucleotides to more than 210 nucleotides. For example, the length of the tracrRNA can be about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 210 or more nucleotides. In some embodiments, the tracrRNA has a length of 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 150, 160, 170, 180, 190, 200, 210 or more nucleotides. In some embodiments, the tracrRNA has a length of about 100 to about 200 nucleotides, including lengths of about 95, about 96, about 97, about 98, about 99, about 100, about 105, about 106, about 107, about 108, about 109 and about 100 nucleotides. In some embodiments, the tracrRNA has a length of 100 to 110 nucleotides, including lengths of 95, 96, 97, 98, 99, 100, 105, 106, 107, 108, 109 and 110 nucleotides.

[0086] A guide RNA forms a complex with an RNA-guided DNA-binding polypeptide (RGN) or an RNA-guided nuclease or a fusion protein containing it to guide the RNA-guided nuclease or fusion protein containing it to bind to a target sequence. If the guide RNA complexes with an RGN, the bound RGN introduces a single-strand or double-strand break at the target sequence. After the target sequence has been cleaved, the break can be repaired, allowing modification of the target DNA sequence during the repair process. This paper provides methods for using mutant variants of RNA-guided nucleases, which are either nuclease-inactive or nickases linked to base-editing polypeptides (e.g., deaminases) or lead-editing polypeptides, to modify target sequences in the host cell's DNA. Mutant variants of RNA-guided nucleases whose nuclease activity is inactivated or significantly reduced may be referred to as RNA-guided DNA-binding polypeptides (RGDBPs) because the polypeptide can bind to the target sequence but does not necessarily cleave it. RNA-guided nucleases that can only cleave single-stranded double-stranded nucleic acid molecules are referred to herein as nickases.

[0087] The target nucleotide sequence is bound by an RNA-guided DNA-binding polypeptide (e.g., RGN) and hybridizes with a guide RNA that is associated with an RDBBP (e.g., RGN). If the RDBBP possesses nuclease activity (i.e., RGN), encompassing activity as a cleavage enzyme, the target sequence can then be cleaved.

[0088] The guide RNA can be a single guide RNA or a dual guide RNA system. A single guide RNA comprises a crRNA and optionally a tracrRNA on a single RNA molecule, while a dual guide RNA system comprises a crRNA and a tracrRNA presented on two different RNA molecules that hybridize to each other by at least a portion of the CRISPR repeat sequence of the crRNA and at least a portion of the tracrRNA, the tracrRNA being fully or partially complementary to the CRISPR repeat sequence of the crRNA. In some embodiments where the guide RNA is a single guide RNA, the crRNA and optionally the tracrRNA are separated by an adapter nucleotide sequence.

[0089] Suitable crRNA repeats, tracrRNAs, and guide RNA sequences for the RGN component of the fusion protein of this disclosure are disclosed in the following publications: International Patent Publications Nos. WO 2019 / 236566, WO 2021 / 030344, WO 2020 / 139783, WO 2021 / 217002, WO 2021 / 138247, WO 2021 / 231437, WO 2023 / 139557, and PCT International Application No. PCT / IB2023 / 058160, filed on August 12, 2023. Each of these publications is incorporated herein by reference in its entirety and / or provided herein in Table 1.

[0090] Generally, the linker nucleotide sequence between crRNA and tracrRNA is a linker nucleotide sequence that does not include complementary bases, in order to avoid the formation of secondary structures within the nucleotides of the linker nucleotide sequence, or to avoid the formation of secondary structures containing the linker nucleotide sequence. In some embodiments, the length of the linker nucleotide sequence between crRNA and tracrRNA is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12 or more nucleotides. In some embodiments, the length of the linker nucleotide sequence between crRNA and tracrRNA is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more nucleotides. In some embodiments, the length of the linker nucleotide sequence of a single guide RNA is at least 4 nucleotides. In some embodiments, the length of the linker nucleotide sequence of a single guide RNA is 4 nucleotides.

[0091] In some embodiments, guide RNA can be introduced as an RNA molecule into target cells, organelles, or embryos. Guide RNA can be transcribed in vitro or chemically synthesized. In some embodiments, a nucleotide sequence encoding the guide RNA is introduced into cells, organelles, or embryos. In some embodiments, the nucleotide sequence encoding the guide RNA is operatively linked to a promoter (e.g., the RNA polymerase III promoter). The promoter can be a natural promoter or heterologous to the nucleotide sequence encoding the guide RNA. In some embodiments, the promoter is selected from any of the promoters disclosed in International Application No. PCT / US2022 / 032940, filed June 10, 2022, which is incorporated herein by reference in its entirety.

[0092] In various embodiments, guide RNA can be introduced into target cells, organelles, or embryos as a ribonucleoprotein complex as described herein, wherein the guide RNA binds to an RNA-guided nuclease polypeptide.

[0093] Guide RNA directs an associated RDBBP (e.g., RGN) to a specific target nucleotide sequence of interest by hybridizing with the guide RNA. The target nucleotide sequence may comprise DNA, RNA, or a combination of both, and may be single-stranded or double-stranded. The target nucleotide sequence may be genomic DNA (i.e., chromosomal DNA), plasmid DNA, or RNA molecules (e.g., messenger RNA, ribosomal RNA, transfer RNA, microRNA, small interfering RNA). The target nucleotide sequence may be bound (and in some embodiments, cleaved) to an RNA-guided DNA-binding polypeptide in vitro or in cells. The chromosomal sequence targeted by the RDBBP (e.g., RGN) may be a nuclear, plasmosome, or mitochondrial chromosomal sequence. In some embodiments, the target nucleotide sequence is unique within the target genome.

[0094] In some embodiments, the target nucleotide sequence is adjacent to a prespacer adjacent motif (PAM). The PAM is typically within about 1 to about 10 nucleotides of the target nucleotide sequence, including about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 nucleotides. In some embodiments, the PAM is within 1 to 10 nucleotides of the target nucleotide sequence, including 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. Unless otherwise stated, the PAM is immediately adjacent to the target nucleotide sequence at its 5' or 3' end. In some embodiments, the PAM is the 3' end of the target sequence. Typically, the PAM is a common sequence of about 2-6 nucleotides, but in some embodiments, its length is 1, 2, 3, 4, 5, 6, 7, 8, 9, or more nucleotides.

[0095] Suitable PAM sequences for the RGN component of the fusion protein used in this disclosure are disclosed in the following documents: International Patent Publication Nos. WO 2019 / 236566, WO 2021 / 030344, WO 2020 / 139783, WO 2021 / 217002, WO 2021 / 138247, WO 2021 / 231437, WO 2023 / 139557, and PCT International Application No. PCT / IB2023 / 058160, filed on 12 August 2023. Each of these documents is incorporated herein by reference in its entirety and / or provided herein in Table 1.

[0096] PAMs restrict which sequences a given RDBBP (e.g., RGN) can target because its PAM needs to be close to the target nucleotide sequence. After recognizing its corresponding PAM sequence, the RGN can cleave the target nucleotide sequence at a specific cleavage site. As used herein, a cleavage site consists of two specific nucleotides within the target nucleotide sequence between which the RGN cleaves the nucleotide sequence. The cleavage site can contain the 1st and 2nd, 2nd and 3rd, 3rd and 4th, 4th and 5th, 5th and 6th, 7th and 8th, or 8th and 9th nucleotides of the PAM in the 5' or 3' direction. Because the RGN can cleave the target nucleotide sequence, resulting in staggered ends, in some embodiments, the cleavage site is defined based on the distance between two nucleotides of the PAM on the positive (+) strand of the polynucleotide and the distance between two nucleotides of the PAM on the negative (-) strand of the polynucleotide.

[0097] RDBBP and RGN can be used to deliver fused peptides, polynucleotides, or small molecule payloads to specific genomic locations.

[0098] In embodiments where the fusion protein includes a meganuclease as a DNA-binding component, the target sequence may comprise a pair of inverted nine-base-pair "half-sites" separated by four base pairs. In the case of a single-stranded meganuclease, the N-terminal domain of the protein contacts the first half-site, and the C-terminal domain of the protein contacts the second half-site. Cleavage by the meganuclease produces a four-base-pair 3' overhang. In embodiments where the DNA-binding polypeptide comprises a compact TALEN, the recognition sequence comprises a first CNNNGN sequence recognized by the I-TevI ​​domain, followed by a nonspecific spacer of 4-16 base pairs, followed by a second sequence of 16-22 bp recognized by the TAL effector domain (this sequence typically has 5'T bases). In embodiments where the DNA-binding polypeptide component of the fusion protein comprises a zinc finger, the DNA-binding domain typically recognizes an 18-bp recognition sequence comprising a pair of nine-base-pair "half-sites" separated by 2-10 base pairs, and cleavage by the nuclease produces a blunt end or a variable-length 5' overhang (typically four base pairs in length).

[0099] III. Deaminase

[0100] Some aspects of the present invention provide adenine deaminases generated through directed evolution and optimization of a previously disclosed adenine deaminase LPG50148 (as shown herein as SEQ ID NO:5, excluding its initiating methionine), which is disclosed in International Application No. WO 2022 / 056254, which is incorporated herein by reference in its entirety. These evolved adenine deaminases are shown in SEQ ID NO:6, 300-414, 596-598, and 720-723 (again, excluding initiating methionine).

[0101] The term "deaminase" refers to an enzyme that catalyzes a deamination reaction. The deaminase of this invention is a nucleobase deaminase, and the terms "deaminase" and "nucleobase deaminase" are used interchangeably herein. The deaminase can be a naturally occurring deaminase or its active fragment or variant. The deaminase can be active against single-stranded nucleic acids (such as ssDNA or ssRNA) or against double-stranded nucleic acids (such as dsDNA or dsRNA). In some embodiments, the deaminase is only capable of deaminating ssDNA and does not act on dsDNA.

[0102] The deaminase of this invention can be used to edit DNA or RNA molecules, and can be used alone as a deaminase or as a component of a fusion protein. In some embodiments, the deaminase can be used to edit ssDNA or ssRNA molecules. Deamination of adenine, adenosine, or deoxyadenosine produces inosine, which is then processed by a polymerase to guanine.

[0103] To date, no naturally occurring adenine deaminase is known to deaminate adenine in DNA. Several methods have been employed for evolution and optimization. do For t RNA (ADAT) protein gland purine strip Aminoaminases, which are active against DNA molecules in mammalian cells (Gaudelli et al., 2017; Koblan, LW et al., 2018, Nat Biotechnol 36, 843-846; Richter, MF et al., 2020, Nat Biotechnol, doi:10.1038 / s41587-020-0562-8, each of which is incorporated herein by reference in its entirety).

[0104] In some embodiments, the adenine deaminase of this disclosure is used in combination with a cytosine deaminase, which catalyzes the hydrolysis and deamination of cytosine, cytidine, or deoxycytidine to uracil, such as those disclosed in International Application Publications WO 2020 / 189783 and WO 2022 / 204093, and U.S. Application Publication 2022 / 0145296, each of which is incorporated herein by reference in its entirety.

[0105] The deaminase disclosed herein, or its active variants or fragments, can be introduced into cells as part of a deaminase DNA-binding polypeptide fusion, and / or can be co-expressed with a DNA-binding polypeptide deaminase fusion to improve the efficiency of introducing desired A>N (where N is C, T, or G) mutations (such as A>G mutations) into target DNA molecules.

[0106] The deaminase disclosed herein comprises an amino acid sequence having about 50% to about 100% identity with any one of SEQ ID NO:5, 6, 300-414, 596-598, and 720-723. In some embodiments, the deaminase has about 50% to about 100% identity with any one of SEQ ID NO:5, 6, 300-414, 596-598, and 720-723, and has at least one of the amino acid residues shown in any one of Tables 2, 4, 6, and 19. In some embodiments, the deaminase has an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 5, 6, 300-414, 596-598, and 720, and has at least one of the amino acid residues shown in any one of Tables 2, 4, 6, and 19. In some embodiments, the deaminase comprises an amino acid sequence selected from SEQ ID NO: 303, 319, 366, 375, 406, 596, 597, and 720. Non-limiting examples of such fusion proteins are described in the Examples section herein.

[0107] The deaminase may have at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or higher sequence identity with SEQ ID NO:5, wherein the deaminase has at least one of the following amino acid residues:

[0108] a) Position C at the position corresponding to position 2 of SEQ ID NO:5;

[0109] b) F or C at the position corresponding to position 22 of SEQ ID NO:5;

[0110] c) Q at the position corresponding to position 23 of SEQ ID NO:5;

[0111] d) N at position 35 corresponding to SEQ ID NO:5;

[0112] e) L at the position corresponding to position 40 of SEQ ID NO:5;

[0113] f) Q at the position corresponding to position 46 of SEQ ID NO:5;

[0114] g) Position A or M corresponding to position 68 of SEQ ID NO:5;

[0115] h) H at the position corresponding to position 72 of SEQ ID NO:5;

[0116] i) W, A, Q, Y or D at the position corresponding to position 75 of SEQ ID NO:5;

[0117] j) G at the position corresponding to position 76 of SEQ ID NO:5;

[0118] k) S at the position corresponding to position 81 of SEQ ID NO:5;

[0119] l) M at the position corresponding to position 105 of SEQ ID NO:5;

[0120] m) H at the position corresponding to position 108 of SEQ ID NO:5;

[0121] n) L at the position corresponding to position 109 of SEQ ID NO:5;

[0122] o) The I or L position corresponding to position 117 of SEQ ID NO:5;

[0123] p) F at the position corresponding to position 120 of SEQ ID NO:5;

[0124] q) The E, T, A, G or V at the position corresponding to position 121 of SEQ ID NO:5;

[0125] r) the I or H position corresponding to position 122 of SEQ ID NO:5;

[0126] s) K at the position corresponding to position 125 of SEQ ID NO:5;

[0127] t) A at position 126 corresponding to SEQ ID NO:5;

[0128] u) H at the position corresponding to position 135 of SEQ ID NO:5;

[0129] v) V at the position corresponding to position 137 of SEQ ID NO:5;

[0130] w) The position Y corresponding to position 138 of SEQ ID NO:5;

[0131] x) L at the position corresponding to position 139 of SEQ ID NO:5;

[0132] y) K or A at the position corresponding to position 142 of SEQ ID NO:5;

[0133] z) A, Q, L or M at the position corresponding to position 145 of SEQ ID NO:5;

[0134] aa) K at the position corresponding to position 148 of SEQ ID NO:5;

[0135] bb) Q at the position corresponding to position 151 of SEQ ID NO:5;

[0136] cc) The position E or R corresponding to position 153 of SEQ ID NO:5;

[0137] dd) W at the position corresponding to position 155 of SEQ ID NO:5;

[0138] The R or V at the position corresponding to position 156 of SEQ ID NO:5;

[0139] ff) F at the position corresponding to position 157 of SEQ ID NO:5;

[0140] The R at position 158 of SEQ ID NO:5 (gg);

[0141] hh) is the position Q corresponding to position 159 of SEQ ID NO:5;

[0142] ii) D or W at the position corresponding to position 160 of SEQ ID NO:5;

[0143] jj) is the W, H, V or A position corresponding to position 162 of SEQ ID NO:5;

[0144] kk) at position R corresponding to position 165 of SEQ ID NO:5; and

[0145] ll) H at the position corresponding to position 166 of SEQ ID NO:5.

[0146] In some embodiments, the deaminase has:

[0147] a) N at position 35 of SEQ ID NO:5 and W at position 162 of SEQ ID NO:5;

[0148] b) N at position 35 of SEQ ID NO:5 and Q at position 46 of SEQ ID NO:5;

[0149] c) S at position 81 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5;

[0150] d) Q at position 46 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5;

[0151] e) R at the position corresponding to position 156 of SEQ ID NO:5 and W at the position corresponding to position 162 of SEQ ID NO:5;

[0152] f) Position A corresponding to position 68 of SEQ ID NO: 5 and position S corresponding to position 81 of SEQ ID NO: 5;

[0153] g) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5;

[0154] h) N at position 35 of SEQ ID NO:5, Q at position 46 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5;

[0155] i) W at position 155 corresponding to SEQ ID NO:5, R at position 156 corresponding to SEQ ID NO:5, and W at position 162 corresponding to SEQ ID NO:5;

[0156] j) N at position 35 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5;

[0157] k) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5;

[0158] l) M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5;

[0159] m) Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5;

[0160] n) M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5;

[0161] o) N at position 35, Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5;

[0162] p) S at position 81 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5;

[0163] q) W at position 75 corresponding to SEQ ID NO:5 and S at position 81 corresponding to SEQ ID NO:5;

[0164] r) S at position 81 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5;

[0165] s) S at position 81 corresponding to SEQ ID NO:5 and A at position 145 corresponding to SEQ ID NO:5;

[0166] t) W at position 75 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5;

[0167] u) W at position 155 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5;

[0168] v) A at position 145 of SEQ ID NO: 5 and W at position 155 of SEQ ID NO: 5;

[0169] w) W at position 75 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5;

[0170] x) W at position 75 of SEQ ID NO:5 and A at position 145 of SEQ ID NO:5;

[0171] y) Position A corresponding to position 145 of SEQ ID NO:5 and position D corresponding to position 160 of SEQ ID NO:5;

[0172] z) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5;

[0173] aa) S at position 81 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5;

[0174] bb) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5;

[0175] cc) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5;

[0176] dd) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and A at position 145 corresponding to SEQ ID NO:5;

[0177] (ee) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5;

[0178] ff) W at position 75 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5;

[0179] gg) W at position 75 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5;

[0180] hh) A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5;

[0181] ii) W at position 75 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5;

[0182] jj) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5;

[0183] kk) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5;

[0184] ll) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5;

[0185] W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; mm)

[0186] The positions corresponding to positions 75 and 145 of SEQ ID NO: 5 are: W, A, W, and D, respectively.

[0187] oo) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5;

[0188] pp) L at position 40 of SEQ ID NO:5 and M at position 105 of SEQ ID NO:5;

[0189] qq) L at position 40 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5;

[0190] rr) L at position 40 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5;

[0191] ss) L at position 40 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5;

[0192] L at position 40 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5;

[0193] uu) M at the position corresponding to position 105 of SEQ ID NO:5 and V at the position corresponding to position 121 of SEQ ID NO:5;

[0194] vv) M at the position corresponding to position 105 of SEQ ID NO:5 and M at the position corresponding to position 145 of SEQ ID NO:5;

[0195] ww) M at position 105 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5;

[0196] xx) M at position 105 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5;

[0197] yy) V at position 121 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5;

[0198] zz) V at the position corresponding to position 121 of SEQ ID NO: 5 and Q at the position corresponding to position 159 of SEQ ID NO: 5;

[0199] aaa) V at position 121 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5;

[0200] bbb) M at position 145 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5;

[0201] ccc) M at position 145 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5;

[0202] ddd) Q at position 159 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5;

[0203] The position S corresponding to position 81 of SEQ ID NO:5, the position M corresponding to position 145 of SEQ ID NO:5, and the position W corresponding to position 160 of SEQ ID NO:5;

[0204] fff) S at position 81 corresponding to SEQ ID NO:5, V at position 121 corresponding to SEQ ID NO:5, and Q at position 159 corresponding to SEQ ID NO:5;

[0205] ggg) A at position 68 of SEQ ID NO:5 and S at position 81 of SEQ ID NO:5;

[0206] hhh) G at the position corresponding to position 76 of SEQ ID NO:5 and S at the position corresponding to position 81 of SEQ ID NO:5;

[0207] iii) S at position 81 corresponding to SEQ ID NO:5 and L at position 117 corresponding to SEQ ID NO:5;

[0208] jjj) A at position 68 of SEQ ID NO:5 and G at position 76 of SEQ ID NO:5;

[0209] kkk) A at position 68 of SEQ ID NO:5 and L at position 117 of SEQ ID NO:5;

[0210] lll) G at the position corresponding to position 76 of SEQ ID NO:5 and L at the position corresponding to position 117 of SEQ ID NO:5;

[0211] A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and S at position 81 of SEQ ID NO:5;

[0212] The positions corresponding to positions 68 and 81 of SEQ ID NO: 5 are A, S and L respectively.

[0213] ooo) G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5;

[0214] A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5;

[0215] A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5;

[0216] The positions corresponding to positions 81 and 109 of SEQ ID NO: 5 are: S, L, E, W, and D. The position corresponding to position 153 of SEQ ID NO: 5.

[0217] The positions corresponding to positions 75, 81, 117, 121, 145, 155, 160, and 160 of SEQ ID NO: 5 are: A, S, L, L, A, W, and D.

[0218] Q at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, I at position 117 of SEQ ID NO:5, Q at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5;

[0219] uuu) Y at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, L at position 117 of SEQ ID NO:5, G at position 121 of SEQ ID NO:5, L at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5;

[0220] vvv) C at position 22 of SEQ ID NO: 5, A at position 68 of SEQ ID NO: 5, Y at position 75 of SEQ ID NO: 5, G at position 76 of SEQ ID NO: 5, S at position 81 of SEQ ID NO: 5, H at position 108 of SEQ ID NO: 5, L at position 117 of SEQ ID NO: 5, H at position 122 of SEQ ID NO: 5, Y at position 138 of SEQ ID NO: 5, L at position 139 of SEQ ID NO: 5, A at position 142 of SEQ ID NO: 5, A at position 145 of SEQ ID NO: 5, R at position 153 of SEQ ID NO: 5, W at position 155 of SEQ ID NO: 5, R at position 156 of SEQ ID NO: 5, D at position 160 of SEQ ID NO: 5, and the position corresponding to SEQ ID NO: 5 The position H corresponding to position 162 of NO:5; or

[0221] The position S at the location corresponding to position 81 of SEQ ID NO:5 (www).

[0222] In some embodiments, the deaminase comprises an amino acid sequence having N at position 35 of SEQ ID NO: 5; S at position 81 of SEQ ID NO: 5; R at position 156 of SEQ ID NO: 5; and W at position 162 of SEQ ID NO: 5. In some embodiments, the deaminase comprises an amino acid sequence having W at position 75 of SEQ ID NO: 5; S at position 81 of SEQ ID NO: 5; W at position 155 of SEQ ID NO: 5; and D at position 160 of SEQ ID NO: 5. In some embodiments, the deaminase comprises an amino acid sequence having S at position 81 of SEQ ID NO: 5; and L at position 117 of SEQ ID NO: 5. In some embodiments, the deaminase comprises an amino acid sequence having A at position 68 of SEQ ID NO: 5; and S at position 81 of SEQ ID NO: 5. In some embodiments, the deaminase comprises an amino acid sequence having a Y at position 75 of SEQ ID NO: 5; an S at position 81 of SEQ ID NO: 5; an L at position 117 of SEQ ID NO: 5; a G at position 121 of SEQ ID NO: 5; an L at position 145 of SEQ ID NO: 5; an F at position 155 of SEQ ID NO: 5; and a D at position 160 of SEQ ID NO: 5. In some embodiments, the deaminase comprises an amino acid sequence having an S at position 81 of SEQ ID NO: 5; a W at position 155 of SEQ ID NO: 5; and a D at position 160 of SEQ ID NO: 5. In some embodiments, the deaminase comprises an amino acid sequence having an S at position 81 of SEQ ID NO:5; an L at position 109 of SEQ ID NO:5; an E at position 153 of SEQ ID NO:5; a W at position 155 of SEQ ID NO:5; and a D at position 160 of SEQ ID NO:5.

[0223] In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:303. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:319. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:365. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:375. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:406. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:596. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:597. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:720.

[0224] In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:303. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:319. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:365. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:375. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:406. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:596. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:597. In some embodiments, the deaminase comprises the amino acid sequence of SEQ ID NO:720.

[0225] When compared with the parental LPG50148 deaminase, the deaminases of the present invention can have improved deaminase activity. Improved deaminase activity can be measured by any method known in the art, either alone or when fused with a DNA-binding polypeptide (e.g., RGN), to measure the deamination of nucleosides (e.g., adenine). In terms of deamination of nucleosides, the deaminases disclosed herein can have an efficiency of 1.1 to 10 times, 1.1 to 20 times, 1.1 to 30 times, or greater, compared with the parental LPG50148 deaminase, including but not limited to about 1.1 times, about 1.5 times, about 2 times, about 2.5 times, about 3 times, about 3.5 times, about 4 times, about 4.5 times, about 5 times, about 5.5 times, about 6 times, etc. Approximately 6.5 times, approximately 7 times, approximately 7.5 times, approximately 8 times, approximately 8.5 times, approximately 9 times, approximately 9.5 times, approximately 10 times, approximately 11 times, approximately 12 times, approximately 13 times, approximately 15 times, approximately 16 times, approximately 17 times, approximately 18 times, approximately 19 times, approximately 20 times, approximately 21 times, approximately 22 times, approximately 23 times, approximately 24 times, approximately 25 times, approximately 26 times, approximately 27 times, approximately 28 times, approximately 29 times, and approximately 30 times.

[0226] IV. Heterologous Peptides

[0227] In some embodiments, the fusion protein of this disclosure comprises a heterologous polypeptide inserted into the RGN. The heterologous polypeptide is any polypeptide that does not naturally bind to the RGN protein. In some embodiments, the heterologous polypeptide comprises a biologically active protein domain. The heterologous polypeptide may comprise a detectable tag, a selectable tag, or a purification tag. In some embodiments, the heterologous polypeptide is a base-editing polypeptide (e.g., a deaminase) or a lead-editing polypeptide.

[0228] A. Lead Editing Peptides

[0229] The heteropeptide of the fusion protein disclosed herein can be a lead editing peptide, such as a reverse transcriptase.

[0230] Leader editing is a general and precise genome editing method that uses a polymerase-associated nucleic acid-programmable DNA-binding protein to directly write new genetic information into a designated DNA site (described, for example, in the following references: US11,447,770B1; WO2021072328; WO2021226558; WO2020156575; WO2021042047; US11193123; each of which is incorporated herein by reference in its entirety). Leader editing systems use a nickase (typically an HNH domain-inactivated nickase) as the RGN, and the system is programmed with a leader editing (PE) guide RNA (“PEgRNA”). PEgRNA is a guide RNA that both specifies the target sequence and provides a template for the polymerization of the edited alternative strand through engineered extensions on the guide RNA (e.g., at the 5' or 3' end, or within the inner portion of the guide RNA). The RGN nickase / lead editing peptide fusion is guided to the target sequence by PEgRNA, creating a gap in the target strand upstream of the sequence to be edited and upstream of the PAM, thereby forming a 3′ flap on the target strand. The pegRNA includes a primer binding site (PBS) complementary to the 3′ flap of the target strand. In some embodiments, the PBS is at least about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides in length. In some embodiments, the pegRNA comprises PBS with a length of at least 5 nucleotides (e.g., at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 28, 19, or 20 nucleotides). In some embodiments, the pegRNA may comprise PBS with a length of at least 8 nucleotides. Hybridization of the PBS with the 3′ lobe of the target strand allows the use of the PEgRNA extension as a template to polymerize the edited alternative strand. The PEgRNA extension may be formed from RNA or DNA. In the case of RNA extension, the polymerase of the leader editor may be an RNA-dependent DNA polymerase (such as reverse transcriptase). In the case of DNA extension, the polymerase of the leader editor may be a DNA-dependent DNA polymerase.

[0231] The alternative strand containing the desired edit (e.g., a single nucleobase substitution) shares the same sequence as the target strand of the target sequence to be edited (except that it includes the desired edit). The target strand of the target sequence is replaced by a newly synthesized alternative strand containing the desired edit via DNA repair and / or replication mechanisms. In some cases, leader editing can be considered a “search and replace” genome editing technique because the leader editor not only searches for and locates the desired target sequence to be edited but also simultaneously encodes an alternative strand containing the desired edit, which is installed to replace the corresponding target strand of the target sequence. Thus, in some embodiments, the guide RNA used in the compositions and methods of this disclosure comprises an extension containing an editing template for leader editing. In some embodiments, the leader editing polypeptide that can be fused with an RGN includes a DNA polymerase (e.g., an RNA-dependent DNA polymerase). In some embodiments, the DNA polymerase is a reverse transcriptase. In some embodiments, the RGN is a nicking enzyme, such as a nicking enzyme with an inactivated HNH domain.

[0232] B. Base-editing peptides

[0233] The heteropeptide of the fusion protein disclosed herein can be a base-editing peptide, such as a deaminase.

[0234] In some embodiments, the deaminase component of the fusion protein of this disclosure (wherein the deaminase is inserted into the RGN) is any one of SEQ ID NO: 5, 6, 229-414, 596, 597 and 720-723, including those deaminases disclosed herein (SEQ ID NO: 300-414, 596-598 and 720-723), or their active fragments or variants. The deaminase component may have a sequence identity between 80% and 99% or higher with any one of SEQ ID NO: 5, 6, 229-414, 596, 597 and 720-723 (which retain deaminase activity), including but not limited to at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or higher. The fusion protein may comprise an active fragment of the deaminase shown in any of SEQ ID NO:5, 6, 229-414, 596, 597, and 720-723, such as an active fragment differing from any of SEQ ID NO:5, 6, 229-414, 596, 597, and 720-723 by as few as 1-15 amino acid residues, as few as 1-10 (e.g., 6-10), as few as 5, as few as 4, as few as 3, as few as 2, or as few as 1 amino acid residue. In some embodiments, the deaminase inserted into the RGN has about 50% to about 100% identity with any of SEQ ID NO:5, 6, 300-414, 596-598, and 720-723, and has at least one of the amino acid residues shown in any of Tables 2, 4, 6, and 19. In some embodiments, the deaminase has an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 5, 6, 300-414, 596-598, and 720-723, and has at least one of the amino acid residues shown in any one of Tables 2, 4, 6, and 19.

[0235] In some embodiments, the active deaminase fragment comprises an N-terminal or C-terminal truncation, which may contain at least the deletion of 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60 or more amino acids from the N- or C-terminus of the polypeptides shown in any one of SEQ ID NO: 5, 6, 229-414, 596, 597 and 720-723. The fusion protein may comprise a deaminase lacking the first and last amino acid residues compared to the parental deaminase from which the deaminase is derived, wherein deaminase activity is retained. The fusion protein may comprise a deaminase lacking the first and last amino acid residues from any one of SEQ ID NO: 5, 6, 229-414, 596, 597 and 720-723, wherein deaminase activity is retained. In some embodiments, the active deaminase variant comprises an internal deletion, which may include at least the deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60 or more amino acids from any one of SEQ ID NO: 5, 6, 229-414, 596, 597 and 720-723. The active fragment of the deaminase may contain at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more consecutive amino acid residues of the amino acid sequence shown in any one of SEQ ID NO: 5, 6, 229-414, 596, 597 and 720-723.

[0236] In some embodiments, the deaminase component of the fusion protein of this disclosure (wherein the deaminase is inserted into the RGN) is any one of SEQ ID NO: 303, 319, 365, 375, 406, 597 and 720 or an active variant or fragment thereof.

[0237] A "base editor" is a fusion protein containing an RGN operatively linked to a deaminase, enabling the fusion protein to deaminate nucleobases of a target nucleic acid sequence (or its neighbor) upon binding to guide RNA. The RGN of a base editor typically contains an RGN nickase (usually a nickase with an inactivated RuvC domain) or a dead RGN lacking nuclease activity. Base editor fusion proteins in which the deaminase is inserted within the RGN are referred to herein as "mosaic base editors."

[0238] The fusion protein disclosed herein may contain adenine deaminases, such as those disclosed herein (e.g., 5, 6, 248-414, 596, 597, and 720-723). Base editors containing RGN and adenine deaminases are referred to herein as “A-based editors,” “adenine base editors,” or “ABE,” and can be used for targeted editing of nucleic acid sequences. To date, no naturally occurring adenine deaminases that deaminate adenine in DNA are known. Several methods have been employed for evolution and optimization. do For t RNA (ADAT) protein gland purine strip Aminoaminases, which are active against DNA molecules in mammalian cells (Gaudelli et al., 2017; Koblan, LW et al., 2018, Nat Biotechnol 36, 843-846; Richter, MF et al., 2020, Nat Biotechnol, doi:10.1038 / s41587-020-0562-8, each of which is incorporated herein by reference in its entirety).

[0239] Non-limiting examples of adenine deaminases that can be used in this invention include, but are not limited to, those disclosed in PCT International Publication No. WO 2022 / 056254, and SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723 or their active variants or fragments.

[0240] Adenine base editors (ABEs) comprise RGN and adenine deaminase. ABEs function by deamination of adenine to inosine on the DNA target molecule (Gaudelli, NM et al., 2017). Inosine is recognized as guanine by the polymerase and allows cytosine to be incorporated into the complementary DNA strand opposite inosine. Following a round of replication after deamination, an A:T to G:C base pair change occurs in the genome. In some embodiments, the adenine deaminase of the fusion protein, or its active variants or fragments, introduces an A>N mutation into the DNA molecule, where N is C, G, or T. In further embodiments, they introduce an A>G mutation into the DNA molecule.

[0241] In some embodiments, the fusion protein of this disclosure comprising an RGN having an inserted deaminase may include a cytosine deaminase that catalyzes the hydrolytic deamination of cytosine, cytidine, or deoxycytidine to uracil. Cytosine deaminases can act on DNA or RNA and generally act on single-stranded nucleic acid molecules. In some embodiments, the cytosine deaminase is a deaminase of the apolipoprotein B mRNA editing complex (APOBEC) family. In some embodiments, the cytosine deaminase is a deaminase of the APOBE11 family. In some embodiments, the cytosine deaminase is an activation-induced cytidine deaminase (AID). In some embodiments, the cytosine deaminase is an ACF1 / ASE deaminase.

[0242] Fusion proteins containing RGN and cytosine deaminase are referred to herein as “C-base editors,” “cytosine base editors,” or “CBEs.” CBEs can convert cytosine to uracil, which can then be converted to thymine via DNA replication or repair. In some embodiments, CBEs convert cytosine to guanine or adenine. Without being bound by any theory or mechanism of action, it is believed that the conversion of cytosine to guanine or adenine by the cytosine base editor is due to the deamination of cytosine to uracil and the subsequent activity of uracil DNA glycosidases during base excision repair of uracil residues. In some embodiments, the cytosine deaminase of the fusion protein, or its active variants or fragments, introduces a C>N mutation into the DNA molecule, where N is A, G, or T. In further embodiments, they introduce a C>G or C>A mutation into the DNA molecule.

[0243] Non-limiting examples of cytosine deaminases that can be used in this invention include, but are not limited to, those disclosed in PCT International Publication Nos. WO 2020 / 139783 and WO 2022 / 204093, as well as SEQ ID NO:229-247 or its active variants or fragments.

[0244] The mutation rate of adenine or cytosine within or adjacent to the target sequence that binds to the RGN deaminase fusion protein can be measured using any method known in the art, including polymerase chain reaction (PCR), restriction fragment length polymorphism (RFLP), or DNA sequencing.

[0245] V. Fusion protein

[0246] This invention provides various types of fusion proteins. Those skilled in the art will understand that, whenever a protein fuses with another protein, the N-terminal methionine (Met) of the C-terminal protein can optionally be omitted from the sequence. Therefore, as used herein, reference to the fusion of a given amino acid sequence with another amino acid sequence explicitly includes such sequences having or optionally not having an N-terminal Met, regardless of whether such sequences contain an N-terminal Met as listed in the sequence listing.

[0247] Some aspects of the present invention relate to fusion proteins comprising heterologous polypeptides inserted into an RGN. The heterologous polypeptide can be inserted into the RGN at its surface (i.e., between amino acid residues on the surface or within a surface loop).

[0248] Heterologous peptides can be inserted into the adapter domain 2, wedge (WED) domain, RuvC domain, HNH domain, Rec-2 domain, or PAM interaction (PI) domain. In some embodiments, the RuvC domain is the RuvCIII domain. The Rec or recognition lobe mediates nucleic acid binding through multiple Rec domains (e.g., Rec1-3) by sensing nucleic acids, regulates the HNH conformational transition, and locks the catalytic HNH domain at the cleavage site. The wedge domain is responsible for recognizing and guiding the RNA scaffold. Unrestricted examples of domains within the RGN include: RuvC-I from amino acid residues 1-54; BH from amino acid residues 55-83; REC1 from amino acid residues 84-244; REC2 from amino acid residues 245-462; RuvC-II from amino acid residues 463-521; L1 from amino acid residues 522-552; HNH from amino acid residues 553-672; L2 from amino acid residues 673-685; RuvC-III from amino acid residues 686-833; WED from amino acid residues 834-938; and PI from amino acid residues 939-1071; all with reference to APG07433.1 as shown in SEQ ID NO: 1. APG05586 (as shown in SEQ ID NO:2) has the following domains: RuvC-I from amino acid residues 1-33; BH from amino acid residues 34-71; REC1 from amino acid residues 72-232; REC2 from amino acid residues 233-468; RuvC-II from amino acid residues 469-517; L1 from amino acid residues 518-552; HNH from amino acid residues 553-672; L2 from amino acid residues 673-687; RuvC-III from amino acid residues 688-837; WED from amino acid residues 838-998; and PI from amino acid residues 999-1150. LPG10145 (as shown in SEQ ID NO:4) has the following domains: RuvC-I from amino acid residues 1-42; BH from amino acid residues 43-79; REC1 from amino acid residues 80-236; REC2 from amino acid residues 237-476; RuvC-II from amino acid residues 477-524; L1 from amino acid residues 525-560; HNH from amino acid residues 561-676; L2 from amino acid residues 677-690; RuvC-III from amino acid residues 691-828; WED from amino acid residues 829-976; and PI from amino acid residues 977-1130.APG01604 (as shown in SEQ ID NO:3) has the following domains: RuvC-I from amino acid residues 1-40; BH from amino acid residues 41-74; REC1 from amino acid residues 75-223; REC2 from amino acid residues 224-430; RuvC-II from amino acid residues 431-483; L1 from amino acid residues 484-516; HNH from amino acid residues 517-631; L2 from amino acid residues 632-651; RuvC-III from amino acid residues 652-775; WED from amino acid residues 776-909; and PI from amino acid residues 910-1052.

[0249] The PAM interaction domain is the domain that binds to the PAM sequence. The general domains of RGN proteins can be determined by structural comparison with RGN proteins that have defined domains.

[0250] In those embodiments where the fusion protein comprises an RGN having at least 90% sequence identity with SEQ ID NO:1 or 49, a heterologous polypeptide (e.g., a deaminase) may be inserted into the RGN immediately following an amino acid position selected from the group consisting of:

[0251] i) The amino acid position corresponding to position 30 of SEQ ID NO:1;

[0252] ii) The amino acid position corresponding to position 642 in SEQ ID NO:1;

[0253] iii) The amino acid position corresponding to position 670 of SEQ ID NO:1;

[0254] iv) The amino acid position corresponding to position 737 of SEQ ID NO:1;

[0255] v) The amino acid position corresponding to position 772 of SEQ ID NO:1;

[0256] vi) The amino acid position corresponding to position 775 of SEQ ID NO:1;

[0257] vii) The amino acid position corresponding to position 778 of SEQ ID NO:1; and

[0258] viii) The amino acid position corresponding to position 802 of SEQ ID NO:1.

[0259] In those embodiments where the fusion protein comprises an RGN having at least 90% sequence identity with SEQ ID NO:2 or 50, the heteropeptide may be inserted into the RGN immediately following an amino acid position selected from the group consisting of:

[0260] i) The amino acid position corresponding to position 678 of SEQ ID NO:2;

[0261] ii) The amino acid position corresponding to position 736 of SEQ ID NO:2;

[0262] iii) The amino acid position corresponding to position 778 of SEQ ID NO:2;

[0263] iv) The amino acid position corresponding to position 788 of SEQ ID NO:2; and

[0264] v) The amino acid position corresponding to position 922 of SEQ ID NO:2.

[0265] In those embodiments where the fusion protein comprises an RGN having at least 90% sequence identity with SEQ ID NO:3 or 51, the heteropeptide may be inserted into the RGN immediately following an amino acid position selected from the group consisting of:

[0266] i) The amino acid position corresponding to position 725 of SEQ ID NO:3;

[0267] ii) The amino acid position corresponding to position 739 in SEQ ID NO:3; and

[0268] iii) The amino acid position corresponding to position 744 of SEQ ID NO:3.

[0269] In those embodiments where the fusion protein comprises an RGN having at least 90% sequence identity with SEQ ID NO:4 or 52, the heteropeptide may be inserted into the RGN immediately following an amino acid position selected from the group consisting of:

[0270] i) The amino acid position corresponding to position 347 in SEQ ID NO:4;

[0271] ii) The amino acid position corresponding to position 524 in SEQ ID NO:4;

[0272] iii) The amino acid position corresponding to position 666 in SEQ ID NO:4;

[0273] iv) The amino acid position corresponding to position 680 of SEQ ID NO:4;

[0274] v) The amino acid position corresponding to position 740 of SEQ ID NO:4;

[0275] vi) The amino acid position corresponding to position 785 of SEQ ID NO:4;

[0276] vii) The amino acid position corresponding to position 910 of SEQ ID NO:4; and

[0277] viii) The amino acid position corresponding to position 1077 of SEQ ID NO:4.

[0278] In those embodiments where the fusion protein comprises an RGN having at least 90% sequence identity with SEQ ID NO: 131 or 698, the heteropeptide may be inserted into the RGN immediately following an amino acid position selected from the group consisting of:

[0279] i) The amino acid position corresponding to position 766 of SEQ ID NO: 131; and

[0280] ii) The amino acid position corresponding to position 806 of SEQ ID NO:131.

[0281] The fusion protein of the present invention may comprise a base editor fusion protein, wherein the heterologous polypeptide is a deaminase, and the deaminase is inserted into an RGN (such as an RGN nicking enzyme or a nuclease-inactive RGN). In some embodiments, the RGN component of the base editor fusion protein comprises an RGN nicking enzyme or a nuclease-inactive RGN having about 60% to about 99.5% identity with any one of SEQ ID NO:49-52 and 698. In some embodiments, the RGN nicking enzyme or nuclease-inactive RGN comprises an amino acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with any one of SEQ ID NO:49-52 and 698.

[0282] In some embodiments, the base editor fusion protein comprises an RGN nicking enzyme or a nuclease-inactive RGN containing an amino acid sequence having a sequence identity between 80% and 99% or higher with any one of SEQ ID NO:49-52 and 698, including but not limited to at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or higher.

[0283] The fusion protein of the present invention may comprise a base editor fusion protein (wherein the heterologous polypeptide is a deaminase, and the deaminase is inserted into an RGN (such as an RGN nicking enzyme or a nuclease-inactive RGN)), wherein the base editor fusion protein has improved editing activity compared to a parental base editor comprising a deaminase fused to the N-terminus of an RGN. The improved editing activity can be measured by any method known in the art. In terms of editing, the base editor fusion protein disclosed herein may have an efficiency of 1.1 to 10 times, 1.1 to 20 times, 1.1 to 30 times or greater compared to a parental end-to-end fusion protein, including but not limited to about 1.1 times, about 1.5 times, about 2 times, about 2.5 times, about 3 times, about 3.5 times, about 4 times, about 4.5 times, about 5 times, about 5.5 times, about 6 times, about 6 times, etc. 0.5 times, approximately 7 times, approximately 7.5 times, approximately 8 times, approximately 8.5 times, approximately 9 times, approximately 9.5 times, approximately 10 times, approximately 11 times, approximately 12 times, approximately 13 times, approximately 15 times, approximately 16 times, approximately 17 times, approximately 18 times, approximately 19 times, approximately 20 times, approximately 21 times, approximately 22 times, approximately 23 times, approximately 24 times, approximately 25 times, approximately 26 times, approximately 27 times, approximately 28 times, approximately 29 times, and approximately 30 times.

[0284] The binding of a base editor fusion protein to a target sequence results in the modification of nucleotides adjacent to the target sequence. The nucleotides adjacent to the target sequence modified by the deaminase can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 base pairs from the 5' or 3' end of the target sequence. The full range of nucleotides that can be edited using a base editor fusion protein (e.g., an RGN fused with a deaminase) is typically represented as the distance of nucleotides from the PAM sequence (e.g., nucleotides at positions 8 and 22 upstream (i.e., 5') of the PAM sequence), referred to herein as the “editing window” of a particular base editor fusion protein. The editing window can be, for example, 1-100 base pairs (5' or 3') of a PAM sequence, including but not limited to 5-50, 5-25, 8-22, 10-20, or 10-21 base pairs (5' or 3') of a PAM sequence.

[0285] Unrestricted examples of mosaic base editor fusion proteins include:

[0286] a) The LPG50274 / LPG50148.2.76 (as shown in SEQ ID NO:375) or its active variants or fragments, such as the sequence shown in SEQ ID NO:599, is inserted after the amino acid position corresponding to position 922 of SEQ ID NO:2 into nAPG05586 (as shown in SEQ ID NO:50) or its active variants or fragments;

[0287] b) Inserting LPG50274 / LPG50148.2.76 (as shown in SEQ ID NO:375) or its active variants or fragments into nLPG10145 (as shown in SEQ ID NO:52) or its active variants or fragments, such as the sequence shown in SEQ ID NO:600, after the amino acid position corresponding to position 910 of SEQ ID NO:4;

[0288] c) LPG50221 / LPG50148.2.20 (as shown in SEQ ID NO:319) or its active variants or fragments, such as the sequence shown in SEQ ID NO:602, inserted after the amino acid position corresponding to position 922 of SEQ ID NO:2;

[0289] d) Inserting LPG50221 / LPG50148.2.20 (as shown in SEQ ID NO:319) or its active variants or fragments into nLPG10145 (as shown in SEQ ID NO:52) or its active variants or fragments, such as the sequence shown in SEQ ID NO:603, after the amino acid position corresponding to position 910 of SEQ ID NO:4;

[0290] e) Inserting LPG50274 / LPG50148.2.76 (as shown in SEQ ID NO:375) or its active variants or fragments into nAPG07433.1 (as shown in SEQ ID NO:49) or its active variants or fragments, such as the sequence shown in SEQ ID NO:606, after the amino acid position corresponding to position 772 of SEQ ID NO:1;

[0291] f) Inserting LPG50274 / LPG50148.2.76 (as shown in SEQ ID NO:375) or its active variants or fragments into nAPG05586 (as shown in SEQ ID NO:50) or its active variants or fragments, such as the sequence shown in SEQ ID NO:610, after the amino acid position corresponding to position 678 of SEQ ID NO:2;

[0292] g) Inserting LPG50221 / LPG50148.2.20 (as shown in SEQ ID NO:319) or its active variants or fragments into nAPG07433.1 (as shown in SEQ ID NO:49) or its active variants or fragments, such as the sequence shown in SEQ ID NO:612, after the amino acid position corresponding to position 772 of SEQ ID NO:1;

[0293] h) Inserting LPG50324 (as shown in SEQ ID NO:597) or its active variants or fragments into nAPG05586 (as shown in SEQ ID NO:50) or its active variants or fragments, such as the sequence shown in SEQ ID NO:614, after the amino acid position corresponding to position 922 of SEQ ID NO:2.

[0294] i) Inserting LPG50221 / LPG50148.2.20 (as shown in SEQ ID NO:319) or its active variants or fragments into nAPG05586 (as shown in SEQ ID NO:50) or its active variants or fragments, such as the sequence shown in SEQ ID NO:615, after the amino acid position corresponding to position 678 of SEQ ID NO:2;

[0295] j) Inserting LPG50319 / LPG50148.2.107 (as shown in SEQ ID NO:406) or its active variants or fragments into nAPG05586 (as shown in SEQ ID NO:50) or its active variants or fragments, such as the sequence shown in SEQ ID NO:617, after the amino acid position corresponding to position 678 of SEQ ID NO:2;

[0296] k) Inserting LPG50324 (as shown in SEQ ID NO:597) or its active variants or fragments, such as the sequence shown in SEQ ID NO:619, after the amino acid position corresponding to position 678 of SEQ ID NO:2.

[0297] l) Inserting LPG50320 (as shown in SEQ ID NO:596) or its active variants or fragments, such as the sequence shown in SEQ ID NO:620, after the amino acid position corresponding to position 678 of SEQ ID NO:2.

[0298] m) Inserting LPG50319 / LPG50148.2.107 (as shown in SEQ ID NO:406) or its active variants or fragments into nLPG10145 (as shown in SEQ ID NO:52) or its active variants or fragments, such as the sequence shown in SEQ ID NO:621, after the amino acid position corresponding to position 910 of SEQ ID NO:4;

[0299] n) Inserting LPG50324 (as shown in SEQ ID NO:597) or its active variants or fragments, such as the sequence shown in SEQ ID NO:622, after the amino acid position corresponding to position 910 of SEQ ID NO:4;

[0300] o) Inserting LPG50319 / LPG50148.2.107 (as shown in SEQ ID NO:406) or its active variants or fragments into nAPG05586 (as shown in SEQ ID NO:50) or its active variants or fragments, such as the sequence shown in SEQ ID NO:623, after the amino acid position corresponding to position 922 of SEQ ID NO:2;

[0301] p) Inserting LPG50320 (as shown in SEQ ID NO:596) or its active variants or fragments, such as the sequence shown in SEQ ID NO:624, after the amino acid position corresponding to position 910 of SEQ ID NO:4;

[0302] q) Inserting LPG50320 (as shown in SEQ ID NO:596) or its active variants or fragments into nAPG05586 (as shown in SEQ ID NO:50) or its active variants or fragments, such as the sequence shown in SEQ ID NO:626, after the amino acid position corresponding to position 922 of SEQ ID NO:2;

[0303] r) Inserting LPG50319 / LPG50148.2.107 (as shown in SEQ ID NO:406) or its active variants or fragments into nAPG07433.1 (as shown in SEQ ID NO:49) or its active variants or fragments, such as the sequence shown in SEQ ID NO:630, after the amino acid position corresponding to position 772 of SEQ ID NO:1;

[0304] s) LPG50320 (as shown in SEQ ID NO: 596) or its active variants or fragments, such as the sequence shown in SEQ ID NO: 631, inserted after the amino acid position corresponding to position 772 of SEQ ID NO: 1; and

[0305] t) Inserting LPG50324 (as shown in SEQ ID NO:597) or its active variants or fragments, such as the sequence shown in SEQ ID NO:632, after the amino acid position corresponding to position 772 of SEQ ID NO:1.

[0306] The fusion protein of the present invention may comprise a base editor fusion protein (wherein the heterologous polypeptide is a deaminase, and the deaminase is inserted within an RGN (such as an RGN nickase or a nuclease-inactive RGN)), wherein the base editor fusion protein has a shifted editing window compared to a parental base editor comprising a deaminase fused to the N-terminus of the RGN. The editing window may be shifted to be narrower or wider compared to the parental base editor. When compared to the parental end-to-end fusion protein, the shifted editing window is a wider or narrower editing window (e.g., differing by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides), and / or when compared to the parental end-to-end fusion protein, the editing window shifts in the 5' or 3' direction of the target molecule, for example, shifting by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides in the 5' or 3' direction.

[0307] In some embodiments, the fusion protein includes one or more peptide linkers between the heteropeptide and the RGN. Those skilled in the art will understand that linkers may optionally be added to the sequence whenever a protein fuses with another protein. Therefore, as used herein, references to a given fusion protein explicitly include such sequences having or optionally not having linkers, regardless of whether such sequences are listed as having linkers. The linker between the deaminase and the RGN can determine the editing window of the fusion protein. Various linker lengths and flexibility can be employed, ranging from very flexible linker forms (GGGGS). n and (G) n To a more rigid joint type (EAAAK) n and (XP) nTo achieve optimal length and stiffness for the deaminase activity of a specific application. As used herein, the term "peptide linker" refers to a peptide that connects two polypeptides. In some embodiments, the linker connects an RNA-guided nuclease and a deaminase. In some embodiments, the linker connects a dead or inactive RGN and a deaminase. In some embodiments, the linker length is 3-100 amino acids, for example, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids. Longer or shorter linkers are also considered. In some embodiments, shorter linkers are preferred to reduce the overall size or length of the fusion protein or its coding sequence. Non-limiting examples of peptide linkers between the RGN and the inserted heteropeptide include GGS, SGG, and GSSG (as shown in SEQ ID NO: 642). The linker may be located at the N-terminus, C-terminus, or both of the inserted heteropeptide.

[0308] In some embodiments, the RGN and the inserted heteropeptide are directly fused to each other without a linker sequence.

[0309] Another aspect of the invention relates to a fusion protein comprising a DNA-binding polypeptide (e.g., an inactive nuclease or a nickase RGN) operatively linked to a deaminase of the invention (e.g., SEQ ID NO: 6, 300-414, 596-598, and 720-723, or an active fragment or variant thereof). In some embodiments, the DNA-binding polypeptide (e.g., an inactive nuclease RGN or a nickase RGN) fused to the deaminase of the invention can be targeted to a specific location on a nucleic acid molecule (i.e., a target nucleic acid molecule) to alter the expression of a desired sequence, which in some embodiments is a specific genomic locus. In some embodiments, binding of the fusion protein to the target sequence results in the deamination of a nucleobase, thereby resulting in the conversion from one nucleobase to another. In some embodiments, binding of the fusion protein to the target sequence results in the deamination of a nucleobase adjacent to the target sequence.

[0310] Some aspects of this disclosure provide fusion proteins comprising (i) a DNA-binding polypeptide (e.g., a nuclease-inactive or nickase RGN polypeptide); (ii) a deaminase polypeptide; and optionally (iii) a second deaminase. The second deaminase may be the same as the first deaminase, or it may be a different deaminase. In some embodiments, both the first and second deaminases are adenine deaminases of the present invention.

[0311] This disclosure provides fusion proteins with various configurations. In some embodiments, the fusion protein is an end-to-end fusion, wherein the deaminase peptide is fused to the N-terminus of a DNA-binding peptide (e.g., an RGN peptide), or the deaminase peptide is fused to the C-terminus of a DNA-binding peptide (e.g., an RGN peptide).

[0312] Non-limiting examples of end-to-end base editor fusion proteins include those comprising an N-terminal fused deaminase or its active variant or fragment thereof, containing a nicking enzyme or an active variant or fragment thereof, as used in the examples herein, including but not limited to:

[0313] a) LPG50221 / LPG50148.2.20 (as shown in SEQ ID NO:319) or its active variants or fragments fused to the N-terminus of nAPG07433.1 (as shown in SEQ ID NO:49) or its active variants or fragments, such as sequences shown in SEQ ID NO:578 or 582;

[0314] b) LPG50274 / LPG50148.2.76 (as shown in SEQ ID NO:375) or its active variants or fragments fused to the N-terminus of nAPG07433.1 (as shown in SEQ ID NO:49) or its active variants or fragments, such as sequences shown in SEQ ID NO:594 or 577;

[0315] c) LPG50319 / LPG50148.2.107 (as shown in SEQ ID NO:406) or its active variants or fragments fused to the N-terminus of nAPG07433.1 (as shown in SEQ ID NO:49) or its active variants or fragments, such as sequences shown in SEQ ID NO:588 or 616;

[0316] d) LPG50320 (as shown in SEQ ID NO: 596) or its active variants or fragments fused to the N-terminus of nAPG07433.1 (as shown in SEQ ID NO: 49) or its active variants or fragments, such as sequences shown in SEQ ID NO: 593 or 604; and

[0317] e) LPG50324 (as shown in SEQ ID NO:597) or its active variants or fragments fused to the N-terminus of nAPG07433.1 (as shown in SEQ ID NO:49) or its active variants or fragments, such as the sequence shown in SEQ ID NO:605.

[0318] In some embodiments, end-to-end deaminases and DNA-binding peptides (e.g., RNA-guided DNA-binding peptides) are fused to each other via a linker. As used herein, the term "linker" refers to a chemical group or molecule that connects two molecules or portions, such as the binding and cleavage domains of a nuclease. In some embodiments, the linker connects an RNA-guided nuclease and a deaminase. In some embodiments, the linker connects a dead or inactive RGN and a deaminase. In further embodiments, the linker connects two deaminases. In some embodiments, the linker connects an RNA-guided nuclease and a USP. In some embodiments, the linker connects a deaminase and a USP. In some embodiments, the linker fuses an RNA-guided nuclease deaminase to a USP. Typically, the linker is located between or on either side of two groups, molecules, or other portions and is covalently linked to each of them, thereby connecting the two groups, molecules, or other portions. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical part. In some embodiments, the linker length is 3-100 amino acids, for example, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids. Longer or shorter linkers are also considered. In some embodiments, shorter linkers are preferred to reduce the overall size or length of the fusion protein or its coding sequence.

[0319] In some embodiments, the linker between the deaminase and the DNA-binding polypeptide in the end-to-end fusion comprises any one of SEQ ID NO: 704, 705, and 707-712. Additional suitable linker motifs and conformations will be apparent to those skilled in the art. In some embodiments, suitable linker motifs and conformations include those described in Chen et al., 2013 (Adv Drug Deliv Rev. 65(10): 1357-69, the entire contents of which are incorporated herein by reference). Additional suitable linker sequences will be apparent to those skilled in the art. In some embodiments, the linker sequence comprises an amino acid sequence as shown in SEQ ID NO: 233 or 234.

[0320] In some embodiments, the overall architecture of the exemplary fusion protein provided herein comprises the following structures: [NH2]-[deaminase]-[DBP]-[COOH]; [NH2]-[DBP]-[deaminase]-[COOH]; [NH2]-[DBP]-[deaminase]-[deaminase]-[COOH]; [NH2]-[deaminase]-[DBP]-[deaminase]-[COOH]; or [NH2]-[deaminase]-[deaminase]-[DBP]-[COOH], wherein DBP is a DNA-binding polypeptide, NH2 is the N-terminus of the fusion protein, and COOH is the C-terminus of the fusion protein. In some embodiments, the fusion protein comprises more than two deaminase polypeptides.

[0321] In some embodiments, the overall architecture of the exemplary fusion proteins provided herein comprises the following structures: [NH2]-[deaminase]-[RGN]-[COOH]; [NH2]-[RGN]-[deaminase]-[COOH]; [NH2]-[RGN]-[deaminase]-[deaminase]-[COOH]; [NH2]-[deaminase]-[RGN]-[deaminase]-[COOH]; or [NH2]-[deaminase]-[deaminase]-[RGN]-[COOH], wherein NH2 is the N-terminus of the fusion protein and COOH is the C-terminus of the fusion protein. In some embodiments, the fusion protein comprises more than two deaminase polypeptides.

[0322] In some embodiments, the fusion protein comprises the following structures: [NH2]-[deaminase]-[nuclease-inactive RGN]-[COOH]; [NH2]-[deaminase]-[deaminase]-[nuclease-inactive RGN]-[COOH]; [NH2]-[nuclease-inactive RGN]-[deaminase]-[COOH]; [NH2]-[deaminase]-[nuclease-inactive RGN]-[deaminase]-[COOH]; or [NH2]-[nuclease-inactive RGN]-[deaminase]-[deaminase]-[COOH]. It should be understood that "nuclease-inactive RGN" represents any RGN, including any CRISPR-Cas protein that has been mutated to nuclease-inactive. In some embodiments, the fusion protein comprises more than two deaminase polypeptides.

[0323] In some embodiments, the fusion protein comprises the following structures: [NH2]-[deaminase]-[RGN nickase]-[COOH]; [NH2]-[deaminase]-[deaminase]-[RGN nickase]-[COOH]; [NH2]-[RGN nickase]-[deaminase]-[COOH]; [NH2]-[deaminase]-[RGN nickase]-[deaminase]-[COOH]; or [NH2]-[RGN nickase]-[deaminase]-[deaminase]-[COOH]. It should be understood that "RGN nickase" represents any RGN, including any CRISPR-Cas protein that has been mutated to have nickase-like activity.

[0324] In some embodiments, the fusion protein provided herein contains the full-length sequence of the deaminase. However, in some embodiments, the fusion protein provided herein does not contain the full-length sequence of the deaminase, but only fragments thereof.

[0325] In some embodiments, the fusion protein of the present invention comprises a DNA-binding polypeptide (e.g., RGN) and a deaminase, wherein the deaminase has an amino acid sequence having about 50% to about 100% identity with any one of SEQ ID NO:6, 300-414, 596-598, and 720-723. In some embodiments, the fusion protein of the present invention comprises a DNA-binding polypeptide (e.g., RGN) and a deaminase, wherein the deaminase has an amino acid sequence having about 50% to about 100% identity with any one of SEQ ID NO:6, 300-414, 596-598, and 720-723, and has at least one of the amino acid residues shown in any one of Tables 2, 4, 6, and 19. In some embodiments, the fusion protein of the present invention comprises a DNA-binding polypeptide (e.g., RGN) and a deaminase, wherein the deaminase has an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 6, 300-414, 596-598, and 720-723, and has at least one of the amino acid residues shown in any one of Tables 2, 4, 6, and 19. Examples of such fusion proteins are described in the Examples section herein.

[0326] In some embodiments, the fusion protein comprises a deaminase polypeptide. In some embodiments, the fusion protein comprises at least two deaminase polypeptides operably linked directly or via a peptide linker. In some embodiments, the fusion protein comprises a deaminase polypeptide, and a second deaminase polypeptide is co-expressed with the fusion protein.

[0327] In some embodiments, the "-" used in the overall architecture above indicates the presence of an optional adapter sequence. In some embodiments, the fusion protein provided herein does not contain an adapter sequence. In some embodiments, at least one of the optional adapter sequences is present.

[0328] Other exemplary features that may be present on the deaminases or fusion proteins of this disclosure are localization sequences, such as nuclear localization sequences, cytoplasmic localization sequences, export sequences, such as nuclear export sequences, or other localization sequences, and sequence tags that can be used to dissolve, purify, or detect the fusion protein. Suitable localization signal sequences and protein tag sequences are provided herein, and include, but are not limited to, biotinylate carboxylase carrier protein (BCCP) tags, myc tags, calmodulin tags, FLAG tags (e.g., 3XFLAG tags), hemagglutinin (HA) tags, multihistidine tags (also referred to as histidine tags or His tags), maltose-binding protein (MBP) tags, nus tags, glutathione-S-transferase (GST) tags, green fluorescent protein (GFP) tags, thioredoxin tags, S tags, Softag (e.g., Softag 1, Softag 3), strep tags, biotin ligase tags, FLAsH tags, V5 tags, and SBP tags. Additional suitable sequences will be apparent to those skilled in the art.

[0329] The compositions and methods disclosed herein can enhance the transport of deaminases or fusion proteins to the cell nucleus by utilizing deaminases or fusion proteins containing at least one nuclear localization signal (NLS). Nuclear localization signals are known in the art and typically comprise a basic amino acid sequence (see, for example, Lange et al., J. Biol. Chem. (2007) 282:5101-5105). In some embodiments, the deaminase or fusion protein contains 2, 3, 4, 5, 6, or more nuclear localization signals. The nuclear localization signal can be a heterologous NLS. Non-limiting examples of nuclear localization signals that can be used for the RGN of this disclosure are nuclear localization signals for the SV40 large T antigen, nucleoplasmic proteins, c-Myc (see, for example, Ray et al. (2015) Bioconjug Chem 26(6):1004-7), POLD1 DNA polymerase δ1, active adenosine deaminase acting on RNA (ADAR2), Interact with RNA polymerase II 1 (Iwr1) (Czeko E. et al., MolCell 2011, which is incorporated herein by reference in its entirety), and OpT (US Publication No. 2020 / 0109382, which is incorporated herein by reference in its entirety). In embodiments, the RGN comprises an NLS sequence as shown in any one of SEQ ID NO:430, 431, and 713-718. The deaminase or fusion protein may contain one or more NLS sequences at its N-terminus, C-terminus, or both. For example, a deaminase or fusion protein may contain two NLS sequences at its N-terminal region and four NLS sequences at its C-terminal region. In some embodiments, the deaminase or fusion protein contains an SV40 NLS (such as the sequence shown in SEQ ID NO:430) at its N-terminus and a nucleoplasmic protein NLS (such as the sequence shown in SEQ ID NO:431) at its C-terminus. In some embodiments, the deaminase or fusion protein contains a c-Myc promoter (such as the sequence shown in SEQ ID NO:718) at both its N-terminus and its C-terminus.

[0330] When an NLS is attached to the N-terminus, C-terminus, or both of a deaminase or fusion protein, an NLS adaptor protein may be present to separate the deaminase or fusion protein from the NLS. This NLS adaptor protein has a length of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more amino acids. In some embodiments, the NLS adaptor protein has a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8 amino acids. In some embodiments, the NLS adaptor protein between the NLS and the deaminase or fusion protein has a sequence as shown in SEQ ID NO: 706, 719, or 728. In some embodiments, the deaminase or fusion protein includes a c-Myc promoter (such as the sequence shown in SEQ ID NO:718) at its N-terminus, which is separated from the deaminase or fusion protein by an NLS adaptor protein having the sequence shown in SEQ ID NO:728; and includes a c-Myc promoter (such as the sequence shown in SEQ ID NO:718) at its C-terminus, which is separated from the deaminase or fusion protein by an NLS adaptor protein having the sequence shown in SEQ ID NO:728.

[0331] In some embodiments, the compositions and methods of this disclosure utilize a deaminase or fusion protein comprising at least one cell-penetrating domain that facilitates cellular uptake of the deaminase or fusion protein. Cell-penetrating domains are known in the art and typically comprise segments of positively charged amino acid residues (i.e., multi-cationic cell-penetrating domains), alternating polar and nonpolar amino acid residues (i.e., amphiphilic cell-penetrating domains), or hydrophobic amino acid residues (i.e., hydrophobic cell-penetrating domains) (see, for example, Milletti F. (2012) Drug Discov Today 17:850-860). A non-limiting example of a cell-penetrating domain is the trans-activator of transcription (TAT) from human immunodeficiency virus 1.

[0332] Nuclear localization signals and / or cell penetration domains may be located at the N-terminus, C-terminus, and / or internal location of deaminases or fusion proteins.

[0333] VI. Nucleotides encoding deaminases, fusion proteins, and / or gRNAs

[0334] This disclosure provides polynucleotides encoding the deaminases, fusion proteins, or gRNAs disclosed herein. In some embodiments, the polynucleotide encodes a fusion protein comprising a deaminase and a DNA-binding polypeptide, such as a meganuclease, zinc finger fusion protein, or TALEN. The invention further provides polynucleotides encoding fusion proteins comprising a deaminase and an RNA-guided DNA-binding polypeptide (RGDBP). Such RNA-guided DNA-binding polypeptides may be RGN or RGN variants. Protein variants may be nuclease-inactive or nickase-inactive. RGN may be a CRISPR-Cas protein or its active variants or fragments. Examples of CRISPR-Cas nucleases are well known in the art, and similar corresponding mutations can produce mutant variants that are also nickase- or nuclease-inactive.

[0335] Embodiments of the present invention provide a polynucleotide encoding a fusion protein comprising the RDBBP (e.g., RGN) described herein and a deaminase. In some embodiments, a second polynucleotide encodes a guide RNA required for the RDBBP to target a nucleotide sequence of interest. In some embodiments, the guide RNA and the fusion protein are encoded by the same polynucleotide.

[0336] In some embodiments, the guide RNA and the fusion protein are encoded by the same polynucleotide. In other embodiments, the guide RNA and the fusion protein are encoded by two separate polynucleotides.

[0337] The use of the term "polynucleotide" is not intended to limit this disclosure to polynucleotides containing DNA, although such DNA polynucleotides are considered. Those skilled in the art will recognize that polynucleotides can comprise ribonucleotides (RNA) (e.g., mRNA) as well as combinations of ribonucleotides and deoxyribonucleotides. Such deoxyribonucleotides and ribonucleotides include both naturally occurring molecules and synthetic analogs. The polynucleotides disclosed herein also encompass all forms of sequences, including but not limited to single-stranded, double-stranded, stem-and-loop structures, circular forms (e.g., including circular RNA), etc.

[0338] An embodiment of the present invention comprises a nucleic acid molecule comprising a sequence encoding a deaminase having about 50% to about 100% identity with any one of SEQ ID NO: 5, 6, 300-414, 596-598, and 720-723. An embodiment of the present invention comprises a nucleic acid molecule comprising a sequence encoding a deaminase having about 50% to about 100% identity with any one of SEQ ID NO: 5, 6, 300-414, 596-598, and 720-723, and having at least one of the amino acid residues shown in any one of Tables 2, 4, 6, and 19. An embodiment of the present invention comprises a nucleic acid molecule comprising a sequence encoding a deaminase having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 5, 6, 300-414, 596-598, and 720-723, and having at least one of the amino acid residues shown in any one of Tables 2, 4, 6, and 19, wherein the nucleic acid molecule encodes a deaminase having deaminase activity. The nucleic acid molecule may further comprise a heterologous promoter or terminator. The nucleic acid molecule may encode a fusion protein wherein the encoded deaminase is operatively linked to a DNA-binding polypeptide and optionally a second deaminase. In some embodiments, the nucleic acid molecule may encode a fusion protein wherein the encoded deaminase is operatively linked to an RGN and optionally a second deaminase.

[0339] In some embodiments, nucleic acid molecules comprising a polynucleotide encoding the deaminase or fusion protein of the present invention are codon-optimized for expression in an organism of interest. A “codon-optimized” coding sequence is a polynucleotide coding sequence whose codon usage frequencies are designed to mimic the frequencies of preferred codon usage or transcriptional conditions in a particular host cell. Expression in a particular host cell or organism is enhanced due to changes in one or more codons at the nucleic acid level, without altering the translated amino acid sequence. Nucleic acid molecules can be codon-optimized in whole or in part. Codon tables and other references providing preference information for various organisms are available in the art (see, for example, Campbell and Gowri (1990) Plant Physiol. 92:1-11, which discusses plant preferred codon usage). Methods for synthesizing plant preferred genes are available in the art. See, for example, U.S. Patent Nos. 5,380,831 and 5,436,391, and Murray et al. (1989) Nucleic Acids Res. 17:477-498, which are incorporated herein by reference.

[0340] In some embodiments, a polynucleotide encoding the deaminase, fusion protein, and / or gRNA described herein is provided in an expression cassette for expression in vitro or in cells, organelles, embryos, or organisms of interest. The expression cassette may include 5' and 3' regulatory sequences operatively linked to the polynucleotide encoding the deaminase, fusion protein, and / or gRNA provided herein, the regulatory sequences allowing expression of the polynucleotide. The expression cassette may additionally contain at least one additional gene or genetic element to be co-transformed into an organism. Where additional genes or elements are included, the components are operatively linked. The term “operatively linked” is intended to mean a functional link between two or more elements. For example, an operative link between a promoter and a coding region of interest (e.g., a region encoding a deaminase, RGN, and / or gRNA) is a functional link that allows expression of the coding region of interest. The operatively linked elements may be contiguous or non-contiguous. When used to refer to the link between two protein-coding regions, “operatively linked” means that the coding regions are located in the same reading frame. In some embodiments, additional genes or elements are provided on multiple expression cassettes. For example, the nucleotide sequence encoding the disclosed deaminase or fusion protein may be present on one expression cassette, while the nucleotide sequence encoding the gRNA may be on a separate expression cassette. Another example may have a first expression cassette encoding the disclosed deaminase separately, a second expression cassette encoding the fusion protein containing the deaminase, and a third expression cassette encoding the gRNA. Such expression cassettes provide multiple restriction sites and / or recombination sites for insertion of transcriptionally regulated polynucleotides in the regulatory region. Expression cassettes containing optional marker genes may also be present.

[0341] The expression cassette may include: a transcription (and in some embodiments, translation) initiation region (i.e., promoter) in the 5'-3' direction of transcription, a deaminase-encoded polynucleotide of the present invention, a fusion protein-encoded polynucleotide of the present invention, and a transcription (and in some embodiments, translation) termination region (i.e., termination region) functional in the organism of interest. Promoters used in the present invention are capable of directing or driving the expression of the coding sequence in the host cell. Regulatory regions (e.g., promoters, transcription regulatory regions, and translation termination regions) may be endogenous or heterologous to the host cell or to each other. As used herein, a “heterologous” sequence is a sequence derived from a foreign species, or, if from the same species, substantially modified from its natural form at constituent and / or genomic loci through intentional human intervention. As used herein, chimeric genes comprise coding sequences operatively linked to transcription initiation regions heterologous to the coding sequence.

[0342] Convenient termination regions can be obtained from Ti plasmids of *Agrobacterium tumefaciens*, such as termination regions for octopine synthase and carmine synthase. See also Guerineau et al. (1991) Mol. Gen. Genet. 262:141-144; Proudfoot (1991) Cell 64:671-674; Sanfacon et al. (1991) Genes Dev. 5:141-149; Mogen et al. (1990) Plant Cell 2:1261-1272; Munroe et al. (1990) Gene 91:151-158; Ballas et al. (1989) Nucleic Acids Res. 17:7891-7903; and Joshi et al. (1987) Nucleic Acids Res. 15:9627-9639.

[0343] Additional regulatory signals include, but are not limited to, transcription start sites, operators, activators, enhancers, other regulatory elements, ribosome binding sites, start codons, and termination signals. See, for example, U.S. Patent Nos. 5,039,523 and 4,853,331; EPO 0480762A2; Sambrook et al. (1992) Molecular Cloning: A Laboratory Manual, edited edition, Maniatis et al. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY), hereinafter referred to as "Sambrook 11"; Davis et al., edited edition, (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, NY, and the references cited therein.

[0344] In preparing expression cassettes, various DNA fragments can be manipulated to provide DNA sequences in the appropriate orientation and, under appropriate conditions, within the appropriate reading frame. For this purpose, adaptors or linkers can be used to ligate the DNA fragments, or other manipulations can be involved to provide convenient restriction sites, remove redundant DNA, and remove restriction sites, etc. This can involve in vitro mutagenesis, primer repair, restriction, annealing, and re-substitution, such as transformation and transversion.

[0345] Many promoters are available for use in the practice of this invention. Promoters can be selected based on the expected results. Nucleic acids can be combined with constitutive, inducible, growth phase-specific, cell type-specific, tissue-preferred, tissue-specific, or other promoters for expression in the organism of interest. See, for example, the promoters shown in WO 99 / 43838 and the following U.S. Patent Nos.: 8,575,425; 7790846; 8147856; 8.586832; 7772369; 7534939; 6072050; 5659026; 5608149; 5608144; 5604121; 5569597; 5466785; 5399680; 5268463; 5608142; and 6,177,611; these documents are incorporated herein by reference.

[0346] For expression in plants, constitutive promoters also include the CaMV 35S promoter (Odell et al. (1985) Nature 313:810-812); rice actin (McElroy et al. (1990) Plant Cell 2:163-171); ubiquitin (Christensen et al. (1989) Plant Mol. Biol. 12:619-632 and Christensen et al. (1992) Plant Mol. Biol. 18:675-689); pEMU (Last et al. (1991) Theor. Appl. Genet. 81:581-588); and MAS (Velten et al. (1984) EMBO J. 3:2723-2730).

[0347] Examples of inducible promoters are the Adh1 promoter, which is induced by hypoxia or cold stress; the Hsp70 promoter, which is induced by heat stress; the PPDK promoter, which is induced by light; and the phosphoenolpyruvate carboxylase promoter, both of which are induced by light. Also useful are chemically inducible promoters, such as the safener-inducible In2-2 promoter (US Patent No. 5,364,780), the auxin-inducible Axig1 promoter (which is tapetum-specific but also active in callus tissue (PCT US01 / 22169)), and steroid-responsive promoters (see, for example, the estrogen-inducible ERE promoter and glucocorticoid-inducible promoters in the following literature: Schena et al. (1991) Proc. Natl. Acad. Sci. USA 88:10421-10425 and McNellis et al. (1998) Plant J.14(2):247-257), and tetracycline-inducible and tetracycline-repressive promoters (see, for example, Gatz et al. (1991) Mol. Gen. Genet. 227:229-237, and U.S. Patent Nos. 5,814,618 and 5,789,156), which are incorporated herein by reference.

[0348] In some embodiments, tissue-specific or tissue-preferred promoters are used to target the expression of an expression construct within a specific tissue. In some embodiments, tissue-specific or tissue-preferred promoters are active in plant tissues. Examples of promoters under developmental control in plants include promoters that preferentially initiate transcription in certain tissues, such as leaves, roots, fruits, seeds, or flowers. A “tissue-specific” promoter is a promoter that initiates transcription only in certain tissues. Unlike constitutive gene expression, tissue-specific expression is the result of several levels of interaction of gene regulation. Therefore, promoters from homologous or closely related plant species can preferably be used to achieve efficient and reliable expression of transgenes in specific tissues. In some embodiments, expression includes tissue-preferred promoters. A “tissue-preferred” promoter is a promoter that preferentially initiates transcription in certain tissues, but not necessarily exclusively or completely in those tissues.

[0349] In some embodiments, the nucleic acid molecule encoding the deaminase or fusion protein described herein includes a cell-type-specific promoter. A “cell-type-specific” promoter is a promoter that is primarily driven to express in certain cell types of one or more organs. Some examples of plant cells include, for example, BETL cells, vascular cells in roots and leaves, stem cells, and stem cells, in which a cell-type-specific promoter that is functional in the plant may have predominant activity. The nucleic acid molecule may also include a cell-type-preferred promoter. A “cell-type-preferred” promoter is a promoter that is predominantly driven to express in most or all certain cell types of one or more organs, but not necessarily exclusively or completely in these tissues. Some examples of plant cells include, for example, BETL cells, vascular cells in roots and leaves, stem cells, and stem cells, in which a cell-type-preferred promoter that is functional in the plant may have preferential activity.

[0350] In some embodiments, nucleic acid sequences encoding deaminases, fusion proteins, and / or gRNAs are operatively linked to a promoter sequence recognized by a bacteriophage RNA polymerase, for example, for use in vitro mRNA synthesis. In such embodiments, the in vitro transcribed RNA can be purified for use in the methods described herein. For example, the promoter sequence can be a T7, T3, or SP6 promoter sequence, or a variation of the T7, T3, or SP6 promoter sequence. In such embodiments, the expressed protein and / or RNA can be purified for use in the genome modification methods described herein.

[0351] In some embodiments, the polynucleotide encoding the deaminase, fusion protein, and / or gRNA is linked to a polyadenylation signal (e.g., SV40 polyA signal and other signals functional in plants) and / or at least one transcription termination sequence. In some embodiments, the sequence encoding the deaminase or fusion protein is linked to a sequence encoding at least one nuclear localization signal, at least one cell-penetrating domain, and / or at least one signal peptide capable of transporting the protein to a specific subcellular location, as described elsewhere herein.

[0352] In some embodiments, the polynucleotide encoding deaminase, fusion protein, and / or gRNA is present in one or more vectors. “Vector” refers to a polynucleotide composition used to transfer, deliver, or introduce nucleic acids into a host cell. Suitable vectors include plasmid vectors, phage particles, granules, artificial / miniature chromosomes, transposons, and viral vectors (e.g., lentiviral vectors, adeno-associated virus vectors, baculovirus vectors). In some embodiments, the vector contains additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences), optional marker sequences (e.g., antibiotic resistance genes), replication origins, etc. Further information can be found in: “Current Protocols in Molecular Biology”, Ausubel et al., John Wiley & Sons, New York, 2003, or “Molecular Cloning: A Laboratory Manual”, Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001.

[0353] In some embodiments, the vector contains selectable marker genes for selecting transformed cells. Selectable marker genes are used to select transformed cells or tissues. Marker genes include genes encoding antibiotic resistance, such as those encoding neomycin phosphotransferase II (NEO) and hygromycin phosphotransferase (HPT), and genes conferring resistance to herbicide compounds such as glufosinate, bromobenzonitrile, imidazolinone, and 2,4-dichlorophenoxyacetate (2,4-D).

[0354] In some embodiments, the expression cassette or vector containing the sequence encoding the fusion protein further contains the sequence encoding gRNA. In some embodiments, the sequence encoding gRNA is operatively linked to at least one transcription control sequence for expression of gRNA in an organism of interest or host cell. For example, the polynucleotide encoding gRNA may be operatively linked to a promoter sequence recognized by RNA polymerase III (PolIII). Examples of suitable PolIII promoters include, but are not limited to, mammalian U6, U3, H1, and 7SL RNA promoters, rice U6 and U3 promoters, and promoters disclosed in PCT International Application No. PCT / US2022 / 032940, filed June 10, 2022, which is incorporated herein by reference in its entirety.

[0355] As indicated, expression constructs comprising nucleotide sequences encoding deaminases, fusion proteins, and / or gRNAs can be used to transform organisms of interest. Methods for transformation involve introducing the nucleotide construct into the organism of interest. The term "introduction" is defined as introducing the nucleotide construct into a host cell in a manner that allows the construct to enter the interior of the host cell. The methods of the present invention do not require a specific method for introducing the nucleotide construct into the host organism; they only require that the nucleotide construct can enter the interior of at least one cell of the host organism. In some embodiments, mRNA encoding a deaminase or fusion protein is introduced into the host cell. In some embodiments where the fusion protein comprises an RGN, mRNA encoding the fusion protein is introduced into the cell, and gRNA is introduced into the cell. The host cell can be a eukaryotic cell or a prokaryotic cell. In some embodiments, the eukaryotic host cell is a plant cell, a mammalian cell, or an insect cell. Methods for introducing the nucleotide construct into plant and other host cells are known in the art, including but not limited to stable transformation methods, transient transformation methods, and virus-mediated methods.

[0356] These methods result in transformed organisms, such as plants, including the whole plant, as well as plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, propagules, embryos, and their offspring. Plant cells can be differentiated or undifferentiated (e.g., callus, suspension culture cells, protoplasts, leaf cells, root cells, phloem cells, pollen).

[0357] "Transgenic organism," "transformed organism," or "stable transformed" refers to an organism, cell, or tissue that has incorporated or integrated polynucleotides encoding the deaminase or fusion protein of this invention. It should be recognized that other exogenous or endogenous nucleic acid sequences or DNA fragments can also be incorporated into host cells. Agrobacterium-mediated transformation and bio-mediated transformation remain the two main approaches used for plant cell transformation. However, host cell transformation can be achieved through infection, transfection, microinjection, electroporation, microinjection, gene gun or particle bombardment, electroporation, silica / carbon fiber, ultrasound-mediated techniques, PEG-mediated techniques, calcium phosphate coprecipitation, polycationic DMSO technology, DEAE dextran procedures, and virus-mediated and liposome-mediated procedures, among others. The introduction of virus-mediated polynucleotides encoding deaminases, fusion proteins, and / or gRNAs includes the introduction and expression mediated by retroviruses, lentiviruses, adenoviruses, and adeno-associated viruses, as well as the use of cauliflower mosaic virus groups (e.g., cauliflower mosaic virus), geminiviruses (e.g., soybean golden mosaic virus or maize stripe virus), and RNA plant viruses (e.g., tobacco mosaic virus).

[0358] Transformation protocols, and protocols for introducing polypeptide or polynucleotide sequences into plants, can vary depending on the type of host cell targeted for transformation (e.g., monocot or dicotyledonous plant cells). Methods for transformation are known in the art and include those shown in U.S. Patent Nos. 8,575,425; 7,692,068; 8,802,934; and 7,541,517, each of which is incorporated herein by reference. See also Rakoczy-Trojanowska, M. (2002) Cell Mol Biol Lett. 7:849-858; Jones et al. (2005) Plant Methods 1:5; Rivera et al. (2012) Physics of Life Reviews 9:308-345; Bartlett et al. (2008) Plant Methods 4:1-12; Bates, GW (1999) Methods in Molecular Biology 111:359-366; Binns and Thomashow (1988) Annual Reviews in Microbiology 42:575-606; Christou, P. (1992) The Plant Journal 2:275-281; ​​Christou, P. (1995) Euphatica 85:13-27; Tzfira et al. (2004) TRENDS in Genetics 20:375-383; Yao et al. (2006) Journal of Experimental Botany 57:3737-3746; Zupan and Zambryski (1995) Plant Physiology 107:1041-1047; Jones et al. (2005) Plant Methods 1:5;

[0359] Transformation can result in either stable or transient incorporation of nucleic acids into the cell. "Stable transformation" refers to a nucleotide construct introduced into the host cell that integrates into the host cell's genome and is inherited by its offspring. "Transient transformation" refers to a polynucleotide introduced into the host cell that does not integrate into the host cell's genome.

[0360] Methods for chloroplast transformation are known in the art. See, for example, Svab et al. (1990) Proc. Natl. Acad. Sci. USA 87:8526-8530; Svab and Maliga (1993) Proc. Natl. Acad. Sci. USA 90:913-917; Svab and Maliga (1993) EMBO J.12:601-606. These methods rely on particle gun delivery of DNA containing selectable markers and targeting the DNA to the plastid genome via homologous recombination. Alternatively, plastid transformation can be achieved by transactivation of plastid-carried transgenes through tissue-optimized expression of nuclear-encoded and plastid-directed RNA polymerases. Such systems have been reported in McBride et al. (1994) Proc. Natl. Acad. Sci. USA 91:7301-7305.

[0361] Transformed cells can be grown into transgenic organisms, such as plants, using conventional methods. See, for example, McCormick et al. (1986) Plant Cell Reports 5:81-84. These plants can then be grown and pollinated with the same or different transforming strains, and heterozygotes possessing deaminases or fusion protein polynucleotides can be identified. Two or more generations can be grown to ensure that the deaminases or fusion protein polynucleotides are stably maintained and inherited, and then the seeds are harvested to ensure the presence of the deaminases or fusion protein polynucleotides. In this way, the present invention provides transformed seeds (also referred to herein as “transgenic seeds”) having nucleotide constructs of the present invention stably incorporated into their genome, such as the expression cassette of the present invention.

[0362] In some embodiments, transformed cells are introduced into an organism. These cells may be derived from an organism, wherein the cells are transformed in an in vitro manner.

[0363] The sequences provided in this article can be used to transform any plant species, including but not limited to monocots and dicots. Examples of plants of interest include, but are not limited to, maize, sorghum, wheat, sunflower, tomato, cruciferous plants, pepper, potato, cotton, rice, soybean, sugar beet, sugarcane, tobacco, barley and rapeseed, Brassica, alfalfa, rye, millet, safflower, peanut, sweet potato, cassava, coffee, coconut, pineapple, citrus, cocoa, tea, banana, avocado, fig, guava, mango, olive, papaya, cashew, macadamia nut, almond, oat, vegetables, ornamental plants, and conifers.

[0364] Vegetables include, but are not limited to, tomatoes, lettuce, green beans, lima beans, peas, and members of the Cucumber genus, such as cucumbers, cantaloupes, and honeydew melons. Ornamental plants include, but are not limited to, azaleas, hydrangeas, hibiscus, roses, tulips, daffodils, petunias, carnations, poinsettias, and chrysanthemums. Preferably, the plants of this invention are crop plants (e.g., corn, sorghum, wheat, sunflowers, tomatoes, cruciferous plants, peppers, potatoes, cotton, rice, soybeans, sugar beets, sugarcane, tobacco, barley, rapeseed, etc.).

[0365] As used herein, the term plant includes plant cells, plant protoplasts, plant cell tissue cultures capable of regenerating plants, plant callus, plant masses, and intact plant cells in a plant or part of a plant (such as embryo, pollen, ovule, seed, leaf, flower, branch, fruit, kernel, spike, rachis, bark, stem, root, root tip, pollen sac, etc.). Cereal is intended to refer to mature seeds produced by commercial growers for purposes other than planting or propagating a species. Progeny, variants, and mutants of regenerated plants are also included within the scope of this invention, provided that these parts contain the introduced polynucleotides. Further, processed plant products or by-products retaining the sequences disclosed herein are provided, including, for example, soybean meal.

[0366] In some embodiments, the polynucleotide encoding deaminase, fusion protein, and / or gRNA is used to transform any eukaryotic species, including but not limited to animals (e.g., mammals, insects, fish, birds, and reptiles), fungi, amoebas, algae, and yeast. In some embodiments, the polynucleotide encoding deaminase, fusion protein, and / or gRNA is used to transform any prokaryotic species, including but not limited to archaea and bacteria (e.g., Bacillus, Klebsiella, Streptomyces, Rhizobium, Escherichia, Pseudomonas, Salmonella, Shigella, Vibrio, Yersinia, Mycoplasma, Agrobacterium, and Lactobacillus).

[0367] In some embodiments, conventional viral and nonviral gene transfer methods are used to introduce nucleic acids into mammalian cells or target tissues. Such methods can be used to administer nucleic acids encoding the deaminases or fusion proteins of the present invention, and optionally gRNA, to cells in culture or in a host organism. Nonviral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of vectors described herein), naked nucleic acids, and nucleic acids complexed with a delivery medium (e.g., liposomes). Viral vector delivery systems include DNA and RNA viruses that have a free or integrated genome after delivery to cells. Non-limiting examples include vectors utilizing the cauliflower mosaic virus group (e.g., cauliflower mosaic virus), geminiviruses (e.g., soybean golden mosaic virus or corn stripe virus), and RNA plant viruses (e.g., tobacco mosaic virus). For reviews of gene therapy procedures, see Anderson, Science 256:808-813 (1992); Nabel & Feigner, TIBTECH 11:211-217 (1993); Mitani & Caskey, TIBTECH 11:162-166 (1993); Dillon, TIBTECH 11:167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10):1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience 8:35-36 (1995); Kremer & Perricaudet, British Medical Bulletin 51(1):31-44 (1995); Haddadada et al., in Current Topics in Microbiology and Immunology, Doerfler and Bohm (edited edition) (1995); and Yu et al., Gene Therapy 1:13-26 (1994).

[0368] Non-viral methods for nucleic acid delivery include lipid transfection, Agrobacterium-mediated transformation, nuclear infection, microinjection, gene gun methods, virions, liposomes, immunoliposomes, multi-cationic or lipid-nucleic acid conjugates, naked DNA, artificial viral particles, and reagent-enhanced DNA uptake. Lipid transfection is described, for example, in U.S. Patent Nos. 5,049,386, 4,946,787; and 4,897,355, and lipid transfection reagents are commercially available (e.g., Transfectam). TM and Lipofectin TMCationic and neutral lipids suitable for effective receptor recognition of polynucleotides for lipid transfection include those mentioned in the following literature: Feigner, WO 91 / 17424; WO 91 / 16024. They can be delivered to cells (e.g., in vitro or ex vivo administration) or target tissues (e.g., in vitro administration). The preparation of lipid:nucleic acid complexes (including targeted liposomes, such as immunoliposome complexes) is well known to those skilled in the art (see, for example, Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Ther. 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad et al., Cancer Res. 52:4817-4820 (1992); U.S. Patent Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787.

[0369] The use of RNA or DNA virus-based systems for nucleic acid delivery leverages highly evolved processes to target viruses to specific cells in vivo and transport viral payloads to the cell nucleus. Viral vectors can be administered directly to patients (in vivo), or they can be used to process cells in vitro, with the modified cells then optionally administered to patients (ex vivo). Conventional virus-based systems can include retroviruses, lentiviruses, adenoviruses, adeno-associated viruses, and herpes simplex virus vectors for gene transfer. Gene transfer methods using retroviruses, lentiviruses, and adeno-associated viruses can integrate into the host genome, which typically leads to long-term expression of the inserted transgene. Furthermore, high transduction efficiency has been observed in many different cell types and target tissues.

[0370] The tropism of retroviruses can be altered by incorporating exogenous envelope proteins, thereby expanding the potential target population of target cells. Lentiviral vectors are retroviral vectors capable of transducing or infecting non-dividing cells and typically produce high viral titers. Therefore, the choice of retroviral gene transfer system may depend on the target tissue. Retroviral vectors contain cis-acting long terminal repeats (LTRs) with up to 6-10 kb of exogenous sequence packaging capability. A minimal cis-acting LTR is sufficient to replicate and package the vector, which is then used to integrate the desired gene into target cells to provide permanent transgenic expression. Widely used retroviral vectors include those based on murine leukemia virus (MuLV), gibberish leukemia virus (GaLV), simmon immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (see, for example, Buchscher et al., J.Virol. 66:2731-2739 (1992); Johann et al., J.Virol. 66:1635-1640 (1992); Sommnerfelt et al., Virol. 176:58-59 (1990); Wilson et al., J.Virol. 63:2374-2378 (1989); Miller et al., J.Virol. 65:2220-2224 (1991); PCT / US94 / 05700).

[0371] In applications where transient expression is preferred, adenovirus-based systems can be used. Adenovirus-based vectors exhibit very high transduction efficiency in many cell types and do not require cell division. High titers and high levels of expression have been obtained using such vectors. These vectors can be mass-produced in relatively simple systems. Adeno-associated virus (“AAV”) vectors can also be used to transduce cells with target nucleic acids, for example, for the in vitro production of nucleic acids and peptides, and for in vivo and in vitro gene therapy procedures (see, for example, West et al., Virology 160:38-47 (1987); US Patent No. 4,797,368; WO93 / 24641; Katin, Human Gene Therapy 5:793-801 (1994); Muzyczka, J. Clin. Invest. 94:1351 (1994). The construction of recombinant AAV vectors has been described in numerous publications, including US Patent No. 5,173,414; Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985); Tratschin et al., Mol. Cell. Biol. 4:2072-2081 (1984); Hermonat & Muzyczka, PNAS 81:6466-6470 (1984); and Samulski et al., J.Virol. 63:03822-3828 (1989). Packaging cells are commonly used to form viral particles capable of infecting host cells. Such cells include 293 cells for packaging adenoviruses, and ψJ2 or PA317 cells for packaging retroviruses.

[0372] Viral vectors used in gene therapy are typically produced by producer cell lines that package nucleic acid vectors into viral particles. The vectors usually contain the minimum viral sequences required for packaging and subsequent integration into the host; other viral sequences are replaced by expression cassettes of the polynucleotides to be expressed. Missing viral functions are typically provided trans-form by the packaging cell line. For example, AAV vectors used in gene therapy typically contain only the ITR sequences from the AAV genome, which are essential for packaging and integration into the host genome. The viral DNA is packaged in a cell line containing helper plasmids encoding other AAV genes, namely rep and cap, but lacking the ITR sequences.

[0373] This cell line can also be infected with adenovirus as a helper. The helper virus promotes the replication of the AAV vector and the expression of the AAV gene from the helper plasmid. Due to the lack of an ITR sequence, the helper plasmid is not packaged in large quantities. Adenovirus contamination can be reduced, for example, by heat treatment; adenovirus is more sensitive to heat treatment than AAV. Additional methods for delivering nucleic acids into cells are known to those skilled in the art. See, for example, US20030087817, which is incorporated herein by reference.

[0374] Ideally, the coding sequence of the fusion protein of the present invention and the corresponding guide RNA for targeting the fusion protein can be packaged entirely into a single AAV vector. The generally accepted size limit for AAV vectors is 4.7 kb, although larger sizes can be considered at the cost of reduced packaging efficiency. To ensure that the expression cassettes for both the fusion protein and its corresponding guide RNA can be packaged into the AAV vector, an active deletion variant of the RGN can be used. In addition to shortening the amino acid sequence, and thus the coding sequences of the RGN and / or deaminase of the fusion protein, the peptide linkers connecting the RGN and the deaminase can also be shortened or removed. Finally, genetic elements, such as promoters, enhancers, and / or terminators, can also be engineered via deletion analysis to determine the minimum size required for each element to be functional. The present invention also teaches a method for targeted base editing using the fusion protein via in vivo AAV vector delivery.

[0375] In some embodiments, host cells are transiently or non-transiently transfected using one or more vectors described herein. In some embodiments, cells are transfected because this occurs naturally in the subject. In some embodiments, the transfected cells are removed from the subject.

[0376] In some embodiments, the transfected cells are eukaryotic cells. In some embodiments, the eukaryotic cells are animal cells (e.g., mammals, insects, fish, birds, and reptiles). In some embodiments, the transfected cells are human cells. In some embodiments, the transfected cells are hematopoietic cells, such as immune cells (i.e., cells of the innate or adaptive immune system), including but not limited to B cells, T cells, natural killer (NK) cells, pluripotent stem cells, induced pluripotent stem cells, chimeric antigen receptor T (CAR-T) cells, monocytes, macrophages, and dendritic cells. The transfected cells can be allogeneic cells (e.g., allogeneic T cells) or autologous cells (e.g., autologous T cells).

[0377] In some embodiments, the cells are derived from cells taken from a subject, such as a cell line. In some embodiments, the cells or cell line are prokaryotes. In some embodiments, the cells or cell line are eukaryotes. In other embodiments, the cells or cell line are derived from insect, bird, plant, or fungal species. In some embodiments, the cells or cell line may be mammalian, such as, for example, human, monkey, mouse, cow, pig, goat, hamster, rat, cat, or dog. Various cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLaS3, Huhl, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Panel, PC-3, TFl, CTLL-2, CIR, Rat6, CVI, RPTE, AlO, T24, 182, A375, ARH-77, Calul, SW480, SW620, SKOV3, SK-UT, CaCo2, P388Dl, SEM-K2, WEHI-231, HB56, TIB55, lurkat, 145.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, Hep G2, HeLa B, and HeLa T4.COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelium, BALB / 3T3 mouse embryonic fibroblasts, 3T3-Swiss, 3T3-Ll, 132-d5 human fetal fibroblasts; 10.1 mouse fibroblasts, 293-T, 3T3, 721, 9L, A2780, A2780ADR, A2780cis, A172, A20, A253, A431, A-549, ALC, B16, B35, BCP-I cells, BEAS-2B, bEnd.3, BHK-21, BR 293. BxPC3, C3H-10Tl / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-Kl, CHO-K2, CHO-T, CHODhfr- / -, COR-L23, COR-L23 / CPR, COR-L235010, CORL23 / R23, COS-7, COV-434, CML Tl, CMT, CT26, D17, DH82, DU145, DuCaP, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA 2. HEK-293, HeLa, Hepalclc7, HL-60, HMEC, HT-29, lurkat, LY cells, K562 cells, Ku812, KCL22, KGl, KYOl, LNCap, Ma-Mel 1-48, MC-38, MCF-7, MCF-10A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCKII, MDCKII, MOR / 0.2R, MONO-MAC 6. MTD-1A, MyEnd, NCI-H69 / CPR, NCI-H69 / LX10, NCI-H69 / LX20, NCI-H69 / LX4, NIH-3T3, NALM-1, NW-145, OPCN / OPCT cell lines, Peer, PNT-1A / PNT 2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THP1 cell lines, U373, U87, U937, VcaP, Vero cells, WM39, WT-49, X63, YAC-1, YAR, and their transgenic variants. Cell lines may be obtained from a variety of sources known to those skilled in the art (see, for example, the United States Type Culture Collection (ATCC) (Manassas, Virginia)).

[0378] In some embodiments, cells transfected with one or more vectors described herein are used to establish new cell lines comprising one or more vector-derived sequences. In some embodiments, cells transiently transfected with the fusion protein and optionally gRNA or the ribonucleoprotein complex of the present invention, and cells whose activity is modified by the fusion protein or ribonucleoprotein complex, are used to establish new cell lines comprising cells containing the modification but lacking any other exogenous sequences. In some embodiments, cells transiently or non-transiently transfected with one or more vectors described herein, or cell lines derived from such cells, are used to evaluate one or more test compounds.

[0379] In some embodiments, one or more vectors described herein are used to produce non-human transgenic animals or transgenic plants. In some embodiments, the transgenic animal is an insect. In further embodiments, the insect is a pest, such as a mosquito or tick. In some embodiments, the insect is a plant pest, such as a corn rootworm or fall armyworm. In some embodiments, the transgenic animal is a bird, such as a chicken, turkey, goose, or duck. In some embodiments, the transgenic animal is a mammal, such as a human, mouse, rat, hamster, monkey, ape, rabbit, pig, cow, horse, goat, sheep, cat, or dog.

[0380] VII. Variations and fragments of polypeptides and polynucleotides

[0381] This disclosure provides deaminases and fusion proteins comprising the same, as well as fusion proteins comprising an RGN active against a DNA molecule and a heterologous polypeptide (e.g., a deaminase). The RGN may comprise any one of SEQ ID NO:1-4, 49-162, 435, 575, 576, 698, and 699, or an active variant or fragment thereof, and a polynucleotide encoding them. The deaminase may comprise any one of SEQ ID NO:6, 229-414, 596, 597, 720, and 723, or an active variant or fragment thereof, and a polynucleotide encoding them.

[0382] While the activity of variants or fragments may change compared to the polynucleotide or peptide of interest, variants and fragments should retain the functionality of the polynucleotide or peptide of interest. For example, variants or fragments may have increased activity, decreased activity, a different activity profile, or any other change in activity when compared to the polynucleotide or peptide of interest.

[0383] If the fragments and variants of the deaminases of the present invention having adenine deaminase activity are part of a fusion protein further comprising a DNA-binding polypeptide or a fragment thereof, they will retain said activity.

[0384] Fragments and variants of the fusion proteins of the present invention having base editing activity, such as fusion protein fragments and variants containing fragments or variants of deaminases or fusion protein fragments and variants containing DNA-binding polypeptides (e.g., RGN), will retain said activity.

[0385] The term "fragment" refers to a portion of the polynucleotide or polypeptide sequence of the present invention. A "fragment" or "bioactive portion" includes a polynucleotide containing a sufficient number of consecutive nucleotides to retain biological activity. A "fragment" or "bioactive portion" includes a polypeptide containing a sufficient number of consecutive amino acid residues to retain biological activity. The fragments of the RGN or deaminase disclosed herein include fragments shorter than the full-length sequence due to the use of an alternative downstream start site. In some embodiments, the bioactive portion of a deaminase or RGN is a polypeptide containing, for example, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160 or more consecutive amino acid residues or variations thereof, of any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698, and 699. In some embodiments, the bioactive portion of the deaminase or RGN is a polypeptide comprising, for example, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160 or more consecutive amino acid residues or variants thereof, of any one of SEQ ID NO: 5, 6, 229-414, 596, 597 and 720-723. Such bioactive portions can be prepared using recombinant techniques, and their activity can be evaluated.

[0386] Generally, "variant" is intended to refer to substantially similar sequences. For polynucleotides, a variant comprises the deletion and / or addition of one or more nucleotides at one or more internal sites within a natural polynucleotide and / or the substitution of one or more nucleotides at one or more sites within a natural polynucleotide. As used herein, "natural" or "wild-type" polynucleotides or polypeptides comprise naturally occurring nucleotide or amino acid sequences, respectively. For polynucleotides, conserved variants include those sequences that encode the natural amino acid sequence of the gene of interest due to the degeneracy of the genetic code. Naturally occurring allelic variants such as these can be identified using well-known molecular biology techniques, such as, for example, polymerase chain reaction (PCR) and hybridization techniques as outlined below. Variant polynucleotides also include synthetically derived polynucleotides, such as those generated, for example, by site-directed mutagenesis but still encoding the polypeptide or polynucleotide of interest. Variants of a particular polynucleotide disclosed herein will have approximately 40% to approximately 99% or higher sequence identity with variants of a particular polynucleotide identified by sequence alignment procedures and parameters as described elsewhere herein. Typically, the specific polynucleotide variants disclosed herein will have at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or higher sequence identity with the specific polynucleotide variants identified by the sequence alignment procedures and parameters described elsewhere herein.

[0387] Variants of the specific polynucleotides disclosed herein (i.e., reference polynucleotides) can also be assessed by comparing the percentage of sequence identity between the polypeptide encoded by the variant polynucleotide and the polypeptide encoded by the reference polynucleotide. The percentage of sequence identity between any two polypeptides can be calculated using the sequence alignment procedures and parameters described elsewhere herein. When assessing them by comparing the percentage of sequence identity shared by two polypeptides encoded by any given pair of polynucleotides disclosed herein, the percentage of sequence identity between the two encoded polypeptides is about 40% to about 99% or higher, and in some embodiments, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or higher.

[0388] In some embodiments, the polynucleotide of this disclosure encodes a deaminase comprising an amino acid sequence having at least about 40% to about 99% or higher identity with the amino acid sequence of any one of SEQ ID NO: 5, 6, 300-414, 596-598 and 720-723, wherein the deaminase comprises at least one of the amino acid residues shown in any one of Tables 2, 4, 6 and 19. In some embodiments, the polynucleotide encoding the present disclosure comprises a deaminase having an amino acid sequence having at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or higher identity with the amino acid sequence of any one of SEQ ID NO: 5, 6, 300-414, 596-598, and 720-723, wherein the deaminase comprises at least one of the amino acid residues shown in any one of Tables 2, 4, 6, and 19.

[0389] The bioactive variants of the adenine deaminase of the present invention may differ by as few as 1-15 amino acid residues, as few as 1-10 (e.g., 6-10), as few as 5, as few as 4, as few as 3, as few as 2, or as few as 1 amino acid residue. In certain embodiments, the polypeptide comprises an N-terminal or C-terminal truncation, which may contain at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, or more amino acid deletions from the N- or C-terminus of the polypeptide. In some embodiments, the polypeptide comprises an internal deletion, which may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, or more amino acid deletions.

[0390] It should be recognized that the RGN and deaminases provided herein can be modified to produce variant proteins and polynucleotides. Human-designed modifications can be introduced by applying site-directed mutagenesis techniques. In some embodiments, naturally occurring, previously unknown, or previously unidentified polynucleotides and / or polypeptides that fall within the scope of the invention and are structurally and / or functionally related to the sequences disclosed herein can also be identified. Conserved amino acid substitutions can be made in non-conserved regions of the polypeptide without altering its function as an adenine deaminase.

[0391] Variant polynucleotides and proteins also encompass sequences and proteins derived from mutagenesis and recombination procedures, such as DNA shuffling. Through such procedures, one or more different RGNs or deaminases are manipulated to produce novel deaminases with desired properties. In this way, libraries of recombinant polynucleotides are generated from a population of related sequence polynucleotides containing sequence regions that have basic sequence identity and can be homologously recombinated in vitro or in vivo. For example, using this protocol, sequence motifs encoding the domain of interest can be shuffled between the sequences provided herein and other subsequently identified genes to obtain genes encoding improved properties of interest (such as increased K in the case of enzymes). m Novel genes for proteins of this type. Strategies for such DNA shuffling are known in the art. See, for example, Stemmer (1994) Proc. Natl. Acad. Sci. USA 91:10747-10751; Stemmer (1994) Nature 370:389-391; Crameri et al. (1997) Nature Biotech. 15:436-438; Moore et al. (1997) J. Mol. Biol. 272:336-347; Zhang et al. (1997) Proc. Natl. Acad. Sci. USA 94:4504-4509; Crameri et al. (1998) Nature 391:288-291; and U.S. Patent Nos. 5,605,793 and 5,837,458. "Recombined" nucleic acids are nucleic acids generated by a recombining procedure (such as any recombining procedure shown herein). Recombined nucleic acids are generated by (physically or virtually) recombining two or more nucleic acids (or strings), for example, manually and optionally recursively. Typically, one or more screening steps are used in the recombining procedure to identify nucleic acids of interest; such screening steps may be performed before or after any recombination step. In some (but not all) recombining embodiments, it is desirable to perform multiple rounds of recombination before selection to increase the diversity of the pool to be screened. Optionally, the entire process of recombination and selection is repeated recursively. Depending on the context, recombining may refer to the entire process of recombination and selection, or alternatively, it may simply refer to the recombination portion of the entire process.

[0392] As used herein, in the context of two polynucleotide or polypeptide sequences, “sequence identity” or “identity” refers to residues in two sequences that are identical when a maximum correspondence alignment is performed on a particular comparison window. It should be recognized that dissimilar residue positions are often due to conserved amino acid substitutions, where an amino acid residue is replaced by another amino acid residue having similar chemical properties (e.g., charge or hydrophobicity) and therefore does not alter the functional properties of the molecule. Protein sequences that are different due to such conserved substitutions are referred to as having “sequence similarity” or “identity.” The means used to measure sequence similarity are well known to those skilled in the art. Typically, this involves scoring conserved substitutions as partial rather than complete mismatches. Thus, for example, where identical amino acids are given a score of 1 and non-conserved substitutions are given a score of zero, conserved substitutions are given a score between zero and 1. The score for conserved substitutions is calculated, for example, as performed in the program PC / GENE (Intelligenetics, Mountain View, California).

[0393] As used herein, “percentage of sequence identity” refers to a value determined by comparing two best-aligned sequences in a comparison window, where the portion of the polynucleotide sequence in the comparison window may include additions or deletions (i.e., gaps) compared to a reference sequence (excluding additions or deletions) to achieve best alignment between the two sequences. The percentage is calculated by determining the number of positions in both sequences where the same nucleic acid base or amino acid residue appears to obtain the number of matching positions, dividing the number of matching positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity.

[0394] Unless otherwise stated, the sequence identity / similarity values ​​provided herein refer to values ​​obtained using GAP version 10 with the following parameters: % identity and % similarity of nucleotide sequences using GAP weights of 50 and length weights of 3, and the nwsgapdna.cmp scoring matrix; % identity and % similarity of amino acid sequences using GAP weights of 8 and length weights of 2, and the BLOSUM62 scoring matrix; or any equivalent procedure thereof. An “equivalent procedure” is intended to be any sequence comparison procedure that, for any two sequences discussed, generates alignments with the same nucleotide or amino acid residue matches and the same percentage of sequence identity when compared to corresponding alignments generated by GAP version 10.

[0395] When aligning two sequences using a defined amino acid substitution matrix (e.g., BLOSUM62), a vacancy presence penalty, and a vacancy expansion penalty, the two sequences are considered "best aligned" to obtain the highest possible score for that pair. Amino acid substitution matrices and their use in quantifying the similarity between two sequences are well-known in the art and are described, for example, in Dayhoff et al. (1978), "A model of evolutionary change in proteins." They are also described in "Atlas of Protein Sequence and Structure," Vol. 5, Supplement 3 (edited by MO Dayhoff), pp. 345-352, Natl. Biomedical Res. Found., Washington, DC, and Henikoff et al. (1992), Proc. Natl. Acad. Sci. USA 89:10915-10919. The BLOSUM62 matrix is ​​commonly used as the default scoring substitution matrix in sequence alignment protocols. A penalty is applied to the introduction of a single amino acid vacancy in one of the sequences, and a penalty is applied to the vacancy expansion for each additional empty amino acid position inserted into an already opened vacancy. Alignment is defined by the amino acid positions of each sequence at the start and end of the alignment, and optionally by inserting vacancies or multiple vacancies in one or both sequences to obtain the highest possible score. While optimal alignment and scoring can be performed manually, the process is facilitated by the use of computer-implemented alignment algorithms, such as BLAST 2.0, described in Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402, and is publicly available on the website of the National Center for Biotechnology Information (www.ncbi.nlm.nih.gov). For example, the best alignment (including multiple alignments) can be prepared using PSI-BLAST, which is available at www.ncbi.nlm.nih.gov and described in the following literature: Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402.

[0396] Regarding the amino acid sequence that best aligns with the reference sequence, the amino acid residues "correspond" to their paired positions in the reference sequence. These "positions" are indicated by numbers that sequentially identify each amino acid in the reference sequence based on its position relative to the N-terminus. Because deletions, insertions, truncations, fusions, etc., must be considered when determining the optimal alignment, the numbers of amino acid residues in the test sequence, typically determined by counting only from the N-terminus, may not be the same as their corresponding positions in the reference sequence. For example, in the case of a deletion in the test sequence, there will be no amino acid corresponding to the deletion site in the reference sequence. In the case of an insertion in the reference sequence, the insertion will not correspond to any amino acid position in the reference sequence. In the case of truncation or fusion, there may be amino acid segments in the reference or aligned sequences that do not correspond to any amino acid in the corresponding sequence.

[0397] VIII. Antibody

[0398] This also covers antibodies against the deaminases, fusion proteins, or ribonucleoproteins of the present invention, including deaminases or their active variants or fragments having the amino acid sequences shown as any one of SEQ ID NO: 6, 300-414, 596-598, and 720-723. Methods for generating antibodies are well known in the art (see, for example, Harlow and Lane (1988) Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY; and U.S. Patent No. 4,196,265). These antibodies can be used in kits for the detection and isolation of the deaminases, fusion proteins, or ribonucleoproteins described herein. Therefore, this disclosure provides kits comprising antibodies that specifically bind to the polypeptides or ribonucleoproteins described herein, including, for example, polypeptides comprising a sequence having at least 85% identity with any one of SEQ ID NO: 5, 6, 300-414, 596-598 and 720-723, wherein the deaminase comprises at least one of the amino acid residues shown in any one of Tables 2, 4, 6 and 19.

[0399] IX. Systems and ribonucleoprotein complexes for binding and / or modifying target sequences of interest, and methods for their manufacture.

[0400] This disclosure provides a system for targeting and modifying a nucleic acid sequence. In some embodiments, the system targets and modifies a nucleic acid sequence with a heterologous polypeptide. In those embodiments where the heterologous polypeptide comprises a base-editing polypeptide (e.g., a deaminase) or a lead-editing polypeptide, the base-editing polypeptide (e.g., a deaminase) or lead-editing polypeptide fused with an RGN (e.g., an RGN nickase or an inactive RGN nuclease) is responsible for modifying the targeted nucleic acid sequence. A guide RNA hybridizes with the target sequence of interest and also forms a complex with the RGN component of the fusion protein, thereby guiding the fusion protein to bind to the target sequence.

[0401] In some embodiments, an RNA-guided DNA-binding polypeptide, such as an RGN (e.g., an RGN nickase or nuclease-active RGN), and gRNA are responsible for targeting the ribonucleoprotein complex to the nucleic acid sequence of interest, and a deaminase polypeptide fused with RDBBP is responsible for modifying the targeting nucleic acid sequence. In those embodiments where the deaminase is an adenine deaminase, a base editor modifies A>N. In some embodiments, the adenine deaminase converts A>G. The guiding RNA hybridizes with the target sequence of interest and also forms a complex with the RNA-guided DNA-binding polypeptide, thereby directing the RNA-guided DNA-binding polypeptide to bind to the target sequence. The RNA-guided DNA-binding polypeptide is part of a fusion protein that also contains a deaminase, such as one described herein. In some embodiments, the RNA-guided DNA-binding polypeptide is an RGN, such as Cas9. Other examples of RNA-guided DNA-binding polypeptides include RGNs, such as those described in International Patent Application Publication Nos. WO 2019 / 236566 and WO 2020 / 139783, each of which is incorporated herein by reference in its entirety. In some embodiments, the RNA-guided DNA-binding polypeptide is a type II CRISPR-Cas polypeptide or its active variant or fragment. In some embodiments, the RNA-guided DNA-binding polypeptide is a type V CRISPR-Cas polypeptide or its active variant or fragment. In some embodiments, the RNA-guided DNA-binding polypeptide is a type VI CRISPR-Cas polypeptide. In some embodiments, the DNA-binding polypeptide of the fusion protein does not require an RNA guide, such as a zinc finger nuclease, TALEN, or meganuclease polypeptide. In some embodiments, the nuclease activity of the DNA-binding polypeptide is partially or completely inactivated. In a further embodiment, the RNA-guided DNA-binding polypeptide comprises the amino acid sequence of RGN, such as, for example, APG07433.1 (SEQ ID NO:1) or its active variant or fragment, such as the nickase nAPG07433.1 (SEQ ID NO:49).

[0402] In some embodiments, the system provided herein for binding and modifying a target sequence of interest is a ribonucleoprotein complex, which is at least one molecule of RNA that binds to at least one protein (i.e., a fusion protein). The ribonucleoprotein complex provided herein comprises at least one guide RNA as an RNA component and a fusion protein comprising an RNA-guided DNA-binding polypeptide (e.g., RGN) and an inserted heterologous polypeptide (e.g., a deaminase) as a protein component. In some embodiments, the ribonucleoprotein complex is purified from cells or organisms transformed with polynucleotides encoding the fusion protein and the guide RNA, and cultured under conditions allowing the expression of the fusion protein and the guide RNA.

[0403] Methods for manufacturing deaminases, fusion proteins, or fusion protein-ribonucleoprotein complexes are provided. Such methods include culturing cells containing nucleotide sequences encoding the deaminase, the fusion protein, and, in some embodiments, the guide RNA, under conditions in which the deaminase or fusion protein (and, in some embodiments, guide RNA) is expressed. The deaminase, fusion protein, or fusion ribonucleoprotein can then be purified from the lysate of the cultured cells.

[0404] Methods for purifying deaminases, fusion proteins, or fusion ribonucleoprotein complexes from lysates of biological samples are known in the art (e.g., size exclusion and / or affinity chromatography, 2D-PAGE, HPLC, reversed-phase chromatography, immunoprecipitation). In certain methods, the deaminases or fusion proteins are recombinantly generated and contain purification tags that facilitate their purification, including but not limited to glutathione S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tags, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, HA, nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, 6xHis, biotinylate carboxyl carrier protein (BCCP), and calmodulin. Typically, immunoprecipitation or other similar methods known in the art are used to purify the labeled deaminase, fusion protein, or fusion ribonucleoprotein complex.

[0405] "Isolated" or "purified" polypeptides or their bioactive portions are substantially or substantially free of components that typically accompany or interact with the polypeptide, such as those found in the environment in which the polypeptide naturally exists. Therefore, isolated or purified polypeptides are substantially free of other cellular material or culture media when produced by recombinant technology, or substantially free of chemical precursors or other chemicals when chemically synthesized. Proteins substantially free of cellular material include preparations of proteins having less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% (by dry weight) of contaminating proteins. When recombinantly producing the proteins of the present invention or their bioactive portions, the optimal culture medium represents less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% (by dry weight) of chemical precursors or non-protein-of-interest chemicals.

[0406] The specific methods provided herein for binding and / or cleaving target sequences of interest involve the use of in vitro assembled ribonucleoprotein complexes. In vitro assembly of ribonucleoprotein complexes can be performed using any method known in the art, wherein a fusion protein containing RDBBP (e.g., RGN) is contacted with the guide RNA under conditions that allow the RDBBP (e.g., RGN) fusion protein to bind to the guide RNA. As used herein, “contact,” “contacting,” and “contacted” mean placing the components of the desired reaction together under conditions suitable for carrying out the desired reaction. RDBBP (e.g., RGN) fusion proteins can be purified from biological samples, cell lysates, or culture media, generated via in vitro translation, or chemically synthesized. Guide RNA can be purified from biological samples, cell lysates, or culture media, transcribed in vivo, or chemically synthesized. RDBBP (e.g., RGN) fusion proteins and guide RNA can be contacted in solution (e.g., buffered saline solution) to allow for the in vitro assembly of the ribonucleoprotein complex.

[0407] X. Methods for localizing or modifying heterologous peptides to target DNA molecules.

[0408] This disclosure provides methods for targeting heterologous polypeptides (e.g., deaminases) to or modifying target nucleic acid molecules of interest (e.g., target DNA molecules). The method includes delivering a fusion protein of this disclosure or a polynucleotide encoding the same to a target sequence or a cell, organelle, or embryo containing the target sequence. In some embodiments, the method includes delivering a system comprising at least one guide RNA or a polynucleotide encoding the same, and at least one fusion protein of this disclosure, to a target sequence or a cell, organelle, or embryo containing the target sequence.

[0409] In some embodiments, the method includes contacting a DNA molecule with: (a) a fusion protein; and (b) gRNA that targets the fusion protein of (a) to a target nucleotide sequence of the DNA molecule; wherein the DNA molecule is contacted with the fusion protein and gRNA in an effective amount and under conditions suitable for the binding of the RGBDP (e.g., RGN) component to the target sequence.

[0410] The target DNA molecule may contain a sequence associated with a disease or symptom, and the editing of at least one nucleotide within the causal mutation may result in a sequence not associated with the disease or symptom. In some embodiments, the disease or symptom affects an animal. In a further embodiment, the disease or symptom affects a mammal, such as a human, cattle, horse, dog, cat, goat, sheep, pig, monkey, rat, mouse, or hamster. In some embodiments, the target DNA sequence is present in the alleles of a crop plant, wherein a specific allele of the trait of interest results in a plant with lower agronomic value. Editing of at least one nucleotide within this region results in an allele that improves the trait and increases the agronomic value of the plant.

[0411] Following delivery of the polynucleotide encoding the guide RNA and / or fusion protein, the cells or embryo can then be cultured under conditions in which the guide RNA and / or fusion protein are expressed. In various embodiments, the method includes contacting the target sequence with a ribonucleoprotein complex comprising the gRNA and the fusion protein. In some embodiments, the method includes introducing the ribonucleoprotein complex of the present invention into cells, organelles, or embryos containing the target sequence. The ribonucleoprotein complex of the present invention can be a ribonucleoprotein complex purified from a biological sample, recombinantly generated and subsequently purified, or assembled in vitro as described herein. In those embodiments where the ribonucleoprotein complex contacted with the target sequence or cell organelle or embryo has been assembled in vitro, the method may further include assembling the complex in vitro prior to contact with the target sequence, cell, organelle, or embryo.

[0412] The purified or in vitro assembled ribonucleoprotein complexes of the present invention can be introduced into cells, organelles, or embryos using any method known in the art (including, but not limited to, electroporation). In some embodiments, fusion proteins (or polynucleotides encoding them) and polynucleotides encoding or containing guide RNA are introduced into cells, organelles, or embryos using any method known in the art (e.g., electroporation).

[0413] Upon delivery to or contact with the target sequence or cells, organelles, or embryos containing the target sequence, the guide RNA directs the fusion protein to bind to the target sequence in a sequence-specific manner. Subsequently, in cases where the fusion protein contains a base-editing polypeptide (e.g., a deaminase) or a lead editing polypeptide, the target sequence can be modified.

[0414] In some embodiments where the fusion protein is a base editor (i.e., containing a base-editing polypeptide, such as a deaminase), the binding of the fusion protein to the target sequence results in the modification of nucleotides adjacent to the target sequence. The nucleotides adjacent to the target sequence modified by the deaminase can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 base pairs from the 5' or 3' end of the target sequence. The full range of nucleotides that can be edited using a base editor fusion protein (e.g., an RGN fused with a deaminase) is typically represented as the distance of nucleotides from the PAM sequence (e.g., nucleotides at positions 8 and 22 upstream (i.e., 5') of the PAM sequence), referred herein to as the “edit window” of a particular base editor fusion protein. The editing window can be, for example, 1-100 base pairs (5' or 3') of a PAM sequence, including but not limited to 5-50, 5-25, 8-22, 10-20, or 10-21 base pairs (5' or 3') of a PAM sequence. Fusion proteins comprising adenine deaminases (such as those disclosed herein) and RNA-guided DNA-binding peptides (e.g., RGN) can introduce targeted A>N mutations into targeted DNA molecules. In some embodiments, the fusion protein introduces targeted A>G mutations into targeted DNA molecules.

[0415] Methods for measuring the binding of fusion proteins to target sequences are known in the art and include chromatin immunoprecipitation assays, gel migration variation assays, DNA pull-down assays, reporter gene assays, and microplate capture and detection assays. Similarly, methods for measuring the cleavage or modification of target sequences are known in the art and include in vitro or in vivo cleavage assays, wherein cleavage is confirmed using PCR, sequencing, or gel electrophoresis, with or without the attachment of appropriate labels (e.g., radioisotopes, fluorescent substances) to the target sequence to facilitate the detection of degradation products. In some embodiments, nick-triggered exponential amplification reaction (NTEXPAR) assays are used (see, for example, Zhang et al. (2016) Chem. Sci. 7:4951-4957). In vivo cleavage can be assessed using a Surveyor assay (Guschin et al. (2010) Methods Mol Biol 649:247-256).

[0416] Methods for measuring the base editing activity of a base editor fusion protein or a lead editor fusion protein are known in the art and include those described herein, as well as any form of sequence analysis of the target nucleic acid molecule after contact with the base editor or lead editor (e.g., PCR, sequencing, or gel electrophoresis with or without appropriate tagging). Base editor fusion proteins containing the deaminases disclosed herein may exhibit improved editing activity (i.e., improved editing efficiency or reduced RNA editing), a larger editing window, or both, compared to similar base editor fusion proteins containing the parental LPG50148 deaminase. Compared to a second base editor fusion protein containing the parental LPG50148 deaminase and the same DNA-binding polypeptide (e.g., RGN) as the first base editor fusion protein, base editor fusion proteins containing the deaminases of the present invention (e.g., SEQ ID NO) may exhibit improved editing activity (e.g., improved editing efficiency or reduced RNA editing), a larger editing window, or both. The base editing efficiency of the first base editor fusion protein (NO: 5, 6, 300-414, 596-598, and 720-723, or an active variant or fragment containing at least one of the amino acid residues shown in any of Tables 2, 4, 6, and 19) can be 1.1 to 10, 1.1 to 20, 1.1 to 30, or greater, including but not limited to about 1.1, about 1.5, about 2, about 2.5, about 3, and about 3.5 times, about 4 times, about 4.5 times, about 5 times, about 5.5 times, about 6 times, about 6.5 times, about 7 times, about 7.5 times, about 8 times, about 8.5 times, about 9 times, about 9.5 times, about 10 times, about 11 times, about 12 times, about 13 times, about 15 times, about 16 times, about 17 times, about 18 times, about 19 times, about 20 times, about 21 times, about 22 times, about 23 times, about 24 times, about 25 times, about 26 times, about 27 times, about 28 times, about 29 times, and about 30 times. When compared to similar base editor fusion proteins containing the parental LPG50148 deaminase, base editor fusion proteins containing the deaminase of the present invention can exhibit reduced RNA editing. Compared to a second base editor fusion protein containing the parental LPG50148 deaminase and the same DNA-binding polypeptide (e.g., RGN) as the first base editor fusion protein, a first base editor fusion protein containing the deaminase of the present invention can exhibit an RNA editing rate of about 1% to about 99% of that of the second base editor fusion protein, including but not limited to about 1%, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, and about 99%.

[0417] Compared to the editing window of a second base editor fusion protein containing the parental LPG50148 deaminase and the same DNA-binding polypeptide (e.g., RGN) as the first base editor fusion protein, the editing window of a first base editor fusion protein containing the deaminase of the present invention (e.g., SEQ ID NO: 5, 6, 300-414, 596-598, and 720-723, or containing at least one of the active variants or fragments of the amino acid residues shown in any of Tables 2, 4, 6, and 19) may have one nucleotide on either side or both sides, including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides on either side or both sides of the parental base editor editing window.

[0418] In some embodiments, the method involves the use of a fusion protein, wherein RDBBP (e.g., RGN) is compounded with more than one guide RNA. The more than one guide RNA may target different regions of a single gene, or may target multiple genes. This multiple targeting enables the deaminase or lead editing peptide of the fusion protein to modify nucleic acids, thereby introducing multiple mutations in the target nucleic acid molecule of interest (e.g., the genome).

[0419] In embodiments where the method involves the use of RNA-guided nucleases (RGNs) (such as nickase RGNs (i.e., those that can only cleave single strands of double-stranded polynucleotides, e.g., nAPG07433.1 (SEQ ID NO:49))), the method may include introducing two different RGNs or RGN variants that target the same or overlapping target sequences and cleave different strands of the polynucleotide. For example, an RGN nickase that cleaves only the positive (+) strand of a double-stranded polynucleotide and a second RGN nickase that cleaves only the negative (-) strand of a double-stranded polynucleotide may be introduced. In some embodiments, two different fusion proteins are provided, each containing a different RGN with a different PAM recognition sequence, allowing for greater diversity of nucleotide sequences to be targeted for mutation.

[0420] Those skilled in the art will understand that any of the methods disclosed herein can be used to target a single target sequence or multiple target sequences. Therefore, the methods include using a fusion protein comprising a single RDBBP (e.g., RGN) in combination with a variety of different guide RNAs, which can target multiple different sequences within a single gene and / or multiple genes. This document also covers methods for introducing combinations of multiple different guide RNAs with multiple different RGN fusion proteins. These guide RNAs and guide RNA / fusion protein systems can target multiple different sequences within a single gene and / or multiple genes.

[0421] In some embodiments, fusion proteins comprising RNA-guided DNA-binding peptides (e.g., RGN) and heterologous peptides (e.g., deaminases) can be used to generate mutations in a targeted gene or region of interest. In some embodiments, the fusion proteins of the present invention can be used for saturation mutagenesis of a targeted gene or region of interest, followed by high-throughput forward genetic screening to identify novel mutations and / or phenotypes. In some embodiments, the fusion proteins described herein can be used to generate mutations at targeted genomic locations, which may or may not contain a coding DNA sequence. Libraries of cell lines generated through the targeted mutagenesis described above can also be used to study gene function or gene expression.

[0422] XI. Target polynucleotides

[0423] In one aspect, the present invention provides a method for modifying target polynucleotides in eukaryotic cells, which may be in vivo, in vitro, or ex vivo. In some embodiments, the method includes sampling cells or cell populations from humans or non-human animals or plants (including microalgae), and modifying the cells or multiple cells. Culture can be performed in vitro at any stage. Cells or multiple cells can even be reintroduced into humans, non-human animals, or plants (including microalgae).

[0424] Utilizing natural variability, plant breeders combine the most useful genes to obtain desired qualities such as yield, quality, uniformity, hardiness, and pest resistance. These desired qualities also include growth, day length preference, temperature requirements, onset date of flowering or reproductive development, fatty acid content, insect resistance, disease resistance, nematode resistance, fungal resistance, herbicide resistance, and tolerance to various environmental factors, including drought, high temperature, humidity, cold, wind, and adverse soil conditions including high salinity. The sources of these useful genes include natural or exotic species, traditional species, wild plant relatives, and induced mutations, such as treatment of plant material with mutagens. Using this invention, plant breeders are provided with a new tool for inducing mutations. Therefore, those skilled in the art can employ this invention to induce an increase in useful genes in a more precise manner than with previous mutagens, thereby accelerating and improving plant breeding programs.

[0425] The target nucleotide of the fusion protein of the present invention can be any polynucleotide that is endogenous or exogenous to eukaryotic cells. For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. In some embodiments, the target polynucleotide is a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). In some embodiments, the target sequence of the fusion protein of the present invention is associated with a PAM (pre-spacer adjacent motif); that is, a short sequence recognized by an RNA-guided DNA-binding polypeptide (e.g., RGN). The precise sequence and length requirements of the PAM vary depending on the RNA-guided DNA-binding polypeptide used, but the PAM is typically a 2-5 base pair sequence adjacent to the pre-spacer (i.e., the target sequence).

[0426] The target polynucleotides of the fusion protein of the present invention may include a variety of disease-related genes and polynucleotides, as well as genes and polynucleotides related to signal transduction biochemical pathways. Examples of target polynucleotides include sequences associated with signal transduction biochemical pathways, such as genes or polynucleotides related to signal transduction biochemical pathways. Examples of target polynucleotides include disease-related genes or polynucleotides. A “disease-related” gene or polynucleotide, compared to a non-disease control tissue or cell, refers to any gene or polynucleotide that produces an abnormal level or abnormal form of its transcriptional or translational product in cells derived from disease-affected tissues. It may be a gene expressed at an abnormally high level; it may be a gene expressed at an abnormally low level, wherein the altered expression is associated with the occurrence and / or progression of the disease. Disease-related genes also refer to genes with mutations or genetic variations that are directly responsible for or in linkage disequilibrium with genes that cause the disease (e.g., causal mutations). The transcriptional or translational product may be known or unknown, and may further be at normal or abnormal levels.

[0427] Non-limiting examples of disease-related genes that can be targeted using the methods and compositions of this disclosure are available from McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.), and are available on the World Wide Web.

[0428] In some embodiments, the method includes contacting a DNA molecule containing a target DNA sequence with a fusion protein of the present invention, wherein the DNA molecule is contacted with the fusion protein in an effective amount and under conditions suitable for modification of the target DNA molecule. In some embodiments, the method includes contacting a DNA molecule containing a target DNA sequence with: (a) the RGN fusion protein of the present invention, which contains a base-editing polypeptide; and (b) gRNA that targets the fusion protein of (a) to a target nucleotide sequence of a DNA strand; wherein the DNA molecule is contacted with the fusion protein and gRNA in an effective amount and under conditions suitable for deamination of at least one nucleobase. In some embodiments, the target DNA sequence contains a sequence associated with a disease or symptom, and wherein deamination of the nucleobase results in a sequence not associated with a disease or symptom. In some embodiments, the target DNA sequence is present in the alleles of a crop plant, wherein a specific allele of the trait of interest results in a plant with lower agronomic value. Deamination of the nucleobase results in an allele that improves the trait and increases the agronomic value of the plant.

[0429] In some embodiments, the target DNA sequence contains a point mutation associated with a disease or condition, and deamination of the mutant base results in a sequence not associated with the disease or condition. In some embodiments, deamination corrects the point mutation in the disease or condition-associated sequence.

[0430] In some embodiments, the sequence associated with a disease or condition encodes a protein, and deamination or leader editing introduces a stop codon into the disease-associated sequence, resulting in a truncation of the encoded protein. In some embodiments, exposure is conducted in a subject who is susceptible, has, or is diagnosed with a disease or condition. In some embodiments, the disease or condition is associated with a point mutation or single-base mutation in the genome. In some embodiments, the disease is a genetic disease, cancer, metabolic disease, or lysosomal storage disease.

[0431] XII. Pharmaceutical Compositions and Treatment Methods

[0432] This document provides a method for treating a disease in a subject with this need. The method includes administering to the subject with the fusion protein of this disclosure or a polynucleotide encoding the fusion protein, gRNA or a polynucleotide encoding the gRNA, a fusion protein system of this disclosure, or cells modified by any of these compositions or cells containing any of these compositions.

[0433] In some embodiments, the treatment includes in vivo gene editing by administering a fusion protein, gRNA, or binding protein system of the present disclosure, or a polynucleotide encoding thereof, to a subject in need. In some embodiments, the treatment includes in vitro gene editing, wherein cells are genetically modified in vitro with a fusion protein, gRNA, or fusion protein system of the present disclosure, or a polynucleotide encoding thereof, and then the modified cells are administered to a subject. In some embodiments, the genetically modified cells are derived from a subject to whom the modified cells are then administered, and the transplanted cells are referred to herein as autologous cells. In some embodiments, the genetically modified cells are derived from a different subject (i.e., a donor) within the same species as the subject to whom the modified cells are administered (i.e., the recipient), and the transplanted cells are referred herein as allogeneic cells. In some examples described herein, the cells may be expanded in culture prior to administration to a subject in need.

[0434] In some embodiments, the disease to be treated with the compositions of this disclosure is a disease that can be treated with immunotherapy, such as chimeric antigen receptor (CAR) T cells. Such diseases include, but are not limited to, cancer.

[0435] In some embodiments, modification of the target sequence (e.g., deamination) leads to the correction of a genetic defect or to the correction of a point mutation that results in the loss of function of a gene product. In some embodiments, the genetic defect is associated with a disease or condition (e.g., lysosomal storage disease or metabolic disease, such as, for example, type 1 diabetes). Therefore, in some embodiments, the disease to be treated with the compositions of this disclosure is associated with the mutated sequence (i.e., the sequence is the cause of the disease or condition or the cause of symptoms associated with the disease or condition) in order to treat the disease or condition or alleviate the symptoms associated with the disease or condition.

[0436] In some embodiments, the disease to be treated with the compositions of this disclosure is associated with a causal mutation. As used herein, a “causal mutation” refers to a specific nucleotide, nucleotide, or nucleotide sequence in the genome that contributes to the severity or presence of a disease or condition in a subject. Correction of the causal mutation results in an improvement in at least one symptom caused by the disease. In some embodiments, correction of the causal mutation results in an improvement in at least one symptom caused by the disease. In some embodiments, the causal mutation is adjacent to a PAM site recognized by the RDBBP (e.g., RGN) of the fusion protein disclosed herein. The causal mutation can be corrected with the fusion protein. Non-limiting examples of disease-related genes and mutations are available from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.), and are available on the World Wide Web.

[0437] In some embodiments, the methods provided herein are used to introduce deactivation point mutations into a gene or allele encoding a gene product associated with a disease or condition. For example, in some embodiments, this document provides a method for introducing deactivation point mutations into oncogenes using fusion proteins (e.g., in the treatment of proliferative diseases). In some embodiments, the deactivation mutation may generate premature stop codons in the coding sequence, resulting in the expression of truncated gene products (e.g., truncated proteins lacking the function of the full-length protein). In some embodiments, the purpose of the methods provided herein is to restore the function of dysfunctional genes via genome editing. The fusion proteins provided herein can be validated in vitro for use in gene-editing-based human therapies, for example, by correcting disease-related mutations in human cell cultures. Those skilled in the art will understand that the fusion proteins provided herein (e.g., fusion proteins comprising an RNA-guided DNA-binding polypeptide and an adenine deaminase polypeptide) can be used to correct any single-point A>G mutation. Deamination of the mutant A to G results in the correction of the mutation.

[0438] As used herein, the terms “treatment,” “treating,” “palliating,” or “ameliorating” are used interchangeably. These terms refer to a method of obtaining a beneficial or desired outcome, including but not limited to therapeutic and / or preventive benefits. A therapeutic benefit means any treatment-related improvement or effect on one or more diseases, conditions, or symptoms during treatment. For preventive benefits, the composition may be administered to subjects at risk of developing a particular disease, condition, or symptom, or to subjects who report one or more physiological symptoms of a disease, even if the disease, condition, or symptom may not yet have appeared.

[0439] The term "effective amount" or "therapeutic effective amount" refers to the amount of a pharmaceutical agent sufficient to achieve a beneficial or desired outcome. Therapeutic effective amounts can vary based on one or more of the following: the subject receiving treatment and the disease condition, the subject's weight and age, the severity of the disease condition, the method of administration, etc., which can be readily determined by a person skilled in the art. Specific dosages can vary based on one or more of the following: the specific pharmaceutical agent selected, the dosing regimen followed, whether it is administered in combination with other compounds, the timing of administration, and the delivery system used.

[0440] The term "application" refers to the placement of an active ingredient into a subject by a method or route that results in the introduced active ingredient being at least partially located at a desired site, such as a site of damage or repair, thereby producing the desired effect. In those embodiments of cell application, the cells can be applied via any suitable route that results in delivery to the desired site in the subject, where at least a portion of the implanted cells or cell components remain viable. The survival time of the cells after application to the subject can be as short as a few hours (e.g., twenty-four hours), as short as a few days, as long as a few years, or even as long as a patient's lifetime, i.e., long-term implantation. For example, in some aspects described herein, an effective amount of photoreceptor cells or retinal progenitor cells is applied via a systemic application route, such as an intraperitoneal or intravenous route.

[0441] In some embodiments, administration includes administration via viral delivery. In some embodiments, administration includes administration via electroporation. In some embodiments, administration includes administration via nanoparticle delivery. In some embodiments, administration includes administration via liposome delivery. In some embodiments, administration includes administration via lipid nanoparticle (LNP) delivery. In some embodiments, LNPs or lipid components thereof are described in the following documents: WO2022173531, WO2022150485, or U.S. Provisional Application No. 63 / 492,537, filed March 28, 2023, each of which is incorporated herein by reference in its entirety. Any effective route of administration may be used to administer an effective amount of the pharmaceutical composition described herein. In some embodiments, administration includes administration by methods selected from the group consisting of intravenous, subcutaneous, intramuscular, oral, rectal, spray, parenteral, ocular, pulmonary, percutaneous, vaginal, ear, nose, and topical administration, or any combination thereof. In some embodiments, for cell delivery, administration by injection or infusion is used.

[0442] As used herein, the term "subject" refers to any individual who requires diagnosis, treatment, or therapy. In some embodiments, a subject is an animal. In some embodiments, a subject is a mammal. In some embodiments, a subject is a human. A human can be an adult, adolescent, child, or infant.

[0443] The efficacy of a treatment can be determined by a skilled clinician. However, a treatment is considered “effective” if any or all of the signs or symptoms of the disease or condition change in a beneficial manner (e.g., a reduction of at least 10%), or if other clinically accepted signs or markers of the disease are improved or alleviated. Efficacy can also be measured by the failure of an individual to deteriorate, such as by hospitalization or the need for medical intervention (e.g., cessation or at least slowing of disease progression). Methods for measuring these indicators are known to those skilled in the art. Treatments include: (1) suppressing the disease, e.g., halting or slowing the progression of symptoms; or (2) alleviating the disease, e.g., causing symptom resolution; and (3) preventing or reducing the likelihood of symptom development.

[0444] A pharmaceutical composition is provided comprising: the fusion protein of the present disclosure or the polynucleotide encoding the fusion protein; the system of the present disclosure; or a cell comprising the fusion protein, the polynucleotide encoding the fusion protein, or any one of the system; and a pharmaceutically acceptable carrier.

[0445] As used herein, a “pharmaceutically acceptable carrier” is a material that does not cause significant irritation to an organism and does not eliminate the activity and properties of the active ingredient (e.g., a deaminase or fusion protein, or a nucleic acid molecule encoding a deaminase or fusion protein). The carrier must be of sufficiently high purity and sufficiently low toxicity to be suitable for administration to a subject receiving treatment. The carrier may be inert, or it may have pharmaceutical benefits. In some embodiments, a pharmaceutically acceptable carrier comprises one or more compatible solid or liquid fillers, diluents, or encapsulating substances suitable for administration to humans or other vertebrates. In some embodiments, the pharmaceutical composition comprises a pharmaceutically acceptable carrier that is not naturally occurring. In some embodiments, the pharmaceutically acceptable carrier and the active ingredient are not found together in nature and are therefore heterogeneous.

[0446] Pharmaceutical compositions used in the methods of this disclosure can be formulated with suitable carriers, excipients, and other agents that provide suitable transfer, delivery, tolerability, etc. A variety of suitable formulations are known to those skilled in the art. See, for example, Remington, The Science and Practice of Pharmacy (21st edition, 2005). Non-limiting examples include sterile diluents such as water for injection, saline solutions, fixative oils, polyethylene glycol, glycerol, propylene glycol, or other synthetic solvents; antibacterial agents such as benzyl alcohol or methylparaben; antioxidants such as ascorbic acid or sodium bisulfite; chelating agents such as ethylenediaminetetraacetic acid; buffers such as acetates, citrates, or phosphates; and agents for adjusting tension such as sodium chloride or glucose. A specific carrier for intravenous administration is physiological saline or phosphate-buffered saline (PBS). Pharmaceutical compositions intended for oral or parenteral use can be formulated into dosage forms in unit doses suitable for containing a specific amount of the active ingredient. Such dosage forms in unit doses include, for example, tablets, pills, capsules, injections (ampoules), suppositories, etc. These compositions may also contain adjuvants, including preservatives, wetting agents, emulsifiers, and dispersants. Antimicrobial activity can be ensured by various antibacterial and antifungal agents, such as parabens, chlorobutanol, phenol, and sorbic acid. Isotonic agents, such as sugars and sodium chloride, are also desirable. Prolonged absorption of injectable drug forms can be achieved by using agents that delay absorption, such as aluminum monostearate and gelatin.

[0447] In some embodiments, cells comprising a fusion protein, system, or polynucleotide encoding thereof, or cells modified with a fusion protein, system, or polynucleotide encoding thereof, are administered to a subject in the form of a suspension along with a pharmaceutically acceptable carrier. Those skilled in the art will recognize that a pharmaceutically acceptable carrier to be used in the cell composition will not, in an amount that substantially interferes with the viability of the cells to be delivered to the subject, include buffers, compounds, cryopreservatives, preservatives, or other formulations. Formulations containing cells may include, for example, permeability buffers that allow cell membrane integrity to be maintained, and optionally nutrients to maintain cell viability or enhance translocation after administration. Such formulations and suspensions are known to those skilled in the art, and / or can be adapted for use with the cells described herein using conventional experiments.

[0448] Cellular compositions may also be emulsified or presented as liposome compositions, provided that the emulsification process does not adversely affect cell viability. Cells and any other active ingredients may be mixed with pharmaceutically acceptable and compatible excipients in amounts suitable for the therapeutic methods described herein.

[0449] Additional formulations included in the cellular composition may include pharmaceutically acceptable salts of its components. Pharmaceutically acceptable salts include acid addition salts (formed with the free amino group of a polypeptide) formed with inorganic acids (such as, for example, hydrochloric acid or phosphoric acid) or organic acids (such as acetic acid, tartaric acid, mandelic acid, etc.). Salts formed from free carboxyl groups may also be derived from inorganic bases (such as, for example, sodium hydroxide, potassium hydroxide, ammonium hydroxide, calcium hydroxide, or iron hydroxide) and organic bases (such as isopropylamine, trimethylamine, 2-ethylamine ethanol, histidine, procaine, etc.).

[0450] XIII. Cells containing polynucleotide genetic modifications

[0451] This article provides cells and organisms containing the target nucleic acid molecule of interest, which has been modified using a fusion protein-mediated process as described herein, optionally with gRNA.

[0452] In some embodiments, the fusion protein comprises a deaminase polypeptide comprising an amino acid sequence or an active variant or fragment thereof of any one of SEQ ID NO:5, 6, 300-414, 596-598, and 720-723, comprising at least one of the amino acid residues shown in any one of Tables 2, 4, 6, and 19. In some embodiments, the fusion protein comprises a deaminase comprising an amino acid sequence having about 50% to about 99% or higher identity with any one of SEQ ID NO:5, 6, 300-414, 596-598, and 720-723. In some embodiments, the fusion protein comprises a deaminase having an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with any of SEQ ID NO: 5, 6, 300-414, 596-598, and 720-723. In some embodiments, the fusion protein comprises a deaminase and a DNA-binding polypeptide (e.g., an RNA-guided DNA-binding polypeptide). In further embodiments, the fusion protein comprises a deaminase and RGN or a variant thereof, such as, for example, APG07433.1 (SEQ ID NO: 1) or a nicking enzyme variant thereof, nAPG07433.1 (SEQ ID NO: 49). In some embodiments, the fusion protein comprises a deaminase and Cas9 or a variant thereof, such as, for example, dCas9 or a nicking enzyme Cas9. In some embodiments, the fusion protein comprises a nuclease-inactive or nickase variant of a type II CRISPR-Cas polypeptide. In some embodiments, the fusion protein comprises a nuclease-inactive or nickase variant of a type V CRISPR-Cas polypeptide. In some embodiments, the fusion protein comprises a nuclease-inactive or nickase variant of a type VI CRISPR-Cas polypeptide.

[0453] The modified cells can be eukaryotic (e.g., mammalian, plant, insect, or avian cells) or prokaryotic. Organelles and embryos containing at least one nucleotide sequence have also been provided, which has been modified using a fusion protein process as described herein. For the modified nucleotide sequence, the genetically modified cells, organisms, organelles, and embryos can be heterozygous or homozygous. Mutations introduced by the fusion protein can result in altered expression (upregulation or downregulation), inactivation, or altered expression of the protein product or integrated sequence. In those cases where the mutation results in gene inactivation or nonfunctional protein product expression, the genetically modified cells, organisms, organelles, or embryos are referred to as “knockout.” The knockout phenotype can be the result of a deletion mutation (i.e., the deletion of at least one nucleotide), an insertion mutation (i.e., the insertion of at least one nucleotide), or a nonsense mutation (i.e., the substitution of at least one nucleotide, resulting in the introduction of a stop codon).

[0454] In some embodiments, mutations introduced by the fusion protein result in the production of a variant protein product. The expressed variant protein product may have at least one amino acid substitution and / or at least one amino acid addition or deletion. When compared to the wild-type protein, the variant protein product may exhibit modified characteristics or activities, including but not limited to altered enzyme activity or substrate specificity.

[0455] In some embodiments, mutations introduced by the fusion protein result in altered protein expression patterns. As a non-limiting example, mutations in the regulatory regions controlling the expression of protein products can lead to overexpression or downregulation of the protein product or altered tissue or temporal expression patterns.

[0456] Modified cells can be grown into organisms, such as plants, using conventional methods. See, for example, McCormick et al. (1986) Plant Cell Reports 5:81-84. These plants can then be grown and pollinated with the same or different strains of the modified organism, and the resulting hybrids will have the genetic modification. This invention provides genetically modified seeds. Progeny, variants, and mutants of regenerated plants are also included within the scope of this invention, provided that these portions contain the genetic modification. Further, processed plant products or byproducts that retain the genetic modification are provided, including, for example, soybean meal.

[0457] The methods described herein can be used to modify any plant species, including but not limited to monocots and dicots. Examples of plants of interest include, but are not limited to, corn, sorghum, wheat, sunflower, tomato, cruciferous plants, pepper, potato, cotton, rice, soybean, sugar beet, sugarcane, tobacco, barley and rapeseed, Brassica, alfalfa, rye, millet, safflower, peanut, sweet potato, cassava, coffee, coconut, pineapple, citrus, cocoa, tea, banana, avocado, fig, guava, mango, olive, papaya, cashew, macadamia nut, almond, oat, vegetables, ornamental plants, and conifers.

[0458] Vegetables include, but are not limited to, tomatoes, lettuce, green beans, lima beans, peas, and members of the Cucumber genus, such as cucumbers, cantaloupes, and honeydew melons. Ornamental plants include, but are not limited to, azaleas, hydrangeas, hibiscus, roses, tulips, daffodils, petunias, carnations, poinsettias, and chrysanthemums. Preferably, the plants of this invention are crop plants (e.g., corn, sorghum, wheat, sunflowers, tomatoes, cruciferous plants, peppers, potatoes, cotton, rice, soybeans, sugar beets, sugarcane, tobacco, barley, rapeseed, etc.).

[0459] The methods described herein can also be used to genetically modify any prokaryotic species, including but not limited to archaea and bacteria (e.g., Bacillus, Klebsiella, Streptomyces, Rhizobium, Escherichia, Pseudomonas, Salmonella, Shigella, Vibrio, Yersinia, Mycoplasma, Agrobacterium, Lactobacillus).

[0460] The methods provided herein can be used to genetically modify any eukaryotic species or cells derived therefrom, including but not limited to animals (e.g., mammals, insects, fish, birds, and reptiles), fungi, amoebas, algae, and yeast. In some embodiments, cells modified by the methods of this disclosure include hematopoietic cells, such as immune cells (i.e., cells of the innate or adaptive immune system), including but not limited to B cells, T cells, natural killer (NK) cells, pluripotent stem cells, induced pluripotent stem cells, chimeric antigen receptor T (CAR-T) cells, monocytes, macrophages, and dendritic cells.

[0461] Modified cells can be introduced into an organism. In the case of autologous cell transplantation, these cells can be derived from the same organism (e.g., a human), where the cells are modified using an ex vivo protocol. In some embodiments, in the case of allogeneic cell transplantation, the cells are derived from another organism within the same species (e.g., another human).

[0462] XIV. Reagent Kit

[0463] Some aspects of this disclosure provide kits comprising the deaminase or fusion protein of the present invention. In some embodiments, this disclosure provides kits comprising the fusion protein and guide RNA. Additionally, in some embodiments, the kit includes suitable reagents, buffers, and / or instructions for using the fusion protein, for example, for in vitro or in vivo DNA or RNA editing. In some embodiments, the kit includes instructions for the design and use of suitable gRNAs for targeted editing of nucleic acid sequences.

[0464] The articles “a” and “a kind” in this text refer to one or more grammatical objects of the article (i.e., at least one). For example, “polypeptide” means one or more polypeptides.

[0465] All publications and patent applications referenced in this specification are intended to be of the skill of a person skilled in the art to which this disclosure pertains. All publications and patent applications are incorporated herein by reference as if each individual publication or patent application were specifically and individually indicated to be incorporated herein by reference.

[0466] Although the invention has been described in detail by way of illustrations and examples for clarity, it will be apparent that certain changes and modifications may be made within the scope of the appended claims.

[0467] Non-limiting embodiments include:

[0468] 1. A fusion protein comprising an RNA-guided nuclease (RGN) and at least one heterologous polypeptide inserted therein, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

[0469] 2. The fusion protein according to Example 1, wherein the RGN comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

[0470] 3. The fusion protein according to Example 1, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

[0471] 4. The fusion protein according to Example 1, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 1-4 and 57-139.

[0472] 5. The fusion protein according to Example 1, wherein the RGN comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 1-4 and 57-139.

[0473] 6. The fusion protein according to Example 1, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NO: 1-4 and 57-139.

[0474] 7. The fusion protein according to any one of Examples 1 to 6, wherein the heterologous polypeptide is inserted into the linker domain 2, wedge domain, RuvC domain, HNH domain, Rec-2 domain or PAM interaction domain of the RGN.

[0475] 8. The fusion protein according to Example 7, wherein the RuvC domain is the RuvCIII domain.

[0476] 9. The fusion protein according to any one of claims 1 to 8, wherein:

[0477] a) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0478] i) The amino acid position corresponding to position 30 of SEQ ID NO:1;

[0479] ii) The amino acid position corresponding to position 642 in SEQ ID NO:1;

[0480] iii) The amino acid position corresponding to position 670 of SEQ ID NO:1;

[0481] iv) The amino acid position corresponding to position 737 of SEQ ID NO:1;

[0482] v) The amino acid position corresponding to position 772 of SEQ ID NO:1;

[0483] vi) The amino acid position corresponding to position 775 of SEQ ID NO:1;

[0484] vii) The amino acid position corresponding to position 778 of SEQ ID NO:1; and

[0485] viii) The amino acid position corresponding to position 802 in SEQ ID NO:1;

[0486] b) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0487] i) The amino acid position corresponding to position 678 of SEQ ID NO:2;

[0488] ii) The amino acid position corresponding to position 736 of SEQ ID NO:2;

[0489] iii) The amino acid position corresponding to position 778 of SEQ ID NO:2;

[0490] iv) The amino acid position corresponding to position 788 of SEQ ID NO:2; and

[0491] v) The amino acid position corresponding to position 922 in SEQ ID NO:2;

[0492] c) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:3, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0493] i) The amino acid position corresponding to position 725 of SEQ ID NO:3;

[0494] ii) The amino acid position corresponding to position 739 in SEQ ID NO:3; and

[0495] iii) The amino acid position corresponding to position 744 of SEQ ID NO:3;

[0496] d) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:4, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0497] i) The amino acid position corresponding to position 347 in SEQ ID NO:4;

[0498] ii) The amino acid position corresponding to position 524 in SEQ ID NO:4;

[0499] iii) The amino acid position corresponding to position 666 in SEQ ID NO:4;

[0500] iv) The amino acid position corresponding to position 680 of SEQ ID NO:4;

[0501] v) The amino acid position corresponding to position 740 of SEQ ID NO:4;

[0502] vi) The amino acid position corresponding to position 785 of SEQ ID NO:4;

[0503] vii) The amino acid position corresponding to position 910 of SEQ ID NO:4; and

[0504] viii) The amino acid position corresponding to position 1077 of SEQ ID NO:4; or

[0505] e) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:131, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0506] i) The amino acid position corresponding to position 766 of SEQ ID NO: 131; and

[0507] ii) The amino acid position corresponding to position 806 of SEQ ID NO:131.

[0508] 10. The fusion protein according to any one of Examples 1 to 9, wherein the heterologous polypeptide comprises a base-editing polypeptide or a lead-editing polypeptide.

[0509] 11. The fusion protein according to Example 10, wherein the lead editing polypeptide comprises a DNA polymerase.

[0510] 12. The fusion protein according to Example 10, wherein the lead editing polypeptide is a reverse transcriptase.

[0511] 13. The fusion protein according to Example 10, wherein the base-editing polypeptide comprises a deaminase.

[0512] 14. The fusion protein according to Example 13, wherein the deaminase is cytosine deaminase or adenine deaminase.

[0513] 15. The fusion protein according to Example 13 or 14, wherein the fusion protein is a base editor fusion protein, and wherein the base editor fusion protein has improved editing activity compared to a parental base editor comprising a deaminase fused to the amino terminus of the RGN.

[0514] 16. The fusion protein according to Examples 13-15, wherein the at least one heteropeptide is inserted into an RGN containing an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1 immediately following an amino acid position selected from the group consisting of:

[0515] a) The amino acid position corresponding to position 642 in SEQ ID NO:1;

[0516] b) The amino acid position corresponding to position 670 of SEQ ID NO:1;

[0517] c) The amino acid position corresponding to position 737 in SEQ ID NO:1;

[0518] d) The amino acid position corresponding to position 772 of SEQ ID NO:1;

[0519] e) The amino acid position corresponding to position 775 of SEQ ID NO: 1; and

[0520] f) The amino acid position corresponding to position 778 of SEQ ID NO: 1; and

[0521] The fusion protein therein is a base editor fusion protein, and the base editor fusion protein has a shifted editing window compared to the parental base editor containing the deaminase fused to the amino terminus of the RGN.

[0522] 17. The fusion protein according to any one of Examples 13 to 16, wherein the deaminase lacks the first amino acid residue and the last amino acid residue compared to the parental deaminase from which the deaminase is derived.

[0523] 18. The fusion protein according to any one of Examples 13 to 17, wherein the deaminase comprises adenine deaminase, and the fusion protein is an adenosine base editor (ABE) fusion protein.

[0524] 19. The fusion protein according to Example 18, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

[0525] 20. The fusion protein according to Example 18, wherein the adenine deaminase comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

[0526] 21. The fusion protein according to Example 18, wherein the adenine deaminase comprises the amino acid sequence of any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

[0527] 22. The fusion protein according to Example 18, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:5 or 6.

[0528] 23. The fusion protein according to Example 18, wherein the adenine deaminase comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO:5 or 6.

[0529] 24. The fusion protein according to Example 18, wherein the adenine deaminase comprises the amino acid sequence of SEQ ID NO: 5 or 6.

[0530] 25. The fusion protein according to Example 18, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

[0531] 26. The fusion protein according to Example 18, wherein the ABE fusion protein comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

[0532] 27. The fusion protein according to Example 18, wherein the ABE fusion protein comprises the amino acid sequence shown in any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

[0533] 28. The fusion protein according to Example 18, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parental ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN.

[0534] 29. The fusion protein according to Example 18, wherein the ABE fusion protein has at least 95% sequence identity with any one of SEQ ID NO: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parental ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN.

[0535] 30. The fusion protein according to Example 18, wherein the ABE fusion protein comprises the amino acid sequence shown in any one of SEQ ID NO: 8-10 and 14-16.

[0536] 31. The fusion protein according to any one of Examples 1 to 30, wherein the RGN is a nicking enzyme or a nuclease-inactive RGN.

[0537] 32. The fusion protein according to Example 31, wherein the RGN nickase comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO:49-56, 698 and 699, and retains nickase activity.

[0538] 33. The fusion protein according to Example 31, wherein the RGN nickase comprises the amino acid sequence of any one of SEQ ID NO:49-56, 698 and 699.

[0539] 34. The fusion protein according to any one of Examples 1 to 33, wherein the RGN and the heteropeptide are directly fused to each other without a linker sequence.

[0540] 35. The fusion protein according to any one of Examples 1 to 34, wherein the fusion protein further comprises at least one nuclear localization signal (NLS).

[0541] 36. A ribonucleoprotein (RNP) complex comprising a fusion protein according to any one of Examples 1 to 35 and 232 to 243 and a guide RNA bound to said fusion protein.

[0542] 37. A nucleic acid molecule comprising a polynucleotide encoding a fusion protein, wherein the fusion protein comprises an RNA-guided nuclease (RGN) and at least one heterologous polypeptide inserted therein, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

[0543] 38. The nucleic acid molecule according to Example 37, wherein the RGN comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

[0544] 39. The nucleic acid molecule according to Example 37, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

[0545] 40. The nucleic acid molecule according to Example 37, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 1-4, 57-139.

[0546] 41. The nucleic acid molecule according to Example 37, wherein the RGN comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO:1-4, 57-139.

[0547] 42. The nucleic acid molecule according to Example 37, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NO: 1-4, 57 and 139.

[0548] 43. The nucleic acid molecule according to any one of Examples 1 to 6, wherein the heterologous polypeptide is inserted into the adapter domain 2, wedge domain, RuvC domain, HNH domain, Rec-2 domain or PAM interaction domain of the RGN.

[0549] 44. The nucleic acid molecule according to Example 7, wherein the RuvC domain is a RuvCIII domain.

[0550] 45. The nucleic acid molecule according to any one of Examples 37 to 44, wherein:

[0551] a) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0552] i) The amino acid position corresponding to position 30 of SEQ ID NO:1;

[0553] ii) The amino acid position corresponding to position 642 in SEQ ID NO:1;

[0554] iii) The amino acid position corresponding to position 670 of SEQ ID NO:1;

[0555] iv) The amino acid position corresponding to position 737 of SEQ ID NO:1;

[0556] v) The amino acid position corresponding to position 772 of SEQ ID NO:1;

[0557] vi) The amino acid position corresponding to position 775 of SEQ ID NO:1;

[0558] vii) The amino acid position corresponding to position 778 of SEQ ID NO:1; and

[0559] viii) The amino acid position corresponding to position 802 in SEQ ID NO:1;

[0560] b) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0561] i) The amino acid position corresponding to position 678 of SEQ ID NO:2;

[0562] ii) The amino acid position corresponding to position 736 of SEQ ID NO:2;

[0563] iii) The amino acid position corresponding to position 778 of SEQ ID NO:2;

[0564] iv) The amino acid position corresponding to position 788 of SEQ ID NO:2; and

[0565] v) The amino acid position corresponding to position 922 in SEQ ID NO:2;

[0566] c) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:3, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0567] i) The amino acid position corresponding to position 725 of SEQ ID NO:3;

[0568] ii) The amino acid position corresponding to position 739 in SEQ ID NO:3; and

[0569] iii) The amino acid position corresponding to position 744 of SEQ ID NO:3;

[0570] d) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:4, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0571] i) The amino acid position corresponding to position 347 in SEQ ID NO:4;

[0572] ii) The amino acid position corresponding to position 524 in SEQ ID NO:4;

[0573] iii) The amino acid position corresponding to position 666 in SEQ ID NO:4;

[0574] iv) The amino acid position corresponding to position 680 of SEQ ID NO:4;

[0575] v) The amino acid position corresponding to position 740 of SEQ ID NO:4;

[0576] vi) The amino acid position corresponding to position 785 of SEQ ID NO:4;

[0577] vii) The amino acid position corresponding to position 910 of SEQ ID NO:4; and

[0578] viii) The amino acid position corresponding to position 1077 of SEQ ID NO:4; or

[0579] e) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:131, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0580] i) The amino acid position corresponding to position 766 of SEQ ID NO: 131; and

[0581] ii) The amino acid position corresponding to position 806 of SEQ ID NO:131.

[0582] 46. ​​The nucleic acid molecule according to Example 45, wherein the heterologous polypeptide comprises a base-editing polypeptide or a lead-editing polypeptide.

[0583] 47. The nucleic acid molecule according to Example 46, wherein the lead editing polypeptide comprises a DNA polymerase.

[0584] 48. The nucleic acid molecule according to Example 46, wherein the lead editing polypeptide comprises a reverse transcriptase.

[0585] 49. The nucleic acid molecule according to Example 46, wherein the base-editing polypeptide comprises a deaminase.

[0586] 50. The nucleic acid molecule according to Example 49, wherein the deaminase is cytosine deaminase or adenine deaminase.

[0587] 51. The nucleic acid molecule according to Example 49 or 50, wherein the fusion protein is a base editor fusion protein, and wherein the base editor fusion protein has improved editing activity compared to a parental base editor comprising a deaminase fused to the amino terminus of the RGN.

[0588] 52. The nucleic acid molecule according to any one of Examples 49 to 51, wherein the at least one heterologous polypeptide is inserted immediately after an amino acid position selected from the group consisting of the following into the RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1:

[0589] a) The amino acid position corresponding to position 642 in SEQ ID NO:1;

[0590] b) The amino acid position corresponding to position 670 of SEQ ID NO:1;

[0591] c) The amino acid position corresponding to position 737 in SEQ ID NO:1;

[0592] d) The amino acid position corresponding to position 772 of SEQ ID NO:1;

[0593] e) The amino acid position corresponding to position 775 of SEQ ID NO: 1; and

[0594] f) The amino acid position corresponding to position 778 of SEQ ID NO: 1; and

[0595] The fusion protein therein is a base editor fusion protein, and the base editor fusion protein has a shifted editing window compared to the parental base editor containing the deaminase fused to the amino terminus of the RGN.

[0596] 53. The nucleic acid molecule according to any one of Examples 49 to 52, wherein the deaminase lacks the first amino acid residue and the last amino acid residue compared to the parental deaminase from which the deaminase is derived.

[0597] 54. The nucleic acid molecule according to any one of Examples 49 to 53, wherein the deaminase comprises adenine deaminase and is an adenosine base editor (ABE) fusion protein.

[0598] 55. The nucleic acid molecule according to Example 54, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

[0599] 56. The nucleic acid molecule according to Example 54, wherein the adenine deaminase comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

[0600] 57. The nucleic acid molecule according to Example 54, wherein the adenine deaminase comprises the amino acid sequence of any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

[0601] 58. The nucleic acid molecule according to Example 54, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:5 or 6.

[0602] 59. The nucleic acid molecule according to Example 54, wherein the adenine deaminase comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO:5 or 6.

[0603] 60. The nucleic acid molecule according to Example 54, wherein the adenine deaminase comprises the amino acid sequence of SEQ ID NO: 5 or 6.

[0604] 61. The nucleic acid molecule according to Example 54, wherein the fusion protein comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

[0605] 62. The nucleic acid molecule according to Example 54, wherein the ABE fusion protein comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

[0606] 63. The nucleic acid molecule according to Example 54, wherein the ABE fusion protein comprises the amino acid sequence shown in any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

[0607] 64. The nucleic acid molecule according to Example 54, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parental ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN.

[0608] 65. The nucleic acid molecule according to Example 54, wherein the ABE fusion protein comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parental ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN.

[0609] 66. The nucleic acid molecule according to Example 54, wherein the ABE fusion protein comprises the amino acid sequence shown in any one of SEQ ID NO: 8-10 and 14-16.

[0610] 67. The nucleic acid molecule according to any one of Examples 37 to 66, wherein the RGN is a nicking enzyme or a nuclease-inactive RGN.

[0611] 68. The nucleic acid molecule according to Example 67, wherein the RGN nickase comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO:49-56, 698 and 699, and retains nickase activity.

[0612] 69. The nucleic acid molecule according to Example 67, wherein the RGN nickase comprises the amino acid sequence of any one of SEQ ID NO:49-56, 698 and 699.

[0613] 70. The nucleic acid molecule according to any one of Examples 37 to 69, wherein the RGN and the heterologous polypeptide are directly fused to each other without a linker sequence.

[0614] 71. The nucleic acid molecule according to any one of Examples 37 to 70, wherein the fusion protein further comprises at least one nuclear localization signal (NLS).

[0615] 72. The nucleic acid molecule according to any one of Examples 37 to 71, wherein the nucleic acid molecule is codon-optimized for expression in eukaryotic cells.

[0616] 73. The nucleic acid molecule according to any one of Examples 37 to 71, wherein the nucleic acid molecule is codon-optimized for expression in prokaryotic cells.

[0617] 74. A vector comprising a nucleic acid molecule according to any one of Examples 37 to 73 and 244 to 255.

[0618] 75. The vector according to Example 74 further comprises at least one nucleotide sequence encoding a guide RNA capable of hybridizing with a non-target strand of the target sequence of the RGN.

[0619] 76. The vector according to Example 74, wherein the guide RNA is a single guide RNA (sgRNA).

[0620] 77. The vector according to Example 74, wherein the guide RNA is a dual guide RNA.

[0621] 78. A cell comprising a fusion protein according to any one of Examples 1 to 35 and 232 to 243 or an RNP complex according to Example 36.

[0622] 79. A cell comprising a fusion protein according to any one of Examples 1 to 35 and 232 to 243, wherein the cell further comprises guide RNA.

[0623] 80. A cell comprising a nucleic acid molecule according to any one of Examples 37 to 73 and 244 to 255.

[0624] 81. A cell comprising a carrier according to any one of Examples 74 to 77.

[0625] 82. The cell according to any one of Examples 78 to 81, wherein the cell is a prokaryotic cell.

[0626] 83. The cell according to any one of Examples 78 to 81, wherein the cell is a eukaryotic cell.

[0627] 84. The cell according to Example 83, wherein the eukaryotic cell is a mammalian cell.

[0628] 85. The cell according to Example 84, wherein the mammalian cell is a human cell.

[0629] 86. The cell according to Example 85, wherein the human cell is an immune cell.

[0630] 87. The cell according to Example 86, wherein the immune cell is a stem cell.

[0631] 88. The cell according to Example 87, wherein the stem cell is an induced pluripotent stem cell.

[0632] 89. The cell according to Example 83, wherein the eukaryotic cell is an insect or avian cell.

[0633] 90. The cell according to Example 83, wherein the eukaryotic cell is a fungal cell.

[0634] 91. The cell according to Example 83, wherein the eukaryotic cell is a plant cell.

[0635] 92. A plant comprising cells according to Example 91.

[0636] 93. A seed comprising cells according to Example 91.

[0637] 94. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a fusion protein according to any one of Examples 1 to 35 and 232 to 243, an RNP complex according to Example 36, a nucleic acid molecule according to any one of Examples 37 to 73 and 244 to 255, a vector according to any one of Examples 74 to 77, or a cell according to any one of Examples 84 to 88.

[0638] 95. A method for manufacturing a fusion protein, the method comprising culturing cells according to any one of Examples 78 to 91 under conditions in which the fusion protein is expressed.

[0639] 96. A method for manufacturing a fusion protein, the method comprising introducing a nucleic acid molecule according to any one of Examples 37 to 73 and 237 to 243 or a vector according to any one of Examples 74 to 77 into a cell, and culturing the cell under conditions in which the fusion protein is expressed.

[0640] 97. The method according to Example 95 or 96, further comprising purifying the fusion protein.

[0641] 98. A method for manufacturing an RGN fusion ribonucleoprotein complex, the method comprising introducing a nucleic acid molecule according to any one of Examples 37 to 73 and 237 to 243, and a nucleic acid molecule comprising an expression cassette encoding guide RNA, or a vector according to any one of Examples 74 to 77, into a cell, and culturing the cell under conditions where the fusion protein and gRNA are expressed and an RGN fusion ribonucleoprotein complex is formed.

[0642] 99. The method according to Example 98, further comprising purifying the RGN fusion ribonucleoprotein complex.

[0643] 100. A system for targeting a heterologous polypeptide to a target DNA molecule containing a target DNA sequence, the system comprising:

[0644] a) a fusion protein or a nucleotide sequence encoding said fusion protein, wherein said fusion protein comprises an RNA-guided nuclease (RGN) and at least one heterologous polypeptide inserted therein; wherein said RGN comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698, and 699; and

[0645] b) one or more guide RNAs capable of hybridizing with the non-target strand of the target DNA sequence, or one or more nucleotide sequences encoding the one or more guide RNAs; and

[0646] The one or more guide RNAs thereon are capable of forming a complex with the fusion protein to guide the fusion protein to bind to the target DNA sequence.

[0647] 101. The system according to Example 100, wherein the RGN comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

[0648] 102. The system according to Example 100, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

[0649] 103. The system according to Example 100, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 1-4 and 57-139.

[0650] 104. The system according to Example 100, wherein the RGN comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 1-4 and 57-139.

[0651] 105. The system according to Example 100, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NO: 1-4 and 57-139.

[0652] 106. The system according to any one of Examples 100 to 105, wherein the heterologous polypeptide is inserted into the connector domain 2, wedge domain, RuvC domain, HNH domain, Rec-2 domain or PAM interaction domain of the RGN.

[0653] 107. The system according to embodiment 106, wherein the RuvC domain is a RuvCIII domain.

[0654] 108. The system according to any one of embodiments 100 to 107, wherein:

[0655] i) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0656] A. The amino acid position corresponding to position 30 of SEQ ID NO:1;

[0657] B. The amino acid position corresponding to position 642 in SEQ ID NO:1;

[0658] C. The amino acid position corresponding to position 670 of SEQ ID NO:1;

[0659] D. The amino acid position corresponding to position 737 in SEQ ID NO:1;

[0660] E. The amino acid position corresponding to position 772 of SEQ ID NO:1;

[0661] F. The amino acid position corresponding to position 775 of SEQ ID NO:1;

[0662] G. The amino acid position corresponding to position 778 of SEQ ID NO:1; and

[0663] H. The amino acid position corresponding to position 802 in SEQ ID NO:1;

[0664] ii) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0665] A. The amino acid position corresponding to position 678 of SEQ ID NO:2;

[0666] B. The amino acid position corresponding to position 736 of SEQ ID NO:2;

[0667] C. The amino acid position corresponding to position 778 of SEQ ID NO:2;

[0668] D. The amino acid position corresponding to position 788 of SEQ ID NO:2; and

[0669] E. The amino acid position corresponding to position 922 in SEQ ID NO:2;

[0670] iii) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:3, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0671] A. The amino acid position corresponding to position 725 of SEQ ID NO:3;

[0672] B. The amino acid position corresponding to position 739 of SEQ ID NO:3; and

[0673] C. The amino acid position corresponding to position 744 in SEQ ID NO:3;

[0674] iv) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:4, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0675] A. The amino acid position corresponding to position 347 in SEQ ID NO:4;

[0676] B. The amino acid position corresponding to position 524 in SEQ ID NO:4;

[0677] C. The amino acid position corresponding to position 666 in SEQ ID NO:4;

[0678] D. The amino acid position corresponding to position 680 of SEQ ID NO:4;

[0679] E. The amino acid position corresponding to position 740 of SEQ ID NO:4;

[0680] F. The amino acid position corresponding to position 785 of SEQ ID NO:4;

[0681] G. The amino acid position corresponding to position 910 of SEQ ID NO:4; and

[0682] H. The amino acid position corresponding to position 1077 of SEQ ID NO:4; or

[0683] v) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:131, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0684] A. The amino acid position corresponding to position 766 of SEQ ID NO: 131; and

[0685] B. The amino acid position corresponding to position 806 of SEQ ID NO:131.

[0686] 109. The system according to any one of Examples 100 to 108, wherein the heterologous polypeptide comprises a base-editing polypeptide or a lead-editing polypeptide.

[0687] 110. The system according to Example 109, wherein the lead editing polypeptide comprises a DNA polymerase.

[0688] 111. The system according to Example 109, wherein the lead editing polypeptide comprises reverse transcriptase.

[0689] 112. The system according to Example 109, wherein the base-editing polypeptide comprises a deaminase.

[0690] 113. The system according to Example 112, wherein the deaminase is cytosine deaminase or adenine deaminase.

[0691] 114. The system according to Example 112 or 113, wherein the fusion protein is a base editor fusion protein, and wherein the base editor fusion protein has improved editing activity compared to a parental base editor comprising a deaminase fused to the amino terminus of the RGN.

[0692] 115. The system according to any one of Examples 112 to 114, wherein the at least one heterologous polypeptide is inserted immediately after an amino acid position selected from the group consisting of the following into the RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1:

[0693] a) The amino acid position corresponding to position 642 in SEQ ID NO:1;

[0694] b) The amino acid position corresponding to position 670 of SEQ ID NO:1;

[0695] c) The amino acid position corresponding to position 737 in SEQ ID NO:1;

[0696] d) The amino acid position corresponding to position 772 of SEQ ID NO:1;

[0697] e) The amino acid position corresponding to position 775 of SEQ ID NO: 1; and

[0698] f) The amino acid position corresponding to position 778 of SEQ ID NO: 1; and

[0699] The fusion protein therein is a base editor fusion protein, and the base editor fusion protein has a shifted editing window compared to the parental base editor containing the deaminase fused to the amino terminus of the RGN.

[0700] 116. The system according to any one of Examples 112 to 115, wherein the deaminase lacks a first amino acid residue and a last amino acid residue compared to the parental deaminase from which the deaminase is derived.

[0701] 117. The system according to any one of Examples 112 to 116, wherein the deaminase comprises adenine deaminase, and the fusion protein is an adenosine base editor (ABE) fusion protein.

[0702] 118. The system according to Example 117, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

[0703] 119. The system according to Example 117, wherein the adenine deaminase comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

[0704] 120. The system according to Example 117, wherein the adenine deaminase comprises the amino acid sequence of any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

[0705] 121. The system according to Example 117, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:5 or 6.

[0706] 122. The system according to Example 117, wherein the adenine deaminase comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO:5 or 6.

[0707] 123. The system according to Example 117, wherein the adenine deaminase comprises the amino acid sequence of SEQ ID NO: 5 or 6.

[0708] 124. The system according to Example 117, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

[0709] 125. The system according to Example 117, wherein the ABE fusion protein comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

[0710] 126. The system according to Example 117, wherein the ABE fusion protein comprises the amino acid sequence shown in any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

[0711] 127. The system according to Example 117, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parental ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN.

[0712] 128. The system according to Example 117, wherein the ABE fusion protein comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parental ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN.

[0713] 129. The system according to Example 117, wherein the ABE fusion protein comprises the amino acid sequence shown in any one of SEQ ID NO: 8-10 and 14-16.

[0714] 130. The system according to any one of Examples 100 to 129, wherein the RGN is a nicking enzyme or a nuclease-inactive RGN.

[0715] 131. The system according to Example 130, wherein the RGN nicking enzyme comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO:49-56, 698 and 699, and retains nicking enzyme activity.

[0716] 132. The system according to Example 130, wherein the RGN nickase comprises the amino acid sequence of any one of SEQ ID NO:49-56, 698 and 699.

[0717] 133. The system according to any one of Examples 100 to 132, wherein the RGN and the heterologous polypeptide are directly fused to each other without a linker sequence.

[0718] 134. The system according to any one of Examples 100 to 133, wherein the fusion protein further comprises at least one nuclear localization signal (NLS).

[0719] 135. The system according to any one of Examples 100 to 134, wherein the nucleotide sequence encoding the fusion protein is codon-optimized for expression in eukaryotic cells.

[0720] 136. The system according to any one of Examples 100 to 135, wherein at least one of the nucleotide sequence encoding the one or more guide RNAs and the nucleotide sequence encoding the fusion protein is operatively linked to a promoter heterologous to the nucleotide sequence.

[0721] 137. The system according to any one of Examples 100 to 136, wherein the target DNA sequence is a eukaryotic target DNA sequence.

[0722] 138. The system according to any one of Examples 100 to 137, wherein the target DNA sequence is located adjacent to the prespacer sequence neighbor motif (PAM) recognized by the RGN.

[0723] 139. The system according to any one of Examples 100 to 138, wherein the nucleotide sequence encoding the one or more guide RNAs and the nucleotide sequence encoding the fusion protein are located on a vector.

[0724] 140. A cell comprising the system according to any one of Examples 100 to 139 and 256 to 267.

[0725] 141. The cell according to Example 140, wherein the cell is a eukaryotic cell.

[0726] 142. The cell according to Example 141, wherein the eukaryotic cell is a mammalian cell.

[0727] 143. The cell according to Example 142, wherein the mammalian cell is a human cell.

[0728] 144. The cell according to Example 143, wherein the human cell is an immune cell.

[0729] 145. The cell according to Example 144, wherein the immune cell is a stem cell.

[0730] 146. The cell according to Example 145, wherein the stem cell is an induced pluripotent stem cell.

[0731] 147. The cell according to Example 141, wherein the eukaryotic cell is an insect cell.

[0732] 148. The cell according to Example 141, wherein the eukaryotic cell is a plant cell.

[0733] 149. A plant comprising cells as described in Example 148.

[0734] 150. A seed comprising cells as described in Example 148.

[0735] 151. The cell according to Example 140, wherein the cell is a prokaryotic cell.

[0736] 152. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a system according to any one of Examples 100 to 139 and 256 to 267 or a cell according to any one of Examples 142 to 146.

[0737] 153. A method for localizing a heterologous polypeptide to a target DNA molecule containing a target DNA sequence, the method comprising delivering a system according to any one of Examples 100 to 139 and 256 to 267 to the target DNA molecule or a cell containing the target DNA molecule.

[0738] 154. A method for modifying a target DNA molecule containing a target DNA sequence, the method comprising delivering a system according to any one of Examples 100 to 139 and 256 to 267 to the target DNA molecule or a cell containing the target DNA molecule.

[0739] 155. The method according to Example 154, wherein the deaminase comprises adenine deaminase, and wherein the modified target DNA molecule comprises an A>N mutation of at least one nucleotide within the target DNA molecule, wherein N is C, G, or T.

[0740] 156. The method according to Example 155, wherein the modified target DNA molecule comprises an A>G mutation of at least one nucleotide within the target DNA molecule.

[0741] 157. A method for localizing a heterologous polypeptide to a target DNA molecule containing a target DNA sequence, the method comprising:

[0742] a) The ribonucleotide complex is assembled in vitro by combining the following items under conditions suitable for its formation:

[0743] i) one or more guide RNAs capable of hybridizing with the non-target strand of the target DNA sequence; and

[0744] ii) A fusion protein comprising an RNA-guided nuclease polypeptide (RGN) and at least one heteropeptide inserted therein, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

[0745] as well as

[0746] b) Contact the target DNA molecule or a cell containing the target DNA molecule with the in vitro assembled ribonucleotide complex;

[0747] The one or more guide RNAs hybridize with the non-target strand of the target DNA sequence, thereby guiding the fusion protein to bind to the target DNA sequence.

[0748] 158. The method according to Example 157, wherein the RGN comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

[0749] 159. The method according to Example 157, wherein the RGN comprises an amino acid sequence having an amino acid sequence having any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

[0750] 160. The method according to Example 157, wherein the RGN contains at least 90% sequence identity with any one of SEQ ID NO: 1-4 and 57-139.

[0751] 161. The method according to Example 157, wherein the RGN comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO:1-4 and 57-139.

[0752] 162. The method according to Example 157, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NO: 1-4 and 57-139.

[0753] 163. The method according to any one of Examples 157 to 162, wherein the heterologous polypeptide is inserted into the linker domain 2, wedge domain, RuvC domain, HNH domain, Rec-2 domain or PAM interaction domain of the RGN.

[0754] 164. The method according to embodiment 163, wherein the RuvC domain is a RuvCIII domain.

[0755] 165. The method according to any one of Examples 157 to 164, wherein:

[0756] a) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0757] I) The amino acid position corresponding to position 30 of SEQ ID NO:1;

[0758] II) The amino acid position corresponding to position 642 in SEQ ID NO:1;

[0759] III) The amino acid position corresponding to position 670 of SEQ ID NO:1;

[0760] IV) The amino acid position corresponding to position 737 in SEQ ID NO:1;

[0761] V) The amino acid position corresponding to position 772 of SEQ ID NO:1;

[0762] VI) The amino acid position corresponding to position 775 of SEQ ID NO:1;

[0763] VII) The amino acid position corresponding to position 778 of SEQ ID NO:1; and

[0764] VIII) The amino acid position corresponding to position 802 in SEQ ID NO:1;

[0765] b) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0766] I) The amino acid position corresponding to position 678 of SEQ ID NO:2;

[0767] II) The amino acid position corresponding to position 736 of SEQ ID NO:2;

[0768] III) The amino acid position corresponding to position 778 of SEQ ID NO:2;

[0769] IV) The amino acid position corresponding to position 788 of SEQ ID NO:2; and

[0770] V) The amino acid position corresponding to position 922 in SEQ ID NO:2;

[0771] c) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:3, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0772] I) The amino acid position corresponding to position 725 of SEQ ID NO:3;

[0773] II) The amino acid position corresponding to position 739 in SEQ ID NO:3;

[0774] III) The amino acid position corresponding to position 744 of SEQ ID NO:3;

[0775] d) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:4, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0776] I) The amino acid position corresponding to position 347 in SEQ ID NO:4;

[0777] II) The amino acid position corresponding to position 524 in SEQ ID NO:4;

[0778] III) The amino acid position corresponding to position 666 in SEQ ID NO:4;

[0779] IV) The amino acid position corresponding to position 680 of SEQ ID NO:4;

[0780] V) The amino acid position corresponding to position 740 of SEQ ID NO:4;

[0781] VI) The amino acid position corresponding to position 785 of SEQ ID NO:4;

[0782] VII) The amino acid position corresponding to position 910 of SEQ ID NO:4; and

[0783] VIII) The amino acid position corresponding to position 1077 of SEQ ID NO:4; or

[0784] e) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:131, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of:

[0785] I) The amino acid position corresponding to position 766 of SEQ ID NO: 131; and

[0786] II) The amino acid position corresponding to position 806 of SEQ ID NO:131.

[0787] 166. The method according to any one of Examples 157 to 165, wherein the heterologous polypeptide comprises a base-editing polypeptide or a lead-editing polypeptide, and wherein the method further comprises modifying the target DNA molecule to generate a modified target DNA molecule.

[0788] 167. The method according to Example 166, wherein the target DNA molecule contains a causal mutation of a disease or condition, and wherein the modification of the target DNA molecule corrects the causal mutation.

[0789] 168. The method according to Example 167, wherein the correction of the causal mutation includes the correction of nonsense mutations.

[0790] 169. The method according to any one of Examples 166 to 168, wherein the lead editing polypeptide comprises a DNA polymerase.

[0791] 170. The method according to any one of Examples 166 to 168, wherein the lead editing polypeptide comprises a reverse transcriptase.

[0792] 171. The method according to any one of Examples 166 to 168, wherein the base-editing polypeptide comprises a deaminase.

[0793] 172. The method according to Example 171, wherein the deaminase is cytosine deaminase or adenine deaminase.

[0794] 173. The method according to Example 171 or 172, wherein the fusion protein is a base editor fusion protein, and wherein the base editor fusion protein has improved editing activity compared to a parental base editor comprising a deaminase fused to the amino terminus of the RGN.

[0795] 174. The method according to any one of Examples 171 to 173, wherein the at least one heterologous polypeptide is inserted into the RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1 immediately following an amino acid position selected from the group consisting of:

[0796] a) The amino acid position corresponding to position 642 in SEQ ID NO:1;

[0797] b) The amino acid position corresponding to position 670 of SEQ ID NO:1;

[0798] c) The amino acid position corresponding to position 737 in SEQ ID NO:1;

[0799] d) The amino acid position corresponding to position 772 of SEQ ID NO:1;

[0800] e) The amino acid position corresponding to position 775 of SEQ ID NO: 1; and

[0801] f) The amino acid position corresponding to position 778 of SEQ ID NO: 1; and

[0802] The fusion protein therein is a base editor fusion protein, and the base editor fusion protein has a shifted editing window compared to the parental base editor containing the deaminase fused to the amino terminus of the RGN.

[0803] 175. The method according to any one of Examples 171 to 174, wherein the deaminase lacks a first amino acid residue and a last amino acid residue compared to the parental deaminase from which the deaminase is derived.

[0804] 176. The method according to any one of Examples 171 to 175, wherein the deaminase comprises adenine deaminase, and the fusion protein is an adenosine base editor (ABE) fusion protein.

[0805] 177. The method according to Example 176, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

[0806] 178. The method according to Example 176, wherein the adenine deaminase comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

[0807] 179. The method according to Example 176, wherein the adenine deaminase comprises the amino acid sequence of any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

[0808] 180. The method according to Example 176, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:5 or 6.

[0809] 181. The method according to Example 176, wherein the adenine deaminase comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO:5 or 6.

[0810] 182. The method according to Example 176, wherein the adenine deaminase comprises the amino acid sequence of SEQ ID NO: 5 or 6.

[0811] 183. The method according to Example 176, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

[0812] 184. The method according to Example 176, wherein the ABE fusion protein comprises an amino acid sequence having at ...

Claims

1. A fusion protein comprising an RNA-guided nuclease (RGN) and at least one heterologous polypeptide inserted therein, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

2. The fusion protein according to claim 1, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

3. The fusion protein according to claim 1 or 2, wherein the heterologous polypeptide is inserted into the linker domain 2, wedge domain, RuvC domain, HNH domain, Rec-2 domain or PAM interaction domain of the RGN.

4. The fusion protein according to any one of claims 1 to 3, wherein: a) The RGN contains an amino acid sequence that has at least 90% sequence identity with SEQ ID NO:1, and The at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: i) The amino acid position corresponding to position 30 of SEQ ID NO:1; ii) The amino acid position corresponding to position 642 in SEQ ID NO:1; iii) The amino acid position corresponding to position 670 of SEQ ID NO:1; iv) The amino acid position corresponding to position 737 of SEQ ID NO:1; v) The amino acid position corresponding to position 772 of SEQ ID NO:1; vi) The amino acid position corresponding to position 775 of SEQ ID NO:1; vii) The amino acid position corresponding to position 778 of SEQ ID NO:1; as well as viii) The amino acid position corresponding to position 802 in SEQ ID NO:1; b) The RGN contains an amino acid sequence that has at least 90% sequence identity with SEQ ID NO:2, and The at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: i) The amino acid position corresponding to position 678 of SEQ ID NO:2; ii) The amino acid position corresponding to position 736 of SEQ ID NO:2; iii) The amino acid position corresponding to position 778 of SEQ ID NO:2; iv) The amino acid position corresponding to position 788 of SEQ ID NO:2; and v) The amino acid position corresponding to position 922 in SEQ ID NO:2; c) The RGN contains an amino acid sequence that has at least 90% sequence identity with SEQ ID NO:3, and The at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: i) The amino acid position corresponding to position 725 of SEQ ID NO:3; ii) The amino acid position corresponding to position 739 in SEQ ID NO:3; and iii) The amino acid position corresponding to position 744 of SEQ ID NO:3; d) The RGN contains an amino acid sequence that has at least 90% sequence identity with SEQ ID NO:4, and The at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: i) The amino acid position corresponding to position 347 in SEQ ID NO:4; ii) The amino acid position corresponding to position 524 in SEQ ID NO:4; iii) The amino acid position corresponding to position 666 in SEQ ID NO:4; iv) The amino acid position corresponding to position 680 of SEQ ID NO:4; v) The amino acid position corresponding to position 740 of SEQ ID NO:4; vi) The amino acid position corresponding to position 785 of SEQ ID NO:4; vii) The amino acid position corresponding to position 910 of SEQ ID NO:4; and viii) The amino acid position corresponding to position 1077 of SEQ ID NO:4; or e) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:131, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: i) The amino acid position corresponding to position 766 of SEQ ID NO: 131; and ii) The amino acid position corresponding to position 806 of SEQ ID NO:

131.

5. The fusion protein according to any one of claims 1 to 4, wherein the heterologous polypeptide comprises a base-editing polypeptide or a lead-editing polypeptide.

6. The fusion protein of claim 5, wherein the base-editing polypeptide comprises a deaminase.

7. The fusion protein of claim 6, wherein the fusion protein is a base editor fusion protein, and wherein the base editor fusion protein has improved editing activity compared to a parental base editor comprising a deaminase fused to the amino terminus of the RGN.

8. The fusion protein according to claim 6 or 7, wherein the at least one heteropeptide is inserted immediately after an amino acid position selected from the group consisting of into an RGN containing an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1: a) The amino acid position corresponding to position 642 in SEQ ID NO:1; b) The amino acid position corresponding to position 670 of SEQ ID NO:1; c) The amino acid position corresponding to position 737 in SEQ ID NO:1; d) The amino acid position corresponding to position 772 of SEQ ID NO:1; e) The amino acid position corresponding to position 775 of SEQ ID NO: 1; and f) The amino acid position corresponding to position 778 of SEQ ID NO:1; and The fusion protein therein is a base editor fusion protein, and the base editor fusion protein has a shifted editing window compared to the parental base editor containing the deaminase fused to the amino terminus of the RGN.

9. The fusion protein according to any one of claims 6 to 8, wherein the deaminase lacks a first amino acid residue and a last amino acid residue compared to the parental deaminase from which the deaminase is derived.

10. The fusion protein according to any one of claims 6 to 9, wherein the deaminase comprises adenine deaminase, and the fusion protein is an adenine base editor (ABE) fusion protein.

11. The fusion protein of claim 10, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

12. The fusion protein of claim 10, wherein the adenine deaminase comprises the amino acid sequence of any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

13. The fusion protein of claim 10, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

14. The fusion protein of claim 10, wherein the ABE fusion protein comprises the amino acid sequence shown in any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

15. The fusion protein of claim 10, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parental ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN.

16. The fusion protein according to any one of claims 1 to 15, wherein the RGN is a nicking enzyme or a nuclease-inactive RGN.

17. The fusion protein of claim 16, wherein the RGN nickase comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO:49-56, 698 and 699, and retains nickase activity.

18. The fusion protein of claim 16, wherein the RGN nickase comprises the amino acid sequence of any one of SEQ ID NO:49-56, 698, and 699.

19. A ribonucleoprotein (RNP) complex comprising a fusion protein according to any one of claims 1-18, 110 and 111 and a guide RNA bound to said fusion protein.

20. A nucleic acid molecule comprising a polynucleotide encoding a fusion protein, wherein the fusion protein comprises an RNA-guided nuclease (RGN) and at least one heterologous polypeptide inserted therein, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

21. The nucleic acid molecule of claim 20, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698, and 699.

22. The nucleic acid molecule according to claim 20 or 21, wherein the heterologous polypeptide is inserted into the linker domain 2, wedge domain, RuvC domain, HNH domain, Rec-2 domain or PAM interaction domain of the RGN.

23. The nucleic acid molecule according to any one of claims 20 to 22, wherein: a) The RGN contains an amino acid sequence that has at least 90% sequence identity with SEQ ID NO:1, and The at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: i) The amino acid position corresponding to position 30 of SEQ ID NO:1; ii) The amino acid position corresponding to position 642 in SEQ ID NO:1; iii) The amino acid position corresponding to position 670 of SEQ ID NO:1; iv) The amino acid position corresponding to position 737 of SEQ ID NO:1; v) The amino acid position corresponding to position 772 of SEQ ID NO:1; vi) The amino acid position corresponding to position 775 of SEQ ID NO:1; vii) The amino acid position corresponding to position 778 of SEQ ID NO:1; as well as viii) The amino acid position corresponding to position 802 in SEQ ID NO:1; b) The RGN contains an amino acid sequence that has at least 90% sequence identity with SEQ ID NO:2, and The at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: i) The amino acid position corresponding to position 678 of SEQ ID NO:2; ii) The amino acid position corresponding to position 736 of SEQ ID NO:2; iii) The amino acid position corresponding to position 778 of SEQ ID NO:2; iv) The amino acid position corresponding to position 788 of SEQ ID NO:2; and v) The amino acid position corresponding to position 922 in SEQ ID NO:2; c) The RGN contains an amino acid sequence that has at least 90% sequence identity with SEQ ID NO:3, and The at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: i) The amino acid position corresponding to position 725 of SEQ ID NO:3; ii) The amino acid position corresponding to position 739 in SEQ ID NO:3; and iii) The amino acid position corresponding to position 744 of SEQ ID NO:3; d) The RGN contains an amino acid sequence that has at least 90% sequence identity with SEQ ID NO:4, and The at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: i) The amino acid position corresponding to position 347 in SEQ ID NO:4; ii) The amino acid position corresponding to position 524 in SEQ ID NO:4; iii) The amino acid position corresponding to position 666 in SEQ ID NO:4; iv) The amino acid position corresponding to position 680 of SEQ ID NO:4; v) The amino acid position corresponding to position 740 of SEQ ID NO:4; vi) The amino acid position corresponding to position 785 of SEQ ID NO:4; vii) The amino acid position corresponding to position 910 of SEQ ID NO:4; and viii) The amino acid position corresponding to position 1077 of SEQ ID NO:4; or e) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:131, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: i) The amino acid position corresponding to position 766 of SEQ ID NO: 131; and ii) The amino acid position corresponding to position 806 of SEQ ID NO:

131.

24. The nucleic acid molecule according to any one of claims 20 to 23, wherein the heterologous polypeptide comprises a base-editing polypeptide or a lead-editing polypeptide.

25. The nucleic acid molecule of claim 24, wherein the base-editing polypeptide comprises a deaminase.

26. The nucleic acid molecule of claim 25, wherein the fusion protein is a base editor fusion protein, and wherein the base editor fusion protein has improved editing activity compared to a parental base editor comprising a deaminase fused to the amino terminus of the RGN.

27. The nucleic acid molecule according to claim 25 or 26, wherein the at least one heteropeptide is inserted immediately after an amino acid position selected from the group consisting of the following into the RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1: a) The amino acid position corresponding to position 642 in SEQ ID NO:1; b) The amino acid position corresponding to position 670 of SEQ ID NO:1; c) The amino acid position corresponding to position 737 in SEQ ID NO:1; d) The amino acid position corresponding to position 772 of SEQ ID NO:1; e) The amino acid position corresponding to position 775 of SEQ ID NO: 1; and f) The amino acid position corresponding to position 778 of SEQ ID NO:1; and The fusion protein therein is a base editor fusion protein, and the base editor fusion protein has a shifted editing window compared to the parental base editor containing the deaminase fused to the amino terminus of the RGN.

28. The nucleic acid molecule according to any one of claims 25 to 27, wherein the deaminase lacks a first amino acid residue and a last amino acid residue compared to the parental deaminase from which the deaminase is derived.

29. The nucleic acid molecule according to any one of claims 25 to 28, wherein the deaminase comprises adenine deaminase and is an adenosine base editor (ABE) fusion protein.

30. The nucleic acid molecule of claim 29, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

31. The nucleic acid molecule of claim 29, wherein the adenine deaminase comprises the amino acid sequence of any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

32. The nucleic acid molecule of claim 29, wherein the fusion protein comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

33. The nucleic acid molecule of claim 29, wherein the ABE fusion protein comprises the amino acid sequence shown in any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

34. The nucleic acid molecule of claim 29, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parental ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN.

35. The nucleic acid molecule according to any one of claims 20 to 34, wherein the RGN is a nicking enzyme or a nuclease-inactive RGN.

36. The nucleic acid molecule of claim 35, wherein the RGN nickase comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO:49-56, 698 and 699, and retains nickase activity.

37. The nucleic acid molecule of claim 35, wherein the RGN nickase comprises the amino acid sequence of any one of SEQ ID NO:49-56, 698 and 699.

38. A vector comprising a nucleic acid molecule according to any one of claims 20 to 37 and 120 to 126.

39. The vector of claim 38, further comprising at least one nucleotide sequence encoding a guide RNA capable of hybridizing with a non-target strand of the target sequence of the RGN.

40. A cell comprising a fusion protein according to any one of claims 1 to 18 and 113 to 119 or an RNP complex according to claim 19.

41. A cell comprising a fusion protein according to any one of claims 1 to 18 and 113 to 119, wherein the cell further comprises guide RNA.

42. A cell comprising a nucleic acid molecule according to any one of claims 20 to 37 and 120 to 126 or a vector according to claim 38 or 39.

43. The cell according to any one of claims 40 to 42, wherein the cell is a mammalian cell.

44. The cell according to any one of claims 40 to 42, wherein the cell is a plant cell.

45. A plant or a seed comprising the cells according to claim 44.

46. ​​A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a fusion protein according to any one of claims 1 to 18 and 113 to 119, an RNP complex according to claim 19, a nucleic acid molecule according to any one of claims 20 to 37 and 120 to 126, a carrier according to claim 38 or 39, or a cell according to claim 43.

47. A method for manufacturing an RGN fusion ribonucleoprotein complex, the method comprising introducing a nucleic acid molecule according to any one of claims 20 to 37 and 120 to 126, and a nucleic acid molecule comprising an expression cassette encoding guide RNA, or a vector according to claim 38 or 39, into a cell, and culturing the cell under conditions where the fusion protein and gRNA are expressed and an RGN fusion ribonucleoprotein complex is formed.

48. A system for targeting a heterologous polypeptide to a target DNA molecule containing a target DNA sequence, the system comprising: a) a fusion protein or a nucleotide sequence encoding said fusion protein, wherein said fusion protein comprises an RNA-guided nuclease (RGN) and at least one heterologous polypeptide inserted therein; wherein said RGN comprises the same as SEQ ID NO: Amino acid sequences having at least 90% sequence identity among any one of 1-4, 49-162, 435, 575, 576, 698, and 699; and b) One or more guide RNAs capable of hybridizing with the non-target strand of the target DNA sequence, or one or more nucleotide sequences encoding the one or more guide RNAs; and The one or more guide RNAs thereon are capable of forming a complex with the fusion protein to guide the fusion protein to bind to the target DNA sequence.

49. The system of claim 48, wherein the RGN comprises the amino acid sequence of any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698, and 699.

50. The system according to any claim 48 or 49, wherein the heterologous polypeptide is inserted into the connector domain 2, wedge domain, RuvC domain, HNH domain, Rec-2 domain or PAM interaction domain of the RGN.

51. The system according to any one of claims 48 to 50, wherein: i) The RGN contains an amino acid sequence that has at least 90% sequence identity with SEQ ID NO:1, and The at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: A. The amino acid position corresponding to position 30 of SEQ ID NO:1; B. The amino acid position corresponding to position 642 in SEQ ID NO:1; C. The amino acid position corresponding to position 670 of SEQ ID NO:1; D. The amino acid position corresponding to position 737 in SEQ ID NO:1; E. The amino acid position corresponding to position 772 of SEQ ID NO:1; F. The amino acid position corresponding to position 775 of SEQ ID NO:1; G. The amino acid position corresponding to position 778 of SEQ ID NO:1; and H. The amino acid position corresponding to position 802 in SEQ ID NO:1; ii) The RGN contains an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2, and The at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: A. The amino acid position corresponding to position 678 of SEQ ID NO:2; B. The amino acid position corresponding to position 736 of SEQ ID NO:2; C. The amino acid position corresponding to position 778 of SEQ ID NO:2; D. The amino acid position corresponding to position 788 of SEQ ID NO:2; and E. The amino acid position corresponding to position 922 in SEQ ID NO:2; iii) The RGN contains an amino acid sequence that has at least 90% sequence identity with SEQ ID NO:3, and The at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: A. The amino acid position corresponding to position 725 of SEQ ID NO:3; B. The amino acid position corresponding to position 739 of SEQ ID NO:3; and C. The amino acid position corresponding to position 744 in SEQ ID NO:3; iv) The RGN contains an amino acid sequence that has at least 90% sequence identity with SEQ ID NO:4, and The at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: A. The amino acid position corresponding to position 347 in SEQ ID NO:4; B. The amino acid position corresponding to position 524 in SEQ ID NO:4; C. The amino acid position corresponding to position 666 in SEQ ID NO:4; D. The amino acid position corresponding to position 680 of SEQ ID NO:4; E. The amino acid position corresponding to position 740 of SEQ ID NO:4; F. The amino acid position corresponding to position 785 of SEQ ID NO:4; G. The amino acid position corresponding to position 910 of SEQ ID NO:4; and H. The amino acid position corresponding to position 1077 of SEQ ID NO:4; or v) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:131, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: A. The amino acid position corresponding to position 766 of SEQ ID NO: 131; and B. The amino acid position corresponding to position 806 of SEQ ID NO:

131.

52. The system according to any one of claims 48 to 51, wherein the heterologous polypeptide comprises a base-editing polypeptide or a lead-editing polypeptide.

53. The system of claim 52, wherein the base-editing polypeptide comprises a deaminase.

54. The system of claim 53, wherein the fusion protein is a base editor fusion protein, and wherein the base editor fusion protein has improved editing activity compared to a parental base editor comprising a deaminase fused to the amino terminus of the RGN.

55. The system according to claim 53 or 54, wherein the at least one heterologous polypeptide is inserted immediately after an amino acid position selected from the group consisting of the following into the RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1: a) The amino acid position corresponding to position 642 in SEQ ID NO:1; b) The amino acid position corresponding to position 670 of SEQ ID NO:1; c) The amino acid position corresponding to position 737 in SEQ ID NO:1; d) The amino acid position corresponding to position 772 of SEQ ID NO:1; e) The amino acid position corresponding to position 775 of SEQ ID NO: 1; and f) The amino acid position corresponding to position 778 of SEQ ID NO:1; and The fusion protein therein is a base editor fusion protein, and the base editor fusion protein has a shifted editing window compared to the parental base editor containing the deaminase fused to the amino terminus of the RGN.

56. The system according to any one of claims 53 to 55, wherein the deaminase lacks a first amino acid residue and a last amino acid residue compared to the parental deaminase from which the deaminase is derived.

57. The system according to any one of claims 53 to 56, wherein the deaminase comprises adenine deaminase, and the fusion protein is an adenosine base editor (ABE) fusion protein.

58. The system of claim 57, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

59. The system of claim 57, wherein the adenine deaminase comprises the amino acid sequence of any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

60. The system of claim 57, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

61. The system of claim 57, wherein the ABE fusion protein comprises the amino acid sequence shown in any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

62. The system of claim 57, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parental ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN.

63. The system according to any one of claims 48 to 62, wherein the RGN is a nicking enzyme or a nuclease-inactive RGN.

64. The system of claim 63, wherein the RGN nicking enzyme comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO:49-56, 698 and 699, and retains nicking enzyme activity.

65. The system of claim 63, wherein the RGN nickase comprises the amino acid sequence of any one of SEQ ID NO:49-56, 698, and 699.

66. The system according to any one of claims 48 to 65, wherein at least one of the nucleotide sequence encoding the one or more guide RNAs and the nucleotide sequence encoding the fusion protein is operatively linked to a promoter heterologous to the nucleotide sequence.

67. The system according to any one of claims 48 to 66, wherein the target DNA sequence is a eukaryotic target DNA sequence.

68. The system according to any one of claims 48 to 67, wherein the target DNA sequence is located adjacent to the prespacer neighbor motif (PAM) recognized by the RGN.

69. The system according to any one of claims 48 to 68, wherein the nucleotide sequence encoding the one or more guide RNAs and the nucleotide sequence encoding the fusion protein are located on a vector.

70. A cell comprising the system according to any one of claims 48 to 69, 117 and 118.

71. The cell of claim 70, wherein the cell is a mammalian cell.

72. The cell of claim 70, wherein the cell is a plant cell.

73. A plant or a seed comprising the cells according to claim 72.

74. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and the system according to any one of claims 48 to 69 and 127 to 133 or the cell according to claim 71.

75. A method for targeting a heterologous polypeptide to a target DNA molecule containing a target DNA sequence, the method comprising delivering a system according to any one of claims 48 to 69 and 127 to 133 to the target DNA molecule or a cell containing the target DNA molecule.

76. A method for modifying a target DNA molecule containing a target DNA sequence, the method comprising delivering a system according to any one of claims 48 to 69 and 127 to 133 to the target DNA molecule or a cell containing the target DNA molecule.

77. A method for localizing a heterologous polypeptide to a target DNA molecule containing a target DNA sequence, the method comprising: a) The ribonucleotide complex is assembled in vitro by combining the following items under conditions suitable for its formation: i) one or more guide RNAs capable of hybridizing with the non-target strand of the target DNA sequence; and ii) A fusion protein comprising an RNA-guided nuclease polypeptide (RGN) and at least one heteropeptide inserted therein, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698, and 699; and b) Contact the target DNA molecule or a cell containing the target DNA molecule with the in vitro assembled ribonucleotide complex; The one or more guide RNAs hybridize with the non-target strand of the target DNA sequence, thereby guiding the fusion protein to bind to the target DNA sequence.

78. The method of claim 77, wherein the RGN comprises an amino acid sequence having an amino acid sequence having any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

79. The method according to claim 77 or 78, wherein the heterologous polypeptide is inserted into the linker domain 2, wedge domain, RuvC domain, HNH domain, Rec-2 domain or PAM interaction domain of the RGN.

80. The method according to any one of claims 77 to 79, wherein: a) The RGN contains an amino acid sequence that has at least 90% sequence identity with SEQ ID NO:1, and The at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: I) The amino acid position corresponding to position 30 of SEQ ID NO:1; II) The amino acid position corresponding to position 642 in SEQ ID NO:1; III) The amino acid position corresponding to position 670 of SEQ ID NO:1; IV) The amino acid position corresponding to position 737 in SEQ ID NO:1; W) The amino acid position corresponding to position 772 of SEQ ID NO:1; VI) The amino acid position corresponding to position 775 of SEQ ID NO:1; VII) The amino acid position corresponding to position 778 of SEQ ID NO:1; as well as VIII) The amino acid position corresponding to position 802 in SEQ ID NO:1; b) The RGN contains an amino acid sequence that has at least 90% sequence identity with SEQ ID NO:2, and The at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: I) The amino acid position corresponding to position 678 of SEQ ID NO:2; II) The amino acid position corresponding to position 736 of SEQ ID NO:2; III) The amino acid position corresponding to position 778 of SEQ ID NO:2; IV) The amino acid position corresponding to position 788 of SEQ ID NO:2; and V) The amino acid position corresponding to position 922 in SEQ ID NO:2; c) The RGN contains an amino acid sequence that has at least 90% sequence identity with SEQ ID NO:3, and The at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: I) The amino acid position corresponding to position 725 of SEQ ID NO:3; II) The amino acid position corresponding to position 739 in SEQ ID NO:3; III) The amino acid position corresponding to position 744 of SEQ ID NO:3; d) The RGN contains an amino acid sequence that has at least 90% sequence identity with SEQ ID NO:4, and The at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: I) The amino acid position corresponding to position 347 in SEQ ID NO:4; II) The amino acid position corresponding to position 524 in SEQ ID NO:4; III) The amino acid position corresponding to position 666 in SEQ ID NO:4; IV) The amino acid position corresponding to position 680 of SEQ ID NO:4; V) The amino acid position corresponding to position 740 of SEQ ID NO:4; VI) The amino acid position corresponding to position 785 of SEQ ID NO:4; VII) The amino acid position corresponding to position 910 of SEQ ID NO:4; and VIII) The amino acid position corresponding to position 1077 of SEQ ID NO:4; or e) The RGN comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:131, and the at least one heteropeptide is inserted into the RGN immediately after an amino acid position selected from the group consisting of: I) The amino acid position corresponding to position 766 of SEQ ID NO: 131; and II) The amino acid position corresponding to position 806 of SEQ ID NO:

131.

81. The method according to any one of claims 77 to 80, wherein the heterologous polypeptide comprises a base-editing polypeptide or a lead-editing polypeptide, and wherein the method further comprises modifying the target DNA molecule to generate a modified target DNA molecule.

82. The method of claim 81, wherein the target DNA molecule comprises a causal mutation of a disease or condition, and wherein the modification of the target DNA molecule corrects the causal mutation.

83. The method of claim 82, wherein the correction of the causal mutation includes the correction of nonsense mutations.

84. The method according to any one of claims 81 to 83, wherein the base-editing polypeptide comprises a deaminase.

85. The method of claim 84, wherein the fusion protein is a base editor fusion protein, and wherein the base editor fusion protein has improved editing activity compared to a parental base editor comprising a deaminase fused to the amino terminus of the RGN.

86. The method according to claim 84 or 85, wherein the at least one heteropeptide is inserted immediately after an amino acid position selected from the group consisting of the following into the RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1: a) The amino acid position corresponding to position 642 in SEQ ID NO:1; b) The amino acid position corresponding to position 670 of SEQ ID NO:1; c) The amino acid position corresponding to position 737 in SEQ ID NO:1; d) The amino acid position corresponding to position 772 of SEQ ID NO:1; e) The amino acid position corresponding to position 775 of SEQ ID NO: 1; and f) The amino acid position corresponding to position 778 of SEQ ID NO:1; and The fusion protein therein is a base editor fusion protein, and the base editor fusion protein has a shifted editing window compared to the parental base editor containing the deaminase fused to the amino terminus of the RGN.

87. The method according to any one of claims 84 to 86, wherein the deaminase lacks a first amino acid residue and a last amino acid residue compared to the parental deaminase from which the deaminase is derived.

88. The method according to any one of claims 84 to 87, wherein the deaminase comprises adenine deaminase, and the fusion protein is an adenosine base editor (ABE) fusion protein.

89. The method of claim 88, wherein the adenine deaminase comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

90. The method of claim 88, wherein the adenine deaminase comprises the amino acid sequence of any one of SEQ ID NO: 5, 6, 248-414, 596, 597 and 720-723.

91. The method of claim 88, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

92. The method of claim 88, wherein the ABE fusion protein comprises the amino acid sequence shown in any one of SEQ ID NO: 8-10, 13-21, 27-31, 36-38, 40-41, 43-48, 647 and 650.

93. The method of claim 88, wherein the ABE fusion protein comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 8-10 and 14-16, and wherein the ABE fusion protein has a shifted editing window compared to a parental ABE base editor comprising the adenine deaminase fused to the amino terminus of the RGN.

94. The method according to any one of claims 77 to 93, wherein the RGN is a nicking enzyme or a nuclease-inactive RGN.

95. The method of claim 94, wherein the RGN nicking enzyme comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO:49-56, 698 and 699, and retains nicking enzyme activity.

96. The method of claim 94, wherein the RGN nickase comprises the amino acid sequence of any one of SEQ ID NO:49-56, 698, and 699.

97. The method according to any one of claims 77 to 96, wherein at least one of the nucleotide sequence encoding the one or more guide RNAs and the nucleotide sequence encoding the fusion protein is operatively linked to a promoter heterologous to the nucleotide sequence.

98. The method according to any one of claims 77 to 97, wherein the target DNA sequence is a eukaryotic target DNA sequence.

99. The method according to any one of claims 77 to 98, wherein the target DNA sequence is located adjacent to the prespacer neighbor motif (PAM) recognized by the RGN.

100. The method according to any one of claims 77 to 99, wherein the target DNA molecule is intracellular.

101. The method of claim 100, further comprising selecting cells containing the modified target DNA molecule.

102. A cell comprising the modified target DNA molecule of the method according to claim 101.

103. The cell of claim 102, wherein the cell is a mammalian cell.

104. The cell of claim 102, wherein the cell is a plant cell.

105. A plant or a seed comprising the cells according to claim 104.

106. A pharmaceutical composition comprising a cellular and pharmaceutically acceptable carrier as described in claim 103.

107. A method for treating a subject suffering from a disease, condition, or illness, or at risk of developing said disease, condition, or illness, the method comprising: The subject is administered the fusion protein according to any one of claims 1 to 18 and 113 to 119, the RNP complex according to claim 19, the nucleic acid molecule according to any one of claims 20 to 37 and 120 to 126, the carrier according to claim 38 or 39, the cell according to any one of claims 43, 71 and 103, the system according to any one of claims 48 to 69 and 127 to 133, or the pharmaceutical composition according to any one of claims 46, 74 and 106.

108. The method of claim 107, wherein the disease is associated with a causal mutation, and the treatment comprises correcting the causal mutation.

109. The use of the fusion protein according to any one of claims 1 to 18 and 113 to 119, the RNP complex according to claim 19, the nucleic acid molecule according to any one of claims 20 to 37 and 120 to 126, the vector according to claim 38 or 39, the cell according to any one of claims 43, 71 and 103, or the system according to any one of claims 48 to 69 and 127 to 133 for the treatment of a subject suffering from a disease, condition or illness, or at risk of developing said disease, condition or illness.

110. The use according to claim 109, wherein the disease is associated with a causal mutation, and the treatment comprises correcting the causal mutation.

111. The use of the fusion protein according to any one of claims 1 to 18 and 113 to 119, the RNP complex according to claim 19, the nucleic acid molecule according to any one of claims 20 to 37 and 120 to 126, the vector according to claim 38 or 39, the cell according to any one of claims 43, 71 and 103, or the system according to any one of claims 48 to 69 and 127 to 133 for the manufacture of a medicament for the treatment of a disease, symptom or condition.

112. The use according to claim 111, wherein the disease is associated with a causal mutation, and an effective amount of the drug corrects the causal mutation.

113. The fusion protein of claim 10, wherein the adenine deaminase comprises an amino acid sequence having at least 85% sequence identity with SEQ ID NO:5, and comprises at least one of the following amino acid residues: a) Position C at the position corresponding to position 2 of SEQ ID NO:5; b) F or C at the position corresponding to position 22 of SEQ ID NO:5; c) Q at the position corresponding to position 23 of SEQ ID NO:5; d) N at position 35 corresponding to SEQ ID NO:5; e) L at the position corresponding to position 40 of SEQ ID NO:5; f) Q at the position corresponding to position 46 of SEQ ID NO:5; g) A or M at the position corresponding to position 68 of SEQ ID NO:5; h) H at the position corresponding to position 72 of SEQ ID NO:5; i) W, A, Q, Y or D at the position corresponding to position 75 of SEQ ID NO:5; j) G at the position corresponding to position 76 of SEQ ID NO:5; k) S at the position corresponding to position 81 of SEQ ID NO:5; l) M at the position corresponding to position 105 of SEQ ID NO:5; m) H at the position corresponding to position 108 of SEQ ID NO:5; n) L at the position corresponding to position 109 of SEQ ID NO:5; o) The I or L position corresponding to position 117 of SEQ ID NO:5; p) F at the position corresponding to position 120 of SEQ ID NO:5; q) The E, T, A, G or V at the position corresponding to position 121 of SEQ ID NO:5; r) the I or H position corresponding to position 122 of SEQ ID NO:5; s) K at the position corresponding to position 125 of SEQ ID NO:5; t) A at position 126 corresponding to SEQ ID NO:5; u) H at the position corresponding to position 135 of SEQ ID NO:5; v) V at the position corresponding to position 137 of SEQ ID NO:5; w) The position Y corresponding to position 138 of SEQ ID NO:5; x) L at the position corresponding to position 139 of SEQ ID NO:5; y) K or A at the position corresponding to position 142 of SEQ ID NO:5; z) A, Q, L or M at the position corresponding to position 145 of SEQ ID NO:5; aa) K at the position corresponding to position 148 of SEQ ID NO:5; bb) Q at the position corresponding to position 151 of SEQ ID NO:5; cc) The position E or R corresponding to position 153 of SEQ ID NO:5; dd) W at the position corresponding to position 155 of SEQ ID NO:5; The R or V at the position corresponding to position 156 of SEQ ID NO:5; ff) F at the position corresponding to position 157 of SEQ ID NO:5; The R at position 158 of SEQ ID NO:5 (gg); hh) is the position Q corresponding to position 159 of SEQ ID NO:5; ii) D or W at the position corresponding to position 160 of SEQ ID NO:5; jj) is the W, H, V or A position corresponding to position 162 of SEQ ID NO:5; kk) at position R corresponding to position 165 of SEQ ID NO:5; and ll) H at the position corresponding to position 166 of SEQ ID NO:

5.

114. The fusion protein of claim 113, wherein the adenine deaminase comprises: a) N at position 35 of SEQ ID NO:5 and W at position 162 of SEQ ID NO:5; b) N at position 35 of SEQ ID NO:5 and Q at position 46 of SEQ ID NO:5; c) S at position 81 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; d) Q at position 46 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; e) R at the position corresponding to position 156 of SEQ ID NO:5 and W at the position corresponding to position 162 of SEQ ID NO:5; f) Position A corresponding to position 68 of SEQ ID NO: 5 and position S corresponding to position 81 of SEQ ID NO: 5; g) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at position 35 of SEQ ID NO:5, Q at position 46 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; i) W at position 155 corresponding to SEQ ID NO:5, R at position 156 corresponding to SEQ ID NO:5, and W at position 162 corresponding to SEQ ID NO:5; j) N at position 35 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; k) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; l) M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; m) Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; n) M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; o) N at position 35, Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; p) S at position 81 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; q) W at position 75 corresponding to SEQ ID NO:5 and S at position 81 corresponding to SEQ ID NO:5; r) S at position 81 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; s) S at position 81 corresponding to SEQ ID NO:5 and A at position 145 corresponding to SEQ ID NO:5; t) W at position 75 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; u) W at position 155 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; v) A at position 145 of SEQ ID NO: 5 and W at position 155 of SEQ ID NO: 5; w) W at position 75 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; x) W at position 75 of SEQ ID NO:5 and A at position 145 of SEQ ID NO:5; y) Position A corresponding to position 145 of SEQ ID NO:5 and position D corresponding to position 160 of SEQ ID NO:5; z) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; aa) S at position 81 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; bb) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; cc) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; dd) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and A at position 145 corresponding to SEQ ID NO:5; (ee) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; ff) W at position 75 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; gg) W at position 75 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; hh) A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; ii) W at position 75 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; jj) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; kk) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; ll) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; mm) The positions corresponding to positions 75 and 145 of SEQ ID NO: 5 are: W, A, W, and D, respectively. oo) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; pp) L at position 40 of SEQ ID NO:5 and M at position 105 of SEQ ID NO:5; qq) L at position 40 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; rr) L at position 40 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; ss) L at the position corresponding to position 40 of SEQ ID NO:5 and Q at the position corresponding to position 159 of SEQ ID NO:5; L at position 40 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; uu) M at position 105 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; vv) M at the position corresponding to position 105 of SEQ ID NO:5 and M at the position corresponding to position 145 of SEQ ID NO:5; ww) M at position 105 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; xx) M at position 105 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; yy) V at position 121 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; zz) V at the position corresponding to position 121 of SEQ ID NO: 5 and Q at the position corresponding to position 159 of SEQ ID NO: 5; aaa) V at position 121 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; bbb) M at position 145 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; ccc) M at position 145 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; ddd) Q at position 159 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; The position S corresponding to position 81 of SEQ ID NO:5, the position M corresponding to position 145 of SEQ ID NO:5, and the position W corresponding to position 160 of SEQ ID NO:5; fff) S at position 81 corresponding to SEQ ID NO:5, V at position 121 corresponding to SEQ ID NO:5, and Q at position 159 corresponding to SEQ ID NO:5; ggg) A at position 68 of SEQ ID NO:5 and S at position 81 of SEQ ID NO:5; hhh) G at the position corresponding to position 76 of SEQ ID NO:5 and S at the position corresponding to position 81 of SEQ ID NO:5; iii) S at position 81 corresponding to SEQ ID NO:5 and L at position 117 corresponding to SEQ ID NO:5; jjj) A at position 68 of SEQ ID NO:5 and G at position 76 of SEQ ID NO:5; kkk) A at position 68 of SEQ ID NO:5 and L at position 117 of SEQ ID NO:5; lll) G at the position corresponding to position 76 of SEQ ID NO:5 and L at the position corresponding to position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and S at position 81 of SEQ ID NO:5; The positions corresponding to positions 68 and 81 of SEQ ID NO: 5 are A, S and L respectively. ooo) G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; The positions corresponding to positions 81 and 109 of SEQ ID NO: 5 are: S, L, E, W, and D. The position corresponding to position 153 of SEQ ID NO:

5. The positions corresponding to positions 75, 81, 117, 121, 145, 155, 160, and 160 of SEQ ID NO: 5 are: A, S, L, L, A, W, and D. Q at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, I at position 117 of SEQ ID NO:5, Q at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; uuu) Y at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, L at position 117 of SEQ ID NO:5, G at position 121 of SEQ ID NO:5, L at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; vvv) C at position 22 of SEQ ID NO: 5, A at position 68 of SEQ ID NO: 5, Y at position 75 of SEQ ID NO: 5, G at position 76 of SEQ ID NO: 5, S at position 81 of SEQ ID NO: 5, H at position 108 of SEQ ID NO: 5, L at position 117 of SEQ ID NO: 5, H at position 122 of SEQ ID NO: 5, Y at position 138 of SEQ ID NO: 5, L at position 139 of SEQ ID NO: 5, A at position 142 of SEQ ID NO: 5, A at position 145 of SEQ ID NO: 5, R at position 153 of SEQ ID NO: 5, W at position 155 of SEQ ID NO: 5, R at position 156 of SEQ ID NO: 5, D at position 160 of SEQ ID NO: 5, and the position corresponding to SEQ ID NO: 5 The position H corresponding to position 162 of NO:5; or The position S at the location corresponding to position 81 of SEQ ID NO:5 (www).

115. The fusion protein according to claim 113 or 114, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO:300-414, 596-598 and 720-723.

116. The fusion protein according to claim 113 or 114, wherein the deaminase has an amino acid sequence as shown in any one of SEQ ID NO: 300-414, 596-598 and 720-723.

117. The fusion protein according to claim 113 or 114, wherein the fusion protein is selected from the group consisting of: a) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 922; b) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375, and is inserted after the amino acid position corresponding to position 910 of SEQ ID NO:4 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:52; c) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is inserted after the amino acid position corresponding to position 922 of SEQ ID NO:2 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50; d) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is inserted after the amino acid position corresponding to position 910 of SEQ ID NO:4 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:52; e) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1 after the amino acid position corresponding to position 772; f) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 678; g) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1 after the amino acid position corresponding to position 772; h) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 922; i) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 678; j) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:2; k) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 678; l) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 678; m) fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:4; n) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:4 after the amino acid position corresponding to position 910; o) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:2; p) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:4 after the amino acid position corresponding to position 910; q) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596, and is inserted after the amino acid position corresponding to position 922 of SEQ ID NO:2 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50; r) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; s) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; and t) fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:

49.

118. The fusion protein according to claim 113 or 114, wherein the fusion protein is selected from the group consisting of: a) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:375, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:

2. b) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:375, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:4 after the amino acid position corresponding to position 910; c) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:319, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:

2. d) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:319, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:4; e) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:375, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1. f) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:375, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:

2. g) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:319, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1. h) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:597, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:

2. i) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:319, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:

2. j) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:

2. k) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:597, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:

2. l) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:596, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:2 after the amino acid position corresponding to position 678; m) fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:4; n) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:597, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:

4. o) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:

2. p) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:596, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:

4. q) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:596, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:2; r) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:406, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 into the RGN having an amino acid sequence as shown in SEQ ID NO:49; s) a fusion protein, wherein the deaminase has the amino acid sequence shown in SEQ ID NO:596, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 into the RGN having the amino acid sequence shown in SEQ ID NO:49; and t) fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:597, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1.

119. The fusion protein of claim 118, wherein the fusion protein has a sequence as shown in any one of SEQ ID NO: 599, 600, 602, 603, 606, 610, 612, 614, 615, 617, 619-624, 626 and 630-632.

120. The nucleic acid molecule of claim 29, wherein the adenine deaminase comprises an amino acid sequence having at least 85% sequence identity with SEQ ID NO:5, and comprises at least one of the following amino acid residues: a) Position C at the position corresponding to position 2 of SEQ ID NO:5; b) F or C at the position corresponding to position 22 of SEQ ID NO:5; c) Q at the position corresponding to position 23 of SEQ ID NO:5; d) N at position 35 corresponding to SEQ ID NO:5; e) L at the position corresponding to position 40 of SEQ ID NO:5; f) Q at the position corresponding to position 46 of SEQ ID NO:5; g) A or M at the position corresponding to position 68 of SEQ ID NO:5; h) H at the position corresponding to position 72 of SEQ ID NO:5; i) W, A, Q, Y or D at the position corresponding to position 75 of SEQ ID NO:5; j) G at the position corresponding to position 76 of SEQ ID NO:5; k) S at the position corresponding to position 81 of SEQ ID NO:5; l) M at the position corresponding to position 105 of SEQ ID NO:5; m) H at the position corresponding to position 108 of SEQ ID NO:5; n) L at the position corresponding to position 109 of SEQ ID NO:5; o) The I or L position corresponding to position 117 of SEQ ID NO:5; p) F at the position corresponding to position 120 of SEQ ID NO:5; q) The E, T, A, G or V at the position corresponding to position 121 of SEQ ID NO:5; r) the I or H position corresponding to position 122 of SEQ ID NO:5; s) K at the position corresponding to position 125 of SEQ ID NO:5; t) A at position 126 corresponding to SEQ ID NO:5; u) H at the position corresponding to position 135 of SEQ ID NO:5; v) V at the position corresponding to position 137 of SEQ ID NO:5; w) The position Y corresponding to position 138 of SEQ ID NO:5; x) L at the position corresponding to position 139 of SEQ ID NO:5; y) K or A at the position corresponding to position 142 of SEQ ID NO:5; z) A, Q, L or M at position 145 of SEQ ID NO:5; aa) K at position 148 of SEQ ID NO:5; bb) Q at the position corresponding to position 151 of SEQ ID NO:5; cc) The position E or R corresponding to position 153 of SEQ ID NO:5; dd) W at the position corresponding to position 155 of SEQ ID NO:5; The R or V at the position corresponding to position 156 of SEQ ID NO:5; ff) F at the position corresponding to position 157 of SEQ ID NO:5; The R at position 158 of SEQ ID NO:5 (gg); hh) is the position Q corresponding to position 159 of SEQ ID NO:5; ii) D or W at the position corresponding to position 160 of SEQ ID NO:5; jj) is the W, H, V or A position corresponding to position 162 of SEQ ID NO:5; kk) at position R corresponding to position 165 of SEQ ID NO:5; and ll) H at the position corresponding to position 166 of SEQ ID NO:

5.

121. The nucleic acid molecule of claim 120, wherein the adenine deaminase comprises: a) N at position 35 of SEQ ID NO:5 and W at position 162 of SEQ ID NO:5; b) N at position 35 of SEQ ID NO:5 and Q at position 46 of SEQ ID NO:5; c) S at position 81 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; d) Q at position 46 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; e) R at the position corresponding to position 156 of SEQ ID NO:5 and W at the position corresponding to position 162 of SEQ ID NO:5; f) Position A corresponding to position 68 of SEQ ID NO: 5 and position S corresponding to position 81 of SEQ ID NO: 5; g) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at position 35 of SEQ ID NO:5, Q at position 46 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; i) W at position 155 corresponding to SEQ ID NO:5, R at position 156 corresponding to SEQ ID NO:5, and W at position 162 corresponding to SEQ ID NO:5; j) N at position 35 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; k) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; l) M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; m) Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; n) M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; o) N at position 35, Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; p) S at position 81 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; q) W at position 75 corresponding to SEQ ID NO:5 and S at position 81 corresponding to SEQ ID NO:5; r) S at position 81 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; s) S at position 81 corresponding to SEQ ID NO:5 and A at position 145 corresponding to SEQ ID NO:5; t) W at position 75 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; u) W at position 155 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; v) A at position 145 of SEQ ID NO: 5 and W at position 155 of SEQ ID NO: 5; w) W at position 75 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; x) W at position 75 of SEQ ID NO:5 and A at position 145 of SEQ ID NO:5; y) Position A corresponding to position 145 of SEQ ID NO:5 and position D corresponding to position 160 of SEQ ID NO:5; z) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; aa) S at position 81 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; bb) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; cc) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; dd) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and A at position 145 corresponding to SEQ ID NO:5; (ee) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; ff) W at position 75 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; gg) W at position 75 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; hh) A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; ii) W at position 75 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; jj) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; kk) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; ll) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; mm) The positions corresponding to positions 75 and 145 of SEQ ID NO: 5 are: W, A, W, and D, respectively. oo) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; pp) L at position 40 of SEQ ID NO:5 and M at position 105 of SEQ ID NO:5; qq) L at position 40 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; rr) L at position 40 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; ss) L at the position corresponding to position 40 of SEQ ID NO:5 and Q at the position corresponding to position 159 of SEQ ID NO:5; L at position 40 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; uu) M at position 105 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; vv) M at the position corresponding to position 105 of SEQ ID NO:5 and M at the position corresponding to position 145 of SEQ ID NO:5; ww) M at position 105 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; xx) M at position 105 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; yy) V at position 121 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; zz) V at the position corresponding to position 121 of SEQ ID NO: 5 and Q at the position corresponding to position 159 of SEQ ID NO: 5; aaa) V at position 121 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; bbb) M at position 145 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; ccc) M at position 145 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; ddd) Q at position 159 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; The position S corresponding to position 81 of SEQ ID NO:5, the position M corresponding to position 145 of SEQ ID NO:5, and the position W corresponding to position 160 of SEQ ID NO:5; fff) S at position 81 corresponding to SEQ ID NO:5, V at position 121 corresponding to SEQ ID NO:5, and Q at position 159 corresponding to SEQ ID NO:5; ggg) A at position 68 of SEQ ID NO:5 and S at position 81 of SEQ ID NO:5; hhh) G at the position corresponding to position 76 of SEQ ID NO:5 and S at the position corresponding to position 81 of SEQ ID NO:5; iii) S at position 81 corresponding to SEQ ID NO:5 and L at position 117 corresponding to SEQ ID NO:5; jjj) A at position 68 of SEQ ID NO:5 and G at position 76 of SEQ ID NO:5; kkk) A at position 68 of SEQ ID NO:5 and L at position 117 of SEQ ID NO:5; lll) G at the position corresponding to position 76 of SEQ ID NO:5 and L at the position corresponding to position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and S at position 81 of SEQ ID NO:5; The positions corresponding to positions 68 and 81 of SEQ ID NO: 5 are A, S and L respectively. ooo) G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; The positions corresponding to positions 81 and 109 of SEQ ID NO: 5 are: S, L, E, W, and D. The position corresponding to position 153 of SEQ ID NO:

5. The positions corresponding to positions 75, 81, 117, 121, 145, 155, 160, and 160 of SEQ ID NO: 5 are: A, S, L, L, A, W, and D. Q at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, I at position 117 of SEQ ID NO:5, Q at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; uuu) Y at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, L at position 117 of SEQ ID NO:5, G at position 121 of SEQ ID NO:5, L at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; vvv) C at position 22 of SEQ ID NO: 5, A at position 68 of SEQ ID NO: 5, Y at position 75 of SEQ ID NO: 5, G at position 76 of SEQ ID NO: 5, S at position 81 of SEQ ID NO: 5, H at position 108 of SEQ ID NO: 5, L at position 117 of SEQ ID NO: 5, H at position 122 of SEQ ID NO: 5, Y at position 138 of SEQ ID NO: 5, L at position 139 of SEQ ID NO: 5, A at position 142 of SEQ ID NO: 5, A at position 145 of SEQ ID NO: 5, R at position 153 of SEQ ID NO: 5, W at position 155 of SEQ ID NO: 5, R at position 156 of SEQ ID NO: 5, D at position 160 of SEQ ID NO: 5, and the position corresponding to SEQ ID NO: 5 The position H corresponding to position 162 of NO:5; or The position S at the location corresponding to position 81 of SEQ ID NO:5 (www).

122. The nucleic acid molecule according to claim 120 or 121, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO:300-414, 596-598 and 720-723.

123. The nucleic acid molecule according to claim 120 or 121, wherein the deaminase has an amino acid sequence as shown in any one of SEQ ID NO: 300-414, 596-598 and 720-723.

124. The nucleic acid molecule according to claim 120 or 121, wherein the fusion protein is selected from the group consisting of: a) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 922; b) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375, and is inserted after the amino acid position corresponding to position 910 of SEQ ID NO:4 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:52; c) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is inserted after the amino acid position corresponding to position 922 of SEQ ID NO:2 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50; d) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is inserted after the amino acid position corresponding to position 910 of SEQ ID NO:4 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:52; e) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1 after the amino acid position corresponding to position 772; f) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 678; g) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1 after the amino acid position corresponding to position 772; h) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 922; i) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 678; j) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:2; k) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 678; l) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 678; m) fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:4; n) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:4 after the amino acid position corresponding to position 910; o) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:2; p) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:4 after the amino acid position corresponding to position 910; q) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596, and is inserted after the amino acid position corresponding to position 922 of SEQ ID NO:2 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50; r) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; s) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; and t) fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:

49.

125. The nucleic acid molecule according to claim 120 or 121, wherein the fusion protein is selected from the group consisting of: a) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:375, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:

2. b) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:375, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:4 after the amino acid position corresponding to position 910; c) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:319, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:

2. d) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:319, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:4; e) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:375, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1. f) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:375, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:

2. g) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:319, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1. h) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:597, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:

2. i) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:319, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:

2. j) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:

2. k) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:597, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:

2. l) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:596, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:2 after the amino acid position corresponding to position 678; m) fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:4; n) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:597, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:

4. o) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:

2. p) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:596, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:

4. q) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:596, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:2; r) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:406, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 into the RGN having an amino acid sequence as shown in SEQ ID NO:49; s) a fusion protein, wherein the deaminase has the amino acid sequence shown in SEQ ID NO:596, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 into the RGN having the amino acid sequence shown in SEQ ID NO:49; and t) fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:597, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1.

126. The nucleic acid molecule of claim 125, wherein the fusion protein has a sequence as shown in any one of SEQ ID NO: 599, 600, 602, 603, 606, 610, 612, 614, 615, 617, 619-624, 626 and 630-632.

127. The system of claim 57, wherein the adenine deaminase comprises an amino acid sequence having at least 85% sequence identity with SEQ ID NO:5, and comprises at least one of the following amino acid residues: a) Position C at the position corresponding to position 2 of SEQ ID NO:5; b) F or C at the position corresponding to position 22 of SEQ ID NO:5; c) Q at the position corresponding to position 23 of SEQ ID NO:5; d) N at position 35 corresponding to SEQ ID NO:5; e) L at the position corresponding to position 40 of SEQ ID NO:5; f) Q at the position corresponding to position 46 of SEQ ID NO:5; g) A or M at the position corresponding to position 68 of SEQ ID NO:5; h) H at the position corresponding to position 72 of SEQ ID NO:5; i) W, A, Q, Y or D at the position corresponding to position 75 of SEQ ID NO:5; j) G at the position corresponding to position 76 of SEQ ID NO:5; k) S at the position corresponding to position 81 of SEQ ID NO:5; l) M at the position corresponding to position 105 of SEQ ID NO:5; m) H at the position corresponding to position 108 of SEQ ID NO:5; n) L at the position corresponding to position 109 of SEQ ID NO:5; o) The I or L position corresponding to position 117 of SEQ ID NO:5; p) F at the position corresponding to position 120 of SEQ ID NO:5; q) The E, T, A, G or V at the position corresponding to position 121 of SEQ ID NO:5; r) the I or H position corresponding to position 122 of SEQ ID NO:5; s) K at the position corresponding to position 125 of SEQ ID NO:5; t) A at position 126 corresponding to SEQ ID NO:5; u) H at the position corresponding to position 135 of SEQ ID NO:5; v) V at the position corresponding to position 137 of SEQ ID NO:5; w) The position Y corresponding to position 138 of SEQ ID NO:5; x) L at the position corresponding to position 139 of SEQ ID NO:5; y) K or A at the position corresponding to position 142 of SEQ ID NO:5; z) A, Q, L or M at the position corresponding to position 145 of SEQ ID NO:5; aa) K at the position corresponding to position 148 of SEQ ID NO:5; bb) Q at the position corresponding to position 151 of SEQ ID NO:5; cc) The position E or R corresponding to position 153 of SEQ ID NO:5; dd) W at the position corresponding to position 155 of SEQ ID NO:5; The R or V at the position corresponding to position 156 of SEQ ID NO:5; ff) F at the position corresponding to position 157 of SEQ ID NO:5; The R at position 158 of SEQ ID NO:5 (gg); hh) is the position Q corresponding to position 159 of SEQ ID NO:5; ii) D or W at the position corresponding to position 160 of SEQ ID NO:5; jj) is the W, H, V or A position corresponding to position 162 of SEQ ID NO:5; kk) at position R corresponding to position 165 of SEQ ID NO:5; and ll) H at the position corresponding to position 166 of SEQ ID NO:

5.

128. The system of claim 127, wherein the adenine deaminase comprises: a) N at position 35 of SEQ ID NO:5 and W at position 162 of SEQ ID NO:5; b) N at position 35 of SEQ ID NO:5 and Q at position 46 of SEQ ID NO:5; c) S at position 81 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; d) Q at position 46 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; e) R at the position corresponding to position 156 of SEQ ID NO:5 and W at the position corresponding to position 162 of SEQ ID NO:5; f) Position A corresponding to position 68 of SEQ ID NO: 5 and position S corresponding to position 81 of SEQ ID NO: 5; g) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at position 35 of SEQ ID NO:5, Q at position 46 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; i) W at position 155 corresponding to SEQ ID NO:5, R at position 156 corresponding to SEQ ID NO:5, and W at position 162 corresponding to SEQ ID NO:5; j) N at position 35 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; k) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; l) M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; m) Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; n) M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; o) N at position 35, Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; p) S at position 81 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; q) W at position 75 corresponding to SEQ ID NO:5 and S at position 81 corresponding to SEQ ID NO:5; r) S at position 81 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; s) S at position 81 corresponding to SEQ ID NO:5 and A at position 145 corresponding to SEQ ID NO:5; t) W at position 75 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; u) W at position 155 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; v) A at position 145 of SEQ ID NO: 5 and W at position 155 of SEQ ID NO: 5; w) W at position 75 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; x) W at position 75 of SEQ ID NO:5 and A at position 145 of SEQ ID NO:5; y) Position A corresponding to position 145 of SEQ ID NO:5 and position D corresponding to position 160 of SEQ ID NO:5; z) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; aa) S at position 81 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; bb) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; cc) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; dd) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and A at position 145 corresponding to SEQ ID NO:5; (ee) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; ff) W at position 75 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; gg) W at position 75 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; hh) A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; ii) W at position 75 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; jj) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; kk) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; ll) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; mm) The positions corresponding to positions 75 and 145 of SEQ ID NO: 5 are: W, A, W, and D, respectively. oo) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; pp) L at position 40 of SEQ ID NO:5 and M at position 105 of SEQ ID NO:5; qq) L at position 40 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; rr) L at position 40 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; ss) L at the position corresponding to position 40 of SEQ ID NO:5 and Q at the position corresponding to position 159 of SEQ ID NO:5; L at position 40 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; uu) M at position 105 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; vv) M at the position corresponding to position 105 of SEQ ID NO:5 and M at the position corresponding to position 145 of SEQ ID NO:5; ww) M at position 105 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; xx) M at position 105 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; yy) V at position 121 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; zz) V at the position corresponding to position 121 of SEQ ID NO: 5 and Q at the position corresponding to position 159 of SEQ ID NO: 5; aaa) V at position 121 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; bbb) M at position 145 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; ccc) M at position 145 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; ddd) Q at position 159 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; The position S corresponding to position 81 of SEQ ID NO:5, the position M corresponding to position 145 of SEQ ID NO:5, and the position W corresponding to position 160 of SEQ ID NO:5; fff) S at position 81 corresponding to SEQ ID NO:5, V at position 121 corresponding to SEQ ID NO:5, and Q at position 159 corresponding to SEQ ID NO:5; ggg) A at position 68 of SEQ ID NO:5 and S at position 81 of SEQ ID NO:5; hhh) G at the position corresponding to position 76 of SEQ ID NO:5 and S at the position corresponding to position 81 of SEQ ID NO:5; iii) S at position 81 corresponding to SEQ ID NO:5 and L at position 117 corresponding to SEQ ID NO:5; jjj) A at position 68 of SEQ ID NO:5 and G at position 76 of SEQ ID NO:5; kkk) A at position 68 of SEQ ID NO:5 and L at position 117 of SEQ ID NO:5; lll) G at the position corresponding to position 76 of SEQ ID NO:5 and L at the position corresponding to position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and S at position 81 of SEQ ID NO:5; The positions corresponding to positions 68 and 81 of SEQ ID NO: 5 are A, S and L respectively. ooo) G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; The positions corresponding to positions 81 and 109 of SEQ ID NO: 5 are: S, L, E, W, and D. The position corresponding to position 153 of SEQ ID NO:

5. The positions corresponding to positions 75, 81, 117, 121, 145, 155, 160, and 160 of SEQ ID NO: 5 are: A, S, L, L, A, W, and D. Q at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, I at position 117 of SEQ ID NO:5, Q at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; uuu) Y at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, L at position 117 of SEQ ID NO:5, G at position 121 of SEQ ID NO:5, L at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; vvv) C at position 22 of SEQ ID NO: 5, A at position 68 of SEQ ID NO: 5, Y at position 75 of SEQ ID NO: 5, G at position 76 of SEQ ID NO: 5, S at position 81 of SEQ ID NO: 5, H at position 108 of SEQ ID NO: 5, L at position 117 of SEQ ID NO: 5, H at position 122 of SEQ ID NO: 5, Y at position 138 of SEQ ID NO: 5, L at position 139 of SEQ ID NO: 5, A at position 142 of SEQ ID NO: 5, A at position 145 of SEQ ID NO: 5, R at position 153 of SEQ ID NO: 5, W at position 155 of SEQ ID NO: 5, R at position 156 of SEQ ID NO: 5, D at position 160 of SEQ ID NO: 5, and the position corresponding to SEQ ID NO: 5 The position H corresponding to position 162 of NO:5; or The position S at the location corresponding to position 81 of SEQ ID NO:5 (www).

129. The system according to claim 127 or 128, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO:300-414, 596-598 and 720-723.

130. The system according to claim 127 or 128, wherein the deaminase has an amino acid sequence as shown in any one of SEQ ID NO: 300-414, 596-598 and 720-723.

131. The system according to claim 127 or 128, wherein the fusion protein is selected from the group consisting of: a) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 922; b) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375, and is inserted after the amino acid position corresponding to position 910 of SEQ ID NO:4 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:52; c) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is inserted after the amino acid position corresponding to position 922 of SEQ ID NO:2 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50; d) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is inserted after the amino acid position corresponding to position 910 of SEQ ID NO:4 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:52; e) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1 after the amino acid position corresponding to position 772; f) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 678; g) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1 after the amino acid position corresponding to position 772; h) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 922; i) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 678; j) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:2; k) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 678; l) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 678; m) fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:4; n) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:4 after the amino acid position corresponding to position 910; o) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:2; p) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:4 after the amino acid position corresponding to position 910; q) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596, and is inserted after the amino acid position corresponding to position 922 of SEQ ID NO:2 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50; r) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; s) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; and t) fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:

49.

132. The system according to claim 127 or 128, a) a fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:375, and is inserted after the amino acid position corresponding to position 922 of SEQ ID NO:2 into the RGN having an amino acid sequence as shown in SEQ ID NO:50; a) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:375, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:4 after the amino acid position corresponding to position 910; b) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:319, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:

2. c) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:319, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:

4. d) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:375, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:49 after the amino acid position corresponding to position 772 of SEQ ID NO:1; e) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:375, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:

2. f) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:319, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1. g) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:597, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:

2. h) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:319, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:

2. i) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:2 after the amino acid position 678. j) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:597, and is inserted after the amino acid position corresponding to position 678 of SEQ ID NO:2 into the RGN having an amino acid sequence as shown in SEQ ID NO:50; k) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:596, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:

2. l) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:

4. m) fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:597, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:4; n) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:

2. o) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:596, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:

4. p) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:596, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:

2. q) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:1 after the amino acid position 772. r) A fusion protein, wherein the deaminase has the amino acid sequence shown in SEQ ID NO:596, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 into the RGN having the amino acid sequence shown in SEQ ID NO:49; and s) fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:597 and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1.

133. The system of claim 132, wherein the fusion protein has a sequence as shown in any one of SEQ ID NO: 599, 600, 602, 603, 606, 610, 612, 614, 615, 617, 619-624, 626 and 630-632.

134. The method of claim 88, wherein the adenine deaminase comprises an amino acid sequence having at least 85% sequence identity with SEQ ID NO:5, and comprises at least one of the following amino acid residues: a) Position C at the position corresponding to position 2 of SEQ ID NO:5; b) F or C at the position corresponding to position 22 of SEQ ID NO:5; c) Q at the position corresponding to position 23 of SEQ ID NO:5; d) N at position 35 corresponding to SEQ ID NO:5; e) L at the position corresponding to position 40 of SEQ ID NO:5; f) Q at the position corresponding to position 46 of SEQ ID NO:5; g) A or M at the position corresponding to position 68 of SEQ ID NO:5; h) H at the position corresponding to position 72 of SEQ ID NO:5; i) W, A, Q, Y or D at the position corresponding to position 75 of SEQ ID NO:5; j) G at the position corresponding to position 76 of SEQ ID NO:5; k) S at the position corresponding to position 81 of SEQ ID NO:5; l) M at the position corresponding to position 105 of SEQ ID NO:5; m) H at the position corresponding to position 108 of SEQ ID NO:5; n) L at the position corresponding to position 109 of SEQ ID NO:5; o) The I or L position corresponding to position 117 of SEQ ID NO:5; p) F at the position corresponding to position 120 of SEQ ID NO:5; q) The E, T, A, G or V at the position corresponding to position 121 of SEQ ID NO:5; r) the I or H position corresponding to position 122 of SEQ ID NO:5; s) K at the position corresponding to position 125 of SEQ ID NO:5; t) A at position 126 corresponding to SEQ ID NO:5; u) H at the position corresponding to position 135 of SEQ ID NO:5; v) V at the position corresponding to position 137 of SEQ ID NO:5; w) The position Y corresponding to position 138 of SEQ ID NO:5; x) L at the position corresponding to position 139 of SEQ ID NO:5; y) K or A at the position corresponding to position 142 of SEQ ID NO:5; z) A, Q, L or M at position 145 of SEQ ID NO:5; aa) K at position 148 of SEQ ID NO:5; bb) Q at the position corresponding to position 151 of SEQ ID NO:5; cc) The position E or R corresponding to position 153 of SEQ ID NO:5; dd) W at the position corresponding to position 155 of SEQ ID NO:5; The R or V at the position corresponding to position 156 of SEQ ID NO:5; ff) F at the position corresponding to position 157 of SEQ ID NO:5; The R at position 158 of SEQ ID NO:5 (gg); hh) is the position Q corresponding to position 159 of SEQ ID NO:5; ii) D or W at the position corresponding to position 160 of SEQ ID NO:5; jj) W, H, V or A at position 162 of SEQ ID NO:5; kk) R at position 165 of SEQ ID NO:5; and ll) H at position 166 of SEQ ID NO:

5.

135. The method of claim 134, wherein the adenine deaminase comprises: a) N at position 35 of SEQ ID NO:5 and W at position 162 of SEQ ID NO:5; b) N at position 35 of SEQ ID NO:5 and Q at position 46 of SEQ ID NO:5; c) S at position 81 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; d) Q at position 46 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; e) R at the position corresponding to position 156 of SEQ ID NO:5 and W at the position corresponding to position 162 of SEQ ID NO:5; f) Position A corresponding to position 68 of SEQ ID NO: 5 and position S corresponding to position 81 of SEQ ID NO: 5; g) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at position 35 of SEQ ID NO:5, Q at position 46 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; i) W at position 155 corresponding to SEQ ID NO:5, R at position 156 corresponding to SEQ ID NO:5, and W at position 162 corresponding to SEQ ID NO:5; j) N at position 35 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; k) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; l) M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; m) Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; n) M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; o) N at position 35, Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; p) S at position 81 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; q) W at position 75 corresponding to SEQ ID NO:5 and S at position 81 corresponding to SEQ ID NO:5; r) S at position 81 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; s) S at position 81 corresponding to SEQ ID NO:5 and A at position 145 corresponding to SEQ ID NO:5; t) W at position 75 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; u) W at position 155 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; v) A at position 145 of SEQ ID NO: 5 and W at position 155 of SEQ ID NO: 5; w) W at position 75 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; x) W at position 75 of SEQ ID NO:5 and A at position 145 of SEQ ID NO:5; y) Position A corresponding to position 145 of SEQ ID NO:5 and position D corresponding to position 160 of SEQ ID NO:5; z) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; aa) S at position 81 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; bb) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; cc) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; dd) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and A at position 145 corresponding to SEQ ID NO:5; (ee) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; ff) W at position 75 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; gg) W at position 75 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; hh) A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; ii) W at position 75 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; jj) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; kk) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; ll) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; mm) The positions corresponding to positions 75 and 145 of SEQ ID NO: 5 are: W, A, W, and D, respectively. oo) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; pp) L at position 40 of SEQ ID NO:5 and M at position 105 of SEQ ID NO:5; qq) L at position 40 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; rr) L at position 40 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; ss) L at the position corresponding to position 40 of SEQ ID NO:5 and Q at the position corresponding to position 159 of SEQ ID NO:5; L at position 40 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; uu) M at position 105 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; vv) M at the position corresponding to position 105 of SEQ ID NO:5 and M at the position corresponding to position 145 of SEQ ID NO:5; ww) M at position 105 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; xx) M at position 105 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; yy) V at position 121 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; zz) V at the position corresponding to position 121 of SEQ ID NO: 5 and Q at the position corresponding to position 159 of SEQ ID NO: 5; aaa) V at position 121 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; bbb) M at position 145 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; ccc) M at position 145 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; ddd) Q at position 159 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; The position S corresponding to position 81 of SEQ ID NO:5, the position M corresponding to position 145 of SEQ ID NO:5, and the position W corresponding to position 160 of SEQ ID NO:5; fff) S at position 81 corresponding to SEQ ID NO:5, V at position 121 corresponding to SEQ ID NO:5, and Q at position 159 corresponding to SEQ ID NO:5; ggg) A at position 68 of SEQ ID NO:5 and S at position 81 of SEQ ID NO:5; hhh) G at the position corresponding to position 76 of SEQ ID NO:5 and S at the position corresponding to position 81 of SEQ ID NO:5; iii) S at position 81 corresponding to SEQ ID NO:5 and L at position 117 corresponding to SEQ ID NO:5; jjj) A at position 68 of SEQ ID NO:5 and G at position 76 of SEQ ID NO:5; kkk) A at position 68 of SEQ ID NO:5 and L at position 117 of SEQ ID NO:5; lll) G at the position corresponding to position 76 of SEQ ID NO:5 and L at the position corresponding to position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and S at position 81 of SEQ ID NO:5; The positions corresponding to positions 68 and 81 of SEQ ID NO: 5 are A, S and L respectively. ooo) G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; The positions corresponding to positions 81 and 109 of SEQ ID NO: 5 are: S, L, E, W, and D. The position corresponding to position 153 of SEQ ID NO:

5. The positions corresponding to positions 75, 81, 117, 121, 145, 155, 160, and 160 of SEQ ID NO: 5 are: A, S, L, L, A, W, and D. Q at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, I at position 117 of SEQ ID NO:5, Q at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; uuu) Y at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, L at position 117 of SEQ ID NO:5, G at position 121 of SEQ ID NO:5, L at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; vvv) C at position 22 of SEQ ID NO: 5, A at position 68 of SEQ ID NO: 5, Y at position 75 of SEQ ID NO: 5, G at position 76 of SEQ ID NO: 5, S at position 81 of SEQ ID NO: 5, H at position 108 of SEQ ID NO: 5, L at position 117 of SEQ ID NO: 5, H at position 122 of SEQ ID NO: 5, Y at position 138 of SEQ ID NO: 5, L at position 139 of SEQ ID NO: 5, A at position 142 of SEQ ID NO: 5, A at position 145 of SEQ ID NO: 5, R at position 153 of SEQ ID NO: 5, W at position 155 of SEQ ID NO: 5, R at position 156 of SEQ ID NO: 5, D at position 160 of SEQ ID NO: 5, and the position corresponding to SEQ ID NO: 5 The position H corresponding to position 162 of NO:5; or The position S at the location corresponding to position 81 of SEQ ID NO:5 (www).

136. The method according to claim 134 or 135, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO:300-414, 596-598 and 720-723.

137. The method according to claim 134 or 135, wherein the deaminase has an amino acid sequence as shown in any one of SEQ ID NO: 300-414, 596-598 and 720-723.

138. The method of claim 134 or 135, wherein the fusion protein is selected from the group consisting of: a) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 922; b) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375, and is inserted after the amino acid position corresponding to position 910 of SEQ ID NO:4 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:52; c) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is inserted after the amino acid position corresponding to position 922 of SEQ ID NO:2 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50; d) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is inserted after the amino acid position corresponding to position 910 of SEQ ID NO:4 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:52; e) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1 after the amino acid position corresponding to position 772; f) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 678; g) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1 after the amino acid position corresponding to position 772; h) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 922; i) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 678; j) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:2; k) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 678; l) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:2 after the amino acid position corresponding to position 678; m) fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:4; n) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:4 after the amino acid position corresponding to position 910; o) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:2; p) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596, and is inserted within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:4 after the amino acid position corresponding to position 910; q) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596, and is inserted after the amino acid position corresponding to position 922 of SEQ ID NO:2 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50; r) A fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 into the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; s) a fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; and t) fusion protein, wherein the deaminase has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 within the RGN having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:

49.

139. The method of claim 134 or 135, wherein the fusion protein is selected from the group consisting of: a) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:375, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:

2. b) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:375, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:4 after the amino acid position corresponding to position 910; c) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:319, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:

2. d) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:319, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:4; e) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:375, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1. f) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:375, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:

2. g) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:319, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1. h) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:597, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:

2. i) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:319, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:

2. j) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:

2. k) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:597, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 678 of SEQ ID NO:

2. l) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:596, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:2 after the amino acid position corresponding to position 678; m) fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:4; n) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:597, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:

4. o) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:406, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:

2. p) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:596, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:52 after the amino acid position corresponding to position 910 of SEQ ID NO:

4. q) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:596, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:50 after the amino acid position corresponding to position 922 of SEQ ID NO:2; r) A fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:406, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 into the RGN having an amino acid sequence as shown in SEQ ID NO:49; s) a fusion protein, wherein the deaminase has the amino acid sequence shown in SEQ ID NO:596, and is inserted after the amino acid position corresponding to position 772 of SEQ ID NO:1 into the RGN having the amino acid sequence shown in SEQ ID NO:49; and t) fusion protein, wherein the deaminase has an amino acid sequence as shown in SEQ ID NO:597, and is inserted into the RGN having an amino acid sequence as shown in SEQ ID NO:49 after the amino acid position corresponding to position 772 of SEQ ID NO:

1.

140. The method of claim 139, wherein the fusion protein has a sequence as shown in any one of SEQ ID NO: 599, 600, 602, 603, 606, 610, 612, 614, 615, 617, 619-624, 626 and 630-632.

141. A deaminase comprising an amino acid sequence having at least 85% sequence identity with SEQ ID NO:5, and comprising at least one of the following amino acid residues: a) Position C at the position corresponding to position 2 of SEQ ID NO:5; b) F or C at the position corresponding to position 22 of SEQ ID NO:5; c) Q at the position corresponding to position 23 of SEQ ID NO:5; d) N at position 35 corresponding to SEQ ID NO:5; e) L at the position corresponding to position 40 of SEQ ID NO:5; f) Q at the position corresponding to position 46 of SEQ ID NO:5; g) A or M at the position corresponding to position 68 of SEQ ID NO:5; h) H at the position corresponding to position 72 of SEQ ID NO:5; i) W, A, Q, Y or D at the position corresponding to position 75 of SEQ ID NO:5; j) G at the position corresponding to position 76 of SEQ ID NO:5; k) S at the position corresponding to position 81 of SEQ ID NO:5; l) M at the position corresponding to position 105 of SEQ ID NO:5; m) H at the position corresponding to position 108 of SEQ ID NO:5; n) L at the position corresponding to position 109 of SEQ ID NO:5; o) The I or L position corresponding to position 117 of SEQ ID NO:5; p) F at the position corresponding to position 120 of SEQ ID NO:5; q) The E, T, A, G or V at the position corresponding to position 121 of SEQ ID NO:5; r) the I or H position corresponding to position 122 of SEQ ID NO:5; s) K at the position corresponding to position 125 of SEQ ID NO:5; t) A at position 126 corresponding to SEQ ID NO:5; u) H at the position corresponding to position 135 of SEQ ID NO:5; v) V at the position corresponding to position 137 of SEQ ID NO:5; w) The position Y corresponding to position 138 of SEQ ID NO:5; x) L at the position corresponding to position 139 of SEQ ID NO:5; y) K or A at the position corresponding to position 142 of SEQ ID NO:5; z) A, Q, L or M at the position corresponding to position 145 of SEQ ID NO:5; aa) K at the position corresponding to position 148 of SEQ ID NO:5; bb) Q at the position corresponding to position 151 of SEQ ID NO:5; cc) The position E or R corresponding to position 153 of SEQ ID NO:5; dd) W at the position corresponding to position 155 of SEQ ID NO:5; The R or V at the position corresponding to position 156 of SEQ ID NO:5; ff) F at the position corresponding to position 157 of SEQ ID NO:5; The R at position 158 of SEQ ID NO:5 (gg); hh) is the position Q corresponding to position 159 of SEQ ID NO:5; ii) D or W at the position corresponding to position 160 of SEQ ID NO:5; jj) is the W, H, V or A position corresponding to position 162 of SEQ ID NO:5; kk) at position R corresponding to position 165 of SEQ ID NO:5; and ll) H at the position corresponding to position 166 of SEQ ID NO:

5.

142. The deaminase according to claim 141, wherein the deaminase comprises: a) N at position 35 of SEQ ID NO:5 and W at position 162 of SEQ ID NO:5; b) N at position 35 of SEQ ID NO:5 and Q at position 46 of SEQ ID NO:5; c) S at position 81 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; d) Q at position 46 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; e) R at the position corresponding to position 156 of SEQ ID NO:5 and W at the position corresponding to position 162 of SEQ ID NO:5; f) Position A corresponding to position 68 of SEQ ID NO: 5 and position S corresponding to position 81 of SEQ ID NO: 5; g) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at position 35 of SEQ ID NO:5, Q at position 46 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; i) W at position 155 corresponding to SEQ ID NO:5, R at position 156 corresponding to SEQ ID NO:5, and W at position 162 corresponding to SEQ ID NO:5; j) N at position 35 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; k) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; l) M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; m) Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; n) M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; o) N at position 35, Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; p) S at position 81 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; q) W at position 75 corresponding to SEQ ID NO:5 and S at position 81 corresponding to SEQ ID NO:5; r) S at position 81 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; s) S at position 81 corresponding to SEQ ID NO:5 and A at position 145 corresponding to SEQ ID NO:5; t) W at position 75 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; u) W at position 155 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; v) A at position 145 of SEQ ID NO: 5 and W at position 155 of SEQ ID NO: 5; w) W at position 75 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; x) W at position 75 of SEQ ID NO:5 and A at position 145 of SEQ ID NO:5; y) Position A corresponding to position 145 of SEQ ID NO:5 and position D corresponding to position 160 of SEQ ID NO:5; z) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; aa) S at position 81 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; bb) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; cc) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; dd) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and A at position 145 corresponding to SEQ ID NO:5; (ee) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; ff) W at position 75 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; gg) W at position 75 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; hh) A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; ii) W at position 75 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; jj) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; kk) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; ll) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; mm) The positions corresponding to positions 75 and 145 of SEQ ID NO: 5 are: W, A, W, and D, respectively. oo) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; pp) L at position 40 of SEQ ID NO:5 and M at position 105 of SEQ ID NO:5; qq) L at position 40 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; rr) L at position 40 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; ss) L at the position corresponding to position 40 of SEQ ID NO:5 and Q at the position corresponding to position 159 of SEQ ID NO:5; L at position 40 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; uu) M at position 105 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; vv) M at the position corresponding to position 105 of SEQ ID NO:5 and M at the position corresponding to position 145 of SEQ ID NO:5; ww) M at position 105 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; xx) M at position 105 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; yy) V at position 121 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; zz) V at the position corresponding to position 121 of SEQ ID NO: 5 and Q at the position corresponding to position 159 of SEQ ID NO: 5; aaa) V at position 121 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; bbb) M at position 145 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; ccc) M at position 145 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; ddd) Q at position 159 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; The position S corresponding to position 81 of SEQ ID NO:5, the position M corresponding to position 145 of SEQ ID NO:5, and the position W corresponding to position 160 of SEQ ID NO:5; fff) S at position 81 corresponding to SEQ ID NO:5, V at position 121 corresponding to SEQ ID NO:5, and Q at position 159 corresponding to SEQ ID NO:5; ggg) A at position 68 of SEQ ID NO:5 and S at position 81 of SEQ ID NO:5; hhh) G at the position corresponding to position 76 of SEQ ID NO:5 and S at the position corresponding to position 81 of SEQ ID NO:5; iii) S at position 81 corresponding to SEQ ID NO:5 and L at position 117 corresponding to SEQ ID NO:5; jjj) A at position 68 of SEQ ID NO:5 and G at position 76 of SEQ ID NO:5; kkk) A at position 68 of SEQ ID NO:5 and L at position 117 of SEQ ID NO:5; lll) G at the position corresponding to position 76 of SEQ ID NO:5 and L at the position corresponding to position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and S at position 81 of SEQ ID NO:5; The positions corresponding to positions 68 and 81 of SEQ ID NO: 5 are A, S and L respectively. ooo) G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; The positions corresponding to positions 81 and 109 of SEQ ID NO: 5 are: S, L, E, W, and D. The position corresponding to position 153 of SEQ ID NO:

5. The positions corresponding to positions 75, 81, 117, 121, 145, 155, 160, and 160 of SEQ ID NO: 5 are: A, S, L, L, A, W, and D. Q at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, I at position 117 of SEQ ID NO:5, Q at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; uuu) Y at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, L at position 117 of SEQ ID NO:5, G at position 121 of SEQ ID NO:5, L at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; vvv) C at position 22 of SEQ ID NO: 5, A at position 68 of SEQ ID NO: 5, Y at position 75 of SEQ ID NO: 5, G at position 76 of SEQ ID NO: 5, S at position 81 of SEQ ID NO: 5, H at position 108 of SEQ ID NO: 5, L at position 117 of SEQ ID NO: 5, H at position 122 of SEQ ID NO: 5, Y at position 138 of SEQ ID NO: 5, L at position 139 of SEQ ID NO: 5, A at position 142 of SEQ ID NO: 5, A at position 145 of SEQ ID NO: 5, R at position 153 of SEQ ID NO: 5, W at position 155 of SEQ ID NO: 5, R at position 156 of SEQ ID NO: 5, D at position 160 of SEQ ID NO: 5, and the position corresponding to SEQ ID NO: 5 The position H corresponding to position 162 of NO:5; or The position S at the location corresponding to position 81 of SEQ ID NO:5 (www).

143. The deaminase according to claim 141 or 142, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 300-414, 596-598, and 720-723.

144. The deaminase according to claim 141 or 142, wherein the deaminase comprises the amino acid sequence of any one of SEQ ID NO: 300-414, 596-598 and 720-723.

145. The deaminase according to any one of claims 141 to 144, wherein the deaminase has improved deaminase activity compared with SEQ ID NO:

5.

146. The deaminase according to any one of claims 141 to 145, wherein the deaminase is an adenine deaminase.

147. A nucleic acid molecule comprising a polynucleotide encoding a deaminase having at least 85% sequence identity with SEQ ID NO:5, wherein said deaminase comprises at least one of the following amino acid residues: a) Position C at the position corresponding to position 2 of SEQ ID NO:5; b) F or C at the position corresponding to position 22 of SEQ ID NO:5; c) Q at the position corresponding to position 23 of SEQ ID NO:5; d) N at position 35 corresponding to SEQ ID NO:5; e) L at the position corresponding to position 40 of SEQ ID NO:5; f) Q at the position corresponding to position 46 of SEQ ID NO:5; g) A or M at the position corresponding to position 68 of SEQ ID NO:5; h) H at the position corresponding to position 72 of SEQ ID NO:5; i) W, A, Q, Y or D at the position corresponding to position 75 of SEQ ID NO:5; j) G at the position corresponding to position 76 of SEQ ID NO:5; k) S at the position corresponding to position 81 of SEQ ID NO:5; l) M at the position corresponding to position 105 of SEQ ID NO:5; m) H at the position corresponding to position 108 of SEQ ID NO:5; n) L at the position corresponding to position 109 of SEQ ID NO:5; o) The I or L position corresponding to position 117 of SEQ ID NO:5; p) F at the position corresponding to position 120 of SEQ ID NO:5; q) The E, T, A, G or V at the position corresponding to position 121 of SEQ ID NO:5; r) the I or H position corresponding to position 122 of SEQ ID NO:5; s) K at the position corresponding to position 125 of SEQ ID NO:5; t) A at position 126 corresponding to SEQ ID NO:5; u) H at the position corresponding to position 135 of SEQ ID NO:5; v) V at the position corresponding to position 137 of SEQ ID NO:5; w) The position Y corresponding to position 138 of SEQ ID NO:5; x) L at the position corresponding to position 139 of SEQ ID NO:5; y) K or A at the position corresponding to position 142 of SEQ ID NO:5; z) A, Q, L or M at position 145 of SEQ ID NO:5; aa) K at position 148 of SEQ ID NO:5; bb) Q at the position corresponding to position 151 of SEQ ID NO:5; cc) The position E or R corresponding to position 153 of SEQ ID NO:5; dd) W at the position corresponding to position 155 of SEQ ID NO:5; The R or V at the position corresponding to position 156 of SEQ ID NO:5; ff) F at the position corresponding to position 157 of SEQ ID NO:5; The R at position 158 of SEQ ID NO:5 (gg); hh) is the position Q corresponding to position 159 of SEQ ID NO:5; ii) D or W at the position corresponding to position 160 of SEQ ID NO:5; jj) is the W, H, V or A position corresponding to position 162 of SEQ ID NO:5; kk) at position R corresponding to position 165 of SEQ ID NO:5; and ll) H at the position corresponding to position 166 of SEQ ID NO:

5.

148. The nucleic acid molecule of claim 147, wherein the deaminase comprises: a) N at position 35 of SEQ ID NO:5 and W at position 162 of SEQ ID NO:5; b) N at position 35 of SEQ ID NO:5 and Q at position 46 of SEQ ID NO:5; c) S at position 81 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; d) Q at position 46 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; e) R at the position corresponding to position 156 of SEQ ID NO:5 and W at the position corresponding to position 162 of SEQ ID NO:5; f) Position A corresponding to position 68 of SEQ ID NO: 5 and position S corresponding to position 81 of SEQ ID NO: 5; g) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at position 35 of SEQ ID NO:5, Q at position 46 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; i) W at position 155 corresponding to SEQ ID NO:5, R at position 156 corresponding to SEQ ID NO:5, and W at position 162 corresponding to SEQ ID NO:5; j) N at position 35 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; k) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; l) M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; m) Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; n) M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; o) N at position 35, Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; p) S at position 81 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; q) W at position 75 corresponding to SEQ ID NO:5 and S at position 81 corresponding to SEQ ID NO:5; r) S at position 81 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; s) S at position 81 corresponding to SEQ ID NO:5 and A at position 145 corresponding to SEQ ID NO:5; t) W at position 75 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; u) W at position 155 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; v) A at position 145 of SEQ ID NO: 5 and W at position 155 of SEQ ID NO: 5; w) W at position 75 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; x) W at position 75 of SEQ ID NO:5 and A at position 145 of SEQ ID NO:5; y) Position A corresponding to position 145 of SEQ ID NO:5 and position D corresponding to position 160 of SEQ ID NO:5; z) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; aa) S at position 81 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; bb) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; cc) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; dd) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and A at position 145 corresponding to SEQ ID NO:5; (ee) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; ff) W at position 75 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; gg) W at position 75 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; hh) A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; ii) W at position 75 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; jj) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; kk) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; ll) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; mm) The positions corresponding to positions 75 and 145 of SEQ ID NO: 5 are: W, A, W, and D, respectively. oo) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; pp) L at position 40 of SEQ ID NO:5 and M at position 105 of SEQ ID NO:5; qq) L at position 40 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; rr) L at position 40 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; ss) L at the position corresponding to position 40 of SEQ ID NO:5 and Q at the position corresponding to position 159 of SEQ ID NO:5; L at position 40 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; uu) M at position 105 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; vv) M at the position corresponding to position 105 of SEQ ID NO:5 and M at the position corresponding to position 145 of SEQ ID NO:5; ww) M at position 105 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; xx) M at position 105 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; yy) V at position 121 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; zz) V at the position corresponding to position 121 of SEQ ID NO: 5 and Q at the position corresponding to position 159 of SEQ ID NO: 5; aaa) V at position 121 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; bbb) M at position 145 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; ccc) M at position 145 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; ddd) Q at position 159 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; The position S corresponding to position 81 of SEQ ID NO:5, the position M corresponding to position 145 of SEQ ID NO:5, and the position W corresponding to position 160 of SEQ ID NO:5; fff) S at position 81 corresponding to SEQ ID NO:5, V at position 121 corresponding to SEQ ID NO:5, and Q at position 159 corresponding to SEQ ID NO:5; ggg) A at position 68 of SEQ ID NO:5 and S at position 81 of SEQ ID NO:5; hhh) G at the position corresponding to position 76 of SEQ ID NO:5 and S at the position corresponding to position 81 of SEQ ID NO:5; iii) S at position 81 corresponding to SEQ ID NO:5 and L at position 117 corresponding to SEQ ID NO:5; jjj) A at position 68 of SEQ ID NO:5 and G at position 76 of SEQ ID NO:5; kkk) A at position 68 of SEQ ID NO:5 and L at position 117 of SEQ ID NO:5; lll) G at the position corresponding to position 76 of SEQ ID NO:5 and L at the position corresponding to position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and S at position 81 of SEQ ID NO:5; The positions corresponding to positions 68 and 81 of SEQ ID NO: 5 are A, S and L respectively. ooo) G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; The positions corresponding to positions 81 and 109 of SEQ ID NO: 5 are: S, L, E, W, and D. The position corresponding to position 153 of SEQ ID NO:

5. The positions corresponding to positions 75, 81, 117, 121, 145, 155, 160, and 160 of SEQ ID NO: 5 are: A, S, L, L, A, W, and D. Q at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, I at position 117 of SEQ ID NO:5, Q at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; uuu) Y at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, L at position 117 of SEQ ID NO:5, G at position 121 of SEQ ID NO:5, L at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; vvv) C at position 22 of SEQ ID NO: 5, A at position 68 of SEQ ID NO: 5, Y at position 75 of SEQ ID NO: 5, G at position 76 of SEQ ID NO: 5, S at position 81 of SEQ ID NO: 5, H at position 108 of SEQ ID NO: 5, L at position 117 of SEQ ID NO: 5, H at position 122 of SEQ ID NO: 5, Y at position 138 of SEQ ID NO: 5, L at position 139 of SEQ ID NO: 5, A at position 142 of SEQ ID NO: 5, A at position 145 of SEQ ID NO: 5, R at position 153 of SEQ ID NO: 5, W at position 155 of SEQ ID NO: 5, R at position 156 of SEQ ID NO: 5, D at position 160 of SEQ ID NO: 5, and the position corresponding to SEQ ID NO: 5 The position H corresponding to position 162 of NO:5; or The position S at the location corresponding to position 81 of SEQ ID NO:5 (www).

149. The nucleic acid molecule according to claim 147 or 148, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO:300-414, 596-598 and 720-723.

150. The nucleic acid molecule according to claim 147 or 148, wherein the deaminase comprises the amino acid sequence of any one of SEQ ID NO: 300-414, 596-598 and 720-723.

151. The nucleic acid molecule according to any one of claims 147 to 150, wherein the deaminase has improved deaminase activity when compared with SEQ ID NO:

5.

152. The nucleic acid molecule according to any one of claims 147 to 151, wherein the deaminase is adenine deaminase.

153. The nucleic acid molecule according to any one of claims 147 to 152, further comprising a heterologous promoter operatively linked to the polynucleotide.

154. A vector comprising a nucleic acid molecule according to any one of claims 147 to 153.

155. A cell comprising a deaminase according to any one of claims 141 to 146, a nucleic acid molecule according to any one of claims 147 to 153, or a vector according to claim 154.

156. The cell of claim 155, wherein the cell is a mammalian cell.

157. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a deaminase according to any one of claims 141 to 146, a nucleic acid molecule according to any one of claims 147 to 153, a carrier according to claim 154, or a cell according to claim 156.

158. A fusion protein comprising a DNA-binding polypeptide and a deaminase having at least 85% sequence identity with SEQ ID NO:5, wherein said deaminase comprises at least one of the following amino acid residues: a) Position C at the position corresponding to position 2 of SEQ ID NO:5; b) F or C at the position corresponding to position 22 of SEQ ID NO:5; c) Q at the position corresponding to position 23 of SEQ ID NO:5; d) N at position 35 corresponding to SEQ ID NO:5; e) L at the position corresponding to position 40 of SEQ ID NO:5; f) Q at the position corresponding to position 46 of SEQ ID NO:5; g) A or M at the position corresponding to position 68 of SEQ ID NO:5; h) H at the position corresponding to position 72 of SEQ ID NO:5; i) W, A, Q, Y or D at the position corresponding to position 75 of SEQ ID NO:5; j) G at the position corresponding to position 76 of SEQ ID NO:5; k) S at the position corresponding to position 81 of SEQ ID NO:5; l) M at the position corresponding to position 105 of SEQ ID NO:5; m) H at the position corresponding to position 108 of SEQ ID NO:5; n) L at the position corresponding to position 109 of SEQ ID NO:5; o) The I or L position corresponding to position 117 of SEQ ID NO:5; p) F at the position corresponding to position 120 of SEQ ID NO:5; q) The E, T, A, G or V at the position corresponding to position 121 of SEQ ID NO:5; r) the I or H position corresponding to position 122 of SEQ ID NO:5; s) K at the position corresponding to position 125 of SEQ ID NO:5; t) A at position 126 corresponding to SEQ ID NO:5; u) H at the position corresponding to position 135 of SEQ ID NO:5; v) V at the position corresponding to position 137 of SEQ ID NO:5; w) The position Y corresponding to position 138 of SEQ ID NO:5; x) L at the position corresponding to position 139 of SEQ ID NO:5; y) K or A at the position corresponding to position 142 of SEQ ID NO:5; z) A, Q, L or M at the position corresponding to position 145 of SEQ ID NO:5; aa) K at the position corresponding to position 148 of SEQ ID NO:5; bb) Q at the position corresponding to position 151 of SEQ ID NO:5; cc) The position E or R corresponding to position 153 of SEQ ID NO:5; dd) W at the position corresponding to position 155 of SEQ ID NO:5; The R or V at the position corresponding to position 156 of SEQ ID NO:5; ff) F at the position corresponding to position 157 of SEQ ID NO:5; The R at position 158 of SEQ ID NO:5 (gg); hh) is the position Q corresponding to position 159 of SEQ ID NO:5; ii) D or W at the position corresponding to position 160 of SEQ ID NO:5; jj) is the W, H, V or A position corresponding to position 162 of SEQ ID NO:5; kk) at position R corresponding to position 165 of SEQ ID NO:5; and ll) H at the position corresponding to position 166 of SEQ ID NO:

5.

159. The fusion protein of claim 158, wherein the deaminase comprises: a) N at position 35 of SEQ ID NO:5 and W at position 162 of SEQ ID NO:5; b) N at position 35 of SEQ ID NO:5 and Q at position 46 of SEQ ID NO:5; c) S at position 81 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; d) Q at position 46 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; e) R at the position corresponding to position 156 of SEQ ID NO:5 and W at the position corresponding to position 162 of SEQ ID NO:5; f) Position A corresponding to position 68 of SEQ ID NO: 5 and position S corresponding to position 81 of SEQ ID NO: 5; g) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at position 35 of SEQ ID NO:5, Q at position 46 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; i) W at position 155 corresponding to SEQ ID NO:5, R at position 156 corresponding to SEQ ID NO:5, and W at position 162 corresponding to SEQ ID NO:5; j) N at position 35 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; k) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; l) M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; m) Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; n) M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; o) N at position 35, Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; p) S at position 81 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; q) W at position 75 corresponding to SEQ ID NO:5 and S at position 81 corresponding to SEQ ID NO:5; r) S at position 81 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; s) S at position 81 corresponding to SEQ ID NO:5 and A at position 145 corresponding to SEQ ID NO:5; t) W at position 75 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; u) W at position 155 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; v) A at position 145 of SEQ ID NO: 5 and W at position 155 of SEQ ID NO: 5; w) W at position 75 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; x) W at position 75 of SEQ ID NO:5 and A at position 145 of SEQ ID NO:5; y) Position A corresponding to position 145 of SEQ ID NO:5 and position D corresponding to position 160 of SEQ ID NO:5; z) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; aa) S at position 81 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; bb) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; cc) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; dd) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and A at position 145 corresponding to SEQ ID NO:5; (ee) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; ff) W at position 75 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; gg) W at position 75 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; hh) A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; ii) W at position 75 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; jj) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; kk) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; ll) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; mm) The positions corresponding to positions 75 and 145 of SEQ ID NO: 5 are: W, A, W, and D, respectively. oo) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; pp) L at position 40 of SEQ ID NO:5 and M at position 105 of SEQ ID NO:5; qq) L at position 40 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; rr) L at position 40 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; ss) L at the position corresponding to position 40 of SEQ ID NO:5 and Q at the position corresponding to position 159 of SEQ ID NO:5; L at position 40 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; uu) M at position 105 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; vv) M at the position corresponding to position 105 of SEQ ID NO:5 and M at the position corresponding to position 145 of SEQ ID NO:5; ww) M at position 105 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; xx) M at position 105 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; yy) V at position 121 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; zz) V at the position corresponding to position 121 of SEQ ID NO: 5 and Q at the position corresponding to position 159 of SEQ ID NO: 5; aaa) V at position 121 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; bbb) M at position 145 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; ccc) M at position 145 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; ddd) Q at position 159 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; The position S corresponding to position 81 of SEQ ID NO:5, the position M corresponding to position 145 of SEQ ID NO:5, and the position W corresponding to position 160 of SEQ ID NO:5; fff) S at position 81 corresponding to SEQ ID NO:5, V at position 121 corresponding to SEQ ID NO:5, and Q at position 159 corresponding to SEQ ID NO:5; ggg) A at position 68 of SEQ ID NO:5 and S at position 81 of SEQ ID NO:5; hhh) G at the position corresponding to position 76 of SEQ ID NO:5 and S at the position corresponding to position 81 of SEQ ID NO:5; iii) S at position 81 corresponding to SEQ ID NO:5 and L at position 117 corresponding to SEQ ID NO:5; jjj) A at position 68 of SEQ ID NO:5 and G at position 76 of SEQ ID NO:5; kkk) A at position 68 of SEQ ID NO:5 and L at position 117 of SEQ ID NO:5; lll) G at the position corresponding to position 76 of SEQ ID NO:5 and L at the position corresponding to position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and S at position 81 of SEQ ID NO:5; The positions corresponding to positions 68 and 81 of SEQ ID NO: 5 are A, S and L respectively. ooo) G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; The positions corresponding to positions 81 and 109 of SEQ ID NO: 5 are: S, L, E, W, and D. The position corresponding to position 153 of SEQ ID NO:

5. The positions corresponding to positions 75, 81, 117, 121, 145, 155, 160, and 160 of SEQ ID NO: 5 are: A, S, L, L, A, W, and D. Q at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, I at position 117 of SEQ ID NO:5, Q at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; uuu) Y at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, L at position 117 of SEQ ID NO:5, G at position 121 of SEQ ID NO:5, L at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; vvv) C at position 22 of SEQ ID NO: 5, A at position 68 of SEQ ID NO: 5, Y at position 75 of SEQ ID NO: 5, G at position 76 of SEQ ID NO: 5, S at position 81 of SEQ ID NO: 5, H at position 108 of SEQ ID NO: 5, L at position 117 of SEQ ID NO: 5, H at position 122 of SEQ ID NO: 5, Y at position 138 of SEQ ID NO: 5, L at position 139 of SEQ ID NO: 5, A at position 142 of SEQ ID NO: 5, A at position 145 of SEQ ID NO: 5, R at position 153 of SEQ ID NO: 5, W at position 155 of SEQ ID NO: 5, R at position 156 of SEQ ID NO: 5, D at position 160 of SEQ ID NO: 5, and the position corresponding to SEQ ID NO: 5 The position H corresponding to position 162 of NO:5; or The position S at the location corresponding to position 81 of SEQ ID NO:5 (www).

160. The fusion protein according to claim 158 or 159, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO:300-414, 596-598 and 720-723.

161. The fusion protein according to claim 158 or 159, wherein the deaminase comprises the amino acid sequence of any one of SEQ ID NO: 300-414, 596-598 and 720-723.

162. The fusion protein according to any one of claims 158 to 161, wherein the deaminase has improved deaminase activity compared with SEQ ID NO:

5.

163. The fusion protein according to any one of claims 158 to 162, wherein the deaminase is adenine deaminase.

164. The fusion protein according to any one of claims 158 to 163, wherein the DNA-binding polypeptide is an RNA-guided nuclease (RGN) polypeptide.

165. The fusion protein of claim 164, wherein the RGN is an RGN cleavage enzyme.

166. The fusion protein of claim 164, wherein the RGN has an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

167. The fusion protein of claim 165, wherein the RGN nickase has an amino acid sequence as shown in any one of SEQ ID NO:49-52 and 698.

168. The fusion protein according to claim 158 or 159, wherein the fusion protein is selected from the group consisting of: a) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; b) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; c) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; d) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; e) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; f) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:303, and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; g) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:366, and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; and h) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:720 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:

49.

169. The fusion protein according to claim 158 or 159, wherein the fusion protein is selected from the group consisting of: a) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:319 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; b) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:375 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; c) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:406 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; d) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:596 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; e) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:597 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; f) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:303 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; g) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:366 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; and h) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence as shown in SEQ ID NO:720 and is fused to the N-terminus of an RGN comprising an amino acid sequence as shown in SEQ ID NO:

49.

170. The fusion protein according to claim 158 or 159, wherein the fusion protein has a sequence as shown in any one of SEQ ID NO: 577, 578, 582, 584, 588, 593, 594, 604, 605, 616 and 729.

171. A ribonucleoprotein (RNP) complex comprising a fusion protein according to any one of claims 158 to 170 and a guide RNA bound to said fusion protein.

172. A nucleic acid molecule comprising a polynucleotide encoding a fusion protein comprising a DNA-binding polypeptide and a deaminase, wherein the deaminase comprises at least one of the following amino acid residues: a) Position C at the position corresponding to position 2 of SEQ ID NO:5; b) F or C at the position corresponding to position 22 of SEQ ID NO:5; c) Q at the position corresponding to position 23 of SEQ ID NO:5; d) N at position 35 corresponding to SEQ ID NO:5; e) L at the position corresponding to position 40 of SEQ ID NO:5; f) Q at the position corresponding to position 46 of SEQ ID NO:5; g) A or M at the position corresponding to position 68 of SEQ ID NO:5; h) H at the position corresponding to position 72 of SEQ ID NO:5; i) W, A, Q, Y or D at the position corresponding to position 75 of SEQ ID NO:5; j) G at the position corresponding to position 76 of SEQ ID NO:5; k) S at the position corresponding to position 81 of SEQ ID NO:5; l) M at the position corresponding to position 105 of SEQ ID NO:5; m) H at the position corresponding to position 108 of SEQ ID NO:5; n) L at the position corresponding to position 109 of SEQ ID NO:5; o) The I or L position corresponding to position 117 of SEQ ID NO:5; p) F at the position corresponding to position 120 of SEQ ID NO:5; q) The E, T, A, G or V at the position corresponding to position 121 of SEQ ID NO:5; r) the I or H position corresponding to position 122 of SEQ ID NO:5; s) K at the position corresponding to position 125 of SEQ ID NO:5; t) A at position 126 corresponding to SEQ ID NO:5; u) H at the position corresponding to position 135 of SEQ ID NO:5; v) V at the position corresponding to position 137 of SEQ ID NO:5; w) The position Y corresponding to position 138 of SEQ ID NO:5; x) L at the position corresponding to position 139 of SEQ ID NO:5; y) K or A at the position corresponding to position 142 of SEQ ID NO:5; z) A, Q, L or M at the position corresponding to position 145 of SEQ ID NO:5; aa) K at the position corresponding to position 148 of SEQ ID NO:5; bb) Q at the position corresponding to position 151 of SEQ ID NO:5; cc) The position E or R corresponding to position 153 of SEQ ID NO:5; dd) W at the position corresponding to position 155 of SEQ ID NO:5; The R or V at the position corresponding to position 156 of SEQ ID NO:5; ff) F at the position corresponding to position 157 of SEQ ID NO:5; The R at position 158 of SEQ ID NO:5 (gg); hh) is the position Q corresponding to position 159 of SEQ ID NO:5; ii) D or W at the position corresponding to position 160 of SEQ ID NO:5; jj) is the W, H, V or A position corresponding to position 162 of SEQ ID NO:5; kk) at position R corresponding to position 165 of SEQ ID NO:5; and ll) H at the position corresponding to position 166 of SEQ ID NO:

5.

173. The nucleic acid molecule according to claim 172, wherein the deaminase comprises: a) N at position 35 of SEQ ID NO:5 and W at position 162 of SEQ ID NO:5; b) N at position 35 of SEQ ID NO:5 and Q at position 46 of SEQ ID NO:5; c) S at position 81 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; d) Q at position 46 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; e) R at the position corresponding to position 156 of SEQ ID NO:5 and W at the position corresponding to position 162 of SEQ ID NO:5; f) Position A corresponding to position 68 of SEQ ID NO: 5 and position S corresponding to position 81 of SEQ ID NO: 5; g) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at position 35 of SEQ ID NO:5, Q at position 46 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; i) W at position 155 corresponding to SEQ ID NO:5, R at position 156 corresponding to SEQ ID NO:5, and W at position 162 corresponding to SEQ ID NO:5; j) N at position 35 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; k) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; l) M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; m) Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; n) M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; o) N at position 35, Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; p) S at position 81 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; q) W at position 75 corresponding to SEQ ID NO:5 and S at position 81 corresponding to SEQ ID NO:5; r) S at position 81 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; s) S at position 81 corresponding to SEQ ID NO:5 and A at position 145 corresponding to SEQ ID NO:5; t) W at position 75 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; u) W at position 155 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; v) A at position 145 of SEQ ID NO: 5 and W at position 155 of SEQ ID NO: 5; w) W at position 75 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; x) W at position 75 of SEQ ID NO:5 and A at position 145 of SEQ ID NO:5; y) Position A corresponding to position 145 of SEQ ID NO:5 and position D corresponding to position 160 of SEQ ID NO:5; z) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; aa) S at position 81 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; bb) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; cc) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; dd) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and A at position 145 corresponding to SEQ ID NO:5; (ee) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; ff) W at position 75 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; gg) W at position 75 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; hh) A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; ii) W at position 75 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; jj) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; kk) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; ll) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; mm) The positions corresponding to positions 75 and 145 of SEQ ID NO: 5 are: W, A, W, and D, respectively. oo) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; pp) L at position 40 of SEQ ID NO:5 and M at position 105 of SEQ ID NO:5; qq) L at position 40 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; rr) L at position 40 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; ss) L at the position corresponding to position 40 of SEQ ID NO:5 and Q at the position corresponding to position 159 of SEQ ID NO:5; L at position 40 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; uu) M at position 105 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; vv) M at the position corresponding to position 105 of SEQ ID NO:5 and M at the position corresponding to position 145 of SEQ ID NO:5; ww) M at position 105 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; xx) M at position 105 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; yy) V at position 121 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; zz) V at the position corresponding to position 121 of SEQ ID NO: 5 and Q at the position corresponding to position 159 of SEQ ID NO: 5; aaa) V at position 121 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; bbb) M at position 145 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; ccc) M at position 145 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; ddd) Q at position 159 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; The position S corresponding to position 81 of SEQ ID NO:5, the position M corresponding to position 145 of SEQ ID NO:5, and the position W corresponding to position 160 of SEQ ID NO:5; fff) S at position 81 corresponding to SEQ ID NO:5, V at position 121 corresponding to SEQ ID NO:5, and Q at position 159 corresponding to SEQ ID NO:5; ggg) A at position 68 of SEQ ID NO:5 and S at position 81 of SEQ ID NO:5; hhh) G at the position corresponding to position 76 of SEQ ID NO:5 and S at the position corresponding to position 81 of SEQ ID NO:5; iii) S at position 81 corresponding to SEQ ID NO:5 and L at position 117 corresponding to SEQ ID NO:5; jjj) A at position 68 of SEQ ID NO:5 and G at position 76 of SEQ ID NO:5; kkk) A at position 68 of SEQ ID NO:5 and L at position 117 of SEQ ID NO:5; lll) G at the position corresponding to position 76 of SEQ ID NO:5 and L at the position corresponding to position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and S at position 81 of SEQ ID NO:5; The positions corresponding to positions 68 and 81 of SEQ ID NO: 5 are A, S and L respectively. ooo) G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; The positions corresponding to positions 81 and 109 of SEQ ID NO: 5 are: S, L, E, W, and D. The position corresponding to position 153 of SEQ ID NO:

5. The positions corresponding to positions 75, 81, 117, 121, 145, 155, 160, and 160 of SEQ ID NO: 5 are: A, S, L, L, A, W, and D. Q at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, I at position 117 of SEQ ID NO:5, Q at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; uuu) Y at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, L at position 117 of SEQ ID NO:5, G at position 121 of SEQ ID NO:5, L at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; vvv) C at position 22 of SEQ ID NO: 5, A at position 68 of SEQ ID NO: 5, Y at position 75 of SEQ ID NO: 5, G at position 76 of SEQ ID NO: 5, S at position 81 of SEQ ID NO: 5, H at position 108 of SEQ ID NO: 5, L at position 117 of SEQ ID NO: 5, H at position 122 of SEQ ID NO: 5, Y at position 138 of SEQ ID NO: 5, L at position 139 of SEQ ID NO: 5, A at position 142 of SEQ ID NO: 5, A at position 145 of SEQ ID NO: 5, R at position 153 of SEQ ID NO: 5, W at position 155 of SEQ ID NO: 5, R at position 156 of SEQ ID NO: 5, D at position 160 of SEQ ID NO: 5, and the position corresponding to SEQ ID NO: 5 The position H corresponding to position 162 of NO:5; or The position S at the location corresponding to position 81 of SEQ ID NO:5 (www).

174. The nucleic acid molecule according to claim 172 or 173, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO:300-414, 596-598 and 720-723.

175. The nucleic acid molecule according to claim 172 or 173, wherein the deaminase comprises the amino acid sequence of any one of SEQ ID NO: 300-414, 596-598 and 720-723.

176. The nucleic acid molecule according to any one of claims 172 to 175, wherein the deaminase has improved deaminase activity when compared with SEQ ID NO:

5.

177. The nucleic acid molecule according to any one of claims 172 to 176, wherein the deaminase is adenine deaminase.

178. The nucleic acid molecule according to any one of claims 172 to 177, wherein the DNA-binding polypeptide is an RNA-guided nuclease (RGN) polypeptide.

179. The nucleic acid molecule according to claim 178, wherein the RGN is an RGN nickase.

180. The nucleic acid molecule of claim 178, wherein the RGN has an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

181. The nucleic acid molecule of claim 179, wherein the RGN nickase has an amino acid sequence as shown in any one of SEQ ID NO:49-52 and 698.

182. The nucleic acid molecule according to claim 172 or 173, wherein the fusion protein is selected from the group consisting of: a) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; b) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; c) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; d) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; e) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; f) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:303, and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; g) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:366, and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; and h) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:720 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:

49.

183. The nucleic acid molecule according to claim 172 or 173, wherein the fusion protein is selected from the group consisting of: a) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:319 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; b) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:375 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; c) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:406 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; d) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:596 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; e) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:597 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; f) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:303 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; g) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:366 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; and h) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence as shown in SEQ ID NO:720 and is fused to the N-terminus of an RGN comprising an amino acid sequence as shown in SEQ ID NO:

49.

184. The nucleic acid molecule according to claim 172 or 173, wherein the fusion protein has a sequence as shown in any one of SEQ ID NO: 577, 578, 582, 584, 588, 593, 594, 604, 605, 616 and 729.

185. A vector comprising a nucleic acid molecule according to any one of claims 172 to 184.

186. A vector comprising a nucleic acid molecule according to any one of claims 172 to 184, said vector further comprising at least one nucleotide sequence encoding a guide RNA capable of hybridizing with a non-target strand of a target sequence.

187. A cell comprising the fusion protein according to any one of claims 158 to 170 or the RNP complex according to claim 171.

188. A cell comprising the fusion protein according to any one of claims 158 to 170, wherein the cell further comprises guide RNA.

189. A cell comprising a nucleic acid molecule according to any one of claims 172 to 184.

190. A cell comprising the carrier according to claims 185 and 186.

191. The cell according to any one of claims 187 to 190, wherein the cell is a mammalian cell.

192. The cell according to any one of claims 187 to 190, wherein the cell is a plant cell.

193. A plant or a seed comprising the cells according to claim 192.

194. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and a fusion protein according to any one of claims 158 to 170, an RNP complex according to claim 171, a nucleic acid molecule according to any one of claims 172 to 184, a carrier according to claim 185 or 186, or a cell according to claim 191.

195. A method for manufacturing an RGN fusion ribonucleoprotein complex, the method comprising introducing a nucleic acid molecule according to any one of claims 172 to 184 or a vector according to claim 185 and a nucleic acid molecule comprising an expression cassette encoding a guide RNA or a vector according to claim 186 into a cell, and culturing the cell under conditions where the fusion protein and gRNA are expressed and an RGN fusion ribonucleoprotein complex is formed.

196. A system for modifying a target DNA molecule comprising a target DNA sequence, the system comprising: a) A fusion protein comprising an RNA-guided nuclease polypeptide and a deaminase, wherein the deaminase has an amino acid sequence having at least 85% sequence identity with SEQ ID NO:5, or a nucleotide sequence encoding the fusion protein, wherein the deaminase comprises at least one of the following amino acid residues: i) Position C at the position corresponding to position 2 of SEQ ID NO:5; ii) F or C at the position corresponding to position 22 of SEQ ID NO:5; iii) Q at the position corresponding to position 23 of SEQ ID NO:5; iv) N at position 35 corresponding to SEQ ID NO:5; v) L at the position corresponding to position 40 of SEQ ID NO:5; vi) Q at the position corresponding to position 46 of SEQ ID NO:5; vii) A or M at the position corresponding to position 68 of SEQ ID NO:5; viii) H at the position corresponding to position 72 of SEQ ID NO:5; ix) the position W, A, Q, Y or D corresponding to position 75 of SEQ ID NO:5; x) G at the position corresponding to position 76 of SEQ ID NO:5; xi) is the position S corresponding to position 81 of SEQ ID NO:5; xii) M at the position corresponding to position 105 of SEQ ID NO:5; xiii) H at the position corresponding to position 108 of SEQ ID NO:5; xiv) L at the position corresponding to position 109 of SEQ ID NO:5; xv) is the I or L position corresponding to position 117 of SEQ ID NO:5; xvi) F at the position corresponding to position 120 of SEQ ID NO:5; xvii) the position E, T, A, G or V corresponding to position 121 of SEQ ID NO:5; xviii) The I or H position corresponding to position 122 of SEQ ID NO:5; xix) is the position K corresponding to position 125 of SEQ ID NO:5; (xx) A at position 126 corresponding to SEQ ID NO:5; xxi) H at the position corresponding to position 135 of SEQ ID NO:5; xxii) V at the position corresponding to position 137 of SEQ ID NO:5; xxiii) The Y at the position corresponding to position 138 of SEQ ID NO:5; xxiv) L at the position corresponding to position 139 of SEQ ID NO:5; The position K or A at position 142 of SEQ ID NO: 5 (xxv) corresponds to the position of SEQ ID NO: 5; (xxvi) The A, Q, L or M position corresponding to position 145 of SEQ ID NO:5; The position K at position 148 of SEQ ID NO: 5 (xxvii) corresponds to position 148. xxviii) Q at the position corresponding to position 151 of SEQ ID NO:5; The position E or R at position 153 of SEQ ID NO: xxix) corresponds to position 153; The W at position 155 of SEQ ID NO: 5 (xxx); The R or V at the position corresponding to position 156 of SEQ ID NO:5; xxxii) F at the position corresponding to position 157 of SEQ ID NO:5; xxxiii) R at the position corresponding to position 158 of SEQ ID NO:5; xxxiv) Q at the position corresponding to position 159 of SEQ ID NO:5; The position D or W corresponding to position 160 of SEQ ID NO: 5; xxxvi) The W, H, V or A position corresponding to position 162 of SEQ ID NO:5; The position R at position 165 corresponding to position xxxvii) of SEQ ID NO:5; as well as xxxviii) H at position 166 corresponding to SEQ ID NO:5; as well as b) one or more guide RNAs capable of hybridizing with the non-target strand of the target DNA sequence, or one or more nucleotide sequences encoding the one or more guide RNAs; and The one or more guide RNAs thereon are capable of forming a complex with the fusion protein to guide the fusion protein to bind to the target DNA sequence and modify the target DNA molecule.

197. The system of claim 196, wherein the deaminase comprises: a) N at position 35 of SEQ ID NO:5 and W at position 162 of SEQ ID NO:5; b) N at position 35 of SEQ ID NO:5 and Q at position 46 of SEQ ID NO:5; c) S at position 81 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; d) Q at position 46 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; e) R at the position corresponding to position 156 of SEQ ID NO:5 and W at the position corresponding to position 162 of SEQ ID NO:5; f) Position A corresponding to position 68 of SEQ ID NO: 5 and position S corresponding to position 81 of SEQ ID NO: 5; g) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at position 35 of SEQ ID NO:5, Q at position 46 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; i) W at position 155 corresponding to SEQ ID NO:5, R at position 156 corresponding to SEQ ID NO:5, and W at position 162 corresponding to SEQ ID NO:5; j) N at position 35 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; k) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; l) M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; m) Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; n) M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; o) N at position 35, Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; p) S at position 81 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; q) W at position 75 corresponding to SEQ ID NO:5 and S at position 81 corresponding to SEQ ID NO:5; r) S at position 81 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; s) S at position 81 corresponding to SEQ ID NO:5 and A at position 145 corresponding to SEQ ID NO:5; t) W at position 75 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; u) W at position 155 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; v) A at position 145 of SEQ ID NO: 5 and W at position 155 of SEQ ID NO: 5; w) W at position 75 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; x) W at position 75 of SEQ ID NO:5 and A at position 145 of SEQ ID NO:5; y) Position A corresponding to position 145 of SEQ ID NO:5 and position D corresponding to position 160 of SEQ ID NO:5; z) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; aa) S at position 81 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; bb) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; cc) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; dd) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and A at position 145 corresponding to SEQ ID NO:5; (ee) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; ff) W at position 75 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; gg) W at position 75 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; hh) A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; ii) W at position 75 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; jj) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; kk) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; ll) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; mm) The positions corresponding to positions 75 and 145 of SEQ ID NO: 5 are: W, A, W, and D, respectively. oo) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; pp) L at position 40 of SEQ ID NO:5 and M at position 105 of SEQ ID NO:5; qq) L at position 40 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; rr) L at position 40 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; ss) L at the position corresponding to position 40 of SEQ ID NO:5 and Q at the position corresponding to position 159 of SEQ ID NO:5; L at position 40 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; uu) M at position 105 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; vv) M at the position corresponding to position 105 of SEQ ID NO:5 and M at the position corresponding to position 145 of SEQ ID NO:5; ww) M at position 105 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; xx) M at position 105 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; yy) V at position 121 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; zz) V at the position corresponding to position 121 of SEQ ID NO: 5 and Q at the position corresponding to position 159 of SEQ ID NO: 5; aaa) V at position 121 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; bbb) M at position 145 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; ccc) M at position 145 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; ddd) Q at position 159 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; The position S corresponding to position 81 of SEQ ID NO:5, the position M corresponding to position 145 of SEQ ID NO:5, and the position W corresponding to position 160 of SEQ ID NO:5; fff) S at position 81 corresponding to SEQ ID NO:5, V at position 121 corresponding to SEQ ID NO:5, and Q at position 159 corresponding to SEQ ID NO:5; ggg) A at position 68 of SEQ ID NO:5 and S at position 81 of SEQ ID NO:5; hhh) G at the position corresponding to position 76 of SEQ ID NO:5 and S at the position corresponding to position 81 of SEQ ID NO:5; iii) S at position 81 corresponding to SEQ ID NO:5 and L at position 117 corresponding to SEQ ID NO:5; jjj) A at position 68 of SEQ ID NO:5 and G at position 76 of SEQ ID NO:5; kkk) A at position 68 of SEQ ID NO:5 and L at position 117 of SEQ ID NO:5; lll) G at the position corresponding to position 76 of SEQ ID NO:5 and L at the position corresponding to position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and S at position 81 of SEQ ID NO:5; The positions corresponding to positions 68 and 81 of SEQ ID NO: 5 are A, S and L respectively. ooo) G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; The positions corresponding to positions 81 and 109 of SEQ ID NO: 5 are: S, L, E, W, and D. The position corresponding to position 153 of SEQ ID NO:

5. The positions corresponding to positions 75, 81, 117, 121, 145, 155, 160, and 160 of SEQ ID NO: 5 are: A, S, L, L, A, W, and D. Q at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, I at position 117 of SEQ ID NO:5, Q at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; uuu) Y at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, L at position 117 of SEQ ID NO:5, G at position 121 of SEQ ID NO:5, L at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; vvv) C at position 22 of SEQ ID NO: 5, A at position 68 of SEQ ID NO: 5, Y at position 75 of SEQ ID NO: 5, G at position 76 of SEQ ID NO: 5, S at position 81 of SEQ ID NO: 5, H at position 108 of SEQ ID NO: 5, L at position 117 of SEQ ID NO: 5, H at position 122 of SEQ ID NO: 5, Y at position 138 of SEQ ID NO: 5, L at position 139 of SEQ ID NO: 5, A at position 142 of SEQ ID NO: 5, A at position 145 of SEQ ID NO: 5, R at position 153 of SEQ ID NO: 5, W at position 155 of SEQ ID NO: 5, R at position 156 of SEQ ID NO: 5, D at position 160 of SEQ ID NO: 5, and the position corresponding to SEQ ID NO: 5 The position H corresponding to position 162 of NO:5; or The position S at the location corresponding to position 81 of SEQ ID NO:5 (www).

198. The system according to claim 196 or 197, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO:300-414, 596-598 and 720-723.

199. The system according to claim 196 or 197, wherein the deaminase comprises the amino acid sequence of any one of SEQ ID NO: 300-414, 596-598 and 720-723.

200. The system according to any one of claims 196 to 199, wherein the deaminase has improved deaminase activity when compared with SEQ ID NO:

5.

201. The system according to any one of claims 196 to 200, wherein the deaminase is adenine deaminase.

202. The system according to any one of claims 196 to 201, wherein at least one of the nucleotide sequence encoding the one or more guide RNAs and the nucleotide sequence encoding the fusion protein is operatively linked to a promoter heterologous to the nucleotide sequence.

203. The system according to any one of claims 196 to 202, wherein the target DNA sequence is a eukaryotic target DNA sequence.

204. The system according to any one of claims 196 to 203, wherein the target DNA sequence is located adjacent to the prespacer sequence neighbor motif (PAM) recognized by the RGN.

205. The system according to any one of claims 196 to 204, wherein the RGN of the fusion protein has an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

206. The system according to any one of claims 196 to 205, wherein the RGN of the fusion protein is an RGN cleavage enzyme.

207. The system of claim 206, wherein the RGN nickase has an amino acid sequence as shown in any one of SEQ ID NO:49-52 and 698.

208. The system of claim 196 or 197, wherein the fusion protein is selected from the group consisting of: a) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; b) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; c) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; d) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; e) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; f) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:303, and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; g) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:366, and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; and h) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:720 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:

49.

209. The system of claim 196 or 197, wherein the fusion protein is selected from the group consisting of: a) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:319 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; b) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:375 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; c) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:406 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; d) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:596 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; e) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:597 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; f) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:303 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; g) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:366 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; and h) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence as shown in SEQ ID NO:720 and is fused to the N-terminus of an RGN comprising an amino acid sequence as shown in SEQ ID NO:

49.

210. The system according to claim 196 or 197, wherein the fusion protein has a sequence as shown in any one of SEQ ID NO: 577, 578, 582, 584, 588, 593, 594, 604, 605, 616 and 729.

211. A cell comprising the system according to any one of claims 196 to 210.

212. The cell of claim 211, wherein the cell is a mammalian cell.

213. The cell of claim 211, wherein the cell is a plant cell.

214. A plant or a seed comprising the cells according to claim 213.

215. A pharmaceutical composition comprising a pharmaceutically acceptable carrier and the system according to any one of claims 196 to 210 or the cell according to claim 212.

216. A method for modifying a target DNA molecule containing a target DNA sequence, the method comprising delivering a system according to any one of claims 196 to 210 to the target DNA molecule or a cell containing the target DNA molecule.

217. A method for modifying a target DNA molecule containing a target sequence, the method comprising: a) The RGN deaminase ribonucleotide complex is assembled in vitro by combining the following items under conditions suitable for the formation of the RGN deaminase ribonucleotide complex: i) one or more guide RNAs capable of hybridizing with the non-target strand of the target DNA sequence; and ii) A fusion protein comprising an RNA-guided nuclease polypeptide (RGN) and at least one deaminase, wherein The deaminase has an amino acid sequence that is at least 85% sequence identical to SEQ ID NO:5; The deaminase said therein has at least one of the following amino acid residues: A) Position C corresponding to position 2 of SEQ ID NO:5; B) F or C at the position corresponding to position 22 of SEQ ID NO:5; C) Q at the position corresponding to position 23 of SEQ ID NO:5; D) N at position 35 corresponding to SEQ ID NO:5; E) L at the position corresponding to position 40 of SEQ ID NO:5; F) Q at the position corresponding to position 46 of SEQ ID NO:5; G) A or M at the position corresponding to position 68 of SEQ ID NO:5; H) H at the position corresponding to position 72 of SEQ ID NO:5; II) The position of W, A, Q, Y or D corresponding to position 75 of SEQ ID NO:5; J) G at the position corresponding to position 76 of SEQ ID NO:5; K) S at the position corresponding to position 81 of SEQ ID NO:5; LI) is the position M corresponding to position 105 of SEQ ID NO:5; MI) H at position 108 corresponding to position 108 of SEQ ID NO:5; N) L at the position corresponding to position 109 of SEQ ID NO:5; O) The I or L position corresponding to position 117 of SEQ ID NO:5; P) F at the position corresponding to position 120 of SEQ ID NO:5; Q) The E, T, A, G or V at the position corresponding to position 121 of SEQ ID NO:5; R) The I or H position corresponding to position 122 of SEQ ID NO:5; S) K at the position corresponding to position 125 of SEQ ID NO:5; T) A at position 126 corresponding to position 126 of SEQ ID NO: 5; U) H at the position corresponding to position 135 of SEQ ID NO:5; V) is located at the position corresponding to position 137 of SEQ ID NO:5; W) The position Y corresponding to position 138 of SEQ ID NO:5; X) L at the position corresponding to position 139 of SEQ ID NO:5; Y) K or A at the position corresponding to position 142 of SEQ ID NO:5; Z) A, Q, L or M at the position corresponding to position 145 of SEQ ID NO:5; AA) K at the position corresponding to position 148 of SEQ ID NO:5; Q at position 151 corresponding to position 151 of SEQ ID NO: 5; The E or R at the position corresponding to position 153 of SEQ ID NO: 5; The position W at the location corresponding to position 155 of SEQ ID NO:5; The R or V at the position corresponding to position 156 of SEQ ID NO: 5; F at the position corresponding to position 157 of SEQ ID NO: 5; R at the position corresponding to position 158 of SEQ ID NO: 5; Q at the position corresponding to position 159 of SEQ ID NO: 5 (HH); II) The position D or W corresponding to position 160 of SEQ ID NO:5; KK) corresponds to the position 162 of SEQ ID NO: 5 with W, H, V or A; The position R at the location corresponding to position 165 of SEQ ID NO: 5; as well as LL) H at position 166 corresponding to SEQ ID NO:5; as well as b) Contact the target DNA molecule or a cell containing the target DNA molecule with the in vitro assembled RGN deaminase ribonucleotide complex; The one or more guide RNAs hybridize with the non-target strand of the target DNA sequence, thereby guiding the fusion protein to bind to the target DNA sequence and causing modification of the target DNA molecule.

218. The method of claim 217, wherein the deaminase comprises: a) N at position 35 of SEQ ID NO:5 and W at position 162 of SEQ ID NO:5; b) N at position 35 of SEQ ID NO:5 and Q at position 46 of SEQ ID NO:5; c) S at position 81 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; d) Q at position 46 of SEQ ID NO:5 and R at position 156 of SEQ ID NO:5; e) R at the position corresponding to position 156 of SEQ ID NO:5 and W at the position corresponding to position 162 of SEQ ID NO:5; f) Position A corresponding to position 68 of SEQ ID NO: 5 and position S corresponding to position 81 of SEQ ID NO: 5; g) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; h) N at position 35 of SEQ ID NO:5, Q at position 46 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; i) W at position 155 corresponding to SEQ ID NO:5, R at position 156 corresponding to SEQ ID NO:5, and W at position 162 corresponding to SEQ ID NO:5; j) N at position 35 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; k) N at position 35 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; l) M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; m) Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; n) M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; o) N at position 35, Q at position 46 of SEQ ID NO:5, M at position 68 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, H at position 108 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, R at position 156 of SEQ ID NO:5, and W at position 162 of SEQ ID NO:5; p) S at position 81 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; q) W at position 75 corresponding to SEQ ID NO:5 and S at position 81 corresponding to SEQ ID NO:5; r) S at position 81 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; s) S at position 81 corresponding to SEQ ID NO:5 and A at position 145 corresponding to SEQ ID NO:5; t) W at position 75 of SEQ ID NO:5 and W at position 155 of SEQ ID NO:5; u) W at position 155 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; v) A at position 145 of SEQ ID NO: 5 and W at position 155 of SEQ ID NO: 5; w) W at position 75 of SEQ ID NO:5 and D at position 160 of SEQ ID NO:5; x) W at position 75 of SEQ ID NO:5 and A at position 145 of SEQ ID NO:5; y) Position A corresponding to position 145 of SEQ ID NO:5 and position D corresponding to position 160 of SEQ ID NO:5; z) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; aa) S at position 81 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; bb) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; cc) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; dd) W at position 75 corresponding to SEQ ID NO:5, S at position 81 corresponding to SEQ ID NO:5, and A at position 145 corresponding to SEQ ID NO:5; (ee) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; ff) W at position 75 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; gg) W at position 75 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, and W at position 155 corresponding to SEQ ID NO:5; hh) A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; ii) W at position 75 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; jj) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; kk) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and W at position 155 of SEQ ID NO:5; ll) S at position 81 corresponding to SEQ ID NO:5, A at position 145 corresponding to SEQ ID NO:5, W at position 155 corresponding to SEQ ID NO:5, and D at position 160 corresponding to SEQ ID NO:5; W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; mm) The positions corresponding to positions 75 and 145 of SEQ ID NO: 5 are: W, A, W, and D, respectively. oo) W at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, A at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; pp) L at position 40 of SEQ ID NO:5 and M at position 105 of SEQ ID NO:5; qq) L at position 40 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; rr) L at position 40 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; ss) L at the position corresponding to position 40 of SEQ ID NO:5 and Q at the position corresponding to position 159 of SEQ ID NO:5; L at position 40 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; uu) M at position 105 of SEQ ID NO:5 and V at position 121 of SEQ ID NO:5; vv) M at the position corresponding to position 105 of SEQ ID NO:5 and M at the position corresponding to position 145 of SEQ ID NO:5; ww) M at position 105 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; xx) M at position 105 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; yy) V at position 121 of SEQ ID NO:5 and M at position 145 of SEQ ID NO:5; zz) V at the position corresponding to position 121 of SEQ ID NO: 5 and Q at the position corresponding to position 159 of SEQ ID NO: 5; aaa) V at position 121 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; bbb) M at position 145 of SEQ ID NO:5 and Q at position 159 of SEQ ID NO:5; ccc) M at position 145 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; ddd) Q at position 159 of SEQ ID NO:5 and W at position 160 of SEQ ID NO:5; The position S corresponding to position 81 of SEQ ID NO:5, the position M corresponding to position 145 of SEQ ID NO:5, and the position W corresponding to position 160 of SEQ ID NO:5; fff) S at position 81 corresponding to SEQ ID NO:5, V at position 121 corresponding to SEQ ID NO:5, and Q at position 159 corresponding to SEQ ID NO:5; ggg) A at position 68 of SEQ ID NO:5 and S at position 81 of SEQ ID NO:5; hhh) G at the position corresponding to position 76 of SEQ ID NO:5 and S at the position corresponding to position 81 of SEQ ID NO:5; iii) S at position 81 corresponding to SEQ ID NO:5 and L at position 117 corresponding to SEQ ID NO:5; jjj) A at position 68 of SEQ ID NO:5 and G at position 76 of SEQ ID NO:5; kkk) A at position 68 of SEQ ID NO:5 and L at position 117 of SEQ ID NO:5; lll) G at the position corresponding to position 76 of SEQ ID NO:5 and L at the position corresponding to position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and S at position 81 of SEQ ID NO:5; The positions corresponding to positions 68 and 81 of SEQ ID NO: 5 are A, S and L respectively. ooo) G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; A at position 68 of SEQ ID NO:5, G at position 76 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, and L at position 117 of SEQ ID NO:5; The positions corresponding to positions 81 and 109 of SEQ ID NO: 5 are: S, L, E, W, and D. The position corresponding to position 153 of SEQ ID NO:

5. The positions corresponding to positions 75, 81, 117, 121, 145, 155, 160, and 160 of SEQ ID NO: 5 are: A, S, L, L, A, W, and D. Q at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, I at position 117 of SEQ ID NO:5, Q at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; uuu) Y at position 75 of SEQ ID NO:5, S at position 81 of SEQ ID NO:5, L at position 117 of SEQ ID NO:5, G at position 121 of SEQ ID NO:5, L at position 145 of SEQ ID NO:5, W at position 155 of SEQ ID NO:5, and D at position 160 of SEQ ID NO:5; vvv) C at position 22 of SEQ ID NO: 5, A at position 68 of SEQ ID NO: 5, Y at position 75 of SEQ ID NO: 5, G at position 76 of SEQ ID NO: 5, S at position 81 of SEQ ID NO: 5, H at position 108 of SEQ ID NO: 5, L at position 117 of SEQ ID NO: 5, H at position 122 of SEQ ID NO: 5, Y at position 138 of SEQ ID NO: 5, L at position 139 of SEQ ID NO: 5, A at position 142 of SEQ ID NO: 5, A at position 145 of SEQ ID NO: 5, R at position 153 of SEQ ID NO: 5, W at position 155 of SEQ ID NO: 5, R at position 156 of SEQ ID NO: 5, D at position 160 of SEQ ID NO: 5, and the position corresponding to SEQ ID NO: 5 The position H corresponding to position 162 of NO:5; or The position S at the location corresponding to position 81 of SEQ ID NO:5 (www).

219. The method according to claim 217 or 218, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO:300-414, 596-598 and 720-723.

220. The method according to claim 217 or 218, wherein the deaminase has the amino acid sequence of any one of SEQ ID NO: 300-414, 596-598 and 720-723.

221. The method according to any one of claims 217 to 220, wherein the deaminase has improved deaminase activity when compared with SEQ ID NO:

5.

222. The method according to any one of claims 217 to 221, wherein the deaminase is adenine deaminase.

223. The method according to any one of claims 217 to 222, wherein the RGN of the fusion protein has an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 1-4, 49-162, 435, 575, 576, 698 and 699.

224. The method according to any one of claims 217 to 222, wherein the RGN of the fusion protein is an RGN cleavage enzyme.

225. The method of claim 224, wherein the RGN nickase has an amino acid sequence as shown in any one of SEQ ID NO:49-52 and 698.

226. The method of claim 217 or 218, wherein the fusion protein is selected from the group consisting of: a) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:319, and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; b) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:375 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; c) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:406, and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; d) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:596 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; e) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:597 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; f) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:303, and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; g) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:366, and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:49; and h) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:720 and is fused to the N-terminus of an RGN comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:

49.

227. The method of claim 217 or 218, wherein the fusion protein is selected from the group consisting of: a) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:319 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; b) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:375 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; c) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:406 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; d) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:596 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; e) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:597 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; f) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:303 and is fused to the N-terminus of an RGN comprising the amino acid shown in SEQ ID NO:49; g) A fusion protein comprising the deaminase, wherein the deaminase comprises the amino acid sequence shown in SEQ ID NO:366 and is fused to the N-terminus of an RGN comprising the amino acid sequence shown in SEQ ID NO:49; and h) A fusion protein comprising the deaminase, wherein the deaminase comprises an amino acid sequence as shown in SEQ ID NO:720 and is fused to the N-terminus of an RGN comprising an amino acid sequence as shown in SEQ ID NO:

49.

228. The method according to claim 217 or 218, wherein the fusion protein has a sequence as shown in any one of SEQ ID NO: 577, 578, 582, 584, 588, 593, 594, 604, 605, 616 and 729.

229. The method according to any one of claims 217 to 228, wherein the target DNA sequence is a eukaryotic target DNA sequence.

230. The method according to any one of claims 217 to 229, wherein the target DNA sequence is located at a neighboring motif (PAM) adjacent to the prespacer sequence.

231. The method according to any one of claims 217 to 230, wherein the target DNA molecule contains a causal mutation of a disease or condition, and wherein the modification of the target DNA molecule corrects the causal mutation.

232. The method of claim 231, wherein the correction of the causal mutation includes the correction of nonsense mutations.

233. The method according to any one of claims 217 to 232, wherein the target DNA molecule is intracellular.

234. The method of claim 233, further comprising selecting cells containing the modified DNA molecule.

235. A cell comprising the modified target DNA sequence of the method according to claim 234.

236. The cell of claim 235, wherein the cell is a mammalian cell.

237. The cell of claim 235, wherein the cell is a plant cell.

238. A plant or a seed comprising the cells according to claim 237.

239. A pharmaceutical composition comprising a cellular and pharmaceutically acceptable carrier as described in claim 236.

240. A method for treating a subject suffering from a disease, condition, or illness, or at risk of developing said disease, condition, or illness, the method comprising: The subject is administered the fusion protein according to any one of claims 158 to 170, the RNP complex according to claim 171, the nucleic acid molecule according to any one of claims 172 to 184, the carrier according to claim 185 or 186, the cell according to any one of claims 191, 212 and 236, the system according to any one of claims 196 to 210, or the pharmaceutical composition according to any one of claims 194, 215 and 239.

241. The method of claim 240, wherein the disease is associated with a causal mutation, and the treatment comprises correcting the causal mutation.

242. Use of the fusion protein according to any one of claims 158 to 170, the RNP complex according to claim 171, the nucleic acid molecule according to any one of claims 172 to 184, the vector according to claim 185 or 186, the cell according to any one of claims 191, 212 and 236, or the system according to any one of claims 196 to 210 for the treatment of a subject suffering from a disease, condition or illness, or at risk of developing a disease, condition or illness.

243. The use according to claim 242, wherein the disease is associated with a causal mutation, and the treatment comprises correcting the causal mutation.

244. Use of the fusion protein according to any one of claims 158 to 170, the RNP complex according to claim 171, the nucleic acid molecule according to any one of claims 172 to 184, the vector according to claim 185 or 186, the cell according to any one of claims 191, 212 and 236, or the system according to any one of claims 196 to 210 for the manufacture of a medicament for the treatment of a disease, symptom or condition.

245. The use according to claim 244, wherein the disease is associated with a causal mutation, and an effective amount of the drug corrects the causal mutation.

Citation Information

Patent Citations

  • Methods and compositions for RNA-directed target DNA modification and for RNA-directed modulation of transcription

    US10000772B2

  • Methods and compositions for RNA-directed target DNA modification and for RNA-directed modulation of transcription

    US10113167B2

  • Methods and compositions for directed genome editing

    US11193123B2

  • Methods and compositions for prime editing nucleotide sequences

    US11447770B1

  • Regulation of endogenous gene expression in cells using zinc finger proteins

    US20030087817A1

Cited By

  • Application of a compact CRISPR / SwCas9 system in maize gene editing

    CN122405593A